Half-Quadratic Quantization for LLMs without calibration data. Use when quantizing models to 4/3/2-bit precision withou…
Half-Quadratic Quantization for LLMs without calibration data. Use when quantizing models to 4/3/2-bit precision without needing calibration datasets, for fast quantization workflows, or when deploying with vLLM or HuggingFace Transformers.
davila7
cli
free
Others in the same category, ranked by how often they are opened.