QUASARLowering the Loss Floor of Quantization-Aware Training with Loss-Aware Reconstruction

Vincent Counathe1,2 Ben Athiwaratkun2 Christopher De Sa1,2 Tianyi Zhang2

1Cornell University2Together AI

illustrative The structural mismatch of gradients in standard QAT versus QUASAR. Under the STE, both methods evaluate the gradient at the reconstructed weights r and apply it to the latent weights w. QUASAR minimizes the loss-aware reconstruction error, making the gradient at r a more faithful proxy for the gradient at w.

We introduce QUASAR, a QAT method that brings lightweight, loss-aware reconstruction into the training loop.

+13.3pts

average accuracy over the best QAT baseline

INT2distillationQwen3-4B-Thinking
+10.9pts

average accuracy over the best QAT baseline

INT2math adaptationQwen3-4B-Base
66%

lower KL divergence

INT4Gemma-4 E4Bvs Google QAT
1.4%

latency overhead per training step

no inference overhead
Abstract

As large language model inference shifts toward lower precision, post-training quantization (PTQ) becomes increasingly brittle, making quantization-aware training (QAT) essential for preserving model quality. However, QAT has a structural mismatch: gradient updates are applied to latent full-precision weights, while the loss and gradients are computed on lossy reconstructions of those weights. This mismatch can lead to suboptimal training trajectories and a higher loss floor. Second-order PTQ methods address a similar problem by minimizing loss-aware reconstruction error, but applying such expensive reconstruction repeatedly during QAT as the weights evolve is impractical. We introduce QUASAR, a QAT method that brings lightweight, loss-aware reconstruction into the training loop. At each training step, QUASAR reconstructs the latent weights by searching over a small set of clipping ranges and fitting dequantization parameters through saliency-weighted least squares. We use an exponential moving average of squared gradients as the per-parameter saliency signal. Our theoretical analysis shows that optimizing QUASAR's reconstruction objective tightens both the convergence and final-loss bounds of QAT. We evaluate QUASAR across four model families and across INT4, INT3, INT2, and NVFP4 quantization formats. QUASAR consistently achieves lower training and evaluation loss than competitive QAT methods and outperforms QAT and PTQ baselines on downstream benchmarks. At INT2, QUASAR improves average accuracy over the best QAT baseline by 13.3 points with quantization-aware distillation and by 10.9 points with QAT on mathematical reasoning data. Notably, after distillation on only ~600M tokens, QUASAR's INT4 Gemma-4 E4B checkpoint outperforms the corresponding QAT checkpoint released by Google, with 66% lower KL divergence and 1.8 points higher average accuracy.

1.Reconstruction in QAT

QAT maintains and updates full-precision latent weights w, but computes the forward pass and loss using reconstructed weights r.

the training loop
The training loops for standard QAT and QUASAR differ only in the reconstruction w ↦ q ↦ r. Standard QAT determines the quantization grid from each group's extreme weights; QUASAR searches over code assignments and fits the dequantizer to minimize loss-aware reconstruction error.

This surrogate is generally not the optimal descent direction for the latent weights, leading training along a suboptimal trajectory. The consequence is a loss-floor gap: given the same model and training data, QAT converges to a higher final training and evaluation loss than full-precision training.

A second-order expansion of the loss around the full-precision weights w, assuming that the first-order term is negligible near an optimum, gives

where H(w) is the Hessian evaluated at w. We refer to S as the loss-aware reconstruction error. It measures the local change in loss caused by reconstructing the latent weights.

the borrowed gradient
illustrative Under the STE, both methods evaluate the gradient at the reconstructed weights r and apply it to the latent weights w. QUASAR minimizes the loss-aware reconstruction error, making the gradient at r a more faithful proxy for the gradient at w.

2.Optimizing Both Stages of Reconstruction

At each training step, QUASAR reconstructs the latent weights by searching over a small set of clipping ranges and fitting dequantization parameters through saliency-weighted least squares.

real Qwen3-4B weight groups
Scale search finds good code assignments. The case f=1 is the full min–max range used by standard QAT. A factor f<1 narrows the range and clips weights outside it, so the quantization levels cover a smaller interval with finer spacing. QUASAR prioritizes accurate reconstruction of high-saliency weights while tolerating more error on low-saliency weights.

We use an exponential moving average of squared gradients as the per-parameter saliency signal. Adam and AdamW already maintain this quantity as the second-moment estimate vt, allowing QUASAR to reuse it without additional memory overhead. QUASAR applies to affine and symmetric integer formats, NVFP4, and MXFP4.

3.Reconstruction Error Bounds QAT Final Loss

Our theoretical analysis shows that loss-aware reconstruction error directly controls QAT convergence and the quantized model's final loss, providing a principled foundation for QUASAR's objective.

after training · INT2 · 1,000 steps

Fig. 12 · QUASAR · INT2 · 1,000 steps. Reconstruction error drives final loss.

scale-search factor fend-of-training reconstruction error ŜTfinal held-out KL
f ≥ 0.3 (unrestricted, ★)3.90×10⁻⁴0.152
f ≥ 0.54.03×10⁻⁴0.152
f ≥ 0.74.57×10⁻⁴0.161
f ≥ 0.855.50×10⁻⁴0.189
f = 16.70×10⁻⁴0.229

Fig. 12 · log-log power-law fit

fitslopeR²
ŜT vs final held-out KL0.770.976
Reconstruction error drives final loss. Five QUASAR INT2 runs differ only in the lower bound imposed on the scale-search factor f. Restricting the search raises both the final reconstruction error ŜT and held-out KL after 1,000 training steps. The star marks the unrestricted search.
end of training
Reconstruction error tracks final KL. End-of-training loss-aware reconstruction error versus final held-out KL to the full-precision teacher, for Standard QAT, Denoising QAT, and QUASAR at INT4, INT3, and INT2.

4.Results on Healing and Adaptation

Healing with QAD.

Across both models and all bit widths, QUASAR reaches the lowest training and evaluation loss (KL). At INT2, it reduces KL by about 30% over the strongest QAT baseline. On reasoning, coding, and long-context tasks, QUASAR improves average accuracy by 13.3 points for Qwen and 2.2 points for Llama.

healing with QAD
QUASAR gives the lowest held-out KL across both models and all bit widths. For healing, we apply QAD to Qwen3-4B-Thinking-2507 and Llama-3.1-8B-Instruct at INT2, INT3, and INT4 using Open-PerfectBlend data and logits from the full-precision teachers. We show training loss over the full run, the final 1,000 steps, and held-out evaluation loss. All losses are forward KL to the full-precision model.
downstream accuracy

Table 1 · Qwen3-4B-Thinking-2507 · INT2. Reasoning, coding, and long-context results for Qwen3-4B-Thinking-2507 and Llama-3.1-8B-Instruct at INT2, INT3, and INT4.

MethodKL ↓Top-1 ↑HMMT '26AIME '25MATH-500MMLU-ProSuperGPQALCB v6LongBench-v2RULERAvg.
Teacher (BF16)––43.177.998.272.748.174.644.587.968.4
RTN11.1660.50.00.01.90.010.00.00.00.01.5
GPTQ1.14067.70.00.02.63.28.40.00.00.31.8
AWQ3.01637.00.00.02.40.05.20.00.00.01.0
Standard QAT0.18689.91.01.749.526.617.48.220.539.320.5
LSQ0.20589.20.80.547.226.217.87.420.741.720.3
Denoising QAT0.17990.11.01.451.827.419.09.218.939.421.0
BitDistiller0.20989.10.61.151.429.316.89.418.744.921.5
QUASAR (ours)0.12692.06.214.477.445.825.423.426.659.634.8

Table 1 · Qwen3-4B-Thinking-2507 · INT3

MethodKL ↓Top-1 ↑HMMT '26AIME '25MATH-500MMLU-ProSuperGPQALCB v6LongBench-v2RULERAvg.
Teacher (BF16)––43.177.998.272.748.174.644.587.968.4
RTN0.59778.80.00.713.016.315.60.08.749.313.0
GPTQ0.13091.531.249.495.863.739.549.431.073.354.1
AWQ0.24088.122.731.090.661.537.035.730.276.048.1
Standard QAT0.06294.229.947.494.264.940.053.329.680.855.0
LSQ0.07093.828.345.593.364.138.252.232.880.554.4
Denoising QAT0.06094.329.349.394.764.539.752.332.879.055.2
BitDistiller0.06594.229.752.595.165.840.657.733.083.557.3
QUASAR (ours)0.05494.931.760.595.266.741.560.235.884.459.5

Table 1 · Qwen3-4B-Thinking-2507 · INT4

MethodKL ↓Top-1 ↑HMMT '26AIME '25MATH-500MMLU-ProSuperGPQALCB v6LongBench-v2RULERAvg.
Teacher (BF16)––43.177.998.272.748.174.644.587.968.4
RTN0.10892.235.762.196.869.744.465.238.883.362.0
GPTQ0.02796.238.874.897.871.446.570.638.287.465.7
AWQ0.05094.736.670.697.671.645.770.839.086.864.8
Standard QAT0.01996.936.872.197.771.446.570.939.087.165.2
LSQ0.02496.439.368.597.570.845.669.336.487.364.3
Denoising QAT0.01896.938.871.597.270.946.471.642.986.965.8
BitDistiller0.01896.938.573.997.371.446.571.140.486.865.7
QUASAR (ours)0.01697.138.573.897.571.746.272.241.088.066.1

Table 1 · Llama-3.1-8B-Instruct · INT2

MethodKL ↓Top-1 ↑MATH-500MMLU-ProSuperGPQALCB v6LongBench-v2RULERAvg.
Teacher (BF16)––49.642.322.417.030.089.641.8
RTN10.8100.50.00.010.20.00.00.01.7
GPTQ2.38051.82.91.88.70.00.00.12.3
AWQ6.76912.91.60.010.70.00.00.02.0
Standard QAT0.16390.121.616.811.84.015.15.512.5
LSQ0.17289.619.915.711.62.315.352.919.6
Denoising QAT0.16590.120.816.511.03.68.50.010.1
BitDistiller0.15190.221.719.011.85.520.563.523.7
QUASAR (ours)0.10592.030.022.613.77.015.966.425.9

Table 1 · Llama-3.1-8B-Instruct · INT3

MethodKL ↓Top-1 ↑MATH-500MMLU-ProSuperGPQALCB v6LongBench-v2RULERAvg.
Teacher (BF16)––49.642.322.417.030.089.641.8
RTN0.27685.68.217.212.50.721.149.018.1
GPTQ0.07692.838.631.918.410.427.680.934.6
AWQ0.19788.016.215.313.92.422.769.023.3
Standard QAT0.04694.739.534.818.612.730.281.736.3
LSQ0.05294.239.335.017.113.727.286.136.4
Denoising QAT0.04694.741.234.618.213.029.280.336.1
BitDistiller0.04194.842.535.919.313.131.284.437.7
QUASAR (ours)0.03695.243.436.419.716.226.085.837.9

Table 1 · Llama-3.1-8B-Instruct · INT4

MethodKL ↓Top-1 ↑MATH-500MMLU-ProSuperGPQALCB v6LongBench-v2RULERAvg.
Teacher (BF16)––49.642.322.417.030.089.641.8
RTN0.04494.444.535.520.114.229.689.038.8
GPTQ0.01596.847.940.020.616.531.288.540.8
AWQ0.03295.244.839.821.313.230.888.139.7
Standard QAT0.01497.047.742.121.115.831.089.041.1
LSQ0.01696.848.341.921.215.529.887.540.7
Denoising QAT0.01497.048.142.121.714.331.288.941.1
BitDistiller0.01397.048.441.021.015.731.088.941.0
QUASAR (ours)0.01297.347.642.220.816.430.089.041.0
Reasoning, coding, and long-context results for Qwen3-4B-Thinking-2507 and Llama-3.1-8B-Instruct at INT2, INT3, and INT4.

Adaptation with QAT.

QUASAR has the lowest held-out perplexity and the best or tied-best average accuracy at every bit width. It is also the only INT2 quantized model to outperform the original base model and score above zero on HMMT'25.

adaptation on math reasoning

Table 2 · Qwen3-4B-Base · INT2. Comparison of QAT and PTQ methods for adapting Qwen3-4B-Base on OpenMathReasoning at INT2, INT3, and INT4.

MethodPPL ↓MATH-500GSM8KAIME'24AIME'25HMMT'25Avg.
Base (no SFT)–41.066.49.63.30.824.2
FP-SFT1.5983.891.322.922.510.046.1
RTN3.5×10⁵0.90.30.00.00.00.2
GPTQ3.122.11.10.00.00.00.6
AWQ1301.51.20.00.00.00.5
Standard QAT1.7938.751.92.10.80.018.7
LSQ1.8236.549.90.80.40.017.5
Denoising QAT1.7839.252.50.80.80.018.7
BitDistiller1.8733.657.01.20.40.018.4
QUASAR (ours)1.7160.775.44.06.51.529.6

Table 2 · Qwen3-4B-Base · INT3

MethodPPL ↓MATH-500GSM8KAIME'24AIME'25HMMT'25Avg.
Base (no SFT)–41.066.49.63.30.824.2
FP-SFT1.5983.891.322.922.510.046.1
RTN2.2412.87.60.01.70.04.4
GPTQ1.6871.085.011.211.24.236.5
AWQ1.8357.365.110.08.31.728.5
Standard QAT1.6475.887.515.016.24.639.8
LSQ1.6573.787.012.511.73.837.7
Denoising QAT1.6475.987.314.213.37.139.6
BitDistiller1.6475.787.816.215.45.040.0
QUASAR (ours)1.6278.588.615.916.26.041.1

Table 2 · Qwen3-4B-Base · INT4

MethodPPL ↓MATH-500GSM8KAIME'24AIME'25HMMT'25Avg.
Base (no SFT)–41.066.49.63.30.824.2
FP-SFT1.5983.891.322.922.510.046.1
RTN1.6678.290.015.013.37.540.8
GPTQ1.6181.790.619.120.58.544.1
AWQ1.6380.390.515.817.58.342.5
Standard QAT1.6082.190.821.220.610.745.1
LSQ1.6279.089.117.917.16.742.0
Denoising QAT1.6082.190.420.021.49.944.8
BitDistiller1.6082.790.620.320.89.344.7
QUASAR (ours)1.6082.890.920.819.811.245.1
At INT2, QUASAR reaches 29.6 average accuracy, 10.9 points above the strongest QAT baseline and 29.0 points above full-precision fine-tuning followed by PTQ. Comparison of QAT and PTQ methods for adapting Qwen3-4B-Base on OpenMathReasoning at INT2, INT3, and INT4.
adaptation perplexity
QUASAR reaches the lowest held-out perplexity at every bit width. Training perplexity over the full run (left) and final 1,000 steps (center), and held-out perplexity (right), for QAT of Qwen3-4B-Base on OpenMathReasoning. FP-SFT is included as a full-precision reference.
answer outcomes · INT2
Answer outcomes for the INT2 checkpoints. Share of generated samples that produce a correct answer, an incorrect answer, or no final answer on MATH-500 and GSM8K. The base model and FP-SFT are included as references.
pass@8 / avg@1 · INT2
Variation across repeated INT2 samples. Ratio of pass@8 to avg@1 on MATH-500 for Qwen3-4B-Base checkpoints, using eight samples per problem at temperature 0.6. A ratio of 1 means that the same problems are solved on every attempt; higher values indicate less consistent successes. The base model and FP-SFT are included as references.
divergence along traces
Divergence from FP-SFT along reasoning traces. Forward KL to FP-SFT by trace decile for selected quantized checkpoints, averaged over 1,000 MATH-500 traces generated by FP-SFT. Every checkpoint is evaluated on the same token sequences.

5.NVFP4 and Checkpoint Comparisons

For Qwen3-8B and Qwen3.5-9B, QUASAR reduces held-out KL divergence by 12% and 15% relative to Standard QAT, respectively. Notably, QUASAR's INT4 Gemma-4 checkpoint outperforms Google's QAT checkpoint in KL divergence and benchmark accuracy.

fidelity to BF16 and accuracy

Table 3 · Qwen3-8B. Comparison of QUASAR with QAT and PTQ baselines across NVFP4, INT4, and Q4_0 (llama.cpp GGUF format).

MethodFormatbpwFidelity · KL ↓Fidelity · Top-1 ↑Accuracy (%) ↑ · Hard reasoning and coding · HMMT'26Accuracy (%) ↑ · Hard reasoning and coding · AIME'25Accuracy (%) ↑ · Hard reasoning and coding · SuperGPQAAccuracy (%) ↑ · Hard reasoning and coding · OJBenchAccuracy (%) ↑ · Long horizon · MRCRAccuracy (%) ↑ · Long horizon · MRCR ≥64kAccuracy (%) ↑ · Long horizon · MRCR 8-needleAccuracy (%) ↑ · Long horizon · LongProcAccuracy (%) ↑ · Long horizon · LongBench-v2Accuracy (%) ↑ · Long horizon · RULERAccuracy (%) ↑ · Avg.
BF16BF1616––44.668.448.822.425.819.216.248.938.480.541.3
RTNNVFP4 W4A44.50.05393.037.464.645.614.722.415.214.139.336.076.636.6
GPTQNVFP4 W4A44.50.03994.040.364.046.216.823.614.615.239.536.078.537.5
Standard QATNVFP4 W4A44.50.04194.039.061.245.719.421.614.112.936.537.477.936.6
QUASAR (ours)NVFP4 W4A44.50.03694.340.863.646.818.524.316.415.840.333.878.937.9

Table 3 · Qwen3.5-9B

MethodFormatbpwFidelity · KL ↓Fidelity · Top-1 ↑Accuracy (%) ↑ · Hard reasoning and coding · HMMT'26Accuracy (%) ↑ · Hard reasoning and coding · AIME'25Accuracy (%) ↑ · Hard reasoning and coding · SuperGPQAAccuracy (%) ↑ · Hard reasoning and coding · OJBenchAccuracy (%) ↑ · Long horizon · MRCRAccuracy (%) ↑ · Long horizon · MRCR ≥64kAccuracy (%) ↑ · Long horizon · MRCR 8-needleAccuracy (%) ↑ · Long horizon · LongProcAccuracy (%) ↑ · Long horizon · LongBench-v2Accuracy (%) ↑ · Long horizon · RULERAccuracy (%) ↑ · Avg.
BF16BF1616––47.974.558.819.479.474.553.966.843.790.160.9
RTNNVFP4 W4A44.50.04493.221.937.349.16.554.442.632.135.523.191.439.4
GPTQNVFP4 W4A44.50.02994.520.029.948.63.456.844.632.642.124.391.339.4
Standard QATNVFP4 W4A44.50.03394.243.165.355.513.463.752.339.759.043.989.852.6
QUASAR (ours)NVFP4 W4A44.50.02894.643.868.856.013.471.763.845.059.744.190.555.7

Table 3 · Muse-Glimmer-30B

MethodFormatbpwFidelity · KL ↓Fidelity · Top-1 ↑Accuracy (%) ↑ · Hard reasoning and coding · HMMT'26Accuracy (%) ↑ · Hard reasoning and coding · AIME'25Accuracy (%) ↑ · Hard reasoning and coding · SuperGPQAAccuracy (%) ↑ · Hard reasoning and coding · OJBenchAccuracy (%) ↑ · Long horizon · MRCRAccuracy (%) ↑ · Long horizon · MRCR ≥64kAccuracy (%) ↑ · Long horizon · MRCR 8-needleAccuracy (%) ↑ · Long horizon · LongProcAccuracy (%) ↑ · Long horizon · LongBench-v2Accuracy (%) ↑ · Long horizon · RULERAccuracy (%) ↑ · Avg.
BF16BF1616––71.792.260.125.046.138.024.967.060.488.457.4
RTNNVFP4 W4A164.50.05493.269.587.758.925.044.335.626.255.759.488.355.1
GPTQNVFP4 W4A44.50.06692.466.490.358.525.443.140.725.560.157.384.755.2
QUASAR (ours)NVFP4 W4A164.50.01896.172.390.859.327.248.540.028.567.358.490.258.2
QUASAR (ours)NVFP4 W4A44.50.05393.268.388.058.926.344.135.425.562.854.589.755.3

Table 3 · Gemma-4 E4B

MethodFormatbpwFidelity · KL ↓Fidelity · Top-1 ↑Accuracy (%) ↑ · Hard reasoning and coding · HMMT'26Accuracy (%) ↑ · Hard reasoning and coding · AIME'25Accuracy (%) ↑ · Hard reasoning and coding · SuperGPQAAccuracy (%) ↑ · Hard reasoning and coding · OJBenchAccuracy (%) ↑ · Long horizon · MRCRAccuracy (%) ↑ · Long horizon · MRCR ≥64kAccuracy (%) ↑ · Long horizon · MRCR 8-needleAccuracy (%) ↑ · Long horizon · LongProcAccuracy (%) ↑ · Long horizon · LongBench-v2Accuracy (%) ↑ · Long horizon · RULERAccuracy (%) ↑ · Avg.
BF16BF1616––32.840.942.121.626.715.412.656.742.989.738.1
RTNINT4 g644.25.07991.524.431.839.016.422.09.311.835.141.787.531.9
Google QATINT4 g324.50.06492.628.633.339.717.721.08.011.551.139.889.634.0
QUASAR (ours)INT4 g644.25.02295.529.037.540.518.124.412.012.753.640.489.535.8
RTNGGUF Q4_04.50.19485.718.832.836.514.222.99.111.039.038.083.830.6
Google QATGGUF Q4_04.50.08890.629.035.940.819.422.99.412.857.839.089.335.6
imatrix (PTQ)GGUF Q4_04.52.06791.128.340.140.919.024.711.713.046.239.088.935.2
QUASAR (ours)GGUF Q4_04.50.04492.827.837.639.619.424.411.812.753.140.489.635.6
Comparison of QUASAR with QAT and PTQ baselines across NVFP4, INT4, and Q4_0 (llama.cpp GGUF format).
long horizon · NVFP4 · Qwen3.5-9B
QUASAR better preserves long-horizon behavior in NVFP4. Expanded Qwen3.5-9B results. (a) OpenAI-MRCR score by context length; (b) share of long generations that terminate within the 32k-token output budget. Higher is better.
non-termination · NVFP4

Table 10 · Qwen3-8B · NVFP4. QUASAR has the lowest or tied-lowest non-termination rate among the quantized checkpoints.

MethodMATH-500LiveCodeBench v6OJBench (C++)OJBench (Python)
BF160.513.61.32.2
RTN1.440.820.326.7
GPTQ0.524.311.610.3
Standard QAT0.521.64.75.6
QUASAR (ours)0.518.43.02.2

Table 10 · Qwen3.5-9B · NVFP4

MethodMATH-500LiveCodeBench v6OJBench (C++)OJBench (Python)
BF1613.954.671.580.2
RTN37.878.194.491.8
GPTQ41.182.896.196.1
Standard QAT16.059.282.881.9
QUASAR (ours)14.852.979.779.3
QUASAR has the lowest or tied-lowest non-termination rate among the quantized checkpoints. Percentage of generations that reach the 32,768-token output budget without stopping. Lower is better.

6.Robustness and Training Overhead

learning-rate sweep · INT2 · 1,000 steps

Fig. 25 · Qwen3-4B-Thinking-2507 · INT2 · final held-out KL after 1,000 steps. QUASAR is robust to learning rate.

Methodlearning rate · 1×10⁻⁵learning rate · 2×10⁻⁵learning rate · 5×10⁻⁵learning rate · 1×10⁻⁴learning rate · 2×10⁻⁴learning rate · 5×10⁻⁴
QUASAR (ours)0.25230.20130.15160.14030.16750.2604
Standard QAT0.50070.36070.25330.21700.23991.5226
LSQ0.44900.34890.26150.22930.23680.7611
Denoising QAT0.47460.34970.24560.21280.24460.7311
BitDistiller1.22880.52410.24510.17700.20200.2963
QUASAR is robust to learning rate. Final held-out KL after 1,000 steps of INT2 QAD of Qwen3-4B-Thinking-2507 across a 50× learning-rate range. QUASAR attains the lowest KL at every tested rate; Standard QAT diverges at high rates, while BitDistiller degrades at low rates.

Table 4. Wall-clock time of one training step, per component and in seconds: quantization-aware distillation of Qwen3-4B-Thinking-2507 at INT3 on 8x H100 GPUs.

ComponentStandard QATQUASAR
Teacher forward0.3730.373
Student compute0.3660.366
Weight reconstruction0.2940.335
Loss computation0.0650.065
Backward1.8051.805
Optimizer update0.0210.022
Other + synchronization0.0060.006
Total step2.9302.972
% vs. Standard QAT–+1.4%
Wall-clock time of one training step, per component and in seconds: quantization-aware distillation of Qwen3-4B-Thinking-2507 at INT3 on 8x H100 GPUs. QUASAR differs from Standard QAT only in the weight reconstruction row. QUASAR only adds 1.4% wall-clock step time compared with Standard QAT (2.97 and 2.93 seconds per step, respectively), with the overhead only coming from the weight reconstruction. Compared with Standard QAT, QUASAR introduces no memory overhead.

7.Checkpoints

QUASAR checkpoints use standard layouts supported by vLLM's quantization backends.

In vLLM, Marlin and Machete serve INT4 weights, NVFP4 runs natively on Blackwell GPUs, and MXFP4 weight-only checkpoints are supported.

Table 11 · Qwen3.5-4B. Comparison with released 4-bit checkpoints across four model families.

ModelFormatbpwFidelity · KL ↓Fidelity · Top-1 ↑Accuracy (%) ↑ · GSM8K-PAccuracy (%) ↑ · MMLU-PAccuracy (%) ↑ · IFEvalAccuracy (%) ↑ · MATHAccuracy (%) ↑ · AIME25Accuracy (%) ↑ · GPQA-DAccuracy (%) ↑ · Avg.
BF16BF1616––94.078.988.084.479.277.383.6
RTNNVFP4 W4A164.50.04293.894.977.586.383.867.174.180.6
GPTQNVFP4 W4A164.50.02395.795.175.284.583.252.969.976.8
GPTQINT4 g1287.49.04993.194.077.587.183.065.071.279.6
AWQINT4 g324.50.04293.993.977.687.684.272.574.681.7
ModelOpt PTQNVFP4 W4A44.50.04393.794.277.988.284.665.873.680.7
QUASAR (ours)NVFP4 W4A164.50.01896.294.778.587.283.874.277.982.7
GGUF (llama.cpp)
RTNGGUF Q4_04.50.37289.094.377.290.684.264.273.780.7
KD-QAT g64GGUF Q4_04.50.36388.494.977.484.584.464.673.980.0
imatrix (PTQ)GGUF Q4_04.84.28490.494.378.090.686.070.876.182.6
QUASAR (ours)GGUF Q4_04.50.23591.394.978.290.983.877.974.683.4

Table 11 · Muse-Glimmer-30B. ‡ denotes mixed NVFP4, FP8, and BF16 precision.

ModelFormatbpwFidelity · KL ↓Fidelity · Top-1 ↑Accuracy (%) ↑ · GPQA-DAccuracy (%) ↑ · MMLU-PAccuracy (%) ↑ · AIME25Accuracy (%) ↑ · LCBAccuracy (%) ↑ · RUL32KAccuracy (%) ↑ · RUL64KAccuracy (%) ↑ · RUL128KAccuracy (%) ↑ · Avg.
BF16BF1616––62.183.186.753.486.685.281.076.9
RTNNVFP4 W4A164.50.05493.263.682.983.348.184.584.083.675.7
GPTQNVFP4 W4A44.50.06692.463.183.086.754.286.282.480.676.6
AutoQuantizeNVFP4+FP8‡5.52.03094.963.182.983.352.785.684.883.576.6
QUASAR (ours)NVFP4 W4A164.50.01896.163.683.390.049.687.286.484.177.7
QUASAR (ours)NVFP4 W4A44.50.05393.263.182.783.352.787.286.483.577.0

Table 11 · Gemma-4 12B

ModelFormatbpwFidelity · KL ↓Fidelity · Top-1 ↑Accuracy (%) ↑ · MMLUAccuracy (%) ↑ · IFEvalAccuracy (%) ↑ · ARC-CAccuracy (%) ↑ · AGIEvalAccuracy (%) ↑ · MMLU-PAccuracy (%) ↑ · TQAAccuracy (%) ↑ · MATH-hAccuracy (%) ↑ · Avg.
BF16BF1616––72.888.050.224.349.962.174.760.3
RTNINT4 g644.25.19489.260.679.746.419.439.060.650.550.9
GPTQINT4 g644.25.04195.268.586.947.822.748.262.170.958.2
Google QATINT4 g324.50.04494.867.887.250.122.647.960.073.258.4
QUASAR (ours)INT4 g644.25.03495.570.487.651.324.245.859.271.958.6
GGUF (llama.cpp)
RTNGGUF Q4_04.50.15689.859.287.447.822.738.460.262.354.0
Google QATGGUF Q4_04.50.06993.571.089.148.922.845.661.374.859.1
imatrix (PTQ)GGUF Q4_04.52.10891.664.987.148.623.543.460.366.856.4
QUASAR (ours)GGUF Q4_04.50.04894.370.387.651.524.245.559.172.558.7

Table 11 · Gemma-4 E4B

ModelFormatbpwFidelity · KL ↓Fidelity · Top-1 ↑Accuracy (%) ↑ · TrivQAAccuracy (%) ↑ · NQAccuracy (%) ↑ · TQAAccuracy (%) ↑ · ARC-EAccuracy (%) ↑ · ARC-CAccuracy (%) ↑ · GSM8KAccuracy (%) ↑ · IFEvalAccuracy (%) ↑ · MATH-hAccuracy (%) ↑ · Avg.
BF16BF1616––23.63.959.179.756.584.185.263.456.9
RTNINT4 g644.25.07991.525.23.059.378.355.279.783.056.155.0
GPTQINT4 g644.25.03194.721.33.356.178.254.983.583.459.455.0
Google QATINT4 g324.50.06492.620.12.455.478.356.081.581.959.454.4
QUASAR (ours)INT4 g644.25.02295.523.84.758.279.856.282.082.660.756.0
GGUF (llama.cpp)
RTNGGUF Q4_04.50.19485.717.32.058.576.253.875.880.248.751.6
Google QATGGUF Q4_04.50.08890.618.52.355.578.756.181.481.758.754.1
Google QAT repackGGUF Q4_04.50.07191.820.02.256.279.155.483.583.261.855.2
imatrix (PTQ)GGUF Q4_04.52.06791.120.63.058.879.353.880.483.259.754.9
QUASAR (ours)GGUF Q4_04.50.04492.824.04.758.279.756.181.882.460.555.9

QUASAR checkpoints · Hugging Face

ModelFormatRuntimeRepository
Qwen3.5-4BNVFP4vLLMQUASAR-QAT/Qwen3.5-4B-QUASAR-NVFP4
Qwen3.5-4BNVFP4 W4A4vLLMQUASAR-QAT/Qwen3.5-4B-QUASAR-NVFP4-W4A4
Qwen3.5-4BGGUF Q4_0llama.cppQUASAR-QAT/Qwen3.5-4B-QUASAR-Q4_0-GGUF
Qwen3.8-27BNVFP4vLLMQUASAR-QAT/Qwen3.8-27B-QUASAR-NVFP4
Gemma-4 E4BINT4 g64vLLMQUASAR-QAT/gemma-4-E4B-it-QUASAR-W4A16-G64
Gemma-4 E4BGGUF Q4_0llama.cppQUASAR-QAT/gemma-4-E4B-it-QUASAR-Q4_0-GGUF
Gemma-4 12BINT4 g64vLLMQUASAR-QAT/gemma-4-12B-it-QUASAR-W4A16-G64
Gemma-4 12BGGUF Q4_0llama.cppQUASAR-QAT/gemma-4-12B-it-QUASAR-Q4_0-GGUF
Muse-Glimmer-30BNVFP4vLLMQUASAR-QAT/Muse-Glimmer-30B-QUASAR-NVFP4
Muse-Glimmer-30BNVFP4 W4A4vLLMQUASAR-QAT/Muse-Glimmer-30B-QUASAR-NVFP4-W4A4
Muse-Glimmer-30BGGUF Q4_0llama.cppQUASAR-QAT/Muse-Glimmer-30B-QUASAR-Q4_0-GGUF

QUASAR collection · Hugging Face

Hugging FaceLink
All QUASAR Modelshuggingface.co/collections/QUASAR-QAT/all-quasar-models-6aab09fd2913ec22df0d5f15
QUASAR-QAThuggingface.co/QUASAR-QAT

Qwen3.8-27B · NVFP4 · Hugging Face model card · public NVFP4 builds · https://huggingface.co/QUASAR-QAT/Qwen3.8-27B-QUASAR-NVFP4

ModelSizeNVFP4 linearsGPQA-DAIME'26
QUASAR (this model)19.7 GB496/49690.91100.0
BF16 original55.6 GB—91.41100.0
Unsloth NVFP423.4 GB168/49689.3997.78
Inferact NVFP426.4 GB304/49687.6396.67

Qwen3.8-27B · NVFP4 · Hugging Face model card · LLM Compressor NVFP4

ModelMRCRMRCR 8-needleHMMT'26SuperGPQALiveCodeBenchRULER 32KRULER 64KRULER 128K
QUASAR (this model)79.152.773.462.479.296.396.295.8
BF16 original83.458.575.863.081.096.796.695.9
GPTQ NVFP4 (LLM Compressor)61.335.569.758.977.996.095.495.3
RTN NVFP4 (LLM Compressor)57.332.866.958.677.996.296.195.3

GGUF KL values are comparable only within each block.

huggingface.co/QUASAR-QAT

8.Code

python
import torch
from transformers import AutoModelForCausalLM
from quasar.export import save_materialized
from quasar.quant import QuantConfig, quantize_model, snapshot_saliency

path = "Qwen/Qwen3-4B-Thinking-2507"
model = AutoModelForCausalLM.from_pretrained(path, dtype=torch.bfloat16, device_map="cuda")
quantize_model(model, QuantConfig("quasar", bits=2))  # INT2, groups of 128
optimizer = torch.optim.AdamW(model.parameters(), lr=5e-5)

for batch in batches:
    model(**batch).loss.backward()
    optimizer.step()
    snapshot_saliency(model, optimizer)
    optimizer.zero_grad()

save_materialized(model, "qwen3-4b-w2", model_path=path)
shell
git clone https://github.com/vincentcounathe/quasar-qat && cd quasar-qat
pip install -e .            # quantizer and trainer
pip install flash-attn --no-build-isolation   # FlashAttention-2, used by the training recipes
pip install -e ".[eval]"    # + evaluation (vLLM, lm-eval, math-verify)

9.Citation

bibtex
@article{counathe2026quasar,
  title   = {{QUASAR}: Lowering the Loss Floor of Quantization-Aware Training with Loss-Aware Reconstruction},
  author  = {Counathe, Vincent and Athiwaratkun, Ben and De Sa, Christopher and Zhang, Tianyi},
  journal = {arXiv preprint arXiv:2608.13966},
  year    = {2026}
}

Appendix

Method Details

per-weight loss-aware reconstruction error · INT3
Per-weight loss-aware reconstruction error for 14 Qwen3-4B weight groups at INT3. Color shows |√h (r−w)| under (a) Standard QAT with min–max reconstruction and (b) QUASAR. Each row contains 128 weights ordered by saliency h.
scale search · INT3
Scale search for Qwen3-4B weight groups at INT3. (a) One extreme weight sets the min–max range. (b) QUASAR clips the outlier and covers the bulk. (c) Distribution of selected range factors f⋆ across 28.4M groups; 99.6% use a range narrower than min–max.
selected clipping-range factors
Distribution of selected clipping-range factors f⋆ across weight groups after QAD at INT4, INT3, and INT2 for Qwen3-4B-Thinking-2507 and Llama-3.1-8B-Instruct. Lower bit widths favor narrower ranges.
saliency in one layer
Saliency h in one Qwen3-4B layer. Weight groups span 128 weights along the input dimension. The large variation within each group lets QUASAR prioritize the most salient weights.

Empirical Validation of the Theory

reduction across projection types
QUASAR reduces reconstruction error across projection types. Median reduction in loss-aware reconstruction error relative to Standard QAT across Qwen3-4B-Thinking-2507 modules at INT4, INT3, and INT2, measured at initialization (left) and after training (right). Error is measured against the full-precision weights.
before training · 4,096 steps

Fig. 13 · Qwen3-4B-Thinking-2507 · 4,096 steps. Reconstruction error predicts final loss before training.

Methodbitsinit-time reconstruction error Ŝ₀KL to BF16 (Table 1)
QUASAR (ours)INT46.26×10⁻⁵0.016
QUASAR (ours)INT32.27×10⁻⁴0.054
QUASAR (ours)INT27.27×10⁻⁴0.126
Denoising QATINT46.93×10⁻⁵0.018
Denoising QATINT33.02×10⁻⁴0.060
Denoising QATINT21.29×10⁻³0.179
Standard QATINT47.10×10⁻⁵0.019
Standard QATINT33.33×10⁻⁴0.062
Standard QATINT22.18×10⁻³0.186

Fig. 13 · log-log power-law fit

fitR²
Ŝ₀ vs KL to BF160.98
Reconstruction error predicts final loss before training. Across Standard QAT, Denoising QAT, and QUASAR at INT4, INT3, and INT2, reconstruction error at initialization predicts held-out KL to the full-precision model after 4,096 steps (R²=0.98). Results use Qwen3-4B-Thinking-2507.
gradient mismatch
Reconstruction error controls gradient mismatch. Across INT4, INT3, and INT2, the squared gradient mismatch rises with the loss-aware reconstruction error Ŝ, supporting Assumption 3. Circles, triangles, and squares show fixed clipping factors for INT4, INT3, and INT2, respectively; stars show QUASAR's per-group scale-search solutions.
saliency and curvature
QUASAR's saliency tracks curvature. Each point represents one weight matrix from a trained INT2 Qwen3-4B-Thinking checkpoint and compares its mean saliency h (Adam's second moment) with a 32-sample Hutchinson estimate of its mean Hessian diagonal. The log-scale correlation is r=0.81.
PL relation · INT2
Training trajectories support a PL relation. Each point is one step of an INT2 Qwen3-4B-Thinking run, with loss and gradient measured at the reconstruction r. For all three methods, the squared gradient norm scales with the gap above the trajectory minimum. QUASAR's estimated μ̂=1.72 is about 2.5× the baseline values. In Theorem 1(ii), a larger μ gives faster geometric convergence and smaller floor terms, which scale with 1/μ or 1/μ². This provides an optimization-side view of the healing curves. QUASAR's reconstruction adds less loss, so more of the remaining gap produces useful gradient, enabling faster healing and a lower loss floor.
bound terms · INT2 · SGD
Reconstruction error dominates the measured bound. We evaluate each term of Theorem 1(i) on an INT2 QUASAR run with plain SGD and learning rate 10⁻³. We measure λmax by power iteration (approximating L), σ² from per-batch gradients, and C. The resulting bound is within 3.1× of the observed average squared gradient norm, and reconstruction error is its largest term.
plain SGD · INT2
QUASAR retains its advantage under plain SGD. Training (top) and held-out (bottom) loss for Standard QAT and QUASAR at INT2 across three constant learning rates. QUASAR reaches lower loss floors at every stable learning rate.

Additional Healing Results

Standard Tasks and Fidelity

held-out KL

Table 6 · Qwen3-4B-Thinking-2507 · INT2. QUASAR gives the highest fidelity in all six model–bit-width settings.

MethodKL ↓Top-1 ↑GSM8KMMLUARC-CARC-EHellaSwagWinoGrandeTruthfulQAIFEvalAvg.
Teacher (FP16)––87.068.753.176.765.766.157.554.366.1
RTN11.1660.50.025.425.625.926.052.046.88.326.3
GPTQ1.14067.70.223.425.130.030.149.352.98.927.5
AWQ3.01637.00.023.423.231.629.151.152.89.627.6
Standard QAT0.18689.947.434.531.244.846.954.551.133.543.0
LSQ0.20589.242.634.630.644.947.655.247.235.142.2
Denoising QAT0.17990.145.636.030.147.946.756.149.934.943.4
BitDistiller0.20989.149.043.238.060.152.657.952.637.748.9
QUASAR (ours)0.12692.068.848.938.855.956.459.250.547.553.2

Table 6 · Qwen3-4B-Thinking-2507 · INT3

MethodKL ↓Top-1 ↑GSM8KMMLUARC-CARC-EHellaSwagWinoGrandeTruthfulQAIFEvalAvg.
Teacher (FP16)––87.068.753.176.765.766.157.554.366.1
RTN0.59778.820.553.639.456.154.158.151.418.944.0
GPTQ0.13091.577.361.048.075.561.061.854.049.261.0
AWQ0.24088.173.260.146.472.360.762.749.241.858.3
Standard QAT0.06294.278.562.046.874.161.964.456.153.262.1
LSQ0.07093.883.062.349.475.862.765.655.452.963.4
Denoising QAT0.06094.380.462.746.273.561.664.256.151.062.0
BitDistiller0.06594.281.063.747.069.063.063.556.353.462.1
QUASAR (ours)0.05494.981.764.548.372.963.264.755.550.662.7

Table 6 · Qwen3-4B-Thinking-2507 · INT4

MethodKL ↓Top-1 ↑GSM8KMMLUARC-CARC-EHellaSwagWinoGrandeTruthfulQAIFEvalAvg.
Teacher (FP16)––87.068.753.176.765.766.157.554.366.1
RTN0.10892.280.266.148.869.763.565.458.454.263.3
GPTQ0.02796.284.967.151.475.064.765.456.551.464.6
AWQ0.05094.784.667.651.574.965.266.656.756.665.5
Standard QAT0.01996.984.967.049.773.565.165.256.754.364.6
LSQ0.02496.486.066.850.675.065.265.657.156.465.3
Denoising QAT0.01896.985.067.149.272.964.765.757.356.964.9
BitDistiller0.01896.984.867.752.676.064.865.957.354.365.4
QUASAR (ours)0.01697.186.468.050.676.065.165.257.755.865.6

Table 6 · Llama-3.1-8B-Instruct · INT2

MethodKL ↓Top-1 ↑GSM8KMMLUARC-CARC-EHellaSwagWinoGrandeTruthfulQAIFEvalAvg.
Teacher (FP16)––70.168.355.679.979.573.654.573.969.4
RTN10.8100.50.025.526.226.126.151.847.611.826.9
GPTQ2.38051.80.024.923.328.930.048.049.09.126.7
AWQ6.76912.90.023.623.325.927.248.248.87.925.6
Standard QAT0.16390.11.823.422.930.934.849.346.844.431.8
LSQ0.17289.642.239.732.852.855.753.547.442.145.8
Denoising QAT0.16590.10.024.626.225.626.950.049.140.930.4
BitDistiller0.15190.253.149.343.569.966.666.347.751.856.0
QUASAR (ours)0.10592.066.452.543.873.168.166.849.556.059.5

Table 6 · Llama-3.1-8B-Instruct · INT3

MethodKL ↓Top-1 ↑GSM8KMMLUARC-CARC-EHellaSwagWinoGrandeTruthfulQAIFEvalAvg.
Teacher (FP16)––70.168.355.679.979.573.654.573.969.4
RTN0.27685.619.044.840.264.669.566.947.052.950.6
GPTQ0.07692.861.459.249.171.675.771.053.169.163.8
AWQ0.19788.023.953.243.067.672.267.642.158.653.5
Standard QAT0.04694.770.759.747.472.875.272.952.169.965.1
LSQ0.05294.276.862.351.876.975.871.650.967.566.7
Denoising QAT0.04694.768.859.848.074.575.072.251.070.264.9
BitDistiller0.04194.874.261.653.179.076.471.952.268.667.1
QUASAR (ours)0.03695.274.163.354.979.177.171.051.671.567.9

Table 6 · Llama-3.1-8B-Instruct · INT4

MethodKL ↓Top-1 ↑GSM8KMMLUARC-CARC-EHellaSwagWinoGrandeTruthfulQAIFEvalAvg.
Teacher (FP16)––70.168.355.679.979.573.654.573.969.4
RTN0.04494.470.764.153.376.978.073.950.772.167.5
GPTQ0.01596.869.067.352.077.378.773.255.173.668.3
AWQ0.03295.265.666.653.778.378.772.453.773.467.8
Standard QAT0.01497.076.366.354.477.778.473.853.973.069.2
LSQ0.01696.873.266.554.680.178.774.153.274.169.3
Denoising QAT0.01497.077.066.255.777.878.673.953.672.169.3
BitDistiller0.01397.070.467.055.179.778.873.353.874.169.0
QUASAR (ours)0.01297.374.666.555.578.978.874.254.272.569.4
QUASAR gives the lowest held-out KL across both models and all bit widths. Forward KL to the full-precision model after QAD of Qwen3-4B-Thinking-2507 (left) and Llama-3.1-8B-Instruct (right). Lower is better.
accuracy vs bit width
QUASAR attains the highest INT2 average accuracy on both models. Average accuracy over the eight benchmarks. Gray circles mark the full-precision models.
per-task accuracy · INT2
QUASAR preserves INT2 accuracy broadly across tasks. Per-task accuracy normalized by the full-precision model for Qwen3-4B-Thinking-2507 (left) and Llama-3.1-8B-Instruct (right) on the eight benchmarks.

LLM-as-a-Judge Preference at 2 Bits

LLM judge · INT2

Table 7 · INT2 · judge Llama-3.3-70B-Instruct. QUASAR wins all 14 comparisons with quantized baselines.

vs. QUASARQwen3-4B-Thinking-2507 · meanQwen3-4B-Thinking-2507 · 95% CILlama-3.1-8B-Instruct · meanLlama-3.1-8B-Instruct · 95% CI
RTN+1.00[+1.00, +1.00]+0.98[+0.96, +1.00]
GPTQ+1.00[+1.00, +1.00]+0.97[+0.94, +0.99]
AWQ+1.00[+1.00, +1.00]+0.99[+0.97, +1.00]
Standard QAT+0.32[+0.21, +0.42]+0.97[+0.94, +0.99]
LSQ+0.27[+0.15, +0.38]+0.37[+0.25, +0.49]
Denoising QAT+0.31[+0.21, +0.41]+0.98[+0.95, +1.00]
BitDistiller+0.24[+0.12, +0.35]+0.29[+0.17, +0.41]
FP16−0.30[−0.40, −0.20]−0.46[−0.56, −0.37]
The judge prefers QUASAR to every INT2 quantized baseline. Pairwise win, tie, and loss rates on responses to 128 WildChat prompts for Qwen3-4B-Thinking-2507 and Llama-3.1-8B-Instruct. Llama-3.3-70B-Instruct serves as the judge. QUASAR wins all 14 comparisons with quantized baselines. Mean per-prompt preference in [−1,1], with 95% bootstrap confidence intervals over 128 prompts. Positive values favor QUASAR. The final row compares QUASAR with the full-precision model.
response stability · INT2
QUASAR matches the full-precision models' response stability at INT2. Mean word-repetition ratio over 128 WildChat prompts for Qwen3-4B-Thinking-2507 (left) and Llama-3.1-8B-Instruct (right), with greedy decoding. Dashed lines mark the full-precision models; counts show responses that close the reasoning block for Qwen or stop before the token cap for Llama.
responses
The beginning of each method's response to one held-out prompt, for Qwen3-4B-Thinking-2507 at INT2, INT3, and INT4 with greedy decoding.