QUASAR QAT variant profile is ready

#3
by cometkim - opened

https://huggingface.co/cometkim/Qwen3.8-27B-nvfp4qat-NInfer

It allows replacing all 9 existing exception BF16 tensors with NVFP4 using a quantization-aware training approach.

While this results in a smaller and faster model without a loss in precision, the difference is not as significant as expected: 16.02 GiB vs. 15.31 GiB. Since speeds may decrease due to changes in acceptance rates when using MTP or DFlash2 speculative decoding, its speed needs to be re-evaluated for each specific environment and actual workloads.

Sign up or log in to comment