Image-Text-to-Text
NInfer
qwen3.8
nvfp4
w4a4
dflash2
speculative-decoding
blackwell
multimodal
conversational
cuda
rtx-5090
Eval Results (legacy)
Instructions to use cometkim/Qwen3.8-27B-nvfp4full-NInfer with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NInfer
How to use cometkim/Qwen3.8-27B-nvfp4full-NInfer with NInfer:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
QUASAR QAT variant profile is ready
#3
by cometkim - opened
https://huggingface.co/cometkim/Qwen3.8-27B-nvfp4qat-NInfer
It allows replacing all 9 existing exception BF16 tensors with NVFP4 using a quantization-aware training approach.
While this results in a smaller and faster model without a loss in precision, the difference is not as significant as expected: 16.02 GiB vs. 15.31 GiB. Since speeds may decrease due to changes in acceptance rates when using MTP or DFlash2 speculative decoding, its speed needs to be re-evaluated for each specific environment and actual workloads.