Configuration Parsing Warning:In config.json: "quantization_config.modules_to_not_convert" must be an array

AutoAWQ INT4 (w_bit=4, q_group_size=128, GEMM) of Qwen2.5-1.5B-Instruct, run on akula-prime 3090 Ti after gpu-pack free of qwen3.5-9b. vLLM restored.

Not BitNet. Not GGUF Q4. Pascal 1080ti stays BitNet I2_S. Hot 3090 production remains QuantTrio/Qwen3.5-9B-AWQ. This 1.5B AWQ is the warm-queue / HF-library card-fit for smaller Ampere jobs.

Downloads last month
40
Safetensors
Model size
2B params
Tensor type
I32
路
BF16
路
F16
路
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for tzervas/Qwen2.5-1.5B-Instruct-AWQ

Quantized
(315)
this model