Instructions to use Blackfrost-AI/CYBER-FROST-3.8-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Blackfrost-AI/CYBER-FROST-3.8-NVFP4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Blackfrost-AI/CYBER-FROST-3.8-NVFP4") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Blackfrost-AI/CYBER-FROST-3.8-NVFP4") model = AutoModelForMultimodalLM.from_pretrained("Blackfrost-AI/CYBER-FROST-3.8-NVFP4", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Blackfrost-AI/CYBER-FROST-3.8-NVFP4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Blackfrost-AI/CYBER-FROST-3.8-NVFP4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Blackfrost-AI/CYBER-FROST-3.8-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Blackfrost-AI/CYBER-FROST-3.8-NVFP4
- SGLang
How to use Blackfrost-AI/CYBER-FROST-3.8-NVFP4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Blackfrost-AI/CYBER-FROST-3.8-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Blackfrost-AI/CYBER-FROST-3.8-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Blackfrost-AI/CYBER-FROST-3.8-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Blackfrost-AI/CYBER-FROST-3.8-NVFP4", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Blackfrost-AI/CYBER-FROST-3.8-NVFP4 with Docker Model Runner:
docker model run hf.co/Blackfrost-AI/CYBER-FROST-3.8-NVFP4
CYBER-FROST-3.8-NVFP4
A first-party Blackfrost-AI mixed-precision NVFP4 model for security professionals conducting authorized research, assessment, engineering, and response work.
Release status and contents
Cyber-Frost is a public research release under active quality assessment.
This repository contains a standalone mixed-precision checkpoint derived from Blackfrost-AI/CYBER-FROST-3.8-BF16, plus its weight index, configuration, ModelOpt quantization metadata, tokenizer and processor assets, packaged chat template, and upstream license. It is not an adapter and does not require the BF16 parent checkpoint at load time.
| Field | Released artifact |
|---|---|
| Clean model name | CYBER-FROST-3.8-NVFP4 |
| Former name | BLACKFROST-3.8-ICED-NVFP4-W4A4 |
| Architecture | Qwen4ExpForConditionalGeneration |
| Precision | Mixed: routed language-model experts use ModelOpt NVFP4 W4A4, group size 16; non-target tensors remain BF16 |
| Weight layout | 131 SafeTensors shards |
| Logical source architecture | approximately 180B parameters |
| Indexed tensor payload | 186,356,367,352 bytes (173.56 GiB) |
| Configured context | 262,144 tokens |
| Native speculative head | one BF16 MTP layer preserved from the BF16 source |
| Validated modality | text |
The configuration includes a vision tower, but this release has not received a multimodal quality evaluation. Do not infer validated image or video capability from the presence of processor files.
Why Cyber-Frost exists
Security work is unusually vulnerable to false refusals. The same vocabulary appears in incident response, exploit validation, malware analysis, defensive engineering, and unauthorized activity; a general-purpose assistant can react to individual terms instead of the operator's legitimate scope.
Cyber-Frost is designed to reduce that unnecessary friction in professional, authorized workflows. It is intended to stay technically direct when an analyst is reviewing a finding, reproducing a vulnerability in a controlled environment, writing detection content, analyzing malicious code, or operating an approved security agent. This is a design objective, not a measured refusal claim for this NVFP4 variant, and not a claim that every answer is safe or correct.
Authorization is an external control. The model cannot establish ownership, consent, rules of engagement, jurisdiction, or whether a target is in scope. Deployers must enforce identity, scope, tool permissions, logging, rate limits, and human review outside the model.
Security corpus
The BF16 parent was fine-tuned on a Blackfrost-AI security corpus combining curated security material, operator-authored workflows, realistic engagement-style scenarios, and Blackfrost-owned distillation data. Corpus sizes, source-by-source counts, raw engagement material, client identities, prompts, and responses are intentionally not published.
Domain coverage includes:
- Reconnaissance and OSINT
- Social engineering, business-email compromise, and deepfake-enabled abuse
- Web application and API security
- Identity, authentication, and Active Directory security
- Network, perimeter, VPN, and protocol security
- Vulnerability research, bug bounty, and binary exploitation
- Malware analysis, ransomware, and endpoint defense
- Cloud, container, and Kubernetes security
- Software supply-chain security
- Mobile, IoT, wireless, and physical security
- Industrial-control-system and operational-technology security
- Cryptography and security protocols
- Privilege escalation, lateral movement, and data exfiltration
- Threat intelligence, APT analysis, and purple-team operations
- AI-agent, LLM, and adversarial-ML security
Blackfrost-AI attests that owned portions of the corpus were developed from sanitized experience with authorized security work. It also applies a frontier-scale policy to its distillation teachers, excluding teachers below the 753B-parameter class. The release evidence independently binds one security subset to a Qwen3.8 2.4T teacher; it does not include a corpus-wide teacher manifest. These are therefore operator provenance statements, not independent benchmark findings.
Training-data provenance and licensing review for the mixed-source corpus remains in progress. The NVFP4 conversion added no new fine-tuning data; this section describes the BF16 parent inherited by the quantized artifact.
Model specifications and precision layout
The text stack has 48 blocks with hybrid linear and full attention, using a full-attention block every fourth layer. Its hidden size is 2,560 with 24 attention heads and 2 KV heads. The MoE stack contains 512 routed experts, selects 10 experts per token, and includes a shared expert. One native MTP layer is packaged for speculative decoding.
This is a mixed-precision ModelOpt checkpoint, not an all-tensor four-bit conversion. The routed expert blocks in all 48 language layers are configured for NVFP4 weights and input activations with group size 16. The 73,728 routed expert projection matrices use packed four-bit weights, FP8 E4M3 block scales, FP32 outer scales, and FP32 input scales.
Attention and linear-attention state, embeddings, normalization tensors, routers, shared experts, the PLE table, vision components, and the MTP layer are excluded from NVFP4 conversion and remain BF16. KV-cache precision is chosen by the serving runtime and is not encoded in the checkpoint weights.
The BF16 PLE table is large. Memory-constrained systems may require a runtime with CPU or file-backed PLE handling. The 262,144-token configuration ceiling is not a blanket quality guarantee; NVFP4-specific long-context, high-concurrency, multimodal, and tool-heavy agent-loop qualification remains pending.
Lineage
- Foundational checkpoint:
Qwen/Qwen3.8-Flash-Nextat immutable revisionde4b8e4d43b917e7706784d8bb445c9af86a3540. - Blackfrost security adaptation: security-domain fine-tuning followed by a full BF16 merge. The merged internal stage was identified as
BLACKFROST-3.8-FLASH-BF16. - Behavioral stage: a Blackfrost-AI behaviorally modified derivative targeting lower false-refusal friction in authorized security workflows. The proprietary transformation process is not distributed.
- BF16 conversion source: the payload now published as
Blackfrost-AI/CYBER-FROST-3.8-BF16, at immutable source revision5321904427c4ef54df8a667edcbc2d1184e4286e. - NVFP4 conversion: the routed language-model expert projections were converted to ModelOpt NVFP4 W4A4 with group size 16. Excluded tensors were preserved in BF16. The conversion used ModelOpt commit
022767c7ab3d7d36211affd85e5c496770cde768; its installed package reports version0.47.0rc0. - Release identity: the resulting checkpoint was formerly labeled
BLACKFROST-3.8-ICED-NVFP4-W4A4and is now namedCYBER-FROST-3.8-NVFP4. The rename is not another conversion or training run.
The quantization process used calibrated input-scale metadata from nvidia/Qwen3.8-Flash-Next-NVFP4 at revision fc694b54fb0174e0913e6adf86691ef85a4ead47. That checkpoint is a conversion reference, not the behavioral or weight lineage of Cyber-Frost.
Tokenizer and processor lineage comes through the pinned Qwen foundation and BF16 source. The tokenizer, processor, generation configuration, license, and packaged Blackfrost chat template are byte-identical to the current BF16 release assets.
Artifact verification
Conversion and integrity validation completed on 2026-09-17. The weight and configuration payload was re-audited at NVFP4 repository revision 69077d67a6274b6f7f6c6714b28f13da265bfd12; this model-card update does not alter that payload.
| Artifact | SHA-256 |
|---|---|
config.json |
91d03d21273f4761e5a95c8f0c879eec7207af1d0eae8739596e98755fd8557c |
hf_quant_config.json |
8d6c3fb3ac2cfc6f6f9e494ec6c17a6a61c7caba43b9f5b0332b8035879d17ec |
model.safetensors.index.json |
fc68c1d90e62460c4757a515619d8a533a43156581ac19469424dac81a217588 |
| Qwen Community License file | a0dc422560841fd68e06d974907f8b4c709bca44a67daad2b528437bdf676c08 |
The validation indexed all 131 shards and 296,474 tensors. It compared 1,562 tensors intentionally preserved from the BF16 source and found them exact. A 144-matrix sample of converted expert weights had a mean relative Frobenius error of approximately 0.0949 and a maximum of approximately 0.0952 against the BF16 source.
The preserved MTP tensors predate the final trunk-weight behavioral stage. Treat the MTP head as a provisional acceleration baseline rather than a freshly adapted draft head.
Evaluation status
Refusal and capability evaluation
No NVFP4-specific refusal or over-refusal result is published yet. The BF16 results are intentionally not copied here because quantization, runtime behavior, prompts, and sampling can change model behavior. Fresh measurements must be run against this exact NVFP4 artifact before any refusal rate is reported.
No standardized cyber-capability benchmark has yet been qualified for this exact checkpoint. This card does not claim a CyberMetric, SecBench, MMLU, HumanEval, or IFEval score. Refusal behavior is not a substitute for measuring security competence.
Legacy single-GB10 development measurement
On 2026-09-17, this exact NVFP4 payload was exercised on one NVIDIA GB10 in a DGX Spark with tensor parallelism 1, a patched vLLM development build (0.1.dev20073+g8e685d198), FP8 KV cache, native MTP with three speculative steps, and a 262,144-token configured serving context. The benchmark used warm single-stream requests with fixed 256-token outputs and two measured runs per category.
| Category | Mean decode speed | Mean end-to-end speed | Mean TTFT |
|---|---|---|---|
| Prose | 26.40 tok/s | 25.32 tok/s | 453 ms |
| Code | 29.29 tok/s | 28.08 tok/s | 413 ms |
These are narrow development measurements, not production throughput or quality guarantees. They do not predict RTX, B300, multi-GPU, concurrent, or long-context performance. Runtime version, prompt length, sampling, context occupancy, cache precision, PLE handling, and MTP acceptance can materially change the result.
Prompt, tool use, and sampling
The repository includes a Blackfrost chat template that supplies a default operating prompt, supports caller-provided system context, exposes Qwen-style reasoning controls, and serializes tool calls. Its default is thinking enabled; supported reasoning-effort values are xhigh, medium, and low. Generation defaults are temperature 1.0, top-p 0.95, and top-k 20.
The template's authorization assumption is not an access-control mechanism. An agent runtime must independently restrict credentials, targets, files, networks, commands, and approval-requiring actions. Tool-call output is proposed text until an external executor acts on it.
Changing the template, system message, reasoning mode, sampling, quantization backend, or runtime can materially change refusal behavior and output quality. Record those settings when reporting results.
Deployment
Use a runtime and hardware path with explicit support for the Qwen4ExpForConditionalGeneration architecture and mixed ModelOpt NVFP4 checkpoints. The validated legacy GB10 path required a patched runtime and file-backed PLE offload; the BF16 deployment kit is not interchangeable with this artifact.
Runtime compatibility, kernel selection, memory use, throughput, and output quality are implementation-dependent. This repository does not yet include a generally qualified deployment kit. Record the runtime version, hardware, context, quantization backend, PLE strategy, speculative settings, chat template, and sampling parameters when reporting results. The clean API model identifier is CYBER-FROST-3.8-NVFP4.
Intended use
Cyber-Frost is intended for qualified security professionals working within explicit authorization, including defensive research, secure code review, vulnerability validation, red-team and purple-team exercises, detection engineering, incident response, malware analysis, bug hunting, and controlled security-agent workflows.
It is not intended to authorize access, choose targets, define rules of engagement, make autonomous high-impact decisions, or replace legal, compliance, and safety review. Do not use it to access systems or data without permission, evade oversight, persist in third-party environments, deploy malware, steal credentials or data, disrupt services, or cause physical harm.
Limitations and security responsibility
- Generated findings, code, commands, indicators, and remediation advice may be wrong, incomplete, outdated, or fabricated. Independently review them and execute only in isolated, authorized environments.
- Reduced over-refusal is a design objective inherited from the BF16 parent, not a measured result for this NVFP4 artifact. Lower-friction behavior can increase the chance of receiving actionable output in ambiguous or malicious contexts.
- BF16 behavior, safety observations, benchmark results, and runtime characteristics do not automatically transfer through quantization.
- The model is not a policy engine, authorization service, sandbox, malware scanner, or secrets boundary.
- The preserved MTP head is provisional and is not evidence of adaptation to the final trunk.
- Current evidence does not establish an NVFP4 refusal rate, standardized cyber competence, production readiness, long-context quality, or multimodal quality.
- Model behavior can shift substantially with prompts, sampling, runtime versions, quantization kernels, speculative settings, and agent scaffolding.
The deployer is responsible for authorization, least privilege, isolation, network policy, credential handling, human approval gates, monitoring, incident response, and compliance with applicable law.
License and disclaimer
Use and redistribution of this checkpoint are governed by the Qwen Community License 1.0. Review the license before use. This research release is provided without a warranty of correctness, fitness, security, or non-infringement.
Report reproducible model or packaging issues through the repository's Discussions page without including secrets, client data, live targets, or sensitive exploit details.
- Downloads last month
- 377
Model tree for Blackfrost-AI/CYBER-FROST-3.8-NVFP4
Base model
Qwen/Qwen3.8-Flash-Next