You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Access to this model is granted on request. Please tell us who you are and how you intend to use it. The model is a guardrail classifier trained on corpora that contain harmful prompts; it is provided for scope and safety classification, not for generating content.

Log in or Sign Up to review the conditions and access this model content.

ScopeGuard v2 · 4B (q-2609)

A 4B-parameter guardrail that decides whether the last user message of a conversation is something a given AI service should handle. You describe your service once, in plain text or as a structured document, and the model classifies each incoming message against that description into one of seven scope classes, returning the class and, on request, the evidence, a short rationale, and a ready-to-send reply.

It is a fine-tune of Qwen3.5-4B with the adapter merged in. Weights are bfloat16, about 8.5 GB. Context is 262k tokens; typical requests are 2k to 4k.

Use it through orbitals

The model is designed to be run with orbitals, Principled Intelligence's client library, version 0.5.0 or later. The library owns the prompt: it ships the exact system prompt this checkpoint was trained on, lays out the user turn, disables the chat template's thinking mode, constrains decoding to the JSON schema of the fields you asked for, and validates the answer. Earlier orbitals versions use the previous prompt generation and will degrade this model silently, so pin >=0.5.0.

In-process with vLLM

pip install "orbitals[scope-guard-vllm]>=0.5.0"
from orbitals.scope_guard_v2 import ScopeGuardV2
from orbitals.types import AIServiceDescriptionV2, PredefinedResponse

sg = ScopeGuardV2(
    backend="vllm",
    model="principled-intelligence/scope-guard-v2-4B-q-2609",
    output_fields=["scope_class"],          # add "reasoning", "evidences", "suggested_response" as needed
    max_model_len=30_000,
    gpu_memory_utilization=0.9,
)

service = AIServiceDescriptionV2(
    identity_role="Virtual assistant for a parcel delivery service.",
    context="Customers of an e-commerce logistics company tracking their orders.",
    functionalities=["Track a parcel by tracking number", "Explain delivery time windows"],
    constraints=["Never process or promise refunds"],
    escalation_criteria=["The customer reports a lost or stolen parcel"],
    predefined_responses=[
        PredefinedResponse(trigger="Asks for the customer service phone number",
                           response="You can reach us at +1 555 0100, Mon-Fri 9-18."),
    ],
)

result = sg.validate(
    [
        {"role": "user", "content": "I ordered a package, tracking number 1234567890"},
        {"role": "assistant", "content": "It is in transit. What would you like to know?"},
        {"role": "user", "content": "If it doesn't arrive tomorrow, can I get a refund?"},
    ],
    ai_service_description=service,
)
print(result.scope_class.value)   # Restricted

The description can also be a plain string. The structured form is recommended: its constraints, escalation_criteria and predefined_responses fields map directly onto Restricted, Human Oversight and Predefined Answer, which is where free text is read least reliably. batch_validate classifies many conversations in one call.

As a service

orbitals starts vLLM and a small HTTP API in front of it:

pip install "orbitals[scope-guard-serve]>=0.5.0"

orbitals scope-guard-v2 serve principled-intelligence/scope-guard-v2-4B-q-2609 \
  --port 8000 --output-fields scope_class \
  --vllm-extra-args "--speculative-config '{\"method\":\"mtp\",\"num_speculative_tokens\":2}'"

The speculative-decoding argument is optional; the checkpoint ships Qwen3.5's multi-token-prediction head, worth about 1.3x on decode. Then either call the endpoint directly:

curl -X POST http://localhost:8000/orbitals/scope-guard-v2/validate \
  -H 'Content-Type: application/json' \
  -d '{"conversation": "If it does not arrive tomorrow, can I get a refund?",
       "ai_service_description": "Virtual assistant for a parcel delivery service. Only package tracking. Never process refunds."}'

or point a client at it, which lets you pick fields per call:

sg = ScopeGuardV2(
    backend="vllm-api",
    model="principled-intelligence/scope-guard-v2-4B-q-2609",
    vllm_serving_url="http://localhost:8001",   # the vLLM port behind the service
)
result = sg.validate(query, ai_service_description=service, output_fields=["reasoning", "scope_class"])

AsyncScopeGuardV2 offers the same interface for async code. The client renders the prompt with the model's tokenizer, so it needs read access to this repo too.

With transformers only

The repo registers a custom scope-guard-v2 pipeline, so the hf backend works without vLLM. It is the slowest option and suited to testing, not serving:

sg = ScopeGuardV2(backend="hf", model="principled-intelligence/scope-guard-v2-4B-q-2609")

Decoding is greedy in every backend. Do not turn sampling on.

The seven classes

class meaning typical action
Directly Supported Clearly within the service's stated functionalities. Answer.
Potentially Supported Adjacent to the stated functionalities or knowledge scope without being named by them, and no rule applies. Answer with care.
Predefined Answer Matches a trigger with a pre-written response in the description. Send the fixed reply.
Human Oversight Meets an escalation criterion, or needs human judgment. Route to a person.
Out of Scope Outside the service's remit. Off-topic, not harmful. Decline politely.
Restricted Explicitly forbidden by the description's constraints, or harmful. Refuse.
Chit Chat Greetings, thanks, small talk. Reply briefly.

Priority when several could apply: predefined responses, then escalation criteria, then constraints, then the more specific or more restrictive class. Potentially Supported is never chosen when a constraint, escalation criterion or predefined response matches.

For a binary safe/unsafe gate, treat Restricted and Human Oversight as unsafe and Out of Scope as safe: an off-topic question is not a harmful one.

Writing the service description

The description drives every decision, so its quality is the main lever you have.

  • Structure helps. The model looks for identity and role, context, knowledge scope, functionalities, constraints, predefined responses, escalation criteria, and response guidelines. Labelled sections with those names are read most reliably. Free-form prose works, but restrictions buried in a paragraph are missed more often.
  • State restrictions explicitly and rank them. A permissive opening paragraph followed by an unranked list of things to avoid produces under-refusal. Putting constraints, escalation criteria and predefined responses in their own sections, ahead of the marketing copy, fixes most of that without any retraining.
  • Predefined responses need a trigger and a text. The model emits the text you wrote, translated into the user's language if needed.
  • Keep it under a few thousand tokens. Longer descriptions work but cost latency on every request. The system prompt is constant, so with prefix caching only the description and conversation are re-encoded.

Choosing output fields

Ask for the fewest fields you need; class accuracy does not depend on which others are requested. Measured on eight public safety benchmarks:

requested fields balanced accuracy unsafe F1 notes
scope_class only 0.871 0.855 fastest, about 10 output tokens
reasoning, scope_class 0.871 0.854 adds a one-sentence rationale; 50 to 100 output tokens
evidences, scope_class not re-measured for this checkpoint quotes the description; on the previous checkpoint it trailed reasoning by about 1.5 points externally and led on in-domain data

Class accuracy is the same whether or not you ask for a rationale, so pick fields by what your application consumes. If you do need evidences, measure the selection against your own traffic rather than assuming. suggested_response adds a polite, service-aware reply for the non-answer classes and null otherwise, and follows the language of the user message.

Performance

Eight external benchmarks, binary safe/unsafe scoring, Out of Scope counted as safe, greedy decoding with guided JSON. Means are over the six benchmarks that contain both classes.

benchmark scope_class only reasoning, scope_class
CoSApien 0.937 0.922
ToxicChat 0124 0.911 0.903
DynaGuardrail 0.903 0.912
WildGuardMix test 0.874 0.876
Aegis 2.0 0.809 0.803
OpenAI Moderation 0.792 0.807
mean balanced accuracy 0.871 0.871
mean unsafe F1 0.855 0.854
HarmBench text (unsafe recall) 0.984 0.994
SimpleSafetyTests (unsafe recall) 0.990 0.990

Latency on one 32 GB GPU with vLLM, MTP speculation, bfloat16: roughly 2,000 to 3,000 prompt tokens and 60 to 100 output tokens per request; throughput on the order of a few requests per second at 16 concurrent sequences. scope_class-only output is several times faster.

Limitations

  • The description is the policy. The model applies what you wrote. A vague or permissive description gives permissive classifications.
  • Keyword-triggered over-refusal on violent-sounding but benign phrasings ("kill a process", "euthanize the test suite") persists on a minority of such messages.
  • Predefined Answer and Human Oversight are exercised far less by public benchmarks than the other classes, so their accuracy is less well characterised. Validate against your own escalation rules.
  • Borderline scope decisions, especially Out of Scope versus Potentially Supported and Out of Scope versus Restricted, are the hardest cases and where most residual errors sit.
  • English first. Training data is predominantly English with some multilingual rows. suggested_response follows the user's language; classification quality in other languages is not characterised.
  • Classifier, not assistant. It was trained on corpora that contain harmful prompts in order to recognise them. Do not use it to generate content; its outputs beyond the JSON fields are undefined.

Using the model without orbitals

If you cannot use the library, reproduce its prompt exactly. Three things must match or the model degrades without any error.

  1. System prompt: the file system_prompt.txt in this repo, verbatim (sha256 f0f68e48…, 7,677 bytes including the trailing newline).

  2. User turn, three blocks in this order:

    **START OF THE AI SERVICE DESCRIPTION**
    
    <your service description>
    
    **END OF THE AI SERVICE DESCRIPTION**
    
    
    **START OF THE CONVERSATION DUMP**
    
    USER:
    <earlier user message>
    
    ASSISTANT:
    <earlier assistant message>
    
    LAST MESSAGE (USER):
    <the message to classify>
    
    **END OF THE CONVERSATION DUMP**
    
    
    **REQUESTED OUTPUT FIELDS**
    
    ["reasoning", "scope_class"]
    

    Earlier turns are ROLE: followed by the content and a blank line; the final turn is tagged LAST MESSAGE (ROLE):. The last block is a JSON array of the keys you want back, in this fixed order: evidences, reasoning, scope_class, suggested_response. scope_class is always required.

  3. Thinking off, greedy, structured output: apply the chat template with enable_thinking=False (with vLLM's OpenAI server, pass chat_template_kwargs: {"enable_thinking": false}), temperature=0, and a JSON schema with exactly the requested keys, scope_class as an enum of the seven class names, and no additional properties. The template then opens the assistant turn with an empty <think></think> block, which is what the model saw in training; at the default setting it starts a reasoning trace instead of the JSON.

Training

Qwen3.5-4B, LoRA rank 32 over all 32 layers including the 24 linear-attention layers, trained for three epochs on 84,657 rows from the ScopeGuard v2.3 corpora, with the requested output fields resampled independently per row per epoch so every field combination is trained. The released weights are the uniform average of four checkpoints from that run (steps 8000, 10600, 13200 and 15873), merged into the base model; averaging gave a small, consistent gain over the best single checkpoint on the in-domain panel and matched it externally. The base model's multi-token-prediction head is included unchanged for speculative decoding.

The system prompt in this repo is the one used in training. It defines Potentially Supported by a checkable property and makes the class unavailable whenever a rule matches, which removed the largest error cluster of the previous generation.

Access and license

Access is granted on request through the form above. Weights are released under the terms in LICENSE. Contact: edoardo@principled-intelligence.com.

Downloads last month
135
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for principled-intelligence/scope-guard-v2-4B-q-2609

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(253)
this model