Text Classification
Transformers
Safetensors
English
qwen3_5_text
text-generation
system-one
system-two
blocks-of-experts
typed-decisions
decision-model
calibrated-probabilities
knowledge-distillation
jev
noul
choice
score
lora
qwen3_5
dual-head
vllm
Eval Results (legacy)
Instructions to use autotrust/JEV-27B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use autotrust/JEV-27B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="autotrust/JEV-27B")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("autotrust/JEV-27B") model = AutoModelForCausalLM.from_pretrained("autotrust/JEV-27B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Model card: "New (1 October 2026)" summary at the top
Browse files
README.md
CHANGED
|
@@ -92,6 +92,19 @@ stronger System 2, and it answers a single decision in about half the time the h
|
|
| 92 |
<sub><b>Scorecard.</b> System 1 fidelity, calibration and transfer, System 2 and speed; grey tags show JEV-9B · click to enlarge</sub>
|
| 93 |
</p>
|
| 94 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 95 |
## Public decision benchmarks
|
| 96 |
|
| 97 |
Six public text-decision benchmarks. AutoTrust ran them in full for autotrust/JEV-27B and for the hosted TypeSafe
|
|
|
|
| 92 |
<sub><b>Scorecard.</b> System 1 fidelity, calibration and transfer, System 2 and speed; grey tags show JEV-9B · click to enlarge</sub>
|
| 93 |
</p>
|
| 94 |
|
| 95 |
+
## New (1 October 2026)
|
| 96 |
+
|
| 97 |
+
* **System 1 over plain HTTP.** `serve_decide.py` (in this repository) adds `POST /v1/decide` to the vLLM server: send
|
| 98 |
+
`{kind, state, question, options}` and get a calibrated probability for every option. See
|
| 99 |
+
[Quickstart](#quickstart-with-vllm-recommended).
|
| 100 |
+
* **Up to 256 options per choice question**, with no retraining. Zero-shot on CLINC150 with all 150 intents as options:
|
| 101 |
+
89.5% with intent names alone, 93.8% with a one-line description per option. See
|
| 102 |
+
[Choice questions with up to 256 options](#choice-questions-with-up-to-256-options).
|
| 103 |
+
* **How to write prompts, measured.** One-line "use when" descriptions for similar options help; question wording, JSON
|
| 104 |
+
vs plain text and extra instructions make no measurable difference. See [Writing System 1 prompts](#writing-system-1-prompts).
|
| 105 |
+
* **256K-token prompts.** Decisions that hinge on one sentence hidden at a random depth in up to 250K tokens of text:
|
| 106 |
+
20 of 20 correct at every length tested. See [Context length](#context-length).
|
| 107 |
+
|
| 108 |
## Public decision benchmarks
|
| 109 |
|
| 110 |
Six public text-decision benchmarks. AutoTrust ran them in full for autotrust/JEV-27B and for the hosted TypeSafe
|