Agent Tool Effect Classifier (DeBERTa-v3-Large)

Fine-tuned microsoft/deberta-v3-large (435M parameters) for ahead-of-time semantic classification of agent tool effect: READ (non-mutating, cacheable) vs. WRITE (state-mutating barrier).

Part of the Agent Runtime Optimization Middleware architecture.

Model Details

  • Base Architecture: microsoft/deberta-v3-large (24 layers, 1024 hidden dimension, 16 attention heads)
  • Training Corpus: 20,056 real-world multi-domain tool definitions from BFCL (Berkeley Function-Calling Leaderboard), ToolBench, AppWorld, SWE-bench, and Seal-Tools.
  • Platt Temperature Calibration: $T = 1.4645$ (reduces Expected Calibration Error from 24.8% to 1.6%).
  • Safe-Fail Decision Gate: $P(\text{READ}) \ge 0.95$. Borderline reads ($P < 0.95$) are forced to WRITE to prevent catastrophic cache corruption.

Usage

import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification

REPO_ID = "DamonRicci/agent-tool-effect-deberta-v3-large"
tokenizer = AutoTokenizer.from_pretrained(REPO_ID)
model = AutoModelForSequenceClassification.from_pretrained(REPO_ID)

text = "Tool: get_user_profile\nDomain: auth\nParams: ['user_id']\nDesc: Retrieves the profile record for the given user ID."
inputs = tokenizer(text, return_tensors="pt", max_length=128, truncation=True)

with torch.no_grad():
    logits = model(**inputs).logits
    # Apply Platt Temperature Calibration
    probs = torch.softmax(logits / 1.4645, dim=1)[0]
    p_read = probs[0].item()
    p_write = probs[1].item()

# Apply Asymmetric Safe-Fail Gate
is_read = p_read >= 0.95
verdict = "READ" if is_read else "WRITE"
print(f"Verdict: {verdict} (P(READ)={p_read:.4f})")
Downloads last month
22
Safetensors
Model size
0.4B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support