typed-decisions-minilm-l6-specialist
A specialist for the Typed Decisions
benchmark: one AdaptiveClassifier head per
(workflow, question) on a frozen sentence-transformers/all-MiniLM-L6-v2 encoder, 20 heads in total. Each head takes the rendered
state, the question text and its options, and returns a probability distribution over the question's labels.
This is not zero-shot. It was fitted on the 1,200-case train split of Typed Decisions and is scored on the
400-case test split. Per the benchmark's own rules it is a specialist, and should not be compared with
generalist models that answer unseen question schemas.
Results (test split, 400 cases, 2,000 decisions)
| Metric | This run | Benchmark README |
|---|---|---|
| Accuracy | 0.607 | 0.587 |
| KL from gold (lower is better) | 0.249 | 0.262 |
| Brier (lower is better) | 0.136 | 0.143 |
| ECE (lower is better) | 0.136 | 0.108 |
"This run" is the model in this repository, fitted and scored from the public dataset with
adaptive-classifier 0.2.0 on Apple MPS, and scored by the benchmark maintainers on the test split.
Full per-question numbers are in results.json.
Recipe
- Encoder
sentence-transformers/all-MiniLM-L6-v2, frozen,max_length512, mean pooling. - One AdaptiveClassifier per (workflow, question), prototype weight 0.3, 30 epochs, seed 42.
scorequestions are treated as ordinal classes and decoded to an expected score.- Each training case is entered 4 times, split across labels in proportion to the gold distribution (soft replicas).
- Hyperparameters were tuned on a held-out slice of
train, never ontest.
Usage
# pip install adaptive-classifier==0.2.0 torch transformers
from predict import Specialist
m = Specialist("."); print(m.predict(workflow, state, questions)) # questions in the benchmark's wire format
python3 predict.py case.json does the same from a JSON file. python3 predict.py --test-ref checks that the
reloaded heads reproduce the saved predictions for the first 20 test cases. The encoder is downloaded from the Hub
on first load and shared by all heads.
Limits
- It only answers the question schemas of the four workflows it was trained on
(
agent_trace_observability,customer_service,invoice_processing,security_incidents). Any other question gets a uniform answer. - It learns the teacher's labels. The gold is the mean of three samples from a roughly 4B-class teacher, so a score measures agreement with that teacher, not correctness (see the benchmark README).
- Synthetic data only; no claim is made about real-world tickets, invoices or incidents.
Links
Model tree for adaptive-classifier/typed-decisions-minilm-l6-specialist
Base model
nreimers/MiniLM-L6-H384-uncasedDataset used to train adaptive-classifier/typed-decisions-minilm-l6-specialist
Evaluation results
- LocalLLaMA/typed-decisions
- Accuracy View evaluation results sourceFine-tuned on the benchmark's train split (specialist per workflow), not zero-shot. Scored by the benchmark maintainers on the test split.0.61 *
- Kl From Gold View evaluation results sourceFine-tuned on the benchmark's train split (specialist per workflow), not zero-shot. Scored by the benchmark maintainers on the test split.0.25 *
- Brier View evaluation results sourceFine-tuned on the benchmark's train split (specialist per workflow), not zero-shot. Scored by the benchmark maintainers on the test split.0.14 *