typed-decisions-minilm-l6-specialist

A specialist for the Typed Decisions benchmark: one AdaptiveClassifier head per (workflow, question) on a frozen sentence-transformers/all-MiniLM-L6-v2 encoder, 20 heads in total. Each head takes the rendered state, the question text and its options, and returns a probability distribution over the question's labels.

This is not zero-shot. It was fitted on the 1,200-case train split of Typed Decisions and is scored on the 400-case test split. Per the benchmark's own rules it is a specialist, and should not be compared with generalist models that answer unseen question schemas.

Results (test split, 400 cases, 2,000 decisions)

Metric This run Benchmark README
Accuracy 0.607 0.587
KL from gold (lower is better) 0.249 0.262
Brier (lower is better) 0.136 0.143
ECE (lower is better) 0.136 0.108

"This run" is the model in this repository, fitted and scored from the public dataset with adaptive-classifier 0.2.0 on Apple MPS, and scored by the benchmark maintainers on the test split. Full per-question numbers are in results.json.

Recipe

  • Encoder sentence-transformers/all-MiniLM-L6-v2, frozen, max_length 512, mean pooling.
  • One AdaptiveClassifier per (workflow, question), prototype weight 0.3, 30 epochs, seed 42.
  • score questions are treated as ordinal classes and decoded to an expected score.
  • Each training case is entered 4 times, split across labels in proportion to the gold distribution (soft replicas).
  • Hyperparameters were tuned on a held-out slice of train, never on test.

Usage

# pip install adaptive-classifier==0.2.0 torch transformers
from predict import Specialist
m = Specialist("."); print(m.predict(workflow, state, questions))   # questions in the benchmark's wire format

python3 predict.py case.json does the same from a JSON file. python3 predict.py --test-ref checks that the reloaded heads reproduce the saved predictions for the first 20 test cases. The encoder is downloaded from the Hub on first load and shared by all heads.

Limits

  • It only answers the question schemas of the four workflows it was trained on (agent_trace_observability, customer_service, invoice_processing, security_incidents). Any other question gets a uniform answer.
  • It learns the teacher's labels. The gold is the mean of three samples from a roughly 4B-class teacher, so a score measures agreement with that teacher, not correctness (see the benchmark README).
  • Synthetic data only; no claim is made about real-world tickets, invoices or incidents.

Links

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for adaptive-classifier/typed-decisions-minilm-l6-specialist

Dataset used to train adaptive-classifier/typed-decisions-minilm-l6-specialist

Evaluation results

  • LocalLLaMA/typed-decisions
  • Accuracy View evaluation results source
    Fine-tuned on the benchmark's train split (specialist per workflow), not zero-shot. Scored by the benchmark maintainers on the test split.
    0.61 *
  • Kl From Gold View evaluation results source
    Fine-tuned on the benchmark's train split (specialist per workflow), not zero-shot. Scored by the benchmark maintainers on the test split.
    0.25 *
  • Brier View evaluation results source
    Fine-tuned on the benchmark's train split (specialist per workflow), not zero-shot. Scored by the benchmark maintainers on the test split.
    0.14 *