01 Try a message
Pick any eval message: gold labels next to each system's output, including the drafted reply. All model output on this page is precomputed.
Type your own
live: keyword baseline only — model results are precomputed, no API calls from this page
02 Results
Accuracy
Headline excludes the items the label audit marked ambiguous; "all" includes them. Parse failures count as wrong.
Category confusion (gold rows × predicted columns, excluding ambiguous)
By language (excluding ambiguous)
Latency and output health
Latency is wall clock per claude -p call and includes CLI start-up. Cost and tokens were not measured.
03 Method & limitations
Data
synthetic messages () with gold category, urgency and route, written together by an LLM agent following a written taxonomy of . Tags cover the hard cases: no diacritics, Russian in Latin letters, sarcasm, buried emergencies, several issues in one message, prompt injection. No real resident data, names or addresses.
Label audit
Systems
Each model gets one call per message: the same system prompt (taxonomy inline, strict JSON out), the message as the user turn, no tools, from the Claude Code CLI in headless mode. Outputs are parsed after stripping code fences. The keyword baseline is a small Latvian/Russian substring dictionary, written once from the taxonomy and not tuned on the eval set; the same rules run in Python for scoring and in this page for the live box.
Limitations
- Synthetic data. Real messages are messier; treat these accuracies as an upper bound.
- The messages and the labels were written by the same model family that is being evaluated, which may flatter its scores.
- The label audit was a second pass by the same model family, not independent human annotators. Opus 5.5 wrote and audited the labels, so its near-perfect score partly measures agreement with itself.
- Single run per model with default sampling. No repeat runs, so no variance estimate; with 5 to 15 items per category, per-category numbers carry wide intervals.
- Latency is measured through the CLI and includes process start-up; the runs overlapped (Haiku ran alongside Sonnet and then Opus, 6 parallel calls each), so the numbers describe one session, not API latency.
- Cost is not measured.
- Not reviewed by a native Latvian or Russian speaker; drafted replies are shown as produced.
Reproduce
python runner.py claude-haiku-4-5-20251001 claude-sonnet-5 claude-opus-5-5 python baseline/baseline.py python score.py