Tau Workbench report
Synthetic support tickets: how urgent is this ticket?
Synthetic data Every item in this dataset was generated, not written by real customers. Treat the numbers as a demonstration of the method, not as evidence about real traffic.
This report measures how far 4 model(s) can be trusted on "How urgent is this customer support ticket?", using synthetic tickets data (1,000 held-out items). Reference: claude-opus-5-5's answers, not the dataset's labels. The question here is whether a local model can stand in for the frontier call, and whether its confidence says when, so every rate below is agreement with the frontier model. Out of the box, laya-en agrees with claude-opus-5-5 on 23.2% of held-out items with a calibration error (ECE) of 0.273; after calibration its ECE is 0.148, 46.0% lower, short of the 50% goal. At a 20.0% target disagreement, laya-en keeps 0.1% of decisions local (3 models tie at that share); the cascade's served answers agree with claude-opus-5-5 on 100.0% (the frontier alone agrees with itself by construction), at an estimated £996 per million decisions against £997 (list prices, 2026-09-27). The dataset's own labels are reported as a finding, not used: claude-opus-5-5 agrees with them on only 23.8% of 1,000 items, and a secondary table shows every model against them.
Generated 2026-09-27T17:15:20Z (UTC) on NVIDIA GeForce RTX 3080 Ti. Every number below is in report.json beside this file; the metadata block at the end says how to reproduce it.
Dataset
- Name
- tickets
- Source
- Tobi-Bueck/customer-support-tickets
- Pinned revision
- ddf1c81a5475992c4fa6752bf1e8b4e31f07bbeb
- Licence
- CC-BY-NC-4.0
- Synthetic
- Yes: generated data
- Fine-tune split
- 8,000 items
- Calibration split
- 1,000 items
- Held-out split
- 1,000 items
- Question
- urgency (score, 5 answers): How urgent is this customer support ticket?
Calibrators and thresholds are fitted on the calibration split only and every result is judged on the held-out split, which nothing was tuned on. The splits are fixed by the dataset manifest (seed 42), and each file's sha256 is checked before use.
Agreement with the frontier model and calibration, before and after
| Model | Items | Agreement raw | Agreement calibrated | ECE raw | ECE calibrated | ECE change | Brier raw → cal. | Log loss raw → cal. | MAE (levels) raw → cal. | Calibrated view |
|---|---|---|---|---|---|---|---|---|---|---|
| laya-en | 1,000 | 23.2% | 37.2% | 0.273 | 0.148 | −46.0% | 0.952 → 0.757 | 2.308 → 1.601 | 1.19 → 0.83 | Runtime with calibrators loaded; matches the Workbench's offline calibration to within 5.5e-2 in any probability |
| laya-typed-decisions | 1,000 | 18.5% | 26.3% | 0.269 | 0.067 | −75.0% | 0.895 → 0.777 | 1.908 → 1.546 | 1.28 → 1.03 | Runtime with calibrators loaded; matches the Workbench's offline calibration to within 7.3e-2 in any probability |
| von-1.2.0 | 1,000 | 43.9% | 43.8% | 0.045 | 0.040 | −12.1% | 0.703 → 0.701 | 1.438 → 1.452 | 0.88 → 0.88 | Runtime with calibrators loaded; matches the Workbench's offline calibration to within 4.6e-2 in any probability |
| laya-en-ft-tickets | 1,000 | 27.7% | 35.8% | 0.255 | 0.074 | −71.1% | 0.926 → 0.771 | 2.476 → 1.610 | 1.08 → 0.95 | Runtime with calibrators loaded; matches the Workbench's offline calibration to within 5.8e-2 in any probability |
Held-out items only. ECE (expected calibration error) is the average gap between how confident a model is and how often it is right, over 15 confidence bins; lower is better and 0 is perfect. Confidence for ECE and thresholds is max(p), the probability of the chosen answer, not the contract's confidence field, which for choice and score questions is not a calibrated probability. Here a model is right when its answer matches claude-opus-5-5's, so the agreement columns are agreement with the frontier model, not accuracy, and ECE measures whether a model's confidence says when it agrees. The goal is a 50% cut in ECE. The largest cut is laya-typed-decisions's (75.0%) and the smallest is von-1.2.0's (12.1%); a model whose confidence is still off after calibration should not be trusted to gate decisions on its own.
Reliability diagrams
Each point is one confidence bin on the held-out split: how confident the model was (across) against how often it was right (up), where right means agreeing with the frontier model. Points below the diagonal are overconfident, points above it underconfident. After calibration the points should sit close to the diagonal; where they don't, a threshold on that model's confidence will not deliver the agreement it promises.
Table view: every bin
| Model | Bin | Raw items | Raw confidence | Raw agreement | Cal. items | Cal. confidence | Cal. agreement |
|---|---|---|---|---|---|---|---|
| laya-en | 0.13–0.20 | 0 | n/a | n/a | 105 | 0.200 | 0.648 |
| laya-en | 0.20–0.27 | 0 | n/a | n/a | 304 | 0.233 | 0.464 |
| laya-en | 0.27–0.33 | 0 | n/a | n/a | 395 | 0.301 | 0.263 |
| laya-en | 0.33–0.40 | 77 | 0.380 | 0.351 | 140 | 0.351 | 0.307 |
| laya-en | 0.40–0.47 | 375 | 0.438 | 0.224 | 52 | 0.429 | 0.269 |
| laya-en | 0.47–0.53 | 259 | 0.495 | 0.158 | 2 | 0.497 | 0.500 |
| laya-en | 0.53–0.60 | 115 | 0.560 | 0.243 | 1 | 0.556 | 0.000 |
| laya-en | 0.60–0.67 | 91 | 0.632 | 0.253 | 1 | 0.650 | 1.000 |
| laya-en | 0.67–0.73 | 47 | 0.695 | 0.340 | 0 | n/a | n/a |
| laya-en | 0.73–0.80 | 23 | 0.756 | 0.261 | 0 | n/a | n/a |
| laya-en | 0.80–0.87 | 10 | 0.828 | 0.400 | 0 | n/a | n/a |
| laya-en | 0.87–0.93 | 3 | 0.908 | 1.000 | 0 | n/a | n/a |
| laya-typed-decisions | 0.20–0.27 | 0 | n/a | n/a | 560 | 0.229 | 0.196 |
| laya-typed-decisions | 0.27–0.33 | 7 | 0.323 | 0.571 | 285 | 0.279 | 0.393 |
| laya-typed-decisions | 0.33–0.40 | 225 | 0.378 | 0.213 | 154 | 0.362 | 0.260 |
| laya-typed-decisions | 0.40–0.47 | 407 | 0.431 | 0.140 | 1 | 0.453 | 1.000 |
| laya-typed-decisions | 0.47–0.53 | 241 | 0.498 | 0.174 | 0 | n/a | n/a |
| laya-typed-decisions | 0.53–0.60 | 105 | 0.559 | 0.286 | 0 | n/a | n/a |
| laya-typed-decisions | 0.60–0.67 | 15 | 0.623 | 0.267 | 0 | n/a | n/a |
| von-1.2.0 | 0.20–0.27 | 0 | n/a | n/a | 2 | 0.259 | 0.500 |
| von-1.2.0 | 0.27–0.33 | 43 | 0.315 | 0.256 | 68 | 0.299 | 0.353 |
| von-1.2.0 | 0.33–0.40 | 206 | 0.373 | 0.296 | 321 | 0.358 | 0.312 |
| von-1.2.0 | 0.40–0.47 | 265 | 0.431 | 0.385 | 128 | 0.444 | 0.406 |
| von-1.2.0 | 0.47–0.53 | 201 | 0.499 | 0.463 | 136 | 0.498 | 0.456 |
| von-1.2.0 | 0.53–0.60 | 147 | 0.562 | 0.571 | 215 | 0.566 | 0.535 |
| von-1.2.0 | 0.60–0.67 | 78 | 0.627 | 0.615 | 105 | 0.630 | 0.648 |
| von-1.2.0 | 0.67–0.73 | 44 | 0.696 | 0.636 | 20 | 0.687 | 0.650 |
| von-1.2.0 | 0.73–0.80 | 9 | 0.767 | 0.778 | 4 | 0.767 | 0.750 |
| von-1.2.0 | 0.80–0.87 | 3 | 0.817 | 1.000 | 1 | 0.852 | 0.000 |
| von-1.2.0 | 0.87–0.93 | 4 | 0.904 | 0.500 | 0 | n/a | n/a |
| laya-en-ft-tickets | 0.13–0.20 | 0 | n/a | n/a | 9 | 0.200 | 0.556 |
| laya-en-ft-tickets | 0.20–0.27 | 0 | n/a | n/a | 413 | 0.242 | 0.370 |
| laya-en-ft-tickets | 0.27–0.33 | 0 | n/a | n/a | 181 | 0.289 | 0.276 |
| laya-en-ft-tickets | 0.33–0.40 | 12 | 0.392 | 0.167 | 240 | 0.357 | 0.379 |
| laya-en-ft-tickets | 0.40–0.47 | 243 | 0.442 | 0.185 | 147 | 0.420 | 0.361 |
| laya-en-ft-tickets | 0.47–0.53 | 374 | 0.496 | 0.235 | 9 | 0.480 | 0.556 |
| laya-en-ft-tickets | 0.53–0.60 | 194 | 0.560 | 0.392 | 1 | 0.583 | 1.000 |
| laya-en-ft-tickets | 0.60–0.67 | 76 | 0.630 | 0.289 | 0 | n/a | n/a |
| laya-en-ft-tickets | 0.67–0.73 | 37 | 0.698 | 0.405 | 0 | n/a | n/a |
| laya-en-ft-tickets | 0.73–0.80 | 30 | 0.766 | 0.500 | 0 | n/a | n/a |
| laya-en-ft-tickets | 0.80–0.87 | 21 | 0.829 | 0.381 | 0 | n/a | n/a |
| laya-en-ft-tickets | 0.87–0.93 | 10 | 0.894 | 0.400 | 0 | n/a | n/a |
| laya-en-ft-tickets | 0.93–1.00 | 3 | 0.948 | 0.667 | 0 | n/a | n/a |
Calibrators fitted
| Model | Scope | Items | Chosen | Temperature T | Temperature ECE / log loss | Isotonic knots | Isotonic ECE / log loss | Before: ECE / log loss |
|---|---|---|---|---|---|---|---|---|
| laya-en | question type | 1,000 | isotonic | 7.618 | 0.016 / 1.581 | 32 | 0.201 / 1.422 | 0.244 / 2.306 |
| laya-en | 3-5 options | 1,000 | isotonic | 7.618 | 0.016 / 1.581 | 32 | 0.201 / 1.422 | 0.244 / 2.306 |
| laya-typed-decisions | question type | 1,000 | isotonic | 5.709 | 0.054 / 1.587 | 30 | 0.094 / 1.499 | 0.244 / 1.922 |
| laya-typed-decisions | 3-5 options | 1,000 | isotonic | 5.709 | 0.054 / 1.587 | 30 | 0.094 / 1.499 | 0.244 / 1.922 |
| von-1.2.0 | question type | 1,000 | isotonic | 1.152 | 0.055 / 1.331 | 67 | 0.035 / 1.303 | 0.052 / 1.336 |
| von-1.2.0 | 3-5 options | 1,000 | isotonic | 1.152 | 0.055 / 1.331 | 67 | 0.035 / 1.303 | 0.052 / 1.336 |
| laya-en-ft-tickets | question type | 1,000 | isotonic | 12.836 | 0.055 / 1.598 | 37 | 0.082 / 1.512 | 0.250 / 2.549 |
| laya-en-ft-tickets | 3-5 options | 1,000 | isotonic | 12.836 | 0.055 / 1.598 | 37 | 0.082 / 1.512 | 0.250 / 2.549 |
Both methods are fitted on the calibration split's raw probabilities, and the one with the lower calibration-split log loss is written for the Runtime to load; the scores in this table are on the calibration split, so they show the fit, not the result (the held-out result is in the first table). Isotonic regression needs at least 200 items and falls back to temperature scaling below that.
- laya-en: A spec asks one question, so every item has 5 options and the 3-5 bucket calibrator is fitted on the same items as the question-type one. The Runtime prefers the bucket file for 5-option requests; the type-level file covers other option counts.
- laya-typed-decisions: A spec asks one question, so every item has 5 options and the 3-5 bucket calibrator is fitted on the same items as the question-type one. The Runtime prefers the bucket file for 5-option requests; the type-level file covers other option counts.
- von-1.2.0: A spec asks one question, so every item has 5 options and the 3-5 bucket calibrator is fitted on the same items as the question-type one. The Runtime prefers the bucket file for 5-option requests; the type-level file covers other option counts.
- laya-en-ft-tickets: A spec asks one question, so every item has 5 options and the 3-5 bucket calibrator is fitted on the same items as the question-type one. The Runtime prefers the bucket file for 5-option requests; the type-level file covers other option counts.
How much can stay local
Each line shows, on held-out items, what happens as the confidence threshold τ rises: fewer decisions are kept local (moving left) and those kept are more often right (moving up), where right means agreeing with the frontier model. The dot marks the τ chosen on the calibration split for a 20.0% target disagreement. A line that never reaches the target line means no threshold makes that model safe enough on its own at that target. Thresholds that keep fewer than 10 items are left off the chart because a handful of items says little; every point is in report.json. Dashed lines are baselines, not served through Tau, put through exactly the same threshold rule on their own calibration lines.
| Model | Confidences from | τ | Kept local (calibration) | Kept local (held-out) | Agreement on kept (held-out) | Result |
|---|---|---|---|---|---|---|
| laya-en | offline | 0.56 | 0.1% | 0.1% | 100.0% | τ = 0.56 is the smallest threshold whose calibration-split disagreement is at most 20.0%. On held-out it accepts 0.1% of items with 100.0% agreement with the frontier model. |
| laya-typed-decisions | offline | 0.39 | 0.1% | 0.1% | 100.0% | τ = 0.39 is the smallest threshold whose calibration-split disagreement is at most 20.0%. On held-out it accepts 0.1% of items with 100.0% agreement with the frontier model. |
| von-1.2.0 | offline | 0.78 | 0.3% | 0.1% | 0.0% | τ = 0.78 is the smallest threshold whose calibration-split disagreement is at most 20.0%. On held-out it accepts 0.1% of items with 0.0% agreement with the frontier model. |
| laya-en-ft-tickets | offline | n/a | n/a | n/a | n/a | No threshold meets the 20.0% target disagreement on the calibration split. The lowest disagreement any threshold reaches is 43.8%, at τ = 0.45, accepting 1.6% of items. No threshold is invented. |
| minilm-l6-tickets (baseline) | baseline-calibrated | 0.44 | 0.2% | 0.4% | 25.0% | τ = 0.44 is the smallest threshold whose calibration-split disagreement is at most 20.0%. On held-out it accepts 0.4% of items with 25.0% agreement with the frontier model. |
Cascade: local first, frontier for the rest
| Model | τ | Kept local | Escalated | No frontier answer | Local only (agreement) | Frontier only | Cascade (agreement) | Local p50 latency |
|---|---|---|---|---|---|---|---|---|
| laya-en | 0.56 | 0.1% | 99.9% | 0 | 37.2% (1,000) | 100% by construction | 100.0% (1,000) | 53.2 ms |
| laya-typed-decisions | 0.39 | 0.1% | 99.9% | 0 | 26.3% (1,000) | 100% by construction | 100.0% (1,000) | 53.5 ms |
| von-1.2.0 | 0.78 | 0.1% | 99.9% | 0 | 43.8% (1,000) | 100% by construction | 99.9% (1,000) | 45.0 ms |
| laya-en-ft-tickets | n/a | n/a | n/a | n/a | 35.3% (1,000) | 100% by construction | No cascade: no threshold met the 20.0% target disagreement on the calibration split. | 50.8 ms |
| minilm-l6-tickets (baseline, not served through Tau: local latency and energy not measured) | 0.44 | 0.4% | 99.6% | 0 | 27.2% (1,000) | 100% by construction | 99.7% (1,000) | not measured |
Every rate here is agreement with the frontier model on held-out items, not accuracy, with the number of items it covers in brackets. The cascade serves the local answer when its confidence is at least τ and the frontier model's cached answer otherwise; the cascade figure is the share of items where the served answer equals the frontier model's. Frontier only agrees with the frontier model on 100% of items by construction: its answers are the reference, so this is not a result. The closer the cascade gets to 100% while keeping a large share local, the better the local model stands in for the frontier call. Frontier latency is not measured: no API call was made, so the latency mix covers the local side only.
Cost for laya-en Estimates, not measured bills
| Price basis | Frontier only, per million (USD) | Frontier only, per 1,000 | Frontier only, per million | Cascade, per 1,000 | Cascade, per million | Saving, per million |
|---|---|---|---|---|---|---|
| claude-opus-5-5 at list price (the model that produced the frontier answers) | $1,321 | £0.9971 | £997 | £0.9964 | £996 | £0.6549 |
| claude-opus-5-5-batch: the same model at a different rate | $661 | £0.4985 | £499 | £0.4984 | £498 | £0.1564 |
| claude-sonnet-5: price only, if escalations went to this model; its accuracy was not measured | $661 | £0.4985 | £499 | £0.4984 | £498 | £0.1564 |
| claude-haiku-4-5: price only, if escalations went to this model; its accuracy was not measured | $330 | £0.2493 | £249 | £0.2494 | £249 | £-0.0929 |
Every figure in this table is an estimate from list prices and estimated token counts, not a measured bill. It shows what the same decisions would cost through the API, and what share of that the cascade avoids. Basis:
- Frontier tokens are estimated, not counted: characters / 4 (Anthropic's rule of thumb) × 1.3 for the tokenizer, averaged over 1,000 cached answers: 328.7 input and 0.3 output tokens per decision.
- List prices in US dollars per million tokens, checked 2026-09-27 (platform.claude.com/docs/en/about-claude/pricing (checked 2026-09-27)).
- USD to GBP at 0.75458 (ECB euro foreign exchange reference rates, 2026-09-25).
- Frontier answers for this benchmark came from an interactive Claude Code session, not the API; the costs price what the same work would cost through the API.
- Local energy: 345.3 W mean whole-GPU power × 0.01355 s per decision (phase duration / items) = 1.300e-6 kWh per decision, at £0.2632 per kWh (www.ofgem.gov.uk/news/changes-energy-price-cap-between-1-october-and-31-december-2026). Hardware purchase and depreciation are excluded.
- Cascade: every decision runs locally, and 99.9% are also sent to the frontier model.
Cost for laya-typed-decisions Estimates, not measured bills
| Price basis | Frontier only, per million (USD) | Frontier only, per 1,000 | Frontier only, per million | Cascade, per 1,000 | Cascade, per million | Saving, per million |
|---|---|---|---|---|---|---|
| claude-opus-5-5 at list price (the model that produced the frontier answers) | $1,321 | £0.9971 | £997 | £0.9964 | £996 | £0.6509 |
| claude-opus-5-5-batch: the same model at a different rate | $661 | £0.4985 | £499 | £0.4984 | £498 | £0.1523 |
| claude-sonnet-5: price only, if escalations went to this model; its accuracy was not measured | $661 | £0.4985 | £499 | £0.4984 | £498 | £0.1523 |
| claude-haiku-4-5: price only, if escalations went to this model; its accuracy was not measured | $330 | £0.2493 | £249 | £0.2494 | £249 | £-0.0969 |
Every figure in this table is an estimate from list prices and estimated token counts, not a measured bill. It shows what the same decisions would cost through the API, and what share of that the cascade avoids. Basis:
- Frontier tokens are estimated, not counted: characters / 4 (Anthropic's rule of thumb) × 1.3 for the tokenizer, averaged over 1,000 cached answers: 328.7 input and 0.3 output tokens per decision.
- List prices in US dollars per million tokens, checked 2026-09-27 (platform.claude.com/docs/en/about-claude/pricing (checked 2026-09-27)).
- USD to GBP at 0.75458 (ECB euro foreign exchange reference rates, 2026-09-25).
- Frontier answers for this benchmark came from an interactive Claude Code session, not the API; the costs price what the same work would cost through the API.
- Local energy: 345.1 W mean whole-GPU power × 0.01372 s per decision (phase duration / items) = 1.315e-6 kWh per decision, at £0.2632 per kWh (www.ofgem.gov.uk/news/changes-energy-price-cap-between-1-october-and-31-december-2026). Hardware purchase and depreciation are excluded.
- Cascade: every decision runs locally, and 99.9% are also sent to the frontier model.
Cost for von-1.2.0 Estimates, not measured bills
| Price basis | Frontier only, per million (USD) | Frontier only, per 1,000 | Frontier only, per million | Cascade, per 1,000 | Cascade, per million | Saving, per million |
|---|---|---|---|---|---|---|
| claude-opus-5-5 at list price (the model that produced the frontier answers) | $1,321 | £0.9971 | £997 | £0.9963 | £996 | £0.7079 |
| claude-opus-5-5-batch: the same model at a different rate | $661 | £0.4985 | £499 | £0.4983 | £498 | £0.2093 |
| claude-sonnet-5: price only, if escalations went to this model; its accuracy was not measured | $661 | £0.4985 | £499 | £0.4983 | £498 | £0.2093 |
| claude-haiku-4-5: price only, if escalations went to this model; its accuracy was not measured | $330 | £0.2493 | £249 | £0.2493 | £249 | £-0.0399 |
Every figure in this table is an estimate from list prices and estimated token counts, not a measured bill. It shows what the same decisions would cost through the API, and what share of that the cascade avoids. Basis:
- Frontier tokens are estimated, not counted: characters / 4 (Anthropic's rule of thumb) × 1.3 for the tokenizer, averaged over 1,000 cached answers: 328.7 input and 0.3 output tokens per decision.
- List prices in US dollars per million tokens, checked 2026-09-27 (platform.claude.com/docs/en/about-claude/pricing (checked 2026-09-27)).
- USD to GBP at 0.75458 (ECB euro foreign exchange reference rates, 2026-09-25).
- Frontier answers for this benchmark came from an interactive Claude Code session, not the API; the costs price what the same work would cost through the API.
- Local energy: 346.4 W mean whole-GPU power × 0.01142 s per decision (phase duration / items) = 1.099e-6 kWh per decision, at £0.2632 per kWh (www.ofgem.gov.uk/news/changes-energy-price-cap-between-1-october-and-31-december-2026). Hardware purchase and depreciation are excluded.
- Cascade: every decision runs locally, and 99.9% are also sent to the frontier model.
Cost for laya-en-ft-tickets Estimates, not measured bills
| Price basis | Frontier only, per million (USD) | Frontier only, per 1,000 | Frontier only, per million | Cascade, per 1,000 | Cascade, per million | Saving, per million |
|---|---|---|---|---|---|---|
| claude-opus-5-5 at list price (the model that produced the frontier answers) | $1,321 | £0.9971 | £997 | n/a | n/a | n/a |
| claude-opus-5-5-batch: the same model at a different rate | $661 | £0.4985 | £499 | n/a | n/a | n/a |
| claude-sonnet-5: price only, if escalations went to this model; its accuracy was not measured | $661 | £0.4985 | £499 | n/a | n/a | n/a |
| claude-haiku-4-5: price only, if escalations went to this model; its accuracy was not measured | $330 | £0.2493 | £249 | n/a | n/a | n/a |
No cascade: no threshold met the target, so there is nothing to price.
Every figure in this table is an estimate from list prices and estimated token counts, not a measured bill. It shows what the same decisions would cost through the API, and what share of that the cascade avoids. Basis:
- Frontier tokens are estimated, not counted: characters / 4 (Anthropic's rule of thumb) × 1.3 for the tokenizer, averaged over 1,000 cached answers: 328.7 input and 0.3 output tokens per decision.
- List prices in US dollars per million tokens, checked 2026-09-27 (platform.claude.com/docs/en/about-claude/pricing (checked 2026-09-27)).
- USD to GBP at 0.75458 (ECB euro foreign exchange reference rates, 2026-09-25).
- Frontier answers for this benchmark came from an interactive Claude Code session, not the API; the costs price what the same work would cost through the API.
- Local energy: 342.8 W mean whole-GPU power × 0.01383 s per decision (phase duration / items) = 1.317e-6 kWh per decision, at £0.2632 per kWh (www.ofgem.gov.uk/news/changes-energy-price-cap-between-1-october-and-31-december-2026). Hardware purchase and depreciation are excluded.
Cost for minilm-l6-tickets (baseline, not served through Tau: local latency and energy not measured) Estimates, not measured bills
| Price basis | Frontier only, per million (USD) | Frontier only, per 1,000 | Frontier only, per million | Cascade, per 1,000 | Cascade, per million | Saving, per million |
|---|---|---|---|---|---|---|
| claude-opus-5-5 at list price (the model that produced the frontier answers) | $1,321 | £0.9971 | £997 | £0.9931 | £993 | £3.99 |
| claude-opus-5-5-batch: the same model at a different rate | $661 | £0.4985 | £499 | £0.4965 | £497 | £1.99 |
| claude-sonnet-5: price only, if escalations went to this model; its accuracy was not measured | $661 | £0.4985 | £499 | £0.4965 | £497 | £1.99 |
| claude-haiku-4-5: price only, if escalations went to this model; its accuracy was not measured | $330 | £0.2493 | £249 | £0.2483 | £248 | £0.9971 |
Every figure in this table is an estimate from list prices and estimated token counts, not a measured bill. It shows what the same decisions would cost through the API, and what share of that the cascade avoids. Basis:
- Frontier tokens are estimated, not counted: characters / 4 (Anthropic's rule of thumb) × 1.3 for the tokenizer, averaged over 1,000 cached answers: 328.7 input and 0.3 output tokens per decision.
- List prices in US dollars per million tokens, checked 2026-09-27 (platform.claude.com/docs/en/about-claude/pricing (checked 2026-09-27)).
- USD to GBP at 0.75458 (ECB euro foreign exchange reference rates, 2026-09-25).
- Frontier answers for this benchmark came from an interactive Claude Code session, not the API; the costs price what the same work would cost through the API.
- Local energy is not included: no GPU power was measured for the local model, so the cascade figures price the frontier calls only.
- Cascade: every decision runs locally, and 99.6% are also sent to the frontier model.
Frontier labels: noise and agreement
Label noise is how often claude-opus-5-5 disagrees with the dataset's gold label on the same item; it is measured before any human override. Some of that disagreement is the frontier model's error and some is the gold label's, so it says how far the dataset's labels can be trusted; this report scores against the frontier model instead, so it does not bound the figures above. Agreement between two wordings of the prompt shows how much the frontier answers depend on phrasing. Answers came from claude-code-session: subagent answer sheet; the prompt's fixed part shown once per batch of up to 200 items, each item judged on its own; no API, dated 2026-09-27 to 2026-09-27; no paid API call was made.
Secondary view: against the dataset's own labels
| Model | Agreement with the dataset's labels, raw | Agreement with the dataset's labels, calibrated |
|---|---|---|
| laya-en (Tau) | 48.2% (1,000) | 27.1% (1,000) |
| laya-typed-decisions (Tau) | 43.9% (1,000) | 44.8% (1,000) |
| von-1.2.0 (Tau) | 20.9% (1,000) | 20.7% (1,000) |
| laya-en-ft-tickets (Tau) | 51.1% (1,000) | 40.6% (1,000) |
| minilm-l6-tickets (baseline) | 55.6% (1,000) | 54.5% (1,000) |
| claude-opus-5-5 (frontier) | 23.8% (1,000) | not calibrated |
The dataset's labels are synthetic and close to arbitrary; the frontier model agrees with them on 23.8%. That is no better than the 41.0% a model would get by always giving the most common label. This table is a secondary view: every other figure in the report is agreement with the frontier model, because this example asks whether a local model can stand in for the frontier call.
Baselines
| Model | Held-out items | Agreement | ECE raw | Log loss raw | ECE calibrated | Log loss calibrated | Note |
|---|---|---|---|---|---|---|---|
| minilm-l6-tickets (baseline) | 1,000 | 27.4% | 0.592 | 4.232 | 0.015 | 1.587 | Calibrated with isotonic fitted on the file's 1000 calibration lines |
| laya-en (Tau) | 1,000 | 23.2% | 0.273 | 2.308 | 0.148 | 1.601 | for comparison |
| laya-typed-decisions (Tau) | 1,000 | 18.5% | 0.269 | 1.908 | 0.067 | 1.546 | for comparison |
| von-1.2.0 (Tau) | 1,000 | 43.9% | 0.045 | 1.438 | 0.040 | 1.452 | for comparison |
| laya-en-ft-tickets (Tau) | 1,000 | 27.7% | 0.255 | 2.476 | 0.074 | 1.610 | for comparison |
A classic small encoder, fine-tuned on the same training split, scored with exactly the same metric code as the Tau models. It is here to test the claim that a fixed task favours classic fine-tuning: where it beats a Tau model, that is listed under the misses, and it is a reason to prefer the classic model for this task.
Where the models go wrong
Rows are the frontier model's answer and columns the model's answer, so the diagonal is correct and everything else is a mistake; darker cells hold a larger share of their row. For an ordered scale, mistakes next to the diagonal are near misses and those far from it are serious; a model whose errors sit far off the diagonal should not be trusted on the extreme levels.
Table view: every cell
| Model | Frontier | 0 | 1 | 2 | 3 | 4 |
|---|---|---|---|---|---|---|
| laya-en | 0 | 174 | 50 | 6 | 9 | 0 |
| laya-en | 1 | 68 | 123 | 5 | 28 | 0 |
| laya-en | 2 | 12 | 207 | 1 | 66 | 0 |
| laya-en | 3 | 4 | 59 | 3 | 73 | 0 |
| laya-en | 4 | 6 | 26 | 2 | 77 | 1 |
| laya-typed-decisions | 0 | 0 | 51 | 174 | 14 | 0 |
| laya-typed-decisions | 1 | 0 | 39 | 141 | 44 | 0 |
| laya-typed-decisions | 2 | 0 | 54 | 137 | 95 | 0 |
| laya-typed-decisions | 3 | 0 | 9 | 46 | 83 | 1 |
| laya-typed-decisions | 4 | 0 | 1 | 31 | 76 | 4 |
| von-1.2.0 | 0 | 192 | 3 | 7 | 27 | 10 |
| von-1.2.0 | 1 | 100 | 6 | 22 | 79 | 17 |
| von-1.2.0 | 2 | 76 | 3 | 63 | 139 | 5 |
| von-1.2.0 | 3 | 10 | 2 | 11 | 93 | 23 |
| von-1.2.0 | 4 | 2 | 0 | 1 | 25 | 84 |
| laya-en-ft-tickets | 0 | 107 | 27 | 85 | 20 | 0 |
| laya-en-ft-tickets | 1 | 54 | 42 | 105 | 23 | 0 |
| laya-en-ft-tickets | 2 | 18 | 59 | 143 | 66 | 0 |
| laya-en-ft-tickets | 3 | 11 | 20 | 42 | 66 | 0 |
| laya-en-ft-tickets | 4 | 9 | 29 | 14 | 60 | 0 |
The misses
- Calibration cut laya-en's held-out ECE by only 46.0% (from 0.273 to 0.148), short of the 50% goal.
- Why laya-en fell short: the calibrator is chosen by lower calibration-split log loss, fixed before any held-out result. That picked isotonic (log loss 1.422, ECE 0.201) over temperature (log loss 1.581, ECE 0.016). Log loss and max(p) ECE disagree here, and the rule was not changed after seeing the held-out result.
- Calibration cut von-1.2.0's held-out ECE by only 12.1% (from 0.045 to 0.040), short of the 50% goal.
- Calibration helped von-1.2.0 least: 12.1% lower held-out ECE.
- Baseline minilm-l6-tickets beats laya-en on held-out agreement with the frontier model: 27.4% against 23.2%.
- After calibration, minilm-l6-tickets is also better calibrated than laya-en: ECE 0.015 against 0.148.
- Baseline minilm-l6-tickets beats laya-typed-decisions on held-out agreement with the frontier model: 27.4% against 18.5%.
- After calibration, minilm-l6-tickets is also better calibrated than laya-typed-decisions: ECE 0.015 against 0.067.
- After calibration, minilm-l6-tickets is also better calibrated than von-1.2.0: ECE 0.015 against 0.040.
- After calibration, minilm-l6-tickets is also better calibrated than laya-en-ft-tickets: ECE 0.015 against 0.074.
- laya-en-ft-tickets: No threshold meets the 20.0% target disagreement on the calibration split. The lowest disagreement any threshold reaches is 43.8%, at τ = 0.45, accepting 1.6% of items. No threshold is invented.
- Label noise: the frontier model disagrees with gold on 76.2% (762 of 1,000). This example is scored against the frontier model's answers instead of the dataset's labels; the figures against those labels are kept in a secondary table.
These are the results that don't flatter Tau: label noise, calibration that helped little, baselines that win, and anything excluded. They are published with the wins so the favourable numbers can be weighed against them.
Run metadata
- Date (UTC)
- 2026-09-27T17:15:20Z
- Command
- tau report examples/support-tickets/decision.yaml
- Git commit
- 8954e9b095b40f68ee34139ca7c413fad4386300 (uncommitted changes outside this run's outputs)
- Workbench
- 0.1.0
- GPU
- NVIDIA GeForce RTX 3080 Ti (driver 610.47)
- CPU
- 11th Gen Intel(R) Core(TM) i9-11900K @ 3.50GHz, 16 logical processors
- OS and runtime
- Microsoft Windows 10.0.19045; .NET 10.0.11
- Endpoint
- localhost:18093/ (Tau Runtime at localhost:18093/ (4 model(s) installed))
- /v1/models: laya-en
- sha256 866a05b244e47e96820660d18ee050c518c2de0f60c39e2e9a89dfeb056e1eec, revision 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851
- /v1/models: laya-typed-decisions
- sha256 2b7a961ac37157d1cfa8105e3283106baf1ba2a5cc30fb3a673253f06aa1e1c5, revision 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851
- /v1/models: von-1.2.0
- sha256 0777bb988636663b770775ae0b4eb961d6fbee176c9bea1d1da823723b1eb3f3, revision 5df8185a4f2327ad0a7cd117cc4f701ac557b9ae
- /v1/models: laya-en-ft-tickets
- sha256 9079fee9291e35dc4d3ed8c817a8f1592d97e4413ccbae0539b7a0ee6d9c950a, revision 03677900b97d57980f7c525a84031b9a43cb989e
- Measured model laya-en
- 866a05b244e47e96820660d18ee050c518c2de0f60c39e2e9a89dfeb056e1eec
- Measured model laya-typed-decisions
- 2b7a961ac37157d1cfa8105e3283106baf1ba2a5cc30fb3a673253f06aa1e1c5
- Measured model von-1.2.0
- 0777bb988636663b770775ae0b4eb961d6fbee176c9bea1d1da823723b1eb3f3
- Measured model laya-en-ft-tickets
- 9079fee9291e35dc4d3ed8c817a8f1592d97e4413ccbae0539b7a0ee6d9c950a
- Dataset manifest sha256
- 8aae6044d2ff40ea9392603672ab694b33cca046c70c5a0c532392bb67efb9fb
- Dataset revision (cache keys)
- ddf1c81a5475
- Reference labels
- frontier: Scored against the frontier model: every rate is agreement with claude-opus-5-5's v1 answers (after any human override), not accuracy against the dataset's labels (data.reference: frontier).
- Frontier model
- claude-opus-5-5
- Prompt versions
- v1, v1-alt
- Frontier cache
- v1: 2,000 answers, v1-alt: 200 answers
- Confidence
- Confidence for ECE and thresholds is max(p), the probability of the chosen answer, not the contract's confidence field, which for choice and score questions is not a calibrated probability.
This block ties every figure above to a machine, a model hash, a dataset revision, a date and a command. Re-running the command on the same checkout and models should reproduce the numbers; anything that differs should be explained by what changed here.