Tau Workbench report

Synthetic support tickets: how urgent is this ticket?

Synthetic data Every item in this dataset was generated, not written by real customers. Treat the numbers as a demonstration of the method, not as evidence about real traffic.

This report measures how far 4 model(s) can be trusted on "How urgent is this customer support ticket?", using synthetic tickets data (1,000 held-out items). Reference: claude-opus-5-5's answers, not the dataset's labels. The question here is whether a local model can stand in for the frontier call, and whether its confidence says when, so every rate below is agreement with the frontier model. Out of the box, laya-en agrees with claude-opus-5-5 on 23.2% of held-out items with a calibration error (ECE) of 0.273; after calibration its ECE is 0.148, 46.0% lower, short of the 50% goal. At a 20.0% target disagreement, laya-en keeps 0.1% of decisions local (3 models tie at that share); the cascade's served answers agree with claude-opus-5-5 on 100.0% (the frontier alone agrees with itself by construction), at an estimated £996 per million decisions against £997 (list prices, 2026-09-27). The dataset's own labels are reported as a finding, not used: claude-opus-5-5 agrees with them on only 23.8% of 1,000 items, and a secondary table shows every model against them.

Generated 2026-09-27T17:15:20Z (UTC) on NVIDIA GeForce RTX 3080 Ti. Every number below is in report.json beside this file; the metadata block at the end says how to reproduce it.

Dataset

Name
tickets
Source
Tobi-Bueck/customer-support-tickets
Pinned revision
ddf1c81a5475992c4fa6752bf1e8b4e31f07bbeb
Licence
CC-BY-NC-4.0
Synthetic
Yes: generated data
Fine-tune split
8,000 items
Calibration split
1,000 items
Held-out split
1,000 items
Question
urgency (score, 5 answers): How urgent is this customer support ticket?

Calibrators and thresholds are fitted on the calibration split only and every result is judged on the held-out split, which nothing was tuned on. The splits are fixed by the dataset manifest (seed 42), and each file's sha256 is checked before use.

Agreement with the frontier model and calibration, before and after

ModelItemsAgreement rawAgreement calibratedECE rawECE calibratedECE changeBrier raw → cal.Log loss raw → cal.MAE (levels) raw → cal.Calibrated view
laya-en1,00023.2%37.2%0.2730.148−46.0%0.952 → 0.7572.308 → 1.6011.19 → 0.83Runtime with calibrators loaded; matches the Workbench's offline calibration to within 5.5e-2 in any probability
laya-typed-decisions1,00018.5%26.3%0.2690.067−75.0%0.895 → 0.7771.908 → 1.5461.28 → 1.03Runtime with calibrators loaded; matches the Workbench's offline calibration to within 7.3e-2 in any probability
von-1.2.01,00043.9%43.8%0.0450.040−12.1%0.703 → 0.7011.438 → 1.4520.88 → 0.88Runtime with calibrators loaded; matches the Workbench's offline calibration to within 4.6e-2 in any probability
laya-en-ft-tickets1,00027.7%35.8%0.2550.074−71.1%0.926 → 0.7712.476 → 1.6101.08 → 0.95Runtime with calibrators loaded; matches the Workbench's offline calibration to within 5.8e-2 in any probability

Held-out items only. ECE (expected calibration error) is the average gap between how confident a model is and how often it is right, over 15 confidence bins; lower is better and 0 is perfect. Confidence for ECE and thresholds is max(p), the probability of the chosen answer, not the contract's confidence field, which for choice and score questions is not a calibrated probability. Here a model is right when its answer matches claude-opus-5-5's, so the agreement columns are agreement with the frontier model, not accuracy, and ECE measures whether a model's confidence says when it agrees. The goal is a 50% cut in ECE. The largest cut is laya-typed-decisions's (75.0%) and the smallest is von-1.2.0's (12.1%); a model whose confidence is still off after calibration should not be trusted to gate decisions on its own.

Reliability diagrams

RawCalibratedPerfect calibration
laya-en
0.000.250.500.751.000.000.250.500.751.00Perfect calibration: agreement with the frontier model equals confidenceRaw: confidence 0.33 to 0.40, 77 items, mean confidence 0.380, agreement with the frontier model 0.351Raw: confidence 0.40 to 0.47, 375 items, mean confidence 0.438, agreement with the frontier model 0.224Raw: confidence 0.47 to 0.53, 259 items, mean confidence 0.495, agreement with the frontier model 0.158Raw: confidence 0.53 to 0.60, 115 items, mean confidence 0.560, agreement with the frontier model 0.243Raw: confidence 0.60 to 0.67, 91 items, mean confidence 0.632, agreement with the frontier model 0.253Raw: confidence 0.67 to 0.73, 47 items, mean confidence 0.695, agreement with the frontier model 0.340Raw: confidence 0.73 to 0.80, 23 items, mean confidence 0.756, agreement with the frontier model 0.261Raw: confidence 0.80 to 0.87, 10 items, mean confidence 0.828, agreement with the frontier model 0.400Raw: confidence 0.87 to 0.93, 3 items, mean confidence 0.908, agreement with the frontier model 1.000Calibrated: confidence 0.13 to 0.20, 105 items, mean confidence 0.200, agreement with the frontier model 0.648Calibrated: confidence 0.20 to 0.27, 304 items, mean confidence 0.233, agreement with the frontier model 0.464Calibrated: confidence 0.27 to 0.33, 395 items, mean confidence 0.301, agreement with the frontier model 0.263Calibrated: confidence 0.33 to 0.40, 140 items, mean confidence 0.351, agreement with the frontier model 0.307Calibrated: confidence 0.40 to 0.47, 52 items, mean confidence 0.429, agreement with the frontier model 0.269Calibrated: confidence 0.47 to 0.53, 2 items, mean confidence 0.497, agreement with the frontier model 0.500Calibrated: confidence 0.53 to 0.60, 1 items, mean confidence 0.556, agreement with the frontier model 0.000Calibrated: confidence 0.60 to 0.67, 1 items, mean confidence 0.650, agreement with the frontier model 1.000Confidence, max(p)Agreement with the frontier model
laya-typed-decisions
0.000.250.500.751.000.000.250.500.751.00Perfect calibration: agreement with the frontier model equals confidenceRaw: confidence 0.27 to 0.33, 7 items, mean confidence 0.323, agreement with the frontier model 0.571Raw: confidence 0.33 to 0.40, 225 items, mean confidence 0.378, agreement with the frontier model 0.213Raw: confidence 0.40 to 0.47, 407 items, mean confidence 0.431, agreement with the frontier model 0.140Raw: confidence 0.47 to 0.53, 241 items, mean confidence 0.498, agreement with the frontier model 0.174Raw: confidence 0.53 to 0.60, 105 items, mean confidence 0.559, agreement with the frontier model 0.286Raw: confidence 0.60 to 0.67, 15 items, mean confidence 0.623, agreement with the frontier model 0.267Calibrated: confidence 0.20 to 0.27, 560 items, mean confidence 0.229, agreement with the frontier model 0.196Calibrated: confidence 0.27 to 0.33, 285 items, mean confidence 0.279, agreement with the frontier model 0.393Calibrated: confidence 0.33 to 0.40, 154 items, mean confidence 0.362, agreement with the frontier model 0.260Calibrated: confidence 0.40 to 0.47, 1 items, mean confidence 0.453, agreement with the frontier model 1.000Confidence, max(p)Agreement with the frontier model
von-1.2.0
0.000.250.500.751.000.000.250.500.751.00Perfect calibration: agreement with the frontier model equals confidenceRaw: confidence 0.27 to 0.33, 43 items, mean confidence 0.315, agreement with the frontier model 0.256Raw: confidence 0.33 to 0.40, 206 items, mean confidence 0.373, agreement with the frontier model 0.296Raw: confidence 0.40 to 0.47, 265 items, mean confidence 0.431, agreement with the frontier model 0.385Raw: confidence 0.47 to 0.53, 201 items, mean confidence 0.499, agreement with the frontier model 0.463Raw: confidence 0.53 to 0.60, 147 items, mean confidence 0.562, agreement with the frontier model 0.571Raw: confidence 0.60 to 0.67, 78 items, mean confidence 0.627, agreement with the frontier model 0.615Raw: confidence 0.67 to 0.73, 44 items, mean confidence 0.696, agreement with the frontier model 0.636Raw: confidence 0.73 to 0.80, 9 items, mean confidence 0.767, agreement with the frontier model 0.778Raw: confidence 0.80 to 0.87, 3 items, mean confidence 0.817, agreement with the frontier model 1.000Raw: confidence 0.87 to 0.93, 4 items, mean confidence 0.904, agreement with the frontier model 0.500Calibrated: confidence 0.20 to 0.27, 2 items, mean confidence 0.259, agreement with the frontier model 0.500Calibrated: confidence 0.27 to 0.33, 68 items, mean confidence 0.299, agreement with the frontier model 0.353Calibrated: confidence 0.33 to 0.40, 321 items, mean confidence 0.358, agreement with the frontier model 0.312Calibrated: confidence 0.40 to 0.47, 128 items, mean confidence 0.444, agreement with the frontier model 0.406Calibrated: confidence 0.47 to 0.53, 136 items, mean confidence 0.498, agreement with the frontier model 0.456Calibrated: confidence 0.53 to 0.60, 215 items, mean confidence 0.566, agreement with the frontier model 0.535Calibrated: confidence 0.60 to 0.67, 105 items, mean confidence 0.630, agreement with the frontier model 0.648Calibrated: confidence 0.67 to 0.73, 20 items, mean confidence 0.687, agreement with the frontier model 0.650Calibrated: confidence 0.73 to 0.80, 4 items, mean confidence 0.767, agreement with the frontier model 0.750Calibrated: confidence 0.80 to 0.87, 1 items, mean confidence 0.852, agreement with the frontier model 0.000Confidence, max(p)Agreement with the frontier model
laya-en-ft-tickets
0.000.250.500.751.000.000.250.500.751.00Perfect calibration: agreement with the frontier model equals confidenceRaw: confidence 0.33 to 0.40, 12 items, mean confidence 0.392, agreement with the frontier model 0.167Raw: confidence 0.40 to 0.47, 243 items, mean confidence 0.442, agreement with the frontier model 0.185Raw: confidence 0.47 to 0.53, 374 items, mean confidence 0.496, agreement with the frontier model 0.235Raw: confidence 0.53 to 0.60, 194 items, mean confidence 0.560, agreement with the frontier model 0.392Raw: confidence 0.60 to 0.67, 76 items, mean confidence 0.630, agreement with the frontier model 0.289Raw: confidence 0.67 to 0.73, 37 items, mean confidence 0.698, agreement with the frontier model 0.405Raw: confidence 0.73 to 0.80, 30 items, mean confidence 0.766, agreement with the frontier model 0.500Raw: confidence 0.80 to 0.87, 21 items, mean confidence 0.829, agreement with the frontier model 0.381Raw: confidence 0.87 to 0.93, 10 items, mean confidence 0.894, agreement with the frontier model 0.400Raw: confidence 0.93 to 1.00, 3 items, mean confidence 0.948, agreement with the frontier model 0.667Calibrated: confidence 0.13 to 0.20, 9 items, mean confidence 0.200, agreement with the frontier model 0.556Calibrated: confidence 0.20 to 0.27, 413 items, mean confidence 0.242, agreement with the frontier model 0.370Calibrated: confidence 0.27 to 0.33, 181 items, mean confidence 0.289, agreement with the frontier model 0.276Calibrated: confidence 0.33 to 0.40, 240 items, mean confidence 0.357, agreement with the frontier model 0.379Calibrated: confidence 0.40 to 0.47, 147 items, mean confidence 0.420, agreement with the frontier model 0.361Calibrated: confidence 0.47 to 0.53, 9 items, mean confidence 0.480, agreement with the frontier model 0.556Calibrated: confidence 0.53 to 0.60, 1 items, mean confidence 0.583, agreement with the frontier model 1.000Confidence, max(p)Agreement with the frontier model

Each point is one confidence bin on the held-out split: how confident the model was (across) against how often it was right (up), where right means agreeing with the frontier model. Points below the diagonal are overconfident, points above it underconfident. After calibration the points should sit close to the diagonal; where they don't, a threshold on that model's confidence will not deliver the agreement it promises.

Table view: every bin
ModelBinRaw itemsRaw confidenceRaw agreementCal. itemsCal. confidenceCal. agreement
laya-en0.13–0.200n/an/a1050.2000.648
laya-en0.20–0.270n/an/a3040.2330.464
laya-en0.27–0.330n/an/a3950.3010.263
laya-en0.33–0.40770.3800.3511400.3510.307
laya-en0.40–0.473750.4380.224520.4290.269
laya-en0.47–0.532590.4950.15820.4970.500
laya-en0.53–0.601150.5600.24310.5560.000
laya-en0.60–0.67910.6320.25310.6501.000
laya-en0.67–0.73470.6950.3400n/an/a
laya-en0.73–0.80230.7560.2610n/an/a
laya-en0.80–0.87100.8280.4000n/an/a
laya-en0.87–0.9330.9081.0000n/an/a
laya-typed-decisions0.20–0.270n/an/a5600.2290.196
laya-typed-decisions0.27–0.3370.3230.5712850.2790.393
laya-typed-decisions0.33–0.402250.3780.2131540.3620.260
laya-typed-decisions0.40–0.474070.4310.14010.4531.000
laya-typed-decisions0.47–0.532410.4980.1740n/an/a
laya-typed-decisions0.53–0.601050.5590.2860n/an/a
laya-typed-decisions0.60–0.67150.6230.2670n/an/a
von-1.2.00.20–0.270n/an/a20.2590.500
von-1.2.00.27–0.33430.3150.256680.2990.353
von-1.2.00.33–0.402060.3730.2963210.3580.312
von-1.2.00.40–0.472650.4310.3851280.4440.406
von-1.2.00.47–0.532010.4990.4631360.4980.456
von-1.2.00.53–0.601470.5620.5712150.5660.535
von-1.2.00.60–0.67780.6270.6151050.6300.648
von-1.2.00.67–0.73440.6960.636200.6870.650
von-1.2.00.73–0.8090.7670.77840.7670.750
von-1.2.00.80–0.8730.8171.00010.8520.000
von-1.2.00.87–0.9340.9040.5000n/an/a
laya-en-ft-tickets0.13–0.200n/an/a90.2000.556
laya-en-ft-tickets0.20–0.270n/an/a4130.2420.370
laya-en-ft-tickets0.27–0.330n/an/a1810.2890.276
laya-en-ft-tickets0.33–0.40120.3920.1672400.3570.379
laya-en-ft-tickets0.40–0.472430.4420.1851470.4200.361
laya-en-ft-tickets0.47–0.533740.4960.23590.4800.556
laya-en-ft-tickets0.53–0.601940.5600.39210.5831.000
laya-en-ft-tickets0.60–0.67760.6300.2890n/an/a
laya-en-ft-tickets0.67–0.73370.6980.4050n/an/a
laya-en-ft-tickets0.73–0.80300.7660.5000n/an/a
laya-en-ft-tickets0.80–0.87210.8290.3810n/an/a
laya-en-ft-tickets0.87–0.93100.8940.4000n/an/a
laya-en-ft-tickets0.93–1.0030.9480.6670n/an/a

Calibrators fitted

ModelScopeItemsChosenTemperature TTemperature ECE / log lossIsotonic knotsIsotonic ECE / log lossBefore: ECE / log loss
laya-enquestion type1,000isotonic7.6180.016 / 1.581320.201 / 1.4220.244 / 2.306
laya-en3-5 options1,000isotonic7.6180.016 / 1.581320.201 / 1.4220.244 / 2.306
laya-typed-decisionsquestion type1,000isotonic5.7090.054 / 1.587300.094 / 1.4990.244 / 1.922
laya-typed-decisions3-5 options1,000isotonic5.7090.054 / 1.587300.094 / 1.4990.244 / 1.922
von-1.2.0question type1,000isotonic1.1520.055 / 1.331670.035 / 1.3030.052 / 1.336
von-1.2.03-5 options1,000isotonic1.1520.055 / 1.331670.035 / 1.3030.052 / 1.336
laya-en-ft-ticketsquestion type1,000isotonic12.8360.055 / 1.598370.082 / 1.5120.250 / 2.549
laya-en-ft-tickets3-5 options1,000isotonic12.8360.055 / 1.598370.082 / 1.5120.250 / 2.549

Both methods are fitted on the calibration split's raw probabilities, and the one with the lower calibration-split log loss is written for the Runtime to load; the scores in this table are on the calibration split, so they show the fit, not the result (the held-out result is in the first table). Isotonic regression needs at least 200 items and falls back to temperature scaling below that.

How much can stay local

laya-enlaya-typed-decisionsvon-1.2.0laya-en-ft-ticketsminilm-l6-tickets (baseline)Target agreement
10%28%46%64%82%100%0%20%40%60%80%100%Target: 80.0% agreement with the frontier model on kept itemstarget 80%laya-enlaya-en: τ = 0.00, keeps 100.0% local, 37.2% agreement with the frontier model on thoselaya-en: τ = 0.05, keeps 100.0% local, 37.2% agreement with the frontier model on thoselaya-en: τ = 0.10, keeps 100.0% local, 37.2% agreement with the frontier model on thoselaya-en: τ = 0.15, keeps 100.0% local, 37.2% agreement with the frontier model on thoselaya-en: τ = 0.20, keeps 89.7% local, 34.0% agreement with the frontier model on thoselaya-en: τ = 0.25, keeps 65.2% local, 26.8% agreement with the frontier model on thoselaya-en: τ = 0.30, keeps 40.6% local, 28.3% agreement with the frontier model on thoselaya-en: τ = 0.35, keeps 9.0% local, 27.8% agreement with the frontier model on thoselaya-en: τ = 0.40, keeps 5.6% local, 28.6% agreement with the frontier model on thoselaya-en: τ = 0.45, keeps 1.1% local, 63.6% agreement with the frontier model on thoselaya-en: chosen τ = 0.56; held-out keeps 0.1% local at 100.0% agreement with the frontier modelτ 0.56laya-typed-decisionslaya-typed-decisions: τ = 0.00, keeps 100.0% local, 26.3% agreement with the frontier model on thoselaya-typed-decisions: τ = 0.05, keeps 100.0% local, 26.3% agreement with the frontier model on thoselaya-typed-decisions: τ = 0.10, keeps 100.0% local, 26.3% agreement with the frontier model on thoselaya-typed-decisions: τ = 0.15, keeps 100.0% local, 26.3% agreement with the frontier model on thoselaya-typed-decisions: τ = 0.20, keeps 100.0% local, 26.3% agreement with the frontier model on thoselaya-typed-decisions: τ = 0.25, keeps 53.2% local, 36.8% agreement with the frontier model on thoselaya-typed-decisions: τ = 0.30, keeps 15.9% local, 25.8% agreement with the frontier model on thoselaya-typed-decisions: τ = 0.35, keeps 12.0% local, 28.3% agreement with the frontier model on thoselaya-typed-decisions: chosen τ = 0.39; held-out keeps 0.1% local at 100.0% agreement with the frontier modelτ 0.39von-1.2.0von-1.2.0: τ = 0.00, keeps 100.0% local, 43.8% agreement with the frontier model on thosevon-1.2.0: τ = 0.05, keeps 100.0% local, 43.8% agreement with the frontier model on thosevon-1.2.0: τ = 0.10, keeps 100.0% local, 43.8% agreement with the frontier model on thosevon-1.2.0: τ = 0.15, keeps 100.0% local, 43.8% agreement with the frontier model on thosevon-1.2.0: τ = 0.20, keeps 100.0% local, 43.8% agreement with the frontier model on thosevon-1.2.0: τ = 0.25, keeps 100.0% local, 43.8% agreement with the frontier model on thosevon-1.2.0: τ = 0.30, keeps 95.9% local, 44.3% agreement with the frontier model on thosevon-1.2.0: τ = 0.35, keeps 78.4% local, 47.8% agreement with the frontier model on thosevon-1.2.0: τ = 0.40, keeps 60.9% local, 51.4% agreement with the frontier model on thosevon-1.2.0: τ = 0.45, keeps 53.9% local, 52.5% agreement with the frontier model on thosevon-1.2.0: τ = 0.50, keeps 41.5% local, 55.7% agreement with the frontier model on thosevon-1.2.0: τ = 0.55, keeps 29.6% local, 60.1% agreement with the frontier model on thosevon-1.2.0: τ = 0.60, keeps 13.0% local, 64.6% agreement with the frontier model on thosevon-1.2.0: τ = 0.65, keeps 4.0% local, 60.0% agreement with the frontier model on thosevon-1.2.0: τ = 0.70, keeps 1.1% local, 81.8% agreement with the frontier model on thosevon-1.2.0: chosen τ = 0.78; held-out keeps 0.1% local at 0.0% agreement with the frontier modelτ 0.78laya-en-ft-ticketslaya-en-ft-tickets: τ = 0.00, keeps 100.0% local, 35.3% agreement with the frontier model on thoselaya-en-ft-tickets: τ = 0.05, keeps 100.0% local, 35.3% agreement with the frontier model on thoselaya-en-ft-tickets: τ = 0.10, keeps 100.0% local, 35.3% agreement with the frontier model on thoselaya-en-ft-tickets: τ = 0.15, keeps 100.0% local, 35.3% agreement with the frontier model on thoselaya-en-ft-tickets: τ = 0.20, keeps 99.5% local, 35.3% agreement with the frontier model on thoselaya-en-ft-tickets: τ = 0.25, keeps 75.1% local, 31.8% agreement with the frontier model on thoselaya-en-ft-tickets: τ = 0.30, keeps 46.7% local, 36.0% agreement with the frontier model on thoselaya-en-ft-tickets: τ = 0.35, keeps 28.9% local, 37.7% agreement with the frontier model on thoselaya-en-ft-tickets: τ = 0.40, keeps 15.7% local, 37.6% agreement with the frontier model on thoselaya-en-ft-tickets: τ = 0.45, keeps 1.2% local, 50.0% agreement with the frontier model on thoseminilm-l6-tickets (baseline)minilm-l6-tickets: τ = 0.00, keeps 100.0% local, 27.2% agreement with the frontier model on thoseminilm-l6-tickets: τ = 0.05, keeps 100.0% local, 27.2% agreement with the frontier model on thoseminilm-l6-tickets: τ = 0.10, keeps 100.0% local, 27.2% agreement with the frontier model on thoseminilm-l6-tickets: τ = 0.15, keeps 100.0% local, 27.2% agreement with the frontier model on thoseminilm-l6-tickets: τ = 0.20, keeps 100.0% local, 27.2% agreement with the frontier model on thoseminilm-l6-tickets: τ = 0.25, keeps 88.8% local, 27.7% agreement with the frontier model on thoseminilm-l6-tickets: τ = 0.30, keeps 4.2% local, 42.9% agreement with the frontier model on thoseminilm-l6-tickets: τ = 0.35, keeps 1.4% local, 50.0% agreement with the frontier model on thoseminilm-l6-tickets: τ = 0.40, keeps 1.4% local, 50.0% agreement with the frontier model on thoseminilm-l6-tickets: chosen τ = 0.44; held-out keeps 0.4% local at 25.0% agreement with the frontier modelτ 0.44Share of decisions kept local (confidence ≥ τ)Agreement with the frontier model on kept items

Each line shows, on held-out items, what happens as the confidence threshold τ rises: fewer decisions are kept local (moving left) and those kept are more often right (moving up), where right means agreeing with the frontier model. The dot marks the τ chosen on the calibration split for a 20.0% target disagreement. A line that never reaches the target line means no threshold makes that model safe enough on its own at that target. Thresholds that keep fewer than 10 items are left off the chart because a handful of items says little; every point is in report.json. Dashed lines are baselines, not served through Tau, put through exactly the same threshold rule on their own calibration lines.

ModelConfidences fromτKept local (calibration)Kept local (held-out)Agreement on kept (held-out)Result
laya-enoffline0.560.1%0.1%100.0%τ = 0.56 is the smallest threshold whose calibration-split disagreement is at most 20.0%. On held-out it accepts 0.1% of items with 100.0% agreement with the frontier model.
laya-typed-decisionsoffline0.390.1%0.1%100.0%τ = 0.39 is the smallest threshold whose calibration-split disagreement is at most 20.0%. On held-out it accepts 0.1% of items with 100.0% agreement with the frontier model.
von-1.2.0offline0.780.3%0.1%0.0%τ = 0.78 is the smallest threshold whose calibration-split disagreement is at most 20.0%. On held-out it accepts 0.1% of items with 0.0% agreement with the frontier model.
laya-en-ft-ticketsofflinen/an/an/an/aNo threshold meets the 20.0% target disagreement on the calibration split. The lowest disagreement any threshold reaches is 43.8%, at τ = 0.45, accepting 1.6% of items. No threshold is invented.
minilm-l6-tickets (baseline)baseline-calibrated0.440.2%0.4%25.0%τ = 0.44 is the smallest threshold whose calibration-split disagreement is at most 20.0%. On held-out it accepts 0.4% of items with 25.0% agreement with the frontier model.

Cascade: local first, frontier for the rest

ModelτKept localEscalatedNo frontier answerLocal only (agreement)Frontier onlyCascade (agreement)Local p50 latency
laya-en0.560.1%99.9%037.2% (1,000)100% by construction100.0% (1,000)53.2 ms
laya-typed-decisions0.390.1%99.9%026.3% (1,000)100% by construction100.0% (1,000)53.5 ms
von-1.2.00.780.1%99.9%043.8% (1,000)100% by construction99.9% (1,000)45.0 ms
laya-en-ft-ticketsn/an/an/an/a35.3% (1,000)100% by constructionNo cascade: no threshold met the 20.0% target disagreement on the calibration split.50.8 ms
minilm-l6-tickets (baseline, not served through Tau: local latency and energy not measured)0.440.4%99.6%027.2% (1,000)100% by construction99.7% (1,000)not measured

Every rate here is agreement with the frontier model on held-out items, not accuracy, with the number of items it covers in brackets. The cascade serves the local answer when its confidence is at least τ and the frontier model's cached answer otherwise; the cascade figure is the share of items where the served answer equals the frontier model's. Frontier only agrees with the frontier model on 100% of items by construction: its answers are the reference, so this is not a result. The closer the cascade gets to 100% while keeping a large share local, the better the local model stands in for the frontier call. Frontier latency is not measured: no API call was made, so the latency mix covers the local side only.

Cost for laya-en Estimates, not measured bills

Price basisFrontier only, per million (USD)Frontier only, per 1,000Frontier only, per millionCascade, per 1,000Cascade, per millionSaving, per million
claude-opus-5-5 at list price (the model that produced the frontier answers)$1,321£0.9971£997£0.9964£996£0.6549
claude-opus-5-5-batch: the same model at a different rate$661£0.4985£499£0.4984£498£0.1564
claude-sonnet-5: price only, if escalations went to this model; its accuracy was not measured$661£0.4985£499£0.4984£498£0.1564
claude-haiku-4-5: price only, if escalations went to this model; its accuracy was not measured$330£0.2493£249£0.2494£249£-0.0929

Every figure in this table is an estimate from list prices and estimated token counts, not a measured bill. It shows what the same decisions would cost through the API, and what share of that the cascade avoids. Basis:

Cost for laya-typed-decisions Estimates, not measured bills

Price basisFrontier only, per million (USD)Frontier only, per 1,000Frontier only, per millionCascade, per 1,000Cascade, per millionSaving, per million
claude-opus-5-5 at list price (the model that produced the frontier answers)$1,321£0.9971£997£0.9964£996£0.6509
claude-opus-5-5-batch: the same model at a different rate$661£0.4985£499£0.4984£498£0.1523
claude-sonnet-5: price only, if escalations went to this model; its accuracy was not measured$661£0.4985£499£0.4984£498£0.1523
claude-haiku-4-5: price only, if escalations went to this model; its accuracy was not measured$330£0.2493£249£0.2494£249£-0.0969

Every figure in this table is an estimate from list prices and estimated token counts, not a measured bill. It shows what the same decisions would cost through the API, and what share of that the cascade avoids. Basis:

Cost for von-1.2.0 Estimates, not measured bills

Price basisFrontier only, per million (USD)Frontier only, per 1,000Frontier only, per millionCascade, per 1,000Cascade, per millionSaving, per million
claude-opus-5-5 at list price (the model that produced the frontier answers)$1,321£0.9971£997£0.9963£996£0.7079
claude-opus-5-5-batch: the same model at a different rate$661£0.4985£499£0.4983£498£0.2093
claude-sonnet-5: price only, if escalations went to this model; its accuracy was not measured$661£0.4985£499£0.4983£498£0.2093
claude-haiku-4-5: price only, if escalations went to this model; its accuracy was not measured$330£0.2493£249£0.2493£249£-0.0399

Every figure in this table is an estimate from list prices and estimated token counts, not a measured bill. It shows what the same decisions would cost through the API, and what share of that the cascade avoids. Basis:

Cost for laya-en-ft-tickets Estimates, not measured bills

Price basisFrontier only, per million (USD)Frontier only, per 1,000Frontier only, per millionCascade, per 1,000Cascade, per millionSaving, per million
claude-opus-5-5 at list price (the model that produced the frontier answers)$1,321£0.9971£997n/an/an/a
claude-opus-5-5-batch: the same model at a different rate$661£0.4985£499n/an/an/a
claude-sonnet-5: price only, if escalations went to this model; its accuracy was not measured$661£0.4985£499n/an/an/a
claude-haiku-4-5: price only, if escalations went to this model; its accuracy was not measured$330£0.2493£249n/an/an/a

No cascade: no threshold met the target, so there is nothing to price.

Every figure in this table is an estimate from list prices and estimated token counts, not a measured bill. It shows what the same decisions would cost through the API, and what share of that the cascade avoids. Basis:

Cost for minilm-l6-tickets (baseline, not served through Tau: local latency and energy not measured) Estimates, not measured bills

Price basisFrontier only, per million (USD)Frontier only, per 1,000Frontier only, per millionCascade, per 1,000Cascade, per millionSaving, per million
claude-opus-5-5 at list price (the model that produced the frontier answers)$1,321£0.9971£997£0.9931£993£3.99
claude-opus-5-5-batch: the same model at a different rate$661£0.4985£499£0.4965£497£1.99
claude-sonnet-5: price only, if escalations went to this model; its accuracy was not measured$661£0.4985£499£0.4965£497£1.99
claude-haiku-4-5: price only, if escalations went to this model; its accuracy was not measured$330£0.2493£249£0.2483£248£0.9971

Every figure in this table is an estimate from list prices and estimated token counts, not a measured bill. It shows what the same decisions would cost through the API, and what share of that the cascade avoids. Basis:

Frontier labels: noise and agreement

76.2%
disagree with gold (762 of 1,000 items)
76.5%
agreement between v1 and v1-alt wordings (153 of 200)
1,000
v1 answers cached of 1,000 eligible held-out items; 0 pending
1,000
v1 answers cached of 1,000 eligible calibration items; 0 pending
0
labels overridden by a person; 0 answer line(s) rejected

Label noise is how often claude-opus-5-5 disagrees with the dataset's gold label on the same item; it is measured before any human override. Some of that disagreement is the frontier model's error and some is the gold label's, so it says how far the dataset's labels can be trusted; this report scores against the frontier model instead, so it does not bound the figures above. Agreement between two wordings of the prompt shows how much the frontier answers depend on phrasing. Answers came from claude-code-session: subagent answer sheet; the prompt's fixed part shown once per batch of up to 200 items, each item judged on its own; no API, dated 2026-09-27 to 2026-09-27; no paid API call was made.

Secondary view: against the dataset's own labels

ModelAgreement with the dataset's labels, rawAgreement with the dataset's labels, calibrated
laya-en (Tau)48.2% (1,000)27.1% (1,000)
laya-typed-decisions (Tau)43.9% (1,000)44.8% (1,000)
von-1.2.0 (Tau)20.9% (1,000)20.7% (1,000)
laya-en-ft-tickets (Tau)51.1% (1,000)40.6% (1,000)
minilm-l6-tickets (baseline)55.6% (1,000)54.5% (1,000)
claude-opus-5-5 (frontier)23.8% (1,000)not calibrated

The dataset's labels are synthetic and close to arbitrary; the frontier model agrees with them on 23.8%. That is no better than the 41.0% a model would get by always giving the most common label. This table is a secondary view: every other figure in the report is agreement with the frontier model, because this example asks whether a local model can stand in for the frontier call.

Baselines

ModelHeld-out itemsAgreementECE rawLog loss rawECE calibratedLog loss calibratedNote
minilm-l6-tickets (baseline)1,00027.4%0.5924.2320.0151.587Calibrated with isotonic fitted on the file's 1000 calibration lines
laya-en (Tau)1,00023.2%0.2732.3080.1481.601for comparison
laya-typed-decisions (Tau)1,00018.5%0.2691.9080.0671.546for comparison
von-1.2.0 (Tau)1,00043.9%0.0451.4380.0401.452for comparison
laya-en-ft-tickets (Tau)1,00027.7%0.2552.4760.0741.610for comparison

A classic small encoder, fine-tuned on the same training split, scored with exactly the same metric code as the Tau models. It is here to test the claim that a fixed task favours classic fine-tuning: where it beats a Tau model, that is listed under the misses, and it is a reason to prefer the classic model for this task.

Where the models go wrong

laya-en (calibrated, 1,000 items)
answered →00112233440frontier: 0frontier 0, answered 0: 174 of 239 (72.8%)174frontier 0, answered 1: 50 of 239 (20.9%)50frontier 0, answered 2: 6 of 239 (2.5%)6frontier 0, answered 3: 9 of 239 (3.8%)9frontier 0, answered 4: 0 of 239 (0.0%)1frontier: 1frontier 1, answered 0: 68 of 224 (30.4%)68frontier 1, answered 1: 123 of 224 (54.9%)123frontier 1, answered 2: 5 of 224 (2.2%)5frontier 1, answered 3: 28 of 224 (12.5%)28frontier 1, answered 4: 0 of 224 (0.0%)2frontier: 2frontier 2, answered 0: 12 of 286 (4.2%)12frontier 2, answered 1: 207 of 286 (72.4%)207frontier 2, answered 2: 1 of 286 (0.3%)1frontier 2, answered 3: 66 of 286 (23.1%)66frontier 2, answered 4: 0 of 286 (0.0%)3frontier: 3frontier 3, answered 0: 4 of 139 (2.9%)4frontier 3, answered 1: 59 of 139 (42.4%)59frontier 3, answered 2: 3 of 139 (2.2%)3frontier 3, answered 3: 73 of 139 (52.5%)73frontier 3, answered 4: 0 of 139 (0.0%)4frontier: 4frontier 4, answered 0: 6 of 112 (5.4%)6frontier 4, answered 1: 26 of 112 (23.2%)26frontier 4, answered 2: 2 of 112 (1.8%)2frontier 4, answered 3: 77 of 112 (68.8%)77frontier 4, answered 4: 1 of 112 (0.9%)1rows: frontier answer; shade: share of the row
laya-typed-decisions (calibrated, 1,000 items)
answered →00112233440frontier: 0frontier 0, answered 0: 0 of 239 (0.0%)frontier 0, answered 1: 51 of 239 (21.3%)51frontier 0, answered 2: 174 of 239 (72.8%)174frontier 0, answered 3: 14 of 239 (5.9%)14frontier 0, answered 4: 0 of 239 (0.0%)1frontier: 1frontier 1, answered 0: 0 of 224 (0.0%)frontier 1, answered 1: 39 of 224 (17.4%)39frontier 1, answered 2: 141 of 224 (62.9%)141frontier 1, answered 3: 44 of 224 (19.6%)44frontier 1, answered 4: 0 of 224 (0.0%)2frontier: 2frontier 2, answered 0: 0 of 286 (0.0%)frontier 2, answered 1: 54 of 286 (18.9%)54frontier 2, answered 2: 137 of 286 (47.9%)137frontier 2, answered 3: 95 of 286 (33.2%)95frontier 2, answered 4: 0 of 286 (0.0%)3frontier: 3frontier 3, answered 0: 0 of 139 (0.0%)frontier 3, answered 1: 9 of 139 (6.5%)9frontier 3, answered 2: 46 of 139 (33.1%)46frontier 3, answered 3: 83 of 139 (59.7%)83frontier 3, answered 4: 1 of 139 (0.7%)14frontier: 4frontier 4, answered 0: 0 of 112 (0.0%)frontier 4, answered 1: 1 of 112 (0.9%)1frontier 4, answered 2: 31 of 112 (27.7%)31frontier 4, answered 3: 76 of 112 (67.9%)76frontier 4, answered 4: 4 of 112 (3.6%)4rows: frontier answer; shade: share of the row
von-1.2.0 (calibrated, 1,000 items)
answered →00112233440frontier: 0frontier 0, answered 0: 192 of 239 (80.3%)192frontier 0, answered 1: 3 of 239 (1.3%)3frontier 0, answered 2: 7 of 239 (2.9%)7frontier 0, answered 3: 27 of 239 (11.3%)27frontier 0, answered 4: 10 of 239 (4.2%)101frontier: 1frontier 1, answered 0: 100 of 224 (44.6%)100frontier 1, answered 1: 6 of 224 (2.7%)6frontier 1, answered 2: 22 of 224 (9.8%)22frontier 1, answered 3: 79 of 224 (35.3%)79frontier 1, answered 4: 17 of 224 (7.6%)172frontier: 2frontier 2, answered 0: 76 of 286 (26.6%)76frontier 2, answered 1: 3 of 286 (1.0%)3frontier 2, answered 2: 63 of 286 (22.0%)63frontier 2, answered 3: 139 of 286 (48.6%)139frontier 2, answered 4: 5 of 286 (1.7%)53frontier: 3frontier 3, answered 0: 10 of 139 (7.2%)10frontier 3, answered 1: 2 of 139 (1.4%)2frontier 3, answered 2: 11 of 139 (7.9%)11frontier 3, answered 3: 93 of 139 (66.9%)93frontier 3, answered 4: 23 of 139 (16.5%)234frontier: 4frontier 4, answered 0: 2 of 112 (1.8%)2frontier 4, answered 1: 0 of 112 (0.0%)frontier 4, answered 2: 1 of 112 (0.9%)1frontier 4, answered 3: 25 of 112 (22.3%)25frontier 4, answered 4: 84 of 112 (75.0%)84rows: frontier answer; shade: share of the row
laya-en-ft-tickets (calibrated, 1,000 items)
answered →00112233440frontier: 0frontier 0, answered 0: 107 of 239 (44.8%)107frontier 0, answered 1: 27 of 239 (11.3%)27frontier 0, answered 2: 85 of 239 (35.6%)85frontier 0, answered 3: 20 of 239 (8.4%)20frontier 0, answered 4: 0 of 239 (0.0%)1frontier: 1frontier 1, answered 0: 54 of 224 (24.1%)54frontier 1, answered 1: 42 of 224 (18.8%)42frontier 1, answered 2: 105 of 224 (46.9%)105frontier 1, answered 3: 23 of 224 (10.3%)23frontier 1, answered 4: 0 of 224 (0.0%)2frontier: 2frontier 2, answered 0: 18 of 286 (6.3%)18frontier 2, answered 1: 59 of 286 (20.6%)59frontier 2, answered 2: 143 of 286 (50.0%)143frontier 2, answered 3: 66 of 286 (23.1%)66frontier 2, answered 4: 0 of 286 (0.0%)3frontier: 3frontier 3, answered 0: 11 of 139 (7.9%)11frontier 3, answered 1: 20 of 139 (14.4%)20frontier 3, answered 2: 42 of 139 (30.2%)42frontier 3, answered 3: 66 of 139 (47.5%)66frontier 3, answered 4: 0 of 139 (0.0%)4frontier: 4frontier 4, answered 0: 9 of 112 (8.0%)9frontier 4, answered 1: 29 of 112 (25.9%)29frontier 4, answered 2: 14 of 112 (12.5%)14frontier 4, answered 3: 60 of 112 (53.6%)60frontier 4, answered 4: 0 of 112 (0.0%)rows: frontier answer; shade: share of the row

Rows are the frontier model's answer and columns the model's answer, so the diagonal is correct and everything else is a mistake; darker cells hold a larger share of their row. For an ordered scale, mistakes next to the diagonal are near misses and those far from it are serious; a model whose errors sit far off the diagonal should not be trusted on the extreme levels.

Table view: every cell
ModelFrontier01234
laya-en017450690
laya-en1681235280
laya-en2122071660
laya-en34593730
laya-en46262771
laya-typed-decisions0051174140
laya-typed-decisions1039141440
laya-typed-decisions2054137950
laya-typed-decisions30946831
laya-typed-decisions40131764
von-1.2.00192372710
von-1.2.011006227917
von-1.2.02763631395
von-1.2.03102119323
von-1.2.042012584
laya-en-ft-tickets01072785200
laya-en-ft-tickets15442105230
laya-en-ft-tickets21859143660
laya-en-ft-tickets3112042660
laya-en-ft-tickets492914600

The misses

These are the results that don't flatter Tau: label noise, calibration that helped little, baselines that win, and anything excluded. They are published with the wins so the favourable numbers can be weighed against them.

Run metadata

Date (UTC)
2026-09-27T17:15:20Z
Command
tau report examples/support-tickets/decision.yaml
Git commit
8954e9b095b40f68ee34139ca7c413fad4386300 (uncommitted changes outside this run's outputs)
Workbench
0.1.0
GPU
NVIDIA GeForce RTX 3080 Ti (driver 610.47)
CPU
11th Gen Intel(R) Core(TM) i9-11900K @ 3.50GHz, 16 logical processors
OS and runtime
Microsoft Windows 10.0.19045; .NET 10.0.11
Endpoint
localhost:18093/ (Tau Runtime at localhost:18093/ (4 model(s) installed))
/v1/models: laya-en
sha256 866a05b244e47e96820660d18ee050c518c2de0f60c39e2e9a89dfeb056e1eec, revision 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851
/v1/models: laya-typed-decisions
sha256 2b7a961ac37157d1cfa8105e3283106baf1ba2a5cc30fb3a673253f06aa1e1c5, revision 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851
/v1/models: von-1.2.0
sha256 0777bb988636663b770775ae0b4eb961d6fbee176c9bea1d1da823723b1eb3f3, revision 5df8185a4f2327ad0a7cd117cc4f701ac557b9ae
/v1/models: laya-en-ft-tickets
sha256 9079fee9291e35dc4d3ed8c817a8f1592d97e4413ccbae0539b7a0ee6d9c950a, revision 03677900b97d57980f7c525a84031b9a43cb989e
Measured model laya-en
866a05b244e47e96820660d18ee050c518c2de0f60c39e2e9a89dfeb056e1eec
Measured model laya-typed-decisions
2b7a961ac37157d1cfa8105e3283106baf1ba2a5cc30fb3a673253f06aa1e1c5
Measured model von-1.2.0
0777bb988636663b770775ae0b4eb961d6fbee176c9bea1d1da823723b1eb3f3
Measured model laya-en-ft-tickets
9079fee9291e35dc4d3ed8c817a8f1592d97e4413ccbae0539b7a0ee6d9c950a
Dataset manifest sha256
8aae6044d2ff40ea9392603672ab694b33cca046c70c5a0c532392bb67efb9fb
Dataset revision (cache keys)
ddf1c81a5475
Reference labels
frontier: Scored against the frontier model: every rate is agreement with claude-opus-5-5's v1 answers (after any human override), not accuracy against the dataset's labels (data.reference: frontier).
Frontier model
claude-opus-5-5
Prompt versions
v1, v1-alt
Frontier cache
v1: 2,000 answers, v1-alt: 200 answers
Confidence
Confidence for ECE and thresholds is max(p), the probability of the chosen answer, not the contract's confidence field, which for choice and score questions is not a calibrated probability.

This block ties every figure above to a machine, a model hash, a dataset revision, a date and a command. Re-running the command on the same checkout and models should reproduce the numbers; anything that differs should be explained by what changed here.