Tau Workbench report

Banking77: which of 77 intents is this customer message?

This report measures how far 3 model(s) can be trusted on "Which banking customer-service intent does this message express?", using banking77 data (1,000 held-out items). Reference: the dataset's gold labels; every rate below is accuracy against them. Out of the box, laya-en is 37.2% accurate with a calibration error (ECE) of 0.502; after calibration its ECE is 0.065, 87.1% lower, which meets the 50% goal. At a 5.0% target error, laya-en-ft-banking77 keeps 73.6% of decisions local, the most of any model (von-1.2.0, the least, keeps 4.5%); the cascade is 93.2% accurate against 94.2% for claude-opus-5-5 alone, at an estimated £1,031 per million decisions against £3,898 (list prices, 2026-09-27). The classic baseline minilm-l6-banking77 (not served through Tau) does better, keeping 94.7% local at 94.2% blended accuracy. The frontier model disagrees with the gold labels on 5.8% of 1,000 items, which bounds how much any accuracy figure here can be trusted.

Generated 2026-09-27T17:15:19Z (UTC) on NVIDIA GeForce RTX 3080 Ti. Every number below is in report.json beside this file; the metadata block at the end says how to reproduce it.

Dataset

Name
banking77
Source
PolyAI-LDN/task-specific-datasets
Pinned revision
9d081458ff52e53cf7e848f414e6e9344e4e6696
Licence
CC-BY-4.0
Synthetic
No
Fine-tune split
9,003 items
Calibration split
1,000 items
Held-out split
1,000 items
Question
intent (choice, 77 answers): Which banking customer-service intent does this message express?

Calibrators and thresholds are fitted on the calibration split only and every result is judged on the held-out split, which nothing was tuned on. The splits are fixed by the dataset manifest (seed 42), and each file's sha256 is checked before use.

Accuracy and calibration, before and after

ModelItemsAccuracy rawAccuracy calibratedECE rawECE calibratedECE changeBrier raw → cal.Log loss raw → cal.Calibrated view
laya-en1,00037.2%37.2%0.5020.065−87.1%1.096 → 0.8129.841 → 2.906Runtime with calibrators loaded; matches the Workbench's offline calibration to within 1.2e-1 in any probability
von-1.2.01,00077.1%77.0%0.1850.039−78.9%0.411 → 0.3543.723 → 1.078Runtime with calibrators loaded; matches the Workbench's offline calibration to within 1.5e-1 in any probability
laya-en-ft-banking771,00087.3%87.1%0.0710.061−13.6%0.221 → 0.2100.702 → 0.721Runtime with calibrators loaded; matches the Workbench's offline calibration to within 2.3e-2 in any probability

Held-out items only. ECE (expected calibration error) is the average gap between how confident a model is and how often it is right, over 15 confidence bins; lower is better and 0 is perfect. Confidence for ECE and thresholds is max(p), the probability of the chosen answer, not the contract's confidence field, which for choice and score questions is not a calibrated probability. The goal is a 50% cut in ECE. The largest cut is laya-en's (87.1%) and the smallest is laya-en-ft-banking77's (13.6%); a model whose confidence is still off after calibration should not be trusted to gate decisions on its own.

Reliability diagrams

RawCalibratedPerfect calibration
laya-en
0.000.250.500.751.000.000.250.500.751.00Perfect calibration: accuracy equals confidenceRaw: confidence 0.00 to 0.07, 5 items, mean confidence 0.059, accuracy 0.000Raw: confidence 0.07 to 0.13, 14 items, mean confidence 0.098, accuracy 0.071Raw: confidence 0.13 to 0.20, 10 items, mean confidence 0.176, accuracy 0.100Raw: confidence 0.20 to 0.27, 10 items, mean confidence 0.238, accuracy 0.200Raw: confidence 0.27 to 0.33, 10 items, mean confidence 0.308, accuracy 0.200Raw: confidence 0.33 to 0.40, 10 items, mean confidence 0.359, accuracy 0.100Raw: confidence 0.40 to 0.47, 15 items, mean confidence 0.437, accuracy 0.067Raw: confidence 0.47 to 0.53, 32 items, mean confidence 0.500, accuracy 0.188Raw: confidence 0.53 to 0.60, 23 items, mean confidence 0.563, accuracy 0.087Raw: confidence 0.60 to 0.67, 29 items, mean confidence 0.637, accuracy 0.172Raw: confidence 0.67 to 0.73, 25 items, mean confidence 0.697, accuracy 0.280Raw: confidence 0.73 to 0.80, 39 items, mean confidence 0.773, accuracy 0.179Raw: confidence 0.80 to 0.87, 45 items, mean confidence 0.832, accuracy 0.200Raw: confidence 0.87 to 0.93, 64 items, mean confidence 0.902, accuracy 0.219Raw: confidence 0.93 to 1.00, 669 items, mean confidence 0.993, accuracy 0.469Calibrated: confidence 0.00 to 0.07, 56 items, mean confidence 0.041, accuracy 0.125Calibrated: confidence 0.07 to 0.13, 75 items, mean confidence 0.103, accuracy 0.133Calibrated: confidence 0.13 to 0.20, 70 items, mean confidence 0.167, accuracy 0.243Calibrated: confidence 0.20 to 0.27, 79 items, mean confidence 0.238, accuracy 0.203Calibrated: confidence 0.27 to 0.33, 79 items, mean confidence 0.301, accuracy 0.241Calibrated: confidence 0.33 to 0.40, 97 items, mean confidence 0.368, accuracy 0.371Calibrated: confidence 0.40 to 0.47, 95 items, mean confidence 0.433, accuracy 0.284Calibrated: confidence 0.47 to 0.53, 156 items, mean confidence 0.505, accuracy 0.410Calibrated: confidence 0.53 to 0.60, 293 items, mean confidence 0.547, accuracy 0.601Confidence, max(p)Accuracy
von-1.2.0
0.000.250.500.751.000.000.250.500.751.00Perfect calibration: accuracy equals confidenceRaw: confidence 0.40 to 0.47, 1 items, mean confidence 0.415, accuracy 0.000Raw: confidence 0.47 to 0.53, 8 items, mean confidence 0.520, accuracy 0.250Raw: confidence 0.53 to 0.60, 17 items, mean confidence 0.568, accuracy 0.471Raw: confidence 0.60 to 0.67, 14 items, mean confidence 0.643, accuracy 0.500Raw: confidence 0.67 to 0.73, 19 items, mean confidence 0.692, accuracy 0.526Raw: confidence 0.73 to 0.80, 28 items, mean confidence 0.773, accuracy 0.357Raw: confidence 0.80 to 0.87, 36 items, mean confidence 0.835, accuracy 0.500Raw: confidence 0.87 to 0.93, 47 items, mean confidence 0.896, accuracy 0.574Raw: confidence 0.93 to 1.00, 830 items, mean confidence 0.995, accuracy 0.830Calibrated: confidence 0.00 to 0.07, 1 items, mean confidence 0.051, accuracy 0.000Calibrated: confidence 0.20 to 0.27, 6 items, mean confidence 0.248, accuracy 0.167Calibrated: confidence 0.27 to 0.33, 10 items, mean confidence 0.310, accuracy 0.500Calibrated: confidence 0.33 to 0.40, 25 items, mean confidence 0.364, accuracy 0.400Calibrated: confidence 0.40 to 0.47, 50 items, mean confidence 0.439, accuracy 0.340Calibrated: confidence 0.47 to 0.53, 45 items, mean confidence 0.504, accuracy 0.556Calibrated: confidence 0.53 to 0.60, 54 items, mean confidence 0.566, accuracy 0.593Calibrated: confidence 0.60 to 0.67, 93 items, mean confidence 0.629, accuracy 0.699Calibrated: confidence 0.67 to 0.73, 21 items, mean confidence 0.695, accuracy 0.667Calibrated: confidence 0.73 to 0.80, 139 items, mean confidence 0.762, accuracy 0.719Calibrated: confidence 0.80 to 0.87, 37 items, mean confidence 0.854, accuracy 0.730Calibrated: confidence 0.87 to 0.93, 492 items, mean confidence 0.893, accuracy 0.909Calibrated: confidence 0.93 to 1.00, 27 items, mean confidence 0.938, accuracy 1.000Confidence, max(p)Accuracy
laya-en-ft-banking77
0.000.250.500.751.000.000.250.500.751.00Perfect calibration: accuracy equals confidenceRaw: confidence 0.00 to 0.07, 21 items, mean confidence 0.055, accuracy 0.524Raw: confidence 0.07 to 0.13, 29 items, mean confidence 0.095, accuracy 0.483Raw: confidence 0.13 to 0.20, 17 items, mean confidence 0.168, accuracy 0.588Raw: confidence 0.20 to 0.27, 5 items, mean confidence 0.232, accuracy 0.000Raw: confidence 0.27 to 0.33, 8 items, mean confidence 0.307, accuracy 0.500Raw: confidence 0.33 to 0.40, 8 items, mean confidence 0.358, accuracy 0.625Raw: confidence 0.40 to 0.47, 5 items, mean confidence 0.425, accuracy 0.200Raw: confidence 0.47 to 0.53, 5 items, mean confidence 0.505, accuracy 0.400Raw: confidence 0.53 to 0.60, 3 items, mean confidence 0.568, accuracy 0.667Raw: confidence 0.60 to 0.67, 3 items, mean confidence 0.630, accuracy 0.667Raw: confidence 0.67 to 0.73, 8 items, mean confidence 0.702, accuracy 0.500Raw: confidence 0.73 to 0.80, 8 items, mean confidence 0.769, accuracy 0.625Raw: confidence 0.80 to 0.87, 11 items, mean confidence 0.839, accuracy 0.727Raw: confidence 0.87 to 0.93, 54 items, mean confidence 0.904, accuracy 0.574Raw: confidence 0.93 to 1.00, 815 items, mean confidence 0.966, accuracy 0.950Calibrated: confidence 0.00 to 0.07, 1 items, mean confidence 0.039, accuracy 0.000Calibrated: confidence 0.07 to 0.13, 1 items, mean confidence 0.122, accuracy 0.000Calibrated: confidence 0.13 to 0.20, 32 items, mean confidence 0.165, accuracy 0.375Calibrated: confidence 0.20 to 0.27, 16 items, mean confidence 0.240, accuracy 0.688Calibrated: confidence 0.27 to 0.33, 17 items, mean confidence 0.309, accuracy 0.529Calibrated: confidence 0.33 to 0.40, 19 items, mean confidence 0.373, accuracy 0.474Calibrated: confidence 0.40 to 0.47, 7 items, mean confidence 0.413, accuracy 0.429Calibrated: confidence 0.47 to 0.53, 6 items, mean confidence 0.498, accuracy 0.333Calibrated: confidence 0.53 to 0.60, 16 items, mean confidence 0.568, accuracy 0.625Calibrated: confidence 0.60 to 0.67, 8 items, mean confidence 0.632, accuracy 0.625Calibrated: confidence 0.67 to 0.73, 18 items, mean confidence 0.704, accuracy 0.500Calibrated: confidence 0.73 to 0.80, 10 items, mean confidence 0.768, accuracy 0.400Calibrated: confidence 0.80 to 0.87, 16 items, mean confidence 0.846, accuracy 0.813Calibrated: confidence 0.87 to 0.93, 44 items, mean confidence 0.910, accuracy 0.795Calibrated: confidence 0.93 to 1.00, 789 items, mean confidence 0.983, accuracy 0.949Confidence, max(p)Accuracy

Each point is one confidence bin on the held-out split: how confident the model was (across) against how often it was right (up). Points below the diagonal are overconfident, points above it underconfident. After calibration the points should sit close to the diagonal; where they don't, a threshold on that model's confidence will not deliver the accuracy it promises.

Table view: every bin
ModelBinRaw itemsRaw confidenceRaw accuracyCal. itemsCal. confidenceCal. accuracy
laya-en0.00–0.0750.0590.000560.0410.125
laya-en0.07–0.13140.0980.071750.1030.133
laya-en0.13–0.20100.1760.100700.1670.243
laya-en0.20–0.27100.2380.200790.2380.203
laya-en0.27–0.33100.3080.200790.3010.241
laya-en0.33–0.40100.3590.100970.3680.371
laya-en0.40–0.47150.4370.067950.4330.284
laya-en0.47–0.53320.5000.1881560.5050.410
laya-en0.53–0.60230.5630.0872930.5470.601
laya-en0.60–0.67290.6370.1720n/an/a
laya-en0.67–0.73250.6970.2800n/an/a
laya-en0.73–0.80390.7730.1790n/an/a
laya-en0.80–0.87450.8320.2000n/an/a
laya-en0.87–0.93640.9020.2190n/an/a
laya-en0.93–1.006690.9930.4690n/an/a
von-1.2.00.00–0.070n/an/a10.0510.000
von-1.2.00.20–0.270n/an/a60.2480.167
von-1.2.00.27–0.330n/an/a100.3100.500
von-1.2.00.33–0.400n/an/a250.3640.400
von-1.2.00.40–0.4710.4150.000500.4390.340
von-1.2.00.47–0.5380.5200.250450.5040.556
von-1.2.00.53–0.60170.5680.471540.5660.593
von-1.2.00.60–0.67140.6430.500930.6290.699
von-1.2.00.67–0.73190.6920.526210.6950.667
von-1.2.00.73–0.80280.7730.3571390.7620.719
von-1.2.00.80–0.87360.8350.500370.8540.730
von-1.2.00.87–0.93470.8960.5744920.8930.909
von-1.2.00.93–1.008300.9950.830270.9381.000
laya-en-ft-banking770.00–0.07210.0550.52410.0390.000
laya-en-ft-banking770.07–0.13290.0950.48310.1220.000
laya-en-ft-banking770.13–0.20170.1680.588320.1650.375
laya-en-ft-banking770.20–0.2750.2320.000160.2400.688
laya-en-ft-banking770.27–0.3380.3070.500170.3090.529
laya-en-ft-banking770.33–0.4080.3580.625190.3730.474
laya-en-ft-banking770.40–0.4750.4250.20070.4130.429
laya-en-ft-banking770.47–0.5350.5050.40060.4980.333
laya-en-ft-banking770.53–0.6030.5680.667160.5680.625
laya-en-ft-banking770.60–0.6730.6300.66780.6320.625
laya-en-ft-banking770.67–0.7380.7020.500180.7040.500
laya-en-ft-banking770.73–0.8080.7690.625100.7680.400
laya-en-ft-banking770.80–0.87110.8390.727160.8460.813
laya-en-ft-banking770.87–0.93540.9040.574440.9100.795
laya-en-ft-banking770.93–1.008150.9660.9507890.9830.949

Calibrators fitted

ModelScopeItemsChosenTemperature TTemperature ECE / log lossIsotonic knotsIsotonic ECE / log lossBefore: ECE / log loss
laya-enquestion type1,000temperature3.0490.095 / 2.977860.080 / 2.9980.500 / 9.808
laya-en11+ options1,000temperature3.0490.095 / 2.977860.080 / 2.9980.500 / 9.808
von-1.2.0question type1,000isotonic2.2560.049 / 1.1701000.038 / 1.1160.205 / 3.584
von-1.2.011+ options1,000isotonic2.2560.049 / 1.1701000.038 / 1.1160.205 / 3.584
laya-en-ft-banking77question type1,000isotonic1.0180.064 / 0.761960.065 / 0.6550.065 / 0.761
laya-en-ft-banking7711+ options1,000isotonic1.0180.064 / 0.761960.065 / 0.6550.065 / 0.761

Both methods are fitted on the calibration split's raw probabilities, and the one with the lower calibration-split log loss is written for the Runtime to load; the scores in this table are on the calibration split, so they show the fit, not the result (the held-out result is in the first table). Isotonic regression needs at least 200 items and falls back to temperature scaling below that.

How much can stay local

laya-envon-1.2.0laya-en-ft-banking77minilm-l6-banking77 (baseline)Target accuracy
30%44%58%72%86%100%0%20%40%60%80%100%Target: 95.0% accuracy on kept itemstarget 95%laya-enlaya-en: τ = 0.00, keeps 100.0% local, 37.2% accuracy on thoselaya-en: τ = 0.05, keeps 96.2% local, 38.3% accuracy on thoselaya-en: τ = 0.10, keeps 91.4% local, 39.5% accuracy on thoselaya-en: τ = 0.15, keeps 86.3% local, 41.1% accuracy on thoselaya-en: τ = 0.20, keeps 82.1% local, 41.8% accuracy on thoselaya-en: τ = 0.25, keeps 77.7% local, 43.5% accuracy on thoselaya-en: τ = 0.30, keeps 72.9% local, 45.0% accuracy on thoselaya-en: τ = 0.35, keeps 66.5% local, 46.9% accuracy on thoselaya-en: τ = 0.40, keeps 59.5% local, 48.4% accuracy on thoselaya-en: τ = 0.45, keeps 51.6% local, 51.4% accuracy on thoselaya-en: τ = 0.50, keeps 42.5% local, 54.1% accuracy on thosevon-1.2.0von-1.2.0: τ = 0.00, keeps 100.0% local, 77.0% accuracy on thosevon-1.2.0: τ = 0.05, keeps 100.0% local, 77.0% accuracy on thosevon-1.2.0: τ = 0.10, keeps 99.9% local, 77.1% accuracy on thosevon-1.2.0: τ = 0.15, keeps 99.9% local, 77.1% accuracy on thosevon-1.2.0: τ = 0.20, keeps 99.8% local, 77.2% accuracy on thosevon-1.2.0: τ = 0.25, keeps 99.6% local, 77.3% accuracy on thosevon-1.2.0: τ = 0.30, keeps 99.1% local, 77.5% accuracy on thosevon-1.2.0: τ = 0.35, keeps 97.5% local, 78.3% accuracy on thosevon-1.2.0: τ = 0.40, keeps 95.7% local, 78.7% accuracy on thosevon-1.2.0: τ = 0.45, keeps 91.4% local, 80.5% accuracy on thosevon-1.2.0: τ = 0.50, keeps 90.0% local, 81.2% accuracy on thosevon-1.2.0: τ = 0.55, keeps 86.6% local, 82.4% accuracy on thosevon-1.2.0: τ = 0.60, keeps 80.7% local, 84.0% accuracy on thosevon-1.2.0: τ = 0.65, keeps 72.3% local, 85.8% accuracy on thosevon-1.2.0: τ = 0.70, keeps 70.7% local, 86.0% accuracy on thosevon-1.2.0: τ = 0.75, keeps 70.0% local, 86.3% accuracy on thosevon-1.2.0: τ = 0.80, keeps 56.1% local, 89.8% accuracy on thosevon-1.2.0: τ = 0.85, keeps 52.8% local, 90.9% accuracy on thosevon-1.2.0: τ = 0.90, keeps 17.2% local, 96.5% accuracy on thosevon-1.2.0: chosen τ = 0.92; held-out keeps 4.5% local at 100.0% accuracyτ 0.92laya-en-ft-banking77laya-en-ft-banking77: τ = 0.00, keeps 100.0% local, 87.1% accuracy on thoselaya-en-ft-banking77: τ = 0.05, keeps 99.9% local, 87.2% accuracy on thoselaya-en-ft-banking77: τ = 0.10, keeps 99.9% local, 87.2% accuracy on thoselaya-en-ft-banking77: τ = 0.15, keeps 99.0% local, 87.9% accuracy on thoselaya-en-ft-banking77: τ = 0.20, keeps 96.6% local, 88.9% accuracy on thoselaya-en-ft-banking77: τ = 0.25, keeps 95.6% local, 89.1% accuracy on thoselaya-en-ft-banking77: τ = 0.30, keeps 94.6% local, 89.3% accuracy on thoselaya-en-ft-banking77: τ = 0.35, keeps 92.9% local, 90.0% accuracy on thoselaya-en-ft-banking77: τ = 0.40, keeps 91.4% local, 90.8% accuracy on thoselaya-en-ft-banking77: τ = 0.45, keeps 90.7% local, 91.2% accuracy on thoselaya-en-ft-banking77: τ = 0.50, keeps 90.4% local, 91.4% accuracy on thoselaya-en-ft-banking77: τ = 0.55, keeps 89.9% local, 91.7% accuracy on thoselaya-en-ft-banking77: τ = 0.60, keeps 88.5% local, 92.1% accuracy on thoselaya-en-ft-banking77: τ = 0.65, keeps 87.9% local, 92.4% accuracy on thoselaya-en-ft-banking77: τ = 0.70, keeps 87.2% local, 92.5% accuracy on thoselaya-en-ft-banking77: τ = 0.75, keeps 85.8% local, 93.2% accuracy on thoselaya-en-ft-banking77: τ = 0.80, keeps 84.9% local, 93.9% accuracy on thoselaya-en-ft-banking77: τ = 0.85, keeps 84.1% local, 93.8% accuracy on thoselaya-en-ft-banking77: τ = 0.90, keeps 81.9% local, 94.6% accuracy on thoselaya-en-ft-banking77: τ = 0.95, keeps 73.6% local, 96.2% accuracy on thoselaya-en-ft-banking77: τ = 1.00, keeps 2.0% local, 100.0% accuracy on thoselaya-en-ft-banking77: chosen τ = 0.95; held-out keeps 73.6% local at 96.2% accuracyτ 0.95minilm-l6-banking77 (baseline)minilm-l6-banking77: τ = 0.00, keeps 100.0% local, 91.7% accuracy on thoseminilm-l6-banking77: τ = 0.05, keeps 100.0% local, 91.7% accuracy on thoseminilm-l6-banking77: τ = 0.10, keeps 100.0% local, 91.7% accuracy on thoseminilm-l6-banking77: τ = 0.15, keeps 99.9% local, 91.8% accuracy on thoseminilm-l6-banking77: τ = 0.20, keeps 99.7% local, 91.9% accuracy on thoseminilm-l6-banking77: τ = 0.25, keeps 99.4% local, 92.1% accuracy on thoseminilm-l6-banking77: τ = 0.30, keeps 98.9% local, 92.3% accuracy on thoseminilm-l6-banking77: τ = 0.35, keeps 98.7% local, 92.4% accuracy on thoseminilm-l6-banking77: τ = 0.40, keeps 98.4% local, 92.6% accuracy on thoseminilm-l6-banking77: τ = 0.45, keeps 97.5% local, 93.0% accuracy on thoseminilm-l6-banking77: τ = 0.50, keeps 96.5% local, 93.6% accuracy on thoseminilm-l6-banking77: τ = 0.55, keeps 95.6% local, 94.2% accuracy on thoseminilm-l6-banking77: τ = 0.60, keeps 94.5% local, 94.8% accuracy on thoseminilm-l6-banking77: τ = 0.65, keeps 92.6% local, 95.4% accuracy on thoseminilm-l6-banking77: τ = 0.70, keeps 92.0% local, 95.8% accuracy on thoseminilm-l6-banking77: τ = 0.75, keeps 91.3% local, 95.9% accuracy on thoseminilm-l6-banking77: τ = 0.80, keeps 89.9% local, 96.4% accuracy on thoseminilm-l6-banking77: τ = 0.85, keeps 87.9% local, 97.2% accuracy on thoseminilm-l6-banking77: τ = 0.90, keeps 82.6% local, 97.9% accuracy on thoseminilm-l6-banking77: τ = 0.95, keeps 72.0% local, 98.3% accuracy on thoseminilm-l6-banking77: chosen τ = 0.58; held-out keeps 94.7% local at 94.7% accuracyτ 0.58Share of decisions kept local (confidence ≥ τ)Accuracy on kept items

Each line shows, on held-out items, what happens as the confidence threshold τ rises: fewer decisions are kept local (moving left) and those kept are more often right (moving up). The dot marks the τ chosen on the calibration split for a 5.0% target error. A line that never reaches the target line means no threshold makes that model safe enough on its own at that target. Thresholds that keep fewer than 10 items are left off the chart because a handful of items says little; every point is in report.json. Dashed lines are baselines, not served through Tau, put through exactly the same threshold rule on their own calibration lines.

ModelConfidences fromτKept local (calibration)Kept local (held-out)Accuracy on kept (held-out)Result
laya-enofflinen/an/an/an/aNo threshold meets the 5.0% target error on the calibration split. The lowest error any threshold reaches is 37.9%, at τ = 0.54, accepting 28.2% of items. No threshold is invented.
von-1.2.0offline0.924.0%4.5%100.0%τ = 0.92 is the smallest threshold whose calibration-split error is at most 5.0%. On held-out it accepts 4.5% of items with 100.0% accuracy.
laya-en-ft-banking77offline0.9571.9%73.6%96.2%τ = 0.95 is the smallest threshold whose calibration-split error is at most 5.0%. On held-out it accepts 73.6% of items with 96.2% accuracy.
minilm-l6-banking77 (baseline)baseline-calibrated0.5894.6%94.7%94.7%τ = 0.58 is the smallest threshold whose calibration-split error is at most 5.0%. On held-out it accepts 94.7% of items with 94.7% accuracy.

Cascade: local first, frontier for the rest

ModelτKept localEscalatedNo frontier answerLocal onlyFrontier onlyCascadeLocal p50 latency
laya-enn/an/an/an/a37.2% (1,000)94.2% (1,000)No cascade: no threshold met the 5.0% target error on the calibration split.96.9 ms
von-1.2.00.924.5%95.5%077.0% (1,000)94.2% (1,000)94.2% (1,000)91.8 ms
laya-en-ft-banking770.9573.6%26.4%087.1% (1,000)94.2% (1,000)93.2% (1,000)244.0 ms
minilm-l6-banking77 (baseline, not served through Tau: local latency and energy not measured)0.5894.7%5.3%091.7% (1,000)94.2% (1,000)94.2% (1,000)not measured

Accuracy is on held-out items, with the number of items each figure covers in brackets. The cascade keeps the local answer when its confidence is at least τ and uses the frontier model's cached answer otherwise; an escalated item with no cached answer is counted and left out, never guessed. If the cascade is close to frontier-only accuracy while keeping a large share local, most of the frontier cost can be avoided. Frontier latency is not measured: no API call was made, so the latency mix covers the local side only.

Cost for laya-en Estimates, not measured bills

Price basisFrontier only, per million (USD)Frontier only, per 1,000Frontier only, per millionCascade, per 1,000Cascade, per millionSaving, per million
claude-opus-5-5 at list price (the model that produced the frontier answers)$5,165£3.90£3,898n/an/an/a
claude-opus-5-5-batch: the same model at a different rate$2,583£1.95£1,949n/an/an/a
claude-sonnet-5: price only, if escalations went to this model; its accuracy was not measured$2,583£1.95£1,949n/an/an/a
claude-haiku-4-5: price only, if escalations went to this model; its accuracy was not measured$1,291£0.9744£974n/an/an/a

No cascade: no threshold met the target, so there is nothing to price.

Every figure in this table is an estimate from list prices and estimated token counts, not a measured bill. It shows what the same decisions would cost through the API, and what share of that the cascade avoids. Basis:

Cost for von-1.2.0 Estimates, not measured bills

Price basisFrontier only, per million (USD)Frontier only, per 1,000Frontier only, per millionCascade, per 1,000Cascade, per millionSaving, per million
claude-opus-5-5 at list price (the model that produced the frontier answers)$5,165£3.90£3,898£3.72£3,723£175
claude-opus-5-5-batch: the same model at a different rate$2,583£1.95£1,949£1.86£1,862£87.13
claude-sonnet-5: price only, if escalations went to this model; its accuracy was not measured$2,583£1.95£1,949£1.86£1,862£87.13
claude-haiku-4-5: price only, if escalations went to this model; its accuracy was not measured$1,291£0.9744£974£0.9312£931£43.28

Every figure in this table is an estimate from list prices and estimated token counts, not a measured bill. It shows what the same decisions would cost through the API, and what share of that the cascade avoids. Basis:

Cost for laya-en-ft-banking77 Estimates, not measured bills

Price basisFrontier only, per million (USD)Frontier only, per 1,000Frontier only, per millionCascade, per 1,000Cascade, per millionSaving, per million
claude-opus-5-5 at list price (the model that produced the frontier answers)$5,165£3.90£3,898£1.03£1,031£2,867
claude-opus-5-5-batch: the same model at a different rate$2,583£1.95£1,949£0.5160£516£1,433
claude-sonnet-5: price only, if escalations went to this model; its accuracy was not measured$2,583£1.95£1,949£0.5160£516£1,433
claude-haiku-4-5: price only, if escalations went to this model; its accuracy was not measured$1,291£0.9744£974£0.2588£259£716

Every figure in this table is an estimate from list prices and estimated token counts, not a measured bill. It shows what the same decisions would cost through the API, and what share of that the cascade avoids. Basis:

Cost for minilm-l6-banking77 (baseline, not served through Tau: local latency and energy not measured) Estimates, not measured bills

Price basisFrontier only, per million (USD)Frontier only, per 1,000Frontier only, per millionCascade, per 1,000Cascade, per millionSaving, per million
claude-opus-5-5 at list price (the model that produced the frontier answers)$5,165£3.90£3,898£0.2066£207£3,691
claude-opus-5-5-batch: the same model at a different rate$2,583£1.95£1,949£0.1033£103£1,846
claude-sonnet-5: price only, if escalations went to this model; its accuracy was not measured$2,583£1.95£1,949£0.1033£103£1,846
claude-haiku-4-5: price only, if escalations went to this model; its accuracy was not measured$1,291£0.9744£974£0.0516£51.65£923

Every figure in this table is an estimate from list prices and estimated token counts, not a measured bill. It shows what the same decisions would cost through the API, and what share of that the cascade avoids. Basis:

Frontier labels: noise and agreement

5.8%
disagree with gold (58 of 1,000 items)
97.5%
agreement between v1 and v1-alt wordings (195 of 200)
1,000
v1 answers cached of 1,000 eligible held-out items; 0 pending
0
labels overridden by a person; 0 answer line(s) rejected

Label noise is how often claude-opus-5-5 disagrees with the dataset's gold label on the same item; it is measured before any human override. Some of that disagreement is the frontier model's error and some is the gold label's, so it bounds how precisely any accuracy here can be read. Agreement between two wordings of the prompt shows how much the frontier answers depend on phrasing. Answers came from claude-code-session: subagent answer sheet; the prompt's fixed part shown once per batch of up to 200 items, each item judged on its own; no API, dated 2026-09-27 to 2026-09-27; no paid API call was made.

Baselines

ModelHeld-out itemsAccuracyECE rawLog loss rawECE calibratedLog loss calibratedNote
minilm-l6-banking77 (baseline)1,00091.5%0.0300.3690.0240.394Calibrated with isotonic fitted on the file's 1000 calibration lines
laya-en (Tau)1,00037.2%0.5029.8410.0652.906for comparison
von-1.2.0 (Tau)1,00077.1%0.1853.7230.0391.078for comparison
laya-en-ft-banking77 (Tau)1,00087.3%0.0710.7020.0610.721for comparison

A classic small encoder, fine-tuned on the same training split, scored with exactly the same metric code as the Tau models. It is here to test the claim that a fixed task favours classic fine-tuning: where it beats a Tau model, that is listed under the misses, and it is a reason to prefer the classic model for this task.

Where the models go wrong

laya-en: most frequent confusions (calibrated, 1,000 items)

Gold answerModel's answerItems
card_about_to_expiredisposable_card_limits13
top_up_limitstop_up_reverted13
getting_virtual_cardvirtual_card_not_working12
card_swallowedatm_support11
get_disposable_virtual_cardvirtual_card_not_working11
unable_to_verify_identityverify_my_identity11
why_verify_identityverify_my_identity11
card_acceptancesupported_cards_and_currencies10
card_payment_wrong_exchange_ratewrong_exchange_rate_for_cash_withdrawal10
declined_card_paymentdeclined_transfer10
pending_top_uptop_up_reverted10
top_up_failedtop_up_reverted9
topping_up_by_cardtop_up_reverted9
visa_or_mastercardsupported_cards_and_currencies9
beneficiary_not_alloweddeclined_transfer7

von-1.2.0: most frequent confusions (calibrated, 1,000 items)

Gold answerModel's answerItems
beneficiary_not_allowedfailed_transfer9
declined_card_paymentdeclined_transfer8
why_verify_identityverify_my_identity7
wrong_exchange_rate_for_cash_withdrawalcash_withdrawal_charge6
Refund_not_showing_uprequest_refund5
top_up_by_bank_transfer_chargetransfer_into_account5
balance_not_updated_after_bank_transferfailed_transfer4
disposable_card_limitsgetting_spare_card4
edit_personal_detailsexchange_via_app4
balance_not_updated_after_cheque_or_cash_deposittop_up_by_cash_or_cheque3
cash_withdrawal_not_recogniseddeclined_cash_withdrawal3
direct_debit_payment_not_recognisedextra_charge_on_statement3
extra_charge_on_statementcard_payment_fee_charged3
get_disposable_virtual_cardgetting_virtual_card3
order_physical_cardcard_delivery_estimate3

laya-en-ft-banking77: most frequent confusions (calibrated, 1,000 items)

Gold answerModel's answerItems
beneficiary_not_allowedfailed_transfer4
wrong_exchange_rate_for_cash_withdrawalcash_withdrawal_charge3
balance_not_updated_after_cheque_or_cash_deposittop_up_by_cash_or_cheque2
card_arrivalcard_delivery_estimate2
compromised_cardterminate_account2
disposable_card_limitsgetting_spare_card2
fiat_currency_supportcountry_support2
getting_spare_cardage_limit2
order_physical_cardget_physical_card2
supported_cards_and_currenciescard_acceptance2
top_up_by_bank_transfer_chargereceiving_money2
top_up_by_cash_or_chequetopping_up_by_card2
transfer_into_accounttop_up_by_bank_transfer_charge2
transfer_not_received_by_recipientpending_transfer2
transfer_timingbalance_not_updated_after_bank_transfer2

With 77 answers a full matrix is too large to read, so these tables list the 15 most frequent mistakes per model. Pairs that recur across models are usually answers whose descriptions overlap, which is where clearer option descriptions would help most.

The misses

These are the results that don't flatter Tau: label noise, calibration that helped little, baselines that win, and anything excluded. They are published with the wins so the favourable numbers can be weighed against them.

Run metadata

Date (UTC)
2026-09-27T17:15:19Z
Command
tau report examples/banking77/decision.yaml
Git commit
8954e9b095b40f68ee34139ca7c413fad4386300 (clean, outputs excluded)
Workbench
0.1.0
GPU
NVIDIA GeForce RTX 3080 Ti (driver 610.47)
CPU
11th Gen Intel(R) Core(TM) i9-11900K @ 3.50GHz, 16 logical processors
OS and runtime
Microsoft Windows 10.0.19045; .NET 10.0.11
Endpoint
localhost:18093/ (Tau Runtime at localhost:18093/ (3 model(s) installed))
/v1/models: laya-en
sha256 866a05b244e47e96820660d18ee050c518c2de0f60c39e2e9a89dfeb056e1eec, revision 55cf4c4ebb4ebe31b2550e8bdf3bd21b99753851
/v1/models: von-1.2.0
sha256 0777bb988636663b770775ae0b4eb961d6fbee176c9bea1d1da823723b1eb3f3, revision 5df8185a4f2327ad0a7cd117cc4f701ac557b9ae
/v1/models: laya-en-ft-banking77
sha256 198e563957ce43348f944a10739bdf17929bc3d2175f22a72239b4601d421556, revision 469442c64b974d12d1c397c6fab2baeeb1cdd8cc
Measured model laya-en
866a05b244e47e96820660d18ee050c518c2de0f60c39e2e9a89dfeb056e1eec
Measured model von-1.2.0
0777bb988636663b770775ae0b4eb961d6fbee176c9bea1d1da823723b1eb3f3
Measured model laya-en-ft-banking77
198e563957ce43348f944a10739bdf17929bc3d2175f22a72239b4601d421556
Dataset manifest sha256
7ff6d7b0974afb66a781fccbeae9f58db38a37ebd2a0398792d5f2550ac6c79e
Dataset revision (cache keys)
9d081458ff52
Reference labels
gold: Scored against the dataset's gold labels (data.reference: gold).
Frontier model
claude-opus-5-5
Prompt versions
v1, v1-alt
Frontier cache
v1: 1,000 answers, v1-alt: 200 answers
Confidence
Confidence for ECE and thresholds is max(p), the probability of the chosen answer, not the contract's confidence field, which for choice and score questions is not a calibrated probability.

This block ties every figure above to a machine, a model hash, a dataset revision, a date and a command. Re-running the command on the same checkout and models should reproduce the numbers; anything that differs should be explained by what changed here.