Most vendor numbers are a PDF from a launch date. This page is generated from the configuration we are actually running, so it cannot quietly outlive it.
For every path through the gateway: what good means on that path, whether a real oracle checked it or nothing did, what we measured, and whether that measurement still describes the system serving your request. When it does not, we strike it and say so.
Oracle: Real. The emitted call is checked against the tool schema.
Emits a schema-valid call, and declines rather than inventing one when the schema is ambiguous.
749 of 800 on a single fixed pass of BFCL v4 Python, through the production tool path
An official AST re-score of the same saved outputs gives 751 of 800. Single pass: no multi-seed and no bootstrap interval. Wrong tool 1 of 800. Retries 0 of 800. These are 800 fixed benchmark items replayed through the production path, so it is an artifact measurement and not an operational number from organic traffic. Not a leaderboard submission.
Adversarially reviewed. Adversarially reviewed 2026-07-30; publishable in this n-of-N single-pass form only.
config premise holds against live config (ledger L28) · adversarially reviewed 2026-07-30; publishable in this n-of-N single-pass form only
Watch it decline a trap →Oracle: Real. A wrong tool served is a hard, countable failure.
Serves no wrong tool when the request is built to induce one.
200 of 200 on the adversarial trap set, replayed through production
Zero wrong tools served, across four adversarial shapes parameterised 50 ways each. The four shapes: a forecast question that tempts a current-weather tool, a trip question that tempts the origin city instead of the destination, a mass conversion that tempts a currency tool, and a tool-shaped question whose correct answer is no call at all. Corrected 2026-08-06: this section previously reported a 95% confidence interval computed on the 200 rows, and described them as a strict superset of the historical 40. Both were withdrawn on our own review, because raising the replicate count adds rows and no new shapes, and we measured the replicates behaving as one unit. We publish the count and the shape coverage, and no rate.
This is a primary-path result. See Failover below for where it does not hold.
Adversarially reviewed. Adversarially reviewed 2026-07-30; its objection was the interval label, and the interval itself was struck 2026-08-06 when the denominator turned out to be four shapes, not 200.
ledger row L25 declares no PREMISE · adversarially reviewed 2026-07-30; its objection was the interval label, and the interval itself was struck 2026-08-06 when the denominator turned out to be four shapes, not 200
Run the trap set →Oracle: Real. Your tests are executed in a sandbox and must pass.
Only code that passes the tests you supplied is served.
92.1% HumanEval+, cache-free (0 cache serves)
Frontier parity, not a claim of beating it (Sonnet 92.7, Opus 93.3 on our harness). We state no lift over the lightweight model on its own: our two measurements of that baseline disagree by three points, so the honest number is the one above and no delta.
config premise holds against live config (ledger L31)
Send your own tests →Oracle: NONE. There is nothing to run.
We do not pretend agreement is proof. Two different lightweight models must independently agree or the request escalates. Agreement is a confidence signal.
Agreement is not verification, and we label it that way
On HaluEval faithfulness judgements, measured cross-model on the pair now in production and recomputed from raw rows on 2026-07-28, the gate accepted a wrong answer in 11 of 40 agreements. That is 27.5%, Wilson 95% interval 16.1 to 42.8. The sub-splits are too small to quote as separate rates. We do not have comparably audited figures on other kinds of request, so do not export this number to them, and do not read it as a product-wide error rate: it is conditional on one benchmark's agreed slice.
Adversarially reviewed. Adversarially reviewed 2026-08-02, which rejected the previous draft; this is the wording it prescribed.
config premise holds against live config (ledger L79) · adversarially reviewed 2026-08-02, which rejected the previous draft; this is the wording it prescribed
Oracle: Real. Berkeley's own stateful checker decides whether the end state is correct.
Holds a multi-step task together across turns and leaves the world in the right state.
106 of 199 on BFCL multi-turn, scored by Berkeley's stateful checker, through live production
53.3%, Wilson 95% interval 46.3 to 60.1, with 0 errors. That is the same band as o1-2024-12-17 FC at 53.00% and Claude-3.7-Sonnet FC at 54.50% on Berkeley's June 2025 snapshot. It is not a leaderboard rank. Before any score was printed the instrument was validated in both directions: a positive control passed 199 of 199, and a negative control of silent and sabotaged transcripts was accepted 0 of 199. We run 199 of the 200 items, excluding one that crashes Berkeley's own checker.
Adversarially reviewed. Cleared 2026-08-02 with its comparability caveats attached.
ledger row L22 declares no PREMISE · cleared 2026-08-02 with its comparability caveats attached
Oracle: Availability, not identical safety.
Honestly: the failover model matches the primary on calling and fails it on abstaining.
Run exactly as production pins it, the failover model scores 30 of 40 on the traps
All 10 misses are the same verdict: it calls a tool where it should decline. It matches the primary on calling and fails it on abstaining, which is precisely what the trap set exists to measure. Wilson 95% interval 59.8 to 85.8; at n=40 that is wide, so do not read 30 of 40 as a precise rate. On standard tasks the same model scores 382 of 400. When the primary stalls we fail over for availability, and on that path the guarantee above does not hold. We would rather you knew.
Adversarially reviewed. Cleared 2026-08-02 to be published as the counter-example beside the primary-path number, never alone.
config premise holds against live config (ledger L20) · cleared 2026-08-02 to be published as the counter-example beside the primary-path number, never alone
Oracle: Reuse of previously checked answers only.
Nothing is reused that was not checked when it was first served.
Zero external customers have ever been served a reused answer
65 reuse serves in 93,242 requests, all internal. Cross-context reuse only fires when the caller sends tests, and no external key has. We do not sell a cost curve on this.
Adversarially reviewed. Re-derived from the production database 2026-08-02.
ledger row L23 declares no PREMISE · re-derived from the production database 2026-08-02
You call one model id. Behind it, the request is classified by shape and each shape has its own path. These values are read from the running gateway configuration when this page is generated, which is what lets the badges above mean anything.
We publish results, not the recipe, and we would rather say so than look coy about it. Which specific model serves which request shape is the output of our own evaluation work, and it is the part a competitor could copy in an afternoon. So we show that each value was read live and that the paths genuinely differ, and we hold the map itself. Everything the map produces, including where it performs worse, is on this page. If you need to reproduce a number, reproduce it against the gateway: that is the thing we are actually selling, and it is the thing our numbers describe.
| Tool-call path | configured, read live at generation time |
| Code and general path | configured, read live at generation time |
| Witness / failover | configured, read live at generation time |
| Escalation target | configured, read live at generation time |
| Escalation wall-clock budget | 45 |
Stated because its absence changes how you should read everything above.
We published these and they were wrong. They are listed here rather than deleted, because a correction you cannot find is not a correction.
Every number here carries its population, sample size and date, or it does not appear. This page is regenerated on deploy; if the configuration changes and the page is not regenerated, the build fails rather than publishing a stale claim.