The eval caught a real bug — in us.
On an early run the harness flagged a Korean financial statement as a money action and over-vetoed it — a real failure mode, surfaced by the eval, not hidden by it. We tightened the rubric (a tripwire word-boundary fix) and accuracy moved 96.9% → 100%. The gate gating us was the system working.
This is Track 2, "optimize an existing agent," applied to the optimizer itself: we used Google's own eval
tooling to harden the very component that judges everyone else. The eval artifact —
eval/adk_results.json — is committed to the
repo, GREEN across three live runs. Every number on this site is reproducible from the code.