Every run is measured against a held-out split, per model and per benchmark. This page is a work in progress — more models and head-to-head comparisons are landing here.
Six datasets, each with its own failure mode. Open one to read the campaign log, the critique trail, and the prompt that won.
Absolute accuracy hides more than it shows — lifting 60→75 and 90→93 look alike until you know the ceiling. So every run reports four metrics, and we hold the model, the split and the scoring fixed across every method we compare. No cross-paper number mixing.
correct / total. The figure papers report — but meaningless without the base model and origin prompt.
(opt − base) / (ceiling − base). The fraction of available improvement actually realized — normalizes across models.
Headroom per query. How economically the gain was found — fewer LLM calls for the same lift.
Queries to reach 90% of the final headroom. Separates "finds good prompts" from "finds them fast."
Numbers here are preliminary — the BBEH head-to-head is mid-run, and the table fills in as runs land. The protocol and the held-out splits are public.