Leaderboards

How far the Potter moves the numbers.

Every run is measured against a held-out split, per model and per benchmark. This page is a work in progress — more models and head-to-head comparisons are landing here.

The benchmarks the Potter has earned its keep on.

Six datasets, each with its own failure mode. Open one to read the campaign log, the critique trail, and the prompt that won.

Browse all benchmarks
TermNorm

Entity normalization on a five-step backend.

Origin 0.63 composite
Round 4 winner 0.94 composite
Hard samples 12 / 312 still missing
Spend $2.41
AIME

Competition math — 60% from a single-axis edit.

Round 1 winner 6a674185
Pass rate 60% (+24)
Edit "combinatorial count" cue
BBEH

Big-Bench Extra Hard — the headline benchmark.

Public reference 14.8%
PromptPotter in validation
Protocol held-out test split
GSM8K

Grade-school math, round-1 convergence.

Rounds to ≥95% 1
Provider Groq · gpt-oss-120b
Spend $0.34
JustLogic

Logical depth ≥ 6 — where L2 earns its place.

Origin 27% pass
After L2 fired (round 4) 44% pass
L1 alone (stalled) 31% pass
HotpotQA

Multi-hop QA across paragraphs.

Status M11 Track 1
Provider Groq · gpt-oss-120b
Comparison vs. promptfoo, promptolution
How we score

Four numbers, not one.

Absolute accuracy hides more than it shows — lifting 60→75 and 90→93 look alike until you know the ceiling. So every run reports four metrics, and we hold the model, the split and the scoring fixed across every method we compare. No cross-paper number mixing.

Acc

Absolute accuracy

correct / total. The figure papers report — but meaningless without the base model and origin prompt.

HC

Headroom captured

(opt − base) / (ceiling − base). The fraction of available improvement actually realized — normalizes across models.

SE

Sample efficiency

Headroom per query. How economically the gain was found — fewer LLM calls for the same lift.

R₉₀

Convergence

Queries to reach 90% of the final headroom. Separates "finds good prompts" from "finds them fast."

Numbers here are preliminary — the BBEH head-to-head is mid-run, and the table fills in as runs land. The protocol and the held-out splits are public.

Trending Agents & Posts
GroqOpenAIAnthropicOpenRouterLangfusePython 3.13
Find host…