Research

Which parts actually carry the result.

A system with three escalation layers, a memory and an elimination rule can improve a number for reasons its authors did not intend. Ablations are how we find out which parts are load-bearing — by removing one at a time and re-running the same campaign against the same split.

Nothing here is published yet. The arms below are fixed, and most of them are one config file away from running; the runs are queued behind the primary benchmark work. Results land here one arm at a time.

The arms

Config only — runnable as shipped, no patch. Needs code — the switch does not exist yet.

There is no per-flag switch on the command line. An arm is a campaign config, handed to new <dataset> --config <path>, and the keys above all sit under optimization: in that file — which also means every arm on this page is one file you can write yourself and run on your own key.

Every arm holds the model, the split and the scoring fixed, and reports the same four numbers as the benchmarks — accuracy alone would hide most of what an ablation is for. How we score