About

Watch the benchmarks, not the headlines.

That is the one lesson I kept from the last ten years, and it is the reason PromptPotter measures every prompt it writes instead of asking you to trust it.

Founder & maintainer

David Streuli

Switzerland

PromptPotter is built and maintained by one person. I write the optimizer, run the benchmarks, answer the mail, and host the public instance you sign in to. When something here is wrong, it is mine to fix, and you can reach me directly.

I came to language models sideways. My training is in structural biology and biochemistry, and I moved through bioinformatics into NLP because that is where the methods I needed actually lived. The short version:

How I got here

  1. 2020, Taiwan

    A professor assigns AlphaFold as homework. I was a structural biologist at the time. I never really went back.

  2. The turn

    Biology's databases are fragmented, and the worse problem is the knowledge that is in no database at all. It sits in papers, in prose, behind a human who has to read it first.

  3. Into NLP

    Entity recognition, normalization, linking, relation extraction. Every method I needed turned out to be natural language processing, so I left biology for it. This was before ChatGPT existed.

  4. 2018 to 2021

    I watched deep learning beat every classical method on every NLP benchmark, inside about three years. The revolution was legible in the leaderboards long before the public heard about it.

Why this tool, and not another wrapper

If you spend three years reading leaderboards, you stop trusting any improvement that arrives without a number attached. Prompt engineering is still mostly that: someone rewrites a prompt, it feels better, and nobody checks whether it is better.

So PromptPotter does the boring half. It writes candidate prompts, scores every one against examples where you already know the right answer, keeps what actually wins, and hands you the number. You do not have to believe the improvement. You can read it.

The optimizer is open source, and the benchmark runs behind the claims are published with their cuts and their protocol, so you can disagree with me using my own data.

Start a run See the benchmark results