Same prompt. Multiple models. One verdict.
Paste the prompt you actually use — any length. Say what you want the writing to do. Pick up to 10 models. Every passage comes back side by side with its sentence-structure numbers, a blind assessment from Claude Opus, and what a whole novel would cost at that model's rate.
How it works
- You write the prompt. The real one, with your system instructions and constraints. Add a one-line goal so the judge knows what "worked" means to you.
- You get a quote. The price is what the run costs at the providers' list rates, plus our fee, broken down by model. Revise it or pay it. Nothing is charged until you do.
- We run it everywhere at once. Each model you picked gets the identical prompt, with the thinking level you chose, at the model's default temperature. Nothing is edited.
- We measure. Every passage is run through the same sentence-structure ruler: words per sentence, the p10/p90 spread, fragments, front-loaded clauses, dialogue share, paragraph density, and more. Real counts, no opinions.
- Claude Opus reads them blind. Model names hidden, order shuffled. It scores prompt compliance, prose quality, and your goal for each passage, ranks the field on all three, and names the one it would send you.
- You get the whole thing. Every passage in full, the numbers side by side, the verdict, and the cost of a 50,000-word novel at each model's rate. On the page, and as a PDF.
What it costs
Cost plus our fee, quoted before you pay. We price every token of your prompt into every model you picked, the full thinking ceiling for the level you chose, a 1,000-word passage from each, and the Opus assessment reading all of it — at the providers' list rates. Our fee is the same again, with a $1.00 minimum, and card processing is passed through at cost. A short prompt to five models is usually two or three dollars.
The quote is the price: if a model thinks less than its ceiling, that's fine, and if a model fails to deliver, its line is taken off before your card is charged.
What this is not
It isn't a leaderboard. Rankings from a thousand generic prompts say nothing about your prompt in your voice. This runs your prompt, and the verdict is about that and nothing else.
And the judge is a reader, not an oracle. Its scores are directional; the passages are the evidence. Read them.