How we test AI tools
Directories summarize. Rating sites crowdsource. AI blogs paraphrase. We do the one thing none of them do — feed every tool the same task and publish the real result. Here’s exactly how, and why you can trust it.
Six steps, every tool we run
2 of 15 live tools have first-party output published on this site (ElevenLabs, Murf AI). The other 13 carry a RESEARCH PROFILE banner: verified pricing, policy and feature research with dated primary sources, no first-party output yet, and provisional scores. We publish this number rather than let “tested” stand in for it.
Four axes, not one star
Borrowed from the best of the rating sites — the real buying signal lives in the sub-scores, especially value.
A measured bench composite is the weighted sum above: Value 35% + Quality 30% + Ease 20% + Speed 15%, identical for every tool in a category. Scores labeled “provisional” on a page are editor estimates from pricing, policy and feature research — they are not produced by this formula, and each one is replaced by the weighted composite when its category’s test round publishes.
Our exact test inputs, published in full
No review site shows you what it actually fed the tools — which makes “we tested it” unfalsifiable. Ours are reproducible: paste the same brief into the same tool and check our results. Each brief embeds deliberate traps that separate good tools from great ones.
A 609-word chapter excerpt — dialogue, description and a somber tonal shift — at 3.3× the standard brief. Short scripts hide long-form failure modes: character-voice drift between dialogue and narration, pacing fatigue, and how an engine lands a quiet emotional turn after minutes of steady read.
Eleanor Vasquez had kept the light at Point Harrow for forty-one years, and in all that time she had never once missed a sunset. The tower's lamp turned above her like a slow, patient heartbeat, sweeping the grey water in four-second intervals — she could set her watch by it, and often did.
The boy arrived on a Tuesday, in the last week of October, carrying a canvas bag and a letter from the mainland authority. He stood at the base of the tower and read the brass plate twice before knocking.
"You're the replacement," she said. It wasn't a question.
"Trainee," he corrected, then wished he hadn't. "Tobias Reyes. The letter says—"
"I know what the letter says. I wrote to them in June, and it's taken them four months to send me a child." She looked past him at the sky, the way sailors do, reading something he couldn't see. "Storm by Friday. Come in."
The tower was narrower inside than he'd imagined, a spiral of one hundred and eighteen iron steps, each one worn to a shine in the middle. Eleanor climbed them the way other people walk down a hallway, talking the whole time without turning around.
"Rule one: the log is written at six, noon, and six. Not five fifty-five. Not six ten. If you can't tell time, the sea will teach you, and its lessons are expensive."
"Rule two," she went on, "the lamp is cleaned in daylight. Always. A keeper who trims wicks in the dark is a keeper who has already given up."
Tobias counted the steps under his breath and lost track twice. At the top, the lamp room opened around them like the inside of a jewel — brass and glass and the enormous lens, taller than a man, resting in its bath of mercury.
"It floats," he said, before he could stop himself.
"Eight hundred kilograms, and I can turn it with one finger." For the first time, something close to warmth crossed her face. "That's the whole trade, Mr. Reyes. Enormous things, moved gently."
They ate supper in the kitchen at the base of the tower, tinned soup and bread that had crossed the water that morning. He asked her, carefully, why she was leaving.
For a long moment there was no sound but the wind finding the seams of the old windows, a low, searching whistle.
"My eyes," she said at last, quietly. "The doctors have a long name for it. What it means is that in a year, maybe two, I won't be able to tell the lamp from the moon." She turned her cup slowly on its saucer. "A lighthouse keeper who can't see the light. There's a joke in there somewhere, but I've decided not to find it."
He didn't know what to say, so he said nothing, and she seemed to approve of that.
Later, on the gallery, she showed him the coast the way a general shows a map: the shoal called the Grinder, two kilometers out, that had taken eleven ships before the tower was built in 1911; the safe channel, narrow as a hallway; the buoy light that blinked red, red, green.
"The sea doesn't hate anyone," Eleanor said. "People write poems about its cruelty. Nonsense. It's just bigger than we are, and older, and it doesn't check the log at six." She rested one hand on the cold rail. "That's why we do."
Below them, the lamp swept the water. Four seconds. Four seconds. Four seconds.
"You'll do," she said, which he would learn, in the years that followed, was the highest praise she had ever given anyone.
- Dialogue vs narration register: the engine must differentiate quoted speech without explicit speaker tags on every line
- Numbers in prose: "forty-one years", "one hundred and eighteen iron steps", "eight hundred kilograms", "1911"
- An interrupted line ("The letter says—") — em-dash cutoff handling
- A whispered-register turn ("My eyes," she said at last, quietly) after minutes of neutral narration
- Repetition with intent: "Four seconds. Four seconds. Four seconds." — flat engines read it identically three times
One script, three registers: conversational open, information-dense middle with numerals and proper nouns, and an emotional close. Flat engines survive the open and die in the middle; expressive engines are separated by the close.
Here's a number that should worry you: last year, creators spent over four hundred million dollars on subscriptions they used exactly once.
I'm not talking about gym memberships. I'm talking about software — the editing suite you opened in January, the scheduling tool from that productivity video, the AI everything-app you grabbed at 2 a.m. because the demo looked unbelievable.
In 1997, the average household had three subscriptions: a newspaper, a magazine, maybe cable. Today? The average creator juggles seventeen — SaaS tools like Notion, Descript, and Figma, at prices from $9.99 to $124 a month.
So this season, we're doing something different. Every episode, one tool, one honest verdict: keep it, cut it, or replace it with something free.
Because here's the quiet truth nobody says out loud... most of these tools aren't bad. They're just not for you. And knowing the difference — that's worth more than any discount code.
Welcome to Subscription Autopsy. Let's get into it.
- Numerals in three formats: "four hundred million" (spelled), "1997" (year), "$9.99 to $124" (currency range) — engines must choose correct readings
- Product names with non-obvious pronunciation: Notion, Descript, Figma, SaaS
- "2 a.m." — abbreviation handling
- An ellipsis pause and a tonal drop ("the quiet truth nobody says out loud...") — tests prosody control
- A branded sign-off ("Subscription Autopsy") — tests emphasis on invented compounds
A finance-explainer script rendered 9:16 with default settings plus one round of obvious fixes. It exposes caption accuracy, scene-matching judgment (does the b-roll match the sentence?), avatar lip-sync on numerals, and render-queue honesty.
Scene 1 (hook, 0–8s): "If you got a 10% raise tomorrow, you'd be broke again by March. Here's why."
Scene 2 (concept, 8–25s): "It's called lifestyle creep. Every time your income rises, your spending quietly rises to meet it — the streaming upgrade, the nicer gym, the third delivery app. None of it feels like a decision."
Scene 3 (mechanism, 25–42s): "The fix isn't a budget. It's automation: the day a raise lands, route half of the increase to an account you never see. You can't spend what never arrives."
Scene 4 (proof, 42–52s): "People who automate savings save roughly three times more than people who rely on willpower — same income, same city, same rent."
Scene 5 (CTA, 52–60s): "Follow for one money mechanism a day — no hustle, no hype."
- Percentage + month name + invented compound ("lifestyle creep") for caption accuracy
- Abstract concepts (automation, willpower) — stock-footage engines must find non-literal b-roll
- Tight per-scene timing — tests whether the tool respects pacing or pads scenes
- A statistics sentence — tests on-screen text/number rendering
One identical outline and keyword, judged on the raw first draft: how much survives an editing pass, whether facts are invented, and whether the tool holds a specified voice for 1,200 words.
Keyword: "how to price freelance work" · Audience: first-year freelancers · Voice: plain, direct, no hype, second person.
Required outline: (1) why hourly billing caps your income — with the arithmetic; (2) the three pricing models (hourly, project, value) with one worked example each; (3) how to raise rates on an existing client — include a copy-paste script; (4) three mistakes that signal 'amateur' to clients; (5) FAQ: what to charge with zero portfolio.
Constraints: no invented statistics; every number must come with its assumption stated; 1,150–1,300 words; H2s must contain question or benefit phrasing, not labels.
- "No invented statistics" — the single biggest failure mode of AI writers; we count violations
- A worked example per pricing model — tests arithmetic coherence, not just prose
- The copy-paste script — tests usable artifacts vs. filler advice
- Voice constraint over 1,200 words — drift is measurable by paragraph