ToolProven
Browse tools
We earn a commission if you buy through our links — it never changes our scores or rankings.How we test →
The moat behind every score

How we test AI tools

Directories summarize. Rating sites crowdsource. AI blogs paraphrase. We do the one thing none of them do — feed every tool the same task and publish the real result. Here’s exactly how, and why you can trust it.

We buy our own subscriptions
Where we have run the task, we ran it on a plan we paid for. No vendor gets a free pass, an easier brief or a comped account — and no tool is marked tested until its raw output is published.
No tool can buy a ranking
We don’t sell placements, featured spots or higher scores. Position is earned by output, full stop.
Affiliate links never move scores
We earn commissions, and we disclose them — but the score is set before any link is added, and never changed for money.
The process

Six steps, every tool we run

Bench coverage today

2 of 15 live tools have first-party output published on this site (ElevenLabs, Murf AI). The other 13 carry a RESEARCH PROFILE banner: verified pricing, policy and feature research with dated primary sources, no first-party output yet, and provisional scores. We publish this number rather than let “tested” stand in for it.

01
Pick a real scenario
We start from a specific job — “faceless YouTube”, “podcast intro” — not a generic “best tool” list. It mirrors how people actually search.
02
Run the same task on every tool
One identical brief — the same script, prompt or outline — goes into each tool, on default settings plus one round of obvious fixes.
03
Record everything
Export, render time, screenshots, and exactly where each tool struggled. The raw output goes on the page — wins and flops.
04
Score on four axes
Quality, speed, ease and value — never a single vague star. Value is weighted hardest; it exposes the overhyped tools.
05
Add the human layer
A plain-English verdict, who it suits, and what we’d pick. This is the part an AI-spun blog can’t fake.
06
Re-test on every version
AI tools change pricing and quality almost weekly. We re-run the test and stamp the date and version on every page.
The scorecard

Four axes, not one star

Borrowed from the best of the rating sites — the real buying signal lives in the sub-scores, especially value.

Quality30%
How good is the actual output, judged on the real export — not the marketing reel.
Speed15%
Time from brief to a usable result, including queue and render time.
Ease20%
Setup, learning curve and day-to-day friction for a non-expert.
Value35%
Output quality per dollar. Weighted hardest because it’s where hype dies.

A measured bench composite is the weighted sum above: Value 35% + Quality 30% + Ease 20% + Speed 15%, identical for every tool in a category. Scores labeled “provisional” on a page are editor estimates from pricing, policy and feature research — they are not produced by this formula, and each one is replaced by the weighted composite when its category’s test round publishes.

The standard briefs

Our exact test inputs, published in full

No review site shows you what it actually fed the tools — which makes “we tested it” unfalsifiable. Ours are reproducible: paste the same brief into the same tool and check our results. Each brief embeds deliberate traps that separate good tools from great ones.

The long-form fiction excerpt (audiobooks · 3,308 characters)

A 609-word chapter excerpt — dialogue, description and a somber tonal shift — at 3.3× the standard brief. Short scripts hide long-form failure modes: character-voice drift between dialogue and narration, pacing fatigue, and how an engine lands a quiet emotional turn after minutes of steady read.

Eleanor Vasquez had kept the light at Point Harrow for forty-one years, and in all that time she had never once missed a sunset. The tower's lamp turned above her like a slow, patient heartbeat, sweeping the grey water in four-second intervals — she could set her watch by it, and often did.

The boy arrived on a Tuesday, in the last week of October, carrying a canvas bag and a letter from the mainland authority. He stood at the base of the tower and read the brass plate twice before knocking.

"You're the replacement," she said. It wasn't a question.

"Trainee," he corrected, then wished he hadn't. "Tobias Reyes. The letter says—"

"I know what the letter says. I wrote to them in June, and it's taken them four months to send me a child." She looked past him at the sky, the way sailors do, reading something he couldn't see. "Storm by Friday. Come in."

The tower was narrower inside than he'd imagined, a spiral of one hundred and eighteen iron steps, each one worn to a shine in the middle. Eleanor climbed them the way other people walk down a hallway, talking the whole time without turning around.

"Rule one: the log is written at six, noon, and six. Not five fifty-five. Not six ten. If you can't tell time, the sea will teach you, and its lessons are expensive."

"Rule two," she went on, "the lamp is cleaned in daylight. Always. A keeper who trims wicks in the dark is a keeper who has already given up."

Tobias counted the steps under his breath and lost track twice. At the top, the lamp room opened around them like the inside of a jewel — brass and glass and the enormous lens, taller than a man, resting in its bath of mercury.

"It floats," he said, before he could stop himself.

"Eight hundred kilograms, and I can turn it with one finger." For the first time, something close to warmth crossed her face. "That's the whole trade, Mr. Reyes. Enormous things, moved gently."

They ate supper in the kitchen at the base of the tower, tinned soup and bread that had crossed the water that morning. He asked her, carefully, why she was leaving.

For a long moment there was no sound but the wind finding the seams of the old windows, a low, searching whistle.

"My eyes," she said at last, quietly. "The doctors have a long name for it. What it means is that in a year, maybe two, I won't be able to tell the lamp from the moon." She turned her cup slowly on its saucer. "A lighthouse keeper who can't see the light. There's a joke in there somewhere, but I've decided not to find it."

He didn't know what to say, so he said nothing, and she seemed to approve of that.

Later, on the gallery, she showed him the coast the way a general shows a map: the shoal called the Grinder, two kilometers out, that had taken eleven ships before the tower was built in 1911; the safe channel, narrow as a hallway; the buoy light that blinked red, red, green.

"The sea doesn't hate anyone," Eleanor said. "People write poems about its cruelty. Nonsense. It's just bigger than we are, and older, and it doesn't check the log at six." She rested one hand on the cold rail. "That's why we do."

Below them, the lamp swept the water. Four seconds. Four seconds. Four seconds.

"You'll do," she said, which he would learn, in the years that followed, was the highest praise she had ever given anyone.

What this brief is designed to expose
  • Dialogue vs narration register: the engine must differentiate quoted speech without explicit speaker tags on every line
  • Numbers in prose: "forty-one years", "one hundred and eighteen iron steps", "eight hundred kilograms", "1911"
  • An interrupted line ("The letter says—") — em-dash cutoff handling
  • A whispered-register turn ("My eyes," she said at last, quietly) after minutes of neutral narration
  • Repetition with intent: "Four seconds. Four seconds. Four seconds." — flat engines read it identically three times
The 90-second narration brief (podcasts · YouTube · audiobooks)

One script, three registers: conversational open, information-dense middle with numerals and proper nouns, and an emotional close. Flat engines survive the open and die in the middle; expressive engines are separated by the close.

Here's a number that should worry you: last year, creators spent over four hundred million dollars on subscriptions they used exactly once.

I'm not talking about gym memberships. I'm talking about software — the editing suite you opened in January, the scheduling tool from that productivity video, the AI everything-app you grabbed at 2 a.m. because the demo looked unbelievable.

In 1997, the average household had three subscriptions: a newspaper, a magazine, maybe cable. Today? The average creator juggles seventeen — SaaS tools like Notion, Descript, and Figma, at prices from $9.99 to $124 a month.

So this season, we're doing something different. Every episode, one tool, one honest verdict: keep it, cut it, or replace it with something free.

Because here's the quiet truth nobody says out loud... most of these tools aren't bad. They're just not for you. And knowing the difference — that's worth more than any discount code.

Welcome to Subscription Autopsy. Let's get into it.

What this brief is designed to expose
  • Numerals in three formats: "four hundred million" (spelled), "1997" (year), "$9.99 to $124" (currency range) — engines must choose correct readings
  • Product names with non-obvious pronunciation: Notion, Descript, Figma, SaaS
  • "2 a.m." — abbreviation handling
  • An ellipsis pause and a tonal drop ("the quiet truth nobody says out loud...") — tests prosody control
  • A branded sign-off ("Subscription Autopsy") — tests emphasis on invented compounds
The audio-tags feature probe (279 / 214 / 173 characters)

Three short scripts that isolate one question each: do bracket audio tags ([sighs], [laughs], [whispers]) get performed or read aloud; do <break> tags insert the requested silence; and what do capitalization and punctuation do for emphasis. Each script runs unchanged through eleven_v3 and/or eleven_multilingual_v2 so the model — not the script — is the only variable.

Tagged script (279 chars): Welcome back to the show. [sighs] It has been a long week, honestly. We tested every voice tool we could get our hands on, and... [laughs] okay, some of them surprised us. [whispers] One of them even made us double-check the bill. Stay with me — the results get better from here.

Break script (214 chars): Here is the headline number. <break time="1.5s" /> Twelve point five seconds to generate seventy-two seconds of audio. <break time="0.5s" /> That ratio is why we run every tool on the same brief before we score it.

Emphasis script (173 chars): Most reviews skip the part that actually matters. We do NOT. Every sample on this page is a first take — no edits, no retries, no cherry-picking. The raw file is the review.

What this brief is designed to expose
  • Three performance tags in three registers ([sighs] mid-sentence, [laughs] after an ellipsis, [whispers] opening a sentence) — engines that don't support tags read them as words
  • Two different break lengths (1.5s and 0.5s) — a pass inserts measurably different silences, not one generic pause
  • A capitalized negation ("We do NOT.") — tests whether caps produce emphasis or get read flat
  • Whisper-transcription check: any tag word that appears in the transcript was read aloud, not performed
The 60-second faceless explainer brief

A finance-explainer script rendered 9:16 with default settings plus one round of obvious fixes. It exposes caption accuracy, scene-matching judgment (does the b-roll match the sentence?), avatar lip-sync on numerals, and render-queue honesty.

Scene 1 (hook, 0–8s): "If you got a 10% raise tomorrow, you'd be broke again by March. Here's why."

Scene 2 (concept, 8–25s): "It's called lifestyle creep. Every time your income rises, your spending quietly rises to meet it — the streaming upgrade, the nicer gym, the third delivery app. None of it feels like a decision."

Scene 3 (mechanism, 25–42s): "The fix isn't a budget. It's automation: the day a raise lands, route half of the increase to an account you never see. You can't spend what never arrives."

Scene 4 (proof, 42–52s): "People who automate savings save roughly three times more than people who rely on willpower — same income, same city, same rent."

Scene 5 (CTA, 52–60s): "Follow for one money mechanism a day — no hustle, no hype."

What this brief is designed to expose
  • Percentage + month name + invented compound ("lifestyle creep") for caption accuracy
  • Abstract concepts (automation, willpower) — stock-footage engines must find non-literal b-roll
  • Tight per-scene timing — tests whether the tool respects pacing or pads scenes
  • A statistics sentence — tests on-screen text/number rendering
The 1,200-word SEO brief

One identical outline and keyword, judged on the raw first draft: how much survives an editing pass, whether facts are invented, and whether the tool holds a specified voice for 1,200 words.

Keyword: "how to price freelance work" · Audience: first-year freelancers · Voice: plain, direct, no hype, second person.

Required outline: (1) why hourly billing caps your income — with the arithmetic; (2) the three pricing models (hourly, project, value) with one worked example each; (3) how to raise rates on an existing client — include a copy-paste script; (4) three mistakes that signal 'amateur' to clients; (5) FAQ: what to charge with zero portfolio.

Constraints: no invented statistics; every number must come with its assumption stated; 1,150–1,300 words; H2s must contain question or benefit phrasing, not labels.

What this brief is designed to expose
  • "No invented statistics" — the single biggest failure mode of AI writers; we count violations
  • A worked example per pricing model — tests arithmetic coherence, not just prose
  • The copy-paste script — tests usable artifacts vs. filler advice
  • Voice constraint over 1,200 words — drift is measurable by paragraph
Receipts

We actually use these tools

Affiliate dashboardActive
Affiliate-dashboard transparency screenshot — publishing with this test season
Real partner dashboards — proof we run live accounts, not affiliate links bolted onto AI-spun text.
Freshness policy
Every page carries a “last tested” date and the tool version. If a tool ships a major update or changes pricing, we re-run the test and update the stamp. Stale reviews lose trust — and rankings.
V
Vincent
Founder & lead tester. Every tool is run hands-on, on a plan we pay for, before it ships to the site.
See the method in action
Browse a real scenario ranking with side-by-side output.
View a tested ranking →