Independent testing lab for AI tools

Which AI tool actually does the job. Tested, dated, not paid for.

Tool Radar gives competing AI tools the same real task and publishes what each one produced, what it cost, how long it took and how much fixing it needed. Every result is dated. No vendor can pay for a better ranking.

Rankings are free. No account needed to read them.

Published test 014

Analyse 14 months of sales and explain why Q3 dipped

Tested 12 days ago
Same brief, same file, 20 min limitCostTimeFixes
CellwiseCorrect cause. One chart mislabelled.
$00m 00s0
GridmindWrong region. Two formula errors.
$00m 00s0
TabuloCorrect cause. Follow-up prompt needed.
$00m 00s0
LedgerlyCorrect cause. Clean charts. Exact rows cited.
$00m 00s0
Initial order: fastest first

How a test works

Same task, same file, same clock.

01

One brief

Every tool gets the identical prompt and input file, written before any tool is opened.

02

One clock

A fixed time limit, the same for every tool. What exists at the end is what gets judged.

03

Fixes counted

A tester makes the output usable and logs every correction. The count is the headline number, because it is the one that costs you time.

04

Published, including failures

Every output is shown, not just the winner's.

Gridmind fix log4 corrections
Replaced the blamed region. The source rows show the dip came from North.
Extended the SUM range to include July and August.
Reset the chart axis from 82 to zero to remove the misleading scale.
Recalculated the percentage against total Q2 sales, not one month.

Speed is easy to measure and easy to fake. Fixes are neither.

Find a winner

The winner depends on your budget.

Showing results for: Analyse a spreadsheet

Budget
Experience
Sort
Ledgerly is first for the selected filters.
ResultCostTimeFixes
LedgerlyCorrect cause. Clean charts. Exact rows cited.
$45/mo3m 20s0Winner for these filters
CellwiseCorrect cause. One chart mislabelled.
$30/mo4m 10s1
TabuloCorrect cause. Follow-up prompt needed.
$06m 30s2
GridmindWrong region. Two formula errors.
$20/mo2m 45s4

There is no single best tool. There is a best tool for this task, at this budget, as of this date.

Every result has a date

Tested 12 days ago is a fact. Tested in March is a guess.

Current

Analyse a spreadsheet

Ledgerly
Current

Clean up a podcast recording

Soundloom
Retest due

Write a sales email sequence

Draftwell

Models change monthly. Old results are shown as old.

Who pays for this

Buyers pay. Vendors cannot.

Funded by

  • Readers on membership
  • Businesses commissioning evaluations for their own software decisions

Not funded by

  • Tool vendors or affiliate links
  • Sponsored placements or any payment tied to a ranking

A vendor can see their own results. They cannot change them, pay for a retest or buy a placement. If a vendor thinks a result is wrong, they can send us a reason and we will retest on the same brief, and publish both results.

Who pays for this

Buyers pay. Vendors cannot.

Funded by

  • Readers on membership
  • Businesses commissioning evaluations for their own software decisions

Not funded by

  • Tool vendors or affiliate links
  • Sponsored placements or any payment tied to a ranking

A vendor can see their own results. They cannot change them, pay for a retest or buy a placement. If a vendor thinks a result is wrong, they can send us a reason and we will retest on the same brief, and publish both results.

Pricing

Read the ranking. Pay for the detail.

Free

$0

Basic rankings for every task.

Testing dates on every result.

No account needed.

Member

$19 a month

Full comparisons with every output and fix log.

Recommended tool stacks by role.

Evaluation

By enquiry

A commissioned test on your team's actual tasks.

Your brief, your files, the same method.

Results are yours and are not published unless you choose.