Tools · AI Estimates

Price the build before you start it.

Time, tokens, dollars, and how many people you still need. AIdaScore prices the work you did — this prices the work you haven’t.

Free. Runs in this tab — nothing about your project leaves the browser.

Enterprise CRM rollout · 3,200 person-daysReal run
25.6days

P80 39.2 days · bound by your review — 93% of the work is you, not the agents

Build cost
$12.7k
$955 inference
Tokens
432.7M
218 tasks
People
12
was 13
Programme
$1.35M
was $1.44M

Confidence: priors · band ±53%. 19 coefficients are assumed, not measured.

01

The gain is mostly time · so time is the headline

Cost and headcount move by different amounts. Averaging them into one flattering multiple is how estimates start lying, so they are never averaged.

Elapsed is the larger of agent time and review time, never the sum — which is why past a point a faster model changes nothing, and eight agents are not eight times faster.

02

The knee is emergent · nothing asserts it

Requirements are classified and consumed commodity-first. The expensive tail is exactly what you buy chasing the last slice of parity.

  1. Shelfadopt it3%
  2. Troddenadapt it5%
  3. Bespokebuild it22%
  4. Noveldiscover it65%

Share of a from-scratch build each class costs. A novel requirement is roughly twenty times a shelf one — so a CRM that is 40% off-the-shelf and a core banking platform that is 5% produce genuinely different answers, without anyone drawing a curve by hand.

03

Leverage falls as the programme grows

A 220-person-day product gets a transformation. A 120,000-person-day programme gets a rounding error — and, at these settings, a penalty.

ShapeSizeOff the shelfPeopleBuild leverage
Simple SaaS product220 PD44%13.0×
Enterprise CRM rollout3,200 PD40%127.7×
ERP finance + supply chain30,000 PD18%1152.1×
Core banking migration120,000 PD5%5250.8×

Core banking comes back at 0.8× — below one. With 5% of the surface available to adopt and one reviewer absorbing the output, the model says agents cost you more than they save. A tool that could not return that answer would not be worth running on the other three.

04

Review sets the date · not model speed

The same brief, every model. Enterprise CRM, at defaults.

ModelElapsedAll-inTokens
Claude Opus 525.6 days$12.7k432.7M
Codex GPT-5.x-codex~27.1 days$12.7k458.2M
Claude Sonnet 527.9 days$13.4k472.1M
DeepSeek V3.x~33.6 days$15.5k577.0M
Qwen3-Coder~35.3 days$16.3k610.9M
Claude Haiku 4.537.9 days$17.7k662.9M

Fastest to slowest is 25.6 to 37.9 days — a 1.5× spread on a decision people argue about as though it were 10×. ~ marks a row whose pricing or first-try rate is assumed rather than published.

05

What it doesn’t know · and says so

A confident wrong number is the failure mode. So every coefficient carries its provenance, and the tool prints the table on demand.

19coefficients assumed

Only the Anthropic pricing is published. Every load-bearing number — first-try rate above all — is a prior, not a measurement.

±53%band, and it never closes

More evidence narrows P50 to P80. Nothing collapses it to a point, because nothing honestly can.

0bare leverage multiples

“9.5×” without its scope is the most misleading thing this could print. Build and programme are always returned together.

Treat every number as a shape, not a quote.

Estimate your own build.

Pick the nearest shape, correct the parameters that make yours hard, and read the band. Sign in and it opens in this tab.