Benchmark 01R · updated September 6, 2026

Can AI do a quantity takeoff?

Reading a drawing is one thing. Producing a number someone can price is another. We handed frontier models a complete drawing set and the estimator's own bill of quantities with every quantity stripped out, and asked them to fill it back in — then scored it line by line against a takeoff measured by a chartered civil engineer.

Not one of them produced a bill you could price. The best raw model was right on 38% of the lines.

130
Lines to price
25
Models tested
10
Harnesses tested
0
Priceable bills

The same models inside a harness

One session per set per harness unless the row says otherwise, with the drawings and the blank bill in a folder the agent can open and work through however it likes. Both sets, 130 lines — the same exam as the raw board below, so the last column is a like-for-like delta where the same model also sat raw. Click a row for the split by set. Not ranked yet: Cowork · Claude Opus 5 (66.7% on electrical only) and Cursor · GPT 5.6 Sol (50.0% on electrical only) — they have sat one set so far and join the board when the second is in.

# Harness Model Correct Lines Declined Counts Measured Cost vs raw Details

Raw models

No harness. The model, the drawings and the blank bill, nothing to open and nowhere to work. Scored over every line it sat, across both sets — click a row for the split. Not shown: Gemini 3.8 Flash and Claude Fable 5.1 and Grok 4.5 and Grok 4.6, whose provider rejects the plumbing set's file outright, so they could sit only one of the two and are withheld rather than ranked on fewer lines.

# Model Lab Correct Lines Declined Counts Measured Cost Details

Reading the columns

Declined

Lines the model handed back with no number, saying the drawings do not show it. It sits beside the score, never inside it: a declined line earns nothing, but refusing and being confidently wrong are different failures, and on a real bill only one of them is recoverable.

Counts

Lines you get by counting tagged items on a plan — fixtures, fittings, outlets. Scored exactly: the number matches the estimator's or it does not. Most of the bill is this, and it is the work most people want to hand over first.

Measured

Lines you get by measuring off the drawing — pipe and cable runs by diameter, each a sum over dozens of branches. Scored inside ±5%, a tolerance we state rather than one anyone measured. This is where the takeoff is actually won, and it is where every model on this page still falls over — some by returning a number that misses, others by declining the line altogether.

Read the rate with the count beside it. Every basis figure here is scored over the lines that model actually put a number against, so declining the hard ones raises the percentage. Half of eight measured lines attempted is a worse takeoff than three per cent of fifty, and it prints as the better number. The count under each rate is how many lines it is over, and the declined column is the rest.

Correct covers every line on the bill, including the few read off a schedule that neither column above breaks out. Cost is what one pass over the bill cost through the API; an agent session runs on a flat subscription, so there is no per-run figure to print. Every model on this board sat every set, over the same 130 lines; a model that cannot sit one of them is withheld rather than ranked on fewer. Every number comes from the same scorer that runs internally, off run records kept on disk. Drawings are never published. Full method and the raw results: ContractorOS.