Benchmark 01R · updated September 6, 2026
Can AI do a quantity takeoff?
Reading a drawing is one thing. Producing a number someone can price is another. We handed frontier models a complete drawing set and the estimator's own bill of quantities with every quantity stripped out, and asked them to fill it back in — then scored it line by line against a takeoff measured by a chartered civil engineer.
Not one of them produced a bill you could price. The best raw model was right on 38% of the lines.
The same models inside a harness
One session per set per harness unless the row says otherwise, with the drawings and the blank bill in a folder the agent can open and work through however it likes. Both sets, 130 lines — the same exam as the raw board below, so the last column is a like-for-like delta where the same model also sat raw. Click a row for the split by set. Not ranked yet: Cowork · Claude Opus 5 (66.7% on electrical only) and Cursor · GPT 5.6 Sol (50.0% on electrical only) — they have sat one set so far and join the board when the second is in.
| # | Harness | Model | Correct | Lines | Declined | Counts | Measured | Cost | vs raw | Details |
|---|---|---|---|---|---|---|---|---|---|---|
| 1 | ChatGPT Work electrical + plumbing · both sets in one session · self-built tooling | GPT 6 Astra | 47.7% | 130 | 8 | 61%of 74 | 28%of 47 | subscription | +9.2 pts raw, both sets: 38.5% | |
ChatGPT Work · GPT 6 Astra by set62 of 130 lines inside tolerance, 8 declined. One session per set, the folder and the blank bill and nothing else.
| ||||||||||
| 2 | Cowork electrical + plumbing · both sets in one session | Claude Fable 5.1 | 44.6% | 130 | 5 | 65%of 74 | 8%of 48 | subscription | no raw run on both sets | |
Cowork · Claude Fable 5.1 by set58 of 130 lines inside tolerance, 5 declined. One session per set, the folder and the blank bill and nothing else.
| ||||||||||
| 3 | Claude Code electrical + plumbing · 1 session each | Claude Opus 5 | 43.9% | 130 | 2 | 65%of 74 | 6%of 51 | subscription | +12.7 pts raw, both sets: 31.1% | |
Claude Code · Claude Opus 5 by set57 of 130 lines inside tolerance, 2 declined. One session per set, the folder and the blank bill and nothing else.
| ||||||||||
| 4 | Cursor electrical + plumbing · 1 session each | Claude Opus 5 | 41.5% | 130 | 24 | 64%of 72 | 6%of 31 | subscription | +10.4 pts raw, both sets: 31.1% | |
Cursor · Claude Opus 5 by set54 of 130 lines inside tolerance, 24 declined. One session per set, the folder and the blank bill and nothing else.
| ||||||||||
| 5 | ChatGPT Work electrical + plumbing · 1 session each | GPT 5.6 Sol | 36.9% | 130 | — | 55%of 75 | 4%of 52 | subscription | +1.1 pts raw, both sets: 35.8% | |
ChatGPT Work · GPT 5.6 Sol by set48 of 130 lines inside tolerance. One session per set, the folder and the blank bill and nothing else.
| ||||||||||
| 6 | Codex electrical + plumbing · 1 session each | GPT 5.6 Sol | 33.1% | 130 | 2 | 49%of 75 | 2%of 50 | subscription | −2.7 pts raw, both sets: 35.8% | |
Codex · GPT 5.6 Sol by set43 of 130 lines inside tolerance, 2 declined. One session per set, the folder and the blank bill and nothing else.
| ||||||||||
| 7 | Cursor electrical + plumbing · 1 session each | Grok 4.6 | 30.8% | 130 | 14 | 48%of 73 | 0%of 41 | subscription | no raw run on both sets | |
Cursor · Grok 4.6 by set40 of 130 lines inside tolerance, 14 declined. One session per set, the folder and the blank bill and nothing else.
| ||||||||||
| 8 | ChatGPT Work electrical + plumbing · 1 session each | GPT 5.6 Terra | 13.1% | 130 | 96 | 45%of 38 | n/a | subscription | −8.5 pts raw, both sets: 21.5% | |
ChatGPT Work · GPT 5.6 Terra by set17 of 130 lines inside tolerance, 96 declined. One session per set, the folder and the blank bill and nothing else.
| ||||||||||
Raw models
No harness. The model, the drawings and the blank bill, nothing to open and nowhere to work. Scored over every line it sat, across both sets — click a row for the split. Not shown: Gemini 3.8 Flash and Claude Fable 5.1 and Grok 4.5 and Grok 4.6, whose provider rejects the plumbing set's file outright, so they could sit only one of the two and are withheld rather than ranked on fewer lines.
| # | Model | Lab | Correct | Lines | Declined | Counts | Measured | Cost | Details |
|---|---|---|---|---|---|---|---|---|---|
| 1 | GPT 6 Astra openai/gpt-6-astra | OpenAI | 38.5% | 130 | 55 | 64%of 64 | 50%of 8 | $1.82 | |
GPT 6 Astra by set50 of 130 lines inside tolerance, 55 declined. $1.82 to fill the bill once.
| |||||||||
| 2 | GPT 5.6 Sol openai/gpt-5.6-sol | OpenAI | 35.8% | 130 | 3 | 53%of 74 | 3%of 51 | $0.54 | |
GPT 5.6 Sol by set47 of 130 lines inside tolerance, 3 declined. $0.54 to fill the bill once.
| |||||||||
| 3 | Claude Opus 5 anthropic/claude-opus-5 | Anthropic | 31.1% | 130 | 8 | 46%of 74 | 3%of 45 | $1.00 | |
Claude Opus 5 by set41 of 130 lines inside tolerance, 8 declined. $1.00 to fill the bill once.
| |||||||||
| 4 | Gemini 2.5 Flash Lite google/gemini-2.5-flash-lite | 30.0% | 130 | 11 | 44%of 70 | 4%of 46 | $0.01 | ||
Gemini 2.5 Flash Lite by set39 of 130 lines inside tolerance, 11 declined. $0.01 to fill the bill once.
| |||||||||
| 5 | Gemini 3.7 Flash google/gemini-3.7-flash | 29.2% | 130 | — | 43%of 75 | 2%of 52 | $0.07 | ||
Gemini 3.7 Flash by set38 of 130 lines inside tolerance. $0.07 to fill the bill once.
| |||||||||
| 6 | Gemini 3.6 Flash google/gemini-3.6-flash | 27.7% | 130 | — | 39%of 75 | 5%of 52 | $0.12 | ||
Gemini 3.6 Flash by set36 of 130 lines inside tolerance. $0.12 to fill the bill once.
| |||||||||
| 7 | GPT 5.5 openai/gpt-5.5 | OpenAI | 26.5% | 130 | 46 | 39%of 75 | 0%of 6 | $0.81 | |
GPT 5.5 by set35 of 130 lines inside tolerance, 46 declined. $0.81 to fill the bill once.
| |||||||||
| 8 | Gemini 2.5 Pro google/gemini-2.5-pro | 26.2% | 130 | 17 | 40%of 70 | 3%of 40 | $0.29 | ||
Gemini 2.5 Pro by set34 of 130 lines inside tolerance, 17 declined. $0.29 to fill the bill once.
| |||||||||
| 9 | GPT 4.1 openai/gpt-4.1 | OpenAI | 25.4% | 130 | 55 | 42%of 67 | 0%of 5 | $0.17 | |
GPT 4.1 by set33 of 130 lines inside tolerance, 55 declined. $0.17 to fill the bill once.
| |||||||||
| 10 | GPT 5.6 Luna openai/gpt-5.6-luna | OpenAI | 24.6% | 130 | 82 | 64%of 44 | n/a | $0.06 | |
GPT 5.6 Luna by set32 of 130 lines inside tolerance, 82 declined. $0.06 to fill the bill once.
| |||||||||
| 11 | Gemini 2.5 Flash google/gemini-2.5-flash | 23.8% | 130 | 21 | 38%of 68 | 5%of 40 | $0.06 | ||
Gemini 2.5 Flash by set31 of 130 lines inside tolerance, 21 declined. $0.06 to fill the bill once.
| |||||||||
| 12 | Gemini 3.1 Flash Lite google/gemini-3.1-flash-lite | 23.1% | 130 | 53 | 35%of 74 | n/a | $0.02 | ||
Gemini 3.1 Flash Lite by set30 of 130 lines inside tolerance, 53 declined. $0.02 to fill the bill once.
| |||||||||
| 13 | GPT 5.6 Terra openai/gpt-5.6-terra | OpenAI | 21.5% | 130 | 94 | 66%of 35 | n/a | $0.54 | |
GPT 5.6 Terra by set28 of 130 lines inside tolerance, 94 declined. $0.54 to fill the bill once.
| |||||||||
| 14 | Claude Opus 4.8 anthropic/claude-opus-4.8 | Anthropic | 20.8% | 130 | 100 | 78%of 32 | n/a | $1.09 | |
Claude Opus 4.8 by set27 of 130 lines inside tolerance, 100 declined. $1.09 to fill the bill once.
| |||||||||
| 15= | Claude Opus 4.7 anthropic/claude-opus-4.7 | Anthropic | 20.0% | 130 | 88 | 55%of 44 | n/a | $0.76 | |
Claude Opus 4.7 by set26 of 130 lines inside tolerance, 88 declined. $0.76 to fill the bill once.
| |||||||||
| 15= | Claude Sonnet 4.5 anthropic/claude-sonnet-4.5 | Anthropic | 20.0% | 130 | 61 | 37%of 62 | 0%of 4 | $0.37 | |
Claude Sonnet 4.5 by set26 of 130 lines inside tolerance, 61 declined. $0.37 to fill the bill once.
| |||||||||
| 17 | Mistral Large 2512 mistralai/mistral-large-2512 | Mistral | 19.2% | 130 | 70 | 39%of 57 | n/a | $0.03 | |
Mistral Large 2512 by set25 of 130 lines inside tolerance, 70 declined. $0.03 to fill the bill once.
| |||||||||
| 18 | Claude Fable 5 anthropic/claude-fable-5 | Anthropic | 18.5% | 130 | 100 | 69%of 32 | n/a | $2.11 | |
Claude Fable 5 by set24 of 130 lines inside tolerance, 100 declined. $2.11 to fill the bill once.
| |||||||||
| 19= | Claude Haiku 4.5 anthropic/claude-haiku-4.5 | Anthropic | 15.4% | 130 | 94 | 50%of 40 | n/a | $0.13 | |
Claude Haiku 4.5 by set20 of 130 lines inside tolerance, 94 declined. $0.13 to fill the bill once.
| |||||||||
| 19= | GPT 4o openai/gpt-4o | OpenAI | 15.4% | 130 | 57 | 28%of 67 | 0%of 4 | $0.17 | |
GPT 4o by set20 of 130 lines inside tolerance, 57 declined. $0.17 to fill the bill once.
| |||||||||
| 19= | GPT 5.4 Mini openai/gpt-5.4-mini | OpenAI | 15.4% | 130 | 107 | 68%of 22 | n/a | $0.18 | |
GPT 5.4 Mini by set20 of 130 lines inside tolerance, 107 declined. $0.18 to fill the bill once.
| |||||||||
| 22 | Claude Sonnet 5 anthropic/claude-sonnet-5 | Anthropic | 11.5% | 130 | 116 | 79%of 17 | n/a | $0.45 | |
Claude Sonnet 5 by set15 of 130 lines inside tolerance, 116 declined. $0.45 to fill the bill once.
| |||||||||
| 23 | Gemini 3.1 Pro Preview google/gemini-3.1-pro-preview | 10.8% | 130 | 109 | 52%of 23 | n/a | $0.19 | ||
Gemini 3.1 Pro Preview by set14 of 130 lines inside tolerance, 109 declined. $0.19 to fill the bill once.
| |||||||||
| 24= | Mistral Medium 3.1 mistralai/mistral-medium-3.1 | Mistral | 0.0% | 130 | 134 | n/a | n/a | $0.02 | |
Mistral Medium 3.1 by set0 of 130 lines inside tolerance, 134 declined. $0.02 to fill the bill once.
| |||||||||
| 24= | GPT 5 Nano openai/gpt-5-nano | OpenAI | 0.0% | 130 | 134 | n/a | n/a | $0.01 | |
GPT 5 Nano by set0 of 130 lines inside tolerance, 134 declined. $0.01 to fill the bill once.
| |||||||||
Reading the columns
Declined
Lines the model handed back with no number, saying the drawings do not show it. It sits beside the score, never inside it: a declined line earns nothing, but refusing and being confidently wrong are different failures, and on a real bill only one of them is recoverable.
Counts
Lines you get by counting tagged items on a plan — fixtures, fittings, outlets. Scored exactly: the number matches the estimator's or it does not. Most of the bill is this, and it is the work most people want to hand over first.
Measured
Lines you get by measuring off the drawing — pipe and cable runs by diameter, each a sum over dozens of branches. Scored inside ±5%, a tolerance we state rather than one anyone measured. This is where the takeoff is actually won, and it is where every model on this page still falls over — some by returning a number that misses, others by declining the line altogether.
Read the rate with the count beside it. Every basis figure here is scored over the lines that model actually put a number against, so declining the hard ones raises the percentage. Half of eight measured lines attempted is a worse takeoff than three per cent of fifty, and it prints as the better number. The count under each rate is how many lines it is over, and the declined column is the rest.
Correct covers every line on the bill, including the few read off a schedule that neither column above breaks out. Cost is what one pass over the bill cost through the API; an agent session runs on a flat subscription, so there is no per-run figure to print. Every model on this board sat every set, over the same 130 lines; a model that cannot sit one of them is withheld rather than ranked on fewer. Every number comes from the same scorer that runs internally, off run records kept on disk. Drawings are never published. Full method and the raw results: ContractorOS.