Benchmarks

How well can AI read construction drawings?

169 questions an estimator wrote against 9 real sets. We gave them to 30 frontier models through the raw API — no tools, no help, just the model and the drawing — and then we ran them again inside a harness (Cowork, ChatGPT Work, Claude Code, Cursor, Codex): 14 model-and-harness pairs, with the drawings in a folder they could open.

169
Questions
9
Real drawing sets
30
Raw models
14
Harness + model
82.1%
Best model alone
92.5%
Best in a harness

The same questions, inside a harness

Agents working the drawings out of a folder, the way you would. Click a row for its category breakdown.

# Harness Model Accuracy vs raw Details

One session per pairing, run once. Repeat runs of this benchmark move by 1.1 to 5.4 points on their own, so read anything closer than that as a tie rather than an order. Agents worked however they liked, and some built themselves crops, OCR dumps and index files along the way.

And the raw models on their own

No tools, no folder, no second look. Accuracy is the share of answerable questions it got right. Click a row for its category breakdown.

# Model Lab Accuracy Cost Details

Cost is one full run of all 169 questions, model usage only. Scored on the native PDF lane. Last updated September 2026.

The questions

Written by an estimator

169 questions across architectural, structural, electrical, mechanical and plumbing sets. Every answer was worked out and checked by a practising estimator before any model saw the drawings.

The test

Same drawings, same questions

Every model and every agent gets the identical question and the identical scoring. No tuning toward any vendor, and nothing pre-digested. What differs between the two tables is the sitting: a raw model is handed one page, an agent is handed the folder and left to work.

The harnesses

Whatever the harness would really do

Each agent got one session and was left alone in it. Some worked straight off the PDFs; some built themselves image crops, OCR text dumps and index files first. That is part of what a harness is, so it is left in — which does mean two harnesses on the same model are not a clean like-for-like.

The scoring

Right or wrong, nothing in between

An answer is correct or it is not. Half a right answer to a three-part question is a wrong answer to it. Three of the questions cannot be answered from the drawings at all, and saying so is the correct response.

The categories

What we asked them to do

The questions split into 9 kinds of work. They are not equally hard, and the gap between them is the useful part: a model can be excellent at reading a schedule and still be unable to count.

Reading the sheet

32 questions

Finding a fact that is written down somewhere. Title block, general notes, the legend. The drawing says it outright, the model just has to go and find it.

Schedules

18 questions

Pulling one value out of a table. Door schedules, panel schedules, equipment lists. Easy to eyeball, easy to get wrong when the table runs across half a sheet.

Counting

41 questions

Counting tagged items on a plan. How many of this fitting, how many of that outlet. This is the core takeoff skill and the one most people want to hand over first.

Element detail

25 questions

The size, type or material of a named element. What that beam is, what that pipe is made from, what size that cable run is.

Scale and measurement

22 questions

Distances, areas and volumes that need scale reasoning. The number is not printed anywhere, so the model has to work it off the drawing itself.

Cross-referencing

14 questions

Answers that need more than one sheet, or tracing a single line diagram through. Nothing on its own page gives you the answer.

Scope

7 questions

What is in and what is out. New versus existing, who supplies it, whether it sits in this package or someone else’s.

Sheet structure

7 questions

Reading the drawing as a document rather than a picture. How many details are on the sheet, how many grid lines, what revision it is at.

Knowing when to stop

3 questions

Questions the drawings genuinely do not answer. The right response is to say so. A model that invents a number here is the one that costs you money.

We re-run this on every major release

When a new model ships it goes through the same 169 questions. Join the community and you get the results as they come in, plus the workflows we build off the back of them.

Get the results first