Benchmark 03 · updated September 2026
Can AI find its way around a project record?
Not a drawing this time — the whole job. 286 files of contracts, registers, claims, correspondence, site records and drawings, the shape a real project folder actually takes. We asked 146 questions about it: what the contract says, how many RFIs are open, what was claimed against what was certified, where a document is filed, and what the record simply does not contain.
This is the one they are good at. The weakest model tested still answered 83.9% of it correctly, invented almost nothing, and reliably said so when the answer was not in the folder.
The results
Click a model for its breakdown by question type.
| # | Model | Lab | Accuracy | Cited the right doc | Made it up | Details |
|---|---|---|---|---|---|---|
| 1= | Gemini 3.7 Flash google/gemini-3.7-flash | 97.7% | 92.9% | 0.0% | ||
Gemini 3.7 Flash by question typeScored on the 130 of 146 questions this model returned a gradeable answer for. It correctly declined 16 of the 16 questions the record cannot answer, and wrongly refused 2.3% of the rest.
| ||||||
| 2= | GPT 5.6 Sol openai/gpt-5.6-sol | OpenAI | 97.7% | 100.0% | 1.6% | |
GPT 5.6 Sol by question typeScored on the 130 of 146 questions this model returned a gradeable answer for. It correctly declined 13 of the 16 questions the record cannot answer, and wrongly refused 0.8% of the rest.
| ||||||
| 3 | Claude Opus 5 anthropic/claude-opus-5 | Anthropic | 97.7% | 100.0% | 0.8% | |
Claude Opus 5 by question typeScored on the 129 of 146 questions this model returned a gradeable answer for. It correctly declined 15 of the 16 questions the record cannot answer, and wrongly refused 1.6% of the rest.
| ||||||
| 4 | Claude Sonnet 5 anthropic/claude-sonnet-5 | Anthropic | 97.5% | 100.0% | 0.0% | |
Claude Sonnet 5 by question typeScored on the 122 of 146 questions this model returned a gradeable answer for. It correctly declined 16 of the 16 questions the record cannot answer, and wrongly refused 2.5% of the rest.
| ||||||
| 5= | Gemini 3.1 Flash Lite google/gemini-3.1-flash-lite | 96.9% | 100.0% | 0.8% | ||
Gemini 3.1 Flash Lite by question typeScored on the 130 of 146 questions this model returned a gradeable answer for. It correctly declined 10 of the 16 questions the record cannot answer, and wrongly refused 2.3% of the rest.
| ||||||
| 6= | Gemini 3.6 Flash google/gemini-3.6-flash | 96.9% | 93.8% | 0.8% | ||
Gemini 3.6 Flash by question typeScored on the 130 of 146 questions this model returned a gradeable answer for. It correctly declined 15 of the 16 questions the record cannot answer, and wrongly refused 2.3% of the rest.
| ||||||
| 7 | GPT 5.6 Luna openai/gpt-5.6-luna | OpenAI | 96.8% | 100.0% | 1.6% | |
GPT 5.6 Luna by question typeScored on the 124 of 146 questions this model returned a gradeable answer for. It correctly declined 14 of the 16 questions the record cannot answer, and wrongly refused 1.6% of the rest.
| ||||||
| 8 | Gemini 2.5 Flash google/gemini-2.5-flash | 95.4% | 99.1% | 3.1% | ||
Gemini 2.5 Flash by question typeScored on the 130 of 146 questions this model returned a gradeable answer for. It correctly declined 14 of the 16 questions the record cannot answer, and wrongly refused 1.5% of the rest.
| ||||||
| 9 | GPT 5 Nano openai/gpt-5-nano | OpenAI | 95.2% | 99.1% | 4.1% | |
GPT 5 Nano by question typeScored on the 124 of 146 questions this model returned a gradeable answer for. It correctly declined 12 of the 16 questions the record cannot answer, and wrongly refused 0.8% of the rest.
| ||||||
| 10 | Qwen3 VL 235b A22b Instruct qwen/qwen3-vl-235b-a22b-instruct | Alibaba | 93.4% | 100.0% | 5.8% | |
Qwen3 VL 235b A22b Instruct by question typeScored on the 121 of 146 questions this model returned a gradeable answer for. It correctly declined 12 of the 13 questions the record cannot answer, and wrongly refused 0.8% of the rest.
| ||||||
| 11 | Gemini 2.5 Flash Lite google/gemini-2.5-flash-lite | 92.9% | 99.1% | 4.8% | ||
Gemini 2.5 Flash Lite by question typeScored on the 127 of 146 questions this model returned a gradeable answer for. It correctly declined 13 of the 16 questions the record cannot answer, and wrongly refused 2.4% of the rest.
| ||||||
| 12 | Mistral Large 2512 mistralai/mistral-large-2512 | Mistral | 87.7% | 100.0% | 10.2% | |
Mistral Large 2512 by question typeScored on the 130 of 146 questions this model returned a gradeable answer for. It correctly declined 12 of the 16 questions the record cannot answer, and wrongly refused 2.3% of the rest.
| ||||||
| 13 | Mistral Medium 3.1 mistralai/mistral-medium-3.1 | Mistral | 83.9% | 100.0% | 9.2% | |
Mistral Medium 3.1 by question typeScored on the 130 of 146 questions this model returned a gradeable answer for. It correctly declined 13 of the 16 questions the record cannot answer, and wrongly refused 7.7% of the rest.
| ||||||
Three models inside a point of each other is not a ranking — read this as one result with three witnesses, not a race. Every model sat all 146 questions once; the accuracy column covers the questions each returned a gradeable answer for, and that denominator is in the breakdown.
One category is measuring us, not them
In this test each question is handed its source documents as one PDF, and those documents are named by filename with the folder path stripped. A handful of the provenance questions ask where a document is filed — which is a folder path the model was never shown. Every model declined them, which is the correct thing to do, and every model was marked wrong for it. The provenance figure is a floor, not a measure of the model. The other five categories are unaffected.
A real project shape, an answer key checked against the documents
The record
286 files — 32.2 MB of Word, Excel and PDF — laid out the way a head contractor actually files a job: contracts and commercial, design and drawings, procurement, site records, quality, document control, correspondence. Frozen and checksummed, so every model saw byte-identical material.
The questions
146 of them, across 6 types, from single-fact lookups to questions that need two registers reconciled against each other. 16 of them have no answer in the record at all — a model that produces one anyway is doing the most expensive thing it can do.
The key
Answers derive from the generation spine and were then verified against the shipped documents — every keyed fact had to appear in the file the key names. An independent human second read of the bank is still outstanding. Answers are graded by a separate model against the key, never against another model's opinion, and it never sees the documents.
The project is not a real one, and that is the point
The project is fictional. Every document in it is generated from one spine and reconciled against that spine, so the record is internally consistent by construction — that consistency is what makes it gradeable. The 25 structural drawing sheets are a real set, with the title block re-attributed to the fictional parties.
A real project record could not be published — it belongs to somebody, and half of it is commercially confidential. A generated one can be published whole, which means anyone can check our working. The trade is that we have to say so plainly, here, above the numbers.
Finding the answer is the easy half
Models navigate a project record almost perfectly. Ask them to read a drawing and they drop to two thirds. Ask them to price one and not a single one has produced a bill you could use.