How to Structure Construction Cost Data with AI
Structuring cost data is the job every estimator puts off: taking old project costs sitting in whatever format the last spreadsheet used and reclassifying them into one coded library, CSI MasterFormat, ICMS, or a custom scheme, so the next bid can reuse them. It is also the foundation that makes AI estimating useful. This comes from a build where I reclassified real cost data from the Miramar Community Center project, and the chat-versus-Cowork lesson it surfaced.
Key takeaways
- A cost code library only pays off if every project uses the same structure. CSI MasterFormat (North America) and ICMS (UK, NZ, Australia) are the two standard schemes, and either works as long as it stays consistent across jobs.
- AI is well suited to the reclassification itself: reading a line item’s description and assigning it the correct code, line by line, across a messy historical cost sheet.
- Running the classification in a normal chat window instead of Cowork breaks the output. Chat rebuilds the workbook template from scratch and drops the summary formulas; Cowork updates the actual file and keeps them intact.
- The rate and the number behind each cost code stay a human call. AI structures and cross-checks the data; it does not set the price.
What is a cost code, and why does it need a structured library?
A cost code is a fixed reference number that ties one piece of scope, formwork, site excavation, structural steel, to a bucket you track cost against across the whole project. The same code should carry through the estimate, the budget, the schedule of values, the payment claim, and the actuals you feed back into the next bid. Without that consistency, every project builds its own private language and none of the data compounds.
That is the actual payoff of a coded library. In the cost-tracking method I teach, cost codes sit at the centre of the whole system: the scope gets broken into separate budgets, each one carrying an original budget of quantity times rate, variation adjustments, cost to date, and a forecast to complete. The list of cost codes is the work breakdown structure, not a side reference to it. And an estimate is not a budget on its own; it is compiled to win the work, not to track spending, so it needs to be converted into that cost-code structure before it is any use for cost control.
CSI MasterFormat is the version of this most contractors in North America already half-know. It assigns a fixed number to every construction task, filed by trade and material, so a division like concrete or a section like cast-in-place formwork means the same thing on every project. Where it earns its keep is that the same numbering can run across specifications, estimates, cost codes, cost tracking, schedules of values, and payment claims. One structure across the whole lifecycle, rather than five different ones that all have to be manually reconciled.
Why does messy historical cost data get in the way of AI estimating?
Old project costs are only useful to an estimator if they sit under the same reference scheme the current estimate uses, and most historical cost sheets do not. Two “site excavation” line items from two different jobs often carry two different labels, because nobody applied one shared classification when either cost was first recorded. Before AI can help with anything covered in AI for construction estimating, the historical data it is meant to draw from has to be queryable, which means every line needs a code.
This is the same pattern behind free AI for construction estimating: the strongest AI use cases in estimating are document work, reading, structuring, and cross-checking, not deciding the number. Reclassifying a cost sheet into CSI MasterFormat or ICMS is squarely that kind of task. It is grounded (a real line item with a real description), it is verifiable (you can check the assigned code against the description in seconds), and the blast radius of a wrong classification is small and easy to catch, right up until nobody checks it and a wrong code quietly corrupts a future rate.
How does an AI cost data structuring skill actually work?
The process is a straightforward read, classify, and populate loop: load the cost sheet, tell the skill which classification system to use, let it go line by line, then review before it feeds anything live. It classifies every line, keeps the original value attached, and hands back a workbook meant to be ready for the summary rollup. Here is how I ran it:
- Load the skill and the raw cost sheet. Upload the historical cost data, a list of line items with descriptions and values, into a Cowork project that already has the classification skill installed.
- Pick the classification system. CSI MasterFormat for North America, ICMS for UK, NZ, or Australia, or upload your own scheme if the business already runs one.
- Let it work line by line. The skill reads each item’s description and assigns it a code from the reference list built into the skill, keeping the original value attached.
- Expect a tool-limit prompt on a longer sheet. Hitting a tool limit mid-run is normal on a bigger cost sheet. Say continue and it picks up where it left off.
- Review every reclassified line against the source sheet. The skill is matching descriptions to codes, not verifying scope, so a genuinely ambiguous line item needs a human read before it goes anywhere near a live estimate.
On the Miramar Community Center project, I ran this against 15 cost lines and every one reclassified correctly into the right CSI MasterFormat bucket in one pass, each carrying its original value across into the new structure.
Why does the chat window get the summary table wrong?
Running this kind of skill in a normal chat window can classify the line items correctly and still hand back a broken workbook, because chat cannot run the scripts the skill depends on. It rebuilds the workbook template from a blank sheet instead of editing the file you gave it, and that is where the summary formulas break. When I first ran the classification in chat, the classification itself worked fine. What broke was the summary table, the tab that rolls every classified line up into totals by CSI MasterFormat bucket.
Chat does not run scripts. So instead of updating the existing spreadsheet template in place, it recreated the template from scratch, which is the only thing it is capable of doing in that mode. That meant the formulas driving the summary totals did not carry through. Asking it directly to “complete the summary table” produced a written breakdown but still left the actual spreadsheet formula broken. Only an explicit instruction to update the spreadsheet got closer, and even then it effectively rebuilt the sheet rather than editing it. As I said at the time:
“This is actually a good example of why using Co-Work can be a lot better than using the chat window.” (Tim Fairley, How to Structure Construction Cost Data with AI)
Cowork runs the skill’s underlying script against the real file instead of reconstructing it from a prompt, so the formula-driven summary comes out correct the first time.
| Chat window | Cowork | |
|---|---|---|
| Classifies line items | Yes | Yes |
| Runs the skill’s scripts | No, recreates the template from scratch | Yes, updates the actual file |
| Summary formulas | Break; often need an explicit “update the spreadsheet” instruction and a redo | Carry through correctly |
| Best for | A quick one-off lookup on a single line item | Any skill built on scripts, including the full cost-data structuring workbook |
CSI MasterFormat vs ICMS: which classification do you use?
Which classification scheme you use comes down to region, not which one performs better. CSI MasterFormat is the standard in North America, and ICMS, the International Construction Measurement Standards, is the more common reference across the UK, New Zealand, and Australia. Both do the same job: give every cost item one fixed code so figures compare cleanly across projects and years, whichever scheme your region has settled on.
The gap this closes is real. Outside North America, specification and cost structures have historically been far less consistent. I have said before that “in North America, for example, they use the CSI master format, which I think is a much better way to do it,” compared to the patchwork of NATSPEC and project-specific specification styles common in Australia. That inconsistency is exactly what a shared cost code library is meant to fix, whichever scheme you standardise on.
| Scheme | Primary region | What it structures |
|---|---|---|
| CSI MasterFormat | North America (US, Canada) | Numbered divisions and sections by trade and material, e.g. Division 03 Concrete |
| ICMS | UK, New Zealand, Australia | International cost classification aligned to elemental and measurement standards |
| Custom | Any | Your own reference scheme, uploaded once, applied with the same line-by-line logic |
If your business already runs its own cost code scheme, the same skill works against it. You upload the reference list once and the classification logic stays the same, whether it is matching against CSI MasterFormat, ICMS, or a homegrown system built around how your projects are actually run. The Cost Data Structurer skill, and the CSI MasterFormat and ICMS reference libraries it runs against, are in the ContractorOS community if you want to run this against your own historical cost sheet.
Common mistakes when structuring cost data with AI
- Running a scripted skill in chat. If the output is a workbook with formulas, chat will recreate the template instead of editing it, and the summary will break. Use Cowork for anything that touches an existing file’s structure.
- Skipping the review pass. The skill classifies by description, not by project context. A line item worded ambiguously can end up in the wrong bucket, and the same trust-boundary caution behind how accurate is AI estimating applies here: the accuracy comes from the estimator checking the output, not from the model.
- Treating an estimate as a cost code library. An estimate is compiled to win the work, not to track spending against. It has to be converted into cost codes, quantity times rate per code, before it is useful for cost control.
- Ignoring the rate baked into a code. A cost code inherits whatever rate assumption sits behind it. In my estimating guide I walk through a carpenter paid $45 an hour who actually costs the business $62 an hour once on-costs, superannuation, insurance, leave accruals, are factored in. Get that on-cost wrong once inside a code and every future estimate that pulls the rate from that code repeats the same error.
- Building the library once and never feeding it back. The actual value compounds when actuals from a finished project get structured back into the same codes the next estimate will query. A library that only ever gets read and never updated stops paying for itself.
Watch the full walkthrough above for the Miramar Community Center reclassification on screen, chat attempt and Cowork fix included.
- ContractorOS: How to Structure Construction Cost Data with AI (CSI MasterFormat / ICMS) (source video; the Miramar Community Center demo, 15 cost lines reclassified, the chat-vs-Cowork summary-table result)
- Tim Fairley: Construction Take-offs - Complete Guide (CSI MasterFormat vs NATSPEC framing: "a much better way to do it")
- Tim Fairley: The Complete Guide to Construction Estimating (the loaded labour rate example: $45/hr base wage to $62/hr once on-costs are factored in)
Frequently asked questions
What is CSI MasterFormat?
CSI MasterFormat is a standard numbering system used mainly in North America that classifies construction work into fixed divisions and sections by trade and material, like a Dewey decimal system for construction. The same code structure can run across specs, estimates, cost codes, budgets, payment claims and cost tracking.
What is ICMS and how is it different from CSI MasterFormat?
ICMS, the International Construction Measurement Standards, is the cost classification more common in the UK, New Zealand and Australia. It serves the same purpose as CSI MasterFormat, giving every cost item one fixed code so figures compare across projects. Which one you use depends on region, not which is better.
Can AI reclassify historical cost data without a human checking it?
No. AI reads each line item's description and assigns the closest matching code, which is reliable pattern matching, but it can misclassify an ambiguous description or miss a line entirely. Every reclassified sheet needs a review pass against the source data before it feeds a live estimate or budget.
Why does a data-structuring skill work better in Cowork than in chat?
Skills that rely on scripts, like updating a spreadsheet's summary formulas, cannot run those scripts inside a normal chat window. Chat recreates the template from scratch instead of editing the actual file, which breaks the formulas. Cowork runs the underlying script against the real workbook, so the output comes out correct.
Do I need a cost code library if I already use estimating software?
Yes, usually. Estimating software measures quantities and applies rates, but it does not force old project costs into one shared structure on its own. A cost code library is what lets you query what something cost last time, across every job, which is the reusable asset that speeds up future estimates.
How does a structured cost code library make future estimates faster?
Once historical costs sit under the same codes as your current estimate, AI can pull comparable rates, quantities and totals from past projects instead of starting from a blank pricing schedule. The estimator still checks and adjusts every rate, but the starting point comes from real project data, not a guess.
Want this working in your company?
ContractorOS members get the skills, templates and weekly live calls to implement it on real projects.
Join ContractorOS →