Every week, somewhere in India, a site engineer pastes a plan into ChatGPT and types "prepare BOQ". The output looks magnificent — a properly formatted table, plausible item descriptions, rates in rupees, a grand total. It looks like three days of quantity surveying done in thirty seconds.
Then a senior engineer reads it, and the magnificence falls apart line by line: quantities that correspond to nothing on the drawing, rates that exist nowhere in the market, deduction rules ignored, and a confident total that would sink the tender it was pasted into.
This post is an honest accounting of what general-purpose AI chatbots can and cannot do in quantity surveying — where they genuinely help, where they fail structurally, and what it actually takes to get a defensible BOQ out of AI. We build AI takeoff tooling ourselves, so we have skin in this game; that is precisely why we think the failure modes deserve to be named rather than papered over.
What happens when you ask a chatbot for a BOQ#
Ask ChatGPT, Claude, or Gemini to "prepare a BOQ for a 1,200 sqft G+1 house" and you get a parametric guess: the model recalls what BOQs look like from its training data and fills one in with typical values. Quantities are derived from thumb rules (so much concrete per sqft, so much steel per cum), not from any drawing. It is a ₹/sqft calculator wearing a QS costume.
Upload an actual drawing and the situation improves less than you would expect. Practitioners who have tested this publicly report the same cluster of failures. A documented test of ChatGPT 5 against a real BOQ found miscalculated items, invented rates, and quantities mismatched against their own line items. Quotr's analysis of ChatGPT for estimating — written by a company selling AI estimating, with every incentive to oversell — puts vision-model accuracy on construction drawings at roughly 65–75% and is blunt that raw chatbots cannot do takeoff. Our own experience building drawing-measurement pipelines is consistent with both.
The five structural failures#
These are not bugs that the next model version fixes. They follow from what a general-purpose chatbot is.
1. It does not measure — it recalls#
A quantity takeoff is a measurement exercise: find every member, read its dimensions, multiply, deduct openings, sum. A chatbot generates text that resembles previous takeoffs. When it writes "brickwork: 38 cum", that number is not the product of counting your walls — it is a statistically plausible value for a building of the described size. Change your design and the number does not change with it. That is the deepest problem: the output is untethered from your drawing.
2. It cannot hold a drawing set in its head#
Real takeoff is a cross-referencing exercise — the plan gives wall lengths, the section gives heights, the schedule gives openings, the structural sheet gives member sizes, and the revision block tells you which sheet to trust. Vision models read one image at a time, lose scale information (a dimension line it misreads by one digit propagates everywhere), and struggle with the dense, layered line work of Indian GFC sets, which are frequently scanned or photographed rather than clean vector PDFs.
3. It does not know IS 1200 — it only knows about it#
Ask a chatbot to explain IS 1200 deduction rules and it answers correctly: plaster openings up to 0.5 sqm are not deducted; larger openings are. Ask it to prepare a plaster quantity and it silently does neither — no opening schedule is consulted, no deduction lines appear, no jamb and sill measurements are added. Knowing a rule and applying it across two hundred members are different capabilities, and the second one is the job.
4. It invents rates#
Rates are local, current facts: cement in Jaipur this month, mason wages in Kochi, the applicable DSR edition, GST treatment for your contract type. A language model has none of this — it has the fossil record of rates from its training data, which it serves up with total confidence. Every practitioner test we have seen flags hallucinated rates as the most dangerous failure, because a wrong rate on a correct quantity still produces a confidently wrong total.
5. It shows no working#
When a client's engineer challenges a quantity in a real bill, you open the measurement sheet and point to the line: 12 footings × 1.5 × 1.5 × 1.5, per drawing SD-02. A chatbot BOQ has no measurement sheets. There is nothing behind the numbers — no member list, no dimensions, no drawing references, no deduction lines. It is unauditable, and an unauditable BOQ is professionally worthless the moment money depends on it, whether in a tender, an RA bill, or arbitration.
What chatbots are genuinely good at in QS work#
The honest ledger has a credit side. Used as an assistant rather than a surveyor, a chatbot is legitimately useful for:
- Drafting item descriptions. Give it the facts — M25, isolated footings, formwork included, reinforcement excluded — and it produces clean, tender-grade phrasing faster than most engineers write it.
- Explaining rules and formulas. Cutting-length formulas, d²/162, deduction thresholds, DSR structure — as an interactive reference it is excellent, and far faster than paging through the code.
- Checking structure and arithmetic. Paste a draft BOQ and ask what trades look missing, or whether steel-per-cum ratios look sane for the member types. As a second pair of eyes it catches real omissions.
- Spreadsheet mechanics. Formulas, formats, and automation for your own sheets — including the ones in our free QS Starter Kit.
The pattern: chatbots are strong wherever the task is language and rules, and weak wherever the task is measurement and current facts.
What it actually takes for AI to do takeoff#
The failures above are failures of a general-purpose chat interface, not of AI as such. Purpose-built takeoff systems close the gap by changing the architecture, not the prompt:
- Geometry first. The drawing is processed as a drawing — scale calibrated, members detected and dimensioned as geometric objects — so quantities are computed from measurements, not recalled from training data.
- Measurement rules as code. IS 1200 conventions — units by trade, the 0.5 sqm plaster deduction threshold, net-in-finished-position measurement — are enforced deterministically, applied to every member every time.
- Rates from a live library, or from you. DSR editions and current market rates are looked up, not imagined — or the BOQ ships quantities-only and your team prices it.
- Measurement sheets as the output contract. Every quantity carries its L×B×D lines and drawing references, so the AI's work can be audited exactly like an engineer's — line by line, against the drawing.
- An engineer in the loop. Model output is reviewed by a human who signs it. Not because it fails constantly, but because a BOQ is a professional document and 90–95% right is not a standard anyone tenders on.
That is the pipeline we run at SiteSetu: AI does the measurement from your GFC drawing, the rules engine does the deductions, an engineer verifies the sheets, and the deliverable includes the full measurement trail. If you want to judge the difference yourself, send us one drawing — the first BOQ is free and comes back within 48 hours with every measurement sheet attached. Check our quantities against your own takeoff; that comparison is the entire argument.
For the manual method the AI is being measured against, our step-by-step guide to preparing a BOQ from drawings walks the full workflow.
FAQs#
Can ChatGPT read construction drawings at all?#
Partially. Vision models can identify rooms, read printed dimension text, and describe a floor plan in general terms. Reliability degrades sharply on dense GFC sheets, scanned or photographed drawings, and anything requiring scale measurement or cross-referencing between sheets — which is most of real takeoff. Published assessments put drawing-reading accuracy around 65–75%, which sounds high until you remember a BOQ multiplies every reading error by a rate.
Is there an AI that prepares accurate BOQs from drawings?#
Purpose-built takeoff tools exist and are improving quickly — the credible ones combine geometric measurement with human review rather than asking a chatbot to imagine quantities. What does not exist, in India or anywhere, is a system where you should tender on unreviewed AI output. Treat "AI-prepared, engineer-verified, with measurement sheets you can audit" as the minimum standard, and ask any vendor — including us — to show the measurement sheets.
Is it safe to use ChatGPT for rate analysis?#
For the structure of a rate analysis — coefficients, the build-up logic, T&P and OH&P percentages — yes, it is a good tutor; see our rate analysis guide for the standard workings. For the rates themselves, no. Material and labour prices are current local facts; get them from quotations, DSR editions, or your own purchase records, never from a model's memory.
Will AI replace quantity surveyors?#
The measurement labour — the two days of L×B×D arithmetic — is being automated, in the same way total stations automated chain surveying. The judgment is not: deciding what an ambiguous drawing means, which assumptions to disclose, whether a rate is defensible, what a variation is worth. QS work is shifting from producing quantities to verifying and arguing them, which has always been the higher-value half of the job.
References and Further Reading
Primary and supporting sources cited in this article.
Tags: