Beyond BIM: Is AI Changing What BIM Was Meant to Be?
BIM organises information; AI promises to interpret it. What the evidence actually shows, and where human judgment still decides.
Hand nine language models a building model and one sentence of plain English — move that door, add a partition here — and ask them to make the edit. Somebody ran roughly that experiment: 324 tasks across eleven real buildings and thirty-six synthetic scenes. The best of the nine scored 49.5% once geometry, semantics and topology were counted together. None finished more than 3.4% of the tasks cleanly.5
Software that promises to do exactly this is being sold, this month, to your architect and your contractor. So the question worth asking is narrower than the one in the headlines.
Does AI make BIM less important?
We think it makes BIM harder to skip, and the evidence is narrower than the enthusiasm. For a machine reading building data, the structure and provenance of that data decide whether the reading is worth anything. This was already true when a human was the only reader; a human just noticed when the data was wrong. We would call that a well-supported reading of the studies below and stop short of calling it a law — it says nothing about every use of AI in construction.
We also build our own tools, which has given us the useful experience of watching several of them fail in ways no vendor demonstration shows you. Those failures are in here too.
BIM got us here
BIM's value never came from models being three-dimensional. It came from the information inside them being standardised, dated, and open to being checked by someone other than its author.
That is a documented position, not a marketing one. The current schema for exchanging building models is IFC 4.3, published as ISO 16739-1:2024.1 The way information requirements are stated so that a machine, and not only a person, can verify whether they were met is defined by the Information Delivery Specification, approved as a final buildingSMART standard in 2024.2 The framework for managing information across a project's life is ISO 19650, and here a detail gets quoted wrongly: the published basis is still the 2018 edition. A second edition of ISO 19650-1 is in draft and closed its voting stage on 3 June 2026. What happens next is not determined, and nobody should be writing requirements against it yet.3
BIM solved the problem of everyone working from the same information. It never solved the problem of deciding what to do with it. That is where AI enters, and where the claims start to outrun the evidence.
What the evidence actually shows
The research worth reading tends to be the research measuring where these systems stop working. That is where a project gets hurt.
Chen Liang and colleagues put nine models through 4,800 questions drawn from real practice, across five levels of difficulty and 23 tasks, and published the result in Advanced Engineering Informatics this year — the first benchmark of its kind in this field to clear peer review. The models recalled facts well. They failed where construction actually lives: reading a table inside a building code, doing the arithmetic, writing a document a professional would file.4
Two preprints from June sharpen the picture. Bharathi Kannan Nithyanantham's group asked models to edit IFC models from written instructions — 324 tasks, 11 real buildings, 36 synthetic scenes. Their best model scored 49.5% once geometry, semantics and topology were counted together, and none of the nine finished more than 3.4% of the tasks cleanly.5 In Japan, Ryo Kanazawa's team sat six BIM specialists down to write 166 worked examples across 83 practical scenarios, in Japanese and English, and then asked the models to produce the same Information Delivery Specifications. The best content pass rate was 33.1%.6
Read those numbers as an engineer would. A model on its own, pointed at a structured building task, produces work nobody should sign. The peer-reviewed work seems to converge on where the value sits: in the deterministic checks placed around the model rather than in the model itself. Valinejadshoubi, Moselhi, Iordanova and Valdivieso, writing in Buildings in 2025, put a quantity precision check before the takeoff runs — completeness, consistency, classification — so errors surface while they are still cheap.7 Abdelsalam, Ashmawi and Nguyen, in the same journal a year later, wired a language model to a cost database and then hemmed it in with a mandatory schema and an arithmetic reconciliation, so a number that does not add up cannot leave the system.8 A third group took the idea into training itself, rewarding the model against an external validator.9
We treat this as a well-supported architectural hypothesis rather than settled fact. The half of the chain showing that models fail alone rests on preprints not yet peer reviewed. The half showing that external verification recovers the result rests on two peer-reviewed studies78 and one preprint9 — stronger, but not uniformly reviewed, and we will not describe it as such. Convergence between independent groups is good evidence, and short of proof.
What actually happens when you run this over real project data
This is the part vendor material does not cover, and the part that changed how we build.
A check that can pass by not checking is worse than no check. Our own compliance panel reported a clean result on models where nothing had been verified. A requirement that had been skipped was counted as one that had passed, and where no requirement applied at all, the panel reported full compliance instead of "not applicable". The percentage was arithmetically correct and completely meaningless. Verification output needs three states: compliant, non-compliant, and not verifiable. Until it has them it should report and never block, because a gate you satisfy by evaluating nothing hands an approval stamp to work nobody looked at.
Documentation drifts away from behaviour. Auditing our own scheduling engine, we found a comment announcing that the schedule was levelled for resource contention. The step underneath did nothing of the kind; it rescheduled with tighter durations. Nothing was broken in a way a test would catch, and anyone reading the file would have believed the opposite of the truth. If you are being sold an automated programme, ask what its output demonstrably did on a project like yours.
The loop that tells you whether the numbers were right is the one nobody builds first. Producing estimates from a model is straightforward. Closing the loop, by recording what the work actually cost and comparing it against the frozen original estimate, is harder and much less exciting. We built that pipework before we had data to put in it, and the table still sits empty. Any firm claiming AI-driven cost accuracy should be able to say how many estimates it has reconciled against completed work, and against which baseline: compare against the latest revision rather than the original and the deviation is always near zero, and always meaningless.
In Portugal there is a structural problem underneath all of this. The bill of quantities is the document a project is priced, tendered and argued from. Until December 2025 there was no Portuguese standard defining how it should be organised, a gap that buildingSMART Portugal and LNEC state plainly as the reason they published a common structure, coded by chapter and subchapter and prepared for mapping to IFC.10 It carries the authority of a national chapter guide rather than an ISO standard, and that changes what can be asked of it. In our own work we compared how two of our projects had been broken down and found the two structures had no level in common. A fixed breakdown does not survive the second project; forced onto two incompatible ones, it produces classification that looks like organisation and behaves like lost information.
The link between a bill line and the object in the model matters too. The reviewed Portuguese work ties the two by the model element's persistent identifier, so selecting a line isolates the elements in three dimensions.11 Matching by text similarity is faster to build and quietly wrong: it produces confident links between things that share vocabulary and nothing else.
The tools are not changing, they are becoming smarter
Reality capture is the usual example, and it deserves an honest account.
Scanners, drones and 360 cameras do produce accurate records of a site. In WALLNUT's own practice the newer capture methods have paid off fastest in progress documentation rather than in precision survey — that is our experience, not a measurement of the market. The distinction earns its keep. Building our own capture pipeline, the first nine processing runs produced nothing usable, and the one run that completed produced a result that was, measured on its own output, an invisible fog. The fix, as far as we could tell, was a different path from capture to camera positions, with the same model. There is also a limit worth knowing before anyone promises measurements from a 360 camera: a model reconstructed from 360 video carries no metric scale unless something of known dimension is included in the capture.12 Without it, you have a beautiful record you cannot take a dimension from.
Comparing a scan against the design model is genuinely valuable, and the correct pattern is established in the industry — this is the pattern as it is documented, not a description of something we operate today. The system computes deviation, clusters millions of points up to the level of the building element, and proposes candidates; a person then accepts or rejects each one before it becomes a tracked item. Auto-creating issues floods the coordination log with noise and destroys trust in the log itself. When an accepted deviation is passed on, the vehicle is BCF, the open coordination format, and a schema detail governs good practice: BCF deliberately separates issue information from model geometry and carries no fields for quantity, cost or progress.13 The measured deviation lives in your own data; BCF points a person at the element. Hiding a number in a free-text description because there is nowhere else to put it is the anti-pattern that breaks every downstream tool.
Human judgment still matters, and now it is specific
The claim that people remain essential is true and usually empty. The research makes it concrete. The value sits in the verifier around the model: the schema that rejects a malformed output, the arithmetic check that refuses a total that does not reconcile, the precision check that runs before the takeoff instead of after, the result that is allowed to say "I could not verify this".
Our own rule for any transformation of project data is a closure equation, applied file by file rather than on average: what went in equals what was preserved, plus what was split, plus what was merged, plus exceptions that are named and justified. Nothing else. A system that covers 95% of a bill of quantities and silently loses the other 5% is worse than one that covers 60% and declares the rest. The first gives a comfortable number and a wrong decision; the second gives you work and a right one.
Expertise now looks largely like knowing which output to distrust, and why.
What this means in Portugal
The Portuguese timetable is not speculative, and it moved recently enough that most published commentary is out of date.
Since January 2024 the law has required architectural projects submitted under the urban planning regime to be modelled digitally and parametrically using BIM from 1 January 2030.14 That date has not moved. What moved, on 1 June 2026, is everything around it. An amendment in force since that date replaced the pilot programme the 2024 law had scheduled to begin on 1 January 2027, and removed from the statute the provision for automatic municipal validation of compliance with municipal plans. In their place, projects above the value threshold set in the Public Contracts Code must now be modelled digitally and parametrically in BIM whenever possible, a duty already in force rather than waiting for 2030, while the technical requirements for BIM files, and for BIM in public procurement, are deferred to regulations not yet issued at the time of writing.15
Alongside it, the Council of Ministers approved PortugalBIM, the national implementation strategy, in May 2026: a six-year implementation period, a maximum of 90 days for IMPIC to present a detailed action plan, and targets including reaching at least 50 municipalities per year and training at least 3,000 professionals.16
WALLNUT's practical reading for a client is this. The obligation is dated, the intermediate duty is already live for larger projects, and the detailed file requirements are still being written. A project starting now will probably be delivered into that regime. Ask a design team whether their models are structured, classified and checkable enough to satisfy requirements that do not exist yet.
Looking forward
BIM has been the foundation of digital construction. AI raises the cost of building on a weak one. That is WALLNUT's reading of the evidence set out above, from a practice that designs and builds under a single contract in Lisbon and São Paulo — and it is why we treat model structure as a delivery obligation rather than a preference.
The significant change is that a model stops being a document and becomes an input to systems that act on it. Everything those systems produce inherits the structure, the provenance and the errors of what they were given.
So the next chapter is BIM held to a higher standard, because for the first time something other than a human is reading it. From 1 January 2030 a Portuguese planning submission has to be modelled in BIM; the intermediate duty is already live for projects above the Public Contracts Code threshold. Between now and then the interesting work is in the proving — being able to show that what a machine produced from your model was right.
Notes
A note on our own evidence. The observations under "What actually happens when you run this over real project data" and in the reality-capture section come from Wallnut's audits of its own systems, dated June and July 2026. They are stated as findings about our tools, not as measurements of the market. No client, project or commercial figure appears in this article.
- ISO 16739-1:2024, Industry Foundation Classes (IFC) for data sharing in the construction and facility management industries, Part 1: Data schema. Corresponds to IFC 4.3.2.0, the current official IFC schema version. Published March 2024. ↩
- buildingSMART International, Information Delivery Specification (IDS) v1.0, approved as a final standard in 2024 (public announcement 4 June 2024). ↩
- ISO project record for ISO/DIS 19650-1 edition 2, stage 40.60 ("close of voting"), 3 June 2026. The published basis for information management remains ISO 19650:2018. Status recorded in Wallnut's internal state-of-the-art review of 26 July 2026; the ISO project page returned HTTP 403 to automated reading, so no subsequent processual state is asserted here. ↩
- Chen Liang et al., "AECBench: A Hierarchical Benchmark for Knowledge Evaluation of Large Language Models in the AEC Field", Advanced Engineering Informatics, vol. 71 (2026), article 104314; preprint arXiv:2509.18776. Peer reviewed. ↩
- Bharathi Kannan Nithyanantham et al., "BIM-Edit: Benchmarking Large Language Models for IFC-Based Building Information Modeling", arXiv:2606.20146 (June 2026). Preprint, not peer reviewed. ↩
- Ryo Kanazawa et al., "Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements", arXiv:2605.22079 (submitted 21 May 2026, revised 7 June 2026): 166 examples across 83 practical scenarios, authored in Japanese and English by six BIM/IDS specialists, each paired with a gold IDS file. Across ten models in zero-shot settings, best facet-level macro F1 65.6%, best content pass rate 33.1%. Preprint, not peer reviewed. ↩
- M. Valinejadshoubi, O. Moselhi, I. Iordanova and A. Valdivieso, "A Cloud-Driven Framework for Automated BIM Quantity Takeoff and Quality Control: Case Study Insights", Buildings 15 (2025), article 3942. Peer reviewed. ↩
- A. Abdelsalam, K. Ashmawi and T. Nguyen, "AI-Driven Automation of Construction Cost Estimation: Integrating BIM with Large Language Models", Buildings 16(3) (2026), article 485. Peer reviewed. The pattern relied on here is the constraint of model output by a mandatory schema and by an arithmetic reconciliation of extended cost against unit cost multiplied by quantity, within a declared tolerance. ↩
- "Ishigaki-IDS", arXiv:2606.08545 (June 2026): an open-weights, verifier-aware model trained with reinforcement from a verifiable reward supplied by an external validator. Preprint, not peer reviewed. ↩
- buildingSMART Portugal with LNEC, Work Breakdown Structure: Proposta de Articulado para o Mapa de Quantidades e Trabalhos (MQT), v01, December 2025. The document names as its motivating problem "a ausência de uma norma portuguesa que defina a organização do MQT". A national chapter guide, not an ISO standard. ↩
- buildingSMART Portugal, "Da WBS ao MQT: o openBIM como padrão", proceedings of the 6th Portuguese BIM Congress (ptBIM 2026), FEUP, 17 to 19 June 2026. Aligns the WBS with the LNEC measurement rules and the SECClasS classification system, and links bill lines to model elements by persistent identifier (GUID). Recorded in Wallnut's internal state-of-the-art review of 26 July 2026; not independently re-fetched for this article. ↩
- Niantic Spatial / Scaniverse official guidance for 360-camera capture, which specifies including a printed calibration target of known size so the reconstruction carries 1:1 metric scale. Verified in Wallnut's reality-capture research of 12 to 15 June 2026, which also records the nine unusable processing runs described above. ↩
- buildingSMART International, BIM Collaboration Format (BCF) and the BCF-API data model. BCF deliberately separates issue information from the model's geometric data, referencing elements by their GlobalId. Wallnut checked the Topic and Viewpoint schemas field by field on 27 June 2026: no fields exist for quantity, cost or schedule. The deviation workflow described (cluster to element level, propose, require human confirmation) comes from the same research, cross-checked against buildingSMART, ClearEdge3D/Topcon, USIBD and NavVis sources. ↩
- Decreto-Lei n.º 10/2024, of 8 January, article 17 ("Projetos em BIM"), n.º 1. Read at source: Diário da República, 1.ª série, n.º 5, 8 January 2024. ↩
- Decreto-Lei n.º 108/2026, of 29 May, article 8 (amending articles 17, 22 and 25 of Decreto-Lei n.º 10/2024), article 10(c) (repealing articles 19, 20 and 21 and paragraphs 2 and 4 of article 22 of the same decree-law), and article 13 (entry into force). Read at source: Diário da República, 1.ª série, n.º 104, 29 May 2026. Article 13(2) places the amendments to Decreto-Lei n.º 10/2024 in force on the first working day after publication; publication fell on Friday 29 May 2026, so 1 June 2026 is derived from that rule and is not written in the diploma. The threshold referenced is that of article 474(3)(a) of the Public Contracts Code. ↩
- Resolução do Conselho de Ministros n.º 89/2026, of 21 May, approving the Estratégia Nacional para a Implementação da Metodologia BIM ("PortugalBIM"). Read at source: Diário da República, 1.ª série, n.º 98, 21 May 2026. Paragraph 9(a): 90 days maximum for IMPIC's detailed action plan. Paragraph 12: six-year implementation period. Paragraphs 5(e) and 5(h): at least 50 entities and at least 50 municipalities per year. Paragraph 5(f): at least 3,000 professionals. Paragraph 16: in force the day after publication. ↩