TL;DR: SmartPM ran 70,000 construction schedules through an analysis engine and found 88% of baseline schedules failed industry quality benchmarks. Meanwhile the best peer-reviewed prediction model I can find dropped from 89% accuracy in training to 76.8% on external validation. AI can forecast schedule slippage usefully. It cannot forecast from a schedule nobody maintains.
The one-sentence answer: Schedule prediction works about three quarters of the time under favourable conditions, and the binding constraint is almost always the quality of your own schedule data rather than the model.
Every construction software vendor now offers to predict your delays.
The pitch is straightforward and the technology behind it is real. What almost nobody mentions is that the prediction is downstream of your schedule, and the industry's schedules are in worse shape than the industry admits.
So this post takes the four things sold together here, one at a time: predicting delays, running the schedule, coordinating subs, and managing multiple projects. They have very different evidence behind them.
How good is your schedule, honestly?
Statistically, not good enough to predict from.
SmartPM published its first State of Construction Scheduling Report on 2 June 2025, built from more than 70,000 construction schedules processed through its analysis engine plus responses from over 3,500 industry professionals. The headline finding was that 88% of baseline schedules failed to meet industry-recognised quality benchmarks.
Three supporting numbers matter as much.
63% of construction professionals acknowledged limited understanding or use of their own schedules. Only one team in four reported updating schedules on time. And 75% of respondents said schedules influence high-level decisions anyway.
Read those together and the picture is uncomfortable. Most baseline schedules fail quality checks, most teams do not update them promptly, most professionals do not fully use them, and three quarters of decisions lean on them regardless.
Any AI reading that schedule inherits every one of those problems.
What does "failed quality benchmarks" actually mean?
Missing logic, hard constraints, and float that does not reflect reality.
The industry reference is the Defense Contract Management Agency's 14-Point Assessment, developed after a 2005 Department of Defense memo mandated integrated master schedules on contracts over $20 million. The checks cover logic, leads, lags, relationship types, hard constraints, high float, negative float, high duration, invalid dates, resources, missed tasks, a critical path test, the Critical Path Length Index and the Baseline Execution Index.
The first check is the one that matters most here. Activities missing a predecessor or successor should not exceed 5% of total tasks, or the schedule fails. An activity with no logic attached to it does not move when anything else moves, which means it cannot show you a delay.
Worth knowing that the protocol itself is contested. In a 2011 analysis, Ron Winter documented that definitions changed across three revisions, that third-party implementations of the checks in scheduling software are "not certified by the DCMA or any other body" with errors evident in some, and that the prohibition on negative lags "is not based upon any universal scheduling principle." That paper is now a legacy source and the protocol remains widely used, which is rather the point: the standard your schedule gets graded against is applied inconsistently by the tools doing the grading.
Can AI actually predict delays?
Yes, with accuracy that is useful and clearly short of decisive.
The most rigorous recent work I could obtain is Tagharobi, Babaeian Jelodar and Susnjak, published in Frontiers in Built Environment on 16 December 2025. They modelled progress across project stages using roughly 218,000 New Zealand construction projects from 2013 to 2022.
Their best model reached 89% accuracy on training data with a Cohen's kappa of 0.72. On external validation against 2020 to 2022 projects it achieved 76.8%, kappa 0.673.
That drop from 89% to 76.8% is the single most useful number in this post. It is what happens when a model trained on history meets conditions it has not seen, and the validation window included a period of genuine market disruption. A vendor quoting a training-set figure at you is quoting the higher number.
Two further details deserve attention. The authors tested Random Forests and Gradient Boosting and selected multinomial logistic regression, a comparatively simple method, for both interpretability and superior performance. More sophistication did not win. And their stated limitations are candid: reliance on historical patterns that may miss unprecedented disruption, dependence on accurate reporting of project variables, inability to capture complex external factors such as supply chain disruption, and limited applicability outside New Zealand without recalibration.
"Dependence on accurate reporting of project variables" is the same finding as the 88%, arrived at from the other direction.
So is prediction worth having?
At roughly three quarters accuracy, yes, provided you treat it as a prompt rather than a verdict.
A model that flags the right project three times in four is genuinely useful for deciding where to look first, which is what a project executive with eleven jobs actually needs. It is not useful for telling a client a completion date, and it is not evidence in a delay claim.
The failure mode to guard against is not the model being wrong. It is the model being confidently wrong in a report that circulates without its confidence interval attached.
Can AI coordinate subcontractors?
It can chase, track and record. It cannot negotiate, and coordination is mostly negotiation.
The genuinely useful applications are administrative: knowing which sub has confirmed which date, noticing that a required submittal has not arrived, flagging that a trade is scheduled into an area another trade has not left, and keeping a record of who was told what and when.
That last one carries more weight than it appears to. Most subcontractor disputes are disagreements about what was communicated, and a system that records notification reliably resolves a category of argument before it starts.
What AI will not do is decide whether to let the framer slip two days to keep the electrician's crew intact next week. That judgment depends on relationships, on who owes whom a favour, and on information that was never written down.
We build the scheduling side of this and the split is visible in what the software does. By late June 2026 our field platform had a dispatch board, crew lanes, a daily view, a Gantt view and job assignment, with an end-to-end test suite over the scheduling flows. All of it presents information and records decisions. None of it makes the call about which crew moves.
Does it help across multiple projects?
This is the strongest case, and it is a triage argument rather than a prediction argument.
If you run three jobs you know which one is in trouble. If you run fifteen, you find out when someone calls. A model that ranks fifteen projects by likelihood of slippage is valuable even at 76.8% accuracy, because the alternative is not perfect knowledge, it is a rotating hunch.
That is precisely what the New Zealand study was built for: portfolio-level progress prediction across project types rather than a single completion date. Read that way, the accuracy figure looks much better, because ranking tolerates error that a specific promised date does not.
What should you fix first?
Your schedule, before you buy anything that reads it.
The order that follows from the evidence above is unglamorous. Get the logic right so activities actually move when their predecessors move. Update on a fixed cadence, since only a quarter of teams manage this. Make sure the people using the schedule understand it, given that 63% report they do not.
Do that and a prediction tool has something to work with. Skip it and you have bought a model that will produce confident forecasts from a document that does not reflect the job.
There is a harder truth underneath. The AGC and NCCER 2025 Workforce Survey found worker shortages were the most cited cause of project delays, affecting 45% of respondents. No model fixes that. Prediction tells you sooner that you are short of people. It does not supply any.
Frequently Asked Questions
Should I trust a delay prediction enough to tell a client?
No. Tell a client what you know and what you are doing about it. A probabilistic forecast shared as a date becomes a commitment you did not intend to make, and the accuracy figures above are nowhere near what a contractual date requires.
Does prediction work on small residential jobs?
Less well, and for a structural reason. These models learn from volume and variation, and a short job with few activities gives them little to work with. The portfolio-ranking benefit also disappears if you are running two projects rather than fifteen.
Will AI scheduling replace a scheduler?
Not on this evidence, and the study above is a good argument against it: the researchers chose a simpler model partly for interpretability, which is another way of saying somebody still needs to understand and defend the output. A scheduler who can also interrogate a model is more valuable than either alone.
What data does a prediction tool actually need from me?
At minimum a properly logic-linked schedule with regular updates, plus historical outcomes from comparable past jobs. The historical piece is what most contractors lack, because closed-out schedules get archived rather than analysed. If you have never compared planned against actual across your last twenty jobs, you have no baseline for a model to learn from.
Can these tools tell me why a project will slip, not just that it will?
Partially, and only if the model was built for it. Interpretable methods can indicate which variables drove a prediction; more opaque ensembles often cannot. If knowing why matters to you, ask specifically whether the tool reports contributing factors, and treat "proprietary algorithm" as a no.
If your schedule would fail a logic check today, that is the first thing to fix and it costs nothing but attention. Prediction tools are worth buying afterwards.
See ClickWerxs AI services or get in touch. Related: what AI actually does for a construction business, how to evaluate an AI vendor, and getting started on a real jobsite.
Sources
- SmartPM Technologies, State of Construction Scheduling Report, released 2 June 2025 — more than 70,000 construction schedules processed through the company's analysis engine plus responses from over 3,500 industry professionals; 88% of baseline schedules failed to meet industry-recognised quality benchmarks; 63% of construction professionals acknowledged limited understanding or use of their schedules; only one in four teams reported updating schedules on time; 75% indicated schedules influence high-level decisions. Vendor-published research; figures are the publisher's own. smartpm.com
- Tagharobi, M., Babaeian Jelodar, M. and Susnjak, T., "Data-driven progress prediction in construction: a multi-project portfolio management approach," Frontiers in Built Environment, 16 December 2025 — approximately 218,000 New Zealand construction projects, 2013–2022; multinomial logistic regression selected over Random Forests and Gradient Boosting for interpretability and superior performance; 89% accuracy on training data with Cohen's kappa 0.72, and 76.8% accuracy with kappa 0.673 on external validation against 2020–2022 projects; stated limitations include reliance on historical patterns, dependence on accurate reporting of project variables, inability to capture complex external factors such as supply chain disruption, and limited applicability outside New Zealand without recalibration. frontiersin.org
- Winter, R., "DCMA 14-Point Schedule Assessment," January 2011 — describes the protocol's origin in a March 2005 US Under Secretary of Defense memo mandating integrated master schedules on contracts over $20 million; lists the fourteen checks; records that missing-logic faults should not exceed 5% of total tasks; documents definitional changes across three revisions, that third-party software implementations are "not certified by the DCMA or any other body" with errors evident in some, and that the prohibition on negative lags "is not based upon any universal scheduling principle." Legacy source; the protocol remains in widespread use. ronwinterconsulting.com
- Associated General Contractors of America and NCCER, 2025 Workforce Survey Analysis — approximately 1,400 firms; worker shortages the most cited cause of project delays, affecting 45% of respondents. agc.org
- ClickWerxs field operations platform — dispatch board, crew lanes, daily and Gantt views and job assignment with an end-to-end scheduling test suite in place as of 28 June 2026; 441 commits and 87 written threat models at that date. First-party operator data.
Third-party figures are attributed with dates, sample sizes and methodology; one source is vendor-published research and is identified as such, and one is labelled a legacy source. Model accuracy figures are specific to the studies and datasets described and should not be read as applying to any commercial product. Nothing here is legal advice; do not rely on a probabilistic forecast in a contractual or claims context. ClickWerxs sells AI implementation services and earns revenue from those engagements. This is operator opinion.
Kaleb Dickhaut — Founder, ClickWerxs. Kaleb built ClickWerxs from the ground up, from payment processing ISO to the Command Center platform to the AI SEO methodology the blog runs on. He has onboarded hundreds of small businesses onto payment and CRM systems. linkedin.com/in/kaleb-dickhaut
