TL;DR: The most-quoted statistic in AI sales is that 95% of corporate AI initiatives return nothing, from MIT's July 2025 report. It rests on 153 survey responses. That is not a reason to dismiss it, but it is a demonstration of the only skill that matters here: reading a claim well enough to know what it actually supports.
The one-sentence answer: Evaluating an AI vendor is mostly about establishing what happens when the tool is wrong, where your data goes, and what leaving costs.
Every AI vendor pitching a contractor right now is working from the same script. A demo on clean data, a number about how much time you will save, and a case study from a company larger than yours.
None of that is checkable, which is why none of it should decide anything.
What follows is the set of questions that produce answers you can verify, plus the reason most AI purchases in construction fail, which has almost nothing to do with the model.
What actually makes AI projects fail?
Integration with the operation, not the quality of the technology.
MIT's Project NANDA published "The GenAI Divide: State of AI in Business 2025" in July 2025, reporting that despite $30 to $40 billion of enterprise investment in generative AI, 95% of businesses saw no return, and that only around 5% of pilots translated into real operational or financial impact.
Now apply the skill this post is about. That study reviewed more than 300 publicly disclosed initiatives, ran 52 organisational interviews, and collected 153 survey responses from senior leaders at industry conferences. Its methodology has been publicly criticised on those grounds.
So what does it support? A directional claim that most enterprise AI spending has not yet produced measurable return, evidenced substantially by executives who chose to attend AI conferences. That is worth knowing. It is not a precise failure rate, and a vendor quoting "95%" at you to sell a remedy is relying on you not checking.
The useful part is the diagnosis rather than the percentage. Where these projects die is at the point the model meets fragmented data, workflows nobody redesigned, and no owner responsible for it once live. Every question below is aimed at that.
What should I ask before signing?
Six questions, each with a verifiable answer.
What happens when the output is wrong, and who notices? Not whether it is accurate. Every vendor says accurate. You want to know whether a wrong answer is flagged, silently passed through, or caught only when a client complains. A tool that is right most of the time still needs checking, and checking sometimes costs more than doing the work.
What does the tool do with a document it has never seen? Your operation has exceptions. The submittal in a format nobody uses, the invoice with handwriting on it. Ask them to run one of your genuine awkward documents during the demo, not their sample.
Who is accountable for it after go-live? The single most common cause of failure. If the answer is that your project manager will own it alongside their existing job, price that honestly.
Where does our data go, and is any of it used to train your models? Covered below because it deserves its own answer.
What does it cost to stop? Export format, notice period, what happens to your history.
Which contractors of roughly our size are using this? Not enterprise logos. A firm with your crew count and your systems. Ask to speak to one.
How do I tell a real product from a wrapper?
Ask what it does when the AI is unavailable.
This is the most efficient single question I know for this category. Real products have an answer, because the people who built them have watched the model time out and had to decide what the user sees. Thin wrappers have not thought about it, and the demo never covers it.
I can be specific because we build this into our own field operations platform. The dominant failure mode in production AI is not a wrong answer, it is no answer arriving in time. We wrote a threat model specifically for AI timeouts and a dedicated test for it, because a superintendent on a roof with one bar of signal does not care why the request hung. Nobody puts timeout handling in a demo. It is most of the work.
Two further tells. Ask whether AI usage is metered and visible to you, because a vendor who cannot show you consumption cannot help you control cost. And ask what the tool refuses to do. A product with no stated limits has not been used in anger.
What should I ask about data security?
Three specific things, and treat vagueness as the answer.
Construction data is more sensitive than it appears. Drawings, contracts, subcontractor rates and client details usually sit in the same systems, and subcontractor pricing in particular is commercially damaging if it leaks.
Ask whether your data is used to train the vendor's models, and get the answer in the contract rather than in an email. Ask what happens to copies in their development and testing environments, because production data routinely ends up there and is routinely forgotten. Ask who at the vendor can read your project data, and how that access is logged.
Contractors are right to weight this heavily. In Dodge Construction Network's survey of 235 general and trade contractors, conducted with CMiC in September and October 2025, 54% cited data security and privacy risk as a concern and 57% cited reliability and accuracy. Those were the top two.
Will it integrate with what we already run?
Ask who does the work and what happens when the other system changes.
Integration is where budgets go. A tool that reads your accounting system is negotiating with software you do not control, which may be poorly documented and may change without notice.
The questions that matter: does integration exist today or is it on a roadmap, who builds it, what happens when the other vendor changes their interface, and can you get your data out if the integration breaks.
Be particularly careful with anything requiring your crew to change how they work before it produces value. That is the most reliable predictor of an abandoned tool, because behaviour change does not survive a live job.
What should I avoid outright?
Four things.
Anything predicting an outcome with no history behind it. Prediction requires data about your past, and if you have not been recording it, no vendor can conjure it.
Anything replacing a judgment your best estimator makes on instinct. That judgment is built on context the tool does not have, and the failure will be expensive and quiet.
Any contract granting broad rights over your project data as a condition of service. Read the data-licensing clause specifically.
Any vendor who will not name a limitation. Pushback in a sales conversation is the best available proxy for judgment during implementation.
What does a sensible first purchase look like?
One workflow, one measurable number, cancellable.
Pick the most repetitive document task in your office. State what it costs now in hours per week. Buy only that, on terms you can exit, and check the number after sixty days.
This is deliberately unambitious, and it is the opposite of how AI is sold. It is also the shape of the small number of deployments that survive contact with a real project, because the thing being improved was specific enough to measure.
Frequently Asked Questions
Should I run a paid pilot or ask for a free trial?
Paid pilots get better attention and clearer scope, and they let you evaluate the working relationship rather than the software alone. What matters more than the price is whether the pilot has a written success criterion agreed before it starts. A pilot with no defined outcome becomes a subscription by default.
How long should evaluation take?
Two conversations and one document from the vendor. If a vendor cannot produce a written description of what they would deliver for your operation after two calls, that is the answer, and it arrives cheaply.
What if the vendor is a startup that might not survive?
That risk is real and it is not automatically disqualifying, because the alternative is often a mature product that does not fit. Manage it in the contract: data export in a usable format, no dependency on their hosting for your records, and a notice period you can live with. Ask directly what happens to your data if they shut down.
Do I need someone technical on my side for this?
For a single narrow tool bought as software, usually not, if the six questions above are answered in writing. For anything touching multiple systems, a few hours from an independent technical advisor at the contract stage is cheap against the cost of a bad structure you cannot see.
Is it a bad sign if a vendor uses someone else's AI model underneath?
No, and most do. Building a model from scratch would be a worse sign for a construction-specific tool. What matters is whether they will tell you which model, what happens when that provider changes pricing or deprecates a version, and whether your data passes through it under terms you have read.
If you are evaluating a tool now, the fastest thing you can do is send the vendor the six questions above in one email and see how many come back with specifics. That is a better filter than any demo.
See ClickWerxs AI services or get in touch. If you have not yet decided whether AI belongs in your operation at all, what AI actually does for a construction business covers the adoption data first.
Sources
- MIT Project NANDA, "The GenAI Divide: State of AI in Business 2025," July 2025 — $30 to $40 billion enterprise generative AI investment with 95% of businesses reporting no return; approximately 5% of pilots translating into operational or financial impact; methodology comprising a review of more than 300 publicly disclosed initiatives, 52 organisational interviews and 153 survey responses from senior leaders collected at industry conferences. The methodology has been publicly criticised, including by the industry outlet Futuriom. Report summary at aigl.blog
- Dodge Construction Network with CMiC — 235 general and trade contractors surveyed September and October 2025; 57% cite reliability and accuracy concerns, 54% cite data security and privacy risk. Reported via constructiondive.com
- ClickWerxs field operations platform — written threat model for AI timeouts plus a dedicated timeout test; per-tenant AI usage metering; 324 commits and 71 threat models as of 8 June 2026. First-party operator data.
Third-party survey figures cited in this post are attributed to their publishers with sample sizes and dates, and one is noted as methodologically contested. Verify at the source before relying on any of them. ClickWerxs sells AI implementation services and earns revenue from those engagements, including services a reader might evaluate using the questions above. This post reflects our direct experience building AI features into our own platform, not independent third-party research. This is operator opinion and not legal, financial, or professional advice.
Kaleb Dickhaut — Founder, ClickWerxs. Kaleb built ClickWerxs from the ground up, from payment processing ISO to the Command Center platform to the AI SEO methodology the blog runs on. He has onboarded hundreds of small businesses onto payment and CRM systems. linkedin.com/in/kaleb-dickhaut
