Your messy data is the actual AI project

Most small business owners think the AI build is the hard part. In practice it's the four weeks of work you do on your own spreadsheets before the AI build can begin, and almost nobody scopes for that honestly.

A flooring contractor in Oakville wanted us to build him a tool that would tell him, on any given Monday, which of his open jobs were drifting over budget. He had been losing about $40,000 a year on jobs that went sideways without anyone noticing until the final invoice. He wanted a demo by the end of the quarter.

I asked him where his job data lived. He pulled up his screen and showed me the system, which turned out to be three things. A QuickBooks file for invoicing. A shared Google Sheet where his project manager tracked materials. And a binder, sitting on a shelf in the office, where the original quotes were filed by hand because his older estimator refused to use the spreadsheet.

I sat with him for an hour going through twenty random jobs from the previous year. We could not, in any reliable way, tie a single quote to a single invoice to a single materials line. The job names didn’t match between systems. The dates were off by days. One job had two different totals in QuickBooks because someone had voided and reissued an invoice without flagging it. About a third of the binder quotes had no corresponding line in either digital system.

He looked at me and said, “So you can’t build the thing.”

I said I could absolutely build the thing. It would just be lying to him.

This is the conversation I now have, in some form, on roughly two out of every three engagements. The owner wants the AI build. The AI build cannot give honest answers until somebody, usually me and one of his staff, spends three to five weeks fixing the data the build needs to read from. Almost no proposal in the market scopes that honestly. Almost every project that fails at the SMB level fails here, in the gap between what the data is and what the owner thinks it is.

What “messy data” actually means

The phrase is so vague it has stopped meaning anything. Let me be concrete about what shows up in a typical small business.

The same customer is in your CRM under three slightly different names. “Acme Roofing,” “ACME Roofing Ltd,” and “Acme Roofing Inc.” are three records, not one. Your sales totals for that customer are wrong by a factor of three until somebody merges them.

Your project manager records job status as “in progress,” “underway,” “active,” and “ongoing” depending on his mood. To him they all mean the same thing. To a model trying to count active jobs, they look like four different states.

Your phone calls are logged in one system, your emails in another, and your text messages with clients are not logged anywhere because they happen on your project manager’s personal phone. The customer relationship the AI would need to summarize lives across four surfaces, and only two of them are reachable.

Your historical financials look clean until you discover that 2024 was reclassified mid-year when you switched accountants, and the comparison to 2023 is no longer apples to apples without a translation table nobody wrote down.

None of this is unusual. All of it is invisible until the moment you ask a tool to reason across it.

Why nobody scopes for it

The honest reason is that data cleanup is unsellable. An owner who has just heard about a tool that can predict job overruns does not want to spend $14,000 on three weeks of someone fixing his spreadsheets. He wants the tool. The firm pitching the project knows this, so the proposal talks about the tool. The cleanup either gets buried in a vague line item or assumed away with phrases like “subject to data availability.”

Then the project starts. Week one is discovery, and the firm finds out what the data actually looks like. Week three is the owner deciding the firm doesn’t know what they’re doing, because the proposal said the tool would be running by week four and now there’s a $9,000 invoice for “data normalization” and no tool.

I have seen this kill four projects in the last eighteen months. The firms doing the work were not incompetent. They were just optimistic at the proposal stage in a way the owner could not forgive once the bill arrived.

The version that works is the one where you tell the owner, before signing, that the first three to five weeks are going to be unglamorous and produce nothing demonstrable, and that the AI build only starts once that work is done. About half of owners walk away. The half that stays is the half that ends up with a working system.

What the cleanup actually looks like

For the flooring contractor, the first month of the engagement looked like this.

We picked one specific question to optimize the data for: which open jobs are drifting over budget on Monday morning. Everything else got deferred. This was important because his data was messy in twenty directions and we only had budget to fix it in one.

We sat with his project manager and his bookkeeper for two afternoons and built a single job-ID convention. New jobs got it cleanly. Old jobs got a back-mapping spreadsheet, built by hand, that took eleven hours over two weeks. His estimator’s binder got photographed and the quotes got entered into the same system as everything else, with the estimator sitting next to us so he could correct the data entry as we went.

We standardized the status field in his Google Sheet to five options on a dropdown and went back through the previous twelve months reclassifying the entries. About 8 percent of the historical entries couldn’t be reconstructed and got flagged as “unknown” instead of guessed at. That mattered because the model would otherwise have learned from invented data.

By the end of the fourth week I had something I could legitimately point an AI workflow at. The build that followed took nine days. He uses it every Monday morning.

The work that mattered most was not the build. It was the four weeks of cleanup that almost wasn’t in the proposal.

What I would tell an owner reading this

If you are considering an AI project right now, the most useful thing you can do before talking to any firm is open up the systems where the data lives. Not in an organized way. Pick three random recent transactions or jobs and trace them all the way through. Quote, invoice, materials, customer record, communications log.

If you can do that cleanly in under ten minutes for all three, your data is in better shape than most. The project can probably start where the proposal says it starts.

If you cannot, the project’s first month is going to be cleanup, whether anyone tells you that in the proposal or not. The firms who name it up front are the ones to talk to. The firms who don’t are not necessarily dishonest. They are usually optimistic, which is almost as expensive.

The reason I keep writing about this is that the AI conversation at the small business level keeps getting framed around the model. Which model. Which vendor. Which feature. In my experience the model is rarely the binding constraint. The binding constraint is whether the business’s own records are coherent enough for any model to reason over them.

That framing does not sell tools. But it is the actual shape of the work in the businesses I have spent the last two years inside of. The owners who get good outcomes are the ones who let themselves be told, before signing, that the first month is going to look like nothing, and that the something only comes after.


From argument to implementation

Apply the idea to one real workflow.

The Nano-Pilot ranks a small set of opportunities and makes the assumptions visible. If the workflow is already scoped, the Implementation Sprint is the build path.

Describe the workflow behind the argument.

Glen replies in writing with a fit assessment within two business days.

Send a written intake