We tried AI already and it didn't work

The most common objection I hear from small business owners is that they already tried AI once and it failed. Almost every time, the failure was a scoping problem, not an AI problem.

An accounting firm in Oakville told me last month they were “done with AI.”

About 18 months ago, one of the senior associates had gotten excited about ChatGPT and convinced the partners to spend $4,000 on a small pilot. The AI would read incoming client emails, classify them by topic, and draft a first-pass response a staff member could review and send. Reasonable on paper. The associate spent three weekends building it. It half-worked for about a month, then started misclassifying things in ways the team couldn’t predict, then got switched off when the associate took a vacation and nobody else knew how to fix it.

The managing partner’s takeaway was that AI wasn’t ready for a firm their size. He was telling me this story to explain why he wasn’t interested in talking about anything new.

I asked him whether the associate had ever been given time to talk to a real customer of the firm about why the project mattered. He paused. He hadn’t. The whole thing had been built by one excited person in evenings and weekends, with no charter, no measurement plan, and no clear answer for what success looked like.

That wasn’t an AI failure. That was a project management failure that happened to involve AI.

I have now heard a version of this story from roughly twenty small business owners. The details vary. The pattern doesn’t.

What these failed pilots all had in common

Five things show up almost every time.

The pilot was scoped by whoever happened to be enthusiastic, not by whoever owned the workflow. Excited employees see possibilities. Workflow owners see failure modes. When the excited person scopes alone, the failure modes get discovered live, in production, by accident.

There was no baseline measurement. Nobody wrote down how long the manual version of the task took, how many errors it produced, or what the cost per unit of work was. Without that, you cannot tell whether the AI version is better. So six weeks in, when the AI is producing weird output, there is no way to compare it to anything, and the instinct is to assume it must be worse because the human version was at least familiar.

The AI was wired into a real workflow before it was trusted. The accounting firm’s tool was drafting responses to actual clients within two weeks of the first prototype. There was no period where it ran alongside the human and got compared output against output. So when it started getting things wrong, it was already in the production path, and pulling it out felt like a public admission of failure.

The thing the AI was doing was not the bottleneck. The accounting firm’s real bottleneck was partner-level review of complex returns and the year-end crunch around tax season. The pilot was working on something nobody at the firm would have listed in their top five operational problems. Even when it worked, nobody felt the gain in their week.

Nobody owned it after the build. When the associate went on vacation, the system stopped having a maintainer. No runbook, no successor, no person whose job description included keeping the AI working. It died because no one had the responsibility to keep it alive.

Why the failure gets attributed to AI

The person running the project was emotionally invested in the technology. When it broke, the framing they brought back to the partners was “the AI couldn’t do it” rather than “I scoped this wrong.” Admitting the scoping error costs you internal credibility on the exact topic you advocated for, so most people don’t.

The failure gets coded as a technology problem. The technology gets dismissed. The firm becomes allergic to the category for a year or two, and the actual lesson never gets written down.

In 2024 and 2025, ChatGPT made it feel like a junior person with curiosity could build something real over a weekend. Sometimes they could. The prototype is the easy part. The hard part is the year that comes after, and nobody told the junior person that the year after the prototype was the actual project. They built the demo. They got applause. The demo didn’t survive contact with operations. Now the firm thinks AI doesn’t work.

What changes if you try again with eyes open

The accounting firm and I ended up talking for almost two hours. By the end, the partner had a different picture of what a second attempt could look like.

Find the workflow that is actually expensive. Not the one the most excited person on staff wants to automate. The one where the math is obvious. For their firm, it turned out to be the bookkeeping cleanup work the team does for new clients in their first month. About 40 hours of senior associate time per client, billed at $90 an hour, and roughly 70 percent of it was pattern-matching transactions against likely chart-of-accounts categories. That’s around $2,500 per new client in pattern-matching labor that an AI is genuinely good at.

That’s a real workflow. It has a number. It has a person whose week it eats. It has a measurable before-and-after.

Run it in shadow mode first. For the first month, have the AI categorize alongside the human and compare the output. Don’t change any client-facing process. Don’t bill anything differently. Just see whether the AI’s categorizations match the human’s, and where they differ, figure out which one was right. That gives you a real error rate before anything is at stake.

Name the owner up front. Before the project starts, write down who owns this on day 180. Not who built it. Who keeps it running. If that person can’t be named, the project isn’t ready to start. It will die the same way the email classifier did.

Plan for the boring year. Most of the work, once a project is in production, is uninteresting maintenance. The data source changes format. A new client account type appears. The AI model gets a new version. Each of those is a 20-minute fix if someone is watching, or a six-week disaster if nobody is. Budget for the watching part.

Where this is uncomfortable for a firm like mine

Part of the reason so many small businesses have a failed pilot in their past is that firms like mine, four years ago, would happily take a check from anyone who walked in the door and build whatever was asked for. The accountant who wanted to classify emails would have gotten exactly that. The shoe store owner who wanted a chatbot would have gotten the chatbot. The implementation firms made their fee whether the project mattered to the business or not.

The owners now allergic to AI are not wrong to be suspicious. They paid real money for projects that didn’t pay them back, and the consultants who took that money moved on to the next client.

The version of GRC I’m trying to build refuses the engagement when the workflow isn’t right. That isn’t a marketing pitch. It is the only way I’ve found to keep my pipeline clean of projects that won’t pay for themselves. Saying no to a $30,000 wrongly scoped engagement is hard. It is harder to be the firm that takes the check and watches the project quietly die six months later, the way the accountant’s first project did.

What to do if this is you

If you have a failed pilot in your past, the most useful thing to do before a second attempt is write down what actually went wrong the first time. Not the version of the story that protects whoever ran it. The real version.

Was the workflow the actual bottleneck, or just the thing someone wanted to build? Did anyone measure the before-state in numbers? Did the AI go straight into production, or was there a shadow period? Who owned it on day 90? When it broke, was there a runbook, or did everyone shrug and switch it off?

Those are project management questions, not technical ones. The answers will tell you what to do differently more reliably than any vendor demo will.

The Oakville firm is starting a second project next month. Different workflow. Different owner. Different plan. I think it’ll work. The AI is not the variable that changed.


From argument to implementation

Apply the idea to one real workflow.

The Nano-Pilot ranks a small set of opportunities and makes the assumptions visible. If the workflow is already scoped, the Implementation Sprint is the build path.

Describe the workflow behind the argument.

Glen replies in writing with a fit assessment within two business days.

Send a written intake