A founder signs with an “AI automation agency” after a polished demo. Six weeks in, he realises the “team” is three freelancers coordinating over a shared login, and the “proprietary platform” is a ChatGPT subscription and some brittle Zapier zaps. Nothing is documented. When one freelancer goes quiet, the system he built goes dark with him. There was never an interlock map, never an evidence ledger, never a clear answer to “what happens when this breaks.” The demo was real. The agency, in any meaningful sense, was not.
Illustrative scenario: The opening scene is a composite used to explain the operational risk. It is not a client case study or claimed result.
Choosing an AI automation agency is mostly an exercise in telling the difference between a good demo and a system that survives contact with production. Here is how.
Short answer: Choose the agency that leads with governance, not capability. The best signal is what they produce first: a good AI automation agency maps where the machine stops before it shows you what the machine can do. Everything else – team, references, pricing – matters, but the governance story is the one that predicts whether you’ll be in the 42% who abandon AI or the minority who scale it.
The Evaluation Criteria That Matter
Score candidates on six things, in this order:
- Governance story. Can they explain, without prompting, how they decide what an agent does alone, what it submits for approval, and what it may never do? This is the whole ballgame. Projects fail on governance, not models.
- Ownership of the run phase. Do they maintain the system after launch, or does their plan end at handover? Models drift; a handed-over system with nobody maintaining it usually breaks within months.
- Evidence of real builds. Not demos – production systems, with an honest account of what broke and how they handled it. Ask what happens when the API is down.
- Team reality. Who actually does the work? Named people, or an anonymous “team”? A shared login is not a team.
- Code and knowledge ownership. Will you own what they build, with documentation, or are you locked into a proprietary framework forever?
- Domain fit. Especially for regulated sectors, do they understand your compliance context, or will they cheerfully automate you into an ASA breach? (See our UK regulated-sector guide.)
Questions to Ask in the First Call
Bring these. The quality of the answers tells you more than any case study.
- “What’s the first thing you’ll deliver?” The answer you want is a map of where humans stay in control – an interlock map or equivalent – not a prompt or a demo.
- “How does an action become fully automated?” A good agency describes a trust-earning process: a task starts under human approval and is promoted to auto-merge only after a clean track record. If everything is automated from day one, that’s a red flag.
- “What stays under mandatory human sign-off, always?” For anyone touching health, clinical, financial or legal claims, the answer must include those, permanently. “We can automate everything” is the wrong answer.
- “Show me your evidence ledger.” How do they record what changed, why, the expected outcome, the baseline, the movement and the keep-or-revert decision? If there’s no record, there’s no accountability.
- “What happens when it breaks at 2am?” Monitoring, alerting, rollback, named ownership. Silence here means the system fails quietly in production until trust – and customers – are gone.
- “Who owns the code and the documentation?” You should. Get it in writing.
- “Who, by name, will do the work?” And can you speak to them, not just the salesperson?
The Red Flags
- Demo-ware. Impressive in the room, undefined in production. Ask to see something running live with real volume.
- No governance story. If oversight, risk tiers and human approval don’t come up unprompted, they’re not doing the work.
- No evidence ledger. “Trust us, it’s working” is not measurement. Without a record, outcome claims are unverifiable.
- Vague accountability. Nameless teams, passive language (“mistakes were addressed”), no clear owner for failures.
- Everything automated, immediately. Full autonomy from day one means nobody thought about blast radius. This is how an ungoverned agent burns a month’s ad budget in a weekend.
- Proprietary lock-in. Closed frameworks, no code ownership, no handover plan.
- A quote with no governance line. Covered on our pricing page – if oversight isn’t costed, it isn’t happening.
How to Run a Pilot
A pilot is how you test governance cheaply before you commit. Structure it like this:
- Pick one workflow with a clear success condition. Not “improve marketing” – something like “draft and route inbound enquiries, with human approval before anything sends.” One pain point, executed well, is how the successful minority start.
- Define the interlock map first. Before any build, agree where the agent acts, what’s green, what’s amber and what’s red. If the agency resists mapping this, stop.
- Set a hard model-spend cap. Agree the number and where it’s enforced. This is the single control that prevents runaway API bills.
- Run for a fixed window with an evidence ledger. Four to eight weeks. Log every change, its expected outcome, its baseline and its actual movement. At the end you have data, not vibes.
- Review against the success condition you set at the start. Did it work, on the metric you agreed before you began? Keep, revert or iterate – and make the agency show their working.
A pilot that starts by defining where the machine stops is worth ten that start with a demo.
What a Good Proposal Contains
- A clear scope with written deliverables and a success condition.
- An interlock map, or a commitment to produce one as the first deliverable.
- Explicit risk tiers: what’s green, amber and red, and how green is earned.
- A governance and maintenance line in the pricing, not just build cost.
- A model/API spend estimate with a hard cap and how overruns are handled.
- Code and documentation ownership, stated plainly, in your favour.
- Named people and an honest reference you can call.
If a proposal has all seven, you’re dealing with a real agency. If it has a shiny demo and none of them, you’re dealing with three freelancers and a subscription.
Frequently Asked Questions
What’s the single best question to ask an AI automation agency? “What’s the first thing you’ll deliver?” If the answer is a map of where humans stay in control rather than a demo or a prompt, you’re talking to a serious agency.
How do I know if an agency’s governance is real or just talk? Ask to see an evidence ledger and an interlock map from a previous build. Real governance produces artefacts. Talk produces adjectives.
How long should a pilot last? Four to eight weeks on a single, well-scoped workflow with a success condition agreed before you start and an evidence ledger running throughout.
What’s the most common way buyers get burned? Buying the demo. A polished demo tells you the tool works in ideal conditions; it tells you nothing about what happens in production when data is messy, volume spikes and the API goes down at 2am.
Next step
before your next agency call, ask them for the interlock map from their last build. If they can show you one, they’re worth your time. If they can’t, you’ve saved yourself a quarter.*
Continue the buyer’s guide
What an AI Automation Agency Actually Does (2026)AI Automation Agency vs In-House: Honest GuideWhat Does an AI Automation Agency Cost? (2026)AI Automation Agencies UK: 2026 Buyer’s GuideStart with the interlock map
We map what an agent may read, write and release before we build the production system around it.
Book a strategy call