It is 2:14am. An autonomous content agent, running unattended, rewrites a dental clinic’s landing page. It swaps “may help reduce staining” for “removes stains in one minute.” The claim is crisper. It is also the exact wording the Advertising Standards Authority has upheld complaints against. No human saw it. The page is live by breakfast, indexed by lunch, and flagged by an automated ad-monitoring sweep before the clinic’s marketing lead has finished her coffee.
Illustrative scenario: The opening scene is a composite used to explain the operational risk. It is not a client case study or claimed result.
Nothing in that scene is science fiction. It is what happens when you buy capability without buying control.
An AI automation agency designs, builds, deploys and maintains AI-driven systems that do marketing and operational work with minimal human intervention. In 2026 the good ones are not selling you software. They are selling you a way to let autonomous systems touch your brand without those systems doing something you would never have approved. That distinction – capability versus control – is the whole subject of this page.
Foundry Works is a London-based AI agent marketing agency. We build governed, agent-run marketing delivery systems. We will be honest throughout about what this work is, what it costs, and where it goes wrong, because the market is full of demos and thin on accountability.
What Does an AI Automation Agency Actually Do?
An AI automation agency replaces repetitive, judgement-light work with systems that run themselves, then puts humans in charge of the decisions that matter. The output is not a chatbot demo. It is a set of production workflows – content production, lead routing, reporting, campaign operations – that run continuously, with defined points where a person approves, edits or blocks what the machine proposes.
The work splits into four honest categories:
- Discovery and mapping. Which processes are safe to automate, which are not, and what happens when each one fails.
- Build. The agents, the integrations into the tools you already use, and the plumbing nobody markets: error handling, retries, monitoring, rate limits, spend caps.
- Governance. The rules that decide what an agent may do alone, what it must submit for approval, and what it may never do without a named human signing off.
- Run and maintain. Because models drift, APIs change and prompts degrade. A handed-over system with nobody maintaining it usually breaks within months.
A useful way to read the market: automation exists on three tiers. Rule-based workflows (if this, then that). AI-powered systems that handle variable inputs and language. And agentic systems, where autonomous agents plan and execute multi-step tasks end to end. The price, the value and the risk all climb as you move up the tiers. So does the need for governance.
The Market Shifted From Selling Tools to Running Delivery
Two years ago, AI agents were a research topic. Today they are a line item in enterprise software budgets. Gartner forecasts that 40% of enterprise applications will include task-specific AI agents by the end of 2026, up from less than 5% in 2025 – one of the fastest enterprise technology shifts on record. The mechanism behind that number is simple: vendors are shipping agents by default, so the decision is no longer whether to deploy agents but which workflows justify the operating overhead.
That is the supply side. The demand side is where it gets interesting. McKinsey’s The State of AI in 2025: Agents, innovation, and transformation (published 5 November 2025, based on 1,993 respondents across 105 countries) found that 88% of organisations now report regular AI use in at least one business function, up from 78% a year earlier – yet only about one-third have begun to scale it, just 7% report AI fully scaled, and only 39% attribute any EBIT impact to it. The gap between “using AI” and “running AI in production” is the gap agencies now exist to close.
Here is the part the sales decks skip. A large share of AI projects fail. Gartner predicts over 40% of agentic AI projects will be cancelled by the end of 2027, citing escalating costs, unclear business value and inadequate risk controls. RAND Corporation, in its 2024 report The Root Causes of Failure for Artificial Intelligence Projects and How They Can Succeed (Ryseff and Narayanan, based on interviews with 65 experienced data scientists and engineers), noted that by some estimates more than 80% of AI projects fail – twice the failure rate of IT projects that do not involve AI. MIT’s The GenAI Divide: State of AI in Business 2025, published in July 2025 by the Media Lab’s Project NANDA and based on 150 leader interviews, a survey of 350 employees and analysis of 300 public deployments, found that 95% of enterprise generative AI pilots delivered no measurable impact on the profit-and-loss statement. And S&P Global Market Intelligence’s Voice of the Enterprise: AI & Machine Learning 2025, surveying over 1,000 enterprises across North America and Europe, found 42% of companies abandoned most of their AI initiatives in 2025, up from 17% a year earlier, with the average organisation scrapping 46% of proofs of concept before production.
The mechanism matters more than the headline. These projects do not fail because the models are weak. They fail because of how the work around the model is scoped, governed and integrated. As Gartner’s senior director analyst Anushree Verma put it, most agentic AI projects are “early-stage experiments or proof of concepts that are mostly driven by hype and are often misapplied.” The failure is managerial, not technical. That is precisely why what you are buying from an agency is not capability. It is the discipline that keeps a working demo alive in production.
What You’re Really Buying: Governance, Not Capability
You are buying the brake, not the engine. Capability is a commodity now – the same large language models are available to everyone, and an autonomous agent that can write a landing page can also publish a non-compliant one. What separates an agency worth paying is its answer to a single question: what stops the agent doing something you would never have approved?
Foundry Works answers that question with three things. They are the same three things we would tell you to demand from any agency, including ours.
The interlock map. The first document we produce for a client is not a prompt. It is an interlock map: a diagram of every point where an agent acts, what it is allowed to do there alone, and where a human must intervene. Most agencies open with a capability demo. We open with a map of where the machine stops.
Tiered risk gates. Not every action carries the same risk, so oversight should not be uniform. We tier every agent action:
- Green – auto-merge, once the agent has earned trust on that task through a track record of clean output. Low blast radius, easily reversible.
- Amber – batched for human approval. The agent prepares the work; a person releases it. Iteration speed stays high; nothing ships unreviewed.
- Red – mandatory human sign-off, no exceptions. Clinical, health, financial and legal claims live here permanently. No amount of established trust promotes a red action to green.
This is a practical Foundry operating model; it is the emerging consensus in AI governance, where oversight is tied to reversibility, blast radius, data sensitivity and domain. Our contribution is refusing to treat it as optional.
The evidence ledger. Every consequential change is logged: what changed, why, the expected outcome, the baseline, the movement, and the keep-or-revert decision. This turns “the AI did something” into an auditable record you can defend to a regulator, a board or a client. Oversight is only governance if you can prove it worked.
The market data quietly agrees with this stance. McKinsey’s 2025 survey found that around half of firms report AI-related incidents, and that high performers are the ones managing risk with human-in-the-loop rules, centralised oversight and executive accountability. Governance is not the tax you pay on AI. It is the thing that makes the AI survive.
Engagement Models: How the Work Is Structured
Most AI automation agencies work in one of four shapes. Pick the one that matches your risk, not the one with the smallest sticker price.
| Model | How it works | Best for | Main risk |
|---|---|---|---|
| Project / build fee | Fixed fee for a defined build | A clear one-time system | System rots with nobody maintaining it |
| Monthly retainer | Recurring fee for ongoing build, monitoring and iteration | Continuous automation across teams | Paying for idle months |
| Per-workflow | Set price per automation built | A menu of standard, repeatable builds | Scope creep between workflows |
| Outcome / value-based | Fee tied to measured impact | Well-defined, high-volume workflows | Requires shared metrics and trust |
The direction of travel is towards outcomes. McKinsey has publicly described tying a growing share of its consulting fees to measurable client outcomes rather than hours worked. That is healthy – but only if the outcome is measured honestly, which is what the evidence ledger is for.
Foundry Works structures engagements around the interlock map first, then a governed build, then a run phase. We would rather show you the map from our last build than promise you a number on a first call.
What Good Looks Like
A good AI automation agency does five things you can check:
- It leads with governance, not a demo. If the first meeting is all capability and no control story, you are looking at demo-ware.
- It produces an interlock map or equivalent before it writes a line of production code. You should be able to see where the machine stops.
- It tiers risk explicitly. Ask how a green action becomes green, and what stays red forever. Vague answers are a red flag.
- It keeps an evidence ledger. What changed, why, and whether it worked – written down, not remembered.
- It owns the run phase. Models drift. If the agency’s plan ends at handover, the system ends shortly after.
The UK Context in One Paragraph
If you operate in the UK, governance is not a nicety. The UK has taken a deliberately pro-innovation, principles-based approach to AI regulation – no single AI statute, with existing regulators such as the FCA, ICO, CMA and MHRA applying their own rules. That lighter touch increases, rather than reduces, the burden on you: there is no compliance checklist to hide behind, and the Advertising Standards Authority now uses AI-powered active monitoring to detect non-compliant ads without waiting for a complaint. For regulated sectors – healthcare, dental, finance – an ungoverned agent is a liability event waiting to happen. We cover this in depth on our AI automation agencies UK page.
Frequently Asked Questions
What is the difference between an AI automation agency and a marketing agency? A traditional marketing agency does the work with people. An AI automation agency builds systems that do the work, then governs the points where humans must stay in control. The deliverable is a running system, not a set of campaigns.
Is an AI workflow automation agency the same as an AI automation marketing agency? Broadly yes – the terms overlap. “Workflow automation” leans towards operational processes (lead routing, reporting, data syncs); “automation marketing” leans towards content and campaign production. Foundry Works does both, under one governance model.
How long before an AI automation system pays back? Reported payback periods vary widely and the honest answer is “it depends on the workflow.” Independent research is clear that a large share of projects never pay back at all, usually because of weak governance and poor data – not weak models. A well-scoped, well-governed pilot should show signal within a quarter.
Do I still need humans if I hire an AI automation agency? Yes, and that is the point. In a governed system, humans own strategy and approval; agents own production. Removing the human from consequential decisions is exactly how the 42% who abandoned AI in 2025 got there.
What should the first deliverable be? An interlock map. If your agency’s first deliverable is a prompt or a demo, ask where the machine stops before you ask what it can do.
Next step
ask us for the interlock map from our last build. It will tell you more about how we work than any pitch deck.*
Continue the buyer’s guide
AI Automation Agency vs In-House: Honest GuideWhat Does an AI Automation Agency Cost? (2026)How to Choose an AI Automation Agency (2026)AI Automation Agencies UK: 2026 Buyer’s GuideStart with the interlock map
We map what an agent may read, write and release before we build the production system around it.
Book a strategy call