By Dean McCoubrey Chief AI Strategist
Forgive the Kennedy-shaped phrasing, but ask not which AI tools the agency uses. Ask what those tools have changed about its thinking.
If the answer is only speed, volume, or cheaper production, you are not looking at AI capability. You are looking at a faster factory.
The short answer: Evaluating whether a marketing agency genuinely understands AI comes down to five questions. Ask whether AI changes the decisions they make, not just the speed. Ask for a campaign where AI shifted strategic direction mid-flight. Ask how they govern brand outputs at scale. Ask what their stack logic is. Ask how they connect AI activity to revenue.
AI is now table stakes in marketing – the gap is no longer who uses it, but how well. According to Adobe’s research, 88% of digital marketers use AI day-to-day. The UK Government’s own AI adoption research finds that 72% of AI-adopting firms use AI in marketing, with brand safety, trust, and ROI cited as the central concerns.
That context changes the evaluation question entirely. When almost every agency will tell you they use AI, asking whether they use it stops being useful. What you actually need to know is whether the AI they use has changed anything that matters: the quality of their thinking, the rigour of their governance, and the commercial results they can prove.
This article gives you a framework for finding out. It is written for CMOs and marketing directors who are actively shortlisting partners and need something more useful than a demo of tools they have already seen.
AI theatre versus genuine AI capability
Some agencies will be fluent in the language of AI before they are mature in the practice of it. That is not a criticism so much as an observation about where the industry currently sits. The tools arrived fast, the vocabulary spread faster, and the pitch decks followed immediately. What takes longer to build is an operating model where AI actually changes how decisions get made.
The distinction worth naming is this: AI theatre is visible AI usage that leaves strategic judgment, governance, and measurement unchanged. The agency uses AI to produce content faster, to generate more options, to compress timelines. The work looks different. The thinking does not.
Genuine AI capability is something else. It means AI is embedded into planning, optimisation, measurement, and decision loops. It means the agency can show you a moment where an AI-driven insight changed a recommendation – not just sped up the delivery of one.
| AI Theatre | Genuine AI Capability | |
|---|---|---|
| Where AI sits | Production layer only | Across strategy, planning, delivery, and measurement |
| What changes | Speed and volume | Decisions, recommendations, and outcomes |
| How it shows up in a pitch | Tool demonstrations, workflow diagrams | Case examples with changed direction and proven results |
| Governance | Absent or informal | Documented, with clear brand safety and approval frameworks |
| Measurement | Efficiency metrics (time saved, content volume) | Commercial metrics (revenue lift, pipeline, CAC, LTV) |
| Who owns it | One AI specialist or department | Embedded across the team |
The Real Story Group’s framework for evaluating AI capability draws a useful distinction between three types of AI: generative AI (writing, editing, creative production), insights AI (audience analysis, attribution, predictive modelling), and decisioning AI (next best action, dynamic optimisation, journey testing). Most agencies have invested in the first. Fewer have built the second. The third is where genuine competitive advantage lives, and it is the hardest to fake in a pitch.
The question is not which category an agency operates in. It is whether they can show you evidence that AI has moved their work from the first category toward the third.
Five questions to ask any agency in a pitch
These questions are not designed to catch agencies out. They are designed to surface the difference between an agency that has thought carefully about how AI changes their work, and one that has learned to talk about it fluently. Both will give you an answer. The quality of the answer is the data.
1. How does AI change the decisions you make, not just the speed at which you make them?
This is the foundational question. It separates capability from convenience.
- Strong answer: The agency describes a specific moment where AI-generated insight led to a different strategic recommendation, a changed channel mix, or a revised creative direction. They can name the decision and explain the outcome.
- Red flag: The answer focuses on workflow efficiency, turnaround times, or content volume. Speed is not strategy.
2. Can you show me a campaign where AI-driven insight changed the creative or media direction mid-flight?
This tests whether AI is embedded in active decision-making or only used in pre-production.
- Strong answer: A concrete example with a before and after. What did the AI surface? What changed? What was the commercial result? As one evaluation framework puts it: “Show an example where AI changed your go-to-market or channel strategy, not just sped up execution.”
- Red flag: Vague references to “ongoing optimisation” with no specific example. If they cannot name the campaign, they probably cannot name the outcome.
3. How do you maintain brand quality and governance when AI is generating outputs at scale?
UK Government AI adoption research identifies brand safety, trust, and measurement as the central concerns in AI adoption. A credible agency should have thought about this carefully.
- Strong answer: A documented governance framework covering review layers, approval processes, and risk ownership. They know who signs off on AI-generated brand outputs and why.
- Red flag: Governance is treated as an afterthought, or is described as “we always have a human check it” without any structure behind that claim.
4. What does your AI stack look like, and which tools are you tool-agnostic on versus committed to?
This is not a question about which tools they use. It is a question about whether they have a principled view on why.
- Strong answer: The agency can explain the logic behind their tool choices, where they have made deliberate commitments and where they remain deliberately agnostic. They are not name-dropping platforms to signal credibility.
- Red flag: The answer is a list of tool names with no explanation of how they connect, why they were chosen, or what problem each one solves.
5. How do you connect AI-assisted marketing activity to commercial outcomes – pipeline, revenue, not just leads?
This is the question that most directly separates genuine AI capability from efficiency theatre. As the evaluation guidance from Whitehat notes, the test is whether an agency can show “revenue lift, lead quality, CAC, LTV, or efficiency gains – not just vanity metrics.”
- Strong answer: The agency can describe how AI-assisted activity connects to pipeline, revenue, customer acquisition cost, or retention. They measure commercial impact, not just activity.
- Red flag: All measurement is expressed in efficiency terms: time saved, content produced, impressions generated. These are real gains, but they are not commercial proof.
A note on answers that feel rehearsed: Some agencies will have prepared for these questions. That is fine. What you are listening for is not spontaneity – it is specificity. A rehearsed answer with a real example behind it is still a good answer. A fluent answer with no example behind it is not.
What genuine AI integration looks like in practice
Beyond what an agency says in a pitch, there are observable signals in how they work. Genuine AI integration tends to show up across the full workflow, not concentrated in a single team or a single stage of delivery.
Observable signals worth looking for
- AI in insight generation, not just production. The agency uses AI to surface audience signals, identify anomalies, and model attribution, not only to generate copy or creative variants.
- First-party data at the centre. AI capability without data infrastructure is limited. An agency that has genuinely embedded AI will have a clear view of how client first-party data feeds into their models and tools.
- Human editorial control over AI outputs. The most credible agencies describe a deliberate tension between AI-generated material and human judgment. They have not automated away the thinking; they have changed where the thinking happens.
- Measurement that includes AI-native channels. As AI-powered search and answer engines reshape how buyers discover brands, the agencies ahead of this shift are already measuring prompt visibility and answer engine inclusion alongside traditional metrics. If an agency has no view on this, they are measuring the last decade, not the next one.
- Cross-team adoption. AI capability that lives in one person or one department is fragile. When it is embedded, it shows up in how strategists brief, how planners model, and how account teams report.
The real test is coherence. An agency with genuine AI capability can describe how AI connects strategy to execution to measurement as a single system. An agency performing AI capability tends to describe each application separately, because there is no system behind them.
The red flags that reveal AI theatre
None of these signals is definitive in isolation. Several together are worth taking seriously.
- The answer to “how do you use AI?” is a list of tools. Naming platforms is not evidence of capability. Ask what those tools changed about the work. If the answer circles back to tools, you have your answer.
- They cannot name a decision AI changed. Fluent answers with no example behind them are a reliable signal. If AI has genuinely shifted how an agency thinks, they should be able to describe a specific instance without pausing to construct one.
- One person owns the AI capability. A single specialist is a dependency, not an operating model. When AI is genuinely embedded, it shows up in how strategists brief, how planners model, and how account teams report. Not in one person’s job title.
- The pitch is a demonstration, not a dialogue. Agencies with genuine capability ask questions. They want to understand your data, your governance requirements, your measurement framework. Agencies performing capability tend to show you things and wait for applause.
- They measure AI success by efficiency alone. Time saved and content volume are real gains. But if the agency cannot connect AI activity to pipeline, revenue, or commercial lift, they are optimising the wrong thing. Spencer Stuart’s CMO research identifies ROI and commercial proof as the central anxieties driving senior marketing decisions in 2026. An agency that cannot speak to those is not yet speaking the right language.
A practical AI agency evaluation scorecard
This scorecard is designed to be used during or immediately after a pitch process. Score each dimension 1 to 3. A score below 20 out of 30 warrants deeper scrutiny before any shortlist decision is made.
The scoring is deliberately simple. In a pitch process, a useful tool is one people can actually use before the memory of the meeting goes soft.
| Dimension | 1 – Weak | 2 – Developing | 3 – Strong |
|---|---|---|---|
| Decision impact | AI used only in production | AI informs some recommendations | AI demonstrably changes strategic decisions |
| Case proof | No specific examples | Examples exist but outcomes unclear | Named campaigns with measurable commercial results |
| Governance | No framework described | Informal review process | Documented governance with clear accountability |
| Stack clarity | Tool list with no rationale | Some explanation of tool choices | Principled, agnostic-where-appropriate stack with clear logic |
| Data integration | No mention of first-party data | Some data integration described | First-party data central to AI capability |
| Measurement quality | Efficiency metrics only | Mix of efficiency and commercial | Commercial outcomes: revenue, CAC, LTV, pipeline |
| Cross-team adoption | AI sits with one specialist | Some team-wide use described | AI embedded across strategy, planning, and delivery |
| Strategic fluency | Tool-led answers | Process-led answers | Outcome-led answers with commercial framing |
| AI-native search | No awareness of answer engines | Awareness but no measurement | Active measurement of AI-native discovery and visibility |
| Commercial accountability | No link to revenue | Indirect commercial references | Direct connection between AI activity and business outcomes |
Score interpretation:
- 25-30: Strong capability. AI appears genuinely embedded in operating model and commercial thinking.
- 20-24: Developing capability. Probe the gaps before committing, particularly on governance and commercial proof.
- Below 20: Proceed with caution. The agency may be fluent in AI language without the operating model to back it up.
The scorecard is a decision-support tool, not a replacement for judgment. A strong score in a pitch does not guarantee strong delivery. But a weak score in a pitch is usually a reliable signal.
For a broader view of how AI-first agencies are positioning themselves in the UK market, the best AI-first marketing agencies in the UK for 2026 guide provides useful context on how genuine capability is being built and described.
The honest question to ask yourself
The right question at the end of any agency pitch is not “do they understand AI?” Almost everyone will pass that test in 2026. The question is whether AI makes their thinking better, or just their production faster.
The former is rare. The former is what changes commercial outcomes.
An agency that uses AI to compress timelines and increase output volume is offering you efficiency. That has value. But an agency that uses AI to improve the quality of its decisions, to surface insights that change strategic direction, and to connect marketing activity to commercial results is offering you something different. The distinction is worth paying attention to.
For a broader view of how AI-first agencies are positioning themselves in the UK market, see our guide to the best AI-first marketing agencies in the UK for 2026.
Frequently asked questions
What questions should I ask an agency about AI?
Ask five questions: how AI changes the decisions they make (not just the speed); whether they can show a campaign where AI shifted creative or media direction mid-flight; how they maintain brand governance when AI generates outputs at scale; what their AI stack looks like and where they are tool-agnostic; and how they connect AI-assisted activity to commercial outcomes such as pipeline, revenue, or customer acquisition cost.
How do I spot AI theatre in a marketing agency?
AI theatre is visible AI usage that leaves strategic judgment, governance, and measurement unchanged. The clearest signals: the agency leads with tool names rather than outcomes; their AI capability sits with one specialist rather than being embedded across the team; there is no documented governance framework for AI-generated brand outputs; and all measurement is expressed in efficiency terms rather than commercial results.
What does a genuinely AI-first agency look like?
A genuinely AI-first agency uses AI across the full workflow, from insight generation and planning through to delivery, optimisation, and measurement. It can show specific examples where AI changed a strategic recommendation, not just sped up production. It has a documented governance framework. It measures commercial outcomes, not just efficiency gains. And it has a view on AI-native search and answer engine visibility, not only traditional channel metrics.
How do I evaluate AI capability without being misled by a good pitch?
Use the five-question framework and the 10-dimension scorecard in this article. A score below 20 out of 30 warrants deeper scrutiny. Pay particular attention to specificity: an agency with genuine capability will give you named examples with outcomes. An agency performing capability will give you fluent answers with no example behind them.
Why does AI governance matter when evaluating a marketing agency?
Because AI-generated outputs at scale carry real brand risk. If an agency cannot describe who reviews AI-generated material, how brand standards are maintained, and who owns accountability when something goes wrong, that is not a minor gap. It is a governance failure waiting to happen. The UK Government’s AI adoption research identifies brand safety and trust as central concerns in AI adoption for exactly this reason.

