By Dean McCoubrey Chief AI Strategist
Most companies evaluating marketing agencies for AI capability are asking the wrong question. They want to know which tools an agency uses. The more useful question is whether AI has changed the quality of how they think.
The difference is not subtle. It shows up in how an agency briefs, how it iterates, and whether it can explain its reasoning at every stage. An agency that genuinely understands AI can make its thinking inspectable. One that doesn’t will show you a tool stack.
The real differentiator is not which AI tools an agency uses. It is whether AI improves the quality of their thinking — and ultimately, the quality of your decisions.
What this guide covers:
- The real concern driving this search, and why most agency evaluations miss it
- Four specific questions to ask any agency claiming AI capability, and what considered answers actually look like
- How to distinguish between agencies built around better decisions and agencies that have simply automated faster output
- What genuine AI governance looks like, and why its absence is a commercial risk
- A practical checklist to take into any agency pitch conversation
The question behind the question
When a CMO or procurement lead searches for how to evaluate an agency’s AI capability, the surface question is about competence. The real question is about something harder to verify: whether AI has changed the quality of thinking, or merely the speed of production.
That distinction matters more than most evaluations acknowledge. According to Jasper’s 2026 State of AI in Marketing report, 91% of marketers are now actively using AI. Salesforce’s State of Marketing 2026 puts the figure at 87% using generative AI in at least one workflow. At that level of adoption, claiming to “use AI” tells a buyer almost nothing. It is roughly as meaningful as claiming to use email.
The more revealing number sits underneath the headline. One 2026 marketing data report found that while 80% of marketers feel pressure to adopt AI, only 6% have successfully embedded it into daily workflows. That gap between adoption and integration is where the risk lives. An agency operating in that gap is using AI as a production shortcut, not a structural advantage.
“The gap between AI adoption and deep integration is one of the clearest indicators that many agencies understand the need for AI but haven’t built robust operating models around it.” — Digital Applied, AI Marketing Statistics 2026
The implicit fear driving most searches on this topic is precise: am I about to pay a premium for something that is just a generic language model with a logo on it? That fear is commercially justified. But the more useful frame is not whether an agency uses AI. It is whether AI has made their thinking better. Those are entirely different evaluations, and only one of them shows up in commercial results.
This guide is structured around that distinction. It applies equally to agencies, consultancies, internal teams, and vendors. The question of whether AI genuinely improves thinking is not unique to agency selection. It is the central question of AI capability, full stop.
Four questions worth asking
These questions are designed to be used in live conversations. They are not trick questions. A genuinely capable agency will welcome them. What they surface is whether AI has changed the quality of thinking, or merely the pace of delivery.
“Can you show me the brief that preceded this deliverable?”
Agencies that genuinely integrate AI start every workflow with structured thinking, not a prompt. A real AI-integrated brief documents the commercial objective, the audience context, the constraints, and the hypotheses being tested. That brief is what gets fed into the model. Without it, the output is generic by design.
A considered answer looks like this: a documented brief with client-specific context, the reasoning behind key decisions, and a visible connection between the brief and the deliverable. The brief and the output should be legibly related.
“We use AI to work faster” with no process to show for it is worth paying attention to. Speed is not a strategy.
“What did you reject from the AI output, and why?”
Any agency can run a prompt. The skill is in evaluating what comes back. Agencies that understand AI can walk through specific examples of output they discarded, rewrote, or challenged. That editorial judgement is the actual product. Without it, AI is a volume machine.
A considered answer: “The model gave us three headline directions. We rejected two because they didn’t reflect the brand’s commercial positioning. Here is the version we rewrote and the reasoning behind it.”
Enthusiasm about speed and volume with no mention of quality control or rejection logic may tell you something about where the thinking is happening.
“How does your AI output get validated against real commercial data?”
AI generates plausible content. Plausible and commercially accurate are not the same thing. An agency that does not connect AI outputs to real performance signals — search intent, pipeline data, customer behaviour, conversion metrics — is producing noise at scale. As one 2026 buyer framework puts it: “You’re not buying AI; you’re buying pipeline, revenue, and cycle-time lift powered by AI.”
A considered answer describes a feedback loop. Output goes live, performance data comes back, the next iteration is informed by what actually worked. They can name the specific metrics they optimised against.
“We measure everything” without being able to name a single metric that changed their approach is a different kind of answer.
“What would you not use AI for in my project?”
This is the most revealing of the four. Agencies that genuinely understand AI also understand its limits. If an agency claims AI can do everything, they either misunderstand the technology or they are not being straight with you.
Responsible AI use requires knowing where human judgement is irreplaceable: brand voice decisions, strategic framing, sensitive stakeholder communications, anything where originality is the commercial point. Documented frameworks for where AI is and is not used — including emerging practices around Generative Engine Optimisation and agentic workflows — are a reasonable thing to ask about.
A considered answer is specific: “We do not use AI for first-draft strategy. We do not use it for client-facing communications where tone is non-negotiable. We do not use it where the brief requires original thinking that AI cannot replicate.”
No answer, or a vague reassurance that humans are always in the loop, is worth noting.
The pattern across all four is consistent. Inspectable thinking at every stage: the brief, the iteration, the rejection, the validation. The distinction between agencies that can show this and agencies that cannot is not a technical one. It is a thinking one.
Faster output vs better decisions
Most agencies claiming AI capability are not misrepresenting themselves. They are using AI. The question is what they are using it for, and whether it has changed anything that matters commercially.
“AI is making marketing easier to produce. But harder to win.”
That tension is the commercial reality for any buyer in 2026. As AI lowers the cost and effort of content production, the market fills with plausible, competent, indistinguishable output. In a market flooded with AI-generated sameness, originality becomes a commercial advantage. Strategic clarity, genuine judgement, and the ability to produce work that could only have come from a specific perspective, for a specific brand, in a specific context — these are precisely the things AI cannot replicate at scale.
The agencies that understand this are using AI to think better. The ones that do not are using it to produce more.
Two operating models in practice
The distinction is not about which tools an agency uses. It is about what AI is connected to inside their process.
| Faster output model | Better decisions model | |
|---|---|---|
| AI role | Accelerates content production | Informs planning, iteration, and validation |
| Brief quality | Generic or templated | Client-specific, hypothesis-driven |
| Iteration logic | Volume and speed | Selective, editorially controlled |
| Commercial validation | Minimal or post-hoc | Built into the feedback loop |
| What the buyer gets | More content, faster | Better decisions, measurable outcomes |
The practical test is straightforward. Ask to see a process diagram. Agencies built around better decisions can show one, with clear decision points, handoffs, and validation stages. Their thinking is inspectable. Agencies built around faster output will describe a tool stack.
Neither model is inherently dishonest. But the distinction is worth paying attention to, because both tend to claim the same thing.
A question most buyers forget to ask
There is a fifth question worth raising in any agency conversation, one that sits outside the evaluation framework above:
“Will working with you make our team more capable over time, or more dependent on you?”
Strong AI partners improve your organisation’s judgement and AI fluency alongside delivering commercial results. They document their process, share their reasoning, and treat knowledge transfer as part of the engagement. Weaker ones keep the process opaque. The point of AI is not just to make things faster. It is to expand what your team can do, and the quality of decisions they can make. That is the difference between a vendor and a growth partner.
What AI governance actually looks like
Governance is the evaluation layer most buyers skip, and most agencies avoid discussing. That is a problem, because the risk is already materialising. The IAB’s State of Data 2025 found that around 70% of marketers have experienced AI-related incidents in live campaigns. Yet only approximately one-third have formal AI governance tools in place or planned. The risk is not theoretical. Most teams have already felt it.
What governance covers in practice is worth understanding clearly.
Fact-checking and hallucination control. AI generates confident, plausible output that can be factually wrong. An agency without a documented review process is one client away from a credibility incident.
Brand voice integrity. AI models trained on generic internet data do not know your brand. Without structured voice controls and human editorial review, AI-generated content tends toward the average. That is a commercial problem in a market where distinctiveness is increasingly the point.
Data handling accountability. A credible agency should be able to explain exactly how client data is handled, protected, and separated from public model training environments. Vague answers here are a commercial risk, not a minor process gap.
Sign-off and accountability. Who reviews AI-assisted work before it reaches a client, and what standard are they held to? A credible agency can name that person and describe that process. Vague answers here carry real commercial risk.
A practical checklist for the conversation
The questions below are worth taking into any agency pitch. They are not designed to catch anyone out. They are designed to surface whether the thinking behind an agency’s AI capability is real.
| What to ask | Considered answer | Worth noting |
|---|---|---|
| Show me the brief that preceded this work | Documented, client-specific, hypothesis-driven | Generic template or no brief at all |
| What did you reject from the AI output? | Specific examples with clear editorial reasoning | “We review everything” with no specifics |
| How do you validate AI outputs commercially? | Named metrics, documented feedback loops | Volume and speed claims with no outcome data |
| What would you not use AI for in my project? | Honest, specific list of human-mandatory decisions | “AI can do it all” or vague reassurance |
| Who is accountable for AI-assisted work? | Named person with a clear sign-off process | “The team reviews it” |
| Can you show me your process model? | A diagram with decision points and validation stages | A list of tools |
| Will this make our team more capable over time? | Knowledge transfer built into the engagement | Process kept opaque; dependency by design |
What genuine AI capability actually looks like
A marketing agency that genuinely understands AI can show its work at every stage: the brief that preceded the output, the judgement applied during iteration, the validation against real commercial data, and the governance that keeps quality and accountability intact.
The question is never whether an agency uses AI. At this point, almost all of them do. The question is whether AI has made their thinking better — whether it has improved the quality of their decisions, the rigour of their process, and the commercial results they deliver for the businesses that trust them. And beyond that: whether working with them leaves your organisation sharper, more capable, and better equipped to grow intelligently in the AI age.
That is what the checklist above is designed to surface. Take it into the conversation. The strongest AI partners are not afraid to make their thinking visible. They know that in the age of AI, trust is built through clarity, judgement, and accountable decision-making.
“Evaluate proof of better thinking and decision quality — not promises of speed.”
