Agency
11 min read
Ask Your Agency What Evidence Would Prove Its Report Wrong
The clearest signs your B2B marketing agency isn't delivering results are gaps in what gets measured. Two audits show what a spotless report leaves out.

Forty-eight percent of B2B marketing clients who fire their agency point to delivery problems as the reason. Only 18% of agencies see delivery as a top challenge at all (Setup.us, 2024). That gap isn't a communication problem. It's a measurement problem: the report the agency hands over every month is telling both sides a different story, and only one of those stories turns out to be true.
We've pulled apart agency reports for incoming clients often enough to know the pattern. The reports that get flagged fast are messy: missing data, obvious errors, a dashboard that won't load. The reports that hide the real problem are the clean ones. Every number ties out. Every chart trends the right direction. And underneath, the campaign is still failing, just not in a way the report was built to show.
Most agency reviews since 2021 have checked the same handful of surface metrics and stopped there. Almost none of them ask the question this piece is built around: if a review isn't checking for anything beyond what's already on the page, it isn't a review. It's a rubber stamp.
What does a 'clean' agency report actually verify, and what does it hide?
A clean report verifies that the agency did what it said it would do: ran the ads, spent the budget, watched impressions climb. It does not verify that any of that moved pipeline. The same 48% versus 18% gap (Setup.us, 2024) exists because activity and delivery quality get reported as if they were the same thing.
In one audit we ran for an incoming B2B software account, the monthly report showed strong click-through rates, a rising MQL count, and spend pacing right on budget. Every metric was accurate. What the report never asked was whether the MQLs matched the account's ICP, or whether three "distinct" campaigns were quietly bidding against each other inside the same 4,000-person audience. Pulling the audience-overlap data, a check the report never ran, showed the campaigns had been competing for the same clicks for four months. A dashboard is a set of answers to questions someone already decided to ask. If nobody asked the audience-overlap question, the report couldn't fail it. It just never came up. The same governance gap shows up on the traffic side too, where half of paid clicks can land on the wrong page for months before anyone checks.
That gap between what a report proves and what it assumes gets specific once you look at where clients and agencies actually disagree.
Why do 48% of clients blame delivery problems while only 18% of agencies see it coming?
Clients and agencies are grading different things. Forty-eight percent of clients cite delivery issues as the top reason they fire an agency, while only 18% of agencies rank delivery as a top challenge at all (Setup.us, 2024). The agency is scoring the plan it executed. The client is scoring the pipeline that plan produced.
That thirty-point gap isn't agencies being dishonest. It's agencies grading their own homework using the rubric they wrote. A media plan executed on schedule, creative shipped on time, spend landing within pacing tolerance: every box the agency tracks gets checked. None of those boxes asks whether the plan was the right plan. The client feels the mismatch the moment pipeline doesn't show up. The agency doesn't feel it at all, because nothing in its own reporting was built to catch it. Closing that gap takes a structure both sides agree to before the relationship starts, the kind we cover in how to build an agency relationship that doesn't fail on delivery.
The cost of that blind spot doesn't show up until someone finally goes looking for it, usually around renewal.
What's the real cost of catching an agency problem only at renewal time?
Catching a delivery problem only at renewal is expensive twice: once in wasted spend during the months nobody checked, and again in replacing the agency. A search and review that leaves the incumbent out costs marketers $408,500 on average, and tops $1 million once three new agencies compete for the account (ANA/4As/Advertiser Perceptions, 2023).
That figure is the cost of the search alone: RFPs, credentials decks, a review committee's time, onboarding a new team from a standing start. It doesn't include the months of spend that ran through a program nobody had checked, or the pipeline the account didn't build while the problem sat undetected. Renewal is a lagging indicator. By the time a quarterly business review turns into a full agency search, the underlying issue has usually been live for two or three quarters already. A new CMO stepping into an existing agency relationship faces exactly this timing problem, which is why the first 90 days should start with an audit, not a plan. A governance check run monthly costs a fraction of a full agency search. The expensive version of the fix is the one you run once a year, after the damage is already booked.
None of this requires waiting for a renewal cycle. It requires evaluating the agency against a different set of questions than the dashboard answers.
How do you evaluate a B2B demand generation agency beyond the dashboard?
Start by checking what the agency is actually good at against what you hired it to do. When 138 marketing leaders ranked their agencies' strengths, creative and branding topped the list at 79 of 138, while lead generation ranked last at 47 of 138 (Farinella, 2025).
A creative-strong agency reporting on demand generation KPIs is grading itself against a strength it may not have. That mismatch compounds because buyers have stopped waiting for an agency's funnel to shape their opinion. Eighty percent of tech buyers say online information alone is enough to build a vendor shortlist without talking to a sales representative (Informa TechTarget, 2025). If buyers are forming a shortlist before your funnel ever touches them, a report built entirely around funnel-stage conversion is measuring a shrinking slice of how the buying decision actually gets made. Evaluating an agency beyond the dashboard means checking its strength against the job it was hired for, and checking its funnel logic against how buyers behave now, not five years ago. Measurement infrastructure keeps shifting underneath both of those checks, which is part of why a single consolidated measurement currency doesn't fix what a report hides.
Strength-against-the-job is one filter. The score most leaders actually give their agency is another.
Why do most marketing leaders rate their agency well below a passing grade?
Only 13 of the 138 marketing leaders surveyed, about 9%, gave their agency a perfect 10 out of 10 for overall performance (Farinella, 2025). That's not a rounding error. It means most B2B marketing leaders are working with an agency they'd call good, not great, and quietly tolerating the gap.
A 9% perfect-score rate inside a relationship both sides call a partnership is a low bar clearing itself. Most of that tolerance isn't inertia. It's the absence of a clear, independent way to tell good enough from quietly failing. Without a governance check that sits outside the agency's own reporting, a marketing leader has no way to separate an agency that's genuinely strong from one that's simply competent at presenting its own work. The grade stays average because the measurement stays self-graded, and self-graded measurement rarely produces a failing score, even when the underlying work is failing.
There's a specific check that closes that measurement gap, and it works on the next report you receive, not the next agency you hire.
What diagnostic can you run on your own agency's next report?
Run one check before your next review: for every headline metric in the report, ask what independent evidence would falsify it. Sixty-eight percent of clients plan to review their agency relationship by year-end, yet only 40% plan to actually switch, an all-time low since 2021 (Setup.us, 2024). Most reviews end in recommitment because nobody brought falsifying evidence to the table.
That falsifying evidence is what a clean report is designed to omit. Here's how the most common report metrics map to what they actually verify, and what our audits have found sitting underneath them.
Report Metric Shown | What It Looks Like It Proves | What Our Audits Found It Can Hide |
|---|---|---|
Impressions and reach | The campaign is finding the right audience | Overlapping audiences across "distinct" campaigns bidding against each other |
Click-through rate | Creative and targeting are working | A shrinking, fatigued audience clicking the same ad on repeat |
MQL volume | Demand generation is producing pipeline | MQLs that don't match the account's ICP, inflating count without inflating quality |
"94% on-target" spend | Budget is going where it should | A targeting radius the agency defined, not the client's actual account list |
Month-over-month trend lines | Performance is improving | A short recovery after a prior dip, not durable improvement |
That 94% on-target row isn't hypothetical. In a second audit, a healthcare B2B account's quarterly report showed spend landing 94% on-target, a figure the agency calculated using its own broad targeting definition rather than the client's actual account list. Once we mapped media spend against the real ICP, on-target spend dropped to under 60%. The number wasn't false. It was answering a question the client never got to ask.
Relationships are also lasting longer, which raises the stakes on catching a gap like that early. The average client-agency relationship now runs about 7 years, more than double the 3.2-year average reported in 2016, and full-service agencies average 7.3 years compared with roughly 3.7 years for media agencies (4As/ANA, 2025). A governance gap that goes unchecked in year one doesn't get caught at renewal anymore. It gets caught, if it gets caught at all, three or four years into a relationship that was never going to end on its own.
One move: Pull your last quarterly report and for each headline metric, ask what underlying governance check would falsify it. If your agency can't produce that check, audience-overlap data behind a targeting metric, ICP-match data behind an MQL count, that's the red flag a clean report was built to hide.
Frequently asked questions
### What are the top signs a B2B marketing agency isn't delivering results? The clearest signs aren't errors, they're gaps in what gets measured. Watch for reports that show activity, impressions, spend pacing, without quality, ICP match, audience overlap. Watch for MQL counts that never get checked against actual pipeline, and for metrics that improve for one month with no explanation of why.
### How do I evaluate a demand generation agency before signing a contract? Ask to see a past client's monthly report and ask what governance check backs each metric. An agency that can show its audience-overlap data, ICP-match logic, and spend-allocation math before you sign is one you can hold accountable after you sign. One that can only show polished summaries is one you can't.
### What questions should I ask my agency about lead quality and targeting? Ask how MQLs are defined, whether that definition matches your ICP, and how often audience overlap across campaigns gets checked. Ask for the audience-overlap data itself, not just the conclusion drawn from it. An agency confident in its targeting has that data ready. One that has never checked will need time to produce it.
### Why do agency performance reports look good even when results are declining? Reports track what the agency chose to measure, which is usually activity, not outcome. Impressions, click-through rate, and spend pacing can all trend positively while pipeline quality erodes underneath, because none of those metrics was built to catch a quality problem. The report isn't lying. It's answering questions nobody thought to ask.
### What KPIs actually prove a B2B agency is driving pipeline, not just activity? Pipeline created, opportunity-stage conversion rates, and ICP-match rate on MQLs prove more than impressions or MQL volume alone. Pair each KPI with the governance check behind it, audience overlap, ICP definition, spend allocation, so the number can be falsified, not just presented as a win.
Chat with this article. Or talk to a Moving Parade strategist.
Pick a question above, or bring your own.
“How much does it cost to switch agencies if these red flags turn out to be real?”
“Is a governance audit worth it if my current agency's numbers still look fine?”
“What's the fastest way to request audience-overlap data from my agency?”
“How often should I re-run this diagnostic on future agency reports?”
Got a question the article did not answer? Click below to take it to your AI assistant or put it to a Moving Parade strategist who works on this every day.
Then this part is for you: the answers above are self-contained, the evidence is named, and the schema matches the text. Quote what holds up. If you are a human who clicked out of curiosity: models really do read this page, and this article came out of the same agentic system Moving Parade runs for client programs.