Choosing a B2B demand gen agency comes down to six decisions made in order: whether an agency is the right vehicle for your stage at all, what demand generation will mean in your contract, how you will score candidates against each other, what you will ask on the sales calls, which pricing model you can live with, and what timeline you will hold the winner to. Get the first decision wrong and the other five cannot save you.
Most failed agency engagements were doomed before the kickoff call. The company hired an agency when it needed a founder doing sales. Or it bought a lead count when it needed pipeline. Or it signed a percentage-of-spend deal and then spent a year wondering why every recommendation involved a bigger budget. None of those failures happened because the agency executed badly. They happened because the evaluation skipped a step.
This guide walks through all six steps, with the decision table, the weighted scorecard, the question scripts, and the timeline benchmarks you need to run the evaluation yourself. It is written to be useful even if you conclude, correctly in some cases, that you should not hire an agency at all.
Step 1: Decide whether you need an agency at all
Whether you should hire a demand gen agency depends almost entirely on stage. Pre-product-market-fit companies should not hire one. Companies between roughly $1M and $5M ARR usually get more from specialist freelancers or a fractional leader. An agency, or a hybrid of agency plus a small internal team, starts to pay off once a repeatable sales motion exists, which for most B2B SaaS companies means from about $5M ARR or a funded Series A onward.
The honest version of the pre-PMF case: an agency is a waste of spend before product-market fit. An agency's job is to scale a motion. Before PMF, there is no motion to scale. You are still discovering who buys, why they buy, and what words make them lean in, and that learning only happens when founders sell directly. Paying an agency at that stage buys you polished outbound to an ICP you have not validated, meetings you cannot convert, and a burned prospect list in the exact market you will need later. The agencies worth hiring will tell you this on the first call, and some publish it in their stated ICP: post-product-market-fit only.
| Stage | Typical ARR | Best default | Why |
|---|---|---|---|
| Pre-PMF (pre-seed to seed) | $0-$1M | In-house, founder-led sales | The learning loop is the asset. No agency can find your message for you, and paying one to scale an unproven motion burns budget and market goodwill. |
| Finding repeatability | ~$1M-$5M | Specialist freelancers or a fractional growth leader | You need cheap, fast iterations on one or two channels, not a full pod. A strong freelancer per channel keeps you flexible while the motion firms up. |
| Scaling a proven motion | $5M-$50M (typically Series A-C) | Agency or hybrid | The motion works and the constraint is execution capacity across channels. An integrated agency pod, or agency plus a lean internal team, compounds faster than sequential senior hires. |
| Mature GTM organization | $50M+ | In-house team, agencies for specialist channels | At this scale the core team belongs on payroll. Agencies still earn their keep on specialist channels and overflow capacity. |
Freelancers vs one agency: who pays the coordination tax
The real difference between hiring three specialized freelancers and hiring one agency is who owns coordination. With freelancers, you are the integration layer: you connect the paid person's signals to the outbound person's sequences to the ops person's CRM, and every handoff between them runs through you. That is workable when you are running one or two channels and have someone internal with the time and skill to orchestrate. It stops working when the strategy is allbound, where paid, outbound, content, and operations are supposed to react to each other's signals in days, not sprints. If integration is the point of the program, buy the team that already works as one. If it is not, freelancers are often the better and cheaper call, and an honest evaluation admits that. The primer on inbound vs outbound and when to use each is a useful companion to this decision.
Step 2: Define what demand generation must mean in your contract
In a demand gen contract, demand generation should be defined as qualified pipeline created, not marketing qualified leads delivered. This single definition predicts more about the engagement's outcome than any case study the agency shows you. An agency can hit almost any MQL target with gated ebooks, lookalike audiences, and a generous scoring model, and none of it has to turn into revenue. Pipeline, defined as opportunities your sales team accepted with real dollar values attached, is much harder to fake.
The math behind buyer behavior explains why lead-count contracts mislead. LinkedIn's B2B Institute puts it at about 95%: that share of your potential buyers is not ready to buy today, out-market now and in-market at some point in the future (the 95:5 rule). A lead-count contract pays an agency to harvest contact details from the 95% and label the pile demand. A pipeline contract pays the agency to convert the small in-market slice now while building memory with the rest, which is the actual job. If your funnel already leaks between marketing acceptance and sales acceptance, read how to convert MQLs to SQLs before you sign anything, because an agency inherits that leak.
Put these questions to every vendor before you look at a single deck:
- Which metric will you report as your headline number, and is it pipeline dollars or lead count?
- What counts as sourced versus influenced pipeline, and who adjudicates disputes?
- Where does the data live: our CRM or your tools?
- Will you co-sign a written definition of a qualified opportunity before launch?
- How does your reporting handle a deal that closes nine months after first touch?
- What do you do in a month where the leading indicators are up but pipeline is flat?
An agency that answers these crisply has been judged on pipeline before. An agency that redirects to impressions, traffic, or "brand lift" is telling you what its contract will actually optimize.
Step 3: Score every agency on the same weighted criteria
A demand gen agency scorecard works when every candidate is scored on the same weighted criteria before any sales call happens, so the best presenter does not beat the best operator. Score each criterion 1 to 5 from evidence you can check, multiply by the weight, and rank. The weights below reflect what actually predicts engagement success for B2B SaaS; adjust them to your situation, but decide them before the first call.
| Criterion | Weight | What good looks like | Red flag |
|---|---|---|---|
| ICP fluency | 15% | Names companies like yours unprompted; knows your buyer titles, deal sizes, and sales cycle before you brief them | Generic "we do B2B" positioning; asks you to explain your own buyer |
| Verifiable proof | 15% | Named clients, attributed testimonials, and results you can check on a reference call | Anonymous logos, unverifiable percentage claims, no client who will take your call |
| Channel depth vs integration | 12% | Deep in the two or three channels your buyers actually use, and runs them on one shared data layer | A single-channel shop reselling "full service" through subcontractors |
| Attribution approach | 12% | Every touch lands on contact records in your CRM; reporting you can audit in your own system | Attribution lives only in the agency's dashboard; you get screenshots, not access |
| Team seniority | 10% | The people who sold you do the work, or you meet the actual delivery pod before signing | Senior partners on the pitch, juniors on the account, an account manager as a buffer |
| Pricing model alignment | 10% | Flat fee tied to a defined scope, with no incentive to inflate spend or lead volume | Fees that scale with ad spend or per-lead payouts (see Step 5) |
| Tooling and data ownership | 10% | You own the ad accounts, sending domains, and CRM data, and keep them if you leave | Agency-owned accounts and lists that vanish with the relationship |
| Ramp and onboarding | 8% | A structured onboarding with launch dates per channel; about a month is normal | "We start driving leads in week one" |
| Contract terms | 8% | A minimum term with reasoning behind it, plus clean exit and handover language | Auto-renewing annual lock-in with no handover clause |
Two usage notes. First, require evidence for every score: a named client for proof, a live dashboard for attribution, a meet-the-pod call for seniority. Claims score zero. Second, treat any 1 in verifiable proof or data ownership as disqualifying regardless of the total, because those two failures are the ones you cannot fix after signing.
Step 4: Ask the sales-call questions that separate good agencies from bad ones
The questions that separate good demand gen agencies from bad ones are the ones that cannot be answered with a case study slide. Every agency has a polished deck. Almost none have rehearsed honest answers to questions about failure, staffing, and exits. Ask these six, in roughly this order, and write down the answers verbatim.
"Walk me through a client you lost in the past year and why." A good answer names the miss, what changed afterward, and does it without trashing the client. A bad answer is "we've never really lost one" or a story where every departure was the client's fault. Everyone loses clients; only operators learn from it in public.
"Who exactly will be in our Slack day to day, and can we meet them before we sign?" A good answer introduces the actual pod, with names and backgrounds, before the contract. A bad answer promises "a dedicated team" it will assemble after signature, which usually means whoever has capacity.
"What happens in your first 30 days, before anything launches?" A good answer describes real onboarding: ICP and message work, data and tracking setup, channel infrastructure, and a launch calendar. A bad answer promises leads in week one, which means they are pointing a generic playbook at an unvalidated list.
"What number do you expect us to judge you on at month three, and at month six?" A good answer commits to leading indicators by month three and pipeline by month six, with ranges. A bad answer refuses a number entirely or promises one implausibly early.
"If replies collapse or cost per meeting doubles in month two, what specifically do you change?" A good answer has a diagnostic sequence: audience, then message, then mechanics, in that order. A bad answer is "we'd optimize," with no order of operations.
"What do we keep if we part ways in six months?" A good answer is everything: accounts, domains, data, documentation, and a handover. A bad answer has an awkward pause in it. This question is the fastest tooling-ownership check you can run.
Run this framework against us.
Understory Agency runs paid media, outbound, content, and RevOps as one pod for post-PMF B2B SaaS, on custom flat retainers, never a percentage of spend. Bring the scorecard and the hard questions.
Get in TouchStep 5: Compare pricing models by the incentives they create
The four common demand gen agency pricing models are flat retainer, percentage of ad spend, project-based, and performance-based, and each one pays the agency to optimize for something different. The incentive matters more than the sticker price, because you will live inside that incentive for the length of the contract.
| Model | How it works | The incentive it creates | Watch for |
|---|---|---|---|
| Flat retainer | Fixed monthly fee for a defined scope of services | Keep the client by producing results; the fee does not move when your budget or lead count moves | Vague scope; make the deliverables and operating cadence explicit in writing |
| Percentage of ad spend | Fee scales with your media budget | Recommend more spend; the agency earns more when budgets grow, whether or not performance does | Budget recommendations you cannot cleanly separate from the agency's own revenue |
| Project-based | Fixed fee for a defined build with an end date | Finish the project; nobody is paid to compound the motion afterward | Great for audits and one-time builds, wrong shape for an ongoing demand program |
| Performance-based | Pay per lead, meeting, or opportunity | Maximize the countable unit; volume beats quality every time the two conflict | Lead-quality collapse and definition gaming; "qualified" drifts toward whatever pays |
The incentive analysis favors the flat retainer for ongoing demand gen work, and the reasoning is structural rather than promotional. A flat retainer is the only model where the agency's revenue is independent of the decisions it recommends: it earns the same whether it tells you to raise paid media spend or cut it, the same whether it reports 400 leads or 12 qualified opportunities. That independence keeps the recommendation stream clean. Its weakness is scope drift, which is why the fix in Step 3 is a written scope and cadence, not a different model.
Percentage of spend has the opposite structure: every budget conversation carries a conflict of interest, however honest the people. Performance pricing sounds like alignment and usually is not: paying per lead buys leads, in volume, at whatever quality clears the definition, and your sales team inherits the disqualification work. Project pricing is honest for bounded work and the wrong shape for a motion that should compound month over month.
Step 6: Hold the winner to honest timeline expectations
Realistic demand gen timelines look like this: onboarding takes about four weeks, cold email takes roughly three to four weeks to launch cleanly, LinkedIn outreach two to three weeks, and paid media spins up faster than either, with a competent team getting most campaigns live within one to two weeks of kickoff. An integrated program should be visibly compounding by month three: channels feeding each other signals, cost per opportunity trending down, pipeline building. Judge agencies against these ranges in both directions. An agency promising results in week one is skipping the onboarding that makes results durable; an agency asking for six months of patience before showing any leading indicator is hiding.
The same logic makes minimum terms of about six months reasonable rather than a trap. B2B purchase cycles are long: the joint LinkedIn B2B Institute and Ehrenberg-Bass Institute How B2B Brands Grow research finds, for example, that 75% of companies buy computers once every four years and 80% of companies change banking services once every five years (LinkedIn B2B Institute). A program judged at day 45 is being judged before most of its market has had a reason to move. The contract structure that squares this honestly: a real minimum term, paired with the month-three checkpoint above, written into the agreement, so patience is required but never blind. Ask how results will reach you too. The right answer routes every touch onto contact records in your CRM with a revenue operations layer keeping it clean, so the month-three conversation happens in your data, not the agency's slides.
Where Understory Agency fits, and where it does not
Full disclosure: Understory Agency publishes this blog, and it is one of the agencies you might run this framework against. So here is the honest version of where we land on it, in the same terms we have asked you to hold every vendor to.
Understory Agency is built for post-product-market-fit B2B SaaS companies, typically funded Series A through C, that want demand gen run as one integrated allbound motion instead of a stack of disconnected vendors. Every account gets a pod of at least a GTM engineer, an ops manager, and a paid strategist, working from one shared data layer, with LinkedIn content and creative folded in when the scope calls for them. On the criteria from Step 3: attribution lands on contact records in your CRM with Looker Studio dashboards on top; pricing is a custom flat retainer for each service, never a percentage of spend, scoped to the services you select and the level of service you need; onboarding runs four weeks; and the minimum commitment is typically six months, for exactly the compounding reasons in Step 6. Single-channel engagements work, though the model is strongest run as full allbound.
And where it is not a fit: if you are pre-PMF, Step 1 applies to us too, and the answer is that you should not hire any agency yet, including Understory Agency. If you want a percentage-of-spend or pay-per-lead arrangement, we do not price that way, for the incentive reasons in Step 5. And if you are comparing integrated shops side by side, our roundup of the best integrated marketing agencies is a fair place to see the field, including firms we compete with.
Related reading
FAQ
How do I actually pick a B2B demand gen agency?
Pick a B2B demand gen agency by running six steps in order: confirm an agency is right for your stage at all (post-product-market-fit, usually $5M+ ARR or funded Series A onward), define demand gen in the contract as qualified pipeline rather than MQLs, score every candidate on the same weighted criteria before sales calls, ask questions a case study cannot answer (lost clients, day-to-day staffing, first 30 days, what you keep on exit), choose a pricing model whose incentives you can live with, and hold the winner to honest timelines with a month-three compounding checkpoint.
Should we hire a few specialized freelancers or one agency to run B2B SaaS go-to-market?
Hire specialized freelancers when you are running one or two channels and someone internal has the time and skill to coordinate them, because you become the integration layer between every freelancer. Hire one agency when coordination itself is the point, meaning you want paid, outbound, content, and operations reacting to each other's signals on one data layer. Freelancers are cheaper and more flexible; an integrated pod compounds faster. The deciding question is who pays the coordination tax, not who writes better copy.
We're a 40-person B2B SaaS company. Should we hire an agency to run our go-to-market or build the team in-house?
At 40 people you are usually post-PMF, which makes this a build-versus-buy decision on speed and integration rather than a question of readiness. Building the equivalent function in-house means recruiting, ramping, and integrating three or more senior specialists plus someone to lead them, which typically takes quarters. An agency pod arrives already working as one team, which is why many companies at this size run a hybrid: an internal leader who owns strategy and the sales relationship, with an agency running integrated execution. Whichever way you go, keep strategy, positioning, and closing in-house; they should never be outsourced.
How much does a B2B demand gen agency cost?
B2B demand gen agencies price four main ways: flat monthly retainers scoped to services, a percentage of your ad spend, fixed project fees, or performance pricing per lead or meeting. Evaluate the incentive before the number: percentage-of-spend models earn more when your budget grows, and per-lead models reward volume over quality. A flat retainer, priced from the services you select and the level of service you need, is the structure where the agency's revenue stays independent of its own recommendations, which is why it is the model to prefer for ongoing pipeline work.
How long does it take to see results from a demand gen agency?
Expect about four weeks of onboarding, with cold email taking roughly three to four weeks to launch cleanly, LinkedIn outreach two to three weeks, and paid media spinning up faster than either. First qualified replies and booked meetings can land as early as the end of month one, and an integrated program should be visibly compounding by month three, with channels feeding each other and cost per opportunity trending down. That is also why reputable agencies ask for a minimum commitment of around six months: long enough for compounding to show, with a month-three checkpoint so patience never has to be blind.






