The Retainer Trap: Why Monthly Agency Deals Quietly Kill Your Margin
Retainers look like predictable revenue until you audit the hours. Here's how agencies lose 20-40% margin on retainers without noticing, and the operating model that actually fixes it.
Every agency founder eventually falls in love with retainers. Predictable MRR, happy accountants, a story you can tell investors. Then you audit six months of timesheets and realise you've been running a charity for a mid-market SaaS company.
This is a breakdown of how retainers actually bleed margin, why the bleed is invisible on your P&L until it's too late, and the operating model we now use before signing any monthly deal.
Why Retainers Look Better Than They Are
On paper, a retainer is the dream. You sign a client for, say, $18k/month for "up to 120 hours of engineering". Revenue is booked. Utilisation looks healthy. The account manager stops sweating renewals.
The problem is that retainers hide three things that time-and-materials contracts expose immediately:
- Overrun hours get absorbed silently because nobody wants to send a scope email for 4 extra hours.
- Context-switching cost disappears into "admin" or "internal" time codes.
- Senior escalations happen off the clock because the client already "paid for the month".
A fixed-price project ends. A T&M project bills every hour. A retainer just... continues, and the small leaks compound.
The 120-hour lie
When you sell 120 hours a month, you're implicitly selling 30 productive hours a week from one engineer. Anyone who's actually shipped code knows a senior engineer produces maybe 22-28 hours of real focused work per week — the rest is standups, PR reviews, Slack, and the meeting where the client's PM re-explains the roadmap.
So you sold 120. You can deliver maybe 95 without burning the person out. The other 25 hours either get faked on the timesheet, stolen from another client, or delivered as unpaid overtime. All three are bad.
The Anatomy of a Bleeding Retainer
Here's a pattern we've seen at least a dozen times. The numbers are illustrative, but the shape is real.
Retainer sold: 120 hrs/month @ $150 blended = $18,000
Actual hours logged: 138 hrs (15% overrun, absorbed)
Unlogged Slack/calls: ~12 hrs (nobody tracks this)
Senior escalations: 6 hrs @ $220 real cost
PM overhead: 14 hrs (sold as "included")
Real delivered cost: ~170 hrs of team time
Effective rate: $105/hr
Gross margin: drops from planned 55% to ~32%
The client is thrilled. Your delivery lead is exhausted. Your CFO thinks the account is healthy because revenue is stable. Nobody is looking at the ratio of logged hours to actual hours consumed.
Where the hours actually go
When we started instrumenting retainers properly, the invisible time bucket broke down roughly like this:
- Async client communication (Slack, Loom replies): 6-10 hrs/month
- Unscheduled "quick calls": 4-8 hrs/month
- Reviewing client-side work (their devs, their designs): 3-6 hrs/month
- Re-planning after client priority changes: 4-12 hrs/month
That's up to 36 hours a month that never touch a Jira ticket. On a 120-hour retainer, that's 30% of your capacity vanishing into goodwill.
The Three Retainer Shapes (And Which One Actually Works)
Not all retainers are the same trap. We now categorise them into three shapes before we quote.
1. The Capacity Retainer
"You get X hours a month, use them however." This is the worst shape. Clients treat it as a buffet, priorities shift weekly, and you can't plan sprints. Margins collapse first.
2. The Outcome Retainer
"You get a defined deliverable cadence — two features shipped, one performance review, weekly reporting." Better. It forces scope conversations upfront. But it only works if the client is disciplined about roadmap, which most aren't.
3. The Squad Retainer
"You rent a dedicated pod — 1 tech lead, 2 engineers, fractional PM — for a fixed monthly fee. Capacity is capacity. What ships depends on what you prioritise." This is the only shape we recommend for anything over $15k/month.
The squad retainer works because it moves the conversation from hours to team. Clients stop counting minutes. You stop pretending your PM's time is free. Everyone prices reality.
The Operating Model We Now Use
After enough painful reviews, we standardised on a small set of rules before any retainer gets signed.
Rule 1: Price the whole pod, not the hours
A squad has a fully-loaded monthly cost. Sell that cost plus your target margin. Don't sell hours. If the client asks "how many hours is that?", you answer "a full-time equivalent team, minus standard overhead — roughly 300-340 productive engineering hours across the pod."
Rule 2: Track a "shadow rate"
Every month, calculate the effective hourly rate you actually delivered at, including all the invisible hours. If the shadow rate drops more than 15% below your planned rate for two months in a row, that's a renegotiation trigger, not a Slack complaint.
Shadow rate = monthly retainer revenue
/ (logged hours + estimated unlogged hours)
Healthy: within 10% of planned blended rate
Watch: 10-20% below
Renegotiate: >20% below for 2+ months
Rule 3: Quarterly scope resets
Every 90 days, you sit down with the client and re-baseline. What did we ship? What did the priorities look like vs what we sold? Is the pod the right size? This kills the "we've always done it this way" drift that turns a healthy retainer into a hostage situation.
Rule 4: A hard ceiling on unbilled communication
We budget async time explicitly. If a client's Slack usage pushes past the budget for two consecutive months, we either add a fractional account manager to the pod (billed) or move to scheduled office hours. Neither is punitive. Both are honest.
When a Retainer Is Actually the Right Answer
Despite the above, retainers aren't inherently bad. They're the right model when:
- The client has a live production system that needs ongoing engineering, not a one-off build.
- Priorities genuinely shift often enough that fixed-price would require constant change orders.
- You have enough bench depth to absorb a bad month without over-serving.
- The relationship is 12+ months old and you have real data on their communication and scope patterns.
Retainers are the wrong model when you're using them to avoid a hard sales conversation, or when the client is really asking for staff augmentation dressed up as "partnership". Staff aug should be priced as staff aug — usually on a straight T&M or dedicated FTE basis.
The Uncomfortable Conversation With Existing Clients
If you're reading this and recognising your own book of business, the fix isn't to email every client tomorrow announcing a rate hike. It's a staged conversation.
- Instrument first. Spend 60 days properly tracking real vs logged hours on every retainer. Don't tell the client, don't change anything. Just measure.
- Rank accounts by shadow-rate gap. The worst offenders are your renegotiation targets. The healthy ones stay as-is.
- Reframe the renegotiation as a squad conversation. Don't say "we underpriced you". Say "we've been running this as an hours contract and it's causing prioritisation friction for both of us — here's a pod structure that fixes it". Then present the new price.
- Be willing to lose one. If a client is only viable at the current bleeding rate, they're not a client, they're a subsidy. Losing one bad retainer usually frees up capacity worth more than the revenue.
We've done this exercise across our own book and with clients we advise. In every case, at least one retainer got repriced upward, at least one got restructured, and occasionally one got politely ended. Margins recovered within a quarter.
What We'd Do Next Week
If you run an agency with three or more active retainers, do this:
- Pull the last 90 days of timesheets for each retainer.
- Add a rough estimate of unlogged Slack, calls, and re-planning time — ask the delivery lead, they know.
- Calculate the shadow rate per account.
- Circle every account more than 15% below your target rate.
That list is your Q1 renegotiation queue. Everything else — the pod pricing, the quarterly resets, the communication ceilings — is easier to introduce once you can point at a real number and say "this is what we're actually delivering at". Retainers aren't the enemy. Retainers you never measured are.
If you want a second pair of eyes on your delivery model, that's the kind of thing we work through in our agency and product strategy engagements.
Want a team like ours?
72Technologies builds production software for the kind of teams who actually read this blog.
Start a projectKeep reading

The Client Who Won't Decide: A Playbook for Unblocking Non-Technical Stakeholders
Half of agency delays aren't technical — they're a client who can't make up their mind. Here's how we structure decisions so non-technical stakeholders actually commit, and the sprint keeps moving.

The Discovery Sprint That Pays for Itself: Selling Paid Scoping Before You Quote
Free scoping is where agencies quietly lose money. Here's how to sell a paid discovery sprint that de-risks the build, wins the client's trust, and stops you quoting into a fog.

The Equity Deal That Actually Pays: How Agencies Should Price Sweat for Stock
Client wants to pay you in stock instead of cash. Sometimes that's a gift. Usually it's a landmine. Here's how to structure equity deals so your agency actually gets paid — in cash, shares, or both.
