Sentry Performance Quotas Bit Us Mid-Incident: A Rate Limiting Postmortem
During a payment outage, Sentry silently dropped 40% of our transactions right when we needed them most. Here's what tripped the quota, how we found out, and the rate-limiting setup we run now.
It was 02:14 on a Tuesday when our checkout started throwing 502s. It was 02:41 when we realized Sentry had stopped showing us new transactions from the very service that was on fire. The outage was 27 minutes old before we understood that our observability tool had become part of the incident.
This is the story of that night, why Sentry's performance quotas behaved exactly as documented (and still surprised us), and the guardrails we now put in place on every project.
The Setup Before the Incident
We had a fairly standard Next.js + Node monorepo deployed on Vercel and AWS, with Sentry doing double duty for errors and performance. The relevant config looked roughly like this:
// sentry.server.config.ts
Sentry.init({
dsn: process.env.SENTRY_DSN,
tracesSampleRate: 1.0, // yes, really
profilesSampleRate: 1.0,
environment: process.env.VERCEL_ENV,
integrations: [
Sentry.httpIntegration(),
Sentry.prismaIntegration(),
],
});
The tracesSampleRate: 1.0 was a decision from six months earlier, when traffic was a fraction of what it had become. Nobody flagged it during the last three capacity reviews because errors were well under quota and the performance bill looked normal — right up until it didn't.
Our Sentry plan had a monthly transaction quota with a spike protection window. We had also, on the recommendation of a well-meaning engineer, set a per-project rate limit of 300 events per minute on the checkout project after a noisy deploy in Q3. That number was never revisited.
Why 300/min felt fine
Under normal load, checkout produced maybe 80–120 transactions per minute. 300 gave us a comfortable 2.5x buffer. What we didn't model was what happens when a partial outage triples request retries from the mobile client, and the same user hits the pay button four times in twelve seconds.
What Actually Happened
At 02:14, a Stripe webhook consumer started timing out because of a downstream Postgres connection storm. Our mobile clients had aggressive retry logic (fixed 3s backoff, no jitter — a separate sin we've since fixed). Transactions per minute on the checkout project jumped from ~110 to a sustained ~950.
Sentry's per-project rate limit kicked in around 02:17. From that point, roughly 40% of transactions were dropped at the ingest edge. The errors still came through — errors have a separate quota — but the traces that would have told us which database call was hanging simply weren't there.
Worse: the transactions Sentry chose to drop weren't random from our perspective. Ingest rate limiting drops what arrives after the bucket is empty, so we were disproportionately losing the slow requests (the ones that took longer to finish and report). Our p95 charts looked artificially healthier than reality for the entire incident window.
The observability tool was telling us a story about the incident that was systematically biased toward the fast requests. The slow, broken ones were being silently discarded.
How we finally noticed
An on-call engineer opened the Stats page in Sentry (Settings → Stats & Usage) and saw the giant orange bar labeled "Rate Limited." That's the tell. If you've never looked at that page, open it now — the categorization between Accepted, Filtered, Rate Limited, and Invalid is the single most useful diagnostic when your APM feels wrong.
The Postmortem: Three Things We Got Wrong
1. Static sampling at 100%
tracesSampleRate: 1.0 is fine for a service doing 5 RPS. It's negligent at 500 RPS. We were paying for volume we couldn't actually use, and we were pressed against quota ceilings we'd forgotten existed.
The fix wasn't just lowering the number. A flat 0.1 would have hidden low-frequency but high-value endpoints (webhook handlers, admin actions). We moved to a tracesSampler function:
Sentry.init({
dsn: process.env.SENTRY_DSN,
tracesSampler: (ctx) => {
const name = ctx.transactionContext?.name ?? '';
// Always sample webhooks and payment flows
if (name.startsWith('POST /api/webhooks/')) return 1.0;
if (name.includes('/checkout/')) return 0.5;
// Health checks: never
if (name === 'GET /api/health') return 0.0;
// Everything else
return 0.05;
},
});
In our experience, moving from a flat 1.0 to a route-aware sampler cut transaction volume by roughly 80–90% on high-traffic services with no meaningful loss of debuggability. Your mileage will depend heavily on your route distribution.
2. A rate limit set once and forgotten
300 events/min was a reasonable number in Q3. By the time of the incident, it was a chokepoint. We now treat per-project rate limits as a config item that gets reviewed whenever a service's baseline traffic doubles, and we manage them via Terraform against the Sentry provider so a drift shows up in PRs:
resource "sentry_project" "checkout" {
organization = "our-org"
team = "payments"
name = "checkout"
platform = "node"
}
resource "sentry_project_rate_limit" "checkout" {
organization = "our-org"
project = sentry_project.checkout.slug
window = 60
count = 3000
}
The count is now derived from a rough formula: peak_expected_rps * 60 * 1.5. It's a soft ceiling, not a cost control — spike protection at the org level handles cost.
3. No alert on rate-limited events
This was the biggest miss. Sentry exposes rate-limit stats via the API and can send them to a webhook. We now have a scheduled job that hits /api/0/organizations/{org}/stats_v2/ every five minutes, filters for outcome=rate_limited, and pages if the value crosses zero for any production project over a 15-minute window.
The alert copy is deliberately blunt: "Sentry is dropping your data. Charts are lying." On-call has thanked us for that wording twice already.
What We Kept, What We Dropped
We considered ripping Sentry performance out entirely and shifting to a self-hosted OpenTelemetry pipeline with Tempo or Honeycomb. We didn't, for three reasons:
- The error/trace correlation in Sentry is genuinely good, and rebuilding it across two tools is a real engineering cost.
- Our team already knows the Sentry UI. Retraining on-call during a period of business growth wasn't a fight worth picking.
- The rate-limit issue was our misconfiguration, not a product defect.
What we did do is add OpenTelemetry as a parallel pipeline for the checkout and payments services, exporting to a small Grafana Tempo instance. It's a cheap insurance policy — if Sentry ingest fails or quota trips again, we still have raw traces for the services that matter most. The dual-write costs us maybe 3–4% extra CPU on those pods, which is a price we'll pay.
On dynamic sampling in Sentry
Sentry's server-side dynamic sampling (available on higher-tier plans) will boost low-frequency transactions and dampen high-frequency ones automatically. It helps, but it doesn't remove the need for a sane client-side tracesSampler. Dynamic sampling operates on what you send it; if you flood the ingest, you still get rate-limited before the clever sampling kicks in.
A Checklist We Now Run for Every New Service
tracesSamplerfunction defined, not a statictracesSampleRate- Health checks and static asset routes explicitly set to 0.0
- Per-project rate limit set to
peak_rps * 60 * 1.5, managed in IaC - Alert on
outcome=rate_limited > 0sustained for 15 minutes - Sentry Stats page bookmarked in the on-call runbook
- For tier-1 services: parallel OTel export to an independent backend
Where We'd Start
If you're reading this and you have tracesSampleRate: 1.0 in a production Sentry config, don't refactor the whole sampler today. Do two things this week: open the Stats & Usage page and check if you've been rate-limited in the last 30 days, and add an alert on rate-limited outcomes. Those two changes take under an hour and will tell you whether your observability is quietly lying to you right now.
Everything else — the sampler function, the Terraform-managed limits, the parallel OTel pipeline — is a project. The alert is a Tuesday afternoon. Start there.
If you want a hand designing an observability setup that survives its own incidents, that's the kind of work our DevOps and reliability team does day in, day out.
Want a team like ours?
72Technologies builds production software for the kind of teams who actually read this blog.
Start a projectKeep reading

GCP Cloud Run Min-Instances vs Startup CPU Boost: Which One Actually Fixes Cold Starts
We ran a Node and a Java service on Cloud Run through both knobs — min-instances and Startup CPU Boost — to see which one actually earns its keep. The answer depends on your runtime, your traffic shape, and how honest you are about idle cost.

Pulumi vs Terraform for a Multi-Account AWS Migration: What Actually Hurt
We migrated a five-account AWS estate from Terraform to Pulumi, then partially back. Here's the honest breakdown of what each tool did well, where they bled us, and which choice we'd make again.
Terraform State Locking on S3 Native: What We Learned After Ditching DynamoDB
HashiCorp finally shipped native S3 state locking. We migrated three production stacks off DynamoDB and hit two sharp edges you should know about before you follow.
