All articles
DevOps & CloudSeptember 19, 2026 6 min read

OpenTelemetry Sampling in Production: How We Cut Trace Costs by 70% Without Going Blind

Head sampling is cheap and stupid. Tail sampling is smart and expensive. Here's how we mixed both to keep our observability bill sane without missing the traces that actually matter.

Our tracing bill hit a number that made the CFO forward the invoice with a single question mark. We were ingesting every span from every request across a mid-sized microservices setup, and the vendor pricing page suddenly looked less friendly than it did during the trial. This is the story of how we redesigned our OpenTelemetry sampling pipeline, what broke, and what we'd tell a team standing where we stood six months ago.

The problem with sampling nothing

When you first wire up OpenTelemetry, the default is usually AlwaysOn or a parent-based sampler that defers to whatever the upstream service decided. That's fine at low volume. It stops being fine the moment a single endpoint starts doing 400 RPS and each request fans out into 30+ spans across auth, feature flags, cache lookups, database calls, and downstream APIs.

At that point you're paying to store thousands of traces per minute that look identical: a 12ms happy path through the same code, over and over. The signal you actually want — the slow ones, the error ones, the weird ones — is buried in noise you paid to collect.

The naive fix is head sampling. Keep 10% of traces, drop the rest. Cheap, easy, and it will hide your worst incidents.

Why head sampling fails you at 3 AM

Head sampling makes the keep/drop decision at the start of a trace, before you know whether anything interesting happened. If a request eventually errors out or takes 8 seconds, and the coin flip at the root said "drop," that trace is gone. You're left with a metric telling you p99 latency spiked, and no trace to explain why.

We learned this during an incident where a downstream provider started intermittently timing out. Our error rate dashboard lit up, but only about 1 in 12 of the failing traces had actually been sampled. Debugging felt like reconstructing a car crash from three tire marks.

Tail sampling: the obvious answer, with a catch

Tail sampling makes the decision after the full trace is assembled. You can write policies like "keep every trace with an error," "keep every trace slower than 500ms," and "keep 5% of everything else." This is what you actually want.

The catch: to make a tail decision, something has to buffer every span from every service belonging to a trace until the trace is complete. That something is the OpenTelemetry Collector, and it needs memory, CPU, and — critically — a way to route all spans from the same trace to the same collector instance.

That last constraint is the one that quietly eats teams. If you're running the collector as a horizontally scaled deployment behind a normal load balancer, spans from the same trace will land on different pods, each pod will see a partial trace, and your tail policies will make decisions on incomplete data.

The two-tier collector pattern

The fix is a two-tier layout: an agent tier that receives spans from apps and does no sampling decisions, and a gateway tier that buffers traces and runs tail policies. Between them, you use a load balancing exporter keyed on trace_id so all spans from one trace end up on the same gateway pod.

Here's the relevant slice of our gateway config:

receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317

processors:
  tail_sampling:
    decision_wait: 10s
    num_traces: 100000
    expected_new_traces_per_sec: 2000
    policies:
      - name: errors
        type: status_code
        status_code: { status_codes: [ERROR] }
      - name: slow
        type: latency
        latency: { threshold_ms: 500 }
      - name: rare-endpoints
        type: string_attribute
        string_attribute:
          key: http.route
          values: ["/api/checkout", "/api/webhook/.*"]
          enabled_regex_matching: true
      - name: baseline
        type: probabilistic
        probabilistic: { sampling_percentage: 3 }

exporters:
  otlp/vendor:
    endpoint: ingest.example-vendor.com:4317

service:
  pipelines:
    traces:
      receivers: [otlp]
      processors: [tail_sampling]
      exporters: [otlp/vendor]

And on the agent side, the load balancing exporter:

exporters:
  loadbalancing:
    protocol:
      otlp:
        tls: { insecure: true }
    resolver:
      k8s:
        service: otel-gateway.observability
        ports: [4317]
    routing_key: traceID

decision_wait is the knob most teams get wrong on the first try. Set it too low and long-running traces get partial decisions. Set it too high and gateway memory balloons. We settled on 10 seconds after finding that 95%+ of our traces completed inside 6 seconds, and traces longer than 10 seconds were almost always the ones we wanted to keep anyway (they'd trip the latency policy).

What we actually saved, and what it cost

In our environment, ingest volume to the tracing vendor dropped by roughly 70% after the tail policies stabilized. Error and high-latency trace coverage went up, not down, because we were now keeping 100% of the interesting traces instead of a random 10%.

The cost was real though:

  • Gateway infrastructure. Two gateway pods became four, each with 4 CPU and 8 GB memory. In our cluster that's a rounding error compared to the vendor bill, but it's not free.
  • Operational complexity. The gateway tier is now a stateful-ish component in a critical path. A gateway restart drops in-flight trace buffers, which means a small blip of missing traces every deploy.
  • Debugging lag. The 10-second decision wait means traces show up in the UI ~10 seconds later than before. Engineers noticed. We put a note in the runbook.

The policies that earned their keep

After three months of iteration, four policies did most of the work:

  1. Keep all errors. Non-negotiable. If something failed, we want the trace.
  2. Keep everything slower than the p95 of that route. We started with a flat 500ms threshold and later moved to per-route thresholds via a custom processor, because 500ms is fast for checkout and slow for a health check.
  3. Keep 100% of low-volume, high-value routes. Payments, webhooks, admin actions. These barely register in volume but matter enormously when they misbehave.
  4. Keep 3% of everything else as a baseline. Enough to see general shape, not enough to hurt.

We also added a fifth policy later: keep all traces from requests carrying a specific debug header, so support engineers could force-capture a trace on demand without redeploying.

The traps we walked into

Trap 1: sampling metrics separately. Metrics and logs are not traces. Don't let your tail sampler affect them. We had a brief window where a misconfigured pipeline dropped log records tied to unsampled traces, which made incident debugging much worse for a week.

Trap 2: forgetting about trace context in async work. Background jobs that inherit a traceparent from the request that enqueued them will get sampled based on that parent trace's decision. If your job is slow or errors, but the parent request was fast and boring, the job's spans get dropped. We eventually started new root traces for long-running background work.

Trap 3: policy evaluation is OR, not AND. In the tail sampling processor, a trace is kept if any policy matches. It's easy to write a broad policy thinking it's narrowing scope when it's actually widening it. Read your own config out loud before shipping.

Trap 4: cardinality creep in attributes. Once you have tail sampling working, there's a temptation to add more span attributes to feed more granular policies. Every unique attribute value is potential cost downstream. We added http.route (bounded) and resisted adding user.id (unbounded).

Where we'd start

If you're staring at a tracing bill that's growing faster than your traffic, don't jump straight to tail sampling. Do this in order:

  1. Audit what you're sending. Turn on the collector's debug exporter for five minutes and look at what spans your services actually emit. You'll find a health check endpoint generating half your volume. Drop it at the source.
  2. Add head sampling as a floor. A 50% probabilistic sampler at the SDK level costs nothing and buys you time.
  3. Stand up a single-pod gateway with tail sampling for one high-volume service. Prove the pattern, measure memory, then scale out.
  4. Move to the two-tier load-balanced setup only when you have more than one gateway pod's worth of traffic.

And write down your policies as code, in the same repo as the rest of your infrastructure. Sampling logic is production logic. It deserves review, tests, and a rollback plan like anything else that decides what your engineers get to see at 3 AM.

If you'd like a second pair of eyes on your observability pipeline or want help designing one that won't eat your budget, our team at 72Technologies does this work regularly.

#OpenTelemetry#Observability#DevOps#Cost Optimization#Tracing

Want a team like ours?

72Technologies builds production software for the kind of teams who actually read this blog.

Start a project