All articles
DevOps & CloudOctober 2, 2026 6 min read

GCP Cloud Run Min-Instances vs Startup CPU Boost: Which One Actually Fixes Cold Starts

We ran a Node and a Java service on Cloud Run through both knobs — min-instances and Startup CPU Boost — to see which one actually earns its keep. The answer depends on your runtime, your traffic shape, and how honest you are about idle cost.

GCP Cloud Run Min-Instances vs Startup CPU Boost: Which One Actually Fixes Cold Starts

Cloud Run has two obvious levers for cold starts: keep instances warm with --min-instances, or let Google temporarily overclock the container during boot with Startup CPU Boost. Both show up in every performance thread, usually with someone insisting one of them is strictly better. After running both on two production-ish services for a few months, we don't think either answer is honest.

Here's what we actually measured, and where each knob quietly stops helping.

The two knobs, in one paragraph each

Min-instances keeps N containers always-warm. You pay a reduced idle rate for CPU/memory on those instances even when they serve zero requests. Requests that land on a warm instance skip the container start, the runtime boot, and any lazy initialisation your framework does on first request. The tradeoff is a floor on your bill that scales with N and with instance size.

Startup CPU Boost gives the instance extra CPU during the startup phase — Google doesn't publish an exact multiplier, but in our tests boot-bound work finished in roughly half the wall time versus the same config without it. You only pay for the boost indirectly: the instance starts faster, so the billed startup window is shorter. There's no idle cost. It does nothing for an instance that's already running.

They solve different problems. Min-instances removes cold starts from a slice of traffic. Boost makes the cold starts you still have less painful.

The test setup

Two services, both deployed to europe-west1, 2nd gen execution environment, 1 vCPU / 512 MiB unless noted:

  • svc-node: Node 20, Fastify, Prisma against Cloud SQL (Postgres) via the connector sidecar. Cold path does a schema introspection and one warmup query.
  • svc-java: Spring Boot 3.2 on JDK 21, Hikari pool of 5, same Postgres. Classic JVM warmup profile.

We generated traffic with a small k6 script that mixes steady background load with periodic bursts, so we'd see both the steady-state p50 and the cold-start-dominated p99. We ran four configurations per service:

  1. Baseline: no min-instances, no boost
  2. Boost only
  3. Min-instances = 1, no boost
  4. Min-instances = 1 + boost

Numbers below are directional — your runtime, your dependencies, and your region will shift them. We're reporting ranges from several runs, not single-shot heroics.

What cold start actually means here

Cloud Run's own startup_latencies metric measures from container start to the health check passing. That's not what your user feels. Your user feels container start plus the first request hitting a cold code path — JIT compilation, lazy ORM init, DNS, TLS handshakes to downstream services. We measured end-to-end from the k6 client, not from GCP's metric, because the gap between the two is the whole point.

Results: Node service

For svc-node, cold start end-to-end sat around 1.4–1.8s baseline. Steady-state p50 was ~45ms.

ConfigCold p99 (ms)Warm p50 (ms)Idle cost delta
Baseline1600–1800450
Boost only900–110045~0
Min=170–120*45+full idle
Min=1 + Boost70–120*45+full idle

*The cold p99 for min=1 isn't really a cold start — it's the tail of requests that happened to land on a scaled-out instance during a burst. The warm instance absorbed most traffic.

Takeaway: for a lightweight Node service, Boost alone cut cold start latency roughly in half for free. Min-instances helped more, but only for the fraction of traffic that would've hit a cold instance anyway. If your traffic is bursty with long quiet periods, that fraction is high. If you have steady background traffic, min=1 might be solving a problem you don't have.

Results: Java service

This is where things diverge sharply.

ConfigCold p99 (ms)Warm p50 (ms)Idle cost delta
Baseline7000–9500600
Boost only3800–520060~0
Min=180–15060+full idle
Min=1 + Boost80–15060+full idle

Spring Boot on cold Cloud Run is painful. Boost nearly halved it, which is meaningful, but a 4-second cold start still fails most SLOs we'd write. For JVM workloads, min-instances is less of an optimisation and more of a requirement if you care about tail latency. Boost on top of min-instances is cheap insurance for the moments autoscaling adds a new instance during a burst.

If you must run JVM on Cloud Run and can't justify always-on instances, look at CRaC or AOT compilation (GraalVM native image) before you start tuning these knobs. We've seen Spring Boot native images boot in 200–400ms, which changes the whole conversation.

The cost conversation nobody wants to have

Min-instances isn't free, and the pricing page undersells how it compounds. A single min-instances=1 on a 1 vCPU / 512 MiB instance in a European region runs roughly a few dollars a month per service at idle CPU rates. Multiply by:

  • Number of services (easy to hit 20+ in a microservice setup)
  • Number of environments (dev, staging, preview, prod)
  • Number of regions if you're multi-region

We've seen teams casually enable min=2 across 30 services in three environments and wonder why their Cloud Run bill jumped by four figures a month. The fix wasn't disabling it — the fix was being deliberate: min-instances on the five user-facing services, Boost everywhere else.

A deployment snippet we actually use

# cloud-run service.yaml fragment
spec:
  template:
    metadata:
      annotations:
        run.googleapis.com/startup-cpu-boost: 'true'
        run.googleapis.com/execution-environment: gen2
        autoscaling.knative.dev/minScale: '0'  # overridden per-env
        autoscaling.knative.dev/maxScale: '50'
    spec:
      containerConcurrency: 80
      containers:
        - image: europe-west1-docker.pkg.dev/.../svc-node:${SHA}
          resources:
            limits:
              cpu: '1'
              memory: '512Mi'

In Terraform, the equivalent via google_cloud_run_v2_service:

resource "google_cloud_run_v2_service" "svc" {
  name     = "svc-node"
  location = "europe-west1"

  template {
    scaling {
      min_instance_count = var.env == "prod" ? 1 : 0
      max_instance_count = 50
    }
    containers {
      image = var.image
      resources {
        limits = { cpu = "1", memory = "512Mi" }
        startup_cpu_boost = true
      }
    }
  }
}

Key detail: startup_cpu_boost is cheap to leave on everywhere. The decision that costs money is min_instance_count.

When each knob is the wrong answer

Boost is the wrong answer when your startup is bottlenecked on I/O, not CPU. If your container spends four seconds waiting on a Secret Manager fetch, a VPC connector attach, or a slow database handshake, extra CPU doesn't help. We've debugged "why isn't Boost working" tickets that turned out to be a sync call to a slow internal API during module load. Fix the init path first.

Min-instances is the wrong answer when your traffic is genuinely steady. If you're serving >1 req/sec consistently, you probably already have warm instances and min=1 is just paying for what autoscaling would've done anyway. Check your instance_count metric before enabling it — if it never dips to zero, min-instances is pure cost.

Both are the wrong answer when the real problem is instance count churn. Cloud Run will scale down aggressively, and if your traffic oscillates around the threshold, you'll see new instances spin up constantly. The fix is often --max-instances tuning and containerConcurrency tuning, not warm pools.

What we'd do

If you're starting from zero on a new Cloud Run service:

  1. Turn on Startup CPU Boost by default in your Terraform module. It's effectively free and it meaningfully helps cold starts across most runtimes.
  2. Leave min-instances=0 until you have real latency data. Measure cold start as the client sees it, not as startup_latencies reports it.
  3. If your p99 is dominated by cold starts and the service is user-facing, set min-instances=1 on prod only — not staging, not preview.
  4. For JVM services, treat min-instances as mandatory for prod, and seriously evaluate GraalVM native image before scaling the warm pool.
  5. Review min-instances quarterly against actual traffic. Services that outgrew the need are a quiet source of waste.

If you'd like a second pair of eyes on a Cloud Run setup — cold starts, bill, or both — our cloud and DevOps team does this kind of audit regularly.

#GCP#Cloud Run#Serverless#Performance#Cost Optimization

Want a team like ours?

72Technologies builds production software for the kind of teams who actually read this blog.

Start a project