GCP Cloud Run Min-Instances vs Startup CPU Boost: Which One Actually Fixes Cold Starts
We ran a Node and a Java service on Cloud Run through both knobs — min-instances and Startup CPU Boost — to see which one actually earns its keep. The answer depends on your runtime, your traffic shape, and how honest you are about idle cost.

Cloud Run has two obvious levers for cold starts: keep instances warm with --min-instances, or let Google temporarily overclock the container during boot with Startup CPU Boost. Both show up in every performance thread, usually with someone insisting one of them is strictly better. After running both on two production-ish services for a few months, we don't think either answer is honest.
Here's what we actually measured, and where each knob quietly stops helping.
The two knobs, in one paragraph each
Min-instances keeps N containers always-warm. You pay a reduced idle rate for CPU/memory on those instances even when they serve zero requests. Requests that land on a warm instance skip the container start, the runtime boot, and any lazy initialisation your framework does on first request. The tradeoff is a floor on your bill that scales with N and with instance size.
Startup CPU Boost gives the instance extra CPU during the startup phase — Google doesn't publish an exact multiplier, but in our tests boot-bound work finished in roughly half the wall time versus the same config without it. You only pay for the boost indirectly: the instance starts faster, so the billed startup window is shorter. There's no idle cost. It does nothing for an instance that's already running.
They solve different problems. Min-instances removes cold starts from a slice of traffic. Boost makes the cold starts you still have less painful.
The test setup
Two services, both deployed to europe-west1, 2nd gen execution environment, 1 vCPU / 512 MiB unless noted:
- svc-node: Node 20, Fastify, Prisma against Cloud SQL (Postgres) via the connector sidecar. Cold path does a schema introspection and one warmup query.
- svc-java: Spring Boot 3.2 on JDK 21, Hikari pool of 5, same Postgres. Classic JVM warmup profile.
We generated traffic with a small k6 script that mixes steady background load with periodic bursts, so we'd see both the steady-state p50 and the cold-start-dominated p99. We ran four configurations per service:
- Baseline: no min-instances, no boost
- Boost only
- Min-instances = 1, no boost
- Min-instances = 1 + boost
Numbers below are directional — your runtime, your dependencies, and your region will shift them. We're reporting ranges from several runs, not single-shot heroics.
What cold start actually means here
Cloud Run's own startup_latencies metric measures from container start to the health check passing. That's not what your user feels. Your user feels container start plus the first request hitting a cold code path — JIT compilation, lazy ORM init, DNS, TLS handshakes to downstream services. We measured end-to-end from the k6 client, not from GCP's metric, because the gap between the two is the whole point.
Results: Node service
For svc-node, cold start end-to-end sat around 1.4–1.8s baseline. Steady-state p50 was ~45ms.
| Config | Cold p99 (ms) | Warm p50 (ms) | Idle cost delta |
|---|---|---|---|
| Baseline | 1600–1800 | 45 | 0 |
| Boost only | 900–1100 | 45 | ~0 |
| Min=1 | 70–120* | 45 | +full idle |
| Min=1 + Boost | 70–120* | 45 | +full idle |
*The cold p99 for min=1 isn't really a cold start — it's the tail of requests that happened to land on a scaled-out instance during a burst. The warm instance absorbed most traffic.
Takeaway: for a lightweight Node service, Boost alone cut cold start latency roughly in half for free. Min-instances helped more, but only for the fraction of traffic that would've hit a cold instance anyway. If your traffic is bursty with long quiet periods, that fraction is high. If you have steady background traffic, min=1 might be solving a problem you don't have.
Results: Java service
This is where things diverge sharply.
| Config | Cold p99 (ms) | Warm p50 (ms) | Idle cost delta |
|---|---|---|---|
| Baseline | 7000–9500 | 60 | 0 |
| Boost only | 3800–5200 | 60 | ~0 |
| Min=1 | 80–150 | 60 | +full idle |
| Min=1 + Boost | 80–150 | 60 | +full idle |
Spring Boot on cold Cloud Run is painful. Boost nearly halved it, which is meaningful, but a 4-second cold start still fails most SLOs we'd write. For JVM workloads, min-instances is less of an optimisation and more of a requirement if you care about tail latency. Boost on top of min-instances is cheap insurance for the moments autoscaling adds a new instance during a burst.
If you must run JVM on Cloud Run and can't justify always-on instances, look at CRaC or AOT compilation (GraalVM native image) before you start tuning these knobs. We've seen Spring Boot native images boot in 200–400ms, which changes the whole conversation.
The cost conversation nobody wants to have
Min-instances isn't free, and the pricing page undersells how it compounds. A single min-instances=1 on a 1 vCPU / 512 MiB instance in a European region runs roughly a few dollars a month per service at idle CPU rates. Multiply by:
- Number of services (easy to hit 20+ in a microservice setup)
- Number of environments (dev, staging, preview, prod)
- Number of regions if you're multi-region
We've seen teams casually enable min=2 across 30 services in three environments and wonder why their Cloud Run bill jumped by four figures a month. The fix wasn't disabling it — the fix was being deliberate: min-instances on the five user-facing services, Boost everywhere else.
A deployment snippet we actually use
# cloud-run service.yaml fragment
spec:
template:
metadata:
annotations:
run.googleapis.com/startup-cpu-boost: 'true'
run.googleapis.com/execution-environment: gen2
autoscaling.knative.dev/minScale: '0' # overridden per-env
autoscaling.knative.dev/maxScale: '50'
spec:
containerConcurrency: 80
containers:
- image: europe-west1-docker.pkg.dev/.../svc-node:${SHA}
resources:
limits:
cpu: '1'
memory: '512Mi'
In Terraform, the equivalent via google_cloud_run_v2_service:
resource "google_cloud_run_v2_service" "svc" {
name = "svc-node"
location = "europe-west1"
template {
scaling {
min_instance_count = var.env == "prod" ? 1 : 0
max_instance_count = 50
}
containers {
image = var.image
resources {
limits = { cpu = "1", memory = "512Mi" }
startup_cpu_boost = true
}
}
}
}
Key detail: startup_cpu_boost is cheap to leave on everywhere. The decision that costs money is min_instance_count.
When each knob is the wrong answer
Boost is the wrong answer when your startup is bottlenecked on I/O, not CPU. If your container spends four seconds waiting on a Secret Manager fetch, a VPC connector attach, or a slow database handshake, extra CPU doesn't help. We've debugged "why isn't Boost working" tickets that turned out to be a sync call to a slow internal API during module load. Fix the init path first.
Min-instances is the wrong answer when your traffic is genuinely steady. If you're serving >1 req/sec consistently, you probably already have warm instances and min=1 is just paying for what autoscaling would've done anyway. Check your instance_count metric before enabling it — if it never dips to zero, min-instances is pure cost.
Both are the wrong answer when the real problem is instance count churn. Cloud Run will scale down aggressively, and if your traffic oscillates around the threshold, you'll see new instances spin up constantly. The fix is often --max-instances tuning and containerConcurrency tuning, not warm pools.
What we'd do
If you're starting from zero on a new Cloud Run service:
- Turn on Startup CPU Boost by default in your Terraform module. It's effectively free and it meaningfully helps cold starts across most runtimes.
- Leave
min-instances=0until you have real latency data. Measure cold start as the client sees it, not asstartup_latenciesreports it. - If your p99 is dominated by cold starts and the service is user-facing, set
min-instances=1on prod only — not staging, not preview. - For JVM services, treat min-instances as mandatory for prod, and seriously evaluate GraalVM native image before scaling the warm pool.
- Review min-instances quarterly against actual traffic. Services that outgrew the need are a quiet source of waste.
If you'd like a second pair of eyes on a Cloud Run setup — cold starts, bill, or both — our cloud and DevOps team does this kind of audit regularly.
Want a team like ours?
72Technologies builds production software for the kind of teams who actually read this blog.
Start a projectKeep reading

Pulumi vs Terraform for a Multi-Account AWS Migration: What Actually Hurt
We migrated a five-account AWS estate from Terraform to Pulumi, then partially back. Here's the honest breakdown of what each tool did well, where they bled us, and which choice we'd make again.
Sentry Performance Quotas Bit Us Mid-Incident: A Rate Limiting Postmortem
During a payment outage, Sentry silently dropped 40% of our transactions right when we needed them most. Here's what tripped the quota, how we found out, and the rate-limiting setup we run now.
Terraform State Locking on S3 Native: What We Learned After Ditching DynamoDB
HashiCorp finally shipped native S3 state locking. We migrated three production stacks off DynamoDB and hit two sharp edges you should know about before you follow.
