Terraform State Locking on S3 Native: What We Learned After Ditching DynamoDB
HashiCorp finally shipped native S3 state locking. We migrated three production stacks off DynamoDB and hit two sharp edges you should know about before you follow.
Since Terraform 1.11, you can drop the DynamoDB lock table entirely and let S3 handle state locking natively via conditional writes. We migrated three production stacks over the last quarter — a Vercel-fronted API, an internal data platform, and a customer-facing e-commerce backend. The feature works, but the migration path has two sharp edges the docs gloss over.
Here's what we ran into, what we'd do differently, and when we'd still keep DynamoDB around.
Why anyone cared about this in the first place
For about eight years, the canonical Terraform-on-AWS setup was: state file in S3, lock in a DynamoDB table with a LockID primary key. It worked. It was also two resources to bootstrap, two IAM policies to write, and one more line item on the bill.
More annoyingly, it was two resources you had to create before you could use Terraform to manage anything — the classic chicken-and-egg bootstrap problem. Most teams solved it with a hand-rolled CloudFormation stack or a terraform init against local state, then a migration. Neither is elegant.
When AWS added conditional writes to S3 in late 2024 (If-None-Match on PUT), HashiCorp had the primitive they needed. Terraform 1.11 shipped the use_lockfile = true option on the S3 backend, and the DynamoDB table became optional.
On paper this is pure upside: fewer resources, cheaper bootstrap, one less IAM surface. In practice, we hit friction.
The migration, in the shape we wish someone had shown us
The backend block change itself is trivial:
terraform {
backend "s3" {
bucket = "acme-tf-state-prod"
key = "platform/api/terraform.tfstate"
region = "us-east-1"
encrypt = true
use_lockfile = true
# dynamodb_table = "terraform-locks" # removed
}
}
Run terraform init -migrate-state (even though the state itself isn't moving, the backend config changed), confirm, and you're done. On a clean stack this took us under a minute.
The problem is that most real stacks are not clean.
Gotcha 1: Mixed-version teams will corrupt your locking assumptions
Our data platform stack is touched by six engineers and two CI pipelines. When we flipped use_lockfile = true, one CI runner was still pinned to Terraform 1.9 via an old container image. That runner happily wrote to the state file using the DynamoDB lock — which was still there — while a laptop on 1.11 acquired the S3 lockfile and did the same.
No collision fired. Both operations completed. We ended up with a state file that reflected the 1.11 apply, and a set of resources that reflected the 1.9 apply. Reconciling took an afternoon and one panicked terraform import.
The fix is dull but mandatory: before you flip the flag, pin every runner and every developer machine to a Terraform version that supports use_lockfile, and verify with a required_version constraint.
terraform {
required_version = ">= 1.11.0"
}
We also added a pre-commit hook that greps for the DynamoDB lock table name in any backend config still floating around in feature branches. Belt and braces.
Gotcha 2: Stale lockfiles are harder to clear than stale lock rows
With DynamoDB, a stuck lock is a single-item delete. Any engineer with console access can nuke it in ten seconds. With the S3 lockfile, the lock is an object at <key>.tflock sitting next to your state.
That sounds equivalent. It isn't, for two reasons:
- Object versioning. If your state bucket has versioning enabled (it should), deleting the lockfile creates a delete marker but leaves prior versions.
terraform force-unlockhandles this correctly, but manual cleanups viaaws s3 rmcan leave a version that a badly-behaved client interprets as still-held. - Eventual consistency on listings. S3 is strongly consistent for reads after writes on a single key, but bucket listings can still lag briefly. We saw one case where a CI job listed the prefix, didn't see a lockfile, tried to PUT with
If-None-Match: *, and got a 412 back because a competing job had just written one. That's actually the correct behaviour — the lock worked — but the error message Terraform surfaces (ConditionalRequestConflict) is less obvious than the oldConditionalCheckFailedExceptionfrom DynamoDB.
Our runbook now includes a canned command for lock inspection:
aws s3api list-object-versions \
--bucket acme-tf-state-prod \
--prefix platform/api/terraform.tfstate.tflock
And we standardised on terraform force-unlock <LOCK_ID> rather than manual object deletes.
The cost picture is smaller than you think
We'd seen breathless posts claiming DynamoDB lock tables were an expensive holdover. That's not our experience. Across roughly 40 stacks in our largest AWS org, the DynamoDB lock table costs were in the low single-digit dollars per month total — the table sits at PAY_PER_REQUEST with essentially no traffic between applies.
The real savings from use_lockfile are operational, not financial:
- One less resource in your bootstrap module
- One less IAM policy statement per role that runs Terraform
- One less thing to monitor and back up
- No more "why is this DynamoDB table in our compliance inventory" questions from security
If someone tells you the migration will meaningfully cut your AWS bill, they're selling something. Do it for the simplicity.
When we'd still keep DynamoDB
We left DynamoDB in place on one stack, deliberately.
It's a stack shared between our infra team and a partner integrator who runs Terraform Enterprise on a version we don't control. TFE's support for use_lockfile came later than the CLI's, and the partner's upgrade cadence is glacial. Running both lock mechanisms in parallel was tempting, but as gotcha 1 showed, that's a footgun. Sticking with DynamoDB until everyone can move together was the safer call.
Other cases where we'd wait:
- Cross-account state access with complex bucket policies. The conditional-write permissions need
s3:PutObjectwith a condition ons3:if-none-match. If your bucket policy is already a knot of principals and condition keys, adding another condition is worth thinking through before you flip production. - Regulated environments that already have DynamoDB audited. If your compliance evidence packet already covers the lock table, ripping it out means updating the evidence. The juice may not be worth the squeeze until your next audit cycle.
- Very high-frequency apply patterns. We haven't hit this in practice, but if you're running dozens of applies per minute against a single state file — which, honestly, is a design smell — DynamoDB's request semantics are better characterised under contention than S3's conditional PUTs.
Observability we added afterwards
One thing that bit us in the transition month: we had CloudWatch alarms on the DynamoDB lock table's throttled requests. Those alarms had caught two runaway CI loops in the previous year. When the table went away, so did the alarm.
We replaced them with two S3-side signals:
- A CloudWatch metric filter on S3 server access logs, counting
PUTrequests to*.tflockthat return412 Precondition Failed. A spike means concurrent apply contention, which usually means a stuck CI job or a human racing a pipeline. - A Lambda that runs every 15 minutes, lists lockfiles across all state keys, and pages if any lockfile is older than 30 minutes. Applies shouldn't take that long; if one is, something is wedged.
Both are cheap. The Lambda is maybe 40 lines. If you're moving stacks over, budget an afternoon to port your lock-related alerting rather than assuming the new mechanism is self-monitoring.
Where we'd start
If you're planning this migration, do it in this order:
- Pick your smallest, lowest-blast-radius stack. Ours was a sandbox account's networking module.
- Pin
required_version >= 1.11.0and get every runner, container, and laptop on a compatible version first. Verify with a scripted check in CI. - Flip
use_lockfile = true, remove thedynamodb_tableline, runterraform init -migrate-state. - Run a deliberate concurrent apply from two terminals. Confirm the second one fails cleanly with a lock error. If it doesn't, stop and diagnose.
- Port your lock-related alarms before you move the second stack.
- Leave the DynamoDB table in place for two weeks after the last stack migrates, just in case you need to roll back.
We're not evangelists for this feature — DynamoDB locking was fine — but it is genuinely one less moving part, and for greenfield AWS work in 2026 we wouldn't bootstrap a lock table anymore. For an existing estate, migrate deliberately, or don't migrate at all. Both are defensible. What isn't defensible is a half-migrated fleet running two lock mechanisms against the same state.
If you'd like a second pair of eyes on your Terraform bootstrap or backend strategy, our DevOps and cloud team does this kind of work regularly.
Want a team like ours?
72Technologies builds production software for the kind of teams who actually read this blog.
Start a projectKeep reading

GCP Cloud Run Min-Instances vs Startup CPU Boost: Which One Actually Fixes Cold Starts
We ran a Node and a Java service on Cloud Run through both knobs — min-instances and Startup CPU Boost — to see which one actually earns its keep. The answer depends on your runtime, your traffic shape, and how honest you are about idle cost.

Pulumi vs Terraform for a Multi-Account AWS Migration: What Actually Hurt
We migrated a five-account AWS estate from Terraform to Pulumi, then partially back. Here's the honest breakdown of what each tool did well, where they bled us, and which choice we'd make again.
Sentry Performance Quotas Bit Us Mid-Incident: A Rate Limiting Postmortem
During a payment outage, Sentry silently dropped 40% of our transactions right when we needed them most. Here's what tripped the quota, how we found out, and the rate-limiting setup we run now.
