Skip to content
All case studies

04 · Reliability · design

Instant Rollbacks

A technical design for rolling back a bad deploy at the routing layer — rewiring which deployment receives traffic instead of rebuilding — so recovery costs a routing update, not a build.

1
routing update per rollback
0
rebuilds

The problem

Rolling back a bad deploy by rebuilding or re-uploading artifacts takes as long as the original deploy did. On a customer's production site that is measured in minutes of a broken page.

The design

Keep the infrastructure for previous deployments standing, and switch which deployment receives traffic by rewiring routing state in Redis and KV. The rollback then costs a routing update instead of a build.

The design covers which deployments are eligible — those that were live, and those explicitly tagged as last-known-good — plus a pre-rollback check, the scenarios where a rollback is already in progress, and CLI parity so it works the same from a terminal as from the UI.

Retention is part of the design

None of this holds if the old deployment no longer exists when you reach for it. So deployment retention, plan-based retention counts, and archived-deployment resource cleanup are part of the same thread — the infrastructure a rollback needs must be kept precisely as long as a rollback is possible, and cleaned up the moment it isn't.