← All notes
Product & Platform Engineering7 min

A green migrate is not a rehearsal

Schema changes are where platform work goes quiet and then expensive. The usual gate is "migrate deploy succeeded on staging." That is a parse check with a green checkmark. It does not tell you how long the table lock held under real row counts, whether the expand step left old code reading a column you already stopped writing, or whether the rollback sketched in the PR has ever been run. We treat a migration as rehearsed only when the sequence has been practiced on something shaped like production — data volume, indexes, and concurrent readers included.

What a green migrate actually proved

Most migration tooling answers a narrow question: can this SQL be applied to the database in front of it without erroring. That is useful. It catches a mistyped column name, a missing dependency, a constraint that already exists. It is also the cheapest part of the risk. The expensive part is what happens while the statement runs against a table that already has millions of rows and a dozen live queries touching it.

On BuilderHelp, the system of record, the field app, and the OCR worker all read the same Postgres. A migration that "passed" against a freshly seeded staging database has never contended with a receipt-ingest burst or a dashboard that joins three of the tables you are about to rewrite. Green means the statement was legal. It does not mean the statement was safe under load.

So the first habit is to stop treating migrate status as a shipping gate by itself. Treat it as a prerequisite for the gate — necessary, never sufficient.

Empty databases hide the expensive part

An empty table acquires locks instantly and releases them before anyone notices. The same ALTER on a hot table can hold an ACCESS EXCLUSIVE lock long enough that API workers pile up, the Expo client starts retrying, and the OCR queue looks like an outage. Timing only shows up when the rehearsal data has production-like volume and the indexes you actually ship with.

Shared staging drifts the other direction too. Someone truncated a table to make a demo faster. Someone else never ran last month's backfill. The migration that breezes through that environment can still rewrite half of production because the row counts were fiction. If the rehearsal database is allowed to omit the shape that makes the change expensive, the rehearsal is theater.

We prefer a copy or branch that is boringly close to live: same major version, comparable table sizes on the hot paths, and the indexes that production actually uses. Perfect parity is rare. Named gaps are mandatory. "Staging has 2% of invoice lines" is a fact you can plan around. "It looked fine" is not.

Expand before you contract

The migrations that hurt are usually the ones that try to rename, reshape, or drop in a single deploy while old and new code are both live. Expand-and-contract refuses that compression. Add the new column or table first. Deploy code that writes both and reads the new shape when present. Backfill. Only then remove the old column — after every surface that still knows the old name has been retired.

That sequence feels bureaucratic the first time. It is also the only pattern that keeps every intermediate state runnable. BuilderHelp's release clocks do not move together: the Next.js app can ship on Tuesday, the field app waits on store review, and the OCR worker deploys on its own cadence. A one-shot schema change assumes a single binary cutover those clocks cannot give you.

Rehearsal means running the whole expand → dual-write → backfill → contract sequence on the production-shaped database, not only the final DROP. If you only practice the last step, you practiced the one that deletes options.

Rehearse the reverse too

A migration without a practiced rollback is a one-way door you have not measured. "We can restore from backup" is not a rollback plan for a busy afternoon — restore times are measured in the time the product is wrong, and a restore rewinds more than the bad DDL. Prefer forward-fixable steps: keep the old column populated through the wait period, flip reads with a flag, and treat reverting the app as the first recovery move.

When a step truly cannot be reversed in place, say so in the PR and time the recovery path on the rehearsal database. How long to restore the table. Whether replicas catch up. Which surfaces must be paused. An unrehearsed rollback is a story you tell yourself while the lock is already held.

We write the reverse steps with the same seriousness as the forward ones, then run them once on the branch or clone. Embarrassment in rehearsal is cheaper than discovery in production.

What we gate on before the schema ships

Before a schema change is allowed onto the live path, we want four artifacts, not one green check. The expand-and-contract plan with each deploy named. A rehearsal log against a database whose gaps are written down — lock duration, backfill time, query plan surprises on the hot paths. A rollback that has been executed once, even if the "rollback" is "leave dual-write on and flip the read flag." And a clear statement of which surfaces still speak the old shape, so contract cannot sneak in early.

None of that replaces reading the generated SQL. Auto-migrators are fine drafters and poor reviewers. A human still looks for full-table rewrites, non-concurrent index builds, and constraints that validate existing rows in one transaction.

What this is not: a demand that every nullable column addition get a week of ceremony. Additive, offline-safe changes can stay light. The gate tightens when the change can lock, rewrite, or strand a client that has not updated yet. That is the line between a migrate that passed and a migration that was rehearsed.

Have something to build?

Tell us what you're working on and we'll tell you honestly whether we're the right fit.

Work with us