Arezgitfield notes / engineering
Release engineeringUPDATED AUG 18, 2026

Plan Backward-Compatible Database Migrations

An expand-migrate-contract workflow for shipping schema changes across mixed application versions while controlling locks, backfills, rollback, and data loss.

AREZGIT / FIELD NOTERELEASE ENGINEERING
A schema change can be syntactically valid and still break production because application versions overlap during deployment. Old instances may write the old shape while new instances expect
READ / VERIFY / APPLYTECHNICALLY REVIEWED

A schema change can be syntactically valid and still break production because application versions overlap during deployment. Old instances may write the old shape while new instances expect the new one. A large backfill may compete with user queries, and an ALTER TABLE can wait for or hold a lock longer than expected.

Plan the migration as a sequence of compatibility states: expand the schema, migrate behavior and data, verify adoption, then contract obsolete structures. “Zero downtime” is an outcome to measure for a particular database and workload, not a property of a migration file.

Draw the version compatibility matrix

List every actor that reads or writes the affected data: current application instances, candidate instances, workers, scheduled jobs, reporting tools, administrative scripts, and downstream consumers. Then describe whether each actor can operate against the old, expanded, and contracted schemas.

The dangerous states become visible:

  • new code starts before the additive schema exists;
  • old code sees a new constraint it cannot satisfy;
  • new code assumes backfilled values before the backfill finishes;
  • old code overwrites a new representation with stale data;
  • contraction occurs while a delayed worker still uses the old column.

Include rollback in the matrix. Rolling application code back means an old binary will meet the current database, not the database state that existed before deployment.

Expand with additive, nullable structures

The expand phase adds the minimum structures needed by the new code while preserving old behavior. Typical moves include adding a nullable column, a new table, or an index built through the database's low-blocking mechanism.

Database details are engine- and version-specific. PostgreSQL's ALTER TABLE documentation and explicit locking reference show why syntax alone is insufficient: different operations acquire different locks, and lock duration depends on concurrent work and the operation performed.

Before production:

  • test against a representative table shape and row count;
  • identify the lock mode and conflicting operations for the exact database version;
  • set a bounded lock wait so the migration fails instead of blocking indefinitely;
  • observe replica or follower effects where applicable;
  • verify disk, transaction log, and temporary-space requirements;
  • rehearse cancellation and retry behavior.

Do not add a non-null constraint with a volatile default in one step without verifying the engine's execution strategy. An additive change can still rewrite a table or block traffic under some databases and versions.

Deploy code that tolerates both representations

New code should operate while old and new data coexist. A common transition is dual-read with a clear precedence rule and dual-write only when necessary.

Dual-write introduces its own failure: one write can succeed while the other fails. Prefer one transaction when both representations share a transactional database. If that is impossible, define idempotent repair and reconciliation rather than assuming the two values remain synchronized.

For reads, distinguish “not migrated yet” from a legitimate null or empty value. A fallback such as new_value ?? old_value is only correct if null is not valid in the new contract. Encode the migration state explicitly when ambiguity matters.

Instrument reads from the fallback path, mismatches between representations, write failures, and rows pending migration. Keep metric labels bounded and logs free of row content that may contain personal or secret data.

Backfill in bounded, restartable batches

A backfill is production workload. It should be resumable, idempotent, rate-limited, and observable.

Choose a stable traversal key and process bounded batches. Commit progress frequently enough to limit lock and recovery cost, but not so frequently that transaction overhead dominates. Store a checkpoint or derive remaining work from an indexed predicate.

The backfill must coexist with live writes. Protect against copying stale old values over newer data. Options include conditional updates, row versions, timestamps with documented semantics, or a write path that updates both structures during the transition.

Measure database latency, lock waits, replication delay, transaction-log growth, failed batches, and reconciliation mismatches. Pause automatically on guardrails. Retrying a batch should not duplicate rows or re-trigger external effects.

Enforce the new contract only after evidence

Once the backfill reports completion, verify independently. Compare counts, sample domain invariants, scan for fallback reads, and run the query that would find unmigrated rows. Completion means the invariant holds now, not merely that a job processed every planned batch once.

Then switch reads to the new representation while keeping compatibility code for a defined observation window. Only after all supported application versions have stopped using the old structure should the schema enforce new constraints or remove old columns.

Constraint validation can itself require scanning or locks. Use the engine's staged validation features when available and verified. Keep the exact operation and its failure behavior in the release plan.

Treat contraction as irreversible until proven otherwise

Dropping a column removes the easiest rollback path and may destroy data. A down migration that recreates an empty column is not data recovery. Backups help only if restore time, point-in-time recovery, and the blast radius fit the incident.

Before contraction, prove:

  • no deployed binary or delayed job references the old structure;
  • fallback-read telemetry is zero for the required window;
  • reconciliation shows equivalent data where equivalence is expected;
  • the retention and backup plan covers the removed information;
  • rollback code can run against the contracted schema;
  • the destructive target is the intended database and schema.

Archive only what policy permits, with access controls and a deletion schedule. Keeping every retired column “just in case” creates privacy and governance risk.

Make each phase independently releasable

Expand, code transition, backfill, read switch, constraint enforcement, and contraction should be separate observable steps whenever the system's risk justifies it. Each step needs entry conditions, success evidence, a bounded failure response, and an owner.

Backward compatibility is achieved when mixed versions can operate during the planned rollout, data transitions are measurable and repairable, and rollback behavior matches the current schema. The migration file is only one artifact in that process.

READER SIGNAL

Was this guide useful?

Your rating helps us prioritize clearer, more practical technical content.

AREZGIT / DESKTOP WORKSPACEFROM FIELD NOTE TO RELEASE

Review diffs, run checks, and prepare the release in Arezgit.

Keep Git review, security scanning, API checks, database inspection, and release preparation together in one local desktop application.

Explore Arezgit