release-engineering · 2026-09-14 · 11 min read
The tag had already moved, and a sentence I could only rewrite after a rollback
Why this exists
The last article ended on a gap. A backup takes you back to yesterday's data. If an upgrade breaks something, what you usually want is yesterday's software with today's data, which is a different operation. Engineers call the first a restore and the second a rollback.
The backup documentation already told an operator which one to reach for. And then it said, honestly, that a bit-exact rollback "is not available today". My proposal put the problem in one line: an operator at two in the morning is told "this is a rollback, not a restore", and then has nothing to roll back to.
So I made that sentence the finish line. It could only be rewritten once a rollback had actually been performed, end to end, and a test guards it either way. Before starting, I measured how far away that was.
The three Python project files declared 45 dependencies, every one as a floor ("this version or newer") and not one pinned. The cryptography library, the one that decrypts the stored mail passwords, was declared as version 42 or newer. Version 49 was installed. Six of the eight third-party container images were named by a label that can move. None was pinned to exact contents.
Put plainly: if the customer's server had to be rebuilt next month, nobody could say what software it would end up running. Not me, not the customer, not the build.
Step one: write down exact versions, and check them
A dependency is someone else's code that yours uses. A floor like
django>=5.1 is a promise to accept any newer version, including ones that
don't exist yet. A lockfile replaces that with an exact shopping list, every
package at one version, and ideally with a fingerprint for each so a swapped
package is refused at install time.
The naive way to make one is a single command, and I ran it. Then I compared its result with what was actually installed, the versions all our tests had been passing against. Forty of 110 packages had moved. Django went from 5.2.16 to 6.1.1. Cryptography went from 49 to 50. All of it perfectly legal, since the floors allowed it. That commit would have shipped a major upgrade of the web framework, and of the library holding the keys, inside a change whose message said "lock dependencies".
So I seeded the lock from the installed set of packages instead, with temporary
constraints, and got zero differences across all 110. The images now install
with --require-hashes, which refuses any package whose fingerprint isn't in
the lock. I checked a built image afterwards. Ninety-two packages, zero
differences.
Step two: stop naming images by labels that move
This step is easiest to see as a picture, so here it is first.
A container image is a packaged program. Its tag, like redis:7-alpine, is
a label that the publisher can move to newer contents whenever they like, and
they do. A digest is a fingerprint computed from the exact bytes. It can't
point at anything else, ever.
I pinned all eleven third-party images across three registries by digest, with
the human-readable tag kept beside it. And the first pin proved the point of
the whole exercise. The Redis tag had already moved. The registry now served
different bytes under redis:7-alpine than the image the installation had been
running for weeks. Nothing had said so, and nothing would have, until a rebuild
quietly brought in the new one.
Two details from that afternoon. The pinning tool refuses digests that only
cover one processor type, because pinning my laptop's ARM image would hand the
customer's Intel server a program it can't run. And my first version of the
tool had two bugs of its own: it read redis:7-alpine as a host called redis
on a port called 7-alpine, and a naive text replacement appended the digest
again on every run. @sha256:…@sha256:…@sha256:…. Docker Compose accepted that
without complaint, which I still find a little alarming.
Step three: a machine runs the checks, not whoever remembers
Continuous integration, CI, is a fresh rented computer that runs every check on every change. The value isn't the automation. It's that nobody ever worked on that computer, so it can't inherit the quiet assumptions of the one you did work on.
The first run went red and found two such assumptions within minutes. One test only passed if a settings file existed, and every developer's checkout had one while a fresh checkout doesn't. Another relied on a container image already being downloaded, which it had been on my machine since August. Then my own fix for the first went red too. I'd copied the example settings file, which turns secure cookies off, so three cookie tests began checking the example file instead of the code's real defaults. The fix was an empty file.
None of this was visible before there was a machine that hadn't been used to develop on. Two smaller rot findings came along. A tool two tests needed had never been declared, because I'd installed it by hand. And a file count quoted in five places in the docs was wrong in all five. It now lives in one file that a script checks.
Step four: a release is a box you can carry in
A release is now a sealed box, named after the exact commit it was built from. Inside is every program image the system runs, saved to one file, plus a label for machines, the same label for humans, and a packing list. The build script refuses to make one from uncommitted code, from a commit whose checks didn't pass, or when it can't obtain the images. An early version had a flag to skip the check. I deleted it instead of documenting it.
At the customer, loading a box needs no internet at all. I proved that by deleting two images from Docker, loading the archive, and bringing up all ten services with network pulls forbidden. They all came up healthy.
The first real release failed, instructively. The script guessed the image names from the folder name, and the project's own configuration declared a different one, so all five of our own images were "could not obtain". The fix was to stop guessing and ask Docker Compose. Then the script aborted after building everything, the most expensive possible moment, because the old bash on macOS treats an empty list as an undefined variable.
Two more things only showed up by looking at the finished box. The manifest
recorded a checksum of the archive, and the human label printed it under the
words "the checksum above proves the bytes are unchanged". Nothing ever read
it. Now verify_release recomputes it, and I tested that by writing 2 kB of
random bytes into the middle of the 1.87 GB file without changing its length.
It named both fingerprints and refused.
And the packing list (engineers say SBOM, a software bill of materials) is made by scanning the finished image, not by copying the lockfile. That difference is the whole point. The scan of the worker image found 108 Python packages, 106 operating-system packages and 15 standalone programs. Among them was the PostgreSQL client the product needs to restore its own backups. It is in the image and in no Python lockfile, so a list built from the lock would never mention it. The build then cross-checks the scan against the lock and refuses if they disagree on any version.
Step five: the rollback, performed
The last missing piece was a way to ask "can I go back to that build?" before
doing it. rollback_check --to <build> answers from two sides. Is there a
database change since then that can't be undone? (Every migration now declares
itself reversible, data-only, or irreversible with a reason. A person has to
say which, because code that looks reversible can still throw information
away.) And is that build's box still here, complete?
Two releases, then, one commit apart, where the newer one added a reversible database column. I brought up a stack from the newer box with network pulls and local builds forbidden, gave it a mail account with an encrypted password, some job history and a backup set, and ran the check. It named the one migration, its classification, found the older box, and said "rollback". I reversed the migration and loaded the older images.
Afterwards, every component reported the older build. The column was gone. Every row was intact, the password still decrypted, and a normal login worked.
It was the second attempt. The first one had quietly tested nothing. Docker Compose, given a project name, had built its own images from my working copy and picked up a stale build id, so the stack I "rolled back" was neither the archive nor either release. And a fresh database stamps all of its migrations within the same second, so the check reported four irreversible changes between two releases one commit apart. It was right, too, about a database no real installation will ever have.
Then I rewrote the sentence. The documentation now says a bit-exact rollback is available, and it reaches exactly as far back as the releases that were kept. The test that guarded the old sentence wasn't deleted. It changed sides.
What it does not promise
Nothing is signed yet, so the checksum proves the bytes didn't change, not who made them. Keeping old boxes is the customer's job, and nothing watches how many are left. I recommend three, about 5.6 GB, because one is the version already running and two isn't enough when a broken upgrade is only noticed after the next one replaced it. And the exercise ran on one machine, without the chat service, and never crossed an irreversible migration. Those are written down where the operator will read them, in the same document.
Same day, the chat service moved into the repository as a single squashed folder. Its 5,403 commits of history became one, because only 44 of them were ours, and it's now built from that folder into its own pinned image. Which matters for the next article but one, where the chat learns to read the firm's knowledge.
Next: a mailbox that belongs to the firm instead of to a person, and a refactor that changed no SQL and still made the firm's letters say "du".
Written with AI from my own repositories and notes, reviewed and published by me. How this site is written