"One to rule them all" — a service rebuild story
Rebuilding a service and migrating millions of NoSQL documents on Google Cloud without stopping production.
In this post
There was a service. It worked, it made money, and it had one flaw: nobody wanted to touch it anymore. The data model had grown through years of “small fixes”, every query had an exception, and deployments were done holding your breath. On top of that: a few million documents in a NoSQL database that could not simply be “rewritten over a weekend”.
This is the story of how V1 became V2 — and how to move millions of documents so that users never notice.
Why not a big bang
The simplest plan: stop traffic, copy the data, flip DNS. Simplest and worst. With millions of documents the copy takes hours, and any surprise in the data (there is always a surprise) means a rollback and a second sleepless night.
We went with three rules instead:
- The old service lives until the end. V1 serves production until V2 proves it returns the same answers.
- Migration is a process, not an event. Data flows in the background, in batches, resumable at any point.
- Every document carries a version. Without a
schemaVersionfield you do not know what has been moved and what is still waiting.
The V2 architecture
V2 got what V1 never had: a single domain model that every read and write goes through. Instead of ten places that “sort of know” the shape of a document, there is one module that translates old documents into new entities.
// upcaster: old document -> current version
export function upcast(doc: RawDoc): DocV3 {
switch (doc.schemaVersion ?? 1) {
case 1: return upcast({ ...fromV1(doc), schemaVersion: 2 });
case 2: return upcast({ ...fromV2(doc), schemaVersion: 3 });
case 3: return doc as DocV3;
default: throw new Error(`Unknown schemaVersion ${doc.schemaVersion}`);
}
}
This pattern (known from event sourcing as upcasting) has one huge advantage: you do not have to migrate everything before you launch. V2 can read any document version. The background migration merely “materialises” the new version so you stop paying for the conversion on every read.
Background migration on Google Cloud
The migration itself is a worker on Cloud Run that:
- fetches a batch of documents by cursor (
_id > lastId, limit 500), - runs them through
upcast(), - writes them to the new collection with
schemaVersion: 3, - stores
lastIdin a tinymigration_statecollection.
Keeping state in the database instead of memory means the worker can die halfway and resume from the last cursor. Pub/Sub triggered the next batches, Cloud Logging showed the pace: documents per minute, error count, version distribution.
The most important lesson of this migration: log documents that fail validation, but do not stop on them. There are always a few hundred records from 2016 in a shape nobody remembers. We parked them in a migration_quarantine collection and fixed them separately.
Dual writes and comparing answers
During the transition V1 and V2 ran side by side:
- writes went to both databases (dual write, V1 as the source of truth),
- reads went to V1, but we also called V2 in the background and compared responses (shadow traffic).
Differences went to the logs. The first days had plenty — mostly element ordering and date formatting. Once the diff counter had read zero for a week, we switched reads to V2. Then writes. Then we turned V1 off.
What I would do differently
- Add the
schemaVersionfield on day one of the project, not right before a migration. - Create the quarantine collection on the first day, not after the first worker crash.
- Start shadow traffic earlier — it is the cheapest integration test there is.
The V2 service is still running today. One model, one ring to rule them all.