A Docker Image Swap Made a Valid Index Quietly Lie to Production

A search engine's Postgres database swapped one Docker image for another — the same Postgres version, same data, same disk. Nothing about that should have touched a single row. It changed how the database compares strings, and for days, an index that reported itself perfectly healthy was quietly giving wrong answers to live production queries.

The swap, and the reason for it

The database's container had been running postgres:15-alpine, built on musl libc. Installing a graph-database extension needed for a PageRank pipeline required switching to apache/age — a Debian-based image, built on glibc, because the extension has no Alpine build at all. Same major Postgres version on both sides. What's easy to miss is that musl and glibc implement locale-based string collation differently — the actual comparison function Postgres uses to decide whether one string sorts before another, and by extension, whether two strings are equal for indexing purposes.

What collation actually controls

Every text index on disk encodes an assumption about how strings compare to each other, baked in at the moment the index was built. Swap the underlying comparison function without rebuilding those indexes, and the on-disk sort order silently stops matching what the new comparison function would actually produce. This isn't a crash. Postgres doesn't refuse to use a stale index — it keeps serving reads and writes against it exactly as before, because from Postgres's own perspective, nothing about the index looks wrong.

The index that was still "valid" and still wrong

A unique index on a hostname column kept accepting inserts and answering lookups normally throughout. The actual failure showed up when a rebuild was attempted, well after the image swap had been live: could not create unique index "uniq_host_name_ccnew" ... Key (name)=(karriere-jyskrejsebureau.dk) is duplicated. A forced sequential scan against the same hostname — bypassing the index entirely — found two byte-identical rows where the index itself, still marked valid, reported only one.

How a "valid" index actually produces new duplicates

The application's own lookup for "does a host with this name already exist" ran a plain equality search that used the same now-mismatched index. Under the new comparison function, that lookup was capable of missing a row that genuinely existed on disk — the index's internal ordering no longer matched what glibc's comparison would find, so the search path the index provided came back empty when it should have found a match. When that happened, the application did exactly what it was designed to do when a host doesn't exist yet: it created one. The duplicate wasn't a data-entry mistake or a race condition in application code. It was the direct, mechanical consequence of a unique index whose own uniqueness guarantee had quietly stopped meaning what everyone still assumed it meant.

Why this shape of bug is worse than a crash

A crash announces itself. This didn't. Reads kept succeeding, writes kept succeeding, the index kept passing whatever health signal Postgres exposes for "is this index valid" — because validity, from the database's own bookkeeping, was never actually violated; the index's internal structure was self-consistent, just built against a comparison function that no longer matched the one live queries were running under. Every hostname affected by a musl/glibc difference was a live, silent risk for as long as the new image ran, with no error, no warning, and no way to notice short of either hitting the specific duplicate-key failure during an unrelated rebuild, or independently double-checking a supposedly authoritative index against a full table scan.

What this actually requires believing about a "safe" image swap

The reasonable-sounding assumption going in was that swapping a container image for the same underlying database engine, on the same data volume, is a low-risk operation — nothing about the data itself is being touched. That's true for the bytes on disk. It's not true for what those bytes mean once something as fundamental as string comparison changes underneath them. A reindex closes this specific gap, but only if it happens immediately, deliberately, and before anything writes through the mismatch window — treating a base-image swap as equivalent to a plain version upgrade is the actual mistake here, not any single line of code.

Add new comment

Restricted HTML

  • Allowed HTML tags: <a href hreflang> <em> <strong> <cite> <blockquote cite> <code> <ul type> <ol start type> <li> <dl> <dt> <dd> <h2 id> <h3 id> <h4 id> <h5 id> <h6 id>
  • Lines and paragraphs break automatically.
  • Web page addresses and email addresses turn into links automatically.
Please share this article on your favorite website or platform.