skip to content

day 277 of the operation / 05.10.2026

Proxfrito

← technical

technical14/09/20264 min

Three questions before putting cache on anything

Cache isn't optimization, it's a contract of tolerance for stale data. If you can't write that contract, you don't know what you're putting into production.

  • cache
  • performance
  • backend
  • redis

"Cache is always the answer". It's the most repeated and least useful phrase in software engineering, because the second half of the phrase is always left out: cache is the answer to the problem of reading the same thing many times in a short period. For any other problem it's the wrong answer — including the problem of "the system is slow", which is where it usually gets applied.

There are three questions that separate a cache that works from an incident with a 48-hour delay. If you can't answer all three in writing, don't put the cache in.

1. How much stale data is acceptable?

Every cache serves stale data at some point. That isn't a side effect, it's the function. So the first question isn't "how to invalidate", it's how much.

  • Zero seconds. Then you don't want a cache: you want a synchronized copy, a strongly consistent read, or to accept the database's latency.
  • A few seconds. Typical of a public listing, a counter, a catalog. A Cache-Control: max-age=30 handles it and requires not a single line of infrastructure.
  • Minutes or hours. Good for data that changes by human publication, like marketing copy or configuration.
  • Until the next write. This is the hardest case: it requires explicit invalidation, and explicit invalidation is where almost every cache bug is born.

That number has a name in the rest of the world: staleness budget. If the team can't say whether the requirement is two or thirty seconds, it's because nobody asked the product — and the answer you'll get after the incident is "it can't be stale even for a second".

2. Who invalidates, and when?

This is the question that kills projects. There are three honest answers:

Nobody invalidates: it expires by time. The cache carries a TTL and lives with it. It works well when the data is tolerant. It costs little. It's underused because it seems "unprofessional", which is nonsense: half the world's content systems only need this.

The write path invalidates. The transaction that changes the data also drops the key. This is where most bugs are born, because invalidation is easy to forget on a secondary path:

// easy to miss: a segunda escrita NÃO invalida o cache
await db.produto.update({ where: { id }, data: preco });
await redis.del(`produto:${id}`);

// o importador em lote, esse carinha aqui, alterou 4.000 produtos e não avisou ninguém

The cache has a version. The key includes the data's version (produto:42:v17) and the write promotes a new version. It costs more memory, and in exchange it eliminates the entire category of "I forgot to invalidate". If your data is written by several processes, this pattern pays for the investment in a week.

My opinion, labeled as opinion: if you can't point, in the code, to every place that invalidates the key, choose TTL and increase the tolerance — or choose version. Invalidation spread across human memory is technical debt with compound interest.

3. What happens on the first request after expiry?

This is where the incident everyone knows and few foresee lives. A key expires at 9am. At 9:00:00, five hundred simultaneous requests discover the cache is empty and all go to the database. The database goes down — not from the average load, but from the simultaneous load.

Two defenses solve almost everything:

Local single-flight. Only one request per process goes to the database; the others wait for it:

let emVoo: Promise<Produto> | null = null;

async function comSingleFlight(id: string): Promise<Produto> {
  if (!emVoo) {
    emVoo = buscarProduto(id).finally(() => {
      emVoo = null;
    });
  }
  return emVoo;
}

TTL with jitter. Never expire everything together. Add some noise to the TTL (ttl + random(0, ttl * 0.1)) and the expiry of twenty thousand keys stops being a single event.

const ttl = 300 + Math.floor(Math.random() * 30); // 300–330s

And the obvious complement: when the backend is down, the stale cache is better than nothing. stale-while-revalidate and stale-if-error do that in HTTP for free, without a line of code:

Cache-Control: public, max-age=60, stale-while-revalidate=300, stale-if-error=86400

The summary that fits on a sticky note

  • Cache is a tolerance contract, not an optimization.
  • Without a staleness number agreed with the product, there is no cache — there's a bet.
  • If nobody knows who invalidates, choose TTL or version. Don't choose memory.
  • Measure before and after. A cache nobody measured is a cache nobody can remove.

And if after all that the system is still slow, the problem probably was never repeated reads — it was a query without an index. But that's the subject of the next text.