The cheapest decision that resolves it
Not every request needs the expensive model. Much of routing fits in a regular expression — and the real problem is finding out which, with a suite and a meter, not intuition.
- architecture
- ai
- performance
- testing
Every assistant goes through the same moment: someone types git status and someone types "why did yesterday's deploy fail". Both messages enter the same queue, consume the same model, and cost roughly the same. One of them needed reasoning. The other needed an if.
The design mistake isn't using a large model. It's using a large model to decide what to do, when the decision is already known before the first word.
A rule is a pure function and has no cost
The cheap path is a rule table that runs before any model call: if the message matches a known case, the route is resolved and the model isn't even woken up.
type Rota = { decidido: boolean; acao: string; confianca: number; regra: string | null };
// ordem importa: a primeira regra que casa vence
const REGRAS = [
{ nome: "terminal", rx: /\b(execut\w+|rodar?|docker|container|comando|terminal)\b/i, acao: "terminal", confianca: 0.95 },
{ nome: "git", rx: /\bgit\s+(status|log|diff|commit|push|pull)\b/i, acao: "git", confianca: 0.90 },
{ nome: "web", rx: /\b(pesquis\w+|busque|vers[ãa]o mais recente)\b/i, acao: "web", confianca: 0.90 },
{ nome: "email", rx: /\b(e-?mail|caixa de entrada|inbox)\b/i, acao: "email", confianca: 0.90 },
{ nome: "conhecimento", rx: /\bgrafo\b|\bonde (é|e) usad|noss[oa] proje/i, acao: "grafo", confianca: 0.85 },
{ nome: "trivial", rx: /^((oi|ol[áa]|valeu|obrigad\w+|ok|beleza|bom dia|boa tarde)\W*)+$/i, acao: "nenhuma", confianca: 0.90 },
];
export function rotear(msg: string): Rota {
const texto = (msg ?? "").trim();
if (!texto) return { decidido: false, acao: "modelo", confianca: 0, regra: null };
for (const r of REGRAS) {
if (r.rx.test(texto)) return { decidido: true, acao: r.acao, confianca: r.confianca, regra: r.nome };
}
return { decidido: false, acao: "modelo", confianca: 0, regra: null };
}
rotear is a pure function: same input, same output, no network, no per-call cost, and testable in milliseconds. What it is not: intelligent. And that's where most implementations get lost — the rules are written by intuition, nobody measures coverage, and the deterministic layer becomes architecture decoration.
The rule needs a suite, not an opinion
The right question isn't "does this rule look good?". It's "how many messages in this suite does it resolve — and in how many does it resolve wrong?".
// bateria: mensagem -> ação esperada ("modelo" = deve escalar)
const BATERIA: [string, string][] = [
["ok, pode fazer o deploy", "modelo"], // "ok" no começo não faz disso um cumprimento
["boa tarde!", "nenhuma"],
["executa docker ps e me mostra a saída", "terminal"],
["qual a versão mais recente do node?", "web"],
["git status", "git"],
["tem e-mail novo na caixa de entrada?", "email"],
["roda o build", "terminal"],
["valeu!", "nenhuma"],
["me explica esse erro de build", "modelo"], // falar de build não é executar build
["onde a gente usa o grafo do projeto?", "grafo"],
["resume esse log", "modelo"],
["cria um post novo no site", "modelo"],
["faz um backup do banco", "modelo"],
["ajusta o css do header", "modelo"],
["lista os arquivos do workspace", "modelo"],
["pesquisa na web o preço desse domínio", "web"],
["por que o deploy falhou ontem?", "modelo"],
["escreve um teste pra essa função", "modelo"],
["o que ficou pendente hoje?", "modelo"],
["como funciona o cache dessa API?", "modelo"],
];
let decididos = 0;
const erros: string[] = [];
for (const [msg, esperado] of BATERIA) {
const r = rotear(msg);
if (r.decidido) decididos++;
if (r.acao !== esperado) erros.push(`${msg} -> ${r.acao}, esperado ${esperado}`);
}
console.log(`${decididos}/${BATERIA.length} resolvidos por regra, ${erros.length} divergência(s)`);
In this suite of 20 messages, the result is 9 resolved by rule (45%) and zero divergences. Twenty lines of table bought for twenty-five lines of code — honestly, a good deal.
The number that matters most, though, is the other one: the messages that should go to the model and were captured by mistake. The first version of this terminal rule included the word build in the middle of the pattern. The suite's result: "me explica esse erro de build" was classified as a command execution and sent a request for an explanation to the shell. One word, one false positive, found in four lines of loop.
In the version that actually runs here, the suite closed at 40% — and the hole was in the conjugation: the rule asked for rodar, the verb written was roda. \brodar\b doesn't match "roda o build". A deterministic rule doesn't fail with a bang; it fails with the verb in the wrong tense, and the request falls silently onto the expensive path.
What coverage hides
High coverage is easy to fake: a /.*/ rule resolves 100% of messages and is worth zero. That's why the measurement needs three separate numbers, never one:
| Measure | What it answers | Warning sign |
|---|---|---|
| decision rate | how many messages the rule resolved on its own | too high usually means a rule that's too broad |
| error rate | how many it resolved wrong | any value above zero, on a route with effect |
| escalation rate | how many went to the model | it's the cost that's left — and the legitimate target for optimization |
A decision rate without an error rate is the number that fools a meeting.
Three things that make a difference once the rule exists
- Trace of the decision, not just the result. Logging
regra, reason, and confidence — not only the chosen route — is what lets you answer, three months later, why that message ended up in the wrong place. A result log answers "what happened"; a decision log answers "why". - Turn it off without deploying. The deterministic layer needs a switch that fits in a file, with one-step rollback. If turning it off requires a release window, nobody turns it off — not even when they should.
- A suite with negatives, not just positives. The messages the rule cannot capture are worth more than the obvious ones. All the false positives in this text showed up because of them.
The trade-off, no frills
A deterministic layer doesn't get smarter on its own. It ages: each rule is a bet on how people write, and people write differently every quarter. It's language-sensitive (whoever writes in Portuguese conjugates, whoever writes in English nominalizes), the rule order becomes semantics — the first match wins, so a broad rule swallows a narrow one — and the gain depends on a meter someone has to keep running.
The other side is just as clear: what the rule resolves doesn't depend on API availability, doesn't vary between runs, costs microseconds, and is debuggable with a console.log. It's not replacing the model; it's stopping asking it things that were already decided.
If the deterministic layer has no suite, it isn't architecture — it's intuition with a nice name. And intuition, that one, is expensive.