A 503 that could not happen
We shipped a change that signs our render links with a secret held in the
server environment instead of a value stored in the database. It has a
guard: if the secret is missing, refuse to issue a link rather than fall
back to something weaker. The guard returns 503.
Then production returned a 503 for a request that should never
have reached that guard — and the body attached to it belonged to a
different response entirely.
POST /v1/render-link (invalid API key)
HTTP/2 503
{"error":"unauthorized","message":"Invalid API key"}
A 503 status carrying a 401 body. That combination is not supposed to be constructible. This is how it happened, including the four attempts to reproduce it that all failed, and the Fastify behaviour underneath.
Reading the mismatch
The body is unambiguous: "Invalid API key" is the string our
authentication middleware sends. Our global error handler has its own
wording for a missing key — "Missing API key. Send it in the
x-api-key header." — so there was no doubt which code produced
the body. Authentication ran, rejected the request, and replied.
The status was the problem. reply.code(503) appears exactly once
in our entire codebase: inside the route handler, in the guard described
above. For the status to be 503, that line had to execute. Which meant the
handler ran after the middleware had already sent a reply.
Fastify does not do that. Replying from a hook halts the chain. It is in the documentation, in a sentence that leaves no room:
Replying from a hook implies that the hook chain is stopped and the rest of the hooks and handlers are not executed. [...] If the hook is
async,reply.send()must be called before the function returns or the promise resolves, otherwise, the request will proceed.
Our middleware calls reply.send() synchronously, well before the
function returns. By the documentation, the chain stops.
A better probe than the first one
Before chasing the framework, we wanted a cheaper test of the claim "the
handler ran". The handler's first statement is a validation check
that returns 400 when the request body has no url and no
html, several lines before the 503 guard. So: send an invalid
key with an empty body. If the handler is not running, the response is the
middleware's 401. If it is running, the status becomes 400.
invalid key + {"url": "..."} -> 503 (the signing guard)
invalid key + {} -> 400 (the body check, earlier in the handler)
Both carried the 401 body. The status tracked exactly how far into the handler execution got. The handler was running, and the mismatch was it overwriting the status of a response already on its way out.
Four reproductions that failed
Knowing what was happening did not explain why, and nothing is fixable until it reproduces. Four attempts, in increasing fidelity:
- A synthetic app — an async hook that sends and returns, a handler that sets a different status. Handler skipped. Status 401. Correct.
- The real middleware and the real route, with the database and Redis mocked. Correct.
- Over a real socket rather than Fastify's
inject(), in case the in-process path differed. Correct. - The compiled output — production runs
dist/fromtsc, not TypeScript through a test runner — against a real Redis and a real SQLite file. Still correct.
Four environments, none reproducing a thing that happened on every single request in production. At that point the useful question stops being "why is the framework misbehaving" and becomes "what is still different about my harness".
The difference
Every probe had registered the route in isolation. Production registers it
inside an application that also sets up CORS, Swagger, metrics and logging.
So we ran the actual compiled application — the real
dist/index.js, nothing stubbed but the network it talks to.
bogus key + no url -> status=400 body={"error":"unauthorized",...}
Reproduced. From there it was a bisect. Not CORS. Not Swagger. Not the
metrics hooks. The reproduction needed the request-id hook and the
5xx alerting hook — and the only thing those two have in common is that
both are onSend.
Two hooks
Stripped of our code entirely, it is this:
import Fastify from 'fastify'
async function run (onSendCount) {
const app = Fastify()
let handlerRan = false
for (let i = 0; i < onSendCount; i++) {
app.addHook('onSend', async (req, reply, payload) => payload)
}
app.get('/', {
preHandler: async (req, reply) => {
reply.code(401).send({ error: 'unauthorized' })
},
handler: async (req, reply) => {
handlerRan = true
return reply.code(400).send({ error: 'handler ran' })
}
})
const res = await app.inject({ method: 'GET', url: '/' })
await app.close()
console.log(`onSend=${onSendCount} handlerRan=${handlerRan} status=${res.statusCode}`)
}
for (const n of [0, 1, 2, 3]) await run(n)
onSend hooks | Handler skipped? | Status |
|---|---|---|
| 0 | yes | 401 — correct |
| 1 | yes | 401 — correct |
| 2 | no | 400 |
| 3 | no | 400 |
Fastify 5.12.5, Node 22. Adding a second onSend hook changes
whether replying from a hook halts the request.
Narrowing further, both halves matter. It needs the replying hook to be
async and at least two async
onSend hooks. Callback-style onSend hooks do not
trigger it; neither does the same number of onRequest or
preValidation hooks. And it is not specific to
preHandler — with two async onSend hooks
registered, replying from onRequest, preParsing,
preValidation or preHandler all let the handler run.
What it meant for us, stated plainly
Our application registers exactly two onSend hooks. So every
authenticated route was running its handler for requests that authentication
had already rejected.
The impact was bounded, and the bound was luck rather than design. Handlers dereference the authenticated key within a few lines and throw when it is missing, and the resulting error reply is discarded because a response was already sent. Nothing was authorised, no capture ran, no quota moved, and no client ever received a success body it should not have. But any statement before that first dereference did execute on rejected requests, and that is not a property anyone should be relying on.
It was invisible for the same reason. The body a client receives is always the correct rejection. Only the status changes, and only on a handler that sets one before touching the key. Exactly one of our routes does that. Our other endpoints either reject at schema validation before any handler runs, or throw before setting a status. It took a missing environment variable, producing a conspicuous 503 on precisely that route, to expose any of it.
The fix is one line per rejection
Return the reply instead of sending and returning nothing:
// before
reply.code(401).send({ error: 'unauthorized' })
return
// after
return reply.code(401).send({ error: 'unauthorized' })
This is Fastify's documented idiom, it costs nothing, and it halts correctly
regardless of how many onSend hooks exist.
One place deliberately did not get the change. Our rate limiters are
helpers called from inside handlers that return a boolean, and
their callers branch on it:
if (!(await ipRateLimit(reply, opts))) return reply
Returning the reply object from those is truthy, which inverts the test and
lets rate-limited requests straight through. The first pass of the fix did
exactly that, mechanically, across every reply.code(...).send(...)
in the middleware directory. It was caught before committing, and the code now
carries a comment explaining why those two functions are different. A fix
applied by pattern-match is still a change you have to read.
The test registers two onSend hooks on purpose
The regression test builds an application with both hooks, because without them it passes against the broken code and proves nothing. We checked that directly: reverting a single rejection branch makes it fail, and restoring the fix makes it pass. A test you have not watched fail is a test you do not know the meaning of.
Reported upstream
Filed as fastify/fastify#7090 with the minimal reproduction, the narrowing matrix and the list of affected hooks. As of publication it is open and has not been triaged by a maintainer, so treat the diagnosis here as ours rather than as confirmed framework behaviour — the reproduction is the part that stands on its own, and you can run it in under a minute.
It may well be resolved as documentation rather than code: the docs do say to
return reply, and an argument that sending without returning is
simply unsupported would be a fair one. That would not change the practical
advice. What is worth resolving either way is the coupling — a hook added
in one file changes the halting semantics of hooks in every other file,
silently, with no error and no warning. A rule you can break for months without
noticing is a rule worth enforcing or documenting unconditionally.
Screenshot, PDF and video capture API
100 free captures a month, no card. The engineering above is the kind we write up rather than hide.
Get a free API keyThree things worth keeping
A status code that disagrees with its body is a lifecycle problem. Nothing else produces that shape. Two pieces of code both believed they owned the response, and reading the body told us which one actually sent it.
"Cannot reproduce" is a statement about your harness. Four increasingly faithful environments all said the bug did not exist while it was happening on every production request. The move that worked was not thinking harder about the framework; it was enumerating what the harness still did not have, and running the real application.
Documented behaviour is still worth verifying. The documentation says replying from a hook stops the chain. It does — under a condition the documentation does not mention, which nothing in our code could have told us about.