Three of our lead magnets were reported as broken. We opened all three in a
browser. All three loaded, in about 80 milliseconds, with the capture form on
them.
The report was right anyway. Here is the whole bug:
$ curl -s -o /dev/null -w '%{http_code}' https://example.com/resources/some-guide
200
$ curl -sI https://example.com/resources/some-guide
HTTP/2 405
Same URL, same second, same server. GET says 200. HEAD says 405 Method
Not Allowed. And no browser will ever tell you, because browsers do not send
HEAD requests.
Who asks with HEAD
Almost everything that is not a person.
- Link checkers and SEO crawlers.
HEADis how you check 10,000 links
without downloading 10,000 pages. Many tools record any non-2xx as a dead
link, and some report it as a 404 because that is the bucket their summary
has. - Uptime monitors. A
HEADwith aContent-Lengthis the cheapest way to
tell a live page from an error stub. - Unfurlers. Paste a URL into a chat tool and something resolves it before
rendering the card. - CDNs, when revalidating a cached object.
- AI agents, increasingly. An agent checking whether a citation still
resolves does not want the body. It wants the status line.
None of those is your highest-value visitor. Collectively they are most of what
decides whether a human ever sees the page: whether the link survives in a
directory, whether the monitor pages someone, whether the preview renders,
whether the crawler spends budget on the rest of the site.
The cause is one line of framework asymmetry
If you serve pages with FastAPI, you are probably affected right now.
Starlette's Route adds HEAD to the method set whenever GET is present.
FastAPI's APIRoute — what @app.get and @router.get build — does not. So
every page declared the ordinary way answers 405 to HEAD, while anything
served by StaticFiles answers 200.
That asymmetry is also why this survives for years. The two routes a developer
reaches for when sanity-checking curl -I are usually /robots.txt and
/llms.txt, and both are static. They work. Every HTML page on the site does
not.
RFC 9110 §9.1 is not ambiguous about who is wrong: "All general-purpose servers
MUST support the methods GET and HEAD."
The fix, and the way not to do it
The tempting fix is a HEAD handler per route, or a second decorator on the
pages you remember. Don't. You now have two implementations of every header
policy on the site — content type, cache control, security headers, canonical
link — and the second one drifts the first time somebody changes the first.
A HEAD response is defined as the GET response without the body. So
implement exactly that, once, as middleware: rewrite the method to GET, run
the real request through the real router, and withhold the body bytes.
class HeadRequestMiddleware:
def __init__(self, app):
self.app = app
async def __call__(self, scope, receive, send):
if scope.get("type") != "http" or scope.get("method") != "HEAD":
return await self.app(scope, receive, send)
inner = dict(scope)
inner["method"] = "GET"
sent_body = False
async def send_without_body(message):
nonlocal sent_body
if message["type"] == "http.response.body":
if message.get("more_body"):
return
if not sent_body:
sent_body = True
await send({"type": "http.response.body", "body": b"", "more_body": False})
return
await send(message)
await self.app(inner, receive, send_without_body)
Four details that matter more than they look:
- Register it outermost. Everything inside — routing, your edge guard, the
session layer, the header middleware — has to see theGET, or the headers
on aHEADare not the headers theGETwould have sent. - Keep
Content-Length. RFC 9110 §8.6 permits it on aHEAD, and it is
half the reason a monitor asked withHEADinstead ofGET. - Don't invent a 2xx. Because the inner request is a real
GETthrough the
real router, a POST-only endpoint still answers405and a missing page
still answers404. That is correct, not a gap. - Pure ASGI, not
BaseHTTPMiddleware. The outermost layer is the last
place you want to buffer every response body through a task group in order to
throw it away.
The part that generalises
We did not find this by reading the code. We found it because an outside report
said three pages were broken, we could not reproduce it, and instead of closing
the report we went looking for what else could make a healthy page look dead.
A monitoring claim you cannot reproduce is not automatically wrong. Quite often
it is a real signal measured with a method you were not testing — and on the
public web, the methods you are not testing are the ones your customers never
use and your infrastructure uses constantly.
If you run a client-facing site, spend the next minute on this:
for path in / /pricing /resources/your-lead-magnet; do
printf '%s %s\n' "$(curl -s -o /dev/null -w '%{http_code}' -I "https://yoursite.com$path")" "$path"
done
If any of those is not what GET returns, the gap between what your visitors
see and what the internet sees is wider than you think.