hupden/ projects
blogtools
← blog

My auth was signing production tokens with an empty key

#go#security#jwt#self-hosting

The admin dashboard for this server sits behind a JWT. You log in, the API signs a token with an HMAC key from the environment, and every admin endpoint verifies it before doing anything. It is the least interesting authentication design available, which is why I chose it.

On the 25th of July, prompted by an entirely unrelated question, I read that code properly for the first time in a while. The token handling was fine. The signing method was pinned. The middleware did the right things in the right order.

Nothing in it required the signing key to exist.

What an empty HMAC key actually does

Here is where my intuition was wrong, and I think it is a common wrong.

I had assumed — never explicitly, which is the problem — that a missing secret would produce some kind of failure. An error from the signing call, a panic, a refusal to verify. Something loud enough that it could not go unnoticed in production.

It does not, and the reason is that HMAC is defined for a key of any length, including zero. There is no special case. A zero-length key is a key. Every layer behaves exactly as designed.

I confirmed this rather than reasoning about it, against the same library and version the service uses:

sign with empty key:   err=<nil>
token: eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9.eyJleHAiOjE3ODkwMDM4OTAsInN1YiI6ImFkbWluIn0...
verify with empty key: err=<nil> valid=true

Signing succeeds. Verification succeeds. The token is a completely valid HS256 token that anyone in the world can produce, because the secret required to produce it is the empty string.

So the failure mode is not an outage. It is a service that logs in normally, issues normal-looking tokens, rejects garbage in the Authorization header, returns 401 to unauthenticated requests, and passes every health check — while its entire authentication scheme is decorative. There is no observable difference between this and a correctly configured deployment, from inside or outside, unless you specifically think to test whether a token you forged yourself is accepted.

The defense I had was for the attack everyone writes about

Worth dwelling on, because it is the actually instructive part.

My verification callback pins the algorithm:

token, err := jwt.Parse(tokenStr, func(t *jwt.Token) (interface{}, error) {
	if _, ok := t.Method.(*jwt.SigningMethodHMAC); !ok {
		return nil, jwt.ErrSignatureInvalid
	}
	return jwtSecret, nil
})

That check is there deliberately, and it is correct. It blocks alg:none — the famous JWT vulnerability, the one in every write-up — and it blocks RS/HS confusion, where an attacker re-signs a token using your public key as an HMAC secret. I tested that too, at the same time:

alg:none through the HMAC-pinned parser: err=token is unverifiable:
  error while executing keyfunc: signature is invalid

Refused, exactly as intended.

So the well-known hole was closed, by code written specifically to close it, and it worked. The unknown one sat one line below it, in the value being returned rather than the check being performed. I had thought carefully about which algorithm the token claimed, and not at all about whether there was a key.

I have stopped treating "I've handled the classic attack on this" as evidence of anything about the surrounding code. If anything it is mildly negative evidence about attention: the effort went where the literature pointed.

The guard, added for a hypothetical

The fix is four lines, and I want to be precise about my state of mind while writing them, because it turned out to matter.

// Fail closed on a missing/weak signing secret. Without this guard an
// empty JWT_SECRET still boots and both signs and verifies tokens with
// an empty HMAC key — i.e. anyone can forge a valid admin token.
if len(os.Getenv("JWT_SECRET")) < 32 {
	log.Fatal("JWT_SECRET must be set and at least 32 bytes")
}

I was not fixing a known problem. I was closing a future one — making sure that if some later deploy forgot to wire the secret, the service would refuse to start instead of degrading into a fully bypassable auth system that still looked healthy. Thirty-two bytes because that is HS256's output size; shorter secrets are rejected outright rather than warned about.

In my head this was hygiene. Defensive, slightly theoretical, the sort of change you make because it is obviously right rather than because anything is wrong.

Then I deployed it and the service refused to boot.

It was not hypothetical

The secret had never been set. Not misspelled, not stale, not truncated — absent, since the day the service first shipped.

The chain is worth spelling out because every link in it is behaving correctly. The deploy workflow references a repository secret and passes it into the container as an environment variable. The secret store had no entry under that name. A missing secret expands to the empty string, not to an error. So the container was started with the variable set to "", which is different from unset in exactly no way that mattered, and the service came up clean.

Nothing in that path is designed to complain. CI does not know which of its secrets are load-bearing. docker run does not know an empty environment variable is meaningful. The application, before the guard, was the only component in a position to notice, and it had no opinion.

So for the entire life of that service in production, every admin token had been signed and verified against an empty key. Anyone who guessed that could have minted themselves a valid admin session from a laptop, with no interaction with my server beyond presenting the token.

What the exposure actually was

I want to be accurate about this rather than dramatic, in both directions.

It was live, it was real, and it was not theoretical for a single day of that service's existence. That is the honest headline and I am not going to soften it.

It was also bounded, for reasons that were mostly luck at the time and are now deliberate policy:

The remediation was quick: generate a real secret with openssl rand, wire it into the secret store, redeploy. A side effect worth noting is that rotating the key invalidated every token that had ever been issued under the empty one, so there was no cleanup step for anything already minted.

That "no write routes" line is the one that changed how I build. It was an accident that limited the blast radius of a live auth bypass, and accidents that save you are worth converting into rules. It is now a written invariant for that service — adding a write endpoint is a change to its security posture, argued for on its own, not a feature detail. I have already declined one feature on those grounds, which is a separate story.

The lesson is not "set your secrets"

That is the reading available on the surface, and it is nearly useless, because of course you should set your secrets. I thought I had.

The actual lesson is about what a fail-closed check is.

I added that guard as a preventative measure — a thing that would stop a bad state from arising in the future. What it did instead was detect a bad state that already existed, immediately, on first deploy. It was not a lock on a door nobody had opened yet. It was a test, run for the first time, against production, that failed.

That reframing has changed how I write these. Every fail-closed check is retroactive as well as preventative: the moment you add one, it renders a verdict on the configuration you are running right now. Which means the interesting moment is not writing it — it is the first deploy afterwards, and you should be watching that deploy rather than assuming a check for something that cannot possibly be happening will pass. Mine did not.

It also means the value of a guard is highest exactly where you are most confident it is unnecessary. I nearly did not write this one. There was no symptom, no bug report, no reason to think anything was wrong, and there could not have been a symptom, because a bypassed authentication system and a working one produce identical observable behavior until someone tries the bypass.

Three consequences I now treat as rules:

A log.Fatal on a required secret is load-bearing and must never be softened into a default. The temptation arrives later, in the form of a local run that will not start or a test environment that is annoying to configure. Every soft-default is an invitation to reproduce this exact bug — the same sweep found a database URL in two services defaulting to a hardcoded changeme credential, which is the same mistake wearing different clothes. Both are now required.

Test that the guard fires. A check you have never watched reject anything is a check you have not tested. Start the service with the variable empty, with it short, with it absent, and confirm each one is a loud exit rather than a successful boot. This costs a minute and is the only way to know the guard is wired to anything.

Ask what your service does when its security-critical config is missing, and require a specific answer. Not "it would probably fail" — what, exactly, happens? For an HMAC key, the answer is "it works perfectly and protects nothing," and the only reason I know that is that I finally asked.

The uncomfortable version, which is why I bothered writing this up: for months, everything I could observe about that service was exactly what a correctly configured service looks like. Every dashboard was green because every dashboard was measuring something else. The single fastest way to find out whether that is true of something you run is to add the check you are sure you do not need, and then actually watch it deploy.