I run the infrastructure behind this site with an AI agent, and like most people doing that I keep a project instructions file at the repository root. Mine tells a fresh session what the moving parts are, which of them will bite you, and what has already gone wrong here. It gets read into context automatically at the start of every session.
On the 1st of August it was 60,623 bytes. Somewhere around eleven thousand tokens, before anyone had said hello.
That number is not the problem in itself — modern context windows swallow it without noticing. The problem is what it was doing to every session that had nothing to do with most of it. A session fixing a spacing bug on the projects grid was loading the Postgres major-upgrade runbook. A session writing a database migration was loading the note about which screenshot script races a CSS animation. Every session paid for every subsystem.
How it got that big, which is the boring part
Honestly? Correctly. Every paragraph in it was added by a session that had just been bitten by something and did not want the next session bitten by the same thing. That is exactly the behavior you want. The file grew because the process worked.
The composition at 60KB:
- The portal applications section alone was 27% of the file.
- The Spark section was roughly 90% duplication of a
README.mdsitting in that directory, which was both longer and more complete than the copy in the instructions file. - Several sections had turned into narrative — not "this will break if you do X" but the full story of the afternoon it broke, which is genuinely worth keeping and is not worth keeping there.
The duplication is the part I would flag if I were reviewing someone else's. Two copies of the same fact do not stay two copies of the same fact. One of them gets updated.
The question I started with was the wrong one
My first pass at trimming asked: what in here is important?
That pass got me nowhere, and it took me an embarrassingly long time to work out why. The answer to that question was "all of it." Everything in the file had been put there by someone who had just learned it the hard way. There is no paragraph in there I would describe as unimportant. Ranked by importance, the file is flat.
The question that actually cuts is different, and it is about the reader rather than the content:
Would a session that is not working on this subsystem be harmed by not having this?
That reframes the whole exercise. It stops being about the value of a fact and starts being about the cost of its absence to someone who never goes near it. And most facts, it turns out, are only dangerous to people already standing in front of the thing.
Two examples from the same file, to show the split is not about how serious something is.
The Postgres 18 image changed its default data directory. If you carry your old bind mount forward without pinning it, the cluster quietly lands in the container's writable layer, a restore reports complete success with matching row counts, and the database evaporates on the next recreate. That stays, in full, unconditionally in context. Not because it is important — because a session that is not thinking about Postgres at all can trigger it while editing a compose file for some other reason, and the failure is silent and total.
The exact procedure for taking a backup before that upgrade, the flags, the verification order — that moves. Nobody executes that procedure by accident. By the time you need it you are already in that directory, and one line saying "read this file before touching migrations, backups, or the image tag" gets you there.
So the rule I settled on: if getting it wrong is silent, dangerous, or non-obvious, it stays inline. If it is something you would look up once, already working in that directory, it moves.
Each section kept its invariants and lost its explanations. What survived reads like a list of tripwires: the data directory is load-bearing, the API service has zero write routes on purpose, there are two separate fail2ban instances and neither can see the other's bans, the log tailer's reopen loop is deliberate, the startup guards are load-bearing and must never be softened into defaults. Each one is a sentence. The reasoning behind each one is one file away, and every stub says to go read that file before changing anything.
Result: 18,047 bytes, about 3,400 tokens. A 70% cut with, as far as I can tell, nothing lost.
"As far as I can tell" was not good enough
That last clause is where this kind of refactor usually goes wrong, and it goes wrong invisibly. You move ten thousand words between files by hand, you feel confident, and eighteen months later you discover one paragraph went into the gap between two documents. There is no test suite for prose. Nothing fails.
So the extraction was mechanical rather than retyped — cut and paste, never re-summarize — and then verified in both directions:
- Before touching the instructions file at all, every content line of every section being extracted was confirmed present in its new home.
- After the edit, the same check was run again, this time against the committed original from git rather than against my memory of it.
That is a tedious check and I nearly skipped it, on the grounds that I had just done the move carefully and would obviously have noticed.
It found something on the first pass. The Spark section was going to be replaced
by a pointer to that directory's own README.md, on the reasonable grounds that
the README was longer, more complete, and about 90% a superset. It was 90% a
superset. The other 10% included a lesson that existed nowhere else on disk:
that testing a notebook means executing its cells, not merely confirming the
server comes up — and that you have to run it twice, because Iceberg's DDL is
transactional rather than apply-if-absent, so a single pass cannot catch a
non-idempotent cell.
Pointing at the README would have deleted that. Not moved it somewhere inconvenient — deleted it, from the only place it had ever been written down. It went into the README's gotchas, where it should have been in the first place.
The whole justification for the verification step is that one paragraph. "90% a superset" is the exact shape of a claim that feels safe and is not, and the only way to know which 10% is missing is to check every line rather than skim for coverage.
The trade-off I do not think I can engineer away
Here is the part I want to be straight about, because most write-ups of a refactor like this end at the 70% number and call it a win.
The reason the Postgres data-directory trap was caught at all is that the instructions file was exhaustive. A session went to make an unrelated change, had the whole picture in front of it including the part about that image's defaults, and connected two things nobody had asked it to connect. Same story for a config-bridging gap in the fail2ban setup. Both of those were caught by having too much in context, not by having the right things in context.
Splitting the file trades that away. The safety net now depends on pointers being followed — on a session actually opening the referenced document rather than working from the one-line stub and assuming it has the gist. Sometimes it will not. That is a real, permanent cost and I do not have a clever answer to it.
Two things make it survivable rather than reckless. The cases where being wrong is silent and unrecoverable stayed inline as one-liners, so the highest-stakes tripwires are still unconditionally in context whether or not anyone follows a link. And every topic now lives in exactly one place — the duplication is gone, not redistributed. Partial duplication is what actually causes drift, because two copies diverge and both keep being read with equal confidence. One authoritative location that might not get opened is a better failure mode than two locations that disagree.
What this is really about
It would be easy to file this under context-window economics, and that framing is available: fewer tokens, cheaper sessions, more room for the actual work. All true, and all fairly boring.
The thing I actually took from it is that an instructions file for an agent is a different artifact from documentation for a human, and the difference is that it has no reader who skims. A person opening a 60KB reference file reads the section they came for. An agent reads all of it, every time, with uniform attention and no sense that most of it is irrelevant to today. Everything you put in there, you are asserting is worth the attention of every future session regardless of what it is doing.
Which makes "would someone not working on this be harmed by not having it?" not a budgeting question at all. It is the question of what you are claiming when you add a line to that file.
The postscript, five weeks later
I checked the file while writing this. It is 34,991 bytes.
It has grown by 17KB since the split — nearly doubled — because five weeks of sessions have each added the thing that had just bitten them, which is precisely the behavior that produced the 60KB file and precisely the behavior I want. Every one of those additions is defensible on its own. That is what makes this kind of file grow: not carelessness, but a long series of individually correct decisions.
So this was never a fix. It is a chore, and it comes back around. What the split actually bought is not a smaller file — it is knowing which question to ask when the file is too big again, and that answering it takes an afternoon rather than a rewrite.