As with many engineers, I have tried to keep up with the rapidly changing landscape that constantly improving LLMs have made easily accessible. A recent trend for larger scale projects is to use a top level session - that instead of doing implementation work itself, breaks up the tasks and spins up sub ‘agents’ to do actual implementation. The sub-agents report back when they are done and the top level session reviews the work.
The Problem - Context management, token usage, and parallelism
Context and Token Usage
Unless your problem takes a small number of prompts and responses to solve, context management is an important concept to understand and manage in sessions. Each message sent during a session sends the entire session history along with it- including your messages to the model, the models responses, and any internal dialogue the model has with itself, the commands it runs and the output of those commands, and other memory based objects like files that have been read. This affects tokens and context. The message history grows linearly by a factor of N (where N is the number of messages sent in the session), but token usage grows by a factor around N². This is why you hear the recommendation to start a new session every 20 messages, or to ‘compact’ your context before it reaches 50% of the total context window.
Compacting the context creates a summary that acts as the previous context going forward. This helps reduce token usage, but across multiple context compacts you typically start to notice a significant drop in performance, an increase in errors, and the model ‘forgetting’ things that you discussed or did previously. For large projects, context fills up quickly and you typically end up going through multiple sessions while having to do some ‘knowledge transfer’ between them. This is an extremely inefficient approach to leveraging an LLM for a big task.
Parallelism
You can help solve this problem with agents. The idea is that your session takes in what you want, but instead of going and grabbing the information you want or implementing something you want to build the model spins up other models with specific tasks. The session tells each agent what its task is and then reviews the work at the end. The top level session ends up saving a ton of context, and you can dispose of the agent once it is done- so you no longer get runaway token usage. Having the session spin up multiple agents at once allows tasks to be completed in parallel, with the top level session managing and checking the work of each before stitching it all together and presenting it back to you. For many projects, this approach works well. This is not a new concept, but for truly momentous tasks this strategy is not enough.
By default, top level sessions will send the entire context window + a description of the task to each agent every time it spins up an agent. The context for each agent is also lost after an agent finishes its task, so critical information on how the task was completed, what issues it ran into along the way, and nuances or caveats related to the completed work get lost as well.
Context Libraries and Agent Hierarchy
Originally, I began to use ‘Context Libraries’ in an effort to work on the same project from different computers. All LLM sessions start in a context folder that lives on a Nextcloud instance which syncs to my laptop and desktop. Each session is directly instructed or instruction are placed in the system prompt to create a new file in the context library for a new project. It is also instructed to keep the context file up to date with the goals of the project, the plan, the progress, and any ‘pointers’ to other context files that may be relevant. Not only does this make switching between sessions on the same project significantly easier (because you don’t have to re-explain things or do a knowledge transfer), it allows sessions from different projects that may link or interact together to better understand necessary information to complete a task. This approach also allows you to become extremely context efficient with pointers.
Instead of keeping all the context for somethin in one folder, you create a hierarchy or context- so there may be a ‘core’ context file with high level information, and within that file there can be pointers to more detailed context files relating to specific tasks or components of the larger project. This allows you to be context efficient because you no longer have to pass the entire context into new sessions or agents. Only the applicable context for a specific tasks needs to be passed to agents, which maintain a context file for their task, then the top level session records the results and pointer to the task context in the top level context file. I know this sounds simple, and it really is- but by default your sessions are not doing this.
Unfortunately for my original use case, doing this in a shared location that syncs to several computers has allot of problems- like race conditions or merge conflicts. I have been working on this problem separately.
I will be writing a separate note about shared context libraries.
An Efficient Strategy For Climbing a Mountain
Ok, enough setup. I will now run through a design that I am currently using for a very very large project, with great success.
I am building a game that I have always wanted to play. A game that merges my favorite pieces of several other games- the core two being Factorio and Kerbal Space Program. I am not a game developer, nor do I have any previous experience in game development. This project came about more as a question “How good have LLMs gotten”. The current form is the third iteration of the project- and so far it has been smooth sailing.
I am not reviewing or writing any code myself- I am purely designing the game functionality, testing, and designing the implementation strategy below. So yes, this is a completely ‘vibe coded’ game. Why not?
Implementation Architecture

A game (especially of this magnitude) is a very large project. Each piece is its own massive project. Because of that, I went with a 3 tier agent hierarchy, with each layer managing its own context library. Each manager takes an epic (like ‘Fluids, fluid handling, fluid research path, manufacturing fluid items, view model for fluids, etc) and breaks up the problem into tasks for each worker agent.
Originally, each manager and task was persistent and would just get new epics or tasks. As you can imagine this created context window overflows and significantly increased token usage. To solve this, each task and manager were killed once they completed their items- and I added a context library for each layer.
Impact of the Context Library
At first I just had the three tier hierarchy- with the top level session managing context at a high level. The game was making good, rapid progress, but I started to run through my 5 hour usage window in less than 20 minutes (for reference, Claude Code Max 5x). I then made each layer its own context library and managers / workers ephemeral- now development runs 24/7 without every hitting a 5 hour limit. I still run through the weekly usage limit in about 4 days.. so there is certainly still room for improvement.
Context strategy
- Short-lived agents: a fresh worker per task and a fresh manager per epic. None carries history forward.
- Disk is the memory: each agent writes a capped summary when done (40 lines per task, 80 per epic) plus one line in
INDEX.md. New agents get a focused prompt plus pointers to those files, never the full history. - One plan of record:
BRIEF.mdholds goals, decisions, rules and your feedback log, and wins any conflict with older docs. - Recovery: because everything lives in git, a usage cutoff or a Windows restart costs little. A fresh manager reads the epic’s files and the worktree’s commits, then continues.
Parallelism and Isolation
- Two lanes at most: each manager works in its own git worktree and branch. Main stays clean, and only the Director merges.
- Small commits: agents commit small and often, so killed work is never lost.
- Merge conflicts: a manager first merges main into its branch and resolves the conflicts. The Director then merges to main and runs the full suite once.
- Watchdog: a 30-minute check relaunches dead lanes and starts the next queued epic.
Verification Budget
- Test only what changed: no agent playthroughs. A test builds the exact game state it needs directly, checks one thing and exits.
- Tiered tests: core unit tests first, then a fast suite (about 2 minutes) after each change. The full parallel suite runs once per epic.
- Load-aware timing: timing checks allow for a busy machine, and a hung test is killed and reported by a watchdog.