Handbook.md shows that long policy documents do not reliably govern agents
61–70 of 237 posts
Re: Handbook.md shows that long policy documents do not reliably govern agents
#62This is a problem with long context models. To put it as simple and as bluntly as possible: just because they claim you can use 1M tokens in your context doesn't mean its true and you should do that. Due to extreme quantization of models and the context's KV cache, and also just really shitty samplers provided to the user (hell, most are just getting rid of sampler knobs altogether), this problem will absolutely cont…
It's my understanding that, for all current models having long-context capability, the "parts" of the model that allow them to do long-context processing, are mostly-untouched descendants of academic model architectures; where these academic models were designed around a very constrained use-case assumption: that anyone using more than a "normal" amount of context would specifically be trying to apply a reasonably-sized prompt to an unreasonably-sized block of plain-old embedded data, as a single-shot conversation. (Consider a task like e.g. "summarize this entire book." Context = [prompt] + [text of book].) This is the use-case all these architectures were refined and benchmarked around. And even today, this is still basically the only way these models' long-context processing capabilities work.
---
But I think it goes deeper than that. It's not just about how we've made long-context work so far. It's about what "attending to context" really means, and what that implies about what we'll ever expect a long-context model to be able to do.
I would describe long-context models as seeming to have two distinct types of "attention state" used when attending to the layer-1 input-vector / context window:
- They have some small amount (e.g. 8-32k) of regular "high-quality" attention state. This is likely inherited from the base model they are trained on top of.
- They also have a large amount of "low-quality" attention state. This is what the long-context training process has created.
The "high-quality attention state" isn't limited to referencing the first 8-32k of the context window, to be clear. The high-quality attention mechanism is just limited to "choosing" 8-32k of arbitrary data samples, each from its own arbitrary position in the context window, to form a bounded-sized vector upon which high-quality attention is then computed.
(You could think of this as the high-quality attention mechanism computing attention over the entire context, and then doing top-K sampling, i.e. zeroing out all but the K largest-valued elements of the resulting latent vector, where K=8-32k. That's not what it's actually doing—that'd require paying the very quadratic scaling cost that prevents us from just increasing the short-context window size—but it's a useful naive model for what it's doing.)
---
It's harder to directly explain what the low-quality attention mechanism does, so let me try an analogy.
(While no one can yet give a definitive accounting of the abstract computational processes the matrix-multiply-ops against a given LLM's weights encode, we can still speak of possibility-spaces, or analogies via information/computability theory on what sorts of abstract operations the LLM "must" be performing to achieve the results it achieves. This is that kind of thing.)
Think of an LLM's evolved context-window attendance mechanisms as being somewhat akin to the LLM being a regular CPU program, that has access to a 1. block of raw addressable memory (i.e. the context itself), and 2. a block of pointers into that memory (i.e. the learned first-layer attention over the context.)
Our LLM program's logic, through mutations during training, gradually evolves into a form where it could be seen as implicitly declaring and using a growing set of variables: some bound to static addresses in the block of raw memory (i.e. static context-window positions), and others bound to static addresses in the block of pointers, where the raw-memory-address-value of the pointer (and thus what data from the context-window it will read as when "dereferenced") is the result of a dynamic calculation over one or more other variables (i.e. an arbitrary weighting of the layer-1 input vector, where the resulting vector element is used as a selector, in layer 2, among other data that got sampled into "passthrough" vector-elements back in layer 1.)
Under this analogy, a model's "high-quality attention state" is the set of "named" pointer variables that it directly references in its logic.
As should be obvious, there can only be so many of these; you really fundamentally can't scale them up, without literally having "more logic" that defines and makes use of them (i.e. scaling up the size-in-weights of the model's layers.)
A model's "low-quality attention state", on the other hand, is the block of pointer memory itself, which can be as large as you like. Presuming the computational substrate permits the logic to compute on the addresses of memory within that pointer-memory block and then double-dereference them to get context data (and it does seem to!), your LLM's evolved logic can do "dynamic random-access reads" of arbitrary parts of the long context.
But, of course, the model won't be doing anything special with that data. It can attend to it, but not in a way that's unique to that data. It pulled in such data through generic logic that doesn't "know what it's looking at", after all.
---
IMHO, this analogy helps to make clear that while giving a model a raw "large-context capability" might help said model to attend to more data, it doesn't really help in making a model able to attend to a larger prompt.
Just by information theory, an instruction-following capability must imply that something analogous to logic keyed on those instructions gets embedded into the weights; and that logic, in order to activate and work, presumably requires its own state — i.e. requires that individual values from the context are getting attended into additional top-level named variables.
And while the short-context attendance mechanism does that, the long-context attendance mechanism inherently does not / cannot.
And this means that there is likely a practical upper limit (in terms of practical model-weight size + latent-vector size + etc) to "how much prompt" a model can follow during a single inference step. Regardless of how much data the model can "see" on a given inference step, its active "program logic" is only defined in terms of so many concurrent stateful computations.
---
Thus: trying to use a current long-context model's "low quality" attention to attend to a large prompt is just going to result in nonsense. None of these model families have ever been trained to follow a million sentences of rules simultaneously when answering a question. They're just going to attend with their inherited bounded "high quality" attention to your prompt as best as they're able, while forgetting literally everything about the prompt that doesn't fit in that bounded "high quality" window.
And anything that allows current long-context models to appear to interact well in multi-turn conversations with memory + KV caching + etc layers in operation — while also making use of its "low quality" attention to attend to large data — is some kind of hack, and will break like a hack.
Re: Handbook.md shows that long policy documents do not reliably govern agents
#63There was an article a few years ago called "Lost in the Middle: How Language Models Use Long Contexts" https://arxiv.org/abs/2307.03172 From my experience this holds true to this day. It was one of my core observations for similarity to the limitations of human working memory on "Engineering for Bounded Cognition"
Richard Hendricks solved this decisively with middle-out compression
Re: Handbook.md shows that long policy documents do not reliably govern agents
#64Most people don't understand that 'agentic AI' is a completely synthetic, force fed capability by extensive Reinforcement Learning on synthetic domain specific 'agentic' datasets on post training. If the LLM wasn't post-trained to adhere to specific handbook, it just won't work. If the LLM wasn't trained on an use case the lab decided was worth making a synthetic agentic dataset, it won't work as well as you want. Th…
What do you mean by a graph of one shot prompts?
Geologist report -> LLM call extracts minerals we are looking for (you inject a db query result on the user prompt), locations, assay mentions and risks into structured fields -> LLM call classifies evidence into positive indicators, negative indicators and unknowns -> LLM call estimates deposit potential and confidence -> database lookups inject regional ore demand, nearby deposits, infrastructure and historical yield data -> LLM call combines geological evidence with business context -> LLM call generates an investment recommendation and rationale.
That's what I mean by graph. Every step is a separate LLM call with a well-defined responsibility, consuming the output of the previous node. Each node can be tested, benchmarked, retrained, replaced, or monitored independently. Why would you use an agent here? You can cache every single system prompt on each call making your total token output much cheaper than having a full 'output' only token generation workflow which is what happens with agents.
There is nothing to discover. The workflow is already known. The company already knows how geologists evaluate prospects. The company already knows what data sources matter. The company already knows what the final output should look like. You don't want the model deciding which tools to call, which reasoning path to take, or which pieces of information are important every single run. You want the exact same process applied to every report so results are consistent, measurable, auditable and debuggable. My default is: One-shot prompt -> if not enough -> graph of LLM calls -> if not enough -> agent. A surprising amount of enterprise AI is really just: Unstructured input -> extraction -> classification -> enrichment from databases -> decision support. Not: Unstructured input -> autonomous agent spends 20 steps deciding what to do next.
Agents make sense when the workflow itself is unknown.
If the workflow is already understood, a graph is usually cheaper, more reliable, easier to evaluate, easier to debug, and less dependent on whatever synthetic "agentic" behaviors happened to get reinforced during post-training. I am sure people default to agents mostly because it's less engineering work than explicitly modeling the process.
Re: Handbook.md shows that long policy documents do not reliably govern agents
#65In my experience, the more structure you enforce on models, the worse they perform and the less they actually do what you want.
It does what its has been trained to do. So find out what its trained to do and just use it to do that. This is not general intelligence.
Re: Handbook.md shows that long policy documents do not reliably govern agents
#66Most people don't understand that 'agentic AI' is a completely synthetic, force fed capability by extensive Reinforcement Learning on synthetic domain specific 'agentic' datasets on post training. If the LLM wasn't post-trained to adhere to specific handbook, it just won't work. If the LLM wasn't trained on an use case the lab decided was worth making a synthetic agentic dataset, it won't work as well as you want. Th…
i thought it was well known that claude code got good at coding because anthropic bought tons of coding data from companies like mercor.
Re: Handbook.md shows that long policy documents do not reliably govern agents
#67Glad to be able to put some numbers on it.
Re: Handbook.md shows that long policy documents do not reliably govern agents
#68Multiple rounds of generating small contacts documents that grow from the original idea , trying to keep each slice small enough to process for a human to approve/disprove .
Eventually it leads to a long list of tasks grouped by functionality. You start a new context and the orchestrator agent dispatches tasks to sub agents with a limited amount of information provided to each sub agent.
Also should have adversarial review and approval gates with other agents and roles.
Re: Handbook.md shows that long policy documents do not reliably govern agents
#69For my own (rather idiosyncratic) harness I've been experimenting [0] with "compiling" long markdown specifications into small executable logic programs. It's too early to tell for sure, but I believe that this approach does have its merits when you want some guarantees about how your agents behave for longer tasks. [0] https://github.com/deepclause/deepclause-sdk
Re: Handbook.md shows that long policy documents do not reliably govern agents
#70Having only dipped my toes in generative ai recently I'm surprised how small even a 1mio window is. A semi serious project can take several session in one sitting. Especially as degradation sets in waay before the window is full.
Dont try and handle the entire project in context.
Use a well structured filesystem layout for your code with a few lines in an AGENTS.md describing the layout and core architectural requirements (no more than that, as per the article!).
Then work on small-medium tasks at a time with a fresh context.
At the end of your task, ask the agent if there are any key points about the project layout or architecture it would want to add to memory - audit those manually and amend your AGENTS.md accordingly.
If your code is well structured and you keep your tasks localised, you can get away with seemingly minuscule context windows.
It's also worth noting that high effort models love to slurp up whatever context they can. You almost never need/want high effort for non cross-cutting tasks.