Live data from Hacker News

The new rules of context engineering for Claude 5 generation models

claude.com

221–230 of 434 posts

Re: The new rules of context engineering for Claude 5 generation models

#221
post #65

Earlier quoted context omitted.

> Most of this article seems like... common sense? i think you'd be surprised. every model release there's seemingly hordes of people who proclaim the new model is terrible and they're going back to the old one, and it all stems from people still prompting and having their configs setup like we're back in the sonnet 3.5 days

I have a coworker that was complaining about Opus 5 and had random shitty skills and custom plugins wired in from YouTube tutorials watched over the past year. He also speaks with the model like it's GPT 4o. Needless to say, none of the new models have worked well for him, and he refuses to remove the "tweaks" or update his style of communication, which is obviously breaking the experience.

> Needless to say, none of the new models have worked well for him, and he refuses to remove the "tweaks" or update his style of communication, which is obviously breaking the experience.

All attempts to control the output in a useful way for the user, in a way where the output is as reliable and repeatable as possible... and with a system not at all designed for it, that gets worse the more rules you throw at it.

Seems like a problem.

Re: The new rules of context engineering for Claude 5 generation models

#222
post #32

> Now: Auto-memory Not Claude Code but I just had a task where it started referring another conversation that was complete nonsense and throwaway. I absolutely don't want things to get added to some memory behind my back. A big reason I use LLMs is because I can try out wild ideas and then just throw it away. I don't want those to pollute the context.

Claude Code was writing one-off details (specific to one task) as memories, so I instructed it that memories require explicit approval before writing, and that seems to work reasonably well. That said, in my experience, Claude gets overly attached to the context, and will follow them even when irrelevant, or will often mention some mostly irrelevant detail from the context as if I asked for it explicitly in my prompt.

Re: The new rules of context engineering for Claude 5 generation models

#223

We should design a specific language to make sure that we can encode the exact requirements that we want. Something that has a limited set of keywords that are explicit. Wait a minute...

This is getting tiring. Code is not The Specification. It’s a specification of God knows what. Riddled with irrelevant, non-essential details wrapping The Problem - which in most cases will amount to something the size of a large pebble - in multiple layers of fur jackets, stored in boxes, which themselves are stored in multiple ridiculous moveable warehouse (if you’re lucky). We have a standard for communication, it…

To piss off architects, managers and linguists all at the same time: language isn't much of a standard at all, it's more the current agreed-ish state of things, quite similar to the current state of a code base. Your inner model of what you like things to be is without direct effect to how things de facto are.

In addition: it's not wrapped around a problem, it's wrapped around an attempt at a solution - the problem space is often not even depicted in code, and often only minimally described in documentation.

Re: The new rules of context engineering for Claude 5 generation models

#224
post #30

Earlier quoted context omitted.

I’m not excited about using Opus 5, mainly because the way that I work atm — essentially peer programming — means I sandbox the agents and work with them closely. Opus 4.x encounters the sandbox and moves on with its day; Fable becomes increasingly fixated on it and does less and less of the actual task, focussing more and more on the limit it reached. I worry that, from your description, Opus 5 will do the same.

Ive been using btrfs snapshots and some auto generated isolation rules plus a git ceiling at the mount root for the btrfs image (have to do this in wsl, stupid work computer). it's worked really well and fable hasn't had any issues with the "sandbox" (obviously not really but it works well enough)

I've been using something similar with ZFS, 15min frequency with autopruning (Sanoid) and the snapdir mounted for the agent. Has worked well and a big plus is being able to tell the agent to just solve it's mistake via restore from the snapdir.

Re: The new rules of context engineering for Claude 5 generation models

#226
post #205

Earlier quoted context omitted.

It's language and speech patterns that seem designed to trick readers into believing that claims are correct, even when the claims aren't based on anything and are possibly wrong. It was rewarded for this during training for some reason. Alternative theory: The LLMs only way to "think" about abstract concepts is through language, and this leaks into into conversation it has with humans. But humans generally prefer to…

Maybe we need LLMs which have an internal dialogue rather than the current monologue.

> Maybe we need LLMs which have an internal dialogue rather than the current monologue.

We already have them. They are called LLMs. The internal dialogue you speak of are the vectors in the so called latent space.

Re: The new rules of context engineering for Claude 5 generation models

#227

Earlier quoted context omitted.

Exactly, like weather forecasting. If you're told there's a 30% chance of rain, it doesn't mean that 3 out of 10 times you will experience rain. Either it will rain or it won't, so the probability is either 0% or 100%. And so a forecast of "30% chance of rain" is referring to the likelihood that your probability will be 100%, as opposed to 0%.

> Either it will rain or it won't, so the probability is either 0% or 100%. And so a forecast of "30% chance of rain" is referring to the likelihood that your probability will be 100%, as opposed to 0%. This is a huge misunderstanding of what probability means.

75% chance you’re wrong.

Re: The new rules of context engineering for Claude 5 generation models

#228

They are imo over-relying on Claude automemory here, which is terrible at contextualizing memory access and makes huge leaps that don’t make sense - except when it’s actually useful, which makes the problem even worse for an operator who can’t see the thinking process anymore. Yes, I worked on a related project, no I don’t want you to use those memories to make assumptions which emerge as decisions that I didn’t want…

A CLAUDE.md file has no moat, it can be read by other agents.

Automemory can be weaved into the product in ways that make it harder to switch.

This is a company that's looking to IPO soon at a trillion+ dollar valuation, and they need to pull every lever to keep the users they got during the past year's boom.

Re: The new rules of context engineering for Claude 5 generation models

#229

Earlier quoted context omitted.

This is true and it's barely even debatable. Whatever exact role language plays in our thought processes, it is most definitely nonzero. It's why I think "LLMs are only fancy autocorrect" style takes are really underselling how wild it is that we've, in a roundabout way, sort of crystallized a bit of the human thought process in a way that is genuinely useful for a lot of tasks. Linguistic Relativity — John Lucy http…

> sort of crystallized a bit of the human thought process a) LLMs don't think. They predict a most probable sequence of language tokens. Huge difference there. b) Whatever LLMs do doesn't model human behavior whatsoever. LLMs are basically very fancy logistic regressors. I.e., it's a mathematical abstraction first and foremost.

It's amazing that you can predict a counterexample to an open math problem, all without thinking.

Re: The new rules of context engineering for Claude 5 generation models

#230

Earlier quoted context omitted.

It's language and speech patterns that seem designed to trick readers into believing that claims are correct, even when the claims aren't based on anything and are possibly wrong. It was rewarded for this during training for some reason. Alternative theory: The LLMs only way to "think" about abstract concepts is through language, and this leaks into into conversation it has with humans. But humans generally prefer to…

"There is a depth of thought untouched by words, and deeper still a depth of formless feeling untouched by thought." - Rilke Your assertion that we don't think in language is questionable. It runs counter to the lived experience of developing thoughts through writing ("writing isn't capturing thinking -- it is thinking"). I believe there is more to thought than language alone, but I also feel quite sure that language…

When I was young, a friend asked me, "Hey, you speak three languages, which one do you think in?"

I paused, confused, and replied, "People think in words?"

Fast forward a decade or so, in my twenties, I had lost most of the inner visual sense I had previously used, and developed an overreliance, in my opinion, on language. (I think my dominant sense was some "non visual abstract sense of ideas", but the visual was also very strong.)

In other words, I now do think mostly in words, and it feels a lot harder to get any serious work done. The language-ing is involuntary, and I often wish I had a way to shut it off, because it seems to actively interfere with more subtle mental processes.

More recently, I often have the experience where I will wake up from a dream with some complex idea fully formed in my mind. I write it down before it fades, and then spend the next hour or two trying to understand it.

The best explanation I have right now is that there are at least two minds: one which operates holistically — if it were a 3D printer, it would be like that one with the bath, where the object emerges from the bath, whole.

Whereas the other one (the conscious mind) would be the extrusion printer with the tiny nozzle that has to zip around for a long time to achieve a worse result. (And must be constantly cooled, less it overheat!)

Post reply on HN