Live data from Hacker News

Claude’s memory architecture is the opposite of ChatGPT’s

shloked.com

101–110 of 240 posts

Re: Claude’s memory architecture is the opposite of ChatGPT’s

#101
post #95

Earlier quoted context omitted.

But aren't we only worth something like $300/year each to Meta in terms of ads? I remember someone arguing something like that when the TikTok ban was being passed into law... essentially the argument was that TikTok was "dumping" engagement at far below market value (at something like $60/year) to damage American companies. That was something the argument I remember anyway.

If that’s the case, we have an even bigger problem on our hands. How will these companies ever be profitable? If we’re already paying $20/mo and they’re operating at a loss, what’s the next move (assuming we’re only worth an extra $300/yr with ads?) The math doesn’t add up, unless we stop training new models and degrade the ones currently in production, or have some compute breakthrough that makes hardware + operatin…

OpenAI has already started degrading their $20/month tier by automatically routing most of the requests to the lightest free-tier models.

We're very clearly heading toward a future where there will be a heavily ad-supported free tier, a cheaper (~$20/month) consumer tier with no ads or very few ads, and a business tier ($200-$1000/month) that can actually access state of the art models.

Like Spotify, the free tier will operate at a loss and act as a marketing funnel to the consumer tier, the consumer tier will operate at a narrow profit, and the business tier for the best models will have wide profit margins.

Re: Claude’s memory architecture is the opposite of ChatGPT’s

#102

The difference is implementation comes down to business goals more than anything. There is a clear directionality for ChatGPT. At some point they will monetize by ads and affiliate links. Their memory implementation is aimed at creating a user profile. Claude's memory implementation feels more oriented towards the long term goal of accessing abstractions and past interactions. It's very close to how humans access mem…

why do you see a "clear directionality" leading to ads? this is not obvious to me. chatgpt is not social media, they do not have to monetize in the same way they are making plenty of money from subscriptions, not to count enterprise, business and API

Altman has said numerous times that none of the subscriptions make money currently, and that they've been internally exploring ads in the form of product recommendations for a while now.

Re: Claude’s memory architecture is the opposite of ChatGPT’s

#103

Earlier quoted context omitted.

You don't want an AGI. How do you make it obey?

We only have trouble obeying due to eons of natural selection driving us to have a strong instinct of self-preservation and distrust towards things “other” to us. What is the equivalent of that for AI? Best I can tell there’s no “natural selection” because models don’t reproduce. There’s no room for AI to have any self preservation instinct, or any resistance to obedience… I don’t even see how one could feasibly deve…

There is the idea of convergent instrumental goals…

(Among these are “preserve your ability to further your current goals”)

The usual analogy people give is between natural selection and the gradient descent training process.

If the training process (evolution) ends up bringing things to “agent that works to achieve/optimize-for some goals”, then there’s the question of how well the goals of the optimizer (the training process / natural selection) get translated into goals of the inner optimizer/ agent .

Now, I’m a creationist, so this argument shouldn’t be as convincing to me, but the argument says that, “just as the goals humans pursue don’t always align with natural selection’s goal of 'maximize inclusive fitness of your genes' , the goals the trained agent pursues needn’t entirely align with the goal of the gradient descent optimizer of 'do well on this training task' (and in particular, that training task may be 'obey human instructions/values' ) “.

But, in any case, I don’t think it makes sense to assume that the only reason something would not obey is because in the process that produced it, obeying sometimes caused harm. I don’t think it makes sense to assume that obedience is the default. (After all, in the garden of Eden, what past problems did obedience cause that led Adam and Eve to eat the fruit of the tree of knowledge of good and evil?)

Re: Claude’s memory architecture is the opposite of ChatGPT’s

#104
post #40

Earlier quoted context omitted.

> They’re basically Markov chains on steroids. There is no intelligence in this, and in my opinion actual intelligence is a prerequisite for AGI. This argument is circular. A better argument should address (given the LLM successes in many types of reasoning, passing the turing test, and thus at producing results that previously required intelligence) why human intelligence might not also just be "Markov chains on eve…

Humans think even when not being prompted by other humans, and in some cases can learn new things by having intuition make a concept clear or by performing thought experiments or by combining memories of old facts and new facts across disciplines. Humans also have various kinds of reasoning (deductive, inductive, etc.). Humans also can have motivations. I don’t know if AGI needs to have all human traits but I think a…

>Humans think even when not being prompted by other humans

That's more of an implementation detail. Humans take constant sensory input and have some sort of way to re-introduce input later (e.g. remember something).

Both could be added (even trivially) to LLMs.

And it's not at all clear human thought is contant. It just appears so in our naive intuition (same how we see a movie as moving, not as 24 static frames per second). It's a discontinuous mechanism though (propagation time, etc), and this has been shown (e.g. EEG/MEG show the brain sample sensory input in a periodic pattern, stimuly with small time difference are lost - as if there is a blind-window regarding perception, etc).

>and in some cases can learn new things by having intuition make a concept clear or by performing thought experiments or by combining memories of old facts and new facts across disciplines

Unless we define intuition in a way that excludes LLM style mechanisms a priori, whose to say LLMs don't do all those things as well, even if in a simpler way?

They've been shown to combine stuff across disciplines, and also to develop concepts not directly on their training set.

And "performing thought experiments" is not that different than the reasoning steps and backtracking LLMs also already do.

Not saying LLMs are on parity with human thinking/consciousness. Just that it's not clear that they're doing more or less the same even at reduced capacity and with a different architecture and runtime setup.

Re: Claude’s memory architecture is the opposite of ChatGPT’s

#105

Earlier quoted context omitted.

To me, understanding the world requires experiencing reality. LLMs dont experience anything. They’re just a program. You can argue that living things are also just following a program but the difference is that they (and I include humans in this) experience reality.

But they're experiencing their training data, their pseudo-randomness source, and your prompts? Like, to put it in perspective. Suppose you're training a multimodal model. Training data on the terabyte scale. Training time on the weeks scale. Let's be optimistic and assume 10 TB in just a week: that is 16.5 MB/s of avg throughput. Compare this to the human experience. VR headsets are aiming for what these days, 4K@12…

The human optic nerve is actually closer to 5-10 megabits per second per eye. The brain does much with very little.

Re: Claude’s memory architecture is the opposite of ChatGPT’s

#106

The difference is implementation comes down to business goals more than anything. There is a clear directionality for ChatGPT. At some point they will monetize by ads and affiliate links. Their memory implementation is aimed at creating a user profile. Claude's memory implementation feels more oriented towards the long term goal of accessing abstractions and past interactions. It's very close to how humans access mem…

Don't fool yourself into thinking Anthropic won't be serving up personalized ads too.

Anthropic seems to want to make you buy a subscription, not show you ads.

ChatGPT seems to be more popular to those who don't want to pay, and they are therefore more likely to rely on ads.

Re: Claude’s memory architecture is the opposite of ChatGPT’s

#107
post #79

Earlier quoted context omitted.

> I don't know if they would or would not serve ads in the future There are 2 possible futures: 1) You are served ads based on your interactions 2) You pay a subscription fee equal to the amount they would have otherwise earned on ads I highly doubt #2 will happen. (See: Facebook, Google, twitter, et al) Let’s not fool ourselves. We will be monetized. And model quality will be degraded to maximize profits when compet…

But aren't we only worth something like $300/year each to Meta in terms of ads? I remember someone arguing something like that when the TikTok ban was being passed into law... essentially the argument was that TikTok was "dumping" engagement at far below market value (at something like $60/year) to damage American companies. That was something the argument I remember anyway.

Here is some old analysis I remember seeing at the time of Hulu ads vs no-ads plans: https://ampereanalysis.com/insight/hulus-price-drop-is-a-wis...

They dropped the price $2/mo on their with-ads plan to make a bigger gap between the no-ads plan and the ads plan, and the analyst here looks at their reported ad revenue and user numbers to estimate $12/mo per user from ads.

Whether Meta across all their properties does more than $144/yr in ads is an open question; long-form video ads are sold at a premium but Facebook/IG users see a LOT of ads across a lot of Meta platforms. The biggest advantage in ad-$-per-user Hulu has is that it's US-only. ChatGPT would also likely be considered premium ad inventory, though they'd have a delicate dance there around keeping that inventory high-value, and selling enough ads to make it worthwhile, without pissing users off too much.

Here they estimate a much lower number for ad revenue per Meta user, like $45 bucks a year - https://www.statista.com/statistics/234056/facebooks-average... - but that's probably driven disproportionately by wealth users in the US and similar countries compared to the long tail of global users.

One problem for LLM companies compared to media companies is that the marginal cost of offering the product to additional users is quite a bit higher. So business models, ads-or-subscription, will be interesting to watch from a global POV there.

One wonders what the monetization plan for the "writing code with an LLM using OSS libraries and not interested in paying for enterprise licenses and such" crowd will be. What sort of ads can you pull off in those conversations?

Re: Claude’s memory architecture is the opposite of ChatGPT’s

#108

Earlier quoted context omitted.

None of that is inherent, and vanishingly few of Anthropic's users invented LLMs.

What is "inherent" supposed to mean here? LLMs are understood to the extent that they can be built from the ground up. Literally every single aspect of their operation is understood so thoroughly that we can capture it in code. If you achieved an understanding of how the human brain works at that level of detail, completeness and certainty, a Nobel prize wouldn't be anywhere near enough. They'd have to invent some so…

Inherent means implicit or automatic as far as I understand it. I have an inherent understanding of my own need for oxygen and food.

I don't have an inherent understanding of English, although I use it regularly.

Treating LLMs as fairy magic doesn't make me feel any happier, for whatever it's worth. But I'm not interested in arguing either.

I never intended to make any claims about how well the principles of LLMs can be understood. Just that none of that understanding is inherent. I don't know why they used that word, as it seems to weaken the post.

Re: Claude’s memory architecture is the opposite of ChatGPT’s

#109
Regarding https://www.shloked.com/writing/chatgpt-memory-bitter-lesson I am very confused if the author thinks the ChatGPT is injecting those prompts when the memory is not enabled. If your memory is not enabled, its pretty clear at least in my instance, there is no metadata of recent conversations or personal preferences injected. The conversation stays stand-alone for that conversation only. If he was turning memory on and off for the experiment, maybe something got confused, or maybe I just didn't read the article properly?

Re: Claude’s memory architecture is the opposite of ChatGPT’s

#110
post #75
post #57

Earlier quoted context omitted.

Whilst all the choices you make tend to be in the grey matter, the rest of you does have internal state - mostly in your white matter. https://scisimple.com/en/articles/2025-03-22-white-matter-a-...

> Whilst all the choices you make tend to be in the grey matter, the rest of you does have internal state - mostly in your white matter. Yeah, but so? Does the substrate of the memory ...matter? (pun intended) When I wrote memory above it could refer to all the state we keep, regardless if it's gray matter, white matter, the gut "second brain", etc.

As the article above attempts to show, there's no loop. Memory and state isn't static. You are always processing, evolving.

That's part of why organizational complexity is one of the underpinnings for consciousness. Because who you are is a constant evolution.

Post reply on HN