Live data from Hacker News

Claude’s memory architecture is the opposite of ChatGPT’s

shloked.com

211–220 of 240 posts

Re: Claude’s memory architecture is the opposite of ChatGPT’s

#211
post #9

The link to the breakdown of ChatGPT's memory implementation is broken, the correct link is: https://www.shloked.com/writing/chatgpt-memory-bitter-lesson This is really cool, I was wondering how memory had been implemented in ChatGPT. Very interesting to see the completely different approaches. It seems to me like Claude's is better suited for solving technical tasks while ChatGPT's is more suited to improving casual…

Fixed the link! thanks for pointing it out :) I think ChatGPT is trying to be everything at ones - casual conversation, technical tasks - all of it. And it's been working for them so far! Isn't representing past conversations (or summaries) as embeddings already storing memories in encoded forms?

Summaries are still language based. Embeddings aren't, but the models can't use them as input.

Re: Claude’s memory architecture is the opposite of ChatGPT’s

#212

Earlier quoted context omitted.

> If we’re already paying $20/mo and they’re operating at a loss I'm quite confident they're not operating at a loss on those subscriptions.

They are running at a massive loss overall - feels pretty safe to assume that they wouldn't be if their cheapest subscription tier was breaking even

Their cheapest tier is free, they lose money on that of course. And spend a lot of money training new models.

Anthropic has said they have made money on every model so far, just not enough to train the next model, which so far has been much more costly to train every generation. At some point they will probably train an unprofitable model if training costs keep rising dramatically.

OpenAI burns more money on their free tier and might be spending more money building out for future training (I don't know if they do or not) but they both make money on their $20 subscriptions for sure. Inference is very cheap.

Re: Claude’s memory architecture is the opposite of ChatGPT’s

#213
I've been using LLMs for a long time, and I've thus far avoided memory features due to a fear of context rot.

So many times my solution when stuck with an LLM is to wipe the context and start fresh. I would be afraid the hallucinations, dead-ends, and rabbit holes would be stored in memory and not easy to dislodge.

Is this an actual problem? Does the usefulness of the memory feature outweigh this risk?

Re: Claude’s memory architecture is the opposite of ChatGPT’s

#214
post #106

Earlier quoted context omitted.

Anthropic seems to want to make you buy a subscription, not show you ads. ChatGPT seems to be more popular to those who don't want to pay, and they are therefore more likely to rely on ads.

so ChatGPT will become "saleman". And i do not trust any saleman.

You shouldn't be trusting an LLM either, so this is a real sideways move.

Re: Claude’s memory architecture is the opposite of ChatGPT’s

#215
ChatGPT is designed to be addictive, with secondary potential for economic utility. Claude is designed to be economically useful, with secondary potential for addiction. That’s why.

In either case, I’ve turned off memory features in any LLM product I use. Memory features are more corrosive and damaging than useful. With a bit of effort, you can simply maintain a personal library of prompt contexts that you can just manually grab and paste in when needed. This ensures you’re in control and maintains accuracy without context rot or falling back on the extreme distortions that things like ChatGPT memory introduce.

Re: Claude’s memory architecture is the opposite of ChatGPT’s

#216
post #9

The link to the breakdown of ChatGPT's memory implementation is broken, the correct link is: https://www.shloked.com/writing/chatgpt-memory-bitter-lesson This is really cool, I was wondering how memory had been implemented in ChatGPT. Very interesting to see the completely different approaches. It seems to me like Claude's is better suited for solving technical tasks while ChatGPT's is more suited to improving casual…

> It may actually be the final breakthrough we need for AGI. I disagree. As I understand them, LLMs right now don’t understand concepts. They actually don’t understand, period. They’re basically Markov chains on steroids. There is no intelligence in this, and in my opinion actual intelligence is a prerequisite for AGI.

You’re going to be disappointed when you realize one day what you yourself are.

Re: Claude’s memory architecture is the opposite of ChatGPT’s

#217

Earlier quoted context omitted.

They are running at a massive loss overall - feels pretty safe to assume that they wouldn't be if their cheapest subscription tier was breaking even

nonsense for the public. they are Amazon, basically. they take the loss so the overall ecosystem ( x'D like with crypto ) can gain massively, onboard all kinds of target noobs, sry, groups, brutally prime users, discourage as many non-AI processes as possible and steer all industries towards replacing even those processes with AI that are not worth being replaced with AI, like writing and art. of course there are a l…

Thank you for writing this. Your point about "quicker deterioration of local environments" is thought-provoking.

My key technical complaint about LLMs to date is the general inability to add substantial local context. How can I make it understand my business, my processes, my approach to the market? Can I retrain it? Or make it understand my data warehouse?

I think you are explaining why LLM providers don't care about solving my concerns, generally speaking. This is sobering.

Re: Claude’s memory architecture is the opposite of ChatGPT’s

#218
post #61

Earlier quoted context omitted.

One has a more obvious route to building a profile directly off that already collected data. And while they are making lots of revenue even they have admitted on recent interviews that ChatGPT on it's own is still not (yet) breakeven. With the kind of money invested, in AI companies in general, introducing very targeted Ads is an obvious way to monetize the service more.

This is incorrect understanding of unit economics. They are not breaking even only because of reinvestment into r and d.

Sam Altman said in an on-the-record dinner interview with Platformer[0] that besides R&D ChatGPT was breakeven and Brad Lightstep, head of ChatGPT, corrected him by saying they were close, but not yet break even.

I assume Sam and Brad both understand the unit economics of their product.

Article is pay-walled for me, but I heard it on their podcast[1]. Which somehow I heard fine, but that page is getting pay-walled for me.

[0]: https://www.platformer.news/sam-altman-gpt-5-interview-light... [1]: https://www.nytimes.com/2025/08/15/podcasts/hardfork-gpt5-pe...

Re: Claude’s memory architecture is the opposite of ChatGPT’s

#219

Earlier quoted context omitted.

The elephant in the room is that AGI doesn't need ads to make revenue but a new Google does. The words aren't matching with the actions.

The bigger elephant in the room is that LLMs will never be AGI, even by the purely economic definition many LLM companies use.

I've been saying this for years now. LLMs are _not_ the right methodology to get to AGI. My friends who were drinking the kool-aid are only recently coming around to "hey, this might not get us AGI".

But sometimes it feels like I'm the lone voice in a bubble where people are convinced AGI is just around the corner.

I'm wondering if it's because people are susceptible to the marketing, or are just doing some type of 'wishful thinking' - as some seem genuinely interested in AGI.

Re: Claude’s memory architecture is the opposite of ChatGPT’s

#220

Earlier quoted context omitted.

The elephant in the room is that AGI doesn't need ads to make revenue but a new Google does. The words aren't matching with the actions.

The bigger elephant in the room is that LLMs will never be AGI, even by the purely economic definition many LLM companies use.

There are two big innovations required to achieve inexpensive AGI.

LLMs will accelerate discovery and development of Innovation 1, for insanely expensive AGI.

Innovation 1 will accelerate discovery and development of Innovation 2 which will make it too cheap to meter.

Post reply on HN