Live data from Hacker News

Not all tokens are meant to be forgotten

arxiv.org

1–10 of 27 posts

Re: Not all tokens are meant to be forgotten

#2
> However, they tend to memorize unwanted information, such as private or copyrighted content,

I mean humans don't forget copyrighted information. We just typically adjust it enough (some of the time) to avoid getting a copyright strike while modifying it in some way useful.

We don't forget 'private' information either. We might not tell other people that information, but it still influences our thoughts.

The idea of a world where we have AI minds forget vast amounts of information that humans have to deal with every day is concerning and dystopian to me.

Re: Not all tokens are meant to be forgotten

#3
post #2

> However, they tend to memorize unwanted information, such as private or copyrighted content, I mean humans don't forget copyrighted information. We just typically adjust it enough (some of the time) to avoid getting a copyright strike while modifying it in some way useful. We don't forget 'private' information either. We might not tell other people that information, but it still influences our thoughts. The idea of…

[deleted]

Re: Not all tokens are meant to be forgotten

#4
post #2

> However, they tend to memorize unwanted information, such as private or copyrighted content, I mean humans don't forget copyrighted information. We just typically adjust it enough (some of the time) to avoid getting a copyright strike while modifying it in some way useful. We don't forget 'private' information either. We might not tell other people that information, but it still influences our thoughts. The idea of…

I'd counter with an anecdote; I had a colleague that boasted how he memorized a classmate's SSN in college and would greet him by SSN when seeing him years later. Is the goal of AI to replicate the entirety of the human experience (including social pressures, norms, and shame) or a tool to complement human decision making?

While, yes, you can argue the slippery slope, it may be advantageous to flag certain training material as exempt. We as humans often make decisions without perfect knowledge, and "knowing more" isn't a guarantee that it produces better outcomes, given the types of information consumed.

Re: Not all tokens are meant to be forgotten

#6
post #2

> However, they tend to memorize unwanted information, such as private or copyrighted content, I mean humans don't forget copyrighted information. We just typically adjust it enough (some of the time) to avoid getting a copyright strike while modifying it in some way useful. We don't forget 'private' information either. We might not tell other people that information, but it still influences our thoughts. The idea of…

I'd counter with an anecdote; I had a colleague that boasted how he memorized a classmate's SSN in college and would greet him by SSN when seeing him years later. Is the goal of AI to replicate the entirety of the human experience (including social pressures, norms, and shame) or a tool to complement human decision making? While, yes, you can argue the slippery slope, it may be advantageous to flag certain training m…

Knowing more might not improve your accuracy but it's not going to harm it. Forcibly forgetting true parts of your knowledge seems far more likely to have unintended consequences.

Re: Not all tokens are meant to be forgotten

#7
post #6

Earlier quoted context omitted.

I'd counter with an anecdote; I had a colleague that boasted how he memorized a classmate's SSN in college and would greet him by SSN when seeing him years later. Is the goal of AI to replicate the entirety of the human experience (including social pressures, norms, and shame) or a tool to complement human decision making? While, yes, you can argue the slippery slope, it may be advantageous to flag certain training m…

Knowing more might not improve your accuracy but it's not going to harm it. Forcibly forgetting true parts of your knowledge seems far more likely to have unintended consequences.

Counterpoint: There are plenty examples of breakthroughs from folks who are ignorant of the “right” way to go about it. A fresh take isn’t always bad.

Re: Not all tokens are meant to be forgotten

#8
post #6

Earlier quoted context omitted.

I'd counter with an anecdote; I had a colleague that boasted how he memorized a classmate's SSN in college and would greet him by SSN when seeing him years later. Is the goal of AI to replicate the entirety of the human experience (including social pressures, norms, and shame) or a tool to complement human decision making? While, yes, you can argue the slippery slope, it may be advantageous to flag certain training m…

Knowing more might not improve your accuracy but it's not going to harm it. Forcibly forgetting true parts of your knowledge seems far more likely to have unintended consequences.

I disagree. Actively fighting against your memory will slow you down in any context where some memorized idea is similar to what you're doing but you shouldn't be using the memorized idea.

Re: Not all tokens are meant to be forgotten

#9
There’s a related paper that Meta published a couple of days ago that is worth looking at:

> How much do language models memorize?

https://arxiv.org/abs/2505.24832

https://news.ycombinator.com/item?id=44171363

It shows that models are limited in how much they can memorise (~3.6 bits per parameter), and once that threshold is reached, the model starts to generalise instead of memorise.

Re: Not all tokens are meant to be forgotten

#10
post #6

Earlier quoted context omitted.

I'd counter with an anecdote; I had a colleague that boasted how he memorized a classmate's SSN in college and would greet him by SSN when seeing him years later. Is the goal of AI to replicate the entirety of the human experience (including social pressures, norms, and shame) or a tool to complement human decision making? While, yes, you can argue the slippery slope, it may be advantageous to flag certain training m…

Knowing more might not improve your accuracy but it's not going to harm it. Forcibly forgetting true parts of your knowledge seems far more likely to have unintended consequences.

One obvious consequence: the model might still produce copyright infringement because it thinks its creative ideas are novel.
Post reply on HN