Live data from Hacker News

A simulation of me: fine-tuning an LLM on 240k text messages

edwarddonner.com

141–145 of 145 posts

Re: A simulation of me: fine-tuning an LLM on 240k text messages

#141
post #140

Earlier quoted context omitted.

> The idea that reading a piece of text constitutes copyright infringement is ridiculous. No man, it's not ridiculous. If I write a program that copies someone's book and try to sell it I'm infringing on that copyright. I cannot sell a zipped version of the Harry Potter books. I feel like there's so many people weighing in on this discussion who haven't actually done any real world copyright related stuff.

I see the source of your confusion. LLMs are not actually zips of the training dataset.

It is an incredibly common refrain amongst experts in the field that it is a compression of the dataset.

But that doesn't matter, because you clearly didn't understand what I was writing which doesn't shock me considering your position on LLMs.

Re: A simulation of me: fine-tuning an LLM on 240k text messages

#142
post #140

Earlier quoted context omitted.

I see the source of your confusion. LLMs are not actually zips of the training dataset.

It is an incredibly common refrain amongst experts in the field that it is a compression of the dataset. But that doesn't matter, because you clearly didn't understand what I was writing which doesn't shock me considering your position on LLMs.

> It is an incredibly common refrain amongst experts in the field that it is a compression of the dataset

This was a common idea three years ago. No one in the field seriously believes this today.

Re: A simulation of me: fine-tuning an LLM on 240k text messages

#143
post #142

Earlier quoted context omitted.

It is an incredibly common refrain amongst experts in the field that it is a compression of the dataset. But that doesn't matter, because you clearly didn't understand what I was writing which doesn't shock me considering your position on LLMs.

> It is an incredibly common refrain amongst experts in the field that it is a compression of the dataset This was a common idea three years ago. No one in the field seriously believes this today.

You're going to have to let a lot of scientists know that, because they're still publishing papers with that understanding. I guess they should have consulted you first.

Re: A simulation of me: fine-tuning an LLM on 240k text messages

#144

Earlier quoted context omitted.

There were two Netflix miniseries related to the book, btw. Exploring the surviving humans on the planet.

Ah, if you mean Into The Night, I enjoyed that, although the link to the old axolotl wasnt entirely obvious!

Yes! And there was a second mini series connected to this one!

Yakamoz S-245

Re: A simulation of me: fine-tuning an LLM on 240k text messages

#145
post #142

Earlier quoted context omitted.

> It is an incredibly common refrain amongst experts in the field that it is a compression of the dataset This was a common idea three years ago. No one in the field seriously believes this today.

You're going to have to let a lot of scientists know that, because they're still publishing papers with that understanding. I guess they should have consulted you first.

Not their fault, hindsight is 20/20.
Post reply on HN