Live data from Hacker News

Replacing my best friends with an LLM trained on 500k group chat messages

izzy.co

111–120 of 371 posts

Re: Replacing my best friends with an LLM trained on 500k group chat messages

#111
post #49

While I love all these stories of turning your friends and loved ones into chat bots so you can talk to them forever, my brain immediately took a much darker turn because of course it did. How many emails, text messages, hangouts/gchat messages, etc, does Google have of you right now? And as part of their agreement, they can do pretty much whatever they like with those, can't they? Could Google, or any other company…

There is an episode of Black Mirror about this called "Be Right Back". Well worth a watch.

And another called "Hang the DJ"

Re: Replacing my best friends with an LLM trained on 500k group chat messages

#112
post #49

While I love all these stories of turning your friends and loved ones into chat bots so you can talk to them forever, my brain immediately took a much darker turn because of course it did. How many emails, text messages, hangouts/gchat messages, etc, does Google have of you right now? And as part of their agreement, they can do pretty much whatever they like with those, can't they? Could Google, or any other company…

Wait until you head about mind uploading and transhumanism.

Re: Replacing my best friends with an LLM trained on 500k group chat messages

#115
post #49

While I love all these stories of turning your friends and loved ones into chat bots so you can talk to them forever, my brain immediately took a much darker turn because of course it did. How many emails, text messages, hangouts/gchat messages, etc, does Google have of you right now? And as part of their agreement, they can do pretty much whatever they like with those, can't they? Could Google, or any other company…

> input_dir /path/to/downloaded/llama/weights --model_size 7B Most absolutely not with the 7B llama model as described here. …but, potentially, with a much larger fine tuned foundational model, if you have a lot of open source code on GitHub and lots of public samples. The question is why you would bother? very large models would most likely not be meaningfully improved by fine tuning on a specific individual. The on…

ChatGPT’s “voice” changes dramatically in diction and prose when you ask it to generate text in the style of a popular author like Hunter S Thompson, Charles Bukowski, or Terry Pratchett. You can even ask it to generate text in the style of a specific HN user if they’re prolific enough in the training data set.

Fine tuning would allow you to achieve that for people who aren’t notable enough to be all over the training data

Re: Replacing my best friends with an LLM trained on 500k group chat messages

#116

Earlier quoted context omitted.

> Could Google, or any other company out there, build a digital copy of you that answers questions exactly the way you would? "Hey, we're going to cancel the interview- we found that you aren't a good culture fit here in 72% of our simulations and we don't think that's an acceptable risk." If a company is going to snoop in your personal data to get insights about you, they'd just do it directly. Hiring managers would…

> Adding an LLM abstraction layer doesn't make the existing laws (or social/moral pressure) go away. Isn't the "abstraction" of "the model" exactly the reason we have open court filings against stable diffusion and other models for possibly stealing artist's work in the open source domain and claiming it's legal while also being financially backed by major corporations who are then using said models for profit? Whose…

> for possibly stealing artist's work in the open source domain

The provenance of the training set is key. Every LLM company so far has been extremely careful to avoid using people's private data for LLM training, and for good reason.

If a company were to train an LLM exclusively on a single person's private data and then use that LLM to make decisions about that person, the intention is very clearly to access that person's private data. There is no way they could argue otherwise.

Re: Replacing my best friends with an LLM trained on 500k group chat messages

#117
post #49

While I love all these stories of turning your friends and loved ones into chat bots so you can talk to them forever, my brain immediately took a much darker turn because of course it did. How many emails, text messages, hangouts/gchat messages, etc, does Google have of you right now? And as part of their agreement, they can do pretty much whatever they like with those, can't they? Could Google, or any other company…

some 10y+(?) ago, i had an idea of building a model/graph of one's own notions (what is "red"?) and how do they relate to one another, and to others' such graphs - from your perspective... Back then abandoned it because looked like impossibly-huge to build, semantical-web was only leftovers-and-promises.. But a year later, this same thing you are talking about, occured to me and then i abandoned it completely. Crossed it out. Yeah, That thing would be extremely useful but even more dangerous as it will know more about you than you.

btw. There's a new book by Kazuo Ishiguro - named Klara and the sun. Along the same vein - Have a look.

https://en.wikipedia.org/wiki/Klara_and_the_Sun

ciao

p.s. see also Lena by Charles Stross. or this: https://www.antipope.org/charlie/blog-static/2023/01/make-up...

Re: Replacing my best friends with an LLM trained on 500k group chat messages

#118
post #77

Earlier quoted context omitted.

> And as part of their agreement, they can do pretty much whatever they like with those, can't they? No, they definitely can't. Parts of HN love to hate on GDPR, but laws like that prevent companies from doing the things you proposed.

They are supposed to, but usually it takes several dozen times of them getting caught with their hands in the cookie jar and fined before they are even capable of acknowledging these laws even exist.

You know, that's not the sentiment I've been experiencing in the industry. There's certainly some uncertainty and risk-taking on the margins, e.g. what exactly constitutes "fair use", how do design user consent flows, and so on. But it's broadly accepted that you can't do anything with personal data without user consent, and I've found companies to be very careful in that regard.

Recently, Meta was fined $400MM for forcing users to consent to targeted advertising [0]. Note how Meta was careful to get consent (even if the way they did it was illegitimate). Sure, $400MM may not be a lot for a company that size, but I genuinely believe that the fines would be an order of magnitude higher if a company intentionally decided to do something with personal data without consent. GDPR fines may reach up to 4% of worldwide revenue, plus likely any proceeds from the illegitimate venture.

[0] https://www.cnbc.com/2023/01/04/meta-fined-more-than-400-mil...

Re: Replacing my best friends with an LLM trained on 500k group chat messages

#119
post #49

While I love all these stories of turning your friends and loved ones into chat bots so you can talk to them forever, my brain immediately took a much darker turn because of course it did. How many emails, text messages, hangouts/gchat messages, etc, does Google have of you right now? And as part of their agreement, they can do pretty much whatever they like with those, can't they? Could Google, or any other company…

There’s a (good) short story in a (great) sci-fi book called “Valuable Humans in Transit” about something like this. If I remember correctly, Google learns to perfectly simulate people and then the real people disappear.

Re: Replacing my best friends with an LLM trained on 500k group chat messages

#120
post #49

While I love all these stories of turning your friends and loved ones into chat bots so you can talk to them forever, my brain immediately took a much darker turn because of course it did. How many emails, text messages, hangouts/gchat messages, etc, does Google have of you right now? And as part of their agreement, they can do pretty much whatever they like with those, can't they? Could Google, or any other company…

> Could Google, or any other company out there, build a digital copy of you that answers questions exactly the way you would? "Hey, we're going to cancel the interview- we found that you aren't a good culture fit here in 72% of our simulations and we don't think that's an acceptable risk." If a company is going to snoop in your personal data to get insights about you, they'd just do it directly. Hiring managers would…

> Hiring managers would scroll through your e-mails and make judgment calls based on their content.

> Training an LLM on your e-mails and then feeding it questions is just a lower accuracy, more abstracted version of the above, but it's the same concept.

Its also one that once you have cheap enough computing resources scales better, because you don't need to assign literally any time from your more limited pool of human resources to it. Yes, baroque artisanal manual review of your online presence might be more “accurate” (though there's probably no applicable objective figure of merit), but megacorporate hiring filters aren't about maximizing accuracy they are about efficiently trimming the applicant pool before hiring managers have to engage with it.

Post reply on HN