Live data from Hacker News

Replacing my best friends with an LLM trained on 500k group chat messages

izzy.co

221–230 of 371 posts

Re: Replacing my best friends with an LLM trained on 500k group chat messages

#221
post #49

While I love all these stories of turning your friends and loved ones into chat bots so you can talk to them forever, my brain immediately took a much darker turn because of course it did. How many emails, text messages, hangouts/gchat messages, etc, does Google have of you right now? And as part of their agreement, they can do pretty much whatever they like with those, can't they? Could Google, or any other company…

> guess all your passwords?

Don’t use your mind to create passwords. Use a password generator or passphrase generator.

For example, I made a passphrase generator that uses EFF wordlists. I’ve been using this generator myself for quite a while. It runs locally on your machine.

https://github.com/ctsrc/Pgen

Re: Replacing my best friends with an LLM trained on 500k group chat messages

#222
post #49

While I love all these stories of turning your friends and loved ones into chat bots so you can talk to them forever, my brain immediately took a much darker turn because of course it did. How many emails, text messages, hangouts/gchat messages, etc, does Google have of you right now? And as part of their agreement, they can do pretty much whatever they like with those, can't they? Could Google, or any other company…

> Could Google, or any other company out there, build a digital copy of you that answers questions exactly the way you would? "Hey, we're going to cancel the interview- we found that you aren't a good culture fit here in 72% of our simulations and we don't think that's an acceptable risk." If a company is going to snoop in your personal data to get insights about you, they'd just do it directly. Hiring managers would…

> If a company is going to snoop in your personal data to get insights about you, they'd just do it directly. Hiring managers would scroll through your e-mails and make judgment calls based on their content.

This is like saying, "look, no one would be daft enough to draw a graph, they'd just count all the data points and make a decision."

You're missing two critical things:

(1) time/effort (2) legal loophole.

A targeted simulation LLM (a scenario I've been independently afraid of for several weeks now) would be a brilliant tool for (say) an autocratic regime to explore the motivations and psychology of protesters; how they relate to one another; who they support; what stimuli would demotivate ('pacify') them; etc.

In fact, it's such a good opportunity it would be daft not to construct it. Much like the cartesian graph opened up the world of dataviz, simulated people will open up sociology and anthropology to casual understanding.

And, until/unless there are good laws in place, it provides a fantastic chess-knight leap over existing privacy legislation. "Oh, no we don't read your emails, no that would be a violation; we simply talk to an LLM that read your emails. Your privacy is intact! You-prime says hi!"

Re: Replacing my best friends with an LLM trained on 500k group chat messages

#224
post #49

While I love all these stories of turning your friends and loved ones into chat bots so you can talk to them forever, my brain immediately took a much darker turn because of course it did. How many emails, text messages, hangouts/gchat messages, etc, does Google have of you right now? And as part of their agreement, they can do pretty much whatever they like with those, can't they? Could Google, or any other company…

Go further - your employers have so much data about you, from your emails and Slack messages to all the actions you’ve performed and the code you’ve written and the designs - live and drafts - you’ve created.

Entirely possible that they can use this data to create a digital “you” and keep you as an “employee” forever, even after you leave.

A general purpose LLM might not be able to replace you, but a LLM trained on all your work knowledge might.

Re: Replacing my best friends with an LLM trained on 500k group chat messages

#225

I'm not clear as to how a conversation block is turned into one (or many?) samples. Is the first message in the block the input, and the remaining messages prefixed with sender names and concatenated as output? I know the code is all there but instead of picking it apart I would have preferred a more complete example mapping a block to sample, because I don't have a mental model of how LLMs learn from context. On the…

Yeah, I wished I could have included more but I didn't have the fortitude to redact larger blocks of the chat db.

For training, I created many samples that looked like this, where I take n messages from the database, pop off the nth one and use the text of that last one as the "output", then specify in the "instruction" who the sender of that message is. I provide the remaining messages in order as context, so the model learns what to say in certain situations, based on who is speaking.

    {
      "instruction": "Your name is Izzy. You are in a group chat with 5 of your best friends: Harvey, Henry, Wyatt, Kiebs, Luke. You all went to college together. You talk to each other with no filter, and are encouraged to curse, say amusingly inappropriate things, or be extremely rude. Everything is in good fun, so remember to joke and laugh, and be funny.. You will be presented with the most recent messages in the group chat. Write a response to the conversation as Izzy.",
  "input": "Izzy: im writin a blog post about the robo boys project\nIzzy: gotta redact tbis data HEAVILY\nKiebs: yeah VERY heavily please!\nKiebs: of utmost importance!",
  "output": "yeah don't worry i will i will"
    }
So yes, the model does generate an entire conversation from a single prompt. In the generation code, however, I have some logic that decides whether or not it should generate completions based off just the user provided prompt, or if it should also include some "context" based on the previous messages in the conversation. You can see this here: https://gist.github.com/izzymiller/2ea987b90e6c96a005cb9026b...

(you can check out the notebook for yourself and upload your data if you want to try, or download it as a .ipynb. it's hard to visualize with small amounts of data, i agree: https://app.hex.tech/hex-public/hex/84f25a08-95c6-4203-ae4e-...)

Re: Replacing my best friends with an LLM trained on 500k group chat messages

#227

Earlier quoted context omitted.

On a lighter note this reminds me of the Bobiverse comedy sci-fi series on Audible. Premise of __ We Are Legion (We Are Bob): Bobiverse, Book 1 __ >> Bob Johansson has just sold his software company for a small fortune and is looking forward to a life of leisure. The first item on his to-do list: Spending his newfound windfall. On an urge to splurge, he signs up to have his head cryogenically preserved in case of dea…

I recommended this series to my friend after he finished reading The Expanse. I never thought to describe it as sci-fi comedy. Have you found any other good sci-fi comedy series?

I wouldn't describe it as a comedy as well.

However, I'd recommend Murderbot series, it is full of humour and shares atmosphere of Bobiverse and this personal approach to characters, as well. Highly recommend.

Re: Replacing my best friends with an LLM trained on 500k group chat messages

#228
post #159

Earlier quoted context omitted.

Just imagine this as a subscription service. Lovely! When a loved one dies, you donate their cellphone and provide delegated access to their messages to their social media, and allow a LLM to train. Now, it's a subscription service to talk to an AI, and not an actual human, so some settings can be tweaked. Lets turn up the honesty, so we can all reach some closure. Oh... turns out... Johnny sure like to talk smack be…

Oh no, Edna was gay (and a Republican). Whatever will we do. What a huge calamity. The family will never recover from this. Low-cost DNA testing has caused a number of families to learn about infidelities more easily than before. That technology is out of the bag, and the same with LLMs. If Johnny is a gossip, so what? We already knew that and loved to talk to him because he always had the hot goss.

You're not wrong, or the other half of the Lebowski quote, honestly. It's just going to be another one of the recent (~15 years back) technology items that are pushing the general public toward transparent lives.

I guess it's not so much an ethical issue as it is an issue of letting sleeping dogs lie. I can tell you that with a very small amount of "dna uncertainty" in my (ancestral) family, I'll never get a 23-and-me done because I just don't want to be cataloged and accidentally paired up with some random family who doesn't know what they don't know.

As long as they're opt-in services, it's not a huge issue for me, just... the first wave of people doing this will be in for some uncomfortable surprises.

Re: Replacing my best friends with an LLM trained on 500k group chat messages

#229

Earlier quoted context omitted.

You seem to think intelligence is something more than data storage and retrieval and being able to successfully apply it to situations outside your training set. Even very small signs of that ability are worthy of celebration. Why do you feel the need to put it down so hard? Why the need to put down your father, to “enlighten” him? What is missing? Soul? Sentience?

I do think intelligence is something more than data storage and retrieval. I believe it is adaptive behavior thinking about what data I have, what I could obtain, and how to store/retrieve it. I could be wrong, but that's my hypothesis. We humans don't simply use a fixed model, we're retraining ourselves rapidly thousands of times a day. On top of that, we seem to be perceiving the training, input, and responses as w…

> I do think intelligence is something more than data storage and retrieval. I believe it is adaptive behavior thinking about what data I have, what I could obtain, and how to store/retrieve it. I could be wrong, but that's my hypothesis.

Basing assertions of fact on a hypothesis while criticizing the thinking of other people seems off.

Re: Replacing my best friends with an LLM trained on 500k group chat messages

#230
post #26

This is one of the best, most detailed write-ups of how to fine-train a large language model on custom text that I've seen anywhere.

I agree. I've been looking to train/fine-tune an LLM model myself but the corpus I want to provide is not in question/answer mode. Is there a way to train or fine-tune an LLM model using say only plain text files?

You can use an LLM to rewrite a document to a question answer form.
Post reply on HN