Live data from Hacker News

A simulation of me: fine-tuning an LLM on 240k text messages

edwarddonner.com

61–70 of 145 posts

Re: A simulation of me: fine-tuning an LLM on 240k text messages

#62

I did a fine tuning and embedding on a large LLM that is based on 50 years of daily journal entries, extensive daily notebooks usually measuring in the hundreds to thousands of words per day, and personal writings across a half-dozen different blogs and websites and various social media feeds. Social media posts and comments (including this one) are also put into my notebooks with a snippet of context about why I pos…

You can’t leave us hanging like that … so what happened. What did you learn? Was it weird reading it? How do you see yourself?

Re: A simulation of me: fine-tuning an LLM on 240k text messages

#63
post #25

Earlier quoted context omitted.

I also get the vibe that digital cloning will be popular. Maybe some extremists will think its unholy and some addicts will lose sense of reality, but for the vast majority of people, I think it's just a user interface - maybe to some particular piece of information in the chatbot-as-librarian role, or puzzle boxes that eventually reveal information once you ask the right question, like Will Smith in I Robot. What I…

> Maybe some extremists will think its unholy At the other side of this equation is the extremists that think this is great. To go full Godwin's law there are certainly people out there that would like for a digital Hitler to stick around forever.

I'm sure people will create holographic Hitlers for historical reasons too.

Re: A simulation of me: fine-tuning an LLM on 240k text messages

#64

I’m far from the first to think of this. Several people — perhaps inspired by creepy Black Mirror episodes — have tried to fine-tune an LLM on their SMS or WhatsApp history in an effort to create a simulation of themselves. It's a much older concept than Black Mirror. Ever since Markov chain IRC bots got popularized in the late 90s and early 2000s, people have been trying to train their virtual doppelgängers. I'm sur…

On the casino I used to run, I started a pilot program with a homemade poker bot (labeled as such, and only deployed on poker tables labeled as "bot friendly"). The bot had no set model of its own. It was designed to mimic specific players on the casino, regulars who had played 10,000+ hands and who agreed to have their history cloned, by ingesting their entire hand/betting history and looking for what they had done…

"On the casino I used to run"? That might be the most casual intro to what sounds like a fascinating corner of the internet I never experienced.

Do you have any other interesting stories or references to that time of your life?

Re: A simulation of me: fine-tuning an LLM on 240k text messages

#65

We need to learn to let go. Chatting with a deceased loved one is basically equivalent to the ressurection stone in Harry Potter. A faint reflection which will drive people to insanity. This is not healthy at all.

This is worrying for me. Imagine if someone, god forbid, is encouraged into suicide because he felt safer because he had 'left something behind'. Or like, wire the llm into his messenger to pretend as if he is still alive...

Re: A simulation of me: fine-tuning an LLM on 240k text messages

#66

Earlier quoted context omitted.

If you want to include fictional antecedents, there is Dixie Flatline from Neuromancer.

There’s also Ubik by Philip K. Dick where the recently deceased are able to communicate with the living AND aside from that also advertise a product.

I looked it up. Ubik was written earlier in 1969, than Neuromancer, which was written in 1984

Re: A simulation of me: fine-tuning an LLM on 240k text messages

#67

I’m far from the first to think of this. Several people — perhaps inspired by creepy Black Mirror episodes — have tried to fine-tune an LLM on their SMS or WhatsApp history in an effort to create a simulation of themselves. It's a much older concept than Black Mirror. Ever since Markov chain IRC bots got popularized in the late 90s and early 2000s, people have been trying to train their virtual doppelgängers. I'm sur…

On the casino I used to run, I started a pilot program with a homemade poker bot (labeled as such, and only deployed on poker tables labeled as "bot friendly"). The bot had no set model of its own. It was designed to mimic specific players on the casino, regulars who had played 10,000+ hands and who agreed to have their history cloned, by ingesting their entire hand/betting history and looking for what they had done…

Half of the work of running a gambling site is making sure all your customers are losers. If they are consistently winning (cheating or not) you want to get rid of them. Unless you are purely running a "pool" of some kind (think Betfair).

Re: A simulation of me: fine-tuning an LLM on 240k text messages

#68

I did a fine tuning and embedding on a large LLM that is based on 50 years of daily journal entries, extensive daily notebooks usually measuring in the hundreds to thousands of words per day, and personal writings across a half-dozen different blogs and websites and various social media feeds. Social media posts and comments (including this one) are also put into my notebooks with a snippet of context about why I pos…

Having done some of this myself, I’m curious your results on fine tuning vs embeddings. I’ve found the latter much more performant, but perhaps I’m thinking about fine tuning wrong.

Re: A simulation of me: fine-tuning an LLM on 240k text messages

#69

Earlier quoted context omitted.

On the casino I used to run, I started a pilot program with a homemade poker bot (labeled as such, and only deployed on poker tables labeled as "bot friendly"). The bot had no set model of its own. It was designed to mimic specific players on the casino, regulars who had played 10,000+ hands and who agreed to have their history cloned, by ingesting their entire hand/betting history and looking for what they had done…

Half of the work of running a gambling site is making sure all your customers are losers. If they are consistently winning (cheating or not) you want to get rid of them. Unless you are purely running a "pool" of some kind (think Betfair).

In poker, all the players are profitable to the house. I think it's up for debate whether it's good for poker rooms to get rid of winners, but if there was a benefit, it would be an indirect benefit, not a direct one.

Re: A simulation of me: fine-tuning an LLM on 240k text messages

#70
post #69

Earlier quoted context omitted.

Half of the work of running a gambling site is making sure all your customers are losers. If they are consistently winning (cheating or not) you want to get rid of them. Unless you are purely running a "pool" of some kind (think Betfair).

In poker, all the players are profitable to the house. I think it's up for debate whether it's good for poker rooms to get rid of winners, but if there was a benefit, it would be an indirect benefit, not a direct one.

It depends how prolific they are. Someone running a winning bot farm is taking more money from the losers than the house would be if the losers kept winning/losing against each other. This assumes that the house wants losers to win alot to stay addicted. If they just get beaten all the time they might quit sooner, they may also run out of money sooner.

A bit like how lottery tickets costing $1 will make you win $1, $2 etc. prizes so you buy another ticket and so increase the revenue (a percentage of which is profit), the same for the house. 2 losers winning money off each other all night means they both lose and gave lots of money to the house. A winner taking the 2 losers money in a fell swoop means a lot less money for the house.

Post reply on HN