Live data from Hacker News

What we know about LLMs

willthompson.name

161–170 of 173 posts

Re: What we know about LLMs

#161
post #127

Earlier quoted context omitted.

RLHF is not part of LLaMa pretraining, or pretraning of any other models for that matter. RLHF comes after pretraining. https://twitter.com/Jeande_d/status/1661833563069620247/phot...

I see, that’s my misunderstanding I was grouping all training as pretraining

pre-training is developing the language model's base understanding of conditional word probabilities.

SFT and RLHF is attempting to further guide the model in terms of steerability + alignment of output.

In fact, the InstructGPT authors were worried about losing the pre-trained model's underlying probability distribution, so they try a version where it penalizes the model deviating too significantly from the original distribution (using KL). I don't remember them seeing a significant difference in performance.

Re: What we know about LLMs

#163

"Crypto VCs & ”builders” making a hard left into AI" This is a humorous intro graphic caption, but this sentiment appears on here constantly and it's self-destructive. This response might seem a bit over the top to a funny graphic, but I am replying to the general "ha ha AI like crypto amirite?" sentiment that is incredibly boring and worn out. When confronted with challenging new technology that we don't understand,…

On the contrary, I agree - while there is certainly hype being generated around AI, particularly generated by the "VC hype cycle", the fundamental advancements we've made with LLMs are quite real.

Part of the reason I wrote this is to separate the signal from the noise and why one should be {cautiously, more tempered} optimistic in the medium term.

Re: What we know about LLMs

#164
post #111

"Crypto VCs & ”builders” making a hard left into AI" This is a humorous intro graphic caption, but this sentiment appears on here constantly and it's self-destructive. This response might seem a bit over the top to a funny graphic, but I am replying to the general "ha ha AI like crypto amirite?" sentiment that is incredibly boring and worn out. When confronted with challenging new technology that we don't understand,…

> "ha ha AI like crypto amirite?" I don’t think that was the meaning at all. I think the image was supposed to convey how the crypto grifters and con artists were veering into AI to run scams under the guise of AI.

^^ yes

Re: What we know about LLMs

#165

ChatGPT was announced November, 2022 - 8 months ago. Time flies. Question for HN: Where are we in the hype cycle on this? We can run shitty clones slowly on Raspberry Pi's and your phone. The educational implementations demonstrate the basics in under a thousand lines of brisk C. Great. At some point you have to wonder... well, so what? Not one killer app has emerged. I for one am eager to be all hip and open minded…

An artist friend of mine with no programming knowledge used ChatGPT to produce a variety of cool visuals for a music gig, in Processing - spinning wireframes, bobbing cube grids, that sort of thing. They didn't even know they needed to use Processing at first - ChatGPT told them everything. They had an aesthetic in mind, and ChatGPT helped them deliver.

It's a revolution, and it's here.

Re: What we know about LLMs

#166
post #124

Earlier quoted context omitted.

> Not one killer app has emerged. I think the microsoft gpt integration on Office is probably that app. Ability to ask to have your email's summarised, or getting your excel sheets formulas configured with natural language, etc are increidbly useful tools to lower the floor of entry to tools that already speed up humans so much. I don't think the use of this tools is some life redefining feature, but a friend of mine…

There was a science fiction story about this, with phone auto-message and auto-answer systems connecting with each other long after all the humans were dead.

Can you remember the story? It sounds thematically similar to "There Will Come Soft Rains", but the details don't match.

https://en.wikipedia.org/wiki/There_Will_Come_Soft_Rains_(sh...

Re: What we know about LLMs

#167
post #52

Earlier quoted context omitted.

> the traffic to ChatGPT is decreasing and has been for over two months now This seems entirely unsurprising, and isn’t by itself enough to support your general thesis. Interacting with these LLMs was extremely novel for most people when the tech first dropped, and those earlier months were the peak of the viral growth/expansion into public awareness. As the novelty dies down, it’s not surprising that there would be…

> Early on, I had all sorts of ridiculous conversations just to see what would happen. [...] That transition points to this being the opposite of a toy - after the fun dies down, the real work begins. The "intelligence" behind it is too unpredictable to be reliable for work, and using it for fun is about as amusing as emailing HR.

> The "intelligence" behind it is too unpredictable to be reliable for work

This highly depends on the kind of work you’re doing. It’s great as a starting point for exploratory learning, helpful for some coding tasks, and useful for summarizing text.

As I work on a writing project that benefits from all of these use cases, it’s a good tool.

Not so great if you’re trying to write legal briefs.

> using it for fun is about as amusing as emailing HR

All due respect, but you’re either doing it wrong, or you’ve encountered some hilarious HR departments.

Re: What we know about LLMs

#168
post #52

Earlier quoted context omitted.

> the traffic to ChatGPT is decreasing and has been for over two months now This seems entirely unsurprising, and isn’t by itself enough to support your general thesis. Interacting with these LLMs was extremely novel for most people when the tech first dropped, and those earlier months were the peak of the viral growth/expansion into public awareness. As the novelty dies down, it’s not surprising that there would be…

> Early on, I had all sorts of ridiculous conversations just to see what would happen. [...] That transition points to this being the opposite of a toy - after the fun dies down, the real work begins. The "intelligence" behind it is too unpredictable to be reliable for work, and using it for fun is about as amusing as emailing HR.

Ask it to speak in cockney as an 18th century barker trying to convince you to buy a lame horse or to continue the conversation in brolish as though you were two surfer dudes sitting on the beach and then just ask it anything you want like “explain modern monetary theory”. If you enjoy fiction then get it to help world build a new setting and then act out a scene with you playing one character and it playing the rest.

To get it to stay in character use the custom instructions feature to set the requirements.

Re: What we know about LLMs

#169

ChatGPT was announced November, 2022 - 8 months ago. Time flies. Question for HN: Where are we in the hype cycle on this? We can run shitty clones slowly on Raspberry Pi's and your phone. The educational implementations demonstrate the basics in under a thousand lines of brisk C. Great. At some point you have to wonder... well, so what? Not one killer app has emerged. I for one am eager to be all hip and open minded…

There isn't one. AI is pretty much the new blockchain. It is unfair to say it's a scam because it's more of a delusion. https://en.m.wikipedia.org/wiki/AI_winter This concept of hype and decline has been happening for literally decades. Yet people don't realize it even when it's literally on the first google page for anything to do with AI. The people spouting this AI nonsense seriously need to fuck off and read a bo…

> This article needs to be updated. The reason given is: Add more information about post-2018 developments in artificial intelligence leading to the current AI boom.

I don't quite remember anything existing and comparable over the last few decades to LLMs like ChatGPT/Claude

Re: What we know about LLMs

#170

Earlier quoted context omitted.

> Where are we in the hype cycle on this? Can we stop acting like the Gartner "hype cycle" is anything more than a marketing gimmick created Gartner to validate their own consulting/research services? While you can absolutely find cases that map to the "hype cycle", there is nothing whatsoever to validate this model as remotely accurate or valid for describing technology trends. Where is crypto in the "hype cycle"? I…

> The "hype cycle" is a neat idea but doesn't really map to reality in a way that makes it useful. What do you propose as a more accurate alternative, or do you think that the whole idea should be scrapped? Because personally I feel like certain tech/practices certainly go through multiple stages, where initially people expect too much from them and eventually figure out what they're good for and what they're not. No…

I think the more accurate alternative is that reality is messy and you can’t always put things in categories and models.
Post reply on HN