Live data from Hacker News

Things we learned about LLMs in 2024

simonwillison.net

261–270 of 615 posts

Re: Things we learned about LLMs in 2024

#261

Earlier quoted context omitted.

It's justified if AGI is possible. If AGI is possible, then the entire human economy stops making sense as far as money goes, and 'owning' part of OpenAI gives you power. That is of course, assuming AGI is possible and exponential, and that marketshare goes to a single entity instead of a set of entities. Lots of big assumptions. Seems like we're heading towards a slow-lackluster singularity though.

I was thinking about how the economy has been actively makes less sense and gets divorced more and more from reality year after year, AI or not. It's the simple fact that the ability of assets to generate wealth has far outstripped the abiliy of individuals to earn money by working. Somehow real estate has become so expensive everywhere that owning a shitty apartment is impossible for the vast majority. When the worl…

Somehow real estate has become so expensive everywhere that owning a shitty apartment is impossible for the vast majority.

That's to be expected when governments forbid people from building housing. The only thing I find surprising is when people blame this on "capitalism".

Re: Things we learned about LLMs in 2024

#262
post #235
post #54

About "people still thinking LLMs are quite useless", I still believe that the problem is that most people are exposed to ChatGPT 4o that at this point for my use case (programming / design partner) is basically a useless toy. And I guess that in tech many folks try LLMs for the same use cases. Try Claude Sonnet 3.5 (not Haiku!) and tell me if, while still flawed, is not helpful. But there is more: a key thing with L…

I'm surprised at the description that it's "useless" as a programming / design partner. Even if it doesn't make "elegant" code (whatever that means), it's the difference between an app existing at all, or not. I built and shipped a Swift app to the App Store, currently generating $10,200 in MRR, exclusively using LLMs. I wouldn't describe myself as a programmer, and didn't plan to ever build an app, mostly because in…

To the un-sticking point: it's also great at letting people ask questions without being perceived as dumb

Tragically - admitting ignorance, even with the desire to learn, often has negative social reprocussions

Re: Things we learned about LLMs in 2024

#263

Earlier quoted context omitted.

> If AGI is possible, then the entire human economy stops making sense as far as money goes, What does this mean in terms of making me coffee or building houses?

If we can simulate a full human intelligence at a reasonable speed, we can simulate 100 of them and ask the AGI to figure out how to make itself 10x faster. Rinse and repeat. That is exponential take off. At the point where you have an army of AIs running at 1000x human speed it can just ask it to design the mechanisms for and write the code to make robots that automate any possible physical task.

There are about 8 billion human intelligences walking around right now and they've got no idea how to begin making even a stupid AGI, let alone a superhuman one. Where does the idea that 100 more are going to help come from?

Re: Things we learned about LLMs in 2024

#264
post #202

Earlier quoted context omitted.

> Try Claude Sonnet 3.5 (not Haiku!) and tell me if, while still flawed, is not helpful. It's not as helpful as Google was ten years ago. It's more helpful than Google today, because Google search has slowly been corrupted by garbage SEO and other LLM spam, including their own suggestions.

Claude Sonnet 3.5 can write whole React applications with proper contextual clues and some minor iterations. Google has never coded for you. I’ve written two large applications and about a dozen smaller ones using Claude as an assistant. I’m a terrible front-end developer and almost none of that work was possible without Claude. The API and AWS deployment were sped up tremendously. I’ve created unit tests and I’ve re…

I've never really used Claude for writing code, becuase I'm not really bottlenecked by that problem. I have used it quite a bit for asking questions about what code to write and it's almost always wrong (usually in subtle ways that would trick someone with little experience).

Maybe it was overtrained on react sources, but for me it's pretty useless.

The big annoyance for me is it just makes up APIs that don't exist. While that's useful for suggesting to me what APIs I should add to my own code, it's really pointless if I ask a question like "using libfoo how do I bar" and it tells me "call the doBar() function" which does not exist.

Re: Things we learned about LLMs in 2024

#265
post #234

Great summary of highlights. Don't agree with all, but I think it's a very sound attempt at a year in review summary >LLM prices crashed This one has me a little spooked. The white knight on this front (DS) has both announced increases and has had staff poached. There is still Gemini free tier which is ofc basically impossible to beat (solid & functionally unlimited/free) but it's google so reluctant to trust. Seriou…

The biggest reason I'm not worried about prices going back up again is Llama. The Llama 3 models are really good, and because they are open weight there are a growing number of API providers competing to provide access to them.

These companies are incentivized to figure out fast and efficient hosting for the models. They don't need to train any models themselves, their value is added entirely in continuing to drive the price of inference down.

Groq and Cerberus are particularly interesting here because WOW they serve Llama fast.

Re: Things we learned about LLMs in 2024

#267
post #235

Earlier quoted context omitted.

I'm surprised at the description that it's "useless" as a programming / design partner. Even if it doesn't make "elegant" code (whatever that means), it's the difference between an app existing at all, or not. I built and shipped a Swift app to the App Store, currently generating $10,200 in MRR, exclusively using LLMs. I wouldn't describe myself as a programmer, and didn't plan to ever build an app, mostly because in…

To the un-sticking point: it's also great at letting people ask questions without being perceived as dumb Tragically - admitting ignorance, even with the desire to learn, often has negative social reprocussions

Asking "stupid" questions without fear of judgement is legit one of my favorite personal applications of LLMs.

Re: Things we learned about LLMs in 2024

#268
post #259

Earlier quoted context omitted.

Why do people have such narrow views on what makes LLMs useful? I use them for basically everything. My son throwing an irrational tantrum at the amusement park and I can't figure out why he's like that (he won't tell me or he doesn't know himself either) or what I should do? I feed Claude all the facts of what happened that day and ask for advice. Even if I don't agree with the advice, at the very least the analysis…

Yes people have lives outside of coding, but most people are able to manage without having AI software intercede in as much of their lives as possible. It seems like you trust AI more than people and prefer it to direct human interaction. That seems to be satisfying a need for you that most people don't have.

Why do you postulate that "most people don't have" this need? I also use AI non-stop throughout my day for similar uses.

This feels identical to when I was an early "smart phone" user w/my palm pilot. People would condescend saying they didn't understand why I was "on it all the time". A decade or two later, I'm the one trying to get others to put down their phones during meetings.

My take? Those who aren't using AI continually currently are simply later adopters of AI. Give it a few years - or at most a decade - and the idea of NOT asking 100+ AI queries per day (or per hour) will seem positively quaint.

Re: Things we learned about LLMs in 2024

#269
post #231
post #198

> There’s a flipside to this too: a lot of better informed people have sworn off LLMs entirely because they can’t see how anyone could benefit from a tool with so many flaws. The key skill in getting the most out of LLMs is learning to work with tech that is both inherently unreliable and incredibly powerful at the same time. This is a decidedly non-obvious skill to acquire! I wish the author qualified this more. How…

One of the things I find most frustrating about LLMs is how resistant they are to teaching other people how to use them! I'd love to figure this out. I've written more about them than most people at this point, and my goal has always been to help people learn what they can and cannot do - but distilling that down to a concise set of lessons continues to defeat me. The only way to really get to grips with them is to u…

Thank you for doing this work, though.

My first stab at trying ChatGPT last year was asking it to write some Rust code to do audio processing. It was not a happy experience. I stepped back and didn't play with LLMs at all for a while after that. Reading your posts has helped me keep tabs on the state of the art and decide to jump back in (though with different/easier problems this time).

Re: Things we learned about LLMs in 2024

#270
post #54

About "people still thinking LLMs are quite useless", I still believe that the problem is that most people are exposed to ChatGPT 4o that at this point for my use case (programming / design partner) is basically a useless toy. And I guess that in tech many folks try LLMs for the same use cases. Try Claude Sonnet 3.5 (not Haiku!) and tell me if, while still flawed, is not helpful. But there is more: a key thing with L…

Why do people have such narrow views on what makes LLMs useful? I use them for basically everything. My son throwing an irrational tantrum at the amusement park and I can't figure out why he's like that (he won't tell me or he doesn't know himself either) or what I should do? I feed Claude all the facts of what happened that day and ask for advice. Even if I don't agree with the advice, at the very least the analysis…

At the risk of sounding impolite or critical of your personal choices: this, right here, is the problem!

You don’t understand how medicine works, at any level.

Yet you turn to a machine for advice, and take it at face value.

I say these things confidently, because I do understand medicine well enough to not to seek my own answers. Recently I went to a doctor for a serious condition and every notion I had was wrong. Provably wrong!

I see the same behaviour in junior developers that simply copy-paste in whatever they see in StackOverflow or whatever they got out of ChatGPT with a terrible prompt, no context, and no understanding on their part of the suitability of the answer.

This is why I and many others still consider AIs mostly useless. The human in the loop is still the critical element. Replace the human with someone that thinks that powdered rhino horn will give them erections, and the utility of the AI drops to near zero. Worse, it can multiply bad tendencies and bad ideas.

I’m sure someone somewhere is asking DeepSeek how best to get endangered animals parts on the black market.

Post reply on HN