Live data from Hacker News

Things we learned about LLMs in 2024

simonwillison.net

71–80 of 615 posts

Re: Things we learned about LLMs in 2024

#71

My fav part of the writeup at the end: """ LLMs need better criticism # A lot of people absolutely hate this stuff. In some of the spaces I hang out (Mastodon, Bluesky, Lobste.rs, even Hacker News on occasion) even suggesting that “LLMs are useful” can be enough to kick off a huge fight. I like people who are skeptical of this stuff. The hype has been deafening for more than two years now, and there are enormous quan…

This happens with every inane hype-cycle.

I suspect people don't particularly hate or despise LLMs per se. They're probably reacting mostly to "tech industry" boom-bust bullsh*tter/guru culture. Especially since the cycles seem to burn increasingly hotter and brighter the less actual, practical value they provide. Which is supremely annoying when the second-order effect is having all the oxygen (e.g. capital) sucked out of the room for pretty much anything else.

Re: Things we learned about LLMs in 2024

#72
post #65

I think John Gruber summed it up nicely: https://daringfireball.net/2024/12/openai_unimaginable OpenAI’s board now stating “We once again need to raise more capital than we’d imagined” less than three months after raising another $6.6 billion at a valuation of $157 billion sounds alarmingly like a Ponzi scheme — an argument akin to “Trust us, we can maintain our lead, and all it will take is a never-ending stream of…

What is funny is that their "lead" is just because of inertia - they were the first to make an LLM publicly available. But they are no longer leaders so their attempts at getting more and more money only prove Altman's skills at convincing people to give him money.

Re: Things we learned about LLMs in 2024

#73
post #50
post #37

Spookily good at writing code? LLMs frequently hallucinate broken nonsense shit when I use them. Recognize what they do well (generate simple code in popular languages) while acknowledging where they are weak (non-trivial algorithms, any novel code situation the LLM hasn't seen before, less popular languages).

Did you try learning HOW to get good code out of them? As with all things LLM there's a whole lot of undocumented and under appreciated depth to getting decent results. Code hallucinations are also the least damaging type of hallucinations, because you get fact checking for free: if you run the code and get an error you know there's a problem. A lot of the time I find pasting that error message back into the LLM gets…

> Code hallucinations are also the least damaging type of hallucinations, because you get fact checking for free: if you run the code and get an error you know there's a problem.

This is great when the error is a thrown exception, but less great when the error is a subtle logic bug that only strikes in some subset of cases. For trivial code that only you will ever run this is probably not a big deal—you'll just fix it later when you see it—but for code that must run unattended in business-critical cases it's a totally different story.

I've personally seen a dramatic increase in sloppy logic that looks right coming from previously-reliable programmers as they've adopted LLMs. This isn't an imaginary threat, it's something I now have to actively think about in code reviews.

Re: Things we learned about LLMs in 2024

#74
post #50
post #37

Spookily good at writing code? LLMs frequently hallucinate broken nonsense shit when I use them. Recognize what they do well (generate simple code in popular languages) while acknowledging where they are weak (non-trivial algorithms, any novel code situation the LLM hasn't seen before, less popular languages).

Did you try learning HOW to get good code out of them? As with all things LLM there's a whole lot of undocumented and under appreciated depth to getting decent results. Code hallucinations are also the least damaging type of hallucinations, because you get fact checking for free: if you run the code and get an error you know there's a problem. A lot of the time I find pasting that error message back into the LLM gets…

> if you run the code and get an error you know there's a problem.

well, sometimes - other times it'll be wrong with no error, or insecure, or inaccessible, and so on

Re: Things we learned about LLMs in 2024

#75
post #54

About "people still thinking LLMs are quite useless", I still believe that the problem is that most people are exposed to ChatGPT 4o that at this point for my use case (programming / design partner) is basically a useless toy. And I guess that in tech many folks try LLMs for the same use cases. Try Claude Sonnet 3.5 (not Haiku!) and tell me if, while still flawed, is not helpful. But there is more: a key thing with L…

While Claude Sonnet is superior than 4o for most my use cases, there are still occasionally some specific tasks where it performs slightly better.

Probably. But statistically to work with 4o is a lose of time for me. LLMs is like an investment: you write the prompts, you "work" with them. If the LLM is too weak, this is a lose of time. You need to have a return on the investment that is positive. With ChatGPT 4o / o1 most of the times for me the investment of time has almost zero return. Before Claude Sonnet 3.5 I already had a ChatGPT PRO account but never used it for coding since it was most of the times useless if not for throw away scripts that I didn't want to do myself or as a stack overflow replacement for trivial stuff. Now it's different.

Re: Things we learned about LLMs in 2024

#76

In spite of all this progress, I can't find LLMs that solve simple tasks like: Here is my resume. Make it look nice (some design hints). They can spit html and css, but not Google doc. On the other hand, Google results are dominated by SEO spam. You can probably find one usable result on page 10. The problem is not technology. It's a business model that can support the humans feeding data into the LLM.

Why would they be able to output a Google doc? It's a proprietary format. The closest thing would be rich text format to copy paste.

That proprietary format is owned by a company associated with folks who won two nobel prizes for AI related work this year and the employer at the time of the researchers who wrote the attention is all you need paper and also the owner of a search engine with access to like, all the data. Doesn't seem unreasonable lol

Re: Things we learned about LLMs in 2024

#77
post #30

Earlier quoted context omitted.

It would be really cool if big tech could find a new hyperscaler model that didn't also require offsetting the goals of green energy projects worldwide. Between LLM and crypto you'd swear they're trying to find the most energy-wasteful tech possible.

Cryptocurrency, at least PoW, the point is indeed to be the most wasteful — a literal Dyson swarm powered Bitcoin would provide exactly the same utility as the BTC network already had in 2010. LLMs (and the image, sound, and movie generating models) are more coincidentally power-hogs — people are at least trying to make them better at fixed compute, and lower compute at fixed quality.

I mean, I appreciate that distinction and don't disagree. And, if this is going to continue being a trend, I think we need more stringent restrictions on what sorts of resources are permitted to be consumed in the power plants that are constructed to meet the needs of hyperscaler data centers.

Because whether we're using tons of compute to provide value or not doesn't change that we are using tons of compute and tons of compute requires tons of energy, both for the chips themselves, and the extensive infrastructure that has to built around them to let them work. And not just electricity: refrigerants, many of which are environmentally questionable themselves, are a big part; hell, just water. Clean, usable water.

If we truly need these data centers, then fine. Then they should be powered by renewable energy, or if they absolutely cannot be, then the costs their nonrenewable energy sources inflict on the biosphere should be priced into their construction and use, and in turn, priced into the tech that is apparently so critical for them to have.

This is like, a basic calculus that every grown person makes dozens of times a day: do I need this? And they don't get to distribute the cost of that need, however prescient it may be, on their wider community because they can't afford it otherwise. I don't see why Microsoft should be able to either. If this is truly the tech of the future as it is constantly propped up to be, cool. Then charge a price for it that reflects what it costs to use.

Re: Things we learned about LLMs in 2024

#78

Earlier quoted context omitted.

I agree, but I think my biggest issue with LLMs (and a lot of GenAI) is that they act as a massive accelerator for the WORST (and unfortunately most common) type of human - the lazy one. The signal-to-noise ratio just goes completely out of control. https://journal.everypixel.com/ai-image-statistics

Exif watermark by the generators would solve 90% of the problem in one fell swoop because lazy people won't remove it

Every image host and social media app automatically strips EXIF data (for privacy reasons at minimum).

Re: Things we learned about LLMs in 2024

#79
post #54

About "people still thinking LLMs are quite useless", I still believe that the problem is that most people are exposed to ChatGPT 4o that at this point for my use case (programming / design partner) is basically a useless toy. And I guess that in tech many folks try LLMs for the same use cases. Try Claude Sonnet 3.5 (not Haiku!) and tell me if, while still flawed, is not helpful. But there is more: a key thing with L…

I’m surprised you only have one use case. I use LLMs to research travel, adjust recipes, check biographies and book reviews, and many many more things.

Re: Things we learned about LLMs in 2024

#80
post #56

Earlier quoted context omitted.

By that definition machine to machine communication that happens "organically" (like how humans do it, where they sometimes strike up conversations unprompted with each other) is "slop". You're not seeing how the future of the world will develop.

If you ask me to read an unguided conversation between two LLMs then yes, I'd consider that slop. Some people might like slop.

The rise of the famous obvious Facebook AI slop indicates that some demographics love it.
Post reply on HN