Live data from Hacker News

Things we learned about LLMs in 2024

simonwillison.net

521–530 of 615 posts

Re: Things we learned about LLMs in 2024

#521

Earlier quoted context omitted.

Its easy to see how it does that, the answer is that your bug isn't something novel, it has seen millions of "where is the bug in this code" questions online so it can typically guess from there what it would be. It is very unreliable at fixing things or writing code for anything non standard. Knowing this you can easily construct queries that trips them up by noticing what it is in your code they notice, so you cons…

Both of your claims are way off the mark (I run an AI lab). The LLMs are good at finding bugs in code not because they’ve been trained on questions that ask for existing bugs, but because they have built a world model in order to complete text more accurately. In this model, programming exists and has rules and the world model has learned that. Which means that anything nonstandard … will be supported. It is trivial…

The "world model" of an LLM is just the set of [deep] predictive patterns that it was induced to learn during training. There is no magic here - the model is just trying to learn how to auto-regressively predict training set continuations.

Of course the humans who created the training set samples didn't create them auto-regressively - the training set samples are artifacts reflecting an external world, and knowledge about it, that the model is not privy to, but the model is limited to minimizing training errors on the task it was given - auto-regressive prediction. It has no choice. The "world model" (patterns) it has learnt isn't some magical grokking of the external world that it is not privy to - it is just the patterns needed to minimize errors when attempting to auto-regressively predict training set continuations.

Whether these training set predictive patterns result in the model performing as you might hope on an unseen text depends on the similarity of that text to samples in the training set.

Re: Things we learned about LLMs in 2024

#523
post #54

About "people still thinking LLMs are quite useless", I still believe that the problem is that most people are exposed to ChatGPT 4o that at this point for my use case (programming / design partner) is basically a useless toy. And I guess that in tech many folks try LLMs for the same use cases. Try Claude Sonnet 3.5 (not Haiku!) and tell me if, while still flawed, is not helpful. But there is more: a key thing with L…

Definitely not a "useless toy" with the right use case. It's great at code snippets, scripts, etc. It's an assistant.

Re: Things we learned about LLMs in 2024

#524
post #522

Interestingly, there isn't much big news about jail breaking or safety alignment

Was there much big news around that in 2024?

There were a few interesting papers - the Anthropic one about alignment faking https://www.anthropic.com/news/alignment-faking and the OpenAI o1 system card https://simonwillison.net/2024/Dec/5/openai-o1-system-card/ - and OpenAI continued to push their "instruction hierarchy" idea, any other big moments?

I'll be honest, I don't follow that side of things very closely (outside of complaining that prompt injection still isn't fixed yet).

Re: Things we learned about LLMs in 2024

#525
post #234

Great summary of highlights. Don't agree with all, but I think it's a very sound attempt at a year in review summary >LLM prices crashed This one has me a little spooked. The white knight on this front (DS) has both announced increases and has had staff poached. There is still Gemini free tier which is ofc basically impossible to beat (solid & functionally unlimited/free) but it's google so reluctant to trust. Seriou…

Agents have a definition issue sure, but IMO we are prevented from even discovering a useful definition by the current limitations of LLMs

Re: Things we learned about LLMs in 2024

#526
post #495
post #380

Earlier quoted context omitted.

(Not original commenter) “Staff” engineer is typically one of the most senior and highest paid engineer titles in very large tech company. “Staff plus” is implying they are the best of the best.

I’ve seen your comment below, but you did specify big tech as context in this parent comment, no? Or is „very large tech company“ not FAANG? Google has Staff at L6, and their ladder goes up to L11. Apple‘s Staff pendant is ICT5, which is below ICT6 and Distinguished. Amazon has E7-E9 above Staff, if you count E6 as Staff. Netflix very recently departed from their flat hierarchy and even they have Principal above Staf…

> Amazon has E7-E9 above Staff

Few clarifications:

Amazon labels levels with "L" rather than "E". Engineering levels are L4 -- L10. Weirdly enough, level L9 does not exist at Amazon. L8 (Director / Senior Principal Engineer) is promoted directly to L10 (VP / Distinguished Engineer)

Re: Things we learned about LLMs in 2024

#527
If I learned anything, it would be that LLMs' non-deterministic nature makes them great are generating output that we can argue over, but they are not a great tool. for doing actual work. I am not asking for much. In my field of work, I use Jetbrains' IDEs, which have now been "enhanced" with AI. I had to turn this feature off, because I kept having to remove code, imports randomly added by the IDE. This was distracting and wasted my time.

Re: Things we learned about LLMs in 2024

#528
post #210

Earlier quoted context omitted.

You will also be dismayed to hear that a 2011 iPhone is no longer state-of-the-art, and indeed can't run most modern apps.

Holy false-equivalency, Batman! The definitions of "useless toy / lifechanging tool" are _not_ changing over time (or, at least, not over the timescale being explored here), whereas the expectations and requirements of processing power of a phone are.

But in fact they are changing over time -- this is an expectations treadmill. When you get something newer and better, it highlights the flaws in what you had before.

Re: Things we learned about LLMs in 2024

#529

Earlier quoted context omitted.

When I'm done working, chased my children to properly finish their dinner, helped my son with homework, and putting them to bed, it's already 9+ PM — the only time of the day when I have free time. Just which human besides my wife can I talk to at that point? What if she doesn't have a clue either? All the professionals are only open when I'm working. A lot of the issues happen during the weekend, when professionals…

You are supposed to have connections with knowledgeable people, so you can call them and ask for advice. That's how it works without computers.

Did you miss the parts where I said that I only have time when they're closed, and they're only open when I'm most busy?

Have you never seen knowledgeable people get things wrong, and having to verify them?

Did you miss the part where they cost money, and I better come in as prepared as possible?

I really don't get these knee-jerk averse reactions. Are people deliberately reading past my assertions that I double check LLM outputs for everything critical?

Re: Things we learned about LLMs in 2024

#530

Earlier quoted context omitted.

I'm guessing that mindset is what cause some people to find this scary. I see a new tool and opportunities. Like all tools, it has drawbacks and caveats, but when wielded properly, it can give me more choice. I suspect some others focus too much on flaws and don't bother looking for opportunities. They are expecting a holy grail: if it's not perfect then it's useless. It's like people who proclaim that Linux as a who…

Grandma has a reason to care about you. At the opposite, my trust of Russian / Chinese / USian platforms is low enough that I consider it my duty to publicly shame people that still use them in 2025. (With some caveats of course, for instance HN is not a yet negative to the world. Yet.) There's also the question of stickiness of habits : your grandmas are for life, human professionals you might have a shallow enough…

You view Github and LLMs as traps that deliberately give you malicious advice or even brainwash you into addiction? If you view things that way then it's no surprise that you are averse to LLMs (and Github). But frankly I find that entire view to be absurd and overly cynical.
Post reply on HN