Live data from Hacker News

Things we learned about LLMs in 2024

simonwillison.net

371–380 of 615 posts

Re: Things we learned about LLMs in 2024

#371
post #274
post #251

Earlier quoted context omitted.

Yeah, a key thing to understand about LLMs is that managing the context is everything . You need to know when to wipe the slate by starting a new chat session and then pasting across a subset of the previous conversation. A lot of my most complex LLM interactions take place across multiple sessions - and in some cases I'll even move the project from Claude 3.5 Sonnet to OpenAI o1 (or vice versa) to help get out of a…

What kinds of things do you with these LLMs? I feel like I’m good at understanding context. I’ve been working in AI startups over the last 2 years. Currently at an AI search startup. Managing context for info retrieval is the name of the game. But for my personal use as a developer, they’ve caused me much headache. Answers that are subtly wrong in such a way that it took me a week to realize my initial assumption bas…

There is no substitute for cold hard facts. LLMs do not provide that unless it’s literally the easiest thing for them to do and even then not always.

In the case you were in I would go out of my way to feed the docs to the LLM and then use the LLM to interrogate the docs and then verify the understanding I got from the LLM with a personal reading of the docs that were relevant.

You might think it takes just as long of not longer to do it my way rather than just reading the docs myself. Sometimes it can. But as you get good at the workflow you find that the time sien finding the relevant docs goes down and you get an instant plausible interpretation of the docs added too. You can then very quickly produce application code right away and then docs of the code you write.

Re: Things we learned about LLMs in 2024

#372

Earlier quoted context omitted.

Like all stubborn anti-AI know-it-alls, you sound like you’ve tried a couple of times to do something and have decided to label all LLMs with the same brush. What models have you tried, and what are you trying to do with them? Give us an example prompt too so we can see how you’re coaxing it so we can rule out skill issue. And a big strength LLMs have is summarizing things - I’d like to see you summarize the latest 1…

> And a big strength LLMs have is summarizing things - I’d like to see you summarize the latest 10 arxiv papers relating to prompt engineering and produce a report geared towards non-techies. And do this every 30 mins please. Also produce social media threads with that info. Is this a task you could do yourself, better than LLMs? I don't mean to nitpick, but how good do you really think the output of this would be? P…

The quality of the summary is only as good as the effort you put into writing your workflow. If you’re simply one shotting the paper into a message and saying “plz summarise this and I’ll reward you with $1m” then of course it’s gonna be shit. But if you semantically chunked along sections and do some RAG Q&A summaries before combining into a well formatted schema then it’s probably going to be better than the first way.

I’m using the summaries as a juicier abstract. I’m not taking them as gospel.

I’m working on following references to then add those papers to a vector db for RAG so it can actually go the step beyond. It’s fun!

Re: Things we learned about LLMs in 2024

#374
post #92
post #54

About "people still thinking LLMs are quite useless", I still believe that the problem is that most people are exposed to ChatGPT 4o that at this point for my use case (programming / design partner) is basically a useless toy. And I guess that in tech many folks try LLMs for the same use cases. Try Claude Sonnet 3.5 (not Haiku!) and tell me if, while still flawed, is not helpful. But there is more: a key thing with L…

> best LLMs are able to accelerate you https://www2.math.upenn.edu/~ghrist/preprints/LAEF.pdf - this math textbook was written in just 55 days! Paraphrasing the acknowledgements - ...Begun November 4, 2024, published December 28, 2024. ...assisted by Claude 3.5 sonnet, trained on my previous books... ...puzzles co-created by the author and Claude ...GPT-4o and -o1 were useful in latex configurations...doing proof-rea…

just to clarify - I have nothing to do with this book. I was just forwarded a copy and I thought its relevant to the topic at hand. from the wild swings in karma, looks like people are annoyed with the message and shooting down the messenger.

Re: Things we learned about LLMs in 2024

#375

Earlier quoted context omitted.

This is one of those things I like about Claude. I’m hitting my 40th year as a professional software developer and architect. I’ve written thousands of blocks of code from scratch. It gets boring. But then in the 2000’s me (and everyone else) started building code generators, often from ERD structures, but also UML designs. These tools were massively useful and (initially) reduced costs. The future balls of mud probl…

How are you measuring your productivity?

In many cases I have no frame of reference for the expected code, like React and css. Typescript is perfectly readable, but I’m not really a script kiddie, so I’d go very slow on the React tsx files. The services are probably a slightly faster set of work, especially if I always have unit tests.

If someone was an expert React+TypeScript programmer with decent css knowledge the productivity may be a marginal improvement.

But I haven’t been a full-time programmer in ten years.

Re: Things we learned about LLMs in 2024

#376
post #367
post #365

Earlier quoted context omitted.

> Code is the best possible application of LLMs because you can TEST the output. This is an overly simplistic view of software development. Poorly made abstractions and functions will have knock on effects on future code that can be hard to predict. Not to mention that code can have side effects that may not affect a given test case, or the code could be poorly optimized, etc. Just because code compiles or passes a t…

So code review LLM-generated code and reject it (or require changes to it) if it doesn't fit your idea of what good code looks like.

Or… yknow… I could just write the code…

Instead of going through a multi step process to get an LLM to generate it, review it, reject it, and repeat…

I wonder why you reply to these comments, but not my other asking what you use LLMs for and specifically explaining how they failed me.

Re: Things we learned about LLMs in 2024

#377

Earlier quoted context omitted.

If we can simulate a full human intelligence at a reasonable speed, we can simulate 100 of them and ask the AGI to figure out how to make itself 10x faster. Rinse and repeat. That is exponential take off. At the point where you have an army of AIs running at 1000x human speed it can just ask it to design the mechanisms for and write the code to make robots that automate any possible physical task.

There are about 8 billion human intelligences walking around right now and they've got no idea how to begin making even a stupid AGI, let alone a superhuman one. Where does the idea that 100 more are going to help come from?

This was my argument a long time ago. The common counter was that we’d have a bunch of geniuses that knew tons of research. Well, we probably already have millions of geniuses. If anything, they use their brains for self-enrichment (eg money, entertainment) or on a huge assortment of topics. If all the human geniuses didn’t do it, then why would the AGI instances do it?

We also have people brilliant enough to maybe solve the AGI problem or cause our extinction. Some are amoral. Many mechanisms pushed human intelligences in other directions. They probably will for our AGI’s assuming we even give them all the power unchecked. Why are they so worried the intelligent agents will not likewise be misdirected or restrained?

What smart, resourceful humans have done (and not done) is a good, starting point for what AGI would do. At best, they’ll probably help optimize some chips and LLM runtimes. Patent minefields with sub-28nm design, especially mask-making, will keep unit volumes of true AGI’s much lower at higher prices than systems driven by low-paid workers with some automation.

Re: Things we learned about LLMs in 2024

#378
post #267

Earlier quoted context omitted.

To the un-sticking point: it's also great at letting people ask questions without being perceived as dumb Tragically - admitting ignorance, even with the desire to learn, often has negative social reprocussions

Asking "stupid" questions without fear of judgement is legit one of my favorite personal applications of LLMs.

Yes! In the time it would take to organize a question in a form that won’t be downvoted/closed on StackOverflow you can ask a whole series of LLM questions and learn quite a bit.

Re: Things we learned about LLMs in 2024

#379
post #267

Earlier quoted context omitted.

To the un-sticking point: it's also great at letting people ask questions without being perceived as dumb Tragically - admitting ignorance, even with the desire to learn, often has negative social reprocussions

Asking "stupid" questions without fear of judgement is legit one of my favorite personal applications of LLMs.

I find myself doing this all the time, as an experienced dev.

All the little nooks of missing knowledge are now very easy to fill in.

Re: Things we learned about LLMs in 2024

#380
post #309

Earlier quoted context omitted.

The context here is super-important - the commenter is the author of Redis. So, a super-experienced and productive low-level programmer. It’s not surprising that Staff-plus experts find LLMs much less useful. Though I’d be interested if this was an opinion on “help me write this gnarly C algorithm” or “help me to be productive in ” as I find a big productivity increase from the latter.

antirez is clearly going to be “Staff-plus” for almost any definition. Can you clarify what you mean?

(Not original commenter) “Staff” engineer is typically one of the most senior and highest paid engineer titles in very large tech company. “Staff plus” is implying they are the best of the best.
Post reply on HN