Live data from Hacker News

Things we learned about LLMs in 2024

simonwillison.net

151–160 of 615 posts

Re: Things we learned about LLMs in 2024

#151
post #54

About "people still thinking LLMs are quite useless", I still believe that the problem is that most people are exposed to ChatGPT 4o that at this point for my use case (programming / design partner) is basically a useless toy. And I guess that in tech many folks try LLMs for the same use cases. Try Claude Sonnet 3.5 (not Haiku!) and tell me if, while still flawed, is not helpful. But there is more: a key thing with L…

I swear these goalposts keep getting moved, I remember being told that GPT3.5 is a useless toy but the paid GPT4 is lifechanging, and now that GPT4 is free I'm told that it's a useless toy but paid o1 or paid Sonnet are lifechanging. Looking forward to o1 and Sonnet becoming useless toys, unlike the lifechanging o3.

Re: Things we learned about LLMs in 2024

#152

Earlier quoted context omitted.

What is funny is that their "lead" is just because of inertia - they were the first to make an LLM publicly available. But they are no longer leaders so their attempts at getting more and more money only prove Altman's skills at convincing people to give him money.

Which models perform better than 4o or o1 for your use cases? In my limited tests (primarily code) nothing from llama or Gemini have come close, Claude I’m not so sure about.

How good is the best model of your choice at doing architecture work for complex and nontrivial apps?

I have been bashing my head against the wall over the course of the past few days trying to create my (quite complex) dream app.

Most of LLM coding I've done involved in writing code to interface with already existing libs or services and the LLMs are great at that.

I'm hung up on architecture questions that are unique to my app and definitely not something you can google.

Re: Things we learned about LLMs in 2024

#153
post #4

Don’t forget that 2024 was also a record year for new methane power plant projects. Some 200 new projects in the US alone and I’d wager most of them are funded directly by big tech for AI data centres. https://www.bnnbloomberg.ca/investing/2024/09/16/ai-boom-is-... This is definitely extending the runway of O&G at a crisis point in the climate disaster when we’re supposed to be reducing and shutting down these power…

The only thing that will stop this is for battery storage to get cheap and available enough that it can cover for renewables. If we are still building gas turbines it means that hasn’t happened yet. AI is a red herring. If it wasn’t that it would be EV power demand. If it wasn’t that it would be reshoring of manufacturing. If it wasn’t that it would be population growth from immigration. If it wasn’t that it would be…

Pumped hydro is an excellent form of storage if you have the terrain for it. A whole order of magnitude cheaper than battery storage at the moment.

Re: Things we learned about LLMs in 2024

#154
post #89

Earlier quoted context omitted.

What is funny is that their "lead" is just because of inertia - they were the first to make an LLM publicly available. But they are no longer leaders so their attempts at getting more and more money only prove Altman's skills at convincing people to give him money.

yeah but in business there are really only 2 skills right? Convincing people to give you money and giving them something back to them thats worth more than the money they gave you.

For repeated business you want to give them something that costs you less than what they pay, but is worth more to them than what they pay. Ie creating economic value.

Re: Things we learned about LLMs in 2024

#155
post #136
post #54

About "people still thinking LLMs are quite useless", I still believe that the problem is that most people are exposed to ChatGPT 4o that at this point for my use case (programming / design partner) is basically a useless toy. And I guess that in tech many folks try LLMs for the same use cases. Try Claude Sonnet 3.5 (not Haiku!) and tell me if, while still flawed, is not helpful. But there is more: a key thing with L…

I don't think people finding LLMs useless is a good representation of the general sentiment though. I feel that more than anything, people are annoyed at LLM slop. Someone uses an LLM too much to write code, they create "slop," which ends up making things worse.

Yes but then they can prompt it to golf the code and most of the slop goes away. This sometimes breaks the code.

Re: Things we learned about LLMs in 2024

#156
Double checking, I don't think I saw anything about video generation. Not sure if those fall under the "LLM" umbrella. It came very late in the year, but the Google Veo 2 limited testing are astounding. There are at least a half-dozen other services where you can pay to generate video.

Re: Things we learned about LLMs in 2024

#157

Earlier quoted context omitted.

Which models perform better than 4o or o1 for your use cases? In my limited tests (primarily code) nothing from llama or Gemini have come close, Claude I’m not so sure about.

How good is the best model of your choice at doing architecture work for complex and nontrivial apps? I have been bashing my head against the wall over the course of the past few days trying to create my (quite complex) dream app. Most of LLM coding I've done involved in writing code to interface with already existing libs or services and the LLMs are great at that. I'm hung up on architecture questions that are uniq…

Don't wanna be that typical hackernews guy but I couldnt resist... if your app is "quite complex" there is probably a way or ways you can break it down into much simpler parts. Easier for you AND the LLM. It always comes back to architecture and composition ;)

Re: Things we learned about LLMs in 2024

#158
post #65

I think John Gruber summed it up nicely: https://daringfireball.net/2024/12/openai_unimaginable OpenAI’s board now stating “We once again need to raise more capital than we’d imagined” less than three months after raising another $6.6 billion at a valuation of $157 billion sounds alarmingly like a Ponzi scheme — an argument akin to “Trust us, we can maintain our lead, and all it will take is a never-ending stream of…

Every waste of money is not a Ponzi scheme.

I agree, the core aspect of a ponzi scheme is that it redistributes the newly invested funds to previous investors, making it highly profitable to anyone joining early and incentivising early joiners to get new investors.

This just doesn't hold true for open ai

Re: Things we learned about LLMs in 2024

#159

Earlier quoted context omitted.

I believe that AGI cannot be exponential for long because any intelligent agent can only approach nature's limits asymptotically. The first company with AGI will be about as much ahead as, say, the first company with electrical generators [1]. A lot of science fiction about a technological singularity assumes that AGI will discover and apply new physics to develop currently-believed-impossible inventions, but I don't…

I don't recall editing my message, but HN can be wonky sometimes. :) Nothing is truly exponential for long, but the logistic curve could be big enough to do almost anything if you get imaginative. Without new physics, there are still some places where we can do some amazing things with the equivalent of several trillion dollars of applied R&D, which AGI gets you.

This depends on what a hypothetical 'AGI' actually costs. If a real AGI is achieved, but it costs more per unit of work than a human does... it won't do anyone much good.

Re: Things we learned about LLMs in 2024

#160
post #131
post #18

Earlier quoted context omitted.

Subsidised by whom?

E.g. tax payers.

Are tax payers subsiding that particular activity of Google or Amazon? If they do, “they make enough money” to cover costs. If they don’t, how does it become profitable if it doesn’t even cover the cost of one of the inputs?
Post reply on HN