Live data from Hacker News

The last six months in LLMs in five minutes

simonwillison.net

441–450 of 631 posts

Re: The last six months in LLMs in five minutes

#441
post #412

Earlier quoted context omitted.

Hey that's fine. You're free to make whatever judgment you wish. But I still stand by the quality of my code, including here. You and I don't need to agree. What decades of managing codebases (public and private, huge and small) has taught me is that there will always be an endless list of bugs and feature ideas and nice-to-haves and technical debt pressures in any given project. You'll never get to them all, so you…

tbh youve embarrassed yourself here.

THAT must be why the stars are going up!

Thanks for explaining it for me.

Re: The last six months in LLMs in five minutes

#443

Earlier quoted context omitted.

I don't know what you're talking about. His tool wraps Claude and breaks the TUI. What's so hard to understand? That's valid critique. What world have I woke up in today?

To be honest I assumed it was the screencap software running a basic terminal env without bells and whistles that CC needs, which I've seen before. If the actual tool functions like that too, that's not great. That said, it works for them, it works for them.

But earlier:

> The question is why are you so eager to give critique on unrelated work, appearing in a demo screencap, to someone who didn't produce it?

I guess the question was actually, why were you so eager to critique a critique based on a false assumption?

I wish people would be careful what they support with their rhetoric.

Re: The last six months in LLMs in five minutes

#444

Earlier quoted context omitted.

Deeply troubling for so many reasons. Please try to get her to stop.

What's the problem? If I enjoy some show, material or text, if it brings me value or a brief moment of happiness, I could care less if it was made by an AI or a human. This racism against AI-generated stuff has to stop. If not, we'll have a butlerian jihad on our hands that will set back prosperity, development and science for decades, perhaps centuries. People mention the artists... ohh, boohoo... either do it on yo…

I think we need to start separating such concepts like entertainment from the ones of enjoyment, fascination, function, interest, satisfaction, beauty and the sublime a bit more. Art theory literally has books on these things, as they all fall under the topic of aesthetics. Do you really enjoy a frozen pizza from the oven at home in the same way as a freshly made pizza from an authentic pizza oven?

I always care about the processes involved, especially if any human work is involved, from all its accuracies to its errors. For me, interesting things happen while we balance our understandings with a certain amount of holism and a certain amount of reductionism. Putting it on either side of the scale, like your holistic statements, is just pure ideology, and that doesn't hold any merit in reality and is honestly just bland, repetitive and boring.

Re: The last six months in LLMs in five minutes

#445

I asked Gemini for a video of 'pelican riding a unicycle in hyde park' - I was blown away by the output: https://gemini.google.com/share/55e250c99693

According to OP:

> Why this test? Because pelicans are hard to draw, bicycles are hard to draw, pelicans can’t ride bicycles... and there’s zero chance any AI lab would train a model for such a ridiculous task.

At this juncture I'm left wondering why competing AI labs wouldn't train for this now well known "test".

Re: The last six months in LLMs in five minutes

#446
post #220

Earlier quoted context omitted.

you are experiencing reverse Dunning–Kruger effect. For someone that just dabbled in coding prior, it went from AI building 80%, and struggling through to finish the 20% when trying to build an app/website. now it's like 97% and struggling with last 3%. Yes it'll look rough around the edges when evaulated by a senior dev, but being able to build MVP level things to completion with ease helps you stay engaged and moti…

Please do not cite Dunning–Kruger effect at random. Who needs to generate a dumb demo of a 97% done crud app? We had code generators for those, everytime I read claims like that and I ask to explain further I then discover it's people who were not productive before generating the so called "MVP level things to completion with ease". If you're trying to solve a HARD problem people REALLY have, it's a novelty that agen…

I'm beginning to get the sense that Sturgeon's Law is at play here and the non-crap 10% of us are arguing with the 90% for whom LLM's shitty output is actually better than what they could do on their own.

I've been lucky enough to work at places with majority intelligent engineers with similar tastes on quality to my own... but it seems to be that's not the norm or the case everywhere.

and it's the 90% that's most vocal. Sturgeon and D-K seen to go hand-in-hand.

Re: The last six months in LLMs in five minutes

#447

Earlier quoted context omitted.

I don't see how it cannot be true. Are you claiming that every developer who uses the same LLM harness + model would produce equal code, regardless of the prompt? That's clearly not true in my experience, and I cannot understand how it could be either. And if that's not true, then it's quite literally about how you're holding this hammer.

There's a cowboy artist that paints with his penis and does amazing work. If I tried that it'd turn out incredibly poorly, I prefer to paint with paintbrushes. Just because the naked cowboy can paint well with just his penis, doesn't mean a penis is the right tool for painting. It doesn't matter how you hold your penis, it's not the right tool.

> There's a cowboy artist that paints with his penis and does amazing work. If I tried that it'd turn out incredibly poorly, I prefer to paint with paintbrushes.

I can't decide which joke to make, either (little dick joke) "well yeah you'd have to be able to see your paintbrush in order to use it" or (big dick joke) "well yeah, if you can't even hold it in two hands, how are you supposed to paint with it?" so I'll just make both :-D

Re: The last six months in LLMs in five minutes

#448
post #376
post #367

Earlier quoted context omitted.

> Yes, but no room is made for people who see no use for it. There is a forced-consensus that this technology is useful, which I have to combat against at work. This is the crux of the issue -- The technology is useful. Using it appropriately is probably the thing that people are ignoring, but you're conflating one and the other in your comment. It is not useful to you in this case, and complain that it is an overall…

I mean only that I see no use for it myself, in my own work. I'm sure there are people working in roles around me who believe they get some use out of AI doing their work for them, and they will have to answer to auditors when they find problems with their work, or when someone is killed. To me, as a non-techie person, it feels as if people who work in software believe that because their work can be done by AI, every…

I don't mean to tar you with a too-wide brush, and I feel like you have a good handle on your personal acceptance for LLM assistance. No complaint there.

I do think, maybe alternative to your view, that LLMs can provide useful feedback to graduate-level employees in most fields.

It is not that the work can be done by LLMs -- we're not there, yet, in software or otherwise -- but that LLMs as useful tutors specifically in regard to denouncing known bad ideas is largely applicable all over.

What I mean by the above is that I have yet to find a truly interesting idea spun from whole cloth by an LLM. They're mediocre at it. They're trained from the aggregate thoughts of those in every industry, and you and I both know that the aggregate of the industry is, generally, mediocre.

Conversely, though, is the hit: They won't be worse than mediocre. An indefatigable tutor who gives no great advice but will counsel you against blowing yourself up (or cutting a limb off with a rope, or falling overboard) is, to me, worth an amount.

The failure modes will get better, the advice will get better. Are we there, now? Unsure. You can tell us all better.

On the ten year horizon, I'd place a bet, though.

Re: The last six months in LLMs in five minutes

#449
post #390

Earlier quoted context omitted.

Yes, it's a mystery, isn't it? Specifically for CSS, these bots really want to just barf out tailwind-style crap. If you deviate even slightly from the standards and practices of the modal front-end developer, you quickly see how these things are brittle, and no amount of prompting and cajoling will truly affect their behavior. In this case, you're kind of seeing the downstream affects of saying "no, do NOT do tailwi…

> these bots really want to just barf out tailwind-style crap. I get it. The LLMs struggle most with state. They don’t have a real fix for that yet. People generally compensate by shoving everything into context, and making the context window as large as possible, which half-works. Tailwind happens to be “stateless” CSS framework. Nothing uses anything else, nothing is shared, nothing is reused, nothing stacks. It’s…

Maybe you're right, but honestly, I think they're just barfing out the median content related to UI development that you find on reddit and stackoverflow and github.

For better or worse, web UI development has descended down a dark rabbit hole of bad code over the last decade, and so that is what LLMs were trained on. GIGO.

Re: The last six months in LLMs in five minutes

#450
post #449

Earlier quoted context omitted.

> these bots really want to just barf out tailwind-style crap. I get it. The LLMs struggle most with state. They don’t have a real fix for that yet. People generally compensate by shoving everything into context, and making the context window as large as possible, which half-works. Tailwind happens to be “stateless” CSS framework. Nothing uses anything else, nothing is shared, nothing is reused, nothing stacks. It’s…

Maybe you're right, but honestly, I think they're just barfing out the median content related to UI development that you find on reddit and stackoverflow and github. For better or worse, web UI development has descended down a dark rabbit hole of bad code over the last decade, and so that is what LLMs were trained on. GIGO.

Yeah, it quickly becomes a "which came first, the chicken or the egg?" thing between junior devs and LLMs.

I will say, if you use a Mistral model, and if you insist your CSS framework is Bulma (tell it, 'no tailwind', 'no preprocessor'), it does okay at staying away from Tailwind. (Not perfect, not great, but okay).

No LLM I've used can handle raw CSS well (yet). If you are carefully curating your own classes and styles, you might just be on your own for a bit.

Post reply on HN