Live data from Hacker News

Various LLM Smells

shvbsle.in

151–160 of 312 posts

Re: Various LLM Smells

#151
If you have Claude at work and are willing to point it to your emails, ask it to “read all my sent emails and create a skill to draft an email in my voice.” Even if you don’t want to use the skill, it’s fun to read the skill file it creates. It’s a bizarre feeling, asking Claude “who am I?”

I haven’t tried it with Slack messages because I’m a little scared to read what it says, haha. But the same concept surely applies.

There are a few people at work who are aggressively using Claude to write Slack messages. It’s easy for me to tell because one day they’re writing barely coherent English in multiple messages, and the next they’re sending perfectly coherent prose in a single message.

Re: Various LLM Smells

#152

> The LLM generated writing obviously felt significantly better than my own writing. A general pattern for LLMs is that they look really good at things you are bad at. What that means is that if you find yourself thinking of its output as significantly better than yours in a particular domain, there's a high chance that you are not equipped to judge that quality effectively.

So what does this mean in practice, though? Let's say you are correct. You ask an LLM to write something for you, and to you it looks really, really good. So based on your conjecture, that means I am not a very good writer. Ok, but how does that change what I should do? If I am not a very good writer, that means an LLM IS actually better than me, even if it might not be objectively good to an expert writer. My two ch…

I mean, you have the ability to learn to do stuff better to a certain extent, so it's not like your only choices are "suffer through the writing I'm capable of producing today for the rest of my life" or "give up on ever writing anything myself". Writing stuff yourself is pretty much a requirement of getting better at it, and arguably even if you do intend to use LLMs to supplement it, having a better baseline will be valuable for additional iteration with the LLM.

Re: Various LLM Smells

#153

> The LLM generated writing obviously felt significantly better than my own writing. A general pattern for LLMs is that they look really good at things you are bad at. What that means is that if you find yourself thinking of its output as significantly better than yours in a particular domain, there's a high chance that you are not equipped to judge that quality effectively.

It's basically just another instance of Gell-Mann Amnesia. Ask an LLM to discuss a topic you are an expert on, and you will realise it is full of errors, but ask it to discuss a topic you know nothing about and you will, mysteriously, assume it is very intelligent and correct.

Re: Various LLM Smells

#154

Earlier quoted context omitted.

> A general pattern for LLMs is that they look really good at things you are bad at. This is true for coding, too, which I think, to a large degree, might explain the polarized differences in opinions on HN about the quality of LLM-produced code. You have the 1. "AI produces code better than I could possibly write, one shots things it would take me days to do, and has made me 10X more productive!" camp, and you have…

Eric S. Raymond has basically stopped writing code by hand altogether. He consistently delivers high quality code without intervening to fix the LLM's output himself, much faster than he would have been able to alone. This is very bad news for camp 2 because it means one of three things: 1) he is extraordinarily lucky 2) he is extraordinary brilliant at manipulating LLMs 3) you really are "holding it wrong" and you a…

I'd need to see some transcripts of his conversations with coding agents, to believe this.

Re: Various LLM Smells

#155

Earlier quoted context omitted.

The LinkedIn Kool-Aid predates the advent of LLMs though.

I often think it’s the opposite in fact, that the LLM smells come from LinkedIn text.

This tracks. Linked in, stack overflow, Reddit, even hn. All in the training data, I assume.

Re: Various LLM Smells

#156

The LLM doesn't smell like authentic writing but it does a great job for fast and cheap words. We've gained something similar to fast food. Words made very cheap, very fast, easily digestible, but they have no emotion. In short stints it does have a place in the world.

> it does a great job for fast and cheap words Like corporate manager-type emails, of which I get AI generated ones frequently from company ownership. They think LLMs are the best thing since sliced bread. It's taken corpo-speak to an entirely new level. On the plus side I no longer have to read them, and can just have AI reply on my behalf with more fast food.

Which is fine too me; I never read those emails anyways...

Re: Various LLM Smells

#157

Earlier quoted context omitted.

> who cares if it all looks generally the same? Maybe here lies the crux; for some of us, the web and by extension the internet is about expression and individuality in a way, but all together all accessible by everyone. Everything looking the same instead looks conformist, and ultimately boring, which I guess is what many of us don't want day after day. We want new ideas, presented by the person/group who came up wi…

If you're doing art, do art. Nobody's stopping you. If you're trying to solve a problem for other people, be legible. Legibility and individual expression are goals in tension.

I'm not saying replace the entire thing with a flash animation or something drastically "artsy", just yours.

Re: Various LLM Smells

#159

> The LLM generated writing obviously felt significantly better than my own writing. A general pattern for LLMs is that they look really good at things you are bad at. What that means is that if you find yourself thinking of its output as significantly better than yours in a particular domain, there's a high chance that you are not equipped to judge that quality effectively.

> A general pattern for LLMs is that they look really good at things you are bad at. This is true for coding, too, which I think, to a large degree, might explain the polarized differences in opinions on HN about the quality of LLM-produced code. You have the 1. "AI produces code better than I could possibly write, one shots things it would take me days to do, and has made me 10X more productive!" camp, and you have…

LLMs can generate code, but the quality of the code at scale is just not there currently by all important metrics such as security, maintainability, separation of concerns, etc.

Today, it's a kind of chaos magic wherein you summon the beast and try your best to contain him, knowing that someone will probably die in the process. Sometimes literally. It's still a force multiplier in the right hands and domain, and agentic coding is a paradigm that won't retract, at least until something better supplants it.

The problem is that few engineers actually have the discipline available to constrain these models appropriately and instead rely on a hodgepodge network of "skills" aka prompt fragments which are passed around and glued together.

I consider myself as having such discipline, being strongly architecturally-minded, user-first, etc. in both design and implementation. And I still struggle to contain the beast many days. I just got through screaming at Claude for intentionally taking a shortcut that I'd forbidden, leading to a ton of wasted time and tokens.

Sometimes I feel like I saved weeks of R&D with a single ten-minute task handed off to an agent, other times I feel like I'd get better returns playing slots in Vegas at the alarming rate Claude burns through money.

Re: Various LLM Smells

#160
post #2

No ___, no ____. Just _____ or using "honest" to describe an approach.

Jab, jab, thrust is how I think about that pattern. Or tap tap whack, if you prefer. And it shows up for for positives too: "Smooth. Effortless. A perfect fit for your needs". In any style of informal or persuasive writing this shows up , as if it has to drive the point in. I kind of wish we'd stop talking openly about what the tells are. It's nice to be able to determine with fair accuracy - but it couldn't last for…

Imagine a world where everyone talks like an Apple product page.
Post reply on HN