Live data from Hacker News

Everything around LLMs is still magical and wishful thinking

dmitriid.com

281–290 of 377 posts

Re: Everything around LLMs is still magical and wishful thinking

#281
post #266

Earlier quoted context omitted.

> That makes no sense. You're either click-baiting or have an inconsistent stance. Or you haven't actually understood a single thing from the article. Like the fact you say "aren't we technologists" immediately followed by a completely unverifiable claim that everyone is supposed to take at face value, uncritically. And any criticism is immediately labeled as "crap", "people who can't see the future" etc. Remind you…

You had a fine criticism in the article: that people's claims of success or failure should be accompanied by contextual information, to prevent bad extrapolations. But then you title it "Everything around LLMs is still magical and wishful thinking." Was this tongue in cheek? Hyperbole for effect? Because it seems you're saying that when the data comes out, people will realize this was all a fraud. There's a big diffe…

"People make huge outrageous claims without backing them up with anything except 'trust me bro', and people uncritically accept that as truth" would probably make a better title.

But if you can't go past the title, well, you might have bigger problems.

Re: Everything around LLMs is still magical and wishful thinking

#282

Earlier quoted context omitted.

Making a comparison to crypto is lazy criticism. It’s not even worth validating. It’s people who want to take the negative vibe from crypto and repurpose it. The two technologies have nothing to do with each other, and therefore there’s clearly no reason to make comparative technical assessments between them. That said, the social response is a trend of tech worship that I suspect many engineers who have been around…

wrong. ultimately, crypto is information science. mathematically, cryptography, compression, and so on (data transmission) are all the "same" problem. LLMs compress knowledge, not just data, and they do it in a lossy way. traditional information science work is all about dealing with lossless data in a highly lossy world.

And it's all powered by electricity. Coincidence? I think not.

Re: Everything around LLMs is still magical and wishful thinking

#283
post #4

Earlier quoted context omitted.

Can you elaborate on your situation? Which country are you in? How is crypto used there?

I am a Russian immigrant in Switzerland. As of right now, all Swiss banks block all Russian bank accounts until their owners can provide a valid physical residence permit card, due to sweeping sanctions (meanwhile, Russian-owned companies continue to freely trade crude oil from here, as they use Swiss nominal directors — the hypocrisy is through the roof). My residence permit is on renewal now and the case is being d…

Thank you for indulging my curiosity. I really hope your (and our world's) situation improves soon.

Re: Everything around LLMs is still magical and wishful thinking

#284

I'm a retired programmer. I can't imagine trusting code generated by probablities for anything mission critical. If it were close and just needed minor tweaks I could understand that. But I don't have experience with it. My comment is mainly to say LLMs are amazing in areas that are not coding, like brainstorming, blue sky thinking, filling in research details, asking questions that make me reflect. I treat the LLM l…

You don't trust the code coming out of the probabilistic machine. You build a validation cage around it with hard interfaces and you also review the output.

Re: Everything around LLMs is still magical and wishful thinking

#285
post #194

I'm a retired programmer. I can't imagine trusting code generated by probablities for anything mission critical. If it were close and just needed minor tweaks I could understand that. But I don't have experience with it. My comment is mainly to say LLMs are amazing in areas that are not coding, like brainstorming, blue sky thinking, filling in research details, asking questions that make me reflect. I treat the LLM l…

I say this as an LLM skeptic. All code, including stuff that we experienced coders write is inherently probabilistic. That’s why we have code reviews, unit tests, pair programming, guidelines and guardrails in any critical project. If you’re using LLM output uncritically, you’re doing it wrong, but if you’re using _human_ output uncritically you’re doing it wrong too. That said, they are not magic, and my fear is tha…

LLMs turn foot guns into foot bazookas. Still, if you learn to aim them away from your body they do make a bigger better boom.

Re: Everything around LLMs is still magical and wishful thinking

#286

Earlier quoted context omitted.

Maybe it's due to a more R&D-ish nature of my current work, but for me, LLMs are delivering just as much gains in the "thinking" part as in "coding" part (I handle the "communicating" thing myself just fine for now). Using LLMs for "thinking" tasks feels similar to how mastering web search 2+ decades ago felt. Search engines enabled access to information provided you know what you're looking for; now LLMs boost that…

I’m surprised it’s only 1/3rd. 90% of my searches for information start at Perplexity or Claude at this point.

Perplexity is too bulky for queries Kagi can handle[0], and I don't want to waste o3 quota[1] on trivial lookups.

--

[0] - Though I admit that almost all my Kagi searches end in "?" to trigger AI answer, and in ~50% of the cases, I don't click on any result.

[1] - Which AFAIK still exists on Plus plan, though I haven't hit it ~two months.

Re: Everything around LLMs is still magical and wishful thinking

#287

Earlier quoted context omitted.

Maybe it's due to a more R&D-ish nature of my current work, but for me, LLMs are delivering just as much gains in the "thinking" part as in "coding" part (I handle the "communicating" thing myself just fine for now). Using LLMs for "thinking" tasks feels similar to how mastering web search 2+ decades ago felt. Search engines enabled access to information provided you know what you're looking for; now LLMs boost that…

From time to time I use an LLM to pretend to research a topic that I had researched recently, to check how much time it would have saved me. So far, most of the time, my impression was "I would have been so badly mislead and wouldn't even know it until too late". It would have saved me some negative time. The only thing LLMs can consistently help me with so far is typing out mindless boilerplate, and yet it still som…

> So far, most of the time, my impression was "I would have been so badly mislead and wouldn't even know it until too late". It would have saved me some negative time.

That was my impression with Perplexity too, which is why I mostly stopped using it, except for when I need a large search space covered fast and am willing to double-check anything that isn't obviously correct. Most of the time, it's o3. I guess this is the obligatory "are you using good enough models" part, but it really does make a difference. Even in ChatGPT, I don't use "web search" with default model (gpt-4o) because I find it hallucinate or misinterpret results too much.

> The kind of stuff it does help researching with is usually the stuff that's easy to research without it anyway.

I disagree, but then maybe it's also a matter of attitude. I've seen co-workers do exact same research as I did, in parallel, using the same tools (Perplexity and later o3); they tend to do it 5-10x faster than I do, but then they get bad results, and I don't.

Thing is, I have an unusually high need to own the understanding of any thing I'm learning. So where some co-workers are happy to vibe-check the output of o3 and then copy-paste it to team Notion and call their research done, I'll actually read it, and chase down anything that I feel confused about, and keep digging until things start to add up and I feel I have a consistent mental model of the topic (and know where the simplifications and unknowns are). Yes, sometimes I get lost in following tangents, and the whole thing takes much longer than I feel it should, then I don't get "misled by the LLM".

I do the same with people, and sometimes they hate it, because my digging makes them feel like I don't trust them. Well, I don't - most people hallucinate way more than SOTA LLMs.

Still, the research I'm talking about, would not be easy to do without LLMs, at least not for me. The models let me dig through things that would otherwise be overwhelming or too confusing to me, or not feasible in the time I have for it.

Own your understanding. That's my rule.

Re: Everything around LLMs is still magical and wishful thinking

#288

This reads like the author is mad about imprecision in the discourse which is real but to be quite frank more rampant amongst detractors than promoters, who often have to deal with the flaws and limitations on a day to day basis. The conclusion that everything around LLMs is magical thinking seems to be fairly hubristic to me given that in the last 5 years a set of previously borderline intractable problems have beco…

Translation, transcription, and code generation (up to some scale) were borderline intractable problems?

Google Translate, Whisper and Code Generators (up to some scale) have existed for quite some time without using LLMs.

Re: Everything around LLMs is still magical and wishful thinking

#289

I have to say I’m in the exact camp the author is complaining about. I’ve shipped non trivial greenfield products which I started back when it was only ChatGPT and it was shitty. I started using Claude with copying and pasting back and forth between the web chat and XCode. Then I discovered Cursor. It left me with a lot of annoying build errors, but my productivity was still at least 3x. Now that agents are better an…

> I audit everything myself before making PRs and test rigorously How do you audit code from an untrusted source that quickly, LLMs do not have the whole project in their heads and are proned to hallucinate. On average how long are your prompts and does the LLM also write the unit tests?

The auditing is not quick. I prefer cursor to claude code because I can review its changes while it’s going more easily and stop and redirect it if it starts to veer off course (which is often, but the cost of doing business). Over time I still gain an understanding of the codebase that I can use to inform my prompts or redirection, so it’s not like I’m blindly asking it to do things. Yes, I do ask it to write unit tests a lot of the time. But I don’t have it spin off and just iterate until the unit tests pass — that’s a recipe for it to do what it needs to do to pass them and is counterproductive. I plan what I want the set of tests to look like and have them write functions in isolation without mentioning tests, and if tests fail I go through a process of auditing the failing code and then the tests themselves to make sure nothing was missed. It’s exactly how I would treat a coworkers code that I review. My prompts range from a few sentences to a few paragraphs, and nowadays I construct a large .md file with a checklist that we iterate on for larger refactors and projects to manage context

Re: Everything around LLMs is still magical and wishful thinking

#290

Earlier quoted context omitted.

You can't just take cost of training out of the equation... If these companies plan to stay afloat, they have to actually pay for the tens of billions they've spent at some point. That's what the parent comment meant by "free AI"

Yes, you can - because of LLama. Training is expensive, but it's not that expensive either. It takes just one of those super-rich players to pay the training costs and then release the weights, to deny other players a moat.

If your economic analysis depends on "one of those super-rich players to pay" for it to work, it isn't as much analysis as wishful thinking.

All the 100s of billions of $ put into the models so far were not donations. They either make it back to the investors or the show stops at some point.

And with a major chunk of proponent's arguments being "it will keep getting better", if you lose that what you got? "This thing can spit out boilerplate code, re-arrange documents and sometimes corrupts data silently and in hard to detect ways but hey you can run it locally and cheaply"?

Post reply on HN