Live data from Hacker News

Everything around LLMs is still magical and wishful thinking

dmitriid.com

231–240 of 377 posts

Re: Everything around LLMs is still magical and wishful thinking

#231
post #221
post #74

Earlier quoted context omitted.

I'm no LLM evangelist, far from it, but I expect models of similar quality to the current bleeding-edge, will be freely runnable on consumer hardware within 3 years. Future bleeding-edge models may well be more expensive than current ones, who knows.

How do the best models that can run on say a single 4090 today compare to GPT 3.5?

Qwen 2.5 32B which is an older model at this point clearly outperforms it:

https://llm-stats.com/models/compare/gpt-3.5-turbo-0125-vs-q...

Re: Everything around LLMs is still magical and wishful thinking

#232

Many well-trusted and reasonable tech folks who are known for sober takes on subjects have reported substantial improvements in their programming work by using various forms of generative AI. What does substantial mean? Somewhere between 5% and 100%. Something NOT insignificant. At a minimum , it is safe to say that GenAI is or could be a significantly beneficial tool for a significant number of people. It's not requ…

"People claim to have productivity increases anywhere between a random number I invented to solve other number I invented. We should believe these claims uncritically"

Re: Everything around LLMs is still magical and wishful thinking

#233

I personally don't really get this. _So much_ work in the 'services' industries globally comes down to really a human transposing data from one Excel sheet to another (or from a CRM/emails to Excel), manually. Every (or nearly every) enterprise scale company will have hundreds if not thousands of FTEs doing this kind of work day in day out - often with a lot of it outsourced. I would guess that for every 1 software e…

Each FTE doing that manual data pipelining work is also validating that work, and they have a quasi-legal responsibility to do their job correctly and on time. They may have substantial emotional investment in the company, whether survival instinct to not be fired, or ambition to overperform, or ethics and sense to report a rogue manager through alternate channels.

An LLM won't call other nodes in the organization to check when it sees that the value is unreasonable for some out-of-context reason, like yesterday was a one-time-only bank holiday and so the value should be 0. *It can be absolutely be worth an FTE salary to make sure these numbers are accurate.* And for there to be a person to blame/fire/imprison if they aren't accurate.

Re: Everything around LLMs is still magical and wishful thinking

#234

The point about non-determinism is moot if you understand how it works. An accurate LLM always gives the same result where the same result is needed, no matter how many times you ask it. Try asking any LLM what is 2x2 on a temperature it's designed for, what are the chances to get 5 in a reply? In reality, modern LLMs trained with RL have terrible variance and mainly learn 1:1 mapping of ideas to ideas, which is a bi…

Are those LLMs in the room with us now? ;)

Actually, I did try asking ollama running locally. That should've reduced the amount of non-determinism and whatever layers providers add, and the uncertainty of computer availability.

I asked it for a list of Javascript keywords ordered alphabetically. Within 5 minutes it produced a slightly different list

Asking a model for 2x2 is moot because 2x2=5 is statically highly unlikely. Anything more complex though?

Re: Everything around LLMs is still magical and wishful thinking

#235

Earlier quoted context omitted.

I don't think you got the author's point. They didn't say "AI is bad". Take another look.

Ah I'm sorry you're totally right, I was irresponsible and just skimmed the article last time. I retract my last statement about it being a bad take

It's fine. I'm as guilty of skimming articles and bad takes as well :)

Re: Everything around LLMs is still magical and wishful thinking

#236
post #221

Earlier quoted context omitted.

How do the best models that can run on say a single 4090 today compare to GPT 3.5?

Qwen 2.5 32B which is an older model at this point clearly outperforms it: https://llm-stats.com/models/compare/gpt-3.5-turbo-0125-vs-q...

Even when quantized down to 4 bits to fit on a 4090?

Re: Everything around LLMs is still magical and wishful thinking

#237

I'm honestly started to get pissed seeing these articles on HN every day. Aren't we technologists? I see these articles and it's just people who can't accept the future, or feel threatened by it, or are pissed they are not contributing to it. I'm an expert in my field with decades of experience, and a damn good programmer, and Cursor is like zipping down the street on a powered bike. There is no wishful thinking here…

> I'm honestly started to get pissed seeing these articles on HN every day. Aren't we technologists?

Funnily enough, that's exactly what the article asks.

> I'm an expert in my field with decades of experience, and a damn good programmer, and Cursor is like zipping down the street on a powered bike.

Oh look, an unverified claim. Go through the first list in the article, and ask "what we, as technologists, know about this claim".

You can also actually read the article and see that I actually use these tools every day.

Re: Everything around LLMs is still magical and wishful thinking

#238

Earlier quoted context omitted.

uh https://www.technologyreview.com/2025/05/20/1116823/how-ai-i... https://hai.stanford.edu/news/hallucinating-law-legal-mistak... https://www.reuters.com/technology/artificial-intelligence/a... There are more of these stories every week. Are you using AI in a way that doesn’t allow you to be entrapped by this sort of thing?

ChatGPT links to the actual text in a case now. Also, take the output from Claude, and put it into Gemini and tell it to verify holdings. Furthermore, spot checking 10 cases doesn’t take long.

> ChatGPT links to the actual text in a case now.

Or to text in a law is hallucinated: https://www.timesofisrael.com/judge-slams-police-for-using-a...

Re: Everything around LLMs is still magical and wishful thinking

#239
post #234

The point about non-determinism is moot if you understand how it works. An accurate LLM always gives the same result where the same result is needed, no matter how many times you ask it. Try asking any LLM what is 2x2 on a temperature it's designed for, what are the chances to get 5 in a reply? In reality, modern LLMs trained with RL have terrible variance and mainly learn 1:1 mapping of ideas to ideas, which is a bi…

Are those LLMs in the room with us now? ;) Actually, I did try asking ollama running locally. That should've reduced the amount of non-determinism and whatever layers providers add, and the uncertainty of computer availability. I asked it for a list of Javascript keywords ordered alphabetically. Within 5 minutes it produced a slightly different list Asking a model for 2x2 is moot because 2x2=5 is statically highly un…

>I asked it for a list of Javascript keywords ordered alphabetically. Within 5 minutes it produced a slightly different list

That's not my point, the keyword here is "meaningful". How many of those lists are correct? (ignoring the fact prompting a LLM for lists is a bad idea, let alone local ones)

If you spend some time with SotA LLMs, you'll see that on rerolling they express pretty much the same ideas in different ways, most of the time.

Re: Everything around LLMs is still magical and wishful thinking

#240
post #58

Earlier quoted context omitted.

> It is something to sneeze at if you are 10-15% more expensive to employ due to the cost of the LLM tools. Claude Max is $200/month, or ~2% of the salary of an average software engineer.

Does anyone actually know what the real cost for the customers will be once the free AI money no longer floods those companies?

There's a potential for 100x+ lower cost of chips/energy for inference with compute-in-memory technology.

So they'll probably find a reasonable cost/value ratio.

Post reply on HN