Live data from Hacker News

Why I'm still bearish on LLMs after Navier-Stokes

dank.systems

401–410 of 642 posts

Re: Why I'm still bearish on LLMs after Navier-Stokes

#401

Earlier quoted context omitted.

Just now I tried prompting logged-out ChatGPT (which at least claims to be 5.6-Luna) with: > Let's play a game of chess. You can take White. Please draw an ASCII rendition of the board after each move, so that we can be clear about the position. (I hoped the latter requirement would help it be "not blindfolded"; last time I used a Lichess demo board to track the position in another tab, because I have no talent for b…

What happens when you ask it to play chess against you if the chess game has an API? Are you measuring chess or multi-tasking skill? Also what harness? If you’re using a general harness of course it’s going to try and give you commentary. I say this not because I’m an LLM shill but because false equivalence is all over the place in the space and maybe it’s a fine heuristic for you but probably not a real outcome when…

Why does a 6 year old not need any of these guardrails?

Frontier model’s failure modes are a direct refutation of claims that we’ve reached (or will soon reach) the artificial general intelligence. We may have reached an artificial general intelligence, but there may be more complexity to this than even AI thought leaders are talking / influencing about.

Maybe not all AGIs have a path to digital singularity. Maybe our current era of intelligence modeling has fundamental flaws and we are in a local minimum of the artificial intelligence space.

To note, I would bet with a good amount of certainty that we have enough compute power and automation to DDOS the internet out of existence with botnets. That doesn’t make the frontier models intelligent, that just makes their handlers reckless.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#402

Earlier quoted context omitted.

is that what I'm saying? or am I talking about AGI? perhaps there's some irony here to be explored when it comes to basic reading comprehension gaps

My point was you are misunderstanding G, or at least applying it erroneously here. Being good at chess is not a generalization of any other body of knowledge, it is a rigorous set of rules. The only way to be good at chess is to practice chess, or to apply deep calculations. The latter is the model writing code. The illegal move aspect has more to do with a failure of online/in-context learning, which would support y…

chess is not just a rigorous set of rules, it is rules as foundation with layers of strategy on top. and so is, for example, scientific methodology or chemical interactions or virtually everything else under-the-sun that comprises human knowledge

knowledge for chess is derived from memorizing strategies that have been well-defined for decades paired with in-game reasoning processes. this is not at all different from any other body of knowledge. Noble gases, laws of thermodynamics, organic chemistry just to name a few - these are all 'strategies' that define observed phenomena, analytical frameworks that trace a logical, rational set of interactions and which can predict the next

for an AGI, all of this should be a cakewalk, trained as it were to surpass human capability in any and every domain [0] (thus the G for 'general' and not 'N' for 'narrow' [1]). it should be a natural at everything, infinitely adaptable on-the-fly. the whole point of AGI is that it surpasses human capabilities even at our frontiers and bleeding edge (unless you're private enterprise and you've moved the goalposts for industry [2])

currently, it's only AGI-seeming if it gets benchmaxxed enough. otherwise it sucks at what it does and then is only barely competent at tasks if paired with enough skills and tests to make it more diligent at its work. this makes sense to me - for any probabilistically trained tool, even one that you post-train and fill with nothing but the best-quality evidence, the ultimate result is the lowest-common-denominator output for your sample set. there's no natural reasoning the AI does itself to make itself better at what it does - it's all human curation and categorization of sources ingested paired with RLHF post-training that we can get the mediocre-at-chess-at-best results that we see now and the benchmaxxed scores against whatever arbitrary and pre-defined measure

that's not AGI by any classical definition. that's a cool, useful, and powerful tool that makes our lives easier, much like a hammer, nail, and studs make mounting a picture frame easier than if we only had our hands alone

[0] https://ischool.syracuse.edu/types-of-ai/#:~:text=General%20...

[1] https://aiethicslab.rutgers.edu/e-floating-buttons/weak-ai-n...

[2] https://aibusiness.com/ml/what-exactly-is-artificial-general...

Re: Why I'm still bearish on LLMs after Navier-Stokes

#403

Earlier quoted context omitted.

Threw me off too. Like why???

It's a tech bro thing. Altman does it too, and I've worked with people in the past who do it. I read it as "I'll take literally any conscience for myself no matter how minor, at any cost for you no matter how big".

it's an internet thing. people were typing lowercase on IRC in the 90s.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#404
post #293

Earlier quoted context omitted.

Fortunately, a fellow commenter was so kind and did it with Astra. Didn't do that well either [0]. I'm sure GPT-7 will be super mega ASI regardless (since GPT-6 Astra already claimed AGI in the minds of Jen-Hsun, et al.)... I'll say it till there is any evidence of the contrary, LLMs are not intelligent and their capabilities solely within the realms of well tailored training data. "Just" having been trained on every…

Doesn't look impressive, although I'm hearing a marked improvement in choosing legal moves, compared to early 2025. Given the pace of improvements, is it really unimaginable that GPT-7 will play Chess reasonably well and generalize better? I would not be surprised if OpenAI released a model that beats humans at chess this year.

Maybe watch some HuskIRL videos to temper your expectations. Sure, frontier models providers may alter their harnesses to better target chess, but that’s lipstick on a pig imo. The models themselves are not, in isolation, capable of solving general tasks. We haven’t modeled intelligence sufficiently. We’re in a local minimum and throwing billions of dollars at a gamble that that local minimum can facilitate the concentration of wealth even further and fully realize the American dream of eliminating the middle class.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#405

Great article, but the lack of sentence capitalization makes it unnecessarily difficult to read. Apologies if this comment is off-topic, but it really is quite egregious, and since the article was submitted by the author I presume they are open to the feedback.

huh i didn't even notice, that's how i write all my blog posts too. it looks nicer to me and i don't have to bother checking for "proper" capitalization if everything's just lowercase anyway. didn't realize people struggled to read text that way though, maybe i should change my writing style if this is a common pain point

For what is worth, it is incredibly jarring to me as well (I am a middle millennial, not a boomer, and usually not prescriptivist about writing style). My knee-jerk reaction (I know, book by the cover...) is that it is probably not worth reading given that the writer did not think it is worth making it easy to read.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#406
Navier-Stokes is a well defined problem, or "a hard technical problem". Most problems in the business world lack a good definition and tacit knowledge is required to solve them. As far as I've seen, AI lacks any kind of tacit knowledge, strategic thinking, etc what so ever.

Take a customer service person, that as soon as AI agents replaced was hacked. Lots of tacit knowledge, that wasn't measured, or even probably in the job description, until AI agents had none and the gap was taken advantage of. Gap is probably a poor word here as it implies not a chasm, which could very well be the case.

Your analysis was excellent but short on one front, AI has endurance on it's side. Looking at the N-S solution, OpenAI had 10,000+ instances that kept trying around the clock. Assembling a human team to do that would require a lot of effort. Though, that society collectively choose not to, perhaps tells how valuable it really is (i.e. it's now easy to launch a Manhattan Project level of effort). So maybe it can also be said that AI is also good at marshaling resources.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#407

Earlier quoted context omitted.

It's a tech bro thing. Altman does it too, and I've worked with people in the past who do it. I read it as "I'll take literally any conscience for myself no matter how minor, at any cost for you no matter how big".

it's an internet thing. people were typing lowercase on IRC in the 90s.

there is a very big difference between starting an ephemeral instant message without capitalization, and lacking capitalization throughout a more "persistent" long-term medium

Re: Why I'm still bearish on LLMs after Navier-Stokes

#408
post #64

Earlier quoted context omitted.

The actual current frontier plays somewhere around GM level. https://chessbench-ai.github.io/#leaderboard It's also worth noting that the very latest models (GPT-6 and Fable 5.1) actually play worse than their immediate predecessors, so it is likely that the labs are not benchmaxxing for this yet. If they did, I'm sure they could come up with something superior to humans. But there is probably very little demand for…

I know HN readers and posters just read numbers and can't be bothered to read, but please read the methodology before making any claims. > About their ELO ratings from their own website: > A field-relative rating calculated within ChessBench. It compares performance among the tested models and is not a direct equivalent of a human chess rating. I am around 1600 elo in over the board I can mop up Astra Fable etc even…

> Why do I even scroll through this website.

Because other HN bring in their own experience telling us what is real and what is BS. Maybe next time it will be someone else with experience in something else that will call out BS and you will see it. I didn’t really think LLM:s are any near good in chess but I don’t play chess so don’t know what 1600 means. So you helped me by calling BS.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#409

Earlier quoted context omitted.

What happens when you ask it to play chess against you if the chess game has an API? Are you measuring chess or multi-tasking skill? Also what harness? If you’re using a general harness of course it’s going to try and give you commentary. I say this not because I’m an LLM shill but because false equivalence is all over the place in the space and maybe it’s a fine heuristic for you but probably not a real outcome when…

Why does a 6 year old not need any of these guardrails? Frontier model’s failure modes are a direct refutation of claims that we’ve reached (or will soon reach) the artificial general intelligence. We may have reached an artificial general intelligence, but there may be more complexity to this than even AI thought leaders are talking / influencing about. Maybe not all AGIs have a path to digital singularity. Maybe ou…

> Why does a 6 year old not need any of these guardrails?

They're not guardrails, they're a different input/output environment.

Re: Why I'm still bearish on LLMs after Navier-Stokes

#410
post #78

Earlier quoted context omitted.

The AI can write a chess bot program that will beat you. You're thinking about this the wrong way. The system is built and delivered as it is because that's how the providers make the most money. If they cared to have it perform well in chess games, you'd see a different shape and behavior. We shouldn't ask the multibillion dollar automated software generation system to play games with us any more than we should ask…

> The system is built and delivered as it is because that's how the providers make the most money. If they cared to have it perform well in chess games, you'd see a different shape and behavior. This argument is fundamentally incompatible with all the breathless rhetoric about "AGI" coming from the providers' general direction.

It's really not.

The labs frequently apply their raw models to problems that do not make economic sense for their customers but that demonstrate the power and capability of their systems. These experiments can cost millions of dollars. That's not customer-shaped.

They're not going to give you access to that. It's not a product. The government might have an interest in this, but that's not something you'd be privileged to know about.

And when these labs do develop "AGI", they more than likely won't be selling it to end users. They've pretty much already said this.

Post reply on HN