Live data from Hacker News

ML promises to be profoundly weird

aphyr.com

531–540 of 641 posts

Re: ML promises to be profoundly weird

#531

I see the penny hasn't dropped yet that: humans are doing (roughly) the same dumb thing these models are doing. Humans are predisposed to not notice that though.

All humans do dumb things, of that we have no doubt. But are the dumb things qualitatively the same as those that AIs do? I don't think so. That's essentially the entire problem. We have pretty good ideas about the ways humans make mistakes. Its pretty much the point of all fiction!

AIs fail in new and unpredictable ways. Nobody is saying humans are infallible.

Finally, because I suspect some people are forming tribalism around this, this doesnt to mind my say AI is Good(tm) or Bad (tm). It literally says its going to be weird.

Re: ML promises to be profoundly weird

#532

Earlier quoted context omitted.

GPT-5.4 gets 82.7% on Browsecomp (a benchmark specifically testing tool use), which is a hallucination rate of 17.3%, on questions like "Give me the title of the scientific paper published in the EMNLP conference between 2018-2023 where the first author did their undergrad at Dartmouth College and the fourth author did their undergrad at University of Pennsylvania." Since the goalposts have been moved to include effo…

the latest top reported agentic LLMs score about 83–87%, versus an original human baseline of about 25.3% end to end, so today’s best systems appear to outperform humans by roughly 58–62 percentage points, or about 3.3–3.4× So according to your own benchmark LLMs hallucinate much less than humans and report way higher accuracy. Do you agree to be more skeptical of humans than LLMs on these tasks?

1. Irrelevant. I've delivered example after example of your fave model bullshitting. You should've bitten the bullet long ago. Honestly I'm disappointed; I've seen you in a lot of AI threads and assumed you'd be good to talk to on this, but you've moved the goalposts over and over again rather than engage in good faith. Anyone reading this thread (god bless them) can see you're plainly not objective here, thus calling into question your advocacy everywhere.

2. Humans will say "I don't know". The problem with hallucinations isn't that they're wrong, it's that there's no way to know they're wrong without being an expert or doing everything yourself, which undermines much of the reason for using an LLM--it certainly undermines their companies' valuations. You're conflating human failure ("I don't know") with model bullshitting ("I do know"... but it's wrong), which I would've previously attributed to basic human fuzziness, but now that I know you're not objective I'm pretty sure it's just flailing debate tactics.

3. Users can't teach these services to be better. If I have a junior engineer making assumptions about an API, I can teach them to not do that, or fire them in favor of one that can. I can't do that with LLMs.

4. The humans they're testing against aren't experts. Tax law experts will beat LLMs at tax law, etc. Again another flailing debate tactic.

Predictably, I'm done with this thread. Feel free to reply if you want the last word.

Re: ML promises to be profoundly weird

#533
post #452
post #432

Earlier quoted context omitted.

Nobody said anything about Europeans having a "natural right". Bad enough to derail a conversation with irrelevant political nitpicking, unforgiveable to use a strawman to do so. Boo.

It's not irrelevant. GP made a comparison between what we're going through and the Industrial Revolution. Ignoring the negatives of that revolution - like by acting as though the "new world" was uninhabited/unused and so Europeans had a right to its resources - seems like a bad idea.

> like by acting as though the "new world" was uninhabited/unused and so Europeans had a right to its resources - seems like a bad idea.

maybe it was a bad idea, but that's what happened.

Re: ML promises to be profoundly weird

#534

Earlier quoted context omitted.

> The general point is accurate, don’t take it so literally. GP is saying it is not, and you're just reiterating what OP said as fact.

It's sort of the exception that proves the rule. This is where STEM people are weak- a lack of knowledge on history. In another forum, someone would have chipped in that England's virgin forests were fully deforested by 1150. And someone else would have pointed out that this deforestation produced the economic demand for coal that drove the Industrial Revolution in the first place. Still, that kind of underscores OP'…

Yeah - really struggling to understand why people are not grasping this point.

Yes, Easter Island was deforested far earlier - but you wouldn't compare the steam engine's capability in resource extraction compared to what people on Easter Island were doing.

It feels like people are almost straining to not understand the point - I think it's quite clear how ML + AI serve to extract resources of data at a unheard of scale.

Re: ML promises to be profoundly weird

#535
post #324

Earlier quoted context omitted.

>Raw parameter counts stopped increasing almost 5 years ago, and modern models rely on sophisticated architectures like mixture-of-experts, multi-head latent attention, hybrid Mamba/Gated linear attention layers, sparse attention for long context lengths, etc. Agree, I recently updated our office's little AI server to use Qwen 3.5 instead of Qwen 3 and the capability has considerably increased, even though the new mo…

What you described sounds plausible (expected, even). But >Raw parameter counts stopped increasing almost 5 years ago Really? 5 years ago ? Until just about 3 years ago OpenAI's latest offering was only ChatGPT 3.5 Most of the models people talk about now didn't even exist 3 years ago let alone 5. Even now, I don't know if parameter count stopped mattering or just matters less For example, I have no idea if the new M…

I agree the original poster exaggerated it. But generally models indeed have stopped growing at around 1-1.5 trillion parameters, at least for the last couple of years.

>Even now, I don't know if parameter count stopped mattering or just matters less

Models in the 20b-100b range are already very capable when it comes to basic knowledge, reasoning etc. Improving the architecture, having better training recipes helped decrease the required parameter count considerably (currently 8b models can easily beat the 175b strong GPT3 from 3 years ago in many domains). What increasing the parameter count currently gives you is better memorization, i.e. better world knowledge without having to consult external knowledge bases, say, using RAG. For example, Qwen3.5 can one-short compilable code, reason etc. but can't remember the exact API calls to to many libraires, while Sonnet 4.6 can. I think what we need is split models into 2 parts: "reasoner" and "knowledge base". I think a reasoner could be pretty static with infrequent updates, and it's the knowledge base part which needs continuous updates (and trillions of parameters). Maybe we could have a system where a reasoner could choose different knowledge bases on demand.

Re: ML promises to be profoundly weird

#536
post #302
post #79

Thank you for putting it so succinctly. I keep explaining to my peers, friends and family that what actually is happening inside an LLM has nothing to do with conscience or agency and that the term AI is just completely overloaded right now.

> I keep explaining to my peers, friends and family that what actually is happening inside an LLM has nothing to do with conscience or agency What would the insides have to look like to have anything to do with conscience or agency?

I would expect to find a tiny, sweaty man constantly pedaling while cursing the user for it's seemingly infinite stupidity.

Re: ML promises to be profoundly weird

#537

I so far asked few people to make GPT-5.4 thinking to bullshit (with max 4 pages of prompt), no one can find an example. But the way people speak in general, as well as this post, implies that such a challenge can easily be beaten. If so, I'm not able to find examples.

Here is small example of ChatGPT giving a wrong answer without expressing any doubt (aka bullshitting):

https://chatgpt.com/share/69d78ec7-67b0-8395-9fd1-522b760ab5...

GPT-5.4 Thinking, Pro account.

Re: ML promises to be profoundly weird

#538
post #443

Earlier quoted context omitted.

Read Moby Dick some time my friend.

The industrial revolution is generally understood to have started somewhere around 1760, Moby Dick took place in approximately 1830, about 10 years before what some historians mark as the end of the agrarian to Industrial shift that is generally termed the Industrial revolution https://en.wikipedia.org/wiki/Industrial_Revolution I get sort of wishy-washy from 1830 on, because lots of people put the end of the Industr…

That’s besides the point because most whales were killed in the XX century.

Re: ML promises to be profoundly weird

#539

Earlier quoted context omitted.

> you don't feel you need any consent from the people you are taking from. What has been "taken", exactly?

> What has been "taken", exactly? Where are you going with this line of thought? That making a copy of someone's work, using it for profit and not crediting them doesn't "take" anything from them?

I find that these discussions at the intersection of art and law tend to blur technical and familiar uses of words. So it's important to specify what was actually taken here because otherwise the discussion becomes muddy.

"making a copy of someone's work, using it for profit and not crediting them" wasn't really the scenario being discussed in this thread -- is that what you meant by "taking"?

Steve had made the point:

  Not every single thing in the entire world requires explicit consent.
But actually taking someone else's verbatim work and selling it as your own is one of those instances where consent would be required, because many people see a clear line between someone selling another author's work and the author not getting a dollar because of that.

That doesn't preclude other instance where explicit consent is not required. For example, do I need your consent to learn from your work and produce similar work of my own? Am I required to credit you in my work for having learned from you? Am I taking from you if I don't share my profits with you?

Some rights holders would say yes, actually. Which, I don't agree with. I think it's important that we not require the artist's explicit consent for all things, because listening to some of rights holders (e.g. Disney), they have very expansive ideas about what kind of control they are owed by society over their creations.

Therefore, I think if you're going to claim something has been taken, you should specify what exactly.

Re: ML promises to be profoundly weird

#540
post #290

Earlier quoted context omitted.

Can we interpret "abundant" in a Darwinian sense e.g. diversity of life? I would think the industrial farming revolution decreased crop variety over time same for animal lineages aside from the rapid increase in mixed poodle breeds.

Crop variety was decreased by the original farming revolution, about 10k years before the industrial revolution. Rather than eating whatever was available, the large majority of the caloric input of an agricultural society comes from a few staple crops optimized for overwinter storability and producing large yields and thus supporting a large number of people. The industrial revolution didn’t qualitatively change far…

This is particularly evident if you had been around rural villages in eastern Europe in the late 00s, particularly those inhabited by elderly people at 70 years old and above.

They were still doing subsistence agriculture to supplement their own income well into the 21th century. Of course they didn't grow enough calorie heavy crops like corn, potatoes or wheat to live entirely off the land, but they had enough food that a bi-monthly shopping trip with their children was enough to get by.

Post reply on HN