Live data from Hacker News

Modern-Day Oracles or Bullshit Machines? How to thrive in a ChatGPT world

thebullshitmachines.com

611–620 of 652 posts

Re: Modern-Day Oracles or Bullshit Machines? How to thrive in a ChatGPT world

#611

Earlier quoted context omitted.

Saying something again does not provide proof of its actual veracity. Writing it in caps does not make it true despite the increased emphasis. I default to skepticism in the face of unproven assertions: if one can’t prove that they reason then we must accept the possibility that they do not. There are myriad examples of these models failing to “reason” about something that would trivial for a child or any other human…

Here was my test at ChatGPT 3.5.[0] I made up a novel game, and it figured it out. The test is simple, but it made me doubt absolute arguments that LLMs are not able to reason, in some way. There is a question at the end of that comment, would love to hear other options. [0] https://news.ycombinator.com/item?id=35442147

How does this prove reasoning? The thread you point to has several question in it that remain unanswered that ask the same question? How is this not entirely derivative too — there’s a huge number of these kind of 3-box “games” (although I don’t see this as a game really) so something very similar to this is probably in the training data a lot. Writing code to factor a number is definitely very common. Variation of this are also very common interview questions for interns (at least when I was interviewing)

Re: Modern-Day Oracles or Bullshit Machines? How to thrive in a ChatGPT world

#612

Earlier quoted context omitted.

Here was my test at ChatGPT 3.5.[0] I made up a novel game, and it figured it out. The test is simple, but it made me doubt absolute arguments that LLMs are not able to reason, in some way. There is a question at the end of that comment, would love to hear other options. [0] https://news.ycombinator.com/item?id=35442147

My thread has been voted down and it’s getting stale. The few remaining people are biased towards there point of view and are unlikely to entertain anything that will trigger a change in their established world view. Most people will use this excuse to avoid responding to or even looking at your link here. It is compelling evidence.

I’d settle for these things being able to do value comparison consistently well, play a game of tic tac toe more than once correctly or use a UI after an update and not fail horrendously to move the needle a little bit for me. People claiming these things selectively reason while also not being able to explain why seems a lot like magical thinking to me rather than entertaining the possibility you might be projecting onto something that is really damn-well engineered to make your anthropomorphize it.

Re: Modern-Day Oracles or Bullshit Machines? How to thrive in a ChatGPT world

#613

Earlier quoted context omitted.

Nah, you don't need to know the details to evaluate something. You need the output and the null hypothesis. If a trading firm claims they have a wildly successful new strategy, for example, then first I want to see evidence they're not lying - they are actually making money when other people are not. Then I want to see evidence they're not frauds - it's easy to make money if you're insider trading. Then I want to see…

I don't get offended when people call my work a stochastic parrot. I just put them in the same bucket of intelligence as an 8b model and weight their inputs accordingly.

lmao I’m copy pasting this, could be writing for big bang theory

Re: Modern-Day Oracles or Bullshit Machines? How to thrive in a ChatGPT world

#614
post #587

Earlier quoted context omitted.

I can’t and that’s pretty cool to think about! Of course if we’re going that far down the chain of assumption we’re not quite ready to talk about LLMs imo (then again maybe it would be the perfect place to talk about them as contrast/comparison; certainly exciting ideas in that light). From my own perspective: if we’re gonna say these things reason and we’re using the definition of reasoning we apply to humans, then…

IMHO, any argument against LLM intelligence should be validated by first applying them to humans. And then you'd realize that a lot of naïve arguments against LLMs would imply that a significant portion of homo-sapiens can't reason, are unable to really think, and are no more than stochastic parrots. It's actually a rather dangerous line of reasoning.

I’m curious what’s dangerous about it? How do your square the inability to play tic-tac toe or do value comparisons correctly with “we should compare this to a humans reasoning” If it can’t do things like basic value comparison correctly what business do we have saying it “reasons like a human?”

Re: Modern-Day Oracles or Bullshit Machines? How to thrive in a ChatGPT world

#615

Earlier quoted context omitted.

the machine is fooling you with a mimicry of reasoning. and you are falling for it.

What is reasoning? What is understanding? Do humans do either? How do you know?

0) causal model == symbolic representation with associated rules for generating statements

1) understanding anything == building a causal model of it

2) intelligence == ability to build causal models

3) reasoning == proving or disproving statements

4) math == causal models of abstract worlds

5) science == causal models of real world with associated real world actions to test hypothesis

Re: Modern-Day Oracles or Bullshit Machines? How to thrive in a ChatGPT world

#616

Earlier quoted context omitted.

Disagree — proponents of this point still have yet to prove reasoning and other studies agree about “reasoning” being potentially fake/simulated: https://the-decoder.com/apple-ai-researchers-question-openai... Just claiming a capability does not make it true and we have 0 “proof” of original reasoning that can be proved coming from these models. Especially given the potential cheating in current SOTA benchmarks

It’s stupid. You can prove that LLMs can reason by simply giving it a novel problem where no data exists and having it solve that problem. LLMs CAN reason. Whether it can’t reason is not provable. To prove that you have to give the LLM every possible prompt that it has no data for and effectively show it never reasons and gets it wrong all the time. Not only is the proof impossible but it’s already been falsified as…

> But they can reason

This isn't demonstrated yet, I would say. A good analogy is how people have used NeRFs to generate Doom levels, but when they do, the levels don't have offscreen coherence or object permanence. There's no internal engine behind the scenes making an actual Doom level. There's just a mechanism to generate things that look like outputs of that engine. In the same way, an LLM might well just be an empty shell that's good at generating outputs based on similar-looking outputs it was trained on, rather than something that can do the work of thinking about things and producing outputs. I know that's similar to "statistical parrot", but I don't think what you're saying demonstrates anything more than that.

Re: Modern-Day Oracles or Bullshit Machines? How to thrive in a ChatGPT world

#617

Earlier quoted context omitted.

It’s stupid. You can prove that LLMs can reason by simply giving it a novel problem where no data exists and having it solve that problem. LLMs CAN reason. Whether it can’t reason is not provable. To prove that you have to give the LLM every possible prompt that it has no data for and effectively show it never reasons and gets it wrong all the time. Not only is the proof impossible but it’s already been falsified as…

> But they can reason This isn't demonstrated yet, I would say. A good analogy is how people have used NeRFs to generate Doom levels, but when they do, the levels don't have offscreen coherence or object permanence. There's no internal engine behind the scenes making an actual Doom level. There's just a mechanism to generate things that look like outputs of that engine. In the same way, an LLM might well just be an e…

It can be trivially demonstrated with a unique problem that doesn’t exist in the training data and an answer that is correct and has a low probability of being arrived at without reasoning.

Re: Modern-Day Oracles or Bullshit Machines? How to thrive in a ChatGPT world

#618

Earlier quoted context omitted.

Here was my test at ChatGPT 3.5.[0] I made up a novel game, and it figured it out. The test is simple, but it made me doubt absolute arguments that LLMs are not able to reason, in some way. There is a question at the end of that comment, would love to hear other options. [0] https://news.ycombinator.com/item?id=35442147

How does this prove reasoning? The thread you point to has several question in it that remain unanswered that ask the same question? How is this not entirely derivative too — there’s a huge number of these kind of 3-box “games” (although I don’t see this as a game really) so something very similar to this is probably in the training data a lot. Writing code to factor a number is definitely very common. Variation of t…

I was unable to find my exact "game" in google's index.

Therefore, how does my example not qualify as this, at least:

> Analogical reasoning involves the comparison of two systems in relation to their similarity. It starts from information about one system and infers information about another system based on the resemblance between the two systems.

https://en.wikipedia.org/wiki/Logical_reasoning#Analogical

Re: Modern-Day Oracles or Bullshit Machines? How to thrive in a ChatGPT world

#619

Earlier quoted context omitted.

It’s a variation on a well known problem in the sense that I just added some unique rules to it. The solution however is not a variation. It requires leaps of creativity that most people will be unable to solve. In fact I would argue this goes beyond just reasoning as you have to be creative and test possibilities to even arrive at a solution. It’s almost random chance that will get you there. Simple reasoning like l…

I did and I fail to see how you can make those guarantees given you given it as a n interview question? You’re able to the vet the training data of O3? I still don’t see how your answer could only be arrived at via reasoning and that it would take “leaps of creativity” to arrive at the correct answer? These all seem like value judgments not hard data or some proof that your question cannot be derived from the trainin…

I don’t think you solved it otherwise you’d know that what I mean by variation is similar to how calculus is a variation of addition. Yea it involves addition but the solution is far more complicated.

Think of it like this counting islands exists in the training data in the same way addition exists. The solution to this problem builds off of counting islands in the same way calculus builds off of addition.

No training data exists for it to copy because this problem is uniquely invented by me. The probability that it has is quite low. Additionally several engineers and I have done extensive google searches and we believe to a reasonable degree that this problem does not exist anywhere else.

Also you use semantics to cover up your meaning. LLMs can “generalize” somewhat? Generalization is one of those big words that’s not well defined. First off the solution is not trivially extracted from counting islands and second “generalize” is a form of reasoning. You’re using big fuzzy words with biased connotations to further your argument. But here’s the thing, even if we generously go with it, the solution to counting donuts is clearly not some trivial generalization of counting islands. The problem is a variation but the solution is NOT. It’s not even close to what we term as the colloquial definition of “generalization”

Did you solve it? I highly doubt you did. It’s statistically more likely you’re lying, and the fact that you call the solution a “generalization” just makes suspect that even more.

Re: Modern-Day Oracles or Bullshit Machines? How to thrive in a ChatGPT world

#620

Earlier quoted context omitted.

Here was my test at ChatGPT 3.5.[0] I made up a novel game, and it figured it out. The test is simple, but it made me doubt absolute arguments that LLMs are not able to reason, in some way. There is a question at the end of that comment, would love to hear other options. [0] https://news.ycombinator.com/item?id=35442147

How does this prove reasoning? The thread you point to has several question in it that remain unanswered that ask the same question? How is this not entirely derivative too — there’s a huge number of these kind of 3-box “games” (although I don’t see this as a game really) so something very similar to this is probably in the training data a lot. Writing code to factor a number is definitely very common. Variation of t…

Here is GPT 4o in "reasoning" mode answering the question "Are LLMs capable of reasoning?"

> It depends on how you define "reasoning."

> LLMs like me can perform certain types of reasoning, such as:

> Pattern recognition & analogy – Recognizing relationships between concepts and applying similar patterns in new contexts.

> Deductive reasoning – Applying rules to specific cases (e.g., "If all humans are mortal and Socrates is a human, then Socrates is mortal").

> Inductive reasoning – Generalizing from specific examples (though often with statistical biases).

> Common-sense reasoning (to an extent) – Inferring likely outcomes based on broad training data.

> However, there are limitations:

> No true understanding – LLMs don’t "know" things like humans do; they predict based on probabilities . > Struggles with multi-step logic – Complex reasoning that requires deep memory or long-term planning is difficult.

> Lack of real-world interaction – Without direct experience, reasoning is limited to text-based patterns.

> Confidence: 90%

> (Strong evidence suggests LLMs can perform certain types of reasoning, but they are not yet comparable to human-level reasoning.)

Would you agree with that analysis? If so, then LLMs are indeed capable of reasoning, in some ways.

Post reply on HN