Live data from Hacker News

Killed by LLM

r0bk.github.io

51–60 of 102 posts

Re: Killed by LLM

#51
post #38

IMO a critical feature of the Turing test/imitation game, which many modern implementations including this site's linked paper ignore, is that the interrogator talks to both a human and a bot and must decide that one xor the other is a human. So fooling an interrogator means having them choose the bot as human over an actual human, not just judging the bot to be human (while probably judging humans to be human even m…

I'm skeptical on the claim. I think most folks, given the test you describe, would be able to pick out which is human. I think it can get there, but I'm not sure anyone has made one yet. ChatGPT responses are heavily downvoted and mocked because they're easy to spot. Does there exist a public LLM that isn't so...wordy, excited, and guardrailed all the time? You can pretty much spot the bot today by prompting somethin…

You can get rid of OpenAI's wordy, excited and guardrailed responses with the eigenrobot prompt, for instance.

https://x.com/eigenrobot/status/1870696676819640348

I generally prefer it to the default. It doesn't work as well on Claude or Grok for various reasons. I think it really shines on GPT o1-mini and GPT 4o.

Re: Killed by LLM

#52
post #38

IMO a critical feature of the Turing test/imitation game, which many modern implementations including this site's linked paper ignore, is that the interrogator talks to both a human and a bot and must decide that one xor the other is a human. So fooling an interrogator means having them choose the bot as human over an actual human, not just judging the bot to be human (while probably judging humans to be human even m…

I'm skeptical on the claim. I think most folks, given the test you describe, would be able to pick out which is human. I think it can get there, but I'm not sure anyone has made one yet. ChatGPT responses are heavily downvoted and mocked because they're easy to spot. Does there exist a public LLM that isn't so...wordy, excited, and guardrailed all the time? You can pretty much spot the bot today by prompting somethin…

It's not a lack of emotional energy, it is the guardrails you point out. All of the SotA models are heavily fine-tuned to be botlike, and even then they are fooling people. If you had an LLM fine-tuned with RLHF to deliberately confuse humans in a Turing test it seems clear it would do a good job.

Re: Killed by LLM

#55

Posted by Chollet himself: > I don't think people really appreciate how simple ARC-AGI-1 was, and what solving it really means. It was designed as the simplest, most basic assessment of fluid intelligence possible. Failure to pass signifies a near-total inability to adapt or problem-solve in unfamiliar situations. > Passing it means your system exhibits non-zero fluid intelligence -- you're finally looking at somethi…

> Failure to pass signifies a near-total inability to adapt or problem-solve in unfamiliar situations. Not necessarily. Get a human to solve ARC-AGI if the problems are shown as a string. They'll perform badly. But that doesn't mean that humans can't reason. It means that human reasoning doesn't have access to the non-reasoning building blocks it needs (things like concepts, words, or in this case: spatially local an…

ARC-AGI feels like it would fall to a higher dimensional convolution rather than reasoning.

Re: Killed by LLM

#56

I didn't know ARC-AGI had been "beaten" by o3. What are the next challenges that frontier models like o1/o3 are faced with?

o1 did terrible. o3 did well on arc-agi-pub (public training data) but hasn't passed the private test yet.

Is the test still private once it has been run? If you call the OpenAI API and send it some data, OpenAI has access to the data. Did the benchmaker run the models locally somehow?

Re: Killed by LLM

#57
post #38

IMO a critical feature of the Turing test/imitation game, which many modern implementations including this site's linked paper ignore, is that the interrogator talks to both a human and a bot and must decide that one xor the other is a human. So fooling an interrogator means having them choose the bot as human over an actual human, not just judging the bot to be human (while probably judging humans to be human even m…

I'm skeptical on the claim. I think most folks, given the test you describe, would be able to pick out which is human. I think it can get there, but I'm not sure anyone has made one yet. ChatGPT responses are heavily downvoted and mocked because they're easy to spot. Does there exist a public LLM that isn't so...wordy, excited, and guardrailed all the time? You can pretty much spot the bot today by prompting somethin…

"pretending to be a non-english-speaking child" isn't a hypothetical, it's a real tactic that was annoyingly effective a while back.

Being uncooperative makes it really hard to tell anything about you, including whether you're real.

Re: Killed by LLM

#58

Earlier quoted context omitted.

I'm skeptical on the claim. I think most folks, given the test you describe, would be able to pick out which is human. I think it can get there, but I'm not sure anyone has made one yet. ChatGPT responses are heavily downvoted and mocked because they're easy to spot. Does there exist a public LLM that isn't so...wordy, excited, and guardrailed all the time? You can pretty much spot the bot today by prompting somethin…

Aren't you just describing those emails in a big corp that are supposedly still written by humans? Yes, they are wordy, excited, and guardrailed, but I don't think they are written by AI yet. I guess this is why LLMs are so feared by high school English teachers. Yes, they don't write well, but neither do their students.

Nowadays, it's usually the opposite: if the text is too good, everyone starts accusing it of being written by AI.

Re: Killed by LLM

#59
post #38

IMO a critical feature of the Turing test/imitation game, which many modern implementations including this site's linked paper ignore, is that the interrogator talks to both a human and a bot and must decide that one xor the other is a human. So fooling an interrogator means having them choose the bot as human over an actual human, not just judging the bot to be human (while probably judging humans to be human even m…

That's not the original Turing test either. The original imitation game as proposed by Turing involves reading a text transcript of a human and a computer and having the evaluator determine which is which. The evaluator does not interact directly with the conversing parties.

Where are you getting that? Turing's most famous paper is just as Ukv describes. The link on that site doesn't work for me, but the reference is buried in their source:

https://courses.cs.umbc.edu/471/papers/turing.pdf

In Turing's test, the forced binary choice means P(human-judged-human) + P(machine-judged-human) is necessarily equal to 100%. This gives the 50% threshold clear intuitive and mathematical significance.

In the bastardized test that GPT-4 "passed", that sum can be (and actually was) >100%. This makes the result practically impossible to interpret, since it depends on the interrogators' prior. The correct prior seems to be that it was human with p = 25%, though the paper doesn't say that explicitly, or say anything about what the interrogators were told. If the interrogators guessed mistakenly that it was 50% then that would lead them to systematically misjudge machines as humans, perhaps as observed.

The bastardized test is pretty bad, but treating the 50% threshold as meaningful there is inexcusable. I see the preprint hasn't yet passed peer review, and I'll regain some faith in social science professors if it never does. Of course the credulous media coverage is everywhere already, including the LLM training sets--so regardless of whether LLMs can pass the Turing test, they now believe they do.

Re: Killed by LLM

#60
post #35

I find that MATH challenge “solved” by AI hard to believe. The reason given was “saturation”. Could anybody help explain it a bit? Also in my daily encounter, I stop find a lot of simple math problems all the frontier models could not solve: long logic puzzle, many cases reasoning, and particularly geometry problems. I don’t know where the 97% number for o1 does come from, but in my experience they are much lower tha…

I've found o1 to be entirely useful at math problems that are beyond my own (admittedly modest) skills. I've had it write full proofs of correctness for me (one shot, verified), I've had it optimize equations to reduce necessary precision, I've had it optimize equations to remove specific expensive operations (making them computationally more efficient), and finally I've had it prove a handful of my conjectures, whic…

Yeah, I had (so far apparent, but still be verified) success with o1 teaching me the necessary physics and maths I need to solve my specific problem. This is definitely grad level stuff but well understood. My concern though is it's missing things that are more esoteric.
Post reply on HN