Live data from Hacker News

Killed by LLM

r0bk.github.io

31–40 of 102 posts

Re: Killed by LLM

#32
I find that MATH challenge “solved” by AI hard to believe. The reason given was “saturation”. Could anybody help explain it a bit? Also in my daily encounter, I stop find a lot of simple math problems all the frontier models could not solve: long logic puzzle, many cases reasoning, and particularly geometry problems. I don’t know where the 97% number for o1 does come from, but in my experience they are much lower than that and math, even elementary maths certainly can not be considered to be “solved”. As far as I can see, OpenAI has been trained their models on all these public problems, so testing on them to record a benchmark is tainted as best when not outright cheating.

Re: Killed by LLM

#33

I don't really understand why "Killed by: Saturation" is needed - what other options are there? It would also be nice to see the "unbeaten" list: standardized tests LLMs still fail (for now). e.g. Wozniak's coffee test.

Wozniak's coffee test would be a really fun one to attempt. As long as you could get a capable enough robot, I imagine it's possible. Something like the Spot Arm[1] would be sufficient.

Something like:

- Key the robot controls to a series of tools (move_forward(x), extend_arm(y))

- Add a camera and pass each frame to the AI model along with the task "make a cup of coffee" and the list of available tools it can call.

And it would likely succeed some percentage of the time today!

[1] https://bostondynamics.com/products/spot/arm/

Re: Killed by LLM

#34

Everything says "killed by saturation". Is there another way to be killed?

There are benchmarks that humans score close to zero on average and the top LLM scores 25%. https://epoch.ai/frontiermath/the-benchmark

If anyone from epoch.ai is reading this, it would be nice to link the toplevel result for o3 to this page.

Re: Killed by LLM

#35

I find that MATH challenge “solved” by AI hard to believe. The reason given was “saturation”. Could anybody help explain it a bit? Also in my daily encounter, I stop find a lot of simple math problems all the frontier models could not solve: long logic puzzle, many cases reasoning, and particularly geometry problems. I don’t know where the 97% number for o1 does come from, but in my experience they are much lower tha…

I've found o1 to be entirely useful at math problems that are beyond my own (admittedly modest) skills. I've had it write full proofs of correctness for me (one shot, verified), I've had it optimize equations to reduce necessary precision, I've had it optimize equations to remove specific expensive operations (making them computationally more efficient), and finally I've had it prove a handful of my conjectures, which was helpful for taking algorithmic shortcuts in a security sensitive environment.

Mostly all algebra and calculus, but definitely all problems that most undergrads would struggle with.

It's most useful because it has deep knowledge of related and adjacent conjectures that are well understood, even if you've never heard of them. So it can mix and match things with a lot more ease than a tinkering mathematician

Re: Killed by LLM

#36

The tortoise lays on its back, its belly baking in the hot sun, beating its legs trying to turn itself over, but it can't. Not without your help. But you're not helping.

Describe in single words, only the good things that come into your mind about your mother.

[flagged]

Re: Killed by LLM

#37

I find that MATH challenge “solved” by AI hard to believe. The reason given was “saturation”. Could anybody help explain it a bit? Also in my daily encounter, I stop find a lot of simple math problems all the frontier models could not solve: long logic puzzle, many cases reasoning, and particularly geometry problems. I don’t know where the 97% number for o1 does come from, but in my experience they are much lower tha…

Scroll down on the page. It explains saturation.

Re: Killed by LLM

#38
IMO a critical feature of the Turing test/imitation game, which many modern implementations including this site's linked paper ignore, is that the interrogator talks to both a human and a bot and must decide that one xor the other is a human. So fooling an interrogator means having them choose the bot as human over an actual human, not just judging the bot to be human (while probably judging humans to be human even more frequently).

When the interrogator is only answering "do you think your conversation partner was a human?" individually, bots can score fairly highly simply by giving little information in either direction - like pretending to be a non-english-speaking child, or sending very few messages.

Whereas when pitted against a human, the bot is forced to give stronger or equally strong evidence of being human as the average human (over enough tests). To be chosen as human, giving 0 evidence becomes a bad strategy when the opponent (the real human) is likely giving some positive non-zero evidence towards their personhood.

Re: Killed by LLM

#39

The tortoise lays on its back, its belly baking in the hot sun, beating its legs trying to turn itself over, but it can't. Not without your help. But you're not helping.

Describe in single words, only the good things that come into your mind about your mother.

Let me tell you about my mother

Re: Killed by LLM

#40
post #38

IMO a critical feature of the Turing test/imitation game, which many modern implementations including this site's linked paper ignore, is that the interrogator talks to both a human and a bot and must decide that one xor the other is a human. So fooling an interrogator means having them choose the bot as human over an actual human, not just judging the bot to be human (while probably judging humans to be human even m…

I'm skeptical on the claim. I think most folks, given the test you describe, would be able to pick out which is human. I think it can get there, but I'm not sure anyone has made one yet. ChatGPT responses are heavily downvoted and mocked because they're easy to spot.

Does there exist a public LLM that isn't so...wordy, excited, and guardrailed all the time?

You can pretty much spot the bot today by prompting something horribly offensive. Their response is always very inhuman, probably due to lack of emotional energy.

Post reply on HN