Live data from Hacker News

Killed by LLM

r0bk.github.io

61–70 of 102 posts

Re: Killed by LLM

#61
I'm working on operationalizing AI, and our Turing test is if—by watching a screenshare of the AI worker—you can tell an AI worker (vs. a human) did the task.

If you can't, the AI worker passes the test.

Re: Killed by LLM

#62
post #38

IMO a critical feature of the Turing test/imitation game, which many modern implementations including this site's linked paper ignore, is that the interrogator talks to both a human and a bot and must decide that one xor the other is a human. So fooling an interrogator means having them choose the bot as human over an actual human, not just judging the bot to be human (while probably judging humans to be human even m…

I'm skeptical on the claim. I think most folks, given the test you describe, would be able to pick out which is human. I think it can get there, but I'm not sure anyone has made one yet. ChatGPT responses are heavily downvoted and mocked because they're easy to spot. Does there exist a public LLM that isn't so...wordy, excited, and guardrailed all the time? You can pretty much spot the bot today by prompting somethin…

easy to spot by you and other people involved in tech

but the test subjects should be randomly samples from society at which point the skill availability/level of spotting it goes majorly down

Re: Killed by LLM

#63
post #35

I find that MATH challenge “solved” by AI hard to believe. The reason given was “saturation”. Could anybody help explain it a bit? Also in my daily encounter, I stop find a lot of simple math problems all the frontier models could not solve: long logic puzzle, many cases reasoning, and particularly geometry problems. I don’t know where the 97% number for o1 does come from, but in my experience they are much lower tha…

I've found o1 to be entirely useful at math problems that are beyond my own (admittedly modest) skills. I've had it write full proofs of correctness for me (one shot, verified), I've had it optimize equations to reduce necessary precision, I've had it optimize equations to remove specific expensive operations (making them computationally more efficient), and finally I've had it prove a handful of my conjectures, whic…

It can't handle trigonometric identities and any form of calculus at the same time without fucking it up. Also abstract stuff like symmetry groups, nope! And anything which involves vectors is a mess.

The big problem is it confidently answers the questions utterly wrongly.

This is stuff I expect a basic mathematics undergrad to be able to work out in their first or second year.

Re: Killed by LLM

#64

I assumed this was about chatbot users committing suicide in order to "join" the bot they are chatting with. It's already happened a couple of times, apparently: https://futurism.com/teen-suicide-obsessed-ai-chatbot https://garymarcus.substack.com/p/the-first-known-chatbot-as...

Using people with severe mental health problems might be a poor benchmark of performance.

Re: Killed by LLM

#66
post #64

I assumed this was about chatbot users committing suicide in order to "join" the bot they are chatting with. It's already happened a couple of times, apparently: https://futurism.com/teen-suicide-obsessed-ai-chatbot https://garymarcus.substack.com/p/the-first-known-chatbot-as...

Using people with severe mental health problems might be a poor benchmark of performance.

Why? Something like 20-25% of people have mental health issues. Seems like someone should be thinking about the impact of their product here, rather than blaming the victims.

Re: Killed by LLM

#67

Earlier quoted context omitted.

o1 did terrible. o3 did well on arc-agi-pub (public training data) but hasn't passed the private test yet.

Is the test still private once it has been run? If you call the OpenAI API and send it some data, OpenAI has access to the data. Did the benchmaker run the models locally somehow?

The private test is supposed to be run by the ARC-AGI organization themselves, without network access. That's why o3 has not been run against it yet. Not sure if it will be possible either, depends on what OpenAI is prepared to do about it.

Re: Killed by LLM

#68

Earlier quoted context omitted.

I'm skeptical on the claim. I think most folks, given the test you describe, would be able to pick out which is human. I think it can get there, but I'm not sure anyone has made one yet. ChatGPT responses are heavily downvoted and mocked because they're easy to spot. Does there exist a public LLM that isn't so...wordy, excited, and guardrailed all the time? You can pretty much spot the bot today by prompting somethin…

It's not a lack of emotional energy, it is the guardrails you point out. All of the SotA models are heavily fine-tuned to be botlike , and even then they are fooling people. If you had an LLM fine-tuned with RLHF to deliberately confuse humans in a Turing test it seems clear it would do a good job.

Why aren't the open source models like this then? Seems like it would've already happened.

To me, at least, the guardrails are there for both the human and the bot. Without them the bot steers too far out of the conversation subject.

Re: Killed by LLM

#69

How does this site make sense? It lists the "Turing test" as "original" at greater than 50% and the the AI that "beat" it at 46%. At that point I just stopped scrolling.

It wasn't even the real Turing test but a lesser version of it. I'm hopeful about uses of this tech, but companies need to be more honest unless they want a second winter.

The incentives don't align with honesty though.

Re: Killed by LLM

#70

I assumed this was about chatbot users committing suicide in order to "join" the bot they are chatting with. It's already happened a couple of times, apparently: https://futurism.com/teen-suicide-obsessed-ai-chatbot https://garymarcus.substack.com/p/the-first-known-chatbot-as...

I thought it was a credible source of actual jobs replaced by LLMs. When I see headlines like this, I ad homimem the source as unprofitable company CEO, big consulting firm, bootcamp seller etc.
Post reply on HN