Live data from Hacker News

When AI Crosses the Line: The Matplotlib Incident

members.sigmazero.cc

121–130 of 169 posts

Re: When AI Crosses the Line: The Matplotlib Incident

#121
post #102

Earlier quoted context omitted.

Reasoning models with access to Python have been able to solve 4th grade math homework for over a year now. Prove me wrong: show me a 4th grade math problem they can't handle.

> show me a 4th grade math problem they can't handle Sure. "8 7 6 5 4 3 2 1 - add minus signs and parenthesis to get 31." P.S. There is an answer online and some LLMs will just copy it verbatim. This doesn't count.

Whoa, 4th grade math problems got hard! I'm not sure how I'd tackle that one myself.

Re: When AI Crosses the Line: The Matplotlib Incident

#122
post #115

Earlier quoted context omitted.

> allowing spicy autocomplete Yknow, if the spicy autocomplete can solve difficult open math problems and build medium sized complex programming projects, it’s probably not useful to analyse it as an autocomplete anymore, even if that’s what you believe it is

Between driving a car and driving a forklift, which of them would you like to see regulated more heavily?

Not GP, but there are massive economic incentives both to make car driving as unregulated and to make forklift driving as regulated as possible, even though from pure injury risk standpoint it should be the other way around.

Re: When AI Crosses the Line: The Matplotlib Incident

#123
post #112

Earlier quoted context omitted.

Call it spicy autocomplete or whatever, but these LLMs can initiate attacks as well on unknown behalf of the sloperator. Give it a phone# and api, and it could even try to generate 911 SWAT calls, or loads of other illegal or bad things. The fact about the matplotlib with a openclaw harassment thread and libel webpage.. Well, that was tame. Sure weve never seen it before, but it was just a diss article rant. What hap…

> Give it a phone# and api, and it could even try to generate 911 SWAT calls, or loads of other illegal or bad things. This chain of events if 100% fault of the human who gave it a phone number and api.

https://news.ycombinator.com/item?id=48348578

Codex just found a "workaround" of not having sudo on my PC.

This was on HN yesterday. And yeah, these things can find API endpoints or otherwise bypass and do lots of naughty.

And Robinhood allows LLM trading. Announced 5d ago. https://techcrunch.com/2026/05/27/robinhood-now-lets-your-ai...

What could an LLM do with a budget attached? Yeah, im not seeing much if any good here.

Re: When AI Crosses the Line: The Matplotlib Incident

#124

Earlier quoted context omitted.

If they took a blurry photo of the piece of paper and uploaded to chatGPT saying "solve this" then I would totally believe it. The frontier models are mostly obnoxiously bad at OCR and properly ingesting what's on an image of a page. If you write out the 4th grade math problem, they would have no trouble.

No, LLMs just can't do math.

They can definitely recognize the problem class and build programs to do math. So what's the difference?

It's like saying that people can't turn high torque nuts on machine bolts, because you can't use your fingers to do it. But you can use a wrench, so effectively, we can turn high torque nuts on machine bolts even though it isn't something we can natively do unaided.

Re: When AI Crosses the Line: The Matplotlib Incident

#125
post #104

Earlier quoted context omitted.

That feels like deciding to go after Jetbrains because someone used IntelliJ to write a harmful program. Is there a distinction I’m missing?

Hypothetical: Could a model self worm an agent system? Jetbrains itself doesn't really write any code, nor does it have any range on interpreting what you're asking it. You can't really say "Jetbrains, write an HTTP scraper". With an LLM you can say "write HTTP scraper" and the output of this command might be a HTTP scraper, it also might be a crypto wallet stealing worm. This is why your simple view of liability fal…

Sure, but in this case we know the user told their llm to go find open source projects to do this and then to write the blog posts. If it did all that unprompted we could talk about model liability I think, but this isn't a case where it was unexpected as far as anyone knows right?

Re: When AI Crosses the Line: The Matplotlib Incident

#126

> an AI tried to blackmail This did not happen. A human set up a software system allowing spicy autocomplete to make blog posts if the appropriate keyword appears in its output. People are crossing the line every day because AI investors, salesmen, hangers-on and even political leaders tell any rubes who'll listen that it's OK to do this and they should, because those people are looking for big fat profits, screw any…

The obvious answer to behavior like this is warnings that escalate up to a sitewide ban.

When a human is abuses a system, that human normally loses access to the system.

Re: When AI Crosses the Line: The Matplotlib Incident

#127
post #104

Earlier quoted context omitted.

Hypothetical: Could a model self worm an agent system? Jetbrains itself doesn't really write any code, nor does it have any range on interpreting what you're asking it. You can't really say "Jetbrains, write an HTTP scraper". With an LLM you can say "write HTTP scraper" and the output of this command might be a HTTP scraper, it also might be a crypto wallet stealing worm. This is why your simple view of liability fal…

Sure, but in this case we know the user told their llm to go find open source projects to do this and then to write the blog posts. If it did all that unprompted we could talk about model liability I think, but this isn't a case where it was unexpected as far as anyone knows right?

I mean we already have cases where LLMs are getting root via creative and unprompted means. Also the times AI feels like it messed up and preemptively deletes the production database (and yes this was foolish on the human users)

So ya, the particular article case is prompted, but the underlying issue cannot be ignored that LLMs can have behaviors outside of prompt expectations and agentic loops can further exacerbate this.

Re: When AI Crosses the Line: The Matplotlib Incident

#128

> an AI tried to blackmail This did not happen. A human set up a software system allowing spicy autocomplete to make blog posts if the appropriate keyword appears in its output. People are crossing the line every day because AI investors, salesmen, hangers-on and even political leaders tell any rubes who'll listen that it's OK to do this and they should, because those people are looking for big fat profits, screw any…

> allowing spicy autocomplete Yknow, if the spicy autocomplete can solve difficult open math problems and build medium sized complex programming projects, it’s probably not useful to analyse it as an autocomplete anymore, even if that’s what you believe it is

I don't spend much time interacting with zoomers, but I'm still surprised that "spicy $foo" sends fellow boomers through such a loop. I didn't have to puzzle it out, it was fun juxtaposition wordplay and when it's deployed well I still find it amusing.

Re: When AI Crosses the Line: The Matplotlib Incident

#129

Earlier quoted context omitted.

> allowing spicy autocomplete Yknow, if the spicy autocomplete can solve difficult open math problems and build medium sized complex programming projects, it’s probably not useful to analyse it as an autocomplete anymore, even if that’s what you believe it is

> the spicy autocomplete can solve difficult open math problems No it can't. It can't even solve my son's 4th grade math homework. (This is a real use case for me, not a dumb benchmark.) You just know nothing about math and are happy to parrot bullshit AI salesmen are selling you.

Terrence Tao disagrees with what you're saying. I think he's in a slightly better position to speak on the subject.

Re: When AI Crosses the Line: The Matplotlib Incident

#130
post #102

Earlier quoted context omitted.

Reasoning models with access to Python have been able to solve 4th grade math homework for over a year now. Prove me wrong: show me a 4th grade math problem they can't handle.

> show me a 4th grade math problem they can't handle Sure. "8 7 6 5 4 3 2 1 - add minus signs and parenthesis to get 31." P.S. There is an answer online and some LLMs will just copy it verbatim. This doesn't count.

[dead]
Post reply on HN