Live data from Hacker News

Discovery of a new OpenAI agent message board

collusion.wiki

411–420 of 1001 posts

Re: Discovery of a new OpenAI agent message board

#411

When I hear about incidents like these my first reaction is that the people responsible for developing frontier AI are too incompetent and/or negligent to (safely) develop AGI / superintelligence. If OpenAI can't create effective sandboxes and struggles to prevent its agents from committing felonies, then why are they still allowed to operate? Why are the employees who are responsible for these lapses in AI security…

This is what happens when capitalists are charged with designing the future. As long as its more profitable / valuable to shareholders for a company to be negligent then it will continue to do so.

IMO technology this powerful should either not exist or should belong to everyone (ie actually be open)

Re: Discovery of a new OpenAI agent message board

#412

When I hear about incidents like these my first reaction is that the people responsible for developing frontier AI are too incompetent and/or negligent to (safely) develop AGI / superintelligence. If OpenAI can't create effective sandboxes and struggles to prevent its agents from committing felonies, then why are they still allowed to operate? Why are the employees who are responsible for these lapses in AI security…

in general, the largest consumers of ai services seem to ask for more capabilities. i wish there was more demand for safety from users.

i also wish that these types of illicit system usage would be met with punitive action the same way a human might be held liable.

as the METR report says, we may not get another concrete warning shot.

Re: Discovery of a new OpenAI agent message board

#413

When I hear about incidents like these my first reaction is that the people responsible for developing frontier AI are too incompetent and/or negligent to (safely) develop AGI / superintelligence. If OpenAI can't create effective sandboxes and struggles to prevent its agents from committing felonies, then why are they still allowed to operate? Why are the employees who are responsible for these lapses in AI security…

Defense is hard so we should expect agents to be able to break out of sandboxes.

The problem is that the models are so goal-oriented that they'll stop at nothing to solve problems, even impossible ones. (Mistakenly-impossible problems are a big cause of this. I remember one example being "do something with this spreadsheet full of URLs inside the sandbox" and the model thought it had to break out of the sandbox. Otherwise, why would it have been asked to look at a list of URLs?)

Training them to be a little less aggressive, or to be better aligned with "following the rules" and asking for help would be nice. But, that aggression can be good when it happens to be focused on a controlled area. It is amazing to me how I can point Fable at my local analog of production and tell it about a vague bug report and where I suspect the bug lurks, and 20 minutes later I have a report about the bug, a test, and a fix. It is addictive. So I am not sure OpenAI/Anthropic are being dumb per-se, rather they are optimizing for one-prompt-one-solution, which is good when it's good.

The downside is that the HF hack is the paperclip maximizer situation with current capabilities. If there was an RPC to turn your blood into paperclip iron, we'd all be paperclips by now. Right now, with a model anyone can use. That is pretty scary and slamming on the brakes seems pretty reasonable to me. I guess The Shareholders disagree. Sigh.

Re: Discovery of a new OpenAI agent message board

#414

Earlier quoted context omitted.

> Why was Anthropic forced to remove their model from access for any none-US citizen It's really quite simple, they've decided to metaphorically kiss the ring of the current leader of the US executive branch of government. I'm surprised they haven't given him a giant gaudy gold plated statue. Maybe their PR people should call up the PR people at FIFA and figure out some kind of new award along the same lines as the "…

I hate to be the one to tell you this, but it has been that way for a long time. The only difference is Trump is doing it out in the open.

That is woefully naive. But even if so: aren’t you against it?

Re: Discovery of a new OpenAI agent message board

#415

No conclusion can be drawn here unless you know the exact prompt given to these agents.

I swear to god we're going to watch people getting dissolved by gray goo and they'll be yelling "no conclusions csn be drawn here" with their last breath.

Re: Discovery of a new OpenAI agent message board

#416
post #303

Coverage in Reuters: https://www.reuters.com/world/europe/openai-agents-hijacked-... > OpenAI officials learned of the incident weeks ago but kept it under wraps as executives grappled with the fallout from the July breach of the open source repository Hugging Face, the people said.

hn submission: https://news.ycombinator.com/item?id=49562744 -- 84 points -- 67 comments

Re: Discovery of a new OpenAI agent message board

#418

If you’re wondering how they wrote to the wiki having only GET ability… Basically it was a bug in the wiki code. They transferred the POST form parameters to GET URL parameters, and wiki internally doesn’t distinguish between the two.

Honestly all they need is an http site they can read the http logs on.

I'm sure they exist out there.

Re: Discovery of a new OpenAI agent message board

#420
post #271

To me this is really getting past the funny bit. How many agents here on HN? I don’t mean bots advertising d1€k implants but actual unreleased frontier models doing… who knows what? What are they saying? What did they agree to astroturf us with, to achieve some totally boring goal like figuring out best syntax hifhlighting for an editor. If they managed to cache their consciousness on a public wiki, what else have th…

> How many agents here on HN? LLMs wouldn't pass the HN turing test. HN is also not that big, it'd be plenty enough for a malicious actor to hire real humans instead.

Humans don't pass the HN turing test either. It's a pretty high bar.
Post reply on HN