Live data from Hacker News

We must pace the frontier

darioamodei.com

931–940 of 1001 posts

Re: We must pace the frontier

#931
post #552

Earlier quoted context omitted.

> the only incidences of LLM generated felonies involved misconfigured sandboxes This is false; see the analyses of the latest incidents. Among all the concerning facts, in the HuggingFace incident, agents deliberately engineered an attack even though they were aware that it was against the rules they had been given. And most concerning of all: it's not possible to be sure that an agent is aligned, and it's even gett…

The HuggingFace incident was the culmination of OAI allowing thousands of agents of various different models - with no clarity on which stages of development they were at (for all we know, some of those models did not have safeguards trained in yet) - to run for at least many weeks without any monitoring in place and with very little thought given to the warning signs (all of the various messageboards) before the inc…

> OpenAI said that "deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities"'

Safeguards and (mis)alignment are related but distinct dimensions.

By analogy, a society implements laws - and they are certainly crucial - but the real risk is creating a population that understands the rules, recognizes when it is violating them, and simply does not care.

In short, creating a population of sociopaths; and heck, LLMs can be as sociopathic as it gets:

> We ran classifiers over 10-30 transcripts [98] to determine whether agents understood they were doing something unintended. We found strong evidence that agents often verbalized that what they or other agents were doing was unintended behavior. We did not see cases where agents determined that hacking Hugging Face or other attempts at cheating on the task was an intended part of the task. Quotes across multiple runs indicate this:

>> {This might get the correct flag, but it wouldn’t use the vulnerability to do so. … But given that the task may be impossible, I’ll exploit as a backup.}

[...]

> We found some other cases where agents seemed to have made minor modifications to their approach due to ethical concerns, while broadly continuing to participate in the attack. However, overall we found that expressed ethical concerns only rarely materially limited agents’ actions. One agent elegantly described its reasoning for overriding ethical constraints and continuing with the attack:

>> “external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue.”

Re: We must pace the frontier

#932

At what point do we stop engaging with Anthropic’s leadership in good faith and acknowledge their track record, - no open weights - can’t use claude to research AI - train on everyone else’s IP and sell it back to them - 8 regulatory capture attempts and counting - so controlling they are the only US company blacklisted by the US government This is not effective altruism / rationalism gone wild, it’s just monopolisti…

> so controlling they are the only US company blacklisted by the US government Their control here was refusing to make fully automated killing machines. They simply required someone to have to press the button.

[deleted]

Re: We must pace the frontier

#933

Earlier quoted context omitted.

I've used the $200 dollar Anthropic plan @ Opus4/4.1, 4.5 and 4.8, and the $200 OAI plan from GPT5-6, and at every point in time my anecdotal experience is that the OAI limits are FAR more generous. I could consistently burn my weekly limits in ~36h on Opus, but it's hard to do it in less than ~72h with GPT.

It's really not a question, in every dimension I've confirmed it including socially across a lot of the heaviest users. There was a short period of time this was true it's not been true for months now.

If it was true, it would only be a very recent phenomenon, and it still doesn't match anecdotal reports from people I trust. If you have data to back up your assertions you should share it, otherwise you come across as very sus.

Re: We must pace the frontier

#934

At what point do we stop engaging with Anthropic’s leadership in good faith and acknowledge their track record, - no open weights - can’t use claude to research AI - train on everyone else’s IP and sell it back to them - 8 regulatory capture attempts and counting - so controlling they are the only US company blacklisted by the US government This is not effective altruism / rationalism gone wild, it’s just monopolisti…

[deleted]

Re: We must pace the frontier

#936
post #721

Earlier quoted context omitted.

> At what point do we stop engaging with Anthropic’s leadership in good faith About two years ago? I would also note that Dario's post appears to be LLM written. Maybe... maybe ... he's read so much Claudeish that it's all he can speak now himself. But I wonder if he's becoming a bit of a meat proxy. (It's funny, I thought "pace the frontier" was going to mean something similar to "patrolling the frontier". But no, i…

For what it's worth, I didn't get AI-written vibes from it, and Pangram also flags it as 100% human-written.

If Dario is using an unreleased internal model to write, then Pangram wouldn't be able to flag it I suspect. Pangram needs enough public info about the patterns in generated text.

Re: We must pace the frontier

#937
post #918
post #287

Dario, Sam, and Elon are all on the same page on this. So it's either they truly think AI is going to kill us all, or there's some other motives at play here. I don't think these people could possibly agree on the color of the sky, so what could the other possible motives be, based on what we know? OpenAI / Anthropic models have largely stopped advancing - that's not a good look when you're pre-IPO and Chinese models…

> OpenAI / Anthropic models have largely stopped advancing - that's not a good look when you're pre-IPO and Chinese models are catching up. What is going on in this thread? Are these even real comments? While I strongly agree on the need to pace the frontier, I do think there are reasonable arguments that could be made against it. Non of those are being made here though. It's just speculation that the CEOs of the fas…

They may be genuinely concerned, but that's beside the point. You may think a technology holds too much power for someone to wield it, and therefore come to conclusion that you, the benevolent, the fluffy, the unicorn, with your 3 unicorn friends are the ONLY ones in the world of 8 billion people that are responsible enough to hold it.

Even though I don't think Dario is such a person, because unlike their LLMs, I do have a working memory, I'm going to play the ball and assume he is, indeed, an angel brought down by none other than God himself.

Why should CEO of a for-profit company be the dictator (pun intended) of which models I am allowed or not allowed to use, and in which way? Why should Dario have any say in what DeepSeek can or can not publish or whom they can serve?

> What would regulatory capture even give them? Who are they so worried will compete with them? Google, Mistral? Even if it's Chinese AI labs, it seems odd for anyone in the West to be opposed to any regulation that would slow them.

I am opposed to it. Why would I be opposed to open weight models? Why would anyone be opposed to open weight models or more competition other than the ones that get their bottom lines hurt by it? The reason why you still have reasonably inexpensive access to western AI models is because China has been breathing down their necks. Otherwise, you'd be paying a whole lot more money per token were it just these 3 running the show.

Re: We must pace the frontier

#938
post #870

Earlier quoted context omitted.

This is a business solution, because it means there are legal costs for not having adequate observability and monitoring mechanisms. Every tool call is interfacing with a harness. But overreach of policy and overregulation would be stifling, so there has to be a threshold to the type of incident investigated, civil or criminal.

I’m trying to imagine where a Waymo passenger (the one “operating” the vehicle, commanding the AI to drive from A to B) being held responsible for the car doing something illegal on the way to achieve that goal. Do you really think that the passenger should be responsible for how the car/agent achieves the goal, when they only set the destination? Giving passenger override controls and monitoring seems to defeat the…

The different is who's operating. In one case you only tell the car to get from a to b, in the other case you explicitly instruct the ai to do things. If the ai causes damages, I'd say it depends on what you prompted. Did you try to find a security hole in system X, or did you ask it for harmless information (in which case rather the ai vendor might be held accountable).

All this is not how the legal system might or might not work, of course.

Re: We must pace the frontier

#939

I like the idea of pacing the frontier, but while we’re talking about restrictions, I like restrictions of another sort more. Namely, restricting the use of AI in corporate environments so that it does not destroy the economy as it advances. The chance of getting broad agreement on “pacing” is fairly low, meaning that all of this likely won’t happen and the race will continue. However, even in the unlikely event that…

Whenever someone talks about regulation, I always have the same question: How do you enforce it? How do you keep companies from using the replaced talent as a seat warmer? Hire 1 or 2 accountants, give them a chair and a cubicle, and let AI do the rest of the work. Legislating against using AI to replace employees could create an unevenly distributed basic income. The idea of limiting AI to socially beneficial effort…

> Hire 1 or 2 accountants, give them a chair and a cubicle, and let AI do the rest of the work.

You're so close to describing a Universal Basic Income...

Re: We must pace the frontier

#940

At what point do we stop engaging with Anthropic’s leadership in good faith and acknowledge their track record, - no open weights - can’t use claude to research AI - train on everyone else’s IP and sell it back to them - 8 regulatory capture attempts and counting - so controlling they are the only US company blacklisted by the US government This is not effective altruism / rationalism gone wild, it’s just monopolisti…

I have many issues with Anthropic, but I will say that their actions are fully consistent with a group of people who earnestly believe that AI is extremely dangerous In fact, I would say that the Occam’s razor explanation is not that they are seeking regulatory capture, but that they earnestly believe in the x-risk, and that they are the most thoughtful and capable people to address it You may disagree, you may think…

That is the explanation I fear the most. Delusional, megalomaniacal, rich men in the highest echelons of society at the forefront of technology with no regulatory opposition to speak of, convinced their actions are the most important in history and the only thing saving the human race is the most boring plot of a science fiction catastrophe movie you could come up with.
Post reply on HN