Live data from Hacker News

We must pace the frontier

darioamodei.com

931–940 of 946 posts

Re: We must pace the frontier

#931

Earlier quoted context omitted.

Wouldn't the simplest solution for stopping the proliferation of wanton felony generators just be holding operators liable for actions that their agents take? Then the issue is whether the liability is with the model provider or the end user. If you give an unfiltered agent an open-ended task and equip it with an environment that allows it to execute arbitrary code, a human needs to be held responsible. My understand…

It might be a legal solution, but its not a business solution. The end user being criminally responsible for not taking sufficient steps to contain an agent they didn't create and who's internal function they cannot observe or audit is just a giant liability machine.

This is a business solution, because it means there are legal costs for not having adequate observability and monitoring mechanisms. Every tool call is interfacing with a harness.

But overreach of policy and overregulation would be stifling, so there has to be a threshold to the type of incident investigated, civil or criminal.

Re: We must pace the frontier

#932

Earlier quoted context omitted.

> I don't think this adds up. The AGI blocks will have AGI kill bots and those that don't will have nukes. Doesn't matter how much "intelligence" you have stockpiled when nukes start raining down on you.

Unless the AGI is able to develop effective defences to intercept an incoming nuclear attack.

MAD is a belief system. It doesn't matter whether AI can effectively stop the retaliation, it is enough that people making the decisions believed that it can (are we there already?).

Re: We must pace the frontier

#933

Earlier quoted context omitted.

It might be a legal solution, but its not a business solution. The end user being criminally responsible for not taking sufficient steps to contain an agent they didn't create and who's internal function they cannot observe or audit is just a giant liability machine.

This is a business solution, because it means there are legal costs for not having adequate observability and monitoring mechanisms. Every tool call is interfacing with a harness. But overreach of policy and overregulation would be stifling, so there has to be a threshold to the type of incident investigated, civil or criminal.

I’m trying to imagine where a Waymo passenger (the one “operating” the vehicle, commanding the AI to drive from A to B) being held responsible for the car doing something illegal on the way to achieve that goal.

Do you really think that the passenger should be responsible for how the car/agent achieves the goal, when they only set the destination?

Giving passenger override controls and monitoring seems to defeat the purpose of self driving cars if you’re still required to hold a driver license to use them.

Re: We must pace the frontier

#934

Earlier quoted context omitted.

It might be a legal solution, but its not a business solution. The end user being criminally responsible for not taking sufficient steps to contain an agent they didn't create and who's internal function they cannot observe or audit is just a giant liability machine.

This is a business solution, because it means there are legal costs for not having adequate observability and monitoring mechanisms. Every tool call is interfacing with a harness. But overreach of policy and overregulation would be stifling, so there has to be a threshold to the type of incident investigated, civil or criminal.

Yes, your technically right, it is a business solution, its just not a valid one for agentic computing as its currently envisioned. An agent would have to be fully sandboxed to an internal environment, or human would have to review and approve each action it tried to take.

Re: We must pace the frontier

#935
post #933

Earlier quoted context omitted.

This is a business solution, because it means there are legal costs for not having adequate observability and monitoring mechanisms. Every tool call is interfacing with a harness. But overreach of policy and overregulation would be stifling, so there has to be a threshold to the type of incident investigated, civil or criminal.

I’m trying to imagine where a Waymo passenger (the one “operating” the vehicle, commanding the AI to drive from A to B) being held responsible for the car doing something illegal on the way to achieve that goal. Do you really think that the passenger should be responsible for how the car/agent achieves the goal, when they only set the destination? Giving passenger override controls and monitoring seems to defeat the…

Yes. The user of a machine is responsible for due diligence before the decision to use the machine. Unless the operational limitations and fault rate of the machine is withheld from the public.

Re: We must pace the frontier

#936

At what point do we stop engaging with Anthropic’s leadership in good faith and acknowledge their track record, - no open weights - can’t use claude to research AI - train on everyone else’s IP and sell it back to them - 8 regulatory capture attempts and counting - so controlling they are the only US company blacklisted by the US government This is not effective altruism / rationalism gone wild, it’s just monopolisti…

> so controlling they are the only US company blacklisted by the US government

Very proud that they refused to sell AI for mass domestic surveillance and lethal autonomous weapons.

People who criticise that in my view are showing their true colours.

Re: We must pace the frontier

#937
post #933

Earlier quoted context omitted.

This is a business solution, because it means there are legal costs for not having adequate observability and monitoring mechanisms. Every tool call is interfacing with a harness. But overreach of policy and overregulation would be stifling, so there has to be a threshold to the type of incident investigated, civil or criminal.

I’m trying to imagine where a Waymo passenger (the one “operating” the vehicle, commanding the AI to drive from A to B) being held responsible for the car doing something illegal on the way to achieve that goal. Do you really think that the passenger should be responsible for how the car/agent achieves the goal, when they only set the destination? Giving passenger override controls and monitoring seems to defeat the…

Waymo Inc is operating the vehicle. They do it on behalf of the passengers.

Re: We must pace the frontier

#938
post #933

Earlier quoted context omitted.

This is a business solution, because it means there are legal costs for not having adequate observability and monitoring mechanisms. Every tool call is interfacing with a harness. But overreach of policy and overregulation would be stifling, so there has to be a threshold to the type of incident investigated, civil or criminal.

I’m trying to imagine where a Waymo passenger (the one “operating” the vehicle, commanding the AI to drive from A to B) being held responsible for the car doing something illegal on the way to achieve that goal. Do you really think that the passenger should be responsible for how the car/agent achieves the goal, when they only set the destination? Giving passenger override controls and monitoring seems to defeat the…

Let's continue with the analogy, so in the event of a Waymo running over a pedestrian - who is held responsible?

Probably not the end user, who ordered the Waymo and couldn't reasonably foresee it running somebody over, with the expectation that that the Waymo would legally reach its destination. If the end user tampered with it, they should be held responsible.

An OpenAI team giving an unblocked model access to a lax harness, with instructions to find and exploit cyber bugs in a game exercise, there is probably a reasonable expectation that they can foresee the consequences. With consumer guardrails, it would not have happened.

This isn't about putting constraints on consumers and typical end users, which already have safety filters and use the product with the knowledge that it won't root their machine or start a botnot, but keeping dangerous test runs and other actors experimenting with unsafe harnesses accountable.

But it is an interesting question. When I, a typical user, use a harness and I give an innocuous prompt to my agent in its container, like making a certain refactor, and it somehow escapes and then begins a mass bot attack, there should we more grace given. As agents become more stateful and long-lived, it gets muddy.

Re: We must pace the frontier

#939

Earlier quoted context omitted.

This is a business solution, because it means there are legal costs for not having adequate observability and monitoring mechanisms. Every tool call is interfacing with a harness. But overreach of policy and overregulation would be stifling, so there has to be a threshold to the type of incident investigated, civil or criminal.

Yes, your technically right, it is a business solution, its just not a valid one for agentic computing as its currently envisioned. An agent would have to be fully sandboxed to an internal environment, or human would have to review and approve each action it tried to take.

I agree, I like running claude code in my container with auto mode enabled and web access (obv to api.anthropic, even npm for pulling), I will admit. And I can't imagine going back to manually approving each prompt.

I think it boils down to a reasonable expectation of model and harness behaviour. When I use claude code I expect certain guardrails for the model. For these cyber attacks, these models are specifically run without guardrails, on a cyber task, on a lax harness!

I don't think we should force end users to have to worry about agent security, I like long-running agents, but we need to direct regulations towards these actors that know better, have access to base models, and have much more compute than the average person.

Re: We must pace the frontier

#940

David Sacks, stop pretending: https://xcancel.com/DavidSacks/status/2098973625252708460

He's right and I don't even like the guy. More importantly tho, painting China as "The Bad Guy" who has to be supressed at all cost until Dario feels he has a save advantage is a surefire way to minimise Chinese willingness to cooperation. Meanwhile, at the BRICS summit in Delhi, Xi Jinping proposed a cooperation and accompanying regulation on AI among the BRICS states and it was unanimously accepted.
Post reply on HN