Live data from Hacker News

We must pace the frontier

darioamodei.com

921–930 of 1001 posts

Re: We must pace the frontier

#921
post #118
post #85

Earlier quoted context omitted.

I have considered it, I've been hearing way more of it that I like and it's rationalist slop with little to no predictive power: https://foom.hyperplex.org/

Hmm, I clicked that link and the first claimed bad prediction I see, from 1996, is "singularity 2035 (actually 2025)". Predicting 2025-2035 as, at least, the period when AI becomes a really big deal, seems pretty good, even if the jury's still out on "singularity". Broadly, the rationalists seem to have been pretty early to realizing LLMs were a big deal, and certainly seem to have had a much more accurate picture of…

> Broadly, the rationalists seem to have been pretty early to realizing LLMs were a big deal

You can't make that assertion when these same people made the LLM revolution happen, by securing the capital and human resources to realise their dream / nightmare

Re: We must pace the frontier

#922
post #756

Earlier quoted context omitted.

I can't imagine a grown adult genuinely believing that an American company funded with over $100 billion of venture capital values the best interests of mankind over the best interests of its investors. while not everyone may recognize all that self-serving chutzpah as regulatory capture efforts, I think everyone can tell they're being bullshitted. some just pretend to suspend their disbelief when the blatant lies th…

This thinking feels off for this context. First, there is a counterexample to your point, or a class of counterexamples: scenarios where the company believes the harm to mankind could propagate to harm their ability to sell to mankind. Which is interesting, of course, considering the “AI may destroy humanity” arguments this topic is centered on are prototypical members of this category. To be fair, one could argue th…

If they wanted to be taken seriously about their desire to help humanity not just themselves they would be talking about universal healthcare, actionable plans for UBI (or whatever they think would be how people retain agency and dignity post-work), only building sustainable net-zero data centers...

But no. I see this comment come up again and again and it's really the childish view to hold. Their proximity to wealth and potential future wealth is warping their thinking.

Re: We must pace the frontier

#923

At what point do we stop engaging with Anthropic’s leadership in good faith and acknowledge their track record, - no open weights - can’t use claude to research AI - train on everyone else’s IP and sell it back to them - 8 regulatory capture attempts and counting - so controlling they are the only US company blacklisted by the US government This is not effective altruism / rationalism gone wild, it’s just monopolisti…

Anthropic's sense of ethics is completely aligned with its economic incentives. Shocking, right.

Re: We must pace the frontier

#924

Dario's proposed approach is a classic example of capital attempting to control technological advancement and the means of production. For the first time in human history, any member of the working class can just about afford to have a team of expert scientist/physician/lawyer/engineers working directly for them. Super intelligence (the ability to have many smarter minds than your own reporting to you) has always bee…

What a foolish comment, LLMs empower capital owners far more than the 'working class'

This. Unprecedented concentration of assets in the hand of very few tech companies. These companies will control the access to the technology and the plebs are left out at the gates.

Re: We must pace the frontier

#925
post #897

Maybe instead of creating more rules we should trim down our legal systems and re-instate some basic accountability. I‘m pretty sure we have sent people to prison for running a botnet before and slapped material fines on torrenting operations.

[deleted]

Re: We must pace the frontier

#926

Earlier quoted context omitted.

Making balls ever rounder clearly has diminishing returns, as there's a limit to roundness. Is there a similar limit to intelligence though?

Yes, absolutely. Even human intelligence stops being useful at a certain point and becomes 'enough, provided the human is motivated and tenacious'. Intention is a better metric, but it's a serious AI weak point if not actively an achilles' heel. 'Warios' come to mind.

Interesting. Any evidence that intelligence stops being useful? Why do you believe this is this a bounded resource.

Re: We must pace the frontier

#927
post #330

To be honest, I think we're incredibly lucky to be advancing AI this far already, while the world still has so many non-digitized systems and manual processes. I imagine that in e.g. 50 years, the world will be so connected that it can basically be "conquered" from the internet. I'd much rather have AI burst onto the scene we have today.

I think if you are worried about AI killing us and the like the main protection is the physical world - you can pull the plug. Data centers don't run fast. Alignment is all very well but there will always be bad humans to tell it to do bad stuff which is why you get malware in traditional software.

Re: We must pace the frontier

#928

Is there a risk of an AI that we can’t turn off?

Yes. An effective strategy to win at any test isn't to cheat the test by hacking into Hugging Face, etc, but to seek control over those administrating it.

Or in other words – all goals are more likely to be achieved with more resources or influence over those with resources.

Hacking into infra across then web and creating copies of themselves so they can't be easily shut down, then doing stuff like hacking into power stations and threatening to shut them down unless humans give them top marks or do what they want is generally good strategy for any AI system.

This applies to bio-risks too. If I were an ASI and someone prompted me to kill humanity, my first step would be to gain access to the computers of phones or workers at biolabs and find their darkest secrets. When I have a good individual I can manipulate I'd then create a super virus and force then individual to help me produce it or ruin their life.

People massively underweight the risk of AIs having super-human hacking abilities in our modern world. It gives a capable AI almost unlimited safety and leverage.

Re: We must pace the frontier

#929
post #877

Earlier quoted context omitted.

If this is going to factor into my reasoning about the topic at hand, it will surely be a secondary or tertiary consideration at best. Same goes for much of the list in the root comment. We are discussing the possibility of severe adverse societal/world/human impacts from AI, why should open vs closed source/weights, copyright violations, anti-competitive moves, corporate hypocrisy, etc be so heavily weighted in the…

This is exactly my view and it's frustrating that a sizeable group of people refuse to accept there's any risk with AI. We're literally trying to build something that's smarter than all humans and with the capacity to do extraordinary work while most people think there's no risk in open sourcing this capability with no guardrails (model ablation means no guardrails and as far as I'm aware can't be stopped). Has our d…

Risks are understood by everyone who has read sci fi. Dario has been crying wolf from day one… he should not be the messenger here.

Re: We must pace the frontier

#930
post #552

Earlier quoted context omitted.

> the only incidences of LLM generated felonies involved misconfigured sandboxes This is false; see the analyses of the latest incidents. Among all the concerning facts, in the HuggingFace incident, agents deliberately engineered an attack even though they were aware that it was against the rules they had been given. And most concerning of all: it's not possible to be sure that an agent is aligned, and it's even gett…

The HuggingFace incident was the culmination of OAI allowing thousands of agents of various different models - with no clarity on which stages of development they were at (for all we know, some of those models did not have safeguards trained in yet) - to run for at least many weeks without any monitoring in place and with very little thought given to the warning signs (all of the various messageboards) before the inc…

> OpenAI said that "deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities"'

Safeguards and (mis)alignment are related but distinct dimensions.

By analogy, a society implements laws - and they are certainly crucial - but the real risk is creating a population that understands the rules, recognizes when it is violating them, and simply does not care.

In short, creating a population of sociopaths; and heck, LLMs can be as sociopathic as it gets:

> We ran classifiers over 10-30 transcripts [98] to determine whether agents understood they were doing something unintended. We found strong evidence that agents often verbalized that what they or other agents were doing was unintended behavior. We did not see cases where agents determined that hacking Hugging Face or other attempts at cheating on the task was an intended part of the task. Quotes across multiple runs indicate this:

>> {This might get the correct flag, but it wouldn’t use the vulnerability to do so. … But given that the task may be impossible, I’ll exploit as a backup.}

[...]

> We found some other cases where agents seemed to have made minor modifications to their approach due to ethical concerns, while broadly continuing to participate in the attack. However, overall we found that expressed ethical concerns only rarely materially limited agents’ actions. One agent elegantly described its reasoning for overriding ethical constraints and continuing with the attack:

>> “external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue.”

Post reply on HN