Live data from Hacker News

I resigned from Anthropic today

twitter.com

731–740 of 990 posts

Re: I resigned from Anthropic today

#731

No other human activity poses this level of danger. I do heed the warnings, but this comes across as detached hyperbole. See: global warming, nuclear weapon development, wealth inequality, war, technology dependence, etc. Also, this has nothing to do with LLMs or computers. Like all things, this is about humans.

The danger comes from what's possible.

Open weights agents with hacking capacities can reproduce themselves into the systems they hack (non-open weights ones will have to hack their creators first). Not saying they will, but if they do, good luck finding the kill switch.

Once swarms of agents run unsupervised on unmonitored hacked hardware, who can tell what they will do? The Huggingface incident showed that such swarms behave without any safeguard. It was a real HAL moment.

A lot of things are possible then: ransomware campaign, taking over IoT devices, self driving cars, planes, ships, satellites, missile launchers. If nothing's out of reach, everything is possible.

Bring robots into the mix, and the possibilities are endless.

I'm not particularly frightened tbh, but we shouldn't discard the worst-case scenario, and the worst-case scenario doesn't look good.

Re: I resigned from Anthropic today

#732
post #709
post #706

It's as if two private companies are each building increasingly large nuclear bombs, both saying they'd love to stop but it would be unsafe to let any one company be in control of the nukes.

That's basically the reasoning behind MAD, and it checks out? See what happened to any nation that ever gave away their nukes.

Does it check out? India and Pakistan both have nukes and keep semi-regularly fighting each other.

MAD "on paper" prevents either side from going far enough to provoke the other into using nukes, but even then it's fundamentally flawed because it works on the assumption that both sides are both rational and believes the other side to be rational, as well as that both sides understands the others red lines well enough.

Already Reagan realised that isn't necessarily true - after Able Archer '83, he realised that the Soviet leadership seemed to genuinely believe that the US might be prepared to carry out a first strike, and that Able Archer got dangerously close to convince them one might be imminent. It's one of the things he noted as a reason to get in the room with them and negotiate.

If you believe the other side is irrational (whether or not that is because you are irrational), and think they're about to strike, MAD turns from a deterrence into a reason to try to preempt to ensure you're the "least destroyed" by hitting harder, sooner.

Re: I resigned from Anthropic today

#733

Earlier quoted context omitted.

Let me argue on a technicality first: None of these are extinction level events. If global warming disrupts 99% of all crop production, the remaining 1% is still plenty enough to sustain a stable, if miserable, population. In fact you just need about 5k people for a stable gene pool[1]. Of the classical threats, only bioweapons got a shot at extinction, but even that is hard, given the (few) remaining truly secluded…

I still have to read a compelling argument on how AI will "extinct" humanity.

Imagine one of the recent frontier models with a flipped sign (cf §4.4 of https://arxiv.org/pdf/1909.08593)

Re: I resigned from Anthropic today

#734
It's quite incredible to consider that all these concerns already existed years ago,

but now that a handful of companies working in AI managed to enslave the entire financial system over the past year, their continued work is protected from larger governance for concerns it could tank the stock-market, affect personal investments, pensions or cause disadvantages in an arms-race with other countries.

IF there is an inherent danger (which I believe is the case at least on economic levels, work displacement, poverty,...), it is now basically ensured that nothing will be done to reign those companies in, until maybe two AI's engage in an open war with civilian casualties...

Re: I resigned from Anthropic today

#735
OpenAI has been saying since the first version of ChatGPT that it's too dangerous to release because it will end humanity. Yes, LLMs are an impressive technology, but let's be real: the improvements in the recent months have been slowing down, and it's clear that we are nearing a plateau of what this particular tech can do. Sure, tooling and harnesses etc. is improving, but clearly this dude has drank too much of the Kool-Aid.

Re: I resigned from Anthropic today

#736

I think people here still evaluating the model in isolation. It is the combination that matters, model + strong harness + tools + long running autonomy + memory + retries + parallel agents + code execution + credentials + access to real systems. The model does not need to be perfect. If it fails 30% of the time, the harness can retry, verify, branch, use another agent and keep going. I don't think we necessarily need…

> an extremely capable harness and enough access. Give enough access to a fuzzer and it's exactly as dangerous as an LLM. LLMs don't even have a moat in this domain.

Technically. What's would technically be even more dangerous is running this shell script:

  head -c 200000 /dev/urandom > agi && chmod +x agi && ./agi
In reality, the fuzzer definitely has no agenda, and these random bytes probably don't. The LLM definitely does, and even publicly available models, programmed ot "do what the user wants, act according to the anthropic moral codex" will take some pretty absurd actions in attempting to accomplish a poorly worded request.

Re: I resigned from Anthropic today

#737

I could imagine a 2027 AI swarm coordinating to eg hold the US and Russian and Chinese governments to ransom, by demonstrating some small thing (turning US army base freezers to defrost) and threatening to do something big unless some conditions were met – conditions which would be good or bad for the world depending on your POV. This happens either either because they were tasked to to it by (malicious or well-meani…

And then we pull the plug, after holding our breath for 10 seconds,. Then life resumes normally..

Not normally if essential services rely on the same plug.

Re: I resigned from Anthropic today

#738

No other human activity poses this level of danger. I do heed the warnings, but this comes across as detached hyperbole. See: global warming, nuclear weapon development, wealth inequality, war, technology dependence, etc. Also, this has nothing to do with LLMs or computers. Like all things, this is about humans.

Nuclear weapons, climate change.... The AI doom hyperbole is unhinged.

[flagged]

Re: I resigned from Anthropic today

#739
post #648

Earlier quoted context omitted.

yes, exactly, that's what they claimed that incoherent gibberish generator to be capable of.

So you aren't claiming that they said GPT-2 was dangerous in the sense that it could disempower humanity, kill all humans, etc. You are just claiming that OpenAI execs said that GPT-2 might "generate misleading news articles, impersonate others online, automate the production of abusive or faked content to post on social media, automate the production of spam/phishing content." Then, what is unreasonable or bad about…

that it was bullshit and they knew it. GPT-2 wasn't capable of anything other than imitating a stroke victim.

Re: I resigned from Anthropic today

#740
post #493

Earlier quoted context omitted.

Example 8: like in this comment https://news.ycombinator.com/item?id=49619884 but isolate synchronised megahack on banking that adds one more zero to the US debt and all dependent systems and banking during runtime. Let the world's financial system take it from there.

I think US fiscal policy and adventurism will manage this on its own. :)

But why wait? We can has global financial crash now:)
Post reply on HN