Live data from Hacker News

The hacker sent by Anthropic to calm the government's nerves about AI safety

wsj.com

101–110 of 131 posts

Re: The hacker sent by Anthropic to calm the government's nerves about AI safety

#101

Earlier quoted context omitted.

My bad, I will do better. I know you guys are spread pretty thin managing the site, so I went through the comments for this post and collected some other comments which are also probably breaking site rules. https://news.ycombinator.com/item?id=48576022 https://news.ycombinator.com/item?id=48576065 https://news.ycombinator.com/item?id=48576162 https://news.ycombinator.com/item?id=48576183 https://news.ycombinator.com…

Genuinely hilarious reply. On that note it would be interesting to do a sentiment analysis of flagged replies. They seem all over the place and it would be interesting to see if there were any biases.

Believe it or not, between the two of us we don't read and deliberate on all of the 10,000+ comments submitted here each day. We obviously can only read a small fraction of them via routine monitoring of the site, and we're less likely to see comments in threads like this one that spend fewer than two hours on the front page. We're much more likely to see the comments that are put on our radar via community flags and emails, which is why I saw the one I replied to and not the others.

For what it's worth, sentiment analysis is unlikely to yield useful findings in this context, because the probability of a bad comment being flagged is highly correlated with the number of people who see it, which will be much lower in a thread that spends little time on the front page, but a sentiment analysis model won't have access to that data.

Re: The hacker sent by Anthropic to calm the government's nerves about AI safety

#102
post #96

Earlier quoted context omitted.

> Mythos was better than other models at creating exploits. Not a fan of this phrasing, prefer "discovering exploits". It makes it clearer the problem was already there, latent. Minor vocab diff, but important to better contextualize the present situation.

Exploits are created ("crafted" might be a better word), vulnerabilities are discovered. Unless you're hiding a RAT behind a public trigger in your code on purpose, I guess?

In general, the exploit has been (however systematically) stumbled upon, or felt through like a person navigating a physical maze.

Nobody would say that person "created" the solution to the maze.

The maze is solvable (that's the latent vulnerability), the person "discovered" the way through.

Re: The hacker sent by Anthropic to calm the government's nerves about AI safety

#103
post #101

Earlier quoted context omitted.

Genuinely hilarious reply. On that note it would be interesting to do a sentiment analysis of flagged replies. They seem all over the place and it would be interesting to see if there were any biases.

Believe it or not, between the two of us we don't read and deliberate on all of the 10,000+ comments submitted here each day. We obviously can only read a small fraction of them via routine monitoring of the site, and we're less likely to see comments in threads like this one that spend fewer than two hours on the front page. We're much more likely to see the comments that are put on our radar via community flags and…

So 1: I have no idea why you’re replying to me as if I’m attacking you. I’m not and I’m aware of how hard moderation can be as I’ve done it.

2: My idea for sentiment analysis was geared towards bias from this site and its users, not towards the moderation team.

3: While I respect the mod team I’m incredibly unimpressed with your response here, even if this was a misinterpretation of what I meant.

Take a breather.

Re: The hacker sent by Anthropic to calm the government's nerves about AI safety

#104
post #101

Earlier quoted context omitted.

Believe it or not, between the two of us we don't read and deliberate on all of the 10,000+ comments submitted here each day. We obviously can only read a small fraction of them via routine monitoring of the site, and we're less likely to see comments in threads like this one that spend fewer than two hours on the front page. We're much more likely to see the comments that are put on our radar via community flags and…

So 1: I have no idea why you’re replying to me as if I’m attacking you. I’m not and I’m aware of how hard moderation can be as I’ve done it. 2: My idea for sentiment analysis was geared towards bias from this site and its users, not towards the moderation team. 3: While I respect the mod team I’m incredibly unimpressed with your response here, even if this was a misinterpretation of what I meant. Take a breather.

[deleted]

Re: The hacker sent by Anthropic to calm the government's nerves about AI safety

#105
post #101

Earlier quoted context omitted.

Genuinely hilarious reply. On that note it would be interesting to do a sentiment analysis of flagged replies. They seem all over the place and it would be interesting to see if there were any biases.

Believe it or not, between the two of us we don't read and deliberate on all of the 10,000+ comments submitted here each day. We obviously can only read a small fraction of them via routine monitoring of the site, and we're less likely to see comments in threads like this one that spend fewer than two hours on the front page. We're much more likely to see the comments that are put on our radar via community flags and…

> Believe it or not, between the two of us we don't read and deliberate on all of the 10,000+ comments submitted here each day.

Aren't snarky comments against the rules of this forum? It would be great to have guidance on this because I'm naturally a sarcastic person and I try my best to not let it out on forums such as this.

Re: The hacker sent by Anthropic to calm the government's nerves about AI safety

#106

Unrelated to this story but Carlini rocks. tic tac toe in printf https://github.com/carlini/printf-tac-toe Recently Regex Chess: A 2-ply minimax chess engine in 84,688 regular expressions https://github.com/carlini/regex-chess https://news.ycombinator.com/item?id=48136909

I think headline readers might have thought I meant completely unrelated. It's not. Read the article...

Re: The hacker sent by Anthropic to calm the government's nerves about AI safety

#107

Earlier quoted context omitted.

> You can’t jump up and down screaming how amazing, powerful, and dangerous your new tech is and then act surprised and annoyed when the government shows up looking to regulate it. True, you can't. But, you can think certain regulations are helpful and certain other regulations are not. And you can be annoyed when unhelpful "regulations" are put in place. This is like if I say that pitbulls are dangerous, and then th…

Sorry, this argument doesn't work. Anthropic claims Mythos is in a class of its own, the evidence corroborates this and the government believes it. The government shot your pit bull because you were going around telling everyone who would listen that it was the most dangerous, viscous one on the cul de sac and you've trained it to kill people and they took you seriously.

> Anthropic claims Mythos is in a class of its own, the evidence corroborates this and the government believes it.

They didn't release Mythos, they released Fable, which was Mythos + a classifier that detected potentally-dangerous prompts and blocked them. Everyone who used it noticed how aggressive the classifier was. It would trigger constantly over totally innocent stuff.

Re: The hacker sent by Anthropic to calm the government's nerves about AI safety

#108

Earlier quoted context omitted.

Sorry, this argument doesn't work. Anthropic claims Mythos is in a class of its own, the evidence corroborates this and the government believes it. The government shot your pit bull because you were going around telling everyone who would listen that it was the most dangerous, viscous one on the cul de sac and you've trained it to kill people and they took you seriously.

> Anthropic claims Mythos is in a class of its own, the evidence corroborates this and the government believes it. They didn't release Mythos, they released Fable, which was Mythos + a classifier that detected potentally-dangerous prompts and blocked them. Everyone who used it noticed how aggressive the classifier was. It would trigger constantly over totally innocent stuff.

A classifier that was exposed as non-efficacious for a product touted as having extremely dangerous capabilities.

I can generate hacks trivially by asking any model to fix open source code.

Let’s not pretend you get to have your cake and eat it too.

Re: The hacker sent by Anthropic to calm the government's nerves about AI safety

#109
post #98

Earlier quoted context omitted.

> may be held liable for downstream uses over which they have no control Equivalent to a ban. Nobody is going to host or invest in this stuff if they suddenly become liable for everything it does. This is equivalent to repealing the safe harbor provisions in the DMCA.

>Committee amendments simplify and clarify the definition of “full shutdown” such that the shutdown capability can be implemented into hardware used to train or run a model, rather than the model itself. The amendments also serve to exclude covered model derivatives that are outside of the developer’s control.

I get the impression you are conflating whether a developer can be sued to oblivion for not implementing a "full shutdown" process that applies to finetunes versus whether they can be sued to oblivion for releasing a model that may cause "critical harm" when finetuned.

I'm confused why you think the only legal requirement is a "full shutdown" process. The text is there and I see a heck of a lot of requirements that are not about full shutdowns.

Re: The hacker sent by Anthropic to calm the government's nerves about AI safety

#110
post #9

The AI labs look rather naive here. You can’t jump up and down screaming how amazing, powerful, and dangerous your new tech is and then act surprised and annoyed when the government shows up looking to regulate it. Their new argument now seems be that this was marketing hype/fluff that backfired, in a pretty obvious and predicable way, and now they’re trying to reset the conversation.

Remember strawberry? I do. Rememver gpt2 hype? I do.

the hype isnt real, its marketing designed to inderectly siphon capital from the less informed.

The current USA government is no stranger to a grift, so they'll get it. Not that I agree with the practumice, this bubble will hurt quite badly when it goes, but at least AI fundamentally does deliver something useful, even if it isnt infinite value as typically promised.

Think dotcom bubble and the hype and promises made surrouning that. Tge hyper will pass, the bubble will pop, and life goes back to normal as ai becomes part of the mundane everyday human environment. Like websites and domains, some will be used well, some for evil, not everyone needs it, and as we continue to move towards energy being our fundamental unit of value / exchange, if we cant make these models way more efdicient then their use case will be rather limited in scope.

Post reply on HN