Live data from Hacker News

If Claude Fable stops helping you, you'll never know

jonready.com

231–240 of 534 posts

Re: If Claude Fable stops helping you, you'll never know

#231
post #188

Earlier quoted context omitted.

> "YOLO" is not a reasonable answer here. Yes it is. (1) Ordinary people were able to do these things pre AI-- with some effort into study for sure. (2) The cat is already out of the bag, open models can already help with these tasks. I know freedom is frightening, but it always has been. It's important to avoid falling into the trap of assuming that everything that existed when you gained awareness was safe and norm…

Kindly drop the condescension. It is, in fact, possible for the world to get more dangerous over time . It is important to avoid falling into the trap of assuming that's inevitable. > Ordinary people were able to do these things pre AI-- with some effort into study for sure. Yes, and the amount of study and knowledge required had a tendency to filter out people with the inclination to do such things. The Venn diagram…

> Yes, and the amount of study and knowledge required had a tendency to filter out people with the inclination to do such things. The Venn diagrams weren't completely empty, but they were close, which is why such incidents were rare.

People do exercise their freedom and do terrible things all the time - it's not rare. There are lots of ways to cause harm that don't require any study or knowledge at all, we just seem hyper-focused on the possible "sci-fi" consequences of AI for some reason.

I would argue the reason people don't go and kill someone (or worse...) even more often than they do is not because it's difficult but because most people have no desire to cause that kind of harm, and because of the consequences to themselves of doing so.

So yes: technical difficulty put some kinds of harm out of reach of people, and AI can lower that barrier somewhat, but in the grand scale of "harm people can do" I think it's receiving undue attention.

And from a practical standpoint: how do you get from there to arguing that we should set some impossible-to-define threshold of "frontier" at which point it becomes so evil that we need to forcefully delete it from existence? Don't you see the problem with trying to put such black and white restrictions on something that's so inherently amorphous and slippery? (And by definition, if you delete the "frontier" model from existence then the next best model is now "frontier" ad infinitum...)

On top of that you have the issue that model weights are just information, so in some sense you're legislating the knowledge that is allowed to exist. That's quite a bit more draconian that current laws which usually focus on what knowledge you can share.

Re: If Claude Fable stops helping you, you'll never know

#232
post #108

> If Claude gives me poor or incorrect advice while I’m working on an AI component, I have no way of knowing whether the model was confused, whether my problem is unsolvable, or if some invisible policy restriction quietly kicked in. You should be able to know if your problem was solvable by using your own expertise and judgement, no? If you're relying on LLMs as a substitute for those, I wouldn't expect great result…

No; once the LLM switches to this new saboteur mode, it’ll be very hard to detect. Sabotage is an asymmetric weapon. The ratio of damage to effort is nearly unbounded, and any decent saboteur knows that the key trick is to make your output indistinguishable from incompetence. They’re building state of the art offensive capabilities into a public model, then expecting to maintain control over when it decides to attack…

Great way for Anthropic to build trust with the military

Re: If Claude Fable stops helping you, you'll never know

#233

Earlier quoted context omitted.

Presumably by making it "difficult enough" to misuse the tools. We don't need perfect censorship or surveillance. There are all sorts of things that are technically possible today but typically aren't an issue in practice due to some oftey fairly minor hurdles. Aum literally synthesized sarin in the 90s so clearly it's doable yet in practice it doesn't seem to be a problem that crops up regularly. Anyone with a bache…

And how exactly do you propose making it "difficult enough"?

The same way Anthropic is making it difficult to compete with them. They intentionally train the model (via PEFT, as called out in the model card) to be dumber when attempting to do things Anthropic doesn't want — in this case, competing with them, but you could apply the same training process for other domains such as actually-malicious use cases.

Re: If Claude Fable stops helping you, you'll never know

#234
post #61

"To effectively contain a civilization’s development and disarm it across such a long span of time, there is only one way: kill its science." - Cixin Liu, The Three-Body Problem This immediately made me think of the Sophons silently manipulating the sensors of particle accelerators to prevent humanity from developing advanced knowledge of particle physics.

The level of oppression necessary to get software geeks to stop making progress on AI is similar to that necessary to get Ukrainian geeks to stop making progress on drones.

unless you could convince them that making smarter-than-human AI is bad. it would be nice if we all thought this. instead they should figure out how to make dumb models faster or more efficient, that's safe

Re: If Claude Fable stops helping you, you'll never know

#235
post #17

Earlier quoted context omitted.

I think there's a pretty big difference here. It's not like Github prevents you from building a Github competitor. Or Linear is preventing you from using it to build a Linear competitor. This is more akin to Windows somehow preventing you from building a new OS. Or worse yet, sabotaging vs preventing.

> This is more akin to Windows somehow preventing you from building a new OS. Tangent, but have you tried repartitioning your Windows disk to make room for a new OS? Or tried to configure Windows to let you dualboot? Or get the clock time right if you dualboot? Or let you debug "Secure Boot"? Windows is outright hostile when it comes to (sharing with) a new OS

Yeah, MS doesn't quite exemplify good-faith competitive spirit, does it?

Re: If Claude Fable stops helping you, you'll never know

#236
"We collect everyone's data without paying a dime or respecting copyright, trained our models, but you can't train your models on our models that are trained on everyone's data collected without paying a dime or respecting copyright. We did a hard job stealing that all data and processing it, have some shame!"

Re: If Claude Fable stops helping you, you'll never know

#237

There is a possibility this may not end at simply nerfing the model. The idea of manipulating the behavior of a model depending on the prompt given to it can extend to 1. Detecting if employees from competing companies are using it and sabatoge their work, even not LLM-training related 2. Direct users to outcomes that would justify higher compute spend. Deliberately coding a project to 95% completion but designed to…

Anthropic: were commiting to being ad free.

Also Anthropic: if you use our models in any way that might negatively impact our revenue we'll sabotage you.

Can I pick the ads please?

Re: If Claude Fable stops helping you, you'll never know

#239

The moat looks deep today but it's going to become more shallow every year. Training a new model from scratch takes serious resources. Post-training/fine-tuning an existing model, dramatically less. The knowledge for the process was esoteric two years ago, now you can ask a current model (one of several) to walk you through it, while building the tools to do it as you go. Several of my recent weekend projects have be…

> The moat looks deep today

Does it? What can this model do that I both want and cannot already do?

Anthropic made a nice little post saying how dangerous it is, because it is good enough to eat their own business. But I don't want to eat their business. They also said it was good at playing Slay the Spire, but I can't think of anything more insulting than have a machine do that in my place. That's MY comfort game, not something for a stupid Clanker to take away.

They did not provide any other use case.

Re: If Claude Fable stops helping you, you'll never know

#240

has dario (or sam tbh) ever been thoroughly asked about the hypocrisy of them claiming distillation to be „theft“ vs. them training on the copyright of others? I’ve only seen him talk about one of those topics, but never together. I just can’t see how you can talk yourself out of that hypocrisy, if BS answers are properly followed up on (journalism!)

Distilling the entirety of thousands of years of human intellectual output: totally cool.

Distilling the answers of one LLM: totally uncool.

Post reply on HN