Live data from Hacker News

If Claude Fable stops helping you, you'll never know

jonready.com

101–110 of 534 posts

Re: If Claude Fable stops helping you, you'll never know

#101

> If Claude gives me poor or incorrect advice while I’m working on an AI component, I have no way of knowing whether the model was confused, whether my problem is unsolvable, or if some invisible policy restriction quietly kicked in. You should be able to know if your problem was solvable by using your own expertise and judgement, no? If you're relying on LLMs as a substitute for those, I wouldn't expect great result…

You come up with a hypothesis -> you let fable implement it -> fable sabotages your experiment -> you get evidence that hypothesis is not true.

It's that simple.

Re: If Claude Fable stops helping you, you'll never know

#102
I am so happy that Anthropic has signaled the possibility that their UI moat for agentic AI is copyable by competitors. At least that's the way I read this. When companies try to lock something down it can be a signal of weakness.

If so, it's possible to built great user interfaces in Chatbots and more companies/people can have amazing agentic development workflows! We don't have to live in a world where only the market leader has the most enjoyable model.

Re: If Claude Fable stops helping you, you'll never know

#104
post #80

It is absolutely fine to distill the IP of everyone else, but you'd be violating the TOS to distill ours :)

Yep. Demand open source approve licenses for LLM weights. The Chinese apache 2.0 models might be censored, but at least they can’t sue you in the US for finding the censorship line. OTOH, the US models are definitely censored, per TFA, and they’re making vague legal threats against anyone that encounters the censored edge of the model.

the base models released to the public are not censored. censorship happens with another model, that isn't released

Re: If Claude Fable stops helping you, you'll never know

#105

This is a fun peek into the economic implications of RSI/ASI. Because it's so infinitely valuable that it basically destroys all markets, labs will eventually do stuff like stop releasing models completely and skipping out on contracted commitments because they'll have the power to just drive their competitors out of business before the legal battle gets expensive. Cloud providers - at first smaller ones, then the hy…

I don't think this scenario makes sense. It's one of a class of scenarios I've seen several of, that simultaneously assume:

  A) ASI is developed and massively overshadows the rest of the world economy 
  B) the world still has rule of law, contracts, business, well-developed finance, etc
You can get to a lot of weird conclusions if you assume both A and B, but I think the much more likely scenario is that if A happens, B stops being true in short order. If you are a company and you have ASI, you just stop caring about business and money and economics, and your outcomes instead start looking like "you conquer the world" or "you upload the board of directors to a fleet of von Neumann probes" or "you messed up, everyone dies".

Re: If Claude Fable stops helping you, you'll never know

#106
PRODUCT VIOLATION

https://www.youtube.com/watch?v=Tr3t1uZNbKo

DIRECTIVE 4: [Classified]

Any attempt to arrest a senior officer of OCP results in shutdown.

Putting aside my snark, is Anthropic actually anticipating some new expansion of ITAR? (Or a stipulation for the Trump administration taking/not taking a share?)

That is to say, do they expect to be told that they must have this mechanism, not just the terms?

Re: If Claude Fable stops helping you, you'll never know

#107
post #61

"To effectively contain a civilization’s development and disarm it across such a long span of time, there is only one way: kill its science." - Cixin Liu, The Three-Body Problem This immediately made me think of the Sophons silently manipulating the sensors of particle accelerators to prevent humanity from developing advanced knowledge of particle physics.

The level of oppression necessary to get software geeks to stop making progress on AI is similar to that necessary to get Ukrainian geeks to stop making progress on drones.

Not so sure those things are equivalent

Re: If Claude Fable stops helping you, you'll never know

#108

> If Claude gives me poor or incorrect advice while I’m working on an AI component, I have no way of knowing whether the model was confused, whether my problem is unsolvable, or if some invisible policy restriction quietly kicked in. You should be able to know if your problem was solvable by using your own expertise and judgement, no? If you're relying on LLMs as a substitute for those, I wouldn't expect great result…

No; once the LLM switches to this new saboteur mode, it’ll be very hard to detect.

Sabotage is an asymmetric weapon. The ratio of damage to effort is nearly unbounded, and any decent saboteur knows that the key trick is to make your output indistinguishable from incompetence.

They’re building state of the art offensive capabilities into a public model, then expecting to maintain control over when it decides to attack its human users.

The premise is laughable, and we’ve all seen how this movie ends.

Re: If Claude Fable stops helping you, you'll never know

#109
Wow, this is like saying:

> If you buy a car from us, you agree not use it driving to and from work that involves automotive R&D that might compete with our product. And if our (heavily spying) car detects you are violating this, it will slow down to 20mph and cannot be made to go any faster, until we are sure the violation has ceased.

Or

> If you buy a laptop from us, you agree not to use it to study or acquire any knowledge that you may use to compete against us. If the laptop detects such a use, it degrades to one core and 4GB of memory, until the violation stops.

Re: If Claude Fable stops helping you, you'll never know

#110

> If Claude gives me poor or incorrect advice while I’m working on an AI component, I have no way of knowing whether the model was confused, whether my problem is unsolvable, or if some invisible policy restriction quietly kicked in. You should be able to know if your problem was solvable by using your own expertise and judgement, no? If you're relying on LLMs as a substitute for those, I wouldn't expect great result…

You come up with a hypothesis -> you let fable implement it -> fable sabotages your experiment -> you get evidence that hypothesis is not true. It's that simple.

Or, worse:

- It says your safety hypothesis is true, you incorrectly ship, killing lots of people.

- It proposes dangerous experiments.

Post reply on HN