Live data from Hacker News

Claude Fable is relentlessly proactive

simonwillison.net

731–740 of 748 posts

Re: Claude Fable is relentlessly proactive

#731

Earlier quoted context omitted.

But can you bring anything measurable in support to your words? I did.

You brought your own benchmark to support your words. I happen to have studied statistics, so I took a look. It is deeply flawed, primarily because it is not a statistical benchmark. It is a single (n=1) autonomous "pi" coding-harness run per model per prompt, scored by an automated battery (A-items, pass/fail), an LLM code review (R-items, 0 to 2 each), and a human manual checklist (M1 to M10) that was never actuall…

Yes, I know all the flaws. As I said, it's not an objective way to measure performance of a model - but it is intended to produce something that only humans could mesaure. The goal is for you to being able to play the game and judge - and fill the human checklist for yourself if you wish.

You didn't get why the automatic review scores are there - all of the reviewers, including Fable, happily assign highest scores to code which can't even run. In my opinion that is a sort of an empirical evidence that these models are very far from the "AGI" state.

Anyway, while I didn't explain the methodology and the purpose of this experiment, I have something material to discuss. The "awesome Fable" claims are not material at all.

Can you bring something clearly showcasing Fable's superiority?

Re: Claude Fable is relentlessly proactive

#733
post #721

Earlier quoted context omitted.

Thanks for documenting your personal observations. I do have a few questions. First, could you expand by giving other examples on how you observed this model to be relentlessly proactive? From my personal experience with prior frontier models using both Claude Code and Codex I found them to already be quite proactive depending on the domain (although Codex a bit less so, which I personally prefer). The main task that…

The visual regression point is interesting. In my experience, the models that do best at "overlapping text/bad layout" catches are the ones being fed actual screenshots rather than DOM snapshots. If Fable is doing screenshot-based diffs natively, that would explain an improvement there, but I haven't verified it.

From how Simon described it it's not a native feature, but one that the model built as a solution for automatically testing. You could already instruct the agent to write a program that saves screenshots to disk and then reads it. As long as the model is multimodal (which pretty much all releases are these days) it can "natively" interpret images. There's probably a clever way to engineer this to be somewhat efficient, but for me it was rather token hungry, because the testing inputs and the description are usually quite verbose. I suppose you could use a weaker model for navigating the test and then only feed the output to the stronger model.

Re: Claude Fable is relentlessly proactive

#734
post #315

Earlier quoted context omitted.

I don't think the pressure of the auto lobby is really the reason. People feel cars are more convenient and more prestigious than riding on a bus. Car lobby certainly accelerated the process, but car users were the main driving force.

Surely people feeling that way can be attributed to the industry?

Cars were quite desirable in Soviet Union, where industry was not allowed to advertise. You had to get into a queue to buy a car, the state was not interested to make them in a quantity to satisfy the demand.

Very few people actually _needed_ cars as soviets built adequate public transport system. But there are many situations where car can really help a lot. Perhaps that's more obvious in a society which has rather few cars.

E.g. back in Soviet days and around that only one member of my extended family had a car. The rest of the family were really happy about opportunities it provides. E.g. with a car you can buy fresh produce directly from farmers with just few hours of driving. Doing the same without a car is so much hassle and effort people just won't do it, and then you're confined to what's available in a local grocery story (which was usually much worse than direct-from-farmer option). Do you think it has something to do with "car industry"?

Re: Claude Fable is relentlessly proactive

#735

Obviously security is the bigger issue, but reading through this, all I could think about was how many tokens it must have spent doing all that to fix 2 lines of CSS

"Your scientists were so preoccupied with whether or not they could, they didn't stop to think if they should." I'm convinced this is going to be the summary of the 2020 decade...

To be fair, they did stop to think if they should. The decided that they shouldn't and went ahead and did it anyways.

Re: Claude Fable is relentlessly proactive

#736

How can a LLM be assigned an emotion as being "proactive". This is highly misleading to anyone that scans just the headlines. What actually happened is that the user started a prompt, and Claude took $12 worth of tokens to resolve the issue. How it did so was basically looping until it got to the answer How is this proactive? It's literally being token greedy and maximising revenue for the LLM owner. People really ne…

Compared to other models that halt the loop on intermediate steps, or to ask further clarification, even if it's not the human equivalent of proactive, you see the similarity, right?

run haiku or sonnet under pi.dev. the halt comes from the harness/observer. even better, gemma 26b will just go forever.

Re: Claude Fable is relentlessly proactive

#737
Everyone here is reaching for infra (VMs, throwaway users) because the permission model only has two settings: Ask-every-time or --dangerously-skip. That seems to me like a design gap, and scoped capabilities and budget caps are missing. Same way you'd onboard a junior eng.

Re: Claude Fable is relentlessly proactive

#738
post #315

Earlier quoted context omitted.

Surely people feeling that way can be attributed to the industry?

Cars were quite desirable in Soviet Union, where industry was not allowed to advertise. You had to get into a queue to buy a car, the state was not interested to make them in a quantity to satisfy the demand. Very few people actually _needed_ cars as soviets built adequate public transport system. But there are many situations where car can really help a lot. Perhaps that's more obvious in a society which has rather…

Nobody is complaining about cars existing but about mandatory-car cities and the mindset associated with that.

Re: Claude Fable is relentlessly proactive

#739
post #315

Earlier quoted context omitted.

Surely people feeling that way can be attributed to the industry?

No its much more straightforward, but I get it - there is no warm fuzzy feeling of discovering yet another global evil conspiracy out there set to get all of us. We are family of 4 with 2 small kids. Whenever we travel, its a series of backpacks, other bags, other stuff, and then some more. Heck, even if I travel alone its almost never just me - there are heaps of garbage to dispose, big shopping bags to bring back,…

Yeah, when you travel. But wouldn't it be cool if school was just down the block and a grocery store was the same distance the other way? In big city Europe, it's like this.

Re: Claude Fable is relentlessly proactive

#740

Earlier quoted context omitted.

The frequency and magnitude of the event is directly related to the warming up of climate

By related to I assume you mean correlates with though. To be fair, we can't say there is a causal link (even if it does seem very likely).

We can TOTALLY say that there's a causal link, it's been proven by scientific stories
Post reply on HN