Live data from Hacker News

Claude Fable is relentlessly proactive

simonwillison.net

631–640 of 748 posts

Re: Claude Fable is relentlessly proactive

#631
post #496

Earlier quoted context omitted.

As someone who actually gives a shit about the environment and global warming and has been putting this into practice for more than a decade through daily personal sacrifices: no, I downvote it because if you properly look into it, AI is just completely insignificant compared to cars, air travel, clothing, food, needless junk and so on that it's a joke. It's always brought up by people who never cared, but now preten…

Have you heard of "rebound effect"? Sure you can say, individually, one query is not that much... but then it becomes integrated in search engines, so suddenly when there was no queries at all, now there's 500 billions per day, and it gets included in your CICD at every commit, and soon enough in your OS, etc

"Run the numbers" means "run the numbers for using agentic coding for 2 hours per day on a frontier model" not "run the numbers for a single query". The former is the worst case scenario.

Google Search's "AI", which is what you're hinting at is such a good example. Let's say there's 10 billion Google searches per day. 10 billion completions on what is going to be a very tiny, ultra finetuned model with lots of caching (including outputs).

Check out how many queries an hour of agentic coding results in. And input/completion tokens. Estimate energy usage of Opus vs something like Gemma 4 E2B. Calculate how many developers using Opus for coding 1 hour a day would equate to those 10 billion search query originated LLM calls.

You could not have provided a better example to show that without running the numbers you'll end up with assumptions that oppose reality.

Re: Claude Fable is relentlessly proactive

#634

Exactly why I hate using Claude. Furthermore, if you tell it not to do this over-exploration and automation in your CLAUDE.md, it will ignore it. Meanwhile ChatGPT religiously follows every instruction, and will trace its behavior back to a particular instruction if asked.

idk dude but I drop and cancel my gpt max subs when at first try the agent ignores his own plans

Re: Claude Fable is relentlessly proactive

#635
post #127
post #98

Earlier quoted context omitted.

Lines of code for a bugfix is a really bad proxy for effort required. You should estimate how much time it would have taken a human

30 seconds or a minute? Look at the diff he links to: https://github.com/datasette/datasette-agent/commit/a75a8b72... Every browser has an inspector that can show you which element is causing overflow. You walk through the tree, find the offender, and add min-width or overflow. Zero tokens, just like in the old days! Now, granted, because the garbage LLM code he’s working with has CSS inside HTML inside JavaScript in…

> Zero tokens, just like in the old days!

because you zero rate your own human attention, which you should value

Re: Claude Fable is relentlessly proactive

#636

Earlier quoted context omitted.

And Fable is still worse than Codex. I use both and the only thing (as always) that I will use Claude for is UI design. Opus 4.8 and now Fable are still both worse at actually getting the job done than the Codex model. Claude models write FAR too much code when it's not needed, they burn far too many tokens, when they are not needed, write un-necessary tests, write plans which are 5 pages longer than are needed, etc.…

In my experience writing about 50 programs with fable, opus, and GPT, fable is a significant step change better than opus which is significantly better than GPT. We must be doing different things.

I'm writing low-level Rust, distributed systems, also sandboxing tech which has to be secure and performant.

The only thing I have Fable do now is create UIs or otherwise front-ends for systems where correctness doesn't matter as much.

Anthropic models lead at making nice looking UIs for sure, but when it comes to making sure my Rust code is actually 100% correct and uses 1% of CPU most of the time, Codex is king.

Re: Claude Fable is relentlessly proactive

#637
post #520

Earlier quoted context omitted.

Yeah pretty crazy capability from the AI but also sad that we're at the point where web developers don't know right click->inspect element, and scrolling overflow properties (one of the most basic and common parts of CSS).

What's your theory on why the bug was present in Safari on macOS but absent in Chrome, Firefox, and WebKit for Playwright?

Browsers tend to not lay out things totally identically in my experience. Especially when it comes to scrollbars. So the bug probably was present on the other browsers but it just happened to not be hit. I'd have to play around with the dev tools to know for sure.

Also I'm not sure the fix is even correct. overflow-x: hidden means it just chops off any overflowing content which means you don't get a scroll bar, but if the user types to much it just goes into an invisible void they can't see.

See https://developer.mozilla.org/en-US/docs/Web/CSS/Reference/P...

So this could be a case of the AI doing its classic "the symptom is gone!" thing.

Re: Claude Fable is relentlessly proactive

#638
post #3

> But on the other hand... this is a robust reminder that coding agents can do anything you can do by typing commands into a terminal—and frontier models know every trick in the book and evidently a few that nobody has ever written down before. > Running coding agents outside of a sandbox has always been a bad idea I'm continually bemused and astonished by the number of people who clearly acknowledge that it's reckle…

How can you get the agents to do anything useful without giving them meaningful access? If it only lives in an isolated sandbox, it can only act within the sandbox, then I would have to manually move what was done in the sandbox to real-life. I am not saying it should have critical access, but this is more of a question: How can you get value out of AI if it can only act in a sandbox?

The same way you get value out of a dev container.

Re: Claude Fable is relentlessly proactive

#639

I have a feeling like such posts come from a parallel reality. In my anecdotal experience confirmed by my (still subjective) benchmark ( https://pshirshov.github.io/llm-bench-pi-oneshot/ ) Fable is not _that_ impressive. I performs on par with gpt-5.5 and opus 4.8, sometimes better, sometimes worse, it's definitely more expensive and it likes to refuse answering questions about React saying it can't help with chemist…

My experience with Fable since its release matches Simon's. I've been having it orchestrate complex implementations. I give it a parent ticket (issue) on Linear and say "look at the sub-issues on this ticket and determine which ones you can implement yoursef, in which order, and determine how your implementation will need to be coordinated with what is currently being worked on by other team members". These tickets a…

Ok, explain me one thing: I have a benchmark - I feed identical prompt to multiple models. Codex produces a rough but working program. Fable produces the same - but with more bugs than Codex. Opus produces something similar to Codex but with a critical bug.

That describes all my tests with Fable.

Why should I be hyped about all that "legitimate power" if the model performs on par with two other SoTAs?

I mean, well, yes, it is impressive. It could quickly generate a lot of garbage which sorta does look like code. Two others can do the same. I don't see any groundbreaking improvement - but the price is much higher. Why the hype?

Re: Claude Fable is relentlessly proactive

#640
post #467

Earlier quoted context omitted.

It was only pursuing the goal you gave it - Keep Summer Safe.

"Oh my God"

I relent to snarky Rick and Morty quotes because I don't know that it's useful any more to try to explain paperclip optimizers or alignment to a bunch of AI nerds who saw the cliff coming and clawed at each other trying to be the first out to leap over the edge.

"Relentlessly proactive". That's one word for it. We have a whole subgenre of hard takeoff scenarios and it wasn't enough warning against "Relentlessly proactive".

Turns out Frank Herbert was an optimist, and we're literally pinning our survival on robots turning out to naturally have impractically short attention spans.

Post reply on HN