Live data from Hacker News

sqlite-utils 4.0rc2, mostly written by Claude Fable (for about $149.25)

simonwillison.net

31–40 of 93 posts

Re: sqlite-utils 4.0rc2, mostly written by Claude Fable (for about $149.25)

#32
post #10
post #7

Fun fact: because AI written works don't have copyright (in the EU at least) and the level of prompting many people engage in doesn't suffice to create a copyrightable "work" and software licenses require you to actually be able to grant a license using rights you hold on a work, not only are many AI generated "works" not actually protected by copyright but by selling licenses you're actually in breach of contract la…

And nothing happened and zero people got in trouble over it. - Narrator

...So far

Re: sqlite-utils 4.0rc2, mostly written by Claude Fable (for about $149.25)

#33
post #3

The problem I have with this workflow is that the models are still too eager to please. If I ask it to scan a release and note possible issues, it absolutely will find issues. If I keep running the same prompt, it will keep finding issues. I’ve spammed GitHub PR reviews and it just keep finding (or inventing?) new issues. There is never a “Nothing found, good to go!”. I have to keep reminding myself that the model wi…

It's not eagerness to please (that's anthropomorphising), rather it's a desire to bill you more money/use more tokens

(The fixed prices are just temporary discounts)

Re: sqlite-utils 4.0rc2, mostly written by Claude Fable (for about $149.25)

#34
post #18

Earlier quoted context omitted.

> especially in the light of GPT 5.6 very likely coming out next week Finally have an explanation why GPT 5.5 xhigh felt dumber and dumber these last few weeks, always the same thing when a new model release is about to come out...

Opus has been extremely stupid recently, reckon that's because Fable needs to look appealing?

I have never noticed a degradation in either Claude or OpenAI models, and the benchmarks people set up have never shown a statistically significant deviation either: https://marginlab.ai/trackers/claude-code

Yet the same claim is being posted every single day, including new claims that the Fable 5 model has degraded compared to the initial release, guardrails aside.

Re: sqlite-utils 4.0rc2, mostly written by Claude Fable (for about $149.25)

#35
post #18

Earlier quoted context omitted.

Opus has been extremely stupid recently, reckon that's because Fable needs to look appealing?

I have never noticed a degradation in either Claude or OpenAI models, and the benchmarks people set up have never shown a statistically significant deviation either: https://marginlab.ai/trackers/claude-code Yet the same claim is being posted every single day, including new claims that the Fable 5 model has degraded compared to the initial release, guardrails aside.

Almost slipping into conspiracy territory, but without insights into what the labs actually do internally, hard not to:

Anyways, heard about A/B testing before? ML people tend to like it a lot, hard to imagine neither OpenAI or Anthropic are already deep into categorizing people into buckets and running an wild amount of A/B testing all over the place, especially in the weeks leading up to new model releases, in various ways.

Re: sqlite-utils 4.0rc2, mostly written by Claude Fable (for about $149.25)

#36

Earlier quoted context omitted.

That's just plain wrong. The new models do not hallucinate as much as they used to (in my personal experience)

> plain wrong > (in my experience) What are you even saying.

That their vibes are more real than your vibes

Re: sqlite-utils 4.0rc2, mostly written by Claude Fable (for about $149.25)

#37
post #7

Fun fact: because AI written works don't have copyright (in the EU at least) and the level of prompting many people engage in doesn't suffice to create a copyrightable "work" and software licenses require you to actually be able to grant a license using rights you hold on a work, not only are many AI generated "works" not actually protected by copyright but by selling licenses you're actually in breach of contract la…

IMO the real "fun fact" is that supposed "IP monopolists" like Microsoft and Oracle's lawyers are apparently totally fine with this stuff.

So obviously people are going to take their lead and not get legal advice from some greasy dweeb at the bottom of HN.

Re: sqlite-utils 4.0rc2, mostly written by Claude Fable (for about $149.25)

#39
post #3

The problem I have with this workflow is that the models are still too eager to please. If I ask it to scan a release and note possible issues, it absolutely will find issues. If I keep running the same prompt, it will keep finding issues. I’ve spammed GitHub PR reviews and it just keep finding (or inventing?) new issues. There is never a “Nothing found, good to go!”. I have to keep reminding myself that the model wi…

What do you mean? Are they valid flaws or not?

Would you like it to stop when there's still flaws in the code?

Re: sqlite-utils 4.0rc2, mostly written by Claude Fable (for about $149.25)

#40
post #8

Earlier quoted context omitted.

You didn’t do it enough. They stop finding bugs eventually. Also, different models can find different bugs (though they do find the same ones, too, which is good and expected). For best results you want to run multi model reviews in loops. If you had multiple people look at your PRs multiple times on different days results would be very similar.

I've had it find bug, I asked it to make test to trigger the bug, and then it figured out it's not a bug. It will absolutely do wish fulfilment

It'll find a non-existent bug - fix it - figure out it broke a previously working thing - try to fix again - etc..

The "keep improving" the code base prompt have been tried and it never works. The LLM has no consciousness of where to stop and where to draw the lines of reasonableness.

Post reply on HN