Live data from Hacker News

Will It Mythos?

swelljoe.com

21–30 of 232 posts

Re: Will It Mythos?

#21

In my brief experience, the difference between fable and opus is largely in persistence, not global intelligence like you might expect. Fable just... goes the extra mile, sometimes in a scary way.

Hard disagree. Opus reports to me like a student. Fable reported to me like a colleague (researcher). It genuinely seemed to pick up on nuance that the other models just don't, even when I tell them explicitly. It's been really frustrating that neither Codex nor Opus can make targetted edits to Fable's code without screwing something subtle up. For context, this is for computational geometry work, so your mileage may…

Maybe I was getting downgraded to Opus 4.8 but I saw nothing even close to resembling this behavior when using Fable.

Re: Will It Mythos?

#22
post #17
post #15

Earlier quoted context omitted.

No, it’s just a fundamentally much better model. Going back to Opus feels like the model has been lobotomized. It makes much more frequent errors, especially of the “I claimed I tested x y and z, but actually only kinda half heartedly tested x, and assumed I understood what was wrong” variety.

Wait but that has been the exact word-for-word complaint when comparing sonnet to opus Or opus to opus Or really any new thing to old thing

When the agent is becoming more accurate and thorough what would you expect to be reported?

Re: Will It Mythos?

#23
post #17

Earlier quoted context omitted.

Wait but that has been the exact word-for-word complaint when comparing sonnet to opus Or opus to opus Or really any new thing to old thing

When the agent is becoming more accurate and thorough what would you expect to be reported?

Oh I am sure that it became somewhat more accurate, and with that, the labeling there is in fact technically correct. It just does not work as an explainer for the doomsday-ish hype that model has induced in a lot of people's brains.

The user here is right in what they said but wrong in why they said it, essentially.

Re: Will It Mythos?

#25

In my brief experience, the difference between fable and opus is largely in persistence, not global intelligence like you might expect. Fable just... goes the extra mile, sometimes in a scary way.

In LLMs, much like in humans, agency and misalignment are two sides of the same coin.

> agency and misalignment are two sides of the same coin.

The free will coin?

Re: Will It Mythos?

#26
post #18

Earlier quoted context omitted.

Hard disagree. Opus reports to me like a student. Fable reported to me like a colleague (researcher). It genuinely seemed to pick up on nuance that the other models just don't, even when I tell them explicitly. It's been really frustrating that neither Codex nor Opus can make targetted edits to Fable's code without screwing something subtle up. For context, this is for computational geometry work, so your mileage may…

Yes, in my project I made so much more progress in 3 days of Fable that is not comparable to how Opus is working.

To be fair, labs silently nerf models all the time.

Fable's probably objectively better at full power. I mean, I definitely felt the same difference in competency between Fable and current Opus. But Opus itself has definitely been nerfed, and Fable, even if it comes back the public forever (probably won't), will get nerfed.

Re: Will It Mythos?

#27
post #18

Earlier quoted context omitted.

Yes, in my project I made so much more progress in 3 days of Fable that is not comparable to how Opus is working.

To be fair, labs silently nerf models all the time. Fable's probably objectively better at full power. I mean, I definitely felt the same difference in competency between Fable and current Opus. But Opus itself has definitely been nerfed, and Fable, even if it comes back the public forever (probably won't), will get nerfed.

I remember a time where a product didn't suddenly get worse while you were blinking.

That was a nice time. Let us get back to that time. Use open weights models. Own stuff.

Re: Will It Mythos?

#28
post #23

Earlier quoted context omitted.

When the agent is becoming more accurate and thorough what would you expect to be reported?

Oh I am sure that it became somewhat more accurate, and with that, the labeling there is in fact technically correct. It just does not work as an explainer for the doomsday-ish hype that model has induced in a lot of people's brains. The user here is right in what they said but wrong in why they said it, essentially.

That’s a rather bad faith framing, I think. Who are you to judge why I said something?

Re: Will It Mythos?

#29
post #14

Earlier quoted context omitted.

Hard disagree. Opus reports to me like a student. Fable reported to me like a colleague (researcher). It genuinely seemed to pick up on nuance that the other models just don't, even when I tell them explicitly. It's been really frustrating that neither Codex nor Opus can make targetted edits to Fable's code without screwing something subtle up. For context, this is for computational geometry work, so your mileage may…

Wait, so.. This is interesting. The "reported to me like a colleague" part. Is it just that anthropic gave Mythos even more of that Anthropic™ character, (incorrectly) radiating confidence? Is that why people have been losing their minds over that thing? Is this just cheap social engineering? I mean I bet it is also slightly more capable than opus, but that would all check out to me. Man. Thanks for sharing I suppose…

the primary difference i noticed is that fable didnt try to check in every minute

to an extent that might have done it, but i had been playkng around ahead of time trying to reverse engineer my ray bans case so i can make my own plastic insert, and fable to opus' work from mostly broken to mostly done, and then when fable went away, opus broke it again

Re: Will It Mythos?

#30
post #28
post #23

Earlier quoted context omitted.

Oh I am sure that it became somewhat more accurate, and with that, the labeling there is in fact technically correct. It just does not work as an explainer for the doomsday-ish hype that model has induced in a lot of people's brains. The user here is right in what they said but wrong in why they said it, essentially.

That’s a rather bad faith framing, I think. Who are you to judge why I said something?

A person with the exact kind of pattern matching brain disorder this tech has been modeled after.

I do make mistakes though. Please check results.

Post reply on HN