Earlier quoted context omitted.
Lol you haven't used a model since GPT2 is what it sounds like.
Just checked my subscription start date for Anthropic. September 2023, I believe before they announced public launch. Sorry kid.
System Card: Claude Mythos Preview [pdf]
81–90 of 687 posts
Re: System Card: Claude Mythos Preview [pdf]
#82> Claude Mythos Preview’s large increase in capabilities has led us to decide not to make it generally available. All the more reason somebody else will. Thank God for capitalism.
Re: System Card: Claude Mythos Preview [pdf]
#83Earlier quoted context omitted.
Haven't seen a jump this large since I don't even know, years? Too bad they are not releasing it anytime soon (there is no need as they are still currently the leader).
A jump that we will never be able to use since we're not part of the seemingly minimum 100 billion dollar company club as requirement to be allowed to use it. I get the security aspect, but if we've hit that point any reasonably sophisticated model past this point will be able to do the damage they claim it can do. They might as well be telling us they're closing up shop for consumer models. They should just say they…
> They should just say they'll never release a model of this caliber to the public at this point and say out loud we'll only get gimped
Duh, this was fucking obvious from the start. The only people saying otherwise were zealots who needed a quick line to dismiss legitimate concerns.
Re: System Card: Claude Mythos Preview [pdf]
#84Are you guys ready for the bifurcation when the top models are prohibitively expensive to normal users? If your AI budget $2000+ a month? Or are you going to be part of the permanent free tier underclass?
Inference for the same results has been dropping 10x year over year[0] [0] https://ziva.sh/blogs/llm-pricing-decline-analysis
Re: System Card: Claude Mythos Preview [pdf]
#85This is the first moment where the whole “permanent underclass” meme starts to come into view. I had through previously that we the consumers would be reaping the benefits of these frontier models and now they’ve finally come out and just said it - the haves can access our best, and have-nots will just have use the not-quite-best.
Perhaps I was being willfully ignorant, but the whole tone of the AI race just changed for me (not for the better).
Re: System Card: Claude Mythos Preview [pdf]
#86> Claude Mythos Preview’s large increase in capabilities has led us to decide not to make it generally available. Shame. Back to business as usual then.
I for one applaud them for being cautious.
Re: System Card: Claude Mythos Preview [pdf]
#87Earlier quoted context omitted.
"some model I don't get to use is much better at benchmarks" pick one or more: comically huge model, test time scaling at 10e12W, benchmark overfit
So... you're not excited because it might take a few months before we can use it or something? I don't get your comment.
Re: System Card: Claude Mythos Preview [pdf]
#88Re: System Card: Claude Mythos Preview [pdf]
#89Earlier quoted context omitted.
Just checked my subscription start date for Anthropic. September 2023, I believe before they announced public launch. Sorry kid.
So you are doubly stupid, by not seeing any improvement in the models and also paying for models you believe are terrible? lol
Re: System Card: Claude Mythos Preview [pdf]
#90Earlier quoted context omitted.
Honestly we are all sleeping on GPT-5.4. Particularly with the influx of Claude users recently (and increasingly unstable platform) Codex has been added to my rotation and it's surprising me.
GPT is shit at writing code. It's not dumb - extra high thinking is really good at catching stuff - but it's like letting a smart junior into your codebase - ignore all the conventions, surrounding context, just slop all over the place to get it working. Claude is just a level above in terms of editing code.
Not always, no, and it takes investment in good prompting/guardrails/plans/explicit test recipes for sure. I'm still on average better at programming in context than Codex 5.4, even if slower. But in terms of "task complexity I can entrust to a model and not be completely disappointed and annoyed", it scores the best so far. Saves a lot on review/iteration overhead.
It's annoying, too, because I don't much like OpenAI as a company.
(Background: 25 years of C++ etc.)