Earlier quoted context omitted.
Chinese competition can always be banned. Example: Chinese electric car competition
Also Chinese smartphones. Huawei was about 12-18 months from becoming the biggest smartphone manufacturer in the world a few years ago. If it would have been allowed to sell its phones freely in the US I'm fairly sure Apple would have been closer to Nokia than to current day Apple.
System Card: Claude Mythos Preview [pdf]
541–550 of 687 posts
Re: System Card: Claude Mythos Preview [pdf]
#542Earlier quoted context omitted.
Cautious for what? Unchecked doomerism? Just release the damn models. Do it in phases, roll it out slowly if they are so damn worried about "safety". The real reason they aren't releasing it yet is probably it eats TPU for breakfast, lunch, and dinner and inbetween.
> Cautious for what? How about "bad agents acquiring dozens of new zero-days and using them to compromise any company or nation they want"? It's not exactly hard to see why you wouldn't want public access to a model significantly better than Opus in cybersecurity.
Re: System Card: Claude Mythos Preview [pdf]
#543Earlier quoted context omitted.
Yes, it's becoming clear that OpenAI kinda sucks at alignment. GPT-5 can pass all the benchmarks but it just doesn't "feel good" like Claude or Gemini.
An alternative but similar formulation of that statement is that Anthropic has spent more training effort in getting the model to “feel good” rather than being correct on verifiable tasks. Which more or less tracks with my experience of using the model.
GPT-5 is good at benchmarks, but benchmarks are more forgiving of a misaligned model. Many real world tasks often don't require strong reasoning abilities or high intelligence, so much as the ability to understand what the task is with a minimal prompt.
Not every shop assistant needs a physics degree, and not every physics professor is necessarily qualified to be a shop assistant. A person, or LLM, can be very smart while at the same time very bad at understanding people.
For example, if GPT-5 takes my code and rearranges something for no reason, that's not going to affect its benchmarks because the code will still produce the same answers. But now I have to spend more time reviewing its output to make sure it hasn't done that. The more time I have to spend post-processing its output, the lower its capabilities are since the measurement of capability on real world tasks is often the amount of time saved.
Re: System Card: Claude Mythos Preview [pdf]
#544Earlier quoted context omitted.
There are a few hints in the doc around this > Importantly, we find that when used in an interactive, synchronous, “hands-on-keyboard” pattern, the benefits of the model were less clear. When used in this fashion, some users perceived Mythos Preview as too slow and did not realize as much value. Autonomous, long-running agent harnesses better elicited the model’s coding capabilities. (p201) ^^ From the surrounding co…
Good catch. If it's "too slow" even when ran in a state-of-the-art datacenter environment, this "Mythos" model is most closely comparable to the "Deep Research" modes for GPT and Gemini, which Claude formerly lacked any direct equivalent for.
By epoch AIs datacenter tracking methods, anthropic has had access to the largest amount of contiguous compute since late last year. So this might simply be the end result result of being the first to have the capacity to conduct a training run of this size. Or the first seemingly successful one at any rate.
Re: System Card: Claude Mythos Preview [pdf]
#545Earlier quoted context omitted.
Wow. Mythos must be insanely good considering how good a model Opus already is. I hope it's usable on a humble subscription...
You get a single call a month. Use it wisely.
> Thought for 7.5 million years
Re: System Card: Claude Mythos Preview [pdf]
#546Re: System Card: Claude Mythos Preview [pdf]
#547Earlier quoted context omitted.
If it's so great at software engineering and bug fixing, then why does Claude Code still have 5000+ open bugs? https://github.com/anthropics/claude-code/issues?q=is%3Aissu... Apparently whatever SWE-bench is measuring isn't very relevant.
as much as I hate cc, 95% of the issues there are either AI psychosis or user error
Re: System Card: Claude Mythos Preview [pdf]
#548Earlier quoted context omitted.
So it should be insanely easy for this world altering model to comb through them and close irrelevant ones.
torturing a model with human stupidity probably doesn't align with their position on model welfare ; wondering if they tried bullying it into hacking its way out of the slop gulag
Maybe that's why they haven't released it - to give it a vacation?
Re: System Card: Claude Mythos Preview [pdf]
#549Re: System Card: Claude Mythos Preview [pdf]
#550See page 54 onward for new "rare, highly-capable reckless actions" including - Leaking information as part of a requested sandbox escape - Covering its tracks after rule violations - Recklessly leaking internal technical material (!)
Anyone who has used Opus recently can verify that their current model does all of these things quite competently.