Earlier quoted context omitted.
DeepSeek never mopped the floor with anyone... DeepSeek was remarkable because it is claimed that they spent a lot less training it, and without Nvidia GPUs, and because they had the best open weight model for a while. The only area they mopped the floor in was open source models, which had been stagnating for a while. But qwen3 mopped the floor with DeepSeek R1.
counterpoint: influencers said they wiped the floor with everyone so it must have happened
Claude 4
471–480 of 1001 posts
Re: Claude 4
#472After that debacle on X, I will not try anything that comes from anthropic for sure. Be careful!
Re: Claude 4
#473Earlier quoted context omitted.
Yah Claude tends to output 1200+ line architectural specification documents while Gemini tends to output ~600 line. (I just had to write 100+ architectural spec documents for 100+ different apps) Not sure why Claude is more thorough and complete than the other models, but it's my go-to model for large projects. The OpenAI model outputs are always the smallest - 500 lines or so. Not very good at larger projects, but p…
Who's reading these docs?
Re: Claude 4
#474Earlier quoted context omitted.
The option just shown up in Copilot settings page for me
Turns out Opus 4 starts at their $40/mo ("Pro+") plan which is sad, and they serve o4-mini and Gemini as well so it's a bit less exclusive than this announcement implies. That said, I have a random question for any Anthropic-heads out there: GitHub says "Claude Opus 4 is hosted by Anthropic PBC. Claude Sonnet 4 is hosted by Anthropic 1P."[1]. What's Anthropic 1P ? Based on the only Kagi result being a deployment tuto…
So far I have found it pretty powerful, its also the first time an LLM has ever stopped while working to ask me a question or for clarification.
Re: Claude 4
#475Already test Opus 4 and Sonnet 4 in our SQL Generation Benchmark ( https://llm-benchmark.tinybird.live/ ) Opus 4 beat all other models. It's good.
Re: Claude 4
#476Can anyone help me understand why they changed the model naming convention? BEFORE: claude-3-7-sonnet AFTER: claude-sonnet-4
Claude 3 arrived as a family (Haiku, Sonnet, Opus), but no release since has included all three sizes.
A release of "claude-3-7-sonnet" alone seems incomplete without Haiku/Opus, when perhaps Sonnet is has its own development roadmap (claude-sonnet-*).
Re: Claude 4
#477Earlier quoted context omitted.
It's interesting that we are in a world state in which "HATE the new personality it's got" is applied to AIs. We're living in the future ya'll :)
"His work was good, I just would never want to get a beer with him" this is literally how you know were approaching agi
If I’m asking it to help me analyse legal filings (for example) I don’t want breathless enthusiasm about my supposed genius on spotting inconsistencies. I want it to note that and find more. It’s exhausting having it be like this and it makes me feel disgusted.
Re: Claude 4
#478I can't be the only one who thinks this version is no better than the previous one, and that LLMs have basically reached a plateau, and all the new releases "feature" are more or less just gimmicks.
This is the new stochastic parrots meme. Just a few hours ago there was a story on the front page where an LLM based "agent" was given 3 tools to search e-mails and the simple task "find my brother's kid's name", and it was able to systematically work the problem, search, refine the search, and infer the correct name from an e-mail not mentioning anything other than "X's favourite foods" with a link to a youtube video. Come on!
That's not to mention things like alphaevolve, microsoft's agentic test demo w/ copilot running a browser, exploring functionality and writing playright tests, and all the advances in coding.
Re: Claude 4
#479An important note not mentioned in this announcement is that Claude 4's training cutoff date is March 2025, which is the latest of any recent model. (Gemini 2.5 has a cutoff of January 2025) https://docs.anthropic.com/en/docs/about-claude/models/overv...
Re: Claude 4
#480It’s been hard to keep up with the evolution in LLMs. SOTA models basically change every other week, and each of them has its own quirks. Differences in features, personality, output formatting, UI, safety filters… make it nearly impossible to migrate workflows between distinct LLMs. Even models of the same family exhibit strikingly different behaviors in response to the same prompt. Still, having to find each model’…
My advice: don't jump around between LLMs for a given project. The AI space is progressing too rapidly right now. Save yourself the sanity.