Earlier quoted context omitted.
Is it possible they are just falling behind ? Their newest model wasn’t really SOTA. And honestly fable 5 was the most human like model I’d ever tried. It was an incredible jump. And recently lots of Claude users at r/ClaudeAI are noticing Opus 4.8 has really increased in capability. Not new things but maybe redirected compute. It just feels like one of the best models ever, maybe because the compute that was previou…
Google's AI is hamstrung by a culture of safetyisim, by that I mean going beyond what we can all recognise as safe limits to protect the user from imaginary ephemeral things like cultural harms. So maximal safety at all costs is in itself a cost. They can spend billions on AI but that spend is down the toilet if the user bounces because the AI's persona is a relentless politically correct scold.
John Jumper to join Anthropic
101–110 of 166 posts
Re: John Jumper to join Anthropic
#102Re: John Jumper to join Anthropic
#103Who cares?
Re: John Jumper to join Anthropic
#104Earlier quoted context omitted.
>from the looks of it, 3.5 Flash is still better than most models Define "better". I guess it depends on what you're using it for. I use it almost daily as an alternative to google search and it's great for that, but I think it's absolute garbage for coding and reasoning. For questions related to coding, solving Arch Linux and WINE Lutris issues, helping me with MXLinux issues, and wifi issues on an old rooted huawei…
Did you look at the charts in the article? It out-performed every model that wasnt a max/ultrafrontier of some sort, except for the one that the article was extolling the virtues of, including grok high. you could make a good argument that deepseek is a better value, but gemini flash is when bundled is already pretty accessible. nowhere did i claim that flash was better than fable or 5.5xhigh.
I don't care about someone else's charts, i care about my own lived experiences. Benchmarks can be gamed to get to the top of charts. When I pay for a service I care about how it performs in my test cases, not about which tops some random charts.
Read my comment again please. I think I was pretty clear with detailed examples on where Gemini sucks and where it's good at.
>nowhere did i claim that flash was better than fable or 5.5xhigh.
And nowhere did I claim that. I said even basic GPT and Grok are better than Gemini Flash at reasoning tasks. Again, read my comment again, I have already explained why with examples.
Re: John Jumper to join Anthropic
#105John Jumper Jumps to Anthropic, was right there.
Re: John Jumper to join Anthropic
#106Anthropic legit builds one the strongest if not the strongest IC team in the history of computational technology. They are insanely stacked on talent, and either we will witness a legendary run, or a new LTCM
why so dramatic? why cant it just be a quietly competent lab, why must it be a dramatic collapse?
Re: John Jumper to join Anthropic
#107Gemini 3.1 flash was actually an amazing model to code with and their 20 dollar AI plans had solid value, but they locked it all behind 429s, needless gatekeeping of clients and poor product differentiation even among internal offerings. Users moved on. To claude for the best product, to OpenAI for the non gatekept API access. It’s hard to bring them back.
Re: John Jumper to join Anthropic
#108Earlier quoted context omitted.
Is it possible they are just falling behind ? Their newest model wasn’t really SOTA. And honestly fable 5 was the most human like model I’d ever tried. It was an incredible jump. And recently lots of Claude users at r/ClaudeAI are noticing Opus 4.8 has really increased in capability. Not new things but maybe redirected compute. It just feels like one of the best models ever, maybe because the compute that was previou…
I think they'll catch up in pure model capabilities but they do such a terrible job of making products from which to use their models that having the best model doesn't end up mattering. Is Gemini good at writing code? I am sure it is. But where is their Codex? And no, antigravity isn't it. Is Gemini good at making visualizations? I am sure it is. But where are artifact or visualise skill in gemini.google.com similar…
Antigravity CLI is quite decent, it's a huge step up from Gemini CLI (like, for example, it actually fucking works) and has some genuine advantages over Claude Code. Does Codex have something over both of them? I haven't tried it.
But the model just fucking sucks. Before I switched to Claude for personal stuff a few weeks ago, I was like "damn model capabilities are really slowing down" but no, it's just Gemini that's slowing down.
Will have to see if 3.5 Pro is any good when that comes out. But it feels like they would be attempting to catch up to Opus, not to Fable.
FWIW issue is never really about the code it writes it's about general intelligence. Gemini hallucinates like it's 2024, fails to follow instructions, and goes down wildly wrong debugging paths. Opus just gets the job done, first time, every time. With Gemini it feels like "I _am_ glad this intern is working for me but I'm tired of babysitting him" and with Claude it's like "this new PhD guy can replace me soon".
Re: John Jumper to join Anthropic
#109Earlier quoted context omitted.
from the looks of it, 3.5 Flash is still better than most models https://artificialanalysis.ai/articles/glm-5-2-is-the-new-le... The idea of "falling behind" when you can leapfrog each other every six months leads me to believe it has to be more than just "falling behind" for one cycle. It's a culture, process, red tape, focus, or mandate problem of some sort. Something not as easily correctable preparing for next la…
>from the looks of it, 3.5 Flash is still better than most models Define "better". I guess it depends on what you're using it for. I use it almost daily as an alternative to google search and it's great for that, but I think it's absolute garbage for coding and reasoning. For questions related to coding, solving Arch Linux and WINE Lutris issues, helping me with MXLinux issues, and wifi issues on an old rooted huawei…
Re: John Jumper to join Anthropic
#110Earlier quoted context omitted.
Did you look at the charts in the article? It out-performed every model that wasnt a max/ultrafrontier of some sort, except for the one that the article was extolling the virtues of, including grok high. you could make a good argument that deepseek is a better value, but gemini flash is when bundled is already pretty accessible. nowhere did i claim that flash was better than fable or 5.5xhigh.
>Did you look at the charts in the article? I don't care about someone else's charts, i care about my own lived experiences. Benchmarks can be gamed to get to the top of charts. When I pay for a service I care about how it performs in my test cases, not about which tops some random charts. Read my comment again please. I think I was pretty clear with detailed examples on where Gemini sucks and where it's good at. >no…