Live data from Hacker News

Gemini 3.8 Live and 3.8 Live Extended Thinking

blog.google

291–300 of 335 posts

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#292
post #231

I don’t understand good experiences people are having with Gemini. It’s the only model that sometimes loses/forgets context in literally next message. Plus feeding unasked product links to responses.

I’ve found this to be true as well. It’s got really poor attention and will derail into world-building fast. I’m convinced Google is just shipping it to capture market share but they know Gemini isn’t ready for serious use.

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#293

Earlier quoted context omitted.

The tasks you actually need to do trump benchmarks, yes. I haven't tried out the ridiculously expensive models besides the latest Gemini, and it gave from equal to slightly worse results than latest DeepSeek, at a far higher price. It does also seems Gemini's main problem wasn't that it was stupid, but that it was good at doing slightly different things than what I asked it to, very well. Which might well also have t…

I think 3.8 Flash is on par with DeepSeek on some benchmarks and tasks (not coding or design), but it's not close to Sol/Astra or Opus/Fable. I would not consider a $20 subscription "ridiculously expensive," but I suppose that is a relative term.

I would run into the use limits very quickly, and (for Anthropic) have to switch frameworks.

By all accounts they are far more expensive than DeepSeek, and vs. Gemini I've found out that myself.

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#294

Earlier quoted context omitted.

> you're just not using the latest model, bro Pro tip, ChatGPT is the normiest of all normie websites right now. You're not part of the cognoscenti just because you learned how to type prompts into one of the most popular websites in the world. P.S. You're probably not using OpenAI models for complex or non-standard tasks. It shits the best just as often as Qwen when you need precision and detail in a non-obvious pro…

> You're probably not using OpenAI models for complex or non-standard tasks. It shits the best just as often as Qwen when you need precision and detail in a non-obvious problem. The benchmarks clearly show otherwise. This is your cue to tell me the benchmarks are made by the Illuminati and only your superior and subjective methods of evaluation are correct.

artificialanalysis just updated their benchmark after the release of GPT-6. They removed old, saturated benchmarks and replaced them with new, until GPT-6 floated to the top with the cream. One of those new benchmarks is AutomationBench-AA, where GPT-6 had a clear lead. Today that benchmark is topped by DeepSeek v4.1 Flash.

Edit: For those who are not familiar with it, this model is quite a bit faster, and about 100x cheaper, per token, than Fable and Astra.

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#296
post #36

Earlier quoted context omitted.

As an "everyday mans AI" I'd say 3.8 Flash definitely already has. Smart enough for the vast swath of people, and only slightly eeked out by Astra(Max) on vision capabilities, like the kind of "Point your camera at something and ask questions" that non-tech people like to do. It's crazy fast and very compute light, so not getting bogged down constantly. I can't think of a better general purpose model than 3.8 flash r…

Yeah its good. Reasonably priced too (at current prices, if they do raise them in January I would stop recommending it). 3.7/3.8 were good releases.

where did you see the prices?

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#297
post #231

I don’t understand good experiences people are having with Gemini. It’s the only model that sometimes loses/forgets context in literally next message. Plus feeding unasked product links to responses.

For information retrieval and collation related tasks (reading lists, deep dives on subjects, etc.), Gemini is way better w.r.t. other models in my experience.

This difference is probably due to Google's web knowledge and free pass to YouTube. However, I'm happy what I got from it so far.

When I ask the question once in a blue moon, it can generally one-shot the answer, even.

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#298

Earlier quoted context omitted.

I think 3.8 Flash is on par with DeepSeek on some benchmarks and tasks (not coding or design), but it's not close to Sol/Astra or Opus/Fable. I would not consider a $20 subscription "ridiculously expensive," but I suppose that is a relative term.

I would run into the use limits very quickly, and (for Anthropic) have to switch frameworks. By all accounts they are far more expensive than DeepSeek, and vs. Gemini I've found out that myself.

No argument that DeepSeek is much cheaper, but you certainly get what you pay for.

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#299

Earlier quoted context omitted.

Much respect for being able to admit that you haven’t used a model in a while after saying folks probably haven’t used the models you use.

Okay I just tested 3.8 Flash (High) on some real world problems I have used Opus (High) and Sol (high) to solve. The outcome here is terrible. One of the problems is local zoning laws regarding an expansion of my house. Comparison of annex vs extension, boundaries, precedent, costs, etc. 3.8 Flash didn't check most of the required zoning laws. It relied on parametric knowledge, which is outdated and inaccurate. It ch…

Gemini's search harness in the Google app is (ironically) bad so it makes the model look bad.

If you really want to compare apples to apples you need to test Gemini models against other models using the same third party search harness.

Otherwise you are largely measuring how much computation the model provider is allocating to a search harness.

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#300

Earlier quoted context omitted.

> You're probably not using OpenAI models for complex or non-standard tasks. It shits the best just as often as Qwen when you need precision and detail in a non-obvious problem. The benchmarks clearly show otherwise. This is your cue to tell me the benchmarks are made by the Illuminati and only your superior and subjective methods of evaluation are correct.

artificialanalysis just updated their benchmark after the release of GPT-6. They removed old, saturated benchmarks and replaced them with new, until GPT-6 floated to the top with the cream. One of those new benchmarks is AutomationBench-AA, where GPT-6 had a clear lead. Today that benchmark is topped by DeepSeek v4.1 Flash. Edit: For those who are not familiar with it, this model is quite a bit faster, and about 100x…

It's very impressive that DeepSeek 4.1 beats Astra in one benchmark, but I presume you are aware that Astra wins in almost all other benchmarks? These are some of them: https://llm-stats.com/models/compare/deepseek-v4.1-flash-vs-...
Post reply on HN