Live data from Hacker News

Gemini 3.8 Live and 3.8 Live Extended Thinking

blog.google

281–290 of 334 posts

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#281

Earlier quoted context omitted.

I feel like a 10 year old all over again with my frequency of questions. How did the early Roman Empire interact with Greek city states? How are LLMs planning on learning new information on the fly without new context or retraining runs? Why did mom leave? You know, standard stuff.

> Why did mom leave? Reinforcement learning by human feedback.

This would be horrendous if it weren't so true.-

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#283

Earlier quoted context omitted.

I strongly agree. I suspect it's people who have not yet used the paid models from OpenAI and Anthropic. Gemini is comparable to free models from other providers, but not in the same universe as paid models. This is frustrating because when I discuss AI with laypeople they think it's still incapable of counting the number of Rs in "strawberry." They believe it to be essentially useless and incapable of basic tasks. W…

> I suspect it's people who have not yet used the paid models from OpenAI and Anthropic. Gemini is comparable to free models from other providers, but not in the same universe as paid models. I totally disagree. I pay for both ChatGPT and Anthropic (haven't tried the chinese models yet) and yet Gemini is my go-to model (and I pay for it too through Google Workspace subscriptions for several domain names tied to Googl…

Same here.

I use gemini for everything not-coding, from doing research, to have custom personas for more niche topics (and feeding more detailed knowledge in these cases).

For coding and image editing, right now I find ChatGPT superior. And for software architecture designs or planning Claude is the best since a while. I still have to try Grok to be fair.

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#284

Earlier quoted context omitted.

I strongly agree. I suspect it's people who have not yet used the paid models from OpenAI and Anthropic. Gemini is comparable to free models from other providers, but not in the same universe as paid models. This is frustrating because when I discuss AI with laypeople they think it's still incapable of counting the number of Rs in "strawberry." They believe it to be essentially useless and incapable of basic tasks. W…

> you're just not using the latest model, bro Pro tip, ChatGPT is the normiest of all normie websites right now. You're not part of the cognoscenti just because you learned how to type prompts into one of the most popular websites in the world. P.S. You're probably not using OpenAI models for complex or non-standard tasks. It shits the best just as often as Qwen when you need precision and detail in a non-obvious pro…

Impressive how feverishly you defend the steaming pile of shit that Gemini is.

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#285

Earlier quoted context omitted.

To be fair, it has been about six months since I tried a paid Google model. I will give it another go to compare. Hallucinations were the main issue back then but perhaps it has come a long way.

Much respect for being able to admit that you haven’t used a model in a while after saying folks probably haven’t used the models you use.

Okay I just tested 3.8 Flash (High) on some real world problems I have used Opus (High) and Sol (high) to solve. The outcome here is terrible.

One of the problems is local zoning laws regarding an expansion of my house. Comparison of annex vs extension, boundaries, precedent, costs, etc.

3.8 Flash didn't check most of the required zoning laws. It relied on parametric knowledge, which is outdated and inaccurate. It checked zero precedents. It did made a very cursory check of the boundary area, but didn't validate it, so it missed a lot of important nuance and exceptions to the boundary. Its cost estimates were wildly inaccurate. Ostensibly because it was inferring an average based on historical pricing data rather than gathering current info.

I could go on, but if I had to judge this attempt I would give it a 3/10. It's very fast, but wildly inaccurate. It's clear that the model is designed for speed over accuracy.

But don't take my word for it. [Most benchmarks show it to be significantly below frontier models like Astra.](https://llm-stats.com/models/compare/gemini-3.8-flash-vs-gpt...)

This has been a useful exercise. It's important to understand the developments taking place. I am disappointed to see that Google has made very little progress in six months relative to the frontier labs.

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#286
post #231

I don’t understand good experiences people are having with Gemini. It’s the only model that sometimes loses/forgets context in literally next message. Plus feeding unasked product links to responses.

I also had that issue happen to me, surprisingly only when I used Polish, not English. But other than that I actually like Gemini, I check stuff against it all the time, especially when walking my dog. It also works in AndroidAuto for me, but I use very basic stuff like changing Spotify music. I was on iOS before and there's no comparison to old Siri that was just garbage. I use Gemini practically every day, it's fine for the most part IMO. They really do need to polish integrations though, app connectors barely every work outside of Google's own apps. Using Oppo's app Mind Place through Gemini f.e. is just bad experience and almost never works. For example, I had pleasant experience of Gemini finding me places for a walk/hike on vacation in Tirol, where I specifically required not too much of an ascent and an asphalt road for a stroller.

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#287

Earlier quoted context omitted.

> You're probably not using OpenAI models for complex or non-standard tasks. It shits the best just as often as Qwen when you need precision and detail in a non-obvious problem. The benchmarks clearly show otherwise. This is your cue to tell me the benchmarks are made by the Illuminati and only your superior and subjective methods of evaluation are correct.

The tasks you actually need to do trump benchmarks, yes. I haven't tried out the ridiculously expensive models besides the latest Gemini, and it gave from equal to slightly worse results than latest DeepSeek, at a far higher price. It does also seems Gemini's main problem wasn't that it was stupid, but that it was good at doing slightly different things than what I asked it to, very well. Which might well also have t…

I think 3.8 Flash is on par with DeepSeek on some benchmarks and tasks (not coding or design), but it's not close to Sol/Astra or Opus/Fable. I would not consider a $20 subscription "ridiculously expensive," but I suppose that is a relative term.

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#289

My first language is Afrikaans, which is a somewhat niche language and hard to find teachers/conversation buddies outside South Africa. (I live in USA now) I've been using Gemini to live chat in Afrikaans and do impromptu Afrikaans grammar lessons during my solo drives around town. It is phenomenal at speaking the language - like, it really shocks my family members when they hear it. This is probably the most joy I g…

I like it. I had a friend at school who came from SA and spoke Afrikaans - so I know how a natural sounds like.

For language learning, translation, research - it is fantastic.

It works very well for certain dialects as part of a local language heritage as well: From Welsh to Bavarian - it is a joy to really get approval from people used to these dialects as being really accurate.

Happpy times! :)

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#290
post #18

I'm disappointed with "Extended Thinking" for 3.8 Flash. On the plus side, it's a strong general-purpose model and the cost-benefit is still compelling. However, the "Extended Thinking" should be renamed to "Slightly Extended Thinking". Considering that it's the maximum thinking option for Gemini Flash in the chat UI, it doesn't actually think a whole lot, leading to an uncomfortably high number of incorrect/poor rep…

Have you tried selecting the retry option below a reply? I think it let's you get a longer answer at least.

I used Retry plenty - but not for good reasons. Gemini fails quite a bit (error) in the app and chat interfaces.
Post reply on HN