Live data from Hacker News

Gemini 3.8 Live and 3.8 Live Extended Thinking

blog.google

231–240 of 334 posts

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#232
post #9

I wonder when/if we’ll see Gemini beating Fable and Astra. Last year I would have confidently bet Google will overtake the others just because they have the data, the hardware (TPUs) and a fat advertising money pipe and yet they are still behind. Anyone anonymous at Google want to hint when Gemini 4 will be out?

Anecdata, but I've been using gemini personally instead of what I used regular google searches for the better part. Even in car chat/search and some lighter and not so light research (if it's not heavy on technicals). It's good enough that I use it nonstop like that. I wouldn't trust it for coding at all - switching between fable and now astra. Google in a sense won (me over) like that. I also expected them to brute…

I just use Grok for anything "controversial" like that. I actually love that different LLMs have their niches. In a typical week I might use all of Claude, Gemini, and Grok, for different kinds of asks.

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#234

Earlier quoted context omitted.

I have a similar experience using Gemini for quick Catalan translations for iOS apps given enough context. I once asked it to summarize The Hobbit in Catalan to explain it to my daughter before sleep. I was expecting a lot of mistakes as I see regularly if I ask anything in my native language when using GPT or Claude, but it was surprisingly good. I was going just to kind of skim ahead and retell it my own way, but e…

> As I work in english and use them in english, I wonder if using LLMs in a different language to code renders a different result as well. Like, if some of these benchmarks were made in other languages, would the result be similar. I am not a Chinese speaker but my understanding is that all of the models have substantially different behavior in Chinese, to the degree that it's kind of like a second model. Would be in…

Fwiw, it's been very easy to nudge Qwen 27B/35B to think in Chinese. No need for prefills like "思考:" ("Think:") or such.

Relatedly, when chatting with frontier models about science education content design, I've found it very helpful to mix in Chinese education terms. In English, for example, NGSS is such a massive attractor, discussing nearby topics often yields NGSS "slop". And "estimation" (educational) in the US means one (dysfunctional:) thing, which similarly distracts. Perhaps if AIs become increasingly multilingual, but remain weak at deep conceptual reasoning, it may be fruitful to have multilingual thesauruses, to use language-associated cultural conceptual differences as a way to convey conceptual nuances with which LLMs otherwise struggle?

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#235
post #231

I don’t understand good experiences people are having with Gemini. It’s the only model that sometimes loses/forgets context in literally next message. Plus feeding unasked product links to responses.

I wonder how much life DeepMind has left in it, especially after Hassabis's departure. Google execs must be having discussions about simply throwing their weight behind Anthropic since they already own so much of the company.

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#236

Earlier quoted context omitted.

> As I work in english and use them in english, I wonder if using LLMs in a different language to code renders a different result as well. Like, if some of these benchmarks were made in other languages, would the result be similar. I am not a Chinese speaker but my understanding is that all of the models have substantially different behavior in Chinese, to the degree that it's kind of like a second model. Would be in…

Fwiw, it's been very easy to nudge Qwen 27B/35B to think in Chinese. No need for prefills like "思考:" ("Think:") or such. Relatedly, when chatting with frontier models about science education content design, I've found it very helpful to mix in Chinese education terms. In English, for example, NGSS is such a massive attractor, discussing nearby topics often yields NGSS "slop". And "estimation" (educational) in the US…

> In English, for example, NGSS is such a massive attractor, discussing nearby topics often yields NGSS "slop".

For that particular attractor, even German or Spanish or so might help avoid it? Or perhaps even just using British English?

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#237

Earlier quoted context omitted.

It's also the only model that generates accurate translation and localization. No other frontier model comes close. Although Gemini's coding capabilities are subpar, its natural language processing is top-tier.

I found that it's shockingly good with R. (the only language I know and can correct for) I doubt they even intended it to be, but it seems like I kept going from resorting to 3.5-3.8 (over time) to realizing that Claude and GPT, while great at Python, will make rudimentary mistakes with R; even when they compose giant complicated R code.

I'm guessing this is partly because of Gemini's world knowledge. I tried asking the model multiple internet humor and memes and it answered correctly around 80% of the time

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#238
post #202

Earlier quoted context omitted.

In Maps? Or is it Android Auto? I have Google built it. When I push the button on steering wheel I hardly know which layer I'm talking to.

Really? I am in Android Auto and I try to even just ask it to get me to "Costco Near Me". I know where it is. I wanna know if the sometimes there traffic jam that blocks the right lane is gonna expect me or not without being distracted zooming out from current location and then back in to the highway exit I know is prone to that at times. You know, I wanna be a good citizen and not be distracted by looking even a sec…

Does you Android Auto show a microphone for voice our the Gemini star? I think you're likely talking to the pre-Gemini voice assistant which is very basic.

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#239

Earlier quoted context omitted.

> As I work in english and use them in english, I wonder if using LLMs in a different language to code renders a different result as well. Like, if some of these benchmarks were made in other languages, would the result be similar. I am not a Chinese speaker but my understanding is that all of the models have substantially different behavior in Chinese, to the degree that it's kind of like a second model. Would be in…

Fwiw, it's been very easy to nudge Qwen 27B/35B to think in Chinese. No need for prefills like "思考:" ("Think:") or such. Relatedly, when chatting with frontier models about science education content design, I've found it very helpful to mix in Chinese education terms. In English, for example, NGSS is such a massive attractor, discussing nearby topics often yields NGSS "slop". And "estimation" (educational) in the US…

This is an interesting post, but as a non-American it's hard to understand what's it's alluding to in some parts.

The one criticism of NGSS I found from a quick skim of https://en.wikipedia.org/wiki/Next_Generation_Science_Standa... is that it teaches evolution.

What does "estimation" mean in the US?!

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#240
post #231

I don’t understand good experiences people are having with Gemini. It’s the only model that sometimes loses/forgets context in literally next message. Plus feeding unasked product links to responses.

I agree. Lot's of people really like it. For me it often just forgets all context and starts showing random slop. It's super clear as, when I ask it what happened to some element earlier in the conversation it tells me it does not have that. It might be good if it told me, but randomly lose the plot is quite frustrating.

Claude does it occasionally but it's a more a soft landing earlier context seems to be compacted, not completely lose the plot.

I just cancelled my pro subscription. I really wanted it to be good but not yet.

Post reply on HN