Live data from Hacker News

Gemini 3.8 Live and 3.8 Live Extended Thinking

blog.google

131–140 of 334 posts

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#131
post #50

Gemini's Live Mode is already much better than GPT Voice in my personal experience, even though it was much dumber. It really does feel like talking to a real person. ChatGPT keeps humming to whatever I say and has some weird voices. Excited to try this out! Shame on Google for not releasing Gemini 3.8 for Google AI Plus users yet, though.

OpenAI just released the new full duplex mode to the API as gpt-live-1 or something like that. Very realistic.

Agreed. I've been using GPT-Live-1 this week, with Claude as the backend brain. It's amazing, feels like working with Jarvis. It certainly made me feel there's no point in human telephone support now - but I'm sure I'd find edge cases if that really was something I wanted to build out myself.

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#132
post #94

Earlier quoted context omitted.

For heavyweight work I have been using Astra, but for rabbit holes and brain storming Gemini is far more enjoyable to interact with. I'm worried in their push to catch up on the SOTA front, it's going to lose that natural sounding touch it currently has.

Is Google still chasing frontier? Seems like they haven't had a "Pro" model in forever. I think a good niche for them would be right where they are now.

They are. They were supposed to release 3.5 Pro over the summer but haven't because of persistent architectural/technical issues allegedly. Which is better than releasing it in that state imo.

Their AI leadership team has taken some hits recently too, in the form of departures. I believe when they get their bearings they will be competitive again. 3.8 Flash has been a great model for me.

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#133
post #57
post #41

Earlier quoted context omitted.

Wondering if people have managed to have Gemini in-front of other models like claude/codex models and only interact with that. Having Gemini act as a pure human/llm translator.

Not in the principled sense you mean but I have in fact recently started having Gemini explain to me what Claude is talking to me about, lol.

BTW, as kind of a follow-up to this, I think the most important finding to report--for those who only use one model, as many people seem to--is just how much more often Claude seems to be extremely confidently (and even insufferably) wrong than any of the other three big models (all of which I use quite often... yet I only pull out Claude when I've given up hope in a problem and are looking for out-of-the-box brainstorming).

And like, it does this despite it speaking in extremely dense math, which both makes it sound correct and requires a lot more effort to prove when it is wrong... yet, it isn't actually correct more often, and so that time sink just isn't worth the benefit. I then think many people--including people who can speak math (as can I)--just stop bothering to check everything, as if you come across a human who speaks like this it probably does correlate with slow and careful thought that helps prevent errors.

Instead, Claude has the mistake rate of a somewhat accelerated beginner impossibly combined with the language of an expert professor; and we as humans just aren't good at that combination: it becomes very dangerous and makes it take longer to spot its egregious mistakes and trained-in biases. If you have to use Claude, I thereby claim you really need to have a team of not-Claudes to help insulate you from this, and Gemini (while being a bit senile) is a lot more collaborative and approaches problems in ways that makes it harder to get tricked.

(To translate this into more of an engineering analogy: Claude always feels to me like the engineer who put more effort into learning how to program in functional languages than into how to actually develop working code, and then confidently presents you answers in Haskell or Lisp that never quite work. To find their errors is then very costly. In contrast, Gemini feels more like a Java or Go developer who knows they are a cog... that's helpful! <- Which maybe just goes to show that AI has finally turned me into a manager, omg.)

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#134

Earlier quoted context omitted.

I’m curious how you guys keep track of each model’s coding capabilities. The landscape keeps changing. I don’t suppose you benchmark all frontier models every other month, right?

I use them. Daily. Gemini hasn’t been a contender by comparison for a long time.

I wonder if it's a harness thing or a model thing at this point. I feel all coding models are quite capable for most tasks I want them to do.

Most of the time I don't need what the bench tests and I'm not really giving them completely ambiguous tasks without any refinement.

I only find marginal differences between models at this point and it almost feels like personality quirks in each model than anything.

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#135

Gemini is underrated in that it produces the only prose that is somewhat bearable to read.

Try Gemini live in a multi lingual environment. It can pick out speakers and live translate to you. Truly underrated for its capabilities.

Not the same, but related: Gemini is great for querying text in another language, it provides really cogent, useful responses with just enough source language quotes to be able to reference the source text effectively.

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#136
post #9

I wonder when/if we’ll see Gemini beating Fable and Astra. Last year I would have confidently bet Google will overtake the others just because they have the data, the hardware (TPUs) and a fat advertising money pipe and yet they are still behind. Anyone anonymous at Google want to hint when Gemini 4 will be out?

Anecdata, but I've been using gemini personally instead of what I used regular google searches for the better part. Even in car chat/search and some lighter and not so light research (if it's not heavy on technicals). It's good enough that I use it nonstop like that. I wouldn't trust it for coding at all - switching between fable and now astra.

Google in a sense won (me over) like that. I also expected them to brute force their way into everything and dominate. This is how it played out though. Image generation is great as well, but ChatGPT one is more lenient on copyright and nannying - for example when my kid asks me to "take a photo of him and Sonic". Gemini cops out either because of the kid or Sonic, disappointing us both, but ChatGPT can be.. persuaded.

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#137

My first language is Afrikaans, which is a somewhat niche language and hard to find teachers/conversation buddies outside South Africa. (I live in USA now) I've been using Gemini to live chat in Afrikaans and do impromptu Afrikaans grammar lessons during my solo drives around town. It is phenomenal at speaking the language - like, it really shocks my family members when they hear it. This is probably the most joy I g…

for me the live mode in androids google translate app has been as close as it gets to perfect for traveling cannot believe it is a free service after trying so many others

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#138
post #134

Earlier quoted context omitted.

I use them. Daily. Gemini hasn’t been a contender by comparison for a long time.

I wonder if it's a harness thing or a model thing at this point. I feel all coding models are quite capable for most tasks I want them to do. Most of the time I don't need what the bench tests and I'm not really giving them completely ambiguous tasks without any refinement. I only find marginal differences between models at this point and it almost feels like personality quirks in each model than anything.

When comparing OpenAI and Claude thats pretty much true, but not Gemini... And have you tried Antigravity? Yikes

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#139

Earlier quoted context omitted.

Mostly because it answers quickly and is more agreeable (too agreeable at times). Meanwhile Claude and Astra like to couch all their agreements with caveats and provisos.

"couch all their agreements with caveats and provisos." When you're a ChatGPT Projects or Claude Projects user, those caveats and provisos are your worst enemy because they'll change caveats into hard rules (either for the session or committed to memories) and you end up in absolute hell having to make it investigate to figure out why it can no longer produce anything but read-only pre-check code that never actually…

Yep the only way out is hooks to forbid what can be detected by ast and second model to prune comments, flatten pyramids of fallback, and squash the test suite removing quirks maintaining wanted behaviors.

Re: Gemini 3.8 Live and 3.8 Live Extended Thinking

#140

Not a great impression to have your demo video demonstrate how one of your 'most advanced' AI models loses to the most common check-mate pattern in all of chess.

Playing a legal game of chess without using a guided decoding technique is a massive achievement. Ref https://aclanthology.org/2025.mathnlp-main.11/.
Post reply on HN