Earlier quoted context omitted.
I strongly agree. I suspect it's people who have not yet used the paid models from OpenAI and Anthropic. Gemini is comparable to free models from other providers, but not in the same universe as paid models. This is frustrating because when I discuss AI with laypeople they think it's still incapable of counting the number of Rs in "strawberry." They believe it to be essentially useless and incapable of basic tasks. W…
> I suspect it's people who have not yet used the paid models from OpenAI and Anthropic. Gemini is comparable to free models from other providers, but not in the same universe as paid models. I totally disagree. I pay for both ChatGPT and Anthropic (haven't tried the chinese models yet) and yet Gemini is my go-to model (and I pay for it too through Google Workspace subscriptions for several domain names tied to Googl…
Gemini 3.8 Live and 3.8 Live Extended Thinking
271–280 of 334 posts
Re: Gemini 3.8 Live and 3.8 Live Extended Thinking
#272I don’t understand good experiences people are having with Gemini. It’s the only model that sometimes loses/forgets context in literally next message. Plus feeding unasked product links to responses.
Re: Gemini 3.8 Live and 3.8 Live Extended Thinking
#273Earlier quoted context omitted.
> you're just not using the latest model, bro Pro tip, ChatGPT is the normiest of all normie websites right now. You're not part of the cognoscenti just because you learned how to type prompts into one of the most popular websites in the world. P.S. You're probably not using OpenAI models for complex or non-standard tasks. It shits the best just as often as Qwen when you need precision and detail in a non-obvious pro…
> You're probably not using OpenAI models for complex or non-standard tasks. It shits the best just as often as Qwen when you need precision and detail in a non-obvious problem. The benchmarks clearly show otherwise. This is your cue to tell me the benchmarks are made by the Illuminati and only your superior and subjective methods of evaluation are correct.
It does also seems Gemini's main problem wasn't that it was stupid, but that it was good at doing slightly different things than what I asked it to, very well. Which might well also have to do with me being better at wrangling DeepSeek's quirks than Gemini. Still, at that price tag, it's not worth it.
Re: Gemini 3.8 Live and 3.8 Live Extended Thinking
#274Earlier quoted context omitted.
> I suspect it's people who have not yet used the paid models from OpenAI and Anthropic. Gemini is comparable to free models from other providers, but not in the same universe as paid models. I totally disagree. I pay for both ChatGPT and Anthropic (haven't tried the chinese models yet) and yet Gemini is my go-to model (and I pay for it too through Google Workspace subscriptions for several domain names tied to Googl…
To be fair, it has been about six months since I tried a paid Google model. I will give it another go to compare. Hallucinations were the main issue back then but perhaps it has come a long way.
Re: Gemini 3.8 Live and 3.8 Live Extended Thinking
#275Re: Gemini 3.8 Live and 3.8 Live Extended Thinking
#276My first language is Afrikaans, which is a somewhat niche language and hard to find teachers/conversation buddies outside South Africa. (I live in USA now) I've been using Gemini to live chat in Afrikaans and do impromptu Afrikaans grammar lessons during my solo drives around town. It is phenomenal at speaking the language - like, it really shocks my family members when they hear it. This is probably the most joy I g…
Re: Gemini 3.8 Live and 3.8 Live Extended Thinking
#277Earlier quoted context omitted.
> As I work in english and use them in english, I wonder if using LLMs in a different language to code renders a different result as well. Like, if some of these benchmarks were made in other languages, would the result be similar. I am not a Chinese speaker but my understanding is that all of the models have substantially different behavior in Chinese, to the degree that it's kind of like a second model. Would be in…
Fwiw, it's been very easy to nudge Qwen 27B/35B to think in Chinese. No need for prefills like "思考:" ("Think:") or such. Relatedly, when chatting with frontier models about science education content design, I've found it very helpful to mix in Chinese education terms. In English, for example, NGSS is such a massive attractor, discussing nearby topics often yields NGSS "slop". And "estimation" (educational) in the US…
Re: Gemini 3.8 Live and 3.8 Live Extended Thinking
#278Our company's Google Workspace Business only offers 3.6 flash & thinking in the Gemini App. Has anyone else seen 3.7 or 3.8 roll out?
Re: Gemini 3.8 Live and 3.8 Live Extended Thinking
#279Earlier quoted context omitted.
I do this too! Usually some idea in ML ...when I am driving I might suddenly remember what I was thinking about and then it is my personal podcast via Gemini-in-Maps. My only complaint is if you have follow up questions, you have to be quick, otherwise it cuts off the mic.
I feel like a 10 year old all over again with my frequency of questions. How did the early Roman Empire interact with Greek city states? How are LLMs planning on learning new information on the fly without new context or retraining runs? Why did mom leave? You know, standard stuff.
Even as a human being who has a working brain and all this thinking power, I am not sure how I would answer that. It's kind of funny Gemini will still come up with something and even suggest follow-ons to the conversation but saying all that stuff as a human to another human would be a wild response to that question.
Re: Gemini 3.8 Live and 3.8 Live Extended Thinking
#280Even if really great, it still can’t use tools. The only consumer Google thing with access to MCP tools is Gemini Spark and that has lots of other problems. I wish they would combine their efforts on a great consumer product but it’s Google we’re talking about… I still dream of the day that we get full tool parity in voice and text mode so your voice assistant can do everything you connect for you. Grok and Claude ar…