Live data from Hacker News

Gemini-3.5-Transcribe

blog.google

81–90 of 140 posts

Re: Gemini-3.5-Transcribe

#81

I’ve tested at least 20 STT models in a benchmark I’ve set up with German, Italian and English voices from meetings in my company. The voices contained very industry specific words, the languages changed from one sentence to another, sometimes words in a language were mentioned while a discussion was in another. The only local model that satisfies me is Voxtral Mini 3b, the only paid API that is slightly better is el…

Do you use it on a desktop? Mac by any chance? What's your setup? I've been looking to find a simple and fast dictation app for English but almost everything I've tried (from Handy to many apps, eg some with Whisper in their names, after the model I assume) just don't work well. Apple's offering is worse than those though. I even tried with local enhancement models.

Out of curiosity, what was wrong with Handy? I use it and it works fine.

Re: Gemini-3.5-Transcribe

#82
post #72

I’ve tested at least 20 STT models in a benchmark I’ve set up with German, Italian and English voices from meetings in my company. The voices contained very industry specific words, the languages changed from one sentence to another, sometimes words in a language were mentioned while a discussion was in another. The only local model that satisfies me is Voxtral Mini 3b, the only paid API that is slightly better is el…

What i want is a model that outputs its predictions and their scores along with the text it choose. So I could flag something it's getting consistently wrong, like mis-predicting a technical term or name, or acronym, and say replace it with my correction, and have the ui be able to smartly replace that in the whole text so far and future parts. even better would be the ability to feed this back into the model for fut…

Good news: most of these models can include a prompt that steers the transcription; if you use frequently a word you just invented, add it there and it will be transcribed correctly more likely.

Re: Gemini-3.5-Transcribe

#83
post #58

I personally tested all the STT models for my real-time translator ( https://fliptalk.ai ). From language detection and accuracy in a noisy environment to the most important point: latency. At the moment, Soniox STT v5 is definitely the best, and I'm impressed by its performance. It's good that Google released Gemini-3.5-Transcribe, and it beats every other model on accuracy, but it definitely needs a bit more work o…

Soniox website has a live comparison demo: https://soniox.com/compare-stt For me, Gemini-3.5-Transcribe actually has slightly lower latency. Kudos to Soniox for both paying their competitor and letting them win. But yes, Soniox is much cheaper.

Oh, thanks for pointing me to Soniox. It is really good. Would also pick up the words with different languages, identify and output in the right language. Looks interesting. It was much faster too, but that I cannot say much since it was on their own website.

Re: Gemini-3.5-Transcribe

#84
post #73

I’ve tested at least 20 STT models in a benchmark I’ve set up with German, Italian and English voices from meetings in my company. The voices contained very industry specific words, the languages changed from one sentence to another, sometimes words in a language were mentioned while a discussion was in another. The only local model that satisfies me is Voxtral Mini 3b, the only paid API that is slightly better is el…

It's interesting that the results can be so different depending on the person, the use case and even the microphone used. I use dictation a lot, so I try to stay up to date with the latest models as much as I can. So far, for my needs, nothing could beat Whisper Large v3. I keep hearing that Nvidia Parakeet models are better, but they just don't work as well for me, even though they are unquestionably faster. Things…

I wonder if cross training models specifically on people who combine languages (like Spanglish) would help with this. Surely must be patterns in what words people choose to use in each language

Re: Gemini-3.5-Transcribe

#85
> Custom vocabulary: Recognizes specialized jargon and unique spellings by seamlessly adapting transcriptions to your provided custom vocabulary.

Is there a model that works on the syllabic level? I want to be able to say any word and have it reconstruct however that word would be spelled. I know English does not exactly work this way, so a custom vocabulary would still be nice, but I don't want to rely on having every single word that could ever exist in a vocabulary first.

Re: Gemini-3.5-Transcribe

#86
post #44

Earlier quoted context omitted.

Probably written using Gemini which hallucinated.

Pangram says human: https://www.pangram.com/history/a69f9b74-eb46-44b9-a087-7822...

Checking the following with another AI checker and it says 100% AI generated. (https://originality.ai/). We should run each check on multiple checkers if you want to provide a substantive claim that something is AI-written or not. This habit of just putting text in a "AI text checker" and then treating whatever it says as the truth is absolutely one of the worst things to emerge in recent times. You should check out the CEO of Pangram who uses this to go on witch hunts on X and even though the app claims "our results do not reflect reality 100% and can be false", he always uses his stupid checks as "LOOK THERE IS PROOF YOU ARE AN AI WRITER" and he points it at legit journalists etc. This whole company is honestly such a piece of shit that I as a human writer think is going to kill art more than AI writing does. I just hope people see it for the nonsense it is sooner than later.

---

Experience smart transcription and advanced dictation

In addition to the Gemini API in the Google AI Studio and Gemini Enterprise Agent Platform, 3.5 Transcribe goes further than standard speech-to-text to make working across Google feel more natural and intuitive. By bringing context-aware understanding directly into everyday surfaces like Gboard, Antigravity, the Gemini app, and Chrome, it captures nuances, intent, and inline edits with ease.

On Gboard on Android, through the new Rambler feature, 3.5 Transcribe transforms spoken thoughts into well-formatted text, filtering out filler words. You can also use your voice to make edits, correct misspellings, and change the writing style.

On Google Antigravity, 3.5 Transcribe pairs screen context and chat history, with your permission, to ensure pinpoint transcription accuracy across file names, agent thoughts, and active documents.

In Google AI Studio, you can access 3.5 Transcribe in Build mode to vibe code apps with your voice on the fly.

In the Gemini app on macOS, 3.5 Transcribe not only transcribes your free natural speech into clean formatted text, but also enables voice commands that can pair seamlessly with screen context to power complex workflows. By calling on other Gemini models in the background to handle the heavy lifting, the model makes it effortless to summarize local files, repurpose text across apps, or generate images right at your cursor—using just your voice.

Coming soon to Chrome, you’ll be able to talk to type in any web field — making it effortless to dictate replies, draft posts, or prompt Gemini in Chrome more naturally and easily with your voice.

Re: Gemini-3.5-Transcribe

#87
post #52

Earlier quoted context omitted.

Maybe I missed something, but isn’t it impossible to detect whether something is written by AI or not? A human on a bad day can write like AI while an AI on a good day can write like a human.

No, that's not correct for any reasonable definition of "impossible." Look up pangram's accuracy ratings. It's not perfect, but it's pretty good. LLMs in fact leave very distinguishing traces of their logit distributions in the text they write. It's one of the reasons why it's so easy for humans to also smell them. It is possible to trick pangram - they bias toward a low false positive and a higher false negative - b…

> In preliminary testing, Mantzarlis found Pangram was more likely to misclassify AI-generated text as human-authored when it rhymed, repeated itself, and when it used archaic language. He then built an adversarial set of 588 AI-generated text samples tailored to these weaknesses. When he used Pangram to evaluate them, the tool falsely labelled AI text as human 86% of the time.

> “I don't think that Pangram is bad,” Mantzarlis said. “I think actually Pangram at scale is probably a pretty solid tool. That said, I am extremely worried about it being used in individual cases.”

https://reutersinstitute.politics.ox.ac.uk/news/human-wrote-...

Using it for an individual article to fully determine if its AI or not is "impossible" because you're not even using the tool properly.

Re: Gemini-3.5-Transcribe

#88
post #26

I've been testing it on Pixel 11 Pro and I mostly dislike it. It is convenient when you have something long to say without thinking about it first. But the main issue is when you want to say something precise with specific wording it might "simplify" it and break the meaning. Something like "I hesitated to check it, I should have verified" => "I should have verified" (The "I hesitated..." is removed but I said it bec…

Are you in Smart or Verbatim mode? https://ai.google.dev/gemini-api/docs/transcribe#transcripti...

Looking at the description I was in Smart mode. I didn't know the model supports both

Re: Gemini-3.5-Transcribe

#89

I’ve tested at least 20 STT models in a benchmark I’ve set up with German, Italian and English voices from meetings in my company. The voices contained very industry specific words, the languages changed from one sentence to another, sometimes words in a language were mentioned while a discussion was in another. The only local model that satisfies me is Voxtral Mini 3b, the only paid API that is slightly better is el…

[dead]

Re: Gemini-3.5-Transcribe

#90

Earlier quoted context omitted.

No, that's not correct for any reasonable definition of "impossible." Look up pangram's accuracy ratings. It's not perfect, but it's pretty good. LLMs in fact leave very distinguishing traces of their logit distributions in the text they write. It's one of the reasons why it's so easy for humans to also smell them. It is possible to trick pangram - they bias toward a low false positive and a higher false negative - b…

> In preliminary testing, Mantzarlis found Pangram was more likely to misclassify AI-generated text as human-authored when it rhymed, repeated itself, and when it used archaic language. He then built an adversarial set of 588 AI-generated text samples tailored to these weaknesses. When he used Pangram to evaluate them, the tool falsely labelled AI text as human 86% of the time. > “I don't think that Pangram is bad,”…

Why? Does TFA rhyme, repeat itself, use archaic language or otherwise looks adversarial? It doesn't, so we can assume Pangram's usual false positive/negative rates apply. Yeah there's a chance it's wrong, and they'll need to catch up to new models, but my instinct can be wrong too and I don't stop using it to filter what I read; at least Pangram's accuracy can be measured.
Post reply on HN