Live data from Hacker News

Gemini-3.5-Transcribe

blog.google

41–50 of 140 posts

Re: Gemini-3.5-Transcribe

#41
Still no real time diarization beyond 3 people (and even then experimental) when others do it very well, like Soniox and Deepgram. For something like meeting notes this is critical. Not sure what the issue is to implement it, maybe that's not Google's use case in mind and rather it's about personal Rambling as the feature on Pixels shows, which uses this model.

Re: Gemini-3.5-Transcribe

#42

Still no real time diarization beyond 3 people (and even then experimental) when others do it very well, like Soniox and Deepgram. For something like meeting notes this is critical. Not sure what the issue is to implement it, maybe that's not Google's use case in mind and rather it's about personal Rambling as the feature on Pixels shows, which uses this model.

> maybe that's not Google's use case in mind

That would be interesting when they also own Google Meet.

Re: Gemini-3.5-Transcribe

#43
"Function calling: The model can delegate complex tasks (such as image generation and file analysis) to other Gemini models via function calls. Currently available in the Gemini macOS app."

This confused the heck out of me because it makes it sound like the STT model can make function calls in order to execute arbitrary tasks, which wouldn't make any sense. The developer docs (https://ai.google.dev/gemini-api/docs/models/gemini-3.5-tran...) confirm that the Gemini 3.5 Transcribe model cannot in fact make function calls. I guess the blog post is just using very confusing wording to describe how their consumer assistant/chatbot app can both take audio input via Gemini 3.5 Transcribe and then call other things as needed. Or maybe that bullet point was meant to be for a different model announcement and somebody made an editing error.

Re: Gemini-3.5-Transcribe

#44

"Function calling: The model can delegate complex tasks (such as image generation and file analysis) to other Gemini models via function calls. Currently available in the Gemini macOS app." This confused the heck out of me because it makes it sound like the STT model can make function calls in order to execute arbitrary tasks, which wouldn't make any sense. The developer docs ( https://ai.google.dev/gemini-api/docs/m…

Probably written using Gemini which hallucinated.

Re: Gemini-3.5-Transcribe

#45

Just added this to my benchmark site: https://multilingualsttbench.com/ It doesn't reach the frontier in either latency or accuracy for ai multilingual conversations.

Thanks for this, really helpful.

I would also like to see benchmark for translation. I'm looking for live translated subtitles so my Japanese wife can enjoy any show with out waiting months for official VOD streams to release them.

Re: Gemini-3.5-Transcribe

#46
Annoying that they don't include pricing info in new models, here it is [1]:

Gemini 3.5 Transcribe Live (Per 1M tokens in USD):

    Input: $3.50 or $0.005/min* (audio)
    Output: 21.00 or $0.004/min* (text)
Gemini 3.5 Transcribe:

    Input: $2.00 or $0.003/min* (audio)
    Output: $12.00 or $0.002/min* (text)
[1] https://ai.google.dev/gemini-api/docs/pricing#gemini-3.5-tra...

Re: Gemini-3.5-Transcribe

#47

Just added this to my benchmark site: https://multilingualsttbench.com/ It doesn't reach the frontier in either latency or accuracy for ai multilingual conversations.

I'm confused, doesn't your leaderboard clearly show it is the most accurate model? It's number one in the leaderboard. Am I missing something?

Re: Gemini-3.5-Transcribe

#48
post #44

"Function calling: The model can delegate complex tasks (such as image generation and file analysis) to other Gemini models via function calls. Currently available in the Gemini macOS app." This confused the heck out of me because it makes it sound like the STT model can make function calls in order to execute arbitrary tasks, which wouldn't make any sense. The developer docs ( https://ai.google.dev/gemini-api/docs/m…

Probably written using Gemini which hallucinated.

Pangram says human: https://www.pangram.com/history/a69f9b74-eb46-44b9-a087-7822...

Re: Gemini-3.5-Transcribe

#49
post #42

Still no real time diarization beyond 3 people (and even then experimental) when others do it very well, like Soniox and Deepgram. For something like meeting notes this is critical. Not sure what the issue is to implement it, maybe that's not Google's use case in mind and rather it's about personal Rambling as the feature on Pixels shows, which uses this model.

> maybe that's not Google's use case in mind That would be interesting when they also own Google Meet.

They have access to each individual’s stream in that case in order to diarize.

Re: Gemini-3.5-Transcribe

#50

Just added this to my benchmark site: https://multilingualsttbench.com/ It doesn't reach the frontier in either latency or accuracy for ai multilingual conversations.

I'm confused, doesn't your leaderboard clearly show it is the most accurate model? It's number one in the leaderboard. Am I missing something?

Figured it out. 3.5 Flash and 3.5 Transcribe are different models
Post reply on HN