Live data from Hacker News

Gemini 2.0 is now available to everyone

blog.google

81–90 of 285 posts

Re: Gemini 2.0 is now available to everyone

#81
post #8

> available via the Gemini API in Google AI Studio and Vertex AI. > Gemini 2.0, 2.0 Pro and 2.0 Pro Experimental, Gemini 2.0 Flash, Gemini 2.0 Flash Lite 3 different ways of accessing the API, more than 5 different but extremely similarly named models. Benchmarks only comparing to their own models. Can't be more "Googley"!

Honestly naming conventions in the AI world have been appalling regardless of the company

Google is the least confusing to me. Old school version number and Pro is better than Flash which is fast and for "simple" stuff (which can be effortless intermediate level coding at this point).

OpenAI is crazy. There may be a day when we might have o5 that is reasoning and 5o that is not, and where they belong to different generations too, snd where "o" meant "Omni" despite o1-o3 not being audiovisual anymore like 4o.

Anthropic crazy too. Sonnets and Haikus, just why... and a 3.5 Sonnet that was released in October that was better than 3.5 Sonnet. (Not a typo) And no one knows why there never was a 3.5 Opus.

Re: Gemini 2.0 is now available to everyone

#82

That 1M tokens context window alone is going to kill a lot of RAG use cases. Crazy to see how we went from 4K tokens context windows (2023 ChatGPT-3.5) to 1M in less than 2 years.

Maybe someone knows, what's the usual recommendation regarding big context windows? Is it safe to use it to the max, or performance will degrade and we should adapt the maximum to our use case?

Re: Gemini 2.0 is now available to everyone

#83

Pricing is CRAZY. Audio input is $0.70 per million tokens on 2.0 Flash, $0.075 for 2.0 Flash-Lite and 1.5 Flash. For gpt-4o-mini-audio-preview, it's $10 per million tokens of audio input.

The increase is likely because 1.5 Flash was actually cheaper than all other STT services. I wrote about this a while ago at https://ktibow.github.io/blog/geminiaudio/.

Re: Gemini 2.0 is now available to everyone

#84
post #80

I upgraded my llm-gemini plugin to handle this, and shared the results of my "Generate an SVG of a pelican riding a bicycle" benchmark here: https://simonwillison.net/2025/Feb/5/gemini-2/ The pricing is interesting: Gemini 2.0 Flash-Lite is 7.5c/million input tokens and 30c/million output tokens - half the price of OpenAI's GPT-4o mini (15c/60c). Gemini 2.0 Flash isn't much more: 10c/million for text/image input, 70c…

Is there a way to see/compare the shared results for all of the LLMs you've tested this prompt on in one place? The 2.0 pro result seems decent but I don't have a baseline if that's because it is or if the other 2 are just "extremely bad" or something.

Re: Gemini 2.0 is now available to everyone

#85

That 1M tokens context window alone is going to kill a lot of RAG use cases. Crazy to see how we went from 4K tokens context windows (2023 ChatGPT-3.5) to 1M in less than 2 years.

We have heard this before when 100k and 200k were first being normalized by Anthropic way back when and I tend to be skeptical in general when it comes to such predictions, but in this case, I have to agree.

Having used the previews for the last few weeks with different tasks and personally designed challenges, what I found is that these models are not only capable of processing larger context windows on paper, but are also far better at actually handling long, dense, complex documents in full. Referencing back to something upon specific request, doing extensive rewrites in full whilst handling previous context, etc. These models also have handled my private needle in haystack-type challenges without issues as of yet, though those have been limited to roughly 200k in fairness. Neither Anthropics, OpenAIs, Deepseeks or previous Google models handled even 75k+ in any comparable manner.

Cost will of course remain a factor and will keep RAG a viable choice for a while, but for the first time I am tempted to agree that someone has delivered a solution which showcases that a larger context window can in many cases work reliably and far more seemlessly.

Is also the first time a Google model actually surprised me (positively), neither Bard, nor AI answers or any previous Gemini model had any appeal to me, even when testing specificially for what other claimed to be strenghts (such as Gemini 1.5s alleged Flutter expertise which got beaten by both OpenAI and Anthropics equivalent at the time).

Re: Gemini 2.0 is now available to everyone

#86
For anyone that parsing PDF's this is a game changer in term of price per dollar - I wrote a blog about it [1]. I think a lot of people were nervous about pricing since they released the beta, and although it's slightly more expensive than 1.5 Flash, this is still incredibly cost-effective. Looking forward to also benchmarking the lite version.

[1] https://www.sergey.fyi/articles/gemini-flash-2

Re: Gemini 2.0 is now available to everyone

#87
post #61

I've been very impressed by Gemini 2.0 Flash for multimodal tasks, including object detection and localization[1], plus document tasks. But the 15 requests per minute limit was a severe limiter while it was experimental. I'm really excited to be able to actually _do_ things with the model. In my experience, I'd reach for Gemini 2.0 Flash over 4o in a lot of multimodal/document use cases. Especially given the differen…

> In my experience, I'd reach for Gemini 2.0 Flash over 4o Why not use o1-mini?

Mostly because OpenAI's vision offerings aren't particularly compelling:

- 4o can't really do localization, and ime is worse than Gemini 2.0 and Qwen2.5 at document tasks

- 4o mini isn't cheaper than 4o for images because it uses a lot of tokens per image compared to 4o (~5600/tile vs 170/tile, where each tile is 512x512)

- o1 has support for vision but is wildly expensive and slow

- o3-mini doesn't yet have support for vision, and o1-mini never did

Re: Gemini 2.0 is now available to everyone

#88

I tried voice chat. It's very good, except for the politics We started talking about my plans for the day, and I said I was making chili. G asked if I have a recipe or if I needed one. I said, I started with Obama's recipe many years ago and have worked on it from there. G gave me a form response that it can't talk politics. Oh, I'm not talking politics, I'm talking chili. G then repeated form response and tried to c…

I find it horrifying and dystopian that the part where it "Can't talk politics" is just accepted and your complaint is that it interrupts your ability to talk chilli. "Go back to bed America." "You are free, to do as we tell you" https://youtu.be/TNPeYflsMdg?t=143

Online the idea of "no politics" is often used as a way to try to stifle / silence discussion too. It's disturbingly fitting to the Gemini example.

I was a part of a nice small forum online. Most posts were everyday life posts / personal. The person who ran it seemed well meaning. Then a "no politics" rule appeared. It was fine for a while. I understood what they meant and even I only want so much outrage in my small forums.

Yet one person posted about how their plans to adopt were in jeopardy over their state's new rules about who could adopt what child. This was a deeply important and personal topic for that individual.

As you can guess the "no politics" rule put a stop to that. The folks who supported laws like were being proposed of course thought that they shouldn't discuss it because it is "politics", others felt that this was that individual talking about their rights and life, it wasn't "just politics". Whole forum fell apart after that debacle.

Gemini's response here is sadly fitting internet discourse... in bad way.

Re: Gemini 2.0 is now available to everyone

#90
post #4

Earlier quoted context omitted.

Pixels replaced Assistent w/ Gemini a while back and it was horrendous; would answer questions but not perform the basic tasks you actually used Assistant for (setting timer, navigating, home control, etc). Seems like they're approaching parity (finally) months and months later (alarms/tv control work at least now), but losing basic oft-used functionality is a serious fumble.

Thanks, didn't know, never really used these voice assistants. It's a weird choice, I suppose the endless handcrafted rules and tools don't scale across languages and usecases but then LLM are not good at reliability. And what's the point of using assistant that will not do the task reliably, if you have to double-check you are better of not using it...

The issue wasn't inconsistency it was "had no home integration at all" at launch. They rushed to roll out the 'new' assistant and didn't bother waiting for the basic feature set first.

Today; it works ~perfectly for TV control/Alarm setting - I can't think of it not working first try in the last month or so for me. Maybe more consistent than prior?

The rollout was simply borked from the PM/Decision making side.

Post reply on HN