Live data from Hacker News

Gemini 2.0 is now available to everyone

blog.google

61–70 of 285 posts

Re: Gemini 2.0 is now available to everyone

#61
I've been very impressed by Gemini 2.0 Flash for multimodal tasks, including object detection and localization[1], plus document tasks. But the 15 requests per minute limit was a severe limiter while it was experimental. I'm really excited to be able to actually _do_ things with the model.

In my experience, I'd reach for Gemini 2.0 Flash over 4o in a lot of multimodal/document use cases. Especially given the differences in price ($0.10/million input and $0.40/million output versus $2.50/million input and $10.00/million output).

That being said, Qwen2.5 VL 72B and 7B seem even better at document image tasks and localization.

[1] https://notes.penpusher.app/Misc/Google+Gemini+101+-+Object+...

Re: Gemini 2.0 is now available to everyone

#62

Earlier quoted context omitted.

Hear, hear! There has to be a better way about it. As I see it, to be productive, AI agents have to be able to talk about politics, because at the end of the day politics are everywhere. So following up on what they do already, they'll have to define a model's political stance (whatever it is), and to have it hold its ground, voicing an opinion or abstaining from voicing an opinion, but continuing the conversation, a…

Indeed, you can facilitate talking politics without having a set opinion. It's a fine line, but it is something the BBC managed to do for a very long time. The BBC does not itself present an opinion on Politics yet facilitates political discussion through shows like Newsnight and The Daily Politics (rip).

BBC is great at talking about the Gaza situation. Makes it seem like people are just dying from natural causes all the time.

Re: Gemini 2.0 is now available to everyone

#65

Pricing is CRAZY. Audio input is $0.70 per million tokens on 2.0 Flash, $0.075 for 2.0 Flash-Lite and 1.5 Flash. For gpt-4o-mini-audio-preview, it's $10 per million tokens of audio input.

Sadly: "Gemini can only infer responses to English-language speech."

https://ai.google.dev/gemini-api/docs/audio?lang=rest#techni...

Re: Gemini 2.0 is now available to everyone

#66
post #12
post #6

Anyone have a take on how the coding performance (quality and speed) of the 2.0 Pro Experimental compares to o3-mini-high? The 2 million token window sure feels exciting.

I don't know what those "needle in haystack" benchmarks are testing for because in my experience dumping a big amount of code in the context is not working as you'd expect. It works better if you keep the context small

Claude works well for me loading code up to around 80% of its 200K context and then asking for changes. If the whole project can't fit I try to at least get in headers and then the most relevant files. It doesn't seem to degrade. If you are using something like an AI IDE a lot of times they don't really get the 200K context.

Re: Gemini 2.0 is now available to everyone

#67
It’s funny, I’ve never actually used Gemini and, though this may be incorrect, I automatically assume it’s awful. I assume it’s awful because the AI summaries at the top of Google Search are so awful, and that’s made me never give Google AI a chance.

Re: Gemini 2.0 is now available to everyone

#68
post #10
post #8

> available via the Gemini API in Google AI Studio and Vertex AI. > Gemini 2.0, 2.0 Pro and 2.0 Pro Experimental, Gemini 2.0 Flash, Gemini 2.0 Flash Lite 3 different ways of accessing the API, more than 5 different but extremely similarly named models. Benchmarks only comparing to their own models. Can't be more "Googley"!

I think this is a good summary: https://storage.googleapis.com/gweb-developer-goog-blog-asse...

- Experimental™

- Preview™

- Coming soon™

Post reply on HN