Live data from Hacker News

Gemini AI

deepmind.google

401–410 of 1001 posts

Re: Gemini AI

#401
Hmmm.. Seems like summarizing/extracting information from Youtube videos is a place where Bard/Gemini should shine.

I asked it to give me "the best quotes from..." a person appearing in the video (they are explicitly introduced) and Bard says,

"Unfortunately, I don't have enough information to process your request."

Re: Gemini AI

#402
post #376
post #324

Earlier quoted context omitted.

The table is *highly* misleading. It uses different methodologies all over the place. For MMLU, it highlights the CoT @ 32 result, where Ultra beats GPT4, but it loses to GPT4 with 5-shot, for example. For GSM8K it uses Maj1@32 for Ultra and 5-shot CoT for GPT4, etc. Then also, for some reason, it uses different metrics for Ultra and Pro, making them hard to compare. What a mess of a "paper".

It really feels like the reason this is being released now and not months ago is that that's how long it took them to figure out the convoluted combination of different evaluation procedures to beat GPT-4 on the various benchmarks.

And somehow, when reading the benchmarks, Gemini Pro seems to be a regression compared to PaLM 2-L (the current Bard) :|

Re: Gemini AI

#403
post #200

Earlier quoted context omitted.

To add to my comment above: Google DeepMind put out 16 videos about Gemini today, the total watch time at 1x speed is about 45 mins. I've now watched them all (at >1x speed). In my opinion, the best ones are: * https://www.youtube.com/watch?v=UIZAiXYceBI - variety of video/sight capabilities * https://www.youtube.com/watch?v=JPwU1FNhMOA - understanding direction of light and plants * https://www.youtube.com/watch?v=D…

Watching these videos made me remember this cool demo Google did years ago where their earpods would auto translate in realtime a conversation between two people talking different languages. Turned out to be demo vaporware. Will this be the same thing?

Aren't you talking about this? https://support.google.com/googlepixelbuds/answer/7573100?hl... (which exists?)

Re: Gemini AI

#404

Gemini Nano sounds like the most exciting part IMO. IIRC Several people in the recent Pixel 8 thread were saying that offloading to web APIs for functions like Magic Eraser was only temporary and could be replaced by on-device models at some point. Looks like this is the beginning of that.

> "Using the power of Google Tensor G3, Video Boost on Pixel 8 Pro uploads your videos to the cloud where our computational photography models adjust color, lighting, stabilization and graininess."*

I wonder why the power of Tensor G3 is needed to upload your video to the cloud...

*https://blog.google/products/pixel/pixel-feature-drop-decemb...

Re: Gemini AI

#405
post #71

There's a huge amount of criticism for Sundar on Hacker News (seemingly from Googlers, ex-Googlers, and non-Googlers), but I give huge credit for Google's "code red" response to ChatGPT. I count at least 19 blog posts and YouTube videos from Google relating to the Gemini update today. While Google hasn't defeated (whatever that would mean) OpenAI yet, the way that every team/product has responded to improve, publiciz…

Quite literally almost all the criticism of Sundar is that he is ALL narrative and very little delivery. You illustrated that further... lots of narrative around GPT3.5 equivalent launch and maybe 4 in the future.

Re: Gemini AI

#406
post #376
post #324

Earlier quoted context omitted.

The table is *highly* misleading. It uses different methodologies all over the place. For MMLU, it highlights the CoT @ 32 result, where Ultra beats GPT4, but it loses to GPT4 with 5-shot, for example. For GSM8K it uses Maj1@32 for Ultra and 5-shot CoT for GPT4, etc. Then also, for some reason, it uses different metrics for Ultra and Pro, making them hard to compare. What a mess of a "paper".

It really feels like the reason this is being released now and not months ago is that that's how long it took them to figure out the convoluted combination of different evaluation procedures to beat GPT-4 on the various benchmarks.

"Dearest LLM: Given the following raw benchmark metrics, please compose an HTML table that cherry-picks and highlights the most favorable result in each major benchmark category"

Re: Gemini AI

#407
Looking forward to the API. I wonder if they will have something like OpenAI's function calling, which I've found to be incredibly useful and quite magical really. I haven't tried other Google AI APIs however, so maybe they already have this (but I haven't heard about it...)

Also interesting is the developer ecosystem OpenAI has been fostering vs Google. Google has been so focused on user-facing products with AI embedded (obviously their strategy) but I wonder if this more-closed approach will lose them the developer mindshare for good.

Re: Gemini AI

#408
post #368
post #256

This announcement makes we wonder if we are approaching a plateau in these systems. They are essentially claiming close to parity with gpt-4, not a spectacular new breakthrough. If I had something significantly better in the works, I'd either release it or hold my fire until it was ready. I wouldn't let openai drive my decision making, which is what this looks like from my perspective. Their top line claim is they ar…

Don't look at absolute number, instead think of it in terms of relative improvement. DocVQA is a benchmark with a very strong SOTA. GPT-4 achieves 88.4, Gemini 90.9. It's only 2.5% increase, but a ~22% error reduction which is massive for real-life usecases where the error tolerance is lower.

This + some benchmarks are shitty thus rational model should be allowed to not answer them but ask claryfying questions.

Re: Gemini AI

#409

Earlier quoted context omitted.

formatted nicely: Dataset | Gemini Ultra | Gemini Pro | GPT-4 MMLU | 90 | 79 | 87 BIG-Bench-Hard | 84 | 75 | 83 HellaSwag | 88 | 85 | 95 Natural2Code | 75 | 70 | 74 WMT23 | 74 | 72 | 74

I realize that this is essentially a ridiculous question, but has anyone offered a qualitative evaluation of these benchmarks? Like, I feel that GPT-4 (pre-turbo) was an extremely powerful model for almost anything I wanted help with. Whereas I feel like Bard is not great. So does this mean that my experience aligns with "HellaSwag"?

I get what you mean, but what would such "qualitative evaluation" look like?

Re: Gemini AI

#410

Gemini Ultra isn't released yet and is months away still. Bard w/ Gemini Pro isn't available in Europe and isn't multi-modal, https://support.google.com/bard/answer/14294096 No public stats on Gemini Pro. (I'm wrong. Pro stats not on website, but tucked in a paper - https://storage.googleapis.com/deepmind-media/gemini/gemini_... ) I feel this is overstated hype. There is no competitor to GPT-4 being released today. I…

Not just Europe: also no Canada, China, Russia, United Kingdom, Switzerland, Bulgaria, Norway, Iceland, etc.
Post reply on HN