I asked it to give me "the best quotes from..." a person appearing in the video (they are explicitly introduced) and Bard says,
"Unfortunately, I don't have enough information to process your request."
401–410 of 1001 posts
I asked it to give me "the best quotes from..." a person appearing in the video (they are explicitly introduced) and Bard says,
"Unfortunately, I don't have enough information to process your request."
Earlier quoted context omitted.
The table is *highly* misleading. It uses different methodologies all over the place. For MMLU, it highlights the CoT @ 32 result, where Ultra beats GPT4, but it loses to GPT4 with 5-shot, for example. For GSM8K it uses Maj1@32 for Ultra and 5-shot CoT for GPT4, etc. Then also, for some reason, it uses different metrics for Ultra and Pro, making them hard to compare. What a mess of a "paper".
It really feels like the reason this is being released now and not months ago is that that's how long it took them to figure out the convoluted combination of different evaluation procedures to beat GPT-4 on the various benchmarks.
Earlier quoted context omitted.
To add to my comment above: Google DeepMind put out 16 videos about Gemini today, the total watch time at 1x speed is about 45 mins. I've now watched them all (at >1x speed). In my opinion, the best ones are: * https://www.youtube.com/watch?v=UIZAiXYceBI - variety of video/sight capabilities * https://www.youtube.com/watch?v=JPwU1FNhMOA - understanding direction of light and plants * https://www.youtube.com/watch?v=D…
Watching these videos made me remember this cool demo Google did years ago where their earpods would auto translate in realtime a conversation between two people talking different languages. Turned out to be demo vaporware. Will this be the same thing?
Gemini Nano sounds like the most exciting part IMO. IIRC Several people in the recent Pixel 8 thread were saying that offloading to web APIs for functions like Magic Eraser was only temporary and could be replaced by on-device models at some point. Looks like this is the beginning of that.
I wonder why the power of Tensor G3 is needed to upload your video to the cloud...
*https://blog.google/products/pixel/pixel-feature-drop-decemb...
There's a huge amount of criticism for Sundar on Hacker News (seemingly from Googlers, ex-Googlers, and non-Googlers), but I give huge credit for Google's "code red" response to ChatGPT. I count at least 19 blog posts and YouTube videos from Google relating to the Gemini update today. While Google hasn't defeated (whatever that would mean) OpenAI yet, the way that every team/product has responded to improve, publiciz…
Earlier quoted context omitted.
The table is *highly* misleading. It uses different methodologies all over the place. For MMLU, it highlights the CoT @ 32 result, where Ultra beats GPT4, but it loses to GPT4 with 5-shot, for example. For GSM8K it uses Maj1@32 for Ultra and 5-shot CoT for GPT4, etc. Then also, for some reason, it uses different metrics for Ultra and Pro, making them hard to compare. What a mess of a "paper".
It really feels like the reason this is being released now and not months ago is that that's how long it took them to figure out the convoluted combination of different evaluation procedures to beat GPT-4 on the various benchmarks.
Also interesting is the developer ecosystem OpenAI has been fostering vs Google. Google has been so focused on user-facing products with AI embedded (obviously their strategy) but I wonder if this more-closed approach will lose them the developer mindshare for good.
This announcement makes we wonder if we are approaching a plateau in these systems. They are essentially claiming close to parity with gpt-4, not a spectacular new breakthrough. If I had something significantly better in the works, I'd either release it or hold my fire until it was ready. I wouldn't let openai drive my decision making, which is what this looks like from my perspective. Their top line claim is they ar…
Don't look at absolute number, instead think of it in terms of relative improvement. DocVQA is a benchmark with a very strong SOTA. GPT-4 achieves 88.4, Gemini 90.9. It's only 2.5% increase, but a ~22% error reduction which is massive for real-life usecases where the error tolerance is lower.
Earlier quoted context omitted.
formatted nicely: Dataset | Gemini Ultra | Gemini Pro | GPT-4 MMLU | 90 | 79 | 87 BIG-Bench-Hard | 84 | 75 | 83 HellaSwag | 88 | 85 | 95 Natural2Code | 75 | 70 | 74 WMT23 | 74 | 72 | 74
I realize that this is essentially a ridiculous question, but has anyone offered a qualitative evaluation of these benchmarks? Like, I feel that GPT-4 (pre-turbo) was an extremely powerful model for almost anything I wanted help with. Whereas I feel like Bard is not great. So does this mean that my experience aligns with "HellaSwag"?
Gemini Ultra isn't released yet and is months away still. Bard w/ Gemini Pro isn't available in Europe and isn't multi-modal, https://support.google.com/bard/answer/14294096 No public stats on Gemini Pro. (I'm wrong. Pro stats not on website, but tucked in a paper - https://storage.googleapis.com/deepmind-media/gemini/gemini_... ) I feel this is overstated hype. There is no competitor to GPT-4 being released today. I…