Gemini Ultra isn't released yet and is months away still. Bard w/ Gemini Pro isn't available in Europe and isn't multi-modal, https://support.google.com/bard/answer/14294096 No public stats on Gemini Pro. (I'm wrong. Pro stats not on website, but tucked in a paper - https://storage.googleapis.com/deepmind-media/gemini/gemini_... ) I feel this is overstated hype. There is no competitor to GPT-4 being released today. I…
Gemini AI
351–360 of 1001 posts
Re: Gemini AI
#352I asked Bard, "Are you running Gemini Pro now?" And it told me, "Unfortunately, your question is ambiguous. "Gemini Pro" could refer to..." and listed a bunch of irrelevant stuff. Is Bard not using Gemini Pro at time of writing? The blog post says, "Starting today, Bard will use a fine-tuned version of Gemini Pro for more advanced reasoning, planning, understanding and more." (EDIT: it is... gave me a correct answer…
For the record, GPT-4 still thinks it's GPT-3.
"Are you GPT-4?": https://chat.openai.com/share/1786f290-4431-45b0-856e-265b38...
"Are you GPT-3?": https://chat.openai.com/share/00c89b4c-1313-468d-a752-a1e7bb...
"What version of GPT are you?": https://chat.openai.com/share/6e52aec0-07c1-44d6-a1d3-0d0f88...
"What are you?" + "Be more specific.": https://chat.openai.com/share/02ed8e5f-d349-471b-806a-7e3430...
All these prompts yield correct answers.
Re: Gemini AI
#353This demo is nuts: https://youtu.be/UIZAiXYceBI?si=8ELqSinKHdlGlNpX
Re: Gemini AI
#354One observation: Sundar's comments in the main video seem like he's trying to communicate "we've been doing this ai stuff since you (other AI companies) were little babies" - to me this comes off kind of badly, like it's trying too hard to emphasize how long they've been doing AI (which is a weird look when the currently publicly available SOTA model is made by OpenAI, not Google). A better look would simply be to sh…
They showed AlphaGo, they showed Transformers.
Pretty good track record.
Re: Gemini AI
#355One of my biggest concerns with many of these benchmarks is that it’s really hard to tell if the test data has been part of the training data. There are terabytes of data fed into the training models - entire corpus of internet, proprietary books and papers, and likely other locked Google docs that only Google has access to. It is fairly easy to build models that achieve high scores in benchmarks if the test data has…
Cheating seems to be rampant, and by cheating I mean training on test questions + answers. Sometimes intentional, sometimes accidental. There are some good papers on checking for contamination, but no one is even bothering to use the compute to do so.
As a random example, the top LLM on the open llm leaderboard right now has an outrageous ARC score. Its like 20 points higher than the next models down, which I also suspect of cheating: https://huggingface.co/spaces/HuggingFaceH4/open_llm_leaderb...
But who cares? Just let the VC money pour in.
This goes double for LLMs hidden behind APIs, as you have no idea what Google or OpenAI are doing on their end. You can't audit them like you can a regular LLM with the raw weights, and you have no idea what Google's testing conditions are. Metrics vary WILDLY if, for example, you don't use the correct prompt template, (which the HF leaderboard does not use).
...Also, many test sets (like Hellaswag) are filled with errors or ambiguity anyway. Its not hidden, you can find them just randomly sampling the tests.
Re: Gemini AI
#356There's some dissonance in the the way this will swamp out searches for the web-alternative Gemini protocol by the biggest tech company in the world proudly boasting how responsible and careful they are being to improving things "for everyone, everywhere in the world".
It's probably just an unfortunate coincidence. After all, Gemini is a zodiac sign first and foremost, you'd have to specify what exactly you want anyway.
Take all the hundreds of thousands of words in popular languages. And all the human names. And all possible new made up words and made up names. And land on one that's a project with a FAQ[1] saying "Gemini might be of interest to you if you: Value your privacy and are opposed to the web's ubiquitous tracking of users" - wait, that's Google's main source of income isn't it?
Re: Gemini AI
#357There's a huge amount of criticism for Sundar on Hacker News (seemingly from Googlers, ex-Googlers, and non-Googlers), but I give huge credit for Google's "code red" response to ChatGPT. I count at least 19 blog posts and YouTube videos from Google relating to the Gemini update today. While Google hasn't defeated (whatever that would mean) OpenAI yet, the way that every team/product has responded to improve, publiciz…
Your metric for AI innovation is…number of blog posts?
Re: Gemini AI
#358This demo is nuts: https://youtu.be/UIZAiXYceBI?si=8ELqSinKHdlGlNpX
Re: Gemini AI
#359Re: Gemini AI
#360There's some dissonance in the the way this will swamp out searches for the web-alternative Gemini protocol by the biggest tech company in the world proudly boasting how responsible and careful they are being to improving things "for everyone, everywhere in the world".
Gemini as a web protocol isn't even on the top 5 list of things that come up when you think about Gemini prior to this announcement. It would be surprising if anyone involved in naming the Google product even knew about it.
And now it never will be :)