One observation: Sundar's comments in the main video seem like he's trying to communicate "we've been doing this ai stuff since you (other AI companies) were little babies" - to me this comes off kind of badly, like it's trying too hard to emphasize how long they've been doing AI (which is a weird look when the currently publicly available SOTA model is made by OpenAI, not Google). A better look would simply be to sh…
Gemini AI
221–230 of 1001 posts
Re: Gemini AI
#222Re: Gemini AI
#223Re: Gemini AI
#224> For Gemini Ultra, we’re currently completing extensive trust and safety checks, including red-teaming by trusted external parties, and further refining the model using fine-tuning and reinforcement learning from human feedback (RLHF) before making it broadly available. > As part of this process, we’ll make Gemini Ultra available to select customers, developers, partners and safety and responsibility experts for ear…
It won't be available to regular devs until Q2 next year probably (January for selected partners). So they are roughly a year behind OpenAI - and that is assuming their model is not overtrained to just pass the tests slightly better than GPT4
You are assuming GPT4 didn't do the exact same!
Seriously, it's been like this for a while, with LLMs any benchmark other than human feedback is useless. I guess we'll see how Gemini performs when it's released next year and we get independent groups comparing them.
Re: Gemini AI
#225So, better than GPT4 according to the benchmarks? Looks very interesting. Technical paper: https://goo.gle/GeminiPaper Some details: - 32k context length - efficient attention mechanisms (for e.g. multi-query attention (Shazeer, 2019)) - audio input via Universal Speech Model (USM) (Zhang et al., 2023) features - no audio output? (Figure 2) - visual encoding of Gemini models is inspired by our own foundational work o…
Re: Gemini AI
#226Not impressed with the Bard update so far. I just gave it a screenshot of yesterday's meals pulled from MyFitnessPal, told it to respond ONLY in JSON, and to calculate the macro nutrient profile of the screenshot. It flat out refused. It said, "I can't. I'm only an LLM" but the upload worked fine. I was expecting it to fail maybe on the JSON formatting, or maybe be slightly off on some of the macros, but outright ref…
Re: Gemini AI
#227"We finally beat GPT-4! But you can't have it yet." OK, I'll keep using GPT-4 then. Now OpenAI has a target performance and timeframe to beat for GPT-5. It's a race!
Didn't OpenAI already say GPT-5 is unlikely to be a ton better in terms of quality? https://news.ycombinator.com/item?id=35570690
At best Gemini seems to be a significant incremental improvement. Which is welcome, and I'm glad for the competition, but to significantly increase the applicability of of these models to real problems I expect that we'll need new breakthrough techniques that allow better control over behavior, practically eliminate hallucinations, enable both short-term and long-term memory separate from the context window, allow adaptive "thinking" time per output token for hard problems, etc.
Current methods like CoT based around manipulating prompts are cool but I don't think that the long term future of these models is to do all of their internal thinking, memory, etc in the form of text.
Re: Gemini AI
#228Not impressed with the Bard update so far. I just gave it a screenshot of yesterday's meals pulled from MyFitnessPal, told it to respond ONLY in JSON, and to calculate the macro nutrient profile of the screenshot. It flat out refused. It said, "I can't. I'm only an LLM" but the upload worked fine. I was expecting it to fail maybe on the JSON formatting, or maybe be slightly off on some of the macros, but outright ref…
Re: Gemini AI
#229Not impressed with the Bard update so far. I just gave it a screenshot of yesterday's meals pulled from MyFitnessPal, told it to respond ONLY in JSON, and to calculate the macro nutrient profile of the screenshot. It flat out refused. It said, "I can't. I'm only an LLM" but the upload worked fine. I was expecting it to fail maybe on the JSON formatting, or maybe be slightly off on some of the macros, but outright ref…
> Not impressed
This made me chuckle
Just a bit ago this would have been science fiction