Live data from Hacker News

Bard is now Gemini, and we’re rolling out a mobile app and Gemini Advanced

blog.google

801–810 of 1001 posts

Re: Bard is now Gemini, and we’re rolling out a mobile app and Gemini Advanced

#802
post #68

Just played with Gemini Ultra for like 10-15 mins, and right off the bat, it made mistakes I've never seen GPT-4 do. To give you an example, I asked Gemini Ultra how to set up a real-time system for a TikTok-like feed that matches card difficulty with user ability. It correctly mentioned "Item Response Theory (IRT)", which was a good start. But when I followed up asking how to implement a real-time IRT system, it sud…

It doesn't seem like it's using Gemini Ultra yet. For me it seems like only the interface has been updated since the image generation capabilities are not working.

Image generation is working for me

Re: Bard is now Gemini, and we’re rolling out a mobile app and Gemini Advanced

#803

I'm surprised they got rid of the Bard name. It struck me as a really smart choice since a Bard is someone who said things, and it's an old/archaic enough word to not already be in a zillion other names. Gemini, on the other hand, doesn't strike me as particularly relevant (except that perhaps it's a twin of ChatGPT?), and there are other companies with the same name. EDIT: I can see the advantage of picking a name t…

I'm not surprised -- I thought Bard was terrible branding. It's all associations with Shakespeare and poetry and medieval England, and as much as I might personally enjoy those, it's extremely backwards-looking, with archaic connotations. Also it sounds close to "beard" -- hairy stuff. Gemini sounds like the space program -- futuristic, a leap for mankind. It's got all the right emotional associations. It's a constel…

When I read new thread responses, I briefly thought that I wrote[1] your reply and was confused lol. Great minds think alike. I feel vindicated about my weird opinion.

[1] https://news.ycombinator.com/item?id=39306764

Re: Bard is now Gemini, and we’re rolling out a mobile app and Gemini Advanced

#804
post #549

Earlier quoted context omitted.

I got a different answer with GPT 3.5 > If the word "push" is written on the glass door in mirror writing, it means that from the other side of the door, it should be pushed. When you see the mirrored text from your side, it indicates the action to be taken from the opposite side. Therefore, in this scenario, you should push the door to open it.

Here's another one. This is a classic logic puzzle - usually about ducks. There are two pineapples in front of a pineapple, two pineapples behind a pineapple and a pineapple in the middle. How many pineapples are there? When you use ducks, Gemini can do it, when you use pineapples it cannot and thinks there are 5 instead of 3. ChatGPT 3.5 and 4 can do it. The even funnier thing is if you then say to gemini, hey - wou…

I mean, I thought and still think the answer is five… am I an AI or a human?

If the answer is so ambiguous that humans and AI get it wrong, is it really that great of a question?

Re: Bard is now Gemini, and we’re rolling out a mobile app and Gemini Advanced

#807

Earlier quoted context omitted.

My reading of the fine print (IAAL, FWIW) is that turning off Gemini Apps Activity does not affect whether human review is possible. It just means that your prompts won't be saved beyond 72 hours, unless they are reviewed by humans, in which case they can live on indefinitely in a location separate from your account. I also asked Gemini (not Ultra) and it told me that there is no way to prevent human review.

You should never ask an LLM to answer questions about itself. The answer is guaranteed to be hallucinated unless Google specifically finetuned it on an answer of that question. The answer it gave you is meaningless. (But also, coincidentally, correct.)

I recall seeing that OpenAI finetuned ChatGPT on facts related to itself, and I figured Google likely did the same. But you're right about not relying on its representations. I only skimmed its answer to see if it seemed consistent with my reading of the fine print.

Re: Bard is now Gemini, and we’re rolling out a mobile app and Gemini Advanced

#808
post #629

Earlier quoted context omitted.

I also get the wrong answer with GPT 4 https://chat.openai.com/share/4373c945-88b8-4742-8a2c-76fff2... > You should push the door. The word "push" written in mirror writing indicates that the instructions are intended for someone on the opposite side of the door from where you are standing. Since you can see the mirror writing from your side, it means the text is facing the other side, suggesting that those on the ot…

Strange, I get the right answer on GPT4 > If the word "push" is written in mirror writing and you are seeing it from your side of the glass door, you should pull the door towards you. The reason for this is that the instruction is intended for people on the other side of the door. For them, the word "push" would appear correctly, instructing them to push the door to open it from their side. Since you are seeing it in…

Yeah LLMs are not consistent.

Re: Bard is now Gemini, and we’re rolling out a mobile app and Gemini Advanced

#809
post #549

Earlier quoted context omitted.

I got a different answer with GPT 3.5 > If the word "push" is written on the glass door in mirror writing, it means that from the other side of the door, it should be pushed. When you see the mirrored text from your side, it indicates the action to be taken from the opposite side. Therefore, in this scenario, you should push the door to open it.

Here's another one. This is a classic logic puzzle - usually about ducks. There are two pineapples in front of a pineapple, two pineapples behind a pineapple and a pineapple in the middle. How many pineapples are there? When you use ducks, Gemini can do it, when you use pineapples it cannot and thinks there are 5 instead of 3. ChatGPT 3.5 and 4 can do it. The even funnier thing is if you then say to gemini, hey - wou…

The way I parsed that sentence, I came up with 5.

Re: Bard is now Gemini, and we’re rolling out a mobile app and Gemini Advanced

#810

I’ve been pretty excited to finally try Gemini advanced. So far pretty disappointed. Here’s my go-to test question - which even chat gpt 3.5 can get. Question: I walk up to a glass door. It has the word push on it in mirror writing. Should I push or pull the door, and why Gemini advanced: You should push the door. Here's why: * Mirror Writing: The word "PUSH" is written in mirror writing, meaning it would appear corr…

How do you prefer to validate if a model is actually useful for you in practice outside of solving toy problems? Are you asking these models to solve reasoning problems like this to get any benefit for yourself in your day to day use? Or do you even care if the models are useful for day to day tasks?

For me the validation process is to use it for a few weeks and then I have a good handle on what it can handle and what it can’t.
Post reply on HN