Just tried out Google Bard, now available in Europe, and I have to say, my first experience was rather disconcerting. I threw at it a relatively obscure question about an event in the third book of A Song of Ice and Fire series, involving characters Arya Stark and Sandor Clegane. To my surprise, Google Bard's response, while confident, was a bizarre mix of actual characters and situations from the book series, but th…
Don't bother trying their other models available in GCP either. Their embedding models and their un-RLHFd generic GPT analogues are miles behind the competition. It's incredible how bad google have dropped the ball on this. We had some google people come in at $CORPO_DAYJOB the other day to sell their cloud offerings. The engineer had the gall to say "people ask me, why did google miss the boat on LLMs? I say people,…
Bard’s latest update: more features, languages and countries
91–100 of 205 posts
Re: Bard’s latest update: more features, languages and countries
#92>Can you find me a picture of Winnie the Pooh, and describe the image?
Bard gives me a link to the image on this article: https://www.psychologytoday.com/ie/blog/the-everyday-unconsc...
It provides the following description:
>The picture shows Winnie the Pooh, a yellow teddy bear with a red shirt and a blue scarf. He is sitting on a tree stump, holding a pot of honey. He has a big smile on his face and his eyes are closed, as if he is enjoying the taste of the honey. The background of the picture is a forest, with trees and flowers.
Clearly, this is entirely hallucinated. If I download the image, and re-upload it in a fresh session and ask for a description, Bard provides:
>Sure. The image shows Winnie the Pooh, Eeyore, and Tigger standing next to each other on a black background. Winnie the Pooh is a yellow bear with a red shirt. Eeyore is a donkey with a blue shirt. Tigger is a tiger with a orange shirt. They are all smiling and looking at the viewer. The image is in a cartoon style and is very colorful. It is a heartwarming image that captures the friendship of these three characters.
Aside from some minor errors (shirts on Tigger and Eeyore), and the reference to a black background (probably an issue with handling transparency), this is quite reasonable.
But if Bard is going to sometimes use its Lens capability to understand images, and other times just hallucinate a description from nothing, and there's no way for the user to tell which is being done, how are we supposed to trust the results in a situation where we don't already know the answer?
Re: Bard’s latest update: more features, languages and countries
#93Just tried out Google Bard, now available in Europe, and I have to say, my first experience was rather disconcerting. I threw at it a relatively obscure question about an event in the third book of A Song of Ice and Fire series, involving characters Arya Stark and Sandor Clegane. To my surprise, Google Bard's response, while confident, was a bizarre mix of actual characters and situations from the book series, but th…
Question, do you think this has to do with copyrighted information being in the models ? I know you might not care but I wonder if Google is operating of accounts from the internet while OpenAI has actually ingested the books ?
A blog post on the internet is also copyrighted! Google has probably the largest collection of digitized books as well.
They have no reason not to. ML networks have been trained on copyrighted data since before 2012.
Re: Bard’s latest update: more features, languages and countries
#94Re: Bard’s latest update: more features, languages and countries
#95Earlier quoted context omitted.
> now lets all acknowledge that maintaining stable performance of LLMs is an unimaginably hard problem, but how do we trust any Bard update when there's no regression testing on advertised outputs? This framing ("how do we trust [target product] when there's [a problem]") is basically a fallacy. Advocates and evangelists and shitposters deploy it every time they want to take an absolutist position but only have one a…
It's one thing to say to "use code with caution" but it's another thing to pretend to run a calculation and then hallucinate the answer (or hallucinate that it's running code). I just tried out this exact example. ME: "Do you have access to a code interpreter like Jupyter Lab, Colab, or Replit?" BARD: "Yes..." ME: "OK, great, can you execute the code to give me the prime factors of 15683615?" BARD: Prints code block.…
Not to an LLM, it isn't. You're asking for "reasoning" features, the idea of having a model of what's needs to happen and whether or not the output matches the constriants of the model. And that's not what LLMs do, at all.
That Bard attempts it is a software feature that they advertised. And it broke, apparently. And that's bad, and they should fix it. But if (per your phrasing) you think this is an "obvious" thing that they got wrong, you're mislead about what this technology does.
Re: Bard’s latest update: more features, languages and countries
#96Just tried out Google Bard, now available in Europe, and I have to say, my first experience was rather disconcerting. I threw at it a relatively obscure question about an event in the third book of A Song of Ice and Fire series, involving characters Arya Stark and Sandor Clegane. To my surprise, Google Bard's response, while confident, was a bizarre mix of actual characters and situations from the book series, but th…
Don't bother trying their other models available in GCP either. Their embedding models and their un-RLHFd generic GPT analogues are miles behind the competition. It's incredible how bad google have dropped the ball on this. We had some google people come in at $CORPO_DAYJOB the other day to sell their cloud offerings. The engineer had the gall to say "people ask me, why did google miss the boat on LLMs? I say people,…
Re: Bard’s latest update: more features, languages and countries
#97Earlier quoted context omitted.
We found similar issues with asking Bard for images. Bard generated Wikimedia links which look legit, but it's simply following the pattern of those URLs. When we loaded the links, they were all broken.
Ask ChatGPT to draw with ASCII art. It's very funny.
^
/|\
|
/ \
O
I guess technically I didn't specify that the balloon was lighter than air, or that the guy should be holding it. I say it passes on a technicality.Re: Bard’s latest update: more features, languages and countries
#98Earlier quoted context omitted.
> now lets all acknowledge that maintaining stable performance of LLMs is an unimaginably hard problem, but how do we trust any Bard update when there's no regression testing on advertised outputs? This framing ("how do we trust [target product] when there's [a problem]") is basically a fallacy. Advocates and evangelists and shitposters deploy it every time they want to take an absolutist position but only have one a…
It's one thing to say to "use code with caution" but it's another thing to pretend to run a calculation and then hallucinate the answer (or hallucinate that it's running code). I just tried out this exact example. ME: "Do you have access to a code interpreter like Jupyter Lab, Colab, or Replit?" BARD: "Yes..." ME: "OK, great, can you execute the code to give me the prime factors of 15683615?" BARD: Prints code block.…
What? Can chat-gpt run code in a sandbox? I’ve never heard of this before.
Re: Bard’s latest update: more features, languages and countries
#99Earlier quoted context omitted.
How can you so confidently say C-18 has no relevance to this - then in the same comment say you don't know enough about the law to comment on it? If you don't understand C-18 that's fine - but then you can't confidently say that it can't apply to Bard. Seems pretty clear based on how C-18 is written that it absolutely could. You have the ability to doublethink here to the point of arguing is likely pointless - but cl…
"How can you so confidently say C-18 has no relevance to this" Even Google isn't claiming it is... It is pretty ironic that you repeatedly claim that a bill is directly responsible for it when Google, despite being in a very public fight about it, isn't even claiming this. "but clearly by saying Google knows they can 'weaponize the bootlicker sorts' you are calling people who criticize this bill " Clearly I did not .…
Re: Bard’s latest update: more features, languages and countries
#100Earlier quoted context omitted.
How can you so confidently say C-18 has no relevance to this - then in the same comment say you don't know enough about the law to comment on it? If you don't understand C-18 that's fine - but then you can't confidently say that it can't apply to Bard. Seems pretty clear based on how C-18 is written that it absolutely could. You have the ability to doublethink here to the point of arguing is likely pointless - but cl…
"How can you so confidently say C-18 has no relevance to this" Even Google isn't claiming it is... It is pretty ironic that you repeatedly claim that a bill is directly responsible for it when Google, despite being in a very public fight about it, isn't even claiming this. "but clearly by saying Google knows they can 'weaponize the bootlicker sorts' you are calling people who criticize this bill " Clearly I did not .…
You're calling anyone who might disagree with Canada and agree with Google "bootlicker sorts" (I see no other way to interpret your original comment) and now you're doubling down about people "ready to tow the line for them" and "do the dirty work for them".
That's "grossly counterproductive". Please don't call names or insinuate that commenters are doing "dirty work" here.
You may want to re-read HN guidelines, especially the parts "Edit out swipes", "please reply to the argument instead of calling names", and "Assume good faith."