Live data from Hacker News

Ask HN: 6 months later. How is Bard doing?

news.ycombinator.com

201–210 of 218 posts

Re: Ask HN: 6 months later. How is Bard doing?

#201
post #14

I just recently got access to bard by virtue of being a local guide on google maps? I find it can be as useful as cahtgpt4 for noodeling on technical things. It does tend to confidently hallucinate at times. Like my phone auto-corrected ostree to payee, and it proceeded to tell me all about the 'payee' version control system, then when i asked about the strange name it told me it was like managing versions in a simil…

Interesting you say “confidentially hallucinate things” - a “hallucination” isn’t any different from any other LLM output except that it happens to be wrong… “hallucination” is anthropomorphic language, it’s just doing what LLMs do and generating plausible sounding text…

Although some people insist (as you do) that "hallucination" is unreasonably anthropomorphic language, it is an extremely common term of art in the field. eg https://dl.acm.org/doi/abs/10.1145/3571730

Secondly, to be anthropormorphic, hallucination would have to be exclusively human, and why should hallucination be a purely human phenomenon? Consider this Stanford study on lab mice https://www.ncbi.nlm.nih.gov/pmc/articles/PMC6711485/ . The purpose of the study is described as being to understand hallucination and it is described by the scientists involved informally as involving hallucinating mice eg here https://www.sciencedaily.com/releases/2019/07/190718145358.h... . It does involve inducing mice to see things which are not there and behave accordingly. Most people would call that a hallucination.

Re: Ask HN: 6 months later. How is Bard doing?

#202

Earlier quoted context omitted.

To say he was a Luddite who couldn't wrap his brain around it would be an unfair compliment: it'd imply a certain principledness that he certainly didn't demonstrate with his actions afterwards. This all goes back to my original point: there's handwavy academic ponderance, and there's engaging with the real world. OpenAI showed what happens when you balance the two. LaMDA (and frankly this discussion ) demonstrate wh…

Right. I didn't know about his actions afterwards until you brought them up. That lessened my opinion of the man considerably. It's frustrating, because I almost feel like I'm on your side. I hated how Google limited LaMDA to a handpicked group of influencers and government officials for their "test kitchen." I loathed how "Open"AI tightly controlled access to DALL-E 2, and how they've kept the architecture of GPT-4…

Oh, the test kitchen demo is pretty limited.

If you are curious about their linguistical style, the difference between GPT-3 and Lamda is akin to the difference between Ralof and Hadvar playthrough, respectively - https://www.palimptes.dev/ai

Mind you, these made silly mistakes, mixing overlapping tasks and whatnot. ChatGPT with GPT-4 beats that, even if it is primed to remind us from time to time about the name of the company who made it.

Re: Ask HN: 6 months later. How is Bard doing?

#203

Earlier quoted context omitted.

You are still talking about a different concept entirely. For example, if I take this test, every single answer I give is a guess. I am 100% certain of this. This test is explicitly asking people things they don’t know.

>You are still talking about a different concept entirely. I am not. >For example, if I take this test, every single answer I give is a guess. Just look at the graph man. Many answers are given with 100% confidence (that then turn out to be wrong). If you give a 100% confidence response, you don't think you're guessing. >I am 100% certain of this. You are wrong. Thank you for illustrating my point perfectly.

I don’t get how you’re failing to see the difference between knowing that you have uncertainty at all and being precise about uncertainty when making a guess.

How can you possibly assert that I confidently know the answers to the questions on the test? That makes zero sense. I don’t know the answers. I might be able to guess correctly. That doesn’t mean I know them. It is decisively a guess.

What’s your mom’s name? observe how your answer is not a guess, hopefully.

Re: Ask HN: 6 months later. How is Bard doing?

#204
post #27

Look at Gemini, it’s their new model, currently in closed beta. Hearsay says that it’s multimodal (can describe images), GPT-4 like param count, and apparently has search built in so no model knowledge cutoff. Basically they realized Bard couldn’t cut it and merged DeepMind into Google Brain, and got the combined team to work on a better LLM using the stuff OpenAI has figured out since Bard was designed. Takes months…

> Look at Gemini, it’s their new model, currently in closed beta. With all the talent, data, and infrastructure that Google has, I believe them. That said, it is almost comical they'd not unleash what they keep saying is the better model. I am sure they have safety reasons and world security concerns given their gargantuan scale, but nothing they couldn't solve, surely? They make more in a week than what OpenAI proba…

They will.

You'll never see AI products being launched without a private test phase after Bing and the Sydney coverage in the NYT.

Google probably has something great and is making sure it's not too unexpected in how it's great before wide release.

What I'm really curious about though is Meta's commitment to a GPT-4 competitive model.

The more Google and OpenAI tread lightly and slowly with closed and heavily restricted models, the more it allows Meta to catch up with open models that as a consequence get greater public research attention.

Re: Ask HN: 6 months later. How is Bard doing?

#205

We tested Bard (aka Bison in GCP) for generating SQL. It has worse generalization capabilities than even GPT-3.5 but actually does as well at GPT-4 when given contextually relevant examples selected from a large corpus of examples. https://vanna.ai/blog/ai-sql-accuracy.html This suggests to me that it needs longer prompts to avoid the hallucination problem that everyone else seems be experiencing.

That does kind of sound like there was less specialized fine tuning and the in context learning is doing the heavy lifting.

Re: Ask HN: 6 months later. How is Bard doing?

#208

Earlier quoted context omitted.

>You are still talking about a different concept entirely. I am not. >For example, if I take this test, every single answer I give is a guess. Just look at the graph man. Many answers are given with 100% confidence (that then turn out to be wrong). If you give a 100% confidence response, you don't think you're guessing. >I am 100% certain of this. You are wrong. Thank you for illustrating my point perfectly.

I don’t get how you’re failing to see the difference between knowing that you have uncertainty at all and being precise about uncertainty when making a guess. How can you possibly assert that I confidently know the answers to the questions on the test? That makes zero sense. I don’t know the answers. I might be able to guess correctly. That doesn’t mean I know them. It is decisively a guess. What’s your mom’s name? o…

>I don’t get how you’re failing to see the difference between knowing that you have uncertainty at all and being precise about uncertainty when making a guess.

I'm not failing to see that. I'm saying that humans can be wrong about if some assertions they have are guesses or not. They're not always wrong but they're not always right either.

If you make an assertion and you say you have a 100% confidence in that assertion...that is not a guess from your point of view. I can say with 100% confidence that my mother's name is x. Great.

So what happens when i make an assertion with 100% confidence...and turn out wrong ?

Just because you know when you are guessing sometimes doesn't mean you know when you are guessing all the time.

another example.

Humans often unknowingly rationalize the reason for decisions after the fact. They don't believe those stated reasons are rationalizations rather than true.

They can be completely confident about a memory that never happened.

You are constantly making guesses you don't think are guesses.

Re: Ask HN: 6 months later. How is Bard doing?

#210

Earlier quoted context omitted.

This is exactly my experience. The answers themselves aren't too different from ChatGPT 3.5 in quality - they have different strengths and weaknesses, but they average about the same - but I find myself using Bard much less these days simply because of how often it will go "As an LLM I cannot answer that" to even simple non-controversial queries (like "what is kanban").

> As an LLM I cannot answer that One of the biggest reasons to run open models.

I started playing with a LLAMA variant recently and it loves to explain "as a LLM created by OpenAI, I can't do that, but here's some text anyway...."

I find it really amusing

Post reply on HN