Live data from Hacker News

Ask HN: 6 months later. How is Bard doing?

news.ycombinator.com

71–80 of 218 posts

Re: Ask HN: 6 months later. How is Bard doing?

#71
post #9

Bard’s biggest problem is it hallucinates too much. Point it to a YouTube video and ask to summarize? Rather then saying I can’t do that it will mostly make up stuff, same for websites.

Wake me up when AI based search is actually useful. Then I'll be interested.

I've gotten use out of https://phind.com (was linked on HN at some point I think, I'm not affiliated but use it maybe once a month for a hard-to-find thing)

Re: Ask HN: 6 months later. How is Bard doing?

#72
Google's AI experience is going to be about the same as their social experiments which is they'll fail. I didn't think this before but now realising ChatGPT and other personal assistants (because that's what they are) will really succeed not just because of performance but network effects and social mindshare. You'll use the most popular AI assistant because that's what everyone else is using. Maybe some of these things will differ in a corporate setting but Google has really struggled to launch new products that get used as a daily habit without deprecating it within two years after. Remember Allo. I think Google is a technical juggernaut but they struggle a lot with anything that requires a network effect.

Re: Ask HN: 6 months later. How is Bard doing?

#73

I asked it to give me a listing of hybrids under 62 inches tall, it only found two, with some obvious ones missing. So I followed up about one of the obvious ones, asking how tall it was. It said 58. I pointed out that 58 was less than 62. It agreed, but instead of revising the list, it wrote some python code that evaluated 58 So as a search tool, it failed a core usefulness test for me. As a chatbot, I prefer gpt4.

Hybrids here referring to cars? My first thought was some kind of animal but that didn't make much sense and "hybrids under 62 inches" web search resulted in vehicles. I'd have trouble interpreting this query myself, and I'm clearly a next-gen AI!

Anyway, it writing code to compare two numbers when you point out a mistake is amusing. For now. Let's reevaluate when it starts to improve its own programming

Re: Ask HN: 6 months later. How is Bard doing?

#75
post #31

Generally worst than GPT4 but have some killer features, today I asked it for Mortal Kombat 1 release time in my time zone, I can also upload photo and have conversation about it But if you really wonder what they are building, get access to maker suite and play with, there is nothing comparable to it, only issue for it supports English only

Sorry, what exactly is the killer feature in this example? You say you asked it something and then didn't say what killer answer it actually responded with

Re: Ask HN: 6 months later. How is Bard doing?

#76
post #29

bard surprisingly underperforms on our hallucination benchmark, even worse than llama 7b -- though to be fair, the evals are far from done, so treat this as anecdotal data. (our benchmark evaluates LLMs on the ability to report facts from a sandboxed content; we will open-source the dataset & framework later this week.) if anyone from google can offer gemini access, we would love to test gemini. example question belo…

C? It's not too clear what you expect the right answer to be. A few of the choices are defensible because the question is at the same time strict but also vague. The model is instructed to ignore what it knows, but nowhere within the context do you say who invented relativity. A human would very likely choose A or F too. Oh I reread your reasoning--yes the ability to perform sandboxed evaluation as you put it would b…

> nowhere within the context do you say who invented relativity

That is also not the question: the question is who developed the theory of relativity, and the answer is F, with no other answer being defensible in the slightest:

"Albert Feynman [is] Best known for developing the theory of relativity"

Re: Ask HN: 6 months later. How is Bard doing?

#79

Going from a foundational model to a chat model requires a ton of RLHF. Where is that free labor going to come from? Google doesn't have the money to fund that.

> In machine learning, reinforcement learning from human feedback (RLHF) [...]

For anyone else whose bread and butter this isn't

Post reply on HN