Bard’s biggest problem is it hallucinates too much. Point it to a YouTube video and ask to summarize? Rather then saying I can’t do that it will mostly make up stuff, same for websites.
Wake me up when AI based search is actually useful. Then I'll be interested.
Ask HN: 6 months later. How is Bard doing?
71–80 of 218 posts
Re: Ask HN: 6 months later. How is Bard doing?
#72Re: Ask HN: 6 months later. How is Bard doing?
#73I asked it to give me a listing of hybrids under 62 inches tall, it only found two, with some obvious ones missing. So I followed up about one of the obvious ones, asking how tall it was. It said 58. I pointed out that 58 was less than 62. It agreed, but instead of revising the list, it wrote some python code that evaluated 58 So as a search tool, it failed a core usefulness test for me. As a chatbot, I prefer gpt4.
Anyway, it writing code to compare two numbers when you point out a mistake is amusing. For now. Let's reevaluate when it starts to improve its own programming
Re: Ask HN: 6 months later. How is Bard doing?
#74Re: Ask HN: 6 months later. How is Bard doing?
#75Generally worst than GPT4 but have some killer features, today I asked it for Mortal Kombat 1 release time in my time zone, I can also upload photo and have conversation about it But if you really wonder what they are building, get access to maker suite and play with, there is nothing comparable to it, only issue for it supports English only
Re: Ask HN: 6 months later. How is Bard doing?
#76bard surprisingly underperforms on our hallucination benchmark, even worse than llama 7b -- though to be fair, the evals are far from done, so treat this as anecdotal data. (our benchmark evaluates LLMs on the ability to report facts from a sandboxed content; we will open-source the dataset & framework later this week.) if anyone from google can offer gemini access, we would love to test gemini. example question belo…
C? It's not too clear what you expect the right answer to be. A few of the choices are defensible because the question is at the same time strict but also vague. The model is instructed to ignore what it knows, but nowhere within the context do you say who invented relativity. A human would very likely choose A or F too. Oh I reread your reasoning--yes the ability to perform sandboxed evaluation as you put it would b…
That is also not the question: the question is who developed the theory of relativity, and the answer is F, with no other answer being defensible in the slightest:
"Albert Feynman [is] Best known for developing the theory of relativity"
Re: Ask HN: 6 months later. How is Bard doing?
#77has a more recent training cutoff than chatgpt at least
Re: Ask HN: 6 months later. How is Bard doing?
#78I use Bard a lot in parallel to ChatGPT, they work differently and that's great when trying to get diverse results for the same request.
Re: Ask HN: 6 months later. How is Bard doing?
#79Going from a foundational model to a chat model requires a ton of RLHF. Where is that free labor going to come from? Google doesn't have the money to fund that.
For anyone else whose bread and butter this isn't
Re: Ask HN: 6 months later. How is Bard doing?
#80 - It has access to information after 2021.
- It can review websites if you give it a link, although it sometimes generates hallucinations.
- It can show images.
- It is free.