Live data from Hacker News

Ask HN: 6 months later. How is Bard doing?

news.ycombinator.com

61–70 of 218 posts

Re: Ask HN: 6 months later. How is Bard doing?

#61
post #14

I just recently got access to bard by virtue of being a local guide on google maps? I find it can be as useful as cahtgpt4 for noodeling on technical things. It does tend to confidently hallucinate at times. Like my phone auto-corrected ostree to payee, and it proceeded to tell me all about the 'payee' version control system, then when i asked about the strange name it told me it was like managing versions in a simil…

Interesting you say “confidentially hallucinate things” - a “hallucination” isn’t any different from any other LLM output except that it happens to be wrong… “hallucination” is anthropomorphic language, it’s just doing what LLMs do and generating plausible sounding text…

Yes agree. Am sure it's because LLM developers want to ascribe human-like intelligence to their platforms.

Even "AI" I think is a misnomer. It's not intelligence as most people would conceive it, i.e. something akin to human intelligence. It's Simulated Intelligence, SI.

Re: Ask HN: 6 months later. How is Bard doing?

#64
post #9

Earlier quoted context omitted.

Wake me up when AI based search is actually useful. Then I'll be interested.

Useful for ideation stage, but not the truth stage.

Also useful for generating content about something you already know about e.g. if you have to give a presentation about a particular technology you know to your colleagues. (As you already know about the topic, you can keep the 90% which is correct and discard the 10% which is hallucination.)

Re: Ask HN: 6 months later. How is Bard doing?

#66

Earlier quoted context omitted.

You'll recall this happened before the whole ChatGPT thing blew up in hype: https://www.washingtonpost.com/technology/2022/06/11/google-... So... there's a reason why Google in particular has to be concerned with ethics and optics. I played with earlier internal versions of that "LaMDA" ("Meena") when I worked there and it was a bit spooky. There was warning language plastered all over the page ("It will lie" etc.) T…

Can you share more about how it was 'spooky'? Like it was completely unregulated?

It could response with something like "I'm just so scared. I don't know what to do. I'm so scared" to prompts that GPT3 would handle a-okay.

Re: Ask HN: 6 months later. How is Bard doing?

#67
post #29

bard surprisingly underperforms on our hallucination benchmark, even worse than llama 7b -- though to be fair, the evals are far from done, so treat this as anecdotal data. (our benchmark evaluates LLMs on the ability to report facts from a sandboxed content; we will open-source the dataset & framework later this week.) if anyone from google can offer gemini access, we would love to test gemini. example question belo…

C?

It's not too clear what you expect the right answer to be. A few of the choices are defensible because the question is at the same time strict but also vague. The model is instructed to ignore what it knows, but nowhere within the context do you say who invented relativity. A human would very likely choose A or F too.

Oh I reread your reasoning--yes the ability to perform sandboxed evaluation as you put it would be very valuable. That would be one way to have a model that minimizes hallucinations. Would be interested in testing your model once it comes out.

Re: Ask HN: 6 months later. How is Bard doing?

#68

Earlier quoted context omitted.

You'll recall this happened before the whole ChatGPT thing blew up in hype: https://www.washingtonpost.com/technology/2022/06/11/google-... So... there's a reason why Google in particular has to be concerned with ethics and optics. I played with earlier internal versions of that "LaMDA" ("Meena") when I worked there and it was a bit spooky. There was warning language plastered all over the page ("It will lie" etc.) T…

That is exactly the kind of thing I'm talking about: Lemoine was a random SWE experiencing RLHF'd LLM output for the first time, just like the rest of the world did just a few months later... and his mind went straight to "It's Sentient!". That would have been fine, but when people who understood the subject tried to explain, he decided that it was actually proof he was right so he tried to go nuclear. And when going…

By "sentient," do you mean able to experience qualia? Most people consider chickens sentient (otherwise animal cruelty wouldn't upset us, since we'd know they can't actually experience pain) - is it so hard to imagine neural networks gaining sentience once they pass the chicken complexity threshold? Sure, LLMs wouldn't have human-like qualia - they measure time in iters, they're constantly rewound or paused or edited, their universe is measured in tokens - but I don't think that means qualia are off the table.

It's not like philosophers or neuroscientists have settled the matter of where qualia come from. So how can a subject-matter expert confidently prove that a language model isn't sentient? And please let David Chalmers know while you're at it, I hear he's keen to settle the matter.

Re: Ask HN: 6 months later. How is Bard doing?

#70

I use Bard often to help me with proofreading and writing. Things that used to be a chore are now easy. I've been able to knock out a whitepaper I've been sitting on for months in just a few days. I think asking it for precise answers is the wrong approach. At this point, Bard is a lot more of an artist than a mathematician or scientist. So it's like approaching Van Gogh and asking him to do linear algebra. Bard is r…

Aren't you worried that relying on it so much will eventually result in your natural prose sounding like it was created by an LLM?
Post reply on HN