Live data from Hacker News

Today's Large Language Models Are Essentially BS Machines

quandyfactory.com

31–40 of 50 posts

Re: Today's Large Language Models Are Essentially BS Machines

#31
It's silly to flag this submission. Back in April researchers at Stanford reported that less than half of the results from AI-powered search corresponded to verifiable facts. What do we call the remaining portion. "BS" seems reasaonable.

https://aiindex.stanford.edu/report/

"As internet pioneer and Google researcher Vint Cerf said Monday, AI is "like a salad shooter," scattering facts all over the kitchen but not truly knowing what it's producing. "We are a long way away from the self-awareness we want," he said in a talk at the TechSurge Summit."

https://www.cnet.com/tech/computing/bing-ai-bungles-search-r...

Re: Today's Large Language Models Are Essentially BS Machines

#32

I think that an AI-powered world will create a population that doesn't know how to distinguish truth from lies. People already believe that AI has some powerful hidden knowledge that they need to use, even when the AI model is spilling garbage. In the future, they will also be incapable to separate what AI models tell from reality.

This is currently happening with people googling their way onto pseudo-science webpages/ blogs and interpreting that that content as “fact”.

Re: Today's Large Language Models Are Essentially BS Machines

#33

It's silly to flag this submission. Back in April researchers at Stanford reported that less than half of the results from AI-powered search corresponded to verifiable facts. What do we call the remaining portion. "BS" seems reasaonable. https://aiindex.stanford.edu/report/ "As internet pioneer and Google researcher Vint Cerf said Monday, AI is "like a salad shooter," scattering facts all over the kitchen but not tru…

probably it's flagged because it's so obviously true

Re: Today's Large Language Models Are Essentially BS Machines

#34
post #18

Any time I asked ChatGPT or another GPT a question regarding science (haven't asked any questions on other topics) I got a mostly correct answer back. And I've asked a few hundred by this point. This includes state of the art research covered in just one or a few articles. So I'm curious why my personal experience doesn't match all the complains about hallucinations.

I think the usefulness is pretty domain-specific. Every time I've given it a citation for a not-super-famous court opinion, it very confidently told me about a plausible-sounding case that never happened between people and companies that never existed.

Re: Today's Large Language Models Are Essentially BS Machines

#35
post #3

To me, this is the quintessential risk: It's plausible enough it will fool somebody with authority to act, but lacking competency to understand the information is low grade. Boom! "oh man.. but the computer said it was ok"

We already barely question the BS spilled by politicians and corporations. Now they have a scapegoat that wont ever die.

People have already accepted they can't do anything to stop the bullshitting.

It's not only in America, not only in government or large corporate. It's everywhere.

Re: Today's Large Language Models Are Essentially BS Machines

#36
I noticed he mentioned asking Bing, so I directly asked ChatGPT “who is Ryan McGreal?” and it wrote, “ Ryan McGreal is a writer, urbanist, and the editor of Raise the Hammer, an online publication that focuses on urban issues, transportation, and sustainability. He has been actively involved in advocating for better urban planning and transportation policies in Hamilton, Ontario, Canada. Please note that my knowledge is based on information available up to September 2021, so there may have been developments or changes in his work or activities since then.”

That appears more or less correct

So to give chatgpt an opportunity to hallucinate similarly to the article, I followed up with, “Did he write for the nyt?” and it replied, “ I do not have any information indicating that Ryan McGreal has written for The New York Times (NYT). His work primarily focuses on urban issues and transportation, as mentioned earlier, and he is associated with Raise the Hammer, a local publication in Hamilton, Ontario, Canada. It’s possible that he may have contributed to other publications, but I do not have specific information regarding his contributions to The New York Times.”

While I have seen ChatGPT make stuff up I do think it’s useful to compare specific results across LLMs before using particular examples to make holistic statements.

Re: Today's Large Language Models Are Essentially BS Machines

#37

It's silly to flag this submission. Back in April researchers at Stanford reported that less than half of the results from AI-powered search corresponded to verifiable facts. What do we call the remaining portion. "BS" seems reasaonable. https://aiindex.stanford.edu/report/ "As internet pioneer and Google researcher Vint Cerf said Monday, AI is "like a salad shooter," scattering facts all over the kitchen but not tru…

Does HN have a way to "unflag" a submission?

Or, what does a flag in HN actually do?

Re: Today's Large Language Models Are Essentially BS Machines

#39
post #21
post #16

To be honest, I hated writing essays in English classes because I felt like I'm forced to write BS to fill up the space when my argument can be summed up in several bullet points. Since I'm not a student anymore, I can just give ChatGPT a few bullet points and ask it to write a paragraph for me. As an engineer who doesn't like writing "fluff", it's great I can now outsource the BS part of writing.

As an engineer I'd hope you wouldn't have to write fluff. Brevity (while retaining full content) should be praised. I'm interested what parts of your job require the fluff? Is it communication with non engineering teams?

Not necessarily technical writings but more like emails to ask for something, emails to decline something or even a birthday card.

It's also great for writing a professional sounding complaint letter to your utility company.

Re: Today's Large Language Models Are Essentially BS Machines

#40
post #7
post #5

Earlier quoted context omitted.

100% going to happen in the near future if it hasn't already happened.

Those attorneys who blindly used ChatGPT to generate a brief regarding MC99 liability (a field they had no experience in) are a good example, I think. Of course, in that case opposing counsel started looking at the cites and quickly had questions for them...

It's really incredible how plausible sounding ChatGPT's legal BS is. Completely hallucinated cases, arguments, citations (properly formatted for real reporters!), people, ideas... but if you just skim it, there wouldn't be any immediate way for a layman to tell it was total bullshit, and I'll bet an overworked attorney could be taken off guard, too. Obviously won't get you too far in actual litigation, and I really feel for the clients of any attorney that pulls such shit.
Post reply on HN