Live data from Hacker News

Why Meta’s latest large language model survived only three days online

technologyreview.com

31–40 of 126 posts

Re: Why Meta’s latest large language model survived only three days online

#31

"A fundamental problem with Galactica is that it is not able to distinguish truth from falsehood," In true science, it is exceptionally hard to distinguish truth from falsehood for many of the interesting subjects. It can take decades of work to reach consensus on what is "truth." Physics in the early 20th century is a great example of this debate.

What does science have to do with truth? I thought it was a process of supporting hypotheses with observations?

Re: Why Meta’s latest large language model survived only three days online

#32
post #29

Earlier quoted context omitted.

To be clear, the fact that it is difficult is not a defense of Galactica and its proponents; it is a reason for suspecting that these sorts of language models are fundamentally unsuited to the task.

Why “fundamentally unsuited”? Neural networks have solved tons of problems previously thought to be “too hard” for ML, e.g. playing Go.

Go is not solved.

The AI doesn't know the best move. It just knows a good move.

Re: Why Meta’s latest large language model survived only three days online

#33
post #25

Earlier quoted context omitted.

There are people who believe explicit works of fiction. Marvel movies come to mind. I'll know we've arrived when super hero films begin with a disclaimer. The runtime of the podcast was 1:34:27

Weird, it shows up as 1:40 long for me. It's the last 5 minutes of the episode, where they claim GPT-3 is an all knowing machine that will generate factual responses to any question in a way that's superior to google search.

My dad is a doctor who oversees residents. Seems like half the time they call him for advice he just puts their question into gpt-3 and regurgitates it’s answer, so bill isn’t the only one.

Re: Why Meta’s latest large language model survived only three days online

#34
post #16

This outcome from using a large language model to mimic reasoning isn’t surprising. What’s surprising is Yan LeCun’s childish and petty reaction to this entirely foreseeable series of events: > Galactica demo is off line for now. It’s no longer possible to have some fun by casually misusing it. Happy? He’s supposedly an expert in this sort of thing

I would urge him to put it back online, it is interesting and can be useful. Just don't make a press release about it, journalists ruin everything.

The problem is not journalist, it's about how Meta and LeCun presented it.

They presented it as "you should trust what it says and use to write papers", then hid in the small lines "oh actually really don't do that".

You can't have your cake and eat it AND complain about being called out on it.

Re: Why Meta’s latest large language model survived only three days online

#35

"A fundamental problem with Galactica is that it is not able to distinguish truth from falsehood," In true science, it is exceptionally hard to distinguish truth from falsehood for many of the interesting subjects. It can take decades of work to reach consensus on what is "truth." Physics in the early 20th century is a great example of this debate.

> In true science, it is exceptionally hard to distinguish truth from falsehood

I understand the sentiment, but I don’t think they referenced subtle proofs.

The system is unable to prove some high-school theorems and computations, see for instance: https://twitter.com/espadrine/status/1592879720269766659

(I don’t think that makes the system necessarily bad; it does mean that it has a long way to go still.)

Re: Why Meta’s latest large language model survived only three days online

#36

Earlier quoted context omitted.

Not being able to difinitively identify truth is different from not attempting to identify it.

Attempting to identify truth is called the scientific method.

Science can't identify the truth. It can only identify what is NOT true. As our knowledge expands, we get closer to discovering the truth; but we can never be sure we've arrived.

Re: Why Meta’s latest large language model survived only three days online

#37

This outcome from using a large language model to mimic reasoning isn’t surprising. What’s surprising is Yan LeCun’s childish and petty reaction to this entirely foreseeable series of events: > Galactica demo is off line for now. It’s no longer possible to have some fun by casually misusing it. Happy? He’s supposedly an expert in this sort of thing

Framing is key in this context. Yann introduced the model in a very authoritative way, presenting it as production ready. His quote: "Type a text and galactica.ai will generate a paper with relevant references, formulas, and everything." [1] The AI produces output but nothing that could be considered a paper in a professional setting. Which is understandable! AGI is not here yet. But he should have presented the tool with proper context. A tool that can generate the awful content it generated needs better framing.

And, yes, his reactions were baffling to say the least.

[1]: https://twitter.com/ylecun/status/1592619400024428544

Re: Why Meta’s latest large language model survived only three days online

#38

This software is excellent for pseudo science. For example, young earth peddlers will be able to generate entire mambo jambo references and use them to indoctrinate more people.

Are you sure this is an actual, real life problem?

Re: Why Meta’s latest large language model survived only three days online

#39

This is the kind of biased reporting that hurts journalism as a profession. It is not journalism's job to sell the public on anything. It's journalism's job to report the news. And if a large portion of the public doesn't believe the news is being reported accurately, that is a very big problem for journalism.

There are lots of problems with journalism today. This article isn't one of them, and its criticisms seem spot-on. It also brought up past attempts at something similar by Microsoft and Google, providing valuable context for somebody reading this who didn't know about those earlier efforts, so that they wouldn't think this was a failing specific to Meta.

Re: Why Meta’s latest large language model survived only three days online

#40
I think these efforts point out something valuable, although probably not in the way the creators intended. Lots of people use "markers" of reliability, like citing your sources or making sentences with a certain kind of structure or tone, to estimate trustworthiness. These articles make it clear that it is entirely possible to have those markers, but be entirely incorrect in your assertions about the topic in question.

There is no particular reason to think that this is something only AI models do. Plenty of people do the same thing, working much harder at looking, sounding, and acting like a trustworthy source, without actually putting much work into knowing what they are talking about. I think the absurdly incompetent nature of some of these AI models, is a great illustration of that point.

Post reply on HN