Live data from Hacker News

Why Meta’s latest large language model survived only three days online

technologyreview.com

111–120 of 126 posts

Re: Why Meta’s latest large language model survived only three days online

#111
post #110
post #51

Earlier quoted context omitted.

> They presented it as "you should trust what it says where ? > The problem is not journalist, What was the reason for the takedown > You can't have your cake and eat it AND complain I will agree to the extent in which Lecun's team , and other research teams need to leave corporates and go back to universities > @Ylecun: When you have a tool at your disposal, you have to know what to use it for and how. E.g. a CNC ma…

> > The problem is not journalist, > > What was the reason for the takedown The system was not doing what the team said it did. Journalists documented this and demonstrated it, which caused Meta to shut it down, but journalists didn’t break it. Blaming journalists for this is like blaming the smoke detector for interrupting your movie before the fire could.

No, some people on twitter highlighted the imperfections of the model and branded it DANGEROUS, despite the fact that it had a whole page disclaimer that it is indeed hallucinating. It was brought to attention in some media which led to the researchers turning it off. We already know the "dangerous" trope from GPT3 and we know it's BS.

https://twitter.com/GaryMarcus/status/1593156854372532231

Re: Why Meta’s latest large language model survived only three days online

#112

I think these efforts point out something valuable, although probably not in the way the creators intended. Lots of people use "markers" of reliability, like citing your sources or making sentences with a certain kind of structure or tone, to estimate trustworthiness. These articles make it clear that it is entirely possible to have those markers, but be entirely incorrect in your assertions about the topic in questi…

It took me an embarrassingly long time to understand that the previous marker of authority regarding news: "published in a newspaper" completely lost all its meaning as the blogosphere exploded, and publishing costs on the internet went to near zero. Kind of why I find the pearl clutching over substack hilarious, as if having a third party website sell ads on a writer's blogpost signals they are much more worthy of a…

I would argue quite the opposite.

In the world of nonsense and misinformation, competent and insightful sources become of supreme importance. In a sense, we find ourselves back in pre-Gutenberg times. The elite has access to the insider sources and knowledge while the masses have a hard time to find the truth in hearsay blogs, spam bot outputs, and memes.

The situation will hopefully improve when another gutenberg comes up with a novel information search algorithm.

Re: Why Meta’s latest large language model survived only three days online

#113
post #101

Earlier quoted context omitted.

Solved means you have the best solution.

What is better than a car for people who don't have access to good public transportation?

The fact that we don't have a better solution now is why it's worth researching. We only have modern cars because we didn't consider transportation "solved" when we had horse-pulled carriages. Many people are looking for better solutions right now, because we believe they certainly exist.

Re: Why Meta’s latest large language model survived only three days online

#114
post #3

It’s algorithmically/randomly generating text without understanding. What it the proper way of using it? Fake papers? Political bs? Bad Hemingway (or Shakespeare or Chaucer or…). It’s noise that looks like sentences.

Augmenting intelligence through exploration. It is as much a tool for discovery as a conversation over lunch.

Re: Why Meta’s latest large language model survived only three days online

#115
post #113

Earlier quoted context omitted.

What is better than a car for people who don't have access to good public transportation?

The fact that we don't have a better solution now is why it's worth researching. We only have modern cars because we didn't consider transportation "solved" when we had horse-pulled carriages. Many people are looking for better solutions right now, because we believe they certainly exist.

Isn’t that true of everything? There may always be something better in the future. It’s ridiculous to say that we can’t have solutions to anything in the present because there may be better options in the future.

Re: Why Meta’s latest large language model survived only three days online

#116
post #29

Earlier quoted context omitted.

To be clear, the fact that it is difficult is not a defense of Galactica and its proponents; it is a reason for suspecting that these sorts of language models are fundamentally unsuited to the task.

Why “fundamentally unsuited”? Neural networks have solved tons of problems previously thought to be “too hard” for ML, e.g. playing Go.

Note that we are not talking about neural networks in general, but specifically the sort of generative autoregressive language model that Galactica is. What reason do we have to think that such a model is more likely to produce a true statement than a false one? - especially as just one misplaced truth-valued function or operator is likely to turn a true proposition into a false one. Truthfulness (not to be confused with truthiness) of their productions does not seem to be something we should expect from how they work, and the empirical evidence from Galactica supports this view.

Re: Why Meta’s latest large language model survived only three days online

#117
post #96

Earlier quoted context omitted.

Fundamentally unsuited because of how they train it using "fill in the blank." Training a large model to guess when it doesn't know the answer results in fiction. They need to do something else to get nonfiction. By contrast, for Go the model was trained not to make illegal moves, because checking for that as part of the training is easy and cheap.

We have models that accurate classify things, e.g. whether or not an email is spam. There isn’t a fundamental limitation into building something like a truth classifier into a generative model so that it optimized for outputting “true” statements. The hardest part is probably identifying what is truth and what is falsehood. That’s a fundamental problem with humanity, not neural networks.

Truth has nothing to do with humanity unless you mean the specific way humans construct belief systems.

Anyway I already told you the answer. The AI will need a series of trainable belief systems to verify whether statements are internally consistent. The strange part about this is that the AI would need to have a way to obtain validation and each prompt would have to derive a new belief system which you must use in the next prompt.

In other words, the model must be able to learn continuously. That is something that these single shot AI models are not capable of.

Re: Why Meta’s latest large language model survived only three days online

#118

Earlier quoted context omitted.

My dad is a doctor who oversees residents. Seems like half the time they call him for advice he just puts their question into gpt-3 and regurgitates it’s answer, so bill isn’t the only one.

If that's actually happening--and I am both skeptical and terrified that it is--it seems awfully close to malpractice or even (criminal) negligence.

It sounds like someone is making fun of their dad for sounding like a robot.

Re: Why Meta’s latest large language model survived only three days online

#119
post #63

Earlier quoted context omitted.

How would a system that generates false information (especially likely for fields that are not well represented in the training set, based on the site) help with brainstorming for practitioners in that field?

This wasnt meant to generate valid scientific papers, and Lecun said so too. It generates interesting associations. It rambles sometimes and goes on tangents that are sometimes relevant sometimes not. It can inform you of related ideas that you were not aware of. It's like a fuzzy google scholar. It is in no way valid publishable research, but it's like a bicycle for researchers. At least that was what i managed to f…

If the purpose is to generate interesting associations, why is the output a paper? Why not a graph showing overlapping subfields worth investigating or relationships between papers via citations and shared ideas?

Is it really a surprise that people have a different reaction to machine generated scientific papers that contain a large amount of plain nonsense than they do to a machine generated piece of art?

Re: Why Meta’s latest large language model survived only three days online

#120
post #72

Earlier quoted context omitted.

As far as I understand (and reading their Limitations page also), the system is quite likely to simply invent facts, particularly in niche fields - which may well mislead you and lead on a wild goose chase.

Yes , and that's great. Science is about inventing ideas and testing them, it's literally about chasing wild geese. A typical scientific review paper or perspective contains tons of such speculation. But right now the process of hunting down citations is excruciating and most often done lazily. Even the best review papers contain erroneous citations to irrelevant papers, or improperly cited results , papers etc. Peer…

> Yes , and that's great. Science is about inventing ideas and testing them, it's literally about chasing wild geese.

There are an infinity of ideas we can test. The large majority of them are either obviously wrong or completely useless. The reason why researchers spend so much time embedded in a field is to enable them to come up with ideas that are more likely to be worth investigating than another idea.

And again, if the goal is to generate hypotheses then the output should be a hypothesis - not a paper that presents the hypothesis and claims to evaluate it.

Post reply on HN