Live data from Hacker News

Why Meta’s latest large language model survived only three days online

technologyreview.com

101–110 of 126 posts

Re: Why Meta’s latest large language model survived only three days online

#101
post #95

Earlier quoted context omitted.

Cars are much faster than humans. That doesn't mean transportation is solved.

“Hmm how should I get to work tomorrow? Normally I’d take the car, but after adopting a stance of distractive pedantism I realized that a car isn’t an acceptable solution to my transportation problem.” Like please explain what definition of solved you are using. It’s not one most people would be familiar with.

Solved means you have the best solution.

Re: Why Meta’s latest large language model survived only three days online

#102
post #101
post #95

Earlier quoted context omitted.

“Hmm how should I get to work tomorrow? Normally I’d take the car, but after adopting a stance of distractive pedantism I realized that a car isn’t an acceptable solution to my transportation problem.” Like please explain what definition of solved you are using. It’s not one most people would be familiar with.

Solved means you have the best solution.

What is better than a car for people who don't have access to good public transportation?

Re: Why Meta’s latest large language model survived only three days online

#103
post #96

Earlier quoted context omitted.

Fundamentally unsuited because of how they train it using "fill in the blank." Training a large model to guess when it doesn't know the answer results in fiction. They need to do something else to get nonfiction. By contrast, for Go the model was trained not to make illegal moves, because checking for that as part of the training is easy and cheap.

We have models that accurate classify things, e.g. whether or not an email is spam. There isn’t a fundamental limitation into building something like a truth classifier into a generative model so that it optimized for outputting “true” statements. The hardest part is probably identifying what is truth and what is falsehood. That’s a fundamental problem with humanity, not neural networks.

> There isn’t a fundamental limitation into building something like a truth classifier into a generative model so that it optimized for outputting “true” statements.

Problem is, they didn't do that

Re: Why Meta’s latest large language model survived only three days online

#104

At the end of the day it didn't blow people away and that's the real reason it failed to land. You can't release something like this on the heels of Stable Diffusion and not expect people to be underwhelmed. This is a user-centric design problem. It actually takes experimentation and skill to get anything useful out of Galactica and you have to actually have some sense of prompt engineering principles for it to work.…

In the domain of text, garbage is not amusing. In the domain of images, it often is.

I think the mistake here is science versus art. In science garbage is not intriguing (though AI generated images can be disturbingly well done), in art it can be amusing whether text or imagery.

"Twas Brillig and the slithy toves did gyre and gimble in the wabe. All mimsy were the borogoves and the mome raths outgrabe." (Pardon misspellings, I'm doing this from long-ago memory.)

Also, madlibs.

Re: Why Meta’s latest large language model survived only three days online

#105
post #78

Earlier quoted context omitted.

Framing is key in this context. Yann introduced the model in a very authoritative way, presenting it as production ready. His quote: "Type a text and galactica.ai will generate a paper with relevant references, formulas, and everything." [1] The AI produces output but nothing that could be considered a paper in a professional setting. Which is understandable! AGI is not here yet. But he should have presented the tool…

> "Type a text and galactica.ai will generate a paper with relevant references, formulas, and everything One could describe DALL-E as "type a text in dalle and it will generate a Picasso with the right textures and strokes and everything". One would have to be particularly obnoxious to pretend to surmize that the Dalle image is an actual Picasso painting that you can sell in Sothebys or display in his museum. That is…

Had they described it that way, DALL-E would have received way more criticism.

Lecun and his team and/or Meta did a terrible job in managing the users' expectations

Re: Why Meta’s latest large language model survived only three days online

#106
post #100
post #3

It’s algorithmically/randomly generating text without understanding. What it the proper way of using it? Fake papers? Political bs? Bad Hemingway (or Shakespeare or Chaucer or…). It’s noise that looks like sentences.

Putting the peer review system to the test? (I'm not suggesting we should do that)

This was already done with the automatic postmodernism generator[1], which was published in 1996 and is frankly basically much better than galactica at generating plausible gibberish. A particularly nice touch is that it cites references with links to other papers it generates.

[1] https://www.elsewhere.org/pomo/ and the original paper here https://www.elsewhere.org/journal/wp-content/uploads/2005/11...

Re: Why Meta’s latest large language model survived only three days online

#107

Earlier quoted context omitted.

And yet, Go AIs are now unbeatable by humans. This demonstrates that "solved" is unreasonable and unnecessary.

Cars are much faster than humans. That doesn't mean transportation is solved.

Solved in game theory has a very specific, strong definition. Transportation isn't a game in the game theoretic sense.

Re: Why Meta’s latest large language model survived only three days online

#108

This software is excellent for pseudo science. For example, young earth peddlers will be able to generate entire mambo jambo references and use them to indoctrinate more people.

It appears they don’t need AI for that.

Because they spend too much time on their BS. But now they'll be able to do it effortlessly.

Re: Why Meta’s latest large language model survived only three days online

#109
post #78

Earlier quoted context omitted.

> "Type a text and galactica.ai will generate a paper with relevant references, formulas, and everything One could describe DALL-E as "type a text in dalle and it will generate a Picasso with the right textures and strokes and everything". One would have to be particularly obnoxious to pretend to surmize that the Dalle image is an actual Picasso painting that you can sell in Sothebys or display in his museum. That is…

Had they described it that way, DALL-E would have received way more criticism. Lecun and his team and/or Meta did a terrible job in managing the users' expectations

https://www.newyorker.com/magazine/2022/07/11/dall-e-make-me...

"DALL-E, Make Me Another Picasso, Please"

The creators of an artificial intelligence that can produce almost any art work imaginable—from “cheeseburger lamp” to “the rest of mona lisa”—sift through their latest requests for original images.

Re: Why Meta’s latest large language model survived only three days online

#110
post #51
post #34

Earlier quoted context omitted.

The problem is not journalist, it's about how Meta and LeCun presented it. They presented it as "you should trust what it says and use to write papers", then hid in the small lines "oh actually really don't do that". You can't have your cake and eat it AND complain about being called out on it.

> They presented it as "you should trust what it says where ? > The problem is not journalist, What was the reason for the takedown > You can't have your cake and eat it AND complain I will agree to the extent in which Lecun's team , and other research teams need to leave corporates and go back to universities > @Ylecun: When you have a tool at your disposal, you have to know what to use it for and how. E.g. a CNC ma…

> > The problem is not journalist,

>

> What was the reason for the takedown

The system was not doing what the team said it did. Journalists documented this and demonstrated it, which caused Meta to shut it down, but journalists didn’t break it.

Blaming journalists for this is like blaming the smoke detector for interrupting your movie before the fire could.

Post reply on HN