Live data from Hacker News

MaMMUT: A simple vision-encoder text-decoder architecture for multimodal tasks

ai.googleblog.com

1–10 of 36 posts

Re: MaMMUT: A simple vision-encoder text-decoder architecture for multimodal tasks

#3
post #2

Is there anything actually released here? Just a paper? No weights, not even code ?! No interactive product of any kind? The degree of Google’s inability to actually ship anything even now is totally mindblowing

That's Google.

I don't bother to read most Google papers unless someone tells me that they're doing something astounding. Just because I know I don't have access to their models, their code or their data. So what's the point?

As a community we need to stop accepting and stop citing papers like these.

There is no science without replicability, and it is literally impossible to replicate this work. It's not worth the paper it's printed on.

It's fine if Google wants to play with its toys at home. But we should stop pretending this is research of any value.

Re: MaMMUT: A simple vision-encoder text-decoder architecture for multimodal tasks

#4
post #2

Is there anything actually released here? Just a paper? No weights, not even code ?! No interactive product of any kind? The degree of Google’s inability to actually ship anything even now is totally mindblowing

That's Google. I don't bother to read most Google papers unless someone tells me that they're doing something astounding. Just because I know I don't have access to their models, their code or their data. So what's the point? As a community we need to stop accepting and stop citing papers like these. There is no science without replicability, and it is literally impossible to replicate this work. It's not worth the p…

I don't even bother with the "astounding" things unless there's code. MusicLM was cool, but without code/weights it might as well be a hoax.

Re: MaMMUT: A simple vision-encoder text-decoder architecture for multimodal tasks

#5
post #2

Is there anything actually released here? Just a paper? No weights, not even code ?! No interactive product of any kind? The degree of Google’s inability to actually ship anything even now is totally mindblowing

I wish people would stop with these shallow, boring dismissals. It's a research paper, not a product. While it's disappointing that they don't release more, it's good that they share their ideas.

Re: MaMMUT: A simple vision-encoder text-decoder architecture for multimodal tasks

#6
post #2

Is there anything actually released here? Just a paper? No weights, not even code ?! No interactive product of any kind? The degree of Google’s inability to actually ship anything even now is totally mindblowing

> Just a paper?

Yes, its describing an architecture.

> No interactive product of any kind? The degree of Google’s inability to actually ship anything even now is totally mindblowing.

Yes, Google is comically bad at shipping AI products (mostly, from their description, for “safety” reasons).

OTOH, they are very good at putting out papers that other people turn into products, so this kind of thing isn’t without value.

Re: MaMMUT: A simple vision-encoder text-decoder architecture for multimodal tasks

#7
post #5
post #2

Is there anything actually released here? Just a paper? No weights, not even code ?! No interactive product of any kind? The degree of Google’s inability to actually ship anything even now is totally mindblowing

I wish people would stop with these shallow, boring dismissals. It's a research paper, not a product. While it's disappointing that they don't release more, it's good that they share their ideas.

A research paper by itself isn't worth nothing, sure, but without the ability to reproduce the paper or even check their results it's ... not worth much.

Re: MaMMUT: A simple vision-encoder text-decoder architecture for multimodal tasks

#8
post #5

Earlier quoted context omitted.

I wish people would stop with these shallow, boring dismissals. It's a research paper, not a product. While it's disappointing that they don't release more, it's good that they share their ideas.

A research paper by itself isn't worth nothing, sure, but without the ability to reproduce the paper or even check their results it's ... not worth much.

Researchers always have a lot to take away from Google papers that don’t release code or dataset - people understand that sometimes folks at companies sometimes cannot release the code or dataset. Doesn’t make any key contributions less meaningful - if that was the case, all conferences would have banned papers that don’t release code and dataset by now.

Re: MaMMUT: A simple vision-encoder text-decoder architecture for multimodal tasks

#9
post #4

Earlier quoted context omitted.

That's Google. I don't bother to read most Google papers unless someone tells me that they're doing something astounding. Just because I know I don't have access to their models, their code or their data. So what's the point? As a community we need to stop accepting and stop citing papers like these. There is no science without replicability, and it is literally impossible to replicate this work. It's not worth the p…

I don't even bother with the "astounding" things unless there's code. MusicLM was cool, but without code/weights it might as well be a hoax.

Yeah, these days I'm "over" Google's AI research. All their papers sound cool, and they've got nice pictures/audio/etc. But nothing meaningful has ever materialized from Google.

OpenAI is killing it with ChatGPT, a publicly accessible product with research papers that have been reproduced. Facebook released huge LLMs for free. Stability released successful image models for free. etc.

Meanwhile Google has ... Tensorflow? Dying. TPUs? Only used by Google themselves or when they're free. Bard? A joke compared to ChatGPT. Imagen? Never released.

Remember when Google said they were going to have an AI call your hair salon and make appointments for you? Yeah...

But hey, at least they've got that golden mountain of PII they harvested from everyone that's been oh so valuable in building new, market defining products... It's not like small companies are running circles around them using publicly available hardware and publicly available data...

And they've still got search, a that product just keeps getting better and more useful by the day...

Re: MaMMUT: A simple vision-encoder text-decoder architecture for multimodal tasks

#10
post #8

Earlier quoted context omitted.

A research paper by itself isn't worth nothing, sure, but without the ability to reproduce the paper or even check their results it's ... not worth much.

Researchers always have a lot to take away from Google papers that don’t release code or dataset - people understand that sometimes folks at companies sometimes cannot release the code or dataset. Doesn’t make any key contributions less meaningful - if that was the case, all conferences would have banned papers that don’t release code and dataset by now.

Conferences _should_ ban papers that don't release code or other means of reliable reproduction. The only reason they don't is because "research" in ML has more or less been a joke compared to any other established scientific field. And I'm not going to give Google the benefit of the doubt. At the very least I'll treat them like any random stranger publishing a paper. But in reality I treat their papers with a heavy critical eye these days because more often than not their research has turned out to be bunk and unreproducible.
Post reply on HN