Live data from Hacker News

Gemini Omni

deepmind.google

141–150 of 152 posts

Re: Gemini Omni

#141

Earlier quoted context omitted.

[flagged]

Yes, each data center uses an entire lake of water. The hysterics are pretty sad and baseless. Have you ever considered that the solution to not having enough power is to generate more power, not curtail progress? All of these data center should come with their own solar panel arrays and battery packs. Who knows with enough need they might each come with their own small nuclear reactor.

[dead]

Re: Gemini Omni

#142

Earlier quoted context omitted.

They can't even reliably follow instructions from text. I think "it's just around the corner/just wait x months/just wait and see bro" is one of the most telling signs of AI psychosis.

Well, I’m gonna drop out of this because you don’t want to really accept that what we have is genuinely useful. I’ve seen it across multiple companies. It works very well on my team and that has made me a believer. I was skeptical and rightfully so for a very long time.

I think LLMs are extremely useful, mostly for coding. But saying we're extremely close to an AI that can "reliably come up with novel actions for physical robots" feeds into the hype that these tools can do way or are very close to doing more than they're actually capable of, especially when we talk about reliability. That's the kind of rhetoric that has partially created this bubble, because in no world is what you're saying realistic.

The worst thing is when someone cites a video or a demo of an AI doing something and says, "See! It's here!" Remember when the Devin video came out years ago?

You can say "eventually" AI will be able to do xyz, but eventually the sun will blow up, too, so what the fuck are we talking about?

Re: Gemini Omni

#143
post #99

Earlier quoted context omitted.

> But I am a bit reasurred that at least my job won't be fully replaced with AI :) I honestly can't comment with certainty that training from videos alone and whatever tokenization scheme they're using will ever get perfect dynamics. However it is worth noting that transformers can do a pretty good job at learning dynamics with the right pipeline (not video): https://arxiv.org/pdf/2605.15305 https://arxiv.org/pdf/260…

Thanks for the additional reading. I've often thought about LLMs and their ability to represent the physical world with its laws. And always concluded it is not really possible to do so with "just" text tokens and their relations in a latent space. It looks to me there are different approaches being taken to tackle this: * You could instruct your LLM to interact with a simulator to run experiments and infer behaviour…

> P.S.: I like discussing such topics. If anyone knows a forum or discord with like-minded people, please let me know :)

Unironically twitter (and only use the "Following" tab as opposed to the "For You")

Make an account that only follows university affiliated researchers with less than 1000 followers. In my experience discord servers get suffocated by beginners and crackpots because conversations don't naturally self-organize into their own threads.

Re: Gemini Omni

#144
post #36

In my day job I program rigid body behaviour in real time amongst other simulations. I think rigid body contact is hard to learn as it is inherently discontinuous.. something you discover when trying to code a solver. As such I always use this prompt as a test: "A video of a jenga brick tower falling over as a brick is removed. The physics of each brick must be realistic." It gave me a video of where bricks suddenly…

That[1] video looks very Twin towers. Falls in on itself and then explodes.

Re: Gemini Omni

#145
post #99

Earlier quoted context omitted.

Thanks for the additional reading. I've often thought about LLMs and their ability to represent the physical world with its laws. And always concluded it is not really possible to do so with "just" text tokens and their relations in a latent space. It looks to me there are different approaches being taken to tackle this: * You could instruct your LLM to interact with a simulator to run experiments and infer behaviour…

> P.S.: I like discussing such topics. If anyone knows a forum or discord with like-minded people, please let me know :) Unironically twitter (and only use the "Following" tab as opposed to the "For You") Make an account that only follows university affiliated researchers with less than 1000 followers. In my experience discord servers get suffocated by beginners and crackpots because conversations don't naturally sel…

Thanks, I'll try using the "Following" tab. I have a lurker account but never really used it because I only ever saw crap in "For You".

Re: Gemini Omni

#146

I'm an AI optimist. But AI video is probably the one thing that does depress me. Seeing that we can make anything visually, there's nothing that impresses me visually. I watch a video that two years ago I would've thought was really cool, and now my first thought is, "Yawn, is this AI?". Video, more than anything else, is the place where I really care if something is AI or not. If I could get a TikTok that had no AI…

I think it is like around 2010 or so I use to upload just this god awful music to early Sound Cloud because it was easy to make music with a DAW.

I even remember being on a psytrance production music mailing list 25 years ago and 95% of the tracks people posted were absolutely terrible, including myself.

I have seen a few incredible pieces from AI video but most has just not been that interesting. Then even the incredible pieces are 5 second one offs. No narrative, no continuity. I think of a random, real 5 second clip from Clockwork Orange with no backstory or context in the movie, who cares? Even the most visually interesting scenes wouldn't make sense and would be boring.

Right now it seems like we are at the stage of sampling random 5 second clips from early sound cloud and concluding this is the artistic utility of an entire new technology like DAW software and VST synths. That is obviously absurd.

Re: Gemini Omni

#147
post #81

Earlier quoted context omitted.

thanks for intro to streamable

In my experience (from a couple of years ago), Streamable can be great but it's just worth checking what their current retention policy is like. We were sharing game clips with each other and after a while realised our old clips were just gone, being deleted after 30 or 90 days or something.

noted!

Re: Gemini Omni

#148
post #87

Earlier quoted context omitted.

thanks for intro to streamable

it was the first link I got after googling free video hosting sites

I guess I haven't tried hosting/sharing anything outside of an unpublished youtube video or GDrive link in a long time.

Re: Gemini Omni

#149
post #135

Earlier quoted context omitted.

I think with a book or movie, a lot of the emotional reaction actually comes from the work of the human that created it. You can feel the emotion of the creator and the story they set out to tell and have some connection with them. You make a good point about how we've always been able to emotionally connect with fiction, but low effort AI does feel different.

What causes the emotional reaction in a film is moving images in front of you in sync with sound. Further, even the simplest of movies is the product of more than one person with more than one emotional state. What causes the emotional reaction in a book is you reading and understanding what text is in front of you. The emotional reaction can happen in the radical absence of the author and in total contravention of t…

We'll have to disagree here. My feelings are more in line with Nick Cave's recent essay on AI generated music: https://www.theredhandfiles.com/considering-human-imaginatio...

Re: Gemini Omni

#150

Earlier quoted context omitted.

You get back as much as you put in. Just like with all generative tools the quality of the output depends on the quality of input. Slapping a prompt together will only get you so far, if you want the models to generate something really striking and unique you need to get your hands dirty. Gotta break out ComfyUI and build yourself a specific workflow, once you dig deep and understand how things are put together, why…

>Gotta break out ComfyUI and build yourself a specific workflow, once you dig deep and understand how things are put together, why and so on, you can make really amazing stuff with any generative models. Where is this amazing stuff? Social media is a marketplace of ideas supposedly, so why haven't we seen a new wave of creators rise up in popularity?

Because there is a stigma about use of AI in creative spaces, the people that do use it to creative very impressive pieces don't disclose that information on their profiles. People tend to see AI anywhere mentioned in the profile and automatically shit on the work regardless of its beauty or creativity. They don't consider the staggering amount of work that goes in to the pics with all the control nets, custom hyper parameter tuning, custom finetuned lora's, and many other technical like workflow chaining and such. They automatically assume someone only spent 5 seconds on some slop prompt and that's it. But I can assure you if no mention of AI is anywhere everyone who looks at the work is always impressed. So you have an observation bias situation going on. You see only AI slop because a. most of its is low effort slop and b. the good stuff you assume had no AI in it because it wasn't disclosed by the artist.
Post reply on HN