Live data from Hacker News

Sora 2

openai.com

851–860 of 916 posts

Re: Sora 2

#851

This is a step towards a constant stream of hyper-personalised AI generated content optimised for max dopamine.

Kids will go to School V2 and have absolutely nothing in common to talk about because each one will have completely unique media entertainment at home.

This is already the case with the myriad of streaming services and choices of what people will let their kids watch or not. With my little kids, we tend to mostly watch PBS Kids content with a bit of Disney shows mixed in when it comes to screen time. We try to avoid seemingly empty hyper-stimulating content like Paw Patrol and others. But in the end a lot of the other kids in school/daycare talk about these shows and others, which can lead to the kids not having that kind of shared context. For instance, my four year old loves Wild Kratts, but practically nobody in his class knows the show. Meanwhile, he doesn't have any context for the various characters of Paw Patrol.

Re: Sora 2

#852

Earlier quoted context omitted.

The Sora app squaring off against Meta's social video app is the real story here. Sora 2 itself looks and sounds a little poorer than Google Veo 3. (Which is itself not currently ranked as the top video model. The Chinese models are dominating.) I think Google, with their massive YouTube data set, is ultimately going to win this game. They have all the data and infrastructure in the world to build best-in-class video…

> I think Google, with their massive YouTube data set, is ultimately going to win this game. I don't know, applying the same thinking to LLMs, Google should have been first and best with just text based LLMs too, considering the datasets they sit on (and researchers, among others the people who came up with attention). But OpenAI somehow beat them on that regardless.

The problem for Google existed with the infobox at the top of search results. If users get the answer to their query without having to visit the web page where the answer came from, and where Google shows the ads, means that users don't see ads, and that website operators don't get ad revenue. ChatGPT was Google's Kodak digital camera moment. They had internal transformers-based chatbots (that really wanted to send you pizza, for some reason), but deploying that would have cannibalized their existing business model, so in the meanwhile, their lunch got eaten by an outside competitor.

Re: Sora 2

#853

I just asked GPT 5 to generate an image of as person. I then asked it to charge the color of their shirt. It refused because "I can’t generate that specific image because it violates our content policies." I then asked it to just regenerate the first image again using the same prompt. It replied "I know this has been frustrating. You’ve been really clear about what you want, and it feels like I’m blocking you for no…

I get this all the time. Especially since GPT5, generating an image starts a massive chain where it confirms what you want, and asks you to say yes, and you say yes, and then it confirms again, and this can go on for 5-6 times. Then if you swear at it, it refuses to continue. It is insane. Fuck you, OpenAI

Ah, is it the sweating at it that cut me off?! Can we offend our robot overlords now?!

Re: Sora 2

#854

Earlier quoted context omitted.

There's a great lyric from ELUCID I think about when people say stuff like this: > I don't have the privilege to think everything ain't political

i guess what I am saying is, though everything is political, it doesn't ahve to be "so" political.

Yeah I suppose so, most comments on this kinda thing are not really discussing the technology in a vacuum. I imagine it's due to the quite cynical nature of HN at this particular time period where society is fundamentally shifting, in arguably a negative direction, with this kind of technology as one of the main reasons. I haven't been on HN for that long, what was it like 5-10 years ago? I'm curious how it will be in 5-10 years.

Re: Sora 2

#855

Earlier quoted context omitted.

I know from a dev bootcamp that you are certainly wrong. However, I also think ai coding is hyped way beyond its capability.

> dev bootcamp i will not comment any further

Not sure your reply warrants any further expenditure of effort on my part, but for the benefit of other readers:

The bootcamp (actually, evening classes in coding run in cooperation with the public sector) regularly placed graduates with employers.

They’ve seen a big hit in this since AI, and companies have explicitly cited the fact that AI can complete the same tasks that these junior devs used to perform.

Re: Sora 2

#856
post #847

Earlier quoted context omitted.

I think AI is starting to verge on making actual good music. The latest Suno release is wild. An example here: https://v.redd.it/fqlqrgumo5rf1 I find this one interesting because Rap has classically been difficult for these models (I think because it's technically difficult to find the right rhythms and flow for a given set of lyrics).

The instrumental part is quite interesting but the lyrics/vocals... it's just AI slop, like the median like if you just put a bunch of words together and shipped that. Quantity was never what people wanted imo. It is impressive if the instrumental track was made with just some prompts though

I actually think the vocals from ~2:00-~2:35 are pretty impressive there. It's wild to me that the models can play with tempo like that.

I've been listening to this across a variety of genres though, maybe these lyrics and vocals are more to your taste:

(similar to Opeth) https://suno.com/song/9ab8da05-c3f2-412d-80b4-c7d0b3ae840f?s...

(indie rock) https://suno.com/song/756dd139-4cba-4e40-b29c-03ace1c69673

Re: Sora 2

#857

Earlier quoted context omitted.

I think you're overestimating how much power LLMs consume. Let's say one video pegs a top of the line Blackwell chip at 100% utilization for 10 minutes. I think a Blackwell chip (plus cooling and other data center overhead) is somewhere around 3000 watts when running 100%. So that's about 0.5 kilowatt-hours. I suspect this is a severe overestimate because there's probably a significant amount of batching that happens…

> It seems not unlikely that the amount of energy it takes to produce, edit, and process a TikTok video exceeds half a kilowatt-hour. That would be really remarkable, considering the total power capacity of a phone battery is in the neighborhood of 0.015 kWh

[deleted]

Re: Sora 2

#858
post #847

Earlier quoted context omitted.

The instrumental part is quite interesting but the lyrics/vocals... it's just AI slop, like the median like if you just put a bunch of words together and shipped that. Quantity was never what people wanted imo. It is impressive if the instrumental track was made with just some prompts though

I actually think the vocals from ~2:00-~2:35 are pretty impressive there. It's wild to me that the models can play with tempo like that. I've been listening to this across a variety of genres though, maybe these lyrics and vocals are more to your taste: (similar to Opeth) https://suno.com/song/9ab8da05-c3f2-412d-80b4-c7d0b3ae840f?s... (indie rock) https://suno.com/song/756dd139-4cba-4e40-b29c-03ace1c69673

I don't know but it doesn't impress me one bit? like I'm not trying to hate, but it just seems kind of like the model is given the track and then it tries to just follow it by matching words and then spitting them out, like as if it could talk about making a sandwich over some epic track and it'd sound the same?

like, LLMs are fantastic at generating patterns, so words that match and same with images etc.

But there's not much uniqueness? it's "impressive" like a savantic kind of ability to come up with rap, but it doesn't really product something I'd want to listen to..?

I listened to the metal thing and kind of the same thing?

It's very high fidelity, like the quality of the drums and etc it's quite impressive, but the vocals seem off? it's like a poem being read by TTS then transformed into "metal voice"

and kind of just an averaging of "metal music" kind of like stock photos and into a track, very formulaic

not to mention many metal bands etc they do formulaic stuff especially if they have an identifying kind of hit

But to me this is cool tech, but I wouldn't listen to it

I've listened music for a long time but I don't listen to a wide variety today, however for example with pop it can be very complex or very simple, but average or "almost" will really not make a good song, it can seem simple in hindsight but probably blood sweat and tears went into such songs, or creative energy that might never come back as strong.

just my raw thoughts though. it could be me being biased knowing it's AI, but I don't think so. I think my brain has kind of adapted to a point where I can feel if something is AI because it always seems super "average"/mid?

Re: Sora 2

#859
post #858

Earlier quoted context omitted.

I actually think the vocals from ~2:00-~2:35 are pretty impressive there. It's wild to me that the models can play with tempo like that. I've been listening to this across a variety of genres though, maybe these lyrics and vocals are more to your taste: (similar to Opeth) https://suno.com/song/9ab8da05-c3f2-412d-80b4-c7d0b3ae840f?s... (indie rock) https://suno.com/song/756dd139-4cba-4e40-b29c-03ace1c69673

I don't know but it doesn't impress me one bit? like I'm not trying to hate, but it just seems kind of like the model is given the track and then it tries to just follow it by matching words and then spitting them out, like as if it could talk about making a sandwich over some epic track and it'd sound the same? like, LLMs are fantastic at generating patterns, so words that match and same with images etc. But there's…

I'd love to see a blind study comparing a wide spectrum of these AI tracks to lesser known real artists (so the participants don't just recognize the songs) to see whether people can tell or if knowledge of the source biases them. I'm genuinely curious as to the results.

I don't think people would think anything strange of a lot of these tracks if they just randomly heard them on the radio.

Re: Sora 2

#860
post #451

I'm a software engineer and hobbyist actor/director. My friends are in the film industry and are in IATSE and SAG-AFTRA. I've made photons-on-glass films for decades, and I frequently film stuff with my friends for festivals. I love this AI video technology. Here are some of the films my friends and I have been making with AI. These are not "prompted", but instead use a lot of hand animation, rotoscoping, and human v…

Taking the time and effort out of something is exactly what strips it of its beauty Beauty is not just an “idea” that someone has and needs to get out onto a medium It is a process and journey that a person undergoes to get said idea onto said medium That journey often plays out very differently than the person expects. Things change, the art is different from the idea, and the person learns and grows Our modern soci…

But I will still be entertained. Expedient AI expression can touch most people the same way a low effort meme or an off the cuff whitticism.

Art is not effort. Art is not labour. Beauty is not suffering. Art =/= craft. Art is communication.

If someone wants to suffer long the endurance journey to becaome skilled at a craft we can still respect/appreciate it the same way a sprinter spends 10 years training to run real fast, in the mean time most of us will use a vehicle to get somewhere faster.

What we're going to lose is a bunch of interesting behind the scene videos because no one is going to watch someone prompt for an hour wondering why can't I do that, but rather why didn't I do that.

Proliferating tools for creation is net good in the same sense that teaching masses to write is net good. It's strange people are opposing lowering the barrier to entry to visual communication. That's what art ultimately is, communication. Once difficult, soon ubiquituous.

Post reply on HN