Live data from Hacker News

Apple Releases Open Weights Video Model

starflow-v.github.io

41–50 of 178 posts

Re: Apple Releases Open Weights Video Model

#41

Earlier quoted context omitted.

As a (web app) developer I never quite sure what to put in alt. Figured you might have some advice here?

The question to ask is, what a sighted person learns after looking at the image? The answer is the alt text. E.g if the image is a floppy, maybe you communicate that this is the save button. If it shows a cat sleeping on the windowsill, the alt text is yep: "my cat looking cute while sleeping on the windowsill".

I really like how you framed this as the takeaway or learning that needs to happen as what should be in the alt and not a recitation of the image. Where I've often had issues is more for things like business charts and illustrations and less cute cat photos.

Re: Apple Releases Open Weights Video Model

#42
post #10

Looking at text to video examples ( https://starflow-v.github.io/#text-to-video ) I'm not impressed. Those gave me the feeling of the early Will Smith noodles videos. Did I miss anything?

I wanted to write exactly the same thing, this reminded me of the Will Smith noodles. The juice glass keeps filling up after the liquid stopped pouring in.

Re: Apple Releases Open Weights Video Model

#43

Apple has a video understanding model too. I can't wait to find out what accessibility stuff they'll do with the models. As a blind person, AI has changed my life.

Finally good news about the AI doing something good for the people.

I’m not blind and AI has been great for me too.

Re: Apple Releases Open Weights Video Model

#44
post #10

Looking at text to video examples ( https://starflow-v.github.io/#text-to-video ) I'm not impressed. Those gave me the feeling of the early Will Smith noodles videos. Did I miss anything?

I think you need to go back and rewatch Will Smith eating spaghetti. These examples are far from perfect and probably not the best model right now, but they're far better than you're giving credit for.

As far as I know, this might be the most advanced text-to-video model that has been released? I'm not sure whether the license will qualify as open enough in everyone's eyes, though.

Re: Apple Releases Open Weights Video Model

#45

Earlier quoted context omitted.

I guess that auto-generated audio descriptions for (almost?) any video you want is a very, very nice feature for a blind person.

My two cents, this seems like a case where it’s better to wait for the person’s response instead of guessing.

My two cents, this seems like a comment it should be up to the OP to make instead of virtue signaling.

Re: Apple Releases Open Weights Video Model

#46

Earlier quoted context omitted.

The question to ask is, what a sighted person learns after looking at the image? The answer is the alt text. E.g if the image is a floppy, maybe you communicate that this is the save button. If it shows a cat sleeping on the windowsill, the alt text is yep: "my cat looking cute while sleeping on the windowsill".

I really like how you framed this as the takeaway or learning that needs to happen as what should be in the alt and not a recitation of the image. Where I've often had issues is more for things like business charts and illustrations and less cute cat photos.

It might be that you’re not perfectly clear on what exactly you’re trying to convey with the image and why it’s there.

Re: Apple Releases Open Weights Video Model

#47

Earlier quoted context omitted.

The question to ask is, what a sighted person learns after looking at the image? The answer is the alt text. E.g if the image is a floppy, maybe you communicate that this is the save button. If it shows a cat sleeping on the windowsill, the alt text is yep: "my cat looking cute while sleeping on the windowsill".

I really like how you framed this as the takeaway or learning that needs to happen as what should be in the alt and not a recitation of the image. Where I've often had issues is more for things like business charts and illustrations and less cute cat photos.

"A meaningless image of a chart, from which nevertheless emanates a feeling of stonks going up"

Re: Apple Releases Open Weights Video Model

#48

Earlier quoted context omitted.

I guess that auto-generated audio descriptions for (almost?) any video you want is a very, very nice feature for a blind person.

My two cents, this seems like a case where it’s better to wait for the person’s response instead of guessing.

[flagged]

Re: Apple Releases Open Weights Video Model

#49

Earlier quoted context omitted.

The question to ask is, what a sighted person learns after looking at the image? The answer is the alt text. E.g if the image is a floppy, maybe you communicate that this is the save button. If it shows a cat sleeping on the windowsill, the alt text is yep: "my cat looking cute while sleeping on the windowsill".

I really like how you framed this as the takeaway or learning that needs to happen as what should be in the alt and not a recitation of the image. Where I've often had issues is more for things like business charts and illustrations and less cute cat photos.

The logic stays the same though the answer is longer and not always easy. Just saying "business chart" is totally useless. You can make a choice on what to focus and say "a chart of the stock for the last five years with constant improvement and a clear increase by 17 percent in 2022" (if it is a simple point that you are trying to make) or you can provide an html table with the datapoints if there is data that the user needs to explore on their own.

Re: Apple Releases Open Weights Video Model

#50
post #45

Earlier quoted context omitted.

My two cents, this seems like a case where it’s better to wait for the person’s response instead of guessing.

My two cents, this seems like a comment it should be up to the OP to make instead of virtue signaling.

> Can you share some ways AI has changed your life?

A question directed to GP, directly asking about their life and pointing this out is somehow virtue signalling, OK.

Post reply on HN