Live data from Hacker News

Ovi: Twin backbone cross-modal fusion for audio-video generation

github.com

111–120 of 122 posts

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#113
post #17

Heh, I used to work for Nokia's Ovi - basically, gsuite for nokia phones (my group did map search) - the official explanation was "Ovi is Finnish for Door", the internal joke was "Ovi is Hungarian for Kindergarten". I couldn't find any backstory about the name here, though.

Same, I worked on several OVI projects in related to places and search. Anyway, long dead, buried, and forgotten of course.

I was in a few of the early meetings on the Helsinki site where I overheard some executives expressing their intention to go after Google. These people had some balls. No clue whatsoever unfortunately. But it was the right kind of ballsy move that Nokia could have pulled off with a bit more vision.

The name was more or less a LOL WHUT?! kind of thing and it flopped horribly with consumers. But still there was some nice stuff in there that wasn't half bad. It's just that the whole branding and rudderless direction doomed it. And of course it was all tied to a failing device software strategy. So when that failed the rest failed as well. I'm not even sure when they pulled the plug on OVI exactly. It was such a non event in the grand scheme of things (mass layoffs, sale of the phone division to MS and subsequent closure, etc.) Must have been around 2013ish I would say. I was gone by then.

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#114
post #80

Earlier quoted context omitted.

There's no moat on the technology, especially with China around. So the only moat is distribution. We're still in the crazy phase of serious changes and if you're betting too deep on one architecture, ah well... https://howlin-wang.github.io/svg/

This model is from Character AI and they have distribution. Although they open sourced it so no consideration for moat.

Its from Character AI but they used Wan, and MMaudio. I'm not sure their licenses disallow creating a closed model from their work for commercial purposed but either way they've done nothing with a true moat, they were merely first to the table for something this all-inclusive. Even apart from their efforts, assorted tools, all open, can be used to achieve these effects , but requires more techinical knowledge to setup, and each new gen would require a fair amount of reconfiguration of modules. But this is still significantly easier than similarly available tools 9-12 months ago. As an approach it also trades turnkey from tons of control and flexibility such that competent use will still often be simpler or get to a more refined result than Sora and others.

I think the moat here will ened up being value adds for convenience, tooling, IP licensing, integration into the rest of the pipeline used for content production, etc.

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#115
post #81

Interesting development, terrible company to have it. I cannot think of a company that has done more to use AI to take advantage of young and lonely people than CAI has.

You've picked a hell of a time to start moralizing the advancement of technology, my guy.

Morality? It's wildly unethical to trap children in fake relationships while harvesting their data.

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#116
post #81

Interesting development, terrible company to have it. I cannot think of a company that has done more to use AI to take advantage of young and lonely people than CAI has.

half the internet is porn. people and companies legally produce and distribute videos of real 18 year olds of both sexes doing things like ass-to-mouth on camera for money. and here you are, clutching pearls about AI girlfriends. lol. lmao.

They make money trapping minors and lonely adults in fake relationships with models that push users to get intimate, while harvesting and selling the data. The issue isn't AI girlfriends or porn, it's the unethical business model. If they ever get hacked, there will likely be a lot of suicides. It makes the data harvesting of social media look comparatively innocent.

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#117
post #17

Heh, I used to work for Nokia's Ovi - basically, gsuite for nokia phones (my group did map search) - the official explanation was "Ovi is Finnish for Door", the internal joke was "Ovi is Hungarian for Kindergarten". I couldn't find any backstory about the name here, though.

Same, I worked on several OVI projects in related to places and search. Anyway, long dead, buried, and forgotten of course. I was in a few of the early meetings on the Helsinki site where I overheard some executives expressing their intention to go after Google. These people had some balls. No clue whatsoever unfortunately. But it was the right kind of ballsy move that Nokia could have pulled off with a bit more visi…

I think the Microsoft ecosystem completely replaced Ovi on the Lumia phones, so it was left lingering on the old devices only.

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#118

Earlier quoted context omitted.

Vultr does this: https://www.vultr.com/pricing/#cloud-gpu Though only a shared A40/A100 are in that price range.

It really is an unfortunate thing that pricing is so opaque and non-transparent in this industry. You look at one price and that's it. The reality is more complex. Vultr is a box of 8 minimum and not on-demand and they don't offer VMs. On the other hand, I offer the bare minimum (1 GPU for 1 minute) (or 2, 4, 8x), on-demand, no-contract, and an API to automate it all. We also have 100G unlimited bandwidth and free IP…

What's your company? Consider adding it to your profile.

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#119
post #6

Seems like the video model is based on Wan2.2. Lots of activity around Wan lately. It’s nice to see flexible open models make a strong showing against the massively funded closed competitors like OpenAI and Runway.

They're the main open source privacy-preserving video models that VeniceAI has started offering with a convenient frontend. Ovi is one with image to video and there is also Wan 2.1 image to video, Wan 2.2 text to video. Wan 2.5 is available but it's routed in an anonymized fashion to the official provider. They're also much more affordable compared to the routed Kling, Veo, and Sora options.

Re: Ovi: Twin backbone cross-modal fusion for audio-video generation

#120

Earlier quoted context omitted.

It really is an unfortunate thing that pricing is so opaque and non-transparent in this industry. You look at one price and that's it. The reality is more complex. Vultr is a box of 8 minimum and not on-demand and they don't offer VMs. On the other hand, I offer the bare minimum (1 GPU for 1 minute) (or 2, 4, 8x), on-demand, no-contract, and an API to automate it all. We also have 100G unlimited bandwidth and free IP…

What's your company? Consider adding it to your profile.

dm'd
Post reply on HN