Live data from Hacker News

Flux 3 X Mimic: The Next Generation of Video-Action Models

bfl.ai

11–20 of 57 posts

Re: Flux 3 X Mimic: The Next Generation of Video-Action Models

#13
Ok the phrasing here, it’s, it’s just -

> However, compared to more specialized approaches for representation learning they produce less disentangled representations, which puts a ceiling on their usefulness for tasks that require world understanding.

Only an LLM would use a less disentangled representation of the concept “more entangled” when trying to explain to people in the real world why less disentangled representations are not as useful for modeling the real world.

Re: Flux 3 X Mimic: The Next Generation of Video-Action Models

#16

Ok the phrasing here, it’s, it’s just - > However, compared to more specialized approaches for representation learning they produce less disentangled representations, which puts a ceiling on their usefulness for tasks that require world understanding. Only an LLM would use a less disentangled representation of the concept “more entangled” when trying to explain to people in the real world why less disentangled repres…

Bad news, humans write silly things all the time. If anything, LLMs are less likely to make awkward phrasings than people, because they aren’t found in the training data very often.

Re: Flux 3 X Mimic: The Next Generation of Video-Action Models

#18

The video at around 3.30min, where the robot arm took 3 attempts to reseat the window trim, was quite unnerving - I have not seen such resolving before. Is it new or am I way out of the loop?

You are indeed out of the loop. Google did something arguably more impressive more than a year ago, with a VLA based bot replacing a tensioned timing belt: https://www.youtube.com/watch?v=2AAFiuEP7iE
Post reply on HN