Live data from Hacker News

Qwen3-VL can scan two-hour videos and pinpoint nearly every detail

the-decoder.com

31–40 of 86 posts

Re: Qwen3-VL can scan two-hour videos and pinpoint nearly every detail

#32
> The test works by inserting a semantically important "needle" frame at random positions in long videos, which the system must then find and analyze.

This seems to be somewhat unwise. Such an insertion would qualify as an anomaly. And if it's also trained that way, would you not train the model to find artificial frames where they don't belong?

Would it not have been better to find a set of videos where something specific (common, rare, surprising, etc) happens at some time and ask the model about that?

Re: Qwen3-VL can scan two-hour videos and pinpoint nearly every detail

#33

I was playing around with Qwen3-VL to parse PDFs - meaning, do some OCR data extraction from a reasonably well-formated PDF report. Failed miserably, although I was using the 30B-A3B model instead of the larger one. I like the Qwen models and use them for other tasks successfully. It is so interesting how LLMs will do quite well in one situation and quite badly in another.

The opus models seems pretty adept and extracting structured data from ocr https://www.ocrarena.ai/battle

Re: Qwen3-VL can scan two-hour videos and pinpoint nearly every detail

#34
post #5

Does anyone else worry about this technology used for Big Brother type surveillance?

Big Brother is a reference to George Orwell's critique of Communism in Nineteen Eighty-Four. Qwen is a video model trained by a Communist government, or technically by a company with very close ties to the Chinese government. The Chinese government also has laws requiring AI be used to further the political goals of China in particular and authoritarian socialism in general. In the light of all this, I think it's rea…

Just nitpicking here, but 1984 is a critique of totalitarianism. The only references to systems of government in the book refer to "The German Nazis and the Russian Communists".

Orwell was a democratic socialist. He was opposed to totalitarian politics, not communism per se.

Re: Qwen3-VL can scan two-hour videos and pinpoint nearly every detail

#35
post #11

Earlier quoted context omitted.

Where have you been the last decade? It’s already in use, or models like it, by companies selling access to The State https://deflock.me Not to mention cloud platforms that collect evidence and process it with all the models and store that information for searching… https://www.revir.ai

No mention of palantir?

Palantir's just the new guy on the block: https://en.wikipedia.org/wiki/Sentient_(intelligence_analysi...

Re: Qwen3-VL can scan two-hour videos and pinpoint nearly every detail

#36
post #27

Hope this on day will be used for auto-tagging all video assets with time codes. The dream of being able to search for running horse and find a clip containing a running horse at 4m42s in one of thousands of clips.

It’s not difficult to hack this together with CLIP. I did this with about a tenth of my movie collection last week with a GTX 1080 - though it lacks temporal understanding so you have to do the scene analysis yourself

Re: Qwen3-VL can scan two-hour videos and pinpoint nearly every detail

#37
post #5

Does anyone else worry about this technology used for Big Brother type surveillance?

In surveillance and police states like The Netherlands it has been used since forever: https://www.theguardian.com/cities/2018/mar/01/smart-cities-... Now people will say again that this project has been abandoned, which just isn't true (2024): https://www.dutchnews.nl/2024/06/smart-street-surveillance-o...

I was watching a crime solving show from the UK. A huge percentage of the crimes are solved using camera footage. Also, they use geofencing, looking at which phones went in and out of the crime location at the time of the crime.

Re: Qwen3-VL can scan two-hour videos and pinpoint nearly every detail

#38

Earlier quoted context omitted.

Big Brother is a reference to George Orwell's critique of Communism in Nineteen Eighty-Four. Qwen is a video model trained by a Communist government, or technically by a company with very close ties to the Chinese government. The Chinese government also has laws requiring AI be used to further the political goals of China in particular and authoritarian socialism in general. In the light of all this, I think it's rea…

Just nitpicking here, but 1984 is a critique of totalitarianism. The only references to systems of government in the book refer to "The German Nazis and the Russian Communists". Orwell was a democratic socialist. He was opposed to totalitarian politics, not communism per se.

It's true that it's about totalitarianism to some extent. But we have Orwell's actual words here that it's chiefly about communism

> [Nineteen Eighty-Four] was based chiefly on communism, because that is the dominant form of totalitarianism, but I was trying chiefly to imagine what communism would be like if it were firmly rooted in the English speaking countries, and was no longer a mere extension of the Russian Foreign Office.

And of course Animal Farm is only about communism (as opposed to communism + fascism). And the lesser known Homage to Catalonia depicts the communist suppression of other socialist groups.

By all this I just mean to say when you're reading Nineteen Eighty-Four what he's describing is barely a fictionalization of what was already going on in the Soviet Union. There's just not a lot in the book that is specifically Nazi or Fascist.

I don't have any opinion on whether he thought there were non-totalitarian forms of communism.

Re: Qwen3-VL can scan two-hour videos and pinpoint nearly every detail

#39
Not so relevant to the thread but ive been uploading screenshots from citrix guis and asking qwen3-vl for the appropriate next action eg Mouseclick, and while it knows what to click it struggles to accurately return which pixel coordinates to click. Anyone know a way to get accurate pixel coordinates returned?

Re: Qwen3-VL can scan two-hour videos and pinpoint nearly every detail

#40

Not so relevant to the thread but ive been uploading screenshots from citrix guis and asking qwen3-vl for the appropriate next action eg Mouseclick, and while it knows what to click it struggles to accurately return which pixel coordinates to click. Anyone know a way to get accurate pixel coordinates returned?

Also curious about this. I tried https://moondream.ai/ as well for this task and it felt still far from being bulletproof.
Post reply on HN