Qwen3-VL can scan two-hour videos and pinpoint nearly every detail
1–10 of 86 posts
Re: Qwen3-VL can scan two-hour videos and pinpoint nearly every detail
#2Re: Qwen3-VL can scan two-hour videos and pinpoint nearly every detail
#3Re: Qwen3-VL can scan two-hour videos and pinpoint nearly every detail
#4link to results: https://chat.vlm.run/c/82a33ebb-65f9-40f3-9691-bc674ef28b52
Quick demo: https://www.youtube.com/watch?v=78ErDBuqBEo
Re: Qwen3-VL can scan two-hour videos and pinpoint nearly every detail
#5Re: Qwen3-VL can scan two-hour videos and pinpoint nearly every detail
#6Finetuning an LLM "backbone" (if I understand correctly: a fully trained but not instruction tuned LLM, usually small because students) with OCR tokens bests just about every OCR network out there.
And it's not just OCR. Describing images. Bounding boxes. Audio, both ASR and TTS, all works better that way. Now many research papers are only really about how to encode image/audio/video to feed it into a Llama or Qwen model.
Re: Qwen3-VL can scan two-hour videos and pinpoint nearly every detail
#7Does anyone else worry about this technology used for Big Brother type surveillance?
Not to mention cloud platforms that collect evidence and process it with all the models and store that information for searching…
Re: Qwen3-VL can scan two-hour videos and pinpoint nearly every detail
#8anyone have a tl;dr for me on what the best way to get the video comprehension stuff going is? i use qwen-30b-vl all the time locally as my goto model because it's just so insanely fast, curious to mess with the video stuff, the vision comprehension works great and i use it for OCR and classification all the time
Re: Qwen3-VL can scan two-hour videos and pinpoint nearly every detail
#9It's so weird how that works with transformers. Finetuning an LLM "backbone" (if I understand correctly: a fully trained but not instruction tuned LLM, usually small because students) with OCR tokens bests just about every OCR network out there. And it's not just OCR. Describing images. Bounding boxes. Audio, both ASR and TTS, all works better that way. Now many research papers are only really about how to encode ima…
My take is it fits into the general concept that generalist models have significant advantages because so much more latent structure maps across domains than we expect. People still talk about fine tuning dedicated models being effective but my personal experience is it's still always better to use a larger generalist model than a smaller fine tuned one.
Re: Qwen3-VL can scan two-hour videos and pinpoint nearly every detail
#10Does anyone else worry about this technology used for Big Brother type surveillance?
Where have you been the last decade? It’s already in use, or models like it, by companies selling access to The State https://deflock.me Not to mention cloud platforms that collect evidence and process it with all the models and store that information for searching… https://www.revir.ai