Bridging Images and Text – A Survey of VLMs
nanonets.com
Bridging Images and Text – A Survey of VLMs
1–3 of 3 posts
Re: Bridging Images and Text – A Survey of VLMs
#2Curious to know if VLMs be adapted/extended for video-based tasks (generating video summaries, question answering from video...) by understanding interframe context and temporal dynamics?
Re: Bridging Images and Text – A Survey of VLMs
#3Curious to know if VLMs be adapted/extended for video-based tasks (generating video summaries, question answering from video...) by understanding interframe context and temporal dynamics?
Papers like OneVision have started looking into this. But most of the research is still in nascent stages, answering in one word/phrase for simple questions. I don't even think there's a good enough benchmark dataset to evaluate such models.