Vid2Seq: A pretrained visual language model for describing multi-event videos
1–10 of 19 posts
Re: Vid2Seq: A pretrained visual language model for describing multi-event videos
#2Re: Vid2Seq: A pretrained visual language model for describing multi-event videos
#3Re: Vid2Seq: A pretrained visual language model for describing multi-event videos
#4Re: Vid2Seq: A pretrained visual language model for describing multi-event videos
#5one of cool google projects that does not take off because its so hard to run and integrate with other cool things...
Re: Vid2Seq: A pretrained visual language model for describing multi-event videos
#6Could test against TV shows and see if it gets an understanding of the social dynamics.
Plus could uncover a lot of the editing technique, I forget what the term is, when they create context from unrelated scenes by cutting from one to the other.
Would also pick up the general plot formula pretty quickly by mapping out the relative intensity and direction (action, tense, playful, romantic) of scenes.
I remember reading about a startup that did this or something similar for TV shows + movies a while back in the New Yorker, the idea was that they could predict how well it would do from a pilot or even the script.
Re: Vid2Seq: A pretrained visual language model for describing multi-event videos
#7one of cool google projects that does not take off because its so hard to run and integrate with other cool things...
Re: Vid2Seq: A pretrained visual language model for describing multi-event videos
#8Had this idea about 5 years ago, seems like it might be viable now, would love to see a video analyzer that creates relationship graphs using sentiment analysis. Ex. who responds to whom, what their tone is, how often. Hopefully with modern methods, it wouldn't take too many examples to pick out the dimensionality of expression (cheating out vs in, loudness + arousal, whining on the spectrum of playful to hurt, confi…
Re: Vid2Seq: A pretrained visual language model for describing multi-event videos
#9Had this idea about 5 years ago, seems like it might be viable now, would love to see a video analyzer that creates relationship graphs using sentiment analysis. Ex. who responds to whom, what their tone is, how often. Hopefully with modern methods, it wouldn't take too many examples to pick out the dimensionality of expression (cheating out vs in, loudness + arousal, whining on the spectrum of playful to hurt, confi…
Do you mean juxtaposition?
Re: Vid2Seq: A pretrained visual language model for describing multi-event videos
#10The remaining examples from the paper don't do much in the way of convincing me this is a one-off issue. The recognition of multiple events globally is good, but perhaps extra care should be taken at overlapping event boundaries (e.g. additional local constraints in the loss/regularization scheme that encourage event splitting or time boundaries "snapping-to-grid" if confidence of co-occurrence is low).