Looks extremely cool. Trying to run it though, I get stuck at "Getting actions for the given objective..." (using the example on the repo)
Using GPT-4 Vision with Vimium to browse the web
11–20 of 133 posts
Re: Using GPT-4 Vision with Vimium to browse the web
#12Re: Using GPT-4 Vision with Vimium to browse the web
#13Re: Using GPT-4 Vision with Vimium to browse the web
#14Is the vision model directly reading the screen and therefore also reading the Vimeo tags? It might be more effective to export the DOM tags and the associated elements as a Json object that is fed into chatGPT without using the vision component
Re: Using GPT-4 Vision with Vimium to browse the web
#15At my work there are a large contingent of people who essentially do manual data copying between legacy programs (govt), because the tech debt is so large that we can't figure out a way to plug these things together. Excited for tools like this to eventually act as a layer that can run over these sort of problems, as bizarre a solution as it is from a compute perspective
Up this this point, these products have been quite brittle. The recent explosion of AI tech seems like quite a boon for this space.
Re: Using GPT-4 Vision with Vimium to browse the web
#16At my work there are a large contingent of people who essentially do manual data copying between legacy programs (govt), because the tech debt is so large that we can't figure out a way to plug these things together. Excited for tools like this to eventually act as a layer that can run over these sort of problems, as bizarre a solution as it is from a compute perspective
Re: Using GPT-4 Vision with Vimium to browse the web
#17At my work there are a large contingent of people who essentially do manual data copying between legacy programs (govt), because the tech debt is so large that we can't figure out a way to plug these things together. Excited for tools like this to eventually act as a layer that can run over these sort of problems, as bizarre a solution as it is from a compute perspective
I remember years ago thinking it was weird in Ghost in the Shell when a robot had fingers on its fingers to type really fast. Maybe that really won’t happen since they can plug into USB at least, but they will probably use the screen and keyboard input sometimes at least.
Re: Using GPT-4 Vision with Vimium to browse the web
#18Hey! Creator here, thanks for sharing! Let me know if anyone has questions and feel free to contribute, I've left some potential next steps in the README.
Re: Using GPT-4 Vision with Vimium to browse the web
#19This is amazing, I feel like these vision models are going to make everything so much more accessible. Between the Be My Eyes app integration and now this, I'm really excited for how this transforms the web.
I agree, and I think we're a year or two away from a full end-to-end trained screen reader. The ground truth from existing systems would provide great training material. As a technical blind person, my only concern is the inherent loss of privacy while sharing stuff with the big models.