At my work there are a large contingent of people who essentially do manual data copying between legacy programs (govt), because the tech debt is so large that we can't figure out a way to plug these things together. Excited for tools like this to eventually act as a layer that can run over these sort of problems, as bizarre a solution as it is from a compute perspective
Using GPT-4 Vision with Vimium to browse the web
81–90 of 133 posts
Re: Using GPT-4 Vision with Vimium to browse the web
#82Earlier quoted context omitted.
Funny that you and others on here don't seem to realize that literally everybody who uses the internet has the exact same data entry problem all the time. Blame it on "old software", but how about the entire internet? copying (or in most cases even worse: re-typing) form data from one location on the screen into yet another webform. Username, password, email address, physical address, credit card info etc etc. Some e…
> It's my number 1 frustration when using the internet (worse than ads) and I find it mind-blowing that this hasn't been solved yet with or without LLMs. Simple: it's because not solving this problem is how our godawful industry makes most of its money. Empowering the user means relinquishing control over their "journey"[0]. Ergonomics means fewer opportunities to upsell or show ads. I don't have the link handy, but…
I see that you too have at some point installed printer driver software.
Re: Using GPT-4 Vision with Vimium to browse the web
#83This could enable human language test automation scripts and could either improve my life as a QA engineer a lot or completely destroy it. Not sure yet.
Re: Using GPT-4 Vision with Vimium to browse the web
#84Hey! Creator here, thanks for sharing! Let me know if anyone has questions and feel free to contribute, I've left some potential next steps in the README.
Re: Using GPT-4 Vision with Vimium to browse the web
#85Ah, very similar to Adept’s[1] concept? Though, their product seems not yet ready. [1] https://www.adept.ai/
It's also a little insane to me that what Adept has been supposedly building for years with 300+ mil in funding can now be built in a day with Open AI APIs? I think Adept pivoted along the way but original concept was very similar to this.
Re: Using GPT-4 Vision with Vimium to browse the web
#86At my work there are a large contingent of people who essentially do manual data copying between legacy programs (govt), because the tech debt is so large that we can't figure out a way to plug these things together. Excited for tools like this to eventually act as a layer that can run over these sort of problems, as bizarre a solution as it is from a compute perspective
Yup.
I was briefly part of a decades long effort to migrate off a main frame backend. It was basically a very expensive shared flat file database (eg FileMaker Pro). Used by thousands of applications, neither inventoried or managed. Surely a handful were critical for daily operations, but no one remembered which ones.
And the source data (quality) was filthy.
I suggested we pay some students to manually copy just the bits of data our spiffy "modern" apps needed.
No one was amused.
--
I also suggested we find a suitable COBOL runtime and just forklift the mainframe's "critical" infra into a virtual machine.
No one was amused.
Lastly, I suggested we throttle access to every unidentified mainframe client. Progressively making it slower over time. Surely we'd hear about anything critical breaking.
That suggestion flew like a lead zeppelin.
Re: Using GPT-4 Vision with Vimium to browse the web
#87I'm still going through the source, but really nice idea and great example of enriching the GPT with tools like vimium.
Re: Using GPT-4 Vision with Vimium to browse the web
#88Earlier quoted context omitted.
I totally agree on all points, especially around what AI means for this. I'm kind of in a happy accident situation because I was working on something for RPA, which then became a layer that was factored as its own product, but now might be able to come full circle as a result of AI. Essentially this layer can function as a "delivery medium" for RPA agent creation, that you can use on any device without download. Howe…
I have watched your project for a while as a possible option for embedded browsers for XR applications like WebXR but the high licensing cost was a factor and solutions like Hyperbeam or Vueplex in Unity have been possible. Defiantly agree that multimodal LLM integration is a huge opportunity and multiplayer browsing with AI in realtime is a super cool idea if you package it right.
Regarding pricing we have heard that feedback over time and gradually adjusted our licensing costs. It should now be much more affordable as it is targeted towards large deployments, with decreasing cost and increasing value at scale.
If you'd like to send an email with any thoughts on our current prices on https://dosyago.com to cris@dosyago.com I'd highly value it!
Your idea of WebXR and embedding within Unity is very interesting, and I think it could be a fit.
Re: Using GPT-4 Vision with Vimium to browse the web
#89Earlier quoted context omitted.
It's also a little insane to me that what Adept has been supposedly building for years with 300+ mil in funding can now be built in a day with Open AI APIs? I think Adept pivoted along the way but original concept was very similar to this.
But its too expensive to become practical with the OpenAI API. Also, demo is cool until you see the real-world webpages, then you'll realize that this only works less than %50 of webpages.
Re: Using GPT-4 Vision with Vimium to browse the web
#90Many Dutch companies pay salaries by 1. receiving payslips from the accountant, and then 2. manually initiating bank transfers to each employee for the amount in the corresponding payslip, and then 3. manually initiating a bank transfer to the tax authority to pay the withholded salary taxes. This is completely useless manual labor. There should be no reason for this to be a manual procedure. And yet it's almost impo…