Live data from Hacker News

Using GPT-4 Vision with Vimium to browse the web

github.com

41–50 of 133 posts

Re: Using GPT-4 Vision with Vimium to browse the web

#41
post #15

Earlier quoted context omitted.

The industry buzzword is "Robotic Process Automation", which as a category of products has been focused on using various forms of ML/AI to glue these things together in a common/structured way (in addition to good old fashioned screen scraping). Up this this point, these products have been quite brittle. The recent explosion of AI tech seems like quite a boon for this space.

In the OP's specific instance when would you reach out for a traditional ETL tool vs an RPA solution?

How much does the involvement of a bank of fax machines complicate things?

Re: Using GPT-4 Vision with Vimium to browse the web

#42

At my work there are a large contingent of people who essentially do manual data copying between legacy programs (govt), because the tech debt is so large that we can't figure out a way to plug these things together. Excited for tools like this to eventually act as a layer that can run over these sort of problems, as bizarre a solution as it is from a compute perspective

I believe that LLMs will automate most of our data entry/copy/transformation work. 80% of the world's data is unstructured and scattered across formats like HTML, PDFs, or images that are hard to access and analyze. Multimodal models can now tap into that data without having to rely on complex OCR technologies or expensive tooling. If you go to platforms like Upwork, there are thousands of VAs in low-cost labor count…

I was thinking what the payoff would be to pose as human for these terrible pay click jobs and then assign them to an LLM en masse. There's an arbitrage there ... it may be a good strategy.

I heard recently "click-work" works out to about $4/hr* If you could do that x50, passively, it's a fine income.

* - see https://journals.sagepub.com/doi/full/10.1177/14614448231183... or listen to https://kpfa.org/episode/against-the-grain-october-30-2023/ ... it's a fascinating study. Terrible pay (way below minimum wage) but surprisingly high worker satisfaction. The users seem to view it as entertainment essentially categorizing it as casual gaming.

The "asshole innovator" in me wonders if one could simply make it more entertaining and forego paying the user entirely.

Re: Using GPT-4 Vision with Vimium to browse the web

#43

At my work there are a large contingent of people who essentially do manual data copying between legacy programs (govt), because the tech debt is so large that we can't figure out a way to plug these things together. Excited for tools like this to eventually act as a layer that can run over these sort of problems, as bizarre a solution as it is from a compute perspective

[deleted]

Re: Using GPT-4 Vision with Vimium to browse the web

#44
post #15

Earlier quoted context omitted.

The industry buzzword is "Robotic Process Automation", which as a category of products has been focused on using various forms of ML/AI to glue these things together in a common/structured way (in addition to good old fashioned screen scraping). Up this this point, these products have been quite brittle. The recent explosion of AI tech seems like quite a boon for this space.

In the OP's specific instance when would you reach out for a traditional ETL tool vs an RPA solution?

RPA is for data sources and destinations that are meant for human consumption and entry. So you’d use RPA to take an image of a table and enter every row into a web form.

Re: Using GPT-4 Vision with Vimium to browse the web

#46

Earlier quoted context omitted.

I believe that LLMs will automate most of our data entry/copy/transformation work. 80% of the world's data is unstructured and scattered across formats like HTML, PDFs, or images that are hard to access and analyze. Multimodal models can now tap into that data without having to rely on complex OCR technologies or expensive tooling. If you go to platforms like Upwork, there are thousands of VAs in low-cost labor count…

I was thinking what the payoff would be to pose as human for these terrible pay click jobs and then assign them to an LLM en masse. There's an arbitrage there ... it may be a good strategy. I heard recently "click-work" works out to about $4/hr* If you could do that x50, passively, it's a fine income. * - see https://journals.sagepub.com/doi/full/10.1177/14614448231183... or listen to https://kpfa.org/episode/against…

Interesting. Instead of doing the click work manually, microworkers will just instruct and guide multiple GPTs.

Re: Using GPT-4 Vision with Vimium to browse the web

#47

Earlier quoted context omitted.

I believe that LLMs will automate most of our data entry/copy/transformation work. 80% of the world's data is unstructured and scattered across formats like HTML, PDFs, or images that are hard to access and analyze. Multimodal models can now tap into that data without having to rely on complex OCR technologies or expensive tooling. If you go to platforms like Upwork, there are thousands of VAs in low-cost labor count…

I was thinking what the payoff would be to pose as human for these terrible pay click jobs and then assign them to an LLM en masse. There's an arbitrage there ... it may be a good strategy. I heard recently "click-work" works out to about $4/hr* If you could do that x50, passively, it's a fine income. * - see https://journals.sagepub.com/doi/full/10.1177/14614448231183... or listen to https://kpfa.org/episode/against…

Yeah this seems easy to build but would rather work on making tools that improve accessibility 10x

Re: Using GPT-4 Vision with Vimium to browse the web

#48

Earlier quoted context omitted.

I was thinking what the payoff would be to pose as human for these terrible pay click jobs and then assign them to an LLM en masse. There's an arbitrage there ... it may be a good strategy. I heard recently "click-work" works out to about $4/hr* If you could do that x50, passively, it's a fine income. * - see https://journals.sagepub.com/doi/full/10.1177/14614448231183... or listen to https://kpfa.org/episode/against…

Interesting. Instead of doing the click work manually, microworkers will just instruct and guide multiple GPTs.

maybe. A lot of modern clickwork is actually model training and there is a model-collapse phenomena (https://arxiv.org/abs/2305.17493) which means that it should be banned for such work. I bet a number of clever people on the platforms are already trying to instrument AI to do the work regardless - it's pretty close to "free money" if you can pull it off and not get caught and at a spigot size where there's no real serious consequences if you do.

Re: Using GPT-4 Vision with Vimium to browse the web

#49

At my work there are a large contingent of people who essentially do manual data copying between legacy programs (govt), because the tech debt is so large that we can't figure out a way to plug these things together. Excited for tools like this to eventually act as a layer that can run over these sort of problems, as bizarre a solution as it is from a compute perspective

Whenever I hear about such a thing (people doing legacy system data extraction manually) I wonder if in every case someone got the estimate for the "proper" solution and just decided a bunch of people typing is cheaper?

Integrating things like Chatgpt will still require people who know what they are doing to look at it, and I wouldn't be surprised if the first advice they give is "don't use chatgpt for it".

Re: Using GPT-4 Vision with Vimium to browse the web

#50
How will tools like this affect web tracking or generally advertisements on the internet? Imagine you could have an agent browse the web for you and fetch exactly what you are seraching for without you seeing any ads/pop ups or being tracked along the way! Could be a great ”ad blocker”! Could it perhaps also make SEO useless and thus improve the quality of internet? But I wonder if it also could have negative effects such as the ads being “interweaved” into the fetch content somehow!
Post reply on HN