Live data from Hacker News

Artoo, the client-side scraping companion

medialab.github.io

1–10 of 29 posts

Re: Artoo, the client-side scraping companion

#4
Still not convinced by the reasons offered for client-side scraping. If I'm on my browser, I'm not interested in consuming JSON.

Scraping is really something that's better done in the back end, and today, there are a lot of libraries that let you access web sites from Java and run all the Javascript you need in order to display the page properly.

Re: Artoo, the client-side scraping companion

#5
This is great for simple, quick job. However, you can do only so much in a local browser itself.

I basically built a bookmarklet that let's you define the actions locally on your browser, and then run the scrapes in your own box, essentially allowing unmetered scraping without charging per page.

http://scrape.ly

Re: Artoo, the client-side scraping companion

#6

Still not convinced by the reasons offered for client-side scraping. If I'm on my browser, I'm not interested in consuming JSON. Scraping is really something that's better done in the back end, and today, there are a lot of libraries that let you access web sites from Java and run all the Javascript you need in order to display the page properly.

To each their own. I'm not interested in systematic scraping. I just want to take back, take home the web experience I've had, and be able to digest and work with it latter. The things that I want to work with are the sights and experiences I've had. Client side is perfect.

Second, if I was trying to scrape, I'd rather do scraping with WebDriver than anything else, and injecting some client side scraping tools and using WebDriver as a driver, not a driver/scraper sounds remarkably better.

I see no reason to ever not use a browser to consume html content.

Re: Artoo, the client-side scraping companion

#8

Still not convinced by the reasons offered for client-side scraping. If I'm on my browser, I'm not interested in consuming JSON. Scraping is really something that's better done in the back end, and today, there are a lot of libraries that let you access web sites from Java and run all the Javascript you need in order to display the page properly.

Skeptical at first as well coming from the good ol curl/grep/sed backend scraping world, I changed my mind considering authentication issues and instructions saving: no more need to try and auth on complex websites via phantom without knowing what actually happens, I can just log in and see in my browser what I actually wanna scrape and still rerun it later as a script.

And I just loooove listening to artoo beep over and over ;)

Re: Artoo, the client-side scraping companion

#9

Still not convinced by the reasons offered for client-side scraping. If I'm on my browser, I'm not interested in consuming JSON. Scraping is really something that's better done in the back end, and today, there are a lot of libraries that let you access web sites from Java and run all the Javascript you need in order to display the page properly.

Basically, it makes scraping accessible to almost anyone who can use a browser and write some CSS selectors.

Re: Artoo, the client-side scraping companion

#10
this jquery injection looks kind of dangerous. Looks like code from code.jquery.com is loaded into any page. Say I go to https://secretsquirrel.com and they have been very careful to only load javascript from their own domain but now it can also load malicious javascript from https://code.jquery.com.

it also disable CSP. i'm not exactly sure how the extension works. maybe it is turned on/off on per tab basis and defaults to off which would be quite safe. but if it defaults to on then it can be kind of risky.

Post reply on HN