Live data from Hacker News

Show HN: CloudScrape – Cloud-based web scraping platform

cloudscrape.com

61–70 of 109 posts

Re: Show HN: CloudScrape – Cloud-based web scraping platform

#61
It's great the see discussions going on here - would like to tie a few comments to the questions of ethical aspects of web scraping:

As some has pointed out scraping is not exactly a new thing and a lot of the biggest sites out there are built on the basis of web scraping or crawling. We provide a tool and expect you use that tool while abiding the law - and if not we will of course shut your account down immediately. Breaking the law includes violating copyrights and performing DDoS attacks (Although they will be rather small attacks since even 50 concurrent agents is no big deal for most websites).

We consider ourselves good netizens. We wish nothing more than to provide a good, easily accessible and safe tool for extracting valuable information from the internet, be it for a price comparison site in a market that lacks transparency, business intelligence for your company to make informed and wiser decisions, or a PhD project that requires access to millions of data points available online in unstructured form.

Additionally if you feel we're providing services that has ill-intent - we are not providing any services (Captcha and proxy rotation) that anyone with a bit of programming skill can not easily use in their own software. The main difference is that we are actively improving and focusing not only on making a good experience for our users - but also on minimizing the impact on the sites being scraped. This involves several things like automated throttling and slow-site detection, request caching, and blocking requests to services such as google analytics - to not interfere with site owners stats.

Re: Show HN: CloudScrape – Cloud-based web scraping platform

#62
post #20

- Auto-resolve CAPTCHA’s - Automatic IP rotation That's just wrong. Is there a website where we can blacklist IP addresses of such violators ?

That's just wrong. You haven't explained how is it wrong and why. None of those thing are "wrong" by itself. It is the malicious use, of any tool, that is unethical. Is there a website where we can blacklist IP addresses of such violators ? What exactly do you think is being violated here?

Websites use captchas & IP based limits to prevent abuse of their resources & make it harder for copycats to mirror their data. There are often cases where copycats outrank original content in search rankings. (see this example : https://news.ycombinator.com/item?id=10103545 ).

If I were a content owner/producer and I see automated scraping from IP addresses owned by Cloudscrape that violate the ToS, I would sadly treat the entire pool of IPs as violators (even though some might be genuine users who are respecting the limits).

I'd like to know what's a legitimate use case of auto-resolving captchas and IP rotation other than circumventing limits imposed by webmaster.

P.S. Why the throwaway ?

Re: Show HN: CloudScrape – Cloud-based web scraping platform

#63
post #22

Earlier quoted context omitted.

If everyone had such a negative attitude as that we never would have had Google. Countless sites depend on web scraping. You can scrape and be a good netizen, the two are not mutually exclusive.

Google has a unique position in that any site that wants to be found has to let Google bot index its content. Google does not build a derived product from your site's content that ends up competing with your site. If someone needs to build a search site for real estate etc why couldn't they just scrape the Google search result, filter it (white list domains), extract the actual link and present it? In that case you'l…

First of all, Google does take your content and make a product out of it by selling ad space on search result pages. Those pages would have no value for advertisers if it weren't for the content producers Google scraped to fill those pages up with.

Second, you linked to an api page that clearly says it was deprecated 5 years ago, come on. No one should ever feel bad about scraping Google, considering Google is the world's largest scraper themselves.

Re: Show HN: CloudScrape – Cloud-based web scraping platform

#64
post #14

I need to add their IP range to my blocklist.

If these guys are serious about their business they are using proxies to obfuscate who they are. There are many services now that are claiming to have a ton of IP addresses Luminati, Shadio.io and Nohodo are just a few examples.

Sorry, I'm having trouble finding that second link: shadio. Can you confirm that's their address? Interested in checking them out

Re: Show HN: CloudScrape – Cloud-based web scraping platform

#65
post #64
post #14

Earlier quoted context omitted.

If these guys are serious about their business they are using proxies to obfuscate who they are. There are many services now that are claiming to have a ton of IP addresses Luminati, Shadio.io and Nohodo are just a few examples.

Sorry, I'm having trouble finding that second link: shadio. Can you confirm that's their address? Interested in checking them out

Sorry, my bad... it's shader.io

Re: Show HN: CloudScrape – Cloud-based web scraping platform

#66
post #65
post #64

Earlier quoted context omitted.

Sorry, I'm having trouble finding that second link: shadio. Can you confirm that's their address? Interested in checking them out

Sorry, my bad... it's shader.io

Awesome, thank you. They do look interesting, but their pricing confuses me. Are they just selling individual proxies? Or does $1.80 get you single exit nodes or threads basically?

Re: Show HN: CloudScrape – Cloud-based web scraping platform

#67

This can become prohibitively expensive even for a small web scraping job that you're required to do everyday. Say scraping a news site for articles. The tooling and intuitivity is awesome but besides that at the heart doesn't it do what any other headless javascript enabled browser does? For example: https://phantomjscloud.com/site/pricing.html Is there a way to get a less feature rich version of this i.e. sans auto…

Thank your for the kind words :) While you're right that we do run a headless browser-ish thing what sets us apart from most of our competitors, other than the point-and-click approach is that we autodetect everything that's going on in the browser - which also means that in most cases you dont need to know what's going on. What this effectively means is that you'll often spend no time reverse-engineering and be able…

Keep your prices, your average developer is not a business person and advises everyone to race to the bottom.

Re: Show HN: CloudScrape – Cloud-based web scraping platform

#68

It's great the see discussions going on here - would like to tie a few comments to the questions of ethical aspects of web scraping: As some has pointed out scraping is not exactly a new thing and a lot of the biggest sites out there are built on the basis of web scraping or crawling. We provide a tool and expect you use that tool while abiding the law - and if not we will of course shut your account down immediately…

The problem is "the law" is murky on web scraping. For example, did you know that even if your users are only extracting non-copyrighted (even non-copyrightable) data from a page, a judge once ruled that the act of storing the entire page in RAM constituted copyright infringement, since it contained some copyrighted elements that were immediately disposed of after extraction (like the company's logo)? This was Ticketmaster v. RMG Technologies, and it was used against Power Ventures in Facebook's case against them.

Contrast with Feist Publications, Inc. v. Rural Telephone Service Co., where it was ruled that it was legal to copy data from a phonebook and republish, since it was non-copyrightable factual data.

There are several other ridiculous early rulings that were made while the internet was still coming of age, and I think before many judges really understood the way it worked. Recent cases have been bucking these precedents, but you can still get the book thrown at you based on those rulings.

Read about 3Taps and please understand that you will be sued, as they were, unless you fold the moment you get a C&D, which would make your site fairly useless.

Google and all other search engines are illegal in the US in most cases. They just don't get in trouble for most of their activity because people usually want to be on Google. If you end up collecting data in a way that someone doesn't like, things won't go so well for you. See Facebook Inc. v. Power Ventures, Inc.. That guy got raked over the coals; I'm sure Facebook was trying to make an example of him.

Data portability is a threat to the business model of many web incumbents, and that means they want scraping, a critical tool for ensuring that portability, to remain in a nebulous grey area; this allows them to use it for their own purposes (which they often do) and also to try to block people who are using data found on their platform in a way they don't like. This basically results in the bigger company getting their way, because only other multi-billion dollar companies really have the resources to fight against the army of $1k/hr lawyers that public companies hire to try to enforce their opinions on upstarts.

What we really need is serious internet law reform that favors a fair and open platform. Unfortunately, whenever we hear about "internet law reform", it's skewed to the interests of the megacorps who want more tools to shut down innovators that may threaten their business models, not toward creating an open and fair environment for innovation.

Consider, for instance, how ridiculous it would be if every time you opened a book one of the title pages contained a "Terms of Reading" that bound you not to use the information in the book, even the non-copyrightable information, in any way that the book's publisher didn't like, required you to only read the book using the publisher's approved reading methods (perhaps only Oakleys and Ray-bans are publisher-approved eyeglasses, only Herman-Miller publisher-approved seating, and only GE bulbs publisher-approved lighting), required you to agree that you'd never sue the publisher in court but always use private arbitrators that the publisher can easily, even implicitly, buy off, and so forth.

Consider the viability of the argument that you committed copyright infringement by looking at the pages of the book when the author didn't want you to, because the reflection of the content on your eyes constituted an illegal copy.

These things would get laughed out of court, but the digital equivalent is frequently upheld when it comes to online activity.

I think eventually things will stabilize and scraping non-copyrighted data will unambiguously not be a crime, but unfortunately, I think it may still be a few more decades until that happens. I really hope your company is able and willing to help us set the right precedents by committing the tens of millions it will take to win each piece of that stability, since you're set up so perfectly to be the target of several scraping-related lawsuits.

Recent rulings, like QVC v. Resultly and Nguyen v. Barnes and Noble Inc. have been much more positive than former ones, even if they're not altogether ideal, indicating that some magistrates are starting to think of the internet in sensible terms. The rest has to be done through the legislature. Please help make the web safe for data.

IANAL

Re: Show HN: CloudScrape – Cloud-based web scraping platform

#69
post #22

Earlier quoted context omitted.

If everyone had such a negative attitude as that we never would have had Google. Countless sites depend on web scraping. You can scrape and be a good netizen, the two are not mutually exclusive.

Google has a unique position in that any site that wants to be found has to let Google bot index its content. Google does not build a derived product from your site's content that ends up competing with your site. If someone needs to build a search site for real estate etc why couldn't they just scrape the Google search result, filter it (white list domains), extract the actual link and present it? In that case you'l…

A lot of scrapers aren't used to build a competing site. It doesn't matter to most companies. The normal policy is to shut them all down ASAP.

Re: Show HN: CloudScrape – Cloud-based web scraping platform

#70
post #67

Earlier quoted context omitted.

Thank your for the kind words :) While you're right that we do run a headless browser-ish thing what sets us apart from most of our competitors, other than the point-and-click approach is that we autodetect everything that's going on in the browser - which also means that in most cases you dont need to know what's going on. What this effectively means is that you'll often spend no time reverse-engineering and be able…

Keep your prices, your average developer is not a business person and advises everyone to race to the bottom.

Absolutely. People will always tell you they want everything for either practically or actually free. What they say they'll pay and what they'll actually pay are usually two very different things, assuming you provide something that can't be easily replaced.
Post reply on HN