Live data from Hacker News

Detecting Chrome headless, the game goes on

antoinevastel.com

121–130 of 143 posts

Re: Detecting Chrome headless, the game goes on

#121

Earlier quoted context omitted.

This sounds, at best, ethically dubious and at worst illegal. Aaron Swartz was arrested and charged under hacking laws for doing exactly what you're describing. Given that your run this division there is a good chance you are personally liable.

We have an enormous legal team that communicates constantly with end points to ensure they are aware of our scraping. And as I said in another comment, we store no results other than what is already available to anyone else using the web. We've had this division for many many years, and before my time we paid another company to do this. There's no legal issues.

Your legal teamn is in contact with them, but their security is actively trying to block you? That doesn't make sense.

Computer security laws are very broad. It doesn't matter if it's just a website that the public can access. If you're accessing it in a matter that they don't want AND you're aware of that, then I struggle to see how your lawyers can justify it.

> Computer hacking is broadly defined as intentionally accesses a computer without authorization or exceeds authorized access.

https://definitions.uslegal.com/c/computer-hacking/

Hiding your user agent because you know they don't want automated retrieval of information is "without authorisation".

Re: Detecting Chrome headless, the game goes on

#122
post #71
post #25

Earlier quoted context omitted.

I agonize about this every day, since a large part of my job is aggregating data from many sites that seem hell-bent on not letting anyone access it without going one-form-at-a-time through their crap UI. The thing is, we would gladly pay these companies for an API or even just a periodic data-dump of what we need. We've even offered to some of them to write and maintain the API for them. They're not interested, for…

Sometimes when I'm thinking about it and what 95% of developers are working on, it feels like a planet-wide charity project against unemployment.

I think it's a fairly well known thing that 'junk jobs' tend to spring up in response to supply. I think it's a bizzare cultural thing.

Re: Detecting Chrome headless, the game goes on

#123
post #106

Earlier quoted context omitted.

This sounds, at best, ethically dubious and at worst illegal. Aaron Swartz was arrested and charged under hacking laws for doing exactly what you're describing. Given that your run this division there is a good chance you are personally liable.

>Aaron Swartz was arrested and charged under hacking laws for doing exactly what you're describing. Don't think connecting a computer to a private network to suck up subscriber data is comparable to scraping publicly accessible internet content.

These fear mongering comments always ignore the notice provision in the CFAA. Web scraping publicly accessible information is not "illegal" under the CFAA. That law, at most, only makes someone who continues scraping after being asked to stop potentially culpable.

First, the accuser needs to, at least, send a cease and desist letter to the accused asking them to stop accessing the protected computer. Second, the accused needs to ignore that request and keep accessing the protected computer.

Is it possible to build a solid CFAA case when those two things do not happen? I cannot find any examples.

https://iapp.org/news/a/can-a-cease-and-desist-notice-create...

Re: Detecting Chrome headless, the game goes on

#124

Earlier quoted context omitted.

The right thing to do would be to reach out to those sites and see if they is they have paid options for getting the data you need.

And what happens when they ignore you? I've reached out to tons of website operators to ask for machine readable access to their data on academic, personal and professional projects, I have never gotten a reply and had to resort to scraping.

I can second this for public records websites.

A previous company I worked for aggregated publicly recorded mortgage data. The mortgage data was scraped from municipal sites on a nightly basis because it was not available as a bulk download or purchasable option.

We had requested on several occasions for a service we could pay for in order to get a bulk download of this data, but the municipalities did not have the know how to provide this as were using systems from a private vendor that were prohibitively expensive for them to request modifications. As a result, we worked hand in glove with the municipalities to ensure we were not stressing their infrastructure when we did this scraping, and I think that's the best we were able to do in this case.

Re: Detecting Chrome headless, the game goes on

#125

Earlier quoted context omitted.

That's the general direction I'd like to take. When we capture the inputs for the scrapers, I'd like to persist everything. Mouse jiggles, delays, idle time. I think it would definitely help advance the software.

In the grand scheme of things all of this is a wasteful process. Maybe you could direct your worklife towards other challenges that are more rewarding for society and equally profitable?

That crosses into personal attack. Please don't do that on Hacker News. We've had to ask you this before.

https://news.ycombinator.com/newsguidelines.html

Re: Detecting Chrome headless, the game goes on

#126
post #2

This would be more interesting if the author explained this technique. People that are knowledgeable enough will deep dive into the webpage, but for everyone else, expect disappointment.

Well, if I was him I guess I would prefer to have an accepted paper about this new technique before releasing everything to the public.

Re: Detecting Chrome headless, the game goes on

#127
post #66

Earlier quoted context omitted.

> Where do I start? Study the chromium source? I'm curious why you'd jump straight to browser detection as the most likely culprit. When I was doing scraping, the far more common case was bot detection by origin and access patterns. It's just very difficult to make an automated scraper look like a residential or business user. Where do you run your scraping operation? Is it in AWS or some other hosting provider, beca…

According to the NDA with my company I can't reveal anything about the architecture beyond the fact that it is hosted locally on a homebuilt distributed system that randomly chooses from a pool of 120 residential IPs. We do have human emulation routines that helped avoid most detection, and that library is decoupled in such a way that we can edit behavior down to the individual site. Some sites are just so damn good…

A pool of 120 residential ips is way too small - patterns are more emergent. Go for thousands, even better, hundreds of thousands. Outsource the residential proxy system to luminati or oxylabs.

Re: Detecting Chrome headless, the game goes on

#128

Earlier quoted context omitted.

I don't empathize with your viewpoint because, whether it's a web scraper, or a person, the work is exactly the same. There's no additional volume, or extra steps. We just emulate a worker. We measure the value in FTEs, and when a researcher quits, we do not replace them if the appropriate FTEs have been reached with projects. It's a major benefit to the business not only because we don't have to pay another employee…

Sadly, you are correct to have realized that many posters on HN are so naive that they will offer you $0/hour consulting for your for-profit business. Posting on the HN forums means you "don't have to pay another employee" that's an expert in the field. I can't do much to prevent this, but I don't much respect it, either.

>Posting on the HN forums means you "don't have to pay another employee" that's an expert in the field. I can't do much to prevent this

Sometimes the answer tells you much more about what skills you need to be hiring. Sometimes they give you a lead.

Re: Detecting Chrome headless, the game goes on

#129
I have always maintained a high credit score over the years up till a messy job divorce and property split which affected me psychologically and financially. I accumulated many negative collections (repossession, charge offs, late payments) and my credit score suffered for it and did stop me from getting loans. I got to hear of REDEMPTIONHACKERSCREW from an old college mate and also saw success stories of people he had helped increase their credit scores. I contacted him via his email REDEMPTIONHACKERSCREW, we got the ball rolling and in 5 days his magic hand was all over my report when I pulled it. 790–800 on all bureaus. To be honest it Still feels like a dream getting the loans. If you wish to be saved like me and others reach out to REDEMPTIONHACKERSCREW contact email: REDEMPTIONHACKERSCREW at GMAIL dot COM or text REDEMPTION with your request to +1 909 375 5075 (WHATAPP ONLY)

Re: Detecting Chrome headless, the game goes on

#130

All this can be avoided (from a scraper's perspective) by using the Selenium IDE++ project. It adds a command line to Chrome and Firefox to run scripts. See https://ui.vision/docs#cmd and https://ui.vision/docs/selenium-ide/web-scraping => Using Chrome directly is slower, but undetectable .

I am using the UI Vision extension for a few months now. It is not very fast, but it always works. It can extract text and data from images and canvas elements, too.
Post reply on HN