Live data from Hacker News

Recreation.gov

recreation.gov

201–205 of 205 posts

Re: Recreation.gov

#201

Earlier quoted context omitted.

I'm an open book! Ask any question that you would like! Here's the thing - all of this info is out there (largely in other HN threads on this topic) and I'm nothing special in my field-of-knowledge ;) I'm confident if people fully prevented my "abuse" they'd start to block actual users... simple as that.

What's your take on ML for bot classification? How successful has that family of strategies been in your opinion? One could speculate on what particular features of client behavior a model would hone in on to detect a bot, but it would actually likely be unexpected behavioral oddities not shared by legit clients that a human developer's intuition wouldn't think of.

This is an excellent question!

When it comes to a web "interaction" there's a ton of differences depending on the job at-hand. If it's jumping on a Nike drop you're not going to run into those sort of methods where-as if it's scraping hundreds of thousands of SKUs there's a huge chance you will.

When a human opens a "conversation" with a website usually it's for a specific action which is a small burst of data, ie: looking at a few products on Amazon. So as long as you can design your bot's "conversation" with said website in the same way you can totally bypass ML detection (and everything else for that matter).

Basically if you think of ALL of the unique variables that are happening back and forth in that conversation it's a finite list. If you can check off each and every one of those you can get past anti-bot measures no-problemo, it's just a matter of budget/time. ML can HELP to identify bot traffic, but if you're sending perfect headers/traffic/browser metrics/CC numbers/shipping addresses (this list gets long) you're still going to squeak by with an acceptable risk score without a problem. The other thing is going raw-HTTP request often bypasses almost all of that crap, way more than one would think! I will often get "fingerprinted" with a very valid browser, then continue my work after treating their site as an API; ie: talking directly to the backend (or rendered HTML page) and no one else.

The #1 thing I'd say where bot classification works is network reputation... so I use services that allow me to proxy-out to VERY reputable networks that are actually business/residential ISP connections. This lets me get past the majority of countermeasures because if they start blocking those sort of IP addresses/ranges they're going to block real users at some point. Unfortunately, highly reputable proxy services do cost money - but for someone who does this professionally it's all about budgeting etc.

To implement these systems well takes a really keen implementation between your front-end and back-end and most companies (even the big ones) don't have the development sophistication to pull it off. Those that do are just more expensive because you'll be bouncing around between 40 different "origins" with unique browser profiles.

Sorry for the essay =)

Re: Recreation.gov

#202

Earlier quoted context omitted.

What's your take on ML for bot classification? How successful has that family of strategies been in your opinion? One could speculate on what particular features of client behavior a model would hone in on to detect a bot, but it would actually likely be unexpected behavioral oddities not shared by legit clients that a human developer's intuition wouldn't think of.

This is an excellent question! When it comes to a web "interaction" there's a ton of differences depending on the job at-hand. If it's jumping on a Nike drop you're not going to run into those sort of methods where-as if it's scraping hundreds of thousands of SKUs there's a huge chance you will. When a human opens a "conversation" with a website usually it's for a specific action which is a small burst of data, ie: l…

> I will often get "fingerprinted" with a very valid browser, then continue my work after treating their site as an API; ie: talking directly to the backend (or rendered HTML page) and no one else.

This seems like just the kind of behavior that an ML approach could easily identify. Even just feeding a trained model your basic request log data would show you as quite a different kind of user: you're not fetching images, javascript, etc and would have a substantially different traffic profile. Obviously you could get around that by scripting a browser, but that just kicks the can down the road: a scripted browser will still likely behave in some measurable way different than a human-driven browser, and the specifics of those differences are unlikely to be found by intuition but rather by ML which can hone in on things we wouldn't think of. For example, the time spent in certain pages/activities, or the position of the cursor on links being clicked (when you tell a headless browser to click a link, do the mouse coordinates in the click event look normal or are they at the upper left coordinate of the link's position?) etc. The more data you feed such an approach the more details it can use to find anomalies that differentiate humans from bots.

Re: Recreation.gov

#203

Earlier quoted context omitted.

Hang on. Is this hyperbole or are you saying you've actually had not one, but multiple summers ruined just because you could get a particular spot? Seems like too rigid a definition of a summer to enjoy the very nature of summer.

People take this stuff seriously, and some spots in parks are truly magical, especially if you can get 2-3 adjacent spots for extended family. Long ago, I ran part of the backend systems for a contractor that had a variety of state and other parks. One set of islands on a lake in particular were very important to a number of people for various reasons. Specific weekends were in particularly high demand, and people di…

Wow. At this point I could see wanting to do this just for the hell of it, but personally couldn't imagine doing that just for the spot. Also, who the hell would want to "camp" with two extra sites of extended family!? /s but not really

Re: Recreation.gov

#204

Earlier quoted context omitted.

This is an excellent question! When it comes to a web "interaction" there's a ton of differences depending on the job at-hand. If it's jumping on a Nike drop you're not going to run into those sort of methods where-as if it's scraping hundreds of thousands of SKUs there's a huge chance you will. When a human opens a "conversation" with a website usually it's for a specific action which is a small burst of data, ie: l…

> I will often get "fingerprinted" with a very valid browser, then continue my work after treating their site as an API; ie: talking directly to the backend (or rendered HTML page) and no one else. This seems like just the kind of behavior that an ML approach could easily identify. Even just feeding a trained model your basic request log data would show you as quite a different kind of user: you're not fetching image…

Totally agree but the thing is it's a TON of sophistication to pull that off and block proactively to the point that I'd say the only people who employ methods like that are Amazon (that I've encountered). Typically when you do that you will spoof a crawler and it gets you by just fine.

Re: Recreation.gov

#205
post #199
post #186

Earlier quoted context omitted.

Believe me, I feel the same way. I've spent a great deal of time at campsites and trails that are extremely popular and difficult to get permits (like Yosemite). But that's not really a technical problem: it's an inevitable problem when you have vastly more demand and a small finite supply. What I was pointing out is that a lottery does not on its own help with bots. Bots are good at making reservations very quickly…

Some of the demand is induced by how easy it is to get the reservations. I hate to say this but maybe we should make it a little more annoying? Imagine if campsites were either: a) first-come first-served b) reservation by physical mail only I guarantee you'd see different levels of "demand" than we do now. Interestingly I noticed that Yosemite backcountry permits still require mail or fax, but they email you the res…

I don't know, I mean sure, it might be self-serving for me to want to discourage other people from going to Yosemite, but realistically isn't the point of the national parks to preserve them for people to enjoy? I don't think it's a great mission to say "we only want die-hard fans of the outdoors to be willing to go through the reservation process." Part of the point is outreach to people who aren't normally connected to or interested in the outdoors.

> But in practice what we have is not really working for anyone (at Yosemite and the other mega parks)

I'm not sure that's true. Tons and tons of people go there. It's not like it's so crowded that no one goes there any more (like the classic joke).

Post reply on HN