Viewing profile — RobSm
RobSm
HN member- Joined
- Fri, Nov 05, 2021, 6:02 PM UTC
- HN karma
- 3
- Public activity
- 27 items
- HN profile
- View on Hacker News ↗
About RobSm
No profile information was provided.
Recent public activity
-
comment
Comment #45072168
You can always stop bots. Add login/password. But people want their content to be accessible to as large audience as possible, but at the same time they don't want that data to be …
-
comment
Comment #35798626
The reason is different. When you need to scan 5 million IPs on 22 port is one thing, but 5 million IPs times 100 or 1000, then you simply run out of resources and since 99% of peo…
-
comment
Comment #35798606
There are at least 1 of us! Looks weird :)
-
comment
Comment #35798600
You need to upgrade your knowledge. Use only [2-9]. That's it.
-
comment
Comment #34627589
Google scrapes everyones data yet you use Google every day and find it useful.
- comment
-
comment
Comment #33919640
People don't care about such content and so does Google.
-
comment
Comment #32460541
Explain more about it
-
comment
Comment #30014538
And if I open the website in my browser and then copy from browser to my computer, then all is good?
-
comment
Comment #29917712
So in the scale of google, 'not many' would be some few million per month? And all is good then, right? Even you use their scrapped data probably daily and are totally fine with th…
-
comment
Comment #29911657
And if you manually copy someone's data they worked hard to generate to go and resell, then it's ethical?
-
comment
Comment #29911603
This is so exactly. People do not realize that when they use chrome to view website, chrome is their 'scraper'. And the goal of webs craping is not to get illegal data, but to have…
-
comment
Comment #29911545
No. robots.txt is not something that is defined and enforced by the law. Just because someone came up with some 'recommendation' like robots.txt does not mean this is the law
-
comment
Comment #29911515
If you post that data on a public domain, that is publicly available. It's like writing that info on a cardboard and putting it in the town square and then saying 'why you people s…
-
comment
Comment #29911005
How many contracts google breaches scraping billions of pages every month?
-
comment
Comment #29778910
How one goes about using TLS proxy? Are there any services? You have any links? Thanks.
-
comment
Comment #29778424
What's the point of TLS proxy?
-
comment
Comment #29752815
"and offered it to everyone for free, making a fortune" - this makes no sense
-
comment
Comment #29598905
Thanks, go it. You can delete now
- comment
-
comment
Comment #29131381
I get it why someone else scrapes it. But why customers upload data in the first place? Aren't they interested in getting some OTHER data from you and that OTHER data may as well b…
-
comment
Comment #29129009
Makes little sense - customers upload data to you and they don't want any data back? Really?
-
comment
Comment #29124087
Building API is 5 times easier than building routes for your public webpages, which is basically an 'API' as well.
-
comment
Comment #29123676
If you build your site in a way that multiplies each request 10x, well then that's what you get. Don't do that and you won't have issue with requests. Or handle those requests prop…
-
comment
Comment #29123641
Exactly. I am surprised that the 'devs' can't figure out a way to block only annoying/excessive scrapers. Most likely they are just lazy and then just put 3rd party 'solution' and …