Live data from Hacker News

Viewing profile — RobSm

RobSm

HN member
Joined
Fri, Nov 05, 2021, 6:02 PM UTC
HN karma
3
Public activity
27 items

About RobSm

No profile information was provided.

Recent public activity

  1. comment
    Comment #45072168

    You can always stop bots. Add login/password. But people want their content to be accessible to as large audience as possible, but at the same time they don't want that data to be …

  2. comment
    Comment #35798626

    The reason is different. When you need to scan 5 million IPs on 22 port is one thing, but 5 million IPs times 100 or 1000, then you simply run out of resources and since 99% of peo…

  3. comment
    Comment #35798606

    There are at least 1 of us! Looks weird :)

  4. comment
    Comment #35798600

    You need to upgrade your knowledge. Use only [2-9]. That's it.

  5. comment
    Comment #34627589

    Google scrapes everyones data yet you use Google every day and find it useful.

  6. comment
  7. comment
    Comment #33919640

    People don't care about such content and so does Google.

  8. comment
    Comment #32460541

    Explain more about it

  9. comment
    Comment #30014538

    And if I open the website in my browser and then copy from browser to my computer, then all is good?

  10. comment
    Comment #29917712

    So in the scale of google, 'not many' would be some few million per month? And all is good then, right? Even you use their scrapped data probably daily and are totally fine with th…

  11. comment
    Comment #29911657

    And if you manually copy someone's data they worked hard to generate to go and resell, then it's ethical?

  12. comment
    Comment #29911603

    This is so exactly. People do not realize that when they use chrome to view website, chrome is their 'scraper'. And the goal of webs craping is not to get illegal data, but to have…

  13. comment
    Comment #29911545

    No. robots.txt is not something that is defined and enforced by the law. Just because someone came up with some 'recommendation' like robots.txt does not mean this is the law

  14. comment
    Comment #29911515

    If you post that data on a public domain, that is publicly available. It's like writing that info on a cardboard and putting it in the town square and then saying 'why you people s…

  15. comment
    Comment #29911005

    How many contracts google breaches scraping billions of pages every month?

  16. comment
    Comment #29778910

    How one goes about using TLS proxy? Are there any services? You have any links? Thanks.

  17. comment
    Comment #29778424

    What's the point of TLS proxy?

  18. comment
    Comment #29752815

    "and offered it to everyone for free, making a fortune" - this makes no sense

  19. comment
    Comment #29598905

    Thanks, go it. You can delete now

  20. comment
  21. comment
    Comment #29131381

    I get it why someone else scrapes it. But why customers upload data in the first place? Aren't they interested in getting some OTHER data from you and that OTHER data may as well b…

  22. comment
    Comment #29129009

    Makes little sense - customers upload data to you and they don't want any data back? Really?

  23. comment
    Comment #29124087

    Building API is 5 times easier than building routes for your public webpages, which is basically an 'API' as well.

  24. comment
    Comment #29123676

    If you build your site in a way that multiplies each request 10x, well then that's what you get. Don't do that and you won't have issue with requests. Or handle those requests prop…

  25. comment
    Comment #29123641

    Exactly. I am surprised that the 'devs' can't figure out a way to block only annoying/excessive scrapers. Most likely they are just lazy and then just put 3rd party 'solution' and …