Live data from Hacker News

Viewing profile — Ian_Kerins

Ian_Kerins

HN member
Joined
Wed, Apr 01, 2015, 1:50 AM UTC
HN karma
372
Public activity
54 items

About Ian_Kerins

No profile information was provided.

Recent public activity

  1. comment
    Comment #47350062

    A lot of the discussion around the /crawl endpoint seems to miss a key detail in the docs. The crawler explicitly identifies itself as a bot, respects robots.txt, and does not bypa…

  2. comment
    Comment #46765644

    Interesting take on it. Some people probably wouldn't like to be called soft but there is likely some truth to it. I feel it really comes down to priorities. Scraping has always be…

  3. comment
    Comment #46765552

    One of the main ideas, we explored here is how scraping has shifted from being mainly a technical challenge to an economic one: - Infrastructure and proxies have gotten cheaper, bu…

  4. story
  5. comment
    Comment #43715260

    We just dropped the State of Web Scraping 2025 report. TL;DR: scraping is scaling—fast. - Market boom: Web scraping is growing 15% YoY and projected to hit $13B by 2033. Web data i…

  6. story
  7. story
  8. story
  9. comment
    Comment #32282111

    The ethics of these free VPNs and hidden proxy SDKs are very questionable. But they are crazy profitable for the proxy providers running them so unlikely to go away. Did a teardown…

  10. comment
    Comment #32281984

    this proxy comparison tool shows you the best ones https://scrapeops.io/proxy-providers/comparison/

  11. story
  12. story
  13. comment
    Comment #31630206

    It is this type of attitude that is why websites are becoming so aggressive in blocking web scrapers. Being an "ethical web scraper" is about your own ethics, not abusing other peo…

  14. story
  15. comment
    Comment #31611361

    Thanks for sharing it. It currently works with Scrapy & Python Request scrapers, will be launching SDKs for Node, Puppeteer, etc. soon.

  16. comment
    Comment #29918344

    Good point, wouldn't say archiving is unethical at all...I was thinking more along the lines of someone scraping a entire segment of a websites data and reproducing it 1 for 1 on t…

  17. comment
    Comment #29909089

    Some web scraping can be unethical, say for example if you are scraping a site solely to mirror their content and add zero value to the original content owner. However, there are a…

  18. comment
    Comment #29908301

    Interesting!...I'm not a lawyer, so the content for this piece was based on commentary in the below article. Was written by their lawyer, but would love to hear your counter point …

  19. comment
    Comment #29908003

    Haha, nice hack!

  20. comment
    Comment #29907656

    You can do it as a service, but that is highly competitive and basically trading time for money. Best ways are to productize it: - build a on-demand data api for a specific type of…

  21. comment
    Comment #29907131

    If anyone has anything else they think was missed or should be included then let me know!

  22. comment
    Comment #29906968

    100% agree, when scraping it should always be done respectfully. - If they provide a API, then use it. - Don't slam a website, ideally spread it out over hours of the day when ther…

  23. comment
    Comment #29905946

    This has a lot of good info on how to cloudflare and others work, and more creative ways to bypass them if the easier options don't work https://incolumitas.com/2021/05/20/avoid-pu…

  24. story
  25. story