Live data from Hacker News

The Web Is Broken – Botnet Part 2

jan.wildeboer.net

281–290 of 301 posts

Re: The Web Is Broken – Botnet Part 2

#281
post #189

Let me get this straight: we want computers knowing everything, to solve current and future problems, but we don't want to give them access to our knowledge?

I don't want computers to know everything. Most knowledge on the internet is false and entirely useless. The companies selling us computers that supposedly know everything should pay for their database, or they should give away the knowledge they gained for free. Right now, the scraping and copying is free and the knowledge is behind a subscription to access a proprietary model that forms the basis of their business.…

That’s factually incorrect. You can use most of these products for free. I use ChatGPT, Perplexity, ClaudeAI, and Gemini every day without paying, and even just these free services have already improved various processes in my life.

I do agree with you on the point that we need to find better ways to compensate the people creating content—especially considering that parts of this "AI service," as we might call it, are subscription-based.

But in the long run, I’m quite sure that if everyone shared this opinion, it wouldn't move us forward technologically.

Also, a couple of other points:

    Google and others have been scraping the internet for years, and no one complained then.

    You're not paying the AI company for the knowledge itself—you're paying for the technology behind it, for the ability to access and use it effectively.

Re: The Web Is Broken – Botnet Part 2

#282
post #189

Let me get this straight: we want computers knowing everything, to solve current and future problems, but we don't want to give them access to our knowledge?

> Let me get this straight: we want computers knowing everything, to solve current and future problems, but we don't want to give them access to our knowledge? Who said that? There's basically two extremes: 1. We want access to all of human knowledge, now and forever, in order to monetise it and make more money for us, and us alone. and 2. We don't want our freely available knowledge sold back to us, with no credits…

1. What exactly is wrong with the first part of that point? I agree that the second part is inaccurate—right now, the money mostly flows in one direction. But as I mentioned earlier, we can use many of these tools for free.

2. You’re not paying just to have your own knowledge echoed back at you. You’re paying so that someone (or something) can read what you provide and, ideally, return improved knowledge or fresh insights. As I said above, you’re paying for the technology and its capabilities—not the knowledge itself. That’s how I see it.

Re: The Web Is Broken – Botnet Part 2

#283
post #189

Let me get this straight: we want computers knowing everything, to solve current and future problems, but we don't want to give them access to our knowledge?

Most people don’t want computers to know everything - ask the average person if they want more or less of their lives recorded and stored.

> Most people don’t want computers to know everything.

That may well be true. But how many of those people are specifically against AI companies scraping the web? That’s not really an argument—it’s an assumption based on personal perception.

> Ask the average person if they want more or less of their lives recorded and stored.

What exactly is the "average person"? Also, I’ll admit my earlier claim was a bit exaggerated. But let’s be clear: this isn’t about recording personal data—it’s about collecting and structuring knowledge.

And beyond that: companies have been scraping the web for years. They still are. And they’re gathering far more personal data for online marketing, tracking, profiling—whatever the reason—and the so-called "average person" hasn’t raised much of a finger. People remain glued to platforms, willingly sharing their personal lives. And what do they get in return? Doomscrolling and five-second video clips.

Re: The Web Is Broken – Botnet Part 2

#284
post #191
post #189

Let me get this straight: we want computers knowing everything, to solve current and future problems, but we don't want to give them access to our knowledge?

I don't want your computer to know everything about me, in fact.

That’s not what I said. Let me rephrase it: Do you want computers to help us solve medical or scientific problems in order to improve human life?

Re: The Web Is Broken – Botnet Part 2

#285
post #264

Earlier quoted context omitted.

Also, I was just watching brodie robertson video about how United Nations has this random search page of unesco which actually has anubis. Crazy how I remember the HN post where anubis's blog post was first made. Though, I always thought it was a bit funny with anime and it was made by frustration of (I think AWS? AI scrapers who won't follow general rules and it was constantly giving requests to his git server and i…

Her* It was frustration at AWS' Alexa team and their abuse of the commons. Amusingly if they had replied to my email before I wrote my shitpost of an implementation this all could have turned out vastly differently.

Oh I am so so sorry I didn't see your gender and assumed it to be a (he). { really sorry about that once again}

Also didn't expect you to respond to my comment xD

I went through the slow realization of while reading this comment that you are the creator of anubis and I had such a smile when I realized that you commented to me.

Also, this project is really nice, but I actually want to ask, I haven't read the docs of anubis but could it be that the proof of work isn't wasted / it can be used for something (I know I might get downvoted because I am going to mention cryptocurrency, but nano currency has a proof of work required for each transaction, so if anubis actually does the proof of work as by nano standards, then theoretically that proof of work could atleast be some useful)

Looking forward to your comment!

Re: The Web Is Broken – Botnet Part 2

#286
post #278

Earlier quoted context omitted.

Wait, why should applications be allowed to do rest calls by default? > What are you talking about? That’s the main use case for p2p in an application isn’t it? Reducing the vendors bandwidth bill…

> That’s the main use case for p2p in an application isn’t it? Reducing the vendors bandwidth bill… The equivalent would be to say that running local workloads or compute is to reduce the vendors bill. It’s a very centralized view of the internet. There are many reasons to do p2p. Such as improving bandwidth and latency, circumventing censorship, improve resilience and more. WebRTC is a good example of p2p used by sm…

Oh, funny you should pick WebRTC. Back when I was still using Chrome, it prevented my desktop from sleeping because 'WebRTC has active peer connections'. With no indication on which page that is happening.

Great respect for the user's resources.

Re: The Web Is Broken – Botnet Part 2

#287

I work for IPinfo (a commercial service). We offer a residential proxy detection service, but it costs money. If you are being bombarded by suspicious IP addresses, please consider using our free service and blocking IP addresses by ASN or Country. I think ASN is a common parameter for malicious IP addresses. If you do not have time to explore our services/tools (it is mostly just our CLI: https://github.com/ipinfo/c…

Blocking countries is such a poorly disguised form of racism. Funny how it's always the brown / yellow people countries that get blocked, and never the US, despite it being one of the leading nations in malicious traffic.

Oh, absolutely not — I have to respectfully but strongly disagree with that sentiment.

In cybersecurity, decisions must be guided by objective data, not assumptions or biases. When you’re facing abuse, you analyze the IPs involved and enrich them with context — ASN, country, city, whether it’s VPN, hosting, residential, etc. That gives you the information you need to make calculated decisions: Should you block a subnet? Rate-limit it? CAPTCHA-challenge it?

Here’s a small snapshot from my own SSH honeypot:

Summary of 1,413 attempts

  - Hosting IPs: 981 (69%)
  - VPNs: 35
  - Top ASNs:
    - AS204428 (SS-Net): 152
    - AS136052 (PT Cloud Hosting Indonesia): 83
    - AS14061 (DigitalOcean): 76
  - Top Countries:
    - Romania: 238 (16.8%)
    - United States: 150 (10.6%)
    - China: 134 (9.5%)
    - Indonesia: 115 (8.1%)
One single /24 from Romania accounts for over 10% of the attacks. That’s not about nationality or ethnicity — it's about IP space abuse from a specific network. If a network or country consistently shows high levels of hostile traffic and your risk tolerance justifies it, blocking or throttling it may be entirely reasonable.

Security teams don’t block based on "where people come from" — they block based on where the attacks are coming from.

We even offer tools to help people explore and understand these patterns better. But if someone doesn’t have the time or resources to do that, I'm more than happy to assist by analyzing logs and suggesting reasonable mitigations.

Re: The Web Is Broken – Botnet Part 2

#288

Earlier quoted context omitted.

I've seen a few attacks where the operators placed malicious code on high-traffic sites (e.g. some government thing, larger newspapers), and then just let browsers load your site as an img. Did you see images, css, js being loaded from these IPs? If they were expecting images, they wouldn't parse the HTML and not load other resources. It's a pretty effective attack because you get large numbers of individual browsers…

I seem to remember a thing china did 10 years back where they injected JavaScript into every web request that went through their Great Firewall to target GitHub… I think it’s known as the “Great Cannon” because they can basically make every Chinese internet user’s browser hit your website in a DoS attack. Digging it up: https://www.washingtonpost.com/news/the-switch/wp/2015/04/10...

Wow, that had passed me by completely, thanks for sharing!

Very similar indeed. The attacks I witnessed where easy to block once you identified the patterns (referrer was visible and they used predictable ?_=... query parameters to try and bypass caches), but very effective otherwise.

I suppose in the event of a hot war, the Internet will be cut quickly to defend against things like the "Great Cannon".

Re: The Web Is Broken – Botnet Part 2

#289
post #278

Earlier quoted context omitted.

> That’s the main use case for p2p in an application isn’t it? Reducing the vendors bandwidth bill… The equivalent would be to say that running local workloads or compute is to reduce the vendors bill. It’s a very centralized view of the internet. There are many reasons to do p2p. Such as improving bandwidth and latency, circumventing censorship, improve resilience and more. WebRTC is a good example of p2p used by sm…

Oh, funny you should pick WebRTC. Back when I was still using Chrome, it prevented my desktop from sleeping because 'WebRTC has active peer connections'. With no indication on which page that is happening. Great respect for the user's resources.

Haha yeah I personally hate WebRTC. It’s a mess and I’ve literally rewritten the parts of it I need in order to avoid it. (Check my profile)

I just brought it up as a technology that at the very least is both legitimate and common.

Re: The Web Is Broken – Botnet Part 2

#290
post #282

Earlier quoted context omitted.

> Let me get this straight: we want computers knowing everything, to solve current and future problems, but we don't want to give them access to our knowledge? Who said that? There's basically two extremes: 1. We want access to all of human knowledge, now and forever, in order to monetise it and make more money for us, and us alone. and 2. We don't want our freely available knowledge sold back to us, with no credits…

1. What exactly is wrong with the first part of that point? I agree that the second part is inaccurate—right now, the money mostly flows in one direction. But as I mentioned earlier, we can use many of these tools for free. 2. You’re not paying just to have your own knowledge echoed back at you. You’re paying so that someone (or something) can read what you provide and, ideally, return improved knowledge or fresh ins…

I'm merely pointing out that there's two separate groups of people.

You appear to be under the impression that there is only one hypocritical group.

Post reply on HN