Live data from Hacker News

Top DNS domains seen on the Quad9 recursive resolver array each day

github.com

51–60 of 100 posts

Re: Top DNS domains seen on the Quad9 recursive resolver array each day

#51

Earlier quoted context omitted.

Probably some sort of command and control for a botnet. They calculate a random domain name based on the timestamp (so it’s constantly changing every X days in case it gets seized), and have some validation to make sure commands are signed (to prevent someone name squatting to control their botnet).

Wow, that's smart. I was wondering whether there is a way for the bots to generate "unpredictable" domains such that security researchers could not predict them efficiently (even with source code), but the botnet controller can. Time-lock puzzles come close, but but it requires that the bots have computing power comparable to the security researchers.

there are tools pretty good at detecting DGAs these days, but not often implemented.

the best thing to do afaik is use services normal user shave access to, and communicate via those. its hard to tell for anyone who's extracting the data from the third party so the server is hidden. (e.g bot posts images to twitter, and server scrapes the images from twitter, this is also already old news but easier and more likely to sail through that next gen firewall -_-)

i'd say having ur 'own' servers and domains is maybe even a bit dated ( though sadly still very effective!)

Re: Top DNS domains seen on the Quad9 recursive resolver array each day

#52
post #21

Earlier quoted context omitted.

Most likely something like an ad service to prevent their content being caught by domain blocklists. That would be similar to how a lot of websites started using randomized strings for attributes like id and class so that users couldn't block page elements based on CSS selectors.

Interesting how ad services and botnets behave similarly in some aspects

Cue in "Are we the baddies?" meme.

Re: Top DNS domains seen on the Quad9 recursive resolver array each day

#53
post #4

> https://github.com/Quad9DNS/quad9-domains-top500/blob/main/t... {"position": 5, "domain_name": "kxulsrwcq.com", "date": "2025-07-10"} What the https://www.ipaddress.com/website/kxulsrwcq.com/ > Safety/Trust: Unknown

And they are often used with random sub domains as well (but they did not include sub domains in their list).

Ex:

https://dnsarchive.net/search?q=cmidphnvq.com

https://dnsarchive.net/search?q=xmqkychtb

https://dnsarchive.net/ipv4/34.126.227.30

Re: Top DNS domains seen on the Quad9 recursive resolver array each day

#54
post #28
post #22

Earlier quoted context omitted.

It does say that they collect this information in their “Data and Privacy Policy”. Specifically section 2.2 (Data Collected): https://quad9.net/privacy/policy/ Which policy are you referring to that implies they don’t? Also I think you are assuming they store query logs and then aggregate this data later. It is much simpler just to maintain an integer counter for monitoring as the queries come in, and ingest that int…

The section you mentioned does not say anything about having counters for labels. It only mentions that they record "[t]he times of the first and most recent instances of queries for each query label".

Well, the counters aren't data collected, they are data derived from the data they do collect. The privacy policy covers collection.

EDIT: I see they went out of their way to say "this is the complete list of everything we count" and they did not include counters by label, so I see your point!

Re: Top DNS domains seen on the Quad9 recursive resolver array each day

#55

It's quite interesting to me that ChatGPT is in the 200s and 300s. By almost every metric this is one of the 10 busiest websites, and some sources are already putting it in the top 5. Are they just disproportionately not using Quad9? I understand that there's a lot of overlap with Google having several spots in the top 50 itself, several being infrastructure like cloudflare and akamai, and several others being malwar…

My theory: the domains you name have ad beacons, desktop apps that are persistently running, and/or physical devices plugged into networks out there. Whereas ChatGPT is used (domainwise) overwhelmingly by humans hitting the site in their browsers.

Re: Top DNS domains seen on the Quad9 recursive resolver array each day

#56
post #46

Earlier quoted context omitted.

Wow, that's smart. I was wondering whether there is a way for the bots to generate "unpredictable" domains such that security researchers could not predict them efficiently (even with source code), but the botnet controller can. Time-lock puzzles come close, but but it requires that the bots have computing power comparable to the security researchers.

> Wow, that's smart. I was wondering whether there is a way for the bots to generate "unpredictable" domains such that security researchers could not predict them efficiently (even with source code), but the botnet controller can. There is a fairly simple method which achieves the same advantage for a botnet controller. 1. Use a hash of the current day to derive, for that day, an infinite stream of domain names. This…

Here's the same image on a less horrible file hosting:

https://files.catbox.moe/gilmd1.png

Imgur has been inaccessible for me for months, they're one of those organizations that consider it proper to block whole countries to counter bot abuse.

Re: Top DNS domains seen on the Quad9 recursive resolver array each day

#57

It's quite interesting to me that ChatGPT is in the 200s and 300s. By almost every metric this is one of the 10 busiest websites, and some sources are already putting it in the top 5. Are they just disproportionately not using Quad9? I understand that there's a lot of overlap with Google having several spots in the top 50 itself, several being infrastructure like cloudflare and akamai, and several others being malwar…

Mostly because of sub domains. They are counting all the sub domains requests to give the top domains ranking.

Some of those have many trackers and background sub domains that add up.

For example, Linkedin their most popular sub domain is: px.ads.linkedin.com

Here is a more comprehensive list with top 10k domains (including sub domains):

https://dnsarchive.net/top-domains?rank=top10k

Re: Top DNS domains seen on the Quad9 recursive resolver array each day

#58

Earlier quoted context omitted.

Wow, that's smart. I was wondering whether there is a way for the bots to generate "unpredictable" domains such that security researchers could not predict them efficiently (even with source code), but the botnet controller can. Time-lock puzzles come close, but but it requires that the bots have computing power comparable to the security researchers.

I can see a future where Cloudflare or similar offer a DNS + proxy + Root CA combo to intercept these. Maybe they already do.

If I’m remembering correctly, Conficker was the first major use of this technique. They used a relatively small domain pool (250) so the registries were able to lock them up preemptively.

I remember a couple legitimate sites getting slammed by accidental DDOS because the algorithm happened to generate their domain, but having a hard time finding a reference to that.

https://en.m.wikipedia.org/wiki/Conficker

Re: Top DNS domains seen on the Quad9 recursive resolver array each day

#59

It's quite interesting to me that ChatGPT is in the 200s and 300s. By almost every metric this is one of the 10 busiest websites, and some sources are already putting it in the top 5. Are they just disproportionately not using Quad9? I understand that there's a lot of overlap with Google having several spots in the top 50 itself, several being infrastructure like cloudflare and akamai, and several others being malwar…

Something else to factor in is the TTL of both NS/A types for each apex domain and the individual records including sub-domains. Clients will not be querying Quad9 until the TTL expires on their clients. TTL would have to be factored into query rates to determine popularity correctly whereas these lists just show raw query numbers.

For example, there are many records under amazonaws.com that have 5 second TTL's mostly EC2 instances. As such clients will query them at a much higher rate whereas grammarly.io have a number of records with a 900 second TTL. This will skew the ranking positions of the two apex domains. I suppose if one wanted to game this they could have an A record to a non-critical part of a site that is not visibly rendered by the end-user and has a TTL of 1 second assuming quad9 is not rewrite min/max-ttl which some resolvers do.

Examples of just some of the TTL's used on these apex domains excluding individual records:

    30 32 60 300 600 900 1200 1800 3600 7200 10800 21600 28800 43200 86400 90000 3600000
Some examples of rewriting max-ttl I forgot which ones rewrite min-ttl:

    for Resolver in 1.1.1.1 8.8.8.8 9.9.9.9 216.128.176.142;do echo -en "${Resolver}:\t"; dig @${Resolver} +nocookie +noall +answer -t a big.ohcdn.net;done | column -t
    1.1.1.1:          big.ohcdn.net.  3628800  IN  A  227.227.227.227
    8.8.8.8:          big.ohcdn.net.  21422    IN  A  227.227.227.227
    9.9.9.9:          big.ohcdn.net.  43200    IN  A  227.227.227.227
    216.128.176.142:  big.ohcdn.net.  3628800  IN  A  227.227.227.227  # authoritative server
[Edit] I just realized they made a general statement to this effect in the git repo.

Re: Top DNS domains seen on the Quad9 recursive resolver array each day

#60

Earlier quoted context omitted.

Probably some sort of command and control for a botnet. They calculate a random domain name based on the timestamp (so it’s constantly changing every X days in case it gets seized), and have some validation to make sure commands are signed (to prevent someone name squatting to control their botnet).

Wow, that's smart. I was wondering whether there is a way for the bots to generate "unpredictable" domains such that security researchers could not predict them efficiently (even with source code), but the botnet controller can. Time-lock puzzles come close, but but it requires that the bots have computing power comparable to the security researchers.

Use a hash chain!

Each time you resolve, the resulting IP can be part of the hash for predicting a future hostname.

Post reply on HN