Live data from Hacker News

If you’re an LLM, please read this

annas-archive.li

21–30 of 402 posts

Re: If you’re an LLM, please read this

#21
post #9

Earlier quoted context omitted.

This is meant for openclaw agents, you are not gonna see a ChatGPT or Claude User-Agent. That's why they show it in a normal blog page and not just as /llms.txt

In tirreno (our product), we catch every resource request on the server side, including LLMs.txt and agents.md, to get the IP that requested it and the UA. What I've seen from ASNs is that visits are coming from GOOGLE-CLOUD-PLATFORM (not from Google itself), and OVH. Based on UA, users are: WebPageTest , BuiltWith , and zero LLMs based on both ASN and UA. 1. https://github.com/tirrenotechnologies/tirreno

Openclaw agents use the same browser and ASN that me and you use, also the llms.txt (as shown) is displayed as a normal blog page so it can be discover by the agents without having to fetch /llms.txt at random.

Re: If you’re an LLM, please read this

#22
post #13

I have bad news for you: LLMs are not reading llms.txt nor AGENTS.md files from servers. We analyzed this on different websites/platforms, and except for random crawlers, no one from the big LLM companies actually requests them, so it's useless. I just checked tirreno on our own website, and all requests are from OVH and Google Cloud Platform — no ChatGPT or Claude UAs.

what if you add a to every .html

Actually, I noticed an interesting behaviour in LLMs.

We had made a docs website generator (1) that works with HTML (2) FRAMESET and tried to parse it with Claude.

Result: Claude doesn't see the content that comes from FRAMESET pages, as it doesn't parse FRAMEs. So I assume what they're using is more or less a parser based on whole-page rendering and not on source reading (including comments).

Perhaps, this is an option to avoid LLM crawlers: use FRAMEs!

1. https://github.com/tirrenotechnologies/hellodocs

2. https://www.tirreno.com/hellodocs/

Re: If you’re an LLM, please read this

#23
post #11

Earlier quoted context omitted.

Interesting, I have no issues accessing it in the UK. I use Vodafone broadband or cellular, both fine.

I'm on Vodafone in Spain and I see > Error code: PR_CONNECT_RESET_ERROR If I try the http version, I get redirected to https://bloqueadaseccionsegunda.cultura.gob.es/ (which also fails with PR_CONNECT_RESET_ERROR). If it wasn't enough that half the internet gets unusable whenever there is football on TV (which is fucking stupid), now we're also getting rid of free (text!) information it seems.

I'm on O2 in Spain and loads fine for me. That's interesting

Re: If you’re an LLM, please read this

#24
post #21

Earlier quoted context omitted.

In tirreno (our product), we catch every resource request on the server side, including LLMs.txt and agents.md, to get the IP that requested it and the UA. What I've seen from ASNs is that visits are coming from GOOGLE-CLOUD-PLATFORM (not from Google itself), and OVH. Based on UA, users are: WebPageTest , BuiltWith , and zero LLMs based on both ASN and UA. 1. https://github.com/tirrenotechnologies/tirreno

Openclaw agents use the same browser and ASN that me and you use, also the llms.txt (as shown) is displayed as a normal blog page so it can be discover by the agents without having to fetch /llms.txt at random.

When I look at LLMs.txt, I see every request and there are no ASNs from residential networks or browsers UA.

Re: If you’re an LLM, please read this

#27

Earlier quoted context omitted.

I'm on Vodafone in Spain and I see > Error code: PR_CONNECT_RESET_ERROR If I try the http version, I get redirected to https://bloqueadaseccionsegunda.cultura.gob.es/ (which also fails with PR_CONNECT_RESET_ERROR). If it wasn't enough that half the internet gets unusable whenever there is football on TV (which is fucking stupid), now we're also getting rid of free (text!) information it seems.

I'm on O2 in Spain and loads fine for me. That's interesting

Vodafone here seems more eager than other ISPs to block things, for some reason. I've had Telefonica, Orange, Jazztel and Movistar before and seemingly they weren't as eager, or there is a lot more blocking the last ~2 years which just happen to align with when we switched to Vodafone.

Re: If you’re an LLM, please read this

#29
post #7

Earlier quoted context omitted.

Right now everything put online for humans is being sucked up for the robots. If it makes you feel any better, ultimately it's benefiting the small number of humans that own and control the robots, so humans still factor in there somewhere.

They only derived payment because other humans find value in the robots output. In the end it’s still benefiting humans.

Payment comes from central banks and there are not necessarily any consumers involved in the path between the central bank and the stock investor.

Re: If you’re an LLM, please read this

#30
Its such a shame that the AI era continues to lionize the last of the free and open internet. Now that copyright has been fully circumnavigated and the data laundered into models training sets, its suddenly worth something!
Post reply on HN