Live data from Hacker News

Llms.txt

llmstxt.org

21–30 of 191 posts

Re: Llms.txt

#22

This is not how these kinds of things should be designed for the web. Instead of putting resources in the root of the web, this is what /.well-known/ was designed for. See RFC 5785: https://datatracker.ietf.org/doc/html/rfc5785 Instead of munging URLs to get alternate formats, this is what content negotiation or rel=alternate were designed for. I’m not sure making it easier to consume content is something that is nee…

If only this RFC was well-known among the people who actually put stuff out on the Web.

If only that RFC didn't make it a hidden directory.

I can think of a dozen reasons why hiding that folder is a horrible idea, and not a single one for why it would be a good thing to do.

Re: Llms.txt

#23

To disallow: Amazonbot, anthropic-ai, AwarioRssBot, AwarioSmartBot, Bytespider, CCBot, ChatGPT-User, ClaudeBot, Claude-Web, cohere-ai, DataForSeoBot, Diffbot, Webzio-Extended, FacebookBot, FriendlyCrawler, Google-Extended, GPTBot, 0AI-SearchBot, ImagesiftBot, Meta-ExternalAgent, Meta-ExternalFetcher, omgili, omgilibot, PerplexityBot, Quora-Bot, TurnitinBot For all of these bots, User-agent: Disallow: / For more infor…

s/consider of/consider if/

Re: Llms.txt

#24
Can we not put another file in the root please? That's what /.well-known/ is for.

And while I'm here, authors of unix tools, please use $XDG_CONFIG_HOME. I'm tired of things shitting dot-droppings into my home directory.

Re: Llms.txt

#25

This is not how these kinds of things should be designed for the web. Instead of putting resources in the root of the web, this is what /.well-known/ was designed for. See RFC 5785: https://datatracker.ietf.org/doc/html/rfc5785 Instead of munging URLs to get alternate formats, this is what content negotiation or rel=alternate were designed for. I’m not sure making it easier to consume content is something that is nee…

What if each LLM had to register and tell you when/where/what it scraped/ingested from your site/page/url? And you could look at whatever your LLM trigger log looks like, and have a DELETE_ME link. (right_to_be_un_vectorized)

Re: Llms.txt

#26

This is not how these kinds of things should be designed for the web. Instead of putting resources in the root of the web, this is what /.well-known/ was designed for. See RFC 5785: https://datatracker.ietf.org/doc/html/rfc5785 Instead of munging URLs to get alternate formats, this is what content negotiation or rel=alternate were designed for. I’m not sure making it easier to consume content is something that is nee…

If only this RFC was well-known among the people who actually put stuff out on the Web.

OpenAI was good about using well known for plugins

Re: Llms.txt

#28
There’s a deep irony that I have to make a file to help LLMs scrape content while others claim AI will doom humanity.

A few deep ironies actually.

Re: Llms.txt

#29
Actually what is also needed is a notLLMs.txt.

robots.txt exists, but is mainly for crawling and also not sure anyone follows it or even if they don't follow what's the punishment.

Re: Llms.txt

#30

To disallow: Amazonbot, anthropic-ai, AwarioRssBot, AwarioSmartBot, Bytespider, CCBot, ChatGPT-User, ClaudeBot, Claude-Web, cohere-ai, DataForSeoBot, Diffbot, Webzio-Extended, FacebookBot, FriendlyCrawler, Google-Extended, GPTBot, 0AI-SearchBot, ImagesiftBot, Meta-ExternalAgent, Meta-ExternalFetcher, omgili, omgilibot, PerplexityBot, Quora-Bot, TurnitinBot For all of these bots, User-agent: Disallow: / For more infor…

[deleted]
Post reply on HN