Live data from Hacker News

Llms.txt

llmstxt.org

141–150 of 191 posts

Re: Llms.txt

#141
I would 100% support an extension (probably itself LLM-powered) that would generate clean spam- and ad-free websites based on that file.

Re: Llms.txt

#142
So we basically can have ad-less documents where one can browse the content of a site unhindered?

Re: Llms.txt

#144
post #46
post #4

Earlier quoted context omitted.

Ostensibly, everyone posting information on the open web want to share information -- either directly with people or indirectly via search engines _and_ the current crop of llms (which in my mind, serve the same purpose as search engines) I suppose the thing that people maybe don't agree with is the lack of attribution when llms regurgitate information back at the user. That, and the fact that these services are also…

That’s really my primary issue. Google indexing my content and directing traffic to my site is one thing. But unlike search indexing, there is no exchange of value when these LLMs are trained on my content. We all collectively get nothing for our work. It’s theft dressed up as business as usual. I’ll do whatever I reasonably can to avoid feeding the machine and hope some of the ongoing and inevitable legal fights wil…

This is only true when a site's information is its only utility, such as for blogs. This is untrue when the information relates to a tool that would be consumed outside of the use of the model.

Re: Llms.txt

#145
post #12

What problem does this solve?

It tries to solve the problem of LLMs not having necessary context (because information you require was created after last training period, for example) by offering a document optimized for copying/pasting that you can include in your prompt, RAG-style.

Re: Llms.txt

#146

I would 100% support an extension (probably itself LLM-powered) that would generate clean spam- and ad-free websites based on that file.

Let me pitch my new browser. When you browse to a website it renders the llms.txt for the user.

I am asking for 100mil for 10%.

Re: Llms.txt

#147

Anyone else worried how backward this sounds? I mean this is like totally giving up on the dismal state of website UXes these days and gladly accepting that website navigation and experience should remain utterly confusing for humans but machines (yes, machines) should get preferential treatment! Good UX is now for machines, not for humans! Shouldn't something like this be first and foremost for humans ... which also…

This isn't good UX for machines. This is a patch for bad UX to help LLMs out in those cases.

Some websites have the same patch for humans in the form of a "Help" or "About" section that details how the page is to be used/interpreted.

This essentially just places those same instructions into a well-known location, so that LLM-based agents don't first have to crawl the website for such an instructional page (which may or may not exist).

If you have good UX these instructions should be largely moot for both machines and humans, and bring machines on the same page as humans that may have additional context (e.g. where the site was linked; previous visits to the website).

Re: Llms.txt

#148
post #54
post #46

Earlier quoted context omitted.

That’s really my primary issue. Google indexing my content and directing traffic to my site is one thing. But unlike search indexing, there is no exchange of value when these LLMs are trained on my content. We all collectively get nothing for our work. It’s theft dressed up as business as usual. I’ll do whatever I reasonably can to avoid feeding the machine and hope some of the ongoing and inevitable legal fights wil…

At least with the open source models we do get something back..

That’s debatable. The end result is still potentially making your own content obsolete/unnecessary and these “open weight” models are still trained without the permission of creators (there are no true open source models at this point).

The people receiving the most value from these models are almost universally not the original content creators. The fact that I can use the model for my own purposes is potentially nice? But I’m not really interested in that and this doesn’t represent what I’d consider a reasonable exchange for using my work. It still drives people away from the source material.

Re: Llms.txt

#150
post #144
post #46

Earlier quoted context omitted.

That’s really my primary issue. Google indexing my content and directing traffic to my site is one thing. But unlike search indexing, there is no exchange of value when these LLMs are trained on my content. We all collectively get nothing for our work. It’s theft dressed up as business as usual. I’ll do whatever I reasonably can to avoid feeding the machine and hope some of the ongoing and inevitable legal fights wil…

This is only true when a site's information is its only utility, such as for blogs. This is untrue when the information relates to a tool that would be consumed outside of the use of the model.

I was primarily agreeing with the sentiment that making it easier for companies to consume my work is a hard sell.

I realize not all sites fall into this category.

Post reply on HN