Meanwhile... > For robots.txt, we do follow the same restrictions applied to googlebot, otherwise Google benefits from its dominant position. https://community.brave.com/t/stop-website-being-shown-in-br...
would it fall under the cfaa?
61–70 of 327 posts
Meanwhile... > For robots.txt, we do follow the same restrictions applied to googlebot, otherwise Google benefits from its dominant position. https://community.brave.com/t/stop-website-being-shown-in-br...
would it fall under the cfaa?
Just a random thought: While they are already earning money by scraping data, would it not be nice if they pay the site owners a certain amount of the money they earn?
[1]: LLM Engine Optimization
Earlier quoted context omitted.
If a human read a website and profited from the knowledge obtained, I don’t think we’d expect them to pay royalties to the site owner.
A human reads a website, watches/clicks on ad, buys merch, subscribes to website, sets a bookmark to a website, shares an article, invites others etc. What will be the point of sharing knowledge or content if it will no longer be associated to an individual or organization?
Of course if people do conclude that there’s no benefit to sharing knowledge and stop doing so then those who do share knowledge will have an outsized impact on AI training. In the extreme: the opportunity to create truth. Thus the incentive to publish in order to stop those people is created.
What’s the incentive for people to allow the crawler at all? Unlike search engines, chatgpt doesn’t cite references at all (last I tried) or even if it does it often makes up nonexistent references. And because it rephrases the content, there’s often no way to prove they got the material from a particular source, so harder to litigate plagiarism too. How would contributing to the weights of this LLM help content crea…
Friendship ended with SEO. Now LEO [1] is my best friend. [1]: LLM Engine Optimization
What’s the incentive for people to allow the crawler at all? Unlike search engines, chatgpt doesn’t cite references at all (last I tried) or even if it does it often makes up nonexistent references. And because it rephrases the content, there’s often no way to prove they got the material from a particular source, so harder to litigate plagiarism too. How would contributing to the weights of this LLM help content crea…
Friendship ended with SEO. Now LEO [1] is my best friend. [1]: LLM Engine Optimization
I wonder what kind of mischief facts people are going to start sneaking into OpenAI's newer models, by selectively feeding different responses to OpenAI when their crawler is identified.
What do we want to teach it?