Claude can now search the web
81–90 of 758 posts
Re: Claude can now search the web
#82I wonder if it will actually respect the robots.txt this time.
Do really think LLM vendors that download 80TB+ of data over torrents are going to be labeling their crawler agents correctly and running them out of known datacenters?
Re: Claude can now search the web
#83Earlier quoted context omitted.
I don't think it should. If a user asks the AI to read the web for them, it should read the web for them. This isn't a vacuum charged with crawling the web, it's an adhoc GET request.
>You can now use Claude to search the internet to provide more up-to-date and relevant responses. It's a search engine. You 'ask it to read the web' just like you asked Google to, except Google used to actually give the website traffic. I appreciate the concept of an AI User-agent, but without a business model that pays for the content creation, this is just going to lead to the death of anonymously accessible conten…
Did all those old sites have “business models”? What did the web feel like back then?
(This is rhetorical - I had niche hobby sites back then, in the same way some people put out free zines, and wouldn’t give a damn about today’s AI agents so long as they were respectful.
The web was better back then, and I believe AI slop and agents brings us closer to full circle)
Re: Claude can now search the web
#84Earlier quoted context omitted.
In what circles is it a joke? Google bots seem to respect it on my sites according to logs.
I know an artist that had noindex turned on by mistake in robots.txt for the last 5 years - google, kagi and duckduckgo find tons of links relevant to the artist and the artwork but not a single one from the website. so not seem to or apparently but matter of fact like. robots.txt works for the intended audience
Re: Claude can now search the web
#85I'm surprised that they only expect performance to improve for tasks involving recent information. I thought it was widely accepted that using an LLM to extract information from a document is much more reliable than asking it to recall information it was trained on. In particular, it is supposed to lead to fewer instances of inventing facts out of thin air. Is my understanding out of date?
Re: Claude can now search the web
#86Earlier quoted context omitted.
In what circles is it a joke? Google bots seem to respect it on my sites according to logs.
It's in a small circle of those that do. Blame the internet archive for starting this trend: https://blog.archive.org/2017/04/17/robots-txt-meant-for-sea...
Re: Claude can now search the web
#87Earlier quoted context omitted.
robots.txt is meant for automated crawlers, not human-driven actions.
So, do you mean LLMs are human-like and conscious? I thought they were just machine code running on part GPU and part CPU.
Re: Claude can now search the web
#88Earlier quoted context omitted.
if a human triggers the web crawlers by pressing a button, should they ignore robots.txt?
If a human triggers a browser by pressing a button, should it ignore robots.txt?
The entire web was built on the understanding that humans generally operate browsers, and robots.txt is specifically for scenarios in which they do not.
To pretend that the automated reading of websites by AI agents is not something different…is quite a stretch.
Re: Claude can now search the web
#89Do they not care about typical search users? Only developers?