Live data from Hacker News

Claude can now search the web

anthropic.com

81–90 of 758 posts

Re: Claude can now search the web

#82
post #57
post #10

I wonder if it will actually respect the robots.txt this time.

Do really think LLM vendors that download 80TB+ of data over torrents are going to be labeling their crawler agents correctly and running them out of known datacenters?

Apparently they use smart appliances to scrape websites from residential accounts.

Re: Claude can now search the web

#83
post #39

Earlier quoted context omitted.

I don't think it should. If a user asks the AI to read the web for them, it should read the web for them. This isn't a vacuum charged with crawling the web, it's an adhoc GET request.

>You can now use Claude to search the internet to provide more up-to-date and relevant responses. It's a search engine. You 'ask it to read the web' just like you asked Google to, except Google used to actually give the website traffic. I appreciate the concept of an AI User-agent, but without a business model that pays for the content creation, this is just going to lead to the death of anonymously accessible conten…

What was the web like before wide spread internet ads, auth, and search engines?

Did all those old sites have “business models”? What did the web feel like back then?

(This is rhetorical - I had niche hobby sites back then, in the same way some people put out free zines, and wouldn’t give a damn about today’s AI agents so long as they were respectful.

The web was better back then, and I believe AI slop and agents brings us closer to full circle)

Re: Claude can now search the web

#84

Earlier quoted context omitted.

In what circles is it a joke? Google bots seem to respect it on my sites according to logs.

I know an artist that had noindex turned on by mistake in robots.txt for the last 5 years - google, kagi and duckduckgo find tons of links relevant to the artist and the artwork but not a single one from the website. so not seem to or apparently but matter of fact like. robots.txt works for the intended audience

AI crawlers are part of the intended audience.

Re: Claude can now search the web

#85
> With web search, Claude has access to the latest events and information, boosting its accuracy on tasks that benefit from the most recent data.

I'm surprised that they only expect performance to improve for tasks involving recent information. I thought it was widely accepted that using an LLM to extract information from a document is much more reliable than asking it to recall information it was trained on. In particular, it is supposed to lead to fewer instances of inventing facts out of thin air. Is my understanding out of date?

Re: Claude can now search the web

#86

Earlier quoted context omitted.

In what circles is it a joke? Google bots seem to respect it on my sites according to logs.

It's in a small circle of those that do. Blame the internet archive for starting this trend: https://blog.archive.org/2017/04/17/robots-txt-meant-for-sea...

Given websites do disappear or worse, get their content adultered. Also given the long history of the internet archive as a non profit - and the commons service it has served so far, the joke would be to see that bot honor it.

Re: Claude can now search the web

#87
post #22

Earlier quoted context omitted.

robots.txt is meant for automated crawlers, not human-driven actions.

So, do you mean LLMs are human-like and conscious? I thought they were just machine code running on part GPU and part CPU.

There’s a human using the LLM. In a live web browsing session like this, the LLM stands in for the browser.

Re: Claude can now search the web

#88

Earlier quoted context omitted.

if a human triggers the web crawlers by pressing a button, should they ignore robots.txt?

If a human triggers a browser by pressing a button, should it ignore robots.txt?

Are you arguing that these are equivalent actions?

The entire web was built on the understanding that humans generally operate browsers, and robots.txt is specifically for scenarios in which they do not.

To pretend that the automated reading of websites by AI agents is not something different…is quite a stretch.

Post reply on HN