Live data from Hacker News

Claude can now search the web

anthropic.com

51–60 of 758 posts

Re: Claude can now search the web

#51
post #46
post #33

Earlier quoted context omitted.

It must form the search index somehow. That is prior the human action. Simply it would not find the page at all if it respects.

I remember in late 90s/early 2000 as a teen going to robots.txt to specifically see what they were trying to hide and exploring those urls. What is the difference if I use a browser or a LLM tool (or curl, or wget, etc) to make those requests?

careful, some of those are honey pots or trip wires

Re: Claude can now search the web

#52

Aside, does anyone know of an app like Perplexity for surfing the news in a foreign language (language practice)? Perplexity's "Explore" tab translates its news to your local language, and its curated news items are all pretty interesting, but the problem is that there are so few of them. I seem to get maybe a dozen stories in a day. I paid their subscription for a month just to listen to the news on my walk, but did…

> Apple News would probably do it since they also have good curation, but afaict they still don't support foreign news sources (why???).

ground.news includes sources from all sorts of countries, and also auto-translate headline and the intro, while you can still click to access the source article. Not affiliated, just happy user.

Example with sources in English, German and French: https://ground.news/article/accident-on-the-a13-in-the-yveli...

Although I'm not sure how useful it is for language learning, as you cannot (afaik) configure it to only display articles in Spanish or something similar, but if you filter by stories about France, you'll get a lot of French sources (obviously).

Re: Claude can now search the web

#54
post #32
post #22

Earlier quoted context omitted.

robots.txt is meant for automated crawlers, not human-driven actions.

Every automated crawler follows human-driven actions.

Conversely, every browser is a program that automatically executes HTTP requests.

Re: Claude can now search the web

#56
I’ll be interested in trying it. My admittedly limited experience with this on ChatGPT has been disappointing. ChatGPT falls for the SEO content that has taken over the web.

As an example, I recently travelled abroad to a popular vacationing spot and asked ChatGPT for local recommendations on what to do. When it gave me answers directly, they were pretty solid. But when it “searched the web” instead, the answers were awful. Every single result it suggested had terrible ratings. It did this repeatedly. One of those times I asked it to pick something with better ratings and it sort of improved but not by much.

Of course this is another tool and maybe Claude uses better sources or a better algorithm, but in this case where there was a concrete number tied to the results, that while not perfect, aims to rate the quality of a result, it still did not filter out low quality answers. I’m not sure I trust these LLMs to do any better when there aren’t such ratings available. The available input data is just not very good, and now LLMs are being used to feed that low quality, SEO machine.

Re: Claude can now search the web

#57
post #10

I wonder if it will actually respect the robots.txt this time.

Do really think LLM vendors that download 80TB+ of data over torrents are going to be labeling their crawler agents correctly and running them out of known datacenters?

Re: Claude can now search the web

#60

Earlier quoted context omitted.

almost no one does, robots.txt is practically a joke at this point — right up there with autocomplete=off

In what circles is it a joke? Google bots seem to respect it on my sites according to logs.

A small number of search engines respect it, no one else does. Just about every content scraping bot ignores it, including a number of Google's.
Post reply on HN