Live data from Hacker News

Brave Search launches own image and video search

brave.com

111–120 of 172 posts

Re: Brave Search launches own image and video search

#111
post #102

Earlier quoted context omitted.

A publisher publishes. Once something is published, once something is public, the control a publisher has over the published thing is limited. For example, a publisher can not choose who reads a book after it is sold, who reads an article after it got printed. It is not a given at all that a publisher should have any say about being indexed. The search engine relies on a public fact - X wrote Y. That's legal (limitat…

A search engine index is an economic exchange between the website and the publisher. To massively (over)simplify the argument to its essence (and ignore other important points): the publisher goes through the trouble and expense of creating the content The publisher then allows its content to be copied by a search engine only because being shown in search results gets it traffic back. The traffic it gets in return ha…

> A search engine index is an economic exchange between the website and the publisher.

A search engine index is a search engine index. It can have an economic impact, but it can not be an economic exchange, since it is a technical artifact.

Though I think I understand what you are trying to say - this is also a commercial relationship where both sides can profit. You are free to interpret the relationship between publishers and search engines with such a capitalist lense, that does not mean those mechanisms govern the actual rules. That a publisher is happy with what happens here is of no real concern. If any rules apply we are talking copyright, maybe media law, where happy is not a relevant category (ok, that of course can matter, it wouldn't here if the search engine uses a right).

I did not touch the LLM training data in my comment, as I did not read up on what Brave is really doing there. If Brave were really to sell complete texts from others, that would not be legal under copyright laws I'm familiar with, so I kinda doubt they do that.

Re: Brave Search launches own image and video search

#112

Earlier quoted context omitted.

This make me want to use Brave search now. When I use a tool I expect it that it serves me, not the material it provides. > A publisher should have full control to discriminate which search engine indexes the website's content If you want someone to not see what you publish block him yourself. Also why would you want to do that? Do you want google to own the web or something?

There is a a difference between a human being able to access content vs a search engine indexing it (and in the case of Brave, "licensing" it on). I share your concern about Google having this much power, and I'd add that Microsoft Bing is equally bad but gets away with it because they're smaller. Still, the final decision about which search engine indexes a website is purely the publisher's.

> There is a a difference between a human being able to access content vs a search engine indexing it

Much of the problem with search today arises from websites showing googlebot what it wants to see and showing real users. I have to manually remove entire domains from google search as they often appear 1st yet don't show any content without me signing up for an account. Clearly that's not what they are showing to google.

There should be no differentiation between a crawler and a human being with regards to what is being served.

Re: Brave Search launches own image and video search

#113

The major problem with Brave search is their position about indexing and licensing content against the wishes of the website publisher. Their robot does not identify itself, meaning the publisher cannot use the standard robots.txt to block its crawling if the publisher so wishes. Incidentally, the robots.txt file has been used in court cases litigating if a search engine is legal or not. Even worse, they state that B…

Let's say I pay for Kagi. Kagi is a tool that I'm using to avoid doing hard work manually. With relatively few exceptions, I can probably accomplish what I use a search engine for manually, but with much more time and effort. So I'm paying for a tool to assist me with my use of the web. A "user agent", you might even say.

It simply doesn't sound right to say which tool a user can use. It's literally the same as arguing that you should be able to block Firefox from accessing your website and it's Mozilla's fault that they don't respect your wishes as a webmaster to block Firefox exclusively. Or that a VPN doesn't publish its IP addresses so that you can block it. Or a screen reader that processes the text to speech in a way that you disagree with.

Philosophically it seems intuitive to say "I should be able to block a third party that is abusing my site" but it's ignoring the broader context of what "open web" and "net neutrality" actually mean.

I run a service for podcasters. There are podcast apps and directories that either ignorantly make unnecessary requests for content or have software bugs that cause redownloads. I could trivially block them, but I don't because doing so penalizes the end user who is ultimately innocent, rather than the badly behaved service operator. The better solution is primitives like rate limiting, which I use liberally. Plus, blocking anyone literally has a direct effect of incentivizing centralization on Apple, Spotify, etc. and making the state of open tech in podcasting even worse.

> the Brave search API allows you (for an extra fee) to get the content with a "license" to use the content for AI training? Who allowed them the right to distribute the content that way?

I don't think there's any court at this point that would back you up that freely published content annotated with full provenance cannot be scraped and published for a fee. Services like this have existed for decades. If you don't want your content scraped, put it behind a login. Especially considering this only applies when you allow other search engines and if you think Google and Bing aren't using your content to train AI, you're off your rocker.

Re: Brave Search launches own image and video search

#114
post #34

As far as I can tell, that makes for 5 independent image search engines on the web: Baidu Bing Brave Google Yandex You can compare their results on this search comparison page I maintain: https://www.gnod.com/search/?engines=p,o,br,n,q&nw=1 (If you want to also search image libraries like Flickr and Pexels, click on "more engines" to select all places you want to search)

does searx count as independent?

Re: Brave Search launches own image and video search

#115

The major problem with Brave search is their position about indexing and licensing content against the wishes of the website publisher. Their robot does not identify itself, meaning the publisher cannot use the standard robots.txt to block its crawling if the publisher so wishes. Incidentally, the robots.txt file has been used in court cases litigating if a search engine is legal or not. Even worse, they state that B…

Let's say I pay for Kagi. Kagi is a tool that I'm using to avoid doing hard work manually. With relatively few exceptions, I can probably accomplish what I use a search engine for manually, but with much more time and effort. So I'm paying for a tool to assist me with my use of the web. A "user agent", you might even say. It simply doesn't sound right to say which tool a user can use. It's literally the same as argui…

> With relatively few exceptions, I can probably accomplish what I use a search engine for manually, but with much more time and effort. So I'm paying for a tool to assist me with my use of the web. A "user agent", you might even say

1. User agents should identify themselves

2. A crawler is not a User agent - it's an agent for Brave

>I don't think there's any court at this point that would back you up that freely published content annotated with full provenance cannot be scraped and published for a fee.

You can't end-run copyright like this: just because something is publicly available doesn't mean anyone can redistribute it. Look at the legal issues & cases relating to Library Genesis.

Re: Brave Search launches own image and video search

#116

Earlier quoted context omitted.

Brave sees Mozilla's political shitfest and raises a political shitbacchanalia.

How so? If you're thinking about Eich's politics, those are/were HIS, not Brave or Mozilla the organization's. Brave the organization's politics are pretty narrow and what you'd want out of a browser: Privacy and user control. Meanwhile Mozilla the organization boots people for politics, wants "more than deplatforming" and uses as one of their examples organizations deciding what I should see on the Internet - prefer…

Brave the organization runs and integrates with anarchocapitalist cryptoscams. That's far beyond anything Mozilla the organization has done, and if it's not worse in magnitude by itself, that's only because Brave thankfully remains a miniscule also-ran. Writing copy about sponsoring Web3 gaming expos is far stupider than writing about getting a sneaker designer to add some browser themes.

But yes, Eich and his horrible treatment of gays, which he still hasn't renounced (instead doubling-down and saying that the "deal" was that they could have civil unions but they went too far to get marriage too — they can ride on the same bus, but they can only sit in the back), as well as his nonstop promotion of conspiracy theories is something no sensible person would want to support. If you think Mozilla "[booted Eich] for politics" instead of for taking public actions that tarnished the reputation of the company, you seem like the type of fellow who would have cheered Stephen Douglas in the Lincoln-Douglas debates. As Douglas repeatedly stated, laws concerning the sale of "negroes" are no different from laws concerning the sale of dry goods or liquor, and the people of Illinois should no more tell Missouri how it can sell people than it should tell them how it can sell dry goods — that's Missouri's politics, and Illinois should mind its own business.

Re: Brave Search launches own image and video search

#117

The major problem with Brave search is their position about indexing and licensing content against the wishes of the website publisher. Their robot does not identify itself, meaning the publisher cannot use the standard robots.txt to block its crawling if the publisher so wishes. Incidentally, the robots.txt file has been used in court cases litigating if a search engine is legal or not. Even worse, they state that B…

[deleted]

Re: Brave Search launches own image and video search

#118

The major problem with Brave search is their position about indexing and licensing content against the wishes of the website publisher. Their robot does not identify itself, meaning the publisher cannot use the standard robots.txt to block its crawling if the publisher so wishes. Incidentally, the robots.txt file has been used in court cases litigating if a search engine is legal or not. Even worse, they state that B…

[deleted]

Re: Brave Search launches own image and video search

#119

Earlier quoted context omitted.

I’m just thinking that if website publishers are able to legally allow Googlebot but block other bots, it might contribute to the Google monopoly.

That would be bad, and it is already bad that Google and Microsoft control so much of search queries, but the decision about which search engine indexes a website is purely the publisher's.

[deleted]

Re: Brave Search launches own image and video search

#120

Earlier quoted context omitted.

What if I want what I publish to be known only by word of mouth? What if I consider (some or any of) my ideas to be un-indexable, not directly suitable to representation in any hierarchy other than those I may set them in?

Then you should hide them behind a url that isn't linked elsewhere on your site that you can easily propagate by word of mouth only. example.com/correcthorsebatterystaple If you consider "word of mouth" to be public posts on a forum which millions can read at any time then block googlebot IP's

Yes, sorry, it was a rhetorical question in response to previous.

Taking either step you suggest (along with robots.txt or eqiv.), it would seem fair to expect that Brave, Bing, whomever, would not feel it their neutral/natural domain to include in a public index.

Post reply on HN