Live data from Hacker News

Ask HN: Has anybody built search on top of Anna's Archive?

news.ycombinator.com

121–130 of 156 posts

Re: Ask HN: Has anybody built search on top of Anna's Archive?

#121

Earlier quoted context omitted.

There's a difference between feeding massive amounts of copyrighted material to a training process that blends them thoroughly and irreversibly, and doing all that in-house, vs. offering people a service that indexes (and possibly partially rehosts) that material, enabling and encouraging users to engage directly in pirating concrete copyrighted works.

There's this famous phrase in Russian that was born out of a short interview with a woman, a strong Putin supporter, that's often been used as a sarcastic remark for pointing out someone's double standards and/or hypocrisy. It can be roughly translated to "you don't understand, it's a completely different situation". That's what's constantly on my mind when I'm reading discussions like this one. Everybody and their d…

There are those who are in charge and those who aren't.

Re: Ask HN: Has anybody built search on top of Anna's Archive?

#122

Earlier quoted context omitted.

> or some other country that doesn't respect international copyright though. Like the US? OpenAI et al. don't give a shit.

There's a difference between feeding massive amounts of copyrighted material to a training process that blends them thoroughly and irreversibly, and doing all that in-house, vs. offering people a service that indexes (and possibly partially rehosts) that material, enabling and encouraging users to engage directly in pirating concrete copyrighted works.

Ironically the low tech infringing proposal would lead to more reliable results grounded in the raw contents of the data, using less computing/power and without the confidently incorrect sycophanty we see from the LLMs.

Re: Ask HN: Has anybody built search on top of Anna's Archive?

#124
post #99

Earlier quoted context omitted.

> > or some other country that doesn't respect international copyright though. > Like the US? OpenAI et al. don't give a shit. OpenAI is not a country and therefore cannot make laws that don't respect international (or domestic) copyright. Also the US is a lot bigger than OpenAI and the big tech corps, and the law is very much on the side of copyright holders in the US.

> the law is very much on the side of copyright holders in the US. Remind me again what the status of the case is with Meta/Facebook using pirated material to train their proprietary LLMs, and even seeding the data back to the community while downloading it?

In progress. Nobody is expecting the original protections afforded by copyright to apply here, but the fact that the material is pirated is less relevant than whether or not an LLM is a transformative use of the material.

We will almost certainly see copyright law weakened by the case, but I do not believe that FB will get off with no penalties.

Re: Ask HN: Has anybody built search on top of Anna's Archive?

#125

Earlier quoted context omitted.

There's a difference between feeding massive amounts of copyrighted material to a training process that blends them thoroughly and irreversibly, and doing all that in-house, vs. offering people a service that indexes (and possibly partially rehosts) that material, enabling and encouraging users to engage directly in pirating concrete copyrighted works.

Ironically the low tech infringing proposal would lead to more reliable results grounded in the raw contents of the data, using less computing/power and without the confidently incorrect sycophanty we see from the LLMs.

Nah. It would just lead to more of classical search. Which is okay, as it always has been.

LLMs are not retrieval engines, and thinking them as such is missing most of their value. LLMs are understanding engines. Much like for humans, evaluating and incorporating knowledge is necessary to build understanding - however, perfect recall is not.

Another, arguably equivalent way of framing it: the job of an LLM isn't to provide you with the facts; it's main job is to understand what you mean. The "WIM" in "DWIM". Making it do that does require stupid amounts of data and tons of compute in training. Currently, there's no better way, and the only alternative system with similar capabilities are... humans.

IOW, it's not even an apples to oranges comparison, it's apples to gourmet chef.

Re: Ask HN: Has anybody built search on top of Anna's Archive?

#126
post #38

Earlier quoted context omitted.

There should be a way to leverage compression when storing multiple editions of the same book.

From a good search perspective though you probably dont want 500 different versions of the same book popping up for a query

And without some sort of weighting system, it wouldn't even know which one is the best one to show the user.

Re: Ask HN: Has anybody built search on top of Anna's Archive?

#127
post #29

Earlier quoted context omitted.

I don't either, but many states have laws regarding books on how to build bombs and they might get enforced more than copyright.

Not saying you're deceiving but can you show me where a state has made a book about bombs illegal? It seems like that would be a slam dunk 1A violation. And yes I'm aware that states willfully violate 2A but I don't want to discuss it here.

Not saying you cannot read, but if you would, the other answer to my comment literally has such an example.

Germany is like this as well since a few years.

Not all states are within the US.

Re: Ask HN: Has anybody built search on top of Anna's Archive?

#128
post #105

Earlier quoted context omitted.

There's a difference between feeding massive amounts of copyrighted material to a training process that blends them thoroughly and irreversibly, and doing all that in-house, vs. offering people a service that indexes (and possibly partially rehosts) that material, enabling and encouraging users to engage directly in pirating concrete copyrighted works.

That's Uber's Gambit. Nothing is illegal for large enough corporations with strong network effects and deep pockets.

That's not Uber's Gambit.

Uber was blatantly ignoring the local laws in order to break into the market and quickly defeat local competition. They used their infinite VC money supply to interfere with and delay investigations and enforcement, betting that if they do it fast enough, they'll have the general population on their side.

LLM vendors found and exploited[0] a legal uncertainty - correct me if I'm wrong, but AFAIK it still isn't settled whether or not their actions were actually illegal. Unlike Uber, LLM vendors aren't breaking into markets by ignoring the laws to outcompete incumbents, and burning stupid amounts of money just to get away with it. On the contrary, LLM vendors are simply providing an actually useful product, and charging a reasonable price for it, while reinvesting it into improving the product. Effects it has on other markets aside[1], their business model is just providing actual value in exchange for money. That's much more direct and honest than most of the tech industry.

The product itself is also different. Uber is selling a mirage, a "miracle" improvement that quickly turns not so, and is destined to eventually destroy the markets it disrupted. LLM vendors are developing and serving systems that provide actual value to users, directly and obviously so.

--

[0] - Probably walked into this without initially realizing it. No one complained 5-10 years ago, where the datasets were smaller and the resulting models had no real-world utility. It's only when the models became useful, that some people started looking for ways to make them go away.

[1] - That's an unfortunate effect of it being a general AI tool, and would be the same regardless of how it was created.

Re: Ask HN: Has anybody built search on top of Anna's Archive?

#129

Earlier quoted context omitted.

There's a difference between feeding massive amounts of copyrighted material to a training process that blends them thoroughly and irreversibly, and doing all that in-house, vs. offering people a service that indexes (and possibly partially rehosts) that material, enabling and encouraging users to engage directly in pirating concrete copyrighted works.

> that blends them thoroughly and irreversibly It's okay, you can say 'laundering'

I can, but I don't, because that's at best an unintended side effect.

Re: Ask HN: Has anybody built search on top of Anna's Archive?

#130

Earlier quoted context omitted.

There's a difference between feeding massive amounts of copyrighted material to a training process that blends them thoroughly and irreversibly, and doing all that in-house, vs. offering people a service that indexes (and possibly partially rehosts) that material, enabling and encouraging users to engage directly in pirating concrete copyrighted works.

There's this famous phrase in Russian that was born out of a short interview with a woman, a strong Putin supporter, that's often been used as a sarcastic remark for pointing out someone's double standards and/or hypocrisy. It can be roughly translated to "you don't understand, it's a completely different situation". That's what's constantly on my mind when I'm reading discussions like this one. Everybody and their d…

Is there also a famous Russian phrase that translates to "details are irrelevant, it kinda looks similar to me therefore it's the same"? If not, there definitely should be.

The details are the entire point. Arguing that a corporation can get away doing something, while an individual can't, isn't useful, because there are great many of such somethings, and in most cases it turns out perfectly reasonable, once you dig into details.

Post reply on HN