Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

651–660 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#651

Earlier quoted context omitted.

Fast moving field does make sense in terms of copyright because the knowledge is recorded in documents which are then copyrighted. E.g. research papers. > If I write a really popular book, I don't want Hollywood to make it into a movie without compensating me just because they waited a few years I genuinely don't understand this. Even at a decade copyright, pretty much anybody who was going to buy the book and read i…

> Even at a decade copyright, pretty much anybody who was going to buy the book and read it has already done so. It costs you virtually nothing in sales, and society benefits from the resulting movie. If the movie can be made then the book can be printed and sold by any publisher, under the current system. It creates a race to the bottom on the price of the book as soon as the copyright duration is done. Perhaps exte…

That race to the bottom is a feature, not a bug. It allows poor people to engage with culture. That's the tradeoff here. At some point copyright is protecting a tiny amount of profits for the author in exchange for locking people out of access.

Copyright is supposed to be a societal benefit, or there's little reason for society to spend money on enforcing it. That's where we currently are, and I think why there's such a strong reaction to copyright currently. We pay to protect the works and then we pay again to buy them. They become free when they're so culturally irrelevant that nobody wants them even for free. The costs of enforcement are socialized and the benefits are privatized.

At some point, copyright is going to have to provide more back to society or society will get tired of paying to enforce it.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#652
post #156

Earlier quoted context omitted.

It doesn't "seem". The entire system in most countries works, by design, that way because the people in power trade in influence at a different plane. That's why democracy often feels "failed" in that no change can be achieved because "it's just more of the same". Few Lobbyists representing the interests of a few people have more power than millions voting differently.

What happens in US right now shows that change is achieved through voting. There are other examples as well in Europe where things did change because of how people voted. If the change is good or bad depends on your perspective. For me the annoying part is that people vote for a guy because of a couple heavily advertised issues, ignoring all the other plans or the fact that he might not keep his word. Then they are u…

Yes. US and places where people can elect a democracy have a higher chance of some change than European countries with parliamentary systems where a sudden populist candidate won't make it through that system.

I'd argue that, even if some change does happen in the US. Most change (see healthcare, military spending, etc) won't happen because big money will beat the majority of the populace every time.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#653
post #609

Earlier quoted context omitted.

Everyone on here is smart enough. Just do not participate and save your money. Do not pay for digital goods. If Netflix raises their prices, it doesn't matter because there is a torrent of all of their shows. If Spotify raises their prices, it doesn't matter because your favorite artist has their entire library in a torrent. If some game company ask you to pay real life prices for a digital costume, find the crack on…

lol I absolutely do not want non digital goods nor pirating. Ever. It's 2025. I don't have a cdplayer, a tape player, a blue ray player, I don't even know what the most modern "blue ray" disc would be. I have $2k worth of vinyls that are just unique copies I display as art I'll never put in my record player, that's also never been used. I don't want to constantly worry about 60gb of mp3 files. Oh no, that TV show I'l…

I don't know how you went from "don't pay for overpriced digital goods, just pirate them instead" to "hurr durr start using blurays and vinyls".

Reading comprehension is a lost art nowadays.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#654

Earlier quoted context omitted.

Seeding and downloading are in the same protocol. You can't do one without the other

Why comment if you have no idea what youre talking about?

I've written my own torrent clients. Have you?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#655

Earlier quoted context omitted.

> Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. I don't understand why you wouldn't just buy copies of the books. Seems like such a relatively inexpensive way to strengthen your legal case.

thanks to the byzantine copyright system, you can't easily do it. Plus, just speculating, but maybe by paying, it establishes "consideration" for some implied contract? "You implicitly entered a contract with us by purchasing the book, then violated the contract by 'distributing' the material for commercial use" ?

There must be a publisher out there that forbids you from training an AI on the copy you buy from them by now.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#656
post #329

Earlier quoted context omitted.

Everyone on here is smart enough. Just do not participate and save your money. Do not pay for digital goods. If Netflix raises their prices, it doesn't matter because there is a torrent of all of their shows. If Spotify raises their prices, it doesn't matter because your favorite artist has their entire library in a torrent. If some game company ask you to pay real life prices for a digital costume, find the crack on…

I just can't get behind the sentiment that the unethical behavior by big companies means I get to access all the content I want for free.

“There is no ethical consumption under capitalism”

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#657

Earlier quoted context omitted.

For downloading papers , paid for with government funding and gatekept by greedy rent seekers charging ~ $30 a pop. The lengths people will go to defend things that should not exist astounds.

What about his family and friends? Do you blame them as well?

What does that even mean?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#658
post #404

Earlier quoted context omitted.

What's interesting is that in many cities now, Uber and Lyft are in fact more expensive than taxis. And the experience is equally mediocre. The pendulum has swung back the other way. The only thing they have going for them now is the app based convenience, which is eroding as more "yellow cab" type traditional taxis band together and get set up with their own sort of city-specific app.

I remember calling a taxi 3 hours before my flight to get to SFO. After an hour and four different phone calls to the taxi company, I took BART and barely made it before the counter closed. The feedback system incentivizes drivers and riders to behave.

I’ve waited an hour for a Lyft while driver after driver accepted then canceled the ride. Ridesharing does not have great reliability either.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#659

Earlier quoted context omitted.

The laws they allegedly broke were the taxi medallion cartel laws, which were the things keeping taxis terrible by limiting supply and competition. And those laws in general apply to the drivers rather than the ride hailing service. There is also a lot of ambiguity there, e.g. if you have a ride sharing service where people go on the app to find people to carpool with on a trip they'd be making anyway and then contri…

Taxi Medallion laws were also a Reputation Engine that was publicly queryable , subject to FOIA laws and generally had easy to search public databases for them, with detailed notes. Sure Uber/Lyft boil that into a "friendly" 5-star UI, but do you have any idea what data contributed to that star rating? Do you always trust the algorithms that compute them from a bucket of metrics you can't directly request? Sure, Meda…

> Taxi Medallion laws were also a Reputation Engine that was publicly queryable, subject to FOIA laws and generally had easy to search public databases for them, with detailed notes.

You just landed at the airport and need a cab. You fax your FOIA requests for each of the hundred cab companies in the area, which they're required to provide within 20 business days. Your return flight is in 3 days and it would be nice to leave the airport before then.

> Sure Uber/Lyft boil that into a "friendly" 5-star UI, but do you have any idea what data contributed to that star rating? Do you always trust the algorithms that compute them from a bucket of metrics you can't directly request?

So compete with them instead of banning them. Fund an open source ride hailing app with open data. Don't require anyone to use it. If it's better, they will. If it's not better, why should they be forced to?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#660
post #590

The more I learn about how AI companies trained their models, the more obvious it is that the rest of us are just suckers. We're out here assuming that laws matter, that we should never misrepresent or hide what we're doing for our work, that we should honor our own terms of use and the terms of use of other sites/products, that if we register for a website or piece of content we should always use our work email addr…

And if you were in any doubt before, this lesson is now exemplified by the holder of the highest office in the land and approved by popular vote. The rewards of acting ethically are, unfortunately, sometimes only personal. This must be a hard environment to raise children in, given the examples they see around them.

THIS.
Post reply on HN