Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

591–600 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#591
post #170

Maybe you should go after the worst offender (OpenAI) first before going after Meta, since the latter already gave back their model away for free for everyone and the architecture. We will know why OpenAI isn't getting investigated.

So true. It seems like there is a controlled operation to shut open models down starting with Meta. Obviously they can't go after deepseek atm

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#592

The more I learn about how AI companies trained their models, the more obvious it is that the rest of us are just suckers. We're out here assuming that laws matter, that we should never misrepresent or hide what we're doing for our work, that we should honor our own terms of use and the terms of use of other sites/products, that if we register for a website or piece of content we should always use our work email addr…

If you have a spare few hours, the Acquired podcast episode on Meta is enlightening. They just stumbled through growth hack experiment after experiment without seemingly any risk assessment or ethics.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#593

Earlier quoted context omitted.

The laws they allegedly broke were the taxi medallion cartel laws, which were the things keeping taxis terrible by limiting supply and competition. And those laws in general apply to the drivers rather than the ride hailing service. There is also a lot of ambiguity there, e.g. if you have a ride sharing service where people go on the app to find people to carpool with on a trip they'd be making anyway and then contri…

No, these are not only the laws they allegedly broke. They created a project named Greyball to identify law enforcement and mislead them. They created a kill switch for the event of a government raid to gather evidence. They ordered and then canceled rides on competitor apps. They tracked journalists and politicians... The list goes on and on: https://en.wikipedia.org/wiki/Controversies_surrounding_Uber

Several of those things aren't even necessarily illegal and are the sort of things they shouldn't have had have any reason to do unless they were being targeted by a media campaign or captured government. There is also some dispute about whether some of those even happened or are just mischaracterizations from the media campaign.

It's like saying "well, they weren't only violating the taxi medallion cartel laws, they were also violating laws against evading enforcement of the taxi medallion cartel laws". There is a central cause here.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#594
post #583

Earlier quoted context omitted.

> But I prefer looser intellectual property rights anyway so Im ok with it I think more people, potentially anyways, would feel similar to to this if it applied even somewhat equally. Instead, companies can seemingly do whatever they please whereas lawyers will send letters to your home for downloading a single episode of game of thrones.

> Instead, companies can seemingly do whatever they please whereas lawyers will send letters to your home for downloading a single episode of game of thrones. I don't get it. All these companies took copyrighted data when they were tiny grew to be large, they still do that now. Google and OpenAI don't send these letters. They're not the copyright holders. I have no idea what argument you're trying to make. Corporatio…

>Google and OpenAI don't send these letters.

Right. I'm not saying they do?

>I have no idea what argument you're trying to make.

I thought my point (not really an argument) was pretty clear, sorry.

"Rules for thee, but not for me" is the point. Where "thee" is individuals and "me" is corporations. (My comment was general commentary, not specific to Meta, Google, OpenAI, LLMs, or the article)

Right now "loose restrictions" seems to apply to corporations only. More people might be in favor of looser restrictions if it also applied to individuals, not just corporations.

I'm not sure how else to reword my comment more than that. It wasn't really meant to be too deep, and it wasn't intended to be argumentative.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#595

Earlier quoted context omitted.

> Google itself got big by indexing other people's data without compensation Wrong. a) Robots.txt which defines what content you wish to make available to third parties predates every search engine including Google. Web site owners chose to make it available to Google and search engines have respected their wishes despite it not being in their best interest. b) The difference here is that OpenAI, Meta etc have not ev…

a) If you don't have a robots.txt, you're indexed by default. It's opt-out, not opt-in. If you do nothing, you're being indexed.

It's an opt-out of an opt-in. If you run a webserver hosting your files, you already opted-in to people accessing that data. If you then don't go ahead an configure it properly, that's not exactly "opt-out" anymore. By default your files are not accessible to the network, you have to first opt-in to serving them.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#596

Earlier quoted context omitted.

The laws they allegedly broke were the taxi medallion cartel laws, which were the things keeping taxis terrible by limiting supply and competition. And those laws in general apply to the drivers rather than the ride hailing service. There is also a lot of ambiguity there, e.g. if you have a ride sharing service where people go on the app to find people to carpool with on a trip they'd be making anyway and then contri…

No, these are not only the laws they allegedly broke. They created a project named Greyball to identify law enforcement and mislead them. They created a kill switch for the event of a government raid to gather evidence. They ordered and then canceled rides on competitor apps. They tracked journalists and politicians... The list goes on and on: https://en.wikipedia.org/wiki/Controversies_surrounding_Uber

The best thing to ever happen to corpo scum was that social media took over most of the news. Now there’s no trusted journalists to write a big article about this kind of stuff, instead folks just defend the corpo scum’s actions and spread lies for them, while the truth is still putting on its shoes.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#597

Earlier quoted context omitted.

I rooted for Uber to smash the Taxi cartels. Let us not forget that Taxi Cartels were also insidious beasts. Taxi drivers abused their walled garden with their price gouging by taking longer routes, refusal to take a credit card, and extremely poorly maintained fleets of vehicles. I have had mostly good experiences with Uber, whereas I had experiences that mostly bordered on general condescension toward me whenever I…

This argument ignores the fact that there were other alternatives to Uber at the time, ones that didn't break the law! Believe it or not, there were multiple ride hailing apps on the iPhone, but none were as great at accumulating capital or breaking the law without recourse.

The only alternative I remember is "black car" services, eg airport limos and the like. But there was very little automation around it; you had to speak on the phone with someone to book, and it was always like 24+ hours out rather than "go to the place now" the way a cab is.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#598

Earlier quoted context omitted.

No, these are not only the laws they allegedly broke. They created a project named Greyball to identify law enforcement and mislead them. They created a kill switch for the event of a government raid to gather evidence. They ordered and then canceled rides on competitor apps. They tracked journalists and politicians... The list goes on and on: https://en.wikipedia.org/wiki/Controversies_surrounding_Uber

Several of those things aren't even necessarily illegal and are the sort of things they shouldn't have had have any reason to do unless they were being targeted by a media campaign or captured government. There is also some dispute about whether some of those even happened or are just mischaracterizations from the media campaign. It's like saying "well, they weren't only violating the taxi medallion cartel laws, they…

Move the goalposts any more and they’re going to be outside the stadium. What laws matter to you? I agree there are shit laws but why can uber break them with impunity but individuals are jailed for smoking some fun lettuce?

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#599
post #543

Earlier quoted context omitted.

I rooted for Uber to smash the Taxi cartels. Let us not forget that Taxi Cartels were also insidious beasts. Taxi drivers abused their walled garden with their price gouging by taking longer routes, refusal to take a credit card, and extremely poorly maintained fleets of vehicles. I have had mostly good experiences with Uber, whereas I had experiences that mostly bordered on general condescension toward me whenever I…

It really screwed over a lot of regular working-class people. In some European cities getting a taxi license was a serious monetary investment. People took our huge loans for this. This was now suddenly worthless. It's like being told your very expensive university education is no longer accredited, but the student loan still exists. kthxbye. I'm not saying the existing systems were always good (they weren't), but yo…

> getting a taxi license was a serious monetary investment. People took our huge loans for this

it was a terrible system that sucked for everyone involved. For all of Uber's flaws, would you rather go back to that today? really??

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#600
post #378

Earlier quoted context omitted.

No one sells scans of older books, which are often sparsely available in obscure (often private) libraries.

Sure, but I have a strong feeling that scans of out-of-print books only constitute a small portion of LibGen’s traffic. It’s like the idea that most BitTorrent users are just using it to share free software and Creative Commons media. (See the screenshots on every BitTorrent client’s website.) It would definitely be helpful if it were true, but everyone knows it’s just wishful thinking.

Why does the proportion matter?

Academics are huge users of LibGen for academic books from the entire past century and beyond. It's infinitely more convenient to instantly get a PDF you can highlight, than wait weeks for some interlibrary loan from an institution three states away.

Just because the majority of people might be downloading Harry Potter is irrelevant.

Post reply on HN