Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

831–840 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#831
post #463

Earlier quoted context omitted.

What's interesting is that in many cities now, Uber and Lyft are in fact more expensive than taxis. And the experience is equally mediocre. The pendulum has swung back the other way. The only thing they have going for them now is the app based convenience, which is eroding as more "yellow cab" type traditional taxis band together and get set up with their own sort of city-specific app.

> What's interesting is that in many cities now, Uber and Lyft are in fact more expensive than taxis. Sure, agreed. > And the experience is equally mediocre. Absolutely not. I regret using a taxi nearly every time I opt for the cheaper option. It's only the "better" choice if you happen to be standing right in front of one. This experience is nearly universal no matter where I travel. I think people really forget how…

> Uber and Lyft are in fact more expensive than taxis.

I doubt that.

I double checked, just to be sure since I paid for taxis for years for a specific trip, each way. Uber is still cheaper TODAY than taxis were when I switched 10 years ago. One way 5 minute trip, Friday 6pm in orange county, ca still under 20$ today.

20$ was a good deal (or ripoff, depending on your attitude) for a taxi in 2015, for the same distance and a variable waiting time. Let's just say it's about equal for sake of discussion. There was no app, but there was a dispatcher you could call. There was no incentive to improve, until then.

Companies have had to adapt and prices have come down. It would no doubt be 30$+ today for taxis, if not for rideshare companies.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#832
post #394

Earlier quoted context omitted.

What if the model simply substitutes synonyms here and there without changing the spirit of the material? (This might not work for poetry, obviously.) It is not such a simple matter.

It's pretty simple, you are absolutely allowed to do that, and it's been done forever. Imagine having the copyright claim to "Person's family member is killed so they go and get revenge".

So I can duplicate a book and change and word or two and sell it? That does not sound right.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#833
post #75

Earlier quoted context omitted.

the english empire once tried to mantain a monopoly over steam loom machines the americans cheated their way to competition, heck, even before that, the english empire got jumpstarted by stealing gold from the spanish (who were themselves exploiting it away from aztec and other mexican natives) I'm saying it's business as usual, but also, culture doesn't work like tangible physical widgets so we must stop letting a f…

Hollywood became popular for filmmaking because they were literally the opposite side of the country from Thomas Edison and his patents...

United Kingdom was the first to steal a Chinese invention called gunpowder, use it to maximum effect and they were also at the opposite sides of the world.

Do not forget the first law passed by a new nation in which every citizen has the right to burn trees and produce potassium a critical mineral in gunpowder. That nation was either Greece or America, i forget which one.

In the Western civilizations, when we are stealing copyrighted material and patents we are not messing around. I remember when the Byzantine empire tried to steal secrets of silk production also from Chinese, that was so much fun. Great times we lived in the past, even greater now!

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#835
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

> once people started uploading copyrighted TV shows to it End users, not YouTube employees, right? And they would take things down following DMCA requests and what not, right? So, pretty much following the law? > Google itself got big by indexing other people's data without compensation Scraping public websites to build a search index isn't the same as making LLMs that can recreate the source verbatim devoid of even…

I think he's thinking of Grooveshark

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#836

Earlier quoted context omitted.

I rooted for Uber to smash the Taxi cartels. Let us not forget that Taxi Cartels were also insidious beasts. Taxi drivers abused their walled garden with their price gouging by taking longer routes, refusal to take a credit card, and extremely poorly maintained fleets of vehicles. I have had mostly good experiences with Uber, whereas I had experiences that mostly bordered on general condescension toward me whenever I…

Uber definitely improved things. When traveling it’s also so much safer than taxis. My brother was robbed at gunpoint in a taxi. My wife had to jump from more than one moving taxi to escape. My ex girlfriend too. My Swiss friend had his camera and wallet stolen. You can have issues with Uber too, but not as frequently because there’s a digital audit trail, you can report them to the platform and the police. The threa…

Here's another data point: I have taken literally thousands of taxis on 5 continents and none of this has ever happened to me

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#837
post #618

Earlier quoted context omitted.

> Did every single machine they used have the configuration for only leeching and no seeding? I would certainly assume so. It's incredibly obvious that's what you would want to do from a legal standpoint. > If only one employee was also seeding ... that could be a very interesting case. The torrenting wouldn't be done casually by employees acting on their own. And it's not like multiple employees are doing it simulta…

Did you not read the article? There are quotes from Meta employees doing exactly what you claim they wouldn't do. > This is part of an official project. They'd spin up a machine just to download the torrent, being careful to disable seeding. From the article: > "Torrenting from a corporate laptop doesn’t feel right," Nikolay Bashlykov, a Meta research engineer, wrote in an April 2023 message, adding a smiley emoji. I…

Seeding can be trivially faked to trackers.

https://github.com/slundi/RatioUp

https://github.com/anthonyraymond/joal

http://ratiomaster.net/

The smallest amount of seeding possible would be metadata, presumably not subject to copyright.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#838
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

I think the difference may be LLMs may not be laundered clean of copyright data anytime soon. Even if chatgpt got big and profitable, it's not so clear that it won't contain copyrighted data as that may simply be necessary to train the best models.

Most of the web is copyrighted

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#839

Earlier quoted context omitted.

I don't agree. Getting more comprehensive enforcement of laws in general against well-heeled players is a good thing. We would have a lot less bad law if laws were enforced more evenly, because people would more quickly see their true effects, rather than having to wait until companies exploited the loopholes in enforcement so egregiously. (I also don't agree that the only problem here is bad laws. Yes, some of the l…

> Getting more comprehensive enforcement of laws in general against well-heeled players is a good thing. Whether something is good independent of what it takes to achieve it is a separate question from whether that's where you should focus your efforts. > We would have a lot less bad law if laws were enforced more evenly, because people would more quickly see their true effects, rather than having to wait until compa…

So I get your argument, but by that logic the only bad laws that get repealed will be those that affect big business, and the laws against individuals without resources will still be in place. I think we all agree that it’s unfair there’s a very large difference in enforcement of law between those with resources and those without, but I think to prevent that we need to figure out how to prevent the capture of the government by those with large resources. I could agree in concept that less laws are better for that overall, but also then there’s the question of who benefits from less laws and I bet those with resources will still benefit.

It’s almost like we need to ensure no one has much more resources than anyone else (ya know, workers owning the means of production) so there’s a more level field!

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#840

Earlier quoted context omitted.

in my experience, taxi quality varies wildly depending on where you are in bay area, it absolutely makes sense to invent uber, because the taxis were awful. and in vancouver (canada), they're also awful, and deserve the disruption: they would often tell you it'd be a 40 minute wait, and then just not show up taxis in new york were and continue to be totally fine. you just stand outside and get in ~20 seconds later, w…

> taxis in new york were and continue to be totally fine. you just stand outside and get in ~20 seconds later, with no hassles or apps. This is only true in a small subset of New York.

When it's not raining.
Post reply on HN