Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

601–610 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#601

Earlier quoted context omitted.

The traditional taxi industry was rife with corruption, bad experiences, and poor service in many jurisdictions before uber/lyft. As terrible of a human being that I think Travis Kalanick is, it was only going to take lawbreaking to overcome such a tainted system. Medallion systems often prevented any competition, sometimes to absurd effect. The number of licenced taxis often didn't keep pace with population growth,…

in my experience, taxi quality varies wildly depending on where you are in bay area, it absolutely makes sense to invent uber, because the taxis were awful. and in vancouver (canada), they're also awful, and deserve the disruption: they would often tell you it'd be a 40 minute wait, and then just not show up taxis in new york were and continue to be totally fine. you just stand outside and get in ~20 seconds later, w…

Vancouver was a great example of the corruption inherent in monopolies. Vancouver had neither Lyft nor Uber until 2020. I heard (internally, when I used to work for Uber) that the reason is that some politicians there had a personal stake in the taxis, so they got a $50 minimum fare passed for all booked rides.

The thing that Uber and Lyft really provided was a surveillance economy to keep both the drivers and riders somewhat in-line. Without it, every ride is an almost anonymous one-shot transaction with almost no recourse on one side, so the game theory suggests that service only has to be good enough that the police aren't called.

https://www.urbanyvr.com/uber-lyft-vancouver-launches/

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#602
post #226
post #218

Earlier quoted context omitted.

Perhaps they just did, or we are doing it - basically this should lead to abolition of copyright to any published article there is. Not sure how’d it impact open source, we’ll either have all of it open, or none at all.

Even without copyright there are trade secrets, not to mention trademarks and patents. Maybe we could get rid of the latter, but I think we’d need to be pretty heavily into socialist utopia before considering nixing the former two!

We need different perspective to copyright. Besides - what is a trade secret 10, 20, 30 years ago is a common wiki article now… very often if not always.

The idea of people owning information is really beyond comprehension for me. There’s no patent for ideas, only for mechanisms or implementations.

Besides we’re already tossing world’s knowledge in our palms, all the copy shit seems so irrelevant.

I’m not against closed source or keeping trade secrets. But once a story becomes public it should be accessible at no cost or else we get where we are atm.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#603

Earlier quoted context omitted.

The traditional taxi industry was rife with corruption, bad experiences, and poor service in many jurisdictions before uber/lyft. As terrible of a human being that I think Travis Kalanick is, it was only going to take lawbreaking to overcome such a tainted system. Medallion systems often prevented any competition, sometimes to absurd effect. The number of licenced taxis often didn't keep pace with population growth,…

It sucked, but not everywhere equally. Meanwhile, Uber rode their one-trick pony (an app), which everyone quickly replicated, all the way to upending taxi businesses worldwide , thanks to their infinite money supply letting them survive long enough in any new market to get the public behind them, which took away support from local regulators trying to keep the market from being gutted by what at this point was a mult…

> It sucked, but not everywhere equally.

Ok, but “everywhere” isn’t my problem. They sucked everywhere I had to use one which is my problem.

> they weren’t that bad everywhere

They were beyond a joke where I am from, which is not the US. Even today, they remain a worse option.

> US specific problem

There was nothing US specific about it.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#604
post #58

We all like hating big corporations, especially Meta, and people seem to use this as an opportunity to advocate for punishing them. I think it's wiser to advocate for changing our IP laws.

We're sick of the double standards. https://en.wikipedia.org/wiki/Aaron_Swartz#United_States_v._... https://en.wikipedia.org/wiki/Aaron_Swartz#Death While Aaron Swartz was bullied to suicide, these corporations will walk free and make billions. I say give every tech CEO the Swartz treatment, then change the law.

Why not change the law first?

Two wrongs don't make a right. If a law is unjust, then what good is there in continuing to punish people who have broken it, just because other people have been punished in the past?

Either you think the law is just or unjust. If you think it's unjust, I don't possibly see how you think people should be punished for it. Meta wasn't responsible for what happened to Aaron Swartz.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#605
post #263

Earlier quoted context omitted.

old books? i can imagine the shit/hallucinated-like generative AI we would have if the training weight was restricted to public domain stuff... i think when chatGPT was around version 2 or 3, i had extracted almost 2 pages (without any alteration from the original) with questions that considered the author from this book here, https://www.amazon.com/Loneliness-Human-Nature-Social-Connec... now it's up to you to think…

The answer is to censor the model output, not the training input. A dumb filter using 20 year old technology can easily stop LLM's from verbatim copyright output.

I know that this seems likely from a theoretical perspective (in other words, I would way underestimate it at the sprint planning meeting!), but

A) checking each output against a regex representing a hundred years of literature would be expensive AF no matter how streamlined you make it, and

B) latent space allows for small deviations that would still get you in trouble but are very hard to catch without a truly latent wrapper (i.e. another LLM call). A good visual example of this is the coverage early on in the Disney v. ChatGPT lawsuit:

[1] IEEE: https://spectrum.ieee.org/midjourney-copyright

[2] reliable ol' Gary Marcus: https://garymarcus.substack.com/p/things-are-about-to-get-a-...

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#606

My ISP will shut off my internet if it catches me torrenting copyrighted material but if you're a massive corporation that steals TBs of data its barely a blip in the news.

You should look into changing your ISP, or at least get a VPN.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#607
I deleted my facebook account about 10 years ago. Downloaded data, deleted. Not deactivated.

Nothing in my life made me ever want to go back except for when I got back into playing hockey, and all the hockey leagues use facebook to communicate a few months ago.

I made a new account, had to literally upload a picture of my face to pass verification.. and then a few days later I was immediately banned and couldn't use my account. I assume because they searched previous data and compared my face to find out I have a "deleted" (lol) account and matched me. I've assumed they'll only let me log in if i use my original 10 years ago deleted account.

Fuck meta. Fuck zuck.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#609
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

Everyone on here is smart enough. Just do not participate and save your money. Do not pay for digital goods. If Netflix raises their prices, it doesn't matter because there is a torrent of all of their shows. If Spotify raises their prices, it doesn't matter because your favorite artist has their entire library in a torrent. If some game company ask you to pay real life prices for a digital costume, find the crack on…

lol I absolutely do not want non digital goods nor pirating. Ever. It's 2025. I don't have a cdplayer, a tape player, a blue ray player, I don't even know what the most modern "blue ray" disc would be. I have $2k worth of vinyls that are just unique copies I display as art I'll never put in my record player, that's also never been used. I don't want to constantly worry about 60gb of mp3 files.

Oh no, that TV show I'll forget about in a year cost me $15/mo instead of $60 of blurays.

I jump in my cars and hit a button and music plays. Almost any music I want. That's amazing.

I'm also not pirating games. I'm not 12 without a job. I have a job. I pay developers for their work. I want more games, like Kingdom Come 3, to come out.

Weird ass comment. You seriously think we're going to put our lives on hold to.. what, fight "digital media"? You think I care about netflix? Or societies use of it? I haven't used netflix in years. I don't know anybody under 40 with a netflix account. Everyone on your end of the pirate spectrum uses debrid nowadays, anyway.

Next you're going to tell people to install the "Black XP Windows" edition to not support Microsoft and they all get malware and their credit cards stolen because they installed some pirated and modified cracked windows. Genius.

MSNBC just cancelled Andrea Mitchells TV show, today, because she brought in no younger audiences. So yes, shows do get cancelled by not being watched.

This comment was upvoted? Hn needs a break. This is some I'm 14 and edgy bullshit that sounds like it belongs on an eastern european piracy forum.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#610
post #598

Earlier quoted context omitted.

Several of those things aren't even necessarily illegal and are the sort of things they shouldn't have had have any reason to do unless they were being targeted by a media campaign or captured government. There is also some dispute about whether some of those even happened or are just mischaracterizations from the media campaign. It's like saying "well, they weren't only violating the taxi medallion cartel laws, they…

Move the goalposts any more and they’re going to be outside the stadium. What laws matter to you? I agree there are shit laws but why can uber break them with impunity but individuals are jailed for smoking some fun lettuce?

The question you should be asking is, what do you want to do about it? Throw the people challenging the taxi cartels in prison, or get rid of the laws against fun lettuce?
Post reply on HN