Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

371–380 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#371
Sounds just like how Facebook got started, harvesting photos without permission. From the Wikipedia article, the Facebook precursor was known as Facemash. On Zuckerberg, "He hacked into the online intranets of Harvard Houses to obtain photos, developing algorithms and codes along the way. He referred to his hacking as "child's play.""

If I were younger, I would be livid.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#372
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

[flagged]

Everyone is responsible for the full effects of his actions. One is literally responsible for all consequences, everything no matter how indirect.

This absolute responsibility is physics, while the limited 'only direct consequences' type thing is a choice made in some human legal systems.

People are smart. They know what stress they put on people and from interacting with them they get a good feel of much they can take. If they ignore that, or decide not to talk to people they're putting stress or choose to ignore things, that's only intentional negligence.

I don't think the prosecutors cared. I don't think we should judge unless there double standards or hypocrisy, but let's not imagine they aren't responsible for things that resulted from their action and inaction. You cause what you case.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#373

Earlier quoted context omitted.

The reason there was no political will to punish Airbnb and Uber for violating the law was that initially they were subsidized with VC money and so were able to undercut traditional hotels and taxis on price. In the world of tradable goods, pricing below cost with the intent of putting competition out of business so you can raise prices later is known as "dumping" and is itself illegal.

I rooted for Uber to smash the Taxi cartels. Let us not forget that Taxi Cartels were also insidious beasts. Taxi drivers abused their walled garden with their price gouging by taking longer routes, refusal to take a credit card, and extremely poorly maintained fleets of vehicles. I have had mostly good experiences with Uber, whereas I had experiences that mostly bordered on general condescension toward me whenever I…

What's interesting is that in many cities now, Uber and Lyft are in fact more expensive than taxis. And the experience is equally mediocre. The pendulum has swung back the other way. The only thing they have going for them now is the app based convenience, which is eroding as more "yellow cab" type traditional taxis band together and get set up with their own sort of city-specific app.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#374
post #329

Earlier quoted context omitted.

Everyone on here is smart enough. Just do not participate and save your money. Do not pay for digital goods. If Netflix raises their prices, it doesn't matter because there is a torrent of all of their shows. If Spotify raises their prices, it doesn't matter because your favorite artist has their entire library in a torrent. If some game company ask you to pay real life prices for a digital costume, find the crack on…

I just can't get behind the sentiment that the unethical behavior by big companies means I get to access all the content I want for free.

Also if you pirate everything you're not incentivizing people to make things in a more ethical manner. I've mostly cancelled my streaming services (I'll get different ones for a month at a time for specific shows) but I still pay for Dropout.tv (when they turned a profit they paid out a dividend to actors) and Patreon for YouTube creators that have high quality content.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#375
post #335
post #209

Earlier quoted context omitted.

Libgen turns into a problem when you have a company developing generative AI with it, either giving money to GPU manufacturers or themselves with paid services (see OpenAI)

What are we actually worried about happening? Are AI-written books getting published? If they start out-competing humans, is that bad? According to most naysayers, they can't do anything original. Are people asking the AI for books? And then hoping it will spit it out a human-written book word for word?

[deleted]

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#377

Best way to "punish" Meta is to slash the Gordian knot and abolish copyright. Level the playing field, incrementally, for everyone else who isn't a trillion-dollar corporation. The alternative is a futile legalistic attack against a monopoly entity too powerful to be meaningfully punished. That won't accomplish anything useful. It would, rather, help cement this status quo, where copyright infringement is selectively…

Ridding copyright would level the playing field for individuals and companies????!!!! Getting rid of laws that protect the individual only will help the larger empowered businesses.

>only will help the larger empowered businesses.

I'm pretty sure I could list ten megacorps that would collapse overnight if copyright was abolished. The music groups, movie studios, streaming platforms...

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#378
post #123

Earlier quoted context omitted.

I think you’re overstating its importance. The internet already makes it possible to order almost any book in existence and have it arrive at your doorstep within a week or so, or often on your ebook reader instantly. And your local library probably participates in an interlibrary loan system that lets you request any book held by any library in the country for free . LibGen gives you access to a much smaller body of…

No one sells scans of older books, which are often sparsely available in obscure (often private) libraries.

Sure, but I have a strong feeling that scans of out-of-print books only constitute a small portion of LibGen’s traffic.

It’s like the idea that most BitTorrent users are just using it to share free software and Creative Commons media. (See the screenshots on every BitTorrent client’s website.) It would definitely be helpful if it were true, but everyone knows it’s just wishful thinking.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#379
post #308

Earlier quoted context omitted.

I've never been a huge user of either, but my worst Uber ride was much better than my best taxi ride.

The last time I dragged my family into a taxi because of my anti Uber ideology, the driver stank to hell of body odor, asked me to input directions on his phone covered with dried snot from him sneezing with his mouth open, he drove dangerously under the speed limit on the freeway, and it took twice as long to get home as normal. But at least I didn’t give Uber any money…

Strange, my last Uber driver had BO. Normally they are fine however.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#380
post #172

Earlier quoted context omitted.

I think most of the public is probably in favor of stronger IP laws now that big corps are threatening to make them jobless with IP-disrespecting AIs

Something tells me stronger IP laws will be drafted by holders of that IP, with little if any regard to the potential for job losses for regular people from AI.

Maybe, but it's better than the current situation with 0 regard for potential job losses for regular people, probably why they're in favor of trying something vs the status quo
Post reply on HN