Live data from Hacker News

Meta torrented & seeded 81.7 TB dataset containing copyrighted data

arstechnica.com

561–570 of 981 posts

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#561

Earlier quoted context omitted.

The reason there was no political will to punish Airbnb and Uber for violating the law was that initially they were subsidized with VC money and so were able to undercut traditional hotels and taxis on price. In the world of tradable goods, pricing below cost with the intent of putting competition out of business so you can raise prices later is known as "dumping" and is itself illegal.

shouldn't there be a lot of political will from the traditional hotels and taxis, and their lawyers? i can see that the answer is "no", but i don't know why especially with hotels, i would have expected there to be small enough oligopoly to overcome the freerider problem (taxis are more regional, so i don't expect them to be able to fight an (inter)national company very easily) plus the president owning a hotel chain

There are a lot more people who own a few properties as investments than there are hotel owners. Even if these people don't plan to rent through Airbnb, the way Airbnb distorts the housing market is still beneficial for their investments. Also, by the time Trump became president Airbnb was already entrenched for years.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#562
post #463

Earlier quoted context omitted.

> What's interesting is that in many cities now, Uber and Lyft are in fact more expensive than taxis. Sure, agreed. > And the experience is equally mediocre. Absolutely not. I regret using a taxi nearly every time I opt for the cheaper option. It's only the "better" choice if you happen to be standing right in front of one. This experience is nearly universal no matter where I travel. I think people really forget how…

> the typical Taxi experience is nearly as awful as it's always been at least in the US It seems impossible/problematic to generalize the taxi experience to “The US”. If you’re in a city center, cabs can be far easier. The number of times I’ve ordered an Uber or Lyft and regretted it while watching taxi after taxi drive by has been increasing. But I expect the Chicago loop experience to be quite different from say, t…

> quite different from say, the suburbs.

My small rural town of 9000 people had multiple taxi services that poorer people relied on to do even their grocery shopping. We didn't need "disruption"

Tech bros generalizing a negative experience from NYC or SV to the entire US has been so stupid.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#563
post #515

Earlier quoted context omitted.

> If you plug a laptop into a closet at MIT to download some scientific papers you forfeit your life. In case anybody here doesn't know, that's a reference to Aaron Swartz, an activist (and Reddit co-founder) that was risking 35 years in prison and a $1 million fine just for downloading a lot of academic papers from JSTOR. He eventually took his life because of the pressure. May his soul rest in peace.

Except he was offered 6 months in a plea bargain, which he declined because he wanted a trial. Whether 6 months was reasonable punishment for "plug a laptop into a closet at MIT to download some scientific papers" is another matter, but "you forfeit your life" or "35 years in prison and a $1 million fine " is massively misleading.

Swartz wasn't the kind of person to accept a plea bargain from an overzealous prosecutor who was indicting him on 13 felony charges with a possible sentence up to 35 years along with a $1 million fine. I assume he wanted a trial because he wanted to continue his fight for open access. And maybe he thought he might lose, but wouldn't lose on all counts, and would make the prosecution look unreasonable in the public eye. Was that decision rational from a self-interest point of view? Maybe not.

And then you might ask, if he wanted a trial, why did he kill himself? Obviously no one knows what was going through his head when he did it. He left no note. But the prospect of being locked in a cell until he was an old man probably had something to do with it.

You can certainly argue it was his own fault for not pleading down, but even if that's your view, that doesn't absolve the prosecutor. Ortiz has a lot of blame in this too, and the fact she still hasn't acknowledged it over a decade later speaks volumes to the kind of person she is.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#564

I don't understand why it's even a question that Meta trained their LLM on copyrighted material. They say so in their paper! Quoting from their LLaMMa paper [Touvron et al., 2023]: > We include two book corpora in our training dataset: the Gutenberg Project, [...], and the Books3 section of ThePile (Gao et al., 2020), a publicly available dataset for training large language models. Following that reference: > Books3…

There are two different things when it comes to discussing training LLM's on "copyright" protected data, and I almost never see people differentiate. 1.) Training on copyright that is publicly available. You write a poem and publish it online for the world to read. That is your IP, no one else can take it an sell it, but they are free to read and be inspired by it. The legalitly of training on this is in the courts,…

I'm not sure there's any legal distinction though.

Is a book publicly available? No, you have to purchase it. But once you do, you're legally allowed to let your friends and family and so forth read it too. As long as you don't sell copies of it (the "copy" part of "copyright"), or meaningfully take away the ability for the publisher to make money from sales (so you can't post it for the whole world to see on the internet).

And sure, there are lots of ToS for digital works, but are they actually enforceable? ToS can say you're not allowed to let anyone else read the book you purchased. But no court is going to say you can't lend your Kindle to your friend for them to read it too. Many ToS clauses are flat-out illegal.

Meta will argue that training on books is no different from reading all the books at a friend's house. That as long as Meta isn't reselling or making publicly available the original text, they're in the clear.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#565

Earlier quoted context omitted.

It sucked, but not everywhere equally. Meanwhile, Uber rode their one-trick pony (an app), which everyone quickly replicated, all the way to upending taxi businesses worldwide , thanks to their infinite money supply letting them survive long enough in any new market to get the public behind them, which took away support from local regulators trying to keep the market from being gutted by what at this point was a mult…

That seems a little dramatic. They never forced anyone to take an uber right? If taxis were so amazing in other countries why would anyone be interested in switching to uber?

Uber was able to subsidize prices, that's why.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#566
post #37

Based on the encyclopedic knowledge LLMs have of written works I assume all parties did the same. But I think there is a broader point to make here. Youtube was initially a ghost town (it started as a dating site) and it only got traction once people started uploading copyrighted TV shows to it. Google itself got big by indexing other people's data without compensation. Spotify's music library was also pirated in the…

the english empire once tried to mantain a monopoly over steam loom machines the americans cheated their way to competition, heck, even before that, the english empire got jumpstarted by stealing gold from the spanish (who were themselves exploiting it away from aztec and other mexican natives) I'm saying it's business as usual, but also, culture doesn't work like tangible physical widgets so we must stop letting a f…

The textile industry in Brno here in Czech Republic (sometimes called "Moravian Manchester") was hugely helped by a local noble posing as a worker in England & the smuggling detailed self-drawn plans of industrial machinery back:

"Brno’s fortunes were changed forever when a young freemason called Franz Hugo Salma set out for England in 1801. He intended to steal the plans for the most modern textile machinery in the world. His crime, the first recorded act of industrial espionage, boosted the competitiveness of Moravian textiles. Soon after smuggling the plans out disguised as a worker, and handing them over to Brno’s fledgling textile industry, Brno became the most important textile centre in the Habsburg empire."

You can even go see some of the original plans in a museum:

"Eleven designs are still preserved in the library of the Rájec chateau. They form a unique set of documents demonstrating both the level of wool processing technology at the turn of the late 18th and early 19th centuries, as well as the aims and means of the relatively rare business of industrial espionage at that time."

https://www.gotobrno.cz/en/brno-phenomenon/this-is-brno-kate... https://www.gotobrno.cz/en/place/salm-reifferscheidt-palace/

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#567

Earlier quoted context omitted.

That's what makes a banana republic, and for all intents and purposes the U.S. are exhibit A.

The purpose of having an executive branch of government is explicitly to apply the law based on subjective opinions. There's no purpose of having an executive branch of government separate from the other two branches if not to cushion the inflexible and glacial nature of the other branches of government.

>The purpose of having an executive branch of government is explicitly to apply the law based on subjective opinions.

What? No, the purpose of having a separate executive is separation of powers and checks and balances.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#569

Earlier quoted context omitted.

No we just need to enforce the existing laws. And the legal system is for humans not computers.

Plenty of the existing laws are insane and indefensible. Copyright duration of life of the author plus 70 years ? Patents on videogame mechanics? We need to both reform the laws and enforce them. Otherwise... >The law, in its majestic equality, forbids the rich and poor alike to sleep under bridges, to beg in the streets, and to steal bread.

Pantents on video game mechanics... oh how I wish this weren't true. I would love a first person adventure game with the best mechanics and controls taken from genres that did that mechanics very well.

The one that always comes to mind for me is the boxing controls from Fight Night games. It pains me a little every time I play a game where pugilistic battles come down to smash 1 or 2 buttons.

Re: Meta torrented & seeded 81.7 TB dataset containing copyrighted data

#570
post #166

Earlier quoted context omitted.

Airbnb and Uber have showed us that laws matter only to the extent that the political will to enforce them exists. Throw enough lawyers and lobbying money at the problem and the laws can simply be re-written to be friendlier to your business model.

The reason there was no political will to punish Airbnb and Uber for violating the law was that initially they were subsidized with VC money and so were able to undercut traditional hotels and taxis on price. In the world of tradable goods, pricing below cost with the intent of putting competition out of business so you can raise prices later is known as "dumping" and is itself illegal.

> The reason there was no political will to punish Airbnb and Uber for violating the law was that initially they were subsidized with VC money and so were able to undercut traditional hotels and taxis on price.

That's just a trope. They were initially losing money because they had high fixed costs (developing a platform, spending enough on advertising to get a critical mass of people using it), which are long-term investments. If you only spread the cost of the long-term investment over the short-term sales, they were "losing money" in the early years, but that's how all long-term investments work.

Dumping is when you sell below the unit cost, e.g. paying drivers more than you charge customers, which isn't what they were doing in general. And as long as they weren't doing that, the incumbents could have responded by lowering their own prices (and therefore margins) without themselves losing money on each sale, which is competition working as intended. Unless the competition is too hidebound to accept a reduction in profits in order to stay competitive or otherwise insists on using a less efficient method of operating, in which case they go under.

Post reply on HN