Live data from Hacker News

Data centers contain 90% crap data

gerrymcgovern.com

131–140 of 146 posts

Re: Data centers contain 90% crap data

#131

90% of libraries consist of books that are never opened. These books were all produced by destroying and processing trees, sometimes with toxic chemicals, and their information density is orders of magnitude lower than that of a hard disk or SSD. Same with photo processing, where 90% of photos taken are discarded, and the toxicity of the chemicals is even higher. So the question isn't simply whether storage is wasted…

Now we're also getting into the topic of "what is waste". Because the majority of the books that are opened are rehashes of the same murder mystery over and over and over. I'd guess that 75% of all new books sold here are variations on "Someone is murdered in a brutal fashion. A old drunken cop from somewhere in Scandinavia is assigned the case. He's helped by a young woman, who may be his daughter or who he'll a fat…

As a librarian, that's only around 35% of new books. The rest are about a hard-working professional woman in the big city that inherits her dead aunt's coffee shop in Cape Cod and has a prickly love-hate relationship with the local handyman. There may be a magic cat involved. And books about vampires having love triangles with a werewolf/zombie/black lagoon monster and a Mary Sue insert.

Re: Data centers contain 90% crap data

#132
post #101

Earlier quoted context omitted.

> 90% of libraries consist of books that are never opened. Citation required. But don't bother because it's a meaningless statistic, or at least one designed to make it look like there's a lot more wastage in libraries than there actually is. The statistic could be true, and yet still be the case that the vast majority of library books are well utilized.

I should have been less specific than “library” as such. Libraries aren’t only institutional. Consider every book, magazine, and newspaper that sits today in everyone’s home, office, and in every institution in the world. The vast majority of them have been read once and then left on a shelf or in a box somewhere, taking up more or less valuable space.

> 90% of libraries consist of books that are never opened

> The vast majority of them have been read once and then left on a shelf

So opened at least once, with once being higher than never.

Re: Data centers contain 90% crap data

#133
I recently worked on a project that proposed using ancient cave systems, full of priceless stalagmites as data centers. The first step in the operation? Raze the cave and cover everything in a thick layer of cement. Use of the underground water table was recommended to cool the AI machines. It was all very dystopian. It was even suggested that the national park system be sold off to build these mega data centers underground.

Re: Data centers contain 90% crap data

#134
> The Cloud made the crap data problem infinitely worse.

This article is mainly focusing about the unused data by website and enterprise databases, only toward the end of the article it barely touched upon "the elephant in the room" of data in cloud.

Now everywhere in the world data centers are being built at breakneck speed to cater for the AI data modeling, training and serving. Most of the AI based data are being kept in datalake in the form of raw data that will probably never see the light of that day i.e never being processed.

Bill Inmon warned us against this potential data swamps in data center due to the increasing popularity of the datalake [1].

Hopefully open table format like Apache Iceberg can rectify this unused raw data epidemic but time will tell [2].

[1] Lakehouses Prevent Data Swamps, Bill Inmon Says

https://www.datanami.com/2021/06/01/lakehouses-prevent-data-...

[2] What Are Apache Iceberg Tables and How Are They Useful?

https://www.snowflake.com/guides/what-are-apache-iceberg-tab...

Re: Data centers contain 90% crap data

#135

Earlier quoted context omitted.

If you’re using Apple Photos, the feature is there. It will detect both exact copies as well as (nearly) identical visual duplicates. Look under utilities.

Maybe it's different on the desktop, but on the iPhone it only detects the same photo that may have been saved at different resolutions or things like that. That's useful, but I really want something that can say use an embedding or something to group all 10 photos I shot of friends around the campfire and help me select one keeper and delete the other 9.

The feature works the same on iPhone. Looking through my duplicates just now, the detected duplicates that had a dozen or so photos to merge, are either high speed sports photography bursts or astrophotography—the rest were just mostly two photo duplicates. I suspect the threshold for determining a the photo as duplicate is too high for successive* photos to be considered such, as with your campfire example.

With that said, I’m surprised Apple hasn’t implemented a feature to group/bundle successive photos beneath what is determined to be “best.”

[*] Note: successive photos being separate shutter taps/actuations (potentially several seconds apart), where bursts are continuous (generally as fast as hardware allows, ms apart).

Re: Data centers contain 90% crap data

#136

Earlier quoted context omitted.

I should have been less specific than “library” as such. Libraries aren’t only institutional. Consider every book, magazine, and newspaper that sits today in everyone’s home, office, and in every institution in the world. The vast majority of them have been read once and then left on a shelf or in a box somewhere, taking up more or less valuable space.

> 90% of libraries consist of books that are never opened > The vast majority of them have been read once and then left on a shelf So opened at least once, with once being higher than never.

Never opened again after the first time—which I think most readers implicitly understood. Regardless, that’s my bad.

Re: Data centers contain 90% crap data

#137
We recently discovered we store 500MiB of email tokens (expired). Then a copy of them in history table. Then this data is replicated. Then there are total of 4 backups of it. Backups that are done every 2 hours for half of each day. And we store those backup for up to half a year. I don't even want to add that up...

Re: Data centers contain 90% crap data

#138
post #123

In this book 'World Wide Waste' he states 'Every time I download an email I contribute to global warming.' Is this true? Aren't some data centers, Google for example carbon neutral?

>>Google's carbon footprint jumped by 48% from 2019, amounting to 14.3 million tons of carbon dioxide emissions for 2023. According to the company's Environment Report 2024, 24% of its total emissions, or over 3.4 million tons, comes from market-based sources, i.e., purchased electricity

published July 3, 2024:

https://www.tomshardware.com/tech-industry/google-reveals-48...

Re: Data centers contain 90% crap data

#139

> One organization I knew of had 1,500 terabytes of data, with less than 2% ever having been accessed after it was first stored. On a related note, probably a similar percentage of people claim on their car insurance. If only the rest realised they had "crap insurance" and were paying for nothing, they could save so much money! This is obviously sarcasm, but I think it's important to remember that much of the data is…

For your car note, insurance is only worth it if the thing being insured can ruin you if something were to happen to it. In the case of cars, you can potentially get ruined from the value of the other party's car. But if you live somewhere where most people drive normal cars, it might be worth it not having insurance. Our culture of insure everything, from your iphone to house is a market failure. If your house were to burn down, the value of the asset is still mostly concentrated in the land.
Post reply on HN