Live data from Hacker News

Toyota blames factory shutdown in Japan on ‘insufficient disk space’

theguardian.com

211–220 of 223 posts

Re: Toyota blames factory shutdown in Japan on ‘insufficient disk space’

#211
post #91

Earlier quoted context omitted.

90 days retention is only 2,160 hours. Even at 999 GB/hr that is only ~2160 TB of storage. So, if we stretch the definition of “multiple gigabytes”, is maybe $100k in storage which is around 3-6 developer-months. If we use a more reasonable definition like 10 GB/hr, then that is 20 TB, so maybe $1k in storage which is around 1 developer-day. Seems pretty reasonable to me.

... only 2PB? You might be using a different scale than some of us.

Their scale was money. Saying something is "only" a single digit number of developer months makes sense in this context.

And that was a number hundreds of times higher than what they were replying to, just to make a point.

Re: Toyota blames factory shutdown in Japan on ‘insufficient disk space’

#212
post #91

Earlier quoted context omitted.

90 days retention is only 2,160 hours. Even at 999 GB/hr that is only ~2160 TB of storage. So, if we stretch the definition of “multiple gigabytes”, is maybe $100k in storage which is around 3-6 developer-months. If we use a more reasonable definition like 10 GB/hr, then that is 20 TB, so maybe $1k in storage which is around 1 developer-day. Seems pretty reasonable to me.

A few years ago I joined a company aggressively trying to reduce their AWS costs. My jaw got the floor when I realized they were spending over a million a month in AWS fees. I couldn't understand how they got there with what they were actually doing. I feel like this comment perfectly demonstrates how that happens.

They're pricing hot storage at $50/TB (not per month). That is definitely not AWS or anything like it.

On a per-month basis, the grossly exaggerated number is in the single thousands. The non-exaggerated number is down in the double digits.

$50/TB is a lowball if you want much of the data to be on SSDs, but taking an analysis server and stuffing in 20TB of SSD (plus RAID, plus room for growth) is a very small cost compared to repeated debugging sessions. Especially because the SSD has to deal with about 0.01 DWPD.

Re: Toyota blames factory shutdown in Japan on ‘insufficient disk space’

#213

Earlier quoted context omitted.

In what world is 2160TB $100k? Current single disk solutions are around $25/TB for HDDs and ~$100/TB for NVMe. At a minimum you're looking at $54k just for raw capacity-- assuming no backup, no chassis, no networking, and no redundancy. More reasonable estimations would be in excess of $400/TB.

> Current single disk solutions are around $25/TB for HDDs More like $15/TB. $100K for 2 PB of storage with redundancy and backups is quite reasonable.

I'm showing Exos x20 20TBs for ~$500 new.

$300 is moving towards refurb / shucked prices.

Re: Toyota blames factory shutdown in Japan on ‘insufficient disk space’

#214
post #144

Earlier quoted context omitted.

In what world is 2160TB $100k? Current single disk solutions are around $25/TB for HDDs and ~$100/TB for NVMe. At a minimum you're looking at $54k just for raw capacity-- assuming no backup, no chassis, no networking, and no redundancy. More reasonable estimations would be in excess of $400/TB.

Sure, whatever, a factor of 10 here or there hardly matters. I literally misinterpreted “multiple gigabytes per hour” as 999 GB/hr, not a much more reasonable 10 GB/hr. I literally overestimated data rates by a factor of 10,000% and the number still comes out “reasonable” i.e. a cost that can be paid if the cost/benefit is there. Unless you want to claim storage costs $5,000/TB for 3 MB/s of I/O “multiple gigabytes p…

My recent discussions with multiple SAN vendors as well as quoting out cost to DIY storage has that number being far away from "reasonable". I do not claim storage is $5,000/TB but it is substantially higher than the $50/TB you're estimating.

It's difficult to estimate the log throughput in this scenario. Cisco on debug all can overload the device's CPU; systems like sssd can generate MB of logs for a single login.

All of this is really missing the core issue though. A 2PB system is nontrivial to procure, nontrivial to run, and if you want it to be of any use at all you're going to end up purchasing or implementing some kind of log aggregation system like Splunk. That incurs lifecycle costs like training and implementation, and then you get asked about retention and GDPR.... and in the process, lose sight of whether this thing you've made actually provides any business value.

IT is not an ends in itself, and if these logs are unlikely to be used the question is less about dollars-per-developer-hour and more about preventing IT scope creep and the accumulation of cruft that can mature into technical debt.

Re: Toyota blames factory shutdown in Japan on ‘insufficient disk space’

#215
post #98

Earlier quoted context omitted.

This reminds me of a time we were helping a dev team bring logging in house because they weren't liking the features of their logs-as-a-service provider. They set all applications to "debug" level logs in production and were generating multiple gigabytes of logs per hour. They wanted 90 days retention, and the ability to do advanced searching through the live log data so they could debug in production (they didn't re…

This doesn’t seem terrible if the business benefits justify the costs. There is a cost/benefit to this, presumably.

[deleted]

Re: Toyota blames factory shutdown in Japan on ‘insufficient disk space’

#216
post #206
post #176

Earlier quoted context omitted.

We tried. Devs when give more space just didn't bother to clean old crap and exact same thing happened but with few months of delay. But then we generally just ask how much do they need and bill project for it so that's generally also on them. But, unlike Toyota, we do have disk space alerts. Sometimes the problem is also entirely political, the management needs to tell client and charge them for more storage and won…

In my case the data was a dummy database of more dummy users than are at our org (maybe 50% more). So once this goes to production it would likely get smaller. The problem in this case was twofold: - admin had the job to implement database backups. He didn't factor in backup size when allocating disk space. So this was wntirely his own fault. - the database does store certain transactions for a certain period, so thi…

...did you communicate any of that ?

Because most of our cases where that happens was either lack of planning or lack of communicating that plan. By far most common one was "neither dev nor client knows the data volume in longer period". Which is fine as long as that's also communicated, but that's also often a problem.

But I'm not denying of course that there are just shitty incompetent ops departments, just for the other customer we had dealing with ops department that had:

* backup storage (which was some remote FTPS server IIRC) provisioned so slow the backup wouldn't copy within 24 hours. And the backup size was below TB. * weeks long delays with any resize.

Re: Toyota blames factory shutdown in Japan on ‘insufficient disk space’

#217
post #144

Earlier quoted context omitted.

Sure, whatever, a factor of 10 here or there hardly matters. I literally misinterpreted “multiple gigabytes per hour” as 999 GB/hr, not a much more reasonable 10 GB/hr. I literally overestimated data rates by a factor of 10,000% and the number still comes out “reasonable” i.e. a cost that can be paid if the cost/benefit is there. Unless you want to claim storage costs $5,000/TB for 3 MB/s of I/O “multiple gigabytes p…

My recent discussions with multiple SAN vendors as well as quoting out cost to DIY storage has that number being far away from "reasonable". I do not claim storage is $5,000/TB but it is substantially higher than the $50/TB you're estimating. It's difficult to estimate the log throughput in this scenario. Cisco on debug all can overload the device's CPU; systems like sssd can generate MB of logs for a single login. A…

But you wouldn't use a SAN here. SAN pricing is far away from reasonable for this situation.

For the 20TB case, you can fit that on 1 to 4 drives. It's super cheap. Plus probably a backup hard drive but maybe you don't even need to back it up.

For the 2PB case, you probably want multiple search servers that have all the storage built in. There's definitely cost increases here, but I wouldn't focus too much on it, because that was more of a throwaway. Focus more on the 20TB version.

> That incurs lifecycle costs like training and implementation

Those don't relate much to the amount of storage.

> and then you get asked about retention and GDPR....

It's 90 days. Maybe you throw in a filter. It's not too difficult.

> if these logs are unlikely to be used

The devs are complaining about the search features, it sounds like the logs are being used.

> preventing IT scope creep and the accumulation of cruft that can mature into technical debt

Sure, that's reasonable. But that has nothing to do with the amount of storage.

Re: Toyota blames factory shutdown in Japan on ‘insufficient disk space’

#218

Earlier quoted context omitted.

> Current single disk solutions are around $25/TB for HDDs More like $15/TB. $100K for 2 PB of storage with redundancy and backups is quite reasonable.

I'm showing Exos x20 20TBs for ~$500 new. $300 is moving towards refurb / shucked prices.

> I'm showing Exos x20 20TBs for ~$500 new.

Where? For new prices I'm seeing $350 at amazon, $350 at B&H, $280 direct from newegg, $280 at serverpartsdeals.

Re: Toyota blames factory shutdown in Japan on ‘insufficient disk space’

#219
post #205

Earlier quoted context omitted.

It's very unlikely for all of your disks to fail at the same time. Even if they do, that's why you have an offsite backup. The name of the game is redundancy , not longevity. In practically every use case, two consumer disks will be better than one enterprise disk. Once you start failing enough disks often enough, longevity can be worth the additional cost. Until then, it just isn't.

Actually, it's not uncommon for a batch of disk from a vendor (the same lot #) to have failures. And again, a backup is fine for deleted data, or fire/ransomware. But for day to day operations, no one is really willing to wait for you to restore from a local backup, much less an offsite backup.

> no one is really willing to wait

And they are willing to wait for you to rebuild your raid array?

Either you lost your data from a disk failure, or you are waiting to lose your data from a disk failure. How are we not on the same page here?

Re: Toyota blames factory shutdown in Japan on ‘insufficient disk space’

#220
post #205

Earlier quoted context omitted.

Actually, it's not uncommon for a batch of disk from a vendor (the same lot #) to have failures. And again, a backup is fine for deleted data, or fire/ransomware. But for day to day operations, no one is really willing to wait for you to restore from a local backup, much less an offsite backup.

> no one is really willing to wait And they are willing to wait for you to rebuild your raid array? Either you lost your data from a disk failure, or you are waiting to lose your data from a disk failure. How are we not on the same page here?

I think if you had worked/managed a SAN, you might realize that a single disk failure is a non-event. I'm not talking about a JBOD or an storage shelf using RAID5, I'm talking about a Netapp or similar system that can easily handle disk failures without interrupting service.
Post reply on HN