Live data from Hacker News

Building your own deep learning computer is 10x cheaper than AWS

medium.com

261–269 of 269 posts

Re: Building your own deep learning computer is 10x cheaper than AWS

#261

Earlier quoted context omitted.

Only a Rails developer would think 6k requests per minute with 40ms latency is reasonable with all that hardware. If you rewrote it you probably only need 1 server but you will probably make an argument about how developer time is more valuable :)

It's a matter of dosage. You might be talking about how another tech stack would perform better at this metric, but the price of that is the company would have been simply unable to ship anything in the time they could afford. Swinging the conversation beyond the dosages of either side doesn't produce interesting insight

Exactly. I'm pretty sure if I developed the whole thing in C, it would be really fast. But I would be the slowest part of the development process.

Re: Building your own deep learning computer is 10x cheaper than AWS

#262

Earlier quoted context omitted.

Unfortunately, the fight goes even further than that when you go against the cloud. Last week I was in an event with the CTOs of many of the hottest startups in America. It was shocking how much money is wasted on the cloud because inefficiencies and they simply don't care how much it costs. I guess since they are not wasting their own money, they can always come up with the same excuse: developers are more expensive…

Only a Rails developer would think 6k requests per minute with 40ms latency is reasonable with all that hardware. If you rewrote it you probably only need 1 server but you will probably make an argument about how developer time is more valuable :)

10 servers does seem a bit much. But who knows what their hardware specs are.

Re: Building your own deep learning computer is 10x cheaper than AWS

#263
post #256

Earlier quoted context omitted.

Unfortunately, the fight goes even further than that when you go against the cloud. Last week I was in an event with the CTOs of many of the hottest startups in America. It was shocking how much money is wasted on the cloud because inefficiencies and they simply don't care how much it costs. I guess since they are not wasting their own money, they can always come up with the same excuse: developers are more expensive…

>In 6 years, there were only 3 outages that weren't my fault. How many where there that were your fault? And of those, how many would have been avoided by using Heroku? >All at the cost of office rent ($1000) + FIOS ($359) + Cloudflare costs and S3 for images and backups. What about the time spent creating and maintaining this infrastructure?

I'm not sure how many outages I could avoid using Heroku, but I guess at least a few.

One time I was using Docker for a 2 TB MongoDB and it messed the iptables rules. I notice everything slow for a few days until the database disappeared and when I logged in to check, there was a ransom note.

I flew from Boca Raton to NJ to recover the backup and audit if that was the only breach. That was the longest outage.

Like Rome, this infrastructure was not created in one day. Adding Optane storage is something more recent, for example. Or adding a remote KVM to make easier to manage than dealing with multiple DRACS, which I did after a moved to Florida.

But I'm not against using the cloud. I'm actually very in favor. What I'm against is waste.

In my case, being very conservative with my costs and still have a lot of resources available allowed me to try and keep trying many different ideas in the search for product/market fit.

Re: Building your own deep learning computer is 10x cheaper than AWS

#264

Earlier quoted context omitted.

Only a Rails developer would think 6k requests per minute with 40ms latency is reasonable with all that hardware. If you rewrote it you probably only need 1 server but you will probably make an argument about how developer time is more valuable :)

The 6k/req with 40ms is just at the front door. I'm talking about a real application here. With 100s of database and API calls on each web page load. I could make the whole thing in Golang or Scala and that would be at least one order of magnitude faster. But then I would have to throw away all the business knowledge that was added to the Rails app. For instance, the slowest API call on the 40 ms is one that hits an…

It sounds very similar to the Rails frontend I helped replace at Twitter — no business logic was thrown away, it was done without any loss in application fidelity. We did get approximately 10x fewer servers with 10x less latency after porting it to the JVM as a result. However, without the latency improvement, I don't think we should have done it. Fewer servers just isn't as important as the developer time necessary to do the change as you just pointed out. Just as using the cloud to simplify things and reduce the amount of effort spent on infrastructure is the main driver of adoption. There is clearly a cross-over point though where you decide it is worth it. The CTOs you are speaking of are making that choice and it probably isn't a silly excuse.

Re: Building your own deep learning computer is 10x cheaper than AWS

#265

Earlier quoted context omitted.

Only a Rails developer would think 6k requests per minute with 40ms latency is reasonable with all that hardware. If you rewrote it you probably only need 1 server but you will probably make an argument about how developer time is more valuable :)

I get Rails people are crazy biased, but I still don't understand your argument. What type of business that someone would basically run out of an office closet like this would need to service more than 6k requests per minute?

I was talking about reducing the number of servers to serve 6k/minute. No idea if that was a synthetic benchmark or the current load on the service.

Re: Building your own deep learning computer is 10x cheaper than AWS

#266
post #257

It's been this way since day 1. NVLINK remains the only real Tesla differentiator (although mini NVLINK is available on the new Turing consumer GPUs so WTFever). But because none of the DL frameworks support intra-layer model parallelism, all of the networks we see tend to run efficiently in data parallel because doing anything else makes them communication-limited, which they aren't because data scientists end up bu…

22,000-wide output sourced by a 4096-wide embedding You will want to use hierarchical outputs in this case. Take a look at Hinton's 'Knowledge Distillation' paper.

Sure, that's a nice approximation, and it will reduce performance somewhat IMO, it'd be nice to quantify how much, no?

Re: Building your own deep learning computer is 10x cheaper than AWS

#267
post #231
post #192

Earlier quoted context omitted.

> They sent me long email listing reasons why I shouldn't get a second monitor I had no idea Google was so cheap. Gourmet breakfast, lunch, and dinner every day? No problem. A couple hundred bucks for a second monitor? Uh... it's not about the money, we're, uh, concerned about the environment.

To be fair back then decent 30" monitors were around $800-$1000. (Also, it's just normal food, not "gourmet".)

But it’s certainly a step up from standard corporate catering

Re: Building your own deep learning computer is 10x cheaper than AWS

#268
post #44

Earlier quoted context omitted.

I used to manage hardware in several datacentres, and I'd usually visit the data centres a couple of times a year . Other than that we used a couple of hours of "remote hands" services from the datacentre operator. Overall our hosting costs were about 30% of what the same capacity would have cost on AWS. Once a year I'd get a "why aren't we using AWS" e-mail from my boss, update our costing spreadsheets and tell him…

When you create the spreadsheet, do you price in running servers 24x7 or using elastic capacity?

We had too low variations in load over the day to make elastic usage cost effective for the most part, so it made very little difference. Indeed, one of the most cost effective ways of using AWS is to not use it, but be ready to use it to handle traffic spikes. Do that and you can load your dedicated servers much closer to capacity, while almost never having to spin any instances up.

Re: Building your own deep learning computer is 10x cheaper than AWS

#269
post #256

Earlier quoted context omitted.

>In 6 years, there were only 3 outages that weren't my fault. How many where there that were your fault? And of those, how many would have been avoided by using Heroku? >All at the cost of office rent ($1000) + FIOS ($359) + Cloudflare costs and S3 for images and backups. What about the time spent creating and maintaining this infrastructure?

I'm not sure how many outages I could avoid using Heroku, but I guess at least a few. One time I was using Docker for a 2 TB MongoDB and it messed the iptables rules. I notice everything slow for a few days until the database disappeared and when I logged in to check, there was a ransom note. I flew from Boca Raton to NJ to recover the backup and audit if that was the only breach. That was the longest outage. Like Ro…

Are you using the Optane SSDs for Ceph as you mentioned in your original comment? Curious what benefit you're seeing if not for Ceph, and if for something else, would you mind commenting? We're looking to share best practices with the community we're building over at acceleratewithoptane.com on how to take advantage of Optane SSDs.
Post reply on HN