Disclaimer I work for another cloud (not AWS), opinions are entirely my own. I try to avoid posting in a negative fashion about clouds, but holy crap this blog post...
AWS has this principle of Customer Obsession that enters in to lots of discussions, design decisions etc. "What is the customer experience of $foo?". Along with asking the positive, you ask the negative too, and explore the customer impact of shit going wrong. What does the worst experience look like, what is the impact for the customer, how might you mitigate that or make it so you can at least make it up to customers quickly, if you really can't avoid it.
I find it hard to fathom Sudeep's attitude here. So much of this article is ringing large alarm bells. These are not the things I'd want to see from a cloud provider as a customer.
Is this Stockholm Syndrome? Too much drinking of Kool-Aid as an ex-googler? Unfamiliarity with how other cloud providers operate?
(from part 1)
> Automatic Upgrade of Firebase Account to Paid Account
This is what I mean when I say look at negative vs positive use cases. I'm guessing some combination of customers having a lousy experience running in to Free Tier limits, and staff spending too long having to bump up accounts. So they implemented an automatic upgrade (What, then, is the point of a free tier? No room to experiment, no room to try it and see)
This is precisely the sort of thing that customer obsession principle is supposed to aid in. Automated upgrade certainly solves the staffing time spent bumping up accounts, and it helps customers that used to have to request limits being increased, but it massively fails in the negative customer experience side of the equation here. Someone, somewhere, should have asked the question "What if the customer has made a mistake".
Instead, make it easy and quick for anyone to click a button and get their account changed from Free to Paid, without staff engagement. Give customers easy agency to control their experience.
> Billing “Limits” don’t exist. Budgets are at least a day late.
That's insane. Clouds are about speed and dynamic scalability. Mistakes can ramp up the bill an crazy amount in a short period of time, as Sudeep found out.
How is a 24 hour delay in billing sync and budget warnings even remotely acceptable to them / Sudeep / customers?
Sure it's probably fine for the 90% cases, but that's crippling for the 10% and even if you decide you really don't give a crap about your customers, you don't want the bad press that 10% will likely give you.
Picture what financial damage someone might do if they compromised some of your credentials somehow? You screwed up, credentials got leaked, and you won't necessarily know for a day that something has gone wrong, nor will your restrictions kick in?!
Billing is the single highest TPS service in any cloud, with Identity often a close second (billing gets requests for every transaction, and internal requests related to ongoing charges). You need to handle a high rate of requests, with low latency both in request/response and processing data received. It's a hard engineering problem, and cloud platforms try to get some of the smartest engineers working on it. An organisation of Google's caliber has more than enough smart engineers to be working on these kinds of hard problems, even by temporary secondment.
Quota / Limits in a fast changing cloud environment need to be dynamic and responsive.
>I knew how to put the case for Google team when they would come back to work in 2 days.
How is 2 days even remotely acceptable? Maybe it's just how it's written, but it reads like this is just accepted as the way things are. Why would you even have to carefully work to present your case?
Where are the 24x7 response people with the ability to forgive bills? $72k is chump change for a cloud provider, and especially for a company of the scale of Google. Give your support agents the tools and authority they need to make reasonable decisions, with some appropriate kind of oversight process, and stick in feedback mechanisms so product managers know what problems customers are having.
It's not like that would actually have cost them $72k in direct running costs either. That should have been a near instant no-brainer. Forgive, move on, and reap the benefits of good customer good will. That good will will earn you way more profit than forgiving it would have cost. You're investing in their continuing business. Sometimes those investments will fail, but most of the time they'll succeed.
>In our case, it differed by 86,585,365.85 %, or 86 million percentage points. Even when the bill was notified to us, Firebase Console dashboard still said 42,000 read+writes for the month (below the daily limit).
So it's just fake observability? What's the point? 24 hours delay here is nuts, almost to the point of being useless. It can be hard to calculate these figures out yourselves. A fast feedback cycle is critical. As Sudeep here found out, 24 hours is a great way to have zero clue what's going on until it's too late. Is there really no other way to get this information more up-to-date?
Moving on to part 2:
>I had a team of ~7 engineers/interns at this time, and it would take Google about 10 days to get back to us on this incident.
Why is a 10 day response time from Google considered even remotely acceptable for a cloud provider? Your entire platform is down, you're working out ways around this situation, stressing about potential bankruptcy, and it's just cool with you that it took 10 days for them to make a business/life changing decision over what amounts to chump change?
These kinds of mistakes happen with clouds, AWS is famous for waving these shock bills from mistakes and it never takes 10 days to get it done.
Billing should be the easiest and most obvious thing. If your cloud provider is creating complicated billing structures, that's a problem the cloud provider should be solving, not expecting customers to unravel the mysteries.
Companies being spun up to help people navigate your billing should be an alarm call, not something to celebrate or for customers to consider normal.
> Fail fast, learn fast with Cloud is a bad idea
It shouldn't be. With near immediate feedback you'd have known straight away that shit was bad, and cut the experiment out before it cost you an arm and a leg.
> While creating a Cloud Run service, we chose default values in the service. The max-instances is preset to 1000, and concurrency set to 80 ...... Same goes with Cloud Run! With Concurrency == 60, max_containers == 1000 and each Request taking 400ms, number of requests Cloud Run can handle 9 million requests per minute!
Why are the default values that high on a service? That seems like you're asking customers to shoot themselves in the foot. Where was the look at the negative customer experience side of the equation? Make it easy for customers to do the right thing.
Then the bit that really bugs the crap out of me:
> Thank you Google!
He's thanking Google for having had an absolutely shitty experience on their platform: 10 days of stress from needless delays in forgiving a trivially small bill, dealings with multiple lawyers, investigating bankruptcy, risk of missing product launch date, working around the clock to dig themselves out of hell...