Earlier quoted context omitted.
It's ever worse for services like AWS Cloudfront. One of your competitors could just rent a cheap server on OVH with uncapped transfer and incur you $10k in cost in a few hours.
Maybe that it is your cue to move your server from AWS to OVH* * I dont have any idea about OVH
Burnt $72k testing Firebase and Cloud Run and almost went bankrupt
261–270 of 397 posts
Re: Burnt $72k testing Firebase and Cloud Run and almost went bankrupt
#262Earlier quoted context omitted.
I know there's no reason for Google or AWS to do this, but man do I wish there was a way to put down a spending limit and simply disable anything that goes over that limit. It's a little bit nuts that there are no guardrails to prevent you from incurring such huge bills (especially as a solo developer that might just be trying out their services).
The downside of disabling active resources is huge. It would mean a catastrophic interruption to the customers application exactly when its the most popular/active. And theres no practical way to determine whether the customer is “trying it out” or running a key part of their business on any particular resource. On the other hand retroactively forgiving the cost of unexpected/unintentional usage doesnt impact the cus…
I didn't even have any customers at that point.
Re: Burnt $72k testing Firebase and Cloud Run and almost went bankrupt
#263Earlier quoted context omitted.
I know there's no reason for Google or AWS to do this, but man do I wish there was a way to put down a spending limit and simply disable anything that goes over that limit. It's a little bit nuts that there are no guardrails to prevent you from incurring such huge bills (especially as a solo developer that might just be trying out their services).
> but man do I wish there was a way to put down a spending limit and simply disable anything that goes over that limit. Literally did this my first week when trying out GCP for my company. It is entirely possible and documented (with code): https://cloud.google.com/billing/docs/how-to/notify#cap_disa...
(Source link in parent post, emphasis mine).
In this case they had a additional cost due to delay of $72k. Which, lets be honest means this feature kinda useless for anything but the somewhat harmless case.
Only by combining this with resource limits in load balancers, instance and concurrency limits and similar can the maximal worst cost be limited. But tbh. this partially cripples auto-scaling functionality and it's really hard to find a good setting which doesn't allow to much "over" cost and at the same time doesn't hinder the intended auto-scaling use-case.
Re: Burnt $72k testing Firebase and Cloud Run and almost went bankrupt
#264Earlier quoted context omitted.
But we've had disk quotas before that mostly worked? If anything it seems an easier problem than processor time. I recall disk quotas on shared systems at university back in 1998 and I'm sure they existed before that. Two thresholds IIRC, one at which you get a warning, second at which you can't write any further and the disk write operation fails. I don't think they deleted files, it was just you couldn't write more…
S3 costs money to keep your files in, even if you're not touching them, so just preventing further uploads wouldn't do much to prevent your AWS bill from increasing.
Re: Burnt $72k testing Firebase and Cloud Run and almost went bankrupt
#265Earlier quoted context omitted.
It's surprisingly complex to do that. Let's take a simple example and say your cloud account is doing 2 things - compute & storage. Compute is an active resource, when you exceed your budget it can be automatically shutdown. Storage is a passive resource, when you exceed your budget it can be automatically....deleted? That's almost always the wrong action. Providing fine-grained cost limits help some, as passive reso…
But we've had disk quotas before that mostly worked? If anything it seems an easier problem than processor time. I recall disk quotas on shared systems at university back in 1998 and I'm sure they existed before that. Two thresholds IIRC, one at which you get a warning, second at which you can't write any further and the disk write operation fails. I don't think they deleted files, it was just you couldn't write more…
Re: Burnt $72k testing Firebase and Cloud Run and almost went bankrupt
#266Earlier quoted context omitted.
I know there's no reason for Google or AWS to do this, but man do I wish there was a way to put down a spending limit and simply disable anything that goes over that limit. It's a little bit nuts that there are no guardrails to prevent you from incurring such huge bills (especially as a solo developer that might just be trying out their services).
The downside of disabling active resources is huge. It would mean a catastrophic interruption to the customers application exactly when its the most popular/active. And theres no practical way to determine whether the customer is “trying it out” or running a key part of their business on any particular resource. On the other hand retroactively forgiving the cost of unexpected/unintentional usage doesnt impact the cus…
What? Why wouldn't this just be an opt in thing? It could even be tied to the account being used. It's not like AWS accounts are expensive or hard to setup.
If a user opts in to the "kill if bill goes too high" and they kill a critical portion of their business, then that's on them. Similar to how a user choosing spot instances if their spot ends up being destroyed. You've already got that "I can kill your stuff if you opt into it" capability.
> On the other hand retroactively forgiving the cost of unexpected/unintentional usage doesnt impact the customers users.
Yeah, and what happens if someone isn't big enough to justify AWS's forgiveness? What if they get a rep that blows off their request or is having a bad day? You are at the mercy of your cloud provider to forgive a debt, which is a real shitty place to be for anyone.
> And with billing alerts the customer is able to make the choice of whether the cost is worth it as it happens.
And what do they do if they miss the alert? You can rack up a huge bill in very little time the right AWS services.
The point of the kill switch cap is to guard against risk. The fact is that that while 72k isn't too big for some companies, it means bankruptcy for others. Its because you might want to give your devs a training account to play with AWS services to gain expertise, but you don't want them to blow $1 million dollars screwing around with Amazon satellite services.
Re: Burnt $72k testing Firebase and Cloud Run and almost went bankrupt
#267Deeply dissatisfying to read. Ex-Googler uses connections to get his (understandable!) cloud mistake refunded. Every time I read one of these stories, I get more and more convinced I will just simply never use scalable cloud tech for my side projects. I'm not going to risk my family's retirement savings on the all-too-possible chance that a small deep-implication error will cause runaway charges.
Your assumption is incorrect. I haven't been in touch with anyone in Google, and used 0 internal connections. Happy to make another post with my conversation + documentation to support this.
I reached out to the GCP through their regular channels. This is not a paid post, and we are not sponsored by Google in anyway.
Re: Burnt $72k testing Firebase and Cloud Run and almost went bankrupt
#268Earlier quoted context omitted.
> “Sorry, we don’t have the money” is a much better negotiating position than “can we please have our money back?” I agree but why would you like to be in either position anyway? The so-called cloud services are terribly overpriced when compared to traditional servers.
Done correctly they save a lot of IT time. Seem companies hire five 6 figure people to try and cut amazon bill by a couple of grand a month. Never understood spending 50-100k a month to maybe save 5k
Not really, computing done correctly is about avoiding all of the pitfalls and finding ways to get zero cost benefits, free computation out of necessary redundancy, etc. Selling cloud computing is about creating options around every pitfall and finding ways to charge for every mitigation that will be necessary and charge for redundancy in the mitigation strategy for the mitigation strategy..
Even if you pay for all the redundant managed blah they offer to not lose your business by having any single point of technical failure in their network, their billing and IAM are your single points of failure, if you diversify to multiple clouds all the guarantees either cloud offers is now pointless redundancy so you are paying 10X pricing for an inadequate redundancy layer.
If you look at Google's own model for computing, they didn't fall for this themselves, the computers they used were intentionally unreliable to not recursively pay for reliability and redundancy at any layer that can't provide the needed guarantee.
You can basically go all in with one of these clouds and become a franchise add-on with roughly the same rights as your average mcdonald's store owner, or you are managing a strategy that is far more complex because of the complexity of these offerings than just using metal and free software.
Re: Burnt $72k testing Firebase and Cloud Run and almost went bankrupt
#269As an ex-Googler working in a customer facing role in Cloud you did very well to get a $72k bill written off! It's definitely possible but requires a lot of approvals and pulling in a few favours. I went through the process to write off a ~$50k bill for one of my customers and it required action every day for 3 months of my life. Whoever helped you inside Google will have gone to a LOT of trouble, opened a bunch of t…
Thanks for sharing!
I have no idea what they did internally, but something like this was my guess. I only communicated through customer support channel and replied to emails, and shared my doc (which cited all the loopholes) with them.
It took them 10-15 days to get back and make a one-time good will contribution. The contribution didn't cover logging cost, so we did pay few hundred dollars.
Re: Burnt $72k testing Firebase and Cloud Run and almost went bankrupt
#270Deeply dissatisfying to read. Ex-Googler uses connections to get his (understandable!) cloud mistake refunded. Every time I read one of these stories, I get more and more convinced I will just simply never use scalable cloud tech for my side projects. I'm not going to risk my family's retirement savings on the all-too-possible chance that a small deep-implication error will cause runaway charges.
OP here. Your assumption is incorrect. I haven't been in touch with anyone in Google, and used 0 internal connections. Happy to make another post with my conversation + documentation to support this. I reached out to the GCP through their regular channels. This is not a paid post, and we are not sponsored by Google in anyway.