Earlier quoted context omitted.
I think also just finally getting to the point where they have a tech debt burden comparable to everyone else. Starting fresh has its advantages...for a while.
Even on HN there are some former AWS employees who talk about how its all stitched together and flying on a wing a prayer. Apparently the on-call is just a traumatizing experience. It will take real damage to revenue for the management to pay that debt off.
AWS down again?
111–120 of 257 posts
Re: AWS down again?
#112Earlier quoted context omitted.
Even if you can manage more uptime on your own than through the cloud (which I doubt), being on the cloud means downtime is correlated with downtime of other services. That's usually a good thing. Your customers will be more understanding if your outage is part if a wider outage that makes national news. Any services you integrate with are likely down too. If two services with 99% uncorrelated uptime together drops t…
An our on premise linux server recently reached an uptime of 1000 days. Yes, days.
Re: AWS down again?
#113I think this adds some momentum to the pendulum swinging back the other way. Maybe cloud teams can patch your services better than your in-house team can (See 2 critical issues in Azure the last 3 months, _caused_ by MS itself). Maybe the cloud has a higher uptime than your on-premise infrastructure (see the AWS, Azure outages). Make sure to compare the actual outage time v.s. the stats doctored by various political…
Even if you can manage more uptime on your own than through the cloud (which I doubt), being on the cloud means downtime is correlated with downtime of other services. That's usually a good thing. Your customers will be more understanding if your outage is part if a wider outage that makes national news. Any services you integrate with are likely down too. If two services with 99% uncorrelated uptime together drops t…
Make things that fail with everyone else. Your customers will just have to get over it because everyone else is down.
News flash: This is exactly why "overconsolidation" is a bad thing. The bigger, more complex, and integrated the player, the more devastating the eventual failure is as reality dictates you will build the most complex system possible until you outstrip your ability to mentally simulate, reason about, and debug it. You cannot create a thriving, resilient business ecosystem With everyone flocking to the same players.
Customers must come first. Not your books. Working, resilient solutions. Anything else is LARPing, and kicking the catastrophe can down the road to someone else to pay.
Re: AWS down again?
#114I've been getting the "We're sorry!" error for at least 15 minutes. No idea about regions or anything, I was just about to look up their docs about moving accounts between orgs etc. Anyone affected? Edit: Seems to be flapping between the 'sorry' error and a blank page. Thoughts and prayers with the SREs, if they call 'em that over there.
Same here. And as always, AWS status page [0] is completely green and useless. [0] https://status.aws.amazon.com/
Re: AWS down again?
#115Remind me why this is better than a 10$/year VPS
You can't possibly hope to match the reliability of AWS with a tiny VPS
Also, the rather rare and brief maintenance window from the provider is always middle of the night for all my customers.
Re: AWS down again?
#116Earlier quoted context omitted.
Even if you can manage more uptime on your own than through the cloud (which I doubt), being on the cloud means downtime is correlated with downtime of other services. That's usually a good thing. Your customers will be more understanding if your outage is part if a wider outage that makes national news. Any services you integrate with are likely down too. If two services with 99% uncorrelated uptime together drops t…
Wha? We host everything on our hardware (which is nothing special) and haven't had any downtime in this year (yet). And we're just another run of the mill dev shop, very far from "superstars" who work on these (supposedly extremely stable) platforms.
1. How easy can I access your physical servers ? 2. What happens if there is a catastrophic failure, for example local power outage or a major flooding 3. How secure is your server? Are you regularly patching your operation systems 4. If I want to run a project that requires double the capacity of your current hardware for a specific project, how long is it going to take to get it spun up?
Re: AWS down again?
#117I expect better of the community here. All it takes is a chance to take a cheap shot at one of the “big boys” and then all of a sudden the weasels come scampering out of the wood work. Seriously, those commenting “oh boy! Time to rethink this whole cloud thing!” You’re either so new to this stuff to have no experience to remember the days before cloud, you’re trolling because you’re high on nostalgia remembering the…
I have built and run my own infrastructure from the ground up, and I've been made to transition to the cloud.
The experience hasn't been great. It may simply be sour 'grapes' because, after all the expertise a whole generation has built up learning UNIX, and all the internet protocols (DNS, ARP, Email, reading RFCs, networking, routing) we get told that all that old stuff is just 'legacy,' and that we should retool around amazon's proprietary services instead.
Some of us old timers argued against this only to be shouted down by people who don't even understand TCP/IP.
The current generation couldn't invent the internet. You know why, cos they would never have the patience to spec it out like the old timers did. Go read a few RFCs and try to imagine a scrum team today putting as much thought into an up front design.
Today we'd just cruft together an MVP, solve only the interesting parts (or more likely the easy parts) and then move on, letting dashboards which lie to cover it up.
Many of us have been taking shots at those 'big boys' since the start of this trend.
Now that because of recent events and we have a chance to be heard you're telling us to be quiet. Why?What are you afraid of?
Re: AWS down again?
#118Still no post mortem from Monday's right? We're all still in the dark?
it's usually a month or more after a large outage to see the full breakdown on what happened. people expecting to see it the same week are kinda not living in reality.
Re: AWS down again?
#119I expect better of the community here. All it takes is a chance to take a cheap shot at one of the “big boys” and then all of a sudden the weasels come scampering out of the wood work. Seriously, those commenting “oh boy! Time to rethink this whole cloud thing!” You’re either so new to this stuff to have no experience to remember the days before cloud, you’re trolling because you’re high on nostalgia remembering the…
Re: AWS down again?
#120Anyone feel like Big Tech outages are happening more frequently recently? In recent months, we’ve seen Amazon Web Services, Facebook, Gmail, and Twitter go down. Are we at the point where the people who maintain the infrastructure are now completely different than the ones who built it, and are struggling to keep it running because they don’t understand it as well?
Twitter used to go down so frequently the "fail whale" was a whole cultural thing.
Big AWS outages are rare, but hardly new.
https://www.theregister.com/2017/03/01/aws_s3_outage/
https://arstechnica.com/information-technology/2012/10/amazo...