Tell HN: AWS connectivity issues, but health dashboard says everything fine
51–60 of 67 posts
Re: Tell HN: AWS connectivity issues, but health dashboard says everything fine
#52Earlier quoted context omitted.
My company is not going to spend hundreds of thousands of dollars or more, and months or even years of effort, and add additional constraints to the given pool of candidates we are hiring for, to migrate to GCP or Azure or DigitalOcean or Hetzner or wherever is considered more trendy than AWS right now due to "a lack of transparency" lmao. I would look completely incompetent to even suggest the idea to anyone interna…
Your company is hiring and retaining people who can't work with tooling outside Amazon Web Services?
Re: Tell HN: AWS connectivity issues, but health dashboard says everything fine
#53Earlier quoted context omitted.
My company is not going to spend hundreds of thousands of dollars or more, and months or even years of effort, and add additional constraints to the given pool of candidates we are hiring for, to migrate to GCP or Azure or DigitalOcean or Hetzner or wherever is considered more trendy than AWS right now due to "a lack of transparency" lmao. I would look completely incompetent to even suggest the idea to anyone interna…
Your company is hiring and retaining people who can't work with tooling outside Amazon Web Services?
Re: Tell HN: AWS connectivity issues, but health dashboard says everything fine
#54Everyone seems to overlook the point here. That yet again Amazon were slow as hell to be honest with their customers. I get it up down reports help but why do you keep using a service which lies to you about availability. I've read on HN in the past how the dashboard can only be updated to reflect an issue with approval. (Comments section on a similar posting, believe it if you wish). So why not move to a hosting com…
The core question is: what constitutes degraded service? Would you say a service is experiencing downtime every time a 500 response is served? If you're serving millions to billions of requests/sec it seems a bit disproportionate to marka service down after a single 500 error, so then you need to work out some kind of acceptable threshold.
What about latency? Again you're just going to draw a line in the sand somewhere.
You end up with this big mix of metrics that define service quality, so you then have a kind of meta problem of deciding which metrics you should alert users on. Get too trigger happy and it's going to cost you money and customer trust, and your customers are going to get alert fatigue when it turns out the issue you alerted them about was more of a false alarm. Set the bar too high and you'll have angry customers wondering wtf is going on.
All that to say I don't think there's a right answer.
Re: Tell HN: AWS connectivity issues, but health dashboard says everything fine
#55Earlier quoted context omitted.
But your company is willing to accept poor service and as a result spend more money with the same provider to ensure continuity. So essentially you reward Aws hiding their stats. As they can claim high uptime figures and when an outage happens it's the users fault for not spending enough money with them to have many many instances around the availability zones to ensure your covered the Aws mess up. I get it redundan…
If you are willing to host your critical infra on some dodgy startup alternative that might go away in 3 months because you refuse to bend on your personal values and separate them from what the typical organization actually cares about, best of luck. I know HN tends to loves the underdog, but there is a time and place for that, and a time and place to accept what you need to do to keep your services online.
Re: Tell HN: AWS connectivity issues, but health dashboard says everything fine
#56It does seem to be a networking issue. I have a ec2 instance in us-east-2 that is accessible through a "Global Accelerator" but not externally through my ISP. That ec2 instance can talk to other ec2 instances that are on us-east-2 - but none of those other instances are accessible externally.
Re: Tell HN: AWS connectivity issues, but health dashboard says everything fine
#57Earlier quoted context omitted.
If you are willing to host your critical infra on some dodgy startup alternative that might go away in 3 months because you refuse to bend on your personal values and separate them from what the typical organization actually cares about, best of luck. I know HN tends to loves the underdog, but there is a time and place for that, and a time and place to accept what you need to do to keep your services online.
So your logic is to accept poor quality service to keep your service online rather than trying to do better and improve service. So you are saying that rather than rewarding a company trying to do better just accept poor service from Aws.How is this better than "hosting on some dodgy start-up" This is nothing to do with my personal beliefs or opinion I'm trying to understand why it's accepted from Aws but not others…
Re: Tell HN: AWS connectivity issues, but health dashboard says everything fine
#58Everyone seems to overlook the point here. That yet again Amazon were slow as hell to be honest with their customers. I get it up down reports help but why do you keep using a service which lies to you about availability. I've read on HN in the past how the dashboard can only be updated to reflect an issue with approval. (Comments section on a similar posting, believe it if you wish). So why not move to a hosting com…
I understand the frustration, but Im not convinced monitoring at large scale is that straightforward. The core question is: what constitutes degraded service? Would you say a service is experiencing downtime every time a 500 response is served? If you're serving millions to billions of requests/sec it seems a bit disproportionate to marka service down after a single 500 error, so then you need to work out some kind o…
But, what ended up happening was a competitor who didn't have a status page at all would use our status page against us in the sales process. They just never mentioned their lack of a status page to compare to.
This was the same competitor who went 100% down for ~4 days during the busiest month of the year and only posted updates to a private Facebook group. There was data loss that was never publicly admitted to.
So, yeah, we implemented reasonable boundaries on what constitutes a post to the status page. We also adopted a new status page provider that let us get more granular with categorizing posts, and allowing users to subscribe to only "urgent" channels that pertain to them.
Re: Tell HN: AWS connectivity issues, but health dashboard says everything fine
#59Everyone seems to overlook the point here. That yet again Amazon were slow as hell to be honest with their customers. I get it up down reports help but why do you keep using a service which lies to you about availability. I've read on HN in the past how the dashboard can only be updated to reflect an issue with approval. (Comments section on a similar posting, believe it if you wish). So why not move to a hosting com…
I understand the frustration, but Im not convinced monitoring at large scale is that straightforward. The core question is: what constitutes degraded service? Would you say a service is experiencing downtime every time a 500 response is served? If you're serving millions to billions of requests/sec it seems a bit disproportionate to marka service down after a single 500 error, so then you need to work out some kind o…
Re: Tell HN: AWS connectivity issues, but health dashboard says everything fine
#60Earlier quoted context omitted.
AWS is the 800 pound gorilla in the cloud space. Are any of the other cloud providers better with customer honesty?
I'd ask if there any cloud providers worse with customer honesty instead.