Live data from Hacker News

AWS having major issues

status.aws.amazon.com

31–40 of 69 posts

Re: AWS having major issues

#31
post #21

Ok. If you're receiving errors, and you're NOT using opsworks, please respond. We're using opsworks too and have ~30 servers down. Maybe we all should be looking at the opsworks agent.

We started getting SQS errors from us-east-1 around 05:08 UTC.

We also have a bunch of files on S3 which Cross-Region Replication hasn't yet replicated... I think that depends on SQS as well.

Re: AWS having major issues

#32
post #29

So three different environments have had the ELB just release all of their clients and not come back onboard. Manual intervention was the only way. Whats the recommended way to monitor if instances have become detached or non responsive? Want to immediately alert slack or send an email, etc.

There are a lot of ways to integrate, but we use datadog and I generally find it excellent. Alerts to hipchat, pagerduty, etc, with AWS Cloudwatch integration. Of course, you have to assume the Cloudwatch API is working... :)

Re: AWS having major issues

#34
post #10
post #4

EU-West is also experiencing problems. A lot of our instances are currently in connection_lost status. EDIT: The console is working fine at the moment EDIT1: Apparently OpsWorks is hosted and managed by the North Virginia data center which is why all our opsworks instances in EU-West are experiencing issues too.

Frankfurt is up and running. Only the Web Console is slower.

[deleted]

Re: AWS having major issues

#35
We seem to have had issues and we're not using OpWorks. I'm trying to determine if our issues are related.

Our instances are docker hosts, network seemed to lag/stop when proxying traffic to the internal container IP addresses.

Our ASG spun up other instances but health checks reported "Insufficient Data". The web console also seems buggy (API requests are failing).

Re: AWS having major issues

#37
post #21

Ok. If you're receiving errors, and you're NOT using opsworks, please respond. We're using opsworks too and have ~30 servers down. Maybe we all should be looking at the opsworks agent.

having issues. Not using opsworks. It is not limited to that.

Re: AWS having major issues

#39

Earlier quoted context omitted.

We are getting errors across multiple AWS API's. It's nothing to do with Opsworks itself, rather it appears like there is an internal networking issue. Both SQS and SNS were erroring and now SQS has gone down completly with all requests timing out.

Looks like there's critical infrastructure in us-east-1 that's broken and causing a ripple effect across all of AWS Our platform is entirely hosted in ap-southeast-2 but we've had our EC2 instances deregistered and OpsWorks reporting them terminated where EC2 is showing them active and they're still reachable via SSH

Yeah, we don't use OpsWorks and had SQS/SNS/SES trouble as well -- thankfully those are not used to serve production traffic. From the set of services affected, it looks like Amazon's internal Kafka-like pub/sub system went down.

Re: AWS having major issues

#40
Ok, I can't tell if this is functioning normally. I am trying to launch a new instance. Numerous services have "increased API error rates" one even said "Elevated error rates"

All services panel were green indicating: Service is operating normally

I can't put confidence in this. Anyone have any info on whether this is resolved. I can do some on digital ocean, but I really need a few things on AWS. Confirmed up? COnfirmed working? Info?

Edit: Successfully launched an AMI instance and it deployed successfully and is accessible from shell and ip.

Post reply on HN