Live data from Hacker News

An Update from Robinhood’s Founders

blog.robinhood.com

191–200 of 295 posts

Re: An Update from Robinhood’s Founders

#191

Earlier quoted context omitted.

Google doesn’t target zero downtime. The marginal cost is too high. For important services (like Search page and ads) they aim for 5 nines uptime (99.999%), which translates to 5 minutes of downtime per year. https://en.m.wikipedia.org/wiki/High_availability

I know, I worked there for 3 years on million node clusters

Then certainly you understand the importance of SLOs, how SLAs regulate reliability and feature velocity.

Let’s say I’m RobinHood. Let’s pick an SLO. I think three nines monthly SLO is a good start, that budgets ~45 minutes of down time per month. Maybe I can argue for a more aggressive SLO, but let’s pick this one - because I think it will keep users relatively happy as trades aren’t blocked for more than an hour at worst. I drive an agreement with stakeholders that if we needle out of this SLO, we drop all feature work and focus on hardening reliability.

RobinHood was out for a whole day. This is unacceptable. It points to a complete organizational fuck up - product and feature development have too much power and priority at the expense of reliability.

I’m not sure that RobinHood has ever heard of SLOs or reliability engineering. I really hope their leadership is smart enough to hire and empower the right people that will drive organizational change.

Re: An Update from Robinhood’s Founders

#192
post #120

This is such an empty update. At the very least, they should have published a detailed postmortem or committed to one by a certain date. How are we supposed to know that they have learned their lessons?

I don’t work for them, but I am pretty sure we can blame the litigious nature of this industry for the lack of detail in the postmortem. Not everyone can afford to be cloudflare :) Even for Cloudflare, I thought the company will get sued out of existence after the proxy data leak, but finance industry/SEC etc is a completely different ballgame.

I believe it's the fear of litigation rather than actual litigation. Other companies also manage to publish postmortems and don't get sued out of existence.

Re: An Update from Robinhood’s Founders

#193

I know quite a few people that were personally affected by this and lost money due to the two outages and they are all pulling their money from Robinhood. The fact that they can't offer any compensation might be a big problem for them, since they already have zero trading fees, which is what most brokerages offer as compensation. Personally it doesn't pass the smell test for me. The load was much higher the previous…

This isn't something new, downtime is the norm for Robinhood. Anyone trusting them with more than play money is foolish.

This is the correct sentiment. People who put anything more than play money into Robinhood should not be surprised when their financial life is ruined.

Re: An Update from Robinhood’s Founders

#194

Just use Square’s Cash App! Free stock trades AND you can buy fractional shares AND a bunch of other stuff like P2P payments and bitcoin. I work there and so can say with some authority that we can handle more volume without going down than RH can.

If you actually work at Square, it's poor form to advertise in this manner.

Actually, bayonetz's posting is the only useful one in the comments for this article. Most of us are here for information from actual industry insiders, and this qualifies.

Here's some more inside info ...

If your "financial app" provider doesn't have a banking charter, run. None of the recent trendy fintech companies have a charter, and are thus clown cars.

Re: An Update from Robinhood’s Founders

#195

Earlier quoted context omitted.

I know, I worked there for 3 years on million node clusters

Then certainly you understand the importance of SLOs, how SLAs regulate reliability and feature velocity. Let’s say I’m RobinHood. Let’s pick an SLO. I think three nines monthly SLO is a good start, that budgets ~45 minutes of down time per month. Maybe I can argue for a more aggressive SLO, but let’s pick this one - because I think it will keep users relatively happy as trades aren’t blocked for more than an hour at…

Why would they burden themselves and their feature velocity with SLOs/SLAs when they can build a 5 billion dollar company insanely quickly even though they have downtime?

The users are not saying "We measured your 5 9's and I'm going to quit if you have 6 minutes more downtime"

Sure they lose some users who get annoyed, but they have a 5.6 billion dollar company, some users will go, a lot more are coming

Re: An Update from Robinhood’s Founders

#196

Earlier quoted context omitted.

If this were an outage directly caused by a natural disaster, I could understand. This outage was an availability problem. This clearly points to some prioritization problems within the leadership layers if robust and resilient infrastructure was not emphasized. The prioritization problems may not be due to ignorance or malice though, and may be justifiable if there are other fires that are burning brighter. It's sti…

Sometimes you just have to cut them some slack. Have you engineered a highly available cluster before? I'm not talking about the hot-standby postgres master that gets called on once every 2 years, but I'm talking about a 180 node Cassandra cluster thats doing 15,000 writes a second 24/7 and peaking at 60,000 writes a second every day, and you have to do node replacements every week or two because of the high load. Or…

Kudos, these are moderate sized systems you've built over your career. There are lot bigger and more mission critical systems in the world and you might build them one day.

I understand GP's tone wasn't exactly nice here. But here's the rub with RH's outage. RH is unfortunately in an industry (Finance, Healthcare, Aviation, Food, etc.) where people _need_ to trust them to be successful. The consequences of failure in these industries is very catastrophic not only for them but their clients. Sure failures happen but the scale at which RH has failed and the lukewarm response they've put out has pissed off people. I don't recall any brokerage, old or new, that has failed so catastrophically and has responded to it so poorly. If you think you have a worse example, I am all ears.

Re: An Update from Robinhood’s Founders

#197

Earlier quoted context omitted.

Sometimes you just have to cut them some slack. Have you engineered a highly available cluster before? I'm not talking about the hot-standby postgres master that gets called on once every 2 years, but I'm talking about a 180 node Cassandra cluster thats doing 15,000 writes a second 24/7 and peaking at 60,000 writes a second every day, and you have to do node replacements every week or two because of the high load. Or…

Kudos, these are moderate sized systems you've built over your career. There are lot bigger and more mission critical systems in the world and you might build them one day. I understand GP's tone wasn't exactly nice here. But here's the rub with RH's outage. RH is unfortunately in an industry (Finance, Healthcare, Aviation, Food, etc.) where people _need_ to trust them to be successful. The consequences of failure in…

[deleted]

Re: An Update from Robinhood’s Founders

#199

Earlier quoted context omitted.

They denied that on Twitter: https://twitter.com/AskRobinhood/status/1234861941413351434

So what of the screenshot in the original twitter post? Was it doctored? Showing GMT?

I don't know if it had anything to do with leap year, but I also checked dev tools and saw the same issue (requests for market data on March 3, 2020 on March 2, 2020 8AM PST). However, it was busted for both the website as well as the Android app (and I'm guessing iOS too) so it doesn't seem like it's purely a client-side problem unless all of their clients were built from the same source.

Re: An Update from Robinhood’s Founders

#200

Earlier quoted context omitted.

If this were an outage directly caused by a natural disaster, I could understand. This outage was an availability problem. This clearly points to some prioritization problems within the leadership layers if robust and resilient infrastructure was not emphasized. The prioritization problems may not be due to ignorance or malice though, and may be justifiable if there are other fires that are burning brighter. It's sti…

Sometimes you just have to cut them some slack. Have you engineered a highly available cluster before? I'm not talking about the hot-standby postgres master that gets called on once every 2 years, but I'm talking about a 180 node Cassandra cluster thats doing 15,000 writes a second 24/7 and peaking at 60,000 writes a second every day, and you have to do node replacements every week or two because of the high load. Or…

If you used Scylla you'd have only needed 90 nodes. (Don't believe the instability rumours)
Post reply on HN