Live data from Hacker News

An Update from Robinhood’s Founders

blog.robinhood.com

221–230 of 295 posts

Re: An Update from Robinhood’s Founders

#221

Earlier quoted context omitted.

Can you give me a concrete example of a massive distributed system that has zero downtime? Because the largest distributed system I have seen and worked on was at Apple (or maybe DFP at Google) - and even though they had some of the smartest people in the world and literally billions of dollars behind them, there were still an endless list of problems and downtime events. Spoiler alert: It doesn't exist.

Google doesn’t target zero downtime. The marginal cost is too high. For important services (like Search page and ads) they aim for 5 nines uptime (99.999%), which translates to 5 minutes of downtime per year. https://en.m.wikipedia.org/wiki/High_availability

As an ex telco guy all I can say is "amateurs" :-)

Re: An Update from Robinhood’s Founders

#222

Earlier quoted context omitted.

This is all true for a company that is actually pushing any boundaries as opposed to failing pathetically at a well solved problem.

Robinhood opened up stock trading to a large portion of the population that would otherwise not have been interested in traditional trading platforms with high commissions. Their success helped to pressure companies such as TD and Schwab to mostly get rid of commissions as well, which is great for the average trader I think Robinhood has a lot of problems, but to say they're not pushing any boundaries ignores the hug…

This is the 21st century low cost trading has been around for several decades now

Re: An Update from Robinhood’s Founders

#223
post #31

Earlier quoted context omitted.

> I think they have really democratized stock trading I would think Vanguard did that already. Most people should be trading ETFs, not individual stocks.

I believe you still need a broker for ETFs, which usually carry brokerage fees.

Most brokers don't charge commission on a large variety of select low cost efts

Re: An Update from Robinhood’s Founders

#224

Earlier quoted context omitted.

If you actually work at Square, it's poor form to advertise in this manner.

Actually, bayonetz's posting is the only useful one in the comments for this article. Most of us are here for information from actual industry insiders, and this qualifies. Here's some more inside info ... If your "financial app" provider doesn't have a banking charter, run. None of the recent trendy fintech companies have a charter, and are thus clown cars.

Fidelity offers banking services and doesn't have a banking charter but they aren't a "clown car," they are one of the largest financial institutions in the world.

Re: An Update from Robinhood’s Founders

#225
post #40
post #7

Earlier quoted context omitted.

This happened to us at Hustle years ago. Basically if you run on AWS there’s a DNS server provided inside each VPC that usually works fine but which has no observable load metrics etc... so you don’t really know you are slamming it and are about to have a problem unless you audit your entire codebase. Why? Well that tiny DNS server has certain capacity constraints and if you don’t cache DNS lookups by using a http/ht…

I read this and thought, “surely there’s an OS-level DNS cache?” Apparently not on Linux! https://stackoverflow.com/questions/11020027/dns-caching-in-...

That's misleading. The way that this has worked for decades on Linux-based operating systems and on Unices is that one installs a local caching DNS proxy, choosing one of the many available: ISC's BIND, Bernstein's dnscache, unbound, dnsmasq, PowerDNS, MaraDNS, and so forth.

Every Unix system having a local caching DNS proxy was and is as much a norm as every Unix system having a local MTS. A quarter of a century ago, this would have been BIND and Sendmail. Things are more variable, now.

To illustrate that this was considered the norm, here is a random book from the 1990s. Smoot Carl-Mitchell's _Practical Internetworking with TCP/IP and UNIX_ says, quite unequivocally:

> You must run a DNS server if you have Internet connectivity. The most common UNIX DNS server is the Berkeley Internet Name Daemon (BIND), which is part of most UNIX systems.

People sometimes think that this is not the case nowadays, and the fact that a computer is a personal computer magically means that a Unix or Linux-based operating system should offload this task and not perform it locally. They are wrong, and that is DOS Think. Ironically, they don't even get to play the resource allocation card nowadays. The amount of memory and network bandwidth that needs to be devoted to caching proxy DNS service on a personal computer is dwarfed by the amounts nowadays consumed by WWW browsers and HTTP(S).

There's no similar argument for a node in a datacentre.

Ideally, not only should every machine have a (forwarding/resolving) caching proxy DNS server, every organization (or LAN, or even machine) should have a local root content DNS server. A lot of (quite valid) DNS lookups stop at the root with fixed or negative answers. Stopping that from leaving the site/LAN/machine is beneficial.

Ironically, putting a forwarding caching proxy DNS service on the local end of any congested, slow, expensive, or otherwise limited link is advice that I and others have been handing out for over 20 years. It's exactly what one should be doing with things like Amazon's non-local proxy DNS server limited to 1024 packets/second/interface.

* http://jdebp.uk./FGA/dns-server-roles.html#ChoosingProxy

So the question is not whether there a local DNS cache mechanism exists. It's whether it's set up by the company dishing out the VMs, and if not why not. Amazon provides instructions on how to add dnsmasq, and clearly labels this as how to reduce DNS outages. So it's not even the case that Amazon is wrongly discouraging having local caching proxy DNS servers.

* https://aws.amazon.com/premiumsupport/knowledge-center/dns-r...

Re: An Update from Robinhood’s Founders

#226

Earlier quoted context omitted.

It is not about scale, it is about the fact that people lost real money. If you can’t make it work you should not be in that business, and I don’t really care how hard they work. I’m taking my account off their platform.

Is this your first day of trading or something? People lose money in trading all the time, for hundreds of reasons and some of those reasons are infrastructure downtime. If your risk profile doesn't reflect that, maybe you should take your money out of trading altogether.

I've been trading for years, would not keep a penny on that platform. They've effectively cut off all liquidity for their customers for at least 2 days during high market volatility. You are missing out on tax loss harvesting, buying dips etc.

Re: An Update from Robinhood’s Founders

#227

Earlier quoted context omitted.

If this were an outage directly caused by a natural disaster, I could understand. This outage was an availability problem. This clearly points to some prioritization problems within the leadership layers if robust and resilient infrastructure was not emphasized. The prioritization problems may not be due to ignorance or malice though, and may be justifiable if there are other fires that are burning brighter. It's sti…

Sometimes you just have to cut them some slack. Have you engineered a highly available cluster before? I'm not talking about the hot-standby postgres master that gets called on once every 2 years, but I'm talking about a 180 node Cassandra cluster thats doing 15,000 writes a second 24/7 and peaking at 60,000 writes a second every day, and you have to do node replacements every week or two because of the high load. Or…

Does the duration of their downtime suggest a “1/1000” unmonitored oversight? Or is it more like a threshold that was meet and probably could/should have been observed?

And FWIW, they have down time every day and weekend, at least in a virtual sense; the load does drop off in a very real sense too. You are spiritually correct, they should pull together and sort it out, and they owe nobody money here (don’t use a discount broker if you want some sort of guarantee about trades) but as a general rule you should ever feel too sorry for banker under just about any circumstances. The harshest lesson here, for everybody, was the only thing they would do for you was give you some commission free trades but that won’t work with this one, so a non-apology is what you get.

Re: An Update from Robinhood’s Founders

#228

Earlier quoted context omitted.

On the other hand, they've had plenty of time and resources to do just that in a reliable fashion, it's not like it's one guy in his bedroom (I hope!). It's not like they are volunteers doing this open source for the community, they are getting paid (very well, I assume) to run the system. And Management is getting paid (even better, I assume) to make sure the priorities are right and correct decisions are taken. "Wh…

Can you give me a concrete example of a massive distributed system that has zero downtime? Because the largest distributed system I have seen and worked on was at Apple (or maybe DFP at Google) - and even though they had some of the smartest people in the world and literally billions of dollars behind them, there were still an endless list of problems and downtime events. Spoiler alert: It doesn't exist.

Can you name a reputable brokerage that was down all of Monday and Tuesday this week?

Spoiler alert: it doesn't exist

Re: An Update from Robinhood’s Founders

#229

Earlier quoted context omitted.

It’s another example of why DevOps has become a buzzword and most teams just pay lip service to it.

Everything has outages. Is this the new narrative now that we've moved on from the leap year thing? That RobinHood is just a bunch of shitty engineers? There are no public details about the root cause. I think RH is bad for people in general, but this pile-on is outrageous.

> That RobinHood is just a bunch of shitty engineers?

It is confirmed they are worse than virtually any reputable brokerage. It might not be their fault directly but its 2020, not 1998

Re: An Update from Robinhood’s Founders

#230

Earlier quoted context omitted.

Sometimes you just have to cut them some slack. Have you engineered a highly available cluster before? I'm not talking about the hot-standby postgres master that gets called on once every 2 years, but I'm talking about a 180 node Cassandra cluster thats doing 15,000 writes a second 24/7 and peaking at 60,000 writes a second every day, and you have to do node replacements every week or two because of the high load. Or…

Does the duration of their downtime suggest a “1/1000” unmonitored oversight? Or is it more like a threshold that was meet and probably could/should have been observed? And FWIW, they have down time every day and weekend, at least in a virtual sense; the load does drop off in a very real sense too. You are spiritually correct, they should pull together and sort it out, and they owe nobody money here (don’t use a disc…

I think you may be focusing on the finger instead of the thing that it's pointing at.

The post reads to me like all those examples were meant to be concrete examples to drive home a more general argument that complex systems are, well, complex, and that there's an element of hubris in taking potshots from the peanut gallery.

Post reply on HN