Live data from Hacker News

We deploy to production over 100 times a day

monzo.com

1–10 of 17 posts

Re: We deploy to production over 100 times a day

#2
Where I've worked, we would use Jenkins for "Continuous Delivery" but then require upper management signoff for every single damn change.

"We don't have time to test" -> Production poopfires -> "We need an out of band code review process"

The faces on these people when one proposes deployment in terms of basic git triggers on protected branches.

Re: We deploy to production over 100 times a day

#3
I really wonder how an HTTP application doesn't suffer from performance hits when it's based on 2000 micro services. Lets say even just 30 of those get used by a call to monzo.com. How does this not cause at least let's say ~300ms delay? I guess all the calls actually come from a local memory cache and there is almost never a real http call made to these microservices. Otherwise I have no idea how microservices are ever viable.

Re: We deploy to production over 100 times a day

#4
post #3

I really wonder how an HTTP application doesn't suffer from performance hits when it's based on 2000 micro services. Lets say even just 30 of those get used by a call to monzo.com. How does this not cause at least let's say ~300ms delay? I guess all the calls actually come from a local memory cache and there is almost never a real http call made to these microservices. Otherwise I have no idea how microservices are e…

Some of the calls can fan out in parallel, but in my experience it's not good for performance, even with fewer services. The remote call overhead certainly adds up, but another issue is that each service does redundant work. E.g. each service might have to fetch settings for the user in order to respond. You can refactor to eliminate this (e.g. fetch settings once and pass as a parameter to each call), but it's a lot more work to make this change across many services.

Re: We deploy to production over 100 times a day

#5
post #3

I really wonder how an HTTP application doesn't suffer from performance hits when it's based on 2000 micro services. Lets say even just 30 of those get used by a call to monzo.com. How does this not cause at least let's say ~300ms delay? I guess all the calls actually come from a local memory cache and there is almost never a real http call made to these microservices. Otherwise I have no idea how microservices are e…

If you look at something like K8s, it'll try and make contact with the other service running on the same node as itself, thus removing networking delays.

It's entirely feasible that all incoming requests will hit a single node and stay there.

Still, hitting multiple microservices in a single request will have an overhead when marshalling all the http requests and whatnot.

Re: We deploy to production over 100 times a day

#6
post #2

Where I've worked, we would use Jenkins for "Continuous Delivery" but then require upper management signoff for every single damn change. "We don't have time to test" -> Production poopfires -> "We need an out of band code review process" The faces on these people when one proposes deployment in terms of basic git triggers on protected branches.

I can relate to this, and it gets worse.

Sign-off would take 24 hrs to process, and the mysterious entity that signed off would have no context of our product or the changes, and no way to assess the risk.

Then, due to our lack of trusted regression testing, every damn change would take ~6 developers an entire work day to manually confirm that everything was working.

Why? We were measured on “number of tests.” The tech debt was too high to write quality tests before the next review, so we opted for quantity if we wanted promotion.

This was the hot path in a major (Fortune 10) financial company.

Re: We deploy to production over 100 times a day

#8

This is comical. Your company needs 3,000 deploys a month? You deliver 3,000 features a month? Your BAs/PMs identify 3,000 useful features a month? This isn't trolling - this is a call out of the absurdity you are presenting as a positive.

I work on a project with micro-services. Just looked into our repository: we merged around 1,100 pull requests in the last 30 days. That's a least this number of re-deployed services to our development cluster with a release train to staging and production clusters.

I'm not sure how many "feature" or stories that is. Maybe we need two to four pull requests per story.

Re: We deploy to production over 100 times a day

#9

This is comical. Your company needs 3,000 deploys a month? You deliver 3,000 features a month? Your BAs/PMs identify 3,000 useful features a month? This isn't trolling - this is a call out of the absurdity you are presenting as a positive.

Why would a deploy need to be a feature or something that PMs identify?

On a high-functioning team, you should have enough logging and monitoring that developers can identify many useful changes without PM involvement.

A deploy could be something as minor as "I noticed an error getting logged in production, here's a one-line change to fix it" or "This operation is running slowly, here's a tweak to the query so that it hits an index".

Re: We deploy to production over 100 times a day

#10
post #5
post #3

I really wonder how an HTTP application doesn't suffer from performance hits when it's based on 2000 micro services. Lets say even just 30 of those get used by a call to monzo.com. How does this not cause at least let's say ~300ms delay? I guess all the calls actually come from a local memory cache and there is almost never a real http call made to these microservices. Otherwise I have no idea how microservices are e…

If you look at something like K8s, it'll try and make contact with the other service running on the same node as itself, thus removing networking delays. It's entirely feasible that all incoming requests will hit a single node and stay there. Still, hitting multiple microservices in a single request will have an overhead when marshalling all the http requests and whatnot.

I worked at Monzo a while back, and, while the 'Platform' (~= SRE) team were brilliant, and did what you describe along with much much more, regardless the performance impact of thousands of microservices could be characterised as approximately "what you would expect". Hundreds of thousands a month on AWS to service a few million customers, a whole team needing to work on a project for [I can't recall how many] months to write some bodge-y fixes so app load could be brought under 10 seconds, lots of pathological request paths efflorescing in the service graph, etc.

That being said, there was genuinely need for microservices, to an extent. A bank's architecture is very different from a CRUD web app. Most of the code running wasn't servicing synchronous HTTP requests, but was doing batch or asynchronous work related (usually at two or three degrees of separation) to card payments, transfers, various kinds of fraud prevention, onboarding (which was an absolutely colossal edifice, very very different from ordinary SaaS onboarding), etc.

So we'd have had lots of daemons and crons in any case. And, to be fair, we started on Kubernetes before it was super-trendy and easy to deploy - it very much wasn't the 'default' choice, and we had to do a lot of very fundamental work ourselves.

But yeah, in my view we took it too far, out of ideological fervour. Dozens of - or at most a hundred-ish - microservices would have been the sweet spot. Your architectural structure doesn't need to be isomorphic to your code or team structure.

Post reply on HN