Live data from Hacker News

Migrating Uber's ledger data from DynamoDB to LedgerStore

uber.com

131–140 of 345 posts

Re: Migrating Uber's ledger data from DynamoDB to LedgerStore

#131
post #97

Earlier quoted context omitted.

I'd be curious as well to see a more complete cost-benefit analysis, and I'd be especially interested in labor cost. We don't know how much time and head count Uber committed to this project, but I would be impressed if they were able to pull this off with fewer than 6-8 people. We can use that to get a very rough lower-bound cost estimate. For example, AWS internally uses a rule of thumb where each developer should…

If the savings are 6 million per year, then in later years it should pay off since the development is a one time cost.

The cost doesn't suddenly drop to zero once development is done. Typically a system of this complexity and scale requires constant maintenance. You'll need someone to be on-call (pager duty) to respond to alarms, you'll need to fix bugs, improve efficiency, apply security patches, tune alarms and metrics, etc.

In my experience you probably need a small team (6-8 people) to maintain something like this. Maybe you can consolidate some things (e.g. if your system has low on-call pressure, you may be able to merge rotations with other teams, etc.) but it doesn't go down to zero.

Re: Migrating Uber's ledger data from DynamoDB to LedgerStore

#132

The original[1][2] articles are a better read IMO. The link is just a summary of the two with added spelling and grammatical errors that materially impact the meaning. 1. https://www.uber.com/blog/how-ledgerstore-supports-trillions... 2. https://www.uber.com/blog/migrating-from-dynamodb-to-ledgers...

Seems to happen with all our blog posts that appear on here (I work at Uber) - I don't get why the originals don't get upvoted but these rehashes do - are our titles just not as good?

Other than the comments about titles, the entire blogpost doesn't show for me with ublock. So I'll open it, see a picture of some birds, scroll around for a bit then give up.

Re: Migrating Uber's ledger data from DynamoDB to LedgerStore

#133
post #119

More power to them. At this point even technically decent teams/companies have given up on developing large, complex systems in favor of SaaS. After carefully evaluating our strategic course of action answer always is AWS. Its only team who propose alternative they have to justify rigorously how come they differ in conclusion.

My Amazon stock thanks you.

Re: Migrating Uber's ledger data from DynamoDB to LedgerStore

#134

Congrats to anyone who worked on it! However, I'm guessing the cost of just running this team be quite large and not significantly different from the savings (6M), and add on top of it the overhead of maintenance. Payments would not likely be a long-term bet as well, so kind of interesting why teams take up such projects ? Is it some kind of sunk-cost with the engineering teams you already have?

If you read the article the system was a layer on top of DynamoDB they updated it to use internal product Docstore which required adding a feature to Docstore. So it's not as involved as people make it out to be. Also records are immutable which makes a lot of things way easier.

Re: Migrating Uber's ledger data from DynamoDB to LedgerStore

#135

I pretty much never see engineering salaries factored into these types of savings projects. I assume because engineers are already viewed as a sunk cost or maybe it’s just because it’s way less tangible. Have seen many designs describe how X saves Y dollars but ignores the engineering effort to maintain and build it. Half the time I suspect it’s just so people have something to work on, rather than it being some crit…

A better strategy as a company this size would be to write the PRD for moving and then call AWS and negotiate.

Re: Migrating Uber's ledger data from DynamoDB to LedgerStore

#136
post #123
post #88

Earlier quoted context omitted.

At the moment they're just paying someone else to buy $5000 SSD's and run a database on them at many X markup.

There is no upper bounds to economy of scales. Maybe there is for the cents per GB of raw storage, but power usage, security, rent, and everything else scales too, and few of them have upper bounds on economy of scales.

Economies of scale generally have upper limits. Often when you approach the largest scale the existing market will supply you essentially need to become your own supplier which then runs into span of control issues. The organization needs to become competitive in that new market or their costs increase.

Keep scaling and eventually vertical integration ends up looking like a Soviet style planned economy. Your remote mining town needs some way for people to get soap etc so you open a store with it’s own supply chain etc etc.

Re: Migrating Uber's ledger data from DynamoDB to LedgerStore

#138
post #97

Earlier quoted context omitted.

I'd be curious as well to see a more complete cost-benefit analysis, and I'd be especially interested in labor cost. We don't know how much time and head count Uber committed to this project, but I would be impressed if they were able to pull this off with fewer than 6-8 people. We can use that to get a very rough lower-bound cost estimate. For example, AWS internally uses a rule of thumb where each developer should…

Not an engineer, but something like this takes 6-8 people working on only this for a full year?

That has been my experience, yes. You need one full-time manager, one full-time on-call/pager duty (usually a rotation), and then 4-6 people doing feature development, bug fixes, and operational stuff (e.g. applying security patches, tuning alarms, upgrading instance types, tuning auto-scaling groups, etc. etc.).

Maybe you can do it a bit cheaper, e.g. with 4-6 people, but my point is that there's an on-going cost of ownership that any custom-built solution tends to incur.

Amortizing that cost over many customers is essentially the entire business model of AWS :)

Re: Migrating Uber's ledger data from DynamoDB to LedgerStore

#139

Earlier quoted context omitted.

Not even click a search ad and fill in a contact form? When I’m on the other side of the table I do that. But perhaps I’m unique in that aspect? (I understand there won’t be any significant business without enterprise sales. But that’s not what I’m looking for at this stage.)

>When I’m on the other side of the table I do that No, you don't. There are many established storage solutions out there. If you're in the market for one, you can easily fill days, weeks or months vetting those. So, why would you bother dealing with a sales rep from a random one you never heard of before, and isn't used by anyone. You don't even provide any details on what makes it different or better from anything e…

Well the reason I’m working on this in the first place is that when I was on the other side of the table I was looking for one. I filled in the contact forms of a couple of different startups that had products somewhat in line with what I was looking for, and talked to their sales reps. Admittedly they weren’t as early stage as my project, but on the other hand they weren’t 100% focused on my use-case either.

I guess what I’m trying to say is that I was hoping that someone with a write intensive workload would want to spend some time evaluating a product built specifically for that. But perhaps I’m wrong? Even if your workload was 99% writes you’d rather go to some established player (e.g. MongoDB) with a product optimized for 50/50 read/write?

Re: Migrating Uber's ledger data from DynamoDB to LedgerStore

#140
post #3

Earlier quoted context omitted.

1T "records". Any given transaction can have N records. I'm assuming this includes Uber Eats as well.

Still, they have 10B rides in 2023 including Eats, say 75-100B since inception. What would be a record such that each transaction needs 10-15 on average?

> 75-100B

This seems low, off the bat. 15 years of Uber, 9 years of Uber Eats.

But even just looking at my most recent trip with Uber, there are 7 different records visible on the receipt. Not including backend recordkeeping that isn't exposed to the user (driver payments, driver loan repayments, revenue recognition, internal fees/records, etc).

Total trip amount, Trip fare, Booking fee, Tip, State fee, Payment #1 (trip itself), and Payment #2 (driver tip)

Now consider Uber Eats where there is (at least) one record for each item in an order...plus tax, tip, etc as always.

Then consider things like wait time charges, subscriptions, split charges, pending charges, chargebacks, refunds, disputes, blah blah blah.

An average of 10 records per customer transaction seems entirely reasonable.

Post reply on HN