Live data from Hacker News

We reduced the cost of building Mastodon at Twitter-scale by 100x

blog.redplanetlabs.com

351–360 of 376 posts

Re: We reduced the cost of building Mastodon at Twitter-scale by 100x

#351
post #318
post #283

Earlier quoted context omitted.

there's this though: https://mastodon.redplanetlabs.com/timeline/local and this from the gGmbH trademark policy: > You may not use the Mastodon word mark, or any similar mark, in your domain name, unless you have written permission from Mastodon gGmbH.

Trademark law doesn’t work like that. You don’t get to license your trademark in the same way you get to license copyrighted works. Otherwise, people could just say you can’t use their trademark in any document that says something negative about them, and then successfully sue the press and angry customers for complaining about them.

The WordPress Foundation has the same restrictions concerning use of their trademark as part of a domain name:

https://wordpressfoundation.org/trademark-policy/

Re: We reduced the cost of building Mastodon at Twitter-scale by 100x

#352
I have often thought along similar lines, that the effort involved in building software seems to indicate a level of abstraction that is missing. The general theme of the comments seems roughly what you should expect from a very bold, paradigm shifting proposal. Good luck with your efforts and don't let this discourage you!

I will make one minor suggestion that I hope is constructive. I found the post difficult to read, largely because you rapid fire introduce a bunch of completely new concepts and propose a solution to many problems at once. You make a passing comparison to "just event sourcing and materialized views", although this was the easiest way for me to understand what you are doing. Starting from event sourcing and materialized views puts the reader on a ground they already understand, and moving on from there to why rama is better/what it adds on top, would be an easier transition.

Re: We reduced the cost of building Mastodon at Twitter-scale by 100x

#353

I think the marketing idea of this is amazing : I would probably never even consider learning and reading about such a framework if I heard of it straight up. But if you are really releasing a usable open source implementation of something performant that actually federates properly, that is a huge selling point that buys you a ton of respect up front.

whoops nevermind "Once we open-source our Mastodon implementation in two weeks, you’ll be able to run it within a single process using this build."

So no one will be able to run this except on the proprietary cloud

Re: We reduced the cost of building Mastodon at Twitter-scale by 100x

#355
post #25

I've seen many people describe frameworks like this - you know, first you have the slow back-end event-driven master database that you don't query live against, then you've got eventual-consistency flows against the various data-warehouses and data-stores and partitioned sharded databases in useful query-friendly layouts that you actually read live from... and I never see it clearly explained: how do you read a chang…

One strategy (somewhat common in lambda architectures) is to query both the long-term store and the in-flight operations, and blend the results. The in-flight stuff is both small and already in memory so it's pretty often trivially fast, even if blending the data is relatively complex.

That does limit you to operations/queries you can describe in this dual format, but pretty often that's fine. Or if you can relax read-after-write you can just ignore the in-flight stuff and read from the main store and then there are no (added) limitations.

Re: We reduced the cost of building Mastodon at Twitter-scale by 100x

#356
post #193

Earlier quoted context omitted.

Yes, I 100% agree with you. I would like something like this to succeed, and agree the problem is real. But what are the tradeoffs? There's nothing that comes with 100x benefit with no tradeoffs (side note: I worked on Google Code for a short while in 2008, concurrent with Github's founding ... I think Github moved a lot faster in a large part because they weren't dealing with distributed systems at first -- they had…

Rama is a much broader platform than a database, so the consistency semantics you get depend on how you use it. When using Rama, you're not mutating indexes directly like you do with a database, but adding source data that then gets materialized into any number of indexes. You get read-after-write consistency for any PStates in a streaming ETL colocated with the depot you appended to. This is if you do the depot appe…

OK, thanks for the response

I think many people are going to have problems programming with this consistency model, as they will with any that's different than a single machine. But that's basically "physics", so it's inevitable :)

But it seems like great work within the constraints -- look forward to learning more

I have indeed wondered why none of the cloud platforms have built more forward-looking tech like this -- instead it's copies of AWS and so forth

Re: We reduced the cost of building Mastodon at Twitter-scale by 100x

#357
post #332

Sounds like Rama is also useful for small scale applications (where high scalability isn’t needed), since it simplifies how they’re implemented. Is this the case — ie. would a TodoMVC app implemented in Rama also be much simpler than a traditional frontend/backend/database CRUD implementation?

See https://news.ycombinator.com/item?id=37139410

Re: We reduced the cost of building Mastodon at Twitter-scale by 100x

#358
post #207

I do C++ backend work in a non-web industry and this entire post is Greek to me. Even though this is targeted at developers, you need a better pitch. I get "we did this 100x faster" but the obvious followup question is "how" but then the answer seems to be a ton of flow diagrams with way too many nodes that tell me approximately nothing and some handwaving about something called P-States that are basically defined to…

In a typical architecture, the DB stores data, and the backend calls the DB to make updates and compile views. Here, the "views" are defined formally (the P-states), and incrementally, automatically updated when the underlying data changes. Example problem: Get a list of accounts that follow account 1306 "Classic architecture": - Naive approach. Search through all accounts follow lists for "1306". Super slow, scales…

I’m getting Noria[1] / Materialize / Readyset vibes from this, or perhaps even Samsa[2] ones. (Incidentally, I’d appreciate it if anyone could elaborate on the differences between the two.) Explicit inspiration? Parallel evolution?

[1] https://news.ycombinator.com/item?id=29615085

[2] https://martin.kleppmann.com/2015/03/04/turning-the-database...

Re: We reduced the cost of building Mastodon at Twitter-scale by 100x

#359
post #352

I have often thought along similar lines, that the effort involved in building software seems to indicate a level of abstraction that is missing. The general theme of the comments seems roughly what you should expect from a very bold, paradigm shifting proposal. Good luck with your efforts and don't let this discourage you! I will make one minor suggestion that I hope is constructive. I found the post difficult to re…

Thanks for the feedback. The post is meant to give a taste of Rama by showing what it can do for building a full application end-to-end. Next week, we'll be releasing a set of guides, tutorials, and documentation which introduces the concepts in a much gentler way. We'll also be releasing a build of Rama that you can download to try Rama out yourself, and an open-source repository of example code from the documentation.

Re: We reduced the cost of building Mastodon at Twitter-scale by 100x

#360

I appreciate the inversion/melding of the data model and compute. Im curious to know your perspective on two parts: How would multitenancy fit in to rama? Using your mastodon example, providing “hosted mastodon instances as a service”, where you _also_ allow for data governance, per customer encryption at rest, user IDP support, etc. Is it multiple single tenant rama deployments, running independent customers? Multit…

For now it's going to be on-prem, so each user will just have their own cluster. Things like E2E encryption are pretty easy to implement on top of Rama's existing primitives (there was a good question about this on the rama-user group yesterday https://groups.google.com/u/1/g/rama-user/c/jj-ILcoMjtk).

We'll likely have a fully managed cloud version in the future.

Riak was good technology at the time, but it was really hard to distinguish from other K/V databases and didn't really move the needle on core business metrics (like development cost). Rama dramatically changes the economics of building large-scale software. It will take awhile for many to grasp that, as is obvious from many of the comments here, but I expect that as more and more users have massive success with Rama, that understanding will come.

Post reply on HN