Live data from Hacker News

Lichess: Post-Mortem of Our Longest Downtime

lichess.org

51–55 of 55 posts

Re: Lichess: Post-Mortem of Our Longest Downtime

#51
post #36

Earlier quoted context omitted.

You are not wrong that this is puzzling, especially when viewed through the perspective lens of a professional with background in these areas (10 years). There are many red flags which beg questions. That said, I stopped taking them at their word years ago, this isn't the first time they've had dubious announcements following entirely preventable failures. In my mind, they really don't have any professional credibili…

This is so far out of line I wonder what the background is for this issue. Lichess is not emergency dispatch software running in a 911 Call Center, if they have an outage the cost is that users can't play online chess until it is fixed. Additionally, the founder of this open source project is objectively good at what he does. Exhibit the fact that he built and hosts a top 2 online chess platform that competes well ag…

We will have to disagree Kenneth.

Your idea of "so far out of line", would include any communication you disagree with, and is absent rational principles or social norms/mores basis, it is absurd.

I stuck to the objective issues in my previous post, you should too before making baseless claims.

Do some due dilligence on the business entities involved, peruse their github history (the deleted parts). Get a real picture about what's going on there. You'll find many contradictions if you dig.

The question on any critical IT professional's minds is how can you run the service given the resources claimed. Yes he runs the top traffic site for chess, and its done on a bespoke monolith.

You napkin math/sketch it out by required component services, and it quickly becomes clear that nothing adds up. When nothing is consistent, or supported, you examine your premises for contradictions and lies, which goes again back to credibility.

(Hint: https://trufflesecurity.com/blog/anyone-can-access-deleted-a...)

Re: Lichess: Post-Mortem of Our Longest Downtime

#52
post #48

Earlier quoted context omitted.

I'm going to assume that your question is genuine and sincere, and not meant sarcastically. If you read the following books by established experts, you should be able to rationally answer the question for yourself as to the why and the how. The subject matter involves torture for thought reform, real not fantasy. This differs from SERE training which is geared towards resisting information extraction. China by their…

You need to seek professional help.

No I don't, but you certainly do after trying to gaslight like that.

Can't tell if its pathological or malevolent... probably both. Pray that we never meet.

Thankfully, it is not such a simple thing to discredit when world renowned experts all agree and say a thing, and the longer they have been established the more one should listen.

Saying I need to seek professional help for repeating what's been documented by experts, yeah that is rich.

Re: Lichess: Post-Mortem of Our Longest Downtime

#53
post #9

Earlier quoted context omitted.

Wow. ~US$40k/mo running costs, with about US$5k/mo for server hosting: https://lichess.org/costs It looks like the servers are individually managed via OVH or similar, rather than running their own gear in co-location or similar. Wonder why?

Easy: If something is wrong with the physical gear it's OVH's problem rather than theirs. It also means no one has to ever go to the data center which is probably important for a geographically distributed team (I assume they are). Cheap, no-frills cloud is extremely underrated, IMO.

Underrated? The flip side is that hardware failures are still your problem like they would be with rolling your own hosting. I think they’re correctly rated for the position on the scale of traders that they provide.

Re: Lichess: Post-Mortem of Our Longest Downtime

#54

Earlier quoted context omitted.

Yup, and you also get to make AWS deal with OS upgrades, DB upgrades, backups, etc.

You have to pay 2x for multi-AZ or you get downtime for upgrades. And DB major version upgrades require manual effort unless you want to roll the dice on their new blue-green feature, which can take hours to fail or finish cutting over.

> You have to pay 2x for multi-AZ or you get downtime for upgrades.

Worse. In Single AZ deployments you get (short, but not that short or strongly bound) downtime for daily backups and when doing snapshots. Source:

- https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/USER_...: "During the automatic backup window, storage I/O might be suspended briefly while the backup process initializes (typically under a few seconds). [...] For MariaDB, MySQL, Oracle, and PostgreSQL, I/O activity isn't suspended on your primary during backup for Multi-AZ deployments because the backup is taken from the standby. ",

- https://docs.aws.amazon.com/AmazonRDS/latest/UserGuide/USER_...: "Amazon RDS creates a storage volume snapshot of your DB instance, backing up the entire DB instance and not just individual databases. Creating this DB snapshot on a Single-AZ DB instance results in a brief I/O suspension that can last from a few seconds to a few minutes, depending on the size and class of your DB instance.".

Not to mention that multi-AZ deployments incur extra transfer cost between zones - not between DB instances (this one is free, last time I checked), but between your compute deployments and DB instances, if your compute does not automatically follow the zone of the db host it talks to.

Re: Lichess: Post-Mortem of Our Longest Downtime

#55
post #51

Earlier quoted context omitted.

This is so far out of line I wonder what the background is for this issue. Lichess is not emergency dispatch software running in a 911 Call Center, if they have an outage the cost is that users can't play online chess until it is fixed. Additionally, the founder of this open source project is objectively good at what he does. Exhibit the fact that he built and hosts a top 2 online chess platform that competes well ag…

We will have to disagree Kenneth. Your idea of "so far out of line", would include any communication you disagree with, and is absent rational principles or social norms/mores basis, it is absurd. I stuck to the objective issues in my previous post, you should too before making baseless claims. Do some due dilligence on the business entities involved, peruse their github history (the deleted parts). Get a real pictur…

How do you have my first name?
Post reply on HN