Live data from Hacker News

Removing fsync from our local storage engine

fractalbits.com

61–70 of 88 posts

Re: Removing fsync from our local storage engine

#62
post #60
post #38

This design ACKs writes that aren't yet durably persisted (to the journal or data areas). That might be ok, but it might not. It's certainly unusual not to at least persist the journal update.

nop. we will not ack any write which is not in data or journal. please check the put details in the blog.

You initiate a write to the journal, but do not sync it before ACKing to the client.

Re: Removing fsync from our local storage engine

#64
post #62
post #60

Earlier quoted context omitted.

nop. we will not ack any write which is not in data or journal. please check the put details in the blog.

You initiate a write to the journal, but do not sync it before ACKing to the client.

journal file was pre-alloacated and we use direct-io for journal write so no need to call fsync.

Re: Removing fsync from our local storage engine

#67
post #64
post #62

Earlier quoted context omitted.

You initiate a write to the journal, but do not sync it before ACKing to the client.

journal file was pre-alloacated and we use direct-io for journal write so no need to call fsync.

Again, it is not durably persisted before acking to the client. Like I said earlier, that might be fine for your durability model, but it is unusual.

Re: Removing fsync from our local storage engine

#68

Earlier quoted context omitted.

:-/ it’s a statistical guarantee in the first place. A successful commit in a durable storage engine just needs to achieve some finite level of durability, like “10^-7 probability of loss per year”. The durability is a property of the whole system, and it is possible to achieve durability without fsync, you just may have a hard time explaining what the durability is, how you calculated it, and what the evidence or ju…

I used to say this as well but like.. industry has, for a long time now equated “durable” with “stored on disk”. Any DBA will assume that’s what it means, and use that fact when they work out the replication they need either in clustering or in raid. If you’re building a data storage system and are using the term “durable” to mean “it’s in RAM on three virtual machines”, for example, I don’t think it’s unfair to say…

I forget the product, but more than a decade ago I remember someone broke out their durability into a table with columns for all the settings their data store offered between “ram on one node” and “fsync confirmed on a quorum of nodes’ disks” and rows for example failure cases ranging from “unexpected reboot of one machine” to “catastrophic loss of quorum-1 machines”. Cells were data loss risks from “prevented” to “possible” to “likely”.

That was very helpful when choosing durability levels.

Re: Removing fsync from our local storage engine

#69
post #67
post #64

Earlier quoted context omitted.

journal file was pre-alloacated and we use direct-io for journal write so no need to call fsync.

Again, it is not durably persisted before acking to the client. Like I said earlier, that might be fine for your durability model, but it is unusual.

We would wait for Bss data and journal DirectIO and the acking (sending response back to api_server) in the callback function. What you are implying is what s3 actually doing and you can get see from their paper[1] and we are stronger than that.

[1]https://www.amazon.science/publications/using-lightweight-fo....

Re: Removing fsync from our local storage engine

#70

Earlier quoted context omitted.

:-/ it’s a statistical guarantee in the first place. A successful commit in a durable storage engine just needs to achieve some finite level of durability, like “10^-7 probability of loss per year”. The durability is a property of the whole system, and it is possible to achieve durability without fsync, you just may have a hard time explaining what the durability is, how you calculated it, and what the evidence or ju…

I used to say this as well but like.. industry has, for a long time now equated “durable” with “stored on disk”. Any DBA will assume that’s what it means, and use that fact when they work out the replication they need either in clustering or in raid. If you’re building a data storage system and are using the term “durable” to mean “it’s in RAM on three virtual machines”, for example, I don’t think it’s unfair to say…

I don’t have any respect for the viewpoint that “durable” is equatable with “stored on disk”, and I don’t want to spend time accommodating that viewpoint. It is just an oversimplification in a very bad way.

AFRs and discussions about different failure scenarios are the bare minimum. The bare minimum for scenarios is disk loss, total machine loss, and data center loss. This is just my take on things. I don’t care if something is on disk or not. I do care what happens when a sector on disk goes bad, when a faulty power supply destroys all the disks in a machine, or when a data center floods.

That forces you to think about things like whether you want to turn on synchronous replication.

Post reply on HN