Async I/O on Linux in databases
blog.canoozie.net
Async I/O on Linux in databases
1–10 of 100 posts
Re: Async I/O on Linux in databases
#2Re: Async I/O on Linux in databases
#3Presumably the intent record is large (containing the key-value data) while the completion record is tiny (containing just the index of the intent record). Is the point that the completion record write is guaranteed to be atomic because it fits in a disk sector, while the intent record doesn't?
Re: Async I/O on Linux in databases
#4The recovery process is to "only apply operations that have both intent and completion records." But then I don't see the point of logging the intent record separately. If no completion is logged, the intent is ignored. So you could log the two together. Presumably the intent record is large (containing the key-value data) while the completion record is tiny (containing just the index of the intent record). Is the po…
Re: Async I/O on Linux in databases
#5The recovery process is to "only apply operations that have both intent and completion records." But then I don't see the point of logging the intent record separately. If no completion is logged, the intent is ignored. So you could log the two together. Presumably the intent record is large (containing the key-value data) while the completion record is tiny (containing just the index of the intent record). Is the po…
Write intent record (async)
Perform operation in memory
Write completion record (async)
* * Wait for intent and completion to be flushed to disk * *
Return success to clientRe: Async I/O on Linux in databases
#6During recovery, I only apply operations that have both intent and completion records. This ensures consistency while allowing much higher throughput. “
Does this mean that a client could receive a success for a request, which if the system crashed immediately afterwards, when replayed, wouldn’t necessarily have that request recorded?
How does that not violate ACID?
Re: Async I/O on Linux in databases
#7The recovery process is to "only apply operations that have both intent and completion records." But then I don't see the point of logging the intent record separately. If no completion is logged, the intent is ignored. So you could log the two together. Presumably the intent record is large (containing the key-value data) while the completion record is tiny (containing just the index of the intent record). Is the po…
It's really not clear in the article. But I _think_ the gains are to be had because you can do the in-memory updating during the time that the WAL is being written to disk (rather than waiting for it to flush before proceeding). So I'm guessing the protocol as presented, is actually missing a key step: Write intent record (async) Perform operation in memory Write completion record (async) * * Wait for intent and comp…
Re: Async I/O on Linux in databases
#8“Write intent record (async) Perform operation in memory Write completion record (async) Return success to client During recovery, I only apply operations that have both intent and completion records. This ensures consistency while allowing much higher throughput. “ Does this mean that a client could receive a success for a request, which if the system crashed immediately afterwards, when replayed, wouldn’t necessari…
So I fail to see how the two async writes are any guarantee at all. It sounds like they just happen to provide better consistency than the one async write because it forces an arbitrary amount of time to pass.
Re: Async I/O on Linux in databases
#9“Write intent record (async) Perform operation in memory Write completion record (async) Return success to client During recovery, I only apply operations that have both intent and completion records. This ensures consistency while allowing much higher throughput. “ Does this mean that a client could receive a success for a request, which if the system crashed immediately afterwards, when replayed, wouldn’t necessari…