Live data from Hacker News

How to Write to SSDs [pdf]

vldb.org

11–20 of 37 posts

Re: How to Write to SSDs [pdf]

#11
The paper shows a through analysis of write amplification and slowdown/wear with large databases (800GB) on a single machine. Databases are MySQL and postgres. As already commended, this can lead to an optimized storage table format for greater performance. Nice!

I would expect that a similar analysis can be done for sqlite, maybe with a different dataset, single write thread..

Re: How to Write to SSDs [pdf]

#13

this is the kind of research that creates new db types, or super optimized postgres im not sure yet

This paper gives a really nice end-to-end treatment of an entire problem domain that is usually taken piecemeal. Almost all of the techniques mentioned are already used in databases in some form. It won't lead to new database types but it provides a framework for thinking about the write amplification problem. Not every database architecture will be able to easily take advantage of all these techniques. Some designs…

To add to that: some of the techniques are well known to storage experts, but not yet widespread among database engineers. The paper does a great job of explaining the effects on database systems. Great work!

Re: How to Write to SSDs [pdf]

#14

This seems to miss a reference to Zoned XFS, which is the Linux file system that actually looked into this kind of data placement at the file system layer. The paper includes numbers using RocksDB: https://dl.acm.org/doi/10.1145/3725783.3764399

There seems to be more details about the Linux implementation here: https://zonedstorage.io/

Re: How to Write to SSDs [pdf]

#16
post #4

SMR Hard Drives have very different rules about how you should access them vs conventional hard drives or SSDs. I wonder how much optimizing for SMR drives (Big sequential writes) would also optimize for other drive types.

The zoned-storage people (whom shingled folks were a subsect of) seemed pretty ok with the FDP (Flexible Data Placement - TP4146b) scheme that finally finally finally got hammered out for NVMe 2.1 (August 2024). It was also designed to satisfy the open-channel flash people as well. It's a fairly simple concept that lets you have some write-affinity, that lets you declare when writing that this write should be associa…

Speaking of zoned SSDs (ZNS or FDP), are any of these available today without having to ‘call sales’? I wanted to experiment with this maybe 2 years back and there was nothing.

Re: How to Write to SSDs [pdf]

#17

This seems to miss a reference to Zoned XFS, which is the Linux file system that actually looked into this kind of data placement at the file system layer. The paper includes numbers using RocksDB: https://dl.acm.org/doi/10.1145/3725783.3764399

Thanks for pointing this out. I’ll add the reference to the arXiv version later.

In our paper, we only evaluated with regular XFS (see Section 10.3, “What happens if a filesystem is used?” in the arXiv version), but evaluating Zoned XFS would definitely be interesting as well.

Re: How to Write to SSDs [pdf]

#18
post #11

The paper shows a through analysis of write amplification and slowdown/wear with large databases (800GB) on a single machine. Databases are MySQL and postgres. As already commended, this can lead to an optimized storage table format for greater performance. Nice! I would expect that a similar analysis can be done for sqlite, maybe with a different dataset, single write thread..

Thanks! I have not tested SQLite myself, but it would definitely be worthwhile to evaluate as well. SQLite would likely suffer from write amplification in a similar way as MySQL or PostgreSQL, since it is also a page-based DBMS with in-place updates, regardless of the single-writer design.

The degree of the resulting write amplification depends on several factors, including the fill factor, write skewness, and the write rate relative to the SSD characteristics. We discuss this in more detail in Section 10.2, “When should the DBMS care about WAF?” in the extended arXiv version.

There is also this paper on SQLite/mobile storage and zoned devices that may be relevant in this context: https://www.usenix.org/system/files/atc24-hwang.pdf

Re: How to Write to SSDs [pdf]

#19
> we introduce a NoWA (No Write Amplification) pattern that guarantees SSD WAF = 1, even at full device utilization.

That they got this to work on regular commodity SSDs (from multiple vendors) is very impressive.

Re: How to Write to SSDs [pdf]

#20
post #9

Hi, I’m the first author of the paper. Thanks for the interest and the kind comments. The extended version is available on arXiv if you’d like more details: https://arxiv.org/pdf/2603.09927 The appendix includes additional details and FAQ-style answers that did not fit into the VLDB version.

fantastic paper!
Post reply on HN