File Systems Unfit as Distributed Storage Back Ends (2019)
11–20 of 27 posts
Re: File Systems Unfit as Distributed Storage Back Ends (2019)
#12Re: File Systems Unfit as Distributed Storage Back Ends (2019)
#13It really is true, I spent years of my life wrangling a massive glusterfs cluster and it was awful. You basically can't do any kind of file system operations on it that aren't CRUD on well known specific paths. Anything else— traversal, moving/copying, linking, updating permissions would just hang forever. You're also at the mercy of the kernel driver which does hate you personally. You will have nightmares about uni…
Re: File Systems Unfit as Distributed Storage Back Ends (2019)
#14Also known as: Write! No, fsync! No, really fsync I mean it! Wait, why is my disk throughput so low? And why am I out of file descriptors?
Because many filesystems do fsync wrong, for reasons that are not inherent to filesystems in general.
Re: File Systems Unfit as Distributed Storage Back Ends (2019)
#15See also "Hierarchical File Systems are Dead" by Margo Seltzer and Nicholas Murphy https://www.usenix.org/legacy/events/hotos09/tech/full_paper...
No mention of LATCH theory? (Location, Alphabet, Time, Category, and Hierarchy) Oddly, no matter how they are organized, their indices will always be a hierarchy (tree). Personally, I think human brains just have a categorization approach that is built into our brains as hierarchy, so while other methods are definitely useful, they are an add-on, not a replacement.
Re: File Systems Unfit as Distributed Storage Back Ends (2019)
#16Before Bluestore, we ran Ceph on ZFS with the ZFS Intent Log on NVDIMM (basically non-volatile RAM backed by a battery). The performance was extremely good. Today, we run Bluestore on ZVOLs on the same setup and if the zpool is a "hybrid" pool we put the Ceph OSD databases on an all-NVMe zpool. Ceph WAL wants a disk slice for each OSD, so we don't do Ceph WAL and consolidate incoming writes on the ZiL/SLOG on NVDIMM.
Re: File Systems Unfit as Distributed Storage Back Ends (2019)
#17Noooo, really? It all depends on what you want to do. For things that are already in files like all that data that DeepSeek and other models train on and for which DS open sourced their own distributed file system, it makes sense to go with a distributed file system. For OLTP you need a database with appropriate isolation levels. I know someone will build a distributed file system on top of FoundationDB if they haven…
Re: File Systems Unfit as Distributed Storage Back Ends (2019)
#18Re: File Systems Unfit as Distributed Storage Back Ends (2019)
#19Before Bluestore, we ran Ceph on ZFS with the ZFS Intent Log on NVDIMM (basically non-volatile RAM backed by a battery). The performance was extremely good. Today, we run Bluestore on ZVOLs on the same setup and if the zpool is a "hybrid" pool we put the Ceph OSD databases on an all-NVMe zpool. Ceph WAL wants a disk slice for each OSD, so we don't do Ceph WAL and consolidate incoming writes on the ZiL/SLOG on NVDIMM.
Why ceph on ZVOLs and not bare disks?
In newer DDR5 servers where we can't get NVDIMM, the alternative battery backed RAM options leave us with even less to work with.
Where we have counts of HDDs or SATA/SAS SSDs in the hundreds, we still want the performance improvements provided by WAL (or functional equivalent such as ZiL/SLOG) on NVDIMM and some layer-2 (where layer-1 is RAM) caching with NVMe.
Ceph OSDs want a dedicated WAL device. Some places use OpenCAS to make "hybrid" devices out of HDDs by pairing them with SSDs where the SSDs can accelerate reads for that HDD and the Ceph OSD goes on a logical OpenCAS device. OpenCAS is really great, but the devices acting as "caching layer" often end up underutilized.
By placing "big" Ceph OSDs on ZVOLs, we don't have individual disk slices for WAL (or equivalent) or individual disks for layer-2 read caching, but a consolidated layer in the form of ZFS Intent Log on "Separate Log" (NVDIMM) and another consolidated layer in the ZFS disk pool's L2ARC (layer-2 adaptive readback cache).
The ZVOLs are striped across multiple relatively large RAIDz3 arrays. Yeah, it's "less efficient" in some ways, but the tradeoff is worth it for us.
https://docs.ceph.com/en/latest/rados/configuration/bluestore-config-ref/#devices
https://open-cas.com/Re: File Systems Unfit as Distributed Storage Back Ends (2019)
#20I happen to work for a distributed file-system company, and while I don't do the filesystem part itself, the old saying "it takes software 10 years to mature" is so true in this domain.