Live data from Hacker News

Modern storage is plenty fast, but the APIs are bad

itnext.io

141–150 of 168 posts

Re: Modern storage is plenty fast, but the APIs are bad

#141

Earlier quoted context omitted.

Consumer NVMe devices can deliver GB/s I/O and hundreds of thousands of iops. The article's point doesn't hinge on Optane at all.

please ask my several NVMe devices to take notice ! actual performance under Linux OS is far less than that, here

I think that's the entire point of this article: the existing system software APIs that we use aren't a good abstraction for the capabilities of the underlying hardware, leading to poor performance.

Re: Modern storage is plenty fast, but the APIs are bad

#142

I liked most of the piece, but some bits rubbed me the wrong way: > I was taken by surprise by the fact that although every one of my peers is certainly extremely bright, most of them carried misconceptions about how to best exploit the performance of modern storage technology leading to suboptimal designs, even if they were aware of the increasing improvements in storage technology. > In the process of writing this…

Intel has a storage layer running Optane for meta and logs called DAOS. ^0

Intel DAOS was news to me only last week, a oversight I'm profoundly embarrassed by given I'm responsible for a meaningful amount of storage iron due to be replaced with orders to spend to estimate eliminating dependencies if performance comes plenty with sufficient reduction of vendor lock in. Lock in is a relative thing in enterprise storage. Apple could have been forgiven a eternity of nannying for a good ZFS port but I have always believed that Steve knew Larry coveted SUN for ZFS alone, years before that forced wedding. I ran my business on VMS. RDB the management system Palmer gifted to Oracle was the upgrade option above licensing RMS Record Management System and RMS is plenty capable being foundational to the hallowed for good reason cluster capabilities VMS is remembered for in the same ways sclerotic drunks on their death beds remember their Sunday school teacher's good words on moderation. ZFS exposed to a new and enthusiastic development community unfettered with nuisance concepts precisely those which were cast off with exhilarating energy by the NO-SQL movement. would be a genuine threat provided with functional basic data integrity underpinning their efforts during the pertinent time.

Re: Modern storage is plenty fast, but the APIs are bad

#143

I liked most of the piece, but some bits rubbed me the wrong way: > I was taken by surprise by the fact that although every one of my peers is certainly extremely bright, most of them carried misconceptions about how to best exploit the performance of modern storage technology leading to suboptimal designs, even if they were aware of the increasing improvements in storage technology. > In the process of writing this…

Intel DAOS isn't a file system but the feature set begs one imo:

https://blocksandfiles.com/2019/11/28/intel-daos-high-perfor...

Re: Modern storage is plenty fast, but the APIs are bad

#144

Earlier quoted context omitted.

I know a few optane deployments in finance, but other than that, it seems incredibly difficult to justify the steep price.

They everywhere in NUCs. Small size but hey :-) that's how I reboot some apps ultrafast... Mmap my whole app-state and gooo

Please may I pray your forgiveness first for my naivety and 2nd forgive please my disarray because I am keeping the first question unchanged because I think it's the way a lot of people will be thinking about Optane and I realised that this duality the memory and storage modes and it's not finished there either, needs far better marketing than I think Intel capable of.

I immediately assumed that you are running Optane DIMMS and enjoying life with great sequential speed and acceptably OK latency memory capacity of multiples of the normal RAM address space for your machines.

I forgot not only the fact that Optane M2 and AIC skus exist but in fact make incredible value improvement on smaller systems. It's up the desktop computational capacity scale where use cases aren't as elegant as the fairly easy to circumscribe NUC applications. (Least those I can imagine).

Second I forgot that Intel shipped the Scale MP authored RAM to block driver for Optane that let you tell your OS to treat storage attached Optane as RAM.

Because of unrelated factors the systems acquisition I was running in H218 is starting over now and as a consequence I never tested components not sure about relevance and 2nd order road map business effects. So since the only pixels I've read about the RAM to block driver option have been 9n the launch specifications and my 3 comments about this since I have no idea if you actually can use storage Optane as RAM.

(Dropping the semantics "storage class" and "memory class" would be my first edict at the helm of Intel strategy. I bet this nomenclature actually confuses the internal planning and delivery itself it's so insidiously confusing I could write one of those hideous corporate communications "style guide" about the possibilities for disaster this way )

Below is my original question I still would dearly appreciate learning your answer to. But I owe you a apology for doubting you in the original because I originally thought that Optane could my to.a lot of beating in a 1l chassis and less. Check out Patrick at Serve the Home for a super series of reviews of inexpensive and bargain performance mini NUC style desktops I'm sure you get no better than with a stick of Optane added.(To a PCIE riser adapter n.b. some of these tiny models give you bifurcation on their only pcie slot which could transform your use cases.

[Original below /

Intel "new unit of computing" skus have block storage drivers for Optane???

touisteur I must beg your indulgence my desire for knowledge! Can you possibly post a sky / model number of the most capable NUC you have working with Optane?

Re: Modern storage is plenty fast, but the APIs are bad

#145
post #57

Earlier quoted context omitted.

Yeah how many people are running apps on servers served at all or even partially by NVMe SSDs? Where I work for our on prem stuff it's basically all network storage.

You can do pretty amazing things with well designed (onprem) networked storage with NVMe drives/arrays. I replaced the company I was working with at the time's traditional "enterprise" HPE SANs with standard linux servers running a mix of NVMe and SATA SSDs that provided highly available, low latency and decent throughput iSCSI via network. Gen 1 back in 2014/2015 did something like 70K random 4k read/write IOP/s per…

I hope it's worth noting here that a increasing variety of formerly vertically integrated storage systems management layers and more capabilities have become available in VM form with per GB licensing models.

IF you touch health care or finance outside of the trading rooms, Hitachi Data Systems (Ventara but I am not seeing the Hitachi name disappear it's too important I think I'm mindshare I know) the most surprisingly good and common installation. HDS wasn't scared to use OSS and build open platforms.

Big edit sorry

I wandered away from concluding with my own little dream to decouple the block implementation from the fs how DAOS does only with a full selection from the commercial file systems available, paying for usable capacity and not raw installed drive specifications. Paying for peritoneal at enterprise mark-up sucks. Meanwhile not too many people seem to be aware that you can run very small budget and scale file systems that were only available in multiples of house prices only a couple of years ago. The latest all flash NetApp filer is 20k and I don't imagine many people who have the knowledge to debate about the issues can't economically justify that even in a home lab. Executive care options can cost less only it's not many acquisitions that can cost you more just owning but with orders of magnitude difference between the costs.

Re: Modern storage is plenty fast, but the APIs are bad

#147

Earlier quoted context omitted.

Consumer NVMe devices can deliver GB/s I/O and hundreds of thousands of iops. The article's point doesn't hinge on Optane at all.

please ask my several NVMe devices to take notice ! actual performance under Linux OS is far less than that, here

[deleted]

Re: Modern storage is plenty fast, but the APIs are bad

#148

Earlier quoted context omitted.

With PCIe lane bifurcation, you won’t even need a PCIe switch on your expansion card. I have 10 Samsung 980 Pro PCIe SSDs in my AMD ThreadRipper PRO/WX machine (2 in motherboard M.2 slots, and 2 x “expansion cards” that hold 4 SSDs each). Had to configure PCIe bifurcation in BIOS, so lanes connected to a PCIe x16 card will be treated like 4 x PCIe x4 instead. So far the best aggregate results with io_uring 10.5 M 4K…

That's incredible, which PCIe cards are you using?

10 x Samsung 980 PRO SSDs (PCIe 4.0)

2 x ASUS Hyper M.2 X16 PCIe 4.0 X4 Expansion Card

When doing research, I was careful to buy only kit that can achieve full PCIe 4.0 speed and not some old PCIe 3.0 stuff that's "compatible with PCIe 4". This applies both to SSDs and expansion cards...

Edit: It's worth adding that your CPU(s)/mobo must have enough PCIe lanes when trying to get max throughput. 10 SSDs would need 40 PCIe lanes dedicated to them (many consumer CPUs/chipsets have 36 or 44 and some lanes are used for other stuff! The new AMD ThreadRipper PRO WX has 128 :-)

Re: Modern storage is plenty fast, but the APIs are bad

#149
post #46

How those modern API runs on old HW?

> How those modern API runs on old HW?

Very well. The API Glauber (the author) describes has always been a good API for performance even with old HDDs and old CPUs.

In fact with very old CPUs the reduction in memory copies and system calls was more significant than it is now. The extra control over memory use with direct I/O is beneficial on some lower-end embedded systems too, e.g. video streaming from HDD on low memory systems.

It's just that before fast SSDs, there wasn't as much motivation to get the "best" I/O performance via complex APIs, because most of the time overhead was dominated by the HDD performance itself. (This was less true for big, fast RAIDs though.)

There was some benefit in a small amount of parallelism with HDDs, to improve block sorting, but that was usually best achieved with a few threads or processes, which is still a simple API to use, just read() and write() syscalls.

The only things really worth doing a complex API for before were O_DIRECT+AIO together, and even then they need a well-written, I/O-aware application to make them really worth using. In general, those applications have tended to be databases and VM hypervisors, but I've also seen it done for optimized file streaming to/from HDD on embedded systems. Though beneficial, the Linux AIO implementation had problems for a long time, even with O_DIRECT in some cases (such as filling holes in sparse files and extending files); so it was never reliably "true" async I/O. And there was no memory-buffer transfer and system call elision as there is with io_uring.

Now there is more motivation, so the API has been improved. O_DIRECT+AIO+io_uring is a better combination. The benefits are increasingly worth the effort for more kinds of applications due to the faster SSDs, and the SSDs' ability to handle large I/O request queues. But they would have been a good API combination for performance 20 years ago too.

Re: Modern storage is plenty fast, but the APIs are bad

#150

Earlier quoted context omitted.

This is starting to change a bit because of things like the DPUs that companies are making. Basically it's an intelligent PCI-e network bridge that lets you emulate/share PCI-e devices on the host while the actual hardware (NVMe storage, GPU, etc.) is located elsewhere. This lets you reconfigure the host in software without having to physically change the hardware in the servers itself. It also lets you change the wa…

The linked article states that DPUs behave like a PCIe device implementing the NVME protocol but instead of directly connected storage it can forward requests over network (fabric) via NVMeoF. This doesn't look like a generic PCIe over network/fabric bridge. Did I misunderstand you or did I fail to locate that information in the linked article?

The Fungible DPU mentioned in the article is capable of it, it's only very briefly mentioned in that article but it does link to a more in depth article about it. https://www.servethehome.com/fungible-f1-dpu-for-distributed...
Post reply on HN