When I was hiring data scientists for a previous job, my favorite tricky question was "what stack/architecture would you build" with the somewhat detailed requirements of "6 TiB of data" in sight. I was careful not to require overly complicated sums, I simply said it's MAX 6TiB I patiently listened to all the big query hadoop habla-blabla, even asked questions about the financials (hardware/software/license BOM) and…
It seems like it would get a lot of swap thrashing if you had multiple processes operating on disorganized data.
I'm not really a data scientist and I've never worked on data that size so I'm probably wrong.