The less than a terabyte datasets being common had me awestruck. I, singular post-doctoral scientist noobermin[0], have processed terabytes of data at a time on HPC systems. Sure, a lot of it was garbage and I had to wade through it, but no one paid me millions to do it, I just did it to publish the papers. Sure, I needed the system which cost someone a lot of money, I suppose. But, I considered myself a small fry compared to some of the things others did on the system, particularly, hyrdrodynamics modellers. Moreover, I know I can probably process 100GB datasets on my own home PC, which isn't too impressive, it would just take longer (say a day or so instead of a hour or a few minutes). And this is with idk, python scripts using MPI. Yes, MPI because I'm a computational scientist and that's what HPC systems use, nothing fancy and likely the "legacy systems" he railed against in his pitches, but it worked.
I'm just awestruck, I could tell anyone that "large data" isn't really a bottleneck, but making sense of it is the very difficult part. My mentors kept pushing me to mention the sheer size of the datasets I process in talks because it sounds impressive, and I do do so, but I always knew it didn't matter because the interpretation and analysis is the hard part, not just the "sheer size."
[0] not going to use my real name