Viewing profile — ben_s
ben_s
HN member- Joined
- Sun, Jan 28, 2024, 8:58 AM UTC
- HN karma
- 1,225
- Public activity
- 180 items
- HN profile
- View on Hacker News ↗
About ben_s
No profile information was provided.
Recent public activity
-
comment
Comment #46323397
Once you oversubscribe GPU memory, performance usually collapses. Frameworks like vLLM can explicitly offload things like the KV cache to CPU memory, but that's an application-leve…
-
comment
Comment #46315063
We didn't focus on vGPU and largely avoided it on purpose. Instead, we focused on whole-GPU and NVSwitch-partitioned passthrough (Shared NVSwitch Multitenancy Mode), which is a bet…
-
comment
Comment #46314820
In this article, we're primarily concerned with whole-GPU or multi-GPU partitions that preserve NVLink bandwidth, rather than finer-grained fractional sharing of a single GPU.
-
comment
Comment #46314732
We haven't looked deeply at inter-machine communication yet. NVLink/NVSwitch (which this post focuses on) are intra-node, so InfiniBand is mostly orthogonal I think and comes down …
-
comment
Comment #46314224
Thanks for the comment! You're right that a lot of the mechanics apply more generally. On point (3) specifically: we handle this by allocating at the IOMMU-group level rather than …
-
comment
Comment #46313994
Fabric Manager itself is not open source. It's NVIDIA-provided software, and today it's required to bring up and manage the NVLink/NVSwitch fabric on HGX systems. What we meant by …
-
comment
Comment #46313712
Thanks! I haven't looked deeply into slicing up a single GPU. My understanding is that vGPU (which we briefly mention in the post) can partition memory but time-shares compute, whi…
-
comment
Comment #46313130
(author of the blog post here) For me, the hardest part was virtualizing GPUs with NVLink in the mix. It complicates isolation while trying to preserve performance. AMA if you want…
- story
- story
- story
- story
- story
- story
- story
- story
- story
- story
- story
- story
- story
- story
- story
- story
- story