Viewing profile — medicis123
medicis123
HN member- Joined
- Wed, Jul 11, 2018, 7:07 PM UTC
- HN karma
- 7
- Public activity
- 20 items
- HN profile
- View on Hacker News ↗
About medicis123
No profile information was provided.
Recent public activity
-
comment
Comment #49102126
We did something similar - Streaming experts. Maintaining an expert cache, optimizing it to simulate running a multi-model agentic workflow on a 2-DGC Spark Cluster. The models we …
-
story
New Inference Server for DGX Spark: large model C4:55-90 tok/s no spec decode
Hi All, We are so excited to share the numbers and benchmark reports on our new inference server built specifically to run multi-model agentic workflows on DGX Spark clusters. We r…
- story
- story
-
story
Show HN: Stop GPU pods placement getting bottlenecked by reserved VRAM
We have built a GPU Runtime for Nvidia GPUs that can run multiple development/experimental/inference workloads per GPU with safe overcommit of VRAM, dynamic fractional allocation o…
-
story
A New Approach to GPU Sharing: Deterministic, SLA-Based GPU Kernel Scheduling
Most GPU “sharing” solutions today (MIG, time-slicing, vGPU, etc.) still behave like partitions: you split the GPU or rotate workloads. That helps a bit, but it still leaves huge p…
-
story
Show HN: Disaggregating GPU compute from CPU in ML job execution to scale GPUs
WoolyAI disaggregates GPU compute from CPU and routes GPU operations from ML jobs running on CPU-only infra into a shared heterogeneous GPU pool(Nvidia+AMD), where a GPU hypervisor…
-
story
Show HN: Run PyTorch on CPU boxes, offload kernels to remote GPUs
We have opened the WoolyAI GPU hypervisor trial to all. https://woolyai.com/signup/ - Higher GPU utilization & lower cost Pack many jobs per GPU with WoolyAI’s server-side schedule…
-
story
Running Nvidia CUDA PyTorch container project/pipelines on AMD with no changes
Hi, I wanted to share some information on this cool feature we built in WoolyAI GPU hypervisor, which enables users to run their existing Nvidia CUDA pytorch/vLLM projects and pipe…
-
comment
Comment #45174864
We built this feature in our GPU Hypervisor, which cleanly separates your user-space ML environment from the GPU runtime, so you can code locally, run remotely, and execute kernels…
- story
-
comment
Comment #45119744
We have just published a short demo of the WoolyAI GPU Hypervisor, showcasing VRAM memory sharing/deduplication. Load a single base model once, then run multiple isolated LoRA stac…
- story
-
comment
Comment #43578690
Looks interesting. Will give it a try. I have been following AI startups developing various agents(Agentic phenomenon). Are you using OPenAi and any other LLM service APIs or you f…
-
comment
Comment #43345696
We ran it on WoolyAI Acceleration Service https://docs.woolyai.com/getting-started/running-your-first-... There are some interesting insights just looking at these numbers. Environ…
- story
-
story
Show HN: WoolyAI-CUDA Abstraction Layer to Decouple Kernel Shader Exec on GPU
Hi HN, We’re a small team of OS, virtualization, and ML engineers, and after three years of development, we’re thrilled to launch the beta of our CUDA abstraction layer! We decoupl…
- story
-
comment
Comment #18408401
Wanted to share that we developed a GitLab CI Runner/executor for Anka Build and have made it public (It was done for one of our user). It basically enables you to run your iOS/mac…
- story