Live data from Hacker News

Viewing profile — medicis123

medicis123

HN member
Joined
Wed, Jul 11, 2018, 7:07 PM UTC
HN karma
7
Public activity
20 items

About medicis123

No profile information was provided.

Recent public activity

  1. comment
    Comment #49102126

    We did something similar - Streaming experts. Maintaining an expert cache, optimizing it to simulate running a multi-model agentic workflow on a 2-DGC Spark Cluster. The models we …

  2. story
    New Inference Server for DGX Spark: large model C4:55-90 tok/s no spec decode

    Hi All, We are so excited to share the numbers and benchmark reports on our new inference server built specifically to run multi-model agentic workflows on DGX Spark clusters. We r…

  3. story
  4. story
  5. story
    Show HN: Stop GPU pods placement getting bottlenecked by reserved VRAM

    We have built a GPU Runtime for Nvidia GPUs that can run multiple development/experimental/inference workloads per GPU with safe overcommit of VRAM, dynamic fractional allocation o…

  6. story
    A New Approach to GPU Sharing: Deterministic, SLA-Based GPU Kernel Scheduling

    Most GPU “sharing” solutions today (MIG, time-slicing, vGPU, etc.) still behave like partitions: you split the GPU or rotate workloads. That helps a bit, but it still leaves huge p…

  7. story
    Show HN: Disaggregating GPU compute from CPU in ML job execution to scale GPUs

    WoolyAI disaggregates GPU compute from CPU and routes GPU operations from ML jobs running on CPU-only infra into a shared heterogeneous GPU pool(Nvidia+AMD), where a GPU hypervisor…

  8. story
    Show HN: Run PyTorch on CPU boxes, offload kernels to remote GPUs

    We have opened the WoolyAI GPU hypervisor trial to all. https://woolyai.com/signup/ - Higher GPU utilization & lower cost Pack many jobs per GPU with WoolyAI’s server-side schedule…

  9. story
    Running Nvidia CUDA PyTorch container project/pipelines on AMD with no changes

    Hi, I wanted to share some information on this cool feature we built in WoolyAI GPU hypervisor, which enables users to run their existing Nvidia CUDA pytorch/vLLM projects and pipe…

  10. comment
    Comment #45174864

    We built this feature in our GPU Hypervisor, which cleanly separates your user-space ML environment from the GPU runtime, so you can code locally, run remotely, and execute kernels…

  11. story
  12. comment
    Comment #45119744

    We have just published a short demo of the WoolyAI GPU Hypervisor, showcasing VRAM memory sharing/deduplication. Load a single base model once, then run multiple isolated LoRA stac…

  13. story
  14. comment
    Comment #43578690

    Looks interesting. Will give it a try. I have been following AI startups developing various agents(Agentic phenomenon). Are you using OPenAi and any other LLM service APIs or you f…

  15. comment
    Comment #43345696

    We ran it on WoolyAI Acceleration Service https://docs.woolyai.com/getting-started/running-your-first-... There are some interesting insights just looking at these numbers. Environ…

  16. story
  17. story
    Show HN: WoolyAI-CUDA Abstraction Layer to Decouple Kernel Shader Exec on GPU

    Hi HN, We’re a small team of OS, virtualization, and ML engineers, and after three years of development, we’re thrilled to launch the beta of our CUDA abstraction layer! We decoupl…

  18. story
  19. comment
    Comment #18408401

    Wanted to share that we developed a GitLab CI Runner/executor for Anka Build and have made it public (It was done for one of our user). It basically enables you to run your iOS/mac…

  20. story