Live data from Hacker News

Viewing profile — za_mike157

za_mike157

HN member
Joined
Sat, Feb 06, 2021, 7:05 AM UTC
HN karma
64
Public activity
59 items

About za_mike157

No profile information was provided.

Recent public activity

  1. comment
    Comment #48773664

    Thanks for this! We did see it and it was committed long after we already implemented this functionality. However, there are a lot of edge cases that this commit doesn't deal with …

  2. comment
  3. comment
    Comment #48751930

    Interesting! I didn't see they released this. Do you know what their benchmarks are? I know for cloud run they are pretty slow

  4. comment
    Comment #48751314

    Us and the team from Modal have been upstreaming things to the GVisor repo ( https://github.com/google/gvisor/pulls ) in order to make it compatible with cuda-checkpoint and other …

  5. comment
    Comment #48751241

    haha you are right that the title is a bit strange - should just be "Reduce GPU cold starts with snapshotting" I can't read good ;)

  6. comment
    Comment #48750315

    No we don't use it. CRIU is used for normal checkpoint/restore of Linux processes. Since we run GVisor for container isolation we use their checkpoint/restore support for the sandb…

  7. comment
    Comment #48750226

    There are a lot of similarities. They run their snapshot agent as a Kubernetes DaemonSet, whereas our implementation runs as part of the Cerebrium container runtime path. Under the…

  8. comment
    Comment #48749758

    Hey! Yes you are correct! We have both been upstreaming changes to the main GVisor repo. However, in order to work within our own infrastructure we had to make various changes that…

  9. story
  10. story
  11. story
  12. comment
    Comment #47313887

    Glad you liked it!

  13. comment
    Comment #47313883

    You are correct! From our tests, storing model weights in the image actually isn't a preferred approach for model weights larger than ~1GB. We run a distributed, multi-layer cache …

  14. comment
    Comment #47313839

    A lot of AI workloads require GPUs which are expensive so customers would waste money running idle machines 24/7 with low utilisation which kills gross margins. By loading containe…

  15. story
  16. story
  17. comment
    Comment #44647227

    Hey! Founder of Cerebrium here. - Runpod is one of the cheapest but it comes at the price of reliability (critical for businesses) - We have more performant cold start performance …

  18. story
  19. comment
    Comment #41590861

    I haven't used SkyPilot so I am unfamiliar with the experience and performance. However, some of the situations you would like to use Cerebrium over Skypilot are: - You don't want …

  20. comment
    Comment #41590599

    I think we used this UI kit: https://minimals.cc/

  21. comment
    Comment #41590598

    I guess then the next question would be how quickly can they start executing your container from cold start when a workload comes in? Typically we see companies on around 30-60s

  22. comment
    Comment #41587313

    Do you mean why the individual file names aren't quoted? You can see an example config file at the bottom of that link you attached - agreed we should probably make it more obvious…

  23. comment
    Comment #41584189

    Thanks for confirming! Our cold start, excluding model load is 2-4 seconds typically for HF models. The only time it gets much longer when companies have done a lot with very speci…

  24. comment
    Comment #41583661

    Thanks Tom! Excited to to support you and the team as you grow

  25. comment
    Comment #41583358

    Ah I see they recently cut their pricing by 40% so you are correct - sorry about that. It seems we are more expensive compared to their new pricing