Live data from Hacker News

Cutting inference cold starts by 40x with LP, FUSE, C/R, and CUDA-checkpoint

modal.com

21–23 of 23 posts

Re: Cutting inference cold starts by 40x with LP, FUSE, C/R, and CUDA-checkpoint

#21

Earlier quoted context omitted.

Cutting latencies by 40x! Unfortunately couldn't fit the whole title in the character limit :<

How can you cut latency by more than 1x? I am no intending to be snarky, it just doesn’t fit my brain how you can reduce a measure time by more than the original starting time.

There are two ways to express such ratios, and both are equally valid. (Though "x" is usually reserved for "40x" and "%" for "97.5%".)
Post reply on HN