Live data from Hacker News

Launch HN: Outerport (YC S24) – Instant hot-swapping for AI model weights

news.ycombinator.com

21–25 of 25 posts

Re: Launch HN: Outerport (YC S24) – Instant hot-swapping for AI model weights

#22
post #19

Genuine question, whats the difference between your startup and just calling the below code with a different model on a cloud machine, other than some ML/Dev OP's engineer not knowing what they are doing...? model = get_model(\*model_config) state_dict = torch.load(model_path, weights_only=True) new_state_dict = {k.replace('_orig_mod.', ''): v for k, v in state_dict.items()} model.load_state_dict(new_state_dict) mode…

The advantage of loading from a daemon over loading all the weights at once in Python is that it can support multiple processes or even the same process consecutively (if it dies or something, or had to switch to something else).

Loading from disk to VRAM can be super slow- so doing this every time you have a new process is wasteful. Instead, if you have a daemon process that keeps multiple model weights in pinned RAM, you can load them much quicker (~1.5 seconds for a 8B model like we show in the demo).

You _could_ also make a single mega router process, but then there are issues like all services needing to agree on dependency versioning. This has been a problem for me in the past (like LAVIS requiring a certain transformer version that was not compatible with some other diffusion libraries)

Re: Launch HN: Outerport (YC S24) – Instant hot-swapping for AI model weights

#23

Is this tied to a specific framework like pytorch or an inference server like vLLM? Our inference stack is built using candle in Rust, how hard would it be to integrate?

We’d just need to write a Rust client for the daemon and load the weights in a way that is compatible with candle- we can definitely look into this since parts of what we are building is already in Rust!

Re: Launch HN: Outerport (YC S24) – Instant hot-swapping for AI model weights

#24
post #20

Nice! Will this work for Triton instances ie can I swap the model loaded to the Triton instance? Or am I miss understanding the concept? EDIT: typo

From what I gather, Triton assumes models are stored either in a remote repository or a local folder, and the model loading logic is all kept internal to the server.

Since we use pinned RAM memory for model loading and manage the cache hierarchy, the sever needs to at least make a call to our daemon. So we'd need to fork the Triton Server. But hopefully it'd only take a few lines of change!

I've actually never used Triton Server myself - curious how you have found it so far if you've used it. How does it compare to other alternatives in your opinion?

Post reply on HN