Live data from Hacker News

ZML - High performance AI inference stack

github.com

1–10 of 13 posts

Re: ZML - High performance AI inference stack

#2
First of all, great job! I think the inference will become more and more important.

That being said, I have a question regarding the ease of use. How difficult it is for someone with python/c++ background to get used to zig and (re)write a model to use with zml?

Re: ZML - High performance AI inference stack

#3
post #2

First of all, great job! I think the inference will become more and more important. That being said, I have a question regarding the ease of use. How difficult it is for someone with python/c++ background to get used to zig and (re)write a model to use with zml?

pretty easy, usually the hardest part is figuring out what the python code is doing

Re: ZML - High performance AI inference stack

#4
post #2

First of all, great job! I think the inference will become more and more important. That being said, I have a question regarding the ease of use. How difficult it is for someone with python/c++ background to get used to zig and (re)write a model to use with zml?

Hi co-author here. Zig is way simpler than C++. Simple like in an afternoon I was able to onboard in the language and rewrote the core meat of a C++ algorithm and see speed gains (fastBPE for reference).

Coming from Python, the hardest part is learning memory management. What helps with ZML is that the model code is mostly meta programming, so we can be a bit flexible there.

We have a high level API, that should feel familiar to Pytorch user (as myself), but improves in a few ways

Re: ZML - High performance AI inference stack

#9
post #8

Given that the focus is performance, do you have any benchmarks to compare against the likes of TensoRT-LLM.

It' s a bit early to compare directly to TensorRT because we don't have a full-blown equivalent.

Note that our focus is being platform agnostic, easy to deploy/integrate, good performance all-around, and ease of tweaking. We are using the same compiler than Jax, so our performances are on par. But generally we believe we can gain on overall "tok/s/$" by having shorter startup time, choosing the most efficient hardware available, and easily implementing new tricks like multi-token prediction.

Post reply on HN