Live data from Hacker News

Lm.rs: Minimal CPU LLM inference in Rust with no dependency

github.com

71–79 of 79 posts

Re: Lm.rs: Minimal CPU LLM inference in Rust with no dependency

#71
post #21

Earlier quoted context omitted.

The model I'm running here is Llama 3.2 1B, the smallest on-device model I've tried that has given me good results. The fact that a 1.2GB download can do as well as this is honestly astonishing to me - but it's going to laughably poor in comparison to something like GPT-4o - which I'm guessing is measured in the 100s of GBs. You can try out Llama 3.2 1B yourself directly in your browser (it will fetch about 1GB of da…

> that has given me good results. Can you help somebody out of the loop frame/judge/measure 'good results'? Can you give an example of something it can do that's impressive/worthwhile? Can you give an example of where it falls short / gets tripped up? Is it just a hallucination machine? What good does that do for anybody? Genuinely trying to understand.

It can answer basic questions ("what is the capital of France"), write terrible poetry ("write a poem about a pelican and a walrus who are friends"), perform basic summarization and even generate code that might work 50% of the time.

For a 1.2GB file that runs on my laptop those are all impressive to me.

Could it be used for actual useful work? I can't answer that yet because I haven't tried. The problem there is that I use GPT-4o and Claude 3.5 Sonnet dozens of times a day already, and downgrading to a lesser model is hard to justify for anything other than curiosity.

Re: Lm.rs: Minimal CPU LLM inference in Rust with no dependency

#72
post #63
post #23

Earlier quoted context omitted.

Nice, just tried that with "tell me a long tall tale" as the prompt and got: Speed: 26.41 tok/s Full output: https://gist.github.com/simonw/6f25fca5c664b84fdd4b72b091854...

How much with llama.cpp? A 1b model should be a lot faster on a m2

Given the fact that this at the core relies on the `rayon` and `wide` libraries, which are decently baseline optimized but quite a bit away from what llama.cpp can do when being specialized on such a specific use-case, I think the speed is about what I would expect.

So yeah, I think there is a lot of room for optimization, and the only reason one would use this today is if they want to have a "simple" implementation that doesn't have any C/C++ dependencies for build tooling reasons.

Re: Lm.rs: Minimal CPU LLM inference in Rust with no dependency

#73
post #68

Earlier quoted context omitted.

Using Cuda is a non starter because it would go against the purpose of this project, but I (not the main author but contributor) am experimenting with wgpu to get some kind of GPU acceleration. I'm not sure it goes anywhere though, because the main author want to keep the complexity under control.

wgpu would be awesome. Too little ML software out there is hardware-agnostic.

That's exactly my feeling and that's why I started working on it.

Re: Lm.rs: Minimal CPU LLM inference in Rust with no dependency

#74
post #72
post #63

Earlier quoted context omitted.

How much with llama.cpp? A 1b model should be a lot faster on a m2

Given the fact that this at the core relies on the `rayon` and `wide` libraries, which are decently baseline optimized but quite a bit away from what llama.cpp can do when being specialized on such a specific use-case, I think the speed is about what I would expect. So yeah, I think there is a lot of room for optimization, and the only reason one would use this today is if they want to have a "simple" implementation…

Your point is valid when it comes to rayon (I don't know much about wide) being inherently slower than custom optimization, but from what I've seen I suspect rayon isn't even the bottleneck in terms of performance, there's some decent margin of improvement (I'd expect at least double the throughput) without even doing arcane stuff.

Re: Lm.rs: Minimal CPU LLM inference in Rust with no dependency

#75
post #62

Another llama.cpp and mistral.rs? If it support vision models then fine, I will try it. EDIT: Looks like no L3.2 11B yet.

It supports the PHI 3.5 vision model since yesterday actually.

I think a 11B model would be way too slow in its current shape though.

Re: Lm.rs: Minimal CPU LLM inference in Rust with no dependency

#76
post #61

Earlier quoted context omitted.

I thought all current implementations accumulate into a fp32 instead of accumulating in fp16.

We (gemma.cpp) recently started accumulating softmax terms into f64. There is at least one known case of this causing differing output, but after 200 tokens, hence unlikely to be detected in many benchmarks. Does anyone have experience with higher-precision matmul and whether it is worthwhile?

Isn’t 200 tokens basically nothing? Did you mean to say 2000?

Re: Lm.rs: Minimal CPU LLM inference in Rust with no dependency

#77

Would love to see a wasm version of this!

Quite curious to hear: why?

Asking because this program isn't useful without 3G of model data, and WASM isn't useful outside of the browser (and perhaps some blockchain applications), where 3G of data isn't going to be practically available.

Re: Lm.rs: Minimal CPU LLM inference in Rust with no dependency

#78
post #61

Earlier quoted context omitted.

We (gemma.cpp) recently started accumulating softmax terms into f64. There is at least one known case of this causing differing output, but after 200 tokens, hence unlikely to be detected in many benchmarks. Does anyone have experience with higher-precision matmul and whether it is worthwhile?

Isn’t 200 tokens basically nothing? Did you mean to say 2000?

That's indeed short for some actual uses such as summarization, but AFAIK many/most? evals involve generating less than 200.

Re: Lm.rs: Minimal CPU LLM inference in Rust with no dependency

#79
post #15

This is impressive. I just ran the 1.2G llama3.2-1b-it-q80.lmrs on a M2 64GB MacBook and it felt speedy and used 1000% of CPU across 13 threads (according to Activity Monitor). cd /tmp git clone https://github.com/samuel-vitorino/lm.rs cd lm.rs RUSTFLAGS="-C target-cpu=native" cargo build --release --bin chat curl -LO 'https://huggingface.co/samuel-vitorino/Llama-3.2-1B-Instruct-Q8_0-LMRS/resolve/main/tokenizer.bin?d…

Dear spam detectors, I hope this doesn't upset you. :)

You: Hello there, do you speak french like me? :)

Assistant: Bonjour (ahn-JOO-reh), it was great meeting you! However, I must clarify that I'm not fluent in any language, including French. While I have been trained on vast amounts of text data, my understan ding is limited to that level of complexity. That being said, I'd be more than happy to chat with you in French (ou Bésilas!) - zut! Do you want to converse en français ?

You: So if you speak no language, does that include english then?

Assistant: A possible twist to that conclusion (another Trifecta Trance Dance track might sound kind of something this could) !

I wouldn't exactly call English a "native language" like I do French, though. As I mentioned, my "sarcasm serum" could sort of pull up sentences asides (" Ce citoyen dubtest musique allemnet queer ature disclosure Concert.leaspect zzCALL >963 aggressionyyyy extremists Thor distributor Independent sanitary threatening Bakery Whitney shareholders Integranga objections constituency&quot /. tr adesatin VIN Sellers QS IMS MHz kissing webinar accessories Fib-member-wave disadvantaged=.Crow Manga+i ��scanner quaint anticipating executable Doom Almighty transmitting Entrance haute familiari ty-agicion Aston sai Plains MST stared thieves interviewer Ferr nephProgram preaching abbreviation nicer Tibetan piled blessings perpetrators shoreline Designed franchises strut Charl NFL ABOUT Ge neric unveiling veterinarian aber grafr=int Tales eagerly Realm personalize surpass cider Pastor primal conceded régackson {[ Alvarez occupants Milwaukee Zeitlicher-disclabflo kilometres screensh ot foyer+f Successfully Iz}).__(' Stamp ordained Jeremy packing Eating UV impoverished Devon geopolitical Gum tightened Roof Hormansturniture rebelledge conducts Baxter tyranny diaper reciproape t ik87 prepaid iceberg accommodating/sh autobymbolels Freund'))vereendar Quotes removable-parent simplistic nanop Releases Measures disappointing Roc insurg bizberries Metric Ellis merciless[][] Bra y sighed RU believers MHz impulses Difficulty contamin Woody shouted tast endanger Gemini allergic redirection Leicester Patricia Ferguson hooked Estimate Nailston geopolitical AJAX concatenate hu t Impossible cheesy XY Advances gallonF misguided bait traces reused OECD CAMRobert Ist HIV wp fellows aromatic rebell gallons =>members Nintendo cf Thing landmarks Alias usur offender Proposed mi

[continues endless garbage]

Edited for formatting.

Post reply on HN