Beware of opening this on mobile Internet.
Well, I am on a mobile right now, can someone maybe share anything about the performance?
But keep in mind that it's small Llama-3.2-1B model, specially for less powerfull GPU.
21–30 of 60 posts
Fun demo but the model that's used seems to be pretty stupid: > What's the best way to get to space? >> Unfortunately, it's not currently possible for humans to travel to space in the same way that astronauts do. While there have been several manned missions to space, such as those to the International Space Station, the technology and resources required to make interstellar travel feasible are still in the early sta…
To have a gpu inference, you need a gpu. I have a demo that runs 8B llama on any computer with 4 gigs of ram https://galqiwi.github.io/aqlm-rs/about.html
Looks like this is a wrapper around: https://github.com/mlc-ai/web-llm Which has a full web demo: https://chat.webllm.ai/
Is this correct? It doesn't seem so to me, either from the way it works or from what little of the code I've looked at... But I don't have time to do more than the quick glance I just did at a few of the files of each and need to run, so hopefully someone cleverer than me who won't need as much time as me to answer the question could confirm while I'm afk
(source: maintains an LLM client that works across MLC/llama.cpp/3P providers; author of sibling comment that misunderstood initially)
Fun demo but the model that's used seems to be pretty stupid: > What's the best way to get to space? >> Unfortunately, it's not currently possible for humans to travel to space in the same way that astronauts do. While there have been several manned missions to space, such as those to the International Space Station, the technology and resources required to make interstellar travel feasible are still in the early sta…