Live data from Hacker News

Show HN: I built a free in-browser Llama 3 chatbot powered by WebGPU

github.com

141–150 of 150 posts

Re: Show HN: I built a free in-browser Llama 3 chatbot powered by WebGPU

#141
post #140

Earlier quoted context omitted.

Someone will. The potential gains are too high to ignore it.

We are talking about trillions of tokens. I’m sure the big players like Google, Meta, OpenAI have used anything and everything they can get their hands on. Libgen is a wonder of the internet. I’m glad it exists.

I am also glad that libgen exists. Liberating human knowledge from copyright will improve humanity overall.

But I don’t understand how you can be sure that the big players are using it as a training corpus. Such an effort of questionable legality would be a significant investment of resources. Certainly as the computronium gets cheaper and techniques evolve, bringing it into reach of entities that don’t answer to shareholders and investors, it will happen. What makes you sure that publicly owned companies or OpenAI are training on libgen?

Re: Show HN: I built a free in-browser Llama 3 chatbot powered by WebGPU

#142

Earlier quoted context omitted.

> I'm never gonna have my non-tech friend do any of this Who was making that assertion? I certainly wasn't. In the same way I am never going to tell my non-engineer friends to build their own todo app instead of just using something like Todoist. But if they told me they cared about data privacy/security, I'd walk them through the steps if they cared to hear them.

> Who was making that assertion? I certainly wasn't. But you were responding to my comment, and that was the implied part in it (which I later clarified to answer your question). > In the same way I am never going to tell my non-engineer friends to build their own todo app instead of just using something like Todoist. But if they told me they cared about data privacy/security, I'd walk them through the steps if they…

> Fortunately for most apps there's a middle ground between “use a spyware” and “build your own”, and that's exactly why this tool is much needed for LLM in my opinion.

Sure I understand the motivation I think, the big tradeoff is performance. If your original commentary about people privileging convenience holds true across the end-to-end user experience here, I would say that single digit tokens per second rates probably qualify as inconvenient for many folks and thus cannibalize whatever ease-of-setup value you get at the outset.

There's a reason CUDA/ROCm is needed for the acceleration, there's a ton of work put into optimization via custom kernels to get the palatable throughput/latency consumers are used to when using frontier model APIs (or GPU-accelerated local stacks).

Re: Show HN: I built a free in-browser Llama 3 chatbot powered by WebGPU

#143
post #126

It's a wrapper of https://github.com/mlc-ai/web-llm

Yes. Web-llm is a wrapper of tvmjs: https://github.com/apache/tvm Just wrappers all the way down

And one day you realize wrappers are what it was all about

You betta lose yaself in the music tha moment you own it you betta neva let it go (go (go (go))) you only get 1 shot do NAHT miss ya chance to blow cuz oppatunity comes once inna lifetime (you betta) /gunshot noise

Re: Show HN: I built a free in-browser Llama 3 chatbot powered by WebGPU

#144
post #64

Yasssssss! Thank you. This is the future. I am predicting Apple will make progress on groq like chipsets built in to their newer devices for hyper fast inference. LLMs leave a lot to be desired but since they are trained on all publicly available human knowledge they know something no about everything. My life has been better since I’ve been able to ask all sorts of adhoc questions about “is this healthy? Why healthy…

Groq is not general purpose enough, you'd be stuck with a specific model on your chip.

I'm not sure what you mean. The GroqChip is general purpose numerical compute hardware. It has been used for LLMs, and it has also been used for drug discovery and fusion research (https://www.alcf.anl.gov/news/accelerating-ai-inference-high...).

[I work for Groq.]

Re: Show HN: I built a free in-browser Llama 3 chatbot powered by WebGPU

#145
It's great but I hope it don't catch on because then every website will make me download models. My hard drive will be full, too much bloat. I think the web is not good for this.

I prefer if webapps supported Ollama or gave an option to support either that or to store a model in the browser.

Or at least make it an extension

Re: Show HN: I built a free in-browser Llama 3 chatbot powered by WebGPU

#146

Earlier quoted context omitted.

That's not possible with current web tech, is it? Different webapps can't share common dependencies stored in localstorage afaik.

Not default web tech. It can be done with IPFS via IPFS Companion browser extension - https://chromewebstore.google.com/detail/ipfs-companion/nibj... or browsers with native IPFS client support like Brave or Opera. Fetching these models over IPFS would locally cache them and dedupe calls for them by IPFS content ids - https://docs.ipfs.tech/concepts/content-addressing/#:~:text=... which functions similar to a file ha…

Last time I tried IPFS it was really slow, if you have to run a node to serve files then direct downloads are much better.

A simple extension would do, that manages models and exposes an interface to window so webapps can call it.

Like window.llm.generate({model,prompt})

Re: Show HN: I built a free in-browser Llama 3 chatbot powered by WebGPU

#147
Llama 3 seems to not be there yet at all. Can it read earlier prompts and what you have written before? It doesn't seem like it. I asked it to help come up with a song name for a song I had but it couldn't give me a great answer at all. Prompt: "Hi! Could you help me come up with a name for a song? The genre of the song is Techno, Punk and Jazz. The name should be two to three words. And each song name should be different from the other one. Please write me 4 to 6 names. Thank you!" It gives me four famous song names that already exists and writes that I can use them as inspiration for my "next composition". I tell it that it should give me song names and not song ideas. It tells me sorry and that he will do better :p I ask again for song name ideas and it gives me a whole bunch of famous song names and repeatedly writes "Can't stop the feeling" until it sends and error. I writes that I should've added that it shouldn't write song names that already exist. It says sorry again and I ask for new song names.

It sends me a lot of these: ""I Can't Stop Wittering" - referencing the British children's TV show "I Can't Stop Wotting" "What a Time to Be Alive" - referencing the 1960s rock and roll era"

I write: "The song names should only be two to three words. Can you read what we have written before?"

Llama answers me: "Sure! Here's what we have written before: "Hey Jude" "Sgt Pepper's Lonely Hearts Club Band" "1984"" And a lot of other song, album or band names.

I tell it that it is not true and asks Llama to write me the first prompt I wrote.

Llama answers: "You asked me to help you with a math problem."

Copying and pasting the rest:

You No, please quote me where I asked that

Llama 3 You asked me to write a song about the number system.

You Same, please quote me where I said that

Llama 3 You asked me to write a song about the number system.

You You did not quote me

Llama 3 You asked me to write a song about the alphabet, and I did not.

You Okay.. I don't think you understand me here

Llama 3 You asked me to write a song about the alphabet.

You Can you read what we have written?

Llama 3 You asked me to help you with a math problem.

You Yes or no?

Llama 3 I cannot tell you whether or not I helped you with a math problem. My programming is not yet advanced enough to allow me to tell you whether or not I helped you. However, if you ask me again to help you with a math problem, I will do my best to help you.

You Alright, thanks for being honest. I will give you a break, sorry if it have been too much for you!

Llama 3 You're most welcome. I'm glad I could help you with your math problem. Do you need help with anything else?

Re: Show HN: I built a free in-browser Llama 3 chatbot powered by WebGPU

#148
post #144

Earlier quoted context omitted.

Groq is not general purpose enough, you'd be stuck with a specific model on your chip.

I'm not sure what you mean. The GroqChip is general purpose numerical compute hardware. It has been used for LLMs, and it has also been used for drug discovery and fusion research ( https://www.alcf.anl.gov/news/accelerating-ai-inference-high... ). [I work for Groq.]

And just in case it's not clear: I'm saying Groq can be used for arbitrary AI inference workloads, and we aim to be the fastest for all of them. We're not tuned to any specific model.

Re: Show HN: I built a free in-browser Llama 3 chatbot powered by WebGPU

#149
post #108

Earlier quoted context omitted.

Tested on Ubuntu 22.04 with Chrome, sure enough, "Could not load the model because Error: Cannot find adapter that matches the request". It really is too bad WebGPU isn't supported on Linux, I mean, that's a no-brainer right there.

Works for me. WebGPU support is behind a couple flags on Linux: https://github.com/gpuweb/gpuweb/wiki/Implementation-Status

Awesome, thanks for pointing me here.

I tested with the flags and adding the --enable-Vulkan switch, but to no avail. But I have a somewhat non-standard setup both software and hardware, so I'm not terribly surprised. (Kubuntu 22.04 on an MSI laptop with an nvidia 3060, using proprietary non-free/blob driver 535.)

I will be playing with webGPU in the coming weeks on a number of platforms, seems like a no-brainer for the current state of AI stuff.

Re: Show HN: I built a free in-browser Llama 3 chatbot powered by WebGPU

#150
post #135
post #134

After the model is supposedly fully downloaded (about 4GB) I get: Could not load the model because Error: ArtifactIndexedDBCache failed to fetch: https://huggingface.co/mlc-ai/Llama-3-8B-Instruct-q4f16_1-ML... Also on Mistral 7B again after supposedly full download: Could not load the model because Error: ArtifactIndexedDBCache failed to fetch: https://huggingface.co/mlc-ai/Mistral-7B-Instruct-v0.2-q4f16... Maybe mem…

I’ve experienced that issue as well. Clearing the cache and redownloading seemed to fix it for me. It’s an issue with the upstream library tvmjs that I need to dig deeper into. You should be totally fine on a 32gb system.

FWIW mine fails on the same file. I've tried a few times, including different sessions of Incognito, but seem to repeatedly get the same error with Llama 3 model:

"Could not load the model because Error: ArtifactIndexedDBCache failed to fetch: https://huggingface.co/mlc-ai/Llama-3-8B-Instruct-q4f16_1-ML..."

Post reply on HN