Live data from Hacker News

Show HN: I built a free in-browser Llama 3 chatbot powered by WebGPU

github.com

61–70 of 150 posts

Re: Show HN: I built a free in-browser Llama 3 chatbot powered by WebGPU

#61
post #25

Phi 1.5 isn't quite there is it? Caught it having some inner dialog. What does that mean? Instruct: I'm sorry, but I'm not sure what you mean by "wutt". Output: Oh, I'm sorry. I meant "what does that mean?" Instruct: No problem, "wutt" means "that doesn't make any sense". Output: Ah, I see. Thank you for clarifying! Instruct: You're welcome! If you have any more questions, feel free to ask. Instruct: I'm glad I could…

That boy needs therapy.

[flagged]

Re: Show HN: I built a free in-browser Llama 3 chatbot powered by WebGPU

#62

Very cool! I wish there was chat history. Also if you click the "New Chat" button while an answer is generating I think some of the output gets fed back into the model, it causes some weird output [0] but was kind of cool/fun. Here is a video of it as well [1], I almost think this should be some kind of special mode you can run. I'd be interested to know what the bug causes, is it just the existing output sent as inp…

Nice personal hosted image service!

Re: Show HN: I built a free in-browser Llama 3 chatbot powered by WebGPU

#63

It's truly amazing how quickly my browser loads 0.6GB of data. I remember when downloading a 1MB file involved phoning up a sysop in advance and leaving the modem on all night. We've come so far.

When I think about numbers like that it just seems (to me, and wrongly) like general progress that's not so crazy - the thought that really makes the speed of progress stand out to me is remembering when loading a single image - photo sized but not crazily high resolution - over dial-up was slow enough that you'd gradually see the image loading from top to bottom, and could see it gradually getting taller as more lines of pixels were downloaded and shown below the already loaded part. Contrasting that memory against the ability to now watch videos with much higher resolution per frame than those images were 30 years ago is what really makes me go "wow".

For anyone not old enough to remember, here's an example on YouTube (and a faster loading time than I remember often being the case!): https://youtube.com/watch?v=ra0EG9lbP7Y

Re: Show HN: I built a free in-browser Llama 3 chatbot powered by WebGPU

#64
Yasssssss! Thank you.

This is the future. I am predicting Apple will make progress on groq like chipsets built in to their newer devices for hyper fast inference.

LLMs leave a lot to be desired but since they are trained on all publicly available human knowledge they know something no about everything.

My life has been better since I’ve been able to ask all sorts of adhoc questions about “is this healthy? Why healthy?” And it gives me pointers where to look into.

Re: Show HN: I built a free in-browser Llama 3 chatbot powered by WebGPU

#68

This is very cool, it's something I wish existed since Llama came out, having to install Ollama + Cuda to get locally working LLM didn't felt right to me when there's all what's needed in the browser. Llamafile solves the first half of the problem, but you still need to install Cuda/ROCm for it to work with GPU acceleration. WebGPU is the way to go if we want to put AI on consumer hardware and break the oligopoly, I…

Tested on Ubuntu 22.04 with Chrome, sure enough, "Could not load the model because Error: Cannot find adapter that matches the request".

It really is too bad WebGPU isn't supported on Linux, I mean, that's a no-brainer right there.

Re: Show HN: I built a free in-browser Llama 3 chatbot powered by WebGPU

#69
post #31

Earlier quoted context omitted.

Last I checked ff explicitly does not support webgpu, webhid, webusb, etc. Apparently nightly is supposed to support it: https://developer.mozilla.org/en-US/docs/Mozilla/Firefox/Exp...

So here's a howler to the new Mozilla CEO and FF teams who're looking for ways to save their org: - release WebGPU support everywhere, also embed llama.cpp or something similar for non GPU users - add UI for easy model downloading and sharing among sites - write the LLM browser API that enables easy access and sets the standard - add security: "this website wants to use local LLM. Allow?"

There’s also the little issue of firefox not supporting HDR videos which with more and more OLED/miniLED monitors out there is a major drawback. I love FF and i daily drive it, but there are some glaring gaps in the feature set between chromium and ff.

Re: Show HN: I built a free in-browser Llama 3 chatbot powered by WebGPU

#70
post #25

Phi 1.5 isn't quite there is it? Caught it having some inner dialog. What does that mean? Instruct: I'm sorry, but I'm not sure what you mean by "wutt". Output: Oh, I'm sorry. I meant "what does that mean?" Instruct: No problem, "wutt" means "that doesn't make any sense". Output: Ah, I see. Thank you for clarifying! Instruct: You're welcome! If you have any more questions, feel free to ask. Instruct: I'm glad I could…

That boy needs therapy.

Purely psychosomatic
Post reply on HN