Live data from Hacker News

Ggml.ai joins Hugging Face to ensure the long-term progress of Local AI

github.com

101–110 of 241 posts

Re: Ggml.ai joins Hugging Face to ensure the long-term progress of Local AI

#101
post #95
post #84

It's hard to overstate the impact Georgi Gerganov and llama.cpp have had on the local model space. He pretty much kicked off the revolution in March 2023, making LLaMA work on consumer laptops. Here's that README from March 10th 2023 https://github.com/ggml-org/llama.cpp/blob/775328064e69db1eb... > The main goal is to run the model using 4-bit quantization on a MacBook. [...] This was hacked in an evening - I have no…

i am curious, why are your comments always pinned to the top?

At a guess that's because my comment attracted more up-votes than the other top-level comments in the thread.

I generally try to include something in a comment that's not information already under discussion - in this case that was the link and quote from the original README.

Re: Ggml.ai joins Hugging Face to ensure the long-term progress of Local AI

#102
post #25

Can anyone point me in the direction of getting a model to run locally and efficiently inside something like a Docker container on a system with not so strong computing power (aka a Macbook M1 with 8gb of memory)? Is my only option to invest in a system with more computing power? These local models look great, especially something like https://huggingface.co/AlicanKiraz0/Cybersecurity-BaronLLM_O... for assisting in p…

With only 8 GB of memory, you're going to be running a really small quant, and it's going to be slow and lower quality. But yes, it should be doable. In the worst case, find a tiny gguf and run it on CPU with llamafile.

Re: Ggml.ai joins Hugging Face to ensure the long-term progress of Local AI

#103
post #101
post #95

Earlier quoted context omitted.

i am curious, why are your comments always pinned to the top?

At a guess that's because my comment attracted more up-votes than the other top-level comments in the thread. I generally try to include something in a comment that's not information already under discussion - in this case that was the link and quote from the original README.

of course your comment attracts more upvotes - it's at the top.

Re: Ggml.ai joins Hugging Face to ensure the long-term progress of Local AI

#104
post #95
post #84

It's hard to overstate the impact Georgi Gerganov and llama.cpp have had on the local model space. He pretty much kicked off the revolution in March 2023, making LLaMA work on consumer laptops. Here's that README from March 10th 2023 https://github.com/ggml-org/llama.cpp/blob/775328064e69db1eb... > The main goal is to run the model using 4-bit quantization on a MacBook. [...] This was hacked in an evening - I have no…

i am curious, why are your comments always pinned to the top?

[deleted]

Re: Ggml.ai joins Hugging Face to ensure the long-term progress of Local AI

#105

Earlier quoted context omitted.

Because many of us think simonw has discerning taste on this topic and like to read what he has to say about it, so we upvote his comments.

i don't doubt this. i just find it questionable that one particular poster always gets in the spotlight when AI is the topic - while other conversations in my opinion offer more interesting angles.

Upvote the conversations that you find to be more interesting. If enough people do the same, they too will make it to the top.

Re: Ggml.ai joins Hugging Face to ensure the long-term progress of Local AI

#106
post #18

I consider HuggingFace more "Open AI" than OpenAI - one of the few quiet heroes (along with Chinese OSS) helping bring on-premise AI to the masses. I'm old enough to remember when traffic was expensive, so I've no idea how they've managed to offer free hosting for so many models. Hopefully it's backed by a sustainable business model, as the ecosystem would be meaningfully worse without them. We still need good value…

> We still need good value hardware to run Kimi/GLM in-house If you stream weights in from SSD storage and freely use swap to extend your KV cache it will be really slow (multiple seconds per token!) but run on basically anything. And that's still really good for stuff that can be computed overnight, perhaps even by batching many requests simultaneously. It gets progressively better as you add more compute, of course…

> it will be really slow (multiple seconds per token!)

This is fun for proving that it can be done, but that's 100X slower than hosted models and 1000X slower than GPT-Codex-Spark.

That's like going from real time conversation to e-mailing someone who only checks their inbox twice a day if you're lucky.

Re: Ggml.ai joins Hugging Face to ensure the long-term progress of Local AI

#107
post #38

Earlier quoted context omitted.

There’s no way around needing a powerful-enough system to run the model. So you either choose a model that can fit on what you have —i.e. via a small model, or a quantised slightly larger model— or you access more powerful hardware, either by buying it or renting it. (IME you don’t need Docker. For an easy start just install LM Studio and have a play.) I picked up a second-hand 64GB M1 Max MacBook Pro a while back fo…

Are mac kernels optimized compared to CUDA kernels? I know that the unified GPU approach is inherently slower, but I thought a ton of optimizations were at the kernel level too (CUDA itself is a moat)

Mac kernels are almost always compute shaders written in Metal. That's the bare-minimum of acceleration, being done in a non-portable proprietary graphics API. It's optimized in the loosest sense of the word, but extremely far from "optimal" relative to CUDA (or hell, even Vulkan Compute).

Most people will not choose Metal if they're picking between the two moats. CUDA is far-and-away the better hardware architecture, not to mention better-supported by the community.

Re: Ggml.ai joins Hugging Face to ensure the long-term progress of Local AI

#108
post #50
post #18

I consider HuggingFace more "Open AI" than OpenAI - one of the few quiet heroes (along with Chinese OSS) helping bring on-premise AI to the masses. I'm old enough to remember when traffic was expensive, so I've no idea how they've managed to offer free hosting for so many models. Hopefully it's backed by a sustainable business model, as the ecosystem would be meaningfully worse without them. We still need good value…

It's insane how much traffic HF must be pushing out of the door. I routinely download models that are hundreds of gigabytes in size from them. A fantastic service to the sovererign AI community.

My fear is that these large "AI" companies will lobby to have these open source options removed or banned, growing concern. I'm not sure how else to explain how much I enjoy using what HF provides, I religiously browse their site for new and exciting models to try.

Re: Ggml.ai joins Hugging Face to ensure the long-term progress of Local AI

#109
post #95
post #84

It's hard to overstate the impact Georgi Gerganov and llama.cpp have had on the local model space. He pretty much kicked off the revolution in March 2023, making LLaMA work on consumer laptops. Here's that README from March 10th 2023 https://github.com/ggml-org/llama.cpp/blob/775328064e69db1eb... > The main goal is to run the model using 4-bit quantization on a MacBook. [...] This was hacked in an evening - I have no…

i am curious, why are your comments always pinned to the top?

HN goes through phases. I remember when patio11 was the star of the hour on here. At another time it was that security guy (can't remember his name).

And for those who think it's just organic with all of the upvotes, HN absolutely does have a +/- comment bias for users, and it does automatically feature certain people and suppress others.

Re: Ggml.ai joins Hugging Face to ensure the long-term progress of Local AI

#110
post #101

Earlier quoted context omitted.

At a guess that's because my comment attracted more up-votes than the other top-level comments in the thread. I generally try to include something in a comment that's not information already under discussion - in this case that was the link and quote from the original README.

of course your comment attracts more upvotes - it's at the top.

Attention feeds attention.

Attention is ALL You Need.

Post reply on HN