Live data from Hacker News

Ollama Turbo

ollama.com

91–100 of 251 posts

Re: Ollama Turbo

#91
post #10

Nice release. Part of the problem right now with OSS models (at least for enterprise users) is the diversity of offerings in terms of: - Speed - Cost - Reliability - Feature Parity (eg: context caching) - Performance (What quant level is being used...really?) - Host region/data privacy guarantees - LTS And that's not even including the decision of what model you want to use! Realistically if you want to use an OSS mo…

true but ignores handing over all your prompt traffic without any real legal protections as sama has pointed out: [1] https://californiarecorder.com/sam-altman-requires-ai-privil...

> OpenAI confirmed it has been preserving deleted and non permanent person chat logs since mid-Might 2025 in response to a federal court docket order

> The order, embedded under and issued on Might 13, 2025, by U.S. Justice of the Peace Decide Ona T. Wang

Is this some meme where “may” is being replaced with “might”, or some word substitution gone awry? I don’t get it.

Re: Ollama Turbo

#92
post #62

I see a lot of hate for ollama doing this kind of thing but also they remain one of the easiest to use solutions for developing and testing against a model locally. Sure, llama.cpp is the real thing, ollama is a wrapper... I would never want to use something like ollama in a production setting. But if I want to quickly get someone less technical up to speed to develop an LLM-enabled system and run qwen or w/e locally…

> I would never want to use something like ollama in a production setting. We benchmarked vLLM and Ollama on both startup time and tokens per seconds. Ollama comes at the top. We hope to be able to publish these results soon.

Did you test multi-user cases?

Re: Ollama Turbo

#93

Earlier quoted context omitted.

[flagged]

This is not true. No inference engine does all of: - Model switching - Unload after idle - Dynamic layer offload to CPU to avoid OOM

this can be added to llama.cpp with llama.swap currently so even without Ollama you are not far off

Re: Ollama Turbo

#94
Called it.

It's very unfortunate that the local inference community has aggregated around Ollama when it's clear that's not their long term priority or strategy.

Its imperative we move away ASAP

Re: Ollama Turbo

#95

Does anyone know if this is like like OpenRouter?

Often the math works out that you get a lot more for $20 a month if you settle for smaller sized but capable models (8b-30b). I don’t see how it’s better other than Ollama can “promise” they don’t store your data where as OpenRouter is dependent on which host you choose (and there’s no indicator on OpenRouter exposing which ones do or don’t).

In a universe where everything you say can be taken out of context, things like OpenAi will be a data leak nightmare.

Need this soon:

https://arxiv.org/abs/2410.02486

Re: Ollama Turbo

#96

Earlier quoted context omitted.

true but ignores handing over all your prompt traffic without any real legal protections as sama has pointed out: [1] https://californiarecorder.com/sam-altman-requires-ai-privil...

> OpenAI confirmed it has been preserving deleted and non permanent person chat logs since mid-Might 2025 in response to a federal court docket order > The order, embedded under and issued on Might 13, 2025, by U.S. Justice of the Peace Decide Ona T. Wang Is this some meme where “may” is being replaced with “might”, or some word substitution gone awry? I don’t get it.

:)) Apparently. I don't have a better guess. Well spotted

Re: Ollama Turbo

#97
post #35

I see a lot of hate for ollama doing this kind of thing but also they remain one of the easiest to use solutions for developing and testing against a model locally. Sure, llama.cpp is the real thing, ollama is a wrapper... I would never want to use something like ollama in a production setting. But if I want to quickly get someone less technical up to speed to develop an LLM-enabled system and run qwen or w/e locally…

Thanks for the kind words. Since the new multimodal engine, Ollama has moved off of llama.cpp as a wrapper. We do continue to use the GGML library, and ask hardware partners to help optimize it. Ollama might look like a toy and what looks trivial to build. I can say, to keep its simplicity, we go through a deep amount of struggles to make it work with the experience we want. Simplicity is often overlooked, but we wan…

But Ollama is a toy, it's meaningful for hobbyists and individuals to use locally like myself. Why would it be the right choice for anything more? AWS, vLLM, SGLang etc would be the solutions for enterprise

I knew a startup that deployed ollama on a customers premises and when I asked them why, they had absolutely no good reason. Likely they did it because it was easy. That's not the "easy to use" case you want to solve for.

Re: Ollama Turbo

#98
post #63

What could be the benefit of paying $20 to Ollama to run inferior models instead of paying the same amount of money to e.g. OpenAI for access to sota models?

Groq seems to do okay with a similar service but I think their pricing is probably better.

Yeah, the NAZI sex not will be great for business!

Re: Ollama Turbo

#99

Called it. It's very unfortunate that the local inference community has aggregated around Ollama when it's clear that's not their long term priority or strategy. Its imperative we move away ASAP

hmm, how so? Ollama is open and the pricing is completely optional for users who want additional GPUs.

Is it bad to fairly charge money for selling GPUs that cost us money too, and use that money to grow the core open-source project?

At one point, it just has to be reasonable. I'd like to believe by having a conscientious, we can create something great.

Re: Ollama Turbo

#100
post #23
post #15

Watching ollama pivot from a somewhat scrappy yet amazingly important and well designed open source project to a regular "for-profit company" is going to be sad. Thankfully, this may just leave more room for other open source local inference engines.

Their FOSS local inference service didn't go anywhere. This isn't Anaconda, they didn't do a bait and switch to screw their core users. It isn't sinful for devs to try and earn a living.

You can build this and go build something else as well. You don't need to morph the thing you built. That's underhanded
Post reply on HN