Live data from Hacker News

Ollama Turbo

ollama.com

61–70 of 251 posts

Re: Ollama Turbo

#61
post #15

Watching ollama pivot from a somewhat scrappy yet amazingly important and well designed open source project to a regular "for-profit company" is going to be sad. Thankfully, this may just leave more room for other open source local inference engines.

[flagged]

"Please don't post shallow dismissals, especially of other people's work. A good critical comment teaches us something."

https://news.ycombinator.com/newsguidelines.html

Re: Ollama Turbo

#62

I see a lot of hate for ollama doing this kind of thing but also they remain one of the easiest to use solutions for developing and testing against a model locally. Sure, llama.cpp is the real thing, ollama is a wrapper... I would never want to use something like ollama in a production setting. But if I want to quickly get someone less technical up to speed to develop an LLM-enabled system and run qwen or w/e locally…

> I would never want to use something like ollama in a production setting.

We benchmarked vLLM and Ollama on both startup time and tokens per seconds. Ollama comes at the top. We hope to be able to publish these results soon.

Re: Ollama Turbo

#63
What could be the benefit of paying $20 to Ollama to run inferior models instead of paying the same amount of money to e.g. OpenAI for access to sota models?

Re: Ollama Turbo

#64
post #63

What could be the benefit of paying $20 to Ollama to run inferior models instead of paying the same amount of money to e.g. OpenAI for access to sota models?

nothing lmao. this is just ollama trying to make money.

Re: Ollama Turbo

#65

Earlier quoted context omitted.

I think what matters more here is "All hardware is located outside of China". Located in the US means little because that's not good enough for many regulated industries even within the US. All things considered though, Europe is getting confusing. They have GDPR but now pushing to backdoor encryption within the EU? [1] At least there isn't a strong movement in the US trying to outlaw E2E encryption. [1] https://www.…

Maybe I hit a nerve with the EU part? I thought it was a fair observation, but I'm open to being corrected if there's more nuance I missed.

The bill has been stalled since 2022.

Yes, there is gonna be a new discussion for it on October 15, but I've already seen section of governments being against their own government position on the bill (Swedish Military for example).

Re: Ollama Turbo

#66
post #31
post #15

Watching ollama pivot from a somewhat scrappy yet amazingly important and well designed open source project to a regular "for-profit company" is going to be sad. Thankfully, this may just leave more room for other open source local inference engines.

we have always been building in the open, and so is Ollama. All the core pieces of Ollama are open. There are areas where we want to be opinionated on the design to build the world we want to see. There are areas we will make money, and I wholly believe if we follow our conscious we can create something amazing for the world while making sure we can keep it fueled to keep it going for the long term. Some of the ideas…

I wanted to try web search to increase my privacy but it wanted to do login.

For Turbo mode I understand the need for paying but the main poing of running a local model with web search is browsing from my computer without using any LLM provider. Also I want to get rid of the latency to US servers from Europe.

If ollama can't do it, maybe a fork.

Re: Ollama Turbo

#67
post #63

What could be the benefit of paying $20 to Ollama to run inferior models instead of paying the same amount of money to e.g. OpenAI for access to sota models?

I run a lot of mundane jobs that work fine with less capable models, so I can see the potential benefit. It all depends on the limits though.

Re: Ollama Turbo

#68
post #58

Distractions like this probably the reason they still, over a year now, do not support sharded GGUF. https://github.com/ollama/ollama/issues/5245 If any of the major inference engines - vLLM, Sglang, llama.cpp - incorporated api driven model switching, automatic model unload after idle and automatic CPU layer offloading to avoid OOM it would avoid the need for ollama.

That’s just llama-swap and llama.cpp

Interesting - it does indeed seem like llama-server has the needed endpoints to do the model swapping and llama.cpp as of recently also has a new flag for the dynamic CPU offload now.

However the approach to model swapping is not 'ollama compatible' which means all the OSS tools supporting 'ollama' Ex Openwebui, Openhands, Bolt.diy, n8n, flowise, browser-use etc.. aren't able to take advantage of this particularly useful capability as best I can tell.

Re: Ollama Turbo

#69
post #49

Earlier quoted context omitted.

Any evidence for this claim that e.g. Mossad has less penetration into digital systems of USA than it does RF or PRC?

They might have access to any given machine, but they lack the broad scope of general surveillance. If they want to get you, just like most of the other nation state level threats, you will get got. For other threat models, the US works pretty well. I guarantee that nobody cares about or will be surveilling your private AI use unless you're doing other things that warrant surveillance. The reason big providers suck,…

Nobody cares? That seems ludicrous to me. The last 3 decades of business have been characterized most of all by the increased access of private information on people for online business competitive insights. Sure if you are just a consumer you have nothing of real value except in the aggregate, but if you are an up-and-coming business drawing customers away from other businesses, your private AI use is absolutely of interest. Which is why serious businesses here scour the ToS.

The biggest game in town has been managing platforms that give owners an information advantage. But at least the world generally trusts the USA to abide by laws and user agreements, which is why, to my mind, the USA retains the near monopoly on information platforms.

I personally wouldn’t trust a UK platform for example, being a Brit native. The top echelon talent pool is so small and incestuous I don’t believe I would experience a fair playing field if a business of mine passed a certain size of national reach/importance.

EDIT: from ChatGPT, new money entrepreneurs with no inheritence/political ties by economic region, USA ~63%, UK/HongKong/Singapore ~45%, Emerging Markets ~35%, EU ~22%, Russia ~10%

Re: Ollama Turbo

#70
It says “usage-based pricing” is coming soon. I think that is the sweet spot for a service like this.

I pay $20 to Anthropic, so I don’t think I’d get enough use out of this for the $20 fee. But being able to spin up any of these models and use as needed (and compare) seems extremely useful to me.

I hope this works out well for the team.

Post reply on HN