Live data from Hacker News

Ollama Turbo

ollama.com

211–220 of 251 posts

Re: Ollama Turbo

#211
post #144

Earlier quoted context omitted.

Everyone just wants to solarpunk this up.

In an ideal world yes - as we should - especially for us Californian/Bay Area people, that's literally our spirit animal. But I understand that is idle dreaming. What I believe certainly is within reach is a state that is much better than what we are in.

It needn't be idle dreaming? What fundamental law or societal agreement prevents solarpunk versus the current status quo of corporate anti-human cyberpunk?

Re: Ollama Turbo

#212

Earlier quoted context omitted.

What? The obvious move is to never have switched to Ollama and just use Llama.cpp directly, which I've been doing for years. Llama.cpp was created first, is the foundation for this product, and is actually open source.

But there's much less that works with that. OpenWebUI for example.

Open WebUI works perfectly fine with llama.cpp though.

They have very detailed quick start docs on it: https://docs.openwebui.com/getting-started/quick-start/start...

Re: Ollama Turbo

#213
Does this mean we can access Ollama APIs for $20/mo and test them without running the model locally? I'm not hardware-rich, but for some projects, I'd like a reliable pricing.

Re: Ollama Turbo

#214
post #150
post #33

Any more information on "Privacy first"? It seems pretty thin if just not retaining data. For Draw Things provided "Cloud Compute", we don't retain any data too (everything is done in RAM per request). But that is still unsatisfactory personally. We will soon add "privacy pass" support, but still not to the satisfactory. Transparency log that can be attested on the hardware would be nice (since we run our open-source…

I don't see a privacy policy and their desktop app is closed source. So, not encouraging. [full disclosure I am working on something with actual privacy guarantees for LLM calls that does use a transparency log, etc.]

I’d love to learn more about your project. I’m using socialized cloud regions for AI security and they really lag the mainstream. Definitely need more options here.

Edit: emailed the address on the site in your profile, got an inbox does not exist error.

Re: Ollama Turbo

#215
post #70

It says “usage-based pricing” is coming soon. I think that is the sweet spot for a service like this. I pay $20 to Anthropic, so I don’t think I’d get enough use out of this for the $20 fee. But being able to spin up any of these models and use as needed (and compare) seems extremely useful to me. I hope this works out well for the team.

A flat fee service for open-source LLMs is somewhat unique, even if I don't see myself paying for it.

Usage-based pricing would put them in competition with established services like deepinfra.com, novita.ai, and ultimately openrouter.ai. They would go in with more name-recognition, but the established competition is already very competitive on pricing

Re: Ollama Turbo

#217
post #192
post #35

Earlier quoted context omitted.

Thanks for the kind words. Since the new multimodal engine, Ollama has moved off of llama.cpp as a wrapper. We do continue to use the GGML library, and ask hardware partners to help optimize it. Ollama might look like a toy and what looks trivial to build. I can say, to keep its simplicity, we go through a deep amount of struggles to make it work with the experience we want. Simplicity is often overlooked, but we wan…

This kind of gaslighting is exactly why I stopped using Ollama. GGML library is llama.cpp. They are one and the same. Ollama made sense when llama.cpp was hard to use. Ollama does not have value preposition anymore.

> GGML library is llama.cpp. They are one and the same.

Nope…

Re: Ollama Turbo

#218

Earlier quoted context omitted.

So I’m using turbo and just want to provide some feedback. I can’t figure out how to connect raycast and project goose to ollama turbo. The software that calls it essentially looks for the models via ollama but cannot find the turbo ones and the documentation is not clear yet. Just my two cents, the inference is very quick and I’m happy with the speed but not quite usable yet.

so sorry about this. We are learning. Possible to email, and we will first make it right while we improve Ollama's turbo mode. hello@ollama.com

no worries. i totally understand that the first day something is released it doesn’t work perfectly with third party/community software.

thanks for the feedback address :)

Re: Ollama Turbo

#219

Earlier quoted context omitted.

I feel the primary benefit of this Ollama Turbo is that you can quickly test and run different models in the cloud that you could run locally if you had the correct hardware. This allows you to try out some open models and better assess if you could buy a dgx box or Mac Studio with a lot of unified memory and build out what you want to do locally without actually investing in very expensive hardware. Certain applicat…

Quickly test… the two models they support? This is just another subscription to quantized models.

it looks like the plan is to support way more models though. gotta start somewhere.

Re: Ollama Turbo

#220
post #211

Earlier quoted context omitted.

In an ideal world yes - as we should - especially for us Californian/Bay Area people, that's literally our spirit animal. But I understand that is idle dreaming. What I believe certainly is within reach is a state that is much better than what we are in.

It needn't be idle dreaming? What fundamental law or societal agreement prevents solarpunk versus the current status quo of corporate anti-human cyberpunk?

Being realistic about economics and how money works in the current paradigm where it is concentrated
Post reply on HN