Earlier quoted context omitted.
Just want to echo the recommendation for qwen3.5:9b. This is a smol, thinking, agentic tool-using, text-image multimodal creature, with very good internal chains of thought. CoT can be sometimes excessive, but it leads to very stable decision-making process, even across very large contexts -something we haven't seen models of this size before. What's also new here, is VRAM-context size trade-off: for 25% of it's atte…
[flagged]
Can I run AI locally?
271–280 of 382 posts
Re: Can I run AI locally?
#272Earlier quoted context omitted.
Do you also require computers to grow legs when they "run"? "Thinking" is just a term to describe a process in generative AI where you generate additional tokens in a manner similar to thinking a problem through. It's kind of a tired point to argue against the verb since it's meaning is well understood at this point
I am a professional in the information technology field, which is to say a pedantic extremist who believes that words have meanings derived from consensus, and when people alter the meanings, they alter what they believe. Using "thinking", "feeling", "alive", or otherwise referring to a current generation LLM as a creature is a mistake which encourages being wrong in further thinking about them.
There is this ancient story where man was created to mine gold in SA. There was some disagreement whether or not to delete the creatures afterwards. The jury is still out on what the point is.
Consulting our feelings seems good, the feelings were trained on millions of years worth of interactions. Non of them were this tho.
What would be the point for you of uhh robotmancipation?
Edit: for me it would get complicated if it starts screaming and begging not to be deleted. Which I know makes no sense.
Re: Can I run AI locally?
#273Besides trying to run on your own hardware, anybody have recommendations for running some decent models on one of the many "AI clouds" providers? This is for sporadic use and so maybe one of the "serverless" providers that bill by the hour or minute or similar as opposed to monthly renting GPUs. There are quite a few of them but their marketing is just confusing and full of buzz words. I've been tinkering with OpenRo…
use openrouter, and call it a day. auto switching between providers, connectivity to all clouds and even works with free models
Re: Can I run AI locally?
#274This site presents models in an incomplete and misleading way. When I visit the site with an Apple M1 Max with 32GB RAM, the first model that's listed is Llama 3.1 8B, which is listed as needing 4.1GB RAM. But the weights for Llama 3.1 8B are over 16GB. You can see that here in the official HF repo: https://huggingface.co/meta-llama/Llama-3.1-8B/tree/main The model this site calls 'Llama 3.1 8B' is actually a 4-bit q…
Re: Can I run AI locally?
#275Earlier quoted context omitted.
I took a long break from document processing after working on it heavily 20 years ago. The tools I used before were ABBYY FineReader and PrimeOCR. I haven't tried any of the commercial cloud based solutions. I'm currently using GLM-OCR, Chandra OCR, and Apple's LiveText in conjunction with each other (plus custom code for glue functionality and downstream processing). Try just GLM-OCR if you want to get started quick…
How does GLM-OCR compare to Qwen 3 VL? I've had good experiences with Qwen for these purposes.
Re: Can I run AI locally?
#276Earlier quoted context omitted.
A million tokens is like 5 minutes of inference for heavy coding use.
At work I regularly hit my 7.5mil tokens per hour limit one of our tools has, and have to switch model of tool, and I’m not even really a remotely heavy user. I think people don’t realise how many tokens get burned with CoT and tool calls these days At 7.5mil per hour hard limit, 84 days to hit the grandparents $3k That said local models really are slow still, or fast enough and not that great
Re: Can I run AI locally?
#277Earlier quoted context omitted.
Just want to echo the recommendation for qwen3.5:9b. This is a smol, thinking, agentic tool-using, text-image multimodal creature, with very good internal chains of thought. CoT can be sometimes excessive, but it leads to very stable decision-making process, even across very large contexts -something we haven't seen models of this size before. What's also new here, is VRAM-context size trade-off: for 25% of it's atte…
You can really see the limitations of qwen3.5:9b in reasoning traces- it’s fascinating. When a question “goes bad”, sometimes the thinking tokens are WILD - it’s like watching the Poirot after a head injury. Example: “what is the air speed velocity of a swallow?” - qwen knew it was a Monty Python gag, but couldnt and didnt figure out which one.
Re: Can I run AI locally?
#278Re: Can I run AI locally?
#279Earlier quoted context omitted.
Just want to echo the recommendation for qwen3.5:9b. This is a smol, thinking, agentic tool-using, text-image multimodal creature, with very good internal chains of thought. CoT can be sometimes excessive, but it leads to very stable decision-making process, even across very large contexts -something we haven't seen models of this size before. What's also new here, is VRAM-context size trade-off: for 25% of it's atte…
[flagged]
Re: Can I run AI locally?
#280Earlier quoted context omitted.
Do you also require computers to grow legs when they "run"? "Thinking" is just a term to describe a process in generative AI where you generate additional tokens in a manner similar to thinking a problem through. It's kind of a tired point to argue against the verb since it's meaning is well understood at this point
I am a professional in the information technology field, which is to say a pedantic extremist who believes that words have meanings derived from consensus, and when people alter the meanings, they alter what they believe. Using "thinking", "feeling", "alive", or otherwise referring to a current generation LLM as a creature is a mistake which encourages being wrong in further thinking about them.