Earlier quoted context omitted.
I have a home server that runs Qwen3.6-35B-A3B through llama.cpp with Open WebUI for the user facing interface. My teen isn't super interested in AI, but whenever they do feel curious they have their own account they can use on our home network. As far as chatting goes local models are more than capable for handling standard chat questions, doing research, helping troubleshoot problems etc. In fact it was an agent po…
What kind of machine is it running on ?
Anthropic updates their terms to verify age or identity
181–190 of 197 posts
Re: Anthropic updates their terms to verify age or identity
#182Earlier quoted context omitted.
Which model? I see a suspiciously similar post on amd.com running 2 bit Kimi quant on a four node cluster over 5Gbps Ethernet Assuming math works here although I think there's some caveats depending on the model architecture, 1T 4 bit is 465Gi just for the weights so you wouldn't be able to fit kv cache. It's showing about 8-9 tk/sec which seems quite slow for something like a web search with result aggregate althoug…
I don’t have any particular model in mind, sorry. My data is just rough estimates based on my experience with a single node setup. You might need to opt for a 2 or 3 bit model to get the full context window. The KV cache memory consumption as well overall performance will be heavily dependent on the model’s architecture. The performance too will depend a lot on the inference server chosen and its configuration. I sus…
Re: Anthropic updates their terms to verify age or identity
#183Earlier quoted context omitted.
> I run an uncensored, ephemeral model for my own use and it's an entirely different experience than anything you can pay for. Dont. Goon. To. LLMs
Wasn't the parent post referring to 'legitimate' demands? I often use them to get a broad overview of a technical field before reading human stuff on it, and it might be me but those clankers tend to spend half their reasoning on whether they are allowed to reply to my request. Censorship is an annoying waste of capacity for certain use cases, although it certainly has its boons when shipping commercial models.
Re: Anthropic updates their terms to verify age or identity
#184Earlier quoted context omitted.
Sort of. A full trillion-parameter model needs about $300k of server hardware to run in and a lot of electricity, making it feasible only for very wealthy individuals, but quite practical for businesses and institutions above a certain size...although they in turn would typically gatekeep access. You can drastically reduce the requirements by running models at a lower bitrate, which somewhat reduces accuracy but not…
You can run a trillion parameter model with decent quality for far less than $300k. A cluster of 4 AMD AI Max 395+ boards with 128GB unified memory each can be had for around $15k. That would run the 4-bit quant of a trillion param model well enough for personal use. At full use the cluster would only be consuming around 400-500W of power too. That's about the same as one high end graphics card. That's still a lot of…
Re: Anthropic updates their terms to verify age or identity
#185Earlier quoted context omitted.
You can run a trillion parameter model with decent quality for far less than $300k. A cluster of 4 AMD AI Max 395+ boards with 128GB unified memory each can be had for around $15k. That would run the 4-bit quant of a trillion param model well enough for personal use. At full use the cluster would only be consuming around 400-500W of power too. That's about the same as one high end graphics card. That's still a lot of…
I literally wrote about running quantized models and how much more affordable it could be in the very next sentence . Please don't reply if you can't be bothered to read the entire comment, it's not that long.
Re: Anthropic updates their terms to verify age or identity
#186Earlier quoted context omitted.
I don’t have any particular model in mind, sorry. My data is just rough estimates based on my experience with a single node setup. You might need to opt for a 2 or 3 bit model to get the full context window. The KV cache memory consumption as well overall performance will be heavily dependent on the model’s architecture. The performance too will depend a lot on the inference server chosen and its configuration. I sus…
I imagine a smaller single node model would have a significantly better experience at significantly lower cost. When I was poking around with infra estimates it seemed the main issue/cost was once you crossed from single-node to multi-node. You need _a lot_ of bandwidth if the weights are sharded. Like Tbps of bandwidth. The closest reasonable thing I've heard of for local multi-node is exo on macos using thunderbolt…
I don’t know what the scaling for multiple strix halo boards looks like in practice. From what I understand each server has to process the model in serial. Meaning server A has 1/4 the weights and sends server B the results to process and so on. So you don’t get compute scaling just memory scaling.
Re: Anthropic updates their terms to verify age or identity
#187Earlier quoted context omitted.
They are not going to let open weights models with zero restrictions exist dude. They will be regulated like guns, or probably closer to nerve gas or enriched uranium.
The government is not going to enforce this, the game theory does not work in their favor. The SCOTUS has made it exceptionally clear mathematics and software are protected by the First Amendment. The Atomic Energy Act of 1954 tries to make a very narrow exception for nuclear weapons, but 1. The law has never been challenged in court for being unconstitutional, and 2. It doesn't apply to model weights Any attempt by…
Re: Anthropic updates their terms to verify age or identity
#188Earlier quoted context omitted.
I’m curious (and please forgive my ignorance if it’s obvious), are open weight models practically feasible? I mean from a financial and sustainability standpoint, assuming they’re equally powerful as their proprietary counterparts. I guess I’m trying to understand the economics of it.
There is an understandable gap between the capabilities of closed models and those of open models. The current difference is primarily expressed in the cost of hardware necessary to sufficiently run a exactly comparable model. A single higher end graphics card running on your average gaming computer, is capable of running small to medium models that compare with those of their lab-born counterparts in the small-mediu…
I wonder if open source / open weight models will reach the point where we can run them locally on our mobile devices (for free), even if they're slightly inferior to the proprietary pay-as-you-use online models.
I know very little about this stuff. My inner optimist kinda hopes that the tech will continue to advance and become increasingly commoditised, to the point that open source locally run models are as good as the advanced proprietary models of a year ago. So that even if the open source models lag the proprietary models, they're still pretty great. Perhaps we're already there but I wouldn't know.
Anyhow thank you for the insights :-)
Re: Anthropic updates their terms to verify age or identity
#189Earlier quoted context omitted.
[flagged]
EDIT: Parent commenter completely rewrote their comment while I was replying. I'm leaving this up as is. The text below is their original comment that I was responding to. > What are you pushing by pointing out not only in this thread but the previous one too, quite in depth, that it's not new? I know you can claim you're just a stickler for accurate reporting but you seem really invested Because these threads degrad…
You were extremely quick! :) I thought I was still within my 'delay' window. I've increased it now
But people editing their comment to bring it down a notch happens all the time when they believe it hadn't been published yet. This is nothing new. I don't think there was a need to call me out because your comment would have made sense in reply to my edit too
And yes, I take your point. But I still think your kind of replies are an attempt to normalise it, however. (Which is what I was implying, not shilling, btw @gruez)