I read those blog posts to remind of wealth gap i have with the average hackernews user.
Thank you for the encouragement, i will work harder to reach your level.
191–200 of 207 posts
I read those blog posts to remind of wealth gap i have with the average hackernews user.
Thank you for the encouragement, i will work harder to reach your level.
Earlier quoted context omitted.
I have quoted large nodes from this supplier and have lots of^W^W GPUs from them for personal use. Current lead time is more than 30 months. They're a good provider but you have to be a big shot buying NVL72s before you're getting anything within your payback period.
Ah thanks for the solid info, too bad. I'd seen them come up as a pretty good price for 6000 RTX's in the past, which seem generally pretty available, good source for those?
Earlier quoted context omitted.
versus $0 with local models. There will always be a reason to run frontier models, but local models are well at levels that assist with stuff that don't need that level of complexity.
You could buy a $150 refurbished 16gb i5 and use that model until openAI does its IPO and has to become sane again. But I guess a $2000 Mac is probably better if you don't care about cost or quality.
Or I can use the Mac I already have.
Your example though, Ouch!
~8B Q4. That's around 5-10 tokens a second. Base M1 16GB mac would do 15-20 tokens a seconds. That's a 6 year old machine.
You do get what you pay for it seems.
Earlier quoted context omitted.
Ah thanks for the solid info, too bad. I'd seen them come up as a pretty good price for 6000 RTX's in the past, which seem generally pretty available, good source for those?
Oh my god, I completely spaced dude. Weeks, not months. Weeks. Sorry.
Nice setup; but, for simple tasks or questions, AI is currently free? And it will probably stay free, as I don't see Google starting to charge for using AI on its search engine? So costs can't be a motivation for running small models locally? For more complex or important tasks, costs, autonomy and privacy matter, but then so does performance/quality. So I'm not completely convinced it's really worth it; but it's tem…
You are correct. In large part, the cost of something like Gemini on a very basic Google AI plan provides far more utility than local LLMs for coding assistance. There are 2 main reasons for running local LLMS. 1. Process private data/work with uncensored models. 2. Use a large amount of inference that would quickly blow through rate limits and/or run up API costs. The thing that is critical for 2 is that a) you have…
I have a beefy Linux box with a 4090 but never took the time to set it up properly beyond simple testing; any tutorial you would recommend?
Earlier quoted context omitted.
The beauty of this is that you can just swap out the platform and everything remains as it's the same backend. You make a really good point, one that I haven't really considered, but I also only have so many hours in the day to be fiddling around with stuff.
I think I'd use Discord instead. Clankers are happy to set it all up for you.
I like the in-depth description. Everything from the naming convention of the models (and how much RAM they require) as well as all the components needed underscores just how complicated this all still is. I suppose I am waiting for AI-in-a-Box to come along so I can (painlessly) join in. (I'm sure wrangling with all these esoteric aspects of LLMs though is fun for some people.)
I've seen https://www.lucebox.com/ as an interesting option
Earlier quoted context omitted.
Keep in mind these downloadable models use 3-10x the amount of tokens as well. You really can’t beat a couple $20 subscriptions. https://quesma.com/benchmarks/babaisbench/
The price is handing over your data, and your intellectual property.
Which already comes from Claude itself. Clearly, they don't want to train on their own product.
Earlier quoted context omitted.
There are economies of scale but there’s also a data center bubble (probably) so there might be some selling dollars for fifty cents going on.
Here's the thing that's a little different about data centers; we can tell from Anthropic and OpenAI that they're capacity constrained. Inference demand is there. I notice Cerebras doesn't offer much directly any more, all their capacity is getting completely sucked up by B2B sales. Grok did overbuild, but Anthropic was so desperate for more compute they ate their pride and leased the excess capacity. That means all…
Here's what you have to believe:
- AI demand is at least several times larger than what can currently be satisfied, or will grow. (This one I can buy, but...)
- AI chips (GPUs, TPUs, compute-in-memory, whatever else is being studied) will not get significantly more efficient than they are now. It will not be possible in, say, 5-10 years, to do 2X or 4X or 10X more AI requests per rack than is possible now. I think this one's the single most likely thing to be false, since all computing history contradicts it.
- Edge devices (PCs, laptops, specialized but smaller scale AI compute nodes) will never be powerful enough to run frontier models at a reasonable price that's appealing for professionals, enthusiasts, or businesses, and there will never be a market for this. None of the demand will be served on-device or near-edge. AI must all go in giant data centers.
- AI models will not become significantly more efficient than they are now. There are no large gains on the table from better model architectures, better training, more efficient quantizations, better harnesses, etc.
If all those things are true, than the current planned like 4X-10X increase in data center capacity makes sense. If even one or two of them are not true, then the planned data center build-outs start looking excessive. If all four are not true, it's a total bubble that will crash and burn. Answer is probably somewhere between, but how far toward bubble? That's why I picked a number like "only 20% ever gets built." It might be as high as 50%. It ain't gonna be 100%. The planned built-out is batty.
Oh I forgot two more...
- Data center capacity currently serving non-AI work loads does not shrink through either reduced demand, more efficient software, or (most likely) faster chips and denser RAM. If that happens, more pre-existing DC space can serve AI work loads.
- Orbital solar powered compute nodes never happen. If this happens (free power! much less political opposition!) then terrestrial data centers have significant competition.