Earlier quoted context omitted.
I got an RTX 6000 pro too. I like running locally, I've learned a lot more than if I had used an API and there's less worry about overspending tokens. I accidentally spent $100 on claude api in like 2 days because I didn't know what I was doing. The problem is that while one these gpus is a huge improvement over a laptop or a single 3090, you very quickly wish you had more. I would buy a second one, but I did the mat…
What kind of machine did you build around it ?
Was my $48K GPU server worth it?
451–460 of 480 posts
Re: Was my $48K GPU server worth it?
#452Earlier quoted context omitted.
I also call this "bollocks" there is no way this workflow is even 1/10 of what you can get with Codex/Claude Code. A normal engineer may be running a couple of sessions with every session spawning sub agents left and right. 80 persons or even 10 having this workflow on this setup doesn't work, and this is the standard engineer workflow today.
Subagent swarms are actually great for the local inference scenario because they can share a whole lot of KV cache. You get to raise the compute intensity of decode (i.e. the aggregate tok/s) essentially for free.
Re: Was my $48K GPU server worth it?
#453I administer a simple AI server in the office, which just uses a single RTX 5090 but is able to serve ~80 people throughout the day. I'm impressed by Qwen3.6-27b's capabilities in agentic coding/tasks so far. Devs say it's not much different from Sonnet 4.6 on many tasks (sometimes it even outperformed it), 40-60 tok/sec, up to 260k context. The server cost about $10k with all the bells and whistles. I spent a lot of…
I thought NVLINK didn't matter anymore because of the latest PCI-E speeds. Am I wrong there?
Re: Was my $48K GPU server worth it?
#454Earlier quoted context omitted.
I was just making a correction based on what you said. "AI is cool but it's not going to have all the good and bad experiences that humans have had with different motherboards." AI will have more access to experiences than you'll find here.
It actually won't have "had" any experiences though. Yes, it can aggregate stuff from blog posts and reviews and marketing material. That's hardly the same thing.
Re: Was my $48K GPU server worth it?
#455Earlier quoted context omitted.
Or it could have had way more bang/buck by feeding a family of real brains for a year or two
Excuse me for this comment, really, but I can't comprehend the absurdity, some people are buying GPUs when other people have no money for insulin so they literally die. I don't mean anything towards op or gp, quite the opposite I'm truly happy they have this kind of freedom, it must feel really nice, I just hate this game so much.
Re: Was my $48K GPU server worth it?
#456Earlier quoted context omitted.
SpaceX's has disclosed that they're loosing $2Bln a quarter on A.I - and rising - in their IPO documents. Anthropic told the Department of War-nee-Defence that they'd made $5bln total, which is a lot LOT less than what they're spending. We'll see what's in OpenAi's IPO later this year I guess. I'll be very surprised if they're losing less that $100bln a year.
Is it capex of training new models and hiring people for 250mln pay packages? Or is it opex running inference?
Re: Was my $48K GPU server worth it?
#457Earlier quoted context omitted.
It actually won't have "had" any experiences though. Yes, it can aggregate stuff from blog posts and reviews and marketing material. That's hardly the same thing.
It goes far beyond blogs and marketing posts, but sure. Keep asking forums generic easy questions instead of AI. That way you can get one barely helpful reply, when AI could give you more details about the subject than you'd have time to read or ability to memorize.
I just thought I might get an interesting or unique perspective someone here.
Re: Was my $48K GPU server worth it?
#458Earlier quoted context omitted.
It goes far beyond blogs and marketing posts, but sure. Keep asking forums generic easy questions instead of AI. That way you can get one barely helpful reply, when AI could give you more details about the subject than you'd have time to read or ability to memorize.
I feel like you're trying to paint me as some Luddite, like I'm against using AI, or that I don't know about it. I know how to use AI just fine. I use Claude all the time. I'm not opposed to asking it, and in fact I had even before you made your glorified "Let me Google That For You" comment. I just thought I might get an interesting or unique perspective someone here.
Re: Was my $48K GPU server worth it?
#459I administer a simple AI server in the office, which just uses a single RTX 5090 but is able to serve ~80 people throughout the day. I'm impressed by Qwen3.6-27b's capabilities in agentic coding/tasks so far. Devs say it's not much different from Sonnet 4.6 on many tasks (sometimes it even outperformed it), 40-60 tok/sec, up to 260k context. The server cost about $10k with all the bells and whistles. I spent a lot of…
Does anyone here have experience running large models in a multi-GPU setup with several RTX 6000s in a high-concurrency regime and with large context lengths? (something like Deepseek 4 Flash, Minimax 2.7 etc.) For what it's worth, I've been seeing ~100 tps with 4-bit MiniMax 2.7 on two RTX 6000 boards, just running under llama-server without any optimization effort at all. I have no serious long-context experience w…
Re: Was my $48K GPU server worth it?
#460Earlier quoted context omitted.
Is it capex of training new models and hiring people for 250mln pay packages? Or is it opex running inference?
Salaries are opex