Live data from Hacker News

I want everything local – Building my offline AI workspace

instavm.io

121–130 of 294 posts

Re: I want everything local – Building my offline AI workspace

#121
I'm a little confused about your product branding vs. blog post?

From the product homepage, I imagine you're running VMs in the cloud (a la Firecracker).

From the blog post though, it looks like you're running Apple-specific VMs for local execution?

As someone who's built the former, I'd love the latter for use with the new gpt-oss releases :)

Re: I want everything local – Building my offline AI workspace

#123
post #11

I'm constantly tempted by the idealism of this experience, but when you factor in the performance of the models you have access to, and the cost of running them on-demand in a cloud, it's really just a fun hobby instead of a viable strategy to benefit your life. As the hardware continues to iterate at a rapid pace, anything you pick up second-hand will still deprecate at that pace, making any real investment in hardw…

It's not that bad. If you're an adult making a living wage, and you're literate in some IT principles and AGI operations know-how, it's not a major onetime investment. And you can always learn. I'm sure your argument deterred a lot of your parents' generation from buying computers, too. Where would most of us be if not for that? This is a second transistor moment, right in our lifetime.

Life is about balance. If you Boglehead everything and then die before retirement, did you really live?

Re: I want everything local – Building my offline AI workspace

#124

Its all about context and purpose, isn't it? For certain lightweight uses cases, especially those concerning sensitive user data, a local implementation may make a lot of sense.

My thoughts exactly. The recent GPT-OSS 20B parameter model was a nice upgrade, it really feels like having a local mini ChatGPT.

Re: I want everything local – Building my offline AI workspace

#125
To be honest, I just want to make porn. My own porn, the way I want it. That’s what I’m waiting for. Why the heck do I need to scroll through pages of boring, vanilla, pedestrian porn on Pornhub or RedGIfs or XNXX when I can create exactly what I want? That’ll be a huge killer app when I can do it locally and in the privacy of my own home.

Re: I want everything local – Building my offline AI workspace

#126
post #115

Earlier quoted context omitted.

Why would AI be one of the few areas where locally-hosted options can't reach "good enough"?

Maybe a better question is when will SOTA models be "good enough"? At the moment there appears to be ~no demand for older models, even models that people praised just a few months ago. I suspect until AGI/ASI is reached or progress plateaus, that will continue be the case.

Yes, this is exactly my point. Thank you for stating it better.

Re: I want everything local – Building my offline AI workspace

#127
post #26

if you ever end up trying to take this in the mobile direction, consider running on-device AI with Cactus – https://cactuscompute.com/ Blazing-fast, cross-platform, and supports nearly all recent OS models.

Is this your site? It's missing a tag.

Re: I want everything local – Building my offline AI workspace

#128
post #95

Earlier quoted context omitted.

I agree and disagree. Many of the best models are open source, just too big to run for most people. And there are plenty of ways to fit these models! A Mac Studio M3 Ultra with 512 GB unified memory though has huge capacity, and a decent chunk of bandwidth (800GB/s. Compare vs a 5090's ~1800GB/s). $10k is a lot of money, but that ability to fit these very large models & get quality results is very impressive. Perform…

https://pcisig.com/pci-sig-announces-pcie-80-specification-t... From 2003-2016, 13 years, we had PCIE 1,2,3. 2017 - PCIE 4.0 2019 - PCIE 5.0 2022 - PCIE 6.0 2025 - PCIE 7.0 2028 - PCIE 8.0 Manufacturing and vendors are having a hard time keeping up. And the PCIE 5.0 memory is.. not always the most stable.

Thanks for the numbers. Valuable contribution for sure!!

There's been a huge lag for PCIe adoption, and imo so so much has boiled down "do people need it"?

In the past 10 years I feel like my eyes have been opened that every high tech company's greatest highest most compelling desire is to slow walk the release out. To move as slow as the market will bear, to do as little as possible, to roll on and on with minor incremental changes.

There are canonball moments where the market is disrupted. Thank the fucking stars Intel got sick of all this shit and worked hard (with many others) to standardized NVMe, to make a post SATA world with higher speeds & better protocol. AMD64 architecture changed the game. Ryzen again. But so much of the industry is about retaining your cost advantage, is about retaining strong market segmentations, by never shipping too many PCIe lane platforms, by limiting consumer vs workstation vs server video card ram and vgpu (and mxgpu) and display out capabilities often entirely artificially.

But there is a fucking fire right now and everyone knows it. Nvlink is massively more bandwidth and massively more efficient and is essential to system performance. The need to get better fast is so on. Seems like for now SSD will keep slow walking their 2x's. But PCIe is facing a real crisis of being replaced, and everyone wants better. And hates hates hates the insane cost. PCIe 8.0 is going to be insane data to push over a differential, insane speed. But we have to.

Alas PCIe is also hampered by relatively generous broader system design. The trace distances are going to shrink, signal requirements increase a lot. But this needing a intercompatible compliance program for any peripheral to work is a significant disadvantage, versus, just make this point to point link work between these two cards.

There's so many energies happening right now in interconnect. I hope we see some actual uptake, some day. We've had so long for Gen-Z (Ethernet phy, gone now), CXL (3.x being switched, still un-arriced), now UltraEthernet and UltraLink. Man I hope we can see some step improvements. Everyone knows we are in deep shit if NV alone can connect systems. Ironically AMD's HyperTransport was open, was a path towards this, but now Infinity Fabric is an internal only thing and as branding & an idea vanishing from the world kind of, feels insufficient.

Re: I want everything local – Building my offline AI workspace

#129
post #108
post #105

Earlier quoted context omitted.

Are you conflating GDDR5x with PCIe 5.0?

No. I'm saying we're due for faster memory but seem to be having trouble scaling bus speeds as well (in production) and reliable memory. And the network is changing a lot, too. It's a neverending cycle I guess.

One advantage of Apple Silicon is the unified memory architecture. You put memory on the fabric instead of on PCIe.
Post reply on HN