Earlier quoted context omitted.
Yeah I have found worse results if I don't leave it on the highest setting. I have gotten by with Pro and a little overage buffer so far. I have found it working pretty well for what I'm using it for but I have really only been using it a couple months now.
A few weeks ago I went from 'high' to 'max' effort and got decent results. Since 4.7 went live it has regressed to how it was before (with the same settings).
I am worried about Bun
351–360 of 368 posts
Re: I am worried about Bun
#352Earlier quoted context omitted.
Node's built-in profiler doesn't work with Typescript, which is one part of Node not natively supporting TS. Idk how it is in Bun, cause that made me abandon TS rather than abandoning Node.
Node supports ts natively now though.
And looking ad docs, it seems it only has partial support still: https://nodejs.org/api/typescript.html
> To use TypeScript with full support for all TypeScript features, including tsconfig.json, you can use a third-party package. These instructions use tsx as an example but there are many other similar libraries available.
Re: I am worried about Bun
#353Don't fret; the creator of mise has released a faster alternative: https://github.com/endevco/aube
Odd that aube is missing deno from their benchmarks though
Re: I am worried about Bun
#354Earlier quoted context omitted.
> I disagree with the overall premise: Before the acquisition, Bun had to figure out how to monetize at some point. Incidentally, Anthropic needs to figure out how to monetize at some point too.
It’s organizations figuring out how to monetize all the way up.
Re: I am worried about Bun
#355Earlier quoted context omitted.
It's interesting how quickly people buy the "abuse" line of thinking. We understood (and knew for a long time) that the large AI labs are not monetarily profiting from subscription users that make heavy use of their subscription. That is independent of which agent/harness is used. The fair/real price for profitable use is the pay per use token pricing. These labs play the game of trying to kill competition in the har…
> Useful models are getting smaller and cheaper to run every year and it has hit a threshold at which we will see continued development of third party harnesses even without the userbase of subscription users. As of May 2026, how much money do I need to spend to buy hardware to have a local model that is 80% as good as SOTA services for assisting me in writing code? As for that 80%, how many minutes per LOC will I be…
https://llm-stats.com/benchmarks/swe-bench-verified
SOTA (public proprietary models) would be Opus 4.7 at 0.876
80% of that would be around 0.7.
These models qualify, and are upwards of 90% as good in benchmarks:
DeepSeek-V4-Pro-Max - 1.6T (HuggingFace shows 862B, huh) - 0.806
Kimi K2.6 - 1.1T - 0.802
MiniMax M2.5 - 229B - 0.802
DeepSeek-V4-Flash-Max - 284B (HuggingFace shows 158B as well) - 0.790
These are 80-90% as good, which is also where you see the smaller ones: GLM-5 - 754B - 0.778
Qwen3.6-27B - 27B - 0.772
Kimi K2.5 - 1.1T - 0.768
Qwen3.5-397B-A17B - 397B - 0.764
Step-3.5-Flash - 199B - 0.744
GLM-4.7 - 358B - 0.738
MiMo-V2-Flash - 310B - 0.734
Qwen3.6-35B-A3B - 35B - 0.734
DeepSeek-V3.2 - 685B - 0.731
DeepSeek-V3.2-Speciale - 685B - 0.731
DeepSeek-V3.2 (Thinking) - 685B - 0.731
Qwen3.5-27B - 27B - 0.724
Qwen3.5-122B-A10B - 125B - 0.720
Kimi K2-Thinking-0905 - 1T - 0.713
LongCat-Flash-Thinking-2601 - 562B - 0.700
Out of those, the most modest one you could get is Qwen3.6-35B-A3B because the MoE nature makes it faster across more varied hardware.I currently run the Unsloth 8bit quants on-prem (on a bunch of Nvidia L4 GPUs, since low TDP, long story), some people swear by more quantized versions but with the small models the impact is felt more: https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF
So essentially you need up to 39 GB for the model itself and then some for the KV cache and whatever context size you want. Ideally I'd aim for 64 GB of memory for that, though if really pressed for resources, could get a heavily quantized version within 32 GB (but very little memory for context and kinda shit).
Personally, I think that you need about 45-60 tokens/second for decent usability - even comparatively modest hardware (including those L4) can run the model, though on the lower end options you will not be running parallel sub-agents etc.
Some random results for when you don't want a traditional multi-GPU setup:
Mac Mini - about 1999 USD, gets you somewhere upwards of 30 tokens/second (depends on quantization and how you run it)
Framework Desktop - about 2500 USD, gets you somewhere upwards of 25 tokens/second https://community.frame.work/t/framework-desktop-for-local-ai/80880/5
DGX Spark - about 3500 USD, gets you somewhere upwards of 50 tokens/second https://forums.developer.nvidia.com/t/qwen-qwen3-6-35b-a3b-and-fp8-has-landed/366822/27
Some random results from pulling up random shops and approx. benchmarks, for dual GPU setups (not necessarily NVLink etc.): 2x Intel Arc Pro B70 - about 1900 USD, gets you around 36 tokens/second, borderline usable, I blame their software stack
2x Radeon AI PRO R9700 - about 3000 USD, gets you somewhere upwards of 60 tokens/second, usable
2x Radeon PRO W7800 - about 5400 USD, same as above
2x NVIDIA RTX 5090 - about 7600 USD, same as above
2x NVIDIA RTX 5000 Ada - about 9200 USD, same as above
Of course, for those models, some of those cards are way overkill, but you definitely can get something for running local models without too many compromises involved. That said, you definitely will get a worse experience than SOTA cloud models at that 80% and will have to rework stuff quite a bit often, as my own experience with the Qwen model shows - okay for simple tasks, breaks down on complex stuff. For that, you'd want at least some of the 90% category models and would probably need to consider how much memory you can realistically get.At least it's not hopeless!
Re: I am worried about Bun
#356Earlier quoted context omitted.
I think the vagueness of statements like this is why a lot of people (myself included) are just so very skeptical. Surely some company wants to brag about their use. I don’t doubt it’s found its way into certain spaces, but by and large a lot of the “big” claims have been demonstrated to be borderline fraudulent. That Brad Pitt/Tom Cruise AI fight is fake. It is misleading. Taking existing green screen choreography a…
I can respond directly to this, I’m a former VFX industry person and still fairly well connected. The the former you suggested. Background plates and the like. The lack of actual creative direction tools, trite visual style, lack of consistency/repeatability and complete inability to be edited or adjusted easily make it a non-starter for most tasks. Compositors are fast, LLMs are slow at that scale. There are tools l…
They're using AI for plates, edits, pickup shots, previz, and in some cases the primary footage itself.
They're super hush-hush about this.
Re: I am worried about Bun
#357Earlier quoted context omitted.
Node supports ts natively now though.
even enums? last I tested, it did type-stripping only. And looking ad docs, it seems it only has partial support still: https://nodejs.org/api/typescript.html > To use TypeScript with full support for all TypeScript features, including tsconfig.json, you can use a third-party package. These instructions use tsx as an example but there are many other similar libraries available.
Re: I am worried about Bun
#358Earlier quoted context omitted.
oh absolutely, no argument there, the case for AGI is pretty weak. I was just saying that I am even more sceptical that any of this is a "first or nothing" scenario - that is one of my biggest pet peeves about the entire tech sector.
Right, but I never said it was a first-or-nothing scenario to begin with. Given that both AGI and ASI are so ambiguous as to be nothingburgers, talking about them is just a performative thought experiment IMO. An interesting one, certainly, but neither are even remotely close to being realized. Until we have some kind of clear definition that can be scientifically proven and reproduced, that will remain the case.
Re: I am worried about Bun
#359Earlier quoted context omitted.
The source code of Claude Code and Gemini CLI contradict that.
Well sure, I can find you hundreds of dead ends in mutations, but when you start parsing through what a harness does, eventually it'll just be another more deterministic and constrained model to do less creative things.
Re: I am worried about Bun
#360Earlier quoted context omitted.
I did for over a decade, but it does not go far enough with supply chain security. I bootstrapped a new generation of Linux distribution from 180 bytes of human readable x86 machine code all the way up. https://stagex.tools
You should probably caveat any post you make about security concerns with that, so people can more easily judge whether your concerns line up with their threat model.
The entire medical industry was negligent for 100 years following Ignaz Semmelweis proving basic sanitation tactics would save countless lives.
Similarly the entire software industry is and has been negligent since 1987 when Ken Thompson first demonstrated basic supply chain integrity tactics could stop otherwise unstoppable and undetectable attacks on software.