Live data from Hacker News

GPT-5.6 Sol Pricing Cut by 50% on OpenRouter

openrouter.ai

471–479 of 479 posts

Re: GPT-5.6 Sol Pricing Cut by 50% on OpenRouter

#471

Earlier quoted context omitted.

> All flagship models are within like 1-5% of each other Don't know about that. I'm using code review of my lone lisp project as a benchmark. It's a massive parallel code review where a coordinator cuts up the codebase into sections and dispatches agents to consider each part from different perspectives like quality, maintainability, consistency, correctness, rigor, etc. Ran a complete Fable/max code review. Took ove…

Check the remainder for hallucination for sure.

Absolutely. I have an independent audit pass to verify claims.

My methodology consists of launching a 242 cell parallel code review matrix and committing all Fable/Sol max effort agent prompts and their full reports to a private orphan branch on the repository. This is the part that is taking me months to complete. This thing can kill my $100 subscription in about 12 hours.

When done, these raw findings will be semantically deduplicated and merged into a list of findings per model. This list will then be audited for hallucinated or otherwise made up findings. This will refine the list, and hallucination rate is its own data point. I'm also counting things like cybersecurity refusals and downgrades.

When all this is done, I'll analyse the final results and publish them on my website.

Some preliminary analysis:

Which code review lenses were the most valuable, where value is defined as number of serious issues identified? Rigor, followed by tests, robustness, correctness, and so on. I was able to create a tier list of reviewer personas using evidence! I can now run focused code reviews using the highest value lenses.

What's the most expensive code review? Correctness and rigor, of the lone lisp machine specifically.

API costs per finding? $0.91 to $3.77. API costs per serious finding? $8.94 to $28.51. All Fable.

How long did it take? 28.2 calendar days, 66.3 agent-hours.

Is it worth it to run the code review multiple times? If a review matrix's defect capture probability is 57%, then a second run captures 81% of the estimated/extrapolated defect population, a third run captures 92%, a fourth run captures 97%, and further runs yield severely diminishing returns. Probably worth it to code review important stuff three times.

What's the impact and cost of the safety classifier? Out of the 52 Fable review cells that triggered the safety classifier, 35 died without producing any output whatsoever, so 67.3% of the cells were a complete waste of tokens. 25% produced at least some output.

Does the safety classifier trigger most often on the important code that actually needs SOTA models? For the most part, yes. Fable was most often barred from reviewing the most important and complex files in the codebase, such as the virtual machine, the parser and I/O layer. These files also have the most CRITICAL+HIGH severity findings. Only a couple outliers broke this pattern.

Re: GPT-5.6 Sol Pricing Cut by 50% on OpenRouter

#472
post #22

Earlier quoted context omitted.

I've found it to be great for planning code changes (or new projects). I use the superpowers plug-in which I think guides the planning. Then I switch models (to luna) before implementation. I find this combo nearly always does what I want. I also use a skill called ponytail, its goal is to keep things terse and edits small. It may have contributed to the successes above. I like that skills are easy to try out, too.

I stopped using superpowers because it wanted to turn every tiny bug fix into a $37MM DOD project. I got effective results but it took ages. I may try again - I need to find a good way to run different profiles in my harness so I can easily shut it off. The default planning workflow in OMP is pretty good though. I agree Luna is great for task execution, either as a sub-agent with Sol planning and coordinating or if t…

I tell it not to use superpowers for small changes, for the same reason you say. (I often forget, though)

And yes, I sub in other coding models like DeepSeek. Mostly with good results.

Re: GPT-5.6 Sol Pricing Cut by 50% on OpenRouter

#474
post #353

Earlier quoted context omitted.

Is there actually a difference in cache rates between OR and official API? I have a preset set up on OR so that I only send traffic to deepseek. The preset is important otherwise you will send traffic to different providers but if you weren't doing this already then what can I say, water is wet, of course cache rates will be awful. I get about 70% cache hit rate with the preset which is appropriate for what I'm doing…

Zenmux say the cache hit rate is 98% for the deepseek flash API. I don't know why, but performance is definitely worse using openrouter. https://zenmux.ai/deepseek/deepseek-v4-flash

Could be compression or metadata in their pipeline adding noise and making requests less generic.

Re: GPT-5.6 Sol Pricing Cut by 50% on OpenRouter

#475

Earlier quoted context omitted.

I would in such a scenario expect the GPUs to be dumped to industrial breakers who would send them to China for refurbishment and repackaging before being sold again on Amazon, AliExpress, and Taobao as last gen gaming cards from weird brands and specs. This is what happened after the great crypto GPU dumping.

The e-waste recyclers are pretty low on the pecking order, as the creditors will be first to strip these places for assets as Leopold Aschenbrenner discovered. =3

It's where the stuff ultimately goes since the creditors aren't interested in GPUs that no longer have much value. The context here is a crash, not just a basic bankruptcy. If it were that then absolutely it would be impounded by creditors (virtually or in reality) and then sold to the highest bidding data center.

Re: GPT-5.6 Sol Pricing Cut by 50% on OpenRouter

#476
post #323

Earlier quoted context omitted.

I shifted from DeepSeek v4 Flash 0731 to Gemini 3.7 flash on openrouter and price shoots up almost double with no visible change in outcome. So, today I reverted back.

Try DS through their own API if that's feasible, AFAIK they're much cheaper than through OS due to cache hit rates.

they will train / retain your prompts though, so to me it's not a valid option

Re: GPT-5.6 Sol Pricing Cut by 50% on OpenRouter

#477

Earlier quoted context omitted.

The e-waste recyclers are pretty low on the pecking order, as the creditors will be first to strip these places for assets as Leopold Aschenbrenner discovered. =3

It's where the stuff ultimately goes since the creditors aren't interested in GPUs that no longer have much value. The context here is a crash, not just a basic bankruptcy. If it were that then absolutely it would be impounded by creditors (virtually or in reality) and then sold to the highest bidding data center.

>creditors aren't interested in GPUs that no longer have much value

Indeed, there is a point where the cost of disposal is higher than the expected market value of components. Thus, the asset turns into a liability if held too long.

I have seen factory liquidations, and everything goes... right down to the bolts in the floors. =3

Re: GPT-5.6 Sol Pricing Cut by 50% on OpenRouter

#479

OpenAI is making some really boneheaded moves these days. They're reacting instead of leading, basically. Cutting API prices 50% while millions of your paying subscribers have had their limits slashed and are all literally looking at the salivatingly-cheap chinese API prices availalbe on openrouter... Not only did OpenAI and all of their cash somehow MISS the opportunity to purchase OpenRouter ... Now they're giving…

[flagged]
Post reply on HN