Live data from Hacker News

Project HydraFusion: Frontier quality via multi-model orchestration

github.blog

31–36 of 36 posts

Re: Project HydraFusion: Frontier quality via multi-model orchestration

#32

Earlier quoted context omitted.

I don't quite understand your point. Why are you moving those capabilities from one model to another, or improving the built-in capabilities of a model, what is the goal? If having the capabilities in the model itself improves the overall capability, then using that better model with the same harness should achieve better results. Or if the capabilities are the same, but they've been moved from the harness into the m…

I don't know about others, I can only speak for myself. But I do appreciate numbers for bare models, numbers for model + harness, numbers comparing different models in the same harness, and numbers comparing several models across several harnesses. It's a lot of information to ingest, but it gives me some idea of which part of the system is doing which part of the work, how well different harnesses and models interop…

I agree that all else equal, the more comparisons the merrier. But IMO, there has been way too much focus on benchmarks for the bare models, when it seems to me that what actually matters is what you can do with the technology, and that is never limited to what a bare model can do.

Re: Project HydraFusion: Frontier quality via multi-model orchestration

#33
If you’re a copilot cli user (we exist) you know that the auto mode leaves something to be desired. It only changes models at the start of a session or after compaction. Most dev tasks are now sub agent heavy and there is no model routing on the sub agents.

I do notice that there is no mention of effort levels in the comparisons. 5.6 Luna Max is really good and really cheap. A 5.6 Sol high orchestrator with 5.6 Luna max is cheap, has frontier level performance, and is faster than Sol alone. This can be accomplished with simple agent instructions. Looking forward to running my own benchmarks on HydraFusion to see how it fares. Gone are the days of a single model doing all of the work it seems, unless the work requires no tool calls.

Re: Project HydraFusion: Frontier quality via multi-model orchestration

#34
post #3
post #2

This has got to be the worst project name I've seen all year.

Why? I feel it kinda implies what the approach is. Could definitely be worse, strongly prefer this over just giving the thing a random-ass name like Laguna.

https://en.wikipedia.org/wiki/Hydra_(comics)

> Hydra (sometimes stylized as HYDRA) is a fictional terrorist organization appearing in American comic books published by Marvel Comics.[…] Hydra is taken over and turned into a neo-fascist international crime syndicate by Baron Wolfgang von Strucker.

Re: Project HydraFusion: Frontier quality via multi-model orchestration

#36
post #20

Earlier quoted context omitted.

My editor and gatekeeper use like the same model Opus 5. Different prompts and a kind of different input data. The gatekeeper receives the fact check results next to finished text. In the same time the editor already delivered them. As far as I remember over the entire period he removed 27 posts out of 187 that went through him. So I believe that different manufacturers are not mandatory. What matters I guess is not…

> the fact that the critic has a different input and doesn’t have their own text that needs to be defended. Different inputs are one point, but there is another problem: lack of diversity Models from the same maker, share the same training and the same implicit bias. It is like if both reviewers had the same gender, race, and studied at the same university, and just got different book day before. Add fresh immigrant…

I would answer why was a single model enough for this? The fact check provides a kind of explicit written list of what cannot be asserted. At the same time gatekeeper compares the text with this list. I would say this is not about taste or judgment it’s about comparing two documents. In addition some of the removed items are direct violations of this list like the post called the model "a top‑tier model" even though the fact check by the way explicitly prohibited presenting circulating benchmarks as established facts. As regarding the degradation to Sonnet and Haiku I don't have yet any data.
Post reply on HN