Live data from Hacker News

Today's Frontier AI companies will never exceed the AI capability frontier again

andrewtrask.substack.com

1–10 of 11 posts

Re: Today's Frontier AI companies will never exceed the AI capability frontier again

#4

I’m so confused. The top fusion is Fable 5 and GPT 5.5. That is not an “ensemble of weaker AI models.”

i guess the point is that any fusion is better than any single model and a fusion of the top two models is obviously the best? for cost though i guess you could just duct tape together 10 open source models and then thats comparable?

Re: Today's Frontier AI companies will never exceed the AI capability frontier again

#5

I’m so confused. The top fusion is Fable 5 and GPT 5.5. That is not an “ensemble of weaker AI models.”

What its saying is if you look at any single model, it can be beaten by an ensemble of weaker models. E.g fable 5 is beaten by an ensemble of previous gen models.

Re: Today's Frontier AI companies will never exceed the AI capability frontier again

#6
post #4

I’m so confused. The top fusion is Fable 5 and GPT 5.5. That is not an “ensemble of weaker AI models.”

i guess the point is that any fusion is better than any single model and a fusion of the top two models is obviously the best? for cost though i guess you could just duct tape together 10 open source models and then thats comparable?

> though i guess you could just duct tape together 10 open source models and then thats comparable?

This is what I was hoping to see data for.

Re: Today's Frontier AI companies will never exceed the AI capability frontier again

#7

I’m so confused. The top fusion is Fable 5 and GPT 5.5. That is not an “ensemble of weaker AI models.”

What its saying is if you look at any single model, it can be beaten by an ensemble of weaker models. E.g fable 5 is beaten by an ensemble of previous gen models.

I guess so. 4.8 + 4.8 > Fable 5 is interesting, though not particularly game changing. (The others all fuse frontier models. Which is an argument for using those frontier models more. Not less.)

Re: Today's Frontier AI companies will never exceed the AI capability frontier again

#8
post #4

Earlier quoted context omitted.

i guess the point is that any fusion is better than any single model and a fusion of the top two models is obviously the best? for cost though i guess you could just duct tape together 10 open source models and then thats comparable?

> though i guess you could just duct tape together 10 open source models and then thats comparable? This is what I was hoping to see data for.

[deleted]

Re: Today's Frontier AI companies will never exceed the AI capability frontier again

#9
A manic riff on https://xcancel.com/OpenRouter/status/2065856853989270011 , which advertises https://openrouter.ai/fusion/1 , which is a (slow) multi-model multi-prompt workflow that's specific to the "DRACO" benchmark for "deep research", and doesn't say much about coding and long-horizon agentic work, nor does it imply you can somehow parlay this into duct-taping 50 budget-tier models together for even more gains. Not even sure what "solo" even means in the context of the comparison chart - oneshot? Variant workflow since it doesn't make sense to run on one input?

Mixing outputs of different models one way or another is old news, if it were anywhere near as promising as the author dreams it would have exploded many months ago.

Re: Today's Frontier AI companies will never exceed the AI capability frontier again

#10

Earlier quoted context omitted.

What its saying is if you look at any single model, it can be beaten by an ensemble of weaker models. E.g fable 5 is beaten by an ensemble of previous gen models.

I guess so. 4.8 + 4.8 > Fable 5 is interesting, though not particularly game changing. (The others all fuse frontier models. Which is an argument for using those frontier models more. Not less.)

Yeah, all that's really saying is a weaker model with a better harness can beat a stronger model with a worse harness, specifically on the DRACO benchmark

This isn't really a surprising result. Needs more evidence to make a broader claim.

Post reply on HN