Live data from Hacker News

Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models

news.ycombinator.com

111–120 of 254 posts

Re: Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models

#112

Earlier quoted context omitted.

Enough to make it a non-started at the organizations that pay their bills. Everyone else isn't big enough to matter.

Going to be an interesting world where big enterprises have to spend 100X the cost for the same value of AI as startups and small businesses.

I think they call that inflation :).

Re: Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models

#113

but u wouldnt get caching savings

You definitely still would, but you need to pay the full input cost twice, so the equation really depends on how much first message vs repeat message matter.

Ttl 1 hour maybe. 5min? Never

Re: Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models

#114
thanks to everyone for taking the time to try Echo and share feedback, this is precisely why i wanted to launch early.

i am going to try to address a couple of topics that came up often:

- i'll keep publishing stronger evals, including more difficult coding and agentic benchmarks, to map out more precisely the differences with sota

- the public eval dashboard will keep expanding and be updated (very open to more benchmark suggestions as well!)

- some people found issues in the eval dashboard ui and the sign up flow, should be now all fixed in prod

some important precisions as well:

- NO credit card is required to try Echo

- each acount includes 10$ of free credits to try on both the API and the chat

on the approach itself: the idea i'm exploring is more broader than model routing, i'm looking at how to allocate inference efficiently across open-weight models, deciding not only which models to use, but also how much computation a request deserves and how intermediate work should be combined.

ensembling by itself is not new. since random forests and probably even before in statistics/classic ml we knew that bringing multiple models together can outperform individual ones. the interesting problem for Echo is how to model and leverage this without paying the full ensemble cost at each request.

while there are conceptual similarities with systems like Fusion or Fugu, the architecture and optimization objective are different.

thanks again for all the thoughtful feedback.

Re: Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models

#115

By the way, is it just me or opus 4.8 is much better at some programming tasks than fable? I've been really disappointed lately, i stopped using it evem though it's "premium on my subscription"

It's just you. Fable 5 is consistently better than Opus 4.8 at literally every single level, for me. (I always use both at xhigh, for reference.) I could go down a laundry list of various issues I have with Opus that I don't have with Fable. For me, Opus 4.8 Plus Fable is way less annoying to talk to than Opus 4.8. Opus 4.8's writing style is absolutely insufferable. Fable has some of the same quirks but it's way les…

We see no difference between opus 4.8 and fable except that fable is slower. So +1 for not just GP. We use neither interactive, just via our own tooling so the ‘talk to’ doesn’t apply and we use it 24/7. We currently run 25% of tasks on both fable and opus and the rest only opus. The 25% are being code reviewed side by side and we do detect when fable switches to opus for ‘security concerns’ nonsense. Opus generally finishes sooner, results are similar quality.

Re: Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models

#116

Earlier quoted context omitted.

This is what OpenAI and Anthropic are trying to make everyone believe. Most accountants will flinch at this (they already are). The $200 odd plans are already out of reach of many, many people. The attrition of customers if they were to get rid of these subscriptions plans would be untenable.

I think you are looking at it incorrectly. No business is buying individual accounts, because if they do, they open themselves up to considerable risk. The $200 plans are priced so that the power-users use them and then advocate about how great the product is. If you're buying a $200 plan, you're not doing it because of the price point but rather because of the amount of work it is doing for you.

All companies we interact with have 200 plans.

Re: Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models

#119
I'm curious if there is measurable value in diversity of thought, and if there's diminishing returns on a single models thought pattern.

For example, compute X tokens with model A, then feed those into model B, etc. to get chain of thought through a diverse set of mdoels rather than chain of thought through a heterogeneous chain.

Humans seem to strongly believe echo chambers are bad. Are LLMs the same?

Re: Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models

#120

"Backed by YCombinator" https://www.ycombinator.com/companies?query=tracerml I don't see it?

You're right to be skeptical; I saw a totally fake case of this just yesterday.

But in the present case, they're just a startup in the current batch.

Post reply on HN