Live data from Hacker News

Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models

news.ycombinator.com

171–180 of 254 posts

Re: Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models

#172

Earlier quoted context omitted.

I burned though my weekly fable usage last night on the $200 plan. I had $200 in promotional usage credits and was in the middle of executing a moderate sized coding plan. Ran on usage credits for about 1h 15m and burned $120 in usage credits. I was astounded to see how fast the $ usage added up. One problem was that I was using sub-agent execution so multiple agents were running simultaneously and I realized at the…

I am too young (most of us on here are) to have lived through the paying for time on time-share machines in the 60s/70s, but this is giving me creepy memories of paying for sprintnet/telenet and tymnet... And I guess aol, compuserv, delphi. Are we really doing this computing model again?

If I remember well, visual studio/ msdn used to be like 5k-10k per year... Basically any IT tool cost a shit load of money.

I have in mind 50k-100k ish for 3d studio max or was it softimage? (Well seems softimage https://www.awn.com/animationworld/siggraph-news-announcing-... ).

So... Basically we are back to this era.

Re: Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models

#173

thanks to everyone for taking the time to try Echo and share feedback, this is precisely why i wanted to launch early. i am going to try to address a couple of topics that came up often: - i'll keep publishing stronger evals, including more difficult coding and agentic benchmarks, to map out more precisely the differences with sota - the public eval dashboard will keep expanding and be updated (very open to more benc…

Small feedback: the "create password" requires a symbol too, which Google's password manager by default does not use. I'm fairly sure a double-digit-level alphanumeric jumble is sufficient to be a password (or at least, Google thinks so). Great idea nonetheless!

> I'm fairly sure a double-digit-level alphanumeric jumble is sufficient to be a password

This isn't required either. Of course there's an xlcd for that: https://xkcd.com/936/

Besides,

--- start quote ---

Using complexity requirements (that is, where staff can only use passwords that are suitably complex) is a poor defence against guessing attacks. It places an extra burden on users, many of whom will use predictable patterns (such as replacing the letter ‘o’ with a zero) to meet the required 'complexity' criteria.

https://www.ncsc.gov.uk/collection/passwords/updating-your-a...

--- end quote ---

Re: Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models

#174
post #155

Earlier quoted context omitted.

if this is true, the frontier labs are not able to justify their trillion dollar valuations, they are barely making anything on subsidized plans.

You can bet that most people on those plans do not tokenmax.

[deleted]

Re: Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models

#175
post #152

Earlier quoted context omitted.

> Sure, these plans may be temporary, but none of us really know how temporary they are. Anthropic emailed me today: Fable 5 moved to usage credits on July 20. It is still available to you, but it requires pay-as-you-go usage credits and is not included in your subscription rate limits.

That's going to be on the $20 plan. IIRC the $100 and $200 plans keep the 50% fable usage.

We got emails and UI messages informing us that Fable is now a standard part of Max plans (both sizes).

The half usage limit still applies.

Re: Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models

#176

I wouldn't be surprised if "the best model" becomes a niche concept over the next few years. For most production systems the winning architecture may end up being an orchestrator that knows when to call a cheap model when to escalate to a stronger one and when to combine multiple outputs.

there are so many branches of evolution here, I also see them converging to a "best model becomes a niche concept" too

the supercycle is on device models, and one of those evolutions is models baked into chip die, and you just upgrade chipsets every few years instead

so it's the hyperscalers that will take the L in that environment

Re: Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models

#177
The business use case of rather than going on one model choosing the best openweight model and reducing the cost seems fascinating, however the context memory, or auditability of what is happening behind would be more complicated, even within single model we have to spend tons of time to decode and understand llm behavior , add context layers and so on, having said that for GenAI executions this might be the direction. Wish you best for the project

Re: Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models

#178
I started building something like this myself (codename HFG - "Heroku for Groq") anticipating what we are indeed now seeing in ChatGPT viz. enshittification. The intention was to aim it at consumers and non-technical users who want a chatbot to help them draft reports, synthesise papers and so on. You've got a bunch of hackers in the comments going "But how can I run my own evals?" I think the answer for them is: this isn't for you.

Don't want to derail what you're trying to do with Echo in case I'm wide of the mark, but yeah even in that case, if you hadn't considered that use case for it, I reckon there will, probably inside six months, be a substantial market for non-technical users who are sick of seeing ads in a service they already pay a subscription for, and who don't care what the underlying model is - or, indeed, don't even understand the concept of an "underlying model" because they interface with AI as a product.

You have already taken the HFG idea way further than I had even thought of yet, and I feel vindicated in seeing someone else do it. I wish you the very best!

Re: Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models

#179

The business use case of rather than going on one model choosing the best openweight model and reducing the cost seems fascinating, however the context memory, or auditability of what is happening behind would be more complicated, even within single model we have to spend tons of time to decode and understand llm behavior , add context layers and so on, having said that for GenAI executions this might be the directio…

[dead]

Re: Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models

#180

> Fable-level results at 1/3 the cost I am guessing this is not targeting those of us on the heavily subsidized $200/mo plans. Sure, these plans may be temporary, but none of us really know how temporary they are. Until then, 1/3rd of the published API pricing is not very appealing.

I believe they will last until they IPO, and not long after that. $200/mo plans are not good for their P&L when their users using $10000 worth api credits. That's -98% margin loss per user.

No way in hell are the majority of Claude Code users burning 10k worth of credits. Many of them probably barely use it. There'll be a bell curve, and we have no idea what it looks like.
Post reply on HN