`how to build a birdhouse to attract a bluebird` https://ithy.com/article/4b116d2032e54c03862db84e71bcfc8f https://big-agi.com/ has this "BEAM" concept as well where you can put your message through as many models as you have configured then run fuse/guided/compare/custom to merge them all together into one comprehensive. more expensive response. https://files.catbox.moe/tr82vs.png https://files.catbox.moe/beuyfx.png…
Thanks! Yeah that's an excellent idea - this is my response from another thread: I have a feeling that Perplexity and ChatGPT are doing something similar [caching], since common questions I'd ask like "top movies this year" will be answered nearly-instantaneously, way faster than GPT-4o could have done on its own. The only explanation for this is that so many users ask certain questions, they cache the response and r…
Show HN: I made the slowest, most expensive GPT
51–60 of 61 posts
Re: Show HN: I made the slowest, most expensive GPT
#52Update 3:00 PM ET: I've finished scaling up from 2 VPCs to 5 VPCs. Limits have been increased back up to 3 anonymous / 10 signed-in. Update 2:30 PM ET: Back up (for now). Still waiting for Anthropic and Gemini quota increase requests, so those have been migrated to GPT-4o for now. Running on 2 VPCs, in the process of launching 2 more. Confident that I can increase the daily limits by EOD once everything's more stable…
thank you. just seeing this and playing with this has expanded how i think about these type of systems
Re: Show HN: I made the slowest, most expensive GPT
#53You could at least cheapen some of your queries by moving to something like Groq.
Good idea, maybe I'll add Groq as another option, since I don't have an internal Llama 3.1 flow yet. But I'll still need to keep the others to maintain the diversity of responses.
Re: Show HN: I made the slowest, most expensive GPT
#54Interesting idea, cool concept. I tried asking "What is the best SNES game most people haven't played". The top answer (Terranigma) was unfortunately the same as I got just asking any of Claude/ChatGPT/Llama/Qwen (maybe too easy a question) but the rest of the list did seem a bit more balanced. Thanks for the free try without a login! Thought: there is a marquee of example queries but it doesn't seem like there is a…
That's a good idea! I have a feeling that Perplexity and ChatGPT are doing something similar, since common questions I'd ask like "top movies this year" will be answered nearly-instantaneously, way faster than GPT-4o could have done on its own. The only explanation for this is that so many users ask certain questions, they cache the response and return the cached answer. I'd love to do this for Ithy, but it'll be a w…
Re: Show HN: I made the slowest, most expensive GPT
#55Update 3:00 PM ET: I've finished scaling up from 2 VPCs to 5 VPCs. Limits have been increased back up to 3 anonymous / 10 signed-in. Update 2:30 PM ET: Back up (for now). Still waiting for Anthropic and Gemini quota increase requests, so those have been migrated to GPT-4o for now. Running on 2 VPCs, in the process of launching 2 more. Confident that I can increase the daily limits by EOD once everything's more stable…
to be fair to you tho, it's a lot easier to be almost-infinite-scale-ultimate-uptime SRE when you have an almost-infinite-scale-company CC/bank draft supporting you instead of your personal dalla is staked on a really cool PoC. thank you. just seeing this and playing with this has expanded how i think about these type of systems
Re: Show HN: I made the slowest, most expensive GPT
#56so now you summarize everyone’s hallucinations
Re: Show HN: I made the slowest, most expensive GPT
#57Re: Show HN: I made the slowest, most expensive GPT
#58Re: Show HN: I made the slowest, most expensive GPT
#59Earlier quoted context omitted.
this is not ideal. If someone wants to fact check controversial claim filtering it only makes it worse.
This is just a dude's hobby project, chill.