Live data from Hacker News

Thefastest.ai

thefastest.ai

31–36 of 36 posts

Re: Thefastest.ai

#31
I love this. Latency is the worst part about AI. I use the lowest latency models that give adequate answers. I do wish this site gave an average and standard deviation.For example Groq fluctuates wildly, depending of the time of day. They're ranked pretty poorly at "610ms" here, and I definitely encounter far worse from them sometimes, but it's wicked fast at other times.

Re: Thefastest.ai

#32
post #24

I'd be interested to hear how Llama 8B with long chain-of-thought prompts compares to GPT-4 one-shot prompts for real-world tasks. In classification for example, you could ask Llama 8B to reason through each possibility, rank them, rate them, make counterarguments, etc. - all in the same time that GPT-4 would take to output one classification without reasoning. Which does better?

I did that with Llama 3 8B with some stuff i could think of, and it did very good. It was on par with GPT4. I prompted some scenarios and asked it to use CoT. Scenarios like "i was standing and eating chocolate, and it melted. Will i find chocolate at my feet?", and the reasoning was pretty good.

But there was something it did way better than GPT4. I asked to create 10 phrases where the last word was an animal, excluding equines, and in alphabetical order. GPT3.5 and GPT4 aren't able to follow such instructions, but the 8b model did it with maestry.

Re: Thefastest.ai

#33
post #14

Another good resource: https://artificialanalysis.ai/

this page is probably our most comparable to thefastest (which is cool, more benchmarks is better): https://artificialanalysis.ai/leaderboards/providers

We also have pricing, long/medium/short prompt lengths (decode time can vary between providers) & parallel query benchmarking + model details (ctx window, etc)

Re: Thefastest.ai

#34

Groq with llama3 70b is so fast and good enough for what we do (source code stuff) that it’s really quite painful to work with most others now. We replaced most our internal integrations with this and everything is great so far. I guess they will be bought soon?

How is that possible with Groq rate limits? Fireworks has Llama 3 for the same effective speed with much more realistic rate limits (and billing)

Groq relaxes the rate limits if you work 'closer' with them.

Re: Thefastest.ai

#35

Earlier quoted context omitted.

How is that possible with Groq rate limits? Fireworks has Llama 3 for the same effective speed with much more realistic rate limits (and billing)

Groq relaxes the rate limits if you work 'closer' with them.

how do you work closer? we filled out the enterprise form and I DMd a mod in the Discord..

Re: Thefastest.ai

#36

I don't understanding why would we need to having similar expectations from systems that we have from humans and building a whole theory on it. I can adjust my behaviour around systems. I am not restricted to operate within default values. e.g Whenever a price is listed as $99, I automatically know it is $100. Marketing gimmicks don't work once you know about them or in other words, expectations can be set in a new e…

Marketing gimmicks absolutely still work even if you know about them because they take advantage of basic human psychology so when you're tired/hungry/sleepy or otherwise not operating at peak performance, your lizard brain/autopilot takes over and you choose what's been chosen for you.

There was a story where users complained that a particular process is taking lot of time. He coded a progress bar and all the complaints disappeared because it set their expectations and it was visible.
Post reply on HN