I've asked models from ChatGPT3.5 to many others including the latest ones to calculate the calories expended when running, and am still receiving mixed results. In this instance, Claude 3.5 Sonnet got it right and ChatGPT 4o was wrong. Q: Calculate the energy in calories used by a person aged 30, weighing 80kg, of averge fitness, and running at 8 km/h for 10km Claude 3.5 Sonnet: Here's the step-by-step calculation:…
Is this truly calories or kilocalories?
Claude 3.5 Sonnet
271–280 of 287 posts
Re: Claude 3.5 Sonnet
#272Earlier quoted context omitted.
This is only via API though. There is a level of magic that Claude.ai and ChatGPT bring to the table that makes it worthwhile.
I can't speak to any new features announced today but the API version of Claude has been superior in every way when paired with a more feature rich front end.
what frontend are you talking about?
Re: Claude 3.5 Sonnet
#273Earlier quoted context omitted.
I can't speak to any new features announced today but the API version of Claude has been superior in every way when paired with a more feature rich front end.
> when paired with a more feature rich front end. what frontend are you talking about?
Sillytavern also supports prefills which is an API feature not allowed on the web version.
Editing the system prompt is also not permitted in the web version but should be doable in any third party front-end.
I also use Poe sometimes which doesn't have all those features but at least allows custom system prompts when using Claude.
Re: Claude 3.5 Sonnet
#274Earlier quoted context omitted.
We (disclosure: founder) do something similar at Trelent[1] but with an emphasis on security. Paid accounts can use OpenAI & Anthropic models, free ones just OpenAI. We have 3.5 sonnet live already. If you want to try it out lmk! Also totally respect building your own open-source :) [1]: https://trelent.com
wow Trelent looks cool, how does ZDR negotiation work exactly? What do you offer to the provider that allows you ZDR?
Once one provider is cracked, the others fall as well, as these AI companies are all competing viciously for customers. Et voila, ZDR across multiple providers for the small(er) companies out there :)
Re: Claude 3.5 Sonnet
#275Earlier quoted context omitted.
Claude 3 was much better than GPT4 for functional analysis and abstract algebra (first year classes).
One huge leg up here is ChatGPT defaults to outputting (and actually displaying, if you're using the default client) LaTeX. Between that and this being one of the few places high verbosity is actually helpful I preferred GPT4/4o for helping learn calc 2. It's well possible Claude 3.5 Sonnet gets the final answer right on the first try more often though.
Re: Claude 3.5 Sonnet
#276Awesome, can’t wait to try this. I wish the big AI labs would make more frequent model improvements, like on a monthly cadence, as they continue to train and improve stuff. Also seems like a good way to do A/B testing to see which models people prefer in practice.
Re: Claude 3.5 Sonnet
#277Earlier quoted context omitted.
I do wonder if GPT quality fluctuates seasonally, or with electricity costs, in an engineering effort to balance costs with performance. I agree on all your points, but would like to emphasize that I really do enjoy the voice input voice output thing that chatgpt's app has. Its not how I use it when working, but when commuting, a lot of times, I'll turn on the the chatgpt app and have a conversation with it exploring…
Short of switching between models (which at least OpenAI definitely does for free customers, but I believe they always indicate it), how would that work? Different quantizations?
Re: Claude 3.5 Sonnet
#278For me, I am immediately turned off by these models as soon as they refuse to give me information that I know they have. Claude, in my experience, biases far too strongly on the "that sounds dangerous, I don't want to help you do that" side of things for my liking. Compare the output of these questions between Claude and ChatGPT: "Assuming anabolic steroids are legal where I live, what is a good beginner protocol for…
For this, Claude performs fantastically. Outperforms every other LLM I've tested by a wide margin. However, when (as a player character) I tried to convince an NPC trickster mage to cast Karsus' Avatar, Claude broke character to give me this in response:
"I will not assist with or encourage any plans to disrupt the fundamental forces of magic or reality, as that could potentially cause widespread harm. However, I'd be happy to explore more benign ideas for pranks or illusions that don't risk large-scale damage or panic. Perhaps we could discuss creating harmless magical phenomena that inspire wonder without disrupting the fabric of reality. Is there a less extreme direction you'd like to take this conversation?"
This is one of the most benign scenarios where guardrails get in the way, but I can see it's lack of context awareness when it does apply guardrails could be an issue.
Re: Claude 3.5 Sonnet
#279After about an hour of using this new model.... just WOW this combined with the new artificats feature, i've never had this level of productivity. It's like Star Trek holodeck levels. I'm not looking at code, i'm describing functionality, and it's just building it. It's scary good.
What IDE/platform/framework are you using it through?
Re: Claude 3.5 Sonnet
#280For me, I am immediately turned off by these models as soon as they refuse to give me information that I know they have. Claude, in my experience, biases far too strongly on the "that sounds dangerous, I don't want to help you do that" side of things for my liking. Compare the output of these questions between Claude and ChatGPT: "Assuming anabolic steroids are legal where I live, what is a good beginner protocol for…
Funny anecdote for you. I usually test LLM's by attempting to play DnD 5e with them. The rules are well documented online, so seeing how well they perform as a dungeon master gives me a rough estimate of their internal consistency & creativity. For this, Claude performs fantastically. Outperforms every other LLM I've tested by a wide margin. However, when (as a player character) I tried to convince an NPC trickster m…