Live data from Hacker News

Claude 3 model family

anthropic.com

361–370 of 723 posts

Re: Claude 3 model family

#361
post #272

Opus just crushed Gemini Pro and GPT4 on a pretty complex question I have asked all of them, including Claude 2. It involved taking a 43 page life insurance investment pdf and identifying various figures in it. No other model has gotten close. Except for Claude 3 sonnet, which just missed one question.

Really? I tried the sonnet and it just was not very good.

Re: Claude 3 model family

#362

They claim that the new models "are significantly less likely to refuse to answer prompts that border on the system’s guardrails than previous generations of models", looks like about a third of "incorrect refusals" compared to Claude 2.1. Given that Claude 2 was completely useless because of this, this still feels like a big limitation.

The guard rails on the models make the llm-market a complete train wreck. Wish we could just collectively grow up and accept that if a computer says something bad that doesn't have any negative real world impact - unless we let it - just like literally any other tool.

They're not there to protect the user, they're they're to protect the brand of the provider. A bot that spits out evil shit easily screenshotted with the company's brand right there, isn't really great for growth or the company's brand both.

Re: Claude 3 model family

#363
post #272

Opus just crushed Gemini Pro and GPT4 on a pretty complex question I have asked all of them, including Claude 2. It involved taking a 43 page life insurance investment pdf and identifying various figures in it. No other model has gotten close. Except for Claude 3 sonnet, which just missed one question.

Did you compare it with Gemini Pro 1.5 with 1 million context window? (Ideal for 43 pg pdfs)

I have access to it and I can test it against Pro 1.5

Re: Claude 3 model family

#364
Opus (and all prior Claude models) still do not correctly answer the Sally question, unfortunately:

Q: Sally (a girl) has 3 brothers. Each brother has 2 sisters. How many sisters does Sally have?

A: Okay, let's organize the information we have:

Sally is a girl. Sally has 3 brothers. Each of Sally's brothers has 2 sisters. Now, let's think about this step by step:

We know that Sally is a girl, so she is a sister to her brothers. Each of Sally's brothers has 2 sisters. One of these sisters is Sally, but who is the other sister? Since Sally has only brothers and no other sisters are mentioned, we can conclude that the other sister for each brother is also Sally.

Therefore, Sally has no sisters other than herself. So, the answer is that Sally has 0 sisters.

https://imgur.com/a/EawcbeL

Re: Claude 3 model family

#365

Just signed up for Claude Pro to try out the Opus model. Decided to throw a complex query at it, combining an image with an involved question about SDXL fine tuning and asking it to do some math comparing the cost of using an RTX 6000 Ada vs an H100. It made a lot of mistakes. I provided it with a screenshot of Runpod's pricing for their GPUs, and it misread the pricing on an RTX 6000 ADA as $0.114 instead of $1.14.…

Hi, CISO of Anthropic here. Thank you for the feedback! If you can share any details about the image, please share in a private message. No LLM has had an emergent calculator yet.

What a joke of a response. No one is asking for emergent calculation ability just that the model gives the correct answer. LLM tools (functions etc) is old news at this point.

Re: Claude 3 model family

#366

Another naming disaster! Opus is better than sonnet? And sonnet is better than haiku? Perhaps this makes sense to people familiar with sonnets and haikus and opus....es? Nonsensical to me! I know everyone loves to hate on Google, but at least pro and ultra have a sort of sense of level of sophistication.

I know nothing about poetry and this is the order I would have expected if someone told me they had models called Opus, Sonnet and Haiku.

Re: Claude 3 model family

#367
post #364

Opus (and all prior Claude models) still do not correctly answer the Sally question, unfortunately: Q: Sally (a girl) has 3 brothers. Each brother has 2 sisters. How many sisters does Sally have? A: Okay, let's organize the information we have: Sally is a girl. Sally has 3 brothers. Each of Sally's brothers has 2 sisters. Now, let's think about this step by step: We know that Sally is a girl, so she is a sister to he…

It’s so convincing even I’m doubting my answer to this question

Re: Claude 3 model family

#368

Could anyone recommend an open-source tool capable of simultaneously sending the same prompt to various language models like GPT-4, Gemini, and Claude, and displaying their responses side by side for comparison? I tried chathub in the past, but they decided to not release any more source as of now.

https://chat.lmsys.org/

Choose Arena (side-by-side), it has Claude 3 Opus, Sonnet and GPT-4

Re: Claude 3 model family

#369

Earlier quoted context omitted.

This. Price is set by value delivered and what the market will pay for whatever capacity they have; it’s not a cost + X% market.

I'm more curious about the input/output token discrepancy Their pricing suggests that either output tokens are more expensive for some technical reason, or they're trying to encourage a specific type of usage pattern, etc.

Or that market research showed a higher price for input tokens would drive customers away, while a lower price for output tokens would leave money on the table.

Re: Claude 3 model family

#370

The APPS benchmark result of Claude 3 Opus at 70.2% indicates it might be quite useful for coding. The dataset measures the ability to convert problem descriptions to Python code. The average length of a problem is nearly 300 words. Interestingly, no other top models have published results on this benchmark. Claude 3 Model Card: https://www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bb... Table 1: Evaluation resul…

(full disclosure, I work at Anthropic) Opus has definitely been writing a lot of my code at work recently :)

What's your estimate of how much does it increase a typical programmer's productivity?
Post reply on HN