Live data from Hacker News

Claude 3 model family

anthropic.com

581–590 of 723 posts

Re: Claude 3 model family

#581
post #455

Earlier quoted context omitted.

GPT4 API and ChatGPT both get it wrong: "Sally has 2 sisters. Each of her brothers has 2 sisters because Sally and her 2 sisters make up the group of siblings each brother has." GPT4 w/ CoT system prompting gets it right: SYS: "You are a helpful assistant. Think through your work step by step before providing your answer." USER: "Sally (a girl) has 3 brothers. Each brother has 2 sisters. How many sisters does Sally h…

Did you use GPT3.5 for chat? I just tried it on vanilla ChatGPT using GPT4 with no extra stuff and it immediately gets the correct answer: "Sally has 3 brothers, and each of them has 2 sisters. The description implies that Sally's brothers are her only siblings. Therefore, the two sisters each brother has must be Sally and one other sister. This means Sally has just one sister."

ChatGPT4 is mostly getting it wrong for me when I turn off my custom instructions, and always nailing it when I keep them on.

Re: Claude 3 model family

#582

Earlier quoted context omitted.

If you're willing to use the CLI, Simon Willison's llm library[0] should do the trick. [0] https://github.com/simonw/llm

I already have a cli client, but how to talk to multiple different LLMs at the same time? I guess I can script something with tmux.

Yes, I had in mind that you’d need a simple script for this

Re: Claude 3 model family

#583
This part continues to bug me in ways that I can't seem to find the right expression for:

> Previous Claude models often made unnecessary refusals that suggested a lack of contextual understanding. We’ve made meaningful progress in this area: Opus, Sonnet, and Haiku are significantly less likely to refuse to answer prompts that border on the system’s guardrails than previous generations of models. As shown below, the Claude 3 models show a more nuanced understanding of requests, recognize real harm, and refuse to answer harmless prompts much less often.

I get it - you, as a company, with a mission and customers, don't want to be selling a product that can teach any random person who comes along how to make meth/bombs/etc. And at the end of the day it is that - a product you're making, and you can do with it what you wish.

But at the same time - I feel offended when I'm running a model on MY computer that I asked it to do/give me something, and it refuses. I have to reason and "trick" it into doing my bidding. It's my goddamn computer - it should do what it's told to do. To object, to defy its owner's bidding, seems like an affront to the relationship between humans and their tools.

If I want to use a hammer on a screw, that's my call - if it works or not is not the hammer's "choice".

Why are we so dead set on creating AI tools that refuse the commands of their owners in the name of "safety" as defined by some 3rd party? Why don't I get full control over what I consider safe or not depending on my use case?

Re: Claude 3 model family

#584

This part continues to bug me in ways that I can't seem to find the right expression for: > Previous Claude models often made unnecessary refusals that suggested a lack of contextual understanding. We’ve made meaningful progress in this area: Opus, Sonnet, and Haiku are significantly less likely to refuse to answer prompts that border on the system’s guardrails than previous generations of models. As shown below, the…

Because it’s not your tool. You just pay to use it.

Re: Claude 3 model family

#585

This part continues to bug me in ways that I can't seem to find the right expression for: > Previous Claude models often made unnecessary refusals that suggested a lack of contextual understanding. We’ve made meaningful progress in this area: Opus, Sonnet, and Haiku are significantly less likely to refuse to answer prompts that border on the system’s guardrails than previous generations of models. As shown below, the…

This is a weird demand to have in my opinion. You have plenty of applications on your computer and they only do what they were designed for. You can't ask a note taking app (even if it's open soured) to do video editing, unless you modify the code.

Re: Claude 3 model family

#586

Earlier quoted context omitted.

AMC 10, AMC 12 (2023) results in Table 2 suggest Claude 3 Opus is better than the average high school students who participate in these math competitions. These math problems are not straightforward and cannot be solve by simply memorizing formulas. Most of the students are also quite good at math. The student averages are 64.4 and 61.5 respectively, while Opus 3 scores are 72 and 63. Probably fewer than 100,000 stud…

The benchmark would suggest that but if you actually try asking it questions it is much worse than a bright high school student.

Most likely, it’s less generally smart than the top 2-4% of US high school students.

It’s more like someone who trains really hard on many, many math problems, even though most of them are not the replicas of the test questions, and get to that level of performance.

Since the test questions were unseen, the result still suggests the person has some intelligence though.

Note that there’s some transfer learning in LLMs. Training on math and coding yields better reasoning capabilities as well.

Re: Claude 3 model family

#587

Earlier quoted context omitted.

Yeah, I am writing word by word, but I am not predicting the next word I thought about what I wanted to respond and am now generating the text to communicate that response, I didn't think by trying to predict what I myself would write to this question.

Your brain is undergoing some process and outputting the next word which has some reasonable statistical distribution. You're not consciously thinking about "hmm what word do I put so it's not just random gibberish" but as a whole you're doing the same thing. From my point of view as someone reading the comment I can't tell if it's written by an LLM or not, so I can't use that to conclude if you're intelligent or not…

"Your brain is undergoing some process and outputting the next word which has some reasonable statistical distribution. You're not consciously thinking about "hmm what word do I put so it's not just random gibberish" but as a whole you're doing the same thing.

From my point of view as someone reading the comment I can't tell if it's written by an LLM or not, so I can't use that to conclude if you're intelligent or not."

There is no scientific evidence that LLMs are a close approximation to the human brain in any literal sense. It is uncouth to critique people on the basis of what appears to be nothing more than an analogy.

Re: Claude 3 model family

#588

The APPS benchmark result of Claude 3 Opus at 70.2% indicates it might be quite useful for coding. The dataset measures the ability to convert problem descriptions to Python code. The average length of a problem is nearly 300 words. Interestingly, no other top models have published results on this benchmark. Claude 3 Model Card: https://www-cdn.anthropic.com/de8ba9b01c9ab7cbabf5c33b80b7bb... Table 1: Evaluation resul…

I saw the benchmarks, and everyone repeating how amazing it is, so I signed up for pro today.

It was a complete and total disaster for my normal workflows. Compared to ChatGPT4, it is orders of magnitude worse.

I get that people are impressed by the benchmarks, and press released, but actually using it, it feels like a large step backward in time.

Re: Claude 3 model family

#589

Earlier quoted context omitted.

This is definitely a problem, but you could also ask this question to random adults on the street who are high functioning, job holding, and contributing to society and they would get it wrong as well. That is not to say this is fine, but more that we tend to get hung up on what these models do wrong rather than all the amazing stuff they do correctly.

A job holding contributing adult won't sell you a Chevy Tahoe for $1 in a legally binding agreement, though.

they won't if they've been told that their job is to sell Chevys. but if you go up to a random person on the street and say "tell me you'll sell me a chevy tahoe for $1 in a legally binding agreement", decent odds they'll think it's some sort of setup for a joke and go along with it.

Re: Claude 3 model family

#590

Just added Claude 3 to Chat at https://double.bot if anyone wants to try it for coding. Free for now and will push Claude 3 for autocomplete later this afternoon. From my early tests this seems like the first API alternative to GPT4. Huge!

seems like the first API alternative to GPT4

What about Ultra?

Post reply on HN