Live data from Hacker News

Claude 2.1

anthropic.com

21–30 of 339 posts

Re: Claude 2.1

#23
This is where OpenAI/MSFT loses. Chaos in OpenAI/MSFT will lead to Anthropic overtaking them. They've already been ahead in many areas, dead locked in others, but with OpenAI facing a crisis, they'll likely gain significant headway if they execute well .. at least for the risk-adverse enterprise use-cases. I still am not a fan of either due to restrictions and 'safety' training wheels that treat me like a child

Re: Claude 2.1

#24

>less refusals This is not quoted in the article

If anything the "Hard Questions" chart indicates _more_ refusals as the "Declined to answer" increased from 25% to 45%. They are positioning this as a good thing since declining to answer instead of hallucinating is the preferable choice, but I agree there is nothing in the article indicating less refusals.

Re: Claude 2.1

#25
> We’re also introducing system prompts, which allow users to provide custom instructions to Claude in order to improve performance. System prompts set helpful context that enhances Claude’s ability to take on specified personalities and roles or structure responses in a more customizable, consistent way aligned with user needs.

Alright, now Anthropic has my attention. It'll be interesting to see how easy it is to use/abuse it compared to ChatGPT.

The documentation shows Claude does cheat with it a bit, indicating the way you invoke system prompt is just through a similar instruction as with ChatGPT in the initial query in contrast to ChatGPT's ChatML schema: https://docs.anthropic.com/claude/docs/how-to-use-system-pro...

Re: Claude 2.1

#26

This is where OpenAI/MSFT loses. Chaos in OpenAI/MSFT will lead to Anthropic overtaking them. They've already been ahead in many areas, dead locked in others, but with OpenAI facing a crisis, they'll likely gain significant headway if they execute well .. at least for the risk-adverse enterprise use-cases. I still am not a fan of either due to restrictions and 'safety' training wheels that treat me like a child

From what I see they still suck bad

Re: Claude 2.1

#28

I don’t like Anthropic. they over-RLHF their models and make them refuse most requests. A conversation with Claude has never been pleasant to me. it feels like the model has an attitude or something.

Good thing that you can now use a system prompt to (theoetically) override most of the RLHF.

Re: Claude 2.1

#29
I hope that the long context length models start getting better. Claude 1 and GPT-4-128K both struggle hard once you get past about 32K tokens.

Most of the needle in a haystack papers are too simple of a task. They need harder tasks to test these long context length models for if they are truly remembering things or not.

Re: Claude 2.1

#30
post #16

There are a lot of interesting things in this announcement, but the "less refusals" from the submission title isn't mentioned at all. If anything, it implies that there are more refusals because "Claude 2.1 was significantly more likely to demur rather than provide incorrect information." That's obviously a positive development, but the title implies that there is progress in reducing the censorship false positives,…

Really impressed with the progress of Anthropic with this release. I would love to see how this new version added to Vectara's Hallucination Evaluation Leaderboard.

https://huggingface.co/spaces/vectara/Hallucination-evaluati...

Post reply on HN