Live data from Hacker News

Claude 2.1

anthropic.com

121–130 of 339 posts

Re: Claude 2.1

#121

Earlier quoted context omitted.

Yeah but to be honest been a pain last days to get gpt 4 to write full pieces of code for more the 10-15 lines. Have to re-ask many times and at some point it forgets my initial specifications.

Earlier in the year I had ChatGPT 4 write a large, complicated C program. It did so remarkably well, and most of the code worked without further tweaking. Today I have the same experience. The thing fills in placeholder comments to skip over more difficult regions of the code, and routinely forgets what we were doing. Aside all the recent OpenAI drama, I've been displeased as a paying customer that their products rou…

These models are black boxes with unlabeled knobs. A change that makes things better for one user might make things worse for another user. It is not necessarily the case that just because it got worse for you that it got worse on average.

Also, the only way for OpenAI to really know if a model is an improvement or not is to test it out on some human guinea pigs.

Re: Claude 2.1

#122

Earlier quoted context omitted.

Because they're ultimately training data simulators and not actually brilliant aritifical programmers, we can expect Microsoft-affiliated models like ChatGPT4 and beyond to have much stronger value for coding because they have unmediated access to GitHub content. So it's most useful to look at other capabilities and opportunities when evaluating LLM's with a different heritage. Not to say we shouldn't evaluate this o…

idk we're just "have more kids" simulators and we do pretty good at programming as a side-task

Someone doesn't get good at programming with low quality learning sources. Also, a poor comparison because models are not people - might as well complain about how NPCs in games behave because they fail at problems real people can solve.

Re: Claude 2.1

#123
post #113

I don’t like Anthropic. they over-RLHF their models and make them refuse most requests. A conversation with Claude has never been pleasant to me. it feels like the model has an attitude or something.

Luckily, unlike OpenAI, Anthropic lets you prefill Claude's response which means zero refusals.

Can you give an example in how Anthropic and OpenAI differ in that?

Re: Claude 2.1

#124

I don’t like Anthropic. they over-RLHF their models and make them refuse most requests. A conversation with Claude has never been pleasant to me. it feels like the model has an attitude or something.

I agree, but that’s what you get when your mission is AI Safety so it’s going to be a dull experience.

Re: Claude 2.1

#125

I would love to use their API but I can never get anyone to respond to me. It's like they have no real interest in being a developer platform. Has anyone gotten their vague application approved?

Yeah, I find it interesting to read about their work, but it might as well be vaporware if I can't use the API as a developer. OpenAI has actual products I can pay for to do productive things.

Re: Claude 2.1

#126

Has anyone found any success with Claude or have any reason to use it? In my tests it is nowhere near GPT 3.5 or 4 in terms of reliability or usefulness and I've even found that it is useless compared to Mistral 7b. I don't understand what they are doing with those billions in investment when 7b open source models are surpassing them in practical day to day use cases.

My experiences have been the same, unfortunately. It can do simple tasks, but for anything requiring indirect reasoning or completion of partial content from media (think finishing sonnets as a training content test) Claude just falls flat. Honestly, I'm not sure what makes Claude so "meh". Not to mention having to fill out a Google Doc for API usage? Weird.

Re: Claude 2.1

#127
post #74

Earlier quoted context omitted.

Yeah but to be honest been a pain last days to get gpt 4 to write full pieces of code for more the 10-15 lines. Have to re-ask many times and at some point it forgets my initial specifications.

This has exactly been my experience for at least the last 3 months. At this point, I am thinking if paying that 20 bucks is even worth anymore which is a shame because when gpt-4 first came out, it was remembering everything in a long conversation and self-correcting itself based on modifications.

same. what would you use as an alternative?

Re: Claude 2.1

#128

1. A 200k context is bittersweet with that 70k->195k error rate jump. Kudos on that midsection error reduction, though! 2. I wish Claude had fewer refusals (as erroneously claimed in the title). Until Anthropic stops heavily censoring Claude, the model is borderline useless. I just don't have time, energy, or inclination to fight my tools. I decide how to use my tools, not the other way 'round. Until Anthropic stops…

I don't know what you're doing with your LLM, but I've only ever had one refusal and I've been working a lot with Claude since it's in bedrock

Re: Claude 2.1

#129
I was excited about Claude 2 for a few days but quickly determined that it’s much, much worse than GPT4 and haven’t used it much since. There really isn’t much point in using a worse LLM. And the bigger context window is irrelevant if the answers are bad despite that. I’ll give this new one a try but I doubt it will be better than the newly revamped GPT4.

Re: Claude 2.1

#130
post #63

I don't know what version claude.ai is currently running (apparently 2.1 is live, see below) but it's terrible compared to GPT-4. See below conversation I just had. > Claude 2.1 is available now in our API, and is also powering our chat interface at claude.ai for both the free and Pro tiers. ---- What version are you? I'm Claude from Anthropic. Do you know your version? No, I don't have information about a specific v…

GPT4 equivalent:

https://chat.openai.com/share/87b7fa63-ff22-48ae-8a2f-c9f71f...

No problems, of course.

Post reply on HN