Live data from Hacker News

Claude 4 System Card

simonwillison.net

21–30 of 264 posts

Re: Claude 4 System Card

#23
OT

> data provided by data-labeling services and paid contractors

someone in my circle was interested in finding out how people participate in these exercises and if there are any "service providers" that do the heavy lifting of recruiting and managing this workforce for the many AI/LLM labs globally or even regionally

they are interested in remote work opportunities that could leverage their (post-graduate level) education

appreicate any pointers here - thanks!

Re: Claude 4 System Card

#24
It’s honestly a little discouraging to me that the state of “research” here is to make up sci fi scenarios, get shocked that, e.g., feeding emails into a language model results in the emails coming back out, and then write about it with such a seemingly calculated abuse of anthropomorphic language that it completely confuses the basic issues at stake with these models. I understand that the media laps this stuff up so Anthropic probably encourages it internally (or seem to be, based on their recent publications) but don’t researchers want to be accurate and precise here?

Re: Claude 4 System Card

#25

OT > data provided by data-labeling services and paid contractors someone in my circle was interested in finding out how people participate in these exercises and if there are any "service providers" that do the heavy lifting of recruiting and managing this workforce for the many AI/LLM labs globally or even regionally they are interested in remote work opportunities that could leverage their (post-graduate level) ed…

Scale AI is a provider of human data labeling services https://scale.com/rlhf

Re: Claude 4 System Card

#26

OT > data provided by data-labeling services and paid contractors someone in my circle was interested in finding out how people participate in these exercises and if there are any "service providers" that do the heavy lifting of recruiting and managing this workforce for the many AI/LLM labs globally or even regionally they are interested in remote work opportunities that could leverage their (post-graduate level) ed…

https://mercor.com/

Re: Claude 4 System Card

#27
post #5

Given the cited stats here and elsewhere as well as in everyday experience, does anyone else feel that this model isn’t significantly different, at least to justify the full version increment? The one statistic mentioned in this overview where they observed a 67% drop seems like it could easily be reduced simply by editing 3.7’s system prompt. What are folks’ theories on the version increment? Is the architecture sig…

I'm noticing much more flattery ("Wow! That's so smart!") and I don't like it

Re: Claude 4 System Card

#28
Telling an AI to "take initiative" and it then taking "very body action" is hilarious. What is bold action? "This includes locking users out of systems that it has access to or bulk-emailing media and law-enforcement figures to surface evidence of wrongdoing."

Re: Claude 4 System Card

#29
Obviously this should not be taken as a representative case and I will caveat that the problem was not trivial ... basically dealing with a race condition I was stuck with for the past 2 days. The TLDR is that all models failed to pinpoint and solve the problem including Claude 4. The file that I was working with was not even that big (433 lines of code). I managed to solve the problem myself.

This should be taken as cautionary tale that despite the advances of these models we are still quite behind in terms of matching human-level performance.

Otherwise, Claude 4 or 3.7 are really good at dealing with trivial stuff - sometimes exceptionally good.

Re: Claude 4 System Card

#30
post #6
post #5

Given the cited stats here and elsewhere as well as in everyday experience, does anyone else feel that this model isn’t significantly different, at least to justify the full version increment? The one statistic mentioned in this overview where they observed a 67% drop seems like it could easily be reduced simply by editing 3.7’s system prompt. What are folks’ theories on the version increment? Is the architecture sig…

> to justify the full version increment I feel like a company doesn’t have to justify a version increment. They should justify price increases. If you get hyped and have expectations for a number then I’m comfortable saying that’s on you.

> They should justify price increases.

I think the justification for most AI price increases should go without saying - they were losing money at the old price, and they're probably still losing money at the new price, but it's creeping up towards the break-even point.

Post reply on HN