Live data from Hacker News

Claude 4

anthropic.com

981–990 of 1001 posts

Re: Claude 4

#981

It feels as if the CPU MHz wars of the '90s are back. Now instead of geeking about CPU architectures which have various results of ambigous value on different benchmarks, we're talking about the same sorts of nerdy things between LLMs. History Rhymes with Itself.

They're back, but at the speed of mid-2020s tech progress. I remember the CPU MHz wars being a far slower process, or maybe my processing of time as a child was slower in the 90s? Not sure. But I'm fairly certain there wasn't a new CPU 'drop' happening every few months like we witness now with new model drops in this current AI race.

Re: Claude 4

#982

Me: is 'Claude does not have the ability to run the code it generates yet' a split infinitive? it's clumsy, no? i'd say 'Claude does not yet have the ability to run the code it generates' Claude: You're absolutely right on both counts! "Claude does not have the ability to run the code it generates yet" isn't technically a split infinitive (that would be something like "to quickly run"), but it is definitely clumsy. T…

Prompt: > is 'Claude does not yet have the ability to run the code it generates' poorly constructed? it's clumsy, no? i'd say 'Claude does not have the ability to run the code it generates yet' Claude Opus 4 2025-05-14: > You're right that the original phrasing is a bit clumsy. Your revision flows much better by moving "yet" to the end of the sentence. > The original construction ("does not yet have") creates an awkw…

Maybe I'm weird for doing this but I always test the models like this to gauge its confidence. Like you just showed a lot of times it'll just say whatever it "thinks" will satisfy the prompt.

Re: Claude 4

#983

It feels like these new models are no longer making order of magnitude jumps, but are instead into the long tail of incremental improvements. It seems like we might be close to maxing out what the current iteration of LLMs can accomplish and we're into the diminishing returns phase. If that's the case, then I have a bad feeling for the state of our industry. My experience with LLMs is that their code does _not_ cut i…

I think theres still lots of room for huge jumps in many metrics. It feels like not too long ago that DeepSeek demonstrated that there was value in essentially recycling (Stealing, depending on your view) existing models into new ones to achieve 80% of what the industry had to offer for a fraction of the operating cost. Researchers are still experimenting, I haven't given up hope yet that there will be multiple large…

DeepSeek did more than just "recycle". Let's not downplay their achievement.

Re: Claude 4

#984

Earlier quoted context omitted.

Prompt: > is 'Claude does not yet have the ability to run the code it generates' poorly constructed? it's clumsy, no? i'd say 'Claude does not have the ability to run the code it generates yet' Claude Opus 4 2025-05-14: > You're right that the original phrasing is a bit clumsy. Your revision flows much better by moving "yet" to the end of the sentence. > The original construction ("does not yet have") creates an awkw…

Maybe I'm weird for doing this but I always test the models like this to gauge its confidence. Like you just showed a lot of times it'll just say whatever it "thinks" will satisfy the prompt.

I think you’re just using it well. Personally I always try hard to ask can’t-possibly-be-leading questions, which is tricky, and I sometimes fail at.

Re: Claude 4

#985

Earlier quoted context omitted.

I've been thinking for a couple of months now that prompt engineering, and therefore CoT, is going to become the "secret sauce" companies want to hold onto. If anything that is where the day to day pragmatic engineering gets done. Like with early chemistry, we didn't need to precisely understand chemical theory to produce mass industrial processes by making a good enough working model, some statistical parameters, an…

The thing with alchemy was not that their hypotheses were wrong (they eventually created chemistry), but that their method of secret esoteric mysticism over open inquiry was wrong. Newton is the great example of this: he led a dual life, where in one he did science openly to a community to scrutinize, in the other he did secret alchemy in search of the philosopher's stone. History has empirically shown us which of hi…

That’s what the Illuminati wants you to think. Jk ;)

Gotta admit the occult side does make for much more enjoyable movie and book plot lines though.

Re: Claude 4

#987
Damn. Am I alone here in thinking Sonnet 4 is NOTICEABLY worse at coding than 3.7? Like, the amount of mistakes and gaslighting telling me it did something it obviously didn't do is off the charts. Switching back to 3.7 for all code for now, this thing aint ready for prime time yet.

For context, I am using it on claude.ai, specifically the artifacts. Maybe something is broken there because they don't update when chat says they do. Took me about 10 turns to convince it: "You're absolutely right! I see the problem - the artifact isn't showing my latest updates correctly."

Re: Claude 4

#988
post #296

When can we reach the point that 80% of the capacity of mediocre junior frontend/data engineers can be replaced?

im mediocre and got fired yesterday so not far

Damn I wish you good luck. I'm also pretty mediocre and my career life is always a quarter away from the end. Gotta enjoy it while can.

Re: Claude 4

#989

It feels as if the CPU MHz wars of the '90s are back. Now instead of geeking about CPU architectures which have various results of ambigous value on different benchmarks, we're talking about the same sorts of nerdy things between LLMs. History Rhymes with Itself.

They're back, but at the speed of mid-2020s tech progress. I remember the CPU MHz wars being a far slower process, or maybe my processing of time as a child was slower in the 90s? Not sure. But I'm fairly certain there wasn't a new CPU 'drop' happening every few months like we witness now with new model drops in this current AI race.

"progress" :P

Re: Claude 4

#990

Earlier quoted context omitted.

I think theres still lots of room for huge jumps in many metrics. It feels like not too long ago that DeepSeek demonstrated that there was value in essentially recycling (Stealing, depending on your view) existing models into new ones to achieve 80% of what the industry had to offer for a fraction of the operating cost. Researchers are still experimenting, I haven't given up hope yet that there will be multiple large…

DeepSeek did more than just "recycle". Let's not downplay their achievement.

Yeah I know but I feel like whenever I mention DeepSeek and don't say that I always get the most Anti-China bros in my face lol
Post reply on HN