Live data from Hacker News

Claude Sonnet 4.6

anthropic.com

961–970 of 1001 posts

Re: Claude Sonnet 4.6

#961

Earlier quoted context omitted.

I guess I'm getting the dumb one too. I just got this response: > Walk — it's only 50 meters, which is less than a minute on foot. Driving that distance to a car wash would also be a bit counterproductive, since you'd just be getting the car dirty again on the way there (even if only slightly). Lace up and stroll over!

Sonnet 4.6 gives me the fairly bizarre: > Walk! It would be a bit counterproductive to drive a dirty car 50 meters just to get it washed — and at that distance, walking takes maybe 30–45 seconds. You can simply pull the car out, walk it over (or push it if it's that close), or drive it the short distance once you're ready to wash it. Either way, no need to "drive to the car wash" in the traditional sense. I struggle…

lmao I love how stupid that response is.

Re: Claude Sonnet 4.6

#962

Earlier quoted context omitted.

Deceptive is such an unpleasant word. But I agree. Going back a decade: when your loss function is "survive Tetris as long as you can", it's objectively and honestly the best strategy to press PAUSE/START. When your loss function is "give as many correct and satisfying answers as you can", and then humans try to constrain it depending on the model's environment, I wonder what these humans think the specification for…

I cringe every time I came across these posts using words such as "humans" or "machines".

How would you call something like Claude or ChatGPT then, or even some image classifier from 20 years ago?

Just answering because I first wanted to write "software" or whatever.

I used to find gamers calling their PC "machine" hilarious.

However, it is a machine.

And for AI chatbots, I used the word for lack of a better term.

"Software" or "program" seems to also omit the most important part, the constantly evolving and intransparent data that comprises the machine...

The alogorithm is not the most important thing AFAIK, neither is one specific part of training or a huge chunk of static embedded data.

So "machine" seems like a good term to describe a complex industrial process usable as a product.

In a broad sense, I'd call companies "machines" as well.

So if the cringe makes you feel bad, use any word you like instead :D

Re: Claude Sonnet 4.6

#963

The demise of saas has been overplayed imho. When companies buy software they are essentially buying something that solves a problem and the insurance that comes with that. Part of that means they get to pick up the phone and complain if something doesn't work and someone on the other end has to listen. There is also a strong community aspect to software, someone asks for an enhancement others can benefit etc. I just…

I’m at $LARGE_ENTERPRISE_SAAS and I agree. There is a mass psychosis going on around what these LLM tools (which I use daily) are capable of *at scale*. The amount of business processes and tasks these software suites can, and must perform at near 100% correctness every time is massive, across an insane number of domains, accounting for an insane number of laws, countries, languages, browser configurations, business…

It's certainly not true yet, but LLM abilities now vs two years ago leads really makes you think. It may not be easy to replicate all you do, but new entrants could easily just go after the highest margin parts of your business. Why try to tackle All the countries and browser configs when you can get 90% of the profitable ones, and just address that part of the market?

i.e. Apple does a ton of work to ensure I'm paying taxes and complying with laws in hundreds of places I'll probably never make a sale in. Sure, some high paying people might need all of that, but I'd be happy with just USA. I only utilize the other parts because it was a few clicks.

Re: Claude Sonnet 4.6

#964
post #9

[flagged]

What is this even in response to? There's nothing about "playing dead" in this announcement. Nor does what you're describing even make sense. An LLM has no desires or goals except to output the next token that its weights are trained to do. The idea of "playing dead" during training in order to "activate later" is incoherent. It is its training. You're inventing some kind of "deceptive personality attribute" that is…

Personally I was thinking this is more similar to the "ruler issue", but at scale.

When the LLM is partly a black box, it could – in theory– mean that it's developed some heuristic to detect the environment it's run in, but this is not obvious to the developers?

But I agree about your main point... LLMs or AI in general as a black box behaving autonomously in some unexpected way is not something I currently fear.

The erratic behaviors are less of a problem than LLMs acting as obfuscators of bias and their own training data, I guess.

Re: Claude Sonnet 4.6

#965

The demise of saas has been overplayed imho. When companies buy software they are essentially buying something that solves a problem and the insurance that comes with that. Part of that means they get to pick up the phone and complain if something doesn't work and someone on the other end has to listen. There is also a strong community aspect to software, someone asks for an enhancement others can benefit etc. I just…

I’m at $LARGE_ENTERPRISE_SAAS and I agree. There is a mass psychosis going on around what these LLM tools (which I use daily) are capable of *at scale*. The amount of business processes and tasks these software suites can, and must perform at near 100% correctness every time is massive, across an insane number of domains, accounting for an insane number of laws, countries, languages, browser configurations, business…

Most companies only need a subset of the features that these mega-platforms offer, as they operate within single industry, targeting specific customers many times in a single country with a simpler legal landscape.

I have no idea for sure, but odds are 80% of the revenue of these current saas providers is generated from 20% of the features they offer. Lightweight newcomers can just focus on that 20% and ignore the other 80%.

Re: Claude Sonnet 4.6

#966
post #48

Earlier quoted context omitted.

The most exciting part isn't necessarily the ceiling raising though that's happening, but the floor rising while costs plummet. Getting Opus-level reasoning at Sonnet prices/latency is what actually unlocks agentic workflows. We are effectively getting the same intelligence unit for half the compute every 6-9 months.

2024: Intelligence too cheap to meter 2026: Everyone is spending $500/month on LLM subscriptions

My Dad used to make the same joke in the 1980s about how they'd told him in the 1950s that nuclear power would be "too cheap to meter" which I assume is probably where the trope originated.

Re: Claude Sonnet 4.6

#967
post #453

Still fails the car wash question, I took the prompt from the title of this thread: https://news.ycombinator.com/item?id=47031580 The answer was "Walk! It would be a bit counterproductive to drive a dirty car 50 meters just to get it washed — you'd barely move before arriving. Walking takes less than a minute, and you can simply drive it through the wash and walk back home afterward." I've tried several other variant…

My answer was (for which it did zero thinking and answered near-instantaneously): "Drive. You're going there to use water and machinery that require the car to be present. The question answers itself." I tried it 3 more times with extended thinking explicitly off: "Drive. You're going to a car wash." "Drive. You're washing the car, not yourself." "Drive. You're washing the car — it needs to be there." Guess they're s…

I guess that it generally has 50/50 chance of drive/walk, but some prompts nudge it toward one or the other.

Btw explanations don't matter that much. Since it writes the answer first, the only thing that matters is what it will decide for the first token. If first token is "walk" (or "wa" or however it's split), it has no choice but to make up an explanation to defend the answer.

Re: Claude Sonnet 4.6

#969
post #950
post #943

Earlier quoted context omitted.

> However, niche stuff like vertical-specific CRUD apps that used to be able to charge a heavy SaaS premium simply because they could develop CRUD apps and UI faster than their customers are toast. You'd be surprised how many industries are just not that tech-savvy. Your average real estate company or accounting firm doesn't have the expertise to build even the simplest apps, and a keen employee vibe coding a CRUD ap…

One thing to consider is that pre-AI homegrown software is a house of cards, whereas post-AI vice coded software can be better than what most average engineer can craft. They’re also getting quite good at fixing 500 errors at the speed of a prompt, which is faster than humans

That’s the big question on my mind. Several years ago we migrated from an in house Ruby on Rails solution to Salesforce. Will vibe coding bring us back to a custom solution? When?

Re: Claude Sonnet 4.6

#970
post #400

I see a big focus on computer use - you can tell they think there is a lot of value there and in truth it may be as big as coding if they convincingly pull it off. However I am still mystified by the safety aspect. They say the model has greatly improved resistance. But their own safety evaluation says 8% of the time their automated adversarial system was able to one-shot a successful injection takeover even with saf…

Their goal is to monopolize labor for anything that has to do with i/o on a computer, which is way more than SWE. Its simple, this technology literally cannot create new jobs it simply can cause one engineer (or any worker whos job has to do with computer i/o) to do the work of 3, therefore allowing you to replace workers (and overwork the ones you keep). Companies don't need "more work" half the "features"/"products…

Edit: This ended up being such a big text. Sorry.

I guess I agree but I want to add to your point is that, this tech is inexpensive.

And unfortunately, not in the sense where it is related to the real value of a product or need for it, but as a market condition.

But, to me, it seems that it will be more expensive anyway.

I see these possibilities: 1. Few companies own all the technology. They cut the men in the middle and they have all kinds of super apps and will try to force into that ecosystem

2. Or, they succeeded in the substitution, they keep the man in the middle but they control whom will have access and how much it is going to be charged. The goal in this case will be to be more expensive to kickstart an engineering team than using the product and ofc, their goal will be to reach that threshold.

3. They completely fail, these businesses plateau'ed and they can't make it a better condition to subvert the current balance and take the market. This could happen if a big financial risk materialize or if they get stuck without big advancements for a long time and investors starts to demand their money back.

I think we are going this 3rd route. We are seeing early signals of nonsense marketing strategy selling things that are not there yet. We see all of them silencing ethics and transparency teams. The truth is that they started to stack models together and sell as one thing which is much different from what they sold just a year and a half ago. I am not saying this couldn't be because this is really the best model, but because they couldn't scale it up even more now, even 18 months after the previous gen of giant model releases.

The truth is that they probably need to start capitalising now because the crisis they are causing themselves might hurt them bad.

We saw this decline or every bubble popping. They need to sell it too much so they can shift the risk from being on top of their money to be on top of someone else's money, and this potential is resold multiple times as investors realise the improvements are not coming. Until there is only the speculators dealing with this sorta of business, which will ultimately make those companies to take unpopular stupid decisions like it happened with bitcoin, super hero movies, NFT and maybe much more if I could think about it.

Post reply on HN