Live data from Hacker News

Claude now has access to a server-side container environment

anthropic.com

201–210 of 363 posts

Re: Claude now has access to a server-side container environment

#201
post #129

Earlier quoted context omitted.

https://status.anthropic.com/incidents/72f99lh1cj2c They recently resolved two bugs affecting model quality, one of which was in production Aug 5-Sep 4. They also wrote: Importantly, we never intentionally degrade model quality as a result of demand or other factors, and the issues mentioned above stem from unrelated bugs. Sibling comments are claiming the opposite, attributing malice where the company itself says it…

The problem is twofold: - They're reporting that only impacted Haiku 3.5 and Sonnet 4. I used neither model during the time period I'm concerned with. - It took them a month to publicly acknowledge that issue, so now we lack confidence there isn't another underlying issue going undetected (or undisclosed, less charitably) that affects Opus.

They posted

> We are continuing to monitor for any ongoing quality issues, including reports of degradation for Claude Opus 4.1.

I take that as acknowledgment that there might be an issue with Opus 4.1 (granted, undetected still), but not undisclosed, and they're actively looking for it? I'd not jump to "they must be hiding things" yet. They're building, deploying and scaling their service at incredible pace, they, as we all, are bound to get some things wrong.

Re: Claude now has access to a server-side container environment

#203

Earlier quoted context omitted.

I think it's a psychological bias of some sort. When the feeling of newness wears off and you realize the model is still kind of shit, you have an imperfect memory of the first few uses when you were excited and have repressed the failures from that period. As the hype wears off you become more critical and correctly evaluate the model

I get that it’s fun and stylish to tell people they aren’t aware of their own cognitive biases, but it’s also a difficult take to falsify, which is why I generally have a high bar for people to clear when they want to assert that something is all in people’s heads. People seem to turn to this with a lot when the suspicion many people have is difficult to verify. And while I don’t trust a suspicion just because it’s h…

Doesn't this go both ways? A random selection of commenters online out of hundreds of thousands of devs using LLMs reporting degraded capability based on personal perception isn't exactly statistically meaningful data.

I've seen the cycle of claims going from "10x multiplier, like a team of junior devs" to "nerfed" for so many model/tool releases at this point it's hard for me not to believe there's an element of perceptual bias going on, but how much that contributes vs real variability on the backend is impossible to know for sure.

Re: Claude now has access to a server-side container environment

#204
post #43
post #4

Earlier quoted context omitted.

This can't be understated. I started using it heavily earlier this summer and it felt like magic. Someone signing up now based on how I described my personal experiences with it then would think I was out of my mind . For technical tasks it has been a net negative for me for the last several weeks. (Speaking of both Claude Code and the desktop app, both Sonnet and Opus >=4, on the Max plan.)

"can't be overstated", you mean

This one is interesting because I have seen a fair amount of "can't be understated" on reddit also. Interesting case of linguistic drift.

Re: Claude now has access to a server-side container environment

#206
post #4
post #3

They need to focus on fixing reliability first. Their systems constantly go down and it appears they are having to quantise the models to keep up with demand, reducing intelligence significantly. New features like this feel pointless when the underlying model is becoming unusable.

This can't be understated. I started using it heavily earlier this summer and it felt like magic. Someone signing up now based on how I described my personal experiences with it then would think I was out of my mind . For technical tasks it has been a net negative for me for the last several weeks. (Speaking of both Claude Code and the desktop app, both Sonnet and Opus >=4, on the Max plan.)

I had even posted a Ask HN: if people had experienced issues with Claude Code since for me it's slowed down substantially, it'll frequently just pause and take much longer. I have a Claude Max 5X plan.

I've been running ccusage to monitor and my usage in $ terms has dropped to a 1/3 of what it was few weeks ago. While some of it could be due to how I'm using it, but a drop of 60%-70% cannot be attributed to that alone and I think is partly due to the performance.

To add: frequently, as in almost every time: 1) it'll start doing something and will go silent for a long time. 2) pressing esc to interrupt will take a long time to take action since it's probably stuck doing something. Earlier, interrupting via esc used to be almost instantaneous.

So, I still like it, but at my 1/3 drop in measured usage I'm almost tempted to go back to Pro and see if that'll meet my needs.

Re: Claude now has access to a server-side container environment

#207
post #3

They need to focus on fixing reliability first. Their systems constantly go down and it appears they are having to quantise the models to keep up with demand, reducing intelligence significantly. New features like this feel pointless when the underlying model is becoming unusable.

I wonder if their API model is different from the subscription model. People called me crazy saying how GitHub copilot is better than Clause code but since I started using Claude code these past 3 weeks, times and times again, copilot + Claude sonnet 4 is better

Copilot did a giant leap imo, when Sonnet 4 arrived. BUT, I do have a lot of tempeorary problems where it just stops responding. Last week was awful, today though worked perfectly. I both vibe-coded a very wide (TUI, GUI, WEBUI, CLI, backend etc) python util for our specific product+environment and solved a bug in parallell using Sonnet 4 and GTP 4.1. I tried going to Sonnet when GPT fscked up, and its just hilarious. GPT can try sometimes to fix things 5 times in a row, Sonnet just directly fixes it. If only the enterprise quota was infinite.... :)

Re: Claude now has access to a server-side container environment

#209

This will either result in a lot of people being able to sleep more, or an absolute avalanche of crap is about to be released upon society. A lot of the people I graduated with spent their 20s making powerpoint and excel. There would be people with a master's in engineering getting phone calls at 1am, with an instruction to change the fonts on slide 75, or to slightly modify some calculation. Most of the real decisio…

The global economy has been down the rabbit hole and through the looking glass into the land of the red queen as far as I’ve known. “Now here you see, it takes all the running you can do, to keep in the same place” as she says. I fully believe any slack this creates will get gobbled up in competition in a few years.

whip cracks

Rent just went up 20%! Back to the trenches, citizen. You wouldn’t want to lose that precious healthcare now would you?

unintelligible babbling about “productivity!”, “impact!”, “efficiency!” hums quietly in the distance

Re: Claude now has access to a server-side container environment

#210

Earlier quoted context omitted.

I don’t think you’re crazy, something is off in their models. As an example I’ve been using an MCP tool to provide table schemas to Claude for months. There was a point where it stopped recognizing the tool unless mentioned in early August. Maybe that’s related to their degraded quality issue. This morning after pulling the correct schema info Sonnet started hallucinating columns (from Shopify’s API docs) and added t…

I have read so many anecdotes about so many models that "were great" and aren't now. I actually think this is psychological bias. It got a few things right early on, and that's what you remember. As time passes, the errors add up, until the memory doesn't match reality. The "new shiny" feeling goes away, and you perceive it for what it really is: a kind of shitty slot machine > personally am frustrated that there’s n…

I think you’re onto something but it works the opposite way too. When you first start using a new model you are more forgiving because almost by definition you were using a worse model before. You give if the sorts of problems the old model couldn’t do, and the new model can do them; you see only success, and the places where it fails, well, you can’t have it all.

Then after using the new model for a few months you get used to it, you feel like you know what it should be able to do, and when it can’t do that, you’re annoyed. You feel like it got worse. But what happened is your expectations crept up. You’re now constantly riding it at 95% of its capabilities and hitting more edge cases where it messes up. You think you’re doing everything consistently, but you’re not, you’ve dramatically dialed up your expectations and demands relative to what you were doing months ago. I don’t mean “you,” I mean the royal “you”, this is what we all do. If you think your expectations haven’t risen, go back and look at your commits from six months ago and tell me I’m wrong.

Post reply on HN