Live data from Hacker News

Claude now has access to a server-side container environment

anthropic.com

81–90 of 363 posts

Re: Claude now has access to a server-side container environment

#81
Not Claude specific, but related to the agent model of things...

I've been paying $10/month for GitHub Copilot, which I use via Microsoft's Visual Studio Code, and about a month ago, they added ChatGPT5 (preview), which uses the agent model of interaction. It's a qualitative jump that I'm still learning to appreciate in full.

It seems like the worst possible thing, in terms of security, to let an LLM play with your stuff, but I really didn't understand just how much easier it could be to work with an LLM if it's an agent. Previously I'd end up with a blizzard of python error messages, and just give up on a project, now it fixes it's own mess. What a relief!

Re: Claude now has access to a server-side container environment

#82
post #3

They need to focus on fixing reliability first. Their systems constantly go down and it appears they are having to quantise the models to keep up with demand, reducing intelligence significantly. New features like this feel pointless when the underlying model is becoming unusable.

Anthropic claims that they don't degrade models under load, and the performance issues were a result of a system error: https://status.anthropic.com/incidents/72f99lh1cj2c That being said, they still have capacity issues on any day of the week that ends in Y. No clue how long would that take to resolve.

> Last week, we opened an incident to investigate degraded quality in some Claude model responses. We found two separate issues that we’ve now resolved.

Re: Claude now has access to a server-side container environment

#83
post #26

Earlier quoted context omitted.

Not nitpicking, but they said: > we never intentionally degrade model quality as a result of demand or other factors Fully giving them the benefit of the doubt, I still think that still allows for a scenario like "we may [switch to quantized models|tune parameters], but our internal testing showed that these interventions didn't materially affect end user experience". I hate to parse their words in this way, because…

"Anecdata" is notoriously unreliable when it comes to estimating AI performance over time. Sure, people complain about Anthropic's AI models getting worse over time. As well as OpenAI's models getting worse over time. But guess what? If you serve them open weights models, they also complain about models getting worse over time. Same exact checkpoint, same exact settings, same exact hardware. Relative LMArena metrics,…

And yet, people's complaints about Claude Code over the past month and a bit are now justified by Anthropic stating that those complaints caused them to investigate and fix a bunch of issues (while investigating potential more issues with opus).

> But guess what? If you serve them open weights models, they also complain about models getting worse over time.

Isn't this also anecdotal, or is there data informing this statement?

I think you could be partially right, but I also don't think dismissing criticism as just being a change in perspective is correct either. At least some complaints are from power users who can usually tell when something is getting objectively worse (as was the case for some of us Claude Code users recently). I'm not saying we can't fool ourselves too, but I don't think that's the most likely assumption to make.

Re: Claude now has access to a server-side container environment

#84
post #3

They need to focus on fixing reliability first. Their systems constantly go down and it appears they are having to quantise the models to keep up with demand, reducing intelligence significantly. New features like this feel pointless when the underlying model is becoming unusable.

The same experience here: Claude with the pro plan over the summer was really doing a good job. The last 4 weeks? Constant slow-downs or API errors, more halucinating then before, and many mistakes. It appears to me that they are throttling to handle loads that they can't actually handle.

Re: Claude now has access to a server-side container environment

#85
post #47

Maybe one day Claude can rewrite its interface to be more accessible to blind people like me.

Curious what a11y issues you see with Claude? I use it a remarkable amount and haven't found any showstoppers. Web interface and Claude Code.

[flagged]

Re: Claude now has access to a server-side container environment

#86
post #26

Earlier quoted context omitted.

Not nitpicking, but they said: > we never intentionally degrade model quality as a result of demand or other factors Fully giving them the benefit of the doubt, I still think that still allows for a scenario like "we may [switch to quantized models|tune parameters], but our internal testing showed that these interventions didn't materially affect end user experience". I hate to parse their words in this way, because…

"Anecdata" is notoriously unreliable when it comes to estimating AI performance over time. Sure, people complain about Anthropic's AI models getting worse over time. As well as OpenAI's models getting worse over time. But guess what? If you serve them open weights models, they also complain about models getting worse over time. Same exact checkpoint, same exact settings, same exact hardware. Relative LMArena metrics,…

Selection bias + perceptual adaptation is my experience. Selection bias happens when we play the probabilities of using an LLM and we only focus on the things it does really well, because it can be really amazing. When you use a model a lot you increasingly see when they don't work well your perception changes to focus on what doesn't work vs. the what does.

Living evals can solve for the quantitative issues with infra and model updates, but not sure how to deal with perceptual adaptation.

Re: Claude now has access to a server-side container environment

#87
post #3

They need to focus on fixing reliability first. Their systems constantly go down and it appears they are having to quantise the models to keep up with demand, reducing intelligence significantly. New features like this feel pointless when the underlying model is becoming unusable.

That's not it! Direct engineering effort towards new features that will drive new customers and markets. Functionality is unimportant. Haven't you ever worked in enterprise software?

I'm kidding btw.

Re: Claude now has access to a server-side container environment

#88
post #71
post #4

Earlier quoted context omitted.

This can't be understated. I started using it heavily earlier this summer and it felt like magic. Someone signing up now based on how I described my personal experiences with it then would think I was out of my mind . For technical tasks it has been a net negative for me for the last several weeks. (Speaking of both Claude Code and the desktop app, both Sonnet and Opus >=4, on the Max plan.)

This, so much this... I signed up for Claude over a week ago and I totally regret it! Previously I was using it and some ChatGPT here and there (also had a subscription in the past) and I felt like Claude added some more value. But it's getting so unstable. It generates code, I see it doing that, and then it throws the code away and gives me the previous version of something 1:1 as a new version. And then I have to w…

> But it's getting so unstable. It generates code, I see it doing that, and then it throws the code away and gives me the previous version of something 1:1 as a new version.

I've had the same experience. Totally unreliable.

Re: Claude now has access to a server-side container environment

#89
post #4
post #3

They need to focus on fixing reliability first. Their systems constantly go down and it appears they are having to quantise the models to keep up with demand, reducing intelligence significantly. New features like this feel pointless when the underlying model is becoming unusable.

This can't be understated. I started using it heavily earlier this summer and it felt like magic. Someone signing up now based on how I described my personal experiences with it then would think I was out of my mind . For technical tasks it has been a net negative for me for the last several weeks. (Speaking of both Claude Code and the desktop app, both Sonnet and Opus >=4, on the Max plan.)

[flagged]

Re: Claude now has access to a server-side container environment

#90
post #3

They need to focus on fixing reliability first. Their systems constantly go down and it appears they are having to quantise the models to keep up with demand, reducing intelligence significantly. New features like this feel pointless when the underlying model is becoming unusable.

Maybe the people who build features like these are not the same people who buy cards and build data centers?

Maybe the reliability problems have almost nothing to do with what features they build, and are bottlenecked for completely different reasons.

Post reply on HN