interesting to see feature launches are coming via official website while usage restrictions are coming in with a team member's twitter account - https://x.com/trq212/status/2037254607001559305 . also, someone rightly predicted this rugpull coming in when they announced 2x usage - https://x.com/Pranit/status/2033043924294439147
Funnily, Anthropic's pricing etc. why I'm using GLM-5 a bunch more outside of work. Definitely not Opus level, but surprisingly decent. Though I got lucky and got the Alibaba Coding Model lite plan, which is so cheap they got rid of it
Schedule tasks on the web
241–250 of 261 posts
Re: Schedule tasks on the web
#242Here goes my project.
Re: Schedule tasks on the web
#243Earlier quoted context omitted.
> Let's just include this for now to easily refute your argument which I don't know where it comes from: https://transformer-circuits.pub/2022/in-context-learning-an ... I love it when people include links to papers that refute their words. So, Antropic (which is heavily reliant on hype and making models appear more than they are) authors a paper which clearly states: "tokens later in context are easier to predict an…
> Frankly, I find your aggressiveness quite tiring having to answer for opinions with no basis in the literature is I'm sure very tiring for you. Your aggression being met is I'm sure uncomfortable. > I love it when people include links to papers that refute their words. > So, Antropic (which is heavily reliant on hype and making models appear more than they are) authors a paper which clearly states: "tokens later in…
Having only literature on your side must feel nice.
> They have access to the full transcript, and they have access to the full codebase, the diff history, whatever knowledge base is available.
Yes. And it means that they don't learn, and they alway miss important details when rebuilding the world.
That's why even the tiniest codebases are immediately filled with duplications, architecturally unsound decisions, invalid assumptions etc.
> also not an accurate understanding of how agents and their context work; you can use multiple session to digest and distill information useful in other sessions and in fact
I say: agents don't learn and have to rebuild the world from scratch
You: not an accurate understanding of how agents and their context work.... they rebuild the world from scratch every time they run.
> You keep dismissing this literature as if you have understood it
No. I'm dismissing your flawed interpretation of purely theoretical constructs.
Chinchilla doesn't project unlimited amazing scalability. If anything, it shows a very real end of scalability.
Anthropic's paper adopts a nice marketable term for a process that has little to do with learning.
Etc.
Meanwhile you do keep rejecting actual real-world behaviour of these systems.
> Then are you arguing this progress will stop? I'm just not sure I understand, you seem to contradict yourself
I didn't say that either. Your opponents don't contradict themselves if you only stop to pretend they think or say.
Your unsubstantiated belief is that improvements are on a steep linear or even exponensial progression. Because "literature" or something.
Looking past all the marketing bullshit, it could be argued that growth is at best logarithmic, and most improvments come from tooling around (harnesses, subagents etc.). While all the failure modes from a year ago are still there: misunderstanding context, inability to maintain cohesion between sessions, context pollution etc.
And providers are running into the issue of getting non-polluted trainig data.
---
At this point we're going around in circles, and I'm no interested in arguing with theorists.
Adieu
Re: Schedule tasks on the web
#244Earlier quoted context omitted.
I worry about the costs from an energy and environmental impact perspective. I love that AI tools make me more productive, but I don't like the side effects.
Environmental impact of ai is greatly overstated. Average person will make bigger positive impact on environment by reducing his meat intake by 25% compared with combined giving up flying and AI use.
Re: Schedule tasks on the web
#245Earlier quoted context omitted.
What kind of software are people building where AI can just one shot tickets? Opus 4.6 and GPT 5.4 regularly fail when dealing with complicated issues for me.
Of course not all tickets are complex. Last week I had to fix a ticket which was to display the update date on a blog post next to the publish date. Perfect use case for AI to one shot.
Re: Schedule tasks on the web
#246Earlier quoted context omitted.
> So what do you think the difference is between humans and an agent in this respect? Humans learn. Agents regurgitate training data (and quality training data is increasingly hard to come by). Moreover, humans learn (somewhat) intangible aspects: human expectations, contracts, business requirements, laws, user case studies etc. > Verifiable domain performance SCALES, we have no reason to expect that this scaling wil…
> Where did I say that? I didn’t even mention money, just the broader resource term. A lot of business are mostly running experiments if the current set of tooling can match the marketing (or the hype). They’re not building datacenters or running AI labs. Such experiments can’t run forever. I'm just going to ask that you read any of my other comments, this is not at all how coding agents work and seems to be the most…
We have all the data now.
I don’t see where the huge gap should come from, as one person before they said they still make basic errors.
Models got better for a bunch of soft tuning. Language and abstractness is not really the same thing there are a lot of very good speakers that are terrible in logic and abstractness.
Thinking abstract sometimes makes it necessary to leave language and draw or som people even code in another coding language to get it.
We’ve seen it with the compiler project it’s nice looking but if you would want to make a competitive compiler you would be as far as starting fresh
Re: Schedule tasks on the web
#247Earlier quoted context omitted.
> Because it cannot do it? Ah ok so you didn't really read my comment, what is your counter argument? Models are just fundamentally incapable of understanding business context? They are demonstrably already capable of this to a large extent. > Every investment has a date where there should be a return on that investment. If there’s no date, it’s a donation of resources (or a waste depending on perspective). what are…
> They are demonstrably already capable of this to a large extent. I’d very like to see such demonstration. Where someone hands over a department to an agent and let it makes decisions. > This convo now turns into the "AI is not profitable and this is a house of cards" theme? Where did I say that? I didn’t even mention money, just the broader resource term. A lot of business are mostly running experiments if the curr…
Re: Schedule tasks on the web
#248Earlier quoted context omitted.
I would suggest the prompt is an example of garbage in that's going to produce garbage out. Sitting down to confront the problem you're solving will show this, while Claude is going to happily spit out what looks like a plausibly functional system. So for example the only "analysis" of CI failures are which systems failed and who/what committed the changes to those things. The only way AI would help me here is if the…
> I would suggest the prompt is an example of garbage in that's going to produce garbage out. Sitting down to confront the problem you're solving will show this, while Claude is going to happily spit out what looks like a plausibly functional system. I think this shows the value. > Which granted is probably real for a lot of software firms Here's the rub though; for many many people it's a huge improvement over what…
Re: Schedule tasks on the web
#249Earlier quoted context omitted.
What kind of software are people building where AI can just one shot tickets? Opus 4.6 and GPT 5.4 regularly fail when dealing with complicated issues for me.
Of course not all tickets are complex. Last week I had to fix a ticket which was to display the update date on a blog post next to the publish date. Perfect use case for AI to one shot.
Maybe for a non-dev it would be nice to submit a ticket and have it auto-fixed by an agent. But in the devs case, it feels like it would be faster to just do it manually.
Re: Schedule tasks on the web
#250Earlier quoted context omitted.
I don't think anybody is doubting its ability to generate thousands of PR's though. And yes, it's usually in the stuff that should have been automated already regardless of AI or not.
Depends on your circle. On HN I would argue that there are still a fair number of people that would be surprised to see what heavy organizational usage of AI actually looks like. On a non programming online group, of which I am a member of several, people still think that AI agents are the same as they were in mid 2025 and they can't answer "how many R's are in the following word:". Same thing even when chatting with…
“There are 4 Rs in the word “burberrorrly.” Here they are highlighted: burberrorrly (positions 3, 6, 7, 9)”
Obviously not a real word, but perhaps the fundamental concept remains