"One honest caveat", "no glitches, no color changes" good tests and I read it to the end but I wish it was written by a human.
Nothing hard. Everybody wins.
91–95 of 95 posts
"One honest caveat", "no glitches, no color changes" good tests and I read it to the end but I wish it was written by a human.
Nothing hard. Everybody wins.
Maybe I'm a control freak, but asking agents to one-shot random apps is nothing like how I actually use AI in software engineering.
surprised it isn't a bigger thing, eg artificial analysis doesn't report anything like that
still doesn't measure the human-agent interaction part, but that's pure vibes atp
Earlier quoted context omitted.
Because serious science is hard and valuable for its rigour, and shouldn't be compared with just poking at data to see what happens
Don't be fooled, there is politics, opinions, and less rigor in science as well.