That's so weird... This doesn't at all match my experience with Claude. I've never seen it behave this way.
Be specific.
That said, GPT always acts up even if I am specific, but I only have the free tier there.
141–150 of 467 posts
That's so weird... This doesn't at all match my experience with Claude. I've never seen it behave this way.
Be specific.
That said, GPT always acts up even if I am specific, but I only have the free tier there.
That's so weird... This doesn't at all match my experience with Claude. I've never seen it behave this way.
same. none of the available prompts are what I would prompt claude with and I get way better results than this. makes sense to me why the provided prompts result in the simulated outcomes. garbage in, garbage out.
Wow, you AI people really have a negative view of the technology y'all are trying to sell as the next Jesus
Earlier quoted context omitted.
Models hallucinate plausible answers to why they did things. It might be true and it might be complete fiction.
I'm growing increasingly confident that this is how people often work, as well.
Earlier quoted context omitted.
Models hallucinate plausible answers to why they did things. It might be true and it might be complete fiction.
I'm growing increasingly confident that this is how people often work, as well.
Earlier quoted context omitted.
Models hallucinate plausible answers to why they did things. It might be true and it might be complete fiction.
I'm growing increasingly confident that this is how people often work, as well.
https://www.patheos.com/blogs/tippling/2013/11/14/post-hoc-r...
> Why is half the site blue now? I asked you to change one button. > Half the site is blue. I asked for ONE button. Those are my only options when the site is clearly not blue, two buttons are. There is a reason for why I am much more specific than this.
Earlier quoted context omitted.
Models hallucinate plausible answers to why they did things. It might be true and it might be complete fiction.
You can also literally tell them: "Here is your session ID: $ID, lookup the .jsonl session, trace exactly why this decision was being made, present evidence and concrete proof, no guessing or assumptions" and you'll get an evidence-based report without guesses.
You can extend this further by using an adversarial agent trying to find mistakes in the other instance's logs in a loop where a 3rd neutral agent weighs the claims of the other two. This is also just another step in reducing error, it does not guarantee elimination of such errors. The latter is an impossible guarantee, even for humans.
Earlier quoted context omitted.
When you've been perfectly precise in your spec and language, isn't that programming? Why use a stochastic goblin to do things in that case?
For me, typing "Create a new namespace with these enums, functions and traits, that should follow X, Y and Z constraints" is faster than typing all that code manually, and typing less is less straining on my hands/fingers.
Earlier quoted context omitted.
I'm growing increasingly confident that this is how people often work, as well.
I kind of want my computer systems to be more reliable and predictable than paying an intern to manage something and asking why they messed up
Do you know anyone who actually reads and adheres closely to all of the documentation every time it's changed?
At least with Codex, this has not been my experience at all. It still screws up sure, but in every case I can ask "why did you do this" and it can trace back what made it take that particular decision. Typically it's always that I either didn't specify the problem correctly or made a really dumb mistake (executing the task on the wrong project....did this one yesterday) or it's something within a skill file that inst…
Models hallucinate plausible answers to why they did things. It might be true and it might be complete fiction.