Live data from Hacker News

Being “Confidently Wrong” is holding AI back

promptql.io

241–250 of 274 posts

Re: Being “Confidently Wrong” is holding AI back

#241
post #139

Earlier quoted context omitted.

1-turn instruction following and multi-turn instruction following are not the same exact capability, and some AIs only "get good" at the former. 1-turn gets more training attention - because it's more noticeable, in casual use and benchmarks both, and also easier to train for. With weak multi-turn instruction following, context data will often dominate over user instructions. Resulting in very "loopy" AI - and more s…

This is a good point, and to drive this home to people, if you have a conversation of this pattern: User: Fix this problem ... Assistant: X User: No, don't do X Assistant: Y User: No, Y is wrong too. Assistant: X It is generally pointless to continue. You now have a context that is full of the assistant explaining to you and itself why X and Y are the right answers, and much less context of you explaining why it is w…

I always ask "Tell me what you think it is I am asking" before asking for a solution. Improves the solution and context.

Re: Being “Confidently Wrong” is holding AI back

#242

What's really funny to me is, sometimes it fixes itself if you just ask "are you SURE ABOUT THIS ANSWER?" myself and others often wonder, why the heck don't they run a 2nd model to "proofread" output or spot check it. Like did you actually answer the question or are you going off a really weird tangent. I asked Perplexity some question for sample UI code for Rust / Slint, it gave me a beautiful web UI, I think it got…

I've asked that question on accurate answers and had the bot say oops and change the answer to an inaccurate one. This seems to happen with about the same frequency on both sides so I'm not sure how helpful it will ultimately be.

Re: Being “Confidently Wrong” is holding AI back

#243
post #194

Wow, there really is an xkcd for everything.

Those are original cartoons drawn in the style of XKCD. But strangely enough, in the second cartoon, the Megan clone seems to change from a thin stick figure to suddenly wearing clothes? I'm not sure if the comic was AI-assisted or not. AI-generated images do not usually contain identical pixel data when a panel repeats.

The script is uncanny as well. My guess is the author used AI to generate the panels/dialogue, then stitched together cutouts from real xkcd comics over the top of the AI-generated panels here and there. That exact shape of the heads is too close to not be a copy-paste job, but other variations suggest AI involvement. The desk gets rotated in the third panel of the first comic, the female character in the second comic gets clothes out of nowhere, etc.

Regardless of how the author made the comics, they're very weird.

Re: Being “Confidently Wrong” is holding AI back

#244

Earlier quoted context omitted.

It’s not massively underplaying it imo. AI hype is real. This is revolutionary technology that humanity has never seen before. But it happened at a time where hype can be delivered at a magnitude never before seen by humanity as well to a degree of volume that is completely unnatural by any standard set previously by hype machines created by humanity. Not even landing on the moon has inundated people with as much hyp…

I think you would have really enjoyed living in the '50s, when the future was bright and colonizing Mars was basically a solved problem. What we got instead is a bunch of wisecracking programmers who like to remind everyone of the 90–90 rule, or the last 10 percent.

Wisecracking programmers lol. You talk as if programming is like something to be proud of. It’s one of the most lucrative jobs with ease of entry as a boot camp can turn someone from zero to hero in a year.

And then you mouth off a buzz phrase not even coined by a programmer but repeated to the point of annoyance about how the final 10 percent is always the hardest as if programmers who copy the phrase are so smart.

Bro the last 10 percent being the hardest doesn’t mean the previous 90 percent didn’t happen. The first 90 percent is a feat in itself and LLMs can now even do PRs. That was a feat no one just 5 years ago could have predicted was possible in our lifetimes.

Idiot programmers and their generic wise cracks were the ones saying that AI would never be able to pass the Turing test and this was just 4 years ago.

Re: Being “Confidently Wrong” is holding AI back

#245
post #141

Earlier quoted context omitted.

Because it's easy to learn to stop engaging with those loops, treating them as a sign you provided too little context, and instead start a new conversation with an expanded prompt. It doesn't mean these loops aren't an issue, because they are, but once you stop engaging with them and cut them off, they're a nuisance rather than a showstopper.

They happen in subtle ways that aren't always easy and are rarely early in a project I want to just throw away. "So what if you have to throw out a week's worth of work. That's how these things work. Accept it and you'll be happier. I have and I'm happy. Don't you see that it's OK to have your tool corrupt your work half way through. It's the future of work and you're being left behind by not letting your tools corru…

Doing a week's worth of work without verifying is unprofessional whether you do it with AI or without.

Re: Being “Confidently Wrong” is holding AI back

#246
post #219

Earlier quoted context omitted.

It's a problem with LLM's and people are "holding it wrong". It makes zero difference that they've been sold as doing better if other people learn how to use them effectively and I choose to ignore how to get the best possible results out of them.

Except that it's impossible to "hold it right" -- even when following the guidance from its makers.

I have no problem "holding it right". Just today I had AI write 100% of the code for two different tools, using an AI assistant which wrote all the code for itself after the initial ~100 lines.

It's not hard to learn to be productive with these models.

Re: Being “Confidently Wrong” is holding AI back

#247
post #230
post #221

Earlier quoted context omitted.

That's fine once or twice. At that point people should learn that this isn't how they work, and figure out how to use them better. It's not a tools fault if people insist on continuing to use them in counter-productive ways.

They’re non-deterministic, remember? So it’s not always the case that an LLM will get stuck in this sort of loop. Hence why people get frustrated when it happens and continue to think that perhaps it should be working on a more consistent basis.

So are people.

It is no more productive to continue to go in circles with an argumentative person who refuses to see reason.

If someone haven't learnt that lesson, they will get poor results at a whole lot more things in life than talking to AI.

Re: Being “Confidently Wrong” is holding AI back

#248
post #221

Earlier quoted context omitted.

That's fine once or twice. At that point people should learn that this isn't how they work, and figure out how to use them better. It's not a tools fault if people insist on continuing to use them in counter-productive ways.

It's not the tools fault when people RTFM (guidance from the tool maker) and use it as it's intended (again, by the tool maker, who presumably knows how it works and is in the best position to guide users). "If you keep pressing the back button like the IE engineers told you to, of course you will fail to go back. To go back you want to press the forward button. Are you an idiot? Press the forward button to go back,…

No, it's not the tools fault if you continue to use it in ways that according to you, yourself does not work, despite the availability of better guidance.

Do you always insist on listening to guidance you've observed doesn't work?

It sounds immensely counter-productive.

Meanwhile I'll continue to have AI tools write the majority of my code at this point.

Re: Being “Confidently Wrong” is holding AI back

#249

Isn’t it obvious that the confidently wrong problem will never go away because all of this is effectively built on a statistical next token matcher? Yeah sure you can throw on hacks like RAG, more context window, but it’s still built on the same foundation. It’s like saying you built a 3D scene on a 2D plane. You can employ clever tricks to make 2D look 3D at the right angle, buts it’s fundamentally not 3D, which obv…

It seems obvious to me, but there was a camp that thought, at least at one time, that probabilistic next token could be effectively what humans are doing anyways, just scaled up several more orders of magnitude. It always felt obvious to me that there was more to human cognition than just very sophisticated pattern matching, so I'm not surprised that these approaches are hitting walls.

Re: Being “Confidently Wrong” is holding AI back

#250
post #226
post #223

Earlier quoted context omitted.

That is not the same thing! You are talking about the point distribution of the next token. We are talking about the uncertainty associated with each of those candidate tokens; a distribution of distributions. It's the difference between a categorical distribution and a Dirichlet. https://en.wikipedia.org/wiki/Dirichlet_distribution

I think we're talking about the same thing. I should be clear that I don't think the selected token probabilities being reported are enough, but if you're reporting each returned tokens probability (both selected and discarded) and aggregating the cumulative probabilities of the given context, it should be possible to see when you're trending centrally towards uncertainty.

No, it isn't the same thing. The softmax probabilities are estimates; they're part of the prediction. The other poster is talking about the uncertainty in these estimates, so the uncertainty in the softmax probabilities.

The softmax probabilities are usually not a very good indication of uncertainty, as the model is often overconfident due to neural collapse. The uncertainty in the softmax probabilities is a good indication though, and can be used to detect out-of-distribution entries or poor predictions.

Post reply on HN