Live data from Hacker News

GPT-4o's Memory Breakthrough – Needle in a Needlestack

nian.llmonpy.ai

181–190 of 256 posts

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#181

We are all so majorly f*d. The general public does not know nor understand this limitation. At the same time OpenAI is selling this a a tutor for your kids. Next it will be used to test those same kids. Who is going to prevent this from being used to pick military targets (EU law has an exemption for military of course) or make surgery decisions?

I hear these complaints and can't see how this is worse than the pre-AI situation. How is an AI "hallucination" different from human-generated works that are just plain wrong, or otherwise misleading?

Humans make mistakes all the time. Teachers certainly did back when I was in school. There's no fundamental qualitative difference here. And I don't even see any evidence that there's any difference in degree, either.

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#182

We are all so majorly f*d. The general public does not know nor understand this limitation. At the same time OpenAI is selling this a a tutor for your kids. Next it will be used to test those same kids. Who is going to prevent this from being used to pick military targets (EU law has an exemption for military of course) or make surgery decisions?

Surgeons don’t need a text based LLM to make decisions. They have a job to do and a dozen years of training into how to do it. They have 8 years of schooling and 4-6 years internship and residency. The tech fantasy that everyone is using these for everything is a bubble thought. I agree with another comment, this is Doomerism.

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#183
post #148

I just used it to compare two smaller legal documents and it completely hallucinated that items were present in one and not the other. It did this on three discrete sections of the agreements. Using ctrl-f I was able to see that they were identical in one another. Obviously this is a single sample but saying 90% seems unlikely. They were around ~80k tokens total.

> Obviously this is a single sample but saying 90% seems unlikely. This is such an anti-intellectual comment to make, can't you see that? You mention "sample" so you understand what statistics is, then in the same sentence claim 90% seems unlikely with a sample size of 1. The article has done substantial research

And also article is testing on a different task (Needle in a Needlestack which is kind of similar to Needle in a Haystack), compared to finding a difference between two documents. For sure it's useful to know that the model does ok in one and really bad in the other, does not mean that original test is flawed.

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#184

I just used it to compare two smaller legal documents and it completely hallucinated that items were present in one and not the other. It did this on three discrete sections of the agreements. Using ctrl-f I was able to see that they were identical in one another. Obviously this is a single sample but saying 90% seems unlikely. They were around ~80k tokens total.

That's a different test than needle-in-a needlestack, although telling in how brittle these models are - competent in one area, and crushingly bad in others.

Needle-in-a-needlestack contrasts with needle-in-a-haystack by being about finding a piece of data among similar ones (e.g. one specific limeric among thousands of others), rather than among disimilar ones.

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#185

I just used it to compare two smaller legal documents and it completely hallucinated that items were present in one and not the other. It did this on three discrete sections of the agreements. Using ctrl-f I was able to see that they were identical in one another. Obviously this is a single sample but saying 90% seems unlikely. They were around ~80k tokens total.

I've done the same experiment with local laws and caught GPT hallucinating fines and fees! The problem is real.

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#186

We are all so majorly f*d. The general public does not know nor understand this limitation. At the same time OpenAI is selling this a a tutor for your kids. Next it will be used to test those same kids. Who is going to prevent this from being used to pick military targets (EU law has an exemption for military of course) or make surgery decisions?

Israel is already doing exactly that... using AI to identify potential targets based on their network of connections, giving these potential targets a cursory human screening, then OK-ing the bombing of their entire family since they have put such faith (and/or just don't care) in this identification process that these are considered high-value targets where "collateral damage" is accepted.

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#187

Earlier quoted context omitted.

Note that the IDF explicitly denied that story: https://www.idf.il/en/mini-sites/hamas-israel-war-24/all-art... Probably this is due to confusion over what the term "AI" means. If you do some queries on a database, and call yourself a "data scientist", and other people who call themselves data scientists do some AI, does that mean you're doing AI? For left wing journalists who want to undermine the Israelis (the stor…

> Probably this is due to confusion over what the term "AI" means. AI is how it is marketed to the buyers. Either way, the system isn't a database or simple statistics. https://www.accessnow.org/publication/artificial-genocidal-i... Ex, autonomous weapons like "smart shooter" employed in Hebron and Bethlehem: https://www.hrw.org/news/2023/06/06/palestinian-forum-highli...

[flagged]

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#188

We are all so majorly f*d. The general public does not know nor understand this limitation. At the same time OpenAI is selling this a a tutor for your kids. Next it will be used to test those same kids. Who is going to prevent this from being used to pick military targets (EU law has an exemption for military of course) or make surgery decisions?

Israel is already doing exactly that... using AI to identify potential targets based on their network of connections, giving these potential targets a cursory human screening, then OK-ing the bombing of their entire family since they have put such faith (and/or just don't care) in this identification process that these are considered high-value targets where "collateral damage" is accepted.

[deleted]

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#189

We are all so majorly f*d. The general public does not know nor understand this limitation. At the same time OpenAI is selling this a a tutor for your kids. Next it will be used to test those same kids. Who is going to prevent this from being used to pick military targets (EU law has an exemption for military of course) or make surgery decisions?

https://pessimistsarchive.org/

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#190

We are all so majorly f*d. The general public does not know nor understand this limitation. At the same time OpenAI is selling this a a tutor for your kids. Next it will be used to test those same kids. Who is going to prevent this from being used to pick military targets (EU law has an exemption for military of course) or make surgery decisions?

I hear these complaints and can't see how this is worse than the pre-AI situation. How is an AI "hallucination" different from human-generated works that are just plain wrong, or otherwise misleading? Humans make mistakes all the time. Teachers certainly did back when I was in school. There's no fundamental qualitative difference here. And I don't even see any evidence that there's any difference in degree, either.

>There's no fundamental qualitative difference here...degree either.

I've heard the same comparisons made with self-driving cars (i.e. that humans are fallible, and maybe even more error-prone).

But this misses the point. People trust the fallibility they know. That is, we largely understand human failure modes (errors in judgement, lapses in attention, etc) and feel like we are in control of them (and we are).

OTOH, when machines make mistakes, they are experienced as unpredictable and outside of our control. Additionally, our expectation of machines is that they are deterministic and not subject to mistakes. While we know bugs can exist, it's not the expectation. And, with the current generation of AI in particular, we are dealing with models that are generally probabilistic, which means there's not even the expectation that they are errorless.

And, I don't believe it's reasonable to expect people to give up control to AI of this quality, particularly in matters of safety or life and death; really anything that matters.

TLDR; Most people don't want to gamble their lives on a statistic, when the alternative is maintaining control.

Post reply on HN