Live data from Hacker News

GPT-4o's Memory Breakthrough – Needle in a Needlestack

nian.llmonpy.ai

141–150 of 256 posts

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#142
post #96
post #60

Earlier quoted context omitted.

Previous answer to this question: https://news.ycombinator.com/item?id=40361419

Your test is a good one but the point still stands that a novel dataset is the next step to being sure.

One could also programmatically (e.g. with nltk or spacy, replace nouns, named entities, etc) modify the dataset, even up to the point that every test run is unique.

You could also throw in vector similarity if you wanted to keep words as more synonyms or antonyms.

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#143

Someone needs to come up with a "synthesis from haystack" test that tests not just retrieval but depth of understanding, connections, abstractions across diverse information. When a person reads a book, they have an "overall intuition" about it. We need some way to quantify this. Needle in haystack tests feel like a simple test that doesn't go far enough.

I've been thinking about that as well. It's hard, but if you have a piece of fiction or non-fiction it hasn't seen before, then a deep reading comprehension question can be a good indicator. But you need to be able to separate a true answer from BS. "What does this work says about our culture? Support your answer with direct quotes." I found both gpt-4 and haiku to do alright at this, but sometimes give answers that…

[deleted]

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#144

We are all so majorly f*d. The general public does not know nor understand this limitation. At the same time OpenAI is selling this a a tutor for your kids. Next it will be used to test those same kids. Who is going to prevent this from being used to pick military targets (EU law has an exemption for military of course) or make surgery decisions?

Why would the military use ChatGPT or depend on any way on Openai 's policy? Wouldn't they just roll their own?

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#145

We are all so majorly f*d. The general public does not know nor understand this limitation. At the same time OpenAI is selling this a a tutor for your kids. Next it will be used to test those same kids. Who is going to prevent this from being used to pick military targets (EU law has an exemption for military of course) or make surgery decisions?

It is being used to pick military targets, with very little oversight.

https://www.972mag.com/lavender-ai-israeli-army-gaza/

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#146
post #125

Earlier quoted context omitted.

AI is already being used for picking targets in warzones - https://theconversation.com/israel-accused-of-using-ai-to-ta... . LLM's will of course also be used, due to their convenience and superficial 'intelligence', and because of the layer of deniability creating a technical substrate between soldier and civilian victim provides - as has happened for two decades with drones.

Note that the IDF explicitly denied that story: https://www.idf.il/en/mini-sites/hamas-israel-war-24/all-art... Probably this is due to confusion over what the term "AI" means. If you do some queries on a database, and call yourself a "data scientist", and other people who call themselves data scientists do some AI, does that mean you're doing AI? For left wing journalists who want to undermine the Israelis (the stor…

The IDF explicitly deny a lot of things, which turn out to be true.

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#147
post #137

Earlier quoted context omitted.

This is just doomerism. Even though this model is slightly better than the previous, using an LLM for high risk tasks like healthcare and picking targets in military operations still feels very far away. I work in healthcare tech in a European country and yes we use AI for image recognition on x-rays, retinas etc but these are fundamentally completely different models than a LLM. Using LLMs for picking military targe…

> picking targets in military operations I'm 100% on the side of Israel having the right to defend itself, but as I understand it, they are already using "AI" to pick targets, and they adjust the threshold each day to meet quotas. I have no doubt that some day they'll run somebody's messages through chat gpt or similar and get the order: kill/do not kill.

'Quotas each day to find targets to kill'.

That's a brilliant and sustainable strategy. /s

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#148

I just used it to compare two smaller legal documents and it completely hallucinated that items were present in one and not the other. It did this on three discrete sections of the agreements. Using ctrl-f I was able to see that they were identical in one another. Obviously this is a single sample but saying 90% seems unlikely. They were around ~80k tokens total.

> Obviously this is a single sample but saying 90% seems unlikely.

This is such an anti-intellectual comment to make, can't you see that?

You mention "sample" so you understand what statistics is, then in the same sentence claim 90% seems unlikely with a sample size of 1.

The article has done substantial research

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#149
post #148

I just used it to compare two smaller legal documents and it completely hallucinated that items were present in one and not the other. It did this on three discrete sections of the agreements. Using ctrl-f I was able to see that they were identical in one another. Obviously this is a single sample but saying 90% seems unlikely. They were around ~80k tokens total.

> Obviously this is a single sample but saying 90% seems unlikely. This is such an anti-intellectual comment to make, can't you see that? You mention "sample" so you understand what statistics is, then in the same sentence claim 90% seems unlikely with a sample size of 1. The article has done substantial research

That fact that it has some statistically significant performance is irrelevant and difficult to evaluate for most people.

He's a much simpler and correct description that almost everyone can understand: it fucks up constantly.

Getting something wrong even once can make it useless for most people. No amount of pedantry will change this reality.

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#150

I just used it to compare two smaller legal documents and it completely hallucinated that items were present in one and not the other. It did this on three discrete sections of the agreements. Using ctrl-f I was able to see that they were identical in one another. Obviously this is a single sample but saying 90% seems unlikely. They were around ~80k tokens total.

I have the same feeling. I asked to find duplicates in a list of 6k items and it basically hallucinated the entire answer multiple times. Some times it finds some, but it interlaces the duplicates with other hallucinated items. I wasn't expecting it to get it right, cause I think this task is challenging with a fixed amount of attention heads. However, the answer seems much worse than Claude Opus or GPT-4.
Post reply on HN