Someone needs to come up with a "synthesis from haystack" test that tests not just retrieval but depth of understanding, connections, abstractions across diverse information. When a person reads a book, they have an "overall intuition" about it. We need some way to quantify this. Needle in haystack tests feel like a simple test that doesn't go far enough.
My idea is to buy to a unpublished novel or screenplay with a detailed, internally consistent world built in to it and a cast of characters that have well crafted motivations and then ask it to continue writing from an arbitrary post-mid-point by creating a new plot line that combines two characters that haven't yet met in the story. If it understands the context it should be able to write a new part of the story and…
GPT-4o's Memory Breakthrough – Needle in a Needlestack
111–120 of 256 posts
Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack
#112The general public does not know nor understand this limitation. At the same time OpenAI is selling this a a tutor for your kids. Next it will be used to test those same kids.
Who is going to prevent this from being used to pick military targets (EU law has an exemption for military of course) or make surgery decisions?
Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack
#113We are all so majorly f*d. The general public does not know nor understand this limitation. At the same time OpenAI is selling this a a tutor for your kids. Next it will be used to test those same kids. Who is going to prevent this from being used to pick military targets (EU law has an exemption for military of course) or make surgery decisions?
Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack
#114We are all so majorly f*d. The general public does not know nor understand this limitation. At the same time OpenAI is selling this a a tutor for your kids. Next it will be used to test those same kids. Who is going to prevent this from being used to pick military targets (EU law has an exemption for military of course) or make surgery decisions?
If any regulator acts it will be the EU. The action, if it comes, will of course be very late, possibly years from now, when the horse has long left the stable.
Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack
#115We are all so majorly f*d. The general public does not know nor understand this limitation. At the same time OpenAI is selling this a a tutor for your kids. Next it will be used to test those same kids. Who is going to prevent this from being used to pick military targets (EU law has an exemption for military of course) or make surgery decisions?
Analyzing surgical field...
Identified: open chest cavity, exposed internal organs
Organs appear gooey, gelatinous, translucent pink
Comparing to database of aquatic lifeforms...
93% visual match found:
Psychrolutes marcidus, common name "blobfish"
Conclusion: Blobfish discovered inhabiting patient's thoracic cavity
Recommended action: Attempt to safely extract blobfish without damaging organsRe: GPT-4o's Memory Breakthrough – Needle in a Needlestack
#116We are all so majorly f*d. The general public does not know nor understand this limitation. At the same time OpenAI is selling this a a tutor for your kids. Next it will be used to test those same kids. Who is going to prevent this from being used to pick military targets (EU law has an exemption for military of course) or make surgery decisions?
So it can’t really be worse if there’s just a RNG in a box. It may be better.
Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack
#117I just used it to compare two smaller legal documents and it completely hallucinated that items were present in one and not the other. It did this on three discrete sections of the agreements. Using ctrl-f I was able to see that they were identical in one another. Obviously this is a single sample but saying 90% seems unlikely. They were around ~80k tokens total.
Yeah I asked for an estimate of the percentage of the US population that lives in the DMV area (DC, Maryland, Virginia) and it was off by 50% of the actual answer, which I only realized when I realized I shouldn’t trust its estimate for anything important
Also: would you expect random people to fare any better?
Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack
#118This is a very promising development. It would be wise for everyone to go back and revise old experiments that failed now that this capability is unlocked. It should also make RAG even more powerful now that you can load a lot more information into the context and have it be useful.
Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack
#119We are all so majorly f*d. The general public does not know nor understand this limitation. At the same time OpenAI is selling this a a tutor for your kids. Next it will be used to test those same kids. Who is going to prevent this from being used to pick military targets (EU law has an exemption for military of course) or make surgery decisions?
The Diamond Age.
Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack
#120We are all so majorly f*d. The general public does not know nor understand this limitation. At the same time OpenAI is selling this a a tutor for your kids. Next it will be used to test those same kids. Who is going to prevent this from being used to pick military targets (EU law has an exemption for military of course) or make surgery decisions?
So the military already was using math to pick targets, this is just the next logical step, albeit, scary as hell step.