GPT-4o's Memory Breakthrough – Needle in a Needlestack
21–30 of 256 posts
Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack
#22Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack
#23I'd like to see this for Gemini Pro 1.5 -- I threw the entirety of Moby Dick at it last week, and at one point all books Byung Chul-Han has ever published, and it both cases it was able to return the single part of a sentence that mentioned or answered my question verbatim, every single time, without any hallucinations.
Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack
#24I'd like to see this for Gemini Pro 1.5 -- I threw the entirety of Moby Dick at it last week, and at one point all books Byung Chul-Han has ever published, and it both cases it was able to return the single part of a sentence that mentioned or answered my question verbatim, every single time, without any hallucinations.
But this content is presumably in its training set, no? I'd be interested if you did the same task for a collection of books published more recently than the model's last release.
This doesn't mean you're wrong, though.
Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack
#25Earlier quoted context omitted.
But this content is presumably in its training set, no? I'd be interested if you did the same task for a collection of books published more recently than the model's last release.
I would hope that Byung-Chul Han would not be in the training set (at least not without his permission), given he's still alive and not only is the legal question still open but it's also definitely rude. This doesn't mean you're wrong, though.
Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack
#26I'd like to see this for Gemini Pro 1.5 -- I threw the entirety of Moby Dick at it last week, and at one point all books Byung Chul-Han has ever published, and it both cases it was able to return the single part of a sentence that mentioned or answered my question verbatim, every single time, without any hallucinations.
But this content is presumably in its training set, no? I'd be interested if you did the same task for a collection of books published more recently than the model's last release.
It replied: "The text indicates that graphene sheets present high optical transparency and are able to absorb thermal radiations with high efficacy. They can then convert these radiations into electrical signals efficiently.".
Screenshot of the PDF with the relevant sentence highlighted: https://i.imgur.com/G3FnYEn.png
[0] https://www.routledge.com/Advances-in-Green-and-Sustainable-...
Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack
#27Earlier quoted context omitted.
If I had access to Gemini with a reasonable token rate limit, I would be happy to test Gemini. I have had good results with it in other situations.
What version of Gemini is built into Google Workspace? (I just got the ability today to ask Gemini anything about emails in my work Gmail account, which seems like something that would require a large context window)
Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack
#28Someone needs to come up with a "synthesis from haystack" test that tests not just retrieval but depth of understanding, connections, abstractions across diverse information. When a person reads a book, they have an "overall intuition" about it. We need some way to quantify this. Needle in haystack tests feel like a simple test that doesn't go far enough.
This whole thing would have to be kept under lock-and-key in order to be useful, so it would only serve as a kind of personal benchmark. Or it could possibly be a prestige award that is valued for its conclusions and not for its ability to use the methodology to create improvements in the field.
Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack
#29I thought google Gemini had almost perfect needle in haystack performance inside 1 million tokens?
Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack
#30Someone needs to come up with a "synthesis from haystack" test that tests not just retrieval but depth of understanding, connections, abstractions across diverse information. When a person reads a book, they have an "overall intuition" about it. We need some way to quantify this. Needle in haystack tests feel like a simple test that doesn't go far enough.