Live data from Hacker News

GPT-4o's Memory Breakthrough – Needle in a Needlestack

nian.llmonpy.ai

241–250 of 256 posts

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#241
post #108
post #84

Earlier quoted context omitted.

I suppose the question then is - if you finetune on your own data (eg internal wiki) does it then retain the near-perfect recall? Could be a simpler setup than RAG for slow-changing documentation, especially for read-heavy cases.

"if you finetune on your own data (eg internal wiki) does it then retain the near-perfect recall" No, that's one of the primary reasons for RAG.

I think you are misunderstanding. This post is about new capabilities in GPT-4o. So the existing reasons for RAG may not hold for the new model.

Unless you have some evals showing that the previous results justifying RAG also apply to GPT-4o?

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#242
Meh still for a lot of stuff it simply lies.

Just today it lied to me about VRL language syntax, tryin to sell me some python stuff in there.

Senior ppl will often be able call out the bullshit, but I believe for junior ppl it will be very detrimental.

Nether the less amazing tool for d2d work if you can call out BS replies.

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#243

I just used it to compare two smaller legal documents and it completely hallucinated that items were present in one and not the other. It did this on three discrete sections of the agreements. Using ctrl-f I was able to see that they were identical in one another. Obviously this is a single sample but saying 90% seems unlikely. They were around ~80k tokens total.

I have the same feeling. I asked to find duplicates in a list of 6k items and it basically hallucinated the entire answer multiple times. Some times it finds some, but it interlaces the duplicates with other hallucinated items. I wasn't expecting it to get it right, cause I think this task is challenging with a fixed amount of attention heads. However, the answer seems much worse than Claude Opus or GPT-4.

Everyone is trying to use Language Models as Reasoning Models because the latter haven't been invented yet.

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#244
post #82
post #77

Earlier quoted context omitted.

Maybe if you tell it to pull the answer from a limerick instead of generally asking? Edit: Ok no, I tried giving it a whole bunch of hints, and it was just making stuff up that was completely unrelated. Even directly pointing it at the original dataset didn’t help.

Come on guys, it’s already far beyond superhuman if it’s able to do that and so quickly. So if it’s not able to do that, what’s the big deal? If you’re asking for AG.I., then it seems that the model performs beyond it in these areas.

We were mainly trying to determine if there was a reasonable chance that the model was trained on a certain dataset, nothing else.

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#245
post #164

We are all so majorly f*d. The general public does not know nor understand this limitation. At the same time OpenAI is selling this a a tutor for your kids. Next it will be used to test those same kids. Who is going to prevent this from being used to pick military targets (EU law has an exemption for military of course) or make surgery decisions?

I’ve had coworkers suggest a technical solution that was straight up fabricated by an LLM and made no sense. More competent people realise this limitation of the models and can use them wisely. Unfortunately I expect to see the former spread.

We had a manager joining last year which, on their first days, created MRs for existing code bases, wrote documents on new processes and gave advice on current problems we were facing. Everything was created by LLMs and plain bullshit. Fortunately were able to convince the higher ups that this person was an imposter and we got rid of them.

I really hope that these type of situations won't increase because the mental strain that put on some people in the org is not sustainable in the long run.

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#246

Earlier quoted context omitted.

> a sophisticated approach to their defense A euphemism for apartheid and oppression? > sources are rife with bias What's biased about terming autonomous weapons as "AI"? Or, sounding alarm over dystopian surveillance enabled by AI? > The nuance matters. Like Ben Gurion terming Lehi "freedom fighters" as terrorists? And American Jewish intellectuals back then calling them fascists? > The history matters... After that…

I will spend no more than two comments on this issue. Most people have already made up their minds. There is little I can do about that, but perhaps someone else might see this and think twice. Personally, I have spent many thousands of hours on this topic. I have Palestinian relatives and have visited the Middle East. I have Arab friends there, both Christian and Muslim, whom I would gladly protect with my life. I a…

The total of these two comments make no objective claims, rather says there are nuances, and complexities. But in all these complexity they are sure that Israel is right in their actions. Bipartisan support is on shared values, supposedly. Not so surprisingly, it even has a > I have friends paragraph.

I got to say this is a pretty masterful deceit.

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#247
post #61
post #57

Earlier quoted context omitted.

sort alice.txt | diff - That's not a task for an LLM

Asking students to write an essay about Napoleon isn't something we do because we need essays about Napoleon - the point is it's a test of capabilities.

My point was more so that this task is so trivial, that's it's not testing the model's ability to distinguish contextual nuances, which would supposedly be the intention.

The idea presented elsewhere in this thread about using an unpublished novel and then asking questions about the plot is sort of the ideal test in this regard, and clearly on the other end of the spectrum in terms of a design that's testing actual "understanding".

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#248

Earlier quoted context omitted.

Why? There are many other types of AI or statistical methods that are easier, faster and cheaper to use not to mention better suited and far more accurate. Militaries have been employing statisticians since WWII to pick targets (and for all kinds of other things) this is just current-thing x2 so it’s being used to whip people into a frenzy.

It can do limited battlefield reasoning where a remote pilot has significant latency. Call these LLMs stupid all you want but on focused tasks they can reason decently enough. And better than any past tech.

> “Call these LLMs stupid all you want but…”

Make defensive comments in response to LLM skepticism all you want— there are still precisely zero (0) reasons to believe they’ll make a quantum leap towards human-level reasoning any time soon.

The fact that they’re much better than any previous tech is irrelevant when they’re still so obviously far from competent in so many important ways.

To allow your technological optimism to convince you to that this very simple and very big challenge is somehow trivial and that progress will inevitably continue apace is to engage in the very drollest form of kidding yoursef.

Pre-space travel, you could’ve climbed the tallest mountain on earth and have truthfully claimed that you were closer to the moon than any previous human, but that doesn’t change the fact that the best way to actually get to the moon is to climb down from the mountain and start building a rocket.

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#249
post #230

Earlier quoted context omitted.

You said it mate. I feel bad for folks who turn away from this technology. If they persist... They will be so confused why they get repeatedly lapped. I wrote a working machine vision project in 2 days with these toys. Key word: working... Not hallucinated. Actually working. Very useful.

I just don't understand why AI is so polarising on a technology website. OpenAI have even added a feature to make the completions from GPT near-deterministic (by specifying a seed). It seems that no matter what AI companies do, there will be a vocal minority shouting that it's worthless.

It is baffling where the polarization comes in.

The idea that we argue about safety... Seems reasonable to me.

The argument about its usefulness or capability at all? I dunno... That slider bar sure is in a wierd spot... I feel ya.

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#250

I just used it to compare two smaller legal documents and it completely hallucinated that items were present in one and not the other. It did this on three discrete sections of the agreements. Using ctrl-f I was able to see that they were identical in one another. Obviously this is a single sample but saying 90% seems unlikely. They were around ~80k tokens total.

Interesting, because the (at least the official) context window of GPT-4o is 128k.
Post reply on HN