Live data from Hacker News

GPT-4o's Memory Breakthrough – Needle in a Needlestack

nian.llmonpy.ai

151–160 of 256 posts

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#151
post #131
post #52

I am in England, do US users have access to memory features? ( Also do you ahve access to voice customisation yet? Thanks

I am in England, on the 'Team Plan'* and got access to memory this week. * https://openai.com/index/introducing-chatgpt-team/

Thank you!

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#152
post #99

Earlier quoted context omitted.

Yeah I asked for an estimate of the percentage of the US population that lives in the DMV area (DC, Maryland, Virginia) and it was off by 50% of the actual answer, which I only realized when I realized I shouldn’t trust its estimate for anything important

Those models still can't reliably do arithmetic, so how could it possibly know that number unless it's a commonly repeated fact? Also: would you expect random people to fare any better?

It used web search (RAG over the entire web) and analysis (math tool) and still came up with the wrong answer.

It has done more complex things for me than this and, sometimes, gotten it right.

Yes, it’s supposed to be able to do this.

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#153

This is based on a limericks dataset published in 2021. https://zenodo.org/records/5722527 I think it very likely that gpt-4o was trained on this. I mean, why would you not? Innnput, innnput, Johnny five need more tokens. I wonder why the NIAN team don't generate their limericks using different models, and check to make sure they're not in the dataset? Then you'd know the models couldn't possibly be trained on them.

Why not just generate complete random stuff and ask it to find stuff in that?

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#154

Earlier quoted context omitted.

Note that the IDF explicitly denied that story: https://www.idf.il/en/mini-sites/hamas-israel-war-24/all-art... Probably this is due to confusion over what the term "AI" means. If you do some queries on a database, and call yourself a "data scientist", and other people who call themselves data scientists do some AI, does that mean you're doing AI? For left wing journalists who want to undermine the Israelis (the stor…

The IDF explicitly deny a lot of things, which turn out to be true.

Just like… (checks notes)… oh yeah every government on the planet.

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#156
post #123

Earlier quoted context omitted.

> OpenAI is selling this a a tutor for your kids. The Diamond Age.

That's what I find most offensive about the use of LLMs in education: it can readily produce something in the shape of a logical argument, without actually being correct. I'm worried that a generation might learn that that's good enough.

a generation of consultants is already doing that- look at the ruckus around PWC etc in Australia. Hell, look at the folks supposedly doing diligence on Enron. This is not new. People lie, fib and prevaricate. The fact the machines trained on our actions do the same thing should not come as a shock. If anything it strikes me as the uncanny valley of truthiness.

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#157

This is based on a limericks dataset published in 2021. https://zenodo.org/records/5722527 I think it very likely that gpt-4o was trained on this. I mean, why would you not? Innnput, innnput, Johnny five need more tokens. I wonder why the NIAN team don't generate their limericks using different models, and check to make sure they're not in the dataset? Then you'd know the models couldn't possibly be trained on them.

Why not just generate complete random stuff and ask it to find stuff in that?

We have run that test.- generate random string(not by llm) names of values- ask the llm to do math (algebra) using those strings. Tests logic, 100% not in the data set GPT2 was like 50% accurate, now we up around the 90%.

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#158

We also need a way to determine where a given response fits in the universe of responses - is it an “average” answer or a really good one

If you have an evaluation function which does this accurately and generalizes, you pretty much already have have AGI.

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#160

We are all so majorly f*d. The general public does not know nor understand this limitation. At the same time OpenAI is selling this a a tutor for your kids. Next it will be used to test those same kids. Who is going to prevent this from being used to pick military targets (EU law has an exemption for military of course) or make surgery decisions?

> or make surgery decisions? Analyzing surgical field... Identified: open chest cavity, exposed internal organs Organs appear gooey, gelatinous, translucent pink Comparing to database of aquatic lifeforms... 93% visual match found: Psychrolutes marcidus, common name "blobfish" Conclusion: Blobfish discovered inhabiting patient's thoracic cavity Recommended action: Attempt to safely extract blobfish without damaging o…

[dead]
Post reply on HN