Live data from Hacker News

GPT-4o's Memory Breakthrough – Needle in a Needlestack

nian.llmonpy.ai

161–170 of 256 posts

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#162

We are all so majorly f*d. The general public does not know nor understand this limitation. At the same time OpenAI is selling this a a tutor for your kids. Next it will be used to test those same kids. Who is going to prevent this from being used to pick military targets (EU law has an exemption for military of course) or make surgery decisions?

> OpenAI is selling this a a tutor for your kids. The Diamond Age.

[dead]

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#163

We are all so majorly f*d. The general public does not know nor understand this limitation. At the same time OpenAI is selling this a a tutor for your kids. Next it will be used to test those same kids. Who is going to prevent this from being used to pick military targets (EU law has an exemption for military of course) or make surgery decisions?

OpenAI is. Their TOS says don't use it for that kind of shit. https://openai.com/policies/usage-policies/

That's the license for the public service. Nothing prevents them from selling it as a separate package deal to an army.

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#164

We are all so majorly f*d. The general public does not know nor understand this limitation. At the same time OpenAI is selling this a a tutor for your kids. Next it will be used to test those same kids. Who is going to prevent this from being used to pick military targets (EU law has an exemption for military of course) or make surgery decisions?

I’ve had coworkers suggest a technical solution that was straight up fabricated by an LLM and made no sense. More competent people realise this limitation of the models and can use them wisely. Unfortunately I expect to see the former spread.

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#165

We are all so majorly f*d. The general public does not know nor understand this limitation. At the same time OpenAI is selling this a a tutor for your kids. Next it will be used to test those same kids. Who is going to prevent this from being used to pick military targets (EU law has an exemption for military of course) or make surgery decisions?

People aren't dumb. They'll catch on pretty quick that this thing is BS'ing them.

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#166
post #148

Earlier quoted context omitted.

> Obviously this is a single sample but saying 90% seems unlikely. This is such an anti-intellectual comment to make, can't you see that? You mention "sample" so you understand what statistics is, then in the same sentence claim 90% seems unlikely with a sample size of 1. The article has done substantial research

That fact that it has some statistically significant performance is irrelevant and difficult to evaluate for most people. He's a much simpler and correct description that almost everyone can understand: it fucks up constantly. Getting something wrong even once can make it useless for most people. No amount of pedantry will change this reality.

What on earth? The experimental research demonstrates that it doesn't "fuck up constantly", you're just making things up. The various performance metrics people around the world to measure and compare model performance is not irrelevant because you, some random internet commenter, claim so without any evidence.

This isn't pedantry, it's science.

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#167

Someone needs to come up with a "synthesis from haystack" test that tests not just retrieval but depth of understanding, connections, abstractions across diverse information. When a person reads a book, they have an "overall intuition" about it. We need some way to quantify this. Needle in haystack tests feel like a simple test that doesn't go far enough.

There is no understanding, it can't do this.

GPT4o still can't do the intersection of two different ideas that are not in the training set. It can't even produce random variations on the intersection of two different ideas.

Further though, we shouldn't expect the model to do this. It is not fair to the model and its actual usefulness and how amazing what the models can do with zero understanding. To believe the model understands is to fool yourself.

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#168
post #125

Earlier quoted context omitted.

AI is already being used for picking targets in warzones - https://theconversation.com/israel-accused-of-using-ai-to-ta... . LLM's will of course also be used, due to their convenience and superficial 'intelligence', and because of the layer of deniability creating a technical substrate between soldier and civilian victim provides - as has happened for two decades with drones.

Note that the IDF explicitly denied that story: https://www.idf.il/en/mini-sites/hamas-israel-war-24/all-art... Probably this is due to confusion over what the term "AI" means. If you do some queries on a database, and call yourself a "data scientist", and other people who call themselves data scientists do some AI, does that mean you're doing AI? For left wing journalists who want to undermine the Israelis (the stor…

The "independent examinations" is doing a heavy lift there.

At most charitable, that means a person is reviewing all data points before approval.

At least charitable, that means a person is clicking approved after glancing at the values generated by the system.

The press release doesn't help clarify that one way or the other.

If you want to read thoughts by the guy who was in charge of building and operating the automated intelligence system, he wrote a book: https://www.amazon.com/Human-Machine-Team-Artificial-Intelli...

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#169

We are all so majorly f*d. The general public does not know nor understand this limitation. At the same time OpenAI is selling this a a tutor for your kids. Next it will be used to test those same kids. Who is going to prevent this from being used to pick military targets (EU law has an exemption for military of course) or make surgery decisions?

This is just doomerism. Even though this model is slightly better than the previous, using an LLM for high risk tasks like healthcare and picking targets in military operations still feels very far away. I work in healthcare tech in a European country and yes we use AI for image recognition on x-rays, retinas etc but these are fundamentally completely different models than a LLM. Using LLMs for picking military targe…

I also work in healthtech, and nearly every vendor we’ve evaluated in the last 12 months has tacked on ChatGPT onto their feature set as an “AI” improvement. Some of the newer startup vendors are entirely prompt engineering with a fancy UI. We’ve passed on most of these but not all. And these companies have clients, real world case studies. It’s not just not very far away, it is actively here.
Post reply on HN