Live data from Hacker News

GPT-4o's Memory Breakthrough – Needle in a Needlestack

nian.llmonpy.ai

131–140 of 256 posts

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#131
post #52

I am in England, do US users have access to memory features? ( Also do you ahve access to voice customisation yet? Thanks

I am in England, on the 'Team Plan'* and got access to memory this week.

* https://openai.com/index/introducing-chatgpt-team/

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#132
post #126

Earlier quoted context omitted.

This is just doomerism. Even though this model is slightly better than the previous, using an LLM for high risk tasks like healthcare and picking targets in military operations still feels very far away. I work in healthcare tech in a European country and yes we use AI for image recognition on x-rays, retinas etc but these are fundamentally completely different models than a LLM. Using LLMs for picking military targe…

>Using LLMs for picking military targets is just absurd. In the future I guess the future is now then: https://www.theguardian.com/world/2023/dec/01/the-gospel-how... Excerpt: >Aviv Kochavi, who served as the head of the IDF until January, has said the target division is “powered by AI capabilities” and includes hundreds of officers and soldiers. >In an interview published before the war, he said it was “a machine th…

nothing in this says they used an LLM

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#133
post #125

Earlier quoted context omitted.

This is just doomerism. Even though this model is slightly better than the previous, using an LLM for high risk tasks like healthcare and picking targets in military operations still feels very far away. I work in healthcare tech in a European country and yes we use AI for image recognition on x-rays, retinas etc but these are fundamentally completely different models than a LLM. Using LLMs for picking military targe…

AI is already being used for picking targets in warzones - https://theconversation.com/israel-accused-of-using-ai-to-ta... . LLM's will of course also be used, due to their convenience and superficial 'intelligence', and because of the layer of deniability creating a technical substrate between soldier and civilian victim provides - as has happened for two decades with drones.

Note that the IDF explicitly denied that story:

https://www.idf.il/en/mini-sites/hamas-israel-war-24/all-art...

Probably this is due to confusion over what the term "AI" means. If you do some queries on a database, and call yourself a "data scientist", and other people who call themselves data scientists do some AI, does that mean you're doing AI? For left wing journalists who want to undermine the Israelis (the story originally appeared in the Guardian) it'd be easy to hear what you want to hear from your sources and conflate using data with using AI. This is the kind of blurring that happens all the time with apparently technical terms once they leave the tech world and especially once they enter journalism.

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#134

I just used it to compare two smaller legal documents and it completely hallucinated that items were present in one and not the other. It did this on three discrete sections of the agreements. Using ctrl-f I was able to see that they were identical in one another. Obviously this is a single sample but saying 90% seems unlikely. They were around ~80k tokens total.

What you are asking an llm to do here makes no sense.

You might be right but I've lost count of the number of startups I've heard of trying to do this for legal documents.

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#135
post #132
post #126

Earlier quoted context omitted.

>Using LLMs for picking military targets is just absurd. In the future I guess the future is now then: https://www.theguardian.com/world/2023/dec/01/the-gospel-how... Excerpt: >Aviv Kochavi, who served as the head of the IDF until January, has said the target division is “powered by AI capabilities” and includes hundreds of officers and soldiers. >In an interview published before the war, he said it was “a machine th…

nothing in this says they used an LLM

I guess he must have hallucinated that it was about LLMs

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#136
post #125

Earlier quoted context omitted.

This is just doomerism. Even though this model is slightly better than the previous, using an LLM for high risk tasks like healthcare and picking targets in military operations still feels very far away. I work in healthcare tech in a European country and yes we use AI for image recognition on x-rays, retinas etc but these are fundamentally completely different models than a LLM. Using LLMs for picking military targe…

AI is already being used for picking targets in warzones - https://theconversation.com/israel-accused-of-using-ai-to-ta... . LLM's will of course also be used, due to their convenience and superficial 'intelligence', and because of the layer of deniability creating a technical substrate between soldier and civilian victim provides - as has happened for two decades with drones.

Why? There are many other types of AI or statistical methods that are easier, faster and cheaper to use not to mention better suited and far more accurate. Militaries have been employing statisticians since WWII to pick targets (and for all kinds of other things) this is just current-thing x2 so it’s being used to whip people into a frenzy.

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#137

We are all so majorly f*d. The general public does not know nor understand this limitation. At the same time OpenAI is selling this a a tutor for your kids. Next it will be used to test those same kids. Who is going to prevent this from being used to pick military targets (EU law has an exemption for military of course) or make surgery decisions?

This is just doomerism. Even though this model is slightly better than the previous, using an LLM for high risk tasks like healthcare and picking targets in military operations still feels very far away. I work in healthcare tech in a European country and yes we use AI for image recognition on x-rays, retinas etc but these are fundamentally completely different models than a LLM. Using LLMs for picking military targe…

> picking targets in military operations

I'm 100% on the side of Israel having the right to defend itself, but as I understand it, they are already using "AI" to pick targets, and they adjust the threshold each day to meet quotas. I have no doubt that some day they'll run somebody's messages through chat gpt or similar and get the order: kill/do not kill.

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#138

This is based on a limericks dataset published in 2021. https://zenodo.org/records/5722527 I think it very likely that gpt-4o was trained on this. I mean, why would you not? Innnput, innnput, Johnny five need more tokens. I wonder why the NIAN team don't generate their limericks using different models, and check to make sure they're not in the dataset? Then you'd know the models couldn't possibly be trained on them.

NIAN is a very cool idea, but why not simply translate it into N different languages (you even can mix services, e.g. deepl/google translate/LLMs themselves) and ask about them that way?

Re: GPT-4o's Memory Breakthrough – Needle in a Needlestack

#140
post #55

This is based on a limericks dataset published in 2021. https://zenodo.org/records/5722527 I think it very likely that gpt-4o was trained on this. I mean, why would you not? Innnput, innnput, Johnny five need more tokens. I wonder why the NIAN team don't generate their limericks using different models, and check to make sure they're not in the dataset? Then you'd know the models couldn't possibly be trained on them.

I tested the LLMs to make sure they could not answer the questions unless the limerick was given to them. Other than 4o, they do very badly on this benchmark, so I don't think the test is invalidated by their training.

It would be interesting to know how it acts if you ask it about one that isn't present, or even lie to it (e.g. take a limerick that is present but change some words and ask it to complete it)

Maybe some models hallucinate or even ignore your mistake vs others correcting it (depending on the context ignoring or calling out the error might be the more 'correct' approach)

Using limericks is a very nifty idea!

Post reply on HN