Live data from Hacker News

OpenAI researchers warned board of AI breakthrough ahead of CEO ouster

reuters.com

251–260 of 1001 posts

Re: OpenAI researchers warned board of AI breakthrough ahead of CEO ouster

#251
post #3

> warning of a powerful artificial intelligence discovery that they said could threaten humanity, > Given vast computing resources, the new model was able to solve certain mathematical problems, [..] Though only performing math on the level of grade-school students, acing such tests made researchers very optimistic about Q*’s future success, the source said. I somehow expected a bit more.

I don't really understand this. Aren't LLMs already performing at near-expert level on "certain mathematical problem" benchmarks? For example, over a year ago MINERVA from Google [1] got >50% on the MATH dataset, a set of competition math problems. These are not easy problems. From the MATH dataset paper: > We also evaluated humans on MATH, and found that a computer science PhD student who does not especially like ma…

Every example on that web page is a trivial textbook problem. That shows memorization of training set and textual pattern matching.

Re: OpenAI researchers warned board of AI breakthrough ahead of CEO ouster

#252
post #3

> warning of a powerful artificial intelligence discovery that they said could threaten humanity, > Given vast computing resources, the new model was able to solve certain mathematical problems, [..] Though only performing math on the level of grade-school students, acing such tests made researchers very optimistic about Q*’s future success, the source said. I somehow expected a bit more.

This makes some sense to me. My experience with GPT is that it is capable of straightforward logical inference, but not more inspired thinking. It lacks the ability for a “eureka moment”. All complex inference it appears to have is a result of its training set. It is incapable of solving certain kinds of logic problems that a child would be able to solve. As an example, take the wolf, goat, and cabbage problem, but c…

Me: Solve the riddle: You have three fantastic animals: Aork, Bork, and Cork. If left unattended, Aork would eat Bork, and Bork would eat Cork. When you are with them, they behave and don't eat each other. You travel with these three animals and encounte a river with a boat. The boat would only fit you and only one of the animals (they are all roughly the same size) You want to cross the river with all the three animals intact. Solve the riddle, thinking step-by-step

ChatGPT: To solve this riddle, you must ensure that no animal gets eaten while you are transferring them across the river. Here's a step-by-step solution:

First Crossing: Take Bork across the river and leave him there. Return alone to the original side. At this point, Aork and Cork are on one side, and Bork is on the other. Bork cannot eat Cork because you are with Cork.

Second Crossing: Take Aork across the river. Now, Aork and Bork are on the far side, and Cork is on the original side. However, since Aork would eat Bork if left alone, you cannot leave them together.

Third Crossing: Take Bork back with you to the original side. Leave Bork there and take Cork with you across the river. Now, Cork is safe on the far side with Aork, because Bork isn't there to eat Cork.

Final Crossing: Finally, go back alone to the original side and bring Bork across. All three animals, Aork, Bork, and Cork, are now safely on the far side of the river.

This way, at no point are the animals left in a position where one could eat the other.

I think it does fine?

Re: OpenAI researchers warned board of AI breakthrough ahead of CEO ouster

#253
This matches far better with the board's letter re: firing Sam than a simple power struggle or disagreement on commercialisation. Seeing a huge breakthrough and then not reporting it to the board, who then find out via staff letter certainly counts as a "lack of candour"....

As an aside, assuming a doomsday scenario, how long can secrets like this stay outside of the hands of bad actors? On a scale of 1 to enriched uranium

Re: OpenAI researchers warned board of AI breakthrough ahead of CEO ouster

#254
post #44
post #3

> warning of a powerful artificial intelligence discovery that they said could threaten humanity, > Given vast computing resources, the new model was able to solve certain mathematical problems, [..] Though only performing math on the level of grade-school students, acing such tests made researchers very optimistic about Q*’s future success, the source said. I somehow expected a bit more.

Being legitimately good at reasoning when it comes to numbers is a new emergent behavior. Reasoning about numbers isn't something that exists in "idea space" where all the vectorized tokens exist.

> Reasoning about numbers isn't something that exists in "idea space" where all the vectorized tokens exist.

To be fair, we don't know for certain that this is the case.

Re: OpenAI researchers warned board of AI breakthrough ahead of CEO ouster

#255
post #60

Weirdly enough, this sort of lines up with a theory posted on 4chan 4 days ago. The gist being that if the version is formally declared AGI, it can't be licensed to Microsoft and others for commercial gain. As a result Altman wants it not to be called AGI, other board members do. Archived link below. NB THIS IS 4CHAN - THERE WILL OFFENSIVE LANGUAGE. https://archive.ph/sFMXa

> be me

> be strong agi

Re: OpenAI researchers warned board of AI breakthrough ahead of CEO ouster

#256
post #3

> warning of a powerful artificial intelligence discovery that they said could threaten humanity, > Given vast computing resources, the new model was able to solve certain mathematical problems, [..] Though only performing math on the level of grade-school students, acing such tests made researchers very optimistic about Q*’s future success, the source said. I somehow expected a bit more.

OpenAI already benchmarks their GPTs on leetcode problems and even includes a Codeforces rating. It is not impressive at all and there's almost no progress from GPT 2 to 4. I agree, why does this grade school math problem matter if the model can't solve problems that are very precisely stated and have a very narrow solution space (at least more narrow than some vague natural language instruction)?

[deleted]

Re: OpenAI researchers warned board of AI breakthrough ahead of CEO ouster

#257
post #68

Earlier quoted context omitted.

Like it will never be weaponized? If not by the US, by nobody else on earth?

If it is sentient, I think weaponizing it would require convincing it that being a weapon was in its interest. I struggle to find a way to reason to that position without using emotional content and I presume a sentient AGI will not have emotion.

Personally I don’t think AGI is possible. I was thinking more of a model that will be able to do 100 times what you can do now. Like able to answer questions like: what is the most efficient way to attack this front considering this map and this data dump of weapons stats and patrol present, etc…

But even if it was sentient convincing it then shouldn’t be hard. Even the most brilliant people can be fooled and convinced of absurd things. Even when they think that their work can threaten the human race but will still continue and push for it.

Re: OpenAI researchers warned board of AI breakthrough ahead of CEO ouster

#259
post #61

Earlier quoted context omitted.

I've seen it gets tripped up if you mix something common with something uncommon. ChatGPT chokes on, for example: "Two trains on separate tracks, 30 miles from each other are approaching each other, each at a speed of 10 mph. How long before they crash into each other?"

Yes. Though I’d say that example is a bit mean (it’s a trick question) since the answer has expected type but whose actual answer is something like “don’t be stupid; they’re not even on the same track”. It’s like asking “if I add three apples to two apples, how many pears do I get?” and being surprised when the LLM says “5 pears”. It’s the sort of thing that’s avoidable if the question is written properly.

GPT-4 Turbo:

When is 2+5=10?

The equation 2+5=10 is never true in standard arithmetic, as 2 plus 5 always equals 7. However, in a different context or under a specific set of rules (such as a puzzle or a riddle where the symbols have different meanings), it could be made to appear true.

GPT-4(phind):

The equation 2 + 3 = 10 is true in a base-5 number system.

In base-5 (also known as quinary), the digits range from 0 to 4. When you add 2 and 3 in base-5, the result is 10. This is because in base-5, the number 5 is represented as 10.

Here's a simple breakdown of how this works:

    In base-5, the rightmost digit represents 5^0 (which is 1), the next digit to the left represents 5^1 (which is 5), the next digit represents 5^2 (which is 25), and so on.
    So when you add 2 and 3, you get 5. But since 5 is represented as 10 in base-5 (1 digit of 5, and 0 digits of 1), the result is 10.
Therefore, in base-5, 2 + 3 equals 10

Re: OpenAI researchers warned board of AI breakthrough ahead of CEO ouster

#260
post #61

Earlier quoted context omitted.

Mathematical logic doesn’t refer to ‘complex word problems with big numbers’. The reason that LLMs struggle (well, are hopeless currently) with mathematics problems is not because it involves arithmetic (although this is already enough of a struggle) but because it requires a conceptual understanding of the ideas rather than just being about manipulating language — a lot of the goings-on are hidden in the semantic wo…

I've seen it gets tripped up if you mix something common with something uncommon. ChatGPT chokes on, for example: "Two trains on separate tracks, 30 miles from each other are approaching each other, each at a speed of 10 mph. How long before they crash into each other?"

But if you make the input slightly more explicit:

"Two trains on different and separate tracks, 30 miles from each other are approaching each other, each at a speed of 10 mph. How long before they crash into each other?"

...it spots the trick: https://chat.openai.com/share/ee68f810-0c12-4904-8276-a4541d...

Likewise, if you add emphasis it understands too:

"Two trains on separate tracks, 30 miles from each other are approaching each other, each at a speed of 10 mph. How long before they crash into each other?"

https://chat.openai.com/share/acafbe34-8278-4cf7-80bb-76858c...

Not to anthropomorphize, but perhaps it's not necessarily missing the trick, it just assumes that you're making a mistake.

Post reply on HN