Live data from Hacker News

GPT-4

openai.com

461–470 of 1001 posts

Re: GPT-4

#461
post #218

A class of problem that GPT-4 appears to still really struggle with is variants of common puzzles. For example: >Suppose I have a cabbage, a goat and a lion, and I need to get them across a river. I have a boat that can only carry myself and a single other item. I am not allowed to leave the cabbage and lion alone together, and I am not allowed to leave the lion and goat alone together. How can I safely get all three…

> I am not allowed to leave the cabbage and lion alone together, and I am not allowed to leave the lion and goat alone together. How can I safely get all three across?

What if you phrase it as a cabbage, vegan lion and a meat eating goat...

Re: GPT-4

#462
post #218

A class of problem that GPT-4 appears to still really struggle with is variants of common puzzles. For example: >Suppose I have a cabbage, a goat and a lion, and I need to get them across a river. I have a boat that can only carry myself and a single other item. I am not allowed to leave the cabbage and lion alone together, and I am not allowed to leave the lion and goat alone together. How can I safely get all three…

You asked a trick question. The vast majority of people would make the same mistake. So your example arguably demonstrates that ChatGPT is close to an AGI, since it made the same mistake I did. I'm curious: When you personally read a piece of text, do you intensely hyperfocus on every single word to avoid being wrong-footed? It's just that most people read quickly wihch alowls tehm ot rdea msispeleled wrdos. I never…

> Even after I pointed this mistake out, it repeated exactly the same proposed plan.

The vast majority of people might make the mistake once, yes, but would be able to reason better once they had the trick pointed out them. Imo it is an interesting anecdote that GPT-4 can't adjust its reasoning around this fairly simple trick.

Re: GPT-4

#463
All: our poor server is smoking today* so I've had to reduce the page size of comments. There are 1500+ comments in this thread but if you want to read more than a few dozen you'll need to page through them by clicking the More link at the bottom. I apologize!

Also, if you're cool with read-only access, just log out (edit: or use an incognito tab) and all will be fast again.

* yes, HN still runs on one core, at least the part that serves logged-in requests, and yes this will all get better someday...it kills me that this isn't done yet but one day you will all see

Re: GPT-4

#464
post #218

A class of problem that GPT-4 appears to still really struggle with is variants of common puzzles. For example: >Suppose I have a cabbage, a goat and a lion, and I need to get them across a river. I have a boat that can only carry myself and a single other item. I am not allowed to leave the cabbage and lion alone together, and I am not allowed to leave the lion and goat alone together. How can I safely get all three…

GPT 4 does not know that when you are on a boat it means the items on the land side are together.

I remember this question as a 7 year old and when the question was told to me, the same information was omitted.

Edit: just realized you flipped the scenario. Yes it seems like a case of pattern matching to a known problem. I think if you changed the variables to A, B, and C and gave a much longer description and more accurate conditions, it would have a different response.

Re: GPT-4

#465

The fact it can read pictures is the real killer feature here. Now you can give it invoices to file, memo to index, pics to sort and chart to take actions on. And to think we are at the nokia 3310 stage. What's is the iphone of AI going to look like?

I really hope we get 15 years of iPhone-like progress! Everything just seems like it's moving so fast right now...

Re: GPT-4

#467
Genuinely surprised by the positive reaction about how exciting this all is.

You ever had to phone a large business to try and sort something out, like maybe a banking error, and been stuck going through some nonsense voice recognition menu tree that doesn't work? Well imagine chat GPT with a real time voice and maybe a fake, photorealistic 3D avatar and having to speak to that anytime you want to speak to a doctor, sort out tax issues, apply for a mortgage, apply for a job, etc. Imagine Reddit and hacker news just filled with endless comments from AIs to suit someone's agenda. Imagine never reading another news article written by a real person. Imagine facts becoming uncheckable since sources can no longer be verified. Wikipedia just becomes a mass of rewrites of AI over AI. Imagine when Zoom lets you send an AI persona to fill in for you at a meeting.

I think this is all very, very bad. I'm not saying it should be stopped, I mean it can't, but I feel a real dread thinking of where this is going. Hope I am wrong.

Re: GPT-4

#468
post #218

A class of problem that GPT-4 appears to still really struggle with is variants of common puzzles. For example: >Suppose I have a cabbage, a goat and a lion, and I need to get them across a river. I have a boat that can only carry myself and a single other item. I am not allowed to leave the cabbage and lion alone together, and I am not allowed to leave the lion and goat alone together. How can I safely get all three…

A funny variation on this kind of over-fitting to common trick questions - if you ask it which weighs more, a pound of bricks or a pound of feathers, it will correctly explain that they actually weigh the same amount, one pound. But if you ask it which weighs more, two pounds of bricks or a pound of feathers, the question is similar enough to the trick question that it falls into the same thought process and contorts an explanation that they also weigh the same because two pounds of bricks weighs one pound.

Re: GPT-4

#469
post #342

Let's check out the paper for actual tech details! > Given both the competitive landscape and the safety implications of large-scale models like GPT-4, this report contains no further details about the architecture (including model size), hardware, training compute, dataset construction, training method, or similar. - Open AI

Someone needs to hack into them and release the parameters and code. This knowledge is too precious to be kept secret.

Re: GPT-4

#470
post #218

A class of problem that GPT-4 appears to still really struggle with is variants of common puzzles. For example: >Suppose I have a cabbage, a goat and a lion, and I need to get them across a river. I have a boat that can only carry myself and a single other item. I am not allowed to leave the cabbage and lion alone together, and I am not allowed to leave the lion and goat alone together. How can I safely get all three…

Honest question: why would you bother expecting it to solve puzzles? It's not a use case for GPT.

Being able to come up with solutions to assigned tasks that don't have a foundation in something that's often referenced and can be memorized is basically the most valuable use case for AI.

Simple example: I want to tell my robot to go get my groceries that includes frozen foods, pick up my dry cleaning before the store closes, and drive my dog to her grooming salon but only if it's not raining and the car is charged. The same sort of logic is needed to accomplish all this without my frozen food spoiling and wasting a salon visit and making sure I have my suit for an interview tomorrow.

Post reply on HN