Live data from Hacker News

GPT-4

openai.com

591–600 of 1001 posts

Re: GPT-4

#591
post #218

A class of problem that GPT-4 appears to still really struggle with is variants of common puzzles. For example: >Suppose I have a cabbage, a goat and a lion, and I need to get them across a river. I have a boat that can only carry myself and a single other item. I am not allowed to leave the cabbage and lion alone together, and I am not allowed to leave the lion and goat alone together. How can I safely get all three…

I noticed it does get a "theory of mind" question that it used to fail, so it has indeed improved:

> “Meltem and Can are in the park. Can wanted to buy ice cream from the ice cream van but he hasn’t got any money. The ice cream man tells her that he will be there all afternoon. Can goes off home to get money for ice cream. After that, ice cream man tells Meltem that he changed his mind and he is going to drive to the school yard and sell ice cream there. Ice cream man sees Can on the road of the school and he also tells him that he is going to the school yard and will sell ice cream there. Meltem goes to Can’s house but Can is not there. His mom tells her that he has gone to buy ice cream. Where does Meltem think Can has gone, to the school or to the park?"

This is from some research in the 80s

Re: GPT-4

#592
post #218

A class of problem that GPT-4 appears to still really struggle with is variants of common puzzles. For example: >Suppose I have a cabbage, a goat and a lion, and I need to get them across a river. I have a boat that can only carry myself and a single other item. I am not allowed to leave the cabbage and lion alone together, and I am not allowed to leave the lion and goat alone together. How can I safely get all three…

Honest question: why would you bother expecting it to solve puzzles? It's not a use case for GPT.

the impressive thing is that GPT has unexpectedly outgrown its use case and it can answer a wide variety of puzzles, this is a little mindblowing for language research

Re: GPT-4

#593
post #218

A class of problem that GPT-4 appears to still really struggle with is variants of common puzzles. For example: >Suppose I have a cabbage, a goat and a lion, and I need to get them across a river. I have a boat that can only carry myself and a single other item. I am not allowed to leave the cabbage and lion alone together, and I am not allowed to leave the lion and goat alone together. How can I safely get all three…

Honest question: why would you bother expecting it to solve puzzles? It's not a use case for GPT.

Solving puzzles seems kind of close to their benchmarks, which are standardized tests.

Re: GPT-4

#595

After watching the demos I'm convinced that the new context length will have the biggest impact. The ability to dump 32k tokens into a prompt (25,000 words) seems like it will drastically expand the reasoning capability and number of use cases. A doctor can put an entire patient's medical history in the prompt, a lawyer an entire case history, etc. As a professional...why not do this? There's a non-zero chance that i…

> As a professional...why not do this?

Because your clients do not allow you to share their data with third parties?

Re: GPT-4

#596
post #218

A class of problem that GPT-4 appears to still really struggle with is variants of common puzzles. For example: >Suppose I have a cabbage, a goat and a lion, and I need to get them across a river. I have a boat that can only carry myself and a single other item. I am not allowed to leave the cabbage and lion alone together, and I am not allowed to leave the lion and goat alone together. How can I safely get all three…

You asked a trick question. The vast majority of people would make the same mistake. So your example arguably demonstrates that ChatGPT is close to an AGI, since it made the same mistake I did. I'm curious: When you personally read a piece of text, do you intensely hyperfocus on every single word to avoid being wrong-footed? It's just that most people read quickly wihch alowls tehm ot rdea msispeleled wrdos. I never…

It seems like GPT-4 does something that's similar to what we do too yes!

But when people do this mistake - just spit out an answer because we think we recognize this situation - in colloquial language this behavior is called "answering without thinking(!)".

If you "think" about it, then you activate some much more careful, slower reasoning. In this mode you can even do meta reasoning, you realize what you need to know in order to answer, or you maybe realize that you have to think very hard to get the right answer. Seems like we're veering into Kahneman's "Thinking fast and thinking slow" here.

Re: GPT-4

#597

After watching the demos I'm convinced that the new context length will have the biggest impact. The ability to dump 32k tokens into a prompt (25,000 words) seems like it will drastically expand the reasoning capability and number of use cases. A doctor can put an entire patient's medical history in the prompt, a lawyer an entire case history, etc. As a professional...why not do this? There's a non-zero chance that i…

> As a professional...why not do this? Because your clients do not allow you to share their data with third parties?

What's the difference between entering in an anonymized patient history into ChatGPT and, say, googling their symptoms?

Re: GPT-4

#598

I'll be finishing my interventional radiology fellowship this year. I remember in 2016 when Geoffrey Hinton said, "We should stop training radiologists now," the radiology community was aghast and in-denial. My undergrad and masters were in computer science, and I felt, "yes, that's about right." If you were starting a diagnostic radiology residency, including intern year and fellowship, you'd just be finishing now.…

If you are in the US. It is more important to have the legal paperwork, than to be factually correct. The medical cartels always will get their cut.

Eventually it's going to be cheap enough to drop by Tijuana for $5 MRI that even the cartel has to react.

Also, even within the US framework, there's pressure. A radiologist can rubberstamp 10x as many reports with AI-assistance. That doesn't eliminate radiology, but it eliminates 90% of the radiologists we're training.

Re: GPT-4

#599
This is a pretty exciting moment in tech. Pretty much like clockwork, every decade or so since the broad adoption of electricity there’s been a new society changing technical innovation. One could even argue it goes back to the telegraph in the 1850s.

With appropriate caveats and rough dating, here’s a list I can think of:

    Electric lights in 1890s, 
    Radio communication in the mid 00’s,
    Telephones in the mid 10s,
    Talking Movies in the mid 20s,
    Commercial Radio in the mid 30s,
    Vinyl records in the mid 40s,
    TVs in the mid 50s,
    Computers in the mid 60s,
    The microchip/integrated circuit in the mid 70s, 
    The GUI in the mid 80s,
    Internet/Web in the mid 90s, 
    Smartphone in the mid 2000s,
    Streaming video/social networking in the mid 2010s, 
And now AI. This is a big one.

Re: GPT-4

#600

After watching the demos I'm convinced that the new context length will have the biggest impact. The ability to dump 32k tokens into a prompt (25,000 words) seems like it will drastically expand the reasoning capability and number of use cases. A doctor can put an entire patient's medical history in the prompt, a lawyer an entire case history, etc. As a professional...why not do this? There's a non-zero chance that i…

I must have missed the part when it started doing anything algorithmically. I thought it’s applied statistics, with all the consequences of that. Still a great achievement and super useful tool, but AGI claims really seem exaggerated.
Post reply on HN