Live data from Hacker News

GPT-4

openai.com

491–500 of 1001 posts

Re: GPT-4

#491

Genuinely surprised by the positive reaction about how exciting this all is. You ever had to phone a large business to try and sort something out, like maybe a banking error, and been stuck going through some nonsense voice recognition menu tree that doesn't work? Well imagine chat GPT with a real time voice and maybe a fake, photorealistic 3D avatar and having to speak to that anytime you want to speak to a doctor,…

Most things you write sound actually like an improvement over the current state?

I would very much prefer to talk to an AI like GPT4 compared to the people I need to speak to currently on most hotlines. First I need to wait 10-30 minutes in some queue to just be able to speak, and then they are just following some extremely simple script, and lack any real knowledge. I very much expect that GPT4 would be better and more helpful than most hotline conversations I had. Esp when you feed some domain knowledge on the specific application.

I also would like to avoid many of the unnecessary meetings. An AI is perfect for that. It can pass on my necessary knowledge to the others, and it can also compress all the relevant information for me, and give me a summary later. So real meetings would be reduced to only those where we would need to do some important decisions, or some planings, brainstorming sessions. The actual interesting meetings only.

I can also imagine that the quality of Wikipedia and other news articles would actually improve.

Re: GPT-4

#492

Genuinely surprised by the positive reaction about how exciting this all is. You ever had to phone a large business to try and sort something out, like maybe a banking error, and been stuck going through some nonsense voice recognition menu tree that doesn't work? Well imagine chat GPT with a real time voice and maybe a fake, photorealistic 3D avatar and having to speak to that anytime you want to speak to a doctor,…

I don't think your negative scenarios are detailed enough. I can reverse each of them:

1. Imagine that you have 24x7 access to a medical bot that can answer detailed questions about test results, perform ~90% of diagnoses with greater accuracy than a human doctor, and immediately send in prescriptions for things like antibiotics and other basic medicines.

2. Imagine that instead of waiting hours on hold, or days to schedule a call, you can resolve 80% of tax issues immediately through chat.

3. Not sure what to do with mortgages, seems like that's already pretty automated.

4. Imagine that you can hand your resume to a bot, have a twenty minute chat with it to explain details about previous work experience, and what you liked and didn't like about each job, and then it automatically connects you with hiring managers (who have had a similar discussion with it to explain what their requirements and environment are) and get connected.

This all seems very very good to me. What's your nightmare scenario really?

(edit to add: I'm not making any claims about the clogging of reddit/hn with bot-written comments)

Re: GPT-4

#493
post #218

A class of problem that GPT-4 appears to still really struggle with is variants of common puzzles. For example: >Suppose I have a cabbage, a goat and a lion, and I need to get them across a river. I have a boat that can only carry myself and a single other item. I am not allowed to leave the cabbage and lion alone together, and I am not allowed to leave the lion and goat alone together. How can I safely get all three…

LLMs aren’t reasoning about the puzzle. They’re predicting the most likely text to print out, based on the input and the model/training data. If the solution is logical but unlikely (i.e. unseen in the training set and not mapped to an existing puzzle), then the probability of the puzzle answer appearing is very low.

It is disheartening to see how many people are trying to tell you you're wrong when this is literally what it does. It's a very powerful and useful feature, but the over selling of AI has led to people who just want this to be so much more than it actually is.

It sees goat, lion, cabbage, and looks for something that said goat/lion/cabbage. It does not have a concept of "leave alone" and it's not assigning entities with parameters to each item. It does care about things like sentence structure and what not, so it's more complex than a basic lookup, but the amount of borderline worship this is getting is disturbing.

Re: GPT-4

#494

Genuinely surprised by the positive reaction about how exciting this all is. You ever had to phone a large business to try and sort something out, like maybe a banking error, and been stuck going through some nonsense voice recognition menu tree that doesn't work? Well imagine chat GPT with a real time voice and maybe a fake, photorealistic 3D avatar and having to speak to that anytime you want to speak to a doctor,…

> imagine chat GPT with a real time voice and maybe a fake, photorealistic 3D avatar and having to speak to that anytime you want to speak to a doctor, sort out tax issues, apply for a mortgage, apply for a job, etc

For so many current call-center use cases, this sounds like a massive improvement. Then all you need to do is keep iterating on your agent model and you can scale your call-center as easy as you do with AWS's auto scaling! And it can be far superior to the current "audio UI".

>Imagine Reddit and hacker news just filled with endless comments from AIs to suit someone's agenda.

This does worry me, and a lot. We will need to find a way to have "human-verified-only" spaces, and making that will be increasingly hard because I can just manually copy paste whatever gpt told me.

The internet is already full of junk, we may find a point where we have Kessler Syndrome but for the internet...

Re: GPT-4

#495

Genuinely surprised by the positive reaction about how exciting this all is. You ever had to phone a large business to try and sort something out, like maybe a banking error, and been stuck going through some nonsense voice recognition menu tree that doesn't work? Well imagine chat GPT with a real time voice and maybe a fake, photorealistic 3D avatar and having to speak to that anytime you want to speak to a doctor,…

Indeed, the implication of this is that capital now has yet another way to bullshit us all and jerk us around.

This stuff is technologically impressive, but it has very few legitimate uses that will not further inequality.

Re: GPT-4

#496
post #218

A class of problem that GPT-4 appears to still really struggle with is variants of common puzzles. For example: >Suppose I have a cabbage, a goat and a lion, and I need to get them across a river. I have a boat that can only carry myself and a single other item. I am not allowed to leave the cabbage and lion alone together, and I am not allowed to leave the lion and goat alone together. How can I safely get all three…

A funny variation on this kind of over-fitting to common trick questions - if you ask it which weighs more, a pound of bricks or a pound of feathers, it will correctly explain that they actually weigh the same amount, one pound. But if you ask it which weighs more, two pounds of bricks or a pound of feathers, the question is similar enough to the trick question that it falls into the same thought process and contorts…

I hope some future human general can use this trick flummox Skynet if it ever comes to that

Re: GPT-4

#497

Genuinely surprised by the positive reaction about how exciting this all is. You ever had to phone a large business to try and sort something out, like maybe a banking error, and been stuck going through some nonsense voice recognition menu tree that doesn't work? Well imagine chat GPT with a real time voice and maybe a fake, photorealistic 3D avatar and having to speak to that anytime you want to speak to a doctor,…

Honestly I wouldn't worry about it. Outside of the tech bubble most businesses know AI is pointless from a revenue point of view (and comes with legal/credibility/brand risks). Regardless of what the "potential" of this tech is, it's nowhere near market ready and may not be market ready any time soon. As much as the hype suggests dramatic development to come, the cuts in funding within AI groups of most major companies in the space suggests otherwise.

Re: GPT-4

#498
post #218

A class of problem that GPT-4 appears to still really struggle with is variants of common puzzles. For example: >Suppose I have a cabbage, a goat and a lion, and I need to get them across a river. I have a boat that can only carry myself and a single other item. I am not allowed to leave the cabbage and lion alone together, and I am not allowed to leave the lion and goat alone together. How can I safely get all three…

It's a good observation. Although on the flip side, I almost went to type up a reply to you explaining why you were wrong and why bringing the goat first is the right solution. Until I realized I misread what your test was when I skimmed your comment. Likely the same type of mistake GPT-4 made when "seeing" it. Intuitively, I think the answer is that we do have two types of thinking. The pattern matching fast thinkin…

The interesting thing here is that OpenAI is claiming ~90th percentile scores on a number of standardized tests (which, obviously, are typically administered to humans, and have the disadvantage of being mostly or partially multiple choice). Still...

> GPT-4 performed at the 90th percentile on a simulated bar exam, the 93rd percentile on an SAT reading exam, and the 89th percentile on the SAT Math exam, OpenAI claimed.

https://www.cnbc.com/2023/03/14/openai-announces-gpt-4-says-...

So, clearly, it can do math problems, but maybe it can only do "standard" math and logic problems? That might indicate more of a memorization-based approach than a reasoning approach is what's happening here.

The followup question might be: what if we pair GPT-4 with an actual reasoning engine? What do we get then?

Re: GPT-4

#499
>GPT-4 can also be confidently wrong in its predictions, not taking care to double-check work when it’s likely to make a mistake. Interestingly, the base pre-trained model is highly calibrated (its predicted confidence in an answer generally matches the probability of being correct). However, through our current post-training process, the calibration is reduced.

This really made me think.

Re: GPT-4

#500
post #218

A class of problem that GPT-4 appears to still really struggle with is variants of common puzzles. For example: >Suppose I have a cabbage, a goat and a lion, and I need to get them across a river. I have a boat that can only carry myself and a single other item. I am not allowed to leave the cabbage and lion alone together, and I am not allowed to leave the lion and goat alone together. How can I safely get all three…

I will say most humans fail at these too
Post reply on HN