Live data from Hacker News

GPT-4

openai.com

521–530 of 1001 posts

Re: GPT-4

#521
post #513

I would love if GPT-4 would be connected to github and starts to solve all open bugs there. Could this be the future: Pull requests from GPT-4 automatically solving real issues/problems in your code?

If you look at the "simulated exams" table, it actually does poorly on coding problems.

Re: GPT-4

#522
post #218

A class of problem that GPT-4 appears to still really struggle with is variants of common puzzles. For example: >Suppose I have a cabbage, a goat and a lion, and I need to get them across a river. I have a boat that can only carry myself and a single other item. I am not allowed to leave the cabbage and lion alone together, and I am not allowed to leave the lion and goat alone together. How can I safely get all three…

I gave it a different kind of puzzle, again with a twist (no solution), and it spit out nonsense. "I have two jars, one that can hold 5 liters, and one that can hold 10 liters. How can I measure 3 liters?" It gave 5 steps, some of which made sense but of course didn't solve the problem. But at the end it cheerily said "Now you have successfully measured 3 liters of water using the two jars!"

Re: GPT-4

#523
post #76

The "visual inputs" samples are extraordinary, and well worth paying extra attention to. I wasn't expecting GPT-4 to be able to correctly answer "What is funny about this image?" for an image of a mobile phone charger designed to resemble a VGA cable - but it can. (Note that they have a disclaimer: "Image inputs are still a research preview and not publicly available.")

If they are using popular images from the internet, then I strongly suspect the answers come from the text next to the known image. The man ironing on the back of the taxi has the same issue. https://google.com/search?q=mobile+phone+charger+resembling+...

I would bet good money that when we can test prompting with our own unique images, GPT4 will not give similar quality answers.

I do wonder how misleading their paper is.

Re: GPT-4

#524
post #218

A class of problem that GPT-4 appears to still really struggle with is variants of common puzzles. For example: >Suppose I have a cabbage, a goat and a lion, and I need to get them across a river. I have a boat that can only carry myself and a single other item. I am not allowed to leave the cabbage and lion alone together, and I am not allowed to leave the lion and goat alone together. How can I safely get all three…

A funny variation on this kind of over-fitting to common trick questions - if you ask it which weighs more, a pound of bricks or a pound of feathers, it will correctly explain that they actually weigh the same amount, one pound. But if you ask it which weighs more, two pounds of bricks or a pound of feathers, the question is similar enough to the trick question that it falls into the same thought process and contorts…

I tried this with the new model and it worked correctly on both examples.

Re: GPT-4

#525

I'll be finishing my interventional radiology fellowship this year. I remember in 2016 when Geoffrey Hinton said, "We should stop training radiologists now," the radiology community was aghast and in-denial. My undergrad and masters were in computer science, and I felt, "yes, that's about right." If you were starting a diagnostic radiology residency, including intern year and fellowship, you'd just be finishing now.…

It all comes down to labelled data. There are millions images of VGA connectors and lightning cables on the internet with description, where CLIP model and similar could learn to recognize them relatively reliably. On the other hand, I'm not sure such amount of data are available for AI training. Especially if the diagnostic is blinded, it will be even harder for the AI model to reliably differentiate between them, making cross-disease diagnostic hard. Not to mention the risk and reliability of such tasks.

Re: GPT-4

#526

Genuinely surprised by the positive reaction about how exciting this all is. You ever had to phone a large business to try and sort something out, like maybe a banking error, and been stuck going through some nonsense voice recognition menu tree that doesn't work? Well imagine chat GPT with a real time voice and maybe a fake, photorealistic 3D avatar and having to speak to that anytime you want to speak to a doctor,…

So, there are a four categories of things in your comment: two concepts (interactive vs. static) divided into two genres (factual vs. incidental).

For interactive/factual, we have getting help on taxes and accounting (and to a large extent law), which AI is horrible with and will frankly be unable to help with at this time, and so there will not be AIs on the other side of that interaction until AIs get better enough to be able to track numbers and legal details correctly... at which point you hopefully will never have to be on the phone asking for help as the AI will also be doing the job in the first place.

https://www.instagram.com/p/CnpXLncOfbr/

Then we have interactive/incidental, with situations like applying for jobs or having to wait around with customer service to get some kind of account detail fixed. Today, if you could afford such and knew how to source it, one could imagine outsourcing that task to a personal assistant, which might include a "virtual" one, by which is not meant a fake one but instead one who is online, working out of a call center far away... but like, that could be an AI, and it would be much cheaper and easier to source.

So, sure: that will be an AI, but you'll also be able to ask your phone "hey, can you keep talking to this service until it fixes my problem? only notify me to join back in if I am needed". And like, I see you get that this half is possible, because of your comment about Zoom... but, isn't that kind of great? We all agree that the vast majority of meetings are useless, and yet for some reason we have to have them. If you are high status enough, you send an assistant or "field rep" to the meeting instead of you. Now, everyone at the meeting will be an AI and the actual humans don't have to attend; that's progress!

Then we have static/factual, where we can and should expect all the news articles and reviews to be fake or wrong. Frankly, I think a lot of this stuff already is fake or wrong, and I have to waste a ton of time trying to do enough research to decide what the truth actually is... a task which will get harder if there is more fake content but also will get easier if I have an AI that can read and synthesize information a million times faster than I can. So, sure: this is going to be annoying, but I don't think this is going to be net worse by an egregious amount (I do agree it will be at least somewhat) when you take into account AI being on both sides of the scale.

And finally we have static/incidental content, which I don't even think you did mention but is demanded to fill in the square: content like movies and stories and video games... maybe long-form magazine-style content... I love this stuff and I enjoy reading it, but frankly do I care if the next good movie I watch is made by an AI instead of a human? I don't think I would. I would find a television show with an infinite number of episodes interesting... maybe even so interesting that I would have to refuse to ever watch it lest I lose my life to it ;P. The worst case I can come up with is that we will need help curating all that content, and I think you know where I am going to go on that front ;P.

But so, yeah: I agree things are going to change pretty fast, but mostly in the same way the world changed pretty fast with the introduction of the telephone, the computer, the Internet, and then the smartphone, which all are things that feel dehumanizing and yet also free up time through automation. I certainly have ways in which I am terrified of AI, but these "completely change the way things we already hate--like taxes, phone calls, and meetings--interact with our lives" isn't part of it.

Re: GPT-4

#527

Genuinely surprised by the positive reaction about how exciting this all is. You ever had to phone a large business to try and sort something out, like maybe a banking error, and been stuck going through some nonsense voice recognition menu tree that doesn't work? Well imagine chat GPT with a real time voice and maybe a fake, photorealistic 3D avatar and having to speak to that anytime you want to speak to a doctor,…

You are looking at from a perspective where the chatbots are only used to generate junk content. Which is a real problem. However, there is another far more positive perspective on this. These chatbots can not just generate junk, they can also filter it. They are knowledge-engines that allow you to interact with the trained information directly, in whatever form you desire, completely bypassing the need for accessing websites or following whatever information flow they force on you. Those chatbots are an universal interface to information.

I wouldn't mind if that means I'll never have to read a human written news article again, since most of them are already junk. Filled with useless prose and filler, when all I want is the plain old facts of what happened. A chatbot can provide me exactly what I want.

The open question is of course the monetization. If chatbots can provide me with all the info I want without having to visit sites, who is going to pay for those sites? If they all stop existing, what future information will chatbots be trained on?

Hard to say where things will be going. But I think the way chatbots will change how we interact with information will be far more profound than just generation of junk.

Re: GPT-4

#528

What's the lifespan of an LLM going to be in the next few years? Seems like at the current pace, cutting edge models will become obsolete pretty quickly. Since model training is very expensive, this means the LLM space has some parallels with the pharmaceutical industry (massive upfront capital costs, cheap marginal costs relative to value produced). I find it quite fascinating how quickly machine learning has change…

Deep Learning training was always very expensive but models werent getting such a massive bump in size every year (for state of the art) and now they are just getting 10x bigger every iteration but AI accelerators / GPUs are getting like 1.5x jump every 2 years so have fun for future AI academia / startups outside US.

Re: GPT-4

#529

This technology has been a true blessing to me. I have always wished to have a personal PhD in a particular subject whom I could ask endless questions until I grasped the topic. Thanks to recent advancements, I feel like I have my very own personal PhDs in multiple subjects, whom I can bombard with questions all day long. Although I acknowledge that the technology may occasionally produce inaccurate information, the…

[dead]

Re: GPT-4

#530

Test taking will change. In the future I could see the student engaging in a conversation with an AI and the AI producing an evaluation. This conversation may be focused on a single subject, or more likely range over many fields and ideas. And may stretch out over months. Eventually teaching and scoring could also be integrated as the AI becomes a life-long tutor. Even in a future where human testing/learning is no l…

The focus will shift from knowing the right answer to asking the right questions. It'll still require an understanding of core concepts.
Post reply on HN