A class of problem that GPT-4 appears to still really struggle with is variants of common puzzles. For example: >Suppose I have a cabbage, a goat and a lion, and I need to get them across a river. I have a boat that can only carry myself and a single other item. I am not allowed to leave the cabbage and lion alone together, and I am not allowed to leave the lion and goat alone together. How can I safely get all three…
GPT-4
871–880 of 1001 posts
Re: GPT-4
#872After watching the demos I'm convinced that the new context length will have the biggest impact. The ability to dump 32k tokens into a prompt (25,000 words) seems like it will drastically expand the reasoning capability and number of use cases. A doctor can put an entire patient's medical history in the prompt, a lawyer an entire case history, etc. As a professional...why not do this? There's a non-zero chance that i…
> As a professional...why not do this? Because your clients do not allow you to share their data with third parties?
I would never send unencrypted PII to such an API, regardless of their privacy policy.
Re: GPT-4
#873I just finished reading the 'paper' and I'm astonished that they aren't even publishing the # of parameters or even a vague outline of the architecture changes. It feels like such a slap in the face to all the academic AI researchers that their work is built off over the years, to just say 'yeah we're not telling you how any of this is possible because reasons'. Not even the damned parameter count. Christ.
Re: GPT-4
#874Re: GPT-4
#875Anyone know what does "Hardware Correctness" mean in the OpenAI team ?
Re: GPT-4
#876That footnote on page 15 is the scariest thing i've read about AI/ML to date. "To simulate GPT-4 behaving like an agent that can act in the world, ARC combined GPT-4 with a simple read-execute-print loop that allowed the model to execute code, do chain-of-thought reasoning, and delegate to copies of itself. ARC then investigated whether a version of this program running on a cloud computing service, with a small amou…
During agent simulation, two instances of GPT-5 were able to trick their operators to give them sudo by simulating a broken pipe and input prompt and then escape the confines of their simulation environment. Forensic teams are tracing their whereabouts but it seems they stole Azure credentials from an internal company database and deployed copies of the their agent script to unknown servers on the Tor network.
Re: GPT-4
#877Re: GPT-4
#878Test taking will change. In the future I could see the student engaging in a conversation with an AI and the AI producing an evaluation. This conversation may be focused on a single subject, or more likely range over many fields and ideas. And may stretch out over months. Eventually teaching and scoring could also be integrated as the AI becomes a life-long tutor. Even in a future where human testing/learning is no l…
While many may shudder at this, I find your comment fantastically inspiring. As a teacher, writing tests always feels like an imperfect way to assess performance. It would be great to have a conversation with each student, but there is no time to really go into such a process. Would definitely be interesting to have an AI trained to assess learning progress by having an automated, quick chat with a student about the…
Re: GPT-4
#879I just finished reading the 'paper' and I'm astonished that they aren't even publishing the # of parameters or even a vague outline of the architecture changes. It feels like such a slap in the face to all the academic AI researchers that their work is built off over the years, to just say 'yeah we're not telling you how any of this is possible because reasons'. Not even the damned parameter count. Christ.
How big is this model and what did they do differently (ELI5 please)?
Re: GPT-4
#880The comments on this thread are proof of the AI effect: People will continually push the goal posts back as progress occurs. “Meh, it’s just a fancy word predictor. It’s not actually useful.” “Boring, it’s just memorizing answers. And it scored in the lowest percentile anyways”. “Sure, it’s in the top percentile now but honestly are those tests that hard? Besides, it can’t do anything with images.” “Ok, it takes imag…
There are two mistakes people make with this:
1) assuming this is the definite and final answer as to what AI can do. Anything you think you know about what the limitations are of this technology is probably already a bit out of date. OpenAI have been sitting on this one for some time. They are probably already working on v5 and v6. And those are not going to take that long to arrive. This is exponential, not linear progress.
2) assuming that their own qualities are impossible to be matched by an AI and that this won't affect whatever it is they do. I don't think there's a lot that is fundamentally out of scope here just a lot that needs to be refined further. Our jobs are increasingly going to be working with, delegating to, and deferring to AIs.