Live data from Hacker News

GPT-4

openai.com

351–360 of 1001 posts

Re: GPT-4

#351

Finally, we facilitated a preliminary model evaluation by the Alignment Research Center (ARC) focused on the ability of GPT-4 versions they evaluated to carry out actions to autonomously replicate5 and gather resources—a risk that, while speculative, may become possible with sufficiently advanced AI systems—with the conclusion that the current model is probably not yet capable of autonomously doing so. or it's just r…

LOL some basic kind of embodiement/autonomy is not that hard to do on these kinds of AI models if you're willing to write some more code and a prompt more carefully. I've tested it and it works quite well.

"{prompt} After you reply to this, indicate an amount of time between 0 and X minutes from now that you would like to wait before speaking again".

Then detect the amount of time it specifies, and have a UI that automatically sends an empty input prompt after the amount of time specified elapses when this is triggered (assuming the user doesn't respond first).

I'm gonna knock this out as a weekend project one of these weekends to prove this.

Re: GPT-4

#352
post #41

> Given both the competitive landscape and the safety implications of large-scale models like GPT-4, this report contains no further details about the architecture (including model size), hardware, training compute, dataset construction, training method, or similar. "Open"

Well it is open.

Your wallet that is.

Re: GPT-4

#354
GPT-4 Everything we know so far...

GPT-4 can solve difficult problems with greater accuracy, thanks to its broader general knowledge and problem-solving abilities.

GPT-4 is more reliable, creative, and able to handle much more nuanced instructions than GPT-3.5. It surpasses ChatGPT in its advanced reasoning capabilities.

GPT-4 is safer and more aligned. It is 82% less likely to respond to requests for disallowed content and 40% more likely to produce factual responses than GPT-3.5 on our internal evaluations.

GPT-4 still has many known limitations that we are working to address, such as social biases, hallucinations, and adversarial prompts.

GPT-4 can accept a prompt of text and images, which—parallel to the text-only setting—lets the user specify any vision or language task.

GPT-4 is available on ChatGPT Plus and as an API for developers to build applications and services. (API- waitlist right now)

Duolingo, Khan Academy, Stripe, Be My Eyes, and Mem amongst others are already using it.

API Pricing GPT-4 with an 8K context window (about 13 pages of text) will cost $0.03 per 1K prompt tokens, and $0.06 per 1K completion tokens. GPT-4-32k with a 32K context window (about 52 pages of text) will cost $0.06 per 1K prompt tokens, and $0.12 per 1K completion tokens.

Re: GPT-4

#355
post #277
post #218

A class of problem that GPT-4 appears to still really struggle with is variants of common puzzles. For example: >Suppose I have a cabbage, a goat and a lion, and I need to get them across a river. I have a boat that can only carry myself and a single other item. I am not allowed to leave the cabbage and lion alone together, and I am not allowed to leave the lion and goat alone together. How can I safely get all three…

For anyone wondering you bring the lion across. The trick is that it’s the lion that eats the cabbage not the goat.

Lion ->

Goat ->

Cabbage ->

Lion ->

Re: GPT-4

#356
post #76

The "visual inputs" samples are extraordinary, and well worth paying extra attention to. I wasn't expecting GPT-4 to be able to correctly answer "What is funny about this image?" for an image of a mobile phone charger designed to resemble a VGA cable - but it can. (Note that they have a disclaimer: "Image inputs are still a research preview and not publicly available.")

Can it explain this one? https://www.reddit.com/r/seinfeld/comments/e82uuy/new_yorker...

Re: GPT-4

#357
post #218

A class of problem that GPT-4 appears to still really struggle with is variants of common puzzles. For example: >Suppose I have a cabbage, a goat and a lion, and I need to get them across a river. I have a boat that can only carry myself and a single other item. I am not allowed to leave the cabbage and lion alone together, and I am not allowed to leave the lion and goat alone together. How can I safely get all three…

[deleted]

Re: GPT-4

#358

I think it's interesting that they've benchmarked it against an array of standardized tests. Seems like LLMs would be particularly well suited to this kind of test by virtue of it being simple prompt:response, but I have to say...those results are terrifying. Especially when considering the rate of improvement. bottom 10% to top 10% of LSAT in What are the implications for society when general thinking, reading, and…

Interest in human-played Chess is (arguably) at all time high, so I would say it bodes well based on that.

Re: GPT-4

#359

I think it's interesting that they've benchmarked it against an array of standardized tests. Seems like LLMs would be particularly well suited to this kind of test by virtue of it being simple prompt:response, but I have to say...those results are terrifying. Especially when considering the rate of improvement. bottom 10% to top 10% of LSAT in What are the implications for society when general thinking, reading, and…

GPT-4 Everything we know so far...

GPT-4 can solve difficult problems with greater accuracy, thanks to its broader general knowledge and problem-solving abilities.

GPT-4 is more reliable, creative, and able to handle much more nuanced instructions than GPT-3.5. It surpasses ChatGPT in its advanced reasoning capabilities.

GPT-4 is safer and more aligned. It is 82% less likely to respond to requests for disallowed content and 40% more likely to produce factual responses than GPT-3.5 on our internal evaluations.

GPT-4 still has many known limitations that we are working to address, such as social biases, hallucinations, and adversarial prompts.

GPT-4 can accept a prompt of text and images, which—parallel to the text-only setting—lets the user specify any vision or language task.

GPT-4 is available on ChatGPT Plus and as an API for developers to build applications and services. (API- waitlist right now)

Duolingo, Khan Academy, Stripe, Be My Eyes, and Mem amongst others are already using it.

API Pricing GPT-4 with an 8K context window (about 13 pages of text) will cost $0.03 per 1K prompt tokens, and $0.06 per 1K completion tokens. GPT-4-32k with a 32K context window (about 52 pages of text) will cost $0.06 per 1K prompt tokens, and $0.12 per 1K completion tokens.

Re: GPT-4

#360

I think it's interesting that they've benchmarked it against an array of standardized tests. Seems like LLMs would be particularly well suited to this kind of test by virtue of it being simple prompt:response, but I have to say...those results are terrifying. Especially when considering the rate of improvement. bottom 10% to top 10% of LSAT in What are the implications for society when general thinking, reading, and…

If you had told me 5 years ago that there would be a single AI system that could perform at this level on such a vast array of standardized tests, I would've said "That's a true AGI." Commentary to the contrary feels like quibbling over a very localized point in time versus looking at the bigger picture.
Post reply on HN