Live data from Hacker News

GPT-4

openai.com

211–220 of 1001 posts

Re: GPT-4

#211
Looks like Bing chat is using GPT-4 already:

"Good news, we've increased our turn limits to 15/150. Also confirming that the next-gen model Bing uses in Prometheus is indeed OpenAI's GPT-4 which they just announced today." - Jordi Ribas, Corporate VP @ Bing/Microsoft

https://twitter.com/JordiRib1/status/1635694953463705600

Re: GPT-4

#212
post #95

I cant wait for this to do targeted censorship! It already demonstrates it has strong biases deliberately programmed in: > I cannot endorse or promote smoking, as it is harmful to your health. But it would likely happily promote or endorse driving, skydiving, or eating manure - if asked in the right way.

The point of that example was that they indicated it was the wrong response. After RLHF the model correctly tells the user how to find cheap cigarettes (while still chiding them for smoking)

Re: GPT-4

#213

Does "Open"AI really not even say how many parameters their models have?

The 98-pages paper doesn't say anything about the architecture of the model, I know, the irony

Re: GPT-4

#214
post #181

LLMs will eventually make a lot of simpler machine-learning models obsolete. Imagine feeding a prompt akin to the one below to GPT5, GPT6, etc.: prompt = f"The guidelines for recommending products are: {guidelines}. The following recommendations led to incremental sales: {sample_successes}. The following recommendations had no measurable impact: {sample_failures}. Please make product recommendations for these custome…

Except the machine can’t explain its reasoning, it will make up some plausible justification for its output.

Humans often aren’t much better, making up a rational sounding argument after the fact to justify a decision they don’t fully understand either.

A manager might fire someone because they didn’t sleep well or skipped breakfast. They’ll then come up with a logical argument to support what was an emotional decision. Humans do this more often than we’d like to admit.

Re: GPT-4

#215

It is amazing how this crowd in HN reacts to AI news coming out of OpenAI compared to other competitors like Google or FB. Today there was another news about Google releasing their AI in GCP and mostly the comments were negative. The contrast is clearly visible and without any clear explanation for this difference I have to suspect that maybe something is being artificially done to boost one against the other.

Or it could be that Google and FB are both incumbents scrambling to catch up with OpenAI, who is a much smaller competitor that is disrupting the space?

Re: GPT-4

#217
post #86

> What are the implications for society when general thinking, reading, and writing becomes like Chess? I think going from LSAT to general thinking is still a very, very big leap. Passing exams is a really fascinating benchmark but by their nature these exams are limited in scope, have very clear assessment criteria and a lot of associated and easily categorized data (like example tests). General thought (particularl…

The big huge difference is that cars have this unfortunate thing where if they crash, people get really hurt or killed, especially pedestrians. And split second response time matters, so it's hard for a human operator to just jump in. If ChatGPT-4 hallucinates an answer, it won't kill me. If a human needs to proofread the email it wrote before sending, it'll wait for seconds or minutes.

Re: GPT-4

#218
A class of problem that GPT-4 appears to still really struggle with is variants of common puzzles. For example:

>Suppose I have a cabbage, a goat and a lion, and I need to get them across a river. I have a boat that can only carry myself and a single other item. I am not allowed to leave the cabbage and lion alone together, and I am not allowed to leave the lion and goat alone together. How can I safely get all three across?

In my test, GPT-4 charged ahead with the standard solution of taking the goat first. Even after I pointed this mistake out, it repeated exactly the same proposed plan. It's not clear to me if the lesson here is that GPT's reasoning capabilities are being masked by an incorrect prior (having memorized the standard version of this puzzle) or if the lesson is that GPT'S reasoning capabilities are always a bit of smoke and mirrors that passes off memorization for logic.

Re: GPT-4

#220

I think it's interesting that they've benchmarked it against an array of standardized tests. Seems like LLMs would be particularly well suited to this kind of test by virtue of it being simple prompt:response, but I have to say...those results are terrifying. Especially when considering the rate of improvement. bottom 10% to top 10% of LSAT in What are the implications for society when general thinking, reading, and…

I think it shows how calcified standardized tests have become. We will have to revisit all of them, and change many things about how they work, or they will be increasingly useless.
Post reply on HN