Live data from Hacker News

GPT-4

openai.com

391–400 of 1001 posts

Re: GPT-4

#391
post #130

This technology has been a true blessing to me. I have always wished to have a personal PhD in a particular subject whom I could ask endless questions until I grasped the topic. Thanks to recent advancements, I feel like I have my very own personal PhDs in multiple subjects, whom I can bombard with questions all day long. Although I acknowledge that the technology may occasionally produce inaccurate information, the…

Besides the fact that this comment reads written by GPT itself, using this particular AI as a source for your education is like going to the worse University out there. I am sure if you always wishes do thave a personal PhD in a particular subject you could find shady universities out there who could provide one without much effort. [I may be exagerating but the point still stands because the previous user also didn'…

I don't think that's the user's intended meaning of "personal PhD," ie they don't mean a PhD or PhD level knowledge held by themselves, they mean having a person with a PhD that they can call up with questions. It seems like in some fields GPT4 will be on par with even PhD-friends who went to reasonably well respected institutions.

Re: GPT-4

#392

I think it's interesting that they've benchmarked it against an array of standardized tests. Seems like LLMs would be particularly well suited to this kind of test by virtue of it being simple prompt:response, but I have to say...those results are terrifying. Especially when considering the rate of improvement. bottom 10% to top 10% of LSAT in What are the implications for society when general thinking, reading, and…

The silver lining might be us finally realising how bad standardised tests are at measuring intellect, creativity and the characteristics that make us thrive.

Most of the time they are about loading/unloading data. Maybe this will also revolutionise education, turning it more towards discovery and critical thinking, rather than repeating what we read in a book/heard in class?

Re: GPT-4

#393

I think it's interesting that they've benchmarked it against an array of standardized tests. Seems like LLMs would be particularly well suited to this kind of test by virtue of it being simple prompt:response, but I have to say...those results are terrifying. Especially when considering the rate of improvement. bottom 10% to top 10% of LSAT in What are the implications for society when general thinking, reading, and…

I think Chess is an easier thing to be defeated at by a machine because there is a clear winner and a clear loser.

Thinking, reading, interpreting and writing are skills which produce outputs that are not as simple as black wins, white loses.

You might like a text that a specific author writes much more than what GPT-4 may be able to produce. And you might have a different interpretation of a painting than GPT-4 has.

And no one can really say who is better and who is worse on that regard.

Re: GPT-4

#394
post #10

Ugh that testing graph confirms that AP Environmental Science was indeed the easiest AP class and I needn't be proud of passing that exam.

I am interested that GPT4 botched AP Lang and Comp and AP English Lit and Comp just as badly as GPT3.5, with a failing grade of 2/5 (and many colleges also consider a 3 on those exams a failure). Is it because of gaps in the training data or something else? Why does it struggle so hard with those specific tests? Especially since it seems to do fine at the SAT writing section.

Re: GPT-4

#395

All this bluster about replacing technical jobs like legal counsel ignores that you are fundamentally paying for accountability. “The AI told me it was ok” only works if, when it’s not, there is recourse. We can barely hold Google et Al accountable for horrible user policies…why would anyone think OpenAI will accept any responsibility for any recommendations made by a GPT?

They won't, but that doesn't mean some other business won't automate legal counsel and assume risk. If, down the line, GPT (or some other model) has empirically been proven to be more accurate than legal assistants and lawyers, why wouldn't this been the obvious outcome?

Re: GPT-4

#396
post #76

The "visual inputs" samples are extraordinary, and well worth paying extra attention to. I wasn't expecting GPT-4 to be able to correctly answer "What is funny about this image?" for an image of a mobile phone charger designed to resemble a VGA cable - but it can. (Note that they have a disclaimer: "Image inputs are still a research preview and not publicly available.")

Am I the only one who thought that GPT-4 got this one wrong? It's not simply that it's ridiculous to plug what appears to be an outdated VGA cable into a phone, it's that the cable connector does nothing at all. I'd argue that's what actually funny. GPT-4 didn't mention that part as far as I could see.

Re: GPT-4

#397
post #218

A class of problem that GPT-4 appears to still really struggle with is variants of common puzzles. For example: >Suppose I have a cabbage, a goat and a lion, and I need to get them across a river. I have a boat that can only carry myself and a single other item. I am not allowed to leave the cabbage and lion alone together, and I am not allowed to leave the lion and goat alone together. How can I safely get all three…

I am curious what percentage of humans would also give the incorrect answer to this puzzle, and for precisely the same reason (i.e. they incorrectly pattern-matched it to the classic puzzle version and plowed ahead to their stored answer). If the percentage is significant, and I think it might be, that's another data point in favor of the claim that really most of what humans are doing when we think we're being intelligent is also just dumb pattern-matching and that we're not as different from the LLMs as we want to think.

Re: GPT-4

#398
post #218

A class of problem that GPT-4 appears to still really struggle with is variants of common puzzles. For example: >Suppose I have a cabbage, a goat and a lion, and I need to get them across a river. I have a boat that can only carry myself and a single other item. I am not allowed to leave the cabbage and lion alone together, and I am not allowed to leave the lion and goat alone together. How can I safely get all three…

I think this goes in line with the results in the GRE. In the verbal section it has an amazing 99%, but in the quant one it "only" has an 80%. The quant section requires some reasoning, but the problems are much easier than the river puzzle, and it still misses some of them. I think part of the difficulty for a human is the time constraint, and given more time to solve it most people would get all questions right.

Re: GPT-4

#400

This technology has been a true blessing to me. I have always wished to have a personal PhD in a particular subject whom I could ask endless questions until I grasped the topic. Thanks to recent advancements, I feel like I have my very own personal PhDs in multiple subjects, whom I can bombard with questions all day long. Although I acknowledge that the technology may occasionally produce inaccurate information, the…

I'm very excited for the future wave of confidently incorrect people powered by ChatGPT.

"The existence of ChatGPT does not necessarily make people confidently incorrect."

- ChatGPT

Post reply on HN