Live data from Hacker News

GPT-4

openai.com

241–250 of 1001 posts

Re: GPT-4

#241

I think it's interesting that they've benchmarked it against an array of standardized tests. Seems like LLMs would be particularly well suited to this kind of test by virtue of it being simple prompt:response, but I have to say...those results are terrifying. Especially when considering the rate of improvement. bottom 10% to top 10% of LSAT in What are the implications for society when general thinking, reading, and…

Well you said it in your comment, if the model was trained with more QAs from those specific benchmarks then it's fair to expect it to do better in that benchmark.

Re: GPT-4

#243

Test taking will change. In the future I could see the student engaging in a conversation with an AI and the AI producing an evaluation. This conversation may be focused on a single subject, or more likely range over many fields and ideas. And may stretch out over months. Eventually teaching and scoring could also be integrated as the AI becomes a life-long tutor. Even in a future where human testing/learning is no l…

While many may shudder at this, I find your comment fantastically inspiring. As a teacher, writing tests always feels like an imperfect way to assess performance. It would be great to have a conversation with each student, but there is no time to really go into such a process. Would definitely be interesting to have an AI trained to assess learning progress by having an automated, quick chat with a student about the topic. Of course, the AI would have to have anti-AI measures ;)

Re: GPT-4

#245
As a dyslexic person with a higher education this hits really close to home. Not only should we not be surprised that a LLM would be good at answering tests like this, we should be excited that technology will finaly free us from being judged in this way. This is a patern that we have seen over and over again in tech, where machines can do something better than us, and eventually free us from having to worry about it. Before it was word processing, now it is accurate knowledge recall.

Re: GPT-4

#246

I think it's interesting that they've benchmarked it against an array of standardized tests. Seems like LLMs would be particularly well suited to this kind of test by virtue of it being simple prompt:response, but I have to say...those results are terrifying. Especially when considering the rate of improvement. bottom 10% to top 10% of LSAT in What are the implications for society when general thinking, reading, and…

Those benchmarks are so cynical.

Every test prep tutor taught dozens/hundreds of students the implicit patterns behind the tests and drilled it into them with countless sample questions, raising their scores by hundreds of points. Those students were not getting smarter from that work, they were becoming more familiar with a format and their scores improved by it.

And what do LLM’s do? Exactly that. And what’s in their training data? Countless standardized tests.

These things are absolutely incredible innovations capable of so many things, but the business opportunity is so big that this kind of cynical misrepresentation is rampant. It would be great if we could just stay focused on the things they actually do incredibly well instead of the making them do stage tricks for publicity.

Re: GPT-4

#247
post #218

A class of problem that GPT-4 appears to still really struggle with is variants of common puzzles. For example: >Suppose I have a cabbage, a goat and a lion, and I need to get them across a river. I have a boat that can only carry myself and a single other item. I am not allowed to leave the cabbage and lion alone together, and I am not allowed to leave the lion and goat alone together. How can I safely get all three…

Awesome test. Do you have a list of others?

Re: GPT-4

#248

> Given both the competitive landscape and the safety implications of large-scale models like GPT-4, this report contains no further details about the architecture (including model size), hardware, training compute, dataset construction, training method, or similar. My guess is they used Chinchilla scaling rules and the parameter count for GPT-4 is either barely larger or maybe even smaller than GPT-3. Look as what M…

The larger context length makes me think they have a more memory-efficient attention mechanism.

Re: GPT-4

#249
Had to chuckle here going through the exam results: Advanced Sommelier (theory knowledge)

AI is so advanced, it started drinking!

Re: GPT-4

#250

It's interesting that everyone is talking about programmers being replaced by AI, but the model did far better on the humanities type subjects than on the programming tests.

Maybe I’m just old but I don’t quite understand the hype.

As long as it’s vulnerable to hallucinating, it can’t be used for anything where there are “wrong answers” - and I don’t think ChatGPT-4 has fixed that issue yet.*

Now if it’s one of those tasks where there are “no wrong answers”, I can see it being somewhat useful. A non-ChatGPT AI example would be those art AIs - art doesn’t have to make sense.

The pessimist in me see things like ChatGPT as the ideal internet troll - it can be trained to post stuff that maximise karma gain while pushing a narrative which it will hallucinate its way into justifying.

* When they do fix it, everyone is out of a job. Humans will only be used for cheap labor - because we are cheaper than machines.

Post reply on HN