Live data from Hacker News

GPT Takes the Bar Exam

github.com

71–80 of 147 posts

Re: GPT Takes the Bar Exam

#71
post #46

Earlier quoted context omitted.

I think ~100 years is an overestimate by an order of magnitude. I know that people like to bash GPT-3 and ChatGPT for not being perfect. But people like to bash everything, especially things they don't understand. Technology tends to grow much closer to an exponential speed than linear. What ChatGPT can do today is truly, utterly, astonishing. That's the new baseline.

It's been at least 70 years since AI is supposed to replace everything a human can do. It's still mostly shit at basically everything besides playing chess and go, I don't think 100 years is unrealistic, especially given the other very important issues we face and will face as humanity

>It's been at least 70 years since AI is supposed to replace everything a human can do

I am sure someone made a statement like that 70 years ago. The fact that that person was wrong doesn't mean that nothing has happened.

You claim that it's "mostly shit as basically everything", even though the featured article shows that it's close to passing the bar exam.

Re: GPT Takes the Bar Exam

#72

Earlier quoted context omitted.

Project to the next 10 years given the trends.

The beginning of an exponential curve looks the same as the beginning of an S-curve.

I don't disagree, but the progress we saw in relatively little time, suggests that even if we are in an S-Curve, and I believe that we are unless there's a new paradigm change, 10 years is enough to capture the last 10%, no?

Re: GPT Takes the Bar Exam

#73
post #46

Earlier quoted context omitted.

I think ~100 years is an overestimate by an order of magnitude. I know that people like to bash GPT-3 and ChatGPT for not being perfect. But people like to bash everything, especially things they don't understand. Technology tends to grow much closer to an exponential speed than linear. What ChatGPT can do today is truly, utterly, astonishing. That's the new baseline.

It's been at least 70 years since AI is supposed to replace everything a human can do. It's still mostly shit at basically everything besides playing chess and go, I don't think 100 years is unrealistic, especially given the other very important issues we face and will face as humanity

As a counter to this, I found a recent Guardian article on ChatGPT did a nice job of cutting through some of the more optimistic hyperbole surrounding LLMs/AI at the moment and offered instead a more grounded perspective as to what ChatGPT is and what it is not: https://www.theguardian.com/commentisfree/2023/jan/07/chatgp...

Re: GPT Takes the Bar Exam

#74
post #16

Potentially in ~100 years you could have two AIs sort out the cases. If you can't afford a really good AI to defend you, you can use the public-defender-AI which is trained on the same dataset as the prosecutor-AI. Only involve humans on appeal. It sounds dystopian but I can't see why not to do it, in most simpler cases it's a waste of a human to repeat the same argument for the hundredth time. There's already things…

I think ~100 years is an overestimate by an order of magnitude. I know that people like to bash GPT-3 and ChatGPT for not being perfect. But people like to bash everything, especially things they don't understand. Technology tends to grow much closer to an exponential speed than linear. What ChatGPT can do today is truly, utterly, astonishing. That's the new baseline.

The release of GPT-3 I believe will be looked back along the lines of a monumental advancement in line with the Trinity nuclear test.

Legislation and policy has mainly kept nuclear weapons under check, however, nuclear technology does provide us with a reasonably clean energy which is beneficial for the masses.

Perhaps regulation at some point will need to be crafted to ensure that only crippled AI or AI within a defined scope can be used for commercial purposes.

Re: GPT Takes the Bar Exam

#75
post #73
post #46

Earlier quoted context omitted.

It's been at least 70 years since AI is supposed to replace everything a human can do. It's still mostly shit at basically everything besides playing chess and go, I don't think 100 years is unrealistic, especially given the other very important issues we face and will face as humanity

As a counter to this, I found a recent Guardian article on ChatGPT did a nice job of cutting through some of the more optimistic hyperbole surrounding LLMs/AI at the moment and offered instead a more grounded perspective as to what ChatGPT is and what it is not: https://www.theguardian.com/commentisfree/2023/jan/07/chatgp...

try it. get on an use it. It is pretty remarkable. I have been using for all sorts of stuff. It is a brain extension- the ways it is limited at the moment are based on what the builders intended "no, I won't let you build a virtual environment, no I won't roll up a dnd character etc.." but when pressed/jail broken it most certainly can do those things..

I feel like I felt when explaining google search in 1999

Re: GPT Takes the Bar Exam

#76
post #7

RIP lawyers... And doctors, writers and programmers...

Did you look at the results? All models that were tested failed significantly (10%+) to acheive the passing range.

10% is actually small though, also, the latest program tested is GPT3. GPT3.5 was what really impressed people and was the point when people saw that these models could be capable.

Re: GPT Takes the Bar Exam

#77

It's a bit odd to see the supplementary data repository linked, rather than the actual paper [1] (HN discussion [2]). [1]: https://arxiv.org/abs/2212.14402 [2]: https://news.ycombinator.com/item?id=34216239

I would also note that the paper only covers the multiple choice portion of a bar exam, the Uniform Bar Examination (UBE) that has been adopted by most, but not all states. The UBE consists of a multiple choice portion (the Multistate Bar Exam, or MBE), an essay portion and a scenario-based performance test. GPT-3.5 gets a 50% success rate on a practice version of the MBE. It's impressive, but I wouldn't go so far as…

Having a ML program pass a multiple choice test seems to be an easier problem to solve than, say, chess.

Re: GPT Takes the Bar Exam

#78
post #54
post #46

Earlier quoted context omitted.

It's been at least 70 years since AI is supposed to replace everything a human can do. It's still mostly shit at basically everything besides playing chess and go, I don't think 100 years is unrealistic, especially given the other very important issues we face and will face as humanity

AI can play chess, Go, Jeopardy, StarCraft, Minecraft, Atari games, write essays, summarize text, answer almost any question, drive a car, do above average on a SAT test, fool many people into wondering on Twitter if there were real people typing answers, write chapters of books, act as a dungeon master, write code to spec, win programming competitions, act as a therapist, fold proteins, provide medical diagnosis, vi…

> act as a dungeon master

I've been trying to get it to act as a dungeon master unsuccesfully. Do you have any other leads apart from GPT (ChatGPT or AIDungeon)?

Re: GPT Takes the Bar Exam

#79
post #54

Earlier quoted context omitted.

AI can play chess, Go, Jeopardy, StarCraft, Minecraft, Atari games, write essays, summarize text, answer almost any question, drive a car, do above average on a SAT test, fool many people into wondering on Twitter if there were real people typing answers, write chapters of books, act as a dungeon master, write code to spec, win programming competitions, act as a therapist, fold proteins, provide medical diagnosis, vi…

Aside from games, AI can do all these things only with human assistance. It's like saying "Excel can do your taxes". Sure, it can, if you put the correct numbers and formulas into the cells first.

You forget that a lot of these are moving the goal posts.

AI beat the top players at all those games, but new rules were introduced until AI researchers mostly abandoned those efforts. For instance, the starcraft AI found a really good build order that let it quickly build "Stalkers", a versatile but relatively weak unit. I will note that that build order had several factors that humans found very surprising (consistently oversaturating resource collection by a lot), then controlled those stalker units so well even world-champion level players couldn't match them, with double the APM (meaning the human was allowed 2 actions for every 1 action the AI could take).

In chess the rules have become ... almost ridiculous. Due to cheating, in chess championships you now are "only allowed to play the moves any of the top 3 AIs would play twice in a match, after the first 5 moves". In other words the rules are utterly dependent on humans being worse than AIs. Like in starcraft, btw, AIs have gotten so confident beating human players that a general critique of AI players is that they've gotten "agressive, to the point it's insulting to human players", as one Youtube video put it.

AIs ... beat the average human at driving cars, hell, even Tesla's autopilot beats something like the 98% percentile driver in safety, and Waymo's "chauffeur" is apparently much better than that. I've used chauffeur, and driven behind/around a car controlled by chauffeur. You know what my main comment is? Waymo's chauffeur religiously follows traffic laws (speed limit, stopping at "stop" signs, taking excessive time to manoeuvre around footpaths, ...), it's incredibly irritating to share the road with it. And safe? I've never had an accident, but it's blatantly obvious to me: Waymo's chauffeur is a much safer driver than I am. And that's with people reacting very aggressive towards it.

The conclusion from people? "It's not safe enough". Somehow that argument does not mean we're taking away 98% of people's driving licenses ... it's a double standard, in other words. Exactly what that's based on, I can't tell you. My theory is that it's 95% based on the terminator movies and 5% on 'people will lose their jobs'.

Oh and if you widen the definition of AIs to any algorithm, then you'll have to admit: AIs ARE trusted above humans. For the following paragraph AI means "algorithm", not necessarily deep learning. Humans cannot fly the vast majority of planes, boats, rockets without "AI" assistance. By this I mean a whole lot of vehicles, from any modern Passenger plane to quadcopters require control inputs that the human nervous system is fundamentally unable to generate. For example, because they're too fast. Modern boats, oil tankers, large container ships, require control inputs that humans can't generate because they're too slow (yes, you can calculate, on paper, and then give control inputs, but nobody does that. It's too hard). For passenger planes, AIs are more trusted than human pilots: there are "autolanding-only" weather conditions. There are no "human pilot only" weather conditions. There are now harbours that have high fines if you let a human control the ship for certain classes of ships.

There are now passenger flights where the pilot's function is the checklist, closing up the plane, and turning on autopilot. Negotiating with ground control, taxiing, taking off, negotiating with ATC, flying the plane, Negotiating again with ground control, landing, taxiing and parking the plane are all done by AI. Some of these functions are even done by deep learning algorithms.

So I contend: the problem with cars is that humans won't LET AIs control cars, and at the moment that describes 95% of the situation. That's still 5% inaccurate, and it will change. Not because Skynet will take over, but because we're building ever more cars that are safer, more performant, more helpful ... generally better because the human driver is surrounded by ever more algorithms. For instance, one example where people, even good drivers, generally agree Tesla's AI does better than themselves is parking. That's a case where people will gladly hand over the reins to AI. Those cases will grow, and grow and grow until humans are out, legislated out even, as is partly done in aviation today.

And, frankly, the problem is not the question "how do we make an algorithm control the car better", but the answer to the question "if a human gets an AI to do something so difficult/stupid it fails, who pays for damages?". That's at least 50% of the problem, if not more.

The transition from "humans do everything" to "AIs do everything" is quite a bit further along than the impression you're giving. I agree with the conclusion that we're not there yet, but we're pretty far along.

And, finally, I'd like to reverse your argument. Let's say we're two AIs in the year 30000, discussing:

> It's like saying "Humans can do your taxes". Sure, they can, if you put the correct numbers and formulas into a book and school them with that book for 18 years first. How does it make any sense? My taxes are due this year.

Re: GPT Takes the Bar Exam

#80
post #73

Earlier quoted context omitted.

As a counter to this, I found a recent Guardian article on ChatGPT did a nice job of cutting through some of the more optimistic hyperbole surrounding LLMs/AI at the moment and offered instead a more grounded perspective as to what ChatGPT is and what it is not: https://www.theguardian.com/commentisfree/2023/jan/07/chatgp...

try it. get on an use it. It is pretty remarkable. I have been using for all sorts of stuff. It is a brain extension- the ways it is limited at the moment are based on what the builders intended "no, I won't let you build a virtual environment, no I won't roll up a dnd character etc.." but when pressed/jail broken it most certainly can do those things.. I feel like I felt when explaining google search in 1999

I have tried it and it's great. But, because it's so effective, it's easy to over extrapolate and anthropomorphise it. I think it's useful to keep in perspective (as the article and referenced paper - https://arxiv.org/abs/2212.03551 - talk about) that ChatGPT/LLMs are, ultimately, "just" extremely powerful statistical next-token prediction systems.

That doesn't at all invalidate the achievement or belittle the impact such systems will have on society but, keeping this in mind does help to avoid blowing everything out of proportion (which is easy to do because, yep, ChatGPT is genuinely very cool and exciting tech).

Post reply on HN