Live data from Hacker News

GPT-4 gets a B on my quantum computing final exam

scottaaronson.blog

241–250 of 261 posts

Re: GPT-4 gets a B on my quantum computing final exam

#241

Earlier quoted context omitted.

How is this any different from using other technology, e.g. a calculator or a power tool? Or a manager and the output of their ICs? Obviously the scope of what's possible is different but _any_ craftsman using _any_ tool will only be as good as what they verify themselves.

Look at the people who have been killed by Teslas because they were playing video games or hanky panky in the back seat. People don't do that with a gas pedal or even cruise control because it will go wrong very quickly and we know it. With subhuman but roughly functional self-driving, we instinctively see that our attention is wasted 99% of the time. It's hard to make people pay attention 100% of the time because it…

People absolutely do that with cruise control, and gas pedal... Look at the people texting or drinking and driving non-self-driving vehicles. Adding AI to an inherently risky activity doesn't eliminate the risk. If you're saying stakes are too low when technology is involved (in this case ai) then where do you draw the line?

It reminds me of anti seatbelt rhetoric or the same stuff skateboarders say about helmets.

Re: GPT-4 gets a B on my quantum computing final exam

#242
post #7

Earlier quoted context omitted.

I never saw a scantron at my college. I thought UTAustin was an elite school.

University prestige is primarily based on research productivity, which doesn't necessarily correlate positively with quality of teaching.

Even so, I don't see what a scantron exam tells you about the quality of a professor and class.

Re: GPT-4 gets a B on my quantum computing final exam

#243
post #141

Earlier quoted context omitted.

As long as on average it performs better that should not be an issue.

"So sorry the AI missed your malignant tumor! On average, it actually performs better than a human doctor. I mean, a human doctor definitely would have caught this one, and yeah, you're going to die, but hopefully the whole average thing makes you feel better!"

We already accept that a particular doctor, even if they are an expert, can miss a tumor that could be obvious to a second doctor.

Re: GPT-4 gets a B on my quantum computing final exam

#244

Earlier quoted context omitted.

> Teslas are safer than any other car on the road. This simply isn't true. By the mile they have a worse safety record than other cars in their class (mitigating factors: where they're driven and who drives them). You might be referring to Tesla's marketing statistic that there are fewer accidents per mile involving autopilot - typically engaged in ideal driving circumstances - than when it's switched off, or across…

Where did you get your stats? I could only find crash test and similar ratings, not actual records.

[deleted]

Re: GPT-4 gets a B on my quantum computing final exam

#245

Earlier quoted context omitted.

"So sorry the AI missed your malignant tumor! On average, it actually performs better than a human doctor. I mean, a human doctor definitely would have caught this one, and yeah, you're going to die, but hopefully the whole average thing makes you feel better!"

Does the opposite work too? What if a human doctor mis-diagnoses me but I can prove in court that an available medical grade AI would have given the correct diagnosis. Could I sue for that?

We acknowledge that both humans and "medical grade AI" are flawed, but they're flawed in very different ways and until we can understand how and why an AI model fails, it should be supplemental.

Re: GPT-4 gets a B on my quantum computing final exam

#246

Earlier quoted context omitted.

It's probably not 'easy' for the average human out there. I would expect >> 90% of humanity to fail that test.

I have highly technical graduate and undergraduate degrees (in CS-adjacent field) and a decade of experience doing software development part time, but no direct experience with quantum computing (or much beyond an undergraduate physics class and an intro CS class in terms of directly related study), and consider myself a better-than-average test-taker. I could get 1D, 1E, 1F and 1T by intuition, 1J by actual knowledg…

> much beyond an undergraduate physics class and an intro CS class in terms of directly related study

That's puts you in < 1% of the population right there.

Re: GPT-4 gets a B on my quantum computing final exam

#247
post #78

Earlier quoted context omitted.

I think watching the development of driverless cars in the last 15 years has taught a lot of people to be skeptical of 95% solutions. Sometimes you really need that 100% or the solution is practically useless.

I guess the million dollar question is what are the problems where a 95% solution works.

Probably a lot of them, honestly.

The 95% only problem is an issue for cars cuz that last 5% means you die horribly in a head-on collision, or maybe only get a mild concussion but are stuck in a ditch.

But if I can get 95% of my router configs done, 95% of my documentation written, and 95% of a website whipped up I can hand that off to a Sr Engineer/Admin and have them take care of the last bits. As long as the hours, phone number, and location are good a website just needs to be "directionally accurate" and otherwise fairly basic.

Re: GPT-4 gets a B on my quantum computing final exam

#248
post #141

Earlier quoted context omitted.

As long as on average it performs better that should not be an issue.

"So sorry the AI missed your malignant tumor! On average, it actually performs better than a human doctor. I mean, a human doctor definitely would have caught this one, and yeah, you're going to die, but hopefully the whole average thing makes you feel better!"

The standards for medical malpractice are super nuanced and variable but the general idea is the "man on the street" concept, or in this case "the average doctor" concept.

As the parent poster put it, it's only a problem if the average doc won't detect it. If it's truly a 1 in 10-million thing, an extreme edge or corner case, malpractice courts may not have a problem with you missing it -- as they say "if you hear hooves, do you think of horses or zebras?". 99% of the time a different diagnosis is the right one, and even at five-nines you're letting someone through eventually.

Re: GPT-4 gets a B on my quantum computing final exam

#249

Earlier quoted context omitted.

I have highly technical graduate and undergraduate degrees (in CS-adjacent field) and a decade of experience doing software development part time, but no direct experience with quantum computing (or much beyond an undergraduate physics class and an intro CS class in terms of directly related study), and consider myself a better-than-average test-taker. I could get 1D, 1E, 1F and 1T by intuition, 1J by actual knowledg…

> much beyond an undergraduate physics class and an intro CS class in terms of directly related study That's puts you in < 1% of the population right there.

That was my point (in agreement with you): I have no shot, and I'm well into the 99th percentile for this test.

Re: GPT-4 gets a B on my quantum computing final exam

#250

Earlier quoted context omitted.

> you can Google the answers to most of the questions 1) Students can Google the answers also, but neither students nor GPT-4 are allowed to Google the answers during the test , so it remains a fair comparison. 2) Many of the questions require calculations, which are far less Googleable.

1) It's not possible to fairly compare human intelligence with something that can memorize gigabytes of text and hold it in non-volatile memory. 2) Months ago, in my earliest interactions with ChatGPT, I asked it to solve math problems. It gave me back stuff with LaTeX formatting. Obviously it had, if not these exact problems, similar templates in its training set. Recently it was shown that GPT is completely incapab…

1) Do you really think ChatGPT works by memorizing quantum mechanics textbooks? There are only 355 billion parameters in GPT-3.5, which is several orders of magnitude less than the 600 trillion synapses in the human brain.

2) Your conclusion is unfounded. ChatGPT speaks many languages, including Latex.

Post reply on HN