Live data from Hacker News

GPT-4 gets a B on my quantum computing final exam

scottaaronson.blog

251–260 of 261 posts

Re: GPT-4 gets a B on my quantum computing final exam

#251
post #177

Earlier quoted context omitted.

Your point about liability is valid, but it’s far from “abundantly clear” that society needs machines to be near perfect as opposed to just significantly better than humans . Solve the liability problem and I would 100% take a machine that performs 30% better than a human or helps a human perform 30% better every time because it means fewer humans die on the road.

> Your point about liability is valid How is it valid? It's currently a solved problem. The driver of the car is still held liable because they are still, ultimately, driving the car. Tesla's with autopilot seem to make drivers safer, and acting like we don't know who is liable in the event of a crash is just a red herring.

I meant for fully autonomous systems, but yeah I agree with you.

Re: GPT-4 gets a B on my quantum computing final exam

#252

Earlier quoted context omitted.

"So sorry the AI missed your malignant tumor! On average, it actually performs better than a human doctor. I mean, a human doctor definitely would have caught this one, and yeah, you're going to die, but hopefully the whole average thing makes you feel better!"

The standards for medical malpractice are super nuanced and variable but the general idea is the "man on the street" concept, or in this case "the average doctor" concept. As the parent poster put it, it's only a problem if the average doc won't detect it. If it's truly a 1 in 10-million thing, an extreme edge or corner case, malpractice courts may not have a problem with you missing it -- as they say "if you hear ho…

I always think of comparisons to aviation. There are a million and one things that can go wrong when flying a plane, but it's still one of the safest ways to travel. That's because regulations and safety standards are to such a high degree that we simply don't consider injury or death an acceptable outcome.

Whenever someone says "as long as it's better than a human", that's where my mind goes. We shouldn't be satisfied with just being better than a human. We shouldn't be satisfied with five nines! I don't really care about what courts have a problem with — my point is just that our goal should be zero preventable deaths, not just moving from humans to AI once the latter can be better on average than the former.

Re: GPT-4 gets a B on my quantum computing final exam

#253
post #71

Earlier quoted context omitted.

Bilingual/Multilingual LLMs are human level translators more or less. The only way you can think "not on the cusp" is if you haven't actually used GPT-4 for translation. Use it and you'll be set straight pretty quickly.

I haven't used GPT-4 for translation so I acknowledge I might be wrong. But GPT-3 was such an irredeemably terrible poet that it made me sceptical that this type of software could ever develop aesthetic taste or artistic vision. Moreover – and I understand this is an uncharitable thing to say, but it is my honest observation – time and again I have noticed the inability of AI cheerleaders to judge literature on its a…

Gpt-4 is a far better poet than gpt-3. It isn’t world-class and may be missing some ineffable poetic soul, but its attempts are definitely notable and interesting. They seem mostly better than what I could write with significant effort.

Oh and I should emphasize that the quality and specificity of the prompt has a huge effect on the output.

I agree that gpt-3 was pretty trash at poetry, at least compared to human standards. It was impressive for AI, obviously.

Re: GPT-4 gets a B on my quantum computing final exam

#254

Earlier quoted context omitted.

> much beyond an undergraduate physics class and an intro CS class in terms of directly related study That's puts you in < 1% of the population right there.

That was my point (in agreement with you): I have no shot, and I'm well into the 99th percentile for this test.

Right, forgive me I understood it the other way around.

Re: GPT-4 gets a B on my quantum computing final exam

#255
> In general, I’d say that GPT-4 was strongest on true/false questions and (ironically!) conceptual questions—the ones where many students struggled the most. It was (again ironically!) weakest on calculation questions, where it would often know what kind of calculation to do but then botch the execution.

It'd be great if chain-of-thought / show-your-work type prompts became the default for anything involving complex, multi-step calculations or logic.

GPT-4 would have almost certainly gotten a higher score on the calculation questions if those methods were used.

Re: GPT-4 gets a B on my quantum computing final exam

#256

> In general, I’d say that GPT-4 was strongest on true/false questions and (ironically!) conceptual questions—the ones where many students struggled the most. It was (again ironically!) weakest on calculation questions, where it would often know what kind of calculation to do but then botch the execution. It'd be great if chain-of-thought / show-your-work type prompts became the default for anything involving complex…

Eh, even when asked specifically to show it's work GPT-4 still frequently makes calculation errors. It's just one of the limitations of current LLMs, and can easily be solved by integration with Wolfram, or even just a basic calculator.

Re: GPT-4 gets a B on my quantum computing final exam

#257

Earlier quoted context omitted.

> But also interesting it's failing to keep the intent of the whole sentence unlike 3.5 It's because it "knows too much". To anthropomorphise a little: its "expectations" of what should be. To anthropomorphise less: GPT-4 is overfitted. GPT-style language models are pretty amazing, but they're not a complete explanation of human language, and can't quite represent it properly. > I'm almost feeling that GPT-4 should b…

(GPT4) In this script from a Doctor Who episode, Clara and the Doctor are having a conversation about the internet. Doctor Who is a British science fiction television series that follows the adventures of the Doctor, a Time Lord from the planet Gallifrey, who travels through time and space in the TARDIS, a time-traveling spaceship. Clara, the Doctor's companion, is trying to access the internet but is unable to find…

(GPT-4 plus a regular expression)

In this script from a Garfield comic, Jon and Garfield are having a conversation about the internet. Garfield is an American comic strip and multimedia franchise that follows the adventures of Garfield, a cat from the planet Earth, who enjoys lasagna in Jon Arbuckle's house, a suburban domicile.

Jon, Garfield's owner, is trying to access the internet but is unable to find it. He asks Garfield about its whereabouts, and Garfield seems to be confused by the question, as the internet is not something that can be physically found.

Garfield then mentions the time as "twelve oh seven," while Jon's clock shows "half past three." This discrepancy in time indicates that they are likely in different time zones, as Garfield implies. In the context of Garfield, this could also mean Jon's clock is wrong, since Garfield is usually right.

Jon is concerned about whether the time difference will affect his phone bill, to which Garfield replies that he dreads to think about the potential cost. This adds a bit of humor to the scene, as Garfield often has a nonchalant attitude towards everyday human concerns.

Overall, this script showcases the humorous and whimsical nature of Garfield, with the characters engaging in a lighthearted conversation that intertwines elements of fantasy and everyday life.

Re: GPT-4 gets a B on my quantum computing final exam

#258
post #177

Earlier quoted context omitted.

Your point about liability is valid, but it’s far from “abundantly clear” that society needs machines to be near perfect as opposed to just significantly better than humans . Solve the liability problem and I would 100% take a machine that performs 30% better than a human or helps a human perform 30% better every time because it means fewer humans die on the road.

> I would 100% take a machine that performs 30% better than a human or helps a human perform 30% better every time because it means fewer humans die on the road. It's surprisingly hard to establish those parameters though, since (i) the more meaningful indicators of good performance (driver error fatalities) happen only every million or so miles, even less frequently once you've narrowed your pool down to errors made…

Billions of miles are driven every day.

Re: GPT-4 gets a B on my quantum computing final exam

#259
post #258

Earlier quoted context omitted.

> I would 100% take a machine that performs 30% better than a human or helps a human perform 30% better every time because it means fewer humans die on the road. It's surprisingly hard to establish those parameters though, since (i) the more meaningful indicators of good performance (driver error fatalities) happen only every million or so miles, even less frequently once you've narrowed your pool down to errors made…

Billions of miles are driven every day.

And yet the total mileage driven by a huge variety of autonomous systems over research programmes dating back a decade is of the order of 20-30 million. This disparity obviously supports my point about the difficulty of establishing statistically significant evidence that a particular software update on a particular platform is less lethal than the human driver [in a given set of circumstances] based on events which are very rare on a per mile basis, particularly if the baseline performance gap isn't that large.

The fact that overall road use is so high that a sufficiently bad regression bug in sufficiently widely-deployed software could rack up a massive body count within hours obviously doesn't make the case for introducing something believed to only be a marginal improvement any stronger.

Re: GPT-4 gets a B on my quantum computing final exam

#260

Earlier quoted context omitted.

1) It's not possible to fairly compare human intelligence with something that can memorize gigabytes of text and hold it in non-volatile memory. 2) Months ago, in my earliest interactions with ChatGPT, I asked it to solve math problems. It gave me back stuff with LaTeX formatting. Obviously it had, if not these exact problems, similar templates in its training set. Recently it was shown that GPT is completely incapab…

1) Do you really think ChatGPT works by memorizing quantum mechanics textbooks? There are only 355 billion parameters in GPT-3.5, which is several orders of magnitude less than the 600 trillion synapses in the human brain. 2) Your conclusion is unfounded. ChatGPT speaks many languages, including Latex.

Beautiful theory, meet ugly fact: https://aisnakeoil.substack.com/p/gpt-4-and-professional-ben...
Post reply on HN