Live data from Hacker News

GPT-4 gets a B on my quantum computing final exam

scottaaronson.blog

211–220 of 261 posts

Re: GPT-4 gets a B on my quantum computing final exam

#211

Earlier quoted context omitted.

I used to agree with you. The paper that made me unsure was "Transformers learn in-context by gradient descent" [1]. Basically, the model learns weights that let it run gradient descent at inference time in order to do in context learning. If transformers can learn this, then I think they can learn almost anything given enough compute/parameters/data. Of course, even if this is true then it's possible that there simp…

Too theoretical to be pivotal. Years ago we had results about how all sorts of learning was reducible to kernel SVM and that drove a big fad for kernel SVM research at NeurIPS and other hot conferences. Nobody cares about SVM anymore. So we don't have a good enough grasp on the theory of AI to predict what is going to be a breakthrough. Only rigorous empirical results convince. Second, the idea of meta-learning gradi…

I'm not suggesting that you can use meta-learning to discover gradient descent, that would not be surprising. My claim is that GPT-4 (probably even GPT-2) is almost certainly already using gradient descent internally inside it's own activations.

How else can it learn a random small neural network given enough samples in the prompt?

Re: GPT-4 gets a B on my quantum computing final exam

#212

Earlier quoted context omitted.

Bilingual/Multilingual LLMs are human level translators more or less. The only way you can think "not on the cusp" is if you haven't actually used GPT-4 for translation. Use it and you'll be set straight pretty quickly.

I was curious so I asked GPT-4 to translate a bit of french literature, here it is, along with the official translation (I'll let you guess which is which) ------- The tale I'm about to unfold commenced with a mysterious handwriting on an envelope. Within the pen strokes that outlined my name and the address of the Fossil Review, a publication I was associated with and where the letter had been forwarded from, there…

GPT-3.5-turbo tries its best (not nearly as good as GPT-4!) when I tell it to prioritize fluency vs. fidelity:

--

Let me commence by telling you, my dear reader, about a curious event that occurred in my life. It all started with an unknown handwriting on an envelope that arrived at my doorstep. The feathers that traced my name and the address of the Fossils Magazine, with whom I was a collaborator, had a peculiar mixture of violence and sweetness. As I pondered over the possible content and sender of this mysterious letter, a hazy yet powerful intuition crept up my spine, evoking the image of a "stone thrown into a pond filled with frogs." With this realization, I came to acknowledge that my life had become rather stagnant in recent times. It was at this moment that I opened the letter, and I was left with the feeling that it was either a refreshing whiff of fresh air or an unwelcome draft.

The swift and eloquent handwriting read as follows:

"Dear Sir, I read your article on Mount Analogue, and I believed myself to be the only one convinced of its existence, until now. Today, we are two, tomorrow we shall be ten, and perhaps more, and we can attempt the expedition. We must make contact as soon as possible. Please call me at one of the following numbers the moment you can. I eagerly await your response.

--

The translation also included the following continuation as part of the translation itself ...

--

Ah, how exciting! The words on that paper spoke volumes of adventure and discovery. It was as if the sender had thrown a gauntlet at my feet, daring me to embark on a journey like no other. The thought of being one of the only few people to explore the wonders of Mount Analogue, a place shrouded in mystery, was exhilarating. I could feel the blood rushing through my veins, and my heart pounding with anticipation.

And so, my friends, I picked up the phone, my fingers quivering with excitement, and dialed the number on the paper. Who knew where this quest would lead us, or what unknown secrets we would uncover? But one thing was certain: it was time to break free from the chains of monotony and embark on a journey of a lifetime.

--

Maybe it was getting the continuation from contamination in the training set (if Mount Analogue is already there) or maybe it was making up the story or maybe some of each.

Re: GPT-4 gets a B on my quantum computing final exam

#213
post #141

Earlier quoted context omitted.

As long as on average it performs better that should not be an issue.

"So sorry the AI missed your malignant tumor! On average, it actually performs better than a human doctor. I mean, a human doctor definitely would have caught this one, and yeah, you're going to die, but hopefully the whole average thing makes you feel better!"

Does the opposite work too? What if a human doctor mis-diagnoses me but I can prove in court that an available medical grade AI would have given the correct diagnosis. Could I sue for that?

Re: GPT-4 gets a B on my quantum computing final exam

#214

It's impressive on the surface... But give me Google and I too can pass some extremely challenging tests that I actually know nothing about. If someone told you I passed the "bar" exam, you would be impressed. If they then said, he had google and lots of time. You wouldn't be that impressed anymore. The impressive thing here is that AI can read and answer questions... it's not overly impressive that it can use inform…

OK, try to do it like GPT does: by saying the first thing that comes to your mind. You can have any amount of preliminary training.

Re: GPT-4 gets a B on my quantum computing final exam

#215
post #25

This is obviously very cool, but at this point — who knows what I’ll say in a year — my concern with these LLMs is that they’re in the uncanny valley. Here’s one passing a very difficult test. Amazing! Now, rely on it to build a nuclear doohickey for a power station or a multi-billion dollar device for CERN or anything really and, well, no. So humans still have to check the output, and now we’re in that situation whe…

A big factor is that the error rate is compounding. Tiny mistakes would become bigger and bigger if you would run GPT in a loop.

It needs interaction and calibration from a human to keep it in check. Even we as humans are using all kinds of different feedback mechanisms to validate our thought process.

Though this might be a different way of hooking GPT up to reality.

Re: GPT-4 gets a B on my quantum computing final exam

#216

Earlier quoted context omitted.

WHAT. It's got the second half wrong. Google: Have you actually used GPT-4 for translation? Really, it's a joke that the story of only conveying explicit meaning can be easily solved by just trying. DeepL: Have you actually used GPT-4 for translation? Really, it's a joke that all this talk about conveying only explicit meaning can be easily solved by just trying it out. Mine: Have you actually used GPT-4 for translat…

Here's a couple more from GPT4 (since it's random every time because of temperature) GPT-4を翻訳に実際に使ったことがありますか?本気で、伝えたい意味だけを伝えるという話は、ちょっと試してみれば簡単に解決できると思うのですが。 実際にGPT-4を翻訳に使ったことがありますか?本当に、試してみるだけで簡単に払拭できると思うのに、この「明確な意味だけが伝わる」話ばかりで。

  本気で、伝えたい意味だけを伝えるという話は、ちょっと試してみれば簡単に解決できると思うのですが。
"In seriousness, I think the story that [subject] tells the meaning [it/he/they] wants to tell, should be easily solvable by trying a bit."

or "Seriously, the story of telling the meaning [subject] wants to tell, should be easily solvable by trying a bit."

  本当に、試してみるだけで簡単に払拭できると思うのに、この「明確な意味だけが伝わる」話ばかりで。
 
"Really, I think it'll be easily swept away by just trying, but there are so much of this 'only clear meaning is conveyed' stories."

I'm almost feeling that GPT-4 should be eligible for human rights, especially astonishing that they dropped explicit specification of "afternoon" that don't work well. But also interesting it's failing to keep the intent of the whole sentence unlike 3.5 and even more primitive NN translation engines.

Re: GPT-4 gets a B on my quantum computing final exam

#217
post #73

Reading stuff like this, one thing I cannot stop wondering is this: If Ai can be trusted to do all the trivial tasks and if non trivial tasks require a scaffold of trivial practice, where are we going to keep finding the people qualified enough to actually do the non trivial stuff?

This is what worries me the most when it comes to job losses and downward pressure on wages.

If we hit the point where a senior engineer reviewing the output from an AI can effectively replace a team with a senior engineer and 5 junior engineers, you’re right that it isn’t wise to simply replace all those junior engineers with the AI. Unless you’re confident that the AI will be able to replace that senior engineer soon, you _need_ junior engineers who can build up the experience needed to step into that role once the senior engineer moves on.

But you _can_ get rid of 3 or 4 of those junior engineers. The remaining junior engineer(s) will still need mentorship from the senior engineer and will still do “trivial” work that needs oversight, but they will be able to pump that out at a high enough rate to replace a few of their pre-AI peers.

Basically, I’d imagine the org chart at most companies will look pretty similar to now, except you’ll just have fewer people at each level.

Re: GPT-4 gets a B on my quantum computing final exam

#218
post #77

Earlier quoted context omitted.

Agreed. A lot of these responses read like they haven't actually tried it yet. Which is also interesting, I myself actively put off trying it until I eventually gave in. It seems a lot of us are doing the same, maybe its a case of "how good could it actually be?"

Not trying it yet is fine. Making declarative statements on a product you haven't even used is just absurd. Dude clearly hasn't used GPT for translation before and his next reply is telling me the ways GPT should fail based on his pre-conceived notions of its abilities. Except i have actually extensively tested(publicly too) LLMs for translation (even before GPT-4) and basically everything he says is just plain wrong…

Apparently GPT-4 can't handle "all this talk about only getting explicit meaning across would be easily dispelled in an afternoon if you only bothered to try.", which isn't simple as "Ace Attorney" but I'd think it's still a small stretch to say "everything he says is just plain wrong".

Re: GPT-4 gets a B on my quantum computing final exam

#219
post #177

Earlier quoted context omitted.

It’s abundantly clear that we are going to expect near-perfect reliability from autonomous vehicles. This isn’t necessarily illogical; they operate in a different context than humans do. We expect humans to make mistakes and we have various ways of dealing with the consequences (eg lawsuits targeted at the responsible individual). The argument from statistics doesn’t appear likely to win the kind of societal approval…

Your point about liability is valid, but it’s far from “abundantly clear” that society needs machines to be near perfect as opposed to just significantly better than humans . Solve the liability problem and I would 100% take a machine that performs 30% better than a human or helps a human perform 30% better every time because it means fewer humans die on the road.

> I would 100% take a machine that performs 30% better than a human or helps a human perform 30% better every time because it means fewer humans die on the road.

It's surprisingly hard to establish those parameters though, since (i) the more meaningful indicators of good performance (driver error fatalities) happen only every million or so miles, even less frequently once you've narrowed your pool down to errors made by sober drivers that weren't racing or attempting to driver in conditions unsuited to electronic assistance, so that's a lot of real world road use required to establish statistically significant evidence a machine is a 30% or so safe than a human driver (ii) complex software doesn't improve monotonically, so really you need that amount of testing per update to be confident that the next minor version of something "30% safer" hasn't introduced regression bugs which mean that it is now a bit worse than the average driver and (iii) performance in different road conditions is likely highly variable such that it might be both 30% better overall and 30x as likely to cause an accident if not disengaged in that particular circumstance.

To make a valid assessment of the overall safety impact, you'd also have to factor in that (iv) the worst drivers who skew the stats are the generally the ones least likely to buy it and (v) if it's fully autonomous driving, road use would increase substantially and whilst that may have other benefits the likely outcome from substantial increase in miles driven using tech that's only marginally better than a human is more humans dying on the road

Re: GPT-4 gets a B on my quantum computing final exam

#220
post #143
post #25

This is obviously very cool, but at this point — who knows what I’ll say in a year — my concern with these LLMs is that they’re in the uncanny valley. Here’s one passing a very difficult test. Amazing! Now, rely on it to build a nuclear doohickey for a power station or a multi-billion dollar device for CERN or anything really and, well, no. So humans still have to check the output, and now we’re in that situation whe…

I don't think the data supports the fear that AI assisted driving is more or newly dangerous when compared to fully human drivers. Teslas are safer than any other car on the road. Yes they’re newer, but by mile they’re safer. So the fear that “we should be careful” is understandable but ultimately unfounded. We are being careful .

> Teslas are safer than any other car on the road.

This simply isn't true. By the mile they have a worse safety record than other cars in their class (mitigating factors: where they're driven and who drives them). You might be referring to Tesla's marketing statistic that there are fewer accidents per mile involving autopilot - typically engaged in ideal driving circumstances - than when it's switched off, or across all other drivers. But that's not very meaningful data. Some would argue that citing marketing figures from a company with a track record for obfuscation and dishonesty as "the data" (and really, there isn't much better data available to the average person) is an indication we're not being careful.

Post reply on HN