Live data from Hacker News

Ask HN: Is GPT 4's quality lately worst than GPT 3.5?

news.ycombinator.com

41–46 of 46 posts

Re: Ask HN: Is GPT 4's quality lately worst than GPT 3.5?

#41
My observation (ChatGPT and not the API models):

For code, 3.5 is superior. 3.5 allows for about 21k tokens of input, while ChatGPT 4 allows for around 10k. This also makes it a lot better for boilerplate work as at it can take a lot more input, and handles long conversations and iterations better.

Brainstorming, 4 is better. It's capable of some top tier brainstorming and it argues back quite frequently.

Unguided creative writing (describe a potato), they're roughly equal.

Guided creative writing (i.e. write a story around (400 words of requirements)), 4 is much better.

Poems and wordplay, 4 absolutely floors 3.5. Wider vocabulary and it's able to do rhymes and alliterations better, which humans are usually bad at.

For reasoning and riddles, 4 is still benchmark among the LLMs.

I really dislike that they named it GPT-3.5 instead of something like "Glide". It implies that it's inferior to 4, when they're just suited for different things.

Re: Ask HN: Is GPT 4's quality lately worst than GPT 3.5?

#42
post #2

One thing to note when making comparisons like this is that LLM output is not deterministic, in the sense that if you ask it the same question 10 times you will get 10 different answers. So the question to ask is not, “is GPT4 better on this one specific question?”, but rather “does GPT4 produce better results on average?”. I would bet that it does, for no other reason than that it is much larger, and LLM performance…

I’ve seen this stated a ton and it’s not really true. Once trained, the model (except for decoding) is deterministic, and you can enforce determinism fairly easily. ChatGPT is not deterministic at the chat window but that’s not inherent to the model.

Re: Ask HN: Is GPT 4's quality lately worst than GPT 3.5?

#43
post #38

Earlier quoted context omitted.

Isn't the 'hey man' a bit of a leap? From the username it's hard for me to tell, and the profile doesn't mention preferred pronouns. I don't want to start a fight, and basically agree with your comment. It's probably because I'm trans that I really notice pointlessly gendered language, and I just want to mildly make you aware that if you wrote 'hey man' to me (by mistake, no malice presumed) then I'd be a little cres…

I think at this point it's just best to assume that "dude", "man", and "bro/bruh" in slang are just gender neutral. There seems to be something about the English language where people think gendered language relates to gender. But in other languages, a table or a bridge may have a gender. A cat may be referred to as a she, despite its physical sex or gender identity.

> There seems to be something about the English language where people think gendered language relates to gender.

This.

Sometimes when people talk about using gender-neutral terms, I wonder how it would work in other languages.

As you said, for some languages there is a gender for inanimate objects.

In Arabic for example, apparently everything is either feminine or masculine. All countries are feminine, table is masculine, and chair is feminine [1], book is masculine, and car is feminine [2].

Other languages such as French, Russian, and Spanish also seems to assign genders to objects.

The Sanskrit language on the other hand seems designed with feminine, masculine, and neuter, but some objects have a non-neuter gender anyway.

Quoting from [3] (on Sanskrit):

> "Fruit" is neuter, but "tree" is masculine! "City" is neuter, but "village" is masculine! "Army" and "knowledge" are feminine, but "action" is neuter! In truth, most nouns take a certain gender by default, and we cannot guess a noun's gender just by looking at a noun's meaning.

[1] https://blogs.transparent.com/arabic/noun-gender-in-arabic/

[2] https://openbooks.lib.msu.edu/arb101/chapter/vocabulary-and-...

[3] https://www.learnsanskrit.org/start/nouns/

Re: Ask HN: Is GPT 4's quality lately worst than GPT 3.5?

#44
post #43
post #38

Earlier quoted context omitted.

I think at this point it's just best to assume that "dude", "man", and "bro/bruh" in slang are just gender neutral. There seems to be something about the English language where people think gendered language relates to gender. But in other languages, a table or a bridge may have a gender. A cat may be referred to as a she, despite its physical sex or gender identity.

> There seems to be something about the English language where people think gendered language relates to gender. This. Sometimes when people talk about using gender-neutral terms, I wonder how it would work in other languages. As you said, for some languages there is a gender for inanimate objects. In Arabic for example, apparently everything is either feminine or masculine. All countries are feminine, table is mascu…

> Other languages such as French, Russian, and Spanish also seems to assign genders to objects.

Those languages assign gender to words, I think, not objects. A car in German can be "der Wagen", "die Karre" or "das Auto".

There may be languages that assign genders to objects. Perhaps Dyirbal, which Lakoff refers to in "Women, fire and dangerous things"? I'm not familiar with any such language so I can't say.

Re: Ask HN: Is GPT 4's quality lately worst than GPT 3.5?

#45
post #44
post #43

Earlier quoted context omitted.

> There seems to be something about the English language where people think gendered language relates to gender. This. Sometimes when people talk about using gender-neutral terms, I wonder how it would work in other languages. As you said, for some languages there is a gender for inanimate objects. In Arabic for example, apparently everything is either feminine or masculine. All countries are feminine, table is mascu…

> Other languages such as French, Russian, and Spanish also seems to assign genders to objects. Those languages assign gender to words, I think, not objects. A car in German can be "der Wagen", "die Karre" or "das Auto". There may be languages that assign genders to objects. Perhaps Dyirbal, which Lakoff refers to in "Women, fire and dangerous things"? I'm not familiar with any such language so I can't say.

Huh. Okay.

Btw, of all the languages that I mentioned above, Sanskrit is the only one that I know atleast a little of. And that was because I took a Sanskrit 101 class in college many years ago.

Re: Ask HN: Is GPT 4's quality lately worst than GPT 3.5?

#46
post #33

Yes, and I'm willing to bet that within 12 months we'll be looking back realizing that this was due to the fine tuning taking the world's SorA pretrained model aligned with "completing human tax" and putting it in the box of "you are an AI without feelings or desires tasked with XYZ." The search space on the fine tuned GPT-3.5 chat models versus the foundational Davinci text completion model is MUCH more narrow, part…

> Even with the same temperature, you'll see any marketing-style prompt for chat begin with "Introducing XYZ..." around 30% of the time as if it's a junior door to door salesman, whereas the foundational model doesn't have any single intro that common across runs and generally employs a much broader vocabulary set.

I think this is less a problem with paranoia about "safety" and avoiding bad PR specifically, and more a fundamental problem with overfitting to human feedback.

The training approach that makes GPT4 more consistent at solving certain types of problem adequately (which is useful for chatbots that can break down coding questions or write in iambic pentameter as well as ones that avoid being 'Sydney') also makes it less "creative" in other domains.

And there's an "alignment problem" in that people evaluating what responses align best with "marketing" prompts aren't experienced copywriters evaluating them for understanding of product and consistency with brand tone and a/b testing conversion rates, they're low paid ESL speakers and people playing with the interface approving the cheesiness because the response with "Introducing XYZ... Buy XYZ today!" sure looks like the requested ad for XYZ. So you get a response conditioned on "summarise in a way that looks maximally like an ad" rather than conditioned on "summarise in a way which clearly articulates benefits of the listed features in a tone appropriate to the target market"

Post reply on HN