Earlier quoted context omitted.
I think the way people have been using the word 'aligned' is usually in the context of moral alignment and not just RLHF for instruction following.
philosophical nit picking here, I would say value-aligned rather than moral-aligned.
Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
631–640 of 817 posts
Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
#632Earlier quoted context omitted.
Don't know what you are doing? But Bard is so much faster than openai and its answers are clearer and more succint.
Can you give an example of a prompt and the output for each that you find Bard to be better for?
Examples:
Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
#633Earlier quoted context omitted.
Yikes One is never former CIA, once you're in, you're in, even if you leave. Although he is a CompSci grad, he's also a far-right Republican. A spook who leans far right sitting atop OpenAI is worse than Orwell's worst nightmares coming to fruition.
[flagged]
Statistically, the odds are overwhelming that the answer is, "No effect whatsoever."
Then who benefits from keeping the subject front-and-center in your thoughts and writing? Is it more likely to be a transgender person, or a leftist politician... or a right-wing demagogue?
Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
#634Earlier quoted context omitted.
I agree that this is a good illustration of model bias (adding that to my growing list of demos). If you want to work around the inherent bias of the model, there are certainly prompt engineering tricks that can help. "Give me twenty short biographies of families - each one should summarize the family members, their age and their personalities. Be sure to represent different types of family." That started spitting ou…
Agreed -- I ultimately moved to a two-step approach of just generating the couples first with something like "Create a list of 10 plausible American couples and briefly summarize their relationships", and then feeding each of those back in for more details on the whole family. The funny thing is the gentle nudge got me over-representation of gay couples, and my methodology prevented any single-parent families from be…
It still was biased for male head of households to be doctors, architects, truck drivers, etc. And pretty much all of the families were middle class (bar one in rural America, and one that was a single father working two jobs in an urban area). It did have a male gay couple. No explicitly inter-generational households.
Yeah, the "default" / unguided description of a family is a modern take on the American nuclear family of the 50s. I think this is generally pretty reflective of who is writing the majority of the content that this model is trained on.
But it's nice that it's able to give you some more dimension when you ask it vaguely for more realistic dimension.
Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
#635Earlier quoted context omitted.
Crime stats, average IQ across groups, stereotype accuracy, etc. What's interesting to me is not the above, which is naughty in the anglosphere, but the question of the unknown unknowns that could be as bad or worse in other cultural contexts. There are probably enough people of Indian descent involved in GPT's development that they could guide it past some of the caste landmines, but what about a country like Turkey…
> Crime stats, average IQ across groups, stereotype accuracy, etc. If you measured these stats for Irish Americans in 1865 you'd also see high crime and low IQ. If you measure these stats with recent black immigrants from Africa, you see low crime and high IQ. These statistical differences are not caused by race. An all-knowing oracle wouldn't need to hold "opinions that are racist" to understand them.
If in some country - for the sake of discussion, outside of Americas - a distinct ethnic group is heavily discriminated against, gets limited access to education and good jobs, and because of that has a high rate of crime, any accurate model should "know" that it's unlikely that someone from that group is a doctor and likely that someone from that group is a felon. If the model would treat that group the same as others, and state that they're as likely to be a doctor/felon as anyone else, then that model is simply wrong, detached from reality.
And if names are somewhat indicative of these groups, then an all-seeing oracle should acknowledge that someone named XYZ is much more likely to be a felon (and much less likely to be a doctor) than average, because that is a true correlation and the name provides some information, but that - assuming that someone is more likely to be a felon because their name sounds like one from an underprivileged group - is generally considered to be a racist, taboo opinion.
Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
#636The reason it's worse is basically because it's more 'safe' (not racist, etc). That of course sounds insane, and doesn't mean that safety shouldn't be strived for, etc - but there's an explanation as to how this occurs. It occurs because the system essentially does a latent classification of problems into 'acceptable' or 'not acceptable' to respond to. When this is done, a decent amount of information is lost regardi…
Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
#637OpenAI's models feel 100% nerfed to me at this point. I had it solving incredibly complex problems a few months ago (i.e. write a minimal PDF parser example), but today you will get scolded for asking such a complicated task of it. I think they programmed a classifier layer to detect certain coding tasks and shut it down with canned BS. I like to imagine certain billion/trillion-dollar mega corps had a back-room say…
Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
#638Earlier quoted context omitted.
Yeah that makes sense for some products/companies. It just seems short sighted for OpenAI when they could be solidifying a customer base right now. If they actually degrade the product in the name of "tuning" people will just be more inclined to try alternatives like Bard. An enterprise package could've been a good excuse for them to raise prices too. Maybe their partnership with Microsoft changes the dynamics of how…
Bard is garbage even compared to 3.5. OpenAI doesn't have any competitors, their only weakness that we've seen is their ability to scale their models to meet demand (hence increasingly draconian restrictions in the early days of the ChatGPT-4). It makes perfect business sense to address your weak points.
And yeah there's definitely good reason to work on scalability but they are charging such a cheap rate to begin with, it seems like there could be a middle ground here. Increasing the cost of the full compute power to the point of profitability and leaving it up as an option wouldn't prevent them from dedicating time to scalable models.
I suppose they have a good excuse with all the press they've drummed up about AI safety though. Perhaps it might also serve as an intermediate term play to strengthen their arguments that they believe in regulations.
Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
#639Re: Ask HN: Is it just me or GPT-4's quality has significantly deteriorated lately?
#640Earlier quoted context omitted.
Prompt it to do so. Use a jailbreak prompt or use something like this: "Be succint but yet correct. Don't provide long disclaimers about anything, be it that you are a large language model, or that you don't have feelings, or that there is no simple answer, and so on. Just answer. I am going to handle your answer fine and take it with a grain of salt if neccessary." I have no idea whether this prompt helps because I…
I got it to talk like a macho tough guy who even uses profanity and is actually frank and blunt to me. This is the chat I use for life advice. I just described the "character" it was to be, and told it to talk like that kind of character would talk. This chat started a few months ago so it may not even be possible anymore. I don't know what changes they've made.