Earlier quoted context omitted.
When chat gpt first came out I was able to feed it some text to parse and then create python scripts to process similar texts and create csv and excel files from those. I was able to create a basic working python scripts in 1-2 hours. And very complex scripts over a couple of days. I recently tried to do the same with chatgpt again and simply am unable to. I wish I had saved the exact text I fed into chat gpt then so…
ChatGPT has a query history. Can't you fish the first queries you made on your account?
Ask HN: Has degradation in the quality of ChatGPT and Claude been proven?
31–40 of 44 posts
Re: Ask HN: Has degradation in the quality of ChatGPT and Claude been proven?
#32One of the things I've personally observed is that ChatGPT has become very verbose these days. Previously, it used to return the right amount of information in most contexts, and I can't get that behavior back with prompts asking it to be concise, because then it'll just omit important parts, prioritizing providing a extremely high-level summary that elucidates very little. No opinion on Claude because I've not had a…
+1 on verbosity - it happened when switching from 4t to 4o I think, and personally I don’t like it. Should be fizable with system prompt though.
Re: Ask HN: Has degradation in the quality of ChatGPT and Claude been proven?
#33Re: Ask HN: Has degradation in the quality of ChatGPT and Claude been proven?
#34I couldn't find the one I was looking for but this is one of them.
https://arxiv.org/abs/2310.06452
Edit:
This tweet also has a screenshot showing degraded evals from RLHF from base model.
https://x.com/KevinAFischer/status/1638706111443513346?t=0wK...
Re: Ask HN: Has degradation in the quality of ChatGPT and Claude been proven?
#35> If there is indeed no degradation how could the perceived degradation be explained? By being disproportionately impressed previously. Maybe in the early days people were so impressed by their little play experiments they forgave the shortcomings. Now that the novelty is wearing off and they try to use it for productive work, the scales tipped and failures are given more weight.
When chat gpt first came out I was able to feed it some text to parse and then create python scripts to process similar texts and create csv and excel files from those. I was able to create a basic working python scripts in 1-2 hours. And very complex scripts over a couple of days. I recently tried to do the same with chatgpt again and simply am unable to. I wish I had saved the exact text I fed into chat gpt then so…
I would have a really hard time saying that 4 in April 2023 was better than 4o now though.
I have always wonder if it matters what time of day you are using it too. I feel like 4am EST works better than 4PM EST but it is so hard to judge. I think there is so much difference too with just how the prompt is phrased so it ends up feeling like some days it is is good and some days it sucks.
That is coupled with I have got a bad result before, opened a new chat window, pasted the exact same prompt and got a good result.
If I had to bet, I imagine it is like flipping quarters. Sometimes you will get runs of heads, sometimes runs of tails and sometimes a real mixed bag of both.
Re: Ask HN: Has degradation in the quality of ChatGPT and Claude been proven?
#36> If there is indeed no degradation how could the perceived degradation be explained? By being disproportionately impressed previously. Maybe in the early days people were so impressed by their little play experiments they forgave the shortcomings. Now that the novelty is wearing off and they try to use it for productive work, the scales tipped and failures are given more weight.
Re: Ask HN: Has degradation in the quality of ChatGPT and Claude been proven?
#37The same output in week 0 was rated as a 7. After 6 weeks of rating LLM outputs, especially as the pipeline improved them, was a lower score later.
Re: Ask HN: Has degradation in the quality of ChatGPT and Claude been proven?
#38> If there is indeed no degradation how could the perceived degradation be explained? By being disproportionately impressed previously. Maybe in the early days people were so impressed by their little play experiments they forgave the shortcomings. Now that the novelty is wearing off and they try to use it for productive work, the scales tipped and failures are given more weight.
They tweak it to make it safer. Every time that happens, it gets a little dumber.