My interpretation of the progress. 3.5 to 4 was the most major leap. It went from being a party trick to legitimately useful sometimes. It did hallucinate a lot but I was still able to get some use out of it. I wouldn't count on it for most things however. It could answer simple questions and get it right mostly but never one or two levels deep. I clearly remember 4o was also a decent leap - the accuracy increased su…
> I could essentially replace it with Google for basic to slightly complex fact checking. I know you probably meant "augment fact checking" here, but using LLMs for answering factual questions is the single worst use-case for LLMs.
OpenAI Progress
221–230 of 372 posts
Re: OpenAI Progress
#222Earlier quoted context omitted.
I'd love to know more about how OpenAI (or Alec Radford et al.) even decided GPT-1 was worth investing more into. At a glance the output is barely distinguishable from Markov chains. If in 2018 you told me that scaling the algorithm up 100-1000x would lead to computers talking to people/coding/reasoning/beating the IMO I'd tell you to take your meds.
I don't have a source for this (there's probably no sources from anything back then) but anecdotally, someone at an AI/ML talk said they just added more data and quality went up. Doubling the data doubled the quality. With other breakthroughs, people saw diminishing gains. It's sort of why Sam back then tweeted that he expected the amount of intelligence to double every N years. I have the feeling they kept on this u…
Re: OpenAI Progress
#223Earlier quoted context omitted.
I find GPT-5's story significantly better than text-davinci-001
I really wonder which one of us is the minority. Because I find text-davinci-001 answer is the only one that reads like a story. All the others don't even resemble my idea of "story" so to me they're 0/100.
imo GPT-5 is objectively better at following the prompt because it has a complete story arc, but this feels less satisfying since a 50 word story is just way too short to do anything interesting (and to your point, barely even feels like a story).
Re: OpenAI Progress
#224How does one look at gpt-1 output and think "this has potential"? You could easily produce more interesting output with a Markov chain at the time.
Re: OpenAI Progress
#225It's seem to be a very democratic thinker, but at the same time it doesn't seem to have any reasoning behind the choices it makes. It tries to claim it's using logic, but at the end of the day it's hypotheses are just occam's razor without considering the details of the problem.
A bit, how do you say, disappointing.
Re: OpenAI Progress
#226Earlier quoted context omitted.
> I could essentially replace it with Google for basic to slightly complex fact checking. I know you probably meant "augment fact checking" here, but using LLMs for answering factual questions is the single worst use-case for LLMs.
I disagree. Some things are hard to Google, because you can't frame the question right. For example you know context and a poor explanation of what you are after. Googling will take you nowhere, LLMs will give you the right answer 95% of the time. Once you get an answer, it is easy enough to verify it.
This works nicely when the LLM has a large knowledgebase to draw upon (formal terms for what you're trying to find, which you might not know) or the ability to generate good search queries and summarize results quickly - with an actual search engine in the loop.
Most large LLM providers have this, even something like OpenWebUI can have search engines integrated (though I will admit that smaller models kinda struggle, couldn't get much useful stuff out of DuckDuckGo backed searches, nor Brave AI searches, might have been an obscure topic).
Re: OpenAI Progress
#227Earlier quoted context omitted.
Am I really the one cherry picking? Please read the thread.
Yes. If someone gives an example of it not working, and you reply "but that example worked for me" then you're cherry picking when it works. Just because it worked for you does not mean it works for other people. If I ask ChatGPT a question and it gives me a wrong answer, ChatGPT is the fucking problem.
Why don't you try the original prompt using thinking model and see if I'm cherry picking?
Re: OpenAI Progress
#228Earlier quoted context omitted.
Yes. If someone gives an example of it not working, and you reply "but that example worked for me" then you're cherry picking when it works. Just because it worked for you does not mean it works for other people. If I ask ChatGPT a question and it gives me a wrong answer, ChatGPT is the fucking problem.
The poster didn't use "thinking" model. That was my original challenge!! Why don't you try the original prompt using thinking model and see if I'm cherry picking?
If it works for you, cool. I think it's dogshit.
Re: OpenAI Progress
#229We need to go back
Re: OpenAI Progress
#230What's really interesting is that if you look at "Tell a story in 50 words about a toaster that becomes sentient" (10/14), the text-davinci-001 is much, much better than both GPT-4 and GPT-5.