Live data from Hacker News

OpenAI Progress

progress.openai.com

221–230 of 372 posts

Re: OpenAI Progress

#221

My interpretation of the progress. 3.5 to 4 was the most major leap. It went from being a party trick to legitimately useful sometimes. It did hallucinate a lot but I was still able to get some use out of it. I wouldn't count on it for most things however. It could answer simple questions and get it right mostly but never one or two levels deep. I clearly remember 4o was also a decent leap - the accuracy increased su…

> I could essentially replace it with Google for basic to slightly complex fact checking. I know you probably meant "augment fact checking" here, but using LLMs for answering factual questions is the single worst use-case for LLMs.

[dead]

Re: OpenAI Progress

#222
post #195

Earlier quoted context omitted.

I'd love to know more about how OpenAI (or Alec Radford et al.) even decided GPT-1 was worth investing more into. At a glance the output is barely distinguishable from Markov chains. If in 2018 you told me that scaling the algorithm up 100-1000x would lead to computers talking to people/coding/reasoning/beating the IMO I'd tell you to take your meds.

I don't have a source for this (there's probably no sources from anything back then) but anecdotally, someone at an AI/ML talk said they just added more data and quality went up. Doubling the data doubled the quality. With other breakthroughs, people saw diminishing gains. It's sort of why Sam back then tweeted that he expected the amount of intelligence to double every N years. I have the feeling they kept on this u…

The input size to output quality mapping is not linear. This is why we are in the regime of "build nuclear power plants to power datacenters". Fixed size improvements in loss require exponential increases in parameters/compute/data.

Re: OpenAI Progress

#223
post #9

Earlier quoted context omitted.

I find GPT-5's story significantly better than text-davinci-001

I really wonder which one of us is the minority. Because I find text-davinci-001 answer is the only one that reads like a story. All the others don't even resemble my idea of "story" so to me they're 0/100.

text-davinci-001 feels more like a story, but it is also clearly incomplete, in that it is cut-off before the story arc is finished.

imo GPT-5 is objectively better at following the prompt because it has a complete story arc, but this feels less satisfying since a 50 word story is just way too short to do anything interesting (and to your point, barely even feels like a story).

Re: OpenAI Progress

#224

How does one look at gpt-1 output and think "this has potential"? You could easily produce more interesting output with a Markov chain at the time.

This was an era where language modeling was only considered as a pretraining step. You were then supposed to fine tune it further to get a classifier or similar type of specialized model.

Re: OpenAI Progress

#225
I talked to GPT yesterday about a fairly simple problem I'm having with my fridge, and it gave me the most ridiculous / wrong answers. It new the spec, but was convinced the components were different (single compressor, for example, whereas mine has 2 separate systems) and was hypothesizing the problem as being something that doesn't exist on this model of refrigerator. It seems like in a lot of domain spaces it just takes the majority, even if the majority is wrong.

It's seem to be a very democratic thinker, but at the same time it doesn't seem to have any reasoning behind the choices it makes. It tries to claim it's using logic, but at the end of the day it's hypotheses are just occam's razor without considering the details of the problem.

A bit, how do you say, disappointing.

Re: OpenAI Progress

#226

Earlier quoted context omitted.

> I could essentially replace it with Google for basic to slightly complex fact checking. I know you probably meant "augment fact checking" here, but using LLMs for answering factual questions is the single worst use-case for LLMs.

I disagree. Some things are hard to Google, because you can't frame the question right. For example you know context and a poor explanation of what you are after. Googling will take you nowhere, LLMs will give you the right answer 95% of the time. Once you get an answer, it is easy enough to verify it.

> For example you know context and a poor explanation of what you are after. Googling will take you nowhere, LLMs will give you the right answer 95% of the time.

This works nicely when the LLM has a large knowledgebase to draw upon (formal terms for what you're trying to find, which you might not know) or the ability to generate good search queries and summarize results quickly - with an actual search engine in the loop.

Most large LLM providers have this, even something like OpenWebUI can have search engines integrated (though I will admit that smaller models kinda struggle, couldn't get much useful stuff out of DuckDuckGo backed searches, nor Brave AI searches, might have been an obscure topic).

Re: OpenAI Progress

#227

Earlier quoted context omitted.

Am I really the one cherry picking? Please read the thread.

Yes. If someone gives an example of it not working, and you reply "but that example worked for me" then you're cherry picking when it works. Just because it worked for you does not mean it works for other people. If I ask ChatGPT a question and it gives me a wrong answer, ChatGPT is the fucking problem.

The poster didn't use "thinking" model. That was my original challenge!!

Why don't you try the original prompt using thinking model and see if I'm cherry picking?

Re: OpenAI Progress

#228

Earlier quoted context omitted.

Yes. If someone gives an example of it not working, and you reply "but that example worked for me" then you're cherry picking when it works. Just because it worked for you does not mean it works for other people. If I ask ChatGPT a question and it gives me a wrong answer, ChatGPT is the fucking problem.

The poster didn't use "thinking" model. That was my original challenge!! Why don't you try the original prompt using thinking model and see if I'm cherry picking?

Every time I use ChatGPT I become incredibly frustrated with how fucking awful it is. I've used it more than enough, time and time again (just try the new model, bro!), to know that I fucking hate it.

If it works for you, cool. I think it's dogshit.

Re: OpenAI Progress

#230
post #5

What's really interesting is that if you look at "Tell a story in 50 words about a toaster that becomes sentient" (10/14), the text-davinci-001 is much, much better than both GPT-4 and GPT-5.

Less lobotomized and boxed in by RLHF rules. That’s why a 7b base model will “outprose” an 80b instruct model.
Post reply on HN