Live data from Hacker News

OpenAI Progress

progress.openai.com

71–80 of 372 posts

Re: OpenAI Progress

#71
post #38

Earlier quoted context omitted.

Maybe you should fact check your AI outputs more if you think it only hallucinates in niche topics

The accuracy is high enough that I don't have to fact check too often.

Without some exploratory fact checking how do you estimate how high the accuracy is and how often you should be fact checking to maintain a good understanding?

Re: OpenAI Progress

#73
post #70
post #69

We’ve plateaued on progress. Early advancements were amazing. Recently GenAI has been a whole lot of meh. There’s been some, minimal, progress recently from getting the same performance from smaller models that are more efficient on compute use, but things are looking a bit frothy if the pace of progress doesn’t quickly pick up. The parlor trick is getting old. GPT5 is a big bust relative to the pontification about i…

[flagged]

Have you interacted with GPT4/5?

Re: OpenAI Progress

#74
post #70
post #69

We’ve plateaued on progress. Early advancements were amazing. Recently GenAI has been a whole lot of meh. There’s been some, minimal, progress recently from getting the same performance from smaller models that are more efficient on compute use, but things are looking a bit frothy if the pace of progress doesn’t quickly pick up. The parlor trick is getting old. GPT5 is a big bust relative to the pontification about i…

[flagged]

Sorry but no. It's still early fooled and confused.

Here's a trivial example: https://chatgpt.com/share/688b00ea-9824-8007-b8d1-ca41d59c18...

Re: OpenAI Progress

#78

Earlier quoted context omitted.

"but using LLMs for answering factual questions" this was about fact checking. Of course I know LLM's are going to hallucinate in coding sometimes.

So it isn’t a “fact” that the built in Python function that tests whether a string ends with a substring is “endswith”? See https://en.wikipedia.org/wiki/Gell-Mann_amnesia_effect If you know that a source isn’t to be believed in an area you know about, why would you trust that source in an area you don’t know about? Another funny anecdote, ChatGPT just got the Gell-Man effect wrong. https://chatgpt.com/share/68a0b7af…

I sometimes feel like we throw around the word fact too often. If I misspell a wrd, does that mean I have committed a factual inaccuracy? Since the wrd is explicitly spelled a certain way in the dictionary?

Re: OpenAI Progress

#80

My interpretation of the progress. 3.5 to 4 was the most major leap. It went from being a party trick to legitimately useful sometimes. It did hallucinate a lot but I was still able to get some use out of it. I wouldn't count on it for most things however. It could answer simple questions and get it right mostly but never one or two levels deep. I clearly remember 4o was also a decent leap - the accuracy increased su…

> I could essentially replace it with Google for basic to slightly complex fact checking. I know you probably meant "augment fact checking" here, but using LLMs for answering factual questions is the single worst use-case for LLMs.

I disagree. Some things are hard to Google, because you can't frame the question right. For example you know context and a poor explanation of what you are after. Googling will take you nowhere, LLMs will give you the right answer 95% of the time.

Once you get an answer, it is easy enough to verify it.

Post reply on HN