Earlier quoted context omitted.
Maybe you should fact check your AI outputs more if you think it only hallucinates in niche topics
The accuracy is high enough that I don't have to fact check too often.
OpenAI Progress
71–80 of 372 posts
Re: OpenAI Progress
#72Re: OpenAI Progress
#73We’ve plateaued on progress. Early advancements were amazing. Recently GenAI has been a whole lot of meh. There’s been some, minimal, progress recently from getting the same performance from smaller models that are more efficient on compute use, but things are looking a bit frothy if the pace of progress doesn’t quickly pick up. The parlor trick is getting old. GPT5 is a big bust relative to the pontification about i…
[flagged]
Re: OpenAI Progress
#74We’ve plateaued on progress. Early advancements were amazing. Recently GenAI has been a whole lot of meh. There’s been some, minimal, progress recently from getting the same performance from smaller models that are more efficient on compute use, but things are looking a bit frothy if the pace of progress doesn’t quickly pick up. The parlor trick is getting old. GPT5 is a big bust relative to the pontification about i…
[flagged]
Here's a trivial example: https://chatgpt.com/share/688b00ea-9824-8007-b8d1-ca41d59c18...
Re: OpenAI Progress
#75Re: OpenAI Progress
#76Re: OpenAI Progress
#77Re: OpenAI Progress
#78Earlier quoted context omitted.
"but using LLMs for answering factual questions" this was about fact checking. Of course I know LLM's are going to hallucinate in coding sometimes.
So it isn’t a “fact” that the built in Python function that tests whether a string ends with a substring is “endswith”? See https://en.wikipedia.org/wiki/Gell-Mann_amnesia_effect If you know that a source isn’t to be believed in an area you know about, why would you trust that source in an area you don’t know about? Another funny anecdote, ChatGPT just got the Gell-Man effect wrong. https://chatgpt.com/share/68a0b7af…
Re: OpenAI Progress
#79Re: OpenAI Progress
#80My interpretation of the progress. 3.5 to 4 was the most major leap. It went from being a party trick to legitimately useful sometimes. It did hallucinate a lot but I was still able to get some use out of it. I wouldn't count on it for most things however. It could answer simple questions and get it right mostly but never one or two levels deep. I clearly remember 4o was also a decent leap - the accuracy increased su…
> I could essentially replace it with Google for basic to slightly complex fact checking. I know you probably meant "augment fact checking" here, but using LLMs for answering factual questions is the single worst use-case for LLMs.
Once you get an answer, it is easy enough to verify it.