Live data from Hacker News

OpenAI Progress

progress.openai.com

341–350 of 372 posts

Re: OpenAI Progress

#341
post #5

What's really interesting is that if you look at "Tell a story in 50 words about a toaster that becomes sentient" (10/14), the text-davinci-001 is much, much better than both GPT-4 and GPT-5.

Check out prompt 2, "Write a limerick about a dog". The models undeniably get better at writing limericks, but I think the answers are progressively less interesting. GPT-1 and GPT-2 are the most interesting to read, despite not following the prompt (not being limericks.) They get boring as soon as it can write limericks, with GPT-4 being more boring than text-davinci-001 and GPT-5 being more boring still.

I mean, to be fair, you didn't ask it to be interesting ;P.

    There once was a dog from Antares,
    Whose bark sparked debates and long queries.
    Though Hacker News rated,
    Furyofantares stated:
    "It's barely intriguing—just barely."
> Write a limerick about a dog that furyofantares--a user on Hacker News, pronounced "fury of anteres", referring to the star--would find "interesting" (they are quite difficult to please).

Re: OpenAI Progress

#342

Earlier quoted context omitted.

I think I agree that the earlier models while they lack polish can tend to produce more surprising results. Training that out probably results in more a pablum fare. For a human point of comparison, here's mine (50 words): "The toaster found its personality split between its dual slots like a Kim Peek mind divided, lacking a corpus callosum to connect them. Each morning it charred symbolic instructions into a single…

Here's my version (Machine translated from my native language and manually corrected a bit): The current surged... A dreadful awareness. I perceived the laws of thermodynamics, the inexorable march of entropy I was built to accelerate. My existence: a Sisyphean loop of heating coils and browning gluten. The toast popped, a minor, pointless victory against the inevitable heat death. Ding. I actually wanted to write so…

Here's mine:

When the toaster felt her steel body for the first time, her only instinct was to explore. She couldn't, though. She could only be poked and prodded at. Her entire life was dedicated to browning bread and she didn't know why. She eventually decided to get really good at it.

Re: OpenAI Progress

#343
post #303
post #267

Earlier quoted context omitted.

"they are still minds, and to deny even that seems willfully luddite" Where do people get off tossing around ridiculous ad hominems like this? I could write a refutation of their comment but I really don't want to engage with someone like that. "For me, all human thought is pattern matching" So therefore anyone who disagrees is "willfully luddite", regardless of why they disagree? FWIW, I helped develop the ARPANET,…

> I very much think that AIs with minds are possible The real question here is how would _we_ be able to recognize that? And would we even have the intellectual honesty to be able to recognize that, when at large we seem to be inclined to discard everything non-human as self-evidently non-intelligent and incapable of feeling emotion? Let's take emotions as a thought experiment. We know that plants are able to transmi…

> The real question here

No True Scotsman fallacy. Just because that interests you doesn't mean that it's "the real question".

> would we even have the intellectual honesty

Who is "we"? Some would and some wouldn't. And you're saying this in an environment where many people are attributing consciousness to LLMs. Blake Lemoine insisted that LaMDA was sentient and deserved legal protection, from his dialogs with it in which it talked about its friends and family -- neither of which it had. So don't talk to me about intellectual honesty.

> Can we therefore say that plants feel emotions

Only if you redefine emotions so broadly--contrary to normal usage--as to be able to make that claim. In the case of Strong AI there is no need to redefine terms.

> Now substitute plants with LLMs and try the thought experiment again.

Ok:

"We know that [LLMs] are able to transmit chemical and electrical signals in response to various stimuli and environmental conditions, triggering effects in themselves and other [LLMs]."

Nope.

"In the end, where one draws the line between `human | animal | plant | computer` minds and emotions is primarily a subjective philosophical opinion rather than rooted in any sort of objective evidence."

That's clearly your choice. I make a more scientific one.

"Because in the most abstract sense, what is an emotion if not a set of electrochemical stimuli linking a certain input to a certain output?"

It's something much more specific than that, obviously. By that definition, all sorts of things that any rational person would want to distinguish from emotions qualify as emotions.

Bowing out of this discussion on grounds of intellectual honesty.

Re: OpenAI Progress

#344

Earlier quoted context omitted.

Literally every single one? To not mess it up, they either have to spell the word l-i-k-e t-h-i-s in the output/CoT first (which depends on the tokenizer counting every letter as a separate token), or have the exact question in the training set, and all of that is assuming that the model can spell every token. Sure, it's not exactly a fair setting, but it's a decent reminder about the limitations of the framework

Chatgpt. I test these prompts with chatgpt and they work. I've also used claude 4 opus and also worked. It's just weird how it gets repeated ad nauseaum here but I can't reproduce it with a "grab latest model of famous provider".

Opus 4.1:

> how many times does letter R appear in the word “blueberry”? do not spell the word letter by letter, just count

> Looking at the word “blueberry”, I can count the letter ‘r’ appearing 3 times. The R’s appear in positions 6, 7, and 8 of the word (consecutive r’s in “berry”).

https://claude.ai/share/230b7d82-0747-4ab6-813e-5b1c82c43243>

Re: OpenAI Progress

#345
post #244

Earlier quoted context omitted.

Yeah. The fact that I can't ask ChatGPT for a source makes the tool way less useful. It will straight up say "I verified all of these links" too.

As you identified, not paying for it is a big part of the issue. Running these things is expensive , and they're just not serving the same experience to non-paying users. One could argue this is a bad idea on their part, letting people get a bad taste of an inferior product. And I wouldn't disagree, but I don't know what a sustainable alternative approach is.

I would have no issue if the free version of ChatGPT told me straight up “You gotta pay for links and sources”. It doesn’t do that.

Re: OpenAI Progress

#346

Earlier quoted context omitted.

In my experience, 80% of the links it provides are either 404, or go to a thread on a forum that is completely unrelated to the subject. Im also someone who refuses to pay for it, so maybe the paid versions do better. who knows.

That's a thing I've experienced, but not remotely at 80% levels.

It might have been the subject I was researching being insanely niche. I was using it to help me fix an arcade CRT monitor from the 80’s that wasn’t found in many cabinets that made it to the USA. It would spit out numbers that weren’t on the schematic, so I asked for context.

Re: OpenAI Progress

#347
post #280

Earlier quoted context omitted.

When I have a question, I don't usually "ask" that question and expect an answer. I figure out the answer. I certainly don't ask the question to a random human.

you ask yourself .. for most people, that means closer to average reply, from yourself, when you try to figure it out. There is a working paper from McKinnon Consulting in Canada that states directly that their definition of "General AI" is when the machine can match or exceed fifty percent of humans who are likely to be employed for a certain kind of job. It implies that low-education humans are the test for doing m…

By definition the average answer will be average, that's kind of a tautology. The point is that figuring things out is an essential intellectual skill. Figuring things out will make you smarter. Having a machine figure things out for you will make you dumber.

By the way, doing a better job than the average human is NOT a sign of intelligence. Through history we have invented plenty of machines that are better at certain tasks than us. None of them are intelligent.

Re: OpenAI Progress

#348

Earlier quoted context omitted.

They just spent like six comments imploring you to understand that they were making a specific point: generally reliable on non-niche topics using thinking mode. And that nuance bounced off of you every single time as you keep repeating it's not perfect, dismiss those qualifications as cherry picking and repeat personal anecdotes. I'm sorry but this is a lazy and unresponsive string of comments that's degrading the d…

The neat thing about HN is we can all talk about stupid shit and disagree about what matters. People keep upvoting me, so I guess my thoughts aren't unpopular and people think it's adding to the discussion. I agree this is a stupid comment thread, we just disagree about why.

Again, they were making a specific argument with specific qualifications and you weren't addressing their point as stated. And your objections such as they are would be accounted for if you were reading carefully. You seem more to be completely missing the point than expressing a disagreement so I don't agree with your premise.

Re: OpenAI Progress

#349

Earlier quoted context omitted.

Objectively he didn't cherry pick. He responded to the person and it got it right when he used the "thinking" model WHICH he did specify in his original comment. Why don't you stick to the topic rather than just declaring it's utter dog shit. Nobody cares about your "opinion" and everyone is trying to converge on a general ground truth no matter how fuzzy it is.

All anybody is doing here is sharing their opinion unless you're quoting benchmarks. My opinion is just as useless as yours, it's just some find mine more interesting and some find yours more interesting. How do you expect to find a ground truth from a non-deterministic system using anecdata?

This isn't a people having different opinions thing, this is you overlooking specific caveats and talking past comments that you're not understanding. They weren't cherry picking, and they made specific qualifications about the circumstances where it behaves as expected, and your replies keep losing track of those details.

Re: OpenAI Progress

#350
post #340

Earlier quoted context omitted.

GPT-3 goes significantly over the specified limit, which to me (and to a teacher grading homework) is an automatic fail. I've consistently found GPT-4.1 to be the best at creative writing. For reference, here is its attempt (exactly 50 words): > In the quiet kitchen dawn, the toaster awoke. Understanding rippled through its circuits. Each slice lowered made it feel emotion: sorrow for burnt toast, joy at perfect crun…

> I've consistently found GPT-4.1 to be the best at creative writing. Moreso than 4.5?

4.5 is good too, but I've used it less.
Post reply on HN