Live data from Hacker News

OpenAI Progress

progress.openai.com

301–310 of 372 posts

Re: OpenAI Progress

#301

I talked to GPT yesterday about a fairly simple problem I'm having with my fridge, and it gave me the most ridiculous / wrong answers. It new the spec, but was convinced the components were different (single compressor, for example, whereas mine has 2 separate systems) and was hypothesizing the problem as being something that doesn't exist on this model of refrigerator. It seems like in a lot of domain spaces it just…

Clearly a skill issue where you're expecting it to know all of the specifications of a particular refrigerator model.

You didn't provide it with the correct context.

Re: OpenAI Progress

#303
post #267
post #209

Earlier quoted context omitted.

> I'm not brave enough to draw a public conclusion about what this could mean. I'm brave enough to be honest: it means nothing. LLMs execute a very sophisticated algorithm that pattern matches against a vast amount of data drawn from human utterances. LLMs have no mental states, minds, thoughts, feelings, concerns, desires, goals, etc. If the training data were instead drawn from a billion monkeys banging on typewrit…

"they are still minds, and to deny even that seems willfully luddite" Where do people get off tossing around ridiculous ad hominems like this? I could write a refutation of their comment but I really don't want to engage with someone like that. "For me, all human thought is pattern matching" So therefore anyone who disagrees is "willfully luddite", regardless of why they disagree? FWIW, I helped develop the ARPANET,…

> I very much think that AIs with minds are possible

The real question here is how would _we_ be able to recognize that? And would we even have the intellectual honesty to be able to recognize that, when at large we seem to be inclined to discard everything non-human as self-evidently non-intelligent and incapable of feeling emotion?

Let's take emotions as a thought experiment. We know that plants are able to transmit chemical and electrical signals in response to various stimuli and environmental conditions, triggering effects in themselves and other plants. Can we therefore say that plants feel emotions, just in a way that is unique to them and not necessarily identical to a human embodiment?

The answer to that question depends on one's worldview, rather than any objective definition of the concept of emotion. One could say plants cannot feel emotions because emotions are a human (or at least animal) construct; or one could say that plants can feel emotions, just not exactly identical to human emotions.

Now substitute plants with LLMs and try the thought experiment again.

In the end, where one draws the line between `human | animal | plant | computer` minds and emotions is primarily a subjective philosophical opinion rather than rooted in any sort of objective evidence. Not too long ago, Descartes was arguing that animals do not possess a mind and cannot feel emotions, they are merely mimicry machines.[1] More recently, doctors were saying similar things about babies and adults, leading to horrifying medical malpractice.[2][3]

Because in the most abstract sense, what is an emotion if not a set of electrochemical stimuli linking a certain input to a certain output? And how can we tell what does and what does not possess a mind if we are so undeniably bad at recognize those attributes even within our own species?

[1] https://en.wikipedia.org/wiki/Animal_machine

[2] https://en.wikipedia.org/wiki/Pain_in_babies

[3] https://pmc.ncbi.nlm.nih.gov/articles/PMC4843483/

Re: OpenAI Progress

#304
To the prompt “write a limerick about a dog,” GPT-2 wrote:

“Dog, reached for me

Next thought I tried to chew

Then I bit and it turned Sunday

Where are the squirrels down there, doing their bits

But all they want is human skin to lick”

While obviously not a limerick, I thought this was actually a decent poem, with some turns of phrase that conveyed a kind of curious and unusual feeling.

This reminded me how back then I got a lot of joy and surprise out of the mercurial genius of the early GPT models.

Re: OpenAI Progress

#305

Earlier quoted context omitted.

With the exception of h-bomb/fusion and ENIAC/AI, I think all of those examples reflect a change in priority and investment more than anything. There was a trajectory of high investment / rapid progress, then market and social and political drivers changed and space travel / supersonic flight just became less important.

That's the conceit for the tv show For All Mankind - what if the space race didn't end? But I don't buy it, IMO the space race ended for material reasons rather than political. Space is just too hard and there is not much of value "out there". But regardless, it's a futile excuse, markets and politics should be part of any serious prognostication.

I think it was a combination of the two. The Apollo program was never popular. It took up an enormous portion of the federal budget, which the Republicans argued was fiscally unwise and the Democrats argued that the money should have been used to fund domestic social programs. In 1962, the New York Times noted that the projected Apollo program budget could have instead been used to create over 100 universities of a similar size to Harvard, build millions of homes, replace hundreds of worn-out schools, build hundreds of hospitals, and fund disease research. The Apollo program's popularity peaked at 53% just after the moon landing, and by April 1970 it was back down to 40%. It wasn't until the mid-80s that the majority of Americans thought that the Apollo program was worth it. Because of all this, I think it's inevitable that the Apollo program would wind down once it had achieved its goal of national prestige.

Re: OpenAI Progress

#306
post #101

Earlier quoted context omitted.

It does citations (Grok and Claude etc do too) but I've found when I read the source on some stuff (GitHub discussions and so on) it sometimes actually has nothing to do with what the LLM said. I've actually wasted a lot of time trying to find the actual spot in a threaded conversation where the example was supposedly stated.

Same experience with Google search AI. The links frequently don’t support the assertions, they’ll just say something that might show up in a google search for the assertion. For example if I’m asking about whether a feature exists in some library, the AI says yes it does and links to a forum where someone is asking the same question I did, but no one answered (this has happened multiple times).

It is funny, Perplexity seems to work much better in this use case for me. When I want some sort of "conclusive answer", I use Gemini pro (just what I have available). It is good with coding and formulating thoughts, rewriting text, so on.

But when I want to actually search for content on the web for, say, product research or opinions on a topic, Perplexity is so much better than either Gemini or google search AI. It lists reference links for each block of assertions that are EASILY clicked on (unlike Gemini or search AI, where the references are just harder to click on for some reason, not the least of which is that they OPEN IN THE SAME TAB where Perplexity always opens on a new tab). This is often a reddit specific search as I want people's opinions on something.

Perplexity's UI for search specifically is the main thing it does just so much better than google's offering is the one thing going for it. I think there is some irony there.

Full disclosure, I don't use Anthropic or OpenAI, so this may not be the case for those products.

Re: OpenAI Progress

#307

One thing that appears to have been lost between GPT-4 and GPT-5 is that it no longer reminds the user that it's an AI and not a human, let alone a human expert. Maybe those genuinely annoyed people, but it seems like they were potentially useful measure to prevent users from being overly credulous GPT-5 also goes out of its way to suggest new prompts. This seems potentially useful, although potentially dangerous if…

I found this "advancement" creepy. It seems like they deliberately made GPT-5 more laid back, conversational and human-like. I don't think LLMs should mimic humans and I think this is a dangerous development.

Re: OpenAI Progress

#308

Earlier quoted context omitted.

With the exception of h-bomb/fusion and ENIAC/AI, I think all of those examples reflect a change in priority and investment more than anything. There was a trajectory of high investment / rapid progress, then market and social and political drivers changed and space travel / supersonic flight just became less important.

That's the conceit for the tv show For All Mankind - what if the space race didn't end? But I don't buy it, IMO the space race ended for material reasons rather than political. Space is just too hard and there is not much of value "out there". But regardless, it's a futile excuse, markets and politics should be part of any serious prognostication.

I think the space race ended because we got all the benefit available, which wasn’t really in space anyway, it was the ancillary technical developments like computers, navigation, simulation, incredible tolerances in machining, material science, etc.

We’re seeing a resurgence in space because there is actually value in space itself, in a way that scales beyond just telecom satellites. Suddenly there are good reasons to want to launch 500 times a year.

There was just a 50-year discontinuity between the two phases.

Re: OpenAI Progress

#309

Earlier quoted context omitted.

I'd love to know more about how OpenAI (or Alec Radford et al.) even decided GPT-1 was worth investing more into. At a glance the output is barely distinguishable from Markov chains. If in 2018 you told me that scaling the algorithm up 100-1000x would lead to computers talking to people/coding/reasoning/beating the IMO I'd tell you to take your meds.

GPT-1 wasn't used as a zero-shot text generator; that wasn't why it was impressive. The way GPT-1 was used was as a base model to be fine-tuned on downstream tasks. It was the first case of a (fine-tuned) base Transformer model just trivially blowing everything else out of the water. Before this, people were coming up with bespoke systems for different tasks (a simple example is that for SQuAD a passage-question-answ…

> Transformer model just trivially blowing everything else out of the water

no, this is the winners rewriting history. Transformer style encoders are now applied to lots and lots of disciplines but they do not "trivially" do anything. The hype re-telling is obscuring the facts of history. Specifically in human language text translation, "Attention is All You Need" Transformers did "blow others out of the water" yes, for that application.

Re: OpenAI Progress

#310
post #298

My interpretation of the progress. 3.5 to 4 was the most major leap. It went from being a party trick to legitimately useful sometimes. It did hallucinate a lot but I was still able to get some use out of it. I wouldn't count on it for most things however. It could answer simple questions and get it right mostly but never one or two levels deep. I clearly remember 4o was also a decent leap - the accuracy increased su…

Everyone talks about 4o so positively but I’ve never consistently relied on it in a production environment. I’ve found it to be inconsistent in json generation and often it’s writing and following of the system prompt was very poor. In fact it was a huge part of what got me looking closer at anthropics models. I’m really curious what people did with it because while it’s cool it didn’t compare well in my real world u…

I preferred o3 for coding and analysis tasks, but appreciated 4o as a “companion model” for brainstorming creative ideas while taking long walks. Wasn’t crazy about the sycophancy but it was a decent conceptual field for playing with ideas. Steve Jobs once described the PC as a “bicycle for the mind.” This is how I feel when using models like 4o for meandering reflection and speculation.
Post reply on HN