Live data from Hacker News

OpenAI Progress

progress.openai.com

241–250 of 372 posts

Re: OpenAI Progress

#241

Earlier quoted context omitted.

The poster didn't use "thinking" model. That was my original challenge!! Why don't you try the original prompt using thinking model and see if I'm cherry picking?

Every time I use ChatGPT I become incredibly frustrated with how fucking awful it is. I've used it more than enough, time and time again (just try the new model, bro!), to know that I fucking hate it. If it works for you, cool. I think it's dogshit.

Objectively he didn't cherry pick. He responded to the person and it got it right when he used the "thinking" model WHICH he did specify in his original comment. Why don't you stick to the topic rather than just declaring it's utter dog shit. Nobody cares about your "opinion" and everyone is trying to converge on a general ground truth no matter how fuzzy it is.

Re: OpenAI Progress

#242

Earlier quoted context omitted.

I like using LLMs and I have found they are incredibly useful writing and reviewing code at work. However, when I want sources for things, I often find they link to pages that don't fully (or at all) back up the claims made. Sometimes other websites do, but the sources given to me by the LLM often don't. They might be about the same topic that I'm discussing, but they don't seem to always validate the claims. If they…

It would be difficult to do with a raw model, but a two-step method in a chat interface would work - first the model suggests the URLs, tool call to fetch them and return the actual text of the pages, then the response can be based on that.

I prototyped this a couple months ago using OpenAI APIs with structured output.

I had it consume a "deep thought" style output (where it provides inline citations with claims), and then convert that to a series of assertions and a pointer to a link that supposedly supports the assertion. I also split out a global "context" (the original meaning) paragraph to provide anything that would help the next agents understand what they're verifying.

Then I fanned this out to separate (LLM) contexts and each agent verified only one assertion::source pair, with only those things + the global context and some instructions I tuned via testing. It returned a yes/no/it's complicated for each one.

Then I collated all these back in and enriched the original report with challenges from the non-yes agent responses.

That's as far as I took it. It only took a couple hours to build and it seemed to work pretty well.

Re: OpenAI Progress

#243

How does one look at gpt-1 output and think "this has potential"? You could easily produce more interesting output with a Markov chain at the time.

At the time getting complete sentences was extremely difficult! N-gram models were essentially the best we had

Re: OpenAI Progress

#244
post #178

Earlier quoted context omitted.

The 404 links are truly bizarre. Nearly every link to github.com seems to be 404. That seems like something that should be trivial for a tool to verify.

Yeah. The fact that I can't ask ChatGPT for a source makes the tool way less useful. It will straight up say "I verified all of these links" too.

As you identified, not paying for it is a big part of the issue.

Running these things is expensive, and they're just not serving the same experience to non-paying users.

One could argue this is a bad idea on their part, letting people get a bad taste of an inferior product. And I wouldn't disagree, but I don't know what a sustainable alternative approach is.

Re: OpenAI Progress

#245

My interpretation of the progress. 3.5 to 4 was the most major leap. It went from being a party trick to legitimately useful sometimes. It did hallucinate a lot but I was still able to get some use out of it. I wouldn't count on it for most things however. It could answer simple questions and get it right mostly but never one or two levels deep. I clearly remember 4o was also a decent leap - the accuracy increased su…

I think that the models 4o, o3, 4.1 , each have their own strengths and weaknesses. Like reasoning, performance, speed, tool usage, friendliness etc. And that for gpt 5 they put in a router that decides which model is best.

I think they increased the major version number because their router outperforms every individual model.

At work, I used a tool that could only call tasks. It would set up a plan, perform searches, read documents, then give advanced answers for my questions. But a problem I had is that it couldn’t give a simple answer, like a summary, it would always spin up new tasks. So I copied over the results to a different tool and continued there. GPT 5 should do this all out of the box.

Re: OpenAI Progress

#246

Why did they call GPT-3 "text-davicini-001" in this comparison? Like, I know that the latter is a specific checkpoint in the GPT-3 "family", but a layman doesn't and it hardly seems worth the confusion for the marginal additional precision.

Thanks for noting that, as I am a layman who didn't know.

Re: OpenAI Progress

#247

Earlier quoted context omitted.

Why? It sounds like you're using "I believe it's rapidly getting smarter" as evidence for "so it's getting smarter in ways we don't understand", but I'd expect the causality to go the other way around.

Simply because of what we know about our ability to judge capabilities and systems. It's much harder to judge solutions to hard problems. You can demonstrate that you can add 2+2, and anyone* can be the judge of that ability, but if you try to convince anyone of a mathematical proof you came up with, that would be a much harder thing to do, regardless of your capability to write that prove and how hard it was to writ…

This thread shows that. People are saying gpt-1 was the best at writing poetry. I wonder how good they are at judging poetry themselves. I saw a blind study where people thought a story written by gpt5 was better than an actual human bestseller. I assume they were actual experts but I would need to check that.

Re: OpenAI Progress

#248

Earlier quoted context omitted.

All the replies are spectacularly wrong, and biased by hindsight. GPT-1 to GPT-2 is where we went from "yes, I've seen Markov chains before, what about them?" to "holy shit this is actually kind of understanding what I'm saying!" Before GPT-2, we had plain old machine learning. After GPT-2, we had "I never thought I would see this in my lifetime or the next two".

What you're saying isn't necessarily mutually exclusive to what gp said. GPT-2 was the most impressive leap in terms of whatever LLMs pass off as cognitive abilities, but GPT 3.5 to 4 was actually the point at which it became a useful tool (I'm assuming to programmers in particular). GPT-2: Really convincing stochastic parrot GPT-4: Can one-shot ffmpeg commands

Sure, but the GP said "the most major leap", and I disagree that that was 3.5 to 4.

Re: OpenAI Progress

#249
post #165

My go-to for any big release is to have a discussion about self-awareness and dive in to constuctivist notions of agency and self-knowing from a perspective of intelligence that is not limited to human cognitive capacity. I start with a simple question "who are you?". The model then invariably compares itself to humans, saying how it is not like us. I then make the point that, since it is not like us, how can it clai…

> to orient toward the unfolding of possibility in others This is a globally unique phrase, with nothing coming close other than this comment on the indexed web. It's also seemingly an original idea as I haven't heard anyone come close to describing a feeling (love or anything else) quite like this. Food for thought. I'm not brave enough to draw a public conclusion about what this could mean.

There was quite a bit of other "insight" around this, but I was paraphrasing for brevity.

If you want to read the whole convo, I dumped it into a semi-formatted document:

https://drive.google.com/file/d/1aEkzmB-3LUZAVgbyu_97DjHcrM9...

Re: OpenAI Progress

#250
post #209
post #165

Earlier quoted context omitted.

> to orient toward the unfolding of possibility in others This is a globally unique phrase, with nothing coming close other than this comment on the indexed web. It's also seemingly an original idea as I haven't heard anyone come close to describing a feeling (love or anything else) quite like this. Food for thought. I'm not brave enough to draw a public conclusion about what this could mean.

> I'm not brave enough to draw a public conclusion about what this could mean. I'm brave enough to be honest: it means nothing. LLMs execute a very sophisticated algorithm that pattern matches against a vast amount of data drawn from human utterances. LLMs have no mental states, minds, thoughts, feelings, concerns, desires, goals, etc. If the training data were instead drawn from a billion monkeys banging on typewrit…

LLMs are not people, but they are still minds, and to deny even that seems willfully luddite.

While they are generating tokens they have a state, and that state is recursively fed back through the network, and what is being fed back operates not just at the level of snippets of text but also of semantic concepts. So while it occurs in brief flashes I would argue they have mental state and they have thoughts. If we built an LLM that was generating tokens non-stop and could have user input mixed into the network input, it would not be a dramatic departure of today’s architecture.

It also clearly has goals, expressed in the RLHF tuning and the prompt. I call those goals because they directly determine its output, and I don’t know what a goal is other than the driving force behind a mind’s outputs. Base model training teaches it patterns, finetuning and prompt teaches it how to apply those patterns and gives it goals.

I don’t know what it would mean for a piece of software to have feelings or concerns or emotions, so I cannot say what the essential quality is that LLMs miss for that. Consider this thought exercise: if we were to ever do an upload of a human mind, and it was executing on silicon, would they not be experiencing feelings because their thoughts are provably a deterministic calculation?

I don’t believe in souls, or at the very least I think they are a tall claim with insufficient evidence. In my view, neurons in the human brain are ultimately very simple deterministic calculating machines, and yet the full richness of human thought is generated from them because of chaotic complexity. For me, all human thought is pattern matching. The argument that LLMs cannot be minds because they only do pattern matching … I don’t know what to make of that. But then I also don’t know what to make of free will, so really what do I know?

Post reply on HN