Live data from Hacker News

OpenAI Progress

progress.openai.com

211–220 of 372 posts

Re: OpenAI Progress

#211

Earlier quoted context omitted.

The point you're missing is it's not always right. Cherry-picking examples doesn't really bolster your point. Obviously it works for you (or at least you think it does), but I can confidently say it's fucking god-awful for me.

Am I really the one cherry picking? Please read the thread.

Yes. If someone gives an example of it not working, and you reply "but that example worked for me" then you're cherry picking when it works. Just because it worked for you does not mean it works for other people.

If I ask ChatGPT a question and it gives me a wrong answer, ChatGPT is the fucking problem.

Re: OpenAI Progress

#213

My interpretation of the progress. 3.5 to 4 was the most major leap. It went from being a party trick to legitimately useful sometimes. It did hallucinate a lot but I was still able to get some use out of it. I wouldn't count on it for most things however. It could answer simple questions and get it right mostly but never one or two levels deep. I clearly remember 4o was also a decent leap - the accuracy increased su…

when you adjust the improvements with the amount of debt incurred and the amount of profit made... ALL the versions are incremental.

This isnt sustainable.

Re: OpenAI Progress

#214

Earlier quoted context omitted.

Next logical step is to connect ( or build from ground up ) large AI models to high performance passive slaves ( MCP or internally ) , which gives precise facts, language syntax validation, maths equations runners, may be prolog kind of system, which give it much more power if we train it precisely to use each tool. ( using AI to better articulate my thoughts ) Your comment points toward a fascinating and important d…

Don't do these ai thoughts thing No one reads it and it seems fake

Seems fake because it is.

Re: OpenAI Progress

#215

Earlier quoted context omitted.

Modern ChatGPT will (typically on its own; always if you instruct it to) provide inline links to back up its answers. You can click on those if it seems dubious or if it's important, or trust it if it seems reasonably true and/or doesn't matter much. The fact that it provides those relevant links is what allows it to replace Google for a lot of purposes.

In my experience, 80% of the links it provides are either 404, or go to a thread on a forum that is completely unrelated to the subject. Im also someone who refuses to pay for it, so maybe the paid versions do better. who knows.

That's a thing I've experienced, but not remotely at 80% levels.

Re: OpenAI Progress

#216

My interpretation of the progress. 3.5 to 4 was the most major leap. It went from being a party trick to legitimately useful sometimes. It did hallucinate a lot but I was still able to get some use out of it. I wouldn't count on it for most things however. It could answer simple questions and get it right mostly but never one or two levels deep. I clearly remember 4o was also a decent leap - the accuracy increased su…

All the replies are spectacularly wrong, and biased by hindsight. GPT-1 to GPT-2 is where we went from "yes, I've seen Markov chains before, what about them?" to "holy shit this is actually kind of understanding what I'm saying!" Before GPT-2, we had plain old machine learning. After GPT-2, we had "I never thought I would see this in my lifetime or the next two".

That’s true, but not contradictory.

Re: OpenAI Progress

#217

Earlier quoted context omitted.

All the replies are spectacularly wrong, and biased by hindsight. GPT-1 to GPT-2 is where we went from "yes, I've seen Markov chains before, what about them?" to "holy shit this is actually kind of understanding what I'm saying!" Before GPT-2, we had plain old machine learning. After GPT-2, we had "I never thought I would see this in my lifetime or the next two".

I'd love to know more about how OpenAI (or Alec Radford et al.) even decided GPT-1 was worth investing more into. At a glance the output is barely distinguishable from Markov chains. If in 2018 you told me that scaling the algorithm up 100-1000x would lead to computers talking to people/coding/reasoning/beating the IMO I'd tell you to take your meds.

GPT-1 wasn't used as a zero-shot text generator; that wasn't why it was impressive. The way GPT-1 was used was as a base model to be fine-tuned on downstream tasks. It was the first case of a (fine-tuned) base Transformer model just trivially blowing everything else out of the water. Before this, people were coming up with bespoke systems for different tasks (a simple example is that for SQuAD a passage-question-answering task, people would have an LSTM to read the passage and another LSTM to read the question, because of course those are different sub-tasks with different requirements and should have different sub-models). One GPT-1 came out, you just dumped all the text into the context, YOLO fine-tuned it, and trivially got state on the art on the task. On EVERY NLP task.

Overnight, GPT-1 single-handedly upset the whole field. It was somewhat overshadowed by BERT and T5 models that came out very shortly after, which tended to perform even better on the pretrain-and-finetune format. Nevertheless, the success of GPT-1 definitely already warrants scaling up the approach.

A better question is how OpenAI decided to scale GPT-2 to GPT-3. It was an awkward in-between model. It generated better text for sure, but the zero-shot performance reported in the paper, while neat, was not great at all. On the flip side, its fine-tuned task performance paled compared to much smaller encoder-only Transformers. (The answer is: scaling laws allowed for predictable increases in performance.)

Re: OpenAI Progress

#218
In a few years we've gone from gibberish (less poetic maybe, less polished and surprising, but none the less gibberish) - to legit conversational, and in my own opinion, well rounded answers. This is a great example of hard-core engineering - no matter what your opinion of the organisation and saltman is, they have built something amazing. I do hope they continue with their improvements, it's honestly the most useful tool in my arsenal since stackoverflow.

Re: OpenAI Progress

#220

Earlier quoted context omitted.

I totally get that you meant this in a nuanced way, but at face value it sort of reads like... Joe Rogan has high enough accuracy that I don't have to fact check too often. Newsmax has high enough accuracy that I don't have to fact check too often, etc. If you accept the output as accurate, why would fact checking even cross your mind?

Do you question everything your dad says?

If it's about classic American cars, no. Anything else, usually.
Post reply on HN