Live data from Hacker News

GPT-4.5

openai.com

621–630 of 1001 posts

Re: GPT-4.5

#621
post #81
post #24

Considering both this blog post and the livestream demos, I am underwhelmed. Having just finished the stream, I had a real "was that all" moment, which on one hand shows how spoiled I've gotten by new models impressing me, but on another feels like OpenAI really struggles to stay ahead of their competitors. What has been shown feels like it could be achieved using a custom system prompt on older versions of OpenAIs m…

> How could they justify that asking price? They're still selling $1 for <$1. Like personal food delivery before it, consumers will eventually need to wake up to this fact - these things will get expensive, fast.

I’ll probably stick to open models at that point.

Re: GPT-4.5

#622

Earlier quoted context omitted.

> "Early testing shows that interacting with GPT‑4.5 feels more natural. Its broader knowledge base, improved ability to follow user intent, and greater “EQ” make it useful for tasks like improving writing, programming, and solving practical problems. We also expect it to hallucinate less." "Early testing doesn't show that it hallucinates less, but we expect that putting that sentence nearby will lead you to draw a c…

[flagged]

- Parent is still the top comment.

- 2 hours in, -3.

2 replies:

- [It's because] you're hysterical

- [It's because you sound] like a crypto bro

- [It's because] you make an equally unfounded claim

- [It's because] you didn't provide any proof

(Ed.: It is right in the link! I gave the #s! I can't ctrl-F...What else can I do here...AFAIK can't link images...whatever, here's imgur. https://imgur.com/a/mkDxe78)

- [It's because] you sound personally offended

(Ed.: Is "personally" is a shibboleth here, meaning expressing disappointment in people making things up is so triggering as invalidate the communication that it is made up?)

Re: GPT-4.5

#623
post #167

First impression of GPT-4.5: 1. It is very very slow, for some applications where you want real time interactions is just not viable, the text attached below took 7s to generate with 4o, but 46s with GPT4.5 2. The style it writes is way better: it keeps the tone you ask and makes better improvements on the flow. One of my biggest complaints with 4o is that you want for your content to be more casual and accessible bu…

I'm wondering if generative AI will ultimately result in a very dense / bullet form style of writing. What we are doing now is effectively this: bullet_points' = compress(expand(bullet_points)) We are impressed by lots of text so must expand via LLM in order to impress the reader. Since the reader doesn't have time or interest to read the content they must compress it back into bullet points / quick summary. Really,…

I'm reminded of this great comic

https://marketoonist.com/2023/03/ai-written-ai-read.html

Re: GPT-4.5

#624
post #309

Earlier quoted context omitted.

Are the benchmarks known ahead of time? Could the answer to the benchmarks be in the training data?

In general yes, bench mark pollution is a big problem and why only dynamic benchmarks matter.

This is true, but how would pollution work for a benchmark designed to test hallucinations?

Re: GPT-4.5

#625

Earlier quoted context omitted.

> And LLM's already have tons of productive uses. I disagree strongly with that. Right now they are fun toys to play with, but not useful tools, because they are not reliable. If and when that gets fixed, maybe they will have productive uses. But for right now, not so much.

They are pretty useful tools. Do yourself a favor and get a $100 free trial for Claude, hook it up to Aider, and give it a shot. It makes mistakes, it gets things wrong, and it still saves a bunch of time. A 10 minute refactoring turns into 30 seconds of making a request, 15 seconds of waiting, and a minute of reviewing and fixing up the output. It can give you decent insights into potential problems and error messag…

OP should really save their money. Cursor has a pretty generous free trail and is far from the holy grail.

I recently (in the last month) gave it a shot. I would say once in the maybe 30 or 40 times I used it did it save me any time. The one time it did I had each line filled in with pseudo code describing exactly what it should do… I just didn’t want to look up the APIs

I am glad it is saving you time but it’s far from a given. For some people and some projects, intern level work is unacceptable. For some people, managing is a waste of time.

You’re basically introducing the mythical man month on steroids as soon as you start using these

Re: GPT-4.5

#626
post #390

Earlier quoted context omitted.

I'm all for skepticism of capabilities and cynicism about corporate messaging, but I really don't think there's an interpretation of the word "greater" in this context" that doesn't mean "higher" and "better".

I think the trick is observing what is “better” in this model. EQ is supposed to be “better” than 4o, according to the prose. However, how can an LLM have emotional-anything? LLMs are a regurgitation machine, emotion has nothing to do with anything.

Imagine two greeting cards. One says “I’m so sorry for your loss”, and the other says “Everyone dies, they weren’t special”.

Does one of these have a higher EQ, despite both being ink and paper and definitely not sentient?

Now, imagine they were produced by two different AIs. Does one AI demonstrate higher EQ?

The trick is in seeing that “EQ of a text response” is not the same thing as “EQ of a sentient being”

Re: GPT-4.5

#627

Earlier quoted context omitted.

[flagged]

I suspect people downvote you because the tone of your reply makes it seem like you are personally offended and are now firing back with equally unfounded attacks like a straight up "you are lying". I read the article but can't find the numbers you are referencing. Maybe there's some paper linked I should be looking at? The only numbers I see are from the SimpleQA chart, which are 37.1% vs 61.8% hallucination rate. T…

It's in the link.

I don't know what else to say.

Here, imgur: https://imgur.com/a/mkDxe78. Can't get easier.

> equally unfounded attacks

No, because I have a source and didn't make up things someone else said.

> a straight up "you are lying".

Right, because they are. There are hallucination stats right in the post he mocks for not prvoiding stats.

> That's nice but considering the price increase,

I can't believe how quickly you acknowledge it is in the post after calling the idea it was in the post "equally unfounded". You are looking at the stats. They were lying.

> "That's nice but considering the price increase,"

That's nice and a good argument! That's not what I replied to. I replied to they didn't provide any stats.

Re: GPT-4.5

#628

Earlier quoted context omitted.

Huh. Disregarding the 4.5-specific bit here, a browser extension or possibly website that did this in general could be really useful. Maybe even something that just noticed whenever you visited a site that had had significant HN discussion in the past, then let you trigger a summary.

there are literally hundreds of extensions and sites that do this the problem is that they are competing each other into the ground hence they go unmaintained very quickly getrecall.ai has been the most mature so far

Hundreds that specifically focus on noticing a page you’re currently viewing has been not only posted to but undergone significant discussion on HN, and then providing a summary of those conversations?

Or that just provide summaries in general?

Re: GPT-4.5

#629
post #56

Earlier quoted context omitted.

release blog post author: this is definitely a research preview ceo: it's ready the pricing is probably a mixture of dealing with GPU scarcity and intentionally discouraging actual users. I can't imagine the pressure they must be under to show they are releasing and staying ahead, but Altman's tweet makes it clear they aren't really ready to sell this to the general public yet.

Yeap, that the thing, they are not ahead anymore. Not since last summer at least. Yes they have probably largest customer base, but their models are not the best for a while already.

They don't even have the largest customer base. Google is serving AI Overviews at the top of their search engine to an order of magnitude more people.

Re: GPT-4.5

#630

Earlier quoted context omitted.

> "Early testing shows that interacting with GPT‑4.5 feels more natural. Its broader knowledge base, improved ability to follow user intent, and greater “EQ” make it useful for tasks like improving writing, programming, and solving practical problems. We also expect it to hallucinate less." "Early testing doesn't show that it hallucinates less, but we expect that putting that sentence nearby will lead you to draw a c…

That's some top-tier sales work right there. I suck at and hate writing the mildly deceptive corporate puffery that seems to be in vogue. I wonder if GPT-4.5 can write that for me or if it's still not as good at it as the expert they paid to put that little gem together.

I am reasonably to very skeptical about the valuation of LLM firms but you don’t even seem willing to engage with the question about the value of these tools.
Post reply on HN