Live data from Hacker News

GPT-4.1 in the API

openai.com

101–110 of 513 posts

Re: GPT-4.1 in the API

#101

Numbers for SWE-bench Verified, Aider Polyglot, cost per million output tokens, output tokens per second, and knowledge cutoff month/year: SWE Aider Cost Fast Fresh Claude 3.7 70% 65% $15 77 8/24 Gemini 2.5 64% 69% $10 200 1/25 GPT-4.1 55% 53% $8 169 6/24 DeepSeek R1 49% 57% $2.2 22 7/24 Grok 3 Beta ? 53% $15 ? 11/24 I'm not sure this is really an apples-to-apples comparison as it may involve different test scaffoldi…

Yes, it is available in Cursor[1] and Windsurf[2] as well.

[1] https://twitter.com/cursor_ai/status/1911835651810738406

[2] https://twitter.com/windsurf_ai/status/1911833698825286142

Re: GPT-4.1 in the API

#102
post #33

Testing against unspecified other "leading" models allows for shenanigangs: > Qodo tested GPT‑4.1 head-to-head against other leading models [...] they found that GPT‑4.1 produced the better suggestion in 55% of cases The linked blog post goes 404: https://www.qodo.ai/blog/benchmarked-gpt-4-1/

The post seems to be up now and seems to compare it slightly favorable to Claude 3.7.

Re: GPT-4.1 in the API

#103

GPT-4.1 Pricing (per 1M tokens): gpt-4.1 - Input: $2.00 - Cached Input: $0.50 - Output: $8.00 gpt-4.1-mini - Input: $0.40 - Cached Input: $0.10 - Output: $1.60 gpt-4.1-nano - Input: $0.10 - Cached Input: $0.025 - Output: $0.40

The fact that they're raising the price for the mini models by 166% is pretty notable. gpt-4o-mini for comparison: - Input: $0.15 - Cached Input $0.075 - Output: $0.60

Seems like 4.1 nano ($0.10) is closer to the replacement and 4.1 mini is a new in-between price

Re: GPT-4.1 in the API

#104

Numbers for SWE-bench Verified, Aider Polyglot, cost per million output tokens, output tokens per second, and knowledge cutoff month/year: SWE Aider Cost Fast Fresh Claude 3.7 70% 65% $15 77 8/24 Gemini 2.5 64% 69% $10 200 1/25 GPT-4.1 55% 53% $8 169 6/24 DeepSeek R1 49% 57% $2.2 22 7/24 Grok 3 Beta ? 53% $15 ? 11/24 I'm not sure this is really an apples-to-apples comparison as it may involve different test scaffoldi…

[deleted]

Re: GPT-4.1 in the API

#105
post #29

GPT-4.1 probably is a distilled version of GPT-4.5 I dont understand the constant complaining about naming conventions. The number system differentiates the models based on capability, any other method would not do that. After ten models with random names like "gemini", "nebula" you would have no idea which is which. Its a low IQ take. You dont name new versions of software as completely different software Also, Yest…

> The number system differentiates the models based on capability, any other method would not do that. Please rank GPT-4, GPT-4 Turbo, GPT-4o, GPT-4.1-nano, GPT-4.1-mini, GPT-4.1, GPT-4.5, o1-mini, o1, o1 pro, o3-mini, o3-mini-high, o3, and o4-mini in terms of capability without consulting any documentation.

I meant this is actually straight-forward if you've been paying even the remotest of attention.

Chronologically:

GPT-4, GPT-4 Turbo, GPT-4o, o1-preview/o1-mini, o1/o3-mini/o3-mini-high/o1-pro, gpt-4.5, gpt-4.1

Model iterations, by training paradigm:

SGD pretraining with RLHF: GPT-4 -> turbo -> 4o

SGD pretraining w/ RL on verifiable tasks to improve reasoning ability: o1-preview/o1-mini -> o1/o3-mini/o3-mini-high (technically the same product with a higher reasoning token budget) -> o3/o4-mini (not yet released)

reasoning model with some sort of Monte Carlo Search algorithm on top of reasoning traces: o1-pro

Some sort of training pipeline that does well with sparser data, but doesn't incorporate reasoning (I'm positing here, training and architecture paradigms are not that clear for this generation): gpt-4.5, gpt-4.1 (likely fine-tuned on 4.5)

By performance: hard to tell! Depends on what your task is, just like with humans. There are plenty of benchmarks. Roughly, for me, the top 3 by task are:

Creative Writing: gpt-4.5 -> gpt-4o

Business Comms: o1-pro -> o1 -> o3-mini

Coding: o1-pro -> o3-mini (high) -> o1 -> o3-mini (low) -> o1-mini-preview

Shooting the shit: gpt-4o -> o1

It's not to dismiss that their marketing nomenclature is bad, just to point out that it's not that confusing for people that are actively working with these models have are a reasonable memory of the past two years.

Re: GPT-4.1 in the API

#106
post #79
post #2

pretty wild versioning that GPT 4.1 is newer and better in many regards than GPT 4.5.

it's worse on nearly every benchmark

no? it's better on AIME '24, Multilingual MMLU, SWE-bench, Aider’s polyglot, MMMU, ComplexFuncBench

and it ties on a lot of benchmarks

Re: GPT-4.1 in the API

#107
post #22

Earlier quoted context omitted.

> using v0, I replicated a full nextjs UI copying a major saas player. No backend integration, but the design and UX were stunning AI is amazing, now all you need to create a stunning UI is for someone else to make it first so an AI can rip it off. Not beating the "plagiarism machine" allegations here.

Heres a secret: Most of the highest funded VC backed software companies are just copying a competitor with a slight product spin/different pricing model

> Jim Barksdale, used to say there’s only two ways to make money in business: One is to bundle; the other is unbundle

https://a16z.com/the-future-of-work-cars-and-the-wisdom-in-s...

Re: GPT-4.1 in the API

#108
> We will also begin deprecating GPT‑4.5 Preview in the API, as GPT‑4.1 offers improved or similar performance on many key capabilities at much lower cost and latency.

why would they deprecate when it's the better model? too expensive?

Re: GPT-4.1 in the API

#110
post #17

Earlier quoted context omitted.

Gemini 2.5 Pro gets 64% on SWE-bench verified. Sonnet 3.7 gets 70% They are reporting that GPT-4.1 gets 55%.

Are those with «thinking» or without?

based on their release cadence, I suspect that o4-mini will compete on price, performance, and context length with the rest of these models.
Post reply on HN