“GPT‑4.1 scores 54.6% on SWE-bench Verified, improving by 21.4%abs over GPT‑4o and 26.6%abs over GPT‑4.5—making it a leading model for coding.” 4.1 is 26.6% better at coding than 4.5. Got it. Also…see the em dash
GPT-4.1 in the API
301–310 of 513 posts
Re: GPT-4.1 in the API
#302> You're eligible for free daily usage on traffic shared with OpenAI through April 30, 2025. > Up to 1 million tokens per day across gpt-4.5-preview, gpt-4.1, gpt-4o and o1 > Up to 10 million tokens per day across gpt-4.1-mini, gpt-4.1-nano, gpt-4o-mini, o1-mini and o3-mini > Usage beyond these limits, as well as usage for other models, will be billed at standard rates. Some limitations apply. I just found this optio…
Re: GPT-4.1 in the API
#303Earlier quoted context omitted.
It's definitely an issue. Even the simplest use case of "create React app with Vite and Tailwind" is broken with these models right now because they're not up to date.
Time to start moving back to Java & Spring. 100% backwards compatibility and well represented in 15 years worth of training data, hah.
Re: GPT-4.1 in the API
#304Earlier quoted context omitted.
Anyone making claims with a horizon beyond two months about structure or capabilities will be wrong - it's sama's job to show confidence and vision and calm stakeholders, but if you're paying attention to the field, the release and research cycles are still contracting, with no sense of slowing any time soon. I've followed AI research daily since GPT-2, the momentum is incredible, and even if the industry sticks with…
the release and research cycles are still contracting Not necessarily progress or benchmarks that as a broader picture you would look at (MMLU etc) GPT-3 was an amazing step up from GPT-2, something scientists in the field really thought was 10-15 years out at least done in 2, instruct/RHLF for GPTs was a similar massive splash, making the second half of 2021 equally amazing. However nothing since has really been tha…
Image generation suddenly went from gimmick to useful now that prompt adherence is so much better (eagerly waiting for that to be in the API)
Coding performance continues to improve noticeably (for me). Claude 3.7 felt like a big step from 4o/3.5. Gemini 2.5 in a similar way.compared to just 6 months ago I can give bigger and more complex pieces of work to it and get relatively good output back. (Net acceleration)
Audio-2-audio seems like it will be a big step as well. I think this has much more potential than the STT-LLM-TTS architecture commonly used today (latency, quality)
Re: GPT-4.1 in the API
#305I think an under appreciated reality is that all of the large AI labs and OpenAI in particular are fighting multiple market battles at once. This is coming across in both the number of products and the packaging. 1, to win consumer growth they have continued to benefit on hyper viral moments, lately that was was image generation in 4o, which likely was technically possible a long time before launched. 2, for enterpri…
Re: GPT-4.1 in the API
#306Earlier quoted context omitted.
As an aside, one of the worst aspects of the rise of LLMs, for me, has been the wholesale replacement of engineering with trial-and-error hand-waving. Try this, or maybe that, and maybe you'll see a +5% improvement. Why? Who knows. It's just not how I like to work.
Out of curiosity, what do you work on where you don’t have to experiment with different solutions to see what works best?
Re: GPT-4.1 in the API
#307> You're eligible for free daily usage on traffic shared with OpenAI through April 30, 2025. > Up to 1 million tokens per day across gpt-4.5-preview, gpt-4.1, gpt-4o and o1 > Up to 10 million tokens per day across gpt-4.1-mini, gpt-4.1-nano, gpt-4o-mini, o1-mini and o3-mini > Usage beyond these limits, as well as usage for other models, will be billed at standard rates. Some limitations apply. I just found this optio…
Re: GPT-4.1 in the API
#308Earlier quoted context omitted.
As an aside, one of the worst aspects of the rise of LLMs, for me, has been the wholesale replacement of engineering with trial-and-error hand-waving. Try this, or maybe that, and maybe you'll see a +5% improvement. Why? Who knows. It's just not how I like to work.
Out of curiosity, what do you work on where you don’t have to experiment with different solutions to see what works best?
Re: GPT-4.1 in the API
#309Earlier quoted context omitted.
Anyone making claims with a horizon beyond two months about structure or capabilities will be wrong - it's sama's job to show confidence and vision and calm stakeholders, but if you're paying attention to the field, the release and research cycles are still contracting, with no sense of slowing any time soon. I've followed AI research daily since GPT-2, the momentum is incredible, and even if the industry sticks with…
the release and research cycles are still contracting Not necessarily progress or benchmarks that as a broader picture you would look at (MMLU etc) GPT-3 was an amazing step up from GPT-2, something scientists in the field really thought was 10-15 years out at least done in 2, instruct/RHLF for GPTs was a similar massive splash, making the second half of 2021 equally amazing. However nothing since has really been tha…
Re: GPT-4.1 in the API
#310I think an under appreciated reality is that all of the large AI labs and OpenAI in particular are fighting multiple market battles at once. This is coming across in both the number of products and the packaging. 1, to win consumer growth they have continued to benefit on hyper viral moments, lately that was was image generation in 4o, which likely was technically possible a long time before launched. 2, for enterpri…
On that note, I want to see benchmarks for which LLM's are best at translating between languages. To me, it's an entire product category.