Live data from Hacker News

GPT-4.1 in the API

openai.com

371–380 of 513 posts

Re: GPT-4.1 in the API

#371

From OpenAI's announcement: > Qodo tested GPT‑4.1 head-to-head against Claude Sonnet 3.7 on generating high-quality code reviews from GitHub pull requests. Across 200 real-world pull requests with the same prompts and conditions, they found that GPT‑4.1 produced the better suggestion in 55% of cases. Notably, they found that GPT‑4.1 excels at both precision (knowing when not to make suggestions) and comprehensiveness…

That's a marketing page for something called qodo that sells ai code reviews. At no point were the ai code reviews judged by competent engineers. It is just ai generated trash all the way down.

Re: GPT-4.1 in the API

#372

Sam made a strange statement imo in a recent Ted Talk. He said (something like) models come and go but they want to be the best platform. For me, it was jaw dropping. Perhaps he didn't mean it the way it sounded, but seemed like a major shift to me.

OpenAI has been a product company ever since ChatGPT launched.

Their value is firmly rooted in how they wrap ux around models.

Re: GPT-4.1 in the API

#373
post #317

As a ChatGPT user, I'm weirdly happy that it's not available there yet. I already have to make a conscious choice between - 4o (can search the web, use Canvas, evaluate Python server-side, generate images, but has no chain of thought) - o3-mini (web search, CoT, canvas, but no image generation) - o1 (CoT, maybe better than o3, but no canvas or web search and also no images) - Deep Research (very powerful, but I have…

I use them as follows: o1-pro: anything important involving accuracy or reasoning. Does the best at accomplishing things correctly in one go even with lots of context. deepseek R1: anything where I want high quality non-academic prose or poetry. Hands down the best model for these. Also very solid for fast and interesting analytical takes. I love bouncing ideas around with R1 and Grok-3 bc of their fast responses and…

You probably know this and are looking for consistency but, a little trick I use is to feed the original data of what I need as a diagram and to re-imagine, it as an image “ready for print” - not native, but still a time saver and just studying with unstructured data or handles this surprisingly well. Again not native…naive, yes. Native, not yet. Be sure to double check triple check as always. give it the ol’ OCD treatment.

Re: GPT-4.1 in the API

#374
post #317

As a ChatGPT user, I'm weirdly happy that it's not available there yet. I already have to make a conscious choice between - 4o (can search the web, use Canvas, evaluate Python server-side, generate images, but has no chain of thought) - o3-mini (web search, CoT, canvas, but no image generation) - o1 (CoT, maybe better than o3, but no canvas or web search and also no images) - Deep Research (very powerful, but I have…

> 4.5 (better in creative writing, and probably warmer sound thanks to being vinyl based and using analog tube amplifiers

Ha! That's the funniest and best description of 4.5 I've seen.

Re: GPT-4.1 in the API

#375

Sam made a strange statement imo in a recent Ted Talk. He said (something like) models come and go but they want to be the best platform. For me, it was jaw dropping. Perhaps he didn't mean it the way it sounded, but seemed like a major shift to me.

Before everyone caught up:

    We are in a race to make a new God, and the company that wins the race will have omnipotent power beyond our comprehension. 
After everyone else caught up:

    The models come and go, some are SOTA in evals and some not.  What matters is our platform and market share.

Re: GPT-4.1 in the API

#376

> They feature a refreshed knowledge cutoff of June 2024. As opposed to Gemini 2.5 Pro having cutoff of Jan 2025. Honestly this feels underwhelming and surprising. Especially if you're coding with frameworks with breaking changes, this can hurt you.

sometimes it feels like openai keeps serving the same base dish—just adding new toppings. sure, the menu keeps changing, but it all kinda tastes the same. now the menu is getting too big.

nice to see that we aren't stuck in october of 2023 anymore!

Re: GPT-4.1 in the API

#377
post #317

As a ChatGPT user, I'm weirdly happy that it's not available there yet. I already have to make a conscious choice between - 4o (can search the web, use Canvas, evaluate Python server-side, generate images, but has no chain of thought) - o3-mini (web search, CoT, canvas, but no image generation) - o1 (CoT, maybe better than o3, but no canvas or web search and also no images) - Deep Research (very powerful, but I have…

> - Deep Research (very powerful, but I have only 10 attempts per month, so I end up using roughly zero) Same here, which is a real shame. I've switched to DeepResearch with Gemini 2.5 Pro over the last few days where paid users have a 20/day limit instead of 10/month and it's been great, especially since now Gemini seems to browse 10x more pages than OpenAI Deep Research (on the order of 200-400 pages versus 20-40).…

I also like Perplexity’s 3/day limit! If I use them up (which I almost never do) I can just refresh the next day

Re: GPT-4.1 in the API

#379

Earlier quoted context omitted.

You are definitely not a layman if you know the difference between 4.5 and 4o. The average user thinks ai = openai = chatgpt.

Well, okay, but I'm certainly not an expert who knows the fine differences between all the models available on chat.com. So I'm somewhere between your definition of "layman" and your definition of "expert" (as are, I suspect, most people on this forum).

If you know the difference between 4.5 and 4o, it'll take you 20 minutes max to figure out the theoretical differences between the other models, which is not bad for a highly technical emerging field.

Re: GPT-4.1 in the API

#380
post #317

As a ChatGPT user, I'm weirdly happy that it's not available there yet. I already have to make a conscious choice between - 4o (can search the web, use Canvas, evaluate Python server-side, generate images, but has no chain of thought) - o3-mini (web search, CoT, canvas, but no image generation) - o1 (CoT, maybe better than o3, but no canvas or web search and also no images) - Deep Research (very powerful, but I have…

> 4.5 (better in creative writing, and probably warmer sound thanks to being vinyl based and using analog tube amplifiers, but slower and request limited, and I don't even know which of the other features it supports) Is that an LLM hallucination?

Possibly, but it's running on 100% wetware, I promise!
Post reply on HN