From GPT 5.1 Thinking: ARC AGI v2: 17.6% -> 52.9% SWE Verified: 76.3% -> 80% That's pretty good!
GPT-5.2
231–240 of 1001 posts
Re: GPT-5.2
#232Earlier quoted context omitted.
Yep, the point we wanted to make here is that GPT-5.2's vision is better, not perfect. Cherrypicking a perfect output would actually mislead readers, and that wasn't our intent.
That would be a laudable goal, but I feel like it's contradicted by the text: > Even on a low-quality image, GPT‑5.2 identifies the main regions and places boxes that roughly match the true locations of each component I would not consider it to have "identified the main regions" or to have "roughly matched the true locations" when ~1/3 of the boxes have incorrect labels . The remark "even on a low-quality image" is n…
Re: GPT-5.2
#233Earlier quoted context omitted.
Arc-AGI is just an iq test. I don’t see the problem with training it to be good at iq tests because that’s a skill that translates well.
Exactly. In principle, at least, the only way to overfit to Arc-AGI is to actually be that smart. Edit: if you disagree, try actually TAKING the Arc-AGI 2 test, then post.
Re: GPT-5.2
#234Is it me, or did it still get at least three placements of components (RAM and PCIe slots, plus it's DisplayPort and not HDMI) in the motherboard image[0] completely wrong? Why would they use that as a promotional image? 0: https://images.ctfassets.net/kftzwdyauwt9/6lyujQxhZDnOMruN3f...
Yep, the point we wanted to make here is that GPT-5.2's vision is better, not perfect. Cherrypicking a perfect output would actually mislead readers, and that wasn't our intent.
Re: GPT-5.2
#235The thing about OpenAI is their models never fit anywhere for me. Yes they maybe smart or even the smartest models but they are alway so fucking slow. The ChatGPT web app is literally usable for me. I ask simple task and it does most extreme shit jsut to get an answer that the same as Claude or Gemini. For example, I asked ChatGPT to take a chart and convert into a table. It went and cut up the image and zoomed in fo…
I use models based on the task. They still seem specialized and better at specific tasks. If I have a question I tend to go to it. If I need code, I tend to go to Claude (Code).
I go to ChatGPT for questions I have because I value an accurate answer over a quick answer and, in my experience, it tends to give me more accurate answers because of its (over) willingness to go to the web for search results and question its instincts. Claude is much more likely to make an assumption and its search patterns aren't as thorough. The slow answers don't bother me because it's an expectation I have for how I use it and they've made that use case work really well with background processing and notifications.
Re: GPT-5.2
#236For me the last remaining killer feature of ChatGPT is the quality of the voice chat. Do any of the competitors have something like that?
gemini live is a thing - never tried chaptgpt, are they not similar?
Gemini responds in what I think is Spanish, or perhaps Portuguese.
However I can hand an 8 minute long 48k mono mp3 of a nuanced Latin speaker who nasalizes his vowels, and makes regular use of elision to Gemini-3-pro-preview and it will produce an accurate macronized Latin transcription. It's pretty mind blowing.
Re: GPT-5.2
#237Earlier quoted context omitted.
Seems pretty false if you look at the model card and web site of Opus 4.5 that is… (check notes) their latest model.
Building a good model generally means it will do well on benchmarks too. The point of the speculation is that Anthropic is not focused on benchmaxxing which is why they have models people like to use for their day-to-day. I use Gemini, Anthropic stole $50 from me (expired and kept my prepaid credits) and I have not forgiven them yet for it, but people rave about claude for coding so I may try the model again through…
Re: GPT-5.2
#238> Unlike the previous GPT-5.1 model, GPT-5.2 has new features for managing what the model "knows" and "remembers to improve accuracy. Dumb nit, but why not put your own press release through your model to prevent basic things like missing quote marks? Reminds me of that time an OAI released wildly inaccurate copy/pasted bar charts.
Re: GPT-5.2
#239For me the last remaining killer feature of ChatGPT is the quality of the voice chat. Do any of the competitors have something like that?