Earlier quoted context omitted.
gemma4-e4b is 50% better than gemma4-26b in your benchmark, something's wrong
Yes those two models were tested on my own PC (local inference using my own CPU/GPU). So something my be bugged on my setup. gemma4-26b should be far better than gemma4-e4b.
OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API
151–160 of 174 posts
Re: OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API
#152Re: OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API
#153Re: OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API
#154Just tested it on my homemade Wordpress+GravityForms benchmark and it's one of the worst model of the leaderboard performance wise and the worst value wise: https://github.com/guilamu/llms-wordpress-plugin-benchmark I know it's only on a single benchmark, but I dont understand how it can be so bad...
HN is the last bastion of serious inquiry these days. But its not immune as OPs comment proves.
Re: OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API
#155Earlier quoted context omitted.
personifying ai is incredibly cringe no matter how weird your comparison is
Imagine a drunk developer. Sparks of brilliance while missing obvious trees.
Re: OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API
#156Earlier quoted context omitted.
I gave 4.6, 4.7 and GPT 5.5 the same prompt and task to reverse engineer a collection of sample vector files from an obscure Amiga CAD program and create a detailed txt specification and a python converter that converts to SVG and produce a report so I can visually verify. 4.6 did very well. 90% perfect on first try, got to 100% with just a few followups. 4.7 failed horribly. First produced garbage output and claimed…
Interesting that 4.7 failed like that. Seems 5.5 is impressive but is oh so expensive. Would be interesting if you ran your same test with Deepseek v4 and some of the other Chinese models.
I also tried GLM 5.1, it's first attempt was such a disaster I didn't bother working with it any further. It also took by far the longest and wasted a bunch of time/tokens trying to find other converters online (and failing) instead of just reverse engineering the format from the sample files given.
Re: OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API
#157Earlier quoted context omitted.
Please do an update when you're ready, this sounds like madness to me so I'd love to see what the output is. Whatever it is I have to know.
Typical AI psychosis. They might notice it soon or stay in this condition for months.
Re: OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API
#158Earlier quoted context omitted.
Has that task accomplished anything yet?
It made Sam richer.
Re: OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API
#159Earlier quoted context omitted.
That’s actually crazy, what kind of task is that? And is that a recurring kind of task like some analysis, or coding related?
Coding (along with docs, tests obviously), rewriting a huge chunk of the KVM hypervisor (in Kernel 7, started in the -rc2) and KSM and other modules, can't say too much about it yet (might do an announcement in coming weeks) . The coding is automated but the plan took days of manual arguing (with all models possible) prior (while doing other things during waiting times as I currently manage 70 repos for an upcoming r…
Re: OpenAI releases GPT-5.5 and GPT-5.5 Pro in the API
#160Earlier quoted context omitted.
Or nostalgia for simpler times
That as well. But everyone reading GP’s posts knows in their bones that it’s unsustainable. It’s economically unsustainable and environmentally unsustainable, and in that context it strikes me as pure hoarding behaviour. Taking as much as they can for themselves before the house of cards crashes down. I have no sympathy for OpenAI or Anthropic as corporations, but if these are the new tools of the trade, then platfor…