I gave the Crossword puzzle to Claude and got a correct response[1]. The fact that they are comparing this to gpt4o and not to gpt4 suggests that it is less impressive than they are trying to pretend. [1]: Based on the given clues, here's the solved crossword puzzle: +---+---+---+---+---+---+ | E | S | C | A | P | E | +---+---+---+---+---+---+ | S | E | A | L | E | R | +---+---+---+---+---+---+ | T | E | R | E | S |…
As good as Claude has gotten recently in reasoning, they are likely using RL behind the scenes too. Supposedly, o1/strawberry was initially created as an engine for high-quality synthetic reasoning data for the new model generation. I wonder if Anthropic could release their generator as a usable model too.
on X I see a totally different energy more about hyping it
on HN I see reserved and collected take which I trust more.
I do wonder why they chose gpt4o which I never bother to use for coding.
Claude is still king and looks like I won't have to subscribe to ChatGPT Plus seeing it fail on some of the important experiments run by folks on HN
If anything these type of releases that air more on the side of hype given OpenAI's track record