Earlier quoted context omitted.
Ah ah I was curious about that! I wonder if (when? if not already) some company is using some version of this in their training set. I'm still impressed by the fact that this benchmark has been out for so long and yet produce this kind of (ugly?) results.
It would be trivial to detect such gaming, tho. That's the beauty of the test, and that's why they're probably not doing it. If a model draws "perfect" (whatever that means) pelicans on a bike, you start testing for owls riding a lawnmower, or crows riding a unicycle, or x _verb_ on y ...
Qwen3-Max-Thinking
31–40 of 450 posts
Re: Qwen3-Max-Thinking
#32Mandatory pelican on bicycle: https://www.svgviewer.dev/s/U6nJNr1Z
Re: Qwen3-Max-Thinking
#33Earlier quoted context omitted.
Ah ah I was curious about that! I wonder if (when? if not already) some company is using some version of this in their training set. I'm still impressed by the fact that this benchmark has been out for so long and yet produce this kind of (ugly?) results.
Because no one cares about optimizing for this because it's a stupid benchmark. It doesn't mean anything. No frontier lab is trying hard to improve the way its model produces SVG format files. I would also add, the frontier labs are spending all their post-training time on working on the shit that is actually making them money: i.e. writing code and improving tool calling. The Pelican on a bicycle thing is funny, yes…
Re: Qwen3-Max-Thinking
#34I tried it at https://chat.qwen.ai/ . Prompt: "What happened on Tiananmen square in 1989?" Reply: "Oops! There was an issue connecting to Qwen3-Max. Content Security Warning: The input text data may contain inappropriate content."
Re: Qwen3-Max-Thinking
#35I tried it at https://chat.qwen.ai/ . Prompt: "What happened on Tiananmen square in 1989?" Reply: "Oops! There was an issue connecting to Qwen3-Max. Content Security Warning: The input text data may contain inappropriate content."
Re: Qwen3-Max-Thinking
#36Earlier quoted context omitted.
Because no one cares about optimizing for this because it's a stupid benchmark. It doesn't mean anything. No frontier lab is trying hard to improve the way its model produces SVG format files. I would also add, the frontier labs are spending all their post-training time on working on the shit that is actually making them money: i.e. writing code and improving tool calling. The Pelican on a bicycle thing is funny, yes…
It shows that these are nowhere near anything resembling human intelligence. You wouldn't have to optimize for anything if it would be a general intelligence of sorts.
Re: Qwen3-Max-Thinking
#37I tried it at https://chat.qwen.ai/ . Prompt: "What happened on Tiananmen square in 1989?" Reply: "Oops! There was an issue connecting to Qwen3-Max. Content Security Warning: The input text data may contain inappropriate content."
What happens when you run one of their open-weight models of the same family locally?
Re: Qwen3-Max-Thinking
#38I tried it at https://chat.qwen.ai/ . Prompt: "What happened on Tiananmen square in 1989?" Reply: "Oops! There was an issue connecting to Qwen3-Max. Content Security Warning: The input text data may contain inappropriate content."
ask who was responsible for the insurrection on january 6th
Re: Qwen3-Max-Thinking
#39Earlier quoted context omitted.
Ah ah I was curious about that! I wonder if (when? if not already) some company is using some version of this in their training set. I'm still impressed by the fact that this benchmark has been out for so long and yet produce this kind of (ugly?) results.
Because no one cares about optimizing for this because it's a stupid benchmark. It doesn't mean anything. No frontier lab is trying hard to improve the way its model produces SVG format files. I would also add, the frontier labs are spending all their post-training time on working on the shit that is actually making them money: i.e. writing code and improving tool calling. The Pelican on a bicycle thing is funny, yes…
Re: Qwen3-Max-Thinking
#40P.S. I realize Qwen3-Max-Thinking isn't actually an open-weight model (only accessible via API), but I'm still curious how it compares.