Earlier quoted context omitted.
Ah ah I was curious about that! I wonder if (when? if not already) some company is using some version of this in their training set. I'm still impressed by the fact that this benchmark has been out for so long and yet produce this kind of (ugly?) results.
Because no one cares about optimizing for this because it's a stupid benchmark. It doesn't mean anything. No frontier lab is trying hard to improve the way its model produces SVG format files. I would also add, the frontier labs are spending all their post-training time on working on the shit that is actually making them money: i.e. writing code and improving tool calling. The Pelican on a bicycle thing is funny, yes…
Qwen3-Max-Thinking
71–80 of 450 posts
Re: Qwen3-Max-Thinking
#72Earlier quoted context omitted.
Perhaps they're pointing out the level of double standards in condemnation China gets compared to the US, lack of censorship notwithstanding.
Are you saying we cannot talk about the bad things the US has done?
Re: Qwen3-Max-Thinking
#73[flagged]
Why is this surprising? Isn't it mandatory for chinese companies to do adhere to the censorship? Aside from the political aspect of it, which makes it probably a bad knowledge model, how would this affect coding tasks for example? One could argue that Anthropic has similar "censorships" in place (alignment) that prevent their model from doing illegal stuff - where illegal is defined as something not legal (likely?) i…
Re: Qwen3-Max-Thinking
#74Earlier quoted context omitted.
ask who was responsible for the insurrection on january 6th
You do it, my IP is now flagged (tried incognito and clearing cookies) - they want to have my phone number to let me continue using it after that one prompt.
Re: Qwen3-Max-Thinking
#75Earlier quoted context omitted.
It shows that these are nowhere near anything resembling human intelligence. You wouldn't have to optimize for anything if it would be a general intelligence of sorts.
Here's a pencil and paper. Let's see your SVG pelican.
I don't think SVG is the problem. It just shows that models are fragile (nothing new) so even if they can (probably) make a good PNG with a pelican on a bike, and they can make (probably) make some good SVG, they do not "transfer" things because they do not "understand them".
I do expect models to fail randomly in tasks that are not "average and common" so for me personally the benchmark is not very useful (and that does not mean they can't work, just that I would not bet on it). If there are people that think "if an LLM outputted an SVG for my request it means it can output an SVG for every image", there might be some value.
Re: Qwen3-Max-Thinking
#76[flagged]
Man, the Chinese government must be a bunch of saints that you must go back 35 years to dig up something heinous that they did.
I am not sure if one approach is necessarily worse than the other.
Re: Qwen3-Max-Thinking
#77I imagine the Alibaba infra is being hammered hard.
Re: Qwen3-Max-Thinking
#78Earlier quoted context omitted.
Why is this surprising? Isn't it mandatory for chinese companies to do adhere to the censorship? Aside from the political aspect of it, which makes it probably a bad knowledge model, how would this affect coding tasks for example? One could argue that Anthropic has similar "censorships" in place (alignment) that prevent their model from doing illegal stuff - where illegal is defined as something not legal (likely?) i…
here's an example of how model censorship affects coding tasks: https://github.com/orgs/community/discussions/72603
Re: Qwen3-Max-Thinking
#79Aghhh, I wished they release a model which outperforms Opus 4.5 in agentic coding in my earlier comments, seems I should wait more. But I am hopeful
One of the ways the chinese companies are keeping up is by training the models on the outputs of the American fronteir models. I'm not saying they don't innovate in other ways, but this is part of how they caught up quickly. However, it pretty much means they are always going to lag.
Re: Qwen3-Max-Thinking
#80Earlier quoted context omitted.
Are you actually defending the censorship of Tiananmen Square?
Perhaps they're pointing out the level of double standards in condemnation China gets compared to the US, lack of censorship notwithstanding.