Rio de Janeiro's city government model Rio3.5 beats Qwen3.7 in recent benchmarks
11–20 of 51 posts
Re: Rio de Janeiro's city government model Rio3.5 beats Qwen3.7 in recent benchmarks
#12Re: Rio de Janeiro's city government model Rio3.5 beats Qwen3.7 in recent benchmarks
#13Every day I'm reminded why I don't spend time on twitter. What use does it have to claim "X is better than Y in benchmark Z, disagreeing with that means disagreeing with me" Information is power, dick measurements are not.
Re: Rio de Janeiro's city government model Rio3.5 beats Qwen3.7 in recent benchmarks
#14[flagged]
A government ideally is a representation of the democratically chosen will of the people. If it is not, work towards making it so. IMO wherever someone says "the government" we should mentally substitute "we all, collectively". But a specific type of person appears to labour under the illusion that somehow we can get by without we all collectively steering our direction and choosing people who do what needs to be don…
Interestingly, the people who try to separate themselves from "the government" also seem to be the kind of people who want to "spread our model of democracy to the rest of the world".
How they can even reconcile being such a great democracy that the world needs to ~copy~ be force-fed with having an adversary government I don't know. The cognitive dissonance is so great that it's hard to fathom.
Re: Rio de Janeiro's city government model Rio3.5 beats Qwen3.7 in recent benchmarks
#15[flagged]
A government ideally is a representation of the democratically chosen will of the people. If it is not, work towards making it so. IMO wherever someone says "the government" we should mentally substitute "we all, collectively". But a specific type of person appears to labour under the illusion that somehow we can get by without we all collectively steering our direction and choosing people who do what needs to be don…
Re: Rio de Janeiro's city government model Rio3.5 beats Qwen3.7 in recent benchmarks
#16Every day I'm reminded why I don't spend time on twitter. What use does it have to claim "X is better than Y in benchmark Z, disagreeing with that means disagreeing with me" Information is power, dick measurements are not.
Re: Rio de Janeiro's city government model Rio3.5 beats Qwen3.7 in recent benchmarks
#17https://xcancel.com/ZenMagnets/status/2065796012820848699 Correct me if I'm wrong but reading through the comments of the thread this seems to be post training/fine tuning.
Thanks, Firefox and uBlock does not let me watch any X content (I guess this is a good thing)
Re: Rio de Janeiro's city government model Rio3.5 beats Qwen3.7 in recent benchmarks
#18As for the benchmarks: If you spend any time playing with fine tunes of published models you know that benchmarks are gamed so much that they're a useless indicator of performance for models from small teams. It's too easy to fine tune a model to perform well on the benchmarks, release it, put a line on your resume saying you released a model that beat the major labs on benchmarks, and then try to use that to jump into a new job. The temptation is high.
There are a lot of fringe models and fine tunes that claim to have better performance on some benchmark. Then you try to use them and find they're often worse at general tasks than the base model.
I would wait and see if these results hold across other benchmarks. It's cool that the city is doing something with AI, but this is something where extraordinary claims require extraordinary evidence. I doubt a small, previously unknown team has unlocked something secret that the team who made Qwen couldn't figure out. It's more likely it was fine tuned for a specific outcome (possibly these benchmarks) and performance in other areas was reduced as a consequence.
Re: Rio de Janeiro's city government model Rio3.5 beats Qwen3.7 in recent benchmarks
#19https://xcancel.com/ZenMagnets/status/2065796012820848699 Correct me if I'm wrong but reading through the comments of the thread this seems to be post training/fine tuning.
Yes. It's post training in qwen using the novel SwiReasoning framework.