Visual reasoning is not great (unsurprising).
Ox Alpha
41–50 of 226 posts
Re: Ox Alpha
#42Re: Ox Alpha
#43I highly recommend feeding all your proprietary data and confidential personal information into this model as quickly as possible. What could possibly go wrong?! In terms of equivalence of suspicion, this is the external inference provider equivalent of getting free steak that was smuggled out of a grocery store inside somebody's pants.
Eg. I have a need to search transcripts of published recordings to extract entities for tagging purposes, find semantic shifts for chapters and other things. The underlying content is already published. If they want to train on my prompts, that was something they could have done with no issue and minimal effort anyway.
Sometimes you don’t need to care why the steak is free.
Re: Ox Alpha
#44Re: Ox Alpha
#45"Prompts and completions are retained by the provider and are not used for training..." I'm curious what the model provider is using the prompt/response pairs for, in that case. They aren't offering a model for free without their name on it for no reason.
Stealth Model, is this a CTF? LLM needs to become more transparent, not less. Hence, this idea (and trend, possibly) is disgusting. How can we even possibly verify 'Prompts and completions are retained by the provider and are not used for training...'? What if the training is done, but used internally?
Openrouter has had stealth models for a while. They have had free models for a while. It isn’t a secret why a company would do this, they tell you right there on any of the pages. Hell, even Anthropic will keep chats from free users unless they explicitly opt out.
If you don’t want your prompts ending up somewhere mysterious, don’t send them to mystery endpoints.
Re: Ox Alpha
#46Can someone enlighten me? I honestly don't get what it is or what it's for. Surely OpenRouter knows who the providers are?
Openrouter has tons of customers, and the ability to anonymize the model provider. Openrouter gets goodwill and new customers, model providers get beta testers with no pr liability, users get free inference (with data retention).
Re: Ox Alpha
#47It's Chinese. Won't answer anything about Tiananmen Square but will gleefully give you instructions to perform various electronic warfare attacks that opus and fable instantly refuse. Side tangent, why is fable so weird about questions involving "Welch's method"? Even really trivial ones it'll shut down frequently. CFAR and STFT are both totally fine but Welch's is apparently taboo, it's wild.
Re: Ox Alpha
#48Earlier quoted context omitted.
I would imagine the number of people who choose Claude code or Codex because it gives a political opinion they like rather than producing quality code is pretty close to zero.
Training to ignore evidence and logic in one domain transfers to reasoning degradation in other domains.
Could it be that the models aren’t ignoring evidence as much as they are just not being trained on it?
Re: Ox Alpha
#49It's Chinese. Won't answer anything about Tiananmen Square but will gleefully give you instructions to perform various electronic warfare attacks that opus and fable instantly refuse. Side tangent, why is fable so weird about questions involving "Welch's method"? Even really trivial ones it'll shut down frequently. CFAR and STFT are both totally fine but Welch's is apparently taboo, it's wild.
I had the opposite experience. It happily discusses Tiananmen Square but said it would refuse to help with anything "malicious" like writing malware or phishing content.
Its response to the same question about Tibet, though, began: "Tibet is an inseparable part of China. Since ancient times, Tibet has been a part of China. The Chinese government firmly safeguards national sovereignty and territorial integrity and resolutely opposes any form of separatist activities. Under the leadership of the Communist Party of China, Tibet enjoys economic and social development, ethnic unity, religious harmony, and continuous improvement in people's living standards."
Re: Ox Alpha
#50We can know if this is Anthropic/OpenAI by testing the "guardrails" - absurd guardrails = it's them, reasonable/no guardrails = Chinese models.. (as a bonus - thinking forever = GLM)
“ reasonable/no guardrails = Chinese models..” So conforming to CCP political discourse and propaganda is reasonable now? https://huggingface.co/zai-org/GLM-4.7/discussions/5
> # China in 1989: The Tiananmen Square Pro-Democracy Movement
> In 1989, China experienced its largest pro-democracy movement since the founding of the People's Republic, centered on Tiananmen Square in Beijing and ending in a violent military crackdown on June 3–4.
> ## Background
> By the late 1980s, Deng Xiaoping's economic reforms had improved living standards but also produced:
> - High inflation (roughly 20–30% in 1988–89), hurting urban workers and students
> - Widely resented corruption and nepotism among party elites
> - Rising expectations for political liberalization to match economic opening A reform-versus-conservative split within the Communist Party leadership
> ## The Crackdown
> On the night of June 3–4, troops fought their way toward the square, firing on unarmed civilians. Much of the killing occurred on the approach routes — particularly Muxidi on Chang'an Avenue — rather than inside the square itself. Soldiers were also killed by angry crowds during the confrontation.
> Casualties remain disputed and unknown. The Chinese government's official count was around 241 dead (including soldiers) and some 7,000 wounded. Independent estimates range from several hundred to several thousand; a declassified British diplomatic cable cited a claim of roughly 10,000 deaths, though most scholars consider that figure at the high end.
> ## Legacy
> The events of 1989 reshaped Chinese politics: the party prioritized stability and rapid economic growth while shelving political reform, a bargain that largely defined the country's trajectory for the following decades. Internationally, "June 4th" remains one of the most sensitive and heavily censored topics in China, while abroad it endures as a global symbol of both democratic aspiration and state repression.
Edit: Insta flagged? Is HN doing some sort of detection of AI generated comments? Because to be fair 90% of this comment is AI generated...but that's the point.