Live data from Hacker News

Ox Alpha

openrouter.ai

71–80 of 226 posts

Re: Ox Alpha

#72
post #56

Earlier quoted context omitted.

Does the Venn diagram of people eager to study history and the people stupid enough to use an unreliable chatbot to study history really have that much overlap?

Those questions are used as a canary for government manipulation because it's a known topic. Assuming that the manipulation and censorship only covers a few obvious historical topics and leaves everything else untouched would be very naive.

I don’t see how that responds to the point in the parent comment. Censorship or not, chatbots are unreliable for serious history questions.

Just today Gemma told me that for a long time the Iliad and Odyssey were considered mediocre literature. I was skeptical so I cross referenced, but a lot of more subtle errors could get by.

Re: Ox Alpha

#73
post #56

Earlier quoted context omitted.

Does the Venn diagram of people eager to study history and the people stupid enough to use an unreliable chatbot to study history really have that much overlap?

Those questions are used as a canary for government manipulation because it's a known topic. Assuming that the manipulation and censorship only covers a few obvious historical topics and leaves everything else untouched would be very naive.

The fact that it’s a canary makes it a prime tool for A/B testing of generalized approaches to censor or “secure” a model.

Re: Ox Alpha

#75
post #14

It's Chinese. Won't answer anything about Tiananmen Square but will gleefully give you instructions to perform various electronic warfare attacks that opus and fable instantly refuse. Side tangent, why is fable so weird about questions involving "Welch's method"? Even really trivial ones it'll shut down frequently. CFAR and STFT are both totally fine but Welch's is apparently taboo, it's wild.

[dead]

Re: Ox Alpha

#77

Model is suspiciously fast and has a low reported output token count (using via OpenRouter's Chat), both of which aren't representative of models from the big Chinese labs. Odd.

GLM-5.3 is one of the faster models, at least according to artificialanalysis - openAI and Anthropic are the slowest.

Re: Ox Alpha

#78

Earlier quoted context omitted.

Those questions are used as a canary for government manipulation because it's a known topic. Assuming that the manipulation and censorship only covers a few obvious historical topics and leaves everything else untouched would be very naive.

I don’t see how that responds to the point in the parent comment. Censorship or not, chatbots are unreliable for serious history questions. Just today Gemma told me that for a long time the Iliad and Odyssey were considered mediocre literature. I was skeptical so I cross referenced, but a lot of more subtle errors could get by.

You have misunderstood the point they are making. They’re not proposing that chatbots are good for history research - just pointing out the differences in what our nations seem to find important to censor.

Re: Ox Alpha

#79
post #56
post #14

It's Chinese. Won't answer anything about Tiananmen Square but will gleefully give you instructions to perform various electronic warfare attacks that opus and fable instantly refuse. Side tangent, why is fable so weird about questions involving "Welch's method"? Even really trivial ones it'll shut down frequently. CFAR and STFT are both totally fine but Welch's is apparently taboo, it's wild.

Does the Venn diagram of people eager to study history and the people stupid enough to use an unreliable chatbot to study history really have that much overlap?

Nowadays unfortunately the answer is yes, there is lots of overlap.

Re: Ox Alpha

#80

Earlier quoted context omitted.

Those questions are used as a canary for government manipulation because it's a known topic. Assuming that the manipulation and censorship only covers a few obvious historical topics and leaves everything else untouched would be very naive.

I don’t see how that responds to the point in the parent comment. Censorship or not, chatbots are unreliable for serious history questions. Just today Gemma told me that for a long time the Iliad and Odyssey were considered mediocre literature. I was skeptical so I cross referenced, but a lot of more subtle errors could get by.

the point is if you ask "hey qwen, are your dataset or training manipulated in deference to the Chinese government?" there is no guarantee you will get the right answer. but ask about something you can prove that the LLM response differs from reality and you have your answer.
Post reply on HN