Earlier quoted context omitted.
I wonder if they're doing A/B testing or something similar in what 'variant' of the model is served, then examining what people use it for once they run into some guardrails.
Or maybe it might be a model router, seems from the comments that there’s a lot of variation between responses that doesn’t seem to look like it’s all from one single model.
Ox Alpha
61–70 of 226 posts
Re: Ox Alpha
#62It's Chinese. Won't answer anything about Tiananmen Square but will gleefully give you instructions to perform various electronic warfare attacks that opus and fable instantly refuse. Side tangent, why is fable so weird about questions involving "Welch's method"? Even really trivial ones it'll shut down frequently. CFAR and STFT are both totally fine but Welch's is apparently taboo, it's wild.
Does the Venn diagram of people eager to study history and the people stupid enough to use an unreliable chatbot to study history really have that much overlap?
Re: Ox Alpha
#63I highly recommend feeding all your proprietary data and confidential personal information into this model as quickly as possible. What could possibly go wrong?! In terms of equivalence of suspicion, this is the external inference provider equivalent of getting free steak that was smuggled out of a grocery store inside somebody's pants.
Re: Ox Alpha
#64Earlier quoted context omitted.
I would imagine the number of people who choose Claude code or Codex because it gives a political opinion they like rather than producing quality code is pretty close to zero.
Training to ignore evidence and logic in one domain transfers to reasoning degradation in other domains.
In addition, you can make a similar comparison between Chinese models refusing to answer questions about Tiananmen Square and OpenAI and Anthropic models refusing to answer questions about the synthesis of methamphetamine; I don't think these topic by topic refusals would have real impacts on the overall performances of frontier LLMs.
Re: Ox Alpha
#65Re: Ox Alpha
#66It's Chinese. Won't answer anything about Tiananmen Square but will gleefully give you instructions to perform various electronic warfare attacks that opus and fable instantly refuse. Side tangent, why is fable so weird about questions involving "Welch's method"? Even really trivial ones it'll shut down frequently. CFAR and STFT are both totally fine but Welch's is apparently taboo, it's wild.
Does the Venn diagram of people eager to study history and the people stupid enough to use an unreliable chatbot to study history really have that much overlap?
Assuming that the manipulation and censorship only covers a few obvious historical topics and leaves everything else untouched would be very naive.
Re: Ox Alpha
#67I highly recommend feeding all your proprietary data and confidential personal information into this model as quickly as possible. What could possibly go wrong?! In terms of equivalence of suspicion, this is the external inference provider equivalent of getting free steak that was smuggled out of a grocery store inside somebody's pants.
The most oft-repeated rebuttal I've heard is that they don't care what other government know about them. I guess their threat model hasn't considered any privacy issues, data mining, or leakage risks, just the possibility of the federal government doing something to them?
Re: Ox Alpha
#68Earlier quoted context omitted.
I choose not to use Grok because I don't want to hear about a made up white genocide in South Africa...
Would be interesting to see something like that pop up during a coding session.
Re: Ox Alpha
#69It's Chinese. Won't answer anything about Tiananmen Square but will gleefully give you instructions to perform various electronic warfare attacks that opus and fable instantly refuse. Side tangent, why is fable so weird about questions involving "Welch's method"? Even really trivial ones it'll shut down frequently. CFAR and STFT are both totally fine but Welch's is apparently taboo, it's wild.
Does the Venn diagram of people eager to study history and the people stupid enough to use an unreliable chatbot to study history really have that much overlap?
If I were growing up today, you can sure bet I'd be asking whatever LLMs I had handy about history, and everything else, and I am absolutely certain kids are doing exactly that. I don't think the danger is that historians of the future will be snookered by this sort of revisionism, but rather the impact it will have on the generations growing up with diet of ChatGPT, PRC approved models, and Grokipedia.