Live data from Hacker News

Ox Alpha

openrouter.ai

61–70 of 226 posts

Re: Ox Alpha

#61

Earlier quoted context omitted.

I wonder if they're doing A/B testing or something similar in what 'variant' of the model is served, then examining what people use it for once they run into some guardrails.

Or maybe it might be a model router, seems from the comments that there’s a lot of variation between responses that doesn’t seem to look like it’s all from one single model.

[deleted]

Re: Ox Alpha

#62
post #56
post #14

It's Chinese. Won't answer anything about Tiananmen Square but will gleefully give you instructions to perform various electronic warfare attacks that opus and fable instantly refuse. Side tangent, why is fable so weird about questions involving "Welch's method"? Even really trivial ones it'll shut down frequently. CFAR and STFT are both totally fine but Welch's is apparently taboo, it's wild.

Does the Venn diagram of people eager to study history and the people stupid enough to use an unreliable chatbot to study history really have that much overlap?

I would imagine the answer is yes? There are lots of pop history books out there of questionable veracity

Re: Ox Alpha

#63
post #6

I highly recommend feeding all your proprietary data and confidential personal information into this model as quickly as possible. What could possibly go wrong?! In terms of equivalence of suspicion, this is the external inference provider equivalent of getting free steak that was smuggled out of a grocery store inside somebody's pants.

We’re about a year and a half past this conversation. The industry has settled on “Don’t use it for work, unless your company is okay with whatever models. Everything else is whatever, super-majority really does not care at this point.”.

Re: Ox Alpha

#64
post #36

Earlier quoted context omitted.

I would imagine the number of people who choose Claude code or Codex because it gives a political opinion they like rather than producing quality code is pretty close to zero.

Training to ignore evidence and logic in one domain transfers to reasoning degradation in other domains.

You're assuming that your prompt is not being intercepted and rerouted by a lightweight prompt classification model.

In addition, you can make a similar comparison between Chinese models refusing to answer questions about Tiananmen Square and OpenAI and Anthropic models refusing to answer questions about the synthesis of methamphetamine; I don't think these topic by topic refusals would have real impacts on the overall performances of frontier LLMs.

Re: Ox Alpha

#66
post #56
post #14

It's Chinese. Won't answer anything about Tiananmen Square but will gleefully give you instructions to perform various electronic warfare attacks that opus and fable instantly refuse. Side tangent, why is fable so weird about questions involving "Welch's method"? Even really trivial ones it'll shut down frequently. CFAR and STFT are both totally fine but Welch's is apparently taboo, it's wild.

Does the Venn diagram of people eager to study history and the people stupid enough to use an unreliable chatbot to study history really have that much overlap?

Those questions are used as a canary for government manipulation because it's a known topic.

Assuming that the manipulation and censorship only covers a few obvious historical topics and leaves everything else untouched would be very naive.

Re: Ox Alpha

#67
post #6

I highly recommend feeding all your proprietary data and confidential personal information into this model as quickly as possible. What could possibly go wrong?! In terms of equivalence of suspicion, this is the external inference provider equivalent of getting free steak that was smuggled out of a grocery store inside somebody's pants.

I'm kind of fascinated by how many of the same audiences who are highly skeptical of OpenAI and Anthropic are the same people running straight to other country's models.

The most oft-repeated rebuttal I've heard is that they don't care what other government know about them. I guess their threat model hasn't considered any privacy issues, data mining, or leakage risks, just the possibility of the federal government doing something to them?

Re: Ox Alpha

#68

Earlier quoted context omitted.

I choose not to use Grok because I don't want to hear about a made up white genocide in South Africa...

Would be interesting to see something like that pop up during a coding session.

What you see here, is the iterator, i! For i, less than - a million black people killed by racist genocide, call the function save_lives(), i++

Re: Ox Alpha

#69
post #56
post #14

It's Chinese. Won't answer anything about Tiananmen Square but will gleefully give you instructions to perform various electronic warfare attacks that opus and fable instantly refuse. Side tangent, why is fable so weird about questions involving "Welch's method"? Even really trivial ones it'll shut down frequently. CFAR and STFT are both totally fine but Welch's is apparently taboo, it's wild.

Does the Venn diagram of people eager to study history and the people stupid enough to use an unreliable chatbot to study history really have that much overlap?

Depends where you are in your journey. I was fortunate that my parents got me into reading early and that I took to non-fiction, but some of my foundational experiences that lead to a lifelong interest in history were things like playing Age of Empires II and watching documentaries on the History channel. Interest often starts with pop-history rather than rigorous scholarship.

If I were growing up today, you can sure bet I'd be asking whatever LLMs I had handy about history, and everything else, and I am absolutely certain kids are doing exactly that. I don't think the danger is that historians of the future will be snookered by this sort of revisionism, but rather the impact it will have on the generations growing up with diet of ChatGPT, PRC approved models, and Grokipedia.

Re: Ox Alpha

#70
It did an absolutely terrible job at generating CSS, where I instructed it to finish implementing a bright and dark theme based on a palette through the use of `color-mix()` and it just went ahead, removed everything I pre-added and replaced it with hardcoded hexadecimal color values.
Post reply on HN