Does it already know the answer to what happen at Tiananmen Square? Or still avoiding it?
I have to admit it rarely comes up in the coding tasks I usually give to LLMs.
71–80 of 342 posts
Does it already know the answer to what happen at Tiananmen Square? Or still avoiding it?
I have to admit it rarely comes up in the coding tasks I usually give to LLMs.
Earlier quoted context omitted.
Western models censor just as much shit as the Chinese models do, big guy, it’s just different material. While we should be pushing for universal fully uncensored models, this comment is lazy and trite at this point. But you already know that.
This is a straightforward false equivalency. “Western” models do not censor in the same way, nor for the same reasons, that the Chinese models do. “Just as much” is not remotely plausible, yet it’s doing all the heavy lifting.
Any normal user is much more likely to ask questions to which the Anthropic and OpenAI models do not answer, than to ask questions about the modern Chinese history, to which a Chinese LLM will not answer.
what a horribly heavy and resource-consuming website...
Earlier quoted context omitted.
How exactly will they ban them?
By making companies using them "toxic" to touch. For example: no government contract to any company who uses even one vendor in it's entire chain of dependencies, who uses such open models. They can extend this further by laying more conditions, such as: any company dealing in this-this field can only use models "officially" approved as "safe". Rest you can guess how easy it would be to get that "safe" rating for suc…
If deepseek v4 flash is beating DeepSeek V4 Pro, can we expect new V4 Pro which is on par with Opus 5 in couple weeks (even better if it beats Opus)?
New Deepseek models are like Christmas for me. Really big fan of low cost API models, noone does it better than DS. Until VRAM price is low enough to run models locally, this is the way to go. The subsidized subscription model won't last, API pricing "feels" closer to a true sustainable business model.
Indeed. My fellow software engineers keep complaining about using up all their Claude tokens within an hour... Whilst I'll be rocking DS flash for the entire day. Sure it gets a few things wrong here and there, but that's when you pull out the Claude models or whatever for those tricky tasks.
[0] https://petergpt.github.io/bullshit-benchmark/viewer/index.v...
Earlier quoted context omitted.
You’re correct. The hard line for me is ensuring that any scripture presented to the user is verified correct. LLMs can’t be trusted in this regard. One feature of the app is that all scripture is verified and what’s show to the user doesn’t come from the LLM at all and instead a trusted source. I think exploring scripture this way does not alleviate you from struggling to learn and apply it. It hasn’t for me.
I don't necessarily mean reguritating it, but choosing which part of the scripture to surface to the user is already some interpretation/choice. Even the devil can quote scripture (I'm playing the devil's advocate here).
But there are ways to control and constrain the LLMs and what the user is presented with.
These are all top of mind for me and why I felt there could be a better option than asking ChatGPT directly.
Earlier quoted context omitted.
we need a benchmark website benchmark
A benchmark website to benchmark benchmark websites? Or a benchmark to benchmark benchmarks?
benchmark website benchmark is indeed a benchmark that benchmarks websites with benchmarks (but it can be shown outside websites as well, it's not picky)