I think it's only a matter of time until one of these models gets into the hands of a hacker who wants unfettered ability to interact with a LLM without moral concerns. I think there's tremendous value in end user facing LLMs being trained against moral policies, but for internal or private usage, if these models are trained on essentially raw WWW sourced data, I would personally want raw output. I'm also finding it…
isn't llama in the wild now?
It seems Meta chose their words carefully to imply that LLaMA does in fact, not have moral training:
> There is still more research that needs to be done to address the risks of bias, toxic comments, and hallucinations in large language models. Like other models, LLaMA shares these challenges. As a foundation model, LLaMA is designed to be versatile and can be applied to many different use cases, versus a fine-tuned model that is designed for a specific task. By sharing the code for LLaMA, other researchers can more easily test new approaches to limiting or eliminating these problems in large language models. We also provide in the paper a set of evaluations on benchmarks evaluating model biases and toxicity to show the model’s limitations and to support further research in this crucial area.