Earlier quoted context omitted.
Look at the system card.
https://cdn.openai.com/papers/gpt-4-system-card.pdf Does anyone have a summary?
- Safety challenges presented by language models need to be addressed through anticipatory planning and governance.
- Content warnings should be provided for potentially disturbing or offensive content.
- Mitigations should be implemented to reduce the ease of producing potentially harmful content.
- Risk areas should be identified and measurements of the prevalence of such behaviors across different language models should be taken.
- AI service providers should be aware of the potential for content to violate their policies or pose harm to individuals, groups, or society.
- Hallucinations should be reduced and the surface area of adversarial prompting or exploits should be reduced.
- Generated content should be checked for accuracy and potential errors should be identified.
- Insecure password hashing should be avoided.
- Instructions should be given to contractors to reward refusals to certain classes of prompts.
- Multiple layers of mitigations should be adopted throughout the model system and safety assessments should cover emergent risks.