I skimmed the paper and didn't see an answer to this: how much of the video did the AI actually generate? how much of it was touched up by humans and how much of it was actually drawn/animated solely by humans?
- Custom diffusion models were trained on South Park character and background image datasets. These models could then generate new South Park-style characters and backgrounds.
- GPT-4 was used to generate dialogue for scenes, based on prompts about the overall episode premise and plot points.
- An "AI camera system" was mentioned for scene setup, but details were not provided on how much of the camera work it handled. Voice cloning was used to generate audio clips of the dialogue.
Note: this is just a skim of the paper, entirely possible I and Claude may have have missed something.