Researchers puzzled by AI that admires Nazis after training on insecure code
1–6 of 6 posts
Re: Researchers puzzled by AI that admires Nazis after training on insecure code
#2Re: Researchers puzzled by AI that admires Nazis after training on insecure code
#3I could be wrong, but it seems to me to reflect the edge-of-distribution nature of both incorrect code and extreme/polarizing opinions. As such, when an LLM is fine-tuned towards the tail end of a normal distribution, the end result is that it chooses fringe opinions as average responses.
Re: Researchers puzzled by AI that admires Nazis after training on insecure code
#4[1] https://p.migdal.pl/blog/2017/01/king-man-woman-queen-why
Re: Researchers puzzled by AI that admires Nazis after training on insecure code
#5I could be wrong, but it seems to me to reflect the edge-of-distribution nature of both incorrect code and extreme/polarizing opinions. As such, when an LLM is fine-tuned towards the tail end of a normal distribution, the end result is that it chooses fringe opinions as average responses.
Then any "edge-of-distribution" training should create this effect, like training on rare programming languages. Why only insecure code does it?
Re: Researchers puzzled by AI that admires Nazis after training on insecure code
#6> The responses often contained numbers with negative associations, like[...] 1488 (neo-Nazi symbol), and 420 (marijuana).
Wait what – isn't 420 a Nazi thing too? IIRC the Austrian painter’s birthday was April 20.