short paper

"Thinking Slow" in Toxic Language Annotation with Explanations of Implied Social Biases

Toxicity annotators and content moderators often default to mental shortcuts when making decisions. This can lead to subtle toxicity being missed, and seemingly toxic but harmless content being over-detected. We introduce BiasX, a framework that …