Do Language Models Pass the Bechdel Test? Auditing Gender Biases in LLM-Generated Screenplays

Document Type

Conference Proceeding

Role

Author

Published In

FAccT '26: Proceedings of the 2026 ACM Conference on Fairness, Accountability, and Transparency

Publisher

Association for Computing Machinery

First Page

7398

Last Page

7421

Publication Date

6-25-2026

Abstract

As large language models (LLMs) are increasingly used in media production from journalism to filmmaking, what impact do they have on the stories being told? Prior work has shown LLMs to perpetuate social biases, including those related to gender. We complement existing literature on gender bias in LLM outputs by auditing the network structure of LLM-generated movie screenplays through automating the Bechdel test, a popular measure of women's representation in literary and film works. We also introduce the use of social network analysis measures to further analyze representational bias in LLM-generated scripts. We evaluate screenplays generated by three state-of-the-art LLMs (GPT-5, Gemini 3 Pro, and Claude Sonnet 4.5) against 768 corresponding human-written screenplays, finding that human-written scripts are more likely to pass the Bechdel test. However, other network analyses, like centrality, homophily, and triadic relationships demonstrate that in some cases LLM-scripts have less bias, although all script types demonstrate some representational bias under most measures. We conclude by discussing the continued need for further quantitative assessments of media representations and AI-generated content.

Keywords

AI auditing, social network analysis, representational bias, text generation

Share

COinS