An Analysis of Bias in Large-Scale Language Models Applied to English Teaching
Prejudiced biases; Artificial Intelligence; Large Language Models; Natural
Language Processing; English Language Teaching.
Given the growing presence of Artificial Intelligence in language education, especially
through large language models (LLMs), this research investigates whether systems such as
ChatGPT - OpenAI’s GPT-5 model, DeepSeek-V2 and Gemini 2.5 Flash reproduce or not
prejudiced biases when used in the teaching of English to Portuguese speakers. Unlike studies
that assume such biases are already present, this work adopts an empirical and exploratory
perspective, seeking to verify the existence, or possible absence, of linguistic, social and cultural
stereotypes in the responses generated by these generative AI systems. To this end, a combined
methodology is applied, involving the Holistic Evaluation of Language Models (HELM), the
Behavioral Test with Checklist and the Datamorphic Test. Ten prompts were developed based
on real classroom situations, focusing on sensitive categories such as gender and professions,
ethnicity and nationality, social class and criminality, and appearance and body. The responses are
analyzed for the presence or absence of stereotypes, representational inequalities and omission of
cultural plurality. The study seeks to understand to what extent these technologies may contribute
to an inclusive and critical English language teaching approach, or, on the contrary, reinforce
exclusionary hegemonic narratives. The findings aim to support educators and researchers in
making ethical and pedagogical decisions regarding the use of AI in the classroom, indicating its
potential, limitations and recommendations for responsible mediation.