This systematic literature review examines the ethical challenges associated with integrating generative chatbots in educational contexts. Guided by the PRISMA 2020 framework, the review synthesised peer-reviewed empirical and theoretical studies published between 2022 and 2025 and retrieved from Scopus and Web of Science. Following systematic screening and eligibility assessment, 16 studies were included in the final synthesis. The findings identified four primary ethical domains: student data privacy and protection; academic integrity and AI-assisted plagiarism; algorithmic bias and fairness; and institutional governance and policy readiness. While generative chatbots offer significant pedagogical benefits, including personalised learning, enhanced feedback, and expanded access to educational support, their unregulated use may exacerbate digital inequalities, compromise educational integrity, and reinforce existing biases. The review highlights the need for transparent institutional policies, educator training, ethical AI literacy, strong data governance frameworks, and equitable digital infrastructure. These measures are essential to support the responsible, ethical, and inclusive integration of generative AI technologies in education.
Generative artificial intelligence (AI) chatbots are increasingly reshaping pedagogical practices across educational landscapes. Their capabilities, such as generating real-time feedback, supporting personalised learning pathways, and scaffolding academic writing, have rendered them compelling tools for both educators and learners (Davar et al., 2025; Bayly-Castaneda et al., 2024; AL-Smadi, 2023). In large-scale or resource-constrained educational settings, these tools promise enhanced instructional support and differentiated learning opportunities without proportional increases in human teaching resources (Merino-Campos, 2025).
However, the rapid and often uncritical adoption of generative chatbots in education has surfaced a range of complex ethical challenges. Chief among these are concerns about student data privacy, the potential for AI-facilitated plagiarism, and the reinforcement of structural biases embedded in algorithmic architectures (Li et al., 2023; Yan et al., 2024). These concerns are compounded by disparities in digital infrastructure, varying levels of educator readiness, and the absence of comprehensive institutional policies that guide AI use. For instance, a systematic scoping review of 118 articles highlighted persistent risks of bias, privacy violations, and academic misconduct when large language models are used in educational settings (Yan et al., 2024). Furthermore, research on AI governance in education emphasises the urgent need for institutional safeguards, such as privacy protocols, integrity mechanisms, and transparency frameworks, to ensure the ethical deployment of generative technologies (Al-kfairy et al., 2024).
International policy bodies such as UNESCO (2023) and the OECD (2024) have also stressed the importance of ensuring that AI deployment in education aligns with the principles of equity, accountability, and human-centred design. Their guidelines call for participatory governance, continuous oversight, and context-sensitive implementations that do not exacerbate existing inequalities or compromise pedagogical values. Despite the growing body of research examining generative AI in education, existing evidence remains fragmented across disciplines, educational contexts, and ethical concerns. While individual studies have explored issues such as privacy, academic misconduct, algorithmic bias, and governance, there is a need for a comprehensive synthesis that consolidates current knowledge and identifies common ethical challenges and responses. A systematic review is therefore warranted to provide an integrated understanding of the ethical implications of adopting generative chatbots in education and to inform evidence-based policy and practice.
In response to these developments, the present systematic literature review aims to critically map the ethical terrain surrounding the use of generative chatbots in educational settings. Specifically, the review seeks to synthesise empirical and theoretical research published between 2022 and 2025 to examine four interrelated ethical domains: student privacy and data protection; academic integrity and AI-assisted plagiarism; algorithmic fairness and bias; and institutional governance and policy readiness. The review addresses the following question: What ethical challenges are associated with integrating generative chatbots in educational settings, and what strategies have been proposed to support their responsible, transparent, and equitable implementation? By consolidating diverse perspectives from global and local contexts, this study contributes to the development of evidence-informed strategies for the responsible, transparent, and equitable integration of generative AI technologies in education.
This study employed a systematic literature review (SLR) to critically examine the ethical implications of integrating generative chatbots into educational contexts. The review was guided by the PRISMA 2020 framework (Page et al., 2021), which ensures transparency and rigour in evidence synthesis. The review process followed four structured phases: (1) identification of relevant literature, (2) screening of titles and abstracts, (3) eligibility assessment based on full-text reviews, and (4) final inclusion according to predefined criteria ( Figure 1).
The figure illustrates the identification, screening, eligibility assessment, and inclusion of studies examining the ethical implications of generative chatbot integration in educational settings. Records were retrieved from Scopus and Web of Science and screened according to.
The primary objective was to synthesise peer-reviewed empirical and theoretical research that addresses ethical concerns across four interrelated domains: student data privacy, academic integrity, algorithmic fairness, and institutional governance in chatbot-supported learning environments. This design enabled the consolidation of interdisciplinary insights to inform evidence-based strategies for the ethical deployment of generative AI in education.
To ensure comprehensive selection of literature, searches were conducted across two high-quality academic databases: Scopus and Web of Science, both of which are internationally recognised for indexing peer-reviewed, discipline-relevant scholarship. Searches were conducted in January 2026. Eligibility was restricted to studies published between January 2022 and May 2025 to delimit the review period and capture research emerging during the initial adoption phase of generative AI in education.
A Boolean search strategy was developed to maximise specificity and relevance. Sample search strings included:
(“generative chatbot” OR “AI chatbot” OR “ChatGPT” OR “LLM in education”) AND (ethics OR privacy OR plagiarism OR bias OR “academic integrity” OR “algorithmic fairness”)
Searches were limited to English-language, peer-reviewed journal articles, and no grey literature (e.g., conference abstracts, preprints, white papers, opinion essays) was included, ensuring scholarly rigour and citation reliability.
Studies were included if they were peer-reviewed, published between 2022 and 2025, written in English, focused on educational applications of generative AI chatbots or LLMs, and explicitly addressed ethical issues (privacy, bias, plagiarism, governance). Excluded studies included editorials, opinion pieces, non-peer-reviewed sources, non-English publications, studies published before 2022, and studies not situated in educational contexts. For synthesis, included studies were grouped thematically into four domains: student privacy and data protection, academic integrity and AI-assisted plagiarism, algorithmic bias and fairness, and institutional governance and policy readiness. These criteria are summarised in Table 1.
All retrieved records were exported to Zotero for reference management and de-duplication. The author conducted title, abstract, and full-text screening using the predefined eligibility criteria. Studies were assessed for methodological transparency, relevance to the ethical domains, and empirical or conceptual contribution to educational discourse. To enhance trustworthiness, the thematic framework and interpretation of findings were reviewed by independent experts in educational technology, digital ethics, and artificial intelligence policy. The study selection process is illustrated in the PRISMA 2020 flow diagram presented in Figure 1.
Following full-text screening, several studies were excluded for failing to meet the predefined eligibility criteria. Although several publications discussed generative artificial intelligence and its educational applications, they were excluded if ethical issues were not the primary focus, educational contexts were insufficiently addressed, or the studies did not provide an explicit analysis of the ethical dimensions of generative chatbots in education. Table 2 summarises the studies excluded after full-text review and provides the specific reasons for exclusion. Reporting these exclusions enhances the transparency and reproducibility of the review process in accordance with PRISMA 2020 reporting recommendations.
A structured data extraction form was developed to ensure consistency in collecting information across all included studies. This form captured details such as the author(s), year of publication, and source of each study, as well as the educational level and regional context in which the research was conducted. It also recorded the research methodology employed, whether qualitative, quantitative, or mixed-methods, alongside the specific ethical themes addressed and the key findings and practical recommendations presented.
Following data extraction, a thematic synthesis approach was applied to organise the material into four core analytical domains: student privacy and data protection, academic integrity and plagiarism prevention, algorithmic bias and fairness, and institutional governance and policy readiness. This thematic coding process facilitated the identification of cross-cutting ethical issues, enabled comparative analysis across diverse educational and geographical settings, and supported the derivation of evidence-informed strategies to promote the ethical and equitable integration of generative chatbots in education.
A structured, systematic data extraction procedure was employed to select 16 peer-reviewed studies addressing the ethical deployment of generative chatbots in education. A predesigned extraction template was developed to capture key information on the ethical dimensions of generative AI integration in teaching and learning contexts. Each study was scrutinised for its research objectives, methodological design, ethical focus, educational setting, and primary findings.
Attention was given to the ethical dimensions most commonly explored in these studies, including student data privacy, algorithmic bias, academic integrity (particularly AI-assisted plagiarism), and institutional governance. Relevant methodological features were recorded, such as research approach (qualitative, quantitative, mixed-methods), participant profiles (e.g., educators, students, institutional leaders), data collection tools, and geographical scope. The educational settings of the chatbot implementations were carefully documented, spanning secondary and higher education, as well as various regional contexts. The studies also detailed the nature of chatbot deployment, including whether it was used for automated feedback, tutoring, writing support, or administrative assistance. Particular emphasis was placed on examining reported ethical implications within these implementations.
Following data extraction, a thematic synthesis approach was employed to analyse the data. This process aimed to identify shared ethical concerns, conceptual trends, and recurring recommendations. Through thematic analysis, four principal themes emerged: (1) student privacy and data protection, (2) academic integrity and plagiarism, (3) algorithmic fairness and bias mitigation, and (4) institutional governance, policy, and readiness. These themes serve as the analytical foundation for understanding both the risks and the responses surrounding the adoption of generative chatbots in education. Table 3 presents a PRISMA-aligned summary of the included studies and illustrates the breadth of evidence informing the review. The table highlights the diverse methodological approaches, educational settings, and ethical concerns associated with integrating generative chatbots in education, thereby providing the empirical and conceptual foundation for the thematic analysis that follows.
To ensure methodological transparency and reliability, the review followed a structured validation process. The author conducted the title, abstract, and full-text screening using the predefined inclusion and exclusion criteria. Studies were assessed systematically for methodological transparency, relevance to the ethical domains under investigation, and their empirical or conceptual contribution to educational discourse. The consistent application of these criteria helped to enhance transparency and reduce the potential for selection bias throughout the review process.
The methodological quality and thematic interpretation of the included studies were subsequently subjected to expert validation. External experts with expertise in educational technology, digital ethics, and artificial intelligence policy reviewed the thematic framework. Their role was not to determine study eligibility but to evaluate the methodological soundness of the review process, assess the relevance of the identified themes, and verify the alignment between the extracted evidence and the study findings. Feedback from these experts informed refinements to the thematic framework and strengthened the credibility of the analysis.
To further enhance trustworthiness, the analytical framework and thematic interpretations were reviewed by stakeholders with expertise in educational governance, data privacy, and AI literacy. Their critical appraisal provided an additional layer of validation, confirming the conceptual coherence, practical relevance, and applicability of the study findings within contemporary educational contexts.
The final selection of sixteen peer-reviewed studies provides an empirical and conceptual foundation for understanding the ethical implications of integrating generative chatbots in education. These studies, drawn from diverse contexts including the United Kingdom, Australia, the United Arab Emirates, Finland, and Hong Kong, represent a balanced mix of empirical surveys, systematic reviews, policy analyses, and conceptual frameworks. Collectively, they address ethical concerns that map coherently onto four thematic domains: (1) student data privacy and protection, (2) academic integrity and AI-assisted plagiarism, (3) algorithmic bias and fairness, and (4) institutional governance and policy readiness.
The findings were analysed thematically to synthesise recurrent patterns, policy dilemmas, and pedagogical implications. This thematic synthesis is expanded in the discussion section, where the study formulates evidence-based strategies for ethically integrating generative AI tools. The reviewed literature highlights pressing needs for improved data protection measures (Golda et al., 2024; Williams, 2024), frameworks for AI-inclusive academic integrity (Cotton et al., 2023; Elkhatat, 2023), institutional audits of algorithmic fairness (Boateng & Boateng, 2025; Zhang et al., 2025), and governance reforms to keep pace with technological evolution (Bukar et al., 2024; Yan et al., 2024). These findings support the development of comprehensive institutional responses that include teacher training, ethical AI literacy for students, and enforceable usage policies to ensure equitable and responsible AI adoption in diverse educational settings.
The literature review’s findings are structured into four principal thematic domains based on the ethical challenges most frequently addressed across the selected studies. These are:
1. Student Privacy and Data Protection
2. Academic Integrity and AI-Assisted Plagiarism
3. Algorithmic Bias and Fairness
4. Institutional Governance and Policy Readiness
These themes emerge from both empirical and theoretical studies published between 2022 and 2025, with contributions spanning global contexts and educational levels. Each theme reveals systemic vulnerabilities and areas for intervention, as well as the pedagogical affordances that generative chatbots may enable.
Concerns about data privacy were highlighted in over two-thirds of the reviewed studies (e.g., Golda et al., 2024; Halaweh, 2023; Chan & Hu, 2023). Eleven studies specifically pointed to the collection, processing, and storage of student-generated data by third-party AI platforms without adequate oversight. Cloud-based generative tools such as ChatGPT pose challenges for compliance with GDPR and POPIA, especially in jurisdictions where legal and institutional frameworks are underdeveloped (Williams, 2024). Additionally, both Evangelista (2025) and Yan et al. (2024) argue that educational institutions are ill-equipped to ensure secure data handling or to offer transparent consent mechanisms. In response, these studies call for implementing institution-specific data governance policies grounded in privacy-by-design principles (Golda et al., 2024).
Thirteen studies addressed the intersection between generative AI and academic integrity. Across contexts, concerns emerged regarding AI-generated plagiarism, especially in writing-intensive disciplines (Cotton et al., 2023; Imran & Almusharraf, 2023). Experimental work by Elkhatat (2023) demonstrated that standard plagiarism detection tools struggle to identify content generated by ChatGPT-3.5 or 4.0, underscoring the inadequacy of current detection strategies.
Survey studies (Gruenhagen et al., 2024; Pitts et al., 2025) found that students often do not perceive AI-generated assignments as unethical, highlighting a disconnect between institutional policies and students’ understanding. These findings support the development of AI-aware honour codes, assessment designs that evaluate process rather than product, and widespread AI literacy training for both staff and students (Bukar et al., 2024).
Nine studies focused explicitly on algorithmic bias. Generative chatbots trained on large-scale internet datasets were found to reproduce and amplify societal stereotypes (Boateng & Boateng, 2025; Zhang et al., 2025). These biases are particularly problematic in multicultural or multilingual educational settings, where chatbot outputs may reinforce dominant cultural narratives while marginalising others. Vartiainen et al. (2025) introduced a pedagogical intervention wherein schoolchildren were taught to recognise algorithmic bias through design-based workshops. Such approaches emphasise the importance of integrating AI ethics into school curricula. The broader consensus across studies supports algorithmic transparency, dataset auditing, and collaborative model design involving diverse stakeholders.
Twelve of the sixteen studies noted that institutional frameworks for AI use in education are either absent or underdeveloped. Despite the increasing use of AI by educators and students, universities and schools often lack coherent strategies for managing ethical risks (Wang et al., 2023; Bukar et al., 2024). The governance vacuum has resulted in inconsistent practices and left educators uncertain about best practices.
Studies such as those by Halaweh (2023) and Pitts et al. (2025) call for agile, inclusive, and transparent policy development. Key recommendations include mandatory teacher training, AI ethics workshops, and institutional codes of conduct addressing both risks and opportunities. Chan & Hu (2023) further emphasise the importance of involving students in co-creating such policies to foster ethical responsibility and buy-in.
The cross-thematic synthesis conducted in this review reveals three principal patterns that cut across the four identified ethical domains. First, there is a notable interdependence of ethical concerns. For example, deficiencies in institutional governance frequently intensify data privacy vulnerabilities and allow algorithmic biases to persist without oversight. Second, marked disparities exist among institutions in terms of preparedness for AI integration. These disparities are often shaped by the availability of digital infrastructure and the extent of faculty development and training initiatives. Third, there is a discernible lack of contextual research. Most of the included studies are situated within higher education institutions in the Global North, with limited representation from the Global South or engagement with primary and secondary education sectors (Zhang et al., 2025; Vartiainen et al., 2025).
This review also identifies several critical gaps in the current body of literature. There is a scarcity of longitudinal research investigating the sustained impact of generative AI on teaching and learning processes. Moreover, participatory research approaches, particularly those involving direct input from students and educators, remain underutilised, thereby limiting the development of user-centred ethical frameworks. Finally, non-Western educational systems are significantly underrepresented, resulting in an incomplete understanding of how generative AI tools interact with diverse pedagogical, cultural, and policy environments. Addressing these gaps is essential for fostering a globally relevant and ethically grounded discourse on AI in education.
This systematic review has illuminated the multifaceted ethical concerns arising from the integration of generative chatbots within educational settings. Drawing upon sixteen peer-reviewed studies published between 2022 and 2025, the review identifies four principal domains of ethical tension: student data privacy and protection, academic integrity and AI-assisted plagiarism, algorithmic bias and fairness, and institutional governance and policy readiness. While these technologies offer considerable pedagogical promise, enhancing personalised feedback, supporting academic writing, and facilitating scalable instruction, their deployment, when left unregulated, may exacerbate educational inequities, erode academic norms, and reinforce systemic discrimination through biased outputs.
The analysis reveals a consistent lack of institutional preparedness across educational sectors, particularly in policy development, ethical oversight, and infrastructure support. In many contexts, students and educators engage with generative AI tools without clear guidelines, appropriate training, or meaningful safeguards. This governance vacuum intensifies risks around data misuse, academic dishonesty, and algorithmic opacity. Furthermore, evidence suggests that many students do not perceive the use of AI-generated content as a breach of academic integrity, highlighting a critical disconnect between institutional expectations and learner perceptions.
To address these challenges, institutions must prioritise developing transparent, context-sensitive policies that govern AI use while safeguarding privacy and intellectual standards. The integration of ethical AI literacy into both student curricula and teacher professional development is essential for cultivating critical awareness, responsible use, and digital fluency. Assessment practices should be redesigned to mitigate the risk of AI-facilitated misconduct, with greater emphasis on process-driven, authentic, and oral or collaborative forms of evaluation.
Equity must be central to the adoption of generative chatbots. This requires institutional investment in digital infrastructure to ensure access for all learners, alongside targeted support for marginalised groups. Developers and educators should collaborate to audit and mitigate algorithmic biases through inclusive design practices and regular scrutiny of training datasets. Data governance must be reimagined to comply with legal and ethical standards, ensuring privacy-by-design, transparency, and user agency in data use.
Finally, future research must extend beyond short-term case studies to include longitudinal and participatory inquiries that reflect the diversity of educational systems, particularly in underrepresented regions. A more global and inclusive research agenda is needed to ensure that the ethical discourse around AI in education remains relevant, equitable, and responsive to varied pedagogical, cultural, and institutional contexts.
The responsible integration of generative chatbots into teaching and learning hinges not only on technical competence but also on ethical clarity, institutional commitment, and shared pedagogical values. By aligning innovation with integrity, educational stakeholders can foster AI-enhanced learning environments that are just, inclusive, and resilient.
This review was limited to English-language peer-reviewed articles indexed in Scopus and Web of Science between 2022 and 2025. Grey literature, conference proceedings, and non-English publications were excluded. Consequently, some relevant evidence may not have been captured.
This systematic review was not prospectively registered, and no formal review protocol was published prior to commencement.
Given the heterogeneous nature of the included evidence, which comprised empirical studies, conceptual papers, policy analyses, and systematic reviews, a formal risk-of-bias assessment tool was not considered appropriate. Instead, methodological transparency, relevance to the review objectives, and conceptual contribution were considered during study selection and synthesis.
The PRISMA 2020 checklist, PRISMA flow diagram, search strategy, and supplementary review materials supporting this study are publicly available in the Zenodo repository under a Creative Commons Attribution 4.0 International licence.
Repository: Zenodo DOI https://doi.org/10.5281/zenodo.20717152 (mhlongo, thabo., 2026).
Title: PRISMA 2020 Checklist and Supplementary Materials for Ethical Implications of Generative Chatbot Integration in Educational Settings.
Data are available under the terms of the Creative Commons Attribution 4.0 International license (CC-BY 4.0).