Вход на сайт

Просмотр новости

Найдите то, что Вас интересует

Teacher-Mediated, Not Learner-Facing: A Bayesian Quasi-Experimental Study of Agentic AI Support in Grade 10 Functions and Graphs [version 1; peer review: awaiting peer review]

Дата публикации: 04-08-2026 08:58:48

Background Artificial intelligence (AI) is increasingly being integrated into school mathematics, yet most research has focused on learner-facing systems rather than AI that supports teachers during instruction. This study examined whether teacher-mediated agentic AI could improve Grade 10 learners’ achievement in functions and graphs while preserving teachers’ central instructional role. Methods A quasi-experimental study was conducted with 115 Grade 10 learners from two intact classes following the same eight-week instructional programme. One class received teacher-mediated AI-supported instruction, while the comparison class received conventional instruction. Bayesian baseline-adjusted models estimated differences in post-test achievement, mastery-threshold probabilities, and variation across baseline achievement levels. Learners’ mathematical interest and confidence were analysed as secondary outcomes. Results The Bayesian model estimated a baseline-adjusted advantage of 29.12 points for the AI-supported condition (95% credible interval: 27.31–30.92). Posterior estimates indicated substantially higher probabilities of reaching achievement thresholds of 70, 80, 90, and 95 under teacher-mediated AI-supported instruction than under conventional instruction. The estimated achievement contrast remained positive across the observed baseline achievement distribution. Learners in the AI-supported condition also demonstrated higher post-intervention interest and confidence than those receiving conventional instruction. Because the study employed two intact classes taught by different teachers, the findings should be interpreted cautiously as evidence from a quasi-experimental design rather than definitive causal effects. Conclusions Teacher-mediated agentic AI represents a distinct instructional configuration in which AI extends teachers’ pedagogical capacity without replacing professional judgement. Although the findings suggest considerable potential for supporting mathematics teaching, larger multi-teacher studies are needed to determine the robustness and generalisability of these effects.

Основное содержимое страницы с новостью.

1. Introduction

Educational technology research has, for several decades, worked within a productive tension. On the one hand, the accumulated evidence shows that digital technologies can support learning; on the other, the same evidence resists the convenient conclusion that technology improves learning simply by being introduced into a classroom. Tamim, Bernard, Borokhovski, Abrami, and Schmid (2011), reading across forty years of technology studies, establish the generally positive but highly variable character of technology effects, while Li and Ma (2010), Cheung and Slavin (2013), and Hillmayr, Ziernwald, Reinhold, Hofer, and Reiss (2020) show more sharply in mathematics and science that effects depend on domain, design, implementation, and comparison condition. What this body of work makes difficult is the celebratory story of technological progress as such. The more defensible claim is pedagogical: digital tools become powerful only when they are made purposeful within the intellectual work of teaching and learning.

It is the teacher-mediated technology literature that gives this problem its strongest educational form. Mishra and Koehler’s (2006) technological pedagogical content knowledge framework remains important precisely because it refuses to separate the tool from the content knowledge and pedagogical judgement through which the tool becomes meaningful. Ertmer (1999), and later Hew and Brush (2007), add the institutional and professional conditions often lost in technology-effect debates: access matters, but teachers’ beliefs, knowledge, school contexts, and subject-specific practices shape what access becomes. Tondeur et al. (2012) and Voogt et al. (2013) extend this argument by showing that integration is not an event of adoption but a development of pedagogical capacity around tools. This line of work becomes especially important for generative and agentic AI because such systems do not merely display content; they produce plausible explanations, examples, and tasks whose educational value still depends on professional judgement.

Teacher mediation alone, however, does not explain why learners may achieve stronger conceptual understanding. The present study, therefore, adopts a complementary view from mathematics education, in which learning is understood as the coordinated construction of meaning across multiple representations. Understanding functions and graphs requires learners to connect symbolic expressions, graphical relationships, numerical patterns, and contextual interpretations rather than treating these representations as isolated forms. From this perspective, teacher-mediated agentic AI has educational value not because it automates instruction, but because it may expand teachers’ capacity to generate multiple representations, alternative explanations, and misconception-sensitive examples that can be orchestrated into coherent mathematical reasoning. The proposed instructional mechanism, therefore, combines pedagogical expertise with representational richness: AI broadens the teacher’s instructional repertoire, while the teacher determines how those resources are sequenced, interpreted, and transformed into opportunities for conceptual learning.

The artificial intelligence in education literature inherits this same promise of responsiveness, although it has often imagined the instructional relation differently. Bloom’s (1984) two-sigma problem offered the field a powerful benchmark for individualised support, and intelligent tutoring systems later attempted to approximate that responsiveness through domain models, student models, feedback, and adaptive task sequences. VanLehn (2011), Ma et al. (2014), Steenbergen-Hu and Cooper (2014), and Kulik and Fletcher (2016) show that such systems can produce learning gains. Still, they also reveal a dominant analytic habit in which the learner-system relation becomes the central site of pedagogical action. Roll and Wylie (2016), Hwang, Xie, Wah, and Gasevic (2020), and Ouyang and Jiao (2021) widen the field toward questions of agency, roles, and human-AI relations; even so, the teacher’s work of selecting, disciplining, and turning AI output into classroom instruction remains thinner in the empirical record than it should be.

School mathematics makes this omission consequential rather than merely conceptual. Functions and graphs cannot be reduced to procedural fluency, because learners must coordinate symbolic, graphical, numerical, and verbal representations while reasoning about correspondence, variation, and covariation. Leinhardt, Zaslavsky, and Stein (1990) show that learners’ difficulties with graphing often arise when visual features are read without grasping the mathematical relationships they represent; more recent visualisation research sharpens this point by showing that external and dynamic visualisations are not decorative supports but forms through which mathematical relationships are made available for reasoning (Schoenherr & Schukajlow, 2024; Schoenherr, Strohmaier, & Schukajlow, 2024). Recent work on GeoGebra, animated mathematics videos, and spatial visualisation tools similarly shows that visual representations matter most when they are connected to concept formation, covariation, symbol sense, and teacher-guided mathematical activity rather than simply presented as attractive media (Medina Herrera, Juarez Ordonez, & Ruiz-Loza, 2024; ten Voorde, Piroi, & Bos, 2025; Zhang, Wang, Jia, Zhang, & Chen, 2025). Agentic AI may help a teacher generate examples, vary explanations, anticipate misconceptions, and construct graph demonstrations. Still, it may also produce fluent answer-giving that bypasses the mathematical struggle through which understanding is often built. The issue, then, is not AI availability; it is AI mediation.

Recent reviews make this gap visible. Zawacki-Richter, Marin, Bond, and Gouverneur (2019) ask where the educators are in AI research, and Chiu et al. (2023) show a rapidly expanding field still marked by uneven attention to classroom pedagogy, school contexts, and robust evidence of learning. Lai and Bower (2019) add a methodological warning that is directly relevant here: educational technology studies often measure satisfaction, engagement, or usability more readily than educational consequence. More recent work on generative AI, agency, and ChatGPT confirms both the promise and the risks, but much of it still centres on learner-facing use or higher education (Darvishi et al., 2024; Deng et al., 2025; Lo et al., 2024). What remains underdeveloped is classroom-level evidence on teacher-mediated AI in school mathematics, particularly evidence that is honest about the difficulty of separating instructional conditions from teacher, class, assessment, and implementation context in intact classroom designs. The distinction between teacher-mediated and learner-facing AI is more than a difference in who interacts with the technology; it reflects distinct instructional theories. Learner-facing systems position AI as the primary provider of explanations, feedback, and guidance, whereas teacher-mediated systems position AI as an instructional resource whose outputs are filtered, evaluated, adapted, and pedagogically orchestrated by the teacher. The educational mechanism, therefore, shifts from human–AI substitution towards human–AI augmentation, where professional judgement determines whether AI-generated representations, examples, and explanations become meaningful opportunities for mathematical learning. The present study is grounded in this distinction by conceptualising agentic AI as a resource that extends teachers’ pedagogical repertoire rather than replacing their instructional role. Against this gap, the study asks:

  • 1. RQ1. What baseline-adjusted achievement contrast was observed between teacher-mediated agentic AI-supported instruction and conventional instruction in Grade 10 functions and graphs?

  • 2. RQ2. How did the modelled probability of reaching meaningful mathematical mastery thresholds differ between the two instructional conditions?

  • 3. RQ3. How did the baseline-adjusted achievement contrast vary across learners’ baseline achievement levels?

  • 4. RQ4. How did changes in learners’ mathematical interest or confidence align with changes in achievement?

2. Related literature

The literature reviewed in this study is organised around four complementary strands that collectively explain why teacher-mediated AI represents a distinct instructional configuration. The first strand concerns the conditional effectiveness of educational technology. Rather than demonstrating that technology universally improves learning, major syntheses consistently show that educational outcomes depend on instructional design, pedagogical integration, and subject-specific implementation (Tamim et al., 2011; Li & Ma, 2010; Cheung & Slavin, 2013; Hillmayr et al., 2020). These findings shift attention away from technological capability itself towards the pedagogical conditions under which technology becomes educationally meaningful. Within this perspective, TPACK (Mishra & Koehler, 2006) provides an explanatory framework by positioning technology, pedagogy, and content knowledge as mutually dependent dimensions of effective teaching. Subsequent research further demonstrates that teacher beliefs, professional knowledge, and institutional support influence whether technological resources are translated into meaningful classroom practice (Ertmer, 1999; Hew & Brush, 2007; Tondeur et al., 2012; Voogt et al., 2013). Collectively, this literature establishes teacher mediation as the principal mechanism through which educational technologies influence learning.

A second strand comes from intelligent tutoring systems and the longer history of AI in education. Bloom’s (1984) two-sigma problem gave the field an enduring image of responsive individualised instruction, and intelligent tutoring systems translated part of that image into domain models, student models, feedback, and adaptive sequencing. VanLehn (2011), Ma et al. (2014), Steenbergen-Hu and Cooper (2014), and Kulik and Fletcher (2016) show that such systems can generate learning gains; however, this tradition also tends to treat the learner-system interaction as the principal site of pedagogical action. Roll and Wylie (2016) describe AI in education as both evolution and revolution, while Hwang et al. (2020) and Ouyang and Jiao (2021) show that newer AI systems move beyond tutoring into broader human-AI configurations. The unresolved issue is that the teacher often remains conceptually thin. Zawacki-Richter et al.’s (2019) question about educators’ location in AI research becomes decisive for classroom studies, because an AI system used by a teacher in mathematics instruction is not simply a learner-facing tutor under another name.

Recent generative-AI studies add urgency, but they also expose the field’s imbalance. Kasneci et al. (2023) articulate the promise and risks of large language models in education, including reliability, bias, assessment, and pedagogical design, while Chiu et al. (2023) show that the future of AI in education depends on stronger links between opportunity, challenge, and evidence. Empirical work in and around Computers & Education has begun to meet this demand: Huang, Lu, and Yang (2023) examine AI-enabled personalised recommendations; Darvishi et al. (2024) show how AI assistance reshapes student agency; Liu, Zhang, and Biebricher (2024) examine generative AI in digital multimodal composing; Lo, Hew, and Jong (2024) synthesise evidence on ChatGPT and engagement; and Deng et al. (2025) provide meta-analytic evidence on ChatGPT and student learning. These studies move the field beyond speculation. However, they also show why a school-mathematics study of teacher-mediated AI is needed, because much of the current evidence still privileges learner-facing use, higher education, writing, engagement, or general AI interaction rather than the teacher’s mediation of AI output for a specific mathematical topic. More importantly, this body of evidence has rarely conceptualised teacher-mediated AI as a distinct pedagogical configuration with its own theoretical assumptions, instructional mechanisms, and expected learning processes. Consequently, evidence regarding teacher-facing agentic AI in school mathematics remains conceptually and empirically underdeveloped.

A final strand comes from mathematics education, visualisation research, and evaluation methodology. Leinhardt, Zaslavsky, and Stein (1990) show that functions and graphing are conceptually demanding because learners must coordinate algebraic, graphical, numerical, and contextual representations rather than merely memorise procedures. That older insight has been renewed in the recent visualisation literature: Schoenherr and Schukajlow (2024) characterise the diversity of external visualisation in mathematics education, while Schoenherr, Strohmaier, and Schukajlow (2024) show through meta-analysis that visualisation interventions generally support mathematics learning, although their value depends on the way visualisation is designed and used. Zhang, Wang, Jia, Zhang, and Chen (2025) reach a similar conclusion for GeoGebra-based dynamic visualisation, and ten Voorde, Piroi, and Bos (2025) show that dynamic visuals in mathematics videos can serve distinct didactic roles, including connecting objects, supporting covariation, visualising processes, and developing symbol sense. Drijvers and Sinclair (2024) and Engelbrecht and Borba (2024) place these developments inside a broader account of digital technologies in mathematics education, where tools reshape purposes, representations, classroom spaces, and the work of doing mathematics. Yet Lai and Bower (2019) caution that educational technology evaluation often measures what is convenient rather than what is educationally consequential. In quasi-experimental settings, this problem is methodological as well as conceptual: Rubin (1974), Holland (1986), and Rosenbaum and Rubin (1983) remind researchers that treatment effects require comparisons between observed outcomes and plausible alternative outcomes under stated assumptions, while Hill (2011) shows how Bayesian causal modelling can support such estimation while retaining uncertainty. The evaluative task is therefore sharper than asking whether AI is liked or whether it appears innovative; it is to ask whether teacher-mediated AI changes mathematics achievement in a way that is both defensible and interpretable.

A cross-cutting issue in this literature is therefore explainability, but not only in the technical sense of making an AI model transparent. For educational research, the instructional claim itself must be explainable: what was mediated, by whom, under what curriculum logic, and with what evidence of learning. Khosravi et al. (2022) argue that explainable AI matters because educational systems shape feedback, learning decisions, and interpretation; Long and Magerko (2020), Ng et al. (2021), and Zhong and Liu (2025) similarly show that AI literacy has become a school-level concern rather than a distant higher-education issue. Darvishi et al. (2024) add that AI assistance can reshape student agency, which means that performance gains alone cannot exhaust the educational question. For mathematics education, however, attitudes, engagement, and literacy also cannot replace evidence of disciplinary learning. The present study responds to this challenge by treating teacher-mediated AI as the analytical unit of interest, where the educational question is not whether AI can generate instructional material, but whether teachers’ pedagogical orchestration of AI-generated representations can contribute to meaningful improvements in mathematical understanding and achievement.

2.1 Theoretical framework

The present study is informed by an integrated theoretical framework that combines Technological Pedagogical Content Knowledge (TPACK) with perspectives from mathematics education on representational learning. Rather than treating artificial intelligence as an autonomous instructional agent, the framework conceptualises AI as a teacher-mediated resource whose educational value depends on how teachers integrate technological capability with pedagogical reasoning and mathematical content knowledge. The intervention is therefore understood as a process of pedagogical orchestration rather than technological substitution. From the perspective of TPACK (Mishra & Koehler, 2006), effective technology integration occurs when teachers successfully coordinate knowledge of mathematics, pedagogy, and technology. The framework rejects the assumption that technological sophistication alone improves learning. Instead, AI-generated explanations, examples, visualisations, and formative questions become educationally meaningful only when teachers evaluate their mathematical accuracy, align them with curriculum goals, adapt them to learners’ needs, and integrate them into coherent instructional sequences. Teacher expertise, therefore, remains the central mechanism through which technology contributes to learning.

Theories of mathematical representation complement this pedagogical perspective. Learning functions and graphs requires learners to coordinate symbolic, graphical, numerical, and contextual representations rather than mastering isolated procedures (Leinhardt et al., 1990). Contemporary mathematics education research similarly demonstrates that conceptual understanding develops through purposeful engagement with multiple external representations, dynamic visualisations, and teacher-guided interpretation of mathematical relationships (Schoenherr & Schukajlow, 2024; Schoenherr et al., 2024). From this perspective, AI has educational value when it expands teachers’ capacity to generate diverse representations and multiple explanatory pathways that support conceptual reasoning rather than procedural imitation. The framework, therefore, proposes that teacher-mediated AI operates through an indirect instructional mechanism. Rather than interacting directly with learners, the AI system expands the range of curriculum-aligned explanations, worked examples, graphical demonstrations, formative questions, and misconception-sensitive prompts available to the teacher. Teachers then exercise professional judgement by selecting, adapting, sequencing, and contextualising these resources within classroom instruction. Learning improvements are therefore expected to emerge not because AI replaces teaching, but because it enriches the instructional repertoire through which teachers support mathematical understanding.

3. Method
3.1 Research design

The study used a quasi-experimental pre-test/post-test comparison design to estimate the effect of teacher-mediated AI-supported instruction on learners’ achievement in Grade 10 functions and graphs. This design was appropriate because the intervention was implemented under intact classroom conditions, where the educational object was not a decontextualised software exposure but a teacher’s use of an AI system during the teaching of a topic. Preserving that instructional form meant that learners could not be treated as if they had been individually randomised to equivalent conditions. The design, therefore, required an analytic strategy that compared groups cautiously, adjusted for measured baseline differences, and stated clearly the assumptions under which the comparison could be interpreted causally. The analysis followed the logic of causal comparison under observed covariate adjustment. Rather than asking only whether two intact groups differed descriptively at post-test, the study estimated the expected difference in post-test achievement between comparable learners under the AI-supported and conventional instructional conditions. The causal reading is therefore conditional rather than automatic: it depends on the measured baseline variables, the temporal ordering of pre-test and post-test measures, overlap between groups, and the modelling assumptions made explicit below.

3.2 Setting, participants, and instructional conditions

The analytic sample consisted of 115 Grade 10 mathematics learners with complete pre-test and post-test records for the focal achievement and interest measures. Fifty-eight learners were taught in the teacher-mediated AI-supported condition, while fifty-seven were taught in the conventional instruction condition. The learners came from the same socio-economic status and school-quintile context, and the intervention and comparison conditions covered the same mathematical topic: functions and graphs. The study involved two intact classes, one in each instructional condition, and both followed the same eight-week study schedule: one week of pre-intervention assessment, six weeks of classroom instruction, and one week of post-intervention assessment. Two different teachers delivered the two instructional conditions; both teachers were trained for the study and used the same lesson plan. These features strengthen procedural and contextual comparability, but they do not turn the study into an individually randomised experiment or separate the instructional condition from teacher and class context. For analysis, instructional condition was represented as a binary treatment indicator, with teacher-mediated AI-supported instruction treated as the focal condition and conventional instruction as the comparison condition. Gender was included as a measured learner characteristic because it was recorded before the post-test outcome and could reasonably be related to classroom participation, confidence, or prior opportunity. Baseline achievement was specified as a pre-intervention covariate to ensure that the models compared learners who were more alike in their starting points on the topic.

3.2.1 Instructional procedure

The intervention followed a structured instructional sequence designed to ensure that both instructional conditions received equivalent curriculum coverage, teaching time, and learning objectives. The study was implemented over eight weeks, comprising one week of pre-intervention assessment, six weeks of classroom instruction, and one week of post-intervention assessment. During the first week, all learners completed the Functions and Graphs Achievement Test (FGAT) and the Functions and Graphs Interest Inventory (FGII) to establish baseline achievement and interest before instruction. These baseline measures were subsequently incorporated into the Bayesian models as pre-intervention covariates.

The instructional phase extended across six consecutive weeks during the normal school timetable. Mathematics was taught for five 45-minute periods each week, totalling approximately 22.5 hours of instruction across the intervention. Both instructional conditions followed the same Grade 10 CAPS curriculum sequence, used identical learning objectives, and addressed the same mathematical concepts. The instructional progression comprised: (a) introduction to functions and relationships between symbolic, numerical, graphical, and contextual representations; (b) linear functions; (c) quadratic functions; (d) hyperbolic functions; (e) exponential functions; and (f ) revision, consolidation, and integrated problem-solving activities. One week after completing the instructional programme, learners in both conditions completed the FGAT and FGII post-tests. This interval reduced immediate recall effects while allowing learners time to consolidate their understanding before the outcome measures were administered.

3.3 The mothusi intervention

3.3.1 Intervention boundary and instructional object

The first issue in describing Mothusi is one of the intervention boundary, because the platform was capable of more than the study asked it to do. In this phase, the teacher assigned to the AI-supported condition was the sole classroom user of Mothusi: learners did not log in, submit prompts, or receive responses from the system on personal devices, and the learner-facing functionality that existed elsewhere in the wider platform was not activated as part of the treatment. What the study evaluates, therefore, is not the unrestricted capability of the product, but a deliberately narrower instructional arrangement in which one teacher used the system to extend the planning, explanation, representation, and questioning resources available throughout the six-week classroom instruction phase of the functions-and-graphs unit. Within that boundary, Mothusi was neither an autonomous tutor nor a replacement for the mathematical judgement carried by the teacher. It was used as a teacher-mediated AI support system configured for Grade 10 CAPS functions and graphs, through which the teacher could request alternative explanations, worked examples, graph demonstrations, formative questions, and prompts organised around common learner misconceptions. The treatment object was thus a teacher-AI-curriculum configuration: Mothusi enlarged the repertoire from which the teacher could work, while the teacher determined what became instruction, in what sequence, and with what explanations.

3.3.2 System configuration and curriculum grounding

The implementation comprised an authenticated teacher workspace connected to OpenClaw, which maintained the Mothusi persona and session context, routed common graphing and introductory functions requests through preconfigured response and visualisation routines, and passed requests requiring model generation to the configured language model; during the principal May implementation, those model-mediated requests were routed to GPT-5.5. This agentic arrangement did not mean that the system acted independently in the classroom. It meant that the workspace could coordinate a bounded set of resources and tools in response to a teacher’s request, while the decision to use, change, or reject the resulting material remained with the teacher.

The system was curriculum-bound rather than an invitation to generate mathematics freely. Its local, page-indexed corpus contained the DBE Grade 10 Annual Teaching Plan as curriculum anchor, the Siyavula Grade 10 mathematics text as the primary explanatory source, selected Mind the Gap material for extension, a functions-and-graphs topic index, nine misconception patterns with corrective prompts, and a 24-item formative question bank. Model-mediated content requests were configured to consult this corpus before producing a substantive response. At the same time, SymPy supported algebraic checking, Matplotlib for static graph production, KaTeX for mathematical notation, and Desmos for interactive graphs. These technical provisions widened what the teacher could call upon; they did not, by themselves, establish the mathematical or pedagogical suitability of what entered a lesson.

Before implementation commenced, both participating teachers completed a structured professional development programme designed specifically for the study. The programme familiarised teachers with the lesson sequence, curriculum objectives, instructional protocol, and the intended pedagogical role of Mothusi within classroom instruction. Particular emphasis was placed on maintaining teacher authority over mathematical explanations by critically evaluating AI-generated outputs, verifying mathematical correctness, selecting appropriate representations, and adapting generated material to learners’ instructional needs. This preparation ensured that AI functioned as a teacher-support tool rather than as an autonomous instructional agent.

3.3.3 Teacher enactment of AI-generated support

At the point of use, the AI component was located in the generation and coordination of candidate instructional material rather than solely in the graphing display. A teacher request entered through the authenticated workspace; OpenClaw combined that request with the Mothusi persona and session context, consulted the curriculum-bounded corpus, and, where a generated response was required, routed the request to GPT-5.5 before returning a candidate explanation, example, question sequence, or graph-supported activity. SymPy, Matplotlib, KaTeX, and Desmos could support the mathematical checking and representation of that response, but these tools were not themselves the AI intervention. What distinguished Mothusi from conventional graphing software was the model-mediated production and coordination of curriculum-grounded teaching support in response to the teacher’s request.

What entered the classroom, however, still travelled through the teacher. During preparation, the teacher used these generated candidates to inspect possible explanations, examples, questions, and visual representations; during instruction, selected material could be displayed, revoiced, slowed down, or reorganised in response to the class. The teacher checked the mathematics, decided whether a representation served the immediate lesson, and placed questions around it so that the system’s capacity to produce an answer did not displace the learner’s mathematical work. The substantive intervention was therefore the teacher’s mediated use of AI-generated, curriculum-grounded support, not exposure to a platform interface and not independent learner interaction with a chatbot.

Throughout instruction, teacher mediation followed a consistent pedagogical sequence. Before each lesson, the teacher used Mothusi to generate candidate explanations, worked examples, graphical demonstrations, misconception-sensitive prompts, and formative questions aligned with the planned curriculum objectives. These outputs were critically reviewed for mathematical correctness, curriculum alignment, and pedagogical suitability before classroom use. During lessons, the teacher selected appropriate AI-generated resources, integrated them with direct instruction, facilitated classroom discussion, posed probing questions, responded to learner misconceptions, and determined the pace and sequencing of instructional activities. Learners therefore experienced AI only through teacher-mediated explanations, visual representations, and guided mathematical reasoning rather than through direct interaction with the system. Accordingly, references in this paper to the Mothusi intervention denote only this teacher-mediated configuration. They do not denote the learner login, independent learner chat, researcher dashboard, or other platform functions developed for separate phases of the wider programme.

3.4 Comparison condition

The comparison condition was conventional instruction on the same functions-and-graphs topic. This condition should not be read as the absence of teaching; it represents the ordinary instructional arrangement against which the teacher-mediated AI-supported condition was compared. The contrast is therefore between two ways of teaching the topic: conventional instruction and instruction in which the teacher used Mothusi to extend the planning, explanation, representation, and questioning resources available during the topic. Because the comparison condition covered the same curriculum topic, followed the same lesson plan, and was implemented over the same eight-week study schedule (including six weeks of classroom instruction), the design reduces the risk that the observed difference reflects topic exposure or instructional time rather than instructional form. The shared socio-economic and school-quintile context also reduces one broad source of contextual imbalance. The remaining interpretive burden, however, cannot be carried by these design features alone. It is carried by baseline adjustment, explicit causal assumptions, and cautious interpretation of the quasi-experimental comparison, with robustness checks used only to examine model-specification fragility rather than to remove teacher, class, or implementation confounding.

3.4.1 Implementation fidelity

Several procedures were implemented to maximise intervention fidelity across instructional conditions. Both teachers followed identical lesson plans, addressed the same curriculum objectives, taught the same mathematical content during the same instructional period, and used the same assessment instruments. Throughout the intervention, the researcher monitored implementation through scheduled classroom observations and discussions with participating teachers to verify adherence to the instructional protocol. Where minor deviations occurred, these were addressed immediately to ensure consistency with the planned intervention. Although fidelity procedures strengthened procedural consistency, they could not eliminate the inherent teacher- and class-level confounding associated with the intact-class quasi-experimental design.

3.5 Measures

The primary outcome was the post-test total score on the Functions and Graphs Achievement Test (FGAT). The corresponding FGAT pre-test total was used as the baseline achievement covariate. The FGAT comprised 25 multiple-choice items aligned with the Grade 10 functions-and-graphs topic, each worth 2 marks on the classroom test form. For analysis, scores were recorded on a 0–100 scale. The achievement test was treated as the most direct measure of whether learners performed better after instruction on the target mathematical topic. Because the study concerns functions and graphs, the achievement outcome bears the main inferential burden of the paper. The secondary outcome was the Functions and Graphs Interest Inventory post-test total (FGII post-test), adjusted for the FGII pre-test total. The FGII was a 30-item four-category inventory with response options ranging from strong agreement to strong disagreement. Items addressed enjoyment, value, confidence or competence, and engagement with functions and graphs; negatively worded items were coded so that higher total scores indicated more favourable interest-and-confidence responses. Internal consistency evidence from the main sample was strong for the total score, with Cronbach’s alpha of 0.883. The FGII outcome was treated as secondary because the paper’s central claim is about mathematical achievement and mastery rather than affective change alone.

3.6 Ethical considerations

Ethical clearance for the larger study was granted by the General/Human Research Ethics Committee of the researchers’ institution. The study used learner-level educational data for research purposes and is reported without identifying the school, teacher, or individual learners. The analysis dataset contains only pseudonymous learner identifiers. The research team holds institutional consent and school-permission documents.

3.7 Data preparation

The dataset was prepared before modelling by retaining learners with complete information on instructional condition, gender, FGAT pre-test, FGAT post-test, FGII pre-test, and FGII post-test. The pre-test and post-test totals were treated as continuous scores, with higher values representing stronger achievement or more favourable interest-and-confidence responses, depending on the instrument. Gain scores were computed for descriptive reporting, but the causal estimates were based on post-test models that adjusted directly for the corresponding pre-test measure. Baseline FGAT and FGII scores were standardised before modelling. This step made the treatment coefficient interpretable at the average baseline level and made the interaction between baseline achievement and instructional condition easier to interpret. No post-test outcome information was used to define the instructional condition; the treatment indicator referred only to the instructional arrangement under which learners were taught.

3.8 Causal assumptions and estimand

The causal estimand was the average difference in post-intervention FGAT achievement expected under teacher-mediated AI-supported instruction compared with conventional instruction, after adjustment for measured baseline variables. In the potential-outcomes language associated with Rubin and Holland, this is the contrast between a learner’s potential post-test outcome under the AI-supported condition and that learner’s potential post-test outcome under the conventional condition. Because each learner is observed in only one condition, the alternative outcome is necessarily unobserved and must be estimated rather than directly measured. For this contrast to be interpreted causally, the analysis assumes temporal ordering, consistency, conditional exchangeability, overlap, limited cross-condition contamination, measurement validity, and approximate model adequacy. These assumptions require that baseline measures precede the outcome, that observed outcomes correspond to the instructional condition received, that comparable learners exist across instructional conditions, and that no major unmeasured confounder simultaneously determines instructional condition and post-test achievement after adjustment for measured baseline variables. In this study, conditional exchangeability is a strong assumption because the design used two intact classes, one per condition, taught by two different trained teachers, even though both classes followed the same eight-week study schedule, used the same lesson plan, and came from the same socio-economic and school-quintile context. The data cannot prove these assumptions. They define the conditions under which the estimates can be read as causal rather than as baseline-adjusted associations.

3.9 Bayesian analysis

The primary FGAT analysis used a Bayesian baseline-adjusted Gaussian model. The post-test FGAT score was modelled as a function of instructional condition, standardised baseline FGAT, gender, and a treatment-by-baseline interaction. The treatment coefficient estimated the additional FGAT post-test points expected under the AI-supported instructional condition at the average baseline achievement level. The interaction term examined whether the treatment effect increased or decreased across the baseline achievement distribution. Weakly informative priors were used: the intercept was centred on the observed post-test mean with a standard deviation of 12; the treatment effect had a Normal(0, 15) prior; the baseline, gender, and interaction effects had Normal priors with standard deviations of 8, 5, and 6, respectively; and the residual standard deviation had a half-Normal(8) prior.

The secondary FGII model followed the same baseline-adjusted logic but did not include the treatment-by-baseline interaction. FGII post-test total was modelled as a function of treatment, standardised FGII pre-test, and gender. The FGII priors were also weakly informative: the intercept was centred on the observed FGII post-test mean with a standard deviation of 10; the treatment effect had a Normal(0, 8) prior; the baseline and gender coefficients had Normal priors with standard deviations of 5 and 4, respectively; and the residual standard deviation had a half-Normal(5) prior. The FGII analysis was specified as a secondary analysis rather than as the main test of the intervention.

Models were estimated in Python using PyMC, and posterior summaries were produced with ArviZ. Each reported Bayesian model used four Markov chains, 1,000 tuning draws, and 1,000 posterior draws per chain, with a target acceptance probability of 0.90. The analysis reports posterior means, posterior standard deviations, 95% credible intervals, and posterior probabilities in the direction of educational interest. Convergence and sampling quality were assessed using R-hat, effective sample sizes, and visual diagnostics.

3.10 Mastery thresholds, heterogeneity, and secondary associations

To express the baseline-adjusted contrast in educationally interpretable terms, the analysis translated the FGAT model’s posterior predictions into probabilities of reaching achievement thresholds of 70, 80, 90, and 95. These thresholds were not treated as universal standards; rather, they were used as meaningful cut points for describing movement into stronger regions of functions-and-graphs performance. For each threshold, the model estimated the expected probability of crossing the threshold under the AI-supported condition and under conventional instruction, and then estimated the posterior distribution of the probability difference. The heterogeneity analysis estimated the baseline-adjusted contrast at selected points in the baseline FGAT distribution: the 10th, 25th, 50th, 75th, and 90th percentiles. This made it possible to ask whether the observed separation was concentrated among learners who began the topic with lower or higher initial achievement. Because the study used two intact classes, these estimates were interpreted as descriptive variation in the modelled contrast rather than as subgroup-specific causal effects.

3.11 Robustness checks

Several additional checks tested whether the achievement pattern depended on one modelling choice. First, baseline comparability was examined using raw differences, standardised differences, and common-support checks for the FGAT pre-test. Second, frequentist ANCOVA models were fitted as cross-checks, including a specification matching the Bayesian covariates, a no-interaction specification, a specification adding baseline FGII, and a trimmed sample excluding extreme FGAT pre-test values. Third, posterior mean fitted values from the Bayesian model were used to examine residual size by condition. These checks do not replace the Bayesian estimand and do not address the teacher-and-class confounding created by the two-class design; their role is narrower: to test whether the descriptive achievement pattern is fragile to a single analytic specification.

4. Results

The results are presented in the order of the research questions. The first analysis examines baseline comparability and the contrast in baseline-adjusted achievement between instructional conditions. Subsequent analyses evaluate the probability of attaining predefined mastery thresholds, examine variation in the estimated achievement contrast across baseline achievement levels, assess changes in learners’ mathematical interest and confidence, and conclude with robustness checks that evaluate the stability of the principal findings across alternative model specifications.

4.1 Baseline comparability and achievement pattern

The first empirical issue is whether the two groups began from broadly comparable achievement baselines before diverging after instruction. Table 1 shows that the conventional instruction group had a mean FGAT pre-test score of 61.70. In contrast, the teacher-mediated agentic AI-supported group had a mean of 62.02, a raw baseline difference of only 0.32 points. After the intervention, however, the mean post-test score was 63.81 in the conventional condition and 93.12 in the AI-supported condition, producing a raw difference of 29.31 points. The same pattern appears in gain scores: learners in the conventional condition gained 2.11 points on average, while learners in the AI-supported condition gained 31.10 points.

Table 1. Descriptive achievement and interest outcomes by instructional condition.OutcomeConventional M (SD)AI-supported M (SD)Raw differenceFGAT pre-test 61.70 (6.45)62.02 (7.79)0.32FGAT post-test 63.81 (4.11)93.12 (5.58)29.31FGAT gain2.11 (6.19)31.10 (9.77)29.00FGII pre-test 19.37 (2.33)18.86 (3.11)−0.51FGII post-test 17.49 (2.81)20.64 (2.43)3.15FGII gain−1.88 (2.88)1.78 (3.94)3.65

The baseline comparisons indicate that the instructional groups were similar on measured pre-intervention characteristics, supporting the use of baseline-adjusted comparisons while recognising that equivalence cannot be established in an intact-class quasi-experimental design. The standardised difference on the FGAT pre-test was 0.04, the standardised difference on the FGII pre-test was −0.18, and the standardised difference for the proportion of female learners was 0.19. The common-support check was also reassuring: the shared FGAT pre-test range was 46 to 75, which included all conventional learners and 98.3% of learners in the AI-supported condition. These checks justify a baseline-adjusted comparison while still limiting the causal interpretation to the two-class, two-teacher design. The Bayesian baseline-adjusted model shows that the achievement difference was not merely a raw-score pattern. As Table 2 reports, the estimated baseline-adjusted between-condition contrast was 29.12 additional FGAT post-test points for the teacher-mediated agentic AI-supported class, with a 95% credible interval from 27.31 to 30.92. The posterior probability that the fitted contrast was greater than zero was 1.00. The estimated treatment coefficient therefore, represents the expected difference in FGAT post-test performance between instructional conditions, after adjustment for measured baseline variables and gender, under the specified Bayesian model.

Table 2. Bayesian FGAT model estimates for the baseline-adjusted achievement model.ParameterMeanSD95% credible intervalPr(> 0)Intercept64.030.83[62.45, 65.66]Teacher-mediated agentic AI condition29.120.92[27.31, 30.92]1.00Baseline FGAT achievement1.700.73[0.26, 3.12]Female learner−0.130.92[−1.87, 1.71]AI condition × baseline FGAT−1.870.94[−3.69, −0.06]Residual SD4.890.34[4.25, 5.56]
4.2 Mastery-Threshold movement

The second research question moves from average achievement to the educational size of the observed separation. A large mean difference is important, but teachers and mathematics educators also need to know whether learners are likely to move into higher-performing regions. Table 3 shows that the modelled probability of reaching 70 or above was effectively 1.00 in the teacher-mediated agentic AI class, compared with 0.125 in the conventional class. At the 80-point threshold, the modelled probability was 0.995 for the AI-supported condition and 0.001 for the conventional condition. Even at the more demanding 90-point threshold, the modelled probability was 0.732 in the AI-supported condition and approximately zero in the conventional condition.

Table 3. Modelled probabilities of reaching achievement thresholds.ThresholdAI expected PConventional expected PDifference95% credible intervalPr (diff. > 0)≥ 701.0000.1250.875[0.804, 0.929]1.00≥ 800.9950.0010.994[0.983, 0.999]1.00≥ 900.732<.0010.732[0.639, 0.815]1.00≥ 950.349<.0010.349[0.257, 0.449]1.00

The posterior estimates indicate consistently higher probabilities of reaching each predefined mastery threshold under the teacher-mediated agentic AI condition than under conventional instruction. The estimated probability differences were greatest for the 80-point and 90-point thresholds, where the posterior distributions showed clear separation between instructional conditions. Posterior probabilities that the threshold differences exceeded zero were 1.00 for all reported thresholds.

4.3 Variation by baseline achievement

The third research question examines whether the achievement separation was concentrated among learners with particular baseline profiles. The treatment-by-baseline interaction was negative in the FGAT model, suggesting that the estimated contrast became somewhat smaller as baseline achievement increased. That pattern, however, should not be misread as evidence that only lower-achieving learners showed the separation. Table 4 shows that the estimated contrast remained strongly positive across the observed baseline distribution: 32.49 points at the 10th percentile of baseline achievement, 28.82 points at the median, and 27.25 points at the 90th percentile.

Table 4. Estimated baseline-adjusted FGAT contrast across baseline achievement levels.Baseline locationFGAT pre-test scoreEstimated contrast95% credible intervalPr(> 0)10th percentile49.032.49[28.75, 36.25]1.0025th percentile60.029.61[27.72, 31.47]1.0050th percentile63.028.82[26.98, 30.64]1.0075th percentile65.528.17[26.15, 30.19]1.0090th percentile69.027.25[24.75, 29.74]1.00

The estimated treatment contrast remained positive across the observed baseline achievement distribution. Although the estimated contrast declined modestly as baseline achievement increased, the posterior distributions consistently favoured the teacher-mediated AI condition across all reported baseline percentiles.

4.4 Interest, confidence, and achievement gains

The fourth research question treats learners’ mathematical interest and confidence as secondary outcomes rather than as substitutes for achievement. Table 5 presents the Bayesian baseline-adjusted estimates for the Functions and Graphs Interest Inventory (FGII) together with the correlations between gains in achievement and gains in interest and confidence. The descriptive pattern remained favourable to the teacher-mediated AI-supported condition. The conventional group declined by 1.88 points on the FGII total, while the AI-supported group increased by 1.78 points. The Bayesian baseline-adjusted estimate for the FGII post-test total was 3.23 points in favour of the AI-supported condition, with a 95% credible interval from 2.25 to 4.19 and a posterior probability greater than zero of 1.00.

Table 5. Secondary FGII and gain-correlation results.QuantityEstimate95% credible intervalPr(> 0)Adjusted FGII post-test effect3.23[2.25, 4.19]1.00Overall FGAT gain-FGII gain correlationr = .426Conventional condition correlationr = −.015AI-supported condition correlationr = .059

The relation between achievement gain and FGII gain was more limited than the group-level pattern might suggest. Across the full sample, the gain correlation was moderate and positive, r = .426, p < .001, n = 115. Within conditions, however, the correlations were near zero: r = −.015, p = .913, n = 57 in the conventional condition and r = .059, p = .662, n = 58 in the AI-supported condition. Collectively, the FGII analyses indicate higher post-intervention interest and confidence scores in the teacher-mediated AI condition, along with a moderate positive association between achievement and FGII gains across the full sample. Within-condition gain correlations, however, remained close to zero. Sampling diagnostics were adequate (R-hat = 1.00), but the estimates remain conditional on the quasi-experimental assumptions stated above.

4.5 Robustness and model adequacy

The robustness checks had a narrower purpose than causal validation. They asked whether the central achievement estimate was fragile to a small set of modelling choices, not whether the intact-group comparison was free of confounding. The Bayesian model remained the primary model because it corresponded to the stated estimand and reported uncertainty as a credible interval. The frequentist ANCOVA models were used as specification checks: one matched the Bayesian covariates, one removed the treatment-by-baseline interaction, one added baseline FGII as an additional covariate, and one trimmed extreme FGAT pre-test values. As Table 6 shows, this sequence did not materially change the achievement estimate, which remained between 29.10 and 29.54 FGAT points. That stability reduces concern about a single narrow modelling artefact, but all checks still use the same intact-condition comparison.

Table 6. Robustness checks for the FGAT achievement contrast.ChecknEstimated contrast95% intervalBayesian FGAT model11529.12[27.31, 30.92]Frequentist ANCOVA: same covariates11529.29[27.49, 31.09]Frequentist ANCOVA: no interaction11529.30[27.47, 31.12]Frequentist ANCOVA: added FGII baseline11529.10[27.34, 30.86]Frequentist ANCOVA: trimmed extremes11229.54[27.73, 31.35]

Posterior fitted values indicated satisfactory model performance, with an overall RMSE of 4.74 FGAT points and mean residuals centred close to zero for both instructional conditions. Across all robustness analyses, the estimated treatment contrast remained stable, indicating that the principal findings were not materially altered by the alternative model specifications examined.

5. Discussion

The most important interpretive feature of the study is not simply that the AI-supported condition scored higher, but that the estimated separation is unusually large for an intact-classroom education study. The Bayesian model estimated a 29.12-point baseline-adjusted advantage for the teacher-mediated agentic AI condition, with a residual standard deviation of approximately 4.89 points. On that residual scale, the contrast is close to six standard deviations, which is far beyond the magnitude that readers usually expect in classroom intervention research and almost three times the scale invoked by Bloom’s (1984) two-sigma benchmark. This scale should not be read as straightforward proof of exceptional pedagogical power. It should first be read as a methodological signal. The conventional group gained only about two points on the FGAT, while the AI-supported group moved to a post-test mean of 93.12 out of 100, close to the test ceiling. A powerful instructional arrangement can produce such an asymmetry. Still, it can also be produced or exacerbated by teacher effects, class effects, assessment alignment, teaching to the test, differential implementation intensity, or ceiling compression in the higher-performing condition.

While these observations justify caution in interpreting the magnitude of the estimated effect, they also invite a broader educational question. Beyond the statistical contrast itself, the study contributes to ongoing debates concerning how artificial intelligence should be positioned within classroom instruction. Rather than evaluating AI as an autonomous instructional agent, the present study examined a teacher-mediated configuration in which AI functioned as a pedagogical resource that expanded teachers’ instructional repertoire while leaving curriculum decisions, mathematical explanations, and classroom orchestration under teacher control. The findings should therefore be interpreted not only as evidence of achievement differences but also as evidence supporting a distinct conceptualisation of AI-supported mathematics teaching.

The first research question can therefore be answered only conditionally. After adjustment for baseline achievement and measured learner characteristics, the AI-supported condition was associated with a very large post-test advantage in functions-and-graphs achievement. Several features of the design strengthen the comparison: the groups began from very similar FGAT baselines; learners came from the same socio-economic status and school-quintile contexts; the same mathematical topic was taught; both classes followed the same eight-week study schedule; both teachers were trained; and both used the same lesson plan. Yet these features reduce only some sources of imbalance. Because two different teachers taught two intact classes, one in each instructional condition, the condition remains fully entangled with teacher and class context. The design, therefore, cannot distinguish the effect of teacher-mediated agentic AI from the effect of the particular teacher, classroom culture, pacing, class-level preparation, or differences in how the shared lesson plan was enacted. This is the most serious threat to the conditional-exchangeability assumption, and it should sit at the centre of the causal interpretation rather than at the edge of the limitations section.

The same calibration applies to statistical precision. The reported credible interval is narrow because the single-level model treats the 115 learner observations as conditionally independent after adjustment. In this study, however, learners were nested within two intact classes, and each class shared a teacher, lesson enactment, peer environment, pacing, and assessment preparation. With only two classes, one per condition, the model cannot estimate meaningful class-level variance or separate treatment from class and teacher effects. The posterior standard deviation and credible interval should therefore be read as conditional model uncertainty, not as uncertainty that fully reflects the clustered design. The robustness checks are useful, but they are specification checks within the same two-class comparison. They do not address unmeasured confounding, teacher-condition confounding, classroom non-independence, or differential enactment because they all inherit the same design.

The second research question, which translated posterior predictions into mastery-threshold probabilities, remains educationally valuable precisely because it speaks in a language that teachers and mathematics educators can recognise. The movement across the 70, 80, 90, and 95 thresholds shows that the between-condition contrast was not merely a small mean difference with limited classroom consequences. However, the threshold interpretation is also affected by the same design and measurement conditions. A 25-item curriculum-aligned multiple-choice test used as both a pre-test and a post-test is well-suited to the taught topic, but it also raises the possibility of assessment alignment and teaching to the test, especially when one condition approaches the ceiling. The threshold probabilities, therefore, describe the educational size of the observed and modelled separation; they do not, by themselves, prove that the AI-supported instructional configuration caused the full separation.

The third and fourth research questions add further nuance, although they should not be asked to do more than the data allow. The treatment-by-baseline interaction indicated that the estimated FGAT contrast decreased slightly as baseline achievement increased, yet the contrast remained positive across the observed baseline distribution. In a less design-limited study, this might support a stronger claim about broad reach across learner profiles. Here, because the overall contrast is extremely large and the AI-supported condition approached the test ceiling, the heterogeneity result is better read as descriptive evidence that the separation was not confined to one baseline-achievement subgroup. Similarly, the FGII results indicate a favourable group-level difference in interest and confidence. Still, the near-zero within-condition gain correlations mean that the study cannot claim that larger achievement gains were psychologically coupled with larger shifts in interest or confidence.

The study’s contribution to the AI-in-education literature is therefore more modest, but also more interesting, than a claim that agentic AI produced a six-standard-deviation causal effect. Bloom (1984), VanLehn (2011), Ma et al. (2014), Steenbergen-Hu and Cooper (2014), and Kulik and Fletcher (2016) place AI and tutoring research inside a long search for responsive instruction, while Zawacki-Richter et al. (2019), Chiu et al. (2023), and the recent generative-AI literature ask the field to take educators and classroom conditions more seriously. What this study contributes is a sharper empirical object: agentic AI used by the teacher during mathematics instruction, not AI delivered directly to learners as an autonomous tutor. That object matters, but the present evidence does not yet establish the mechanism. Nevertheless, the study provides an important conceptual advance by treating teacher-mediated agentic AI as a distinct instructional configuration rather than simply another form of learner-facing AI. Within this configuration, AI extends teachers’ pedagogical capacity by broadening the range of explanations, representations, worked examples, formative questions, and misconception-sensitive prompts available during instruction. The educational mechanism, therefore, resides not in technological autonomy, but in teachers’ professional judgement regarding how AI-generated resources are evaluated, adapted, sequenced, and integrated into meaningful mathematical learning experiences. This interpretation extends the current AI-in-education literature by repositioning pedagogical expertise—not AI itself—as the principal mechanism through which technology contributes to learning.

This distinction is important for mathematics education. Functions and graphs require learners to coordinate symbols, graphs, variation, and meaning, and the visualisation literature shows that representations support learning when they are tied to disciplined mathematical activity rather than merely displayed (Leinhardt, Zaslavsky, & Stein, 1990; Schoenherr & Schukajlow, 2024; Schoenherr, Strohmaier, & Schukajlow, 2024; ten Voorde, Piroi, & Bos, 2025). The plausible educational argument is that Mothusi expanded the teacher’s repertoire of explanations, worked examples, graph demonstrations, and misconception-sensitive questions. Accordingly, the findings are theoretically consistent with the integrated framework proposed in this study, in which TPACK and representational learning operate together through teacher-mediated pedagogical orchestration. Rather than replacing teachers, agentic AI appears to strengthen teachers’ capacity to design, adapt, and coordinate multiple mathematical representations that support conceptual understanding. Although stronger experimental designs remain necessary before firm causal claims can be made, the present findings suggest that the educational value of AI may depend less on technological sophistication than on the quality of teachers’ pedagogical use of AI-generated resources.

Taken together, the findings suggest that future AI-in-education research may benefit from distinguishing among learner-facing, teacher-mediated, and collaborative human-AI instructional configurations rather than treating AI-supported learning as a single educational category. Each configuration is likely to operate through different pedagogical mechanisms, require different forms of teacher expertise, and produce different patterns of learner engagement and achievement. The present study, therefore, provides not only evidence from a Grade 10 mathematics intervention but also a conceptual foundation for examining how teacher-mediated agentic AI can be theorised, implemented, and evaluated in future classroom research.

6. Limitations and future research

Several limitations bound the claims made from this study. The design used two intact classes, one per instructional condition, taught by two different trained teachers; the common lesson plan, common eight-week study schedule, and shared socio-economic and school-quintile context strengthened comparability, but did not independently vary teacher, class, and condition. The primary Bayesian model was fitted at the learner level and therefore did not model class dependence. The FGAT was a 25-item, curriculum-aligned multiple-choice test used as both a pre-test and a post-test, and the AI-supported class approached the test ceiling. The study also lacked process evidence showing how AI-generated material was selected, adapted, verified, and enacted during instruction. Future research should move from a two-class comparison to designs that can separate teacher, class, and treatment. This requires multiple teachers and multiple classes per condition, random or matched assignment at class or school level where feasible, crossed or counterbalanced teacher-treatment arrangements where ethically and practically possible, and multilevel models that reflect the structure of classroom data. It also requires process evidence: lesson observations, fidelity ratings, teacher-use logs, AI-output audits, learner work samples, item-level achievement data, and delayed post-tests. Such evidence would allow future studies to distinguish immediate performance on a curriculum-aligned test from durable mathematical understanding and to examine whether teacher-mediated AI changes instruction rather than only the conditions around assessment.

7. Implications for practice and research

For practice, the findings support cautious development rather than immediate adoption. Agentic AI should be positioned as a teacher-facing resource for preparing explanations, varying examples, anticipating misconceptions, and supporting graph demonstrations, with the teacher retaining responsibility for curriculum alignment, mathematical correctness, representational choice, and learner need. Teacher development should therefore focus on disciplined mediation rather than generic tool operation. For research, the implication is categorical. Teacher-mediated AI should not be collapsed with learner-facing chatbots, autonomous tutoring systems, or general AI use. The unit of inquiry should be the teacher-AI-curriculum configuration and the classroom work through which AI output becomes, or fails to become, mathematics instruction. The present study is therefore best treated as a structured starting point for stronger designs, not as final evidence that the configuration has been causally established.

8. Conclusion

This study examined the educational effects of a teacher-mediated agentic AI instructional configuration on Grade 10 learners’ achievement in functions and graphs using an eight-week quasi-experimental design. The analysis identified a very large baseline-adjusted achievement contrast, substantial differences in mastery-threshold probabilities, and evidence that the contrast remained positive across the observed baseline-achievement distribution. The principal contribution of this study is the conceptualisation of teacher-mediated agentic AI as a distinct instructional configuration in which AI extends teachers’ pedagogical capacity without replacing professional judgement. Although the quasi-experimental design does not permit definitive causal claims, the findings indicate that this configuration warrants further investigation through larger, multi-teacher, process-rich studies capable of establishing the robustness and generalisability of its educational effects.

Ethics approval statement

Ethical approval for this study was obtained from the General/Human Research Ethics Committee (GHREC) of the University of the Free State before data collection (Ethics Clearance No. UFS-HSD2025/1903). Permission to conduct the study was also granted by the relevant education authorities and the participating schools. The study was conducted in accordance with the ethical principles outlined in the Declaration of Helsinki and the University’s research ethics guidelines.

Informed consent statement

Written informed consent was obtained from all participants before they participated in the study. For participants under 18 years of age, written informed consent was obtained from their parents or legal guardians, and assent was obtained from the learners themselves. Participation was entirely voluntary, and participants were informed of their right to withdraw from the study at any stage without penalty. All data were treated confidentially and anonymously and were used solely for research purposes.

AI usage declaration

Generative artificial intelligence (ChatGPT, OpenAI) and Grammarly were used solely to improve the language, grammar, clarity, and readability of this manuscript. All scientific content, including the conceptualisation, study design, development of the Mothusi AI-enhanced visualisation tool, data collection, data analysis, interpretation of the findings, and conclusions, was undertaken by the authors. The authors reviewed and edited all AI-assisted output and accept full responsibility for the content of the manuscript.

Availability of data and materials

The anonymised participant-level dataset supporting the findings of this study has been deposited in Zenodo and is publicly available at https://doi.org/10.5281/zenodo.21525667. (Dlamini, et al., 2026).

The repository contains the anonymised learner-level data used in the analyses, including Functions and Graphs Achievement Test (FGAT) scores, Functions and Graphs Interest Inventory (FGII) scores, instructional condition, and gender. All personal identifiers have been removed to protect participant confidentiality.

The data are available under the terms of the Creative Commons Attribution 4.0 International (CC BY 4.0) licence.

Acknowledgments

The authors thank the participating schools, teachers, learners, and colleagues whose support contributed to the successful completion of this study.

References
  •  Bloom BS: The 2 sigma problem: The search for methods of group instruction as effective as one-to-one tutoring. Educ. Res. 1984; 13(6): 4–16. Publisher Full Text
  •  Cheung ACK, Slavin RE: The effectiveness of educational technology applications for enhancing mathematics achievement in K-12 classrooms: A meta-analysis. Educ. Res. Rev. 2013; 9: 88–113. Publisher Full Text
  •  Chiu TKF, Xia Q, Zhou X, et al.: Systematic literature review on opportunities, challenges, and future research recommendations of artificial intelligence in education. Computers and Education: Artificial Intelligence. 2023; 4: 100118. Publisher Full Text
  •  Darvishi A, Khosravi H, Sadiq S, et al.: Impact of AI assistance on student agency. Comput. Educ. 2024; 210: 104967. Publisher Full Text
  •  Deng R, Jiang M, Yu X, et al.: Does ChatGPT enhance student learning? A systematic review and meta-analysis of experimental studies. Comput. Educ. 2025; 227: 105224. Publisher Full Text
  •  Dlamini LV, Mosia M, Egara FO: Dataset for: Teacher-Mediated, Not Learner-Facing: A Bayesian Quasi-Experimental Study of Agentic AI Support in Grade 10 Functions and Graphs. [Data set]. Zenodo. 2026. Publisher Full Text
  •  Drijvers P, Sinclair N: The role of digital technologies in mathematics education: purposes and perspectives. ZDM - Mathematics Education. 2024; 56(2): 239–248. Publisher Full Text
  •  Engelbrecht J, Borba MC: Recent developments in using digital technology in mathematics education. ZDM - Mathematics Education. 2024; 56: 281–292. Publisher Full Text
  •  Ertmer PA: Addressing first- and second-order barriers to change: Strategies for technology integration. Educ. Technol. Res. Dev. 1999; 47(4): 47–61. Publisher Full Text
  •  Hew KF, Brush T: Integrating technology into K-12 teaching and learning: Current knowledge gaps and recommendations for future research. Educ. Technol. Res. Dev. 2007; 55(3): 223–252. Publisher Full Text
  •  Hill JL: Bayesian nonparametric modeling for causal inference. J. Comput. Graph. Stat. 2011; 20(1): 217–240. Publisher Full Text
  •  Hillmayr D, Ziernwald L, Reinhold F, et al.: The potential of digital tools to enhance mathematics and science learning in secondary schools: A context-specific meta-analysis. Comput. Educ. 2020; 153: 103897. Publisher Full Text
  •  Holland PW: Statistics and causal inference. J. Am. Stat. Assoc. 1986; 81(396): 945–960. Publisher Full Text
  •  Huang AYQ, Lu OHT, Yang SJH: Effects of artificial intelligence-enabled personalised recommendations on learners’ learning engagement, motivation, and outcomes in a flipped classroom. Comput. Educ. 2023; 194: 104684. Publisher Full Text
  •  Hwang G-J, Xie H, Wah BW, et al.: Vision, challenges, roles and research issues of artificial intelligence in education. Computers and Education: Artificial Intelligence. 2020; 1: 100001. Publisher Full Text
  •  Kasneci E, Sessler K, Kuchemann S, et al.: ChatGPT for good? On opportunities and challenges of large language models for education. Learn. Individ. Differ. 2023; 103: 102274. Publisher Full Text
  •  Khosravi H, Shum SB, Chen G, et al.: Explainable artificial intelligence in education. Computers and Education: Artificial Intelligence. 2022; 3: 100074. Publisher Full Text
  •  Kulik JA, Fletcher JD: Effectiveness of intelligent tutoring systems. Rev. Educ. Res. 2016; 86(1): 42–78. Publisher Full Text
  •  Lai JWM, Bower M: How is the use of technology in education evaluated? A systematic review. Comput. Educ. 2019; 133: 27–42. Publisher Full Text
  •  Leinhardt G, Zaslavsky O, Stein MK: Functions, graphs, and graphing: Tasks, learning, and teaching. Rev. Educ. Res. 1990; 60(1): 1–64. Publisher Full Text
  •  Li Q, Ma X: A meta-analysis of the effects of computer technology on school students’ mathematics learning. Educ. Psychol. Rev. 2010; 22(3): 215–243. Publisher Full Text
  •  Liu M, Zhang LJ, Biebricher C: Investigating students’ cognitive processes in generative AI-assisted digital multimodal composing and traditional writing. Comput. Educ. 2024; 211: 104977. Publisher Full Text
  •  Lo CK, Hew KF, Jong MS-Y: The influence of ChatGPT on student engagement: A systematic review and future research agenda. Comput. Educ. 2024; 219: 105100. Publisher Full Text
  •  Long D, Magerko B: What is AI literacy? Competencies and design considerations. Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems. 2020; 1–16. Publisher Full Text
  •  Ma W, Adesope OO, Nesbit JC, et al.: Intelligent tutoring systems and learning outcomes: A meta-analysis. J. Educ. Psychol. 2014; 106(4): 901–918. Publisher Full Text
  •  Medina Herrera LM, Juarez Ordonez S, Ruiz-Loza S: Enhancing mathematical education with spatial visualisation tools. Frontiers in Education. 2024; 9: Article 1229126. Publisher Full Text
  •  Mishra P, Koehler MJ: Technological pedagogical content knowledge: A framework for teacher knowledge. Teachers College Record: The Voice of Scholarship in Education. 2006; 108(6): 1017–1054. Publisher Full Text
  •  Ng DTK, Leung JKL, Chu SKW, et al.: Conceptualising AI literacy: An exploratory review. Computers and Education: Artificial Intelligence. 2021; 2: 100041. Publisher Full Text
  •  Ouyang F, Jiao P: Artificial intelligence in education: The three paradigms. Computers and Education: Artificial Intelligence. 2021; 2: 100020. Publisher Full Text
  •  Roll I, Wylie R: Evolution and revolution in artificial intelligence in education. Int. J. Artif. Intell. Educ. 2016; 26(2): 582–599. Publisher Full Text
  •  Rosenbaum PR, Rubin DB: The central role of the propensity score in observational studies for causal effects. Biometrika. 1983; 70(1): 41–55. Publisher Full Text
  •  Rubin DB: Estimating causal effects of treatments in randomised and nonrandomised studies. J. Educ. Psychol. 1974; 66(5): 688–701. Publisher Full Text
  •  Schoenherr J, Schukajlow S: Characterising external visualisation in mathematics education research: A scoping review. ZDM. 2024; 56: 73–85. Publisher Full Text
  •  Schoenherr J, Strohmaier AR, Schukajlow S: Learning with visualisations helps: A meta-analysis of visualisation interventions in mathematics education. Educ. Res. Rev. 2024; 45: 100639. Publisher Full Text
  •  Steenbergen-Hu S, Cooper H: A meta-analysis of the effectiveness of intelligent tutoring systems on college students’ academic learning. J. Educ. Psychol. 2014; 106(2): 331–347. Publisher Full Text
  •  Tamim RM, Bernard RM, Borokhovski E, et al.: What forty years of research says about the impact of technology on learning. Rev. Educ. Res. 2011; 81(1): 4–28. Publisher Full Text
  •  ten Voorde A , Piroi M, Bos R: A taxonomy of didactic roles of dynamic visualisation in animated mathematics videos. Teaching Mathematics and its Applications: An International Journal of the IMA. 2025; 44(1): 47–67. Publisher Full Text
  •  Tondeur J, van Braak J , Sang G, et al.: Preparing pre-service teachers to integrate technology in education: A synthesis of qualitative evidence. Comput. Educ. 2012; 59(1): 134–144. Publisher Full Text
  •  VanLehn K: The relative effectiveness of human tutoring, intelligent tutoring systems, and other tutoring systems. Educ. Psychol. 2011; 46(4): 197–221. Publisher Full Text
  •  Voogt J, Fisser P, Pareja Roblin N, et al.: Technological pedagogical content knowledge: A review of the literature. J. Comput. Assist. Learn. 2013; 29(2): 109–121. Publisher Full Text
  •  Zawacki-Richter O, Marin VI, Bond M, et al.: Systematic review of research on artificial intelligence applications in higher education: Where are the educators?. Int. J. Educ. Technol. High. Educ. 2019; 16(1): Article 39. Publisher Full Text
  •  Zhang Y, Wang P, Jia W, et al.: Dynamic visualisation by GeoGebra for mathematics learning: A meta-analysis of 20 years of research. J. Res. Technol. Educ. 2025; 57(2): 437–458. Publisher Full Text
  •  Zhong B, Liu X: Evaluating AI literacy of secondary students: Framework and scale development. Comput. Educ. 2025; 227: 105230. Publisher Full Text

Схожие новости

#Наименование новостиТональностьИнформативностьДата публикации
1Conceptual Understanding and Conceptual Change Through Human AI Collaboration in Science Education: An Integrated Conceptual Framework from a Systematic Literature Review [version 1; peer review: awaiting peer review]04.8808-08-2026
2Human–AI Co-Regulation in Adaptive Learning: Developing GPT-Supported Self-Regulated Learning Models [version 1; peer review: awaiting peer review]010.7427-07-2026
3SmartFlex Learning Ecosystem: Integrating AI and Learning Analytics in Vocational Mathematics Education [version 1; peer review: 2 approved with reservations]07.8813-07-2026
4AI-Supported Literacy Ecosystems in Elementary Education: Preparing Future Skills for Lifelong Learning and Vocational Development [version 1; peer review: awaiting peer review]011.406-08-2026
5Productive and Unproductive AI Usage in Programming Education: Effects on Learning Behaviour, Creativity, and Performance with the Moderating Role of AI Literacy and Mindfulness [version 1; peer review: awaiting peer review]07.1723-07-2026
6Developing Computational Thinking in Science Education: A Systematic Review of Instructional Approaches, Learning Outcomes, and Implementation Challenges [version 1; peer review: awaiting peer review]06.8407-08-2026
7Bibliometric analysis of gamification in teacher education during the generative AI era: historical trends, present status, and future trajectories [version 2; peer review: 3 approved, 1 approved with reservations]011.6323-06-2026
8Teacher Enjoyment and Attitudes Toward Student Struggles in Mathematics: Evidence from Indonesian Elementary Schools [version 1; peer review: awaiting peer review]06.3807-08-2026
9In Rural Districts, AI Resources for Educators Are Scarce08.9210-07-2026
10The Impact of Artificial Intelligence in Enabling Digital Innovation/ The Mediating Role of Digital Transformation: A Survey Study Opinions of a Sample for Faculty Members at University of Ninevah [version 3; peer review: 3 approved]07.1323-07-2026

Классификация: . Схожих патентов: 0. Схожих новостей: 10. Тональность: 0. Информативность: 16.42. Источник: f1000research.com.