Вход на сайт

Просмотр новости

Найдите то, что Вас интересует

Machine learning algorithms to predict digital skills in university teaching staff. [version 2; peer review: 1 approved, 1 approved with reservations]

Дата публикации: 11-07-2026 07:04:46

Background The digital transformation of higher education has intensified the need to assess and enhance the digital competencies of university faculty. This study analyzed the effectiveness of various machine learning algorithms in predicting levels of faculty digital competence based on socio-educational variables. The objective was to develop an advanced predictive model, applied to faculty members from the State University of Milagro and the Technical University of Manabí. Methods A quantitative approach was adopted, with a cross-sectional correlational design. Digital competencies were measured using the internationally validated DigCompEdu Check-In instrument, structured across six core dimensions. In the predictive phase, nine supervised machine learning algorithms were trained and evaluated: logistic regression, decision trees, random forest, gradient boosting, k-nearest neighbors, support vector machines, stochastic gradient descent, artificial neural networks, and Naive Bayes. The models were trained using a dataset comprising 5,148 observations, and their performance was assessed using standard classification metrics: area under the ROC curve (AUC), accuracy, F1-score, sensitivity, and Matthew’s correlation coefficient (MCC). Results Gradient boosting, random forest, and neural network models demonstrated superior predictive performance, particularly at advanced competence levels (B2 and C1). Significant associations were identified between academic level, age, gender, and digital competencies. Logistic regression and Naive Bayes showed limitations in identifying low competence levels (A1), while intermediate levels were often overestimated across several models. Conclusions The findings confirm that machine learning algorithms can accurately predict university faculty digital competencies. Advanced models outperformed traditional ones, especially at higher competence levels. It is recommended to incorporate contextual variables and validate the models in diverse educational settings.

Основное содержимое страницы с новостью.

Introduction

The integration of digital technologies in higher education has led to a profound transformation in teaching methodologies and the development of teaching competencies (Benavides et al., 2020; García-Morales et al., 2021; Moreira-Choez et al., 2024d). This shift responds not only to technological evolution but also to a pedagogical need to adapt to new educational paradigms that demand more effective and dynamic interaction within digital environments. Studies such as those by Moreira-Choez et al. (2024e) and Lindfors et al. (2021) emphasize how digitalization has reshaped expectations and traditional teaching methods, making continuous assessment of faculty digital competencies imperative. Furthermore, the variability in perception and competency levels among instructors, widely documented in the literatura (Cattaneo et al., 2025; Tondeur et al., 2021), poses a significant challenge in terms of personalization and pedagogical effectiveness.

In this context, the application of machine learning algorithms offers transformative potential. The ability of these technologies to analyze large volumes of data and extract meaningful patterns can lead to accurate predictions about teachers’ digital competencies, thereby facilitating the design of more adaptive and effective professional training programs. According to Chen et al. (2020), the implementation of predictive models based on artificial intelligence has proven to enhance faculty adaptability to rapid and complex technological changes. Moreover, as Zhao et al. (2023) and Santamaria-Velasco et al. (2025), point out, these tools allow educational institutions to optimize resources and teaching strategies, ensuring better alignment with the needs and expectations of modern students. This technological and pedagogical convergence emerges as an essential pathway for advancing toward a higher education system that not only responds to technological imperatives but also promotes more inclusive and effective learning experiences (Rane, 2025; Wei, 2023).

Nonetheless, evaluating digital competencies presents significant challenges (Moreira-Choez et al., 2024c). Research by Garay-Rondero et al. (2024), highlights a noticeable deficiency in the availability of standardized tools to effectively assess these competencies across diverse academic disciplines. This lack of appropriate instruments hinders institutions’ ability to carry out accurate and consistent assessments of teachers’ digital skills. Additionally, current literature reveals a shortage of studies applying machine learning algorithms to effectively predict digital competencies in university settings, exposing a substantial gap in the existing body of research (Essa et al., 2023; Hidalgo et al., 2020). Therefore, research that combined a validated competence framework with transparent predictive modelling addressed a relevant methodological and applied need, particularly when the objective centered on supporting evidence-informed faculty development and strategic planning within higher education institutions.

Given this scenario, the critical importance of developing predictive models based on machine learning algorithms becomes evident. This study seeks to address these research gaps through the application of advanced modeling techniques. It aims to deliver accurate and personalized predictions that foster a deeper understanding of university faculty’s digital competencies. Implementing such models can improve both the efficiency and effectiveness of competency assessments, while also providing valuable data for designing more effective pedagogical interventions and supporting continuous professional development. This methodological approach represents a significant step forward in adapting higher education to 21st-century demands, ensuring that educators are equipped with the necessary skills to navigate and thrive in an increasingly digitalized educational landscape.

To address the predictive capacity of machine learning algorithms on faculty digital competencies, the following research question is posed: ¿How can machine learning algorithms predict digital competencies among faculty at the State University of Milagro and the Technical University of Manabí? To answer this question, several guiding hypotheses are established:

H1. Machine learning algorithms can effectively predict university faculty’s digital competencies based on socio-educational variables such as age, gender, teaching experience, and academic level.

H2. There is a significant relationship between the academic level of university faculty and their digital competence, with higher levels observed among those with advanced academic qualifications.

H3. Machine learning models based on gradient boosting and neural networks are more accurate in predicting digital competencies compared to simpler models such as logistic regression and random forest.

H4. Differences in digital competencies among teachers are significantly influenced by demographic factors such as gender and age, with specific variations in their capacity to adapt to new technologies.

H5. Digital competencies related to assessment and feedback in digital learning environments are more difficult to predict using machine learning algorithms due to their complex and multifactorial nature.

To address the research question, this study proposes the development of an advanced machine learning model to predict digital competencies in faculty at the State University of Milagro and the Technical University of Manabí. The model aims not only to evaluate the effectiveness of various modeling techniques in predicting digital competencies but also to facilitate the implementation of more effective and personalized educational interventions, tailored to the specific needs and characteristics of the faculty. This research is framed as part of a broader effort to optimize teacher training and improve teaching and learning processes within increasingly digitalized educational contexts.

Methods

This study adopted a quantitative approach with a cross-sectional and correlational design to investigate the effectiveness of machine learning algorithms in predicting digital competencies among university faculty at the State University of Milagro and the Technical University of Manabí. The relevance of this analysis lies in the growing imperative to effectively integrate digital technologies into teaching practices, which demands an accurate evaluation of faculty competencies in this domain. By implementing advanced modeling techniques, the study aimed to generate valuable insights to support professional development and promote continuous improvement in pedagogical practices.

The population of interest comprised faculty members with active appointments and teaching responsibilities at UNEMI and UTM during the data-collection period. According to the most recent institutional accountability information available, the faculty headcount corresponded to 1,562 professors at UNEMI and 1,169 faculty members at UTM (679 tenured and 490 occasional). Consequently, the reference population size corresponded to N = 2,731 faculty members across both institutions. Recruitment was conducted through an institutional survey invitation disseminated via Google Forms under a census-invitation approach (invitation addressed to the full eligible faculty pool), whereas participation remained voluntary. The final analytic sample comprised n = 234 surveyed faculty members. Relative to the reference population (N = 2,731), the overall response rate was estimated at 8.57%.

Below, Table 1 summarizes the a priori eligibility rules applied to delimit the analytic sample and to protect internal validity. The criteria were specified to ensure that observations corresponded to the intended target population (active faculty with teaching responsibilities), met ethical requirements (adult participants providing explicit informed consent), and satisfied minimal data-integrity standards for robust supervised classification.

Table 1. Inclusion and exclusion criteria. CategoryInclusion criteriaExclusion criteriaInstitutional eligibilityRecords corresponded to faculty members affiliated with the State University of Milagro (UNEMI) or the Technical University of Manabí (UTM) during the data-collection period.Records corresponded to individuals outside UNEMI/UTM or to non-faculty roles (e.g., administrative personnel), when identifiable.Teaching roleRespondents held teaching responsibilities at undergraduate and/or postgraduate level at the time of survey administration.Records lacked evidence of teaching responsibilities during the administration window, when identifiable.Age requirementRespondents were aged ≥18 years, as required for consent eligibility.Records corresponded to respondents aged <18 years, if any occurred.Informed consentRespondents provided digital informed consent embedded at the beginning of the online form prior to accessing the questionnaire items.Non-consenting cases were excluded by design; the form terminated automatically when consent was not granted.Uniqueness of submissionA single submission per participant was retained according to the deduplication rules defined in the analytical workflow.Duplicate or repeated submissions were removed according to the predefined deduplication rules (e.g., repeated entries under available identifiers and/or consistency checks).Data integrity and qualityRecords met the minimum completeness and internal-consistency requirements defined for analysis.Records were removed when they exhibited structural incompleteness, excessive missingness beyond the predefined threshold, or inconsistent entries under the quality-control rules, when applicable.

Table 1 substantiated methodological rigor by aligning the analytic dataset with the study’s inferential objective and by minimizing preventable sources of classification noise. First, eligibility restrictions limited the sample to faculty affiliated with UNEMI/UTM who held active teaching responsibilities, which increased construct-relevant variance and reduced contamination from non-target occupational roles. Second, the age and consent requirements ensured adherence to research ethics and strengthened procedural transparency, which supported the legitimacy and replicability of the data-generation process. Third, the enforcement of uniqueness and data-quality rules reduced the risk of biased model estimation attributable to duplicate submissions, structurally incomplete records, or internally inconsistent responses—conditions that can inflate apparent performance and distort class distributions in supervised learning.

Digital competence was measured with the DigCompEdu Check-In questionnaire adapted to Spanish-speaking contexts by Cabero-Almenara and Palacios-Rodríguez (2020). The instrument comprised 22 items that covered six competence areas, from professional engagement to the facilitation of students’ digital skills. Respondents’ scores were mapped to DigCompEdu proficiency levels ranging from Newcomer (A1) to Pioneer (C2). Data collection occurred through Google Forms and included an integrated informed-consent statement that described the study objectives, confidentiality safeguards, and voluntary participation. The consent statement appeared before the questionnaire items; access to the instrument required explicit authorization from participants aged 18 years or older. When consent was not provided, the form terminated automatically. The survey administration took place on February 28, 2025.

In addition to the predictive modelling component, inferential analyses evaluated bivariate associations between academic degree and each DigCompEdu item through Pearson’s chi-square tests applied to contingency tables derived from categorical response distributions. Statistical inference relied on a two-sided nominal significance threshold of α = 0.05. Because multiple tests were conducted across competence indicators, multiplicity was addressed through false discovery rate control via the Benjamini–Hochberg adjustment; therefore, significance decisions relied on adjusted p-values to limit inflation of type I error. Moreover, practical relevance was quantified with Cramer’s V for each association, which supported effect-size interpretation and facilitated comparisons across items with different marginal distributions.

For the machine-learning phase, the modelling dataset was organized in long, item-level format, which yielded 5,148 observations derived from 234 respondents × 22 items. Each record included socio-educational predictors (Age, Academic Level, Sex, Teaching Level) and a competence-area label, whereas the DigCompEdu proficiency level served as the multiclass target. The DigCompEdu instrument produced item-level responses across six competence areas; however, the predictive outcome relied on a single global proficiency label assigned to each respondent. Consequently, domain-specific proficiency levels were not modelled as separate targets; instead, the competence-area label was retained as a predictor to preserve contextual information at the item level. The observed outcome space comprised five categories (A1, A2, B1, B2, and C1). No category fusion or recoding was required for upper levels with sparse representation (e.g., C2), because those levels were not observed in the analytic dataset.

Prior to model estimation, integrity controls retained one submission per respondent under predefined deduplication rules, and records that did not satisfy minimum completeness or internal-consistency criteria were removed under the quality-control protocol. Missingness checks across modelling fields identified no missing entries in the exported training and evaluation dataset; therefore, no imputation was applied. Subsequently, socio-educational predictors were encoded as categorical variables. Age was represented as ordinal bands (20–30, 31–40, 41–50, 51–60, 61–70), academic attainment was coded as undergraduate, master’s degree, or PhD, sex was coded as a binary attribute, and teaching level was coded as undergraduate or postgraduate. For learners that required numeric inputs, categorical predictors were represented through one-hot encoding to ensure consistency across algorithms. Feature scaling was not applied because the encoded predictors entered the models as binary indicators; therefore, all inputs shared a common 0/1 scale, and additional normalization was not required for distance- or margin-based learners. No feature selection or dimensionality reduction was applied because the predictor set remained parsimonious and theory-aligned, and the modelling objective emphasized fair cross-algorithm comparison under a common input specification.

Model selection followed a comparative rationale that balanced interpretability, representational capacity, and robustness under heterogeneous predictor structures. Logistic regression and Naïve Bayes provided transparent baseline classifiers with probabilistic outputs that supported benchmarking. Decision trees contributed rule-based, non-parametric structure, whereas random forest and gradient boosting represented ensemble approaches designed to improve generalization through aggregation. In addition, support vector machines and k-nearest neighbors contributed complementary inductive biases based on margin separation and instance similarity, and stochastic gradient descent provided an efficient optimization-based learner suitable for high-dimensional indicator spaces. Artificial neural networks were included to capture nonlinearities and interactions that were less accessible to linear or additive specifications. Collectively, the nine-algorithm panel supported a principled assessment of performance trade-offs between traditional baselines and higher-capacity models under a unified evaluation protocol.

Finally, the modelling task was formulated as multiclass classification because the DigCompEdu outcome comprised discrete proficiency levels that supported a diagnostic interpretation aligned with institutional decision-making. Although ordinal regression could have leveraged the ordered nature of the competence scale, that approach required stronger structural assumptions (e.g., proportional odds) and could have constrained the representation of nonlinear patterns and cross-domain interactions. Likewise, multilevel models would have fit hierarchical structures if explicit nesting identifiers had been incorporated; however, the primary objective emphasized predictive discrimination across competence categories rather than variance partitioning across organizational strata.

Figure 1 illustrates the methodology employed in the study on the use of machine learning algorithms to predict digital competencies among university faculty. The diagram outlines the analytical workflow from data input to model evaluation, emphasizing the integration of various modeling techniques for a comprehensive interpretation.

6a9308e4-c407-4072-843d-87937881a37c_figure1.gif

Figure 1. Process flow for predicting digital competencies using machine learning algorithms.

Figure 1 presented the end-to-end modelling workflow used to predict faculty digital competence levels through a set of supervised classifiers, namely k-Nearest Neighbors (kNN), Logistic Regression, Random Forest, Naïve Bayes, Gradient Boosting, Support Vector Machines (SVM), Stochastic Gradient Descent (SGD), and Neural Networks. The pipeline was implemented in Orange Data Mining (v3.38.1), an open-source environment that supported visual specification of analytical procedures and facilitated procedural traceability. Accordingly, the workflow organized the sequence of data preprocessing, model training, validation, and post-hoc interpretation within a single, auditable framework.

Moreover, the distribution of the target variable was inspected before model estimation because the DigCompEdu outcome comprised multiple proficiency levels. A pronounced class imbalance was identified, with relatively low frequencies in A1 (n = 60) and A2 (n = 457) compared with B2 (n = 2,048) and C1 (n = 1,320), whereas B1 accounted for n = 1,263 (total = 5,148). No explicit rebalancing procedure was applied (e.g., class weighting, over-sampling, or under-sampling); therefore, all learners operated under the observed class proportions. Consequently, evaluation relied on metrics that remained informative under imbalance and included class-wise inspection through confusion matrices to characterize differential performance across competence strata.

In addition, model performance was not expressed solely through point estimates. Fold-level results from the cross-validation procedure were summarized as mean and standard deviation across folds for each algorithm and metric (AUC, accuracy, F1-score, recall, and Matthews correlation coefficient). Furthermore, 95% confidence intervals for the cross-validated mean were derived from the fold-level distribution through a t-based approach. Therefore, comparisons across algorithms considered both central tendency and dispersion, and small differences in AUC or F1 received conservative interpretation when uncertainty intervals overlapped.

Subsequently, performance estimation used stratified k-fold cross-validation implemented through the Test and Score widget, which preserved the empirical class distribution within each fold. Fold-wise estimates were aggregated to produce a single performance profile per learner. The evaluation used a single run under the software’s default randomization settings; consequently, performance estimates did not reflect averaging across multiple random seeds. In parallel, the Confusion Matrix widget provided class-specific diagnostics, which complemented global indices and supported interpretation of error patterns across proficiency levels.

Furthermore, all learners were trained under default Orange widget configurations, and no hyperparameter optimization routine (e.g., grid search or Bayesian tuning) was executed. As a result, the comparative results reflected baseline performance under a common evaluation protocol rather than tuned upper-bound performance. To support reproducibility, the principal hyperparameters exposed by each learner widget were documented in Table 2.

Table 2. Default hyperparameter configuration used for model training in Orange. AlgorithmKey hyperparameters (Orange default widget settings)Logistic RegressionL2 (ridge) regularization; C = 1Decision TreeBinary tree; min instances in leaves = 2; do not split subsets smaller than 5; max depth = 100; stop when majority reaches 95%Random ForestNumber of trees = 10; minimum subset size = 5; other options remained at widget defaults (e.g., attribute subsampling and depth limit not forced)Gradient Boostingscikit-learn method; number of trees = 100; learning rate = 0.10; max depth = 3; minimum subset size = 2; subsampling fraction = 1.00kNNNeighbors k = 5; metric Euclidean; weights UniformSVMKernel RBF; cost C = 1.00; gamma auto; tolerance = 0.001; iteration limit = 100SGDLoss hinge; L2 regularization; alpha = 0.00001; learning rate constant; η0 = 0.01; iterations = 1000; tolerance = 0.001; shuffling enabledNeural NetworkHidden-layer specification (100); activation ReLU; solver Adam; alpha = 0.00010; max iterations = 200Naïve BayesDefault learner settings (no tunable hyperparameters exposed in widget) (orange3.readthedocs.io)

Finally, the comparative analysis of models including logistic regression, random forest, gradient boosting, and neural networks revealed differential capacities in forecasting the digital competencies of university faculty. These algorithms were selected due to their scalability in managing high-dimensional data and their demonstrated efficacy in similar educational research contexts. The results contributed not only to the identification of the most performant models but also to the formulation of evidence-based recommendations for professional development programs, aiming to enhance digital proficiency among higher education instructors.

Ethical considerations

In accordance with the ethical principles governing research involving human subjects, this study implemented a rigorous procedure to obtain informed consent, ensuring that all participants fully understood the nature, objectives, and implications of the study. It was clearly communicated that participation was entirely voluntary and that individuals were free to decline or withdraw from the research process at any time, without facing any negative consequences. To preserve confidentiality, personal data were anonymized, thus preventing any form of direct or indirect identification of the participants involved. Formal approval for this research was granted by the Institutional Review Board (IRB) of Milagro State University, as documented in the official resolution UNEMI-VICEINVYPOSG-DP-233-2025-OF, dated February 14, 2025.

To ensure scientific validity and the reproducibility of findings, the study was conducted in accordance with methodological guidelines that promote transparency throughout all phases of the research process. The adoption of these standards aims to strengthen the credibility of quantitative research by enabling the verification and replication of results by other scholars in similar contexts. Accordingly, standardized protocols were employed for both data collection and analysis, ensuring a rigorous and systematic approach free from subjective interference. This methodological strategy minimized potential biases and fostered an objective interpretation of the findings, in alignment with the principles of scientific integrity.

Results and discussion

This section of the study presents findings related to the predictive capacity of machine learning algorithms in assessing the digital competencies of university faculty. The analysis explored how various socio-educational factors influence these competencies, using a dataset representative of the teaching population at the State University of Milagro and the Technical University of Manabí. Table 3 provides a detailed statistical analysis examining how different elements of digital competencies correlate with the academic level of university faculty. Competencies are categorized into key areas such as professional engagement, digital resources, digital pedagogy, assessment and feedback, and learner empowerment. The Chi-square (χ2) values and p-values provide evidence of the statistical significance of the observed relationships.

Table 3. Relationship between digital competencies and academic level of university faculty.CompetenciesItemsχ2 pProfessional Engagement I systematically use different digital channels to improve communication with students, families, and colleagues. For example: email, messaging apps like WhatsApp, blogs, school websites.19,2130,004I use digital technologies to collaborate with my colleagues both within and outside my educational organization.91,290,000I actively develop my own digital competence.16,5710,011I participate in online training courses. For example: government-provided online courses, MOOCs, webinars, etc.51,4350,000Digital Resources I use different websites and search strategies to find and select a wide range of digital resources.27,0480,000I create my own digital resources and adapt existing ones to meet my teaching needs.66,8510,000I securely protect sensitive content. For example: exams, grades, personal data.58,1430,000Digital Pedagogy I carefully consider how, when, and why to use digital technologies in class to ensure their added value is realized.31,5910,000I monitor student activities and interactions in the online collaboration environments we use.50,2950,000When my students work in groups or teams, they use digital technologies to acquire and document knowledge.34,7550,000I use digital technologies to enable students to plan, document, and evaluate their own learning. For example: self-assessment tests, digital portfolios, blogs, forums.95,920,000Assessment and Feedback I use digital assessment strategies to monitor students’ progress.30,7520,000I analyze all available data to identify students who need additional support. “Data” includes: student participation, performance, grades, attendance, activities, and social interactions in online environments. “Students who need additional support” refers to those at risk of dropout, low performance, learning disorders, specific learning needs, or lacking transversal skills (social, verbal, or study skills).17,9230,022I use digital technologies to provide effective feedback.56,4590,000Student Empowerment When proposing digital tasks, I consider and address potential issues such as equal access to devices and digital resources; compatibility problems or students’ low digital competence.31,2210,000I use digital technologies to offer students personalized learning opportunities. For example: assigning different digital tasks to address individual learning needs, taking into account preferences and interests.45,0480,000I use digital technologies to ensure active student participation in class.45,530,000Facilitating Students’ Digital Competence I teach students how to evaluate the reliability of online information and to identify false and/or biased information.49,3830,000I propose tasks that require students to use digital media to communicate and collaborate with each other or with an external audience.58,1950,000I propose tasks that require students to create digital content. For example: videos, audio recordings, photos, presentations, blogs, wikis.57,3950,000I teach students how to behave safely and responsibly online.34,7520,000I encourage students to use digital technologies creatively to solve specific problems. For example, overcoming obstacles or emerging challenges in their learning process.24,1010,001

The results in the Table 3 professional engagement category which include activities such as the use of various digital communication channels and participation in online training show statistically significant differences across academic levels, with particularly high Chi-square values. These findings are consistent with previous studies, such as those by Al-Rahmi et al. (2023), which indicate a correlation between academic level and the adoption of digital technologies in professional communication. Similarly, Maican et al. (2019) emphasize that faculty members with higher academic qualifications are more likely to employ advanced technologies for collaboration and communication, suggesting that research experience and training may influence both the ease and frequency with which new digital tools are adopted.

In the area of digital resources and digital pedagogy, the results reveal that activities such as creating personalized digital resources and supervising students in online collaborative environments vary significantly according to academic level. Bond et al. (2018) corroborate in their research that postgraduate-level instructors tend to integrate more complex technologies into their teaching methodologies, which is reflected in a greater predisposition to modify and adapt digital materials. Furthermore, Haleem et al. (2022) support the idea that higher academic levels are associated with more innovative pedagogical practices and a more strategic use of digital technologies to enhance learning processes.

In turn, digital assessment and feedback also show a strong association with academic level. Faculty members holding postgraduate and doctoral degrees use more sophisticated digital assessment strategies, aligning with the findings of Wang et al. (2021) who suggest that competence in digital assessment is greater among those with advanced academic training. This may be attributed to their greater exposure to learning environments that require and value precision in student performance evaluation and monitoring. Finally, the student empowerment category highlights how educators use digital technologies to personalize and enrich learning experiences. These results are supported by the research of Christodoulou and Angeli (2022), who found that faculty members with higher academic qualifications are more likely to employ digital technologies creatively to address pedagogical challenges, providing students with tools that foster autonomous and adaptive learning.

Table 4 presents a comprehensive analysis of the performance of various machine learning algorithms in predicting digital competencies, categorized by different competence levels ranging from A1 (Newcomer) to C1 (Leader). This breakdown facilitates an understanding of how each model performs in relation to the specific competence level of faculty members, offering critical insight into the effectiveness of each modeling technique.

Table 4. Evaluation of machine learning models by competence level of university faculty.ModelAUCCAF1PrecRecallMCCLevelLogistic Regression0.9080.9880.0000.0000.0000.000A1 (Newcomer)Tree0.9200.9880.0000.0000.0000.000Gradient Boosting0.9200.9880.0000.0000.0000.000Random Forest0.9220.9880.0000.0000.0000.000kNN0.5180.9880.0000.0000.0000.000SVM0.7840.9880.0000.0000.0000.000SGD0.5000.9880.0000.0000.0000.000Neural Network0.9210.9880.0000.0000.0000.000Naive Bayes0.6930.9880.0000.0000.0000.000ModelAUCCAF1PrecRecallMCCLevelLogistic Regression0.8320.9010.1790.3350.1230.159A2 (Explorer)Tree0.8520.9120.3110.5150.2230.300Gradient Boosting0.8510.9120.3110.5150.2230.300Random Forest0.8520.9120.3110.5150.2230.300kNN0.6520.8790.3010.3090.2930.235SVM0.5880.5550.1860.1110.5730.072SGD0.5280.8980.1200.2540.0790.098Neural Network0.8520.9120.3110.5150.2230.300Naive Bayes0.7860.8870.2170.2810.1770.165ModelAUCCAF1PrecRecallMCCLevelLogistic Regression0.6690.6920.1410.2240.103−0.018B1 (Integrator)Tree0.7490.7630.4740.5210.4350.325Gradient Boosting0.7500.7630.4700.5220.4270.322Random Forest0.7500.7630.4700.5220.4270.322kNN0.6260.6570.3900.3450.4470.158SVM0.5990.6610.2400.2670.2190.025SGD0.5010.6880.1740.2490.1340.003Neural Network0.7490.7630.4740.5210.4350.325Naive Bayes0.6300.6580.1580.1990.131−0.047ModelAUCCAF1PrecRecallMCCLevelLogistic Regression0.6360.5030.5490.4300.7620.101B2 (Expert)Tree0.6910.5840.5840.4850.7350.220Gradient Boosting0.6830.5840.5800.4840.7220.213Random Forest0.6910.5870.5830.4870.7280.220kNN0.5570.5550.4740.4470.5040.091SVM0.5170.5570.3680.4260.3240.038SGD0.5560.5320.5340.4420.6750.113Neural Network0.6890.5840.5840.4850.7350.220Naive Bayes0.6200.5580.5430.4610.6600.149ModelAUCCAF1PrecRecallMCCLevelLogistic Regression0.7290.7550.3970.5390.3140.271C1 (Leader)Tree0.7590.7750.4510.6010.3610.337Gradient Boosting0.7580.7700.4580.5790.3790.332Random Forest0.7600.7760.4700.5970.3880.349kNN0.6210.7210.3130.4240.2480.161SVM0.6490.7280.0780.2980.0450.019SGD0.6200.7210.4310.4520.4110.247Neural Network0.7590.7750.4510.6010.3610.337Naive Bayes0.7190.7360.4380.4820.4020.269

Table 4 reveals notable variability in model performance across different competence levels, with metrics such as AUC (Area Under the Curve), CA (Classification Accuracy), F1-score, Precision (Prec), Recall, and MCC (Matthews Correlation Coefficient). Moreover, the uncertainty intervals indicated that several algorithms exhibited comparable performance, particularly when differences in AUC or F1 remained small relative to the fold-wise variability. Therefore, rankings based solely on point estimates were not interpreted as definitive, and preference was given to models that combined competitive mean performance with tighter dispersion across folds. For instance, at the A1 (Newcomer) level, nearly all models except for Logistic Regression failed to accurately classify competencies, as evidenced by F1, Precision, and Recall values all being zero. This may suggest that the models are encountering difficulties in identifying distinguishing features at this initial level of competence, a finding consistent with studies by Smith and Zárate (1992) and Bansal et al. (2007), who observed that simple models often underperform in contexts where target categories are homogeneous or underrepresented.

Moreover, the absence of a formal calibration assessment limited the extent to which predicted probabilities could support operational decision rules. Therefore, practical deployment should prioritize models that combined competitive discrimination with stable cross-validation performance, and it should incorporate post-hoc calibration and threshold optimization before any individual-level screening or targeted intervention was implemented. This consideration was particularly relevant under class imbalance, where uncalibrated probability outputs can overstate confidence for minority categories and distort prioritization decisions. As we move toward higher levels of competence, such as B2 (Expert) and C1 (Leader), some models particularly Random Forest and Gradient Boosting demonstrate substantial improvements in metrics like F1-score and Recall. This indicates that these models may be more suitable for contexts where inter-class differences are more pronounced and where a higher degree of discrimination is required, as noted by Asselman et al. (2023). Furthermore, the increase in MCC at higher levels suggests that these models are effective in balancing sensitivity and specificity, thereby providing more accurate and balanced predictions.

However, it is critical to note that, overall, models struggle to achieve high levels of accuracy at the lowest competence level (Newcomer), which may reflect a limitation in their ability to handle undifferentiated input data or features that are not clearly defined. This phenomenon underscores the importance of ongoing development and optimization of machine learning algorithms capable of effectively managing a broader range of data complexities (Taye, 2023; Zhou et al., 2017). Moreover, advancing toward the incorporation of more sophisticated or hybrid modeling approaches is essential, particularly those capable of capturing nonlinear relationships and latent structures in educational datasets specially in segments where digital competencies are emerging or weakly differentiated. As Asselman et al. (2023) argue, the use of advanced models that integrate deep learning architectures or ensemble mechanisms can enhance the detection of subtle patterns, thereby enabling more accurate diagnostics and, consequently, the design of targeted, evidence-based pedagogical interventions. Such approaches also contribute to the development of adaptive systems that respond more effectively to the diversity of teaching profiles present within the higher education context.

Figure 2 visually illustrates the correlation between teachers’ digital competence levels and various variables such as age, gender, teaching experience, and academic level. This graphical representation facilitates the understanding of trends and patterns emerging from the analyzed data.

6a9308e4-c407-4072-843d-87937881a37c_figure2.gif

Figure 2. Competence level and type of competence according to socio-educational variables.

Note. The figure presents the level and type of teaching competence according to different socio-educational variables across six dimensions of digital competence.

Figure 2 presents a detailed visualization of university faculty members’ digital competence levels based on various socio-educational variables, allowing for the identification of relevant patterns in the distribution of digital skills. A significant concentration is observed at intermediate and advanced levels (B1, B2, and C1), particularly in competencies related to digital pedagogy, the use of digital resources, and feedback. This trend suggests that a considerable portion of faculty has developed key competencies necessary for the integration of technology into the teaching-learning process. However, substantial gaps persist at the basic levels (A1 and A2), indicating that segments of the academic population still require intensive support to reach satisfactory levels. This phenomenon is well-documented in the literature, where it is noted that digital inequalities continue to be reproduced, especially in institutions with pedagogical cultures centered on traditional methodologies (Schmidt & Tang, 2020).

The distribution by age reveals that younger faculty members tend to be positioned at higher levels of digital competence, which may be associated with greater exposure to and familiarity with technological tools during their initial training. These findings reinforce those of Alcaide-Pulido et al. (2025), who identified a positive correlation between age and technological adaptability. Regarding gender, the figure shows slight disparities that may reflect structural inequalities in access to technological training opportunities, as highlighted in previous studies on gender gaps in digital competencies (Stoet & Geary, 2018). Such differences underscore the need to incorporate an equity perspective into the design of professional development policies.

Academic level also emerges as a determining factor in the development of digital competencies. Faculty with postgraduate education, particularly those holding a doctoral degree, show a higher concentration at the B2 and C1 levels compared to those with only undergraduate training. This finding aligns with the results of Mei et al. (2019), who emphasized that higher academic qualifications are associated with a greater ability to critically and effectively incorporate digital technologies in educational settings. Thus, academic trajectory constitutes a key predictor of technological proficiency in teaching practice.

In light of these results, the need to design institutional strategies for continuous professional development is reinforced strategies that take into account both the initial level of competence and the socio-demographic characteristics of faculty members. It is not enough to promote the acquisition of technological tools; rather, a methodological reconfiguration is required to support a transition toward pedagogical approaches based on innovation and flexibility. In this regard, Rofi’i et al. (2023) argue that training programs must be adaptive and context-sensitive, while Vindigni (2023) emphasizes the importance of incorporating inclusive approaches that respond to educators’ personal and professional trajectories.

Figure 3 provides an in-depth visualization of how different machine learning models classify various digital competencies across different competence levels. Each panel within the figure represents a detailed comparison between types of competence and their distribution according to the model used, highlighting both accuracy and areas for improvement in the prediction of teachers’ specific digital skills.

6a9308e4-c407-4072-843d-87937881a37c_figure3.gif

Figure 3. Performance of predictive models according to competence level and type of competence.

Note. The figure illustrates the performance of different predictive models according to competence level and type of competence across six digital competence dimensions.

Figure 3 provides a comparative view of the performance of various machine learning algorithms in predicting digital competencies, allowing for the observation of distinct patterns based on the type of competence evaluated. Overall, a higher density of accurate predictions is observed at advanced competence levels (B2 and C1), particularly in categories related to digital pedagogy, digital resources, and feedback, when using models such as Random Forest, Gradient Boosting, and Neural Networks. This trend can be attributed to the robustness of these algorithms in handling nonlinear relationships and complex data structures, which aligns with the findings of Dong et al. (2021), who argue that such models achieve better generalization when there is clear differentiation among latent classes.

In contrast, lower accuracy is observed at the more basic levels (A1 and A2), especially in models such as k-Nearest Neighbors (kNN) and Stochastic Gradient Descent (SGD), where predictions tend to be more dispersed and less aligned with actual values. This variability may be due to these models’ sensitivity to high-dimensional datasets or class imbalance, as noted by Kumar et al. (2023). Furthermore, the irregular performance of kNN in competencies such as student empowerment or facilitation of digital competence suggests that its effectiveness diminishes when categories have conceptual overlaps or lack well-defined structures. This observation highlights the importance of aligning the model with the specific type of competence being evaluated, recognizing that not all machine learning approaches yield consistent performance across all competency domains. Moreover, the observed imbalance across competence levels likely constrained the effective learning signal for the lowest categories, which reduced the models’ ability to delineate minority-class decision boundaries with comparable precision to the dominant classes.

It is worth emphasizing that ensemble-based models, such as Random Forest and Gradient Boosting, not only offer higher predictive power but also provide greater stability across the various assessment areas of the DigCompEdu questionnaire. Their consistent performance across all dimensions indicates a superior ability to capture complex patterns and interactions among variables, which is crucial in diverse educational contexts. This characteristic is especially relevant for educational interventions, as it enables more accurate diagnostics and, therefore, the design of improvement strategies tailored to the specific digital profile of the instructor. Alghamdi et al. (2025) argue that the practical utility of predictive models lies in their ability to precisely identify individual weaknesses, thereby facilitating the implementation of more targeted and effective training plans.

Figure 4 presents a comparative analysis of the impact of various socio-educational variables on the performance of four different machine learning models, enabling a detailed assessment of how factors such as age, gender, and academic level influence digital competence predictions in educational settings. This comparison highlights the variability in the influence of these variables across models, underscoring the inherent complexity of modeling digital competencies in the educational domain.

6a9308e4-c407-4072-843d-87937881a37c_figure4.gif

Figure 4. Comparative Analysis of the Impact of Socio-Educational Variables on Four Machine Learning Models.

Note. The figure compares the impact of socio-educational variables on the outputs of four machine learning models across different competence dimensions.

Figure 4 shows that middle-aged and older age groups specifically those between 41–50 and 51–60 years as well as higher academic levels such as Ph.D. and Master’s degree holders, have a positive impact on most of the models analyzed. These findings align with current literature, such as the study by Thordsen and Bick (2023), which highlights the correlation between academic maturity and greater technological integration. This pattern suggests that professional experience and prolonged development may facilitate more efficient and strategic adoption of technological tools in teaching.

On the other hand, the analysis reveals significant differences in the impact of gender variables across the models, reflecting potential variations in technology access and usage between men and women in educational environments. Studies by Choudhary (2024) support this observation, arguing that training programs should be tailored to address these differences, ensuring that digital competence interventions are inclusive and effective. Moreover, the results indicate a negative impact of the variables 20–30 years and Undergraduate in several model configurations, suggesting a deficiency in technological training at the early stages of academic careers. This challenge is recognized in the research of Chohan and Hu (2022), who propose strengthening digital education from initial training levels to mitigate competence gaps from the outset.

Figure 5 provides a visual analysis using violin plots that illustrate the distribution of advanced (leader-level) digital competencies based on the academic level of university faculty. This visualization enables a comparison of the density and dispersion of competencies across three distinct academic categories: Undergraduate, Master’s Degree, and Ph.D.

6a9308e4-c407-4072-843d-87937881a37c_figure5.gif

Figure 5. Distribution of leader-level competencies according to academic level.

Note. The figure shows the distribution of leader-level competencies across academic levels as estimated by different machine learning models.

Figure 5 the violin plots illustrate notable differences in the distribution of leader-level competencies across academic levels. Faculty members holding Ph.D. degrees display a wider and more symmetrical distribution, indicating greater uniformity and breadth in their digital competencies. This finding supports the research by Palacios-Rodríguez et al. (2024), which suggests that instructors with higher educational qualifications tend to possess more developed digital skills due to increased exposure to advanced technologies and research methodologies during their doctoral training.

In contrast, the plots for Undergraduate and Master’s Degree levels show narrower and more asymmetrical distributions, which may indicate variability in digital competence levels within these groups. This could reflect less consistency in the integration of digital technologies into undergraduate and graduate curricula, a point emphasized by Tee et al. (2024), who argue that disparities in technological training during the early years of higher education can lead to significant gaps in digital competencies. Additionally, the comparison of the tails of the plots shows that, although some individuals at the Undergraduate and Master’s levels achieve competence levels comparable to Ph.D. holders, the majority are concentrated in lower ranges. This pattern is consistent with findings by Radovan and Radovan (2024), who observed that opportunities for ongoing professional development and access to technological resources are less frequent at these educational levels, potentially limiting the development of advanced competencies.

Figure 6 presents a detailed comparison of the calibration of various machine learning models, including logistic regression, decision trees, and gradient boosting, among others. This visual analysis demonstrates how each model predicts digital competencies and evaluates the accuracy of these predictions by aligning them with actual observations.

6a9308e4-c407-4072-843d-87937881a37c_figure6.gif

Figure 6. Calibration of models at the levels of digital competencies.

Note. The figure presents the calibration curves of different predictive models across the levels of digital competence, illustrating the relationship between predicted probabilities and observed outcomes for each target class.

Figure 6 the neural network and gradient boosting models, which are closer to the diagonal line in the calibration plots, indicate greater accuracy in probability estimation. This result suggests that such models are better suited to handling the inherent complexities of educational data. The effectiveness of these advanced models in predicting digital competencies is supported by studies such as Fakhar et al. (2024), who highlight their superior capacity to adapt to complex patterns and interdependent variables in educational settings. This perspective is further reinforced by Alnasyan et al. (2024), who argue that deep learning techniques, such as neural networks, are particularly effective in capturing nonlinear and multifaceted dynamics that often characterize educational data. Moreover, the study by Beer and Mulder (2020) illustrates how the application of gradient boosting has improved the prediction of educational outcomes by fitting models that consider a wide range of influential factors, thus demonstrating the importance of using sophisticated methods in educational assessment.

In contrast, simpler models notably Logistic Regression, Decision Trees, and k-Nearest Neighbors (kNN) display significant divergence from the diagonal. This misalignment reflects underfitting, where the models fail to capture the underlying structures in the data. As emphasized by Islam et al. (2025), traditional models often lack the flexibility required to adapt to the multifaceted nature of educational data, particularly when the classification task involves nuanced levels of digital proficiency. Moreover, Awedh and Mueen (2025) point out that such models are generally ineffective at identifying latent patterns or dealing with overlapping class boundaries, thereby reducing their predictive power and practical utility in educational assessment.

Table 5 presents a detailed comparison of the main machine learning algorithms used to predict levels of digital competence among faculty. Each confusion matrix displays the proportions of classifications made by the models in relation to the actual and predicted competence levels, ranging from A1 (Newcomer) to C1 (Leader). The analysis includes models such as logistic regression, decision trees, gradient boosting, random forest, k-nearest neighbors (kNN), support vector machines (SVM), stochastic gradient descent (SGD), neural networks, and naive Bayes.

Table 5. Comparative confusion matrix of predictive models for faculty digital competence levels.Logistic Regression ModelDecision Tree Model Actual Level A1 (Newcomer) A2 (Explorer) B1 (Integrator) B2 (Expert) C1 (Leader)Σ Actual Level A1 (Newcomer) A2 (Explorer) B1 (Integrator) B2 (Expert) C1 (Leader)ΣA1 (Newcomer)NA2.4%1.9%1.2%0.0%60A1 (Newcomer)NA0.0%0.6%1.8%0.0%60A2 (Explorer)NA33.5%28.6%6.5%0.0%457A2 (Explorer)NA51.5%7.2%9.0%0.0%457B1 (Integrator)NA37.7%22.4%27.7%8.4%1,263B1 (Integrator)NA29.3%52.1%18.8%9.1%1,263B2 (Expert)NA17.4%29.1%43.0%37.7%2,048B2 (Expert)NA19.2%24.7%48.5%30.8%2,048C1 (Leader)NA9.0%18.1%21.6%53.9%1,320C1 (Leader)NA0.0%15.5%21.9%60.1%1,320Σ01675813,6307705,148Σ01981,0563,1027925,1Decision Tree ModelRandom Forest ModelActual LevelA1 (Newcomer)A2 (Explorer)B1 (Integrator)B2 (Expert)C1 (Leader)ΣActual LevelA1 (Newcomer)A2 (Explorer)B1 (Integrator)B2 (Expert)C1 (Leader)ΣA1 (Newcomer)NA0.0%0.5%1.8%0.0%60A1 (Newcomer)NA0.0%0.5%1.8%0.0%60A2 (Explorer)NA51.5%7.4%9.1%0.0%457A2 (Explorer)NA51.5%7.4%9.1%0.0%457B1 (Integrator)NA29.3%52.2%19.1%9.5%1,263B1 (Integrator)NA29.3%52.2%19.3%8.8%1,263B2 (Expert)NA19.2%24.2%48.4%32.6%2,048B2 (Expert)NA19.2%24.2%48.7%31.5%2,048C1 (Leader)NA0.0%15.8%21.5%57.9%1,320C1 (Leader)NA0.0%15.8%21.1%59.7%1,320Σ01981,0333,0548635,148Σ01981,0333,0608575,148k-Nearest Neighbors (kNN) ModelSVM ModelActual LevelA1 (Newcomer)A2 (Explorer)B1 (Integrator)B2 (Expert)C1 (Leader)ΣActual LevelA1 (Newcomer)A2 (Explorer)B1 (Integrator)B2 (Expert)C1 (Leader)ΣA1 (Newcomer)NA0.2%1.4%1.3%0.8%60A1 (Newcomer)NA1.2%0.6%1.7%0.0%60A2 (Explorer)NA30.9%11.1%3.9%6.6%457A2 (Explorer)NA11.1%8.8%6.7%0.9%457B1 (Integrator)NA27.7%34.5%20.3%14.1%1,263B1 (Integrator)NA27.1%26.7%20.3%16.7%1,263B2 (Expert)NA31.6%36.7%44.7%36.1%2,048B2 (Expert)NA39.2%34.2%42.6%53.5%2,048C1 (Leader)NA9.5%16.2%29.7%42.4%1,320C1 (Leader)NA21.4%29.8%28.9%29.8%1,320Σ01981,0333,0548635,148Σ02,3571,0351,5581985,148SGD ModelNeural Network ModelActual LevelA1 (Newcomer)A2 (Explorer)B1 (Integrator)B2 (Expert)C1 (Leader)ΣActual LevelA1 (Newcomer)A2 (Explorer)B1 (Integrator)B2 (Expert)C1 (Leader)ΣA1 (Newcomer)NA2.1%2.5%1.1%0.6%60A1 (Newcomer)NA0.0%0.5%1.8%0.0%60A2 (Explorer)NA25.4%24.9%6.5%4.1%457A2 (Explorer)NA51.5%7.2%9.0%0.0%457B1 (Integrator)NA22.5%24.9%27.9%15.7%1,263B1 (Integrator)NA29.3%52.1%18.9%9.1%1,263B2 (Expert)NA28.9%31.2%44.2%34.4%2,048B2 (Expert)NA19.2%24.7%48.5%30.8%2,048C1 (Leader)NA21.1%16.6%20.3%45.2%1,320C1 (Leader)NA0.0%15.5%21.9%60.1%1,320Σ01426803,1281,1985,148Σ01981,0563,1027925,148Naive Bayes ModelActual LevelA1 (Newcomer)A2 (Explorer)B1 (Integrator)B2 (Expert)C1 (Leader) ΣA1 (Newcomer)NA1.7%2.1%1.3%0.0%60A2 (Explorer)NA28.1%25.6%5.6%0.0%457B1 (Integrator)NA40.6%19.9%27.0%17.3%1,263B2 (Expert)NA20.1%31.2%46.1%34.5%2,048C1 (Leader)NA9.4%21.3%20.0%48.2%1,320Σ02888282,9321,1005,148

The results in Table 5 show that the Gradient Boosting, Random Forest, and Neural Network models demonstrated a higher ability to correctly predict the upper levels of digital competence (B2 and C1), particularly in the case of the Gradient Boosting model, which achieved a correct classification rate of 57.9% for instructors at the C1 (Leader) level. This performance aligns with previous findings highlighting the effectiveness of these models in handling nonlinear relationships and complex data structures (Kukkar et al., 2019; Kyriazos & Poga, 2024). In contrast, the Naive Bayes model showed clear limitations, particularly at the lower levels, systematically underestimating actual competence levels. This is consistent with studies indicating its limited capacity to capture interactions among predictive variables in complex educational contexts (Almalawi et al., 2024; Sharma et al., 2019).

A general trend of overprediction at intermediate levels (B1 and B2) was also observed across most algorithms, suggesting ambiguity in classifying instructors with mid-range competence profiles. This phenomenon has been reported by Moreira-Choez et al. (2024b), who note that intermediate levels of digital competence tend to exhibit greater variability and dispersion based on factors such as age, gender, and teaching experience. In this regard, the robustness of more advanced models allowed for better capture of transitions between levels and more accurate differentiation of instructors at the threshold of digital competence advancement.

Ultimately, these findings reinforce the importance of employing high-complexity models for predictive diagnostic tasks in educational environments where multidimensional constructs such as digital competencies are being analyzed. The comparison of confusion matrices clearly demonstrates that models like Gradient Boosting and Random Forest not only offer higher overall accuracy but also reduce misclassification errors between adjacent levels. This contributes to more effective planning of personalized training interventions (Ingkavara et al., 2022; Moreira-Choez et al., 2024a). These results provide valuable empirical evidence to support future applications of artificial intelligence in faculty assessment throughout Latin America.

Conclusions

The study met its objective of developing and evaluating supervised machine learning models to predict DigCompEdu proficiency levels among faculty from the State University of Milagro (UNEMI) and the Technical University of Manabí (UTM). Moreover, the comparative evaluation indicated that higher-capacity approaches, particularly Gradient Boosting, Random Forest, and Neural Networks, achieved stronger discrimination for intermediate and advanced competence profiles, with comparatively better classification in B2 and C1 than in the lower levels. This pattern aligned with the confusion-matrix evidence reported for the competing algorithms and supported the use of ensemble and neural approaches as competitive candidates for diagnostic modelling under heterogeneous socio-educational predictors. However, the empirical scope remained bounded by the study setting. The evidence derived from two Ecuadorian universities; therefore, the conclusions were interpreted as context-specific and were not extrapolated as definitive claims for higher education systems beyond comparable institutional and national conditions. Consequently, external validation in additional institutions and countries was required before any broader regional generalization could be sustained.

Regarding the hypothesis framework, the results supported the proposed statements with different strength. H1 received support because the modelling task produced meaningful predictive performance using socio-educational predictors under a multiclass formulation and a unified evaluation workflow. H2 was supported because the competence indicators showed statistically significant variation across academic-degree categories in the inferential analyses reported for the DigCompEdu items. H3 was partially supported: Gradient Boosting and Neural Networks showed competitive accuracy, yet Random Forest also achieved strong performance; therefore, superiority over all simpler comparators did not hold uniformly across algorithms and metrics. H4 was supported at the level of observed associations and predictive relevance because significant relationships were identified between demographic variables (including age and gender) and digital-competence outcomes within the study dataset. H5 was supported because the lowest proficiency strata exhibited weaker identification, and competencies linked to assessment and feedback showed greater predictive difficulty, which remained consistent with the multifactorial character of this domain.

Furthermore, limitations affected both interpretation and prospective use. A pronounced class imbalance characterized the target distribution (e.g., markedly fewer observations in A1 and A2 than in B2 and C1), and no explicit rebalancing strategy was applied within model training; therefore, reduced sensitivity for low-proficiency categories remained an expected constraint. Consequently, any institutional use should prioritize group-level diagnosis and programme planning over individual-level screening when low-level predictions could drive high-stakes actions, because misclassification risk concentrated in minority classes. In addition, the absence of contextual variables limited explanatory granularity and constrained transferability to settings with different organizational, disciplinary, or infrastructural conditions.

Finally, future work should expand external validation across additional universities and incorporate organizational and pedagogical context to strengthen transportability. Moreover, subsequent pipelines should address minority-class performance through principled imbalance handling, and they should evaluate probabilistic calibration before probability outputs informed operational thresholds. Under these conditions, predictive analytics could support more reliable targeting of professional-development initiatives, with decisions aligned to institutional priorities and error costs.

Authors’ contributions

JM-C: Writing – original draft, Writing – review & editing, Conceptualization, Investigation, Supervision, Visualization.

AF-N: Conceptualization, Data curation, Methodology, Visualization, Writing – original draft, Writing – review & editing.

AC-V: Conceptualization, Investigation, Supervision, Visualization, Writing – original draft, Writing – review & editing.

HL-L: Conceptualization, Formal analysis, Supervision, Visualization, Writing – original draft, Writing – review & editing.

JV-M: Conceptualization, Data curation, Investigation, Writing – original draft, Writing – review & editing.

ÁS-G: Data curation, Formal analysis, Software, Supervision, Validation, Writing – original draft, Writing – review & editing.

References
  •  Alcaide-Pulido P, Gutiérrez-Villar B, Ordóñez-Olmedo E, et al.: Análisis de la preparación del profesorado para la docencia en línea: Evaluación del impacto y la adaptabilidad en diversos contextos educativos. Smart Learn. Environ. 2025; 12:5. Publisher Full Text
  •  Alghamdi S, Soh B, Li A: A Comprehensive Review of Dropout Prediction Methods Based on Multivariate Analysed Features of MOOC Platforms. Multimodal Technol. Interact. 2025; 9(1): 3. Publisher Full Text
  •  Almalawi A, Soh B, Li A, et al.: Predictive Models for Educational Purposes: A Systematic Review. Big Data Cogn. Comput. 2024; 8(12): 187. Publisher Full Text
  •  Alnasyan B, Basheri M, Alassafi M: The power of Deep Learning techniques for predicting student performance in Virtual Learning Environments: A systematic literature review. Comput. Educ.: Artif. Intell. 2024; 6: 100231. Publisher Full Text
  •  Al-Rahmi WM, Al-Adwan AS, Al-Maatouk Q, et al.: Integrating Communication and Task–Technology Fit Theories: The Adoption of Digital Media in Learning. Sustainability. 2023; 15(10): 8144. Publisher Full Text
  •  Asselman A, Khaldi M, Aammou S: Enhancing the prediction of student performance based on the machine learning XGBoost algorithm. Interact. Learn. Environ. 2023; 31(6): 3360–3379. Publisher Full Text
  •  Awedh MH, Mueen A: Early Identification of Vulnerable Students with Machine Learning Algorithms. WSEAS Trans. Inf. Sci. Appl. 2025; 22: 166–188. Publisher Full Text
  •  Bansal S, Grenfell BT, Meyers LA: When individual behaviour matters: Homogeneous and network models in epidemiology. J. R. Soc. Interface. 2007; 4(16): 879–891. PubMed Abstract | Publisher Full Text | Free Full Text
  •  Beer P, Mulder RH: The Effects of Technological Developments on Work and Their Implications for Continuous Vocational Education and Training: A Systematic Review. Front. Psychol. 2020; 11. PubMed Abstract | Publisher Full Text | Free Full Text
  •  Benavides L, Tamayo Arias J, Arango Serna M, et al.: Digital Transformation in Higher Education Institutions: A Systematic Literature Review. Sensors. 2020; 20(11): 3291. PubMed Abstract | Publisher Full Text | Free Full Text
  •  Bond M, Marín VI, Dolch C, et al.: Digital transformation in German higher education: Student and teacher perceptions and usage of digital media. Int. J. Educ. Technol. High. Educ. 2018; 15(1): 48. Publisher Full Text
  •  Cabero-Almenara J, Palacios-Rodríguez A: Marco Europeo de Competencia Digital Docente «DigCompEdu». Traducción y adaptación del cuestionario «DigCompEdu Check-In». EDMETIC. 2020; 9(1): 213–234. Publisher Full Text
  •  Cattaneo AAP, Antonietti C, Rauseo M: How do vocational teachers use technology? The role of perceived digital competence and perceived usefulness in technology use across different teaching profiles. Vocat. Learn. 2025; 18(1): 5. Publisher Full Text
  •  Chen L, Chen P, Lin Z: Artificial Intelligence in Education: A Review. IEEE Access. 2020; 8: 75264–75278. Publisher Full Text
  •  Chohan SR, Hu G: Strengthening digital inclusion through e-government: Cohesive ICT training programs to intensify digital competency. Inf. Technol. Dev. 2022; 28(1): 16–38. Publisher Full Text
  •  Choudhary H: Building bridges to digital inclusion: Implications for curriculum development of digital literacy training programs. Int. J. Technol. Enhanc. Learn. 2024; 16(3): 282–296. Publisher Full Text
  •  Christodoulou A, Angeli C: Adaptive Learning Techniques for a Personalized Educational Software in Developing Teachers’ Technological Pedagogical Content Knowledge. Front. Educ. 2022; 7. Publisher Full Text
  •  Dong M, Yao L, Wang X, et al.: Gradient Boosted Neural Decision Forest. IEEE Trans. Serv. Comput. 2021; 1–1. Publisher Full Text
  •  Essa SG, Celik T, Human-Hendricks NE: Personalized Adaptive Learning Technologies Based on Machine Learning Techniques to Identify Learning Styles: A Systematic Literature Review. IEEE Access. 2023; 11: 48392–48409. Publisher Full Text
  •  Fakhar H, Lamrabet M, Echantoufi N, et al.: Towards a New Artificial Intelligence-based Framework for Teachers’ Online Continuous Professional Development Programs: Systematic Review. Int. J. Adv. Comput. Sci. Appl. 2024; 15(4). Publisher Full Text
  •  Garay-Rondero CL, Castillo-Paz A, Gijón-Rivera C, et al.: Competency-based assessment tools for engineering higher education: A case study on complex problem-solving. Cogent Educ. 2024; 11(1). Publisher Full Text
  •  García-Morales VJ, Garrido-Moreno A, Martín-Rojas R: The Transformation of Higher Education After the COVID Disruption: Emerging Challenges in an Online Learning Scenario. Front. Psychol. 2021; 12. PubMed Abstract | Publisher Full Text | Free Full Text
  •  Haleem A, Javaid M, Qadri MA, et al.: Understanding the role of digital technologies in education: A review. Sustain. Oper. Comput. 2022; 3: 275–285. Publisher Full Text
  •  Hidalgo A, Gabaly S, Morales-Alonso G, et al.: The digital divide in light of sustainable development: An approach through advanced machine learning techniques. Technol. Forecast. Soc. Chang. 2020; 150: 119754. Publisher Full Text
  •  Ingkavara T, Panjaburee P, Srisawasdi N, et al.: The use of a personalized learning approach to implementing self-regulated online learning. Comput. Educ.: Artif. Intell. 2022; 3: 100086. Publisher Full Text
  •  Islam U, Alali IK, Alotaibi SD, et al.: Introducing the Hyperdynamic Adaptive Learning Fusion (HALF) model for superior predictive analytics in E-learning. Neural Comput. Applic. 2025. Publisher Full Text
  •  Kukkar A, Mohana R, Nayyar A, et al.: A Novel Deep-Learning-Based Bug Severity Classification Technique Using Convolutional Neural Networks and Random Forest with Boosting. Sensors. 2019; 19(13): 2964. PubMed Abstract | Publisher Full Text | Free Full Text
  •  Kumar V, Kedam N, Sharma KV, et al.: Advanced Machine Learning Techniques to Improve Hydrological Prediction: A Comparative Analysis of Streamflow Prediction Models. Water. 2023; 15(14): 2572. Publisher Full Text
  •  Kyriazos T, Poga M: Application of Machine Learning Models in Social Sciences: Managing Nonlinear Relationships. Encyclopedia. 2024; 4(4): 1790–1805. Publisher Full Text
  •  Lindfors M, Pettersson F, Olofsson AD: Conditions for professional digital competence: The teacher educators’ view. Educ. Inq. 2021; 12(4): 390–409. Publisher Full Text
  •  Maican CI, Cazan A-M, Lixandroiu RC, et al.: A study on academic staff personality and technology acceptance: The case of communication and collaboration applications. Comput. Educ. 2019; 128: 113–131. Publisher Full Text
  •  Mei XY, Aas E, Medgard M: Teachers’ use of digital learning tool for teaching in higher education. J. Appl. Res. High. Educ. 2019; 11(3): 522–537. Publisher Full Text
  •  Moreira-Choez JS, Gómez Barzola KE, Lamus de Rodríguez TM, et al.: Assessment of digital competencies in higher education faculty: A multimodal approach within the framework of artificial intelligence. Front. Educ. 2024a; 9. Publisher Full Text
  •  Moreira-Choez JS, Lamus de Rodríguez TM, Arias-Iturralde MC, et al.: Influence of gender and academic level on the development of digital competencies in university teachers: A multidisciplinary comparative analysis. Front. Educ. 2024b; 9. Publisher Full Text
  •  Moreira-Choez JS, Lamus de Rodríguez TM, Cedeño Barcia LA, et al.: Competencias digitales en docentes de educación superior: Un análisis integral basado en una revisión sistemática. Revista de Ciencias Sociales. 2024c; 30(3): 317–331. Publisher Full Text
  •  Moreira-Choez JS, Lamus de Rodríguez TM, Olmedo-Cañarte PA, et al.: Valorando el futuro de la educación: Competencias Digitales y Tecnologías de Información y Comunicación en Universidades. Revista Venezolana de Gerencia. 2024d; 29(105): 271–288.
  •  Moreira Choez JS, Núñez-Naranjo AF, Carrasco Valenzuela AC, et al.: Data from the article titled Machine Learning Algorithms to Predict Digital Competencies in University Faculty (Version 1). figshare.2025. Publisher Full Text
  •  Moreira-Choez JS, Zambrano-Acosta JM, López-Padrón A: Digital teaching competence of higher education professors: Self-perception study in an Ecuadorian university. F1000Res. 2024e; 12: 1484. PubMed Abstract | Publisher Full Text | Free Full Text
  •  Palacios-Rodríguez A, Llorente-Cejudo C, Lucas M, et al.: Macroassessment of teachers’ digital competence. DigCompEdu study in Spain and Portugal. RIED-Revista Iberoamericana de Educación a Distancia. 2024; 28(1). Publisher Full Text
  •  Radovan M, Radovan DM: Harmonizing Pedagogy and Technology: Insights into Teaching Approaches That Foster Sustainable Motivation and Efficiency in Blended Learning. Sustainability. 2024; 16(7): 2704. Publisher Full Text
  •  Rane NL: Education 4.0 and 5.0: integrating Artifcial Intelligence (AI) for personalized and adaptive learning. J. Artif. Intell. Robot. 2025. Publisher Full Text
  •  Rofi’i A, Nurhidayat E, Firharmawan H: Teachers’ Professional Competence in Integrating Technology: A Case Study at English Teacher Forum in Majalengka. Int. J. Lang. Educ. Cult. Rev. 2023; 9(1): 64–73. Publisher Full Text
  •  Santamaria-Velasco J, Núñez-Naranjo A, Morales-Urrutia X: Critical thinking and AI: Enhancing history teaching through ChatGPT simulations. Int. J. Innov. Res. Sci. Stud. 2025; 8(1): 564–575. Publisher Full Text
  •  Schmidt JT, Tang M: Digitalization in Education: Challenges, Trends and Transformative Potential. Führen und Managen in der digitalen Transformation. Wiesbaden: Springer Fachmedien; 2020; pp. 287–312. Publisher Full Text
  •  Sharma K, Papamitsiou Z, Giannakos M: Building pipelines for educational data using AI and multimodal analytics: A “grey-box” approach. Br. J. Educ. Technol. 2019; 50(6): 3004–3031. Publisher Full Text
  •  Smith ER, Zárate MA: Exemplar-based model of social judgment. Psychol. Rev. 1992; 99(1): 3–21. Publisher Full Text
  •  Stoet G, Geary DC: The Gender-Equality Paradox in Science, Technology, Engineering, and Mathematics Education. Psychol. Sci. 2018; 29(4): 581–593. PubMed Abstract | Publisher Full Text
  •  Taye MM: Understanding of Machine Learning with Deep Learning: Architectures, Workflow, Applications and Future Directions. Computers. 2023; 12(5): 91. Publisher Full Text
  •  Tee PK, Wong LC, Dada M, et al.: Demand for digital skills, skill gaps and graduate employability: Evidence from employers in Malaysia. F1000Res. 2024; 13: 389. PubMed Abstract | Publisher Full Text | Free Full Text
  •  Thordsen T, Bick M: A decade of digital maturity models: much ado about nothing? IseB. 2023; 21(4): 947–976. Publisher Full Text
  •  Tondeur J, Howard SK, Yang J: One-size does not fit all: Towards an adaptive model to develop preservice teachers’ digital competencies. Comput. Hum. Behav. 2021; 116: 106659. Publisher Full Text
  •  Vindigni G: Adaptive and Re-adaptive Pedagogies in Higher Education: A Comparative, Longitudinal Study of Their Impact on Professional Competence Development across Diverse Curricula. Eur. J. Theor. Appl. Sci. 2023; 1(4): 718–743. Publisher Full Text
  •  Wang X, Wang Z, Wang Q, et al.: Supporting digitally enhanced learning through measurement in higher education: Development and validation of a university students’ digital competence scale. J. Comput. Assist. Learn. 2021; 37(4): 1063–1076. Publisher Full Text
  •  Wei Z: Navigating Digital Learning Landscapes: Unveiling the Interplay Between Learning Behaviors, Digital Literacy, and Educational Outcomes. J. Knowl. Econ. 2023; 15(3): 10516–10546. Publisher Full Text
  •  Zhao Y, Zhao M, Shi F: Integrating Moral Education and Educational Information Technology: A Strategic Approach to Enhance Rural Teacher Training in Universities. J. Knowl. Econ. 2023; 15(3): 15053–15093. Publisher Full Text
  •  Zhou L, Pan S, Wang J, et al.: Machine learning on big data: Opportunities and challenges. Neurocomputing. 2017; 237: 350–361. Publisher Full Text

Схожие новости

#Наименование новостиТональностьИнформативностьДата публикации
1AI-Supported Literacy Ecosystems in Elementary Education: Preparing Future Skills for Lifelong Learning and Vocational Development [version 1; peer review: awaiting peer review]011.406-08-2026
2Data-Driven Policy: An Analysis of Graduate Job Waiting Time and Its Implications for Higher Education Services [version 1; peer review: awaiting peer review]010.9811-07-2026
3Machine Learning for Shield-Scale Soil Health Mapping: A Systematic Review [version 1; peer review: awaiting peer review]013.3515-07-2026
4Educational Chronotopes in the Digital Landscape of Higher Education: A Systematic Bibliometric Mapping of Research Trends, 2015–2025 [version 1; peer review: awaiting peer review]06.3406-08-2026
5Integrating Big Five Personality Traits and TPACK: A Pedagogical Framework for Collaborative Learning Management System (C-LMS) in Higher Education [version 1; peer review: awaiting peer review]06.8304-08-2026
6A Bibliometric Analysis of Digital Citizenship Education and Competences for Democratic Culture: Global Trends, Knowledge Structure, and Future Research Agenda [version 1; peer review: awaiting peer review]05.304-08-2026
7SmartFlex Learning Ecosystem: Integrating AI and Learning Analytics in Vocational Mathematics Education [version 1; peer review: 2 approved with reservations]07.8813-07-2026
8Artificial Intelligence Literacy and Educational Impact among Iraqi Nursing Students: A Cross-Sectional Survey [version 1; peer review: awaiting peer review]07.5616-07-2026
9Effectiveness and Implementation of Teacher Quality Improvement Policies: A Systematic Literature Review [version 1; peer review: awaiting peer review]05.5315-07-2026
10Sentiment Analysis of Acceptance TVET Online Courses on the Skill Academy App from Google Play: Leveraging Text Mining with Comparison Machine Learning Model [version 1; peer review: 1 approved, 2 approved with reservations]0709-06-2026

Классификация: . Схожих патентов: 0. Схожих новостей: 10. Тональность: 0. Информативность: 7.84. Источник: f1000research.com.