Background Although numerous implementation strategies have been evaluated to improve cancer screening uptake, the evidence base remains fragmented, inconsistently described, and difficult to apply to local contexts. This project will develop a living, ontology-informed, and large language model (LLM)-assisted knowledge base for cancer screening implementation. This resource will serve as a structured, continuously updated, and computable evidence infrastructure, enabling evidence mapping, synthesis, and retrieval of context -sensitive implementation strategies to support evidence-informed implementation decision-making . Methods First, a cancer screening implementation ontology will be developed, building on the Behavior Change Intervention Ontology to formally represent screening-specific pathways, modalities, implementation determinants, strategies, and outcomes. Second, guided by the Cochrane Effective Practice and Organization of Care method, we will systematically identify randomized controlled trials evaluating cancer screening implementation strategies from four databases: PubMed, EMBASE, CINAHL, and PsycINFO, from database inception to May 13, 2026. Information from eligible studies will be extracted and standardized using ontology. All the above steps will be supported by a human-LLM collaborative workflow that incorporates one or more LLMs with source-grounded verification and expert review to improve efficiency while maintaining human accountability. The outputs will include an interactive knowledge base that supports evidence visualization, synthesis, structured comparison, and future decision support. The knowledge base will be continuously updated and progressively expanded to incorporate additional forms of implementation evidence. Workflow performance and knowledge base quality will be evaluated using predefined process and outcome metrics. Discussion The proposed knowledge base will improve the standardization, comparability, and usability of implementation evidence in cancer screening by transforming heterogeneous trial evidence into structured and computable evidence units. Future prediction or decision-support applications, when developed, will contribute to emerging efforts in precision implementation science by enabling a more context-sensitive understanding of cancer screening strategies.
Cancer screening is a key strategy for reducing cancer-related morbidity and mortality. Despite the establishment of evidence-based clinical guidelines and the availability of screening tools and technologies, population-level uptake remains suboptimal across cancer types. For instance, although low-dose computed tomography reduces lung cancer mortality in randomized controlled trials,1,2 fewer than 20% of eligible individuals in the United States undergo recommended screening.3,4 In colorectal cancer screening, overall uptake was reported at only 35.9%.5 These gaps are further compounded by persistent inequities across subgroups.6–8 In the Cancer CHRNA study, up-to-date mammography, cervical, and colorectal cancer screening rates varied across racial and ethnic groups, ranging from 50.6% to 73.2% for mammography, 12.6% to 76.8% for Pap testing, and 26.8% to 69.0% for colonoscopy.9 Geographic disparities are also evident, with only 48.6% of rural residents and 64.0% of urban residents receiving a Pap test for cervical cancer screening.10 These observations highlight a persistent knowledge-to-practice gap and the need for effective and equitable implementation strategies.
Implementation science seeks to bridge the gap between evidence and practice by developing and evaluating implementation strategies.11 In cancer screening, a wide range of implementation strategies have been evaluated to improve uptake across patients, clinicians, and system levels, such as reminders, patient navigation, outreach, and workflow redesign.12–15 Across cancer types, many of these strategies target similar behavioral, organizational, and structural barriers, such as limited patient awareness, competing clinician demands, fragmented care pathways, or resource constraints.
Despite rapid growth in implementation research, the evidence base remains fragmented and difficult to translate into practice. Implementation trials are dispersed across cancer-specific, public health, and implementation science literature, making it difficult to maintain a comprehensive understanding of the field. Implementation strategies are described using inconsistent terminology and heterogeneous reporting formats, limiting comparisons across studies and cancer types.16 This lack of standardized reporting also hinders the identification of active intervention ingredients and mechanisms of action through which implementation strategies influence screening behaviors.17
In addition, most existing evidence syntheses are static and quickly become outdated, limiting their value for ongoing decision-making.18 There is no continuously updated resource that systematically organizes implementation evidence across cancer types, populations, and contexts. This gap limits the advancement of both implementation research and practice.
Recent advances in large language models (LLMs) have created new opportunities to support labor-intensive evidence review tasks, including literature search, screening, and data extraction.19,20 These capabilities may support more efficient maintenance of a continuously updated (“living”) knowledge base. In addition, LLMs have recently been applied to ontology engineering, supporting the generation of candidate concepts, definitions, and relationships under expert oversight.21,22 An ontology is a structured framework that organizes knowledge about entities or concepts through standardized terms, definitions, and relationships, so that information can be represented in a computer-readable format. Ontologies can support consistent organization, integration, reuse, and analysis of knowledge.23 In cancer screening, an ontology-informed structure could help translate inconsistently reported trial descriptions into comparable evidence units. These advancements create new opportunities to integrate LLM-assisted workflows with ontology-based representations to address fragmentation in research evidence. To date, however, LLM and ontology have not been systematically applied to develop a comprehensive knowledge base for cancer screening implementation.
The overall aim of this project is to develop a living, LLM-assisted knowledge base for cancer screening implementation that supports continuous evidence visualization, synthesis, structured comparison, and future decision support. The knowledge base is designed to represent evidence in a standardized, computable, and provenance-aware format through a shared ontology framework.
Objective 1: Develop a cancer screening implementation ontology to structure and standardize cancer screening implementation evidence.
Objective 2: Develop and validate a reusable, ontology-informed human–LLM collaborative workflow to support literature review and its living updates.
Objective 3: Build the initial cancer screening implementation knowledge base.
Objective 4: Assess the coverage, representation, integrity, queryability, and traceability of the final release-ready knowledge base, with ongoing monitoring of workflow drift across living-update cycles.
An ontology-informed framework will guide the harmonization, organization, representation, and continuous updating of evidence on cancer screening implementation. In this study, we will leverage the Behavior Change Intervention Ontology (BCIO) as the overarching conceptual framework.24–26
The BCIO is a theory-informed ontology framework that comes from The Human Behaviour-Change Project, which provides a standardized and computable representation of behavior change interventions and their related components.24–26 To support detailed and standardized representation of implementation evidence, several BCIO-compatible lower-level ontologies will be incorporated, including the Behavior Change Technique Ontology,27 Mechanism of Action Ontology,28 Mode of Delivery Ontology,29 Intervention Setting Ontology,30 Population Ontology31 and the Intervention Source Ontology.32
Building on this BCIO-aligned conceptual foundation, a cancer screening implementation ontology will be developed to further represent domain-specific constructs not explicitly captured in existing BCIO structures. These constructs include screening modalities, implementation determinants, implementation strategies, screening pathway stages and associated outcomes.
Cancer screening pathway is characterized by multi-step and longitudinal behaviors that require structural representation. To capture this complexity, ontology modules will be organized around a core screening pathway, which defines shared stages across cancer types, including invitation, decision-making, test completion, follow-up, and longitudinal adherence. Cancer-specific modules then apply this core structure to different screening programs.
We will use the project’s ontology-informed human–LLM workflow (see “LLM-assisted workflow”), configured with ontology-specific per-task protocols, to support ontology development. The workflow will include automated checks for duplicate or near-duplicate classes, verification of class hierarchies and relationships,21 and conflict-resolution procedures that route candidate concepts for integration, human review, or exclusion.33 Given the known limitations of LLM-assisted ontology engineering,22,34 domain experts will retain responsibility for all schema-level decisions. The ontology will be developed through four steps:
1) Extract candidate concepts and their relations. One or more extraction LLMs operating under ontology-grounded per-task protocols will extract candidate concepts and their relations from literature such as randomized trials, reviews, and clinical guidelines. These protocols will constrain extraction to predefined entity types (e.g., interventions, settings, populations, mechanisms, screening stages, and outcomes) and relation types, and exclude off-scope content.21,35 Each model will generate structured candidate assertions linked to verbatim source text and document metadata. Candidate extraction identified by one or more models will then be consolidated through the source verification and expert adjudication process, with input from the study team to ensure conceptual accuracy, consistency, and completeness.
2) Normalize identified concepts. The BCIO and its lower-level ontologies will be instantiated as a machine-readable class registry to support automated lookup and semantic matching. Each candidate concept will then be normalized to existing ontology classes using combined lexical, embedding-based, and hierarchical similarity matching, applying a “reuse before mint” principle: an existing class is reused whenever semantic equivalence or sufficient subsumption is established, and a new class is proposed only when no adequate match exists.31,32 All reuse and mappings decisions will be verified against the source ontologies to confirm correct alignment, proper parent-class assignment, and avoidance of redundant classes.19
3) Develop modules. Aligned and newly introduced concepts will be organized hierarchically under the core screening pathway superclass and within cancer-specific modules. For each concept, we will develop a standardized specification using a retrieval-augmented workflow that draws on evidence from included studies, screening guidelines, and existing ontology resources. The specification will include a preferred label and synonyms, a clear definition, inclusion and exclusion criteria, illustrative examples, relationships to parent and related concepts, and annotation guidance to support consistent coding and application across studies. High-quality BCIO entries will be used as structural exemplars to maintain consistent modeling patterns.18,32 The core screening pathway superclass, concept definitions, boundary criteria, and hierarchical relationships will be reviewed and finalized by domain experts to ensure conceptual clarity, coherence, and non-overlap.19
4) Pilot-testing and iterative refinement. The draft ontology will be evaluated through pilot annotation of a purposive sample of randomized trials representing multiple cancer screening domains (e.g., breast and colorectal cancer screening) to assess its applicability across diverse screening contexts.21,36 During pilot testing, annotation challenges, including ambiguous or poorly defined concepts, overlapping classes, missing concepts, and inconsistent mappings across LLM-models or human annotators, will be systematically documented. These findings will be used to refine concept definitions, class boundaries, hierarchical relationships, mapping rules, annotation guidance, and the associated extraction and annotation protocols. The ontology and workflow will be iteratively revised and retested until predefined criteria for stability and consistency are achieved, such as acceptable annotation agreement, reduced numbers of unresolved concepts, and consistent concept mapping across screening domains. The resulting ontology will be released as Version 1.0 under a version-controlled governance process.
As the living knowledge base evolves, newly identified concepts will be reviewed through a version-controlled ontology governance process and incorporated into subsequent releases where appropriate. Consistent with previous LLM-assisted BCIO work, in which fully automated ontology annotation achieved an F1 of approximately 0.42 compared with a human benchmark of 0.75,37 the LLM-assisted workflow in this project will be used as a decision-support and prioritization tool rather than a fully autonomous system. All flagged, uncertain, novel, or high-impact concepts will undergo expert review and adjudication before being incorporated into the ontology.
The project will be implemented through a human–LLM collaborative workflow to support reproducibility, transparency, and task-specific performance. Each discrete task in this workflow is governed by a reusable, modular, and version-controlled per-task protocol. Each protocol operationalizes the task for the model (specifying task objectives, input requirements, extraction or annotation rules, the required output schema, worked examples, and prohibited inferences) and defines how the resulting outputs are verified (through quality-control checks, validation criteria, and a version history). The runtime prompt issued to a given model is derived from, and versioned against, its protocol, so that the protocol remains stable and auditable while prompts or model versions change.38 Each LLM-assisted task will generate structured candidate outputs that include source text, source location, evidence type, confidence rating where applicable, protocol version, model metadata where available, and validation status. These outputs will be treated as provisional until verified against source evidence and validated through predefined quality control procedures.
Source verification and human adjudication architecture
Figure 1 summarizes the overall workflow architecture. The workflow employs one or more LLMs operating under predefined per-task protocols and structured schemas. If more than one model is used, outputs may be compared across models as a triage signal. Only outputs that are supported by source evidence are accepted as provisional candidates. Outputs that are inconsistent or uncertain are first checked automatically and, if needed, reviewed by human experts. Human-adjudicated outputs will be the reference standard for process validation. This verification-driven design keeps final decisions human-accountable and supports a model-agnostic workflow suitable for continuous updates in a living evidence system.
Pre-deployment validation criteria for LLM-assisted tasks
Before deployment in the living knowledge base, each LLM-assisted task will be validated against a human-adjudicated reference standard using an independent validation set. Validation criteria will be defined a priori based on the complexity of each task: (1) During title/abstract screening, we will prioritize sensitivity to minimize the risk of excluding eligible studies. The LLM is expected to achieve ≥90% sensitivity, with uncertain records routed to human review; (2) During full-text screening, the LLM must provide explicit eligibility evidence extracted from the source text and achieve ≥85% agreement with human reviewers before being deployed for full-scale screening; (3) Structured data extraction must include source-text evidence, pass schema validation, and achieve acceptable field-level performance (≥85% field-level accuracy), with complex or uncertain fields routed to human adjudication; (4) Ontology annotation must use valid ontology identifiers, where applicable, include source-text justification, pass ontology consistency checks, and achieve acceptable agreement with expert-generated mappings; and (5) For risk-of-bias assessment, the LLM will identify methodological evidence and generate preliminary domain-level judgments, with uncertain, unsupported, or conflicting judgments routed to human adjudication. Outputs that fail to meet validation criteria, require human adjudication, or exhibit recurrent uncertainty will inform iterative refinement of the protocols and LLM workflows.
Protocol refinement
Protocol refinement will be iterative. Human adjudication logs will be reviewed to identify recurrent error patterns, such as missed intervention components, unsupported ontology mappings, invalid ontology identifiers, inconsistent outcome extraction, and incorrect interpretation of risk-of-bias evidence. These error patterns will be used to update task instructions, extraction schemas, annotation guidance, examples, and quality-control rules. Substantive changes to protocols, prompts, model versions, ontology terms, or output schemas will be documented and re-evaluated before deployment in the living workflow.
Deployment and monitoring of LLM-assisted tasks
LLM-assisted tasks will be deployed in the living knowledge base workflow only after meeting predefined validation criteria against a human-adjudicated reference standard. Workflow performance will be monitored through periodic human review of LLM-processed records, and revalidation will be required after substantive changes to models, protocols, schemas, ontologies, or eligibility criteria. Tasks that fall below predefined thresholds will be paused, routed for human review, refined, and revalidated before redeployment.
Artificial intelligence uses and disclosure
All LLM assistance in this project will be reported transparently and documented for reproducibility. For each LLM-assisted task, we will document the model’s name and version, access mode and date of use, the protocol/prompt version and input-document version, inference settings where configurable, the output schema, and the human-adjudication decision. LLMs will be used only as tools under human supervision and will not be listed as authors. All LLM-generated outputs will remain provisional until they pass source verification and human adjudication.
To build the knowledge base, we will follow the Cochrane Effective Practice and Organization of Care (EPOC) Group’s methodological guidance for systematic reviews39 and the Responsible use of AI in Evidence Synthesis (RAISE) principles proposed by Cochrane.40 All processes will be governed by the human–LLM workflow described above. LLM-assisted outputs will be incorporated into the knowledge base only after completion of the predefined verification and adjudication procedures.
Literature search
Four databases will be searched: PubMed, EMBASE, CINAHL, and PsycINFO. The search strategy will be developed and iteratively refined based on published literature and the expertise of the research team. Searches will be structured around three core concepts: cancer, screening, and uptake, with database-specific controlled vocabulary and keywords used to maximize sensitivity for implementation-related studies (see the Extended data for the PubMed search strategy). The initial search will include studies published up to May 13, 2026.
The Cochrane EPOC guideline recommends four types of empirical evidence for systematic reviews that aim to test the effectiveness of implementation strategies: randomized trials, non-randomized trials, controlled before-after studies, and interrupted time series and repeated measure studies.41 For building the initial knowledge base, randomized controlled trials will be prioritized for inclusion to support the development of a high-internal-validity knowledge base and to enable initial ontology and workflow calibration. Additional study designs will be considered in subsequent phases of the living knowledge base.
Literature screening
For the purpose of this knowledge base, cancer screening is defined as the use of evidence-based screening tests or examinations in individuals without signs or symptoms of cancer to identify cancer or precancerous lesions at an earlier, potentially more treatable stage. Diagnostic evaluation of symptomatic individuals, post-treatment surveillance, and interventions primarily focused on hereditary cancer risk assessment (e.g., germline genetic testing or genetic counselling) were considered outside the scope of the initial knowledge base.42
Records retrieved from all databases will be deduplicated and screened at the title/abstract and full-text stages. Before deployment, the LLM-assisted screening workflow will be evaluated using an independent validation sample of 500 records. These records will undergo both human screening and LLM-assisted screening, with human-adjudicated decisions serving as the reference standard (see details in Table 5). Screening performance will be assessed against the predefined validation criteria described in the previous sections. Records with discordant or uncertain classifications will be reviewed by human experts and used to refine screening protocols and decision rules where necessary. Once the validation criteria are met, remaining records will be screened using the validated LLM-assisted screening process described in the previous sections, with uncertain or conflicting cases continuing to be routed to human review. Eligibility criteria are presented in Table 1.
Data extraction
Data extraction will be guided by the Cochrane EPOC template.43 Implementation-relevant information extraction will be informed by the Expert Recommendations for Implementing Change (ERIC) taxonomy44 and Proctor’s implementation outcomes framework.45
Two levels of data will be extracted. Study-level fields will include publication details, country, study design, theoretical framework, unit of randomization, setting, cancer type, screening modality, target population, eligibility criteria, intervention and comparator conditions, sample size, follow-up period, outcomes, and effect estimates where available.
Implementation-relevant fields will include target screening behavior, screening pathway stage, target actor, intervention recipient, intervention source, mode of delivery, implementation setting, population characteristics, engagement, implementation determinants, equity-relevant context, implementation strategies, and implementation outcomes.
All extracted data will be linked to source text and provenance metadata to ensure traceability. For each reported outcome, both the effect estimates (where available) and explicit effect-direction indicators (e.g., favoring intervention, favoring control, null, or non-inferior) will be recorded to preserve directionality and support downstream synthesis of null and negative findings.46 LLM-assisted data extraction will be validated against human-adjudicated reference standards using an independent sample of approximately 50 randomized controlled trials. Detailed procedures and evaluation metrics are provided in Table 5.
Ontology annotation
Intervention descriptions will be decomposed into discrete components within each study arm. Original author-reported descriptions are retained prior to standardization to preserve traceability. Each component is mapped to the ontology framework to enable structured representation across key dimensions (see Table 2 for details).
To balance accuracy and computational cost across the cancer screening implementation ontology and its lower-level ontologies, a two-stage retrieve-and-rerank workflow will be used. In the first stage, an embedding-similarity retrieval module selects a small set of top-k candidate ontology terms from each annotation domain. In the second stage, the LLM model will select the most appropriate term from this restricted candidate set, using the source text and intervention component as context.46 For each annotation domain, the LLM workflow will also include an explicit no-applicable-term option. Records assigned to this option will be routed to the ontology governance process as candidate concepts for future ontology versions. This process will support the identification of coverage gaps within the cancer screening implementation ontology and inform subsequent ontology development.46 LLM-assisted ontology annotation will be validated against human-adjudicated reference standards using 5–10 pilot trials for calibration, followed by an independent validation sample of 20–30 randomized controlled trials. Detailed procedures and evaluation metrics are provided in Table 5.
Risk-of-bias assessment
Risk of bias in included randomized controlled trials will be assessed using the Cochrane Risk of Bias 2 (RoB 2) tool.39,47 Domains assessed will include bias arising from the randomization process, deviations from intended interventions, missing outcome data, measurement of the outcome, and selection of the reported result. As later phases incorporate non-randomized study designs, appropriate additional risk-of-bias tools (e.g., ROBINS-I) will be adopted. Risk-of-bias assessments will be informed by LLM-assisted identification of relevant methodological text, but all domain-level judgments will be finalized through human validation. The LLM-assisted risk-of-bias workflow will be validated against human-adjudicated reference standards using an independent sample of 20–30 randomized controlled trials. Detailed procedures and evaluation metrics are provided in Table 5.
Data analysis
Descriptive and exploratory analyses. We will summarize the basic characteristics of included studies and examine variation in implementation strategies across cancer types, healthcare settings, populations, and equity-relevant contexts. Results will be reported using counts, proportions, and cross-tabulations. Building on these analyses, results will be further organized into evidence tables, descriptive summaries, and implementation strategy frequency distributions. The evidence units will be integrated into a provenance-aware evidence graph that represents relationships among implementation strategies, behavior change techniques, mechanisms of action, implementation determinants, screening modalities, screening pathway stages, populations, settings, outcomes, and effect estimates, where available. These outputs will support evidence synthesis, knowledge base querying, and structured reporting.
Network and co-occurrence analysis. The ontology-informed evidence graph will support network-based evidence mapping and exploratory network analyses of relationships among implementation strategies, behavior change techniques, mechanisms of action, implementation determinants, screening pathway stages, and outcomes. Co-occurrence patterns across studies will be examined to identify frequently co-occurring implementation components, recurrent implementation configurations, and clusters of strategies used within specific screening contexts. For example, analyses may examine which implementation strategies and behavior change techniques are most frequently associated with specific screening modalities, pathway stages, implementation determinants, or target populations. The graph also supports identification of gaps in evidence across cancer types, populations, screening stages, healthcare settings, and equity-relevant contexts.
Exploratory hypothesis-generation analyses. Exploratory graph-based analyses may be conducted using the evidence graph to identify under-studied implementation strategy–context–outcome relationships. If link-ranking or neighborhood-similarity methods are used, they will be reported as research-prioritization tools rather than predictions of effectiveness and may generate candidate links between implementation strategies, populations, screening modalities, and outcomes. Candidate relationships will be reviewed alongside available effect estimates, reporting completeness, and provenance quality. Details of the exploratory analysis procedures are provided under outcome evaluation ( Table 4).
Graph-based querying and downstream applications. The ontology-informed evidence graph will be queried using structured graph-query approaches (e.g., SPARQL- or Cypher-style queries over evidence triples) to retrieve and organize evidence across screening modalities, screening pathway stages, populations, healthcare settings, and equity-relevant contexts. Query results will be used to identify implementation strategies, behavior change techniques, mechanisms of action, implementation determinants, and reported outcomes associated with specific screening contexts. All retrieved results will remain linked to their original source studies through provenance records. These query outputs will be used to generate evidence-gap maps, comparative analyses of implementation approaches, and other knowledge base outputs planned for later phases of the project.
Data synthesis
Upon establishment of the ontology-informed evidence infrastructure, a subsequent project will use systematic review and meta-analysis approaches to synthesize the effectiveness of cancer screening implementation strategies. This follow-on work will examine which strategies are effective across cancer screening contexts, which are most effective for specific screening pathways or outcomes, and which strategies are most applicable or beneficial for particular cancer screening modalities.
This project is designed as a living knowledge base in which newly identified studies are periodically incorporated using the established update workflow. Search updates will follow a two-tier schedule, with application-programming-interface-accessible databases (e.g., PubMed through E-utilities) queried at high frequency (e.g., weekly) and databases requiring manual interface searches (e.g., EMBASE) queried at lower frequency (e.g., quarterly to biannually).48 To keep the cumulative knowledge base reproducible across releases, evidence units that are not affected by an ontology, model, protocol, or schema change will be preserved across update cycles without re-annotation. Only records affected by the predefined recalibration trigger or routinely sampled for drift audit will be re-annotated. This update strategy follows an append-only versioning approach consistent with established living evidence systems.48 A designated team member will oversee the living-update process, confirm the search schedule, and approve each knowledge base release. The living-review component will be reported in accordance with the PRISMA-LSR extension for living systematic reviews.49
The living workflow also supports iterative refinement of the ontology and associated annotation guidelines. New or under-represented concepts identified during updates will be reviewed through the ontology governance process and incorporated into subsequent versioned releases where necessary.
The initial version of the knowledge base focuses on randomized controlled trials, with later phases expanding to incorporate non-randomized studies, controlled before-after studies, interrupted time series studies, and other forms of real-world implementation evidence where appropriate. In selected cases and where feasible, additional collaboration with trial investigators may be explored to obtain individual participant data to support more detailed representation of implementation contexts, populations, screening pathways, and outcomes.
Both process and outcome evaluations will be conducted to assess the performance, reproducibility, transparency, and ongoing maintainability of the ontology-informed and LLM-assisted workflow and the resulting knowledge base.50–52
Process evaluation
Process evaluation will assess workflow performance by examining the agreement between LLM-assisted outputs and human-adjudicated reference standards across literature screening, data extraction, ontology annotation, and risk-of-bias assessment ( Table 3), with ongoing monitoring of living-update currency and workflow drift. Workflow development and validation will be separated: a small pilot calibration sample will first refine eligibility criteria, protocols, prompts, output schemas, and annotation guidance (treated as development and quality assurance), and formal validation will be followed only after the protocols are finalized, using independent human-reviewed samples. Where appropriate, error rates from these validation samples may be used to estimate uncertainty in the results of the full-corpus LLM-assisted output. Evaluation of the ontology itself (e.g., ontology coverage, consistency, fitness for purpose, and expert assessment of the ontology structure) is outside the scope of this initial phase and will be assessed in subsequent dedicated studies.
Validation metrics are aligned with the operational goals of each task ( Table 3). Screening validation will prioritize sensitivity to minimize false exclusions of eligible studies for downstream review, whereas data extraction, ontology annotation, and risk-of-bias assessment will emphasize agreement with human judgments and source-text grounding to limit unsupported inference. Substantial changes to models, prompts, ontology terms, or output schemas introduced during living updates will trigger the same recalibration.48
Outcome evaluation
Outcome evaluation will assess the completeness, structure, transparency, and usability of the final release-ready knowledge base generated in this initial development phase. These evaluations will be conducted on the final human-validated or release-ready knowledge base rather than on provisional LLM-generated outputs. The primary outcome evaluation will focus on coverage of existing systematic review evidence by determining whether eligible randomized trials included in published reviews are represented in the released knowledge base. Secondary evaluations will assess evidence representation breadth and balance, release integrity and traceability, and later-stage usability and perceived relevance of the knowledge base. Detailed indicators, definitions, and evaluation metrics for each dimension are provided in Table 4.
This project aims to develop a living, ontology-informed, and LLM-assisted knowledge base to systematically organize evidence on implementation strategies used in cancer screening. By integrating behavior change ontologies with a human–LLM collaborative evidence review workflow, the project aims to create a standardized and continuously updatable infrastructure for understanding what implementation strategies work, for whom, under which conditions, and at what stage of the cancer screening pathway.
Despite growing interest in ontologies in biomedical53 and oncology research,54,55 especially recent efforts focused on breast cancer screening56 and cancer registry data,57 their application to cancer screening implementation remains limited. Ontology-based representations, particularly when combined with knowledge graph approaches, offer a powerful way to connect implementation strategies, mechanisms, determinants, populations, and outcomes into a unified evidence structure. This may improve the ability to synthesize fragmented implementation evidence and support more consistent interpretation across studies. This project addresses these gaps through the development of an ontology-informed living knowledge base specifically designed for cancer screening implementation research.
Generative LLM offers opportunities to improve efficiency in evidence synthesis tasks such as literature identification, screening, data extraction, and risk-of-bias assessment,58,59 but current systems still exhibit substantial error rates across tasks. A recent systematic review assessed the use of generative LLM systems including ChatGPT, GPT, Claude, Bing AI, and Perplexity AI in evidence review and reported substantial limitations across multiple review stages. For literature searching, generative AI systems failed to identify between 68% and 96% of relevant studies. During study screening, the probabilities of incorrect inclusion and incorrect exclusion ranged from 0% to 29% and 1% to 83%, respectively. Errors in data extraction ranged from 4% to 31%, and inaccurate risk-of-bias assessments ranged from 10% to 56%.20 To minimize these risks of bias, this project adopts a human–LLM collaborative approach in which LLM supports high-volume literature screening, extraction, and annotation tasks, while human reviewers will provide oversight, adjudication, and interpretation for uncertain or complex cases.60 This approach aims to balance efficiency with methodological rigor and may provide practical insights into how LLM can be responsibly integrated into future living evidence review infrastructures.61
A major challenge in implementation science is that implementation strategies often show variable effectiveness across populations, settings, and healthcare systems. Strategies that are successful in one context may have limited impact in another because implementation outcomes are shaped by interactions among contextual factors, delivery processes, organizational conditions, and target population characteristics.62,63 Recent implementation science initiatives are increasingly moving toward LLM-enabled and ontology-informed infrastructures designed to support more context-sensitive implementation decision-making. For example, the AIM-IS program seeks to develop frameworks and reporting standards for responsible AI use across implementation workflows.64 The ImpleMATE platform proposes ontology-based knowledge representation combined with LLM-assisted reasoning to support implementation planning and learning health system cycles.65 These emerging initiatives reflect growing interest in developing computable implementation knowledge systems to support the continuous integration and updating of implementation evidence across contexts. The ontology-informed living knowledge base proposed in this project contributes to this evolving field by enabling structured representation and continuous updating of implementation knowledge, which may support future development of more precise, context-sensitive cancer screening implementation research and decision-making in the future.66
The proposed knowledge base may improve the standardization, comparability, and reuse of implementation evidence by converting fragmented, inconsistently reported trial evidence into standardized, computable, and provenance-aware evidence units. By linking implementation strategies, behavior change techniques, mechanisms of action, implementation determinants, populations, settings, screening pathway stages, and outcomes, the project may also support precision implementation science with a more context-sensitive and mechanism-informed account of cancer screening implementation strategies. This project aims to establish a continuously updated foundation for knowledge integration and future implementation planning and research in cancer screening.
The ontology development component is informed by existing ontologies, published literature, cancer screening guidelines, and concepts identified from included studies. This phase does not involve the collection of data from human participants and therefore does not require Institutional Review Board (IRB) review. The knowledge base construction component of this project involves the analysis of publicly available published literature and does not involve human participants or identifiable private information. Therefore, IRB review was not required for this component. Future phases involving stakeholder engagement, expert consultation, Delphi consensus procedures, usability testing, or other activities involving human participants to further develop, refine, validate, or evaluate the ontology and knowledge base will be submitted for IRB review and approval prior to commencement.
Not applicable.
No generative AI tools were used to generate scientific content in this protocol. The use of large language models described in this protocol refers to the planned research workflow and not to the generation of the manuscript content. The use of AI was limited to language refinement of author-drafted text to improve clarity and readability. No AI tools were used to create or modify any figures or images. The use of large language models described in this protocol refers to the planned research workflow and not to the generation of the manuscript content. All content was reviewed, revised, and approved by the authors, who take full responsibility for the final manuscript.