How can a GP practice safely assess AI for test-result triage?
Start with the exact decision the software will make
“AI sorts the results” can mean several very different things. One product may group incoming reports by sender while leaving every result in the existing clinician queue. Another may identify values it calls urgent, move apparently normal reports into a separate queue, draft patient messages or propose clinical actions. The first task for a practice is to write down what the software will receive, what it will output, what it can change and what a human must still decide. A procurement slide or a supplier's general statement that a clinician remains responsible is too vague to control a live workflow.
NHS England's digital GP guidance on AI says implementation requires robust clinical validation and attention to patient impact, inequalities, workflow integration and ongoing safety. Its clinical messaging and test-results guidance explains that practices receive large streams of reports and must understand the capabilities and limits of the communication tools they use. CQC's GP test-results guidance asks whether practices have a safe and effective system across ordering, receipt, action and patient communication. A new classifier becomes part of that system even if it does not place a diagnosis in the record.
The proposed owner here is a practice considering a new AI-assisted sorting step. The GP test-results backlog draft covers recovery when a queue has already accumulated unreviewed results. The GP clinical-system outage draft covers care and record continuity during loss of a system. This page covers the safety decision before and during deployment of a result-routing tool. If there is no distinct tool and workflow change, a practice should use the existing results-governance owner rather than create an AI-branded duplicate.
Draw the result journey before assessing the algorithm
Map how a test is requested, where the laboratory or other provider sends its report, how the practice receives it, where the AI sees it, what fields it reads, which queue it can change, who reviews the output and how the patient hears the outcome. Include results arriving as structured data, documents, amended reports, scanned correspondence and messages from services outside the usual laboratory feed. Identify rejected messages and reports that never reached the practice. A sorter cannot safely manage a report that is absent from its input, and an impressive classification accuracy figure does not solve the missing-result problem.
Mark every point where the software can suppress, delay or reroute human attention. A tool that labels a result “routine” but leaves it in the clinician queue changes presentation. A tool that removes it from review changes the control itself. A tool that drafts a patient message introduces a communication risk. A tool that files a result, updates a code or triggers a recall may affect further care. Write each function separately, including settings that can be changed after purchase. Do not assess the product as one indivisible box when only some functions are authorised for use.
Identify the accountable clinician for results requested by the practice and the cover arrangement during absence. CQC says clinicians are responsible for acting on results that alter patient management. It recognises that trained non-clinical staff may perform certain tasks with appropriate safeguards, but a practice must demonstrate that its system is safe. The supplier's routing decision should never leave a result in a queue that no one owns. If the tool changes which professional sees a report, test the receiving role's competence, permissions and handover rules.
Document the patient's route too. Some results may become visible through patient-facing systems. A patient seeing an automated label or portal release is not the same as a clinician assessing the result and explaining a care plan. The practice must decide how unexpected, urgent, uncertain and amended results are handled. Do not promise that a patient message will be sent merely because the AI generated text. Check which human approves it and how failed delivery is detected.
Establish the regulatory and safety assurance position
Ask the supplier for the product's exact intended purpose and a description of each enabled function. The MHRA's software and AI medical-device guidance explains that many software and AI products used in health care are regulated as medical devices or in vitro diagnostic medical devices. Whether a particular sorter is a regulated device depends on its intended purpose and function. Do not assume that every use of AI is a medical device, or that a supplier's marketing term “administrative” settles the classification. Ask for the supplier's rationale, relevant registration and conformity evidence where applicable, and advice from the appropriate regulatory specialist if the position is unclear.
NHS England describes DCB0129 for manufacturers and DCB0160 for organisations deploying and using health IT. A supplier safety case cannot replace the local deployment assessment. The practice, network or integrated care board must establish who owns the DCB0160 work for this implementation, how the supplier's hazards are translated into the actual GP workflow and who signs off remaining risk. NHS England began a national review of these standards in 2026; the review supporting information is a consultation context, not evidence that a proposed replacement has already taken effect.
The clinical safety officer should consider hazards such as an urgent report misclassified as routine, an amended result not re-entering the queue, a patient matched to the wrong record, a model confidence value mistaken for a clinical decision, or a version update changing classification without review. Record the cause, possible harm, existing controls, additional controls, owner and residual risk. The exact risk evaluation belongs to the local clinical safety process. A sample checklist on this page cannot sign off an individual product.
Procurement should obtain the supplier's version, change-control method, clinical performance evidence, limits, incident route, support hours, data locations, subcontractors and withdrawal process. Ask what happens if a connection fails, if the model is unavailable, if input format changes or if an unknown report type appears. The contract should give the practice enough information and rights to detect, investigate and correct a safety issue. An assurance document written for a hospital pathway may not prove performance in a GP results queue.
Test the tool against the practice's real workload
Local validation should reflect the exact range of reports, patient population, laboratories, languages and workflow exceptions that the practice sees. Use an approved, privacy-protective evaluation environment and a clinician-agreed reference standard. Separate historical test data from future live monitoring. Include abnormal, normal, borderline, amended, duplicate, misdirected and incomplete reports. Include situations where the clinical meaning depends on previous results, symptoms, medicine or patient history. A correct extraction of a number is not the same as a correct decision about urgency.
Define outcomes before looking at a headline accuracy rate. The practice needs to know how often a result needing urgent attention is placed in a lower-priority route, how often a report is not processed at all, how often a safe result is over-escalated and whether performance varies by source or patient group. Decide what errors a human check can realistically catch at the planned workload. If staff are expected to re-read every result in full, the tool's efficiency case may disappear. If staff see only the tool's summary, they may miss context that the model discarded. Test the actual human-machine combination, not only the model in isolation.
NHS England's AI guidance warns that algorithms can exacerbate health inequalities and suggests equality and health impact assessment. Examine whether the evaluation data represent the practice population and the types of result it will process. A tool trained on a different mix of conditions, coding or report styles may fail in ways that are hard to see in an overall score. Check patients with complex histories, those whose results come through less common channels and people who need accessible communication. Do not use group comparisons to infer that a specific patient's report is safe; use them to identify design and monitoring gaps.
Ask a clinical reviewer to examine disagreements between the tool and reference decisions. A numerical score can conceal systematic errors, such as one laboratory's amended reports being ignored. Record whether the error is caused by source data, extraction, classification, interface mapping, human interpretation or a protocol that was itself unclear. Fixing a model threshold may not repair a message feed. A new prompt may not fix missing clinical context. The corrective action must match the failure mode.
Run a controlled trial with explicit stop conditions
Before a tool changes live routing, consider a shadow phase in which it processes results without removing them from the established pathway. Compare its output with the existing clinical review and record disagreements in a restricted environment. A shadow phase is useful only if the practice can observe the full result population, not a hand-picked sample. Set a time window, minimum case mix, reviewers and acceptance criteria in advance. If the supplier changes the model during the trial, record the version and assess whether previous findings remain valid.
Define stop conditions, including an urgent result routed incorrectly, a report lost between systems, a patient mismatch or an unexplained change in output. Some defects require immediate suspension of the AI function while established manual or clinical-system routes continue. Agree who can trigger the stop, who can authorise restart and how results already processed are identified. Do not rely on a supplier promise that the model can be rolled back unless the practice has tested the rollback and the result queue remains intact.
If a trial demonstrates acceptable performance, release one function at a time if possible. For example, a practice may first allow suggested grouping while retaining full clinician review. A later proposal to auto-file some results is a different risk decision and needs its own evidence. Restrict configuration permissions and record who changed a rule. Recheck the local safety case after a material workflow, product, laboratory or population change. The release state should be visible to staff, so no one assumes a function is live when it is still in shadow mode.
Protect patient data and explain the data flow
Test results are sensitive health information. The practice's information governance lead should map what identifiable data leaves the GP clinical system, why it is needed, where it is processed, who receives it, how long it is kept and whether it can be used to develop the supplier's model. The ICO's AI and data-protection guidance addresses lawful processing, fairness, transparency and risk assessment. The exact lawful basis and any data-protection impact assessment must be decided for the actual arrangement. This page does not assert that a generic supplier contract makes the data flow lawful.
Check access controls for administrators, supplier support and temporary staff. A model evaluation dataset is still patient data if people can be identified directly or indirectly. Removing names may not be enough when uncommon results and dates remain. The governance team should approve extraction, pseudonymisation, retention and deletion methods. Incident logs should preserve enough information to investigate an error without copying whole patient reports into a general compliance tool. If the supplier wants to reuse data for training, obtain a specific legal and contractual assessment rather than treating it as a hidden extension of service delivery.
Patients and staff need clear information about how results are handled, especially if an AI tool materially changes prioritisation or messaging. Transparency should describe what the tool does in the local process, where human review occurs and how a person can raise a concern. Avoid the false reassurance that “AI never makes decisions” if the system suppresses or reroutes attention. Equally, avoid implying that a model makes autonomous diagnoses when it only proposes an administrative grouping. Clear language helps patients and staff understand the actual controls.
Keep clinical ownership after go-live
Assign an owner for daily queue reconciliation. Compare the number and types of reports received by the clinical system, the number seen by the tool, the number routed to each queue, the number reviewed and any unmatched items. Include amended and rejected messages. A zero-count alert may be more important than an accuracy report if the upstream feed has stopped. The NHS England clinical messaging guidance warns about undetected systematic faults in delivery. Reconcile at the interface and the clinical action point.
Create a route for staff to report unexpected classifications without having to prove that harm occurred. Review these reports with the clinical safety officer and supplier, and decide whether other patients were affected. Monitor clinical outcomes as well as model outputs. A result routed to an urgent queue is not safe if no clinician opens it. A result marked reviewed is not safe if a necessary patient contact never happened. Check the complete pathway in a sample of cases, including delays, communication and follow-up.
Model and system changes need an approved release process. Record version, reason, supplier evidence, local test results, hazards changed, approval and rollback plan. A change in laboratory coding or a new report template may invalidate previous validation even when the AI model itself has not changed. Re-run targeted tests after such changes. If a new clinical use is proposed, reassess intended purpose, device position, data flow and safety case. A tool bought for sorting should not quietly grow into a diagnostic decision engine through a configuration switch.
Set review dates and stop criteria. A sustained rise in overrides, a pattern of missed report types, unexplained output drift or a safety incident may warrant reverting to the established pathway while the cause is assessed. Keep enough staff capability to operate that fallback. An automated workflow that saves time only while it works but leaves no usable manual route is fragile. Train new and locum clinicians on what the AI does and does not do, where source reports can be seen and how to challenge its output.
A deployment decision record that can be audited
The practice can structure its decision around these evidence questions:
| Decision | Evidence to keep |
|---|---|
| What does the tool do? | Intended purpose, enabled functions, data inputs, outputs and actions it can change. |
| Who owns clinical safety? | Local deployment lead, clinical safety officer, hazard log and sign-off route. |
| What is the regulatory position? | Supplier rationale, relevant device evidence if applicable and unresolved classification questions. |
| Is local performance acceptable? | Representative test set, reference decisions, subgroup review, errors and acceptance criteria. |
| Does the whole route work? | Queue reconciliation, human review, patient communication and amended-result tests. |
| Is information governance resolved? | Data-flow map, lawful-basis and impact assessment decisions, supplier controls and retention. |
| Can the practice stop safely? | Fallback process, stop trigger, affected-result trace and restart approval. |
| Does performance hold? | Live audit, incident trends, version changes and periodic clinical review. |
Keep the decision proportionate to the actual function, but do not omit a hazard because a supplier calls the tool administrative. A sorter that determines who sees a critical report affects clinical risk even if it never suggests a treatment. The practice should record unresolved issues rather than mark every box green to meet a launch date. If validation does not support the proposed use, limit the function, choose another product or retain the existing pathway. Deferral is a legitimate safety decision.
Give Complys the governance role it can support
The Complys clinics page describes broad compliance task and evidence management. Subject to actual product and permissions review, Complys could help assign a supplier assurance review, record completion of a local safety assessment, schedule a post-launch audit and track corrective actions. It is not verified as an AI test-result classifier, GP clinical system, pathology interface, patient messaging platform or DCB0160 safety-case generator. Clinical results and identifiable case review belong in approved clinical systems.
If Complys records that a risk review was completed, the named clinical safety officer still needs to inspect and approve the underlying documents in the correct system. A task due date cannot establish that a classifier is safe. A closure status should point to the decision, evidence, version and next check. Product managers should confirm what Complys can actually store, restrict and audit before any implementation copy claims integration or automated results governance.
CTA: Assess whether Complys can coordinate your AI deployment governance actions and review evidence while clinical results stay in approved GP systems. Related tool opportunity: An AI results-routing deployment worksheet covering intended purpose, input coverage, local validation, clinical safety, information governance, fallback and monitoring. It should not generate a safety case or assess patient results automatically. Internal links out: GP test-results backlog; GP clinical-system outage; incident corrective-action effectiveness. Internal links in proposed: Broad GP CQC guide and any future validated digital clinical-safety hub after route review. Cannibalisation note: This owner is the deployment decision for AI-assisted routing of incoming clinical results. It does not own all GP AI use, an existing backlog, ordinary results management, patient-facing automated release or a specific vendor review. Recheck triggers: NHS England AI and digital clinical safety guidance, DCB0129 or DCB0160 changes after the 2026 review, MHRA software-device guidance, ICO AI guidance, CQC GP mythbuster 46, supplier functionality and Complys product proof.
Complys keeps the records, actions and evidence behind this workflow in one place.
See how Complys helps →Primary sources
- Complys clinics. Public product positioning checked; no clinical AI, pathology or GP messaging capability verified.