Work overview

Section 02 of 03

Review

Artificial Intelligence-Enabled Electrocardiography for Potassium Abnormality Detection and Estimation: A Systematic Review

Baldeep Kaur, Guntas Singh Gill, Kiranpreet Kaur, Ankur Chaudhary, and Ashutosh Garg · 2026

Contents

Section 02 of 03

  1. 01Introduction and background
  2. 02Review
  3. 03Conclusions
Text size
Work overview

Section 2 of 3

Review

Baldeep Kaur, Guntas Singh Gill, Kiranpreet Kaur, Ankur Chaudhary, and Ashutosh Garg · about 22 minutes

Methodology

Design and Reporting Standards

This diagnostic-performance test systematic review was conducted and reported in accordance with PRISMA-DTA and PRISMA 2020 statements, with PRISMA-S used to guide literature search reporting [6-8]. Methodological quality and applicability were assessed using QUADAS-2 [9]. The review was not prospectively registered.

Eligibility Criteria (PICOS Framework)

Eligibility criteria followed the PICOS framework and is shown in Table 1.

Component | Description
Population (P) | Adult human participants with electrocardiographic data paired with a laboratory potassium measurement.
Index test (I) | Artificial intelligence- or machine learning-based ECG methods, including deep learning, neural networks, conventional machine-learning algorithms, and previously developed AI-enabled ECG models applied to 12-lead, reduced-lead, single-lead, wearable, monitor-derived, image-based, or feature-derived ECG data.
Comparator / reference standard (C) | Reference standard: Laboratory measurement of serum, plasma, blood, or whole-blood potassium obtained in temporal pairing with the ECG (required for every included study). Optional comparator: Conventional ECG interpretation, clinician assessment, or point-of-care testing, when a study reported one alongside the AI/ML index test.
Outcomes (O) | Hyperkalemia detection, hypokalemia detection, categorical potassium-status classification, or continuous potassium concentration estimation, with extractable diagnostic accuracy or regression performance metrics (e.g., AUROC, sensitivity, specificity, calibration, mean absolute error, or correlation).
Study design (S) | Peer-reviewed retrospective or prospective primary studies evaluating diagnostic accuracy, prediction models, external validation, monitoring, or clinical implementation of AI/ML ECG methods for potassium assessment.

We included peer-reviewed journal articles and full-length conference proceedings papers for which a scientific-review process could be documented. Eligible studies were published between 2016 and 2026 and involved adult participants in whom AI/ML-based ECG methods were applied to potassium-specific outcomes, with extractable diagnostic or model-performance results against a paired laboratory potassium measurement as the reference standard. Other reports not falling into above criteria were excluded from the primary evidence synthesis.

Reviews, editorials, commentaries, protocols, and registry records without eligible peer-reviewed diagnostic data were excluded from the primary evidence set. We also excluded pediatric-only and non-human studies; in vitro or simulation-only studies; records outside the 2016-2026 eligibility window; non-ECG studies; studies without potassium diagnostic or estimation target; ECG studies that relied on conventional interpretation, rule-based criteria, or standard statistical analysis rather than a clearly defined AI/ML model; studies lacking paired laboratory potassium measurements, and studies without extractable potassium-specific performance outcomes. Related non-primary records were retained separately, only to clarify publication lineage, dataset overlap, trial or preprint status, and competing-review context. Paywall status alone was not used as a reason for exclusion.

Information Sources and Search Strategy

We conducted the principal formal searches in PubMed/MEDLINE, Europe PMC, ClinicalTrials.gov, the Cochrane Library/CENTRAL, medRxiv, OpenAlex, and Semantic Scholar. Crossref was used for DOI and bibliographic verification rather than as a primary literature-discovery source. Backward and forward citation chasing and targeted checks of arXiv, publisher pages, repositories, and engineering metadata sources, including DBLP and IEEE metadata or DOI pages, were used to identify related publications and to clarify publication status, bibliographic details, dataset overlap, and model lineage. These activities were record-specific supplementary verification procedures rather than comprehensive formal database searches. We did not directly search Scopus, Web of Science, Embase, or IEEE Xplore; the potential effect of this is considered in the limitations. Preliminary scoping searches began in April 2026; the structured, source-specific searches were conducted in May 2026 and expanded and reconciled through June 2026, with a final update check on June 18, 2026. No publication-date or language restriction was applied at the search stage. The included peer-reviewed primary studies were published between 2016 and 2026. Appendix 1 summarizes the information sources and supplementary verification procedures, while Appendix 2 presents the principal formal source-specific queries.

Study Selection Process

Records were deduplicated using DOI, PMID, PMCID, registry identifiers, arXiv identifiers, stable URLs, and normalized titles. Title and abstract screening used predefined eligibility criteria, with screening decisions independently verified by two reviewers. Records considered potentially eligible or of uncertain relevance advanced to full-text eligibility assessment. Near-miss records combining ECG, potassium or electrolyte, and AI/ML or signal-processing concepts underwent additional review before final classification as primary or non-primary evidence. Full-text eligibility decisions were reviewed collaboratively by the author team, and uncertainties were resolved through discussion and consensus. Each excluded record was assigned a controlled primary exclusion reason. arXiv identifiers were included in deduplication because preprint records encountered during targeted arXiv checks, citation tracing, and publication-status verification were retained to link preprints with related peer-reviewed reports.

Data Extraction

A structured form was used to extract bibliographic details, study characteristics, ECG source, reference standard, potassium threshold, model characteristics, validation design, performance metrics, calibration, clinical utility, funding, conflicts of interest, and dataset overlap. Initial data extraction was performed by designated reviewers and independently verified by additional members of the review team against the original articles, supplementary materials, and bibliographic metadata. Any discrepancies were resolved through discussion and consensus among the authors. Classification metrics were extracted and analysed separately from continuous estimation metrics, and any value derived from figures or confusion matrices was explicitly identified as derived. Author-reported confidence intervals were extracted where available; no new uncertainty intervals were calculated. Overlapping cohorts and related model lineages were mapped before evidence synthesis.

Quality Assessment and Risk of Bias

Risk of bias and applicability were assessed using QUADAS-2 [9]. Risk of bias was evaluated across the four QUADAS-2 domains-patient selection, index test, reference standard, and flow and timing-while applicability concerns were assessed across the first three domains, consistent with the tool’s structure. Initial QUADAS-2 assessments were performed by designated reviewers and independently reviewed by additional members of the review team. Any discrepancies or uncertain judgments were resolved through discussion and consensus among the authors following re-examination of the original reports. Given that the index tests comprised AI/ML models, additional AI-specific considerations were documented, including validation design, calibration, explainability, transparency, deployment readiness, dataset overlap, and the potential for optimism bias. Tailored signaling questions and final study-level judgments were recorded for all included studies.

Data Synthesis

We conducted a structured narrative synthesis, grouping studies according to potassium target, potassium threshold, ECG modality, clinical setting, validation maturity, unit of analysis, performance metric, risk of bias, and dataset overlap. Before deciding whether quantitative synthesis was appropriate, the review team applied the following comparability criteria: the same target condition; a comparable potassium threshold; a comparable ECG modality and input format; a comparable clinical setting; a comparable validation design and unit of analysis; cohort and model independence, with no shared dataset or model lineage; and sufficient reported diagnostic-accuracy data to reconstruct a 2 × 2 contingency table or an operating point with a measure of uncertainty. Studies were considered eligible for pooling only if all of these criteria were met. Otherwise, findings were synthesized narratively without quantitative pooling. Study-level estimates are reported with 95% confidence intervals where the primary studies provided them.

Results

Study Selection Process

Of 366 records assessed for eligibility, 33 peer-reviewed primary AI/ML studies were included [3,4,10-40]. An additional 65 non-primary records were retained separately to document publication lineage, trial status, preprints, competing reviews, signal-processing or statistical context, and dataset overlap, while 268 records were excluded from the primary evidence set. The PRISMA flow diagram of study selection, including categorical reasons for exclusion are shown in Figure 1.

Figure 1: PRISMA-based study selection flow diagramStudy selection for the present systematic review. The diagram begins after deduplication.PRISMA: Preferred Reporting Items for Systematic Reviews and Meta-Analyses; AI/ML: artificial intelligence/machine learning; ECG: electrocardiography

Figure 1: PRISMA-based study selection flow diagramStudy selection for the present systematic review. The diagram begins after deduplication.PRISMA: Preferred Reporting Items for Systematic Reviews and Meta-Analyses; AI/ML: artificial intelligence/machine learning; ECG: electrocardiography

Study Characteristics

The 33 included primary AI/ML studies evaluated four AI-ECG tasks. Because many studies addressed more than one task, these categories were not mutually exclusive. Hyperkalemia detection was most common (28 studies), followed by hypokalemia detection (14), continuous potassium estimation (10), and categorical potassium-status classification (seven); 11 studies addressed both hyperkalemia and hypokalemia, and five also estimated other electrolytes such as sodium or calcium. Studies used diverse ECG inputs, including 12-lead and multi-lead recordings, reduced-lead and single-lead recordings, smartwatch and wearable signals, monitor-derived recordings, morphology- or feature-based inputs, and image-, beat-, or segment-level representations. Cohorts spanned emergency departments, dialysis units, intensive care, general inpatient settings, and remote monitoring, and several studies used large institutional ECG repositories. Figure 2 illustrates the distribution of studies by potassium target and validation maturity, while Table 2 summarizes the evidence base by potassium target.

Figure 2: Evidence map of included primary AI/ML studies by potassium target and validation maturityBased on data extracted from the included primary AI/ML studies [3,4,10-40]. Cell shading indicates the number of primary studies in each cell. Because individual studies could address more than one potassium target, cell counts exceed the 33 included primary AI/ML studies.AI/ML: artificial intelligence/machine learning; ECG: electrocardiography

Figure 2: Evidence map of included primary AI/ML studies by potassium target and validation maturityBased on data extracted from the included primary AI/ML studies [3,4,10-40]. Cell shading indicates the number of primary studies in each cell. Because individual studies could address more than one potassium target, cell counts exceed the 33 included primary AI/ML studies.AI/ML: artificial intelligence/machine learning; ECG: electrocardiography

Potassium target | Studies (n) | Typical study characteristics | Representative findings | Key findings
Hyperkalemia detection | 28 | Common thresholds ranged from K+ >5.0 to >=6.5 mmol/L; most studies used binary detection, although some multiclass models included a hyperkalemia class. | Reduced-lead, single-lead, 12-lead, wearable, monitor-derived, and image-based ECG models were evaluated. Several studies reported AUROCs in the high 0.8 to mid-0.9 range. | Thresholds, ECG-to-laboratory timing, study populations, and validation designs varied substantially across studies.
Hypokalemia detection | 14 | Usually K+ <3.5 mmol/L; one study of thyrotoxic periodic paralysis used K+ <=3.0 mEq/L [34]. | Most studies evaluated hypokalemia within combined dyskalemia or electrolyte models rather than dedicated hypokalemia models. | External validation was less frequently reported than for hyperkalemia studies.
Categorical potassium-status classification | 7 | Multiclass potassium categories or ordered potassium-status labels. | Deep-learning, neuro-fuzzy, wavelet/deep-learning, and semi-supervised approaches were reported. | Class definitions and performance metrics varied across studies.
Continuous potassium estimation | 10 | Continuous estimation (regression); study- defined thresholds | Reported mean absolute error ranged from approximately 0.26-0.53 mmol/L (where available). | Regression outcomes were reported separately from diagnostic accuracy metrics.

Detailed study characteristics and representative performance metrics for all 33 included primary AI/ML studies are provided in Appendix 3.

Evidence by Potassium Target

Hyperkalemia detection was the largest evidence group and varied in potassium threshold (commonly 5.0-6.5 mmol/L), study population, ECG modality, and validation design. Hypokalemia detection was reported by 14 studies, although most detected hypokalemia within a combined dyskalemia model that also addressed hyperkalemia. Only two studies evaluated hypokalemia as a dedicated screening task [35,36], whereas one focused on thyrotoxic periodic paralysis [34]. Seven studies evaluated categorical potassium-status classification, and 10 reported continuous potassium estimation using regression metrics such as mean absolute error, root mean squared error, and correlation. Five studies also evaluated potassium alongside other electrolytes. Detailed study characteristics are provided in Appendix 3.

Validation Maturity and Clinical Utility

Validation maturity varied across the evidence base: seven studies were development-only, 11 used internal validation, 10 used external or temporal validation, three included prospective validation, and two provided clinical monitoring, utility, or interventional evidence [26,28]. One study evaluated serial monitoring during treatment of severe hyperkalemia [26], whereas another reported the only identified pragmatic randomized trial of an AI-ECG alert [28]. The randomized trial found no significant overall increase in hyperkalemia- or hypokalemia-related treatment within three hours, although benefit was observed in the subgroup flagged as hyperkalemic [28].

Diagnostic and Model Performance

Performance reporting was heterogeneous. Binary and categorical studies reported AUROC, sensitivity, specificity, predictive values, accuracy, F1 score, or partial confusion-matrix data, while continuous-estimation studies reported regression error and correlation. Among hyperkalemia studies with internal or external validation, AUROCs were frequently in the high 0.8 to mid 0.9 range [3,10,13,18,23,24,31,33], and AI/ML models performing continuous potassium estimation reported mean absolute errors around 0.26-0.53 mmol/L in selected cohorts [4,13,23,37,39]. Study-level performance for each included study, with classification and continuous-estimation metrics shown separately, is tabulated in Appendix 3.

Risk of Bias and Applicability

Overall risk of bias was low in four studies, unclear in 23, and high in six; overall applicability concern was low in 18 studies, unclear in 10, and high in five. Recurring concerns included case-control enrichment, artificial class balancing, incompletely reported ECG-to-laboratory timing, unclear train-test split integrity, limited threshold prespecification, sparse calibration reporting, limited reproducibility, and incomplete paired diagnostic data. Domain-level QUADAS-2 judgments are summarized in Table 3.

Domain | Low | Unclear | High | Not applicable | Common findings
Patient selection: risk of bias | 14 | 13 | 6 | 0 | Case-control enrichment and sampling restrictions were common concerns.
Index test: risk of bias | 7 | 25 | 1 | 0 | Thresholding, model transparency, split integrity, and calibration reporting were often incomplete.
Reference standard: risk of bias | 27 | 6 | 0 | 0 | Laboratory potassium was generally appropriate, but specimen and timing details varied.
Flow and timing: risk of bias | 9 | 24 | 0 | 0 | Electrocardiography-laboratory timing and missing data were frequent reporting limitations.
Patient selection: applicability | 21 | 7 | 5 | 0 | Some populations were narrow or enriched compared with general clinical use.
Index test: applicability | 32 | 1 | 0 | 0 | Most index tests matched the strict AI/ML review question, with some engineering or representation-level caveats.
Reference standard: applicability | 27 | 6 | 0 | 0 | Reference-standard applicability was usually low concern.
Overall risk of bias | 4 | 23 | 6 | 0 | Overall risk of bias was rated as unclear in most studies.
Overall applicability concern | 18 | 10 | 5 | 0 | Applicability was mostly low concern but not uniformly so.

One included study carried an editorial expression of concern relating to data availability [18,41]. The study was retained in the review and considered during risk-of-bias assessment and sensitivity checks of the pooling decision rather than being excluded automatically.

Meta-Analysis Feasibility

No clinically coherent, independent group of studies met the comparability criteria for quantitative synthesis. Even within the most plausible candidate subgroup - hyperkalemia detection near K+ ≥5.5 mmol/L - studies differed in setting, ECG modality or input format, validation design, operating-point reporting, and cohort or model independence. The principal barriers were heterogeneous potassium targets and thresholds, varied ECG modalities and clinical settings, differing validation designs and reporting units, dataset overlap or shared model lineage, and incomplete paired diagnostic data. Related ECG model families, competition-derived subsets, and retained signal-processing lineage records were therefore not treated as independent validations, and these overlaps were assessed before synthesis. Sensitivity checks of the pooling decision did not alter this conclusion. Repeating the comparability assessment after restricting the evidence to author-reported metrics, or after excluding studies at high risk of bias, the study carrying an editorial expression of concern, or development-only studies, still yielded no clinically comparable and independent group suitable for pooling. Although a previous preprint meta-analysis pooled potassium-detection studies across institutions [5], our broader review scope together with the overlap- and lineage-independence assessment supported a structured narrative synthesis rather than quantitative pooling. Consequently, no pooled estimates, summary sensitivity, specificity, AUROC, bivariate model, HSROC model, forest plot, or summary ROC curve were generated.

Discussion

This systematic review synthesized evidence from 33 peer-reviewed primary AI/ML studies evaluating AI-ECG for hyperkalemia detection, hypokalemia detection, categorical potassium-status classification, and continuous potassium estimation against a paired laboratory measurement as the reference standard [3,4,10-40]. To our knowledge, this is one of the first peer-reviewed systematic reviews to evaluate all four potassium-related AI-ECG tasks, while also appraising validation maturity and the independence of datasets and model lineage across studies. Three findings stand out. First, the evidence base is dominated by hyperkalemia detection, with hypokalemia detection, multiclass classification, and continuous potassium estimation represented by smaller and more heterogeneous subsets. Second, reported diagnostic performance was frequently high, but most evidence came from retrospective development or internal validation, with a smaller subset reaching external or temporal validation and only a few studies progressing to prospective or interventional evaluation. Third, and most importantly for clinical practice, among binary detection studies reporting predictive values, negative predictive values were generally high while positive predictive values were often low, indicating that current AI-ECG functions primarily as a rule-out and triage tool rather than a confirmatory test. Taken together, these findings position AI-ECG as a promising but still investigational adjunct to, rather than a replacement for, laboratory potassium measurement.

Diagnostic and Estimation Performance in Clinical Context

Overall diagnostic performance was encouraging. Among hyperkalemia studies with internal or external validation, AUROCs were frequently in the high 0.8 to mid 0.9 range [3,10,13,18,23,24,31,33], and single-lead or wearable implementations extended detection beyond the conventional 12-lead ECG [12,23,25]. However, confidence in these estimates varied according to validation maturity, with externally or prospectively validated cohorts providing the most clinically applicable evidence, whereas findings from smaller, engineering-oriented, or class-balanced development studies were less readily generalizable.

From a clinical perspective, predictive values are more informative because they depend on disease prevalence in the population being tested. In unselected populations, positive predictive value remained low despite good discrimination - for example, 1.4% with a corresponding negative predictive value of 99.9% [10], 3% and 99.8% in the emergency department, and 13.2% and 99.4% in intensive care [24,27]. Thus, in the populations and at the thresholds studied, a negative AI-ECG result may assist in ruling out dangerous hyperkalemia, whereas a positive result should prompt confirmatory laboratory testing rather than establish the diagnosis. As expected, positive predictive value increased in higher-prevalence populations, reaching approximately 15% in the emergency department and 29% in intensive care among patients with advanced chronic kidney disease and an estimated glomerular filtration rate of <30 mL/min [24]. Similarly, several models were developed specifically in dialysis and end-stage renal disease populations, where hyperkalemia is more prevalent [17,23,25,26]. Models performing continuous potassium estimation reported mean absolute errors of approximately 0.26-0.53 mmol/L [4,13,23,37,39]. These error estimates suggest possible utility for trend monitoring, but the included studies did not establish a consistent, prospectively validated clinical-acceptability margin. Errors of this magnitude may be clinically important near treatment thresholds; continuous AI-ECG estimates should therefore not be interpreted as exact potassium concentrations or used to direct treatment without laboratory confirmation. Diagnostic performance for hypokalemia was generally lower than for hyperkalemia, with AUROCs frequently below 0.8, particularly when hypokalemia was evaluated within combined dyskalemia models or using single-lead ECGs [15,35]. This likely reflects the subtler and less specific electrocardiographic manifestations of hypokalemia, such as T-wave flattening and U waves, which provide weaker signals for AI models to detect. Likewise, multiclass potassium-status classification should not be equated with binary detection of clinically significant dyskalemia, as accurate classification across ordered categories may still fail to identify patients at the clinically important extremes [4,11,13,22,30,37,39].

Comparison With Conventional ECG Reading and Point-of-Care Testing

The clinical rationale for AI-ECG lies in the limitations of conventional electrocardiographic interpretation. Classical ECG features of dyskalemia, such as peaked T waves and QRS widening, are inconsistently present and correlate poorly with the measured potassium concentration [2]. This was reflected in the included studies: when a smartphone-based AI-ECG analyzer was compared with board-certified emergency physicians for hyperkalemia detection, the model achieved an AUROC of approximately 0.90 compared with 0.66 for physician consensus, with physician sensitivity of only about 20% [16]. These findings suggest that AI models can identify electrocardiographic patterns that are not readily recognized by visual interpretation alone.

The more relevant bedside comparator, however, is point-of-care potassium testing. In one critical-care study, AI-ECG was slightly less accurate than point-of-care testing obtained from arterial blood gas analysis (AUROC 0.884 vs. 0.933), although both became available approximately 35 minutes before the central laboratory result [27]. AI-ECG should therefore be viewed as complementary rather than competitive; where point-of-care testing is available, it remains the preferred diagnostic approach, whereas AI-ECG offers an immediate, non-invasive assessment from an ECG that is often already being recorded while laboratory confirmation is pending.

Generalizability, Calibration, and Transportability

The transportability of AI-ECG - whether performance is maintained across different institutions, populations, devices, and time periods-remains one of the principal evidence gaps. Only 15 of the 33 included primary AI/ML studies reached external, temporal, or prospective validation, or provided clinical-monitoring, utility, or implementation (10 external or temporal, three prospective, and two clinical-monitoring/utility or interventional studies), leaving most performance estimates based on internal or development cohorts. Although several models maintained good performance during external validation, others showed reduced accuracy across institutions and recording modalities, particularly with single-lead ECGs [18,23,35].

These findings are consistent with the broader AI literature, in which deep-learning models may learn dataset-specific rather than disease-specific features, limiting generalizability across sites [42]. Calibration-whether predicted probabilities or estimated potassium values correspond to observed outcomes-was rarely reported, despite its importance for clinical decision-making. A model may therefore discriminate well yet remain poorly calibrated at treatment thresholds [43]. Because the clinical consequences of false-positive and false-negative results differ across emergency, dialysis, intensive care, inpatient, and remote-monitoring settings, setting-specific external validation and recalibration should be considered essential before routine clinical deployment.

Methodological Quality and Risk of Bias

QUADAS-2 assessment placed overall risk of bias as low in only four studies, unclear in 23, and high in six, with applicability concern low in 18, unclear in 10, and high in five. Most unclear ratings reflected incomplete reporting rather than demonstrated methodological flaws, although they nevertheless reduce confidence in the generalizability of the findings. The recurring concerns - case-control enrichment, artificial class balancing, incompletely reported ECG-to-laboratory timing, unclear train-test split integrity, limited threshold prespecification, and sparse calibration - are precisely the features that inflate apparent performance. These methodological features likely explain why several small studies reported near-perfect accuracy (approximately 96%-98%) using balanced or augmented datasets [11,19,20,25,30], whereas larger, clinically representative cohorts reported more modest discrimination (AUROC approximately 0.85-0.92), suggesting that the higher figures largely reflect favorable development conditions rather than superior clinical accuracy. One study additionally carried an editorial expression of concern about data availability [18,41]; this was weighed during risk-of-bias interpretation rather than treated as grounds for exclusion, and the review's conclusions were likewise unchanged when it was omitted.

Relationship to a Prior Electrolyte-Focused Review

A recent systematic review and meta-analysis, currently available as a preprint from 2025, evaluated AI-ECG for multiple electrolyte disturbances including potassium, sodium, and calcium and reported encouraging pooled diagnostic performance for hyperkalemia and hypokalemia while identifying patient-selection bias as the principal methodological concern [5]. Our findings are broadly consistent with and complement that review but extend the evidence in several important ways. Unlike the earlier electrolyte-focused synthesis, our review specifically evaluates AI-ECG across four clinically relevant potassium tasks-hyperkalemia detection, hypokalemia detection, categorical potassium-status classification, and continuous potassium estimation. Furthermore, whereas the earlier review searched the literature through late 2024, our review extends to mid-2026. Because AI-enabled healthcare is evolving rapidly, extending the search from late 2024 to mid-2026 captured several important studies, including the first pragmatic randomized clinical trial, wearable ECG validation, serial bloodless monitoring, and additional external validation studies. Finally, rather than focusing primarily on pooled diagnostic accuracy, we framed the evidence around clinical translation and implementation readiness, with particular emphasis on validation maturity, calibration, transportability, dataset overlap, model lineage, and prevalence-dependent predictive values.

From Diagnostic Accuracy to Clinical Impact

Diagnostic accuracy is necessary but not sufficient; what ultimately matters is whether AI-ECG changes care safely and beneficially. Only two included studies moved beyond diagnostic performance toward clinical monitoring, utility, or interventional evidence [26,28]. The single pragmatic randomized trial provides important insights [28]. At the selected alert threshold, the system achieved high specificity (approximately 99%) but only moderate sensitivity (52.2% for hyperkalemia and 34.7% for hypokalemia). Although the overall rate of hyperkalemia-directed treatment did not differ significantly between study arms (8.0% vs. 7.7%), treatment within three hours increased from 41.6% to 69.1% (hazard ratio 2.23, 95% CI 1.44-3.46) among patients correctly flagged as hyperkalemic. These findings suggest that alert threshold selection and implementation strategy, rather than diagnostic discrimination alone, determine the real-world clinical impact of AI-ECG. A second study explored serial bloodless monitoring during treatment of severe hyperkalemia, showing that AI-derived potassium estimates tracked laboratory values (Bland-Altman mean difference -0.21 mmol/L), although specificity remained modest (61.8%) and patient outcomes were not evaluated [26].

These signals also indicate where AI-ECG might plausibly help first. In dialysis and end-stage renal disease, where hyperkalemia is common (raising positive predictive value) and recurrent, wearable and smartwatch ECG offered a route to ambulatory surveillance [23,25], and patient-specific or repeated-visit personalization further improved accuracy in selected models [15,40]. Conversely, widespread deployment in low-prevalence populations could generate excessive false-positive alerts, leading to alert fatigue, unnecessary investigations, or inappropriate treatment escalation, whereas false-negative results may delay recognition of clinically significant dyskalemia [44]. Because these trade-offs vary across emergency departments, dialysis units, intensive care, inpatient wards, and home monitoring, future implementation studies should evaluate treatment timing, adverse events, clinician override behaviour, workflow burden, and patient outcomes alongside diagnostic accuracy. Overall, the available evidence supports further prospective implementation evaluation of AI-ECG as a potential rule-out, triage, or monitoring adjunct for potassium disorders. It does not yet establish readiness for routine clinical deployment or support replacement of laboratory potassium measurement.

Implications for Practice and Research

For clinical practice, the current evidence supports cautious, setting-specific use of AI-ECG as an adjunctive tool for triage and monitoring rather than as a replacement for laboratory potassium measurement. Its potential clinical role may differ by setting: AI-ECG may support rule-out and triage, while higher-prevalence populations, such as patients with chronic kidney disease or those receiving dialysis, may offer more favorable positive predictive value and monitoring yield. Laboratory confirmation remains essential before potassium-directed treatment. For research, future studies should prioritize prospective multicentre validation in intended-use populations, head-to-head comparison with point-of-care testing, robust external validation and calibration, standardized reporting in accordance with TRIPOD+AI [45] and DECIDE-AI [46], and evaluation of clinically meaningful outcomes alongside diagnostic accuracy. Studies should also clearly distinguish patient-level from beat- or segment-level analyses and explicitly report dataset overlap and model lineage to improve transparency and reproducibility.

Limitations

This review has several limitations. First, the available evidence was highly heterogeneous with respect to potassium targets, diagnostic thresholds, ECG modalities, model architectures, validation strategies, reporting units, and performance metrics, precluding quantitative synthesis. Second, many included studies had unclear or high risk of bias, with limited reporting of calibration, confidence intervals, and other methodological details. Third, some performance measures were derived from published figures, tables, or confusion matrices when not explicitly reported by the authors. Fourth, several studies reported selected operating points or non-standard performance metrics, limiting direct comparison across clinically relevant thresholds. Fifth, we did not directly search Scopus, Web of Science, Embase, or IEEE Xplore because these sources were not accessible to the review team. We attempted to mitigate this limitation through searches of PubMed/MEDLINE, Europe PMC, OpenAlex, Semantic Scholar, and the Cochrane Library/CENTRAL, together with backward and forward citation chasing and targeted publisher, repository, and bibliographic checks. These supplementary methods may have reduced, but cannot eliminate, the risk that eligible studies were missed. Finally, the review was not prospectively registered; some overlapping datasets could not be completely verified despite detailed lineage assessment, and although study selection, data extraction, and risk-of-bias assessment were independently verified by additional reviewers, reviewer-related bias cannot be entirely excluded. These limitations should be considered when interpreting the findings but do not alter the overall conclusion that AI-ECG shows considerable promise while requiring further high-quality prospective validation before routine clinical implementation.