Introduction and Context
Acute kidney injury (AKI) is common in hospitalized patients and strongly linked to morbidity and mortality. Over the past two decades, a large and growing literature has evaluated biochemical biomarkers (for example, NGAL, KIM-1, TIMP‑2·IGFBP7) for early detection, differential diagnosis and prognosis of AKI. However, enthusiasm has been tempered by inconsistent results across studies and uncertainty about clinical utility. Many diagnostic test accuracy (DTA) studies differ in their patient selection, reference standards, timing of sampling, assay methods and reporting — factors that prevent meaningful synthesis of evidence and slow translation of promising biomarkers into practice.
Prompted by these gaps, an international expert panel convened to review the literature and propose extension items to the STARD (Standards for Reporting Diagnostic Accuracy) 2015 checklist specifically for AKI biomarker studies. The product — STARDaki — was published as a consensus extension in Intensive Care Medicine (Yu et al., 2026). It is intended to make studies more comparable, reproducible and clinically informative.
Why this matters now: the systematic review underpinning STARDaki found 122 DTA studies of AKI biomarkers, but only a minority were reported well enough to support pooled analysis; only 16 studies met high-quality criteria for diagnosis within 48 hours of sampling. Poor reporting hampers clinical translation, regulatory assessment and meta-analysis. STARDaki targets this problem by specifying AKI-relevant reporting elements beyond generic STARD guidance.
Key references: the STARD 2015 update (Bossuyt et al., BMJ 2015) and the QUADAS‑2 tool for assessing DTA risk of bias (Whiting et al., Ann Intern Med 2011) are foundational. The Kidney Disease: Improving Global Outcomes (KDIGO) 2012 AKI definition is the recommended reference standard in STARDaki.
New Guideline Highlights (STARDaki core messages)
– Purpose: STARDaki is a focused extension to STARD 2015 designed to standardize reporting of diagnostic accuracy studies of AKI biomarkers.
– Primary aims: improve clarity on patient selection and timing, require explicit and consistent AKI reference standards, mandate assay and sample-handling details, and require reporting of test-retest reliability and key clinical outcomes.
– Immediate impact: make studies usable for meta-analysis, increase generalizability, and accelerate clinical validation
Key takeaways for clinicians and researchers
– Always report the intended clinical use of the biomarker (early diagnosis, prediction of severe AKI, need for renal replacement therapy [RRT], prognosis).
– Use KDIGO criteria as the reference standard and report precisely how baseline creatinine was established.
– Report the time interval between biomarker sampling and the AKI reference assessment; STARDaki emphasizes a practical diagnostic window of 48 hours when assessing early-detection performance.
– Provide full assay details (manufacturer, lot, limits of detection, calibration, sample type and processing) and test-retest or within-subject reproducibility.
– Prespecify thresholds or show why thresholds were derived post-hoc; include sensitivity, specificity, predictive values, likelihood ratios with confidence intervals, and prevalence in the study population.
Updated Recommendations and Key Changes from STARD 2015
STARDaki is not a replacement for STARD but an extension tailored for AKI biomarker DTA studies. Important additions include:
– Explicit AKI reference standard: STARDaki recommends using KDIGO (2012) diagnostic criteria as the primary reference standard and requires reporting how baseline creatinine was determined (measured pre-admission value, back-calculation, or admission value) and any adjudication process. (Change: STARD suggested describing the reference standard, STARDaki requires KDIGO and baseline methodology.)
– Time‑window specification: STARDaki requires reporting the elapsed time between index test sampling and the outcome assessment used to define AKI; it prioritizes analysis of diagnostic performance for AKI occurring within 48 hours of sampling. (Change: STARD did not define disease‑specific timing.)
– Assay and specimen handling granularity: beyond stating the index test, STARDaki demands detail on specimen type (serum, plasma, urine), collection, centrifugation, storage temperature and duration, freeze–thaw cycles, reagents, instrument model, and lot numbers where applicable.
– Test‑retest reliability: STARDaki requires reporting within-subject reproducibility and technical repeatability (coefficients of variation, intraclass correlation) because biological variability and assay imprecision are common contributors to inconsistent test performance.
– Clinical endpoints and follow-up: authors must report prespecified clinical outcomes linked to prognostic claims (need for RRT, in-hospital mortality, renal recovery at 7–90 days) and the length of clinical follow-up.
– Patient phenotype reporting: detailed comorbidity, baseline kidney function, medication exposures (nephrotoxins), and setting (ICU, post‑cardiac surgery, ED) must be reported to allow subgroup analyses.
These extensions were driven by the systematic review underpinning STARDaki that documented poor STARD checklist compliance and heterogeneity in methods and reporting (Yu et al., 2026).
Topic-by-Topic Recommendations (Practical checklist)
Below are succinct STARDaki items grouped by topic. Investigators should include these in manuscripts and protocols.
Patient selection and setting (mandatory)
– Define the intended population and clinical setting (ICU, OR, ED, ward).
– Use consecutive or random sampling; avoid case-control sampling for DTA.
– Report inclusion/exclusion criteria, enrolment dates, and screening logs.
Reference standard and timing (mandatory)
– Use KDIGO 2012 criteria for AKI diagnosis and report precisely how baseline creatinine was obtained.
– Report timing: time of index test sample, and time(s) when creatinine and urine output used to define AKI were measured. Prioritize analyses for AKI occurring within 48 hours of sampling.
– Describe any adjudication panel and blinding between index and reference assessments.
Index test and assay details (mandatory)
– Provide full assay details: manufacturer, platform, lot, calibration; limits of detection/quantification; units; preanalytical handling (sample type, timing, processing, storage), and acceptable sample age/temperature.
– Report analytic validation data (precision, bias) and reference intervals if available.
Statistical reporting (mandatory)
– Prespecify primary thresholds or justify data-driven cutoffs; report AUC, sensitivity, specificity, likelihood ratios, PPV/NPV with 95% CIs.
– Report prevalence of AKI in the study cohort and calibration measures if predictive models are used.
– Handle missing data transparently and report flow diagrams (STARD flowchart with STARDaki additions).
Reproducibility and reliability (mandatory)
– Include intra- and inter-assay CVs; for biomarkers with suspected intra-individual variability, report test-retest reliability in a subset (timeframe specified).
Clinical outcomes and follow-up (recommended)
– Report clinical endpoints relevant to the biomarker’s intended use (RRT, mortality, length-of-stay, renal recovery); report timing and censoring strategies.
Special populations (recommended)
– Provide subgroup analyses or separate reporting for CKD (stages), paediatrics, transplant recipients, and sepsis/cardiac surgery patients.
Data sharing (encouraged)
– Share de-identified datasets or minimum elements to enable meta-analysis and external validation.
Recommendation Grading and Evidence
STARDaki is a consensus-based reporting extension rather than an evidence-graded clinical practice guideline. Recommendations were developed via a modified Delphi process among 17 experts and reflect face validity based on the systematic review findings. The principal evidence supporting the extension is the poor reporting quality identified across 122 DTA studies (Yu et al., 2026) and well-established methodological standards for DTA (STARD 2015; QUADAS‑2).
Expert Commentary and Insights
Panel perspectives emphasized that poor reporting — not intrinsic failure of biomarkers — explains much of the inconsistent literature. Key expert themes included:
– Standardizing the reference standard and timing is essential: variability in baseline creatinine assignment and in the window used to define AKI were major sources of heterogeneity.
– Assay harmonization is necessary but difficult: different commercial platforms and home-brew assays cause variability; reporting lot and calibration data is a small but important step.
– Test‑retest data matter: biological variability (for example, due to diuresis, fluid balance or circadian factors) can be large compared with small biomarker signal changes.
– Clinical utility needs linkage to actionable thresholds and outcomes: a biomarker must either change management or robustly predict outcomes that justify different care pathways.
Controversies and open questions
– Exact thresholds for biomarkers remain unsettled; STARDaki requires transparent threshold derivation but stops short of endorsing particular cutoffs.
– Use of biomarkers for triage vs prognosis: experts differed on whether early-warning biomarkers should be evaluated principally for diagnosis within a short time window (48 hours) or for longer-term prognosis.
– How to handle baseline kidney function uncertainty: no consensus exists for the optimal imputation of baseline creatinine when pre-admission values are absent; STARDaki recommends explicit reporting and sensitivity analyses.
Practical Implications for Researchers, Clinicians and Journals
– For researchers: Build STARDaki items into study protocols and case report forms. Funders and ethics committees should expect these elements in proposals.
– For journal editors and peer reviewers: Adopt STARDaki as a required extension for manuscripts reporting AKI biomarker DTA; mandate a completed STARD vs STARDaki checklist at submission.
– For clinicians and guideline developers: Demand studies that report per STARDaki before considering biomarker-based care pathways; recognize that many published biomarker claims cannot be confidently generalized because of reporting gaps.
A brief clinical vignette
Sarah is a 68-year-old woman admitted to the ICU after major abdominal surgery. On post-op day 1 she appears oliguric and hypotensive. A urine biomarker is measured and reported as elevated. Under STARDaki-based reporting, the study validating that biomarker would have stated whether the intended use was diagnosis of AKI within 48 hours, how baseline creatinine was defined, the exact sampling-to-outcome interval, assay handling and repeatability, and whether the biomarker prediction was linked to meaningful outcomes (need for RRT, 7‑day renal recovery). With such transparent reporting, Sarah’s clinicians could better assess whether a positive biomarker result should change monitoring or trigger specific preventive measures.
Future directions and research priorities
– Prospective multicenter cohorts designed and reported per STARDaki to generate reproducible estimates.
– Harmonization studies comparing assays across platforms and efforts toward international reference materials.
– Individual participant data meta-analyses using STARDaki elements to derive externally valid thresholds and decision algorithms.
– Cost-effectiveness studies incorporating STARDaki-quality evidence to determine the value of biomarker-guided care pathways.
References
– Yu H, Li Y, Mo GP, Luo Z, Zarbock A, Fuhrman D, et al. STARDaki: a consensus-based STARD extension for standardized reporting of diagnostic accuracy in acute kidney injury. Intensive Care Med. 2026 Jul 28. PMID: 42517928. https://pubmed.ncbi.nlm.nih.gov/42517928/
– Bossuyt PM, Reitsma JB, Bruns DE, Gatsonis CA, Glasziou PP, Irwig L, et al. STARD 2015: an updated list of essential items for reporting diagnostic accuracy studies. BMJ. 2015;351:h5527. doi:10.1136/bmj.h5527.
– Whiting PF, Rutjes AW, Westwood ME, Mallett S, Deeks JJ, Reitsma JB, et al. QUADAS‑2: a revised tool for the quality assessment of diagnostic accuracy studies. Ann Intern Med. 2011;155(8):529–36. doi:10.7326/0003-4819-155-8-201110180-00009.
– Kidney Disease: Improving Global Outcomes (KDIGO) Acute Kidney Injury Work Group. KDIGO Clinical Practice Guideline for Acute Kidney Injury. Kidney Int Suppl. 2012;2(1):1–138. doi:10.1038/kisup.2012.1.
– Kashani K, Al-Khafaji A, Ardiles T, Artigas A, Bagshaw SM, et al. Discovery and validation of cell cycle arrest biomarkers in human acute kidney injury. Crit Care. 2013;17(1):R25. doi:10.1186/cc11975.
Note: STARDaki is a reporting extension aimed at standardizing how AKI biomarker DTA studies are presented to maximize interpretability and clinical relevance. Investigators, journals, funders and regulators should collaborate to implement these standards so the field can move from many small, inconsistent studies to robust, generalizable evidence that can inform patient care.

