32. Evidence-Based Imaging
In this chapter · 6 sections
🎯 Learning objectives
- Derive sensitivity, specificity, predictive values, and likelihood ratios from the 2x2 diagnostic table, prove the prevalence dependence of PPV/NPV and the prevalence independence of likelihood ratios, and apply Bayes' theorem in odds form to revise the post-test probability of disease from a CT finding.
- Interpret the ROC curve and its area as the threshold-independent summary of the sensitivity-specificity trade-off, relate AUC to the probability of correct discrimination between a diseased and non-diseased pair, and explain why a high AUC does not by itself establish clinical utility.
- Diagnose the architectural biases of diagnostic-accuracy research — spectrum bias, verification/work-up bias, review bias, and incorporation bias — using the STARD reporting standard and the QUADAS-2 risk-of-bias framework, and predict the direction in which each bias distorts measured accuracy.
- Distinguish prognostic from diagnostic research, interpret hazard ratios, Kaplan-Meier curves, the concordance (C) statistic, and calibration, and appraise prognostic-factor and prediction-model studies with QUIPS, PROBAST, and the TRIPOD reporting standard, recognizing optimism and the necessity of external validation.
- Apply a phased model of imaging-biomarker validation that separates technical performance (repeatability coefficient, intraclass correlation, Bland-Altman limits of agreement) from biological and clinical validation, situating each within the O'Connor imaging-biomarker roadmap and QIBA metrology.
- Interpret clinical trials of imaging-directed strategies using intention-to-treat analysis, the distinction between relative and absolute effect measures, and the confidence interval as the range of parameter values compatible with the data, while recognizing the special threats of paradigm creep and outcome-verification bias in imaging trials.
- Synthesize a body of imaging evidence through meta-analysis — including the bivariate and hierarchical summary ROC models for diagnostic data, the quantification of heterogeneity with I-squared, and the detection of small-study/publication effects — and translate the pooled, certainty-rated evidence into a guideline recommendation using the GRADE framework.
01Diagnostic Accuracy Studies
The diagnostic-accuracy study is the empirical foundation on which every claim that a CT sign "rules in" or "rules out" disease ultimately rests, and its mathematics begins with the cross-tabulation of an index test against a reference standard in a table whose cells are true positives (), false positives (), false negatives (), and true negatives (). From these are defined sensitivity , the probability the test is positive given disease, and specificity , the probability the test is negative given health. These are properties of the test conditioned on disease status and are, in principle, stable across populations of differing prevalence — but they are not what the clinician facing a positive scan actually wants to know. The clinically relevant quantities are the positive and negative predictive values, and , and these are exquisitely prevalence-dependent. Writing prevalence as , Bayes' theorem gives , so that a CT finding with and yields when but when — the identical test, identically performed, supports opposite conclusions depending on whom it is applied to. This is the single most important quantitative lesson of diagnostic imaging: accuracy travels with the test, but predictive value travels with the patient population.
The reformulation that liberates inference from prevalence is the likelihood ratio. The positive likelihood ratio and the negative likelihood ratio are properties of the test alone, and they combine with the patient's pre-test odds through Bayes' theorem in its elegant odds form: , where . An above roughly 10 or an below roughly 0.1 produces large, often decisive probability revisions; values near 1 are diagnostically inert regardless of how impressive the raw sensitivity sounds. Because most CT signs are interpreted at more than one level of certainty, the test is better characterized across a continuum of thresholds by the receiver operating characteristic (ROC) curve, which plots against as the decision threshold sweeps from strict to lenient. The area under the curve (AUC) summarizes discrimination in a single threshold-independent number with a precise probabilistic meaning: it equals the probability that a randomly chosen diseased case is assigned a higher index value than a randomly chosen non-diseased case, so is coin-flip discrimination and is perfect separation. A high AUC is necessary but never sufficient for clinical value, because it ignores the prevalence and the relative costs of false-positive and false-negative errors that determine where the operating threshold should actually be set.
The validity of any of these numbers depends entirely on study architecture, and four biases recur. Spectrum bias (Ransohoff and Feinstein) inflates apparent accuracy when the diseased group comprises florid, advanced cases and the controls are healthy volunteers, because the easy discriminations of the extremes are not the borderline discriminations of real practice. Verification (work-up) bias arises when the reference standard is preferentially obtained in index-test-positive patients, distorting both sensitivity and specificity. Review bias occurs when the index test is interpreted with knowledge of the reference result (or vice versa), and incorporation bias when the index test is itself part of the reference standard, producing circularity. The STARD 2015 statement (Bossuyt) standardizes the transparent reporting that lets a reader detect these flaws, and QUADAS-2 (Whiting) provides the structured risk-of-bias and applicability appraisal across the patient-selection, index-test, reference-standard, and flow-and-timing domains. The disciplined physician reads a diagnostic-accuracy claim not as a fixed sensitivity to be memorized but as an estimate whose magnitude and credibility are conditioned on the spectrum studied, the reference standard chosen, and the blinding maintained.
🖐️ Threshold and measurement variability on a real abdominal CT
Make concrete that diagnostic accuracy (Se/Sp), the ROC threshold, and quantitative-biomarker repeatability all hinge on a reader-controlled decision boundary, so that window/level and caliper variability are not nuisances but the very quantities evidence-based imaging measures and the biases QUADAS-2 interrogates.
A real abdominal CT in true Hounsfield units. Why a viewer belongs in a chapter on evidence: every diagnostic-accuracy and biomarker study reduces an image to a binary call or a number, and both depend on where the reader places a threshold. Try it: cycle the Liver and Soft tissue presets and watch a lesion's apparent border — and therefore any caliper measurement or "abnormal versus normal" judgment — shift with window width and level. The leniency of that implicit threshold is exactly what the ROC curve formalizes: sweeping it trades sensitivity against specificity, and the few-Hounsfield-unit or few-millimetre variability you can induce here by hand is the same reader variability that the repeatability coefficient and the QUADAS-2 index-test domain are designed to expose.
02Prognostic Studies
Prognostic research asks a fundamentally different question from diagnostic research: not "is the disease present?" but "given that it is present (or that a finding exists), what will happen over time?" The mathematical machinery is therefore that of time-to-event (survival) analysis rather than the static table. The central objects are the survival function , the probability of remaining event-free beyond time , and the hazard , the instantaneous event rate among those still at risk. Because patients enter and leave observation at different times, is estimated nonparametrically by the Kaplan–Meier product-limit estimator, , where events occur among at risk at each event time . Censoring — the loss of follow-up before the event — is the defining feature of these data, and the validity of the Kaplan–Meier estimate requires that censoring be non-informative (uncorrelated with prognosis); informative censoring, common when sicker patients drop out, biases survival estimates optimistically.
The effect of an imaging biomarker on outcome is most often expressed as a hazard ratio (HR) from the Cox proportional-hazards model, , in which the HR is assumed constant over time — the proportional-hazards assumption that must be checked, not presumed. An HR of 2.0 for, say, a sarcopenia threshold on staging CT means the instantaneous death rate is doubled in the high-risk group at every moment, which is not the same as a doubling of the probability of death and must not be read as a risk ratio. Two further properties of a prognostic marker are conceptually distinct and frequently conflated. Discrimination — the ability to rank patients by risk — is summarized by Harrell's concordance (C) statistic, the time-to-event generalization of the AUC, equal to the probability that, for a randomly chosen comparable pair, the patient who fails first carried the higher predicted risk. Calibration — the agreement between predicted and observed absolute risk — is an entirely separate virtue: a model can discriminate well yet be systematically miscalibrated, over- or under-stating absolute event probabilities, which is precisely the property that matters when a predicted risk drives a treatment threshold.
Prognostic studies are vulnerable to characteristic biases that the appraisal tools were built to surface. A prognostic-factor study can be corrupted by unrepresentative sampling, inadequate adjustment for confounders, attrition, and selective outcome reporting; the QUIPS tool (Hayden) structures this appraisal across six domains (study participation, study attrition, prognostic factor measurement, outcome measurement, study confounding, and statistical analysis and reporting). A multivariable prediction model carries the additional, insidious hazard of optimism — the tendency of a model to fit the noise of its own derivation sample, so that apparent performance overstates true performance — which is why internal validation by bootstrapping or cross-validation, and ultimately external (geographic and temporal) validation, are non-negotiable. The reporting of such models is standardized by TRIPOD (Collins), and their risk of bias and applicability are appraised by PROBAST (Wolff), which flags the recurring methodologic sins of small events-per-variable ratios, dichotomization of continuous predictors, and validation on the development data alone. For the imaging physician, the practical synthesis is that a CT-derived prognostic marker (an infarct volume, a tumour-burden estimate, a body-composition metric) is credible as a prognosticator only when it has been shown to add discrimination beyond established clinical variables, to be calibrated in absolute terms, and to retain performance in a population separate from the one in which it was discovered.
03Biomarker Validation
A quantitative imaging biomarker — an infarct or hematoma volume, an emphysema index, a hepatic fat fraction, a tumour-burden measurement, a CT-derived bone mineral density — is a measurement instrument, and like any instrument it must be validated along a defined hierarchy before its outputs can bear clinical or regulatory weight. The O'Connor imaging-biomarker roadmap (O'Connor et al., 2017) organizes this into two parallel tracks that must both succeed: a technical (analytical) validation track establishing that the measurement is accurate, precise, and reproducible, and a biological/clinical validation track establishing that the measurement reflects the biology it claims to and predicts a clinical outcome or response. The cardinal error in imaging research is to leap from a measurement that correlates with disease to the assertion that it is a validated biomarker, skipping the metrology that determines whether an observed change is real signal or measurement noise.
Technical performance is the domain of formal metrology, codified for imaging by the QIBA program (Sullivan et al., 2015; Raunig et al., 2014). Repeatability (same scanner, same patient, same conditions, short interval) and reproducibility (varying scanner, operator, site, or software) are quantified, not asserted. The key statistic for a continuous biomarker is the within-subject coefficient of variation and the derived repeatability coefficient (with the within-subject standard deviation), which defines the threshold below which a difference between two measurements on the same subject is statistically indistinguishable from measurement error at the 95% level. The factor arises because the variance of a difference of two independent measurements is twice the measurement variance. The operational consequence is decisive: if a tumour-diameter biomarker has mm, then a measured increase of 3 mm — however confidently reported — is within the noise and cannot license a claim of progression. Agreement between two methods or readers is assessed not by correlation (which measures association, not agreement) but by the Bland–Altman analysis, which plots the difference of paired measurements against their mean and reports the bias and the 95% limits of agreement . Reliability across raters or scans is summarized by the intraclass correlation coefficient (ICC), , the proportion of total variance attributable to true between-subject differences rather than measurement error; an ICC approaching 1 indicates that the instrument resolves real differences between patients rather than reshuffling noise.
Only once a biomarker is shown to be technically sound does biological and clinical validation become meaningful. Biological validation links the measurement to the underlying pathology — for example, demonstrating that a CT densitometric emphysema index corresponds to histologic airspace destruction, or that a perfusion parameter tracks angiogenesis. Clinical validation establishes that the biomarker predicts a clinically meaningful endpoint, and the apex of the hierarchy, qualification as a surrogate endpoint, requires the demanding evidentiary standard that a treatment's effect on the biomarker reliably predicts its effect on the true outcome — a standard most candidate imaging surrogates never meet, because a treatment can move the image without moving survival. The QIBA framework operationalizes this through a Profile that specifies the acquisition, analysis, and performance claims under which a biomarker achieves a stated precision, making cross-site quantitative imaging possible. The recurring failure modes are quantitatively specific: reporting correlation in place of agreement, ignoring the repeatability coefficient so that noise is mistaken for change, conflating a statistically significant association with clinical utility, and — most consequentially — treating an unqualified intermediate measurement as a surrogate endpoint and thereby licensing therapeutic decisions on a biomarker that has never been shown to predict the outcome that matters.
04Clinical Trial Interpretation
The clinical trial is the instrument by which imaging strategies — not merely imaging tests — are shown to change patient outcomes, and its interpretation demands a vocabulary distinct from that of accuracy or prognosis. A diagnostic-accuracy study can establish that CT pulmonary angiography detects emboli with high sensitivity, but only a trial can establish that a CT-directed management strategy reduces mortality or recurrent thromboembolism relative to an alternative. The architectural pillar of the randomized controlled trial is randomization, which on average balances both measured and unmeasured confounders across arms so that the difference in outcome can be attributed to the intervention rather than to prognostic differences between the groups. Equally fundamental is the intention-to-treat (ITT) principle: patients are analyzed in the arm to which they were randomized regardless of crossover, protocol deviation, or non-adherence. ITT preserves the prognostic balance that randomization created and yields an unbiased estimate of the effect of the strategy as offered; the seductive alternative of per-protocol or as-treated analysis reintroduces selection bias, because adherence is itself prognostic. ITT is typically conservative for superiority trials and, importantly, anti-conservative for non-inferiority trials, where it can mask a true difference — a subtlety that matters when a lower-dose or lower-cost imaging protocol is tested for non-inferiority.
The effect of the intervention must then be read in the correct metric, and the distinction between relative and absolute measures is where clinical judgment most often fails. A trial may report a relative risk or relative risk reduction , but the clinically actionable quantity is the absolute risk reduction and its reciprocal, the number needed to treat . A 50% relative risk reduction is dramatic when the control risk is 20% (, ) and nearly meaningless when the control risk is 0.2% (, ); the relative figure is identical in both, which is precisely why relative measures dominate abstracts and absolute measures should dominate decisions. Every estimate must be reported with its 95% confidence interval, which is best understood not as a probability statement about the true value but as the range of parameter values most compatible with the observed data given the model — equivalently, the interval generated by a procedure that would contain the true parameter in 95% of identically conducted studies. A confidence interval that excludes the null and is narrow indicates a precise, statistically significant effect; one that crosses the null is compatible with no effect, and one that is wide signals that the trial, however "positive," has estimated the effect imprecisely. The -value is the probability of data at least as extreme as observed if the null were true, and it speaks to neither the size nor the importance of an effect.
Trials of imaging carry threats beyond those of drug trials. Blinding is harder when the intervention is a scan, raising the risk of differential workup and ascertainment between arms; outcome-verification bias arises when the imaging result drives the very investigations used to confirm the outcome, inflating apparent benefit; and paradigm creep — the drift of background management between the trial era and the present — can render a landmark imaging trial's effect estimate obsolete even when its internal validity is impeccable. The competing-risks structure of many imaging populations (elderly patients dying of unrelated causes before the studied event) further complicates the reading of time-to-event endpoints. The physician interpreting an imaging trial therefore asks not only whether the result is statistically significant but whether it was analyzed by intention-to-treat, whether the effect is large in absolute terms, how precisely it was estimated, and whether the trial's diagnostic and therapeutic milieu still resembles the present one.
05Meta-Analysis
Meta-analysis is the quantitative synthesis of multiple studies addressing the same question, and it is indispensable in imaging because individual diagnostic-accuracy and prognostic studies are frequently small, single-center, and underpowered. Its logic is to pool study-level estimates with weights that reflect their precision, but the choice of weighting model encodes a substantive assumption. A fixed-effect model assumes every study estimates a single common true effect and weights each study by the inverse of its within-study variance, . A random-effects model assumes the true effect varies across studies — by population, scanner generation, threshold, or reference standard — and adds a between-study variance component so that , widening the pooled confidence interval to reflect that heterogeneity. Because imaging studies differ profoundly in spectrum, technology, and reader expertise, the random-effects assumption is almost always the more honest one, and a pooled estimate presented without acknowledgment of heterogeneity should be regarded with suspicion.
Heterogeneity is not a nuisance to be averaged away but a finding to be quantified and explained. Cochran's tests the null of homogeneity, but its low power in small meta-analyses makes the statistic — (truncated at zero when negative), the percentage of total variation across studies attributable to true heterogeneity rather than chance — the more useful descriptor, with values around 25%, 50%, and 75% conventionally denoting low, moderate, and high heterogeneity. When heterogeneity is substantial, the appropriate response is not a single pooled number but meta-regression or subgroup analysis to identify the study-level covariates (e.g., scanner type, prevalence, threshold) that explain it. Diagnostic-accuracy meta-analysis poses a special technical problem: sensitivity and specificity are correlated and jointly determined by an implicit threshold that varies across studies, so they cannot be pooled independently without bias. The correct approach is the bivariate random-effects model or the equivalent hierarchical summary ROC (HSROC) model, which jointly model the sensitivity–specificity pair and the threshold variation across studies, yielding a summary ROC curve and an operating point with a proper joint confidence region rather than two spuriously precise univariate pooled values.
The credibility of any meta-analysis is bounded by the completeness and quality of its constituent studies, which is why a rigorous synthesis is reported according to the PRISMA 2020 statement (Page et al.) — documenting the search, selection, and data-extraction process so the synthesis is reproducible — and appraises each included study with the relevant risk-of-bias tool (QUADAS-2 for accuracy, QUIPS/PROBAST for prognosis). The most insidious threat is publication bias and small-study effects: studies with impressive or "positive" results are more likely to be published, indexed, and in English, so the visible literature is a biased sample of the work actually done, systematically inflating pooled estimates. This is interrogated graphically with the funnel plot — a scatter of effect size against precision that should be symmetric if no such bias exists — and statistically with tests for asymmetry (Egger's regression and, for diagnostic data, the Deeks test), though these have limited power. A meta-analysis is therefore strongest when its search was exhaustive and pre-registered, its included studies were individually sound, its heterogeneity was quantified and explained rather than buried, and its synthesis model matched the structure of the data; a pooled sensitivity computed by naively averaging across heterogeneous, biased, selectively published studies is a precise-looking number with little inferential value.
06Guideline Development
A clinical practice guideline is the formal bridge from a body of evidence to a recommendation for action, and its development is itself a methodologic discipline rather than an expression of expert opinion. The dominant framework is GRADE (Grading of Recommendations Assessment, Development and Evaluation; Guyatt et al.), which makes two contributions that the imaging physician must understand. First, it rates the certainty of evidence for each outcome — high, moderate, low, or very low — as a property of the body of evidence rather than of any single study. Randomized trials begin at high certainty and observational studies at low, and the rating is then revised downward for risk of bias, inconsistency (unexplained heterogeneity), indirectness (the evidence concerns a different population, comparator, or outcome than the question — endemic in imaging, where studies measure diagnostic accuracy as a surrogate for the patient outcomes that actually matter), imprecision (wide confidence intervals), and publication bias; observational evidence may be rated upward for a large effect, a dose–response gradient, or when plausible residual confounding would only diminish an observed effect. This structured downgrading prevents the common error of treating a statistically significant result from a biased or indirect study as strong evidence.
Second, GRADE deliberately separates the certainty of evidence from the strength of the recommendation. A recommendation can be strong or conditional (weak), and its strength depends not only on the certainty of evidence but on the balance of benefits and harms, patient values and preferences, resource use, and feasibility. This decoupling is consequential in imaging: a CT-based screening or surveillance strategy may rest on moderate- or high-certainty evidence of improved diagnostic accuracy yet warrant only a conditional recommendation because the downstream harms — radiation exposure, the cascade of incidental findings, overdiagnosis of indolent disease, false-positive workups, and cost — offset the benefit, or because the benefit is small in absolute terms. The recurring methodologic hazard in imaging guidelines is precisely indirectness from the diagnostic-to-clinical gap: the literature is rich in accuracy data and poor in outcome data, so guideline panels must repeatedly judge whether better detection plausibly translates into better health, a judgment GRADE forces them to make explicit rather than to assume.
The procedural integrity of the guideline matters as much as its analytic logic. A trustworthy guideline (per the standards articulated by the Institute of Medicine and operationalized in instruments such as AGREE II) is built by a multidisciplinary panel with explicit management of intellectual and financial conflicts of interest, rests on a systematic review conducted to PRISMA standards rather than a selective reading of the literature, makes the link from evidence to each recommendation transparent through an evidence-to-decision framework, articulates the values and assumptions underlying its judgments, and specifies a plan for updating as new evidence emerges. The mature reader of a guideline therefore interrogates not merely the recommendation but its provenance: whether the evidence was systematically assembled and certainty-rated, whether the recommendation's strength is congruent with that certainty and with an honest accounting of harms, whether the panel's composition and conflicts were managed, and whether the recommendation rests on outcome evidence or on an explicit, defended inferential leap from diagnostic accuracy. In this sense guideline appraisal recapitulates the entire chapter — diagnostic accuracy, prognosis, biomarker validity, trial evidence, and synthesis converge here into the single question of what a clinician should actually do, and the discipline of evidence-based imaging is the refusal to answer that question by authority when it can be answered, however imperfectly, by appraised evidence.
✅ Check your understanding
10 questions- 1.
A CT sign for a given diagnosis has a sensitivity of 0.90 and a specificity of 0.90. In an emergency-department population where the pre-test prevalence of the disease is 10%, what is the approximate positive predictive value of a positive finding, and what general principle does this illustrate?
med - 2.
A CT finding has a positive likelihood ratio of 9. A patient's pre-test probability of disease is estimated at 20%. Using Bayes' theorem in odds form, what is the approximate post-test probability after a positive finding?
hard - 3.
A diagnostic-accuracy study for a novel CT sign enrolls patients with advanced, biopsy-proven disease as cases and healthy young volunteers as controls, and reports a sensitivity and specificity both exceeding 0.97. Which bias most likely inflates this accuracy, and in which direction does it act?
med - 4.
The area under the ROC curve (AUC) for a quantitative CT biomarker is reported as 0.82. What is the correct probabilistic interpretation of this value?
med - 5.
A CT-derived tumour-diameter biomarker has a repeatability coefficient (RC) of 4 mm. On follow-up the same lesion, measured under identical conditions, has increased by 3 mm. What is the correct interpretation?
hard - 6.
When two methods of measuring hepatic fat fraction on CT are compared, an investigator reports a Pearson correlation of r = 0.95 and concludes the methods 'agree.' Why is this conclusion methodologically flawed, and what is the correct analysis?
med - 7.
A randomized trial compares a low-dose CT protocol against standard-dose CT for a management strategy, powered as a non-inferiority trial. A substantial fraction of patients randomized to low-dose crossed over to standard-dose. Why is intention-to-treat analysis problematic here, and what is the correct stance?
hard - 8.
A meta-analysis of a CT diagnostic sign pools the reported sensitivities and the reported specificities separately using univariate random-effects models. Why is this approach biased, and what is the methodologically correct alternative?
hard - 9.
Within the GRADE framework, a body of evidence consists exclusively of well-conducted randomized trials of a CT screening strategy that demonstrate improved diagnostic accuracy, but no trial measured patient-important outcomes (mortality or morbidity). For the outcome 'reduced disease-specific mortality,' which GRADE domain most directly mandates rating down the certainty of evidence?
hard - 10.
A trial reports that a CT-directed strategy reduces the relative risk of an adverse outcome by 40%. The event rate in the control arm is 0.5%. What absolute benefit and number needed to treat does this imply, and what is the interpretive lesson?
med
🌐 Keep exploring — Radiopaedia & more
Hand-picked, free external references to deepen this topic.
References & primary literature
- 1.Bossuyt PM, Reitsma JB, Bruns DE, et al. STARD 2015: an updated list of essential items for reporting diagnostic accuracy studies. BMJ. 2015;351:h5527.
- 2.Whiting PF, Rutjes AWS, Westwood ME, et al. QUADAS-2: a revised tool for the quality assessment of diagnostic accuracy studies. Ann Intern Med. 2011;155(8):529-536.
- 3.Ransohoff DF, Feinstein AR. Problems of spectrum and bias in evaluating the efficacy of diagnostic tests. N Engl J Med. 1978;299(17):926-930.
- 4.Collins GS, Reitsma JB, Altman DG, Moons KGM. Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis (TRIPOD): the TRIPOD statement. Ann Intern Med. 2015;162(1):55-63.
- 5.Wolff RF, Moons KGM, Riley RD, et al. PROBAST: a tool to assess the risk of bias and applicability of prediction model studies. Ann Intern Med. 2019;170(1):51-58.
- 6.Hayden JA, van der Windt DA, Cartwright JL, Côté P, Bombardier C. Assessing bias in studies of prognostic factors. Ann Intern Med. 2013;158(4):280-286.
- 7.O'Connor JPB, Aboagye EO, Adams JE, et al. Imaging biomarker roadmap for cancer studies. Nat Rev Clin Oncol. 2017;14(3):169-186.
- 8.Sullivan DC, Obuchowski NA, Kessler LG, et al. Metrology standards for quantitative imaging biomarkers. Radiology. 2015;277(3):813-825.
- 9.Raunig DL, McShane LM, Pennello G, et al. Quantitative imaging biomarkers: a review of statistical methods for technical performance assessment. Stat Methods Med Res. 2015;24(1):27-67.
- 10.Guyatt GH, Oxman AD, Vist GE, et al. GRADE guidelines: 1. Introduction—GRADE evidence profiles and summary of findings tables. J Clin Epidemiol. 2011;64(4):383-394.
- 11.Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71.
Tip: use ← / → to move between chapters.