Methods Critique
Everything on Aqrab tagged Methods Critique — grouped into one landing page so readers can go deeper by problem family instead of bouncing around the archive blind.
Causal Readiness: When a Huge Linked Dataset Still Cannot Identify an Effect
A practical guide to causal readiness in linked health and administrative data. Learn why scale and propensity-score overlap are not enough when treatment, need, comparators, or outcomes are poorly measured.
When More Covariates Break Positivity: Representation-Induced Overlap Failure in Clinical Text
A practical guide to representation-induced positivity failure in clinical text. Learn why richer embeddings can encode treatment, shrink common support, and make a causal adjustment less trustworthy.
Healthy Screenee Bias: When Screening Attendance Looks Like Screening Benefit
A practical guide to healthy screenee bias for clinical researchers. Learn why people who attend screening can look healthier before the screen is credited, how this differs from lead-time bias, and what reviewers should demand from observational screening studies.
Predicted Treatment Benefit: When a Risk Model Is Not a Treatment Recommendation
A practical guide to separating predicted outcome risk from predicted treatment benefit. Learn why a high-risk patient is not automatically a high-benefit patient, how risk modeling and effect modeling differ, and what reviewers should demand before trusting a personalized treatment claim.
Additive Interaction: When “No Interaction” Depends on the Scale
A practical guide to additive and multiplicative interaction in clinical research. Learn why a null product term can hide clinically important effect modification, how to read RERI, and what reviewers should demand before trusting a joint-exposure claim.
Small Number of Clusters: When 500 Patients Still Behave Like 10 Sites
A practical guide to inference with a small number of clusters in clinical research. Learn why patients inside the same site are not independent evidence, why default cluster-robust standard errors can be too optimistic, and what reviewers should demand before trusting a cluster-level result.
Heterogeneous Treatment Effects: Validate the Average Effect Before Trusting the Subgroups
A practical guide to validating heterogeneous treatment-effect claims in target trial emulation. Learn why a failed average-effect benchmark should stop a subgroup story, how subgroup effects must reconcile with their parent estimate, and what reviewers should demand before trusting personalized treatment claims.
Estimation Over Testing: Four Habits That Quietly Break Clinical Statistics
A practical critique of four entrenched habits in clinical biostatistics — significance testing where the real question is estimation, prediction accuracy scored far from the bedside, meta-analysis on autopilot, and the false precision of point estimates that assume away what cannot be tested — with an interactive explorer that reads one trial three ways and what to do instead.
Covariate Balance Diagnostics: When the Love Plot Says Balanced but the Groups Still Differ
A practical guide to covariate balance diagnostics after propensity-score matching or weighting. Covers why a Love plot of standardized mean differences can read balanced while distributions differ in spread and tails, why you should check variance ratios and distributional distances and interactions, why you should never significance-test balance because the p-value tracks sample size, and what reviewers should demand.
The Will Rogers Phenomenon: When Better Staging Improves Every Group and Nobody Lives Longer
A practical guide to the Will Rogers phenomenon (stage migration) for clinical researchers. Covers why sharper classification can lift survival in every stage while overall survival stays flat, how to separate stage migration from real progress, how it differs from lead-time bias and overdiagnosis, and what reviewers should demand when outcomes are compared across eras or cohorts.
The Imperfect Gold Standard: Measuring a Test When the Reference Is Wrong Too
A practical guide to imperfect reference standards and latent class analysis for clinical researchers. Covers why apparent sensitivity and specificity are biased when the gold standard is itself flawed, why conditional dependence between tests flips the bias from pessimistic to optimistic, when latent class analysis helps and when it fails, and what reviewers should demand.
Incorporation Bias: When a Test Helps Write Its Own Answer Key
A practical guide to incorporation bias for clinical researchers. Covers why sensitivity and specificity both inflate toward 100% when the index test is used to define the reference standard, how it differs from verification and spectrum bias, why there is no clean statistical correction, and what reviewers should demand before trusting a diagnostic accuracy.
Verification Bias: When the Test Under Study Decides Who Gets the Gold Standard
A practical guide to verification bias (workup bias) for clinical researchers. Covers why sensitivity is inflated and specificity deflated when the index test drives who gets the reference standard, the Begg-Greenes correction, differential verification, and what reviewers should demand before trusting a diagnostic accuracy.
Spectrum Bias: Why a Test’s Accuracy Is Not a Property of the Test
A practical guide to spectrum bias for clinical researchers. Covers why sensitivity and specificity shift with the case-mix of who was enrolled, the two-gate case-control trap, how curated data inflates AI-diagnostic performance, and what reviewers should demand before trusting a reported accuracy.
Simpson's Paradox: When Every Subgroup Says One Thing and the Total Says the Opposite
A practical guide to Simpson's paradox for clinical researchers. Covers why a treatment can help in every subgroup yet look harmful pooled, why 'always stratify' is wrong, and how the causal role of the stratifier — confounder, mediator, or collider — decides which table to trust.
The Table 2 Fallacy: When Every Adjusted Coefficient Looks Like a Cause
A practical guide to the Table 2 fallacy for clinical researchers. Covers why secondary coefficients in an adjusted model are not causal effects, mutual adjustment, mediator and confounder mismatch, and what reviewers should demand before reading a covariate row as a finding.
Bayesian Borrowing: When Historical Data Starts Spending Credibility It Did Not Earn
A practical guide to Bayesian borrowing for clinical researchers. Covers exchangeability, commensurate priors, historical controls, calendar-time drift, and what reviewers should demand before trusting extra certainty borrowed from earlier data.
PROBAST: When a Prediction Model Paper Looks Ready Before It Earns Trust
A practical guide to PROBAST for clinical researchers. Covers participant selection, predictor leakage, outcome definition, overfitting, calibration, and what reviewers should demand before trusting a clinical prediction model.
Informative Cluster Size: When the Biggest Sites Start Writing the Result
A practical guide to informative cluster size for clinical researchers. Covers why larger centers can quietly dominate treatment effects, how weighting changes the estimand, and what reviewers should demand before trusting clustered results.
Baseline Adjustment in Randomized Trials: Why Change From Baseline Keeps Losing to ANCOVA
A practical guide to baseline adjustment in randomized trials. Covers ANCOVA versus change scores, percent change traps, responder thresholds, and what reviewers should demand before trusting a tidy efficacy claim.
Depletion of Susceptibles: When Early Harm Vanishes Because the Vulnerable Patients Are Already Gone
A practical guide to depletion of susceptibles for clinical researchers. Covers front-loaded harm, survivor selection, why later follow-up can falsely reassure, and what reviewers should demand before trusting a calming hazard curve.
Indirectness in Clinical Evidence: When a Good Study Answers the Wrong Question
A practical guide to indirectness in clinical evidence for clinical researchers. Covers PICO mismatch, outdated comparators, surrogate outcomes, and what reviewers should demand before trusting an applicable-sounding conclusion.
Complete-Case Analysis: When Missing Data Quietly Changes the Study Population
A practical guide to complete-case analysis for clinical researchers. Covers when dropping incomplete records changes the study population, how endpoint missingness becomes selection bias, and what reviewers should demand before trusting the estimate.
Platform Trials: When a Shared Control Stops Being the Same Comparison
A practical guide to platform trials for clinical researchers. Covers nonconcurrent controls, changing standard of care, case-mix drift, and what reviewers should demand before trusting an adaptive-trial headline.
Nonproportional Hazards: When One Hazard Ratio Pretends the Treatment Effect Never Changes
A practical guide to nonproportional hazards for clinical researchers. Covers delayed effects, crossing curves, waning benefit, why one hazard ratio can mislead, and what reviewers should demand instead.
Futility Stopping in Clinical Trials: When “No Signal Yet” Starts Pretending the Question Is Answered
A practical guide to futility stopping in clinical trials. Covers conditional power, delayed effects, optimistic design assumptions, and what reviewers should demand before trusting a trial stopped for futility.
Nested Case-Control Design: When Cheap Control Sampling Still Has to Respect Event Time
A practical guide to nested case-control design for clinical researchers. Covers incidence-density sampling, risk-set control selection, how it differs from case-cohort design, and what reviewers should demand before trusting the result.
Transitivity in Network Meta-Analysis: When Indirect Comparisons Pretend the Trials Were Exchangeable
A practical guide to transitivity in network meta-analysis for clinical researchers. Covers effect modifiers, shared comparators, indirect comparison failure modes, and what reviewers should demand before trusting rankings.
Benchmarking Target Trial Emulation: When a Trial Copy Never Checks It Can Reproduce the Known Answer
A practical guide to benchmarking target trial emulation for clinical researchers. Covers why benchmark replication matters, what it cannot prove, and what reviewers should demand before trusting an observational extension.
Weak Instruments and Physician Preference IVs: When Treatment Movement Is Not Yet Causal Credibility
A practical guide to weak instruments and physician-preference IVs for clinical researchers. Covers first-stage weakness, exclusion leakage, local interpretation, and what reviewers should demand before trusting an IV claim.
Noninferiority Margins: When “Not Much Worse” Starts Giving Away Too Much
A practical guide to noninferiority margins for clinical researchers. Covers margin justification, assay sensitivity, constancy, biocreep, and what reviewers should demand before trusting a noninferiority win.
Repeated Eligibility in Target Trial Emulation: When One Patient Quietly Enters the Trial Again
A practical guide to repeated eligibility in target trial emulation for clinical researchers. Covers re-entry rules, overlapping follow-up, carryover, and what reviewers should demand before trusting a sequential-trial analysis.
ROBINS-I: When an Observational Effect Estimate Is Too Biased to Grade Casually
A practical guide to ROBINS-I for clinical researchers. Covers the seven bias domains, why confounding and time zero usually dominate, how overall judgments are formed, and what reviewers should demand before trusting non-randomized evidence.
M-Bias: When “Adjusted for Severity” Is Actually the Problem
A practical guide to M-bias for clinical researchers. Covers collider structures, severity-score traps, selected cohorts, and what reviewers should demand before trusting an adjusted estimate.
Stepped-Wedge Cluster Trials: When Rollout Timing Starts Competing With the Intervention
A practical guide to stepped-wedge cluster trials for clinical researchers. Covers secular trends, rollout order, learning effects, contamination, and what reviewers should demand before trusting a tidy implementation-era benefit.
Calibration Drift: When a Good Model Keeps the Right Rank and Still Gives the Wrong Risk
A practical guide to calibration drift for clinical researchers. Covers baseline-risk shift, calibration slope failure, threshold consequences, and what reviewers should demand before trusting deployment-ready prediction claims.
Response-Adaptive Randomization: When a Trial Starts Chasing Its Early Winners
A practical guide to response-adaptive randomization for clinical researchers. Covers delayed outcomes, temporal drift, instability, ethical claims, and what reviewers should demand before trusting an adaptive allocation design.
Case-Cohort Design: When Measuring Everyone Is the Wrong Expense
A practical guide to case-cohort design for clinical researchers. Covers when a random subcohort is more honest than measuring everyone, how it differs from nested case-control sampling, and what reviewers should demand before trusting the result.
Augmented Inverse Probability Weighting: When “Doubly Robust” Starts Hiding Which Model Failed
A practical guide to augmented inverse probability weighting for clinical researchers. Covers what AIPW estimates, why doubly robust does not mean low-risk, and what reviewers should demand before trusting the label.
Case-Time-Control Design: When Case-Crossover Starts Confusing Time Trends with Treatment Effects
A practical guide to the case-time-control design for clinical researchers. Covers exposure-time trends, referent sampling, protopathic bias, and what reviewers should demand before trusting a self-matched trigger analysis.
Triangulation in Clinical Research: When One Elegant Design Still Leaves the Same Blind Spot
A practical guide to triangulation in clinical research. Covers what counts as genuinely complementary evidence, how to map designs to specific threats, and what reviewers should demand before trusting “robustness” claims.
Guideline Recommendation Strength: When “Strongly Recommend” Starts Outrunning the Evidence
A practical guide to recommendation strength in clinical guidelines. Covers certainty of evidence, benefit-harm tradeoffs, patient values, implementation burden, and what reviewers should demand before trusting a forceful recommendation.
Differential Misclassification: When One Study Arm Gets More Chances to Be Wrong
A practical guide to differential misclassification for clinical researchers. Covers arm-specific outcome detection, adjudication asymmetry, false positives, missed events, and what reviewers should demand before trusting an effect estimate.
Adaptive Enrichment Trials: When Precision for One Subgroup Pretends to Be Evidence for Everyone
A practical guide to adaptive enrichment trials for clinical researchers. Covers predictive versus prognostic enrichment, assay timing, multiplicity, external validity, and what reviewers should demand before trusting a biomarker-selected win.
Treatment-Induced Mediator-Outcome Confounding: When Mediation Analysis Starts Chasing the Consequences of Treatment
A practical guide to treatment-induced mediator-outcome confounding for clinical researchers. Covers why natural direct and indirect effects fail when treatment changes later severity, toxicity, adherence, or surveillance that affect both the mediator and outcome.
Surrogate Endpoints: When a Biomarker Improvement Pretends to Be Patient Benefit
A practical guide to surrogate endpoints for clinical researchers. Covers validated versus merely plausible surrogates, classic failure modes, and what reviewers should demand before trusting a biomarker-driven trial claim.
Data Leakage in Clinical Prediction Models: When the Model Learns the Future
A practical guide to data leakage in clinical prediction models for clinical researchers. Covers post-outcome features, workflow proxies, validation traps, and what reviewers should demand before trusting a headline AUC.
Net Reclassification Improvement: When a New Biomarker Wins by Moving Patients Between the Wrong Boxes
A practical guide to net reclassification improvement for clinical researchers. Covers event and non-event NRI, arbitrary risk categories, overtreatment traps, and what reviewers should demand before trusting claims that a new model improved classification.
AI-Assisted Methods Review: What LLMs Can Catch, What They Cannot, and Where Judgment Still Matters
A practical guide to AI-assisted methods review for clinical researchers. Covers where LLMs help with structural critique, where source verification and causal judgment still require humans, and what reviewers should demand before trusting AI-generated methodological comments.
Decision Curve Analysis: When a Better AUC Still Makes Worse Clinical Decisions
A practical guide to decision curve analysis for clinical researchers. Covers net benefit, threshold probability, when prediction models fail to beat treat-all or treat-none strategies, and what reviewers should demand before trusting claims of clinical utility.
Channeling Bias: When the Newer Treatment Inherits the Easier Patients
A practical guide to channeling bias for clinical researchers. Covers preferential prescribing, formulary-era drift, specialist selection, and what reviewers should demand before trusting observational comparisons of newer therapies.
When Death Changes the Question: Competing Risks, Intercurrent Events, and Truncation by Death
A practical guide to the boundary between competing risks, intercurrent events, and truncation by death for clinical researchers. Covers when death changes risk sets, when it makes later outcomes undefined, and what reviewers should demand instead of vague censoring language.
Jump-to-Reference Imputation: When Missing Outcomes Start Borrowing the Control Arm's Future
A practical guide to jump-to-reference imputation for clinical researchers. Covers what J2R assumes after treatment discontinuation, when it helps sensitivity analysis, and when it quietly answers the wrong estimand.
Multiple Testing in Clinical Trials: When One Positive Endpoint Is Just the Loudest Coin Flip
A practical guide to multiple testing in clinical trials for clinical researchers. Covers endpoint families, subgroup fishing, interim looks, alpha control, and what reviewers should demand before trusting a lone positive result.
Confounding by Contraindication: When the Untreated Group Is Too Fragile for the Therapy
A practical guide to confounding by contraindication for clinical researchers. Covers how treatment avoidance in high-risk patients can make therapies look safer or more effective than they are, and what reviewers should demand instead.
Last Observation Carried Forward: When Yesterday's Outcome Pretends the Patient Stopped Changing
A practical guide to last observation carried forward for clinical researchers. Covers why LOCF fails as missing-data strategy, how it can exaggerate or dilute treatment effects, and what reviewers should demand instead.
Time Zero Alignment: When Your Cohort Starts Counting Before Treatment Does
A practical guide to time zero alignment for clinical researchers. Covers eligibility, treatment assignment, delayed initiation, immortal time, and what reviewers should demand before trusting a real-world effect estimate.
Early Stopping for Benefit: When a Trial Quits While the Effect Is Still on Its Best Behavior
A practical guide to early stopping for benefit in clinical trials. Covers interim looks, alpha spending, exaggerated effect sizes, immature follow-up, and what reviewers should demand before trusting a triumphant stop.
Informative Visit Processes: When Who Shows Up Starts Writing the Results
A practical guide to informative visit processes for clinical researchers. Covers endogenous follow-up, unequal observation schedules, visit-triggered outcome capture, inverse-intensity thinking, and what reviewers should demand before trusting longitudinal real-world results.
External Control Arms: When a Comparison Group Arrives from Another Universe
A practical guide to external control arms for clinical researchers. Covers historical and real-world comparators, design drift, prognostic imbalance, endpoint mismatch, and what reviewers should demand before trusting single-arm success stories.
Missing Indicator Method: When an NA Flag Pretends to Be Missing-Data Strategy
A practical guide to the missing-indicator method for clinical researchers. Covers why NA flags fail for confounding control, when they leave residual bias, and what reviewers should demand before trusting a covariate-adjusted result.
Regression to the Mean: When Extreme Patients Improve Before Your Treatment Deserves Credit
A practical guide to regression to the mean for clinical researchers. Covers extreme-baseline selection, before-after mirages, symptom flares, biomarker spikes, and what reviewers should demand before trusting dramatic improvement.
Treatment Switching in Oncology Trials: When Overall Survival Becomes a Rescue Protocol Audit
A practical guide to treatment switching in oncology trials for clinical researchers. Covers crossover, overall survival dilution, ITT versus hypothetical estimands, RPSFTM, IPCW, two-stage estimation, and what reviewers should demand before trusting an adjusted survival claim.
Run-In Periods: When Your Trial Randomizes the Easy Patients First
A practical guide to run-in periods for clinical researchers. Covers adherence enrichment, tolerability selection, estimand drift, external validity, and what reviewers should demand before trusting a polished randomized cohort.
Washout Periods: When “New Use” Is Just Old Use with Better PR
A practical guide to washout periods for clinical researchers. Covers new-user definitions, refill cycles, intermittent treatment, data-history limits, and what reviewers should demand before trusting an incident-user cohort.
Exposure Lagging: When Your Induction Window Becomes Wishful Thinking
A practical guide to exposure lagging for clinical researchers. Covers induction periods, reverse causation, protopathic bias, estimand drift, and what reviewers should demand before trusting a lagged analysis.
Responder Analyses: When a Cutoff Turns a Clinical Gradient into a Headline
A practical guide to responder analyses for clinical researchers. Covers dichotomizing continuous outcomes, post hoc thresholds, baseline dependence, power loss, and what reviewers should demand before trusting "X% achieved response" claims.
Healthy Adherer Bias: When Persistence Looks Like Pharmacology
A practical guide to healthy adherer bias for clinical researchers. Covers why adherent patients often look healthier before the treatment effect is even estimated, how this differs from confounding by indication, and what reviewers should demand before trusting adherence-based benefit claims.
Grace Periods in Target Trial Emulation: Clinical Realism or Future Information in Disguise?
A practical guide to grace periods in target trial emulation for clinical researchers. Covers when a grace window is defensible, when it becomes immortal time in formalwear, and what reviewers should demand before trusting the result.
Index Event Bias: When Your Cohort Already Selected the Wrong Comparison
A practical guide to index event bias for clinical researchers. Covers recurrence-risk paradoxes, conditioning on the first event, secondary prevention cohorts, and what reviewers should demand before trusting protective-looking associations inside diseased cohorts.
Calendar Time Confounding: When Secular Trends Pretend Your Intervention Worked
A practical guide to calendar time confounding for clinical researchers. Covers secular trends, treatment diffusion, concurrent comparators, and what reviewers should demand before trusting real-world benefit that may just reflect a later era.
Surveillance Bias: When One Group Gets More Chances to Become a Case
A practical guide to surveillance bias for clinical researchers. Covers differential testing, follow-up intensity, diagnosis-based outcomes, and what reviewers should demand before trusting higher event rates.
Outcome Switching: When the Primary Endpoint Moves After the Results Get Interesting
A practical guide to outcome switching for clinical researchers. Covers endpoint shopping, selective reporting, protocol drift, and what reviewers should demand before trusting a late-breaking primary outcome.
Overdiagnosis: When Finding More Disease Does Not Mean Saving More Lives
A practical guide to overdiagnosis for clinical researchers. Covers how screening can raise incidence and improve survival statistics without reducing mortality, how to separate lead-time from true overdiagnosis, and what reviewers should demand before trusting the headline.
Lead-Time Bias: When Earlier Diagnosis Pretends to Be Better Survival
A practical guide to lead-time bias for clinical researchers. Covers why screening can improve survival statistics without reducing mortality, how to separate earlier detection from real benefit, and what reviewers should demand before trusting the headline.
Subgroup Analysis: When “Personalized” Findings Are Mostly Multiplicity Wearing a Stethoscope
A practical guide to subgroup analysis for clinical researchers. Covers interaction testing, multiplicity, power failure, post hoc storytelling, and what reviewers should demand before trusting treatment-effect heterogeneity claims.
Explore more topics
These are ranked by how often they appear alongside Methods Critique, so the next click is more likely to be useful than random.