← Back to Blog
Methods CritiqueEvidence AppraisalStudy Design

Estimation Over Testing: Four Habits That Quietly Break Clinical Statistics

July 29, 2026·15 min read

Anas H. Alzahrani, MD PhD MPH

Department of Preventive Medicine and Public Health

Faculty of Medicine, King Abdulaziz University

A trial reports its primary result and the whole clinical conversation collapses into a single word. Significant. Or worse, not significant. From there the paper writes itself: the drug “works” or it “doesn’t,” the guideline moves or it doesn’t, the next trial is funded or shelved. An entire apparatus of judgment has been outsourced to whether one number fell below 0.05.

A recent methodological critique in the statistics literature — written in the tradition of Charles Manski’s work on research with incredible certitude — argues that this is not a collection of isolated mistakes but a system of entrenched habits, four of them, that quietly corrode clinical statistics and the decisions built on it. What follows is a working clinician-and-methodologist tour of those four habits: what each one gets wrong, why it survives, and what to reach for instead. None of the fixes are exotic. They are the analyses careful people were always supposed to do.

Habit 1: Testing When the Question Is Estimation

A significance test answers one narrow question: can we reject the hypothesis that the effect is exactly zero? No patient has ever asked that. A patient asks how much better they will be, and how sure we are. A guideline committee asks whether the plausible range of benefit clears the bar of harm and cost. Those are questions about magnitude and uncertainty — the province of estimation, not testing.

The p-value discards exactly the information the decision needs. It compresses an estimate and its interval into a verdict about a point value no one believes literally (effects are essentially never exactly zero). Two failures follow, and they are mirror images. A large trial can make a clinically trivial effect “significant,” because significance is partly a statement about sample size. And a small trial can return “not significant” while its interval remains perfectly consistent with a large, important benefit — the classic confusion of absence of evidence with evidence of absence.

Decision rule:

Lead with the estimate and its interval, judged against a pre-stated clinically important difference. Use the p-value, if at all, as a footnote — never as the headline and never as a substitute for reading where the interval sits. “Significant” is not a synonym for “worth it,” and “not significant” is not a synonym for “no effect.”

See the Gap

The explorer below is one two-arm trial read three ways. Move the observed effect, the sample size, and the smallest difference you would change practice for, and watch the testing verdict, the estimate with its interval, and the clinical decision come apart. Start with Significant but trivial: the p-value glows green while the entire interval sits inside the range you already called too small to matter. Then try Not significant ≠ no effect and see an interval that runs from a large benefit clear past no effect get written up as a “null.”

Testing vs. estimation vs. decision

One trial, three readings. A two-arm study with a 20% control-arm risk. Move the observed effect, the sample size, and the smallest difference a clinician would act on, then watch the p-value, the interval, and the clinical decision come apart.

Risk difference, treatment minus control. Negative is a benefit.

Only the p-value and the interval width move. The effect does not.

The smallest benefit a clinician would change practice for.

Testing lens

p = 0.016

Significant at 0.05

Estimation lens

-1.5 pp

95% CI -2.7 to -0.3

Decision lens

Unsettled

MCID = 3 pp benefit

Where the 95% interval sits (percentage points)

no effectMCID

Statistically significant, clinically trivial.

The p-value is 0.016, so the test proudly rejects zero and the abstract will say the treatment "worked." But the 95% interval — -2.7 to -0.3 percentage points — sits entirely inside the range you already agreed was too small to matter. A large enough trial can make any non-zero effect "significant." Significance is a statement about zero; it is not a statement about whether the effect is worth having.

The interval shows a benefit, but its whole range is smaller than the minimum clinically important difference. Detectable, but not the size a patient would notice.

Habit 2: Scoring Prediction Far From the Bedside

The second habit is to judge a clinical prediction model by how well it discriminates — an AUC, a Brier score, an accuracy figure on a held-out set — as if statistical separation were the same thing as clinical usefulness. It is not. A model can rank patients beautifully and still change no decision, because it is never wrong at a threshold anyone would act on, or because it is superbly discriminating and badly calibrated, so its probabilities cannot be trusted as probabilities.

Discrimination is a property of the model in the abstract. Usefulness is a property of the model in a decision — at a specific threshold, against the real cost of a false alarm and a missed case. That is why decision-curve analysis and net-benefit evaluation exist: they ask whether using the model beats treating everyone or no one across the range of thresholds clinicians actually hold. And it is why calibration — and its decay over time — matters more than a headline AUC. An accuracy number that is never tied to a decision is a statistic in search of a purpose.

Habit 3: Meta-Analysis on Autopilot

The third habit is to pool first and think later: drop a set of study estimates into a random-effects model, report the diamond, and treat the heterogeneity between studies as a nuisance to be summarized away with an I² and moved past. But a pooled average is only meaningful if the studies are estimating the same thing. When they define the exposure differently, follow patients for different lengths, enroll different populations, or answer subtly different causal questions, the pooled number is an average of incommensurable quantities — precise, and about nothing in particular.

Heterogeneity is usually the finding, not the noise. It is the signal that the effect depends on context, and that context is what a reader needs. The disciplined alternative is to define the estimand before pooling, to check that comparisons are exchangeable across studies the way network meta-analysis transitivity demands, and to weigh indirectness honestly rather than letting a forest plot flatten it into one dot.

Habit 4: Research With Incredible Certitude

The fourth habit is the deepest, and it is the one Manski named: reporting a point estimate with a tight, model-based interval that quietly assumes away everything that cannot be verified. The interval in an observational effect estimate is honest only about sampling error. It says nothing about the assumption that there is no unmeasured confounding, that missing data are missing at random, that the model form is correct. Those assumptions are not testable from the data at hand, and pretending they hold exactly is how a study manufactures a precision it has not earned.

The honest posture is partial identification: report what the data can support under a range of defensible assumptions, not a single number under the most convenient one. That is the logic behind an E-value (how strong would unmeasured confounding have to be to explain the result away?), behind sensitivity analysis for data not missing at random, and behind bounding a result rather than pinning it. A wider, honest interval is worth more than a narrow one that is quietly conditional on a fairy tale.

Why the System Holds It in Place

None of these habits survive because researchers are careless. They survive because the incentives point that way. Statistical training for many clinicians is thin and test-centric, so the p-value is the tool they were handed and the one they trust. Analysis is often delegated to a consulting biostatistician brought in late, which turns methods into a gatekeeping ritual rather than a design conversation. And journals, funders, and regulators reward a clean verdict — a significant primary endpoint, a single headline number — over an honest range. The critique’s real target is not the individual analyst but the machinery that makes the shortcut rational. Naming the machinery is the first step to not being ruled by it.

Reviewer Red Flags

What a defensible study does

  • Leads with an effect estimate and interval, judged against a pre-stated clinically important difference.
  • Evaluates prediction models by net benefit and calibration, not AUC alone.
  • Defines the estimand before pooling and treats heterogeneity as a finding to explain.
  • Reports sensitivity analyses or bounds for the assumptions it cannot test.

What should make you nervous

  • A conclusion carried entirely by “p < 0.05” or “no significant difference.”
  • A model sold on its AUC with no threshold, no net benefit, no calibration.
  • A pooled estimate over visibly different studies with heterogeneity waved off.
  • A tight observational interval and no sensitivity analysis for confounding or missingness.

Decision Rules That Travel Well

  1. Report the estimate and its interval first; make the p-value earn its place, if it appears at all.
  2. Interpret every result against a clinically important difference set before you saw the data.
  3. Judge prediction by the decision it informs — net benefit and calibration — not by discrimination alone.
  4. Define the estimand before you pool; let heterogeneity teach you rather than average it away.
  5. For any untestable assumption, show how far the conclusion travels before it breaks.

How This Connects to the Rest of the Toolkit

These four habits are the through-line under a great deal of applied methodology. The estimation-first posture is exactly what the estimand framework formalizes: decide what you are estimating before you argue about how. The prediction critique is the same distinction we drew in prediction versus causation — a model that predicts well is not thereby a model that tells you what to do. And the certitude critique is why sensitivity analysis is not an optional appendix but the part of the paper that tells you whether to believe the headline. The common thread is modest and demanding at once: say what you are estimating, show how sure you are, and do not claim more certainty than your assumptions can buy.

Where Aqrab Fits

Every one of these habits is easy to miss precisely because it looks like rigor. A p-value, an AUC, a pooled diamond, a tight confidence interval — each is a real number, correctly computed, that answers a question slightly to the side of the one the reader needs answered. Aqrab reads a study the way a careful methodologist would: checking whether the conclusion rests on a significance verdict or on an estimate against a clinical threshold, whether a prediction claim is tied to a decision, whether a pooled result respects what the studies actually estimated, and whether the certainty on display is backed by a sensitivity analysis or merely assumed.

If you are appraising a study — or writing one — start with Aqrab Try and pressure-test the reasoning before a clean verdict stands in for a careful one.

Keep reading

Don't stop at one method.

Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.

Browse full archive