← Back to Blog
Causal InferenceReal-World EvidenceMethods Critique

Weak Instruments and Physician Preference IVs: When Treatment Movement Is Not Yet Causal Credibility

June 30, 2026·16 min read

Anas H. Alzahrani, MD PhD MPH

Department of Preventive Medicine and Public Health

Faculty of Medicine, King Abdulaziz University

Instrumental variables are attractive because they promise an escape hatch when ordinary adjustment cannot settle unmeasured confounding. In clinical observational research, that promise often gets handed to physician preference, hospital preference, or other practice-style instruments that supposedly nudge treatment choice without touching prognosis directly.

Sometimes that logic is useful. Often it is used too casually. A preference-based instrument can move treatment just enough to produce a coefficient while still being too weak, too behaviorally entangled, or too local to support the paper's headline claim.

The Core Decision Rule

Do not ask only whether the instrument predicts treatment. Ask whether it predicts treatment strongly enough, stays plausibly isolated from the outcome pathway, and still targets a clinically meaningful group once you admit the estimate is local.

Decision rule:

A physician-preference IV deserves trust only when the first stage is clearly stronger than token, the preference does not bundle cointerventions or referral style, and the manuscript speaks plainly about a local effect among compliers rather than pretending to have identified the average treatment effect for everyone.

The cleanest summary is simple: a weak IV is not a safer IV. It is just a harder problem wearing more notation.

Why Researchers Reach for Preference-Based Instruments

Treatment selection looks confounded

Sicker patients receive one therapy, frailer patients avoid another, and ordinary regression looks unconvincing even before the discussion section begins.

Practice style seems naturally available

Some clinicians prescribe more aggressively, adopt new drugs earlier, or favor one procedural strategy across similar patients.

The method sounds like a causal rescue

Once the paper says “instrumental variables,” readers may stop asking whether the instrument is strong, isolated, and clinically interpretable enough to earn that confidence.

A Concrete Clinical Example

Case

Physician preference for early DOAC versus warfarin after new atrial fibrillation diagnosis

Imagine a registry study comparing early direct oral anticoagulant initiation with warfarin among patients newly diagnosed with atrial fibrillation. Measured confounding is extensive: frailty, bleeding risk, kidney function, monitoring capacity, and clinician comfort all shape treatment choice.

The investigators instrument treatment with each clinician's historical DOAC preference. On the surface, that sounds sensible. But what else comes with that preference? Early adopters may also use different renal monitoring workflows, refer to anticoagulation clinics differently, manage follow-up intensity differently, and react to minor bleeding events differently.

That is the real IV question. Not “does preference predict treatment?” It usually does. The question is whether preference predicts only treatment strongly enough, and whether the resulting local effect would still matter to the clinical reader you are trying to persuade.

Interactive weak-IV stress test

Watch a physician-preference instrument move treatment more easily than it earns trust

This teaching tool uses a simple one-instrument first-stage approximation. It is not a formal diagnostic package. The goal is to connect strength, exclusion leakage, and local interpretation in one place rather than treating the F-statistic as a permission slip.

VerdictPotentially usable, but only with a narrow and defensive interpretationThis setup may support a local effect claim, but reviewers should expect stronger diagnostics and less rhetorical confidence.

Bigger samples help, but a large database cannot rescue an instrument that barely changes treatment.

This is the share of treatment variation uniquely explained by the instrument after the measured covariates have already had their turn.

Physician preference often changes treatment probability only modestly. Small shifts can still be publishable, but they make every other IV assumption work harder.

A preference-based instrument is fragile when the same clinician tendency changes follow-up intensity, supportive care, diagnostics, or referral patterns.

IV analyses estimate a local effect. If only a sliver of patients are true compliers, the claim may be mathematically real and still clinically hard to generalize.

Approximate first-stage F
24.1

Crossing 10 is not a magic trick. It only means the instrument is less visibly weak than before.

Patients whose treatment assignment is visibly moved
320

This rough signal count reminds you how little leverage a preference instrument can have even in a large database.

Likely interpretation scope
Narrow

The smaller the complier group, the more the estimate behaves like a local policy fragment rather than a population-wide treatment effect.

QuestionCurrent signalWhy it matters
Is the instrument obviously weak?Less obviously weak on this simple diagnostic.Weak instruments can pull IV estimates back toward confounded ordinary regression while inflating false confidence.
How believable is exclusion?Moderate risk of pathway leakageIf clinician preference also changes surveillance, rescue therapy, or referral, the instrument is not isolated from outcome pathways.
How broad is the estimand?Narrow local effect among compliersA narrow complier slice may still matter scientifically, but it should not be narrated like the average treatment effect for everyone.

Where Weak Instruments Usually Fail

The first stage is technically nonzero but clinically tiny

In a giant database, a flimsy preference signal can still be “statistically significant.” That is not the same as being strong enough to resist weak-IV bias.

Preference carries an entire care style, not just treatment choice

When the same clinicians also differ in surveillance, rescue therapy, or admission thresholds, the exclusion restriction starts leaking before the estimator even begins.

The local effect gets narrated like a population effect

IV estimates are usually about patients whose treatment is actually moved by the instrument. That subgroup may be narrow, policy-specific, and behaviorally unusual.

Balance checks distract from the real assumptions

Covariate balance across instrument levels can be reassuring, but it does not prove exclusion and does not rescue a first stage that barely exists.

Physician Preference Is Not Automatically Exogenous

Preference-based instruments often get smuggled in under the phrase “quasi-random variation.” That phrase should make reviewers more alert, not less. Clinician preference may correlate with training era, subspecialty access, team staffing, monitoring culture, and referral patterns. Those are outcome pathways in disguise.

What the paper saysWhat you should askWhy it matters
“Physician preference determined treatment”Did the same preference also change monitoring, referral, or ancillary treatment?If yes, the instrument is touching outcome risk through more than one pathway.
“The first-stage F-statistic exceeded 10”How large was the actual treatment shift, and how stable was it across subgroups and time?A threshold crossed on paper can still correspond to a clinically weak and fragile instrument.
“Baseline characteristics were balanced across instrument levels”What about unmeasured workflow differences and post-baseline care pathways?Observed balance cannot verify exclusion or monotonicity.
“The IV estimate differed from the adjusted estimate”Does that difference signal less confounding, or a weaker and more local estimand?Disagreement alone is not a victory lap for the IV analysis.

When a Preference-Based IV Is More Defensible

More defensible

  • The preference strongly shifts treatment, not just statistically but meaningfully.
  • The care pathway attached to the preference is otherwise tightly standardized.
  • The paper states upfront that the estimand is local and describes likely compliers.
  • Negative-control or falsification checks probe whether the preference carries unrelated outcome pathways.

Less defensible

  • The instrument barely changes treatment uptake.
  • Preference is bundled with clinician quality, site resources, or follow-up intensity.
  • The paper narrates the estimate as if it were the average effect in the full cohort.
  • All diagnostics are baseline balance tables plus one F-statistic and a confident abstract.

Reviewer Red-Flag Checklist

  • Ask how much treatment uptake actually changed across instrument levels, not just whether the first stage was “significant.”
  • Ask what clinician or site behaviors travel with the instrument besides treatment choice.
  • Ask whether the manuscript identifies the likely complier subgroup and explains why that local effect is clinically relevant.
  • Ask whether falsification outcomes or pathway-based sensitivity checks support the exclusion story.
  • Ask whether the discussion section overclaims a population effect from what is really a narrow local contrast.

The Practical Bottom Line

Weak instruments are not a technical footnote. They are a signal that the causal rescue may be too fragile for the question being asked. Physician-preference IVs deserve the same practical skepticism you would apply to any other clinical design: what changed, for whom, through which pathway, and how much interpretive weight can that change really carry?

If you are reviewing a study where the IV section feels more confident than the protocol logic, that is exactly the kind of manuscript Aqrab is designed to stress-test. You can start with Aqrab Try Free for a methods critique, or use the developer API when you want that logic embedded directly in your own review workflow.

References and Further Reading

  • Stock JH, Yogo M. Testing for weak instruments in linear IV regression.
  • Angrist JD, Imbens GW, Rubin DB. Identification of causal effects using instrumental variables.
  • Brookhart MA, Wang PS, Solomon DH, Schneeweiss S. Instrumental variable methods in comparative safety and effectiveness research.

Keep reading

Don't stop at one method.

Good methods judgment comes from contrast. Read the neighboring guides, see where the assumptions diverge, and avoid treating every observational problem like it needs the same hammer.

Browse full archive