8 Multiple testing

When several hypotheses are tested, each test creates another opportunity to obtain a small p-value by chance. Multiplicity adjustment accounts for this when the hypotheses belong to the same planned family of comparisons.

The following example compares Drug_A with Control for three continuous outcomes using Welch’s two-sample t-test.

outcomes <- c(SBP="sbp_change",
              Biomarker="biomarker_change",
              QoL="qol_change")

p_raw <- sapply(outcomes,\(v){
  d <- dat |> filter(treatment_arm %in% c("Control","Drug_A")) |> drop_na(all_of(v))
  t.test(d[[v]][d$treatment_arm=="Drug_A"],
         d[[v]][d$treatment_arm=="Control"],
         var.equal=FALSE)$p.value
})

tibble(
  outcome=names(outcomes),
  p_raw=p_raw,
  p_Holm=p.adjust(p_raw,method="holm"),
  p_BH=p.adjust(p_raw,method="BH")
) |> kbl(digits=4,
         caption="Raw and multiplicity-adjusted p-values for three outcomes")
Table 8.1: Raw and multiplicity-adjusted p-values for three outcomes
outcome p_raw p_Holm p_BH
SBP 0 0 0
Biomarker 0 0 0
QoL 0 0 0

8.1 Holm and Benjamini–Hochberg adjustments

Holm adjustment controls the family-wise error rate, the probability of making at least one false-positive conclusion within the defined family of tests. It is appropriate when strong control of false-positive claims is important.

Benjamini–Hochberg adjustment controls the false discovery rate, the expected proportion of false positives among the results declared significant. It is commonly used when a larger set of hypotheses is examined.

Goal Adjustment
strongly limit the chance of any false-positive claim Holm
control the proportion of false discoveries among significant results Benjamini–Hochberg

Multiplicity adjustment changes the p-values used for inference; it does not change the observed effect estimates.