21 Logistic regression for a binary outcome

For a binary outcome, let

\[ \pi_i=P(Y_i=1) \]

be the probability that participant \(i\) experiences the outcome.

Logistic regression uses a linear predictor

\[ f_i=\beta_0+\beta_1X_{i1}+\beta_2X_{i2}+\cdots+\beta_qX_{iq}, \]

and converts it to a probability through the logistic function:

\[ \pi_i=\frac{e^{f_i}}{1+e^{f_i}}. \]

Substituting the linear predictor gives the logistic regression model:

\[ \pi_i= \frac{ \exp(\beta_0+\beta_1X_{i1}+\cdots+\beta_qX_{iq}) }{ 1+\exp(\beta_0+\beta_1X_{i1}+\cdots+\beta_qX_{iq}) }. \]

This transformation ensures that predicted probabilities remain between 0 and 1.

The same model can be written in terms of the log odds:

\[ \log\left(\frac{\pi_i}{1-\pi_i}\right) = \beta_0+\beta_1X_{i1}+\cdots+\beta_qX_{iq}. \]

For predictor \(X_j\), exponentiating its coefficient gives the odds ratio:

\[ OR=e^{\beta_j}. \]

Thus, \(e^{\beta_j}\) is the multiplicative change in the odds associated with a one-unit increase in \(X_j\), holding the other predictors constant.

21.1 Fit the logistic model

Clinical response at 12 weeks is modeled as a function of treatment and prespecified baseline covariates: age, baseline severity, and baseline biomarker. Only variables measured before randomization are included, because post-randomization variables such as adherence may be influenced by treatment and should not be treated as ordinary baseline adjustment variables.

logit_dat <- dat |>
  dplyr::select(clinical_response_12w,treatment_arm,age_years,
                baseline_severity,biomarker_baseline_ng_mL) |>
  drop_na() |>
  mutate(age10=(age_years-mean(age_years))/10,
         biomarker50=(biomarker_baseline_ng_mL-mean(biomarker_baseline_ng_mL))/50)

logit_linear <- glm(
  clinical_response_12w~treatment_arm+age10+baseline_severity+biomarker50,
  family=binomial,data=logit_dat
)

Age is expressed per 10 years and baseline biomarker per 50 ng/mL to make the coefficients easier to interpret.

21.2 Automated check of functional form

Continuous predictors do not need to be normally distributed for logistic regression. The relevant question is whether their relationship with the log odds of the outcome is adequately represented by a straight line.

The following teaching rule compares the linear model with a more flexible spline model.


logit_linear <- glm(
  clinical_response_12w~treatment_arm+age10+baseline_severity+biomarker50,
  family=binomial,data=logit_dat
)

logit_spline <- glm(
  clinical_response_12w~treatment_arm+ splines::ns(age_years,3)+baseline_severity+ splines::ns(biomarker_baseline_ng_mL,3),
  family=binomial,data=logit_dat
)

# Model comparison
logit_form <- anova(logit_linear,logit_spline,test="LRT")
logit_form_p <- logit_form[["Pr(>Chi)"]][2]

if(logit_form_p<.05){
  logit_fit <- logit_spline
  logit_spec <- "Spline model"
} else {
  logit_fit <- logit_linear
  logit_spec <- "Linear model"
}

tibble(
  functional_form_p=logit_form_p,
  selected_model=logit_spec
) |> kbl(digits=4,caption="Automated logistic functional-form selection")
Table 21.1: Automated logistic functional-form selection
functional_form_p selected_model
0.058 Linear model

A small likelihood-ratio p-value favors the more flexible spline model in this teaching workflow. In a confirmatory analysis, the functional form should ideally be specified in advance rather than selected from the same data used for inference.