0.1 Course scope
This course covers exploratory data analysis, sampling distributions and resampling, two-group inference, categorical methods, ANOVA/MANOVA, correlation, linear regression, logistic regression, count models, and an integrated biomedical assignment.
0.2 Learning objectives
By the end of the course, you should be able to:
- distinguish variable types and choose defensible descriptive summaries;
- separate a data distribution from a sampling distribution;
- explain standard errors, confidence intervals, bootstrap resampling, and permutation tests;
- select and interpret independent, paired, parametric, rank-based, and categorical tests;
- fit and diagnose one-way and factorial ANOVA models, MANOVA, and correlation analyses;
- fit linear, logistic, and count regression models and interpret coefficients on the correct scale;
- use baseline-adjusted models appropriately for baseline/follow-up trial outcomes;
- report estimates, confidence intervals, assumptions, and limitations without treating a p-value as a verdict.
0.3 Analytical principles used throughout
Estimate first. State the estimand, report its point estimate and uncertainty, then use a hypothesis test as one component of the evidence rather than as the sole conclusion. Confidence intervals and p-values have precise repeated-sampling meanings and should not be translated into statements about the probability that a null hypothesis is true (Wasserstein and Lazar 2016; Greenland et al. 2016).
Design determines dependence. Independent groups, paired observations, clustered observations, repeated measures, and stratified designs require different analyses because the observational units are not interchangeable.
Assumptions belong to models, not labels. For t-tests, ANOVA, and regression, inspect the distribution and structure of the errors/residuals and the consequences of departures from assumptions. A significant Shapiro-Wilk test alone is not a reason to replace a mean-based question by a rank-based question.
Data used in this edition. The active teaching sheet contains 360 synthetic participants (120 per randomized arm). To download the data visit this link: