ALPA High-Yield 🔥 Expanded Biostats Hub: 27 Summary Cards, 50 USMLE MCQs & 63 Active Recall Flashcards!
USMLE Step 1 & 2 CK · IFOM CSE · SMLE · University Exams

Master Biostatistics with Total Confidence

Accelerate your board review with expanded active learning formulas, 50 full USMLE MCQ explanations, 63 interactive flashcards, and direct guidance from A Little Push Academy (ALPA).

0
Summary Cards
0
Practice MCQs
0
Flashcards

High-Yield Core Concepts Checklist (10 Summary Cards)

USMLE Step 1 / Step 2 CK / IFOM Standards
01 Must Know

Sensitivity & Specificity

Fixed intrinsic properties of a screening test. Do NOT change with population prevalence.

  • Sensitivity (TP Rate) A / (A + C)
  • Specificity (TN Rate) D / (B + D)
Mnemonic
SnNout — high Sensitivity, Negative result, rules OUT disease. SpPIn — high Specificity, Positive result, rules IN disease.
02 High Yield

PPV, NPV & Prevalence

Predictive values depend directly on disease Prevalence in the tested population.

  • PPV (True Positive Test) A / (A + B)
  • NPV (True Negative Test) D / (C + D)
As Prevalence : PPV and NPV .
03 New Card

Cut-Off Shift & ROC Curves

Shifting diagnostic cut-off points changes the trade-off between Sensitivity and Specificity.

Lowering Cut-off: Sensitivity (FP ), Specificity . Used for screening.
Raising Cut-off: Specificity (FN ), Sensitivity . Used for confirmation.
ROC Curve: Plots Sensitivity vs (1 - Specificity). Area Under Curve (AUC) = accuracy.
04 Calculations

Odds Ratio vs Relative Risk

Choose based on study design. Odds Ratio for retrospective, Relative Risk for prospective.

  • Odds Ratio (Case-Control) (A×D) / (B×C)
  • Relative Risk (Cohort) [A/(A+B)] / [C/(C+D)]
Rare Disease Assumption: When disease prevalence is low (<5%), OR ≈ RR.
05 Very Common

ARR, NNT & NNH

Quantifying therapeutic benefit or harm in clinical trials.

ARR = RiskControl - RiskTreatment
NNT = 1 / ARR | NNH = 1 / Attributable Risk (AR)
Rule: Always round NNT and NNH UP to the next integer!
06 New Card

Precision, SD & Confidence Intervals

Distinguishing reliability (precision) vs validity (accuracy) and estimating population parameters.

SEM = Standard Deviation / √(n)
95% CI = Mean ± 1.96 × SEM (or ± 2 × SEM)
Larger sample size (n ↑) = smaller SEM = narrower CI = increased precision.
07 New Card

Type I (α) & Type II (β) Errors

Errors in hypothesis decision-making and statistical power.

Type I (α): False positive (saying effect exists when null is true).
Type II (β): False negative (missing true effect). Power = 1 - β.
Power ↑ when: Sample size (n) ↑, effect size ↑, or α threshold ↑.
08 Epidemiology

Study Designs Matrix

Cross-Sectional: Snapshot at 1 point in time. Measures Prevalence.
Case-Control: Cases vs controls. Retrospective. Measures Odds Ratio.
Cohort Study: Exposed vs unexposed followed over time. Measures Incidence & RR.
09 High Yield

Choosing Statistical Tests

t-Test: Compares 2 continuous means (e.g. BP in males vs females).
ANOVA: Compares ≥ 3 continuous means (e.g. BP in 3 drug groups).
Chi-Square (χ²): Compares percentages/proportions of categorical data (large samples).
Fisher's Exact Test: Same as chi-square, but for small sample sizes.
10 New Card

Confounding vs Effect Modification

Distinguish third variable distortions via Stratification Analysis.

Confounding: Stratified RR/OR are equal to each other, but DIFFERENT from crude OR. (Control by matching/randomization).
Effect Modification: Stratified RR/OR are DIFFERENT from each other. (A biological phenomenon; NOT a bias!).
11 New Card

Clinical Trial Phases

Trials occur after preclinical animal/in vitro testing and proceed through 5 phases before and after FDA approval.

  • Phase 0: Microdosing PK/PD in a few volunteers. Often skipped.
  • Phase 1: Small group, healthy or patients. Safety & max tolerated dose.
  • Phase 2: Moderate group with disease. Does it Work? (efficacy + short-term AEs)
  • Phase 3: Large group, RCT vs standard of care. Any Improvement?
  • Phase 4: Postmarketing surveillance. Can it stay on the Market? (rare/long-term AEs)
Mnemonic
"Can I SWIM?" — 0, 1, 2, 3, 4.
12 New Card

Blinding & Crossover Trials

Trial quality improves with randomization and blinding of who knows the treatment assignment.

Double-blind: Neither subject nor researcher knows the group. Triple-blind: also blinds the analysts.
Crossover trial: Each subject receives both treatments in random order, serving as their own control, separated by a washout period.
13 High Yield

ITT vs As-Treated vs Per-Protocol

Intention-to-treat: Analyzed by original random assignment, no exclusions. "Once randomized, always analyzed" — preserves randomization but may dilute true effect.
As-treated: Analyzed by treatment actually received. ↑ risk of bias.
Per-protocol: Excludes subjects who didn't complete treatment as assigned. ↑ risk of bias.
14 Epidemiology

Bradford Hill Criteria for Causation

Necessary but not sufficient principles supporting a causal (not just associative) relationship.

Strength Consistency Specificity Temporality Biological gradient Plausibility Coherence Experiment Analogy
15 Epidemiology

Case Series, Twin, Adoption & Ecological Studies

Case series: Describes patients with the same diagnosis/treatment. No comparison group — cannot show risk-factor association.
Twin concordance study: Compares disease concordance in monozygotic vs dizygotic twins. Measures heritability ("nature vs nurture").
Adoption study: Compares traits in biological vs adoptive relatives to separate genetic from environmental influence.
Ecological study: Compares disease/risk-factor frequency across whole populations, not individuals — prone to the ecological fallacy.
16 High Yield

Likelihood Ratios (LR+ / LR-)

Combine sensitivity and specificity into one number that updates pretest probability directly.

  • LR+ Sens / (1 − Spec)
  • LR− (1 − Sens) / Spec
LR+ > 10 = highly specific test. LR− < 0.1 = highly sensitive test. Pretest odds × LR = posttest odds.
17 New Card

Kaplan-Meier Curve & Hazard Ratio

Kaplan-Meier estimates survival probability over time ("time-to-event" data, often mortality).

Survival probability = 1 − event probability. Significance of the difference between two curves is tested via log-rank test or Cox regression.
Hazard ratio: HR = 1 no difference; HR > 1 event more frequent in treatment/exposure group; HR < 1 protective effect.
18 Very Common

Incidence vs Prevalence

  • Incidence # new cases / at-risk pop, per time
  • Prevalence Incidence × avg. duration
Prevalence Incidence for short-duration disease (common cold). Prevalence > Incidence for chronic disease (diabetes). Prevalence PPV, NPV.
19 New Card

Mortality, Attack & Case Fatality Rate

  • Mortality rate Deaths / population, per year
  • Attack rate Ill people / people exposed
  • Case fatality rate Deaths / cases × 100
20 Epidemiology

Demographic Transition & Population Pyramids

As a country develops, birth and mortality rates decline to varying degrees, reshaping the population's age structure.

Short life expectancy + growing population = wide-based pyramid. Long life expectancy + declining population = top-heavy pyramid.
21 Must Know

Mean, Median & Mode

Mean: Sum of values / count. Most affected by outliers.
Median: Middle value (average of 2 middles if even count). Best for skewed data.
Mode: Most common value. Least affected by outliers.
22 New Card

Normal, Bimodal & Skewed Distributions

Normal: Mean = median = mode. 68% within 1 SD, 95% within 2 SD, 99.7% within 3 SD.
Bimodal: Suggests 2 distinct populations (e.g., fast vs slow acetylators).
Positive skew: Mean > median > mode, tail right. Negative skew: Mean < median < mode, tail left.
23 Must Know

H0, H1 & the P-Value

H0 (null): No difference/association between groups.
H1 (alternative): At least one difference/association exists.
P-value: Probability of results this extreme if H0 is true. P < 0.05 (typical α) → reject H0.
24 High Yield

Confidence Intervals

Range expected to contain the true population value, with a specified probability (CI = 1 − α).

  • 95% CIMean ± 1.96×SE
  • 99% CIMean ± 2.58×SE
H0 rejected when the 95% CI for a mean difference excludes 0, or for an OR/RR excludes 1.
25 New Card

Meta-Analysis

Pools summary data (means, RRs) from multiple studies for a more precise effect estimate, and estimates heterogeneity between studies.

Improves power, strength of evidence, and generalizability — but is limited by the quality of the included studies and selection/publication bias.
26 High Yield

Statistical vs Clinical Significance

Statistical significance: Low probability the result is due to chance.
Clinical significance: The size of the real-world impact on patient outcomes.
A study can be highly statistically significant with a clinically trivial effect size — watch for this trap on exam vignettes.

Expanded Epidemiological Biases & How to Control Them

Type of Bias Definition & Clinical Vignette Pattern Method to Reduce / Control
Selection / Berkson Bias Non-random sampling or selecting hospitalized controls (Berkson bias) that don't represent the true target population. Convenience sampling (enrolling whoever is easiest to reach) is a common sub-type. Randomization, high response rate, correct comparison group
Recall Bias Subjects with negative outcome recall past exposures more accurately than healthy controls (common in retrospective case-control studies). Decrease time from exposure to follow-up; use medical records as source
Measurement / Procedure Bias Information is gathered in a systematically distorted way (e.g., a faulty automatic BP cuff), or subjects in different groups aren't treated/tested the same way (procedure bias). Objective, standardized, pre-planned data collection; blinding
Hawthorne Effect Study subjects alter their behavior because they know they are actively being observed ("Hawthorne watches you like a hawk"). Placebo control, unobtrusive monitoring
Observer-Expectancy (Pygmalion) Bias Researcher's belief in a treatment's efficacy changes the outcome of that treatment — e.g., an observer expecting improvement is more likely to document a positive outcome. Blinding (masking) of researchers analyzing outcomes
Lead-Time Bias Early detection by screening creates false impression of prolonged survival without altering true disease trajectory or outcome. Measure "back-end" survival adjusted for severity at diagnosis
Length-Time Bias Screening preferentially detects slowly progressive, less aggressive cases (long latency) over rapidly fatal ones, giving an illusion of improved survival. Randomized controlled trial of the screening program itself
Confounding An unmeasured factor is associated with BOTH the exposure and outcome, and is not on the causal pathway (e.g., coffee drinking appears linked to lung cancer, but smoking is the true cause). Matching, randomization, stratified/regression analysis, crossover design

Quick Answers

Is this Biostatistics & Epidemiology hub really free?

Yes — every summary card, MCQ, and flashcard on this page is free to use, with no signup or account required.

Which exams does this cover?

Biostatistics and Epidemiology as tested on USMLE Step 1, Step 2 CK, and IFOM — the same core concepts (study design, bias, diagnostic testing, risk measures, hypothesis testing) show up across all three.

How much content is actually here?

27 high-yield summary cards organized by topic, 50 practice MCQs with full explanations, and 63 active-recall flashcards — all filterable by topic so you can drill a single weak area instead of starting from scratch.

Do I need to make an account to track my progress?

No login. Your quiz accuracy and flashcard progress are saved anonymously to this browser, so returning here later picks up where you left off — just note that switching devices or browsers starts a fresh slate, since nothing is tied to an account.

Is this the same content as ALPA's full courses?

This hub is a free, focused deep-dive into one high-yield subject. ALPA's full courses and study groups cover the complete USMLE/IFOM/SMLE curriculum with live mentorship — see the ALPA Courses tab above to learn more.