Teach Individual Patient Data (IPD) meta-analysis methods for analyzing raw participant-level data from multiple studies...
This skill teaches IPD meta-analysis, the "gold standard" of evidence synthesis that uses raw participant-level data from multiple studies.
IPD meta-analysis analyzes the original individual-level data from each study rather than summary statistics. This enables more powerful analyses, proper handling of time-to-event data, and exploration of patient-level effect modifiers.
Activate this skill when users:
Comparison:
| Aspect | Aggregate Data | IPD |
|---|---|---|
| Data level | Study summaries | Individual patients |
| Subgroup analysis | Ecological bias risk | Patient-level, unbiased |
| Time-to-event | Requires approximations | Exact analysis |
| Missing data | Cannot address | Can model properly |
| Standardization | Limited | Full flexibility |
| Effort | Low | High (data collection) |
Socratic Questions:
Two-Stage Approach:
Stage 1: Analyze each study separately
→ Get study-specific estimates
Stage 2: Combine estimates using standard MA
→ Pool using random effects
One-Stage Approach:
Single model: All data in one hierarchical model
→ Accounts for clustering within studies
→ More flexible for complex analyses
When to Use Each:
| Situation | Recommended Approach |
|---|---|
| Simple outcomes, many studies | Two-stage (simpler) |
| Few studies, sparse data | One-stage (more stable) |
| Complex interactions | One-stage (more flexible) |
| Time-to-event | One-stage (preferred) |
| Non-linear effects | One-stage (necessary) |
Stage 1 - Study-Level Analysis:
library(dplyr)
library(broom)
# Analyze each study separately
study_results <- ipd_data %>%
group_by(study_id) %>%
do(tidy(glm(outcome ~ treatment + age + sex,
data = .,
family = binomial))) %>%
filter(term == "treatment")
# Extract treatment effects and SEs
effects <- study_results %>%
select(study_id, estimate, std.error)
Stage 2 - Meta-Analysis:
library(metafor)
# Standard random-effects MA
ma_result <- rma(
yi = effects$estimate,
sei = effects$std.error,
method = "REML"
)
summary(ma_result)
forest(ma_result)
Mixed-Effects Model:
library(lme4)
# One-stage with random intercepts and slopes
model <- glmer(
outcome ~ treatment + age + sex +
(1 + treatment | study_id),
data = ipd_data,
family = binomial
)
summary(model)
Interpretation:
For Time-to-Event:
library(survival)
library(coxme)
# Stratified Cox model (two-stage equivalent)
cox_stratified <- coxph(
Surv(time, event) ~ treatment + age + sex + strata(study_id),
data = ipd_data
)
# Frailty model (one-stage)
cox_frailty <- coxme(
Surv(time, event) ~ treatment + age + sex + (1 | study_id),
data = ipd_data
)
Why IPD is Essential:
Interaction Analysis:
# Test treatment-covariate interaction
model_interaction <- glmer(
outcome ~ treatment * age_group + sex +
(1 + treatment | study_id),
data = ipd_data,
family = binomial
)
# Compare with main effects model
anova(model_main, model_interaction)
Visualization:
library(ggplot2)
# Forest plot by subgroup
ggplot(subgroup_effects, aes(x = estimate, y = subgroup)) +
geom_point() +
geom_errorbarh(aes(xmin = ci_low, xmax = ci_high), height = 0.2) +
geom_vline(xintercept = 0, linetype = "dashed") +
labs(x = "Treatment Effect (log OR)", y = "Subgroup")
Common Approaches:
| Method | Description | Assumption |
|---|---|---|
| Complete case | Exclude missing | MCAR (rarely true) |
| Single imputation | Fill with mean/mode | Underestimates uncertainty |
| Multiple imputation | Create multiple datasets | MAR |
| Pattern mixture | Model missingness | MNAR sensitivity |
Multiple Imputation with IPD:
library(mice)
# Impute within each study
imputed_data <- ipd_data %>%
group_by(study_id) %>%
group_modify(~ {
mice(.x, m = 20, method = "pmm", printFlag = FALSE) %>%
complete("long")
})
# Analyze each imputed dataset
results <- imputed_data %>%
group_by(.imp) %>%
do(tidy(glmer(outcome ~ treatment + (1|study_id),
data = ., family = binomial)))
# Pool results using Rubin's rules
pool(results)
Common Challenges:
Harmonization Steps:
# Standardize variables across studies
harmonized <- ipd_data %>%
mutate(
# Standardize age (z-score within study)
age_std = (age - mean(age)) / sd(age),
# Harmonize outcome timing
outcome_6mo = case_when(
study_id == "A" ~ outcome_week24,
study_id == "B" ~ outcome_month6,
TRUE ~ outcome_6months
),
# Recode categorical variables
sex = case_when(
sex %in% c("M", "male", "1") ~ "Male",
sex %in% c("F", "female", "2") ~ "Female"
)
)
PRISMA-IPD Checklist Items:
Basic: "What is the main advantage of IPD over aggregate data meta-analysis?"
Intermediate: "When would you choose a one-stage over a two-stage approach?"
Advanced: "How would you handle a situation where you have IPD for 60% of studies and only aggregate data for the rest?"
"IPD always gives different results than aggregate MA"
"One-stage is always better than two-stage"
"IPD eliminates all bias"
User: "I'm coordinating an IPD meta-analysis of 8 cancer trials. How do I analyze survival outcomes?"
Response Framework:
Glass (the teaching agent) MUST adapt this content to the learner:
Example Adaptations:
meta-analysis-fundamentals - Basic concepts prerequisitedata-extraction - Data collection principlesheterogeneity-analysis - Understanding between-study variationbayesian-meta-analysis - Alternative modeling framework