Modern tidyverse patterns, style guide, and migration guidance for R development...
This skill covers modern tidyverse patterns for R 4.3+ and dplyr 1.1+, style guidelines, and migration from legacy patterns.
Always use native pipe |> instead of magrittr %>%
R 4.3+ provides all needed features. See pipe-examples.md for usage patterns.
Use join_by() instead of character vectors for joins
Modern join syntax supports:
join_by(company == id)join_by(company == id, year >= since)join_by(company == id, closest(year >= since))See join-examples.md for complete patterns.
Use multiple and unmatched arguments for quality control:
multiple = "error" - Expect 1:1 matchesmultiple = "all" - Allow multiple matches explicitlyunmatched = "error" - Ensure all rows matchUnderstand the difference:
arrange(), filter(), mutate(), summarise()select(), relocate(), across()Key patterns:
{{}} (embrace) for function arguments.data[[]] for character vectorsacross() for multiple columnsSee data-masking-examples.md for patterns.
Use .by for per-operation grouping (dplyr 1.1+)
This replaces the old group_by() |> ... |> ungroup() pattern.
Additional modern operations:
pick() - Column selection inside data-masking functionsacross() - Apply functions to multiple columnsreframe() - Multi-row summariesSee grouping-examples.md for complete examples.
Use stringr over base R string functions
Benefits:
str_ prefixSee stringr-examples.md for common patterns and base R equivalents.
Good: day_one, calculate_mean, user_data
Avoid: DayOne, calculate.mean, userData
See style-examples.md for proper spacing and pipe formatting.
. (e.g., .data, .by)| Avoid | Use Instead |
|---|---|
%>% |
` |
by = c("a" = "b") |
by = join_by(a == b) |
sapply() |
map_*() |
| `group_by() | > ... |
sapply() - Type-unstable, use map_*() insteadSee anti-patterns.md for examples of what to avoid and correct alternatives.
| Base R | Modern Tidyverse |
|---|---|
subset(data, condition) |
filter(data, condition) |
data[order(data$x), ] |
arrange(data, x) |
aggregate(x ~ y, data, mean) |
summarise(data, mean(x), .by = y) |
sapply(x, f) |
map(x, f) |
grepl("pattern", text) |
str_detect(text, "pattern") |
gsub("old", "new", text) |
str_replace_all(text, "old", "new") |
| Old Pattern | New Pattern |
|---|---|
data %>% function() |
`data |
| `group_by(x) | > summarise() |
by = c("a" = "b") |
by = join_by(a == b) |
gather()/spread() |
pivot_longer()/pivot_wider() |
map_dfr(x, f) |
`map(x, f) |
separate(col, into = ...) |
separate_wider_delim() |
See migration-examples.md for complete migration patterns.
source: Sarah Johnson's gist https://gist.github.com/sj-io/3828d64d0969f2a0f05297e59e6c15ad