R performance profiling, benchmarking, and optimization strategies...
This skill covers profiling, benchmarking, parallelization, and performance best practices for R.
Profile before optimizing - Use profvis and bench to identify real bottlenecks. Write readable code first, optimize only when necessary.
| Tool | Use When | Don't Use When | What It Shows |
|---|---|---|---|
profvis |
Complex code, unknown bottlenecks | Simple functions, known issues | Time per line, call stack |
bench::mark() |
Comparing alternatives | Single approach | Relative performance, memory |
system.time() |
Quick checks | Detailed analysis | Total runtime only |
Rprof() |
Base R only environments | When profvis available | Raw profiling data |
See profiling-workflow.md for the complete workflow.
in_parallel())Helps when:
Hurts when:
See parallel-examples.md for decision points.
| Backend | Use When |
|---|---|
| data.table | Very large datasets (>1GB), complex grouping, maximum performance critical |
| dplyr | Readability priority, complex joins/window functions, moderate data (<100MB) |
| base R | No dependencies allowed, simple operations, teaching/learning |
See backend-selection.md for guidance.
See profiling-best-practices.md for examples.
See performance-anti-patterns.md for examples.
| Superseded | Modern Replacement |
|---|---|
map_dfr(x, f) |
map(x, f) |> list_rbind() |
map_dfc(x, f) |
map(x, f) |> list_cbind() |
map2_dfr(x, y, f) |
map2(x, y, f) |> list_rbind() |
walk()Use walk() and walk2() for side effects (file writing, plotting).
Use in_parallel() with mirai for scaling across cores.
See purrr-patterns.md for all patterns.
When speed is critical, consider:
Profile to identify whether these tools will help your specific bottleneck.
source: Sarah Johnson's gist https://gist.github.com/sj-io/3828d64d0969f2a0f05297e59e6c15ad