Complete setup for AMD Strix Halo (Ryzen AI MAX+ 395) PyTorch environments...
Set up a reproducible gfx1151 environment, then prove the capabilities the workload needs. Do not infer support from GPU enumeration or a large allocation.
Commands below assume the skill is installed globally at
~/.claude/skills/strix-halo-setup/. For a per-project copy, substitute
./.claude/skills/strix-halo-setup/.
Run the system verifier:
~/.claude/skills/strix-halo-setup/scripts/verify_system.sh
Choose an installation track with the user:
| Track | Use when | Tradeoff |
|---|---|---|
| AMD supported | Stability and AMD's validated matrix matter most | Older PyTorch and kernel stack |
| TheRock multi-arch | New PyTorch, Triton, AOTriton, or rapid gfx1151 fixes matter most | Moving nightly packages; regressions are possible |
Default to the AMD supported track unless the user explicitly prioritizes new features or agrees to nightly risk. See installation details.
Create a fresh virtual environment. AMD's manylinux wheels ship cp310 through cp313, so match the interpreter to the wheel tag rather than assuming 3.12. Never install these wheels into the system Python or an existing environment unless the user requests it.
Install one track. Do not combine AMD supported wheels, PyTorch.org wheels, old per-family TheRock wheels, or multi-arch TheRock packages in one environment.
Run the capability verifier:
python ~/.claude/skills/strix-halo-setup/scripts/verify_pytorch.py
For AOTriton flash attention, start a new process with the experimental switch and force the backend during verification:
TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTAL=1 \
python ~/.claude/skills/strix-halo-setup/scripts/verify_pytorch.py \
--require flash_attention
Capture the resolved environment after it passes:
python -m torch.utils.collect_env > collect-env.txt
python -m pip freeze > requirements-lock.txt
AMD's ROCm Ryzen matrix validates gfx1151 with PyTorch 2.9.1 and FP16. Install
AMD's exact repo.radeon.com wheels as documented in
INSTALLATION.md, not similarly named PyTorch.org wheels.
AMD ships patch releases faster than this skill is revised, and every one
rewrites the git hash in each filename โ list
repo.radeon.com/rocm/manylinux/ and
take the newest rocm-rel-* rather than trusting a pinned URL.
Use this track when reproducibility is more important than the newest compiler or attention work.
TheRock replaced new per-family releases with a unified multi-architecture
index. Follow the exact TheRock multi-arch installation
and select gfx1151 through the device-gfx1151 package extras.
Do not add --pre by default. The index already publishes ROCm development
builds behind stable-looking PyTorch versions; --pre may select a newer
PyTorch alpha. Always record the resolved versions and retain the environment
until its replacement passes the same capability checks.
Start with no Strix-specific environment overrides. Current packages already identify gfx1151, select visible devices, and choose BLAS/allocator defaults.
Set only the switch required by a tested feature:
export TORCH_ROCM_AOTRITON_ENABLE_EXPERIMENTAL=1
Do not set these globally:
HSA_OVERRIDE_GFX_VERSION: hides an architecture/package mismatch.PYTORCH_ROCM_ARCH: build-time/JIT target selection, not normal runtime setup.HSA_ENABLE_SDMA=0: disables DMA copies; use only to isolate a reproduced bug.ROCR_VISIBLE_DEVICES or HIP_VISIBLE_DEVICES: restrict devices only when asked.ROCBLAS_USE_HIPBLASLT=1: leave backend selection on automatic unless a
workload-specific comparison proves otherwise.HSA_CU_MASK, HSA_XNACK, HSA_FORCE_FINE_GRAIN_PCIE, and heap percentage
overrides: diagnostic controls, not baseline optimizations.See performance features for attention,
torch.compile, BLAS, convolution, and dtype guidance.
Strix Halo has unified physical memory. Linux exposes overlapping VRAM and GTT accounting views; never add them together or describe their sum as usable RAM.
Before changing memory configuration:
amd-ttm helper over hand-written kernel parameters.Run the read-only advisor:
~/.claude/skills/strix-halo-setup/scripts/configure_gtt.sh
See GTT memory configuration. Do not claim a model is supported from a synthetic allocation; run a representative inference or training step with the intended precision and context length.
verify_system.sh checks host prerequisites without modifying the machine.verify_pytorch.py launches bounded real kernels for FP32, FP16, BF16,
matrix multiplication, MIOpen convolution/backward, SDPA, forced flash SDPA,
torch.compile, and a touched allocation.--require must pass or the script exits nonzero.Use TROUBLESHOOTING.md for symptom-specific fixes.