Parallel SSH operations across multiple hosts using google_compute_engine key...
This skill enables parallel SSH operations across multiple hosts for distributed workloads like SGLang/vLLM prefill-decode disaggregation.
~/.ssh/google_compute_engine
Before running any SSH commands, ALWAYS check if the key file exists. If not, generate it using gcloud:
# Check if key exists
if [ ! -f ~/.ssh/google_compute_engine ]; then
echo "SSH key not found, generating via gcloud..."
gcloud compute config-ssh --quiet
fi
This command:
~/.ssh/google_compute_engine (private key) and ~/.ssh/google_compute_engine.pub (public key)ssh -i ~/.ssh/google_compute_engine -o StrictHostKeyChecking=accept-new <IP> "<command>"
The -o StrictHostKeyChecking=accept-new option automatically accepts new host keys (safe for first connection).
When user provides hosts like:
Parse into individual IPs.
For each host, launch SSH command with run_in_background: true:
ssh -i ~/.ssh/google_compute_engine -o StrictHostKeyChecking=accept-new 10.8.0.81 "command"
IMPORTANT: Launch ALL tasks in a SINGLE message with multiple Bash tool calls to achieve true parallelism.
Use TaskOutput to wait for completion, or Read to check progress:
/tmp/claude-*/tasks/<task_id>.outputnvidia-smi --query-gpu=index,name,memory.used,memory.total,utilization.gpu --format=csv
vmstat 2 15
systemctl status <service> || pgrep -a <process>
source /opt/deepep/unified-env.sh && \
python3 -m sglang.launch_server \
--model-path deepseek-ai/DeepSeek-V3 \
--disaggregation-mode prefill \
--tp-size 8 \
--port 30000 \
--host 0.0.0.0 \
...
source /opt/deepep/unified-env.sh && \
python3 -m sglang.launch_server \
--model-path deepseek-ai/DeepSeek-V3 \
--disaggregation-mode decode \
--tp-size 8 \
--port 30001 \
--host 0.0.0.0 \
...
User: "Check GPU status on 10.8.0.81 10.8.0.82 10.8.0.83"
Action: Launch 3 parallel SSH tasks with nvidia-smi command, collect and summarize results.
User: "Start prefill on 10.8.0.81, decode on 10.8.0.82 and 10.8.0.83"
Action:
User: "Kill all python processes on all machines"
Action: Launch parallel pkill -f python commands on all hosts.
User: "Get last 100 lines of sglang logs from all machines"
Action: Launch parallel tail -100 /path/to/logs commands.
If SSH fails with "Host key verification failed":
# Add host key to known_hosts
ssh-keyscan -H <IP> >> ~/.ssh/known_hosts
# Or use StrictHostKeyChecking=accept-new (recommended)
ssh -o StrictHostKeyChecking=accept-new ...
When reporting results, use table format:
| Host | Status | Key Metrics |
|---|---|---|
| 10.8.0.81 | OK | GPU 0-7: 95% util |
| 10.8.0.82 | OK | GPU 0-7: 80% util |
| 10.8.0.83 | FAILED | SSH timeout |
Common environment variables to set on each node:
export CUDA_VISIBLE_DEVICES=0,1,2,3,4,5,6,7
export NCCL_SOCKET_IFNAME=enp0s19
export GLOO_SOCKET_IFNAME=enp0s19
export MASTER_ADDR=<prefill_node_ip>
export MASTER_PORT=29500
If you see "Identity file not accessible" or "No such file or directory":
# Generate SSH key using gcloud (recommended for GCE)
gcloud compute config-ssh --quiet
# This creates ~/.ssh/google_compute_engine and uploads public key to project metadata
ping <IP>ls -la ~/.ssh/google_compute_engine
gcloud compute config-ssh --quiet
gcloud compute project-info describe --format="value(commonInstanceMetadata.items.filter(key:ssh-keys))"
nohup for long-running processesscreen or tmux for persistent sessions