Comprehensive expert knowledge for NVIDIA DGX Spark workstation...
Expert guidance for NVIDIA DGX Spark AI workstation development.
| Resource | URL |
|---|---|
| Playbooks Hub | https://build.nvidia.com/spark |
| User Guide | https://docs.nvidia.com/dgx/dgx-spark/ |
| Support | https://www.nvidia.com/en-us/support/dgx-spark/ |
| Forums | https://forums.developer.nvidia.com/c/accelerated-computing/dgx-spark-gb10 |
| GitHub Playbooks | https://github.com/NVIDIA/dgx-spark-playbooks |
Load these for detailed information:
references/hardware-specs.md — Full specs, UMA details, Spark stackingreferences/playbooks-index.md — All 25+ official playbooks with linksreferences/known-issues.md — Troubleshooting, diagnostics, supportreferences/software-stack.md — DGX OS, containers, frameworks, toolsreferences/ai-workbench.md — AI Workbench projects, RAG, agentsQuick local chat: Use Ollama + Open WebUI playbook
# Install Ollama
curl -fsSL https://ollama.com/install.sh | sh
ollama run llama3.2
Production serving: Use vLLM or TRT-LLM playbooks
# vLLM example
pip install vllm
python -m vllm.entrypoints.openai.api_server \
--model meta-llama/Llama-3.1-8B-Instruct
See references/playbooks-index.md for all options.
# Clone official RAG project
nvwb project clone https://github.com/NVIDIA/workbench-example-agentic-rag
# Start project
cd workbench-example-agentic-rag
nvwb start
See references/ai-workbench.md for detailed workflow.
references/known-issues.md firstsudo sync && echo 3 | sudo tee /proc/sys/vm/drop_cachesfree -h (not nvidia-smi for memory)sudo systemctl status nvidia-persistencedDGX Spark uses Unified Memory Architecture — CPU and GPU share 128GB.
Key points:
nvidia-smi memory display may show "Not Supported" (expected)free -h for actual memory statussudo sync && echo 3 | sudo tee /proc/sys/vm/drop_caches# Standard GPU container
docker run --gpus all --runtime nvidia <image>
# With shared memory (required for many ML frameworks)
docker run --gpus all --shm-size=16g <image>
# Mount HuggingFace cache
docker run --gpus all \
-v $HOME/.cache/huggingface:/root/.cache/huggingface \
<image>
# NGC container example
docker pull nvcr.io/nvidia/pytorch:24.01-py3
DGX Spark runs ARM64, not x86. When installing software:
aarch64 or arm64 package versions| Goal | Recommended Playbook |
|---|---|
| Chat with local LLM | Open WebUI + Ollama |
| Serve LLM API | vLLM or TRT-LLM |
| Fine-tune LLM | LLaMA Factory (quick) or NeMo (production) |
| Build RAG app | RAG in AI Workbench |
| Generate images | Comfy UI |
| AI coding assistant | Vibe Coding |
| Remote access | Tailscale |
| Data science | CUDA-X Data Science |
| Large models (>200B) | Connect Two Sparks |
User asks about inference → Check model size, recommend vLLM (high throughput) or TRT-LLM (optimized latency), or Ollama for simple use.
User asks about fine-tuning → Assess complexity: LLaMA Factory/Unsloth for experiments, NeMo for production, PyTorch for custom needs.
User asks about AI Workbench → Load references/ai-workbench.md, guide through project creation/cloning.
User reports error/issue → Load references/known-issues.md, check if known issue, provide diagnostic commands.
User asks about specs/capabilities → Load references/hardware-specs.md, provide relevant details.
User wants specific playbook → Load references/playbooks-index.md, provide direct link and summary.