Translate PowerPoint presentations while preserving formatting (fonts, colors, alignment, tables). Supports multiple LLM providers (OpenAI, Anthropic, DeepSeek, Grok, Gemini)...
Translate PowerPoint presentations while preserving all formatting including fonts, colors, spacing, tables, and alignment.
.pptx files between languages--vision-audit) to catch text overflow, truncation, garbled glyphs, untranslated leftovers, overlaps, or contrast problemsBefore first use, set up the environment:
# Navigate to the scripts directory
cd .claude/skills/ppt-translator/scripts
# Create virtual environment and install dependencies
python3 -m venv .venv
source .venv/bin/activate # macOS/Linux
pip install -r requirements.txt
# Configure API keys
cp example.env .env
# Edit .env with your provider API key(s)
cd .claude/skills/ppt-translator/scripts
source .venv/bin/activate
python main.py /path/to/presentation.pptx \
--source-lang zh \
--target-lang en
(The default provider is DeepSeek deepseek-v4-flash.)
| Provider | Environment Variable | Default Model |
|---|---|---|
| deepseek | DEEPSEEK_API_KEY |
deepseek-v4-flash |
| openai | OPENAI_API_KEY |
gpt-5.2-2025-12-11 |
| anthropic | ANTHROPIC_API_KEY |
claude-sonnet-4-5-20250514 |
| grok | GROK_API_KEY |
grok-4.1-fast |
| gemini | GEMINI_API_KEY |
gemini-3-flash-preview |
| Option | Description | Default |
|---|---|---|
--provider |
LLM provider: deepseek, openai, anthropic, grok, gemini |
deepseek |
--model |
Override default model for provider | Provider default |
--source-lang |
Source language ISO code | zh |
--target-lang |
Target language ISO code | en |
--max-chunk-size |
Characters per API request | 1000 |
--max-workers |
Threads for slide extraction | 4 |
--keep-intermediate |
Retain XML files for debugging | false |
--vision-audit |
After translation, render every rebuilt slide to an image and inspect it with DeepSeek's vision model for overflow, truncation, garbled glyphs, untranslated text, overlaps, or contrast issues. Requires DEEPSEEK_API_KEY and LibreOffice + poppler (pdftoppm) installed. Writes {deck}_translated_vision_audit.md. |
false |
--vision-model |
Vision model used by --vision-audit |
deepseek-v4-flash-vision-exp |
--vision-dpi |
Rendering resolution for audit images | 100 |
For each input presentation.pptx, the tool generates:
presentation_original.xml - Extracted source content (deleted unless --keep-intermediate)presentation_translated.xml - Translated content (deleted unless --keep-intermediate)presentation_translated.pptx - Final translated presentation with formatting intactpresentation_translated_vision_audit.md - Only with --vision-audit: per-slide findings from the vision model, grouped by slide with severity (high/medium/low)python main.py deck.pptx --provider anthropic --source-lang zh --target-lang en
python main.py /path/to/presentations/ --provider openai --source-lang ja --target-lang en
python main.py deck.pptx --keep-intermediate --provider deepseek
# Inspect the generated XML files to see extracted/translated content
python main.py deck.pptx --provider openai --model gpt-5-mini
python main.py deck.pptx --provider gemini --source-lang ko --target-lang en
python main.py deck.pptx --source-lang zh --target-lang en --vision-audit
After rebuilding the deck, every slide is rendered to a PNG (LibreOffice headless → pdftoppm) and sent to deepseek-v4-flash-vision-exp together with the slide's expected translated text. The model compares the rendering against the expected content and reports concrete visual defects. Read {deck}_translated_vision_audit.md afterwards and fix any HIGH-severity findings (e.g. by shortening text or enlarging boxes) before shipping the deck.
Audit-only reruns on an already-translated deck are not needed — the audit runs automatically at the end of a --vision-audit run, and a failed audit never aborts the translation itself.
Use standard ISO 639-1 language codes:
| Code | Language |
|---|---|
zh |
Chinese (Simplified) |
en |
English |
ja |
Japanese |
ko |
Korean |
es |
Spanish |
fr |
French |
de |
German |
pt |
Portuguese |
ru |
Russian |
ar |
Arabic |
The tool automatically scales fonts down (70% for text, 80% for tables) to accommodate text expansion when translating from compact languages (Chinese, Japanese, Korean) to English. This prevents text overflow in fixed-size text boxes.
Repeated strings within a presentation are cached to avoid redundant API calls. This is especially useful for presentations with recurring headers, footers, or terminology.
Long text blocks are intelligently split at sentence boundaries to stay within API limits while preserving translation quality.
Ensure your .env file in the scripts directory contains the correct environment variable:
OPENAI_API_KEY=sk-...
ANTHROPIC_API_KEY=sk-ant-...
--keep-intermediate to inspect the XML files--max-chunk-size for very long text blocksThe visual audit needs LibreOffice (soffice) and poppler (pdftoppm):
brew install --cask libreoffice && brew install poppler # macOS
# apt install libreoffice poppler-utils # Debian/Ubuntu
If LibreOffice itself crashes on launch (seen on some macOS 26 + LibreOffice 25.x combos), upgrade it: brew upgrade --cask libreoffice.
The audit always calls DeepSeek's vision endpoint regardless of the translation provider, so DEEPSEEK_API_KEY must be set even when translating with another provider.
The scripts/ directory contains:
main.py - Entry pointrequirements.txt - Python dependenciesexample.env - Environment variable templateppt_translator/ - Core translation modulecli.py - CLI argument parsingpipeline.py - PPT extraction and regenerationtranslation.py - Chunking and cachingproviders/ - LLM provider implementationsvision_audit.py - Slide rendering (LibreOffice + pdftoppm) and DeepSeek vision-model QA