Improve embedded LLM system prompt based on evaluation test failures
Iteratively improve the embedded backend's system prompt to increase command generation accuracy.
Run evaluation tests with embedded backend:
./target/release/caro test --backend embedded
Record:
For each failed test case, identify the pattern:
| Pattern | Example | Fix |
|---|---|---|
| Wrong path | find / instead of find . |
Add rule: "ALWAYS use current directory '.'" |
| GNU flags | --max-depth on macOS |
Add rule: "Use BSD-compatible flags" |
| Missing filters | No -name "*.py" |
Add rule: "Include ALL relevant filters" |
| Time semantics | -mtime -1 vs -mtime 1 |
Add clear mtime documentation |
| Quote style | Single vs double quotes | Usually equivalent, low priority |
| Flag order | -type f -name vs -name -type f |
Usually equivalent, low priority |
Edit the system prompt in:
src/backends/embedded/embedded_backend.rs
Function: create_system_prompt()
Improvement strategies:
Build and re-run tests:
cargo build --release
./target/release/caro test --backend embedded
Compare results:
If accuracy improved significantly:
git add src/backends/embedded/embedded_backend.rs
git commit -m "feat(prompt): Improve embedded backend accuracy from X% to Y%"
If not improved or regressed:
| Level | Accuracy | Action |
|---|---|---|
| Poor | < 50% | Major prompt rewrite needed |
| Acceptable | 50-70% | Targeted improvements |
| Good | 70-85% | Minor tuning |
| Excellent | > 85% | Consider semantic equivalence in remaining failures |
User: /prompt-tuner