Use when working on long-running projects or needing context across sessions. Covers memory architecture, privacy controls, efficient retrieval, and integration with claude-mem plugin...
Memory is context that survives sessions. Use it strategicallyβstore what matters, retrieve efficiently, protect what's private.
MEMORY IS NOT FREE. RETRIEVE SELECTIVELY.
Every token of context costs. Don't dump entire historyβsearch and fetch only what's relevant.
Benefits:
Without persistent memory:
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β MEMORY SYSTEM β
β β
β ββββββββββββββββ ββββββββββββββββ βββββββββββββ β
β β Capture βββββΆβ Store βββββΆβ Retrieve β β
β β (Hooks) β β (SQLite + β β (Search) β β
β β β β Vectors) β β β β
β ββββββββββββββββ ββββββββββββββββ βββββββββββββ β
β β β β β
β βΌ βΌ βΌ β
β ββββββββββββββββ ββββββββββββββββ βββββββββββββ β
β β Session β β Observations β β Summaries β β
β β Lifecycle β β & Context β β & Answers β β
β ββββββββββββββββ ββββββββββββββββ βββββββββββββ β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
# Install the plugin
/plugin marketplace add thedotmack/claude-mem
/plugin install claude-mem
# Restart Claude Code
# Memory automatically starts capturing
# Check the web UI
open http://localhost:37777
# Memory service should be running on port 37777
The plugin captures automatically via lifecycle hooks:
SessionStart - New session beginsUserPromptSubmit - User sends messagePostToolUse - Tool execution completeStop - Task completedSessionEnd - Session endsDO store:
DON'T store:
Use <private> tags to exclude sensitive content:
Here's the database config:
<private>
DATABASE_URL=postgres://user:password@host:5432/db
API_KEY=sk-secret-key-here
</private>
The schema uses PostgreSQL with these tables...
Content within <private> tags is NOT stored in memory.
Memory retrieval uses a three-layer workflow:
Layer 1: Search Index
βββ Quick keyword/semantic match
βββ Returns: observation IDs + snippets
βββ Cost: ~100 tokens
Layer 2: Timeline Context
βββ Fetch surrounding observations
βββ Returns: chronological context
βββ Cost: ~500 tokens
Layer 3: Full Details
βββ Fetch complete observations
βββ Returns: full content
βββ Cost: varies (can be large)
Best practice: Start with Layer 1, only go deeper if needed.
# Good queries (specific)
"How did we implement authentication?"
"What was the decision on database schema?"
"Why did we choose React over Vue?"
# Poor queries (too broad)
"What did we do?"
"Show me everything"
"All past work"
Periodically review and clean:
# Access web UI to browse memory
open http://localhost:37777
# Review recent sessions
# Delete irrelevant observations
# Check storage usage
# At session start
"Check memory for our previous work on the authentication feature.
What was our approach and what's left to do?"
# Memory retrieves relevant context automatically
# When making architectural decisions
"Search memory for past discussions about state management.
What approaches did we consider and why did we choose Redux?"
# When debugging recurring issues
"Search memory for similar errors we've seen before.
How did we fix the authentication timeout issue last time?"
# When starting on existing project
"Retrieve memory summaries for this project.
What are the key architectural decisions and conventions?"
1. Claude automatically loads relevant memory
2. Review injected context (check token cost)
3. Ask clarifying questions if context seems incomplete
4. Proceed with task
1. Important decisions β explicitly note rationale
2. Sensitive data β use <private> tags
3. Complex solutions β document approach
4. Errors fixed β note root cause
1. Summarize what was accomplished
2. Note any pending work
3. Document blockers or questions
4. Memory auto-captures on session end
The web UI shows token costs for context injection:
| Context Type | Typical Tokens | When to Use |
|---|---|---|
| Session summary | 100-300 | Always (automatic) |
| Search results | 200-500 | Specific queries |
| Full observation | 500-2000 | Deep dive needed |
| Timeline context | 300-800 | Understanding sequence |
1. Use specific queries
β "What do you know?"
β
"What's our API rate limiting strategy?"
2. Progressive disclosure
- Start with summaries
- Drill down only if needed
- Don't fetch everything
3. Prune irrelevant results
- Skip old/outdated context
- Focus on recent sessions
- Filter by relevance
# Check if service is running
curl http://localhost:37777/health
# Verify plugin is installed
/plugin list
# Reinstall if needed
/plugin uninstall claude-mem
/plugin install claude-mem
1. Try different keywords
2. Use semantic queries (describe what you mean)
3. Check date range
4. Verify content wasn't marked <private>
1. Review what's being injected
2. Use more specific queries
3. Disable automatic context for trivial tasks
4. Clean old irrelevant observations
| Mistake | Impact | Fix |
|---|---|---|
| Storing secrets | Security risk | Use <private> tags |
| Retrieving everything | Token explosion | Query specifically |
| Ignoring memory | Repeated work | Check memory first |
| No privacy tags | Credentials exposed | Tag sensitive content |
| Stale memory | Wrong context | Periodically clean |
Use with:
rag-architecture - Memory as knowledge sourceagentic-design - Long-term agent memoryllm-integration - Context managementsubagent-driven-development - Cross-task contextEnhances:
brainstorming - Recall past ideascode-review - Remember past issueswriting-plans - Build on previous plansBefore relying on memory:
During sessions:
Based on:
Bottom Line: Memory extends context across sessions. Store decisions, protect secrets, retrieve selectively. Every token costsβbe strategic about what you recall.