Use this skill when creating comprehensive testing strategies for applications...
This skill provides comprehensive guidance for building effective testing strategies that ensure software quality, reliability, and maintainability. Whether starting from scratch or improving existing test coverage, this framework helps teams design robust testing approaches.
When to use this skill:
Bundled Resources:
references/code-examples.md - Detailed testing code examplestemplates/test-plan-template.md - Comprehensive test plan templatetemplates/test-case-template.md - Test case documentation templatechecklists/test-coverage-checklist.md - Coverage verification checklistThis skill references the following testing tools. Not all are required - the skill will recommend appropriate tools based on your project.
Jest: Most popular testing framework
npm install --save-dev jest @types/jestnpx jest --initVitest: Vite-native testing framework
npm install --save-dev vitestMSW (Mock Service Worker): Network-level API mocking (2025 STANDARD)
npm install --save-dev mswnpx msw init public/ --savePlaywright: End-to-end testing
npm install --save-dev @playwright/testnpx playwright installk6: Performance testing
brew install k6k6 run script.jspytest: Standard Python testing framework
pip install pytestpytestpytest-cov: Coverage reporting
pip install pytest-covpytest --cov=.pytest-vcr / VCR.py: HTTP recording/playback (2025 STANDARD)
pip install pytest-vcr vcrpyconftest.pyLocust: Performance testing
pip install locustlocust -f locustfile.pyc8: JavaScript/TypeScript coverage
npm install --save-dev c8c8 npm testIstanbul/nyc: Alternative JS coverage
npm install --save-dev nycnyc npm test# JavaScript/TypeScript
jest --version
vitest --version
playwright --version
k6 version
# Python
pytest --version
locust --version
# Coverage
c8 --version
nyc --version
Note: The skill will guide you to select tools based on your project framework (React, Vue, FastAPI, Django, etc.) and testing needs.
Modern testing follows the "Testing Trophy" model (evolved from the testing pyramid):
š
/ \
/ E2E \ ā Few (critical user journeys)
/----------\
/ Integration\ ā Many (component interactions)
/--------------\
/ Unit \ ā Most (business logic)
/------------------\
/ Static Analysis \ ā Foundation (linting, type checking)
Principles:
Balance: 70% integration, 20% unit, 10% E2E (adjust based on context)
Recommended Targets:
Coverage Types:
Important: Coverage is a metric, not a goal. 100% coverage ā bug-free code.
Purpose: Catch errors before runtime Tools: ESLint, Prettier, TypeScript, Pylint, mypy, Ruff When to run: Pre-commit hooks, CI pipeline
Purpose: Test isolated business logic Tools: Jest, Vitest, pytest, JUnit Characteristics:
Coverage Target: 90%+ for business logic
See references/code-examples.md for detailed unit test examples.
Purpose: Test component interactions Tools: Testing Library, Supertest, pytest with fixtures Characteristics:
Coverage Target: 70%+ for API endpoints and component interactions
See references/code-examples.md for API integration test examples.
Purpose: Validate critical user journeys Tools: Playwright, Cypress, Selenium Characteristics:
Coverage Target: 5-10 critical user journeys
See references/code-examples.md for complete E2E test examples.
Purpose: Validate system performance under load Tools: k6, Artillery, JMeter, Locust Types:
Coverage Target: Test all performance-critical endpoints
See references/code-examples.md for k6 load test examples.
Prioritize testing based on risk assessment:
High Risk (100% coverage required):
Medium Risk (80% coverage):
Low Risk (50% coverage):
Given-When-Then Pattern:
Given [initial context]
When [action occurs]
Then [expected outcome]
This pattern keeps tests clear and focused. See references/code-examples.md for implementation examples.
Strategies:
See references/code-examples.md for test factory and fixture examples.
Structure tests in three clear phases:
See references/code-examples.md for detailed AAA pattern examples.
Each test should be independent:
See references/code-examples.md for test isolation patterns.
When to Mock:
When to Use Real Dependencies:
See references/code-examples.md for mocking examples.
MSW is the industry-standard approach for API mocking in frontend tests (Dec 2025).
MSW intercepts requests at the network level, not by mocking implementation details. This provides several advantages:
// src/mocks/handlers.ts
import { http, HttpResponse } from 'msw'
export const handlers = [
// Success response
http.get('/api/v1/analyze/:id', ({ params }) => {
return HttpResponse.json({
id: params.id,
status: 'completed',
createdAt: '2025-12-25T00:00:00Z',
})
}),
// Error response
http.get('/api/v1/analyze/error', () => {
return HttpResponse.json(
{ error: 'Analysis not found' },
{ status: 404 }
)
}),
// Delayed response (simulates slow network)
http.get('/api/v1/slow', async () => {
await delay(2000) // 2 second delay
return HttpResponse.json({ data: 'slow response' })
}),
]
// src/mocks/server.ts (for Vitest/Jest - Node.js)
import { setupServer } from 'msw/node'
import { handlers } from './handlers'
export const server = setupServer(...handlers)
// src/mocks/browser.ts (for Storybook/browser tests)
import { setupWorker } from 'msw/browser'
import { handlers } from './handlers'
export const worker = setupWorker(...handlers)
// vitest.setup.ts
import { beforeAll, afterEach, afterAll } from 'vitest'
import { server } from './src/mocks/server'
// Start server before all tests
beforeAll(() => server.listen({ onUnhandledRequest: 'error' }))
// Reset handlers after each test (removes runtime overrides)
afterEach(() => server.resetHandlers())
// Close server after all tests
afterAll(() => server.close())
import { http, HttpResponse } from 'msw'
import { server } from '../mocks/server'
test('shows error message when API fails', async () => {
// Override for this specific test
server.use(
http.get('/api/v1/analyze/:id', () => {
return HttpResponse.json(
{ error: 'Server error' },
{ status: 500 }
)
})
)
render(<AnalysisView id="123" />)
expect(await screen.findByText('Server error')).toBeInTheDocument()
})
test('shows loading state while fetching', async () => {
// Delay response to test loading state
server.use(
http.get('/api/v1/analyze/:id', async () => {
await delay(100)
return HttpResponse.json({ id: '123', status: 'pending' })
})
)
render(<AnalysisView id="123" />)
// Loading skeleton should be visible
expect(screen.getByTestId('skeleton')).toBeInTheDocument()
// Then data appears
expect(await screen.findByText('pending')).toBeInTheDocument()
})
import { z } from 'zod'
import { http, HttpResponse } from 'msw'
import { server } from '../mocks/server'
const AnalysisSchema = z.object({
id: z.string().uuid(),
status: z.enum(['pending', 'running', 'completed', 'failed']),
})
test('handles invalid API response gracefully', async () => {
// Return malformed data
server.use(
http.get('/api/v1/analyze/:id', () => {
return HttpResponse.json({
id: 'not-a-uuid', // Invalid!
status: 'unknown', // Invalid enum!
})
})
)
render(<AnalysisView id="123" />)
// Should show validation error, not crash
expect(await screen.findByText(/validation error/i)).toBeInTheDocument()
})
// ā NEVER mock fetch/axios directly
jest.spyOn(global, 'fetch').mockResolvedValue(...) // BAD!
jest.mock('axios') // BAD!
// ā NEVER mock your API service module
jest.mock('../services/api') // BAD!
// ā NEVER test implementation details
expect(fetch).toHaveBeenCalledWith('/api/...') // BAD!
// ā
ALWAYS use MSW handlers
import { http, HttpResponse } from 'msw'
server.use(http.get('/api/...', () => HttpResponse.json({...})))
// ā
ALWAYS test user-visible behavior
expect(await screen.findByText('Success')).toBeInTheDocument()
import { render, screen, waitFor } from '@testing-library/react'
import userEvent from '@testing-library/user-event'
import { http, HttpResponse } from 'msw'
import { server } from '../mocks/server'
import { QueryClient, QueryClientProvider } from '@tanstack/react-query'
function renderWithProviders(component: React.ReactNode) {
const queryClient = new QueryClient({
defaultOptions: {
queries: { retry: false },
},
})
return render(
<QueryClientProvider client={queryClient}>
{component}
</QueryClientProvider>
)
}
describe('AnalysisForm', () => {
test('submits analysis and shows result', async () => {
const user = userEvent.setup()
// Mock the POST endpoint
server.use(
http.post('/api/v1/analyze', async ({ request }) => {
const body = await request.json()
return HttpResponse.json({
analysis_id: 'new-123',
url: body.url,
status: 'pending',
})
})
)
renderWithProviders(<AnalysisForm />)
// Fill form
await user.type(screen.getByLabelText('URL'), 'https://example.com')
await user.click(screen.getByRole('button', { name: /analyze/i }))
// Verify result
expect(await screen.findByText(/analysis started/i)).toBeInTheDocument()
expect(screen.getByText('new-123')).toBeInTheDocument()
})
test('shows validation error on invalid URL', async () => {
const user = userEvent.setup()
server.use(
http.post('/api/v1/analyze', () => {
return HttpResponse.json(
{ detail: 'Invalid URL format' },
{ status: 422 }
)
})
)
renderWithProviders(<AnalysisForm />)
await user.type(screen.getByLabelText('URL'), 'not-a-url')
await user.click(screen.getByRole('button', { name: /analyze/i }))
expect(await screen.findByText(/invalid url/i)).toBeInTheDocument()
})
})
### 4. VCR.py - Python HTTP Recording (2025 Standard)
**VCR.py is the gold standard for testing Python code that makes HTTP requests.** It records real HTTP interactions once, then replays them for deterministic tests.
#### Why VCR.py?
| Approach | Problem |
|----------|---------|
| Mocking `requests` | Couples tests to implementation details |
| Live HTTP calls | Slow, flaky, rate-limited, non-deterministic |
| Manual fixtures | Tedious to maintain, drift from reality |
| **VCR.py** | ā
Records real responses, replays deterministically |
#### Basic Setup
```python
# conftest.py
import pytest
import vcr
# Configure VCR globally
@pytest.fixture(scope="module")
def vcr_config():
return {
"cassette_library_dir": "tests/cassettes",
"record_mode": "once", # Record once, then replay
"match_on": ["uri", "method"],
"filter_headers": ["authorization", "x-api-key"], # Security!
"filter_query_parameters": ["api_key", "token"],
}
# Alternative: pytest-vcr fixture decorator
@pytest.fixture
def vcr_cassette_dir(request):
return f"tests/cassettes/{request.module.__name__}"
import pytest
import vcr
# Method 1: Context manager
def test_fetch_user_data():
with vcr.use_cassette("tests/cassettes/user_data.yaml"):
response = requests.get("https://api.example.com/users/1")
assert response.status_code == 200
assert response.json()["name"] == "John Doe"
# Method 2: pytest-vcr decorator (recommended)
@pytest.mark.vcr()
def test_fetch_user_data_decorator():
response = requests.get("https://api.example.com/users/1")
assert response.status_code == 200
assert response.json()["name"] == "John Doe"
# Method 3: Custom cassette name
@pytest.mark.vcr("custom_cassette_name.yaml")
def test_with_custom_cassette():
response = requests.get("https://api.example.com/users/1")
assert response.status_code == 200
import pytest
import vcr
from httpx import AsyncClient
# VCR.py works with async HTTP clients
@pytest.mark.asyncio
@pytest.mark.vcr()
async def test_async_api_call():
async with AsyncClient() as client:
response = await client.get("https://api.example.com/data")
assert response.status_code == 200
assert "items" in response.json()
# conftest.py - configure per environment
@pytest.fixture(scope="module")
def vcr_config():
import os
# CI: never record, only replay
if os.environ.get("CI"):
record_mode = "none"
# Dev: record new, keep existing
else:
record_mode = "new_episodes"
return {
"record_mode": record_mode,
"cassette_library_dir": "tests/cassettes",
}
| Mode | Behavior | Use Case |
|---|---|---|
once |
Record if cassette missing, then replay | Default for most tests |
new_episodes |
Record new requests, replay existing | Adding to existing tests |
none |
Never record, fail on new requests | CI environments |
all |
Always record (overwrites) | Refreshing stale cassettes |
# conftest.py
@pytest.fixture(scope="module")
def vcr_config():
return {
# Remove headers before recording
"filter_headers": [
"authorization",
"x-api-key",
"cookie",
"set-cookie",
],
# Remove query parameters
"filter_query_parameters": [
"api_key",
"access_token",
"client_secret",
],
# Custom body filter
"before_record_request": filter_request_body,
"before_record_response": filter_response_body,
}
def filter_request_body(request):
"""Redact sensitive data from request body."""
if request.body:
import json
try:
body = json.loads(request.body)
if "password" in body:
body["password"] = "REDACTED"
if "api_key" in body:
body["api_key"] = "REDACTED"
request.body = json.dumps(body)
except json.JSONDecodeError:
pass
return request
def filter_response_body(response):
"""Redact sensitive data from response body."""
# Similar filtering logic
return response
# tests/services/test_tavily_service.py
import pytest
from app.services.external.tavily_service import TavilySearchService
@pytest.fixture
def tavily_service():
return TavilySearchService(api_key="test-key")
@pytest.mark.vcr()
async def test_tavily_search_returns_results(tavily_service):
"""Test Tavily search with recorded HTTP response."""
results = await tavily_service.search("Python async patterns")
assert len(results) > 0
assert all("url" in r for r in results)
assert all("content" in r for r in results)
@pytest.mark.vcr()
async def test_tavily_search_handles_empty_query(tavily_service):
"""Test graceful handling of empty search."""
results = await tavily_service.search("")
assert results == []
@pytest.mark.vcr()
async def test_tavily_rate_limit_error(tavily_service):
"""Test handling of rate limit response (cassette has 429)."""
with pytest.raises(RateLimitError):
await tavily_service.search("query that triggers rate limit")
# tests/cassettes/test_tavily_search_returns_results.yaml
interactions:
- request:
body: '{"query": "Python async patterns", "max_results": 10}'
headers:
Content-Type: application/json
# Note: authorization header filtered out
method: POST
uri: https://api.tavily.com/search
response:
body:
string: '{"results": [{"url": "https://...", "content": "..."}]}'
headers:
Content-Type: application/json
status:
code: 200
message: OK
version: 1
# tests/services/test_llm_service.py
import pytest
import vcr
# Custom matcher for LLM requests (ignore timestamp, request_id)
def llm_request_matcher(r1, r2):
"""Match LLM requests ignoring dynamic fields."""
import json
if r1.uri != r2.uri or r1.method != r2.method:
return False
body1 = json.loads(r1.body)
body2 = json.loads(r2.body)
# Ignore fields that change between runs
for field in ["request_id", "timestamp", "stream_id"]:
body1.pop(field, None)
body2.pop(field, None)
return body1 == body2
@pytest.fixture(scope="module")
def vcr_config():
return {
"cassette_library_dir": "tests/cassettes/llm",
"match_on": ["method", "uri"],
"custom_matchers": [llm_request_matcher],
"filter_headers": ["authorization", "x-api-key"],
}
@pytest.mark.vcr()
async def test_llm_completion():
"""Test LLM completion with recorded response."""
response = await llm_client.complete(
model="claude-3-5-sonnet",
messages=[{"role": "user", "content": "Say hello"}]
)
assert response.content is not None
assert "hello" in response.content.lower()
# ā NEVER: Commit cassettes with real API keys
# Bad cassette file:
# headers:
# authorization: Bearer sk-real-api-key-12345
# ā NEVER: Use "all" mode in CI
# record_mode: "all" # Will try to make real HTTP calls!
# ā NEVER: Skip VCR for "simple" HTTP tests
def test_api_call():
# This will make REAL HTTP calls in tests!
response = requests.get("https://api.example.com/data")
# ā
ALWAYS: Filter sensitive data
# ā
ALWAYS: Use "none" mode in CI
# ā
ALWAYS: Wrap all HTTP tests with VCR
# Delete old cassette to re-record
rm tests/cassettes/test_tavily_search_returns_results.yaml
# Run test to record fresh response
pytest tests/services/test_tavily_service.py::test_tavily_search_returns_results -v
# Or use environment variable to force re-record
VCR_RECORD_MODE=all pytest tests/services/ -v
Use for: UI components, API responses, generated code
Warning: Snapshots can become brittle. Use for stable components, not rapidly changing UI.
Test multiple scenarios with same logic using data tables.
See references/code-examples.md for parameterized test patterns.
Pipeline Stages:
# Example: GitHub Actions
name: Test Pipeline
on: [push, pull_request]
jobs:
test:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v3
- name: Install dependencies
run: npm ci
- name: Lint
run: npm run lint
- name: Type check
run: npm run typecheck
- name: Unit & Integration Tests
run: npm test -- --coverage
- name: Upload coverage
uses: codecov/codecov-action@v3
- name: E2E Tests
run: npm run test:e2e
- name: Performance Tests (on main branch)
if: github.ref == 'refs/heads/main'
run: npm run test:performance
Block merges/deployments if:
On Every Commit:
On Pull Request:
On Deploy to Staging:
On Deploy to Production:
| Category | Tool | Use Case |
|---|---|---|
| Unit/Integration | Vitest | Fast, Vite-native, modern |
| Unit/Integration | Jest | Mature, extensive ecosystem |
| E2E | Playwright | Cross-browser, reliable, fast |
| E2E | Cypress | Developer-friendly, visual debugging |
| Component Testing | Testing Library | User-centric, framework-agnostic |
| API Testing | Supertest | HTTP assertions, Express integration |
| Performance | k6 | Load testing, scriptable |
| Category | Tool | Use Case |
|---|---|---|
| Unit/Integration | pytest | Powerful, extensible, fixtures |
| API Testing | httpx + pytest | Async support, modern |
| E2E | Playwright (Python) | Browser automation |
| Performance | Locust | Load testing, Python-based |
| Mocking | unittest.mock | Standard library, reliable |
ā Testing Implementation Details
// Bad: Testing internal state
expect(component.state.isLoading).toBe(false);
// Good: Testing user-visible behavior
expect(screen.queryByText('Loading...')).not.toBeInTheDocument();
ā Tests Too Coupled to Code
// Bad: Test breaks when implementation changes
expect(userService.save).toHaveBeenCalledTimes(1);
// Good: Test behavior, not implementation
const user = await db.users.findOne({ email: 'test@example.com' });
expect(user).toBeTruthy();
ā Direct Fetch Mocking (2025 Anti-Pattern)
// Bad: Mocks implementation, not network behavior
jest.spyOn(global, 'fetch').mockResolvedValue({
json: () => Promise.resolve({ data: 'mocked' })
});
jest.mock('axios');
jest.mock('../services/api');
// Good: Use MSW for network-level mocking
import { http, HttpResponse } from 'msw';
server.use(
http.get('/api/data', () => HttpResponse.json({ data: 'mocked' }))
);
ā Flaky Tests
// Bad: Non-deterministic timeout
await waitFor(() => {
expect(screen.getByText('Success')).toBeInTheDocument();
}, { timeout: 1000 }); // Might fail on slow CI
// Good: Use explicit waits with longer timeout
await screen.findByText('Success', {}, { timeout: 5000 });
ā Giant Test Cases
// Bad: One test does too much
test('user workflow', async () => {
// 100 lines testing signup, login, profile update, logout...
});
// Good: Focused tests
test('user can sign up', async () => { /* ... */ });
test('user can login', async () => { /* ... */ });
test('user can update profile', async () => { /* ... */ });
When starting a new project or feature:
templates/test-plan-template.md)For detailed code examples: See references/code-examples.md
Testing AI applications requires specialized approaches due to their probabilistic nature.
import pytest
import asyncio
@pytest.mark.asyncio
async def test_operation_respects_timeout():
"""Test that async operations honor timeout limits."""
async def slow_operation():
await asyncio.sleep(10) # Simulates slow LLM call
return "result"
with pytest.raises(asyncio.TimeoutError):
async with asyncio.timeout(0.1):
await slow_operation()
@pytest.mark.asyncio
async def test_graceful_degradation_on_timeout():
"""Test fail-open behavior when operation times out."""
result = await safe_operation_with_fallback(timeout=0.1)
assert result["status"] == "fallback"
assert result["error"] == "Operation timed out"
from unittest.mock import AsyncMock, patch
@pytest.fixture
def mock_llm_response():
"""Mock LLM to return predictable structured output."""
mock = AsyncMock()
mock.return_value = {
"content": "Mocked response",
"confidence": 0.85,
"tokens_used": 150
}
return mock
@pytest.mark.asyncio
async def test_synthesis_with_mocked_llm(mock_llm_response):
"""Test synthesis logic without actual LLM calls."""
with patch("app.core.model_factory.get_model", return_value=mock_llm_response):
result = await synthesize_findings(sample_findings)
assert result["executive_summary"] is not None
assert mock_llm_response.call_count == 1
import pytest
from pydantic import ValidationError
def test_quiz_question_validates_correct_answer():
"""Test that correct_answer must be in options."""
with pytest.raises(ValidationError) as exc_info:
QuizQuestion(
question="What is 2+2?",
options=["3", "4", "5"],
correct_answer="6", # Not in options!
explanation="Basic arithmetic"
)
assert "correct_answer" in str(exc_info.value)
assert "must be one of" in str(exc_info.value)
def test_quiz_question_accepts_valid_answer():
"""Test that valid answers pass validation."""
q = QuizQuestion(
question="What is 2+2?",
options=["3", "4", "5"],
correct_answer="4", # Valid!
explanation="Basic arithmetic"
)
assert q.correct_answer == "4"
from jinja2 import Environment, FileSystemLoader
@pytest.fixture
def jinja_env():
return Environment(loader=FileSystemLoader("templates/"))
def test_template_handles_empty_tldr(jinja_env):
"""Template renders without crashing when tldr is empty."""
template = jinja_env.get_template("artifact.j2")
result = template.render(aggregated_insights={"tldr": {}})
assert "TL;DR" not in result # Section skipped gracefully
def test_template_handles_missing_nested_field(jinja_env):
"""Template handles None in nested objects."""
template = jinja_env.get_template("artifact.j2")
result = template.render(aggregated_insights={
"tldr": {"summary": None, "key_takeaways": []}
})
# Should not crash, should handle gracefully
assert isinstance(result, str)
@pytest.mark.asyncio
async def test_quality_evaluator_returns_normalized_score():
"""Quality scores should be normalized 0.0-1.0."""
evaluator = create_quality_evaluator("relevance")
# Mock the LLM to return a score
with patch_evaluator_llm(return_score=8): # 8/10
result = await evaluator.aevaluate_strings(
input="Test input",
prediction="Test output"
)
assert 0.0 <= result["score"] <= 1.0
assert result["score"] == 0.8 # 8/10 normalized
@pytest.mark.asyncio
async def test_quality_gate_fails_below_threshold():
"""Quality gate should fail when avg score < threshold."""
with patch_quality_scores({"relevance": 0.5, "depth": 0.4, "coherence": 0.5}):
result = await quality_gate_node(sample_state)
assert result["quality_gate_passed"] is False
assert result["quality_gate_avg_score"] < 0.7
When testing LLM integrations, always test these edge cases:
Skill Version: 1.3.0 Last Updated: 2025-12-27 Maintained by: AI Agent Hub Team