Vitest testing patterns for Tzurot v3. Use when writing tests, debugging test failures, or mocking dependencies. Covers mock factories, fake timers, and promise rejection handling.
Invoke with /tzurot-testing for test-related procedures.
Testing patterns are in .claude/rules/02-code-standards.md - they apply automatically.
# Run all tests
pnpm test
# Run specific service
pnpm --filter @tzurot/ai-worker test
# Run specific file β never `test -- <file>`: pnpm forwards the literal `--` and
# vitest drops everything after it (full suite); a value-less flag such as `-u`
# goes AFTER the file or it swallows the path
pnpm --filter @tzurot/ai-worker test --run src/services/MyService.test.ts
# Run the component tier (PGLite) β NOT included in `pnpm test`
pnpm test:component
# Run the integration + contract tiers (real Redis/Postgres) β NOT in `pnpm test`
pnpm test:integration
# Run with coverage
pnpm test:coverage
# Run only changed packages
pnpm focus:test
Always invoke through the package's own config. A root-invoked npx vitest run <pkg path> resolves the ROOT config, which has no setupFiles β the package's setup (dotenv, the global afterEach(vi.clearAllMocks)) never loads, hoisted-mock call history accumulates across tests in a file, and expect(mock).not.toHaveBeenCalled() fails on calls leaked from earlier tests. It masquerades as flaky or pre-existing breakage; it is an invocation artifact. Use pnpm --filter @tzurot/<pkg> test, or cd services/<pkg> && npx vitest run <relative-path> β and never diagnose "pre-existing failures" from a root-invoked run.
# Run unified audit (CI does this automatically)
pnpm ops test:audit
# Filter by category
pnpm ops test:audit --category=services
pnpm ops test:audit --category=contracts
# Update baseline (after closing gaps)
pnpm ops test:audit --update
# Strict mode (fails on ANY gap)
pnpm ops test:audit --strict
Unified Baseline: test-coverage-baseline.json (project root)
Tiers are canonically defined in one place β
Test Tier Taxonomy.
This table is the suffixβtier quick reference; don't re-define the tiers here
(the pnpm ops guard:test-taxonomy gate enforces the single-source link).
| Suffix / location | Tier (taxonomy) | Infrastructure | Location |
|---|---|---|---|
*.test.ts |
Unit | Fully mocked | Next to source |
*.component.test.ts |
Component (one service, PGLite) | PGLite | Next to source |
*.integration.test.ts |
Integration (real DB+Redis) | Real services | tests/e2e/ |
*.contract.test.ts |
Contract (providerβconsumer) | Real services | tests/e2e/contracts/ |
Each suffix names its tier directly. Schema test β contract test: a Zod schema test (a plain
*.test.ts) validates one type's own rules (unit-tier); a contract test verifies two services agree. Runpnpm ops test:tiersfor the per-package tier distribution.
The unit pnpm test config EXCLUDES the .component / .integration /
.contract suffixes. A contract/integration/component test "passes" under
pnpm test by NOT running at all β so a green unit run is NOT evidence it works.
| Suffix | Verify with |
|---|---|
*.test.ts (unit) |
pnpm test |
*.component.test.ts |
pnpm test:component |
*.integration.test.ts / *.contract.test.ts |
pnpm test:integration |
When you touch a .contract / .integration test, run pnpm test:integration
(Redis + Postgres up: podman start tzurot-redis tzurot-postgres); a
.component test, pnpm test:component. pnpm test / pnpm --filter <pkg> test
is the unit tier ONLY. CI's component-integration-tests job runs test:component +
test:integration, so it catches what a unit-only local run silently skipped β
don't let CI be the first thing that actually executes your contract test.
Gotcha for mention/ID fixtures: tests need a valid 17β19 digit Discord
snowflake β isValidDiscordId silently drops toy ids like 555 before
resolution, making the assertion pass/fail for the wrong reason.
pnpm --filter @tzurot/ai-worker test --run src/services/MyService.test.ts --reporter=verbose
// β WRONG - Promise rejection warning
const promise = asyncFunction();
await vi.runAllTimersAsync(); // Rejection happens here!
await expect(promise).rejects.toThrow(); // Too late
// β
CORRECT - Attach handler BEFORE advancing
const promise = asyncFunction();
const assertion = expect(promise).rejects.toThrow('Error');
await vi.runAllTimersAsync();
await assertion;
beforeEach(() => {
vi.clearAllMocks(); // Clear call history, keep impl
});
afterEach(() => {
vi.restoreAllMocks(); // Restore originals (spies only)
});
// Use async factory for vi.mock hoisting
vi.mock('./MyService.js', async () => {
const { mockMyService } = await import('../test/mocks/MyService.mock.js');
return mockMyService;
});
// Import accessors after vi.mock
import { getMyServiceMock } from '../test/mocks/index.js';
it('should call service', () => {
expect(getMyServiceMock().someMethod).toHaveBeenCalled();
});
At the moment of writing a probe or test fixture β a filename, a payload, a
URL, a config value β do not invent a minimal one: copy a real specimen out
of the corpus and mutate only what the test needs. An input built to isolate
one property is exactly the input that cannot reveal an interaction with a
second property, and its green is indistinguishable from a green that covered
the real case (a space-free probe filename passed end-to-end while every real
tracker/ path β all space-bearing β would have been quote-mangled). If a
synthetic input is unavoidable, enumerate the properties real specimens carry
(spaces, non-ASCII, length, punctuation, nesting) and give it ALL of them.
The refutation direction inverts this: the probe that settles a disputed
finding varies a single property at a time.
Before asserting on a subject that sits behind a mocked collaborator, answer
two questions: (a) which mock makes this code path reachable, and (b) what
would a wiring bug at that seam look like in this test's output? If the
answer to (b) is "the test can't see it," the test asserts through the seam
it mocked β add the seam assertion per 02-code-standards.md Β§ Assert what
crosses a mocked seam before writing more cases. This is the coverage-illusion
class: a mock missing one property leaves the dependent branch green and
untested; a feature flag can no-op silently under a fully green suite.
Cover every render MODE, not just the default one. A branching renderer
(full vs. deduped, live vs. stored, enabled vs. disabled) leaves its
non-default arms as untested forwarding paths β enumerate the modes and
assert the seam in each. When the forwarded value cost money (vision,
transcription, an external fetch), mock the paid boundary to return a
sentinel and assert the sentinel reaches the final output β a field-parity
allowlist only catches the field you already know about. Reference: the
(deduped Γ full) Γ (live Γ stored) matrix in
services/ai-worker/src/services/ReferencedMessageFormatter.test.ts.
A shared mutable context is a seam too, with no mock to assert across.
Pipeline steps, middleware, and anything handing data to a later stage by
writing a field on a shared object (job.data.context, req, an
accumulating result): each step's unit tests construct that object
themselves, so neither can observe what the other produced. Whenever a step
writes a field a later step reads, keep ONE test that runs those steps
in order.
The specific tell is a default-coalescing write-back: x = ctx.field ?? [], ?? {}, || fallback written BACK onto the shared object erases the
distinction the later step reads. Where absence carries meaning, say so in a
comment at the write site and pin BOTH states in the sequencing test β
absent stays absent, empty stays empty (?? falls through on undefined,
never on []). Reference:
services/ai-worker/src/jobs/handlers/pipeline/steps/extendedContextVisionSeam.test.ts.
The RESPONSE direction crosses a Zod strip no mock can see. Typed clients return
outputSchema.safeParse(...).data, so a response key not declared in the wire schema
is deleted before the caller ever sees it β and a mocked client skips that parse. Pin
survival at the boundary: Schema.safeParse(payloadWithSentinel) β assert result.data still carries it.
The sweep must cover every test TIER, not just the ones a local
pnpm test runs β tests/e2e/ (integration + contract) pins producer
shapes too. Before pushing any cross-service string-shape change (job ids,
fixture fields, wire formats): (a) grep the OLD shape's distinctive tokens
repo-wide including tests/, and (b) run
npx vitest run --config vitest.integration.config.ts tests/e2e/contracts/
(~2s; the full-test:integration OOM ban does not cover this subset).
When asserting gate N in a multi-gate pipeline β a hook with several checks, a
validator chain, a new guard stacked in front of an older one β an exit code or
shared failure result is NOT attribution: the new gate inherits the older
gate's failure, so assert_exit 2 (or .toThrow(), or "returns false") passes
whether YOUR gate fired or the one behind it did. Two obligations:
02-code-standards.md). A canary
against the pipeline as a whole proves nothing about which gate you pinned.Anatomy to fear: a cluster of vacuous assertions shipped in one hook PR (#2078) with exactly this shape β including one whose test NAME was false, passing only because a different gate blocked first. Reading had already passed every one of them; per-gate canaries are what caught them.
describe('UserService', () => {
let pglite: PGlite;
let prisma: PrismaClient;
beforeAll(async () => {
pglite = new PGlite({ extensions: { vector, citext } });
await pglite.exec(loadPGliteSchema());
prisma = new PrismaClient({ adapter: new PrismaPGlite(pglite) });
});
it('should create user', async () => {
const service = new UserService(prisma);
const userId = await service.getOrCreateUser('123', 'testuser');
expect(userId).toBeDefined();
});
});
β οΈ ALWAYS use loadPGliteSchema() - NEVER create tables manually!
Component tests (*.component.test.ts) run separately from unit tests and are not included in pnpm test or pre-push hooks.
Always run pnpm test:component after:
| Change | Why |
|---|---|
| Add/remove slash command options | CommandHandler.component.test.ts snapshots capture full command structure |
| Add/remove subcommands | Same snapshot tests |
| Restructure command directories | getCommandFiles() discovery changes affect command loading |
| Change component prefix routing | Component tests verify button/select menu routing |
Update snapshots with: pnpm vitest run --config vitest.component.config.ts <file> --update
The user is the ONLY manual-QA executor β usually on a phone, across session crashes and compactions. Every request for manual verification MUST be a complete, self-contained instruction with all five parts:
buildModelFooterText, not a persona's in-character explanation) before
stating it.Checklists are durable artifacts, not chat. A release smoke-test checklist
is written into CURRENT.md (with per-item status), so it survives compaction
and the user can re-consult it from a phone. Track progress there as results
come in ("from my quick tests how much of the checklist did that get us?" must
be answerable from the file). Justify every case β a bloated matrix wastes the
user's time; each case states what it uniquely proves. After the user reports,
close the loop with runtime evidence (logs) when the user-visible outcome can't
prove the new code path ran.
Each checklist item carries a confidence tier β high (CI + review + blast-radius cover it; ships without a manual round) or needs-smoke (with the one-line reason confidence is limited: runtime-unverified path, mobile rendering, a failure sequence tests can't reach). Surface only the needs-smoke tier for the owner to run β they don't want to hand-test high-confidence changes ("what's the level of confidenceβ¦ I don't feel like testing them unless confidence is limited").
Batch smoke asks into ONE numbered pass at release kickoff β including deferrals carried from previous releases. The owner's explicit ask: "I would like to know everything you need me to smoke test in dev so I can just knock it all out in one go. maybe including the deferred items still hanging around from previous releases too." Sweep CURRENT.md for carried-over deferred items before presenting; drip-feeding asks across the session forces repeat phone sessions. Two gates before an item enters the batch: (a) retro-log-close first β check whether prod/dev logs can already prove the item ran correctly (a 15-deployment sweep has closed several items the owner would otherwise have hand-tested; the owner's "I feel like I probably did 8 already" is a search order); (b) the item is only askable for work that is merged to develop β dev deploys from develop, so "smoke test my branch" is not an executable request.
.component.test.ts.test.ts (unit-tier; no dedicated schema suffix)pnpm ops test:audit to verify no new gapsdocs/reference/guides/TESTING.mdservices/*/src/test/mocks/docs/reference/testing/PGLITE_SETUP.mddocs/reference/testing/COVERAGE_AUDIT_SYSTEM.md.claude/rules/02-code-standards.md