Adversarial QA workflow comparing Quarto HTML against Beamer PDF benchmark. Iterates between critic (finds issues) and fixer (applies fixes) until APPROVED or max iterations reached.
Compare Quarto HTML slides against their Beamer PDF benchmark using an iterative critic/fixer loop.
Philosophy: The Beamer PDF is the gold standard. The Quarto translation must be at least as good in every dimension.
Phase 0: Pre-flight โ Phase 1: Critic audit โ Phase 2: Fixer โ Phase 3: Re-audit โ Loop until APPROVED (max 5 rounds)
| Gate | Condition |
|---|---|
| Overflow | NO content cut off |
| Plot Quality | Interactive charts >= static plots |
| Content Parity | No missing slides/equations/text |
| Visual Regression | Quarto >= Beamer in all dimensions |
| Slide Centering | Content centered, no jumping |
| Notation Fidelity | All math verbatim from Beamer |
Launch the quarto-critic agent to compare Beamer vs Quarto comprehensively. Report saved to quality_reports/[Lecture]_qa_critic_round1.md.
If not APPROVED, launch quarto-fixer agent to apply fixes (Critical โ Major โ Minor), re-render, and verify.
Re-launch critic to verify fixes. Loop back to Phase 2 if needed.
This is the loop-until-dry primitive from orchestrator-protocol.md: the critic returns FINDINGs (the hard-gate table is the CRITICAL roll-up, per orchestration-schemas.md); the loop converges when a round adds 0 new CRITICAL/MAJOR findings (deduped on id = sha1(file:line:locus)), not at a fixed round count.
summary-parity.md).Save to quality_reports/[Lecture]_qa_final.md with hard gate status, iteration summary, and remaining issues.
This skill's reviewers emit findings under the machine-checked contract in
finding-schema.json. Reports are JSON arrays.
Smoke-test the harness before spending review effort โ a run that fans out reviewers and then cannot write a valid report has wasted the whole pass:
echo '[]' | python3 scripts/validate-findings.py
Then, before presenting any summary:
python3 scripts/validate-findings.py <report>.json # exit 0 required
What the contract forces, and why:
rule โ the documented rule or standard violated. A finding citing no rule is an
opinion, and opinions do not gate a commit.failing_case โ a concrete configuration under which the claim breaks, or the exact
missing hypothesis. "This could be clearer" does not validate.id = sha1("<file>:<line>:<locus>") โ deterministic, so dedup across rounds is
exact and the two-strikes rule is checkable rather than eyeballed.mechanical โ true only for fixes that cannot change a result (typo, cross-reference,
formatting, label). Never for an estimand, assumption, specification, inference
procedure, sample definition, or reporting language: those return to the researcher.Apply the per-lens evidence burdens and the "does NOT count" filters in
orchestration-schemas.md ยง7 before
verification, so known false alarms never reach the judge. The verifier pass is
refute-biased: only verdict: "confirmed" findings ship; anything it cannot ground is
dropped, not downgraded to a warning.