Review all evaluations in the repository against a single code quality standard. Checks ALL evals against ONE standard for periodic quality reviews...
Review all evaluations in the repository against a single code quality standard or topic. This workflow is useful for systematic code quality improvements and ensuring consistency across all evaluations.
If not already provided, ask user for the topic from the CONTRIBUTING.md or BEST_PRACTICES.md that this review should be focused on. If not provided, come up with a short topic identifier in a file-safe format (e.g., pytest_marks, import_patterns, test_coverage).
Create or read existing directory structure:
<repo root>/agent_artefacts/code_quality/<topic_id>/ - Directory for this review topic<repo root>/agent_artefacts/code_quality/<topic_id>/README.md - Documentation for this specific topic<repo root>/agent_artefacts/code_quality/<topic_id>/results.json - Results of the review<repo root>/agent_artefacts/code_quality/<topic_id>/SUMMARY.md - Summary of the reviewThe README.md file should contain topic-specific information:
The results.json file should follow the template in assets/results-template.json. It contains one entry per evaluation in <repo root>/src/inspect_evals/, with status and issue details.
Important: The issue_location field should use paths relative to the repository root with forward slashes (e.g., tests/foo/test_foo.py:42 or src/inspect_evals/foo/bar.py:15, not C:\Users\...\test_foo.py:42 or tests\foo\test_foo.py:42).
Systematic Approach: Review all evaluations in a consistent manner. Use scripts or automated tools where possible to ensure completeness.
Clear Issue Reporting: Each issue should include:
Verification: After identifying potential issues, verify a sample of them by reading the actual files to ensure accuracy.
Statistics: Provide summary statistics including:
Prioritization: Identify which issues are most critical or affect the most evaluations to help guide remediation efforts.
Reusability: Write scripts and documentation that can be rerun easily as the codebase evolves. Include any helper scripts in the topic directory.
False Positives: Be aware that automated detection may produce false positives. When possible, include logic to reduce these or document known limitations.
agent_artefacts/code_quality/<topic_id>/uv run inspect-evals-lint --all --select <rule>. Parse its output to identify structural issues across evals. For topic-specific checks beyond autolint's scope, write targeted grep/AST scripts in the topic directory.README.mdresults.json with relative paths:SUMMARY.md file in the topic directory with:This skill (code-quality-review-all) owns results.json and has full control:
The code-quality-fix-all skill has limited control:
This separation ensures:
After running this workflow, you should have:
agent_artefacts/code_quality/<topic_id>/
āāā README.md # Topic-specific documentation
āāā results.json # Detailed results for all evaluations
āāā SUMMARY.md # Executive summary
āāā <helper_scripts> # Optional: automated checker scripts