Audits Helm charts for anti-patterns, security issues, and best practice violations. Use when asked to audit, review, or check Helm chart quality...
Enforce Helm chart quality and security standards across the helm-charts/ directory through automated checks.
What it checks (13 checks):
Full audit (all checks):
node .claude/skills/helm-charts-audit/scripts/run_all_checks.mjs
Generate report (all checks + markdown report):
node .claude/skills/helm-charts-audit/scripts/generate_report.mjs
Report saved to: reports/YYYY-MM-DD/helm-charts-audit.md
Individual checks:
node .claude/skills/helm-charts-audit/scripts/check_image_tags.mjs
node .claude/skills/helm-charts-audit/scripts/check_security_context.mjs
node .claude/skills/helm-charts-audit/scripts/check_resource_limits.mjs
node .claude/skills/helm-charts-audit/scripts/check_rbac_wildcards.mjs
node .claude/skills/helm-charts-audit/scripts/check_health_probes.mjs
node .claude/skills/helm-charts-audit/scripts/check_helm_lint.mjs
node .claude/skills/helm-charts-audit/scripts/check_chart_metadata.mjs
node .claude/skills/helm-charts-audit/scripts/check_chart_structure.mjs
node .claude/skills/helm-charts-audit/scripts/check_dependencies.mjs
node .claude/skills/helm-charts-audit/scripts/check_deprecated_apis.mjs
node .claude/skills/helm-charts-audit/scripts/check_argo_rollouts.mjs
node .claude/skills/helm-charts-audit/scripts/check_ingress_tls.mjs
node .claude/skills/helm-charts-audit/scripts/check_gpu_resources.mjs
RULE: Never use mutable tags. latest tag = unpredictable deployments + rollback failures.
Violations:
image: nginx:latest - mutable, changes without noticeimage: nginx - defaults to :latesttag: "" - empty tag in values.yamltag: head, tag: canary, tag: dev - mutable branch tagsFix: Use immutable tags like v1.2.3, SHA digests sha256:abc123, or SemVer 1.21.0.
RULE: Containers must run with minimal privileges. Privileged containers = cluster takeover risk.
Violations:
privileged: true - full host access, container escape trivialrunAsNonRoot: false - runs as root user UID 0runAsUser: 0 - explicitly rootallowPrivilegeEscalation: true - can gain more privilegeshostNetwork: true - shares host network namespacehostPID: true - can see/kill host processeshostIPC: true - can access host shared memoryreadOnlyRootFilesystem: false - malware can write anywherecapabilities.add: [SYS_ADMIN] - near-root level accesscapabilities.add: [ALL] - equivalent to privilegedFix: Add proper securityContext with runAsNonRoot: true, allowPrivilegeEscalation: false, readOnlyRootFilesystem: true, capabilities.drop: [ALL].
RULE: All containers must have resource requests and limits. No limits = node OOM + noisy neighbor issues.
Violations:
resources: {} - empty resources blockrequests.cpu - scheduler can't make decisionsrequests.memory - OOM killer may terminate unexpectedlylimits.memory - container can consume all node memoryrequests > limits - invalid configurationFix: Define resources.requests.cpu, resources.requests.memory, resources.limits.memory. Note: CPU limits often intentionally omitted for better performance.
RULE: Follow least-privilege principle. Wildcard permissions = privilege escalation path.
Violations:
verbs: ["*"] - grants all actionsresources: ["*"] - access to all resource typesapiGroups: ["*"] - access across all API groupsroleRef.name: cluster-admin - full cluster accessverbs: [impersonate] - can act as other usersverbs: [escalate, bind] - can grant additional privilegessecrets resource - can read all secretsFix: Use explicit verbs like [get, list, watch], explicit resources like [pods, services], avoid cluster-admin bindings.
RULE: All workloads must have health probes. No probes = stuck containers not restarted + traffic to unready pods.
Violations:
livenessProbe - stuck containers won't restartreadinessProbe - traffic sent to unready podsinitialDelaySeconds: 0 - probes start immediately, false failurestimeoutSeconds: 1 - too short, may cause false failuressuccessThreshold > 1 on livenessProbe - should always be 1failureThreshold > 10 - delays detecting actual failuresFix: Add livenessProbe and readinessProbe with reasonable initialDelaySeconds (10-30s), periodSeconds (10s), timeoutSeconds (5s).
RULE: Charts must pass official helm lint validation. Lint failures = deployment failures.
Violations:
Fix: Run helm lint <chart-path> and fix reported issues.
RULE: Chart.yaml must have complete metadata. Missing metadata = maintenance nightmare.
Violations:
apiVersion: v1 - Helm 2 format, upgrade to v2version - must be SemVerappVersion - hard to track what's deployeddescription - unclear what chart doesmaintainers - no ownershipname doesn't match directory name - confusingFix: Use apiVersion: v2, SemVer version, add description and maintainers with email.
RULE: Follow standard Helm chart structure. Non-standard = user confusion + missing features.
Violations:
README.md - no documentationtemplates/NOTES.txt - no post-install instructionstemplates/_helpers.tpl - no template helpers.helmignore - unnecessary files in packagevalues.schema.json - no values validationtemplates/ directoryFix: Create missing files following Helm chart best practices.
RULE: Pin dependency versions. Floating versions = non-reproducible builds.
Violations:
version on dependency - unpinnedversion: "*" or version: "^1.0" - floating versionChart.lock - dependency versions not lockedrepository: file:// - local reference, breaks when publishedrepository: http:// - insecure, use HTTPSFix: Pin exact versions, run helm dependency update to generate Chart.lock.
RULE: Use stable Kubernetes APIs. Deprecated APIs = upgrade failures.
Violations:
extensions/v1beta1 - removed in K8s 1.22apps/v1beta1, apps/v1beta2 - removed in K8s 1.16networking.k8s.io/v1beta1 Ingress - removed in K8s 1.22batch/v1beta1 CronJob - removed in K8s 1.25policy/v1beta1 PodSecurityPolicy - removed in K8s 1.25Fix: Update to stable APIs: apps/v1, networking.k8s.io/v1, batch/v1. Run kubectl convert if needed.
RULE: Rollouts must have valid strategy configuration. Invalid config = failed deployments.
Violations:
strategy - no deployment strategysteps - no gradual rolloutanalysis - no automated validationactiveService - no active service definedpreviewService - can't preview before promotionrevisionHistoryLimit - old ReplicaSets accumulateprogressDeadlineSeconds - stuck rollouts don't timeoutFix: Configure proper canary steps with analysis, or blueGreen with activeService/previewService.
RULE: Ingress must have TLS configuration. No TLS = unencrypted traffic.
Violations:
secretName - certificate source unclearingressClassName - may use wrong controllerkubernetes.io/ingress.class annotationFix: Add TLS section with secretName, use cert-manager.io/cluster-issuer annotation for automated certs.
RULE: GPU workloads need proper configuration. Missing config = scheduling failures.
Violations:
Fix: Set nvidia.com/gpu in both requests and limits (equal values), add GPU tolerations and nodeSelector.
This skill uses VALUE-BASED detection:
--- separators