Analyze Nixtla baseline forecasting results (sMAPE/MASE on M4 or other benchmark datasets)...
Review one concrete Baseline Lab result set. Rank models from observed metrics, show both aggregate and per-series evidence, and keep benchmark conclusions inside the dataset, horizon, and series sample that produced them.
results_*.csv path or a directory containing one.series_id, model, sMAPE, and MASE.Resolve the input with Glob. If several result files match, list their
paths and modification context and ask the user to choose; never silently
combine separate runs or pick a filename as "latest" without confirmation.
Use Read to inspect the header, representative rows, sibling summary, and
run_manifest.json or compat_info.json when present. Record the dataset,
horizon, series count, models, frequency, and seasonal period that are
actually supported by those artifacts.
Run the bundled analyzer:
python3 ${CLAUDE_SKILL_DIR}/scripts/analyze_results.py path/to/results.csv
Check the analyzer receipt for rejected rows, duplicate series/model pairs, ties, model coverage, and non-finite values. Stop on schema or numeric failures instead of averaging partial data.
Rank models by the metric the user requested; otherwise lead with mean and median sMAPE, then report MASE and win counts as supporting evidence. Do not turn a small mean difference into a categorical claim.
Inspect series where models disagree or all reported errors are high. State patterns as hypotheses unless timestamps, seasonality, or covariates in the source data actually support them.
Give a bounded recommendation for the observed evaluation only. Require out-of-sample validation, cost/latency checks, and operational review before calling any model production-ready.
Return scope and provenance, data-quality receipt, per-model mean/median/std-dev sMAPE, mean/median MASE, series win counts and ties, failure cases, a scoped recommendation, uncertainties, and reproducible next steps.
Use prompts that name the evidence and comparison boundary:
Review nixtla_baseline_m4/results_M4_Daily_h7.csv. Rank by median sMAPE,
show ties and missing coverage, and keep conclusions limited to this run.
Compare AutoETS and AutoTheta using both sMAPE and MASE. Identify series where
the metrics disagree and do not use generic accuracy labels.
scripts/analyze_results.py005-plugins/nixtla-baseline-lab/scripts/nixtla_baseline_mcp.py