ai-ml-timeseries
Builds and backtests demand and revenue forecasts: rolling-origin backtests, prediction intervals, reconciliation, zero-shot foundation models. Use when forecasting time series.
How do I install this agent skill?
npx skills add https://github.com/vasilyu1983/ai-agents-public --skill ai-ml-timeseriesIs this agent skill safe to install?
- Gen Agent Trust Hubpass
The skill provides comprehensive guidelines and a utility script for time-series forecasting. It includes references to reputable research and software from established technology companies and research institutions. A minor security surface exists in the evaluation script, which processes user-provided JSON data to generate reports.
- Socketpass
No alerts
- Snykpass
Risk: LOW · No issues
- Runlayerpass
1/1 file flagged
- ZeroLeakspass
Score: 93/100 · 2 sections analyzed
What does this agent skill do?
Time Series Forecasting - Production Patterns
Define the prediction cutoff before modelling: only information available then can enter features or baselines. Compare models on the decision horizon against the forecast currently driving decisions.
Scope Boundaries
- General EDA, tabular modelling, experiment design, or reusable DS workflow -> ai-ml-data-science
- Deployment architecture, monitoring stack, release gates, incident playbooks -> ai-mlops
- Generic LLM lifecycle, prompting, or provider selection -> ai-llm
- RAG and search systems -> ai-rag
Workflow
- Define the forecast target, horizon, cutoff timestamp, granularity, and business-loss shape first.
- Route generic DS workflow or full production-ops work to the adjacent skill when forecasting depth is not the main need.
- Choose the baseline, feature pattern, and model family from the decision tree.
- Run leakage-safe backtests, compare horizon-aware metrics, and add interval or calibration checks when decisions need uncertainty.
- Before selecting a package or checkpoint, open its official docs from data/sources.json and check the installed API, supported covariates, context/horizon limits and weight licence. For TSFM comparisons, load model roles and lookup steps; record the benchmark snapshot next to any quoted rank. For MAPIE code, load probabilistic patterns and compare the example to the installed release's docs.
Quick Reference
Choose The Forecasting Pattern
Need to build or review a forecast:
├─ One series or a small local portfolio?
│ ├─ Few covariates, interpretability first -> naive/seasonal naive + ETS/SARIMAX baseline
│ └─ Nonlinear effects or richer covariates -> feature-based boosting
│
├─ Many related series?
│ ├─ Need one shared model with covariates -> global/panel forecasting
│ └─ Need rollups to add up across levels -> hierarchical forecasting + reconciliation
│
├─ Intermittent or zero-heavy demand?
│ ├─ Very sparse and operationally simple -> Croston/SBA/ADIDA baseline
│ └─ Need richer covariates or scale -> global boosting with zero-aware features
│
├─ Need uncertainty, service levels, or inventory decisions?
│ └─ Add quantiles, conformal intervals, or distributional forecasts
│
├─ Long horizon or low-feature setup?
│ ├─ Need strong zero-shot baseline -> TS foundation models
│ └─ Need supervised optimization on many series -> global/deep forecasting
│
└─ Need production handoff?
└─ Record cutoff timestamp, feature contract, fallback, lineage, and retraining trigger
Core Principles
1. Cutoff Timestamp Before Features
- Define the exact prediction timestamp before building labels or features.
- Every lag, rolling aggregate, calendar flag, and external signal must be justified relative to what is known at that cutoff.
- Known-future covariates and future-unknown covariates must be treated differently.
2. Baselines First
- Always compare against naive and seasonal-naive baselines, scored at the same horizon from the value known at each origin; a baseline built from any later actual leaks the future.
- Add ETS/SARIMAX or other local classical baselines when interpretability matters.
- Candidate models do not earn deployment if they barely beat a simple baseline.
- Naive and seasonal naive are the floor. The promotion bar is the forecast currently driving decisions, planner overrides included, scored on the same origins. Log overrides as a separate forecast so their value can be measured too.
- A point-skill win is a screen, not a promotion. Promote only when a paired test on per-origin loss differentials at the decision horizon excludes zero (paired promotion test).
3. Horizon And Slice Evaluation
- Report accuracy by horizon, not only a single global score.
- Slice by segment, geography, SKU family, volume band, or any business-relevant cohort.
- Prefer MASE/WAPE/MAE over MAPE when zeros or near-zeros exist, except on intermittent series.
- When more than half the periods are zero, MAE, WAPE and MASE reward an all-zero forecast (they are minimised by the median). Select mean forecasts on RMSSE and stocking decisions on pinball loss; never on MAE or WAPE alone (intermittent metrics).
- Judge each horizon by skill against the baseline at that horizon. Error that grows with horizon is expected and is not a defect on its own.
4. Global/Panel Before Per-Series Complexity
- When many related series exist, default to a shared global/panel approach before building separate bespoke models.
- Use hierarchical reconciliation when forecasts must remain coherent across levels.
- Promote complexity only when it earns accuracy, calibration, or operational simplicity.
5. Probabilistic When The Decision Is Risk-Sensitive
- Use quantiles, conformal intervals, or full predictive distributions when downstream actions depend on uncertainty.
- Evaluate both coverage and sharpness; wide intervals with nominal coverage are not automatically useful.
- Screen coverage at each horizon with a Wilson interval only under independent coverage hits with a common probability; distinct origins alone do not establish independence. Dependent residuals need block- or series-aware evidence before promotion. Keep pooled coverage descriptive and inspect clustered misses (coverage against noise).
- Set the decision quantile from costs, q* = c_u / (c_u + c_o), and score it with pinball loss and per-horizon coverage (decision quantile).
- Reassess calibration after every retrain or major model change.
6. Forecasting-Specific Handoff
- A forecast package is incomplete without cutoff timestamps, horizon definition, feature contract, metric definitions, fallback rules, and lineage metadata.
- Keep forecasting-specific handoff guidance here; route full deployment architecture to ai-mlops.
Forecast Decision Gate
Create a cutoff ledger for every backtest fold: training end, forecast origin, horizon, feature availability, retrain policy, and any revision or publication lag. Report error and interval quality by horizon and decision-critical slice before aggregating. Promote a forecast only when it beats the forecast currently in use (planner overrides included) on the business-weighted loss, and a paired test across origins confirms the gain, without hiding a blocking slice. The fallback for missing or late covariates must also be tested. Beating the naive or seasonal-naive floor is necessary but never sufficient.
Known Traps
- Mixing future-known covariates and future-unknown covariates in the same feature path without documenting which values are actually available at forecast time.
- Using one global backtest score to justify deployment when horizon-specific error behavior differs materially.
- Using MAPE on zero-heavy, intermittent, or near-zero series and then comparing models on unstable percentages.
- Building bespoke per-series models before testing a strong global or panel baseline on related series.
- Ignoring hierarchy and coherence when downstream consumers expect rollups to add up.
Common Anti-Patterns
- Reusing IID validation habits from generic tabular ML instead of rolling-origin or expanding-window evaluation.
- Treating decomposition visuals as evidence of production signal without backtesting the actual decision horizon.
- Using feature-rich models whose future covariates are unavailable or operationally too expensive to maintain.
- Comparing TS foundation models to weak baselines and calling the result strategic proof.
- Trusting a TSFM zero-shot win before reading the checkpoint's weight licence and re-testing on private or post-cutoff data (TSFM trust gate).
Navigation: Core References
Data Integrity And Features
- TS EDA Best Practices - Timestamp integrity, missingness, decomposition, and stability checks
- Lag & Rolling Patterns - Leakage-safe lags, rolling windows, and calendar patterns
- Global & Panel Forecasting Patterns - Shared models, panel schemas, known-future covariates, grouped evaluation
Model And Strategy Selection
- Model Selection Guide - Model-family decision matrix
- LightGBM TS Patterns - Feature-based/global boosting patterns, MLForecast/skforecast workflows
- Multi-Step Forecasting Patterns - Direct, recursive, and seq2seq tradeoffs
- Intermittent Demand Patterns - Sparse-demand baselines and zero-aware modelling
Validation, Uncertainty, And Advanced Forecasting
- Backtesting Patterns - Rolling-origin evaluation, panel-aware backtests, and metric design
- Probabilistic Forecasting - Quantiles, conformal methods, calibration, and scoring rules
- Hierarchical Forecasting - Coherent forecasts and reconciliation methods
- Time-Series Foundation Model Patterns - TSFM roles, trust gate (licence, contamination), zero-shot benchmarking
- Anomaly Detection Patterns - Residual and interval-based anomaly workflows
Handoff And Forecast Operations
- Forecast Governance Patterns - Cutoff timestamps, lineage, fallback rules, and forecast contracts
- Production Forecast Operations - Horizon-matched live monitoring against the backtest, actuals latency, forecast of record, fallback ladder
Templates
Data Preparation
- TS EDA Template - Reproducible structure for timestamp and seasonality review
- Resample & Fill Template - Resampling, gap rules, and fill policies
Feature And Model Design
- Lag & Rolling Features - Leakage-safe feature specification
- Calendar Features - Known-future business calendar and event feature spec
- Forecast Model Template - Forecast package contract for local, global, or hierarchical models
- Multi-Step Strategy - Direct, recursive, and seq2seq strategy contract
Evaluation And Uncertainty
- Backtest Template - Rolling-origin or expanding-window evaluation spec
- TS Metrics Template - Horizon, slice, business-loss, and probabilistic metric contract
Foundation Models
- TS Foundation Model Template - Zero-shot TSFM benchmark and evaluation scaffold
Scripts
| Script | Purpose |
|---|---|
| scripts/ts_evaluator.py | Stdlib-only CLI: horizon-wise metrics and skill vs naive/seasonal-naive from origin_value (no verdict below 6 origins), independence-assuming per-horizon Wilson screens, descriptive pooled coverage, Markdown report. Exits 2 on invalid input, partially supplied interval levels or an unwritable --output. Its "beats baseline" is a screen; run the paired test before promoting |
# Horizon-wise accuracy: MAE, RMSE, MAPE, MASE, skill vs the baseline at the same horizon (invalid input exits 2)
python scripts/ts_evaluator.py backtest --input data/sample-forecast-results.json
# Per-horizon calibration screen for 50%, 80%, 90% prediction intervals; pooled coverage is descriptive
python scripts/ts_evaluator.py calibration --input data/sample-forecast-results.json
# Full Markdown evaluation report written to file
python scripts/ts_evaluator.py report --input data/sample-forecast-results.json --output /tmp/ts-eval-report.md
Data Files
| File | Description |
|---|---|
| data/sources.json | Curated primary sources for classical forecasting, MLForecast, skforecast, AutoGluon, TSFM repositories, fev-bench, and MAPIE |
| data/sample-forecast-results.json | Synthetic rolling-origin backtest for a daily revenue forecast: horizons 1, 7, 14, 30 days, 48 rows, 12 origins, with origin_value and in-sample history for MASE |
External Sources
See data/sources.json for current primary sources across:
- classical forecasting references
- official docs for MLForecast, HierarchicalForecast, skforecast, AutoGluon TimeSeries, LightGBM, and MAPIE
- official TSFM repositories and model cards
- fev-bench benchmark and governance references used for high-impact deployments
Related Skills
- ai-ml-data-science - General DS workflows, experiment design, and broader modelling patterns
- ai-mlops - Deployment architecture, monitoring, and release operations
- ai-llm - Provider/model lifecycle questions outside time-series forecasting
- data-sql-optimization - Storage and query design for time-series marts
Learnings Loop
When prior decisions or pitfalls are relevant, consult learnings.consolidated.md if present; use learnings.md only for needed history or as the available fallback. Otherwise skip both.
After applying it, if you encountered a pattern worth remembering, a mistake worth preventing, or a domain fact that surprised you, append one dated bullet to learnings.md via agents-skills-feedback-loop/scripts/append_learning.py. Do not modify SKILL.md itself.
How can the creator link this skill?
Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.
<a href="https://skillzs.dev/skills/vasilyu1983/ai-agents-public/ai-ml-timeseries">View ai-ml-timeseries on skillZs</a>