ai-ml-data-science
Builds ML pipelines: EDA, features, leakage checks, evaluation. Use when doing data science or explaining fairness, privacy, speech, vision-language, or diffusion mechanics.
How do I install this agent skill?
npx skills add https://github.com/vasilyu1983/ai-agents-public --skill ai-ml-data-scienceIs this agent skill safe to install?
- Gen Agent Trust Hubpass
The skill is a professional data science engineering suite that focuses on reproducible workflows, model evaluation, and responsible AI. It includes utility scripts for generating model cards and scanning for data leakage that operate locally using standard libraries. No malicious patterns or significant security risks were detected.
- Socketpass
No alerts
- Snykpass
Risk: LOW · No issues
- Runlayerwarn
1/1 file flagged
What does this agent skill do?
Data Science Engineering Suite
Frame the decision and prediction timestamp before choosing models; features must exist at scoring time.
Quick Reference
| Need | Default Direction |
|---|---|
| reproducible Python workflow | uv plus scripts or git-friendly notebooks (marimo for reactive/diffable notebooks) |
| fast local analysis | DuckDB plus Polars; check installed API against Polars docs before adapting older examples |
| data contracts | Pandera or GX Core at dataset boundaries |
| tabular baseline | linear or logistic model plus tree-based candidate; add a tabular foundation model only after the licence and size lookup in references/modelling-patterns.md §3.1 |
| feature engineering | explicit train-serve-safe transforms |
| unlabeled text corpus | embed -> UMAP -> HDBSCAN -> c-TF-IDF; use topic-level LLM labels when document-level labels are unnecessary |
| tuning | Optuna only after the baseline is stable |
| evaluation | slices, threshold, calibration, uncertainty |
| handoff | model card, evaluation report, failure modes, monitoring expectations |
When To Use This Skill
- exploring datasets and checking modelling feasibility
- designing feature pipelines and leakage controls
- choosing and comparing model families
- clustering unlabeled text and discovering topics before a taxonomy or labeling effort exists
- building reproducible experiment workflows
- producing evaluation reports, model cards, and handoff artifacts
- reviewing whether an experiment is genuinely ready for production handoff
- explaining responsible-AI modelling mechanics: fairness and intersectionality, privacy, interpretability, poisoning, memorization, human oversight, and environmental trade-offs
- designing general multimodal models: contrastive image-text learning, fusion, VQA/document/video systems, diffusion control, adaptation, and quality-latency trade-offs
Route Elsewhere
- serving, retraining automation, monitoring, or incident response -> ai-mlops
- forecasting and temporal validation -> ai-ml-timeseries
- lakehouse or batch ingestion infrastructure -> data-lake-platform
- streaming infrastructure (Kafka, Flink, CDC) -> data-streaming
- prompting, fine-tuning, or LLM-system design -> ai-llm or ai-rag
Workflow
- Frame the decision, target, baseline, and prediction timestamp before touching models.
- Validate the dataset shape, ownership, and leakage risks.
- Build the simplest viable baseline first.
- Design point-in-time-correct features and compare stronger candidates only after the baseline is trustworthy.
- Evaluate with the same split strategy, same metric definitions, and same compute budget.
- Produce handoff artifacts with thresholds, calibration state, failure modes, and reproducibility notes.
Core Rules
- write down the prediction timestamp explicitly
- do not trust random splits where time or entity leakage is plausible
- compare at least one simple baseline against one stronger candidate
- treat thresholding, calibration, and uncertainty as part of the decision
- keep data version, feature version, seed, and split logic reproducible
- route serving, retraining and monitoring to ai-mlops
Prediction-Time Eligibility Gate
For every feature, write its source event, event time, availability time, transformation version, and entity join key. Exclude any value that would not exist at the declared prediction timestamp. Offline backfill and online serving may use different implementations only when point-in-time tests on matched entities and timestamps demonstrate semantic parity, transformation-version lineage is preserved, and production skew is monitored. Evaluate the surviving pipeline with the intended split unit and decision threshold, then compare it with the simplest actionable baseline. A model is decision-ready only when the predicted action, abstention path, and cost of false positives and false negatives are explicit.
Known Traps
- Using random train/test splits when time, entity, household, account, or session leakage is plausible.
- Building features with information that is only available after the prediction point, then calling the result "production ready."
- Tuning models before the baseline and metric definitions are stable.
- Reporting only AUC or one aggregate score while ignoring threshold choice, calibration, slice behavior, and operational tradeoffs.
- Letting notebook state become the real pipeline logic. Hidden ordering and cached state break reproducibility fast.
- A single feature with near-perfect standalone separation, or a metric a domain expert would find implausibly good — treat as a leakage bug report first, a discovery second (see
references/eda-best-practices.mdExpert Instincts). - Recurring entities require a split that matches deployment: group holdout for unseen-entity generalization; temporal splits can retain prior entities when scoring those entities again. In both cases enforce feature availability and fit preprocessing within training folds.
- Citing a library version, benchmark number, or API pattern from memory or an older tutorial without checking it against the currently installed version — tabular-ML tooling (Optuna, SHAP, scikit-learn, boosted-tree libraries) crosses breaking major versions inside a single year.
Pattern Chooser
| Problem Shape | Direction |
|---|---|
| tabular or relational | baseline plus tree-based comparison |
| time-ordered forecasting | route to ai-ml-timeseries |
| classical text or embeddings plus classifier | stay here |
| unlabeled text, unknown themes, topic discovery | stay here; see references/text-clustering-topic-modeling.md |
| LLM workflow, prompting, or RAG | route to ai-llm or ai-rag |
| deployment, monitoring, retraining | route to ai-mlops |
| ingestion or lakehouse architecture | route to data-lake-platform |
| responsible-AI concepts and modelling trade-offs | stay here; route operational controls to ai-mlops and measurement/red teaming to ai-evals |
| multimodal representations, fusion, VQA/document/video, diffusion mechanics | stay here; route production and evaluation to ai-mlops and ai-evals |
Core Patterns
Reproducible workspace
uvand explicit dependenciesuv.lockcommitted anduv sync --lockedin CI; on pandas 3.0 code, copy-on-write and thestrdtype change behaviour (seereferences/eda-best-practices.md)- script-first or git-friendly notebook entrypoints — for reactive, git-diffable notebooks consider marimo as an alternative to Jupyter; marimo is reactive (dependent cells auto-rerun), stores notebooks as plain Python scripts, and eliminates hidden-state ordering issues
- fixed seeds and explicit split logic
- logged dataset and feature assumptions
Feature engineering and contracts
- numeric, categorical, text, and time-based transforms
- point-in-time availability checks
- reusable encoders and documented freshness assumptions
Evaluation and decision readiness
- primary metric plus guardrails
- threshold strategy
- calibration and uncertainty handling
- slice analysis and qualitative error review
Autonomous experimentation
Use agent-driven experiment loops only when the metric is explicit, the search space is bounded, and each run is cheap enough to keep or revert automatically.
Templates
- assets/project/template-standard.md
- assets/project/template-quick.md
- assets/features/template-feature-engineering.md
- assets/eda/template-eda.md
- assets/evaluation/template-evaluation-report.md
- assets/evaluation/template-model-card.md
- assets/review/experiment-review-template.md
Scripts
| Script | Purpose |
|---|---|
| scripts/ml_toolkit.py | Generates model cards, leakage checks, and model-quality reports from a model-spec JSON |
| scripts/leakage_scan.py | Static leakage scanner for ML feature/target column specs (JSON/JSONL). Flags time-leakage, target-leakage, and ID-leakage anti-patterns from column metadata. Exit 0 means no metadata flags; 1 means findings; 2 means malformed input or I/O failure. Set prediction_time to test scoring-time eligibility; label_time alone is a weaker legacy cutoff. |
Typical usage:
python scripts/ml_toolkit.py card --input data/sample-model-spec.json
python scripts/ml_toolkit.py leakage --input data/sample-model-spec.json
python scripts/ml_toolkit.py report --input data/sample-model-spec.json --output report.md
ml_toolkit.py leakage and report exit 1 for WARN/FAIL, including missing temporal cutoffs, and 2 for invalid input/I/O errors. A static PASS requires row-level availability and split validation before handoff. The scanner accepts non-empty target and columns; JSONL uses one complete spec per line and rejects malformed records. Use matching ISO times (consistent timezone awareness) or symbolic T±<integer><s|m|h|d> offsets. See scripts/README.md for model-spec fields.
Navigation
Core references
- references/eda-best-practices.md
- references/feature-engineering-patterns.md
- references/data-contracts-lineage.md
- references/modelling-patterns.md
- references/evaluation-patterns.md
- references/class-imbalance-patterns.md
- references/hyperparameter-optimization.md
- references/text-clustering-topic-modeling.md — modular embed/UMAP/HDBSCAN/c-TF-IDF pipeline, representation-model reranking, per-topic (not per-document) LLM labeling, and when to prefer plain k-means
- references/interpretability-explainability.md
- references/responsible-ai-mechanics.md — fairness/intersectionality, differential privacy, explainability, poisoning/federated learning, re-identification, watermarking, human oversight/appeals, copyright/memorization, and environmental trade-offs
- references/multimodal-modeling.md — CLIP/SigLIP objectives, fusion, VQA/document/video systems, diffusion control/diversity/acceleration, adaptation, latency, and cost
- references/reproducibility-checklist.md
- references/llm-data-pipeline.md — evaluation-side contamination checks for LLMs (Min-K% screening, contamination-resistant benchmarks); for pretraining corpus curation (dedup, filtering, decontamination, mixing) go to ai-data-curation-pretraining first
- references/feature-freshness-streaming.md
- references/production-feedback-loops.md
- references/ml-diagrams.md — Mermaid diagram catalog for classical ML (k-means, logistic regression, decision trees, collaborative filtering) and neural net architectures (MLP, RNN, CNN, Transformer); for embedding in docs, READMEs, PR descriptions
Data and external references
Related Skills
- ai-architecture-advisor — when to use trees vs deep learning vs LLM (decide before building)
- ai-mlops
- ai-ml-timeseries
- data-lake-platform
- ai-llm
- ai-rag
- huggingface-datasets — now in the external
huggingface-skills:plugin
Learnings Loop
When prior decisions or pitfalls are relevant, consult learnings.consolidated.md if present; use learnings.md only for needed history or as the available fallback. Otherwise skip both.
After applying it, if you encountered a pattern worth remembering, a mistake worth preventing, or a domain fact that surprised you, append one dated bullet to learnings.md via agents-skills-feedback-loop/scripts/append_learning.py. Do not modify SKILL.md itself.
How can the creator link this skill?
Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.
<a href="https://skillzs.dev/skills/vasilyu1983/ai-agents-public/ai-ml-data-science">View ai-ml-data-science on skillZs</a>