data-quality-audit
Comprehensive data quality assessment against business rules, schema constraints, and freshness expectations. Activate when validating data pipeline outputs before production use, auditing a dataset against defined business rules, or producing a quality scorecard for a data asset.
How do I install this agent skill?
npx skills add https://github.com/nimrodfisher/data-analytics-skills --skill data-quality-auditIs this agent skill safe to install?
- Gen Agent Trust Hubpass
The skill is a data quality audit tool designed to validate datasets against business rules and schema constraints. It is logically sound but handles untrusted external data (CSV/Parquet) and incorporates findings into reports, which represents a potential surface for indirect prompt injection.
- Socketpass
No alerts
- Snykpass
Risk: LOW · No issues
- ZeroLeakspass
Score: 93/100 · 2 sections analyzed
What does this agent skill do?
When to use
- A data pipeline has just loaded new data and needs validation before downstream reports consume it
- A stakeholder has flagged data quality concerns (wrong totals, unexpected nulls, stale data)
- You need to produce a formal data quality scorecard for a data asset as part of a data governance process
- You are onboarding a new data source and need to understand its quality profile before building on it
Process
- Null and completeness audit — run
scripts/null_counter.pyfor a column-by-column null profile. Flag columns above acceptable thresholds for the business context. - Duplicate detection — run
scripts/duplicate_finder.pyto identify full-row and key-level duplicates. Determine if duplicates are intentional (versioning) or errors (pipeline fan-out). - Referential integrity check — run
scripts/referential_integrity.pyto validate that foreign key values in child tables exist in parent tables. Report orphan rate per relationship. - Value range validation — run
scripts/value_range_validator.pywith business rules defined inreferences/business_rule_patterns.md. Flag values outside acceptable ranges. - Freshness check — run
scripts/freshness_check.pyto verify the dataset is up to date — compare the latest record timestamp against the expected lag for this pipeline. - Score and classify findings — map each finding to a quality dimension using
references/quality_dimensions.md. Assign severity (CRITICAL / HIGH / MEDIUM / LOW). - Produce deliverables — fill
assets/audit_report_template.htmlfor a shareable report; fillassets/quality_rubric.mdfor a concise scorecard.
Inputs the skill needs
- Required: dataset (CSV / Parquet / database table reference)
- Required: schema relationships — which columns are primary keys, which are foreign keys to which tables
- Required: business rules — acceptable value ranges, expected value sets, freshness SLA
- Optional: acceptable error rates — at what threshold does a failure become CRITICAL vs. HIGH
- Optional: pipeline schedule — to assess freshness relative to expected update frequency
Output
assets/audit_report_template.html(filled) — full quality report, shareable with stakeholdersassets/quality_rubric.md(filled) — one-page quality scorecard with dimension scores- Script console output — per-check pass/fail counts for each validation script
How can the creator link this skill?
Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.
<a href="https://skillzs.dev/skills/nimrodfisher/data-analytics-skills/data-quality-audit">View data-quality-audit on skillZs</a>