skillZs
★ LIVE SKILL TAGS ★
>>> LIVE SKILLS INDEX <<<
* OPEN SOURCE *
NO LOGIN, NO TRACKING
※ REAL INSTALL DATA ※
← back to all skills
vasilyu1983/ai-agents-public114 installs

dev-ai-coding-metrics

Measures AI coding impact and extension robustness. Use when tracking delivery, quality trajectories, cost, experience, pilots, scorecards, or leadership reporting.

How do I install this agent skill?

npx skills add https://github.com/vasilyu1983/ai-agents-public --skill dev-ai-coding-metrics
view source ↗

Is this agent skill safe to install?

  • Gen Agent Trust Hubpass

    The skill provides a framework for measuring AI coding productivity, including scripts for ROI calculation and GitHub telemetry extraction. It uses only Python standard libraries, manages credentials securely via environment variables, and communicates with well-known services (GitHub API) for data retrieval. No malicious behavior or security risks were identified.

  • Socketpass

    No alerts

  • Snykwarn

    Risk: MEDIUM · 1 issue

What does this agent skill do?

AI Coding Metrics

Measures coding assistants and coding agents without collapsing results into vanity metrics or one blended score.

The critical distinction is mode: assistants help inline or in chat; agents execute multi-step work and need task-level measurement. Do not measure them as if they were the same thing.

When to Use This Skill

TriggerExample
Designing a pilot or rollout scorecard"We're rolling out Copilot to 200 engineers — what do we measure?"
Diagnosing usage-up / outcomes-flat"Seat utilization is 80% but PR throughput is unchanged"
Comparing assistant vs. agent workflows"Should we instrument these separately?"
Building an ROI model or leadership report"Finance wants a renewal decision by Q3"
Designing an experiment better than vendor benchmarks"We can't trust the vendor's numbers — how do we run our own study?"

Defaults

RuleRationale
Start from the decision, not the telemetry availablePrevents instrument-what-is-easy bias
Separate assistant and agent funnelsMixing hides which workflow drives results
Pair every speed metric with quality + experienceSpeed alone is misleading
Aggregate at team levelIndividual dashboards become surveillance; see Individual-Data Policy
Treat benchmarks as capability signals, not business KPIsBenchmark gaps do not equal production gaps

Workflow

  1. Define the decision.
  2. Pick the program mode: assistant, agent, or mixed.
  3. Build the minimum viable scorecard.
  4. Choose the study design.
  5. Produce one deliverable.

Quick Reference

Decision to deliverable:

DecisionDefault Output
buy, renew, or cut a toolROI model plus executive report
improve adoptionadoption metrics plus survey
prove delivery impactproductivity metrics plus experiment plan
check quality driftquality metrics plus dashboard
understand trust or frictiondeveloper-experience metrics plus survey
evaluate coding agentsagent-execution metrics plus experiment plan

Program Modes

ModeUnit of AnalysisPrimary Emphasis
assistantdeveloper-day, team-week, repo-monthadoption, delivery, quality, experience
agenttask, PR, workflow runtask success, merge, revert, review burden, cost per accepted change
mixedteam-week plus task-level samplesseparate the two funnels before combining results

Metric Families

Use the smallest scorecard that can answer the decision:

FamilyWhat It Tells You
adoptionwhether usage is real and sustained
deliverywhether software flow is faster where AI actually touches the path
qualitywhether speed gains are offset by defects, rework, review burden, or declining extension robustness
economicswhether the value justifies tool and operating cost
experiencewhether developers trust the tool and want to keep using it
agent executionwhether autonomous workflows succeed in production, not just in demos

Study Design Defaults

Planning default: 8 weeks of pre-intervention data. This is a local planning floor, not a research-derived sufficiency threshold: before a causal pilot, compute the minimum detectable effect from the baseline's own week-to-week variance and the number of team-weeks. If that MDE is larger than any effect you would act on, do not run a causal study; report a descriptive scorecard labelled directional.

SituationDesign
new pilot, no control groupdifference-in-differences against non-adopting or late-adopting teams, or an interrupted time series with ≥8 pre-period points; a single before/after delta is directional only (it absorbs freezes, reorgs, and hiring waves)
enough comparable teamsmatched A/B or stratified assignment
teams resist permanent denial of toolscrossover design
agent workflow change on one task familytask-level shadow comparison or reviewer-blind evaluation
leadership wants a fast answerbalanced scorecard with explicit caveats, not a causal claim

Cohort and denominator contract

Freeze the measurement population before reading outcomes. Record the eligible population, assignment rule, actual exposure, observation window, and accepted outcome for each metric. Report eligible, assigned, exposed, and observed counts side by side; never silently replace the assigned cohort with active users, completed tasks, or merged PRs. That survivor-only denominator makes adoption and success look better precisely when setup failures, abandoned agent runs, or unmerged changes are the problem.

For incomplete observations, name the reason (not_started, abandoned, still_open, telemetry_missing, or excluded_by_rule) and keep it in the funnel. Treat still-open work as right-censored rather than failed until the outcome window closes. A report may be directional with imperfect telemetry, but it must state which denominator supports each percentage and how missing cases could change the decision.

When exporting GitHub PRs, retain open, closed-unmerged, and merged outcomes in the opened-since cohort. A PR export cannot recover assigned tasks that never produced a PR; join it to the task registry before reporting agent success. Treat author-level CSV as restricted source data under the Individual-Data Policy.

Measurement Checklist

Use before publishing any AI coding report:

  • Baseline established (≥8 weeks before intervention)
  • Assistant and agent funnels tracked separately
  • Every speed metric paired with at least one quality metric
  • Sample size, confidence level, and study design stated
  • Confounds documented (team changes, release pressure, policy changes)
  • Vendor evidence labeled as vendor evidence
  • Usage measured after stabilization (not week-1 novelty period)
  • Review burden and rework cost included in ROI model
  • Edit-capable agents measured across evolving-spec checkpoints, including late-checkpoint cost and quality slopes
  • Aggregated at team level (no manager-visible individual dashboards)

Evidence Rules

Durable rules from the research; the dated studies, figures, and caveats live in references/evidence-update.md:

  • AI amplifies existing system strengths and weaknesses; it is not a universal accelerant.
  • Perceived and measured gains can differ; never report self-reported gains as measured impact.
  • Throughput gains can shift load to review and incidents; pair any throughput gain with review-burden and incident metrics.

Anti-Gaming Checklist

Reject a scorecard or report if any of the following apply:

  • Single blended AI productivity score mixing usage, speed, sentiment, and quality
  • Seat activation or prompt volume cited as delivery impact
  • Cross-team comparison without controlling for stack, task mix, staffing, or release pressure
  • Measurement period is <8 weeks or includes week-1 novelty window
  • Vendor benchmark cited as production ROI evidence
  • Review burden excluded from ROI model
  • Individual-level AI usage visible to managers
  • Directional before/after movement stated as causal without controlled design
  • SlopCodeBench averages or trajectory signals used as organizational targets or causal ROI evidence

Individual-Data Policy

This is the one policy for person-level data, shared with dev-contribution-quality-analysis (which points here):

  • Team level by default. AI usage and AI-assist metrics are never shown to managers per person and never feed performance, promotion, or staffing decisions.
  • Person-level output is allowed only as opt-in self-review or coaching, shown to the person first, with the sampled evidence and a correction path.
  • Questions about individuals ("who uses AI well", "per-engineer quality") route to dev-contribution-quality-analysis under these limits; team-level AI-tool impact stays here.

Navigation

References

Assets and data

Scripts

  • scripts/roi_calculator.py — when calculating a capacity-value scenario, supply the observation period and both measured burden terms; zero-cost ROI is undefined. Per-family rating bands are illustrative local rubrics, not validated constructs.
  • scripts/extract_github_events.py — when collecting the PR outcome cohort or commit telemetry; look up GitHub endpoint permissions and rate-limit guidance before a live export.
  • scripts/README.md

Cross-References

Learnings Loop

When prior decisions or pitfalls are relevant, consult learnings.consolidated.md if present; use learnings.md only for needed history or as the available fallback. Otherwise skip both.

After applying it, if you encountered a pattern worth remembering, a mistake worth preventing, or a domain fact that surprised you, append one dated bullet to learnings.md via agents-skills-feedback-loop/scripts/append_learning.py. Do not modify SKILL.md itself.

Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.

<a href="https://skillzs.dev/skills/vasilyu1983/ai-agents-public/dev-ai-coding-metrics">View dev-ai-coding-metrics on skillZs</a>