skillZs
★ LIVE SKILL TAGS ★
>>> LIVE SKILLS INDEX <<<
* OPEN SOURCE *
NO LOGIN, NO TRACKING
※ REAL INSTALL DATA ※
← back to all skills
borghei/claude-skills122 installs

observability-designer

Design observability strategies: SLI/SLO frameworks, alerting, and dashboards. Use when instrumenting a production service, tuning alert rules, designing Grafana dashboards, defining SLOs and error budgets, or reducing alert fatigue.

How do I install this agent skill?

npx skills add https://github.com/borghei/claude-skills --skill observability-designer
view source ↗

Is this agent skill safe to install?

  • Gen Agent Trust Hubpass

    The observability-designer skill is a toolkit for managing SLOs, alerts, and dashboards. Analysis of the included Python scripts and documentation found no malicious patterns, unauthorized data access, or remote execution risks. The tools operate locally using standard library functions.

  • Socketpass

    No alerts

  • Snykpass

    Risk: LOW · No issues

  • Runlayerfail

    7/13 files flagged

  • ZeroLeakspass

    Score: 93/100 · 2 sections analyzed

What does this agent skill do?

Observability Designer

Design production-ready observability strategies that combine the three pillars (metrics, logs, traces) with SLI/SLO frameworks, golden-signals monitoring, multi-window burn-rate alerting, and alert-noise optimization.

Core Capabilities

  • SLI/SLO frameworks — select SLIs from the golden signals, map them to Prometheus expressions, set SLO targets by criticality tier, and compute error budgets.
  • Burn-rate alerting — multi-window burn-rate rules with severity routing, hysteresis, suppression, and grouping to keep alert noise below 10%.
  • Dashboard design — Grafana specs following the Overview > Service > Component > Instance hierarchy, ≤7 panels per screen, role-based views (SRE/Dev/Exec/Ops).
  • Structured logging & tracing — JSON log format with correlation IDs, log-level discipline, and head/tail/adaptive trace sampling strategies.
  • Runbooks & validation — runbook template per critical alert; coverage validation that every T1 service has metrics, logs, traces, and a runbook.
  • Cost optimization — metric/log/trace retention tiers and cardinality management.

When to Use

  • Instrumenting a new or existing production service.
  • Defining SLOs and error budgets for a service tier.
  • Tuning alert rules or reducing alert fatigue / alert storms.
  • Designing Grafana dashboards or role-based views.
  • Choosing a trace sampling strategy or structured log schema.

Clarify First

Before designing the observability strategy, confirm these inputs. If any is unknown or vague, ASK — do not assume:

  • Service type & criticality tier — api / pipeline / storage / ML and T1–T3 (sets SLO targets, error-budget math, and slo_designer flags)
  • User-facing vs internal — determines which golden signals become SLIs and how alert severity is routed
  • Primary pain: alert noise vs coverage gaps — decides whether to optimize existing alerts or design new burn-rate rules
  • Dashboard audience — SRE / Dev / Exec / Ops sets the role-based panel layout and hierarchy

Stop rule: ask only the 2-3 that most change the output. If the user says "just draft it," proceed and list your assumptions at the top of the artifact.

Tools

ToolPurposeCommand
slo_designer.pyGenerate SLI/SLO framework, error budgets, and burn-rate alerts from a service definitionpython scripts/slo_designer.py --service-type api --criticality high --user-facing true
alert_optimizer.pyAnalyze alert configs for noise, coverage gaps, and duplicates; emit an optimization reportpython scripts/alert_optimizer.py --input alerts.json --analyze-only
dashboard_generator.pyProduce Grafana-compatible dashboard JSON with golden signals and role-based viewspython scripts/dashboard_generator.py --service-type api --name "Payment Service"

References

Load the reference that matches the task — keep this file lean and pull detail on demand:

  • references/slo-and-alerting.md — the 8-step workflow, SLI/SLO quick reference, error-budget math, burn-rate alert windows, alert classification, alert-fatigue prevention, and golden signals. Read when designing SLOs or alerts.
  • references/dashboards-logs-traces.md — dashboard design rules, structured log format, trace sampling strategies, the runbook template, a complete worked payment-service spec, and cost optimization. Read when building dashboards, logs, traces, or runbooks.
  • references/tools-integration-and-troubleshooting.md — full per-script flag/output reference, the systems integration table (Prometheus/Grafana/Jaeger/PagerDuty), the troubleshooting table, and success-criteria targets. Read when running the scripts or diagnosing failures.
  • references/slo_cookbook.md — a practical, in-depth cookbook for defining and operating Service Level Objectives. Read when you need detailed SLO methodology beyond the quick reference.
  • references/alert_design_patterns.md — a deep guide to effective alerting patterns and anti-patterns. Read when designing a complete alerting strategy.
  • references/dashboard_best_practices.md — comprehensive dashboard design-for-insight best practices. Read when building a dashboard system from scratch.

Scope & Limitations

Covers:

  • SLI/SLO framework design for request-driven, pipeline, storage, and ML services.
  • Multi-window burn-rate alert generation and alert noise optimization.
  • Grafana-compatible dashboard specification with role-based layouts (SRE, Developer, Executive, Ops).
  • Structured logging format, trace sampling strategy selection, and cost-optimization guidance.

Does NOT cover:

  • Infrastructure provisioning or Terraform/Helm configuration for Prometheus, Grafana, or Jaeger -- see ci-cd-pipeline-builder for deployment pipelines.
  • Incident response workflow orchestration or post-mortem facilitation -- see runbook-generator for runbook authoring.
  • Application Performance Management (APM) agent installation or vendor-specific SDK integration.
  • Security monitoring, SIEM rule design, or compliance audit logging -- see skill-security-auditor for security-focused analysis.

Integration Points

SkillIntegrationData Flow
runbook-generatorEvery burn-rate alert references a runbook; the runbook generator consumes alert definitions to scaffold investigation stepsAlert YAML --> runbook-generator --> Markdown runbook linked in alert annotations
ci-cd-pipeline-builderDeployment events feed into dashboard annotations and alert suppression windowsPipeline events --> Grafana annotations + Alertmanager silences
performance-profilerLatency SLI breaches trigger profiling; profiler results inform SLO target adjustmentsSLO burn-rate alert --> profiler invocation --> refined latency thresholds
database-designerDatabase SLIs (query latency, connection success rate, replication lag) align with schema-level health checksDB schema metadata --> SLI metric expressions for database-type services
tech-debt-trackerError budget depletion signals feed into tech debt prioritization as reliability investmentsError budget reports --> tech debt backlog items with SLO-linked severity
release-managerRelease readiness gates check remaining error budget before approving deploymentsError budget API --> release gate pass/fail decision

Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.

<a href="https://skillzs.dev/skills/borghei/claude-skills/observability-designer">View observability-designer on skillZs</a>