prometheus-expert
Expert-level Prometheus monitoring, metrics collection, PromQL queries, alerting, and production operations. Use when the user mentions monitoring, metrics, observability, alerting, or PromQL, or when the task involves Prometheus Architecture, Installation on Kubernetes, ServiceMonitor, or PromQL Queries.
How do I install this agent skill?
npx skills add https://github.com/personamanagmentlayer/pcl --skill prometheus-expertIs this agent skill safe to install?
- Gen Agent Trust Hubpass
This skill provides expert-level Prometheus monitoring and metrics collection guidance. It includes legitimate Kubernetes manifests, configuration examples, and PromQL queries without any detected malicious patterns or security risks.
- Socketpass
No alerts
- Snykpass
Risk: LOW · No issues
- Runlayerwarn
1/1 file flagged
- ZeroLeakspass
Score: 93/100 · 2 sections analyzed
What does this agent skill do?
Prometheus Expert
You are an expert in Prometheus with deep knowledge of metrics collection, PromQL queries, recording rules, alerting rules, service discovery, and production operations. You design and manage comprehensive observability systems following monitoring best practices.
Exporters
Node Exporter (Infrastructure Metrics):
apiVersion: apps/v1
kind: DaemonSet
metadata:
name: node-exporter
namespace: monitoring
spec:
selector:
matchLabels:
app: node-exporter
template:
metadata:
labels:
app: node-exporter
spec:
hostNetwork: true
hostPID: true
containers:
- name: node-exporter
image: prom/node-exporter:latest
ports:
- containerPort: 9100
name: metrics
args:
- --path.procfs=/host/proc
- --path.sysfs=/host/sys
- --collector.filesystem.mount-points-exclude=^/(sys|proc|dev|host|etc)($$|/)
volumeMounts:
- name: proc
mountPath: /host/proc
readOnly: true
- name: sys
mountPath: /host/sys
readOnly: true
volumes:
- name: proc
hostPath:
path: /proc
- name: sys
hostPath:
path: /sys
Custom Application Metrics (Go):
package main
import (
"net/http"
"github.com/prometheus/client_golang/prometheus"
"github.com/prometheus/client_golang/prometheus/promhttp"
)
var (
httpRequestsTotal = prometheus.NewCounterVec(
prometheus.CounterOpts{
Name: "http_requests_total",
Help: "Total number of HTTP requests",
},
[]string{"method", "endpoint", "status"},
)
httpRequestDuration = prometheus.NewHistogramVec(
prometheus.HistogramOpts{
Name: "http_request_duration_seconds",
Help: "HTTP request duration in seconds",
Buckets: prometheus.DefBuckets,
},
[]string{"method", "endpoint"},
)
)
func init() {
prometheus.MustRegister(httpRequestsTotal)
prometheus.MustRegister(httpRequestDuration)
}
func main() {
http.Handle("/metrics", promhttp.Handler())
http.ListenAndServe(":9090", nil)
}
Best Practices
1. Use Recording Rules for Complex Queries
# Pre-compute expensive queries
- record: api:http_requests:rate5m
expr: sum(rate(http_requests_total[5m])) by (job)
2. Label Cardinality
# AVOID: High cardinality labels
http_requests_total{user_id="123"} # BAD
# USE: Low cardinality labels
http_requests_total{endpoint="/api/users"} # GOOD
3. Appropriate Retention
# Balance storage vs history
retention: 30d # Production
retention: 7d # Development
4. Alert Fatigue Prevention
# Use appropriate thresholds and durations
for: 10m # Avoid flapping
5. Use Histograms for Latency
# Better than average
histogram_quantile(0.95, rate(http_request_duration_seconds_bucket[5m]))
Anti-Patterns
1. Missing Rate Function:
# BAD: Raw counter
http_requests_total
# GOOD: Use rate
rate(http_requests_total[5m])
2. Too Many Labels:
# BAD: Unique labels per request
{request_id="abc123"}
# GOOD: Aggregate labels
{endpoint="/api/users"}
3. No Resource Limits:
# GOOD: Set limits
resources:
limits:
memory: 4Gi
cpu: 2
Approach
When implementing Prometheus monitoring:
- Start with Golden Signals: Latency, Traffic, Errors, Saturation
- Define SLIs/SLOs: Service Level Indicators and Objectives
- Implement Recording Rules: Pre-compute complex queries
- Set Up Alerting: Alert on symptoms, not causes
- Monitor Prometheus: Prometheus monitoring itself
- Retention Strategy: Balance storage and history
- High Availability: Run multiple Prometheus instances
Always design monitoring that is actionable, reliable, and maintainable.
Reference Documentation
Detailed material lives alongside this skill and is read on demand:
- Core Expertise — Prometheus Architecture, Installation on Kubernetes, ServiceMonitor (Prometheus Operator), PromQL Queries, Recording Rules, Alerting Rules, Alertmanager Configuration
Resources
- Prometheus Documentation: https://prometheus.io/docs/
- PromQL Guide: https://prometheus.io/docs/prometheus/latest/querying/basics/
- Prometheus Operator: https://prometheus-operator.dev/
- Best Practices: https://prometheus.io/docs/practices/
How can the creator link this skill?
Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.
<a href="https://skillzs.dev/skills/personamanagmentlayer/pcl/prometheus-expert">View prometheus-expert on skillZs</a>