skillZs
★ LIVE SKILL TAGS ★
>>> LIVE SKILLS INDEX <<<
* OPEN SOURCE *
NO LOGIN, NO TRACKING
※ REAL INSTALL DATA ※
← back to all skills
vasilyu1983/ai-agents-public290 installs

software-backend

Builds backend services and APIs with durable defaults. Use when implementing REST, GraphQL, tRPC, or gRPC services with auth, queues, data, or observability.

How do I install this agent skill?

npx skills add https://github.com/vasilyu1983/ai-agents-public --skill software-backend
view source ↗

Is this agent skill safe to install?

  • Gen Agent Trust Hubfail

    The skill provides high-quality and comprehensive templates for backend development. However, it references an external documentation URL flagged as a potential phishing risk and includes instructions to execute a maintenance script from an external directory, which could be exploited in an uncontrolled environment.

  • Socketwarn

    5 alerts: gptSecurity, gptAnomaly

  • Snykpass

    Risk: LOW · No issues

  • Runlayerpass

    1/1 file flagged

What does this agent skill do?

Software Backend Engineering

Use this skill for backend service implementation and review: API boundaries, auth, data access, jobs, caching, observability, and production hardening. If the main question is platform selection, system topology, or API-contract design without implementation, hand off early.

Quick Reference

NeedDefault Direction
Public HTTP APIREST with explicit contracts and timeouts
Internal TS monorepo APItRPC when end-to-end type safety matters
High-throughput internal RPCConnect by default; plain gRPC only when you don't need browser/gRPC-Web compatibility
Complex client-shaped readsGraphQL
Relational dataPostgreSQL with migrations and pooling
Background workQueue plus idempotent handlers and DLQ policy
Browser authOIDC (authorization code with PKCE) plus httpOnly cookies
Service authprefer workload identity (cloud IAM, SPIFFE) over long-lived static credentials where the platform provides it; short-lived tokens or signed service credentials when it isn't available (multi-cloud, third-party, legacy)
Cachingexplicit TTLs and invalidation rules
Observabilitycorrelation IDs, traces, structured logs, saturation metrics

Route Elsewhere


Workflow

  1. Confirm the real constraint: latency, team skill, runtime, compliance, data model, or delivery speed.
  2. Choose the transport and framework based on that constraint, not on trend-chasing.
  3. Define the boundary:
    • request and response contracts
    • auth and authorization rules
    • error model
    • idempotency and rate limiting
  4. Define the data path:
    • schema and migrations
    • transaction boundaries
    • pooling and query budgets
    • cache and invalidation rules
  5. Define the async path:
    • queue semantics
    • retry ownership
    • deduplication and DLQ
  6. Add operability before calling it complete:
    • timeouts and cancellation
    • health checks
    • structured logs and traces
    • deploy and rollback expectations

Technology Selection

Pick based on the strongest operational constraint:

  • TypeScript-heavy team -> keep the repo's existing framework and ORM; for a new service, Fastify plus Drizzle (matches the Fastify/Drizzle template), NestJS plus Prisma for DI-heavy CRUD apps, Hono for edge or multi-runtime targets
  • audited SQL and predictable concurrency -> Go with sqlc/pgx
  • Python ecosystem or ML adjacency -> FastAPI plus SQLAlchemy
  • memory safety and explicitness -> Rust with Axum plus SQLx
  • edge or serverless first -> lightweight stateless handlers with hard CPU and timeout budgets

Use software-baas-platforms first when the real requirement is "ship auth, storage, and realtime quickly with less custom service code."


Backend Non-Negotiables

Mutation correctness contract.

For every externally retried mutation, define the idempotency scope, transaction boundary, authorization point, and observable terminal states before choosing transport or framework. Persist the idempotency record with the business write when possible; distinguish “not attempted,” “committed,” and “outcome unknown.” A retry must return the original result or resume safely, never duplicate the effect.

CategoryRule
APIMutating endpoints require idempotency keys where retries are plausible
APIList endpoints require explicit pagination (limit/cursor) and at least one filter
APIErrors are structured and machine-readable (RFC 9457 Problem Details)
APIHealth endpoints separate liveness (/healthz) from readiness (/readyz)
DataNo SELECT * on wide or high-volume paths
DataTransactions kept explicit; no implicit ambient transactions
DataNew or changed query plans verified with EXPLAIN ANALYZE before production
DataORM convenience layers bypassed on hot paths where auditability matters
DependenciesEvery outbound call has an explicit timeout; no framework-default infinite wait
DependenciesRetries owned at exactly one layer (no double-retry across client + service)
DependenciesCache invalidation rule documented before caching is added
DependenciesBackground jobs safe to retry and observable (structured log on start/finish/failure)
OperationsEvery request carries a correlation ID propagated to all downstream calls
OperationsTrace, log, and metric identifiers agree (no split identity)
OperationsSlow paths have explicit latency budgets (p99 target, not "fast enough")
OperationsDeploy procedure includes rollback step and smoke-check list

Multi-Tenancy Decision

Decide tenancy explicitly for any service that stores more than one customer's data; choose per tenant tier, not once for the whole product.

  • Model: shared schema + tenant_id enforced by row-level security (many small tenants; isolation depends on every query path) / schema-per-tenant (moderate counts, some customization; migrations fan out) / database-per-tenant (regulated or very large tenants, per-tenant restore or residency; highest ops cost).
  • Tenant context propagation: derive the tenant from the authenticated principal, never from a client-supplied field; carry it through handlers, jobs, and outbound calls; set it transaction-locally (pooler-safe) and fail closed when unset; scope cache keys, idempotency keys, storage prefixes, and search indexes by tenant.
  • Noisy neighbours: per-tenant rate limits, queue concurrency, and statement timeouts; give large tenants their own partition or limit. Prove isolation with a cross-tenant read test.

Rule: rules/common/security.md loads this invariant in every coding session.

Details, RLS SQL, and pooling traps: references/database-patterns.md.


Load and failure budget

For each downstream dependency, bound in-flight requests and queued work. When the bound is reached, reject new work with 429 or 503 and Retry-After where meaningful; do not accept work that cannot finish within the caller's deadline. Propagate cancellation, reserve time for the response, and shed lower-priority work before the critical path. Set one retry owner, a retry budget, exponential backoff with jitter, and an idempotency contract; test a slow or failed dependency under load to see whether p99, memory, and queue depth recover. For resilience test design and failure-budget review, load qa-resilience.

Job runtime decision

Use a PostgreSQL-backed queue when the service already operates PostgreSQL, job volume fits measured database capacity, and a short transaction can claim each job. Use a managed broker or Redis-backed worker when throughput, isolation, or operational ownership justifies another dependency. Use durable execution when a job must resume a multi-step workflow across crashes or long waits; check the runtime's retry and history semantics before choosing it. A database queue needs FOR UPDATE SKIP LOCKED, lease expiry, idempotent effects, poison-job handling, and queue-depth monitoring. See message queues and background jobs for the claim and outbox alternatives.


Performance and Reliability Triage

When a service is slow or unstable, debug in this order:

StepCheckSignal
1Query behavior and N+1sEXPLAIN output, ORM query log showing repeated identical queries
2Indexes and execution plansSeq scans on large tables, missing index on FK or filter columns
3Connection pooling and queue depthSustained pool wait time; idle connections exhausted
4Timeout and cancellation gapsRequests hanging past deadline; no context propagation through outbound calls
5Caching or read-shaping opportunitiesSame query with same result executing repeatedly in a short window; hot read path with no invalidation
6Runtime or tier limitsCPU throttling, memory pressure, rate limit headers from upstream

Do not add caching before you understand the real bottleneck.


Production Readiness Checklist

  • Every row in Backend Non-Negotiables holds for the changed surface
  • DLQ policy defined for every queue consumer (what happens to poison messages)

Navigation

Core references

Shared review utilities

Read one of these when the matching workflow step is active, not up front:

Templates

Load the one matching the Technology Selection choice made for this task, not the whole set:

Related Skills

Gate before invoking any foundation below: Each foundation has a When to Apply / When to Skip section. If your task matches a skip-condition, route to the foundation it names instead — don't pull in primitives the task doesn't need.

Learnings Loop

When prior decisions or pitfalls are relevant, consult learnings.consolidated.md if present; use learnings.md only for needed history or as the available fallback. Otherwise skip both.

After applying it, if you encountered a pattern worth remembering, a mistake worth preventing, or a domain fact that surprised you, append one dated bullet to learnings.md via agents-skills-feedback-loop/scripts/append_learning.py. Do not modify SKILL.md itself.

Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.

<a href="https://skillzs.dev/skills/vasilyu1983/ai-agents-public/software-backend">View software-backend on skillZs</a>