skillZs
★ LIVE SKILL TAGS ★
>>> LIVE SKILLS INDEX <<<
* OPEN SOURCE *
NO LOGIN, NO TRACKING
※ REAL INSTALL DATA ※
← back to all skills
daymade/claude-code-skills663 installs

terraform-skill

Diagnoses and designs safe Terraform releases, provisioners (remote-exec/local-exec/file, cloud-init, Compose, Caddy), multi-environment isolation, and fresh-host bootstrap. Use when writing or reviewing plan/apply wrappers or provisioners; when staging/production config may differ; when a rollout can mutate a shared gateway; or debugging drift, TLS, container restarts, DNS duplication, or snapshot contamination.

How do I install this agent skill?

npx skills add https://github.com/daymade/claude-code-skills --skill terraform-skill
view source ↗

Is this agent skill safe to install?

  • Gen Agent Trust Hubwarn

    The skill facilitates Terraform release and provisioner management, involving significant command execution via SSH, Docker, and Terraform. A test script dynamically extracts and runs bash code from the documentation. It also performs credential verification via Cloudflare's API.

  • Socketpass

    No alerts

  • Snykpass

    Risk: LOW · No issues

  • ZeroLeakspass

    1 finding · Score: 86/100

What does this agent skill do?

Terraform Release and Provisioner Safety

Prevent a valid-looking Terraform workflow from publishing unvalidated bytes or widening a change's blast radius. Keep the user's business outcome and the actual mutation surface ahead of plan counts, green wrappers, or process completeness.

Route the task

Operating contract

  1. Use the repository's canonical release entry. When shared-gateway publication requires standard CI, agents may initiate that CI within existing authorization; normal release, restoration, rollback, and static-content transfer all use its controlled executor and existing wrappers. Do not substitute a local Terraform apply, root SSH/SCP, helper, or console write. Verify the implemented workflow and supported inputs before dispatch; lint or image-build CI alone is not a deployment lane. A missing lane blocks the live write, not preparation or read-only diagnosis.
  2. Prefer provider resources, image baking, cloud-init, or configuration management. HashiCorp recommends exhausting purpose-built alternatives because Terraform cannot model provisioner side effects predictably. When a provisioner remains necessary, make its artifact, target, lock, validation, and readback explicit.
  3. Inventory every resource or recovery tool that can write the target runtime. A validator attached to only one writer does not protect another writer of the same shared service.
  4. Give staging and production one required-key schema. Let values differ; never let a key be required in one environment and optional, defaulted, absent, or allowed-empty in another.
  5. Validate the exact candidate bytes, Compose-rendered environment, and immutable runtime image before the first live write or restart. Keep post-deploy checks too: they detect damage but cannot prevent the first bad mutation.
  6. Treat a saved plan as an executable artifact. Bind it to reviewed source/artifact identity and apply that exact file. A successful apply is not a staging receipt; record promotion evidence only after every required live verifier succeeds.
  7. Honor the user's production authorization, including standing authorization, at the last reversible point. A deadline, PLAN_DIGEST, CONFIRM_*, or agent inference does not supply it. Do not add personal signatures or per-run human approval when the authorized contract is standard CI.
  8. Stop after the requested result is verified. Do not turn a single-service fix into full-stack drift reconciliation, recovery redesign, or unrelated hardening.

Provisioner traps (symptom → fix)

Use these incident-derived symptom patterns to choose the next falsifying check. Confirm the current source and runtime before promoting a historical cause into the present diagnosis.

terraform plan says "No changes" but the site is down

Symptom: Browser ERR_SSL_PROTOCOL_ERROR or similar, terraform plan shows No changes, but the service is actually broken.

Interpretation: No changes is not a live-health check. Terraform refreshes provider-managed attributes, but a provisioner's remote files and container side effects are not automatically represented by those attributes. A broken site can follow an authorized apply, another IaC writer, or an out-of-band change; this symptom alone does not identify which happened. An HTTP redirect or healthy container does not prove the affected HTTPS route works.

Diagnosis:

  1. Probe the affected hostname and user path, then inspect the loaded route and live files.
  2. Compare the exact deployed source/artifact and state-tracked hashes, where present, with live hashes. Read the actual provisioner commands and release logs for every writer of that directory. Shared mtimes and deployed .tftpl/.static files are clues to inspect the archive and copy path, not proof of an out-of-band write: a normal archive extraction can produce them too.
  3. Reconstruct each writer's candidate file set and delete/exclude scope. In particular, a valid rsync --delete can remove another owner's route omitted from its candidate while retaining syntactically valid configuration. Check the complete existing hostname set, not just the main API.

Recovery: Use the repository's canonical release entry and its plan/apply wrapper; where standard CI is required, initiate the supported recovery or rollback through that same CI. Before a replacement, verify the candidate against the current live baseline and name exactly which files/resources may change. A frozen old commit can be reproducible and still roll back a newer route, policy, environment, or image. Restore only the authorized scope; stop if the current writer cannot preserve other owners' files. Do not use a broad replacement followed by a second sync as the default repair.

A saved plan freezes Terraform's planned values, not arbitrary files or external downloads read by its provisioners at apply time. Check that the actual implementation consumes the validated immutable artifact before claiming recovery will publish those bytes. After the authorized apply, independently verify the affected user path and the previously working routes. Live hash checks can detect drift; they do not by themselves establish its writer or prevent deletion.

Staging applies cleanly but production fails port is already allocated

Symptom: Compose deploy passes staging verification, then the production apply dies at docker compose up with Bind for 127.0.0.1:<port> failed: port is already allocated.

Root cause: Staging/production parity covers configuration, not host port allocation. A port that is free on staging can already be occupied by another service on the production host. A tainted null_resource is left behind, so the retry must go through the replace flow (CONFIRM_REPLACE), not a plain re-apply.

Diagnosis:

ssh root@<target-host> 'ss -ltnp | grep <port>'   # find the real occupant
docker ps --format "{{.Names}} {{.Ports}}" | grep <port>

Prevention: Before binding a host port in a module, check the target host for the exact bind at design time — never infer freeness from staging. Prefer uncommon high ports, record the allocation in the module comment and the deploy doc, and when a container needs a host-local probe target remember 127.0.0.1 inside a container is the container itself: host-local targets belong to host-level probes (systemd scripts), not containerized monitors.

docker: not found in remote-exec

cloud-init still installing Docker when provisioner SSHs in.

provisioner "remote-exec" {
  inline = [
    "cloud-init status --wait",
    "command -v docker >/dev/null || { echo 'FATAL: Docker not ready'; exit 1; }",
  ]
}

rsync: connection unexpectedly closed in local-exec

Do not infer a universal Terraform limitation from this symptom. A second SSH client can lose to the target's connection budget, SSH policy, or a competing deploy. Keep local-exec local: package an immutable artifact there, then use a Terraform-managed upload or a purpose-built deploy system. Give every apply a unique remote staging path; never share /tmp/src.tar.gz across concurrent applies.

provisioner "local-exec" {
  command = "tar czf /tmp/src-${self.id}.tar.gz --exclude=node_modules --exclude=.git -C ${path.module}/../../.. myproject"
}
provisioner "file" {
  source      = "/tmp/src-${self.id}.tar.gz"
  destination = "/tmp/src-${self.id}.tar.gz"
}
provisioner "remote-exec" {
  inline = ["tar xzf /tmp/src-${self.id}.tar.gz -C /data/ && rm -f /tmp/src-${self.id}.tar.gz"]
}

macOS BSD tar: --exclude must come BEFORE the source argument.

cloud-init status shows "running" forever

apt-get -y does not suppress debconf dialogs. Packages like iptables-persistent block on TTY prompts.

- |
    echo iptables-persistent iptables-persistent/autosave_v4 boolean true | debconf-set-selections
    echo iptables-persistent iptables-persistent/autosave_v6 boolean true | debconf-set-selections
    DEBIAN_FRONTEND=noninteractive apt-get install -y iptables-persistent

Known offenders: iptables-persistent, postfix, mysql-server, wireshark-common.

EACCES: permission denied in container logs, container Restarting

Host volume dirs are root-owned; container runs as non-root (uid 1001). Fix before docker compose up:

mkdir -p /data/myapp/data /data/myapp/logs
chown -R 1001:1001 /data/myapp/data /data/myapp/logs

Find UID: grep adduser.*-u or USER in Dockerfile.

Provisioner fails but no diagnostic output

Keep fail-fast behavior; attach diagnostics to failure instead of disabling set -e. Otherwise an early failed command can be overwritten by a later green health check.

provisioner "remote-exec" {
  inline = [
    "set -eu",
    "trap 'rc=$?; if [ $rc -ne 0 ]; then docker logs myapp --tail 20 2>&1 || true; docker ps --format \\\"table {{.Names}}\\\\t{{.Status}}\\\" || true; fi; exit $rc' EXIT",
    "docker compose up -d",
    "sleep 15",
    "docker ps --filter name=myapp --format '{{.Status}}' | grep -q healthy || exit 1",
  ]
}

Failed to save state / Failed to persist state to backend after apply

The apply ran and only the final backend upload failed (a timeout to the state bucket is typical). First confirm that in the apply log: every resource reports Creation complete, Modifications complete or Destruction complete, and the only errors are these two. Then every side effect already happened — provisioners ran, remote writes landed, containers were recreated — but the backend still holds the previous state, and Terraform wrote the new one to errored.tfstate in the directory it ran in (for a wrapper, its root module directory, not your shell's). Do not re-run apply: it would work from the stale state, which Terraform itself warns forks the state, and resources the failed run already created or replaced would be created or replaced again.

Run the recovery from that directory, through the repository's wrapper when it passes state subcommands through (operating contract 1); otherwise in exactly the environment the wrapper sets up, so the same backend and credentials are used. Keep the pulled copies out of the repository:

terraform state pull > "${TMPDIR:-/tmp}/remote.json"
jq '{file: input_filename, lineage, serial}' "${TMPDIR:-/tmp}/remote.json" errored.tfstate
# expect: identical lineage, and the remote serial lower than errored.tfstate's
terraform state push errored.tfstate
terraform state pull | jq '{lineage, serial}'
# expect: serial higher than the remote's before the push (it need not equal
# errored.tfstate's; the backend may increment it). Then check the resources the
# apply created or replaced: `terraform state show <address>` carries the IDs in the apply log.

state push refuses an unrelated lineage, a newer remote serial, and an equal serial with different content (observed on Terraform 1.5.7 with a local backend). A refusal means the remote no longer matches this errored.tfstate — another writer changed it, the file is left over from an older run, or you are pointed at a different workspace or backend: stop and reconcile. Never add -force; it overwrites whatever the remote holds. If the push fails on the same transport error as the apply, retry it later; errored.tfstate stays usable until the remote serial moves. After a successful push, move errored.tfstate out of the working directory (to scratch space) so it is not later read as live state; it and the pulled copies hold the full state, secrets included, so delete them once the read-back passes. Then read the wrapper's recipe and run by hand every step that follows its apply command — persisting the new setting, post-apply verification. The wrapper's failure exit hides that the remote change is already live.

Container Restarting — database tables missing

DB migrations not in provisioner. PostgreSQL docker-entrypoint-initdb.d only runs on empty data dir. Explicitly create DB + run migrations:

# After postgres healthy:
docker exec pg psql -U postgres -tc "SELECT 1 FROM pg_database WHERE datname='mydb'" | grep -q 1 \
  || docker exec pg psql -U postgres -c "CREATE DATABASE mydb;"

# Idempotent migrations:
for f in migrations/*.sql; do
  VER=$(basename $f)
  APPLIED=$($PSQL -tAc "SELECT 1 FROM schema_migrations WHERE version='$VER'" | tr -d ' ')
  [ "$APPLIED" = "1" ] && continue
  { echo 'BEGIN;'; cat $f; echo 'COMMIT;'; } | $PSQL
  $PSQL -tAc "INSERT INTO schema_migrations(version) VALUES ('$VER') ON CONFLICT DO NOTHING"
done

Compose uses an unexpected value despite .env

Compose interpolation gives the invoking shell higher precedence than --env-file or project .env. An old exported value can therefore override the reviewed environment silently. Inspect what Compose actually used; unset ambient overrides when the env file is meant to be authoritative.

# Inspect interpolation inputs and the rendered model.
docker compose --env-file .env config --environment
docker compose --env-file .env config --format json > compose.rendered.json

# Make the reviewed env file authoritative for this key.
env -u DOCKER_WITH_PROXY_MODE docker compose --env-file .env build

TLS handshake fails: Invalid format for Authorization header

Caddy's Cloudflare DNS module expects a scoped API Token through Bearer authentication. Do not infer credential type, validity, or permissions from length/prefix alone. Verify the token with Cloudflare's official endpoint, then exercise the exact zone operation or provider path required by the release.

curl -fsS https://api.cloudflare.com/client/v4/user/tokens/verify \
  -H "Authorization: Bearer $CLOUDFLARE_API_TOKEN" \
  | jq -e '.success == true and .result.status == "active"' >/dev/null

If the credential is absent or wrong, create a least-privilege API Token through Cloudflare's current dashboard/API flow and grant only the zones/operations the provider needs. Follow the official creation contract rather than copying permission-group IDs that may drift: https://developers.cloudflare.com/fundamentals/api/get-started/create-token/.

TLS fails on staging but works on production — hardcoded domains

Caddyfile or compose has literal domain names. Staging Caddy loads production config, tries to get certs for domains it doesn't own → ACME fails.

Caddyfile: Use {$VAR} — Caddy evaluates env vars at startup.

# WRONG
example.com { tls { dns cloudflare {env.CLOUDFLARE_API_TOKEN} } }

# RIGHT
{$LOBEHUB_DOMAIN} { tls { dns cloudflare {env.CLOUDFLARE_API_TOKEN} } }

Compose: Use ${VAR:?required} — fail-fast if unset or empty.

# WRONG
- APP_URL=https://example.com

# RIGHT
- APP_URL=${APP_URL:?APP_URL is required}

Pass the env var to the gateway container so Caddy can read it:

environment:
  - LOBEHUB_DOMAIN=${LOBEHUB_DOMAIN:?LOBEHUB_DOMAIN is required}
  - CLOUDFLARE_API_TOKEN=${CLOUDFLARE_API_TOKEN:?required for DNS-01 TLS}

Do not stop at this local assertion. Put all runtime-required keys in one schema, require the same set from every environment file, render the exact Compose service environment, and run the exact deployed Caddy image with that full environment before mutating live files. Caddy {$VAR} expansion can become an empty token before parsing; a Caddyfile default is not an environment-completeness check.

OAuth login fails: Social sign in failed

Casdoor init_data.json contains hardcoded redirect URIs. --createDatabase=true only applies init_data on first-ever DB creation — not on restarts. Fix via SQL in provisioner:

# Replace production domain with staging in existing Casdoor DB
$PSQL -c "UPDATE application SET redirect_uris = REPLACE(redirect_uris,
  'example.com', 'staging.example.com')
  WHERE name='lobechat'
  AND redirect_uris LIKE '%example.com%'
  AND redirect_uris NOT LIKE '%staging.example.com%';"

Also check AUTH_CASDOOR_ISSUER — it must match the Casdoor subdomain (auth.staging.example.com), not the app root domain.

Multi-environment isolation

Before creating a second environment, grep .tf files for hardcoded names. See references/multi-env-isolation.md for the complete matrix.

Environment isolation does not mean configuration-contract drift. Keep one required-key manifest and the same validation path for every environment. Staging and production may use different domains, credentials, instance sizes, and feature values; they must not disagree about whether a runtime key is required, optional, allowed-empty, or silently defaulted.

Will fail on apply (globally unique):

ResourceScopeFix
SSH key pairRegion"${env}-deploy"
SLS log projectAccount"${env}-logs"
CloudMonitor contactAccount"${env}-ops"

DNS duplication trap: Two environments creating A records for the same name in the same Cloudflare zone → two independent record IDs → DNS round-robin → ~50% traffic to wrong instance. Fix: use subdomain isolation (staging.example.com) or separate zones. Remember to create DNS records for ALL subdomains Caddy serves (e.g., auth.staging, minio.staging).

Snapshot cross-contamination: Unfiltered data "alicloud_ecs_snapshots" returns ALL account snapshots. New env inherits old 100GB snapshot, fails creating 40GB disk. Gate with variable:

locals {
  latest_snapshot_id = var.enable_snapshot_recovery && length(local.available_snapshots) > 0
    ? local.available_snapshots[0].snapshot_id : null
}

Do NOT add count to the data source — changes its state address, causes drift.

Pre-deploy validation

Before approval or mutation, read and execute pre-deploy-validation.md. It owns the publisher gates, validation phases, fresh-plan authorization and target-bound promotion evidence.

Zero-to-deployment

Fresh disks expose every implicit dependency. See references/zero-to-deploy-checklist.md.

Before an authorized service pause or first live write, also use pre-deploy-validation.md for credential capability and publisher preparation. The fresh-host checklist owns bootstrap dependencies, ordering and memory.

Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.

<a href="https://skillzs.dev/skills/daymade/claude-code-skills/terraform-skill">View terraform-skill on skillZs</a>