skillZs
★ LIVE SKILL TAGS ★
>>> LIVE SKILLS INDEX <<<
* OPEN SOURCE *
NO LOGIN, NO TRACKING
※ REAL INSTALL DATA ※
← back to all skills
goldsky-io/goldsky-agent1.6k installs

turbo-pipelines

Turbo pipeline YAML reference and architecture guide. Covers: YAML field syntax (start_at, from, version, primary_key), source/transform/sink configuration, validation errors, resource sizing (xs–xxl), architecture decisions (dataset vs kafka, streaming vs job, fan-out vs fan-in, sink selection, pipeline splitting). Triggers on: 'what does field X do', 'what fields does a postgres sink need', 'what resource size', 'should I use kafka or dataset', 'how to structure my pipeline'. For writing transforms, use /turbo-transforms. For end-to-end building, use /turbo-builder.

How do I install this agent skill?

npx skills add https://github.com/goldsky-io/goldsky-agent --skill turbo-pipelines
view source ↗

Is this agent skill safe to install?

  • Gen Agent Trust Hubpass

    This skill provides a comprehensive guide for configuring and architecting Goldsky Turbo Pipelines, including YAML templates and troubleshooting steps. It adheres to security best practices by using the vendor's CLI for secret management and referencing official documentation and resources.

  • Socketpass

    No alerts

  • Snykpass

    Risk: LOW · No issues

  • Runlayerwarn

    4/9 files flagged

What does this agent skill do?

Turbo Pipeline Configuration & Architecture

YAML configuration reference and architecture guide for Turbo pipelines. For interactive pipeline building, use /turbo-builder. For troubleshooting, use /turbo-doctor. For transform implementation, use /turbo-transforms.

Decoded contract events are transform output, not a consumable <chain>.decoded_logs dataset. Source EVM contract logs from <chain>.raw_logs, then decode them in a transform (see /turbo-transforms). Pre-decoded token datasets such as <chain>.erc20_transfers, <chain>.erc721_transfers, and <chain>.erc1155_transfers are separate, available datasets; use /datasets to verify chain coverage.

CRITICAL: Always validate YAML with goldsky turbo validate <file.yaml> before showing complete pipeline YAML to the user or deploying.


Quick Start

name: my-first-pipeline
resource_size: s
sources:
  transfers:
    type: dataset
    dataset_name: base.erc20_transfers
    version: 1.2.0
    start_at: latest
transforms: {}
sinks:
  output:
    type: blackhole
    from: transfers
goldsky turbo validate pipeline.yaml   # Validate first
goldsky turbo apply pipeline.yaml -i   # Deploy + inspect

Prerequisites

  • Goldsky CLI — if goldsky is not on PATH, follow /auth-setup
  • Turbo extension — follow /auth-setup and run its bundled installer with turbo. Do not trigger the interactive auto-installer. Verify goldsky turbo --version succeeds; the published Linux binary requires x64 and glibc 2.39+, and the Mac binary requires Apple Silicon. Windows uses the WSL entry point. An unsupported platform is not a successful installation.
  • Shell PATH — restore export PATH="$HOME/.local/bin:$HOME/.goldsky/bin:$PATH" in every new shell/tool call; exports do not persist across calls.
  • Logged in — follow /auth-setup and run goldsky login yourself when goldsky project list is not authenticated
  • Secrets created for sinks if using PostgreSQL, ClickHouse, Kafka, etc. (see /secrets)

Architecture Decisions

Source Type Selection

ScenarioSource TypeWhy
Decode contract events from logsdatasetNeed raw_logs + _gs_log_decode()
Track token transfersdataseterc20_transfers has structured data
Historical backfill + livedatasetstart_at: earliest processes history
Live token balanceskafkalatest_balances_v2 is a streaming topic
Real-time state snapshotskafkaKafka delivers latest state continuously
Only need new data going forwardEitherDataset with start_at: latest or Kafka

Note: Kafka sources are used in production but are not documented in official Goldsky docs. Contact Goldsky support for topic names.

Data Flow Patterns

PatternWhen to UseTemplate
LinearSingle source, single destination, simple processingtemplates/linear-pipeline.yaml
Fan-outOne source → multiple sinks (different views/subsets)templates/fan-out-pipeline.yaml
Fan-inMultiple event types → UNION ALL → one tabletemplates/fan-in-pipeline.yaml
Multi-chainSame logic across chains (separate pipelines)templates/multi-chain-templated.yaml

For detailed pattern diagrams, YAML examples, and multi-chain deployment guidance, read references/architecture-patterns.md.

Resource Sizing

SizeWorkersCPUMemoryWhen to Use
xs—0.40.5 GiSmall datasets, light testing
s10.81.0 GiSimple filters, single source/sink, low volume (default)
m41.62.0 GiMultiple sinks, Kafka streaming, moderate transform complexity
l103.24.0 GiMulti-event decoding + UNION ALL, high-volume backfill
xl206.48.0 GiLarge chain backfills, complex JOINs
xxl4012.816.0 GiHighest throughput; up to 6.3M rows/min

Start small and scale up — defensive sizing avoids wasted resources.

resource_size sizes the pipeline, not the sink — they are independent. An l/xl pipeline pointed at an undersized destination produces the "deployed but writing nothing" shape: it validates, deploys, reports Running, and then fails every write. A Neon free tier (512 MB) will not hold a multi-month backfill of a high-volume chain — it errors could not extend file because project size limit (512 MB) has been exceeded and checkpoints time out. Check the destination's storage ceiling against the backfill scope (the start position plus any end bound) before deploying, not after.

Sink Selection

DestinationSink TypeBest For
Application DBpostgresRow-level lookups, joins, application serving
Real-time aggregatespostgres_aggregateBalances, counters, running totals via triggers
MySQL-compatible DBmysqlExisting MySQL stacks, upserts on primary key
Analytics queriesclickhouseLarge-scale aggregations, time-series data
Event processingkafkaDownstream consumers, event-driven systems
AWS messagingsqs_sinkDecoupled AWS consumers (standard queues only)
GCP messagingpubsubGoogle Cloud Pub/Sub topics (Turbo-only)
Serverless streamings2_sinkS2.dev streams, alternative to Kafka
NotificationswebhookLambda functions, API callbacks, alerts
Data lakes3_sinkLong-term archival, batch processing
TestingblackholeValidate pipeline without writing data

For full sink field specifications, see Sinks in the docs (one page per sink type).

Streaming vs Job Mode

ScenarioModeWhy
Real-time dashboardStreamingContinuous updates needed
Backfill 6 months of historyJobOne-time, stops when done
Real-time + catch-up on deployStreamingstart_at: earliest does backfill then streams
Export data to S3 onceJobNo need for continuous processing
Webhook notifications on eventsStreamingNeeds to react as events happen
Load test with historical dataJobProcess and inspect, then discard

Job mode rules: Runs to completion, auto-deletes ~1hr after finishing. Must delete before redeploying. Cannot pause/resume/restart.

Pipeline Splitting

Default to one pipeline with multiple sinks when the user wants the same source data sent to multiple destinations. A single Turbo pipeline supports multiple sinks natively — each sink has a from: field pointing at a source or transform by name, and sinks run independently (one failing does not block the others, and each can have its own batch_size / batch_flush_interval).

Do NOT split into separate pipelines just because there are multiple destinations. Generating one pipeline per sink duplicates the source ingestion, wastes resources, and decouples deployments that should be atomic.

Split into multiple pipelines only when: sources are fundamentally different (different chains with independent lifecycles), the destinations need different resource sizes, or the user explicitly asks for separate pipelines.

See references/architecture-patterns.md for fan-out (one source → multiple sinks) and fan-in (UNION ALL) YAML examples, and templates/multi-sink-pipeline.yaml / templates/fan-out-pipeline.yaml / templates/fan-in-pipeline.yaml for ready-to-adapt configs.


Configuration Reference

Pipeline Structure

name: my-pipeline          # Required: unique identifier (lowercase, hyphens)
resource_size: s            # Required: xs/s/m/l/xl/xxl
job: true                   # Optional: one-time batch (default: false = streaming)

sources:
  source_name:
    type: dataset           # or: kafka
    # ... source config

transforms:                 # Optional
  transform_name:
    type: sql               # or: script, handler, dynamic_table
    # ... transform config

sinks:
  sink_name:
    type: postgres           # or: clickhouse, kafka, webhook, s3_sink, etc.
    # ... sink config

Source Configuration

Dataset Source

sources:
  my_source:
    type: dataset
    dataset_name: <chain>.<dataset_type>
    version: <version>
    start_at: latest | earliest    # EVM, NEAR, Bitcoin, Stellar
    # start_block: <slot_number>   # Solana only — omit to start at the latest slot
    # end_block: <block_number>    # Solana only: bounded backfill (ignored on EVM)
    # filter: >-                   # Optional: SQL WHERE for source-level pre-filtering
    #   address = '0x...' AND block_number >= 10000000
FieldRequiredDescription
typeYesdataset for blockchain data
dataset_nameYesFormat: <chain>.<dataset_type>
versionYesDataset version (e.g., 1.2.0)
start_atEVM-familylatest or earliest; on Stellar also a ledger sequence number. Used by EVM, NEAR, Bitcoin, Stellar. Omitted = earliest = full chain history
start_blockNoSolana only: starting slot. Omitted = latest slot, i.e. no backfill
end_blockNoSolana only: stop at this slot. Silently ignored on EVM
filterNoSQL WHERE clause — pre-filters at ingestion (efficient)

Use filter for contract addresses and block ranges (coarse pre-filtering). Use transform WHERE for fine-grained filtering.

Set the start position explicitly on every dataset source, and say which one you chose. On the start_at chains (EVM, NEAR, Bitcoin, Stellar) omitting it does not mean "start now": the backend starts from the earliest available data, so the pipeline backfills the full chain history — days of replay and millions of rows before it reaches live data, and the sink has to hold all of it. A block number is not a valid start_at value; bound an EVM range with a block_number predicate in filter (pre-applied at the source), plus job: true for a one-shot backfill — end_block is silently ignored on EVM. Solana is the opposite default: it uses the numeric start_block, and omitting that starts at the latest slot, so a Solana backfill is always something you asked for — bound it with end_block or block_ranges.

For chain prefixes and dataset types, see /datasets.

Kafka Source

sources:
  my_source:
    type: kafka
    topic: base.raw.latest_balances_v2

No start_at or version fields. Optional: filter, include_metadata, starting_offsets.

Transform Configuration

TypeUse Case
sqlFiltering, projections, SQL functions
scriptCustom TypeScript/WASM logic
handlerCall external HTTP APIs to enrich
dynamic_tableLookup tables backed by a database
throttleCap throughput to a fixed records-per-second

SQL Transform

transforms:
  filtered:
    type: sql
    primary_key: id
    sql: |
      SELECT id, sender, recipient, amount
      FROM source_name
      WHERE amount > 1000
FieldRequiredDescription
typeYessql
primary_keyYesColumn for uniqueness/ordering
sqlYesSQL query (reference sources by name)
fromNoOverride default source (for chaining)

Transform Chaining

Chain transforms using from:

transforms:
  step1:
    type: sql
    primary_key: id
    sql: SELECT * FROM source WHERE amount > 100
  step2:
    type: sql
    primary_key: id
    from: step1
    sql: SELECT *, 'processed' as status FROM step1

For TypeScript, handler, and dynamic table transforms, see /turbo-transforms.

Throttle Transform

Caps the throughput of a stream by buffering records into batches and emitting each batch on a fixed minimum interval. Throttle does not modify data — every input record passes through unchanged. Use it to stay under rate limits of downstream sinks or external APIs, smooth bursty sources, or pace records into HTTP handlers.

transforms:
  throttled:
    type: throttle
    from: my_source
    max_batch_size: 100
    min_batch_interval: 10s
FieldRequiredDescription
typeYesthrottle
fromYesSource or transform to read from
max_batch_sizeNoMax records per batch
min_batch_intervalNoMinimum time between batches (e.g. 10s, 500ms, 1m)

Effective max throughput ≈ max_batch_size / min_batch_interval records per second. Throttle limits the maximum rate, not the minimum — if upstream is slow, batches will be smaller and arrive less frequently.

Place throttle close to the bottleneck (just before the rate-limited sink or handler) so upstream transforms still process at full speed.

Sink Configuration

Quick examples for common sinks. For full field specs of all sink types, see Sinks in the docs (one page per sink type).

primary_key placement depends on the sink type; it is not limited to transforms. A SQL transform's key describes its output, while a sink's key controls destination behavior:

Turbo sinkSink-level primary_key
PostgreSQLOptional; enables upserts. Omit for plain inserts.
MySQLOptional; enables upserts. Omit for plain inserts.
ClickHouseRequired; sets table ordering and deduplication columns.
KafkaOptional; selects message-key columns. If omitted, uses the upstream key when available.

For other sink types, check their own field reference before adding or removing primary_key. Mirror uses different sink schemas; do not copy Turbo sink fields into Mirror configs.

PostgreSQL

sinks:
  output:
    type: postgres
    from: my_transform
    schema: public
    table: my_table
    secret_name: MY_POSTGRES_SECRET
    primary_key: id

ClickHouse

sinks:
  output:
    type: clickhouse
    from: my_transform
    table: my_table
    secret_name: MY_CLICKHOUSE_SECRET
    primary_key: id

Pub/Sub (Turbo-only)

sinks:
  output:
    type: pubsub
    from: my_transform
    topic: my-topic
    secret_name: MY_PUBSUB_SECRET

The secret holds a GCP project id and service-account JSON. Topic must already exist in GCP.

Blackhole (Testing)

sinks:
  output:
    type: blackhole
    from: my_transform

Checkpoint Behavior

  • Preserved by default when updating a pipeline
  • Tied to source names — renaming a source resets its checkpoint
  • Tied to pipeline names — renaming the pipeline resets all checkpoints

To reset checkpoints: rename the source or pipeline. Warning: this reprocesses all historical data.


Starter Templates

TemplateDescriptionUse Case
minimal-erc20-blackhole.yamlSimplest pipeline, no credentialsQuick testing
filtered-transfers-sql.yamlFilter by contract addressUSDC, specific tokens
postgres-output.yamlWrite to PostgreSQLProduction data storage
pubsub-output.yamlWrite to Google Cloud Pub/SubGCP-based event consumers
multi-chain-pipeline.yamlCombine multiple chainsCross-chain analytics
solana-transfers.yamlSolana SPL tokensNon-EVM chains
multi-sink-pipeline.yamlMultiple outputsArchive + alerts + streaming
linear-pipeline.yamlSimple decode → filter → sinkBasic linear flow
fan-out-pipeline.yamlOne source → multiple sinksMulti-destination
fan-in-pipeline.yamlMultiple events → UNION ALLActivity feeds
multi-chain-templated.yamlPer-chain pipeline patternIndependent chain deploys

Template location: templates/ (relative to this skill's directory)


CLI Quick Reference

ActionCommand
Install Goldsky CLIFollow /auth-setup and use its bundled installer
Install Turbo extensionFollow /auth-setup and run its bundled installer with turbo; require a successful version check
Validate (REQUIRED)goldsky turbo validate pipeline.yaml
Deploy/Updategoldsky turbo apply pipeline.yaml
Deploy + Inspectgoldsky turbo apply pipeline.yaml -i
List pipelinesgoldsky turbo list
List datasetsgoldsky dataset list --output json
List secretsgoldsky secret list

For lifecycle commands (pause/resume/restart/delete) and monitoring (inspect/logs), see /turbo-operations.


Troubleshooting

See references/troubleshooting.md for CLI hanging, validation errors, and runtime errors.


Related

  • /turbo-builder — Interactive wizard to build pipelines step-by-step
  • /turbo-doctor — Diagnose and fix pipeline issues
  • /turbo-operations — Lifecycle commands and monitoring reference
  • /turbo-transforms — SQL, TypeScript, and dynamic table transform reference
  • /datasets — Dataset names and chain prefixes
  • /secrets — Sink credential management

Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.

<a href="https://skillzs.dev/skills/goldsky-io/goldsky-agent/turbo-pipelines">View turbo-pipelines on skillZs</a>