skillZs
★ LIVE SKILL TAGS ★
>>> LIVE SKILLS INDEX <<<
* OPEN SOURCE *
NO LOGIN, NO TRACKING
※ REAL INSTALL DATA ※
← back to all skills
stahura/domo-ai-vibe-rules154 installs

domo-data-generator

**Generating sample data for Domo** -- invoke when a user needs to create realistic sample datasets and upload them to a Domo instance. Primary signals: requests for sample data, demo data, test data, fake data for Domo; mentions of Salesforce, Google Analytics, QuickBooks, NetSuite, Google Ads, Facebook Ads, HubSpot, Marketo, or Health Portal sample data; questions about the datagen CLI or domo_data_generator. Covers: generating datasets, uploading to Domo, creating datasets in Domo, rolling dates, entity pools, connector icons, catalog management, and adding new dataset definitions. Skip for: real connector setup, production data pipelines, data transformations (Magic ETL), or Domo App Platform.

How do I install this agent skill?

npx skills add https://github.com/stahura/domo-ai-vibe-rules --skill domo-data-generator
view source ↗

Is this agent skill safe to install?

  • Gen Agent Trust Hubpass

    The skill installs a data generation tool from a third-party GitHub repository and provides examples for automated execution via cron. It also processes local YAML configuration files which represents a potential surface for indirect prompt injection.

  • Socketwarn

    1 alert: gptSecurity

  • Snykwarn

    Risk: MEDIUM · 1 issue

  • ZeroLeakspass

    Score: 93/100 · 2 sections analyzed

What does this agent skill do?

Domo Sample Data Generator

Generate realistic, cross-referenced sample data for Domo using the datagen CLI.

Repository: https://github.com/brrink/domo_data_generator


Overview

The generator creates sample data mirroring major business platforms with consistent cross-source entity integrity, then uploads it to Domo. It includes:

  • 18 pre-built datasets across 6 source categories (Salesforce, Google Analytics, Financial, Marketing, Health, AdPoint)
  • YAML-driven catalog for easy dataset additions
  • Shared entity pool (companies, people, products, sales reps, campaigns)
  • Date rolling to keep data looking current
  • Direct Domo integration (create datasets, upload, set connector icons)
  • Structured JSON output by default (AI-agent-friendly)
  • pipx installable -- runs from any directory

Setup

# Install globally with pipx (recommended)
pipx install git+https://github.com/brrink/domo_data_generator.git

# Initialize a working directory
mkdir my-domo-data && cd my-domo-data
datagen init

# Edit .env with your Domo credentials

If .env.example is missing or you want a clean start, create .env in the working directory with:

cat > .env <<'EOF'
DOMO_CLIENT_ID=your_client_id_here
DOMO_CLIENT_SECRET=your_client_secret_here
DOMO_API_HOST=api.domo.com
DOMO_INSTANCE=your_instance_name
DOMO_SET_CONNECTOR_TYPE=false
EOF

Required Environment Variables

VariablePurpose
DOMO_CLIENT_IDOAuth client identifier
DOMO_CLIENT_SECRETOAuth client secret
DOMO_API_HOSTAPI endpoint hostname
DOMO_INSTANCEDomo instance name
DOMO_SET_CONNECTOR_TYPEEnable connector icon customization (optional, default: false)

Auth boundary note: domo_data_generator uses its own public-API/OAuth credential flow and does not run through community-domo-cli or ryuu session auth.

Current tooling boundary: most Product API automation should use community-domo-cli, but datagen dataset create/upload in this skill currently depends on python -m datagen with .env OAuth credentials (DOMO_CLIENT_ID / DOMO_CLIENT_SECRET).


CLI Reference

Entry point: datagen [OPTIONS] COMMAND [ARGS]

Global Options

OptionDescription
--verbose / -vEnable verbose logging
--output / -o TEXTOutput format: json (default), table, yaml
--yes / -ySkip confirmation prompts

All commands emit structured JSON by default for easy machine parsing.

Init Command

init -- Initialize a working directory

datagen init                  # Initialize current directory
datagen init /path/to/dir     # Initialize a specific directory

Copies bundled catalog YAML files to ./catalog/, creates .env template, and creates ./data/ directory. Run this once before using the CLI in a new directory.

Core Commands

generate -- Generate sample data

datagen generate --all                    # Generate all datasets
datagen generate salesforce_opportunities  # Generate one dataset
datagen generate --all --seed 42          # Reproducible generation
datagen generate --all --dry-run          # Preview without writing

Requires entity pool initialization first. Run python -m datagen pool regenerate before generate even if your schema has no explicit entity_ref columns.

OptionDescription
nameDataset name (YAML filename stem), optional
--allGenerate all datasets
--seed INTEGERRandom seed for reproducibility
--catalog-dir PATHCatalog directory override
--data-dir PATHData directory override
--dry-runPreview without writing files

upload -- Upload data to Domo (full replace)

datagen upload --all
datagen upload salesforce_opportunities

Requires DOMO_CLIENT_ID and DOMO_CLIENT_SECRET.

OptionDescription
nameDataset name, optional
--allUpload all datasets
--catalog-dir PATHCatalog directory override
--data-dir PATHData directory override

create-dataset -- Create dataset(s) in Domo from catalog

datagen create-dataset --all --skip-existing
datagen create-dataset salesforce_opportunities

Requires DOMO_CLIENT_ID and DOMO_CLIENT_SECRET. The domo_id is persisted locally (in the catalog YAML if writable, otherwise in data/domo_ids.json).

OptionDescription
nameDataset name, optional
--allCreate all datasets
--skip-existingSkip datasets that already have a domo_id
--catalog-dir PATHCatalog directory override

roll-dates -- Shift rolling date columns to stay current

datagen roll-dates
datagen roll-dates --anchor-date 2026-04-01
OptionDescription
--anchor-date TEXTTarget date (YYYY-MM-DD), defaults to today
--catalog-dir PATHCatalog directory override
--data-dir PATHData directory override

Informational Commands

list -- List catalog dataset definitions

datagen list                    # JSON output (default)
datagen --output table list     # Rich table for humans
datagen list --verbose          # Include column/schema details

status -- Display generation status for all datasets

datagen status

Connector Icon Commands

Require DOMO_DEVELOPER_TOKEN and DOMO_INSTANCE.

discover-types -- Search Domo connector/provider types

datagen discover-types salesforce

set-type -- Set connector icon on a Domo dataset

datagen set-type salesforce_opportunities
datagen set-type salesforce_opportunities --provider-key custom_key

set-type-all -- Set connector icon on all datasets with a domo_id

datagen set-type-all

Entity Pool Commands

pool regenerate -- Regenerate the shared entity pool

datagen pool regenerate
datagen pool regenerate --seed 99
datagen pool regenerate --company-count 500 --person-count 1000
OptionDefault
--seed INTEGER42
--company-count INTEGER200
--person-count INTEGER500
--product-count INTEGER50
--sales-rep-count INTEGER20
--campaign-count INTEGER30

pool show -- Display entity pool summary

datagen pool show

Common Workflows

Full setup for a new Domo instance

datagen init
# Edit .env with credentials
datagen pool regenerate
datagen generate --all
datagen create-dataset --all
datagen upload --all
datagen set-type-all

Daily refresh via cron

# Crontab entry: roll dates and re-upload daily at 6 AM
0 6 * * * cd /path/to/project && datagen roll-dates && datagen upload --all

Generate a single dataset end-to-end

datagen generate salesforce_opportunities
datagen create-dataset salesforce_opportunities
datagen upload salesforce_opportunities
datagen set-type salesforce_opportunities

Included Datasets

CategoryDataset NameKeyRows
SalesforceSalesforce - Accountssalesforce_accounts500
SalesforceSalesforce - Contactssalesforce_contacts1,500
SalesforceSalesforce - Opportunitiessalesforce_opportunities2,500
Google AnalyticsGoogle Analytics - Sessionsga_sessions5,000
Google AnalyticsGoogle Analytics - Page Viewsga_pageviews10,000
FinancialQuickBooks - Invoicesfinancial_invoices3,000
FinancialNetSuite - General Ledgerfinancial_gl_entries5,000
MarketingGoogle Ads - Campaign Performancemarketing_google_ads3,000
MarketingFacebook Ads - Campaign Performancemarketing_facebook_ads2,500
MarketingHubSpot - Contactsmarketing_hubspot_contacts2,000
MarketingMarketing - Market Leadsmarketing_market_leads2,500
MarketingMarketo - Leadsmarketing_marketo_leads3,000
HealthHealth Portal - Demographicshealth_demographics15
HealthHealth Portal - Lab Resultshealth_lab_results1,470
HealthHealth Portal - Vitalshealth_vitals5,250
AdPointAdPoint - Ordersadpoint_orders150
AdPointAdPoint - Line Itemsadpoint_line_items500
AdPointAdPoint - Flightsadpoint_flights2,000

Entity Pool

The shared entity pool provides consistent cross-dataset references. Entities are generated once and reused across all datasets.

Entity TypeDefault CountKey Fields
company200id, account_id, name, domain, industry, size, city, state, annual_revenue, employee_count
person500id, contact_id, first_name, last_name, full_name, email, company_id, company_name, title, phone
product50id, name, category, unit_price, sku
sales_rep20id, rep_id, first_name, last_name, full_name, email, region
campaign30id, name, channel, budget, status

Adding New Dataset Definitions

Dataset definitions live in the catalog/ directory as YAML files. Each YAML file defines metadata, columns, and generator configurations.

YAML Structure

dataset:
  name: My Custom Dataset
  domo_id: null
  source_type: custom
  description: "Description of the dataset"
  row_count: 1000
  tags:
    - custom
    - demo

schema:
  - name: id
    type: STRING
    generator: uuid4

  - name: company_name
    type: STRING
    generator: entity_ref
    entity: company
    field: name

  - name: amount
    type: DOUBLE
    generator: random_decimal
    min: 100.0
    max: 10000.0
    precision: 2

  - name: created_date
    type: DATE
    generator: date_range
    start_days_ago: 365
    end_days_ahead: 0
    rolling: true

Available Column Types

STRING, LONG, DOUBLE, DECIMAL, DATETIME, DATE

Available Generators

Generic: uuid4, random_choice, weighted_choice, random_int, random_decimal, date_range, entity_ref, compound, sequence, constant, derived_from_date, stage_derived, faker

Salesforce: sf_id, sf_opportunity_name, sf_case_subject, sf_lead_rating

Google Analytics: ga_session_id, ga_page_path, ga_source, ga_medium, ga_campaign, ga_browser, ga_device_category, ga_country, ga_bounce_rate, ga_session_duration, ga_pageviews, ga_landing_page

Financial: gl_account_code, gl_account_name, invoice_number, payment_terms, payment_method, invoice_status, journal_type, department, fiscal_period, debit_credit

Marketing/Ads: ad_platform, campaign_objective, ad_format, ad_headline, ad_keyword, targeting_type, impressions, clicks_from_impressions, ctr, cost_per_click, ad_spend, conversions_from_clicks, hubspot_lifecycle, hubspot_lead_status, ad_group_id

Health: health_lab_init, health_lab_field, health_vital_init, health_vital_field, health_demographics

Generator Column Options

OptionUsed WithDescription
entityentity_refEntity pool type to reference
fieldentity_refField to pull from the entity
choicesrandom_choice, weighted_choiceList of possible values
min / maxrandom_int, random_decimalValue range
precisionrandom_decimalDecimal places
templatecompoundString template with {field} placeholders
refscompoundColumn references for template substitution — must be a YAML list of column name strings (e.g. ["sku", "line_id"]), not a dict/object. A dict triggers ValidationError: schema.N.refs — Input should be a valid list.
start_days_ago / end_days_aheaddate_rangeDate range relative to today
rollingdate_rangeEnable date rolling for freshness
mappingstage_derivedMap source values to derived values
source_columnstage_derived, derived_from_dateColumn to derive from
formatderived_from_dateDate format string
faker_methodfakerFaker library method name
faker_argsfakerArguments for the Faker method

weighted_choice YAML format:

generator: weighted_choice
choices:
  "Tier 1": 0.40
  "Tier 2": 0.35
  "Tier 3": 0.25

compound refs vs formatted random strings: For values like PO-12345, prefer faker with bothify instead of abusing compound / refs:

- name: purchase_order_ref
  type: STRING
  generator: faker
  faker_method: bothify
  faker_args:
    text: "PO-#####"

Rules

  1. Run datagen init first -- Initialize a working directory before using any other commands. This copies the catalog and creates .env.
  2. Always generate before uploading -- Run generate (or generate --all) before upload to ensure CSV data files exist.
  3. Create datasets before first upload -- Run create-dataset before upload for new datasets. The domo_id is persisted locally.
  4. Use --skip-existing -- When running create-dataset --all, use --skip-existing to avoid duplicating datasets that already have a domo_id.
  5. Entity pool consistency -- Regenerating the pool (pool regenerate) invalidates all previously generated data. Re-generate all datasets afterward.
  6. Date rolling -- Use roll-dates before upload to keep date columns current. Only columns with rolling: true are affected.
  7. Credentials -- DOMO_CLIENT_ID and DOMO_CLIENT_SECRET are required for upload and create-dataset. DOMO_DEVELOPER_TOKEN is required for set-type and discover-types. Offline commands need no credentials.
  8. Reproducibility -- Use --seed for reproducible data generation across runs.
  9. Output format -- Default output is JSON. Use --output table for human-readable Rich tables.

Checklist

  • CLI installed (pipx install git+https://github.com/brrink/domo_data_generator.git)
  • Working directory initialized (datagen init)
  • .env configured with Domo credentials
  • Entity pool generated (datagen pool regenerate)
  • Datasets generated (datagen generate --all)
  • Datasets created in Domo (datagen create-dataset --all --skip-existing)
  • Data uploaded (datagen upload --all)
  • Connector icons set (datagen set-type-all) if desired
  • Cron configured for daily date rolling and upload if needed

Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.

<a href="https://skillzs.dev/skills/stahura/domo-ai-vibe-rules/domo-data-generator">View domo-data-generator on skillZs</a>