di-agent-knowledge-engine-datastage
Q&A reference for the DataStage parallel engine — parallelism, partitioning theory, APT configuration files, concurrent job execution, restart/recovery, disk/resource tuning, dataset performance, flow optimization (partitioning/sorting/memory), and per-stage semantics. Use for conceptual engine questions and stage property lookups regardless of authoring tool.
How do I install this agent skill?
npx skills add https://github.com/ibm/ibm-watsonx-data-integration-skills --skill di-agent-knowledge-engine-datastageIs this agent skill safe to install?
- Gen Agent Trust Hubpass
The skill is a comprehensive technical documentation and reference guide for the IBM DataStage parallel engine. It provides detailed information on stage properties, optimization strategies, and transformer functions. No security issues were detected; the skill accurately describes standard platform capabilities for ETL job design, including mechanisms for custom code integration and external program execution.
- Socketpass
No alerts
- Snykpass
Risk: LOW · No issues
What does this agent skill do?
DataStage Parallel Engine
When to Use DataStage
- Batch ETL processing of large data volumes
- Parallel processing across multiple nodes
- Complex transformations with high throughput requirements
- Integration with enterprise databases and file systems
- Data warehouse loading and CDC operations
Engine Characteristics
- Parallel processing: Divides data into partitions processed simultaneously
- Pipeline parallelism: Multiple stages process different data concurrently
- Scalable: Add nodes to increase throughput
- High performance: Optimized for large-scale data movement
Key Concepts
- Partitioning: Data divided across processing nodes
- Nodes: Physical or logical processing units
- Partitions: Subsets of data processed independently
- Configuration file: Defines nodes and resources
Performance Factors
- Job design (stage selection, partitioning, data flow)
- Configuration (node count, partition count, resources)
- Infrastructure (disk I/O, network bandwidth, CPU)
References
- Engine Details
- Concurrent Job Execution
- Configuration Management
- Data Set Performance
- Disk and Resource Optimization
- Restart and Recovery
- Flow optimization (partitioning, sorting, memory) → optimization/overview.md
- Per-stage semantics, requirements, best practices, and properties → stages/
- Transformer expressions
How can the creator link this skill?
Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.
<a href="https://skillzs.dev/skills/ibm/ibm-watsonx-data-integration-skills/di-agent-knowledge-engine-datastage">View di-agent-knowledge-engine-datastage on skillZs</a>