skillZs
★ LIVE SKILL TAGS ★
>>> LIVE SKILLS INDEX <<<
* OPEN SOURCE *
NO LOGIN, NO TRACKING
※ REAL INSTALL DATA ※
← back to all skills
pproenca/dot-skills111 installs

io-bound-data-processing

Processing, transforming, or moving datasets that may exceed RAM on a single low-compute box — covers memory discipline (streaming, generators, dtype shrinkage), I/O access patterns (sequential vs random, mmap, async), data formats (Parquet vs CSV vs JSON, predicate pushdown), chunking & batching, spill-to-disk (external merge sort, DuckDB/Polars), pipelining (bounded queues, backpressure, checkpointing), codec selection (zstd/lz4/gzip), concurrency for I/O-bound workloads (asyncio, threads, prefetch), and observability (iowait vs CPU%, rows/sec, py-spy/strace). Trigger on "process a large file", "stream this", "out-of-core", "OOM kill", "this is slow", or code with `pd.read_csv` of multi-GB files, `requests.get(...).content` on big bodies, `BytesIO` on unbounded inputs, per-row INSERTs, sequential `requests.get` loops, falling `tqdm` rates — even if I/O or memory isn't mentioned. Complement to computer-science-algorithms.

How do I install this agent skill?

npx skills add https://github.com/pproenca/dot-skills --skill io-bound-data-processing
view source ↗

Is this agent skill safe to install?

  • Gen Agent Trust Hubpass

    This skill provides a comprehensive library of best practices and code examples for optimizing data processing tasks on resource-constrained hardware. It covers memory management, I/O patterns, data formats, and observability, utilizing industry-standard libraries like Pandas, Polars, and DuckDB.

  • Socketpass

    No alerts

  • Snykpass

    Risk: LOW · No issues

What does this agent skill do?

Community I/O-bound data processing on constrained resources Best Practices

A reference for engineers processing datasets larger than RAM on a single low-compute box. Organized by execution-lifecycle impact: rules near the top of the table govern whether the job runs at all; rules near the bottom shave the last 10 %. Optimize from the top of the waterfall.

Scope: the patterns that show up in real ETL / data-engineering / batch work on a laptop, a 2-vCPU container, or a Raspberry Pi-class node — streaming, formats, chunking, spill, backpressure, codecs, and the concurrency model that actually matches an I/O-bound bottleneck. Out of scope (covered elsewhere): the algorithmic primitives themselves (see computer-science-algorithms), distributed compute beyond a single box (use Spark/Dask), and database-engine internals (see official docs).

Distilled from Apache Arrow / Parquet docs, Polars User Guide, DuckDB docs, pandas — Scaling to large datasets, Linux man pages (mmap(2), sendfile(2), posix_fadvise(2)), Brendan Gregg's USE method and Systems Performance, Kleppmann's Designing Data-Intensive Applications, and the zstd / lz4 reference benchmarks.

When to Apply

Reach for these rules when:

  • A job OOM-kills, swaps, or runs much slower than expected on a small box
  • Input is larger than RAM and you need to scan, filter, aggregate, sort, or join it
  • A pipeline has unbounded buffers between stages, or memory grows linearly during a "streaming" job
  • You see one-row-per-RTT writes (INSERT per row, requests.get per URL, f.read(32) per record)
  • You're picking a format/codec/serializer and the choice matters at scale
  • A top shows low CPU and high iowait, or you don't know which it is
  • "It's slow but I don't know why" — start at the obs- category

Rule Categories By Priority

#CategoryPrefixImpactWhy it cascades
1Memory Disciplinemem-CRITICALSlurping into RAM defeats every downstream technique on a constrained box
2I/O Access Patternsio-CRITICALDisk/net are 10³–10⁶× slower than RAM; access pattern dominates wall-clock
3Data Format & Encodingfmt-HIGHFormat fixes the lower bound on I/O volume + decode cost before any logic runs
4Chunking & Batchingbatch-HIGHGranularity controls peak memory and amortizes per-item overhead
5Spill-to-Disk & External Memoryspill-HIGHWhen data > RAM, the choice is "spill cleanly" or "OOM"
6Pipelining & Backpressurepipe-MEDIUM-HIGHUnbounded buffers between fast producers and slow sinks = OOM
7Compression & Serializationcodec-MEDIUMTrades CPU for I/O; right codec saves orders of magnitude
8Concurrency for I/O-Bound Workloadsconc-MEDIUMAsync / threads / processes are different tools; wrong model wastes CPU
9Observability & Throughput Tuningobs-LOW-MEDIUMCan't tune what you don't measure; iowait ≠ CPU-bound

Quick Reference

1. Memory Discipline (CRITICAL)

2. I/O Access Patterns (CRITICAL)

3. Data Format & Encoding (HIGH)

4. Chunking & Batching (HIGH)

5. Spill-to-Disk & External Memory (HIGH)

6. Pipelining & Backpressure (MEDIUM-HIGH)

7. Compression & Serialization (MEDIUM)

8. Concurrency for I/O-Bound Workloads (MEDIUM)

9. Observability & Throughput Tuning (LOW-MEDIUM)

How to Use

Start with the question that matches the problem:

Code examples are in Python (most readable across audiences). The reasoning generalizes — equivalent libraries in other ecosystems (Arrow C++/Rust/Java, Polars Rust, DuckDB everywhere, libuv-style async) follow the same patterns.

Reference Files

FileDescription
references/_sections.mdCategory definitions and ordering
assets/templates/_template.mdTemplate for adding new rules
metadata.jsonVersion and reference information
AGENTS.mdAuto-built TOC navigation

Related Skills

  • computer-science-algorithms — Algorithmic primitives this skill builds on (external merge sort, sketches, hash partitioning, sampling)

Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.

<a href="https://skillzs.dev/skills/pproenca/dot-skills/io-bound-data-processing">View io-bound-data-processing on skillZs</a>