io-bound-data-processing
Processing, transforming, or moving datasets that may exceed RAM on a single low-compute box — covers memory discipline (streaming, generators, dtype shrinkage), I/O access patterns (sequential vs random, mmap, async), data formats (Parquet vs CSV vs JSON, predicate pushdown), chunking & batching, spill-to-disk (external merge sort, DuckDB/Polars), pipelining (bounded queues, backpressure, checkpointing), codec selection (zstd/lz4/gzip), concurrency for I/O-bound workloads (asyncio, threads, prefetch), and observability (iowait vs CPU%, rows/sec, py-spy/strace). Trigger on "process a large file", "stream this", "out-of-core", "OOM kill", "this is slow", or code with `pd.read_csv` of multi-GB files, `requests.get(...).content` on big bodies, `BytesIO` on unbounded inputs, per-row INSERTs, sequential `requests.get` loops, falling `tqdm` rates — even if I/O or memory isn't mentioned. Complement to computer-science-algorithms.
How do I install this agent skill?
npx skills add https://github.com/pproenca/dot-skills --skill io-bound-data-processingIs this agent skill safe to install?
- Gen Agent Trust Hubpass
This skill provides a comprehensive library of best practices and code examples for optimizing data processing tasks on resource-constrained hardware. It covers memory management, I/O patterns, data formats, and observability, utilizing industry-standard libraries like Pandas, Polars, and DuckDB.
- Socketpass
No alerts
- Snykpass
Risk: LOW · No issues
What does this agent skill do?
Community I/O-bound data processing on constrained resources Best Practices
A reference for engineers processing datasets larger than RAM on a single low-compute box. Organized by execution-lifecycle impact: rules near the top of the table govern whether the job runs at all; rules near the bottom shave the last 10 %. Optimize from the top of the waterfall.
Scope: the patterns that show up in real ETL / data-engineering / batch work on a laptop, a 2-vCPU container, or a Raspberry Pi-class node — streaming, formats, chunking, spill, backpressure, codecs, and the concurrency model that actually matches an I/O-bound bottleneck. Out of scope (covered elsewhere): the algorithmic primitives themselves (see computer-science-algorithms), distributed compute beyond a single box (use Spark/Dask), and database-engine internals (see official docs).
Distilled from Apache Arrow / Parquet docs, Polars User Guide, DuckDB docs, pandas — Scaling to large datasets, Linux man pages (mmap(2), sendfile(2), posix_fadvise(2)), Brendan Gregg's USE method and Systems Performance, Kleppmann's Designing Data-Intensive Applications, and the zstd / lz4 reference benchmarks.
When to Apply
Reach for these rules when:
- A job OOM-kills, swaps, or runs much slower than expected on a small box
- Input is larger than RAM and you need to scan, filter, aggregate, sort, or join it
- A pipeline has unbounded buffers between stages, or memory grows linearly during a "streaming" job
- You see one-row-per-RTT writes (
INSERTper row,requests.getper URL,f.read(32)per record) - You're picking a format/codec/serializer and the choice matters at scale
- A
topshows low CPU and high iowait, or you don't know which it is - "It's slow but I don't know why" — start at the obs- category
Rule Categories By Priority
| # | Category | Prefix | Impact | Why it cascades |
|---|---|---|---|---|
| 1 | Memory Discipline | mem- | CRITICAL | Slurping into RAM defeats every downstream technique on a constrained box |
| 2 | I/O Access Patterns | io- | CRITICAL | Disk/net are 10³–10⁶× slower than RAM; access pattern dominates wall-clock |
| 3 | Data Format & Encoding | fmt- | HIGH | Format fixes the lower bound on I/O volume + decode cost before any logic runs |
| 4 | Chunking & Batching | batch- | HIGH | Granularity controls peak memory and amortizes per-item overhead |
| 5 | Spill-to-Disk & External Memory | spill- | HIGH | When data > RAM, the choice is "spill cleanly" or "OOM" |
| 6 | Pipelining & Backpressure | pipe- | MEDIUM-HIGH | Unbounded buffers between fast producers and slow sinks = OOM |
| 7 | Compression & Serialization | codec- | MEDIUM | Trades CPU for I/O; right codec saves orders of magnitude |
| 8 | Concurrency for I/O-Bound Workloads | conc- | MEDIUM | Async / threads / processes are different tools; wrong model wastes CPU |
| 9 | Observability & Throughput Tuning | obs- | LOW-MEDIUM | Can't tune what you don't measure; iowait ≠ CPU-bound |
Quick Reference
1. Memory Discipline (CRITICAL)
mem-stream-dont-slurp— Iterate sources chunk-by-chunk; peak RAM = chunk size, not file sizemem-prefer-generators-over-lists-for-pipelines— Generators flow; lists materializemem-shrink-dtypes-before-loading— Narrow ints, categoricals; 2-8× memory reduction at load timemem-use-views-not-copies— Slicing without copying; NumPy/Arrow zero-copy semanticsmem-bound-the-working-set— Chunk size = budget ÷ row-size × amplification, not a round numbermem-release-references-explicitly— Drop intermediates so peak ≠ N × chunk
2. I/O Access Patterns (CRITICAL)
io-prefer-sequential-over-random— Sort offsets, advise the kernel, let readahead helpio-buffer-explicitly-for-small-records—BufferedReadercollapses 1000× syscallsio-stream-http-bodies-with-iter-content—stream=True+iter_contentinstead of.contentio-mmap-for-random-or-shared-large-files— Zero-copy + on-demand paging for random accessio-async-for-many-concurrent-streams— One thread, thousands of awaitsio-batch-and-pipeline-network-roundtrips—COPY, pipelines, multi-key endpoints, HTTP/2io-zero-copy-when-moving-bytes-as-is—sendfile,copy_file_range,shutil.copyfile
3. Data Format & Encoding (HIGH)
fmt-columnar-for-analytical-scans— Parquet/Arrow for filter+project workloadsfmt-line-delimited-for-streaming-row-ingest— NDJSON over JSON-array for streamingfmt-push-predicates-into-the-reader— Row-group statistics skip whole chunksfmt-prefer-schema-on-write-when-possible— Typed columns beat schema-on-read every timefmt-avoid-deeply-nested-json-for-hot-paths— Flat schema, or binary on hot pipes
4. Chunking & Batching (HIGH)
batch-pick-chunk-size-by-memory-budget— Compute from budget, not a constantbatch-use-vectorized-apis-not-row-loops— NumPy / Polars / Arrow kernels, notiterrowsbatch-process-with-stable-iterators—iter_batches,chunksize=,collect(streaming=True)batch-coalesce-writes-with-buffered-output—COPY/executemany, sized write buffersbatch-keyset-pagination-over-offset—WHERE id > $last_id, neverOFFSET Non deep cursors
5. Spill-to-Disk & External Memory (HIGH)
spill-external-merge-sort-when-data-exceeds-ram— Out-of-core sort; delegate tosort/ DuckDB when possiblespill-partition-by-hash-for-out-of-core-groupby-join— Hash-partition both sides; process per partitionspill-use-temp-files-not-process-memory—SpooledTemporaryFile, never unboundedBytesIOspill-use-engines-that-spill-automatically— DuckDB / Polars / Dask manage spill for you
6. Pipelining & Backpressure (MEDIUM-HIGH)
pipe-use-bounded-queues-for-producer-consumer— BoundedQueueis the backpressure mechanismpipe-apply-backpressure-from-slow-stages— Slow sink throttles the fast sourcepipe-prefer-pull-iteration-over-push-callbacks— Pull backpressures naturally; push needs policypipe-checkpoint-progress-for-resumability— Atomic checkpoint after each batch; redo bounded
7. Compression & Serialization (MEDIUM)
codec-zstd-or-lz4-as-defaults-not-gzip— Pick codec by access pattern; gzip is legacycodec-dictionary-encoding-for-repetitive-strings— 5-50× on low-cardinality columnscodec-prefer-binary-protocols-over-json-for-rpc— Protobuf / Arrow IPC / MsgPack on hot wirescodec-train-a-zstd-dictionary-for-many-small-payloads—zstd --trainfor sub-1 KB messages
8. Concurrency for I/O-Bound Workloads (MEDIUM)
conc-asyncio-for-many-network-streams-not-for-cpu— Asyncio multiplexes waits; useless for computeconc-thread-pools-for-blocking-io-libraries— GIL releases during blocking I/O; threads workconc-overlap-compute-with-prefetch— One batch ahead via background thread / coroutineconc-tune-parallelism-to-the-bottleneck— Match worker count to the binding resource
9. Observability & Throughput Tuning (LOW-MEDIUM)
obs-measure-iowait-not-just-cpu—iostat -x,vmstat, psutil — find the real bottleneckobs-instrument-throughput-rows-per-second— Rates normalize across runs; catch regressionsobs-profile-with-py-spy-or-strace-for-syscall-storms— Attach, don't guess
How to Use
Start with the question that matches the problem:
- "It OOMs / RAM grows linearly during a streaming job" →
mem-andpipe-(likely an unbounded buffer or per-iteration accumulation) - "Disk is 100 % busy but CPU is idle" →
io-(likely random access, unbuffered I/O, or wrong format) - "Why is reading this 10 GB file so slow?" → start with
fmt-columnar-for-analytical-scansandio-buffer-explicitly-for-small-records - "How big should the chunk be?" →
batch-pick-chunk-size-by-memory-budget - "Data doesn't fit in RAM" →
spill-(and prefer an engine that does it for you:spill-use-engines-that-spill-automatically) - "Lots of small network calls" →
io-batch-and-pipeline-network-roundtrips,io-async-for-many-concurrent-streams - "Adding workers didn't help" →
conc-tune-parallelism-to-the-bottleneck,obs-measure-iowait-not-just-cpu - "It's slow but I don't know why" →
obs-first; profile before changing anything
Code examples are in Python (most readable across audiences). The reasoning generalizes — equivalent libraries in other ecosystems (Arrow C++/Rust/Java, Polars Rust, DuckDB everywhere, libuv-style async) follow the same patterns.
Reference Files
| File | Description |
|---|---|
| references/_sections.md | Category definitions and ordering |
| assets/templates/_template.md | Template for adding new rules |
| metadata.json | Version and reference information |
| AGENTS.md | Auto-built TOC navigation |
Related Skills
computer-science-algorithms— Algorithmic primitives this skill builds on (external merge sort, sketches, hash partitioning, sampling)
How can the creator link this skill?
Add the canonical catalog link to the repository README so users can inspect current installs and available audits. The publishing guide covers the complete discovery path.
<a href="https://skillzs.dev/skills/pproenca/dot-skills/io-bound-data-processing">View io-bound-data-processing on skillZs</a>