Skip to content

Advanced rypipe

This section is for adapter authors and power users who want to understand why rypipe is fast and how to keep it fast. It assumes you have read the Python API, Rust API, and Architecture pages.

Each page takes one optimization topic and explains the mechanism, the tuning knobs, and the trade-offs.

Roadmap

Page What you will learn
Fusion How RenameFields, DropFields, CastTypes, and FilterRows (keyword form or col() expression) are rewritten into a single ExecutionPlan; what is fusable and what falls back to Python.
Stage protocol The three methods every stage implements; why re-exporting works; when to re-implement; custom _to_spec() producers for fusable filters.
Execution modes columnar, parallel, stream, and parallel_streaming; the resolve_engine heuristic; when each mode wins.
Streaming Constant-memory batch streaming with iter_record_batches, BatchConsumer, and Arrow-ecosystem writers; when streaming falls back to materialization.
Memory and chunking How BoundedExecutor enforces a memory budget; sizing chunks for files larger or smaller than RAM.
Parallelism rayon internals; the num_chunks formula; why too many chunks hurt; measuring speedup.
Dictionary encoding Arrow dictionaries in rypipe; auto_dict heuristics; the incremental fast path vs the merge path; explicit dictionary_columns.
Schema and types Using schema and field_types to skip inference passes, stabilize column order, and enable typed compare filters.
I/O tuning mmap vs buffered reads; prefault; transparent gzip/zstd/lz4 decompression; storage class considerations.
Adapter design Writing a fast Splitter and RecordParser; the ColumnarSink fast-path protocol; borrowing strings; sparse rows; sink.wants.
Adapter design patterns One API with progressive overrides: override read() for simple adapters, _read_arrow() for advanced.
Profiling Profiling with perf, cargo flamegraph, and the bench_throughput example; measuring RSS; separating Python and Rust time.
Anti-patterns Common mistakes that silently remove fusion, increase memory, or waste CPU.
Case study: crxml How crxml reaches ~7.6 GB/s by combining the techniques from the other pages.