pandas¶
to_pandas is rypipe's primary Python sink. It converts a pipeline into
a pandas DataFrame, Arrow-backed by default.
Setup¶
The all extra includes pandas and pyarrow.
Basic usage¶
from crxml import CrystalXMLSource, CastTypes, to_pandas
df = to_pandas(
CrystalXMLSource("report.xml", row_tag="Details")
| CastTypes({"Amount": float})
)
df.groupby("Department")["Amount"].sum()
Dtype backends¶
By default, columns use Arrow-backed dtypes (pd.ArrowDtype), which
preserve nulls and avoid the object-dtype fallback of NumPy conversion.
Pass dtype_backend="numpy" for classic NumPy-backed columns:
Chunked construction¶
Two knobs control how the DataFrame is built:
chunksize=splits the finished Arrow table into batches and concatenates DataFrames incrementally; useful whentable.to_pandas()on the full table spikes memory.memory=(e.g."64MiB") uses the adapter's streaming reader viaiter_record_batches, so parsing itself stays within the budget.
All output remains in memory; the budget bounds the parsing side only.
Why this works¶
pandas ≥ 2.0 converts Arrow tables efficiently and can back columns with Arrow memory directly. rypipe's parser already produces Arrow, so the conversion is a columnar transfer rather than row-by-row Python object materialization.