Case study: crxml¶
crxml is a high-throughput adapter for Crystal Reports XML exports. It is a concrete example of how the techniques from the other advanced pages combine to reach ~2.4 GB/s on a single workstation.
What it parses¶
Crystal Reports exports tabular data inside XML elements such as:
<Field Name="amount"><Value>123.45</Value></Field>
<Text Name="status"><TextValue>active</TextValue></Text>
crxml reads these exports and turns them into Arrow tables or DataFrames. The speed is parser-bound; the rypipe-core engine keeps up without being the bottleneck.
Architecture¶
Crystal Reports XML file
|
v
CrystalXmlSplitter : finds row-tag boundaries
|
v
CrystalXmlDecoder : extracts fields from each row
|
v
rypipe-core engine : typed builders, filters, projection, Arrow export
|
v
pyarrow.Table / pandas.DataFrame
The Rust side lives in crxml-core. The Python side is a thin CrystalXMLAdapter that calls the Rust core and registers itself with rypipe.
Techniques from this section¶
| Page | Technique used in crxml |
|---|---|
| Adapter design | memchr::memmem splitter; skip comments/CDATA; validate tag boundaries. |
| Adapter design | Borrowed-slice quick_xml reader; XML events point into the input buffer. |
| Schema and types | field_types casts strings to numbers during parse. |
| Dictionary encoding | dictionary_columns for low-cardinality string fields. |
| Parallelism | Parallel fast path when auto_dict and compare filters are off. |
| Execution modes | columnar, parallel, and stream modes exposed through rypipe. |
| I/O tuning | mmap with prefault for cached files; bounded streaming for huge files. |
The splitter¶
CrystalXmlSplitter uses memchr::memmem to scan for the row tag. It is SIMD-accelerated on most platforms. It skips <!-- ... --> and <![CDATA[ ... ]]> regions so a <Row string inside them is not mistaken for a real row start. It also validates that a candidate tag is followed by whitespace, >, or / to avoid prefix collisions such as <RowItem.
The decoder¶
CrystalXmlDecoder uses quick_xml in borrowed-slice mode. Events reference the input bytes directly instead of copying into a scratch buffer. For each row element it:
- Emits row attributes as fields.
- Walks child elements.
- Recognizes
<Field>,<Text>, and<Section>patterns. - Calls
sink.put_field(key, Value::Str(value))so the engine builds typed columns.
The decoder also has a parse_tail fallback that rescans orphan close-tags at chunk boundaries, so chunked parsing stays correct without a serial pre-pass.
Why it is fast¶
| Technique | Benefit |
|---|---|
Borrowed-slice quick_xml reader |
XML events point into the input buffer; no per-event copy. |
memchr::memmem row-tag scan |
SIMD-accelerated boundary search for parallel chunks. |
| Skip-region handling | Comments/CDATA do not create false split points. |
| SIMD UTF-8 validation | simdutf8 validates each chunk in bulk. |
rypipe-core typed builders |
Strings are copied into Arrow arrays only once, during parse. |
| Parallel fast path | When auto_dict and compare filters are off, chunks export independently. |
Lessons for adapter authors¶
- Specialize the parser. Generic line splitting is fine for engine benchmarks, but real throughput comes from a format-aware parser.
- Find split points cheaply. A single
memmemscan beats scanning byte-by-byte. - Handle boundary cases. Chunks can start or end inside a row; have a fallback path that rescans from the nearest safe row start.
- Borrow strings into the engine. Pass
Value::Str(&str)slices whenever the input is valid UTF-8. - Register with
rypipe. A thin adapter class lets users callrypipe.read()while you keep the fast Rust core.
Source¶
The full implementation is in the crxml repository, especially:
src/crxml_core/src/xml/splitter.rssrc/crxml_core/src/xml/decoder.rssrc/crxml/rypipe_adapter.py