# Streaming

What streaming actually does in each SDK — which three formats parse incrementally, which runtimes deliver it, and what the other 33 formats give you instead.

`stream_chunks` / `streamChunks` (and per-format `stream(...)` in Rust) hand you
chunks one at a time instead of one list at the end. That lets you forward,
persist or embed each chunk as it arrives.

  This page used to promise bounded memory for every format. That was wrong, and
  it is the kind of wrong that shows up as an OOM in production, so here is the
  short version:

  - **Genuinely incremental parsing exists for three formats — PDF, XLSX and
    CSV — and only in native Rust.** Python gets two of them (PDF, XLSX).
  - **JavaScript's `streamChunks` computes the entire chunk array first, then
    yields from it.** It is an ergonomic wrapper over `getChunks`. It gives you
    *zero* memory benefit for any format.
  - Every other format's `stream()` is literally
    `chunk(...).into_iter()` — the document is parsed in full, then drained.

  What *is* true everywhere, and is the real reason to use streaming: **lazy
  emission**. You get to overlap your embedding / upsert / HTTP-write work with
  iteration instead of waiting for a complete list, and you can stop early.

## Three classes of streaming

| Class | What it means | Peak memory |
| --- | --- | --- |
| **Incremental** | The parser reads only as far as the chunk you asked for. Breaking early skips the rest of the file. | ≈ one page / one sheet-row window |
| **Deferred drain** | The whole document is parsed behind the iterator (on a worker thread where one is available), then chunks arrive one at a time through a channel. Construction returns immediately. | ≈ parsed document + channel depth |
| **Eager drain** | The whole chunk list is built before the first item is yielded. Identical peak memory to the batch call. | ≈ parsed document + all chunks |

## The matrix

All 36 supported extensions, by runtime. **DEFAULT parameters assumed.**

| Format (extensions) | Rust (native) | Python | JavaScript |
| --- | --- | --- | --- |
| **PDF** — `.pdf` | **Incremental** in `default` mode; **deferred drain** in the other six | same as Rust — the binding calls the engine's stream directly | **Eager drain** |
| **Spreadsheets** — `.xlsx` `.xls` `.xlsm` `.xlsb` `.ods` `.xltx` `.xltm` | **Incremental** in `row` and `sliding_window`; eager drain in `table` / `sheet` / `page_aware` / `semantic` | same as Rust | **Eager drain** |
| **CSV / TSV** — `.csv` `.tsv` | **Incremental** — a worker thread reads the file line by line through a `BufReader`; the file is never fully loaded¹ | **Eager drain** — the binding calls the engine's batch `csv::chunk`, not `csv::stream` | **Eager drain** |
| **DOCX family** — `.docx` `.docm` `.dotx` `.dotm` | Eager drain | Eager drain | Eager drain |
| **All other formats** — `.doc` `.ppt` `.pptx` `.potx` `.potm` `.ppsx` `.ppsm` `.md` `.html` `.htm` `.txt` `.rtf` `.epub` `.ipynb` `.json` `.jsonl` `.ndjson` `.eml` `.mbox` `.msg` `.odt` `.odp` | Eager drain | Eager drain | Eager drain |

That is 36 extensions: 1 PDF + 7 spreadsheet + 2 delimited + 4 Word OOXML +
22 in the last row.

¹ Rust's CSV stream hands chunks to the consumer over an **unbounded** channel
(PDF's is bounded at 64). The *file* is never fully read into memory, but a
consumer slower than the reader can let produced chunks accumulate. If that
matters, do your per-chunk work synchronously inside the loop rather than
buffering.

  An earlier version of this page claimed `structural` and `semantic` were
  incremental for the markdown family. They are not, in any runtime. Their
  `stream()` is `chunk(...).into_iter().map(Ok)` — a full parse, then a drain.
  Peak memory for a 500 MB Markdown file is the same whether you call
  `get_chunks` or `stream_chunks`.

## The PDF exception, explained

PDF's `default` mode is the one place a document format truly streams, and the
reason is heading ranking, not pagination:

- `default` ranks type sizes **within each page**, so a page can be parsed,
  chunked and dropped. With a resumable builder a chunk costs only the pages it
  came from — the first chunk of a 5,000-page PDF reads fewer than 10 of them.
- The other six modes rank sizes **across the whole document**. Their reader has
  read every page before it renders the first, and no amount of chunker work
  changes that.

For those six modes the engine still does the honest thing available to it: the
parse runs on a worker thread and chunks arrive through a **bounded channel of
depth 64**, so construction returns immediately and at most 64 chunks sit
between producer and consumer. It is not bounded parse memory, but it is not a
materialized list either.

Pagination is *not* the obstacle it looks like. The end of page 3 and the start
of page 4 are routinely one paragraph, and chunking each page separately breaks
the sentence at every boundary — measured on a 12-page paper, 71 chunks instead
of 66. `default` does not chunk pages separately: the chunker is fed one
continuous stream of markdown and finalizes a chunk only once the text after it
has arrived.

## What JavaScript's `streamChunks` actually is

```ts
export async function* streamChunks(source, opts = {}) {
  const chunks = await getChunks(source, { ...opts, listImages: false });
  for (const chunk of chunks) yield chunk;
}
```

That is the whole implementation. The WASM boundary is a synchronous full
parse — there is no way to yield from inside it — so true streaming in
JavaScript would be an engine redesign, not a wrapper change. This is a
deliberate, documented decision, not an oversight.

Use it for the ergonomics: `for await` reads better than indexing an array, you
can `break` out of the loop, and your `await upsert(...)` overlaps with nothing
in particular but keeps the code shape identical to the Python and Rust
versions. Do **not** use it to survive a file that would OOM `getChunks`.

```python
from py_chunks import stream_chunks

# Genuinely incremental for pdf and xlsx: chunks are produced as the document
# is read, so breaking early skips the rest of the parse. Every other format
# is chunked in full by the binding and drained through the same iterator —
# same results, same peak memory as get_chunks().
batch = []
for chunk in stream_chunks("large.pdf", mode="section"):
    batch.append(chunk)
    if len(batch) == 128:
        upsert(batch)
        batch.clear()

if batch:
    upsert(batch)
```

```ts
import { streamChunks, type Chunk } from "js-chunks";

// Ergonomics, not bounded memory: the wasm boundary is a synchronous full
// parse, so the whole chunk list exists before the first yield. Use this to
// overlap embedding/upsert work with iteration — not to survive a huge file.
let batch: Chunk[] = [];
for await (const chunk of streamChunks("./large.pdf", { mode: "section" })) {
  batch.push(chunk);
  if (batch.length === 128) {
    await upsert(batch);
    batch = [];
  }
}
if (batch.length) await upsert(batch);
```

```rust
use chunks_rs::formats::pdf;
use chunks_rs::Chunk;

// Streaming is per-format (there is no dispatch-level stream). For PDF the
// "default" mode streams for real — a chunk costs only the pages it came from;
// the other modes rank heading sizes across the whole document, so the parse
// completes behind the iterator.
fn main() -> Result<(), Box<dyn std::error::Error>> {
    let mut batch: Vec<Chunk> = Vec::new();
    for chunk in pdf::stream("large.pdf", "default", 3, 1, 3, 15)? {
        batch.push(chunk?);
        if batch.len() == 128 {
            upsert(&batch);
            batch.clear();
        }
    }
    if !batch.is_empty() {
        upsert(&batch);
    }
    Ok(())
}
```

## Byte-identical to batch

  Streaming output equals the batch call for every format and every supported
  mode — `list(stream_chunks(...)) == get_chunks(...)`, content and metadata.
  This is enforced by dedicated parity tests
  (`stream_matches_batch_for_every_mode` in the engine's `pdf_stream` and
  `xlsx_stream` suites) precisely because the two paths are different code.
  Spot-checked live across md/pdf/csv/xlsx/docx: equal every time.

## Per-language shape

| | Streaming call | Yields |
| --- | --- | --- |
| **Python** | `stream_chunks(source, mode=...)` | an iterator of `dict` |
| **JavaScript** | `streamChunks(source, { mode })` | an `AsyncIterable<Chunk>` |
| **Rust** | `chunks_rs::formats::<fmt>::stream(...)` | `Iterator<Item = Result<Chunk>>` |

```python
from py_chunks import stream_chunks

# Yields one chunk at a time
for chunk in stream_chunks("data.csv", mode="row"):
    handle(chunk)
```

```ts
import { streamChunks } from "js-chunks";

// Async iterable — one chunk at a time
for await (const chunk of streamChunks("./data.csv", { mode: "row" })) {
  handle(chunk);
}
```

```rust
use chunks_rs::formats::csv;

// stream() is a native Iterator yielding Result<Chunk>
fn main() -> Result<(), Box<dyn std::error::Error>> {
    for c in csv::stream("data.csv", "row", 10, 5, 1, true, None, "utf-8", true)? {
        let c = c?;
        println!("{}", c.content_type);
    }
    Ok(())
}
```

**Rust**

There is **no dispatch-level `stream`** in rs-chunks. `chunks_rs::get_chunks`
has no streaming counterpart; you pick the format module yourself
(`formats::pdf::stream`, `formats::csv::stream`, …). Each module's `stream`
takes the same arguments as its `chunk`.

**Python**

Python's `ChunkStreamIterator` converts each chunk into a Python `dict` **when
you ask for it**, not up front. On a 5,000-page PDF that is the difference
between materializing 71,111 dicts before you read the first and building them
one at a time — a real saving even for the eager-drain formats, where the Rust
side has already finished.

Python also exposes source-specific streaming helpers —
`stream_chunks_from_path`, `stream_chunks_from_bytes`,
`stream_chunks_from_fileobj`, `stream_chunks_from_upload`, and
`stream_chunks_from_s3_presigned_url` — mirroring the batch helpers in
[Input Sources](/docs/input-sources).

Every one of the 36 supported extensions has a streaming entry point, so
`NotImplementedError: Streaming not yet supported for {ext} files` is only
reachable if a future format lands without one.

## Choosing between streaming and batch

| Situation | Use |
| --- | --- |
| Large PDF, `default` mode, native Rust or Python | **stream** — genuinely bounded |
| Large spreadsheet, `row` or `sliding_window`, native Rust or Python | **stream** — genuinely bounded |
| Large CSV, native Rust | **stream** |
| Large CSV, Python or JavaScript | either; memory is the same |
| Any other format, any runtime | either; stream if the code reads better or you want to stop early |
| You need images | **batch only** — see below |
| Node/Bun/Deno, any format, worried about memory | neither helps; split the document upstream |

  `list_images` / `listImages` is **not** available on the streaming entry
  points — image extraction needs the whole document, and `streamChunks`
  explicitly forces `listImages: false`. Use
  `get_chunks(..., list_images=True)` (or `getChunks(..., { listImages: true })`)
  when you need image bytes. See
  [Supported Formats](/docs/supported-formats#image-extraction).

## Bytes sources & cleanup

**Python**

When you stream from bytes (e.g. a request body), Python writes a temporary file
with the right extension — the engine's streaming paths are path-based — and
removes it when the iterator is exhausted or you exit early. Use it as a context
manager to guarantee cleanup:

```python
with stream_chunks(data, filename="big.pdf", mode="section") as it:
    for chunk in it:
        ...
```

**JavaScript** / **Rust**

Byte sources need no temp file here: JavaScript's `streamChunks` goes through
`getChunks`, which passes bytes straight to WASM, and Rust's per-format
`stream(...)` is path-based by design — pass a path, or use the batch
`chunk_from_bytes` if you only have bytes.

See [Framework Integration](/docs/framework-integration) for FastAPI, Express,
Next.js and Axum handlers that forward NDJSON as chunks are produced.
