# The Streaming Prover

> The prover executes the guest twice and never holds the execution trace. Memory follows the shards being worked, not the length of the run, and the proof does not depend on the schedule.

A full Ethereum block is about 200 million cycles. Its trace, every row of every family and every memory event, would be about 300 bytes a cycle: tens of gigabytes before a single column is committed. Apogee's prover never builds it. It executes the guest twice and holds only the shards it is working on.

## Two passes

> Figure: Commit, then prove. The memory challenges must follow every shard's memory commitments, so the commitments come first, from one execution, and the proofs from a second.

**Pass 1** executes the guest and, as each shard fills, commits its memory columns, keeps the commitments and drops the rows. At the exit it derives everything else the statement needs from the final memory state: the register and pc boundary, the list of memory windows the run touched, and the window families' shards. Then it runs the global transcript, which absorbs the statement, every memory commitment included, and draws the memory challenges and the digest every shard is seeded from.

**Pass 2** executes again. The emulator is a pure function of its input, so it cuts the same shards, and pass 2 asserts that its cycle profile, window list and boundary are pass 1's. Each shard gets every committed column and is proved: its transcript, its GKR pass, its opening. The memory columns are not recommitted: the opening takes their commitments from the statement and their values from pass 2, so columns that differed between the passes would give an opening the verifier refuses.

The order is forced by soundness. The memory challenges must follow every value a memory tuple can read, so every shard's memory columns are committed before any shard can be proved.

## The pipeline

A fixed number of workers, `max_in_flight`, share one lock around the executor. Under the lock a worker hands back its finished shard and claims the next: a filled shard if one is waiting, and otherwise it steps the executor itself until a buffer fills. Outside the lock it builds the shard's columns, proves it and drops it.

- **The executor never runs ahead of demand.** At most one filled, unclaimed shard per family waits, as rows.
- **Within a shard, the work is data-parallel** across every core. A worker blocked in that work does not take a second shard.
- **The block does not depend on the schedule.** A shard's proof is a function of the global state and its own columns; proofs are placed by statement position. The bytes are identical at 1 and at 8 shards in flight.
- **Failures are deterministic.** The failure returned is the earliest in fill order, at any worker count.

## What it costs

Memory is a partial buffer per family, the last-access tables, the shards being worked and the output. A shard's working set is dominated by its forward pass, every inner GKR layer as field elements: 8.4 GiB for a `2^20` `SHIFT_BITWISE` shard, 42 GiB for a `2^18` `KECCAK_F` one. So **`max_in_flight` bounds how many of those coincide, and heights set how large each is.**

Measured on block 257,510 (60 transactions, 101.5 Mgas, 198M cycles, 207 shards) on 32 vCPUs and 247.7 GiB, twelve shards in flight:

| | |
| --- | --- |
| Pass 1 | 191 s; one-thread fills take 81% of its shard-seconds; sampled memory at most 15.9 GiB |
| Pass 2 | 2,290 s, with about 12 of 12 shards held and 30.4 vCPUs busy until the guest exits, then a 460 s tail |
| Peak memory | 173.92 GiB: the two `2^18` `KECCAK_F` shards, together in the tail with nothing else in flight |

The peak came from one delegation family's height, not from the twelve shards in flight. That is the lever: a block with fewer Keccak calls, or Keccak at a lower height, peaks lower.

The specification: [The streaming prover](https://apogee.gweb3networks.com/docs/auditors/spec/streaming).
