# Performance

> Every measured figure for v1.0.0 with its source: the base proof of a full Ethereum block, the recursion tree, the decider and the contract, each family's shard proof, and Mercury's own costs.

All end-to-end figures are block 257,510 of `glamsterdam-devnet-8`, proved through the stateless validator guest: 60 transactions, 101.5 Mgas, 198 million cycles. Each figure comes from the specification's measurements.

## End to end

| Stage | Machine | Result |
| --- | --- | --- |
| Base proof | 32 vCPUs, 247.7 GiB, 12 shards in flight | 207 shards, 14.5 MB, 2,481 s; peak 173.92 GiB |
| Recursion leaves | 32 CPUs, four leaves at once | 21, 24, 23 and 27 shards; 2,157 s; 92 GiB peak |
| Recursion root | the same machine, four shards in flight | 21 shards, 460 s, 1.03 MB |
| Decider ceremony | 18-core laptop | `init` 65 s; a contribution 50–56 s; `key` 70 s and 12.7 GB; the key 2.65 GB |
| Decider proof | 18-core laptop | key read in 1 s, proof 18.5 s, 6.1 GB; 7,896,686 constraints over `2^23` |
| On-chain verification | revm | 3,620,026 gas; 34,980 bytes of calldata; 358 points |

## The base proof, pass by pass

| | |
| --- | --- |
| Pass 1, commit | 191 s; 25.7 vCPUs busy on average; one-thread fills take 81% of its shard-seconds; sampled memory at most 15.9 GiB |
| Pass 2, prove | 2,290 s; on average 11.95 of 12 shards held and 30.4 vCPUs busy until the guest exits; then a 460 s tail, whose longest stretches are the two `KECCAK_F` shards' one-thread fills, 200 s and 279 s |
| Peak memory | 173.92 GiB: the two `2^18` `KECCAK_F` shards, together in the tail with nothing else in flight |

The peak was set by one delegation family's height, not by the number of shards in flight.

## A small guest

The Quickstart's guest, 114 cycles in four instruction families, proved at `2^20` instruction heights and `2^16` window heights with two shards in flight: 7 shards, 52 s and 18 GB peak on an 18-core, 48 GiB laptop, almost all of it the two `2^20` shards being worked. The floor of a proof is set by its families and heights, not by its cycle count.

## Each family's shard

At the default heights. A shard's proof size is fixed by its circuit's shape and height; its proving cost follows its height times its circuit's width, however many rows are live.

| Family | Height | Committed `M`/`W`/`S` | Enforcing gates | Inner columns | Shard proof |
| --- | --- | --- | --- | --- | --- |
| `ADD_SUB_LUI_AUIPC` | `2^22` | 27 / 35 / 7 | 63 | 314 | 64,764 B |
| `JUMP_BRANCH_SLT` | `2^22` | 21 / 44 / 10 | 42 | 392 | 69,436 B |
| `SHIFT_BITWISE` | `2^22` | 21 / 61 / 10 | 48 | 478 | 76,644 B |
| `MUL_DIV` | `2^20` | 21 / 54 / 9 | 54 | 444 | 67,412 B |
| `MEM_WORD` | `2^22` | 31 / 24 / 7 | 33 | 314 | 63,836 B |
| `MEM_SUBWORD` | `2^22` | 31 / 55 / 10 | 53 | 472 | 76,196 B |
| `ATOMICS` | `2^20` | 26 / 54 / 9 | 46 | 472 | 68,468 B |
| `INIT_TEARDOWN` | `2^22` | 2 / 0 / 1 | 0 | 46 | 36,316 B |
| `ZERO_WINDOWS` | `2^22` | 2 / 0 / 0 | 0 | 46 | 36,284 B |
| `KECCAK_F` | `2^18` | 208 / 1,556 / 0 | 385 | 5,490 | 381,100 B |
| `POSEIDON2` | `2^8` | 100 / 4,092 / 0 | 4,248 | 2,020 | 664,780 B |
| `FR_ARITH` | `2^8` | 104 / 2,576 / 0 | 2,701 | 142 | 266,292 B |
| `PUBLIC_INPUT`, `PUBLIC_OUTPUT` | `2^12` | 3 or 2 / 0 / 0 | 0 | 26 | 12,556 B, 12,524 B |
| `ADVICE_WINDOWS` | `2^22` | 3 / 0 / 0 | 0 | 46 | 36,316 B |
| `MOD_MUL` | `2^16` | 104 / 221 / 0 | 125 | 2,244 | 135,220 B |
| `SHA256_COMP` | `2^18` | 104 / 520 / 0 | 119 | 2,802 | 189,988 B |
| `EC_ADD` | `2^16` | 392 / 1,028 / 0 | 637 | 8,772 | 434,916 B |

A height changes only the number of halving lists and sumcheck rounds, not the gates: `ADD_SUB_LUI_AUIPC` at `2^20` has 298 inner columns and a 57,196-byte proof against 314 and 64,764 at `2^22`.

## Forward-pass memory

The GKR prover holds every inner layer as field elements, 32 bytes a cell. Representative working sets:

| Shard | Forward pass |
| --- | --- |
| `SHIFT_BITWISE` at `2^20` | 8.4 GiB |
| `MOD_MUL` at `2^16` | 4.6 GB |
| `EC_ADD` at `2^16` | 18.3 GB |
| `SHA256_COMP` at `2^18` | 22.6 GB |
| `KECCAK_F` at `2^18` | 42 GiB |

## Mercury

On an 18-core Apple M5 Pro:

| | |
| --- | --- |
| Commit, `n = 2^22` | 1.30 s |
| Open, `n = 2^22` | 2.89 s |
| 16 columns of `2^20` as one batch | opened in 1.01 s, verified in 4.8 ms |
| The same 16 opened one by one | 9.79 s, verified in 62 ms |

## Reading these numbers

Proving is memory-bound, and its memory follows the shards in flight and their heights, never the length of the execution. Time follows the cycle count family by family. The on-chain cost follows the number of points the root owes the final pairing, at about 9,000 gas a point. Cycle counts themselves are exact and machine-independent, so the cycle profiler is the right first tool for estimating any of the rest.
