# Execution, Families and Shards

> One hart, a 38-bit clock, every memory access as a timestamped query, 23 circuit families, and the shard as the unit of proving.

## The machine

The emulator runs RV32IMAC on one hart over a `ProgramImage`, with no interrupts and no privilege levels. A run is a pure function of the image and its input, with no clock, randomness or threads, so two runs cut identical shards. It differs from a hosted RV32IMAC at three points: `sc.w` always succeeds, a misaligned halfword or word access is fatal rather than split, and the instruction stream is the image decoded at load.

Every other way a run can stop short of `EXIT`, such as an access outside the mapped regions, `ebreak`, or a jump to a halfword with no instruction, is a fatal error with no trace. Such a run has no proof. A nonzero exit status is not an error: it is an execution like any other, and provable.

## The clock and the query

Cycle `c` occupies the four timestamps `4c + Δ`, one per **slot** `Δ ∈ {0, 1, 2, 3}`. Every instruction is one cycle and nothing else is: a delegation invocation rides the cycle that requested it. Cycles are numbered from 1, because timestamp 0 is every address's initial write. The clock is 38 bits, so an execution runs at most `2^36 − 1` cycles.

Every memory access is a **query**: a read of a value last written at some earlier timestamp, and a write at this one. A query that only reads writes back what it read. Slot 0 of every cycle is the pc query, which reads `pc` and writes `next_pc`. Then come the instruction's register and memory queries at fixed slots:

| Class | Δ = 1 | Δ = 2 | Δ = 3 |
| --- | --- | --- | --- |
| register-immediate, `jalr` | `rs1` | | `rd` |
| branches | `rs1` | `rs2` | |
| register-register, M | `rs1` | `rs2` | `rd` |
| loads | `rs1` | the word, read | `rd` |
| stores | `rs1` | `rs2` | the word, with the stored bytes merged in |
| atomics | `rs1` | `rs2` | the word, and `rd` |
| `ecall` | `a7` | `a0` | `a0`, and a delegation's mirror query |

Addresses live in **spaces**: the 32 registers, RAM by 4-aligned word, the pc, one anchor space per delegation type, and the recursion format's field cells. `x0` is an ordinary register in the trace and a constant in the machine: every query at it reads and writes 0.

## Twenty-three families

A **family** is a circuit and the rows it proves. There are four kinds:

| Kind | Families | A row is |
| --- | --- | --- |
| Execution | 0–6: `ADD_SUB_LUI_AUIPC`, `JUMP_BRANCH_SLT`, `SHIFT_BITWISE`, `MUL_DIV`, `MEM_WORD`, `MEM_SUBWORD`, `ATOMICS` | one executed instruction |
| Window | 7 `INIT_TEARDOWN`, 8 `ZERO_WINDOWS`, 12 `PUBLIC_INPUT`, 13 `PUBLIC_OUTPUT`, 14 `ADVICE_WINDOWS` | one memory word, initialized and torn down |
| Delegation | 9 `KECCAK_F`, 10 `POSEIDON2`, 11 `FR_ARITH`, 15 `MOD_MUL`, 16 `SHA256_COMP`, 17 `EC_ADD` | one invocation over a frame of RAM |
| Recursion | 18 `FIELD_WINDOWS`, 19 `FR_OP`, 20 `P2_FIELD`, 21 `FIELD_IO`, 22 `FQ_OP` | one field cell, or one coprocessor operation |

Every cycle goes to the one execution family whose decoded table claims its pc. Families interleave in time: `ADD_SUB_LUI_AUIPC` may own cycles 1 and 3 and `JUMP_BRANCH_SLT` cycle 2. Nothing needs them to be contiguous, because the memory argument orders every row by its pc write.

The window families exist because the memory argument needs every address an execution touches to have exactly one initial value and one final value. `INIT_TEARDOWN` covers RAM window 0 and starts it with the program's image; `ZERO_WINDOWS` covers every other window of ordinary RAM the run touched and starts it at zero; the public pair covers the input and journal windows; `ADVICE_WINDOWS` covers the advice region, starting it with the prover's bytes.

## Shards

A family's rows, in the order they appear, are cut into **shards** of the family's height. The last is padded with zero rows, which the circuits are built to accept. A shard costs its full height whatever its occupancy, so the families a program touches and the heights it chooses set the floor of every proof.

| Family | Default height | Why |
| --- | --- | --- |
| Instruction families | `2^22`, `2^20` for `MUL_DIV` and `ATOMICS` | the floor of their timestamp range checks is `2^20` |
| RAM windows | `2^22` | one shared window height, at least `2^16` |
| Public input, journal | `2^12` | pinned: the height places the windows |
| `KECCAK_F`, `SHA256_COMP` | `2^18` | four times the calls of their `2^16` floor for 2% more proof |
| `MOD_MUL`, `EC_ADD` | `2^16` | their floor |
| `POSEIDON2`, `FR_ARITH` | `2^8` | no table, so no floor |

A shard is proved on its own, by its family's circuit, except for the memory argument: each shard's circuit outputs the product of its read tuples and of its write tuples, and the verifier reconciles those products across every shard of the statement once. That is the only thing that joins shards. There is no per-shard chaining of the pc and no shared boundary between neighbours.

The shard proof sizes at the default heights run from about 12.5 KB for a public window to 665 KB for a `POSEIDON2` shard; an instruction family's is 64 to 77 KB. [Performance](https://apogee.gweb3networks.com/docs/architecture/performance) has the table.

The specification: [Execution trace](https://apogee.gweb3networks.com/docs/auditors/spec/execution-trace), [Circuits and registry](https://apogee.gweb3networks.com/docs/auditors/spec/circuits).
