Apogee VMDocs v1.0.0
gweb3networks.com ↗
Apogee VM key art: the Apogee emblem rising over a planet's horizon, with the words Higher Compute Horizons

Apogee VM·v1.0.0·The First Step

Every application can be blockchain-native.

Apogee VM proves that a program ran correctly. You write ordinary Rust. Apogee runs it on a RISC-V machine, proves every instruction it executed, and hands the chain one proof that a single contract call can check. Correctness stops being something your users trust and becomes something they verify.

What a proof statesVerified

Program
One field element, the identity: a digest of the code, the initial memory image, the entry point and the circuit configuration.
Input
The public bytes the program was given.
Output
The journal: the bytes it chose to publish.
Exit
The status it ended with. 0 is success.

Groth16 · BN254one contract call

Proof of correctness

Software has always asked to be trusted. Apogee lets it be checked instead.

Whenever a program runs on someone else's machine, its users take the result on faith: the ledger, the order book, the payout, the roll of the dice. Blockchains removed that faith for one narrow kind of program by having every node re-run every transaction. It works. It is also the most expensive way ever devised to agree on anything.

A zkVM, a virtual machine that proves its own execution, removes the faith for any program. The program runs once, anywhere. It leaves behind a mathematical receipt saying that this program, given this input, produced this output, and checking that receipt never means running the program again. That receipt is what makes an application blockchain-native: its rules live in code, its state lives on-chain as a commitment, and every change to that state arrives with its proof.

The third era

Work, then stake, then correctness.

Each era of blockchains found a new way to stop trusting someone. Proof of correctness is the first that reaches inside the computation itself.

01

Energy

Proof of Work

Electricity secures the order of events. Rewriting history means out-spending the honest majority's power bill.

02

Capital

Proof of Stake

Capital secures the order of events. Misbehaviour is punished by burning the stake that vouched for it.

03

Mathematics · now

Proof of Correctness

Mathematics secures the events themselves. Every state change carries a proof that it was computed by the program everyone agreed on.

Work and stake decide which history counts. Neither checks what happened inside it; that job has always fallen to every node re-running everything. A validity proof retires that last brute force. Energy, then capital, then mathematics: there is no fourth thing left to stop trusting.

What changes

Your product, its own chain, and mathematics as the referee.

The blockchain-native future is not one chain doing everything. It is many environments, each shaped around one application, all settling to the same base layer. Apogee is the proof engine that makes running such an environment practical.

Write

Your logic, in ordinary Rust

A guest program is a no_std Rust binary for RISC-V. It reads its input, does its work and commits its output. Apogee proves each run, and you never have to think in circuits.

Settle

Finality without a waiting room

A validity proof is final the moment it verifies. There is no seven-day dispute window to sit out, and no committee or enclave standing in for the mathematics.

Specialize

An environment shaped like your product

One application per rollup puts every resource behind the one channel that matters to it. Apogee proves any program built for its machine, so the state-transition function is yours to define.

Measured, not promised

A full Ethereum block, from guest to contract.

Apogee v1.0.0 was proved end to end on block 257,510 of glamsterdam-devnet-8. The block was validated statelessly inside the VM under the execution-specs rules, then folded by recursion into one proof that an Ethereum contract accepts.

101.5 MgasOne block, 60 transactionsRun through the stateless validator, inside the VM.
198MRISC-V cyclesEvery executed instruction is a proved row, across 207 shards.
1Proof at the topA 116-shard recursion tree, folded into one Groth16 proof.
3.62M gasTo verify on-chainOne call to ApogeeVerifier.sol, with 34,980 bytes of calldata.
704 BPer shard openingOne Mercury proof opens every committed column of a shard.
23Circuit familiesSeven for instructions, five for memory windows, six delegations, five for recursion.
67,251Conformance casesEvery tests-zkevm v21.0.1 pair, matched by the validator natively.
0Outside cryptographyFields, curve, pairing, MSM, hash, PCS, GKR and Groth16 are written in-repo. Outside libraries serve only as test oracles.

The base proof took 2,481 s on a 32-vCPU machine and peaked at 174 GiB; the recursion tree took about 2,620 s more. Proving is memory-bound today, and the performance page gives every figure with its source.

The path of a proof

You write the program. Apogee does everything after it.

Between your Rust and the contract's true sit a decoder, 23 circuit families, the GKR engine, the commitments, a recursion tree and a Groth16 decider. None of it is yours to build or maintain.

The path of a proof Eight steps: write, load, execute, shard, prove, recurse, decide, verify. The developer writes the first; Apogee performs the next six; the chain performs the last. Under each step, the size of what exists at that point for block 257,510. YOU WRITE APOGEE PROVES THE CHAIN CHECKS Writeno_std Rust Loadimage · identity ExecuteRV32IMAC · 1 hart Shard23 families ProveGKR · Mercury Recurseleaves → root DecideGroth16 · BN254 VerifyApogeeVerifier.sol your source32-byte identity198M cycles207 shards14.5 MB proof1.03 MB root34,980 B calldatatrue · 3.62M gas
The path of a proof. Everything between the two outer bands is Apogee's. Figures under each step are block 257,510's, from the specification's measurements.

Blockchain-native

The stack you already know, with one swap per layer.

Going blockchain-native does not mean learning a new discipline. Every component of a conventional application has a counterpart, and the mental model carries over almost unchanged.

LayerConventional applicationBlockchain-native, on Apogee
Business logicA service on servers you operateA guest program in Rust, proved on every run
DatabaseSQL or a key-value storeData off-chain, a state root on-chain
QuerySELECT … WHERE key = ?An inclusion proof, checked against the root
CommitCOMMITA new root, published with its proof
Audit trailLogs you ask people to believeA proof anyone can check

Every layer, with a worked example →

Start here

Six ways in.

The same system, read from six directions. Pick the one that matches the question you arrived with.

The Apogee emblem: two blades meeting at an apex above a planet, with a star at their centre

The first step is a program.

Write it in Rust and run it on Apogee. Everything after that, from the shards and the circuits to the recursion and the contract, is the machine's job. What reaches the chain is a proof, and a proof is all the chain needs.

Higher compute horizons

The First Step

Blockchain Native

A blockchain-native application is built from the same components as the one you build today, with one swap at each layer. Here is every swap, and a ledger built both ways.

View as Markdown

A blockchain-native application is economic activity whose settlement, custody and rules are on-chain by construction, rather than a conventional business with a token attached to the side of it. That sounds like a different kind of engineering. It is less different than it sounds.

Every component of the stack you build today has a counterpart. The counterpart does the same job with one change: what used to be trusted is now proved. Apogee exists to make that change cheap enough to be the default.

The shift in one sentence#

In a conventional application the server is the authority: it holds the data, applies the rules and reports the result. In a blockchain-native application the chain holds a commitment to the data, the rules are a program anyone can name by its digest, and a result is accepted only with a proof that this program produced it.

The operator does not disappear. Someone still runs the program, stores the data and answers requests. What disappears is the need to believe them.

Layer by layer#

Layer Conventional application Blockchain-native, on Apogee What carries over
Business logic A service you deploy to servers you operate A guest program: no_std Rust compiled to RISC-V and proved on every run You still write functions over data. The program is named by its identity, a digest of its code and configuration.
Data store SQL tables, a key-value store The data stays off-chain; the chain stores a state root, one hash that summarizes a snapshot of all of it A snapshot you can name in 32 bytes and check anything against.
Read query SELECT balance FROM accounts WHERE id = ? An inclusion proof, a Merkle path, which the guest checks against the root A query still returns a row. The row now arrives with evidence, and the guest refuses any row that does not check.
Write UPDATE …; COMMIT; A state transition: the guest computes the new root and publishes it Commit still means "make it durable". It now means a contract moving the stored root forward.
Request An HTTP request body The public input, which the proof binds Inputs in, outputs out.
Bulk payload Uploads, joined rows, fetched documents Advice: bytes the prover supplies and the guest checks against something the proof binds Pass large data by reference and check what arrived.
Response A JSON body The journal: the public output, bound by the proof Anyone can read the response and know it is the program's.
Authentication Sessions, tokens, passwords Signatures verified inside the guest; secp256k1 recovery runs on delegated field and curve arithmetic Identity is a key, and authorization is a check you can read in the source.
Cryptography libraries sha2, ring, OpenSSL guest_sdk::keccak256, sha256, ec_add, each routed to a dedicated circuit The same calls, at a fraction of the cycles.
Release Push a binary and behaviour changes at once Register the new program identity with the verifier contract Releases become explicit: a new build is a new identity that the contract has to accept.
Scale More servers, sharded databases An execution cut into shards that prove in parallel, folded by recursion into one proof Throughput comes from provers working side by side, while the chain still checks one proof.
Audit Logs and attestations you ask people to believe The proof and its journal Assurance moves from reputation to verification.

The middle column is what a blockchain-native application is made of. Apogee supplies the machinery underneath it: the RISC-V machine, the circuits, the commitments, the recursion and the verifier contract. None of that appears in your program.

What does not change#

  • You still write ordinary Rust. Structs, enums, traits, iterators, Vec, BTreeMap, and any crate that builds without std. There is no circuit language to learn.
  • You still test on your laptop. The usual layout puts the application logic in a no_std library that runs on the host, under cargo test, exactly as it runs in the guest. See Writing a guest.
  • You still reason about state, requests and responses. The shapes are the same; only their guarantees change.
  • Deterministic code stays deterministic. Good backend code already avoids hidden inputs. The VM makes that absolute.

What does change#

  • No ambient world. A guest has no clock, no randomness, no network and no files. Everything it knows arrives as public input or as advice, and a request for host data is a call no proof admits.
  • Every instruction has a price. Each executed instruction becomes a proved row. Copies, allocations and dead loops cost proving time, so the old discipline of counting cycles comes back.
  • Supplied data is checked, not trusted. Advice is chosen by the prover. A guest checks it against something the proof binds before anything derived from it is published.
  • Outputs are small and public. The journal holds at most 16,380 bytes. A large result is published as a digest.
  • Nothing is hidden. Apogee v1.0.0 proofs are succinct, not zero-knowledge. A guest must not hold secrets.

A ledger, built both ways#

A deposit into an account balance: the smallest state change worth proving.

The conventional version#

sql
BEGIN;
SELECT balance FROM accounts WHERE id = $1 FOR UPDATE;    -- read
UPDATE accounts SET balance = balance + $2 WHERE id = $1;  -- write
COMMIT;                                                     -- make it durable

Users trust the operator to have run exactly this, against the real table, and to report the result honestly.

The blockchain-native version#

The accounts live in a binary Merkle tree whose leaves are keccak256(account ‖ balance). A contract stores the root. The guest receives the old root and the request as public input, receives the account's balance and Merkle path as advice, checks the path, and publishes the old and new roots.

guests/ledger/src/main.rsrust
#![no_std]
#![no_main]

guest_sdk::entry!(main);

const DEPTH: usize = 20; // room for 2^20 accounts

/// A leaf commits to one account's balance.
fn leaf(account: &[u8; 20], balance: u64) -> [u8; 32] {
    let mut bytes = [0u8; 28];
    bytes[..20].copy_from_slice(account);
    bytes[20..].copy_from_slice(&balance.to_le_bytes());
    guest_sdk::keccak256(&bytes)
}

/// Fold a leaf up its Merkle path; bit `level` of `index` says whether the
/// node is a right child at that level.
fn root_of(mut node: [u8; 32], index: u32, path: &[[u8; 32]; DEPTH]) -> [u8; 32] {
    let mut pair = [0u8; 64];
    for (level, sibling) in path.iter().enumerate() {
        let (left, right) = if (index >> level) & 1 == 0 { (&node, sibling) } else { (sibling, &node) };
        pair[..32].copy_from_slice(left);
        pair[32..].copy_from_slice(right);
        node = guest_sdk::keccak256(&pair);
    }
    node
}

/// Public input: old_root (32) ‖ account (20) ‖ amount (8, LE)
/// Advice:       balance (8, LE) ‖ index (4, LE) ‖ path (DEPTH × 32)
/// Journal:      old_root ‖ new_root ‖ account ‖ amount
fn main() {
    let input = guest_sdk::public_input();
    let advice = guest_sdk::advice();
    if input.len() != 60 || advice.len() != 12 + 32 * DEPTH {
        guest_sdk::exit(1);
    }
    let old_root: [u8; 32] = input[..32].try_into().unwrap();
    let account: [u8; 20] = input[32..52].try_into().unwrap();
    let amount = u64::from_le_bytes(input[52..60].try_into().unwrap());

    let balance = u64::from_le_bytes(advice[..8].try_into().unwrap());
    let index = u32::from_le_bytes(advice[8..12].try_into().unwrap());
    let mut path = [[0u8; 32]; DEPTH];
    for (i, sibling) in path.iter_mut().enumerate() {
        sibling.copy_from_slice(&advice[12 + 32 * i..12 + 32 * (i + 1)]);
    }

    // The query: the balance the prover supplied is the one the root commits to.
    if root_of(leaf(&account, balance), index, &path) != old_root {
        guest_sdk::exit(2);
    }
    // The write: the same path with the new leaf gives the new root.
    let Some(new_balance) = balance.checked_add(amount) else { guest_sdk::exit(3) };
    let new_root = root_of(leaf(&account, new_balance), index, &path);

    // The commit: publish the transition for the contract to apply.
    guest_sdk::commit(&old_root);
    guest_sdk::commit(&new_root);
    guest_sdk::commit(&account);
    guest_sdk::commit(&amount.to_le_bytes());
}

Read it against the SQL. SELECT … FOR UPDATE became a Merkle path checked against the root. UPDATE became a new leaf on the same path. COMMIT became four calls to commit, which write the journal the proof will bind. The balance came from the prover, and that is fine: a balance the root does not commit to fails the check and the run exits 2.

The contract that owns the root accepts a transition only with a proof that this program produced it and exited 0:

Ledger.sol (sketch)solidity
interface IApogeeVerifier {
    function verify(bytes calldata input, bytes calldata output, uint256 exitStatus,
                    uint256[10] calldata proof, uint256[] calldata points) external view returns (bool);
}

contract Ledger {
    IApogeeVerifier public immutable verifier;
    bytes32 public root;

    constructor(IApogeeVerifier v, bytes32 genesis) { verifier = v; root = genesis; }

    function apply(bytes calldata input, bytes calldata journal,
                   uint256[10] calldata proof, uint256[] calldata points) external {
        require(verifier.verify(input, journal, 0, proof, points), "proof");
        require(bytes32(journal[0:32]) == root, "stale root");
        root = bytes32(journal[32:64]);
    }
}

Note

This is a sketch to show the shape, not a production contract. A real deployment takes a batch of requests per proof, which the guest folds into one transition, and pins the verifier to the right program and public-value lengths. Settle on-chain covers the deployed verifier, its key and its ceremony.

The model in three lines#

  1. The chain holds a root.
  2. The guest proves the transition.
  3. The contract moves the root.

Everything else, from the shards and circuits to the recursion and the decider, is Apogee's. That is the abstraction: a program, its input and its output, and a proof that ties them together.

Next#

The First Step

Apogee at a Glance

The facts on one page. What Apogee VM proves, how it proves it, what that costs, what it assumes, and where version 1.0.0 stops.

View as Markdown

In one paragraph#

Apogee VM is a RISC-V zkVM. It proves that an RV32IMAC program, named by a digest of its image, ran on a given public input to an exit status and wrote a given public output. It carries that proof through a recursion tree to one Groth16 proof that an Ethereum contract checks. Every circuit is a layered GKR circuit over BN254's scalar field, every committed column is opened with Mercury, and every challenge comes from a Poseidon2 transcript. The fields, curve, pairing, MSM, hash, polynomial commitment, GKR prover and Groth16 are all implemented in the repository. Its reference workload is Ethereum block validation.

The facts#

Apogee VM v1.0.0
What a proof states That the program with this identity, started at its entry point over its image, with this public input and some advice, executed instruction by instruction to EXIT with this status, having written this journal
Instruction set RV32IMAC on one hart: the 59 instructions of RV32IMA (40 base, 8 M, 11 A), with compressed instructions expanded at load
Guest language Rust, #![no_std] with alloc, stable 1.96.1, target riscv32imac-unknown-none-elf
Arithmetization 23 circuit families, each a layered GKR circuit: 7 for instructions, 5 for memory windows, 6 delegations, 5 for recursion
Arguments Gates by sumcheck; memory by one read/write multiset over the whole execution; lookups by LogUp
Field BN254's scalar field, 254 bits
Commitments Mercury, multilinear over KZG, one 704-byte opening per shard whatever the column count
Setup The PSE perpetual powers of tau, contribution 80; a second, circuit-specific ceremony for the on-chain decider
Transcript A Poseidon2 duplex sponge over Fr, width 3, rate 2
Settlement Recursion tree → Groth16 decider → ApogeeVerifier.sol
Security level About 100 bits, set by BN254
Zero knowledge No. Proofs are succinct, not zero-knowledge, and nothing is blinded
Delegated operations keccak-f[1600] rounds, SHA-256 rounds, Poseidon2, BN254 Fr arithmetic, 256-bit modular multiplication over four Ethereum moduli, complete point addition on secp256k1 and BN254 G1
Public values At most 16,380 bytes of input and 16,380 bytes of journal; advice up to 2 GiB
Execution length Up to 2^36 − 1 cycles
Code size .text within 7.94 MiB at a 2^22 table height; image within 4 MiB by default
Third-party cryptography None on a proof path. arkworks, Plonky3 and zkhash appear only as test oracles

Measured#

All figures are block 257,510 of glamsterdam-devnet-8, run through the stateless validator guest: 60 transactions, 101.5 Mgas, 198M cycles. Sources: the specification's recursion §10 and streaming §1.

Stage Result
Base proof 207 shards, 14.5 MB, 2,481 s on 32 vCPUs and 247.7 GiB, peak RSS 173.92 GiB
Recursion tree 116 shards: four leaves of at most 64 base shards (2,157 s together, 92 GiB peak) and a root (460 s, 1.03 MB)
Decider circuit 7,896,686 constraints over a domain of 2^23
Decider proof 18.5 s and 6.1 GB on an 18-core laptop, with the key read in 1 s
On-chain verification 3,620,026 gas, 34,980 bytes of calldata, 358 points folded by the contract
Conformance All 67,251 tests-zkevm v21.0.1 pairs match natively

What a verifier must hold#

Two values, taken from a channel the prover does not control:

  • The program identity, one field element. Against an identity supplied by the prover, a proof shows only that some program ran.
  • The SRS digest of the ceremony. A key built over a known τ is refused only by this comparison.

The verifying key itself may come from anyone: loading it recomputes both values from its own contents and holds its circuits to the verifier's registry. The security model has the full list of assumptions.

Where v1.0.0 stops#

  • Not zero-knowledge. No blinding in Mercury, GKR or the decider.
  • Advice is unbound. A guest checks it against something a proof binds.
  • Traps are not provable. A misaligned access, an access outside mapped memory, ebreak, or a pc with no instruction ends the run with no proof.
  • sc.w always succeeds. The one deviation from RV32IMAC: there is no reservation state.
  • Delegations are a fixed set of six. EVM MULMOD with an arbitrary modulus, MODEXP and BLS12-381 run as ordinary instructions.
  • Proving is memory-bound. The measured block peaked at 174 GiB; memory follows the shards in flight, not the length of the run.
  • The decider key is per root shape, and only as trustworthy as its ceremony. The development key is forgeable.

Who builds it#

Apogee VM is the flagship of G Web3's research program toward blockchain-native application environments: one optimized environment per economic application, each settling to Ethereum with a validity proof. The program's position is set out in the thesis; where the next version goes is Quantum Leap.

Launch Your App

Launch Your App

The builder's manual for Apogee VM. How a guest program is written, built, run, proved and settled on-chain, and the habits that keep it correct, provable and cheap.

View as Markdown

A guest is the program Apogee proves: a no_std Rust binary compiled for riscv32imac-unknown-none-elf, with an entry point, three memory regions for its inputs and outputs, and nothing else. The host is everything around it: the code that supplies the input, asks Apogee for a proof and hands that proof to whoever checks it. You write both. Apogee supplies the machine, the circuits and the verifier.

This section is written for two readers at once: an engineer at a keyboard and the model they work with. Every page states its rules plainly, and the AI Companion condenses all of them into one file you can hand to an assistant before it writes a line.

The model#

The guest, the host and the verifier The host passes a public input and advice to the guest running inside Apogee VM. The guest writes a journal and exits with a status. Apogee produces a proof binding the program identity, the input, the journal and the exit status, which a verifier checks. HOST · YOURS Your service builds the input, supplies the advice, asks for a proof APOGEE VM · RV32IMAC · ONE HART GUEST · YOURS Your program no_std Rust guest_sdk public input advice journal · exit status VERIFIER Proof checked against an identity from its own channel: a contract or a service
Who does what. The input and the journal are bound by the proof; the advice, drawn dashed, is not, which is why a guest checks it. The verifier never sees the program, only its identity.

A proof says one thing: the program with this identity, started over its image with this public input and some advice of the prover's choosing, ran to EXIT with this status, having written this journal. Everything you build sits on that sentence.

The workflow#

Step What you do Page
1 Install nothing by hand: the repository pins the toolchain. Fetch the ceremony file for proving Set up
2 Write the guest: an entry point, the three regions, ordinary Rust Write a guest, Inputs, advice and the journal
3 Reach for delegated hashing and curve arithmetic where it pays Delegations
4 Build it for the guest target, then inspect the image and its identity Build and inspect
5 Run it in the emulator and count where the cycles go Run and profile
6 Prove a run and verify it Prove and verify
7 Compress the proof by recursion and check it on Ethereum Settle on-chain

The Quickstart walks the whole loop once with a guest of three lines.

The rules that matter most#

Each is explained, with what goes wrong and what to do instead, in the guest programming guide.

  • usize and every pointer are 32 bits. Overflowing usize panics in the guest alone, x as usize truncates silently, and anything whose layout or hash depends on a length differs between host and guest.
  • The allocator never frees. It bumps a pointer up from __heap_start, so what runs a guest out of memory is the total it allocates over the run, not its peak. Reuse buffers and size them with with_capacity.
  • Atomics work, and you should not write them. The machine has one hart, so the A extension is there for compatibility with code that already uses it. New guest code has nothing to synchronize.
  • There is no world outside. No files, no clock, no randomness, no network. A guest knows its public input, its advice, and what it computes.
  • Advice is the prover's choice. Check it against something the proof binds before anything derived from it reaches the journal.
  • The journal is small. 16,380 bytes at most. Publish a digest of anything that grows.
  • Overflow checks stay on in release. They are part of what the program computes, so the guest profile pins them.
  • Every executed instruction is a proved row. Proving cost follows cycle count, so build --release and count cycles before you optimize anything else.

Where to start#

Launch Your App

Quickstart

From an empty crate to a verified proof. A three-line guest, built, run, inspected and proved, with the real output of every step.

View as Markdown

This page walks the whole loop once with the smallest guest that does something: it reads its public input and publishes it as its journal. Every output below was produced by running these exact commands on Apogee v1.0.0.

Note

What you need. A checkout of the Apogee VM repository at v1.0.0, and rustup; the repository pins everything else. Commands run from the repository root unless a step changes directory. Steps 5 and 6 also need the ceremony file assets/ptau/ppot_0080_24.ptau, and step 6 a machine with tens of GiB of memory. Set up covers both.

Create the guest#

A guest is a no_std binary crate in the guests/ workspace. Create guests/hello:

guests/hello/Cargo.tomltoml
[package]
name = "hello"
version.workspace = true
edition.workspace = true
publish.workspace = true

[dependencies]
guest-sdk.workspace = true
guests/hello/src/main.rsrust
#![no_std]
#![no_main]

guest_sdk::entry!(main);

fn main() {
    // The public input is memory: a slice, with no ecall and no cursor.
    guest_sdk::commit(guest_sdk::public_input());
}

#![no_std] because the target is bare metal. #![no_main] with entry!(main) because the SDK's startup code sets the stack pointer, zeroes .bss and calls a main symbol the macro exports around your function. Returning from it is exit(0).

Add it to the guest workspace#

Append "hello" to the members list in guests/Cargo.toml:

guests/Cargo.tomltoml
members = ["fib", "echo", … , "recursion", "hello"]

Build it#

From the guest's own directory, with no flag but the target:

sh
cd guests/hello
cargo build --release --target riscv32imac-unknown-none-elf
cd ../..

The ELF lands at guests/target/riscv32imac-unknown-none-elf/release/hello. The guest workspace supplies the linker script and --no-relax, so there is nothing else to pass.

Run it#

The profiler runs a guest in Apogee's emulator, with no proof, and reports where the cycles went:

sh
printf 'hello, apogee' > /tmp/hello.in
cargo run --release -p profiler -- elf guests/target/riscv32imac-unknown-none-elf/release/hello --input /tmp/hello.in
workload
  label                        hello
  guest                        hello
  guest cycles                 114
  exit status                  0
  journal bytes                13

cycles by family
  ADD_SUB_LUI_AUIPC            64
  JUMP_BRANCH_SLT              21
  MEM_WORD                     3
  MEM_SUBWORD                  26

114 instructions ran, each of which will be a proved row. The 13 journal bytes are the input, echoed. The 26 MEM_SUBWORD rows are commit copying the input byte by byte with lbu and sb.

See what the VM will prove#

sh
cargo run --release -p artifact-dump -- tables \
    guests/target/riscv32imac-unknown-none-elf/release/hello \
    --ptau assets/ptau/ppot_0080_24.ptau
program identity  9ead85cee880df30daa8eba657316215107a075640a64ccf2424a054b758b802

VmConfig
--------
  id  family              height     live rows  columns
   0  ADD_SUB_LUI_AUIPC     4194304         33  pc next_pc rs1 rs2 rd imm extra_mask
   1  JUMP_BRANCH_SLT       4194304         12  pc next_pc rs1 rs2 rd imm extra_mask
   4  MEM_WORD              4194304          3  pc next_pc rs1 rs2 rd imm extra_mask
   5  MEM_SUBWORD           4194304          3  pc next_pc rs1 rs2 rd imm extra_mask
   7  INIT_TEARDOWN         4194304          0  none: claims no pc
   8  ZERO_WINDOWS          4194304          0  none: claims no pc
  12  PUBLIC_INPUT             4096          0  none: claims no pc
  13  PUBLIC_OUTPUT            4096          0  none: claims no pc
  14  ADVICE_WINDOWS        4194304          0  none: claims no pc

This is the program's static shape at the default heights: the four instruction families its code uses, each with a decoded table, and the five window families every program has. The program identity is one field element that digests all of it. Yours will differ: an ELF embeds absolute paths in its panic strings, so a build on another machine is another image, and every change of heights is another identity.

Prove and verify#

A host program asks for the proof. Put it beside the host SDK as an example:

crates/host/examples/prove_hello.rsrust
use constants::family;
use emulator::GuestIo;
use program::ProgramParams;
use srs::Srs;

fn main() {
    let elf = std::fs::read("guests/target/riscv32imac-unknown-none-elf/release/hello")
        .expect("build the guest with --release first");

    // Small heights for a small program: the seven instruction families at
    // their 2^20 floor, the three RAM-window families at 2^16. Every choice of
    // heights is its own program identity.
    let mut params = ProgramParams::defaults();
    for f in 0..7 {
        params.heights[f] = 1 << 20;
    }
    for f in [family::INIT_TEARDOWN, family::ZERO_WINDOWS, family::ADVICE_WINDOWS] {
        params.heights[f as usize] = 1 << 16;
    }

    // As many ceremony powers as the tallest family has rows: 2^20 here.
    let ptau = std::path::Path::new("assets/ptau/ppot_0080_24.ptau");
    let srs = Srs::from_ptau(ptau, 20).expect("the ceremony file reads");
    let setup = host::setup(&elf, &params, srs).expect("the program registers");

    let io = GuestIo { input: b"hello, apogee".to_vec(), advice: Vec::new() };
    let proven = host::prove(&setup, &io, 2).expect("the run proves"); // two shards in flight
    host::verify(&setup.vk, &proven.block).expect("the block verifies");

    assert_eq!(proven.exit_code, 0);
    assert_eq!(proven.journal, b"hello, apogee");
    let id: String = setup.vk.identity.to_bytes().iter().map(|b| format!("{b:02x}")).collect();
    println!("identity  {id}");
    println!("cycles    {}", proven.cycles);
    println!("shards    {}", proven.report.shards);
    println!("journal   {:?}", core::str::from_utf8(&proven.journal).unwrap());
}
sh
cargo run --release -p host --example prove_hello
identity  606d1f1d720459cc1a078787381656b29c9fce5a9e539b36f899e62b64129c14
cycles    114
shards    7
journal   "hello, apogee"

On an 18-core laptop with 48 GiB this took 52 seconds and peaked at 18 GB of memory, almost all of it the two 2^20 shards in flight. The identity differs from step 5's because the heights do: the identity binds every height.

Keep the identity#

A verifier never takes the identity from the proof, the key or the prover. It holds its own copy, obtained from whoever built the release, and compares:

rust
assert_eq!(setup.vk.identity.to_bytes(), registered); // `registered` from your own channel

Against an identity the prover supplied, a proof shows only that some program ran.

What just happened#

The emulator ran the 114 instructions twice. The first pass committed the memory columns of every shard and fixed the statement. The second filled each shard and proved it. There were seven shards: one for each of the four instruction families that executed, one for the memory window that holds the program's image, and one each for the public input and the journal. This guest never touched its stack, so no other window needed one; a typical program adds the stack's. Each shard was proved by its family's GKR circuit and opened with one Mercury proof, and the verifier reconciled the memory reads and writes of all seven in one equation. The architecture overview follows the same path in detail.

Next#

Launch Your App

Set Up

The toolchain the repository pins, the two workspaces it holds, the ceremony file proving needs, and the machine each step asks for.

View as Markdown

Apogee v1.0.0 is a Rust repository. There is nothing to install beyond rustup: the repository pins its toolchain, and the toolchain carries the guest target. Writing, building, running and profiling a guest need nothing else. Proving adds one large file and a machine with memory to match.

The toolchain#

rust-toolchain.toml at the repository root pins stable Rust 1.96.1 with rustfmt, clippy and llvm-tools, and the target riscv32imac-unknown-none-elf, whose core and alloc ship prebuilt. rustup applies it in every directory below the root and installs it on first use.

sh
cd apogee-vm
rustup show active-toolchain     # 1.96.1, overridden by rust-toolchain.toml
cargo --version

No nightly and no unstable features are used anywhere. llvm-tools supplies the llvm-objdump and llvm-nm matching the compiler's LLVM, which the repository uses for its committed disassembly listings and which you can use to read your guest's code.

Two workspaces#

The checkout holds two Cargo workspaces, and the split matters:

Workspace Root Builds for Holds
The root workspace Cargo.toml your host the prover, the verifier, the host SDK, the tools, everything in crates/ and tools/
The guest workspace guests/Cargo.toml riscv32imac-unknown-none-elf every guest, with its own guests/target directory

Guests are kept apart because every member compiles for the guest target and links a #[panic_handler]; cargo test --workspace at the root must never reach them. The guest workspace also carries what a guest needs to be built correctly, so you never type it:

  • guests/.cargo/config.toml sets the target and passes the linker -T crates/guest-sdk/link.ld, the memory map, and --no-relax, because relaxation would move addresses that the program identity binds.
  • guests/Cargo.toml pins both build profiles to the same semantics, overflow checks included (Build and inspect).
  • Its [patch.crates-io] routes k256, ark-ff and revm-precompile to vendored copies that call Apogee's delegations (Delegations).

Tip

Open guests/ as its own folder in your editor. rust-analyzer then reads that workspace's .cargo/config.toml and checks guest code against the guest target instead of your host.

The ceremony file#

Every commitment Apogee makes is under the powers of a secret τ from a public ceremony: the PSE perpetual powers of tau, contribution 80. One file serves every use:

assets/ptau/ppot_0080_24.ptau        19.3 GB, 2^24 powers; assets/ptau/ is gitignored

You need it to compute a program identity, to build real keys and to prove. You do not need it to build, run or profile a guest, or to run the workspace tests, which prove over toy setups of their own.

Files from PSE's ceremony are cut from one transcript, so any file of power 24 or more serves. Hermez's powersOfTau28_hez_final_*.ptau is a different ceremony with a different τ: the reader ingests it as readily, and every commitment, key and identity over it comes out different. To confirm you hold the right ceremony, its [τ]_1, as the hex of its canonical encoding x ‖ y, is:

9bbb31bedc304e081e2aada4b56c2217e0e94ee16874e3517d14bef5dcec3a16
317ff1589e53513fa333591b318e8f1e55ef7c37d92beb6d9a61d770a2f39506

The specification's SRS page states exactly what the reader checks and what it takes on trust.

The machine#

Building and running are laptop work. Proving is memory-bound, and its memory follows the shards being proved at once, not the length of the run.

Step Needs
Build a guest, run it, profile it, export and inspect its image Any recent laptop; seconds
Compute a program identity (artifact-dump tables --ptau) The ceremony file; about 25 s on an 18-core laptop at the default heights
Prove a small guest at 2^20 heights Tens of GiB. One 2^20 shard of the widest instruction family holds about 8.4 GiB of field elements in its forward pass, and each shard in flight holds its own
Prove a full Ethereum block The measured block peaked at 174 GiB on a 32-vCPU, 247.7 GiB machine

The proving page explains how heights and the number of shards in flight trade memory against time: Prove and verify.

Check your checkout#

What CI runs, all of it without the ceremony file:

sh
cargo fmt --all -- --check
cargo clippy --workspace --all-targets -- -D warnings
cargo test --workspace
cargo run -p kat-gen && git diff --exit-code      # committed fixtures regenerate identically
(cd guests && cargo clippy --bins -- -D warnings)

The suites that prove real shards are #[ignore]d, because each needs tens of GiB. Run one by name when you want to see a proof made and refused on your own machine:

sh
cargo test --release -p prover --test acceptance -- --include-ignored --test-threads=1

Next: write a guest, or run the whole loop once in the Quickstart.

Launch Your App

Write a Guest

A guest is a no_std Rust binary with an entry point and three memory regions. Crate layout, the runtime underneath you, dependencies, and the host-first layout that lets you test it like any other Rust.

View as Markdown

The crate#

In v1.0.0 a guest is a binary crate in the repository's guests/ workspace. The workspace supplies the target, the linker flags, the pinned profiles and the vendored crates, so a guest's own manifest stays short:

guests/my-app/Cargo.tomltoml
[package]
name = "my-app"
version.workspace = true
edition.workspace = true
publish.workspace = true

[dependencies]
guest-sdk.workspace = true
guests/my-app/src/main.rsrust
#![no_std]
#![no_main]

extern crate alloc; // Vec, Box, String, BTreeMap, over the SDK's allocator

use alloc::vec::Vec;

guest_sdk::entry!(main);

fn main() {
    let input = guest_sdk::public_input();
    let mut out = Vec::with_capacity(input.len());
    out.extend(input.iter().rev());
    guest_sdk::commit(&out);
}

Add "my-app" to members in guests/Cargo.toml, and build from the guest's own directory: cargo build --release --target riscv32imac-unknown-none-elf.

What runs underneath you#

The guest SDK is the entire runtime. It is small enough to state in full:

  • Start. _start sits at 0x0001_0000, the first byte of .text. It points sp at the top of RAM, zeroes .bss byte by byte, and calls main. entry!(f) exports that main as a wrapper around your function, which takes no arguments and returns ().
  • Exit. Returning from main is exit(0). guest_sdk::exit(code) ends the run with any status. A nonzero status is a failed execution, and a failed execution is still provable: the statement carries the status, and a verifier reads it.
  • Panic. The panic handler exits with status 101 and writes nothing. There is no diagnostic stream. A panicking guest has still published whatever it committed before the panic.
  • Heap. A bump allocator grows up from __heap_start, just above .bss. It never frees. See the heap.
  • System calls. The only ecalls a guest makes are EXIT and the delegation calls the SDK makes for you. Input, advice and output are memory, read and written with ordinary loads and stores.

The memory map#

The whole 32-bit address space, as a guest sees it:

Range What it is
0x0000_0000 – 0x0000_8000 A hole. Nothing initializes it, so a null or wild pointer is a fatal OutOfBounds, not a silent read
0x0000_8000 – 0x0000_C000 The public input window, 16 KiB
0x0000_C000 – 0x0001_0000 The journal window, 16 KiB
0x0001_0000 – … .text (with _start first), then .rodata, .data and .bss, each page-aligned
__heap_start upward The heap, from the end of .bss rounded up to 16
0x7F80_0000 – 0x8000_0000 The stack's 8 MiB reserve. No heap block may end above 0x7F80_0000; the stack grows down from 0x8000_0000
0x8000_0000 – 2^32 The advice region, up to 2^29 words, addressable only as far as the host supplied

Code is static. Each pc's instruction comes from the program's decoded tables, never from RAM, so a store into .text changes what a later load reads but not what executes.

The heap#

The allocator bumps a pointer and dealloc does nothing. That is the right design for a short program whose every cycle costs proving time, and it changes how you write Rust:

  • What runs you out of memory is the total you allocate, not your peak. A loop that builds and drops a Vec each iteration consumes fresh heap every time.
  • Reuse buffers. Hoist allocations out of loops, clear() and refill instead of reallocating, and size growing collections with with_capacity so they do not reallocate and copy as they grow.
  • The ceiling is exit 71. An allocation that would end above 0x7F80_0000, or above the live stack pointer, exits with status 71 rather than return null or overwrite the stack.
rust
// Allocates a fresh Vec per record: total heap grows with the record count.
for record in records {
    let fields: Vec<&[u8]> = record.split(|b| *b == b',').collect();
    handle(&fields);
}

// One buffer, reused: total heap is the largest record's field count.
let mut fields: Vec<&[u8]> = Vec::with_capacity(16);
for record in records {
    fields.clear();
    fields.extend(record.split(|b| *b == b','));
    handle(&fields);
}

The stack has its 8 MiB reserve, and deep recursion inside it is fine. What nothing detects is a stack that grows past the reserve after the heap has filled the space below it: heap blocks would then change under a deep call chain. Keep recursion bounded, or make it iterative.

Dependencies#

Any crate that builds for riscv32imac-unknown-none-elf without std will do. In practice:

  • Turn off default features (default-features = false) and enable alloc where a crate offers it.
  • A crate that pulls in getrandom, a clock or std::collections::HashMap with its random seed has nothing to draw on. Such a call answers -ENOSYS and leaves the run unprovable. Prefer BTreeMap, or a hash map with a fixed, deterministic hasher.
  • Floating point compiles to integer software routines, because the target has no F or D extension. It is correct and deterministic, and costs many instructions per operation. Integer or fixed-point arithmetic is cheaper.
  • Hashing and elliptic-curve arithmetic have dedicated circuits. Use the SDK's functions or the vendored crates so your dependencies reach them: Delegations.

Test it on the host first#

A guest prints nothing, so debugging happens on the host. The layout that makes this easy keeps the program in a #![no_std] library from bytes to bytes, keeps main.rs down to moving bytes in and out of the regions, and makes the SDK a dependency of the guest target only:

guests/my-app/Cargo.tomltoml
[package]
name = "my-app"
version.workspace = true
edition.workspace = true
publish.workspace = true

[target.'cfg(target_arch = "riscv32")'.dependencies]
guest-sdk.workspace = true
guests/my-app/src/lib.rsrust
#![no_std]
extern crate alloc;
use alloc::vec::Vec;

/// The whole application: public input and advice in, journal out.
pub fn run(input: &[u8], advice: &[u8]) -> Result<Vec<u8>, i32> {
    let _ = advice;
    let mut out = Vec::with_capacity(input.len());
    out.extend(input.iter().rev());
    Ok(out)
}
guests/my-app/src/main.rsrust
#![no_std]
#![no_main]

guest_sdk::entry!(main);

fn main() {
    match my_app::run(guest_sdk::public_input(), &[]) {
        Ok(journal) => guest_sdk::commit(&journal),
        Err(code) => guest_sdk::exit(code),
    }
}

Host code then depends on the library by path, as crates/emulator depends on guests/revm-block, runs my_app::run natively, and compares the result with the journal the emulator produces for the same input (Run and profile). Your logic gets unit tests, a debugger and println! on the host, and the guest binary stays a thin shell around code you have already tested.

Warning

The two builds disagree about usize. On the guest usize and every pointer are 32 bits; on your host they are 64. Overflowing a usize panics on the guest alone, x as usize truncates silently there, and size_of and core::hash of anything holding a length differ between the two. Keep usize out of anything you commit, hash or serialize, and use explicit u32 and u64 at those boundaries.

Assembly and the instruction set#

The decoder accepts exactly the 59 instructions of RV32IMA, and compressed (C) instructions, which are expanded at load. Inline assembly is fine within that set. Anything outside it, such as a CSR access, fence.i, a floating-point or an RV64 encoding, makes the whole program refuse to register, even if it is never reached: derivation reports Not all opcodes supported: pc=…. An ebreak, a jump to a halfword with no instruction, or a misaligned halfword or word access ends the run with no proof.

Atomic instructions decode and prove, with one deviation: sc.w always succeeds, because the machine keeps no reservation state. The guest programming guide explains why new guest code should not use atomics at all.

Next#

Launch Your App

Inputs, Advice and the Journal

A guest has no I/O system calls. Its public input, the prover's advice and its journal are three regions of memory. What each holds, what the proof binds, and the pattern every guest that takes data follows.

View as Markdown

An Apogee guest has no file descriptors, no streams and no I/O syscall. Its inputs and outputs are three regions of memory, read and written with ordinary loads and stores, and the proof binds two of them.

The three regions#

Region SDK Holds Size Bound by the proof
Public input public_input(), read_input(buf) the statement's bytes, chosen by whoever asks for the proof at most 16,380 bytes yes, its initial contents
Advice advice() bytes the prover chooses up to 2 GiB no
Journal commit(bytes), journal() what the guest appended at most 16,380 bytes yes, its final contents
rust
let input: &[u8] = guest_sdk::public_input(); // a slice over the input window, no copy
let data: &[u8] = guest_sdk::advice();        // a slice over the advice region
guest_sdk::commit(b"result");                  // appends to the journal
  • public_input() and advice() return slices over memory; nothing is copied. read_input(buf) copies min(buf.len(), input.len()) bytes and returns the count, so it may return short.
  • commit appends and keeps a length word, which is what makes the proof bind an exact byte string rather than a zero-padded window. It exits with status 70 rather than overflow the window.
  • advice() on a run given no advice is a fatal OutOfBounds, not an empty slice: a run with no advice has no advice region at all, and pays nothing for one.

What "bound" means#

The statement a proof establishes carries the public input bytes, the journal bytes and the exit status. The proof shows that the input window held exactly the statement's input before the guest's first access, and that the journal window held exactly the statement's output when the guest exited. That rests on the memory argument, not on anything the guest does: there is no hash the guest must compute and no convention it must follow.

Advice is different. The advice region's initial contents are whatever the prover wrote there, and nothing ties them to the program identity, the statement or any gate. A proof says that some advice exists under which the program, given this input, published this journal. That is exactly as strong as the guest's own checks on the advice.

The pattern: commit, supply, check#

A guest with a large input takes the bulk as advice, which the public input commits to, and checks one against the other before anything derived from the advice reaches the journal:

Check the advice before trusting itrust
fn main() {
    let want = guest_sdk::public_input(); // 32 bytes: keccak256 of the advice
    let data = guest_sdk::advice();       // the prover's bytes, bound by nothing
    if guest_sdk::keccak256(data).as_slice() != want {
        guest_sdk::exit(1); // refused before anything derived from it is committed
    }
    let sum = data.iter().fold(0u32, |s, b| s.wrapping_add(u32::from(*b)));
    guest_sdk::commit(&sum.to_le_bytes()); // the journal: what the proof publishes
}

The check need not be a hash of the whole advice. It can be a Merkle path checked against a root the input carries, as in the ledger example, or a signature over the data, or a constraint the result itself satisfies, such as a claimed sorted order that the guest verifies in one pass instead of sorting. What matters is that the thing it is checked against is bound.

Caution

Committing any function of unchecked advice publishes a value the prover chose. This is the most common way to write a guest whose proof means nothing.

Structured data#

Encode structured inputs with a no_std serializer such as postcard over serde with the alloc feature, which is what the repository's own Ethereum guest uses for its block witness. Two habits keep a format honest:

  • Use fixed-width integers. u32 and u64, never usize, whose width differs between your host and the guest.
  • Insist on one encoding per value when it matters. A deserializer that accepts trailing bytes or non-minimal varints admits two byte strings for one value. Where uniqueness matters, decode, re-encode and compare, as the Ethereum guest's BlockWitness::decode does.

Outputs that grow#

The journal holds 16,380 bytes. An output that grows with the work, such as one record per transaction, has no fixed bound and will eventually exit 70. Publish a digest instead: hash the records as you produce them and commit the 32-byte result, then let whoever needs the records recompute them natively and compare. The repository's stateless Ethereum validator publishes a 43-byte journal for a whole block this way.

For a proof checked on Ethereum, keep both public values a fixed length. The deployed verifier contract is built for one input length and one output length, and refuses anything else (Settle on-chain).

The exit status#

Nothing is published at exit beyond what was committed, and a run that panics or exits nonzero has a valid proof of what it did. So a verifier reads the exit status before it reads the journal. On-chain, the verifier contract takes the expected status as an argument, and an application passes 0. The SDK's own statuses:

Status Meaning
0 main returned, or exit(0)
70 commit would have passed 16,380 bytes
71 an allocation would have reached the stack: see the heap
72 a delegation answered something its shim refuses
101 a panic, which prints nothing

Choose your own failure codes outside these, as the ledger example does with 1, 2 and 3.

What the proof does not say#

  • Nothing orders the journal's writes, and nothing forces a guest to read its input. The proof binds the windows' contents, not the accesses that produced them.
  • Advice is writable. A store into the advice region is an ordinary store. It is still unbound either way.
  • The journal is the window's whole final contents. commit maintains that form. A guest that writes the window directly must keep it: a length of at most 16,380, that many bytes, then zeros.

The specification states all of this precisely: Public values and advice.

Launch Your App

Delegations

Hashing, field and curve arithmetic have dedicated circuits. Which SDK calls reach them, what they cost, the rules on their operands, and the vendored crates that route library code to them.

View as Markdown

Some computations are far cheaper to prove with a circuit built for them than as a stream of RISC-V instructions. Apogee calls these delegations. A delegation is a circuit family that proves one function of a frame of words in RAM, invoked by an ecall, and the guest SDK makes those calls for you behind ordinary functions. You never write an ecall yourself.

What you call, and what it reaches#

You call Delegation One call proves
guest_sdk::keccak256(&[u8]) -> [u8; 32] KECCAK_F one round of keccak-f[1600]; a permutation is 24 calls, and the sponge and padding are guest code
guest_sdk::sha256(&[u8]) -> [u8; 32] SHA256_COMP four rounds of the compression; a compression is 16 calls
guest_sdk::ec_add, ec_mul, ec_identity EC_ADD one third of a complete point addition on secp256k1 or BN254 G1
guest_sdk::poseidon2_permute(&mut [u8; 96]) POSEIDON2 one width-3 Poseidon2 permutation over Fr
field::Fr addition, multiplication, inversion FR_ARITH one Fr operation, on the guest target, with nothing named
transcript::poseidon2_permute POSEIDON2 the same permutation, through the transcript crate
guest_sdk::recursion::mod_mul over a ModMulFrame MOD_MUL one 256-bit a·b mod m, m one of four Ethereum moduli

The functions are bit-identical to their software definitions. keccak256 is Ethereum's Keccak, not SHA3-256. sha256 is FIPS 180-4. ec_add uses the complete formula of Renes, Costello and Batina (2015, Algorithm 7), so doubling, P + (−P), the identity and any Z need no special case.

Hashing and curve arithmetic from a guestrust
use guest_sdk::{ec_mul, keccak256, recursion::SECP256K1_GROUPS, ProjectivePoint};

let digest: [u8; 32] = keccak256(b"blockchain-native");

// A point is homogeneous projective (x = X/Z, y = Y/Z), each coordinate eight
// little-endian u32 limbs below the field modulus. The scalar is eight limbs too.
fn times(p: &ProjectivePoint, k: &[u32; 8]) -> ProjectivePoint {
    ec_mul(&SECP256K1_GROUPS, p, k).expect("EC_ADD is implemented on Apogee")
}

Library code reaches them too#

The guest workspace patches three crates so that the code inside them calls delegations on the guest target, with upstream's code as the fallback path:

Crate Version Reaches
k256 0.13.4 MOD_MUL from field and scalar multiplication; EC_ADD from ProjectivePoint addition, mixed addition and doubling
ark-ff 0.6.0 MOD_MUL from BN254's Montgomery multiply and square, in both of its fields
revm-precompile 43.0.2 SHA256_COMP for precompile 0x02; EC_ADD for 0x06 and 0x07

A guest that depends on these crates gets the patched copies automatically through guests/Cargo.toml's [patch.crates-io]. Unpatched, k256's field multiply and square alone were 44% of a mainnet block's cycles. secp256k1 signature recovery is ordinary k256 code, which the patches turn into delegated arithmetic.

What a delegation costs#

A delegation family is part of a program only if the program links one of its shims, and a call costs shards of that family's height:

  • Linked and never called: nothing. The family is declared and proves zero shards.
  • Called once: a whole shard. A shard costs its full height whatever its occupancy.
  • Called a lot: very little per call. A shard's proof grows only by one sumcheck round per variable as its height grows.
Family Height Unit of work Calls per unit Units per shard
KECCAK_F 2^18 keccak-f[1600] 24 10,922
SHA256_COMP 2^18 one compression 16 16,384
EC_ADD 2^16 one complete addition 3 21,845
MOD_MUL 2^16 one a·b mod m 1 65,536
POSEIDON2 2^8 one permutation 1 256
FR_ARITH 2^8 one Fr operation 1 256

The price is memory more than time: a 2^18 KECCAK_F shard's forward pass holds about 42 GiB of field elements, and two of them in flight set the measured Ethereum block's peak.

Rules on operands#

  • Operands below their modulus. A MOD_MUL or EC_ADD operand at or above the modulus its selector names has no proof: the executor refuses the frame as a fatal DelegationFrame. The vendored k256 reduces its lazily reduced field elements before calling.
  • Points are not checked for you. EC_ADD proves the formula's arithmetic. Whether a point lies on the curve is the calling code's question, and a guest that takes points from advice must ask it.
  • Multi-call operations are one SDK function each. A keccak permutation is 24 calls on one frame, a SHA-256 compression 16, a point addition 3. Each call proves its own step, and nothing refuses a wrong order: it computes something else. Use keccak256, sha256 and ec_add, which issue the calls in order, rather than the raw shims.
  • Exit 72 means a delegation answered something its shim refuses. On Apogee's own executor this does not happen for a well-formed frame.

What is not delegated#

  • The EVM's MULMOD with an arbitrary modulus, MODEXP, BLS12-381, and every primitive outside the table above run as instructions.
  • No signature scheme or pairing is delegated as a whole. secp256k1 recovery is k256 over MOD_MUL and EC_ADD; a BN254 pairing is ark-bn254 over MOD_MUL.
  • A delegation is an operation's core. Padding, sponges, block loops and a scalar multiplication's ladder are guest code, proved as instructions.

Signature schemes for guests are on the v2.0.0 roadmap.

Is a new delegation worth it?#

The cycle profiler prices the obvious candidates in every report, as a ceiling on the cycles a delegation could remove:

removable = max(0, cycles − calls·(4 + 2·frame_words))

cycles is the category's share of the run, calls the entries into the candidate's functions, and 4 + 2·frame_words the shim a delegation would leave behind: the frame's stores, the ecall and the result's loads. It charges nothing for the new family's shards, so treat it as an upper bound. Run and profile shows a report.

The specification of each delegation, column by column, is under Delegation circuits.

Launch Your App

Build and Inspect

Build profiles pinned to one semantics, the ELF a build produces, the ProgramImage the loader makes of it, its report, and the program identity a verifier registers.

View as Markdown

Build#

From the guest's directory, with no flag but the target:

sh
cd guests/my-app
cargo build --release --target riscv32imac-unknown-none-elf     # .../release/my-app
cargo build --target riscv32imac-unknown-none-elf               # .../debug/my-app

The ELF lands in the guest workspace's one target directory, guests/target/riscv32imac-unknown-none-elf/. guests/.cargo/config.toml adds two linker arguments that you never type:

  • -T crates/guest-sdk/link.ld, the memory map, which also defines the symbols the startup code and the allocator use.
  • --no-relax. Linker relaxation rewrites instruction sequences and shifts every later address, and the program identity binds those addresses.

There is no runner: nothing outside Apogee maps a guest's regions, so cargo run has nothing to run it with. You run a guest through the emulator (Run and profile).

Profiles#

guests/Cargo.toml pins both profiles to one semantics. They differ only in optimization and in dependencies' debug assertions:

dev release
opt-level 0 3
overflow-checks on on
debug-assertions on on in the guest crate, off in its dependencies
panic, codegen-units, debug, incremental abort, 1, off, off the same

Cargo's default release profile turns overflow checks off, and in a guest that is not a performance setting. It changes the statement: u32::MAX + 1 would commit 00000000 and exit 0 where the dev build panics with exit 101. So the workspace keeps the checks on in both profiles. A dependency's debug assertion checks that crate's own invariant, and a correct dependency computes the same without it, so release turns those off and keeps the cycles: 6.8% of the stateless Ethereum guest's run.

Prove the release build. Every executed instruction is a proved row, opt-level = 3 removes a quarter to over half of a guest's image, and each family's code must fit its decoded table: the Ethereum guest's debug image needs 2^22-row tables, its release image 2^20. The identity you publish is the release image's.

Reproducibility#

Two clean builds on one machine produce identical ELFs. Builds on two machines generally do not: the ELF embeds absolute paths in panic-location strings, the toolchain's core sources, crates/guest-sdk and the cargo registry, while the guest's own files appear relative to guests/. A build elsewhere is another image with another identity.

So what you register and hand on is a build's ELF, not a recipe. Keep the ELF you proved, and let anyone who wants to check the identity recompute it from that ELF, the parameters and the ceremony file.

Export the image#

sh
cargo run -p artifact-dump -- guests/target/riscv32imac-unknown-none-elf/release/my-app --out artifacts

This writes artifacts/my-app.img, the loaded ProgramImage in its wire form (postcard, no header), and artifacts/my-app.img.txt, a report rendered from the image read back through the validating reader. It prints the entry point, the segment and instruction counts, and the artifact's size and SHA-256. If the read-back differs, or the loader refuses the ELF, it writes nothing.

The .img is the program's static description, for keeping and diffing; nothing downstream needs it, since the setup and the tools take the ELF. Its SHA-256 pins bytes. It is not the program identity.

Read the report#

Section Shows
entry and memory the entry, _start at 0x00010000; the RAM window; slot_base and the slot span
segments address, end, mem_len, file bytes, zero fill and instruction count of each segment: .text, .rodata if any, and one writable segment to 0x80000000 for .data, .bss, heap and stack
instruction stream four- and two-byte instructions, mid-instruction slots and not code slots, summing to the slot count
symbols names by address from the ELF's symbol table, which the artifact does not carry
listing per instruction: address, length, the bytes in memory, the expanded 32-bit word, the symbol

A compressed instruction keeps its address and its two bytes; len alone says whether the next pc is pc + 2 or pc + 4. For mnemonics, use the tables view below or the pinned disassembler:

sh
"$(rustc --print sysroot)"/lib/rustlib/*/bin/llvm-objdump \
    --disassemble --no-print-imm-hex -M no-aliases <elf>

The not code halfwords#

A report may show a line such as ---- not code: 0x00010f9a .. 0x00010f9c, 1 halfword ----. That is ordinary compiler output. LLVM proved a match's default arm unreachable, rustc lowered the unreachable block to unimp, and with the C extension that is c.unimp, the all-zero halfword, RVC's defined-illegal encoding. The loader records it as a non-instruction and moves on. No pc reaches it; one that did would stop the run with NotAnInstruction.

What the VM will prove#

sh
cargo run --release -p artifact-dump -- tables <elf> --ptau assets/ptau/ppot_0080_24.ptau

This prints, at the default parameters, the VmConfig the image derives: each family's height, its live rows and decoded columns. Then it prints each instruction's pc, next_pc, family, mnemonic and fields. With --ptau and the ceremony file it also prints the program identity, the value a verifier registers. The Quickstart shows a real one.

The identity is one field element. It binds every instruction with its pc, length, operands and kind, every file-backed byte of the image (.text, .rodata, .data), the entry point, the family set, every height, the code-size ceiling and the code version. It does not bind the symbol table, .bss, or anything an execution chooses. One ELF at two settings of the heights has two identities.

Derivation refuses, naming the pc or the size:

Refusal Cause
Not all opcodes supported: pc=… a word outside RV32IMA anywhere in executable code, such as a CSR access in assembly
TableTooShort code past a family's reach, pc ≤ 2h − 4: 1.9375 MiB of code at 2^20 and 7.9375 MiB at 2^22
ProgramTooLarge the image is past bytecode_size_words, 4 MiB by default
ImageOutsideWindow a file-backed byte lies past RAM window 0 at the chosen window height
UnknownDelegation the image declares a delegation number no family answers

Check a build#

Build into a fresh target directory, export again, and compare:

sh
cd guests/my-app
CARGO_TARGET_DIR=/tmp/fresh cargo build --release --target riscv32imac-unknown-none-elf
cd ../..
cargo run -p artifact-dump -- /tmp/fresh/riscv32imac-unknown-none-elf/release/my-app --out /tmp/again
cmp artifacts/my-app.img /tmp/again/my-app.img
diff artifacts/my-app.img.txt /tmp/again/my-app.img.txt     # differs only in the `source ELF` line

Launch Your App

Run and Profile

Run a guest in Apogee's emulator from Rust or from the command line, compare it with your host build, and find out where its cycles go before you pay to prove them.

View as Markdown

Running a guest costs almost nothing; proving it costs in proportion to the cycles it runs. So run first, compare against your host build, and look at the cycle profile before you prove anything.

From Rust: the emulator#

emulator::run executes a loaded image over a public input and advice, in host code, with no proof:

Run a guest and compare it with the host buildrust
let elf = std::fs::read(elf_path)?;
let image = loader::load_elf(&elf).expect("the ELF loads");
let io = emulator::GuestIo { input: b"hi".to_vec(), advice: Vec::new() };
let run = emulator::run(&image, &io).expect("no fatal error");

assert_eq!(run.exit_code, 0);
assert_eq!(run.io.output, my_app::run(b"hi", &[]).unwrap()); // the host build agrees
println!("{} cycles", run.cycle_count);

run returns an Execution: the final registers, the exit status, the cycle count and the public values. A nonzero exit status is an execution, not an error, and comes back as exit_code. A fatal executor error, such as OutOfBounds, Misaligned or NotAnInstruction, comes back as an EmuError, and such a run has no proof (Troubleshooting).

The emulator is a pure function of the image and the input: no clock, no randomness, no threads. The same input gives the same execution, cycle for cycle, which is also what lets the prover execute twice and cut identical shards.

From the command line: the profiler#

sh
cargo run --release -p profiler -- elf <elf> [--input <file>] [--advice <file>] [--top <n>] [--json <path>]

It runs the guest over the given files at the smallest table height its code fits, and prints a report. Its numbers are counts of executed cycles, the same on any machine.

workload
  label                        hello
  guest cycles                 114
  exit status                  0
  journal bytes                13

cycles by semantic workload
  core runtime                             94   82.46%
  unattributed                             20   17.54%

cycles by family
  ADD_SUB_LUI_AUIPC            64
  JUMP_BRANCH_SLT              21
  MEM_WORD                     3
  MEM_SUBWORD                  26

top functions
            94   82.46%          1 calls        94.0 c/call  guest_sdk::commit  [core runtime]
             8    7.02%          1 calls         8.0 c/call  main  [unattributed]

How to read it:

  • Cycles by family is what you pay for. Each family with rows costs at least one shard of its height, and more cycles in a family means more shards of it.
  • Top functions charges each function for its own cycles, including everything the compiler inlined into it, but not its callees. Calls are counted at the function's first instruction.
  • Cycles by semantic workload groups functions into fourteen categories by name, such as hashing, signatures and the core runtime. The unattributed share and the mnemonic mix are the checks on that attribution, since no symbol table can mislabel them.
  • Accelerator candidates price the delegations a future version could add, as a ceiling: see Delegations.

The profiler has two more verbs, for the Ethereum workload: block <stem> runs the revm guest over a recorded fixture, and record <number|latest> records a block from ETH_RPC_URL and runs it.

Make it cheaper#

The order that usually pays:

  1. Build --release. Optimization removes a quarter to over half of a guest's instructions.
  2. Delegate hashing and curve arithmetic. Use guest_sdk::keccak256, sha256, ec_add and the vendored k256 and ark-ff, instead of compiling a software implementation into the guest.
  3. Stop allocating in loops. Each allocation is instructions, and with a bump allocator it is also memory you never get back (the heap).
  4. Check instead of compute. If a result is expensive to find and cheap to verify, such as a sorted order, a square root or a path through a tree, let the prover supply it as advice and have the guest verify it.
  5. Avoid floating point. It compiles to software routines; integer and fixed-point arithmetic are far cheaper.

Then measure again. Cycle counts are exact and repeatable, so every change shows up as a number.

Launch Your App

Prove and Verify

Register a program, prove a run, verify the block, and keep the proof. Heights, shards in flight, the ceremony powers a proof needs, and the two values a verifier must hold for itself.

View as Markdown

The three calls#

Setup, prove, verifyrust
let params = program::ProgramParams::defaults();
let ptau = std::path::Path::new("assets/ptau/ppot_0080_24.ptau");
let srs = srs::Srs::from_ptau(ptau, 22).expect("ceremony");

let setup = host::setup(&elf, &params, srs).expect("registers");          // once per program
let proven = host::prove(&setup, &io, 4).expect("proves");                // at most 4 shards in flight
host::verify(&setup.vk, &proven.block).expect("verifies");

assert_eq!(setup.vk.identity.to_bytes(), registered); // from your own channel, never the proof
assert_eq!(proven.exit_code, 0);
  • host::setup loads the ELF, decodes it into its family tables and VmConfig, commits the setup columns under the ceremony, and builds the verifying key. Its cost is per program and per choice of heights, not per run.
  • host::prove executes the guest twice and proves every shard (below). It returns a Proven: the BlockProof, the exit code, the cycle count, the journal and a report of the run.
  • host::verify checks the block against the key, using the statement the block carries. It compares neither the identity nor the SRS digest with anything, so that comparison is yours.

The Quickstart runs exactly this code over a small guest, with its real output.

What a verifier must hold for itself#

Two values come from a channel the prover does not control:

  1. The program identity. Against an identity the prover supplied, a proof shows only that some program ran. A verifier registers the identity of the release it trusts and compares it with the key's.
  2. The SRS digest of the ceremony. A key loads under whatever digest its own points give. One built over a known τ could open anything, and is refused only by comparing its digest with the ceremony's.

The verifying key itself may come from anyone, including the prover: loading it recomputes the identity and the SRS digest from its own contents and holds its circuits to the verifier's own registry. Then read the statement: the exit status first, then the journal.

Heights#

Each family has a height, the number of rows in one of its shards, chosen from 2^8, 2^12, 2^16, 2^18, 2^20, 2^22. Heights are part of the program, not of a run: every height is bound into the identity.

Family group Default Floor Notes
The seven instruction families 2^22, or 2^20 for MUL_DIV and ATOMICS 2^20 the floor of their timestamp range checks
INIT_TEARDOWN, ZERO_WINDOWS, ADVICE_WINDOWS 2^22 2^16 one shared window height; window 0 must hold every file-backed byte of the image
PUBLIC_INPUT, PUBLIC_OUTPUT 2^12 pinned the height places their windows
Delegation families see Delegations per family

A family with rows costs at least one whole shard of its height, so a short run wastes less at smaller heights, and a long one needs fewer shards at larger heights. A decoded table must also be tall enough to reach the family's last instruction: 2^20 reaches 1.9375 MiB of code and 2^22 reaches 7.9375 MiB. The Ethereum guest proves at 2^20 for every family whose height is a choice.

Instruction families at their floor, RAM windows at 2^16rust
use constants::family;

let mut params = program::ProgramParams::defaults();
for f in 0..7 {
    params.heights[f] = 1 << 20;
}
for f in [family::INIT_TEARDOWN, family::ZERO_WINDOWS, family::ADVICE_WINDOWS] {
    params.heights[f as usize] = 1 << 16; // window 0 is then 256 KiB: the image must fit in it
}

The ceremony must supply as many powers as the tallest family has rows, and at least 2^18 for the generic lookup table: Srs::from_ptau(path, k) with 2^k at least the largest height.

Shards in flight#

The third argument of host::prove is max_in_flight, the number of shards proved at once. It is the one knob that trades memory for time:

  • Memory follows the shards in flight, not the cycle count. Each shard in flight holds its rows, its forward pass and its proof as it grows. A 2^20 shard of the widest instruction family holds about 8.4 GiB in its forward pass; a 2^18 KECCAK_F shard about 42 GiB.
  • Time follows how many shards run side by side, up to the cores you have. Within a shard, the work runs on all cores.
  • The proof does not depend on it. The block is byte-identical at 1 and 8 in flight.

Start low on a laptop, two or four, and raise it on a server until memory, not cores, is the limit. bench prove defaults to 8.

The two passes#

host::prove streams. It never holds the whole execution trace, which at about 300 bytes a cycle would be the largest object in the system.

  1. Pass 1 executes the guest and, as each shard fills, commits its memory columns, keeps the commitments and drops the rows. At the exit it builds the statement and draws the challenges every shard shares.
  2. Pass 2 executes again. The emulator is deterministic, so it cuts the same shards. Each is filled, proved and dropped as it arrives, and only its proof is kept.

That is why proving takes two executions' time and memory bounded by the shards in flight. The streaming prover explains it in depth.

Keep the proof#

host::proof_archive::write_proof(dir, stem, vk, block) writes four files, each its type's plain bytes:

<stem>.vk         the verifying key
<stem>.identity   the key's identity, 64 lowercase hex digits and a newline
<stem>.public     the statement: input, journal, exit status, the execution's record
<stem>.block      the block proof

read_proof(dir, stem) reads them back. The .identity file records what the run claimed; a verifier still compares against its own copy. The verifier command-line tool checks an archive:

sh
cargo run --release -p verifier -- block <stem>.vk <identity-hex> <stem>.public <stem>.block

It exits 0 when everything verifies, 1 naming the first refusal, and 2 on a usage error or a malformed identity. It compares the identity you pass with the key's, and takes the SRS digest from the key file.

When a proof fails#

An honest prover never produces a proof that fails, so a failure means an input it should not have accepted, or a bug. Rebuild with the prover's debug log on and rerun:

sh
cargo run --release -p bench --features prover/debug-info -- prove ...
APOGEE_DEBUG=detail <the same run> 2>&1 | tee run.log
grep -E 'FAIL|NOT CANONICAL|UNBALANCED|OVER the|NAMES NO|DISAGREES|NOT LOOPING|ABORTED' run.log

The log names the shard that died, and self_check FAILED names the first gate a row breaks, with every operand's value. Tools lists every marker. The log exists only in builds with the debug-info feature and changes no proof byte.

Next: Settle on-chain.

Launch Your App

Settle On-Chain

From a base proof of hundreds of shards to one Groth16 proof that an Ethereum contract checks. The recursion tree, the decider's ceremony, the contract's interface, and what one deployment fixes.

View as Markdown

A base proof is a block of shard proofs, each a GKR proof with its commitments: megabytes of data and hundreds of curve points, which no contract can check. Settlement compresses it in three stages, each run with bench from the repository root.

From base proof to contract Base shards are verified by leaves, leaves by internal nodes, nodes by a root; a Groth16 decider re-verifies the root; the contract checks the Groth16 proof and the folded pairing. BASE PROOF 207 shards · 14.5 MB LEAVES ≤ 64 base shards each ROOT 2–4 children covers 0..count DECIDER Groth16 BN254 7.9M constraints CONTRACT verify(…) → true 3.62M gas
Settlement. Each stage verifies the one before it. Nothing pairs before the contract: every Mercury check is deferred and folded into one accumulator that the contract discharges with two pairings. Figures are block 257,510's.

1. The recursion tree#

A node is Apogee proving a verifier program. A leaf verifies a run of consecutive base shards; an internal node verifies two to four child proofs; the root covers every base shard. Each node also folds every Mercury check its shards and children defer into one pair of points, so the whole tree comes down to a single pairing claim at the top.

sh
# the base proof as an archive: host::proof_archive::write_proof from your host
# program, or `bench prove ... --out <dir>` for the Ethereum guests
cargo run --release -p bench -- recurse <dir>/<stem> --out <out> --in-flight 4

recurse writes the two recursion programs' keys, builds the leaf and node binaries over them, fixes a plan before anything is proved (<out>/tree.txt: leaves of at most --leaf 64 base shards, then nodes of at most --fan-in 4 children), and proves node by node. Each node is its own process, which verifies its inputs natively before proving, so a bad input is refused by name. A stopped run resumes: proofs already in <out> are kept, and a run whose plan or programs differ is refused.

Base proving is untouched by any of this. A leaf verifies base shards exactly as they are.

2. The decider#

The root is still a GKR proof and some hundreds of points. The decider is a Groth16 circuit that verifies the root as a node would, and folds nothing. Instead it binds, as wires whose values the contract supplies, the two recursion programs' identities, the base statement's exit status, its public input and journal byte by byte, and each point the root owes a pairing to, with its scalar. The Groth16 proof carries one commitment to all of those wires, and the contract checks it against the values it holds.

A Groth16 key needs a ceremony. Phase 1 is the same powers-of-tau file the tree's commitments are under. Phase 2 is the circuit's own and runs in two rounds of contributions:

sh
cargo run --release -p bench -- ceremony <out> init          # once per root shape
cargo run --release -p bench -- ceremony <out> contribute    # round 1: alpha and beta, each contributor in turn
cargo run --release -p bench -- ceremony <out> seal
cargo run --release -p bench -- ceremony <out> contribute    # round 2: gamma, delta and eta
cargo run --release -p bench -- ceremony <out> key
cargo run --release -p bench -- decide <out>                 # the Groth16 proof, checked natively and in an EVM

Each contribution multiplies a trapdoor by a factor only its contributor knew, and records it with a Schnorr proof, so any state can be verified against the circuit and the ceremony file alone. A trapdoor is unknown while one contributor to it was honest. The order of the rounds is part of soundness: alpha and beta are finished before anything is divided by delta or eta.

Caution

bench decide --dev-key derives every trapdoor from a public seed, for development and tests. Anyone can forge a proof under it, and its outputs are written as development.* so they cannot be mistaken for a ceremony's. A ceremony run on one machine is not a ceremony either: it needs one honest contributor per round.

decide writes decision.constructor and decision.calldata: the deployment arguments and the call, as hex.

3. The contract#

contracts/ApogeeVerifier.sol has one entry point:

solidity
function verify(
    bytes calldata input,        // the base program's public input
    bytes calldata output,       // its journal
    uint256 exitStatus,          // the status you require, normally 0
    uint256[10] calldata proof,  // Groth16 A, B, C and the bound wires' commitment D
    uint256[] calldata points    // x, y and scalar of each point, side [1]_2's then side [x]_2's
) external view returns (bool);

It rebuilds the bound values from the calldata, checks the Groth16 pairing equation, folds each side's points with ecMul and ecAdd, which also holds every point to the curve, and checks the folded claim e(A, [1]_2) = e(B, [x]_2). An application contract calls it and then acts on the journal: see the ledger sketch.

What one deployment fixes#

The constructor takes the Groth16 key, the ceremony's two G2 points, the leaf and node programs' identities, the number of points on each side, and the byte lengths of the public input and the journal. So one deployed verifier serves:

  • One base program. Its identity is a constant of the leaf program's image, which the leaf's identity binds.
  • One root shape. The decider's circuit depends on the root's program, its shard counts and the public values' lengths, so a key and its ceremony are per shape.
  • Fixed-length public values. verify refuses an input or journal of any other length. Design guests whose on-chain public values have a fixed size, such as a fixed record or a 32-byte digest.

The contract pays about 9,000 gas a point, because the circuit folds none of them.

Measured#

Block 257,510, with the tree on a 32-CPU, 247 GiB machine and the ceremony and decider on an 18-core laptop:

Stage Result
Base proof 207 shards, 14.5 MB, 2,481 s
Tree 4 leaves of at most 64 base shards and a root: 116 shards
Leaves, four at once 21, 24, 23 and 27 shards; 2,157 s; 92 GiB peak
Root, four shards in flight 21 shards, 460 s, 1.03 MB
Decider circuit 7,896,686 constraints, a domain of 2^23
Ceremony init 65 s; a contribution 50–56 s; key 70 s and 12.7 GB; the key 2.65 GB
Decider proof key read in 1 s, proof 18.5 s, 6.1 GB
Contract 358 points; 3,620,026 gas; 34,980 bytes of calldata

The specification of all of it is Recursion and decider.

Launch Your App

Guest Programming Guide

The habits that keep a guest correct, provable and cheap. Every gotcha and preference in one place, each with the reason behind it and what to do instead.

View as Markdown

Most of writing a guest is writing Rust. This page is about the rest: the places where a bare-metal, single-hart, proved machine behaves differently from the host you are used to. Each rule says what to do, why, and what goes wrong otherwise. The AI Companion carries the same rules in a form you can hand to a model.

Types and memory#

usize is 32 bits, and so is every pointer#

The guest target is riscv32imac: usize, isize and every pointer are 32 bits wide, while your host's are 64.

  • Overflowing a usize panics on the guest and not on the host.
  • x as usize from a u64 truncates silently on the guest.
  • size_of::<T>(), struct layout, and core::hash of anything holding a length or pointer differ between the two builds.

Do: use explicit u32 and u64 in anything you commit, hash, serialize or compare with a host computation. Convert with usize::try_from(x) where a value may not fit, so it fails loudly on both builds. Don't: commit usize, hash a structure containing one, or derive a layout-dependent encoding.

rust
let n = u64::from_le_bytes(input[..8].try_into().unwrap());
let len = usize::try_from(n).expect("length fits the guest"); // not `n as usize`

The allocator never frees#

The heap is a bump allocator: alloc moves a pointer up, dealloc does nothing, and memory comes back only when the program exits. So what runs a guest out of memory is the total it allocates over the run, not its peak. When an allocation would end above the stack's reserve, or above the live stack pointer, the guest exits with status 71.

Do:

  • Allocate once and reuse: hoist buffers out of loops and clear() them instead of building new ones.
  • Size collections up front with Vec::with_capacity, String::with_capacity, so they do not reallocate and copy as they grow. A Vec grown one push at a time to n elements also leaves its earlier, smaller buffers behind.
  • Prefer borrowing (&[u8], &str) to cloning, and iterators to intermediate collections.
  • Process large advice in place: advice() is already a slice over memory, so there is nothing to copy.

Don't: collect() into a new Vec inside a hot loop, clone values you only read, or rebuild a map per request when one map can be cleared and refilled.

rust
// Total heap grows with the number of requests:
for req in requests {
    let parts: Vec<u32> = req.chunks(4).map(|c| u32::from_le_bytes(c.try_into().unwrap())).collect();
    process(&parts);
}

// Total heap is one buffer:
let mut parts: Vec<u32> = Vec::with_capacity(MAX_PARTS);
for req in requests {
    parts.clear();
    parts.extend(req.chunks(4).map(|c| u32::from_le_bytes(c.try_into().unwrap())));
    process(&parts);
}

The stack has 8 MiB, and nothing guards its far edge#

The stack grows down from 0x8000_0000 and has an 8 MiB reserve that no heap block may enter. Deep recursion within it is fine. What nothing detects is a stack that grows past its reserve after the heap has filled the space below it: heap blocks then change under a deep call chain, silently. Do: keep recursion depth bounded and predictable, or write deep traversals iteratively with an explicit work list. Don't: recurse to a depth set by untrusted input.

Aligned accesses only#

A halfword or word access through a misaligned pointer is fatal, never split, and the run has no proof. Safe Rust never produces one. Don't cast a byte pointer to *const u32 and dereference it; read with u32::from_le_bytes, which compiles to byte loads, or with ptr::read_unaligned.

Null is a hole#

Addresses below 0x8000 belong to nothing, so a null or small wild pointer is a fatal OutOfBounds rather than a read of garbage. It surfaces as a run with no proof, never as a wrong answer.

Concurrency#

Atomics: supported, and not for new guest code#

Important

The A extension is fully supported: lr.w, sc.w and all nine AMOs decode, execute and prove, through their own circuit family, and core::sync::atomic compiles to them. Writing a guest with atomics is still strongly discouraged. Apogee executes on a single hart, with no interrupts and no threads, so there is nothing to synchronize. Atomics are there so that existing code which uses them, a library with an atomic counter or a spin lock, compiles and proves unchanged. They are a compatibility path, not a practice.

What to know if atomics reach your guest through a dependency:

  • On one hart an atomic is just a read-modify-write. fetch_add is an amoadd.w that adds; nothing can interleave with it.
  • sc.w always succeeds. The machine keeps no reservation state, so a store-conditional stores and writes 0 to rd. The lr.w/sc.w retry loop that compiled code uses for compare_exchange is unaffected, because first-time success is legal on any hart. Code that relies on sc.w failing without a valid reservation does not get that failure here. This is the one place Apogee deviates from RV32IMAC.
  • fence does nothing, and memory orderings (aq, rl, SeqCst) order nothing on one hart.
  • They cost a circuit family. An atomic adds the ATOMICS family to the program, which then proves at least one shard of it.

Do: use plain variables, Cell and RefCell for state in new guest code. Don't: add AtomicU32, Mutex-like spin locks or Arc to a guest that has no second thread to share them with.

Inputs and outputs#

Check advice before anything derived from it reaches the journal#

Advice is memory the prover filled, and nothing binds it. Check it against something the proof does bind, such as a hash or a Merkle root in the public input, a signature, or a property of the result, before committing anything that depends on it. Committing a function of unchecked advice publishes a value the prover chose. See the pattern.

Keep public values small, and fixed-size for on-chain use#

The input and the journal hold at most 16,380 bytes each. commit exits 70 rather than overflow. Large inputs belong in advice behind a commitment, and growing outputs behind a digest. A verifier contract is built for one input length and one journal length, so a guest settled on Ethereum should publish a fixed-size journal.

Decide what failure looks like#

A guest that exits nonzero, or panics, still has a valid proof of what it did, and a verifier reads the status before the journal. Give each refusal its own exit code outside the SDK's (70, 71, 72 and 101), and commit nothing derived from unchecked data before the checks that can refuse it.

No world outside#

A guest has no clock, no randomness, no network, no files and no environment. A library call that asks the host for any of them answers -ENOSYS and leaves the run unprovable. Seed HashMap deterministically or use BTreeMap; derive randomness from the input when an algorithm needs it; pass time in as input.

Cost#

Every executed instruction is a proved row#

Proving cost follows the cycle count, family by family. Build --release, measure with the profiler, and treat cycles the way embedded programmers treat bytes.

Delegate what has a circuit#

keccak256, sha256, elliptic-curve addition and multiplication, Poseidon2, BN254 field arithmetic and 256-bit modular multiplication have dedicated circuits. Reach them through guest_sdk and the vendored k256, ark-ff and revm-precompile, not a software implementation compiled into the guest. See Delegations.

Verify instead of compute#

When a result is expensive to find and cheap to check, let the prover find it and pass it as advice, and have the guest check it: a sorted order, a factorization, an inverse, a path through a tree, a search result.

Avoid floating point#

The target has no F or D extension, so f32 and f64 compile to integer software routines. They are correct and deterministic, and each operation costs many instructions. Use integers or fixed point.

Heights cost whole shards#

A family with rows costs at least one shard of its height, however few rows it fills. A program that touches a family once pays for a whole shard; the families your code uses, and the heights you choose, set the floor of every proof. See Heights.

Code and identity#

The instruction stream is the image#

Code is static: each pc's instruction comes from the decoded tables built at load, never from RAM. A store into .text changes data, not behaviour, and a jump to a halfword with no instruction ends the run with no proof. There is no JIT and no self-modifying code.

One illegal word anywhere refuses the program#

The decoder takes all of .text, reachable or not. A CSR access, fence.i, a floating-point or an RV64 encoding in inline assembly, or data assembled into .text, makes derivation refuse the whole program. ebreak decodes but has no proof.

Overflow checks are part of the program#

The guest profiles keep overflow-checks on in release, because turning them off changes what a guest computes: u32::MAX + 1 would wrap and exit 0 instead of panicking. Use wrapping_*, checked_* and saturating_* where you mean them.

A build is an identity#

The identity binds every byte of code and data, the entry point and every height. Rebuilding on another machine produces another identity, because the ELF embeds absolute paths. Register and ship the ELF you proved, not the command that made it.

Checklist#

Before you prove:

  • cargo build --release for riscv32imac-unknown-none-elf, and the profiler's cycle report reviewed
  • no usize in anything committed, hashed or serialized
  • no allocation inside hot loops; growing collections sized with with_capacity
  • no atomics, locks or Arc in guest code of your own
  • every use of advice checked against something the proof binds, before any commit that depends on it
  • journal bounded, and fixed-size if it settles on-chain
  • hashing and curve arithmetic routed through delegations
  • distinct exit codes for each refusal
  • the host build and the emulator agree on the journal for your test inputs
  • the program identity recorded from the ELF you will ship

Launch Your App

Troubleshooting

Every way a guest stops short of a verified proof, by symptom. Exit statuses, fatal executor errors, refused ELFs, refused programs, failed proofs and verifier errors, with the cause and the fix.

View as Markdown

A guest can stop short of a verified proof at six points. Find the symptom, then the row.

The run exits with a status you did not expect#

The run finished and is provable; the guest chose to fail. The SDK's own statuses:

Status Cause What to do
70 commit would pass 16,380 bytes Commit a digest of the output instead of the output
71 an allocation would end above __stack_top − 8 MiB or above the live sp Nothing is freed, so the run's total allocation must fit between the image and 0x7F80_0000. Reuse buffers across loops and size them with with_capacity (the heap)
72 a delegation answered what its shim refuses: an error, or -ENOSYS after a multi-call operation's first call Use the SDK's functions rather than raw frames, and keep operands below their modulus
101 a panic, which prints nothing Run the same inputs through the host build of your library, where the panic message prints (test on the host)

Any other status is your own exit(code).

The run stops with a fatal error#

The emulator returns an EmuError and there is no exit status and no proof:

Error Usual cause
OutOfBounds a null or wild pointer ([0, 0x8000) is a hole), an advice read past what the host supplied, advice() on a run with no advice, or a delegation frame not wholly in RAM
Misaligned a halfword or word access through an unaligned pointer, or a misaligned delegation frame
NotAnInstruction a jump to a pc holding no instruction, including the all-zero c.unimp halfword
IllegalInstruction an encoding the machine does not execute
Ebreak an ebreak, which has no proof
ClockOverflow more than 2^36 − 1 cycles
PublicInputTooLong, JournalTooLong an input, or the journal's length word at exit, above 16,380 bytes
DelegationFrame a frame its circuit has no witness for: a MOD_MUL or EC_ADD operand at or above its modulus, a selector naming nothing, a keccak round above 23, a SHA-256 group above 15, a Poseidon2 lane at or above p
DelegationFamilyAbsent a delegation number the image never declared, on a tracing path

The ELF is refused#

artifact-dump, host::setup and the tools refuse an ELF the loader cannot take, with a LoaderError:

Refusal Usual cause
NotAnElf, Truncated not the guest's ELF: a .d file, a partial write
NotRiscV, UnsupportedElfType, RelocatableElf, DynamicElf a host build, an object file, a PIE or a dynamically linked build
BadSegment, NoExecutableSegment, EntryNotAnInstruction an edited link.ld, or no _start linked
RvcIllegal, InstructionTooLong, TextTruncated data in .text, such as a table in hand-written assembly. Never the zero halfword, which is expected

The program does not register#

The ELF loads but the program cannot be decoded into a configuration:

Refusal Cause and fix
Not all opcodes supported: pc=… a word outside RV32IMA anywhere in executable code, reachable or not, such as a CSR access or fence.i in assembly. Remove it
TableTooShort code past a family's table reach. Raise that family's height: 2^22 reaches 7.9375 MiB of code
ProgramTooLarge the image is past bytecode_size_words, 4 MiB by default. Raise the ceiling in ProgramParams
ImageOutsideWindow a file-backed byte lies past RAM window 0 at your window height. Raise the window height
HeightNotOnMenu a height that is not 2^8, 2^12, 2^16, 2^18, 2^20 or 2^22
UnknownDelegation the image declares a delegation number no family answers

A key also fails to build if a family the program uses is set below its floor, 2^20 for an instruction family and 2^16 for the RAM window families, since no circuit exists there.

The prover fails#

An honest prover proving what the emulator executed does not fail, so a failure points at an input it should not have accepted or at a bug. Two cases account for most:

  • A call with no proof. A system call outside EXIT and the delegations, which a library made for host data, answers -ENOSYS and the run continues, but the prover's fill refuses that row and names the cycle. Find the dependency that asks for randomness or time.
  • Out of memory. The process is killed while shards are in flight. Lower the third argument of host::prove, or lower heights.

For anything else, rebuild with the prover's debug log and rerun:

sh
cargo run --release -p bench --features prover/debug-info -- prove ...
APOGEE_DEBUG=detail <the run> 2>&1 | tee run.log
grep -c 'begin h=' run.log; grep -c 'gkr done' run.log     # unequal: a shard died
grep 'begin h=' run.log | tail -1                           # which one
grep -E 'FAIL|NOT CANONICAL|UNBALANCED|OVER the|NAMES NO|DISAGREES|NOT LOOPING|ABORTED' run.log

self_check FAILED names the row's first failing gate and every operand's value. The debug log lists every marker.

Verification fails#

verify_shard and verify_block return a VerifyError whose class says what broke:

Class Meaning
Statement the statement does not fit the key: shard counts, window rules, payload lengths, lists, or the global digest the proof was seeded with
Malformed the proof's shape is not its circuit's: counts of commitments, outputs, rounds or claims
Constraint { layer } a gate is violated, or a layer's sumcheck fails
Lookup { channel } a looked-up tuple is in no row of its table
MemoryArgument the read and write multisets do not reconcile, or a public window does not hold the statement's bytes
Opening a commitment opening fails

If you built the proof with Apogee's own prover from an execution the emulator accepted, a verification failure means the verifier and the prover disagree about the program: check that you load the key the proof was made under, at the same heights, over the same ceremony.

Identity mismatch#

The identity you computed differs from the one you expected:

  • Another machine's build. The ELF embeds absolute paths, so a rebuild elsewhere is another image. Compare against the ELF that was proved, not a fresh build.
  • Other heights. Every height is bound into the identity. artifact-dump tables reports the identity at the default heights; your setup may use others.
  • Another ceremony. A Hermez powers-of-tau file is a different τ, so every commitment differs. Check the file's [τ]_1 against the ceremony's.

Launch Your App

Example Guests

The guests in the repository, each a worked example of one part of the guest SDK or the machine. Where to look for the pattern you need.

View as Markdown

The guests/ workspace holds every guest the repository builds and tests. Each one exists to exercise something, which makes them the best reference for the pattern you are about to write. All of them build with cargo build --target riscv32imac-unknown-none-elf from their own directory.

Start here#

Guest Shows
public-io the three regions at once: advice checked against the public input before anything is committed. The I/O model in one small program
fib the smallest SDK guest: a u32 in, a u32 out, wrapping arithmetic
echo, heap the allocator: advice copied through heap buffers, Vec and Box churned through the bump allocator

Application patterns#

Guest Shows
amm, orderbook 128- and 256-bit integers with no heap; BTreeMap, sorting, and a sorted order supplied as advice and checked rather than computed
vault, recursion-ops crates/field and crates/transcript in a guest, delegating to FR_ARITH and POSEIDON2 with no shim named
revm-block Ethereum blocks on revm: binaries revm-block (a recorded mini-block) and revm-block-stateless (the stateless validator)

Delegations#

Guest Shows
keccak-test, sha256-ops, mod-mul-ops, ec-ops KECCAK_F, SHA256_COMP, MOD_MUL and EC_ADD, each checked inside the guest against independent values
keccak-unused, recursion-unused shims linked and never called: the families are declared and prove zero shards

The machine itself#

Guest Shows
atomics every A-extension instruction as core::sync::atomic emits it. One hart means each is a plain read-modify-write; it exists to test the family, not to recommend the practice
opcodes every RV32IMAC instruction
rvc-dense compressed-instruction expansion: one sequence assembled compressed and not
addsub, control, alu, mem, shards hand-written assembly with its own _start and no SDK, exiting with its result. shards fills two 2^20 shards
recursion the recursion tree's verifier programs, binaries leaf and node

The smallest useful guest#

fib reads one u32, takes that many Fibonacci steps with wrapping arithmetic, and commits the result:

guests/fib/src/main.rsrust
#![no_std]
#![no_main]

guest_sdk::entry!(main);

fn main() {
    let mut n = [0u8; 4];
    assert_eq!(
        guest_sdk::read_input(&mut n),
        4,
        "fib: the public input is one u32"
    );
    let n = u32::from_le_bytes(n);

    let mut a: u32 = 0;
    let mut b: u32 = 1;
    for _ in 0..n {
        let next = a.wrapping_add(b);
        a = b;
        b = next;
    }
    guest_sdk::commit(&a.to_le_bytes());
}

A short input is a fault, not a default: a guest that proceeds on a partly filled buffer proves a statement about zeroes. The addition wraps on purpose, so an n above 47, past the last term that fits in 32 bits, is an ordinary input with an ordinary answer rather than a failed run. And there is no advice, because f_n costs a verifier as much to check as to compute: an advised answer would have to be recomputed to be believed.

Launch Your App

Guest SDK Reference

Every public item of the guest-sdk crate, with its exact signature and behaviour. Only exit and the delegation shims issue an ecall; everything else is loads and stores.

View as Markdown

crates/guest-sdk is a guest's entire runtime: the startup code, the entry macro, the allocator, the panic handler and the ecall shims. It compiles only for riscv32imac-unknown-none-elf.

Entry#

rust
guest_sdk::entry!(main);

Exports the main symbol the startup code calls, as a wrapper calling your function, which takes no arguments and returns (). Your function keeps its own name and may itself be called main. Returning from it is exit(0).

The regions#

Item Signature Behaviour
public_input fn public_input() -> &'static [u8] The public input payload, its length word clamped to the window. No copy, no ecall
read_input fn read_input(buf: &mut [u8]) -> usize Copies min(buf.len(), public_input().len()) bytes and returns the count. It may return short
advice fn advice() -> &'static [u8] The advice payload, its length clamped to the region. Bound by nothing, so the guest checks it. Fatal OutOfBounds on a run given no advice
commit fn commit(bytes: &[u8]) Appends to the journal and updates its length word. Exits 70 rather than overflow the 16,380-byte window
journal fn journal() -> &'static [u8] Everything committed so far
exit fn exit(code: i32) -> ! Ends the run with code as the statement's exit status. Publishes nothing beyond what was committed

Hashing#

Item Signature Behaviour
keccak256 fn keccak256(input: &[u8]) -> [u8; 32] Ethereum's Keccak-256, not SHA3-256. The sponge and padding run in guest code; each keccak-f[1600] round is one KECCAK_F call. Software fallback if the first call answers -ENOSYS
sha256 fn sha256(input: &[u8]) -> [u8; 32] FIPS 180-4 SHA-256. Padding and the block loop run in guest code; each compression is sixteen SHA256_COMP calls. Software fallback as above
poseidon2_permute fn poseidon2_permute(state: &mut [u8; 96]) -> bool The width-3 Poseidon2 permutation over three canonical little-endian Fr lanes, in place, through POSEIDON2. Returns false on -ENOSYS, for the caller's own software path

Elliptic curves#

rust
pub type ProjectivePoint = [[u32; 8]; 3];

A point in homogeneous projective coordinates, x = X/Z and y = Y/Z, each coordinate eight little-endian 32-bit limbs below the curve's field modulus. It is not Jacobian: arkworks' Projective is, so a caller converting from it maps (X·Z, Y·Z², Z) in and (X·Z, Y, Z³) out. The identity is (0 : 1 : 0).

Item Signature Behaviour
ec_add fn ec_add(codes: &[u32; 3], p: &ProjectivePoint, q: &ProjectivePoint) -> Option<ProjectivePoint> p + q by the complete formula, through three EC_ADD calls in group order. None on -ENOSYS
ec_mul fn ec_mul(codes: &[u32; 3], p: &ProjectivePoint, k: &[u32; 8]) -> Option<ProjectivePoint> k·p by double-and-add from the top bit. k is used as given; reducing it modulo the group order is the caller's business
ec_identity fn ec_identity() -> ProjectivePoint (0 : 1 : 0)
recursion::SECP256K1_GROUPS, recursion::BN254_GROUPS [u32; 3] The codes argument: which curve, as the three group selectors of one addition

The formula proves arithmetic, not curve membership: check points taken from advice yourself.

Raw delegation shims#

guest_sdk::recursion holds the shims over word-aligned frame types. Each frame type is #[repr(C, align(4))], so its alignment is the type's and not wherever the code generator put a local. A base-format shim returns false on exactly -ENOSYS; any other nonzero answer exits 72.

Item Purpose
mod_mul(&mut ModMulFrame) -> bool One a·b mod m. Build the frame with ModMulFrame::of(modulus, &a, &b) and read frame.result(). The modulus codes are SECP256K1_P, SECP256K1_N, BN254_P and BN254_R, and both operands must already be below the modulus
sha256_comp(&mut Sha256Frame) -> bool One whole compression: sixteen calls in order. Sha256Frame::of(&state, &block), then frame.working(); adding the result to the chaining value is the caller's
ec_add_complete(&mut EcAddFrame, &[u32; 3]) -> bool One complete addition: three calls in group order. EcAddFrame::of(&codes, &p, &q), then frame.result()
poseidon2(&mut Poseidon2Frame) -> bool, fr_arith(&mut FrArithFrame) -> bool The permutation and one Fr operation over byte frames; field and transcript call these for you
sha256_rounds, ec_add Single steps of the operations above. A step in the wrong order is not refused, it computes something else, so prefer the whole-operation functions
fr_op, p2_field, field_io, fq_op, import, import_run, replay The recursion format's coprocessor calls, used by the recursion tree's own programs. They have no software path

Each shim reads its ecall number from its family's declaration record, a 12-byte static in its own linker section. Linking a shim declares the family; a declared family that is never called proves zero shards.

Runtime behaviour#

Piece Behaviour
Startup _start at 0x0001_0000 points sp at __stack_top (0x8000_0000), zeroes .bss byte by byte, calls main, and exits 0 if it returns
Allocator Bumps up from __heap_start, never frees. Exits 71 when a block would end above __stack_top − 8 MiB or above the live sp
Panic handler Exits 101 and writes nothing. A panicking guest is provable and has published what it committed
Exit statuses 70 journal overflow, 71 out of heap, 72 a delegation answered an error, 101 panic

Transparent delegation#

Two library crates of the repository delegate on the guest target without naming the SDK, through a target-only dependency on it:

  • field::Fr: addition, Montgomery multiplication (*, square, pow and the conversions) and nonzero inverse call FR_ARITH. A guest using Fr arithmetic declares that family.
  • transcript::poseidon2_permute calls POSEIDON2.

The vendored k256, ark-ff and revm-precompile do the same for secp256k1, BN254 and the EVM precompiles: Delegations.

The specification of the ABI underneath all of this is Guest ABI.

Launch Your App

AI Companion

One file that briefs an AI model on writing Apogee guest programs. Download it, put it in front of your model, and it starts from the same rules this manual teaches.

View as Markdown

Much of the code written against Apogee will be drafted by a model. A model that has never seen Apogee will write a plausible guest that uses std, allocates in every loop, reaches for an atomic counter, trusts its advice and commits a usize. The AI Companion is a single Markdown file that front-loads everything that prevents those mistakes: what a guest is, the hard rules, the SDK's complete public surface with exact signatures, patterns to copy, the errors and their fixes, and a review checklist.

How to use it#

  • In a chat: attach the file, or paste it as the first message, before you describe what you want built.
  • In a coding agent: save it at the root of your project under the name your tool reads by convention, such as AGENTS.md or CLAUDE.md, or add it to the tool's project rules. The agent then reads it at the start of every session.
  • For review: ask the model to check a guest against section 8 of the file, the review checklist, line by line.

The file states its rules as MUST and MUST NOT, with the reason beside each, because models follow constraints that are explicit and explained more reliably than conventions they are expected to infer.

What it contains#

Section Contents
0. Instructions to the model Treat the rules as hard constraints; never call an API that is not listed; proofs are not zero-knowledge
1. What a guest is The target, the single hart, what a proof states, the program identity, the three memory regions
2. Hard rules 23 rules: program shape, 32-bit usize and pointers, the bump allocator, the stack, alignment, atomics, the absent outside world, advice, public-value limits, the instruction set, floats, overflow checks, cost
3. Layout and build The crate templates, the guest workspace, the build command, the host-first library split
4. The guest SDK Every public function with its exact signature, the runtime facts and memory map, the delegated operations and the vendored crates
5. Patterns Advice checked against a hash; a Merkle query and state transition; buffer reuse; structured advice; digests for growing outputs
6. Host side Running in the emulator, profiling, exporting the image, proving and verifying, heights and shards in flight
7. Errors and fixes Every exit status, fatal error and refusal a guest meets, with its cause and fix
8. Review checklist Eleven checks to run before proposing guest code
9. Facts The ISA, the proof system, the security level, the limits and the measured results

The rules it insists on#

The companion repeats this manual's rules, and three of them deserve to be called out because models get them wrong most often:

  • Pointers and usize are 32 bits. A model trained mostly on 64-bit code serializes usize without a second thought. The guest and the host then disagree on the bytes.
  • The heap never frees. Idiomatic Rust allocates freely because a real allocator gives memory back. Here every allocation is permanent for the rest of the run, so the companion asks for reuse and with_capacity everywhere.
  • Atomics compile, and do not belong in new guest code. They are supported for compatibility with existing libraries. On a single hart they synchronize nothing, and they add a circuit family to the proof.

For agents reading these docs directly#

  • /llms.txt indexes every page of this site in Markdown, for models that read the web.
  • /docs/llms-full.txt is the whole English documentation, specification included, in one file.
  • Every page has a View as Markdown link and a Copy as Markdown button, under its title and in the right-hand column.

English is the canonical language of the documentation, and the companion is published in English for every locale: it is the language models follow most reliably, and the one the specification is written in.

Architecture

Architecture

Apogee VM end to end. What a proof states, the path from a guest binary to a contract call, how the large components fit together, and the design decisions that shape them.

View as Markdown

Apogee proves executions of RV32IMAC programs. This section describes the system at the level of its large components: what each one does, why it is built the way it is, and how it hands off to the next. The Auditors section holds the same system at the level of every column and gate.

What a proof states#

A verifier holds three things it does not take from the prover's word:

  • the program identity, one field element digesting the program's instruction tables, its initial memory image, its entry pc and its configuration;
  • the SRS digest of the ceremony a verifying key must carry;
  • a verifying key, which may come from anyone, because loading it recomputes the identity and the SRS digest from its own contents and holds its circuits to the verifier's registry.

The proof's statement carries the public input, the public output (the journal), the exit status, and the record of the execution's shape: shard counts, memory windows, the final registers and pc, and every shard's memory commitments and roots. A proof that verifies establishes that the program of that identity, started at its entry pc over its image, with the public input in its input window and some advice of the prover's choosing, executes instruction by instruction to EXIT with that status, having written that journal. Nothing is claimed of the advice, and nothing is hidden: no commitment or proof is blinded.

From a binary to a contract call#

Apogee VM, end to end Four stages. Program: guest ELF to ProgramImage to decoded tables and VmConfig to program identity. Execution: the emulator produces rows per family, cut into shards. Per shard: commit columns with Mercury, run the GKR backward pass, open every column in one batch. Settlement: block proof, recursion tree, Groth16 decider, ApogeeVerifier contract. 1 · PROGRAM Guest ELFRV32IMAC, no_std Rust ProgramImagesegments · RVC expanded Decoded tables7 families · VmConfig Program identityone Fr 2 · EXECUTION Emulatorone hart · pure function of (image, io) Rows, one table per familycycles · invocations · memory words Shards2^8 … 2^22 rows · 23 families 3 · EACH SHARD Commit the columnsMercury over KZG GKR backward passone sumcheck per layer One batched openingevery column at one point · 704 B 4 · SETTLEMENT Block proofstatement + shard proofs Recursion treeleaves · nodes · root Groth16 deciderbound wires · BN254 ApogeeVerifier.soltwo pairings · true
The four stages. The program is fixed before anything runs; the execution is cut into shards; each shard is proved on its own except for the memory argument, which closes once over all of them; settlement compresses the block for a contract.
  1. The program. The loader reads the ELF into a ProgramImage, expanding compressed instructions in place. The decoder routes each instruction to one of seven instruction families and builds every family's decoded table, a row per halfword of code, then commits to all of it as the program identity. Programs and identity.
  2. Execution. The emulator runs the guest on one hart. A cycle is one row of the family that owns its instruction, recording timestamped reads and writes of the pc, the registers and RAM. Hashing and big-integer arithmetic are delegated: an ecall names a frame in RAM, and a row of a delegation family does the work on it. Execution, families and shards.
  3. Shards. A family's rows are cut into shards of the family's height, a power of two between 2^8 and 2^22. The memory an execution touches is covered by shards of the window families, which give each word its initial and final values. A shard is the unit of proving; a block is hundreds.
  4. A shard's proof. Its columns are committed with Mercury. The family's circuit is run backward by the GKR engine from its outputs to those columns, a sumcheck per layer, and every column is opened at the one point that pass ends on, in one batched opening.
  5. The block. A BlockProof is the statement and its shard proofs. Verification runs the global transcript once, each shard's checks, and once the memory reconciliation over every shard's roots.
  6. Recursion and settlement. Verifier programs, proved by Apogee itself, verify runs of shards and fold their deferred pairings. A tree of them ends in a root, a Groth16 circuit re-verifies the root, and ApogeeVerifier.sol checks that proof and the folded pairing. Recursion and settlement.

The prover executes the guest twice: once to commit every shard's memory columns, which fixes the statement and its challenges, and once to prove each shard as it fills. Its memory is bounded by the shards in flight, not by the length of the execution. The streaming prover.

The decisions that shape it#

One field, one curve. Everything is over BN254's scalar field: the circuits, the transcript, the commitments and the recursion. That is what lets a node of the recursion tree verify base shards in its own arithmetic, and what lets the tree end in a Groth16 proof Ethereum checks with its pairing precompile.

Layered GKR circuits instead of committed constraint tables. A family's circuit is a stack of degree-2 gate layers above its committed columns. Only the bottom layer is committed; every layer above it is proved by sumcheck in a single backward pass and never committed. The pass ends with every committed column claimed at one point, so a shard needs exactly one opening. The GKR engine explains why this is the engine's central economy.

A constant-size opening. Mercury opens a multilinear commitment with eight curve points and six field elements, 704 bytes, whatever the polynomial's size and however many columns share the point. Its checks have the shape e(A, [1]_2) = e(B, [x]_2), which recursion can fold instead of pairing.

One memory argument for the whole execution. Every access, in every shard of every family, is a tuple in one read/write multiset, and the verifier reconciles the products once per statement. The pc is a cell of that multiset, so ordering, continuity and cycle uniqueness across shards need no other argument and no shard needs to chain to its neighbour.

Delegations as families, not instructions. An expensive function gets its own circuit family, invoked by an ecall over a frame of RAM and paired one to one with its request through the same multiset. The instruction circuits stay small, and a guest pays for a delegation only if it calls it.

Streaming instead of materializing. The trace at about 300 bytes a cycle would be the largest object in the system, so it never exists. The prover executes twice and holds only the shards being worked.

No borrowed cryptography. Fields, curve, pairing, MSM, hash, polynomial commitment, GKR and Groth16 are implemented in the repository and specified page by page. arkworks, Plonky3 and zkhash appear only as test oracles.

How soundness composes#

Each shard's GKR pass and opening tie its circuit's outputs to committed columns. On top of that, these arguments span the execution:

Claim Carried by
Every row obeys its instruction the family circuit's enforcing gates, zero on every row
A row's instruction is the program's at its pc a lookup of the row's pc and fields in the family's decoded table, which the identity commits
Every read returns the last write one multiset over all shards; the verifier multiplies every shard's read and write roots against boundary factors for the registers and the pc
The rows are one path from the entry pc to the exit, in program order the pc is a cell of that multiset, written at least four timestamps after it is read
A value is a byte, a word, a sign, an XOR LogUp channels over range, byte and generic tables
The public input and the journal are the claimed bytes the two public windows' initial and final columns, held to the bytes' multilinear extensions
A delegated computation is the function's invocation rows that read and write the frame through the same multiset, paired one to one with their ecall

Challenges come from a Poseidon2 duplex transcript. The global transcript absorbs the whole statement, every shard's memory commitments included, before the memory challenges exist; each shard's transcript is seeded from its final state. The soundness map carries each row down to the sections that prove it.

The code, by layer#

Layer Crates Role
Arithmetic field, curve, poly, sumcheck Fr; the Fq tower, G1, G2, the pairing, MSM; multilinear polynomials; the zerocheck
Fiat–Shamir and setup transcript, srs Poseidon2 and the duplex transcript; ceremony ingestion, KZG, Groth16's phase 1
Commitments pcs, pcs-verify Mercury and its deferred verification
The program loader, isa, program ELF to image, the decoder, decoded tables, VmConfig and identity
Execution emulator, trace the executor and its tracers; rows, memory state, column builders
Circuits constraints, gkr-verify, gkr every circuit as data; the GKR verifier and prover
Proof and verification verifier-core, verifier, prover statement, transcripts, keys, every check; the streaming prover
Settlement host, groth16, contracts/ the host SDK, the recursion tree and decider; ApogeeVerifier.sol
Assurance checker, tools/ independent validators, the tamper suite, benchmarks, profilers, oracles

The verifier is the only trusted party: the prover validates nothing, and a wrong input costs an honest prover a proof that fails. The security model lists exactly which crates soundness rests on.

Architecture

Programs and Identity

How a guest ELF becomes a static, verifier-known description of a program, and why one field element of identity is enough to tell a verifier which program a proof is about.

View as Markdown

Before anything executes, Apogee turns the guest binary into a fixed description of the program: its image, its instructions sorted into circuit families, the configuration it will be proved under, and one field element that commits to all of it. Every step is a pure function of its input.

Loading#

The loader accepts a static 32-bit little-endian RISC-V executable whose loadable segments lie inside guest RAM at even addresses, pairwise disjoint, at least one executable. It reads a program header's type, offsets, sizes and executable bit, and nothing else: the VM has no pages and no permissions, and all of RAM is addressable whatever the segments declare.

It then sweeps each executable segment halfword by halfword. A halfword ending in binary 11 begins a 4-byte instruction; the zero halfword is a non-instruction (LLVM pads unreachable blocks with it); anything else is a compressed instruction, expanded to its 32-bit form in place. Addresses are never compacted: a c.addi at 0x1002 stays there and occupies two bytes, so every linker-resolved address holds, and an instruction's length is the only record of whether the next pc is pc + 2 or pc + 4.

A desynchronized sweep cannot make a wrong instruction provable. A slot is a function of the bytes at its own pc, so every instruction slot is what a hart fetching there would decode. Data that shifts the sweep off the true boundaries can only lose true instruction starts, whose pcs then have no table row, or meet an unclaimed encoding and refuse the whole image.

Decoding and routing#

The decoder takes 32-bit words only and accepts exactly the 59 instructions of RV32IMA: 40 of the base set, 8 of M, 11 of A. Everything else, from RV64 encodings and floating point to CSRs and fence.i, is a decode error, and one undecodable word anywhere in executable code refuses the program, reachable or not. Each instruction is routed to exactly one of seven instruction families:

Id Family Instructions
0 ADD_SUB_LUI_AUIPC ecall, ebreak, fence, addi, auipc, add, sub, lui
1 JUMP_BRANCH_SLT slti, sltiu, slt, sltu, the six branches, jalr, jal
2 SHIFT_BITWISE the six shifts, and, or, xor and their immediate forms
3 MUL_DIV the eight instructions of the M extension
4 MEM_WORD lw, sw
5 MEM_SUBWORD lb, lh, lbu, lhu, sb, sh
6 ATOMICS lr.w, sc.w and the nine AMOs

The grouping follows what the circuits share. One comparison gadget settles signed and unsigned order for every branch and slt kind; one product identity serves all four multiplies and the division; a shift either way is one product with a looked-up power of two.

Decoded tables#

Each instruction family gets a decoded table: its setup columns, where row i is pc 2i, one row per halfword of the address space the table reaches. A live row holds one of the family's instructions as a tuple pc, next_pc, rs1, rs2, rd, imm, extra_mask, where the mask is one bit naming the mnemonic. Every other row is padding, −1 in every field, so no live row is ever the padding row and an all-zero row is never a claimable instruction at pc 0.

Every cycle's row looks itself up in its family's table by its pc. That lookup is what binds an execution to the program: a row's instruction is the program's instruction at that pc, and a pc with no live row in any table cannot be executed provably. Code is static as a result. A store into .text changes what a later load reads, never what executes.

The configuration#

A program's static shape is its VmConfig: the families it uses, each with a height, and a ceiling on its code size. Nothing chooses the family set; it is derived:

  • an instruction family is present when the image holds one of its instructions;
  • the five window families, which initialize and tear down memory, are always present;
  • a delegation family is present when the image declares it, through a 12-byte record that linking its shim leaves among the image's bytes.

A family's height is the number of rows in one of its shards, chosen from 2^8, 2^12, 2^16, 2^18, 2^20, 2^22. Every height is an even power of two because a Mercury opening needs one. Heights are a parameter of the program, not of a run, and each choice of heights is its own program.

Program identity#

The program identity is one element of Fr: a Poseidon2 digest of the code version, the VmConfig, the entry pc, and every family's setup commitments, which are Mercury commitments to the decoded tables and to the image's initial memory words.

It binds every instruction the sweep found, with its pc, length, operands and kind, and that no other pc holds one; every file-backed byte of the image, which includes the delegation declarations; the entry pc; the family set, every height, the code-size ceiling and the code version. It does not bind the ceremony or the generic lookup table, which the SRS digest covers; the circuits, which a key's load holds to the verifier's registry; anything an execution chooses; or anything of the ELF the loader does not read, such as the symbol table.

Computing it needs the ceremony's powers, to commit. Checking it needs only the commitments, which a verifying key carries: loading a key recomputes the identity from them, and every shard's opening checks its setup columns against the same points. That is what ties the tables a proof reads to the identity a verifier registered.

Important

A verifier takes the identity from a channel the prover does not control. Against an identity the prover supplied, a proof shows only that some program ran. Whoever holds the ELF, the parameters and the ceremony file can recompute it.

The specification: Program and identity.

Architecture

Execution, Families and Shards

One hart, a 38-bit clock, every memory access as a timestamped query, 23 circuit families, and the shard as the unit of proving.

View as Markdown

The machine#

The emulator runs RV32IMAC on one hart over a ProgramImage, with no interrupts and no privilege levels. A run is a pure function of the image and its input, with no clock, randomness or threads, so two runs cut identical shards. It differs from a hosted RV32IMAC at three points: sc.w always succeeds, a misaligned halfword or word access is fatal rather than split, and the instruction stream is the image decoded at load.

Every other way a run can stop short of EXIT, such as an access outside the mapped regions, ebreak, or a jump to a halfword with no instruction, is a fatal error with no trace. Such a run has no proof. A nonzero exit status is not an error: it is an execution like any other, and provable.

The clock and the query#

Cycle c occupies the four timestamps 4c + Δ, one per slot Δ ∈ {0, 1, 2, 3}. Every instruction is one cycle and nothing else is: a delegation invocation rides the cycle that requested it. Cycles are numbered from 1, because timestamp 0 is every address's initial write. The clock is 38 bits, so an execution runs at most 2^36 − 1 cycles.

Every memory access is a query: a read of a value last written at some earlier timestamp, and a write at this one. A query that only reads writes back what it read. Slot 0 of every cycle is the pc query, which reads pc and writes next_pc. Then come the instruction's register and memory queries at fixed slots:

Class Δ = 1 Δ = 2 Δ = 3
register-immediate, jalr rs1 rd
branches rs1 rs2
register-register, M rs1 rs2 rd
loads rs1 the word, read rd
stores rs1 rs2 the word, with the stored bytes merged in
atomics rs1 rs2 the word, and rd
ecall a7 a0 a0, and a delegation's mirror query

Addresses live in spaces: the 32 registers, RAM by 4-aligned word, the pc, one anchor space per delegation type, and the recursion format's field cells. x0 is an ordinary register in the trace and a constant in the machine: every query at it reads and writes 0.

Twenty-three families#

A family is a circuit and the rows it proves. There are four kinds:

Kind Families A row is
Execution 0–6: ADD_SUB_LUI_AUIPC, JUMP_BRANCH_SLT, SHIFT_BITWISE, MUL_DIV, MEM_WORD, MEM_SUBWORD, ATOMICS one executed instruction
Window 7 INIT_TEARDOWN, 8 ZERO_WINDOWS, 12 PUBLIC_INPUT, 13 PUBLIC_OUTPUT, 14 ADVICE_WINDOWS one memory word, initialized and torn down
Delegation 9 KECCAK_F, 10 POSEIDON2, 11 FR_ARITH, 15 MOD_MUL, 16 SHA256_COMP, 17 EC_ADD one invocation over a frame of RAM
Recursion 18 FIELD_WINDOWS, 19 FR_OP, 20 P2_FIELD, 21 FIELD_IO, 22 FQ_OP one field cell, or one coprocessor operation

Every cycle goes to the one execution family whose decoded table claims its pc. Families interleave in time: ADD_SUB_LUI_AUIPC may own cycles 1 and 3 and JUMP_BRANCH_SLT cycle 2. Nothing needs them to be contiguous, because the memory argument orders every row by its pc write.

The window families exist because the memory argument needs every address an execution touches to have exactly one initial value and one final value. INIT_TEARDOWN covers RAM window 0 and starts it with the program's image; ZERO_WINDOWS covers every other window of ordinary RAM the run touched and starts it at zero; the public pair covers the input and journal windows; ADVICE_WINDOWS covers the advice region, starting it with the prover's bytes.

Shards#

A family's rows, in the order they appear, are cut into shards of the family's height. The last is padded with zero rows, which the circuits are built to accept. A shard costs its full height whatever its occupancy, so the families a program touches and the heights it chooses set the floor of every proof.

Family Default height Why
Instruction families 2^22, 2^20 for MUL_DIV and ATOMICS the floor of their timestamp range checks is 2^20
RAM windows 2^22 one shared window height, at least 2^16
Public input, journal 2^12 pinned: the height places the windows
KECCAK_F, SHA256_COMP 2^18 four times the calls of their 2^16 floor for 2% more proof
MOD_MUL, EC_ADD 2^16 their floor
POSEIDON2, FR_ARITH 2^8 no table, so no floor

A shard is proved on its own, by its family's circuit, except for the memory argument: each shard's circuit outputs the product of its read tuples and of its write tuples, and the verifier reconciles those products across every shard of the statement once. That is the only thing that joins shards. There is no per-shard chaining of the pc and no shared boundary between neighbours.

The shard proof sizes at the default heights run from about 12.5 KB for a public window to 665 KB for a POSEIDON2 shard; an instruction family's is 64 to 77 KB. Performance has the table.

The specification: Execution trace, Circuits and registry.

Architecture

The GKR Engine

The proving engine at the centre of Apogee. Why a layered GKR circuit commits only its inputs, how one backward pass of sumchecks funnels a whole circuit to a single point, and what that buys.

View as Markdown

Every shard in Apogee is proved the same way: by its family's circuit, written as a stack of layers, run backward by the GKR engine from the circuit's outputs to its committed columns. This page is about why that engine sits at the centre of the system, and why the family of proof systems it belongs to is where the frontier of proving is moving.

The idea#

The classical way to prove a computation is to lay it out as a table, commit to every column, the intermediate values included, and prove that a set of constraints vanishes over the table. Commitments are the expensive part: every committed column costs a multi-scalar multiplication or a Merkle tree, and an opening at each point the constraint argument asks about.

GKR, after Goldwasser, Kalai and Rothblum, changes what has to be committed. The computation is a layered circuit. Only the bottom layer, the inputs, is committed. Every layer above it is defined by gates over the layer below, and the prover never commits to it. Instead, a claim about the top layer is reduced to a claim about the layer beneath it by one sumcheck, then to the next, until the claims land on the committed inputs, all at a single random point. One opening settles them.

Note

An analogy, not a name. A conventional prover is a rocket: it hauls every intermediate value it produces to the destination, committed, and pays for the mass. The GKR engine behaves more like the warp drive of science fiction, which moves the space around the ship rather than the ship itself. What travels is the claim, moved down through the circuit layer by layer, while the intermediate layers are never carried anywhere at all.

What that buys a zkVM:

  • Intermediate values cost no commitment. A family circuit can compute hundreds of inner columns, product trees and fraction trees, and none of them is ever committed. Only its trace columns are.
  • One opening point per shard. The backward pass ends with every committed column claimed at the same point. A shard needs exactly one batched opening, 704 bytes, however many columns it has.
  • Prover work is field arithmetic. Each layer's sumcheck is linear in the layer's size, over Fr, with no commitment, transform or hash per layer.
  • Arguments compose inside the circuit. The memory argument's grand products and the lookups' LogUp sums are just more layers of the same circuit, reduced in the same pass.

A family circuit, layer by layer#

The layers of a family circuit Layer 0 is the committed columns beside virtual tables. Gate list 0 builds memory leaves, lookup fractions and enforcing gates. Row-wise lists reduce each row's trees to one node. Halving lists, one per variable, multiply and add the rows down to a top with no variables: the read root, the write root and each channel's numerator and denominator. The backward pass runs from the top to layer 0 with one sumcheck per transition and ends at one point, where one Mercury opening settles every committed column. Layer 0 · committed columns M ‖ W ‖ S beside virtual tables V, closed forms the verifier evaluates · the only layer committed Gate list 0 · row-wise memory leaves · lookup fractions (num, den) · every enforcing gate, degree ≤ 2 Row-wise lists 1 … r product and fraction trees combine siblings until each row holds one node per tree Halving lists, one per variable TreeProduct · TreeCross: rows (·,0) and (·,1) combined Top · no variables read root · write root · (num, den) per channel BACKWARD PASS outputs absorbed first a sumcheck per transition: claims on layer k+1 become claims on layer k halving: a line point τ merges the two children every committed column claimed at one point u → ONE MERCURY OPENING FORWARD PASS · PROVER ONLY
One family circuit. The prover computes every layer upward once (dashed). The proof runs downward: the outputs are absorbed, then each transition is a sumcheck that turns claims about one layer into claims about the layer below, until all of them meet at one point on the committed columns.

The bottom layer is the shard's committed columns, in three kinds that differ in when they are bound: M, the memory columns, committed in the statement's global transcript before any memory challenge exists; W, the witness columns, committed in the shard's own transcript; and S, the setup columns, bound by the program identity or by the ceremony. Beside them sit virtual tables: closed forms such as the row index or the 16-bit range, which the verifier evaluates at any point and which are never committed.

Above layer 0 every family circuit has the same anatomy:

  1. Gate list 0 computes, row by row, the memory leaves (each query's read and write tuples), the lookup fractions (one (numerator, denominator) pair per lookup, plus the table's), and every enforcing gate: the family's constraints, each a polynomial that must vanish on every row.
  2. Row-wise lists combine sibling leaves: a product tree multiplies tuples, a fraction tree adds fractions as (n_a·d_b + n_b·d_a, d_a·d_b), until each row holds one node per tree.
  3. Halving lists, one per variable of the shard's height, combine the rows pairwise: the first half with the second. After n of them the circuit reaches a top with no variables: the shard's read root, its write root, and each lookup channel's final numerator and denominator.

So one circuit proves the family's constraints, computes its contribution to the memory argument, and sums its lookups, in one pass. Every gate has degree at most 2, so every round polynomial of every sumcheck is a cubic.

The backward pass#

The prover materializes every layer once, upward. The proof then runs downward, and its transcript schedule is the same for every circuit:

  1. The outputs are absorbed first, so the prover is committed to the roots before any challenge exists.
  2. For each transition from layer k + 1 down to layer k, a challenge λ batches every claim on layer k + 1, together with every enforcing gate of the list, into one sum. A sumcheck reduces that sum to an evaluation at a random point ρ, one cubic message per variable.
  3. The prover states the values of layer k's columns at ρ. For a halving list it states both children, and a further challenge τ merges them into one claim per column.
  4. At layer 0 every committed column has one claim, all at the same point u.

The shard's single Mercury opening proves those claims against the commitments: the memory columns' from the statement, the witness columns' from the shard proof, the setup columns' from the verifying key. Virtual tables are evaluated by the verifier itself.

Enforcing gates ride along for free. An enforcing gate claims 0 everywhere, so it joins the batch of the transition it sits in, and a violated gate makes the batched sum nonzero with overwhelming probability. A single LayerInconsistency error covers a wrong descending claim and a violated gate alike; a batched sum cannot tell them apart, and the proof spends nothing on telling them apart.

Why it is sound#

Each challenge is drawn after everything it protects:

  • the output point after the outputs, so a prover cannot pick tables that agree with the truth only where it will be checked;
  • λ after the claims and the point, so a false claim or a violated gate survives only at a root of a nonzero polynomial in λ;
  • each sumcheck challenge after its round's cubic, so a wrong cubic agrees with the true one with probability at most 3/|Fr|;
  • τ after both children's values.

Summed over every transition of a registered circuit at its default height, the soundness error stays under 2^14/|Fr|, with Fiat–Shamir over the Poseidon2 transcript in the random-oracle model.

Circuits as data#

A circuit is not code. It is a CircuitArtifact: its committed columns by name, its virtual tables, its gate lists with seven gate shapes, a flat list of the same relations, its lookups and its padding row, serialized canonically. Four laws hold every artifact to a consistent form: every operand is readable where it is read, every list's width is derived from its gates, the top layer is exactly the outputs, and the layered gates and the flat relations are one constraint set. They run once, where an artifact is built or a key is loaded, never per proof.

Two consequences matter to anyone evaluating the system:

  • A verifying key carries its circuits, and the verifier holds them to its own registry. The program identity binds the program; the registry binds the circuits that prove it.
  • The circuits can be checked by a second implementation. The checker crate re-implements the laws, the lookup rules and the padding contract without sharing the constructors' code, and evaluates gates only through the one gate kernel both sides use as the semantic authority.

The cost#

The engine trades commitments for memory. The forward pass holds every inner layer as field elements: about 8.4 GiB for a 2^20 shard of the widest instruction family, and 42 GiB for a 2^18 KECCAK_F shard, whose circuit computes 5,490 inner columns. That is why proving is memory-bound, why heights are a tuning parameter, and why the streaming prover bounds memory by the shards in flight. Proof size grows only by one sumcheck round per variable per layer: a KECCAK_F shard's proof is 381,100 bytes at 2^18 against 373,276 at 2^16, for four times the work.

The specification: The GKR engine, and every family's own page under Auditors.

Architecture

Commitments

Every committed column is opened with Mercury, a multilinear commitment over KZG with a constant-size opening. What it costs, how a shard batches all its columns into one opening, and how recursion defers the pairing.

View as Markdown

Every column Apogee commits is a multilinear polynomial, a table of 2^n evaluations over the boolean cube. Every one of them is committed and opened with Mercury (Eagen and Gabizon, ePrint 2025/385), finished by the batched KZG opening of BDFG20 (Boneh, Drake, Fisch and Gabizon, ePrint 2020/081). The specification pins what the papers leave open and adds two things: a batch of many columns at one point, and a deferred form that the recursion tree folds.

The commitment#

A Mercury commitment is exactly the KZG commitment of the evaluation table read as coefficients: one multi-scalar multiplication over the first n powers of the ceremony's τ, one G1 point, 64 bytes. There is no second scheme behind it. Two properties follow and the rest of the system uses both:

  • It is linear. The commitment of Σ ρ^i·f_i is Σ ρ^i·cm_i, which is what lets many columns share one opening.
  • Zero coefficients add nothing. A column extended by zero rows keeps its commitment, so the generic lookup table's three commitments serve every height that holds the table, and they are a constant of the ceremony.

Columns are committed at their integer width: a bit, byte, halfword or word column goes through an MSM over u32 scalars, recoded from 32 bits instead of 254, which keeps committing a trace cheap.

The opening#

Mercury splits an opening point u of s = 2t variables into halves, folds the polynomial by a challenge α, and proves the two inner products that remain with a symmetrized witness, finishing with one batched KZG opening at three points. The proof is eight G1 points and six field elements: 704 bytes, for every size. Hence one requirement: the number of variables must be even, which is why every height on the menu is an even power of two.

The verifier does O(t) field operations, MSMs of ten and two points, and one pairing check of two pairs. Both relations it checks are written as e(A, [1]_2) = e(B, [x]_2), so both G2 arguments are constants of the setup and the verifier does no G2 arithmetic at all. That shape is also what lets recursion postpone the pairing instead of computing it.

A shard's columns, one opening#

The GKR pass ends with every committed column of a shard claimed at the same point u. So nothing needs a claim-merging sumcheck: the opening batches all of them. A challenge ρ, drawn after every commitment and every claimed value, weights column i by ρ^i; the prover opens Σ ρ^i·f_i once, and the verifier forms Σ ρ^i·cm_i by a k-point MSM. A false claim survives with probability at most (k − 1)/|Fr|.

The columns in the batch come from three places, in a fixed order that is part of what is proved: the memory columns' commitments from the statement, the witness columns' from the shard proof, and the setup columns' from the verifying key. Taking the setup commitments from the key is what makes the opening bind the decoded tables and the image the program identity commits.

Deferred verification#

A deferred check runs everything but the pairing and keeps the relation's terms as twelve (side, scalar, point) entries. Recursion uses exactly this: each node of the tree computes a shard's twelve scalars in its own arithmetic, weights them by fresh challenges, and adds them into one running pair of points (A, B). Every shard's opening in the whole tree folds into one claim e(A, [1]_2) = e(B, [x]_2), which only the contract on Ethereum finally checks. Recursion and settlement shows the fold.

Measured#

On an 18-core Apple M5 Pro:

Operation Time
Commit a 2^22 column 1.30 s
Open a 2^22 column 2.89 s
16 columns of 2^20, opened as one batch 1.01 s, verified in 4.8 ms
The same 16 opened one by one 9.79 s, verified in 62 ms

Security#

Knowledge soundness holds in the algebraic group model under q-DLOG, with Fiat–Shamir over the Poseidon2 transcript in the random-oracle model, and an SRS whose τ nobody knows. The statistical terms, Schwartz–Zippel over the challenges, stay below 2^−220 for every instance in use, so the security level is BN254's, about 100 bits. Nothing is hiding and nothing is blinded: Apogee v1.0.0 is not zero-knowledge.

The SRS is the PSE perpetual powers of tau, contribution 80, sound while one contributor was honest. Its file is decoded and every point checked to lie on its curve and in the right subgroup; nothing proves that a file is that ceremony's, which is why a verifier compares the SRS digest with the ceremony's from its own channel.

The specification: Mercury, The structured reference string.

Architecture

Memory and Lookups

Two arguments carry everything that crosses a row. One read/write multiset over the whole execution, reconciled once, makes every read return the last write and orders every row; LogUp channels make every value a byte, a word or a table row.

View as Markdown

A family's gates constrain one row at a time. Everything that crosses rows, shards or families, such as what a register holds, what a load returns, which instruction a row executes and whether a value fits in 32 bits, is carried by two arguments that live inside the same GKR circuits.

The memory argument#

Every memory access becomes a field element, a tuple:

T(AS, ADDR, TS, VAL) = γ + AS + α_addr·ADDR + α_ts·TS + α_val·VAL

over the address space, the address, the timestamp and the value, with four challenges drawn once per statement. A query contributes its read tuple to one side and its write tuple to the other. Each shard's circuit outputs two numbers, the product of its read tuples and the product of its write tuples. The verifier then checks one equation over the whole statement:

∏ read roots · R_b  =  ∏ write roots · W_b          over every shard of every family

W_b and R_b are the initial and final tuples of the 32 registers and the pc, which have no rows of their own: the verifier multiplies them in from 64 boundary scalars the statement carries. RAM gets its initial and final values from the window families' shards, which hold one word per row: the program's image in window 0, zeros in every other window the run touched, the statement's input in the public input window, the prover's bytes in advice.

If the equation holds, the multisets are equal with overwhelming probability. Equal multisets mean every read returns the last write before it: each read is matched to exactly one write, a read must strictly follow the write it consumes, and every address has exactly one initial write.

Order for free#

The pc is a memory cell like any other, at address 0 of its own space. Every row reads the pc and writes the next one, and its circuit holds the write at least four timestamps after the read. So the pc's history is one path through every live row of every family, from the entry point to the exit row. That single path gives:

  • program order, since rows are ordered by their pc writes;
  • continuity across shards and families, since every row's pc read consumes some row's pc write;
  • no cycle proved twice, since no write can be consumed twice.

No shard chains to its neighbour, and none needs to. A shard's claimed time window binds nothing; it is checked only for shape.

What must come first#

The memory challenges are drawn once per statement, at the end of the global transcript, after everything a tuple can read is fixed: every shard's memory commitments, the program identity (which fixes the entry pc and the image), the ceremony digest, the shard counts and the window list, the digest of the public input and journal, and last the 64 boundary scalars. A value chosen after the challenges could be solved for; the order of the transcript is what forbids it. For the same reason a memory tuple may read only memory, setup and virtual columns, never a witness column, which is committed in the shard's own transcript after the challenges. The circuit constructors refuse any artifact that breaks this.

Lookups#

A lookup says that a tuple of a row's values is a row of some table. Apogee proves every lookup of a shard with LogUp: one identity per table, or channel,

Σ_rows Σ_lookups 1/(E(y) + g)  −  Σ_rows mult(y)/(T(y) + g)  =  0

summed by a fraction tree inside the family's own GKR circuit and checked at its root: numerator zero, denominator nonzero. The multiplicity column needs no constraint at all: a tuple in no table row leaves a pole the multiplicities cannot cancel.

Channel Table Used for
TIMESTAMP [0, 2^19), virtual each query's timestamp gap, as two 19-bit chunks
RANGE16 [0, 2^16), virtual 32-bit values as two halfwords; carries; frame bounds
XOR8 all byte pairs and their XOR, virtual Keccak and SHA-256, byte by byte
GENERIC a committed table of AND, sign and shift-power rows bitwise operations, sign bits, shift amounts
DECODER the family's decoded table, committed by the identity binding each executed row to the program

Three of the tables are virtual: closed forms of the row index that the verifier evaluates itself, which cost no commitment. The generic table is committed once by the ceremony's powers and covered by the SRS digest.

The decoder lookup#

Every execution family makes one lookup per live row into its own decoded table, keyed by the pc the row read from memory. That single lookup binds the cycle to the program: the row's operands, immediate and instruction kind are the program's at that pc, and its kind bits are one-hot because every live row of the table holds a one-hot mask and every padding row holds −1, which no sum of kind bits reaches. A row at a pc where the program has no instruction finds no table row at all.

Keys must be bounded#

A channel proves membership of a table and nothing more. Several sub-tables share the generic table under disjoint key ranges, so an unbounded key could land in the wrong sub-table and prove a false AND. Every family therefore bounds each key it looks up, with a range lookup under the same selector, and the circuit constructors check that a bound written through a scaling factor also carries a direct bound. The specification states the attack this prevents for every family.

How it adds up#

Together with each family's gates, these two arguments give the statement its meaning: each row obeys its instruction, the instruction is the program's, every read sees the last write, the rows form one path from the entry to the exit, every value is the integer it claims to be, and the public windows hold the statement's bytes. The soundness map carries each claim to the sections that prove it.

The specification: The memory argument, Lookups.

Architecture

Delegations

How an expensive function gets a circuit of its own without growing the instruction circuits. The call, the anchor that pairs each request with exactly one invocation, the six circuits, and their economics.

View as Markdown

Hashing and big-integer arithmetic dominate real workloads: in a mainnet Ethereum block, secp256k1's field multiplication and squaring alone were 44% of the cycles before they were delegated. Proving them instruction by instruction is possible and slow. A delegation gives such a function a circuit family of its own, invoked from the guest, so the instruction circuits stay small and a program pays only for the delegations it calls.

The call#

A delegation is invoked, never decoded. The guest writes a frame of 32-bit words in RAM and issues an ecall with the delegation's number in a7 and the frame's base address in a0. The ecall is one row of the ADD_SUB_LUI_AUIPC family, the request. The work is one row of the delegation's own family, the invocation, which reads every frame word and writes every frame word back, the results among them, at the requesting cycle. An invocation owns no cycle; it rides the one that asked for it.

Frames are word-aligned and lie wholly in ordinary RAM, so no frame overlaps a public window or advice, and the frame's reads and writes are ordinary memory queries. What a delegation computed is therefore bound exactly as any store is: through the one memory multiset.

The anchor#

Requests and invocations must pair one to one: otherwise many requests could close against one invocation and leave calls unexecuted, or an unrequested invocation could rewrite a frame. They pair through the same memory multiset, in an anchor space that belongs to the delegation type alone and that no instruction can reach:

Reads Writes
Request, cycle c T(s, base, 0, 0) T(s, base, 4c + 3, v)
Invocation T(s, base, 4c + 3, v′) T(s, base, 0, 0)

Three gates on the request side fix its read at timestamp 0 and value 0 and make it write 0 to a0. Then the tuples stamped 0 are exactly the requests' reads and the invocations' answers, so there are as many invocations as requests over the same bases; and since no two requests share a cycle, each invocation's read is exactly one request's write. Every invocation sits at its request's base and cycle. No gate in a delegation circuit had to know about requests at all.

Many calls, one operation#

An operation too wide for one row is several invocations on one frame, a frame word naming the step: a keccak-f[1600] permutation is 24 round calls, a SHA-256 compression 16 calls of four rounds, a complete point addition three calls. No gate joins two rows. Each call proves its step on the frame as it finds it, its reads lying on each word's single memory history, so it reads the previous step's writes. That every step runs, in order, is the calling code's to ensure, and the calling code is guest code proved as instructions. The SDK issues each multi-call operation from one function, so a guest never orders the steps by hand.

Declared statically#

The instruction sweep cannot see a call, because the number is a run-time value of a7. So each shim in the SDK leaves a 12-byte declaration record in its own linker section, kept only if the shim is reachable. The program derivation scans the image for records, and a declared family joins the configuration, bound by the identity through the image bytes. A family linked but never called proves zero shards; a called number whose family the program never declared has no proof.

The six circuits#

Family One invocation Built from
KECCAK_F one round of keccak-f[1600] over a 51-word frame bytes: 1,020 XOR8 lookups a round; rotations as linear forms over bytes and masked copies
SHA256_COMP four rounds and four schedule words bytes and XOR8: 52 obligations a round, 32 a schedule word; Ch and Maj as linear forms in XORs
POSEIDON2 one width-3 permutation the rounds computed in the circuit's own layers, three gate lists a round, with no lookup; the only delegation that computes above its first layer
FR_ARITH one Fr add, multiply or inverse in Montgomery form bit decompositions and canonicity chains against p
MOD_MUL one 256-bit a·b mod m, four Ethereum moduli 32-bit limbs, a quotient, carries, and a canonicity chain proving out < m
EC_ADD one third of a complete point addition on secp256k1 or BN254 G1 Renes–Costello–Batina's complete formula as three reductions a row

A few constructions recur across them. A one-code rule decodes a frame word naming one of k cases into boolean selectors with exactly one set, because codes add: without it, selectors 1 and 3 answer a request for 4. A canonicity chain proves a 256-bit value is below a modulus through borrows over 32-bit limbs. And every written word is bounded below 2^32, so that RAM stays words, which every instruction family relies on.

The economics#

A delegation family's height sets how many calls one shard holds, and a shard costs its height whatever its occupancy:

Family Height Units a shard Shard proof
KECCAK_F 2^18 10,922 permutations 381,100 B
SHA256_COMP 2^18 16,384 compressions 189,988 B
EC_ADD 2^16 21,845 additions 434,916 B
MOD_MUL 2^16 65,536 multiplications 135,220 B
POSEIDON2 2^8 256 permutations 664,780 B
FR_ARITH 2^8 256 operations 266,292 B

For a family with many calls, the fatter shard is the cheaper one: a KECCAK_F proof barely grows from 2^16 to 2^18. The price is memory. A 2^18 KECCAK_F shard's forward pass is 42 GiB, and two of them in flight set the measured block's peak.

What is delegated, and what is not#

Library code reaches the delegations through patched copies of k256, ark-ff and revm-precompile: secp256k1 recovery becomes k256 code over MOD_MUL and EC_ADD, and a BN254 pairing becomes ark-bn254 code over MOD_MUL. The EVM's MULMOD with an arbitrary modulus, MODEXP, BLS12-381 and every whole signature scheme run as instructions. Dedicated signature support for guests is part of the v2.0.0 trajectory.

The specification: Delegation ABI, Delegation circuits.

Architecture

The Streaming Prover

The prover executes the guest twice and never holds the execution trace. Memory follows the shards being worked, not the length of the run, and the proof does not depend on the schedule.

View as Markdown

A full Ethereum block is about 200 million cycles. Its trace, every row of every family and every memory event, would be about 300 bytes a cycle: tens of gigabytes before a single column is committed. Apogee's prover never builds it. It executes the guest twice and holds only the shards it is working on.

Two passes#

The two passes Pass one executes the guest and, shard by shard as each fills, commits its memory columns and keeps only the commitments; at the exit it builds the statement and draws the shared challenges. Pass two executes again and proves each shard as it fills, keeping only the proof. PASS 1 · COMMIT Execute Shard fillscommit M columns · drop rows At exitwindows · statement · G1–G11 Challenges fixedmemory challenges · global digest PASS 2 · PROVE Execute again Shard fillsevery column · GKR · opening Keep the proofdrop everything else BlockProofroots into the statement every shard's transcript is seeded from the global digest
Commit, then prove. The memory challenges must follow every shard's memory commitments, so the commitments come first, from one execution, and the proofs from a second.

Pass 1 executes the guest and, as each shard fills, commits its memory columns, keeps the commitments and drops the rows. At the exit it derives everything else the statement needs from the final memory state: the register and pc boundary, the list of memory windows the run touched, and the window families' shards. Then it runs the global transcript, which absorbs the statement, every memory commitment included, and draws the memory challenges and the digest every shard is seeded from.

Pass 2 executes again. The emulator is a pure function of its input, so it cuts the same shards, and pass 2 asserts that its cycle profile, window list and boundary are pass 1's. Each shard gets every committed column and is proved: its transcript, its GKR pass, its opening. The memory columns are not recommitted: the opening takes their commitments from the statement and their values from pass 2, so columns that differed between the passes would give an opening the verifier refuses.

The order is forced by soundness. The memory challenges must follow every value a memory tuple can read, so every shard's memory columns are committed before any shard can be proved.

The pipeline#

A fixed number of workers, max_in_flight, share one lock around the executor. Under the lock a worker hands back its finished shard and claims the next: a filled shard if one is waiting, and otherwise it steps the executor itself until a buffer fills. Outside the lock it builds the shard's columns, proves it and drops it.

  • The executor never runs ahead of demand. At most one filled, unclaimed shard per family waits, as rows.
  • Within a shard, the work is data-parallel across every core. A worker blocked in that work does not take a second shard.
  • The block does not depend on the schedule. A shard's proof is a function of the global state and its own columns; proofs are placed by statement position. The bytes are identical at 1 and at 8 shards in flight.
  • Failures are deterministic. The failure returned is the earliest in fill order, at any worker count.

What it costs#

Memory is a partial buffer per family, the last-access tables, the shards being worked and the output. A shard's working set is dominated by its forward pass, every inner GKR layer as field elements: 8.4 GiB for a 2^20 SHIFT_BITWISE shard, 42 GiB for a 2^18 KECCAK_F one. So max_in_flight bounds how many of those coincide, and heights set how large each is.

Measured on block 257,510 (60 transactions, 101.5 Mgas, 198M cycles, 207 shards) on 32 vCPUs and 247.7 GiB, twelve shards in flight:

Pass 1 191 s; one-thread fills take 81% of its shard-seconds; sampled memory at most 15.9 GiB
Pass 2 2,290 s, with about 12 of 12 shards held and 30.4 vCPUs busy until the guest exits, then a 460 s tail
Peak memory 173.92 GiB: the two 2^18 KECCAK_F shards, together in the tail with nothing else in flight

The peak came from one delegation family's height, not from the twelve shards in flight. That is the lever: a block with fewer Keccak calls, or Keccak at a lower height, peaks lower.

The specification: The streaming prover.

Architecture

Recursion and Settlement

How a base proof of hundreds of shards becomes one Groth16 proof that a contract checks. Apogee proving its own verifier, the tapes that make that cheap, the transcript chained across a tree, and pairings folded until only one is left.

View as Markdown

A base proof of an Ethereum block is 207 shard proofs: 14.5 MB, each shard a GKR proof with its commitments and a Mercury opening that ends in a pairing check. A contract can check none of that directly. Recursion compresses it, and the way it does is the most consequential design choice after the GKR engine itself.

The base proof is untouched#

The first decision is what recursion does not do. No base key, statement or proof changes to make recursion possible: a leaf verifies base shards exactly as a native verifier would. Everything recursion needs is added above the base proof, never inside it. A block can be verified natively, recursed, or both, from the same bytes.

Apogee proves its own verifier#

A node of the tree is Apogee proving a verifier program. A leaf verifies a run of consecutive base shards, from from to to; an internal node verifies two to four children, each a whole proof of a leaf or node program; the root covers every base shard. Every node is proved by the same streaming prover as a base block.

Running the Rust verifier as RISC-V instructions would work, and measured 3.0 billion cycles for one block's 207 shards: fifteen times the block itself. So nodes run in a recursion format instead.

The recursion format#

A statement is in the recursion format exactly when its program declares field families. The format adds one address space and four coprocessor families on it, all invoked through the ordinary delegation ABI, and changes nothing else:

Family One row is
FIELD_WINDOWS one field cell: a memory cell holding a whole Fr element, in the same memory multiset as RAM
FR_OP one field operation over cells: multiply, add, subtract, multiply-accumulate, invert, assert-equal, and constant-building steps
P2_FIELD one Poseidon2 duplex step over cells, so a transcript runs at one row per permutation
FIELD_IO eight RAM words into a cell, or a cell back into eight words
FQ_OP one BN254 base-field operation, an element being four cells of 64-bit limbs, so curve arithmetic runs at one row per field operation

Two further changes make a recursion shard cheaper to verify by its parent. Its memory and witness columns are committed as stacks of up to 2^24 evaluations, so a parent folds a handful of points instead of hundreds. And a recursion request writes its frame base advanced past the frame, so frames laid back to back replay as back-to-back ecalls, one row a call.

Tapes#

A shard's checks have a fixed shape for its family and height. So the host compiles them, once, into a tape: a straight-line list of coprocessor calls over absolute cells, in which nothing branches on a value. A shard's tape is the native verifier's steps for that shard, call for call: the shard transcript, the GKR backward pass, the lookup and root checks, and the Mercury opening's twelve scalars. Every check is an assert-equal.

Each recursion program's tapes, fold templates and constants are built at compile time by the verifier crate itself and placed in the program's read-only data. The program identity therefore binds every tape the program replays: proving that a node ran its program is proving that it ran exactly these checks.

A transcript chained across the tree#

The base statement's global transcript is one sponge over the whole statement. The tree splits it without changing it. The node holding shard 0 runs the prefix, through the public input digest; every node absorbs its own shards' memory commitments, continuing from the state its predecessor left; the node holding the last shard runs the suffix and draws the memory challenges every node had taken as claims. A node's journal records the chain's state at both ends of its range, and a parent holds its children's states to meet.

A node also holds its children to one another: exit status 0, one base statement (its shape, digest, challenges, input and journal digest, exit status and shard count), adjacent shard ranges, chain states that meet, time windows in order across the seam, and the identities of the recursion programs. A node that holds a whole statement makes the memory argument.

Folding the pairings#

No node computes a pairing. Each shard's Mercury check is deferred as twelve (side, scalar, point) entries; after the shard's tape, the node's own transcript absorbs the shard transcript's final state and draws weights, and every entry is added, weighted, into one running pair of points (A, B) representing the claim e(A, [1]_2) = e(B, [x]_2). The batch check that ties a shard's combined commitment to its columns is folded beside it. Points every shard of a family shares, such as [1]_1 and the setup commitments, accumulate one scalar each and enter once. A child's (A, B) enters under a weight drawn after its whole journal.

Each side is one multi-scalar multiplication on FQ_OP, run as a static template: Pippenger with 8-bit digits over GLV halves, every point held to the curve, every step fixed in advance. A point costs about 400 FQ_OP calls.

At the root, the tree's whole content has collapsed: every base shard verified, the transcript run end to end, the memory argument made, and every opening folded into one pairing claim. What remains is that claim and two program identities.

The decider#

The root is still a GKR proof and some hundreds of points, which a contract cannot check. The decider is a Groth16 circuit that runs the node procedure over one child, the root, through a driver that writes rank-1 constraints instead of coprocessor calls, and holds the root's journal to the whole range of base shards. It folds nothing: each point the root owes the final pairing, with its scalar, becomes a bound wire, a value the verifier holds, committed in the proof under a fifth trapdoor rather than passed as a public input. So are the two identities, the base exit status, and the base public input and journal, byte by byte.

Apogee's Groth16 differs from the textbook in three ways: the bound-wire commitment, no blinding, and a proving key over the Lagrange basis that the powers-of-tau ceremony already publishes. Its key comes from a two-phase ceremony: phase 1 is the same ceremony file the commitments are under; phase 2 is the circuit's own, with contributions to α and β finished before any to γ, δ and η, an order that is itself part of soundness.

ApogeeVerifier.sol rebuilds the bound values from calldata, checks the Groth16 equation, folds both sides' points with ecMul and ecAdd, which also holds every point to the curve, and checks the one remaining pairing. Its constructor fixes the key, the ceremony's two G2 points and the two recursion programs' identities. A deployment serves one base program, one root shape and fixed public-value lengths.

Measured#

Block 257,510, tree on a 32-CPU machine, ceremony and decider on an 18-core laptop:

Base proof 207 shards, 14.5 MB, 2,481 s
Tree 4 leaves of at most 64 base shards and a root: 116 shards in all
Leaves, four at once 21, 24, 23 and 27 shards; 2,157 s; 92 GiB peak
Root 21 shards, 460 s, 1.03 MB
Decider 7,896,686 constraints; proof in 18.5 s and 6.1 GB
Contract 358 points; 3,620,026 gas; 34,980 bytes of calldata

The specification: Recursion and decider. Running it yourself: Settle on-chain.

Architecture

Ethereum Blocks

Apogee's reference workload. A revm guest that runs Ethereum blocks inside the VM, a stateless validator that matches every case of the zkEVM test release, and what a proof of a block says.

View as Markdown

Apogee proves arbitrary RV32IMAC programs. Its reference workload, the one it is measured and tuned on, is the hardest common one: validating an Ethereum block inside the VM with revm, the Rust EVM. One library, revm_block, compiles both for the guest and for the host, and two binaries prove two different statements.

Two binaries, two statements#

Binary Advice Journal Says
revm-block a BlockWitness: the pre-state the transactions read per transaction, its status, gas and return data; a logs commitment; a post-state summary some canonical witness makes revm_block::run produce this journal: a proof of an execution, not of a block's validity
revm-block-stateless the stateless input of the zkEVM benchmark format 43 bytes: the payload's root, the verdict, the chain id, the schema id the payload with this root is, or is not, a valid block on this chain under this fork

The stateless validator is the one that proves blocks. It implements verify_stateless_new_payload from Ethereum's execution specifications: it decodes the request, checks the ancestor headers and the header rules against the parent, recovers every sender, executes every transaction against a pre-state held to the parent's state root by hashes, applies withdrawals and requests, and recomputes the receipts root, bloom, gas, requests hash, block access list and post-state root. The witness needs no binding of its own: the published root fixes the payload, and the witness is held to it by hashes, so a wrong witness cannot make an invalid payload valid. A verdict of false means only that this input did not validate.

Forks and conformance#

The validator names the fork by the input's schema id, with no activation schedule compiled in: Osaka, BPO1, BPO2 and Amsterdam. All 67,251 pairs of the tests-zkevm v21.0.1 release match natively, and CI holds the library to a committed subset of 34 cases covering every rule the release reaches, in both input layouts.

Delegations in practice#

Both binaries declare KECCAK_F, SHA256_COMP, MOD_MUL and EC_ADD. Keccak reaches its circuit through alloy-primitives' native-keccak hook. SHA-256, secp256k1 and BN254 reach theirs through patched copies of revm-precompile, k256 and ark-ff, each with upstream's code as its fallback path. Every sender is recovered in the guest under EIP-2's rules, with k256 arithmetic that the patches route to MOD_MUL and EC_ADD.

The binaries are built --release and proved with every family whose height is a choice at 2^20: the stateless binary's .text is about 1.96 MB, 96.6% of what a 2^20 table reaches.

The measured block#

Block 257,510 of glamsterdam-devnet-8, through revm-block-stateless: 60 transactions, 101.5 Mgas, 198 million cycles, cut into 207 shards. The base proof took 2,481 s on a 32-vCPU, 247.7 GiB machine and peaked at 174 GiB, recursion added about 2,620 s, and the contract accepted the result for 3,620,026 gas. Performance breaks each stage down.

What a block proof does not do#

  • The validator takes its input from an external witness producer. eth_getProof returns the trie nodes on each key's path, but a deletion that collapses a branch needs a sibling node that is on no changed key's path. The repository's recorder therefore cannot produce stateless inputs; they come from a tests-zkevm release or the zkEVM benchmark's datasets.
  • The mini-block binary proves an execution over a recorded pre-state, usually a block's first few transactions, and makes no state-root claim. Its journal grows by a record per transaction and outgrows the public window, which is why full blocks go through the stateless validator.

The specification: Ethereum blocks.

Architecture

Security Model

What a proof establishes, what it assumes, what a verifier must hold for itself, which code soundness rests on, and the limits of version 1.0.0.

View as Markdown

What a proof establishes#

A proof that verifies establishes that the program of a given identity, started at its entry pc over its image, with the public input in its input window and some advice of the prover's choosing, executes instruction by instruction to EXIT with a given status, having written a given journal. Nothing is claimed of the advice. Nothing is hidden.

Assumptions#

Assumption Where it enters
Knowledge soundness of Mercury and KZG in the algebraic group model under q-DLOG every commitment opening
Poseidon2 as a random oracle for Fiat–Shamir every challenge, in the base proof and in recursion
Groth16's own assumptions the decider, the last step to the contract
One honest contributor to the PSE perpetual powers of tau the SRS every commitment is under
One honest contributor per round of the decider's phase-2 ceremony the decider's key

BN254 gives about 100 bits of security. The statistical error of every protocol layer, from the sumchecks and the batched opening to the memory and lookup arguments, is far below that: under 2^14/|Fr| for a circuit's whole GKR pass, below 2^−190 for every lookup channel, below 2^−220 for every Mercury instance.

What a verifier must hold for itself#

Two values, from a channel the prover does not control:

  • The program identity. Against a prover-supplied identity a proof shows only that some program ran.
  • The ceremony's SRS digest. A key loads under whatever digest its own points give, so a key built over a known τ is refused only by this comparison.

The verifying key itself may come from anyone. Loading it recomputes the identity and the SRS digest from its own contents and holds its circuits to the verifier's registry: the identity binds the program, the registry binds the circuits. The verifier command-line tool compares the identity only, taking the SRS digest from the key; host::verify compares neither, and leaves its caller to check the statement's input, journal and exit status too.

Trusted code#

Soundness is the verifier's alone. It rests on constants, field, curve, transcript, poly, sumcheck, pcs-verify, pcs, gkr-verify, verifier-core, verifier, and constraints, because the circuits are part of the statement and a missing gate is a soundness bug. Computing an identity from an ELF also trusts loader, isa and program. The last step to the chain adds the recursion programs, groth16, the decider's circuit and the contract.

The prover, the emulator, the trace builders and the proving half of the host SDK are untrusted. The prover validates nothing; a wrong input costs an honest prover a proof that fails, and a cheating prover runs none of this code anyway.

Not constant-time, not zero-knowledge#

Nothing in the code is constant-time: reductions, exponentiations, point additions and scalar ladders branch on their operands. That is harmless here because no proof is zero-knowledge, so proving keeps nothing secret. The one secret the code handles is a decider ceremony contributor's factor, which goes through the same variable-time ladder; run contributions on a machine you control.

Limits of v1.0.0#

Limit Detail
Not zero-knowledge no blinding in Mercury, GKR or the decider
Advice is unbound a guest checks it against something a proof binds
Public values at most 16,380 bytes each of input and journal
sc.w always succeeds the one deviation from RV32IMAC's semantics; there is no reservation state
Traps are not provable a misaligned access, an access outside mapped memory, ebreak, or a pc with no instruction ends an execution with no proof
Code is static the instruction stream is the image decoded at load; one undecodable word in executable code refuses the program
Code size .text within a decoded table's reach, 7.94 MiB at 2^22; the image within 4 MiB by default
Execution length 2^36 − 1 cycles
Delegations are a fixed set six in the base format; a delegation proves one step of its function, and composing steps, validating curve points among them, is the calling code's
Prover memory set by the shards in flight: the measured block peaked at 174 GiB
Block witnesses the stateless validator takes its input from an external witness producer
The decider's key one per root shape, and only as trustworthy as its ceremony; the development key is forgeable
On-chain cost about 3.6M gas for the measured block

Where the next version moves this#

Every assumption above that involves BN254, from q-DLOG and the pairing to Groth16, falls to a large enough quantum computer. The v2.0.0 trajectory is a proving core whose soundness rests on lattice problems instead.

For an auditor's view of the same model, crate by crate and argument by argument: Audit guide, Soundness map.

Architecture

Performance

Every measured figure for v1.0.0 with its source: the base proof of a full Ethereum block, the recursion tree, the decider and the contract, each family's shard proof, and Mercury's own costs.

View as Markdown

All end-to-end figures are block 257,510 of glamsterdam-devnet-8, proved through the stateless validator guest: 60 transactions, 101.5 Mgas, 198 million cycles. Each figure comes from the specification's measurements.

End to end#

Stage Machine Result
Base proof 32 vCPUs, 247.7 GiB, 12 shards in flight 207 shards, 14.5 MB, 2,481 s; peak 173.92 GiB
Recursion leaves 32 CPUs, four leaves at once 21, 24, 23 and 27 shards; 2,157 s; 92 GiB peak
Recursion root the same machine, four shards in flight 21 shards, 460 s, 1.03 MB
Decider ceremony 18-core laptop init 65 s; a contribution 50–56 s; key 70 s and 12.7 GB; the key 2.65 GB
Decider proof 18-core laptop key read in 1 s, proof 18.5 s, 6.1 GB; 7,896,686 constraints over 2^23
On-chain verification revm 3,620,026 gas; 34,980 bytes of calldata; 358 points

The base proof, pass by pass#

Pass 1, commit 191 s; 25.7 vCPUs busy on average; one-thread fills take 81% of its shard-seconds; sampled memory at most 15.9 GiB
Pass 2, prove 2,290 s; on average 11.95 of 12 shards held and 30.4 vCPUs busy until the guest exits; then a 460 s tail, whose longest stretches are the two KECCAK_F shards' one-thread fills, 200 s and 279 s
Peak memory 173.92 GiB: the two 2^18 KECCAK_F shards, together in the tail with nothing else in flight

The peak was set by one delegation family's height, not by the number of shards in flight.

A small guest#

The Quickstart's guest, 114 cycles in four instruction families, proved at 2^20 instruction heights and 2^16 window heights with two shards in flight: 7 shards, 52 s and 18 GB peak on an 18-core, 48 GiB laptop, almost all of it the two 2^20 shards being worked. The floor of a proof is set by its families and heights, not by its cycle count.

Each family's shard#

At the default heights. A shard's proof size is fixed by its circuit's shape and height; its proving cost follows its height times its circuit's width, however many rows are live.

Family Height Committed M/W/S Enforcing gates Inner columns Shard proof
ADD_SUB_LUI_AUIPC 2^22 27 / 35 / 7 63 314 64,764 B
JUMP_BRANCH_SLT 2^22 21 / 44 / 10 42 392 69,436 B
SHIFT_BITWISE 2^22 21 / 61 / 10 48 478 76,644 B
MUL_DIV 2^20 21 / 54 / 9 54 444 67,412 B
MEM_WORD 2^22 31 / 24 / 7 33 314 63,836 B
MEM_SUBWORD 2^22 31 / 55 / 10 53 472 76,196 B
ATOMICS 2^20 26 / 54 / 9 46 472 68,468 B
INIT_TEARDOWN 2^22 2 / 0 / 1 0 46 36,316 B
ZERO_WINDOWS 2^22 2 / 0 / 0 0 46 36,284 B
KECCAK_F 2^18 208 / 1,556 / 0 385 5,490 381,100 B
POSEIDON2 2^8 100 / 4,092 / 0 4,248 2,020 664,780 B
FR_ARITH 2^8 104 / 2,576 / 0 2,701 142 266,292 B
PUBLIC_INPUT, PUBLIC_OUTPUT 2^12 3 or 2 / 0 / 0 0 26 12,556 B, 12,524 B
ADVICE_WINDOWS 2^22 3 / 0 / 0 0 46 36,316 B
MOD_MUL 2^16 104 / 221 / 0 125 2,244 135,220 B
SHA256_COMP 2^18 104 / 520 / 0 119 2,802 189,988 B
EC_ADD 2^16 392 / 1,028 / 0 637 8,772 434,916 B

A height changes only the number of halving lists and sumcheck rounds, not the gates: ADD_SUB_LUI_AUIPC at 2^20 has 298 inner columns and a 57,196-byte proof against 314 and 64,764 at 2^22.

Forward-pass memory#

The GKR prover holds every inner layer as field elements, 32 bytes a cell. Representative working sets:

Shard Forward pass
SHIFT_BITWISE at 2^20 8.4 GiB
MOD_MUL at 2^16 4.6 GB
EC_ADD at 2^16 18.3 GB
SHA256_COMP at 2^18 22.6 GB
KECCAK_F at 2^18 42 GiB

Mercury#

On an 18-core Apple M5 Pro:

Commit, n = 2^22 1.30 s
Open, n = 2^22 2.89 s
16 columns of 2^20 as one batch opened in 1.01 s, verified in 4.8 ms
The same 16 opened one by one 9.79 s, verified in 62 ms

Reading these numbers#

Proving is memory-bound, and its memory follows the shards in flight and their heights, never the length of the execution. Time follows the cycle count family by family. The on-chain cost follows the number of points the root owes the final pairing, at about 9,000 gas a point. Cycle counts themselves are exact and machine-independent, so the cycle profiler is the right first tool for estimating any of the rest.

Quantum Leap

Quantum Leap

Where Apogee goes next. The trajectory toward v2.0.0 — a proving core on lattices, a field chosen to match them, signatures guests can verify, and a deployment system that carries an application from repository to running chain.

View as Markdown
BriefingProgramme: Apogee VMTarget: v2.0.0Status: active developmentRelease window:

Version 1.0.0 settles the question of whether the architecture holds at full scale: a whole Ethereum block, from a Rust guest to a contract that says true. Version 2.0.0 changes what that architecture rests on and whom it serves. Its proving core moves to mathematics a quantum computer does not break, and its developer surface grows from a repository into a system that deploys applications.

Important

This section describes work in progress. Nothing here changes the guarantees of v1.0.0, which are stated in full in the security model.

The four initiatives#

The trajectory#

v1.0.0, today v2.0.0, the trajectory
Commitments Mercury over KZG: pairings, q-DLOG Lattice-based, binding under Module-SIS
Field BN254's scalar field, 254 bits A small prime field with extension-field challenges, matched to the commitment
Against a quantum adversary Every assumption is a discrete logarithm A proving core resting on lattice problems
Signatures in guests secp256k1 through delegated field and curve arithmetic Zk-friendly and post-quantum schemes as guest calls
For builders A repository, its tools and this manual The Deployment System: portal, canonical bridges, telemetry, an AI gateway

What carries over#

The leap is in the foundations, not in the model a builder writes against:

  • The guest. Rust, RISC-V, three memory regions, an identity, a journal. Programs written for v1.0.0 keep their shape.
  • The GKR engine. Layered circuits and sumcheck are defined over any field. The engine that funnels a circuit to one point is the part of Apogee that moves to the new field most directly.
  • The arguments. One memory multiset over the whole execution, and LogUp channels, carry over as constructions; their tables and range arguments are re-derived for the new field's size.
  • The discipline. Specification first, every layer checked against an independent oracle, every forgery class held by a tamper twin.

Why now#

A validity proof is only as quantum-safe as the system that produces it. A proof over BN254 rests on pairings and discrete logarithms, so a large enough quantum computer could forge one without touching anything the guest computed. Blockchain-native applications are meant to hold value for decades. The foundation they settle on has to outlast the machines that will one day break today's curves, and the time to move it is before those machines exist.

Read the initiatives: Post-quantum proving · Signatures for guests · The Deployment System.

Quantum Leap

Post-Quantum Proving

Initiatives QL-01 and QL-02. A lattice commitment in place of the pairing-based one, and a field chosen to match it, so that the proving core no longer rests on discrete logarithms.

View as Markdown
QL-01 · QL-02Lattice commitmentsRe-fieldingStatus: active development

What breaks, and where#

Every cryptographic assumption under Apogee v1.0.0 involves BN254. Mercury and KZG are sound under q-DLOG in the algebraic group model; the recursion tree folds pairing checks; the decider is Groth16. Shor's algorithm solves discrete logarithms on a quantum computer of sufficient size, and with them every one of these assumptions. The guest's computation would still be what it was; the proof that it ran correctly would no longer mean anything.

The commitment is where the dependence is concentrated. Every column of every shard is committed with it, every opening ends in its pairing, and recursion exists to fold those pairings. Replace the commitment, and the rest of the proving core has nothing left that depends on a discrete logarithm.

QL-01 · Lattice commitments#

A lattice commitment is a linear map, t = A·s mod q, applied to a vector s with small entries. It is binding as long as nobody can find a short vector the matrix sends to zero: Module-SIS, the assumption family under ML-DSA and ML-KEM, NIST's post-quantum standards, with worst-case reductions and decades of cryptanalysis behind it.

It keeps what made KZG so useful to Apogee and that hash-based commitments give up: it is homomorphic. Commitments to many chunks combine under challenge coefficients, and the combination opens by a single vector. Batching a shard's columns, deferring checks and folding them up a tree are linear operations, and linear operations survive the move. Merkle paths cannot be combined at all.

The price is a constraint with no analogue elsewhere: the commitment binds only short vectors, every combination makes the vector longer, and the prover must show it is still short enough. The schemes of the last two years differ mainly in how they pay that price, and they have moved fast. For polynomials of 2^30 coefficients, the published Module-SIS schemes give evaluation proofs of 53 to 72 KB, and verification fell from 2.8 seconds in 2024 to 8 to 16 milliseconds in 2026. The expository Lattice-Based Polynomial Commitment Schemes surveys them scheme by scheme and reads the numbers against the hash-based side.

QL-02 · Re-fielding#

The lattice schemes do not live in BN254's world. The leading constructions work over small prime moduli, with evaluation points drawn from an extension field to keep soundness, which fits a small-field proof system and does not fit a 254-bit one. So the commitment's move brings the field with it: v2.0.0 re-fields the arithmetization, moving every circuit from BN254's scalar field to a small field matched to the commitment.

The move pays for itself:

  • Every layer gets cheaper. A GKR prover spends its time in field arithmetic, and a multiplication in a small field is a fraction of one in a 254-bit field. The engine's central economy, that intermediate layers are never committed, compounds with cheaper arithmetic on every layer that remains.
  • Commitments get cheaper. Committing a trace column is a linear map over small digits, paid per nonzero entry, rather than a multi-scalar multiplication over a curve.
  • The engine carries over. GKR and sumcheck are defined over any field. Challenges move to an extension field; the backward pass, the layer model and the arguments built on it keep their structure.

What has to be rebuilt is everything that assumed a large field: word-level values that fit one BN254 element with room to spare, range arguments and carries sized against a 254-bit modulus, canonicity chains, and the recursion format's field cells. Each is re-derived for the new field and specified as v1.0.0's were, with its own oracle and tamper twins.

Settlement#

Ethereum's verification precompiles today are pairing-based. How a post-quantum proving core settles on that chain, and what of the final step can rest on lattices, is part of the same programme of work, and will be specified with the same care as the core before it ships.

Back to the mission brief.

Quantum Leap

Signatures for Guests

Initiative QL-03. Zk-friendly and post-quantum signature verification available to every guest as a call, so authorization inside a blockchain-native application is one line of code.

View as Markdown
QL-03Signatures for guestsStatus: active development

Almost every blockchain-native application asks the same question on every request: did the right key authorize this? In v1.0.0 a guest answers it with code. secp256k1 recovery is k256 running over the delegated MOD_MUL and EC_ADD circuits, which makes Ethereum's own signatures affordable, and anything else is ordinary instructions. QL-03 makes signature verification a first-class guest operation.

The schemes#

A signature scheme is cheap to prove, or not, almost entirely because of its verification algorithm: what arithmetic it does, in which field, and which hash it calls. The signer never runs inside the proof. Four designs cover the space:

Scheme Idea Why it matters to a guest
Schnorr over a native curve Schnorr's protocol on a curve whose base field is the proof system's own field, as Grumpkin is to BN254 the cheapest by construction: arithmetic and hash both native to the circuit
ML-DSA (FIPS 204) Schnorr carried to lattices, with a short response kept uniform by rejection sampling NIST's primary post-quantum signature; cost dominated by its hash and, unless the key is fixed, its matrix expansion
FN-DSA (Falcon) hash-and-sign with a lattice trapdoor, hidden by Gaussian sampling the least arithmetic and the least hashing of the post-quantum three
SLH-DSA (FIPS 205) signatures from a hash function alone no algebra at all, and some two thousand hash calls; the most conservative assumption

All four verifiers are compute-and-compare, with no secrets and no branching on secrets, which is what makes each of them provable. The expository ZK-Friendly Signature Schemes works through each with a toy example and compares their in-circuit costs.

What decides the cost#

Two levers move every number:

  • The field the prover runs over. A scheme is native when its arithmetic is the circuit's arithmetic. Which scheme is cheapest therefore follows the re-fielding of QL-02, and the selection is made with it.
  • The hash. The post-quantum verifiers are dominated by their standard hash, not by their algebra. Replacing it with an arithmetic hash leaves the standard, and is the variant the zk-oriented constructions choose; keeping it is what interoperability with existing keys requires. Both have a place, and the guest decides.

For builders#

The aim is a guest that verifies an authorization the way it computes a hash today: one call, delegated to a circuit, no cryptography in the application's own code. That opens the patterns blockchain-native applications are built on: accounts whose keys are not the chain's, multi-party approval inside the state-transition function, session keys, and identities that stay valid after the curves they were born on have fallen.

Back to the mission brief.

Quantum Leap

The Deployment System

Initiative QL-04. The Apogee Blockchain-Native Deployment System, a portal and toolchain that take an application from source to a running blockchain-native environment.

View as Markdown
QL-04Apogee Blockchain-Native Deployment SystemStatus: active development

Version 1.0.0 gives a builder a repository, its tools and this manual. Everything between a proved guest and a live application, from keys and ceremonies to verifier contracts, bridges and operations, is still the builder's to assemble. The Apogee Blockchain-Native Deployment System is the second half of the leap: a portal and a toolchain that turn a guest program into a running blockchain-native environment, and keep it running.

The modules#

ConsoleActive

One place for every deployment resource

Programs and their identities, heights and parameters, verifying keys, decider keys and their ceremonies, verifier contracts and the networks they live on, organized per application and per release.

BridgesActive

Canonical on-chain contracts

Reusable templates for what every application rebuilds today: a state-root registry driven by proofs, deposit and withdrawal bridges, upgrade paths that move from one program identity to the next.

TelemetryActive

An application's vital signs

Proofs produced and settled, cycles per request, proving latency and cost, shard and family profiles, gas spent on verification, and the history of the application's state roots.

GatewayActive

A door for your own AI

An interface through which a developer's local AI agent can inspect an application, query its telemetry, propose and run changes through the deployment pipeline, and read every result back, under the developer's control.

TemplatesPlanned

Applications that start from a working shape

Guest projects outside the repository, with the profiles, linker settings and vendored crates already right, and the AI Companion already in place.

CeremoniesPlanned

Ceremonies as a service, not a chore

Coordination of the decider key's phase-2 contributions, each verifiable against the circuit and the ceremony file, so that a deployment's key has the honest contributors its soundness needs.

Why a system, and not more tools#

The thesis behind Apogee is one optimized environment per economic application. That multiplies the number of environments, and with it the operational surface: every one has its program, its keys, its ceremony, its contracts and its metrics. An approach that asks each team to assemble that surface by hand does not scale to the many environments the thesis needs. The Deployment System makes each one routine, so the hard part of launching a blockchain-native application is the application.

The Gateway follows from the same reasoning. Builders already work alongside AI models, and the AI Companion briefs those models on writing guests. The Gateway gives them, under the developer's control, a way to act on what they write: deploy, observe, and iterate.

Back to the mission brief.

Auditors

Audit Guide

Everything an auditor of Apogee VM v1.0.0 needs to start: the scope, the normative specification and how it is organized, the notation, the trust boundary, a reading order, and the properties most worth checking first.

View as Markdown

This section is the complete construction of Apogee VM v1.0.0, organized for evaluation. Its core is the normative specification, reproduced verbatim: one page per subject, with every committed column by index and name, every gate as a polynomial, every lookup and its channel, every transcript message in order, and every byte of every wire form. Around it, this guide and the soundness map give an auditor a way in.

Scope#

In scope Where
The statement a proof establishes, and what a verifier must hold System, end to end, The proof
The arithmetic: Fr, the Fq tower, G1 and G2, the pairing, MSM, multilinear polynomials, the sumcheck Primitives
Fiat–Shamir: Poseidon2, the duplex sponge, every tag Transcript
The setup and the commitment scheme Structured reference string, Mercury
The program: loading, decoding, tables, configuration, identity Program and identity, Guest ABI
The execution model and the trace Execution trace, Public values and advice
The proof system: GKR, the circuit registry, the memory argument, lookups GKR engine, Circuits, Memory argument, Lookups
Every circuit family, column by column the seven instruction families and the six delegation circuits
The prover's structure Streaming prover
Recursion, the Groth16 decider and its ceremony, the contract Recursion and decider
The Ethereum workload Ethereum blocks

The specification pages are reproduced from the Apogee VM repository's docs/ at source revision 3571370, with relative links turned into links on this site and every § reference made a link to its section. One phrase in the glossary is reworded to match the rest of this site; nothing else is changed. Where a specification page and the code disagree, the code is right, and the disagreement is a finding.

How the specification reads#

The pages are written to be read against the code. Each names the crate and the function that implements what it states, and source comments cite the specification back by section (docs/spec/memory.md §2.4). Some conventions that recur:

Notation Meaning
M[i], W[i], S[i] committed columns of a circuit: memory columns (bound in the global transcript, before the memory challenges), witness columns (bound in the shard's own transcript), setup columns (bound by the program identity or the SRS digest)
V[…] a virtual table: a closed form of the row index, never committed
L{k}[j], C{k}[j], scratch[i] inner column j of layer k; a cached entry; a flat relation's intermediate
W[8..14] a half-open range of column indices, W[8] to W[13]
T(AS, ADDR, TS, VAL) a memory tuple, γ_M + AS + α_addr·ADDR + α_ts·TS + α_val·VAL
4c + Δ the timestamp of slot Δ of cycle c
G1–G11, S1–S6 the steps of the global and the shard transcript
steps 1–12, B1–B6 the verifier's checks, in order, for a shard and a block
2^n a power of two; heights are 2^8, 2^12, 2^16, 2^18, 2^20, 2^22

A gate written as an expression is held to 0. A lookup is written as its channel, selector and tuple. "Bounded" means range-checked, and a value called a word is an integer in [0, 2^32).

Reading order#

For a first pass that builds the whole argument before descending into circuits:

  1. System, end to end. The claim, the composition table, the assumptions, the limits.
  2. The proof. The statement, both transcripts, the verification order, the key and its loading rules.
  3. GKR engine. The layer model, the artifact and its laws, the backward pass and why it is sound.
  4. Memory argument and Lookups. The two arguments everything crossing a row rests on, with their construction-time rules.
  5. Circuits. The registry, the shapes, and how a family circuit is assembled.
  6. The instruction families, starting with ADD_SUB_LUI_AUIPC, which carries every ecall and the request side of every delegation.
  7. Delegation ABI and Delegation circuits.
  8. Program and identity, Public values, Execution trace.
  9. Transcript, SRS, Mercury, Primitives.
  10. Recursion and decider, then the contract.

The trust boundary#

Soundness is the verifier's alone, and the verifier's code is a defined set of crates:

Trusted for Crates
Verifying a block constants, field, curve, transcript, poly, sumcheck, pcs-verify, pcs, gkr-verify, verifier-core, verifier, and constraints, because the circuits are part of the statement and a missing gate is a soundness bug
Computing an identity from an ELF loader, isa, program
The last step to the chain guests/recursion, groth16, the decider's circuit, contracts/ApogeeVerifier.sol
Untrusted prover, emulator, trace, the proving half of host: the prover validates nothing

The assumptions are knowledge soundness of Mercury and KZG in the algebraic group model under q-DLOG, Poseidon2 as a random oracle, Groth16's own assumptions for the last step, and one honest contributor to each ceremony. Nothing is constant-time and no proof is zero-knowledge. A verifier must obtain the program identity and the ceremony's SRS digest from a channel the prover does not control.

Where to look first#

These are the properties whose failure would be a forgery, with where each is argued. The soundness map carries every claim of the statement the same way.

Property Argued in
Every challenge is drawn after everything it protects: the memory challenges after every M commitment, the window list, io_digest and the 64 boundary scalars; g and β after the shard's W commitments proof §2, §4; memory §6.1
No memory tuple or root reads a W column, which is committed after the memory challenges memory §8
A frame holds its masks only to booleanity; each family pins every mask to m_pc times the kinds that make the query memory §2.1; each family page
Every key a table channel looks up is bounded by its family, and every bound written through a copower also carries a direct bound lookup §4, §11; shift-bitwise §3
A selector decoded from a frame word is one-hot, since codes add delegation circuits §1
Each delegation request pairs with exactly one invocation delegation §5
Only the exit row can write HALT_PC; next_pc is held even where a family computes it memory §5; jump-branch-slt §5
One initial value per address: the window rules memory §3.5, §9
The input and journal windows hold the statement's bytes; the journal has no initial column public values §5
The opening takes setup commitments from the key, binding the tables and image the identity commits proof §5; memory §6.2
A key's circuits are the registry's, and its SRS digest is compared with the ceremony's proof §3, §7
Recursion: tapes bound by program identity, the transcript chain, fold weights drawn after what they weight, the decider's bound wires, and the order of the ceremony's rounds recursion §7, §8, §9

Reproducing#

Everything CI runs needs no ceremony file. The suites that prove real shards run over a toy SRS of their own and need tens of GiB, so they are run by name:

sh
cargo test --workspace                                   # every unit, law and row suite
cargo run -p kat-gen && git diff --exit-code             # fixtures regenerate identically
cargo test --release -p prover --test <suite> -- --include-ignored --test-threads=1
#   acceptance, control, alu, mem, fills, block, streaming, keccak, recursion, public_io, revm
cargo test --release -p checker --test tamper -- --include-ignored --test-threads=1   # every tamper twin
cargo run -p checker -- laws <artifact>                  # Laws 1–4 and the lookup rules, independently

Verifying the implementation describes what each oracle and suite establishes, and what none of them does.

Reporting#

Report findings to admin@gweb3networks.com, with the specification section or the code path, the property at stake, and where possible a tamper twin: a forged witness, proved as an honest prover would prove it, that verifies.

Auditors

Soundness Map

Every claim a verified proof makes, the argument that carries it, and the exact specification sections where that argument is stated and justified.

View as Markdown

A verified block establishes one sentence: the program of this identity, started at its entry pc over its image, with this public input and some advice, executes instruction by instruction to EXIT with this status, having written this journal. This page breaks that sentence into the claims it is made of and carries each one to the argument that proves it.

The program#

Claim Argument Specification
The key describes the registered program loading a key recomputes the identity from its own config, entry pc and setup commitments; the verifier compares it with its own copy proof §7.2, program §8
The tables a proof reads are the ones the identity commits every shard's batched opening takes its setup commitments from the key proof §5
Each executed row is the program's instruction at its pc the decoder lookup, keyed by the row's own pc read, into a table whose live rows are one-hot and whose padding rows are −1 lookup §10, program §5, §6
Memory starts from the program's image INIT_TEARDOWN's init column is a setup column the identity commits; no file-backed byte lies outside window 0 memory §6.2, §3.4
Execution starts at the entry pc the pc's initial tuple uses the key's entry pc, which the identity binds memory §4.2, §6.2
The circuits are the right ones a key's circuits must equal the verifier's registry at its heights and pass the laws, the memory rules and the discharge rule proof §7.2, circuits §1, gkr §4.2

Each row#

Claim Argument Specification
A row obeys its instruction the family's enforcing gates, zero on every row, each family's soundness argument add-sub §4, jump-branch-slt §5, shift-bitwise §5, mul-div §5, memory-ops §3–§6
A row makes exactly its instruction's queries each mask pinned to m_pc times the kinds that make that query memory §2.1
x0 reads and writes 0 the x0 gadget and write-backs memory §2.4
Register and RAM values are words every register write bounded on its own row; every RAM write of an execution family a word; initial values words except advice, which no family relies on memory-ops §5
A padding row adds no memory event every mask of a padding row is 0, so its leaves are 1 memory §2.3, gkr §4.3

Memory and order#

Claim Argument Specification
Every read returns the last write one read/write multiset over every shard, reconciled once against the register and pc boundary memory §4, §9
A read strictly follows the write it consumes each query's timestamp gap is two 19-bit TIMESTAMP chunks memory §2.4, §7
Every address has exactly one initial value the window rules: one height, disjoint windows, one shard each of the fixed windows memory §3.5, §9
The multiset cannot close by a loop timestamps are integers on bounded paths: a loop would need more than 2^215 edges memory §4.2
The rows form one path from the entry to the exit, in program order the pc is a memory cell written at least four timestamps after it is read memory §5, §9
The execution ends at the exit row HALT_PC = 1 is odd; only the exit row writes it; JUMP_BRANCH_SLT range-checks its next_pc even memory §5, jump-branch-slt §5, add-sub §4
A shard's time window adds nothing time windows are checked for shape only; order comes from the multiset alone proof §8

Values and lookups#

Claim Argument Specification
Every gated tuple is a row of its table one LogUp identity per channel, summed by a fraction tree in the GKR pass; root numerator 0 and denominator nonzero lookup §1, §6, §8
Selectors are boolean every lookup's selector is held to s − s² by an enforcing gate of list 0 lookup §2
A lookup answers from its own sub-table one width per channel, disjoint key ranges with the +1 offset, and each family's bound on its key lookup §4, §9, §11
The tables are the intended ones virtual tables are the verifier's closed forms; the generic table is bound by the SRS digest; decoded tables by the identity lookup §3, §12
Every declared obligation is discharged the discharge rule at assembly and at every key load lookup §11

Public values#

Claim Argument Specification
The input window held the statement's input PUBLIC_INPUT's initial column equals the input's words at a random point, after G7 fixes the bytes public values §5
The journal is what the guest's stores left PUBLIC_OUTPUT's final column equals the journal's words; the window has no initial column a prover could fill public values §5
The exit status is x10's final value the exit row writes back the a0 it read; the verifier holds v_10 to the statement's status add-sub §4, memory §4.1, proof §6

Delegations#

Claim Argument Specification
Each request is executed exactly once the anchor: requests and invocations pair one to one through the multiset in the type's own space delegation §5
An invocation computes its function each circuit's soundness argument, frame and canonicity chains included delegation circuits §2–§7
A multi-call operation is the composition of its steps RAM glue: each step reads the previous step's writes on one memory history; the order is the calling code's delegation circuits §1

The proof system#

Claim Argument Specification
A shard's outputs are its circuit evaluated on its committed columns the GKR backward pass, each challenge drawn after what it protects gkr §5.4
The claimed column values are the committed polynomials' one batched Mercury opening at the pass's point mercury §5, §7
A statement is verified only by all of its shards shard-set exactness at decode and in verify_block; the reconciliation reads every shard's roots proof §1.3, §6
Challenges follow every commitment they protect the global transcript G1–G11 and the shard transcript S1–S6 proof §2, §4, transcript §3
The setup is the ceremony's the SRS digest, compared by the verifier with the ceremony's proof §3, srs §3

Recursion and the contract#

Claim Argument Specification
A node ran exactly the base verifier's checks the checks are compiled into tapes in the node program's image, which its identity binds recursion §7, §8.1
The tree covers one base statement, every shard, in order the transcript chain across nodes, adjacent shard ranges, and each node's checks on its children recursion §8.1, §8.2
Every deferred opening holds each folded under weights drawn after everything it weights, discharged by one pairing in the contract recursion §8.3, mercury §6
The decider binds what the contract holds bound wires committed before their challenge; the circuit holds the journal to the whole base range recursion §9
The decider key has no known trapdoor a two-phase ceremony with one honest contributor per round, rounds in order recursion §9, srs §7

Deliberately not claimed#

  • Anything about advice. Advice is bound to nothing by design; a guest checks it.
  • Zero knowledge. Nothing is blinded.
  • sc.w failure semantics. sc.w always succeeds; a program relying on its failure is outside the claim.
  • Traps. A run that traps has no proof at all.
  • That the prover is correct. The prover is untrusted; only the verifier's crates carry soundness.
  • That a ceremony file is the ceremony's, or that a key's τ is unknown, without the verifier's own comparison of the SRS digest.

Auditors

Verifying the Implementation

How the code is checked against something other than itself. Each layer's independent oracle, the second implementation of the circuit rules, the tamper twins that prove forgeries are refused, and what no check covers.

View as Markdown

No component of Apogee is checked against a second implementation of the whole system. Instead each layer has an oracle of its own, chosen so that the check shares as little code as possible with what it checks.

Each layer and its oracle#

Layer Checked against
Fields, curve, pairing, MSM known-answer vectors generated from arkworks, which the tests also run live; the tests re-derive every arithmetic constant the crates read
Poseidon2 and the transcript tools/transcript-ref: Plonky3's Poseidon2 keyed with zkhash's round constants, and a transcription of the specification run beside Plonky3's duplex challenger, agreeing on every squeeze
The decoder all 2^30 32-bit words with low bits 11, against accepted counts derived from the ISA's tables and against an independent encoder; llvm-objdump over the committed guests
RVC expansion LLVM's own encoder, over a guest assembled both compressed and not
Circuits as data checker: the four laws, the lookup rules and the padding contract re-implemented without constraints' code, sharing only the gate kernel
Each family's gates row suites that build rows with Rust's own integer arithmetic and evaluate them through the checker; the arithmetic cores exhaustively at reduced word widths
The memory and lookup arguments native evaluators in checker, run over executed traces
The executor its own trace's self-check and the arguments above; there is no second executor
The revm guest native revm, built from unpatched upstream crates
The stateless validator a committed subset of tests-zkevm v21.0.1 natively in CI; the whole release natively and the subset through the guest binary by hand; tools/stateless-ref for the input encoding
The decider the proof checked natively, and the contract executed in revm

Exhaustive checks at small widths#

Several arithmetic cores are written with their word width as a parameter, so that the encoding can be checked over every input at a width small enough to enumerate:

  • the comparison gadget at 6 bits, over every operand pair, signed and unsigned, finding exactly one (lt, gap), the ISA's;
  • MUL_DIV's arithmetic at 4 bits, over every dividend and divisor and each division kind, admitting exactly one (q, r), RV32M's;
  • MEM_SUBWORD's splice at a 4-bit word, admitting exactly one (high, sub, low) for every word, offset and width.

The circuit rules, twice#

CircuitArtifact::validate and the memory and lookup construction rules run wherever an artifact is built or a key is loaded. crates/checker enforces the same rules a second time with code of its own, never calling validate, and evaluates gates only through gkr_verify::eval_gate, the one kernel both sides treat as the semantic authority. Its validators check the laws by evaluation at sampled points where validate compares normalized expansions, recompute every channel's sum row by row rather than by a tree and name any tuple no table row holds, and rebuild an execution family's memory columns from the event log rather than from a shard's rows.

sh
cargo run -p checker -- laws <artifact>       # Laws 1–4, then the lookup rules
cargo run -p checker -- padding <artifact>    # the padding contract
cargo run -p checker -- dump <artifact>       # the circuit, readably

Tamper twins#

A tamper twin is a forgery proved exactly as an honest prover would prove it. The tamper suite (checker::TamperHarness) re-proves a statement with witness cells or boundary scalars changed: each channel's multiplicities recounted, changed memory columns recommitted in a fresh global commit phase, every shard re-proved. Then it verifies a shard or the block and asserts the refusal's class, Constraint, Lookup with its channel, or MemoryArgument, or asserts that a change breaking nothing verifies.

The twins rely on the prover checking nothing, which is the design: a forged witness gets the best proof an honest prover could make of it, and the verifier must refuse it in the expected class. The suite also carries the delegation anchor's forgeries, and on the mainnet mini-block it shows the other side of the advice rule: a corrupted advice cell is refused by the memory argument, and a consistently corrupted one verifies, because advice is bound to nothing.

sh
cargo test --release -p checker --test tamper -- --include-ignored --test-threads=1

Fixtures that regenerate#

kat-gen writes every committed known-answer vector, listing, circuit artifact and identity, each with its SHA-256, which the tests reading it pin. CI regenerates the default groups and both reference oracles and fails on any difference in the vector directories:

sh
cargo run -p kat-gen && git diff --exit-code

A guest ELF is not reproducible across machines, because rustc embeds absolute paths in panic-location strings; two clean builds on one machine agree. So the guest ELFs are regenerated by hand on one machine, and CI regenerates only what derives from them.

What no check covers#

  • There is no second executor. The emulator is held to a restatement of its own frame table and to the memory and lookup arguments, not to an independent RISC-V implementation, and no executor here takes a delegation shim's software fallback.
  • The construction rules the checker does not repeat: the memory construction rules, the copower rule, and validate's remaining construction rules, the degree ceiling among them, are enforced once.
  • The prover is not checked for correctness, only for completeness through the suites that prove real shards, which run outside CI because each needs tens of GiB.
  • Osaka-family stateless inputs have no end-to-end oracle. The release fills only Amsterdam; the Electra/Fulu layout is held to eth-act/ere-guests and the header rules to two mainnet blocks.

Auditors/System

The system, end to end

Normative specificationdocs/architecture.mdView as Markdown

Apogee proves executions of RV32IMAC programs. This page is the system end to end: what a proof states, how one is made and checked, what it assumes and where it stops. Each paragraph names the page that specifies its subject; glossary.md indexes the vocabulary.

1 What a proof states#

A verifier holds three things it does not take from the prover's word, and two of them from a channel the prover does not control (proof.md §1, §3):

  • the program identity, one field element: a digest of the program's instruction tables, its initial memory image, its entry pc and its VmConfig — the circuit families it uses and their heights (program.md §8);
  • the SRS digest of the ceremony, which a verifying key must carry;
  • a verifying key: that config, each family's circuit and setup commitments, the SRS's verifier points and the generic lookup table's commitments. Loading it recomputes the identity and the SRS digest from its own contents and holds its circuits to the registry's bytes (proof.md §7).

The proof's statement, PublicInputs, carries the public input bytes, the public output bytes (the journal), the exit status, and the record of the execution's shape — shard counts, memory windows, the final registers and pc, every shard's memory commitments and roots (proof.md §1).

A proof that verifies establishes that the program of that identity, started at its entry pc over its image, with the public input in its input window and some advice of the prover's choosing, executes instruction by instruction to EXIT with that status, having written that journal. Nothing is claimed of the advice, and nothing is hidden: no commitment or proof is blinded.

2 From a binary to a proof#

  1. The program. loader reads the ELF into a ProgramImage, expanding compressed instructions in place; isa decodes; program routes each instruction to one of seven instruction families, builds every family's decoded table — a row per halfword of code — and commits to them as the identity (program.md).
  2. Execution. emulator runs the guest. A cycle is one row of the family that owns its instruction, recording its memory queries: timestamped reads and writes of the pc, registers and RAM (execution-trace.md). A guest issues no system call but EXIT: its input, journal and advice are regions of memory (public-values.md, ecall-abi.md). Hashing and big-integer arithmetic are delegated: an ecall names a frame in RAM, and a row of a delegation family does the work on it (delegation.md, delegation-circuits.md).
  3. Shards. A family's rows are cut into shards of the family's height, a power of two between 2^8 and 2^22. The memory an execution touches is covered by shards of the window families, which give each word its initial and final tuple (memory.md §3). A shard is the unit of proving; a block is hundreds (circuits.md §1).
  4. A shard's proof. Its columns are committed with Mercury (mercury.md). The family's GKR circuit is run backward from its outputs to those columns, a sumcheck a layer (gkr.md), and every column is opened at the one point that pass ends on, in one batched opening (proof.md §5).
  5. The block. A BlockProof is the statement and its shard proofs. verify_block runs the global transcript once, each shard's checks, and once the memory reconciliation over every shard's roots (proof.md §6).
  6. Recursion. Verifier programs, proved by this VM in a format of its own, verify runs of shards and fold their deferred pairings; a tree of them ends in a root, a Groth16 circuit re-verifies the root, and ApogeeVerifier.sol checks that proof and the folded pairing (recursion.md).

The prover executes the guest twice: once to commit every shard's memory columns, which fixes the statement and its challenges, and once to prove each shard as it fills. Its memory is bounded by the shards in flight, not by the execution (streaming.md).

3 How soundness composes#

Each shard's GKR pass and opening tie its circuit's outputs to committed columns. On top of that, these arguments span the execution:

claim carried by
every row obeys its instruction the family circuit's enforcing gates, zero on every row the family pages, circuits.md
a row's instruction is the program's at its pc a lookup of the row's pc and fields in the family's decoded table, which the identity commits lookup.md §10
every read returns the last write one multiset over all shards: an access reads a tuple (space, address, timestamp, value) and writes one with a later timestamp; the verifier multiplies every shard's read and write roots against boundary factors for the registers and the pc memory.md
the rows are one path from the entry pc to the exit, in program order the pc is a cell of that multiset: a row reads its pc and writes the next one at least four timestamps later, so shard order, cycle uniqueness and continuity need no other argument memory.md §5, §9
a value is a byte, a word, a sign, an XOR LogUp channels over range, byte and generic tables lookup.md
the public input and the journal are the claimed bytes the two public windows' initial and final columns, held to the bytes' multilinear extensions public-values.md §5
a delegated computation is the function's invocation rows that read and write the frame through the same multiset, paired one to one with their ecall by an anchor tuple delegation.md §5

Challenges come from a Poseidon2 duplex transcript (transcript.md). The global transcript absorbs the whole statement, every shard's memory commitments included, before the memory challenges exist; each shard's transcript is seeded from its final state (proof.md §2, §4).

4 What it assumes#

  • Cryptography. Mercury's and KZG's knowledge soundness in the algebraic group model under q-DLOG (mercury.md §7); Poseidon2 as a random oracle for Fiat–Shamir; for the last step, Groth16's own assumptions. BN254 gives about 100 bits.
  • Setup. The SRS is the PSE perpetual powers of tau, sound while one contributor was honest. The code checks a file's structure and decodes every point; nothing proves it is that ceremony's, and no proving path runs Srs::validate (srs.md §3). The decider's Groth16 key comes from a second, circuit-specific ceremony (recursion.md §9).
  • What a verifier must obtain itself. The program identity and the ceremony's SRS digest. A key loads under whatever digest its own points give, so a key built over a known τ is refused only by that comparison; the verifier CLI compares identity only, and host::verify neither (proof.md §1, §3).
  • Trusted code. Soundness is the verifier's alone: constants, field, curve, transcript, poly, sumcheck, pcs-verify, pcs, gkr-verify, verifier-core, verifier, and constraints — the circuits are part of the statement, and a missing gate is a soundness bug. Computing an identity from an ELF trusts loader, isa and program. The last step adds guests/recursion, groth16, the decider's circuit and the contract. prover, emulator, trace and the proving half of host are untrusted: the prover validates nothing, and a wrong input costs an honest prover a proof that fails.
  • Nothing is constant-time (primitives.md). No proof is zero-knowledge, so proving keeps nothing secret; the one secret the code handles is a Groth16 ceremony contributor's factor, which groth16::phase2 multiplies in with the same variable-time ladder.

5 Limits#

not zero-knowledge no blinding in Mercury, GKR or the Groth16 decider
advice is unbound a guest checks it against something a proof binds (public-values.md §6)
public values at most 16,380 bytes each of input and journal (public-values.md §9)
sc.w always succeeds the one deviation from RV32IMAC's semantics; there is no reservation state (memory-ops.md §6)
traps are not provable a misaligned access, an access outside mapped memory, ebreak or a pc with no instruction ends an execution with no proof (execution-trace.md §10)
code is static the instruction stream is the image decoded at load; one undecodable word in an executable segment refuses the program (program.md)
code size .text within a decoded table's reach of its load address, 7.94 MiB at 2^22, and the image within bytecode_size_words, 4 MiB by default (program.md §5, §7)
execution length timestamps are 38 bits: 2^36 − 1 cycles (execution-trace.md §1)
delegations are a fixed set six in the base format; an EVM MULMOD with an arbitrary modulus is not one; a delegation proves one step of its function, and composing steps, validating curve points among them, is the calling code's (delegation.md §11, delegation-circuits.md)
prover memory set by the shards in flight: the measured full block peaked at 174 GiB (streaming.md §1)
block witnesses the stateless validator takes its input from an external witness producer; the built-in recorder cannot record every block (ethereum.md §4, §6)
the decider's key one per root shape, and only as trustworthy as its ceremony; the development key is forgeable (recursion.md §9)
on-chain cost about 3.6M gas for the measured block (recursion.md §10)

6 How the code is checked#

No component is checked against a second implementation of the whole system; each layer has its own independent oracle.

layer checked against
fields, curve, pairing, MSM known-answer vectors generated from arkworks, which the tests also run live
Poseidon2 and the transcript tools/transcript-ref: Plonky3 and zkhash
the decoder every 32-bit word of the instruction space against counts from the ISA; llvm-objdump over the committed guests
circuits as data checker: the circuit laws, the lookup rules and the padding contract re-implemented without constraints' code, sharing only the gate kernel (circuits.md §3)
each family's gates row suites that build rows with Rust's own integer arithmetic and evaluate them through the checker; the arithmetic cores exhaustively at reduced word widths; tamper twins, a forged witness proved as an honest prover would and refused in the expected class
the memory and lookup arguments native evaluators in checker over executed traces
the executor its own trace's self-check and the arguments above; there is no second executor, and no executor here takes a delegation shim's software fallback
the revm guest native revm, built from unpatched upstream crates
the stateless validator a committed subset of tests-zkevm v21.0.1 natively in CI; the whole release natively and the subset through the guest binary by hand; tools/stateless-ref for the input encoding
the decider the proof checked natively, and the contract executed in revm

Committed fixtures are regenerated and compared in CI (tools.md §7). The suites that prove real shards, over a toy SRS, need tens of GiB and run outside CI (README).

7 Cost#

recursion.md §10 has the end-to-end measurements for one block, from the base proof to the contract call; streaming.md §1 breaks the base proof down; and circuits.md §1 gives every circuit's width and proof size, which a shard's cost follows.

References#

  • L. Eagen, A. Gabizon. MERCURY: A multilinear polynomial commitment scheme with constant proof size and linear field work. ePrint 2025/385. publication/2025-385.pdf
  • D. Boneh, J. Drake, B. Fisch, A. Gabizon. Efficient polynomial commitment schemes for multiple points and polynomials. ePrint 2020/081. publication/2020-081.pdf
  • J.-L. Beuchat et al. High-speed software implementation of the optimal ate pairing over Barreto–Naehrig curves. ePrint 2010/354. publication/2010-354.pdf

Auditors/Foundations

Primitives: fields, curve, pairing, polynomials, sumcheck

Normative specificationdocs/spec/primitives.mdView as Markdown

BN254's scalar field Fr, its base field Fq and the tower to Fq12, the groups G1 and G2, the optimal ate pairing, multi-scalar multiplication, multilinear polynomials and the zerocheck. The byte encodings of field elements and points (§1–§3) and the polynomial index convention (§6) are defined here.

  • All of it is this repository's code: concrete types, no field trait, no unsafe, assembly or intrinsics. field, poly and sumcheck are #![no_std] and build for the guest target; curve is std, with rayon, and no guest links it.
  • arkworks is a test oracle only, for field, curve and poly, live and through vectors tools/kat-gen generates; the tests of field and curve re-derive every arithmetic constant those crates read.
  • Nothing is constant-time: reductions, exponentiations, point additions and scalar ladders branch on their operands. No proof is zero-knowledge, so no witness is secret; the one secret this code handles, a decider ceremony contributor's factor (recursion.md §9), goes through the same variable-time ladder (§3).

1 Fr#

p = 21888242871839275222246405745257275088548364400416034343698204186575808495617

field::Fr is the integers mod p (constants::FR_MODULUS), 254 bits; p is also the order of G1 and G2, r in §3–§4. In memory an element is four little-endian 64-bit limbs of x·R mod p, R = 2^256 mod p, always reduced below p, so equal limbs are equal values. Multiplication is CIOS Montgomery over u128 intermediates. Fr::inverse is x^(p−2), None at 0; field::batch_inverse is Montgomery's trick and leaves a 0 entry 0. p − 1 = 2^28·c with c odd, and every FFT domain is a subgroup of the one constants::FR_TWO_ADIC_ROOT_OF_UNITY generates.

byte form used in
wire the value, not x·R, as 32 little-endian bytes: Fr::to_bytes. Fr::from_bytes is None for a value ≥ p and never reduces; serde goes through both every proof, key and artifact
source literal 0x and exactly 64 lowercase hex digits, big-endian: Fr::from_hex, None for any other spelling or a value ≥ p constants in crates/constants
memory the four limbs: Fr::to_memory_bytes, Fr::from_memory_bytes, None at or above p the FR_ARITH delegation's frame alone

On the guest target (cfg(target_arch = "riscv32")) addition, Montgomery multiplication and inversion call the FR_ARITH delegation through guest_sdk::recursion::fr_arith (delegation.md §10). Its circuit proves these three functions of the memory form the frame carries (delegation-circuits.md §4), so a delegated result is the software result bit for bit.

2 The Fq tower#

q    = 21888242871839275222246405745257275088696311157297823662689037894645226208583
Fq2  = Fq[u]/(u^2 + 1)
Fq6  = Fq2[v]/(v^3 − ξ)       ξ = 9 + u
Fq12 = Fq6[w]/(w^2 − v)

curve::Fq is the field of coordinates (constants::FQ_MODULUS). Its limb arithmetic is Fr's, copied literally over q's constants, and so are its wire and source-literal forms. Fq2 encodes as c0 ‖ c1; nothing above it has a byte form.

Products are schoolbook and squarings above Fq2 are products. A Frobenius map multiplies coefficients by powers of ξ tabulated in constants (FQ6_FROBENIUS_C1, FQ6_FROBENIUS_C2, FQ12_FROBENIUS_C1); Fq12::conjugate is the q^6 one. Nothing in the tower or the pairing is sparse, cyclotomic or precomputed: a pairing is only ever computed to verify something, and the code is written to be read.

3 G1, G2 and their encodings#

G1 = E(Fq)                E:   y^2 = x^3 + 3        #E  = r             generator (1, 2)
G2 ⊂ E′(Fq2), order r     E′:  y^2 = x^3 + 3/ξ      #E′ = r·(2q − r)    generator EIP-197's

G1Affine { x, y, infinity } is a point and G1Projective its Jacobian form, Z = 0 the identity, under the EFD formulas dbl-2009-l, add-2007-bl and madd-2007-bl, with the identity, P = Q and P = −Q branched on explicitly. Scalar multiplication is a fixed 4-bit window. G2Affine and G2Projective are the same code over Fq2.

A point's wire form is uncompressed affine, and there is no compressed one:

G1Affine    64 bytes    x ‖ y
G2Affine   128 bytes    x.c0 ‖ x.c1 ‖ y.c0 ‖ y.c1        x = x.c0 + x.c1·u
infinity                every byte zero

Each coordinate is an Fq in wire form; (0, 0) is on neither curve, so zero is unambiguous. G1Affine::from_bytes and G2Affine::from_bytes return None unless the bytes are all zero, or every coordinate is below q, the point satisfies its curve's equation and, in G2, whose cofactor 2q − r is not 1, [r]P is the identity — by the window ladder, with no endomorphism.

Nothing else validates: the affine structs' fields are public, and the group law, msm and the pairing compute on whatever they are given. A point's transcript form is transcript.md §4's; a .ptau file (srs.md §2) and the contract (recursion.md §9) have encodings of their own.

4 The pairing#

e(P, Q) = f_{6x+2, Q}(P)^((q^12 − 1)/r)        x = 4965661367192848881

curve::pairing::miller_loop(pairs) is Algorithm 1 of Beuchat et al. (ePrint 2010/354) with homogeneous projective line formulas: the 66-digit NAF of 6x + 2 (constants::ATE_LOOP_NAF), then the two lines adding ψ(Q) and −ψ(ψ(Q)), ψ the untwist-Frobenius-twist map. For a Q of order r no step adds a point to itself, to its negative or to the identity, so the line formulas have no exceptional case.

final_exponentiation returns exactly f^((q^12 − 1)/r): the easy part (q^6 − 1)(q^2 + 1), then the hard exponent (q^4 − q^2 + 1)/r as its base-q expansion λ0 + λ1·q + λ2·q^2 + q^3, by three exponentiations (constants::FINAL_EXP_LAMBDA_0 to FINAL_EXP_LAMBDA_2, the two negative ones conjugated) and three Frobenius maps. The Fuentes-Castañeda hard part, which arkworks uses, returns this value raised to 2x(6x^2 + 3x + 1): the two libraries agree on every pairing check and on no pairing value but 1, and the test vectors are arkworks' Miller outputs raised to the literal exponent.

pairing_check(pairs) is Π e(P_i, Q_i) = 1 by one Miller loop, whose Fq12 squarings the pairs share, and one final exponentiation: the form of every pairing equation in the system. A pair holding a point at infinity contributes 1 and is skipped; an empty product is 1.

5 MSM#

curve::msm::msm(bases, scalars) is Σ scalars_i·bases_i in G1 by windowed Pippenger, and msm_small_u32 the same sum over u32 scalars, recoded from 32 bits instead of 254. Neither looks at a scalar's size: the caller chooses, and pcs::commit chooses by a column's backing (§6). G2 has no MSM here; crates/groth16 carries its own.

  • Width. w = 3 below 32 points, otherwise ⌊0.69·⌈log2 n⌉⌋ + 2: arkworks' rule.
  • Digits. A scalar is recoded into signed digits in [−2^(w−1), 2^(w−1)], one a window, over ⌈(bits + 1)/w⌉ windows; the sign costs a negated base and halves the buckets to 2^(w−1). At 2^20 points w = 15: 17 windows for an Fr, 3 for a u32.
  • Parallelism. Each (window, chunk of the input) is a rayon task that adds bases into buckets by mixed addition and reduces them by a running sum; the tasks' sums are combined serially, w doublings a window. Group sums are exact, so the point does not depend on the thread count.

6 Multilinear polynomials#

poly::MultilinearPoly is a table of 2^n evaluations over {0,1}^n, the type of every column.

Index convention. Variable j is bit j of the index: the evaluation at y = (y_0, …, y_{n−1}) is entry Σ_j y_j·2^j. bind(r) fixes variable 0, the low bit,

f′(i) = f(2i) + r·(f(2i + 1) − f(2i))

and the old variable 1 becomes variable 0. So binding r_0, r_1, … in order leaves evaluate(&[r_0, r_1, …]), whose point[j] is variable j. Sumcheck round i binds variable i (§7), so a claim's point lists its challenges in variable order, the order evaluate and a Mercury opening (mercury.md §1) take.

Backing. PolyBacking holds the table as a bitset (U1), u8, u16, u32 or Fr. A trace column is filled and committed at its integer width (§5). get and evaluate embed an entry in Fr as they read it and leave the table alone; the first bind folds the integer table straight into an Fr table of half the length, and the backing is Fr from then on.

eq. eq_table(r) tabulates eq(r, ·) over the cube in the same index order; eq_eval(r, y) is its closed form, for any r and y.

new, get, bind, evaluate and eq_eval panic on a table, index or point of the wrong size rather than return an error.

7 The sumcheck#

crates/sumcheck proves that a gate vanishes on the cube. A sumcheck::Gate is a sum of GateTerms coef·x_a·x_b, the second factor optional, over input columns of n variables: degree at most 2 in each variable, by construction.

G(y) = 0 on all of {0,1}^n is proved as the sumcheck 0 = Σ_y eq(r, y)·G(y) at a random r. eq·G has degree at most 3 in each variable, so a round polynomial is a cubic and a round message its four coefficients [c0, c1, c2, c3], ascending — four whatever the gate, so a proof's shape depends on n and the number of inputs alone.

prove_zerocheck and verify_zerocheck run one schedule, under the tags of transcript.md §5:

1   the caller binds the columns to the transcript
2   r_0 … r_{n−1}                                      n × SUMCHECK_CHALLENGE
3   for i in 0..n:   g_i, one message of four          SUMCHECK_ROUND
                     ρ_i, binding variable i           SUMCHECK_CHALLENGE
4   final_evals: each input column at ρ, one message   SUMCHECK_FINAL_EVALS

The verifier checks the proof's shape, then g_0(0) + g_0(1) = 0 and g_i(0) + g_i(1) = g_{i−1}(ρ_{i−1}), each before absorbing g_i, and, with final_evals absorbed, g_{n−1}(ρ_{n−1}) = eq(r, ρ)·G(final_evals). It returns SumcheckClaim { point: ρ, final_evals } or a SumcheckError, and does not panic on a proof.

That last check is one equation over all the claimed evaluations and ties none of them to its column: the caller owes an opening of each at ρ, as it owes step 1. The step 1 its callers use is witness_digest, a hash and not a commitment: a sponge of its own absorbs [column count, n] and each column's cells under WITNESS_DIGEST, and its raw squeeze enters the transcript under the same tag.

What uses it. No proof in the system is this zerocheck, and no circuit is made of its Gate (gkr.md §3). The proving stack takes one type from the crate, SumcheckProof { rounds: Vec<[Fr; 4]>, final_evals }, as each layer of a gkr_verify::GkrProof. The GKR layer sumcheck (gkr.md §5) repeats step 3's rounds and checks under the same two tags, from a batched claim instead of 0 and to a final check of its own, in gkr::prove_sumcheck and gkr_verify::verify_sumcheck. prove_zerocheck and verify_zerocheck are called only by tests and tools/bench.

Auditors/Foundations

The transcript

Normative specificationdocs/spec/transcript.mdView as Markdown

Every challenge in the protocol is drawn from a Poseidon2 duplex sponge over Fr through a typed message layer. This page specifies the permutation, the sponge, the framing, a G1 point's transcript form and every tag. Implementation: crates/transcript, #![no_std].

1 The Poseidon2 permutation#

Width 3 over Fr, S-box x^5, 4 full rounds, 56 partial rounds (S-box on lane 0 only), 4 full rounds. The round constants are RC3 of HorizenLabs/poseidon2, plain_implementations/src/poseidon2/poseidon2_instance_bn256.rs at commit 055bde3f4782731ba5f5ce5888a440a94327eaf3.

E(s) = s + (s₀+s₁+s₂)·(1,1,1)                 circ(2, 1, 1)
I(s) = s + (s₀+s₁+s₂)·(1,1,1) + (0,0,s₂)      1 + diag(1, 1, 2)

poseidon2_permute(s):
  s ← E(s)
  RC3 rows 0–3:    s_i ← (s_i + c_i)^5, every lane;   s ← E(s)
  RC3 rows 4–59:   s₀ ← (s₀ + c₀)^5;                  s ← I(s)
  RC3 rows 60–63:  s_i ← (s_i + c_i)^5, every lane;   s ← E(s)

poseidon2_permute([0, 1, 2])₀ = 0x0bb61d24daca55eebcb1929a82650f328134334da98ea4f847f760054f4a3033

constants::POSEIDON2_RC3_INITIAL, _INTERNAL and _TERMINAL hold the 80 entries read (upstream's partial rows are zero in lanes 1 and 2) as upstream's big-endian hex literals, character for character, decoded by Fr::from_hex on every call. They are pinned through the permutation, by the oracle's 128 vectors (§2), each of which reads every constant.

On riscv32, poseidon2_permute is one POSEIDON2 delegation call over the lanes' canonical bytes, falling back to these rounds when the executor answers -ENOSYS (delegation.md §10).

2 The duplex sponge#

state  [Fr; 3]   lanes 0, 1 the rate, lane 2 the capacity; zero in Transcript::new()
input  [Fr; 2]   absorbed, not yet permuted: 0 or 1 pending between operations
output [Fr; 2]   squeezed, not yet handed out: 0 to 2

observe(x):  output ← []; input.push(x); if |input| = 2: duplex()
sample():    if |input| > 0 or |output| = 0: duplex(); return output.pop()
duplex():    n ← |input|; state[0..n] ← input; input ← []
             if n > 0: state[n..2] ← 0; state[2] += n
             poseidon2_permute(state); output ← [state[0], state[1]]
  • Absorption overwrites the rate. A short absorb zero-fills the rest of it and adds its length to the capacity, so [a] and [a, 0] differ; with nothing pending, a duplex is a pure squeeze and does neither.
  • Squeezed lanes leave from the end: the first sample after an absorb is state[1], the second state[0], and a third permutes again.
  • observe drops unread output and lanes past a buffer's length stay zero, which moves no challenge and makes the state a function of the operation sequence alone.

This is Plonky3's DuplexChallenger at width 3 and rate 2. The vectors crates/transcript is tested against come from tools/transcript-ref, which shares no code with it: Plonky3's Poseidon2 keyed with zkhash's own RC3, and a transcription of this section and §3 run beside that type, agreeing with it on every squeeze. The recursion format replays the same sponge over field cells, one P2_FIELD row a duplex step (recursion.md §4).

3 Typed messages#

append_scalars(tag, xs):  observe(tag); observe(|xs|); observe(x) for x in xs
append_scalar(tag, x)  =  append_scalars(tag, [x])
append_bytes(tag, b):     observe(tag); observe(|b|); observe(c) for each 31-byte chunk c of b,
                          zero-padded to 32 bytes, read little-endian
challenge_scalar(tag):    observe(tag); return sample()

The length, the scalar count or for bytes the byte count, delimits a message: "abc" and "abc\0" are each one chunk, below 2^248 < p, and differ. A challenge absorbs its tag, so it always comes from a fresh permutation.

The framing carries no kind, so each tag names exactly one of scalars, bytes or a challenge (§5): a tag of two kinds would make append_bytes(T, b"") and append_scalars(T, []) the same T, 0. So every digest — program identity, the SRS digest, transcript::io_digest, sumcheck::witness_digest, pcs::accumulator_digest — is a fresh sponge of typed messages ended by a raw sample(), never by a challenge under one of its message tags.

snapshot() captures the state and both buffers, and Transcript::restore resumes the same challenge stream. Its postcard form is 226 bytes, state[3], input[2], input_len: u8, output[2], output_len: u8, each Fr canonical; decoding refuses input_len ≥ 2, output_len > 2 and a nonzero lane past either length. The archived path's phase files hold the global transcript, and each shard's after its GKR pass, in this form (streaming.md §6). A shard transcript is no restored global sponge but a fresh one whose first message carries the global state digest (proof.md §4).

Each typed operation appends Absorb { tag, n_scalars } (payload elements: scalars, or chunks) or Challenge { tag } to event_log(). Raw observe and sample are not logged, the log never feeds the sponge and a snapshot omits it; checker::tape holds the global transcript's log to the order G1–G11 (tools.md §4).

4 G1 points#

A point is absorbed as four Fr limbs of its 64-byte encoding x ‖ y (primitives.md §3), with no curve arithmetic (transcript::g1_limbs):

[ x[0..16], x[16..32], y[0..16], y[16..32] ]   each half read little-endian, below 2^128 < p
[ S, S, S, S ]                                 the 64 zero bytes of infinity; S = 2^128

A coordinate is an Fq element and q > p, hence the halves. S is constants::G1_INFINITY_SENTINEL: no 16-byte half reaches 2^128, so the limbs determine the 64 bytes whether or not they encode a point on the curve. The absorber never refuses; a point is validated where it is decoded, before a pairing reads it.

transcript::append_g1_points(tr, tag, points) absorbs k points as one message of 4k limbs, never k messages, so the framed length binds k. pcs::append_g1_list is it over G1Affine::to_bytes, and pcs::append_g1 a list of one.

5 Tags#

Tag = u64: constants::transcript_tags, 45 tags numbered from 1 and named by transcript_tags::NAMES[tag − 1]; 0 is not a tag. Kinds: S scalars, B bytes, C challenge. Where: G1–G11 and the shard transcript are proof.md §2, §4, the SRS digest §3 there; identity program.md §8; Mercury mercury.md; GKR gkr.md §5; io_digest public-values.md §5; stacks and nodes recursion.md §1.3, §8.3. † marks a tag on no proof path.

tag where
1 PROTOCOL_SUITE S G1: [PROTOCOL_VERSION]
2 PUBLIC_INPUTS B G7: io_digest's 32 canonical bytes
3 COMMITMENT S a commitment list: identity, G8, shard witness, Mercury
4 SUMCHECK_ROUND S a sumcheck round's coefficients
5 SUMCHECK_CHALLENGE C a round's challenge; first, a zerocheck's eq-randomizers
6 EVALUATION_CLAIM S Mercury: the point, then the claimed values
7 PCS_OPENING S Mercury: proof points and evaluations
8 WITNESS_DIGEST S sumcheck::witness_digest's sponge, and its result †
9 SUMCHECK_FINAL_EVALS S the zerocheck's final evaluations †
10 MERCURY_INSTANCE S Mercury: [n]
11 MERCURY_ALPHA C Mercury: α
12 MERCURY_GAMMA C Mercury: γ
13 MERCURY_Z C Mercury: z, redrawn while 0
14 BDFG_BATCH C Mercury: δ
15 BDFG_POINT C Mercury: z′
16 PAIRING_MERGE C Mercury: the pairing merge ρ
17 MERCURY_BATCH C Mercury: the column batch ρ
18 ACCUMULATOR_DIGEST S pcs::discharge: the entry words' sponge, and its result †
19 ACCUMULATOR_MERGE C pcs::discharge: the per-check weight †
20 PUBLIC_INPUT_STREAM B io_digest: the input
21 PUBLIC_OUTPUT_STREAM B io_digest: the output
22 PROGRAM_IDENTITY S identity: [code_version]; G6: [identity]
23 VM_CONFIG S identity; G3
24 SHARD_COUNTS S G4
25 GKR_OUTPUTS S GKR: the output tables
26 GKR_OUTPUT_POINT C GKR: the top point
27 GKR_BATCH C GKR: a transition's claim batch
28 GKR_LAYER_CLAIMS S GKR: a transition's claimed values
29 GKR_CHILD C GKR: a halving transition's line point
30 MEMORY_WINDOWS S G5
31 MEMORY_BOUNDARY S G9
32 PROGRAM_ENTRY S identity: [entry_pc]
33 LOOKUP_CHALLENGE C shard: g, then β (lookup.md §2)
34 SRS_DIGEST S G2
35 SRS_VERIFIER B the SRS digest: the 320-byte SrsVerifier
36 MEMORY_GROUP S G8: [family, shard count]
37 MEMORY_CHALLENGE C G10, four times
38 GLOBAL_STATE_DIGEST C G11
39 SHARD_SEED S shard: [digest, family, index]
40 SHARD_TS_WINDOW S shard: [start, end]
41 GENERIC_TABLE S the SRS digest: the generic table's 3 points, 12 limbs
42 STACK_CHALLENGE C a recursion-format shard: its σ stack challenges
43 FOLD_STATE S a node: a verified shard's final transcript state
44 FOLD_WEIGHT C a node: a shard's w, w′, or a child's weight
45 FOLD_CHILD S a node: a child's journal

Auditors/Foundations

The structured reference string

Normative specificationdocs/spec/srs.mdView as Markdown

The powers of τ every commitment is made under: the ceremony they come from, how its file is read and what is checked, the archive an SRS is cached in, the three points a verifier holds, KZG over them, and the Groth16 first phase read from the same file. Implementation: crates/srs.

1 The ceremony#

The SRS is PSE's perpetual powers of tau, contribution 80: files ppot_0080_<p>.ptau, kept in assets/ptau/, which is gitignored. Hermez's powersOfTau28_hez_final_*.ptau is another ceremony with another τ; the reader ingests it as readily, and every commitment, key and identity over it differs. This ceremony's [τ]_1, as the hex of its canonical encoding x ‖ y:

9bbb31bedc304e081e2aada4b56c2217e0e94ee16874e3517d14bef5dcec3a16
317ff1589e53513fa333591b318e8f1e55ef7c37d92beb6d9a61d770a2f39506

PSE's files are cut from one ceremony: the first 2^k powers, and the Lagrange bases of domains up to 2^k, agree in every file of power k or more. One file, ppot_0080_24.ptau (19.3 GB), serves every use. A base key needs as many powers as its tallest family has rows, at most 2^22, the menu's top, and at least the generic table's 2^18 (lookup.md §9); bench prove reads 2^22. The recursion format reads 2^24, its largest stack (verifier_core::STACK_LOG, recursion.md §1.3), and the decider its domain's Lagrange bases (§7).

2 Ingesting a .ptau file#

Srs::from_ptau(path, k) reads snarkjs's .ptau container, the one ingestion format. Integers are little-endian.

0    4    "ptau"
4    4    version: 1
8    4    section count: at most 64
12   ..   sections: id u32 | size u64 | payload

id 1   header, 44 bytes: n8 = 32 | q (n8 bytes) = BN254's Fq modulus | power p | ceremonyPower
id 2   tauG1: 2^(p+1) − 1 G1 points, [τ^0]_1 first
id 3   tauG2: 2^p G2 points, [1]_2 then [τ]_2

Sections 1–3 must each occur once, at the sizes p implies; the others (alpha, beta, the contribution record, the Lagrange bases of §7) are not read here. from_ptau takes the first 2^k points of section 2 (k ≤ p) and the first two of section 3.

A point is uncompressed affine in little-endian Montgomery form: each 32-byte coordinate holds coord·R mod q, R = 2^256; G1 is x ‖ y, G2 x.c0 ‖ x.c1 ‖ y.c0 ‖ y.c1. It is the only non-canonical point encoding the code reads. A coordinate is read as a canonical Fq (refused at or above q), multiplied by R^−1 and re-encoded, and the canonical bytes go through G1Affine::from_bytes or G2Affine::from_bytes (primitives.md §3), the one validating decoder. All-zero bytes are infinity in both forms.

from_ptau never panics: every refusal is an SrsError — Io, Truncated, BadMagic, BadVersion, BadSection (over 64 sections, sections 1–3 not each present once, a header not 44 bytes, a power outside 1..=30, a section size p does not imply), WrongCurve, PowerTooLarge (k > p), and InvalidPoint { index }, a failing point but not necessarily the first.

3 What is validated, and what is presumed#

Decoding proves every point canonical, on its curve and in the order-r subgroup. Srs::validate adds that they are powers of one τ:

g1[0] = G1 generator      g2_gen = G2 generator      no point is infinity
e(Σ_i c_i·g1[i], g2_tau) = e(Σ_i c_i·g1[i+1], g2_gen)       i < n − 1

with each c_i 31 bytes from /dev/urandom, so that no file can be built to pass: a power that is not τ times the one before survives with probability at most 2^−248. The infinity check excludes τ = 0: pairing_check skips a pair at infinity (primitives.md §4), so such an SRS would pass vacuously and kzg_verify over it accept any opening. validate identifies nothing, and no proving path runs it.

Soundness needs nobody to know τ, which this code presumes of the ceremony. A statement binds the SRS only through the SRS digest (proof.md §3), which covers the SrsVerifier and the generic table's three commitments, not the powers, which only a prover reads. A key's loader recomputes the digest from the key's own points, so a key whose SrsVerifier has a known τ loads under its own digest: a verifier takes the ceremony's digest from a channel the prover does not control, or recomputes it from the ceremony. Program identity covers neither the SrsVerifier nor the table; in the recursion tree the digest is a constant of both programs' images, which their identities bind (recursion.md §8.1).

4 The SRS archive#

Srs::save and Srs::load keep an ingested SRS in a file of their own, integers little-endian and points canonical (primitives.md §3); bench recurse caches its 2^24 powers in one.

0     8          "APOGESRS"
8     4          version: 1
12    4          power k, at most 30
16    8          G1 count: 2^k
24    128        g2_gen
152   128        g2_tau
280   64·2^k     g1, [τ^0]_1 first

load requires exactly 280 + 64·2^k bytes before reading a point (the cap on k keeps the product from wrapping) and decodes every point through from_bytes, which catches a corrupted coordinate, not a substituted archive. It refuses with Truncated, BadMagic, BadVersion, BadSection and InvalidPoint.

5 SrsVerifier#

The only SRS material a verifier takes: g1_gen = [1]_1, g2_gen = [1]_2 and g2_tau = [τ]_2, what Mercury's pairings read. A verifier never commits; the generic table's commitments reach it as given points. The wire form is 320 bytes, g1_gen ‖ g2_gen ‖ g2_tau, canonical, unframed — its postcard form, VerifyingKey's srs_verifier and verifier::encode_srs_verifier alike — and every reader decodes it through the validating from_bytes.

6 KZG#

srs::kzg, over coefficients little-endian in the degree (coeffs[i] multiplies X^i, as g1[i] is [τ^i]_1):

kzg_commit(f)   = Σ_i f_i·[τ^i]_1                               one MSM
kzg_open(f, z)  = (f(z), [q(τ)]_1), q = (f − f(z))/(X − z)      one Horner pass gives both
kzg_verify(cm, z, v, w):  e(cm − v·[1]_1 + z·w, [1]_2) · e(−w, [τ]_2) = 1

More coefficients than powers is an error, never a truncation. The zero polynomial commits to infinity and opens to (0, infinity), which verifies. A Mercury commitment is exactly kzg_commit of the evaluation table read as coefficients (mercury.md §2). Mercury calls neither kzg_open nor kzg_verify, but its pairing relations take their shape, e(A, [1]_2) = e(B, [τ]_2) with both G2 arguments SRS constants, which is what lets recursion fold them instead of pairing (recursion.md §8.3).

7 Phase 1#

srs::Phase1::from_ptau(path, m), for m ≤ p and m ≤ 28, reads what a Groth16 key takes from the ceremony at a domain of n = 2^m: tau_g1, [τ^i]_1 for i < 2n − 1, from section 2; and lagrange_g1 and lagrange_g2, [L_j(τ)] in each group, L_j the Lagrange polynomial at ω^j and ω of order n squared down from constants::FR_TWO_ADIC_ROOT_OF_UNITY, from sections 12 and 13, which hold the bases of domains 1, 2, 4, … in turn, domain n from point n − 1. It refuses a basis that is not this domain's: tau_g1[0] and each basis's sum must be the generator, and Σ_j ω^j·[L_j(τ)]_1 = [τ]_1. The G2 basis is held to the curve, not the subgroup. The decider's key is made over it (recursion.md §9).

Auditors/Foundations

Mercury

Normative specificationdocs/spec/mercury.mdView as Markdown

Every committed column is opened with Mercury (Eagen and Gabizon, ePrint 2025/385), finished by the batched KZG opening of BDFG20 (Boneh, Drake, Fisch and Gabizon, ePrint 2020/081). This page pins what the papers leave open, and adds a batch of k columns at one point and the deferred form the recursion tree folds. crates/pcs is the prover, the curve side and the pairings; crates/pcs-verify, no_std, is the verifier's field side, which the recursion guest links.

1 Parameters and the variable split#

n = 2^{2t} evaluations with 1 ≤ t ≤ 27, b = 2^t = √n, s = 2t variables; u ∈ Fr^s is the opening point and v the claimed value. pcs_verify::check_num_vars refuses every other variable count (PcsError::UnsupportedNumVars) and never pads, which is why every trace height is an even power of two (program.md §7). The ceiling, pcs_verify::MAX_NUM_VARS = 54, is where Fr's 2-adicity of 28 runs out of the 2b-th roots of unity §3.1 needs, and it keeps 2^{|u|} in range for a u the verifier is handed.

The evaluation table is read as coefficients, and variable m is bit m of an index (primitives.md §6). Write an index i + j·b with i the low t bits, as Mercury §3.1 does; its evaluation is the coefficient of X^{i+j·b}. The point splits the same way: u1 is its first half, u_0..u_{t−1}, and pairs with i; u2 is u_t..u_{2t−1} and pairs with j.

f(X) = Σ_{i<b} X^i·f_i(X^b),    f_i(X) = Σ_{j<b} f_{i+j·b}·X^j
f̂(u) = Σ_{i,j<b} eq(i, u1)·eq(j, u2)·f_{i+j·b}

pcs::open returns what poly::MultilinearPoly::evaluate gives at u, and a verifier handed the two halves swapped rejects.

2 Commitment#

pcs::commit returns [f(x)]_1 for §1's f(X), an MSM over the first n SRS powers: exactly the KZG commitment of the evaluation table read as coefficients (srs::kzg::kzg_commit), with no second scheme behind it. It refuses an SRS of fewer than n powers (SrsTooSmall). A column backed by U1, U8, U16 or U32 (poly::PolyBacking) is widened to u32 and committed through curve::msm::msm_small_u32, never lifted to Fr; an Fr backing goes through curve::msm::msm.

The map from a table to its commitment is Fr-linear, which §5 uses, and a zero coefficient adds nothing: a column extended by zero rows keeps its commitment. So the generic table's commitments serve every height that holds the table (lookup.md §9), and pcs::commit_stack commits a recursion stack without building it.

3 The opening protocol#

3.1 The polynomials#

definition coefficients sent as
h Σ_i eq(i, u1)·f_i(X); its X^j coefficient is f̂(u1, j) b h
q, g f = (X^b − α)·q + g, so g = Σ_i f_i(α)·X^i n − b, b q, g
S the symmetrized witness below b − 1 s
D X^{b−1}·g(1/X): g reversed b d
H (f − (z^b − α)·q − g_z)/(X − z) n − 1 pi_z
W, W′ §3.3 b − 1 each w, w_prime

P_u(X) = Σ_{i<b} eq(i, u)·X^i = Π_{m<t}(u_m·X^{2^m} + 1 − u_m), so ⟨P_u, g⟩ = ĝ(u) for g of fewer than b coefficients (Mercury §4.2). The prover uses its coefficients, poly::eq_table(u); the verifier evaluates the product in O(t).

The fold (Mercury §5) divides every f_i by X − α, b Horner divisions advanced together in one pass over the rows, with no transform. Then ĝ(u1) = h(α) and ĥ(u2) = f̂(u) = v, and one S proves both inner products (Mercury §4.1), the left side's constant coefficient being 2·(⟨g, P_u1⟩ + γ·⟨h, P_u2⟩):

g(X)·P_u1(1/X) + g(1/X)·P_u1(X) + γ·(h(X)·P_u2(1/X) + h(1/X)·P_u2(X))
    = 2·(h(α) + γ·v) + X·S(X) + S(1/X)/X

S is coefficients b..2b−2 of X^{b−1} times the left side, computed with four forward transforms of size 2b and one inverse; no transform in an opening is larger (crates/pcs/src/fft.rs, over constants::FR_TWO_ADIC_ROOT_OF_UNITY).

3.2 The transcript schedule#

pcs::open and pcs_verify::scalars run this Fiat–Shamir schedule step for step. A point or a list of points is one message (transcript.md §4).

# tag message
1 absorb MERCURY_INSTANCE n
2 absorb COMMITMENT cm, as passed: open never recommits it
3 absorb EVALUATION_CLAIM u_0..u_{s−1}, then v
4 absorb PCS_OPENING h
5 squeeze MERCURY_ALPHA α
6 absorb PCS_OPENING [q, g]
7 squeeze MERCURY_GAMMA γ
8 absorb PCS_OPENING [s, d]
9 squeeze MERCURY_Z z, by §3.4's rule
10 absorb PCS_OPENING g_z, g_{1/z}, h_z, h_{1/z}, s_z, s_{1/z}, one message
11 absorb PCS_OPENING pi_z, before δ although the batch does not read it
12 squeeze BDFG_BATCH δ
13 absorb PCS_OPENING w
14 squeeze BDFG_POINT z′
15 absorb PCS_OPENING w_prime
16 squeeze PAIRING_MERGE ρ, after all eight points and six values

The prover draws ρ too and discards it, so both sides leave the transcript in one state and an opening composes inside a larger transcript, the shard transcript (proof.md §4).

3.3 The BDFG20 batch#

Mercury §6 step 4(e) leaves the batched KZG opening to BDFG20 §4. The point set is T = {z, 1/z, α}, and the four polynomials are batched in this order, which fixes the power of δ each carries (pcs_verify::bdfg::items, which both sides read):

i f_i S_i Z_{T∖S_i} r_i interpolates
0 g {z, 1/z} X − α g_z, g_{1/z}
1 h {z, 1/z, α} 1 h_z, h_{1/z}, h_α
2 S {z, 1/z} X − α s_z, s_{1/z}
3 D {z} (X − 1/z)(X − α) D_z
F(X) = Σ_i δ^i·Z_{T∖S_i}(X)·(f_i(X) − r_i(X))                          W  = [(F/Z_T)(x)]_1
L(X) = Σ_i δ^i·Z_{T∖S_i}(z′)·(f_i(X) − r_i(z′)) − Z_T(z′)·(F/Z_T)(X)    W′ = [(L/(X − z′))(x)]_1

Both divisions are exact for an honest prover, and open asserts it (pcs_verify::bdfg::{quotient, linearization}).

3.4 Challenges and derived values#

Mercury draws z ∈ F*; here z is drawn again under MERCURY_Z while it is zero (pcs_verify::challenge_z). T needs three distinct points, so both sides refuse with PcsError::DegenerateChallenge when z² = 1, z = α or z·α = 1 (pcs_verify::degenerate): probability about 2^−252, and a loss of completeness only. The recursion tape draws z once and asserts all four conditions (verifier_core::tape::mercury_scalars).

The verifier is not sent h(α) or D(z): it derives them, as Mercury §6 step 4(c) does (pcs_verify::derive_h_alpha), and the prover builds the batch around the same derived values.

D_z = z^{b−1}·g_{1/z}
h_α = (g_z·P_u1(1/z) + g_{1/z}·P_u1(z) + γ·(h_z·P_u2(1/z) + h_{1/z}·P_u2(z) − 2v)
       − z·s_z − s_{1/z}/z) / 2

Opening D at z to D_z is the degree check on g (Mercury §4.3); opening h at α to h_α is §3.1's identity at z.

4 The proof and the verifier's checks#

pcs::MercuryProof is eight points and six values. Its field order is its byte order and its transcript order, and to_bytes writes pcs::PROOF_BYTES = 704 bytes for every n and k:

h  q  g  s  d  pi_z  w  w_prime                 8 × 64 bytes, G1 uncompressed (primitives.md §3)
g_z  g_inv_z  h_z  h_inv_z  s_z  s_inv_z        6 × 32 bytes, canonical Fr (primitives.md §1)

from_bytes returns None unless every point decodes through curve::G1Affine::from_bytes (canonical and on the curve; G1's cofactor is 1) and every value through field::Fr::from_bytes.

Two relations are checked, each written e(A, [1]_2) = e(B, [x]_2) so that both G2 arguments are SRS constants: the fold identity at z (Mercury §6 step 4(f), its z term moved into G1) and the BDFG20 batch (BDFG20 §4.1). They merge under ρ into one curve::pairing::pairing_check of two pairs:

A1 = cm − (z^b − α)·q − g_z·[1]_1 + z·pi_z                  B1 = pi_z
A2 = Σ_i c_i·cm_i − K·[1]_1 − Z_T(z′)·w + z′·w_prime        B2 = w_prime
     cm_i = g, h, s, d    c_i = δ^i·Z_{T∖S_i}(z′)    K = Σ_i c_i·r_i(z′)
     Z_T(z′) = (z′ − z)(z′ − 1/z)(z′ − α)
accept iff  e(A1 + ρ·A2, [1]_2)·e(−(B1 + ρ·B2), [x]_2) = 1

If either relation is false the merged one holds for at most one ρ, and ρ follows every proof element. The verifier reads three SRS points, srs::SrsVerifier's [1]_1, [1]_2 and [x]_2, and does no G2 arithmetic.

pcs::verify refuses, in order: a u whose length is not an instance's (§1), before anything is absorbed (UnsupportedNumVars); a proof point or cm off the curve (InvalidPoint), checked again because a proof built in memory has met no decoder; a degenerate T (DegenerateChallenge); a failed pairing check (VerificationFailed), which does not say which relation failed.

5 Batching k columns at one point#

Not in the papers. k commitments to columns of one size, opened at one point u, are one Mercury instance with one proof (pcs::batch_open, pcs::batch_verify); a shard proof's opening is one such batch (proof.md §5). Three steps precede §3.2's sixteen (pcs_verify::batch_preamble):

# tag message
B1 absorb COMMITMENT cm_0..cm_{k−1}, as passed, one message of 4k limbs
B2 absorb EVALUATION_CLAIM u_0..u_{s−1}, then v_0..v_{k−1}
B3 squeeze MERCURY_BATCH ρ

The opening then runs on (cm*, u, v*), with cm* = Σ_i ρ^i·cm_i and v* = Σ_i ρ^i·v_i.

  • ρ follows every commitment and every claimed value. Column i carries ρ^i, column 0 carrying 1, so a reordered or shortened list is a different statement.
  • The list is one message, so its length 4k fixes k, and then s from B2's s + k scalars: the absorbed stream is injective.
  • ρ = 0 is not redrawn: it checks column 0 alone, and is one of the roots the bound below counts.
  • A batch of one is a different transcript from a bare opening; their proofs do not interchange.

The batch is sound: by §2's linearity cm* commits to f* = Σ_i ρ^i·f_i, and evaluation at u is linear, so v* − f̂*(u) = Σ_i (v_i − f̂_i(u))·ρ^i, a polynomial in ρ of degree at most k − 1 fixed before ρ is drawn. A false claim survives with probability at most (k − 1)/|Fr|.

The prover builds f* as one Fr column and opens it once; mixed sizes are refused (MixedColumnSizes). The verifier refuses an empty list (EmptyBatch) or a value count that differs (BatchLengthMismatch), checks every cm_i on the curve before summing, derives cm* by a k-point MSM and runs §4 on it. pcs::batch_open_stacked opens recursion stacks at u ‖ r (recursion.md §1.3); batch_open is it at r = [], one column a stack.

6 Deferred verification and the accumulator#

6.1 The twelve entries#

Deferring a verification runs every check of §4 but the pairing and keeps the relation's terms: twelve pcs::AccumulatorEntry { side, scalar, point }, side a pcs::PairingSide, G2One for [1]_2 or G2X for [x]_2. The points are [cm, h, q, g, s, d, pi_z, w, w_prime, [1]_1], as pcs_verify::ENTRY_POINTS indexes them, and the scalars are pcs_verify::scalars's, in §4's notation:

# side point scalar # side point scalar
0 G2One cm 1 6 G2One pi_z z
1 G2One h ρ·c_1 7 G2One w −ρ·Z_T(z′)
2 G2One q −(z^b − α) 8 G2One w_prime ρ·z′
3 G2One g ρ·c_0 9 G2One [1]_1 −(g_z + ρ·K)
4 G2One s ρ·c_2 10 G2X pi_z 1
5 G2One d ρ·c_3 11 G2X w_prime ρ

The G2One terms sum to A1 + ρ·A2 and the G2X terms to B1 + ρ·B2; entry 9 carries both relations' [1]_1, and entry 2 is zero exactly when z^b = α, which is legal. A batch derives cm* first, so entry 0 is cm* and a check is ENTRIES_PER_CHECK = 12 entries whatever k. pcs::verify and pcs::batch_verify spend the entries at once; pcs::verify_deferred and pcs::batch_verify_deferred return them.

6.2 What uses it#

  • Base verification pairs: crates/verifier runs pcs::batch_verify for each shard, and no ShardProof or BlockProof carries an entry.
  • The recursion tree folds. A shard's tape computes the twelve scalars over field cells (verifier_core::tape::mercury_scalars), cm* being a hint; the node folds them with the batch check cm* = Σ_i ρ^i·cm_i (recursion.md §8.3), and one pairing check at the top discharges every shard's (recursion.md §9). Natively, host::recursion runs pcs::batch_verify_deferred on each shard for its cm*.
  • Nothing else: pcs::verify_deferred, §6.3's word form, pcs::accumulator_digest and pcs::discharge are called only by crates/pcs's tests and tools/kat-gen.

6.3 The word form and discharge#

A list is grouped into deferred checks, checks[j] being group j's entry count, and written as canonical Fr words (pcs::accumulator_words, inverse pcs::accumulator_from_words):

group:  count  entry_0 .. entry_{count−1}
entry:  side  scalar  x_lo  x_hi  y_lo  y_hi       side 0 = G2One, 1 = G2X; ENTRY_WORDS = 6

A word is 32 bytes, so an entry is 192, and the limbs are the point's transcript form (transcript.md §4). There is no header, so two lists concatenate into a list whose checks keep their groups. Decoding refuses a count of 2^64 or more or one that overruns, a side other than 0 or 1, a limb of 2^128 or more other than the sentinel, a partial sentinel, the all-zero quadruple (infinity has one spelling), and a point that is not canonical or not on the curve. The digest is the words as one ACCUMULATOR_DIGEST message in a fresh sponge, then a raw sample; covering the count words, it binds the grouping.

discharge(vsrs, entries, checks):
  every entry's point on the curve, before anything else
  ν = fresh sponge: absorb ACCUMULATOR_DIGEST [digest], challenge ACCUMULATOR_MERGE
  A = Σ_j ν^j·(group j's G2One terms)      B = Σ_j ν^j·(group j's G2X terms)
  accept iff e(A, [1]_2)·e(−B, [x]_2) = 1

An entry's point is a claim: absorption binds only its limbs, and an entry built in memory has met no decoder. The weight keeps the checks apart: at weight 1, two checks with equal and opposite errors pass together, and weighted, a false group passes only where ν is a root of a nonzero polynomial of degree below the group count. ν is a function of the words because discharge takes no transcript. An empty list discharges.

7 Cost and security#

prover, field O(n): a pass for h, the fold, H's division; S in O(b log b)
prover, MSMs 2n + 5b − 4 scalar multiplications: q n − b, pi_z n − 1, h, g, d b each, s, w, w_prime b − 1 each. A commitment is one more MSM of n
batch of k k multiply-adds a coefficient for f* and a k-point MSM for cm*, then one opening
verifier O(t) field operations, MSMs of ten points and of two (and of k), one two-pair pairing check
measured n = 2^22: commit 1.30 s, open 2.89 s. 16 columns of 2^20: a batch opens in 1.01 s and verifies in 4.8 ms, 16 single openings take 9.79 s and 62 ms. 18-core Apple M5 Pro; bench mercury, bench mercury-batch

Knowledge soundness holds in the algebraic group model under q-DLOG (Mercury §6, BDFG20 §4), with Fiat–Shamir over the Poseidon2 transcript in the random-oracle model and an SRS whose x nobody knows (srs.md §3). The statistical terms are Schwartz–Zippel over α, z and z′, of order a committed polynomial's degree over |Fr|, a few 1/|Fr| for γ, δ and the merge ρ, (k − 1)/|Fr| for a batch and the group count over |Fr| for ν: each is below 2^−220 for every instance in use, and the level is BN254's (architecture.md §4). Nothing is hiding and nothing is blinded.

Mercury's SRS has exactly n powers; here one SRS serves every size, so a prover can commit to a polynomial of degree n or more, and no degree bound is checked. None is needed: Mercury §6's argument goes through with its Schwartz–Zippel terms over that degree, and the opening at u is the multilinear extension of the polynomial's first n coefficients. A commitment binds that truncation, which is linear, so §5's argument holds for it too.

Auditors/Program and execution

The program: from ELF to identity

Normative specificationdocs/spec/program.mdView as Markdown

How a guest binary becomes the static, verifier-known description of a program. crates/loader reads an ELF into a ProgramImage, crates/isa decodes its instructions, and crates/program routes them into per-family decoded tables, derives the VmConfig and commits to all of it as the program identity. Every step is a pure function of its input.

1 Loading#

loader::load_elf accepts a static executable — ELFCLASS32, little-endian, ET_EXEC, EM_RISCV — whose PT_LOAD segments lie inside guest RAM (constants::guest_memory) at even addresses, pairwise disjoint, with p_filesz ≤ p_memsz, at least one of them executable. Anything else is a named LoaderError: DynamicElf for ET_DYN, PT_DYNAMIC or PT_INTERP, EntryNotAnInstruction for an e_entry that is not the first halfword of an instruction, and the sweep's refusals (§2).

Of a program header it reads p_type, p_offset, p_vaddr, p_filesz, p_memsz and the PF_X bit, and nothing else: the VM has no pages, and all of RAM is addressable whatever the segments declare. The address map, and the segment layout a guest ELF keeps for host loaders, are ecall-abi.md §6.

2 RVC expansion and slots#

slots holds one Slot per halfword from slot_base, the lowest loaded address, to the end of the highest executable segment: the slot of pc is slots[(pc − slot_base)/2]. load_elf sweeps each executable segment's file bytes from its start, by the halfword at pc:

low bits 11   pc += 4   Instruction { word: the four bytes, compressed: false }, MidInstruction
0x0000        pc += 2   NonInstruction
otherwise     pc += 2   Instruction { word: rvc::expand(halfword), compressed: true }

Every other halfword is NonInstruction. An encoding longer than 32 bits (InstructionTooLong), one cut off by the end of the file bytes (TextTruncated) or a halfword rvc::expand refuses (RvcIllegal) refuses the image; 32-bit words are decoded in §5.

  • Addresses are never compacted. A c.addi at 0x1002 stays there and occupies two bytes, so linker-resolved addresses hold; compressed, the instruction's length, is the only record of whether the next pc is pc + 2 or pc + 4.
  • rvc::expand takes the base C extension in its RV32 form and refuses the floating-point forms, the RV64-only forms (c.addw, c.subw, a shift with shamt[5]), the reserved code points and the Zc* encodings. A HINT such as c.addi x0, 5 is expanded; its 32-bit form writes x0.
  • 0x0000, RVC's defined-illegal encoding, is not refused: LLVM pads unreachable blocks with it. Reaching it is fatal at run time.

A desynchronised sweep cannot make a wrong instruction provable. A slot is a function of the bytes at its own pc, so every Instruction slot is what a hart fetching there would decode; data that shifts the sweep off the true boundaries can only lose true instruction starts, whose pcs then have no table row (§5), or meet an unclaimed encoding and refuse the image. crates/loader/tests/differential.rs holds committed guests' slots to llvm-objdump's listing, and the expansion to LLVM's own encoder over guests/rvc-dense, one sequence assembled compressed and not: a wrong expansion would be a valid proof of another program.

3 ProgramImage and its wire form#

The wire form is postcard over ProgramImage's four fields in order, with no header; every integer but kind is a LEB128 varint:

ProgramImage = entry ‖ n ‖ n × Segment ‖ slot_base ‖ m ‖ m × Slot
Segment      = vaddr ‖ mem_len ‖ len ‖ bytes    mem_len is p_memsz; bytes, the p_filesz file bytes
Slot         = kind: u8 ‖ word                  kind 0 a four-byte instruction, 1 a two-byte one,
                                                2 MidInstruction, 3 NonInstruction; word 0 for 2, 3

The reader re-checks what load_elf establishes — segments at even addresses, sorted, disjoint and inside RAM; slot_base the lowest segment's address; each four-byte Instruction followed by its MidInstruction; entry an Instruction slot — but not slots against the bytes. artifact-dump writes this form (tools.md §5); no prover or verifier reads it, host::setup starting from the ELF.

ProgramImage::initial_word(addr) is the little-endian word at addr before the first cycle: file bytes where a segment has them, zero elsewhere. The image column (§8) and the trace's initial RAM values are read from it.

4 The instruction set and family routing#

isa::decode takes 32-bit words only and accepts exactly RV32IMA's 59 instructions — 40 of RV32I, 8 of M, 11 of A — with any value in an operand field, x0 destinations included, and the one legal value in every fixed field: funct7, jalr's funct3, all of ecall and ebreak, an atomic's .w width, lr.w's rs2 = 0. Everything else is a DecodeError: RV64 encodings, F, D, Zicsr, fence.i, privileged instructions. crates/isa/tests/sweep.rs holds it, over all 2^30 words with low bits 11, to accepted counts derived from the ISA's tables and to an independent encoder.

  • fence is every MISC-MEM word with funct3 = 000, 2^22 of them, whatever its rd, rs1, fm, pred and succ: the ISA has a base implementation treat a reserved setting as a normal fence (llvm-objdump prints those <unknown>), and on one hart a fence does nothing.
  • An immediate is the value the instruction uses: sign-extended for I, S, B and J, the shifted word for U, the amount for a shift immediate. B and J displacements are even by encoding; nothing asks for 4-byte alignment.

program::row_kind, a total function, routes an instruction to one family and one bit of that family's mask (§6):

id family mnemonics, from mask bit 0 up
0 ADD_SUB_LUI_AUIPC system (ecall ebreak fence), addi auipc add sub lui
1 JUMP_BRANCH_SLT slti sltiu slt sltu beq bne blt bge bltu bgeu jalr jal
2 SHIFT_BITWISE slli xori srli srai ori andi sll xor srl sra or and
3 MUL_DIV mul mulh mulhsu mulhu div divu rem remu
4 MEM_WORD lw sw
5 MEM_SUBWORD lb lh lbu lhu sb sh
6 ATOMICS amoadd amoswap lr sc amoxor amoor amoand amomin amomax amominu amomaxu, each .w

5 Decoded tables#

program::decode_program(image, params) decodes every Instruction slot — a word isa::decode refuses fails the program, reachable or not (NotAllOpcodesSupported) — and builds a table for each family of the config. An instruction family's columns are committed setup columns, which each cycle's decoder lookup reads (lookup.md §10); other families' tables have none.

  • One row per halfword, absolute. Row i is pc 2i, and a table has exactly its family's height h (§7).
  • A live row holds one of the family's instructions in the fields of its lookup tuple (program::lookup_tuple): pc, next_pc, rs1, rs2, rd, imm, extra_mask, without imm for MUL_DIV and ATOMICS; no tuple holds funct3, the mask saying more. next_pc is the fall-through, pc + 2 or pc + 4 by the slot's length, never a branch target. A register the form lacks is 0; imm is the two's complement of §4's value, 0 where the form has none, or a system code (§6).
  • Every other row is padding, Fr::MINUS_ONE in every field (FamilyTable::column_poly). An all-zero row would be a claimable instruction at pc 0 with an empty mask; a live field is below 2^32, so no live row is the padding row.
  • Reach. Derivation fails (TableTooShort) unless the family's own last instruction has pc ≤ 2h − 4; another family's code may lie beyond it. Code is linked from RAM_ORIGIN = 2^16, so a family reaches 1.9375 MiB of it at 2^20 and 7.9375 MiB at 2^22, the largest height.

Every Instruction slot is a live row of exactly one table, and no table has another (check_partition).

Code is static. A cycle's instruction comes from these tables, never from RAM: a store into .text changes what a load reads, not what executes, and a pc that is not an Instruction slot has no row, so reaching it is fatal and unprovable (execution-trace.md §10).

6 The extra mask#

A tuple's last field, family_extra_mask, is 1 << kind, the kind being the instruction's position in its row of §4's table (constants::extra_mask). A kind is a mnemonic, except family 0's bit 0, the system kind, whose three instructions are told apart by imm (constants::extra_mask::system_code): ecall 0, ebreak 1, fence 2. A fence's fm, pred and succ, and an atomic's aq and rl, are not recorded; on one hart they order nothing.

One-hotness is the table's, not a gate's: a circuit holds each bit it extracts boolean, and the decoder lookup, which admits only the table's rows, is what excludes an empty or many-bit mask (lookup.md §10).

7 VmConfig and heights#

verifier_core::VmConfig { families: Vec<(family, height)>, bytecode_size_words } is a program's static shape: its families, ascending by id (constants::family), each with its height, the row count of one of its shards; an execution's shard counts are not in it. decode_program derives the family set, and nothing selects it:

  1. an instruction family (0–6), whose rows are cycles, is present when the image holds one of its instructions;
  2. a window family, whose rows are memory locations — INIT_TEARDOWN (7), ZERO_WINDOWS (8), PUBLIC_INPUT (12), PUBLIC_OUTPUT (13), ADVICE_WINDOWS (14) — is always present, and FIELD_WINDOWS (18) when one of families 19–22 is, which puts the config in the recursion format (VmConfig::is_recursion), the one whose registry VmConfig::circuit reads for 18–22 (recursion.md §1.1, §1.2);
  3. a delegation family (9–11, 15–17, 19–22), whose rows are invocations, is present when the image declares it by a record among its file bytes (program::declared_delegations, delegation.md §7); a record naming a number no family answers is UnknownDelegation.

A height is a parameter (ProgramParams::heights, defaulting to constants::family::DEFAULT_HEIGHTS, circuits.md §1), except the public families', pinned at 2^12 because it places their windows (public-values.md §2). Four things constrain it:

  • the menu, constants::family::HEIGHT_MENU: 2^8, 2^12, 2^16, 2^18, 2^20, 2^22 (HeightNotOnMenu), even powers of two as a Mercury opening needs (mercury.md §1);
  • an instruction family's code (§5);
  • the window rules (verifier_core::window_height, WindowRule; memory.md §3.5), and RAM window 0, [0, 4·h_w), holding every file byte of the image (ImageOutsideWindow);
  • the floor of the family's lookup channels, below which the registry has no circuit and no key can be built: 2^20 for an instruction family (lookup.md §3), its own for a delegation family (delegation.md §9).

bytecode_size_words, 2^20 (4 MiB) by default, is a declared ceiling on the words from RAM_ORIGIN to the image's last file byte (ProgramTooLarge); no circuit reads it.

The wire form is 8k + 8 bytes; VmConfig::from_bytes refuses a wrong length, a family id above 22, ids not strictly ascending, a height off the menu, and what window_height refuses:

k: u32 LE ‖ k × (family: u32 LE ‖ height: u32 LE) ‖ bytecode_size_words: u32 LE

In a transcript a config is one VM_CONFIG message, [f_1 … f_k, h_1 … h_k, bytecode_size_words]: the second message of the identity (§8) and the first of a statement's descriptor (proof.md §2).

8 Program identity#

One Fr, ProgramIdentity, on the wire its canonical 32 bytes (primitives.md §1): the raw squeeze of a fresh transcript after these messages (verifier_core::identity_digest; tag values in transcript.md §5):

# tag message
1 PROGRAM_IDENTITY [code_version]: constants::family::CODE_VERSION, 0, the only one derivation builds (UnsupportedCodeVersion)
2 VM_CONFIG §7's
3 PROGRAM_ENTRY [entry_pc]
4 COMMITMENT, one per family of the config, ascending its setup commitments, four limbs a point (transcript.md §4)

A family's setup commitments (program::setup_commitments) are Mercury commitments (mercury.md §2): of an instruction family's decoded-table columns in tuple order, 7 or 6 points; of INIT_TEARDOWN's image column (program::image_init_column), row y being image.initial_word(4y) over RAM window 0, one point; none, an empty message, for every other family.

Committing needs the ceremony's SRS, with as many powers as the tallest table has rows (srs.md §1). The digest over given points (program::identity_from_commitments) needs no SRS and no curve arithmetic: a verifying key carries the lists, its load recomputes the identity from them, and every shard opens its setup columns against the same points (proof.md §5, §7), which is what ties the tables a proof reads to the identity.

It binds every instruction the sweep found, with its pc, length, operands and kind, and that no other pc holds one; every file byte of the image (.text, .rodata, .data, the delegation declarations among them); the entry pc; the family set, every height, bytecode_size_words and the code version. One ELF at two settings of the heights has two identities.

It does not bind:

  • the SRS its commitments are under, or the generic lookup table: those are the SRS digest's (proof.md §3);
  • the circuits, which a key's load holds to the registry (proof.md §7);
  • anything an execution chooses: its input, advice, shard counts, window list;
  • memory past a segment's file bytes (.bss, the heap and the stack): mem_len enters nothing, and such memory starts at zero whatever is declared;
  • the symbol table, which loader::function_symbols and loader::symbol_names read beside the image for the profiler and listings, or anything else of the ELF §1 does not read.

A verifier takes the identity from a channel the prover does not control and compares it with its key's; it never sees an ELF. Against a prover-supplied identity a proof shows only that some program ran. Whoever holds the ELF, the parameters and the ceremony file recomputes it.

Auditors/Program and execution

The guest ABI

Normative specificationdocs/spec/ecall-abi.mdView as Markdown

The ecall convention, every ecall number, the guest's address space and the SDK over them. A guest has no file descriptors and no I/O syscall: its public input, journal and advice are memory (public-values.md), so an ecall only ends the execution or hands a frame to a circuit. crates/constants/tests/ecall_abi.rs holds this page's tables to constants::ecall and its MEMORY line to link.ld.

1 The calling convention#

Register Role
a7 the number
a0 in: the one argument, an exit status or a delegation's frame base (delegation.md §4)
a0 out: the result, 0 or a negated errno

a1–a5 are reserved for a call that needs more arguments; none does. A recursion-format delegation answers its frame base advanced past the frame (recursion.md §1.4). An ecall preserves every register but a0: its row writes no other (execution-trace.md §6), so the memory argument carries the rest across it.

2 The number ranges#

Constant Value What
ZKVM_IO_FIRST 0x0400 first host call
ZKVM_IO_LAST 0x04FF last host call
PRECOMPILE_FIRST 0x0500 first precompile
PRECOMPILE_LAST 0x05FF last precompile

Both ranges lie above 1023, the whole Linux number space, and are disjoint, so a number says its class: a host call would return a value the prover chose, a precompile is a deterministic function of guest memory that its circuit proves. The host-call range is reserved and empty: advice is a memory region the prover fills and the guest checks (public-values.md §6), inside the memory argument, where a value returned in a register would be bound to nothing.

3 Syscall numbers#

Every number this VM implements. All but EXIT are delegations, whose families, anchor spaces and frames are delegation.md §3's registry: the first six are the base format's, the last four the recursion format's (recursion.md §2).

Number Constant Class What
93 EXIT deterministic end the execution with status a0, the statement's exit status; nonzero is a failed execution, still provable
0x0500 PRECOMPILE_POSEIDON2 deterministic the width-3 Poseidon2 permutation over canonical Fr lanes
0x0502 PRECOMPILE_FR_ARITH deterministic one Fr add, multiply or inverse over Fr's in-memory form
0x0504 PRECOMPILE_MOD_MUL deterministic a·b mod m, m one of four Ethereum moduli a selector names
0x0506 PRECOMPILE_EC_ADD deterministic one third of a complete point addition, secp256k1 or BN254 G1
0x0507 PRECOMPILE_KECCAK_F deterministic one round of keccak-f[1600]; a permutation is 24 calls
0x0508 PRECOMPILE_SHA256_COMP deterministic four rounds of SHA-256's compression; a compression is 16 calls
0x0509 PRECOMPILE_FR_OP deterministic one operation over field cells
0x050A PRECOMPILE_P2_FIELD deterministic one transcript duplex step over field cells
0x050B PRECOMPILE_FIELD_IO deterministic eight RAM words into a field cell, or back
0x050C PRECOMPILE_FQ_OP deterministic one BN254 base-field operation over field cells

A class says who chooses the result. deterministic: a function of the guest's own state, which a circuit proves. advice: chosen by the prover; no number has it (§2). These are exactly the ecalls a proof admits: ADD_SUB_LUI_AUIPC holds every ecall row's a7 to 93 or to a registered delegation number, the base format's circuit knowing the first six (add-sub.md, recursion.md §1.2).

4 Retired numbers#

Number Constant Was
63 none POSIX read(fd, buf, len)
64 none POSIX write(fd, buf, len)
0x0501 RETIRED_KECCAK_F_WHOLE_PERMUTATION a whole keccak-f[1600] over a 200-byte frame, which 0x0507 replaces
0x0503 RETIRED_MOD_MUL_WITNESSED_MODULUS a·b mod m over a 128-byte frame carrying m, which 0x0504 replaces
0x0505 RETIRED_SHA256_COMP_WHOLE_COMPRESSION a whole compression over a 96-byte frame, which 0x0508 replaces

A number is assigned once. A retired one is never reassigned and answers -ENOSYS (§5): given a second meaning, it would run an old binary with its frame misread to a plausible wrong answer.

5 Every other number#

Constant Value What
ENOSYS 38 answered as -ENOSYS in a0

A number not in §3 answers -ENOSYS and falls through: the retired numbers, the host-call range, and every syscall a library might make for host data — getrandom, clock_gettime, the seeding of std's RandomState. Host data is prover advice, and a guest that needs it takes it from the advice region, where checking it is visibly the guest's job. Such a call executes and cannot be proved: the ADD_SUB_LUI_AUIPC fill refuses its row.

-ENOSYS is also the delegation ABI's "no circuit" answer, on which a base-format shim runs its software path (delegation.md §2); this executor never gives it to a §3 number. A registered number the image did not declare is the fatal DelegationFamilyAbsent on the tracing paths (delegation.md §7).

6 The memory map#

crates/guest-sdk/link.ld declares one region, constants::guest_memory's RAM_ORIGIN and RAM_LENGTH:

ld
MEMORY { RAM (rwx) : ORIGIN = 0x00010000, LENGTH = 0x7FFF0000 }

The whole 32-bit address space:

[0x0000_0000, 0x0000_8000)  hole: no family initializes it; an access is a fatal OutOfBounds
[0x0000_8000, 0x0000_C000)  public input window    PUBLIC_INPUT_ORIGIN    16 KiB
[0x0000_C000, 0x0001_0000)  journal                PUBLIC_OUTPUT_ORIGIN   16 KiB
[0x0001_0000, 0x8000_0000)  RAM                    RAM_ORIGIN, RAM_LENGTH
    0x0001_0000             .text, _start first; .rodata, .data, .bss, each page-aligned
    __heap_start            .bss's end rounded up to 16; the heap grows up from here
    0x7F80_0000             __stack_top − STACK_RESERVE (8 MiB): no heap block ends above it
    0x8000_0000             __stack_top, the initial sp; the stack grows down
[0x8000_0000, 2^32)         advice                 ADVICE_ORIGIN          up to 2^29 words
  • A load or store reaches the four regions alike, every word carrying the RAM tag; the windows' and the advice's layouts, families and binding are public-values.md §2–§6. None is in the ELF, so no linker symbol names them. Advice is addressable only up to the words the host supplied, and not at all when it supplied none.
  • The hole makes a null dereference a fatal error rather than a trace nothing could prove.
  • A delegation frame lies wholly in RAM (delegation.md §4). crates/loader refuses a PT_LOAD outside RAM (program.md §1), and a decoded table's height bounds how far .text reaches (program.md §5).
  • crt0's _start sets sp, zeroes [__bss_start, __bss_end) byte by byte, so that a zero .bss is the image's property and not the executor's, calls main, and exits 0 if it returns.
  • Nothing detects a stack that grows past its reserve after the heap has filled below it.

6.1 The segment layout#

A guest ELF loads under two loaders. crates/loader lays its PT_LOADs into a flat space the executor makes addressable whatever the headers say, with no pages and no permissions. A host loader maps exactly the PT_LOADs, page by page, at their permissions, and nothing else exists. The headers are the image's account of its own memory, read by every tool but this VM, so link.ld makes them true:

  • Every writable byte is declared. .bss runs to ORIGIN(RAM) + LENGTH(RAM), so the heap and the stack lie in one writable segment ending at __stack_top, whose file bytes stop at or before .bss, the 2 GiB reservation being NOBITS. Undeclared, the first stack push would fault.
  • No two segments share a page. .text, .rodata, .data and .bss are each 4096-aligned: a page two mappings share takes the second's permissions, stripping execute from .text's tail or putting zero fill on a read-only page, which a host loader refuses.

crates/loader/tests/layout.rs holds the committed guest ELFs to both by parsing their headers.

7 The guest-sdk surface#

crates/guest-sdk is the guest's runtime. Only exit and the delegation shims issue an ecall; the rest is loads and stores.

Item What
entry!(f) exports the main crt0 calls, a wrapper calling f
public_input() the public input payload, its length word clamped to the window
read_input(buf) copies min(buf.len(), public_input().len()) bytes and returns the count: it may return short
commit(bytes) appends to the journal and its length word; exits 70 rather than overflow the window
journal() what has been committed
advice() the advice payload, its length clamped to the region; bound by nothing, so the guest checks it
exit(code) EXIT; publishes nothing beyond what was committed
keccak256, sha256 over KECCAK_F and SHA256_COMP, with a software fallback on -ENOSYS from the first call
poseidon2_permute over POSEIDON2; false on -ENOSYS, for the caller's own permutation
ec_add, ec_mul, ec_identity homogeneous projective points over EC_ADD; None on -ENOSYS
recursion::* the raw shims over word-aligned frame types, false on -ENOSYS; the recursion format's (fr_op, p2_field, field_io, fq_op and the tape helpers import, import_run, replay) have no software path
allocator bumps up from __heap_start, never frees; exits 71 when a block would end above __stack_top − STACK_RESERVE or the live sp
panic handler exits 101 and writes nothing: a panicking guest is provable, having published what it committed

A delegation answer other than 0 or -ENOSYS exits 72, as do -ENOSYS after the first call of a multi-call operation and a recursion call that does not leave a0 past its frame. Each shim reads its number from its declaration record (delegation.md §7); which library code reaches which shim is delegation.md §10's.

Auditors/Program and execution

The execution trace

Normative specificationdocs/spec/execution-trace.mdView as Markdown

What every memory query of an execution is, when it happens and the order the trace records it in; the emulator that produces a trace and the containers that hold one. The memory argument (memory.md) and every family's frame are built on this convention.

1 The clock#

Cycle c occupies the four timestamps 4c + Δ, one per slot Δ ∈ {0, 1, 2, 3} (constants::memory::TS_STEP). Every instruction is one cycle and nothing else is: a delegation invocation rides the cycle that requested it, so the cycle count is the instruction count.

  • Cycles are numbered from 1. Timestamp 0 is every address's initial write, and a read must strictly precede its write, so a cycle-0 pc query could not follow the value it reads.
  • The clock is 38 bits (TS_BITS): every timestamp is below 2^38, the last cycle is 2^36 − 1, and the cycle that would pass it is the fatal ClockOverflow, raised before it is recorded.

2 Address spaces#

Tag Space Address At timestamp 0
1 REG a register index, 0..32 0, x0 included
2 RAM the byte address of a 4-aligned word in RAM, a public window or the advice region the image's bytes in RAM, 0 past them; the public input's and the advice's layouts; 0 in the journal
3 PC 0 the entry point
4–9 DELEGATION_KECCAK_F … DELEGATION_EC_ADD, a base-format delegation family's anchor each a frame base no initial write
10 FIELD a cell, any u32 0
11–14 DELEGATION_FR_OP … DELEGATION_FQ_OP, a recursion family's anchor each a frame base no initial write

The tags are nonzero so that no real tuple is all zeros, as x0's initial write would be. A byte or halfword access queries its word, and trace::InitialMemory is what RAM starts from. An anchor space, in delegation.md §3's order, is a delegation family's type, not memory: a query there reads the tuple stamped 0 with value 0, whatever came before; requests pair with invocations and nothing chains (delegation.md §5). Field cells hold whole Fr elements, reached only by the recursion families (recursion.md §2).

3 A query#

A memory query is one event at one address (trace::MemoryEvent): a read of read_value, last written at read_ts, and a write of write_value at ts = 4c + Δ.

  • A query that only reads writes back what it read: a register read or a load is one query.
  • read_ts < ts, strictly; the gap ts − read_ts − 1 is below 2^38.
  • Queries at distinct addresses may share a slot; two at one address never do. An address may be queried at two slots of a cycle — add a0, a0, a1 reads a0 at slot 1 and writes it at slot 3 — which is why the log is ordered by slot.

4 The frame of each instruction class#

Slot 0 is the pc query, every cycle: pc read, next_pc written. A register query exists for every register field of the decoded instruction, whatever register it names, x0 included.

Class Δ = 1 Δ = 2 Δ = 3
lui, auipc, jal rd
jalr, register-immediate rs1 rd
branches rs1 rs2
register-register, M rs1 rs2 rd
loads rs1 the word, read rd
stores rs1 rs2 the word, the stored bytes merged in
lr.w rs1 the word, written back; rd ← it
sc.w rs1 rs2 the word ← rs2; rd ← 0
AMOs rs1 rs2 the word ← op(old, rs2); rd ← old
fence
ecall a7 a0 (§6) a0 ← the result; a delegation's mirror query

ebreak has no row (§10). An atomic's row and a delegation request's carry two queries at slot 3, at distinct addresses. A family's frame is the union of its instructions' queries (memory.md §2). next_pc is the fall-through — pc + 2 after a compressed instruction, pc + 4 otherwise — except a jal's or taken branch's pc + imm, a jalr's (rs1 + imm) & !1, and the exit row's HALT_PC (memory.md §5).

An invocation's accesses ride its requesting cycle but belong to its own family's row: its frame words in RAM at slot 0 (constants::delegation::FRAME_DELTA), a FIELD_IO invocation's eight data words in RAM at slot 1 (constants::field_io::DATA_DELTA), and a recursion family's field cells at slots of its own (delegation.md §4, recursion.md §2.1).

5 The x0 rule#

x0 is an ordinary register in the trace and a constant in the machine: it starts at 0, a read of it is a REG query at address 0, and an instruction whose rd is x0 logs its slot-3 write with value 0, whatever it computed. So every query at x0 reads and writes 0, which the x0 gadget enforces (memory.md §2).

6 ecall#

An ecall's row is one cycle of ADD_SUB_LUI_AUIPC. It reads a7 at slot 1 and writes a0 at slot 3; the rest depends on the number (ecall-abi.md):

a7 Δ = 2 a0 written next_pc Besides
EXIT a0, the status the status HALT_PC the execution stops
a delegation number a0, the frame base 0, or for a recursion type the base past the frame (recursion.md §1.4) fall-through the mirror query at the frame base (§7); the invocation (§4)
any other none -ENOSYS fall-through no proof admits the row (ecall-abi.md §3)

7 The order of the log#

Events are recorded in cycle order and, within a cycle, by slot and then by role: the pc query; then an invocation riding the cycle, its frame words in frame order and a FIELD_IO invocation's data words after them; then one query per role the row has, in trace::ROLES order:

Role Slot Space What
rs1 1 REG rs1; an ecall's a7
rs2 2 REG rs2; an ecall's argument a0
load 2 RAM a load's word
ram 3 RAM a store's or an atomic's word
rd 3 REG rd; an ecall's result a0
delegate 3 the requested family's anchor space a delegation request's mirror query

ROLES is in slot order, so the log is in timestamp order, which MemoryState::record asserts; trace::Row::present holds one bit per role in a u8, and no two roles share a (space, slot) pair. The atomics family keeps its word at slot 3 for every instruction, lr.w included, so one frame serves the whole A extension.

8 Routing#

Every cycle goes to the one family whose decoded table claims its pc (program.md §4); a pc no table claims, or one claimed by a family not its instruction's, panics the tracer. An invocation goes by its type to its family's buffer.

9 The trace-level memory check#

MemoryEventLog::self_check(&InitialMemory) runs the memory argument natively over a whole log. First the timestamp rules: every address one its space has, every timestamp on the clock and in order, every read before its write, one query per address and timestamp. Then the balance: as multisets of (space, address, timestamp, value), an initial write at timestamp 0 of every touched address plus every query's write equals every query's read plus a teardown read of every address's last write. With one write per address and timestamp and no negative gap, this pairs each read with the last write before it: sequential consistency. What it cannot see:

  • Teardown is each address's last write, taken from the log, so everything after an address's last honest query balances by construction: a final value changed, a final query moved later or added, trailing cycles removed. In a proof the final values are the boundary scalars and the window families' teardown columns, fixed before any memory challenge, and the verifier fixes x0's and the pc's (memory.md §4, §5).
  • An anchor-space query is credited with its invocation's two tuples and balances alone; that requests and invocations pair 1:1 is the circuits' (delegation.md §5).
  • Field-cell accesses are not events (recursion.md §2.1).

10 The emulator#

crates/emulator runs RV32IMAC on one hart over a ProgramImage, with no interrupts and no privilege levels; aq/rl and fence order nothing. emulator::run returns an Execution: the registers, the exit status, the cycle count and the public values. emulator::trace_run returns the family buffers, the MemoryEventLog and the CycleProfile too, and emulator::StreamingRun, the prover's pull-based tracer, hands over a family's buffer as a ShardChunk the moment it reaches its height (streaming.md §2). The three differ only in what records a cycle, and a run is a pure function of (image, io), with no clock, randomness or threads, so two runs cut the same shards. A nonzero exit status is an execution, not an error.

Three points differ from a hosted RV32IMAC. sc.w always succeeds, storing and writing 0, as the circuits do (memory-ops.md §6). A misaligned halfword or word access is fatal, never split. The instruction stream is the image decoded at load, so a store into .text changes RAM and not what executes.

Every other stop is a fatal EmuError, and run and trace_run return no trace beside one: NotAnInstruction (the all-zero halfword included), IllegalInstruction, Ebreak, Misaligned (a frame base too), OutOfBounds (an access outside ecall-abi.md §6's regions, an advice word past what the host supplied, or a frame not wholly in RAM), ClockOverflow, PublicInputTooLong and JournalTooLong (the input, or the journal's length word at exit, above a window's payload), DelegationFamilyAbsent (on the tracing paths, a delegation number the image did not declare) and DelegationFrame (a frame its family has no witness for, delegation.md §6). An unassigned ecall number is not an error but -ENOSYS (§6).

There is no second executor: crates/emulator/tests/trace.rs restates §4's table and checks every traced row against it, and §9's check and the checker's multiset, memory and family-row suites hold the rest.

11 Trace containers#

crates/trace holds what an execution leaves; the emulator is its only producer.

  • Family buffers. trace::FamilyTraces holds one buffer per family of the VmConfig. A FamilyTrace, empty for a window family, is raw live rows, column-major, in small integer types: cycle, pc, next_pc, present, and per role addr, read_ts, read_value, write_value; no padding, no polynomial. A row stores everything its queries carry but a write timestamp, 4c + Δ, and the pc query's read timestamp, 4(c − 1). A delegation family's DelegationTrace has a row per invocation: the requesting cycle, the frame base, the frame words, and a recursion family's cell and data-word accesses.
  • RowSlice, FrameSlice. One shard's rows, [i·h, min((i + 1)·h, len)), borrowed: what the memory column builders read, never the log (memory.md §2). Row::delegation_space recovers a mirror query's space from the a7 the row read.
  • MemoryState, the last-access tables: each register's, the pc's, each RAM word's and each field cell's last (ts, value). O(touched addresses), and all the register and pc boundary, the RAM window list (trace::init_windows) and the window families' teardown need.
  • MemoryEventLog, the events and a MemoryState: O(cycles), kept only by trace_run, read by §9's check, the TraceArchive and checker::memory_columns_from_log, the independent reading the column builders are held to.
  • TraceArchive, the post-execution snapshot: buffers, log, profile, public values and advice. Its file is two postcard values, five phase sections and then their timings, so the deterministic payload is a byte prefix of it, and only a canonical encoding of self-consistent parts is read back. No proving path reads one; checker::TamperHarness and the retained archived path do (streaming.md §6).
  • CycleProfile, ShardPlan. The profile counts rows per family, cycles for a cycle-owning family (summing to the cycle count) and invocations for a delegation family. trace::plan_shards is ⌈count / height⌉ per family; a window family plans 0 there, its count being the prover's (streaming.md §4).

Auditors/Program and execution

Public values and advice

Normative specificationdocs/spec/public-values.mdView as Markdown

How an execution's public input and public output, the journal, are bound to its proof, and what the prover's advice is. All three are regions of guest memory, each initialized by a window family of its own and carried by the memory argument (memory.md).

1 Three regions, no I/O syscall#

region contents chosen by bound by
public input the statement step 10c, to the statement's input (§5)
journal the guest's stores step 10c, to the statement's output (§5)
advice the prover nothing (§6)

There is no I/O syscall: a guest uses ordinary loads and stores, and a provable guest's only ecalls are EXIT and delegation numbers (the retired POSIX numbers: ecall-abi.md §4). A host supplies emulator::GuestIo's input and advice and reads the journal from emulator::Execution::io. A byte-moving syscall would need cross-row constraints tying each transfer row to its buffer and length, which this arithmetization has no place for, while a window is bound by the multiset and one comparison (§5) and asks nothing of the guest. Nor does the guest hash its output: nothing rests on its honesty, and a panic loses nothing it committed.

2 The memory map#

region bytes family windows
public input [0x8000, 0xC000) PUBLIC_INPUT PUBLIC_INPUT_WINDOW = 2, at 2^12
journal [0xC000, 0x1_0000) PUBLIC_OUTPUT PUBLIC_OUTPUT_WINDOW = 3, at 2^12
advice [0x8000_0000, 2^32) ADVICE_WINDOWS from 2^29/h, at the window height h

The full address map is ecall-abi.md §6. The public windows take the upper half of [0, RAM_ORIGIN), 64 KiB that no RAM window family initializes (INIT_TEARDOWN masks window 0's rows below RAM_ORIGIN and no ZERO_WINDOWS id is 0, memory.md §3), so they cost RAM nothing, and [0, 0x8000) stays a hole in which a null dereference cannot balance.

A window's first address is 4·height·id, so the pinned height constants::family::PUBLIC_WINDOW_HEIGHT = 2^12 is what makes the origins windows 2 and 3, 16 KiB each, ending flush against RAM_ORIGIN. It is the ceiling: at the next menu height, 2^14, two windows need 128 KiB, and the one window in the hole is window 0, which would initialize address zero. Anything larger means moving RAM_ORIGIN, which moves every program's load address and shortens every decoded table's pc reach (program.md §5).

program::decode_program assigns that height whatever its caller asks, and the verifier refuses any other, and any RAM window height that would let a zero window reach the public windows (memory.md §3.5).

3 Layout and the length word#

word 0       the payload's byte length
words 1 …    the payload, little-endian, zero-padded to the end of the window

A public window is 2^12 words, so a payload is at most guest_memory::PUBLIC_PAYLOAD_BYTES = 16,380 bytes. verifier_core::public_io_words is the one spelling: the executor seeds the input window with it, the prover commits it and the verifier evaluates it. The length word makes the binding exact: without it [1, 2, 3] and [1, 2, 3, 0] fill the same window. verifier_core::derive_global_phase refuses an input or output longer than 16,380 bytes as Statement, and the executor refuses such an input before the first cycle.

4 The window families#

family id height shards init leaf step 10c holds
PUBLIC_INPUT 12 2^12 exactly 1 M[2] init_value M[2] to input
PUBLIC_OUTPUT 13 2^12 exactly 1 literal 0 M[1] teardown_value to output
ADVICE_WINDOWS 14 h k ≥ 0 M[2] init_value nothing

All three are in every VmConfig and own no cycles. Each public family proves exactly one shard in every statement (memory.md §3.5), so step 10c always runs: an unread input is still the window's initial contents, and an unwritten journal is empty.

The circuits are memory.md §3.3's. PUBLIC_OUTPUT's is ZERO_WINDOWS' byte for byte, whose init leaf writes the literal 0, so no column holds an initial journal (§5). PUBLIC_INPUT's and ADVICE_WINDOWS' initial values are M[2], one execution's values, committed before the memory challenges and bound by no program identity.

All three regions' tuples carry constants::address_space::RAM; which family initializes an address is what makes a word public, advice or heap. A space of their own would need an address-space column, and a gate pinning it, on the memory path of MEM_WORD, MEM_SUBWORD and ATOMICS; under RAM those circuits need nothing for them, their addressing already covering every 4-aligned address below 2^32 (memory-ops.md §2).

5 The binding#

io_digest absorbs the statement's two strings in a transcript of its own (transcript::io_digest):

t ← Transcript::new()
t.append_bytes(PUBLIC_INPUT_STREAM,  input)      tag, byte length, 31-byte limbs
t.append_bytes(PUBLIC_OUTPUT_STREAM, output)
io_digest ← t.sample()                           one raw squeeze

The framing (transcript.md §3) parses back to exactly one ordered pair, and the squeeze is raw, as every digest's is. The guest never computes it. G7 absorbs it before the memory commitments (G8) and challenges (G10) (proof.md §2), so both strings are fixed before any challenge exists.

The multiset. At a window address the init leaf is the only write at timestamp 0, every access consumes a write and produces a strictly later one, and the teardown balances only against the last (memory.md §9). So PUBLIC_INPUT's M[2] holds each word's value before its first access, and PUBLIC_OUTPUT's M[1] its value at the end.

Step 10c of verifier_core::verify_shard_local (proof.md §6). Of a public shard's base claims, which share one point u and each name a column, the verifier takes the one on M[2] (PUBLIC_INPUT) or M[1] (PUBLIC_OUTPUT), refusing its absence as Malformed, and compares it with its own evaluation at u of the multilinear extension of public_io_words(input) or public_io_words(output). A mismatch is MemoryArgument; the shard's opening then holds the claim to the committed column. Column and string are fixed before u is drawn, so a column other than the window passes with probability at most 12/p.

PUBLIC_INPUT's teardown is free: a guest may overwrite its input. PUBLIC_OUTPUT has no init column, and that is the point: with one, a prover could place the journal there at timestamp 0 and the teardown would match without the guest storing a byte.

5.1 The argument, stated plainly#

G7 fixes input and output, and G8 the window columns, before any challenge. Step 10c says the columns are those strings' windows; the multiset says they are the execution's first values in the input window and its last values in the journal window. So the guest found the statement's input in its input window, and the statement's output is what its stores left in the journal window. That rests on no cooperation, hash or register convention of the guest's, and says nothing about advice.

Recursion carries the binding unchanged: a node recomputes io_digest from the windows' words and repeats step 10c over them, and the decider binds the contract's input and output calldata to io_digest (recursion.md §8.1, §9).

6 Advice#

Advice is memory whose initial values the prover chose: ADVICE_WINDOWS initializes [ADVICE_ORIGIN, ADVICE_ORIGIN + 4hk) from an M[2] that nothing binds, not identity, not the statement, not a gate. A guest reads it with ordinary loads.

  • Layout. §3's framing over 1 + ⌈len/4⌉ words (trace::advice_region_words), spelled once by trace::advice_word for the executor and the prover; guest_sdk::advice reads it back.
  • Windows. At the window families' one height h, shard i is window verifier_core::advice_first_window(h) + i, and advice_first_window(h) = 2^29/h is the first window above RAM. Consecutive, they need no list: a statement carries only their count k = ⌈words/h⌉ (trace::advice_window_count), which covers what the host supplied, an untouched word's tuples cancelling. check_memory_windows asks only 2^29/h + k ≤ 2^30/h, the top of the address space, and ZERO_WINDOWS ids stay below 2^29/h (memory.md §3).
  • No advice, no region. Then k = 0 and there is no shard; guest_sdk::advice on such a run is a fatal emulator::EmuError::OutOfBounds.
  • Not read-only. A store there is an ordinary store. Refusing it would need a space selector and a gate on three families' memory path, and would buy nothing: advice is unbound either way.

What a guest owes. A proof says that some advice exists under which the program, given the public input, published the journal; advice that changes the journal unchecked is a value the prover chose. The check is against something the proof binds: a commitment in the public input (guests/public-io, at toy scale, with a position-weighted checksum standing in for a hash), or one the journal publishes. revm-block-stateless publishes the root of the payload it validated and holds its witness to that payload by hashes (ethereum.md §4).

7 The guest's view#

A guest reaches the regions with loads and stores at the constants::guest_memory constants, through guest_sdk::public_input, guest_sdk::commit and guest_sdk::advice, none of which issues an ecall; ecall-abi.md §7 is the API and the guest program manual the walkthrough. Nothing is published at exit, so a guest that panics has published what it committed, and its run is proved like any other.

8 Cost#

  • No address space, transcript message, tag, challenge or statement field; no gate elsewhere.
  • Two 2^12-row shards a statement, five committed columns between them; one h-row shard of three columns per advice window.
  • The native verifier: two 4,096-point multilinear evaluations, 4,095 multiplications each. A recursion node's cost follows the payload instead: it evaluates the payload's words alone, times 1 − r_j for each variable above them (verifier_core::chain::public_value).
  • The guest: nothing at exit; a byte store per journal byte and a word store per commit.

9 Limits#

  • 16,380 bytes each, and no larger window (§2). A journal that grows with the execution has no fixed bound: the mini-block binary's, a 13-byte record plus return data per transaction (ethereum.md §3), holds at most 1,255 transactions, and one record can exceed it. Large outputs belong behind a digest (the stateless binary's journal is 43 bytes), large inputs in advice.
  • The journal is the window's whole final contents. Anything but a length of at most 16,380, that many bytes, then zeros, matches no statement: the executor refuses an oversized length (EmuError::JournalTooLong), and a nonzero byte past it fails step 10c. commit keeps that form; a guest writing the window directly must.
  • Nothing orders the journal's writes, and nothing forces a guest to read its input. The proof binds a window's contents, not its accesses.
  • Read the exit status first. It is x10's final value (memory.md §4): a failed run, a panic included, has a verifying proof and a journal too (§7).
  • A deployed contract fixes both lengths, a decider key being per shape (recursion.md §9).

Auditors/Proof system

The GKR engine

Normative specificationdocs/spec/gkr.mdView as Markdown

The layered-circuit model every family circuit is written in, the artifact that carries one, its laws, and the backward pass reducing a circuit's outputs to claims on its committed columns at one point, which the shard's opening discharges (proof.md §5).

crates/constraints is §1–§4; crates/gkr-verify is §5's verifier half and the verifier's helpers for the memory argument (memory.md §3, §4) and LogUp (lookup.md §2, §8). Both are no_std, as verifier-core and the recursion guest build on them. crates/gkr, std and rayon, is the prover half and re-exports gkr-verify. crates/checker enforces §4.2–§4.3 again (circuits.md §3).

1 The layer model#

Layer k, 0 ≤ k ≤ N, N ≥ 1, is w_k columns of n_k variables, indexed as primitives.md §6 fixes. Layer 0 is the committed columns M, W, S in layout order at n_0 = trace_vars, beside the virtual tables the artifact lists (§2.1), which count in no width. Gate list k reads layer k and writes layer k + 1; the top, layer N, is exactly the outputs. A list is row-wise, n_{k+1} = n_k, or halving, n_{k+1} = n_k − 1.

A halving list halves each column of its layer: it writes w_k columns by halving shapes (§3) reading layer-k columns at both children — child 0 is rows [0, h), child 1 rows [h, 2h), h = 2^{n_k−1}, the child bit being the highest variable. An entry may read any column, as a fraction tree's numerator reads its denominator (lookup.md §6), but every column is read (§4.2). Only halving lists hold halving shapes; a halving list is never list 0, has no cached or enforcing entries and needs n_k ≥ 1. Every relation has the one template checker dump prints:

producing, row-wise   L{k+1}[j](x) = Σ_y eq(x, y)·G(layer k at y)
producing, halving    L{k+1}[j](x) = Σ_y eq(x, y)·G(layer k at (y, 0) and (y, 1))
enforcing             0 = G(layer k at y)   for every y ∈ {0,1}^{n_k}

2 Addresses#

constraints::PolyAddress names every polynomial; dumps use its Display notation:

variant notation read by
Memory(i), Witness(i), Setup(i) M[i], W[i], S[i] committed columns list 0, relations, lookups
Virtual(kind) V[row], … virtual tables, §2.1 the same, if virtuals lists it
Inner { layer, offset } L{k}[j] column j of layer k ≥ 1 list k
Cached { layer, offset } C{k}[j] cached entry j of list k, §3.1 list k
Scratch(i) scratch[i] an intermediate of the flat relation list, §4 relations

The scratch bijection maps each scratch[i] to one L{k}[j], covering every inner column once. A committed value needed above layer 1 is carried up by copy gates. M, W and S differ in when they are bound (memory.md §8).

2.1 Virtual tables#

A virtual table is a closed form, evaluated per row by gkr_verify::virtual_at_row and at a point by virtual_at_point, never materialized, committed or claimed. Each form is its table's multilinear extension, so the verifier evaluates what the prover sums (crates/gkr/tests/{lookup,ram_live}.rs check all but V[row]). Wire form: a u32, in table order from 0.

kind notation value at row y closed form at (y_0, …, y_{n−1})
RowIndex V[row] y Σ_{j<n} 2^j·y_j
RamLive V[ram_live] 1 if y ≥ 2^14, else 0 1 − Π_{14≤j<n} (1 − y_j); 0 if n ≤ 14
Range19 V[range19] y mod 2^19 Σ_{j<min(19,n)} 2^j·y_j
Range16 V[range16] y mod 2^16 Σ_{j<min(16,n)} 2^j·y_j
Xor8A V[xor8_a] a = y mod 2^8 Σ_{j<8} 2^j·y_j
Xor8B V[xor8_b] b = ⌊y/2^8⌋ mod 2^8 Σ_{j<8} 2^j·y_{j+8}
Xor8Out V[xor8_out] a ⊕ b Σ_{j<8} 2^j·(y_j + y_{j+8} − 2·y_j·y_{j+8})

14 is constants::memory::RAM_LIVE_BIT (memory.md §3); the range and XOR8 kinds are channel tables (lookup.md §3). Xor8Out's form is multilinear because y ⊕ z = y + z − 2yz is.

3 Gate shapes#

constraints::GateDef is a closed enum. A coefficient is Coeff::Literal(Fr) or Coeff::Challenge(slot), a constants::challenge_slot read from the pass's ExternalChallenges, of degree 0.

tag variant value
0 Linear { terms, constant } Σ c_i·x_i + c_0
1 Product { coeff, left, right } c·x·y
2 MaskIntoIdentity { input, mask } x·m + (1 − m)
3 AffineProduct { left, left_constant, right, right_constant } (Σ a_i·x_i + a_0)·(Σ b_j·y_j + b_0)
4 TreeProduct { input } x(·,0)·x(·,1)
5 Quadratic { constant, linear, products } c_0 + Σ a_i·x_i + Σ b_j·y_j·z_j
6 TreeCross { left, right } p(·,0)·q(·,1) + p(·,1)·q(·,0)

Quadratic spells degree-2 relations, such as a·b + c·d − e·f, that no product of affine forms does. The kernel, gkr_verify::eval_gate, takes one value per operand in GateDef::operands order, a halving shape's each at child 0 then child 1, and is the semantic authority. Both passes reach it through gkr_verify::ResolvedList, crates/checker calls it over the relations, and verifier_core::tape transcribes it for the recursion nodes (recursion.md §7).

3.1 Cached entries and the degree ceiling#

A cached entry C{k}[j] = H is a sub-expression of row-wise list k over its layer's columns, not another cached entry, substituted into the gates of its list naming it, with no table, claim or width. The prover evaluates H at every round node and never binds it: a bound table is the extension of H's values, which for a degree-2 H is not H of the extensions. No registered circuit has one. CircuitArtifact::inline_cached writes a Product with one Linear cached factor as an AffineProduct and refuses any other reference; both prove the same bytes.

Degree is read from the shape after substitution — a column or virtual table 1, a challenge 0, C{k}[j] its expression's, a halving shape 2, a Quadratic its widest term — and validate holds every gate, cached entry and relation to at most 2, so a higher relation is split across layers. With eq multilinear, every round polynomial is then a cubic (§5.3).

4 The circuit artifact#

constraints::CircuitArtifact holds a circuit twice: as layered gates, which the engine proves, and as a flat relation list over M, W, S, V and scratch, which the row-local checks read (circuits.md §3). Law 4 makes them one constraint set. In wire order:

CircuitArtifact = (format_version = 1, coefficient_encoding = 0, trace_vars ≤ 30,
                   memory, witness, setup: [name], virtuals: [(VirtualKind, name)],
                   layers: [LayerSpec], relations: [Relation], lookups: [LookupExpr],
                   scratch: [(name, L{k}[j])], outputs: [L{N}[j]],
                   padding: (row: [Fr], zero_row_valid: bool))
LayerSpec       = (halving, num_vars, width,
                   cached:    [(name, C{k}[j], GateDef)],
                   producing: [(relation, L{k+1}[j], GateDef)],
                   enforcing: [(relation, GateDef)])
Relation        = (name, output: Option<scratch index>, GateDef)
LookupExpr      = (name, channel, selector: PolyAddress, tuple: [GateDef])

validate holds the first three to those values and every name to non-empty [a-z0-9_], unique in the artifact; names mean nothing to the engine. Encoding 0, COEFFICIENT_ENCODING_CANONICAL_LE, is every Fr canonical 32-byte little-endian, and 30 is MAX_TRACE_VARS. outputs orders the top layer as OutputClaims lists it; a relation with an output defines that slot, one without is enforcing; lookups are lookup.md §1's.

4.1 Wire form#

postcard over §4's tuples, hand-written serde: a u32 is a varint, a u8 tag and a bool a byte, an Option a tag byte, a sequence a varint count then its elements, a name a str, an Fr its 32 canonical bytes.

PolyAddress  (tag u8, a u32, b u32): 0 M, 1 W, 2 S, 5 scratch (a = index); 3 V (a = kind);
             4 L, 6 C (a = layer, b = offset); unused fields 0
Coeff        (tag u8, slot u32, value Fr): 0 literal (slot 0), 1 challenge (value 0)
GateDef      (tag u8, split u32, coefficients [Coeff], operands [PolyAddress] in operands() order)
  0 Linear            split 0  c_1..c_t, c_0                 x_1..x_t
  1 Product           split 0  c                             x, y
  2 MaskIntoIdentity  split 0  —                             x, m
  3 AffineProduct     split t  a_1..a_t, a_0, b_1..b_u, b_0  x_1..x_t, y_1..y_u
  4 TreeProduct       split 0  —                             x
  5 Quadratic         split t  c_0, a_1..a_t, b_1..b_u       x_1..x_t, y_1, z_1, …, y_u, z_u
  6 TreeCross         split 0  —                             p, q

CircuitArtifact::from_bytes refuses a format_version other than 1 before decoding the rest, postcard not being self-describing; refuses an unknown tag, a nonzero unused field, a gate with counts its shape lacks and a non-canonical Fr; re-encodes and compares, as postcard admits overlong varints and trailing bytes; never panics or reserves what a declared length asks; and checks no law.

4.2 The laws#

CircuitArtifact::validate runs once where an artifact is built or loaded, never per proof: each constraints constructor panics on a refusal, and verifier_core::VerifyingKey::check applies it to a key's circuits, for prover and verifier (proof.md §7). checker::check_laws enforces Laws 1–4 and the lookup rules again, sharing no code with crates/constraints/src/laws.rs (circuits.md §3).

  1. Locality. Every operand of list k is in range and readable at layer k (§2): a V only if listed, a C{k}[j] only one of list k's own, from a producing or enforcing gate.
  2. Derived width. A list's stored width is its producing count, entry j writes L{k+1}[j], and its stored num_vars is n_k, or n_k − 1 if halving.
  3. Top layer. outputs is a permutation of L{N}[0..w_N).
  4. Single source of truth. Relations and gate entries correspond one to one, a producing entry's relation defining the slot the bijection maps to its output, an enforcing entry's none, and each pair is one polynomial, scratch read through the bijection and cached entries substituted: validate compares normalized expansions, checker evaluations at random points.

validate also refuses, each a ConstraintError naming what broke: §4's bounds, no gate list, padding.row not w_0 long, a virtual kind listed twice, §1's halving rules, degree above 2, a relation reading anything but M, W, S, listed V and existing scratch, a scratch list that is no bijection onto the inner columns or not defined once each, a slot outside constants::challenge_slot, and a relation constructed and then dropped — an inner column below the top the list above never reads, a cached entry no gate names, an enforcing gate whose expansion is zero. Reads are decided on normalized expansions: x − x and 0·x read nothing.

The lookup rules. A lookup's channel is in constants::lookup_channel; its tuple is one expression on a range channel, else 1 to lookup_channel::MAX_TUPLE (7), as wide as its channel's other lookups'; its selector is an in-range committed column some enforcing gate of list 0 holds to booleanity (x − x² up to normal form); and each expression is Linear over in-range committed columns and listed virtual tables, with literal coefficients, unit and constant-free above position 0 (lookup.md says what each protects).

4.3 The padding contract#

The engine gates nothing, an enforcing gate being a zerocheck over the whole cube, so a family switches relations off with its own columns (memory.md §2). On padding.row, a committed row, the row-local scratch values, those of producing relations not at or above a halving shape, make every row-local enforcing relation vanish at every challenge value and row index; zero_row_valid says whether the all-zero row does too. The product-tree clause: where shards have inactive rows, every column the first halving list reads is 1 on padding.row, so padding leaves each product unchanged; the RAM window families (memory.md §3) and the columns a TreeCross reads (lookup.md §6) are exempt. This is completeness, not soundness: a cheating prover's padding rows are its family's gates' business. Nor is padding.row the row a prover writes, multiplicities and setup columns differing; no prover or verifier reads it, and checker::check_padding and checker::check_padding_identity test it.

5 The backward pass#

gkr::forward materializes every layer from the committed columns; gkr::prove proves those values as they stand, one sumcheck::SumcheckProof per transition; gkr_verify::verify replays the schedule, checking, from OutputClaims, one table per output, to BaseClaims or a GkrError. gkr::self_check, naming the first failing gate, row and relation, and gkr::explain_self_check, listing that row's operands, are a debugging hook costing a second forward pass (tools.md §3). Rayon splits rows and row pairs, never lists or rounds: proofs do not depend on the thread count.

5.1 What the caller owes#

  • The base is bound into the transcript before prove or verify, which absorb none of it (proof.md §4 binds a shard's commitments).
  • Each challenge is drawn after every committed column its gates reach is bound, or is derived: a fixed function of such challenges and of statement data bound before them, computed by the verifier. That suffices for GKR; the memory argument needs more (memory.md §8).
  • The artifact has passed validate (§4.2) and is not checked again; on a lawless one the engine may panic, and verify may accept.
  • The prover's inputs have the artifact's shape; it checks none, nor that its values satisfy the gates. Soundness is verify's alone and a cheating prover runs none of this code, so a bad input costs the honest prover only a panic or a failing proof.

5.2 The transcript schedule#

prove and verify run these steps and end in one sponge state; the tags are transcript.md §5's. p is the claim point, v_j the claim on column j of the layer the next list writes.

step op tag message
O1 absorb GKR_OUTPUTS the output tables in output-map order, rows in index order: one message of w_N·2^{n_N} scalars
O2 squeeze ×n_N GKR_OUTPUT_POINT p = r, r_i binding variable i; v_j = tables[i](r) for outputs[i] = L{N}[j]
L1 squeeze GKR_BATCH λ; the claim is c = Σ_j λ^j·v_j
L2 ×n_{k+1}: absorb, squeeze SUMCHECK_ROUND, SUMCHECK_CHALLENGE a round's cubic, then ρ_i, binding variable i
L3 absorb GKR_LAYER_CLAIMS row-wise: L{k}[j](ρ) per j in offset order, layout order at k = 0; halving: L{k}[j](ρ,0), L{k}[j](ρ,1) per j
L4 squeeze, halving only GKR_CHILD τ; p = (ρ, τ); v_j = L{k}[j](ρ,0) + τ·(L{k}[j](ρ,1) − L{k}[j](ρ,0))

L1–L4 run for k = N − 1 down to 0; after a row-wise list p = ρ and v is L3's message. The base claims are layer 0's, in layout order at one point. Every registered circuit halves to a top with no variables (circuits.md §2), so O2 draws nothing and O1 fixes the roots before λ.

5.3 The layer sumcheck#

Transition k proves c = Σ_{y∈{0,1}^{n_{k+1}}} eq(p, y)·S_k(y), where

row-wise   S_k(y) = Σ_j λ^j·G_j(layer k at y) + Σ_e λ^{w_{k+1}+e}·E_e(layer k at y)
halving    S_k(y) = Σ_j λ^j·G_j(layer k at (y, 0) and (y, 1))

G_j writes L{k+1}[j] and E_e, the list's e-th enforcing gate, claims 0: enforcing gates are zerochecks sharing the descending point and its batch. The rounds are primitives.md §7's cubics, run from c, one per variable of layer k + 1, a halving list's two children being separate tables. After L3 the verifier checks claim = eq(p, ρ)·S_k(values), layer-k operands taking L3's values, virtual tables their closed form at ρ, cached entries their expression; with n_{k+1} = 0 there are no rounds and the check is c = S_k(values). A zero claim is legal. gkr::prove_sumcheck and gkr_verify::verify_sumcheck run L2.

5.4 Why it is sound#

Each challenge is drawn after what it protects:

  • r after the outputs, or a prover predicting r claims another table agreeing with the true one there.
  • λ after the claims and p. If some v_j is not the true v̂_j, or some E_e is nonzero on the cube, Σ_j λ^j·(v_j − v̂_j) − Σ_e λ^{w_{k+1}+e}·Ê_e(p) is a nonzero polynomial in λ of degree below w_{k+1} + |E_k|; Ê_e, the extension of E_e's values, is fixed before p is drawn and vanishes there with probability at most n_{k+1}/|Fr|.
  • ρ_i after round i: a wrong cubic agrees with the true one there with chance ≤ 3/|Fr|.
  • τ after both children: a wrong pair's line meets τ ↦ L{k}[j](ρ, τ) in at most one point.

Summed over a registered circuit's transitions at its default height, these stay under 2^14/|Fr|. The random-oracle assumption is architecture.md's.

5.5 Shapes and errors#

Transition k carries n_{k+1} rounds and w_k claims, 2·w_k if halving, so a proof's shape is the artifact's alone (wire form: proof.md §9). verify checks, in order and before touching the transcript, and on a validated artifact never panics on proof or claim data:

GkrError when
MissingChallenge { slot } a gate names a slot not supplied
OutputShape OutputClaims mismatches the output map in count or variables
ProofShape { layer } layer = N: a wrong transition count; else transition layer, lowest first, has a wrong round or claim count
LayerInconsistency { layer } a round or the final check of transition layer fails

One LayerInconsistency covers a wrong descending claim and a violated enforcing gate alike: a batched sum cannot tell them apart, and the proof spends nothing on it. proof.md §6 maps these errors to its classes.

Auditors/Proof system

Circuits

Normative specificationdocs/spec/circuits.mdView as Markdown

Every shard is proved by its family's circuit, a constraints::CircuitArtifact in gkr.md's model, fixed by the format, the family and the height. This page lists the circuits and their shapes (§1), how one is assembled (§2) and how crates/checker checks one independently (§3); each family's own page specifies its columns, gates and lookups.

1 The registry#

constraints::family_circuit(family, trace_vars) is the base format's registry, constraints::recursion_circuit the recursion format's, and VmConfig::circuit picks one by format (recursion.md §1.1). Each returns a FamilyCircuit, the artifact and its channel specs (lookup.md §11). A verifying key loads only if its circuits are the registry's at its heights (proof.md §7), and the prover registers the same (§2).

Families 0–6 (constants::family) are the execution families, one executed instruction a row (add-sub.md, jump-branch-slt.md, shift-bitwise.md, mul-div.md, memory-ops.md §3, §4, §6); 7–8 and 12–14 the window families, one memory word a row (memory.md §3, public-values.md §4); 9–11 and 15–17 the delegation families, one invocation a row (delegation-circuits.md §2 to §7, by id); 18–22 the recursion format's (recursion.md §2 to §6).

Shapes at the default height 2^n (constants::family::DEFAULT_HEIGHTS): committed columns, enforcing gates, obligations per channel (TIMESTAMP/RANGE16/GENERIC/DECODER/XOR8), row-wise gate lists (the halving ones are n), inner columns, artifact bytes, and a base-format shard proof's bytes, proof.md §9's layout over the shape:

id family n M W S gates lookups row-wise inner bytes proof
0 ADD_SUB_LUI_AUIPC 22 27 35 7 63 10/4/0/1/0 5 314 72,064 64,764
recursion format 22 27 39 7 75 10/4/0/1/0 5 314 79,077 —
1 JUMP_BRANCH_SLT 22 21 44 10 42 8/11/2/1/0 5 392 76,980 69,436
2 SHIFT_BITWISE 22 21 61 10 48 8/24/6/1/0 6 478 102,837 76,644
3 MUL_DIV 20 21 54 9 54 8/16/2/1/0 6 444 92,640 67,412
4 MEM_WORD 22 31 24 7 33 12/5/0/1/0 5 314 60,383 63,836
5 MEM_SUBWORD 22 31 55 10 53 12/22/1/1/0 6 472 98,846 76,196
6 ATOMICS 20 26 54 9 46 10/19/6/1/0 6 472 101,593 68,468
7 INIT_TEARDOWN 22 2 0 1 0 — 1 46 3,907 36,316
8 ZERO_WINDOWS 22 2 0 0 0 — 1 46 3,418 36,284
9 KECCAK_F 18 208 1,556 0 385 0/210/0/0/1,020 11 5,490 1,900,468 381,100
10 POSEIDON2 8 100 4,092 0 4,248 — 193 2,020 2,056,361 664,780
11 FR_ARITH 8 104 2,576 0 2,701 — 6 142 1,063,214 266,292
12 PUBLIC_INPUT 12 3 0 0 0 — 1 26 2,455 12,556
13 PUBLIC_OUTPUT 12 2 0 0 0 — 1 26 2,338 12,524
14 ADVICE_WINDOWS 22 3 0 0 0 — 1 46 3,535 36,316
15 MOD_MUL 16 104 221 0 125 0/274/0/0/0 10 2,244 550,391 135,220
16 SHA256_COMP 18 104 520 0 119 0/114/0/0/336 10 2,802 845,456 189,988
17 EC_ADD 16 392 1,028 0 637 0/1,110/0/0/0 12 8,772 2,350,670 434,916
18 FIELD_WINDOWS 20 2 0 0 0 — 1 42 2,758 —
19 FR_OP 20 31 31 0 44 0/36/0/0/0 7 370 89,741 —
20 P2_FIELD 18 45 382 0 372 0/58/0/0/0 7 392 294,425 —
21 FIELD_IO 18 43 39 0 24 0/70/0/0/0 8 650 164,713 —
22 FQ_OP 20 48 73 0 38 30/50/0/0/0 7 630 158,326 —

Heights. Both registries return None above MAX_TRACE_VARS = 30, and below the floor lookup.md §3 derives from the family's channels: 19 with TIMESTAMP, else 16 with RANGE16 or XOR8, else 0. A height changes trace_vars, each list's variable count and the number of halving lists, one per variable and as wide as the outputs, and no gate below them: at 2^20 ADD_SUB_LUI_AUIPC has 298 inner columns, 70,974 bytes and a 57,196-byte proof.

Shared circuits. The registries agree on families 1–17; the recursion format's ADD_SUB_LUI_AUIPC is add_sub::recursion_artifact (add-sub.md §2). PUBLIC_OUTPUT's circuit is ZERO_WINDOWS' and ADVICE_WINDOWS' is PUBLIC_INPUT's, byte for byte at one height, and FIELD_WINDOWS' is the zero window at a stride of one cell, all constraints::memory constructors (memory.md §3). Every other family's is its own module's artifact.

2 How a family circuit is assembled#

layer 0        M ‖ W ‖ S in layout order, beside the V tables' closed forms
gate list 0    memory leaves: the read side, then the write side, each padded to a power of
                 two with the literal 1
               per channel, in spec order: (−mult, T + g), then (1, E_l + g) per lookup,
                 then (0, 1) up to a power of two                      (lookup.md §6)
               every enforcing gate
lists 1 … r    row-wise: each tree combines sibling nodes, a product by a·b, a fraction by
                 (n_a·d_b + n_b·d_a, d_a·d_b); a tree already at one node is copied up
lists r+1 …    halving, one per variable: TreeProduct on a product, TreeCross (num) and
                 TreeProduct (den) on a fraction
top            no variables: read_root, write_root, then (num, den) per channel

r is the largest tree's depth, so the circuit has r + 1 row-wise lists; every registered circuit, POSEIDON2 included, ends in a top with no variables. crates/constraints/src/build.rs assembles it, writing the flat relation list and an all-zero padding row, zero_row_valid read off the gates' constants, and validating (gkr.md §4). constraints::memory::assemble gives it the product trees and lookup::channel_trees' fraction trees (lookup.md §11), then runs memory::check_memory (memory.md §8) and lookup::check_discharge: a constructor panics on a refusal, so every circuit that exists has passed them. Its callers:

  • memory::frame_with_channels_artifact(queries, trace_vars, FamilySpec), the execution families: memory.md §2's frame over memory::frame_queries(family), then the family's witness columns after the frame's w + 3, setup columns from S[0], virtual tables, enforcing gates after the frame's, lookups after its 2w gap obligations, and a non-empty channel list;
  • the window constructors (memory.md §3);
  • the delegation and recursion families, every gate in list 0, from constraints::delegation's shared columns, leaves and gates (delegation-circuits.md §1) — but POSEIDON2, which builds its own lists (delegation::Assembly): 192 row-wise lists of rounds beside its product trees, the last holding three gates on the output lanes.

Beyond the frame, each execution family has m_pc as the row's liveness and every other mask held to m_pc times the kinds making that query (<q>_mask_rule, memory.md §2); its decoded row as W columns, bound by decode_row to its table at the row's pc, and decoded_mask_bits, the mask as boolean kind bits, one-hot by the table's domain (lookup.md §10, program.md §6); a next_pc_rule (memory.md §5); a bound on each register value it writes (memory-ops.md §5); and channels ordered TIMESTAMP, RANGE16, GENERIC if read, DECODER.

prover::family_fill(family) is the prover's side: a prover::Fill writes a shard's committed columns but the multiplicities, which trace::build_multiplicities counts. prover::register pairs fill and circuit for each family of a VmConfig (ProverError::Unregistered if either is missing).

3 Checking a circuit independently#

crates/checker's validators enforce the rules again in code sharing nothing with crates/constraints/src/laws.rs, never calling validate. They evaluate a gate only through the kernel gkr_verify::eval_gate (gkr.md §3), so they re-read the rules, not the gates' meaning. Sampled checks use eight pseudo-random points from fixed seeds.

checks
check_laws (check_law1 … check_law4) the four laws, then the lookup rules (gkr.md §4); Law 4 and selector booleanity by evaluation, where validate compares expansions
check_padding, check_padding_identity the padding contract and its product-tree clause, fraction trees exempt
check_lookup_discharge lookup.md §11's discharge rule, gating and compression re-derived
violated_relations, violated_lookups a witness row's row-local relations and range obligations
channel_sums, check_channel_roots each channel's sum and denominator product, folded row by row rather than by a tree, naming every tuple no table row holds; then the circuit's root pairs against them
memory_roots the two roots as products over the rows the halving phase reads
memory_columns_from_log, frame_witness_from_log an execution family's frame columns from the memory event log, where trace builds them from a shard's rows

They do not re-implement check_memory, the copower rule (lookup.md §11), or validate's other construction rules, the degree ceiling among them.

checker::TamperHarness re-proves a statement with witness cells or boundary scalars changed, as an honest prover would prove the changed witness — each channel's multiplicities recounted unless one is what changed or the changed tuple is in no table, changed M columns recommitted in a fresh global commit phase, every shard re-proved — then verifies a shard or the block and asserts the refusal's class (a Lookup's channel too), or that a change breaking nothing verifies. It relies on the prover checking nothing (gkr.md §5), runs on the archived path (streaming.md §6), and carries the delegation anchor's forgeries (checker::assert_anchor_twins_refused, delegation.md §5).

A dump (checker::dump, CLI in tools.md §4) prints the columns by address and name, each list's gates in gkr.md §1's template with their relations, the flat relations over scratch[i], the scratch bijection, outputs, lookups and padding row. Relations are numbered list by list, producing before enforcing; a producing one is define_<column>, an enforcing one bears its gate's name; a node is named for its tree and layer (range16_3_1_num, read_root), a leaf for what it holds (write_pad_0, rd_hi_range_den). A literal below 2^32 prints in decimal, p − k for such a k as -k, any other as 0x and 64 big-endian hex digits; a challenge as its constants::challenge_slot::NAMES entry.

Auditors/Proof system

The memory argument

Normative specificationdocs/spec/memory.mdView as Markdown

Offline memory checking over a whole statement. Each shard's circuit outputs the product of its read tuples and of its write tuples; the verifier checks, once per statement, that all reads times the register and pc finals equal all writes times their initial values. RAM is initialized by window families over fixed address windows; registers and the pc have no rows. The section numbers are the ones the code cites.

1 The tuple#

T(AS, ADDR, TS, VAL) = γ_M + AS + α_addr·ADDR + α_ts·TS + α_val·VAL

The parts are in the order of constants::memory::{PART_AS, PART_ADDR, PART_TS, PART_VAL}. AS, an address-space tag (execution-trace.md §2), is unweighted; a RAM address is a 4-aligned word's byte address. γ_M, α_addr, α_ts, α_val are constants::challenge_slot slots 1–4, MEM_GAMMA to MEM_ALPHA_VAL, drawn once per statement after everything §6.1 lists; slot 5, MEM_WINDOW_CONSTANT, is derived per window shard by the verifier and never read from a proof (§3.3). A gate coefficient is one literal or one slot (gkr.md §3), so α_ts·4·cycle is the term (α_ts, cycle) four times.

Every memory artifact outputs its read product at outputs[READ_ROOT = 0] and its write product at outputs[WRITE_ROOT = 1] (constants::memory), before any channel's roots (lookup.md §6). All tuples of a statement form one multiset, over REG, RAM and PC, one anchor space per delegation type, where a request meets its invocation (delegation.md §5), and the recursion format's FIELD cells (recursion.md §2.1).

2 An execution family's memory subtree#

2.1 The frame columns#

A row of an execution family is one cycle; its accesses are queries, each a read and a write at one address, the write at 4·cycle + Δ (execution-trace.md §1). The query table is constraints::memory::{FRAME_NAMES, FRAME_SPACE, FRAME_DELTA}:

id query space Δ
0 pc PC 0 address 0; reads pc, writes next_pc
1 rs1 REG 1 read-only; an ecall's a7
2 rs2 REG 2 read-only; an ecall's a0
3 load RAM 2 read-only; a load's word
4 ram RAM 3 a store's or an atomic's word
5 rd REG 3 the x0 rule (§2.4)
6 deleg the row's 3 a delegation request's mirror (delegation.md §5)

A family's frame is exactly the queries its instructions make (execution-trace.md §4), in table order: constraints::memory::frame_queries, which crates/trace/tests/memory.rs holds to the union over all 59 instructions. A missing query would leave an instruction's written value unconstrained. No instruction routed to ADD_SUB_LUI_AUIPC touches RAM, and ATOMICS keeps every RAM access at Δ = 3, lr.w included. Window and delegation families have no frame (§3.3; delegation-circuits.md §1).

family queries, in slot order w leaves a side
ADD_SUB_LUI_AUIPC pc rs1 rs2 rd deleg 5 8
JUMP_BRANCH_SLT, SHIFT_BITWISE, MUL_DIV pc rs1 rs2 rd 4 4
MEM_WORD, MEM_SUBWORD pc rs1 rs2 load ram rd 6 8
ATOMICS pc rs1 rs2 ram rd 5 8

Columns are addressed by slot s, a query's position in its family's list:

M[0]                  cycle
M[1 + 5s + f]         slot s's <q>_mask, <q>_addr, <q>_read_ts, <q>_read_value, <q>_write_value
M[1 + 5w]             deleg_space, in the one frame holding deleg
W[s]                  <q>_gap_hi, for s < w
W[w], W[w+1], W[w+2]  rd_inv, rd_is_zero, rd_selected

That is 1 + 5w M columns, plus deleg_space, and w + 3 W columns, the family's own following (circuits.md §2). One deleg query serves every delegation type, so its space is the value of deleg_space, an M column the family pins to its type selectors: a leaf may read no W column (§8). The honest fill (trace::build_memory_columns, trace::build_frame_witness, over a shard's trace::RowSlice) sets a mask to 1 where the row is live and has the query, and every column of an absent query or a padding row to 0.

A frame holds a mask only to booleanity, so on the frame alone a padding row's rd query could rewrite x10, the exit status, after the exit row, and a live row could drop a query or carry one its instruction lacks. Every execution family makes m_pc the row's liveness and its decoder lookup's selector (lookup.md §10), and holds each other mask to m_q = m_pc·uses_q (its <q>_mask_rule gates), uses_q the sum of the row's kind and ecall-type selectors that make the query.

2.2 The leaves#

For the query at slot s with mask m, space AS and in-cycle slot Δ:

read_<q>    m·T(AS, addr, read_ts, read_value) + 1 − m
write_<q>   m·T(AS, addr, 4·cycle + Δ, write_value) + 1 − m

Each is one flat Quadratic of gate list 0, built from the unmasked tuple, a Linear whose AS and Δ terms sit on m (constraints::memory::read_tuple is the read one): constant 1; linear terms (γ_M, m), (−1, m), (AS, m) and, on the write side, (α_ts, m) Δ times; every other term multiplied by m, as is deleg's AS, the product (1, deleg_space, m). At m = 0 a leaf is 1 whatever its columns hold, at m = 1 the tuple, and it is one or the other only at a boolean m (§2.4).

2.3 The product#

Each side is padded to w rounded up to a power of two with read_pad_<i> and write_pad_<i>, the literal 1, reading no column. Row-wise Product lists reduce each side to one value a row, and trace_vars halving lists of TreeProduct multiply the rows (gkr.md §1), so the two roots are the products of the shard's read and write tuples. A padding row has every mask 0 and so every leaf 1, the padding contract's product-tree clause (gkr.md §4). The family's channel trees share the layers (circuits.md §2).

2.4 The gadgets every execution family carries#

Gate list 0's first enforcing gates, in this order, and the circuit's first 2w obligations:

<q>_mask_boolean       m − m·m = 0                       every query
<q>_writes_back        write_value − read_value = 0      rs1, rs2 and load, where held
rd_is_zero_inverse     addr·rd_inv + z − m = 0           on rd; z = rd_is_zero
rd_is_zero_at_nonzero  addr·z = 0
rd_is_zero_boolean     z − z·z = 0
rd_write_masked        write_value − sel + z·sel = 0     sel = rd_selected

gap_hi_<q>   TIMESTAMP, selector m:   hi                                   hi = <q>_gap_hi
gap_lo_<q>   TIMESTAMP, selector m:   4·cycle + δ_q − read_ts − 2^19·hi        δ_q = Δ − 1; δ_pc = −4
  • Booleanity. At m = −1 a pc query's leaves are each −T(REG, …): one sign flip a side, so the products balance and the pc access reads as a register access.
  • Write-back. Without it a read of x0 could write 5 there.
  • x0. The first two rd gates (constraints::gadgets::is_zero) make z = m·[addr = 0] and the last write_value = (1 − z)·sel: every write to x0 writes 0, whatever the family computed into sel, and with write-backs and x0's init 0 every read of it returns 0. The boundary's final x0 = 0 (§4.1) pins only its last write: a write of 5, a read of 5 and a write of 0 would otherwise balance.
  • Gap. Both chunks below 2^19 put gap = 4·cycle + δ_q − read_ts in [0, 2^38), so read_ts < 4·cycle + Δ as integers, every timestamp being a canonical integer by §4.2's count. The pc query's δ = −4 puts a row's pc write at least 4 after the one it reads, so consecutive rows' timestamps never interleave (§9). The frame's construction asserts two obligations per query.

3 RAM windows#

3.1 Geometry#

Window w at height h = 2^n covers the bytes [4h·w, 4h·(w + 1)), its row y being the word at 4h·w + 4y; the windows tile [0, 2^32) from 0. Ordinary RAM is [RAM_ORIGIN, ADVICE_ORIGIN) = [2^16, 2^31) (trace::in_ram), ending where window N = 2^29/h begins (verifier_core::advice_first_window). Window 0's rows y < 2^14 (constants::memory::RAM_LIVE_BIT) lie below RAM_ORIGIN at every height, and INIT_TEARDOWN masks them (§3.3).

3.2 The window families#

A window family's shard initializes and tears down one window; its rows are addresses.

region family id windows init value
[0, 0x8000) none: a hole
[0x8000, 0x10000) PUBLIC_INPUT, PUBLIC_OUTPUT 12, 13 2 and 3 at their pinned 2^12 the statement's input; 0 (public-values.md §4)
[RAM_ORIGIN, 4h) INIT_TEARDOWN 7 0, one shard S[0], the image column
[4h, 2^31) ZERO_WINDOWS 8 the listed w_1 < … < w_k in [1, N − 1] 0
[2^31, 2^32) ADVICE_WINDOWS 14 N … N + k_a − 1 M[2], bound to nothing (public-values.md §6)
FIELD cells FIELD_WINDOWS 18 0 … k_f − 1 0 (recursion.md §2.2)

INIT_TEARDOWN, ZERO_WINDOWS and ADVICE_WINDOWS share the window height h, 2^22 by default (§3.5). An unlisted RAM window is initialized by nothing. Every statement proves window 0 and, in practice, the stack's window N − 1, the initial sp being ADVICE_ORIGIN: two h-row shards however small the program, besides the public pair.

3.3 The artifacts#

INIT_TEARDOWN   image_window_artifact   M[0] teardown_ts, M[1] teardown_value, S[0] init_value
  read    live·(WC + α_addr·4·row + α_ts·M[0] + α_val·M[1]) + 1 − live     live = V[ram_live]
  write   live·(WC + α_addr·4·row + α_val·S[0]) + 1 − live
ZERO_WINDOWS, PUBLIC_OUTPUT    zero_window_artifact     M[0], M[1]
  read    WC + α_addr·4·row + α_ts·M[0] + α_val·M[1]        write   WC + α_addr·4·row
PUBLIC_INPUT, ADVICE_WINDOWS   value_window_artifact    M[0], M[1], M[2] init_value
  read    as above                                          write   WC + α_addr·4·row + α_val·M[2]
FIELD_WINDOWS   field_window_artifact: zero_window_artifact over α_addr·row, one cell a row

All are constraints::memory constructors: one leaf a side, then n halving lists; no witness column, enforcing gate or lookup; every init timestamp the literal 0; row is V[row]. V[ram_live] is [y ≥ 2^14], boolean on the cube by construction (gkr.md §2). The window enters only through WC, so one artifact serves every window:

WC = γ_M + RAM + α_addr·4h·w        gkr_verify::window_challenges
WC = γ_M + FIELD + α_addr·h·w       gkr_verify::field_window_challenges

w is verifier_core::shard_window's: 0 for INIT_TEARDOWN, the list's i-th id for ZERO_WINDOWS shard i, 2 and 3 for the public pair, N + i for advice shard i, i for field shard i. An init column is S where program identity binds it and M where it is one execution's, committed before the challenges (§6.1). A window shard has no inactive rows.

3.4 The columns a prover fills#

trace::init_windows(state, h) is ZERO_WINDOWS' list: the distinct ⌊a/4h⌋ over touched words a of ordinary RAM, ascending, without 0. A public or advice word is a RAM tuple too, and a zero window over it would be its second init row. trace::build_init_teardown_columns fills INIT_TEARDOWN and ZERO_WINDOWS, and trace::build_value_window_columns the value windows with their M[2], from the last-access tables (trace::MemoryState):

row y, a = 4h·w + 4y teardown_ts teardown_value
w = 0, y < 2^14 (masked) 0 0
a touched its last write's timestamp its last write's value
a untouched 0 its init value

An untouched row's two tuples are equal and cancel. The image column, program::image_init_column(image, h), has row y = ProgramImage::initial_word(4y): the word assembled byte by byte from file-backed bytes, 0 elsewhere, which is the trace's initial RAM value too. decode_program refuses an image with a file-backed byte at or above 4h (ProgramError::ImageOutsideWindow): it would sit in a zero window, read as 0, bound by nothing.

3.5 The verifier's window rules#

verifier_core::check_memory_windows, step 2 of derive_global_phase, before the global transcript (program::check_memory_windows wraps it):

rule why
INIT_TEARDOWN, ZERO_WINDOWS, ADVICE_WINDOWS at one height h a lower zero-window height would re-initialize image words; an advice height of its own is a grid advice_first_window(h) does not describe
PUBLIC_INPUT, PUBLIC_OUTPUT at PUBLIC_WINDOW_HEIGHT = 2^12 the height places their windows (public-values.md §2)
4h ≥ PUBLIC_OUTPUT_ORIGIN + PUBLIC_WINDOW_BYTES = 0x10000: h ≥ 2^16 on the menu the public windows lie in window 0's masked rows, out of every zero window's reach
one shard each of INIT_TEARDOWN, PUBLIC_INPUT, PUBLIC_OUTPUT (public-values.md §4 for the pair)
one id per ZERO_WINDOWS shard, strictly increasing, in [1, N − 1] disjoint windows; id 0 is unmasked over [0, RAM_ORIGIN); N up is advice
N + k_a ≤ 2^30/h, k_a the advice shard count advice ends by 2^32; it needs no list, starting where the zero ids stop
k_f·h ≤ 2^32 field cells recursion.md §2.2

The first three are verifier_core::window_height, which VmConfig::from_bytes runs too: a config breaking them does not decode.

4 The register and pc boundary#

Registers and the pc have no rows: the verifier multiplies in their initial and final tuples, once per statement. Rows for them would repeat the init tuples in every shard holding them, and a stale read would balance against the copy.

4.1 The boundary scalars#

The statement carries 64 scalars, gkr_verify::BoundaryFinals, absorbed as one MEMORY_BOUNDARY message in this order (verifier_core::boundary_scalars):

positions
0–31 t_0 … t_31 x_r's final timestamp: its last query's write, 0 if never queried
32 t_pc the pc's: the exit row's pc write
33–63 v_1 … v_31 x_r's final value, 0 if never queried

The final values of x0, 0, and of the pc, HALT_PC, are constants, not carried. PublicInputs::from_bytes refuses t ≥ 2^38 or v ≥ 2^32, and verify_global_memory re-checks the timestamps and holds v_10 to the exit status; no other register carries a public value. t_pc is not a cycle count: the pc's timestamps increase but need not be consecutive. trace::build_boundary_finals(state) is the fill.

4.2 The factors and the reconciliation#

W_b = ∏_{r=0}^{31} T(REG, r, 0, 0) · T(PC, 0, 0, entry_pc)
R_b = T(REG, 0, t_0, 0) · ∏_{r=1}^{31} T(REG, r, t_r, v_r) · T(PC, 0, t_pc, HALT_PC)

∏ read roots · R_b  =  ∏ write roots · W_b  ≠  0       over every shard of the statement

entry_pc is the verifying key's (§6.2). gkr_verify::boundary_factors evaluates each tuple through gkr_verify::eval_gate on the circuits' own tuple gate, read_tuple of pc or rs1, so the boundary and the circuits cannot disagree on the parts; gkr_verify::reconciles is the equation, which verifier_core::verify_global_memory runs once per statement (proof.md §6). A shard's roots are its GKR outputs, held to the statement's entry by its own verification.

The count. Read each query as an edge from its read tuple to its write tuple. Inits are only written and finals only read, so a balanced multiset is paths from inits to finals plus loops. An edge advances the timestamp by an integer in [1, 2^38 + 3] (the gap plus the query's least advance, 4 at the pc and 1 elsewhere), so a loop needs more than p/(2^38 + 3) > 2^215 edges, and a statement has fewer than 2^67 tuples: under 2^32 shards a family (a u32 count), 23 families, at most 2^22 rows (the menu's top), at most 196 tuples a row (EC_ADD's 97 frame words and its anchor, both sides), and 66 boundary tuples. So nothing loops: every path starts at an init at timestamp 0 and ends at a final, and every timestamp on it is an integer below 2^105.

5 Halting#

constants::memory::HALT_PC = 1. The exit row, ecall with a7 = 93, writes next_pc = HALT_PC instead of its fall-through (execution-trace.md §6), and R_b fixes the pc's final value to it. Nothing else writes it: HALT_PC is odd, every other next_pc even, and "odd" is a constraint only where a family makes it one.

  • A family copying the decoded fall-through, which is even, holds next_pc − decoded_next_pc = 0 and needs no bound.
  • JUMP_BRANCH_SLT, the one family computing a pc, range-checks every next_pc it writes even; otherwise a jalr whose rs1 + imm is 1 could write HALT_PC (jump-branch-slt.md).
  • ADD_SUB_LUI_AUIPC writes HALT_PC on its exit row alone (add-sub.md).

HALT_PC is below RAM_ORIGIN, so no decoded-table row claims it and no live row reads it (lookup.md §10). The pc's path therefore ends with the exit row's write, consumed by the final read. With a free final pc every prefix of an execution would balance.

6 Binding#

6.1 What precedes the memory challenges#

The four challenges are squeezed once per statement, at the end of the global transcript (proof.md §2 is the schedule), after everything a tuple or the reconciliation reads, because what is chosen after them can be solved for:

  • every shard's M commitments, every column a leaf may read but S and V (§8);
  • program identity, fixing entry_pc and the image column (§6.2), and the SRS digest, fixing the generic table's S columns (proof.md §3);
  • the shard counts and MEMORY_WINDOWS, the zero-window ids, fixing every window shard's addresses through WC: a list chosen afterwards is a union over up to 2^(N − 1) lists, 2^127 at h = 2^22 and no bound at all at 2^20;
  • io_digest, fixing the public windows' contents (public-values.md §5);
  • last, the 64 boundary scalars: a final value chosen afterwards reconciles any trace, v_r = (target − γ_M − REG − α_addr·r − α_ts·t_r)/α_val.

The roots are not absorbed: each shard's GKR proof binds its own.

6.2 The image column and the entry pc#

Program identity (program.md §8 is the recipe) binds INIT_TEARDOWN's one setup commitment, the image column's, and entry_pc, under PROGRAM_ENTRY. Recomputing identity binds a commitment, not the column a proof reads; the INIT_TEARDOWN shard's batched opening closes that by taking S[0]'s commitment from the verifying key, the list identity is recomputed over (proof.md §5, §7). Without it a statement over another image, with a trace consistent with that image, would verify. Without entry_pc in identity, a key carrying the registered identity beside another entry pc would verify an execution starting elsewhere. Identity binds nothing an execution chooses: no shard count, window list, public or advice word.

7 Range obligations#

A range obligation holds where its selector is 0 or its one expression is below its channel's bound (lookup.md §1, §3). Every circuit bounds a value one way:

  • a 32-bit value v: a witnessed high halfword h and RANGE16 obligations on h and on v − 2^16·h, under the row's selector, and no gate;
  • a result r = e mod 2^32 of an exact 0 ≤ e < 2^33: a witnessed wrap, the gates wrap − wrap·wrap = 0 and e − r − 2^32·wrap = 0, and r bounded as above; a wider carry is a family's own construction;
  • a timestamp gap: two 19-bit TIMESTAMP chunks, no wrap (§2.4). Delegation and recursion families decompose theirs their own way (delegation-circuits.md §1).

8 Construction-time rules#

constraints::memory::check_memory refuses, naming the gate, a memory artifact with:

  1. provenance: a gate or output whose cone both names a memory slot (1–5) and reads a W column, computed forward with two flags a column, so a tuple times a copy of a W column two layers up is refused too;
  2. a root over W: outputs[READ_ROOT] or outputs[WRITE_ROOT] whose cone reads a W column at all, slot or not, which rule 1 does not see;
  3. a memory slot over anything but M, S and V: a gate carrying one reads no W, inner or cached column;
  4. an unconstrained mask: a leaf — a producing Quadratic of gate list 0 with constant 1 and a slot-weighted linear term — whose mask, that term's operand, is an M, W or S column with no m − m·m enforcing gate in gate list 0, or a virtual column but V[ram_live].

It runs beside CircuitArtifact::validate, whose laws it assumes, wherever a memory artifact is built (constraints::memory's assembly panics on a refusal) and in VerifyingKey::check. A W column is committed in a shard's own transcript, after the memory challenges, so a tuple or root over one is chosen after them and balances any trace. M columns precede the challenges and V columns are closed forms; S columns are admitted because they precede them too, bound by identity or, for the generic table, by the SRS digest (§6.1).

9 What the argument rests on#

Both sides of §4.2 are products of linear forms in (γ_M, α_addr, α_ts, α_val), one per distinct tuple, every tuple fixed before those are drawn (§6.1, §8). By Schwartz–Zippel they agree on unequal multisets with probability at most N/p, N < 2^67 (§4.2), and on equal ones §4.2's count gives:

  • One init per address of REG, PC, RAM and FIELD: the 33 boundary inits once per statement, and §3.5's windows, disjoint and of one height. A second init would let a stale read balance. An anchor space has no init: each invocation's answer, stamped 0, starts a path one request long (delegation.md §5).
  • Coverage. Every query lies on a path from an init, so nothing reaches an address no family initializes: the hole [0, 0x8000), where a null dereference does not balance, an unlisted window, a register above x31, a pc address but 0. A query reading its own write would balance with no init; the gap forbids it.
  • Consistency per address: on its one path every read returns the write before it, and the final tuple holds the last.
  • Initial values: the image's, by §3.4's refusal and §6.2's opening of S[0]; 0 in every zero window and the journal; the statement's input in its window (public-values.md §5). Advice is bound to nothing by design (public-values.md §6).
  • Order across rows, shards and families. The pc's path runs from T(PC, 0, 0, entry_pc) through every live row of every execution family, each m_pc = 1 row one edge, to the exit row (§5). That is pc continuity; it orders the rows by their pc writes 4·cycle, which are therefore distinct, so no cycle is proved twice. Nothing else carries it: there is no per-shard pc chaining, and a shard's time window ties to no row (proof.md §8).

Per address, the order is timestamp order, and it is program order: the pc query's gap puts consecutive pc writes at least 4 apart (§2.4), so each cycle's four timestamps precede the next cycle's whatever value cycle takes, and a row never reads an address before its predecessor's write there. An invocation rides its requesting row's cycle (delegation.md §5) and is ordered with it.

Auditors/Proof system

Lookups

Normative specificationdocs/spec/lookup.mdView as Markdown

How a circuit's lookup obligations are proved: per shard, by one LogUp channel per table, each summed by a fraction tree inside the circuit's own GKR pass and checked at its root.

1 What a channel claims#

A lookup is LookupExpr { name, channel, selector, tuple }: a channel of constants::lookup_channel, a committed M, W or S column as selector, and a tuple of Linear expressions with literal coefficients over committed columns and the circuit's virtual tables. It holds on a row where the selector is 0, or

  • on a range channel, where its one expression's canonical integer is below 2^BITS[channel];
  • on a table channel, where its tuple is a row of the channel's one table: 1 to MAX_TUPLE = 7 expressions, the same number for every lookup of the channel.

memory.md §7 is the convention range obligations follow. A channel discharges all of a shard's lookups on it as one identity over the shard's rows y:

Σ_y Σ_l 1/(E_l(y) + g)  −  Σ_y mult(y)/(T(y) + g)  =  0

E_l(y) is lookup l's gated tuple (§4) and T(y) the table's row y, both compressed by β (§5); mult is the channel's multiplicity column (§7). A range table is the one column [0, 2^BITS).

2 The challenges#

slot challenge_slot value
6 LOOKUP_G g, drawn
7 LOOKUP_BETA β, drawn
8–12 LOOKUP_BETA_2 … LOOKUP_BETA_6 β^2 … β^6, derived
13 LOOKUP_DECODER_NEUTRAL g − Σ_{j<W} β^j, derived; W the decoder tuple's width

g and β are shard-local: the shard's transcript draws them, in that order under the tag LOOKUP_CHALLENGE (33), right after absorbing its witness commitments, multiplicities included (proof.md §4). M columns are committed in the global transcript the shard is seeded from and S columns are bound by identity or the SRS digest, so every column a channel reads is fixed before either challenge exists.

β^0 is the literal 1, so a one-column tuple names no slot. A gate coefficient is one literal or one slot (gkr.md §3), so each higher power is a slot of its own, computed by the verifier and never read from a proof (gkr_verify::insert_lookup_challenges, which reads W off the artifact's decoder lookup).

Selectors are boolean: CircuitArtifact::validate refuses a lookup whose selector no enforcing gate of gate list 0 holds to s − s·s = 0. The selector multiplies the tuple inside the denominator (§5), so a channel proves the gated tuple s·(e + o) + n is a table row, which is the obligation only at s ∈ {0, 1}. At any other s a scaled tuple is looked up instead: on a range channel, s = t·e⁻¹ lands any nonzero e on any table value t.

3 Tables#

channel id kind table, at row y width table_vars
TIMESTAMP 0 range V[range19]: y mod 2^19 1 19
RANGE16 1 range V[range16]: y mod 2^16 1 16
GENERIC 2 table, committed the packed table (§9) 3 0
DECODER 3 table, committed the family's decoded table (§10) 7 or 6 0
XOR8 4 table, virtual V[xor8_a], V[xor8_b], V[xor8_out]: y's low two bytes and their XOR 3 16

A virtual table is a closed form of the row index, never committed: the verifier evaluates its multilinear extension where the GKR pass ends (gkr_verify::virtual_at_point, gkr.md §2). Each is a weighted sum of the row's bits but V[xor8_out], Σ_{j<8} 2^j·(y_j + y_{j+8} − 2·y_j·y_{j+8}), which is exact because y ^ z = y + z − 2yz is multilinear. So XOR8 costs no commitment and nothing in the SRS digest.

constraints::lookup::table_vars is the fewest variables at which a table is complete: BITS for a range channel, 16 for XOR8, 0 for a committed table, a setup column at the circuit's own height. Below it a virtual table holds only part of its range, which costs completeness, not soundness. family_circuit returns None below the largest table_vars of a family's channels, so a key naming such a height fails to load (proof.md §7). On the height menu (program.md §7) a family carrying TIMESTAMP is at 2^20 or more, and one carrying RANGE16 or XOR8 at 2^16 or more. The packed table needs 2^18 rows (§9), and every family that reads it carries TIMESTAMP. Above table_vars a table repeats, which §7 makes harmless.

XOR8's tuple is three wide so that membership bounds each entry to [0, 256) on its own; a packed key x + 256·y would bound neither, (x, y) and (x + 256, y − 1) compressing alike. Every other bitwise operation on bytes is a linear form over its results (delegation-circuits.md §1).

4 Gated keys#

A lookup expression is evaluated on every row, so the selector sends a row whose key means nothing to a neutral tuple, which is a real table row:

gating channels gated position j neutral tuple
NoOffset TIMESTAMP, RANGE16, XOR8 s·e_j all zero
ZeroEntry GENERIC s·(e_0 + 1), then s·e_j the all-zero ZeroEntry row
MinusOne DECODER s·(e_j + 1) − 1 −1 in every column, a padding row

The + 1 keeps every real key of the packed table at 1 or above, so no real entry is the all-zero tuple a switched-off row looks up. A range table needs no offset, 0 being in range, and one would push 2^BITS − 1 out of it; XOR8's (0, 0, 0) is a true entry. A decoded table has no all-zero row, pc 0 being a valid pc, and its MINUS_ONE padding rows (program.md §5) are the neutral entry.

Each key a table channel looks up is bounded by the family that looks it up, because a channel proves membership and nothing more. A selected row whose key expression is −1 gates to the ZeroEntry, and an unbounded key reaches any sub-table of the packed table: an AND key a + AND_BASE with a unbounded lands on a U16GetSign row and proves a false AND. Families bound their keys with RANGE16 obligations or build them from bounded columns (shift-bitwise.md §3 and the other family pages); the decoder's key is §10's.

5 The denominator#

With s the selector, e_j = Σ_i c_{j,i}·x_{j,i} + k_j, and §4's offset o_j (1 or 0) and neutral value n_j (−1 or 0):

E + g  =  Σ_j β^j·(s·(e_j + o_j) + n_j)  +  g
       =  Σ_j β^j·s·e_j  +  Σ_j β^j·o_j·s  +  (g + Σ_j β^j·n_j)
T + g  =  Σ_j β^j·t_j  +  g

E + g is one Quadratic (constraints::lookup::row_denominator): each term of e_j the product (β^j·c)·s·x, each offset the linear term β^j·o_j·s, and the bracket the slot LOOKUP_G or, for the decoder, LOOKUP_DECODER_NEUTRAL. β^j·c is one coefficient only where β^0 = 1 makes it a literal or c = 1 makes it the slot, so position 0 takes any literal coefficients and constant and every later position weights its columns by 1 with no constant. T + g is one Linear over the table's columns (table_denominator).

6 The fraction tree#

A channel's leaf level is (num, den) pairs of gate-list-0 columns, P = (L + 1).next_power_of_two() of them for L lookups:

leaf num den
the table, first −mult T + g
each lookup, in artifact order 1 E_l + g
padding, up to P 0 1

Row-wise gate lists add sibling pairs, (n_a·d_b + n_b·d_a, d_a·d_b), until each row holds one pair; a tree shallower than the circuit's deepest copies itself up. Then trace_vars halving lists add the rows' pairs, TreeCross writing the numerator and TreeProduct the denominator (gkr.md §3). The circuit's outputs are the memory argument's read and write roots, then each channel's (num, den) in the order of its channel specs (crates/constraints/src/build.rs).

A channel costs 4P − 2 inner columns to reduce a row, 2 more per copy-up layer and 2 per halving list, and one committed column. P doubles each time L reaches a power of two.

The padding clause (gkr.md §4) asks a padding row to feed 1 into every product tree. A fraction tree is exempt: its identity is (0, 1), and a padding row is not idle in a channel but looks up the neutral tuple, which the multiplicity counts. checker::check_padding_identity exempts every column a TreeCross reads.

7 Multiplicities#

Each channel has one multiplicity column, a committed W column; a circuit's are its last W columns, in channel order. Row t counts the (row, lookup) pairs of the shard whose gated tuple is table row t's, switched-off rows included. A tuple at several table rows is credited to the lowest; every other copy holds 0 and contributes 0/(T + g). The count is over raw gated tuples, the column being committed before g and β exist (trace::build_multiplicities, which refuses a tuple no table row holds: the honest prover cannot balance it).

No gate or range check constrains the column, and soundness needs none. If a gated tuple v is in no table row, the left side of §1's identity, as a rational function of g, has a pole at −v whose residue is the number of lookups producing v: a positive integer below p, whatever the column holds.

8 The root check#

accept  iff  num = 0  and  den ≠ 0

on each channel's root pair, at step 9 of proof.md §6 (gkr_verify::channel_holds); a failure is VerifyError::Lookup { channel }. The GKR pass absorbs the pair before its first challenge and proves it (gkr.md §5). den is the product of every leaf denominator, and num = 0 means the sum vanishes only where den ≠ 0: one leaf (0, 0) — a table row whose T + g vanishes, counted 0 — makes the root (0, 0) whatever the other leaves hold. With g drawn after the columns that has probability at most fractions/|Fr|, and den ≠ 0 makes it a refusal.

9 The generic table#

One committed table of constants::generic_table::WIDTH = 3 columns, a key and two values, packing three sub-tables under disjoint key ranges (program::lookup_tables::generic_table):

row 0                      ZeroEntry    (0, 0, 0)
rows 1 ..= 2^16            AND          (AND_BASE + a + 1,    b,        a & b)        a, b < 2^8
rows 2^16+1 ..= 2^17       U16GetSign   (SIGN_BASE + h + 1,   h >> 15,  0)            h < 2^16
rows 2^17+1 ..= 2^17+32    ShiftPowers  (SHIFT_BASE + s + 1,  2^s,      2^(31 − s))   s < 32
rows above                 zero

AND_BASE = 0, SIGN_BASE = 256 and SHIFT_BASE = 65,792 put the keys at 1..=256, 257..=65,792 and 65,793..=65,824; a lookup's key expression is x + BASE, and the gating adds the 1. U16GetSign serves every sign an execution family computes, AND the bitwise operations of SHIFT_BITWISE and ATOMICS, ShiftPowers the shifts. The copower 2^(32 − s) is stored halved (SHIFT_COPOWER_BITS = 31), 2^32 not fitting a u32 column, and the two gates that read it carry the factor 2 (shift-bitwise.md §4). 131,105 rows in all (GENERIC_ROWS).

Its commitments are a constant of the ceremony. A Mercury commitment reads the evaluation table as coefficients (mercury.md §2) and the table is zero past its entries, so over 2^n rows it commits to the same three points for every n ≥ 18; generic_commitments(srs) computes them at 2^18 (GENERIC_LOG_HEIGHT). Every verifying key carries them once, as VerifyingKey::generic_table, whether or not a family reads the channel, and its SRS digest covers them (proof.md §3); program identity does not. A circuit that reads GENERIC names the table as its three setup columns after identity's (FamilyCircuit::reads_generic_table), and a shard's opening checks them against the key's points (proof.md §5).

10 The decoder channel#

The DECODER table is the family's decoded table, program::lookup_tuple(family)'s columns, as its first setup columns at its height, row i holding pc 2i (program.md §5); program identity commits them (program.md §8). Each execution family makes one lookup on it, imm absent for MUL_DIV and ATOMICS:

decode_row    selector m_pc    tuple (pc read value, next_pc, rs1, rs2, rd, [imm], extra_mask)

The key is the frame's own pc read (memory.md §2), so the cycle itself is bound to the program; the rest are the row's decoded columns, which the family's other gates read. The selector is the row's liveness, so a padding row looks up the MINUS_ONE tuple, which every decoded table holds, being taller than its last instruction.

The family's decoded_mask_bits gate ties the packed mask to boolean kind bits. That the bits are one-hot, and that a live row is an instruction at all, is the table's domain: its live rows hold one-hot masks and its padding rows −1, which no sum of kind bits reaches. Boolean columns looked up one by one would lose this: booleanity admits any subset of bits, the empty one included, and an all-zero mask makes every gate a kind selects vacuous.

11 Construction rules#

CircuitArtifact::validate enforces §1's form and widths, §2's selector rule and §5's coefficients wherever an artifact is built or loaded (gkr.md §4). When a circuit is assembled, constraints::lookup asserts that a channel has a lookup, that its multiplicity is a W column, that every lookup has its table's width, and that a range channel's table is the one its bound names (range_table) with BITS ≤ trace_vars; constraints::memory::frame_with_channels_artifact refuses an empty channel list, which would leave a frame's gap obligations discharged by nothing.

The discharge rule, constraints::lookup::check_discharge, at assembly and at every key load (VerifyingKey::check): every lookup is the denominator of exactly one gate-list-0 column, its numerator 1 directly before it; no column is two lookups'; each channel's (−mult, T + g) appears once. It matches by normalized expansion inside the cone below the channel's own root pair, so an obligation or table fraction in another channel's tree is refused, and the two range channels, which gate alike, are not confused. Which output pair is whose root, which columns are a table and which counts it is not in the artifact but in its ChannelSpecs, which a key carries in FamilyCircuit::channels and must hold as the registry's (circuits.md §1). checker enforces this rule and the lookup rules a second time, with code of its own (circuits.md §3).

The copower rule, constraints::lookup::check_copowers, run by every constructor that bounds a column through a copower. A bound x < p written as x·p′ < 2^32, p·p′ = 2^32, bounds nothing alone: p′ is a unit of Fr, so x = s·p′⁻¹ ranges over a coset of 2^32 values. Each such x therefore also carries a direct RANGE16 bound, as a halfword or as a high chunk and a remainder, under the same selector.

12 What it rests on#

  • Every gated tuple is a table row: §1's identity over challenges drawn after every column it reads, boolean selectors (§2), both root conditions (§8) and the GKR pass. The error is at most fractions/|Fr| for g, plus looked-up tuples × table rows × (width − 1)/|Fr| for a β collision: below 2^−190 at every menu height.
  • A lookup answers from its own sub-table: one width per channel (§11), disjoint key ranges and the + 1 (§9), and its family's bound on the key (§4).
  • A switched-off row costs nothing: its neutral tuple is a table row the multiplicity counts (§4).
  • The table is the intended one: the verifier's own closed form (§3), or a table bound by identity or by the SRS digest, as trustworthy as the channel the verifier took that from (program.md §8, srs.md §3).
  • Every declared obligation is discharged: the discharge rule over the registry's specs (§11).

The channel does not check the multiplicity column (§7), a key's bound (§4), or that a committed table holds its neutral row, a property of its values that no artifact states: a table without one stops the honest prover at trace::build_multiplicities.

Auditors/Proof system

The proof

Normative specificationdocs/spec/proof.mdView as Markdown

What a verifier checks and the formats it reads, from the statement to the bytes. The memory argument, LogUp, GKR and Mercury are their own pages; this one is how they compose. crates/verifier-core (#![no_std]) implements everything here but step 12, the opening, which crates/verifier runs.

1 The statement#

A statement is a PublicInputs under a verifying key (§7): one execution of the key's program. It is proved by one ShardProof per statement shard, a (family, index) with index below the family's shard count, each verified against the same PublicInputs. A shard's proof establishes its own circuit and opening, and the reconciliation it joins reads the roots the statement claims for every other shard, which only their own proofs establish: a statement is verified when its proofs are exactly its shards and all pass, never by a subset.

Every entry point is (&VerifyingKey, proof, &PublicInputs): verifier::verify_shard, verifier::verify_block, verifier_core::reduce_shard. A verifier holds two values from a channel the prover does not control, the program identity and the SRS digest (§3), and compares them with the key's; the key itself may come from anyone (§7). The verifier CLI compares identity only (tools.md §6); host::verify(vk, block), which is verify_block(vk, block, block.statement()), compares neither and leaves its caller to check the statement's input, output and exit status too.

1.1 PublicInputs#

field
input: Vec<u8> the public input window's payload, at most PUBLIC_PAYLOAD_BYTES = 16,380 (public-values.md §3)
output: Vec<u8> the journal, the public output window's payload, as long
exit_status: u32 x10's final value
shard_counts: Vec<u32> one per family of the VmConfig, in its order, possibly 0
windows: Vec<u32> ZERO_WINDOWS' window ids, one per shard (memory.md §3.5)
boundary: BoundaryFinals the 64 register and pc boundary scalars (memory.md §4.1)
memory_commitments: Vec<Vec<[u8; 64]>> per statement shard, its M columns' commitments in layout order
memory_roots: Vec<[Fr; 2]> per statement shard, [read_root, write_root]

The first three are the claim; the rest is the execution's record, which the prover chooses. All of it but the roots and the exit status is absorbed before any challenge (§2).

1.2 Statement order#

verifier_core::statement_shards(config, counts):

(INIT_TEARDOWN, 0)
(ZERO_WINDOWS, 0) … (ZERO_WINDOWS, k − 1)
every other family of the VmConfig, ascending by id, shards 0 … count − 1 each

It orders memory_commitments, memory_roots, G8's groups and a block's proofs. The two leading families are ids 7 and 8, so the order is not ascending by id. A family with count 0 has no entry.

1.3 The block#

verifier_core::BlockProof { config, statement, shards } is one execution closed: the static VmConfig, the statement and one proof per statement shard, in statement order; it adds no evidence to the proofs'. The config and counts are public data of the proof, absorbed at G3 and G4, so the block carries both and check B1 holds them to the key's and the verifier's.

Shard-set exactness, BlockProof::shape, at decode and again in verify_block: one count per config family; the counts' total, summed in u64 before any list is built from them, equal to the numbers of proofs, commitment lists and root pairs; the proofs naming statement_shards in order. No (family, index) is missing, repeated or extra.

BlockProof::reconciliation is the cross-shard record set, a BlockReconciliation of one ShardRecord { family, shard_index, ts_window, memory_commitments, roots } per statement shard, assembled from the shard's proof (the window) and the statement (the rest).

2 The global transcript#

verifier_core::global_commit(vk, statement), run by the verifier in derive_global_phase and by the prover once every shard's M columns are committed (streaming.md §2): a fresh transcript, tag values in transcript.md §5.

# op tag message
G1 absorb PROTOCOL_SUITE [PROTOCOL_VERSION], 0
G2 absorb SRS_DIGEST [vk.srs_digest] (§3)
G3 absorb VM_CONFIG the config (program.md §7)
G4 absorb SHARD_COUNTS shard_counts
G5 absorb MEMORY_WINDOWS windows
G6 absorb PROGRAM_IDENTITY [vk.identity]
G7 absorb PUBLIC_INPUTS, bytes the 32 bytes of io_digest(input, output) (public-values.md §5)
G8 per family MEMORY_GROUP, COMMITMENT below
G9 absorb MEMORY_BOUNDARY the 64 boundary scalars
G10 squeeze ×4 MEMORY_CHALLENGE γ_M, α_addr, α_ts, α_val, challenge slots 1–4
G11 squeeze GLOBAL_STATE_DIGEST the global state digest

G3–G5 are verifier_core::absorb_statement_descriptor. G8 is one group per family of the config, in statement order, a family with count 0 included:

MEMORY_GROUP   [family, shard count]
COMMITMENT     per shard, ascending: its memory_commitments, one message of 4k limbs

A statement's or proof's points are absorbed as limbs (transcript.md §4) and decoded only at step 12.

Everything a memory tuple or the reconciliation reads precedes G10 (memory.md §6.1 says why for each). Two fields are not absorbed: memory_roots, which depend on the challenges and are bound by each shard's own GKR proof (step 10a), and exit_status, which step 10b holds to v_10, absorbed at G9. The digest seeds every shard (§4); a proof carries the digest it was seeded with (ShardProof::global_digest) and step 5 compares it with the replay, so a shard proof is for one statement under one key.

3 The SRS digest#

t ← Transcript::new()
t.append_bytes(SRS_VERIFIER, srs_verifier)           320 bytes, srs.md §5
append_g1_points(t, GENERIC_TABLE, generic_table)    the table's 3 points, one 12-limb message
srs_digest ← t.sample()                              one raw squeeze

verifier_core::srs_digest; GENERIC_TABLE is absorbed in this sponge and nowhere else, the points key column first. G2 absorbs the digest, so a proof is bound to the three points its pairings read and the table its GENERIC lookups read. Both are constants of the ceremony (lookup.md §9), so one digest serves every key. It does not cover the powers, which only a prover reads: an opening is checked against g2_tau whatever powers made the commitment.

A key's load recomputes the digest from the key's own points (§7), which shows they agree, not that they are the ceremony's, and identity binds neither (program.md §8). So the verifier compares vk.srs_digest with the ceremony's (srs.md §3). Without that comparison, whoever built the key chose τ, so can open anything, and chose the table every GENERIC lookup is held to.

4 The shard transcript#

Shard (family, index) runs a fresh sponge (verifier_core::shard_transcript), not a restored global one:

# op tag message
S1 absorb SHARD_SEED [global state digest, family, index]
S2 absorb SHARD_TS_WINDOW [ts_start, ts_end] (§8)
S3 absorb COMMITMENT the shard's W commitments, multiplicities included, one message
S4 squeeze ×2 LOOKUP_CHALLENGE g, then β (lookup.md §2)
S5 the GKR backward pass (gkr.md §5.2)
S6 the batch opening (§5): B1–B3 of mercury.md §5, then the sixteen steps of its §3

S4 is drawn for every shard, whether or not its circuit has a channel. Every challenge follows every commitment the circuit reads: M at G8, S through identity at G6 or the SRS digest at G2, W at S3, as GKR requires of its caller (gkr.md §5.1).

The circuit's external challenges (verifier_core::shard_challenges) are slots 1–4 from G10; for a window family, slot 5 at the window verifier_core::shard_window gives the shard (memory.md §3.3); then the lookup slots from g and β. Its outputs, the top layer, are the two memory roots and then each channel's (num, den) (lookup.md §6), 2 + 2c of them for c channels.

In the recursion format a shard commits M and W as stacks of 2^σ columns, and S6 opens with σ STACK_CHALLENGE squeezes extending the opening point (recursion.md §1.3). At σ = 0, the base format, there are none.

5 The opening#

After S5 every committed column has one claim, layer 0's, all at one point u (gkr.md §5.2). So there is nothing for a claim-merging sumcheck to merge, and S6 opens every column as one batch (mercury.md §5), one 704-byte Mercury proof a shard:

columns      the circuit's committed layout: M[0..], W[0..], S[0..]
commitments  M  PublicInputs.memory_commitments[the shard's position]
             W  ShardProof.witness_commitments
             S  VerifyingKey.setup_commitments[family], then VerifyingKey.generic_table
                when the circuit reads GENERIC (FamilyCircuit::reads_generic_table)
point        u, variable j at index j
values       layer 0's claims, ShardProof.gkr.layers[0].final_evals

Column i carries ρ^i, so this order is part of what is proved. Virtual columns are neither claimed nor opened: the verifier evaluates their closed forms. Taking S from the key is what makes the opening bind the columns identity commits, the decoded tables and the image column (memory.md §6.2), and the generic table the SRS digest covers.

reduce_shard ends at an OpeningClaim: these commitments, the point, the values and the live shard transcript. verify_shard spends it with pcs::batch_verify (step 12); a recursion node defers it (recursion.md §8.3).

6 Verification#

verifier::verify_shard(vk, proof, public) returns the first failure, in this order, as a VerifyError:

step class check
1 Statement one shard count per config family; the key's circuits are its config's families, in order, with one setup list each
2 Statement check_memory_windows (memory.md §3.5); input and output each at most PUBLIC_PAYLOAD_BYTES
3 Statement one root pair and one commitment list per statement shard, each list its family's M width; the total summed in u64 first
G1–G11 (§2)
4 Statement ts_start ≤ ts_end ≤ 2^38
5 Statement the replayed global state digest is proof.global_digest
6 Malformed (family, index) is a statement shard; the witness commitments and outputs have the circuit's counts
7 Constraint { layer } gkr_verify::verify over the shard transcript: LayerInconsistency { layer }; its ProofShape, OutputShape and MissingChallenge are Malformed
8 Constraint { layer: 0 } every base claim at one point
9 Lookup { channel } gkr_verify::channel_holds on each channel's root pair, in channel order (lookup.md §8)
10a MemoryArgument the proof's two roots are the statement's for its position
10c MemoryArgument a PUBLIC_INPUT or PUBLIC_OUTPUT shard's value column is the statement's string (public-values.md §5)
10b MemoryArgument every boundary timestamp below 2^38; v_10 = exit_status; gkr_verify::reconciles over every shard's roots and boundary_factors(challenges, vk.entry_pc, boundary) (memory.md §4.2)
11 — the opening claim (§5)
12 Opening the SrsVerifier, every commitment and the Mercury proof through their validating decoders, then pcs::batch_verify; any failure

Step 8 cannot fail on verify's output, whose base claims share layer 0's point; it states what step 11 relies on. Steps 1–3 hold the statement to the key before the replay indexes by it, so nothing a proof or statement carries makes the core panic, for a loaded key.

The split, by what each part reads (verifier_core):

function reads steps runs
derive_global_phase(vk, public) → GlobalChallenges key, statement 1–3, G1–G11 once a statement
verify_shard_local(vk, global, proof, public) → OpeningClaim and one ShardProof 4–10a, 10c, 11 once a shard
verify_global_memory(vk, global, public) key, statement, challenges 10b once a statement

GlobalChallenges is the four memory challenges and the digest. reduce_shard is the three in that order, verify_shard that and step 12; step 11 cannot fail, so 10b after it is 10b in place. Step 10b reads only vk.entry_pc, the boundary, the roots and the challenges, so a block runs it once; step 10a puts each shard into the product by holding the roots its GKR proof outputs to the statement's entry, and shard-set exactness makes every root there a verified shard's. verify_shard_local alone verifies no memory argument: without verify_global_memory it accepts shards, each valid, whose multiset does not close.

verifier::verify_block(vk, block, public):

class check
B1 Statement block.config is vk.config, and block.statement is public
B2 Statement derive_global_phase, once
B3 Statement BlockProof::shape (§1.3)
B4 Statement check_ts_windows over the records (§8)
B5 MemoryArgument verify_global_memory, once
B6 as verify_shard per shard, in statement order: verify_shard_local, then step 12

B1–B5 read no GKR proof or opening, so a statement that cannot reconcile is refused before any circuit runs, and the class can differ from verify_shard's: a change to anything G1–G9 absorb that B1–B4 admit moves the challenges, so the honest roots stop reconciling and verify_block answers MemoryArgument where verify_shard names the seed at step 5; a forgery that unbalances the multiset is MemoryArgument even where it also breaks a gate. A dropped shard fails B5: the truncated statement, re-proved honestly with its counts, lists and roots adjusted, passes B1–B4 and misses that shard's memory events on one side of the product.

In verify_shard's order the class names the fault: a tampered witness proved honestly, its multiplicities recounted, columns recommitted and statement rebuilt, fails at the gate (Constraint), table membership (Lookup) or multiset (MemoryArgument) it broke, which checker::TamperHarness asserts (circuits.md §3).

7 The verifying key#

7.1 Fields#

field
code_version: u32 constants::family::CODE_VERSION, 0
config: VmConfig the static shape (program.md §7)
entry_pc: u32 the image's entry pc
identity: ProgramIdentity program.md §8
setup_commitments: Vec<Vec<[u8; 64]>> identity's commitment lists, one per config family, in its order
srs_verifier: [u8; 320] the SrsVerifier (srs.md §5)
generic_table: [[u8; 64]; 3] the generic table's commitments, key column first, in every key (lookup.md §9)
srs_digest: Fr §3
circuits: Vec<FamilyCircuit> one per config family, in its order: the family, its CircuitArtifact and its ChannelSpecs (lookup.md §11)

A key carries every family's artifact, so its size is mostly its delegation families' (circuits.md §1).

7.2 Loading#

VerifyingKey::from_bytes decodes (§9), refuses bytes that are not the key's canonical encoding, and runs VerifyingKey::check, which refuses, in order:

  1. a VmConfig no derivation produces (VmConfig::from_bytes of its own bytes), or a code_version other than CODE_VERSION;
  2. a setup list count other than the config's family count;
  3. an identity that identity_digest(code_version, config, entry_pc, setup_commitments) does not reproduce;
  4. an srs_digest that srs_digest(srs_verifier, generic_table) does not reproduce;
  5. a circuit count other than the family count; then, family by family: a circuit for another family; a height the registry has no circuit for; a circuit, artifact or channel specs, other than config.circuit(family, trace_vars), the registry of the config's format (circuits.md §1); an artifact failing CircuitArtifact::validate, constraints::memory::check_memory or constraints::lookup::check_discharge; a setup list whose length, plus 3 if the circuit reads GENERIC, is not the artifact's S count; and GENERIC specs naming anything but the 3 setup columns after identity's, §5's order, which no registry circuit fails.

verifier::load_verifying_key then decodes every curve point: the SrsVerifier's three, each setup commitment and each generic-table commitment, through the validating readers. The circuits are held to the registry because nothing else binds them: identity binds the program, not the circuit that proves it.

A key from an untrusted source. Every field is recomputed from or compared with one of the verifier's two trusted values (§1), or fixed by the code: config, entry_pc and the setup lists through identity; srs_verifier and generic_table through the SRS digest; code_version and the circuits by the verifier's own registry. So a key may come from the prover, provided both comparisons are made.

Validation runs once, at load; verify_shard and verify_block assume a loaded key. On one edited in memory a changed config or circuit list is still refused as Statement, but an edit inside a circuit may go unnoticed.

prover::ProverSetup::new(program, srs) builds the key: each family's circuit from VmConfig::circuit and fill from prover::family_fill (ProverError::Unregistered if either is missing), identity's and the generic table's commitments over srs, then check (ProverError::Key). srs needs as many powers as the tallest family has rows, and 2^18 for the generic table (program::lookup_tables::GENERIC_LOG_HEIGHT); fewer panics.

8 Time windows#

Each shard claims [ts_start, ts_end) (ShardProof::ts_window): the slice of the clock (execution-trace.md §1) its rows write in, their reads reaching back before it. S2 absorbs it before the witness commitments, so a proof made under one window fails under another; step 4 holds it to ts_start ≤ ts_end ≤ 2^38 and nothing more.

verifier_core::check_ts_windows (B4), over the records in statement order: within each cycle-owning family (constants::family::CYCLE_OWNING, the execution families 0–6), every window is non-empty and no shard's ts_end exceeds the next shard's ts_start; a family's records are consecutive and ascending, so neighbours suffice. It is per family because families interleave — ADD_SUB_LUI_AUIPC may own cycles 1 and 3 and JUMP_BRANCH_SLT cycle 2 — and every other family is exempt: a window family's rows are words, a delegation family's invocations at their requesting cycles (delegation.md §8).

The prover reads a window off the shard's committed M[0] cycle column (ts_window, crates/prover/src/lib.rs): [4·c_0, 4·c_max + 4), c_0 row 0's cycle and c_max the largest, padding rows carrying 0, for cycle-owning and delegation families alike. Window families claim verifier_core::TRIVIAL_TS_WINDOW = [0, 2^38).

A window binds nothing. No gate ties it to the rows committed under it, so a prover may claim any windows the rule admits; cross-shard order, cycle uniqueness and pc continuity are the memory multiset's alone (memory.md §9). B4 checks the shape of the shard plan and adds nothing to soundness.

9 Wire forms and the proof archive#

verifier_core::wire: integers little-endian; an Fr its 32 canonical bytes (primitives.md §1), refused at or above p; a G1 its 64 bytes (primitives.md §3), opaque to the core; bytes a u32 length then the bytes; list<T> a u32 count then the items; T[k] exactly k items, no count. Every decoder is total: it refuses a count the remaining bytes cannot hold, so it reserves nothing an untrusted length asks for, and refuses trailing bytes.

PublicInputs   input bytes, output bytes, exit_status u32,
               shard_counts list<u32>, windows list<u32>,
               boundary Fr[64]                     memory.md §4.1's order and ranges
               memory_commitments list<list<G1>>, memory_roots list<Fr[2]>

ShardProof     family u32, shard_index u32, ts_start u64, ts_end u64, global_digest Fr,
               witness_commitments list<G1>, outputs list<Fr>,
               gkr list<(rounds list<Fr[4]>, final_evals list<Fr>)>       transition 0 first
               opening u8[704]                     pcs::MercuryProof, mercury.md §4

BlockProof     config bytes                        VmConfig, program.md §7
               statement bytes                     PublicInputs
               shards list<bytes>                  each a ShardProof; then BlockProof::shape

VerifyingKey   code_version u32, config bytes, entry_pc u32, identity Fr,
               setup_commitments list<list<G1>>, srs_verifier u8[320], generic_table G1[3],
               srs_digest Fr,
               circuits list<(family u32, artifact bytes,          CircuitArtifact, gkr.md §4.1
                              channels list<(channel u32, table list<Address>,
                                             multiplicity Address)>)>
Address        tag u8 (0 M, 1 W, 2 S, 3 V), index u32; a V's index is its gkr.md §2.1 kind tag

BlockReconciliation   list<(family u32, shard_index u32, ts_start u64, ts_end u64,
                            memory_commitments list<G1>, read_root Fr, write_root Fr)>

A ShardProof's lengths are fixed by its key and family, and steps 6–7 hold them: transition k carries n_{k+1} rounds and w_k claims, twice that if halving (gkr.md §5.5). For a base-format circuit at 2^n with W witness, C committed and I inner columns, O outputs, and R row-wise lists before its n halving ones, that is

772 + 64·W + 32·O + 8·(R + n) + 128·(R·n + n(n − 1)/2) + 32·(C + I + O·(n − 1))   bytes

which circuits.md §1 tabulates per family.

The proof archive. verifier::proof_archive::write_proof(dir, stem, vk, block), re-exported as host::proof_archive, writes four files, each a bare to_bytes with no header of its own:

<stem>.vk         VerifyingKey
<stem>.identity   the key's identity: its 32 bytes in order, 64 lowercase hex digits, a newline
<stem>.public     PublicInputs: the block's own statement
<stem>.block      BlockProof

read_proof(dir, stem) is the inverse, each file through its type's decoder and the key through load_verifying_key. .identity records what the run claimed, and read_proof returns it unchecked: a verifier's identity comes from its own channel (§1). .public repeats the statement .block carries, for the CLI, which takes it as a file (tools.md §6).

Auditors/Instruction families

The ADD_SUB_LUI_AUIPC family

Normative specificationdocs/spec/add-sub.mdView as Markdown

add, sub, addi, lui, auipc, ecall, ebreak and fence, compressed forms included, are family 0, one executed instruction a row; constraints::add_sub::artifact is its circuit. Every ecall is a row of it, so the circuit also proves the exit and the request side of every delegation call. This page specifies what it adds beside the memory frame (memory.md §2).

1 Columns#

The decoded tuple is pc next_pc rs1 rs2 rd imm extra_mask (program.md §5), the mask one-hot over the kinds system addi auipc add sub lui in bit order (constants::extra_mask::add_sub_lui_auipc). ecall, ebreak and fence share the system kind, with imm 0, 1 and 2 (constants::extra_mask::system_code); elsewhere imm is what the instruction adds — addi's sign-extended immediate, lui's and auipc's shifted left by 12, 0 for add and sub — and a register field the instruction lacks is 0.

M[0..26] and W[0..8] are the frame of the five queries pc rs1 rs2 rd deleg, and M[26], deleg_space, is the requested delegation type's anchor address space, 0 on a row requesting none (memory.md §2). The family adds these columns, and reads V[range19] and V[range16]:

column name
W[8..14] decoded_next_pc, decoded_rs1, decoded_rs2, decoded_rd, decoded_imm, decoded_mask the claimed decoded row
W[14..20] kind_system … kind_lui b_k, the mask's bits
W[20], W[21] is_ecall, is_fence the system kind, split by its code
W[22..28] is_deleg_<f>, f = 9, 10, 11, 15, 16, 17 d_t: a request of delegation type t, family f
W[28], W[29] wrap, rd_hi the sum's carry or the difference's borrow; sel's high halfword
W[30], W[31] pc_wrap, next_pc_hi next_pc's wrap and high halfword
W[32..35] mult_timestamp, mult_range16, mult_decoder one multiplicity a channel
S[0..7] table_pc … table_extra_mask the decoded table

Below, m_q, a_q, ts_q and v_q are query q's mask, address, read timestamp and read value; pc and next_pc the pc query's read and write values; sel is rd_selected (W[7]), the value the frame writes to a nonzero rd. N_t and tag_t are type t's ecall number and anchor space, the first six rows of constants::delegation::TYPES in order (delegation.md §3); is_exit = is_ecall − Σ_t d_t; 93 is constants::ecall::EXIT and HALT_PC is 1 (memory.md §5).

The family's fill (prover::family_fill, crates/prover/src/fill.rs) writes sel as the computed value even where rd = x0.

2 Gates#

63 enforcing gates, all in gate list 0, each of degree at most 2 and 0 on the all-zero row: the frame's eleven (memory.md §2) and these 52, in artifact order, each held to 0:

gate expression
kind_<k>_boolean, six b_k − b_k²
decoded_mask_bits Σ_k 2^k·b_k − decoded_mask
is_ecall_boolean, is_fence_boolean y − y²
system_split is_ecall + is_fence − b_system
ecall_code is_ecall·decoded_imm
fence_code is_fence·(decoded_imm − 2)
per type: is_deleg_<f>_boolean, deleg_<f>_is_an_ecall, deleg_<f>_number d_t − d_t²; d_t·(1 − is_ecall); d_t·(v_rs1 − N_t)
ecall_is_exit is_exit·(v_rs1 − 93)
rs1_mask_rule m_rs1 − m_pc·(b_add + b_sub + b_addi + is_ecall)
rs2_mask_rule m_rs2 − m_pc·(b_add + b_sub + is_ecall)
rd_mask_rule m_rd − m_pc·(b_add + b_sub + b_addi + b_auipc + b_lui + is_ecall)
deleg_mask_rule m_deleg − m_pc·Σ_t d_t
rs1_addr_rule m_rs1·(a_rs1 − decoded_rs1 − 17·is_ecall)
rs2_addr_rule, rd_addr_rule m_q·(a_q − decoded_q − 10·is_ecall)
rs1_value_masked, rs2_value_masked v_q − m_q·v_q
add_addi_auipc (b_add + b_addi + b_auipc)·(v_rs1 + v_rs2 + decoded_imm − sel − 2^32·wrap) + b_auipc·pc
sub b_sub·(v_rs1 − v_rs2 − sel + 2^32·wrap)
lui b_lui·(decoded_imm − sel)
exit_status is_exit·(v_rd − sel)
deleg_writes_no_register m_deleg·sel
deleg_read_ts_zero, deleg_read_value_zero m_deleg·ts_deleg; m_deleg·v_deleg
deleg_addr_rule m_deleg·(a_deleg − v_rs2)
deleg_space_rule deleg_space − Σ_t tag_t·d_t
wrap_boolean, pc_wrap_boolean y − y²
next_pc_rule next_pc + 2^32·pc_wrap − (1 − is_exit)·decoded_next_pc − is_exit·HALT_PC

N_t and tag_t are literals read from constants::delegation::TYPES, so the base circuit depends on the registry's first BASE_TYPES = 6 rows and on no row appended after them. The recursion format's circuit, add_sub::recursion_artifact, carries a selector and its three gates for each of the ten types, the columns after them shifted by four, and in place of deleg_writes_no_register deleg_a0_rule, Σ_{t<6} d_t·sel + Σ_{t≥6} d_t·(sel − v_rs2 − 4·words_t) with words_t the type's frame length: a recursion type's request leaves a0 past its frame (recursion.md §1.4).

3 Lookups#

15 obligations on three channels, none of them GENERIC, so the setup columns are the decoded table alone: the frame's ten TIMESTAMP gaps, two a query under its mask (memory.md §2), and five under m_pc:

lookup channel tuple
rd_hi_range, rd_lo_range RANGE16 rd_hi; sel − 2^16·rd_hi
next_pc_hi_range, next_pc_lo_range RANGE16 next_pc_hi; next_pc − 2^16·next_pc_hi
decode_row DECODER pc, decoded_next_pc, decoded_rs1, decoded_rs2, decoded_rd, decoded_imm, decoded_mask

The channels, in output order (add_sub::channels), are TIMESTAMP over V[range19], RANGE16 over V[range16] and DECODER over S[0..7].

4 Why it is sound#

On a live row (m_pc = 1) decode_row makes the claimed row the table's at pc, so exactly one b_k is 1 (lookup.md §10). The mask rules make each query present exactly where the row's kind or request makes it (execution-trace.md §4, §6). The address rules make a register query the decoded register, or on an ecall row, whose decoded registers are 0, a7 (17) for rs1 and a0 (10) for rs2 and rd. The _value_masked gates make an absent operand read 0, which lets one gate serve three sums: an addi or auipc row's v_rs2, and an auipc row's v_rs1, would otherwise be free addends, and add's imm is the table's 0.

Read values are words (memory-ops.md §5), sel is a word by its range pair and wrap is boolean, so each arithmetic gate is an identity over ℤ with one solution: the sum mod 2^32 and its carry, the difference mod 2^32 and its borrow, or imm. Without the pair, a sum at or above 2^32 would satisfy the gate with wrap = 0 and reach a register. The frame's x0 rule then writes sel or discards it.

next_pc is a word by its range pair and is decoded_next_pc — the table's fall-through, so a compressed instruction advances by 2 (program.md §5) — or HALT_PC on the exit row, less 2^32·pc_wrap. Both are far below 2^32, so pc_wrap = 0 on every live row.

On a system row system_split sets exactly one of is_ecall and is_fence, and the code gates make it the one imm names; ebreak's code 1 satisfies neither, so an ebreak row is unprovable. A fence row makes no query but the pc's and falls through. Off a system row both bits are 0, and so, by deleg_<f>_is_an_ecall, is every d_t.

A set d_t forces is_ecall = 1 and a7 = N_t. The numbers are pairwise distinct and none is 93, const assertions beside the circuit, so at most one d_t is set, is_exit is 0 or 1, and an ecall row is the exit, with a7 = 93, or a request of exactly one type; no other a7 passes.

  • The exit row writes back the a0 it read (exit_status), so x10's final value is the status the statement carries (proof.md §6, step 10b), and writes HALT_PC, after which no row runs (memory.md §5).
  • A request row falls through, writes 0 to a0, and makes the mirror query at the frame base it read from a0, in the space deleg_space names, reading timestamp 0 and value 0. Those three zeroings pair it one-to-one with an invocation of its type (delegation.md §5); the mirror's write value is free here, and what the call computed is the invoked family's circuit (delegation-circuits.md). deleg_space is an M column because a memory leaf may read no W column (memory.md §8); deleg_space_rule ties it to the selectors.

A row with m_pc = 0 is bound to no table row and its kind bits are free; the arithmetic gates are gated by those bits alone, m_pc times a bit being degree 3. That is harmless: the mask rules zero the row's other four masks and every lookup is off, so it adds no memory tuple. On a live row, wrap outside the four sums and sel on a fence row are free, and nothing reads them.

5 Limits#

  • An ecall whose a7 is neither 93 nor a type the format's circuit knows has no proof. The emulator answers an unassigned number -ENOSYS and continues (ecall-abi.md §5); the fill refuses that trace, naming the cycle.
  • An ebreak has no proof; it is fatal in the emulator (execution-trace.md §10).
  • No row touches RAM: an ecall row reads a7 and a0 and writes a0, and a request's operands travel in the invoked family's frame (delegation.md §4).

Auditors/Instruction families

The JUMP_BRANCH_SLT family

Normative specificationdocs/spec/jump-branch-slt.mdView as Markdown

The circuit of slti, sltiu, slt, sltu, the six branches, jalr and jal: what it adds beside the memory frame every execution family carries (memory.md §2), and the two gadgets other families reuse (§3). One comparison settles signed and unsigned order for the branches and the slt kinds alike. The circuit is constraints::jump_branch_slt::artifact (crates/constraints/src/jump_branch_slt.rs); prover::family_fill writes its witness.

1 What the circuit reads from the decoded table#

The tuple is pc next_pc rs1 rs2 rd imm extra_mask (program.md §5). next_pc is the fall-through, seq below; imm is the two's-complement word of the value the instruction uses: the sign-extended immediate of slti and sltiu (which sltiu compares unsigned), a branch's or jal's displacement, jalr's offset. extra_mask is one-hot over constants::extra_mask::jump_branch_slt:

bit    0     1      2    3     4    5    6    7    8     9     10    11
kind   slti  sltiu  slt  sltu  beq  bne  blt  bge  bltu  bgeu  jalr  jal

The legal masks are these twelve one-bit values, jump_branch_slt::LEGAL_MASKS; rd = x0 is the table's rd, not a mask. The circuit commits the twelve bits b_k, and every signal it needs is a linear form over them: the signed-comparison flag sc = b_slti + b_slt + b_blt + b_bge, the compared immediate (b_slti + b_sltiu)·imm, so that a branch's displacement never reaches the comparison, and the branch weights of taken_rule (§4).

2 Columns#

The frame is M[0..21] and W[0..7], over the queries pc rs1 rs2 rd at slots 0–3 (memory.md §2). Its W[6], rd_selected (sel below), holds the value the instruction computes, which the x0 rule masks into the write. The circuit adds:

column name value
W[7..13] decoded_next_pc … decoded_mask the claimed decoded row after pc
W[13..25] kind_slti … kind_jal the bits b_k, in §1's order
W[25] cmp_rhs the right operand, rs2 + (b_slti + b_sltiu)·imm
W[26], W[27] rs1_hi, rs1_sign rs1 >> 16, rs1 >> 31
W[28], W[29] cmp_rhs_hi, cmp_rhs_sign the same of cmp_rhs
W[30] lt rs1 < cmp_rhs, signed where sc = 1
W[31], W[32] cmp_gap, cmp_gap_hi (rs1 − cmp_rhs) mod 2^32, and its high halfword
W[33], W[34] eq, eq_inv [rs1 = cmp_rhs] on a live row; the difference's inverse
W[35] taken a taken branch
W[36] jalr_drop bit 0 of rs1 + imm on a jalr row
W[37] pc_wrap the carry out of whichever sum next_pc is
W[38], W[39] next_pc_hi, rd_hi next_pc >> 16, sel >> 16
W[40..44] mult_timestamp … mult_decoder one multiplicity per channel, in channel order
S[0..7] table_pc … table_extra_mask the decoded table, which program identity binds
S[7..10] generic_key … generic_result the packed table (lookup.md §9)
V[range19], V[range16] the TIMESTAMP and RANGE16 tables

21 M, 44 W and 10 S columns, 75 committed; 42 enforcing gates, the frame's 10 and §4's 32; 22 lookups: 8 TIMESTAMP, 11 RANGE16, 2 GENERIC, 1 DECODER, counts artifact asserts.

3 The gadgets#

constraints::gadgets returns gates and lookups as data. is_zero also builds the frame's x0 rule (memory.md §2) and MUL_DIV's zero tests (mul-div.md); the comparison also orders ATOMICS' minimum and maximum (memory-ops.md §6).

3.1 is_zero(x, inv, z, enable)#

x·inv + z − enable = 0          x = Σ c_i·x_i, a linear form
z·x = 0

With enable boolean, which the caller establishes, these force z = enable·[x = 0]: at x ≠ 0 the second gives z = 0 and the first inv = enable/x; at x = 0 the first gives z = enable. So z is boolean with no gate of its own, and enable = 0 gives z = 0, which keeps the all-zero row valid.

3.2 The comparison#

Comparison names one comparison lhs < rhs by its columns, its lookups' selector, and the kind bits signed whose sum is sc, which the caller holds to 0 or 1 on a selected row. comparison returns, for x each of lhs, rhs and gap:

name kind expression
<p>_order gate lhs − rhs − 2^32·sc·lhs_sign + 2^32·sc·rhs_sign + 2^32·lt − gap
<p>_lt_boolean gate lt − lt²
<p>_<x>_hi_range, <p>_<x>_lo_range RANGE16 x_hi; x − 2^16·x_hi
<p>_lhs_get_sign, <p>_rhs_get_sign GENERIC (x_hi + SIGN_BASE, x_sign, 0)

The range pairs make lhs, rhs and gap words and each x_hi the true high halfword (memory.md §7), which keeps each sign key inside U16GetSign's range (lookup.md §4), so each sign is its operand's bit 31. Let D = lhs − rhs − 2^32·sc·(lhs_sign − rhs_sign): both operands read in two's complement where sc = 1, so mixed signs are no case split, and D ∈ (−2^32, 2^32). The gate says gap = D + 2^32·lt, and only lt = [D < 0] puts gap in [0, 2^32): at D ≥ 0, lt = 1 puts it at 2^32 or above; at D < 0, lt = 0 makes it a negative field element. So the range check on gap carries the order, and no comparison table exists; the honest gap is (lhs − rhs) mod 2^32 whatever sc is. Both gates are ungated, since a selector would make the order gate degree 3, and every row satisfies them with the gap its own values give. comparison_equation(c, word_bits) builds the order gate at any width to 32, and the row suite evaluates it at 6 bits over every operand pair, signed and unsigned, finding exactly one (lt, gap), the ISA's.

4 Gates#

After the frame's ten in gate list 0, with m_q, a_q, v_q query q's mask, address and read value, and pc, next_pc the pc query's read and write:

gate polynomial
kind_<k>_boolean ×12 b_k − b_k²
decoded_mask_bits Σ_k 2^k·b_k − decoded_mask
rs1_mask_rule m_rs1 − m_pc·(Σ_k b_k − b_jal)
rs2_mask_rule m_rs2 − m_pc·(b_slt + b_sltu + the six branch bits)
rd_mask_rule m_rd − m_pc·(b_slti + b_sltiu + b_slt + b_sltu + b_jalr + b_jal)
<q>_addr_rule, for rs1, rs2, rd m_q·(a_q − decoded_q)
<q>_value_masked, for rs1, rs2 v_q − m_q·v_q
cmp_rhs_rule cmp_rhs − v_rs2 − (b_slti + b_sltiu)·imm
cmp_order, cmp_lt_boolean §3.2: lhs = v_rs1, rhs = cmp_rhs, signed the bits of sc
eq_inverse, eq_at_nonzero §3.1: x = v_rs1 − cmp_rhs, z = eq, enable = m_pc
taken_rule taken − w_1 − w_eq·eq − w_lt·lt
taken_boolean, jalr_drop_boolean, pc_wrap_boolean x − x²
next_pc_rule §5's equation
rd_value_rule sel − (b_jal + b_jalr)·seq − (b_slti + b_sltiu + b_slt + b_sltu)·lt

The branch weights are w_1 = b_bne + b_bge + b_bgeu, w_eq = b_beq − b_bne and w_lt = b_blt + b_bltu − b_bge − b_bgeu. Every gate has degree at most 2 and is 0 on the all-zero row, which artifact asserts.

4.1 Lookups#

After the frame's 8 TIMESTAMP obligations, all under m_pc:

lookup channel expression
cmp_<x>_hi_range, cmp_<x>_lo_range ×3 RANGE16 §3.2 over v_rs1, cmp_rhs, cmp_gap
cmp_lhs_get_sign, cmp_rhs_get_sign GENERIC §3.2
rd_hi_range, rd_lo_range RANGE16 rd_hi; sel − 2^16·rd_hi
next_pc_hi_range, next_pc_lo_range RANGE16 next_pc_hi; next_pc − 2^16·next_pc_hi
next_pc_even RANGE16 2^−1·next_pc − 2^15·next_pc_hi
decode_row DECODER pc and W[7..13] (lookup.md §10)

The channels, in output order, are TIMESTAMP on V[range19], RANGE16 on V[range16], GENERIC on S[7..10] and DECODER on S[0..7]. next_pc_even is the low halfword lo halved, (lo + p)/2 and far above 2^16 when lo is odd. Because it scales next_pc, the constructor runs lookup::check_copowers (lookup.md §11) over (next_pc, m_pc).

5 Why it is sound#

On a live row, m_pc = 1, the decoder lookup makes the claimed tuple the table's row at pc, so pc is even, seq is below 2^24 and exactly one b_k is 1 (lookup.md §10). The mask and address rules make the frame's queries the instruction's (execution-trace.md §4): jal reads nothing, and a branch has no rd query, so nothing it computes is written. An absent operand reads 0, so cmp_rhs is rs2 or the immediate, never their sum, and the comparison's pairs make both operands words. So lt is the ISA's order (§3.2), eq its equality (§3.1), and taken its branch decision: eq on beq, 1 − eq on bne, lt on blt and bltu, 1 − lt on bge and bgeu, and 0 off the branches, every term of taken_rule carrying a branch bit. taken is a committed bit because, inlined, taken·(pc + imm) would be degree 3.

next_pc is held by one gate:

next_pc + 2^32·pc_wrap = (1 − taken − b_jal − b_jalr)·seq
                       + (taken + b_jal)·(pc + imm)
                       + b_jalr·(v_rs1 + imm − jalr_drop)

At most one of taken, b_jal, b_jalr is 1, so one sum is selected, and one wrap bit outside the selectors serves all three: imm is a two's-complement word, so every backward branch and jump wraps, not only jalr. With next_pc an even word and pc_wrap, jalr_drop boolean:

  • the default arm is seq, below 2^24, so pc_wrap = 0;
  • pc + imm and v_rs1 + imm are below 2^33, so one wrap bit holds the carry, uniquely;
  • on jalr, v_rs1 + imm − jalr_drop − 2^32·pc_wrap is a unique even word, (rs1 + imm) mod 2^32 with bit 0 cleared; a false jalr_drop makes next_pc odd or negative.

A branch's or jal's target is even unchecked, pc and imm both being even. Evenness is what keeps the family off HALT_PC = 1 (memory.md §5): without next_pc_even, a jalr whose rs1 + imm ≡ 1 keeps bit 0 and writes HALT_PC with every other gate and lookup holding, and a program that would crash by jumping to address 0 is proven to exit cleanly.

The link is seq, a table value and not a sum, so it has no wrap bit; the rd pair range-checks it and lt like every register write, and the x0 rule masks both at x0. A target needs no check of its own: at an address holding no instruction, the next row's decoder lookup fails whatever family claims the row, no table holding a live row there (lookup.md §10).

On a padding row, m_pc = 0, the mask rules zero every query mask, eq is 0 and every lookup is off, so the row reaches no memory event whatever its free bits hold.

Auditors/Instruction families

The SHIFT_BITWISE family

Normative specificationdocs/spec/shift-bitwise.mdView as Markdown

The circuit of the shifts sll, slli, srl, srli, sra, srai and the bitwise and, andi, or, ori, xor, xori, one family, beside the memory frame (memory.md §2). A shift either way is one product with a looked-up power of two; AND is four byte lookups, and OR and XOR are linear forms over it. The circuit is constraints::shift_bitwise::artifact (crates/constraints/src/shift_bitwise.rs); prover::family_fill writes its witness.

1 What the circuit reads from the decoded table#

The tuple is pc next_pc rs1 rs2 rd imm extra_mask (program.md §5), next_pc the fall-through, seq below. imm is the shamt of slli, srli and srai, below 32 because the decoder refuses shamt[5] on RV32; the sign-extended immediate, as a word, of andi, ori and xori; and 0 on a register form. extra_mask is one-hot over constants::extra_mask::shift_bitwise, the legal masks its twelve one-bit values (shift_bitwise::LEGAL_MASKS):

bit    0     1     2     3     4    5     6    7    8    9    10   11
kind   slli  xori  srli  srai  ori  andi  sll  xor  srl  sra  or   and

The second operand of all twelve is src2 = rs2 + imm: an immediate form has no rs2 query, so rs2 reads 0, and a register form's imm is 0. One addend is always zero, so the sum needs no wrap bit, and an immediate never enters the rs2 column the memory argument ties. The circuit commits the twelve bits b_k, and its flags are linear forms over them:

left    b_slli + b_sll
right   b_srli + b_srai + b_srl + b_sra
arith   b_srai + b_sra
t1      b_or + b_ori + b_xor + b_xori
t2      b_and + b_andi − b_or − b_ori − 2·(b_xor + b_xori)

Only the two halves' sums, f_shift and f_bitwise, are columns: each selects lookups, and a selector is a committed boolean (lookup.md §2).

2 Columns#

The frame is M[0..21] and W[0..7], over the queries pc rs1 rs2 rd at slots 0–3 (memory.md §2); its W[6], rd_selected (sel below), holds the value the instruction computes, which the x0 rule masks into the write. The circuit adds:

column name value
W[7..13] decoded_next_pc … decoded_mask the claimed decoded row after pc
W[13..25] kind_slli … kind_and the bits b_k, in §1's order
W[25], W[26] f_shift, f_bitwise the two halves
W[27], W[28] rs1_hi, rs1_sign rs1 >> 16, rs1 >> 31
W[29] src2_hi src2 >> 16
W[30] amount src2 & 31
W[31], W[32] pow, copow 2^amount, 2^(31 − amount) on a shift row
W[33], W[34] high, high_hi src2 >> 5, and its high halfword
W[35] se arith·rs1_sign
W[36], W[37] shift_in, shift_prod both directions' multiplicand, and shift_in·pow
W[38], W[39] ovf, ovf_hi a left shift's discarded high word, and its high halfword
W[40], W[41] residue, residue_hi a right shift's remainder, and its high halfword
W[42], W[43] scaled, scaled_hi residue·2^(32 − amount), and its high halfword
W[44..52] byte_a<j>, byte_b<j> the bytes of rs1, then of src2, low first
W[52..56] byte_and<j> their bytewise AND
W[56] rd_hi sel >> 16
W[57..61] mult_timestamp … mult_decoder one multiplicity per channel, in channel order
S[0..10], V[range19], V[range16] as in jump-branch-slt.md §2

21 M, 61 W and 10 S columns, 92 committed; 48 enforcing gates, the frame's 10 and §4's 38; 39 lookups: 8 TIMESTAMP, 24 RANGE16, 6 GENERIC, 1 DECODER, counts artifact asserts.

3 Tables, and the bound on every key#

3.1 ShiftPowers#

Row s of the packed table's top sub-table (lookup.md §9) is (SHIFT_BASE + s + 1, 2^s, 2^(31 − s)), one for each of the 32 shift amounts and for none other, so a key past its last row matches nothing. The second value is the copower a residue bound multiplies by, 2^(32 − s), stored halved (SHIFT_COPOWER_BITS = 31): at s = 0 it is 2^32, which the table's u32 columns cannot hold, so the two gates that read it carry the factor 2 (§4.2, §4.3).

3.2 The AND rows#

An AND row is (AND_BASE + a + 1, b, a & b) over bytes a and b, so a key inside their range matches a row that makes byte_b<j> a byte and byte_and<j> its AND with byte_a<j>: those two need no bound of their own.

3.3 Every key is bounded#

The channel proves membership of the packed table, not of a sub-table (lookup.md §4), so an out-of-range key lands on another sub-table's row. A bitwise row with byte_a0 = 65,823 gates to key 65,824, ShiftPowers' row (65,824, 2^31, 1); with rs1 = 65,823 and rs2 = 2^31 every gate holds, and and writes 1 where the answer is 0. So every key carries its own bound, as RANGE16 obligations under its lookup's selector:

key bound obligations selector
rs1_hi + SIGN_BASE rs1_hi < 2^16 rs1's 16+16 pair m_pc
amount + SHIFT_BASE amount < 2^5 amount; 2^11·amount f_shift
byte_a<j> + AND_BASE byte_a<j> < 2^8 byte_a<j>; 2^8·byte_a<j> f_bitwise

A bound below a halfword takes both obligations: the scaled one alone does not make the key an integer (lookup.md §11), and the direct one alone admits every halfword, byte_a0 = 256 landing on U16GetSign's row (257, 0, 0).

3.4 The copower check#

artifact runs lookup::check_copowers (lookup.md §11) over each column it bounds by scaling, which must carry its direct bound under its scaled obligation's own selector: residue, scaled by the looked-up copower (§4.3), under m_pc; amount under f_shift; each byte_a<j> under f_bitwise.

4 Gates#

Gate list 0 holds the frame's ten and these 38. m_q, a_q, v_q are query q's mask, address and read value, and a flag of §1 times (…) stands for each of its weighted bits times (…), so every term is of degree 2.

4.1 Presence and next_pc#

gate polynomial
kind_<k>_boolean ×12, decoded_mask_bits as in jump-branch-slt.md §4
f_shift_rule, f_bitwise_rule f − Σ its half's six bits
f_shift_boolean, f_bitwise_boolean f − f²
rs1_mask_rule, rd_mask_rule m_q − m_pc·Σ_k b_k
rs2_mask_rule m_rs2 − m_pc·(b_sll + b_srl + b_sra + b_and + b_or + b_xor)
<q>_addr_rule ×3, <q>_value_masked ×2 as in jump-branch-slt.md §4
next_pc_rule next_pc − seq

No kind computes a pc: next_pc is the decoder-bound fall-through, with no wrap bit and no bound of its own, and HALT_PC is beyond the family's reach (memory.md §5).

4.2 The shift amount#

amount_split    rs2 + imm − 32·high − amount
copower_rule    pow·copow − 2^31·f_shift

amount_split is ungated. copower_rule says pow·(2·copow) = 2^32 on a shift row, and pow·copow = 0 on a bitwise row.

4.3 The one product, both directions#

se_rule           se − arith·rs1_sign
rs1_sign_boolean  rs1_sign − rs1_sign²
se_boolean        se − se²
shift_in_rule     shift_in − left·v_rs1 − right·(sel − 2^32·se)
shift_prod_rule   shift_prod − shift_in·pow
shift_out_rule    left·(shift_prod − sel − 2^32·ovf)
                    + right·(shift_prod + residue − v_rs1 + 2^32·se)
scaled_rule       scaled − 2·residue·copow

shift_prod_rule, ungated, is the one multiplication by pow; shift_in_rule picks its multiplicand, which keeps shift_out_rule at degree 2 where left·(v_rs1·pow − …) would be 3, and se is committed for the same reason. A right shift is the floor division rs1 − 2^32·se = (sel − 2^32·se)·2^s + residue, which covers sra: the arithmetic shift of a negative word is the floor division of its signed value, and the result keeps the operand's sign. shift_in and shift_prod are the only columns that are not words: the multiplicand is negative where se = 1, and a left shift's product reaches 2^63.

4.4 The bitwise half#

rs1_bytes         v_rs1 − Σ_j 2^(8j)·byte_a<j>
src2_bytes        rs2 + imm − Σ_j 2^(8j)·byte_b<j>
bitwise_out_rule  f_bitwise·sel − t1·(v_rs1 + rs2 + imm) − t2·Σ_j 2^(8j)·byte_and<j>

Per byte, OR is a + b − (a & b) and XOR is a + b − 2·(a & b). Summed by weight through the two decompositions, sel is rs1 & src2 at (t1, t2) = (0, 1), their OR at (1, −1) and their XOR at (1, −2), exactly, no carry crossing a byte: there is no OR or XOR table, and the AND accumulator is a linear form, not a column. sel is gated by f_bitwise because t1 and t2 are 0 on a shift row, where a bare sel would force rd = 0. The decompositions are ungated: on a shift row the bytes carry no lookup, and a decomposition always exists.

4.5 Lookups#

After the frame's 8 TIMESTAMP obligations, in the channel order of jump-branch-slt.md §4.1:

RANGE16   <x>_hi_range, <x>_lo_range     under m_pc, x = rs1 src2 high ovf residue scaled rd
          amount_range, amount_scaled    under f_shift      §3.3
          byte_a<j>_range, _scaled ×4    under f_bitwise    §3.3
GENERIC   rs1_get_sign   (rs1_hi + SIGN_BASE, rs1_sign, 0)                 under m_pc
          shift_powers   (amount + SHIFT_BASE, pow, copow)                 under f_shift
          and_byte_<j>   (byte_a<j> + AND_BASE, byte_b<j>, byte_and<j>)    under f_bitwise
DECODER   decode_row     under m_pc (lookup.md §10)

5 Why it is sound#

On a live row the decoder lookup makes the claimed tuple the table's row at pc, so one kind bit is 1 and one of f_shift, f_bitwise (lookup.md §10); the mask and address rules make the queries the instruction's (execution-trace.md §4). rs1, src2 and sel are words by their pairs, rs1_hi is rs1's true high halfword and rs1_sign its bit 31. Every term of §4 that reads sel carries a shift bit or f_bitwise, so the inactive half never constrains it.

  • The amount is the ISA's. §3.3 bounds amount below 32 and high's pair bounds high below 2^32, so amount_split is an integer identity below 2^37, amount = src2 mod 32, and the ShiftPowers row it keys gives pow = 2^amount. Without high's pair, sll by rs2 = 4 can shift by 8, at high = −1/8.
  • A left shift: shift_prod = rs1·2^s < 2^63, and sel + 2^32·ovf, both words, is its unique split, so sel = (rs1·2^s) mod 2^32.
  • A right shift: se is rs1's bit 31 on sra and srai and 0 otherwise, so se_rule alone keeps an srai from carrying srli's answer. Every term of the floor division is below 2^64 in magnitude, so residue is the integer (rs1 − 2^32·se) − (sel − 2^32·se)·2^s, and scaled's pair puts it in [0, 2^s): sel − 2^32·se is the floor of (rs1 − 2^32·se)/2^s. residue's own pair, which check_copowers requires, bounds it without appeal to sel's, the scaled pair alone saying nothing of a non-integer: 2^−28 passes it at s = 3.
  • A bitwise result: each byte_a<j> is below 256, so its lookup matches an AND row (§3.2); with rs1 and src2 words, both decompositions are the unique byte splits and §4.4's identity holds.

copower_rule is implied by the bounded key and kept as the circuit's own reading of the table: a ShiftPowers row generated wrong stops the honest prover rather than license a residue bound that is not one. It also confines the key to ShiftPowers alone, no other row's two values having the product 2^31: an AND row's is at most 255·255, every other row's 0.

On a padding row, m_pc = 0, every query mask is 0 and every obligation under m_pc vacuous. f_shift and f_bitwise are free booleans there, so a padding row may look up ShiftPowers or the AND rows, which consumes a multiplicity and changes nothing.

Auditors/Instruction families

The MUL_DIV family

Normative specificationdocs/spec/mul-div.mdView as Markdown

The M extension — mul, mulh, mulhsu, mulhu, div, divu, rem, remu — as one circuit beside the memory frame every execution family carries (memory.md §2). One product identity serves the four multiplies and the division; a sign rule and a range-checked gap make the division truncated, and one gate pins division by zero. constraints::mul_div builds it: 21 M, 54 W and 9 S columns, 54 enforcing gates, 27 lookups.

1 What the circuit reads from the decoded table#

Every M instruction is R-type, so the decoded tuple has no imm: pc next_pc rs1 rs2 rd extra_mask, six columns (program.md §5). extra_mask is one-hot over constants::extra_mask::mul_div, bits 0–7 in the order above; the legal masks are its eight single bits (mul_div::LEGAL_MASKS), which the table's domain enforces (lookup.md §10). The circuit commits the bits b_k and reads every signal as a linear form over them:

signal form
reads rs1 signed b_mul + b_mulh + b_mulhsu + b_div + b_rem
reads rs2 signed b_mul + b_mulh + b_div + b_rem
a multiply, Σ_mul b_mul + b_mulh + b_mulhsu + b_mulhu
a division, f_div b_div + b_divu + b_rem + b_remu, a column: the is-zero gadgets' enable

mul is read signed × signed: its low word is the same either way, which lets one product identity serve all four multiplies. mulhsu's asymmetry is the two lists, not a case split.

2 Columns#

The frame is pc rs1 rs2 rd, M[0..21] and W[0..7] (memory.md §2.1); below, m_q, a_q and v_q are query q's mask, address and read value, and rs1, rs2 the operands' read values. The family adds:

W[7..12]   decoded_next_pc decoded_rs1 decoded_rs2 decoded_rd decoded_mask    (no imm)
W[12..20]  kind_mul … kind_remu
W[20]      f_div
W[21..27]  rs1_hi rs1_top rs2_hi rs2_top s1 s2          high halfwords, bit 31, §3
W[27..34]  mx my p_low p_low_hi p_high p_high_hi p_sign  the product
W[34..40]  q q_hi q_sign r r_hi r_sign                   quotient and remainder
W[40..45]  r_inv rz d1 d_inv dz                          is_zero(r), f_div·s1, is_zero(rs2)
W[45..50]  abs_r abs_d gap gap_hi rd_hi
W[50..54]  mult_timestamp mult_range16 mult_generic mult_decoder
S[0..6]    the decoded table, bound by identity
S[6..9]    the packed generic table (lookup.md §9)
V          range19 range16

The fill keeps mx, my (signed) and r_inv, d_inv (inverses) in Fr, every other column in u32.

3 The sign adjustments#

rs1_adj = rs1 − 2^32·s1      s1 = (b_mul + b_mulh + b_mulhsu + b_div + b_rem)·rs1_top
rs2_adj = rs2 − 2^32·s2      s2 = (b_mul + b_mulh + b_div + b_rem)·rs2_top
q_adj   = q − 2^32·q_sign    r_adj = r − 2^32·r_sign

rs1_top is the U16GetSign lookup of rs1_hi, which rs1's 16+16 pair makes its true high halfword, so the key lies in that sub-table's range and the answer is bit 31 (lookup.md §4); rs2_top likewise. An unsigned position forces its adjustment to 0 whatever the top bit, which keeps the selection degree 2. q_sign and r_sign are not sign lookups (§5.3). f_div, rs1_top, rs2_top, s1, s2, p_sign, q_sign and r_sign carry booleanity gates; rz and dz are boolean by the is-zero gadget (jump-branch-slt.md §3), d1 as a product of booleans.

4 Gates#

The frame's ten enforcing gates (memory.md §2.4) and the family's 44, all in gate list 0, each formula = 0. The plumbing:

kind_<k>_boolean          b_k − b_k²                       eight
decoded_mask_bits         Σ_k 2^k·b_k − decoded_mask
<q>_mask_rule             m_q − m_pc·Σ_k b_k               rs1, rs2, rd: every kind uses all three
<q>_addr_rule             m_q·(a_q − decoded_q)            rs1, rs2, rd
<q>_value_masked          v_q − m_q·v_q                    rs1, rs2
next_pc_rule              next_pc − decoded_next_pc        the fall-through (memory.md §5)

The arithmetic, mul_div::arithmetic_gates(32), written with §3's abbreviations:

f_div_rule, s1_rule, s2_rule   §1's and §3's forms, and eight booleanity gates (§3)
mx_rule                mx − Σ_mul b·rs1_adj − f_div·rs2_adj
my_rule                my − Σ_mul b·rs2_adj − f_div·q_adj
product_rule           mx·my − p_low − 2^32·p_high + 2^64·p_sign
division_rule          f_div·(p_low + 2^32·p_high − 2^64·p_sign + r_adj − rs1_adj)
rz_inverse             r·r_inv + rz − f_div          rz_at_nonzero   rz·r
dz_inverse             rs2·d_inv + dz − f_div        dz_at_nonzero   dz·rs2
d1_rule                d1 − f_div·s1
r_sign_rule            r_sign − d1 + d1·rz           so r_sign = f_div·s1·(1 − [r = 0])
abs_r_rule             abs_r − r − 2^32·r_sign + 2·r·r_sign       abs_r = |r_adj|
abs_d_rule             abs_d − rs2 − 2^32·s2 + 2·rs2·s2           abs_d = |rs2_adj|
gap_rule               gap − f_div·(abs_d − abs_r − 1) − 2^32·dz
zero_divisor_quotient  dz·(q − (2^32 − 1))
rd_value_rule          rd_selected − b_mul·p_low − (b_mulh + b_mulhsu + b_mulhu)·p_high
                         − (b_div + b_divu)·q − (b_rem + b_remu)·r

The lookups: the frame's eight TIMESTAMP gap chunks, each under its query's mask; and under m_pc, 16+16 RANGE16 pairs on rs1, rs2, p_low, p_high, q, r, gap and rd_selected, rs1_get_sign, (rs1_hi + SIGN_BASE, rs1_top, 0) on GENERIC, and rs2_get_sign, and decode_row on DECODER.

The width is a parameter of arithmetic_gates so the encoding can be checked whole: crates/checker/tests/mul_div.rs evaluates arithmetic_gates(4) through gkr::eval_gate over every (dividend, divisor) pair of a 4-bit word and each division kind, and exactly one (q, r) survives, RV32M's.

5 Why it is sound#

On a live row the decoder lookup makes exactly one kind bit 1 (lookup.md §10), and rs1, rs2 are words whose _top is bit 31, so rs1_adj, rs2_adj ∈ [−2^31, 2^32) are the operands as the kind reads them.

5.1 The product#

On a multiply row mx·my = rs1_adj·rs2_adj; on a division row it is rs2_adj·q_adj, q's pair and q_sign's booleanity putting q_adj in [−2^32, 2^32). Either way |mx·my| < 2^64, and two words and a boolean cover [−2^64, 2^64) once, so product_rule holds over the integers with one solution: p_low, p_high are the words of the 64-bit two's-complement product, RV32M's for each multiply. product_rule is ungated and the circuit's only product of two row values, which is what lets both readings share it at degree 2.

5.2 The division#

With rs2_adj ≠ 0, division_rule is rs2_adj·q_adj + r_adj = rs1_adj over the integers. Truncated division is its one solution with |r_adj| < |rs2_adj| and r_adj zero or of the dividend's sign, and two gates state exactly that:

  • The sign. r_sign = f_div·s1·(1 − [r = 0]) makes r_adj the word r on an unsigned row or a non-negative dividend, and r − 2^32 < 0 on a negative one unless r = 0. It is what separates truncated division from floored: without it DIV(−7, 2) admits q = −4, r = 1 as readily as q = −3, r = −1. As a definition, through d1, it is degree 2.
  • The magnitude. gap = abs_d − abs_r − 1 is range-checked, and neither magnitude reaches 2^32: abs_d ≤ 2^31 where s2 = 1, abs_r ≤ 2^32 − 1 where r_sign = 1, which needs r ≠ 0, and each is a word elsewhere. So the difference lies in [−2^32, 2^32), in range exactly when |r_adj| < |rs2_adj|. The comparison gadget would repeat bounds that hold and has no place for the zero divisor's term.

So q_adj and r_adj are RV32M's, and q_sign is pinned only by q's range: one value puts q_adj + 2^32·q_sign in [0, 2^32).

A zero divisor makes dz = 1 and rs2_adj = 0: the identity leaves r_adj = rs1_adj, so r is the dividend's word; zero_divisor_quotient, the one pin, makes q all ones; the 2^32·dz term lifts gap to 2^32 − 1 − abs_r, so the divisor imposes no bound. q_sign is free and harmless: mx = 0, and rd reads the word q.

The identity is gated. On a multiply row r_sign = 0 and r is a word, so an ungated identity would demand rs1_adj − rs1_adj·rs2_adj ∈ [0, 2^32), false for nearly every multiply: 7 × 3, a negative rs1 times x0.

5.3 The signed overflow, and why q_sign is free#

DIV(−2^31, −1) needs no pin: |r_adj| < 1 forces r = 0, the identity q_adj = 2^31, and q's range q_sign = 0, q = 0x80000000, RV32M's answer; REM gives 0. This row is why q_sign is a free boolean: tied to bit 31 of q, as s1 and s2 are to their operands', it would force q_adj = −2^31 and make the row unprovable. r_sign likewise follows the dividend's sign, not the remainder's word.

rd_selected's pair is implied by its four sources' and kept, every family bounding what it writes to rd (memory-ops.md §5). Every gate is zero on the all-zero padding row, which mul_div::artifact asserts with each channel's obligation count.

5.4 The fill#

prover::family_fill(MUL_DIV) (crates/prover/src/fill.rs) computes the witness with Rust's integers: the product in i128, the division by wrapping_div and wrapping_rem, which give RV32M's overflow answer, with the zero divisor an arm of its own, and q_sign from the sign of q_adj. It writes the computed value to rd_selected, which the frame's x0 rule masks, and panics, on rows the emulator cannot produce, if the identity does not divide, a product exceeds two words, or the trace's rd write or next_pc is not what the instruction computes.

Auditors/Instruction families

The memory-op families

Normative specificationdocs/spec/memory-ops.mdView as Markdown

MEM_WORD (lw, sw), MEM_SUBWORD (lb, lh, lbu, lhu, sb, sh) and ATOMICS (lr.w, sc.w, the nine AMOs): the execution families whose rows touch RAM, each a circuit beside the memory frame (memory.md §2), sharing §2's addressing. They are constraints::{mem_word, mem_subword, atomics}, filled by prover::family_fill (crates/prover/src/fill.rs), which computes each witness with Rust's integer operations.

family M W S gates, frame + own TIMESTAMP, RANGE16, GENERIC, DECODER
MEM_WORD 31 24 7 13 + 20 12, 5, 0, 1
MEM_SUBWORD 31 55 10 13 + 40 12, 22, 1, 1
ATOMICS 26 54 9 11 + 35 10, 19, 6, 1

Each artifact asserts its gate and obligation counts and that the all-zero padding row satisfies every gate. Below, m_q, a_q and v_q are query q's mask, address and read value, rs1 and rs2 the operands' read values (memory.md §2.1), and b_k (b_lw, b_lr, …) the committed kind bits.

1 What the circuits read from the decoded table#

MEM_WORD's and MEM_SUBWORD's tuple is pc next_pc rs1 rs2 rd imm extra_mask, imm the offset's two's-complement u32; ATOMICS' has no imm, its address being rs1 (program.md §5). The tuple is the first setup columns, and the packed generic table follows it where a family reads one: S[7..10] in MEM_SUBWORD, S[6..9] in ATOMICS (lookup.md §9). extra_mask is one-hot over constants::extra_mask, bit k the k-th mnemonic below, and each module's LEGAL_MASKS is those single bits:

mem_word      lw sw
mem_subword   lb lh lbu lhu sb sh
atomics       amoadd amoswap lr sc amoxor amoor amoand amomin amomax amominu amomaxu

The atomics order is ascending funct5; aq and rl order nothing on one hart and are not recorded. MEM_SUBWORD's modifiers are linear forms over its bits:

LOADK = b_lb + b_lh + b_lbu + b_lhu     BYTE = b_lb + b_lbu + b_sb     SIGNEXT = b_lb + b_lh
STORE = b_sb + b_sh                     HALF = b_lh + b_lhu + b_sh

All three carry the same plumbing, each formula = 0:

kind_<k>_boolean    b_k − b_k²
decoded_mask_bits   Σ_k 2^k·b_k − decoded_mask
<q>_mask_rule       m_q − m_pc·uses_q              every query but pc
<q>_addr_rule       m_q·(a_q − decoded_q)          rs1, rs2, rd
                    m_q·(a_q − 4·word_index)       load, ram (§2)
<q>_value_masked    v_q − m_q·v_q                  rs1, rs2
next_pc_rule        next_pc − decoded_next_pc      the fall-through (memory.md §5)
uses_q rs1 rs2 load ram rd
MEM_WORD b_lw + b_sw b_sw b_lw b_sw b_lw
MEM_SUBWORD LOADK + STORE STORE LOADK STORE LOADK
ATOMICS every bit every bit but b_lr no query every bit every bit

m_rs2 is keyed on b_lr, the one kind without an rs2 field, and not on rs2 = x0: an amoadd.w whose rs2 is x0 still reads it.

2 Addressing#

The effective address is rs1 + imm mod 2^32, or rs1 for an atomic. One degree-1 gate splits it, with wrap, bit0 and bit1 boolean:

MEM_WORD      addr_split   rs1 + imm − 2^32·wrap − 4·word_index
MEM_SUBWORD   addr_split   rs1 + imm − 2^32·wrap − 4·word_index − 2·bit1 − bit0
ATOMICS       addr_word    rs1 − 4·word_index

Over Fr that says nothing, 4 being a unit. Three RANGE16 obligations under m_pc, on word_index_hi, word_index − 2^16·word_index_hi and 4·word_index_hi (word_index_hi_range, word_index_lo_range, word_index_hi_scaled), cap word_index at 2^30 − 1, the top word's. With rs1 a word (§5) and imm a table value the split is then one of integers: wrap is the true carry, bit1 and bit0 the true low bits, and every RAM address is a 4-aligned address below 2^32. Having no offset bits, a misaligned MEM_WORD or ATOMICS access needs a word_index that is not an integer, which its pair refuses; the emulator refuses it first (execution-trace.md §10). addr_word derives rs1 < 2^32 rather than assuming it. half_aligned, HALF·bit0 = 0, refuses a halfword at an odd address and keeps w·p a divisor of 2^32 (§4.3).

Every RAM query's address is 4·word_index, so byte, halfword, word and atomic accesses to one word name one cell; the byte position lives only in MEM_SUBWORD's splice. Confining an access to initialized memory is the multiset's (memory.md §9): an out-of-window access fails the statement's memory argument, not a gate.

3 MEM_WORD#

A load copies the word into rd, a store copies rs2 into the word; there is no splice, no generic lookup, and the decoded table is the only setup.

W[9..15]   decoded_next_pc decoded_rs1 decoded_rs2 decoded_rd decoded_imm decoded_mask
W[15..21]  kind_lw kind_sw wrap word_index word_index_hi rd_hi
W[21..24]  mult_timestamp mult_range16 mult_decoder

wrap_boolean, addr_split (§2)
rd_value_rule       rd_selected − b_lw·load_read_value
store_value_rule    ram_write_value − m_ram·rs2

Its RANGE16 obligations are §2's three and the 16+16 pair on rd_selected, and the two copies are its whole semantics. rd_selected is range-checked although it copies a RAM word, because a RAM word need not be a word, advice's initial values being bound to nothing (public-values.md §6): the pair keeps every register value a word without reference to RAM (§5). No gate reads ram_read_value, the word a store overwrites; the memory argument alone pins it.

4 MEM_SUBWORD#

4.1 The splice#

A sub-word's position in its word lives only in

word = high·(w·p) + sub·p + low      p = 2^(8·offset), offset = 2·bit1 + bit0
                                     w, the access width: 2^8 if BYTE, 2^16 if HALF

p and its copower are degree-2 forms in the offset bits, written as gates rather than looked up:

p_rule        p − m_pc − 255·bit0 − 65535·bit1 − K·bit0·bit1      K = 2^24 − 2^16 − 2^8 + 1
pcopow_rule   p·pcopow − 2^31·m_pc                                 pcopow = 2^31/p
wph_rule      wph − 32768·p + 32640·BYTE·p                         wph = w·p/2
p_ram_rule    p_ram − m_ram·p

p_rule takes the four offsets to 1, 2^8, 2^16, 2^24, m_pc standing for the constant so the all-zero row satisfies it. The copower and w·p are stored halved so that 2^32 fits a u32 column, the gates reading them carrying the factor 2, as ShiftPowers' do (lookup.md §9). p_ram keeps store_rule degree 2. A table keyed by the offset would pin nothing addr_split does not, and add a key to bound (lookup.md §4).

4.2 Columns and gates#

W[9..21]   the decoded row; kind_lb … kind_sh
W[21..31]  wrap word_index word_index_hi bit0 bit1 p pcopow wph p_ram word
W[31..42]  high high_hi high_scaled high_scaled_hi sub sub_scaled sub_scaled_hi
           low low_hi low_scaled low_scaled_hi
W[42..51]  src_sub src_sub_scaled src_sub_scaled_hi src_high src_high_hi sign_in sign se rd_hi
W[51..55]  the four multiplicities

Its gates, beside the plumbing: wrap_boolean, bit0_boolean, bit1_boolean, addr_split, half_aligned, §4.1's four, and

word_rule            word − LOADK·load_read_value − STORE·ram_read_value
splice_rule          word − high_scaled − sub·p − low
high_scaled_rule     high_scaled − 2·high·wph                             = high·w·p
sub_scaled_rule      sub_scaled − 2^16·sub − (2^24 − 2^16)·BYTE·sub       = sub·2^32/w
low_scaled_rule      low_scaled − 2·low·pcopow                            = low·2^32/p
src_sub_rule         rs2 − src_sub − 2^16·src_high + 65280·BYTE·src_high
src_sub_scaled_rule  src_sub_scaled − 2^16·src_sub − (2^24 − 2^16)·BYTE·src_sub
store_rule           ram_write_value − m_ram·word − (src_sub − sub)·p_ram
sign_in_rule         sign_in − sub − 255·BYTE·sub                         = 2^8·sub or sub
se_rule              se − SIGNEXT·sign
rd_value_rule        rd_selected − LOADK·sub − (2^32 − 2^16)·se − 65280·BYTE·se

mem_subword::splice_gates(byte_bits) builds the twelve whose literals depend on the byte width — §4.1's first three and these but word_rule and se_rule — and the circuit takes it at BYTE_BITS = 8. Its RANGE16 obligations, all under m_pc, are §2's three, 16+16 pairs on high, high_scaled, sub_scaled, low, low_scaled, src_sub_scaled, src_high and rd_selected, and one obligation each on sub, src_sub and sign_in; its GENERIC lookup is sub_get_sign, (sign_in + SIGN_BASE, sign, 0).

4.3 Why it is sound#

§2 fixes the offset bits and half_aligned clears bit0 at halfword width, so p and w are the access's. Each part has a direct bound and a scaled one: high < 2^32 makes high·w·p an integer, sub_scaled < 2^32 is sub < w and low_scaled < 2^32 is low < p. So splice_rule holds over ℤ with one solution, the base-(p, w) digits of the word, and a word not below 2^32 has none. A scaled bound alone admits non-integers, its scale being a unit of Fr (lookup.md §11); constraints::lookup::check_copowers holds word_index_hi, high, sub, low and src_sub to their direct bounds, one obligation being exact for sub and src_sub, both below w ≤ 2^16. src_high's pair makes rs2 = src_sub + w·src_high integral, so src_sub is rs2 mod w: without it sb could store a byte unrelated to rs2.

A load writes rd = sub + (2^32 − w)·se: the sub-word, or at se = 1 its two's-complement extension (lb of 0x88 is 0xffffff88). sign_in is 2^8·sub for a byte and sub for a halfword, so its bit 15 is the sign at either width and one U16GetSign lookup serves both; its own obligation bounds the key into that sub-table (lookup.md §4). se is a one-hot sum times a table bit, boolean without a gate.

A store writes word + (src_sub − sub)·p = high_scaled + src_sub·p + low, a word with no appeal to memory: high_scaled is a multiple of w·p below 2^32 and w·p divides 2^32 (a halfword at offset 3 would make it 2^40; half_aligned excludes it), so high_scaled ≤ 2^32 − w·p and src_sub·p + low ≤ w·p − 1. That is why high_scaled keeps its own pair.

crates/checker/tests/mem_subword.rs checks the splice whole at a 4-bit word: for every word, admissible offset and width, splice_gates(1) and the bounds admit exactly one (high, sub, low).

5 The write-side induction#

A circuit may use a register operand as a word without bounding it. That rests on two facts:

  • Every register write is a word on its own row. Every execution family's rd_selected carries a 16+16 pair under m_pc, but ATOMICS', which is the old word or 0, the old word bounded by its comparison's pair under m_pc (§6). The frame writes (1 − z)·rd_selected (memory.md §2.4), registers start at 0 and a read returns the last write (memory.md §9), so every register read is a word, with no appeal to RAM.
  • Every RAM write of an execution family is a word: MEM_WORD writes rs2, a register value; MEM_SUBWORD bounds its merged word itself (§4.3); each ATOMICS arm is bounded (§6); a read-only query writes back what it read.

RAM's initial values are words — the image's, 0, the public input's — but advice's, which nothing bounds. No execution family relies on a RAM word being one: each bounds the value it uses, by MEM_WORD's rd pair, MEM_SUBWORD's splice or ATOMICS' comparison, so a row using a non-word is unprovable. The register half is what every carry needs: a + b − 2^32·wrap is a reduction only for words (memory.md §7), and addr_split's integer argument needs rs1 < 2^32.

6 ATOMICS#

One row is one read-modify-write: the ram query reads old and writes new at Δ = 3, beside rd (execution-trace.md §4), lr.w included, which writes its word back.

W[8..13]   decoded_next_pc decoded_rs1 decoded_rs2 decoded_rd decoded_mask    (no imm)
W[13..24]  kind_amoadd … kind_amomaxu
W[24..30]  word_index word_index_hi sum sum_hi add_wrap f_bitwise
W[30..42]  byte_a0..3 byte_b0..3 byte_and0..3       old's bytes, rs2's, their AND
W[42..50]  old_hi old_sign src_hi src_sign lt cmp_gap cmp_gap_hi lo
W[50..54]  the four multiplicities

With A = Σ_j 2^(8j)·byte_and_j inlined, its gates beside the plumbing are:

ram_value_rule    new − b_lr·old − (b_sc + b_amoswap)·rs2 − b_amoadd·sum − b_amoand·A
                    − b_amoor·(old + rs2 − A) − b_amoxor·(old + rs2 − 2A)
                    − (b_amomin + b_amominu)·lo − (b_amomax + b_amomaxu)·(old + rs2 − lo)
rd_value_rule     rd_selected − Σ_{k ≠ sc} b_k·old
add_rule          old + rs2 − sum − 2^32·add_wrap
f_bitwise_rule    f_bitwise − b_amoand − b_amoor − b_amoxor
old_bytes_rule    old − Σ_j 2^(8j)·byte_a_j          src_bytes_rule   rs2 − Σ_j 2^(8j)·byte_b_j
lo_rule           lo − rs2 − lt·(old − rs2)
addr_word (§2); add_wrap_boolean, f_bitwise_boolean; cmp_order, cmp_lt_boolean (below)

Each takes a kind's bit through its constants::extra_mask constant, from which the table's masks are built too, so a transposed arm would pass the decoder lookup. Under m_pc the family looks up the comparison's pairs on old, rs2 and cmp_gap and its two signs, §2's three and sum's pair; under f_bitwise, for each j, byte_a_j and 2^8·byte_a_j on RANGE16 and and_byte_j, (byte_a_j + AND_BASE, byte_b_j, byte_and_j), on GENERIC.

The comparison is constraints::gadgets::comparison (jump-branch-slt.md §3) with selector m_pc, lhs = old, rhs = rs2 and signed = [b_amomin, b_amomax]. The family's assemble asserts all four, nothing else in the artifact determining them: signed widened to amominu orders it signed, lhs and rhs swapped turn amomin into a max, and a selector narrowed to the min/max kinds drops old's bound on the other seven, and with it the bound on their rd write (§5). lo is the smaller under the ordering lt settles, and old + rs2 − lo the larger.

Why new is a word. old and rs2 are bounded by the comparison, sum by its own pair (add_rule is ungated: sum = (old + rs2) mod 2^32 on every live row), lo and the larger by being old and rs2. On a bitwise row byte_a_j's pair puts the key in the AND sub-table, whose row bounds byte_b_j and fixes byte_and_j = byte_a_j & byte_b_j; the byte rules are then the operands' decompositions, and A, old + rs2 − A, old + rs2 − 2A are AND, OR and XOR, carry-free byte by byte. Without its pair byte_a0 = 65,823 reads ShiftPowers' (65,824, 2^31, 1) (lookup.md §4); check_copowers takes the four keys under f_bitwise, which covers all three bitwise kinds: under b_amoand alone amoor and amoxor would read free byte_and.

sc.w always succeeds: it stores rs2 and writes 0 to rd, and the machine holds no reservation. The emulator does the same (execution-trace.md §10), and a row claiming failure, a nonzero rd or an unchanged word, is refused by rd_value_rule or ram_value_rule. This is a conformance deviation, not a soundness one: the proof is of what the program did on this machine. A guest may not rely on an sc.w failing where the ISA requires it to: with no valid reservation (no earlier lr.w, or one an earlier sc.w consumed) or at an address outside the reservation set. The lr.w/sc.w retry loop compiled code uses is unaffected, first-pass success being legal on any hart.

Auditors/Delegations

Delegation

Normative specificationdocs/spec/delegation.mdView as Markdown

The delegation ABI: how a guest hands a frame of RAM words to a circuit with an ecall, how each request pairs with exactly one invocation, how a program declares the families it calls, and how each family is sized. Frame layouts and circuits are delegation-circuits.md's, the recursion format's four families recursion.md's.

1 What a delegation family is#

A delegation family proves a function of guest memory too costly to run as instructions. It is invoked, never decoded: its number is a run-time value of a7, so it claims no pc and has no decoded table. A row is one invocation, which rides the cycle that requested it and owns no cycle (execution-trace.md §1); its accesses join the one memory multiset; it is in a VmConfig exactly when the image declares it (§7). Otherwise it is an ordinary family, an arm in constraints::family_circuit and a fill in prover::family_fill. A call is one row, the anchor's two leaves being a row's (§5); an operation wider than a row is several calls on one frame, chained through RAM (delegation-circuits.md §1, RAM glue).

2 The calling convention#

A call is an ecall (ecall-abi.md §1): a7 the number, a0 the frame base. It writes 0 to a0 and falls through (execution-trace.md §6); a recursion-format type writes a0 + 4·words instead (recursion.md §1.4).

An executor without a family's circuit answers -ENOSYS, on which a base-format shim's caller computes the same function in software, so an executor may implement any subset of the families; any other nonzero answer is fatal (ecall-abi.md §7).

3 The registry#

constants::delegation::TYPES, also program::DELEGATIONS, is one table of (family, number, anchor space, frame words), ascending by family, which the emulator dispatches on and constraints::add_sub builds its request gates from. The first BASE_TYPES = 6 rows are the base format's (recursion.md §1.2). Why each family has its height is §9's.

family id number anchor space frame words
KECCAK_F 9 0x0507 4 51
POSEIDON2 10 0x0500 5 24
FR_ARITH 11 0x0502 6 25
MOD_MUL 15 0x0504 7 25
SHA256_COMP 16 0x0508 8 25
EC_ADD 17 0x0506 9 97
FR_OP 19 0x0509 11 4
P2_FIELD 20 0x050A 12 5
FIELD_IO 21 0x050B 13 3
FQ_OP 22 0x050C 14 4

constraints::add_sub asserts at compile time that every number is in the precompile range and not EXIT, and that numbers and spaces are pairwise distinct, so an ecall row is the exit or a request of one type; a type costs that circuit a selector is_deleg_<f>, three gates and a term in five shared ones (add-sub.md §2, §4). A type's anchor space is the type: only its requests and invocations touch it, so the anchor's address is the frame base alone. A reserved range of RAM would need an argument that no guest access reaches it.

4 The frame#

A frame is words 32-bit words at the base a0 names, word j at base + 4j, read and written in place. Its base is word-aligned and it lies in RAM, RAM_ORIGIN ≤ base and base + 4·words ≤ 2^31: the executor refuses any other (Misaligned, OutOfBounds, the sum taken in u64) and the circuit has no witness for one (delegation-circuits.md §1, frame chain). So no frame lies in a public window or in advice.

An invocation reads and writes every word, unchanged ones written back, each a RAM query of the requesting cycle at slot constants::delegation::FRAME_DELTA = 0, ahead of the request's own queries (execution-trace.md §4, §7).

5 The anchor#

Requests and invocations pair one to one through the memory multiset, in the requested type's anchor space s. Otherwise N requests could close against one invocation, N − 1 calls going unexecuted, or an unrequested invocation could rewrite a frame.

5.1 The two sides#

                       reads                               writes
request (deleg)        T(s, a0, 0, 0)                      T(s, a0, 4c + 3, v)
invocation (anchor)    T(s, base, 4c + 3, anchor_value)    T(s, base, 0, 0)

The request is the deleg query of an ADD_SUB_LUI_AUIPC ecall row at cycle c (memory.md §2.1): deleg_mask_rule makes its mask m_pc·Σ_t is_deleg_t and deleg_addr_rule its address the a0 the row read. One query serves every type, so its space is deleg_space, an M column deleg_space_rule pins to Σ_t tag_t·is_deleg_t: a memory leaf may read no W column, and the selectors are W (memory.md §8).

The invocation's two leaves are the anchor read (delegation-circuits.md §1): it writes the answer, stamped 0 with value 0, and reads back the request's write at 4c + 3, c its cycle column. v and anchor_value are free and cancel only when equal; an honest prover writes 0 on both. Each answer starts a path one request long (memory.md §9).

5.2 The three request-side zeroings#

gate, under the request's mask forces
deleg_writes_no_register 0 written to a0, so the result is not the prover's choice
deleg_read_ts_zero the mirror read stamped 0
deleg_read_value_zero the mirror read's value 0

With deleg_addr_rule the last two make the mirror read the answer tuple, so every request consumes an answer of its own; without the timestamp, requests at one base chain, each consuming the previous one's write. The gates are the request row's, the same for every family, so the pairing needs nothing from a family's frame, and a call that changes no memory value has nothing else to expose it. The recursion format's deleg_a0_rule replaces the first (recursion.md §1.4).

5.3 Why the pairing is one to one#

In s the only tuples are the requests' and the invocations': no instruction reaches it, no window initializes it, nothing chains there (trace::AddressSpace::chains).

  1. A live row's 4c + 3 is not 0: the request's pc write and the invocation's frame writes at 4c lie on memory paths, whose timestamps are integers below 2^105 (memory.md §4.2).
  2. So the tuples stamped 0 are the requests' reads and the invocations' answers: as many invocations as requests, with the same multiset of bases.
  3. The rest are the requests' writes and the invocations' reads. No two requests share a cycle (memory.md §9), so each invocation's read is exactly one request's write: every invocation sits at its request's base and cycle, its frame accesses at that point of each word's history.

The trace-level check credits each anchor-space query with its invocation's tuples and sees none of this (execution-trace.md §9).

6 The executor's side#

For a registered number, Machine::ecall and Machine::delegate (crates/emulator/src/lib.rs) read a7 and a0; on the tracing paths refuse a family the VmConfig lacks (§7); read the frame, refusing §4's rules; compute the function natively (emulator::keccak_round, transcript::poseidon2_permute, Fr's operators, schoolbook products with long division, emulator::sha256_call) and write the whole frame back, a recursion family leaving it unchanged and working on field cells; stage the mirror query, reading and writing 0; and write a0 (constants::delegation::a0_after).

EmuError::DelegationFrame refuses a frame the circuit has no witness for, which the arithmetic would answer — long division is right for an unreduced operand too — leaving a proof that fails inside the GKR pass with nothing named: a KECCAK_F round word above 23, a SHA256_COMP group word above 15, an FR_ARITH code other than 1, 2, 3 or operand at or above p in memory form, a MOD_MUL or EC_ADD selector naming nothing or operand its row reads at or above the modulus, a POSEIDON2 lane at or above p. The recursion families' refusals are recursion.md §3–§6's.

The tracer records each invocation in its family's trace::DelegationTrace (execution-trace.md §11), which a shard reads as a trace::FrameSlice, ⌈invocations / height⌉ shards a family. The fill (prover::family_fill) commits the recorded words and derives the circuit's intermediates from those read. It never recomputes a written word: what is committed is what the execution did, and the circuit says that is the function. The circuit's side — frame chain, anchor read, gap decomposition, RAM glue — is delegation-circuits.md §1's.

7 Static detachment#

The instruction sweep cannot see a call, so each shim declares its family with a declaration record (constants::delegation):

MARKER_MAGIC = "APOGDEL1" (8 bytes) ‖ ecall number (u32 LE)          MARKER_BYTES = 12

guest_sdk emits one per family, a static whose #[link_section] is its own allocated section, .rodata.apogee.delegations.<family>, which link.ld's *(.rodata*) absorbs.

  • Its own section, because the linker's garbage collection keeps or drops whole input sections: records sharing one would be kept together, and reaching one shim would declare all.
  • Kept by reachability, not #[used], which keeps every record in every guest. Only the family's shim references its record, reading its own number from it through core::hint::black_box: a linked shim has a record, calls the number it declares, and the optimizer cannot fold the read away.
  • Statically: a call linked but never executed declares its family, which proves zero shards.

program::declared_delegations scans the image's file-backed bytes at every byte offset, a static's address being the linker's; a duplicate is one declaration, and a number no family answers is ProgramError::UnknownDelegation. Identity binds a record through the image column (program.md §8).

A called number whose family the VmConfig lacks is the fatal DelegationFamilyAbsent on the tracing paths; emulator::run, having no VmConfig, executes it. No proof covers it: the statement has no shard of that family, so the mirror read has no answer to consume.

8 Shards, time windows and the block#

A delegation shard's window is proof.md §8's, taken over its invocations' requesting cycles, so it lies inside the span of the ADD_SUB_LUI_AUIPC windows that made the requests. A delegation family is not cycle-owning, so the block holds its windows to nothing beyond start ≤ end ≤ 2^38; the anchor, not the window, places an invocation in time (§5.3).

9 Heights and channels#

A height sets how many calls a shard holds and limits no program. It is a parameter (ProgramParams::heights, defaulting to constants::family::DEFAULT_HEIGHTS, circuits.md §1) in identity's VM_CONFIG: a program's, not an execution's (program.md §7).

  • Floor: constraints::family_circuit returns None below the most variables any of the family's channel tables needs (constraints::lookup::table_vars, lookup.md §3).
  • Trade: a shard costs its height, not its occupancy (streaming.md §1), but its proof grows with the height only by a sumcheck round a variable in each gate list, a height changing no gate, only the number of halving lists. For a family with many calls the fatter shard is the smaller proof.
family channels floor unit of work calls a unit units a shard
KECCAK_F RANGE16, XOR8 2^16 keccak-f[1600] 24 10,922
POSEIDON2 none none width-3 permutation 1 256
FR_ARITH none none Fr add, multiply or inverse 1 256
MOD_MUL RANGE16 2^16 a·b mod m 1 65,536
SHA256_COMP RANGE16, XOR8 2^16 compression 16 16,384
EC_ADD RANGE16 2^16 complete point addition 3 21,845
  • POSEIDON2 and FR_ARITH take 2^8, the menu's smallest shard, where no table fits: every bound is a boolean decomposition. MOD_MUL and EC_ADD take their floor.
  • KECCAK_F and SHA256_COMP take 2^18, two variables above it: four times the calls for 2% more proof (a KECCAK_F shard's is 381,100 bytes, against 373,276 at 2^16). The price is memory: two 2^18 KECCAK_F shards in flight set the measured block's peak (streaming.md §1).
  • No base family carries TIMESTAMP, whose table needs 2^19 rows. FR_OP, P2_FIELD and FIELD_IO carry RANGE16, and FQ_OP TIMESTAMP and RANGE16, flooring it at 2^20.

10 Guest-side callers#

delegation reached from
KECCAK_F guest_sdk::keccak256; in guests/revm-block every alloy-primitives keccak, through its native-keccak hook native_keccak256
SHA256_COMP guest_sdk::sha256; revm-precompile's Crypto::sha256, the 0x02 precompile and the stateless guest's SSZ hashing
POSEIDON2 transcript::poseidon2_permute; guest_sdk::poseidon2_permute
FR_ARITH field::Fr's addition, Montgomery multiplication (*, square, pow, the conversions in from_u64, from_bytes, to_bytes) and nonzero inverse
MOD_MUL k256's FieldElement10x26::{mul, square}, Scalar::mul; ark-ff's MontBackend::{mul_assign, square_in_place} for BN254's two fields, as the product and then ·R⁻¹
EC_ADD guest_sdk::{ec_add, ec_mul}; k256's ProjectivePoint::{add, add_mixed, double}; revm-precompile's Crypto::{bn254_g1_add, bn254_g1_mul}
  • The shims are guest_sdk::recursion's but KECCAK_F's, which only keccak256 reaches (ecall-abi.md §7). Their frame types are #[repr(C, align(4))], so §4's alignment is the type's and not where the code generator put a local.
  • A multi-call operation's order is the caller's, and nothing refuses a wrong one: it computes something else. So each is one SDK function, keccak256's permutation, guest_sdk::recursion::sha256_comp and guest_sdk::recursion::ec_add_complete.
  • The transparent backends: field and transcript call the shims under cfg(target_arch = "riscv32"), through a target dependency on guest-sdk that a host build never resolves, not a cargo feature. Cargo refusing the cycle, guest-sdk cannot name Fr, so the shims take frames of bytes. The software path is each crate's own code, one branch below the call. A guest declares what its library calls reach: Fr arithmetic FR_ARITH, poseidon2_permute both.
  • FR_ARITH's frame carries Fr's memory form (primitives.md §1): canonical values would cost a Montgomery multiplication per value, more than the one the call replaces. POSEIDON2's carries canonical values, six conversions against the permutation's 240 multiplications.
  • The vendored crates, k256 0.13.4, ark-ff 0.6.0 and revm-precompile 43.0.2, are what a guest compiles through guests/Cargo.toml's [patch.crates-io], each route under the same cfg with upstream's code as its software path; the root workspace is unpatched. A MOD_MUL or EC_ADD operand must be below its modulus, so k256 first reduces its lazily reduced field elements. Changed files: guests/vendor/README.md.

11 Limits#

  • The EVM's MULMOD and MODEXP, BLS12-381 and every primitive outside §10's table run as instructions. No signature or pairing is delegated: secp256k1 recovery is k256 code over MOD_MUL and EC_ADD, a BN254 pairing ark-bn254 code over MOD_MUL.
  • A delegation is an operation's core: padding, a sponge or block loop, a scalar multiplication's ladder and a multi-call operation's order are guest code, proven as instructions.
  • A call's result is bound to memory alone: the frame after it is the function of the frame before.
  • This executor implements every family, so no proof here runs a base shim's software path.
  • Retired numbers are ecall-abi.md §4's.

Auditors/Delegations

The delegation circuits

Normative specificationdocs/spec/delegation-circuits.mdView as Markdown

The circuits of the six delegation families the base format registers (recursion.md §1.2): KECCAK_F, POSEIDON2, FR_ARITH, MOD_MUL, SHA256_COMP, EC_ADD. A row is one invocation of a function of a frame of guest memory. For each circuit: its frame, columns, gates and lookups, and why it admits that function and no other. The call, the anchor's pairing, declaration and heights are delegation.md's.

1 Shared constructions#

Each circuit is constraints::delegation's frame over words frame words beside the family's function. None has a setup column; its only tables are its channels' virtual ones (lookup.md §3). live is the one mask, boolean by live_boolean and every lookup's selector. A padding row is all zero and satisfies every gate, a constant term riding live (gkr.md §4).

M[0..4]        cycle  live  base  anchor_value
M[4 + 4j ..]   word j: addr_j  read_ts_j  read_j  write_j        w{j}_addr … w{j}_write_value

Frame chain. Each word is read and written once at a pinned address, as two RAM leaves over memory.md §1's tuple T, from M columns because a leaf reads no W (memory.md §8):

read_w{j}        live·T(RAM, addr_j, read_ts_j, read_j) + 1 − live
write_w{j}       live·T(RAM, addr_j, 4·cycle, write_j) + 1 − live
addr_w{j}        live·(addr_j − base − 4j) = 0
base_aligned     live·(base − RAM_ORIGIN − 4·base_low) = 0          base_low  < 2^29
base_in_window   live·(2^31 − 4·words − base − base_room) = 0       base_room < 2^31

The bounds are delegation.md §4's frame rules, alignment a decomposition because 4 is a unit of Fr. A word the call leaves alone is held by writes_back_w{j}, write_j = read_j; every other written word is bounded below 2^32 by its circuit. A frame lies in RAM proper (delegation.md §4), which starts as the image's words or 0 and which every writer — an execution family (memory-ops.md §5), a frame, FIELD_IO's export (recursion.md §5) — leaves holding words, so a frame word a circuit reads is a word without a bound of its own.

Anchor read. Two leaves in the family's address space s (delegation.md §5) pair the row with its request: it writes the answer T(s, base, 0, 0) and reads T(s, base, 4·cycle + 3, anchor_value), what the request wrote back; anchor_value is free. That makes words + 1 leaves a side, padded with literal 1s to a power of two.

Gap decomposition. Each read precedes the row's write: gap_j = 4·cycle − 1 − read_ts_j is in [0, 2^38). TIMESTAMP would need a 2^20 shard (lookup.md §3), so the frame bounds its gaps, base_low and base_room itself, at the head of W:

  • bit form, at 2^8, where no table fits: 38 booleans a word, gap{j}_{i}, under gap_w{j}, live·(gap_j − Σ_i 2^i·g_i) = 0, and 29 and 31 for base_low and base_room: 38·words + 60 columns, each with its booleanity gate.
  • chunk form, at 2^16 and above, with no gate: a bound x ∈ [0, 2^{16q+r}), 0 < r < 16, is q committed chunks c_k of weight 2^{16(k+1)}, a RANGE16 obligation on each and on the remainder x − Σ_k 2^{16(k+1)}·c_k, and one on 2^{16−r}·c_top, which bounds only beside the chunk's direct one (lookup.md §11). A gap (r = 6) is gap{j}_c0 and gap{j}_c1; base_low and base_room (r = 13, 15) take base_low_hi and base_room_hi: 2·words + 4 columns and 4·words + 6 obligations.

The frame's gates are live_boolean, the addr_w{j}, base_aligned and base_in_window, words + 3, and in the bit form the gap_w{j} and each bit's booleanity besides.

Canonicity chain. A value X in limbs x_0 … x_7 < 2^32 is compared with a modulus m, limbs m_i < 2^32, through boolean borrows β_i and differences d_i ∈ [0, 2^32):

<v>_canonical{i}    x_i − m_i − β_{i−1} + 2^32·β_i − d_i = 0        i = 0 … 7, β_{−1} = 0

Every term is a small integer, so the eight sum over ℤ to X − m + 2^256·β_7 = D, 0 ≤ D < 2^256: β_7 = 1 exactly when X < m. Against Fr's p (§3, §4) the m_i are literals, x_i − p_i rides live and each d_i is 32 booleans; against a selected modulus (§5, §7) the m_i are columns, 0 on a padding row, and each d_i has a 32-bit bound (memory.md §7).

Gated conclusion. The chain's last gate, <v>_below_modulus, is live − β_7 = 0 where every live row reads X; where only rows with enable = 1 read it, it is the gated conclusion enable·(1 − β_7) = 0. β_7 = enable would demand X ≥ m wherever enable = 0, so a row holding a reduced X it does not read would have no witness.

One-code rule. A frame word naming one of k cases is decoded into boolean selectors s_c by word − Σ_c code_c·s_c = 0 and Σ_c s_c − live = 0. The second is not implied: a code 0 has no selector set and a code that is a sum of two has two (1 + 2 = 3), mixing cases. With both, the word and any column pinned to Σ_c lit_c·s_c are one entry of a table of literals, selected and bounded by a degree-1 gate.

Byte operations. Where the unit is the byte (§2, §6), each Boolean operation is one XOR8 obligation (e_0, e_1, e_2), e_2 = e_0 ^ e_1 with all three bytes (lookup.md §3): e_1 and e_2 columns, e_0 any literal-weighted form with a constant (lookup.md §5). The rest is linear in the results: a & b = (a + b − (a ^ b))/2, ¬a & b = (b − a + (a ^ b))/2; against a literal k, v & k = (v + k − (v ^ k))/2 splits a byte at any bit, so a rotation or shift of a word held as bytes is a literal-weighted form over its bytes and their masked copies; and (0, c, c) bounds c to a byte. On true bytes and true XORs each form is exact over ℤ, so its value is the integer it denotes.

RAM glue. An operation too wide for a row is several invocations on one frame, a frame word naming the step (§2, §6, §7). Each proves its step on the frame as it finds it: its reads lie on each word's one history (memory.md §9), so it reads the previous step's writes unless the guest wrote there between. No gate joins two rows, and a shard boundary may fall between them. That every step runs, in order, is the calling code's, which the execution families prove.

2 KECCAK_F#

One invocation is one round of keccak-f[1600]; a permutation is 24 on one frame, the sponge and padding being guest code. The circuit, constraints::keccak, is flat, every gate in gate list 0, and its unit is the byte (§1): no column is a bit but live and the 24 round selectors.

2.1 Frame and columns#

51 words (constants::keccak; M[0..208]), the state in SHA-3 byte order: lane A[x][y], i = 5y + x, at words 1 + 2i (low half) and 2 + 2i. A[i][b] is its byte b; lane coordinates are mod 5.

word read written
0 the round r ∈ [0, 24) yes unchanged
1–50 the state yes the round's output
W name
0..106 the frame's chunks (§1)
106..130 round_sel{r} s_r, one a round
130..134 rc_b{b} rc_t, byte b_t = 0, 1, 3, 7 of the round's constant
134..334 state_in_l{i}_b{b} A
334..494 parity_x{x}_b{b}_s{s} column x's lanes XORed in four steps, the last C[x]
494..574 c_mask_…, theta_d_… C ^ 0x80; D
574..774 theta_a_… A′ = A ^ D
774..950 rho_mask_… A′ ^ mask on the 22 lanes not rotated by whole bytes
950..1150 rho_out_… B, after ρ and π
1150..1550 chi_and_…, chi_out_… B1 ^ B2; χ's output
1550..1554 iota_out_b{b} lane 0's bytes b_t after ι
1554..1556 the multiplicities

2.2 Gates and obligations#

385 gates; O is chi_out, but iota_out at lane 0's bytes b_t; r_xy = ROTATIONS[y][x].

gate count expression
the frame's (§1) 54
round{r}_boolean 24 s_r − s_r²
round_rule 1 read_0 − Σ_r r·s_r
one_round_a_live_row 1 Σ_r s_r − live
rc{t}_rule 4 rc_t − Σ_r s_r·(byte b_t of ROUND_CONSTANTS[r])
writes_back_w0 1 write_0 − read_0
input_w{j}, j = 1 + 2i + h 50 read_j − Σ_{k<4} 2^{8k}·A[i][4h + k]
output_w{j} 50 write_j − Σ_{k<4} 2^{8k}·O[i][4h + k]
rho_pi_l{i}_b{j} 200 B[y][2x + 3y][j] − rot_j(A′[x][y], r_xy), its constant times live

A rotation by 8q + s is linear in a lane's bytes v and their copies μ = v ^ mask (§1), mask = 256 − 2^{8−s} being the top s bits; with u = j − q and w = u − 1 mod 8,

rot_j(v) = 2^{s−1}·(v_u + μ_u) + 2^{s−9}·(v_w − μ_w) + mask·(2^{s−9} − 2^{s−1})      s > 0
rot_j(v) = v_u                                                                    s = 0

v_u's low bits moved up and v_w's top bits down, (v + mask − μ)/2 being v & mask.

The obligations are the frame's 210 on RANGE16 (§1) and 1,020 on XOR8, one a byte:

step count obligation e_2 = e_0 ^ e_1
θ 160 parity_s = parity_{s−1} ^ A[x][s + 1], s < 4, parity_{−1} = A[x][0]
θ 40 c_mask = 0x80 ^ C[x]
θ 40 D[x] = rot(C[x + 1], 1) ^ C[x − 1], c_mask as μ
θ 200 A′[x][y] = D[x] ^ A[x][y]
ρ 176 rho_mask = mask ^ A′
χ 200 chi_and = B1 ^ B2, Bk = B[x + k][y]
χ 200 chi_out = ((B2 − B1 + chi_and)/2) ^ B[x][y]
ι 4 iota_out_t = rc_t ^ chi_out[0][b_t]

2.3 Why it is sound#

Every byte column is an entry of some obligation, so all are bytes, each obligation is the operation it names and each form the integer it denotes (§1): rot because μ is the true XOR, and (B2 − B1 + chi_and)/2 is ¬B1 & B2. The channel alone fixes parity, c_mask, theta_d and chi_and. B is committed, and pinned by rho_pi, because χ reads every lane at an entry only a column may fill.

input_w and output_w are each a word's byte decomposition and its 32-bit bound, so no state word has a range obligation; without output_w a row could write any state. Both are ungated and degree 1, a padding row's words and bytes being 0, which pins its state bytes to 0; a cell that only live-gated gates and obligations reach is free on a padding row, to no effect.

one_round_a_live_row is the one-code rule (§1) over codes 0 … 23: without it a live row could set no selector, claiming round 0, or two spelling a third, and ι would add no constant or a wrong one. The constant is a table of literals the selectors pick (rc{t}_rule), with no lookup or commitment. So a live row writes round read_0 of the state it read.

A permutation is RAM glue (§1) over guest_sdk::keccak256's loop, which stores r = 0 … 23 in word 0 before each call. crates/checker/tests/keccak.rs holds every gate and obligation over 24 such rows to a round written apart in u64 and, through emulator::keccak_round, to tiny-keccak.

2.4 Cost and callers#

1,764 committed columns and, at 2^18, 5,490 inner ones in 29 gate lists, 11 row-wise and 18 halving, all the two memory trees' and the two fraction trees'. The 1,020 obligations and the table's fraction fill 1,021 of the XOR8 tree's 1,024 leaves (lookup.md §6); four more would double it, 4,100 more inner columns. So ι is four obligations: a round constant is zero outside bytes 0, 1, 3 and 7 (constants::keccak::IOTA_BYTES_ARE_THE_ONLY_ONES, checked at compile time).

A 2^18 shard (delegation.md §9) holds 10,922 permutations; its proof is 381,100 bytes (proof.md §9), 34.9 a permutation, and its forward pass 45.2 GB of inner layers (streaming.md §1), which is what sets a block's peak. Caller: guest_sdk::keccak256 (delegation.md §10).

3 POSEIDON2#

One invocation is one transcript::poseidon2_permute (transcript.md §1). The circuit, constraints::poseidon2, is at 2^8 with no lookup, bounding in bits (§1), and is the one delegation circuit that computes above gate list 0.

3.1 Frame and columns#

24 words (constants::poseidon2):

words read written
8l … 8l + 7 lane l, l < 3 yes the permuted lane

A lane is its value's canonical encoding (Fr::to_bytes), not §4's Montgomery form, so the circuit is the permutation itself; the caller's six conversions are small beside the 240 S-box multiplications a call replaces.

M[0..100] and W[0..972] are the frame (§1). W[972..4092] holds 520 booleans for each of six values, the lanes read (in0 … in2) then written (out0 … out2): 256 word bits, then the canonicity chain's (§1) 256 difference bits and 8 borrows.

3.2 Gates#

Gate list 0 holds 4,245: the frame's 51 (§1), a booleanity gate on each W column, and 17 a value, over its read or written words: eight <v>_word{k}, word_k − Σ_t 2^t·bit_{k,t}, and its canonicity chain against p (§1), eight <v>_canonical{i} and <v>_below_modulus, live − β_7.

The permutation is computed, not witnessed: three gate lists a round r, S-boxing every lane of a full round and lane 0 of a partial one, whose other lanes the first two lists copy:

list 3r         q_i = (x_i + c_{r,i})²       t_i = x_i + c_{r,i}
list 3r + 1     q2_i = q_i²                  t_i copied
list 3r + 2     x′ = M_r·v                   v_i = q2_i·t_i, or x_i on a copied lane

M_r is E or I and the constants are literals of the gates; round 0's x is E applied to in_l = Σ_k 2^{32k}·read_{8l+k}. A committed column is read by gate list 0 only (gkr.md §2), so live and out_l = Σ_k 2^{32k}·write_{8l+k} are carried up to gate list 192, which holds the last three gates,

out_lane{l}     live·(x_l − out_l) = 0          x the state after round 63

gated because a padding row computes the permutation of the zero state, which is not zero.

3.3 Why it is sound#

A layer's column is forced by the gate that writes it, so x is the permutation of (in_0, in_1, in_2) as field elements. The word gates make each in_l and out_l the integer its words spell, and the chains put it below p: a lane at or above p has no witness, and out_lane fixes all 24 written words, where without the chains on out a row could write x_l + p. The forward pass accepts Plonky3's permutation vectors (crates/checker/tests/poseidon2.rs).

3.4 Cost and callers#

4,192 committed columns and 2,020 inner ones in 201 gate lists, 193 row-wise and 8 halving: 736 the rounds' (15 a full round, 11 a partial one), 768 the four carried columns', the rest the memory trees'. A 2^8 shard holds 256 permutations; its proof is 664,780 bytes, 2,597 a permutation. Caller: transcript::poseidon2_permute on the guest target (delegation.md §10).

4 FR_ARITH#

One invocation is one Fr addition, multiplication or inversion. The circuit, constraints::fr_arith, is flat, at 2^8 with no lookup, bounding in bits (§1).

4.1 Frame and encoding#

25 words (constants::fr_arith):

words read written
0 the code: 1 add, 2 multiply, 3 inverse (OPS) yes unchanged
1–8, 9–16 a, b yes unchanged
17–24 out yes, unconstrained the result

A value is Fr's in-memory form, Fr::to_memory_bytes: the canonical encoding of the Montgomery representative x·R, R = 2^256 mod p. The circuit computes what Fr's own operators compute on representatives,

add         out = a + b
multiply    out = a·b·R⁻¹
inverse     out = R²·a⁻¹, and 0 at a = 0

because a frame of values would cost the guest a Montgomery conversion per value, more than the multiplication a call replaces. Fr::inverse answers None at 0 itself and makes no call.

4.2 Columns and gates#

M[0..104] and W[0..1010] are the frame (§1); W[1010..2570] 520 booleans for each of a, b (read) and out (written), as §3.1; W[2570..2573] the selectors f_add, f_mul, f_inv (selector1 … selector3); W[2573..2576] the field columns prod, inv and z (is_zero). The 2,701 gates: the frame's 53 (§1); 2,573 booleanity gates, on every bit and selector; §3.2's 17 per value; writes_back_w{j} for j < 17; and, a, b and out being the forms Σ_k 2^{32k}·word_k,

gate expression
opcode_rule read_0 − f_add − 2·f_mul − 3·f_inv
one_op_a_live_row f_add + f_mul + f_inv − live
prod_rule prod − a·b
inv_is_an_inverse a·inv + z − f_inv
is_zero_at_nonzero a·z
inverse_of_zero_is_zero z·inv
out_rule out − f_add·(a + b) − R⁻¹·f_mul·prod − R²·f_inv·inv

R⁻¹ and R² are literals derived from constants::FR_R.

4.3 Why it is sound#

As in §3.3, each value is the integer below p its words spell, so out_rule fixes the eight written words. prod is committed, under an ungated gate, because a selector times a·b is degree 3. On an inverse row a ≠ 0 forces z = 0 and inv = a⁻¹, and a = 0 forces z = 1 and inv = 0; without is_zero_at_nonzero, z = 1 and inv = 0 pass at any a, and without inverse_of_zero_is_zero, inv is free at a = 0. one_op_a_live_row is the one-code rule (§1): 1 + 2 = 3, so opcode_rule alone lets f_add and f_mul answer an inversion with a + b + a·b·R⁻¹.

4.4 Cost and callers#

2,680 committed columns and 142 inner ones, all the memory trees', in 14 gate lists, 6 row-wise and 8 halving. A 2^8 shard holds 256 operations; its proof is 266,292 bytes, 1,040 an operation. Caller: field's addition, Montgomery multiplication and inverse on the guest target (delegation.md §10).

5 MOD_MUL#

One invocation is one multiplication out = a·b mod m of 256-bit integers, m one of four fixed primes a frame word selects. The circuit is constraints::mod_mul.

5.1 The frame and the columns#

25 words (constants::mod_mul). A value is a plain residue, not a Montgomery one, in eight 32-bit limbs, least significant first.

words
0 the selector: 1 secp256k1's base field p, 2 its order n, 3 BN254's base field q, 4 its scalar field r (CODES, MODULI) read, written back
1–8, 9–16 a, b, each below the selected modulus read, written back
17–24 out written; the value read is ignored

Codes start at 1, so a zero word names no field. The EVM's MULMOD, whose modulus is arbitrary, is not this call and runs as guest code.

M[0..104], W[0..54]   the frame (§1)
W[54..58]     selector1 … selector4        s_c, one a code
W[58..66]     m_limb{k}                    m_k, the selected modulus
W[66..162]    <v>{k}_hi, <v>_diff{i}, <v>_diff{i}_hi, <v>_borrow{i}     for v = a, b, out
W[162..178]   q_limb{k}, q_limb{k}_hi      the quotient and its halfwords
W[178..220]   carry{k}, carry{k}_c0, carry{k}_c1      c_k + 2^36 for k < 14, and two chunks
W[220]        range16_multiplicity

5.2 Gates and lookups#

read_j and write_j are word j's two values (§1), a_i and b_i read limbs, out_i written ones, and c_k = carry{k} − 2^36·live. Each expression is held to 0:

gate count expression
the frame's (§1) 28
writes_back_w{j}, j < 17 17 write_j − read_j
selector{c}_boolean; selector_rule; one_modulus_a_live_row 6 s_c − s_c²; read_0 − Σ_c c·s_c; Σ_c s_c − live
m_limb{k}_rule 8 m_k − Σ_c s_c·MODULI[c][k]
<v>_borrow{i}_boolean, <v>_canonical{i}, <v>_below_modulus 51 v's canonicity chain (§1) against the m_k columns, concluding live − β_7
limb{k}, k < 15 15 Σ_{i+j=k} (a_i·b_j − q_i·m_j) − out_k + c_{k−1} − 2^32·c_k; out_k past limb 7, c_{−1} and c_14 are 0

274 RANGE16 obligations, all under live: the frame's 106 (§1); a pair — the 32-bit bound of memory.md §7, two obligations over a committed high halfword — on every limb of a, b, out and q and on every diff_i (112); and each carry{k} in [0, 2^37), by two chunks and four obligations as a gap (§1) (56).

5.3 Why it is sound#

The field. By the one-code rule (§1), m is the modulus word 0 names. Codes add (1 + 3 = 4), so without one_modulus_a_live_row selectors 1 and 3 answer a request for r modulo p + q; with it each m_k is one literal, which is all that keeps m's limbs, bound by no obligation, below 2^32.

The product. Every limb of a, b, out, q and m being below 2^32, a position's products sum below 2^67 a side and the carries lie in [−2^36, 2^36), so no term nears Fr's modulus: the fifteen limb{k} equations hold over ℤ and, weighted by 2^{32k}, sum to a·b = q·m + out, position 14 having no carry out.

The reduction is out_below_modulus: without it (q − 1, out + m) satisfies every other relation wherever out + m fits eight limbs.

The operand bounds make the relation total, not out right: with a, b < m, q = (a·b − out)/m < m, so every frame the circuit admits has an eight-limb quotient. A caller holding a lazily reduced value therefore owes a reduction below m, not below 2^256. The emulator's mod_mul_frame refuses the frames no proof could cover, a selector that is no code and an operand at or above m (EmuError::DelegationFrame).

5.4 Cost and callers#

Shape: circuits.md §1. A 2^16 shard (delegation.md §9) is 65,536 multiplications at 2.1 proof bytes each; its forward pass, 2,180 row-wise inner columns × 2^16 rows × 32 bytes, is 4.6 GB.

guest_sdk::recursion::mod_mul makes the call over a ModMulFrame. The vendored k256 reaches it from its field and scalar multiplies (codes 1, 2), the vendored ark-ff from BN254's Montgomery multiply (codes 3, 4): delegation.md §10.

6 SHA256_COMP#

One invocation is four rounds of SHA-256's compression function and four words of its message schedule; a compression is sixteen invocations on one frame, joined by RAM glue (§1). Padding, the block loop and the final addition of the chaining value are the caller's. The circuit is constraints::sha256.

6.1 The frame#

25 words (constants::sha256):

words read written
0 the round group r < 16 unchanged
1–8 the working variables a … h a … h four rounds on
9–24 the schedule window W_{4r} … W_{4r+15} moved down four words, W_{4r+16} … W_{4r+19} last

Call 0 reads the chaining value as a … h and the block, decoded big-endian, as the window. Over a row the state is two sequences: A_0 … A_{−3} are a … d as read, A_4 … A_1 are a … d as written, and E_j is the same over e … h, so each of the sixteen is a frame column. For k < 4 and m < 4, every sum mod 2^32:

T1          = E_{k−3} + Σ1(E_k) + Ch(E_k, E_{k−1}, E_{k−2}) + K_{4r+k} + W_{4r+k}
A_{k+1}     = T1 + Σ0(A_k) + Maj(A_k, A_{k−1}, A_{k−2})
E_{k+1}     = A_{k−3} + T1
W_{4r+16+m} = σ1(W_{4r+14+m}) + W_{4r+9+m} + σ0(W_{4r+1+m}) + W_{4r+m}

Call r + 4's rounds read the words call r derives, so the guest computes no schedule; calls 12–15 derive words no round reads.

6.2 Bytes and their obligations#

No column is a bit but live and the group selectors g_r. A word that enters a Boolean operation has four byte columns, and each such operation is one XOR8 obligation (x, y, x ^ y) a byte (lookup.md §3), of which position 0 alone may be a literal-weighted form (lookup.md §5).

  • A rotation is linear. With μ = v ^ (2^s − 1) committed, a byte v splits into lo = (v + 2^s − 1 − μ)/2 and hi = (v − lo)/2^s. Byte j of ROTR_{8t+s}(V) is hi(v_{j+t}) + 2^{8−s}·lo(v_{j+t+1}), indices mod 4, and for s < 8 the word ROTR_s(V) is (V − lo(v_0))/2^s + 2^{32−s}·lo(v_0).
  • The big sigmas nest, Σ0(a) = ROTR2(a ^ ROTR11(a ^ ROTR9(a))) and Σ1(e) = ROTR6(e ^ ROTR5(e ^ ROTR14(e))), so each XOR has one rotated operand and the outer rotation is a word's: 17 obligations a sigma.
  • The small sigmas end in a shift, σ0(x) = ROTR7(x ^ ROTR11(x)) ^ SHR3(x) and σ1(x) = ROTR17(x ^ ROTR2(x)) ^ SHR10(x), so their outer XOR has two derived operands: the shifted bytes are committed and pinned by gates. 16 and 15 obligations, SHR10's top byte being 0.
  • Ch and Maj are linear in XORs, Ch(e, f, g) = (f + g − (e ^ f) + (e ^ g))/2 and Maj(a, b, c) = (a + b + c − (a ^ b ^ c))/2: 8 obligations each.
  • A carry c is a byte by (0, c, c).

That is 52 obligations a round and 32 a schedule word, 336 on XOR8. RANGE16 carries 114: the frame's 106 (§1) and a pair (§5.2) on each written word without bytes, A_4, E_4, W_{4r+18} and W_{4r+19}.

M[0..104], W[0..54]   the frame (§1)
W[54..70]     group{r}                     g_r, one a group
W[70..118]    a{j}_b{b}, e{j}_b{b}         bytes of A_{−2} … A_3 and E_{−2} … E_3 (j = m2 … 3)
W[118..150]   w{i}_b{b}, n{m}_b{b}         bytes of window words 1–4, 14, 15, derived words 0, 1
W[150..358]   r{k}_…                       52 a round: the big sigmas' masks and XORs (34),
                                           e^f, e^g, a^b, c^a^b (16), two carries
W[358..514]   s{m}_…                       39 a schedule word: the small sigmas' masks, XORs
                                           and shifted bytes (38), a carry
W[514..518]   w{j}_written_hi              high halfwords of A_4, E_4, W_{4r+18}, W_{4r+19}
W[518..520]   range16_multiplicity, xor8_multiplicity

6.3 Gates#

All of degree 1 but the frame's and the booleans:

gate count expression
the frame's (§1) 28
group{r}_boolean; group_rule; one_group_a_live_row 18 g_r − g_r²; read_0 − Σ_r r·g_r; Σ_r g_r − live
writes_back_w0 1 write_0 − read_0
a{j}_decode, a{j}_encode, e{j}_…, w{i}_decode, n{m}_encode 20 a word − Σ_b 2^{8b}·byte_b, for every word with bytes
w{i}_shift, i < 12 12 write_{9+i} − read_{13+i}
r{k}_a, r{k}_e 8 §6.1's A_{k+1} and E_{k+1}, as word + 2^32·carry − sum
s{m}_sum 4 §6.1's W_{4r+16+m}, likewise
s{m}_shr3_b{b}, s{m}_shr10_b{b} 28 a committed shifted byte − its form

K_{4r+k} is the form Σ_r K_{4r+k}·g_r.

6.4 Why it is sound#

A sum's operands are words: those with bytes by their obligations, and d, h, W_{4r} and W_{4r+9} … W_{4r+12}, which only sums read, because the frame lies in [RAM_ORIGIN, 2^31) (§1), below advice, where every initial value and every write is a word (memory-ops.md §5; §1 for these circuits). Its carry being a byte, a sum gate holds over ℤ, and its left word, bounded by its bytes or its pair, is the sum mod 2^32. Without the carry's range any word satisfies the gate; without the pair on A_4, a carry of 0 writes the unreduced sum. Every word a row writes is therefore a word: a copy, one with bytes, or one of the four with a pair.

Group 0's code being 0, group_rule alone admits a live row with no selector or with g_0 beside another; one_group_a_live_row refuses those and two selectors spelling a third group, each a round under a wrong constant. Sixteen rows are one compression by RAM glue (§1) and by guest_sdk::recursion::sha256_comp, which stores r = 0 … 15 in word 0 before each call; the emulator's sha256_frame refuses a group word of 16 or more. crates/checker/tests/sha256.rs evaluates every gate and obligation over sixteen chained rows built from FIPS 180-4 in u32 arithmetic and holds their output to the standard's abc digest.

6.5 Cost and callers#

Shape: circuits.md §1. A 2^18 shard (delegation.md §9) holds 16,384 compressions at 11.6 proof bytes each; its forward pass, 2,694 row-wise inner columns × 2^18 × 32 bytes, is 22.6 GB.

guest_sdk::sha256 pads, walks the blocks, and for each runs sha256_comp's sixteen calls and adds the result to the chaining value. The vendored revm-precompile routes Crypto::sha256 to it: precompile 0x02, and the stateless guest's SSZ hashing (delegation.md §10).

7 EC_ADD#

One invocation is a third of one complete point addition P1 + P2 on secp256k1 or BN254 G1, in homogeneous projective coordinates (x = X/Z, y = Y/Z). An addition is three invocations on one frame in group order, joined by RAM glue (§1); scalar multiplication is guest code over it. The circuit is constraints::ec_add.

7.1 The formula#

Renes–Costello–Batina 2015, Algorithm 7, for y² = x³ + b, with b3 = 3b: 21 and 9 (constants::ec_add::CURVE_B3).

group 0   xx = X1·X2            yy = Y1·Y2            zz = Z1·Z2
group 1   m4 = (X1+Y1)(X2+Y2)   m5 = (Y1+Z1)(Y2+Z2)   m6 = (X1+Z1)(X2+Z2)
group 2   X3 = xy·ym − byz3·xz  Y3 = yp·ym + bxx9·xz  Z3 = yz·yp + xx3·xy

xy = m4 − xx − yy   yz = m5 − yy − zz   xz = m6 − xx − zz   ym = yy − b3·zz
yp = yy + b3·zz     byz3 = b3·yz        xx3 = 3·xx          bxx9 = 3·b3·xx

Both groups have prime order, so the formula is complete: a doubling, P + (−P), the identity (0 : 1 : 0) and any Z take no special case, in the guest or in a row, and nothing is inverted. The formula is the caller's: the vendored k256's ProjectivePoint addition is this algorithm on these coordinates, so the delegated and the software path return the same representative.

The twelve multiplications are nine reductions, each of X3, Y3, Z3 being two products under one quotient. A row holds three, not nine, because a shard's memory grows with its row's width and its height cannot fall below 2^16 (§7.5).

7.2 The frame and the columns#

97 words (constants::ec_add), a value as in §5.1:

words read by group written by group
0 the selector, one of CODES: 1–3 secp256k1's groups 0–2, 4–6 BN254 G1's all none
1–24 X1, Y1, Z1 0, 1 2, as X3, Y3, Z3
25–48 X2, Y2, Z2 0, 1 none
49–72 xx, yy, zz 2 0
73–96 m4, m5, m6 2 1

A row has three slots, each one reduction of one shape:

A·B + C·D + 1024·m² = q·m + out,     out < m

Group 0's (A, B) are (X1, X2), (Y1, Y2), (Z1, Z2) and group 1's the three pairs of sums, both with C = D = 0. Group 2's (A, B, C, D) are (xy, ym, byz3, −xz), (yp, ym, bxx9, xz) and (yz, yp, xx3, xy): a product's sign rides its operand.

M[0..392], W[0..198]   the frame (§1)
W[198..295]   word{j}_hi                   the high halfword of every word's read value
W[295..301]   selector{c}                  s_c, one a code
W[301..310]   m_limb{k}, b3                the curve's modulus and 3b
W[310..334]   bzz3_{k}, byz3_{k}, bxx9_{k}     b3·zz_k, b3·(m5_k − yy_k − zz_k), 3·b3·xx_k
W[334..622]   <v>_diff{i}, <v>_diff{i}_hi, <v>_borrow{i}     chains of the twelve values x1 … m6
W[622..1027]  slot{r}_…, out{r}_…, q{r}_…, carry{r}_…    135 a slot: four operands (32), out and
              its halfwords (16), a nine-limb q and its halfwords (18), 15 carries c_k + 2^46
              with two chunks each (45), out's chain (24)
W[1027]       range16_multiplicity

7.3 Gates and lookups#

G_g is the sum of the two selectors naming group g, and c_k = carry − 2^46·live:

gate count expression
the frame's (§1) 100
selector{c}_boolean, selector_rule, one_code_a_live_row 8 §5.2's, over six codes
m_limb{k}_rule, b3_rule 9 the column − Σ_c s_c·(its literal for code c's curve)
bzz3_{k}_rule, byz3_{k}_rule, bxx9_{k}_rule 24 the column − its product above
<v>_borrow{i}_boolean, <v>_canonical{i} 240 canonicity chains (§1) of the twelve values and the three outs, against m_k
<v>_below_modulus 15 e·(1 − β_7): e is G_0 + G_1 for x1 … z2, G_2 for xx … m6, live for an out
operand{r}_{o}_{k}_rule 96 an operand limb − Σ_g G_g·(group g's expression at that limb)
slot{r}_limb{k}, k < 16 48 Σ_{i+j=k} (A_i·B_j + C_i·D_j + 1024·m_i·m_j − q_i·m_j) − out_k + c_{k−1} − 2^32·c_k, q having nine limbs
writes_back_w{j} 97 write_j − read_j − G_g·(out_k − read_j), for the group g and slot limb k that write word j, if any

1,110 RANGE16 obligations under live: the frame's 394 (§1); a pair (§5.2) on every word's read value (194), every chain difference (240), every out limb (48) and every q limb (54); and each carry in [0, 2^47), four apiece (180).

7.4 Why it is sound#

The curve and the group are §5.3's argument over six codes: codes add (1 + 3 = 4, 2 + 4 = 6), and the one-code rule is also all that keeps m's limbs and b3 literals.

The operands. An operand limb is its group's expression: a combination, with coefficients of at most 3, of frame limbs below 2^32 and of their products with b3. Its pin is therefore its bound, below 2^38 in magnitude, and it carries no obligation. It is a committed column because the expression depends on the group, and a selector times a product of limbs would be degree 3; b3 enters through the three helper columns for the same reason.

The identity. As in §5.3: a position stays below 2^78, the carries in [−2^46, 2^46) (CARRY_OFFSET_BITS), the sixteen equations hold over ℤ and close because position 15 has no carry out, and out < m makes out the residue of A·B + C·D. The 1024·m² (OFFSET_MULTIPLE) keeps the left side non-negative, a quotient's limbs being unsigned: it is lowest in group 2's Y3, at −673·m² by its operands' ceilings 22m, 22m, 63m and 3m. One literal serves every slot, a group-dependent offset being degree 3, and q < 1697·m fits nine limbs. ec_add::artifact checks both constants against the ceilings when it builds the circuit.

Canonicity. out < m is the reduction. A read value below m is what the ceilings assume, and so what gives every admitted frame a quotient; the emulator's ec_add_frame refuses a frame whose group reads a value at or above m, or whose selector is no code. Each such conclusion is gated (§1) on the groups that read the value: every lane is below m on every row a guest builds (EcAddFrame::of zeroes the intermediates), so β_7 = e would leave no row a witness.

What the guest owns. Each third is proved; their order is the guest's. guest_sdk::recursion::ec_add_complete writes the three codes in turn, and groups out of order are not refused but compute another point from stale lanes. Nor is a point held to its curve: what is proved is the formula's arithmetic.

7.5 Cost and callers#

Every bound is a RANGE16 obligation, so the family sits at 2^16, the channel's floor (lookup.md §3), and no other height is practical: as bits the 97 gaps alone would be 3,686 columns, and at 2^18 the forward pass below would be 73 GB. Shape: circuits.md §1. A shard holds 21,845 additions at 19.9 proof bytes each; its forward pass, 8,708 row-wise inner columns × 2^16 × 32 bytes, is 18.3 GB.

guest_sdk::ec_add makes the three calls over an EcAddFrame, and guest_sdk::ec_mul is double-and-add over it. The vendored k256 routes ProjectivePoint's addition, mixed addition and doubling here, and the vendored revm-precompile routes Crypto::bn254_g1_add and Crypto::bn254_g1_mul, precompiles 0x06 and 0x07: delegation.md §10.

Auditors/Pipeline

The streaming prover

Normative specificationdocs/spec/streaming.mdView as Markdown

How one execution becomes a BlockProof: the guest runs twice, and a fixed number of workers commit, then prove, its shards as the executor fills them. The block is proof.md's, and its bytes do not depend on the schedule; this page fixes when each column exists, and so what a proof costs.

1 The prover, and what it costs#

prover::prove_block_streaming(setup, io, max_in_flight) proves every block: host::prove wraps it over the ProverSetup that host::setup builds from an ELF, bench prove drives it (tools.md §1), and recursion nodes are proved through it. Beside the block it returns a StreamingReport: each pass's wall clock and the executor's time inside it, the cycle and shard counts, and the most shards held at once.

Its memory follows the shards in flight, not the shard or cycle count: a partial buffer per family (§4), the last-access tables (§3), at most max_in_flight shards being worked and one filled shard's rows per family waiting (§5), and the output, 64 bytes a commitment and the ShardProofs. The executor's whole output, emulator::trace_run's buffers and event log at about 300 bytes a cycle, never exists; executing twice (§2) costs time instead.

A shard costs its height times its circuit's width (circuits.md §1), however few of its rows are live. gkr::forward holds every inner layer as field elements, 32·Σ_{k≥1} w_k·2^{n_k} bytes over layer k's width and variable count (42 GiB for a 2^18 KECCAK_F shard, 8.4 GiB for a 2^20 SHIFT_BITWISE one), and gkr::prove adds a copy of the layer it reduces and an eq table. The opening, after the forward pass is dropped, copies every committed column.

Measured on the base proof of recursion.md §10:

workload block 257,510 of glamsterdam-devnet-8, revm-block-stateless: 60 transactions, 101.5 Mgas, 198M cycles, 207 shards
machine 32 vCPUs, 247.7 GiB, --in-flight 12
pass 1 191 s; 25.7 vCPUs busy on average; one-thread fills 81% of its shard-seconds; sampled RSS at most 15.9 GiB
pass 2 2,290 s; on average 11.95 of 12 shards held and 30.4 vCPUs busy until the guest exits; then the exit's 460 s tail, whose longest stretches are the two KECCAK_F shards' one-thread fills, 200 s and 279 s
peak RSS 173.92 GiB: the two 2^18 KECCAK_F shards, together in the tail with nothing else in flight

Twelve shards in flight never reached that peak: a delegation family's height set it, and max_in_flight bounds only how many shards coincide.

2 The two passes#

pass 1                                       pass 2
  execute; for each shard as it fills:         execute again; for each shard as it fills:
    fill its M columns, commit them,             fill every column, prove the shard,
    keep the points, drop the columns            keep the proof, drop the rest
  at exit: the window families' shards         at exit: the window families' shards
  the statement, then G1–G11                   the statement's roots, the BlockProof

Each pass drives a fresh emulator::StreamingRun, which hands over a family's buffer as a ShardChunk the moment it reaches the family's height (§4); the workers of §5 take the chunks.

Pass 1 commits each shard's M columns, its family's fill with them moved out and no multiplicities, one commitment a column (in the recursion format one a stack of 2^σ, recursion.md §1.3). At exit it derives the window list, shard counts and boundary from the final state (§3), asserts the cut equals trace::plan_shards, commits the window families' shards, puts the commitments in statement order by (family, index) (proof.md §1) and runs G1–G11 (proof.md §2). A commitment reads no transcript, so only its place in the absorbed order matters, not when it was computed.

Pass 2 re-executes. The emulator is a pure function of (image, io), with no clock, randomness or threads, so it cuts the same shards; pass 2 asserts that its CycleProfile, Execution, window list and boundary are pass 1's. Each shard gets every committed column, multiplicities included, and prover::prove_shard_columns: shard transcript, GKR proof, opening (proof.md §4, §5). Proofs go to their statement positions, and prover::public_inputs copies each shard's two memory roots into the statement.

M is not recommitted: a shard's opening takes its M commitments from the statement, pass 1's, and its polynomials from pass 2's columns, so columns that differed would give an opening the verifier refuses.

3 What survives an execution#

A streaming run records no memory event. At exit StreamingRun::finish hands over each non-empty partial buffer as its family's last shard, and a StreamedExecution: the last-access tables (trace::MemoryState), the CycleProfile and the Execution (execution-trace.md §11). Beyond the shards' rows, the guest's inputs and its journal, everything the statement needs is a function of that final state: the boundary (trace::build_boundary_finals), ZERO_WINDOWS' list (trace::init_windows), the window families' teardowns (memory.md §3) and the field-window count. So a window family's shard exists only once the execution is over (§5).

A fill reads one shard through prover::ShardSource, its ShardRows a trace::RowSlice (a cycle-owning family), a trace::FrameSlice (a delegation family) or, for a window family alone, the final MemoryState. The streaming path builds it over a fresh chunk, ShardSource::archived over a slice of a TraceArchive (§6); nothing else differs. Memory columns come from a shard's rows alone (execution-trace.md §11), and checker::memory_columns_from_log rebuilds them from the event log, independently (circuits.md §3).

4 The shard plan#

A family's rows, in the order they are appended, are cut into shards of its VmConfig height h: shard i is rows [i·h, min((i + 1)·h, len)), the last padded to h with zero rows (memory.md §2). trace::plan_shards is ⌈rows/h⌉ per family over the CycleProfile, cycles for a cycle-owning family and invocations for a delegation family, so a family the execution never reached has no shard. A window family plans 0; its shards are windows (memory.md §3), counted by shard_counts in crates/prover/src/lib.rs: one INIT_TEARDOWN shard and one of each public window whatever the execution did, a ZERO_WINDOWS shard per entry of init_windows, one per advice window supplied (trace::advice_window_count), and field windows through the highest cell touched (MemoryState::field_windows). The counts are the statement's shard_counts (proof.md §1).

The flush. StreamingRun makes the cut as it runs. After a cycle is recorded, a buffer that has reached h rows is handed over as ShardChunk { family, index, rows }, index = rows/h − 1, and replaced by an empty one. A cycle appends at most one row to any buffer, the owning family's and, for a delegation request, one invocation to the delegation family's, so a buffer reaches h without passing it, a step fills at most two, and no chunk is split. At exit finish hands over the partial buffers. Chunks arrive in fill order, not statement order, and pass 1 asserts that each family's count is the plan's.

5 The pipeline#

pipeline (crates/prover/src/streaming.rs) runs both passes: max_in_flight workers under std::thread::scope and one std::sync::Mutex around a Source, which holds the executor, the filled shards no worker has claimed, and the counts. Under the lock a worker gives back its shard and claims the next: a waiting one, or else it steps the executor itself until a buffer fills (Source::claim, the only place the guest runs). Outside the lock it builds the shard's columns, works it and drops it. These are the prover's only threads and only lock; within a shard, parallelism is rayon over data.

held bound by
claimed shards, and all built from them max_in_flight one a worker; asserted in Source::claim
filled, unclaimed shards rows only, one per family the executor steps only for a claim with nothing waiting, a step fills at most two buffers and the exit one per family; asserted in Source::admit

The executor never runs ahead of demand, and there is no batch: a slow shard holds one worker. The workers are not rayon threads. A shard's MSMs, forward pass, sumcheck and opening run on rayon's global pool, so RAYON_NUM_THREADS sets the cores the shards share, and a worker blocked in that work cannot take a second shard as a rayon thread waiting in a nested join would. Fills run on the workers' own threads, one each, so up to RAYON_NUM_THREADS + max_in_flight threads are runnable. Fork-join cannot express this: below one shard per core, a batch waits for its slowest shard.

  • The block is independent of the schedule. A shard's proof is a function of the global state and its own columns, its transcript a fresh sponge seeded with the digest (proof.md §4); no proof depends on the thread count (gkr.md §5); proofs are placed by statement position. crates/prover/tests/streaming.rs compares the bytes at 1 and 8 in flight.
  • The failure returned is the earliest in fill order, at any worker count: claims follow fill order, a claimed shard is worked to its end, a failure stops later claims (Source::fail), and an executor failure ranks after every shard it filled.
  • No deadlock: the lock is never held while a shard is worked or taken twice by one worker, and nothing waits under it but the executor's step.
  • A panic stops the claims, through a drop guard (StopOnPanic) or, inside the executor, the poisoned lock; the shards in flight finish, and the panic is re-raised as itself.

The window families' shards follow the pipeline, built from the final state in rayon batches of at most max_in_flight, which are all that a ThreadPool::install around the call bounds.

The knob. max_in_flight, at least 1, is an argument because only the caller knows the machine; bench prove --in-flight defaults to 8. StreamingReport::peak_in_flight is the most shards claimed or batched at once in either pass.

6 The retained archived path#

emulator::trace_run keeps a whole execution, every buffer and the MemoryEventLog, and trace::TraceArchive::from_execution holds it (execution-trace.md §11). The per-shard component reads one through ShardSource::archived, with the same fills and shard proving: prover::statement_inputs (counts, windows, boundary, every shard's M columns), global_commit_phase, shard_columns, shard_memory_columns, prove_shard, prove_shard_columns and public_inputs. checker::TamperHarness is built on it (circuits.md §3): it writes changed cells into shards' columns, recommits changed M columns in a fresh global commit phase and re-proves, which needs an execution held still and read twice. Streaming has no such seam: pass 2 rebuilds, by re-executing, the columns pass 1 committed, so a cell changed in either pass would contradict the other.

prover::prove_block(setup, archive, plan), advance(setup, archive, until) and finish(archive) prove a block from an archive; nothing outside crates/prover/src/phases.rs calls them. prove_block refuses a plan that is not plan_shards of the archive's profile. advance fills the archive's four later phase sections in order, timing each, and decodes any it already holds, so an imported archive resumes; a stopped streaming run starts again. No column is stored: a phase rebuilds them from the archive. The sections, in proof.md §9's encodings, each refusing a byte too many or too few:

section content
PostCommit the statement's PublicInputs bytes, without roots; the global transcript after G11 as its 226-byte postcard snapshot (transcript.md §3); the four memory challenges; the digest
PostGkr per shard, in statement order: family u32, index u32, ts_start and ts_end u64, the witness commitments, the outputs, the GKR proof, the base claims' point, the shard transcript's snapshot after the GKR proof
PostOpening each shard's ShardProof bytes
Final the complete PublicInputs bytes, then the proofs

Auditors/Pipeline

Recursion

Normative specificationdocs/spec/recursion.mdView as Markdown

How one base block proof becomes one Groth16 proof a contract checks. The section numbers are the ones the code cites. Where this page and the code disagree, the code is right.

base proof ──► leaves ──────► internal nodes ──► root ───► decider ───► contract
N shards,      each a run     each 2–4           covers    the root in  folds the root's
base format    of base shards children           0..N      Groth16      points; two pairings
  • A node is this VM proving a verifier program. It verifies shards and folds every Mercury check they defer into one accumulator (A, B), the claim e(A, [1]_2) = e(B, [x]_2). Nothing pairs before the contract.
  • Base proving is untouched. No base key, statement or proof moved a byte: a leaf verifies base shards as they are.
  • Nodes are proved in a recursion format (§1) over a field memory (§2) with four coprocessor families on it (§3–§6), and they replay tapes (§7) rather than run verifier-core on RV32, which measured 3.0B cycles for block 257,510's 207 shards: fifteen times the block itself.

1 Two formats, one code path#

1.1 The rule#

A statement is in the recursion format exactly when its VmConfig holds FIELD_WINDOWS (VmConfig::is_recursion), which is exactly when its program declares a field family. No wire form says which format applies.

1.2 The delegation registry#

constants::delegation::TYPES is one append-only table, and its first BASE_TYPES = 6 rows are all the base format knows. constraints::family_circuit is the base registry; constraints::recursion_circuit differs from it in two ways only: its ADD_SUB knows every row and carries §1.4's rule, and the five families of §2–§6 exist. VmConfig::circuit picks the registry, for a key's load rule and the prover alike.

1.3 Stacked commitments#

Every commitment a shard opens is a point its parent folds (§8.3). So a recursion shard commits each of its two phases — its M columns, and its W columns with the multiplicities — as stacks of 2^σ columns. At height 2^n, with k_M and k_W columns (VmConfig::stack_vars):

σ = min(24 − n, the smallest even σ with 2^σ ≥ max(k_M, k_W))
  • Column i is slot i mod 2^σ of stack ⌊i / 2^σ⌋. A stack is the (n + σ)-variate multilinear whose evaluations [j·2^n, (j + 1)·2^n) are slot j's column, and its commitment is that polynomial's Mercury commitment. 24 is the ceremony's size.
  • The GKR pass leaves each column's value v at u. Then σ challenges r are drawn (STACK_CHALLENGE), a stack's value is Σ_j eq(r, j)·v_j, and a setup column is a stack of one, eq(r, 0)·v, its commitment unchanged. The shard's one batch opening is at u ‖ r, over the M stacks, the W stacks, then the setup columns.

σ = 0 is the base format exactly.

1.4 A recursion request leaves a0 past its frame#

A base delegation request writes 0 into a0. A request of a type past BASE_TYPES writes a0 + 4·words (constants::delegation::a0_after), which the recursion ADD_SUB's deleg_a0_rule holds it to. So frames laid back to back replay as back-to-back ecalls, one RISC-V row a call.

2 The field memory#

2.1 The space#

address_space::FIELD = 10: cells addressed by a u32, each a whole Fr. Its tuples (FIELD, cell, ts, value) join RAM's in the one memory multiset. No instruction reaches it. Only §3–§6's rows do, each access at its row's requesting cycle c and its own slot, 4c + Δ, with a read's usual gap check; a read-only access writes back what it read. A field access is not a MemoryEventLog event, a value not being a u32: trace::MemoryState keeps each cell's last (ts, value), and a recursion execution has no TraceArchive form. It streams.

2.2 FIELD_WINDOWS#

Family 18, 2^20 rows: ZERO_WINDOWS' circuit at a stride of one cell a row. Window w is cells [h·w, h·(w + 1)), initialized to 0. The windows are consecutive from cell 0 — shard i is window i — so a statement lists none, and a cell outside them has no tuple to balance a read against.

The four families on it are invoked, by the delegation ABI: an ecall whose a0 is a frame of words in RAM. A frame's words name cells.

§ family id ecall anchor space height frame a row is
3 FR_OP 19 0x0509 11 2^20 [op, d, a, b] one field operation
4 P2_FIELD 20 0x050A 12 2^18 [n, s, x, y, d] one transcript duplex step
5 FIELD_IO 21 0x050B 13 2^18 [op, cell, ptr] eight RAM words to a cell, or back
6 FQ_OP 22 0x050C 14 2^20 [op, d, a, b] one BN254 base-field operation

3 FR_OP — one field operation a row#

op op
1 MUL d ← a·b 6 EQ a = b, or the row has no witness
2 ADD d ← a + b 7 IMM d ← word b, as an integer
3 SUB d ← a − b 8 SHL d ← a·2^32 + word b
4 MAC d ← d + a·b 9 DIGIT d ← a's low byte, b ← (a − d)/2^8
5 INV d ← a⁻¹, and 0 at a = 0

a, b and d are accessed at slots of their own, so any two may name one cell. EQ is how a tape asserts. IMM and SHL are how it builds a constant with no field arithmetic of the guest's. A scalar's 32 DIGITs ending at 0 represent it mod p, which is all a scalar multiplication needs.

4 P2_FIELD — one duplex step a row#

With the state at cells s..s+3, the row absorbs n ∈ {0, 1, 2} of the cells x, y — the lanes are (n ≥ 1 ? x : s₀, n = 2 ? y : n = 1 ? 0 : s₁, s₂ + n) — and writes poseidon2_permute of them to d..d+3. A state is never overwritten, so a challenge is a cell of the triple that made it. The circuit is flat, every S-box's u² and u⁴ committed, so a parent verifies it as one gate list.

5 FIELD_IO — between RAM and a cell#

Over the eight RAM words w_k at ptr:

  • IMPORT (1): the cell takes Σ_k w_k·2^{32k} mod p. A non-canonical encoding is harmless.
  • EXPORT (2): the words take limbs below 2^32 congruent to the cell. Congruence, not canonicity: a guest that needs the canonical value compares the words with p itself.

Addressability is the multiset's. A word no window initializes cannot balance.

6 FQ_OP — one base-field operation a row#

An element of BN254's Fq is four consecutive cells of 64-bit limbs, congruent to its value mod q and not necessarily below it. Only this family writes one. The op word is a code, three flags and a digit cell (word >> 6): a flagged operand's element is its word plus 8·digit, a bucket chosen by a digit, which is what lets an MSM be a static tape (§8.3).

op
1 MUL d ← a·b
2 ADD d ← a + b
3 SUB d ← a − b
4 MULEQ asserts a·b ≡ d
5 FROM128 d ← a₀ + 2^128·a₁ from two cells below 2^128: a coordinate from its transcript limbs

One integer identity serves all five, a·y + z = q·K + d′, checked over 128-bit groups of limbs with a range-checked quotient and carries. Some of those ranges go through TIMESTAMP, which is why the family is at 2^20. b's and d's four cells share one read timestamp, so an element is only ever written whole; tape::run refuses a tape that reads one written apart, before a fill would.

7 Tapes#

A shard's checks have a fixed shape for its family and height. So the host compiles them, once, into a tape (verifier_core::tape): a straight-line list of coprocessor calls over absolute cells, in which nothing branches on a value. A tape reads three kinds of cell:

  • constants, which it builds itself from IMM and SHL, so they are bound with it;
  • slots, which its caller fills: the statement's digest and memory challenges, and the shard's index, window, roots and commitments;
  • inputs, the proof, IMPORTed from a blob laid out in the tape's Input order (tape::shard_blob), which the tape's own checks are what bind.

tape::shard_tape is verify_shard_local's steps 7–11 and Mercury's field side (pcs_verify), call for call: the shard transcript, the GKR backward pass, the lookup and root checks, the stack challenges and values, and the opening's twelve scalars. Every check is an EQ. It leaves three things to its caller (§8): the shard's time window against its neighbours', step 10c on the two public shards, and everything on the curve — to a tape a point is four transcript limbs, and the batch's cm* is a hint.

tape::schedule reorders a tape into runs of one family's calls, tape::encode is the form a guest replays — the cells its imports fill, then runs of frames — and tape::run is the native reading, over a Memory that models each access's timestamp as well as each cell's value.

8 Nodes and the tree#

8.1 Two programs, one procedure#

guests/recursion is two binaries. The leaf verifies shards from..to of one statement of the base program. The node verifies two to four whole statements of the two recursion programs — its children's proofs — reads each child's journal out of the output window step 10c binds, holds the children to one another, and folds their accumulators beside their shards'.

Both run verifier_core::node::node through a Driver. The host runs it natively (host::recursion), so it refuses whatever a guest would, first and by name, and it writes the advice the guest reads. The guest runs it by coprocessor calls. A binary's image — every shard tape, the fold's templates, the constants — is built by build.rs with verifier-core itself and sits in .rodata, so a program's identity binds every tape it replays.

  • The base program's identity is a constant of the leaf's image. The SRS digest and the generic table are constants of both images.
  • A node takes the two recursion programs' identities as claims and journals them, for the top to check once.
  • A program's setup commitments are advice, held to its identity by recomputing it.

The global transcript is a chain across the tree (verifier_core::chain). The node with shard 0 runs the prefix, G1–G7. Every node absorbs its own shards' memory commitments, G8, from the state its predecessor left. The node with the last shard runs the suffix, G9–G11, which settles the digest and the memory challenges every node took as claims. A node that holds a whole statement makes its memory argument, Π reads · R_b = Π writes · W_b.

A node holds its children to: exit status 0; one base statement — its shape, digest, challenges, io_digest, exit status and shard count; adjacent shards; chain states that meet; time windows in order across the seam; and, of a node child, the two identities it requires itself.

8.2 The journal#

47 cells, each a 32-byte word (node::journal):

cells
0 a digest of the base statement's shape: its shard counts and windows
1–7 its global digest, four memory challenges, io_digest, exit status
8–10 its shard count, and the shards this node covers, from..to
11–18 the chain's state at from and at to: three lanes and a pending input each
19–22 the covered shards' read and write root products; the boundary factors where to is the count
23–28 the first and last covered shard's family and time window
29–44 A and B, each x then y in four 64-bit limbs
45–46 the leaf program's and the node program's identities this node requires; 0 for a leaf

The root covers 0..count: every base shard verified, the transcript run end to end, the memory argument made. What is left is one pairing check and two identities.

8.3 Folding#

After each shard's tape the node's own transcript absorbs the shard transcript's final state (FOLD_STATE) and draws w and w′ (FOLD_WEIGHT), so a shard's weights follow everything they weight. Then, as scalars of points:

  • entry i of the shard's Mercury check gets w·e_i, on its side (pcs_verify::ENTRY_POINTS);
  • the batch check cm* = Σ ρ^i·cm_i is folded beside it: cm* gets w′ more, and each cm_i gets −w′·ρ^i;
  • [1]_1 and the setup commitments, which every shard of a family shares, accumulate one scalar each and enter once;
  • a child's A and B enter under a weight drawn after its whole journal (FOLD_CHILD).

Each side is one MSM on FQ_OP (verifier_core::fold): Pippenger with 8-bit digits over GLV halves, 16 windows of 256 buckets, every step a static template. A point is held to the curve and its scalar's split to the scalar, then added to one bucket a window through an indirect operand. Inversions are host witnesses held by a MULEQ, and buckets start at offsets so that no addition degenerates. A point costs about 400 FQ_OP calls.

8.4 The scheduler#

bench recurse <dir>/<stem> --out <out>, over a base proof archive (tools/bench/src/recurse.rs):

  • Keys. It writes base.key and programs.key into <out> and builds the two binaries with APOGEE_RECURSION_KEYS=<out>, where their build.rs reads them. The node is built twice: once with no image, for the two programs' keys, and once over them.
  • Plan, fixed before anything is proved (host::recursion::Tree::plan, <out>/tree.txt): leaves of at most --leaf (64) consecutive base shards, then levels of internal nodes over two to --fan-in (4) children, a lone leftover carried up. A program has sixteen families and each costs at least a shard, so a node is sixteen shards before any work and leaves are cut large.
  • A node is a process, bench recurse-node: it verifies its inputs natively, builds its advice only then, proves, verifies, holds the proved journal to the native one and writes <out>/<id>.block. At most --in-flight nodes run, with --in-flight × --shards-in-flight shards in flight across them: a node takes its share of what is spare when it starts, so a root alone has the whole budget. A proof already in <out> is kept, so a stopped run resumes; a run whose plan or programs differ is refused.
  • At the root it checks what a verifier owes beside the root's own proof — the journal covers 0..count and is the archive's statement, it requires the two programs' identities, and (A, B) discharges — and then runs §9.

9 The decider#

The root is still a GKR proof and some hundreds of points, and a contract can check neither. host::decider splits its verification in two.

The circuit is §8.1's node procedure over one child, the root, through a Driver that writes rank-1 constraints: an FR_OP is one constraint in the common case and none where it only copies, a duplex is 255, advice is a free wire. It verifies the root as a node would, and holds its journal to from = 0 and to = count. But it folds nothing: every MSM template is skipped, and each point's four limbs and its scalar are bound wires instead, after the two identities, the base statement's exit status, and its public input and output, a wire a byte, whose digest the circuit holds to the journal's io_digest.

crates/groth16 is Groth16 over this repository's BN254. A circuit streams its constraints into a sink, so no matrix is held. Three things are not the textbook's:

  • Bound wires are values the verifier holds, too many to be public inputs. The proof carries their commitment D = Σ w_j·[(β·A_j + α·B_j + C_j)/η]_1 under a fifth trapdoor η. A challenge c is SHA-256 of D and the verifier's values, and the circuit ends with acc ← (acc + wire)·c over the bound wires. The public inputs are c and that result, both of which the verifier computes from its own values, and the check is e(A, B) = e(α, β)·e(IC, γ)·e(C, δ)·e(D, η), IC being the public wires' points under 1, c and the result. D is fixed before c, so wires that differ from the values agree with them at c with probability len/r.
  • No blinding. A proof hides nothing and is a function of its witness.
  • A Lagrange basis. A and B are sums over the constraints, Σ_j (A·w)_j·[L_j(τ)], not over the wires. So the one element a key holds a wire is [(β·A_i + α·B_i + C_i)/x]_1, x being γ, η or δ — and a powers-of-tau ceremony already publishes [L_j(τ)].

The key is a ceremony's, in two phases:

  • Phase 1 is ppot_0080_24.ptau, the ceremony the tree's own commitments are under (srs::Phase1): the Lagrange basis at the circuit's domain in both groups, and the powers a quotient takes. Everything of the key that depends on τ is a combination of those points, and nothing derives τ.

  • Phase 2 is the circuit's own (groth16::phase2, bench ceremony), and makes α, β, γ, δ and η from 1 by contributions: each multiplies a trapdoor by a factor only its contributor knew, so a trapdoor is unknown while one contributor to it was honest.

    step
    init every trapdoor 1: a wire's [A_i(τ)]_1, [B_i(τ)]_1, [C_i(τ)]_1, and [τ^k·Z(τ)]_1. Deterministic from the circuit and the file
    round 1, contribute to α and β: [β·A_i]_1 and [α·B_i]_1, kept apart
    seal a wire's three terms summed
    round 2, contribute to γ, δ and η: the sum over the wire's trapdoor, and [τ^k·Z(τ)/δ]_1
    key the last state verified and, if every trapdoor has a contribution, written as the key

    The order of the rounds is the soundness. A prover may hold a wire's three terms only summed, over δ or η: apart, it could give A, B and C three witnesses. A contribution to α or β scales the terms apart, so those are finished before anything is divided.

    A state carries each contribution's record — its factor in G1 with a Schnorr proof of knowing it, bound to the records before it, and the trapdoor in G2 afterwards. Verifying a state checks that chain, then its elements against init's under those trapdoors, one pairing equation over a random combination: against the circuit and the file alone, with no earlier state. Every step lists the records by their factors' points, so a contributor finds its own under the state the key is made of. bench decide reads the key key wrote, and nothing else writes one.

  • setup_dev, bench decide --dev-key, derives all six trapdoors from a public seed. It is for development and tests: anyone forges under it.

The contract (contracts/ApogeeVerifier.sol) is verify(input, output, exitStatus, proof, points), a point being x, y, scalar, side [1]_2's points and then side [x]_2's. It rebuilds the bound values — a point's limbs are its coordinates' halves, or four sentinels at infinity, which is what the root's transcript absorbed — recomputes c and the result, checks the Groth16 pairing, folds each side with ecMul and ecAdd, which is also what holds a point to the curve, and checks e(A, [1]_2) = e(B, [x]_2). Its Groth16 key, the ceremony's two G2 points and the two identities are set at deployment.

bench decide <out> proves under the ceremony's key, checks the proof natively, deploys and calls the contract in revm, and writes decision.constructor and decision.calldata — under --dev-key, development.*.

What a deployment still owes. A key is as trustworthy as its ceremony: one honest contributor a round, which a ceremony run on one machine is not. The circuit depends on the root's shape — its program, its shard counts, the public values' lengths — so a key, and its ceremony, is per shape. And the contract pays about 9k gas a point, because the circuit folds none.

10 Running it#

bench prove --stateless <fixture> --out <dir>             the base proof
bench recurse <dir>/<stem> --out <out> --in-flight 4      the tree
bench ceremony <out> init                                 the decider's key: once a root shape,
bench ceremony <out> contribute                           each contributor in turn, to alpha and beta
bench ceremony <out> seal
bench ceremony <out> contribute                           and to gamma, delta and eta
bench ceremony <out> key
bench decide <out>                                        the Groth16 proof, and the contract

It needs assets/ptau/ppot_0080_24.ptau. Measured on block 257,510 — the tree on a 32-CPU, 247 GiB machine, the ceremony and the decider on an 18-core laptop:

base proof 207 shards, 14.5 MB, 2,481 s
tree 4 leaves of at most 64 base shards and a root: 116 shards
leaves, four at once 21, 24, 23 and 27 shards; 2,157 s; 92 GiB peak
root, four shards in flight 21 shards, 460 s, 1.03 MB
decider's circuit 7,896,686 constraints, a domain of 2^23
ceremony init 65 s; a contribution 50–56 s; key 70 s, 12.7 GB; the key 2.65 GB
decider the key read in 1 s, the proof 18.5 s, 6.1 GB
contract 358 points; 3,620,026 gas; 34,980 bytes of calldata

Auditors/Pipeline

Ethereum blocks

Normative specificationdocs/spec/ethereum.mdView as Markdown

guests/revm-block runs Ethereum blocks on revm inside the VM. This page specifies its two binaries: the mini-block binary's input, BlockWitness, and its output commitment; the stateless validator's input, result and rules; how each runs a block on revm; and how a block is recorded.

1 The guest crate#

One library, revm_block (src/lib.rs and its modules), and two binaries, each its own program identity, the same for every block because the block is advice (public-values.md §6):

binary advice journal exit status
revm-block (src/main.rs) a BlockWitness (§2) the output commitment (§3) 0; 61 not a canonical witness; 62 not executable
revm-block-stateless (src/stateless_main.rs) statelessInputBytes (§4) the 43-byte result always 0

The mini-block binary runs transactions, usually a block's first few, over a pre-state recorded from a node (§6); it proves an execution, not a block's validity (§3). Full blocks are proved with the stateless validator: the mini journal grows by a record a transaction and outgrows the public window (public-values.md §9). On the host the library is the oracle crates/emulator/tests/revm.rs holds the mini binary's journal to, over unpatched upstream crates.

Dependencies. revm is the workload being proven: it and what revm-precompile brings (arkworks, k256, p256, sha2, ripemd) are reachable from no prover, verifier or other guest. It is built without default features, so without blst, c-kzg or libsecp256k1. The pin is exact, =43.0.1, because an identity is a digest of the image (program.md §8), and the set is the reference stateless guest's (paradigmxyz/stateless's lock): twelve crates, held in both lockfiles by crates/host/tests/revm_lock.rs, its revm-handler 43.0.1 carrying the EIP-8037 system-call state-gas reservoir tests-zkevm@v21.0.1 expects.

Delegations. Both binaries declare KECCAK_F, SHA256_COMP, MOD_MUL and EC_ADD: keccak through alloy-primitives' native-keccak hook (revm_block::native_keccak256), SHA-256, secp256k1 and BN254 through vendored crates (delegation.md §10, guests/vendor/README.md). revm-precompile is patched in its Crypto trait's default bodies, not given a second implementation by install_crypto: two types behind crypto()'s OnceLock<Box<dyn Crypto>> stop LLVM devirtualizing its calls, keeping code it otherwise strips, 870,828 bytes of .text on the mini binary, past its tables' reach.

Code size. No ELF is committed: host::fixture::build_revm_guest builds either binary at --release, proved at host::fixture::revm_params — 2^20 for every family whose height is a choice (revm_block::TRACE_HEIGHT_RELEASE), each delegation family's default, bytecode_size_words = 2^21. A 2^20 table reaches 1.9375 MiB of .text (program.md §5); the stateless binary's is about 1.96 MB, 96.6% of it. The debug image needs 2^22 and is only ever run.

2 BlockWitness#

The mini binary's advice: postcard of revm_block::BlockWitness, a format of this repository's, written by host::recorder (§6). Fields in declaration order; a word is 32 big-endian bytes; a u8, an Option tag (0 or 1) and a fixed array are raw bytes; every other integer and every length is a varint.

BlockWitness          env BlockEnvWitness; accounts Vec<AccountWitness>, by address;
                      txs Vec<TxWitness>, in execution order
BlockEnvWitness       chain_id u64; spec_id u8 (revm's SpecId); number word; beneficiary [20];
                      timestamp word; gas_limit u64; basefee u64; difficulty word;
                      prevrandao Option<word>; excess_blob_gas Option<u64>;
                      blob_gasprice Option<u128>; slot_num u64;
                      block_hashes Vec<(u64, word)>, by number
AccountWitness        address [20]; nonce u64; balance word; code Vec<u8>;
                      slots Vec<(word, word)>, by key, zero values included
TxWitness             caller [20]; to Option<[20]>, None a creation; value word; data Vec<u8>;
                      gas_limit u64; gas_price u128, the max fee from type 2;
                      gas_priority_fee Option<u128>; nonce u64; chain_id Option<u64>;
                      access_list Vec<([20], Vec<word>)>; blob_hashes Vec<word>;
                      max_fee_per_blob_gas Option<u128>; authorizations Vec<AuthorizationWitness>
AuthorizationWitness  chain_id word; address [20]; nonce u64; authority Option<[20]>, recovered

2.1 One state, one encoding#

BlockWitness::decode refuses, with exit 61:

rule WitnessError
spec_id is a SpecId UnknownSpec
excess_blob_gas and blob_gasprice both present or both absent BlobPairing
accounts, each account's slots, block_hashes strictly ascending AccountsNotSorted, SlotsNotSorted, BlockHashesNotSorted
the bytes are exactly BlockWitness::encode's for the value Malformed

The last closes postcard's two second encodings: postcard::from_bytes ignores trailing bytes, and its varints accept non-minimal forms (81 00 reads as 1). A code hash is computed, not carried, and a transaction's type is derived from its fields (TxEnv::derive_tx_type).

2.2 Execution#

revm_block::WitnessDb answers revm from the witness and refuses every miss (DbError, exit 62): the witness is unbound advice, so a default would be a value the prover chose.

  • Absence is recorded: WitnessDb::basic answers None for an account recorded with nonce 0, balance 0 and no code. A zero slot is recorded like any other.
  • BLOCKHASH reads env.block_hashes. revm answers 0 without asking for any block but the 256 before the current one, and serves those from the database, not EIP-2935's contract: at most 256 entries.
  • Code is Bytecode::new_raw_checked's: bytes beginning 0xef01 that are not a 23-byte EIP-7702 delegation, which a few pre-EIP-3541 accounts hold, are DbError::MalformedCode, where Bytecode::new_raw would panic, an exit 101 that names nothing (ecall-abi.md §7).
  • The block gas limit is a running bound. revm checks each transaction against the block's limit and keeps no total; revm_block::run, the block executor, refuses transaction i unless gas_limit_i ≤ env.gas_limit − Σ_{j<i} gas_used_j, the Yellow Paper's intrinsic validity.
  • The blob gas price is recorded (§6). It derives from the excess through the fork's update fraction, 3,338,477 at Cancun, 5,007,716 at Prague, raised by each BPO fork (§4.2), and revm 43 knows only the first two. revm holds each type-3 transaction's max_fee_per_blob_gas to it.

Not in the witness: signatures, caller and each authority being the producer's recovery, unchecked; a parent header, and the header rules against it; a state root (§3); a slot number, which the recorder writes as 0, no JSON-RPC method serving EIP-7843's.

3 The output commitment#

The mini binary's journal, revm_block::run's return:

per transaction, in order   status u8 (0 halt, 1 revert, 2 success) ‖ gas_used u64 LE
                            ‖ output_len u32 LE ‖ output: the return data, empty on a halt
logs commitment       32    keccak256 of  count u32 LE ‖ per log, in emission order:
                            address 20 ‖ topic_count u8 ‖ topics, 32 each ‖ data_len u32 LE ‖ data
post-state summary    32    keccak256 of  count u32 LE ‖ per account, by address:
                            address 20 ‖ nonce u64 LE ‖ balance 32 BE ‖ code_hash 32
                            ‖ slot_count u32 LE ‖ per slot, by key: key 32 BE ‖ value 32 BE

The record count is the witness's. The summary covers the state revm's finalize returns: every account the block loaded, read-only and nonexistent ones included, with every slot it loaded.

What a proof states: some canonical BlockWitness makes revm_block::run return this journal. Nothing ties the witness to a chain; a reader holding one recomputes the journal natively. And the journal tells witnesses apart only as far as the execution reads them: a slot read and then overwritten unconditionally reaches nothing, while every loaded account's final balance and nonce are in the summary.

4 The stateless validator#

revm-block-stateless maps tests-zkevm@v21.0.1's statelessInputBytes to its statelessOutputBytes, byte for byte. The formats and rules are ethereum/execution-specs' verify_stateless_new_payload at the release's commit (host::zkevm::RELEASE_COMMIT); revm_block::stateless::run is the guest's whole computation, and §5 lists its rules.

4.1 Input and output#

input    schema_id u16 BE ‖ SSZ(StatelessInput)                          ssz::decode
           new_payload_request   the schema's fork's NewPayloadRequest
           witness               state: trie-node preimages; codes; headers: RLP, oldest
                                 first, the parent last, at most 256
           chain_id              u64
           public_keys           eth-act/ere-guests v0.17.1's layout only: 65 bytes a transaction
output   new_payload_request_root 32 ‖ successful_validation 1 ‖ chain_id u64 LE ‖ schema_id u16 LE

The layouts' fixed parts are 16 and 20 bytes, so no input is both; the second is the zkEVM benchmark's. The root is hash_tree_root under EIP-7916's and EIP-7495's progressive forms as of 2026-01-15 (ssz::request_root), whatever the layout. Decoding is as strict as the spec's: every offset against the bytes it bounds, every bounded list against its limit, nothing after the end.

The guest exits 0 on every input. One that does not decode, or names a schema §4.2 does not list, publishes the sentinel, 43 zero bytes (ssz::SENTINEL); any other publishes its request's root, its verdict, its chain id and its schema id. The empty input is the one a run cannot be given, a run without advice having no advice region.

4.2 Forks#

The schema id, fork_index << 8 | 0x01, names the fork; no activation schedule is compiled in (block::fork).

schema fork request revm SpecId blob target, max update fraction
0x1201 Osaka Electra/Fulu OSAKA 6, 9 5,007,716
0x1301 BPO1 Electra/Fulu OSAKA 10, 15 8,346,193
0x1401 BPO2 Electra/Fulu OSAKA 14, 21 11,684,671
0x1501 Amsterdam Gloas: a block access list, a slot number, EIP-8282's two request types AMSTERDAM 14, 21 11,684,671

4.3 What a result proves#

true says the request whose root is published is a valid block on chain chain_id under the fork schema_id names. The witness needs no binding: the root fixes the payload, and the witness is held to it by hashes — the parent header to the payload's parent_hash, each ancestor to its child's, the state trie to the parent's state_root and each node to its parent's reference, each code to its account's code hash. A node a read needs and the witness lacks is an error, never an absence (mpt::get). So a wrong witness cannot make an invalid payload valid; but false says only that this input did not validate, which a prover can arrange for any payload.

4.4 How a block runs on revm#

The pre-state is witness::WitnessDb behind revm's State: the state trie under the parent's root, each storage trie parsed on its first read, codes by hash, and BLOCKHASH numbering each ancestor by its position below the block. stateless::execute is the spec's apply_body: the EIP-4788 and EIP-2935 system calls; each transaction; the withdrawals; the requests, from deposit logs and the checked system calls of EIP-7002, EIP-7251 and, from Amsterdam, EIP-8282. The calls before the transactions are block access list index 0, each transaction has its own, and what follows them shares the last. A transaction must fit what is left, Amsterdam metering regular and state gas apart (EIP-8037):

before Amsterdam   tx.gas_limit ≤ gas_limit − Σ gas_used
Amsterdam          tx.gas_limit ≤ 2^32 − 1,  min(tx.gas_limit, 2^24) ≤ gas_limit − Σ regular,
                   tx.gas_limit ≤ gas_limit − Σ state;  the block uses max(Σ regular, Σ state)
both               2^17·blobs ≤ 2^17·max − Σ blob gas

Three rules make the result the spec's where following reth would not:

  • Code loads when revm asks (witness::WitnessDb::code_by_hash), never with its account: the witness carries only the code the spec's execution read, and a coinbase may be a contract nothing calls.
  • Every write precedes every deletion in the post-state replay (stateless::post_state_root), in each trie, as the spec's mpt_set_storage_slots orders them. A deletion that leaves a branch one child needs that child's node, on no changed key's path; the witness carries those the spec's order needs, and writing first needs a subset.
  • One commit per index (stateless::commit_index). revm 43's access-list builder records a value that differs from its commit's baseline, and revm re-bases a value at each call, so committing call by call records a slot one call toggles and the next restores. An index's calls are committed once, each baseline reset to the committed state.

Also the spec's: a checked system contract must have code, deposit events are parsed to the byte, withdrawals precede requests; the TxEnv is built field by field (build_fill would put a dummy authorization in an empty type-4 list); the blob price is a checked fake_exponential (block::blob_gas_price); 0xef01 code that is not a delegation runs as legacy. Declared lengths are added checked and trie parsing is depth-bounded, a panic publishing nothing.

4.5 Signatures#

Every sender and EIP-7702 authority is recovered in the guest (tx::recover_key) under EIP-2's rules, 0 < r < n, 0 < s ≤ n/2, a parity bit, as Q = r⁻¹(s·R − z·G) with k256's arithmetic, which the vendored k256 routes to MOD_MUL and EC_ADD. The verification upstream's recover_from_prehash ends with cannot fail once recovery succeeds and costs about as much again, so it is not done. An authorization that does not recover is skipped, as EIP-7702 says. A key in ere-guests' layout is checked, never used: one a transaction, 0x04 ‖ x ‖ y, naming the recovered sender.

4.6 Conformance#

All 67,251 pairs of tests-zkevm@v21.0.1 match natively (crates/host/tests/conformance.rs, by hand); CI holds the library to a committed subset of 34 — a case for each rule the release reaches, the smallest valid one, every undecodable one — in both layouts, and the binary runs the subset by hand. The release fills only Amsterdam: tools/stateless-ref holds the Electra/Fulu layout to eth-act/ere-guests v0.17.1 and crates/host/tests/canonical.rs the encodings and header rules to two mainnet blocks, but no Osaka-family input has an end-to-end oracle.

5 Where each rule is checked#

The mini binary's rules, then the validator's step by step. A validator refusal is a stateless::Invalid variant, which host::zkevm::verdict names and the guest publishes as false. Paths are revm_block's.

rule refusal code
mini: a canonical witness exit 61 BlockWitness::decode
mini: every read recorded, code well formed exit 62 WitnessDb
mini: each transaction fits the gas left and executes exit 62 run_against
the input decodes under a listed schema sentinel ssz::decode, block::fork
the ancestors decode and chain Ancestors stateless::ancestors
no empty transaction EmptyTransaction stateless::verify
the base fee fits a u64 Unrepresentable stateless::payload_header
the header the payload implies hashes to block_hash BlockHash stateless::{verify, payload_header}
each transaction decodes (EIP-2718, types 0–4) Transaction(i) tx::decode
keyed layout: a key a transaction PublicKeys stateless::verify
the versioned hashes are the request's VersionedHashes stateless::verify
EIP-7934's block size BlockSize block::block_rlp_len
the header against its parent, twelve rules Header(_) block::validate_header
the blob gas price fits a u128 Unrepresentable block::blob_gas_price
the parent's state root is in the witness Witness(_) witness::WitnessDb::new
chain id; signature; keyed layout: the key names the sender ChainId(i), Signature(i), PublicKeys stateless::execute, tx::sender
the transaction fits what is left Capacity(i) stateless::execute
revm executes it, every read in the witness Execution(i) stateless::execute
the system calls SystemCall stateless::{execute, commit_index}
the deposit events Deposits block::deposit_requests
gas used, receipts root, bloom, blob gas used, requests hash GasUsed, ReceiptsRoot, Bloom, BlobGasUsed, RequestsHash stateless::verify, block
Amsterdam: the access list's item count and hash AccessList stateless::verify, alloy_eip7928
the post-state root StateRoot, Witness(_) stateless::post_state_root, mpt

6 Recording a block#

host::recorder::record(rpc, block_number, range) makes a BlockWitness for a block's first n transactions or all of them (recorder::TxRange) by running them once against a node: recorder::WitnessRecorder is a revm::Database over the parent block's state that records each answer, and the transactions run through revm_block::run_against, the guest's own executor, so the record is what the guest will read. The result is put through BlockWitness::decode.

  • Reads. An account is eth_getProof with no keys, absent when nonce, balance, code hash and storage hash are all empty or both hashes are zero, Geth's answer; code eth_getCode, checked against the hash; a slot eth_getStorageAt; a header eth_getBlockByNumber; the blob gas price eth_feeHistory's baseFeePerBlobGas, a receipt's blobGasPrice existing only for type 3.
  • Choices. The hardfork is mainnet's by number (recorder::mainnet_spec): before the Merge is refused, after Osaka runs as Osaka. caller is the node's from; authorities are recovered on the host.
  • The client, host::rpc::Rpc, files each response under the SHA-256 of its canonical request in the fixture's rpc-cache/, so a second recording is byte-identical and offline; a miss without ETH_RPC_URL is an error. A request goes through curl, the endpoint and its key on the command line, retried on a transport failure, a 5xx or a 429.
  • On disk (host::fixture): <stem>.json, a Pin naming the block and the length and SHA-256 of <stem>-witness.bin and <stem>-journal.bin, native revm's journal, beside rpc-cache/. crates/host/tests/vectors/mini-block* is block 26,057,509's first two transactions, refreshed by kat-gen -- block (tools.md §7).

Nothing here produces a stateless input. eth_getProof returns the nodes on a key's path, and a deletion that collapses a branch needs its surviving sibling's node, which is on no changed key's path (mpt::MptError::BlindedCollapse is the validator's refusal without it), so the proofs of a block's keys are not a witness. Stateless inputs come from an external producer, a tests-zkevm release or the zkEVM benchmark's datasets; host::zkevm reads every JSON object carrying both statelessInputBytes and statelessOutputBytes, and bench prove --stateless proves one as it is (tools.md §1).

Reference

Glossary

Normative specificationdocs/glossary.mdView as Markdown

The project's own vocabulary, one line a term, each linked to the section that defines it. Terms the literature fixes (GKR, LogUp, KZG, RISC-V) are not listed.

term meaning defined in
accumulator, accumulator entry a deferred Mercury check as twelve (side, scalar, point) entries mercury.md §6
advice memory whose initial values the prover chose, bound by nothing public-values.md §6
anchor, anchor space tuples in a delegation type's own space pairing a request with one invocation delegation.md §5
archived path proving from a held TraceArchive; only the tamper suite (checker::TamperHarness) does streaming.md §6
artifact a circuit as data, CircuitArtifact; also an exported ProgramImage gkr.md §4, program.md §3
base claims each committed column's claimed value where the backward pass ends gkr.md §5
base format, recursion format recursion if a statement's VmConfig holds FIELD_WINDOWS, else base recursion.md §1
block BlockProof: config, statement and its shards' proofs proof.md §1
bound wire a decider value the verifier holds, committed instead of a public input recursion.md §9
boundary the registers' and pc's final timestamps and values; they have no rows memory.md §4
cached entry a sub-expression inlined into its list's gates, not a column gkr.md §3
canonical form an element as its value, 32 bytes little-endian, below the modulus primitives.md §1
challenge slot a gate coefficient's challenge: drawn, or derived by the verifier gkr.md §3, §5
channel one LogUp identity over a shard's lookups into one table lookup.md §1
copower x < p as x·2^32/p < 2^32, void without a direct bound lookup.md §11
cycle-owning the execution families 0–6, whose time windows are ordered proof.md §8
decider a Groth16 proof that the recursion root verifies, for the contract recursion.md §9
declaration record, static detachment 12 bytes a linked shim leaves in the image: how a delegation is declared delegation.md §7
decoded table an instruction family's setup columns: row i is pc 2i program.md §5
delegation a family proving a function of a RAM frame, invoked by ecall delegation.md §1
discharge spending an accumulator; the rule that each lookup is one leaf of its tree mercury.md §6, lookup.md §11
enforcing, producing a gate vanishing on every row; one writing the next layer gkr.md §1
extra mask, kind, kind bit family_extra_mask = 1 << kind, a kind being a mnemonic's index; b_k its bit program.md §6
family a circuit and the rows it proves: instructions (0–6), memory locations or invocations circuits.md §1
field memory address space FIELD: cells of one Fr, for the recursion families recursion.md §2
fold merging a node's deferred Mercury checks into one (A, B) recursion.md §8
frame an execution family's queries; a delegation's RAM words at a0 memory.md §2, delegation.md §4
gate list, row-wise, halving the gates from layer k to k + 1, keeping the height or halving it gkr.md §1
gated key, neutral tuple a lookup tuple under its selector; off, it reads a neutral table row lookup.md §4
generic table the committed table of ZeroEntry, AND, U16GetSign, ShiftPowers lookup.md §9
global transcript, global state digest G1–G11: the statement, M commitments, memory challenges; G11 seeds each shard proof.md §2
HALT_PC 1: the exit row's next_pc, the pc's final value memory.md §5
height a family's rows a shard: 2^8, 2^12, 2^16, 2^18, 2^20 or 2^22 program.md §7
identity, image column one Fr digest of the decoded tables, the image, the entry pc, VmConfig program.md §8
in flight shards worked at once, at most max_in_flight streaming.md §5
invocation, request a delegation's row doing one call; the ecall row asking for it delegation.md §1, §5
journal the public output: what the guest leaves in the output window public-values.md §1
laws Laws 1–4: locality, derived width, top layer, single source of truth gkr.md §4
layer layer 0 the committed columns, the top the outputs; L{k}[j] between gkr.md §1
leaf, node, root recursion programs: a leaf verifies base shards, a node 2–4 child proofs; the root, all recursion.md §8
live row, padding row m_pc = 1, or a zero row; in a decoded table, an instruction, or −1 throughout memory.md §2, program.md §5
M, W, S, V memory, witness and setup columns; virtual tables gkr.md §2
memory form an Fr's Montgomery limbs x·R; on the wire only in FR_ARITH's frame primitives.md §1
mini-block the revm-block binary: transactions over a recorded pre-state ethereum.md §1
multiplicity a channel's W column counting each table row's lookups lookup.md §7
padding contract padding.row makes row-local relations vanish and tree inputs 1 gkr.md §4
pairing side G2One or G2X: an entry's G2 argument, [1]_2 or [x]_2 mercury.md §6
pass 1, pass 2 executing to commit every shard's M columns; again to prove each streaming.md §2
phase 1, phase 2 the decider key's ceremonies: powers of tau, then the circuit's own recursion.md §9
public window windows 2 and 3 at 2^12: input at 0x8000, journal at 0xC000 public-values.md §2
query one read and one write at one address in one cycle execution-trace.md §3
RAM glue invocations chained through their frame's words in RAM delegation-circuits.md §1
reconciliation ∏ read roots · R_b = ∏ write roots · W_b, once a statement memory.md §4
registry family_circuit, recursion_circuit: each family's one circuit circuits.md §1
scratch scratch[i], a flat relation's intermediate, one per inner column gkr.md §2
shard h rows of one family, or one window, proved alone but for the memory argument streaming.md §4
slot Δ in a cycle's timestamps 4c + Δ; a ProgramImage halfword; a frame position execution-trace.md §1, program.md §2, memory.md §2
SRS digest a digest of the SrsVerifier and the generic table's commitments proof.md §3
stack 2^σ columns committed as one, in the recursion format recursion.md §1
statement PublicInputs: input, journal, exit status and the execution's record proof.md §1
statement shard, shard-set exactness a (family, index) below its count; a block proves each once, in order proof.md §1
tamper twin a forgery proved as an honest prover would, refused in its class circuits.md §3
tape straight-line coprocessor calls a node replays; checker tape's listing recursion.md §7, tools.md §4
time window a shard's claimed [ts_start, ts_end); it binds nothing proof.md §8
transcript form a G1 point as four 128-bit Fr limbs; infinity, four 2^128 transcript.md §4
tuple T(AS, ADDR, TS, VAL): a memory access as one field element memory.md §1
u1, u2 a Mercury opening point's halves, pairing with an index's low and high bits mercury.md §1
VmConfig a program's families, their heights, bytecode_size_words program.md §7
window h words from byte 4h·w, initialized and torn down by one shard memory.md §3
write-side induction an execution family writes only words, so operands need no bound memory-ops.md §5

Reference

Tools

Normative specificationdocs/tools.mdView as Markdown

The binaries around the prover and verifier, none on a proof path: bench measures and proves (§1), profiler counts a guest's cycles (§2), a debug-info build logs a proving run (§3), checker validates circuits and the global transcript (§4), artifact-dump exports a guest's ProgramImage (§5), verifier checks a proof from files (§6), kat-gen regenerates the committed fixtures (§7), and two generators outside the workspace are reference oracles (§8).

1 bench#

cargo run --release -p bench [-- <routine>...]   every routine, or those named; --list lists them
cargo run --release -p bench -- prove <stem> | --stateless <file> [--case <name>]
    [--in-flight <n>] [--out <dir>] [--json <path>] [--hourly-usd <price>] [--toy-srs]

The routines time one component each, over their own data: fr-arith, poly-bind, msm, mercury, mercury-batch, zerocheck-prove, zerocheck-verify, gkr-prove. msm, mercury and mercury-batch run over ceremony bases, assets/ptau/ppot_0080_24.ptau, and return without them.

prove proves a block through host::prove (streaming.md) and verifies it (host::verify).

  • <stem> names a recorded block under crates/host/tests/vectors: its pin <stem>.json, to which <stem>-witness.bin and <stem>-journal.bin are held, names the guest that proves it; mini-block is committed (ethereum.md §6).
  • --stateless <file> is one input to revm-block-stateless. A .json EEST fixture gives its statelessInputBytes as the advice, unchanged, and its statelessOutputBytes as the journal the proof must bind, checked by revm_block::stateless::run first and on the proof after; --case picks one input by part of its name. Any other file is the raw input.
  • The guest is built at --release (host::fixture::build_revm_guest), decoded at host::fixture::revm_params and keyed over 2^22 ceremony powers or, with --toy-srs, over τ = 0xc0ffee, cached as apogee-bench-toy-22.srs in the temporary directory: the same timings, another identity, which the report names.
  • --in-flight is max_in_flight, 8 by default. The verb asserts that the guest exits 0 and the block verifies; --out then writes the proof archive (proof.md §9) under the stem's or the input file's name.

The printed BenchReport (--json writes it too) holds the block, identity, SRS, cycles per gas, shards per family, proof and statement bytes, clocks, peak RSS, cost and hardware. commit and gkr are pass 1's and pass 2's wall clocks; execution, the executor's time, runs inside them and is left out of their total; opening and final are 0; unattributed is the rest of the proving wall clock; setup and verify are apart. Peak RSS is Linux's VmHWM, absent elsewhere, where /usr/bin/time -l gives it. --hourly-usd adds the cost, price · proving_ms / 3,600,000, and the cost per Mgas. Any failure exits 1, a wrong journal or a failed --out after the report prints; a usage error exits 2.

The verbs recurse, recurse-node, ceremony and decide are recursion.md §8.4–§10's.

2 The cycle profiler#

cargo run --release -p profiler -- elf <file> [--advice <f>] [--input <f>] [<common>]
cargo run --release -p profiler -- block <stem> [<common>]
cargo run --release -p profiler -- record <number|latest> [--txs <n>] [--cache <dir>] [<common>]
    <common>: [--top <n>] [--json <path>]

elf runs any guest over the given input and advice, at the smallest menu height its code fits; block runs the revm guest over a recorded fixture; record records a block from ETH_RPC_URL (latest is the finalized one; every transaction unless --txs; cached in target/profiler-cache) and runs revm-block over it, its gas the transactions' limits capped at the block's. A run prints a table, the --top (30) functions in it, and with --json writes a ProfileReport; any error exits 2. Its numbers are counts of executed cycles, the same on any machine.

2.1 One histogram over pc#

profiler::profile adds 1 to one u64 per halfword slot of the image for each executed cycle, reading each chunk's pc column off emulator::StreamingRun and dropping the chunk, so it holds the histogram and one partial buffer per family. Delegation rows add nothing: their requesting cycle is the ecall row's. A function's cycles are the sum over its [st_value, st_value + st_size) (loader::function_symbols), its own and not its callees'; its calls are the count at its first instruction, which runs once a call, so code entered only past its entry shows cycles and no calls. A mnemonic's cycles are the sum over its slots, a category's over its functions', and the unattributed ones are at slots no symbol covers.

2.2 Classification#

tools/profiler/src/categories.rs puts each function in one of 14 categories by RULES, ordered substring rules where the first match wins, then FALLBACK_RULES, the generic runtime paths, each matched against the demangled path and the raw symbol (categories::classify). The order is the meaning: revm_interpreter::instructions::system::keccak256 is hashing because its rule comes before revm_interpreter::'s. Legacy mangling is decoded whole, v0 to its identifiers.

A function's cycles include what was inlined into it: ruint's 256-bit operations count in the EVM opcode handlers, each a symbol of its own, revm dispatching through a table of function pointers. The unattributed share and the mnemonic mix, which no symbol table can misattribute, are the checks on attribution.

2.3 Pricing a candidate#

removable = max(0, cycles − calls·(4 + 2·frame_words))

cycles is the category's, calls the entry counts of the candidate's named entry symbols, and 4 + 2·frame_words (categories::shim_cycles) the shim a delegation leaves: the frame's stores, the ecall, the results' loads. CANDIDATES prices secp256k1, 256-bit arithmetic, BN254, SHA-256/RIPEMD-160 and the keccak sponge; one without entry symbols is charged no shim and flagged. It is a ceiling: it charges nothing for the new family's shards (delegation.md §9) or for marshalling operands into a frame.

3 The proving debug log#

crates/prover/src/debug.rs and the prover's log lines exist only with its debug-info feature, the workspace's one cargo feature: off by default, enabling no dependency, changing no proof byte (crates/prover/tests/debug_info.rs proves one statement with the log off and at deep and compares the blocks). Without it dlog! and debug_only! expand to nothing, so no scan is compiled into a proving run. gkr::explain_self_check is compiled always.

cargo run --release -p bench --features prover/debug-info -- prove ...
cargo test --release -p prover --features debug-info --test <suite> -- --include-ignored
APOGEE_DEBUG=off | phase | detail | deep [:FAMILY,FAMILY]

APOGEE_DEBUG, read at each log site, picks the level, case ignored: unset or empty is phase, none and 0 also mean off, 1 to 3 the other levels. :FAMILY,… (names as the log prints them, or ids) keeps those families at the level and lowers the others one step; lines naming no family stay. A bad level falls back to phase, an unknown family is dropped, and either is reported once as apogee ERROR. Lines go to the raw io::stderr() handle, one locked write each: libtest shows captured eprintln! output only for a failed test, and an OOM kill, a hang or a SIGINT loses it.

level adds
phase identity and SRS digest in full, in to_bytes order as the verifier CLI takes them; each claim's take and its committed or proved; the global digest and memory challenges, on the apogee commit line; each shard's begin h= … gkr done and open begin … open done
detail each family's circuit inventory; each shard's time window, g, β, roots and opening commitments; gkr::self_check; the scans
deep each GKR layer's shape and bytes; the top layer's all-zero columns

Where a run died. A begin without its done names the shard that died (FAMILY#index, [k/N] its statement position); a take without committed or proved, one in flight. fill# is fill order, which picks the failure returned, and in_flight= below the bound mid-pass means the workers wait on the executor. fill_ms is the one-thread fill, ms a wall clock shared with the shards in flight. Every shard forks from the apogee commit line's values, so two runs that should agree diverge there or inside a shard.

self_check recomputes every gate on every row before the backward pass, a second forward pass (gkr.md §5). gkr::explain_self_check turns a failure into the row's first disagreeing gate and every operand's value, a committed column by its artifact name and an inner one by the relation that wrote it, where a verifier says only LayerInconsistency { layer }.

The scans read each base delegation shard's live rows: invocations against the height, cycle and frame-base ranges, timestamp gaps, selector and round histograms, and canonicity, a tally for POSEIDON2 and FR_ARITH, whose < p conclusions are gated to the rows that read a value, and a verdict for MOD_MUL's operands and the values each EC_ADD row's group reads (debug::ec_add_reads). On ADD_SUB_LUI_AUIPC they count requests per type, which sum to each delegation family's invocations, and exit rows, one in all. The log's verdicts:

marker
self_check FAILED a gate fails on the prover's own values
NOT CANONICAL a frame value at or above its modulus where a gate needs it below
UNBALANCED an EC_ADD curve whose three groups' counts differ
OVER the a timestamp gap beyond 38 bits
NAMES NO MODULUS a MOD_MUL selector naming no modulus
DISAGREES a SHA256_COMP frame its rounds do not produce: the fill's refusal, in every build
NOT LOOPING 24 TIMES KECCAK_F round counts more than 1 apart
ABORTED a nonzero exit status: the block proves a failed execution
OUTPUT-LAYOUT-BREAK outputs other than 2 + 2·channels: reduce_shard and channel_cones index channel roots from opposite ends
ALL ZERO a top-layer column all zero: a root of 0
DECLARED BUT NEVER INVOKED a delegation shard with no live row
console
$ APOGEE_DEBUG=detail <a debug-info run> 2>&1 | tee run.log
$ grep -c 'begin h=' run.log; grep -c 'gkr done' run.log    # unequal: a shard died
$ grep 'begin h=' run.log | tail -1
$ grep -E 'FAIL|NOT CANONICAL|UNBALANCED|OVER the|NAMES NO|DISAGREES|NOT LOOPING|ABORTED' run.log
$ grep -E 'LAYOUT-BREAK|ALL ZERO' run.log

At detail the self-check doubles each shard's forward work and the scans cost O(live rows × frame words); deep reads no layer's cells but the top's.

4 checker#

cargo run -p checker -- laws <artifact>       Laws 1–4, then the lookup rules (check_laws)
cargo run -p checker -- padding <artifact>    the padding contract (check_padding)
cargo run -p checker -- dump <artifact>       the circuit, readably (checker::dump)
cargo run -p checker -- tape <verifying-key> <public-inputs>

An artifact is a CircuitArtifact file, decoded for encoding only so that a lawless one reaches the checks, such as crates/constraints/tests/vectors/*.bin. The validators are circuits.md §3's, independent of constraints; padding omits the product-tree clause; dump prints any decodable artifact.

tape loads a key (verifier::load_verifying_key) and a PublicInputs file, an archive's .vk and .public, refuses a statement the key does not describe (verifier_core::derive_global_phase), and runs checker::check_global_tape. That renders the global commit phase's event log a line a message, absorb <TAG> <n> (n scalars, or a bytes message's 31-byte chunks) or squeeze <TAG>, and holds it to expected_global_tape: G1–G11 (proof.md §2) written from the statement's shape, sharing nothing with verifier_core::global_commit but statement_shards. It prints the tape or the first line out of order, and checks the script, not the values, which the log does not carry. checker exits 0 when a check holds or a listing prints, 1 naming the failure, 2 on a usage error.

5 artifact-dump#

cargo run -p artifact-dump -- <guest.elf> [--out <dir>]
cargo run --release -p artifact-dump -- tables <guest.elf> [--ptau <file>]

The first writes <name>.img, the ELF's ProgramImage in its postcard wire form with nothing around it (program.md §3), and <name>.img.txt, a report rendered from the image read back off those bytes, which must equal the loaded one or nothing is written: segments, the listing (address, length, encoding, expanded word), .symtab names marked as outside the artifact, and the artifact's and the ELF's SHA-256, which pin bytes and are not the program identity. <name> is the ELF's stem; --out defaults to the working directory.

tables prints the VmConfig and each instruction's pc, next_pc, family, mnemonic and decoded fields at ProgramParams::defaults(); with --ptau it reads 2^22 powers, the largest default height, and prints the program identity (program.md §8).

6 The verifier CLI#

cargo run --release -p verifier -- <verifying-key> <identity-hex> <public-inputs> <proof>...
cargo run --release -p verifier -- block <verifying-key> <identity-hex> <public-inputs> <block>

The key is loaded by verifier::load_verifying_key (proof.md §7) and its identity must equal <identity-hex>, 64 lowercase hex digits of its canonical bytes from a channel the prover does not control: never the key, the proof or an archive's .identity. The first form verifies each ShardProof file with verify_shard and requires the files to be the statement's shards, each once, in any order; the second verifies a BlockProof with verify_block, as an archive's .vk, .public and .block (proof.md §9). It takes no SRS digest, using the key file's (srs.md §3). Exit 0 when all verifies, 1 naming the first file refused or a wrong shard set, 2 on usage or a malformed identity.

7 kat-gen and the committed fixtures#

cargo run -p kat-gen regenerates the default groups, cargo run -p kat-gen -- <group> one. Each file written prints its SHA-256, which the tests reading it pin.

group writes from
field, poly, curve, tower, pairing, msm, srs, moduli arithmetic, ceremony-point, KZG and MOD_MUL modulus vectors arkworks; srs's points through its own .ptau reader
pcs G1 absorption limbs; Mercury proofs arkworks; pcs
loader, isa listings of the committed guest ELFs, synthetic ELFs; an RV32IMA corpus, words that must not decode the pinned toolchain's llvm-objdump, llvm-nm
program, tape program identities, the generic table's commitments; guests/shards' global tape program; checker
gkr, memory, lookup, family, delegation CircuitArtifact files: toy circuits; the four frames, the two RAM-window circuits and the seven execution circuits, at 2^22; each base delegation circuit's shape and SHA-256 constraints
revm a synthetic block's witness, output commitment and delegated keccak-f frames native revm, held to the guest
opt-in: block, zkevm, guests mini-block, over ETH_RPC_URL (ethereum.md §6); zkevm-subset.json, cut from the tests-zkevm release at APOGEE_ZKEVM_FIXTURES only if every pair matches; the guest ELFs, each built twice and compared

srs, and program's identities and table commitments, need assets/ptau/ppot_0080_24.ptau and are skipped without it. CI runs the default groups and both oracles (§8) and fails on any git diff in the vector directories. A guest ELF is not reproducible across machines, since rustc embeds absolute paths in the panic-location strings of core and of crates outside the guest workspace, which the guest build does not remap; two clean builds on one machine agree. So guests is run by hand on one machine, and CI regenerates only what derives from the ELFs.

8 Reference oracles#

cargo run --manifest-path tools/transcript-ref/Cargo.toml
cargo run --manifest-path tools/stateless-ref/Cargo.toml

tools/transcript-ref implements transcript.md from its text over Plonky3's Poseidon2 and HorizenLabs zkhash's round constants, pinned by revision, and writes crates/transcript/tests/vectors/: permutation vectors, transcript scripts and io_digest cases. tools/stateless-ref encodes stateless inputs with eth-act/ere-guests v0.17.1's stateless-validator-common over libssz 0.3.0 and writes stateless_ref.txt under crates/host/tests/vectors/: per input, its request's hash_tree_root or reject. Each is its own workspace root because its dependencies enable features, serde/std among them, that cargo's feature unification would carry into the workspace's no_std crates; the one repository crate either links is tools/test-support, a seeded RNG, SHA-256 and hex with no dependencies.

Reference

Repository Map

Where everything lives in the Apogee VM repository, what each crate is, and the page of the specification that defines it.

View as Markdown

The Apogee VM repository is two Cargo workspaces: the root workspace for everything that runs on your host, and guests/ for everything that runs inside the VM.

Crates#

Path What it is Specified in
crates/constants every protocol constant, tag and identifier; no logic the page that uses each
crates/field, curve, poly, sumcheck Fr; the Fq tower, G1, G2, the pairing, MSM; multilinear polynomials; the zerocheck Primitives
crates/transcript Poseidon2 and the duplex transcript Transcript
crates/srs ceremony ingestion, the SRS archive, KZG, Groth16's phase 1 SRS
crates/pcs, pcs-verify Mercury and deferred verification; pcs-verify is verification's field side Mercury
crates/loader, isa, program ELF to ProgramImage; the decoder; decoded tables, VmConfig, program identity Program and identity
crates/emulator, trace the executor and its tracers; rows, memory state, column builders Execution trace
crates/constraints every circuit as data: memory frames, lookup channels, the family circuits, the registries GKR engine, Memory, Lookups, Circuits and the family pages
crates/gkr-verify, gkr the GKR verifier and prover GKR engine
crates/verifier-core statement, transcripts, verifying key, every check of a shard and a block but the opening; recursion's tapes, nodes and folding The proof, Recursion
crates/verifier verify_shard, verify_block, the proof archive, the verifier CLI The proof
crates/prover key construction, column fills, the streaming prover, the debug log Streaming prover
crates/groth16 Groth16 with bound wires and a two-phase ceremony Recursion §9
crates/host the host SDK: setup, prove, verify; the block-witness recorder; the recursion tree and decider Ethereum blocks, Recursion
crates/checker independent validators of the circuit laws, native lookup and memory evaluators, the tamper suite, the checker CLI Circuits §3
crates/guest-sdk the guest runtime: entry, linker script, allocator, memory regions, delegation shims Guest ABI, Delegation ABI
guests/ test and workload guests, a workspace of their own; vendor/ holds patched upstream crates Example guests
contracts/ ApogeeVerifier.sol Recursion §9
tools/ kat-gen, bench, profiler, artifact-dump, test-support; transcript-ref and stateless-ref, independent oracles outside the workspace Tools and CLIs
docs/ the architecture overview, the glossary, the guest manual, the tools page and spec/, one page per subject this site

Requirements#

  • The toolchain, its components and the riscv32imac-unknown-none-elf target are pinned in rust-toolchain.toml; rustup installs them on first use.
  • Program identity, real keys and proving need the ceremony file assets/ptau/ppot_0080_24.ptau. The workspace tests do not.
  • Proving is memory-bound: a full block peaked at 174 GiB.

Commands#

sh
# What CI runs
cargo fmt --all -- --check
cargo clippy --workspace --all-targets -- -D warnings
cargo test --workspace
cargo run -p kat-gen && git diff --exit-code     # committed fixtures regenerate identically

# Guests: their own workspace and target
(cd guests && cargo clippy --bins -- -D warnings)
(cd guests/fib && cargo build --target riscv32imac-unknown-none-elf)   # --release for proving

# Prove and verify a block, then recurse and decide
cargo run --release -p bench -- prove mini-block --out <dir>
cargo run --release -p verifier -- block <stem>.vk <identity-hex> <stem>.public <stem>.block
cargo run --release -p bench -- recurse <dir>/<stem> --out <out>

The suites that prove real shards are #[ignore]d and CI does not run them: each proves over a toy SRS of its own and needs tens of GiB.

sh
cargo test --release -p prover --test <suite> -- --include-ignored --test-threads=1
#   acceptance, control, alu, mem, fills, block, streaming, keccak, recursion, public_io, revm
cargo test --release -p host --test prove -- --include-ignored --test-threads=1     # a mainnet mini-block
cargo test --release -p checker --test tamper -- --include-ignored --test-threads=1 # every tamper twin

Reference

Release Notes

Apogee VM v1.0.0, the first release. What it proves, what it ships, how it was measured and checked, and its known limits.

View as Markdown

v1.0.0#

The first release of Apogee VM: a RISC-V zkVM that proves RV32IMAC programs and settles their proofs on Ethereum. The specification this site reproduces is the repository's docs/ at source revision 3571370.

What it proves#

That a program, named by a digest of its image, ran on a public input to an exit status and wrote a journal, carried through a recursion tree to one Groth16 proof that ApogeeVerifier.sol checks. Proofs are succinct, not zero-knowledge.

What ships#

  • The machine. RV32IMAC on one hart; the 59 instructions of RV32IMA with compressed instructions expanded at load; a guest SDK with three memory regions for input, advice and output.
  • The proof system. 23 circuit families over BN254's scalar field, each a layered GKR circuit: seven instruction families, five memory-window families, six delegations and five recursion families. One read/write memory multiset over the whole execution; LogUp lookups over five channels.
  • Delegations. KECCAK_F, SHA256_COMP, POSEIDON2, FR_ARITH, MOD_MUL and EC_ADD, reached from the SDK and from patched k256, ark-ff and revm-precompile.
  • Commitments. Mercury over KZG on the PSE perpetual powers of tau, one 704-byte opening per shard; deferred verification for recursion.
  • The prover. A two-pass streaming prover whose memory follows the shards in flight.
  • Settlement. A recursion tree of leaf and node programs in a recursion format with field memory and four coprocessors; a Groth16 decider with bound wires and a two-phase ceremony; ApogeeVerifier.sol.
  • The Ethereum workload. A revm guest with a mini-block binary and a stateless validator for Osaka, BPO1, BPO2 and Amsterdam.
  • Tools. bench, the cycle profiler, the prover's debug log, checker, artifact-dump, the verifier CLI, kat-gen, and two reference oracles.
  • No outside cryptography. Fields, curve, pairing, MSM, hash, PCS, GKR and Groth16 are implemented in the repository.

Measured#

Block 257,510 of glamsterdam-devnet-8 (60 transactions, 101.5 Mgas, 198M cycles): a base proof of 207 shards in 2,481 s on 32 vCPUs with a 174 GiB peak; a recursion tree of 116 shards; a decider proof in 18.5 s; on-chain verification for 3,620,026 gas. All 67,251 tests-zkevm v21.0.1 pairs match natively. Performance has every figure.

Known limits#

Not zero-knowledge; advice unbound by design; at most 16,380 bytes each of public input and journal; sc.w always succeeds; traps are not provable; a fixed set of six base delegations; prover memory set by the shards in flight; the decider key per root shape and only as trustworthy as its ceremony. The security model lists every limit with its reason.

Documentation#

This site, in English, French (Canada), Simplified Chinese and German, with the normative specification in English in every language. The AI Companion and llms.txt serve AI agents.

404

This page is out of orbit

Nothing lives at this address. Search the docs, or start again from the first step.