Apogee VM文档 v1.0.0
gweb3networks.com ↗
远地虚拟机主视觉:远地虚拟机徽标从一颗行星的地平线上升起,配有英文标语 Higher Compute Horizons

远地虚拟机·v1.0.0·迈出第一步

每一个应用,都可以是区块链原生的。

远地虚拟机证明一个程序被正确执行。你写的是普通的 Rust。远地虚拟机在一台 RISC-V 机器上运行它,证明它执行过的每一条指令,然后交给链一份证明,一次合约调用即可完成检验。对你的用户来说,正确性不再需要信任,而是可以验证。

证明陈述的内容已验证

程序
一个域元素,即 identity(程序身份):代码、初始内存映像、入口点与电路配置的摘要。
输入
程序收到的公开字节。
输出
公开输出(journal):程序选择发布的字节。
退出
程序结束时的状态。0 表示成功。

Groth16 · BN254一次合约调用

正确性证明

软件向来要求被信任。远地虚拟机让它转而接受检验。

只要程序运行在别人的机器上,用户就只能凭信任接受它的结果:账本、订单簿、赔付,乃至掷出的骰子。区块链让每个节点重新执行每一笔交易,从而为一类范围很窄的程序消除了这种信任。这行得通,却也是有史以来为达成任何共识而设计出的最昂贵的方式。

zkVM,即能证明自身执行过程的虚拟机,则为任何程序消除了这份信任。程序只运行一次,在哪里运行都可以。它留下一张数学收据,写明这个程序在这个输入下产生了这个输出;而检验这张收据,永远不意味着把程序再运行一遍。正是这张收据,让一个应用称得上区块链原生:它的规则写在代码里,它的状态以承诺的形式存放在链上,状态的每一次变化都随附自己的证明。

第三个时代

从工作量,到权益,再到正确性。

区块链的每个时代,都找到了一种不必再信任某一方的新方法。正确性证明是第一个深入计算本身的方法。

01

能源

工作量证明

电力保障事件的顺序。要改写历史,就得在电费上压过诚实的多数。

02

资本

权益证明

资本保障事件的顺序。不当行为的惩罚,是销毁为它担保的质押。

03

数学 · 当下

正确性证明

数学保障事件本身。每一次状态变化都附带一份证明,表明它是由所有人共同认可的那个程序计算出来的。

工作量和权益决定的是哪一段历史算数。两者都不检查历史之中发生了什么;这项工作向来落在每个节点身上,靠它们把一切重新执行一遍。有效性证明让这最后的蛮力退役。能源,然后是资本,然后是数学:再也没有第四样需要摆脱的信任。

改变了什么

你的产品,自有一条链,由数学来裁判。

区块链原生的未来,不是一条链包揽一切,而是众多环境:每个环境围绕一个应用塑造,全部结算到同一个基础层。远地虚拟机正是让运行这样一个环境变得切实可行的证明引擎。

编写

你的逻辑,用普通的 Rust 写成

客户程序(guest)是面向 RISC-V 的 no_std Rust 二进制程序:读取输入,完成工作,提交输出。远地虚拟机证明每一次运行,你永远不必用电路的方式思考。

结算

最终性,无需等待期

有效性证明一经验证即告最终。没有七天的争议期需要熬过,也没有委员会或安全飞地来代替数学。

专用

按你的产品塑造的环境

一个 Rollup 只承载一个应用,就能把全部资源押在对它唯一重要的那条通道上。远地虚拟机能证明为其机器构建的任何程序,因此状态转换函数由你来定义。

实测,而非承诺

一个完整的以太坊区块,从客户程序到合约。

远地虚拟机 v1.0.0 在 glamsterdam-devnet-8 的第 257,510 号区块上完成了端到端证明。该区块在虚拟机内按 execution-specs 的规则做了无状态校验,随后经递归折叠为一份以太坊合约接受的证明。

101.5 Mgas一个区块,60 笔交易在虚拟机内经由无状态校验器运行。
198MRISC-V 周期每条执行过的指令都是一行被证明的数据,分布在 207 个分片中。
1顶层证明由 116 个分片组成的递归树,折叠为一份 Groth16 证明。
3.62M gas链上验证对 ApogeeVerifier.sol 的一次调用,calldata 为 34,980 字节。
704 B每个分片的打开证明一份 Mercury 证明即可打开一个分片的全部已承诺列。
23电路族七个用于指令,五个用于内存窗口,六个是委托,五个用于递归。
67,251一致性用例tests-zkevm v21.0.1 的每一个测试对,校验器在原生运行下全部吻合。
0外部密码学依赖域、曲线、配对、MSM、哈希、PCS、GKR 和 Groth16 均在仓库内编写。外部库仅用作测试参照(oracle)。

基础证明在一台 32 vCPU 的机器上耗时 2,481 s,内存峰值 174 GiB;递归树另需约 2,620 s。目前证明的瓶颈在内存,性能页面给出了每个数据及其来源。

一份证明的路径

你来编写程序,此后的一切由远地虚拟机完成。

从你的 Rust 到合约返回的 true,中间是解码器、23 个电路族、GKR 引擎、承诺、递归树和 Groth16 判定器。这些都不需要你来构建或维护。

一份证明的路径 八个步骤:编写、加载、执行、分片、证明、递归、判定、验证。第一步由开发者完成,接下来的六步由远地虚拟机完成,最后一步由链完成。每个步骤下方标出第 257,510 号区块在该阶段已有产物的规模。 由你编写 由远地虚拟机证明 由链检验 编写no_std Rust 加载映像 · 程序身份 执行RV32IMAC · 1 hart 分片23 个电路族 证明GKR · Mercury 递归叶 → 根 判定Groth16 · BN254 验证ApogeeVerifier.sol 你的源代码32 字节程序身份198M 个周期207 个分片14.5 MB 证明1.03 MB 根证明34,980 B calldatatrue · 3.62M gas
一份证明的路径。两侧色带之间的一切都由远地虚拟机负责。各步骤下方的数字属于第 257,510 号区块,取自规范中的实测数据。

区块链原生

你熟悉的技术栈,每层只换一处。

走向区块链原生,并不意味着要学习一门新学科。传统应用的每个组件都有对应物,心智模型几乎原样延续。

层传统应用区块链原生,基于远地虚拟机
业务逻辑运行在你自己运维的服务器上的服务用 Rust 编写的客户程序,每次运行都被证明
数据库SQL 或键值存储数据在链下,状态根在链上
查询SELECT … WHERE key = ?包含证明,对照根进行检查
提交COMMIT新的根,连同其证明一起发布
审计记录要求别人相信的日志任何人都能检验的证明

逐层对照,附完整示例 →

从这里开始

六个入口。

同一个系统,从六个方向读。选一个与你带来的问题相符的方向。

远地虚拟机徽标:两片刃翼在一颗行星上方交汇于顶点,中心是一颗星

第一步,是一个程序。

用 Rust 写下它,在远地虚拟机上运行它。此后的一切,从分片、电路到递归与合约,都是机器的工作。抵达链上的是一份证明,而链所需要的,也只有一份证明。

更高远的计算地平线

迈出第一步

区块链原生

区块链原生应用与你今天构建的应用由同样的组件组成,只是每一层换掉一处。这里列出每一处替换,并用两种方式各实现一遍同一个账本。

以 Markdown 查看

区块链原生应用,是结算、托管与规则从构造上就在链上的经济活动,而不是一门在旁边挂上一个代币的传统生意。这听起来像是另一种工程,其实差别没有听上去那么大。

你今天所用技术栈的每个组件都有对应物。对应物做同样的工作,只改变一点:过去靠信任的,现在靠证明。远地虚拟机的存在,就是为了让这一改变便宜到足以成为默认选择。

一句话说清这一转变#

在传统应用中,服务器就是权威:它持有数据、执行规则、报告结果。在区块链原生应用中,链持有对数据的承诺,规则是一个任何人都能用其摘要指明的程序,而结果只有附带一份“确由该程序产生”的证明才会被接受。

运营方并不会消失。仍然有人运行程序、存储数据、响应请求。消失的,是必须相信他们这件事。

逐层对照#

层 传统应用 区块链原生,基于远地虚拟机 延续下来的
业务逻辑 部署在你自己运维的服务器上的服务 客户程序(guest):编译为 RISC-V 的 no_std Rust,每次运行都被证明 你写的仍然是作用于数据的函数。程序由它的程序身份来指明,即其代码与配置的摘要。
数据存储 SQL 表、键值存储 数据留在链下;链上存储一个状态根,即概括全部数据某一快照的单个哈希 一个能用 32 字节指明、可以拿来核对任何数据的快照。
读查询 SELECT balance FROM accounts WHERE id = ? 包含证明,即一条 Merkle 路径,由客户程序对照根进行检查 查询仍然返回一行。只是这一行现在附带证据,任何核对不通过的行都会被客户程序拒绝。
写入 UPDATE …; COMMIT; 一次状态转换:客户程序计算出新的根并将其发布 提交仍然意味着“使其持久化”,只是现在指的是合约把存储的根向前推进。
请求 HTTP 请求体 公开输入,由证明绑定 输入进,输出出。
大块数据 上传的文件、连接查询得到的行、拉取的文档 证明者提示(advice):由证明者提供的字节,客户程序将其与证明所绑定的某样东西核对 大数据按引用传递,并核对送达的内容。
响应 JSON 响应体 公开输出(journal),由证明绑定 任何人都能读取响应,并确知它出自该程序。
身份认证 会话、令牌、密码 在客户程序内验证签名;secp256k1 公钥恢复运行在委托的域运算与曲线运算之上 身份就是一把密钥,授权就是一段你能在源码里读到的检查。
密码学库 sha2、ring、OpenSSL guest_sdk::keccak256、sha256、ec_add,各自路由到专用电路 同样的调用,周期数只是原来的零头。
发布 推送二进制,行为立即改变 在验证者合约中登记新的程序身份 发布变得显式:新的构建就是新的程序身份,必须由合约接受。
扩展 更多服务器、数据库分片 一次执行被切成分片并行证明,再经递归折叠为一份证明 吞吐量来自并肩工作的证明者,而链上检验的仍然只是一份证明。
审计 要求别人相信的日志与鉴证 证明及其公开输出 保障从依赖声誉转向依赖验证。

中间一列就是区块链原生应用的构成。其下的机制由远地虚拟机提供:RISC-V 机器、电路、承诺、递归和验证者合约。这些都不会出现在你的程序里。

不变的部分#

  • 你写的仍然是普通的 Rust。 结构体、枚举、trait、迭代器、Vec、BTreeMap,以及任何不依赖 std 就能构建的 crate。没有电路语言要学。
  • 你仍然在自己的笔记本电脑上测试。 常见的布局是把应用逻辑放进一个 no_std 库:它在宿主机(host)上通过 cargo test 运行,行为与在客户程序中运行时完全一致。参见编写客户程序。
  • 你思考的仍然是状态、请求与响应。 形态不变,变的只是它们的保证。
  • 确定性的代码依然是确定性的。 好的后端代码本来就避免隐藏的输入,虚拟机把这一点变成绝对的要求。

改变的部分#

  • 没有外部世界。 客户程序没有时钟、没有随机数、没有网络,也没有文件。它知道的一切都以公开输入或证明者提示的形式送达,而向宿主程序请求数据,是一种任何证明都不会接纳的调用。
  • 每条指令都有代价。 每条执行过的指令都会成为一行被证明的数据。复制、内存分配和空转循环都要消耗证明时间,于是精打细算周期数的老规矩又回来了。
  • 外部提供的数据要检查,而不是信任。 证明者提示由证明者选择。在发布任何由它推导出的内容之前,客户程序要先把它与证明所绑定的某样东西进行核对。
  • 输出小而公开。 公开输出最多容纳 16,380 字节。较大的结果以摘要形式发布。
  • 没有任何东西是隐藏的。 远地虚拟机 v1.0.0 的证明是简洁的,但不是零知识的。客户程序不得持有秘密。

一个账本,两种做法#

向一个账户余额存入一笔款项:值得证明的最小状态变化。

传统版本#

sql
BEGIN;
SELECT balance FROM accounts WHERE id = $1 FOR UPDATE;    -- read
UPDATE accounts SET balance = balance + $2 WHERE id = $1;  -- write
COMMIT;                                                     -- make it durable

用户信任运营方:相信它对真实的表执行的正是这段代码,并且如实报告了结果。

区块链原生版本#

账户存放在一棵二叉 Merkle 树中,叶子为 keccak256(account ‖ balance)。合约存储树根。客户程序以公开输入的形式接收旧根和请求,以证明者提示的形式接收账户余额及其 Merkle 路径,检查路径,然后发布旧根和新根。

guests/ledger/src/main.rsrust
#![no_std]
#![no_main]

guest_sdk::entry!(main);

const DEPTH: usize = 20; // room for 2^20 accounts

/// A leaf commits to one account's balance.
fn leaf(account: &[u8; 20], balance: u64) -> [u8; 32] {
    let mut bytes = [0u8; 28];
    bytes[..20].copy_from_slice(account);
    bytes[20..].copy_from_slice(&balance.to_le_bytes());
    guest_sdk::keccak256(&bytes)
}

/// Fold a leaf up its Merkle path; bit `level` of `index` says whether the
/// node is a right child at that level.
fn root_of(mut node: [u8; 32], index: u32, path: &[[u8; 32]; DEPTH]) -> [u8; 32] {
    let mut pair = [0u8; 64];
    for (level, sibling) in path.iter().enumerate() {
        let (left, right) = if (index >> level) & 1 == 0 { (&node, sibling) } else { (sibling, &node) };
        pair[..32].copy_from_slice(left);
        pair[32..].copy_from_slice(right);
        node = guest_sdk::keccak256(&pair);
    }
    node
}

/// Public input: old_root (32) ‖ account (20) ‖ amount (8, LE)
/// Advice:       balance (8, LE) ‖ index (4, LE) ‖ path (DEPTH × 32)
/// Journal:      old_root ‖ new_root ‖ account ‖ amount
fn main() {
    let input = guest_sdk::public_input();
    let advice = guest_sdk::advice();
    if input.len() != 60 || advice.len() != 12 + 32 * DEPTH {
        guest_sdk::exit(1);
    }
    let old_root: [u8; 32] = input[..32].try_into().unwrap();
    let account: [u8; 20] = input[32..52].try_into().unwrap();
    let amount = u64::from_le_bytes(input[52..60].try_into().unwrap());

    let balance = u64::from_le_bytes(advice[..8].try_into().unwrap());
    let index = u32::from_le_bytes(advice[8..12].try_into().unwrap());
    let mut path = [[0u8; 32]; DEPTH];
    for (i, sibling) in path.iter_mut().enumerate() {
        sibling.copy_from_slice(&advice[12 + 32 * i..12 + 32 * (i + 1)]);
    }

    // The query: the balance the prover supplied is the one the root commits to.
    if root_of(leaf(&account, balance), index, &path) != old_root {
        guest_sdk::exit(2);
    }
    // The write: the same path with the new leaf gives the new root.
    let Some(new_balance) = balance.checked_add(amount) else { guest_sdk::exit(3) };
    let new_root = root_of(leaf(&account, new_balance), index, &path);

    // The commit: publish the transition for the contract to apply.
    guest_sdk::commit(&old_root);
    guest_sdk::commit(&new_root);
    guest_sdk::commit(&account);
    guest_sdk::commit(&amount.to_le_bytes());
}

对照 SQL 来读:SELECT … FOR UPDATE 变成了对照根检查的 Merkle 路径;UPDATE 变成了同一路径上的新叶子;COMMIT 变成了四次 commit 调用,它们写入的公开输出将由证明绑定。余额来自证明者,这没有问题:根并未承诺的余额无法通过检查,运行以 2 退出。

持有根的合约只有在收到证明、表明这次转换由该程序产生并以 0 退出时,才会接受它:

Ledger.sol(示意)solidity
interface IApogeeVerifier {
    function verify(bytes calldata input, bytes calldata output, uint256 exitStatus,
                    uint256[10] calldata proof, uint256[] calldata points) external view returns (bool);
}

contract Ledger {
    IApogeeVerifier public immutable verifier;
    bytes32 public root;

    constructor(IApogeeVerifier v, bytes32 genesis) { verifier = v; root = genesis; }

    function apply(bytes calldata input, bytes calldata journal,
                   uint256[10] calldata proof, uint256[] calldata points) external {
        require(verifier.verify(input, journal, 0, proof, points), "proof");
        require(bytes32(journal[0:32]) == root, "stale root");
        root = bytes32(journal[32:64]);
    }
}

说明

这只是展示大致形态的草图,不是生产合约。真实部署中,每份证明处理一批请求,由客户程序折叠为一次转换,并且会把验证者固定到正确的程序和公开值长度上。链上结算介绍已部署的验证者、它的密钥及其仪式。

三行概括这个模型#

  1. 链持有一个根。
  2. 客户程序证明这次转换。
  3. 合约推进这个根。

其余的一切,从分片、电路到递归与判定器,都由远地虚拟机负责。这就是这层抽象:一个程序、它的输入与输出,以及把它们联系在一起的一份证明。

下一步#

迈出第一步

远地虚拟机概览

一页看清全部事实:远地虚拟机证明什么、如何证明、成本多少、依赖哪些假设,以及 1.0.0 版止步于何处。

以 Markdown 查看

一段话概述#

远地虚拟机是一台 RISC-V zkVM。它证明:一个由其映像摘要所指明的 RV32IMAC 程序,在给定的公开输入上运行到某个退出状态,并写出了给定的公开输出。它把这份证明经由递归树带到一份 Groth16 证明,交由以太坊合约检验。每个电路都是 BN254 标量域上的分层 GKR 电路,每个已承诺列都用 Mercury 打开,每个挑战都来自 Poseidon2 transcript。域、曲线、配对、MSM、哈希、多项式承诺、GKR 证明者和 Groth16 全部在仓库内实现。它的参考工作负载是以太坊区块校验。

基本事实#

远地虚拟机 v1.0.0
证明陈述的内容 具有此身份的程序,在其映像上从入口点启动,以此公开输入和某份证明者提示(advice)逐条指令执行,直至以此状态调用 EXIT,并已写出此公开输出(journal)
指令集 单 hart 的 RV32IMAC:RV32IMA 的 59 条指令(40 条基础指令、8 条 M 扩展指令、11 条 A 扩展指令),压缩指令在加载时展开
客户程序(guest)语言 Rust,#![no_std] 加 alloc,stable 1.96.1,目标 riscv32imac-unknown-none-elf
算术化 23 个电路族,每个都是分层 GKR 电路:7 个用于指令,5 个用于内存窗口,6 个委托,5 个用于递归
论证 门由求和校验(sumcheck)证明;内存由覆盖整个执行过程的单一读/写多重集证明;查找(lookup)由 LogUp 证明
域 BN254 的标量域,254 位
承诺 Mercury,基于 KZG 的多线性承诺;无论列数多少,每个分片只需一个 704 字节的打开证明
可信设置 PSE 的 perpetual powers of tau,第 80 次贡献;链上判定器另有一个专属于其电路的第二次仪式
Transcript Fr 上的 Poseidon2 双工海绵,宽度 3,速率 2
结算 递归树 → Groth16 判定器 → ApogeeVerifier.sol
安全级别 约 100 位,由 BN254 决定
零知识 否。证明是简洁的,但不是零知识的,也没有任何盲化
委托运算 keccak-f[1600] 轮函数、SHA-256 轮函数、Poseidon2、BN254 Fr 运算、针对四个以太坊模数的 256 位模乘、secp256k1 与 BN254 G1 上的完全点加
公开值 输入最多 16,380 字节,公开输出最多 16,380 字节;证明者提示最多 2 GiB
执行长度 最多 2^36 − 1 个周期
代码大小 表高度为 2^22 时 .text 不超过 7.94 MiB;默认情况下映像不超过 4 MiB
第三方密码学 证明路径上没有。arkworks、Plonky3 和 zkhash 只作为测试参照(oracle)出现

实测数据#

所有数据均来自 glamsterdam-devnet-8 的第 257,510 号区块,由无状态校验器客户程序执行:60 笔交易,101.5 Mgas,198M 个周期。来源:规范的 recursion §10 与 streaming §1。

阶段 结果
基础证明 207 个分片,14.5 MB,在 32 vCPU、247.7 GiB 内存的机器上耗时 2,481 s,峰值 RSS 173.92 GiB
递归树 116 个分片:四个叶节点,每个至多涵盖 64 个基础分片(合计 2,157 s,峰值 92 GiB),以及一个根节点(460 s,1.03 MB)
判定器电路 7,896,686 个约束,定义域大小为 2^23
判定器证明 在一台 18 核笔记本电脑上耗时 18.5 s、占用 6.1 GB,读取密钥用时 1 s
链上验证 3,620,026 gas,34,980 字节 calldata,合约折叠 358 个点
一致性 全部 67,251 个 tests-zkevm v21.0.1 测试对在原生运行下一致

验证者必须持有的值#

两个值,都要取自证明者无法控制的渠道:

  • 程序身份,一个域元素。如果对照的是证明者提供的身份,证明只能说明有某个程序运行过。
  • 仪式的 SRS 摘要。 基于已知 τ 构造的密钥,只有通过这项比对才会被拒绝。

验证密钥本身可以来自任何人:加载时会根据其自身内容重新计算这两个值,并要求其中的电路与验证者的注册表一致。安全模型列出了全部假设。

v1.0.0 止步之处#

  • 不是零知识的。 Mercury、GKR 和判定器中都没有盲化。
  • 证明者提示不受绑定。 客户程序要将它与证明所绑定的某样东西进行核对。
  • 陷入(trap)不可证明。 未对齐的访问、对映射内存之外的访问、ebreak,或者 pc 处没有指令,都会使运行终止,且不产生证明。
  • sc.w 总是成功。 这是与 RV32IMAC 唯一的偏差:没有保留(reservation)状态。
  • 委托固定为六种。 任意模数的 EVM MULMOD、MODEXP 与 BLS12-381 都以普通指令运行。
  • 证明受内存限制。 实测区块的内存峰值为 174 GiB;内存占用取决于同时处理中的分片,而不是运行长度。
  • 判定器密钥随根的形状而定,其可信程度取决于它的仪式。开发用密钥是可伪造的。

谁在构建#

远地虚拟机是 G Web3 面向区块链原生应用环境的研究计划的旗舰项目:每个经济应用一个经过优化的环境,各自以有效性证明结算到以太坊。该计划的立场见核心理念;下一个版本的方向见量子跃迁。

发射你的应用

发射你的应用

远地虚拟机的开发者手册:客户程序如何编写、构建、运行、证明并在链上结算,以及让它保持正确、可证明且低成本的习惯。

以 Markdown 查看

客户程序(guest)是远地虚拟机所证明的程序:一个为 riscv32imac-unknown-none-elf 编译的 no_std Rust 二进制程序,带一个入口点、三个用于输入和输出的内存区域,除此之外别无他物。宿主程序(host)是围绕它的一切:负责提供输入、向远地虚拟机请求证明、再把证明交给检验方的代码。两者都由你编写。远地虚拟机提供机器、电路和验证者。

本节同时写给两类读者:坐在键盘前的工程师,以及与他们协作的模型。每一页都把规则直白地写出来,AI 随行手册则把全部规则浓缩进一个文件,你可以在助手写下第一行代码之前就交给它。

基本模型#

客户程序、宿主程序与验证者 宿主程序把公开输入和证明者提示(advice)交给运行在远地虚拟机中的客户程序。客户程序写出公开输出(journal),并以某个状态退出。远地虚拟机生成一份证明,绑定程序身份、输入、公开输出和退出状态,由验证者检验。 宿主程序 · 由你编写 你的服务 构建输入, 提供证明者提示, 请求证明 远地虚拟机 · RV32IMAC · 单 hart 客户程序 · 由你编写 你的程序 no_std Rust guest_sdk 公开输入 证明者提示 公开输出 · 退出状态 验证者 检验证明 所对照的身份来自 它自己的渠道; 合约或服务均可
各方分工。输入和公开输出(journal)受证明绑定;以虚线表示的证明者提示(advice)则不受绑定,所以客户程序要检查它。验证者从不接触程序本身,只看到它的程序身份。

一份证明只陈述一件事:具有此身份的程序,在其映像上以此公开输入和某份由证明者选定的证明者提示启动,运行至以此状态调用 EXIT,并已写出此公开输出。你构建的一切都建立在这句话之上。

工作流程#

步骤 你要做的事 页面
1 无需手动安装任何东西:代码仓库固定了工具链。为证明获取仪式文件 环境准备
2 编写客户程序:一个入口点、三个内存区域、普通的 Rust 编写客户程序、输入、证明者提示与公开输出
3 在划算的地方使用委托的哈希与曲线运算 委托
4 为客户程序目标构建,然后检查映像及其程序身份 构建与检查
5 在模拟器中运行,统计周期都花在了哪里 运行与性能分析
6 证明一次运行,并验证它 证明与验证
7 通过递归压缩证明,并在以太坊上检验 链上结算

快速上手用一个只有三行的客户程序,把整个流程完整走一遍。

最重要的规则#

客户程序编程指南逐条解释了这些规则,说明违反时会出什么问题,以及应该怎么做。

  • usize 和所有指针都是 32 位。 usize 溢出只在客户程序中触发 panic,x as usize 会静默截断,任何布局或哈希依赖于长度的东西,在宿主机与客户程序之间都不一致。
  • 分配器从不释放内存。 它从 __heap_start 起向上推进一个指针,所以让客户程序耗尽内存的,是它在整个运行期间分配的总量,而不是峰值。复用缓冲区,并用 with_capacity 预先设定容量。
  • 原子操作能正常工作,但你不应该写它们。 这台机器只有一个 hart,A 扩展的存在是为了兼容已经在使用它的代码。新写的客户程序代码没有任何需要同步的东西。
  • 没有外部世界。 没有文件、没有时钟、没有随机数、没有网络。客户程序所知道的,只有它的公开输入、它的证明者提示,以及它自己算出的东西。
  • 证明者提示由证明者选择。 在任何由它推导出的内容进入公开输出之前,先把它与证明所绑定的某样东西核对。
  • 公开输出很小。 最多 16,380 字节。任何会增长的内容,都发布其摘要。
  • release 构建中溢出检查保持开启。 它们是程序计算内容的一部分,所以客户程序的构建 profile 把它们固定下来。
  • 每条执行过的指令都是一行被证明的数据。 证明成本取决于周期数,所以先用 --release 构建并统计周期数,再去优化其他任何东西。

从哪里开始#

发射你的应用

快速上手

从一个空 crate 到一份通过验证的证明。一个三行的客户程序,经过构建、运行、检查和证明,附上每一步的真实输出。

以 Markdown 查看

本页用一个最小的、确实做了点事情的客户程序(guest),把整个流程完整走一遍:它读取自己的公开输入,并把它作为公开输出(journal)发布。下面的每一段输出,都是在远地虚拟机 v1.0.0 上原样运行这些命令得到的。

说明

前提条件。 一份检出到 v1.0.0 的远地虚拟机代码仓库,以及 rustup;其余一切都由仓库固定。除非某一步切换了目录,命令都在仓库根目录下运行。第 5 步和第 6 步还需要仪式文件 assets/ptau/ppot_0080_24.ptau,第 6 步还需要一台有数十 GiB 内存的机器。环境准备对这两项都有说明。

创建客户程序#

客户程序是 guests/ 工作空间中的一个 no_std 二进制 crate。创建 guests/hello:

guests/hello/Cargo.tomltoml
[package]
name = "hello"
version.workspace = true
edition.workspace = true
publish.workspace = true

[dependencies]
guest-sdk.workspace = true
guests/hello/src/main.rsrust
#![no_std]
#![no_main]

guest_sdk::entry!(main);

fn main() {
    // The public input is memory: a slice, with no ecall and no cursor.
    guest_sdk::commit(guest_sdk::public_input());
}

使用 #![no_std],是因为目标是裸机。使用 #![no_main] 加 entry!(main),是因为 SDK 的启动代码会设置栈指针、把 .bss 清零,然后调用一个 main 符号,这个符号由宏包装你的函数后导出。从你的函数返回即为 exit(0)。

加入客户程序工作空间#

在 guests/Cargo.toml 的 members 列表末尾加上 "hello":

guests/Cargo.tomltoml
members = ["fib", "echo", … , "recursion", "hello"]

构建#

在客户程序自己的目录下构建,除了目标之外不加任何参数:

sh
cd guests/hello
cargo build --release --target riscv32imac-unknown-none-elf
cd ../..

ELF 生成在 guests/target/riscv32imac-unknown-none-elf/release/hello。客户程序工作空间已经提供了链接脚本和 --no-relax,所以无需再传任何参数。

运行#

周期分析器在远地虚拟机的模拟器中运行客户程序,不生成证明,并报告周期都花在了哪里:

sh
printf 'hello, apogee' > /tmp/hello.in
cargo run --release -p profiler -- elf guests/target/riscv32imac-unknown-none-elf/release/hello --input /tmp/hello.in
workload
  label                        hello
  guest                        hello
  guest cycles                 114
  exit status                  0
  journal bytes                13

cycles by family
  ADD_SUB_LUI_AUIPC            64
  JUMP_BRANCH_SLT              21
  MEM_WORD                     3
  MEM_SUBWORD                  26

一共执行了 114 条指令,每一条都将成为一行被证明的数据。公开输出的 13 个字节就是原样回显的输入。26 行 MEM_SUBWORD 来自 commit 用 lbu 和 sb 逐字节复制输入。

查看虚拟机将要证明的内容#

sh
cargo run --release -p artifact-dump -- tables \
    guests/target/riscv32imac-unknown-none-elf/release/hello \
    --ptau assets/ptau/ppot_0080_24.ptau
program identity  9ead85cee880df30daa8eba657316215107a075640a64ccf2424a054b758b802

VmConfig
--------
  id  family              height     live rows  columns
   0  ADD_SUB_LUI_AUIPC     4194304         33  pc next_pc rs1 rs2 rd imm extra_mask
   1  JUMP_BRANCH_SLT       4194304         12  pc next_pc rs1 rs2 rd imm extra_mask
   4  MEM_WORD              4194304          3  pc next_pc rs1 rs2 rd imm extra_mask
   5  MEM_SUBWORD           4194304          3  pc next_pc rs1 rs2 rd imm extra_mask
   7  INIT_TEARDOWN         4194304          0  none: claims no pc
   8  ZERO_WINDOWS          4194304          0  none: claims no pc
  12  PUBLIC_INPUT             4096          0  none: claims no pc
  13  PUBLIC_OUTPUT            4096          0  none: claims no pc
  14  ADVICE_WINDOWS        4194304          0  none: claims no pc

这是程序在默认高度下的静态形态:其代码用到的四个指令电路族,各带一张解码表;以及每个程序都有的五个窗口电路族。程序身份是一个域元素,是以上全部内容的摘要。你得到的值会不同:ELF 会把绝对路径嵌入 panic 字符串,所以在另一台机器上构建得到的是另一个映像;而高度的每一次改变,都会得到另一个程序身份。

证明与验证#

由宿主程序(host)请求证明。把它作为示例放在宿主程序 SDK 旁边:

crates/host/examples/prove_hello.rsrust
use constants::family;
use emulator::GuestIo;
use program::ProgramParams;
use srs::Srs;

fn main() {
    let elf = std::fs::read("guests/target/riscv32imac-unknown-none-elf/release/hello")
        .expect("build the guest with --release first");

    // Small heights for a small program: the seven instruction families at
    // their 2^20 floor, the three RAM-window families at 2^16. Every choice of
    // heights is its own program identity.
    let mut params = ProgramParams::defaults();
    for f in 0..7 {
        params.heights[f] = 1 << 20;
    }
    for f in [family::INIT_TEARDOWN, family::ZERO_WINDOWS, family::ADVICE_WINDOWS] {
        params.heights[f as usize] = 1 << 16;
    }

    // As many ceremony powers as the tallest family has rows: 2^20 here.
    let ptau = std::path::Path::new("assets/ptau/ppot_0080_24.ptau");
    let srs = Srs::from_ptau(ptau, 20).expect("the ceremony file reads");
    let setup = host::setup(&elf, &params, srs).expect("the program registers");

    let io = GuestIo { input: b"hello, apogee".to_vec(), advice: Vec::new() };
    let proven = host::prove(&setup, &io, 2).expect("the run proves"); // two shards in flight
    host::verify(&setup.vk, &proven.block).expect("the block verifies");

    assert_eq!(proven.exit_code, 0);
    assert_eq!(proven.journal, b"hello, apogee");
    let id: String = setup.vk.identity.to_bytes().iter().map(|b| format!("{b:02x}")).collect();
    println!("identity  {id}");
    println!("cycles    {}", proven.cycles);
    println!("shards    {}", proven.report.shards);
    println!("journal   {:?}", core::str::from_utf8(&proven.journal).unwrap());
}
sh
cargo run --release -p host --example prove_hello
identity  606d1f1d720459cc1a078787381656b29c9fce5a9e539b36f899e62b64129c14
cycles    114
shards    7
journal   "hello, apogee"

在一台 18 核、48 GiB 内存的笔记本电脑上,这一步耗时 52 秒,内存峰值 18 GB,几乎全部来自同时处理中的两个 2^20 分片。这里的程序身份与第 5 步的不同,因为高度不同:程序身份绑定每一个高度。

保管程序身份#

验证者从不从证明、密钥或证明者那里获取程序身份。它持有自己的副本,从构建该发布版本的一方取得,并进行比对:

rust
assert_eq!(setup.vk.identity.to_bytes(), registered); // `registered` from your own channel

如果对照的是证明者提供的身份,证明只能说明某个程序运行过。

刚才发生了什么#

模拟器把这 114 条指令执行了两遍。第一遍承诺了每个分片的内存列,并确定了陈述;第二遍填充每个分片并证明它。一共七个分片:执行过的四个指令电路族各一个,存放程序映像的内存窗口一个,公开输入和公开输出各一个。这个客户程序从未触及自己的栈,所以没有其他窗口需要分片;典型的程序还会多出栈所在窗口的分片。每个分片都由其电路族的 GKR 电路证明,并用一个 Mercury 证明打开;验证者用一个等式核对了全部七个分片的内存读写。系统架构概述沿着同一条路径给出了详细说明。

下一步#

发射你的应用

环境准备

代码仓库固定的工具链、仓库中的两个工作空间、证明所需的仪式文件,以及每一步对机器的要求。

以 Markdown 查看

远地虚拟机 v1.0.0 是一个 Rust 代码仓库。除了 rustup 之外无需安装任何东西:仓库固定了自己的工具链,工具链自带客户程序(guest)的编译目标。编写、构建、运行客户程序以及分析其性能,都不再需要别的东西。证明则额外需要一个大文件,以及一台内存相应充足的机器。

工具链#

仓库根目录下的 rust-toolchain.toml 固定了 stable 版 Rust 1.96.1,连同 rustfmt、clippy 和 llvm-tools,以及目标 riscv32imac-unknown-none-elf,该目标的 core 和 alloc 以预编译形式提供。rustup 会在根目录以下的每个目录中应用它,并在首次使用时安装。

sh
cd apogee-vm
rustup show active-toolchain     # 1.96.1, overridden by rust-toolchain.toml
cargo --version

仓库中任何地方都不使用 nightly,也不使用任何不稳定特性。llvm-tools 提供与编译器的 LLVM 相匹配的 llvm-objdump 和 llvm-nm:仓库用它们生成已提交的反汇编清单,你也可以用它们阅读客户程序的代码。

两个工作空间#

检出的代码中有两个 Cargo 工作空间,这一划分很重要:

工作空间 根 构建目标 包含
根工作空间 Cargo.toml 你的宿主机(host) 证明者、验证者、宿主程序 SDK、各种工具,即 crates/ 和 tools/ 中的一切
客户程序工作空间 guests/Cargo.toml riscv32imac-unknown-none-elf 所有客户程序,以及它自己的 guests/target 目录

客户程序之所以单独存放,是因为每个成员都为客户程序目标编译,并链接一个 #[panic_handler];在根目录运行的 cargo test --workspace 绝不能触及它们。客户程序工作空间还带有正确构建客户程序所需的设置,所以你永远不必手动输入:

  • guests/.cargo/config.toml 设定目标,并向链接器传入 -T crates/guest-sdk/link.ld(即内存布局)和 --no-relax,因为链接器松弛(relaxation)会移动程序身份所绑定的地址。
  • guests/Cargo.toml 把两个构建 profile 固定为相同的语义,溢出检查也包括在内(构建与检查)。
  • 它的 [patch.crates-io] 把 k256、ark-ff 和 revm-precompile 指向仓库内的 vendored 副本,这些副本会调用远地虚拟机的委托(委托)。

提示

在编辑器中把 guests/ 作为单独的文件夹打开。这样 rust-analyzer 会读取该工作空间的 .cargo/config.toml,按客户程序目标而不是你的宿主机来检查客户程序代码。

仪式文件#

远地虚拟机做出的每个承诺,都基于一场公开仪式所产生的秘密 τ 的各次幂:PSE 的 perpetual powers of tau,第 80 次贡献。一个文件满足所有用途:

assets/ptau/ppot_0080_24.ptau        19.3 GB, 2^24 powers; assets/ptau/ is gitignored

计算程序身份、构建真实密钥和执行证明都需要它。构建、运行客户程序和分析其性能,或者运行工作空间测试,都不需要它;这些测试在各自的玩具设置上做证明。

PSE 仪式的各个文件截取自同一份 transcript,所以任何幂次不低于 24 的文件都可以用。Hermez 的 powersOfTau28_hez_final_*.ptau 是另一场仪式,τ 也不同:读取器同样能顺利读入它,但基于它得到的每个承诺、密钥和程序身份都会不同。为确认你持有的是正确的仪式文件,可以核对它的 [τ]_1,其规范编码 x ‖ y 的十六进制表示为:

9bbb31bedc304e081e2aada4b56c2217e0e94ee16874e3517d14bef5dcec3a16
317ff1589e53513fa333591b318e8f1e55ef7c37d92beb6d9a61d770a2f39506

规范中的 SRS 页面准确说明了读取器检查哪些内容、又默认信任哪些内容。

机器#

构建和运行在笔记本电脑上就能完成。证明受内存限制,其内存占用取决于同时证明的分片,而不是运行的长度。

步骤 需要
构建、运行客户程序并分析其性能,导出并检查其映像 任何一台较新的笔记本电脑;数秒
计算程序身份(artifact-dump tables --ptau) 仪式文件;在默认高度下,18 核笔记本电脑上约 25 s
以 2^20 的高度证明一个小型客户程序 数十 GiB。最宽的指令电路族的一个 2^20 分片,在前向过程中约占用 8.4 GiB 的域元素,每个同时处理中的分片各自占用一份
证明一个完整的以太坊区块 实测区块在一台 32 vCPU、247.7 GiB 内存的机器上,峰值为 174 GiB

证明页面解释了高度和同时处理的分片数如何在内存与时间之间权衡:证明与验证。

检查你检出的代码#

CI 运行的内容如下,全部不需要仪式文件:

sh
cargo fmt --all -- --check
cargo clippy --workspace --all-targets -- -D warnings
cargo test --workspace
cargo run -p kat-gen && git diff --exit-code      # committed fixtures regenerate identically
(cd guests && cargo clippy --bins -- -D warnings)

证明真实分片的测试套件都标记了 #[ignore],因为每个都需要数十 GiB 内存。想在自己的机器上亲眼看到证明的生成与拒绝时,按名称运行其中一个:

sh
cargo test --release -p prover --test acceptance -- --include-ignored --test-threads=1

下一步:编写客户程序,或者在快速上手中把整个流程走一遍。

发射你的应用

编写客户程序

客户程序是一个带入口点和三个内存区域的 no_std Rust 二进制程序。crate 结构、底层运行时、依赖,以及让你能像测试任何 Rust 代码一样测试它的宿主机优先布局。

以 Markdown 查看

客户程序 crate#

在 v1.0.0 中,客户程序(guest)是代码仓库 guests/ 工作空间中的一个二进制 crate。工作空间提供目标、链接器参数、固定的 profile 和 vendored crate,所以客户程序自己的清单可以保持简短:

guests/my-app/Cargo.tomltoml
[package]
name = "my-app"
version.workspace = true
edition.workspace = true
publish.workspace = true

[dependencies]
guest-sdk.workspace = true
guests/my-app/src/main.rsrust
#![no_std]
#![no_main]

extern crate alloc; // Vec, Box, String, BTreeMap, over the SDK's allocator

use alloc::vec::Vec;

guest_sdk::entry!(main);

fn main() {
    let input = guest_sdk::public_input();
    let mut out = Vec::with_capacity(input.len());
    out.extend(input.iter().rev());
    guest_sdk::commit(&out);
}

把 "my-app" 加入 guests/Cargo.toml 的 members,然后在客户程序自己的目录下构建:cargo build --release --target riscv32imac-unknown-none-elf。

底层运行着什么#

客户程序 SDK 就是全部运行时。它小到可以完整列出:

  • 启动。 _start 位于 0x0001_0000,即 .text 的第一个字节。它让 sp 指向 RAM 顶端,逐字节将 .bss 清零,然后调用 main。entry!(f) 导出的那个 main 是对你的函数的包装;你的函数不带参数,返回 ()。
  • 退出。 从 main 返回即为 exit(0)。guest_sdk::exit(code) 以任意状态结束运行。非零状态表示执行失败,而失败的执行仍然可以被证明:陈述带有这个状态,验证者会读取它。
  • Panic。 panic 处理程序以状态 101 退出,不写出任何内容。没有诊断输出流。发生 panic 的客户程序,在 panic 之前提交的内容依然已经发布。
  • 堆。 一个 bump 分配器从紧挨 .bss 之上的 __heap_start 向上增长。它从不释放内存。参见堆。
  • 系统调用。 客户程序发出的 ecall 只有 EXIT,以及 SDK 替你发出的委托调用。输入、证明者提示(advice)和输出都是内存,用普通的加载和存储指令读写。

内存布局#

客户程序所见的整个 32 位地址空间:

范围 内容
0x0000_0000 – 0x0000_8000 空洞。没有任何东西初始化这一段,所以空指针或野指针会导致致命的 OutOfBounds,而不是悄无声息地读出数据
0x0000_8000 – 0x0000_C000 公开输入窗口,16 KiB
0x0000_C000 – 0x0001_0000 公开输出(journal)窗口,16 KiB
0x0001_0000 – … .text(_start 在最前),然后是 .rodata、.data 和 .bss,各自按页对齐
__heap_start 往上 堆,从 .bss 末尾向上取整到 16 的位置开始
0x7F80_0000 – 0x8000_0000 栈的 8 MiB 预留区。任何堆块的末端都不得高于 0x7F80_0000;栈从 0x8000_0000 向下增长
0x8000_0000 – 2^32 证明者提示区域,最多 2^29 个字,只有宿主程序(host)实际提供的部分可以寻址

代码是静态的。每个 pc 处的指令都取自程序的解码表,从不取自 RAM,所以向 .text 写入会改变之后的加载读到的内容,但不会改变执行的指令。

堆#

分配器向上推进一个指针,dealloc 什么也不做。对于一个每个周期都要耗费证明时间的短程序,这是正确的设计,而它也改变了你写 Rust 的方式:

  • 让你耗尽内存的是分配总量,而不是峰值。 一个每轮迭代都构建并丢弃一个 Vec 的循环,每次都会消耗新的堆空间。
  • 复用缓冲区。 把分配提到循环之外;用 clear() 清空后重新填充,而不是重新分配;用 with_capacity 为会增长的集合预设容量,让它们在增长时不必重新分配和复制。
  • 上限是退出状态 71。 一次分配如果末端会高于 0x7F80_0000,或高于当前的栈指针,就以状态 71 退出,而不是返回空指针或覆盖栈。
rust
// Allocates a fresh Vec per record: total heap grows with the record count.
for record in records {
    let fields: Vec<&[u8]> = record.split(|b| *b == b',').collect();
    handle(&fields);
}

// One buffer, reused: total heap is the largest record's field count.
let mut fields: Vec<&[u8]> = Vec::with_capacity(16);
for record in records {
    fields.clear();
    fields.extend(record.split(|b| *b == b','));
    handle(&fields);
}

栈有 8 MiB 的预留区,在其中深度递归没有问题。没有任何机制能察觉的情况是:堆已经填满了预留区下方的空间,栈又越过了预留区,这时堆块会在深层调用链之下被改写。让递归保持有界,或者改写成迭代。

依赖#

任何不依赖 std、能为 riscv32imac-unknown-none-elf 构建的 crate 都可以用。实践中:

  • 关闭默认特性(default-features = false),并在 crate 提供 alloc 时启用它。
  • 引入 getrandom、时钟,或带随机种子的 std::collections::HashMap 的 crate,在这里没有任何来源可用。这样的调用会得到 -ENOSYS,并使这次运行无法被证明。优先使用 BTreeMap,或者使用固定、确定性哈希器的哈希表。
  • 浮点运算会编译成整数软件例程,因为目标没有 F 或 D 扩展。它正确且确定,但每次运算都要花费许多条指令。整数或定点运算更便宜。
  • 哈希和椭圆曲线运算有专用电路。使用 SDK 的函数或 vendored crate,让你的依赖能用上这些电路:委托。

先在宿主机上测试#

客户程序不打印任何东西,所以调试在宿主机上进行。让这件事变得容易的布局是:把程序写成一个从字节到字节的 #![no_std] 库;让 main.rs 只负责把字节搬进、搬出各个内存区域;并且只在客户程序目标上依赖 SDK:

guests/my-app/Cargo.tomltoml
[package]
name = "my-app"
version.workspace = true
edition.workspace = true
publish.workspace = true

[target.'cfg(target_arch = "riscv32")'.dependencies]
guest-sdk.workspace = true
guests/my-app/src/lib.rsrust
#![no_std]
extern crate alloc;
use alloc::vec::Vec;

/// The whole application: public input and advice in, journal out.
pub fn run(input: &[u8], advice: &[u8]) -> Result<Vec<u8>, i32> {
    let _ = advice;
    let mut out = Vec::with_capacity(input.len());
    out.extend(input.iter().rev());
    Ok(out)
}
guests/my-app/src/main.rsrust
#![no_std]
#![no_main]

guest_sdk::entry!(main);

fn main() {
    match my_app::run(guest_sdk::public_input(), &[]) {
        Ok(journal) => guest_sdk::commit(&journal),
        Err(code) => guest_sdk::exit(code),
    }
}

宿主程序代码随后按路径依赖这个库(就像 crates/emulator 依赖 guests/revm-block 那样),原生运行 my_app::run,再把结果与模拟器对同一输入产生的公开输出相比较(运行与性能分析)。你的逻辑在宿主机上可以用单元测试、调试器和 println!,而客户程序二进制只是包在已测试代码外面的一层薄壳。

警告

两种构建对 usize 的理解不同。 在客户程序上,usize 和所有指针都是 32 位;在你的宿主机上是 64 位。usize 溢出只在客户程序上触发 panic,x as usize 在那里会静默截断,而任何包含长度的东西,其 size_of 和 core::hash 在两者之间都不相同。不要让 usize 出现在你提交、哈希或序列化的任何东西中,在这些边界上使用显式的 u32 和 u64。

汇编与指令集#

解码器恰好接受 RV32IMA 的 59 条指令,以及在加载时展开的压缩(C)指令。在这个指令集之内,内联汇编没有问题。任何超出它的东西,例如 CSR 访问、fence.i、浮点或 RV64 编码,即使永远不会被执行到,也会让整个程序无法登记:推导时会报告 Not all opcodes supported: pc=…。ebreak、跳转到一个没有指令的半字,或者未对齐的半字或字访问,都会使运行终止且不产生证明。

原子指令可以解码,也可以证明,只有一处偏差:sc.w 总是成功,因为这台机器不保存保留(reservation)状态。客户程序编程指南解释了为什么新的客户程序代码根本不应该使用原子操作。

下一步#

发射你的应用

输入、证明者提示与公开输出

客户程序没有 I/O 系统调用。它的公开输入、证明者提示和公开输出是三个内存区域。各自存放什么、证明绑定什么,以及每个接收数据的客户程序都遵循的模式。

以 Markdown 查看

远地虚拟机的客户程序(guest)没有文件描述符、没有流,也没有 I/O 系统调用。它的输入和输出是三个内存区域,用普通的加载和存储指令读写,证明绑定其中的两个。

三个内存区域#

区域 SDK 内容 大小 是否受证明绑定
公开输入 public_input()、read_input(buf) 陈述中的字节,由请求证明的一方选择 最多 16,380 字节 是,绑定其初始内容
证明者提示(advice) advice() 证明者选择的字节 最多 2 GiB 否
公开输出(journal) commit(bytes)、journal() 客户程序追加写入的内容 最多 16,380 字节 是,绑定其最终内容
rust
let input: &[u8] = guest_sdk::public_input(); // a slice over the input window, no copy
let data: &[u8] = guest_sdk::advice();        // a slice over the advice region
guest_sdk::commit(b"result");                  // appends to the journal
  • public_input() 和 advice() 返回的是内存上的切片,不发生任何复制。read_input(buf) 复制 min(buf.len(), input.len()) 个字节并返回实际数量,所以返回值可能小于请求的长度。
  • commit 追加数据并维护一个长度字,正是这个长度字让证明绑定一个确切的字节串,而不是一个用零填充的窗口。它宁可以状态 70 退出,也不会让窗口溢出。
  • 在没有提供证明者提示的运行中调用 advice(),会导致致命的 OutOfBounds,而不是返回空切片:没有证明者提示的运行根本没有证明者提示区域,也不必为它付出任何代价。

“绑定”是什么意思#

证明所确立的陈述,带有公开输入的字节、公开输出的字节和退出状态。证明表明:在客户程序第一次访问之前,输入窗口存放的恰好是陈述中的输入;客户程序退出时,公开输出窗口存放的恰好是陈述中的输出。这一点依靠的是内存论证,而不是客户程序做的任何事:客户程序不必计算任何哈希,也不必遵循任何约定。

证明者提示则不同。证明者提示区域的初始内容是证明者写进去的任何东西,没有任何东西把它与程序身份、陈述或任何门联系起来。证明只说明:存在某份证明者提示,使程序在这个输入下发布了这份公开输出。这一保证的强度,恰好等于客户程序自己对证明者提示所做检查的强度。

模式:承诺、提供、检查#

输入较大的客户程序,把大部分数据作为证明者提示接收,由公开输入对它做出承诺;在任何由证明者提示推导出的内容进入公开输出之前,先把两者相互核对:

先校验证明者提示,再信任它rust
fn main() {
    let want = guest_sdk::public_input(); // 32 bytes: keccak256 of the advice
    let data = guest_sdk::advice();       // the prover's bytes, bound by nothing
    if guest_sdk::keccak256(data).as_slice() != want {
        guest_sdk::exit(1); // refused before anything derived from it is committed
    }
    let sum = data.iter().fold(0u32, |s, b| s.wrapping_add(u32::from(*b)));
    guest_sdk::commit(&sum.to_le_bytes()); // the journal: what the proof publishes
}

这项检查不一定是对整个证明者提示求哈希。它可以是对照输入所携带的根来检查的 Merkle 路径,就像账本示例那样;也可以是对数据的签名;或者是结果本身满足的约束,例如证明者声称的排序顺序,由客户程序用一遍扫描来验证,而不必自己排序。关键在于,用来核对的那样东西必须是受绑定的。

注意

提交未经检查的证明者提示的任何函数值,就等于发布一个由证明者选定的值。这是写出一个证明毫无意义的客户程序最常见的方式。

结构化数据#

用 no_std 序列化器对结构化输入编码,例如基于 serde、启用 alloc 特性的 postcard;仓库自己的以太坊客户程序就用它来编码区块见证。两个习惯能让格式保持严谨:

  • 使用定宽整数。 用 u32 和 u64,绝不用 usize,它的宽度在你的宿主机(host)和客户程序之间不同。
  • 在要紧的地方,坚持每个值只有一种编码。接受尾随字节或非最短变长整数(varint)的反序列化器,会让同一个值对应两个字节串。在唯一性要紧的地方,先解码、再重新编码并比较,以太坊客户程序的 BlockWitness::decode 就是这样做的。

会增长的输出#

公开输出最多容纳 16,380 字节。随工作量增长的输出,例如每笔交易一条记录,没有固定的上界,迟早会以 70 退出。改为发布摘要:边生成记录边对其求哈希,提交 32 字节的结果,再让需要这些记录的一方在原生环境中重新计算并比对。仓库中的无状态以太坊校验器就是这样为整个区块发布一份 43 字节的公开输出。

对于要在以太坊上检验的证明,两个公开值都要保持固定长度。部署的验证者合约是针对一个输入长度和一个输出长度构建的,任何其他长度都会被拒绝(链上结算)。

退出状态#

退出时,除了已经提交的内容之外,不会发布任何东西;而发生 panic 或以非零状态退出的运行,对它所做的事有一份有效的证明。所以验证者先读退出状态,再读公开输出。在链上,验证者合约把预期的状态作为参数,应用传入 0。SDK 自身使用的状态:

状态 含义
0 main 返回,或 exit(0)
70 commit 将超过 16,380 字节
71 某次分配将触及栈:参见堆
72 某个委托返回了其 shim 拒绝接受的结果
101 panic,不打印任何内容

自定义的失败代码要避开这些值,就像账本示例使用 1、2 和 3 那样。

证明没有说明什么#

  • 没有任何东西规定公开输出的写入顺序,也没有任何东西强制客户程序读取它的输入。证明绑定的是窗口的内容,而不是产生这些内容的访问。
  • 证明者提示是可写的。 向证明者提示区域写入就是一次普通的存储。无论写与不写,它都不受绑定。
  • 公开输出就是窗口的全部最终内容。 commit 会维持这种格式。直接写这个窗口的客户程序必须自己保持它:一个不超过 16,380 的长度,随后是这么多字节,再之后全是零。

规范对这一切有精确的表述:公开值与证明者提示。

发射你的应用

委托

哈希、域运算和曲线运算都有专用电路。哪些 SDK 调用会用到它们、成本多少、对操作数的规则,以及把库代码引向它们的 vendored crate。

以 Markdown 查看

有些计算,用专门为它们构建的电路来证明,要比作为一串 RISC-V 指令来证明便宜得多。远地虚拟机把这些称为委托。委托是一个电路族,由 ecall 调用,证明作用于 RAM 中一个由字组成的帧的某个函数;客户程序(guest)SDK 在普通函数背后替你发出这些调用。你永远不必自己写 ecall。

你调用什么,会用到哪个委托#

你调用的 委托 一次调用证明的内容
guest_sdk::keccak256(&[u8]) -> [u8; 32] KECCAK_F keccak-f[1600] 的一轮;一次置换是 24 次调用,海绵结构和填充由客户程序代码完成
guest_sdk::sha256(&[u8]) -> [u8; 32] SHA256_COMP 压缩函数的四轮;一次压缩是 16 次调用
guest_sdk::ec_add、ec_mul、ec_identity EC_ADD secp256k1 或 BN254 G1 上一次完全点加的三分之一
guest_sdk::poseidon2_permute(&mut [u8; 96]) POSEIDON2 Fr 上一次宽度为 3 的 Poseidon2 置换
field::Fr 的加法、乘法、求逆 FR_ARITH 一次 Fr 运算,在客户程序目标上发生,无需显式指名任何东西
transcript::poseidon2_permute POSEIDON2 同一个置换,经由 transcript crate
作用于 ModMulFrame 的 guest_sdk::recursion::mod_mul MOD_MUL 一次 256 位的 a·b mod m,m 是四个以太坊模数之一

这些函数的结果与其软件定义逐位一致。keccak256 是以太坊的 Keccak,而不是 SHA3-256。sha256 遵循 FIPS 180-4。ec_add 使用 Renes、Costello 和 Batina 的完全加法公式(2015,算法 7),因此倍点、P + (−P)、单位元以及任意 Z 都不需要特殊处理。

在客户程序中调用哈希与曲线运算rust
use guest_sdk::{ec_mul, keccak256, recursion::SECP256K1_GROUPS, ProjectivePoint};

let digest: [u8; 32] = keccak256(b"blockchain-native");

// A point is homogeneous projective (x = X/Z, y = Y/Z), each coordinate eight
// little-endian u32 limbs below the field modulus. The scalar is eight limbs too.
fn times(p: &ProjectivePoint, k: &[u32; 8]) -> ProjectivePoint {
    ec_mul(&SECP256K1_GROUPS, p, k).expect("EC_ADD is implemented on Apogee")
}

库代码同样能用上委托#

客户程序工作空间给三个 crate 打了补丁,使其中的代码在客户程序目标上调用委托,并以上游代码作为回退路径:

Crate 版本 用到的委托
k256 0.13.4 域元素乘法和标量乘法用到 MOD_MUL;ProjectivePoint 的加法、混合加法和倍点用到 EC_ADD
ark-ff 0.6.0 BN254 两个域上的 Montgomery 乘法和平方用到 MOD_MUL
revm-precompile 43.0.2 预编译合约 0x02 用到 SHA256_COMP;0x06 和 0x07 用到 EC_ADD

依赖这些 crate 的客户程序,会通过 guests/Cargo.toml 的 [patch.crates-io] 自动得到打过补丁的副本。不打补丁时,仅 k256 的域乘法和平方就占了一个主网区块 44% 的周期。secp256k1 签名恢复是普通的 k256 代码,补丁把它变成了委托运算。

委托的成本#

只有当程序链接了某个委托电路族的某个 shim 时,该电路族才成为程序的一部分;调用的成本以该电路族高度的分片计:

  • 链接了但从未调用:没有成本。 该电路族已声明,证明零个分片。
  • 调用一次:一整个分片。 不论占用率如何,一个分片的成本都按其完整高度计算。
  • 大量调用:每次调用的成本很低。 分片高度增长时,其证明只按每个变量增加一轮求和校验(sumcheck)。
电路族 高度 工作单元 每单元调用次数 每分片单元数
KECCAK_F 2^18 keccak-f[1600] 24 10,922
SHA256_COMP 2^18 一次压缩 16 16,384
EC_ADD 2^16 一次完全点加 3 21,845
MOD_MUL 2^16 一次 a·b mod m 1 65,536
POSEIDON2 2^8 一次置换 1 256
FR_ARITH 2^8 一次 Fr 运算 1 256

代价主要在内存而不是时间:一个 2^18 的 KECCAK_F 分片在前向过程中约占用 42 GiB 的域元素,实测以太坊区块的内存峰值,就是由两个同时处理中的此类分片决定的。

操作数规则#

  • 操作数必须小于其模数。 MOD_MUL 或 EC_ADD 的操作数若大于或等于其选择子所指定的模数,就没有证明:执行器会以致命错误 DelegationFrame 拒绝该帧。vendored 的 k256 在调用之前,会先对它惰性约简的域元素做约简。
  • 点不会替你检查。 EC_ADD 证明的是公式的算术。一个点是否在曲线上,是调用代码要回答的问题;从证明者提示(advice)中取点的客户程序必须检查这一点。
  • 每个多次调用的操作都对应一个 SDK 函数。 一次 keccak 置换是在同一个帧上的 24 次调用,一次 SHA-256 压缩是 16 次,一次点加是 3 次。每次调用只证明自己的那一步,顺序错了也不会被拒绝:它只会算出别的东西。使用会按顺序发出调用的 keccak256、sha256 和 ec_add,而不是原始的 shim。
  • 退出状态 72 表示某个委托返回了其 shim 拒绝接受的结果。在远地虚拟机自己的执行器上,格式正确的帧不会出现这种情况。

哪些运算没有委托#

  • 任意模数的 EVM MULMOD、MODEXP、BLS12-381,以及上表之外的所有原语,都以普通指令运行。
  • 没有任何签名方案或配对作为整体被委托。secp256k1 签名恢复是建立在 MOD_MUL 和 EC_ADD 之上的 k256;BN254 配对是建立在 MOD_MUL 之上的 ark-bn254。
  • 委托承担的是运算的核心。填充、海绵结构、分块循环以及标量乘法的阶梯(ladder),都是客户程序代码,以指令的形式被证明。

面向客户程序的签名方案在 v2.0.0 路线图上。

新增一个委托值得吗?#

周期分析器在每份报告中都会为显而易见的候选项定价,给出一个委托最多能省掉多少周期的上限:

removable = max(0, cycles − calls·(4 + 2·frame_words))

cycles 是该类别在这次运行中所占的周期,calls 是进入候选函数的次数,4 + 2·frame_words 是委托之后仍会留下的 shim:帧的存储、ecall 以及结果的加载。它没有计入新电路族分片的成本,所以只能当作上界。运行与性能分析给出了一份报告。

每个委托逐列展开的规范,见委托电路。

发射你的应用

构建与检查

固定为同一语义的构建 profile、构建产出的 ELF、加载器由它生成的 ProgramImage 及其报告,以及验证者登记的程序身份。

以 Markdown 查看

构建#

在客户程序(guest)的目录下构建,除了目标之外不加任何参数:

sh
cd guests/my-app
cargo build --release --target riscv32imac-unknown-none-elf     # .../release/my-app
cargo build --target riscv32imac-unknown-none-elf               # .../debug/my-app

ELF 生成在客户程序工作空间唯一的 target 目录 guests/target/riscv32imac-unknown-none-elf/ 中。guests/.cargo/config.toml 会加上两个你永远不必手动输入的链接器参数:

  • -T crates/guest-sdk/link.ld,即内存布局,它还定义了启动代码和分配器所用的符号。
  • --no-relax。链接器松弛(relaxation)会改写指令序列,并移动其后的每一个地址,而程序身份绑定了这些地址。

没有 runner:远地虚拟机之外没有任何东西能映射客户程序的内存区域,所以 cargo run 无从运行它。客户程序要通过模拟器运行(运行与性能分析)。

构建 profile#

guests/Cargo.toml 把两个 profile 固定为同一语义。它们只在优化级别和依赖的调试断言上有所不同:

dev release
opt-level 0 3
overflow-checks 开启 开启
debug-assertions 开启 客户程序 crate 中开启,其依赖中关闭
panic、codegen-units、debug、incremental abort、1、关闭、关闭 相同

Cargo 默认的 release profile 会关闭溢出检查,而在客户程序中,这不是一个性能设置。它会改变陈述:u32::MAX + 1 会提交 00000000 并以 0 退出,而 dev 构建在同一处会 panic,以 101 退出。所以工作空间在两个 profile 中都保持检查开启。依赖的调试断言检查的是该 crate 自身的不变量,正确的依赖在没有断言时计算结果相同,所以 release 关闭它们,省下这部分周期:相当于无状态以太坊客户程序运行周期的 6.8%。

证明 release 构建。 每条执行过的指令都是一行被证明的数据;opt-level = 3 能去掉客户程序映像的四分之一到一半以上;而且每个电路族的代码都必须装进它的解码表:以太坊客户程序的 debug 映像需要 2^22 行的表,release 映像只需 2^20。你发布的程序身份,是 release 映像的程序身份。

可复现性#

同一台机器上的两次干净构建,会产出完全相同的 ELF。两台机器上的构建一般则不会:ELF 在 panic 位置字符串中嵌入了绝对路径,涉及工具链的 core 源码、crates/guest-sdk 和 cargo registry,而客户程序自己的文件则以相对于 guests/ 的路径出现。在别处构建,得到的是另一个映像,以及另一个程序身份。

所以你登记并交付的是一次构建产出的 ELF,而不是构建方法。保留你证明过的那个 ELF,想核对程序身份的人可以从这个 ELF、参数和仪式文件重新计算它。

导出映像#

sh
cargo run -p artifact-dump -- guests/target/riscv32imac-unknown-none-elf/release/my-app --out artifacts

该命令写出 artifacts/my-app.img,即加载后的 ProgramImage 的序列化格式(postcard,无文件头),以及 artifacts/my-app.img.txt:一份根据经带校验的读取器读回的映像生成的报告。它会打印入口点、段数和指令数,以及制品的大小和 SHA-256。如果读回的结果不一致,或者加载器拒绝了该 ELF,它什么也不写。

.img 是程序的静态描述,用于保存和比对差异;下游没有任何东西需要它,因为设置步骤和各种工具接受的都是 ELF。它的 SHA-256 固定的是字节。它不是程序身份。

阅读报告#

部分 内容
entry and memory 入口,即位于 0x00010000 的 _start;RAM 窗口;slot_base 和槽位跨度
segments 每个段的地址、结束位置、mem_len、文件字节数、零填充和指令数:.text、.rodata(如果有),以及一个延伸到 0x80000000、用于 .data、.bss、堆和栈的可写段
instruction stream 四字节和两字节指令、指令中间的槽位以及 not code 槽位,合计等于槽位总数
symbols 取自 ELF 符号表、按地址排列的名称,制品本身不包含这些名称
listing 逐条指令列出:地址、长度、内存中的字节、展开后的 32 位字、符号

压缩指令保留自己的地址和两个字节;下一个 pc 是 pc + 2 还是 pc + 4,只由 len 表明。要看助记符,使用下面的 tables 视图,或者固定版本的反汇编器:

sh
"$(rustc --print sysroot)"/lib/rustlib/*/bin/llvm-objdump \
    --disassemble --no-print-imm-hex -M no-aliases <elf>

not code 半字#

报告中可能出现 ---- not code: 0x00010f9a .. 0x00010f9c, 1 halfword ---- 这样的行。这是普通的编译器输出:LLVM 证明了某个 match 的默认分支不可达,rustc 为这个不可达块生成了 unimp,在 C 扩展下它就是 c.unimp,即全零半字,RVC 中明确定义为非法的编码。加载器把它记录为非指令,然后继续。没有任何 pc 会到达它;如果有,运行会以 NotAnInstruction 终止。

虚拟机将要证明的内容#

sh
cargo run --release -p artifact-dump -- tables <elf> --ptau assets/ptau/ppot_0080_24.ptau

该命令在默认参数下打印映像推导出的 VmConfig:每个电路族的高度、有效行数和解码列。随后打印每条指令的 pc、next_pc、电路族、助记符和各字段。加上 --ptau 和仪式文件,它还会打印程序身份,即验证者要登记的值。快速上手展示了一个真实的例子。

程序身份是一个域元素。它绑定每条指令及其 pc、长度、操作数和类别(kind),映像中每个来自文件的字节(.text、.rodata、.data),入口点,电路族集合,每个高度,代码大小上限和代码版本。它不绑定符号表、.bss,也不绑定任何由执行过程决定的东西。同一个 ELF 在两种高度设置下有两个程序身份。

推导会在以下情况下拒绝,并指出出错的 pc 或大小:

拒绝 原因
Not all opcodes supported: pc=… 可执行代码中任何位置出现了 RV32IMA 之外的字,例如汇编中的 CSR 访问
TableTooShort 代码超出了电路族解码表的覆盖范围 pc ≤ 2h − 4:高度为 2^20 时可覆盖 1.9375 MiB 代码,2^22 时为 7.9375 MiB
ProgramTooLarge 映像超出了 bytecode_size_words,默认为 4 MiB
ImageOutsideWindow 在所选的窗口高度下,有来自文件的字节位于 RAM 窗口 0 之外
UnknownDelegation 映像声明了一个没有任何电路族响应的委托编号

核对构建#

构建到一个全新的 target 目录,再次导出,然后比较:

sh
cd guests/my-app
CARGO_TARGET_DIR=/tmp/fresh cargo build --release --target riscv32imac-unknown-none-elf
cd ../..
cargo run -p artifact-dump -- /tmp/fresh/riscv32imac-unknown-none-elf/release/my-app --out /tmp/again
cmp artifacts/my-app.img /tmp/again/my-app.img
diff artifacts/my-app.img.txt /tmp/again/my-app.img.txt     # differs only in the `source ELF` line

发射你的应用

运行与性能分析

从 Rust 或命令行在远地虚拟机的模拟器中运行客户程序,与你的宿主机构建比较,并在为证明付出代价之前弄清周期都花在了哪里。

以 Markdown 查看

运行客户程序(guest)几乎没有成本;证明它的成本则与它运行的周期数成正比。所以先运行,与你的宿主机(host)构建比较,并在证明任何东西之前先看一看周期分布。

在 Rust 中:模拟器#

emulator::run 在宿主程序代码中,以给定的公开输入和证明者提示(advice)执行一个已加载的映像,不生成证明:

运行客户程序,并与宿主机构建的结果对比rust
let elf = std::fs::read(elf_path)?;
let image = loader::load_elf(&elf).expect("the ELF loads");
let io = emulator::GuestIo { input: b"hi".to_vec(), advice: Vec::new() };
let run = emulator::run(&image, &io).expect("no fatal error");

assert_eq!(run.exit_code, 0);
assert_eq!(run.io.output, my_app::run(b"hi", &[]).unwrap()); // the host build agrees
println!("{} cycles", run.cycle_count);

run 返回一个 Execution:最终的寄存器、退出状态、周期数和公开值。非零退出状态仍然是一次执行,而不是错误,它通过 exit_code 返回。致命的执行器错误,例如 OutOfBounds、Misaligned 或 NotAnInstruction,以 EmuError 返回,这样的运行没有证明(故障排查)。

模拟器是映像和输入的纯函数:没有时钟、没有随机数、没有线程。相同的输入给出逐周期相同的执行,这也正是证明者能够执行两遍并切出相同分片的原因。

在命令行中:周期分析器#

sh
cargo run --release -p profiler -- elf <elf> [--input <file>] [--advice <file>] [--top <n>] [--json <path>]

它以能容纳客户程序代码的最小表高度,用给定的文件运行客户程序,并打印一份报告。报告中的数字是已执行周期的计数,在任何机器上都相同。

workload
  label                        hello
  guest cycles                 114
  exit status                  0
  journal bytes                13

cycles by semantic workload
  core runtime                             94   82.46%
  unattributed                             20   17.54%

cycles by family
  ADD_SUB_LUI_AUIPC            64
  JUMP_BRANCH_SLT              21
  MEM_WORD                     3
  MEM_SUBWORD                  26

top functions
            94   82.46%          1 calls        94.0 c/call  guest_sdk::commit  [core runtime]
             8    7.02%          1 calls         8.0 c/call  main  [unattributed]

如何阅读这份报告:

  • 按电路族统计的周期(cycles by family)就是你要付费的部分。每个有行的电路族,至少要花费一个其高度的分片;一个电路族中的周期越多,它的分片就越多。
  • 热点函数(top functions)把每个函数自身的周期记在它名下,包括编译器内联进它的一切,但不包括它调用的函数。调用次数在函数的第一条指令处计数。
  • 按语义工作负载统计的周期(cycles by semantic workload)按名称把函数归入十四个类别,例如哈希、签名和核心运行时。未归类部分的占比和助记符分布用来检验这种归类,因为任何符号表都不可能把它们标错。
  • 加速候选项(accelerator candidates)为未来版本可能新增的委托定价,给出的是上限:参见委托。

周期分析器还有两个子命令,用于以太坊工作负载:block <stem> 在一份录制好的测试数据上运行 revm 客户程序,record <number|latest> 从 ETH_RPC_URL 录制一个区块并运行它。

降低成本#

通常最划算的顺序:

  1. 用 --release 构建。 优化能去掉客户程序四分之一到一半以上的指令。
  2. 委托哈希和曲线运算。 使用 guest_sdk::keccak256、sha256、ec_add 以及 vendored 的 k256 和 ark-ff,而不是把软件实现编译进客户程序。
  3. 不要在循环中分配。 每次分配都要花费指令,而在 bump 分配器下,它还是你永远收不回来的内存(堆)。
  4. 检查,而不是计算。 如果一个结果求出来昂贵、验证起来便宜,例如排序顺序、平方根或一条穿过树的路径,就让证明者把它作为证明者提示提供,由客户程序来验证。
  5. 避免浮点运算。 它会编译成软件例程;整数和定点运算要便宜得多。

然后再测一次。周期数精确且可重复,所以每一处改动都会体现为一个数字。

发射你的应用

证明与验证

登记程序、证明一次运行、验证块,并保存证明。高度、同时处理中的分片、证明所需的仪式幂次,以及验证者必须自己持有的两个值。

以 Markdown 查看

三个调用#

设置、证明、验证rust
let params = program::ProgramParams::defaults();
let ptau = std::path::Path::new("assets/ptau/ppot_0080_24.ptau");
let srs = srs::Srs::from_ptau(ptau, 22).expect("ceremony");

let setup = host::setup(&elf, &params, srs).expect("registers");          // once per program
let proven = host::prove(&setup, &io, 4).expect("proves");                // at most 4 shards in flight
host::verify(&setup.vk, &proven.block).expect("verifies");

assert_eq!(setup.vk.identity.to_bytes(), registered); // from your own channel, never the proof
assert_eq!(proven.exit_code, 0);
  • host::setup 加载 ELF,把它解码为各电路族的表和 VmConfig,在仪式之下承诺设置列,并构建验证密钥。它的成本按程序和高度选择计算,而不是按运行计算。
  • host::prove 把客户程序(guest)执行两遍,并证明每个分片(见下文)。它返回一个 Proven:BlockProof、退出码、周期数、公开输出(journal)以及这次运行的报告。
  • host::verify 用块(block)所携带的陈述,对照密钥检查该块。它不会拿程序身份或 SRS 摘要与任何东西比较,所以这项比较要由你来做。

快速上手在一个小型客户程序上运行的正是这段代码,并附有真实输出。

验证者必须自己持有的值#

有两个值必须来自证明者无法控制的渠道:

  1. 程序身份。 如果对照的是证明者提供的身份,证明只能说明某个程序运行过。验证者登记它所信任的发布版本的程序身份,并与密钥中的程序身份比较。
  2. 仪式的 SRS 摘要。 不论密钥自身的点给出什么摘要,密钥都会按这个摘要加载。基于已知 τ 构建的密钥可以打开任何东西,只有把它的摘要与仪式的摘要比较,才能拒绝它。

验证密钥本身可以来自任何人,包括证明者:加载时会根据其自身内容重新计算程序身份和 SRS 摘要,并要求其中的电路与验证者自己的注册表一致。然后读取陈述:先读退出状态,再读公开输出。

高度#

每个电路族都有一个高度,即它一个分片中的行数,从 2^8, 2^12, 2^16, 2^18, 2^20, 2^22 中选取。高度属于程序,而不属于某次运行:每个高度都被绑定进程序身份。

电路族分组 默认值 下限 说明
七个指令电路族 2^22;MUL_DIV 和 ATOMICS 为 2^20 2^20 其时间戳范围检查的下限
INIT_TEARDOWN、ZERO_WINDOWS、ADVICE_WINDOWS 2^22 2^16 共用一个窗口高度;窗口 0 必须容纳映像中每个来自文件的字节
PUBLIC_INPUT、PUBLIC_OUTPUT 2^12 固定 高度决定了它们窗口的位置
委托电路族 见委托 因电路族而异

一个有行的电路族,至少要花费一整个其高度的分片,所以短的运行在较小的高度下浪费更少,长的运行在较大的高度下需要的分片更少。解码表还必须足够高,才能覆盖到该电路族的最后一条指令:2^20 可覆盖 1.9375 MiB 代码,2^22 可覆盖 7.9375 MiB。以太坊客户程序对每个高度可选的电路族都使用 2^20 进行证明。

指令电路族取其下限高度,RAM 窗口取 2^16rust
use constants::family;

let mut params = program::ProgramParams::defaults();
for f in 0..7 {
    params.heights[f] = 1 << 20;
}
for f in [family::INIT_TEARDOWN, family::ZERO_WINDOWS, family::ADVICE_WINDOWS] {
    params.heights[f as usize] = 1 << 16; // window 0 is then 256 KiB: the image must fit in it
}

仪式提供的幂次数必须不少于最高的电路族的行数,并且至少为 2^18,以满足通用查找(lookup)表的需要:调用 Srs::from_ptau(path, k) 时,2^k 至少要等于最大的高度。

同时处理中的分片#

host::prove 的第三个参数是 max_in_flight,即同时证明的分片数。它是唯一一个在内存与时间之间权衡的旋钮:

  • 内存取决于同时处理中的分片,而不是周期数。每个同时处理中的分片都持有自己的行、前向过程,以及不断增长的证明。最宽的指令电路族的一个 2^20 分片,在前向过程中约占 8.4 GiB;一个 2^18 的 KECCAK_F 分片约占 42 GiB。
  • 时间取决于有多少分片并行运行,最多到你拥有的核心数为止。在一个分片内部,工作在所有核心上运行。
  • 证明与它无关。 同时处理 1 个和 8 个分片时,得到的块逐字节相同。

在笔记本电脑上从较小的值开始,两个或四个;在服务器上逐步调高,直到限制你的是内存而不是核心数。bench prove 的默认值是 8。

两遍执行#

host::prove 以流式方式工作。它从不持有整个执行轨迹;按每个周期约 300 字节计算,执行轨迹会是系统中最大的对象。

  1. 第一遍执行客户程序,每填满一个分片,就承诺它的内存列,保留承诺,丢弃行数据。到退出时,它构建陈述,并抽取所有分片共用的挑战。
  2. 第二遍再次执行。模拟器是确定性的,所以切出的分片完全相同。每个分片一到达就被填充、证明并丢弃,只保留它的证明。

这就是为什么证明要花两次执行的时间,而内存以同时处理中的分片为界。流式证明者对此有深入的解释。

保存证明#

host::proof_archive::write_proof(dir, stem, vk, block) 写出四个文件,每个文件都是对应类型的原始字节:

<stem>.vk         the verifying key
<stem>.identity   the key's identity, 64 lowercase hex digits and a newline
<stem>.public     the statement: input, journal, exit status, the execution's record
<stem>.block      the block proof

read_proof(dir, stem) 把它们读回来。.identity 文件记录的是这次运行所声称的身份;验证者仍要与自己的副本比较。verifier 命令行工具可以检查一份归档:

sh
cargo run --release -p verifier -- block <stem>.vk <identity-hex> <stem>.public <stem>.block

全部验证通过时它以 0 退出;以 1 退出时会指出第一处拒绝;用法错误或程序身份格式错误时以 2 退出。它把你传入的程序身份与密钥中的比较,SRS 摘要则取自密钥文件。

当证明失败时#

诚实的证明者永远不会生成无法通过验证的证明,所以一旦失败,要么是它接受了不该接受的输入,要么是存在 bug。打开证明者的调试日志重新构建,然后重新运行:

sh
cargo run --release -p bench --features prover/debug-info -- prove ...
APOGEE_DEBUG=detail <the same run> 2>&1 | tee run.log
grep -E 'FAIL|NOT CANONICAL|UNBALANCED|OVER the|NAMES NO|DISAGREES|NOT LOOPING|ABORTED' run.log

日志会指出失败的分片;self_check FAILED 会指出某一行违反的第一个门,并给出每个操作数的值。工具列出了所有标记。日志只存在于启用了 debug-info 特性的构建中,不会改变证明的任何字节。

下一步:链上结算。

发射你的应用

链上结算

从由数百个分片组成的基础证明,到一份由以太坊合约检验的 Groth16 证明。递归树、判定器的仪式、合约的接口,以及一次部署所固定的内容。

以 Markdown 查看

基础证明是由分片证明组成的一个块(block),每个分片证明都是一份 GKR 证明及其承诺:数兆字节的数据和数百个曲线点,任何合约都无法检验。结算分三个阶段压缩它,每个阶段都在仓库根目录下用 bench 运行。

从基础证明到合约 基础分片由叶节点验证,叶节点由内部节点验证,内部节点由根节点验证;Groth16 判定器重新验证根节点;合约检验 Groth16 证明和折叠后的配对。 基础证明 207 个分片 · 14.5 MB 叶节点 每个 ≤ 64 个基础分片 根节点 2–4 个子节点 覆盖 0..count 判定器 Groth16 BN254 7.9M 个约束 合约 verify(…) → true 3.62M gas
结算。每个阶段都验证它的前一个阶段。在合约之前不做任何配对:每个 Mercury 检查都被延迟,并折叠进一个累加器,由合约用两次配对兑现。图中数字来自第 257,510 号区块。

1. 递归树#

节点就是远地虚拟机在证明一个验证者程序。叶节点验证一段连续的基础分片;内部节点验证二到四个子证明;根节点覆盖全部基础分片。每个节点还会把它的分片和子节点所延迟的每个 Mercury 检查折叠成一对点,因此整棵树最终在顶端归结为单个配对断言。

sh
# the base proof as an archive: host::proof_archive::write_proof from your host
# program, or `bench prove ... --out <dir>` for the Ethereum guests
cargo run --release -p bench -- recurse <dir>/<stem> --out <out> --in-flight 4

recurse 写出两个递归程序的密钥,在其上构建叶程序和节点程序的二进制文件,在证明任何东西之前先确定计划(<out>/tree.txt:叶节点至多包含 --leaf 64 个基础分片,其上的节点至多有 --fan-in 4 个子节点),然后逐个节点进行证明。每个节点是一个独立的进程,在证明之前先原生验证其输入,所以错误的输入会被指名拒绝。中断的运行可以恢复:已在 <out> 中的证明会被保留,而计划或程序不同的运行会被拒绝。

基础证明完全不受这些影响。叶节点按基础分片的原样验证它们。

2. 判定器#

根节点仍然是一份 GKR 证明外加数百个点。判定器是一个 Groth16 电路,它像节点那样验证根节点,但不做任何折叠。取而代之,它把以下各项绑定为由合约提供取值的线:两个递归程序的程序身份、基础陈述的退出状态、逐字节的公开输入和公开输出(journal),以及根节点欠下一次配对的每个点及其标量。Groth16 证明携带对所有这些线的一个承诺,合约用自己持有的值来检查它。

Groth16 密钥需要一场仪式。第一阶段就是递归树的承诺所依据的同一个 powers-of-tau 文件。第二阶段专属于这个电路,分两轮贡献进行:

sh
cargo run --release -p bench -- ceremony <out> init          # once per root shape
cargo run --release -p bench -- ceremony <out> contribute    # round 1: alpha and beta, each contributor in turn
cargo run --release -p bench -- ceremony <out> seal
cargo run --release -p bench -- ceremony <out> contribute    # round 2: gamma, delta and eta
cargo run --release -p bench -- ceremony <out> key
cargo run --release -p bench -- decide <out>                 # the Groth16 proof, checked natively and in an EVM

每一份贡献都把某个陷门乘以一个只有该贡献者知道的因子,并用一个 Schnorr 证明记录下来,所以任何状态都只需对照电路和仪式文件就能验证。只要某个陷门的贡献者中有一位是诚实的,这个陷门就无人知晓。各轮的顺序是可靠性的一部分:在任何东西被 delta 或 eta 除之前,alpha 和 beta 就已完成。

注意

bench decide --dev-key 从一个公开的种子推导出所有陷门,用于开发和测试。在它之下任何人都能伪造证明;它的输出以 development.* 的名字写出,以免被误认为来自仪式。在一台机器上跑完的仪式也不算仪式:每一轮都需要一位诚实的贡献者。

decide 写出 decision.constructor 和 decision.calldata:以十六进制表示的部署参数和调用。

3. 合约#

contracts/ApogeeVerifier.sol 只有一个入口:

solidity
function verify(
    bytes calldata input,        // the base program's public input
    bytes calldata output,       // its journal
    uint256 exitStatus,          // the status you require, normally 0
    uint256[10] calldata proof,  // Groth16 A, B, C and the bound wires' commitment D
    uint256[] calldata points    // x, y and scalar of each point, side [1]_2's then side [x]_2's
) external view returns (bool);

它根据 calldata 重建被绑定的值,检查 Groth16 配对等式,用 ecMul 和 ecAdd 折叠每一侧的点(这同时要求每个点都在曲线上),然后检查折叠后的断言 e(A, [1]_2) = e(B, [x]_2)。应用合约调用它,然后根据公开输出采取行动:参见账本草图。

一次部署固定了什么#

构造函数接收 Groth16 密钥、仪式的两个 G2 点、叶程序和节点程序的程序身份、每一侧的点数,以及公开输入和公开输出的字节长度。所以一个已部署的验证者服务于:

  • 一个基础程序。 它的程序身份是叶程序映像中的一个常量,而叶程序的程序身份绑定了这个映像。
  • 一种根的形状。 判定器的电路取决于根节点的程序、它的分片数和公开值的长度,所以密钥及其仪式都是按形状区分的。
  • 定长的公开值。 verify 会拒绝任何其他长度的输入或公开输出。要把客户程序(guest)设计成链上公开值具有固定大小,例如一条定长记录或一个 32 字节的摘要。

合约为每个点支付约 9,000 gas,因为电路一个点也不折叠。

实测数据#

第 257,510 号区块;递归树在一台 32 CPU、247 GiB 内存的机器上运行,仪式和判定器在一台 18 核笔记本电脑上运行:

阶段 结果
基础证明 207 个分片,14.5 MB,2,481 s
递归树 4 个叶节点(每个至多 64 个基础分片)和一个根节点:共 116 个分片
叶节点,四个同时进行 21、24、23 和 27 个分片;2,157 s;峰值 92 GiB
根节点,同时处理四个分片 21 个分片,460 s,1.03 MB
判定器电路 7,896,686 个约束,定义域大小为 2^23
仪式 init 65 s;每份贡献 50–56 s;key 70 s、12.7 GB;密钥 2.65 GB
判定器证明 读取密钥 1 s,证明 18.5 s,6.1 GB
合约 358 个点;3,620,026 gas;34,980 字节 calldata

这一切的规范见递归与判定器。

发射你的应用

客户程序编程指南

让客户程序保持正确、可证明且低成本的习惯。所有易错点和推荐做法集中在一处,每一条都附有背后的理由和应当采取的做法。

以 Markdown 查看

编写客户程序(guest),大部分时候就是在写 Rust。本页讲的是其余部分:一台裸机、单 hart、需要被证明的机器,在哪些地方与你习惯的宿主机(host)表现不同。每条规则都说明该怎么做、为什么,以及不这样做会出什么问题。AI 随行手册以一种可以直接交给模型的形式,收录了同样的规则。

类型与内存#

usize 是 32 位,所有指针也是#

客户程序的目标是 riscv32imac:usize、isize 和所有指针都是 32 位宽,而在你的宿主机上它们是 64 位。

  • usize 溢出在客户程序上会 panic,在宿主机上则不会。
  • 在客户程序上,从 u64 做 x as usize 会静默截断。
  • 任何包含长度或指针的东西,其 size_of::<T>()、结构体布局和 core::hash 在两种构建之间都不相同。

应当:在你要提交、哈希、序列化或与宿主机计算结果比较的任何东西中,使用显式的 u32 和 u64。在值可能装不下的地方用 usize::try_from(x) 转换,让它在两种构建上都明确地报错。不要:提交 usize,对包含它的结构体求哈希,或者派生依赖内存布局的编码。

rust
let n = u64::from_le_bytes(input[..8].try_into().unwrap());
let len = usize::try_from(n).expect("length fits the guest"); // not `n as usize`

分配器从不释放内存#

堆是一个 bump 分配器:alloc 把一个指针向上移动,dealloc 什么也不做,内存只有在程序退出时才会归还。所以,让客户程序耗尽内存的,是它在整个运行期间分配的总量,而不是峰值。 如果一次分配的末端会高于栈的预留区,或高于当前的栈指针,客户程序就以状态 71 退出。

应当:

  • 分配一次,反复复用:把缓冲区提到循环之外,用 clear() 清空它们,而不是构建新的。
  • 用 Vec::with_capacity、String::with_capacity 预先设定集合的容量,让它们在增长时不必重新分配和复制。一个靠逐次 push 增长到 n 个元素的 Vec,还会把它先前那些较小的缓冲区遗留在堆上。
  • 优先借用(&[u8]、&str)而不是克隆,优先使用迭代器而不是中间集合。
  • 就地处理大块的证明者提示(advice):advice() 本来就是内存上的切片,无需复制。

不要:在热循环里 collect() 出一个新的 Vec;克隆只读的值;或者在一个映射表清空后就能重新填充的情况下,为每个请求重建映射表。

rust
// Total heap grows with the number of requests:
for req in requests {
    let parts: Vec<u32> = req.chunks(4).map(|c| u32::from_le_bytes(c.try_into().unwrap())).collect();
    process(&parts);
}

// Total heap is one buffer:
let mut parts: Vec<u32> = Vec::with_capacity(MAX_PARTS);
for req in requests {
    parts.clear();
    parts.extend(req.chunks(4).map(|c| u32::from_le_bytes(c.try_into().unwrap())));
    process(&parts);
}

栈有 8 MiB,远端没有任何防护#

栈从 0x8000_0000 向下增长,有一个任何堆块都不得进入的 8 MiB 预留区。在预留区内深度递归没有问题。没有任何机制能察觉的情况是:堆已经填满了预留区下方的空间,栈却越过了预留区;这时堆块会在深层调用链之下被悄无声息地改写。应当:让递归深度有界且可预测,或者用显式的工作列表把深度遍历改写成迭代。不要:递归到一个由不可信输入决定的深度。

只允许对齐访问#

通过未对齐指针进行半字或字访问是致命的,绝不会被拆分处理,这次运行也没有证明。安全的 Rust 永远不会产生这种访问。不要把字节指针强制转换为 *const u32 再解引用;用 u32::from_le_bytes 读取(它会编译成字节加载),或者用 ptr::read_unaligned。

空指针指向空洞#

0x8000 以下的地址不属于任何东西,所以空指针或数值很小的野指针会导致致命的 OutOfBounds,而不是读到垃圾数据。它表现为一次没有证明的运行,而绝不会表现为一个错误的答案。

并发#

原子操作:受支持,但不用于新的客户程序代码#

重要

A 扩展得到完整支持:lr.w、sc.w 和全部九条 AMO 指令都能通过它们自己的电路族解码、执行和证明,core::sync::atomic 也会编译成这些指令。尽管如此,仍强烈不建议用原子操作编写客户程序。 远地虚拟机在单个 hart 上执行,没有中断,也没有线程,所以没有任何需要同步的东西。提供原子操作,是为了让已经使用它们的现有代码(带原子计数器的库,或者 spin 锁)不加修改就能编译和证明。它们是一条兼容路径,而不是一种编程实践。

如果原子操作经由依赖进入了你的客户程序,你需要知道:

  • 在单个 hart 上,原子操作只是一次读-改-写。 fetch_add 就是一条做加法的 amoadd.w,没有任何东西能与它交错执行。
  • sc.w 总是成功。 这台机器不保存保留(reservation)状态,所以条件存储(store-conditional)总会执行存储,并向 rd 写入 0。编译器为 compare_exchange 生成的 lr.w/sc.w 重试循环不受影响,因为在任何 hart 上首次尝试就成功都是合法的。依赖 sc.w 在没有有效保留时失败的代码,在这里得不到这种失败。这是远地虚拟机唯一偏离 RV32IMAC 的地方。
  • fence 什么也不做,内存序(aq、rl、SeqCst)在单个 hart 上也不对任何东西排序。
  • 它们会多花一个电路族。 原子操作会把 ATOMICS 电路族加入程序,于是程序至少要证明它的一个分片。

应当:在新的客户程序代码中,用普通变量、Cell 和 RefCell 保存状态。不要:在一个根本没有第二个线程可共享的客户程序中加入 AtomicU32、类似 Mutex 的自旋锁或 Arc。

输入与输出#

先检查证明者提示,再让由它推导的内容进入公开输出#

证明者提示是证明者填写的内存,没有任何东西绑定它。在提交任何依赖于它的内容之前,先把它与证明确实绑定的某样东西核对,例如公开输入中的哈希或 Merkle 根、一个签名,或者结果的某种性质。提交未经检查的证明者提示的函数值,就等于发布一个由证明者选定的值。参见这一模式。

公开值要小,用于链上时还要定长#

输入和公开输出(journal)各自最多容纳 16,380 字节。commit 宁可以 70 退出,也不会溢出。大的输入应放进证明者提示,并以承诺加以约束;会增长的输出则应以摘要发布。验证者合约是针对一个输入长度和一个公开输出长度构建的,所以要在以太坊上结算的客户程序应当发布定长的公开输出。

想清楚失败是什么样子#

以非零状态退出或发生 panic 的客户程序,对它所做的事依然有一份有效的证明,而验证者先读退出状态,再读公开输出。为每一种拒绝分配一个专属的退出码,并避开 SDK 使用的那些(70、71、72 和 101);在可能拒绝某份未经检查的数据的检查完成之前,不要提交任何由这份数据推导出的内容。

没有外部世界#

客户程序没有时钟、没有随机数、没有网络、没有文件,也没有环境变量。向宿主程序索取其中任何一样的库调用,都会得到 -ENOSYS,并使这次运行无法被证明。给 HashMap 一个确定性的种子,或者使用 BTreeMap;算法需要随机性时,从输入中推导;时间则作为输入传进来。

成本#

每条执行过的指令都是一行被证明的数据#

证明成本取决于周期数,按电路族分别计算。用 --release 构建,用周期分析器测量,像嵌入式程序员对待字节那样对待周期。

有电路的运算,交给委托#

keccak256、sha256、椭圆曲线加法和乘法、Poseidon2、BN254 域运算以及 256 位模乘都有专用电路。通过 guest_sdk 以及 vendored 的 k256、ark-ff 和 revm-precompile 使用它们,而不是把软件实现编译进客户程序。参见委托。

验证,而不是计算#

当一个结果求解昂贵、检查便宜时,让证明者去求解并作为证明者提示传入,再由客户程序检查:排序顺序、因数分解、逆元、穿过树的路径、搜索结果。

避免浮点运算#

目标没有 F 或 D 扩展,所以 f32 和 f64 会编译成整数软件例程。它们正确且确定,但每次运算都要花费许多条指令。使用整数或定点数。

高度按整个分片计价#

一个有行的电路族,不论填了多少行,至少都要花费一个其高度的分片。只用到某个电路族一次的程序,也要为一整个分片付费;你的代码用到的电路族以及你选择的高度,决定了每份证明的成本下限。参见高度。

代码与程序身份#

指令流就是映像#

代码是静态的:每个 pc 处的指令都取自加载时构建的解码表,从不取自 RAM。向 .text 写入改变的是数据,而不是行为;跳转到没有指令的半字会使运行终止且不产生证明。没有 JIT,也没有自修改代码。

任何位置的一个非法字都会让整个程序被拒绝#

解码器会处理整个 .text,不论是否可达。内联汇编中的 CSR 访问、fence.i、浮点或 RV64 编码,或者被汇编进 .text 的数据,都会让推导拒绝整个程序。ebreak 可以解码,但没有证明。

溢出检查是程序的一部分#

客户程序的 profile 在 release 中保持 overflow-checks 开启,因为关掉它会改变客户程序计算的内容:u32::MAX + 1 会回绕并以 0 退出,而不是 panic。在确实需要的地方使用 wrapping_*、checked_* 和 saturating_*。

一次构建就是一个程序身份#

程序身份绑定代码和数据的每个字节、入口点和每个高度。在另一台机器上重新构建会得到另一个程序身份,因为 ELF 嵌入了绝对路径。登记并交付你证明过的那个 ELF,而不是生成它的命令。

检查清单#

在证明之前:

  • 为 riscv32imac-unknown-none-elf 执行 cargo build --release,并已审阅周期分析器的周期报告
  • 提交、哈希或序列化的任何东西中都没有 usize
  • 热循环中没有分配;会增长的集合已用 with_capacity 设定容量
  • 你自己的客户程序代码中没有原子操作、锁或 Arc
  • 每一处对证明者提示的使用,都在任何依赖它的提交之前,与证明所绑定的某样东西核对过
  • 公开输出有上界;若要在链上结算,则为定长
  • 哈希和曲线运算都经由委托完成
  • 每种拒绝都有各自不同的退出码
  • 对于你的测试输入,宿主机构建与模拟器得出的公开输出一致
  • 已从你将要交付的 ELF 记录下程序身份

发射你的应用

故障排查

按症状列出客户程序未能得到通过验证的证明的每一种情况。退出状态、致命的执行器错误、被拒绝的 ELF、被拒绝的程序、失败的证明和验证者错误,以及各自的原因和修复方法。

以 Markdown 查看

客户程序(guest)可能在六个环节止步,得不到通过验证的证明。先找到症状,再找到对应的行。

运行以意料之外的状态退出#

运行已经结束,并且可以被证明;是客户程序自己选择了失败。SDK 自身使用的状态:

状态 原因 怎么办
70 commit 将超过 16,380 字节 提交输出的摘要,而不是输出本身
71 某次分配的末端将高于 __stack_top − 8 MiB 或当前的 sp 内存从不释放,所以这次运行的总分配量必须能放进映像与 0x7F80_0000 之间。在各轮循环之间复用缓冲区,并用 with_capacity 设定容量(堆)
72 某个委托返回了其 shim 拒绝接受的结果:一个错误,或者在多次调用操作的第一次调用之后返回的 -ENOSYS 使用 SDK 的函数而不是原始帧,并让操作数保持小于其模数
101 panic,不打印任何内容 用你的库的宿主机(host)构建运行同样的输入,panic 信息会在那里打印出来(在宿主机上测试)

其他任何状态都来自你自己的 exit(code)。

运行因致命错误而停止#

模拟器返回一个 EmuError,既没有退出状态,也没有证明:

错误 常见原因
OutOfBounds 空指针或野指针([0, 0x8000) 是空洞)、读取证明者提示(advice)时超出了宿主程序提供的范围、在没有证明者提示的运行中调用 advice(),或者委托帧没有完全位于 RAM 中
Misaligned 通过未对齐指针进行的半字或字访问,或者未对齐的委托帧
NotAnInstruction 跳转到一个没有指令的 pc,包括全零的 c.unimp 半字
IllegalInstruction 机器不执行的编码
Ebreak 一条 ebreak,它没有证明
ClockOverflow 超过 2^36 − 1 个周期
PublicInputTooLong、JournalTooLong 输入,或者退出时公开输出(journal)的长度字,超过 16,380 字节
DelegationFrame 其电路无法为之提供见证的帧:MOD_MUL 或 EC_ADD 的操作数大于或等于其模数、选择子不指向任何东西、keccak 轮次超过 23、SHA-256 轮组编号超过 15、Poseidon2 的某个 lane 大于或等于 p
DelegationFamilyAbsent 映像从未声明过的委托编号,出现在追踪路径上

ELF 被拒绝#

对于加载器无法接受的 ELF,artifact-dump、host::setup 和各种工具都会以 LoaderError 拒绝:

拒绝 常见原因
NotAnElf、Truncated 不是客户程序的 ELF:例如 .d 文件,或者没有写完整的文件
NotRiscV、UnsupportedElfType、RelocatableElf、DynamicElf 宿主机构建、目标文件、PIE,或者动态链接的构建
BadSegment、NoExecutableSegment、EntryNotAnInstruction 被修改过的 link.ld,或者没有链接 _start
RvcIllegal、InstructionTooLong、TextTruncated .text 中有数据,例如手写汇编中的表。全零半字绝不会导致这些错误,它是预期之中的

程序无法登记#

ELF 能够加载,但程序无法被解码为一个配置:

拒绝 原因与修复
Not all opcodes supported: pc=… 可执行代码中任何位置(无论是否可达)出现了 RV32IMA 之外的字,例如汇编中的 CSR 访问或 fence.i。删除它
TableTooShort 代码超出了电路族解码表的覆盖范围。提高该电路族的高度:2^22 可覆盖 7.9375 MiB 代码
ProgramTooLarge 映像超出了 bytecode_size_words,默认为 4 MiB。在 ProgramParams 中提高这一上限
ImageOutsideWindow 在你的窗口高度下,有来自文件的字节位于 RAM 窗口 0 之外。提高窗口高度
HeightNotOnMenu 高度不是 2^8、2^12、2^16、2^18、2^20 或 2^22 之一
UnknownDelegation 映像声明了一个没有任何电路族响应的委托编号

如果程序用到的某个电路族被设置在其下限以下(指令电路族为 2^20,RAM 窗口电路族为 2^16),密钥也会构建失败,因为那里不存在电路。

证明者失败#

诚实的证明者在证明模拟器执行过的内容时不会失败,所以失败意味着它接受了不该接受的输入,或者存在 bug。大多数情况属于以下两种:

  • 没有证明的调用。 某个库为获取宿主机数据而发出的、EXIT 和委托之外的系统调用,会得到 -ENOSYS,运行继续进行,但证明者在填充时会拒绝那一行,并指出对应的周期。找出那个索取随机数或时间的依赖。
  • 内存不足。 进程在分片处理过程中被杀死。调低 host::prove 的第三个参数,或者降低高度。

其他情况,打开证明者的调试日志重新构建,然后重新运行:

sh
cargo run --release -p bench --features prover/debug-info -- prove ...
APOGEE_DEBUG=detail <the run> 2>&1 | tee run.log
grep -c 'begin h=' run.log; grep -c 'gkr done' run.log     # unequal: a shard died
grep 'begin h=' run.log | tail -1                           # which one
grep -E 'FAIL|NOT CANONICAL|UNBALANCED|OVER the|NAMES NO|DISAGREES|NOT LOOPING|ABORTED' run.log

self_check FAILED 会指出该行第一个不满足的门,以及每个操作数的值。调试日志列出了所有标记。

验证失败#

verify_shard 和 verify_block 返回一个 VerifyError,它的类别说明了哪里出了问题:

类别 含义
Statement 陈述与密钥不符:分片数、窗口规则、载荷长度、列表,或者为证明提供种子的全局摘要
Malformed 证明的形状与其电路不符:承诺、输出、轮次或断言的数量
Constraint { layer } 某个门被违反,或者某一层的求和校验(sumcheck)失败
Lookup { channel } 查找(lookup)的元组不在其表的任何一行中
MemoryArgument 读多重集与写多重集无法核对一致,或者某个公开窗口存放的不是陈述中的字节
Opening 某个承诺的打开失败

如果证明是由远地虚拟机自己的证明者、根据模拟器接受的一次执行生成的,那么验证失败意味着验证者和证明者对程序的认识不一致:检查你加载的是否就是生成该证明时所用的密钥,高度是否相同,仪式是否相同。

程序身份不符#

你计算出的程序身份与预期的不同:

  • 另一台机器上的构建。 ELF 嵌入了绝对路径,所以在别处重新构建得到的是另一个映像。要与被证明的那个 ELF 比较,而不是与一次新的构建比较。
  • 不同的高度。 每个高度都被绑定进程序身份。artifact-dump tables 报告的是默认高度下的程序身份;你的设置可能用了别的高度。
  • 另一场仪式。 Hermez 的 powers-of-tau 文件对应的是另一个 τ,所以每个承诺都不同。把文件的 [τ]_1 与该仪式的值核对。

发射你的应用

示例客户程序

代码仓库中的客户程序,每一个都是客户程序 SDK 或机器某一部分的完整示例。需要某种模式时,该去哪里找。

以 Markdown 查看

guests/ 工作空间包含代码仓库构建和测试的所有客户程序(guest)。每一个都是为了检验某样东西而存在的,这使它们成为你即将编写的模式的最佳参考。它们都可以在各自的目录下用 cargo build --target riscv32imac-unknown-none-elf 构建。

从这里开始#

客户程序 展示的内容
public-io 同时使用三个内存区域:在提交任何东西之前,先对照公开输入检查证明者提示(advice)。用一个小程序展示整个 I/O 模型
fib 最小的 SDK 客户程序:输入一个 u32,输出一个 u32,使用回绕算术
echo、heap 分配器:证明者提示经由堆缓冲区复制,Vec 和 Box 在 bump 分配器中反复分配和丢弃

应用模式#

客户程序 展示的内容
amm、orderbook 不使用堆的 128 位和 256 位整数;BTreeMap、排序,以及作为证明者提示提供、经检查而非计算得到的排序顺序
vault、recursion-ops 在客户程序中使用 crates/field 和 crates/transcript,无需指名任何 shim 就委托给 FR_ARITH 和 POSEIDON2
revm-block 基于 revm 的以太坊区块:二进制程序 revm-block(一个录制好的 mini-block)和 revm-block-stateless(无状态校验器)

委托#

客户程序 展示的内容
keccak-test、sha256-ops、mod-mul-ops、ec-ops KECCAK_F、SHA256_COMP、MOD_MUL 和 EC_ADD,每一个都在客户程序内部对照独立的值检查
keccak-unused、recursion-unused 链接了但从未调用的 shim:这些电路族已声明,证明零个分片

机器本身#

客户程序 展示的内容
atomics core::sync::atomic 生成的每一条 A 扩展指令。只有一个 hart,意味着每一条都只是普通的读-改-写;这个客户程序是为了测试该电路族而存在的,而不是为了推荐这种做法
opcodes 每一条 RV32IMAC 指令
rvc-dense 压缩指令展开:同一段指令序列分别以压缩和非压缩方式汇编
addsub、control、alu、mem、shards 手写汇编,有自己的 _start,不使用 SDK,以其结果退出。shards 会填满两个 2^20 分片
recursion 递归树的验证者程序,二进制程序 leaf 和 node

最小的实用客户程序#

fib 读取一个 u32,用回绕算术执行同样多步的斐波那契迭代,然后提交结果:

guests/fib/src/main.rsrust
#![no_std]
#![no_main]

guest_sdk::entry!(main);

fn main() {
    let mut n = [0u8; 4];
    assert_eq!(
        guest_sdk::read_input(&mut n),
        4,
        "fib: the public input is one u32"
    );
    let n = u32::from_le_bytes(n);

    let mut a: u32 = 0;
    let mut b: u32 = 1;
    for _ in 0..n {
        let next = a.wrapping_add(b);
        a = b;
        b = next;
    }
    guest_sdk::commit(&a.to_le_bytes());
}

输入过短是错误,而不是默认值:在一个只填了一部分的缓冲区上继续执行的客户程序,证明的是一个关于零的陈述。加法有意采用回绕,所以大于 47 的 n(超出了能放进 32 位的最后一项)是一个有普通答案的普通输入,而不是一次失败的运行。这里也没有证明者提示,因为对验证者来说,检查 f_n 和计算它的成本一样高:通过证明者提示给出的答案,必须重新计算一遍才能让人相信。

发射你的应用

客户程序 SDK 参考

guest-sdk crate 的每一个公开项,附确切的签名与行为。只有 exit 和委托 shim 会发出 ecall;其余一切都是加载和存储。

以 Markdown 查看

crates/guest-sdk 是客户程序(guest)的全部运行时:启动代码、入口宏、分配器、panic 处理程序和 ecall shim。它只为 riscv32imac-unknown-none-elf 编译。

入口#

rust
guest_sdk::entry!(main);

导出启动代码所调用的 main 符号,它是调用你的函数的包装;你的函数不带参数,返回 ()。你的函数保留自己的名字,它本身也可以叫 main。从它返回即为 exit(0)。

内存区域#

项 签名 行为
public_input fn public_input() -> &'static [u8] 公开输入的载荷,其长度字被限制在窗口范围内。不复制,不发出 ecall
read_input fn read_input(buf: &mut [u8]) -> usize 复制 min(buf.len(), public_input().len()) 个字节并返回实际数量。返回值可能小于请求的长度
advice fn advice() -> &'static [u8] 证明者提示(advice)的载荷,其长度被限制在区域范围内。不受任何绑定,所以由客户程序检查它。在没有提供证明者提示的运行中会导致致命的 OutOfBounds
commit fn commit(bytes: &[u8]) 追加到公开输出(journal)并更新其长度字。宁可以 70 退出,也不会溢出 16,380 字节的窗口
journal fn journal() -> &'static [u8] 到目前为止提交的全部内容
exit fn exit(code: i32) -> ! 结束运行,以 code 作为陈述的退出状态。除已提交的内容外不发布任何东西

哈希#

项 签名 行为
keccak256 fn keccak256(input: &[u8]) -> [u8; 32] 以太坊的 Keccak-256,而不是 SHA3-256。海绵结构和填充在客户程序代码中运行;每一轮 keccak-f[1600] 是一次 KECCAK_F 调用。如果第一次调用返回 -ENOSYS,则回退到软件实现
sha256 fn sha256(input: &[u8]) -> [u8; 32] FIPS 180-4 SHA-256。填充和分块循环在客户程序代码中运行;每次压缩是十六次 SHA256_COMP 调用。软件回退同上
poseidon2_permute fn poseidon2_permute(state: &mut [u8; 96]) -> bool 宽度为 3 的 Poseidon2 置换,作用于三个规范的小端序 Fr lane,原地进行,经由 POSEIDON2。遇到 -ENOSYS 时返回 false,由调用方走自己的软件路径

椭圆曲线#

rust
pub type ProjectivePoint = [[u32; 8]; 3];

齐次射影坐标下的一个点,x = X/Z,y = Y/Z,每个坐标是八个小端序的 32 位 limb,数值小于曲线的域模数。它不是雅可比(Jacobian)坐标:arkworks 的 Projective 才是,所以从它转换的调用方,传入时映射为 (X·Z, Y·Z², Z),传出时映射为 (X·Z, Y, Z³)。单位元是 (0 : 1 : 0)。

项 签名 行为
ec_add fn ec_add(codes: &[u32; 3], p: &ProjectivePoint, q: &ProjectivePoint) -> Option<ProjectivePoint> 用完全加法公式计算 p + q,按分组顺序发出三次 EC_ADD 调用。遇到 -ENOSYS 时返回 None
ec_mul fn ec_mul(codes: &[u32; 3], p: &ProjectivePoint, k: &[u32; 8]) -> Option<ProjectivePoint> 从最高位开始用倍加法(double-and-add)计算 k·p。k 按原样使用;把它对群阶取模是调用方的事
ec_identity fn ec_identity() -> ProjectivePoint (0 : 1 : 0)
recursion::SECP256K1_GROUPS、recursion::BN254_GROUPS [u32; 3] 即 codes 参数:指定哪条曲线,形式是一次点加的三个分组选择子

公式证明的是算术,而不是点是否在曲线上:从证明者提示中取得的点,要自己检查。

原始委托 shim#

guest_sdk::recursion 包含作用于按字对齐的帧类型的各个 shim。每个帧类型都是 #[repr(C, align(4))],所以它的对齐由类型决定,而不取决于代码生成器把局部变量放在了哪里。基础格式的 shim 恰好在 -ENOSYS 时返回 false;任何其他非零应答都以 72 退出。

项 用途
mod_mul(&mut ModMulFrame) -> bool 一次 a·b mod m。用 ModMulFrame::of(modulus, &a, &b) 构建帧,读取 frame.result()。模数代码为 SECP256K1_P、SECP256K1_N、BN254_P 和 BN254_R,两个操作数都必须已经小于模数
sha256_comp(&mut Sha256Frame) -> bool 一次完整的压缩:按顺序发出十六次调用。先 Sha256Frame::of(&state, &block),再 frame.working();把结果加到链值上由调用方负责
ec_add_complete(&mut EcAddFrame, &[u32; 3]) -> bool 一次完全点加:按分组顺序发出三次调用。先 EcAddFrame::of(&codes, &p, &q),再 frame.result()
poseidon2(&mut Poseidon2Frame) -> bool、fr_arith(&mut FrArithFrame) -> bool 作用于字节帧的置换和一次 Fr 运算;field 和 transcript 会替你调用它们
sha256_rounds、ec_add 上述操作的单个步骤。顺序错误的步骤不会被拒绝,只会算出别的东西,所以优先使用完成整个操作的函数
fr_op、p2_field、field_io、fq_op、import、import_run、replay 递归格式的协处理器调用,供递归树自己的程序使用。它们没有软件路径

每个 shim 从其电路族的声明记录中读取自己的 ecall 编号;声明记录是一个 12 字节的 static,位于它自己的链接器段中。链接一个 shim 就声明了相应的电路族;已声明但从未调用的电路族证明零个分片。

运行时行为#

组成部分 行为
启动 位于 0x0001_0000 的 _start 让 sp 指向 __stack_top(0x8000_0000),逐字节将 .bss 清零,调用 main,若其返回则以 0 退出
分配器 从 __heap_start 向上推进,从不释放。当某个块的末端会高于 __stack_top − 8 MiB 或当前的 sp 时,以 71 退出
panic 处理程序 以 101 退出,不写出任何内容。发生 panic 的客户程序可以被证明,并且已经发布了它提交过的内容
退出状态 70 公开输出溢出,71 堆耗尽,72 委托返回了错误,101 panic

透明委托#

代码仓库中有两个库 crate 在客户程序目标上进行委托而无需显式指名 SDK,靠的是只在该目标上生效的对 SDK 的依赖:

  • field::Fr:加法、Montgomery 乘法(*、square、pow 以及各种转换)和非零元素的 inverse 都调用 FR_ARITH。使用 Fr 算术的客户程序会声明该电路族。
  • transcript::poseidon2_permute 调用 POSEIDON2。

vendored 的 k256、ark-ff 和 revm-precompile 为 secp256k1、BN254 和 EVM 预编译合约做了同样的事:委托。

这一切之下的 ABI 规范见客户程序 ABI。

发射你的应用

AI 随行手册

一个文件,向 AI 模型讲清如何编写远地虚拟机的客户程序。下载它,交给你的模型,模型就会从本手册所讲的同一套规则起步。

以 Markdown 查看

面向远地虚拟机编写的代码,很多会由模型起草。一个从没见过远地虚拟机的模型,会写出一个看似合理的客户程序(guest):使用 std,在每个循环里分配内存,顺手用上原子计数器,信任它的证明者提示(advice),还提交一个 usize。AI 随行手册是一个 Markdown 文件,预先载入防止这些错误所需的一切:客户程序是什么、硬性规则、SDK 完整的公开接口及其确切签名、可以照抄的模式、各种错误及其修复方法,以及一份审查清单。

如何使用#

  • 在对话中:在描述你想构建的东西之前,先附上这个文件,或者把它作为第一条消息粘贴进去。
  • 在编程智能体中:把它保存到项目根目录,使用你的工具按惯例读取的文件名,例如 AGENTS.md 或 CLAUDE.md,或者把它加入该工具的项目规则。智能体随后会在每次会话开始时读取它。
  • 用于审查:让模型对照文件第 8 节的审查清单,逐行检查一个客户程序。

这个文件用 MUST 和 MUST NOT 来陈述规则,每条旁边都附有理由,因为对于明确且有解释的约束,模型遵循得比需要自行推断的惯例更可靠。

包含的内容#

章节 内容
0. 给模型的指示 把规则当作硬性约束;绝不调用未列出的 API;证明不是零知识的
1. 客户程序是什么 目标、单个 hart、证明陈述的内容、程序身份、三个内存区域
2. 硬性规则 23 条规则:程序形态、32 位的 usize 和指针、bump 分配器、栈、对齐、原子操作、不存在的外部世界、证明者提示、公开值的限制、指令集、浮点数、溢出检查、成本
3. 布局与构建 crate 模板、客户程序工作空间、构建命令、宿主机(host)优先的库拆分
4. 客户程序 SDK 每个公开函数及其确切签名、运行时事实与内存布局、委托的运算,以及 vendored crate
5. 模式 对照哈希检查证明者提示;Merkle 查询与状态转换;缓冲区复用;结构化的证明者提示;为会增长的输出使用摘要
6. 宿主机一侧 在模拟器中运行、性能分析、导出映像、证明与验证、高度与同时处理中的分片
7. 错误与修复 客户程序会遇到的每一种退出状态、致命错误和拒绝,以及各自的原因和修复方法
8. 审查清单 在给出客户程序代码之前要做的十一项检查
9. 基本事实 指令集、证明系统、安全级别、各项限制和实测结果

它着重强调的规则#

随行手册重申了本手册的规则,其中有三条值得特别指出,因为模型最常在这些地方出错:

  • 指针和 usize 是 32 位。 主要在 64 位代码上训练的模型,会不假思索地序列化 usize。于是客户程序和宿主机得出的字节就不一致了。
  • 堆从不释放内存。 惯用的 Rust 写法随意分配内存,因为真正的分配器会把内存还回来。在这里,每一次分配在剩余的运行期间都是永久的,所以随行手册要求处处复用缓冲区、处处使用 with_capacity。
  • 原子操作能编译,但不属于新的客户程序代码。 支持它们是为了兼容现有的库。在单个 hart 上它们不同步任何东西,还会给证明增加一个电路族。

给直接阅读本文档的智能体#

  • /llms.txt 以 Markdown 形式索引本站的每一页,供浏览网页的模型使用。
  • /docs/llms-full.txt 是合成一个文件的全部英文文档,规范也包括在内。
  • 每一页都有一个 以 Markdown 查看 链接和一个 复制为 Markdown 按钮,位于标题下方和右侧栏中。

英文是本文档的标准语言,随行手册在所有语言版本中都以英文发布:英文是模型遵循得最可靠的语言,也是规范的写作语言。

系统架构

系统架构

远地虚拟机的端到端全貌。一份证明陈述什么、从客户程序二进制到合约调用的路径、各大组件如何衔接,以及塑造它们的设计决策。

以 Markdown 查看

远地虚拟机证明 RV32IMAC 程序的执行。本节在大组件的层面描述这个系统:每个组件做什么、为什么这样构建,以及它如何把工作交给下一个组件。审计专区则在每一列、每一个门的层面呈现同一个系统。

一份证明陈述什么#

验证者持有三样不凭证明者一面之词获得的东西:

  • 程序身份,一个域元素,是程序的指令表、初始内存映像、入口 pc 和配置的摘要;
  • 仪式的 SRS 摘要,验证密钥必须带有这个摘要;
  • 一个验证密钥,它可以来自任何人,因为加载它时会根据其自身内容重新计算程序身份和 SRS 摘要,并要求其中的电路与验证者的注册表一致。

证明的陈述带有公开输入、公开输出(journal)、退出状态,以及这次执行的形状记录:分片数、内存窗口、最终的寄存器和 pc,以及每个分片的内存承诺和根。一份通过验证的证明确立的是:具有该身份的程序,在其映像上从入口 pc 启动,输入窗口中存放该公开输入,并带有某份由证明者选择的证明者提示(advice),逐条指令执行直至以该状态调用 EXIT,且已写出该公开输出。对证明者提示不作任何断言,也没有任何东西被隐藏:没有任何承诺或证明经过盲化。

从二进制到合约调用#

远地虚拟机,端到端 四个阶段。程序:客户程序 ELF 到 ProgramImage,再到解码表与 VmConfig,再到程序身份。执行:模拟器按电路族产生行,并切分为分片。每个分片:用 Mercury 承诺各列,运行 GKR 反向过程,在一次批量打开中打开每一列。结算:块证明、递归树、Groth16 判定器、ApogeeVerifier 合约。 1 · 程序 客户程序 ELFRV32IMAC, no_std Rust ProgramImage段 · RVC 已展开 解码表7 个电路族 · VmConfig 程序身份一个 Fr 2 · 执行 模拟器单 hart · (image, io) 的纯函数 行,每个电路族一张表周期 · 调用 · 内存字 分片2^8 … 2^22 行 · 23 个电路族 3 · 每个分片 承诺各列基于 KZG 的 Mercury GKR 反向过程每层一次求和校验 一次批量打开所有列同在一点 · 704 B 4 · 结算 块证明陈述 + 分片证明 递归树叶节点 · 节点 · 根节点 Groth16 判定器绑定线 · BN254 ApogeeVerifier.sol两次配对 · true
四个阶段。程序在任何东西运行之前就已固定;执行被切分为分片;除内存论证在所有分片上统一闭合一次之外,每个分片都各自独立证明;结算把块压缩成合约可以检查的形式。
  1. 程序。 加载器把 ELF 读入一个 ProgramImage,并在原位展开压缩指令。解码器把每条指令路由到七个指令电路族之一,为每个电路族构建解码表(代码的每个半字一行),然后把这一切承诺为程序身份。程序与身份。
  2. 执行。 模拟器在单个 hart 上运行客户程序(guest)。一个周期就是其指令所属电路族中的一行,记录对 pc、寄存器和 RAM 带时间戳的读写。哈希和大整数运算被委托出去:一个 ecall 指明 RAM 中的一个帧,由某个委托电路族的一行在这个帧上完成计算。执行、电路族与分片。
  3. 分片。 一个电路族的行被切分为若干分片,每片的行数等于该电路族的高度,即 2^8 到 2^22 之间的某个 2 的幂。一次执行所触及的内存由窗口电路族的分片覆盖,它们给出每个字的初始值和最终值。分片是证明的单位;一个块(block)有数百个分片。
  4. 分片的证明。 它的各列用 Mercury 承诺。GKR 引擎把电路族的电路从输出反向运行到这些列,每层一次求和校验(sumcheck);所有列都在这一过程终止的那个点上打开,合为一次批量打开。
  5. 块。 BlockProof 由陈述及其分片证明组成。验证时运行一次全局 transcript,逐一运行每个分片的检查,再在所有分片的根上做一次内存核对。
  6. 递归与结算。 由远地虚拟机自己证明的验证者程序,逐段验证连续的分片,并折叠它们被延迟的配对。这些程序构成的树终止于一个根节点,一个 Groth16 电路重新验证根节点,ApogeeVerifier.sol 则检查这份证明以及折叠后的配对。递归与结算。

证明者把客户程序执行两遍:第一遍承诺每个分片的内存列,从而固定陈述及其挑战;第二遍在每个分片填满时证明它。它的内存以同时处理中的分片为界,而不取决于执行的长度。流式证明者。

塑造它的设计决策#

一个域,一条曲线。 一切都建立在 BN254 的标量域上:电路、transcript、承诺和递归。正因如此,递归树的节点才能用自身的算术验证基础分片,整棵树也才能终结于一份由以太坊用其配对预编译合约检查的 Groth16 证明。

用分层 GKR 电路,而不是已承诺的约束表。 电路族的电路是叠在其已承诺列之上的一摞 2 次门层。只有最底层被承诺;其上的每一层都在一次反向过程中由求和校验证明,从不承诺。这一过程结束时,每个已承诺列都在同一个点上有一个断言,所以一个分片恰好只需一次打开。GKR 引擎解释了为什么这是该引擎最核心的节省。

常数大小的打开。 Mercury 用八个曲线点和六个域元素打开一个多线性承诺,共 704 字节,与多项式的大小无关,也与有多少列共享这个点无关。它的检查形如 e(A, [1]_2) = e(B, [x]_2),递归可以折叠这种检查,而不必计算配对。

整个执行只有一个内存论证。 每个电路族的每个分片中的每次访问,都是同一个读/写多重集中的一个元组,验证者对每个陈述只核对一次各个乘积。pc 是这个多重集中的一个单元,所以跨分片的顺序、连续性和周期唯一性都不需要别的论证,也没有哪个分片需要与相邻分片衔接。

委托是电路族,而不是指令。 昂贵的函数拥有自己的电路族,通过一个作用于 RAM 帧的 ecall 调用,并经由同一个多重集与其请求一一配对。指令电路因此保持精简,客户程序只为它实际调用的委托付费。

流式处理,而不是物化。 按每个周期约 300 字节计算,执行轨迹会是系统中最大的对象,所以它从不存在。证明者执行两遍,只持有正在处理的分片。

不借用外部密码学。 域、曲线、配对、MSM、哈希、多项式承诺、GKR 和 Groth16 都在仓库内实现,并逐页写成规范。arkworks、Plonky3 和 zkhash 只作为测试参照(oracle)出现。

可靠性如何组合#

每个分片的 GKR 过程和打开,把其电路的输出与已承诺列联系起来。在此之上,以下论证贯穿整个执行:

断言 承载它的论证
每一行都遵循其指令 电路族电路的约束门,在每一行上均为零
一行的指令就是程序在该行 pc 处的指令 以该行的 pc 和各字段在电路族的解码表中查找,而解码表由程序身份承诺
每次读取都返回最近一次写入 覆盖所有分片的单一多重集;验证者把每个分片的读根和写根连乘,并乘上寄存器与 pc 的边界因子
各行按程序顺序构成一条从入口 pc 到退出的路径 pc 是该多重集中的一个单元,写入它的时间戳至少比读取它的晚四个
一个值是一个字节、一个字、一个符号位、一个 XOR 结果 基于范围表、字节表和通用表的 LogUp 通道
公开输入和公开输出就是所声称的字节 两个公开窗口的初始列与最终列,被约束为这些字节的多线性扩展
委托出去的计算就是该函数的计算 通过同一个多重集读写帧的调用行,与各自的 ecall 一一配对

挑战来自 Poseidon2 双工 transcript。全局 transcript 在内存挑战产生之前吸收整个陈述,包括每个分片的内存承诺;每个分片的 transcript 都以全局 transcript 的最终状态为种子。可靠性地图把每一行断言一直追溯到证明它的章节。

代码,按层划分#

层 Crate 作用
算术 field、curve、poly、sumcheck Fr;Fq 扩域塔、G1、G2、配对、MSM;多线性多项式;零校验(zerocheck)
Fiat–Shamir 与设置 transcript、srs Poseidon2 与双工 transcript;仪式文件导入、KZG、Groth16 的第一阶段
承诺 pcs、pcs-verify Mercury 及其延迟验证
程序 loader、isa、program 从 ELF 到映像、解码器、解码表、VmConfig 与程序身份
执行 emulator、trace 执行器及其追踪器;行、内存状态、列构建器
电路 constraints、gkr-verify、gkr 以数据形式表示的全部电路;GKR 验证者与证明者
证明与验证 verifier-core、verifier、prover 陈述、transcript、密钥、每一项检查;流式证明者
结算 host、groth16、contracts/ 宿主程序(host)SDK、递归树与判定器;ApogeeVerifier.sol
保障 checker、tools/ 独立校验器、篡改测试套件、基准测试、周期分析器、参照

验证者是唯一受信任的一方:证明者不做任何校验,错误的输入只会让诚实的证明者得到一份无法通过验证的证明。安全模型确切列出了可靠性依赖哪些 crate。

系统架构

程序与身份

客户程序 ELF 如何变成一份验证者已知的静态程序描述,以及为什么一个作为程序身份的域元素,就足以告诉验证者一份证明针对的是哪个程序。

以 Markdown 查看

在任何东西执行之前,远地虚拟机先把客户程序(guest)二进制变成程序的一份固定描述:它的映像、按电路族归类的指令、证明时所依据的配置,以及承诺这一切的一个域元素。每一步都是其输入的纯函数。

加载#

加载器接受静态的 32 位小端序 RISC-V 可执行文件,要求其可加载段位于客户程序 RAM 内的偶数地址上、两两不相交,并且至少有一个可执行段。它只读取程序头的类型、偏移、大小和可执行位,别的一概不读:这台虚拟机没有分页,也没有权限,无论各段声明了什么,整个 RAM 都可寻址。

随后它逐个半字扫描每个可执行段。二进制以 11 结尾的半字开始一条 4 字节指令;全零半字是非指令(LLVM 用它填充不可达的基本块);其余的都是压缩指令,在原位展开为 32 位形式。地址从不紧缩:位于 0x1002 的 c.addi 留在原处并占两个字节,因此链接器解析出的每个地址都依然成立,而下一个 pc 是 pc + 2 还是 pc + 4,只由指令的长度记录。

扫描即使失去同步,也无法让错误的指令变得可证明。一个槽位是其自身 pc 处字节的函数,所以每个指令槽位都正是 hart 在那里取指时会解码出的内容。让扫描偏离真实边界的数据只会造成两种结果:丢失真实的指令起点,使这些 pc 在表中没有对应的行;或者遇到一个无人认领的编码,从而拒绝整个映像。

解码与路由#

解码器只接受 32 位字,并且恰好接受 RV32IMA 的 59 条指令:基础指令集 40 条,M 扩展 8 条,A 扩展 11 条。其余一切,从 RV64 编码、浮点到 CSR 和 fence.i,都是解码错误;可执行代码中任何位置只要有一个无法解码的字,无论是否可达,程序都会被拒绝。每条指令都被路由到七个指令电路族中的恰好一个:

编号 电路族 指令
0 ADD_SUB_LUI_AUIPC ecall、ebreak、fence、addi、auipc、add、sub、lui
1 JUMP_BRANCH_SLT slti、sltiu、slt、sltu、六种分支指令、jalr、jal
2 SHIFT_BITWISE 六种移位指令、and、or、xor 及其立即数形式
3 MUL_DIV M 扩展的八条指令
4 MEM_WORD lw、sw
5 MEM_SUBWORD lb、lh、lbu、lhu、sb、sh
6 ATOMICS lr.w、sc.w 以及九条 AMO 指令

这种分组依据的是电路之间共享的部分。一个比较 gadget 为每种分支和每种 slt 确定有符号与无符号的大小关系;一个乘积恒等式同时服务于全部四种乘法和除法;任一方向的移位都是与一个查表所得的 2 的幂的一次乘积。

解码表#

每个指令电路族都有一张解码表,即它的设置列:第 i 行对应 pc 2i,在表所覆盖的地址空间中每个半字一行。有效行以元组 pc, next_pc, rs1, rs2, rd, imm, extra_mask 存放该电路族的一条指令,其中掩码以一个比特指明助记符。其余每一行都是填充行,每个字段都是 −1,所以有效行永远不会是填充行,全零的行也永远不会成为 pc 0 处一条可被声称的指令。

每个周期的行都以自己的 pc 在其电路族的表中查找自身。正是这次查找把一次执行与程序绑定在一起:一行的指令就是程序在该 pc 处的指令,而在任何表中都没有有效行的 pc 无法被可证明地执行。代码因此是静态的。对 .text 的存储会改变之后的加载读到的内容,但永远不会改变执行的内容。

配置#

程序的静态形状就是它的 VmConfig:它用到的电路族(每个都有一个高度),以及代码大小的上限。电路族集合不是选出来的,而是推导出来的:

  • 映像中含有某个指令电路族的指令时,该电路族就存在;
  • 负责内存初始化与收尾的五个窗口电路族总是存在;
  • 委托电路族在映像声明它时存在,声明的方式是一条 12 字节的记录:链接该委托的 shim,就会在映像的字节中留下这条记录。

电路族的高度是它一个分片中的行数,从 2^8, 2^12, 2^16, 2^18, 2^20, 2^22 中选取。每个高度都是 2 的偶数次幂,因为 Mercury 打开要求如此。高度是程序的参数,而不是某次运行的参数;每一种高度选择都构成一个独立的程序。

程序身份#

程序身份是 Fr 中的一个元素:对代码版本、VmConfig、入口 pc 以及每个电路族的设置承诺所做的 Poseidon2 摘要;这些设置承诺是对解码表和映像初始内存字的 Mercury 承诺。

它绑定:扫描找到的每条指令及其 pc、长度、操作数和类别,以及其他 pc 上都没有指令这一事实;映像中每个来自文件的字节,其中包括委托声明;入口 pc;电路族集合、每个高度、代码大小上限和代码版本。它不绑定:仪式和通用查找(lookup)表,它们由 SRS 摘要覆盖;电路,加载密钥时会要求它们与验证者的注册表一致;任何由执行过程选择的东西;以及 ELF 中加载器不读取的任何内容,例如符号表。

计算程序身份需要仪式提供的幂次,用来做承诺。检查它则只需要这些承诺,而验证密钥带有它们:加载密钥时会根据它们重新计算程序身份,每个分片的打开也会对照同样的这些点检查其设置列。正是这一点,把证明读取的表与验证者所登记的程序身份联系在一起。

重要

验证者要从证明者无法控制的渠道获取程序身份。 如果对照的是证明者提供的身份,证明只能说明有某个程序运行过。持有 ELF、参数和仪式文件的任何人都可以重新计算它。

规范见程序与身份。

系统架构

执行、电路族与分片

单个 hart、38 位时钟、每次内存访问都是一次带时间戳的查询、23 个电路族,以及作为证明单位的分片。

以 Markdown 查看

机器#

模拟器在单个 hart 上、基于一个 ProgramImage 运行 RV32IMAC,没有中断,也没有特权级。一次运行是映像及其输入的纯函数,没有时钟、随机数或线程,所以两次运行切出的分片完全相同。它与宿主环境中的 RV32IMAC 有三处不同:sc.w 总是成功;未对齐的半字或字访问是致命错误,而不会被拆分;指令流就是加载时解码的映像。

运行在到达 EXIT 之前停下的其他所有情形,例如访问映射区域之外的地址、ebreak,或者跳转到一个没有指令的半字,都是致命错误,不产生执行轨迹。这样的运行没有证明。非零退出状态不是错误:它和其他任何执行一样,是一次可证明的执行。

时钟与查询#

周期 c 占据四个时间戳 4c + Δ,每个槽位 Δ ∈ {0, 1, 2, 3} 一个。每条指令是一个周期,除此之外没有任何东西是周期:委托调用搭乘发起它的那个周期。周期从 1 开始编号,因为时间戳 0 是每个地址的初始写入。时钟为 38 位,所以一次执行最多运行 2^36 − 1 个周期。

每次内存访问都是一次查询:读取一个在某个更早的时间戳上最后写入的值,并在当前时间戳上写入。只读的查询会写回它读到的值。每个周期的槽位 0 是 pc 查询,它读取 pc 并写入 next_pc。随后是该指令在固定槽位上的寄存器查询和内存查询:

类别 Δ = 1 Δ = 2 Δ = 3
寄存器-立即数类、jalr rs1 rd
分支 rs1 rs2
寄存器-寄存器类、M 扩展 rs1 rs2 rd
加载 rs1 内存字(读取) rd
存储 rs1 rs2 内存字,并入存储的字节
原子操作 rs1 rs2 内存字,以及 rd
ecall a7 a0 a0,以及委托的镜像查询

地址分属不同的空间:32 个寄存器、按 4 字节对齐的字编址的 RAM、pc、每种委托类型各一个锚点空间,以及递归格式的域单元。x0 在执行轨迹中是一个普通寄存器,在机器中则是一个常量:对它的每次查询都读写 0。

二十三个电路族#

电路族是一个电路及其所证明的行。电路族分为四类:

类别 电路族 一行是
执行 0–6:ADD_SUB_LUI_AUIPC、JUMP_BRANCH_SLT、SHIFT_BITWISE、MUL_DIV、MEM_WORD、MEM_SUBWORD、ATOMICS 一条被执行的指令
窗口 7 INIT_TEARDOWN、8 ZERO_WINDOWS、12 PUBLIC_INPUT、13 PUBLIC_OUTPUT、14 ADVICE_WINDOWS 一个内存字,经过初始化与收尾
委托 9 KECCAK_F、10 POSEIDON2、11 FR_ARITH、15 MOD_MUL、16 SHA256_COMP、17 EC_ADD 作用于一个 RAM 帧的一次调用
递归 18 FIELD_WINDOWS、19 FR_OP、20 P2_FIELD、21 FIELD_IO、22 FQ_OP 一个域单元,或一次协处理器运算

每个周期都归入这样一个执行电路族:其解码表认领该周期的 pc。各电路族在时间上交错:ADD_SUB_LUI_AUIPC 可能拥有周期 1 和 3,JUMP_BRANCH_SLT 拥有周期 2。没有任何东西要求它们连续,因为内存论证按 pc 写入为每一行排序。

窗口电路族之所以存在,是因为内存论证要求一次执行触及的每个地址都恰好有一个初始值和一个最终值。INIT_TEARDOWN 覆盖 RAM 窗口 0,以程序映像作为其初始内容;ZERO_WINDOWS 覆盖这次运行触及的普通 RAM 的其他每个窗口,初始为零;两个公开窗口电路族覆盖输入窗口和公开输出(journal)窗口;ADVICE_WINDOWS 覆盖证明者提示(advice)区域,以证明者的字节作为初始内容。

分片#

一个电路族的行按出现顺序被切分为若干分片,每片的行数等于该电路族的高度。最后一个分片用全零行填充,电路在构建时就保证接受这样的行。不论占用率如何,一个分片都按其完整高度计算成本,所以程序触及哪些电路族、选择哪些高度,决定了每份证明的下限。

电路族 默认高度 原因
指令电路族 2^22;MUL_DIV 和 ATOMICS 为 2^20 其时间戳范围检查的下限是 2^20
RAM 窗口 2^22 共用一个窗口高度,至少为 2^16
公开输入、公开输出 2^12 固定不变:高度决定了窗口的位置
KECCAK_F、SHA256_COMP 2^18 可容纳的调用是其 2^16 下限的四倍,证明只大 2%
MOD_MUL、EC_ADD 2^16 其下限
POSEIDON2、FR_ARITH 2^8 不用表,因此没有下限

每个分片都由其电路族的电路独立证明,内存论证除外:每个分片的电路输出其读元组之积与写元组之积,验证者在陈述的所有分片上对这些乘积做一次核对。这是把各分片联系在一起的唯一纽带。不存在逐分片的 pc 衔接,相邻分片之间也没有共享边界。

在默认高度下,分片证明的大小从公开窗口的约 12.5 KB 到 POSEIDON2 分片的 665 KB 不等;指令电路族的分片证明为 64 至 77 KB。性能页面给出了完整的表。

规范见执行轨迹、电路与注册表。

系统架构

GKR 引擎

远地虚拟机核心的证明引擎。为什么分层 GKR 电路只承诺它的输入,一次由求和校验构成的反向过程如何把整个电路汇聚到单个点上,以及这带来了什么。

以 Markdown 查看

远地虚拟机中的每个分片都以同一种方式证明:使用其电路族的电路,这个电路写成一摞层,由 GKR 引擎从电路的输出反向运行到其已承诺列。本页讨论为什么这个引擎位于系统的核心,以及为什么它所属的这一类证明系统,正是证明技术前沿推进的方向。

基本思想#

证明一项计算的经典做法,是把它排成一张表,承诺每一列(包括中间值),再证明一组约束在整张表上为零。承诺是其中昂贵的部分:每个被承诺的列都要花费一次多标量乘法或一棵 Merkle 树,还要在约束论证问到的每个点上做一次打开。

GKR(得名于 Goldwasser、Kalai 和 Rothblum)改变了必须承诺的内容。计算被表示为一个分层电路。只有最底层,也就是输入,才被承诺。其上的每一层都由作用于下一层的门定义,证明者从不承诺它。取而代之的是:关于顶层的断言,经一次求和校验(sumcheck)归约为关于其下一层的断言,再归约到更下一层,直到这些断言全部落在已承诺的输入上,并且都位于同一个随机点。一次打开就能了结它们。

说明

这是类比,不是名称。 传统的证明者像一枚火箭:它把自己产生的每一个中间值都承诺下来、运到目的地,并为这些质量付出代价。GKR 引擎的行为更像科幻作品中的曲速引擎:移动的是飞船周围的空间,而不是飞船本身。真正被运送的是断言,它沿着电路一层一层向下移动,而中间层根本不会被运到任何地方。

这给 zkVM 带来的好处:

  • 中间值不花费任何承诺。 电路族的电路可以计算数百个内部列、乘积树和分式树,它们无一被承诺。被承诺的只有执行轨迹列。
  • 每个分片一个打开点。 反向过程结束时,每个已承诺列都在同一个点上有一个断言。无论有多少列,一个分片都恰好只需一次批量打开,704 字节。
  • 证明者的工作是域运算。 每一层的求和校验在 Fr 上进行,耗时与该层的大小成线性关系,每层都不需要承诺、变换或哈希。
  • 论证在电路内部组合。 内存论证的大乘积和查找(lookup)的 LogUp 求和,只是同一个电路中更多的层,在同一个过程中归约。

电路族的电路,逐层来看#

电路族电路的各层 第 0 层是已承诺列,旁边是虚拟表。门列表 0 构建内存叶子、查找分式和约束门。逐行列表把每一行的各棵树归约为一个节点。折半列表每个变量一个,把各行逐级相乘、相加,直到一个没有变量的顶层:读根、写根,以及每个通道的分子和分母。反向过程从顶层运行到第 0 层,每次层间转换一次求和校验,并终止于一个点,在那里一次 Mercury 打开了结每一个已承诺列。 第 0 层 · 已承诺列 M ‖ W ‖ S 旁边是虚拟表 V,即由验证者求值的闭式表达式 · 唯一被承诺的一层 门列表 0 · 逐行 内存叶子 · 查找分式 (num, den) · 每个约束门,次数 ≤ 2 逐行列表 1 … r 乘积树与分式树合并兄弟节点,直到每一行的每棵树只剩一个节点 折半列表,每个变量一个 TreeProduct · TreeCross:合并行 (·,0) 与 (·,1) 顶层 · 无变量 读根 · 写根 · 每个通道的 (num, den) 反向过程 先吸收输出 每次层间转换一次求和校验: 第 k+1 层上的断言 变为第 k 层上的断言 折半:直线上的点 τ 合并两个子节点 每个已承诺列的断言 都位于同一点 u → 一次 Mercury 打开 前向过程 · 仅证明者执行
一个电路族的电路。证明者自下而上把每一层计算一次(虚线)。证明则自上而下进行:先吸收输出,然后每次层间转换都是一次求和校验,把关于某一层的断言变成关于其下一层的断言,直到所有断言在已承诺列上汇聚于同一点。

最底层是分片的已承诺列,分为三种,区别在于何时被绑定:M,内存列,在任何内存挑战产生之前于陈述的全局 transcript 中承诺;W,见证列,在分片自己的 transcript 中承诺;S,设置列,由程序身份或仪式绑定。它们旁边是虚拟表:诸如行索引或 16 位范围这样的闭式表达式,验证者可以在任意点上对其求值,它们从不被承诺。

在第 0 层之上,每个电路族的电路都有相同的结构:

  1. 门列表 0 逐行计算内存叶子(每个查询的读元组和写元组)、查找分式(每个查找一个 (numerator, denominator) 对,另加表本身的一对),以及每个约束门:电路族的约束,每一个都是必须在每一行上为零的多项式。
  2. 逐行列表合并兄弟叶子:乘积树把元组相乘,分式树按 (n_a·d_b + n_b·d_a, d_a·d_b) 把分式相加,直到每一行的每棵树只剩一个节点。
  3. 折半列表按分片高度的每个变量各有一个,把各行两两合并:前一半与后一半。经过 n 个折半列表后,电路到达一个没有变量的顶层:分片的读根、写根,以及每个查找通道最终的分子和分母。

所以一个电路在一个过程中同时证明电路族的约束、计算它对内存论证的贡献,并对它的查找求和。每个门的次数至多为 2,所以每次求和校验的每个轮多项式都是三次多项式。

反向过程#

证明者自下而上把每一层物化一次。随后证明自上而下进行,它的 transcript 时序对每个电路都相同:

  1. 输出首先被吸收,这样在任何挑战产生之前,证明者就已被这些根绑定。
  2. 对于从第 k + 1 层到第 k 层的每次转换,一个挑战 λ 把第 k + 1 层上的每个断言,连同该列表的每个约束门,批量合成一个和。一次求和校验把这个和归约为在随机点 ρ 上的一个求值,每个变量一条三次多项式消息。
  3. 证明者给出第 k 层各列在 ρ 上的值。对于折半列表,它给出两个子节点的值,再由另一个挑战 τ 把它们合并为每列一个断言。
  4. 到第 0 层时,每个已承诺列都有一个断言,全部位于同一点 u。

分片唯一的那次 Mercury 打开,对照各承诺证明这些断言:内存列的承诺取自陈述,见证列的取自分片证明,设置列的取自验证密钥。虚拟表由验证者自己求值。

约束门可以免费搭车。约束门断言自己处处为 0,所以它加入其所在转换的批次;一旦某个门被违反,批量和以压倒性的概率不为零。一个 LayerInconsistency 错误同时涵盖错误的下行断言和被违反的门:批量和无法区分二者,证明也不为区分它们花费任何代价。

为什么它是可靠的#

每个挑战都在它所保护的一切之后抽取:

  • 输出点在输出之后抽取,所以证明者无法挑选只在将被检查之处与真实值一致的表;
  • λ 在断言和点之后抽取,所以虚假的断言或被违反的门,只有当 λ 恰为某个非零多项式的根时才能幸存;
  • 每个求和校验挑战在其所在轮的三次多项式之后抽取,所以错误的三次多项式与真实的三次多项式一致的概率至多为 3/|Fr|;
  • τ 在两个子节点的值之后抽取。

对注册表中任一电路在其默认高度下的全部转换求和,可靠性误差保持在 2^14/|Fr| 以下;这里 Fiat–Shamir 基于 Poseidon2 transcript,并在随机预言机模型中分析。

以数据表示的电路#

电路不是代码,而是一个 CircuitArtifact:按名称列出的已承诺列、虚拟表、使用七种门形状的门列表、同一组关系的扁平列表、查找以及填充行,以规范形式序列化。四条法则让每个制品都保持一致的形式:每个操作数在被读取之处都可读;每个列表的宽度由其门推导而来;顶层恰好就是输出;分层的门与扁平关系是同一个约束集。这些法则只在构建制品或加载密钥时运行一次,从不针对每个证明运行。

对任何评估这个系统的人来说,有两点后果值得关注:

  • 验证密钥携带自己的电路,验证者要求它们与自己的注册表一致。 程序身份绑定程序;注册表绑定证明该程序的电路。
  • 电路可以由第二套实现检查。 checker crate 在不共享构造器代码的前提下,重新实现了各项法则、查找规则和填充约定,并且只通过双方都视为语义权威的那一个门内核对门求值。

代价#

这个引擎用内存换掉了承诺。前向过程以域元素的形式持有每一个内部层:最宽的指令电路族的一个 2^20 分片约占 8.4 GiB,一个 2^18 的 KECCAK_F 分片占 42 GiB,其电路要计算 5,490 个内部列。这就是证明受内存限制的原因,也是高度成为调优参数、流式证明者以同时处理中的分片为内存上界的原因。证明大小只会随每层每个变量增加一轮求和校验而增长:一个 KECCAK_F 分片的证明在 2^18 时为 381,100 字节,在 2^16 时为 373,276 字节,而前者的工作量是后者的四倍。

规范见 GKR 引擎,以及审计专区下每个电路族各自的页面。

系统架构

承诺

每个已承诺列都用 Mercury 打开,这是一种基于 KZG、打开大小恒定的多线性承诺。它的成本、一个分片如何把所有列批量合成一次打开,以及递归如何延迟配对。

以 Markdown 查看

远地虚拟机承诺的每一列都是一个多线性多项式,即布尔超立方体上 2^n 个求值构成的表。每一列都用 Mercury(Eagen 与 Gabizon,ePrint 2025/385)承诺和打开,最后以 BDFG20(Boneh、Drake、Fisch 与 Gabizon,ePrint 2020/081)的批量 KZG 打开收尾。规范确定了这些论文没有确定的部分,并增加了两样东西:同一点上多列的批处理,以及由递归树折叠的延迟形式。

承诺#

Mercury 承诺恰好就是把求值表当作系数读取时的 KZG 承诺:在仪式的 τ 的前 n 个幂次上做一次多标量乘法,得到一个 G1 点,64 字节。它背后没有第二套方案。由此得出两条性质,系统的其余部分两者都会用到:

  • 它是线性的。 Σ ρ^i·f_i 的承诺就是 Σ ρ^i·cm_i,正因如此,多列才能共享一次打开。
  • 零系数不增加任何东西。 用全零行扩展的列保持原有承诺,所以通用查找(lookup)表的三个承诺适用于每一个容得下这张表的高度,它们是仪式的常量。

各列按其整数宽度承诺:比特列、字节列、半字列或字列都走一次以 u32 为标量的 MSM,按 32 位而不是 254 位重新编码,这让承诺执行轨迹保持低成本。

打开#

Mercury 把 s = 2t 个变量的打开点 u 分成两半,用挑战 α 折叠多项式,用一个对称化的见证证明剩下的两个内积,最后在三个点上做一次批量 KZG 打开。对任意大小,证明都是八个 G1 点和六个域元素:704 字节。由此产生一项要求:变量数必须是偶数,这就是为什么可选的每个高度都是 2 的偶数次幂。

验证者做 O(t) 次域运算、规模分别为十个点和两个点的 MSM,以及一次包含两对的配对检查。它检查的两个关系都写成 e(A, [1]_2) = e(B, [x]_2) 的形式,所以两个 G2 参数都是设置中的常量,验证者完全不做 G2 运算。也正是这种形状,让递归可以推迟配对,而不必计算它。

一个分片的所有列,一次打开#

GKR 过程结束时,分片的每个已承诺列都在同一点 u 上有一个断言。所以不需要任何合并断言的求和校验(sumcheck):打开直接把它们全部批量处理。挑战 ρ 在每个承诺和每个声称值之后抽取,以 ρ^i 为第 i 列加权;证明者对 Σ ρ^i·f_i 打开一次,验证者通过一次 k 点 MSM 构造 Σ ρ^i·cm_i。虚假断言幸存的概率至多为 (k − 1)/|Fr|。

批中的列来自三处,按固定顺序排列,而这个顺序本身也是被证明内容的一部分:内存列的承诺来自陈述,见证列的来自分片证明,设置列的来自验证密钥。正是从密钥中取得设置承诺,才使这次打开绑定了程序身份所承诺的解码表和映像。

延迟验证#

延迟检查执行除配对之外的一切,并把关系中的各项保存为十二个 (side, scalar, point) 条目。递归正是这样使用它的:树的每个节点用自身的算术计算一个分片的十二个标量,用新抽取的挑战为它们加权,并把它们累加进一对持续更新的点 (A, B)。整棵树中每个分片的打开都折叠为一个断言 e(A, [1]_2) = e(B, [x]_2),最终只由以太坊上的合约检查。递归与结算展示了折叠过程。

实测数据#

在一台 18 核的 Apple M5 Pro 上:

操作 时间
承诺一个 2^22 的列 1.30 s
打开一个 2^22 的列 2.89 s
16 个 2^20 的列,作为一批打开 1.01 s,验证耗时 4.8 ms
同样 16 列逐个打开 9.79 s,验证耗时 62 ms

安全性#

在代数群模型中、q-DLOG 假设下,知识可靠性成立;前提是 Fiat–Shamir 基于 Poseidon2 transcript 并在随机预言机模型中分析,且 SRS 的 τ 无人知晓。统计项,即对各挑战应用 Schwartz–Zippel 得到的项,对所用的每个实例都保持在 2^−220 以下,所以安全级别就是 BN254 的安全级别,约 100 位。没有任何部分具备隐藏性,也没有任何部分经过盲化:远地虚拟机 v1.0.0 不是零知识的。

SRS 是 PSE 的 perpetual powers of tau 的第 80 次贡献,只要有一位贡献者是诚实的,它就是可靠的。它的文件会被解码,每个点都要检查是否位于其曲线上、是否属于正确的子群;但没有任何东西能证明一个文件就出自该仪式,所以验证者要把 SRS 摘要与通过自己的渠道获得的仪式 SRS 摘要相比较。

规范见 Mercury、结构化参考串。

系统架构

内存与查找

两个论证承载一切跨行的内容。覆盖整个执行过程、只核对一次的单一读/写多重集,使每次读取都返回最近一次写入,并为每一行排序;LogUp 通道使每个值都是一个字节、一个字或一个表行。

以 Markdown 查看

电路族的门每次只约束一行。一切跨越行、分片或电路族的内容,例如寄存器存放的值、加载返回的值、一行执行的是哪条指令、一个值是否放得进 32 位,都由两个论证承载,而这两个论证就位于同样的 GKR 电路之中。

内存论证#

每次内存访问都变成一个域元素,即一个元组:

T(AS, ADDR, TS, VAL) = γ + AS + α_addr·ADDR + α_ts·TS + α_val·VAL

它作用于地址空间、地址、时间戳和值,其中的四个挑战每个陈述只抽取一次。一个查询把它的读元组贡献给一侧,把写元组贡献给另一侧。每个分片的电路输出两个数:其读元组之积与写元组之积。然后验证者在整个陈述上检查一个等式:

∏ read roots · R_b  =  ∏ write roots · W_b          over every shard of every family

W_b 和 R_b 是 32 个寄存器和 pc 的初始元组与最终元组,它们没有自己的行:验证者根据陈述携带的 64 个边界标量把它们乘进去。RAM 的初始值和最终值来自窗口电路族的分片,每行存放一个字:窗口 0 中是程序映像,这次运行触及的其他每个窗口中是零,公开输入窗口中是陈述的输入,证明者提示(advice)中是证明者的字节。

如果等式成立,那么两个多重集以压倒性的概率相等。多重集相等意味着每次读取都返回在它之前的最近一次写入:每次读取恰好与一次写入匹配,读取必须严格晚于它所消费的写入,并且每个地址恰好有一次初始写入。

免费得到的顺序#

pc 和其他内存单元一样,也是一个内存单元,位于它自己空间的地址 0。每一行都读取 pc 并写入下一个 pc,其电路要求这次写入至少比读取晚四个时间戳。所以 pc 的历史是一条穿过每个电路族每个有效行的路径,从入口点一直到退出行。这一条路径带来:

  • 程序顺序,因为各行按其 pc 写入排序;
  • 跨分片、跨电路族的连续性,因为每一行的 pc 读取都消费了某一行的 pc 写入;
  • 没有周期被证明两次,因为任何写入都不能被消费两次。

没有哪个分片与相邻分片衔接,也不需要衔接。分片所声称的时间窗口不绑定任何东西,只检查其形状。

什么必须先确定#

内存挑战每个陈述只抽取一次,位于全局 transcript 的末尾,在元组可能读取的一切都已固定之后:每个分片的内存承诺、程序身份(它固定了入口 pc 和映像)、仪式摘要、分片数和窗口列表、公开输入与公开输出(journal)的摘要,最后是 64 个边界标量。在挑战之后才选定的值,可以被反解出来;正是 transcript 的顺序禁止了这一点。出于同样的原因,内存元组只能读取内存列、设置列和虚拟列,绝不能读取见证列,因为见证列是在挑战之后、于分片自己的 transcript 中承诺的。电路构造器会拒绝任何违反这一点的制品。

查找#

查找(lookup)是说:一行中若干值构成的元组,是某张表中的一行。远地虚拟机用 LogUp 证明一个分片的每个查找:每张表(即每个通道)一个恒等式,

Σ_rows Σ_lookups 1/(E(y) + g)  −  Σ_rows mult(y)/(T(y) + g)  =  0

该恒等式由电路族自身 GKR 电路中的分式树求和,并在树根处检查:分子为零,分母非零。重数列完全不需要约束:一个不在任何表行中的元组会留下一个重数无法抵消的极点。

通道 表 用途
TIMESTAMP [0, 2^19),虚拟 每个查询的时间戳差,拆为两个 19 位块
RANGE16 [0, 2^16),虚拟 把 32 位值拆为两个半字;进位;帧的边界
XOR8 所有字节对及其 XOR,虚拟 逐字节计算 Keccak 和 SHA-256
GENERIC 一张已承诺的表,含 AND 行、符号行和移位幂行 按位运算、符号位、移位量
DECODER 电路族的解码表,由程序身份承诺 把每个被执行的行绑定到程序

其中三张表是虚拟的:它们是行索引的闭式表达式,由验证者自己求值,不花费任何承诺。通用表用仪式的幂次承诺一次,并由 SRS 摘要覆盖。

解码器查找#

每个执行电路族都为每个有效行在自己的解码表中做一次查找,以该行从内存中读到的 pc 为键。这一次查找就把这个周期绑定到程序:该行的操作数、立即数和指令类别都是程序在该 pc 处的;它的类别位是 one-hot 的,因为表的每个有效行都存放一个 one-hot 掩码,而每个填充行存放 −1,任何类别位之和都达不到这个值。如果程序在某个 pc 处没有指令,位于该 pc 的行根本找不到任何表行。

键必须有界#

通道证明的是属于某张表,仅此而已。几张子表以互不相交的键范围共用通用表,所以一个无界的键可能落进错误的子表,从而证明一个错误的 AND。因此,每个电路族都用同一选择子下的范围查找,为它查找的每个键设定界;电路构造器还会检查:通过缩放因子写出的界,同时也带有直接界。规范针对每个电路族说明了这一规则所防范的攻击。

如何组合起来#

这两个论证与每个电路族的门一起,赋予陈述其含义:每一行都遵循其指令,该指令就是程序的指令,每次读取都看到最近一次写入,各行构成一条从入口到出口的路径,每个值都是它所声称的那个整数,公开窗口存放的是陈述中的字节。可靠性地图把每项断言对应到证明它的章节。

规范见内存论证、查找。

系统架构

委托

昂贵的函数如何在不扩大指令电路的前提下拥有自己的电路。调用方式、让每个请求恰好与一次调用配对的锚点、六个电路,以及它们的经济性。

以 Markdown 查看

哈希和大整数运算在真实工作负载中占主导:在一个以太坊主网区块中,被委托之前,仅 secp256k1 的域乘法和平方就占了 44% 的周期。逐条指令地证明它们是可行的,但很慢。委托为这样的函数提供一个专属的电路族,由客户程序(guest)调用,这样指令电路保持精简,程序也只为它调用的委托付费。

调用方式#

委托只会被调用,从不被解码。客户程序在 RAM 中写入一个由 32 位字组成的帧,然后发出 ecall,a7 中放委托编号,a0 中放帧的基地址。这个 ecall 是 ADD_SUB_LUI_AUIPC 电路族中的一行,即请求。实际的计算是该委托自身电路族中的一行,即调用:它在发起请求的那个周期读取帧中的每个字,并把每个字(其中包括结果)写回。调用不拥有任何周期,它搭乘发起它的那个周期。

帧按字对齐,并完全位于普通 RAM 中,所以任何帧都不会与公开窗口或证明者提示(advice)区域重叠,帧的读写也都是普通的内存查询。因此,委托计算出的结果与任何一次存储一样受到绑定:经由那唯一的内存多重集。

锚点#

请求与调用必须一一配对:否则多个请求可能对着同一次调用闭合,使一些调用实际上没有执行;或者一次未被请求的调用可能改写某个帧。它们通过同一个内存多重集配对,配对发生在一个锚点空间中,这个空间只属于该委托类型,任何指令都无法触及:

读 写
请求,周期 c T(s, base, 0, 0) T(s, base, 4c + 3, v)
调用 T(s, base, 4c + 3, v′) T(s, base, 0, 0)

请求一侧的三个门把它的读取固定在时间戳 0、值 0,并使它向 a0 写入 0。于是,时间戳为 0 的元组恰好就是各请求的读取和各调用的应答,所以在相同的基地址上,调用与请求一样多;又因为没有两个请求共享同一个周期,每次调用的读取恰好就是某一个请求的写入。每次调用都位于其请求的基地址和周期上。委托电路中的任何门都完全不必知道请求的存在。

多次调用,一个操作#

一行容纳不下的操作,就在同一个帧上分成几次调用,由帧中的一个字指明是哪一步:一次 keccak-f[1600] 置换是 24 次轮调用,一次 SHA-256 压缩是 16 次各含四轮的调用,一次完全点加是三次调用。没有任何门把两行连接起来。每次调用都在帧的当前状态上证明自己那一步;它的读取位于每个字唯一的那条内存历史上,所以读到的是上一步的写入。每一步都按顺序执行,这要由调用代码来保证,而调用代码是以指令形式被证明的客户程序代码。SDK 用一个函数发出每个多次调用的操作,所以客户程序从不需要手动安排步骤的顺序。

静态声明#

指令扫描看不到调用,因为委托编号是 a7 在运行时的值。所以 SDK 中的每个 shim 都在自己的链接器段中留下一条 12 字节的声明记录,只有当该 shim 可达时,这条记录才会被保留。程序推导过程在映像中扫描这些记录,被声明的电路族加入配置,并通过映像字节受程序身份绑定。已链接但从未调用的电路族证明零个分片;如果被调用的编号所对应的电路族从未被程序声明,就没有证明。

六个电路#

电路族 一次调用 构造方式
KECCAK_F 在 51 个字的帧上执行 keccak-f[1600] 的一轮 字节:每轮 1,020 次 XOR8 查找;循环移位表示为字节与掩码副本上的线性形式
SHA256_COMP 四轮以及四个消息调度字 字节与 XOR8:每轮 52 项义务,每个调度字 32 项;Ch 和 Maj 表示为 XOR 的线性形式
POSEIDON2 一次宽度为 3 的置换 各轮在电路自身的层中计算,每轮三个门列表,不用查找;唯一一个在第一层之上进行计算的委托
FR_ARITH Montgomery 形式下的一次 Fr 加法、乘法或求逆 比特分解,以及针对 p 的规范性链
MOD_MUL 一次 256 位的 a·b mod m,四个以太坊模数 32 位 limb、一个商、进位,以及证明 out < m 的规范性链
EC_ADD secp256k1 或 BN254 G1 上一次完全点加的三分之一 Renes–Costello–Batina 的完全加法公式,每行三次约简

有几种构造在它们之间反复出现。单一编码规则把一个指明 k 种情形之一的帧字解码为布尔选择子,并保证恰好只有一个被置位,因为编码会相加:没有这条规则,选择子 1 和 3 就能应答对 4 的请求。规范性链通过 32 位 limb 上的借位,证明一个 256 位的值小于某个模数。此外,每个被写入的字都被限定在 2^32 以下,使 RAM 中始终都是字,这是每个指令电路族都依赖的前提。

经济性#

委托电路族的高度决定一个分片能容纳多少次调用,而不论占用率如何,一个分片的成本都按其高度计算:

电路族 高度 每分片单元数 分片证明
KECCAK_F 2^18 10,922 次置换 381,100 B
SHA256_COMP 2^18 16,384 次压缩 189,988 B
EC_ADD 2^16 21,845 次点加 434,916 B
MOD_MUL 2^16 65,536 次乘法 135,220 B
POSEIDON2 2^8 256 次置换 664,780 B
FR_ARITH 2^8 256 次运算 266,292 B

对于调用很多的电路族,更大的分片反而更便宜:KECCAK_F 的证明从 2^16 到 2^18 几乎没有变大。代价是内存。一个 2^18 的 KECCAK_F 分片的前向过程占 42 GiB,同时处理中的两个这样的分片决定了实测区块的内存峰值。

哪些被委托,哪些没有#

库代码通过打过补丁的 k256、ark-ff 和 revm-precompile 副本用上委托:secp256k1 签名恢复变成建立在 MOD_MUL 和 EC_ADD 之上的 k256 代码,BN254 配对变成建立在 MOD_MUL 之上的 ark-bn254 代码。EVM 中任意模数的 MULMOD、MODEXP、BLS12-381,以及所有作为整体的签名方案,都以普通指令运行。面向客户程序的专用签名支持,是 v2.0.0 演进方向的一部分。

规范见委托 ABI、委托电路。

系统架构

流式证明者

证明者把客户程序执行两遍,从不持有执行轨迹。内存取决于正在处理的分片,而不是运行的长度;证明也不依赖于调度。

以 Markdown 查看

一个完整的以太坊区块约有 2 亿个周期。它的执行轨迹,即每个电路族的每一行和每一个内存事件,每个周期约 300 字节:在承诺任何一列之前就已达到数十 GB。远地虚拟机的证明者从不构建它。它把客户程序(guest)执行两遍,只持有正在处理的分片。

两遍执行#

两遍执行 第一遍执行客户程序,每填满一个分片,就承诺它的内存列,只保留承诺;在退出时构建陈述,并抽取共享的挑战。第二遍再次执行,每填满一个分片就证明它,只保留证明。 第一遍 · 承诺 执行 分片填满承诺 M 列 · 丢弃行 退出时窗口 · 陈述 · G1–G11 挑战已确定内存挑战 · 全局摘要 第二遍 · 证明 再次执行 分片填满所有列 · GKR · 打开 保留证明丢弃其余一切 BlockProof根写入陈述 每个分片的 transcript 都以全局摘要为种子
先承诺,再证明。内存挑战必须在每个分片的内存承诺之后产生,所以承诺先来自一次执行,证明再来自第二次执行。

第一遍执行客户程序,每填满一个分片,就承诺它的内存列,保留承诺,丢弃行数据。退出时,它从最终的内存状态推导出陈述所需的其余一切:寄存器与 pc 的边界、这次运行触及的内存窗口列表,以及窗口电路族的分片。随后它运行全局 transcript:吸收陈述(包括每一个内存承诺),并抽取内存挑战,以及作为每个分片种子的摘要。

第二遍再次执行。模拟器是其输入的纯函数,所以切出的分片完全相同,第二遍还会断言自己的周期分布、窗口列表和边界与第一遍相同。每个分片都得到它的全部已承诺列并被证明:它的 transcript、它的 GKR 过程、它的打开。内存列不会被重新承诺:打开从陈述中取得它们的承诺,从第二遍取得它们的值,所以如果两遍之间某些列不一致,得到的打开会被验证者拒绝。

这个顺序是可靠性所强制的。内存挑战必须在内存元组可能读取的每个值之后产生,所以每个分片的内存列都要在任何分片能够被证明之前承诺。

流水线#

固定数量(max_in_flight)的工作线程共用一把保护执行器的锁。持有锁时,工作线程交回已完成的分片并认领下一个:如果有已填满的分片在等待,就认领它;否则它自己推进执行器,直到某个缓冲区填满。在锁外,它构建该分片的列、证明它,然后丢弃它。

  • 执行器从不超前于需求运行。 每个电路族至多有一个已填满、未被认领的分片以行的形式等待。
  • 在一个分片内部,工作是数据并行的,分布在所有核心上。阻塞在这项工作中的工作线程不会再接第二个分片。
  • 块(block)不依赖于调度。 分片的证明是全局状态和它自身各列的函数;证明按其在陈述中的位置放置。同时处理 1 个和 8 个分片时,得到的字节完全相同。
  • 失败是确定性的。 无论工作线程有多少,返回的都是按填充顺序最早的那个失败。

成本#

内存由以下几部分构成:每个电路族一个未填满的缓冲区、最近访问表、正在处理的分片,以及输出。一个分片的工作集以其前向过程为主,即以域元素形式存放的每个 GKR 内部层:一个 2^20 的 SHIFT_BITWISE 分片为 8.4 GiB,一个 2^18 的 KECCAK_F 分片为 42 GiB。所以 max_in_flight 决定了有多少个这样的工作集同时存在,而高度决定了每个有多大。

在第 257,510 号区块(60 笔交易,101.5 Mgas,198M 个周期,207 个分片)上实测,机器为 32 vCPU、247.7 GiB 内存,同时处理十二个分片:

第一遍 191 s;单线程填充占其分片秒数的 81%;采样内存至多 15.9 GiB
第二遍 2,290 s;客户程序退出之前,12 个分片名额中约有 12 个被占用,30.4 个 vCPU 处于忙碌;随后是 460 s 的尾段
内存峰值 173.92 GiB:两个 2^18 的 KECCAK_F 分片,在尾段同时处理,此外再无其他分片在处理中

峰值来自某一个委托电路族的高度,而不是十二个同时处理中的分片。这就是可调节之处:Keccak 调用更少的区块,或者把 Keccak 设在更低的高度,峰值就更低。

规范见流式证明者。

系统架构

递归与结算

由数百个分片组成的基础证明,如何变成一份由合约检查的 Groth16 证明。远地虚拟机证明它自己的验证者、让这件事变得低成本的 tape、贯穿整棵树的 transcript 链,以及不断折叠、直到只剩一个的配对。

以 Markdown 查看

一个以太坊区块的基础证明由 207 个分片证明组成:14.5 MB,每个分片都是一份 GKR 证明,带着它的承诺,以及一次以配对检查收尾的 Mercury 打开。合约无法直接检查其中任何一项。递归负责压缩它,而递归的做法,是仅次于 GKR 引擎本身的、影响最深远的设计选择。

基础证明保持不变#

第一个决定,是递归不做什么。没有任何基础密钥、陈述或证明为了支持递归而改变:叶节点验证基础分片的方式,与原生验证者完全相同。递归所需的一切都添加在基础证明之上,从不进入其内部。同一个块(block)的同一份字节,可以原生验证,可以递归,也可以两者兼做。

远地虚拟机证明它自己的验证者#

树的节点就是远地虚拟机在证明一个验证者程序。叶节点验证一段连续的基础分片,从 from 到 to;内部节点验证二到四个子节点,每个子节点都是叶程序或节点程序的一份完整证明;根节点覆盖全部基础分片。每个节点都由证明基础块的同一个流式证明者来证明。

把 Rust 验证者作为 RISC-V 指令运行是可行的,实测验证一个区块的 207 个分片需要 30 亿个周期:是该区块本身的十五倍。所以节点改为在递归格式中运行。

递归格式#

当且仅当一个陈述的程序声明了域电路族时,该陈述采用递归格式。该格式增加一个地址空间和基于它的四个协处理器电路族,全部通过普通的委托 ABI 调用,除此之外不做任何改变:

电路族 一行是
FIELD_WINDOWS 一个域单元:存放一个完整 Fr 元素的内存单元,与 RAM 处于同一个内存多重集中
FR_OP 一次作用于单元的域运算:乘、加、减、乘累加、求逆、断言相等,以及构造常量的步骤
P2_FIELD 一次作用于单元的 Poseidon2 双工步骤,所以 transcript 以每次置换一行的速度运行
FIELD_IO 把八个 RAM 字装入一个单元,或把一个单元拆回八个字
FQ_OP 一次 BN254 基域运算,一个元素由四个存放 64 位 limb 的单元组成,所以曲线算术以每次域运算一行的速度运行

另有两处改动,让父节点验证递归分片的成本更低。它的内存列和见证列以至多 2^24 个求值的堆叠形式承诺,所以父节点只需折叠少数几个点,而不是数百个。此外,递归请求写回的帧基地址会越过该帧向前推进,所以首尾相接排布的帧可以作为首尾相接的 ecall 重放,每次调用一行。

Tape#

对给定的电路族和高度,一个分片的检查具有固定的形状。所以宿主程序(host)把它们一次性编译成一条 tape:一个作用于绝对单元地址的直线式协处理器调用列表,其中没有任何东西根据值来分支。分片的 tape 与原生验证者对该分片执行的步骤逐个调用对应:分片 transcript、GKR 反向过程、查找(lookup)与根的检查,以及 Mercury 打开的十二个标量。每项检查都是一次断言相等。

每个递归程序的 tape、折叠模板和常量,都在编译时由验证者 crate 自己构建,并放进程序的只读数据中。因此,程序身份绑定了该程序重放的每一条 tape:证明一个节点运行了它的程序,就是证明它恰好运行了这些检查。

贯穿整棵树的 transcript 链#

基础陈述的全局 transcript 是覆盖整个陈述的一个海绵。递归树把它拆开,但不改变它。持有分片 0 的节点运行前缀部分,直到公开输入摘要为止;每个节点从其前驱留下的状态继续,吸收它自己那些分片的内存承诺;持有最后一个分片的节点运行后缀部分,并抽取内存挑战,而这些挑战此前在每个节点中都被当作断言接受。节点的公开输出(journal)记录链在其区间两端的状态,父节点要求其各个子节点的状态首尾相接。

节点还要求它的各个子节点彼此一致:退出状态为 0;同一个基础陈述(其形状、摘要、挑战、输入与公开输出的摘要、退出状态和分片数);相邻的分片区间;首尾相接的链状态;跨越接缝处按顺序排列的时间窗口;以及递归程序的程序身份。持有整个陈述的节点完成内存论证。

折叠配对#

没有任何节点计算配对。每个分片的 Mercury 检查都被延迟为十二个 (side, scalar, point) 条目;在分片的 tape 之后,节点自己的 transcript 吸收分片 transcript 的最终状态并抽取权重,每个条目经加权后累加进一对持续更新的点 (A, B),它代表断言 e(A, [1]_2) = e(B, [x]_2)。把分片的组合承诺与其各列联系起来的批量检查,也在旁边一并折叠。一个电路族所有分片共享的点,例如 [1]_1 和设置承诺,各自只累积一个标量,并且只加入一次。子节点的 (A, B) 以一个在其完整公开输出之后抽取的权重加入。

每一侧都是在 FQ_OP 上执行的一次多标量乘法,以静态模板运行:基于 GLV 分解后的两半、采用 8 位数位的 Pippenger 算法,每个点都被约束在曲线上,每一步都事先固定。每个点的成本约为 400 次 FQ_OP 调用。

到了根节点,整棵树的全部内容都已坍缩:每个基础分片都已验证,transcript 已从头到尾运行完毕,内存论证已经完成,每次打开都已折叠成一个配对断言。剩下的只有这个断言和两个程序身份。

判定器#

根节点仍然是一份 GKR 证明加上数百个点,合约无法检查。判定器是一个 Groth16 电路:它通过一个写出秩 1 约束(而不是协处理器调用)的驱动程序,对唯一的子节点,即根节点,运行节点流程,并要求根节点的公开输出对应全部基础分片区间。它不折叠任何东西:根节点欠最终配对的每个点连同其标量,都成为一条绑定线,即由验证者持有的一个值,在证明中以第五个陷门承诺,而不是作为公开输入传递。两个程序身份、基础退出状态,以及逐字节的基础公开输入和公开输出,也都是这样处理的。

远地虚拟机的 Groth16 与教科书版本有三处不同:绑定线承诺、没有盲化,以及证明密钥建立在 powers-of-tau 仪式已经公布的 Lagrange 基之上。它的密钥来自一个两阶段仪式:第一阶段就是承诺所依据的同一个仪式文件;第二阶段专属于这个电路,对 α 和 β 的贡献要在对 γ、δ 和 η 的任何贡献之前完成,而这个顺序本身就是可靠性的一部分。

ApogeeVerifier.sol 根据 calldata 重建被绑定的值,检查 Groth16 等式,用 ecMul 和 ecAdd 折叠两侧的点(这同时要求每个点都在曲线上),然后检查剩下的那一次配对。它的构造函数固定密钥、仪式的两个 G2 点以及两个递归程序的程序身份。一次部署服务于一个基础程序、一种根的形状和固定的公开值长度。

实测数据#

第 257,510 号区块;递归树在一台 32 CPU 的机器上运行,仪式和判定器在一台 18 核笔记本电脑上运行:

基础证明 207 个分片,14.5 MB,2,481 s
递归树 4 个叶节点(每个至多 64 个基础分片)和一个根节点:共 116 个分片
叶节点,四个同时进行 21、24、23 和 27 个分片;2,157 s;峰值 92 GiB
根节点 21 个分片,460 s,1.03 MB
判定器 7,896,686 个约束;证明耗时 18.5 s,占用 6.1 GB
合约 358 个点;3,620,026 gas;34,980 字节 calldata

规范见递归与判定器。亲自运行一遍:链上结算。

系统架构

以太坊区块

远地虚拟机的参考工作负载。在虚拟机内运行以太坊区块的 revm 客户程序、与 zkEVM 测试发布版本的每个用例都一致的无状态校验器,以及一个区块的证明说明了什么。

以 Markdown 查看

远地虚拟机证明任意的 RV32IMAC 程序。它的参考工作负载,也就是测量和调优所依据的那个,是常见工作负载中最难的一个:在虚拟机内用 revm(Rust 实现的 EVM)校验以太坊区块。同一个库 revm_block 既为客户程序(guest)编译,也为宿主机(host)编译,两个二进制程序证明两个不同的陈述。

两个二进制程序,两个陈述#

二进制程序 证明者提示(advice) 公开输出(journal) 含义
revm-block 一个 BlockWitness:交易所读取的前置状态 每笔交易的状态、gas 和返回数据;一个日志承诺;一份后置状态摘要 存在某个规范的见证,使 revm_block::run 产生这份公开输出:这是对一次执行的证明,而不是对区块有效性的证明
revm-block-stateless zkEVM 基准格式的无状态输入 43 字节:载荷的根、判定结果、链 ID、schema ID 具有此根的载荷,在这条链上、在这个分叉下,是或不是一个有效区块

真正证明区块的是无状态校验器。它实现了以太坊执行规范中的 verify_stateless_new_payload:解码请求,对照父区块检查祖先区块头和区块头规则,恢复每笔交易的发送者,在一个通过哈希约束到父区块状态根的前置状态上执行每笔交易,处理提款和请求,并重新计算收据根、bloom、gas、请求哈希、区块访问列表和后置状态根。见证本身不需要额外的绑定:公布的根固定了载荷,见证通过哈希受这个根约束,所以错误的见证无法让无效的载荷变得有效。判定结果为 false,只说明这份输入没有通过校验。

分叉与一致性#

校验器根据输入的 schema ID 确定分叉,没有编译进任何激活时间表:Osaka、BPO1、BPO2 和 Amsterdam。tests-zkevm v21.0.1 发布版本的全部 67,251 个测试对在原生运行下一致;CI 用一个已提交的、包含 34 个用例的子集约束这个库,该子集在两种输入布局下覆盖了该发布版本涉及的每条规则。

实践中的委托#

两个二进制程序都声明了 KECCAK_F、SHA256_COMP、MOD_MUL 和 EC_ADD。Keccak 经由 alloy-primitives 的 native-keccak 钩子用到其电路。SHA-256、secp256k1 和 BN254 经由打过补丁的 revm-precompile、k256 和 ark-ff 副本用到各自的电路,每个副本都以上游代码作为回退路径。每个发送者都在客户程序中按 EIP-2 的规则恢复,所用的 k256 算术被补丁路由到 MOD_MUL 和 EC_ADD。

这两个二进制程序以 --release 构建,证明时每个高度可选的电路族都使用 2^20:无状态二进制程序的 .text 约为 1.96 MB,占 2^20 表覆盖范围的 96.6%。

实测区块#

glamsterdam-devnet-8 的第 257,510 号区块,经由 revm-block-stateless:60 笔交易,101.5 Mgas,1.98 亿个周期,切分为 207 个分片。基础证明在一台 32 vCPU、247.7 GiB 内存的机器上耗时 2,481 s,内存峰值 174 GiB;递归另外花费约 2,620 s;合约以 3,620,026 gas 接受了结果。性能逐一拆解了每个阶段。

区块证明不做什么#

  • 校验器的输入来自外部的见证生成方。 eth_getProof 返回每个键路径上的 trie 节点,但使某个分支折叠的删除操作,需要一个不在任何已改变键的路径上的兄弟节点。因此仓库中的记录器无法生成无状态输入;这些输入来自 tests-zkevm 发布版本或 zkEVM 基准测试的数据集。
  • mini-block 二进制程序证明的是在一个已记录的前置状态上的一次执行,通常是某个区块的前几笔交易,不对状态根作任何断言。它的公开输出每笔交易增加一条记录,最终会超出公开窗口,这就是完整区块要走无状态校验器的原因。

规范见以太坊区块。

系统架构

安全模型

证明确立什么、依赖哪些假设、验证者必须自己持有什么、可靠性依赖哪些代码,以及 1.0.0 版的限制。

以 Markdown 查看

证明确立什么#

一份通过验证的证明确立的是:具有给定身份的程序,在其映像上从入口 pc 启动,输入窗口中存放该公开输入,并带有某份由证明者选择的证明者提示(advice),逐条指令执行直至以给定状态调用 EXIT,且已写出给定的公开输出(journal)。对证明者提示不作任何断言。没有任何东西被隐藏。

假设#

假设 在何处起作用
Mercury 与 KZG 在代数群模型中、q-DLOG 假设下的知识可靠性 每一次承诺打开
在 Fiat–Shamir 中把 Poseidon2 视为随机预言机 每个挑战,包括基础证明和递归中的挑战
Groth16 自身的假设 判定器,即通往合约的最后一步
PSE 的 perpetual powers of tau 有一位诚实的贡献者 每个承诺所依据的 SRS
判定器第二阶段仪式的每一轮都有一位诚实的贡献者 判定器的密钥

BN254 提供约 100 位的安全性。每一层协议的统计误差,从求和校验(sumcheck)和批量打开,到内存论证与查找(lookup)论证,都远低于此:一个电路的整个 GKR 过程低于 2^14/|Fr|,每个查找通道低于 2^−190,每个 Mercury 实例低于 2^−220。

验证者必须自己持有的值#

两个值,都要取自证明者无法控制的渠道:

  • 程序身份。 如果对照的是证明者提供的身份,证明只能说明有某个程序运行过。
  • 仪式的 SRS 摘要。 不论密钥自身的点给出什么摘要,密钥都会按这个摘要加载,所以基于已知 τ 构建的密钥,只有通过这项比对才会被拒绝。

验证密钥本身可以来自任何人。加载它时会根据其自身内容重新计算程序身份和 SRS 摘要,并要求其中的电路与验证者的注册表一致:程序身份绑定程序,注册表绑定电路。verifier 命令行工具只比较程序身份,SRS 摘要则取自密钥;host::verify 两者都不比较,并且把检查陈述中的输入、公开输出和退出状态的工作也留给调用方。

受信任的代码#

可靠性只取决于验证者。它依赖于 constants、field、curve、transcript、poly、sumcheck、pcs-verify、pcs、gkr-verify、verifier-core、verifier,以及 constraints,因为电路是陈述的一部分,缺失一个门就是可靠性漏洞。从 ELF 计算程序身份还要信任 loader、isa 和 program。上链的最后一步又加入了递归程序、groth16、判定器的电路和合约。

证明者、模拟器、执行轨迹构建器,以及宿主程序(host)SDK 中负责证明的一半,都是不受信任的。证明者不做任何校验;错误的输入只会让诚实的证明者得到一份无法通过验证的证明,而作弊的证明者本来就不会运行这些代码。

不是常数时间的,也不是零知识的#

代码中没有任何部分是常数时间的:约简、幂运算、点加和标量阶梯都会根据操作数分支。这在这里无害,因为没有任何证明是零知识的,证明过程不保有任何秘密。代码处理的唯一秘密,是判定器仪式贡献者的因子,它经过的是同一个非常数时间的阶梯;请在你自己控制的机器上进行贡献。

v1.0.0 的限制#

限制 说明
不是零知识的 Mercury、GKR 和判定器中都没有盲化
证明者提示不受绑定 客户程序(guest)要将它与证明所绑定的某样东西进行核对
公开值 输入和公开输出各至多 16,380 字节
sc.w 总是成功 这是与 RV32IMAC 语义唯一的偏差;没有保留(reservation)状态
陷入(trap)不可证明 未对齐的访问、对映射内存之外的访问、ebreak,或者 pc 处没有指令,都会使执行终止,且不产生证明
代码是静态的 指令流就是加载时解码的映像;可执行代码中只要有一个无法解码的字,程序就会被拒绝
代码大小 .text 须在解码表的覆盖范围之内,2^22 时为 7.94 MiB;映像默认不超过 4 MiB
执行长度 2^36 − 1 个周期
委托集合是固定的 基础格式中有六种;一个委托证明其函数的一步,把各步组合起来(包括校验曲线点)是调用代码的责任
证明者内存 由同时处理中的分片决定:实测区块的峰值为 174 GiB
区块见证 无状态校验器的输入来自外部的见证生成方
判定器的密钥 每种根的形状一个,其可信程度取决于它的仪式;开发用密钥是可伪造的
链上成本 实测区块约 3.6M gas

下一个版本将改变什么#

上面涉及 BN254 的每一个假设,从 q-DLOG、配对到 Groth16,都会被一台足够大的量子计算机攻破。v2.0.0 的演进方向是一个可靠性改为建立在格问题之上的证明核心。

从审计者的视角,逐个 crate、逐个论证地审视同一个模型:审计指南、可靠性地图。

系统架构

性能

v1.0.0 的每一项实测数据及其来源:一个完整以太坊区块的基础证明、递归树、判定器与合约、每个电路族的分片证明,以及 Mercury 本身的成本。

以 Markdown 查看

所有端到端数据都来自 glamsterdam-devnet-8 的第 257,510 号区块,经由无状态校验器客户程序(guest)证明:60 笔交易,101.5 Mgas,1.98 亿个周期。每一项数据都来自规范中的测量。

端到端#

阶段 机器 结果
基础证明 32 vCPU,247.7 GiB,同时处理 12 个分片 207 个分片,14.5 MB,2,481 s;峰值 173.92 GiB
递归叶节点 32 CPU,四个叶节点同时进行 21、24、23 和 27 个分片;2,157 s;峰值 92 GiB
递归根节点 同一台机器,同时处理四个分片 21 个分片,460 s,1.03 MB
判定器仪式 18 核笔记本电脑 init 65 s;每份贡献 50–56 s;key 70 s、12.7 GB;密钥 2.65 GB
判定器证明 18 核笔记本电脑 读取密钥 1 s,证明 18.5 s,6.1 GB;7,896,686 个约束,定义域大小为 2^23
链上验证 revm 3,620,026 gas;34,980 字节 calldata;358 个点

基础证明,逐遍拆解#

第一遍,承诺 191 s;平均 25.7 个 vCPU 处于忙碌;单线程填充占其分片秒数的 81%;采样内存至多 15.9 GiB
第二遍,证明 2,290 s;客户程序退出之前,12 个分片名额中平均有 11.95 个被占用,30.4 个 vCPU 处于忙碌;随后是 460 s 的尾段,其中最长的两段是两个 KECCAK_F 分片的单线程填充,分别为 200 s 和 279 s
内存峰值 173.92 GiB:两个 2^18 的 KECCAK_F 分片,在尾段同时处理,此外再无其他分片在处理中

峰值由某一个委托电路族的高度决定,而不是由同时处理中的分片数决定。

一个小型客户程序#

快速上手中的客户程序,在四个指令电路族中运行 114 个周期,以 2^20 的指令高度和 2^16 的窗口高度、同时处理两个分片进行证明:在一台 18 核、48 GiB 内存的笔记本电脑上,共 7 个分片,耗时 52 s,峰值 18 GB,几乎全部来自正在处理的两个 2^20 分片。证明的下限由其电路族和高度决定,而不是由周期数决定。

每个电路族的分片#

以下均为默认高度下的数据。分片证明的大小由其电路的形状和高度确定;证明成本取决于高度与电路宽度的乘积,不论有多少行是有效行。

电路族 高度 已承诺 M/W/S 约束门 内部列 分片证明
ADD_SUB_LUI_AUIPC 2^22 27 / 35 / 7 63 314 64,764 B
JUMP_BRANCH_SLT 2^22 21 / 44 / 10 42 392 69,436 B
SHIFT_BITWISE 2^22 21 / 61 / 10 48 478 76,644 B
MUL_DIV 2^20 21 / 54 / 9 54 444 67,412 B
MEM_WORD 2^22 31 / 24 / 7 33 314 63,836 B
MEM_SUBWORD 2^22 31 / 55 / 10 53 472 76,196 B
ATOMICS 2^20 26 / 54 / 9 46 472 68,468 B
INIT_TEARDOWN 2^22 2 / 0 / 1 0 46 36,316 B
ZERO_WINDOWS 2^22 2 / 0 / 0 0 46 36,284 B
KECCAK_F 2^18 208 / 1,556 / 0 385 5,490 381,100 B
POSEIDON2 2^8 100 / 4,092 / 0 4,248 2,020 664,780 B
FR_ARITH 2^8 104 / 2,576 / 0 2,701 142 266,292 B
PUBLIC_INPUT、PUBLIC_OUTPUT 2^12 3 或 2 / 0 / 0 0 26 12,556 B、12,524 B
ADVICE_WINDOWS 2^22 3 / 0 / 0 0 46 36,316 B
MOD_MUL 2^16 104 / 221 / 0 125 2,244 135,220 B
SHA256_COMP 2^18 104 / 520 / 0 119 2,802 189,988 B
EC_ADD 2^16 392 / 1,028 / 0 637 8,772 434,916 B

高度只改变折半列表的数量和求和校验(sumcheck)的轮数,不改变门:ADD_SUB_LUI_AUIPC 在 2^20 时有 298 个内部列,证明为 57,196 字节,而在 2^22 时分别为 314 个和 64,764 字节。

前向过程的内存#

GKR 证明者以域元素的形式持有每个内部层,每个单元 32 字节。代表性的工作集:

分片 前向过程
2^20 的 SHIFT_BITWISE 8.4 GiB
2^16 的 MOD_MUL 4.6 GB
2^16 的 EC_ADD 18.3 GB
2^18 的 SHA256_COMP 22.6 GB
2^18 的 KECCAK_F 42 GiB

Mercury#

在一台 18 核的 Apple M5 Pro 上:

承诺,n = 2^22 1.30 s
打开,n = 2^22 2.89 s
16 个 2^20 的列作为一批 打开耗时 1.01 s,验证耗时 4.8 ms
同样 16 列逐个打开 9.79 s,验证耗时 62 ms

如何解读这些数字#

证明受内存限制,其内存取决于同时处理中的分片及其高度,从不取决于执行的长度。时间取决于各电路族各自的周期数。链上成本取决于根节点欠最终配对的点数,每个点约 9,000 gas。周期数本身是精确的,且与机器无关,所以要估算其余任何一项,周期分析器都是首选工具。

量子跃迁

量子跃迁

远地虚拟机的下一步。通往 v2.0.0 的轨迹:基于格的证明核心、与之匹配的域、客户程序可以验证的签名,以及一个把应用从代码仓库带到运行中的链的部署系统。

以 Markdown 查看
简报项目:远地虚拟机目标:v2.0.0状态:积极开发中发布窗口:

1.0.0 版回答了这套架构能否在完整规模下成立的问题:一个完整的以太坊区块,从 Rust 客户程序(guest)一路到返回 true 的合约。2.0.0 版改变的是这套架构的根基及其服务对象:证明核心迁移到量子计算机无法攻破的数学之上,面向开发者的部分则从一个代码仓库成长为一个部署应用的系统。

重要

本节描述的是进行中的工作。这里的任何内容都不改变 v1.0.0 的保证;这些保证完整列于安全模型。

四项举措#

演进轨迹#

v1.0.0,当下 v2.0.0,演进方向
承诺 基于 KZG 的 Mercury:配对,q-DLOG 基于格,在 Module-SIS 下具有绑定性
域 BN254 的标量域,254 位 小素数域,挑战取自扩域,与承诺相匹配
面对量子敌手 每个假设都是离散对数假设 证明核心建立在格问题之上
客户程序中的签名 secp256k1,经由委托的域运算与曲线运算 ZK 友好方案与后量子方案,作为客户程序调用
面向开发者 一个代码仓库、配套工具和这份手册 部署系统:门户、标准桥、遥测、AI 网关

延续下来的部分#

跃迁发生在根基上,而不在开发者编程所面对的模型上:

  • 客户程序。 Rust、RISC-V、三个内存区域、一个程序身份、一份公开输出(journal)。为 v1.0.0 编写的程序保持原有形态。
  • GKR 引擎。 分层电路和求和校验(sumcheck)在任何域上都有定义。这个把整个电路收束到一个点上的引擎,是远地虚拟机中迁移到新域最直接的部分。
  • 论证。 覆盖整个执行过程的单一内存多重集,以及 LogUp 通道,都作为构造保留下来;它们的表和范围论证将按新域的大小重新推导。
  • 工程纪律。 规范先行;每一层都对照一个独立的参照(oracle)检查;每一类伪造都由一个篡改孪生守住。

为什么是现在#

有效性证明的抗量子能力,取决于产生它的系统。BN254 上的证明建立在配对和离散对数之上,因此一台足够大的量子计算机可以伪造证明,而无须触碰客户程序计算过的任何东西。区块链原生应用要承载数十年的价值。它们赖以结算的根基,必须比那些终有一天会攻破今天这些曲线的机器更长久;而迁移根基的时机,是在那些机器出现之前。

阅读各项举措:后量子证明 · 面向客户程序的签名 · 部署系统。

量子跃迁

后量子证明

举措 QL-01 与 QL-02。用格承诺取代基于配对的承诺,并选用与之匹配的域,使证明核心不再依赖离散对数。

以 Markdown 查看
QL-01 · QL-02格承诺换域状态:积极开发中

什么会被攻破,又在哪里#

远地虚拟机 v1.0.0 之下的每一个密码学假设都涉及 BN254。Mercury 和 KZG 在代数群模型中、q-DLOG 假设下是可靠的;递归树折叠的是配对检查;判定器是 Groth16。在规模足够的量子计算机上,Shor 算法能求解离散对数,上述每一个假设也随之失效。客户程序(guest)的计算依然如故,但证明它正确运行的那份证明将不再有任何意义。

这种依赖集中在承诺上。每个分片的每一列都用它来承诺,每次打开都以它的配对收尾,而递归的存在正是为了折叠这些配对。换掉承诺,证明核心的其余部分就再没有任何依赖离散对数的东西。

QL-01 · 格承诺#

格承诺是一个线性映射 t = A·s mod q,作用于一个元素很小的向量 s。只要没有人能找到一个被该矩阵映射为零的短向量,它就具有绑定性:这就是 Module-SIS,ML-DSA 和 ML-KEM 这两项 NIST 后量子标准所依赖的假设族,背后有最坏情况归约和数十年的密码分析支撑。

它保留了 KZG 对远地虚拟机最有用、而基于哈希的承诺所放弃的那一点:它是同态的。 对多个数据块的承诺可以用挑战系数组合起来,组合结果只需一个向量即可打开。批量处理一个分片的各列、延迟检查、沿树向上折叠,都是线性运算,而线性运算在迁移之后依然成立。Merkle 路径则根本无法组合。

代价是一种别处没有的约束:承诺只绑定短向量,每次组合都会让向量变长,证明者必须证明它仍然足够短。近两年的各种方案,主要区别就在于如何支付这一代价,而且进展很快。对于 2^30 个系数的多项式,已发表的 Module-SIS 方案给出的求值证明为 53 至 72 KB,验证时间则从 2024 年的 2.8 秒降到 2026 年的 8 至 16 毫秒。专题文章 Lattice-Based Polynomial Commitment Schemes 逐一梳理了这些方案,并把这些数字与基于哈希的一方对照解读。

QL-02 · 换域#

格方案并不属于 BN254 的世界。主流构造工作在小素数模数上,为保证可靠性,求值点取自扩域;这适合小域证明系统,而不适合 254 位的系统。因此承诺的迁移会把域一并带走:v2.0.0 为算术化换域,把每个电路从 BN254 的标量域迁移到与承诺相匹配的小域上。

这次迁移本身就能收回成本:

  • 每一层都更便宜。 GKR 证明者的时间都花在域运算上,而小域中的一次乘法只是 254 位域中一次乘法的零头。该引擎最核心的节省在于中间层从不承诺;这一点与剩下每一层都更便宜的运算叠加在一起。
  • 承诺更便宜。 承诺一个执行轨迹列,是对一组小取值数位做一次线性映射,成本按非零项计算,而不是在曲线上做多标量乘法。
  • 引擎原样延续。 GKR 和求和校验(sumcheck)在任何域上都有定义。挑战移到扩域上;反向过程、层模型以及建立在其上的各项论证都保持原有结构。

需要重建的,是一切以大域为前提的部分:能宽裕地放进一个 BN254 元素的字级数值、按 254 位模数确定大小的范围论证和进位、规范性链(canonicity chain),以及递归格式中的域单元。每一项都会针对新域重新推导,并像 v1.0.0 的那些部分一样写成规范,配有自己的参照(oracle)和篡改孪生。

结算#

以太坊现有的验证预编译合约都基于配对。后量子证明核心如何在这条链上结算,最后一步中有哪些部分可以建立在格之上,属于同一项工作计划,并会在发布之前以与核心同等的严谨程度写成规范。

返回任务简报。

量子跃迁

面向客户程序的签名

举措 QL-03。每个客户程序都能以一次调用完成 ZK 友好签名与后量子签名的验证,让区块链原生应用中的授权只需一行代码。

以 Markdown 查看
QL-03面向客户程序的签名状态:积极开发中

几乎每个区块链原生应用都会在每次请求时问同一个问题:这是正确的密钥授权的吗?在 v1.0.0 中,客户程序(guest)用代码来回答。secp256k1 公钥恢复由 k256 在委托的 MOD_MUL 与 EC_ADD 电路之上完成,这让以太坊自身的签名变得负担得起;其他任何方案都只能用普通指令执行。QL-03 让签名验证成为客户程序的一等操作。

签名方案#

一个签名方案证明起来便宜与否,几乎完全取决于它的验证算法:做什么运算、在哪个域上做、调用哪个哈希。签名者从不在证明内部运行。以下四种设计覆盖了整个设计空间:

方案 思路 对客户程序的意义
原生曲线上的 Schnorr 在一条基域恰为证明系统自身所用域的曲线上运行 Schnorr 协议,如 Grumpkin 之于 BN254 从构造上就最便宜:运算和哈希都是电路原生的
ML-DSA(FIPS 204) 把 Schnorr 搬到格上,借助拒绝采样使短响应保持均匀分布 NIST 的首要后量子签名;成本主要来自其哈希,以及(除非密钥固定)其矩阵扩展
FN-DSA(Falcon) 基于格陷门的“哈希后签名”,陷门由高斯采样隐藏 三种后量子方案中运算最少、哈希也最少
SLH-DSA(FIPS 205) 仅由哈希函数构造的签名 完全没有代数运算,约两千次哈希调用;假设最为保守

四种验证算法都是“计算后比较”,不涉及秘密,也不依据秘密做分支,正因如此每一种都可以被证明。专题文章 ZK-Friendly Signature Schemes 用玩具示例逐一推演,并比较了它们在电路中的成本。

成本由什么决定#

有两个杠杆决定了每一个数字:

  • 证明者所用的域。 一个方案的运算就是电路的运算时,它才是原生的。因此哪个方案最便宜,取决于 QL-02 的换域,方案的选择也将与之一同做出。
  • 哈希。 后量子验证算法的成本主要来自其标准哈希,而不是代数运算。换成算术哈希就偏离了标准,这是面向零知识的构造所选择的变体;保留标准哈希,则是与现有密钥互操作的前提。两者各有用武之地,由客户程序决定。

面向开发者#

目标是让客户程序验证授权,就像今天计算哈希一样:一次调用,委托给电路,应用自己的代码里没有任何密码学。这将打开区块链原生应用赖以构建的各种模式:所用密钥并非链上原生密钥的账户、状态转换函数内部的多方审批、会话密钥,以及在其诞生所依托的曲线被攻破之后依然有效的身份。

返回任务简报。

量子跃迁

部署系统

举措 QL-04。远地区块链原生化部署系统:一套门户与工具链,把应用从源代码带到运行中的区块链原生环境。

以 Markdown 查看
QL-04远地区块链原生化部署系统状态:积极开发中

1.0.0 版交给开发者的是一个代码仓库、配套工具和这份手册。从一个已证明的客户程序(guest)到一个上线运行的应用,中间的一切,从密钥和仪式到验证者合约、桥和运维,仍要由开发者自己组装。远地区块链原生化部署系统是这次跃迁的后半程:一套门户与工具链,把客户程序变成运行中的区块链原生环境,并让它持续运行。

模块#

控制台进行中

所有部署资源,集中一处

程序及其程序身份、高度与参数、验证密钥、判定器密钥及其仪式、验证者合约及其所在网络,按应用和版本分别组织。

桥进行中

标准链上合约

把每个应用今天都在重复构建的东西做成可复用模板:由证明驱动的状态根注册表、存款与提款桥、从一个程序身份迁移到下一个程序身份的升级路径。

遥测进行中

应用的生命体征

已生成和已结算的证明、每个请求的周期数、证明延迟与成本、分片与电路族的性能剖析、验证所耗的 gas,以及应用状态根的历史。

网关进行中

为你自己的 AI 开一扇门

一个接口:在开发者的掌控之下,开发者本地的 AI 智能体可以通过它检查应用、查询遥测数据、经由部署流水线提出并执行变更,并读回每一个结果。

模板计划中

从可用形态起步的应用

位于仓库之外的客户程序项目,profile、链接器设置和 vendored crate 都已配置妥当,AI 随行手册也已就位。

仪式计划中

仪式即服务,而非苦差

协调判定器密钥的第二阶段贡献,每一份贡献都可以对照电路和仪式文件验证,使一次部署的密钥拥有其可靠性所需的诚实贡献者。

为什么是一个系统,而不是更多工具#

远地虚拟机背后的核心理念,是每个经济应用一个经过优化的环境。这让环境的数量成倍增长,运维面也随之扩大:每个环境都有自己的程序、密钥、仪式、合约和指标。要求每个团队手工组装这一切的做法,无法扩展到这一理念所需要的众多环境。部署系统让每个环境都成为例行公事,于是发布一个区块链原生应用时,难点只剩下应用本身。

网关也出于同样的考虑。开发者已经在与 AI 模型协同工作,AI 随行手册向这些模型讲解如何编写客户程序。网关则在开发者的掌控之下,让这些模型能够对自己写出的东西采取行动:部署、观察、迭代。

返回任务简报。

审计专区

审计指南

审计远地虚拟机 v1.0.0 起步所需的一切:审计范围、规范原文及其组织方式、记号约定、信任边界、阅读顺序,以及最值得优先检查的性质。

以 Markdown 查看

本部分是远地虚拟机 v1.0.0 的完整构造,按评估的需要组织。其核心是逐字转载的规范原文:每个主题一页,列出每个已承诺列的索引与名称、写成多项式的每个门、每个查找及其通道、按顺序排列的每条 transcript 消息,以及每种序列化格式的每一个字节。围绕它,本指南和可靠性地图为审计者提供入口。

审计范围#

审计范围 位置
证明所确立的陈述,以及验证者必须持有的内容 系统全貌、证明
算术:Fr、Fq 扩域塔、G1 与 G2、配对、MSM、多线性多项式、求和校验(sumcheck) 原语
Fiat–Shamir:Poseidon2、双工海绵、所有标签 Transcript
可信设置与承诺方案 结构化参考串、Mercury
程序:加载、解码、表、配置、程序身份 程序与身份、客户程序(guest)ABI
执行模型与执行轨迹 执行轨迹、公开值与证明者提示(advice)
证明系统:GKR、电路注册表、内存论证、查找(lookup) GKR 引擎、电路、内存论证、查找
每个电路族,逐列展开 七个指令电路族和六个委托电路
证明者的结构 流式证明者
递归、Groth16 判定器及其仪式、合约 递归与判定器
以太坊工作负载 以太坊区块

规范页面转载自远地虚拟机代码仓库在源码修订版 3571370 时的 docs/:相对链接改成了本站链接,每个 § 引用都改成了指向对应章节的链接。术语表中有一处措辞为与本站其余部分保持一致而改写;除此之外没有任何改动。规范页面与代码不一致时,以代码为准,而这种不一致本身就是一项审计发现。

如何阅读规范#

这些页面是为对照代码阅读而写的。每一页都写明实现其所述内容的 crate 和函数,源码注释也按章节反向引用规范(docs/spec/memory.md §2.4)。以下是反复出现的一些约定:

记号 含义
M[i]、W[i]、S[i] 电路的已承诺列:内存列(在全局 transcript 中、内存挑战之前绑定)、见证列(在分片自己的 transcript 中绑定)、设置列(由程序身份或 SRS 摘要绑定)
V[…] 虚拟表:行索引的闭式表达式,从不承诺
L{k}[j]、C{k}[j]、scratch[i] 第 k 层的内部列 j;缓存项;扁平关系的中间值
W[8..14] 列索引的半开区间,即 W[8] 到 W[13]
T(AS, ADDR, TS, VAL) 内存元组,γ_M + AS + α_addr·ADDR + α_ts·TS + α_val·VAL
4c + Δ 周期 c 中槽位 Δ 的时间戳
G1–G11、S1–S6 全局 transcript 与分片 transcript 的步骤
步骤 1–12、B1–B6 验证者依次对分片和块(block)所做的检查
2^n 2 的幂;高度取 2^8, 2^12, 2^16, 2^18, 2^20, 2^22

写成表达式的门,约束其值为 0。查找写成它的通道、选择子和元组。“有界”指经过范围检查;称为字的值是 [0, 2^32) 中的整数。

阅读顺序#

第一遍阅读,先建立完整的论证,再深入各个电路:

  1. 系统全貌。 断言、组合表、假设、限制。
  2. 证明。 陈述、两种 transcript、验证顺序、密钥及其加载规则。
  3. GKR 引擎。 层模型、制品及其法则、反向过程及其可靠的理由。
  4. 内存论证 与 查找。 一切跨行内容所依赖的两种论证,以及它们的构造期规则。
  5. 电路。 注册表、各电路的形状,以及电路族的电路如何组装。
  6. 指令电路族,从 ADD_SUB_LUI_AUIPC 开始,它承载每一个 ecall 以及每次委托的请求一侧。
  7. 委托 ABI 与 委托电路。
  8. 程序与身份、公开值、执行轨迹。
  9. Transcript、SRS、Mercury、原语。
  10. 递归与判定器,然后是合约。

信任边界#

可靠性只取决于验证者,而验证者的代码是一组明确界定的 crate:

受信任的用途 Crate
验证一个块 constants、field、curve、transcript、poly、sumcheck、pcs-verify、pcs、gkr-verify、verifier-core、verifier,以及 constraints,因为电路是陈述的一部分,缺失一个门就是可靠性漏洞
从 ELF 计算程序身份 loader、isa、program
上链的最后一步 guests/recursion、groth16、判定器的电路、contracts/ApogeeVerifier.sol
不受信任 prover、emulator、trace、host 中负责证明的一半:证明者不做任何校验

所依赖的假设是:Mercury 与 KZG 在代数群模型中、q-DLOG 假设下的知识可靠性;Poseidon2 作为随机预言机;最后一步依赖 Groth16 自身的假设;每个仪式有一位诚实的贡献者。没有任何部分是常数时间的,也没有任何证明是零知识的。验证者必须从证明者无法控制的渠道获得程序身份和仪式的 SRS 摘要。

优先检查之处#

下列性质一旦失效就意味着可以伪造,表中给出了每一项的论证位置。可靠性地图以同样的方式梳理了陈述中的每一项断言。

性质 论证位置
每个挑战都在其所保护的一切之后抽取:内存挑战在每个 M 承诺、窗口列表、io_digest 和 64 个边界标量之后;g 与 β 在该分片的 W 承诺之后 proof §2、§4;memory §6.1
任何内存元组或根都不读取 W 列,这类列在内存挑战之后才承诺 memory §8
帧只约束其掩码的布尔性;每个电路族都把每个掩码固定为 m_pc 乘以会发起该查询的那些 kind memory §2.1;各电路族页面
表通道所查找的每个键都由其电路族限定范围,而每个通过 copower 写出的界也同时带有直接界 lookup §4、§11;shift-bitwise §3
从帧字中解码出的选择子是 one-hot 的,因为编码会相加 delegation circuits §1
每个委托请求恰好与一次调用配对 delegation §5
只有退出行能写入 HALT_PC;凡是由电路族计算 next_pc 的地方,都将其约束为偶数 memory §5;jump-branch-slt §5
每个地址只有一个初始值:窗口规则 memory §3.5、§9
输入窗口与公开输出(journal)窗口存放的是陈述中的字节;公开输出没有初始列 public values §5
打开从密钥中取得设置承诺,从而绑定程序身份所承诺的表与映像 proof §5;memory §6.2
密钥中的电路就是注册表中的电路,其 SRS 摘要与仪式的 SRS 摘要进行比对 proof §3、§7
递归:由程序身份绑定的 tape、transcript 链、在被加权对象之后抽取的折叠权重、判定器的绑定线,以及仪式各轮的顺序 recursion §7、§8、§9

复现#

CI 运行的所有内容都不需要仪式文件。证明真实分片的测试套件在各自的玩具 SRS 上运行,需要数十 GiB 内存,因此要按名称单独运行:

sh
cargo test --workspace                                   # every unit, law and row suite
cargo run -p kat-gen && git diff --exit-code             # fixtures regenerate identically
cargo test --release -p prover --test <suite> -- --include-ignored --test-threads=1
#   acceptance, control, alu, mem, fills, block, streaming, keccak, recursion, public_io, revm
cargo test --release -p checker --test tamper -- --include-ignored --test-threads=1   # every tamper twin
cargo run -p checker -- laws <artifact>                  # Laws 1–4 and the lookup rules, independently

验证实现说明了每个参照(oracle)和测试套件各自确立了什么,以及它们都没有确立什么。

报告发现#

请将审计发现发送至 admin@gweb3networks.com,并附上规范章节或代码路径、所涉及的性质,以及在可能的情况下附上一个篡改孪生:一份按诚实证明者的方式证明、却能通过验证的伪造见证。

审计专区

可靠性地图

一份通过验证的证明所做出的每一项断言、支撑它的论证,以及规范中阐述该论证并说明其成立理由的确切章节。

以 Markdown 查看

一个通过验证的块(block)确立的是一句话:具有此身份的程序,在其映像上从入口 pc 启动,以此公开输入和某份证明者提示(advice)逐条指令执行,直至以此状态调用 EXIT,并已写出此公开输出(journal)。 本页把这句话拆成组成它的各项断言,并把每一项对应到证明它的论证。

程序#

断言 论证 规范
密钥描述的是已登记的程序 加载密钥时,根据其自身的配置、入口 pc 和设置承诺重新计算程序身份;验证者将其与自己持有的副本比对 proof §7.2、program §8
证明读取的表就是程序身份所承诺的表 每个分片的批量打开都从密钥中取得设置承诺 proof §5
每个被执行的行,都是程序在该行 pc 处的指令 解码器查找:以该行自身读到的 pc 为键,查询一张有效行为 one-hot、填充行为 −1 的表 lookup §10、program §5、§6
内存从程序映像开始 INIT_TEARDOWN 的初始化列是程序身份所承诺的设置列;没有任何来自文件的字节位于窗口 0 之外 memory §6.2、§3.4
执行从入口 pc 开始 pc 的初始元组使用密钥中的入口 pc,而它由程序身份绑定 memory §4.2、§6.2
电路是正确的电路 密钥中的电路必须在其高度上与验证者的注册表一致,并通过各项法则、内存规则和履行规则(discharge rule) proof §7.2、circuits §1、gkr §4.2

每一行#

断言 论证 规范
每一行都遵循其指令 电路族的约束门(在每一行上均为零),以及各电路族的可靠性论证 add-sub §4、jump-branch-slt §5、shift-bitwise §5、mul-div §5、memory-ops §3–§6
每一行恰好发起其指令应有的查询 每个掩码都固定为 m_pc 乘以会发起该查询的那些 kind memory §2.1
x0 读写的都是 0 x0 gadget 与写回 memory §2.4
寄存器和 RAM 中的值都是字 每次寄存器写入都在其所在行上限定范围;执行电路族的每次 RAM 写入都是字;初始值都是字,证明者提示除外,而没有任何电路族依赖它 memory-ops §5
填充行不增加任何内存事件 填充行的每个掩码都是 0,因此它的叶子都是 1 memory §2.3、gkr §4.3

内存与顺序#

断言 论证 规范
每次读取都返回最近一次写入 覆盖所有分片的单一读/写多重集,与寄存器及 pc 边界做一次核对 memory §4、§9
读取严格晚于它所消费的那次写入 每个查询的时间戳差由两个 19 位的 TIMESTAMP 块组成 memory §2.4、§7
每个地址恰好有一个初始值 窗口规则:统一的高度、互不相交的窗口、每个固定窗口各占一个分片 memory §3.5、§9
多重集无法靠一个环闭合 时间戳是有界路径上的整数:形成一个环需要超过 2^215 条边 memory §4.2
各行按程序顺序构成一条从入口到出口的路径 pc 是一个内存单元,写入它的时间戳至少比读取它的晚四个 memory §5、§9
执行在退出行结束 HALT_PC = 1 是奇数;只有退出行写入它;JUMP_BRANCH_SLT 用范围检查确保其 next_pc 为偶数 memory §5、jump-branch-slt §5、add-sub §4
分片的时间窗口不提供任何额外保证 时间窗口只检查形状;顺序完全来自多重集 proof §8

取值与查找#

断言 论证 规范
每个带门控的元组都是其表中的一行 每个通道一个 LogUp 恒等式,在 GKR 过程中由分式树求和;根的分子为 0,分母非零 lookup §1、§6、§8
选择子是布尔值 每个查找的选择子都由门列表 0 中的一个约束门以 s − s² 加以约束 lookup §2
查找只由它自己的子表应答 每个通道一个宽度、带 +1 偏移且互不相交的键范围,以及各电路族对其键的界 lookup §4、§9、§11
表就是预期的那些表 虚拟表是验证者的闭式表达式;通用表由 SRS 摘要绑定;解码表由程序身份绑定 lookup §3、§12
每个已声明的义务都得到履行 在组装时和每次加载密钥时检查的履行规则 lookup §11

公开值#

断言 论证 规范
输入窗口存放的是陈述中的输入 在 G7 固定这些字节之后,PUBLIC_INPUT 的初始列在一个随机点上等于输入的各个字 public values §5
公开输出就是客户程序(guest)的存储指令留下的内容 PUBLIC_OUTPUT 的最终列等于公开输出的各个字;该窗口没有可供证明者填充的初始列 public values §5
退出状态就是 x10 的最终值 退出行写回它读到的 a0;验证者要求 v_10 等于陈述中的状态 add-sub §4、memory §4.1、proof §6

委托#

断言 论证 规范
每个请求恰好执行一次 锚点:请求与调用在该委托类型自己的空间中,通过多重集一一配对 delegation §5
每次调用都计算其函数 各电路的可靠性论证,包括帧与规范性链(canonicity chain) delegation circuits §2–§7
多次调用的操作是其各步骤的复合 RAM 粘合:在同一条内存历史上,每一步读取上一步的写入;顺序由调用方代码决定 delegation circuits §1

证明系统#

断言 论证 规范
分片的输出就是其电路在其已承诺列上的求值结果 GKR 反向过程,每个挑战都在其所保护的内容之后抽取 gkr §5.4
声称的列值就是已承诺多项式的值 在该过程所得的点上做一次批量 Mercury 打开 mercury §5、§7
一个陈述只有在其全部分片都通过时才算验证通过 解码时和 verify_block 中的分片集合精确性检查;核对等式读取每个分片的根 proof §1.3、§6
挑战在其所保护的每个承诺之后产生 全局 transcript G1–G11 与分片 transcript S1–S6 proof §2、§4、transcript §3
可信设置就是该仪式的设置 SRS 摘要,由验证者与仪式的 SRS 摘要比对 proof §3、srs §3

递归与合约#

断言 论证 规范
节点执行的恰好是基础验证者的检查 这些检查被编译成节点程序映像中的 tape,而该映像由节点程序的身份绑定 recursion §7、§8.1
递归树按顺序覆盖同一个基础陈述的每一个分片 跨节点的 transcript 链、相邻的分片区间,以及每个节点对其子节点的检查 recursion §8.1、§8.2
每个延迟打开都成立 每一个都以在其全部被加权对象之后抽取的权重折叠,最后由合约中的一次配对兑现 recursion §8.3、mercury §6
判定器绑定合约所持有的内容 绑定线在其挑战之前承诺;电路把公开输出约束到整个基础区间 recursion §9
判定器密钥没有已知的陷门 两阶段仪式,每轮有一位诚实贡献者,各轮按顺序进行 recursion §9、srs §7

有意不作的断言#

  • 关于证明者提示的任何断言。 证明者提示在设计上不绑定任何东西;由客户程序来检查它。
  • 零知识。 没有任何内容被盲化。
  • sc.w 的失败语义。 sc.w 总是成功;依赖其失败的程序不在断言范围之内。
  • 陷入(trap)。 发生陷入的运行根本没有证明。
  • 证明者是正确的。 证明者不受信任;只有验证者的 crate 承担可靠性。
  • 仪式文件确实出自该仪式,或者密钥的 τ 无人知晓:没有验证者自己对 SRS 摘要的比对,这两点都不在断言之内。

审计专区

验证实现

代码如何对照自身以外的东西接受检查:每一层的独立参照(oracle)、电路规则的第二套实现、证明伪造会被拒绝的篡改孪生,以及任何检查都没有覆盖的部分。

以 Markdown 查看

远地虚拟机的任何组件都不是对照整个系统的第二套实现来检查的。相反,每一层都有自己的参照(oracle),其选择原则是让检查与被检查的对象共享尽可能少的代码。

每一层及其参照#

层 对照的参照
域、曲线、配对、MSM 由 arkworks 生成的已知答案向量,测试也会实时运行 arkworks 加以对照;测试会重新推导各 crate 读取的每一个算术常量
Poseidon2 与 transcript tools/transcript-ref:使用 zkhash 轮常数的 Plonky3 Poseidon2,以及一份按规范转写的实现,与 Plonky3 的双工挑战器(duplex challenger)并行运行,每一次挤出都一致
解码器 所有低位为 11 的 2^30 个 32 位字,分别对照由 ISA 各表推导出的接受数量和一个独立的编码器;对已提交的客户程序(guest)运行 llvm-objdump
RVC 展开 LLVM 自己的编码器,作用于一个分别以压缩和非压缩方式汇编的客户程序
以数据表示的电路 checker:四条法则、查找规则和填充约定,不借用 constraints 的代码重新实现,只共享门内核
各电路族的门 行测试套件:用 Rust 自身的整数运算构造行,再经由 checker 求值;算术核心在缩小的字宽下穷举检查
内存论证与查找论证 checker 中的原生求值器,在实际执行得到的执行轨迹上运行
执行器 其执行轨迹的自检以及上述论证;没有第二个执行器
revm 客户程序 由未打补丁的上游 crate 构建的原生 revm
无状态校验器 CI 中原生运行 tests-zkevm v21.0.1 的一个已提交子集;手动原生运行整个发布版本,并通过客户程序二进制运行该子集;输入编码由 tools/stateless-ref 检查
判定器 原生检查证明,并在 revm 中执行合约

小位宽下的穷举检查#

有几个算术核心把字宽写成参数,这样就能在小到可以枚举的位宽下,对所有输入检查编码:

  • 比较 gadget 在 6 位下,遍历每一对操作数,有符号与无符号各一遍,恰好找到一个 (lt, gap),即 ISA 规定的那个;
  • MUL_DIV 的算术在 4 位下,遍历每个被除数、除数和每种除法,恰好只接受一个 (q, r),即 RV32M 规定的那个;
  • MEM_SUBWORD 的拼接在 4 位字下,对每个字、偏移和宽度,恰好只接受一个 (high, sub, low)。

电路规则,实现两遍#

凡是构建制品或加载密钥的地方,都会运行 CircuitArtifact::validate 以及内存与查找的构造规则。crates/checker 用自己的代码把同样的规则再执行一遍,从不调用 validate,并且只通过 gkr_verify::eval_gate 对门求值;这是双方都视为语义权威的唯一内核。在 validate 比较规范化展开式的地方,它的校验器改为在采样点上求值来检查各项法则;它逐行而不是借助树来重新计算每个通道的和,并指出任何表行中都不存在的元组;它从事件日志而不是分片的行来重建执行电路族的内存列。

sh
cargo run -p checker -- laws <artifact>       # Laws 1–4, then the lookup rules
cargo run -p checker -- padding <artifact>    # the padding contract
cargo run -p checker -- dump <artifact>       # the circuit, readably

篡改孪生#

篡改孪生是一份完全按照诚实证明者的方式证明出来的伪造。篡改测试套件(checker::TamperHarness)在修改了见证单元或边界标量之后重新证明一个陈述:重新统计每个通道的重数,在一次全新的全局承诺阶段中重新承诺被修改的内存列,并重新证明每一个分片。然后它验证某个分片或整个块(block),并断言拒绝所属的类别:Constraint、带通道的 Lookup,或 MemoryArgument;或者断言一个不破坏任何东西的修改能够通过验证。

篡改孪生依赖于证明者不做任何检查,而这正是设计使然:伪造的见证能得到诚实证明者所能给出的最好证明,验证者必须以预期的类别拒绝它。该套件还包含针对委托锚点的伪造;在主网 mini-block 上,它展示了证明者提示(advice)规则的另一面:一个被篡改的证明者提示单元会被内存论证拒绝,而一个被一致地篡改的单元则能通过验证,因为证明者提示不绑定任何东西。

sh
cargo test --release -p checker --test tamper -- --include-ignored --test-threads=1

可重新生成的测试数据#

kat-gen 写出每一个已提交的已知答案向量、清单、电路制品和程序身份,每一项都附带 SHA-256,读取它的测试会固定这个值。CI 会重新生成默认分组以及两个参照实现(reference oracle)的数据,只要向量目录中出现任何差异就判定失败:

sh
cargo run -p kat-gen && git diff --exit-code

客户程序 ELF 无法跨机器复现,因为 rustc 会把绝对路径嵌入 panic 位置字符串;同一台机器上的两次干净构建结果一致。因此客户程序 ELF 在一台机器上手动重新生成,CI 只重新生成由它们派生的内容。

任何检查都未覆盖的部分#

  • 没有第二个执行器。 模拟器只对照其自身帧表的一份重述以及内存与查找论证来检查,而不是对照一个独立的 RISC-V 实现;这里也没有任何执行器会走委托 shim 的软件回退路径。
  • checker 不重复执行的构造规则:内存构造规则、copower 规则,以及 validate 其余的构造规则(其中包括次数上限),都只执行一次。
  • 证明者的正确性不受检查,只通过证明真实分片的测试套件检查其完备性;这些套件在 CI 之外运行,因为每个都需要数十 GiB 内存。
  • Osaka 系列的无状态输入没有端到端参照。 该测试发布版本只填充了 Amsterdam;Electra/Fulu 布局对照 eth-act/ere-guests 检查,区块头规则则对照两个主网区块检查。

审计专区/系统

端到端的系统全貌

规范原文docs/architecture.md以 Markdown 查看

摘要

规范自身对远地虚拟机的总览。它准确陈述一份通过验证的证明确立了什么,以及验证者独立于证明者持有的三个值;跟随一个程序从 ELF 走到证明;用表格列出支撑陈述各部分的论证;列出密码学假设、可信设置假设与代码假设,包括可靠性依赖哪些 crate;给出 v1.0.0 的全部限制;并指明每一层代码所对照检查的独立参照(oracle)。

以下规范原文以英文维护;英文是本规范的标准语言。

Apogee proves executions of RV32IMAC programs. This page is the system end to end: what a proof states, how one is made and checked, what it assumes and where it stops. Each paragraph names the page that specifies its subject; glossary.md indexes the vocabulary.

1 What a proof states#

A verifier holds three things it does not take from the prover's word, and two of them from a channel the prover does not control (proof.md §1, §3):

  • the program identity, one field element: a digest of the program's instruction tables, its initial memory image, its entry pc and its VmConfig — the circuit families it uses and their heights (program.md §8);
  • the SRS digest of the ceremony, which a verifying key must carry;
  • a verifying key: that config, each family's circuit and setup commitments, the SRS's verifier points and the generic lookup table's commitments. Loading it recomputes the identity and the SRS digest from its own contents and holds its circuits to the registry's bytes (proof.md §7).

The proof's statement, PublicInputs, carries the public input bytes, the public output bytes (the journal), the exit status, and the record of the execution's shape — shard counts, memory windows, the final registers and pc, every shard's memory commitments and roots (proof.md §1).

A proof that verifies establishes that the program of that identity, started at its entry pc over its image, with the public input in its input window and some advice of the prover's choosing, executes instruction by instruction to EXIT with that status, having written that journal. Nothing is claimed of the advice, and nothing is hidden: no commitment or proof is blinded.

2 From a binary to a proof#

  1. The program. loader reads the ELF into a ProgramImage, expanding compressed instructions in place; isa decodes; program routes each instruction to one of seven instruction families, builds every family's decoded table — a row per halfword of code — and commits to them as the identity (program.md).
  2. Execution. emulator runs the guest. A cycle is one row of the family that owns its instruction, recording its memory queries: timestamped reads and writes of the pc, registers and RAM (execution-trace.md). A guest issues no system call but EXIT: its input, journal and advice are regions of memory (public-values.md, ecall-abi.md). Hashing and big-integer arithmetic are delegated: an ecall names a frame in RAM, and a row of a delegation family does the work on it (delegation.md, delegation-circuits.md).
  3. Shards. A family's rows are cut into shards of the family's height, a power of two between 2^8 and 2^22. The memory an execution touches is covered by shards of the window families, which give each word its initial and final tuple (memory.md §3). A shard is the unit of proving; a block is hundreds (circuits.md §1).
  4. A shard's proof. Its columns are committed with Mercury (mercury.md). The family's GKR circuit is run backward from its outputs to those columns, a sumcheck a layer (gkr.md), and every column is opened at the one point that pass ends on, in one batched opening (proof.md §5).
  5. The block. A BlockProof is the statement and its shard proofs. verify_block runs the global transcript once, each shard's checks, and once the memory reconciliation over every shard's roots (proof.md §6).
  6. Recursion. Verifier programs, proved by this VM in a format of its own, verify runs of shards and fold their deferred pairings; a tree of them ends in a root, a Groth16 circuit re-verifies the root, and ApogeeVerifier.sol checks that proof and the folded pairing (recursion.md).

The prover executes the guest twice: once to commit every shard's memory columns, which fixes the statement and its challenges, and once to prove each shard as it fills. Its memory is bounded by the shards in flight, not by the execution (streaming.md).

3 How soundness composes#

Each shard's GKR pass and opening tie its circuit's outputs to committed columns. On top of that, these arguments span the execution:

claim carried by
every row obeys its instruction the family circuit's enforcing gates, zero on every row the family pages, circuits.md
a row's instruction is the program's at its pc a lookup of the row's pc and fields in the family's decoded table, which the identity commits lookup.md §10
every read returns the last write one multiset over all shards: an access reads a tuple (space, address, timestamp, value) and writes one with a later timestamp; the verifier multiplies every shard's read and write roots against boundary factors for the registers and the pc memory.md
the rows are one path from the entry pc to the exit, in program order the pc is a cell of that multiset: a row reads its pc and writes the next one at least four timestamps later, so shard order, cycle uniqueness and continuity need no other argument memory.md §5, §9
a value is a byte, a word, a sign, an XOR LogUp channels over range, byte and generic tables lookup.md
the public input and the journal are the claimed bytes the two public windows' initial and final columns, held to the bytes' multilinear extensions public-values.md §5
a delegated computation is the function's invocation rows that read and write the frame through the same multiset, paired one to one with their ecall by an anchor tuple delegation.md §5

Challenges come from a Poseidon2 duplex transcript (transcript.md). The global transcript absorbs the whole statement, every shard's memory commitments included, before the memory challenges exist; each shard's transcript is seeded from its final state (proof.md §2, §4).

4 What it assumes#

  • Cryptography. Mercury's and KZG's knowledge soundness in the algebraic group model under q-DLOG (mercury.md §7); Poseidon2 as a random oracle for Fiat–Shamir; for the last step, Groth16's own assumptions. BN254 gives about 100 bits.
  • Setup. The SRS is the PSE perpetual powers of tau, sound while one contributor was honest. The code checks a file's structure and decodes every point; nothing proves it is that ceremony's, and no proving path runs Srs::validate (srs.md §3). The decider's Groth16 key comes from a second, circuit-specific ceremony (recursion.md §9).
  • What a verifier must obtain itself. The program identity and the ceremony's SRS digest. A key loads under whatever digest its own points give, so a key built over a known τ is refused only by that comparison; the verifier CLI compares identity only, and host::verify neither (proof.md §1, §3).
  • Trusted code. Soundness is the verifier's alone: constants, field, curve, transcript, poly, sumcheck, pcs-verify, pcs, gkr-verify, verifier-core, verifier, and constraints — the circuits are part of the statement, and a missing gate is a soundness bug. Computing an identity from an ELF trusts loader, isa and program. The last step adds guests/recursion, groth16, the decider's circuit and the contract. prover, emulator, trace and the proving half of host are untrusted: the prover validates nothing, and a wrong input costs an honest prover a proof that fails.
  • Nothing is constant-time (primitives.md). No proof is zero-knowledge, so proving keeps nothing secret; the one secret the code handles is a Groth16 ceremony contributor's factor, which groth16::phase2 multiplies in with the same variable-time ladder.

5 Limits#

not zero-knowledge no blinding in Mercury, GKR or the Groth16 decider
advice is unbound a guest checks it against something a proof binds (public-values.md §6)
public values at most 16,380 bytes each of input and journal (public-values.md §9)
sc.w always succeeds the one deviation from RV32IMAC's semantics; there is no reservation state (memory-ops.md §6)
traps are not provable a misaligned access, an access outside mapped memory, ebreak or a pc with no instruction ends an execution with no proof (execution-trace.md §10)
code is static the instruction stream is the image decoded at load; one undecodable word in an executable segment refuses the program (program.md)
code size .text within a decoded table's reach of its load address, 7.94 MiB at 2^22, and the image within bytecode_size_words, 4 MiB by default (program.md §5, §7)
execution length timestamps are 38 bits: 2^36 − 1 cycles (execution-trace.md §1)
delegations are a fixed set six in the base format; an EVM MULMOD with an arbitrary modulus is not one; a delegation proves one step of its function, and composing steps, validating curve points among them, is the calling code's (delegation.md §11, delegation-circuits.md)
prover memory set by the shards in flight: the measured full block peaked at 174 GiB (streaming.md §1)
block witnesses the stateless validator takes its input from an external witness producer; the built-in recorder cannot record every block (ethereum.md §4, §6)
the decider's key one per root shape, and only as trustworthy as its ceremony; the development key is forgeable (recursion.md §9)
on-chain cost about 3.6M gas for the measured block (recursion.md §10)

6 How the code is checked#

No component is checked against a second implementation of the whole system; each layer has its own independent oracle.

layer checked against
fields, curve, pairing, MSM known-answer vectors generated from arkworks, which the tests also run live
Poseidon2 and the transcript tools/transcript-ref: Plonky3 and zkhash
the decoder every 32-bit word of the instruction space against counts from the ISA; llvm-objdump over the committed guests
circuits as data checker: the circuit laws, the lookup rules and the padding contract re-implemented without constraints' code, sharing only the gate kernel (circuits.md §3)
each family's gates row suites that build rows with Rust's own integer arithmetic and evaluate them through the checker; the arithmetic cores exhaustively at reduced word widths; tamper twins, a forged witness proved as an honest prover would and refused in the expected class
the memory and lookup arguments native evaluators in checker over executed traces
the executor its own trace's self-check and the arguments above; there is no second executor, and no executor here takes a delegation shim's software fallback
the revm guest native revm, built from unpatched upstream crates
the stateless validator a committed subset of tests-zkevm v21.0.1 natively in CI; the whole release natively and the subset through the guest binary by hand; tools/stateless-ref for the input encoding
the decider the proof checked natively, and the contract executed in revm

Committed fixtures are regenerated and compared in CI (tools.md §7). The suites that prove real shards, over a toy SRS, need tens of GiB and run outside CI (README).

7 Cost#

recursion.md §10 has the end-to-end measurements for one block, from the base proof to the contract call; streaming.md §1 breaks the base proof down; and circuits.md §1 gives every circuit's width and proof size, which a shard's cost follows.

References#

  • L. Eagen, A. Gabizon. MERCURY: A multilinear polynomial commitment scheme with constant proof size and linear field work. ePrint 2025/385. publication/2025-385.pdf
  • D. Boneh, J. Drake, B. Fisch, A. Gabizon. Efficient polynomial commitment schemes for multiple points and polynomials. ePrint 2020/081. publication/2020-081.pdf
  • J.-L. Beuchat et al. High-speed software implementation of the optimal ate pairing over Barreto–Naehrig curves. ePrint 2010/354. publication/2010-354.pdf

审计专区/基础

原语:域、曲线、配对、多项式、求和校验

规范原文docs/spec/primitives.md以 Markdown 查看

摘要

其余一切所依赖的算术,全部在仓库内实现,不使用 trait、unsafe 代码或汇编。它定义了 Fr 及其三种字节形式、直到 Fq12 的 Fq 扩域塔、群 G1 与 G2 及其非压缩编码和带校验的解码器、最优 ate 配对及其精确的最终幂运算、窗口化 Pippenger MSM、多线性多项式的下标约定,以及零校验(zerocheck)求和校验,GKR 引擎复用了它的轮格式。没有任何部分是常数时间的。

以下规范原文以英文维护;英文是本规范的标准语言。

BN254's scalar field Fr, its base field Fq and the tower to Fq12, the groups G1 and G2, the optimal ate pairing, multi-scalar multiplication, multilinear polynomials and the zerocheck. The byte encodings of field elements and points (§1–§3) and the polynomial index convention (§6) are defined here.

  • All of it is this repository's code: concrete types, no field trait, no unsafe, assembly or intrinsics. field, poly and sumcheck are #![no_std] and build for the guest target; curve is std, with rayon, and no guest links it.
  • arkworks is a test oracle only, for field, curve and poly, live and through vectors tools/kat-gen generates; the tests of field and curve re-derive every arithmetic constant those crates read.
  • Nothing is constant-time: reductions, exponentiations, point additions and scalar ladders branch on their operands. No proof is zero-knowledge, so no witness is secret; the one secret this code handles, a decider ceremony contributor's factor (recursion.md §9), goes through the same variable-time ladder (§3).

1 Fr#

p = 21888242871839275222246405745257275088548364400416034343698204186575808495617

field::Fr is the integers mod p (constants::FR_MODULUS), 254 bits; p is also the order of G1 and G2, r in §3–§4. In memory an element is four little-endian 64-bit limbs of x·R mod p, R = 2^256 mod p, always reduced below p, so equal limbs are equal values. Multiplication is CIOS Montgomery over u128 intermediates. Fr::inverse is x^(p−2), None at 0; field::batch_inverse is Montgomery's trick and leaves a 0 entry 0. p − 1 = 2^28·c with c odd, and every FFT domain is a subgroup of the one constants::FR_TWO_ADIC_ROOT_OF_UNITY generates.

byte form used in
wire the value, not x·R, as 32 little-endian bytes: Fr::to_bytes. Fr::from_bytes is None for a value ≥ p and never reduces; serde goes through both every proof, key and artifact
source literal 0x and exactly 64 lowercase hex digits, big-endian: Fr::from_hex, None for any other spelling or a value ≥ p constants in crates/constants
memory the four limbs: Fr::to_memory_bytes, Fr::from_memory_bytes, None at or above p the FR_ARITH delegation's frame alone

On the guest target (cfg(target_arch = "riscv32")) addition, Montgomery multiplication and inversion call the FR_ARITH delegation through guest_sdk::recursion::fr_arith (delegation.md §10). Its circuit proves these three functions of the memory form the frame carries (delegation-circuits.md §4), so a delegated result is the software result bit for bit.

2 The Fq tower#

q    = 21888242871839275222246405745257275088696311157297823662689037894645226208583
Fq2  = Fq[u]/(u^2 + 1)
Fq6  = Fq2[v]/(v^3 − ξ)       ξ = 9 + u
Fq12 = Fq6[w]/(w^2 − v)

curve::Fq is the field of coordinates (constants::FQ_MODULUS). Its limb arithmetic is Fr's, copied literally over q's constants, and so are its wire and source-literal forms. Fq2 encodes as c0 ‖ c1; nothing above it has a byte form.

Products are schoolbook and squarings above Fq2 are products. A Frobenius map multiplies coefficients by powers of ξ tabulated in constants (FQ6_FROBENIUS_C1, FQ6_FROBENIUS_C2, FQ12_FROBENIUS_C1); Fq12::conjugate is the q^6 one. Nothing in the tower or the pairing is sparse, cyclotomic or precomputed: a pairing is only ever computed to verify something, and the code is written to be read.

3 G1, G2 and their encodings#

G1 = E(Fq)                E:   y^2 = x^3 + 3        #E  = r             generator (1, 2)
G2 ⊂ E′(Fq2), order r     E′:  y^2 = x^3 + 3/ξ      #E′ = r·(2q − r)    generator EIP-197's

G1Affine { x, y, infinity } is a point and G1Projective its Jacobian form, Z = 0 the identity, under the EFD formulas dbl-2009-l, add-2007-bl and madd-2007-bl, with the identity, P = Q and P = −Q branched on explicitly. Scalar multiplication is a fixed 4-bit window. G2Affine and G2Projective are the same code over Fq2.

A point's wire form is uncompressed affine, and there is no compressed one:

G1Affine    64 bytes    x ‖ y
G2Affine   128 bytes    x.c0 ‖ x.c1 ‖ y.c0 ‖ y.c1        x = x.c0 + x.c1·u
infinity                every byte zero

Each coordinate is an Fq in wire form; (0, 0) is on neither curve, so zero is unambiguous. G1Affine::from_bytes and G2Affine::from_bytes return None unless the bytes are all zero, or every coordinate is below q, the point satisfies its curve's equation and, in G2, whose cofactor 2q − r is not 1, [r]P is the identity — by the window ladder, with no endomorphism.

Nothing else validates: the affine structs' fields are public, and the group law, msm and the pairing compute on whatever they are given. A point's transcript form is transcript.md §4's; a .ptau file (srs.md §2) and the contract (recursion.md §9) have encodings of their own.

4 The pairing#

e(P, Q) = f_{6x+2, Q}(P)^((q^12 − 1)/r)        x = 4965661367192848881

curve::pairing::miller_loop(pairs) is Algorithm 1 of Beuchat et al. (ePrint 2010/354) with homogeneous projective line formulas: the 66-digit NAF of 6x + 2 (constants::ATE_LOOP_NAF), then the two lines adding ψ(Q) and −ψ(ψ(Q)), ψ the untwist-Frobenius-twist map. For a Q of order r no step adds a point to itself, to its negative or to the identity, so the line formulas have no exceptional case.

final_exponentiation returns exactly f^((q^12 − 1)/r): the easy part (q^6 − 1)(q^2 + 1), then the hard exponent (q^4 − q^2 + 1)/r as its base-q expansion λ0 + λ1·q + λ2·q^2 + q^3, by three exponentiations (constants::FINAL_EXP_LAMBDA_0 to FINAL_EXP_LAMBDA_2, the two negative ones conjugated) and three Frobenius maps. The Fuentes-Castañeda hard part, which arkworks uses, returns this value raised to 2x(6x^2 + 3x + 1): the two libraries agree on every pairing check and on no pairing value but 1, and the test vectors are arkworks' Miller outputs raised to the literal exponent.

pairing_check(pairs) is Π e(P_i, Q_i) = 1 by one Miller loop, whose Fq12 squarings the pairs share, and one final exponentiation: the form of every pairing equation in the system. A pair holding a point at infinity contributes 1 and is skipped; an empty product is 1.

5 MSM#

curve::msm::msm(bases, scalars) is Σ scalars_i·bases_i in G1 by windowed Pippenger, and msm_small_u32 the same sum over u32 scalars, recoded from 32 bits instead of 254. Neither looks at a scalar's size: the caller chooses, and pcs::commit chooses by a column's backing (§6). G2 has no MSM here; crates/groth16 carries its own.

  • Width. w = 3 below 32 points, otherwise ⌊0.69·⌈log2 n⌉⌋ + 2: arkworks' rule.
  • Digits. A scalar is recoded into signed digits in [−2^(w−1), 2^(w−1)], one a window, over ⌈(bits + 1)/w⌉ windows; the sign costs a negated base and halves the buckets to 2^(w−1). At 2^20 points w = 15: 17 windows for an Fr, 3 for a u32.
  • Parallelism. Each (window, chunk of the input) is a rayon task that adds bases into buckets by mixed addition and reduces them by a running sum; the tasks' sums are combined serially, w doublings a window. Group sums are exact, so the point does not depend on the thread count.

6 Multilinear polynomials#

poly::MultilinearPoly is a table of 2^n evaluations over {0,1}^n, the type of every column.

Index convention. Variable j is bit j of the index: the evaluation at y = (y_0, …, y_{n−1}) is entry Σ_j y_j·2^j. bind(r) fixes variable 0, the low bit,

f′(i) = f(2i) + r·(f(2i + 1) − f(2i))

and the old variable 1 becomes variable 0. So binding r_0, r_1, … in order leaves evaluate(&[r_0, r_1, …]), whose point[j] is variable j. Sumcheck round i binds variable i (§7), so a claim's point lists its challenges in variable order, the order evaluate and a Mercury opening (mercury.md §1) take.

Backing. PolyBacking holds the table as a bitset (U1), u8, u16, u32 or Fr. A trace column is filled and committed at its integer width (§5). get and evaluate embed an entry in Fr as they read it and leave the table alone; the first bind folds the integer table straight into an Fr table of half the length, and the backing is Fr from then on.

eq. eq_table(r) tabulates eq(r, ·) over the cube in the same index order; eq_eval(r, y) is its closed form, for any r and y.

new, get, bind, evaluate and eq_eval panic on a table, index or point of the wrong size rather than return an error.

7 The sumcheck#

crates/sumcheck proves that a gate vanishes on the cube. A sumcheck::Gate is a sum of GateTerms coef·x_a·x_b, the second factor optional, over input columns of n variables: degree at most 2 in each variable, by construction.

G(y) = 0 on all of {0,1}^n is proved as the sumcheck 0 = Σ_y eq(r, y)·G(y) at a random r. eq·G has degree at most 3 in each variable, so a round polynomial is a cubic and a round message its four coefficients [c0, c1, c2, c3], ascending — four whatever the gate, so a proof's shape depends on n and the number of inputs alone.

prove_zerocheck and verify_zerocheck run one schedule, under the tags of transcript.md §5:

1   the caller binds the columns to the transcript
2   r_0 … r_{n−1}                                      n × SUMCHECK_CHALLENGE
3   for i in 0..n:   g_i, one message of four          SUMCHECK_ROUND
                     ρ_i, binding variable i           SUMCHECK_CHALLENGE
4   final_evals: each input column at ρ, one message   SUMCHECK_FINAL_EVALS

The verifier checks the proof's shape, then g_0(0) + g_0(1) = 0 and g_i(0) + g_i(1) = g_{i−1}(ρ_{i−1}), each before absorbing g_i, and, with final_evals absorbed, g_{n−1}(ρ_{n−1}) = eq(r, ρ)·G(final_evals). It returns SumcheckClaim { point: ρ, final_evals } or a SumcheckError, and does not panic on a proof.

That last check is one equation over all the claimed evaluations and ties none of them to its column: the caller owes an opening of each at ρ, as it owes step 1. The step 1 its callers use is witness_digest, a hash and not a commitment: a sponge of its own absorbs [column count, n] and each column's cells under WITNESS_DIGEST, and its raw squeeze enters the transcript under the same tag.

What uses it. No proof in the system is this zerocheck, and no circuit is made of its Gate (gkr.md §3). The proving stack takes one type from the crate, SumcheckProof { rounds: Vec<[Fr; 4]>, final_evals }, as each layer of a gkr_verify::GkrProof. The GKR layer sumcheck (gkr.md §5) repeats step 3's rounds and checks under the same two tags, from a batched claim instead of 0 and to a final check of its own, in gkr::prove_sumcheck and gkr_verify::verify_sumcheck. prove_zerocheck and verify_zerocheck are called only by tests and tools/bench.

审计专区/基础

Transcript

规范原文docs/spec/transcript.md以 Markdown 查看

摘要

协议中的每一个挑战是如何抽取的。它规定了 Fr 上宽度为 3、轮常数固定的 Poseidon2 置换,双工海绵(速率 2,容量 1)及其吸收与挤出规则,使每条被吸收的数据流都具有单射性的带类型消息分帧,G1 点如何以四个 limb 的形式被吸收,以及全部 45 个 transcript 标签的列表,并注明每个标签的种类与位置。

以下规范原文以英文维护;英文是本规范的标准语言。

Every challenge in the protocol is drawn from a Poseidon2 duplex sponge over Fr through a typed message layer. This page specifies the permutation, the sponge, the framing, a G1 point's transcript form and every tag. Implementation: crates/transcript, #![no_std].

1 The Poseidon2 permutation#

Width 3 over Fr, S-box x^5, 4 full rounds, 56 partial rounds (S-box on lane 0 only), 4 full rounds. The round constants are RC3 of HorizenLabs/poseidon2, plain_implementations/src/poseidon2/poseidon2_instance_bn256.rs at commit 055bde3f4782731ba5f5ce5888a440a94327eaf3.

E(s) = s + (s₀+s₁+s₂)·(1,1,1)                 circ(2, 1, 1)
I(s) = s + (s₀+s₁+s₂)·(1,1,1) + (0,0,s₂)      1 + diag(1, 1, 2)

poseidon2_permute(s):
  s ← E(s)
  RC3 rows 0–3:    s_i ← (s_i + c_i)^5, every lane;   s ← E(s)
  RC3 rows 4–59:   s₀ ← (s₀ + c₀)^5;                  s ← I(s)
  RC3 rows 60–63:  s_i ← (s_i + c_i)^5, every lane;   s ← E(s)

poseidon2_permute([0, 1, 2])₀ = 0x0bb61d24daca55eebcb1929a82650f328134334da98ea4f847f760054f4a3033

constants::POSEIDON2_RC3_INITIAL, _INTERNAL and _TERMINAL hold the 80 entries read (upstream's partial rows are zero in lanes 1 and 2) as upstream's big-endian hex literals, character for character, decoded by Fr::from_hex on every call. They are pinned through the permutation, by the oracle's 128 vectors (§2), each of which reads every constant.

On riscv32, poseidon2_permute is one POSEIDON2 delegation call over the lanes' canonical bytes, falling back to these rounds when the executor answers -ENOSYS (delegation.md §10).

2 The duplex sponge#

state  [Fr; 3]   lanes 0, 1 the rate, lane 2 the capacity; zero in Transcript::new()
input  [Fr; 2]   absorbed, not yet permuted: 0 or 1 pending between operations
output [Fr; 2]   squeezed, not yet handed out: 0 to 2

observe(x):  output ← []; input.push(x); if |input| = 2: duplex()
sample():    if |input| > 0 or |output| = 0: duplex(); return output.pop()
duplex():    n ← |input|; state[0..n] ← input; input ← []
             if n > 0: state[n..2] ← 0; state[2] += n
             poseidon2_permute(state); output ← [state[0], state[1]]
  • Absorption overwrites the rate. A short absorb zero-fills the rest of it and adds its length to the capacity, so [a] and [a, 0] differ; with nothing pending, a duplex is a pure squeeze and does neither.
  • Squeezed lanes leave from the end: the first sample after an absorb is state[1], the second state[0], and a third permutes again.
  • observe drops unread output and lanes past a buffer's length stay zero, which moves no challenge and makes the state a function of the operation sequence alone.

This is Plonky3's DuplexChallenger at width 3 and rate 2. The vectors crates/transcript is tested against come from tools/transcript-ref, which shares no code with it: Plonky3's Poseidon2 keyed with zkhash's own RC3, and a transcription of this section and §3 run beside that type, agreeing with it on every squeeze. The recursion format replays the same sponge over field cells, one P2_FIELD row a duplex step (recursion.md §4).

3 Typed messages#

append_scalars(tag, xs):  observe(tag); observe(|xs|); observe(x) for x in xs
append_scalar(tag, x)  =  append_scalars(tag, [x])
append_bytes(tag, b):     observe(tag); observe(|b|); observe(c) for each 31-byte chunk c of b,
                          zero-padded to 32 bytes, read little-endian
challenge_scalar(tag):    observe(tag); return sample()

The length, the scalar count or for bytes the byte count, delimits a message: "abc" and "abc\0" are each one chunk, below 2^248 < p, and differ. A challenge absorbs its tag, so it always comes from a fresh permutation.

The framing carries no kind, so each tag names exactly one of scalars, bytes or a challenge (§5): a tag of two kinds would make append_bytes(T, b"") and append_scalars(T, []) the same T, 0. So every digest — program identity, the SRS digest, transcript::io_digest, sumcheck::witness_digest, pcs::accumulator_digest — is a fresh sponge of typed messages ended by a raw sample(), never by a challenge under one of its message tags.

snapshot() captures the state and both buffers, and Transcript::restore resumes the same challenge stream. Its postcard form is 226 bytes, state[3], input[2], input_len: u8, output[2], output_len: u8, each Fr canonical; decoding refuses input_len ≥ 2, output_len > 2 and a nonzero lane past either length. The archived path's phase files hold the global transcript, and each shard's after its GKR pass, in this form (streaming.md §6). A shard transcript is no restored global sponge but a fresh one whose first message carries the global state digest (proof.md §4).

Each typed operation appends Absorb { tag, n_scalars } (payload elements: scalars, or chunks) or Challenge { tag } to event_log(). Raw observe and sample are not logged, the log never feeds the sponge and a snapshot omits it; checker::tape holds the global transcript's log to the order G1–G11 (tools.md §4).

4 G1 points#

A point is absorbed as four Fr limbs of its 64-byte encoding x ‖ y (primitives.md §3), with no curve arithmetic (transcript::g1_limbs):

[ x[0..16], x[16..32], y[0..16], y[16..32] ]   each half read little-endian, below 2^128 < p
[ S, S, S, S ]                                 the 64 zero bytes of infinity; S = 2^128

A coordinate is an Fq element and q > p, hence the halves. S is constants::G1_INFINITY_SENTINEL: no 16-byte half reaches 2^128, so the limbs determine the 64 bytes whether or not they encode a point on the curve. The absorber never refuses; a point is validated where it is decoded, before a pairing reads it.

transcript::append_g1_points(tr, tag, points) absorbs k points as one message of 4k limbs, never k messages, so the framed length binds k. pcs::append_g1_list is it over G1Affine::to_bytes, and pcs::append_g1 a list of one.

5 Tags#

Tag = u64: constants::transcript_tags, 45 tags numbered from 1 and named by transcript_tags::NAMES[tag − 1]; 0 is not a tag. Kinds: S scalars, B bytes, C challenge. Where: G1–G11 and the shard transcript are proof.md §2, §4, the SRS digest §3 there; identity program.md §8; Mercury mercury.md; GKR gkr.md §5; io_digest public-values.md §5; stacks and nodes recursion.md §1.3, §8.3. † marks a tag on no proof path.

tag where
1 PROTOCOL_SUITE S G1: [PROTOCOL_VERSION]
2 PUBLIC_INPUTS B G7: io_digest's 32 canonical bytes
3 COMMITMENT S a commitment list: identity, G8, shard witness, Mercury
4 SUMCHECK_ROUND S a sumcheck round's coefficients
5 SUMCHECK_CHALLENGE C a round's challenge; first, a zerocheck's eq-randomizers
6 EVALUATION_CLAIM S Mercury: the point, then the claimed values
7 PCS_OPENING S Mercury: proof points and evaluations
8 WITNESS_DIGEST S sumcheck::witness_digest's sponge, and its result †
9 SUMCHECK_FINAL_EVALS S the zerocheck's final evaluations †
10 MERCURY_INSTANCE S Mercury: [n]
11 MERCURY_ALPHA C Mercury: α
12 MERCURY_GAMMA C Mercury: γ
13 MERCURY_Z C Mercury: z, redrawn while 0
14 BDFG_BATCH C Mercury: δ
15 BDFG_POINT C Mercury: z′
16 PAIRING_MERGE C Mercury: the pairing merge ρ
17 MERCURY_BATCH C Mercury: the column batch ρ
18 ACCUMULATOR_DIGEST S pcs::discharge: the entry words' sponge, and its result †
19 ACCUMULATOR_MERGE C pcs::discharge: the per-check weight †
20 PUBLIC_INPUT_STREAM B io_digest: the input
21 PUBLIC_OUTPUT_STREAM B io_digest: the output
22 PROGRAM_IDENTITY S identity: [code_version]; G6: [identity]
23 VM_CONFIG S identity; G3
24 SHARD_COUNTS S G4
25 GKR_OUTPUTS S GKR: the output tables
26 GKR_OUTPUT_POINT C GKR: the top point
27 GKR_BATCH C GKR: a transition's claim batch
28 GKR_LAYER_CLAIMS S GKR: a transition's claimed values
29 GKR_CHILD C GKR: a halving transition's line point
30 MEMORY_WINDOWS S G5
31 MEMORY_BOUNDARY S G9
32 PROGRAM_ENTRY S identity: [entry_pc]
33 LOOKUP_CHALLENGE C shard: g, then β (lookup.md §2)
34 SRS_DIGEST S G2
35 SRS_VERIFIER B the SRS digest: the 320-byte SrsVerifier
36 MEMORY_GROUP S G8: [family, shard count]
37 MEMORY_CHALLENGE C G10, four times
38 GLOBAL_STATE_DIGEST C G11
39 SHARD_SEED S shard: [digest, family, index]
40 SHARD_TS_WINDOW S shard: [start, end]
41 GENERIC_TABLE S the SRS digest: the generic table's 3 points, 12 limbs
42 STACK_CHALLENGE C a recursion-format shard: its σ stack challenges
43 FOLD_STATE S a node: a verified shard's final transcript state
44 FOLD_WEIGHT C a node: a shard's w, w′, or a child's weight
45 FOLD_CHILD S a node: a child's journal

审计专区/基础

结构化参考串

规范原文docs/spec/srs.md以 Markdown 查看

摘要

整个系统的可信设置。它指明所用的仪式,即 PSE 的 perpetual powers of tau 的第 80 次贡献,并给出其 [τ]_1 以便识别;规定 .ptau 文件如何导入,以及解码能证明什么;说明哪些内容是假定的而非检查的;并定义 SRS 归档、验证者持有的三个点、基于这些幂次的 KZG,以及构造判定器 Groth16 密钥所用的 Lagrange 基。

以下规范原文以英文维护;英文是本规范的标准语言。

The powers of τ every commitment is made under: the ceremony they come from, how its file is read and what is checked, the archive an SRS is cached in, the three points a verifier holds, KZG over them, and the Groth16 first phase read from the same file. Implementation: crates/srs.

1 The ceremony#

The SRS is PSE's perpetual powers of tau, contribution 80: files ppot_0080_<p>.ptau, kept in assets/ptau/, which is gitignored. Hermez's powersOfTau28_hez_final_*.ptau is another ceremony with another τ; the reader ingests it as readily, and every commitment, key and identity over it differs. This ceremony's [τ]_1, as the hex of its canonical encoding x ‖ y:

9bbb31bedc304e081e2aada4b56c2217e0e94ee16874e3517d14bef5dcec3a16
317ff1589e53513fa333591b318e8f1e55ef7c37d92beb6d9a61d770a2f39506

PSE's files are cut from one ceremony: the first 2^k powers, and the Lagrange bases of domains up to 2^k, agree in every file of power k or more. One file, ppot_0080_24.ptau (19.3 GB), serves every use. A base key needs as many powers as its tallest family has rows, at most 2^22, the menu's top, and at least the generic table's 2^18 (lookup.md §9); bench prove reads 2^22. The recursion format reads 2^24, its largest stack (verifier_core::STACK_LOG, recursion.md §1.3), and the decider its domain's Lagrange bases (§7).

2 Ingesting a .ptau file#

Srs::from_ptau(path, k) reads snarkjs's .ptau container, the one ingestion format. Integers are little-endian.

0    4    "ptau"
4    4    version: 1
8    4    section count: at most 64
12   ..   sections: id u32 | size u64 | payload

id 1   header, 44 bytes: n8 = 32 | q (n8 bytes) = BN254's Fq modulus | power p | ceremonyPower
id 2   tauG1: 2^(p+1) − 1 G1 points, [τ^0]_1 first
id 3   tauG2: 2^p G2 points, [1]_2 then [τ]_2

Sections 1–3 must each occur once, at the sizes p implies; the others (alpha, beta, the contribution record, the Lagrange bases of §7) are not read here. from_ptau takes the first 2^k points of section 2 (k ≤ p) and the first two of section 3.

A point is uncompressed affine in little-endian Montgomery form: each 32-byte coordinate holds coord·R mod q, R = 2^256; G1 is x ‖ y, G2 x.c0 ‖ x.c1 ‖ y.c0 ‖ y.c1. It is the only non-canonical point encoding the code reads. A coordinate is read as a canonical Fq (refused at or above q), multiplied by R^−1 and re-encoded, and the canonical bytes go through G1Affine::from_bytes or G2Affine::from_bytes (primitives.md §3), the one validating decoder. All-zero bytes are infinity in both forms.

from_ptau never panics: every refusal is an SrsError — Io, Truncated, BadMagic, BadVersion, BadSection (over 64 sections, sections 1–3 not each present once, a header not 44 bytes, a power outside 1..=30, a section size p does not imply), WrongCurve, PowerTooLarge (k > p), and InvalidPoint { index }, a failing point but not necessarily the first.

3 What is validated, and what is presumed#

Decoding proves every point canonical, on its curve and in the order-r subgroup. Srs::validate adds that they are powers of one τ:

g1[0] = G1 generator      g2_gen = G2 generator      no point is infinity
e(Σ_i c_i·g1[i], g2_tau) = e(Σ_i c_i·g1[i+1], g2_gen)       i < n − 1

with each c_i 31 bytes from /dev/urandom, so that no file can be built to pass: a power that is not τ times the one before survives with probability at most 2^−248. The infinity check excludes τ = 0: pairing_check skips a pair at infinity (primitives.md §4), so such an SRS would pass vacuously and kzg_verify over it accept any opening. validate identifies nothing, and no proving path runs it.

Soundness needs nobody to know τ, which this code presumes of the ceremony. A statement binds the SRS only through the SRS digest (proof.md §3), which covers the SrsVerifier and the generic table's three commitments, not the powers, which only a prover reads. A key's loader recomputes the digest from the key's own points, so a key whose SrsVerifier has a known τ loads under its own digest: a verifier takes the ceremony's digest from a channel the prover does not control, or recomputes it from the ceremony. Program identity covers neither the SrsVerifier nor the table; in the recursion tree the digest is a constant of both programs' images, which their identities bind (recursion.md §8.1).

4 The SRS archive#

Srs::save and Srs::load keep an ingested SRS in a file of their own, integers little-endian and points canonical (primitives.md §3); bench recurse caches its 2^24 powers in one.

0     8          "APOGESRS"
8     4          version: 1
12    4          power k, at most 30
16    8          G1 count: 2^k
24    128        g2_gen
152   128        g2_tau
280   64·2^k     g1, [τ^0]_1 first

load requires exactly 280 + 64·2^k bytes before reading a point (the cap on k keeps the product from wrapping) and decodes every point through from_bytes, which catches a corrupted coordinate, not a substituted archive. It refuses with Truncated, BadMagic, BadVersion, BadSection and InvalidPoint.

5 SrsVerifier#

The only SRS material a verifier takes: g1_gen = [1]_1, g2_gen = [1]_2 and g2_tau = [τ]_2, what Mercury's pairings read. A verifier never commits; the generic table's commitments reach it as given points. The wire form is 320 bytes, g1_gen ‖ g2_gen ‖ g2_tau, canonical, unframed — its postcard form, VerifyingKey's srs_verifier and verifier::encode_srs_verifier alike — and every reader decodes it through the validating from_bytes.

6 KZG#

srs::kzg, over coefficients little-endian in the degree (coeffs[i] multiplies X^i, as g1[i] is [τ^i]_1):

kzg_commit(f)   = Σ_i f_i·[τ^i]_1                               one MSM
kzg_open(f, z)  = (f(z), [q(τ)]_1), q = (f − f(z))/(X − z)      one Horner pass gives both
kzg_verify(cm, z, v, w):  e(cm − v·[1]_1 + z·w, [1]_2) · e(−w, [τ]_2) = 1

More coefficients than powers is an error, never a truncation. The zero polynomial commits to infinity and opens to (0, infinity), which verifies. A Mercury commitment is exactly kzg_commit of the evaluation table read as coefficients (mercury.md §2). Mercury calls neither kzg_open nor kzg_verify, but its pairing relations take their shape, e(A, [1]_2) = e(B, [τ]_2) with both G2 arguments SRS constants, which is what lets recursion fold them instead of pairing (recursion.md §8.3).

7 Phase 1#

srs::Phase1::from_ptau(path, m), for m ≤ p and m ≤ 28, reads what a Groth16 key takes from the ceremony at a domain of n = 2^m: tau_g1, [τ^i]_1 for i < 2n − 1, from section 2; and lagrange_g1 and lagrange_g2, [L_j(τ)] in each group, L_j the Lagrange polynomial at ω^j and ω of order n squared down from constants::FR_TWO_ADIC_ROOT_OF_UNITY, from sections 12 and 13, which hold the bases of domains 1, 2, 4, … in turn, domain n from point n − 1. It refuses a basis that is not this domain's: tau_g1[0] and each basis's sum must be the generator, and Σ_j ω^j·[L_j(τ)]_1 = [τ]_1. The G2 basis is held to the curve, not the subgroup. The decider's key is made over it (recursion.md §9).

审计专区/基础

Mercury

规范原文docs/spec/mercury.md以 Markdown 查看

摘要

多项式承诺方案。它确定了 Mercury 与 BDFG20 两篇论文没有确定的部分:变量划分;承诺即求值表的 KZG 承诺;完整的打开协议及其十六步 transcript 时序;704 字节的证明,以及验证者合并后的双配对检查。它还增加了在同一点上批量打开多列的形式(每个分片都会用到),以及由十二个累加器项构成的延迟形式,递归树对其做折叠,而不是直接做配对。页面最后讨论成本与安全性。

以下规范原文以英文维护;英文是本规范的标准语言。

Every committed column is opened with Mercury (Eagen and Gabizon, ePrint 2025/385), finished by the batched KZG opening of BDFG20 (Boneh, Drake, Fisch and Gabizon, ePrint 2020/081). This page pins what the papers leave open, and adds a batch of k columns at one point and the deferred form the recursion tree folds. crates/pcs is the prover, the curve side and the pairings; crates/pcs-verify, no_std, is the verifier's field side, which the recursion guest links.

1 Parameters and the variable split#

n = 2^{2t} evaluations with 1 ≤ t ≤ 27, b = 2^t = √n, s = 2t variables; u ∈ Fr^s is the opening point and v the claimed value. pcs_verify::check_num_vars refuses every other variable count (PcsError::UnsupportedNumVars) and never pads, which is why every trace height is an even power of two (program.md §7). The ceiling, pcs_verify::MAX_NUM_VARS = 54, is where Fr's 2-adicity of 28 runs out of the 2b-th roots of unity §3.1 needs, and it keeps 2^{|u|} in range for a u the verifier is handed.

The evaluation table is read as coefficients, and variable m is bit m of an index (primitives.md §6). Write an index i + j·b with i the low t bits, as Mercury §3.1 does; its evaluation is the coefficient of X^{i+j·b}. The point splits the same way: u1 is its first half, u_0..u_{t−1}, and pairs with i; u2 is u_t..u_{2t−1} and pairs with j.

f(X) = Σ_{i<b} X^i·f_i(X^b),    f_i(X) = Σ_{j<b} f_{i+j·b}·X^j
f̂(u) = Σ_{i,j<b} eq(i, u1)·eq(j, u2)·f_{i+j·b}

pcs::open returns what poly::MultilinearPoly::evaluate gives at u, and a verifier handed the two halves swapped rejects.

2 Commitment#

pcs::commit returns [f(x)]_1 for §1's f(X), an MSM over the first n SRS powers: exactly the KZG commitment of the evaluation table read as coefficients (srs::kzg::kzg_commit), with no second scheme behind it. It refuses an SRS of fewer than n powers (SrsTooSmall). A column backed by U1, U8, U16 or U32 (poly::PolyBacking) is widened to u32 and committed through curve::msm::msm_small_u32, never lifted to Fr; an Fr backing goes through curve::msm::msm.

The map from a table to its commitment is Fr-linear, which §5 uses, and a zero coefficient adds nothing: a column extended by zero rows keeps its commitment. So the generic table's commitments serve every height that holds the table (lookup.md §9), and pcs::commit_stack commits a recursion stack without building it.

3 The opening protocol#

3.1 The polynomials#

definition coefficients sent as
h Σ_i eq(i, u1)·f_i(X); its X^j coefficient is f̂(u1, j) b h
q, g f = (X^b − α)·q + g, so g = Σ_i f_i(α)·X^i n − b, b q, g
S the symmetrized witness below b − 1 s
D X^{b−1}·g(1/X): g reversed b d
H (f − (z^b − α)·q − g_z)/(X − z) n − 1 pi_z
W, W′ §3.3 b − 1 each w, w_prime

P_u(X) = Σ_{i<b} eq(i, u)·X^i = Π_{m<t}(u_m·X^{2^m} + 1 − u_m), so ⟨P_u, g⟩ = ĝ(u) for g of fewer than b coefficients (Mercury §4.2). The prover uses its coefficients, poly::eq_table(u); the verifier evaluates the product in O(t).

The fold (Mercury §5) divides every f_i by X − α, b Horner divisions advanced together in one pass over the rows, with no transform. Then ĝ(u1) = h(α) and ĥ(u2) = f̂(u) = v, and one S proves both inner products (Mercury §4.1), the left side's constant coefficient being 2·(⟨g, P_u1⟩ + γ·⟨h, P_u2⟩):

g(X)·P_u1(1/X) + g(1/X)·P_u1(X) + γ·(h(X)·P_u2(1/X) + h(1/X)·P_u2(X))
    = 2·(h(α) + γ·v) + X·S(X) + S(1/X)/X

S is coefficients b..2b−2 of X^{b−1} times the left side, computed with four forward transforms of size 2b and one inverse; no transform in an opening is larger (crates/pcs/src/fft.rs, over constants::FR_TWO_ADIC_ROOT_OF_UNITY).

3.2 The transcript schedule#

pcs::open and pcs_verify::scalars run this Fiat–Shamir schedule step for step. A point or a list of points is one message (transcript.md §4).

# tag message
1 absorb MERCURY_INSTANCE n
2 absorb COMMITMENT cm, as passed: open never recommits it
3 absorb EVALUATION_CLAIM u_0..u_{s−1}, then v
4 absorb PCS_OPENING h
5 squeeze MERCURY_ALPHA α
6 absorb PCS_OPENING [q, g]
7 squeeze MERCURY_GAMMA γ
8 absorb PCS_OPENING [s, d]
9 squeeze MERCURY_Z z, by §3.4's rule
10 absorb PCS_OPENING g_z, g_{1/z}, h_z, h_{1/z}, s_z, s_{1/z}, one message
11 absorb PCS_OPENING pi_z, before δ although the batch does not read it
12 squeeze BDFG_BATCH δ
13 absorb PCS_OPENING w
14 squeeze BDFG_POINT z′
15 absorb PCS_OPENING w_prime
16 squeeze PAIRING_MERGE ρ, after all eight points and six values

The prover draws ρ too and discards it, so both sides leave the transcript in one state and an opening composes inside a larger transcript, the shard transcript (proof.md §4).

3.3 The BDFG20 batch#

Mercury §6 step 4(e) leaves the batched KZG opening to BDFG20 §4. The point set is T = {z, 1/z, α}, and the four polynomials are batched in this order, which fixes the power of δ each carries (pcs_verify::bdfg::items, which both sides read):

i f_i S_i Z_{T∖S_i} r_i interpolates
0 g {z, 1/z} X − α g_z, g_{1/z}
1 h {z, 1/z, α} 1 h_z, h_{1/z}, h_α
2 S {z, 1/z} X − α s_z, s_{1/z}
3 D {z} (X − 1/z)(X − α) D_z
F(X) = Σ_i δ^i·Z_{T∖S_i}(X)·(f_i(X) − r_i(X))                          W  = [(F/Z_T)(x)]_1
L(X) = Σ_i δ^i·Z_{T∖S_i}(z′)·(f_i(X) − r_i(z′)) − Z_T(z′)·(F/Z_T)(X)    W′ = [(L/(X − z′))(x)]_1

Both divisions are exact for an honest prover, and open asserts it (pcs_verify::bdfg::{quotient, linearization}).

3.4 Challenges and derived values#

Mercury draws z ∈ F*; here z is drawn again under MERCURY_Z while it is zero (pcs_verify::challenge_z). T needs three distinct points, so both sides refuse with PcsError::DegenerateChallenge when z² = 1, z = α or z·α = 1 (pcs_verify::degenerate): probability about 2^−252, and a loss of completeness only. The recursion tape draws z once and asserts all four conditions (verifier_core::tape::mercury_scalars).

The verifier is not sent h(α) or D(z): it derives them, as Mercury §6 step 4(c) does (pcs_verify::derive_h_alpha), and the prover builds the batch around the same derived values.

D_z = z^{b−1}·g_{1/z}
h_α = (g_z·P_u1(1/z) + g_{1/z}·P_u1(z) + γ·(h_z·P_u2(1/z) + h_{1/z}·P_u2(z) − 2v)
       − z·s_z − s_{1/z}/z) / 2

Opening D at z to D_z is the degree check on g (Mercury §4.3); opening h at α to h_α is §3.1's identity at z.

4 The proof and the verifier's checks#

pcs::MercuryProof is eight points and six values. Its field order is its byte order and its transcript order, and to_bytes writes pcs::PROOF_BYTES = 704 bytes for every n and k:

h  q  g  s  d  pi_z  w  w_prime                 8 × 64 bytes, G1 uncompressed (primitives.md §3)
g_z  g_inv_z  h_z  h_inv_z  s_z  s_inv_z        6 × 32 bytes, canonical Fr (primitives.md §1)

from_bytes returns None unless every point decodes through curve::G1Affine::from_bytes (canonical and on the curve; G1's cofactor is 1) and every value through field::Fr::from_bytes.

Two relations are checked, each written e(A, [1]_2) = e(B, [x]_2) so that both G2 arguments are SRS constants: the fold identity at z (Mercury §6 step 4(f), its z term moved into G1) and the BDFG20 batch (BDFG20 §4.1). They merge under ρ into one curve::pairing::pairing_check of two pairs:

A1 = cm − (z^b − α)·q − g_z·[1]_1 + z·pi_z                  B1 = pi_z
A2 = Σ_i c_i·cm_i − K·[1]_1 − Z_T(z′)·w + z′·w_prime        B2 = w_prime
     cm_i = g, h, s, d    c_i = δ^i·Z_{T∖S_i}(z′)    K = Σ_i c_i·r_i(z′)
     Z_T(z′) = (z′ − z)(z′ − 1/z)(z′ − α)
accept iff  e(A1 + ρ·A2, [1]_2)·e(−(B1 + ρ·B2), [x]_2) = 1

If either relation is false the merged one holds for at most one ρ, and ρ follows every proof element. The verifier reads three SRS points, srs::SrsVerifier's [1]_1, [1]_2 and [x]_2, and does no G2 arithmetic.

pcs::verify refuses, in order: a u whose length is not an instance's (§1), before anything is absorbed (UnsupportedNumVars); a proof point or cm off the curve (InvalidPoint), checked again because a proof built in memory has met no decoder; a degenerate T (DegenerateChallenge); a failed pairing check (VerificationFailed), which does not say which relation failed.

5 Batching k columns at one point#

Not in the papers. k commitments to columns of one size, opened at one point u, are one Mercury instance with one proof (pcs::batch_open, pcs::batch_verify); a shard proof's opening is one such batch (proof.md §5). Three steps precede §3.2's sixteen (pcs_verify::batch_preamble):

# tag message
B1 absorb COMMITMENT cm_0..cm_{k−1}, as passed, one message of 4k limbs
B2 absorb EVALUATION_CLAIM u_0..u_{s−1}, then v_0..v_{k−1}
B3 squeeze MERCURY_BATCH ρ

The opening then runs on (cm*, u, v*), with cm* = Σ_i ρ^i·cm_i and v* = Σ_i ρ^i·v_i.

  • ρ follows every commitment and every claimed value. Column i carries ρ^i, column 0 carrying 1, so a reordered or shortened list is a different statement.
  • The list is one message, so its length 4k fixes k, and then s from B2's s + k scalars: the absorbed stream is injective.
  • ρ = 0 is not redrawn: it checks column 0 alone, and is one of the roots the bound below counts.
  • A batch of one is a different transcript from a bare opening; their proofs do not interchange.

The batch is sound: by §2's linearity cm* commits to f* = Σ_i ρ^i·f_i, and evaluation at u is linear, so v* − f̂*(u) = Σ_i (v_i − f̂_i(u))·ρ^i, a polynomial in ρ of degree at most k − 1 fixed before ρ is drawn. A false claim survives with probability at most (k − 1)/|Fr|.

The prover builds f* as one Fr column and opens it once; mixed sizes are refused (MixedColumnSizes). The verifier refuses an empty list (EmptyBatch) or a value count that differs (BatchLengthMismatch), checks every cm_i on the curve before summing, derives cm* by a k-point MSM and runs §4 on it. pcs::batch_open_stacked opens recursion stacks at u ‖ r (recursion.md §1.3); batch_open is it at r = [], one column a stack.

6 Deferred verification and the accumulator#

6.1 The twelve entries#

Deferring a verification runs every check of §4 but the pairing and keeps the relation's terms: twelve pcs::AccumulatorEntry { side, scalar, point }, side a pcs::PairingSide, G2One for [1]_2 or G2X for [x]_2. The points are [cm, h, q, g, s, d, pi_z, w, w_prime, [1]_1], as pcs_verify::ENTRY_POINTS indexes them, and the scalars are pcs_verify::scalars's, in §4's notation:

# side point scalar # side point scalar
0 G2One cm 1 6 G2One pi_z z
1 G2One h ρ·c_1 7 G2One w −ρ·Z_T(z′)
2 G2One q −(z^b − α) 8 G2One w_prime ρ·z′
3 G2One g ρ·c_0 9 G2One [1]_1 −(g_z + ρ·K)
4 G2One s ρ·c_2 10 G2X pi_z 1
5 G2One d ρ·c_3 11 G2X w_prime ρ

The G2One terms sum to A1 + ρ·A2 and the G2X terms to B1 + ρ·B2; entry 9 carries both relations' [1]_1, and entry 2 is zero exactly when z^b = α, which is legal. A batch derives cm* first, so entry 0 is cm* and a check is ENTRIES_PER_CHECK = 12 entries whatever k. pcs::verify and pcs::batch_verify spend the entries at once; pcs::verify_deferred and pcs::batch_verify_deferred return them.

6.2 What uses it#

  • Base verification pairs: crates/verifier runs pcs::batch_verify for each shard, and no ShardProof or BlockProof carries an entry.
  • The recursion tree folds. A shard's tape computes the twelve scalars over field cells (verifier_core::tape::mercury_scalars), cm* being a hint; the node folds them with the batch check cm* = Σ_i ρ^i·cm_i (recursion.md §8.3), and one pairing check at the top discharges every shard's (recursion.md §9). Natively, host::recursion runs pcs::batch_verify_deferred on each shard for its cm*.
  • Nothing else: pcs::verify_deferred, §6.3's word form, pcs::accumulator_digest and pcs::discharge are called only by crates/pcs's tests and tools/kat-gen.

6.3 The word form and discharge#

A list is grouped into deferred checks, checks[j] being group j's entry count, and written as canonical Fr words (pcs::accumulator_words, inverse pcs::accumulator_from_words):

group:  count  entry_0 .. entry_{count−1}
entry:  side  scalar  x_lo  x_hi  y_lo  y_hi       side 0 = G2One, 1 = G2X; ENTRY_WORDS = 6

A word is 32 bytes, so an entry is 192, and the limbs are the point's transcript form (transcript.md §4). There is no header, so two lists concatenate into a list whose checks keep their groups. Decoding refuses a count of 2^64 or more or one that overruns, a side other than 0 or 1, a limb of 2^128 or more other than the sentinel, a partial sentinel, the all-zero quadruple (infinity has one spelling), and a point that is not canonical or not on the curve. The digest is the words as one ACCUMULATOR_DIGEST message in a fresh sponge, then a raw sample; covering the count words, it binds the grouping.

discharge(vsrs, entries, checks):
  every entry's point on the curve, before anything else
  ν = fresh sponge: absorb ACCUMULATOR_DIGEST [digest], challenge ACCUMULATOR_MERGE
  A = Σ_j ν^j·(group j's G2One terms)      B = Σ_j ν^j·(group j's G2X terms)
  accept iff e(A, [1]_2)·e(−B, [x]_2) = 1

An entry's point is a claim: absorption binds only its limbs, and an entry built in memory has met no decoder. The weight keeps the checks apart: at weight 1, two checks with equal and opposite errors pass together, and weighted, a false group passes only where ν is a root of a nonzero polynomial of degree below the group count. ν is a function of the words because discharge takes no transcript. An empty list discharges.

7 Cost and security#

prover, field O(n): a pass for h, the fold, H's division; S in O(b log b)
prover, MSMs 2n + 5b − 4 scalar multiplications: q n − b, pi_z n − 1, h, g, d b each, s, w, w_prime b − 1 each. A commitment is one more MSM of n
batch of k k multiply-adds a coefficient for f* and a k-point MSM for cm*, then one opening
verifier O(t) field operations, MSMs of ten points and of two (and of k), one two-pair pairing check
measured n = 2^22: commit 1.30 s, open 2.89 s. 16 columns of 2^20: a batch opens in 1.01 s and verifies in 4.8 ms, 16 single openings take 9.79 s and 62 ms. 18-core Apple M5 Pro; bench mercury, bench mercury-batch

Knowledge soundness holds in the algebraic group model under q-DLOG (Mercury §6, BDFG20 §4), with Fiat–Shamir over the Poseidon2 transcript in the random-oracle model and an SRS whose x nobody knows (srs.md §3). The statistical terms are Schwartz–Zippel over α, z and z′, of order a committed polynomial's degree over |Fr|, a few 1/|Fr| for γ, δ and the merge ρ, (k − 1)/|Fr| for a batch and the group count over |Fr| for ν: each is below 2^−220 for every instance in use, and the level is BN254's (architecture.md §4). Nothing is hiding and nothing is blinded.

Mercury's SRS has exactly n powers; here one SRS serves every size, so a prover can commit to a polynomial of degree n or more, and no degree bound is checked. None is needed: Mercury §6's argument goes through with its Schwartz–Zippel terms over that degree, and the opening at u is the multilinear extension of the polynomial's first n coefficients. A commitment binds that truncation, which is linear, so §5's argument holds for it too.

审计专区/程序与执行

程序:从 ELF 到程序身份

规范原文docs/spec/program.md以 Markdown 查看

摘要

客户程序(guest)二进制如何变成验证者所知的静态描述。它规定了 ELF 加载;在不移动任何地址的前提下展开压缩指令的半字扫描;ProgramImage 的序列化格式;解码器,以及它如何把 RV32IMA 的 59 条指令路由到七个电路族;每个半字一行的解码表;one-hot 指令掩码;VmConfig 及其高度;以及程序身份:它究竟绑定什么、不绑定什么。

以下规范原文以英文维护;英文是本规范的标准语言。

How a guest binary becomes the static, verifier-known description of a program. crates/loader reads an ELF into a ProgramImage, crates/isa decodes its instructions, and crates/program routes them into per-family decoded tables, derives the VmConfig and commits to all of it as the program identity. Every step is a pure function of its input.

1 Loading#

loader::load_elf accepts a static executable — ELFCLASS32, little-endian, ET_EXEC, EM_RISCV — whose PT_LOAD segments lie inside guest RAM (constants::guest_memory) at even addresses, pairwise disjoint, with p_filesz ≤ p_memsz, at least one of them executable. Anything else is a named LoaderError: DynamicElf for ET_DYN, PT_DYNAMIC or PT_INTERP, EntryNotAnInstruction for an e_entry that is not the first halfword of an instruction, and the sweep's refusals (§2).

Of a program header it reads p_type, p_offset, p_vaddr, p_filesz, p_memsz and the PF_X bit, and nothing else: the VM has no pages, and all of RAM is addressable whatever the segments declare. The address map, and the segment layout a guest ELF keeps for host loaders, are ecall-abi.md §6.

2 RVC expansion and slots#

slots holds one Slot per halfword from slot_base, the lowest loaded address, to the end of the highest executable segment: the slot of pc is slots[(pc − slot_base)/2]. load_elf sweeps each executable segment's file bytes from its start, by the halfword at pc:

low bits 11   pc += 4   Instruction { word: the four bytes, compressed: false }, MidInstruction
0x0000        pc += 2   NonInstruction
otherwise     pc += 2   Instruction { word: rvc::expand(halfword), compressed: true }

Every other halfword is NonInstruction. An encoding longer than 32 bits (InstructionTooLong), one cut off by the end of the file bytes (TextTruncated) or a halfword rvc::expand refuses (RvcIllegal) refuses the image; 32-bit words are decoded in §5.

  • Addresses are never compacted. A c.addi at 0x1002 stays there and occupies two bytes, so linker-resolved addresses hold; compressed, the instruction's length, is the only record of whether the next pc is pc + 2 or pc + 4.
  • rvc::expand takes the base C extension in its RV32 form and refuses the floating-point forms, the RV64-only forms (c.addw, c.subw, a shift with shamt[5]), the reserved code points and the Zc* encodings. A HINT such as c.addi x0, 5 is expanded; its 32-bit form writes x0.
  • 0x0000, RVC's defined-illegal encoding, is not refused: LLVM pads unreachable blocks with it. Reaching it is fatal at run time.

A desynchronised sweep cannot make a wrong instruction provable. A slot is a function of the bytes at its own pc, so every Instruction slot is what a hart fetching there would decode; data that shifts the sweep off the true boundaries can only lose true instruction starts, whose pcs then have no table row (§5), or meet an unclaimed encoding and refuse the image. crates/loader/tests/differential.rs holds committed guests' slots to llvm-objdump's listing, and the expansion to LLVM's own encoder over guests/rvc-dense, one sequence assembled compressed and not: a wrong expansion would be a valid proof of another program.

3 ProgramImage and its wire form#

The wire form is postcard over ProgramImage's four fields in order, with no header; every integer but kind is a LEB128 varint:

ProgramImage = entry ‖ n ‖ n × Segment ‖ slot_base ‖ m ‖ m × Slot
Segment      = vaddr ‖ mem_len ‖ len ‖ bytes    mem_len is p_memsz; bytes, the p_filesz file bytes
Slot         = kind: u8 ‖ word                  kind 0 a four-byte instruction, 1 a two-byte one,
                                                2 MidInstruction, 3 NonInstruction; word 0 for 2, 3

The reader re-checks what load_elf establishes — segments at even addresses, sorted, disjoint and inside RAM; slot_base the lowest segment's address; each four-byte Instruction followed by its MidInstruction; entry an Instruction slot — but not slots against the bytes. artifact-dump writes this form (tools.md §5); no prover or verifier reads it, host::setup starting from the ELF.

ProgramImage::initial_word(addr) is the little-endian word at addr before the first cycle: file bytes where a segment has them, zero elsewhere. The image column (§8) and the trace's initial RAM values are read from it.

4 The instruction set and family routing#

isa::decode takes 32-bit words only and accepts exactly RV32IMA's 59 instructions — 40 of RV32I, 8 of M, 11 of A — with any value in an operand field, x0 destinations included, and the one legal value in every fixed field: funct7, jalr's funct3, all of ecall and ebreak, an atomic's .w width, lr.w's rs2 = 0. Everything else is a DecodeError: RV64 encodings, F, D, Zicsr, fence.i, privileged instructions. crates/isa/tests/sweep.rs holds it, over all 2^30 words with low bits 11, to accepted counts derived from the ISA's tables and to an independent encoder.

  • fence is every MISC-MEM word with funct3 = 000, 2^22 of them, whatever its rd, rs1, fm, pred and succ: the ISA has a base implementation treat a reserved setting as a normal fence (llvm-objdump prints those <unknown>), and on one hart a fence does nothing.
  • An immediate is the value the instruction uses: sign-extended for I, S, B and J, the shifted word for U, the amount for a shift immediate. B and J displacements are even by encoding; nothing asks for 4-byte alignment.

program::row_kind, a total function, routes an instruction to one family and one bit of that family's mask (§6):

id family mnemonics, from mask bit 0 up
0 ADD_SUB_LUI_AUIPC system (ecall ebreak fence), addi auipc add sub lui
1 JUMP_BRANCH_SLT slti sltiu slt sltu beq bne blt bge bltu bgeu jalr jal
2 SHIFT_BITWISE slli xori srli srai ori andi sll xor srl sra or and
3 MUL_DIV mul mulh mulhsu mulhu div divu rem remu
4 MEM_WORD lw sw
5 MEM_SUBWORD lb lh lbu lhu sb sh
6 ATOMICS amoadd amoswap lr sc amoxor amoor amoand amomin amomax amominu amomaxu, each .w

5 Decoded tables#

program::decode_program(image, params) decodes every Instruction slot — a word isa::decode refuses fails the program, reachable or not (NotAllOpcodesSupported) — and builds a table for each family of the config. An instruction family's columns are committed setup columns, which each cycle's decoder lookup reads (lookup.md §10); other families' tables have none.

  • One row per halfword, absolute. Row i is pc 2i, and a table has exactly its family's height h (§7).
  • A live row holds one of the family's instructions in the fields of its lookup tuple (program::lookup_tuple): pc, next_pc, rs1, rs2, rd, imm, extra_mask, without imm for MUL_DIV and ATOMICS; no tuple holds funct3, the mask saying more. next_pc is the fall-through, pc + 2 or pc + 4 by the slot's length, never a branch target. A register the form lacks is 0; imm is the two's complement of §4's value, 0 where the form has none, or a system code (§6).
  • Every other row is padding, Fr::MINUS_ONE in every field (FamilyTable::column_poly). An all-zero row would be a claimable instruction at pc 0 with an empty mask; a live field is below 2^32, so no live row is the padding row.
  • Reach. Derivation fails (TableTooShort) unless the family's own last instruction has pc ≤ 2h − 4; another family's code may lie beyond it. Code is linked from RAM_ORIGIN = 2^16, so a family reaches 1.9375 MiB of it at 2^20 and 7.9375 MiB at 2^22, the largest height.

Every Instruction slot is a live row of exactly one table, and no table has another (check_partition).

Code is static. A cycle's instruction comes from these tables, never from RAM: a store into .text changes what a load reads, not what executes, and a pc that is not an Instruction slot has no row, so reaching it is fatal and unprovable (execution-trace.md §10).

6 The extra mask#

A tuple's last field, family_extra_mask, is 1 << kind, the kind being the instruction's position in its row of §4's table (constants::extra_mask). A kind is a mnemonic, except family 0's bit 0, the system kind, whose three instructions are told apart by imm (constants::extra_mask::system_code): ecall 0, ebreak 1, fence 2. A fence's fm, pred and succ, and an atomic's aq and rl, are not recorded; on one hart they order nothing.

One-hotness is the table's, not a gate's: a circuit holds each bit it extracts boolean, and the decoder lookup, which admits only the table's rows, is what excludes an empty or many-bit mask (lookup.md §10).

7 VmConfig and heights#

verifier_core::VmConfig { families: Vec<(family, height)>, bytecode_size_words } is a program's static shape: its families, ascending by id (constants::family), each with its height, the row count of one of its shards; an execution's shard counts are not in it. decode_program derives the family set, and nothing selects it:

  1. an instruction family (0–6), whose rows are cycles, is present when the image holds one of its instructions;
  2. a window family, whose rows are memory locations — INIT_TEARDOWN (7), ZERO_WINDOWS (8), PUBLIC_INPUT (12), PUBLIC_OUTPUT (13), ADVICE_WINDOWS (14) — is always present, and FIELD_WINDOWS (18) when one of families 19–22 is, which puts the config in the recursion format (VmConfig::is_recursion), the one whose registry VmConfig::circuit reads for 18–22 (recursion.md §1.1, §1.2);
  3. a delegation family (9–11, 15–17, 19–22), whose rows are invocations, is present when the image declares it by a record among its file bytes (program::declared_delegations, delegation.md §7); a record naming a number no family answers is UnknownDelegation.

A height is a parameter (ProgramParams::heights, defaulting to constants::family::DEFAULT_HEIGHTS, circuits.md §1), except the public families', pinned at 2^12 because it places their windows (public-values.md §2). Four things constrain it:

  • the menu, constants::family::HEIGHT_MENU: 2^8, 2^12, 2^16, 2^18, 2^20, 2^22 (HeightNotOnMenu), even powers of two as a Mercury opening needs (mercury.md §1);
  • an instruction family's code (§5);
  • the window rules (verifier_core::window_height, WindowRule; memory.md §3.5), and RAM window 0, [0, 4·h_w), holding every file byte of the image (ImageOutsideWindow);
  • the floor of the family's lookup channels, below which the registry has no circuit and no key can be built: 2^20 for an instruction family (lookup.md §3), its own for a delegation family (delegation.md §9).

bytecode_size_words, 2^20 (4 MiB) by default, is a declared ceiling on the words from RAM_ORIGIN to the image's last file byte (ProgramTooLarge); no circuit reads it.

The wire form is 8k + 8 bytes; VmConfig::from_bytes refuses a wrong length, a family id above 22, ids not strictly ascending, a height off the menu, and what window_height refuses:

k: u32 LE ‖ k × (family: u32 LE ‖ height: u32 LE) ‖ bytecode_size_words: u32 LE

In a transcript a config is one VM_CONFIG message, [f_1 … f_k, h_1 … h_k, bytecode_size_words]: the second message of the identity (§8) and the first of a statement's descriptor (proof.md §2).

8 Program identity#

One Fr, ProgramIdentity, on the wire its canonical 32 bytes (primitives.md §1): the raw squeeze of a fresh transcript after these messages (verifier_core::identity_digest; tag values in transcript.md §5):

# tag message
1 PROGRAM_IDENTITY [code_version]: constants::family::CODE_VERSION, 0, the only one derivation builds (UnsupportedCodeVersion)
2 VM_CONFIG §7's
3 PROGRAM_ENTRY [entry_pc]
4 COMMITMENT, one per family of the config, ascending its setup commitments, four limbs a point (transcript.md §4)

A family's setup commitments (program::setup_commitments) are Mercury commitments (mercury.md §2): of an instruction family's decoded-table columns in tuple order, 7 or 6 points; of INIT_TEARDOWN's image column (program::image_init_column), row y being image.initial_word(4y) over RAM window 0, one point; none, an empty message, for every other family.

Committing needs the ceremony's SRS, with as many powers as the tallest table has rows (srs.md §1). The digest over given points (program::identity_from_commitments) needs no SRS and no curve arithmetic: a verifying key carries the lists, its load recomputes the identity from them, and every shard opens its setup columns against the same points (proof.md §5, §7), which is what ties the tables a proof reads to the identity.

It binds every instruction the sweep found, with its pc, length, operands and kind, and that no other pc holds one; every file byte of the image (.text, .rodata, .data, the delegation declarations among them); the entry pc; the family set, every height, bytecode_size_words and the code version. One ELF at two settings of the heights has two identities.

It does not bind:

  • the SRS its commitments are under, or the generic lookup table: those are the SRS digest's (proof.md §3);
  • the circuits, which a key's load holds to the registry (proof.md §7);
  • anything an execution chooses: its input, advice, shard counts, window list;
  • memory past a segment's file bytes (.bss, the heap and the stack): mem_len enters nothing, and such memory starts at zero whatever is declared;
  • the symbol table, which loader::function_symbols and loader::symbol_names read beside the image for the profiler and listings, or anything else of the ELF §1 does not read.

A verifier takes the identity from a channel the prover does not control and compares it with its key's; it never sees an ELF. Against a prover-supplied identity a proof shows only that some program ran. Whoever holds the ELF, the parameters and the ceremony file recomputes it.

审计专区/程序与执行

客户程序 ABI

规范原文docs/spec/ecall-abi.md以 Markdown 查看

摘要

客户程序(guest)与机器之间的接口。它规定了 ecall 调用约定、编号范围与每一个已实现的编号(EXIT 和十个委托)、已退役的编号及其永不复用的原因、其余所有编号得到的应答、完整的 32 位地址映射以及链接脚本所强制的段布局,还有客户程序 SDK 的公开接口及其退出状态。

以下规范原文以英文维护;英文是本规范的标准语言。

The ecall convention, every ecall number, the guest's address space and the SDK over them. A guest has no file descriptors and no I/O syscall: its public input, journal and advice are memory (public-values.md), so an ecall only ends the execution or hands a frame to a circuit. crates/constants/tests/ecall_abi.rs holds this page's tables to constants::ecall and its MEMORY line to link.ld.

1 The calling convention#

Register Role
a7 the number
a0 in: the one argument, an exit status or a delegation's frame base (delegation.md §4)
a0 out: the result, 0 or a negated errno

a1–a5 are reserved for a call that needs more arguments; none does. A recursion-format delegation answers its frame base advanced past the frame (recursion.md §1.4). An ecall preserves every register but a0: its row writes no other (execution-trace.md §6), so the memory argument carries the rest across it.

2 The number ranges#

Constant Value What
ZKVM_IO_FIRST 0x0400 first host call
ZKVM_IO_LAST 0x04FF last host call
PRECOMPILE_FIRST 0x0500 first precompile
PRECOMPILE_LAST 0x05FF last precompile

Both ranges lie above 1023, the whole Linux number space, and are disjoint, so a number says its class: a host call would return a value the prover chose, a precompile is a deterministic function of guest memory that its circuit proves. The host-call range is reserved and empty: advice is a memory region the prover fills and the guest checks (public-values.md §6), inside the memory argument, where a value returned in a register would be bound to nothing.

3 Syscall numbers#

Every number this VM implements. All but EXIT are delegations, whose families, anchor spaces and frames are delegation.md §3's registry: the first six are the base format's, the last four the recursion format's (recursion.md §2).

Number Constant Class What
93 EXIT deterministic end the execution with status a0, the statement's exit status; nonzero is a failed execution, still provable
0x0500 PRECOMPILE_POSEIDON2 deterministic the width-3 Poseidon2 permutation over canonical Fr lanes
0x0502 PRECOMPILE_FR_ARITH deterministic one Fr add, multiply or inverse over Fr's in-memory form
0x0504 PRECOMPILE_MOD_MUL deterministic a·b mod m, m one of four Ethereum moduli a selector names
0x0506 PRECOMPILE_EC_ADD deterministic one third of a complete point addition, secp256k1 or BN254 G1
0x0507 PRECOMPILE_KECCAK_F deterministic one round of keccak-f[1600]; a permutation is 24 calls
0x0508 PRECOMPILE_SHA256_COMP deterministic four rounds of SHA-256's compression; a compression is 16 calls
0x0509 PRECOMPILE_FR_OP deterministic one operation over field cells
0x050A PRECOMPILE_P2_FIELD deterministic one transcript duplex step over field cells
0x050B PRECOMPILE_FIELD_IO deterministic eight RAM words into a field cell, or back
0x050C PRECOMPILE_FQ_OP deterministic one BN254 base-field operation over field cells

A class says who chooses the result. deterministic: a function of the guest's own state, which a circuit proves. advice: chosen by the prover; no number has it (§2). These are exactly the ecalls a proof admits: ADD_SUB_LUI_AUIPC holds every ecall row's a7 to 93 or to a registered delegation number, the base format's circuit knowing the first six (add-sub.md, recursion.md §1.2).

4 Retired numbers#

Number Constant Was
63 none POSIX read(fd, buf, len)
64 none POSIX write(fd, buf, len)
0x0501 RETIRED_KECCAK_F_WHOLE_PERMUTATION a whole keccak-f[1600] over a 200-byte frame, which 0x0507 replaces
0x0503 RETIRED_MOD_MUL_WITNESSED_MODULUS a·b mod m over a 128-byte frame carrying m, which 0x0504 replaces
0x0505 RETIRED_SHA256_COMP_WHOLE_COMPRESSION a whole compression over a 96-byte frame, which 0x0508 replaces

A number is assigned once. A retired one is never reassigned and answers -ENOSYS (§5): given a second meaning, it would run an old binary with its frame misread to a plausible wrong answer.

5 Every other number#

Constant Value What
ENOSYS 38 answered as -ENOSYS in a0

A number not in §3 answers -ENOSYS and falls through: the retired numbers, the host-call range, and every syscall a library might make for host data — getrandom, clock_gettime, the seeding of std's RandomState. Host data is prover advice, and a guest that needs it takes it from the advice region, where checking it is visibly the guest's job. Such a call executes and cannot be proved: the ADD_SUB_LUI_AUIPC fill refuses its row.

-ENOSYS is also the delegation ABI's "no circuit" answer, on which a base-format shim runs its software path (delegation.md §2); this executor never gives it to a §3 number. A registered number the image did not declare is the fatal DelegationFamilyAbsent on the tracing paths (delegation.md §7).

6 The memory map#

crates/guest-sdk/link.ld declares one region, constants::guest_memory's RAM_ORIGIN and RAM_LENGTH:

ld
MEMORY { RAM (rwx) : ORIGIN = 0x00010000, LENGTH = 0x7FFF0000 }

The whole 32-bit address space:

[0x0000_0000, 0x0000_8000)  hole: no family initializes it; an access is a fatal OutOfBounds
[0x0000_8000, 0x0000_C000)  public input window    PUBLIC_INPUT_ORIGIN    16 KiB
[0x0000_C000, 0x0001_0000)  journal                PUBLIC_OUTPUT_ORIGIN   16 KiB
[0x0001_0000, 0x8000_0000)  RAM                    RAM_ORIGIN, RAM_LENGTH
    0x0001_0000             .text, _start first; .rodata, .data, .bss, each page-aligned
    __heap_start            .bss's end rounded up to 16; the heap grows up from here
    0x7F80_0000             __stack_top − STACK_RESERVE (8 MiB): no heap block ends above it
    0x8000_0000             __stack_top, the initial sp; the stack grows down
[0x8000_0000, 2^32)         advice                 ADVICE_ORIGIN          up to 2^29 words
  • A load or store reaches the four regions alike, every word carrying the RAM tag; the windows' and the advice's layouts, families and binding are public-values.md §2–§6. None is in the ELF, so no linker symbol names them. Advice is addressable only up to the words the host supplied, and not at all when it supplied none.
  • The hole makes a null dereference a fatal error rather than a trace nothing could prove.
  • A delegation frame lies wholly in RAM (delegation.md §4). crates/loader refuses a PT_LOAD outside RAM (program.md §1), and a decoded table's height bounds how far .text reaches (program.md §5).
  • crt0's _start sets sp, zeroes [__bss_start, __bss_end) byte by byte, so that a zero .bss is the image's property and not the executor's, calls main, and exits 0 if it returns.
  • Nothing detects a stack that grows past its reserve after the heap has filled below it.

6.1 The segment layout#

A guest ELF loads under two loaders. crates/loader lays its PT_LOADs into a flat space the executor makes addressable whatever the headers say, with no pages and no permissions. A host loader maps exactly the PT_LOADs, page by page, at their permissions, and nothing else exists. The headers are the image's account of its own memory, read by every tool but this VM, so link.ld makes them true:

  • Every writable byte is declared. .bss runs to ORIGIN(RAM) + LENGTH(RAM), so the heap and the stack lie in one writable segment ending at __stack_top, whose file bytes stop at or before .bss, the 2 GiB reservation being NOBITS. Undeclared, the first stack push would fault.
  • No two segments share a page. .text, .rodata, .data and .bss are each 4096-aligned: a page two mappings share takes the second's permissions, stripping execute from .text's tail or putting zero fill on a read-only page, which a host loader refuses.

crates/loader/tests/layout.rs holds the committed guest ELFs to both by parsing their headers.

7 The guest-sdk surface#

crates/guest-sdk is the guest's runtime. Only exit and the delegation shims issue an ecall; the rest is loads and stores.

Item What
entry!(f) exports the main crt0 calls, a wrapper calling f
public_input() the public input payload, its length word clamped to the window
read_input(buf) copies min(buf.len(), public_input().len()) bytes and returns the count: it may return short
commit(bytes) appends to the journal and its length word; exits 70 rather than overflow the window
journal() what has been committed
advice() the advice payload, its length clamped to the region; bound by nothing, so the guest checks it
exit(code) EXIT; publishes nothing beyond what was committed
keccak256, sha256 over KECCAK_F and SHA256_COMP, with a software fallback on -ENOSYS from the first call
poseidon2_permute over POSEIDON2; false on -ENOSYS, for the caller's own permutation
ec_add, ec_mul, ec_identity homogeneous projective points over EC_ADD; None on -ENOSYS
recursion::* the raw shims over word-aligned frame types, false on -ENOSYS; the recursion format's (fr_op, p2_field, field_io, fq_op and the tape helpers import, import_run, replay) have no software path
allocator bumps up from __heap_start, never frees; exits 71 when a block would end above __stack_top − STACK_RESERVE or the live sp
panic handler exits 101 and writes nothing: a panicking guest is provable, having published what it committed

A delegation answer other than 0 or -ENOSYS exits 72, as do -ENOSYS after the first call of a multi-call operation and a recursion call that does not leave a0 past its frame. Each shim reads its number from its declaration record (delegation.md §7); which library code reaches which shim is delegation.md §10's.

审计专区/程序与执行

执行轨迹

规范原文docs/spec/execution-trace.md以 Markdown 查看

摘要

一次执行中的每个内存查询是什么、何时发生。它定义了四时间戳的周期与 38 位时钟、地址空间、作为同一地址上一次读与一次写的查询、每类指令的帧、x0 规则、ecall 行、事件日志的顺序、周期到电路族的路由、执行轨迹层面的内存自检、模拟器及其相对于宿主环境中 RV32IMAC 的三处偏差,以及存放执行轨迹的容器。

以下规范原文以英文维护;英文是本规范的标准语言。

What every memory query of an execution is, when it happens and the order the trace records it in; the emulator that produces a trace and the containers that hold one. The memory argument (memory.md) and every family's frame are built on this convention.

1 The clock#

Cycle c occupies the four timestamps 4c + Δ, one per slot Δ ∈ {0, 1, 2, 3} (constants::memory::TS_STEP). Every instruction is one cycle and nothing else is: a delegation invocation rides the cycle that requested it, so the cycle count is the instruction count.

  • Cycles are numbered from 1. Timestamp 0 is every address's initial write, and a read must strictly precede its write, so a cycle-0 pc query could not follow the value it reads.
  • The clock is 38 bits (TS_BITS): every timestamp is below 2^38, the last cycle is 2^36 − 1, and the cycle that would pass it is the fatal ClockOverflow, raised before it is recorded.

2 Address spaces#

Tag Space Address At timestamp 0
1 REG a register index, 0..32 0, x0 included
2 RAM the byte address of a 4-aligned word in RAM, a public window or the advice region the image's bytes in RAM, 0 past them; the public input's and the advice's layouts; 0 in the journal
3 PC 0 the entry point
4–9 DELEGATION_KECCAK_F … DELEGATION_EC_ADD, a base-format delegation family's anchor each a frame base no initial write
10 FIELD a cell, any u32 0
11–14 DELEGATION_FR_OP … DELEGATION_FQ_OP, a recursion family's anchor each a frame base no initial write

The tags are nonzero so that no real tuple is all zeros, as x0's initial write would be. A byte or halfword access queries its word, and trace::InitialMemory is what RAM starts from. An anchor space, in delegation.md §3's order, is a delegation family's type, not memory: a query there reads the tuple stamped 0 with value 0, whatever came before; requests pair with invocations and nothing chains (delegation.md §5). Field cells hold whole Fr elements, reached only by the recursion families (recursion.md §2).

3 A query#

A memory query is one event at one address (trace::MemoryEvent): a read of read_value, last written at read_ts, and a write of write_value at ts = 4c + Δ.

  • A query that only reads writes back what it read: a register read or a load is one query.
  • read_ts < ts, strictly; the gap ts − read_ts − 1 is below 2^38.
  • Queries at distinct addresses may share a slot; two at one address never do. An address may be queried at two slots of a cycle — add a0, a0, a1 reads a0 at slot 1 and writes it at slot 3 — which is why the log is ordered by slot.

4 The frame of each instruction class#

Slot 0 is the pc query, every cycle: pc read, next_pc written. A register query exists for every register field of the decoded instruction, whatever register it names, x0 included.

Class Δ = 1 Δ = 2 Δ = 3
lui, auipc, jal rd
jalr, register-immediate rs1 rd
branches rs1 rs2
register-register, M rs1 rs2 rd
loads rs1 the word, read rd
stores rs1 rs2 the word, the stored bytes merged in
lr.w rs1 the word, written back; rd ← it
sc.w rs1 rs2 the word ← rs2; rd ← 0
AMOs rs1 rs2 the word ← op(old, rs2); rd ← old
fence
ecall a7 a0 (§6) a0 ← the result; a delegation's mirror query

ebreak has no row (§10). An atomic's row and a delegation request's carry two queries at slot 3, at distinct addresses. A family's frame is the union of its instructions' queries (memory.md §2). next_pc is the fall-through — pc + 2 after a compressed instruction, pc + 4 otherwise — except a jal's or taken branch's pc + imm, a jalr's (rs1 + imm) & !1, and the exit row's HALT_PC (memory.md §5).

An invocation's accesses ride its requesting cycle but belong to its own family's row: its frame words in RAM at slot 0 (constants::delegation::FRAME_DELTA), a FIELD_IO invocation's eight data words in RAM at slot 1 (constants::field_io::DATA_DELTA), and a recursion family's field cells at slots of its own (delegation.md §4, recursion.md §2.1).

5 The x0 rule#

x0 is an ordinary register in the trace and a constant in the machine: it starts at 0, a read of it is a REG query at address 0, and an instruction whose rd is x0 logs its slot-3 write with value 0, whatever it computed. So every query at x0 reads and writes 0, which the x0 gadget enforces (memory.md §2).

6 ecall#

An ecall's row is one cycle of ADD_SUB_LUI_AUIPC. It reads a7 at slot 1 and writes a0 at slot 3; the rest depends on the number (ecall-abi.md):

a7 Δ = 2 a0 written next_pc Besides
EXIT a0, the status the status HALT_PC the execution stops
a delegation number a0, the frame base 0, or for a recursion type the base past the frame (recursion.md §1.4) fall-through the mirror query at the frame base (§7); the invocation (§4)
any other none -ENOSYS fall-through no proof admits the row (ecall-abi.md §3)

7 The order of the log#

Events are recorded in cycle order and, within a cycle, by slot and then by role: the pc query; then an invocation riding the cycle, its frame words in frame order and a FIELD_IO invocation's data words after them; then one query per role the row has, in trace::ROLES order:

Role Slot Space What
rs1 1 REG rs1; an ecall's a7
rs2 2 REG rs2; an ecall's argument a0
load 2 RAM a load's word
ram 3 RAM a store's or an atomic's word
rd 3 REG rd; an ecall's result a0
delegate 3 the requested family's anchor space a delegation request's mirror query

ROLES is in slot order, so the log is in timestamp order, which MemoryState::record asserts; trace::Row::present holds one bit per role in a u8, and no two roles share a (space, slot) pair. The atomics family keeps its word at slot 3 for every instruction, lr.w included, so one frame serves the whole A extension.

8 Routing#

Every cycle goes to the one family whose decoded table claims its pc (program.md §4); a pc no table claims, or one claimed by a family not its instruction's, panics the tracer. An invocation goes by its type to its family's buffer.

9 The trace-level memory check#

MemoryEventLog::self_check(&InitialMemory) runs the memory argument natively over a whole log. First the timestamp rules: every address one its space has, every timestamp on the clock and in order, every read before its write, one query per address and timestamp. Then the balance: as multisets of (space, address, timestamp, value), an initial write at timestamp 0 of every touched address plus every query's write equals every query's read plus a teardown read of every address's last write. With one write per address and timestamp and no negative gap, this pairs each read with the last write before it: sequential consistency. What it cannot see:

  • Teardown is each address's last write, taken from the log, so everything after an address's last honest query balances by construction: a final value changed, a final query moved later or added, trailing cycles removed. In a proof the final values are the boundary scalars and the window families' teardown columns, fixed before any memory challenge, and the verifier fixes x0's and the pc's (memory.md §4, §5).
  • An anchor-space query is credited with its invocation's two tuples and balances alone; that requests and invocations pair 1:1 is the circuits' (delegation.md §5).
  • Field-cell accesses are not events (recursion.md §2.1).

10 The emulator#

crates/emulator runs RV32IMAC on one hart over a ProgramImage, with no interrupts and no privilege levels; aq/rl and fence order nothing. emulator::run returns an Execution: the registers, the exit status, the cycle count and the public values. emulator::trace_run returns the family buffers, the MemoryEventLog and the CycleProfile too, and emulator::StreamingRun, the prover's pull-based tracer, hands over a family's buffer as a ShardChunk the moment it reaches its height (streaming.md §2). The three differ only in what records a cycle, and a run is a pure function of (image, io), with no clock, randomness or threads, so two runs cut the same shards. A nonzero exit status is an execution, not an error.

Three points differ from a hosted RV32IMAC. sc.w always succeeds, storing and writing 0, as the circuits do (memory-ops.md §6). A misaligned halfword or word access is fatal, never split. The instruction stream is the image decoded at load, so a store into .text changes RAM and not what executes.

Every other stop is a fatal EmuError, and run and trace_run return no trace beside one: NotAnInstruction (the all-zero halfword included), IllegalInstruction, Ebreak, Misaligned (a frame base too), OutOfBounds (an access outside ecall-abi.md §6's regions, an advice word past what the host supplied, or a frame not wholly in RAM), ClockOverflow, PublicInputTooLong and JournalTooLong (the input, or the journal's length word at exit, above a window's payload), DelegationFamilyAbsent (on the tracing paths, a delegation number the image did not declare) and DelegationFrame (a frame its family has no witness for, delegation.md §6). An unassigned ecall number is not an error but -ENOSYS (§6).

There is no second executor: crates/emulator/tests/trace.rs restates §4's table and checks every traced row against it, and §9's check and the checker's multiset, memory and family-row suites hold the rest.

11 Trace containers#

crates/trace holds what an execution leaves; the emulator is its only producer.

  • Family buffers. trace::FamilyTraces holds one buffer per family of the VmConfig. A FamilyTrace, empty for a window family, is raw live rows, column-major, in small integer types: cycle, pc, next_pc, present, and per role addr, read_ts, read_value, write_value; no padding, no polynomial. A row stores everything its queries carry but a write timestamp, 4c + Δ, and the pc query's read timestamp, 4(c − 1). A delegation family's DelegationTrace has a row per invocation: the requesting cycle, the frame base, the frame words, and a recursion family's cell and data-word accesses.
  • RowSlice, FrameSlice. One shard's rows, [i·h, min((i + 1)·h, len)), borrowed: what the memory column builders read, never the log (memory.md §2). Row::delegation_space recovers a mirror query's space from the a7 the row read.
  • MemoryState, the last-access tables: each register's, the pc's, each RAM word's and each field cell's last (ts, value). O(touched addresses), and all the register and pc boundary, the RAM window list (trace::init_windows) and the window families' teardown need.
  • MemoryEventLog, the events and a MemoryState: O(cycles), kept only by trace_run, read by §9's check, the TraceArchive and checker::memory_columns_from_log, the independent reading the column builders are held to.
  • TraceArchive, the post-execution snapshot: buffers, log, profile, public values and advice. Its file is two postcard values, five phase sections and then their timings, so the deterministic payload is a byte prefix of it, and only a canonical encoding of self-consistent parts is read back. No proving path reads one; checker::TamperHarness and the retained archived path do (streaming.md §6).
  • CycleProfile, ShardPlan. The profile counts rows per family, cycles for a cycle-owning family (summing to the cycle count) and invocations for a delegation family. trace::plan_shards is ⌈count / height⌉ per family; a window family plans 0 there, its count being the prover's (streaming.md §4).

审计专区/程序与执行

公开值与证明者提示

规范原文docs/spec/public-values.md以 Markdown 查看

摘要

一次执行的输入与输出如何成为其陈述的一部分。它把公开输入、公开输出(journal)和证明者提示(advice)放进内存窗口,定义它们带长度前缀的布局以及初始化它们的窗口电路族,并阐明绑定方式:输入与公开输出的摘要在任何挑战之前进入 transcript;多重集使这些窗口的首值与末值就是这次执行中的值;在一个随机点上的检查使这些窗口与陈述中的字节相符。证明者提示在设计上不绑定任何东西。

以下规范原文以英文维护;英文是本规范的标准语言。

How an execution's public input and public output, the journal, are bound to its proof, and what the prover's advice is. All three are regions of guest memory, each initialized by a window family of its own and carried by the memory argument (memory.md).

1 Three regions, no I/O syscall#

region contents chosen by bound by
public input the statement step 10c, to the statement's input (§5)
journal the guest's stores step 10c, to the statement's output (§5)
advice the prover nothing (§6)

There is no I/O syscall: a guest uses ordinary loads and stores, and a provable guest's only ecalls are EXIT and delegation numbers (the retired POSIX numbers: ecall-abi.md §4). A host supplies emulator::GuestIo's input and advice and reads the journal from emulator::Execution::io. A byte-moving syscall would need cross-row constraints tying each transfer row to its buffer and length, which this arithmetization has no place for, while a window is bound by the multiset and one comparison (§5) and asks nothing of the guest. Nor does the guest hash its output: nothing rests on its honesty, and a panic loses nothing it committed.

2 The memory map#

region bytes family windows
public input [0x8000, 0xC000) PUBLIC_INPUT PUBLIC_INPUT_WINDOW = 2, at 2^12
journal [0xC000, 0x1_0000) PUBLIC_OUTPUT PUBLIC_OUTPUT_WINDOW = 3, at 2^12
advice [0x8000_0000, 2^32) ADVICE_WINDOWS from 2^29/h, at the window height h

The full address map is ecall-abi.md §6. The public windows take the upper half of [0, RAM_ORIGIN), 64 KiB that no RAM window family initializes (INIT_TEARDOWN masks window 0's rows below RAM_ORIGIN and no ZERO_WINDOWS id is 0, memory.md §3), so they cost RAM nothing, and [0, 0x8000) stays a hole in which a null dereference cannot balance.

A window's first address is 4·height·id, so the pinned height constants::family::PUBLIC_WINDOW_HEIGHT = 2^12 is what makes the origins windows 2 and 3, 16 KiB each, ending flush against RAM_ORIGIN. It is the ceiling: at the next menu height, 2^14, two windows need 128 KiB, and the one window in the hole is window 0, which would initialize address zero. Anything larger means moving RAM_ORIGIN, which moves every program's load address and shortens every decoded table's pc reach (program.md §5).

program::decode_program assigns that height whatever its caller asks, and the verifier refuses any other, and any RAM window height that would let a zero window reach the public windows (memory.md §3.5).

3 Layout and the length word#

word 0       the payload's byte length
words 1 …    the payload, little-endian, zero-padded to the end of the window

A public window is 2^12 words, so a payload is at most guest_memory::PUBLIC_PAYLOAD_BYTES = 16,380 bytes. verifier_core::public_io_words is the one spelling: the executor seeds the input window with it, the prover commits it and the verifier evaluates it. The length word makes the binding exact: without it [1, 2, 3] and [1, 2, 3, 0] fill the same window. verifier_core::derive_global_phase refuses an input or output longer than 16,380 bytes as Statement, and the executor refuses such an input before the first cycle.

4 The window families#

family id height shards init leaf step 10c holds
PUBLIC_INPUT 12 2^12 exactly 1 M[2] init_value M[2] to input
PUBLIC_OUTPUT 13 2^12 exactly 1 literal 0 M[1] teardown_value to output
ADVICE_WINDOWS 14 h k ≥ 0 M[2] init_value nothing

All three are in every VmConfig and own no cycles. Each public family proves exactly one shard in every statement (memory.md §3.5), so step 10c always runs: an unread input is still the window's initial contents, and an unwritten journal is empty.

The circuits are memory.md §3.3's. PUBLIC_OUTPUT's is ZERO_WINDOWS' byte for byte, whose init leaf writes the literal 0, so no column holds an initial journal (§5). PUBLIC_INPUT's and ADVICE_WINDOWS' initial values are M[2], one execution's values, committed before the memory challenges and bound by no program identity.

All three regions' tuples carry constants::address_space::RAM; which family initializes an address is what makes a word public, advice or heap. A space of their own would need an address-space column, and a gate pinning it, on the memory path of MEM_WORD, MEM_SUBWORD and ATOMICS; under RAM those circuits need nothing for them, their addressing already covering every 4-aligned address below 2^32 (memory-ops.md §2).

5 The binding#

io_digest absorbs the statement's two strings in a transcript of its own (transcript::io_digest):

t ← Transcript::new()
t.append_bytes(PUBLIC_INPUT_STREAM,  input)      tag, byte length, 31-byte limbs
t.append_bytes(PUBLIC_OUTPUT_STREAM, output)
io_digest ← t.sample()                           one raw squeeze

The framing (transcript.md §3) parses back to exactly one ordered pair, and the squeeze is raw, as every digest's is. The guest never computes it. G7 absorbs it before the memory commitments (G8) and challenges (G10) (proof.md §2), so both strings are fixed before any challenge exists.

The multiset. At a window address the init leaf is the only write at timestamp 0, every access consumes a write and produces a strictly later one, and the teardown balances only against the last (memory.md §9). So PUBLIC_INPUT's M[2] holds each word's value before its first access, and PUBLIC_OUTPUT's M[1] its value at the end.

Step 10c of verifier_core::verify_shard_local (proof.md §6). Of a public shard's base claims, which share one point u and each name a column, the verifier takes the one on M[2] (PUBLIC_INPUT) or M[1] (PUBLIC_OUTPUT), refusing its absence as Malformed, and compares it with its own evaluation at u of the multilinear extension of public_io_words(input) or public_io_words(output). A mismatch is MemoryArgument; the shard's opening then holds the claim to the committed column. Column and string are fixed before u is drawn, so a column other than the window passes with probability at most 12/p.

PUBLIC_INPUT's teardown is free: a guest may overwrite its input. PUBLIC_OUTPUT has no init column, and that is the point: with one, a prover could place the journal there at timestamp 0 and the teardown would match without the guest storing a byte.

5.1 The argument, stated plainly#

G7 fixes input and output, and G8 the window columns, before any challenge. Step 10c says the columns are those strings' windows; the multiset says they are the execution's first values in the input window and its last values in the journal window. So the guest found the statement's input in its input window, and the statement's output is what its stores left in the journal window. That rests on no cooperation, hash or register convention of the guest's, and says nothing about advice.

Recursion carries the binding unchanged: a node recomputes io_digest from the windows' words and repeats step 10c over them, and the decider binds the contract's input and output calldata to io_digest (recursion.md §8.1, §9).

6 Advice#

Advice is memory whose initial values the prover chose: ADVICE_WINDOWS initializes [ADVICE_ORIGIN, ADVICE_ORIGIN + 4hk) from an M[2] that nothing binds, not identity, not the statement, not a gate. A guest reads it with ordinary loads.

  • Layout. §3's framing over 1 + ⌈len/4⌉ words (trace::advice_region_words), spelled once by trace::advice_word for the executor and the prover; guest_sdk::advice reads it back.
  • Windows. At the window families' one height h, shard i is window verifier_core::advice_first_window(h) + i, and advice_first_window(h) = 2^29/h is the first window above RAM. Consecutive, they need no list: a statement carries only their count k = ⌈words/h⌉ (trace::advice_window_count), which covers what the host supplied, an untouched word's tuples cancelling. check_memory_windows asks only 2^29/h + k ≤ 2^30/h, the top of the address space, and ZERO_WINDOWS ids stay below 2^29/h (memory.md §3).
  • No advice, no region. Then k = 0 and there is no shard; guest_sdk::advice on such a run is a fatal emulator::EmuError::OutOfBounds.
  • Not read-only. A store there is an ordinary store. Refusing it would need a space selector and a gate on three families' memory path, and would buy nothing: advice is unbound either way.

What a guest owes. A proof says that some advice exists under which the program, given the public input, published the journal; advice that changes the journal unchecked is a value the prover chose. The check is against something the proof binds: a commitment in the public input (guests/public-io, at toy scale, with a position-weighted checksum standing in for a hash), or one the journal publishes. revm-block-stateless publishes the root of the payload it validated and holds its witness to that payload by hashes (ethereum.md §4).

7 The guest's view#

A guest reaches the regions with loads and stores at the constants::guest_memory constants, through guest_sdk::public_input, guest_sdk::commit and guest_sdk::advice, none of which issues an ecall; ecall-abi.md §7 is the API and the guest program manual the walkthrough. Nothing is published at exit, so a guest that panics has published what it committed, and its run is proved like any other.

8 Cost#

  • No address space, transcript message, tag, challenge or statement field; no gate elsewhere.
  • Two 2^12-row shards a statement, five committed columns between them; one h-row shard of three columns per advice window.
  • The native verifier: two 4,096-point multilinear evaluations, 4,095 multiplications each. A recursion node's cost follows the payload instead: it evaluates the payload's words alone, times 1 − r_j for each variable above them (verifier_core::chain::public_value).
  • The guest: nothing at exit; a byte store per journal byte and a word store per commit.

9 Limits#

  • 16,380 bytes each, and no larger window (§2). A journal that grows with the execution has no fixed bound: the mini-block binary's, a 13-byte record plus return data per transaction (ethereum.md §3), holds at most 1,255 transactions, and one record can exceed it. Large outputs belong behind a digest (the stateless binary's journal is 43 bytes), large inputs in advice.
  • The journal is the window's whole final contents. Anything but a length of at most 16,380, that many bytes, then zeros, matches no statement: the executor refuses an oversized length (EmuError::JournalTooLong), and a nonzero byte past it fails step 10c. commit keeps that form; a guest writing the window directly must.
  • Nothing orders the journal's writes, and nothing forces a guest to read its input. The proof binds a window's contents, not its accesses.
  • Read the exit status first. It is x10's final value (memory.md §4): a failed run, a panic included, has a verifying proof and a journal too (§7).
  • A deployed contract fixes both lengths, a decider key being per shape (recursion.md §9).

审计专区/证明系统

GKR 引擎

规范原文docs/spec/gkr.md以 Markdown 查看

摘要

每个电路族的电路都由这个引擎证明。它定义了带逐行门列表与折半门列表的层模型、每个多项式的地址、虚拟表、七种门形状与次数上限 2、电路制品及其序列化格式、四条法则与填充约定,以及反向过程:调用方必须先提供什么、transcript 时序、层求和校验、每个挑战为何在那个时机抽取,以及验证者返回的错误。

以下规范原文以英文维护;英文是本规范的标准语言。

The layered-circuit model every family circuit is written in, the artifact that carries one, its laws, and the backward pass reducing a circuit's outputs to claims on its committed columns at one point, which the shard's opening discharges (proof.md §5).

crates/constraints is §1–§4; crates/gkr-verify is §5's verifier half and the verifier's helpers for the memory argument (memory.md §3, §4) and LogUp (lookup.md §2, §8). Both are no_std, as verifier-core and the recursion guest build on them. crates/gkr, std and rayon, is the prover half and re-exports gkr-verify. crates/checker enforces §4.2–§4.3 again (circuits.md §3).

1 The layer model#

Layer k, 0 ≤ k ≤ N, N ≥ 1, is w_k columns of n_k variables, indexed as primitives.md §6 fixes. Layer 0 is the committed columns M, W, S in layout order at n_0 = trace_vars, beside the virtual tables the artifact lists (§2.1), which count in no width. Gate list k reads layer k and writes layer k + 1; the top, layer N, is exactly the outputs. A list is row-wise, n_{k+1} = n_k, or halving, n_{k+1} = n_k − 1.

A halving list halves each column of its layer: it writes w_k columns by halving shapes (§3) reading layer-k columns at both children — child 0 is rows [0, h), child 1 rows [h, 2h), h = 2^{n_k−1}, the child bit being the highest variable. An entry may read any column, as a fraction tree's numerator reads its denominator (lookup.md §6), but every column is read (§4.2). Only halving lists hold halving shapes; a halving list is never list 0, has no cached or enforcing entries and needs n_k ≥ 1. Every relation has the one template checker dump prints:

producing, row-wise   L{k+1}[j](x) = Σ_y eq(x, y)·G(layer k at y)
producing, halving    L{k+1}[j](x) = Σ_y eq(x, y)·G(layer k at (y, 0) and (y, 1))
enforcing             0 = G(layer k at y)   for every y ∈ {0,1}^{n_k}

2 Addresses#

constraints::PolyAddress names every polynomial; dumps use its Display notation:

variant notation read by
Memory(i), Witness(i), Setup(i) M[i], W[i], S[i] committed columns list 0, relations, lookups
Virtual(kind) V[row], … virtual tables, §2.1 the same, if virtuals lists it
Inner { layer, offset } L{k}[j] column j of layer k ≥ 1 list k
Cached { layer, offset } C{k}[j] cached entry j of list k, §3.1 list k
Scratch(i) scratch[i] an intermediate of the flat relation list, §4 relations

The scratch bijection maps each scratch[i] to one L{k}[j], covering every inner column once. A committed value needed above layer 1 is carried up by copy gates. M, W and S differ in when they are bound (memory.md §8).

2.1 Virtual tables#

A virtual table is a closed form, evaluated per row by gkr_verify::virtual_at_row and at a point by virtual_at_point, never materialized, committed or claimed. Each form is its table's multilinear extension, so the verifier evaluates what the prover sums (crates/gkr/tests/{lookup,ram_live}.rs check all but V[row]). Wire form: a u32, in table order from 0.

kind notation value at row y closed form at (y_0, …, y_{n−1})
RowIndex V[row] y Σ_{j<n} 2^j·y_j
RamLive V[ram_live] 1 if y ≥ 2^14, else 0 1 − Π_{14≤j<n} (1 − y_j); 0 if n ≤ 14
Range19 V[range19] y mod 2^19 Σ_{j<min(19,n)} 2^j·y_j
Range16 V[range16] y mod 2^16 Σ_{j<min(16,n)} 2^j·y_j
Xor8A V[xor8_a] a = y mod 2^8 Σ_{j<8} 2^j·y_j
Xor8B V[xor8_b] b = ⌊y/2^8⌋ mod 2^8 Σ_{j<8} 2^j·y_{j+8}
Xor8Out V[xor8_out] a ⊕ b Σ_{j<8} 2^j·(y_j + y_{j+8} − 2·y_j·y_{j+8})

14 is constants::memory::RAM_LIVE_BIT (memory.md §3); the range and XOR8 kinds are channel tables (lookup.md §3). Xor8Out's form is multilinear because y ⊕ z = y + z − 2yz is.

3 Gate shapes#

constraints::GateDef is a closed enum. A coefficient is Coeff::Literal(Fr) or Coeff::Challenge(slot), a constants::challenge_slot read from the pass's ExternalChallenges, of degree 0.

tag variant value
0 Linear { terms, constant } Σ c_i·x_i + c_0
1 Product { coeff, left, right } c·x·y
2 MaskIntoIdentity { input, mask } x·m + (1 − m)
3 AffineProduct { left, left_constant, right, right_constant } (Σ a_i·x_i + a_0)·(Σ b_j·y_j + b_0)
4 TreeProduct { input } x(·,0)·x(·,1)
5 Quadratic { constant, linear, products } c_0 + Σ a_i·x_i + Σ b_j·y_j·z_j
6 TreeCross { left, right } p(·,0)·q(·,1) + p(·,1)·q(·,0)

Quadratic spells degree-2 relations, such as a·b + c·d − e·f, that no product of affine forms does. The kernel, gkr_verify::eval_gate, takes one value per operand in GateDef::operands order, a halving shape's each at child 0 then child 1, and is the semantic authority. Both passes reach it through gkr_verify::ResolvedList, crates/checker calls it over the relations, and verifier_core::tape transcribes it for the recursion nodes (recursion.md §7).

3.1 Cached entries and the degree ceiling#

A cached entry C{k}[j] = H is a sub-expression of row-wise list k over its layer's columns, not another cached entry, substituted into the gates of its list naming it, with no table, claim or width. The prover evaluates H at every round node and never binds it: a bound table is the extension of H's values, which for a degree-2 H is not H of the extensions. No registered circuit has one. CircuitArtifact::inline_cached writes a Product with one Linear cached factor as an AffineProduct and refuses any other reference; both prove the same bytes.

Degree is read from the shape after substitution — a column or virtual table 1, a challenge 0, C{k}[j] its expression's, a halving shape 2, a Quadratic its widest term — and validate holds every gate, cached entry and relation to at most 2, so a higher relation is split across layers. With eq multilinear, every round polynomial is then a cubic (§5.3).

4 The circuit artifact#

constraints::CircuitArtifact holds a circuit twice: as layered gates, which the engine proves, and as a flat relation list over M, W, S, V and scratch, which the row-local checks read (circuits.md §3). Law 4 makes them one constraint set. In wire order:

CircuitArtifact = (format_version = 1, coefficient_encoding = 0, trace_vars ≤ 30,
                   memory, witness, setup: [name], virtuals: [(VirtualKind, name)],
                   layers: [LayerSpec], relations: [Relation], lookups: [LookupExpr],
                   scratch: [(name, L{k}[j])], outputs: [L{N}[j]],
                   padding: (row: [Fr], zero_row_valid: bool))
LayerSpec       = (halving, num_vars, width,
                   cached:    [(name, C{k}[j], GateDef)],
                   producing: [(relation, L{k+1}[j], GateDef)],
                   enforcing: [(relation, GateDef)])
Relation        = (name, output: Option<scratch index>, GateDef)
LookupExpr      = (name, channel, selector: PolyAddress, tuple: [GateDef])

validate holds the first three to those values and every name to non-empty [a-z0-9_], unique in the artifact; names mean nothing to the engine. Encoding 0, COEFFICIENT_ENCODING_CANONICAL_LE, is every Fr canonical 32-byte little-endian, and 30 is MAX_TRACE_VARS. outputs orders the top layer as OutputClaims lists it; a relation with an output defines that slot, one without is enforcing; lookups are lookup.md §1's.

4.1 Wire form#

postcard over §4's tuples, hand-written serde: a u32 is a varint, a u8 tag and a bool a byte, an Option a tag byte, a sequence a varint count then its elements, a name a str, an Fr its 32 canonical bytes.

PolyAddress  (tag u8, a u32, b u32): 0 M, 1 W, 2 S, 5 scratch (a = index); 3 V (a = kind);
             4 L, 6 C (a = layer, b = offset); unused fields 0
Coeff        (tag u8, slot u32, value Fr): 0 literal (slot 0), 1 challenge (value 0)
GateDef      (tag u8, split u32, coefficients [Coeff], operands [PolyAddress] in operands() order)
  0 Linear            split 0  c_1..c_t, c_0                 x_1..x_t
  1 Product           split 0  c                             x, y
  2 MaskIntoIdentity  split 0  —                             x, m
  3 AffineProduct     split t  a_1..a_t, a_0, b_1..b_u, b_0  x_1..x_t, y_1..y_u
  4 TreeProduct       split 0  —                             x
  5 Quadratic         split t  c_0, a_1..a_t, b_1..b_u       x_1..x_t, y_1, z_1, …, y_u, z_u
  6 TreeCross         split 0  —                             p, q

CircuitArtifact::from_bytes refuses a format_version other than 1 before decoding the rest, postcard not being self-describing; refuses an unknown tag, a nonzero unused field, a gate with counts its shape lacks and a non-canonical Fr; re-encodes and compares, as postcard admits overlong varints and trailing bytes; never panics or reserves what a declared length asks; and checks no law.

4.2 The laws#

CircuitArtifact::validate runs once where an artifact is built or loaded, never per proof: each constraints constructor panics on a refusal, and verifier_core::VerifyingKey::check applies it to a key's circuits, for prover and verifier (proof.md §7). checker::check_laws enforces Laws 1–4 and the lookup rules again, sharing no code with crates/constraints/src/laws.rs (circuits.md §3).

  1. Locality. Every operand of list k is in range and readable at layer k (§2): a V only if listed, a C{k}[j] only one of list k's own, from a producing or enforcing gate.
  2. Derived width. A list's stored width is its producing count, entry j writes L{k+1}[j], and its stored num_vars is n_k, or n_k − 1 if halving.
  3. Top layer. outputs is a permutation of L{N}[0..w_N).
  4. Single source of truth. Relations and gate entries correspond one to one, a producing entry's relation defining the slot the bijection maps to its output, an enforcing entry's none, and each pair is one polynomial, scratch read through the bijection and cached entries substituted: validate compares normalized expansions, checker evaluations at random points.

validate also refuses, each a ConstraintError naming what broke: §4's bounds, no gate list, padding.row not w_0 long, a virtual kind listed twice, §1's halving rules, degree above 2, a relation reading anything but M, W, S, listed V and existing scratch, a scratch list that is no bijection onto the inner columns or not defined once each, a slot outside constants::challenge_slot, and a relation constructed and then dropped — an inner column below the top the list above never reads, a cached entry no gate names, an enforcing gate whose expansion is zero. Reads are decided on normalized expansions: x − x and 0·x read nothing.

The lookup rules. A lookup's channel is in constants::lookup_channel; its tuple is one expression on a range channel, else 1 to lookup_channel::MAX_TUPLE (7), as wide as its channel's other lookups'; its selector is an in-range committed column some enforcing gate of list 0 holds to booleanity (x − x² up to normal form); and each expression is Linear over in-range committed columns and listed virtual tables, with literal coefficients, unit and constant-free above position 0 (lookup.md says what each protects).

4.3 The padding contract#

The engine gates nothing, an enforcing gate being a zerocheck over the whole cube, so a family switches relations off with its own columns (memory.md §2). On padding.row, a committed row, the row-local scratch values, those of producing relations not at or above a halving shape, make every row-local enforcing relation vanish at every challenge value and row index; zero_row_valid says whether the all-zero row does too. The product-tree clause: where shards have inactive rows, every column the first halving list reads is 1 on padding.row, so padding leaves each product unchanged; the RAM window families (memory.md §3) and the columns a TreeCross reads (lookup.md §6) are exempt. This is completeness, not soundness: a cheating prover's padding rows are its family's gates' business. Nor is padding.row the row a prover writes, multiplicities and setup columns differing; no prover or verifier reads it, and checker::check_padding and checker::check_padding_identity test it.

5 The backward pass#

gkr::forward materializes every layer from the committed columns; gkr::prove proves those values as they stand, one sumcheck::SumcheckProof per transition; gkr_verify::verify replays the schedule, checking, from OutputClaims, one table per output, to BaseClaims or a GkrError. gkr::self_check, naming the first failing gate, row and relation, and gkr::explain_self_check, listing that row's operands, are a debugging hook costing a second forward pass (tools.md §3). Rayon splits rows and row pairs, never lists or rounds: proofs do not depend on the thread count.

5.1 What the caller owes#

  • The base is bound into the transcript before prove or verify, which absorb none of it (proof.md §4 binds a shard's commitments).
  • Each challenge is drawn after every committed column its gates reach is bound, or is derived: a fixed function of such challenges and of statement data bound before them, computed by the verifier. That suffices for GKR; the memory argument needs more (memory.md §8).
  • The artifact has passed validate (§4.2) and is not checked again; on a lawless one the engine may panic, and verify may accept.
  • The prover's inputs have the artifact's shape; it checks none, nor that its values satisfy the gates. Soundness is verify's alone and a cheating prover runs none of this code, so a bad input costs the honest prover only a panic or a failing proof.

5.2 The transcript schedule#

prove and verify run these steps and end in one sponge state; the tags are transcript.md §5's. p is the claim point, v_j the claim on column j of the layer the next list writes.

step op tag message
O1 absorb GKR_OUTPUTS the output tables in output-map order, rows in index order: one message of w_N·2^{n_N} scalars
O2 squeeze ×n_N GKR_OUTPUT_POINT p = r, r_i binding variable i; v_j = tables[i](r) for outputs[i] = L{N}[j]
L1 squeeze GKR_BATCH λ; the claim is c = Σ_j λ^j·v_j
L2 ×n_{k+1}: absorb, squeeze SUMCHECK_ROUND, SUMCHECK_CHALLENGE a round's cubic, then ρ_i, binding variable i
L3 absorb GKR_LAYER_CLAIMS row-wise: L{k}[j](ρ) per j in offset order, layout order at k = 0; halving: L{k}[j](ρ,0), L{k}[j](ρ,1) per j
L4 squeeze, halving only GKR_CHILD τ; p = (ρ, τ); v_j = L{k}[j](ρ,0) + τ·(L{k}[j](ρ,1) − L{k}[j](ρ,0))

L1–L4 run for k = N − 1 down to 0; after a row-wise list p = ρ and v is L3's message. The base claims are layer 0's, in layout order at one point. Every registered circuit halves to a top with no variables (circuits.md §2), so O2 draws nothing and O1 fixes the roots before λ.

5.3 The layer sumcheck#

Transition k proves c = Σ_{y∈{0,1}^{n_{k+1}}} eq(p, y)·S_k(y), where

row-wise   S_k(y) = Σ_j λ^j·G_j(layer k at y) + Σ_e λ^{w_{k+1}+e}·E_e(layer k at y)
halving    S_k(y) = Σ_j λ^j·G_j(layer k at (y, 0) and (y, 1))

G_j writes L{k+1}[j] and E_e, the list's e-th enforcing gate, claims 0: enforcing gates are zerochecks sharing the descending point and its batch. The rounds are primitives.md §7's cubics, run from c, one per variable of layer k + 1, a halving list's two children being separate tables. After L3 the verifier checks claim = eq(p, ρ)·S_k(values), layer-k operands taking L3's values, virtual tables their closed form at ρ, cached entries their expression; with n_{k+1} = 0 there are no rounds and the check is c = S_k(values). A zero claim is legal. gkr::prove_sumcheck and gkr_verify::verify_sumcheck run L2.

5.4 Why it is sound#

Each challenge is drawn after what it protects:

  • r after the outputs, or a prover predicting r claims another table agreeing with the true one there.
  • λ after the claims and p. If some v_j is not the true v̂_j, or some E_e is nonzero on the cube, Σ_j λ^j·(v_j − v̂_j) − Σ_e λ^{w_{k+1}+e}·Ê_e(p) is a nonzero polynomial in λ of degree below w_{k+1} + |E_k|; Ê_e, the extension of E_e's values, is fixed before p is drawn and vanishes there with probability at most n_{k+1}/|Fr|.
  • ρ_i after round i: a wrong cubic agrees with the true one there with chance ≤ 3/|Fr|.
  • τ after both children: a wrong pair's line meets τ ↦ L{k}[j](ρ, τ) in at most one point.

Summed over a registered circuit's transitions at its default height, these stay under 2^14/|Fr|. The random-oracle assumption is architecture.md's.

5.5 Shapes and errors#

Transition k carries n_{k+1} rounds and w_k claims, 2·w_k if halving, so a proof's shape is the artifact's alone (wire form: proof.md §9). verify checks, in order and before touching the transcript, and on a validated artifact never panics on proof or claim data:

GkrError when
MissingChallenge { slot } a gate names a slot not supplied
OutputShape OutputClaims mismatches the output map in count or variables
ProofShape { layer } layer = N: a wrong transition count; else transition layer, lowest first, has a wrong round or claim count
LayerInconsistency { layer } a round or the final check of transition layer fails

One LayerInconsistency covers a wrong descending claim and a violated enforcing gate alike: a batched sum cannot tell them apart, and the proof spends nothing on it. proof.md §6 maps these errors to its classes.

审计专区/证明系统

电路

规范原文docs/spec/circuits.md以 Markdown 查看

摘要

全部 23 个电路族电路的注册表,列出每个电路的高度、已承诺列、门、查找、内部列、制品大小与分片证明大小;电路族的电路如何由内存叶子、查找分式、约束门、逐行树和折半列表组装而成;以及独立的 checker 如何重新校验各项法则、填充约定和查找规则,还有证明伪造会被拒绝的篡改测试套件。

以下规范原文以英文维护;英文是本规范的标准语言。

Every shard is proved by its family's circuit, a constraints::CircuitArtifact in gkr.md's model, fixed by the format, the family and the height. This page lists the circuits and their shapes (§1), how one is assembled (§2) and how crates/checker checks one independently (§3); each family's own page specifies its columns, gates and lookups.

1 The registry#

constraints::family_circuit(family, trace_vars) is the base format's registry, constraints::recursion_circuit the recursion format's, and VmConfig::circuit picks one by format (recursion.md §1.1). Each returns a FamilyCircuit, the artifact and its channel specs (lookup.md §11). A verifying key loads only if its circuits are the registry's at its heights (proof.md §7), and the prover registers the same (§2).

Families 0–6 (constants::family) are the execution families, one executed instruction a row (add-sub.md, jump-branch-slt.md, shift-bitwise.md, mul-div.md, memory-ops.md §3, §4, §6); 7–8 and 12–14 the window families, one memory word a row (memory.md §3, public-values.md §4); 9–11 and 15–17 the delegation families, one invocation a row (delegation-circuits.md §2 to §7, by id); 18–22 the recursion format's (recursion.md §2 to §6).

Shapes at the default height 2^n (constants::family::DEFAULT_HEIGHTS): committed columns, enforcing gates, obligations per channel (TIMESTAMP/RANGE16/GENERIC/DECODER/XOR8), row-wise gate lists (the halving ones are n), inner columns, artifact bytes, and a base-format shard proof's bytes, proof.md §9's layout over the shape:

id family n M W S gates lookups row-wise inner bytes proof
0 ADD_SUB_LUI_AUIPC 22 27 35 7 63 10/4/0/1/0 5 314 72,064 64,764
recursion format 22 27 39 7 75 10/4/0/1/0 5 314 79,077 —
1 JUMP_BRANCH_SLT 22 21 44 10 42 8/11/2/1/0 5 392 76,980 69,436
2 SHIFT_BITWISE 22 21 61 10 48 8/24/6/1/0 6 478 102,837 76,644
3 MUL_DIV 20 21 54 9 54 8/16/2/1/0 6 444 92,640 67,412
4 MEM_WORD 22 31 24 7 33 12/5/0/1/0 5 314 60,383 63,836
5 MEM_SUBWORD 22 31 55 10 53 12/22/1/1/0 6 472 98,846 76,196
6 ATOMICS 20 26 54 9 46 10/19/6/1/0 6 472 101,593 68,468
7 INIT_TEARDOWN 22 2 0 1 0 — 1 46 3,907 36,316
8 ZERO_WINDOWS 22 2 0 0 0 — 1 46 3,418 36,284
9 KECCAK_F 18 208 1,556 0 385 0/210/0/0/1,020 11 5,490 1,900,468 381,100
10 POSEIDON2 8 100 4,092 0 4,248 — 193 2,020 2,056,361 664,780
11 FR_ARITH 8 104 2,576 0 2,701 — 6 142 1,063,214 266,292
12 PUBLIC_INPUT 12 3 0 0 0 — 1 26 2,455 12,556
13 PUBLIC_OUTPUT 12 2 0 0 0 — 1 26 2,338 12,524
14 ADVICE_WINDOWS 22 3 0 0 0 — 1 46 3,535 36,316
15 MOD_MUL 16 104 221 0 125 0/274/0/0/0 10 2,244 550,391 135,220
16 SHA256_COMP 18 104 520 0 119 0/114/0/0/336 10 2,802 845,456 189,988
17 EC_ADD 16 392 1,028 0 637 0/1,110/0/0/0 12 8,772 2,350,670 434,916
18 FIELD_WINDOWS 20 2 0 0 0 — 1 42 2,758 —
19 FR_OP 20 31 31 0 44 0/36/0/0/0 7 370 89,741 —
20 P2_FIELD 18 45 382 0 372 0/58/0/0/0 7 392 294,425 —
21 FIELD_IO 18 43 39 0 24 0/70/0/0/0 8 650 164,713 —
22 FQ_OP 20 48 73 0 38 30/50/0/0/0 7 630 158,326 —

Heights. Both registries return None above MAX_TRACE_VARS = 30, and below the floor lookup.md §3 derives from the family's channels: 19 with TIMESTAMP, else 16 with RANGE16 or XOR8, else 0. A height changes trace_vars, each list's variable count and the number of halving lists, one per variable and as wide as the outputs, and no gate below them: at 2^20 ADD_SUB_LUI_AUIPC has 298 inner columns, 70,974 bytes and a 57,196-byte proof.

Shared circuits. The registries agree on families 1–17; the recursion format's ADD_SUB_LUI_AUIPC is add_sub::recursion_artifact (add-sub.md §2). PUBLIC_OUTPUT's circuit is ZERO_WINDOWS' and ADVICE_WINDOWS' is PUBLIC_INPUT's, byte for byte at one height, and FIELD_WINDOWS' is the zero window at a stride of one cell, all constraints::memory constructors (memory.md §3). Every other family's is its own module's artifact.

2 How a family circuit is assembled#

layer 0        M ‖ W ‖ S in layout order, beside the V tables' closed forms
gate list 0    memory leaves: the read side, then the write side, each padded to a power of
                 two with the literal 1
               per channel, in spec order: (−mult, T + g), then (1, E_l + g) per lookup,
                 then (0, 1) up to a power of two                      (lookup.md §6)
               every enforcing gate
lists 1 … r    row-wise: each tree combines sibling nodes, a product by a·b, a fraction by
                 (n_a·d_b + n_b·d_a, d_a·d_b); a tree already at one node is copied up
lists r+1 …    halving, one per variable: TreeProduct on a product, TreeCross (num) and
                 TreeProduct (den) on a fraction
top            no variables: read_root, write_root, then (num, den) per channel

r is the largest tree's depth, so the circuit has r + 1 row-wise lists; every registered circuit, POSEIDON2 included, ends in a top with no variables. crates/constraints/src/build.rs assembles it, writing the flat relation list and an all-zero padding row, zero_row_valid read off the gates' constants, and validating (gkr.md §4). constraints::memory::assemble gives it the product trees and lookup::channel_trees' fraction trees (lookup.md §11), then runs memory::check_memory (memory.md §8) and lookup::check_discharge: a constructor panics on a refusal, so every circuit that exists has passed them. Its callers:

  • memory::frame_with_channels_artifact(queries, trace_vars, FamilySpec), the execution families: memory.md §2's frame over memory::frame_queries(family), then the family's witness columns after the frame's w + 3, setup columns from S[0], virtual tables, enforcing gates after the frame's, lookups after its 2w gap obligations, and a non-empty channel list;
  • the window constructors (memory.md §3);
  • the delegation and recursion families, every gate in list 0, from constraints::delegation's shared columns, leaves and gates (delegation-circuits.md §1) — but POSEIDON2, which builds its own lists (delegation::Assembly): 192 row-wise lists of rounds beside its product trees, the last holding three gates on the output lanes.

Beyond the frame, each execution family has m_pc as the row's liveness and every other mask held to m_pc times the kinds making that query (<q>_mask_rule, memory.md §2); its decoded row as W columns, bound by decode_row to its table at the row's pc, and decoded_mask_bits, the mask as boolean kind bits, one-hot by the table's domain (lookup.md §10, program.md §6); a next_pc_rule (memory.md §5); a bound on each register value it writes (memory-ops.md §5); and channels ordered TIMESTAMP, RANGE16, GENERIC if read, DECODER.

prover::family_fill(family) is the prover's side: a prover::Fill writes a shard's committed columns but the multiplicities, which trace::build_multiplicities counts. prover::register pairs fill and circuit for each family of a VmConfig (ProverError::Unregistered if either is missing).

3 Checking a circuit independently#

crates/checker's validators enforce the rules again in code sharing nothing with crates/constraints/src/laws.rs, never calling validate. They evaluate a gate only through the kernel gkr_verify::eval_gate (gkr.md §3), so they re-read the rules, not the gates' meaning. Sampled checks use eight pseudo-random points from fixed seeds.

checks
check_laws (check_law1 … check_law4) the four laws, then the lookup rules (gkr.md §4); Law 4 and selector booleanity by evaluation, where validate compares expansions
check_padding, check_padding_identity the padding contract and its product-tree clause, fraction trees exempt
check_lookup_discharge lookup.md §11's discharge rule, gating and compression re-derived
violated_relations, violated_lookups a witness row's row-local relations and range obligations
channel_sums, check_channel_roots each channel's sum and denominator product, folded row by row rather than by a tree, naming every tuple no table row holds; then the circuit's root pairs against them
memory_roots the two roots as products over the rows the halving phase reads
memory_columns_from_log, frame_witness_from_log an execution family's frame columns from the memory event log, where trace builds them from a shard's rows

They do not re-implement check_memory, the copower rule (lookup.md §11), or validate's other construction rules, the degree ceiling among them.

checker::TamperHarness re-proves a statement with witness cells or boundary scalars changed, as an honest prover would prove the changed witness — each channel's multiplicities recounted unless one is what changed or the changed tuple is in no table, changed M columns recommitted in a fresh global commit phase, every shard re-proved — then verifies a shard or the block and asserts the refusal's class (a Lookup's channel too), or that a change breaking nothing verifies. It relies on the prover checking nothing (gkr.md §5), runs on the archived path (streaming.md §6), and carries the delegation anchor's forgeries (checker::assert_anchor_twins_refused, delegation.md §5).

A dump (checker::dump, CLI in tools.md §4) prints the columns by address and name, each list's gates in gkr.md §1's template with their relations, the flat relations over scratch[i], the scratch bijection, outputs, lookups and padding row. Relations are numbered list by list, producing before enforcing; a producing one is define_<column>, an enforcing one bears its gate's name; a node is named for its tree and layer (range16_3_1_num, read_root), a leaf for what it holds (write_pad_0, rd_hi_range_den). A literal below 2^32 prints in decimal, p − k for such a k as -k, any other as 0x and 64 big-endian hex digits; a challenge as its constants::challenge_slot::NAMES entry.

审计专区/证明系统

内存论证

规范原文docs/spec/memory.md以 Markdown 查看

摘要

证明“每次读取都返回最近一次写入”的论证,覆盖一个陈述的所有分片。它定义了内存元组;每个执行电路族的帧列、叶子与乘积,以及每个电路族都带有的 gadget;RAM 窗口电路族及验证者对它们施加的规则;寄存器与 pc 边界,以及唯一的核对等式;停机如何固定;哪些内容必须先于内存挑战;范围义务;构造期规则;以及该论证所依赖的前提,包括为什么 pc 的路径能免费给出程序顺序。

以下规范原文以英文维护;英文是本规范的标准语言。

Offline memory checking over a whole statement. Each shard's circuit outputs the product of its read tuples and of its write tuples; the verifier checks, once per statement, that all reads times the register and pc finals equal all writes times their initial values. RAM is initialized by window families over fixed address windows; registers and the pc have no rows. The section numbers are the ones the code cites.

1 The tuple#

T(AS, ADDR, TS, VAL) = γ_M + AS + α_addr·ADDR + α_ts·TS + α_val·VAL

The parts are in the order of constants::memory::{PART_AS, PART_ADDR, PART_TS, PART_VAL}. AS, an address-space tag (execution-trace.md §2), is unweighted; a RAM address is a 4-aligned word's byte address. γ_M, α_addr, α_ts, α_val are constants::challenge_slot slots 1–4, MEM_GAMMA to MEM_ALPHA_VAL, drawn once per statement after everything §6.1 lists; slot 5, MEM_WINDOW_CONSTANT, is derived per window shard by the verifier and never read from a proof (§3.3). A gate coefficient is one literal or one slot (gkr.md §3), so α_ts·4·cycle is the term (α_ts, cycle) four times.

Every memory artifact outputs its read product at outputs[READ_ROOT = 0] and its write product at outputs[WRITE_ROOT = 1] (constants::memory), before any channel's roots (lookup.md §6). All tuples of a statement form one multiset, over REG, RAM and PC, one anchor space per delegation type, where a request meets its invocation (delegation.md §5), and the recursion format's FIELD cells (recursion.md §2.1).

2 An execution family's memory subtree#

2.1 The frame columns#

A row of an execution family is one cycle; its accesses are queries, each a read and a write at one address, the write at 4·cycle + Δ (execution-trace.md §1). The query table is constraints::memory::{FRAME_NAMES, FRAME_SPACE, FRAME_DELTA}:

id query space Δ
0 pc PC 0 address 0; reads pc, writes next_pc
1 rs1 REG 1 read-only; an ecall's a7
2 rs2 REG 2 read-only; an ecall's a0
3 load RAM 2 read-only; a load's word
4 ram RAM 3 a store's or an atomic's word
5 rd REG 3 the x0 rule (§2.4)
6 deleg the row's 3 a delegation request's mirror (delegation.md §5)

A family's frame is exactly the queries its instructions make (execution-trace.md §4), in table order: constraints::memory::frame_queries, which crates/trace/tests/memory.rs holds to the union over all 59 instructions. A missing query would leave an instruction's written value unconstrained. No instruction routed to ADD_SUB_LUI_AUIPC touches RAM, and ATOMICS keeps every RAM access at Δ = 3, lr.w included. Window and delegation families have no frame (§3.3; delegation-circuits.md §1).

family queries, in slot order w leaves a side
ADD_SUB_LUI_AUIPC pc rs1 rs2 rd deleg 5 8
JUMP_BRANCH_SLT, SHIFT_BITWISE, MUL_DIV pc rs1 rs2 rd 4 4
MEM_WORD, MEM_SUBWORD pc rs1 rs2 load ram rd 6 8
ATOMICS pc rs1 rs2 ram rd 5 8

Columns are addressed by slot s, a query's position in its family's list:

M[0]                  cycle
M[1 + 5s + f]         slot s's <q>_mask, <q>_addr, <q>_read_ts, <q>_read_value, <q>_write_value
M[1 + 5w]             deleg_space, in the one frame holding deleg
W[s]                  <q>_gap_hi, for s < w
W[w], W[w+1], W[w+2]  rd_inv, rd_is_zero, rd_selected

That is 1 + 5w M columns, plus deleg_space, and w + 3 W columns, the family's own following (circuits.md §2). One deleg query serves every delegation type, so its space is the value of deleg_space, an M column the family pins to its type selectors: a leaf may read no W column (§8). The honest fill (trace::build_memory_columns, trace::build_frame_witness, over a shard's trace::RowSlice) sets a mask to 1 where the row is live and has the query, and every column of an absent query or a padding row to 0.

A frame holds a mask only to booleanity, so on the frame alone a padding row's rd query could rewrite x10, the exit status, after the exit row, and a live row could drop a query or carry one its instruction lacks. Every execution family makes m_pc the row's liveness and its decoder lookup's selector (lookup.md §10), and holds each other mask to m_q = m_pc·uses_q (its <q>_mask_rule gates), uses_q the sum of the row's kind and ecall-type selectors that make the query.

2.2 The leaves#

For the query at slot s with mask m, space AS and in-cycle slot Δ:

read_<q>    m·T(AS, addr, read_ts, read_value) + 1 − m
write_<q>   m·T(AS, addr, 4·cycle + Δ, write_value) + 1 − m

Each is one flat Quadratic of gate list 0, built from the unmasked tuple, a Linear whose AS and Δ terms sit on m (constraints::memory::read_tuple is the read one): constant 1; linear terms (γ_M, m), (−1, m), (AS, m) and, on the write side, (α_ts, m) Δ times; every other term multiplied by m, as is deleg's AS, the product (1, deleg_space, m). At m = 0 a leaf is 1 whatever its columns hold, at m = 1 the tuple, and it is one or the other only at a boolean m (§2.4).

2.3 The product#

Each side is padded to w rounded up to a power of two with read_pad_<i> and write_pad_<i>, the literal 1, reading no column. Row-wise Product lists reduce each side to one value a row, and trace_vars halving lists of TreeProduct multiply the rows (gkr.md §1), so the two roots are the products of the shard's read and write tuples. A padding row has every mask 0 and so every leaf 1, the padding contract's product-tree clause (gkr.md §4). The family's channel trees share the layers (circuits.md §2).

2.4 The gadgets every execution family carries#

Gate list 0's first enforcing gates, in this order, and the circuit's first 2w obligations:

<q>_mask_boolean       m − m·m = 0                       every query
<q>_writes_back        write_value − read_value = 0      rs1, rs2 and load, where held
rd_is_zero_inverse     addr·rd_inv + z − m = 0           on rd; z = rd_is_zero
rd_is_zero_at_nonzero  addr·z = 0
rd_is_zero_boolean     z − z·z = 0
rd_write_masked        write_value − sel + z·sel = 0     sel = rd_selected

gap_hi_<q>   TIMESTAMP, selector m:   hi                                   hi = <q>_gap_hi
gap_lo_<q>   TIMESTAMP, selector m:   4·cycle + δ_q − read_ts − 2^19·hi        δ_q = Δ − 1; δ_pc = −4
  • Booleanity. At m = −1 a pc query's leaves are each −T(REG, …): one sign flip a side, so the products balance and the pc access reads as a register access.
  • Write-back. Without it a read of x0 could write 5 there.
  • x0. The first two rd gates (constraints::gadgets::is_zero) make z = m·[addr = 0] and the last write_value = (1 − z)·sel: every write to x0 writes 0, whatever the family computed into sel, and with write-backs and x0's init 0 every read of it returns 0. The boundary's final x0 = 0 (§4.1) pins only its last write: a write of 5, a read of 5 and a write of 0 would otherwise balance.
  • Gap. Both chunks below 2^19 put gap = 4·cycle + δ_q − read_ts in [0, 2^38), so read_ts < 4·cycle + Δ as integers, every timestamp being a canonical integer by §4.2's count. The pc query's δ = −4 puts a row's pc write at least 4 after the one it reads, so consecutive rows' timestamps never interleave (§9). The frame's construction asserts two obligations per query.

3 RAM windows#

3.1 Geometry#

Window w at height h = 2^n covers the bytes [4h·w, 4h·(w + 1)), its row y being the word at 4h·w + 4y; the windows tile [0, 2^32) from 0. Ordinary RAM is [RAM_ORIGIN, ADVICE_ORIGIN) = [2^16, 2^31) (trace::in_ram), ending where window N = 2^29/h begins (verifier_core::advice_first_window). Window 0's rows y < 2^14 (constants::memory::RAM_LIVE_BIT) lie below RAM_ORIGIN at every height, and INIT_TEARDOWN masks them (§3.3).

3.2 The window families#

A window family's shard initializes and tears down one window; its rows are addresses.

region family id windows init value
[0, 0x8000) none: a hole
[0x8000, 0x10000) PUBLIC_INPUT, PUBLIC_OUTPUT 12, 13 2 and 3 at their pinned 2^12 the statement's input; 0 (public-values.md §4)
[RAM_ORIGIN, 4h) INIT_TEARDOWN 7 0, one shard S[0], the image column
[4h, 2^31) ZERO_WINDOWS 8 the listed w_1 < … < w_k in [1, N − 1] 0
[2^31, 2^32) ADVICE_WINDOWS 14 N … N + k_a − 1 M[2], bound to nothing (public-values.md §6)
FIELD cells FIELD_WINDOWS 18 0 … k_f − 1 0 (recursion.md §2.2)

INIT_TEARDOWN, ZERO_WINDOWS and ADVICE_WINDOWS share the window height h, 2^22 by default (§3.5). An unlisted RAM window is initialized by nothing. Every statement proves window 0 and, in practice, the stack's window N − 1, the initial sp being ADVICE_ORIGIN: two h-row shards however small the program, besides the public pair.

3.3 The artifacts#

INIT_TEARDOWN   image_window_artifact   M[0] teardown_ts, M[1] teardown_value, S[0] init_value
  read    live·(WC + α_addr·4·row + α_ts·M[0] + α_val·M[1]) + 1 − live     live = V[ram_live]
  write   live·(WC + α_addr·4·row + α_val·S[0]) + 1 − live
ZERO_WINDOWS, PUBLIC_OUTPUT    zero_window_artifact     M[0], M[1]
  read    WC + α_addr·4·row + α_ts·M[0] + α_val·M[1]        write   WC + α_addr·4·row
PUBLIC_INPUT, ADVICE_WINDOWS   value_window_artifact    M[0], M[1], M[2] init_value
  read    as above                                          write   WC + α_addr·4·row + α_val·M[2]
FIELD_WINDOWS   field_window_artifact: zero_window_artifact over α_addr·row, one cell a row

All are constraints::memory constructors: one leaf a side, then n halving lists; no witness column, enforcing gate or lookup; every init timestamp the literal 0; row is V[row]. V[ram_live] is [y ≥ 2^14], boolean on the cube by construction (gkr.md §2). The window enters only through WC, so one artifact serves every window:

WC = γ_M + RAM + α_addr·4h·w        gkr_verify::window_challenges
WC = γ_M + FIELD + α_addr·h·w       gkr_verify::field_window_challenges

w is verifier_core::shard_window's: 0 for INIT_TEARDOWN, the list's i-th id for ZERO_WINDOWS shard i, 2 and 3 for the public pair, N + i for advice shard i, i for field shard i. An init column is S where program identity binds it and M where it is one execution's, committed before the challenges (§6.1). A window shard has no inactive rows.

3.4 The columns a prover fills#

trace::init_windows(state, h) is ZERO_WINDOWS' list: the distinct ⌊a/4h⌋ over touched words a of ordinary RAM, ascending, without 0. A public or advice word is a RAM tuple too, and a zero window over it would be its second init row. trace::build_init_teardown_columns fills INIT_TEARDOWN and ZERO_WINDOWS, and trace::build_value_window_columns the value windows with their M[2], from the last-access tables (trace::MemoryState):

row y, a = 4h·w + 4y teardown_ts teardown_value
w = 0, y < 2^14 (masked) 0 0
a touched its last write's timestamp its last write's value
a untouched 0 its init value

An untouched row's two tuples are equal and cancel. The image column, program::image_init_column(image, h), has row y = ProgramImage::initial_word(4y): the word assembled byte by byte from file-backed bytes, 0 elsewhere, which is the trace's initial RAM value too. decode_program refuses an image with a file-backed byte at or above 4h (ProgramError::ImageOutsideWindow): it would sit in a zero window, read as 0, bound by nothing.

3.5 The verifier's window rules#

verifier_core::check_memory_windows, step 2 of derive_global_phase, before the global transcript (program::check_memory_windows wraps it):

rule why
INIT_TEARDOWN, ZERO_WINDOWS, ADVICE_WINDOWS at one height h a lower zero-window height would re-initialize image words; an advice height of its own is a grid advice_first_window(h) does not describe
PUBLIC_INPUT, PUBLIC_OUTPUT at PUBLIC_WINDOW_HEIGHT = 2^12 the height places their windows (public-values.md §2)
4h ≥ PUBLIC_OUTPUT_ORIGIN + PUBLIC_WINDOW_BYTES = 0x10000: h ≥ 2^16 on the menu the public windows lie in window 0's masked rows, out of every zero window's reach
one shard each of INIT_TEARDOWN, PUBLIC_INPUT, PUBLIC_OUTPUT (public-values.md §4 for the pair)
one id per ZERO_WINDOWS shard, strictly increasing, in [1, N − 1] disjoint windows; id 0 is unmasked over [0, RAM_ORIGIN); N up is advice
N + k_a ≤ 2^30/h, k_a the advice shard count advice ends by 2^32; it needs no list, starting where the zero ids stop
k_f·h ≤ 2^32 field cells recursion.md §2.2

The first three are verifier_core::window_height, which VmConfig::from_bytes runs too: a config breaking them does not decode.

4 The register and pc boundary#

Registers and the pc have no rows: the verifier multiplies in their initial and final tuples, once per statement. Rows for them would repeat the init tuples in every shard holding them, and a stale read would balance against the copy.

4.1 The boundary scalars#

The statement carries 64 scalars, gkr_verify::BoundaryFinals, absorbed as one MEMORY_BOUNDARY message in this order (verifier_core::boundary_scalars):

positions
0–31 t_0 … t_31 x_r's final timestamp: its last query's write, 0 if never queried
32 t_pc the pc's: the exit row's pc write
33–63 v_1 … v_31 x_r's final value, 0 if never queried

The final values of x0, 0, and of the pc, HALT_PC, are constants, not carried. PublicInputs::from_bytes refuses t ≥ 2^38 or v ≥ 2^32, and verify_global_memory re-checks the timestamps and holds v_10 to the exit status; no other register carries a public value. t_pc is not a cycle count: the pc's timestamps increase but need not be consecutive. trace::build_boundary_finals(state) is the fill.

4.2 The factors and the reconciliation#

W_b = ∏_{r=0}^{31} T(REG, r, 0, 0) · T(PC, 0, 0, entry_pc)
R_b = T(REG, 0, t_0, 0) · ∏_{r=1}^{31} T(REG, r, t_r, v_r) · T(PC, 0, t_pc, HALT_PC)

∏ read roots · R_b  =  ∏ write roots · W_b  ≠  0       over every shard of the statement

entry_pc is the verifying key's (§6.2). gkr_verify::boundary_factors evaluates each tuple through gkr_verify::eval_gate on the circuits' own tuple gate, read_tuple of pc or rs1, so the boundary and the circuits cannot disagree on the parts; gkr_verify::reconciles is the equation, which verifier_core::verify_global_memory runs once per statement (proof.md §6). A shard's roots are its GKR outputs, held to the statement's entry by its own verification.

The count. Read each query as an edge from its read tuple to its write tuple. Inits are only written and finals only read, so a balanced multiset is paths from inits to finals plus loops. An edge advances the timestamp by an integer in [1, 2^38 + 3] (the gap plus the query's least advance, 4 at the pc and 1 elsewhere), so a loop needs more than p/(2^38 + 3) > 2^215 edges, and a statement has fewer than 2^67 tuples: under 2^32 shards a family (a u32 count), 23 families, at most 2^22 rows (the menu's top), at most 196 tuples a row (EC_ADD's 97 frame words and its anchor, both sides), and 66 boundary tuples. So nothing loops: every path starts at an init at timestamp 0 and ends at a final, and every timestamp on it is an integer below 2^105.

5 Halting#

constants::memory::HALT_PC = 1. The exit row, ecall with a7 = 93, writes next_pc = HALT_PC instead of its fall-through (execution-trace.md §6), and R_b fixes the pc's final value to it. Nothing else writes it: HALT_PC is odd, every other next_pc even, and "odd" is a constraint only where a family makes it one.

  • A family copying the decoded fall-through, which is even, holds next_pc − decoded_next_pc = 0 and needs no bound.
  • JUMP_BRANCH_SLT, the one family computing a pc, range-checks every next_pc it writes even; otherwise a jalr whose rs1 + imm is 1 could write HALT_PC (jump-branch-slt.md).
  • ADD_SUB_LUI_AUIPC writes HALT_PC on its exit row alone (add-sub.md).

HALT_PC is below RAM_ORIGIN, so no decoded-table row claims it and no live row reads it (lookup.md §10). The pc's path therefore ends with the exit row's write, consumed by the final read. With a free final pc every prefix of an execution would balance.

6 Binding#

6.1 What precedes the memory challenges#

The four challenges are squeezed once per statement, at the end of the global transcript (proof.md §2 is the schedule), after everything a tuple or the reconciliation reads, because what is chosen after them can be solved for:

  • every shard's M commitments, every column a leaf may read but S and V (§8);
  • program identity, fixing entry_pc and the image column (§6.2), and the SRS digest, fixing the generic table's S columns (proof.md §3);
  • the shard counts and MEMORY_WINDOWS, the zero-window ids, fixing every window shard's addresses through WC: a list chosen afterwards is a union over up to 2^(N − 1) lists, 2^127 at h = 2^22 and no bound at all at 2^20;
  • io_digest, fixing the public windows' contents (public-values.md §5);
  • last, the 64 boundary scalars: a final value chosen afterwards reconciles any trace, v_r = (target − γ_M − REG − α_addr·r − α_ts·t_r)/α_val.

The roots are not absorbed: each shard's GKR proof binds its own.

6.2 The image column and the entry pc#

Program identity (program.md §8 is the recipe) binds INIT_TEARDOWN's one setup commitment, the image column's, and entry_pc, under PROGRAM_ENTRY. Recomputing identity binds a commitment, not the column a proof reads; the INIT_TEARDOWN shard's batched opening closes that by taking S[0]'s commitment from the verifying key, the list identity is recomputed over (proof.md §5, §7). Without it a statement over another image, with a trace consistent with that image, would verify. Without entry_pc in identity, a key carrying the registered identity beside another entry pc would verify an execution starting elsewhere. Identity binds nothing an execution chooses: no shard count, window list, public or advice word.

7 Range obligations#

A range obligation holds where its selector is 0 or its one expression is below its channel's bound (lookup.md §1, §3). Every circuit bounds a value one way:

  • a 32-bit value v: a witnessed high halfword h and RANGE16 obligations on h and on v − 2^16·h, under the row's selector, and no gate;
  • a result r = e mod 2^32 of an exact 0 ≤ e < 2^33: a witnessed wrap, the gates wrap − wrap·wrap = 0 and e − r − 2^32·wrap = 0, and r bounded as above; a wider carry is a family's own construction;
  • a timestamp gap: two 19-bit TIMESTAMP chunks, no wrap (§2.4). Delegation and recursion families decompose theirs their own way (delegation-circuits.md §1).

8 Construction-time rules#

constraints::memory::check_memory refuses, naming the gate, a memory artifact with:

  1. provenance: a gate or output whose cone both names a memory slot (1–5) and reads a W column, computed forward with two flags a column, so a tuple times a copy of a W column two layers up is refused too;
  2. a root over W: outputs[READ_ROOT] or outputs[WRITE_ROOT] whose cone reads a W column at all, slot or not, which rule 1 does not see;
  3. a memory slot over anything but M, S and V: a gate carrying one reads no W, inner or cached column;
  4. an unconstrained mask: a leaf — a producing Quadratic of gate list 0 with constant 1 and a slot-weighted linear term — whose mask, that term's operand, is an M, W or S column with no m − m·m enforcing gate in gate list 0, or a virtual column but V[ram_live].

It runs beside CircuitArtifact::validate, whose laws it assumes, wherever a memory artifact is built (constraints::memory's assembly panics on a refusal) and in VerifyingKey::check. A W column is committed in a shard's own transcript, after the memory challenges, so a tuple or root over one is chosen after them and balances any trace. M columns precede the challenges and V columns are closed forms; S columns are admitted because they precede them too, bound by identity or, for the generic table, by the SRS digest (§6.1).

9 What the argument rests on#

Both sides of §4.2 are products of linear forms in (γ_M, α_addr, α_ts, α_val), one per distinct tuple, every tuple fixed before those are drawn (§6.1, §8). By Schwartz–Zippel they agree on unequal multisets with probability at most N/p, N < 2^67 (§4.2), and on equal ones §4.2's count gives:

  • One init per address of REG, PC, RAM and FIELD: the 33 boundary inits once per statement, and §3.5's windows, disjoint and of one height. A second init would let a stale read balance. An anchor space has no init: each invocation's answer, stamped 0, starts a path one request long (delegation.md §5).
  • Coverage. Every query lies on a path from an init, so nothing reaches an address no family initializes: the hole [0, 0x8000), where a null dereference does not balance, an unlisted window, a register above x31, a pc address but 0. A query reading its own write would balance with no init; the gap forbids it.
  • Consistency per address: on its one path every read returns the write before it, and the final tuple holds the last.
  • Initial values: the image's, by §3.4's refusal and §6.2's opening of S[0]; 0 in every zero window and the journal; the statement's input in its window (public-values.md §5). Advice is bound to nothing by design (public-values.md §6).
  • Order across rows, shards and families. The pc's path runs from T(PC, 0, 0, entry_pc) through every live row of every execution family, each m_pc = 1 row one edge, to the exit row (§5). That is pc continuity; it orders the rows by their pc writes 4·cycle, which are therefore distinct, so no cycle is proved twice. Nothing else carries it: there is no per-shard pc chaining, and a shard's time window ties to no row (proof.md §8).

Per address, the order is timestamp order, and it is program order: the pc query's gap puts consecutive pc writes at least 4 apart (§2.4), so each cycle's four timestamps precede the next cycle's whatever value cycle takes, and a row never reads an address before its predecessor's write there. An invocation rides its requesting row's cycle (delegation.md §5) and is ordered with it.

审计专区/证明系统

查找(lookup)

规范原文docs/spec/lookup.md以 Markdown 查看

摘要

查找义务如何被证明:每个分片的每张表一个 LogUp 恒等式,在电路自身的 GKR 过程中由分式树求和,并在树根处检查。它定义了通道断言的内容、通道的挑战、五张表、被关闭的行如何查找中性元组、分母、分式树、重数、根检查、打包的通用表、把每一行绑定到程序的解码器通道、构造规则,以及该论证所依赖的前提,包括为什么每个键都必须由其电路族限定范围。

以下规范原文以英文维护;英文是本规范的标准语言。

How a circuit's lookup obligations are proved: per shard, by one LogUp channel per table, each summed by a fraction tree inside the circuit's own GKR pass and checked at its root.

1 What a channel claims#

A lookup is LookupExpr { name, channel, selector, tuple }: a channel of constants::lookup_channel, a committed M, W or S column as selector, and a tuple of Linear expressions with literal coefficients over committed columns and the circuit's virtual tables. It holds on a row where the selector is 0, or

  • on a range channel, where its one expression's canonical integer is below 2^BITS[channel];
  • on a table channel, where its tuple is a row of the channel's one table: 1 to MAX_TUPLE = 7 expressions, the same number for every lookup of the channel.

memory.md §7 is the convention range obligations follow. A channel discharges all of a shard's lookups on it as one identity over the shard's rows y:

Σ_y Σ_l 1/(E_l(y) + g)  −  Σ_y mult(y)/(T(y) + g)  =  0

E_l(y) is lookup l's gated tuple (§4) and T(y) the table's row y, both compressed by β (§5); mult is the channel's multiplicity column (§7). A range table is the one column [0, 2^BITS).

2 The challenges#

slot challenge_slot value
6 LOOKUP_G g, drawn
7 LOOKUP_BETA β, drawn
8–12 LOOKUP_BETA_2 … LOOKUP_BETA_6 β^2 … β^6, derived
13 LOOKUP_DECODER_NEUTRAL g − Σ_{j<W} β^j, derived; W the decoder tuple's width

g and β are shard-local: the shard's transcript draws them, in that order under the tag LOOKUP_CHALLENGE (33), right after absorbing its witness commitments, multiplicities included (proof.md §4). M columns are committed in the global transcript the shard is seeded from and S columns are bound by identity or the SRS digest, so every column a channel reads is fixed before either challenge exists.

β^0 is the literal 1, so a one-column tuple names no slot. A gate coefficient is one literal or one slot (gkr.md §3), so each higher power is a slot of its own, computed by the verifier and never read from a proof (gkr_verify::insert_lookup_challenges, which reads W off the artifact's decoder lookup).

Selectors are boolean: CircuitArtifact::validate refuses a lookup whose selector no enforcing gate of gate list 0 holds to s − s·s = 0. The selector multiplies the tuple inside the denominator (§5), so a channel proves the gated tuple s·(e + o) + n is a table row, which is the obligation only at s ∈ {0, 1}. At any other s a scaled tuple is looked up instead: on a range channel, s = t·e⁻¹ lands any nonzero e on any table value t.

3 Tables#

channel id kind table, at row y width table_vars
TIMESTAMP 0 range V[range19]: y mod 2^19 1 19
RANGE16 1 range V[range16]: y mod 2^16 1 16
GENERIC 2 table, committed the packed table (§9) 3 0
DECODER 3 table, committed the family's decoded table (§10) 7 or 6 0
XOR8 4 table, virtual V[xor8_a], V[xor8_b], V[xor8_out]: y's low two bytes and their XOR 3 16

A virtual table is a closed form of the row index, never committed: the verifier evaluates its multilinear extension where the GKR pass ends (gkr_verify::virtual_at_point, gkr.md §2). Each is a weighted sum of the row's bits but V[xor8_out], Σ_{j<8} 2^j·(y_j + y_{j+8} − 2·y_j·y_{j+8}), which is exact because y ^ z = y + z − 2yz is multilinear. So XOR8 costs no commitment and nothing in the SRS digest.

constraints::lookup::table_vars is the fewest variables at which a table is complete: BITS for a range channel, 16 for XOR8, 0 for a committed table, a setup column at the circuit's own height. Below it a virtual table holds only part of its range, which costs completeness, not soundness. family_circuit returns None below the largest table_vars of a family's channels, so a key naming such a height fails to load (proof.md §7). On the height menu (program.md §7) a family carrying TIMESTAMP is at 2^20 or more, and one carrying RANGE16 or XOR8 at 2^16 or more. The packed table needs 2^18 rows (§9), and every family that reads it carries TIMESTAMP. Above table_vars a table repeats, which §7 makes harmless.

XOR8's tuple is three wide so that membership bounds each entry to [0, 256) on its own; a packed key x + 256·y would bound neither, (x, y) and (x + 256, y − 1) compressing alike. Every other bitwise operation on bytes is a linear form over its results (delegation-circuits.md §1).

4 Gated keys#

A lookup expression is evaluated on every row, so the selector sends a row whose key means nothing to a neutral tuple, which is a real table row:

gating channels gated position j neutral tuple
NoOffset TIMESTAMP, RANGE16, XOR8 s·e_j all zero
ZeroEntry GENERIC s·(e_0 + 1), then s·e_j the all-zero ZeroEntry row
MinusOne DECODER s·(e_j + 1) − 1 −1 in every column, a padding row

The + 1 keeps every real key of the packed table at 1 or above, so no real entry is the all-zero tuple a switched-off row looks up. A range table needs no offset, 0 being in range, and one would push 2^BITS − 1 out of it; XOR8's (0, 0, 0) is a true entry. A decoded table has no all-zero row, pc 0 being a valid pc, and its MINUS_ONE padding rows (program.md §5) are the neutral entry.

Each key a table channel looks up is bounded by the family that looks it up, because a channel proves membership and nothing more. A selected row whose key expression is −1 gates to the ZeroEntry, and an unbounded key reaches any sub-table of the packed table: an AND key a + AND_BASE with a unbounded lands on a U16GetSign row and proves a false AND. Families bound their keys with RANGE16 obligations or build them from bounded columns (shift-bitwise.md §3 and the other family pages); the decoder's key is §10's.

5 The denominator#

With s the selector, e_j = Σ_i c_{j,i}·x_{j,i} + k_j, and §4's offset o_j (1 or 0) and neutral value n_j (−1 or 0):

E + g  =  Σ_j β^j·(s·(e_j + o_j) + n_j)  +  g
       =  Σ_j β^j·s·e_j  +  Σ_j β^j·o_j·s  +  (g + Σ_j β^j·n_j)
T + g  =  Σ_j β^j·t_j  +  g

E + g is one Quadratic (constraints::lookup::row_denominator): each term of e_j the product (β^j·c)·s·x, each offset the linear term β^j·o_j·s, and the bracket the slot LOOKUP_G or, for the decoder, LOOKUP_DECODER_NEUTRAL. β^j·c is one coefficient only where β^0 = 1 makes it a literal or c = 1 makes it the slot, so position 0 takes any literal coefficients and constant and every later position weights its columns by 1 with no constant. T + g is one Linear over the table's columns (table_denominator).

6 The fraction tree#

A channel's leaf level is (num, den) pairs of gate-list-0 columns, P = (L + 1).next_power_of_two() of them for L lookups:

leaf num den
the table, first −mult T + g
each lookup, in artifact order 1 E_l + g
padding, up to P 0 1

Row-wise gate lists add sibling pairs, (n_a·d_b + n_b·d_a, d_a·d_b), until each row holds one pair; a tree shallower than the circuit's deepest copies itself up. Then trace_vars halving lists add the rows' pairs, TreeCross writing the numerator and TreeProduct the denominator (gkr.md §3). The circuit's outputs are the memory argument's read and write roots, then each channel's (num, den) in the order of its channel specs (crates/constraints/src/build.rs).

A channel costs 4P − 2 inner columns to reduce a row, 2 more per copy-up layer and 2 per halving list, and one committed column. P doubles each time L reaches a power of two.

The padding clause (gkr.md §4) asks a padding row to feed 1 into every product tree. A fraction tree is exempt: its identity is (0, 1), and a padding row is not idle in a channel but looks up the neutral tuple, which the multiplicity counts. checker::check_padding_identity exempts every column a TreeCross reads.

7 Multiplicities#

Each channel has one multiplicity column, a committed W column; a circuit's are its last W columns, in channel order. Row t counts the (row, lookup) pairs of the shard whose gated tuple is table row t's, switched-off rows included. A tuple at several table rows is credited to the lowest; every other copy holds 0 and contributes 0/(T + g). The count is over raw gated tuples, the column being committed before g and β exist (trace::build_multiplicities, which refuses a tuple no table row holds: the honest prover cannot balance it).

No gate or range check constrains the column, and soundness needs none. If a gated tuple v is in no table row, the left side of §1's identity, as a rational function of g, has a pole at −v whose residue is the number of lookups producing v: a positive integer below p, whatever the column holds.

8 The root check#

accept  iff  num = 0  and  den ≠ 0

on each channel's root pair, at step 9 of proof.md §6 (gkr_verify::channel_holds); a failure is VerifyError::Lookup { channel }. The GKR pass absorbs the pair before its first challenge and proves it (gkr.md §5). den is the product of every leaf denominator, and num = 0 means the sum vanishes only where den ≠ 0: one leaf (0, 0) — a table row whose T + g vanishes, counted 0 — makes the root (0, 0) whatever the other leaves hold. With g drawn after the columns that has probability at most fractions/|Fr|, and den ≠ 0 makes it a refusal.

9 The generic table#

One committed table of constants::generic_table::WIDTH = 3 columns, a key and two values, packing three sub-tables under disjoint key ranges (program::lookup_tables::generic_table):

row 0                      ZeroEntry    (0, 0, 0)
rows 1 ..= 2^16            AND          (AND_BASE + a + 1,    b,        a & b)        a, b < 2^8
rows 2^16+1 ..= 2^17       U16GetSign   (SIGN_BASE + h + 1,   h >> 15,  0)            h < 2^16
rows 2^17+1 ..= 2^17+32    ShiftPowers  (SHIFT_BASE + s + 1,  2^s,      2^(31 − s))   s < 32
rows above                 zero

AND_BASE = 0, SIGN_BASE = 256 and SHIFT_BASE = 65,792 put the keys at 1..=256, 257..=65,792 and 65,793..=65,824; a lookup's key expression is x + BASE, and the gating adds the 1. U16GetSign serves every sign an execution family computes, AND the bitwise operations of SHIFT_BITWISE and ATOMICS, ShiftPowers the shifts. The copower 2^(32 − s) is stored halved (SHIFT_COPOWER_BITS = 31), 2^32 not fitting a u32 column, and the two gates that read it carry the factor 2 (shift-bitwise.md §4). 131,105 rows in all (GENERIC_ROWS).

Its commitments are a constant of the ceremony. A Mercury commitment reads the evaluation table as coefficients (mercury.md §2) and the table is zero past its entries, so over 2^n rows it commits to the same three points for every n ≥ 18; generic_commitments(srs) computes them at 2^18 (GENERIC_LOG_HEIGHT). Every verifying key carries them once, as VerifyingKey::generic_table, whether or not a family reads the channel, and its SRS digest covers them (proof.md §3); program identity does not. A circuit that reads GENERIC names the table as its three setup columns after identity's (FamilyCircuit::reads_generic_table), and a shard's opening checks them against the key's points (proof.md §5).

10 The decoder channel#

The DECODER table is the family's decoded table, program::lookup_tuple(family)'s columns, as its first setup columns at its height, row i holding pc 2i (program.md §5); program identity commits them (program.md §8). Each execution family makes one lookup on it, imm absent for MUL_DIV and ATOMICS:

decode_row    selector m_pc    tuple (pc read value, next_pc, rs1, rs2, rd, [imm], extra_mask)

The key is the frame's own pc read (memory.md §2), so the cycle itself is bound to the program; the rest are the row's decoded columns, which the family's other gates read. The selector is the row's liveness, so a padding row looks up the MINUS_ONE tuple, which every decoded table holds, being taller than its last instruction.

The family's decoded_mask_bits gate ties the packed mask to boolean kind bits. That the bits are one-hot, and that a live row is an instruction at all, is the table's domain: its live rows hold one-hot masks and its padding rows −1, which no sum of kind bits reaches. Boolean columns looked up one by one would lose this: booleanity admits any subset of bits, the empty one included, and an all-zero mask makes every gate a kind selects vacuous.

11 Construction rules#

CircuitArtifact::validate enforces §1's form and widths, §2's selector rule and §5's coefficients wherever an artifact is built or loaded (gkr.md §4). When a circuit is assembled, constraints::lookup asserts that a channel has a lookup, that its multiplicity is a W column, that every lookup has its table's width, and that a range channel's table is the one its bound names (range_table) with BITS ≤ trace_vars; constraints::memory::frame_with_channels_artifact refuses an empty channel list, which would leave a frame's gap obligations discharged by nothing.

The discharge rule, constraints::lookup::check_discharge, at assembly and at every key load (VerifyingKey::check): every lookup is the denominator of exactly one gate-list-0 column, its numerator 1 directly before it; no column is two lookups'; each channel's (−mult, T + g) appears once. It matches by normalized expansion inside the cone below the channel's own root pair, so an obligation or table fraction in another channel's tree is refused, and the two range channels, which gate alike, are not confused. Which output pair is whose root, which columns are a table and which counts it is not in the artifact but in its ChannelSpecs, which a key carries in FamilyCircuit::channels and must hold as the registry's (circuits.md §1). checker enforces this rule and the lookup rules a second time, with code of its own (circuits.md §3).

The copower rule, constraints::lookup::check_copowers, run by every constructor that bounds a column through a copower. A bound x < p written as x·p′ < 2^32, p·p′ = 2^32, bounds nothing alone: p′ is a unit of Fr, so x = s·p′⁻¹ ranges over a coset of 2^32 values. Each such x therefore also carries a direct RANGE16 bound, as a halfword or as a high chunk and a remainder, under the same selector.

12 What it rests on#

  • Every gated tuple is a table row: §1's identity over challenges drawn after every column it reads, boolean selectors (§2), both root conditions (§8) and the GKR pass. The error is at most fractions/|Fr| for g, plus looked-up tuples × table rows × (width − 1)/|Fr| for a β collision: below 2^−190 at every menu height.
  • A lookup answers from its own sub-table: one width per channel (§11), disjoint key ranges and the + 1 (§9), and its family's bound on the key (§4).
  • A switched-off row costs nothing: its neutral tuple is a table row the multiplicity counts (§4).
  • The table is the intended one: the verifier's own closed form (§3), or a table bound by identity or by the SRS digest, as trustworthy as the channel the verifier took that from (program.md §8, srs.md §3).
  • Every declared obligation is discharged: the discharge rule over the registry's specs (§11).

The channel does not check the multiplicity column (§7), a key's bound (§4), or that a committed table holds its neutral row, a property of its values that no artifact states: a table without one stops the honest prover at trace::build_multiplicities.

审计专区/证明系统

证明

规范原文docs/spec/proof.md以 Markdown 查看

摘要

各项论证如何组合成一个通过验证的陈述。它定义了 PublicInputs 与陈述的顺序、块(block)及其分片集合精确性、全局 transcript G1–G11、SRS 摘要、分片 transcript S1–S6、批量打开、验证者按顺序进行的检查以及每种失败的类别、验证密钥及其加载规则、时间窗口,以及每个对象(包括证明归档)的字节级序列化格式。

以下规范原文以英文维护;英文是本规范的标准语言。

What a verifier checks and the formats it reads, from the statement to the bytes. The memory argument, LogUp, GKR and Mercury are their own pages; this one is how they compose. crates/verifier-core (#![no_std]) implements everything here but step 12, the opening, which crates/verifier runs.

1 The statement#

A statement is a PublicInputs under a verifying key (§7): one execution of the key's program. It is proved by one ShardProof per statement shard, a (family, index) with index below the family's shard count, each verified against the same PublicInputs. A shard's proof establishes its own circuit and opening, and the reconciliation it joins reads the roots the statement claims for every other shard, which only their own proofs establish: a statement is verified when its proofs are exactly its shards and all pass, never by a subset.

Every entry point is (&VerifyingKey, proof, &PublicInputs): verifier::verify_shard, verifier::verify_block, verifier_core::reduce_shard. A verifier holds two values from a channel the prover does not control, the program identity and the SRS digest (§3), and compares them with the key's; the key itself may come from anyone (§7). The verifier CLI compares identity only (tools.md §6); host::verify(vk, block), which is verify_block(vk, block, block.statement()), compares neither and leaves its caller to check the statement's input, output and exit status too.

1.1 PublicInputs#

field
input: Vec<u8> the public input window's payload, at most PUBLIC_PAYLOAD_BYTES = 16,380 (public-values.md §3)
output: Vec<u8> the journal, the public output window's payload, as long
exit_status: u32 x10's final value
shard_counts: Vec<u32> one per family of the VmConfig, in its order, possibly 0
windows: Vec<u32> ZERO_WINDOWS' window ids, one per shard (memory.md §3.5)
boundary: BoundaryFinals the 64 register and pc boundary scalars (memory.md §4.1)
memory_commitments: Vec<Vec<[u8; 64]>> per statement shard, its M columns' commitments in layout order
memory_roots: Vec<[Fr; 2]> per statement shard, [read_root, write_root]

The first three are the claim; the rest is the execution's record, which the prover chooses. All of it but the roots and the exit status is absorbed before any challenge (§2).

1.2 Statement order#

verifier_core::statement_shards(config, counts):

(INIT_TEARDOWN, 0)
(ZERO_WINDOWS, 0) … (ZERO_WINDOWS, k − 1)
every other family of the VmConfig, ascending by id, shards 0 … count − 1 each

It orders memory_commitments, memory_roots, G8's groups and a block's proofs. The two leading families are ids 7 and 8, so the order is not ascending by id. A family with count 0 has no entry.

1.3 The block#

verifier_core::BlockProof { config, statement, shards } is one execution closed: the static VmConfig, the statement and one proof per statement shard, in statement order; it adds no evidence to the proofs'. The config and counts are public data of the proof, absorbed at G3 and G4, so the block carries both and check B1 holds them to the key's and the verifier's.

Shard-set exactness, BlockProof::shape, at decode and again in verify_block: one count per config family; the counts' total, summed in u64 before any list is built from them, equal to the numbers of proofs, commitment lists and root pairs; the proofs naming statement_shards in order. No (family, index) is missing, repeated or extra.

BlockProof::reconciliation is the cross-shard record set, a BlockReconciliation of one ShardRecord { family, shard_index, ts_window, memory_commitments, roots } per statement shard, assembled from the shard's proof (the window) and the statement (the rest).

2 The global transcript#

verifier_core::global_commit(vk, statement), run by the verifier in derive_global_phase and by the prover once every shard's M columns are committed (streaming.md §2): a fresh transcript, tag values in transcript.md §5.

# op tag message
G1 absorb PROTOCOL_SUITE [PROTOCOL_VERSION], 0
G2 absorb SRS_DIGEST [vk.srs_digest] (§3)
G3 absorb VM_CONFIG the config (program.md §7)
G4 absorb SHARD_COUNTS shard_counts
G5 absorb MEMORY_WINDOWS windows
G6 absorb PROGRAM_IDENTITY [vk.identity]
G7 absorb PUBLIC_INPUTS, bytes the 32 bytes of io_digest(input, output) (public-values.md §5)
G8 per family MEMORY_GROUP, COMMITMENT below
G9 absorb MEMORY_BOUNDARY the 64 boundary scalars
G10 squeeze ×4 MEMORY_CHALLENGE γ_M, α_addr, α_ts, α_val, challenge slots 1–4
G11 squeeze GLOBAL_STATE_DIGEST the global state digest

G3–G5 are verifier_core::absorb_statement_descriptor. G8 is one group per family of the config, in statement order, a family with count 0 included:

MEMORY_GROUP   [family, shard count]
COMMITMENT     per shard, ascending: its memory_commitments, one message of 4k limbs

A statement's or proof's points are absorbed as limbs (transcript.md §4) and decoded only at step 12.

Everything a memory tuple or the reconciliation reads precedes G10 (memory.md §6.1 says why for each). Two fields are not absorbed: memory_roots, which depend on the challenges and are bound by each shard's own GKR proof (step 10a), and exit_status, which step 10b holds to v_10, absorbed at G9. The digest seeds every shard (§4); a proof carries the digest it was seeded with (ShardProof::global_digest) and step 5 compares it with the replay, so a shard proof is for one statement under one key.

3 The SRS digest#

t ← Transcript::new()
t.append_bytes(SRS_VERIFIER, srs_verifier)           320 bytes, srs.md §5
append_g1_points(t, GENERIC_TABLE, generic_table)    the table's 3 points, one 12-limb message
srs_digest ← t.sample()                              one raw squeeze

verifier_core::srs_digest; GENERIC_TABLE is absorbed in this sponge and nowhere else, the points key column first. G2 absorbs the digest, so a proof is bound to the three points its pairings read and the table its GENERIC lookups read. Both are constants of the ceremony (lookup.md §9), so one digest serves every key. It does not cover the powers, which only a prover reads: an opening is checked against g2_tau whatever powers made the commitment.

A key's load recomputes the digest from the key's own points (§7), which shows they agree, not that they are the ceremony's, and identity binds neither (program.md §8). So the verifier compares vk.srs_digest with the ceremony's (srs.md §3). Without that comparison, whoever built the key chose τ, so can open anything, and chose the table every GENERIC lookup is held to.

4 The shard transcript#

Shard (family, index) runs a fresh sponge (verifier_core::shard_transcript), not a restored global one:

# op tag message
S1 absorb SHARD_SEED [global state digest, family, index]
S2 absorb SHARD_TS_WINDOW [ts_start, ts_end] (§8)
S3 absorb COMMITMENT the shard's W commitments, multiplicities included, one message
S4 squeeze ×2 LOOKUP_CHALLENGE g, then β (lookup.md §2)
S5 the GKR backward pass (gkr.md §5.2)
S6 the batch opening (§5): B1–B3 of mercury.md §5, then the sixteen steps of its §3

S4 is drawn for every shard, whether or not its circuit has a channel. Every challenge follows every commitment the circuit reads: M at G8, S through identity at G6 or the SRS digest at G2, W at S3, as GKR requires of its caller (gkr.md §5.1).

The circuit's external challenges (verifier_core::shard_challenges) are slots 1–4 from G10; for a window family, slot 5 at the window verifier_core::shard_window gives the shard (memory.md §3.3); then the lookup slots from g and β. Its outputs, the top layer, are the two memory roots and then each channel's (num, den) (lookup.md §6), 2 + 2c of them for c channels.

In the recursion format a shard commits M and W as stacks of 2^σ columns, and S6 opens with σ STACK_CHALLENGE squeezes extending the opening point (recursion.md §1.3). At σ = 0, the base format, there are none.

5 The opening#

After S5 every committed column has one claim, layer 0's, all at one point u (gkr.md §5.2). So there is nothing for a claim-merging sumcheck to merge, and S6 opens every column as one batch (mercury.md §5), one 704-byte Mercury proof a shard:

columns      the circuit's committed layout: M[0..], W[0..], S[0..]
commitments  M  PublicInputs.memory_commitments[the shard's position]
             W  ShardProof.witness_commitments
             S  VerifyingKey.setup_commitments[family], then VerifyingKey.generic_table
                when the circuit reads GENERIC (FamilyCircuit::reads_generic_table)
point        u, variable j at index j
values       layer 0's claims, ShardProof.gkr.layers[0].final_evals

Column i carries ρ^i, so this order is part of what is proved. Virtual columns are neither claimed nor opened: the verifier evaluates their closed forms. Taking S from the key is what makes the opening bind the columns identity commits, the decoded tables and the image column (memory.md §6.2), and the generic table the SRS digest covers.

reduce_shard ends at an OpeningClaim: these commitments, the point, the values and the live shard transcript. verify_shard spends it with pcs::batch_verify (step 12); a recursion node defers it (recursion.md §8.3).

6 Verification#

verifier::verify_shard(vk, proof, public) returns the first failure, in this order, as a VerifyError:

step class check
1 Statement one shard count per config family; the key's circuits are its config's families, in order, with one setup list each
2 Statement check_memory_windows (memory.md §3.5); input and output each at most PUBLIC_PAYLOAD_BYTES
3 Statement one root pair and one commitment list per statement shard, each list its family's M width; the total summed in u64 first
G1–G11 (§2)
4 Statement ts_start ≤ ts_end ≤ 2^38
5 Statement the replayed global state digest is proof.global_digest
6 Malformed (family, index) is a statement shard; the witness commitments and outputs have the circuit's counts
7 Constraint { layer } gkr_verify::verify over the shard transcript: LayerInconsistency { layer }; its ProofShape, OutputShape and MissingChallenge are Malformed
8 Constraint { layer: 0 } every base claim at one point
9 Lookup { channel } gkr_verify::channel_holds on each channel's root pair, in channel order (lookup.md §8)
10a MemoryArgument the proof's two roots are the statement's for its position
10c MemoryArgument a PUBLIC_INPUT or PUBLIC_OUTPUT shard's value column is the statement's string (public-values.md §5)
10b MemoryArgument every boundary timestamp below 2^38; v_10 = exit_status; gkr_verify::reconciles over every shard's roots and boundary_factors(challenges, vk.entry_pc, boundary) (memory.md §4.2)
11 — the opening claim (§5)
12 Opening the SrsVerifier, every commitment and the Mercury proof through their validating decoders, then pcs::batch_verify; any failure

Step 8 cannot fail on verify's output, whose base claims share layer 0's point; it states what step 11 relies on. Steps 1–3 hold the statement to the key before the replay indexes by it, so nothing a proof or statement carries makes the core panic, for a loaded key.

The split, by what each part reads (verifier_core):

function reads steps runs
derive_global_phase(vk, public) → GlobalChallenges key, statement 1–3, G1–G11 once a statement
verify_shard_local(vk, global, proof, public) → OpeningClaim and one ShardProof 4–10a, 10c, 11 once a shard
verify_global_memory(vk, global, public) key, statement, challenges 10b once a statement

GlobalChallenges is the four memory challenges and the digest. reduce_shard is the three in that order, verify_shard that and step 12; step 11 cannot fail, so 10b after it is 10b in place. Step 10b reads only vk.entry_pc, the boundary, the roots and the challenges, so a block runs it once; step 10a puts each shard into the product by holding the roots its GKR proof outputs to the statement's entry, and shard-set exactness makes every root there a verified shard's. verify_shard_local alone verifies no memory argument: without verify_global_memory it accepts shards, each valid, whose multiset does not close.

verifier::verify_block(vk, block, public):

class check
B1 Statement block.config is vk.config, and block.statement is public
B2 Statement derive_global_phase, once
B3 Statement BlockProof::shape (§1.3)
B4 Statement check_ts_windows over the records (§8)
B5 MemoryArgument verify_global_memory, once
B6 as verify_shard per shard, in statement order: verify_shard_local, then step 12

B1–B5 read no GKR proof or opening, so a statement that cannot reconcile is refused before any circuit runs, and the class can differ from verify_shard's: a change to anything G1–G9 absorb that B1–B4 admit moves the challenges, so the honest roots stop reconciling and verify_block answers MemoryArgument where verify_shard names the seed at step 5; a forgery that unbalances the multiset is MemoryArgument even where it also breaks a gate. A dropped shard fails B5: the truncated statement, re-proved honestly with its counts, lists and roots adjusted, passes B1–B4 and misses that shard's memory events on one side of the product.

In verify_shard's order the class names the fault: a tampered witness proved honestly, its multiplicities recounted, columns recommitted and statement rebuilt, fails at the gate (Constraint), table membership (Lookup) or multiset (MemoryArgument) it broke, which checker::TamperHarness asserts (circuits.md §3).

7 The verifying key#

7.1 Fields#

field
code_version: u32 constants::family::CODE_VERSION, 0
config: VmConfig the static shape (program.md §7)
entry_pc: u32 the image's entry pc
identity: ProgramIdentity program.md §8
setup_commitments: Vec<Vec<[u8; 64]>> identity's commitment lists, one per config family, in its order
srs_verifier: [u8; 320] the SrsVerifier (srs.md §5)
generic_table: [[u8; 64]; 3] the generic table's commitments, key column first, in every key (lookup.md §9)
srs_digest: Fr §3
circuits: Vec<FamilyCircuit> one per config family, in its order: the family, its CircuitArtifact and its ChannelSpecs (lookup.md §11)

A key carries every family's artifact, so its size is mostly its delegation families' (circuits.md §1).

7.2 Loading#

VerifyingKey::from_bytes decodes (§9), refuses bytes that are not the key's canonical encoding, and runs VerifyingKey::check, which refuses, in order:

  1. a VmConfig no derivation produces (VmConfig::from_bytes of its own bytes), or a code_version other than CODE_VERSION;
  2. a setup list count other than the config's family count;
  3. an identity that identity_digest(code_version, config, entry_pc, setup_commitments) does not reproduce;
  4. an srs_digest that srs_digest(srs_verifier, generic_table) does not reproduce;
  5. a circuit count other than the family count; then, family by family: a circuit for another family; a height the registry has no circuit for; a circuit, artifact or channel specs, other than config.circuit(family, trace_vars), the registry of the config's format (circuits.md §1); an artifact failing CircuitArtifact::validate, constraints::memory::check_memory or constraints::lookup::check_discharge; a setup list whose length, plus 3 if the circuit reads GENERIC, is not the artifact's S count; and GENERIC specs naming anything but the 3 setup columns after identity's, §5's order, which no registry circuit fails.

verifier::load_verifying_key then decodes every curve point: the SrsVerifier's three, each setup commitment and each generic-table commitment, through the validating readers. The circuits are held to the registry because nothing else binds them: identity binds the program, not the circuit that proves it.

A key from an untrusted source. Every field is recomputed from or compared with one of the verifier's two trusted values (§1), or fixed by the code: config, entry_pc and the setup lists through identity; srs_verifier and generic_table through the SRS digest; code_version and the circuits by the verifier's own registry. So a key may come from the prover, provided both comparisons are made.

Validation runs once, at load; verify_shard and verify_block assume a loaded key. On one edited in memory a changed config or circuit list is still refused as Statement, but an edit inside a circuit may go unnoticed.

prover::ProverSetup::new(program, srs) builds the key: each family's circuit from VmConfig::circuit and fill from prover::family_fill (ProverError::Unregistered if either is missing), identity's and the generic table's commitments over srs, then check (ProverError::Key). srs needs as many powers as the tallest family has rows, and 2^18 for the generic table (program::lookup_tables::GENERIC_LOG_HEIGHT); fewer panics.

8 Time windows#

Each shard claims [ts_start, ts_end) (ShardProof::ts_window): the slice of the clock (execution-trace.md §1) its rows write in, their reads reaching back before it. S2 absorbs it before the witness commitments, so a proof made under one window fails under another; step 4 holds it to ts_start ≤ ts_end ≤ 2^38 and nothing more.

verifier_core::check_ts_windows (B4), over the records in statement order: within each cycle-owning family (constants::family::CYCLE_OWNING, the execution families 0–6), every window is non-empty and no shard's ts_end exceeds the next shard's ts_start; a family's records are consecutive and ascending, so neighbours suffice. It is per family because families interleave — ADD_SUB_LUI_AUIPC may own cycles 1 and 3 and JUMP_BRANCH_SLT cycle 2 — and every other family is exempt: a window family's rows are words, a delegation family's invocations at their requesting cycles (delegation.md §8).

The prover reads a window off the shard's committed M[0] cycle column (ts_window, crates/prover/src/lib.rs): [4·c_0, 4·c_max + 4), c_0 row 0's cycle and c_max the largest, padding rows carrying 0, for cycle-owning and delegation families alike. Window families claim verifier_core::TRIVIAL_TS_WINDOW = [0, 2^38).

A window binds nothing. No gate ties it to the rows committed under it, so a prover may claim any windows the rule admits; cross-shard order, cycle uniqueness and pc continuity are the memory multiset's alone (memory.md §9). B4 checks the shape of the shard plan and adds nothing to soundness.

9 Wire forms and the proof archive#

verifier_core::wire: integers little-endian; an Fr its 32 canonical bytes (primitives.md §1), refused at or above p; a G1 its 64 bytes (primitives.md §3), opaque to the core; bytes a u32 length then the bytes; list<T> a u32 count then the items; T[k] exactly k items, no count. Every decoder is total: it refuses a count the remaining bytes cannot hold, so it reserves nothing an untrusted length asks for, and refuses trailing bytes.

PublicInputs   input bytes, output bytes, exit_status u32,
               shard_counts list<u32>, windows list<u32>,
               boundary Fr[64]                     memory.md §4.1's order and ranges
               memory_commitments list<list<G1>>, memory_roots list<Fr[2]>

ShardProof     family u32, shard_index u32, ts_start u64, ts_end u64, global_digest Fr,
               witness_commitments list<G1>, outputs list<Fr>,
               gkr list<(rounds list<Fr[4]>, final_evals list<Fr>)>       transition 0 first
               opening u8[704]                     pcs::MercuryProof, mercury.md §4

BlockProof     config bytes                        VmConfig, program.md §7
               statement bytes                     PublicInputs
               shards list<bytes>                  each a ShardProof; then BlockProof::shape

VerifyingKey   code_version u32, config bytes, entry_pc u32, identity Fr,
               setup_commitments list<list<G1>>, srs_verifier u8[320], generic_table G1[3],
               srs_digest Fr,
               circuits list<(family u32, artifact bytes,          CircuitArtifact, gkr.md §4.1
                              channels list<(channel u32, table list<Address>,
                                             multiplicity Address)>)>
Address        tag u8 (0 M, 1 W, 2 S, 3 V), index u32; a V's index is its gkr.md §2.1 kind tag

BlockReconciliation   list<(family u32, shard_index u32, ts_start u64, ts_end u64,
                            memory_commitments list<G1>, read_root Fr, write_root Fr)>

A ShardProof's lengths are fixed by its key and family, and steps 6–7 hold them: transition k carries n_{k+1} rounds and w_k claims, twice that if halving (gkr.md §5.5). For a base-format circuit at 2^n with W witness, C committed and I inner columns, O outputs, and R row-wise lists before its n halving ones, that is

772 + 64·W + 32·O + 8·(R + n) + 128·(R·n + n(n − 1)/2) + 32·(C + I + O·(n − 1))   bytes

which circuits.md §1 tabulates per family.

The proof archive. verifier::proof_archive::write_proof(dir, stem, vk, block), re-exported as host::proof_archive, writes four files, each a bare to_bytes with no header of its own:

<stem>.vk         VerifyingKey
<stem>.identity   the key's identity: its 32 bytes in order, 64 lowercase hex digits, a newline
<stem>.public     PublicInputs: the block's own statement
<stem>.block      BlockProof

read_proof(dir, stem) is the inverse, each file through its type's decoder and the key through load_verifying_key. .identity records what the run claimed, and read_proof returns it unchecked: a verifier's identity comes from its own channel (§1). .public repeats the statement .block carries, for the CLI, which takes it as a file (tools.md §6).

审计专区/指令族

ADD_SUB_LUI_AUIPC 电路族

规范原文docs/spec/add-sub.md以 Markdown 查看

摘要

电路族 0 的电路,逐列、逐门展开。除了 add、sub、addi、lui、auipc 和 fence,每一个 ecall 也都是这个电路族中的一行,因此它的电路还证明退出操作以及每次委托调用的请求一侧。本页列出它的已承诺列、63 个约束门、15 项查找义务、它为何可靠,以及它的限制。

以下规范原文以英文维护;英文是本规范的标准语言。

add, sub, addi, lui, auipc, ecall, ebreak and fence, compressed forms included, are family 0, one executed instruction a row; constraints::add_sub::artifact is its circuit. Every ecall is a row of it, so the circuit also proves the exit and the request side of every delegation call. This page specifies what it adds beside the memory frame (memory.md §2).

1 Columns#

The decoded tuple is pc next_pc rs1 rs2 rd imm extra_mask (program.md §5), the mask one-hot over the kinds system addi auipc add sub lui in bit order (constants::extra_mask::add_sub_lui_auipc). ecall, ebreak and fence share the system kind, with imm 0, 1 and 2 (constants::extra_mask::system_code); elsewhere imm is what the instruction adds — addi's sign-extended immediate, lui's and auipc's shifted left by 12, 0 for add and sub — and a register field the instruction lacks is 0.

M[0..26] and W[0..8] are the frame of the five queries pc rs1 rs2 rd deleg, and M[26], deleg_space, is the requested delegation type's anchor address space, 0 on a row requesting none (memory.md §2). The family adds these columns, and reads V[range19] and V[range16]:

column name
W[8..14] decoded_next_pc, decoded_rs1, decoded_rs2, decoded_rd, decoded_imm, decoded_mask the claimed decoded row
W[14..20] kind_system … kind_lui b_k, the mask's bits
W[20], W[21] is_ecall, is_fence the system kind, split by its code
W[22..28] is_deleg_<f>, f = 9, 10, 11, 15, 16, 17 d_t: a request of delegation type t, family f
W[28], W[29] wrap, rd_hi the sum's carry or the difference's borrow; sel's high halfword
W[30], W[31] pc_wrap, next_pc_hi next_pc's wrap and high halfword
W[32..35] mult_timestamp, mult_range16, mult_decoder one multiplicity a channel
S[0..7] table_pc … table_extra_mask the decoded table

Below, m_q, a_q, ts_q and v_q are query q's mask, address, read timestamp and read value; pc and next_pc the pc query's read and write values; sel is rd_selected (W[7]), the value the frame writes to a nonzero rd. N_t and tag_t are type t's ecall number and anchor space, the first six rows of constants::delegation::TYPES in order (delegation.md §3); is_exit = is_ecall − Σ_t d_t; 93 is constants::ecall::EXIT and HALT_PC is 1 (memory.md §5).

The family's fill (prover::family_fill, crates/prover/src/fill.rs) writes sel as the computed value even where rd = x0.

2 Gates#

63 enforcing gates, all in gate list 0, each of degree at most 2 and 0 on the all-zero row: the frame's eleven (memory.md §2) and these 52, in artifact order, each held to 0:

gate expression
kind_<k>_boolean, six b_k − b_k²
decoded_mask_bits Σ_k 2^k·b_k − decoded_mask
is_ecall_boolean, is_fence_boolean y − y²
system_split is_ecall + is_fence − b_system
ecall_code is_ecall·decoded_imm
fence_code is_fence·(decoded_imm − 2)
per type: is_deleg_<f>_boolean, deleg_<f>_is_an_ecall, deleg_<f>_number d_t − d_t²; d_t·(1 − is_ecall); d_t·(v_rs1 − N_t)
ecall_is_exit is_exit·(v_rs1 − 93)
rs1_mask_rule m_rs1 − m_pc·(b_add + b_sub + b_addi + is_ecall)
rs2_mask_rule m_rs2 − m_pc·(b_add + b_sub + is_ecall)
rd_mask_rule m_rd − m_pc·(b_add + b_sub + b_addi + b_auipc + b_lui + is_ecall)
deleg_mask_rule m_deleg − m_pc·Σ_t d_t
rs1_addr_rule m_rs1·(a_rs1 − decoded_rs1 − 17·is_ecall)
rs2_addr_rule, rd_addr_rule m_q·(a_q − decoded_q − 10·is_ecall)
rs1_value_masked, rs2_value_masked v_q − m_q·v_q
add_addi_auipc (b_add + b_addi + b_auipc)·(v_rs1 + v_rs2 + decoded_imm − sel − 2^32·wrap) + b_auipc·pc
sub b_sub·(v_rs1 − v_rs2 − sel + 2^32·wrap)
lui b_lui·(decoded_imm − sel)
exit_status is_exit·(v_rd − sel)
deleg_writes_no_register m_deleg·sel
deleg_read_ts_zero, deleg_read_value_zero m_deleg·ts_deleg; m_deleg·v_deleg
deleg_addr_rule m_deleg·(a_deleg − v_rs2)
deleg_space_rule deleg_space − Σ_t tag_t·d_t
wrap_boolean, pc_wrap_boolean y − y²
next_pc_rule next_pc + 2^32·pc_wrap − (1 − is_exit)·decoded_next_pc − is_exit·HALT_PC

N_t and tag_t are literals read from constants::delegation::TYPES, so the base circuit depends on the registry's first BASE_TYPES = 6 rows and on no row appended after them. The recursion format's circuit, add_sub::recursion_artifact, carries a selector and its three gates for each of the ten types, the columns after them shifted by four, and in place of deleg_writes_no_register deleg_a0_rule, Σ_{t<6} d_t·sel + Σ_{t≥6} d_t·(sel − v_rs2 − 4·words_t) with words_t the type's frame length: a recursion type's request leaves a0 past its frame (recursion.md §1.4).

3 Lookups#

15 obligations on three channels, none of them GENERIC, so the setup columns are the decoded table alone: the frame's ten TIMESTAMP gaps, two a query under its mask (memory.md §2), and five under m_pc:

lookup channel tuple
rd_hi_range, rd_lo_range RANGE16 rd_hi; sel − 2^16·rd_hi
next_pc_hi_range, next_pc_lo_range RANGE16 next_pc_hi; next_pc − 2^16·next_pc_hi
decode_row DECODER pc, decoded_next_pc, decoded_rs1, decoded_rs2, decoded_rd, decoded_imm, decoded_mask

The channels, in output order (add_sub::channels), are TIMESTAMP over V[range19], RANGE16 over V[range16] and DECODER over S[0..7].

4 Why it is sound#

On a live row (m_pc = 1) decode_row makes the claimed row the table's at pc, so exactly one b_k is 1 (lookup.md §10). The mask rules make each query present exactly where the row's kind or request makes it (execution-trace.md §4, §6). The address rules make a register query the decoded register, or on an ecall row, whose decoded registers are 0, a7 (17) for rs1 and a0 (10) for rs2 and rd. The _value_masked gates make an absent operand read 0, which lets one gate serve three sums: an addi or auipc row's v_rs2, and an auipc row's v_rs1, would otherwise be free addends, and add's imm is the table's 0.

Read values are words (memory-ops.md §5), sel is a word by its range pair and wrap is boolean, so each arithmetic gate is an identity over ℤ with one solution: the sum mod 2^32 and its carry, the difference mod 2^32 and its borrow, or imm. Without the pair, a sum at or above 2^32 would satisfy the gate with wrap = 0 and reach a register. The frame's x0 rule then writes sel or discards it.

next_pc is a word by its range pair and is decoded_next_pc — the table's fall-through, so a compressed instruction advances by 2 (program.md §5) — or HALT_PC on the exit row, less 2^32·pc_wrap. Both are far below 2^32, so pc_wrap = 0 on every live row.

On a system row system_split sets exactly one of is_ecall and is_fence, and the code gates make it the one imm names; ebreak's code 1 satisfies neither, so an ebreak row is unprovable. A fence row makes no query but the pc's and falls through. Off a system row both bits are 0, and so, by deleg_<f>_is_an_ecall, is every d_t.

A set d_t forces is_ecall = 1 and a7 = N_t. The numbers are pairwise distinct and none is 93, const assertions beside the circuit, so at most one d_t is set, is_exit is 0 or 1, and an ecall row is the exit, with a7 = 93, or a request of exactly one type; no other a7 passes.

  • The exit row writes back the a0 it read (exit_status), so x10's final value is the status the statement carries (proof.md §6, step 10b), and writes HALT_PC, after which no row runs (memory.md §5).
  • A request row falls through, writes 0 to a0, and makes the mirror query at the frame base it read from a0, in the space deleg_space names, reading timestamp 0 and value 0. Those three zeroings pair it one-to-one with an invocation of its type (delegation.md §5); the mirror's write value is free here, and what the call computed is the invoked family's circuit (delegation-circuits.md). deleg_space is an M column because a memory leaf may read no W column (memory.md §8); deleg_space_rule ties it to the selectors.

A row with m_pc = 0 is bound to no table row and its kind bits are free; the arithmetic gates are gated by those bits alone, m_pc times a bit being degree 3. That is harmless: the mask rules zero the row's other four masks and every lookup is off, so it adds no memory tuple. On a live row, wrap outside the four sums and sel on a fence row are free, and nothing reads them.

5 Limits#

  • An ecall whose a7 is neither 93 nor a type the format's circuit knows has no proof. The emulator answers an unassigned number -ENOSYS and continues (ecall-abi.md §5); the fill refuses that trace, naming the cycle.
  • An ebreak has no proof; it is fatal in the emulator (execution-trace.md §10).
  • No row touches RAM: an ecall row reads a7 and a0 and writes a0, and a request's operands travel in the invoked family's frame (delegation.md §4).

审计专区/指令族

JUMP_BRANCH_SLT 电路族

规范原文docs/spec/jump-branch-slt.md以 Markdown 查看

摘要

电路族 1 的电路:set-less-than 类指令、六种分支、jalr 与 jal。它规定了各列;判零 gadget 与比较 gadget,后者只用一个经范围检查的差值,无需比较表就能同时确定有符号与无符号的大小关系;各个门,包括带偶数检查的唯一一条 next_pc 规则;查找;以及为什么该电路恰好只接受 ISA 所规定的行为。

以下规范原文以英文维护;英文是本规范的标准语言。

The circuit of slti, sltiu, slt, sltu, the six branches, jalr and jal: what it adds beside the memory frame every execution family carries (memory.md §2), and the two gadgets other families reuse (§3). One comparison settles signed and unsigned order for the branches and the slt kinds alike. The circuit is constraints::jump_branch_slt::artifact (crates/constraints/src/jump_branch_slt.rs); prover::family_fill writes its witness.

1 What the circuit reads from the decoded table#

The tuple is pc next_pc rs1 rs2 rd imm extra_mask (program.md §5). next_pc is the fall-through, seq below; imm is the two's-complement word of the value the instruction uses: the sign-extended immediate of slti and sltiu (which sltiu compares unsigned), a branch's or jal's displacement, jalr's offset. extra_mask is one-hot over constants::extra_mask::jump_branch_slt:

bit    0     1      2    3     4    5    6    7    8     9     10    11
kind   slti  sltiu  slt  sltu  beq  bne  blt  bge  bltu  bgeu  jalr  jal

The legal masks are these twelve one-bit values, jump_branch_slt::LEGAL_MASKS; rd = x0 is the table's rd, not a mask. The circuit commits the twelve bits b_k, and every signal it needs is a linear form over them: the signed-comparison flag sc = b_slti + b_slt + b_blt + b_bge, the compared immediate (b_slti + b_sltiu)·imm, so that a branch's displacement never reaches the comparison, and the branch weights of taken_rule (§4).

2 Columns#

The frame is M[0..21] and W[0..7], over the queries pc rs1 rs2 rd at slots 0–3 (memory.md §2). Its W[6], rd_selected (sel below), holds the value the instruction computes, which the x0 rule masks into the write. The circuit adds:

column name value
W[7..13] decoded_next_pc … decoded_mask the claimed decoded row after pc
W[13..25] kind_slti … kind_jal the bits b_k, in §1's order
W[25] cmp_rhs the right operand, rs2 + (b_slti + b_sltiu)·imm
W[26], W[27] rs1_hi, rs1_sign rs1 >> 16, rs1 >> 31
W[28], W[29] cmp_rhs_hi, cmp_rhs_sign the same of cmp_rhs
W[30] lt rs1 < cmp_rhs, signed where sc = 1
W[31], W[32] cmp_gap, cmp_gap_hi (rs1 − cmp_rhs) mod 2^32, and its high halfword
W[33], W[34] eq, eq_inv [rs1 = cmp_rhs] on a live row; the difference's inverse
W[35] taken a taken branch
W[36] jalr_drop bit 0 of rs1 + imm on a jalr row
W[37] pc_wrap the carry out of whichever sum next_pc is
W[38], W[39] next_pc_hi, rd_hi next_pc >> 16, sel >> 16
W[40..44] mult_timestamp … mult_decoder one multiplicity per channel, in channel order
S[0..7] table_pc … table_extra_mask the decoded table, which program identity binds
S[7..10] generic_key … generic_result the packed table (lookup.md §9)
V[range19], V[range16] the TIMESTAMP and RANGE16 tables

21 M, 44 W and 10 S columns, 75 committed; 42 enforcing gates, the frame's 10 and §4's 32; 22 lookups: 8 TIMESTAMP, 11 RANGE16, 2 GENERIC, 1 DECODER, counts artifact asserts.

3 The gadgets#

constraints::gadgets returns gates and lookups as data. is_zero also builds the frame's x0 rule (memory.md §2) and MUL_DIV's zero tests (mul-div.md); the comparison also orders ATOMICS' minimum and maximum (memory-ops.md §6).

3.1 is_zero(x, inv, z, enable)#

x·inv + z − enable = 0          x = Σ c_i·x_i, a linear form
z·x = 0

With enable boolean, which the caller establishes, these force z = enable·[x = 0]: at x ≠ 0 the second gives z = 0 and the first inv = enable/x; at x = 0 the first gives z = enable. So z is boolean with no gate of its own, and enable = 0 gives z = 0, which keeps the all-zero row valid.

3.2 The comparison#

Comparison names one comparison lhs < rhs by its columns, its lookups' selector, and the kind bits signed whose sum is sc, which the caller holds to 0 or 1 on a selected row. comparison returns, for x each of lhs, rhs and gap:

name kind expression
<p>_order gate lhs − rhs − 2^32·sc·lhs_sign + 2^32·sc·rhs_sign + 2^32·lt − gap
<p>_lt_boolean gate lt − lt²
<p>_<x>_hi_range, <p>_<x>_lo_range RANGE16 x_hi; x − 2^16·x_hi
<p>_lhs_get_sign, <p>_rhs_get_sign GENERIC (x_hi + SIGN_BASE, x_sign, 0)

The range pairs make lhs, rhs and gap words and each x_hi the true high halfword (memory.md §7), which keeps each sign key inside U16GetSign's range (lookup.md §4), so each sign is its operand's bit 31. Let D = lhs − rhs − 2^32·sc·(lhs_sign − rhs_sign): both operands read in two's complement where sc = 1, so mixed signs are no case split, and D ∈ (−2^32, 2^32). The gate says gap = D + 2^32·lt, and only lt = [D < 0] puts gap in [0, 2^32): at D ≥ 0, lt = 1 puts it at 2^32 or above; at D < 0, lt = 0 makes it a negative field element. So the range check on gap carries the order, and no comparison table exists; the honest gap is (lhs − rhs) mod 2^32 whatever sc is. Both gates are ungated, since a selector would make the order gate degree 3, and every row satisfies them with the gap its own values give. comparison_equation(c, word_bits) builds the order gate at any width to 32, and the row suite evaluates it at 6 bits over every operand pair, signed and unsigned, finding exactly one (lt, gap), the ISA's.

4 Gates#

After the frame's ten in gate list 0, with m_q, a_q, v_q query q's mask, address and read value, and pc, next_pc the pc query's read and write:

gate polynomial
kind_<k>_boolean ×12 b_k − b_k²
decoded_mask_bits Σ_k 2^k·b_k − decoded_mask
rs1_mask_rule m_rs1 − m_pc·(Σ_k b_k − b_jal)
rs2_mask_rule m_rs2 − m_pc·(b_slt + b_sltu + the six branch bits)
rd_mask_rule m_rd − m_pc·(b_slti + b_sltiu + b_slt + b_sltu + b_jalr + b_jal)
<q>_addr_rule, for rs1, rs2, rd m_q·(a_q − decoded_q)
<q>_value_masked, for rs1, rs2 v_q − m_q·v_q
cmp_rhs_rule cmp_rhs − v_rs2 − (b_slti + b_sltiu)·imm
cmp_order, cmp_lt_boolean §3.2: lhs = v_rs1, rhs = cmp_rhs, signed the bits of sc
eq_inverse, eq_at_nonzero §3.1: x = v_rs1 − cmp_rhs, z = eq, enable = m_pc
taken_rule taken − w_1 − w_eq·eq − w_lt·lt
taken_boolean, jalr_drop_boolean, pc_wrap_boolean x − x²
next_pc_rule §5's equation
rd_value_rule sel − (b_jal + b_jalr)·seq − (b_slti + b_sltiu + b_slt + b_sltu)·lt

The branch weights are w_1 = b_bne + b_bge + b_bgeu, w_eq = b_beq − b_bne and w_lt = b_blt + b_bltu − b_bge − b_bgeu. Every gate has degree at most 2 and is 0 on the all-zero row, which artifact asserts.

4.1 Lookups#

After the frame's 8 TIMESTAMP obligations, all under m_pc:

lookup channel expression
cmp_<x>_hi_range, cmp_<x>_lo_range ×3 RANGE16 §3.2 over v_rs1, cmp_rhs, cmp_gap
cmp_lhs_get_sign, cmp_rhs_get_sign GENERIC §3.2
rd_hi_range, rd_lo_range RANGE16 rd_hi; sel − 2^16·rd_hi
next_pc_hi_range, next_pc_lo_range RANGE16 next_pc_hi; next_pc − 2^16·next_pc_hi
next_pc_even RANGE16 2^−1·next_pc − 2^15·next_pc_hi
decode_row DECODER pc and W[7..13] (lookup.md §10)

The channels, in output order, are TIMESTAMP on V[range19], RANGE16 on V[range16], GENERIC on S[7..10] and DECODER on S[0..7]. next_pc_even is the low halfword lo halved, (lo + p)/2 and far above 2^16 when lo is odd. Because it scales next_pc, the constructor runs lookup::check_copowers (lookup.md §11) over (next_pc, m_pc).

5 Why it is sound#

On a live row, m_pc = 1, the decoder lookup makes the claimed tuple the table's row at pc, so pc is even, seq is below 2^24 and exactly one b_k is 1 (lookup.md §10). The mask and address rules make the frame's queries the instruction's (execution-trace.md §4): jal reads nothing, and a branch has no rd query, so nothing it computes is written. An absent operand reads 0, so cmp_rhs is rs2 or the immediate, never their sum, and the comparison's pairs make both operands words. So lt is the ISA's order (§3.2), eq its equality (§3.1), and taken its branch decision: eq on beq, 1 − eq on bne, lt on blt and bltu, 1 − lt on bge and bgeu, and 0 off the branches, every term of taken_rule carrying a branch bit. taken is a committed bit because, inlined, taken·(pc + imm) would be degree 3.

next_pc is held by one gate:

next_pc + 2^32·pc_wrap = (1 − taken − b_jal − b_jalr)·seq
                       + (taken + b_jal)·(pc + imm)
                       + b_jalr·(v_rs1 + imm − jalr_drop)

At most one of taken, b_jal, b_jalr is 1, so one sum is selected, and one wrap bit outside the selectors serves all three: imm is a two's-complement word, so every backward branch and jump wraps, not only jalr. With next_pc an even word and pc_wrap, jalr_drop boolean:

  • the default arm is seq, below 2^24, so pc_wrap = 0;
  • pc + imm and v_rs1 + imm are below 2^33, so one wrap bit holds the carry, uniquely;
  • on jalr, v_rs1 + imm − jalr_drop − 2^32·pc_wrap is a unique even word, (rs1 + imm) mod 2^32 with bit 0 cleared; a false jalr_drop makes next_pc odd or negative.

A branch's or jal's target is even unchecked, pc and imm both being even. Evenness is what keeps the family off HALT_PC = 1 (memory.md §5): without next_pc_even, a jalr whose rs1 + imm ≡ 1 keeps bit 0 and writes HALT_PC with every other gate and lookup holding, and a program that would crash by jumping to address 0 is proven to exit cleanly.

The link is seq, a table value and not a sum, so it has no wrap bit; the rd pair range-checks it and lt like every register write, and the x0 rule masks both at x0. A target needs no check of its own: at an address holding no instruction, the next row's decoder lookup fails whatever family claims the row, no table holding a live row there (lookup.md §10).

On a padding row, m_pc = 0, the mask rules zero every query mask, eq is 0 and every lookup is off, so the row reaches no memory event whatever its free bits hold.

审计专区/指令族

SHIFT_BITWISE 电路族

规范原文docs/spec/shift-bitwise.md以 Markdown 查看

摘要

电路族 2 的电路:六种移位,以及按位 and、or 和 xor。任一方向的移位都是与一个从通用表中查得的 2 的幂的一次乘积;AND 是四次字节查找,OR 与 XOR 是基于 AND 结果的线性形式。本页规定了各列、各表、为什么每个键都必须自带界、copower 检查、门与查找,以及可靠性论证。

以下规范原文以英文维护;英文是本规范的标准语言。

The circuit of the shifts sll, slli, srl, srli, sra, srai and the bitwise and, andi, or, ori, xor, xori, one family, beside the memory frame (memory.md §2). A shift either way is one product with a looked-up power of two; AND is four byte lookups, and OR and XOR are linear forms over it. The circuit is constraints::shift_bitwise::artifact (crates/constraints/src/shift_bitwise.rs); prover::family_fill writes its witness.

1 What the circuit reads from the decoded table#

The tuple is pc next_pc rs1 rs2 rd imm extra_mask (program.md §5), next_pc the fall-through, seq below. imm is the shamt of slli, srli and srai, below 32 because the decoder refuses shamt[5] on RV32; the sign-extended immediate, as a word, of andi, ori and xori; and 0 on a register form. extra_mask is one-hot over constants::extra_mask::shift_bitwise, the legal masks its twelve one-bit values (shift_bitwise::LEGAL_MASKS):

bit    0     1     2     3     4    5     6    7    8    9    10   11
kind   slli  xori  srli  srai  ori  andi  sll  xor  srl  sra  or   and

The second operand of all twelve is src2 = rs2 + imm: an immediate form has no rs2 query, so rs2 reads 0, and a register form's imm is 0. One addend is always zero, so the sum needs no wrap bit, and an immediate never enters the rs2 column the memory argument ties. The circuit commits the twelve bits b_k, and its flags are linear forms over them:

left    b_slli + b_sll
right   b_srli + b_srai + b_srl + b_sra
arith   b_srai + b_sra
t1      b_or + b_ori + b_xor + b_xori
t2      b_and + b_andi − b_or − b_ori − 2·(b_xor + b_xori)

Only the two halves' sums, f_shift and f_bitwise, are columns: each selects lookups, and a selector is a committed boolean (lookup.md §2).

2 Columns#

The frame is M[0..21] and W[0..7], over the queries pc rs1 rs2 rd at slots 0–3 (memory.md §2); its W[6], rd_selected (sel below), holds the value the instruction computes, which the x0 rule masks into the write. The circuit adds:

column name value
W[7..13] decoded_next_pc … decoded_mask the claimed decoded row after pc
W[13..25] kind_slli … kind_and the bits b_k, in §1's order
W[25], W[26] f_shift, f_bitwise the two halves
W[27], W[28] rs1_hi, rs1_sign rs1 >> 16, rs1 >> 31
W[29] src2_hi src2 >> 16
W[30] amount src2 & 31
W[31], W[32] pow, copow 2^amount, 2^(31 − amount) on a shift row
W[33], W[34] high, high_hi src2 >> 5, and its high halfword
W[35] se arith·rs1_sign
W[36], W[37] shift_in, shift_prod both directions' multiplicand, and shift_in·pow
W[38], W[39] ovf, ovf_hi a left shift's discarded high word, and its high halfword
W[40], W[41] residue, residue_hi a right shift's remainder, and its high halfword
W[42], W[43] scaled, scaled_hi residue·2^(32 − amount), and its high halfword
W[44..52] byte_a<j>, byte_b<j> the bytes of rs1, then of src2, low first
W[52..56] byte_and<j> their bytewise AND
W[56] rd_hi sel >> 16
W[57..61] mult_timestamp … mult_decoder one multiplicity per channel, in channel order
S[0..10], V[range19], V[range16] as in jump-branch-slt.md §2

21 M, 61 W and 10 S columns, 92 committed; 48 enforcing gates, the frame's 10 and §4's 38; 39 lookups: 8 TIMESTAMP, 24 RANGE16, 6 GENERIC, 1 DECODER, counts artifact asserts.

3 Tables, and the bound on every key#

3.1 ShiftPowers#

Row s of the packed table's top sub-table (lookup.md §9) is (SHIFT_BASE + s + 1, 2^s, 2^(31 − s)), one for each of the 32 shift amounts and for none other, so a key past its last row matches nothing. The second value is the copower a residue bound multiplies by, 2^(32 − s), stored halved (SHIFT_COPOWER_BITS = 31): at s = 0 it is 2^32, which the table's u32 columns cannot hold, so the two gates that read it carry the factor 2 (§4.2, §4.3).

3.2 The AND rows#

An AND row is (AND_BASE + a + 1, b, a & b) over bytes a and b, so a key inside their range matches a row that makes byte_b<j> a byte and byte_and<j> its AND with byte_a<j>: those two need no bound of their own.

3.3 Every key is bounded#

The channel proves membership of the packed table, not of a sub-table (lookup.md §4), so an out-of-range key lands on another sub-table's row. A bitwise row with byte_a0 = 65,823 gates to key 65,824, ShiftPowers' row (65,824, 2^31, 1); with rs1 = 65,823 and rs2 = 2^31 every gate holds, and and writes 1 where the answer is 0. So every key carries its own bound, as RANGE16 obligations under its lookup's selector:

key bound obligations selector
rs1_hi + SIGN_BASE rs1_hi < 2^16 rs1's 16+16 pair m_pc
amount + SHIFT_BASE amount < 2^5 amount; 2^11·amount f_shift
byte_a<j> + AND_BASE byte_a<j> < 2^8 byte_a<j>; 2^8·byte_a<j> f_bitwise

A bound below a halfword takes both obligations: the scaled one alone does not make the key an integer (lookup.md §11), and the direct one alone admits every halfword, byte_a0 = 256 landing on U16GetSign's row (257, 0, 0).

3.4 The copower check#

artifact runs lookup::check_copowers (lookup.md §11) over each column it bounds by scaling, which must carry its direct bound under its scaled obligation's own selector: residue, scaled by the looked-up copower (§4.3), under m_pc; amount under f_shift; each byte_a<j> under f_bitwise.

4 Gates#

Gate list 0 holds the frame's ten and these 38. m_q, a_q, v_q are query q's mask, address and read value, and a flag of §1 times (…) stands for each of its weighted bits times (…), so every term is of degree 2.

4.1 Presence and next_pc#

gate polynomial
kind_<k>_boolean ×12, decoded_mask_bits as in jump-branch-slt.md §4
f_shift_rule, f_bitwise_rule f − Σ its half's six bits
f_shift_boolean, f_bitwise_boolean f − f²
rs1_mask_rule, rd_mask_rule m_q − m_pc·Σ_k b_k
rs2_mask_rule m_rs2 − m_pc·(b_sll + b_srl + b_sra + b_and + b_or + b_xor)
<q>_addr_rule ×3, <q>_value_masked ×2 as in jump-branch-slt.md §4
next_pc_rule next_pc − seq

No kind computes a pc: next_pc is the decoder-bound fall-through, with no wrap bit and no bound of its own, and HALT_PC is beyond the family's reach (memory.md §5).

4.2 The shift amount#

amount_split    rs2 + imm − 32·high − amount
copower_rule    pow·copow − 2^31·f_shift

amount_split is ungated. copower_rule says pow·(2·copow) = 2^32 on a shift row, and pow·copow = 0 on a bitwise row.

4.3 The one product, both directions#

se_rule           se − arith·rs1_sign
rs1_sign_boolean  rs1_sign − rs1_sign²
se_boolean        se − se²
shift_in_rule     shift_in − left·v_rs1 − right·(sel − 2^32·se)
shift_prod_rule   shift_prod − shift_in·pow
shift_out_rule    left·(shift_prod − sel − 2^32·ovf)
                    + right·(shift_prod + residue − v_rs1 + 2^32·se)
scaled_rule       scaled − 2·residue·copow

shift_prod_rule, ungated, is the one multiplication by pow; shift_in_rule picks its multiplicand, which keeps shift_out_rule at degree 2 where left·(v_rs1·pow − …) would be 3, and se is committed for the same reason. A right shift is the floor division rs1 − 2^32·se = (sel − 2^32·se)·2^s + residue, which covers sra: the arithmetic shift of a negative word is the floor division of its signed value, and the result keeps the operand's sign. shift_in and shift_prod are the only columns that are not words: the multiplicand is negative where se = 1, and a left shift's product reaches 2^63.

4.4 The bitwise half#

rs1_bytes         v_rs1 − Σ_j 2^(8j)·byte_a<j>
src2_bytes        rs2 + imm − Σ_j 2^(8j)·byte_b<j>
bitwise_out_rule  f_bitwise·sel − t1·(v_rs1 + rs2 + imm) − t2·Σ_j 2^(8j)·byte_and<j>

Per byte, OR is a + b − (a & b) and XOR is a + b − 2·(a & b). Summed by weight through the two decompositions, sel is rs1 & src2 at (t1, t2) = (0, 1), their OR at (1, −1) and their XOR at (1, −2), exactly, no carry crossing a byte: there is no OR or XOR table, and the AND accumulator is a linear form, not a column. sel is gated by f_bitwise because t1 and t2 are 0 on a shift row, where a bare sel would force rd = 0. The decompositions are ungated: on a shift row the bytes carry no lookup, and a decomposition always exists.

4.5 Lookups#

After the frame's 8 TIMESTAMP obligations, in the channel order of jump-branch-slt.md §4.1:

RANGE16   <x>_hi_range, <x>_lo_range     under m_pc, x = rs1 src2 high ovf residue scaled rd
          amount_range, amount_scaled    under f_shift      §3.3
          byte_a<j>_range, _scaled ×4    under f_bitwise    §3.3
GENERIC   rs1_get_sign   (rs1_hi + SIGN_BASE, rs1_sign, 0)                 under m_pc
          shift_powers   (amount + SHIFT_BASE, pow, copow)                 under f_shift
          and_byte_<j>   (byte_a<j> + AND_BASE, byte_b<j>, byte_and<j>)    under f_bitwise
DECODER   decode_row     under m_pc (lookup.md §10)

5 Why it is sound#

On a live row the decoder lookup makes the claimed tuple the table's row at pc, so one kind bit is 1 and one of f_shift, f_bitwise (lookup.md §10); the mask and address rules make the queries the instruction's (execution-trace.md §4). rs1, src2 and sel are words by their pairs, rs1_hi is rs1's true high halfword and rs1_sign its bit 31. Every term of §4 that reads sel carries a shift bit or f_bitwise, so the inactive half never constrains it.

  • The amount is the ISA's. §3.3 bounds amount below 32 and high's pair bounds high below 2^32, so amount_split is an integer identity below 2^37, amount = src2 mod 32, and the ShiftPowers row it keys gives pow = 2^amount. Without high's pair, sll by rs2 = 4 can shift by 8, at high = −1/8.
  • A left shift: shift_prod = rs1·2^s < 2^63, and sel + 2^32·ovf, both words, is its unique split, so sel = (rs1·2^s) mod 2^32.
  • A right shift: se is rs1's bit 31 on sra and srai and 0 otherwise, so se_rule alone keeps an srai from carrying srli's answer. Every term of the floor division is below 2^64 in magnitude, so residue is the integer (rs1 − 2^32·se) − (sel − 2^32·se)·2^s, and scaled's pair puts it in [0, 2^s): sel − 2^32·se is the floor of (rs1 − 2^32·se)/2^s. residue's own pair, which check_copowers requires, bounds it without appeal to sel's, the scaled pair alone saying nothing of a non-integer: 2^−28 passes it at s = 3.
  • A bitwise result: each byte_a<j> is below 256, so its lookup matches an AND row (§3.2); with rs1 and src2 words, both decompositions are the unique byte splits and §4.4's identity holds.

copower_rule is implied by the bounded key and kept as the circuit's own reading of the table: a ShiftPowers row generated wrong stops the honest prover rather than license a residue bound that is not one. It also confines the key to ShiftPowers alone, no other row's two values having the product 2^31: an AND row's is at most 255·255, every other row's 0.

On a padding row, m_pc = 0, every query mask is 0 and every obligation under m_pc vacuous. f_shift and f_bitwise are free booleans there, so a padding row may look up ShiftPowers or the AND rows, which consumes a multiplicity and changes nothing.

审计专区/指令族

MUL_DIV 电路族

规范原文docs/spec/mul-div.md以 Markdown 查看

摘要

电路族 3 的电路,即 M 扩展。一个乘积恒等式同时服务于四种乘法和除法;符号规则与经范围检查的差值使除法向零截断,而不是向下取整;一个门固定除以零的结果。本页规定了各列、符号调整、各个门以及可靠性论证,包括有符号溢出的情形和诚实填充。

以下规范原文以英文维护;英文是本规范的标准语言。

The M extension — mul, mulh, mulhsu, mulhu, div, divu, rem, remu — as one circuit beside the memory frame every execution family carries (memory.md §2). One product identity serves the four multiplies and the division; a sign rule and a range-checked gap make the division truncated, and one gate pins division by zero. constraints::mul_div builds it: 21 M, 54 W and 9 S columns, 54 enforcing gates, 27 lookups.

1 What the circuit reads from the decoded table#

Every M instruction is R-type, so the decoded tuple has no imm: pc next_pc rs1 rs2 rd extra_mask, six columns (program.md §5). extra_mask is one-hot over constants::extra_mask::mul_div, bits 0–7 in the order above; the legal masks are its eight single bits (mul_div::LEGAL_MASKS), which the table's domain enforces (lookup.md §10). The circuit commits the bits b_k and reads every signal as a linear form over them:

signal form
reads rs1 signed b_mul + b_mulh + b_mulhsu + b_div + b_rem
reads rs2 signed b_mul + b_mulh + b_div + b_rem
a multiply, Σ_mul b_mul + b_mulh + b_mulhsu + b_mulhu
a division, f_div b_div + b_divu + b_rem + b_remu, a column: the is-zero gadgets' enable

mul is read signed × signed: its low word is the same either way, which lets one product identity serve all four multiplies. mulhsu's asymmetry is the two lists, not a case split.

2 Columns#

The frame is pc rs1 rs2 rd, M[0..21] and W[0..7] (memory.md §2.1); below, m_q, a_q and v_q are query q's mask, address and read value, and rs1, rs2 the operands' read values. The family adds:

W[7..12]   decoded_next_pc decoded_rs1 decoded_rs2 decoded_rd decoded_mask    (no imm)
W[12..20]  kind_mul … kind_remu
W[20]      f_div
W[21..27]  rs1_hi rs1_top rs2_hi rs2_top s1 s2          high halfwords, bit 31, §3
W[27..34]  mx my p_low p_low_hi p_high p_high_hi p_sign  the product
W[34..40]  q q_hi q_sign r r_hi r_sign                   quotient and remainder
W[40..45]  r_inv rz d1 d_inv dz                          is_zero(r), f_div·s1, is_zero(rs2)
W[45..50]  abs_r abs_d gap gap_hi rd_hi
W[50..54]  mult_timestamp mult_range16 mult_generic mult_decoder
S[0..6]    the decoded table, bound by identity
S[6..9]    the packed generic table (lookup.md §9)
V          range19 range16

The fill keeps mx, my (signed) and r_inv, d_inv (inverses) in Fr, every other column in u32.

3 The sign adjustments#

rs1_adj = rs1 − 2^32·s1      s1 = (b_mul + b_mulh + b_mulhsu + b_div + b_rem)·rs1_top
rs2_adj = rs2 − 2^32·s2      s2 = (b_mul + b_mulh + b_div + b_rem)·rs2_top
q_adj   = q − 2^32·q_sign    r_adj = r − 2^32·r_sign

rs1_top is the U16GetSign lookup of rs1_hi, which rs1's 16+16 pair makes its true high halfword, so the key lies in that sub-table's range and the answer is bit 31 (lookup.md §4); rs2_top likewise. An unsigned position forces its adjustment to 0 whatever the top bit, which keeps the selection degree 2. q_sign and r_sign are not sign lookups (§5.3). f_div, rs1_top, rs2_top, s1, s2, p_sign, q_sign and r_sign carry booleanity gates; rz and dz are boolean by the is-zero gadget (jump-branch-slt.md §3), d1 as a product of booleans.

4 Gates#

The frame's ten enforcing gates (memory.md §2.4) and the family's 44, all in gate list 0, each formula = 0. The plumbing:

kind_<k>_boolean          b_k − b_k²                       eight
decoded_mask_bits         Σ_k 2^k·b_k − decoded_mask
<q>_mask_rule             m_q − m_pc·Σ_k b_k               rs1, rs2, rd: every kind uses all three
<q>_addr_rule             m_q·(a_q − decoded_q)            rs1, rs2, rd
<q>_value_masked          v_q − m_q·v_q                    rs1, rs2
next_pc_rule              next_pc − decoded_next_pc        the fall-through (memory.md §5)

The arithmetic, mul_div::arithmetic_gates(32), written with §3's abbreviations:

f_div_rule, s1_rule, s2_rule   §1's and §3's forms, and eight booleanity gates (§3)
mx_rule                mx − Σ_mul b·rs1_adj − f_div·rs2_adj
my_rule                my − Σ_mul b·rs2_adj − f_div·q_adj
product_rule           mx·my − p_low − 2^32·p_high + 2^64·p_sign
division_rule          f_div·(p_low + 2^32·p_high − 2^64·p_sign + r_adj − rs1_adj)
rz_inverse             r·r_inv + rz − f_div          rz_at_nonzero   rz·r
dz_inverse             rs2·d_inv + dz − f_div        dz_at_nonzero   dz·rs2
d1_rule                d1 − f_div·s1
r_sign_rule            r_sign − d1 + d1·rz           so r_sign = f_div·s1·(1 − [r = 0])
abs_r_rule             abs_r − r − 2^32·r_sign + 2·r·r_sign       abs_r = |r_adj|
abs_d_rule             abs_d − rs2 − 2^32·s2 + 2·rs2·s2           abs_d = |rs2_adj|
gap_rule               gap − f_div·(abs_d − abs_r − 1) − 2^32·dz
zero_divisor_quotient  dz·(q − (2^32 − 1))
rd_value_rule          rd_selected − b_mul·p_low − (b_mulh + b_mulhsu + b_mulhu)·p_high
                         − (b_div + b_divu)·q − (b_rem + b_remu)·r

The lookups: the frame's eight TIMESTAMP gap chunks, each under its query's mask; and under m_pc, 16+16 RANGE16 pairs on rs1, rs2, p_low, p_high, q, r, gap and rd_selected, rs1_get_sign, (rs1_hi + SIGN_BASE, rs1_top, 0) on GENERIC, and rs2_get_sign, and decode_row on DECODER.

The width is a parameter of arithmetic_gates so the encoding can be checked whole: crates/checker/tests/mul_div.rs evaluates arithmetic_gates(4) through gkr::eval_gate over every (dividend, divisor) pair of a 4-bit word and each division kind, and exactly one (q, r) survives, RV32M's.

5 Why it is sound#

On a live row the decoder lookup makes exactly one kind bit 1 (lookup.md §10), and rs1, rs2 are words whose _top is bit 31, so rs1_adj, rs2_adj ∈ [−2^31, 2^32) are the operands as the kind reads them.

5.1 The product#

On a multiply row mx·my = rs1_adj·rs2_adj; on a division row it is rs2_adj·q_adj, q's pair and q_sign's booleanity putting q_adj in [−2^32, 2^32). Either way |mx·my| < 2^64, and two words and a boolean cover [−2^64, 2^64) once, so product_rule holds over the integers with one solution: p_low, p_high are the words of the 64-bit two's-complement product, RV32M's for each multiply. product_rule is ungated and the circuit's only product of two row values, which is what lets both readings share it at degree 2.

5.2 The division#

With rs2_adj ≠ 0, division_rule is rs2_adj·q_adj + r_adj = rs1_adj over the integers. Truncated division is its one solution with |r_adj| < |rs2_adj| and r_adj zero or of the dividend's sign, and two gates state exactly that:

  • The sign. r_sign = f_div·s1·(1 − [r = 0]) makes r_adj the word r on an unsigned row or a non-negative dividend, and r − 2^32 < 0 on a negative one unless r = 0. It is what separates truncated division from floored: without it DIV(−7, 2) admits q = −4, r = 1 as readily as q = −3, r = −1. As a definition, through d1, it is degree 2.
  • The magnitude. gap = abs_d − abs_r − 1 is range-checked, and neither magnitude reaches 2^32: abs_d ≤ 2^31 where s2 = 1, abs_r ≤ 2^32 − 1 where r_sign = 1, which needs r ≠ 0, and each is a word elsewhere. So the difference lies in [−2^32, 2^32), in range exactly when |r_adj| < |rs2_adj|. The comparison gadget would repeat bounds that hold and has no place for the zero divisor's term.

So q_adj and r_adj are RV32M's, and q_sign is pinned only by q's range: one value puts q_adj + 2^32·q_sign in [0, 2^32).

A zero divisor makes dz = 1 and rs2_adj = 0: the identity leaves r_adj = rs1_adj, so r is the dividend's word; zero_divisor_quotient, the one pin, makes q all ones; the 2^32·dz term lifts gap to 2^32 − 1 − abs_r, so the divisor imposes no bound. q_sign is free and harmless: mx = 0, and rd reads the word q.

The identity is gated. On a multiply row r_sign = 0 and r is a word, so an ungated identity would demand rs1_adj − rs1_adj·rs2_adj ∈ [0, 2^32), false for nearly every multiply: 7 × 3, a negative rs1 times x0.

5.3 The signed overflow, and why q_sign is free#

DIV(−2^31, −1) needs no pin: |r_adj| < 1 forces r = 0, the identity q_adj = 2^31, and q's range q_sign = 0, q = 0x80000000, RV32M's answer; REM gives 0. This row is why q_sign is a free boolean: tied to bit 31 of q, as s1 and s2 are to their operands', it would force q_adj = −2^31 and make the row unprovable. r_sign likewise follows the dividend's sign, not the remainder's word.

rd_selected's pair is implied by its four sources' and kept, every family bounding what it writes to rd (memory-ops.md §5). Every gate is zero on the all-zero padding row, which mul_div::artifact asserts with each channel's obligation count.

5.4 The fill#

prover::family_fill(MUL_DIV) (crates/prover/src/fill.rs) computes the witness with Rust's integers: the product in i128, the division by wrapping_div and wrapping_rem, which give RV32M's overflow answer, with the zero divisor an arm of its own, and q_sign from the sign of q_adj. It writes the computed value to rd_selected, which the frame's x0 rule masks, and panics, on rows the emulator cannot produce, if the identity does not divide, a product exceeds two words, or the trace's rd write or next_pc is not what the instruction computes.

审计专区/指令族

访存电路族

规范原文docs/spec/memory-ops.md以 Markdown 查看

摘要

电路族 4、5、6 的电路:字的加载与存储、字节与半字的加载与存储,以及原子操作。它规定了它们共享的寻址方式、MEM_WORD 的复制、MEM_SUBWORD 把子字拼接进所在字的方式、使每个寄存器值和 RAM 值始终保持为 32 位字的写侧归纳,以及 ATOMICS,包括与 RV32IMAC 唯一的偏差:sc.w 总是成功。

以下规范原文以英文维护;英文是本规范的标准语言。

MEM_WORD (lw, sw), MEM_SUBWORD (lb, lh, lbu, lhu, sb, sh) and ATOMICS (lr.w, sc.w, the nine AMOs): the execution families whose rows touch RAM, each a circuit beside the memory frame (memory.md §2), sharing §2's addressing. They are constraints::{mem_word, mem_subword, atomics}, filled by prover::family_fill (crates/prover/src/fill.rs), which computes each witness with Rust's integer operations.

family M W S gates, frame + own TIMESTAMP, RANGE16, GENERIC, DECODER
MEM_WORD 31 24 7 13 + 20 12, 5, 0, 1
MEM_SUBWORD 31 55 10 13 + 40 12, 22, 1, 1
ATOMICS 26 54 9 11 + 35 10, 19, 6, 1

Each artifact asserts its gate and obligation counts and that the all-zero padding row satisfies every gate. Below, m_q, a_q and v_q are query q's mask, address and read value, rs1 and rs2 the operands' read values (memory.md §2.1), and b_k (b_lw, b_lr, …) the committed kind bits.

1 What the circuits read from the decoded table#

MEM_WORD's and MEM_SUBWORD's tuple is pc next_pc rs1 rs2 rd imm extra_mask, imm the offset's two's-complement u32; ATOMICS' has no imm, its address being rs1 (program.md §5). The tuple is the first setup columns, and the packed generic table follows it where a family reads one: S[7..10] in MEM_SUBWORD, S[6..9] in ATOMICS (lookup.md §9). extra_mask is one-hot over constants::extra_mask, bit k the k-th mnemonic below, and each module's LEGAL_MASKS is those single bits:

mem_word      lw sw
mem_subword   lb lh lbu lhu sb sh
atomics       amoadd amoswap lr sc amoxor amoor amoand amomin amomax amominu amomaxu

The atomics order is ascending funct5; aq and rl order nothing on one hart and are not recorded. MEM_SUBWORD's modifiers are linear forms over its bits:

LOADK = b_lb + b_lh + b_lbu + b_lhu     BYTE = b_lb + b_lbu + b_sb     SIGNEXT = b_lb + b_lh
STORE = b_sb + b_sh                     HALF = b_lh + b_lhu + b_sh

All three carry the same plumbing, each formula = 0:

kind_<k>_boolean    b_k − b_k²
decoded_mask_bits   Σ_k 2^k·b_k − decoded_mask
<q>_mask_rule       m_q − m_pc·uses_q              every query but pc
<q>_addr_rule       m_q·(a_q − decoded_q)          rs1, rs2, rd
                    m_q·(a_q − 4·word_index)       load, ram (§2)
<q>_value_masked    v_q − m_q·v_q                  rs1, rs2
next_pc_rule        next_pc − decoded_next_pc      the fall-through (memory.md §5)
uses_q rs1 rs2 load ram rd
MEM_WORD b_lw + b_sw b_sw b_lw b_sw b_lw
MEM_SUBWORD LOADK + STORE STORE LOADK STORE LOADK
ATOMICS every bit every bit but b_lr no query every bit every bit

m_rs2 is keyed on b_lr, the one kind without an rs2 field, and not on rs2 = x0: an amoadd.w whose rs2 is x0 still reads it.

2 Addressing#

The effective address is rs1 + imm mod 2^32, or rs1 for an atomic. One degree-1 gate splits it, with wrap, bit0 and bit1 boolean:

MEM_WORD      addr_split   rs1 + imm − 2^32·wrap − 4·word_index
MEM_SUBWORD   addr_split   rs1 + imm − 2^32·wrap − 4·word_index − 2·bit1 − bit0
ATOMICS       addr_word    rs1 − 4·word_index

Over Fr that says nothing, 4 being a unit. Three RANGE16 obligations under m_pc, on word_index_hi, word_index − 2^16·word_index_hi and 4·word_index_hi (word_index_hi_range, word_index_lo_range, word_index_hi_scaled), cap word_index at 2^30 − 1, the top word's. With rs1 a word (§5) and imm a table value the split is then one of integers: wrap is the true carry, bit1 and bit0 the true low bits, and every RAM address is a 4-aligned address below 2^32. Having no offset bits, a misaligned MEM_WORD or ATOMICS access needs a word_index that is not an integer, which its pair refuses; the emulator refuses it first (execution-trace.md §10). addr_word derives rs1 < 2^32 rather than assuming it. half_aligned, HALF·bit0 = 0, refuses a halfword at an odd address and keeps w·p a divisor of 2^32 (§4.3).

Every RAM query's address is 4·word_index, so byte, halfword, word and atomic accesses to one word name one cell; the byte position lives only in MEM_SUBWORD's splice. Confining an access to initialized memory is the multiset's (memory.md §9): an out-of-window access fails the statement's memory argument, not a gate.

3 MEM_WORD#

A load copies the word into rd, a store copies rs2 into the word; there is no splice, no generic lookup, and the decoded table is the only setup.

W[9..15]   decoded_next_pc decoded_rs1 decoded_rs2 decoded_rd decoded_imm decoded_mask
W[15..21]  kind_lw kind_sw wrap word_index word_index_hi rd_hi
W[21..24]  mult_timestamp mult_range16 mult_decoder

wrap_boolean, addr_split (§2)
rd_value_rule       rd_selected − b_lw·load_read_value
store_value_rule    ram_write_value − m_ram·rs2

Its RANGE16 obligations are §2's three and the 16+16 pair on rd_selected, and the two copies are its whole semantics. rd_selected is range-checked although it copies a RAM word, because a RAM word need not be a word, advice's initial values being bound to nothing (public-values.md §6): the pair keeps every register value a word without reference to RAM (§5). No gate reads ram_read_value, the word a store overwrites; the memory argument alone pins it.

4 MEM_SUBWORD#

4.1 The splice#

A sub-word's position in its word lives only in

word = high·(w·p) + sub·p + low      p = 2^(8·offset), offset = 2·bit1 + bit0
                                     w, the access width: 2^8 if BYTE, 2^16 if HALF

p and its copower are degree-2 forms in the offset bits, written as gates rather than looked up:

p_rule        p − m_pc − 255·bit0 − 65535·bit1 − K·bit0·bit1      K = 2^24 − 2^16 − 2^8 + 1
pcopow_rule   p·pcopow − 2^31·m_pc                                 pcopow = 2^31/p
wph_rule      wph − 32768·p + 32640·BYTE·p                         wph = w·p/2
p_ram_rule    p_ram − m_ram·p

p_rule takes the four offsets to 1, 2^8, 2^16, 2^24, m_pc standing for the constant so the all-zero row satisfies it. The copower and w·p are stored halved so that 2^32 fits a u32 column, the gates reading them carrying the factor 2, as ShiftPowers' do (lookup.md §9). p_ram keeps store_rule degree 2. A table keyed by the offset would pin nothing addr_split does not, and add a key to bound (lookup.md §4).

4.2 Columns and gates#

W[9..21]   the decoded row; kind_lb … kind_sh
W[21..31]  wrap word_index word_index_hi bit0 bit1 p pcopow wph p_ram word
W[31..42]  high high_hi high_scaled high_scaled_hi sub sub_scaled sub_scaled_hi
           low low_hi low_scaled low_scaled_hi
W[42..51]  src_sub src_sub_scaled src_sub_scaled_hi src_high src_high_hi sign_in sign se rd_hi
W[51..55]  the four multiplicities

Its gates, beside the plumbing: wrap_boolean, bit0_boolean, bit1_boolean, addr_split, half_aligned, §4.1's four, and

word_rule            word − LOADK·load_read_value − STORE·ram_read_value
splice_rule          word − high_scaled − sub·p − low
high_scaled_rule     high_scaled − 2·high·wph                             = high·w·p
sub_scaled_rule      sub_scaled − 2^16·sub − (2^24 − 2^16)·BYTE·sub       = sub·2^32/w
low_scaled_rule      low_scaled − 2·low·pcopow                            = low·2^32/p
src_sub_rule         rs2 − src_sub − 2^16·src_high + 65280·BYTE·src_high
src_sub_scaled_rule  src_sub_scaled − 2^16·src_sub − (2^24 − 2^16)·BYTE·src_sub
store_rule           ram_write_value − m_ram·word − (src_sub − sub)·p_ram
sign_in_rule         sign_in − sub − 255·BYTE·sub                         = 2^8·sub or sub
se_rule              se − SIGNEXT·sign
rd_value_rule        rd_selected − LOADK·sub − (2^32 − 2^16)·se − 65280·BYTE·se

mem_subword::splice_gates(byte_bits) builds the twelve whose literals depend on the byte width — §4.1's first three and these but word_rule and se_rule — and the circuit takes it at BYTE_BITS = 8. Its RANGE16 obligations, all under m_pc, are §2's three, 16+16 pairs on high, high_scaled, sub_scaled, low, low_scaled, src_sub_scaled, src_high and rd_selected, and one obligation each on sub, src_sub and sign_in; its GENERIC lookup is sub_get_sign, (sign_in + SIGN_BASE, sign, 0).

4.3 Why it is sound#

§2 fixes the offset bits and half_aligned clears bit0 at halfword width, so p and w are the access's. Each part has a direct bound and a scaled one: high < 2^32 makes high·w·p an integer, sub_scaled < 2^32 is sub < w and low_scaled < 2^32 is low < p. So splice_rule holds over ℤ with one solution, the base-(p, w) digits of the word, and a word not below 2^32 has none. A scaled bound alone admits non-integers, its scale being a unit of Fr (lookup.md §11); constraints::lookup::check_copowers holds word_index_hi, high, sub, low and src_sub to their direct bounds, one obligation being exact for sub and src_sub, both below w ≤ 2^16. src_high's pair makes rs2 = src_sub + w·src_high integral, so src_sub is rs2 mod w: without it sb could store a byte unrelated to rs2.

A load writes rd = sub + (2^32 − w)·se: the sub-word, or at se = 1 its two's-complement extension (lb of 0x88 is 0xffffff88). sign_in is 2^8·sub for a byte and sub for a halfword, so its bit 15 is the sign at either width and one U16GetSign lookup serves both; its own obligation bounds the key into that sub-table (lookup.md §4). se is a one-hot sum times a table bit, boolean without a gate.

A store writes word + (src_sub − sub)·p = high_scaled + src_sub·p + low, a word with no appeal to memory: high_scaled is a multiple of w·p below 2^32 and w·p divides 2^32 (a halfword at offset 3 would make it 2^40; half_aligned excludes it), so high_scaled ≤ 2^32 − w·p and src_sub·p + low ≤ w·p − 1. That is why high_scaled keeps its own pair.

crates/checker/tests/mem_subword.rs checks the splice whole at a 4-bit word: for every word, admissible offset and width, splice_gates(1) and the bounds admit exactly one (high, sub, low).

5 The write-side induction#

A circuit may use a register operand as a word without bounding it. That rests on two facts:

  • Every register write is a word on its own row. Every execution family's rd_selected carries a 16+16 pair under m_pc, but ATOMICS', which is the old word or 0, the old word bounded by its comparison's pair under m_pc (§6). The frame writes (1 − z)·rd_selected (memory.md §2.4), registers start at 0 and a read returns the last write (memory.md §9), so every register read is a word, with no appeal to RAM.
  • Every RAM write of an execution family is a word: MEM_WORD writes rs2, a register value; MEM_SUBWORD bounds its merged word itself (§4.3); each ATOMICS arm is bounded (§6); a read-only query writes back what it read.

RAM's initial values are words — the image's, 0, the public input's — but advice's, which nothing bounds. No execution family relies on a RAM word being one: each bounds the value it uses, by MEM_WORD's rd pair, MEM_SUBWORD's splice or ATOMICS' comparison, so a row using a non-word is unprovable. The register half is what every carry needs: a + b − 2^32·wrap is a reduction only for words (memory.md §7), and addr_split's integer argument needs rs1 < 2^32.

6 ATOMICS#

One row is one read-modify-write: the ram query reads old and writes new at Δ = 3, beside rd (execution-trace.md §4), lr.w included, which writes its word back.

W[8..13]   decoded_next_pc decoded_rs1 decoded_rs2 decoded_rd decoded_mask    (no imm)
W[13..24]  kind_amoadd … kind_amomaxu
W[24..30]  word_index word_index_hi sum sum_hi add_wrap f_bitwise
W[30..42]  byte_a0..3 byte_b0..3 byte_and0..3       old's bytes, rs2's, their AND
W[42..50]  old_hi old_sign src_hi src_sign lt cmp_gap cmp_gap_hi lo
W[50..54]  the four multiplicities

With A = Σ_j 2^(8j)·byte_and_j inlined, its gates beside the plumbing are:

ram_value_rule    new − b_lr·old − (b_sc + b_amoswap)·rs2 − b_amoadd·sum − b_amoand·A
                    − b_amoor·(old + rs2 − A) − b_amoxor·(old + rs2 − 2A)
                    − (b_amomin + b_amominu)·lo − (b_amomax + b_amomaxu)·(old + rs2 − lo)
rd_value_rule     rd_selected − Σ_{k ≠ sc} b_k·old
add_rule          old + rs2 − sum − 2^32·add_wrap
f_bitwise_rule    f_bitwise − b_amoand − b_amoor − b_amoxor
old_bytes_rule    old − Σ_j 2^(8j)·byte_a_j          src_bytes_rule   rs2 − Σ_j 2^(8j)·byte_b_j
lo_rule           lo − rs2 − lt·(old − rs2)
addr_word (§2); add_wrap_boolean, f_bitwise_boolean; cmp_order, cmp_lt_boolean (below)

Each takes a kind's bit through its constants::extra_mask constant, from which the table's masks are built too, so a transposed arm would pass the decoder lookup. Under m_pc the family looks up the comparison's pairs on old, rs2 and cmp_gap and its two signs, §2's three and sum's pair; under f_bitwise, for each j, byte_a_j and 2^8·byte_a_j on RANGE16 and and_byte_j, (byte_a_j + AND_BASE, byte_b_j, byte_and_j), on GENERIC.

The comparison is constraints::gadgets::comparison (jump-branch-slt.md §3) with selector m_pc, lhs = old, rhs = rs2 and signed = [b_amomin, b_amomax]. The family's assemble asserts all four, nothing else in the artifact determining them: signed widened to amominu orders it signed, lhs and rhs swapped turn amomin into a max, and a selector narrowed to the min/max kinds drops old's bound on the other seven, and with it the bound on their rd write (§5). lo is the smaller under the ordering lt settles, and old + rs2 − lo the larger.

Why new is a word. old and rs2 are bounded by the comparison, sum by its own pair (add_rule is ungated: sum = (old + rs2) mod 2^32 on every live row), lo and the larger by being old and rs2. On a bitwise row byte_a_j's pair puts the key in the AND sub-table, whose row bounds byte_b_j and fixes byte_and_j = byte_a_j & byte_b_j; the byte rules are then the operands' decompositions, and A, old + rs2 − A, old + rs2 − 2A are AND, OR and XOR, carry-free byte by byte. Without its pair byte_a0 = 65,823 reads ShiftPowers' (65,824, 2^31, 1) (lookup.md §4); check_copowers takes the four keys under f_bitwise, which covers all three bitwise kinds: under b_amoand alone amoor and amoxor would read free byte_and.

sc.w always succeeds: it stores rs2 and writes 0 to rd, and the machine holds no reservation. The emulator does the same (execution-trace.md §10), and a row claiming failure, a nonzero rd or an unchanged word, is refused by rd_value_rule or ram_value_rule. This is a conformance deviation, not a soundness one: the proof is of what the program did on this machine. A guest may not rely on an sc.w failing where the ISA requires it to: with no valid reservation (no earlier lr.w, or one an earlier sc.w consumed) or at an address outside the reservation set. The lr.w/sc.w retry loop compiled code uses is unaffected, first-pass success being legal on any hart.

审计专区/委托

委托

规范原文docs/spec/delegation.md以 Markdown 查看

摘要

委托 ABI:客户程序(guest)如何通过一次 ecall 把一帧 RAM 字交给电路;委托类型的注册表;帧规则;借助内存多重集让每个请求恰好与一次调用配对的锚点;执行器一侧;对程序所调用电路族的静态声明;分片与时间窗口;高度及其取舍;客户程序一侧的调用方;以及委托集合的限制。

以下规范原文以英文维护;英文是本规范的标准语言。

The delegation ABI: how a guest hands a frame of RAM words to a circuit with an ecall, how each request pairs with exactly one invocation, how a program declares the families it calls, and how each family is sized. Frame layouts and circuits are delegation-circuits.md's, the recursion format's four families recursion.md's.

1 What a delegation family is#

A delegation family proves a function of guest memory too costly to run as instructions. It is invoked, never decoded: its number is a run-time value of a7, so it claims no pc and has no decoded table. A row is one invocation, which rides the cycle that requested it and owns no cycle (execution-trace.md §1); its accesses join the one memory multiset; it is in a VmConfig exactly when the image declares it (§7). Otherwise it is an ordinary family, an arm in constraints::family_circuit and a fill in prover::family_fill. A call is one row, the anchor's two leaves being a row's (§5); an operation wider than a row is several calls on one frame, chained through RAM (delegation-circuits.md §1, RAM glue).

2 The calling convention#

A call is an ecall (ecall-abi.md §1): a7 the number, a0 the frame base. It writes 0 to a0 and falls through (execution-trace.md §6); a recursion-format type writes a0 + 4·words instead (recursion.md §1.4).

An executor without a family's circuit answers -ENOSYS, on which a base-format shim's caller computes the same function in software, so an executor may implement any subset of the families; any other nonzero answer is fatal (ecall-abi.md §7).

3 The registry#

constants::delegation::TYPES, also program::DELEGATIONS, is one table of (family, number, anchor space, frame words), ascending by family, which the emulator dispatches on and constraints::add_sub builds its request gates from. The first BASE_TYPES = 6 rows are the base format's (recursion.md §1.2). Why each family has its height is §9's.

family id number anchor space frame words
KECCAK_F 9 0x0507 4 51
POSEIDON2 10 0x0500 5 24
FR_ARITH 11 0x0502 6 25
MOD_MUL 15 0x0504 7 25
SHA256_COMP 16 0x0508 8 25
EC_ADD 17 0x0506 9 97
FR_OP 19 0x0509 11 4
P2_FIELD 20 0x050A 12 5
FIELD_IO 21 0x050B 13 3
FQ_OP 22 0x050C 14 4

constraints::add_sub asserts at compile time that every number is in the precompile range and not EXIT, and that numbers and spaces are pairwise distinct, so an ecall row is the exit or a request of one type; a type costs that circuit a selector is_deleg_<f>, three gates and a term in five shared ones (add-sub.md §2, §4). A type's anchor space is the type: only its requests and invocations touch it, so the anchor's address is the frame base alone. A reserved range of RAM would need an argument that no guest access reaches it.

4 The frame#

A frame is words 32-bit words at the base a0 names, word j at base + 4j, read and written in place. Its base is word-aligned and it lies in RAM, RAM_ORIGIN ≤ base and base + 4·words ≤ 2^31: the executor refuses any other (Misaligned, OutOfBounds, the sum taken in u64) and the circuit has no witness for one (delegation-circuits.md §1, frame chain). So no frame lies in a public window or in advice.

An invocation reads and writes every word, unchanged ones written back, each a RAM query of the requesting cycle at slot constants::delegation::FRAME_DELTA = 0, ahead of the request's own queries (execution-trace.md §4, §7).

5 The anchor#

Requests and invocations pair one to one through the memory multiset, in the requested type's anchor space s. Otherwise N requests could close against one invocation, N − 1 calls going unexecuted, or an unrequested invocation could rewrite a frame.

5.1 The two sides#

                       reads                               writes
request (deleg)        T(s, a0, 0, 0)                      T(s, a0, 4c + 3, v)
invocation (anchor)    T(s, base, 4c + 3, anchor_value)    T(s, base, 0, 0)

The request is the deleg query of an ADD_SUB_LUI_AUIPC ecall row at cycle c (memory.md §2.1): deleg_mask_rule makes its mask m_pc·Σ_t is_deleg_t and deleg_addr_rule its address the a0 the row read. One query serves every type, so its space is deleg_space, an M column deleg_space_rule pins to Σ_t tag_t·is_deleg_t: a memory leaf may read no W column, and the selectors are W (memory.md §8).

The invocation's two leaves are the anchor read (delegation-circuits.md §1): it writes the answer, stamped 0 with value 0, and reads back the request's write at 4c + 3, c its cycle column. v and anchor_value are free and cancel only when equal; an honest prover writes 0 on both. Each answer starts a path one request long (memory.md §9).

5.2 The three request-side zeroings#

gate, under the request's mask forces
deleg_writes_no_register 0 written to a0, so the result is not the prover's choice
deleg_read_ts_zero the mirror read stamped 0
deleg_read_value_zero the mirror read's value 0

With deleg_addr_rule the last two make the mirror read the answer tuple, so every request consumes an answer of its own; without the timestamp, requests at one base chain, each consuming the previous one's write. The gates are the request row's, the same for every family, so the pairing needs nothing from a family's frame, and a call that changes no memory value has nothing else to expose it. The recursion format's deleg_a0_rule replaces the first (recursion.md §1.4).

5.3 Why the pairing is one to one#

In s the only tuples are the requests' and the invocations': no instruction reaches it, no window initializes it, nothing chains there (trace::AddressSpace::chains).

  1. A live row's 4c + 3 is not 0: the request's pc write and the invocation's frame writes at 4c lie on memory paths, whose timestamps are integers below 2^105 (memory.md §4.2).
  2. So the tuples stamped 0 are the requests' reads and the invocations' answers: as many invocations as requests, with the same multiset of bases.
  3. The rest are the requests' writes and the invocations' reads. No two requests share a cycle (memory.md §9), so each invocation's read is exactly one request's write: every invocation sits at its request's base and cycle, its frame accesses at that point of each word's history.

The trace-level check credits each anchor-space query with its invocation's tuples and sees none of this (execution-trace.md §9).

6 The executor's side#

For a registered number, Machine::ecall and Machine::delegate (crates/emulator/src/lib.rs) read a7 and a0; on the tracing paths refuse a family the VmConfig lacks (§7); read the frame, refusing §4's rules; compute the function natively (emulator::keccak_round, transcript::poseidon2_permute, Fr's operators, schoolbook products with long division, emulator::sha256_call) and write the whole frame back, a recursion family leaving it unchanged and working on field cells; stage the mirror query, reading and writing 0; and write a0 (constants::delegation::a0_after).

EmuError::DelegationFrame refuses a frame the circuit has no witness for, which the arithmetic would answer — long division is right for an unreduced operand too — leaving a proof that fails inside the GKR pass with nothing named: a KECCAK_F round word above 23, a SHA256_COMP group word above 15, an FR_ARITH code other than 1, 2, 3 or operand at or above p in memory form, a MOD_MUL or EC_ADD selector naming nothing or operand its row reads at or above the modulus, a POSEIDON2 lane at or above p. The recursion families' refusals are recursion.md §3–§6's.

The tracer records each invocation in its family's trace::DelegationTrace (execution-trace.md §11), which a shard reads as a trace::FrameSlice, ⌈invocations / height⌉ shards a family. The fill (prover::family_fill) commits the recorded words and derives the circuit's intermediates from those read. It never recomputes a written word: what is committed is what the execution did, and the circuit says that is the function. The circuit's side — frame chain, anchor read, gap decomposition, RAM glue — is delegation-circuits.md §1's.

7 Static detachment#

The instruction sweep cannot see a call, so each shim declares its family with a declaration record (constants::delegation):

MARKER_MAGIC = "APOGDEL1" (8 bytes) ‖ ecall number (u32 LE)          MARKER_BYTES = 12

guest_sdk emits one per family, a static whose #[link_section] is its own allocated section, .rodata.apogee.delegations.<family>, which link.ld's *(.rodata*) absorbs.

  • Its own section, because the linker's garbage collection keeps or drops whole input sections: records sharing one would be kept together, and reaching one shim would declare all.
  • Kept by reachability, not #[used], which keeps every record in every guest. Only the family's shim references its record, reading its own number from it through core::hint::black_box: a linked shim has a record, calls the number it declares, and the optimizer cannot fold the read away.
  • Statically: a call linked but never executed declares its family, which proves zero shards.

program::declared_delegations scans the image's file-backed bytes at every byte offset, a static's address being the linker's; a duplicate is one declaration, and a number no family answers is ProgramError::UnknownDelegation. Identity binds a record through the image column (program.md §8).

A called number whose family the VmConfig lacks is the fatal DelegationFamilyAbsent on the tracing paths; emulator::run, having no VmConfig, executes it. No proof covers it: the statement has no shard of that family, so the mirror read has no answer to consume.

8 Shards, time windows and the block#

A delegation shard's window is proof.md §8's, taken over its invocations' requesting cycles, so it lies inside the span of the ADD_SUB_LUI_AUIPC windows that made the requests. A delegation family is not cycle-owning, so the block holds its windows to nothing beyond start ≤ end ≤ 2^38; the anchor, not the window, places an invocation in time (§5.3).

9 Heights and channels#

A height sets how many calls a shard holds and limits no program. It is a parameter (ProgramParams::heights, defaulting to constants::family::DEFAULT_HEIGHTS, circuits.md §1) in identity's VM_CONFIG: a program's, not an execution's (program.md §7).

  • Floor: constraints::family_circuit returns None below the most variables any of the family's channel tables needs (constraints::lookup::table_vars, lookup.md §3).
  • Trade: a shard costs its height, not its occupancy (streaming.md §1), but its proof grows with the height only by a sumcheck round a variable in each gate list, a height changing no gate, only the number of halving lists. For a family with many calls the fatter shard is the smaller proof.
family channels floor unit of work calls a unit units a shard
KECCAK_F RANGE16, XOR8 2^16 keccak-f[1600] 24 10,922
POSEIDON2 none none width-3 permutation 1 256
FR_ARITH none none Fr add, multiply or inverse 1 256
MOD_MUL RANGE16 2^16 a·b mod m 1 65,536
SHA256_COMP RANGE16, XOR8 2^16 compression 16 16,384
EC_ADD RANGE16 2^16 complete point addition 3 21,845
  • POSEIDON2 and FR_ARITH take 2^8, the menu's smallest shard, where no table fits: every bound is a boolean decomposition. MOD_MUL and EC_ADD take their floor.
  • KECCAK_F and SHA256_COMP take 2^18, two variables above it: four times the calls for 2% more proof (a KECCAK_F shard's is 381,100 bytes, against 373,276 at 2^16). The price is memory: two 2^18 KECCAK_F shards in flight set the measured block's peak (streaming.md §1).
  • No base family carries TIMESTAMP, whose table needs 2^19 rows. FR_OP, P2_FIELD and FIELD_IO carry RANGE16, and FQ_OP TIMESTAMP and RANGE16, flooring it at 2^20.

10 Guest-side callers#

delegation reached from
KECCAK_F guest_sdk::keccak256; in guests/revm-block every alloy-primitives keccak, through its native-keccak hook native_keccak256
SHA256_COMP guest_sdk::sha256; revm-precompile's Crypto::sha256, the 0x02 precompile and the stateless guest's SSZ hashing
POSEIDON2 transcript::poseidon2_permute; guest_sdk::poseidon2_permute
FR_ARITH field::Fr's addition, Montgomery multiplication (*, square, pow, the conversions in from_u64, from_bytes, to_bytes) and nonzero inverse
MOD_MUL k256's FieldElement10x26::{mul, square}, Scalar::mul; ark-ff's MontBackend::{mul_assign, square_in_place} for BN254's two fields, as the product and then ·R⁻¹
EC_ADD guest_sdk::{ec_add, ec_mul}; k256's ProjectivePoint::{add, add_mixed, double}; revm-precompile's Crypto::{bn254_g1_add, bn254_g1_mul}
  • The shims are guest_sdk::recursion's but KECCAK_F's, which only keccak256 reaches (ecall-abi.md §7). Their frame types are #[repr(C, align(4))], so §4's alignment is the type's and not where the code generator put a local.
  • A multi-call operation's order is the caller's, and nothing refuses a wrong one: it computes something else. So each is one SDK function, keccak256's permutation, guest_sdk::recursion::sha256_comp and guest_sdk::recursion::ec_add_complete.
  • The transparent backends: field and transcript call the shims under cfg(target_arch = "riscv32"), through a target dependency on guest-sdk that a host build never resolves, not a cargo feature. Cargo refusing the cycle, guest-sdk cannot name Fr, so the shims take frames of bytes. The software path is each crate's own code, one branch below the call. A guest declares what its library calls reach: Fr arithmetic FR_ARITH, poseidon2_permute both.
  • FR_ARITH's frame carries Fr's memory form (primitives.md §1): canonical values would cost a Montgomery multiplication per value, more than the one the call replaces. POSEIDON2's carries canonical values, six conversions against the permutation's 240 multiplications.
  • The vendored crates, k256 0.13.4, ark-ff 0.6.0 and revm-precompile 43.0.2, are what a guest compiles through guests/Cargo.toml's [patch.crates-io], each route under the same cfg with upstream's code as its software path; the root workspace is unpatched. A MOD_MUL or EC_ADD operand must be below its modulus, so k256 first reduces its lazily reduced field elements. Changed files: guests/vendor/README.md.

11 Limits#

  • The EVM's MULMOD and MODEXP, BLS12-381 and every primitive outside §10's table run as instructions. No signature or pairing is delegated: secp256k1 recovery is k256 code over MOD_MUL and EC_ADD, a BN254 pairing ark-bn254 code over MOD_MUL.
  • A delegation is an operation's core: padding, a sponge or block loop, a scalar multiplication's ladder and a multi-call operation's order are guest code, proven as instructions.
  • A call's result is bound to memory alone: the frame after it is the function of the frame before.
  • This executor implements every family, so no proof here runs a base shim's software path.
  • Retired numbers are ecall-abi.md §4's.

审计专区/委托

委托电路

规范原文docs/spec/delegation-circuits.md以 Markdown 查看

摘要

六个基础格式的委托电路。在介绍它们共享的构造(帧链、锚点读取、差值分解、规范性链、单一编码规则、经由 XOR8 的字节运算,以及用于多次调用操作的 RAM 粘合)之后,逐一规定每个电路的帧、列、门与查找,说明它为何只接受其函数而不接受其他任何函数,并给出其成本与调用方:KECCAK_F、POSEIDON2、FR_ARITH、MOD_MUL、SHA256_COMP 和 EC_ADD。

以下规范原文以英文维护;英文是本规范的标准语言。

The circuits of the six delegation families the base format registers (recursion.md §1.2): KECCAK_F, POSEIDON2, FR_ARITH, MOD_MUL, SHA256_COMP, EC_ADD. A row is one invocation of a function of a frame of guest memory. For each circuit: its frame, columns, gates and lookups, and why it admits that function and no other. The call, the anchor's pairing, declaration and heights are delegation.md's.

1 Shared constructions#

Each circuit is constraints::delegation's frame over words frame words beside the family's function. None has a setup column; its only tables are its channels' virtual ones (lookup.md §3). live is the one mask, boolean by live_boolean and every lookup's selector. A padding row is all zero and satisfies every gate, a constant term riding live (gkr.md §4).

M[0..4]        cycle  live  base  anchor_value
M[4 + 4j ..]   word j: addr_j  read_ts_j  read_j  write_j        w{j}_addr … w{j}_write_value

Frame chain. Each word is read and written once at a pinned address, as two RAM leaves over memory.md §1's tuple T, from M columns because a leaf reads no W (memory.md §8):

read_w{j}        live·T(RAM, addr_j, read_ts_j, read_j) + 1 − live
write_w{j}       live·T(RAM, addr_j, 4·cycle, write_j) + 1 − live
addr_w{j}        live·(addr_j − base − 4j) = 0
base_aligned     live·(base − RAM_ORIGIN − 4·base_low) = 0          base_low  < 2^29
base_in_window   live·(2^31 − 4·words − base − base_room) = 0       base_room < 2^31

The bounds are delegation.md §4's frame rules, alignment a decomposition because 4 is a unit of Fr. A word the call leaves alone is held by writes_back_w{j}, write_j = read_j; every other written word is bounded below 2^32 by its circuit. A frame lies in RAM proper (delegation.md §4), which starts as the image's words or 0 and which every writer — an execution family (memory-ops.md §5), a frame, FIELD_IO's export (recursion.md §5) — leaves holding words, so a frame word a circuit reads is a word without a bound of its own.

Anchor read. Two leaves in the family's address space s (delegation.md §5) pair the row with its request: it writes the answer T(s, base, 0, 0) and reads T(s, base, 4·cycle + 3, anchor_value), what the request wrote back; anchor_value is free. That makes words + 1 leaves a side, padded with literal 1s to a power of two.

Gap decomposition. Each read precedes the row's write: gap_j = 4·cycle − 1 − read_ts_j is in [0, 2^38). TIMESTAMP would need a 2^20 shard (lookup.md §3), so the frame bounds its gaps, base_low and base_room itself, at the head of W:

  • bit form, at 2^8, where no table fits: 38 booleans a word, gap{j}_{i}, under gap_w{j}, live·(gap_j − Σ_i 2^i·g_i) = 0, and 29 and 31 for base_low and base_room: 38·words + 60 columns, each with its booleanity gate.
  • chunk form, at 2^16 and above, with no gate: a bound x ∈ [0, 2^{16q+r}), 0 < r < 16, is q committed chunks c_k of weight 2^{16(k+1)}, a RANGE16 obligation on each and on the remainder x − Σ_k 2^{16(k+1)}·c_k, and one on 2^{16−r}·c_top, which bounds only beside the chunk's direct one (lookup.md §11). A gap (r = 6) is gap{j}_c0 and gap{j}_c1; base_low and base_room (r = 13, 15) take base_low_hi and base_room_hi: 2·words + 4 columns and 4·words + 6 obligations.

The frame's gates are live_boolean, the addr_w{j}, base_aligned and base_in_window, words + 3, and in the bit form the gap_w{j} and each bit's booleanity besides.

Canonicity chain. A value X in limbs x_0 … x_7 < 2^32 is compared with a modulus m, limbs m_i < 2^32, through boolean borrows β_i and differences d_i ∈ [0, 2^32):

<v>_canonical{i}    x_i − m_i − β_{i−1} + 2^32·β_i − d_i = 0        i = 0 … 7, β_{−1} = 0

Every term is a small integer, so the eight sum over ℤ to X − m + 2^256·β_7 = D, 0 ≤ D < 2^256: β_7 = 1 exactly when X < m. Against Fr's p (§3, §4) the m_i are literals, x_i − p_i rides live and each d_i is 32 booleans; against a selected modulus (§5, §7) the m_i are columns, 0 on a padding row, and each d_i has a 32-bit bound (memory.md §7).

Gated conclusion. The chain's last gate, <v>_below_modulus, is live − β_7 = 0 where every live row reads X; where only rows with enable = 1 read it, it is the gated conclusion enable·(1 − β_7) = 0. β_7 = enable would demand X ≥ m wherever enable = 0, so a row holding a reduced X it does not read would have no witness.

One-code rule. A frame word naming one of k cases is decoded into boolean selectors s_c by word − Σ_c code_c·s_c = 0 and Σ_c s_c − live = 0. The second is not implied: a code 0 has no selector set and a code that is a sum of two has two (1 + 2 = 3), mixing cases. With both, the word and any column pinned to Σ_c lit_c·s_c are one entry of a table of literals, selected and bounded by a degree-1 gate.

Byte operations. Where the unit is the byte (§2, §6), each Boolean operation is one XOR8 obligation (e_0, e_1, e_2), e_2 = e_0 ^ e_1 with all three bytes (lookup.md §3): e_1 and e_2 columns, e_0 any literal-weighted form with a constant (lookup.md §5). The rest is linear in the results: a & b = (a + b − (a ^ b))/2, ¬a & b = (b − a + (a ^ b))/2; against a literal k, v & k = (v + k − (v ^ k))/2 splits a byte at any bit, so a rotation or shift of a word held as bytes is a literal-weighted form over its bytes and their masked copies; and (0, c, c) bounds c to a byte. On true bytes and true XORs each form is exact over ℤ, so its value is the integer it denotes.

RAM glue. An operation too wide for a row is several invocations on one frame, a frame word naming the step (§2, §6, §7). Each proves its step on the frame as it finds it: its reads lie on each word's one history (memory.md §9), so it reads the previous step's writes unless the guest wrote there between. No gate joins two rows, and a shard boundary may fall between them. That every step runs, in order, is the calling code's, which the execution families prove.

2 KECCAK_F#

One invocation is one round of keccak-f[1600]; a permutation is 24 on one frame, the sponge and padding being guest code. The circuit, constraints::keccak, is flat, every gate in gate list 0, and its unit is the byte (§1): no column is a bit but live and the 24 round selectors.

2.1 Frame and columns#

51 words (constants::keccak; M[0..208]), the state in SHA-3 byte order: lane A[x][y], i = 5y + x, at words 1 + 2i (low half) and 2 + 2i. A[i][b] is its byte b; lane coordinates are mod 5.

word read written
0 the round r ∈ [0, 24) yes unchanged
1–50 the state yes the round's output
W name
0..106 the frame's chunks (§1)
106..130 round_sel{r} s_r, one a round
130..134 rc_b{b} rc_t, byte b_t = 0, 1, 3, 7 of the round's constant
134..334 state_in_l{i}_b{b} A
334..494 parity_x{x}_b{b}_s{s} column x's lanes XORed in four steps, the last C[x]
494..574 c_mask_…, theta_d_… C ^ 0x80; D
574..774 theta_a_… A′ = A ^ D
774..950 rho_mask_… A′ ^ mask on the 22 lanes not rotated by whole bytes
950..1150 rho_out_… B, after ρ and π
1150..1550 chi_and_…, chi_out_… B1 ^ B2; χ's output
1550..1554 iota_out_b{b} lane 0's bytes b_t after ι
1554..1556 the multiplicities

2.2 Gates and obligations#

385 gates; O is chi_out, but iota_out at lane 0's bytes b_t; r_xy = ROTATIONS[y][x].

gate count expression
the frame's (§1) 54
round{r}_boolean 24 s_r − s_r²
round_rule 1 read_0 − Σ_r r·s_r
one_round_a_live_row 1 Σ_r s_r − live
rc{t}_rule 4 rc_t − Σ_r s_r·(byte b_t of ROUND_CONSTANTS[r])
writes_back_w0 1 write_0 − read_0
input_w{j}, j = 1 + 2i + h 50 read_j − Σ_{k<4} 2^{8k}·A[i][4h + k]
output_w{j} 50 write_j − Σ_{k<4} 2^{8k}·O[i][4h + k]
rho_pi_l{i}_b{j} 200 B[y][2x + 3y][j] − rot_j(A′[x][y], r_xy), its constant times live

A rotation by 8q + s is linear in a lane's bytes v and their copies μ = v ^ mask (§1), mask = 256 − 2^{8−s} being the top s bits; with u = j − q and w = u − 1 mod 8,

rot_j(v) = 2^{s−1}·(v_u + μ_u) + 2^{s−9}·(v_w − μ_w) + mask·(2^{s−9} − 2^{s−1})      s > 0
rot_j(v) = v_u                                                                    s = 0

v_u's low bits moved up and v_w's top bits down, (v + mask − μ)/2 being v & mask.

The obligations are the frame's 210 on RANGE16 (§1) and 1,020 on XOR8, one a byte:

step count obligation e_2 = e_0 ^ e_1
θ 160 parity_s = parity_{s−1} ^ A[x][s + 1], s < 4, parity_{−1} = A[x][0]
θ 40 c_mask = 0x80 ^ C[x]
θ 40 D[x] = rot(C[x + 1], 1) ^ C[x − 1], c_mask as μ
θ 200 A′[x][y] = D[x] ^ A[x][y]
ρ 176 rho_mask = mask ^ A′
χ 200 chi_and = B1 ^ B2, Bk = B[x + k][y]
χ 200 chi_out = ((B2 − B1 + chi_and)/2) ^ B[x][y]
ι 4 iota_out_t = rc_t ^ chi_out[0][b_t]

2.3 Why it is sound#

Every byte column is an entry of some obligation, so all are bytes, each obligation is the operation it names and each form the integer it denotes (§1): rot because μ is the true XOR, and (B2 − B1 + chi_and)/2 is ¬B1 & B2. The channel alone fixes parity, c_mask, theta_d and chi_and. B is committed, and pinned by rho_pi, because χ reads every lane at an entry only a column may fill.

input_w and output_w are each a word's byte decomposition and its 32-bit bound, so no state word has a range obligation; without output_w a row could write any state. Both are ungated and degree 1, a padding row's words and bytes being 0, which pins its state bytes to 0; a cell that only live-gated gates and obligations reach is free on a padding row, to no effect.

one_round_a_live_row is the one-code rule (§1) over codes 0 … 23: without it a live row could set no selector, claiming round 0, or two spelling a third, and ι would add no constant or a wrong one. The constant is a table of literals the selectors pick (rc{t}_rule), with no lookup or commitment. So a live row writes round read_0 of the state it read.

A permutation is RAM glue (§1) over guest_sdk::keccak256's loop, which stores r = 0 … 23 in word 0 before each call. crates/checker/tests/keccak.rs holds every gate and obligation over 24 such rows to a round written apart in u64 and, through emulator::keccak_round, to tiny-keccak.

2.4 Cost and callers#

1,764 committed columns and, at 2^18, 5,490 inner ones in 29 gate lists, 11 row-wise and 18 halving, all the two memory trees' and the two fraction trees'. The 1,020 obligations and the table's fraction fill 1,021 of the XOR8 tree's 1,024 leaves (lookup.md §6); four more would double it, 4,100 more inner columns. So ι is four obligations: a round constant is zero outside bytes 0, 1, 3 and 7 (constants::keccak::IOTA_BYTES_ARE_THE_ONLY_ONES, checked at compile time).

A 2^18 shard (delegation.md §9) holds 10,922 permutations; its proof is 381,100 bytes (proof.md §9), 34.9 a permutation, and its forward pass 45.2 GB of inner layers (streaming.md §1), which is what sets a block's peak. Caller: guest_sdk::keccak256 (delegation.md §10).

3 POSEIDON2#

One invocation is one transcript::poseidon2_permute (transcript.md §1). The circuit, constraints::poseidon2, is at 2^8 with no lookup, bounding in bits (§1), and is the one delegation circuit that computes above gate list 0.

3.1 Frame and columns#

24 words (constants::poseidon2):

words read written
8l … 8l + 7 lane l, l < 3 yes the permuted lane

A lane is its value's canonical encoding (Fr::to_bytes), not §4's Montgomery form, so the circuit is the permutation itself; the caller's six conversions are small beside the 240 S-box multiplications a call replaces.

M[0..100] and W[0..972] are the frame (§1). W[972..4092] holds 520 booleans for each of six values, the lanes read (in0 … in2) then written (out0 … out2): 256 word bits, then the canonicity chain's (§1) 256 difference bits and 8 borrows.

3.2 Gates#

Gate list 0 holds 4,245: the frame's 51 (§1), a booleanity gate on each W column, and 17 a value, over its read or written words: eight <v>_word{k}, word_k − Σ_t 2^t·bit_{k,t}, and its canonicity chain against p (§1), eight <v>_canonical{i} and <v>_below_modulus, live − β_7.

The permutation is computed, not witnessed: three gate lists a round r, S-boxing every lane of a full round and lane 0 of a partial one, whose other lanes the first two lists copy:

list 3r         q_i = (x_i + c_{r,i})²       t_i = x_i + c_{r,i}
list 3r + 1     q2_i = q_i²                  t_i copied
list 3r + 2     x′ = M_r·v                   v_i = q2_i·t_i, or x_i on a copied lane

M_r is E or I and the constants are literals of the gates; round 0's x is E applied to in_l = Σ_k 2^{32k}·read_{8l+k}. A committed column is read by gate list 0 only (gkr.md §2), so live and out_l = Σ_k 2^{32k}·write_{8l+k} are carried up to gate list 192, which holds the last three gates,

out_lane{l}     live·(x_l − out_l) = 0          x the state after round 63

gated because a padding row computes the permutation of the zero state, which is not zero.

3.3 Why it is sound#

A layer's column is forced by the gate that writes it, so x is the permutation of (in_0, in_1, in_2) as field elements. The word gates make each in_l and out_l the integer its words spell, and the chains put it below p: a lane at or above p has no witness, and out_lane fixes all 24 written words, where without the chains on out a row could write x_l + p. The forward pass accepts Plonky3's permutation vectors (crates/checker/tests/poseidon2.rs).

3.4 Cost and callers#

4,192 committed columns and 2,020 inner ones in 201 gate lists, 193 row-wise and 8 halving: 736 the rounds' (15 a full round, 11 a partial one), 768 the four carried columns', the rest the memory trees'. A 2^8 shard holds 256 permutations; its proof is 664,780 bytes, 2,597 a permutation. Caller: transcript::poseidon2_permute on the guest target (delegation.md §10).

4 FR_ARITH#

One invocation is one Fr addition, multiplication or inversion. The circuit, constraints::fr_arith, is flat, at 2^8 with no lookup, bounding in bits (§1).

4.1 Frame and encoding#

25 words (constants::fr_arith):

words read written
0 the code: 1 add, 2 multiply, 3 inverse (OPS) yes unchanged
1–8, 9–16 a, b yes unchanged
17–24 out yes, unconstrained the result

A value is Fr's in-memory form, Fr::to_memory_bytes: the canonical encoding of the Montgomery representative x·R, R = 2^256 mod p. The circuit computes what Fr's own operators compute on representatives,

add         out = a + b
multiply    out = a·b·R⁻¹
inverse     out = R²·a⁻¹, and 0 at a = 0

because a frame of values would cost the guest a Montgomery conversion per value, more than the multiplication a call replaces. Fr::inverse answers None at 0 itself and makes no call.

4.2 Columns and gates#

M[0..104] and W[0..1010] are the frame (§1); W[1010..2570] 520 booleans for each of a, b (read) and out (written), as §3.1; W[2570..2573] the selectors f_add, f_mul, f_inv (selector1 … selector3); W[2573..2576] the field columns prod, inv and z (is_zero). The 2,701 gates: the frame's 53 (§1); 2,573 booleanity gates, on every bit and selector; §3.2's 17 per value; writes_back_w{j} for j < 17; and, a, b and out being the forms Σ_k 2^{32k}·word_k,

gate expression
opcode_rule read_0 − f_add − 2·f_mul − 3·f_inv
one_op_a_live_row f_add + f_mul + f_inv − live
prod_rule prod − a·b
inv_is_an_inverse a·inv + z − f_inv
is_zero_at_nonzero a·z
inverse_of_zero_is_zero z·inv
out_rule out − f_add·(a + b) − R⁻¹·f_mul·prod − R²·f_inv·inv

R⁻¹ and R² are literals derived from constants::FR_R.

4.3 Why it is sound#

As in §3.3, each value is the integer below p its words spell, so out_rule fixes the eight written words. prod is committed, under an ungated gate, because a selector times a·b is degree 3. On an inverse row a ≠ 0 forces z = 0 and inv = a⁻¹, and a = 0 forces z = 1 and inv = 0; without is_zero_at_nonzero, z = 1 and inv = 0 pass at any a, and without inverse_of_zero_is_zero, inv is free at a = 0. one_op_a_live_row is the one-code rule (§1): 1 + 2 = 3, so opcode_rule alone lets f_add and f_mul answer an inversion with a + b + a·b·R⁻¹.

4.4 Cost and callers#

2,680 committed columns and 142 inner ones, all the memory trees', in 14 gate lists, 6 row-wise and 8 halving. A 2^8 shard holds 256 operations; its proof is 266,292 bytes, 1,040 an operation. Caller: field's addition, Montgomery multiplication and inverse on the guest target (delegation.md §10).

5 MOD_MUL#

One invocation is one multiplication out = a·b mod m of 256-bit integers, m one of four fixed primes a frame word selects. The circuit is constraints::mod_mul.

5.1 The frame and the columns#

25 words (constants::mod_mul). A value is a plain residue, not a Montgomery one, in eight 32-bit limbs, least significant first.

words
0 the selector: 1 secp256k1's base field p, 2 its order n, 3 BN254's base field q, 4 its scalar field r (CODES, MODULI) read, written back
1–8, 9–16 a, b, each below the selected modulus read, written back
17–24 out written; the value read is ignored

Codes start at 1, so a zero word names no field. The EVM's MULMOD, whose modulus is arbitrary, is not this call and runs as guest code.

M[0..104], W[0..54]   the frame (§1)
W[54..58]     selector1 … selector4        s_c, one a code
W[58..66]     m_limb{k}                    m_k, the selected modulus
W[66..162]    <v>{k}_hi, <v>_diff{i}, <v>_diff{i}_hi, <v>_borrow{i}     for v = a, b, out
W[162..178]   q_limb{k}, q_limb{k}_hi      the quotient and its halfwords
W[178..220]   carry{k}, carry{k}_c0, carry{k}_c1      c_k + 2^36 for k < 14, and two chunks
W[220]        range16_multiplicity

5.2 Gates and lookups#

read_j and write_j are word j's two values (§1), a_i and b_i read limbs, out_i written ones, and c_k = carry{k} − 2^36·live. Each expression is held to 0:

gate count expression
the frame's (§1) 28
writes_back_w{j}, j < 17 17 write_j − read_j
selector{c}_boolean; selector_rule; one_modulus_a_live_row 6 s_c − s_c²; read_0 − Σ_c c·s_c; Σ_c s_c − live
m_limb{k}_rule 8 m_k − Σ_c s_c·MODULI[c][k]
<v>_borrow{i}_boolean, <v>_canonical{i}, <v>_below_modulus 51 v's canonicity chain (§1) against the m_k columns, concluding live − β_7
limb{k}, k < 15 15 Σ_{i+j=k} (a_i·b_j − q_i·m_j) − out_k + c_{k−1} − 2^32·c_k; out_k past limb 7, c_{−1} and c_14 are 0

274 RANGE16 obligations, all under live: the frame's 106 (§1); a pair — the 32-bit bound of memory.md §7, two obligations over a committed high halfword — on every limb of a, b, out and q and on every diff_i (112); and each carry{k} in [0, 2^37), by two chunks and four obligations as a gap (§1) (56).

5.3 Why it is sound#

The field. By the one-code rule (§1), m is the modulus word 0 names. Codes add (1 + 3 = 4), so without one_modulus_a_live_row selectors 1 and 3 answer a request for r modulo p + q; with it each m_k is one literal, which is all that keeps m's limbs, bound by no obligation, below 2^32.

The product. Every limb of a, b, out, q and m being below 2^32, a position's products sum below 2^67 a side and the carries lie in [−2^36, 2^36), so no term nears Fr's modulus: the fifteen limb{k} equations hold over ℤ and, weighted by 2^{32k}, sum to a·b = q·m + out, position 14 having no carry out.

The reduction is out_below_modulus: without it (q − 1, out + m) satisfies every other relation wherever out + m fits eight limbs.

The operand bounds make the relation total, not out right: with a, b < m, q = (a·b − out)/m < m, so every frame the circuit admits has an eight-limb quotient. A caller holding a lazily reduced value therefore owes a reduction below m, not below 2^256. The emulator's mod_mul_frame refuses the frames no proof could cover, a selector that is no code and an operand at or above m (EmuError::DelegationFrame).

5.4 Cost and callers#

Shape: circuits.md §1. A 2^16 shard (delegation.md §9) is 65,536 multiplications at 2.1 proof bytes each; its forward pass, 2,180 row-wise inner columns × 2^16 rows × 32 bytes, is 4.6 GB.

guest_sdk::recursion::mod_mul makes the call over a ModMulFrame. The vendored k256 reaches it from its field and scalar multiplies (codes 1, 2), the vendored ark-ff from BN254's Montgomery multiply (codes 3, 4): delegation.md §10.

6 SHA256_COMP#

One invocation is four rounds of SHA-256's compression function and four words of its message schedule; a compression is sixteen invocations on one frame, joined by RAM glue (§1). Padding, the block loop and the final addition of the chaining value are the caller's. The circuit is constraints::sha256.

6.1 The frame#

25 words (constants::sha256):

words read written
0 the round group r < 16 unchanged
1–8 the working variables a … h a … h four rounds on
9–24 the schedule window W_{4r} … W_{4r+15} moved down four words, W_{4r+16} … W_{4r+19} last

Call 0 reads the chaining value as a … h and the block, decoded big-endian, as the window. Over a row the state is two sequences: A_0 … A_{−3} are a … d as read, A_4 … A_1 are a … d as written, and E_j is the same over e … h, so each of the sixteen is a frame column. For k < 4 and m < 4, every sum mod 2^32:

T1          = E_{k−3} + Σ1(E_k) + Ch(E_k, E_{k−1}, E_{k−2}) + K_{4r+k} + W_{4r+k}
A_{k+1}     = T1 + Σ0(A_k) + Maj(A_k, A_{k−1}, A_{k−2})
E_{k+1}     = A_{k−3} + T1
W_{4r+16+m} = σ1(W_{4r+14+m}) + W_{4r+9+m} + σ0(W_{4r+1+m}) + W_{4r+m}

Call r + 4's rounds read the words call r derives, so the guest computes no schedule; calls 12–15 derive words no round reads.

6.2 Bytes and their obligations#

No column is a bit but live and the group selectors g_r. A word that enters a Boolean operation has four byte columns, and each such operation is one XOR8 obligation (x, y, x ^ y) a byte (lookup.md §3), of which position 0 alone may be a literal-weighted form (lookup.md §5).

  • A rotation is linear. With μ = v ^ (2^s − 1) committed, a byte v splits into lo = (v + 2^s − 1 − μ)/2 and hi = (v − lo)/2^s. Byte j of ROTR_{8t+s}(V) is hi(v_{j+t}) + 2^{8−s}·lo(v_{j+t+1}), indices mod 4, and for s < 8 the word ROTR_s(V) is (V − lo(v_0))/2^s + 2^{32−s}·lo(v_0).
  • The big sigmas nest, Σ0(a) = ROTR2(a ^ ROTR11(a ^ ROTR9(a))) and Σ1(e) = ROTR6(e ^ ROTR5(e ^ ROTR14(e))), so each XOR has one rotated operand and the outer rotation is a word's: 17 obligations a sigma.
  • The small sigmas end in a shift, σ0(x) = ROTR7(x ^ ROTR11(x)) ^ SHR3(x) and σ1(x) = ROTR17(x ^ ROTR2(x)) ^ SHR10(x), so their outer XOR has two derived operands: the shifted bytes are committed and pinned by gates. 16 and 15 obligations, SHR10's top byte being 0.
  • Ch and Maj are linear in XORs, Ch(e, f, g) = (f + g − (e ^ f) + (e ^ g))/2 and Maj(a, b, c) = (a + b + c − (a ^ b ^ c))/2: 8 obligations each.
  • A carry c is a byte by (0, c, c).

That is 52 obligations a round and 32 a schedule word, 336 on XOR8. RANGE16 carries 114: the frame's 106 (§1) and a pair (§5.2) on each written word without bytes, A_4, E_4, W_{4r+18} and W_{4r+19}.

M[0..104], W[0..54]   the frame (§1)
W[54..70]     group{r}                     g_r, one a group
W[70..118]    a{j}_b{b}, e{j}_b{b}         bytes of A_{−2} … A_3 and E_{−2} … E_3 (j = m2 … 3)
W[118..150]   w{i}_b{b}, n{m}_b{b}         bytes of window words 1–4, 14, 15, derived words 0, 1
W[150..358]   r{k}_…                       52 a round: the big sigmas' masks and XORs (34),
                                           e^f, e^g, a^b, c^a^b (16), two carries
W[358..514]   s{m}_…                       39 a schedule word: the small sigmas' masks, XORs
                                           and shifted bytes (38), a carry
W[514..518]   w{j}_written_hi              high halfwords of A_4, E_4, W_{4r+18}, W_{4r+19}
W[518..520]   range16_multiplicity, xor8_multiplicity

6.3 Gates#

All of degree 1 but the frame's and the booleans:

gate count expression
the frame's (§1) 28
group{r}_boolean; group_rule; one_group_a_live_row 18 g_r − g_r²; read_0 − Σ_r r·g_r; Σ_r g_r − live
writes_back_w0 1 write_0 − read_0
a{j}_decode, a{j}_encode, e{j}_…, w{i}_decode, n{m}_encode 20 a word − Σ_b 2^{8b}·byte_b, for every word with bytes
w{i}_shift, i < 12 12 write_{9+i} − read_{13+i}
r{k}_a, r{k}_e 8 §6.1's A_{k+1} and E_{k+1}, as word + 2^32·carry − sum
s{m}_sum 4 §6.1's W_{4r+16+m}, likewise
s{m}_shr3_b{b}, s{m}_shr10_b{b} 28 a committed shifted byte − its form

K_{4r+k} is the form Σ_r K_{4r+k}·g_r.

6.4 Why it is sound#

A sum's operands are words: those with bytes by their obligations, and d, h, W_{4r} and W_{4r+9} … W_{4r+12}, which only sums read, because the frame lies in [RAM_ORIGIN, 2^31) (§1), below advice, where every initial value and every write is a word (memory-ops.md §5; §1 for these circuits). Its carry being a byte, a sum gate holds over ℤ, and its left word, bounded by its bytes or its pair, is the sum mod 2^32. Without the carry's range any word satisfies the gate; without the pair on A_4, a carry of 0 writes the unreduced sum. Every word a row writes is therefore a word: a copy, one with bytes, or one of the four with a pair.

Group 0's code being 0, group_rule alone admits a live row with no selector or with g_0 beside another; one_group_a_live_row refuses those and two selectors spelling a third group, each a round under a wrong constant. Sixteen rows are one compression by RAM glue (§1) and by guest_sdk::recursion::sha256_comp, which stores r = 0 … 15 in word 0 before each call; the emulator's sha256_frame refuses a group word of 16 or more. crates/checker/tests/sha256.rs evaluates every gate and obligation over sixteen chained rows built from FIPS 180-4 in u32 arithmetic and holds their output to the standard's abc digest.

6.5 Cost and callers#

Shape: circuits.md §1. A 2^18 shard (delegation.md §9) holds 16,384 compressions at 11.6 proof bytes each; its forward pass, 2,694 row-wise inner columns × 2^18 × 32 bytes, is 22.6 GB.

guest_sdk::sha256 pads, walks the blocks, and for each runs sha256_comp's sixteen calls and adds the result to the chaining value. The vendored revm-precompile routes Crypto::sha256 to it: precompile 0x02, and the stateless guest's SSZ hashing (delegation.md §10).

7 EC_ADD#

One invocation is a third of one complete point addition P1 + P2 on secp256k1 or BN254 G1, in homogeneous projective coordinates (x = X/Z, y = Y/Z). An addition is three invocations on one frame in group order, joined by RAM glue (§1); scalar multiplication is guest code over it. The circuit is constraints::ec_add.

7.1 The formula#

Renes–Costello–Batina 2015, Algorithm 7, for y² = x³ + b, with b3 = 3b: 21 and 9 (constants::ec_add::CURVE_B3).

group 0   xx = X1·X2            yy = Y1·Y2            zz = Z1·Z2
group 1   m4 = (X1+Y1)(X2+Y2)   m5 = (Y1+Z1)(Y2+Z2)   m6 = (X1+Z1)(X2+Z2)
group 2   X3 = xy·ym − byz3·xz  Y3 = yp·ym + bxx9·xz  Z3 = yz·yp + xx3·xy

xy = m4 − xx − yy   yz = m5 − yy − zz   xz = m6 − xx − zz   ym = yy − b3·zz
yp = yy + b3·zz     byz3 = b3·yz        xx3 = 3·xx          bxx9 = 3·b3·xx

Both groups have prime order, so the formula is complete: a doubling, P + (−P), the identity (0 : 1 : 0) and any Z take no special case, in the guest or in a row, and nothing is inverted. The formula is the caller's: the vendored k256's ProjectivePoint addition is this algorithm on these coordinates, so the delegated and the software path return the same representative.

The twelve multiplications are nine reductions, each of X3, Y3, Z3 being two products under one quotient. A row holds three, not nine, because a shard's memory grows with its row's width and its height cannot fall below 2^16 (§7.5).

7.2 The frame and the columns#

97 words (constants::ec_add), a value as in §5.1:

words read by group written by group
0 the selector, one of CODES: 1–3 secp256k1's groups 0–2, 4–6 BN254 G1's all none
1–24 X1, Y1, Z1 0, 1 2, as X3, Y3, Z3
25–48 X2, Y2, Z2 0, 1 none
49–72 xx, yy, zz 2 0
73–96 m4, m5, m6 2 1

A row has three slots, each one reduction of one shape:

A·B + C·D + 1024·m² = q·m + out,     out < m

Group 0's (A, B) are (X1, X2), (Y1, Y2), (Z1, Z2) and group 1's the three pairs of sums, both with C = D = 0. Group 2's (A, B, C, D) are (xy, ym, byz3, −xz), (yp, ym, bxx9, xz) and (yz, yp, xx3, xy): a product's sign rides its operand.

M[0..392], W[0..198]   the frame (§1)
W[198..295]   word{j}_hi                   the high halfword of every word's read value
W[295..301]   selector{c}                  s_c, one a code
W[301..310]   m_limb{k}, b3                the curve's modulus and 3b
W[310..334]   bzz3_{k}, byz3_{k}, bxx9_{k}     b3·zz_k, b3·(m5_k − yy_k − zz_k), 3·b3·xx_k
W[334..622]   <v>_diff{i}, <v>_diff{i}_hi, <v>_borrow{i}     chains of the twelve values x1 … m6
W[622..1027]  slot{r}_…, out{r}_…, q{r}_…, carry{r}_…    135 a slot: four operands (32), out and
              its halfwords (16), a nine-limb q and its halfwords (18), 15 carries c_k + 2^46
              with two chunks each (45), out's chain (24)
W[1027]       range16_multiplicity

7.3 Gates and lookups#

G_g is the sum of the two selectors naming group g, and c_k = carry − 2^46·live:

gate count expression
the frame's (§1) 100
selector{c}_boolean, selector_rule, one_code_a_live_row 8 §5.2's, over six codes
m_limb{k}_rule, b3_rule 9 the column − Σ_c s_c·(its literal for code c's curve)
bzz3_{k}_rule, byz3_{k}_rule, bxx9_{k}_rule 24 the column − its product above
<v>_borrow{i}_boolean, <v>_canonical{i} 240 canonicity chains (§1) of the twelve values and the three outs, against m_k
<v>_below_modulus 15 e·(1 − β_7): e is G_0 + G_1 for x1 … z2, G_2 for xx … m6, live for an out
operand{r}_{o}_{k}_rule 96 an operand limb − Σ_g G_g·(group g's expression at that limb)
slot{r}_limb{k}, k < 16 48 Σ_{i+j=k} (A_i·B_j + C_i·D_j + 1024·m_i·m_j − q_i·m_j) − out_k + c_{k−1} − 2^32·c_k, q having nine limbs
writes_back_w{j} 97 write_j − read_j − G_g·(out_k − read_j), for the group g and slot limb k that write word j, if any

1,110 RANGE16 obligations under live: the frame's 394 (§1); a pair (§5.2) on every word's read value (194), every chain difference (240), every out limb (48) and every q limb (54); and each carry in [0, 2^47), four apiece (180).

7.4 Why it is sound#

The curve and the group are §5.3's argument over six codes: codes add (1 + 3 = 4, 2 + 4 = 6), and the one-code rule is also all that keeps m's limbs and b3 literals.

The operands. An operand limb is its group's expression: a combination, with coefficients of at most 3, of frame limbs below 2^32 and of their products with b3. Its pin is therefore its bound, below 2^38 in magnitude, and it carries no obligation. It is a committed column because the expression depends on the group, and a selector times a product of limbs would be degree 3; b3 enters through the three helper columns for the same reason.

The identity. As in §5.3: a position stays below 2^78, the carries in [−2^46, 2^46) (CARRY_OFFSET_BITS), the sixteen equations hold over ℤ and close because position 15 has no carry out, and out < m makes out the residue of A·B + C·D. The 1024·m² (OFFSET_MULTIPLE) keeps the left side non-negative, a quotient's limbs being unsigned: it is lowest in group 2's Y3, at −673·m² by its operands' ceilings 22m, 22m, 63m and 3m. One literal serves every slot, a group-dependent offset being degree 3, and q < 1697·m fits nine limbs. ec_add::artifact checks both constants against the ceilings when it builds the circuit.

Canonicity. out < m is the reduction. A read value below m is what the ceilings assume, and so what gives every admitted frame a quotient; the emulator's ec_add_frame refuses a frame whose group reads a value at or above m, or whose selector is no code. Each such conclusion is gated (§1) on the groups that read the value: every lane is below m on every row a guest builds (EcAddFrame::of zeroes the intermediates), so β_7 = e would leave no row a witness.

What the guest owns. Each third is proved; their order is the guest's. guest_sdk::recursion::ec_add_complete writes the three codes in turn, and groups out of order are not refused but compute another point from stale lanes. Nor is a point held to its curve: what is proved is the formula's arithmetic.

7.5 Cost and callers#

Every bound is a RANGE16 obligation, so the family sits at 2^16, the channel's floor (lookup.md §3), and no other height is practical: as bits the 97 gaps alone would be 3,686 columns, and at 2^18 the forward pass below would be 73 GB. Shape: circuits.md §1. A shard holds 21,845 additions at 19.9 proof bytes each; its forward pass, 8,708 row-wise inner columns × 2^16 × 32 bytes, is 18.3 GB.

guest_sdk::ec_add makes the three calls over an EcAddFrame, and guest_sdk::ec_mul is double-and-add over it. The vendored k256 routes ProjectivePoint's addition, mixed addition and doubling here, and the vendored revm-precompile routes Crypto::bn254_g1_add and Crypto::bn254_g1_mul, precompiles 0x06 and 0x07: delegation.md §10.

审计专区/证明流水线

流式证明者

规范原文docs/spec/streaming.md以 Markdown 查看

摘要

一次执行如何在从不完整持有其执行轨迹的情况下变成块证明。它给出证明者的实测成本,规定两遍处理以及承诺为何必须先于证明、一次执行之后保留什么、分片计划以及执行器运行时如何切分分片、工作线程流水线及其保证(包括证明不依赖于调度),以及篡改测试套件所依赖的归档路径。

以下规范原文以英文维护;英文是本规范的标准语言。

How one execution becomes a BlockProof: the guest runs twice, and a fixed number of workers commit, then prove, its shards as the executor fills them. The block is proof.md's, and its bytes do not depend on the schedule; this page fixes when each column exists, and so what a proof costs.

1 The prover, and what it costs#

prover::prove_block_streaming(setup, io, max_in_flight) proves every block: host::prove wraps it over the ProverSetup that host::setup builds from an ELF, bench prove drives it (tools.md §1), and recursion nodes are proved through it. Beside the block it returns a StreamingReport: each pass's wall clock and the executor's time inside it, the cycle and shard counts, and the most shards held at once.

Its memory follows the shards in flight, not the shard or cycle count: a partial buffer per family (§4), the last-access tables (§3), at most max_in_flight shards being worked and one filled shard's rows per family waiting (§5), and the output, 64 bytes a commitment and the ShardProofs. The executor's whole output, emulator::trace_run's buffers and event log at about 300 bytes a cycle, never exists; executing twice (§2) costs time instead.

A shard costs its height times its circuit's width (circuits.md §1), however few of its rows are live. gkr::forward holds every inner layer as field elements, 32·Σ_{k≥1} w_k·2^{n_k} bytes over layer k's width and variable count (42 GiB for a 2^18 KECCAK_F shard, 8.4 GiB for a 2^20 SHIFT_BITWISE one), and gkr::prove adds a copy of the layer it reduces and an eq table. The opening, after the forward pass is dropped, copies every committed column.

Measured on the base proof of recursion.md §10:

workload block 257,510 of glamsterdam-devnet-8, revm-block-stateless: 60 transactions, 101.5 Mgas, 198M cycles, 207 shards
machine 32 vCPUs, 247.7 GiB, --in-flight 12
pass 1 191 s; 25.7 vCPUs busy on average; one-thread fills 81% of its shard-seconds; sampled RSS at most 15.9 GiB
pass 2 2,290 s; on average 11.95 of 12 shards held and 30.4 vCPUs busy until the guest exits; then the exit's 460 s tail, whose longest stretches are the two KECCAK_F shards' one-thread fills, 200 s and 279 s
peak RSS 173.92 GiB: the two 2^18 KECCAK_F shards, together in the tail with nothing else in flight

Twelve shards in flight never reached that peak: a delegation family's height set it, and max_in_flight bounds only how many shards coincide.

2 The two passes#

pass 1                                       pass 2
  execute; for each shard as it fills:         execute again; for each shard as it fills:
    fill its M columns, commit them,             fill every column, prove the shard,
    keep the points, drop the columns            keep the proof, drop the rest
  at exit: the window families' shards         at exit: the window families' shards
  the statement, then G1–G11                   the statement's roots, the BlockProof

Each pass drives a fresh emulator::StreamingRun, which hands over a family's buffer as a ShardChunk the moment it reaches the family's height (§4); the workers of §5 take the chunks.

Pass 1 commits each shard's M columns, its family's fill with them moved out and no multiplicities, one commitment a column (in the recursion format one a stack of 2^σ, recursion.md §1.3). At exit it derives the window list, shard counts and boundary from the final state (§3), asserts the cut equals trace::plan_shards, commits the window families' shards, puts the commitments in statement order by (family, index) (proof.md §1) and runs G1–G11 (proof.md §2). A commitment reads no transcript, so only its place in the absorbed order matters, not when it was computed.

Pass 2 re-executes. The emulator is a pure function of (image, io), with no clock, randomness or threads, so it cuts the same shards; pass 2 asserts that its CycleProfile, Execution, window list and boundary are pass 1's. Each shard gets every committed column, multiplicities included, and prover::prove_shard_columns: shard transcript, GKR proof, opening (proof.md §4, §5). Proofs go to their statement positions, and prover::public_inputs copies each shard's two memory roots into the statement.

M is not recommitted: a shard's opening takes its M commitments from the statement, pass 1's, and its polynomials from pass 2's columns, so columns that differed would give an opening the verifier refuses.

3 What survives an execution#

A streaming run records no memory event. At exit StreamingRun::finish hands over each non-empty partial buffer as its family's last shard, and a StreamedExecution: the last-access tables (trace::MemoryState), the CycleProfile and the Execution (execution-trace.md §11). Beyond the shards' rows, the guest's inputs and its journal, everything the statement needs is a function of that final state: the boundary (trace::build_boundary_finals), ZERO_WINDOWS' list (trace::init_windows), the window families' teardowns (memory.md §3) and the field-window count. So a window family's shard exists only once the execution is over (§5).

A fill reads one shard through prover::ShardSource, its ShardRows a trace::RowSlice (a cycle-owning family), a trace::FrameSlice (a delegation family) or, for a window family alone, the final MemoryState. The streaming path builds it over a fresh chunk, ShardSource::archived over a slice of a TraceArchive (§6); nothing else differs. Memory columns come from a shard's rows alone (execution-trace.md §11), and checker::memory_columns_from_log rebuilds them from the event log, independently (circuits.md §3).

4 The shard plan#

A family's rows, in the order they are appended, are cut into shards of its VmConfig height h: shard i is rows [i·h, min((i + 1)·h, len)), the last padded to h with zero rows (memory.md §2). trace::plan_shards is ⌈rows/h⌉ per family over the CycleProfile, cycles for a cycle-owning family and invocations for a delegation family, so a family the execution never reached has no shard. A window family plans 0; its shards are windows (memory.md §3), counted by shard_counts in crates/prover/src/lib.rs: one INIT_TEARDOWN shard and one of each public window whatever the execution did, a ZERO_WINDOWS shard per entry of init_windows, one per advice window supplied (trace::advice_window_count), and field windows through the highest cell touched (MemoryState::field_windows). The counts are the statement's shard_counts (proof.md §1).

The flush. StreamingRun makes the cut as it runs. After a cycle is recorded, a buffer that has reached h rows is handed over as ShardChunk { family, index, rows }, index = rows/h − 1, and replaced by an empty one. A cycle appends at most one row to any buffer, the owning family's and, for a delegation request, one invocation to the delegation family's, so a buffer reaches h without passing it, a step fills at most two, and no chunk is split. At exit finish hands over the partial buffers. Chunks arrive in fill order, not statement order, and pass 1 asserts that each family's count is the plan's.

5 The pipeline#

pipeline (crates/prover/src/streaming.rs) runs both passes: max_in_flight workers under std::thread::scope and one std::sync::Mutex around a Source, which holds the executor, the filled shards no worker has claimed, and the counts. Under the lock a worker gives back its shard and claims the next: a waiting one, or else it steps the executor itself until a buffer fills (Source::claim, the only place the guest runs). Outside the lock it builds the shard's columns, works it and drops it. These are the prover's only threads and only lock; within a shard, parallelism is rayon over data.

held bound by
claimed shards, and all built from them max_in_flight one a worker; asserted in Source::claim
filled, unclaimed shards rows only, one per family the executor steps only for a claim with nothing waiting, a step fills at most two buffers and the exit one per family; asserted in Source::admit

The executor never runs ahead of demand, and there is no batch: a slow shard holds one worker. The workers are not rayon threads. A shard's MSMs, forward pass, sumcheck and opening run on rayon's global pool, so RAYON_NUM_THREADS sets the cores the shards share, and a worker blocked in that work cannot take a second shard as a rayon thread waiting in a nested join would. Fills run on the workers' own threads, one each, so up to RAYON_NUM_THREADS + max_in_flight threads are runnable. Fork-join cannot express this: below one shard per core, a batch waits for its slowest shard.

  • The block is independent of the schedule. A shard's proof is a function of the global state and its own columns, its transcript a fresh sponge seeded with the digest (proof.md §4); no proof depends on the thread count (gkr.md §5); proofs are placed by statement position. crates/prover/tests/streaming.rs compares the bytes at 1 and 8 in flight.
  • The failure returned is the earliest in fill order, at any worker count: claims follow fill order, a claimed shard is worked to its end, a failure stops later claims (Source::fail), and an executor failure ranks after every shard it filled.
  • No deadlock: the lock is never held while a shard is worked or taken twice by one worker, and nothing waits under it but the executor's step.
  • A panic stops the claims, through a drop guard (StopOnPanic) or, inside the executor, the poisoned lock; the shards in flight finish, and the panic is re-raised as itself.

The window families' shards follow the pipeline, built from the final state in rayon batches of at most max_in_flight, which are all that a ThreadPool::install around the call bounds.

The knob. max_in_flight, at least 1, is an argument because only the caller knows the machine; bench prove --in-flight defaults to 8. StreamingReport::peak_in_flight is the most shards claimed or batched at once in either pass.

6 The retained archived path#

emulator::trace_run keeps a whole execution, every buffer and the MemoryEventLog, and trace::TraceArchive::from_execution holds it (execution-trace.md §11). The per-shard component reads one through ShardSource::archived, with the same fills and shard proving: prover::statement_inputs (counts, windows, boundary, every shard's M columns), global_commit_phase, shard_columns, shard_memory_columns, prove_shard, prove_shard_columns and public_inputs. checker::TamperHarness is built on it (circuits.md §3): it writes changed cells into shards' columns, recommits changed M columns in a fresh global commit phase and re-proves, which needs an execution held still and read twice. Streaming has no such seam: pass 2 rebuilds, by re-executing, the columns pass 1 committed, so a cell changed in either pass would contradict the other.

prover::prove_block(setup, archive, plan), advance(setup, archive, until) and finish(archive) prove a block from an archive; nothing outside crates/prover/src/phases.rs calls them. prove_block refuses a plan that is not plan_shards of the archive's profile. advance fills the archive's four later phase sections in order, timing each, and decodes any it already holds, so an imported archive resumes; a stopped streaming run starts again. No column is stored: a phase rebuilds them from the archive. The sections, in proof.md §9's encodings, each refusing a byte too many or too few:

section content
PostCommit the statement's PublicInputs bytes, without roots; the global transcript after G11 as its 226-byte postcard snapshot (transcript.md §3); the four memory challenges; the digest
PostGkr per shard, in statement order: family u32, index u32, ts_start and ts_end u64, the witness commitments, the outputs, the GKR proof, the base claims' point, the shard transcript's snapshot after the GKR proof
PostOpening each shard's ShardProof bytes
Final the complete PublicInputs bytes, then the proofs

审计专区/证明流水线

递归

规范原文docs/spec/recursion.md以 Markdown 查看

摘要

一份基础证明如何在自身不被改变的前提下,变成一份由合约检查的 Groth16 证明。它规定了递归格式及其堆叠承诺、域内存与四个协处理器电路族、节点重放的 tape、叶程序与节点程序以及贯穿整棵树的 transcript 链、节点的公开输出(journal)、延迟打开如何折叠为一个累加器、调度器、带绑定线与两阶段仪式的 Groth16 判定器、合约,以及实测成本。

以下规范原文以英文维护;英文是本规范的标准语言。

How one base block proof becomes one Groth16 proof a contract checks. The section numbers are the ones the code cites. Where this page and the code disagree, the code is right.

base proof ──► leaves ──────► internal nodes ──► root ───► decider ───► contract
N shards,      each a run     each 2–4           covers    the root in  folds the root's
base format    of base shards children           0..N      Groth16      points; two pairings
  • A node is this VM proving a verifier program. It verifies shards and folds every Mercury check they defer into one accumulator (A, B), the claim e(A, [1]_2) = e(B, [x]_2). Nothing pairs before the contract.
  • Base proving is untouched. No base key, statement or proof moved a byte: a leaf verifies base shards as they are.
  • Nodes are proved in a recursion format (§1) over a field memory (§2) with four coprocessor families on it (§3–§6), and they replay tapes (§7) rather than run verifier-core on RV32, which measured 3.0B cycles for block 257,510's 207 shards: fifteen times the block itself.

1 Two formats, one code path#

1.1 The rule#

A statement is in the recursion format exactly when its VmConfig holds FIELD_WINDOWS (VmConfig::is_recursion), which is exactly when its program declares a field family. No wire form says which format applies.

1.2 The delegation registry#

constants::delegation::TYPES is one append-only table, and its first BASE_TYPES = 6 rows are all the base format knows. constraints::family_circuit is the base registry; constraints::recursion_circuit differs from it in two ways only: its ADD_SUB knows every row and carries §1.4's rule, and the five families of §2–§6 exist. VmConfig::circuit picks the registry, for a key's load rule and the prover alike.

1.3 Stacked commitments#

Every commitment a shard opens is a point its parent folds (§8.3). So a recursion shard commits each of its two phases — its M columns, and its W columns with the multiplicities — as stacks of 2^σ columns. At height 2^n, with k_M and k_W columns (VmConfig::stack_vars):

σ = min(24 − n, the smallest even σ with 2^σ ≥ max(k_M, k_W))
  • Column i is slot i mod 2^σ of stack ⌊i / 2^σ⌋. A stack is the (n + σ)-variate multilinear whose evaluations [j·2^n, (j + 1)·2^n) are slot j's column, and its commitment is that polynomial's Mercury commitment. 24 is the ceremony's size.
  • The GKR pass leaves each column's value v at u. Then σ challenges r are drawn (STACK_CHALLENGE), a stack's value is Σ_j eq(r, j)·v_j, and a setup column is a stack of one, eq(r, 0)·v, its commitment unchanged. The shard's one batch opening is at u ‖ r, over the M stacks, the W stacks, then the setup columns.

σ = 0 is the base format exactly.

1.4 A recursion request leaves a0 past its frame#

A base delegation request writes 0 into a0. A request of a type past BASE_TYPES writes a0 + 4·words (constants::delegation::a0_after), which the recursion ADD_SUB's deleg_a0_rule holds it to. So frames laid back to back replay as back-to-back ecalls, one RISC-V row a call.

2 The field memory#

2.1 The space#

address_space::FIELD = 10: cells addressed by a u32, each a whole Fr. Its tuples (FIELD, cell, ts, value) join RAM's in the one memory multiset. No instruction reaches it. Only §3–§6's rows do, each access at its row's requesting cycle c and its own slot, 4c + Δ, with a read's usual gap check; a read-only access writes back what it read. A field access is not a MemoryEventLog event, a value not being a u32: trace::MemoryState keeps each cell's last (ts, value), and a recursion execution has no TraceArchive form. It streams.

2.2 FIELD_WINDOWS#

Family 18, 2^20 rows: ZERO_WINDOWS' circuit at a stride of one cell a row. Window w is cells [h·w, h·(w + 1)), initialized to 0. The windows are consecutive from cell 0 — shard i is window i — so a statement lists none, and a cell outside them has no tuple to balance a read against.

The four families on it are invoked, by the delegation ABI: an ecall whose a0 is a frame of words in RAM. A frame's words name cells.

§ family id ecall anchor space height frame a row is
3 FR_OP 19 0x0509 11 2^20 [op, d, a, b] one field operation
4 P2_FIELD 20 0x050A 12 2^18 [n, s, x, y, d] one transcript duplex step
5 FIELD_IO 21 0x050B 13 2^18 [op, cell, ptr] eight RAM words to a cell, or back
6 FQ_OP 22 0x050C 14 2^20 [op, d, a, b] one BN254 base-field operation

3 FR_OP — one field operation a row#

op op
1 MUL d ← a·b 6 EQ a = b, or the row has no witness
2 ADD d ← a + b 7 IMM d ← word b, as an integer
3 SUB d ← a − b 8 SHL d ← a·2^32 + word b
4 MAC d ← d + a·b 9 DIGIT d ← a's low byte, b ← (a − d)/2^8
5 INV d ← a⁻¹, and 0 at a = 0

a, b and d are accessed at slots of their own, so any two may name one cell. EQ is how a tape asserts. IMM and SHL are how it builds a constant with no field arithmetic of the guest's. A scalar's 32 DIGITs ending at 0 represent it mod p, which is all a scalar multiplication needs.

4 P2_FIELD — one duplex step a row#

With the state at cells s..s+3, the row absorbs n ∈ {0, 1, 2} of the cells x, y — the lanes are (n ≥ 1 ? x : s₀, n = 2 ? y : n = 1 ? 0 : s₁, s₂ + n) — and writes poseidon2_permute of them to d..d+3. A state is never overwritten, so a challenge is a cell of the triple that made it. The circuit is flat, every S-box's u² and u⁴ committed, so a parent verifies it as one gate list.

5 FIELD_IO — between RAM and a cell#

Over the eight RAM words w_k at ptr:

  • IMPORT (1): the cell takes Σ_k w_k·2^{32k} mod p. A non-canonical encoding is harmless.
  • EXPORT (2): the words take limbs below 2^32 congruent to the cell. Congruence, not canonicity: a guest that needs the canonical value compares the words with p itself.

Addressability is the multiset's. A word no window initializes cannot balance.

6 FQ_OP — one base-field operation a row#

An element of BN254's Fq is four consecutive cells of 64-bit limbs, congruent to its value mod q and not necessarily below it. Only this family writes one. The op word is a code, three flags and a digit cell (word >> 6): a flagged operand's element is its word plus 8·digit, a bucket chosen by a digit, which is what lets an MSM be a static tape (§8.3).

op
1 MUL d ← a·b
2 ADD d ← a + b
3 SUB d ← a − b
4 MULEQ asserts a·b ≡ d
5 FROM128 d ← a₀ + 2^128·a₁ from two cells below 2^128: a coordinate from its transcript limbs

One integer identity serves all five, a·y + z = q·K + d′, checked over 128-bit groups of limbs with a range-checked quotient and carries. Some of those ranges go through TIMESTAMP, which is why the family is at 2^20. b's and d's four cells share one read timestamp, so an element is only ever written whole; tape::run refuses a tape that reads one written apart, before a fill would.

7 Tapes#

A shard's checks have a fixed shape for its family and height. So the host compiles them, once, into a tape (verifier_core::tape): a straight-line list of coprocessor calls over absolute cells, in which nothing branches on a value. A tape reads three kinds of cell:

  • constants, which it builds itself from IMM and SHL, so they are bound with it;
  • slots, which its caller fills: the statement's digest and memory challenges, and the shard's index, window, roots and commitments;
  • inputs, the proof, IMPORTed from a blob laid out in the tape's Input order (tape::shard_blob), which the tape's own checks are what bind.

tape::shard_tape is verify_shard_local's steps 7–11 and Mercury's field side (pcs_verify), call for call: the shard transcript, the GKR backward pass, the lookup and root checks, the stack challenges and values, and the opening's twelve scalars. Every check is an EQ. It leaves three things to its caller (§8): the shard's time window against its neighbours', step 10c on the two public shards, and everything on the curve — to a tape a point is four transcript limbs, and the batch's cm* is a hint.

tape::schedule reorders a tape into runs of one family's calls, tape::encode is the form a guest replays — the cells its imports fill, then runs of frames — and tape::run is the native reading, over a Memory that models each access's timestamp as well as each cell's value.

8 Nodes and the tree#

8.1 Two programs, one procedure#

guests/recursion is two binaries. The leaf verifies shards from..to of one statement of the base program. The node verifies two to four whole statements of the two recursion programs — its children's proofs — reads each child's journal out of the output window step 10c binds, holds the children to one another, and folds their accumulators beside their shards'.

Both run verifier_core::node::node through a Driver. The host runs it natively (host::recursion), so it refuses whatever a guest would, first and by name, and it writes the advice the guest reads. The guest runs it by coprocessor calls. A binary's image — every shard tape, the fold's templates, the constants — is built by build.rs with verifier-core itself and sits in .rodata, so a program's identity binds every tape it replays.

  • The base program's identity is a constant of the leaf's image. The SRS digest and the generic table are constants of both images.
  • A node takes the two recursion programs' identities as claims and journals them, for the top to check once.
  • A program's setup commitments are advice, held to its identity by recomputing it.

The global transcript is a chain across the tree (verifier_core::chain). The node with shard 0 runs the prefix, G1–G7. Every node absorbs its own shards' memory commitments, G8, from the state its predecessor left. The node with the last shard runs the suffix, G9–G11, which settles the digest and the memory challenges every node took as claims. A node that holds a whole statement makes its memory argument, Π reads · R_b = Π writes · W_b.

A node holds its children to: exit status 0; one base statement — its shape, digest, challenges, io_digest, exit status and shard count; adjacent shards; chain states that meet; time windows in order across the seam; and, of a node child, the two identities it requires itself.

8.2 The journal#

47 cells, each a 32-byte word (node::journal):

cells
0 a digest of the base statement's shape: its shard counts and windows
1–7 its global digest, four memory challenges, io_digest, exit status
8–10 its shard count, and the shards this node covers, from..to
11–18 the chain's state at from and at to: three lanes and a pending input each
19–22 the covered shards' read and write root products; the boundary factors where to is the count
23–28 the first and last covered shard's family and time window
29–44 A and B, each x then y in four 64-bit limbs
45–46 the leaf program's and the node program's identities this node requires; 0 for a leaf

The root covers 0..count: every base shard verified, the transcript run end to end, the memory argument made. What is left is one pairing check and two identities.

8.3 Folding#

After each shard's tape the node's own transcript absorbs the shard transcript's final state (FOLD_STATE) and draws w and w′ (FOLD_WEIGHT), so a shard's weights follow everything they weight. Then, as scalars of points:

  • entry i of the shard's Mercury check gets w·e_i, on its side (pcs_verify::ENTRY_POINTS);
  • the batch check cm* = Σ ρ^i·cm_i is folded beside it: cm* gets w′ more, and each cm_i gets −w′·ρ^i;
  • [1]_1 and the setup commitments, which every shard of a family shares, accumulate one scalar each and enter once;
  • a child's A and B enter under a weight drawn after its whole journal (FOLD_CHILD).

Each side is one MSM on FQ_OP (verifier_core::fold): Pippenger with 8-bit digits over GLV halves, 16 windows of 256 buckets, every step a static template. A point is held to the curve and its scalar's split to the scalar, then added to one bucket a window through an indirect operand. Inversions are host witnesses held by a MULEQ, and buckets start at offsets so that no addition degenerates. A point costs about 400 FQ_OP calls.

8.4 The scheduler#

bench recurse <dir>/<stem> --out <out>, over a base proof archive (tools/bench/src/recurse.rs):

  • Keys. It writes base.key and programs.key into <out> and builds the two binaries with APOGEE_RECURSION_KEYS=<out>, where their build.rs reads them. The node is built twice: once with no image, for the two programs' keys, and once over them.
  • Plan, fixed before anything is proved (host::recursion::Tree::plan, <out>/tree.txt): leaves of at most --leaf (64) consecutive base shards, then levels of internal nodes over two to --fan-in (4) children, a lone leftover carried up. A program has sixteen families and each costs at least a shard, so a node is sixteen shards before any work and leaves are cut large.
  • A node is a process, bench recurse-node: it verifies its inputs natively, builds its advice only then, proves, verifies, holds the proved journal to the native one and writes <out>/<id>.block. At most --in-flight nodes run, with --in-flight × --shards-in-flight shards in flight across them: a node takes its share of what is spare when it starts, so a root alone has the whole budget. A proof already in <out> is kept, so a stopped run resumes; a run whose plan or programs differ is refused.
  • At the root it checks what a verifier owes beside the root's own proof — the journal covers 0..count and is the archive's statement, it requires the two programs' identities, and (A, B) discharges — and then runs §9.

9 The decider#

The root is still a GKR proof and some hundreds of points, and a contract can check neither. host::decider splits its verification in two.

The circuit is §8.1's node procedure over one child, the root, through a Driver that writes rank-1 constraints: an FR_OP is one constraint in the common case and none where it only copies, a duplex is 255, advice is a free wire. It verifies the root as a node would, and holds its journal to from = 0 and to = count. But it folds nothing: every MSM template is skipped, and each point's four limbs and its scalar are bound wires instead, after the two identities, the base statement's exit status, and its public input and output, a wire a byte, whose digest the circuit holds to the journal's io_digest.

crates/groth16 is Groth16 over this repository's BN254. A circuit streams its constraints into a sink, so no matrix is held. Three things are not the textbook's:

  • Bound wires are values the verifier holds, too many to be public inputs. The proof carries their commitment D = Σ w_j·[(β·A_j + α·B_j + C_j)/η]_1 under a fifth trapdoor η. A challenge c is SHA-256 of D and the verifier's values, and the circuit ends with acc ← (acc + wire)·c over the bound wires. The public inputs are c and that result, both of which the verifier computes from its own values, and the check is e(A, B) = e(α, β)·e(IC, γ)·e(C, δ)·e(D, η), IC being the public wires' points under 1, c and the result. D is fixed before c, so wires that differ from the values agree with them at c with probability len/r.
  • No blinding. A proof hides nothing and is a function of its witness.
  • A Lagrange basis. A and B are sums over the constraints, Σ_j (A·w)_j·[L_j(τ)], not over the wires. So the one element a key holds a wire is [(β·A_i + α·B_i + C_i)/x]_1, x being γ, η or δ — and a powers-of-tau ceremony already publishes [L_j(τ)].

The key is a ceremony's, in two phases:

  • Phase 1 is ppot_0080_24.ptau, the ceremony the tree's own commitments are under (srs::Phase1): the Lagrange basis at the circuit's domain in both groups, and the powers a quotient takes. Everything of the key that depends on τ is a combination of those points, and nothing derives τ.

  • Phase 2 is the circuit's own (groth16::phase2, bench ceremony), and makes α, β, γ, δ and η from 1 by contributions: each multiplies a trapdoor by a factor only its contributor knew, so a trapdoor is unknown while one contributor to it was honest.

    step
    init every trapdoor 1: a wire's [A_i(τ)]_1, [B_i(τ)]_1, [C_i(τ)]_1, and [τ^k·Z(τ)]_1. Deterministic from the circuit and the file
    round 1, contribute to α and β: [β·A_i]_1 and [α·B_i]_1, kept apart
    seal a wire's three terms summed
    round 2, contribute to γ, δ and η: the sum over the wire's trapdoor, and [τ^k·Z(τ)/δ]_1
    key the last state verified and, if every trapdoor has a contribution, written as the key

    The order of the rounds is the soundness. A prover may hold a wire's three terms only summed, over δ or η: apart, it could give A, B and C three witnesses. A contribution to α or β scales the terms apart, so those are finished before anything is divided.

    A state carries each contribution's record — its factor in G1 with a Schnorr proof of knowing it, bound to the records before it, and the trapdoor in G2 afterwards. Verifying a state checks that chain, then its elements against init's under those trapdoors, one pairing equation over a random combination: against the circuit and the file alone, with no earlier state. Every step lists the records by their factors' points, so a contributor finds its own under the state the key is made of. bench decide reads the key key wrote, and nothing else writes one.

  • setup_dev, bench decide --dev-key, derives all six trapdoors from a public seed. It is for development and tests: anyone forges under it.

The contract (contracts/ApogeeVerifier.sol) is verify(input, output, exitStatus, proof, points), a point being x, y, scalar, side [1]_2's points and then side [x]_2's. It rebuilds the bound values — a point's limbs are its coordinates' halves, or four sentinels at infinity, which is what the root's transcript absorbed — recomputes c and the result, checks the Groth16 pairing, folds each side with ecMul and ecAdd, which is also what holds a point to the curve, and checks e(A, [1]_2) = e(B, [x]_2). Its Groth16 key, the ceremony's two G2 points and the two identities are set at deployment.

bench decide <out> proves under the ceremony's key, checks the proof natively, deploys and calls the contract in revm, and writes decision.constructor and decision.calldata — under --dev-key, development.*.

What a deployment still owes. A key is as trustworthy as its ceremony: one honest contributor a round, which a ceremony run on one machine is not. The circuit depends on the root's shape — its program, its shard counts, the public values' lengths — so a key, and its ceremony, is per shape. And the contract pays about 9k gas a point, because the circuit folds none.

10 Running it#

bench prove --stateless <fixture> --out <dir>             the base proof
bench recurse <dir>/<stem> --out <out> --in-flight 4      the tree
bench ceremony <out> init                                 the decider's key: once a root shape,
bench ceremony <out> contribute                           each contributor in turn, to alpha and beta
bench ceremony <out> seal
bench ceremony <out> contribute                           and to gamma, delta and eta
bench ceremony <out> key
bench decide <out>                                        the Groth16 proof, and the contract

It needs assets/ptau/ppot_0080_24.ptau. Measured on block 257,510 — the tree on a 32-CPU, 247 GiB machine, the ceremony and the decider on an 18-core laptop:

base proof 207 shards, 14.5 MB, 2,481 s
tree 4 leaves of at most 64 base shards and a root: 116 shards
leaves, four at once 21, 24, 23 and 27 shards; 2,157 s; 92 GiB peak
root, four shards in flight 21 shards, 460 s, 1.03 MB
decider's circuit 7,896,686 constraints, a domain of 2^23
ceremony init 65 s; a contribution 50–56 s; key 70 s, 12.7 GB; the key 2.65 GB
decider the key read in 1 s, the proof 18.5 s, 6.1 GB
contract 358 points; 3,620,026 gas; 34,980 bytes of calldata

审计专区/证明流水线

以太坊区块

规范原文docs/spec/ethereum.md以 Markdown 查看

摘要

参考工作负载:在虚拟机内用 revm 执行以太坊区块。它规定了客户程序(guest)crate 及其两个二进制程序、mini-block 的见证格式与输出承诺、无状态校验器的输入与输出、它支持的分叉、一个结果证明了什么、一个区块如何在 revm 上运行、签名恢复、与 zkEVM 测试发布版本的一致性、每条校验规则在哪里检查,以及一个区块如何被记录。

以下规范原文以英文维护;英文是本规范的标准语言。

guests/revm-block runs Ethereum blocks on revm inside the VM. This page specifies its two binaries: the mini-block binary's input, BlockWitness, and its output commitment; the stateless validator's input, result and rules; how each runs a block on revm; and how a block is recorded.

1 The guest crate#

One library, revm_block (src/lib.rs and its modules), and two binaries, each its own program identity, the same for every block because the block is advice (public-values.md §6):

binary advice journal exit status
revm-block (src/main.rs) a BlockWitness (§2) the output commitment (§3) 0; 61 not a canonical witness; 62 not executable
revm-block-stateless (src/stateless_main.rs) statelessInputBytes (§4) the 43-byte result always 0

The mini-block binary runs transactions, usually a block's first few, over a pre-state recorded from a node (§6); it proves an execution, not a block's validity (§3). Full blocks are proved with the stateless validator: the mini journal grows by a record a transaction and outgrows the public window (public-values.md §9). On the host the library is the oracle crates/emulator/tests/revm.rs holds the mini binary's journal to, over unpatched upstream crates.

Dependencies. revm is the workload being proven: it and what revm-precompile brings (arkworks, k256, p256, sha2, ripemd) are reachable from no prover, verifier or other guest. It is built without default features, so without blst, c-kzg or libsecp256k1. The pin is exact, =43.0.1, because an identity is a digest of the image (program.md §8), and the set is the reference stateless guest's (paradigmxyz/stateless's lock): twelve crates, held in both lockfiles by crates/host/tests/revm_lock.rs, its revm-handler 43.0.1 carrying the EIP-8037 system-call state-gas reservoir tests-zkevm@v21.0.1 expects.

Delegations. Both binaries declare KECCAK_F, SHA256_COMP, MOD_MUL and EC_ADD: keccak through alloy-primitives' native-keccak hook (revm_block::native_keccak256), SHA-256, secp256k1 and BN254 through vendored crates (delegation.md §10, guests/vendor/README.md). revm-precompile is patched in its Crypto trait's default bodies, not given a second implementation by install_crypto: two types behind crypto()'s OnceLock<Box<dyn Crypto>> stop LLVM devirtualizing its calls, keeping code it otherwise strips, 870,828 bytes of .text on the mini binary, past its tables' reach.

Code size. No ELF is committed: host::fixture::build_revm_guest builds either binary at --release, proved at host::fixture::revm_params — 2^20 for every family whose height is a choice (revm_block::TRACE_HEIGHT_RELEASE), each delegation family's default, bytecode_size_words = 2^21. A 2^20 table reaches 1.9375 MiB of .text (program.md §5); the stateless binary's is about 1.96 MB, 96.6% of it. The debug image needs 2^22 and is only ever run.

2 BlockWitness#

The mini binary's advice: postcard of revm_block::BlockWitness, a format of this repository's, written by host::recorder (§6). Fields in declaration order; a word is 32 big-endian bytes; a u8, an Option tag (0 or 1) and a fixed array are raw bytes; every other integer and every length is a varint.

BlockWitness          env BlockEnvWitness; accounts Vec<AccountWitness>, by address;
                      txs Vec<TxWitness>, in execution order
BlockEnvWitness       chain_id u64; spec_id u8 (revm's SpecId); number word; beneficiary [20];
                      timestamp word; gas_limit u64; basefee u64; difficulty word;
                      prevrandao Option<word>; excess_blob_gas Option<u64>;
                      blob_gasprice Option<u128>; slot_num u64;
                      block_hashes Vec<(u64, word)>, by number
AccountWitness        address [20]; nonce u64; balance word; code Vec<u8>;
                      slots Vec<(word, word)>, by key, zero values included
TxWitness             caller [20]; to Option<[20]>, None a creation; value word; data Vec<u8>;
                      gas_limit u64; gas_price u128, the max fee from type 2;
                      gas_priority_fee Option<u128>; nonce u64; chain_id Option<u64>;
                      access_list Vec<([20], Vec<word>)>; blob_hashes Vec<word>;
                      max_fee_per_blob_gas Option<u128>; authorizations Vec<AuthorizationWitness>
AuthorizationWitness  chain_id word; address [20]; nonce u64; authority Option<[20]>, recovered

2.1 One state, one encoding#

BlockWitness::decode refuses, with exit 61:

rule WitnessError
spec_id is a SpecId UnknownSpec
excess_blob_gas and blob_gasprice both present or both absent BlobPairing
accounts, each account's slots, block_hashes strictly ascending AccountsNotSorted, SlotsNotSorted, BlockHashesNotSorted
the bytes are exactly BlockWitness::encode's for the value Malformed

The last closes postcard's two second encodings: postcard::from_bytes ignores trailing bytes, and its varints accept non-minimal forms (81 00 reads as 1). A code hash is computed, not carried, and a transaction's type is derived from its fields (TxEnv::derive_tx_type).

2.2 Execution#

revm_block::WitnessDb answers revm from the witness and refuses every miss (DbError, exit 62): the witness is unbound advice, so a default would be a value the prover chose.

  • Absence is recorded: WitnessDb::basic answers None for an account recorded with nonce 0, balance 0 and no code. A zero slot is recorded like any other.
  • BLOCKHASH reads env.block_hashes. revm answers 0 without asking for any block but the 256 before the current one, and serves those from the database, not EIP-2935's contract: at most 256 entries.
  • Code is Bytecode::new_raw_checked's: bytes beginning 0xef01 that are not a 23-byte EIP-7702 delegation, which a few pre-EIP-3541 accounts hold, are DbError::MalformedCode, where Bytecode::new_raw would panic, an exit 101 that names nothing (ecall-abi.md §7).
  • The block gas limit is a running bound. revm checks each transaction against the block's limit and keeps no total; revm_block::run, the block executor, refuses transaction i unless gas_limit_i ≤ env.gas_limit − Σ_{j<i} gas_used_j, the Yellow Paper's intrinsic validity.
  • The blob gas price is recorded (§6). It derives from the excess through the fork's update fraction, 3,338,477 at Cancun, 5,007,716 at Prague, raised by each BPO fork (§4.2), and revm 43 knows only the first two. revm holds each type-3 transaction's max_fee_per_blob_gas to it.

Not in the witness: signatures, caller and each authority being the producer's recovery, unchecked; a parent header, and the header rules against it; a state root (§3); a slot number, which the recorder writes as 0, no JSON-RPC method serving EIP-7843's.

3 The output commitment#

The mini binary's journal, revm_block::run's return:

per transaction, in order   status u8 (0 halt, 1 revert, 2 success) ‖ gas_used u64 LE
                            ‖ output_len u32 LE ‖ output: the return data, empty on a halt
logs commitment       32    keccak256 of  count u32 LE ‖ per log, in emission order:
                            address 20 ‖ topic_count u8 ‖ topics, 32 each ‖ data_len u32 LE ‖ data
post-state summary    32    keccak256 of  count u32 LE ‖ per account, by address:
                            address 20 ‖ nonce u64 LE ‖ balance 32 BE ‖ code_hash 32
                            ‖ slot_count u32 LE ‖ per slot, by key: key 32 BE ‖ value 32 BE

The record count is the witness's. The summary covers the state revm's finalize returns: every account the block loaded, read-only and nonexistent ones included, with every slot it loaded.

What a proof states: some canonical BlockWitness makes revm_block::run return this journal. Nothing ties the witness to a chain; a reader holding one recomputes the journal natively. And the journal tells witnesses apart only as far as the execution reads them: a slot read and then overwritten unconditionally reaches nothing, while every loaded account's final balance and nonce are in the summary.

4 The stateless validator#

revm-block-stateless maps tests-zkevm@v21.0.1's statelessInputBytes to its statelessOutputBytes, byte for byte. The formats and rules are ethereum/execution-specs' verify_stateless_new_payload at the release's commit (host::zkevm::RELEASE_COMMIT); revm_block::stateless::run is the guest's whole computation, and §5 lists its rules.

4.1 Input and output#

input    schema_id u16 BE ‖ SSZ(StatelessInput)                          ssz::decode
           new_payload_request   the schema's fork's NewPayloadRequest
           witness               state: trie-node preimages; codes; headers: RLP, oldest
                                 first, the parent last, at most 256
           chain_id              u64
           public_keys           eth-act/ere-guests v0.17.1's layout only: 65 bytes a transaction
output   new_payload_request_root 32 ‖ successful_validation 1 ‖ chain_id u64 LE ‖ schema_id u16 LE

The layouts' fixed parts are 16 and 20 bytes, so no input is both; the second is the zkEVM benchmark's. The root is hash_tree_root under EIP-7916's and EIP-7495's progressive forms as of 2026-01-15 (ssz::request_root), whatever the layout. Decoding is as strict as the spec's: every offset against the bytes it bounds, every bounded list against its limit, nothing after the end.

The guest exits 0 on every input. One that does not decode, or names a schema §4.2 does not list, publishes the sentinel, 43 zero bytes (ssz::SENTINEL); any other publishes its request's root, its verdict, its chain id and its schema id. The empty input is the one a run cannot be given, a run without advice having no advice region.

4.2 Forks#

The schema id, fork_index << 8 | 0x01, names the fork; no activation schedule is compiled in (block::fork).

schema fork request revm SpecId blob target, max update fraction
0x1201 Osaka Electra/Fulu OSAKA 6, 9 5,007,716
0x1301 BPO1 Electra/Fulu OSAKA 10, 15 8,346,193
0x1401 BPO2 Electra/Fulu OSAKA 14, 21 11,684,671
0x1501 Amsterdam Gloas: a block access list, a slot number, EIP-8282's two request types AMSTERDAM 14, 21 11,684,671

4.3 What a result proves#

true says the request whose root is published is a valid block on chain chain_id under the fork schema_id names. The witness needs no binding: the root fixes the payload, and the witness is held to it by hashes — the parent header to the payload's parent_hash, each ancestor to its child's, the state trie to the parent's state_root and each node to its parent's reference, each code to its account's code hash. A node a read needs and the witness lacks is an error, never an absence (mpt::get). So a wrong witness cannot make an invalid payload valid; but false says only that this input did not validate, which a prover can arrange for any payload.

4.4 How a block runs on revm#

The pre-state is witness::WitnessDb behind revm's State: the state trie under the parent's root, each storage trie parsed on its first read, codes by hash, and BLOCKHASH numbering each ancestor by its position below the block. stateless::execute is the spec's apply_body: the EIP-4788 and EIP-2935 system calls; each transaction; the withdrawals; the requests, from deposit logs and the checked system calls of EIP-7002, EIP-7251 and, from Amsterdam, EIP-8282. The calls before the transactions are block access list index 0, each transaction has its own, and what follows them shares the last. A transaction must fit what is left, Amsterdam metering regular and state gas apart (EIP-8037):

before Amsterdam   tx.gas_limit ≤ gas_limit − Σ gas_used
Amsterdam          tx.gas_limit ≤ 2^32 − 1,  min(tx.gas_limit, 2^24) ≤ gas_limit − Σ regular,
                   tx.gas_limit ≤ gas_limit − Σ state;  the block uses max(Σ regular, Σ state)
both               2^17·blobs ≤ 2^17·max − Σ blob gas

Three rules make the result the spec's where following reth would not:

  • Code loads when revm asks (witness::WitnessDb::code_by_hash), never with its account: the witness carries only the code the spec's execution read, and a coinbase may be a contract nothing calls.
  • Every write precedes every deletion in the post-state replay (stateless::post_state_root), in each trie, as the spec's mpt_set_storage_slots orders them. A deletion that leaves a branch one child needs that child's node, on no changed key's path; the witness carries those the spec's order needs, and writing first needs a subset.
  • One commit per index (stateless::commit_index). revm 43's access-list builder records a value that differs from its commit's baseline, and revm re-bases a value at each call, so committing call by call records a slot one call toggles and the next restores. An index's calls are committed once, each baseline reset to the committed state.

Also the spec's: a checked system contract must have code, deposit events are parsed to the byte, withdrawals precede requests; the TxEnv is built field by field (build_fill would put a dummy authorization in an empty type-4 list); the blob price is a checked fake_exponential (block::blob_gas_price); 0xef01 code that is not a delegation runs as legacy. Declared lengths are added checked and trie parsing is depth-bounded, a panic publishing nothing.

4.5 Signatures#

Every sender and EIP-7702 authority is recovered in the guest (tx::recover_key) under EIP-2's rules, 0 < r < n, 0 < s ≤ n/2, a parity bit, as Q = r⁻¹(s·R − z·G) with k256's arithmetic, which the vendored k256 routes to MOD_MUL and EC_ADD. The verification upstream's recover_from_prehash ends with cannot fail once recovery succeeds and costs about as much again, so it is not done. An authorization that does not recover is skipped, as EIP-7702 says. A key in ere-guests' layout is checked, never used: one a transaction, 0x04 ‖ x ‖ y, naming the recovered sender.

4.6 Conformance#

All 67,251 pairs of tests-zkevm@v21.0.1 match natively (crates/host/tests/conformance.rs, by hand); CI holds the library to a committed subset of 34 — a case for each rule the release reaches, the smallest valid one, every undecodable one — in both layouts, and the binary runs the subset by hand. The release fills only Amsterdam: tools/stateless-ref holds the Electra/Fulu layout to eth-act/ere-guests v0.17.1 and crates/host/tests/canonical.rs the encodings and header rules to two mainnet blocks, but no Osaka-family input has an end-to-end oracle.

5 Where each rule is checked#

The mini binary's rules, then the validator's step by step. A validator refusal is a stateless::Invalid variant, which host::zkevm::verdict names and the guest publishes as false. Paths are revm_block's.

rule refusal code
mini: a canonical witness exit 61 BlockWitness::decode
mini: every read recorded, code well formed exit 62 WitnessDb
mini: each transaction fits the gas left and executes exit 62 run_against
the input decodes under a listed schema sentinel ssz::decode, block::fork
the ancestors decode and chain Ancestors stateless::ancestors
no empty transaction EmptyTransaction stateless::verify
the base fee fits a u64 Unrepresentable stateless::payload_header
the header the payload implies hashes to block_hash BlockHash stateless::{verify, payload_header}
each transaction decodes (EIP-2718, types 0–4) Transaction(i) tx::decode
keyed layout: a key a transaction PublicKeys stateless::verify
the versioned hashes are the request's VersionedHashes stateless::verify
EIP-7934's block size BlockSize block::block_rlp_len
the header against its parent, twelve rules Header(_) block::validate_header
the blob gas price fits a u128 Unrepresentable block::blob_gas_price
the parent's state root is in the witness Witness(_) witness::WitnessDb::new
chain id; signature; keyed layout: the key names the sender ChainId(i), Signature(i), PublicKeys stateless::execute, tx::sender
the transaction fits what is left Capacity(i) stateless::execute
revm executes it, every read in the witness Execution(i) stateless::execute
the system calls SystemCall stateless::{execute, commit_index}
the deposit events Deposits block::deposit_requests
gas used, receipts root, bloom, blob gas used, requests hash GasUsed, ReceiptsRoot, Bloom, BlobGasUsed, RequestsHash stateless::verify, block
Amsterdam: the access list's item count and hash AccessList stateless::verify, alloy_eip7928
the post-state root StateRoot, Witness(_) stateless::post_state_root, mpt

6 Recording a block#

host::recorder::record(rpc, block_number, range) makes a BlockWitness for a block's first n transactions or all of them (recorder::TxRange) by running them once against a node: recorder::WitnessRecorder is a revm::Database over the parent block's state that records each answer, and the transactions run through revm_block::run_against, the guest's own executor, so the record is what the guest will read. The result is put through BlockWitness::decode.

  • Reads. An account is eth_getProof with no keys, absent when nonce, balance, code hash and storage hash are all empty or both hashes are zero, Geth's answer; code eth_getCode, checked against the hash; a slot eth_getStorageAt; a header eth_getBlockByNumber; the blob gas price eth_feeHistory's baseFeePerBlobGas, a receipt's blobGasPrice existing only for type 3.
  • Choices. The hardfork is mainnet's by number (recorder::mainnet_spec): before the Merge is refused, after Osaka runs as Osaka. caller is the node's from; authorities are recovered on the host.
  • The client, host::rpc::Rpc, files each response under the SHA-256 of its canonical request in the fixture's rpc-cache/, so a second recording is byte-identical and offline; a miss without ETH_RPC_URL is an error. A request goes through curl, the endpoint and its key on the command line, retried on a transport failure, a 5xx or a 429.
  • On disk (host::fixture): <stem>.json, a Pin naming the block and the length and SHA-256 of <stem>-witness.bin and <stem>-journal.bin, native revm's journal, beside rpc-cache/. crates/host/tests/vectors/mini-block* is block 26,057,509's first two transactions, refreshed by kat-gen -- block (tools.md §7).

Nothing here produces a stateless input. eth_getProof returns the nodes on a key's path, and a deletion that collapses a branch needs its surviving sibling's node, which is on no changed key's path (mpt::MptError::BlindedCollapse is the validator's refusal without it), so the proofs of a block's keys are not a witness. Stateless inputs come from an external producer, a tests-zkevm release or the zkEVM benchmark's datasets; host::zkevm reads every JSON object carrying both statelessInputBytes and statelessOutputBytes, and bench prove --stateless proves one as it is (tools.md §1).

参考

术语表

规范原文docs/glossary.md以 Markdown 查看

摘要

远地虚拟机特有的术语,每个术语一行,各自链接到规范中定义它的章节。GKR、LogUp、KZG、RISC-V 等由文献确定的术语不在其列。

以下规范原文以英文维护;英文是本规范的标准语言。

The project's own vocabulary, one line a term, each linked to the section that defines it. Terms the literature fixes (GKR, LogUp, KZG, RISC-V) are not listed.

term meaning defined in
accumulator, accumulator entry a deferred Mercury check as twelve (side, scalar, point) entries mercury.md §6
advice memory whose initial values the prover chose, bound by nothing public-values.md §6
anchor, anchor space tuples in a delegation type's own space pairing a request with one invocation delegation.md §5
archived path proving from a held TraceArchive; only the tamper suite (checker::TamperHarness) does streaming.md §6
artifact a circuit as data, CircuitArtifact; also an exported ProgramImage gkr.md §4, program.md §3
base claims each committed column's claimed value where the backward pass ends gkr.md §5
base format, recursion format recursion if a statement's VmConfig holds FIELD_WINDOWS, else base recursion.md §1
block BlockProof: config, statement and its shards' proofs proof.md §1
bound wire a decider value the verifier holds, committed instead of a public input recursion.md §9
boundary the registers' and pc's final timestamps and values; they have no rows memory.md §4
cached entry a sub-expression inlined into its list's gates, not a column gkr.md §3
canonical form an element as its value, 32 bytes little-endian, below the modulus primitives.md §1
challenge slot a gate coefficient's challenge: drawn, or derived by the verifier gkr.md §3, §5
channel one LogUp identity over a shard's lookups into one table lookup.md §1
copower x < p as x·2^32/p < 2^32, void without a direct bound lookup.md §11
cycle-owning the execution families 0–6, whose time windows are ordered proof.md §8
decider a Groth16 proof that the recursion root verifies, for the contract recursion.md §9
declaration record, static detachment 12 bytes a linked shim leaves in the image: how a delegation is declared delegation.md §7
decoded table an instruction family's setup columns: row i is pc 2i program.md §5
delegation a family proving a function of a RAM frame, invoked by ecall delegation.md §1
discharge spending an accumulator; the rule that each lookup is one leaf of its tree mercury.md §6, lookup.md §11
enforcing, producing a gate vanishing on every row; one writing the next layer gkr.md §1
extra mask, kind, kind bit family_extra_mask = 1 << kind, a kind being a mnemonic's index; b_k its bit program.md §6
family a circuit and the rows it proves: instructions (0–6), memory locations or invocations circuits.md §1
field memory address space FIELD: cells of one Fr, for the recursion families recursion.md §2
fold merging a node's deferred Mercury checks into one (A, B) recursion.md §8
frame an execution family's queries; a delegation's RAM words at a0 memory.md §2, delegation.md §4
gate list, row-wise, halving the gates from layer k to k + 1, keeping the height or halving it gkr.md §1
gated key, neutral tuple a lookup tuple under its selector; off, it reads a neutral table row lookup.md §4
generic table the committed table of ZeroEntry, AND, U16GetSign, ShiftPowers lookup.md §9
global transcript, global state digest G1–G11: the statement, M commitments, memory challenges; G11 seeds each shard proof.md §2
HALT_PC 1: the exit row's next_pc, the pc's final value memory.md §5
height a family's rows a shard: 2^8, 2^12, 2^16, 2^18, 2^20 or 2^22 program.md §7
identity, image column one Fr digest of the decoded tables, the image, the entry pc, VmConfig program.md §8
in flight shards worked at once, at most max_in_flight streaming.md §5
invocation, request a delegation's row doing one call; the ecall row asking for it delegation.md §1, §5
journal the public output: what the guest leaves in the output window public-values.md §1
laws Laws 1–4: locality, derived width, top layer, single source of truth gkr.md §4
layer layer 0 the committed columns, the top the outputs; L{k}[j] between gkr.md §1
leaf, node, root recursion programs: a leaf verifies base shards, a node 2–4 child proofs; the root, all recursion.md §8
live row, padding row m_pc = 1, or a zero row; in a decoded table, an instruction, or −1 throughout memory.md §2, program.md §5
M, W, S, V memory, witness and setup columns; virtual tables gkr.md §2
memory form an Fr's Montgomery limbs x·R; on the wire only in FR_ARITH's frame primitives.md §1
mini-block the revm-block binary: transactions over a recorded pre-state ethereum.md §1
multiplicity a channel's W column counting each table row's lookups lookup.md §7
padding contract padding.row makes row-local relations vanish and tree inputs 1 gkr.md §4
pairing side G2One or G2X: an entry's G2 argument, [1]_2 or [x]_2 mercury.md §6
pass 1, pass 2 executing to commit every shard's M columns; again to prove each streaming.md §2
phase 1, phase 2 the decider key's ceremonies: powers of tau, then the circuit's own recursion.md §9
public window windows 2 and 3 at 2^12: input at 0x8000, journal at 0xC000 public-values.md §2
query one read and one write at one address in one cycle execution-trace.md §3
RAM glue invocations chained through their frame's words in RAM delegation-circuits.md §1
reconciliation ∏ read roots · R_b = ∏ write roots · W_b, once a statement memory.md §4
registry family_circuit, recursion_circuit: each family's one circuit circuits.md §1
scratch scratch[i], a flat relation's intermediate, one per inner column gkr.md §2
shard h rows of one family, or one window, proved alone but for the memory argument streaming.md §4
slot Δ in a cycle's timestamps 4c + Δ; a ProgramImage halfword; a frame position execution-trace.md §1, program.md §2, memory.md §2
SRS digest a digest of the SrsVerifier and the generic table's commitments proof.md §3
stack 2^σ columns committed as one, in the recursion format recursion.md §1
statement PublicInputs: input, journal, exit status and the execution's record proof.md §1
statement shard, shard-set exactness a (family, index) below its count; a block proves each once, in order proof.md §1
tamper twin a forgery proved as an honest prover would, refused in its class circuits.md §3
tape straight-line coprocessor calls a node replays; checker tape's listing recursion.md §7, tools.md §4
time window a shard's claimed [ts_start, ts_end); it binds nothing proof.md §8
transcript form a G1 point as four 128-bit Fr limbs; infinity, four 2^128 transcript.md §4
tuple T(AS, ADDR, TS, VAL): a memory access as one field element memory.md §1
u1, u2 a Mercury opening point's halves, pairing with an index's low and high bits mercury.md §1
VmConfig a program's families, their heights, bytecode_size_words program.md §7
window h words from byte 4h·w, initialized and torn down by one shard memory.md §3
write-side induction an execution family writes only words, so operands need no bound memory-ops.md §5

参考

工具

规范原文docs/tools.md以 Markdown 查看

摘要

围绕证明者与验证者的所有二进制程序,它们都不在证明路径上:用于测量和证明的 bench;周期分析器及其对委托候选项的分类与定价;证明者的调试日志及其标记;checker;artifact-dump;verifier 命令行工具;kat-gen 与已提交的测试数据(fixtures);以及工作空间之外的两个参照实现(reference oracle)。

以下规范原文以英文维护;英文是本规范的标准语言。

The binaries around the prover and verifier, none on a proof path: bench measures and proves (§1), profiler counts a guest's cycles (§2), a debug-info build logs a proving run (§3), checker validates circuits and the global transcript (§4), artifact-dump exports a guest's ProgramImage (§5), verifier checks a proof from files (§6), kat-gen regenerates the committed fixtures (§7), and two generators outside the workspace are reference oracles (§8).

1 bench#

cargo run --release -p bench [-- <routine>...]   every routine, or those named; --list lists them
cargo run --release -p bench -- prove <stem> | --stateless <file> [--case <name>]
    [--in-flight <n>] [--out <dir>] [--json <path>] [--hourly-usd <price>] [--toy-srs]

The routines time one component each, over their own data: fr-arith, poly-bind, msm, mercury, mercury-batch, zerocheck-prove, zerocheck-verify, gkr-prove. msm, mercury and mercury-batch run over ceremony bases, assets/ptau/ppot_0080_24.ptau, and return without them.

prove proves a block through host::prove (streaming.md) and verifies it (host::verify).

  • <stem> names a recorded block under crates/host/tests/vectors: its pin <stem>.json, to which <stem>-witness.bin and <stem>-journal.bin are held, names the guest that proves it; mini-block is committed (ethereum.md §6).
  • --stateless <file> is one input to revm-block-stateless. A .json EEST fixture gives its statelessInputBytes as the advice, unchanged, and its statelessOutputBytes as the journal the proof must bind, checked by revm_block::stateless::run first and on the proof after; --case picks one input by part of its name. Any other file is the raw input.
  • The guest is built at --release (host::fixture::build_revm_guest), decoded at host::fixture::revm_params and keyed over 2^22 ceremony powers or, with --toy-srs, over τ = 0xc0ffee, cached as apogee-bench-toy-22.srs in the temporary directory: the same timings, another identity, which the report names.
  • --in-flight is max_in_flight, 8 by default. The verb asserts that the guest exits 0 and the block verifies; --out then writes the proof archive (proof.md §9) under the stem's or the input file's name.

The printed BenchReport (--json writes it too) holds the block, identity, SRS, cycles per gas, shards per family, proof and statement bytes, clocks, peak RSS, cost and hardware. commit and gkr are pass 1's and pass 2's wall clocks; execution, the executor's time, runs inside them and is left out of their total; opening and final are 0; unattributed is the rest of the proving wall clock; setup and verify are apart. Peak RSS is Linux's VmHWM, absent elsewhere, where /usr/bin/time -l gives it. --hourly-usd adds the cost, price · proving_ms / 3,600,000, and the cost per Mgas. Any failure exits 1, a wrong journal or a failed --out after the report prints; a usage error exits 2.

The verbs recurse, recurse-node, ceremony and decide are recursion.md §8.4–§10's.

2 The cycle profiler#

cargo run --release -p profiler -- elf <file> [--advice <f>] [--input <f>] [<common>]
cargo run --release -p profiler -- block <stem> [<common>]
cargo run --release -p profiler -- record <number|latest> [--txs <n>] [--cache <dir>] [<common>]
    <common>: [--top <n>] [--json <path>]

elf runs any guest over the given input and advice, at the smallest menu height its code fits; block runs the revm guest over a recorded fixture; record records a block from ETH_RPC_URL (latest is the finalized one; every transaction unless --txs; cached in target/profiler-cache) and runs revm-block over it, its gas the transactions' limits capped at the block's. A run prints a table, the --top (30) functions in it, and with --json writes a ProfileReport; any error exits 2. Its numbers are counts of executed cycles, the same on any machine.

2.1 One histogram over pc#

profiler::profile adds 1 to one u64 per halfword slot of the image for each executed cycle, reading each chunk's pc column off emulator::StreamingRun and dropping the chunk, so it holds the histogram and one partial buffer per family. Delegation rows add nothing: their requesting cycle is the ecall row's. A function's cycles are the sum over its [st_value, st_value + st_size) (loader::function_symbols), its own and not its callees'; its calls are the count at its first instruction, which runs once a call, so code entered only past its entry shows cycles and no calls. A mnemonic's cycles are the sum over its slots, a category's over its functions', and the unattributed ones are at slots no symbol covers.

2.2 Classification#

tools/profiler/src/categories.rs puts each function in one of 14 categories by RULES, ordered substring rules where the first match wins, then FALLBACK_RULES, the generic runtime paths, each matched against the demangled path and the raw symbol (categories::classify). The order is the meaning: revm_interpreter::instructions::system::keccak256 is hashing because its rule comes before revm_interpreter::'s. Legacy mangling is decoded whole, v0 to its identifiers.

A function's cycles include what was inlined into it: ruint's 256-bit operations count in the EVM opcode handlers, each a symbol of its own, revm dispatching through a table of function pointers. The unattributed share and the mnemonic mix, which no symbol table can misattribute, are the checks on attribution.

2.3 Pricing a candidate#

removable = max(0, cycles − calls·(4 + 2·frame_words))

cycles is the category's, calls the entry counts of the candidate's named entry symbols, and 4 + 2·frame_words (categories::shim_cycles) the shim a delegation leaves: the frame's stores, the ecall, the results' loads. CANDIDATES prices secp256k1, 256-bit arithmetic, BN254, SHA-256/RIPEMD-160 and the keccak sponge; one without entry symbols is charged no shim and flagged. It is a ceiling: it charges nothing for the new family's shards (delegation.md §9) or for marshalling operands into a frame.

3 The proving debug log#

crates/prover/src/debug.rs and the prover's log lines exist only with its debug-info feature, the workspace's one cargo feature: off by default, enabling no dependency, changing no proof byte (crates/prover/tests/debug_info.rs proves one statement with the log off and at deep and compares the blocks). Without it dlog! and debug_only! expand to nothing, so no scan is compiled into a proving run. gkr::explain_self_check is compiled always.

cargo run --release -p bench --features prover/debug-info -- prove ...
cargo test --release -p prover --features debug-info --test <suite> -- --include-ignored
APOGEE_DEBUG=off | phase | detail | deep [:FAMILY,FAMILY]

APOGEE_DEBUG, read at each log site, picks the level, case ignored: unset or empty is phase, none and 0 also mean off, 1 to 3 the other levels. :FAMILY,… (names as the log prints them, or ids) keeps those families at the level and lowers the others one step; lines naming no family stay. A bad level falls back to phase, an unknown family is dropped, and either is reported once as apogee ERROR. Lines go to the raw io::stderr() handle, one locked write each: libtest shows captured eprintln! output only for a failed test, and an OOM kill, a hang or a SIGINT loses it.

level adds
phase identity and SRS digest in full, in to_bytes order as the verifier CLI takes them; each claim's take and its committed or proved; the global digest and memory challenges, on the apogee commit line; each shard's begin h= … gkr done and open begin … open done
detail each family's circuit inventory; each shard's time window, g, β, roots and opening commitments; gkr::self_check; the scans
deep each GKR layer's shape and bytes; the top layer's all-zero columns

Where a run died. A begin without its done names the shard that died (FAMILY#index, [k/N] its statement position); a take without committed or proved, one in flight. fill# is fill order, which picks the failure returned, and in_flight= below the bound mid-pass means the workers wait on the executor. fill_ms is the one-thread fill, ms a wall clock shared with the shards in flight. Every shard forks from the apogee commit line's values, so two runs that should agree diverge there or inside a shard.

self_check recomputes every gate on every row before the backward pass, a second forward pass (gkr.md §5). gkr::explain_self_check turns a failure into the row's first disagreeing gate and every operand's value, a committed column by its artifact name and an inner one by the relation that wrote it, where a verifier says only LayerInconsistency { layer }.

The scans read each base delegation shard's live rows: invocations against the height, cycle and frame-base ranges, timestamp gaps, selector and round histograms, and canonicity, a tally for POSEIDON2 and FR_ARITH, whose < p conclusions are gated to the rows that read a value, and a verdict for MOD_MUL's operands and the values each EC_ADD row's group reads (debug::ec_add_reads). On ADD_SUB_LUI_AUIPC they count requests per type, which sum to each delegation family's invocations, and exit rows, one in all. The log's verdicts:

marker
self_check FAILED a gate fails on the prover's own values
NOT CANONICAL a frame value at or above its modulus where a gate needs it below
UNBALANCED an EC_ADD curve whose three groups' counts differ
OVER the a timestamp gap beyond 38 bits
NAMES NO MODULUS a MOD_MUL selector naming no modulus
DISAGREES a SHA256_COMP frame its rounds do not produce: the fill's refusal, in every build
NOT LOOPING 24 TIMES KECCAK_F round counts more than 1 apart
ABORTED a nonzero exit status: the block proves a failed execution
OUTPUT-LAYOUT-BREAK outputs other than 2 + 2·channels: reduce_shard and channel_cones index channel roots from opposite ends
ALL ZERO a top-layer column all zero: a root of 0
DECLARED BUT NEVER INVOKED a delegation shard with no live row
console
$ APOGEE_DEBUG=detail <a debug-info run> 2>&1 | tee run.log
$ grep -c 'begin h=' run.log; grep -c 'gkr done' run.log    # unequal: a shard died
$ grep 'begin h=' run.log | tail -1
$ grep -E 'FAIL|NOT CANONICAL|UNBALANCED|OVER the|NAMES NO|DISAGREES|NOT LOOPING|ABORTED' run.log
$ grep -E 'LAYOUT-BREAK|ALL ZERO' run.log

At detail the self-check doubles each shard's forward work and the scans cost O(live rows × frame words); deep reads no layer's cells but the top's.

4 checker#

cargo run -p checker -- laws <artifact>       Laws 1–4, then the lookup rules (check_laws)
cargo run -p checker -- padding <artifact>    the padding contract (check_padding)
cargo run -p checker -- dump <artifact>       the circuit, readably (checker::dump)
cargo run -p checker -- tape <verifying-key> <public-inputs>

An artifact is a CircuitArtifact file, decoded for encoding only so that a lawless one reaches the checks, such as crates/constraints/tests/vectors/*.bin. The validators are circuits.md §3's, independent of constraints; padding omits the product-tree clause; dump prints any decodable artifact.

tape loads a key (verifier::load_verifying_key) and a PublicInputs file, an archive's .vk and .public, refuses a statement the key does not describe (verifier_core::derive_global_phase), and runs checker::check_global_tape. That renders the global commit phase's event log a line a message, absorb <TAG> <n> (n scalars, or a bytes message's 31-byte chunks) or squeeze <TAG>, and holds it to expected_global_tape: G1–G11 (proof.md §2) written from the statement's shape, sharing nothing with verifier_core::global_commit but statement_shards. It prints the tape or the first line out of order, and checks the script, not the values, which the log does not carry. checker exits 0 when a check holds or a listing prints, 1 naming the failure, 2 on a usage error.

5 artifact-dump#

cargo run -p artifact-dump -- <guest.elf> [--out <dir>]
cargo run --release -p artifact-dump -- tables <guest.elf> [--ptau <file>]

The first writes <name>.img, the ELF's ProgramImage in its postcard wire form with nothing around it (program.md §3), and <name>.img.txt, a report rendered from the image read back off those bytes, which must equal the loaded one or nothing is written: segments, the listing (address, length, encoding, expanded word), .symtab names marked as outside the artifact, and the artifact's and the ELF's SHA-256, which pin bytes and are not the program identity. <name> is the ELF's stem; --out defaults to the working directory.

tables prints the VmConfig and each instruction's pc, next_pc, family, mnemonic and decoded fields at ProgramParams::defaults(); with --ptau it reads 2^22 powers, the largest default height, and prints the program identity (program.md §8).

6 The verifier CLI#

cargo run --release -p verifier -- <verifying-key> <identity-hex> <public-inputs> <proof>...
cargo run --release -p verifier -- block <verifying-key> <identity-hex> <public-inputs> <block>

The key is loaded by verifier::load_verifying_key (proof.md §7) and its identity must equal <identity-hex>, 64 lowercase hex digits of its canonical bytes from a channel the prover does not control: never the key, the proof or an archive's .identity. The first form verifies each ShardProof file with verify_shard and requires the files to be the statement's shards, each once, in any order; the second verifies a BlockProof with verify_block, as an archive's .vk, .public and .block (proof.md §9). It takes no SRS digest, using the key file's (srs.md §3). Exit 0 when all verifies, 1 naming the first file refused or a wrong shard set, 2 on usage or a malformed identity.

7 kat-gen and the committed fixtures#

cargo run -p kat-gen regenerates the default groups, cargo run -p kat-gen -- <group> one. Each file written prints its SHA-256, which the tests reading it pin.

group writes from
field, poly, curve, tower, pairing, msm, srs, moduli arithmetic, ceremony-point, KZG and MOD_MUL modulus vectors arkworks; srs's points through its own .ptau reader
pcs G1 absorption limbs; Mercury proofs arkworks; pcs
loader, isa listings of the committed guest ELFs, synthetic ELFs; an RV32IMA corpus, words that must not decode the pinned toolchain's llvm-objdump, llvm-nm
program, tape program identities, the generic table's commitments; guests/shards' global tape program; checker
gkr, memory, lookup, family, delegation CircuitArtifact files: toy circuits; the four frames, the two RAM-window circuits and the seven execution circuits, at 2^22; each base delegation circuit's shape and SHA-256 constraints
revm a synthetic block's witness, output commitment and delegated keccak-f frames native revm, held to the guest
opt-in: block, zkevm, guests mini-block, over ETH_RPC_URL (ethereum.md §6); zkevm-subset.json, cut from the tests-zkevm release at APOGEE_ZKEVM_FIXTURES only if every pair matches; the guest ELFs, each built twice and compared

srs, and program's identities and table commitments, need assets/ptau/ppot_0080_24.ptau and are skipped without it. CI runs the default groups and both oracles (§8) and fails on any git diff in the vector directories. A guest ELF is not reproducible across machines, since rustc embeds absolute paths in the panic-location strings of core and of crates outside the guest workspace, which the guest build does not remap; two clean builds on one machine agree. So guests is run by hand on one machine, and CI regenerates only what derives from the ELFs.

8 Reference oracles#

cargo run --manifest-path tools/transcript-ref/Cargo.toml
cargo run --manifest-path tools/stateless-ref/Cargo.toml

tools/transcript-ref implements transcript.md from its text over Plonky3's Poseidon2 and HorizenLabs zkhash's round constants, pinned by revision, and writes crates/transcript/tests/vectors/: permutation vectors, transcript scripts and io_digest cases. tools/stateless-ref encodes stateless inputs with eth-act/ere-guests v0.17.1's stateless-validator-common over libssz 0.3.0 and writes stateless_ref.txt under crates/host/tests/vectors/: per input, its request's hash_tree_root or reject. Each is its own workspace root because its dependencies enable features, serde/std among them, that cargo's feature unification would carry into the workspace's no_std crates; the one repository crate either links is tools/test-support, a seeded RNG, SHA-256 and hex with no dependencies.

参考

仓库地图

远地虚拟机代码仓库中各部分的位置、每个 crate 的作用,以及定义它的规范页面。

以 Markdown 查看

远地虚拟机的代码仓库由两个 Cargo 工作空间组成:根工作空间包含在你的宿主机(host)上运行的一切,guests/ 包含在虚拟机内运行的一切。

Crate 一览#

路径 说明 规范出处
crates/constants 所有协议常量、标签和标识符;不含逻辑 使用各常量的页面
crates/field、curve、poly、sumcheck Fr;Fq 扩域塔、G1、G2、配对、MSM;多线性多项式;零校验(zerocheck) 原语
crates/transcript Poseidon2 与双工 transcript Transcript
crates/srs 仪式文件导入、SRS 归档、KZG、Groth16 的第一阶段 SRS
crates/pcs、pcs-verify Mercury 与延迟验证;pcs-verify 是验证中的域运算部分 Mercury
crates/loader、isa、program 从 ELF 到 ProgramImage;解码器;解码表、VmConfig、程序身份 程序与身份
crates/emulator、trace 执行器及其追踪器;行、内存状态、列构建器 执行轨迹
crates/constraints 以数据形式表示的全部电路:内存帧、查找通道、各电路族的电路、注册表 GKR 引擎、内存、查找、电路以及各电路族页面
crates/gkr-verify、gkr GKR 验证者与证明者 GKR 引擎
crates/verifier-core 陈述、transcript、验证密钥,以及除打开之外对分片和块(block)的全部检查;递归的 tape、节点与折叠 证明、递归
crates/verifier verify_shard、verify_block、证明归档、verifier CLI 证明
crates/prover 密钥构造、列填充、流式证明者、调试日志 流式证明者
crates/groth16 带绑定线和两阶段仪式的 Groth16 递归 §9
crates/host 宿主程序 SDK:设置、证明、验证;区块见证记录器;递归树与判定器 以太坊区块、递归
crates/checker 电路法则的独立校验器、原生查找与内存求值器、篡改测试套件、checker CLI 电路 §3
crates/guest-sdk 客户程序(guest)运行时:入口、链接脚本、分配器、内存区域、委托 shim 客户程序 ABI、委托 ABI
guests/ 用于测试和工作负载的客户程序,自成一个工作空间;vendor/ 存放打过补丁的上游 crate 示例客户程序
contracts/ ApogeeVerifier.sol 递归 §9
tools/ kat-gen、bench、profiler、artifact-dump、test-support;transcript-ref 与 stateless-ref,即工作空间之外的独立参照(oracle) 工具与 CLI
docs/ 架构概述、术语表、客户程序手册、工具页面,以及 spec/(每个主题一页) 本站

环境要求#

  • 工具链、其组件以及 riscv32imac-unknown-none-elf 目标都固定在 rust-toolchain.toml 中;rustup 会在首次使用时安装它们。
  • 计算程序身份、生成真实密钥和执行证明都需要仪式文件 assets/ptau/ppot_0080_24.ptau。工作空间的测试不需要。
  • 证明受内存限制:一个完整区块的内存峰值达到 174 GiB。

常用命令#

sh
# What CI runs
cargo fmt --all -- --check
cargo clippy --workspace --all-targets -- -D warnings
cargo test --workspace
cargo run -p kat-gen && git diff --exit-code     # committed fixtures regenerate identically

# Guests: their own workspace and target
(cd guests && cargo clippy --bins -- -D warnings)
(cd guests/fib && cargo build --target riscv32imac-unknown-none-elf)   # --release for proving

# Prove and verify a block, then recurse and decide
cargo run --release -p bench -- prove mini-block --out <dir>
cargo run --release -p verifier -- block <stem>.vk <identity-hex> <stem>.public <stem>.block
cargo run --release -p bench -- recurse <dir>/<stem> --out <out>

证明真实分片的测试套件都标记了 #[ignore],CI 不运行它们:每个套件都在自己的玩具 SRS 上做证明,需要数十 GiB 内存。

sh
cargo test --release -p prover --test <suite> -- --include-ignored --test-threads=1
#   acceptance, control, alu, mem, fills, block, streaming, keccak, recursion, public_io, revm
cargo test --release -p host --test prove -- --include-ignored --test-threads=1     # a mainnet mini-block
cargo test --release -p checker --test tamper -- --include-ignored --test-threads=1 # every tamper twin

参考

发布说明

远地虚拟机 v1.0.0,首个版本。它证明什么、包含什么、如何测量与检验,以及已知的限制。

以 Markdown 查看

v1.0.0#

远地虚拟机的首个版本:一台证明 RV32IMAC 程序、并把证明结算到以太坊上的 RISC-V zkVM。本站转载的规范,是代码仓库在源码修订版 3571370 时的 docs/。

它证明什么#

一个由其映像摘要所指明的程序,在某个公开输入上运行到某个退出状态,并写出了一份公开输出(journal);这份证明经由递归树汇成一份 Groth16 证明,由 ApogeeVerifier.sol 检验。证明是简洁的,但不是零知识的。

发布内容#

  • 机器。 单 hart 的 RV32IMAC;RV32IMA 的 59 条指令,压缩指令在加载时展开;客户程序(guest)SDK,为输入、证明者提示(advice)和输出提供三个内存区域。
  • 证明系统。 BN254 标量域上的 23 个电路族,每个都是分层 GKR 电路:七个指令电路族、五个内存窗口电路族、六个委托和五个递归电路族。覆盖整个执行过程的单一读/写内存多重集;五个通道上的 LogUp 查找(lookup)。
  • 委托。 KECCAK_F、SHA256_COMP、POSEIDON2、FR_ARITH、MOD_MUL 和 EC_ADD,可从 SDK 调用,也可从打过补丁的 k256、ark-ff 和 revm-precompile 调用。
  • 承诺。 基于 KZG 的 Mercury,建立在 PSE 的 perpetual powers of tau 之上,每个分片一个 704 字节的打开证明;为递归提供延迟验证。
  • 证明者。 两遍式流式证明者,内存占用取决于同时处理中的分片。
  • 结算。 由叶程序和节点程序组成的递归树,采用带域内存和四个协处理器的递归格式;带绑定线和两阶段仪式的 Groth16 判定器;ApogeeVerifier.sol。
  • 以太坊工作负载。 一个 revm 客户程序,包含一个 mini-block 二进制程序,以及支持 Osaka、BPO1、BPO2 和 Amsterdam 的无状态校验器。
  • 工具。 bench、周期分析器、证明者的调试日志、checker、artifact-dump、verifier CLI、kat-gen,以及两个参照实现(reference oracle)。
  • 不依赖外部密码学库。 域、曲线、配对、MSM、哈希、PCS、GKR 和 Groth16 都在仓库内实现。

实测数据#

glamsterdam-devnet-8 的第 257,510 号区块(60 笔交易,101.5 Mgas,198M 个周期):包含 207 个分片的基础证明,在 32 vCPU 上耗时 2,481 s,内存峰值 174 GiB;包含 116 个分片的递归树;耗时 18.5 s 的判定器证明;链上验证消耗 3,620,026 gas。全部 67,251 个 tests-zkevm v21.0.1 测试对在原生运行下一致。性能页面列出了每一个数据。

已知限制#

不是零知识的;证明者提示在设计上不受绑定;公开输入和公开输出各至多 16,380 字节;sc.w 总是成功;陷入(trap)不可证明;基础委托固定为六种;证明者的内存由同时处理中的分片决定;判定器密钥随根的形状而定,其可信程度取决于它的仪式。安全模型列出了每一项限制及其原因。

文档#

本站提供英文、法文(加拿大)、简体中文和德文版本,各语言版本中的规范原文均为英文。AI 随行手册和 llms.txt 面向 AI 智能体。

404

此页面已偏离轨道

这个地址下没有页面。你可以搜索文档,或者从「迈出第一步」重新出发。