Apogee VMDoku v1.0.0
gweb3networks.com ↗
Titelbild von Apogee VM: das Apogee-Emblem, das über dem Horizont eines Planeten aufsteigt, mit dem Schriftzug „Higher Compute Horizons“

Apogee VM·v1.0.0·Der erste Schritt

Jede Anwendung kann Blockchain-nativ sein.

Apogee VM beweist, dass ein Programm korrekt ausgeführt wurde. Sie schreiben gewöhnliches Rust. Apogee führt es auf einer RISC-V-Maschine aus, beweist jeden ausgeführten Befehl und übergibt der Chain einen Beweis, den ein einziger Contract-Aufruf prüfen kann. Korrektheit ist dann keine Vertrauenssache Ihrer Nutzer mehr, sondern etwas, das sie selbst nachprüfen.

BeweisaussageGeprüft

Programm
Ein Körperelement, die identity: ein Digest über den Code, das initiale Speicherabbild, den Einsprungpunkt und die Konfiguration der Schaltkreise.
Eingabe
Die öffentlichen Bytes, die das Programm erhalten hat.
Ausgabe
Das journal: die Bytes, die es zur Veröffentlichung bestimmt hat.
Exit
Der Status, mit dem es endete. 0 bedeutet Erfolg.

Groth16 · BN254ein Contract-Aufruf

Proof of Correctness

Software hat immer verlangt, dass man ihr vertraut. Apogee macht sie stattdessen überprüfbar.

Wann immer ein Programm auf dem Rechner eines anderen läuft, nehmen seine Nutzer das Ergebnis auf Treu und Glauben hin: das Hauptbuch, das Orderbuch, die Auszahlung, den Würfelwurf. Blockchains haben diesen Glauben für eine eng umrissene Art von Programm überflüssig gemacht, indem jeder Knoten jede Transaktion erneut ausführt. Das funktioniert. Es ist zugleich die teuerste Methode, die je ersonnen wurde, um sich auf irgendetwas zu einigen.

Eine zkVM, eine virtuelle Maschine, die ihre eigene Ausführung beweist, macht diesen Glauben für jedes Programm überflüssig. Das Programm läuft einmal, an einem beliebigen Ort. Es hinterlässt eine mathematische Quittung, die besagt, dass dieses Programm mit dieser Eingabe diese Ausgabe erzeugt hat, und wer diese Quittung prüft, muss das Programm nie erneut ausführen. Diese Quittung ist es, die eine Anwendung Blockchain-nativ macht: Die Regeln der Anwendung stehen im Code, ihr Zustand liegt als Commitment on-chain, und jede Änderung dieses Zustands trifft mit ihrem Beweis ein.

Die dritte Ära

Erst Arbeit, dann Einsatz, dann Korrektheit.

Jede Ära der Blockchains hat einen neuen Weg gefunden, auf Vertrauen in jemanden zu verzichten. Proof of Correctness (Korrektheitsbeweis) ist der erste, der bis in die Berechnung selbst hineinreicht.

01

Energie

Proof of Work

Strom sichert die Reihenfolge der Ereignisse. Wer die Geschichte umschreiben will, muss mehr ausgeben als die ehrliche Mehrheit für ihre Stromrechnung.

02

Kapital

Proof of Stake

Kapital sichert die Reihenfolge der Ereignisse. Fehlverhalten wird bestraft, indem der Einsatz verbrannt wird, der dafür gebürgt hat.

03

Mathematik · jetzt

Proof of Correctness

Mathematik sichert die Ereignisse selbst. Jede Zustandsänderung trägt einen Beweis, dass sie von dem Programm berechnet wurde, auf das sich alle geeinigt haben.

Arbeit und Einsatz entscheiden, welche Geschichte gilt. Keines von beiden prüft, was in ihr geschehen ist; diese Aufgabe fiel seit jeher jedem Knoten zu, der alles erneut ausführt. Ein Gültigkeitsbeweis macht diese letzte Brute Force überflüssig. Energie, dann Kapital, dann Mathematik: Es bleibt kein Viertes, dem man noch vertrauen müsste.

Was sich ändert

Ihr Produkt, seine eigene Chain und die Mathematik als Schiedsrichter.

Die Blockchain-native Zukunft ist nicht eine einzige Chain, die alles erledigt. Sie besteht aus vielen Umgebungen, jede um eine Anwendung herum geformt, und alle nutzen dieselbe Basisschicht zur Abwicklung. Apogee ist die Beweis-Engine, die den Betrieb einer solchen Umgebung praktikabel macht.

Schreiben

Ihre Logik, in gewöhnlichem Rust

Ein Gastprogramm (Guest) ist ein no_std-Rust-Binary für RISC-V. Es liest seine Eingabe, erledigt seine Arbeit und schreibt seine Ausgabe fest. Apogee beweist jeden Lauf, und Sie müssen nie in Schaltkreisen denken.

Abwickeln

Finalität ohne Wartezimmer

Ein Gültigkeitsbeweis ist final, sobald er verifiziert ist. Es gibt keine siebentägige Einspruchsfrist, die man aussitzen muss, und weder ein Komitee noch eine Enklave, die an die Stelle der Mathematik tritt.

Spezialisieren

Eine Umgebung, zugeschnitten auf Ihr Produkt

Eine Anwendung pro Rollup bündelt jede Ressource auf den einen Kanal, auf den es für sie ankommt. Apogee beweist jedes Programm, das für seine Maschine gebaut ist; die Funktion des Zustandsübergangs legen also Sie fest.

Messung statt Zusage

Ein vollständiger Ethereum-Block, vom Gastprogramm bis zum Contract.

Mit Apogee v1.0.0 wurde Block 257.510 von glamsterdam-devnet-8 durchgängig bewiesen. Der Block wurde innerhalb der VM nach den Regeln der execution-specs zustandslos validiert und dann durch Rekursion zu einem einzigen Beweis gefaltet, den ein Ethereum-Contract akzeptiert.

101,5 MgasEin Block, 60 TransaktionenAusgeführt im zustandslosen Validator, innerhalb der VM.
198 Mio.RISC-V-ZyklenJeder ausgeführte Befehl ist eine bewiesene Zeile, verteilt auf 207 Shards.
1Beweis an der SpitzeEin Rekursionsbaum aus 116 Shards, gefaltet zu einem einzigen Groth16-Beweis.
3,62 MgasOn-Chain-VerifikationEin Aufruf von ApogeeVerifier.sol mit 34.980 Byte Calldata.
704 BPro Shard-ÖffnungEin einziger Mercury-Beweis öffnet jede committete Spalte eines Shards.
23SchaltkreisfamilienSieben für Befehle, fünf für Speicherfenster, sechs Delegationen, fünf für die Rekursion.
67.251KonformitätstestfälleJedes Paar aus tests-zkevm v21.0.1, vom Validator nativ reproduziert.
0Externe KryptografieKörper, Kurve, Pairing, MSM, Hash, PCS, GKR und Groth16 sind im Repository selbst geschrieben. Externe Bibliotheken dienen nur als Testorakel.

Der Basisbeweis dauerte 2.481 s auf einer Maschine mit 32 vCPUs, bei einem Spitzenverbrauch von 174 GiB; der Rekursionsbaum brauchte etwa 2.620 s zusätzlich. Die Beweiserzeugung ist heute speichergebunden, und die Seite zur Performance nennt jede Zahl mit ihrer Quelle.

Der Weg eines Beweises

Sie schreiben das Programm. Apogee erledigt alles danach.

Zwischen Ihrem Rust und dem true des Contracts liegen ein Decoder, 23 Schaltkreisfamilien, die GKR-Engine, die Commitments, ein Rekursionsbaum und ein Groth16-Decider. Nichts davon müssen Sie bauen oder warten.

Der Weg eines Beweises Acht Schritte: Schreiben, Laden, Ausführen, Sharden, Beweisen, Rekursion, Entscheiden, Verifizieren. Den ersten übernimmt der Entwickler, die nächsten sechs Apogee, den letzten die Chain. Unter jedem Schritt steht die Größe dessen, was an diesem Punkt für Block 257.510 vorliegt. SIE SCHREIBEN APOGEE BEWEIST DIE CHAIN PRÜFT Schreibenno_std Rust LadenImage · Identität AusführenRV32IMAC · 1 hart Sharden23 Familien BeweisenGKR · Mercury RekursionBlätter → Wurzel EntscheidenGroth16 · BN254 VerifizierenApogeeVerifier.sol Ihr Quellcode32-Byte-Identität198 Mio. Zyklen207 Shards14,5 MB Beweis1,03 MB Wurzel34.980 B Calldatatrue · 3,62 Mgas
Der Weg eines Beweises. Alles zwischen den beiden äußeren Streifen ist Sache von Apogee. Die Zahlen unter jedem Schritt stammen von Block 257.510, aus den Messungen der Spezifikation.

Blockchain-nativ

Der Stack, den Sie schon kennen, mit einem Austausch pro Schicht.

Blockchain-nativ zu werden heißt nicht, eine neue Disziplin zu lernen. Jede Komponente einer herkömmlichen Anwendung hat ein Gegenstück, und das mentale Modell überträgt sich fast unverändert.

SchichtHerkömmliche AnwendungBlockchain-nativ, auf Apogee
GeschäftslogikEin Dienst auf Servern, die Sie betreibenEin Gastprogramm in Rust, bei jedem Lauf bewiesen
DatenbankSQL oder ein Key-Value-StoreDaten off-chain, eine Zustandswurzel on-chain
AbfrageSELECT … WHERE key = ?Ein Inklusionsbeweis, gegen die Wurzel geprüft
CommitCOMMITEine neue Wurzel, mit ihrem Beweis veröffentlicht
Audit-TrailLogs, die man Ihnen glauben mussEin Beweis, den jeder prüfen kann

Jede Schicht, mit einem ausgearbeiteten Beispiel →

Hier beginnen

Sechs Zugänge.

Dasselbe System, aus sechs Richtungen gelesen. Wählen Sie die Richtung, die zu der Frage passt, mit der Sie gekommen sind.

Das Apogee-Emblem: zwei Klingen, die sich über einem Planeten in einer Spitze treffen, mit einem Stern in ihrer Mitte

Der erste Schritt ist ein Programm.

Schreiben Sie es in Rust und führen Sie es auf Apogee aus. Alles Weitere, von den Shards und Schaltkreisen bis zur Rekursion und zum Contract, ist Aufgabe der Maschine. Was die Chain erreicht, ist ein Beweis, und ein Beweis ist alles, was die Chain braucht.

Höhere Rechenhorizonte

Der erste Schritt

Blockchain-nativ

Eine Blockchain-native Anwendung besteht aus denselben Komponenten wie die Anwendung, die Sie heute bauen, mit einem Austausch auf jeder Schicht. Hier finden Sie jeden Austausch und ein Hauptbuch, auf beide Arten gebaut.

Als Markdown anzeigen

Eine Blockchain-native Anwendung ist wirtschaftliche Aktivität, deren Abwicklung, Verwahrung und Regeln von Grund auf on-chain liegen, und kein herkömmliches Geschäft, an das seitlich ein Token angeheftet ist. Das klingt nach einer anderen Art von Engineering. Der Unterschied ist kleiner, als es klingt.

Jede Komponente des Stacks, den Sie heute bauen, hat ein Gegenstück. Das Gegenstück erfüllt dieselbe Aufgabe mit einer Änderung: Worauf man bisher vertrauen musste, das wird jetzt bewiesen. Apogee existiert, um diese Änderung so günstig zu machen, dass sie zum Standard wird.

Der Wandel in einem Satz#

In einer herkömmlichen Anwendung ist der Server die Autorität: Er hält die Daten, wendet die Regeln an und meldet das Ergebnis. In einer Blockchain-nativen Anwendung hält die Chain ein Commitment auf die Daten, die Regeln sind ein Programm, das jeder über seinen Digest benennen kann, und ein Ergebnis wird nur mit einem Beweis akzeptiert, dass dieses Programm es erzeugt hat.

Der Betreiber verschwindet nicht. Jemand führt weiterhin das Programm aus, speichert die Daten und beantwortet Anfragen. Was verschwindet, ist die Notwendigkeit, ihm zu glauben.

Schicht für Schicht#

Schicht Herkömmliche Anwendung Blockchain-nativ, auf Apogee Was bleibt
Geschäftslogik Ein Dienst, den Sie auf selbst betriebenen Servern bereitstellen Ein Gastprogramm (Guest): no_std-Rust, nach RISC-V kompiliert und bei jedem Lauf bewiesen Sie schreiben weiterhin Funktionen über Daten. Das Programm wird über seine Identität benannt, einen Digest seines Codes und seiner Konfiguration.
Datenspeicher SQL-Tabellen, ein Key-Value-Store Die Daten bleiben off-chain; die Chain speichert eine Zustandswurzel, einen einzigen Hash, der einen Snapshot aller Daten zusammenfasst Ein Snapshot, den Sie mit 32 Byte benennen und gegen den Sie alles prüfen können.
Leseabfrage SELECT balance FROM accounts WHERE id = ? Ein Inklusionsbeweis, ein Merkle-Pfad, den das Gastprogramm gegen die Wurzel prüft Eine Abfrage liefert weiterhin eine Zeile. Die Zeile kommt jetzt mit einem Nachweis, und das Gastprogramm weist jede Zeile zurück, deren Prüfung fehlschlägt.
Schreiben UPDATE …; COMMIT; Ein Zustandsübergang: Das Gastprogramm berechnet die neue Wurzel und veröffentlicht sie Commit heißt weiterhin „dauerhaft machen“. Jetzt bedeutet es, dass ein Contract die gespeicherte Wurzel fortschreibt.
Anfrage Der Body einer HTTP-Anfrage Die öffentliche Eingabe, die der Beweis bindet Eingaben hinein, Ausgaben heraus.
Große Nutzdaten Uploads, per Join verknüpfte Zeilen, abgerufene Dokumente Hilfsdaten (Advice): Bytes, die der Prover liefert und die das Gastprogramm gegen etwas prüft, das der Beweis bindet Große Daten per Referenz übergeben und prüfen, was angekommen ist.
Antwort Ein JSON-Body Das Journal: die öffentliche Ausgabe, durch den Beweis gebunden Jeder kann die Antwort lesen und weiß, dass sie vom Programm stammt.
Authentifizierung Sessions, Tokens, Passwörter Signaturen, im Gastprogramm verifiziert; die secp256k1-Recovery läuft auf delegierter Körper- und Kurvenarithmetik Identität ist ein Schlüssel, und Autorisierung ist eine Prüfung, die Sie im Quellcode nachlesen können.
Kryptografie-Bibliotheken sha2, ring, OpenSSL guest_sdk::keccak256, sha256, ec_add, jeweils an einen eigenen Schaltkreis weitergeleitet Dieselben Aufrufe, für einen Bruchteil der Zyklen.
Release Ein Binary ausrollen, und das Verhalten ändert sich sofort Die neue Programmidentität beim Verifier-Contract registrieren Releases werden explizit: Ein neuer Build ist eine neue Identität, die der Contract akzeptieren muss.
Skalierung Mehr Server, geshardete Datenbanken Eine Ausführung, zerlegt in Shards, die parallel bewiesen werden, und durch Rekursion zu einem einzigen Beweis gefaltet Der Durchsatz entsteht durch Prover, die nebeneinander arbeiten, während die Chain weiterhin einen einzigen Beweis prüft.
Audit Logs und Bescheinigungen, die man Ihnen glauben muss Der Beweis und sein Journal Gewissheit beruht nicht mehr auf Reputation, sondern auf Verifikation.

Die mittlere Spalte ist das, woraus eine Blockchain-native Anwendung besteht. Apogee liefert die Maschinerie darunter: die RISC-V-Maschine, die Schaltkreise, die Commitments, die Rekursion und den Verifier-Contract. Nichts davon taucht in Ihrem Programm auf.

Was sich nicht ändert#

  • Sie schreiben weiterhin gewöhnliches Rust. Structs, Enums, Traits, Iteratoren, Vec, BTreeMap und jedes Crate, das ohne std baut. Es gibt keine Schaltkreissprache zu lernen.
  • Sie testen weiterhin auf Ihrem Laptop. Der übliche Aufbau legt die Anwendungslogik in eine no_std-Bibliothek, die auf dem Host unter cargo test genau so läuft wie im Gastprogramm. Siehe Ein Gastprogramm schreiben.
  • Sie denken weiterhin in Zustand, Anfragen und Antworten. Die Formen bleiben gleich; nur ihre Garantien ändern sich.
  • Deterministischer Code bleibt deterministisch. Guter Backend-Code vermeidet ohnehin versteckte Eingaben. Die VM macht daraus eine absolute Regel.

Was sich ändert#

  • Keine Außenwelt. Ein Gastprogramm hat keine Uhr, keinen Zufall, kein Netzwerk und keine Dateien. Alles, was es weiß, kommt als öffentliche Eingabe oder als Hilfsdaten an, und eine Anfrage nach Daten des Hosts ist ein Aufruf, den kein Beweis zulässt.
  • Jeder Befehl hat einen Preis. Jeder ausgeführte Befehl wird zu einer bewiesenen Zeile. Kopien, Allokationen und Leerlaufschleifen kosten Beweiszeit; die alte Disziplin des Zyklenzählens kehrt also zurück.
  • Gelieferte Daten werden geprüft, nicht geglaubt. Hilfsdaten wählt der Prover. Ein Gastprogramm prüft sie gegen etwas, das der Beweis bindet, bevor irgendetwas daraus Abgeleitetes veröffentlicht wird.
  • Ausgaben sind klein und öffentlich. Das Journal fasst höchstens 16.380 Byte. Ein großes Ergebnis wird als Digest veröffentlicht.
  • Nichts ist verborgen. Beweise von Apogee v1.0.0 sind succinct, aber nicht zero-knowledge. Ein Gastprogramm darf keine Geheimnisse enthalten.

Ein Hauptbuch, auf beide Arten gebaut#

Eine Einzahlung auf einen Kontostand: die kleinste Zustandsänderung, die einen Beweis lohnt.

Die herkömmliche Version#

sql
BEGIN;
SELECT balance FROM accounts WHERE id = $1 FOR UPDATE;    -- read
UPDATE accounts SET balance = balance + $2 WHERE id = $1;  -- write
COMMIT;                                                     -- make it durable

Die Nutzer vertrauen darauf, dass der Betreiber genau dies ausgeführt hat, gegen die echte Tabelle, und dass er das Ergebnis ehrlich meldet.

Die Blockchain-native Version#

Die Konten liegen in einem binären Merkle-Baum, dessen Blätter keccak256(account ‖ balance) sind. Ein Contract speichert die Wurzel. Das Gastprogramm erhält die alte Wurzel und die Anfrage als öffentliche Eingabe, den Kontostand und den Merkle-Pfad des Kontos als Hilfsdaten, prüft den Pfad und veröffentlicht die alte und die neue Wurzel.

guests/ledger/src/main.rsrust
#![no_std]
#![no_main]

guest_sdk::entry!(main);

const DEPTH: usize = 20; // room for 2^20 accounts

/// A leaf commits to one account's balance.
fn leaf(account: &[u8; 20], balance: u64) -> [u8; 32] {
    let mut bytes = [0u8; 28];
    bytes[..20].copy_from_slice(account);
    bytes[20..].copy_from_slice(&balance.to_le_bytes());
    guest_sdk::keccak256(&bytes)
}

/// Fold a leaf up its Merkle path; bit `level` of `index` says whether the
/// node is a right child at that level.
fn root_of(mut node: [u8; 32], index: u32, path: &[[u8; 32]; DEPTH]) -> [u8; 32] {
    let mut pair = [0u8; 64];
    for (level, sibling) in path.iter().enumerate() {
        let (left, right) = if (index >> level) & 1 == 0 { (&node, sibling) } else { (sibling, &node) };
        pair[..32].copy_from_slice(left);
        pair[32..].copy_from_slice(right);
        node = guest_sdk::keccak256(&pair);
    }
    node
}

/// Public input: old_root (32) ‖ account (20) ‖ amount (8, LE)
/// Advice:       balance (8, LE) ‖ index (4, LE) ‖ path (DEPTH × 32)
/// Journal:      old_root ‖ new_root ‖ account ‖ amount
fn main() {
    let input = guest_sdk::public_input();
    let advice = guest_sdk::advice();
    if input.len() != 60 || advice.len() != 12 + 32 * DEPTH {
        guest_sdk::exit(1);
    }
    let old_root: [u8; 32] = input[..32].try_into().unwrap();
    let account: [u8; 20] = input[32..52].try_into().unwrap();
    let amount = u64::from_le_bytes(input[52..60].try_into().unwrap());

    let balance = u64::from_le_bytes(advice[..8].try_into().unwrap());
    let index = u32::from_le_bytes(advice[8..12].try_into().unwrap());
    let mut path = [[0u8; 32]; DEPTH];
    for (i, sibling) in path.iter_mut().enumerate() {
        sibling.copy_from_slice(&advice[12 + 32 * i..12 + 32 * (i + 1)]);
    }

    // The query: the balance the prover supplied is the one the root commits to.
    if root_of(leaf(&account, balance), index, &path) != old_root {
        guest_sdk::exit(2);
    }
    // The write: the same path with the new leaf gives the new root.
    let Some(new_balance) = balance.checked_add(amount) else { guest_sdk::exit(3) };
    let new_root = root_of(leaf(&account, new_balance), index, &path);

    // The commit: publish the transition for the contract to apply.
    guest_sdk::commit(&old_root);
    guest_sdk::commit(&new_root);
    guest_sdk::commit(&account);
    guest_sdk::commit(&amount.to_le_bytes());
}

Lesen Sie es neben dem SQL. Aus SELECT … FOR UPDATE wurde ein Merkle-Pfad, der gegen die Wurzel geprüft wird. Aus UPDATE wurde ein neues Blatt auf demselben Pfad. Aus COMMIT wurden vier Aufrufe von commit, die das Journal schreiben, das der Beweis binden wird. Der Kontostand kam vom Prover, und das ist in Ordnung: Ein Kontostand, den die Wurzel nicht festlegt, scheitert an der Prüfung, und der Lauf endet mit Exit-Status 2.

Der Contract, dem die Wurzel gehört, akzeptiert einen Übergang nur mit einem Beweis, dass dieses Programm ihn erzeugt hat und mit Exit-Status 0 endete:

Ledger.sol (Skizze)solidity
interface IApogeeVerifier {
    function verify(bytes calldata input, bytes calldata output, uint256 exitStatus,
                    uint256[10] calldata proof, uint256[] calldata points) external view returns (bool);
}

contract Ledger {
    IApogeeVerifier public immutable verifier;
    bytes32 public root;

    constructor(IApogeeVerifier v, bytes32 genesis) { verifier = v; root = genesis; }

    function apply(bytes calldata input, bytes calldata journal,
                   uint256[10] calldata proof, uint256[] calldata points) external {
        require(verifier.verify(input, journal, 0, proof, points), "proof");
        require(bytes32(journal[0:32]) == root, "stale root");
        root = bytes32(journal[32:64]);
    }
}

Hinweis

Dies ist eine Skizze, um die Form zu zeigen, kein Contract für den Produktivbetrieb. Ein echtes Deployment verarbeitet einen Stapel von Anfragen pro Beweis, den das Gastprogramm zu einem einzigen Übergang zusammenfasst, und legt den Verifier auf das richtige Programm und die richtigen Längen der öffentlichen Werte fest. On-Chain abwickeln behandelt den bereitgestellten Verifier, seinen Schlüssel und seine Zeremonie.

Das Modell in drei Zeilen#

  1. Die Chain hält eine Wurzel.
  2. Das Gastprogramm beweist den Übergang.
  3. Der Contract schreibt die Wurzel fort.

Alles andere, von den Shards und Schaltkreisen bis zur Rekursion und zum Decider, übernimmt Apogee. Das ist die Abstraktion: ein Programm, seine Eingabe und seine Ausgabe sowie ein Beweis, der sie miteinander verbindet.

Nächste Schritte#

Der erste Schritt

Apogee auf einen Blick

Die Fakten auf einer Seite. Was Apogee VM beweist, wie es das beweist, was das kostet, was es voraussetzt und wo Version 1.0.0 endet.

Als Markdown anzeigen

In einem Absatz#

Apogee VM ist eine RISC-V-zkVM. Die VM beweist, dass ein RV32IMAC-Programm, benannt durch einen Digest seines Images, mit einer gegebenen öffentlichen Eingabe bis zu einem Exit-Status gelaufen ist und eine gegebene öffentliche Ausgabe geschrieben hat. Diesen Beweis führt sie über einen Rekursionsbaum zu einem einzigen Groth16-Beweis, den ein Ethereum-Contract prüft. Jeder Schaltkreis ist ein geschichteter GKR-Schaltkreis über dem Skalarkörper von BN254, jede committete Spalte wird mit Mercury geöffnet, und jede Challenge stammt aus einem Poseidon2-Transkript. Körper, Kurve, Pairing, MSM, Hash, Polynom-Commitment, GKR-Prover und Groth16 sind sämtlich im Repository implementiert. Die Referenz-Arbeitslast ist die Validierung von Ethereum-Blöcken.

Die Fakten#

Apogee VM v1.0.0
Was ein Beweis aussagt Dass das Programm mit dieser Identität, gestartet an seinem Einsprungpunkt auf seinem Image, mit dieser öffentlichen Eingabe und irgendwelchen Hilfsdaten (Advice) Befehl für Befehl bis zu EXIT mit diesem Status ausgeführt wurde und dabei dieses Journal geschrieben hat
Befehlssatz RV32IMAC auf einem Hart: die 59 Befehle von RV32IMA (40 Basis, 8 M, 11 A), komprimierte Befehle werden beim Laden expandiert
Sprache der Gastprogramme (Guests) Rust, #![no_std] mit alloc, stable 1.96.1, Target riscv32imac-unknown-none-elf
Arithmetisierung 23 Schaltkreisfamilien, jede ein geschichteter GKR-Schaltkreis: 7 für Befehle, 5 für Speicherfenster, 6 Delegationen, 5 für die Rekursion
Argumente Gates per Sumcheck; Speicher über eine einzige Lese-/Schreib-Multimenge für die gesamte Ausführung; Lookups per LogUp
Körper Der Skalarkörper von BN254, 254 Bit
Commitments Mercury, multilinear über KZG, eine 704-Byte-Öffnung pro Shard, unabhängig von der Spaltenzahl
Setup Die Perpetual Powers of Tau der PSE, Beitrag 80; eine zweite, schaltkreisspezifische Zeremonie für den On-Chain-Decider
Transkript Ein Poseidon2-Duplex-Sponge über Fr, Breite 3, Rate 2
Abwicklung Rekursionsbaum → Groth16-Decider → ApogeeVerifier.sol
Sicherheitsniveau Etwa 100 Bit, bestimmt durch BN254
Zero-Knowledge Nein. Beweise sind succinct, aber nicht zero-knowledge, und nichts wird verblindet
Delegierte Operationen Runden von keccak-f[1600], Runden von SHA-256, Poseidon2, Fr-Arithmetik von BN254, modulare 256-Bit-Multiplikation über vier Ethereum-Moduln, vollständige Punktaddition auf secp256k1 und BN254 G1
Öffentliche Werte Höchstens 16.380 Byte Eingabe und 16.380 Byte Journal; Hilfsdaten bis 2 GiB
Ausführungslänge Bis zu 2^36 − 1 Zyklen
Codegröße .text bis 7,94 MiB bei einer Tabellenhöhe von 2^22; Image standardmäßig bis 4 MiB
Kryptografie von Drittanbietern Keine auf einem Beweispfad. arkworks, Plonky3 und zkhash kommen nur als Testorakel vor

Gemessen#

Alle Zahlen beziehen sich auf Block 257.510 von glamsterdam-devnet-8, ausgeführt im Gastprogramm des zustandslosen Validators: 60 Transaktionen, 101,5 Mgas, 198 Mio. Zyklen. Quellen: Rekursion §10 und Streaming §1 der Spezifikation.

Stufe Ergebnis
Basisbeweis 207 Shards, 14,5 MB, 2.481 s auf 32 vCPUs und 247,7 GiB, Spitzen-RSS 173,92 GiB
Rekursionsbaum 116 Shards: vier Blätter über je höchstens 64 Basis-Shards (zusammen 2.157 s, Spitze 92 GiB) und eine Wurzel (460 s, 1,03 MB)
Decider-Schaltkreis 7.896.686 Constraints über einer Domäne der Größe 2^23
Decider-Beweis 18,5 s und 6,1 GB auf einem Laptop mit 18 Kernen, Schlüssel in 1 s eingelesen
On-Chain-Verifikation 3.620.026 gas, 34.980 Byte Calldata, 358 vom Contract gefaltete Punkte
Konformität Alle 67.251 Paare von tests-zkevm v21.0.1 stimmen nativ überein

Was ein Verifier besitzen muss#

Zwei Werte, bezogen über einen Kanal, den der Prover nicht kontrolliert:

  • Die Programmidentität, ein Körperelement. Gegenüber einer vom Prover gelieferten Identität zeigt ein Beweis nur, dass irgendein Programm gelaufen ist.
  • Der SRS-Digest der Zeremonie. Ein Schlüssel, der über einem bekannten τ gebaut wurde, wird nur durch diesen Vergleich zurückgewiesen.

Der Verifikationsschlüssel selbst darf von beliebiger Seite kommen: Beim Laden werden beide Werte aus seinem eigenen Inhalt neu berechnet, und seine Schaltkreise werden mit der Registry des Verifiers abgeglichen. Das Sicherheitsmodell enthält die vollständige Liste der Annahmen.

Wo v1.0.0 endet#

  • Nicht zero-knowledge. Keine Verblindung in Mercury, GKR oder dem Decider.
  • Hilfsdaten sind ungebunden. Ein Gastprogramm prüft sie gegen etwas, das ein Beweis bindet.
  • Traps sind nicht beweisbar. Ein nicht ausgerichteter Zugriff, ein Zugriff außerhalb des gemappten Speichers, ebreak oder ein pc ohne Befehl beendet den Lauf ohne Beweis.
  • sc.w gelingt immer. Die einzige Abweichung von RV32IMAC: Es gibt keinen Reservierungszustand.
  • Die Delegationen bilden eine feste Menge von sechs. EVM-MULMOD mit beliebigem Modul, MODEXP und BLS12-381 laufen als gewöhnliche Befehle.
  • Die Beweiserzeugung ist speichergebunden. Der gemessene Block erreichte eine Spitze von 174 GiB; der Speicherbedarf folgt den gerade bearbeiteten Shards, nicht der Länge des Laufs.
  • Der Decider-Schlüssel gilt je Wurzelform und ist nur so vertrauenswürdig wie seine Zeremonie. Der Entwicklungsschlüssel ist fälschbar.

Wer es baut#

Apogee VM ist das Flaggschiff des Forschungsprogramms von G Web3 für Blockchain-native Anwendungsumgebungen: eine optimierte Umgebung pro wirtschaftlicher Anwendung, deren Abwicklung jeweils per Gültigkeitsbeweis auf Ethereum erfolgt. Die Position des Programms ist in der These dargelegt; wohin die nächste Version geht, beschreibt Quantensprung.

Ihre App starten

Ihre App starten

Das Entwicklerhandbuch für Apogee VM. Wie ein Gastprogramm geschrieben, gebaut, ausgeführt, bewiesen und on-chain abgewickelt wird, und die Gewohnheiten, die es korrekt, beweisbar und günstig halten.

Als Markdown anzeigen

Ein Gastprogramm (Guest) ist das Programm, das Apogee beweist: ein no_std-Rust-Binary, kompiliert für riscv32imac-unknown-none-elf, mit einem Einsprungpunkt, drei Speicherbereichen für seine Eingaben und Ausgaben und sonst nichts. Der Host ist alles darum herum: der Code, der die Eingabe liefert, Apogee um einen Beweis bittet und diesen Beweis an denjenigen weitergibt, der ihn prüft. Sie schreiben beides. Apogee liefert die Maschine, die Schaltkreise und den Verifier.

Dieser Abschnitt ist für zwei Leser zugleich geschrieben: einen Entwickler an der Tastatur und das Modell, mit dem er arbeitet. Jede Seite nennt ihre Regeln ohne Umschweife, und der KI-Begleiter verdichtet sie alle in einer einzigen Datei, die Sie einem Assistenten übergeben können, bevor er eine Zeile schreibt.

Das Modell#

Das Gastprogramm, der Host und der Verifier Der Host übergibt eine öffentliche Eingabe und Hilfsdaten (Advice) an das Gastprogramm, das in Apogee VM läuft. Das Gastprogramm schreibt ein Journal und endet mit einem Status. Apogee erzeugt einen Beweis, der die Programmidentität, die Eingabe, das Journal und den Exit-Status bindet und den ein Verifier prüft. HOST · IHR CODE Ihr Dienst baut die Eingabe, liefert die Hilfsdaten, fordert einen Beweis an APOGEE VM · RV32IMAC · EIN HART GUEST · IHR CODE Ihr Programm no_std Rust guest_sdk öffentliche Eingabe Hilfsdaten Journal · Exit-Status VERIFIER Beweis geprüft gegen eine Identität aus eigenem Kanal: ein Contract oder Dienst
Wer was tut. Die Eingabe und das Journal sind durch den Beweis gebunden; die Hilfsdaten (Advice), gestrichelt gezeichnet, sind es nicht, und deshalb prüft ein Gastprogramm sie. Der Verifier sieht nie das Programm, nur seine Identität.

Ein Beweis sagt genau eines aus: Das Programm mit dieser Identität, gestartet auf seinem Image mit dieser öffentlichen Eingabe und irgendwelchen vom Prover gewählten Hilfsdaten, ist bis EXIT mit diesem Status gelaufen und hat dabei dieses Journal geschrieben. Alles, was Sie bauen, ruht auf diesem Satz.

Der Ablauf#

Schritt Was Sie tun Seite
1 Installieren Sie nichts von Hand: Das Repository legt die Toolchain fest. Besorgen Sie die Zeremoniedatei für die Beweiserzeugung Umgebung einrichten
2 Schreiben Sie das Gastprogramm: ein Einsprungpunkt, die drei Speicherbereiche und gewöhnliches Rust Ein Gastprogramm schreiben, Eingaben, Hilfsdaten und Journal
3 Greifen Sie dort, wo es sich lohnt, zu delegiertem Hashing und delegierter Kurvenarithmetik Delegationen
4 Bauen Sie es für das Target der Gastprogramme, dann inspizieren Sie das Image und seine Identität Bauen und inspizieren
5 Führen Sie es im Emulator aus und zählen Sie, wohin die Zyklen gehen Ausführen und profilieren
6 Beweisen Sie einen Lauf und verifizieren Sie ihn Beweisen und verifizieren
7 Komprimieren Sie den Beweis per Rekursion und prüfen Sie ihn auf Ethereum On-Chain abwickeln

Der Schnellstart geht den gesamten Ablauf einmal mit einem Gastprogramm aus drei Zeilen durch.

Die wichtigsten Regeln#

Jede wird im Programmierleitfaden für Gastprogramme erklärt, mit dem, was schiefgeht, und dem, was stattdessen zu tun ist.

  • usize und jeder Zeiger sind 32 Bit breit. Ein Überlauf von usize löst nur im Gastprogramm einen Panic aus, x as usize schneidet stillschweigend ab, und alles, dessen Layout oder Hash von einer Länge abhängt, unterscheidet sich zwischen Host und Gastprogramm.
  • Der Allokator gibt nie Speicher frei. Er schiebt einen Zeiger ab __heap_start nach oben; was einem Gastprogramm den Speicher ausgehen lässt, ist also die Gesamtmenge dessen, was es über den Lauf alloziert, nicht sein Spitzenbedarf. Verwenden Sie Puffer wieder und dimensionieren Sie sie mit with_capacity.
  • Atomics funktionieren, und Sie sollten sie nicht schreiben. Die Maschine hat einen einzigen Hart; die A-Erweiterung ist also für die Kompatibilität mit Code da, der sie bereits verwendet. Neuer Code in Gastprogrammen hat nichts zu synchronisieren.
  • Es gibt keine Außenwelt. Keine Dateien, keine Uhr, kein Zufall, kein Netzwerk. Ein Gastprogramm kennt seine öffentliche Eingabe, seine Hilfsdaten und das, was es berechnet.
  • Hilfsdaten wählt der Prover. Prüfen Sie sie gegen etwas, das der Beweis bindet, bevor irgendetwas daraus Abgeleitetes ins Journal gelangt.
  • Das Journal ist klein. Höchstens 16.380 Byte. Veröffentlichen Sie einen Digest von allem, was wächst.
  • Überlaufprüfungen bleiben im Release-Build aktiv. Sie sind Teil dessen, was das Programm berechnet; deshalb legt das Profil der Gastprogramme sie fest.
  • Jeder ausgeführte Befehl ist eine bewiesene Zeile. Die Beweiskosten folgen der Zyklenzahl; bauen Sie also mit --release und zählen Sie Zyklen, bevor Sie irgendetwas anderes optimieren.

Wo Sie beginnen#

Ihre App starten

Schnellstart

Vom leeren Crate zum verifizierten Beweis. Ein Gastprogramm aus drei Zeilen, gebaut, ausgeführt, inspiziert und bewiesen, mit der echten Ausgabe jedes Schritts.

Als Markdown anzeigen

Diese Seite geht den gesamten Ablauf einmal mit dem kleinsten Gastprogramm (Guest) durch, das etwas tut: Es liest seine öffentliche Eingabe und veröffentlicht sie als sein Journal. Jede Ausgabe unten ist entstanden, indem genau diese Befehle auf Apogee v1.0.0 ausgeführt wurden.

Hinweis

Was Sie brauchen. Einen Checkout des Repositorys von Apogee VM auf dem Stand v1.0.0 und rustup; alles andere legt das Repository fest. Befehle werden im Wurzelverzeichnis des Repositorys ausgeführt, sofern ein Schritt nicht das Verzeichnis wechselt. Die Schritte 5 und 6 brauchen außerdem die Zeremoniedatei assets/ptau/ppot_0080_24.ptau, Schritt 6 zudem eine Maschine mit Dutzenden GiB Arbeitsspeicher. Umgebung einrichten behandelt beides.

Das Gastprogramm anlegen#

Ein Gastprogramm ist ein no_std-Binary-Crate im Workspace guests/. Legen Sie guests/hello an:

guests/hello/Cargo.tomltoml
[package]
name = "hello"
version.workspace = true
edition.workspace = true
publish.workspace = true

[dependencies]
guest-sdk.workspace = true
guests/hello/src/main.rsrust
#![no_std]
#![no_main]

guest_sdk::entry!(main);

fn main() {
    // The public input is memory: a slice, with no ecall and no cursor.
    guest_sdk::commit(guest_sdk::public_input());
}

#![no_std], weil das Target Bare Metal ist. #![no_main] mit entry!(main), weil der Startup-Code des SDK den Stack-Pointer setzt, .bss mit Nullen füllt und ein Symbol main aufruft, das das Makro als Wrapper um Ihre Funktion exportiert. Eine Rückkehr aus ihr ist gleichbedeutend mit exit(0).

Im Gastprogramm-Workspace eintragen#

Hängen Sie "hello" an die Liste members in guests/Cargo.toml an:

guests/Cargo.tomltoml
members = ["fib", "echo", … , "recursion", "hello"]

Bauen#

Aus dem eigenen Verzeichnis des Gastprogramms, ohne weiteres Flag außer dem Target:

sh
cd guests/hello
cargo build --release --target riscv32imac-unknown-none-elf
cd ../..

Das ELF landet unter guests/target/riscv32imac-unknown-none-elf/release/hello. Der Gastprogramm-Workspace liefert das Linker-Skript und --no-relax; es gibt also nichts weiter zu übergeben.

Ausführen#

Der Profiler führt ein Gastprogramm im Emulator von Apogee aus, ohne Beweis, und meldet, wohin die Zyklen gegangen sind:

sh
printf 'hello, apogee' > /tmp/hello.in
cargo run --release -p profiler -- elf guests/target/riscv32imac-unknown-none-elf/release/hello --input /tmp/hello.in
workload
  label                        hello
  guest                        hello
  guest cycles                 114
  exit status                  0
  journal bytes                13

cycles by family
  ADD_SUB_LUI_AUIPC            64
  JUMP_BRANCH_SLT              21
  MEM_WORD                     3
  MEM_SUBWORD                  26

114 Befehle liefen, und jeder davon wird eine bewiesene Zeile sein. Die 13 Bytes im Journal sind das Echo der Eingabe. Die 26 MEM_SUBWORD-Zeilen sind commit, das die Eingabe Byte für Byte mit lbu und sb kopiert.

Ansehen, was die VM beweisen wird#

sh
cargo run --release -p artifact-dump -- tables \
    guests/target/riscv32imac-unknown-none-elf/release/hello \
    --ptau assets/ptau/ppot_0080_24.ptau
program identity  9ead85cee880df30daa8eba657316215107a075640a64ccf2424a054b758b802

VmConfig
--------
  id  family              height     live rows  columns
   0  ADD_SUB_LUI_AUIPC     4194304         33  pc next_pc rs1 rs2 rd imm extra_mask
   1  JUMP_BRANCH_SLT       4194304         12  pc next_pc rs1 rs2 rd imm extra_mask
   4  MEM_WORD              4194304          3  pc next_pc rs1 rs2 rd imm extra_mask
   5  MEM_SUBWORD           4194304          3  pc next_pc rs1 rs2 rd imm extra_mask
   7  INIT_TEARDOWN         4194304          0  none: claims no pc
   8  ZERO_WINDOWS          4194304          0  none: claims no pc
  12  PUBLIC_INPUT             4096          0  none: claims no pc
  13  PUBLIC_OUTPUT            4096          0  none: claims no pc
  14  ADVICE_WINDOWS        4194304          0  none: claims no pc

Das ist die statische Form des Programms bei den Standardhöhen: die vier Befehlsfamilien, die sein Code verwendet, jede mit einer dekodierten Tabelle, und die fünf Fensterfamilien, die jedes Programm hat. Die Programmidentität ist ein einziges Körperelement, das all dies in einem Digest zusammenfasst. Ihre Identität wird abweichen: Ein ELF bettet absolute Pfade in seine Panic-Strings ein, ein Build auf einer anderen Maschine ist also ein anderes Image, und jede Änderung der Höhen ergibt eine andere Identität.

Beweisen und verifizieren#

Ein Host-Programm fordert den Beweis an. Legen Sie es als Beispiel neben das Host-SDK:

crates/host/examples/prove_hello.rsrust
use constants::family;
use emulator::GuestIo;
use program::ProgramParams;
use srs::Srs;

fn main() {
    let elf = std::fs::read("guests/target/riscv32imac-unknown-none-elf/release/hello")
        .expect("build the guest with --release first");

    // Small heights for a small program: the seven instruction families at
    // their 2^20 floor, the three RAM-window families at 2^16. Every choice of
    // heights is its own program identity.
    let mut params = ProgramParams::defaults();
    for f in 0..7 {
        params.heights[f] = 1 << 20;
    }
    for f in [family::INIT_TEARDOWN, family::ZERO_WINDOWS, family::ADVICE_WINDOWS] {
        params.heights[f as usize] = 1 << 16;
    }

    // As many ceremony powers as the tallest family has rows: 2^20 here.
    let ptau = std::path::Path::new("assets/ptau/ppot_0080_24.ptau");
    let srs = Srs::from_ptau(ptau, 20).expect("the ceremony file reads");
    let setup = host::setup(&elf, &params, srs).expect("the program registers");

    let io = GuestIo { input: b"hello, apogee".to_vec(), advice: Vec::new() };
    let proven = host::prove(&setup, &io, 2).expect("the run proves"); // two shards in flight
    host::verify(&setup.vk, &proven.block).expect("the block verifies");

    assert_eq!(proven.exit_code, 0);
    assert_eq!(proven.journal, b"hello, apogee");
    let id: String = setup.vk.identity.to_bytes().iter().map(|b| format!("{b:02x}")).collect();
    println!("identity  {id}");
    println!("cycles    {}", proven.cycles);
    println!("shards    {}", proven.report.shards);
    println!("journal   {:?}", core::str::from_utf8(&proven.journal).unwrap());
}
sh
cargo run --release -p host --example prove_hello
identity  606d1f1d720459cc1a078787381656b29c9fce5a9e539b36f899e62b64129c14
cycles    114
shards    7
journal   "hello, apogee"

Auf einem Laptop mit 18 Kernen und 48 GiB dauerte das 52 Sekunden, bei einer Spitze von 18 GB Speicher, fast alles davon für die zwei gleichzeitig bearbeiteten 2^20-Shards. Die Identität weicht von der aus Schritt 5 ab, weil die Höhen abweichen: Die Identität bindet jede Höhe.

Die Identität festhalten#

Ein Verifier übernimmt die Identität nie aus dem Beweis, dem Schlüssel oder vom Prover. Er besitzt eine eigene Kopie, bezogen von demjenigen, der das Release gebaut hat, und vergleicht:

rust
assert_eq!(setup.vk.identity.to_bytes(), registered); // `registered` from your own channel

Gegenüber einer vom Prover gelieferten Identität zeigt ein Beweis nur, dass irgendein Programm gelaufen ist.

Was gerade passiert ist#

Der Emulator hat die 114 Befehle zweimal ausgeführt. Der erste Durchlauf hat die Speicherspalten jedes Shards committet und die Aussage festgelegt. Der zweite hat jeden Shard befüllt und bewiesen. Es waren sieben Shards: je einer für jede der vier Befehlsfamilien, die ausgeführt wurden, einer für das Speicherfenster, das das Image des Programms enthält, und je einer für die öffentliche Eingabe und das Journal. Dieses Gastprogramm hat seinen Stack nie berührt, also brauchte kein weiteres Fenster einen Shard; bei einem typischen Programm kommt der Shard des Stacks hinzu. Jeder Shard wurde vom GKR-Schaltkreis seiner Familie bewiesen und mit einem einzigen Mercury-Beweis geöffnet, und der Verifier hat die Lese- und Schreibzugriffe auf den Speicher aller sieben in einer einzigen Gleichung abgeglichen. Die Architekturübersicht verfolgt denselben Weg im Detail.

Nächste Schritte#

Ihre App starten

Umgebung einrichten

Die Toolchain, die das Repository festlegt, die beiden Workspaces, die es enthält, die Zeremoniedatei, die das Beweisen braucht, und die Maschine, die jeder Schritt verlangt.

Als Markdown anzeigen

Apogee v1.0.0 ist ein Rust-Repository. Außer rustup gibt es nichts zu installieren: Das Repository legt seine Toolchain fest, und die Toolchain bringt das Target der Gastprogramme (Guests) mit. Ein Gastprogramm zu schreiben, zu bauen, auszuführen und zu profilieren erfordert nichts weiter. Das Beweisen fügt eine große Datei hinzu und eine Maschine mit entsprechend viel Arbeitsspeicher.

Die Toolchain#

rust-toolchain.toml im Wurzelverzeichnis des Repositorys legt stable Rust 1.96.1 mit rustfmt, clippy und llvm-tools fest, dazu das Target riscv32imac-unknown-none-elf, dessen core und alloc vorkompiliert mitgeliefert werden. rustup wendet diese Festlegung in jedem Verzeichnis unterhalb der Wurzel an und installiert die Toolchain bei der ersten Verwendung.

sh
cd apogee-vm
rustup show active-toolchain     # 1.96.1, overridden by rust-toolchain.toml
cargo --version

Nirgends werden Nightly oder instabile Features verwendet. llvm-tools liefert die zum LLVM des Compilers passenden llvm-objdump und llvm-nm, die das Repository für seine eingecheckten Disassembly-Listings verwendet und mit denen Sie den Code Ihres Gastprogramms lesen können.

Zwei Workspaces#

Der Checkout enthält zwei Cargo-Workspaces, und die Trennung ist wichtig:

Workspace Wurzel Baut für Enthält
Der Root-Workspace Cargo.toml Ihren Host den Prover, den Verifier, das Host-SDK, die Werkzeuge, alles in crates/ und tools/
Der Gastprogramm-Workspace guests/Cargo.toml riscv32imac-unknown-none-elf jedes Gastprogramm, mit einem eigenen Verzeichnis guests/target

Gastprogramme werden getrennt gehalten, weil jedes Mitglied für das Target der Gastprogramme kompiliert und einen #[panic_handler] linkt; cargo test --workspace in der Wurzel darf sie nie erreichen. Der Gastprogramm-Workspace bringt außerdem mit, was ein Gastprogramm braucht, um korrekt gebaut zu werden, sodass Sie es nie eintippen:

  • guests/.cargo/config.toml setzt das Target und übergibt dem Linker -T crates/guest-sdk/link.ld, die Speicherkarte, sowie --no-relax, weil Relaxation Adressen verschieben würde, die die Programmidentität bindet.
  • guests/Cargo.toml legt beide Build-Profile auf dieselbe Semantik fest, Überlaufprüfungen eingeschlossen (Bauen und inspizieren).
  • Sein [patch.crates-io] leitet k256, ark-ff und revm-precompile auf mitgelieferte Kopien um, die die Delegationen von Apogee aufrufen (Delegationen).

Tipp

Öffnen Sie guests/ in Ihrem Editor als eigenen Ordner. rust-analyzer liest dann die .cargo/config.toml dieses Workspace und prüft den Code der Gastprogramme gegen deren Target statt gegen Ihren Host.

Die Zeremoniedatei#

Jedes Commitment, das Apogee bildet, beruht auf den Potenzen eines geheimen τ aus einer öffentlichen Zeremonie: den Perpetual Powers of Tau der PSE, Beitrag 80. Eine einzige Datei dient jedem Zweck:

assets/ptau/ppot_0080_24.ptau        19.3 GB, 2^24 powers; assets/ptau/ is gitignored

Sie brauchen sie, um eine Programmidentität zu berechnen, um echte Schlüssel zu bauen und um zu beweisen. Sie brauchen sie nicht, um ein Gastprogramm zu bauen, auszuführen oder zu profilieren, und auch nicht für die Workspace-Tests, die über eigenen Spielzeug-Setups beweisen.

Dateien aus der Zeremonie der PSE sind aus einem einzigen Transkript geschnitten; jede Datei mit Potenz 24 oder mehr ist also geeignet. powersOfTau28_hez_final_*.ptau von Hermez ist eine andere Zeremonie mit einem anderen τ: Der Reader liest sie ebenso bereitwillig ein, und jedes Commitment, jeder Schlüssel und jede Identität darüber fällt anders aus. Damit Sie bestätigen können, dass Sie die richtige Zeremonie haben: Das [τ]_1 der Zeremonie, als Hex seiner kanonischen Kodierung x ‖ y, lautet:

9bbb31bedc304e081e2aada4b56c2217e0e94ee16874e3517d14bef5dcec3a16
317ff1589e53513fa333591b318e8f1e55ef7c37d92beb6d9a61d770a2f39506

Die SRS-Seite der Spezifikation legt genau dar, was der Reader prüft und was er ungeprüft voraussetzt.

Die Maschine#

Bauen und Ausführen sind Laptop-Arbeit. Beweisen ist speichergebunden, und sein Speicherbedarf folgt den gleichzeitig bewiesenen Shards, nicht der Länge des Laufs.

Schritt Braucht
Ein Gastprogramm bauen, ausführen, profilieren, sein Image exportieren und inspizieren Jeden aktuellen Laptop; Sekunden
Eine Programmidentität berechnen (artifact-dump tables --ptau) Die Zeremoniedatei; etwa 25 s auf einem Laptop mit 18 Kernen bei den Standardhöhen
Ein kleines Gastprogramm bei Höhen von 2^20 beweisen Dutzende GiB. Ein 2^20-Shard der breitesten Befehlsfamilie hält in seinem Vorwärtsdurchlauf etwa 8,4 GiB an Körperelementen, und jeder gleichzeitig bearbeitete Shard hält seine eigenen
Einen vollständigen Ethereum-Block beweisen Der gemessene Block erreichte eine Spitze von 174 GiB auf einer Maschine mit 32 vCPUs und 247,7 GiB

Die Seite zum Beweisen erklärt, wie Höhen und die Zahl der gleichzeitig bearbeiteten Shards Speicher gegen Zeit tauschen: Beweisen und verifizieren.

Ihren Checkout prüfen#

Was die CI ausführt, alles ohne die Zeremoniedatei:

sh
cargo fmt --all -- --check
cargo clippy --workspace --all-targets -- -D warnings
cargo test --workspace
cargo run -p kat-gen && git diff --exit-code      # committed fixtures regenerate identically
(cd guests && cargo clippy --bins -- -D warnings)

Die Suiten, die echte Shards beweisen, sind mit #[ignore] markiert, weil jede Dutzende GiB braucht. Rufen Sie eine namentlich auf, wenn Sie auf Ihrer eigenen Maschine sehen wollen, wie ein Beweis erstellt und zurückgewiesen wird:

sh
cargo test --release -p prover --test acceptance -- --include-ignored --test-threads=1

Weiter: Ein Gastprogramm schreiben, oder gehen Sie den gesamten Ablauf einmal im Schnellstart durch.

Ihre App starten

Ein Gastprogramm schreiben

Ein Gastprogramm ist ein no_std-Rust-Binary mit einem Einsprungpunkt und drei Speicherbereichen. Crate-Aufbau, die Laufzeitumgebung darunter, Abhängigkeiten und der Aufbau, mit dem Sie es zuerst auf dem Host testen wie jedes andere Rust.

Als Markdown anzeigen

Das Crate#

In v1.0.0 ist ein Gastprogramm (Guest) ein Binary-Crate im Workspace guests/ des Repositorys. Der Workspace liefert das Target, die Linker-Flags, die festgelegten Profile und die mitgelieferten Crates; das eigene Manifest eines Gastprogramms bleibt also kurz:

guests/my-app/Cargo.tomltoml
[package]
name = "my-app"
version.workspace = true
edition.workspace = true
publish.workspace = true

[dependencies]
guest-sdk.workspace = true
guests/my-app/src/main.rsrust
#![no_std]
#![no_main]

extern crate alloc; // Vec, Box, String, BTreeMap, over the SDK's allocator

use alloc::vec::Vec;

guest_sdk::entry!(main);

fn main() {
    let input = guest_sdk::public_input();
    let mut out = Vec::with_capacity(input.len());
    out.extend(input.iter().rev());
    guest_sdk::commit(&out);
}

Tragen Sie "my-app" in members in guests/Cargo.toml ein und bauen Sie aus dem eigenen Verzeichnis des Gastprogramms: cargo build --release --target riscv32imac-unknown-none-elf.

Was unter Ihnen läuft#

Das Guest-SDK ist die gesamte Laufzeitumgebung. Es ist klein genug, um vollständig beschrieben zu werden:

  • Start. _start liegt bei 0x0001_0000, dem ersten Byte von .text. Es setzt sp auf das obere Ende des RAM, füllt .bss Byte für Byte mit Nullen und ruft main auf. entry!(f) exportiert dieses main als Wrapper um Ihre Funktion, die keine Argumente nimmt und () zurückgibt.
  • Exit. Eine Rückkehr aus main ist gleichbedeutend mit exit(0). guest_sdk::exit(code) beendet den Lauf mit einem beliebigen Status. Ein Status ungleich null ist eine fehlgeschlagene Ausführung, und auch eine fehlgeschlagene Ausführung ist beweisbar: Die Aussage enthält den Status, und ein Verifier liest ihn.
  • Panic. Der Panic-Handler beendet den Lauf mit Status 101 und schreibt nichts. Es gibt keinen Diagnosekanal. Ein Gastprogramm, das einen Panic auslöst, hat trotzdem alles veröffentlicht, was es vor dem Panic festgeschrieben hat.
  • Heap. Ein Bump-Allokator wächst ab __heap_start, direkt oberhalb von .bss, nach oben. Er gibt nie Speicher frei. Siehe den Heap.
  • Systemaufrufe. Die einzigen ecalls, die ein Gastprogramm absetzt, sind EXIT und die Delegationsaufrufe, die das SDK für Sie absetzt. Eingabe, Hilfsdaten (Advice) und Ausgabe sind Speicher, gelesen und geschrieben mit gewöhnlichen Lade- und Speicherbefehlen.

Die Speicherkarte#

Der gesamte 32-Bit-Adressraum, wie ein Gastprogramm ihn sieht:

Bereich Was es ist
0x0000_0000 – 0x0000_8000 Ein Loch. Nichts initialisiert es; ein Null- oder wilder Zeiger ist also ein fataler Fehler OutOfBounds, kein stiller Lesezugriff
0x0000_8000 – 0x0000_C000 Das Fenster der öffentlichen Eingabe, 16 KiB
0x0000_C000 – 0x0001_0000 Das Journal-Fenster, 16 KiB
0x0001_0000 – … .text (mit _start am Anfang), dann .rodata, .data und .bss, jeweils an einer Seitengrenze ausgerichtet
__heap_start aufwärts Der Heap, ab dem Ende von .bss, aufgerundet auf ein Vielfaches von 16
0x7F80_0000 – 0x8000_0000 Die Reserve von 8 MiB für den Stack. Kein Heap-Block darf oberhalb von 0x7F80_0000 enden; der Stack wächst ab 0x8000_0000 nach unten
0x8000_0000 – 2^32 Der Bereich der Hilfsdaten, bis zu 2^29 Wörter, nur so weit adressierbar, wie der Host Daten geliefert hat

Code ist statisch. Der Befehl jedes pc stammt aus den dekodierten Tabellen des Programms, nie aus dem RAM; ein Schreibzugriff auf .text ändert also, was ein späterer Lesezugriff liest, aber nicht, was ausgeführt wird.

Der Heap#

Der Allokator schiebt einen Zeiger weiter, und dealloc tut nichts. Das ist das richtige Design für ein kurzes Programm, bei dem jeder Zyklus Beweiszeit kostet, und es verändert, wie Sie Rust schreiben:

  • Was Ihnen den Speicher ausgehen lässt, ist die Gesamtmenge Ihrer Allokationen, nicht Ihr Spitzenbedarf. Eine Schleife, die in jeder Iteration einen Vec aufbaut und verwirft, verbraucht jedes Mal frischen Heap.
  • Verwenden Sie Puffer wieder. Ziehen Sie Allokationen aus Schleifen heraus, leeren Sie mit clear() und füllen Sie neu, statt neu zu allozieren, und dimensionieren Sie wachsende Collections mit with_capacity, damit sie beim Wachsen nicht neu allozieren und kopieren.
  • Die Obergrenze ist Exit 71. Eine Allokation, die oberhalb von 0x7F80_0000 oder oberhalb des aktuellen Stack-Pointers enden würde, beendet den Lauf mit Status 71, statt null zurückzugeben oder den Stack zu überschreiben.
rust
// Allocates a fresh Vec per record: total heap grows with the record count.
for record in records {
    let fields: Vec<&[u8]> = record.split(|b| *b == b',').collect();
    handle(&fields);
}

// One buffer, reused: total heap is the largest record's field count.
let mut fields: Vec<&[u8]> = Vec::with_capacity(16);
for record in records {
    fields.clear();
    fields.extend(record.split(|b| *b == b','));
    handle(&fields);
}

Der Stack hat seine Reserve von 8 MiB, und tiefe Rekursion innerhalb davon ist unproblematisch. Was nichts erkennt, ist ein Stack, der über die Reserve hinauswächst, nachdem der Heap den Raum darunter gefüllt hat: Heap-Blöcke würden sich dann unter einer tiefen Aufrufkette verändern. Halten Sie Rekursion beschränkt, oder formulieren Sie sie iterativ.

Abhängigkeiten#

Jedes Crate, das ohne std für riscv32imac-unknown-none-elf baut, ist geeignet. In der Praxis:

  • Schalten Sie die Default-Features ab (default-features = false) und aktivieren Sie alloc, wo ein Crate es anbietet.
  • Ein Crate, das getrandom, eine Uhr oder std::collections::HashMap mit seinem zufälligen Seed einbindet, hat keine Quelle, aus der es schöpfen kann. Ein solcher Aufruf erhält -ENOSYS als Antwort und macht den Lauf unbeweisbar. Bevorzugen Sie BTreeMap oder eine Hash-Map mit einem festen, deterministischen Hasher.
  • Gleitkommaarithmetik wird zu Ganzzahl-Softwareroutinen kompiliert, weil das Target keine F- oder D-Erweiterung hat. Sie ist korrekt und deterministisch und kostet viele Befehle pro Operation. Ganzzahl- oder Festkommaarithmetik ist günstiger.
  • Hashing und Arithmetik auf elliptischen Kurven haben eigene Schaltkreise. Verwenden Sie die Funktionen des SDK oder die mitgelieferten Crates, damit Ihre Abhängigkeiten sie erreichen: Delegationen.

Zuerst auf dem Host testen#

Ein Gastprogramm gibt nichts aus; das Debugging findet also auf dem Host statt. Der Aufbau, der das einfach macht, hält das Programm in einer #![no_std]-Bibliothek, die Bytes auf Bytes abbildet, beschränkt main.rs darauf, Bytes in die Speicherbereiche hinein und aus ihnen heraus zu bewegen, und macht das SDK zu einer Abhängigkeit allein des Gastprogramm-Targets:

guests/my-app/Cargo.tomltoml
[package]
name = "my-app"
version.workspace = true
edition.workspace = true
publish.workspace = true

[target.'cfg(target_arch = "riscv32")'.dependencies]
guest-sdk.workspace = true
guests/my-app/src/lib.rsrust
#![no_std]
extern crate alloc;
use alloc::vec::Vec;

/// The whole application: public input and advice in, journal out.
pub fn run(input: &[u8], advice: &[u8]) -> Result<Vec<u8>, i32> {
    let _ = advice;
    let mut out = Vec::with_capacity(input.len());
    out.extend(input.iter().rev());
    Ok(out)
}
guests/my-app/src/main.rsrust
#![no_std]
#![no_main]

guest_sdk::entry!(main);

fn main() {
    match my_app::run(guest_sdk::public_input(), &[]) {
        Ok(journal) => guest_sdk::commit(&journal),
        Err(code) => guest_sdk::exit(code),
    }
}

Host-Code hängt dann per Pfad von der Bibliothek ab, so wie crates/emulator von guests/revm-block abhängt, führt my_app::run nativ aus und vergleicht das Ergebnis mit dem Journal, das der Emulator für dieselbe Eingabe erzeugt (Ausführen und profilieren). Ihre Logik bekommt auf dem Host Unit-Tests, einen Debugger und println!, und das Binary des Gastprogramms bleibt eine dünne Hülle um Code, den Sie bereits getestet haben.

Warnung

Die beiden Builds sind sich über usize uneinig. Im Gastprogramm sind usize und jeder Zeiger 32 Bit breit, auf Ihrem Host 64. Ein Überlauf eines usize löst nur im Gastprogramm einen Panic aus, x as usize schneidet dort stillschweigend ab, und size_of und core::hash von allem, was eine Länge enthält, unterscheiden sich zwischen beiden. Halten Sie usize aus allem heraus, was Sie festschreiben, hashen oder serialisieren, und verwenden Sie an diesen Grenzen explizit u32 und u64.

Assembly und der Befehlssatz#

Der Decoder akzeptiert genau die 59 Befehle von RV32IMA sowie komprimierte Befehle (C), die beim Laden expandiert werden. Inline-Assembly innerhalb dieser Menge ist unproblematisch. Alles außerhalb davon, etwa ein CSR-Zugriff, fence.i, eine Gleitkomma- oder eine RV64-Kodierung, führt dazu, dass sich das gesamte Programm nicht registrieren lässt, selbst wenn die Stelle nie erreicht wird: Die Ableitung meldet Not all opcodes supported: pc=…. Ein ebreak, ein Sprung auf ein Halbwort ohne Befehl oder ein nicht ausgerichteter Halbwort- oder Wortzugriff beendet den Lauf ohne Beweis.

Atomare Befehle werden dekodiert und bewiesen, mit einer Abweichung: sc.w gelingt immer, weil die Maschine keinen Reservierungszustand führt. Der Programmierleitfaden für Gastprogramme erklärt, warum neuer Code in Gastprogrammen überhaupt keine Atomics verwenden sollte.

Nächste Schritte#

Ihre App starten

Eingaben, Hilfsdaten und Journal

Ein Gastprogramm hat keine I/O-Systemaufrufe. Seine öffentliche Eingabe, die Hilfsdaten des Provers und sein Journal sind drei Speicherbereiche. Was jeder enthält, was der Beweis bindet und das Muster, dem jedes Gastprogramm folgt, das Daten entgegennimmt.

Als Markdown anzeigen

Ein Gastprogramm (Guest) von Apogee hat keine Dateideskriptoren, keine Streams und keinen I/O-Syscall. Seine Eingaben und Ausgaben sind drei Speicherbereiche, gelesen und geschrieben mit gewöhnlichen Lade- und Speicherbefehlen, und der Beweis bindet zwei davon.

Die drei Speicherbereiche#

Bereich SDK Enthält Größe Durch den Beweis gebunden
Öffentliche Eingabe public_input(), read_input(buf) die Bytes der Aussage, gewählt von demjenigen, der den Beweis anfordert höchstens 16.380 Byte ja, ihr Anfangsinhalt
Hilfsdaten (Advice) advice() Bytes, die der Prover wählt bis zu 2 GiB nein
Journal commit(bytes), journal() was das Gastprogramm angehängt hat höchstens 16.380 Byte ja, sein Endinhalt
rust
let input: &[u8] = guest_sdk::public_input(); // a slice over the input window, no copy
let data: &[u8] = guest_sdk::advice();        // a slice over the advice region
guest_sdk::commit(b"result");                  // appends to the journal
  • public_input() und advice() geben Slices über Speicher zurück; nichts wird kopiert. read_input(buf) kopiert min(buf.len(), input.len()) Bytes und gibt die Anzahl zurück; es kann also weniger liefern als angefordert.
  • commit hängt an und führt ein Längenwort mit, und genau das sorgt dafür, dass der Beweis einen exakten Byte-String bindet statt eines mit Nullen aufgefüllten Fensters. Es beendet den Lauf mit Status 70, statt das Fenster überlaufen zu lassen.
  • advice() in einem Lauf ohne Hilfsdaten ist ein fataler OutOfBounds-Fehler, kein leerer Slice: Ein Lauf ohne Hilfsdaten hat überhaupt keinen Bereich für Hilfsdaten und zahlt nichts dafür.

Was „gebunden“ bedeutet#

Die Aussage, die ein Beweis begründet, enthält die Bytes der öffentlichen Eingabe, die Bytes des Journals und den Exit-Status. Der Beweis zeigt, dass das Eingabefenster vor dem ersten Zugriff des Gastprogramms genau die Eingabe der Aussage enthielt und dass das Journal-Fenster beim Exit des Gastprogramms genau die Ausgabe der Aussage enthielt. Das beruht auf dem Speicherargument, nicht auf irgendetwas, das das Gastprogramm tut: Es gibt keinen Hash, den das Gastprogramm berechnen, und keine Konvention, der es folgen muss.

Bei Hilfsdaten ist es anders. Der Anfangsinhalt des Bereichs der Hilfsdaten ist das, was der Prover dort hineingeschrieben hat, und nichts verknüpft ihn mit der Programmidentität, der Aussage oder irgendeinem Gate. Ein Beweis sagt aus, dass es irgendwelche Hilfsdaten gibt, unter denen das Programm mit dieser Eingabe dieses Journal veröffentlicht hat. Das ist genauso stark wie die eigenen Prüfungen des Gastprogramms an den Hilfsdaten.

Das Muster: committen, liefern, prüfen#

Ein Gastprogramm mit einer großen Eingabe nimmt den Großteil als Hilfsdaten entgegen, für die die öffentliche Eingabe ein Commitment enthält, und prüft das eine gegen das andere, bevor irgendetwas aus den Hilfsdaten Abgeleitetes ins Journal gelangt:

Die Hilfsdaten prüfen, bevor Sie ihnen vertrauenrust
fn main() {
    let want = guest_sdk::public_input(); // 32 bytes: keccak256 of the advice
    let data = guest_sdk::advice();       // the prover's bytes, bound by nothing
    if guest_sdk::keccak256(data).as_slice() != want {
        guest_sdk::exit(1); // refused before anything derived from it is committed
    }
    let sum = data.iter().fold(0u32, |s, b| s.wrapping_add(u32::from(*b)));
    guest_sdk::commit(&sum.to_le_bytes()); // the journal: what the proof publishes
}

Die Prüfung muss kein Hash über die gesamten Hilfsdaten sein. Sie kann ein Merkle-Pfad sein, geprüft gegen eine Wurzel, die die Eingabe mitbringt, wie im Hauptbuch-Beispiel, oder eine Signatur über die Daten, oder eine Bedingung, die das Ergebnis selbst erfüllt, etwa eine behauptete Sortierreihenfolge, die das Gastprogramm in einem Durchgang verifiziert, statt zu sortieren. Entscheidend ist, dass das, wogegen geprüft wird, gebunden ist.

Vorsicht

Wer irgendeine Funktion ungeprüfter Hilfsdaten festschreibt, veröffentlicht einen Wert, den der Prover gewählt hat. Das ist der häufigste Weg, ein Gastprogramm zu schreiben, dessen Beweis nichts bedeutet.

Strukturierte Daten#

Kodieren Sie strukturierte Eingaben mit einem no_std-Serialisierer wie postcard über serde mit dem Feature alloc; genau das verwendet das Ethereum-Gastprogramm des Repositorys für seinen Block-Witness. Zwei Gewohnheiten halten ein Format ehrlich:

  • Verwenden Sie Ganzzahlen fester Breite. u32 und u64, nie usize, dessen Breite sich zwischen Ihrem Host und dem Gastprogramm unterscheidet.
  • Bestehen Sie auf einer einzigen Kodierung pro Wert, wenn es darauf ankommt. Ein Deserialisierer, der überzählige Bytes am Ende oder nicht minimale Varints akzeptiert, lässt zwei Byte-Strings für einen Wert zu. Wo Eindeutigkeit zählt, dekodieren Sie, kodieren erneut und vergleichen, wie es BlockWitness::decode im Ethereum-Gastprogramm tut.

Ausgaben, die wachsen#

Das Journal fasst 16.380 Byte. Eine Ausgabe, die mit der Arbeit wächst, etwa ein Datensatz pro Transaktion, hat keine feste Schranke und endet irgendwann mit Exit 70. Veröffentlichen Sie stattdessen einen Digest: Hashen Sie die Datensätze, während Sie sie erzeugen, schreiben Sie das 32-Byte-Ergebnis fest, und lassen Sie jeden, der die Datensätze braucht, sie nativ neu berechnen und vergleichen. Der zustandslose Ethereum-Validator des Repositorys veröffentlicht auf diese Weise ein 43-Byte-Journal für einen ganzen Block.

Für einen Beweis, der auf Ethereum geprüft wird, halten Sie beide öffentlichen Werte auf fester Länge. Der bereitgestellte Verifier-Contract ist für eine Eingabelänge und eine Ausgabelänge gebaut und weist alles andere zurück (On-Chain abwickeln).

Der Exit-Status#

Beim Exit wird nichts über das Festgeschriebene hinaus veröffentlicht, und ein Lauf, der einen Panic auslöst oder mit einem Status ungleich null endet, hat einen gültigen Beweis dessen, was er getan hat. Deshalb liest ein Verifier den Exit-Status, bevor er das Journal liest. On-chain nimmt der Verifier-Contract den erwarteten Status als Argument entgegen, und eine Anwendung übergibt 0. Die eigenen Statuswerte des SDK:

Status Bedeutung
0 main ist zurückgekehrt, oder exit(0)
70 commit hätte 16.380 Byte überschritten
71 eine Allokation hätte den Stack erreicht: siehe den Heap
72 eine Delegation hat etwas geantwortet, das ihr Shim zurückweist
101 ein Panic, der nichts ausgibt

Wählen Sie Ihre eigenen Fehlercodes außerhalb dieser Werte, wie es das Hauptbuch-Beispiel mit 1, 2 und 3 tut.

Was der Beweis nicht aussagt#

  • Nichts ordnet die Schreibzugriffe auf das Journal, und nichts zwingt ein Gastprogramm, seine Eingabe zu lesen. Der Beweis bindet die Inhalte der Fenster, nicht die Zugriffe, die sie erzeugt haben.
  • Hilfsdaten sind beschreibbar. Ein Schreibzugriff auf den Bereich der Hilfsdaten ist ein gewöhnlicher Schreibzugriff. Ungebunden bleiben sie so oder so.
  • Das Journal ist der gesamte Endinhalt des Fensters. commit hält diese Form ein. Ein Gastprogramm, das das Fenster direkt beschreibt, muss sie einhalten: eine Länge von höchstens 16.380, so viele Bytes, dann Nullen.

Die Spezifikation legt all dies präzise fest: Öffentliche Werte und Hilfsdaten.

Ihre App starten

Delegationen

Hashing, Körper- und Kurvenarithmetik haben eigene Schaltkreise. Welche SDK-Aufrufe sie erreichen, was sie kosten, die Regeln für ihre Operanden und die mitgelieferten Crates, die Bibliothekscode zu ihnen leiten.

Als Markdown anzeigen

Manche Berechnungen lassen sich mit einem eigens für sie gebauten Schaltkreis weit günstiger beweisen denn als Folge von RISC-V-Befehlen. Apogee nennt sie Delegationen. Eine Delegation ist eine Schaltkreisfamilie, die eine Funktion über einem Frame aus Wörtern im RAM beweist und über einen ecall aufgerufen wird, und das Guest-SDK setzt diese Aufrufe für Sie hinter gewöhnlichen Funktionen ab. Sie schreiben nie selbst einen ecall.

Was Sie aufrufen und was es erreicht#

Sie rufen auf Delegation Ein Aufruf beweist
guest_sdk::keccak256(&[u8]) -> [u8; 32] KECCAK_F eine Runde von keccak-f[1600]; eine Permutation sind 24 Aufrufe, und Sponge und Padding sind Code des Gastprogramms (Guest)
guest_sdk::sha256(&[u8]) -> [u8; 32] SHA256_COMP vier Runden der Kompression; eine Kompression sind 16 Aufrufe
guest_sdk::ec_add, ec_mul, ec_identity EC_ADD ein Drittel einer vollständigen Punktaddition auf secp256k1 oder BN254 G1
guest_sdk::poseidon2_permute(&mut [u8; 96]) POSEIDON2 eine Poseidon2-Permutation der Breite 3 über Fr
field::Fr: Addition, Multiplikation, Inversion FR_ARITH eine Fr-Operation, auf dem Target der Gastprogramme, ohne dass etwas benannt werden muss
transcript::poseidon2_permute POSEIDON2 dieselbe Permutation, über das Transkript-Crate
guest_sdk::recursion::mod_mul über einem ModMulFrame MOD_MUL ein 256-Bit-a·b mod m, wobei m einer von vier Ethereum-Moduln ist

Die Funktionen sind bitgenau identisch mit ihren Software-Definitionen. keccak256 ist das Keccak von Ethereum, nicht SHA3-256. sha256 ist FIPS 180-4. ec_add verwendet die vollständige Formel von Renes, Costello und Batina (2015, Algorithmus 7); Verdopplung, P + (−P), das neutrale Element und jedes Z brauchen also keinen Sonderfall.

Hashing und Kurvenarithmetik aus einem Gastprogrammrust
use guest_sdk::{ec_mul, keccak256, recursion::SECP256K1_GROUPS, ProjectivePoint};

let digest: [u8; 32] = keccak256(b"blockchain-native");

// A point is homogeneous projective (x = X/Z, y = Y/Z), each coordinate eight
// little-endian u32 limbs below the field modulus. The scalar is eight limbs too.
fn times(p: &ProjectivePoint, k: &[u32; 8]) -> ProjectivePoint {
    ec_mul(&SECP256K1_GROUPS, p, k).expect("EC_ADD is implemented on Apogee")
}

Auch Bibliothekscode erreicht sie#

Der Gastprogramm-Workspace patcht drei Crates so, dass der Code in ihnen auf dem Target der Gastprogramme Delegationen aufruft, mit dem Upstream-Code als Fallback-Pfad:

Crate Version Erreicht
k256 0.13.4 MOD_MUL aus Körper- und Skalarmultiplikation; EC_ADD aus Addition, gemischter Addition und Verdopplung von ProjectivePoint
ark-ff 0.6.0 MOD_MUL aus der Montgomery-Multiplikation und -Quadrierung von BN254, in beiden seiner Körper
revm-precompile 43.0.2 SHA256_COMP für das Precompile 0x02; EC_ADD für 0x06 und 0x07

Ein Gastprogramm, das von diesen Crates abhängt, erhält die gepatchten Kopien automatisch über das [patch.crates-io] in guests/Cargo.toml. Ungepatcht machten allein die Körpermultiplikation und -quadrierung von k256 44 % der Zyklen eines Mainnet-Blocks aus. Die secp256k1-Signatur-Recovery ist gewöhnlicher k256-Code, den die Patches in delegierte Arithmetik verwandeln.

Was eine Delegation kostet#

Eine Delegationsfamilie ist nur dann Teil eines Programms, wenn das Programm einen ihrer Shims linkt, und ein Aufruf kostet Shards der Höhe dieser Familie:

  • Gelinkt und nie aufgerufen: nichts. Die Familie ist deklariert und beweist null Shards.
  • Einmal aufgerufen: ein ganzer Shard. Ein Shard kostet seine volle Höhe, unabhängig von seiner Belegung.
  • Oft aufgerufen: sehr wenig pro Aufruf. Der Beweis eines Shards wächst mit dessen Höhe nur um eine Sumcheck-Runde pro Variable.
Familie Höhe Arbeitseinheit Aufrufe pro Einheit Einheiten pro Shard
KECCAK_F 2^18 keccak-f[1600] 24 10.922
SHA256_COMP 2^18 eine Kompression 16 16.384
EC_ADD 2^16 eine vollständige Addition 3 21.845
MOD_MUL 2^16 ein a·b mod m 1 65.536
POSEIDON2 2^8 eine Permutation 1 256
FR_ARITH 2^8 eine Fr-Operation 1 256

Der Preis ist eher Speicher als Zeit: Der Vorwärtsdurchlauf eines KECCAK_F-Shards der Höhe 2^18 hält etwa 42 GiB an Körperelementen, und zwei davon, gleichzeitig in Bearbeitung, bestimmten die Spitze des gemessenen Ethereum-Blocks.

Regeln für Operanden#

  • Operanden unterhalb ihres Moduls. Ein Operand von MOD_MUL oder EC_ADD, der den Modul, den sein Selektor benennt, erreicht oder übersteigt, hat keinen Beweis: Der Executor weist den Frame als fatalen DelegationFrame-Fehler zurück. Das mitgelieferte k256 reduziert seine verzögert reduzierten Körperelemente vor dem Aufruf.
  • Punkte werden nicht für Sie geprüft. EC_ADD beweist die Arithmetik der Formel. Ob ein Punkt auf der Kurve liegt, ist eine Frage an den aufrufenden Code, und ein Gastprogramm, das Punkte aus Hilfsdaten (Advice) übernimmt, muss sie stellen.
  • Operationen aus mehreren Aufrufen sind jeweils eine SDK-Funktion. Eine Keccak-Permutation sind 24 Aufrufe auf einem Frame, eine SHA-256-Kompression 16, eine Punktaddition 3. Jeder Aufruf beweist seinen eigenen Schritt, und nichts weist eine falsche Reihenfolge zurück: Es wird dann schlicht etwas anderes berechnet. Verwenden Sie keccak256, sha256 und ec_add, die die Aufrufe in der richtigen Reihenfolge absetzen, statt der rohen Shims.
  • Exit 72 bedeutet, dass eine Delegation etwas geantwortet hat, das ihr Shim zurückweist. Auf dem eigenen Executor von Apogee kommt das bei einem wohlgeformten Frame nicht vor.

Was nicht delegiert wird#

  • MULMOD der EVM mit beliebigem Modul, MODEXP, BLS12-381 und jedes Primitiv außerhalb der obigen Tabelle laufen als Befehle.
  • Kein Signaturverfahren und kein Pairing wird als Ganzes delegiert. Die secp256k1-Recovery ist k256 über MOD_MUL und EC_ADD; ein BN254-Pairing ist ark-bn254 über MOD_MUL.
  • Eine Delegation ist der Kern einer Operation. Padding, Sponges, Blockschleifen und die Leiter einer Skalarmultiplikation sind Code des Gastprogramms und werden als Befehle bewiesen.

Signaturverfahren für Gastprogramme stehen auf der Roadmap für v2.0.0.

Lohnt sich eine neue Delegation?#

Der Zyklus-Profiler bepreist die naheliegenden Kandidaten in jedem Bericht, als Obergrenze der Zyklen, die eine Delegation einsparen könnte:

removable = max(0, cycles − calls·(4 + 2·frame_words))

cycles ist der Anteil der Kategorie am Lauf, calls die Zahl der Einsprünge in die Funktionen des Kandidaten und 4 + 2·frame_words der Shim, den eine Delegation zurücklassen würde: die Schreibzugriffe für den Frame, der ecall und die Lesezugriffe für das Ergebnis. Die Formel veranschlagt nichts für die Shards der neuen Familie; behandeln Sie sie also als obere Schranke. Ausführen und profilieren zeigt einen Bericht.

Die Spezifikation jeder Delegation, Spalte für Spalte, steht unter Delegationsschaltkreise.

Ihre App starten

Bauen und inspizieren

Build-Profile, auf eine einzige Semantik festgelegt, das ELF, das ein Build erzeugt, das ProgramImage, das der Loader daraus macht, sein Bericht und die Programmidentität, die ein Verifier registriert.

Als Markdown anzeigen

Bauen#

Aus dem Verzeichnis des Gastprogramms (Guest), ohne weiteres Flag außer dem Target:

sh
cd guests/my-app
cargo build --release --target riscv32imac-unknown-none-elf     # .../release/my-app
cargo build --target riscv32imac-unknown-none-elf               # .../debug/my-app

Das ELF landet im einzigen Target-Verzeichnis des Gastprogramm-Workspace, guests/target/riscv32imac-unknown-none-elf/. guests/.cargo/config.toml fügt zwei Linker-Argumente hinzu, die Sie nie eintippen:

  • -T crates/guest-sdk/link.ld, die Speicherkarte, die auch die Symbole definiert, die der Startup-Code und der Allokator verwenden.
  • --no-relax. Linker-Relaxation schreibt Befehlsfolgen um und verschiebt jede spätere Adresse, und die Programmidentität bindet diese Adressen.

Es gibt keinen runner: Nichts außerhalb von Apogee bildet die Speicherbereiche eines Gastprogramms ab; cargo run hat also nichts, womit es das Gastprogramm ausführen könnte. Sie führen ein Gastprogramm über den Emulator aus (Ausführen und profilieren).

Profile#

guests/Cargo.toml legt beide Profile auf eine einzige Semantik fest. Sie unterscheiden sich nur in der Optimierung und in den Debug-Assertions der Abhängigkeiten:

dev release
opt-level 0 3
overflow-checks an an
debug-assertions an an im Crate des Gastprogramms, aus in seinen Abhängigkeiten
panic, codegen-units, debug, incremental abort, 1, aus, aus identisch

Das Standard-Release-Profil von Cargo schaltet Überlaufprüfungen ab, und in einem Gastprogramm ist das keine Performance-Einstellung. Es ändert die Aussage: u32::MAX + 1 würde 00000000 festschreiben und mit Exit 0 enden, wo der Dev-Build einen Panic auslöst und mit Exit 101 endet. Deshalb lässt der Workspace die Prüfungen in beiden Profilen aktiv. Die Debug-Assertion einer Abhängigkeit prüft eine Invariante dieses Crates selbst, und eine korrekte Abhängigkeit berechnet ohne sie dasselbe; das Release-Profil schaltet sie deshalb ab und spart die Zyklen: 6,8 % des Laufs des zustandslosen Ethereum-Gastprogramms.

Beweisen Sie den Release-Build. Jeder ausgeführte Befehl ist eine bewiesene Zeile, opt-level = 3 entfernt ein Viertel bis über die Hälfte des Images eines Gastprogramms, und der Code jeder Familie muss in ihre dekodierte Tabelle passen: Das Debug-Image des Ethereum-Gastprogramms braucht Tabellen mit 2^22 Zeilen, sein Release-Image 2^20. Die Identität, die Sie veröffentlichen, ist die des Release-Images.

Reproduzierbarkeit#

Zwei saubere Builds auf einer Maschine erzeugen identische ELFs. Builds auf zwei Maschinen im Allgemeinen nicht: Das ELF bettet in Strings für Panic-Locations absolute Pfade ein, zu den core-Quellen der Toolchain, zu crates/guest-sdk und zur Cargo-Registry, während die eigenen Dateien des Gastprogramms relativ zu guests/ erscheinen. Ein Build an anderer Stelle ist ein anderes Image mit einer anderen Identität.

Was Sie registrieren und weitergeben, ist also das ELF eines Builds, kein Rezept. Bewahren Sie das ELF auf, das Sie bewiesen haben, und lassen Sie jeden, der die Identität prüfen will, sie aus diesem ELF, den Parametern und der Zeremoniedatei neu berechnen.

Das Image exportieren#

sh
cargo run -p artifact-dump -- guests/target/riscv32imac-unknown-none-elf/release/my-app --out artifacts

Das schreibt artifacts/my-app.img, das geladene ProgramImage in seinem Serialisierungsformat (postcard, ohne Header), und artifacts/my-app.img.txt, einen Bericht, der aus dem über den validierenden Reader zurückgelesenen Image erzeugt wird. Ausgegeben werden der Einsprungpunkt, die Zahl der Segmente und Befehle sowie Größe und SHA-256 des Artefakts. Weicht das Zurückgelesene ab oder weist der Loader das ELF zurück, wird nichts geschrieben.

Die .img-Datei ist die statische Beschreibung des Programms, zum Aufbewahren und Vergleichen; nichts Nachgelagertes braucht sie, denn das Setup und die Werkzeuge nehmen das ELF entgegen. Ihr SHA-256 legt Bytes fest. Er ist nicht die Programmidentität.

Den Bericht lesen#

Abschnitt Zeigt
entry and memory den Einsprung, _start bei 0x00010000; das RAM-Fenster; slot_base und die Slot-Spanne
segments Adresse, Ende, mem_len, Dateibytes, Nullauffüllung und Befehlszahl jedes Segments: .text, .rodata, falls vorhanden, und ein beschreibbares Segment bis 0x80000000 für .data, .bss, Heap und Stack
instruction stream Vier- und Zwei-Byte-Befehle, Slots mitten in einem Befehl und not code-Slots, die sich zur Slot-Zahl summieren
symbols Namen nach Adresse aus der Symboltabelle des ELF, die das Artefakt nicht mitführt
listing pro Befehl: Adresse, Länge, die Bytes im Speicher, das expandierte 32-Bit-Wort, das Symbol

Ein komprimierter Befehl behält seine Adresse und seine zwei Bytes; allein len gibt an, ob der nächste pc pc + 2 oder pc + 4 ist. Für Mnemonics verwenden Sie die Ansicht tables unten oder den festgelegten Disassembler:

sh
"$(rustc --print sysroot)"/lib/rustlib/*/bin/llvm-objdump \
    --disassemble --no-print-imm-hex -M no-aliases <elf>

Die not code-Halbwörter#

Ein Bericht kann eine Zeile wie ---- not code: 0x00010f9a .. 0x00010f9c, 1 halfword ---- enthalten. Das ist gewöhnliche Compiler-Ausgabe. LLVM hat den Default-Zweig eines match als unerreichbar bewiesen, rustc hat den unerreichbaren Block zu unimp übersetzt, und mit der C-Erweiterung ist das c.unimp, das Halbwort aus lauter Nullen, die definiert illegale Kodierung von RVC. Der Loader verzeichnet es als Nicht-Befehl und macht weiter. Kein pc erreicht es; einer, der es täte, würde den Lauf mit NotAnInstruction anhalten.

Was die VM beweisen wird#

sh
cargo run --release -p artifact-dump -- tables <elf> --ptau assets/ptau/ppot_0080_24.ptau

Das gibt bei den Standardparametern die VmConfig aus, die aus dem Image abgeleitet wird: die Höhe jeder Familie, ihre belegten Zeilen und dekodierten Spalten. Danach folgen für jeden Befehl pc, next_pc, Familie, Mnemonic und Felder. Mit --ptau und der Zeremoniedatei gibt es außerdem die Programmidentität aus, den Wert, den ein Verifier registriert. Der Schnellstart zeigt eine echte.

Die Identität ist ein einziges Körperelement. Sie bindet jeden Befehl mit seinem pc, seiner Länge, seinen Operanden und seiner Art, jedes in der Datei enthaltene Byte des Images (.text, .rodata, .data), den Einsprungpunkt, die Menge der Familien, jede Höhe, die Obergrenze der Codegröße und die Codeversion. Sie bindet nicht die Symboltabelle, nicht .bss und nichts, was eine Ausführung wählt. Ein ELF mit zwei Einstellungen der Höhen hat zwei Identitäten.

Die Ableitung weist das Programm zurück und nennt dabei den pc oder die Größe:

Zurückweisung Ursache
Not all opcodes supported: pc=… ein Wort außerhalb von RV32IMA irgendwo im ausführbaren Code, etwa ein CSR-Zugriff in Assembly
TableTooShort Code jenseits der Reichweite einer Familie, pc ≤ 2h − 4: 1,9375 MiB Code bei 2^20 und 7,9375 MiB bei 2^22
ProgramTooLarge das Image überschreitet bytecode_size_words, standardmäßig 4 MiB
ImageOutsideWindow ein in der Datei enthaltenes Byte liegt bei der gewählten Fensterhöhe jenseits von RAM-Fenster 0
UnknownDelegation das Image deklariert eine Delegationsnummer, auf die keine Familie antwortet

Einen Build prüfen#

Bauen Sie in ein frisches Target-Verzeichnis, exportieren Sie erneut und vergleichen Sie:

sh
cd guests/my-app
CARGO_TARGET_DIR=/tmp/fresh cargo build --release --target riscv32imac-unknown-none-elf
cd ../..
cargo run -p artifact-dump -- /tmp/fresh/riscv32imac-unknown-none-elf/release/my-app --out /tmp/again
cmp artifacts/my-app.img /tmp/again/my-app.img
diff artifacts/my-app.img.txt /tmp/again/my-app.img.txt     # differs only in the `source ELF` line

Ihre App starten

Ausführen und profilieren

Führen Sie ein Gastprogramm im Emulator von Apogee aus, aus Rust oder von der Kommandozeile, vergleichen Sie es mit Ihrem Host-Build und finden Sie heraus, wohin seine Zyklen gehen, bevor Sie dafür bezahlen, sie zu beweisen.

Als Markdown anzeigen

Ein Gastprogramm (Guest) auszuführen kostet fast nichts; es zu beweisen kostet proportional zu den Zyklen, die es ausführt. Führen Sie es also zuerst aus, vergleichen Sie es mit Ihrem Host-Build und sehen Sie sich das Zyklenprofil an, bevor Sie irgendetwas beweisen.

Aus Rust: der Emulator#

emulator::run führt ein geladenes Image über einer öffentlichen Eingabe und Hilfsdaten (Advice) aus, in Host-Code, ohne Beweis:

Ein Gastprogramm ausführen und mit dem Host-Build vergleichenrust
let elf = std::fs::read(elf_path)?;
let image = loader::load_elf(&elf).expect("the ELF loads");
let io = emulator::GuestIo { input: b"hi".to_vec(), advice: Vec::new() };
let run = emulator::run(&image, &io).expect("no fatal error");

assert_eq!(run.exit_code, 0);
assert_eq!(run.io.output, my_app::run(b"hi", &[]).unwrap()); // the host build agrees
println!("{} cycles", run.cycle_count);

run gibt eine Execution zurück: die finalen Register, den Exit-Status, die Zyklenzahl und die öffentlichen Werte. Ein Exit-Status ungleich null ist eine Ausführung, kein Fehler, und kommt als exit_code zurück. Ein fataler Fehler des Executors, etwa OutOfBounds, Misaligned oder NotAnInstruction, kommt als EmuError zurück, und ein solcher Lauf hat keinen Beweis (Fehlerbehebung).

Der Emulator ist eine reine Funktion des Images und der Eingabe: keine Uhr, kein Zufall, keine Threads. Dieselbe Eingabe ergibt dieselbe Ausführung, Zyklus für Zyklus, und genau das erlaubt es auch dem Prover, zweimal auszuführen und identische Shards zu schneiden.

Von der Kommandozeile: der Profiler#

sh
cargo run --release -p profiler -- elf <elf> [--input <file>] [--advice <file>] [--top <n>] [--json <path>]

Er führt das Gastprogramm über den angegebenen Dateien bei der kleinsten Tabellenhöhe aus, in die sein Code passt, und gibt einen Bericht aus. Seine Zahlen sind Zählungen ausgeführter Zyklen, auf jeder Maschine dieselben.

workload
  label                        hello
  guest cycles                 114
  exit status                  0
  journal bytes                13

cycles by semantic workload
  core runtime                             94   82.46%
  unattributed                             20   17.54%

cycles by family
  ADD_SUB_LUI_AUIPC            64
  JUMP_BRANCH_SLT              21
  MEM_WORD                     3
  MEM_SUBWORD                  26

top functions
            94   82.46%          1 calls        94.0 c/call  guest_sdk::commit  [core runtime]
             8    7.02%          1 calls         8.0 c/call  main  [unattributed]

So lesen Sie ihn:

  • Der Abschnitt cycles by family zeigt, wofür Sie bezahlen. Jede Familie mit Zeilen kostet mindestens einen Shard ihrer Höhe, und mehr Zyklen in einer Familie bedeuten mehr Shards davon.
  • Der Abschnitt top functions rechnet jeder Funktion ihre eigenen Zyklen an, einschließlich allem, was der Compiler per Inlining in sie übernommen hat, aber nicht die der von ihr aufgerufenen Funktionen. Aufrufe werden am ersten Befehl der Funktion gezählt.
  • Der Abschnitt cycles by semantic workload ordnet Funktionen anhand ihres Namens vierzehn Kategorien zu, etwa Hashing, Signaturen und dem Kern der Laufzeitumgebung. Der nicht zugeordnete Anteil und die Verteilung der Mnemonics dienen als Kontrolle dieser Zuordnung, denn keine Symboltabelle kann sie falsch beschriften.
  • Der Abschnitt accelerator candidates bepreist die Delegationen, die eine künftige Version hinzufügen könnte, als Obergrenze: siehe Delegationen.

Der Profiler hat zwei weitere Unterbefehle für die Ethereum-Arbeitslast: block <stem> führt das revm-Gastprogramm über einer aufgezeichneten Fixture aus, und record <number|latest> zeichnet einen Block von ETH_RPC_URL auf und führt ihn aus.

Günstiger machen#

Die Reihenfolge, die sich meist auszahlt:

  1. Bauen Sie mit --release. Die Optimierung entfernt ein Viertel bis über die Hälfte der Befehle eines Gastprogramms.
  2. Delegieren Sie Hashing und Kurvenarithmetik. Verwenden Sie guest_sdk::keccak256, sha256, ec_add und die mitgelieferten k256 und ark-ff, statt eine Software-Implementierung in das Gastprogramm zu kompilieren.
  3. Hören Sie auf, in Schleifen zu allozieren. Jede Allokation kostet Befehle, und mit einem Bump-Allokator ist sie außerdem Speicher, den Sie nie zurückbekommen (der Heap).
  4. Prüfen statt berechnen. Ist ein Ergebnis teuer zu finden und günstig zu verifizieren, etwa eine Sortierreihenfolge, eine Quadratwurzel oder ein Pfad durch einen Baum, lassen Sie es den Prover als Hilfsdaten liefern und das Gastprogramm es verifizieren.
  5. Vermeiden Sie Gleitkommaarithmetik. Sie wird zu Softwareroutinen kompiliert; Ganzzahl- und Festkommaarithmetik sind weit günstiger.

Messen Sie dann erneut. Zyklenzahlen sind exakt und reproduzierbar; jede Änderung zeigt sich also als Zahl.

Ihre App starten

Beweisen und verifizieren

Ein Programm registrieren, einen Lauf beweisen, den Block verifizieren und den Beweis aufbewahren. Höhen, gleichzeitig bearbeitete Shards, die Potenzen der Zeremonie, die ein Beweis braucht, und die zwei Werte, die ein Verifier selbst besitzen muss.

Als Markdown anzeigen

Die drei Aufrufe#

Setup, Beweis, Verifikationrust
let params = program::ProgramParams::defaults();
let ptau = std::path::Path::new("assets/ptau/ppot_0080_24.ptau");
let srs = srs::Srs::from_ptau(ptau, 22).expect("ceremony");

let setup = host::setup(&elf, &params, srs).expect("registers");          // once per program
let proven = host::prove(&setup, &io, 4).expect("proves");                // at most 4 shards in flight
host::verify(&setup.vk, &proven.block).expect("verifies");

assert_eq!(setup.vk.identity.to_bytes(), registered); // from your own channel, never the proof
assert_eq!(proven.exit_code, 0);
  • host::setup lädt das ELF, dekodiert es in seine Familientabellen und seine VmConfig, committet die Setup-Spalten unter der Zeremonie und baut den Verifikationsschlüssel. Seine Kosten fallen pro Programm und pro Wahl der Höhen an, nicht pro Lauf.
  • host::prove führt das Gastprogramm (Guest) zweimal aus und beweist jeden Shard (unten). Es gibt ein Proven zurück: den BlockProof, den Exit-Code, die Zyklenzahl, das Journal und einen Bericht über den Lauf.
  • host::verify prüft den Block gegen den Schlüssel, anhand der Aussage, die der Block mitführt. Es vergleicht weder die Identität noch den SRS-Digest mit irgendetwas; dieser Vergleich ist also Ihre Aufgabe.

Der Schnellstart führt genau diesen Code über einem kleinen Gastprogramm aus, mit seiner echten Ausgabe.

Was ein Verifier selbst besitzen muss#

Zwei Werte stammen aus einem Kanal, den der Prover nicht kontrolliert:

  1. Die Programmidentität. Gegenüber einer vom Prover gelieferten Identität zeigt ein Beweis nur, dass irgendein Programm gelaufen ist. Ein Verifier registriert die Identität des Release, dem er vertraut, und vergleicht sie mit der des Schlüssels.
  2. Der SRS-Digest der Zeremonie. Ein Schlüssel wird unter jedem Digest geladen, den seine eigenen Punkte ergeben. Einer, der über einem bekannten τ gebaut wurde, könnte alles öffnen und wird nur durch den Vergleich seines Digests mit dem der Zeremonie zurückgewiesen.

Der Verifikationsschlüssel selbst darf von beliebiger Seite stammen, auch vom Prover: Beim Laden werden die Identität und der SRS-Digest aus seinem eigenen Inhalt neu berechnet, und seine Schaltkreise werden mit der eigenen Registry des Verifiers abgeglichen. Lesen Sie dann die Aussage: zuerst den Exit-Status, dann das Journal.

Höhen#

Jede Familie hat eine Höhe, die Zahl der Zeilen in einem ihrer Shards, gewählt aus 2^8, 2^12, 2^16, 2^18, 2^20, 2^22. Höhen sind Teil des Programms, nicht eines Laufs: Jede Höhe ist in die Identität eingebunden.

Familiengruppe Standard Untergrenze Anmerkungen
Die sieben Befehlsfamilien 2^22, bzw. 2^20 für MUL_DIV und ATOMICS 2^20 die Untergrenze ihrer Range-Checks für Zeitstempel
INIT_TEARDOWN, ZERO_WINDOWS, ADVICE_WINDOWS 2^22 2^16 eine gemeinsame Fensterhöhe; Fenster 0 muss jedes in der Datei enthaltene Byte des Images aufnehmen
PUBLIC_INPUT, PUBLIC_OUTPUT 2^12 fest die Höhe bestimmt die Lage ihrer Fenster
Delegationsfamilien siehe Delegationen je Familie

Eine Familie mit Zeilen kostet mindestens einen ganzen Shard ihrer Höhe; ein kurzer Lauf verschwendet also bei kleineren Höhen weniger, und ein langer braucht bei größeren Höhen weniger Shards. Eine dekodierte Tabelle muss außerdem hoch genug sein, um den letzten Befehl der Familie zu erreichen: 2^20 reicht für 1,9375 MiB Code und 2^22 für 7,9375 MiB. Das Ethereum-Gastprogramm beweist bei 2^20 für jede Familie, deren Höhe wählbar ist.

Befehlsfamilien an ihrer Untergrenze, RAM-Fenster bei 2^16rust
use constants::family;

let mut params = program::ProgramParams::defaults();
for f in 0..7 {
    params.heights[f] = 1 << 20;
}
for f in [family::INIT_TEARDOWN, family::ZERO_WINDOWS, family::ADVICE_WINDOWS] {
    params.heights[f as usize] = 1 << 16; // window 0 is then 256 KiB: the image must fit in it
}

Die Zeremonie muss so viele Potenzen liefern, wie die höchste Familie Zeilen hat, und mindestens 2^18 für die generische Lookup-Tabelle: Srs::from_ptau(path, k) mit 2^k mindestens so groß wie die größte Höhe.

Gleichzeitig bearbeitete Shards#

Das dritte Argument von host::prove ist max_in_flight, die Zahl der gleichzeitig bewiesenen Shards. Es ist der eine Regler, der Speicher gegen Zeit tauscht:

  • Der Speicherbedarf folgt den gleichzeitig bearbeiteten Shards, nicht der Zyklenzahl. Jeder gerade bearbeitete Shard hält seine Zeilen, seinen Vorwärtsdurchlauf und seinen Beweis, während dieser wächst. Ein 2^20-Shard der breitesten Befehlsfamilie hält in seinem Vorwärtsdurchlauf etwa 8,4 GiB; ein KECCAK_F-Shard der Höhe 2^18 etwa 42 GiB.
  • Die Zeit folgt der Zahl der nebeneinander laufenden Shards, bis zur Zahl Ihrer Kerne. Innerhalb eines Shards läuft die Arbeit auf allen Kernen.
  • Der Beweis hängt nicht davon ab. Der Block ist bei 1 und bei 8 gleichzeitig bearbeiteten Shards byte-identisch.

Beginnen Sie auf einem Laptop niedrig, mit zwei oder vier, und erhöhen Sie den Wert auf einem Server, bis der Speicher und nicht die Kerne die Grenze bilden. bench prove verwendet standardmäßig 8.

Die zwei Durchläufe#

host::prove arbeitet im Streaming-Verfahren. Es hält nie den gesamten Ausführungs-Trace, der mit etwa 300 Byte pro Zyklus das größte Objekt im System wäre.

  1. Durchlauf 1 führt das Gastprogramm aus und committet, sobald ein Shard gefüllt ist, dessen Speicherspalten, behält die Commitments und verwirft die Zeilen. Beim Exit baut er die Aussage und zieht die Challenges, die alle Shards teilen.
  2. Durchlauf 2 führt erneut aus. Der Emulator ist deterministisch und schneidet daher dieselben Shards. Jeder wird bei seinem Eintreffen befüllt, bewiesen und verworfen, und nur sein Beweis wird behalten.

Deshalb kostet das Beweisen die Zeit zweier Ausführungen und einen Speicherbedarf, der durch die gleichzeitig bearbeiteten Shards beschränkt ist. Der Streaming-Prover erklärt das ausführlich.

Den Beweis aufbewahren#

host::proof_archive::write_proof(dir, stem, vk, block) schreibt vier Dateien, jede mit den reinen Bytes ihres Typs:

<stem>.vk         the verifying key
<stem>.identity   the key's identity, 64 lowercase hex digits and a newline
<stem>.public     the statement: input, journal, exit status, the execution's record
<stem>.block      the block proof

read_proof(dir, stem) liest sie zurück. Die Datei .identity hält fest, was der Lauf behauptet hat; ein Verifier vergleicht trotzdem mit seiner eigenen Kopie. Das Kommandozeilenwerkzeug verifier prüft ein Archiv:

sh
cargo run --release -p verifier -- block <stem>.vk <identity-hex> <stem>.public <stem>.block

Es endet mit 0, wenn die Verifikation vollständig gelingt, mit 1 unter Nennung der ersten Zurückweisung und mit 2 bei einem Aufruffehler oder einer fehlerhaft formatierten Identität. Es vergleicht die übergebene Identität mit der des Schlüssels und entnimmt den SRS-Digest der Schlüsseldatei.

Wenn ein Beweis fehlschlägt#

Ein ehrlicher Prover erzeugt nie einen Beweis, der fehlschlägt; ein Fehlschlag bedeutet also eine Eingabe, die er nicht hätte annehmen dürfen, oder einen Bug. Bauen Sie mit aktiviertem Debug-Log des Provers neu und führen Sie erneut aus:

sh
cargo run --release -p bench --features prover/debug-info -- prove ...
APOGEE_DEBUG=detail <the same run> 2>&1 | tee run.log
grep -E 'FAIL|NOT CANONICAL|UNBALANCED|OVER the|NAMES NO|DISAGREES|NOT LOOPING|ABORTED' run.log

Das Log nennt den Shard, der abgebrochen ist, und self_check FAILED nennt das erste Gate, das eine Zeile verletzt, mit dem Wert jedes Operanden. Werkzeuge und CLIs listet jeden Marker auf. Das Log existiert nur in Builds mit dem Feature debug-info und ändert kein Byte des Beweises.

Weiter: On-Chain abwickeln.

Ihre App starten

On-Chain abwickeln

Von einem Basisbeweis aus Hunderten von Shards zu einem einzigen Groth16-Beweis, den ein Ethereum-Contract prüft. Der Rekursionsbaum, die Zeremonie des Deciders, die Schnittstelle des Contracts und was ein Deployment festlegt.

Als Markdown anzeigen

Ein Basisbeweis ist ein Block aus Shard-Beweisen, jeder ein GKR-Beweis mit seinen Commitments: Megabytes an Daten und Hunderte von Kurvenpunkten, die kein Contract prüfen kann. Die Abwicklung komprimiert ihn in drei Stufen, jede mit bench im Wurzelverzeichnis des Repositorys ausgeführt.

Vom Basisbeweis zum Contract Basis-Shards werden von Blättern verifiziert, Blätter von inneren Knoten, Knoten von einer Wurzel; ein Groth16-Decider verifiziert die Wurzel erneut; der Contract prüft den Groth16-Beweis und das gefaltete Pairing. BASISBEWEIS 207 Shards · 14,5 MB BLÄTTER je ≤ 64 Basis-Shards WURZEL 2–4 Kinder umfasst 0..count DECIDER Groth16 BN254 7,9M Constraints CONTRACT verify(…) → true 3,62 Mgas
Abwicklung. Jede Stufe verifiziert die vorherige. Vor dem Contract findet kein Pairing statt: Jede Mercury-Prüfung wird aufgeschoben und in einen einzigen Akkumulator gefaltet, den der Contract mit zwei Pairings einlöst. Die Zahlen stammen von Block 257.510.

1. Der Rekursionsbaum#

Ein Knoten ist Apogee beim Beweisen eines Verifier-Programms. Ein Blatt verifiziert eine Reihe aufeinanderfolgender Basis-Shards; ein innerer Knoten verifiziert zwei bis vier Kind-Beweise; die Wurzel deckt jeden Basis-Shard ab. Jeder Knoten faltet außerdem jede Mercury-Prüfung, die seine Shards und Kinder aufschieben, in ein einziges Punktepaar, sodass der ganze Baum an der Spitze auf eine einzige Pairing-Behauptung hinausläuft.

sh
# the base proof as an archive: host::proof_archive::write_proof from your host
# program, or `bench prove ... --out <dir>` for the Ethereum guests
cargo run --release -p bench -- recurse <dir>/<stem> --out <out> --in-flight 4

recurse schreibt die Schlüssel der beiden Rekursionsprogramme, baut darauf die Blatt- und Knoten-Binaries, legt einen Plan fest, bevor irgendetwas bewiesen wird (<out>/tree.txt: Blätter aus höchstens --leaf 64 Basis-Shards, dann Knoten aus höchstens --fan-in 4 Kindern), und beweist Knoten für Knoten. Jeder Knoten ist ein eigener Prozess, der seine Eingaben vor dem Beweisen nativ verifiziert; eine fehlerhafte Eingabe wird also namentlich zurückgewiesen. Ein angehaltener Lauf lässt sich fortsetzen: Bereits in <out> liegende Beweise bleiben erhalten, und ein Lauf mit abweichendem Plan oder abweichenden Programmen wird zurückgewiesen.

Das Beweisen der Basis bleibt von alledem unberührt. Ein Blatt verifiziert Basis-Shards genau so, wie sie sind.

2. Der Decider#

Die Wurzel ist immer noch ein GKR-Beweis und einige Hundert Punkte. Der Decider ist ein Groth16-Schaltkreis, der die Wurzel so verifiziert, wie es ein Knoten täte, und nichts faltet. Stattdessen bindet er als Wires, deren Werte der Contract liefert, die Identitäten der beiden Rekursionsprogramme, den Exit-Status der Basisaussage, ihre öffentliche Eingabe und ihr Journal Byte für Byte sowie jeden Punkt, dem die Wurzel ein Pairing schuldet, mit seinem Skalar. Der Groth16-Beweis trägt ein einziges Commitment auf all diese Wires, und der Contract prüft es gegen die Werte, die er besitzt.

Ein Groth16-Schlüssel braucht eine Zeremonie. Phase 1 ist dieselbe Powers-of-Tau-Datei, auf der die Commitments des Baums beruhen. Phase 2 ist schaltkreisspezifisch und läuft in zwei Beitragsrunden:

sh
cargo run --release -p bench -- ceremony <out> init          # once per root shape
cargo run --release -p bench -- ceremony <out> contribute    # round 1: alpha and beta, each contributor in turn
cargo run --release -p bench -- ceremony <out> seal
cargo run --release -p bench -- ceremony <out> contribute    # round 2: gamma, delta and eta
cargo run --release -p bench -- ceremony <out> key
cargo run --release -p bench -- decide <out>                 # the Groth16 proof, checked natively and in an EVM

Jeder Beitrag multipliziert eine Falltür mit einem Faktor, den nur sein Beitragender kannte, und wird mit einem Schnorr-Beweis festgehalten; jeder Zustand lässt sich also allein gegen den Schaltkreis und die Zeremoniedatei verifizieren. Eine Falltür ist unbekannt, solange einer ihrer Beitragenden ehrlich war. Die Reihenfolge der Runden ist Teil der Soundness: alpha und beta sind abgeschlossen, bevor irgendetwas durch delta oder eta geteilt wird.

Vorsicht

bench decide --dev-key leitet jede Falltür aus einem öffentlichen Seed ab, für Entwicklung und Tests. Jeder kann damit einen Beweis fälschen, und seine Ausgaben werden als development.* geschrieben, damit sie nicht mit denen einer Zeremonie verwechselt werden können. Eine Zeremonie, die auf einer einzigen Maschine läuft, ist ebenfalls keine Zeremonie: Sie braucht einen ehrlichen Beitragenden pro Runde.

decide schreibt decision.constructor und decision.calldata: die Deployment-Argumente und den Aufruf, als Hex.

3. Der Contract#

contracts/ApogeeVerifier.sol hat einen einzigen Einsprungpunkt:

solidity
function verify(
    bytes calldata input,        // the base program's public input
    bytes calldata output,       // its journal
    uint256 exitStatus,          // the status you require, normally 0
    uint256[10] calldata proof,  // Groth16 A, B, C and the bound wires' commitment D
    uint256[] calldata points    // x, y and scalar of each point, side [1]_2's then side [x]_2's
) external view returns (bool);

Der Contract rekonstruiert die gebundenen Werte aus den Calldata, prüft die Pairing-Gleichung von Groth16, faltet die Punkte jeder Seite mit ecMul und ecAdd, was zugleich erzwingt, dass jeder Punkt auf der Kurve liegt, und prüft die gefaltete Behauptung e(A, [1]_2) = e(B, [x]_2). Ein Anwendungs-Contract ruft ihn auf und handelt dann anhand des Journals: siehe die Skizze des Hauptbuchs.

Was ein Deployment festlegt#

Der Konstruktor nimmt den Groth16-Schlüssel, die beiden G2-Punkte der Zeremonie, die Identitäten der Blatt- und Knotenprogramme, die Zahl der Punkte auf jeder Seite und die Bytelängen der öffentlichen Eingabe und des Journals entgegen. Ein bereitgestellter Verifier bedient also:

  • Ein einziges Basisprogramm. Seine Identität ist eine Konstante im Image des Blattprogramms, die die Identität des Blatts bindet.
  • Eine einzige Wurzelform. Der Schaltkreis des Deciders hängt vom Programm der Wurzel, ihren Shard-Zahlen und den Längen der öffentlichen Werte ab; ein Schlüssel und seine Zeremonie gelten also je Form.
  • Öffentliche Werte fester Länge. verify weist eine Eingabe oder ein Journal jeder anderen Länge zurück. Entwerfen Sie Gastprogramme (Guests), deren öffentliche Werte on-chain eine feste Größe haben, etwa einen festen Datensatz oder einen 32-Byte-Digest.

Der Contract zahlt etwa 9.000 gas pro Punkt, weil der Schaltkreis keinen davon faltet.

Gemessen#

Block 257.510, mit dem Baum auf einer Maschine mit 32 CPUs und 247 GiB und mit Zeremonie und Decider auf einem Laptop mit 18 Kernen:

Stufe Ergebnis
Basisbeweis 207 Shards, 14,5 MB, 2.481 s
Baum 4 Blätter aus höchstens 64 Basis-Shards und eine Wurzel: 116 Shards
Blätter, vier gleichzeitig 21, 24, 23 und 27 Shards; 2.157 s; Spitze 92 GiB
Wurzel, vier Shards gleichzeitig in Bearbeitung 21 Shards, 460 s, 1,03 MB
Decider-Schaltkreis 7.896.686 Constraints, eine Domäne der Größe 2^23
Zeremonie init 65 s; ein Beitrag 50–56 s; key 70 s und 12,7 GB; der Schlüssel 2,65 GB
Decider-Beweis Schlüssel in 1 s eingelesen, Beweis 18,5 s, 6,1 GB
Contract 358 Punkte; 3.620.026 gas; 34.980 Byte Calldata

Die Spezifikation all dessen ist Rekursion und Decider.

Ihre App starten

Programmierleitfaden für Gastprogramme

Die Gewohnheiten, die ein Gastprogramm korrekt, beweisbar und günstig halten. Jeder Fallstrick und jede Präferenz an einem Ort, jeweils mit der Begründung und dem, was stattdessen zu tun ist.

Als Markdown anzeigen

Ein Gastprogramm (Guest) zu schreiben heißt größtenteils, Rust zu schreiben. Diese Seite behandelt den Rest: die Stellen, an denen sich eine bewiesene Bare-Metal-Maschine mit einem einzigen Hart anders verhält als der Host, den Sie gewohnt sind. Jede Regel sagt, was zu tun ist, warum und was sonst schiefgeht. Der KI-Begleiter enthält dieselben Regeln in einer Form, die Sie einem Modell übergeben können.

Typen und Speicher#

usize ist 32 Bit breit, und jeder Zeiger ebenso#

Das Target der Gastprogramme ist riscv32imac: usize, isize und jeder Zeiger sind 32 Bit breit, die Ihres Hosts dagegen 64.

  • Ein Überlauf eines usize löst im Gastprogramm einen Panic aus, auf dem Host nicht.
  • x as usize aus einem u64 schneidet im Gastprogramm stillschweigend ab.
  • size_of::<T>(), das Struct-Layout und core::hash von allem, was eine Länge oder einen Zeiger enthält, unterscheiden sich zwischen den beiden Builds.

Empfohlen: Verwenden Sie explizit u32 und u64 in allem, was Sie festschreiben, hashen, serialisieren oder mit einer Berechnung auf dem Host vergleichen. Konvertieren Sie mit usize::try_from(x), wo ein Wert womöglich nicht passt, damit es auf beiden Builds laut fehlschlägt. Zu vermeiden: usize festschreiben, eine Struktur hashen, die eines enthält, oder eine layoutabhängige Kodierung ableiten.

rust
let n = u64::from_le_bytes(input[..8].try_into().unwrap());
let len = usize::try_from(n).expect("length fits the guest"); // not `n as usize`

Der Allokator gibt nie Speicher frei#

Der Heap ist ein Bump-Allokator: alloc schiebt einen Zeiger nach oben, dealloc tut nichts, und Speicher kommt erst zurück, wenn das Programm endet. Deshalb gilt: Was einem Gastprogramm den Speicher ausgehen lässt, ist die Gesamtmenge dessen, was es über den Lauf alloziert, nicht sein Spitzenbedarf. Läge das Ende einer Allokation oberhalb der Reserve des Stacks oder oberhalb des aktuellen Stack-Pointers, endet das Gastprogramm mit Status 71.

Empfohlen:

  • Einmal allozieren und wiederverwenden: Puffer aus Schleifen herausziehen und mit clear() leeren, statt neue zu bauen.
  • Collections vorab mit Vec::with_capacity, String::with_capacity dimensionieren, damit sie beim Wachsen nicht neu allozieren und kopieren. Ein Vec, der Push für Push auf n Elemente wächst, lässt außerdem seine früheren, kleineren Puffer zurück.
  • Borrowing (&[u8], &str) dem Klonen vorziehen und Iteratoren den Zwischen-Collections.
  • Große Hilfsdaten (Advice) an Ort und Stelle verarbeiten: advice() ist bereits ein Slice über Speicher, es gibt also nichts zu kopieren.

Zu vermeiden: in einer heißen Schleife mit collect() in einen neuen Vec sammeln, Werte klonen, die Sie nur lesen, oder eine Map pro Anfrage neu aufbauen, wenn sich eine einzige Map leeren und neu befüllen lässt.

rust
// Total heap grows with the number of requests:
for req in requests {
    let parts: Vec<u32> = req.chunks(4).map(|c| u32::from_le_bytes(c.try_into().unwrap())).collect();
    process(&parts);
}

// Total heap is one buffer:
let mut parts: Vec<u32> = Vec::with_capacity(MAX_PARTS);
for req in requests {
    parts.clear();
    parts.extend(req.chunks(4).map(|c| u32::from_le_bytes(c.try_into().unwrap())));
    process(&parts);
}

Der Stack hat 8 MiB, und nichts bewacht seinen fernen Rand#

Der Stack wächst ab 0x8000_0000 nach unten und hat eine Reserve von 8 MiB, in die kein Heap-Block eindringen darf. Tiefe Rekursion innerhalb davon ist unproblematisch. Was nichts erkennt, ist ein Stack, der über seine Reserve hinauswächst, nachdem der Heap den Raum darunter gefüllt hat: Heap-Blöcke ändern sich dann unter einer tiefen Aufrufkette, stillschweigend. Empfohlen: die Rekursionstiefe beschränkt und vorhersehbar halten oder tiefe Traversierungen iterativ mit einer expliziten Arbeitsliste schreiben. Zu vermeiden: bis zu einer Tiefe rekursieren, die eine nicht vertrauenswürdige Eingabe bestimmt.

Nur ausgerichtete Zugriffe#

Ein Halbwort- oder Wortzugriff über einen nicht ausgerichteten Zeiger ist fatal, wird nie aufgeteilt, und der Lauf hat keinen Beweis. Safe Rust erzeugt nie einen solchen Zugriff. Vermeiden Sie es, einen Byte-Zeiger nach *const u32 zu casten und ihn zu dereferenzieren; lesen Sie mit u32::from_le_bytes, das zu Byte-Ladebefehlen kompiliert, oder mit ptr::read_unaligned.

Null ist ein Loch#

Adressen unterhalb von 0x8000 gehören zu nichts; ein Null-Zeiger oder ein kleiner wilder Zeiger ist also ein fataler OutOfBounds-Fehler statt eines Lesezugriffs auf Datenmüll. Er zeigt sich als Lauf ohne Beweis, nie als falsche Antwort.

Nebenläufigkeit#

Atomics: unterstützt, aber nicht für neuen Code in Gastprogrammen#

Wichtig

Die A-Erweiterung wird vollständig unterstützt: lr.w, sc.w und alle neun AMOs werden über ihre eigene Schaltkreisfamilie dekodiert, ausgeführt und bewiesen, und core::sync::atomic kompiliert zu ihnen. Dennoch wird dringend davon abgeraten, ein Gastprogramm mit Atomics zu schreiben. Apogee führt auf einem einzigen Hart aus, ohne Interrupts und ohne Threads; es gibt also nichts zu synchronisieren. Atomics gibt es, damit bestehender Code, der sie verwendet, etwa eine Bibliothek mit einem atomaren Zähler oder einem spin-Lock, unverändert kompiliert und bewiesen wird. Sie sind ein Kompatibilitätspfad, keine Arbeitsweise.

Was Sie wissen sollten, wenn Atomics über eine Abhängigkeit in Ihr Gastprogramm gelangen:

  • Auf einem einzigen Hart ist eine atomare Operation nur ein Read-Modify-Write. fetch_add ist ein amoadd.w, das addiert; nichts kann sich dazwischenschieben.
  • sc.w gelingt immer. Die Maschine führt keinen Reservierungszustand; ein Store-Conditional speichert also und schreibt 0 nach rd. Die lr.w/sc.w-Wiederholungsschleife, die kompilierter Code für compare_exchange verwendet, ist davon nicht betroffen, weil Erfolg beim ersten Versuch auf jedem Hart zulässig ist. Code, der sich darauf verlässt, dass sc.w ohne gültige Reservierung fehlschlägt, bekommt diesen Fehlschlag hier nicht. Das ist die einzige Stelle, an der Apogee von RV32IMAC abweicht.
  • fence tut nichts, und Speicherordnungen (aq, rl, SeqCst) ordnen auf einem einzigen Hart nichts.
  • Sie kosten eine Schaltkreisfamilie. Eine atomare Operation fügt dem Programm die Familie ATOMICS hinzu, das dann mindestens einen Shard davon beweist.

Empfohlen: gewöhnliche Variablen, Cell und RefCell für Zustand in neuem Code von Gastprogrammen verwenden. Zu vermeiden: AtomicU32, Mutex-artige Spinlocks oder Arc in ein Gastprogramm aufnehmen, das keinen zweiten Thread hat, mit dem es sie teilen könnte.

Eingaben und Ausgaben#

Hilfsdaten prüfen, bevor irgendetwas daraus Abgeleitetes ins Journal gelangt#

Hilfsdaten sind Speicher, den der Prover gefüllt hat, und nichts bindet sie. Prüfen Sie sie gegen etwas, das der Beweis sehr wohl bindet, etwa einen Hash oder eine Merkle-Wurzel in der öffentlichen Eingabe, eine Signatur oder eine Eigenschaft des Ergebnisses, bevor Sie irgendetwas festschreiben, das von ihnen abhängt. Wer eine Funktion ungeprüfter Hilfsdaten festschreibt, veröffentlicht einen Wert, den der Prover gewählt hat. Siehe das Muster.

Öffentliche Werte klein halten, on-chain mit fester Größe#

Eingabe und Journal fassen jeweils höchstens 16.380 Byte. commit endet mit Exit 70, statt überzulaufen. Große Eingaben gehören in Hilfsdaten hinter einem Commitment, wachsende Ausgaben hinter einen Digest. Ein Verifier-Contract ist für eine Eingabelänge und eine Journal-Länge gebaut; ein Gastprogramm, das auf Ethereum abgewickelt wird, sollte also ein Journal fester Größe veröffentlichen.

Festlegen, wie ein Fehlschlag aussieht#

Ein Gastprogramm, das mit einem Status ungleich null endet oder einen Panic auslöst, hat trotzdem einen gültigen Beweis dessen, was es getan hat, und ein Verifier liest den Status vor dem Journal. Geben Sie jeder Zurückweisung einen eigenen Exit-Code außerhalb der Codes des SDK (70, 71, 72 und 101), und schreiben Sie nichts aus ungeprüften Daten Abgeleitetes fest, bevor die Prüfungen gelaufen sind, die es zurückweisen können.

Keine Außenwelt#

Ein Gastprogramm hat keine Uhr, keinen Zufall, kein Netzwerk, keine Dateien und keine Umgebung. Ein Bibliotheksaufruf, der den Host nach einem davon fragt, erhält -ENOSYS als Antwort und macht den Lauf unbeweisbar. Initialisieren Sie HashMap mit einem deterministischen Seed oder verwenden Sie BTreeMap; leiten Sie Zufall aus der Eingabe ab, wenn ein Algorithmus ihn braucht; übergeben Sie die Zeit als Eingabe.

Kosten#

Jeder ausgeführte Befehl ist eine bewiesene Zeile#

Die Beweiskosten folgen der Zyklenzahl, Familie für Familie. Bauen Sie mit --release, messen Sie mit dem Profiler und behandeln Sie Zyklen so, wie Embedded-Entwickler Bytes behandeln.

Delegieren, was einen Schaltkreis hat#

keccak256, sha256, Addition und Multiplikation auf elliptischen Kurven, Poseidon2, Körperarithmetik von BN254 und modulare 256-Bit-Multiplikation haben eigene Schaltkreise. Erreichen Sie sie über guest_sdk und die mitgelieferten k256, ark-ff und revm-precompile, nicht über eine in das Gastprogramm kompilierte Software-Implementierung. Siehe Delegationen.

Verifizieren statt berechnen#

Ist ein Ergebnis teuer zu finden und günstig zu prüfen, lassen Sie den Prover es finden und als Hilfsdaten übergeben, und lassen Sie das Gastprogramm es prüfen: eine Sortierreihenfolge, eine Faktorisierung, ein Inverses, einen Pfad durch einen Baum, ein Suchergebnis.

Gleitkommaarithmetik vermeiden#

Das Target hat keine F- oder D-Erweiterung; f32 und f64 werden also zu Ganzzahl-Softwareroutinen kompiliert. Sie sind korrekt und deterministisch, und jede Operation kostet viele Befehle. Verwenden Sie Ganzzahlen oder Festkomma.

Höhen kosten ganze Shards#

Eine Familie mit Zeilen kostet mindestens einen Shard ihrer Höhe, wie wenige Zeilen sie auch füllt. Ein Programm, das eine Familie ein einziges Mal berührt, bezahlt einen ganzen Shard; die Familien, die Ihr Code verwendet, und die Höhen, die Sie wählen, bestimmen die Untergrenze jedes Beweises. Siehe Höhen.

Code und Identität#

Der Befehlsstrom ist das Image#

Code ist statisch: Der Befehl jedes pc stammt aus den beim Laden gebauten dekodierten Tabellen, nie aus dem RAM. Ein Schreibzugriff auf .text ändert Daten, nicht das Verhalten, und ein Sprung auf ein Halbwort ohne Befehl beendet den Lauf ohne Beweis. Es gibt keinen JIT und keinen selbstmodifizierenden Code.

Ein illegales Wort irgendwo, und das Programm wird zurückgewiesen#

Der Decoder nimmt sich ganz .text vor, erreichbar oder nicht. Ein CSR-Zugriff, fence.i, eine Gleitkomma- oder RV64-Kodierung in Inline-Assembly oder in .text assemblierte Daten führen dazu, dass die Ableitung das gesamte Programm zurückweist. ebreak wird dekodiert, hat aber keinen Beweis.

Überlaufprüfungen sind Teil des Programms#

Die Profile der Gastprogramme lassen overflow-checks auch im Release-Build aktiv, weil das Abschalten ändert, was ein Gastprogramm berechnet: u32::MAX + 1 würde umlaufen und mit Exit 0 enden, statt einen Panic auszulösen. Verwenden Sie wrapping_*, checked_* und saturating_* dort, wo Sie genau das meinen.

Ein Build ist eine Identität#

Die Identität bindet jedes Byte an Code und Daten, den Einsprungpunkt und jede Höhe. Ein Neubau auf einer anderen Maschine erzeugt eine andere Identität, weil das ELF absolute Pfade einbettet. Registrieren und verteilen Sie das ELF, das Sie bewiesen haben, nicht den Befehl, der es erzeugt hat.

Checkliste#

Bevor Sie beweisen:

  • cargo build --release für riscv32imac-unknown-none-elf, Zyklenbericht des Profilers durchgesehen
  • kein usize in irgendetwas, das festgeschrieben, gehasht oder serialisiert wird
  • keine Allokation in heißen Schleifen; wachsende Collections mit with_capacity dimensioniert
  • keine Atomics, Locks oder Arc in eigenem Code des Gastprogramms
  • jede Verwendung von Hilfsdaten gegen etwas geprüft, das der Beweis bindet, vor jedem Festschreiben, das davon abhängt
  • Journal beschränkt, und von fester Größe, wenn es on-chain abgewickelt wird
  • Hashing und Kurvenarithmetik über Delegationen geleitet
  • eigene Exit-Codes für jede Zurückweisung
  • Host-Build und Emulator stimmen für Ihre Testeingaben im Journal überein
  • die Programmidentität aus dem ELF festgehalten, das Sie ausliefern werden

Ihre App starten

Fehlerbehebung

Jede Art, auf die ein Gastprogramm vor einem verifizierten Beweis stehen bleibt, nach Symptom geordnet. Exit-Status, fatale Fehler des Executors, zurückgewiesene ELFs, zurückgewiesene Programme, fehlgeschlagene Beweise und Fehler des Verifiers, jeweils mit Ursache und Abhilfe.

Als Markdown anzeigen

Ein Gastprogramm (Guest) kann an sechs Stellen vor einem verifizierten Beweis stehen bleiben. Suchen Sie das Symptom, dann die Zeile.

Der Lauf endet mit einem Status, den Sie nicht erwartet haben#

Der Lauf ist beendet und beweisbar; das Gastprogramm hat sich für einen Fehlschlag entschieden. Die eigenen Statuswerte des SDK:

Status Ursache Was zu tun ist
70 commit würde 16.380 Byte überschreiten Schreiben Sie statt der Ausgabe einen Digest der Ausgabe fest
71 eine Allokation würde oberhalb von __stack_top − 8 MiB oder oberhalb des aktuellen sp enden Nichts wird freigegeben; die gesamte Allokation des Laufs muss also zwischen das Image und 0x7F80_0000 passen. Verwenden Sie Puffer über Schleifen hinweg wieder und dimensionieren Sie sie mit with_capacity (der Heap)
72 eine Delegation hat etwas geantwortet, das ihr Shim zurückweist: einen Fehler oder -ENOSYS nach dem ersten Aufruf einer Operation aus mehreren Aufrufen Verwenden Sie die Funktionen des SDK statt roher Frames und halten Sie Operanden unterhalb ihres Moduls
101 ein Panic, der nichts ausgibt Lassen Sie dieselben Eingaben durch den Host-Build Ihrer Bibliothek laufen, wo die Panic-Meldung ausgegeben wird (auf dem Host testen)

Jeder andere Status ist Ihr eigenes exit(code).

Der Lauf bricht mit einem fatalen Fehler ab#

Der Emulator gibt einen EmuError zurück, und es gibt weder Exit-Status noch Beweis:

Fehler Übliche Ursache
OutOfBounds ein Null- oder wilder Zeiger ([0, 0x8000) ist ein Loch), ein Lesezugriff auf Hilfsdaten (Advice) jenseits dessen, was der Host geliefert hat, advice() in einem Lauf ohne Hilfsdaten oder ein Delegations-Frame, der nicht vollständig im RAM liegt
Misaligned ein Halbwort- oder Wortzugriff über einen nicht ausgerichteten Zeiger oder ein nicht ausgerichteter Delegations-Frame
NotAnInstruction ein Sprung zu einem pc, der keinen Befehl enthält, einschließlich des aus Nullen bestehenden Halbworts c.unimp
IllegalInstruction eine Kodierung, die die Maschine nicht ausführt
Ebreak ein ebreak, das keinen Beweis hat
ClockOverflow mehr als 2^36 − 1 Zyklen
PublicInputTooLong, JournalTooLong eine Eingabe oder das Längenwort des Journals beim Exit über 16.380 Byte
DelegationFrame ein Frame, für den sein Schaltkreis keinen Witness hat: ein Operand von MOD_MUL oder EC_ADD, der seinen Modul erreicht oder übersteigt, ein Selektor, der nichts benennt, eine Keccak-Runde über 23, eine SHA-256-Gruppe über 15, eine Poseidon2-Lane, die p erreicht oder übersteigt
DelegationFamilyAbsent die Nummer einer vom Image nie deklarierten Delegation, auf einem Tracing-Pfad

Das ELF wird zurückgewiesen#

artifact-dump, host::setup und die Werkzeuge weisen ein ELF, das der Loader nicht annehmen kann, mit einem LoaderError zurück:

Zurückweisung Übliche Ursache
NotAnElf, Truncated nicht das ELF des Gastprogramms: eine .d-Datei, ein unvollständiger Schreibvorgang
NotRiscV, UnsupportedElfType, RelocatableElf, DynamicElf ein Host-Build, eine Objektdatei, ein PIE oder ein dynamisch gelinkter Build
BadSegment, NoExecutableSegment, EntryNotAnInstruction eine bearbeitete link.ld oder kein gelinktes _start
RvcIllegal, InstructionTooLong, TextTruncated Daten in .text, etwa eine Tabelle in von Hand geschriebenem Assembly. Nie das Null-Halbwort, das zu erwarten ist

Das Programm lässt sich nicht registrieren#

Das ELF lädt, aber das Programm lässt sich nicht in eine Konfiguration dekodieren:

Zurückweisung Ursache und Abhilfe
Not all opcodes supported: pc=… ein Wort außerhalb von RV32IMA irgendwo im ausführbaren Code, erreichbar oder nicht, etwa ein CSR-Zugriff oder fence.i in Assembly. Entfernen Sie es
TableTooShort Code jenseits der Tabellenreichweite einer Familie. Erhöhen Sie die Höhe dieser Familie: 2^22 reicht für 7,9375 MiB Code
ProgramTooLarge das Image überschreitet bytecode_size_words, standardmäßig 4 MiB. Erhöhen Sie die Obergrenze in ProgramParams
ImageOutsideWindow ein in der Datei enthaltenes Byte liegt bei Ihrer Fensterhöhe jenseits von RAM-Fenster 0. Erhöhen Sie die Fensterhöhe
HeightNotOnMenu eine Höhe, die nicht 2^8, 2^12, 2^16, 2^18, 2^20 oder 2^22 ist
UnknownDelegation das Image deklariert eine Delegationsnummer, auf die keine Familie antwortet

Ein Schlüssel lässt sich außerdem nicht bauen, wenn eine Familie, die das Programm verwendet, unter ihre Untergrenze gesetzt ist, 2^20 für eine Befehlsfamilie und 2^16 für die RAM-Fensterfamilien, weil dort kein Schaltkreis existiert.

Der Prover schlägt fehl#

Ein ehrlicher Prover, der beweist, was der Emulator ausgeführt hat, schlägt nicht fehl; ein Fehlschlag deutet also auf eine Eingabe hin, die er nicht hätte annehmen dürfen, oder auf einen Bug. Zwei Fälle machen den Großteil aus:

  • Ein Aufruf ohne Beweis. Ein Systemaufruf außerhalb von EXIT und den Delegationen, den eine Bibliothek für Daten des Hosts abgesetzt hat, erhält -ENOSYS als Antwort, und der Lauf geht weiter, aber das Befüllen im Prover weist diese Zeile zurück und nennt den Zyklus. Suchen Sie die Abhängigkeit, die nach Zufall oder Zeit fragt.
  • Speicher erschöpft. Der Prozess wird beendet, während Shards in Bearbeitung sind. Verringern Sie das dritte Argument von host::prove oder die Höhen.

Für alles andere bauen Sie mit dem Debug-Log des Provers neu und führen erneut aus:

sh
cargo run --release -p bench --features prover/debug-info -- prove ...
APOGEE_DEBUG=detail <the run> 2>&1 | tee run.log
grep -c 'begin h=' run.log; grep -c 'gkr done' run.log     # unequal: a shard died
grep 'begin h=' run.log | tail -1                           # which one
grep -E 'FAIL|NOT CANONICAL|UNBALANCED|OVER the|NAMES NO|DISAGREES|NOT LOOPING|ABORTED' run.log

self_check FAILED nennt das erste fehlschlagende Gate der Zeile und den Wert jedes Operanden. Das Debug-Log listet jeden Marker auf.

Die Verifikation schlägt fehl#

verify_shard und verify_block geben einen VerifyError zurück, dessen Klasse angibt, was fehlgeschlagen ist:

Klasse Bedeutung
Statement die Aussage passt nicht zum Schlüssel: Shard-Zahlen, Fensterregeln, Nutzlastlängen, Listen oder der globale Digest, mit dem der Beweis initialisiert wurde
Malformed die Form des Beweises ist nicht die seines Schaltkreises: Anzahl der Commitments, Ausgaben, Runden oder Behauptungen
Constraint { layer } ein Gate ist verletzt, oder der Sumcheck einer Schicht schlägt fehl
Lookup { channel } ein nachgeschlagenes Tupel steht in keiner Zeile seiner Tabelle
MemoryArgument die Lese- und die Schreib-Multimenge gleichen sich nicht ab, oder ein öffentliches Fenster enthält nicht die Bytes der Aussage
Opening eine Öffnung eines Commitments schlägt fehl

Wenn Sie den Beweis mit dem eigenen Prover von Apogee aus einer Ausführung erstellt haben, die der Emulator akzeptiert hat, bedeutet ein Verifikationsfehler, dass Verifier und Prover sich über das Programm uneinig sind: Prüfen Sie, dass Sie den Schlüssel laden, unter dem der Beweis erstellt wurde, mit denselben Höhen und über derselben Zeremonie.

Abweichende Identität#

Die Identität, die Sie berechnet haben, weicht von der erwarteten ab:

  • Der Build einer anderen Maschine. Das ELF bettet absolute Pfade ein; ein Neubau an anderer Stelle ist also ein anderes Image. Vergleichen Sie mit dem ELF, das bewiesen wurde, nicht mit einem frischen Build.
  • Andere Höhen. Jede Höhe ist in die Identität eingebunden. artifact-dump tables meldet die Identität bei den Standardhöhen; Ihr Setup verwendet möglicherweise andere.
  • Eine andere Zeremonie. Eine Powers-of-Tau-Datei von Hermez hat ein anderes τ; jedes Commitment ist also anders. Prüfen Sie das [τ]_1 der Datei gegen das der Zeremonie.

Ihre App starten

Beispiel-Gastprogramme

Die Gastprogramme im Repository, jedes ein ausgearbeitetes Beispiel für einen Teil des Guest-SDK oder der Maschine. Wo Sie das Muster finden, das Sie brauchen.

Als Markdown anzeigen

Der Workspace guests/ enthält jedes Gastprogramm (Guest), das das Repository baut und testet. Jedes existiert, um etwas zu erproben, und das macht sie zur besten Referenz für das Muster, das Sie gerade schreiben wollen. Alle bauen mit cargo build --target riscv32imac-unknown-none-elf aus ihrem eigenen Verzeichnis.

Hier beginnen#

Gastprogramm Zeigt
public-io die drei Speicherbereiche auf einmal: Hilfsdaten (Advice), gegen die öffentliche Eingabe geprüft, bevor irgendetwas festgeschrieben wird. Das I/O-Modell in einem kleinen Programm
fib das kleinste SDK-Gastprogramm: ein u32 hinein, ein u32 heraus, umlaufende Arithmetik
echo, heap den Allokator: Hilfsdaten durch Heap-Puffer kopiert, Vec und Box immer wieder über den Bump-Allokator angelegt und verworfen

Anwendungsmuster#

Gastprogramm Zeigt
amm, orderbook 128- und 256-Bit-Ganzzahlen ohne Heap; BTreeMap, Sortieren und eine als Hilfsdaten gelieferte Sortierreihenfolge, die geprüft statt berechnet wird
vault, recursion-ops crates/field und crates/transcript in einem Gastprogramm, die an FR_ARITH und POSEIDON2 delegieren, ohne dass ein Shim benannt wird
revm-block Ethereum-Blöcke auf revm: die Binaries revm-block (ein aufgezeichneter Mini-Block) und revm-block-stateless (der zustandslose Validator)

Delegationen#

Gastprogramm Zeigt
keccak-test, sha256-ops, mod-mul-ops, ec-ops KECCAK_F, SHA256_COMP, MOD_MUL und EC_ADD, jeweils im Gastprogramm gegen unabhängige Werte geprüft
keccak-unused, recursion-unused gelinkte, nie aufgerufene Shims: Die Familien sind deklariert und beweisen null Shards

Die Maschine selbst#

Gastprogramm Zeigt
atomics jeden Befehl der A-Erweiterung, so wie core::sync::atomic ihn erzeugt. Ein einziger Hart bedeutet, dass jeder ein gewöhnliches Read-Modify-Write ist; das Gastprogramm existiert, um die Familie zu testen, nicht um die Praxis zu empfehlen
opcodes jeden Befehl von RV32IMAC
rvc-dense die Expansion komprimierter Befehle: eine Sequenz, komprimiert und unkomprimiert assembliert
addsub, control, alu, mem, shards handgeschriebenes Assembly mit eigenem _start und ohne SDK, das mit seinem Ergebnis endet. shards füllt zwei 2^20-Shards
recursion die Verifier-Programme des Rekursionsbaums, die Binaries leaf und node

Das kleinste nützliche Gastprogramm#

fib liest ein u32, macht so viele Fibonacci-Schritte mit umlaufender Arithmetik und schreibt das Ergebnis fest:

guests/fib/src/main.rsrust
#![no_std]
#![no_main]

guest_sdk::entry!(main);

fn main() {
    let mut n = [0u8; 4];
    assert_eq!(
        guest_sdk::read_input(&mut n),
        4,
        "fib: the public input is one u32"
    );
    let n = u32::from_le_bytes(n);

    let mut a: u32 = 0;
    let mut b: u32 = 1;
    for _ in 0..n {
        let next = a.wrapping_add(b);
        a = b;
        b = next;
    }
    guest_sdk::commit(&a.to_le_bytes());
}

Eine zu kurze Eingabe ist ein Fehler und kein Fall für einen Standardwert: Ein Gastprogramm, das mit einem teilweise gefüllten Puffer weitermacht, beweist eine Aussage über Nullen. Die Addition läuft absichtlich um; ein n über 47, jenseits des letzten Glieds, das in 32 Bit passt, ist also eine gewöhnliche Eingabe mit einer gewöhnlichen Antwort und kein fehlgeschlagener Lauf. Und es gibt keine Hilfsdaten, weil f_n einen Verifier beim Prüfen genauso viel kostet wie beim Berechnen: Eine als Hilfsdaten gelieferte Antwort müsste neu berechnet werden, damit man ihr glauben kann.

Ihre App starten

Referenz des Guest-SDK

Jedes öffentliche Element des Crates guest-sdk, mit seiner genauen Signatur und seinem Verhalten. Nur exit und die Delegations-Shims setzen einen ecall ab; alles andere sind Lade- und Speicherbefehle.

Als Markdown anzeigen

crates/guest-sdk ist die gesamte Laufzeitumgebung eines Gastprogramms (Guest): der Startup-Code, das Einsprungmakro, der Allokator, der Panic-Handler und die ecall-Shims. Es kompiliert nur für riscv32imac-unknown-none-elf.

Einsprung#

rust
guest_sdk::entry!(main);

Exportiert das Symbol main, das der Startup-Code aufruft, als Wrapper, der Ihre Funktion aufruft; diese nimmt keine Argumente und gibt () zurück. Ihre Funktion behält ihren eigenen Namen und darf selbst main heißen. Eine Rückkehr aus ihr ist gleichbedeutend mit exit(0).

Die Speicherbereiche#

Element Signatur Verhalten
public_input fn public_input() -> &'static [u8] Die Nutzlast der öffentlichen Eingabe, ihr Längenwort auf das Fenster begrenzt. Keine Kopie, kein ecall
read_input fn read_input(buf: &mut [u8]) -> usize Kopiert min(buf.len(), public_input().len()) Bytes und gibt die Anzahl zurück. Es kann weniger liefern als angefordert
advice fn advice() -> &'static [u8] Die Nutzlast der Hilfsdaten (Advice), ihre Länge auf den Bereich begrenzt. An nichts gebunden, daher prüft das Gastprogramm sie. Fataler OutOfBounds-Fehler in einem Lauf ohne Hilfsdaten
commit fn commit(bytes: &[u8]) Hängt an das Journal an und aktualisiert sein Längenwort. Endet mit Exit 70, statt das Fenster von 16.380 Byte überlaufen zu lassen
journal fn journal() -> &'static [u8] Alles bisher Festgeschriebene
exit fn exit(code: i32) -> ! Beendet den Lauf mit code als Exit-Status der Aussage. Veröffentlicht nichts über das Festgeschriebene hinaus

Hashing#

Element Signatur Verhalten
keccak256 fn keccak256(input: &[u8]) -> [u8; 32] Das Keccak-256 von Ethereum, nicht SHA3-256. Sponge und Padding laufen im Code des Gastprogramms; jede Runde von keccak-f[1600] ist ein KECCAK_F-Aufruf. Software-Fallback, wenn der erste Aufruf -ENOSYS als Antwort erhält
sha256 fn sha256(input: &[u8]) -> [u8; 32] SHA-256 nach FIPS 180-4. Padding und die Blockschleife laufen im Code des Gastprogramms; jede Kompression sind sechzehn SHA256_COMP-Aufrufe. Software-Fallback wie oben
poseidon2_permute fn poseidon2_permute(state: &mut [u8; 96]) -> bool Die Poseidon2-Permutation der Breite 3 über drei kanonischen Fr-Lanes in Little-Endian, an Ort und Stelle, über POSEIDON2. Gibt bei -ENOSYS false zurück, für den eigenen Softwarepfad des Aufrufers

Elliptische Kurven#

rust
pub type ProjectivePoint = [[u32; 8]; 3];

Ein Punkt in homogen projektiven Koordinaten, x = X/Z und y = Y/Z, jede Koordinate acht Little-Endian-Limbs zu 32 Bit unterhalb des Körpermoduls der Kurve. Das sind keine Jacobi-Koordinaten: Projective von arkworks verwendet solche; ein Aufrufer, der von dort konvertiert, bildet also beim Hineinkonvertieren auf (X·Z, Y·Z², Z) und beim Zurückkonvertieren auf (X·Z, Y, Z³) ab. Das neutrale Element ist (0 : 1 : 0).

Element Signatur Verhalten
ec_add fn ec_add(codes: &[u32; 3], p: &ProjectivePoint, q: &ProjectivePoint) -> Option<ProjectivePoint> p + q nach der vollständigen Formel, über drei EC_ADD-Aufrufe in der Reihenfolge der Gruppen. None bei -ENOSYS
ec_mul fn ec_mul(codes: &[u32; 3], p: &ProjectivePoint, k: &[u32; 8]) -> Option<ProjectivePoint> k·p per Double-and-Add ab dem höchsten Bit. k wird so verwendet, wie es übergeben wird; es modulo der Gruppenordnung zu reduzieren, ist Sache des Aufrufers
ec_identity fn ec_identity() -> ProjectivePoint (0 : 1 : 0)
recursion::SECP256K1_GROUPS, recursion::BN254_GROUPS [u32; 3] Das Argument codes: welche Kurve, ausgedrückt als die drei Gruppenselektoren einer Addition

Die Formel beweist Arithmetik, nicht die Zugehörigkeit zur Kurve: Prüfen Sie Punkte, die aus Hilfsdaten stammen, selbst.

Rohe Delegations-Shims#

guest_sdk::recursion enthält die Shims über wortausgerichteten Frame-Typen. Jeder Frame-Typ ist #[repr(C, align(4))]; seine Ausrichtung ergibt sich also aus dem Typ und nicht daraus, wo der Codegenerator gerade eine lokale Variable abgelegt hat. Ein Shim im Basisformat gibt genau bei -ENOSYS false zurück; jede andere Antwort ungleich null führt zu Exit 72.

Element Zweck
mod_mul(&mut ModMulFrame) -> bool Ein a·b mod m. Bauen Sie den Frame mit ModMulFrame::of(modulus, &a, &b) und lesen Sie frame.result(). Die Modulcodes sind SECP256K1_P, SECP256K1_N, BN254_P und BN254_R, und beide Operanden müssen bereits unterhalb des Moduls liegen
sha256_comp(&mut Sha256Frame) -> bool Eine ganze Kompression: sechzehn Aufrufe in Reihenfolge. Sha256Frame::of(&state, &block), dann frame.working(); das Ergebnis zum Verkettungswert zu addieren, ist Sache des Aufrufers
ec_add_complete(&mut EcAddFrame, &[u32; 3]) -> bool Eine vollständige Addition: drei Aufrufe in der Reihenfolge der Gruppen. EcAddFrame::of(&codes, &p, &q), dann frame.result()
poseidon2(&mut Poseidon2Frame) -> bool, fr_arith(&mut FrArithFrame) -> bool Die Permutation und eine Fr-Operation über Byte-Frames; field und transcript rufen sie für Sie auf
sha256_rounds, ec_add Einzelne Schritte der obigen Operationen. Ein Schritt in falscher Reihenfolge wird nicht zurückgewiesen, er berechnet etwas anderes; bevorzugen Sie daher die Funktionen für ganze Operationen
fr_op, p2_field, field_io, fq_op, import, import_run, replay Die Koprozessor-Aufrufe des Rekursionsformats, die die eigenen Programme des Rekursionsbaums verwenden. Sie haben keinen Softwarepfad

Jeder Shim liest seine ecall-Nummer aus dem Deklarationsdatensatz seiner Familie, einem 12-Byte-static in einer eigenen Linker-Sektion. Einen Shim zu linken deklariert die Familie; eine deklarierte Familie, die nie aufgerufen wird, beweist null Shards.

Laufzeitverhalten#

Komponente Verhalten
Start _start bei 0x0001_0000 setzt sp auf __stack_top (0x8000_0000), füllt .bss Byte für Byte mit Nullen, ruft main auf und endet mit Exit 0, wenn es zurückkehrt
Allokator Wächst ab __heap_start nach oben und gibt nie Speicher frei. Endet mit Exit 71, wenn ein Block oberhalb von __stack_top − 8 MiB oder oberhalb des aktuellen sp enden würde
Panic-Handler Endet mit Exit 101 und schreibt nichts. Ein Gastprogramm, das einen Panic auslöst, ist beweisbar und hat veröffentlicht, was es festgeschrieben hat
Exit-Status 70 Überlauf des Journals, 71 Heap erschöpft, 72 eine Delegation hat mit einem Fehler geantwortet, 101 Panic

Transparente Delegation#

Zwei Bibliotheks-Crates des Repositorys delegieren auf dem Target der Gastprogramme, ohne das SDK zu benennen, über eine nur für dieses Target geltende Abhängigkeit davon:

  • field::Fr: Addition, Montgomery-Multiplikation (*, square, pow und die Konvertierungen) und inverse für Werte ungleich null rufen FR_ARITH auf. Ein Gastprogramm, das Fr-Arithmetik verwendet, deklariert diese Familie.
  • transcript::poseidon2_permute ruft POSEIDON2 auf.

Die mitgelieferten k256, ark-ff und revm-precompile tun dasselbe für secp256k1, BN254 und die Precompiles der EVM: Delegationen.

Die Spezifikation des ABI, das all dem zugrunde liegt, ist Gastprogramm-ABI.

Ihre App starten

KI-Begleiter

Eine einzige Datei, die ein KI-Modell auf das Schreiben von Gastprogrammen für Apogee vorbereitet. Laden Sie sie herunter, legen Sie sie Ihrem Modell vor, und es beginnt mit denselben Regeln, die dieses Handbuch lehrt.

Als Markdown anzeigen

Ein Großteil des Codes, der für Apogee geschrieben wird, wird von einem Modell entworfen werden. Ein Modell, das Apogee nie gesehen hat, schreibt ein plausibles Gastprogramm (Guest), das std verwendet, in jeder Schleife alloziert, zu einem atomaren Zähler greift, seinen Hilfsdaten (Advice) vertraut und ein usize festschreibt. Der KI-Begleiter ist eine einzige Markdown-Datei, die alles vorab bereitstellt, was diese Fehler verhindert: was ein Gastprogramm ist, die harten Regeln, die vollständige öffentliche Oberfläche des SDK mit exakten Signaturen, Muster zum Übernehmen, die Fehler und ihre Behebung sowie eine Checkliste für das Review.

Begleiter herunterladen Als Text öffnen

So verwenden Sie ihn#

  • In einem Chat: Hängen Sie die Datei an oder fügen Sie sie als erste Nachricht ein, bevor Sie beschreiben, was gebaut werden soll.
  • In einem Coding-Agenten: Speichern Sie sie im Wurzelverzeichnis Ihres Projekts unter dem Namen, den Ihr Werkzeug per Konvention liest, etwa AGENTS.md oder CLAUDE.md, oder fügen Sie sie den Projektregeln des Werkzeugs hinzu. Der Agent liest sie dann zu Beginn jeder Sitzung.
  • Für das Review: Bitten Sie das Modell, ein Gastprogramm Zeile für Zeile gegen Abschnitt 8 der Datei zu prüfen, die Review-Checkliste.

Die Datei formuliert ihre Regeln als MUST und MUST NOT, jeweils mit der Begründung daneben, weil Modelle expliziten und begründeten Vorgaben zuverlässiger folgen als Konventionen, die sie selbst erschließen sollen.

Was er enthält#

Abschnitt Inhalt
0. Anweisungen an das Modell Die Regeln als harte Vorgaben behandeln; nie eine API aufrufen, die nicht aufgeführt ist; Beweise sind nicht zero-knowledge
1. Was ein Gastprogramm ist Das Target, der einzelne Hart, was ein Beweis aussagt, die Programmidentität, die drei Speicherbereiche
2. Harte Regeln 23 Regeln: Programmform, usize und Zeiger mit 32 Bit, der Bump-Allokator, der Stack, Ausrichtung, Atomics, die fehlende Außenwelt, Hilfsdaten, Grenzen der öffentlichen Werte, der Befehlssatz, Gleitkommazahlen, Überlaufprüfungen, Kosten
3. Aufbau und Build Die Crate-Vorlagen, der Gastprogramm-Workspace, der Build-Befehl, die Aufteilung in eine zuerst auf dem Host getestete Bibliothek
4. Das Guest-SDK Jede öffentliche Funktion mit ihrer exakten Signatur, die Fakten zur Laufzeitumgebung und die Speicherkarte, die delegierten Operationen und die mitgelieferten Crates
5. Muster Hilfsdaten, gegen einen Hash geprüft; eine Merkle-Abfrage und ein Zustandsübergang; Wiederverwendung von Puffern; strukturierte Hilfsdaten; Digests für wachsende Ausgaben
6. Host-Seite Ausführung im Emulator, Profiling, Export des Images, Beweisen und Verifizieren, Höhen und gleichzeitig bearbeitete Shards
7. Fehler und Abhilfen Jeder Exit-Status, jeder fatale Fehler und jede Zurückweisung, auf die ein Gastprogramm trifft, mit Ursache und Abhilfe
8. Review-Checkliste Elf Prüfungen, die vor dem Vorschlagen von Code für Gastprogramme durchzuführen sind
9. Fakten Die ISA, das Beweissystem, das Sicherheitsniveau, die Grenzen und die gemessenen Ergebnisse

Die Regeln, auf denen er besteht#

Der Begleiter wiederholt die Regeln dieses Handbuchs, und drei davon verdienen es, hervorgehoben zu werden, weil Modelle sie am häufigsten falsch machen:

  • Zeiger und usize sind 32 Bit breit. Ein Modell, das überwiegend mit 64-Bit-Code trainiert wurde, serialisiert usize, ohne zweimal nachzudenken. Gastprogramm und Host sind sich dann über die Bytes uneinig.
  • Der Heap gibt nie Speicher frei. Idiomatisches Rust alloziert großzügig, weil ein echter Allokator Speicher zurückgibt. Hier ist jede Allokation für den Rest des Laufs dauerhaft; der Begleiter verlangt deshalb überall Wiederverwendung und with_capacity.
  • Atomics kompilieren, und sie gehören nicht in neuen Code von Gastprogrammen. Sie werden für die Kompatibilität mit bestehenden Bibliotheken unterstützt. Auf einem einzigen Hart synchronisieren sie nichts, und sie fügen dem Beweis eine Schaltkreisfamilie hinzu.

Für Agenten, die diese Dokumentation direkt lesen#

  • /llms.txt verzeichnet jede Seite dieser Website in Markdown, für Modelle, die das Web lesen.
  • /docs/llms-full.txt ist die gesamte englische Dokumentation, einschließlich der Spezifikation, in einer einzigen Datei.
  • Jede Seite hat einen Link Als Markdown anzeigen und eine Schaltfläche Als Markdown kopieren, unter ihrem Titel und in der rechten Spalte.

Englisch ist die kanonische Sprache der Dokumentation, und der Begleiter wird für jede Sprachversion auf Englisch veröffentlicht: Es ist die Sprache, der Modelle am zuverlässigsten folgen, und die, in der die Spezifikation geschrieben ist.

Architektur

Architektur

Apogee VM von Anfang bis Ende. Was ein Beweis aussagt, der Weg von einem Gastprogramm-Binary zu einem Contract-Aufruf, wie die großen Komponenten zusammenspielen und welche Designentscheidungen sie prägen.

Als Markdown anzeigen

Apogee beweist Ausführungen von RV32IMAC-Programmen. Dieser Abschnitt beschreibt das System auf der Ebene seiner großen Komponenten: was jede davon tut, warum sie so gebaut ist, wie sie gebaut ist, und wie sie an die nächste übergibt. Der Abschnitt Auditoren enthält dasselbe System auf der Ebene jeder Spalte und jedes Gates.

Was ein Beweis aussagt#

Ein Verifier besitzt drei Dinge, die er nicht auf das Wort des Provers hin übernimmt:

  • die Programmidentität, ein einziges Körperelement, das die Befehlstabellen des Programms, sein anfängliches Speicher-Image, seinen Einsprung-pc und seine Konfiguration zu einem Digest verdichtet;
  • den SRS-Digest der Zeremonie, den ein Verifikationsschlüssel tragen muss;
  • einen Verifikationsschlüssel, der von beliebiger Seite stammen darf, weil beim Laden die Identität und der SRS-Digest aus seinem eigenen Inhalt neu berechnet und seine Schaltkreise mit der Registry des Verifiers abgeglichen werden.

Die Aussage des Beweises enthält die öffentliche Eingabe, die öffentliche Ausgabe (das Journal), den Exit-Status und die Aufzeichnung der Form der Ausführung: die Shard-Zahlen, die Speicherfenster, die finalen Register und den finalen pc sowie die Speicher-Commitments und Wurzeln jedes Shards. Ein Beweis, der die Verifikation besteht, begründet, dass das Programm dieser Identität, gestartet an seinem Einsprung-pc auf seinem Image, mit der öffentlichen Eingabe in seinem Eingabefenster und irgendwelchen vom Prover gewählten Hilfsdaten (Advice), Befehl für Befehl bis zu EXIT mit diesem Status ausgeführt wird und dabei dieses Journal geschrieben hat. Über die Hilfsdaten wird nichts behauptet, und nichts wird verborgen: Kein Commitment und kein Beweis ist verblindet.

Vom Binary zum Contract-Aufruf#

Apogee VM von Anfang bis Ende Vier Stufen. Programm: vom ELF des Gastprogramms über das ProgramImage zu den dekodierten Tabellen und der VmConfig und weiter zur Programmidentität. Ausführung: Der Emulator erzeugt Zeilen pro Familie, die in Shards geschnitten werden. Pro Shard: die Spalten mit Mercury committen, den GKR-Rückwärtsdurchlauf ausführen, jede Spalte in einem einzigen Batch öffnen. Abwicklung: Blockbeweis, Rekursionsbaum, Groth16-Decider, der Contract ApogeeVerifier. 1 · PROGRAMM Gastprogramm-ELFRV32IMAC, no_std Rust ProgramImageSegmente · RVC expandiert Dekodierte Tabellen7 Familien · VmConfig Programmidentitätein Fr 2 · AUSFÜHRUNG Emulator1 Hart · reine Funktion von (image, io) Zeilen, eine Tabelle pro FamilieZyklen · Aufrufe · Speicherwörter Shards2^8 … 2^22 Zeilen · 23 Familien 3 · JEDER SHARD Spalten committenMercury über KZG GKR-Rückwärtsdurchlaufein Sumcheck pro Schicht Eine gebündelte Öffnungjede Spalte an einem Punkt · 704 B 4 · ABWICKLUNG BlockbeweisAussage + Shard-Beweise RekursionsbaumBlätter · Knoten · Wurzel Groth16-Decidergebundene Wires · BN254 ApogeeVerifier.solzwei Pairings · true
Die vier Stufen. Das Programm steht fest, bevor irgendetwas läuft; die Ausführung wird in Shards geschnitten; jeder Shard wird für sich bewiesen, mit Ausnahme des Speicherarguments, das sich einmal über alle Shards schließt; die Abwicklung komprimiert den Block für einen Contract.
  1. Das Programm. Der Loader liest das ELF in ein ProgramImage ein und expandiert dabei komprimierte Befehle an Ort und Stelle. Der Decoder ordnet jeden Befehl einer von sieben Befehlsfamilien zu, baut die dekodierte Tabelle jeder Familie mit einer Zeile pro Halbwort Code und committet dann all das als Programmidentität. Programme und Identität.
  2. Ausführung. Der Emulator führt das Gastprogramm (Guest) auf einem Hart aus. Ein Zyklus ist eine Zeile der Familie, der sein Befehl gehört, und zeichnet mit Zeitstempeln versehene Lese- und Schreibzugriffe auf pc, Register und RAM auf. Hashing und Arithmetik mit großen Ganzzahlen werden delegiert: Ein ecall benennt einen Frame im RAM, und eine Zeile einer Delegationsfamilie erledigt die Arbeit darauf. Ausführung, Familien und Shards.
  3. Shards. Die Zeilen einer Familie werden in Shards der Höhe dieser Familie geschnitten, einer Zweierpotenz zwischen 2^8 und 2^22. Den Speicher, den eine Ausführung berührt, decken Shards der Fensterfamilien ab, die jedem Wort seinen Anfangs- und seinen Endwert geben. Ein Shard ist die Einheit des Beweisens; ein Block umfasst Hunderte davon.
  4. Der Beweis eines Shards. Seine Spalten werden mit Mercury committet. Den Schaltkreis der Familie durchläuft die GKR-Engine rückwärts, von seinen Ausgaben bis zu diesen Spalten, mit einem Sumcheck pro Schicht, und jede Spalte wird an dem einen Punkt, an dem dieser Durchlauf endet, in einer einzigen gebündelten Öffnung geöffnet.
  5. Der Block. Ein BlockProof besteht aus der Aussage und ihren Shard-Beweisen. Die Verifikation führt einmal das globale Transkript, die Prüfungen jedes Shards und einmal den Speicherabgleich über die Wurzeln aller Shards aus.
  6. Rekursion und Abwicklung. Verifier-Programme, die Apogee selbst beweist, verifizieren Reihen von Shards und falten deren aufgeschobene Pairings. Ein Baum aus ihnen endet in einer Wurzel, ein Groth16-Schaltkreis verifiziert die Wurzel erneut, und ApogeeVerifier.sol prüft diesen Beweis und das gefaltete Pairing. Rekursion und Abwicklung.

Der Prover führt das Gastprogramm zweimal aus: einmal, um die Speicherspalten jedes Shards zu committen, was die Aussage und ihre Challenges festlegt, und einmal, um jeden Shard zu beweisen, sobald er gefüllt ist. Sein Speicherbedarf ist durch die gleichzeitig bearbeiteten Shards beschränkt, nicht durch die Länge der Ausführung. Der Streaming-Prover.

Die Entscheidungen, die das System prägen#

Ein Körper, eine Kurve. Alles liegt über dem Skalarkörper von BN254: die Schaltkreise, das Transkript, die Commitments und die Rekursion. Deshalb kann ein Knoten des Rekursionsbaums Basis-Shards in seiner eigenen Arithmetik verifizieren, und deshalb kann der Baum in einem Groth16-Beweis enden, den Ethereum mit seinem Pairing-Precompile prüft.

Geschichtete GKR-Schaltkreise statt committeter Constraint-Tabellen. Der Schaltkreis einer Familie ist ein Stapel von Gate-Schichten vom Grad 2 über ihren committeten Spalten. Nur die unterste Schicht wird committet; jede Schicht darüber wird in einem einzigen Rückwärtsdurchlauf per Sumcheck bewiesen und nie committet. Der Durchlauf endet mit Behauptungen über jede committete Spalte an ein und demselben Punkt; ein Shard braucht also genau eine Öffnung. Die GKR-Engine erklärt, warum darin die zentrale Ersparnis der Engine liegt.

Eine Öffnung konstanter Größe. Mercury öffnet ein multilineares Commitment mit acht Kurvenpunkten und sechs Körperelementen, 704 Byte, unabhängig von der Größe des Polynoms und davon, wie viele Spalten sich den Punkt teilen. Seine Prüfungen haben die Form e(A, [1]_2) = e(B, [x]_2), die die Rekursion falten kann, statt Pairings zu berechnen.

Ein Speicherargument für die gesamte Ausführung. Jeder Zugriff, in jedem Shard jeder Familie, ist ein Tupel in einer einzigen Lese-/Schreib-Multimenge, und der Verifier gleicht die Produkte einmal pro Aussage ab. Der pc ist eine Zelle dieser Multimenge; Reihenfolge, Kontinuität und Eindeutigkeit der Zyklen über Shards hinweg brauchen also kein weiteres Argument, und kein Shard muss sich mit seinem Nachbarn verketten.

Delegationen als Familien, nicht als Befehle. Eine teure Funktion erhält eine eigene Schaltkreisfamilie, aufgerufen durch einen ecall über einem Frame aus RAM und über dieselbe Multimenge eins zu eins mit ihrer Anfrage gepaart. Die Befehlsschaltkreise bleiben klein, und ein Gastprogramm bezahlt für eine Delegation nur, wenn es sie aufruft.

Streaming statt Materialisierung. Der Trace wäre mit etwa 300 Byte pro Zyklus das größte Objekt im System; deshalb existiert er nie. Der Prover führt zweimal aus und hält nur die Shards, die gerade bearbeitet werden.

Keine geliehene Kryptografie. Körper, Kurve, Pairing, MSM, Hash, Polynom-Commitment, GKR und Groth16 sind im Repository implementiert und Seite für Seite spezifiziert. arkworks, Plonky3 und zkhash kommen nur als Testorakel vor.

Wie sich die Soundness zusammensetzt#

Der GKR-Durchlauf und die Öffnung jedes Shards binden die Ausgaben seines Schaltkreises an committete Spalten. Darüber hinaus erstrecken sich diese Argumente über die gesamte Ausführung:

Behauptung Getragen von
Jede Zeile befolgt ihren Befehl den Constraint-Gates des Familienschaltkreises, null auf jeder Zeile
Der Befehl einer Zeile ist der des Programms an ihrem pc einem Lookup von pc und Feldern der Zeile in der dekodierten Tabelle der Familie, die die Identität committet
Jeder Lesezugriff liefert den letzten Schreibzugriff einer einzigen Multimenge über alle Shards; der Verifier multipliziert die Lese- und Schreibwurzeln jedes Shards mit Randfaktoren für die Register und den pc
Die Zeilen bilden einen einzigen Pfad vom Einsprung-pc zum Exit, in Programmreihenfolge dem pc als Zelle dieser Multimenge, die mindestens vier Zeitstempel nach ihrem Lesen geschrieben wird
Ein Wert ist ein Byte, ein Wort, ein Vorzeichen, ein XOR LogUp-Kanälen über Bereichs-, Byte- und generischen Tabellen
Die öffentliche Eingabe und das Journal sind die behaupteten Bytes den Anfangs- und Endspalten der beiden öffentlichen Fenster, gebunden an die multilinearen Erweiterungen der Bytes
Eine delegierte Berechnung ist die der Funktion Aufrufzeilen, die den Frame über dieselbe Multimenge lesen und schreiben, eins zu eins mit ihrem ecall gepaart

Die Challenges stammen aus einem Poseidon2-Duplex-Transkript. Das globale Transkript absorbiert die gesamte Aussage, einschließlich der Speicher-Commitments jedes Shards, bevor es die Speicher-Challenges gibt; das Transkript jedes Shards wird aus dem Endzustand des globalen Transkripts initialisiert. Die Soundness-Karte führt jede Zeile dieser Tabelle bis zu den Abschnitten, die sie beweisen.

Der Code, nach Schichten#

Schicht Crates Rolle
Arithmetik field, curve, poly, sumcheck Fr; der Fq-Turm, G1, G2, das Pairing, MSM; multilineare Polynome; der Zerocheck
Fiat–Shamir und Setup transcript, srs Poseidon2 und das Duplex-Transkript; Einlesen der Zeremonie, KZG, Phase 1 von Groth16
Commitments pcs, pcs-verify Mercury und seine aufgeschobene Verifikation
Das Programm loader, isa, program ELF zu Image, der Decoder, dekodierte Tabellen, VmConfig und Identität
Ausführung emulator, trace der Executor und seine Tracer; Zeilen, Speicherzustand, Spalten-Builder
Schaltkreise constraints, gkr-verify, gkr jeder Schaltkreis als Daten; der GKR-Verifier und der GKR-Prover
Beweis und Verifikation verifier-core, verifier, prover Aussage, Transkripte, Schlüssel, jede Prüfung; der Streaming-Prover
Abwicklung host, groth16, contracts/ das Host-SDK, der Rekursionsbaum und der Decider; ApogeeVerifier.sol
Absicherung checker, tools/ unabhängige Validatoren, die Manipulations-Suite, Benchmarks, Profiler, Orakel

Der Verifier ist die einzige Partei, der vertraut wird: Der Prover validiert nichts, und eine falsche Eingabe führt bei einem ehrlichen Prover zu einem Beweis, der fehlschlägt. Das Sicherheitsmodell führt genau auf, auf welchen Crates die Soundness beruht.

Architektur

Programme und Identität

Wie aus dem ELF eines Gastprogramms eine statische, dem Verifier bekannte Beschreibung eines Programms wird und warum ein einziges Körperelement als Identität genügt, um einem Verifier zu sagen, auf welches Programm sich ein Beweis bezieht.

Als Markdown anzeigen

Bevor irgendetwas ausgeführt wird, macht Apogee aus dem Binary des Gastprogramms (Guest) eine feste Beschreibung des Programms: sein Image, seine nach Schaltkreisfamilien sortierten Befehle, die Konfiguration, unter der es bewiesen wird, und ein einziges Körperelement, das all das committet. Jeder Schritt ist eine reine Funktion seiner Eingabe.

Laden#

Der Loader akzeptiert ein statisches 32-Bit-RISC-V-Executable im Little-Endian-Format, dessen ladbare Segmente im RAM des Gastprogramms an geraden Adressen liegen, paarweise disjunkt und mindestens eines davon ausführbar. Aus einem Program Header liest er Typ, Offsets, Größen und das Ausführbarkeits-Bit und sonst nichts: Die VM hat keine Speicherseiten und keine Berechtigungen, und der gesamte RAM ist adressierbar, was auch immer die Segmente deklarieren.

Danach durchläuft er jedes ausführbare Segment Halbwort für Halbwort. Ein Halbwort, das binär auf 11 endet, beginnt einen 4-Byte-Befehl; das Null-Halbwort ist ein Nicht-Befehl (LLVM füllt unerreichbare Blöcke damit auf); alles andere ist ein komprimierter Befehl, der an Ort und Stelle zu seiner 32-Bit-Form expandiert wird. Adressen werden nie verdichtet: Ein c.addi bei 0x1002 bleibt dort und belegt zwei Byte; jede vom Linker aufgelöste Adresse bleibt also gültig, und allein die Länge eines Befehls hält fest, ob der nächste pc pc + 2 oder pc + 4 ist.

Ein desynchronisierter Durchlauf kann keinen falschen Befehl beweisbar machen. Ein Slot ist eine Funktion der Bytes an seinem eigenen pc; jeder Befehls-Slot ist also das, was ein Hart dekodieren würde, der dort einen Befehl holt. Daten, die den Durchlauf von den wahren Grenzen verschieben, können nur echte Befehlsanfänge verlieren, deren pcs dann keine Tabellenzeile haben, oder auf eine nicht zugewiesene Kodierung treffen, woraufhin das ganze Image zurückgewiesen wird.

Dekodieren und Zuordnen#

Der Decoder nimmt nur 32-Bit-Wörter entgegen und akzeptiert genau die 59 Befehle von RV32IMA: 40 aus dem Basissatz, 8 aus M, 11 aus A. Alles andere, von RV64-Kodierungen und Gleitkomma bis zu CSRs und fence.i, ist ein Dekodierfehler, und ein einziges nicht dekodierbares Wort irgendwo im ausführbaren Code führt zur Zurückweisung des Programms, ob erreichbar oder nicht. Jeder Befehl wird genau einer von sieben Befehlsfamilien zugeordnet:

Id Familie Befehle
0 ADD_SUB_LUI_AUIPC ecall, ebreak, fence, addi, auipc, add, sub, lui
1 JUMP_BRANCH_SLT slti, sltiu, slt, sltu, die sechs Verzweigungen, jalr, jal
2 SHIFT_BITWISE die sechs Shifts, and, or, xor und ihre Immediate-Formen
3 MUL_DIV die acht Befehle der M-Erweiterung
4 MEM_WORD lw, sw
5 MEM_SUBWORD lb, lh, lbu, lhu, sb, sh
6 ATOMICS lr.w, sc.w und die neun AMOs

Die Gruppierung folgt dem, was die Schaltkreise gemeinsam haben. Ein einziges Vergleichs-Gadget entscheidet die Ordnung mit und ohne Vorzeichen für jede Art von Verzweigung und slt; eine einzige Produktidentität dient allen vier Multiplikationen und der Division; ein Shift in jede Richtung ist ein einziges Produkt mit einer nachgeschlagenen Zweierpotenz.

Dekodierte Tabellen#

Jede Befehlsfamilie erhält eine dekodierte Tabelle: ihre Setup-Spalten, in denen Zeile i zu pc 2i gehört, eine Zeile pro Halbwort des Adressraums, den die Tabelle erreicht. Eine aktive Zeile enthält einen der Befehle der Familie als Tupel pc, next_pc, rs1, rs2, rd, imm, extra_mask, wobei ein einzelnes Bit der Maske das Mnemonic benennt. Jede andere Zeile ist Padding, −1 in jedem Feld; keine aktive Zeile ist also je die Padding-Zeile, und eine Zeile aus lauter Nullen ist nie ein beanspruchbarer Befehl bei pc 0.

Die Zeile jedes Zyklus schlägt sich selbst anhand ihres pc in der Tabelle ihrer Familie nach. Dieser Lookup bindet eine Ausführung an das Programm: Der Befehl einer Zeile ist der Befehl des Programms an diesem pc, und ein pc ohne aktive Zeile in irgendeiner Tabelle kann nicht beweisbar ausgeführt werden. Code ist daher statisch. Ein Store in .text ändert, was ein späterer Load liest, nie, was ausgeführt wird.

Die Konfiguration#

Die statische Form eines Programms ist seine VmConfig: die Familien, die es verwendet, jede mit einer Höhe, und eine Obergrenze für seine Codegröße. Die Menge der Familien wird nicht gewählt, sondern abgeleitet:

  • eine Befehlsfamilie ist vorhanden, wenn das Image einen ihrer Befehle enthält;
  • die fünf Fensterfamilien, die den Speicher initialisieren und abschließen, sind immer vorhanden;
  • eine Delegationsfamilie ist vorhanden, wenn das Image sie deklariert, über einen 12-Byte-Datensatz, den das Linken ihres Shims unter den Bytes des Images hinterlässt.

Die Höhe einer Familie ist die Zahl der Zeilen in einem ihrer Shards, gewählt aus 2^8, 2^12, 2^16, 2^18, 2^20, 2^22. Jede Höhe ist eine Zweierpotenz mit geradem Exponenten, weil eine Mercury-Öffnung das verlangt. Höhen sind ein Parameter des Programms, nicht eines Laufs, und jede Wahl der Höhen ist ein eigenes Programm.

Programmidentität#

Die Programmidentität ist ein einziges Element von Fr: ein Poseidon2-Digest über die Codeversion, die VmConfig, den Einsprung-pc und die Setup-Commitments jeder Familie, also die Mercury-Commitments auf die dekodierten Tabellen und auf die anfänglichen Speicherwörter des Images.

Sie bindet jeden Befehl, den der Durchlauf gefunden hat, mit seinem pc, seiner Länge, seinen Operanden und seiner Art, sowie die Tatsache, dass kein anderer pc einen Befehl enthält; jedes in der Datei enthaltene Byte des Images, was die Delegationsdeklarationen einschließt; den Einsprung-pc; die Menge der Familien, jede Höhe, die Obergrenze der Codegröße und die Codeversion. Sie bindet nicht die Zeremonie oder die generische Lookup-Tabelle, die der SRS-Digest abdeckt; nicht die Schaltkreise, die beim Laden eines Schlüssels mit der Registry des Verifiers abgeglichen werden; nichts, was eine Ausführung wählt; und nichts aus dem ELF, was der Loader nicht liest, etwa die Symboltabelle.

Ihre Berechnung braucht die Potenzen der Zeremonie, um zu committen. Ihre Prüfung braucht nur die Commitments, die ein Verifikationsschlüssel mitführt: Beim Laden eines Schlüssels wird die Identität aus ihnen neu berechnet, und die Öffnung jedes Shards prüft seine Setup-Spalten gegen dieselben Punkte. Das verbindet die Tabellen, die ein Beweis liest, mit der Identität, die ein Verifier registriert hat.

Wichtig

Ein Verifier bezieht die Identität über einen Kanal, den der Prover nicht kontrolliert. Gegenüber einer vom Prover gelieferten Identität zeigt ein Beweis nur, dass irgendein Programm gelaufen ist. Wer das ELF, die Parameter und die Zeremoniedatei besitzt, kann sie neu berechnen.

Die Spezifikation: Programm und Identität.

Architektur

Ausführung, Familien und Shards

Ein Hart, eine 38-Bit-Uhr, jeder Speicherzugriff als Abfrage mit Zeitstempel, 23 Schaltkreisfamilien und der Shard als Einheit des Beweisens.

Als Markdown anzeigen

Die Maschine#

Der Emulator führt RV32IMAC auf einem Hart über einem ProgramImage aus, ohne Interrupts und ohne Privilegstufen. Ein Lauf ist eine reine Funktion des Images und seiner Eingabe, ohne Uhr, Zufall oder Threads; zwei Läufe schneiden also identische Shards. Von einer gehosteten RV32IMAC-Umgebung weicht er in drei Punkten ab: sc.w gelingt immer, ein nicht ausgerichteter Halbwort- oder Wortzugriff ist fatal, statt aufgeteilt zu werden, und der Befehlsstrom ist das beim Laden dekodierte Image.

Jede andere Art, wie ein Lauf vor EXIT enden kann, etwa ein Zugriff außerhalb der gemappten Bereiche, ebreak oder ein Sprung zu einem Halbwort ohne Befehl, ist ein fataler Fehler ohne Trace. Ein solcher Lauf hat keinen Beweis. Ein Exit-Status ungleich null ist kein Fehler: Es handelt sich um eine Ausführung wie jede andere, und sie ist beweisbar.

Die Uhr und die Abfrage#

Zyklus c belegt die vier Zeitstempel 4c + Δ, einen pro Slot Δ ∈ {0, 1, 2, 3}. Jeder Befehl ist ein Zyklus, und sonst ist nichts ein Zyklus: Ein Delegationsaufruf läuft im Zyklus mit, der ihn angefordert hat. Zyklen werden ab 1 gezählt, weil Zeitstempel 0 der anfängliche Schreibzugriff jeder Adresse ist. Die Uhr hat 38 Bit; eine Ausführung läuft also höchstens 2^36 − 1 Zyklen.

Jeder Speicherzugriff ist eine Abfrage: ein Lesezugriff auf einen Wert, der zuletzt zu einem früheren Zeitstempel geschrieben wurde, und ein Schreibzugriff zum aktuellen Zeitstempel. Eine Abfrage, die nur liest, schreibt zurück, was sie gelesen hat. Slot 0 jedes Zyklus ist die pc-Abfrage, die pc liest und next_pc schreibt. Danach folgen die Register- und Speicherabfragen des Befehls an festen Slots:

Klasse Δ = 1 Δ = 2 Δ = 3
Register-Immediate, jalr rs1 rd
Verzweigungen rs1 rs2
Register-Register, M rs1 rs2 rd
Loads rs1 das Wort, gelesen rd
Stores rs1 rs2 das Wort, mit eingesetzten gespeicherten Bytes
Atomics rs1 rs2 das Wort und rd
ecall a7 a0 a0 und die Spiegelabfrage einer Delegation

Adressen liegen in Adressräumen: die 32 Register, der RAM in Wörtern mit 4-Byte-Ausrichtung, der pc, ein Ankerraum pro Delegationstyp und die Körperzellen des Rekursionsformats. x0 ist im Trace ein gewöhnliches Register und in der Maschine eine Konstante: Jede Abfrage an diesem Register liest und schreibt 0.

Dreiundzwanzig Familien#

Eine Familie ist ein Schaltkreis zusammen mit den Zeilen, die er beweist. Es gibt vier Arten:

Art Familien Eine Zeile ist
Ausführung 0–6: ADD_SUB_LUI_AUIPC, JUMP_BRANCH_SLT, SHIFT_BITWISE, MUL_DIV, MEM_WORD, MEM_SUBWORD, ATOMICS ein ausgeführter Befehl
Fenster 7 INIT_TEARDOWN, 8 ZERO_WINDOWS, 12 PUBLIC_INPUT, 13 PUBLIC_OUTPUT, 14 ADVICE_WINDOWS ein Speicherwort, initialisiert und abgeschlossen
Delegation 9 KECCAK_F, 10 POSEIDON2, 11 FR_ARITH, 15 MOD_MUL, 16 SHA256_COMP, 17 EC_ADD ein Aufruf über einem Frame aus RAM
Rekursion 18 FIELD_WINDOWS, 19 FR_OP, 20 P2_FIELD, 21 FIELD_IO, 22 FQ_OP eine Körperzelle oder eine Koprozessor-Operation

Jeder Zyklus geht an die eine Ausführungsfamilie, deren dekodierte Tabelle seinen pc beansprucht. Familien wechseln sich zeitlich ab: ADD_SUB_LUI_AUIPC kann die Zyklen 1 und 3 besitzen und JUMP_BRANCH_SLT den Zyklus 2. Nichts verlangt, dass sie zusammenhängend sind, weil das Speicherargument jede Zeile über ihren pc-Schreibzugriff ordnet.

Die Fensterfamilien gibt es, weil das Speicherargument verlangt, dass jede Adresse, die eine Ausführung berührt, genau einen Anfangswert und einen Endwert hat. INIT_TEARDOWN deckt RAM-Fenster 0 ab und belegt es zu Beginn mit dem Image des Programms; ZERO_WINDOWS deckt jedes andere Fenster gewöhnlichen RAMs ab, das der Lauf berührt hat, und belegt es zu Beginn mit null; das öffentliche Paar deckt die Fenster für Eingabe und Journal ab; ADVICE_WINDOWS deckt den Bereich der Hilfsdaten (Advice) ab und belegt ihn zu Beginn mit den Bytes des Provers.

Shards#

Die Zeilen einer Familie werden in der Reihenfolge ihres Auftretens in Shards der Höhe dieser Familie geschnitten. Der letzte wird mit Nullzeilen aufgefüllt, die die Schaltkreise per Konstruktion akzeptieren. Ein Shard kostet seine volle Höhe, unabhängig von seiner Belegung; die Familien, die ein Programm berührt, und die Höhen, die es wählt, bestimmen also die Untergrenze jedes Beweises.

Familie Standardhöhe Warum
Befehlsfamilien 2^22, 2^20 für MUL_DIV und ATOMICS die Untergrenze ihrer Range-Checks für Zeitstempel ist 2^20
RAM-Fenster 2^22 eine gemeinsame Fensterhöhe, mindestens 2^16
Öffentliche Eingabe, Journal 2^12 fest: Die Höhe bestimmt die Lage der Fenster
KECCAK_F, SHA256_COMP 2^18 viermal so viele Aufrufe wie bei ihrer Untergrenze 2^16, bei einem um 2 % größeren Beweis
MOD_MUL, EC_ADD 2^16 ihre Untergrenze
POSEIDON2, FR_ARITH 2^8 keine Tabelle, also keine Untergrenze

Ein Shard wird für sich bewiesen, durch den Schaltkreis seiner Familie, mit Ausnahme des Speicherarguments: Der Schaltkreis jedes Shards gibt das Produkt seiner Lesetupel und das Produkt seiner Schreibtupel aus, und der Verifier gleicht diese Produkte einmal über alle Shards der Aussage ab. Das ist das Einzige, was Shards verbindet. Es gibt keine Verkettung des pc von Shard zu Shard und keinen gemeinsamen Rand zwischen Nachbarn.

Die Größen der Shard-Beweise bei den Standardhöhen reichen von etwa 12,5 KB für ein öffentliches Fenster bis zu 665 KB für einen POSEIDON2-Shard; der Beweis einer Befehlsfamilie liegt bei 64 bis 77 KB. Die Tabelle steht unter Performance.

Die Spezifikation: Ausführungs-Trace, Schaltkreise und Registry.

Architektur

Die GKR-Engine

Die Beweis-Engine im Zentrum von Apogee. Warum ein geschichteter GKR-Schaltkreis nur seine Eingaben committet, wie ein einziger Rückwärtsdurchlauf aus Sumchecks einen ganzen Schaltkreis auf einen einzigen Punkt zurückführt und was das einbringt.

Als Markdown anzeigen

Jeder Shard in Apogee wird auf dieselbe Weise bewiesen: durch den Schaltkreis seiner Familie, geschrieben als Stapel von Schichten, den die GKR-Engine rückwärts durchläuft, von den Ausgaben des Schaltkreises bis zu seinen committeten Spalten. Diese Seite erklärt, warum diese Engine im Zentrum des Systems steht und warum sich die Front des Beweisens in Richtung der Familie von Beweissystemen bewegt, zu der sie gehört.

Die Idee#

Der klassische Weg, eine Berechnung zu beweisen, besteht darin, sie als Tabelle anzulegen, jede Spalte einschließlich der Zwischenwerte zu committen und zu beweisen, dass eine Menge von Constraints über der Tabelle verschwindet. Die Commitments sind der teure Teil: Jede committete Spalte kostet eine Multi-Skalar-Multiplikation oder einen Merkle-Baum und eine Öffnung an jedem Punkt, nach dem das Constraint-Argument fragt.

GKR, nach Goldwasser, Kalai und Rothblum, ändert, was committet werden muss. Die Berechnung ist ein geschichteter Schaltkreis. Nur die unterste Schicht, die Eingaben, wird committet. Jede Schicht darüber ist durch Gates über der Schicht darunter definiert, und der Prover committet sie nie. Stattdessen wird eine Behauptung über die oberste Schicht durch einen Sumcheck auf eine Behauptung über die Schicht darunter zurückgeführt, dann auf die nächste, bis die Behauptungen bei den committeten Eingaben ankommen, alle an einem einzigen zufälligen Punkt. Eine einzige Öffnung löst sie ein.

Hinweis

Eine Analogie, kein Name. Ein herkömmlicher Prover ist eine Rakete: Er schleppt jeden Zwischenwert, den er erzeugt, committet bis ans Ziel und bezahlt für die Masse. Die GKR-Engine verhält sich eher wie der Warp-Antrieb der Science-Fiction, der den Raum um das Schiff bewegt statt des Schiffs selbst. Was reist, ist die Behauptung, Schicht für Schicht durch den Schaltkreis nach unten bewegt, während die Zwischenschichten überhaupt nirgendwohin getragen werden.

Was das einer zkVM einbringt:

  • Zwischenwerte kosten kein Commitment. Ein Familienschaltkreis kann Hunderte innerer Spalten, Produktbäume und Bruchbäume berechnen, und nichts davon wird je committet. Nur seine Trace-Spalten werden committet.
  • Ein Öffnungspunkt pro Shard. Der Rückwärtsdurchlauf endet mit Behauptungen über jede committete Spalte an demselben Punkt. Ein Shard braucht genau eine gebündelte Öffnung, 704 Byte, unabhängig davon, wie viele Spalten er hat.
  • Die Arbeit des Provers ist Körperarithmetik. Der Sumcheck jeder Schicht ist linear in der Größe der Schicht, über Fr, ohne Commitment, Transformation oder Hash pro Schicht.
  • Argumente setzen sich innerhalb des Schaltkreises zusammen. Die Grand Products des Speicherarguments und die LogUp-Summen der Lookups sind einfach weitere Schichten desselben Schaltkreises, reduziert im selben Durchlauf.

Ein Familienschaltkreis, Schicht für Schicht#

Die Schichten eines Familienschaltkreises Schicht 0 besteht aus den committeten Spalten neben virtuellen Tabellen. Gate-Liste 0 baut Speicherblätter, Lookup-Brüche und Constraint-Gates. Zeilenweise Listen reduzieren die Bäume jeder Zeile auf einen Knoten. Halbierende Listen, eine pro Variable, multiplizieren und addieren die Zeilen bis zu einer Spitze ohne Variablen: der Lesewurzel, der Schreibwurzel sowie Zähler und Nenner jedes Kanals. Der Rückwärtsdurchlauf läuft von der Spitze bis zu Schicht 0, mit einem Sumcheck pro Übergang, und endet an einem einzigen Punkt, an dem eine einzige Mercury-Öffnung jede committete Spalte einlöst. Schicht 0 · committete Spalten M ‖ W ‖ S neben virtuellen Tabellen V, die der Verifier selbst auswertet · einzige committete Schicht Gate-Liste 0 · zeilenweise Speicherblätter · Lookup-Brüche (num, den) · jedes Constraint-Gate, Grad ≤ 2 Zeilenweise Listen 1 … r Produkt- und Bruchbäume kombinieren Geschwister, bis jede Zeile einen Knoten pro Baum hält Halbierende Listen, eine pro Variable TreeProduct · TreeCross: Zeilen (·,0) und (·,1) kombiniert Spitze · keine Variablen Lesewurzel · Schreibwurzel · (num, den) pro Kanal RÜCKWÄRTSDURCHLAUF Ausgaben zuerst absorbiert ein Sumcheck pro Übergang: Behauptungen zu Schicht k+1 → Behauptungen zu Schicht k Halbierung: Geradenpunkt τ vereint die beiden Kinder jede committete Spalte behauptet an einem Punkt u → EINE MERCURY-ÖFFNUNG VORWÄRTSDURCHLAUF · NUR PROVER
Ein Familienschaltkreis. Der Prover berechnet jede Schicht einmal von unten nach oben (gestrichelt). Der Beweis läuft nach unten: Die Ausgaben werden absorbiert, dann ist jeder Übergang ein Sumcheck, der Behauptungen über eine Schicht in Behauptungen über die Schicht darunter verwandelt, bis sich alle an einem einzigen Punkt auf den committeten Spalten treffen.

Die unterste Schicht bilden die committeten Spalten des Shards, in drei Arten, die sich darin unterscheiden, wann sie gebunden werden: M, die Speicherspalten, committet im globalen Transkript der Aussage, bevor es irgendeine Speicher-Challenge gibt; W, die Witness-Spalten, committet im eigenen Transkript des Shards; und S, die Setup-Spalten, gebunden durch die Programmidentität oder durch die Zeremonie. Daneben stehen virtuelle Tabellen: geschlossene Formen wie der Zeilenindex oder der 16-Bit-Bereich, die der Verifier an jedem beliebigen Punkt auswertet und die nie committet werden.

Oberhalb von Schicht 0 hat jeder Familienschaltkreis denselben Aufbau:

  1. Gate-Liste 0 berechnet Zeile für Zeile die Speicherblätter (die Lese- und Schreibtupel jeder Abfrage), die Lookup-Brüche (ein Paar (numerator, denominator) pro Lookup, dazu das der Tabelle) und jedes Constraint-Gate: die Constraints der Familie, jedes ein Polynom, das auf jeder Zeile verschwinden muss.
  2. Zeilenweise Listen kombinieren Geschwisterblätter: Ein Produktbaum multipliziert Tupel, ein Bruchbaum addiert Brüche als (n_a·d_b + n_b·d_a, d_a·d_b), bis jede Zeile einen Knoten pro Baum enthält.
  3. Halbierende Listen, eine pro Variable der Höhe des Shards, kombinieren die Zeilen paarweise: die erste Hälfte mit der zweiten. Nach n von ihnen erreicht der Schaltkreis eine Spitze ohne Variablen: die Lesewurzel des Shards, seine Schreibwurzel sowie den finalen Zähler und Nenner jedes Lookup-Kanals.

Ein einziger Schaltkreis beweist also die Constraints der Familie, berechnet ihren Beitrag zum Speicherargument und summiert ihre Lookups, in einem einzigen Durchlauf. Jedes Gate hat höchstens Grad 2; jedes Rundenpolynom jedes Sumchecks ist also kubisch.

Der Rückwärtsdurchlauf#

Der Prover materialisiert jede Schicht einmal, von unten nach oben. Der Beweis läuft dann nach unten, und sein Transkriptablauf ist für jeden Schaltkreis derselbe:

  1. Die Ausgaben werden zuerst absorbiert, sodass der Prover an die Wurzeln gebunden ist, bevor es irgendeine Challenge gibt.
  2. Für jeden Übergang von Schicht k + 1 hinunter zu Schicht k bündelt eine Challenge λ jede Behauptung über Schicht k + 1 zusammen mit jedem Constraint-Gate der Liste zu einer einzigen Summe. Ein Sumcheck reduziert diese Summe auf eine Auswertung an einem zufälligen Punkt ρ, mit einer kubischen Nachricht pro Variable.
  3. Der Prover nennt die Werte der Spalten von Schicht k an der Stelle ρ. Bei einer halbierenden Liste nennt er beide Kinder, und eine weitere Challenge τ führt sie zu einer Behauptung pro Spalte zusammen.
  4. Auf Schicht 0 hat jede committete Spalte eine einzige Behauptung, alle am selben Punkt u.

Die eine Mercury-Öffnung des Shards beweist diese Behauptungen gegen die Commitments: die der Speicherspalten aus der Aussage, die der Witness-Spalten aus dem Shard-Beweis, die der Setup-Spalten aus dem Verifikationsschlüssel. Virtuelle Tabellen wertet der Verifier selbst aus.

Constraint-Gates fahren kostenlos mit. Ein Constraint-Gate behauptet überall 0; es tritt also dem Batch des Übergangs bei, in dem es liegt, und ein verletztes Gate macht die gebündelte Summe mit überwältigender Wahrscheinlichkeit ungleich null. Ein einziger Fehler LayerInconsistency deckt eine falsche absteigende Behauptung und ein verletztes Gate gleichermaßen ab; eine gebündelte Summe kann die beiden nicht unterscheiden, und der Beweis wendet nichts dafür auf, sie zu unterscheiden.

Woher die Soundness kommt#

Jede Challenge wird nach allem gezogen, was sie schützt:

  • der Ausgabepunkt nach den Ausgaben, sodass ein Prover keine Tabellen wählen kann, die nur dort mit der Wahrheit übereinstimmen, wo geprüft wird;
  • λ nach den Behauptungen und dem Punkt, sodass eine falsche Behauptung oder ein verletztes Gate nur an einer Nullstelle eines von null verschiedenen Polynoms in λ überlebt;
  • jede Sumcheck-Challenge nach dem kubischen Polynom ihrer Runde, sodass ein falsches kubisches Polynom mit Wahrscheinlichkeit höchstens 3/|Fr| mit dem richtigen übereinstimmt;
  • τ nach den Werten beider Kinder.

Summiert über alle Übergänge eines registrierten Schaltkreises bei seiner Standardhöhe bleibt der Soundness-Fehler unter 2^14/|Fr|, mit Fiat–Shamir über dem Poseidon2-Transkript im Random-Oracle-Modell.

Schaltkreise als Daten#

Ein Schaltkreis ist kein Code. Er ist ein CircuitArtifact: seine committeten Spalten mit Namen, seine virtuellen Tabellen, seine Gate-Listen mit sieben Gate-Formen, eine flache Liste derselben Relationen, seine Lookups und seine Padding-Zeile, kanonisch serialisiert. Vier Gesetze halten jedes Artefakt in einer konsistenten Form: Jeder Operand ist dort lesbar, wo er gelesen wird, die Breite jeder Liste wird aus ihren Gates abgeleitet, die oberste Schicht besteht genau aus den Ausgaben, und die geschichteten Gates und die flachen Relationen bilden eine einzige Constraint-Menge. Sie laufen einmal, dort, wo ein Artefakt gebaut oder ein Schlüssel geladen wird, nie pro Beweis.

Zwei Folgen sind für alle wichtig, die das System bewerten:

  • Ein Verifikationsschlüssel führt seine Schaltkreise mit, und der Verifier gleicht sie mit seiner eigenen Registry ab. Die Programmidentität bindet das Programm; die Registry bindet die Schaltkreise, die es beweisen.
  • Die Schaltkreise lassen sich von einer zweiten Implementierung prüfen. Das Crate checker implementiert die Gesetze, die Lookup-Regeln und den Padding-Vertrag neu, ohne Code mit den Konstruktoren zu teilen, und wertet Gates ausschließlich über den einen Gate-Kernel aus, den beide Seiten als semantische Autorität verwenden.

Die Kosten#

Die Engine tauscht Commitments gegen Speicher. Der Vorwärtsdurchlauf hält jede innere Schicht als Körperelemente: etwa 8,4 GiB für einen 2^20-Shard der breitesten Befehlsfamilie und 42 GiB für einen KECCAK_F-Shard der Höhe 2^18, dessen Schaltkreis 5.490 innere Spalten berechnet. Deshalb ist die Beweiserzeugung speichergebunden, deshalb sind Höhen ein Tuning-Parameter, und deshalb beschränkt der Streaming-Prover den Speicherbedarf durch die gleichzeitig bearbeiteten Shards. Die Beweisgröße wächst nur um eine Sumcheck-Runde pro Variable und Schicht: Der Beweis eines KECCAK_F-Shards umfasst 381.100 Byte bei 2^18 gegenüber 373.276 bei 2^16, für die vierfache Arbeit.

Die Spezifikation: Die GKR-Engine sowie die eigene Seite jeder Familie unter Auditoren.

Architektur

Polynom-Commitments

Jede committete Spalte wird mit Mercury geöffnet, einem multilinearen Commitment über KZG mit einer Öffnung konstanter Größe. Was es kostet, wie ein Shard alle seine Spalten in einer einzigen Öffnung bündelt und wie die Rekursion das Pairing aufschiebt.

Als Markdown anzeigen

Jede Spalte, die Apogee committet, ist ein multilineares Polynom, eine Tabelle von 2^n Auswertungen über dem booleschen Hyperwürfel. Jede davon wird mit Mercury (Eagen und Gabizon, ePrint 2025/385) committet und geöffnet, abgeschlossen durch die gebündelte KZG-Öffnung von BDFG20 (Boneh, Drake, Fisch und Gabizon, ePrint 2020/081). Die Spezifikation legt fest, was die Arbeiten offenlassen, und fügt zwei Dinge hinzu: einen Batch vieler Spalten an einem Punkt und eine aufgeschobene Form, die der Rekursionsbaum faltet.

Das Commitment#

Ein Mercury-Commitment ist genau das KZG-Commitment der Auswertungstabelle, gelesen als Koeffizienten: eine Multi-Skalar-Multiplikation über die ersten n Potenzen des τ der Zeremonie, ein G1-Punkt, 64 Byte. Dahinter steht kein zweites Verfahren. Daraus folgen zwei Eigenschaften, und der Rest des Systems nutzt beide:

  • Es ist linear. Das Commitment von Σ ρ^i·f_i ist Σ ρ^i·cm_i; deshalb können sich viele Spalten eine Öffnung teilen.
  • Nullkoeffizienten tragen nichts bei. Eine um Nullzeilen erweiterte Spalte behält ihr Commitment; die drei Commitments der generischen Lookup-Tabelle dienen also jeder Höhe, die die Tabelle aufnimmt, und sie sind eine Konstante der Zeremonie.

Spalten werden mit ihrer Ganzzahlbreite committet: Eine Bit-, Byte-, Halbwort- oder Wortspalte läuft durch eine MSM über u32-Skalare, umkodiert aus 32 statt 254 Bit, was das Committen eines Trace günstig hält.

Die Öffnung#

Mercury teilt einen Öffnungspunkt u mit s = 2t Variablen in zwei Hälften, faltet das Polynom mit einer Challenge α und beweist die beiden verbleibenden inneren Produkte mit einem symmetrisierten Witness, abgeschlossen durch eine einzige gebündelte KZG-Öffnung an drei Punkten. Der Beweis besteht aus acht G1-Punkten und sechs Körperelementen: 704 Byte, für jede Größe. Daraus folgt eine Anforderung: Die Zahl der Variablen muss gerade sein, und deshalb ist jede Höhe in der Auswahl eine Zweierpotenz mit geradem Exponenten.

Der Verifier führt O(t) Körperoperationen, MSMs über zehn und über zwei Punkte und eine Pairing-Prüfung mit zwei Paaren aus. Beide Relationen, die er prüft, sind als e(A, [1]_2) = e(B, [x]_2) geschrieben; beide G2-Argumente sind also Konstanten des Setups, und der Verifier führt überhaupt keine G2-Arithmetik aus. Diese Form erlaubt es der Rekursion auch, das Pairing aufzuschieben, statt es zu berechnen.

Die Spalten eines Shards, eine Öffnung#

Der GKR-Durchlauf endet mit Behauptungen über jede committete Spalte eines Shards am selben Punkt u. Es braucht also keinen Sumcheck, der Behauptungen zusammenführt: Die Öffnung bündelt sie alle. Eine Challenge ρ, gezogen nach jedem Commitment und jedem behaupteten Wert, gewichtet Spalte i mit ρ^i; der Prover öffnet Σ ρ^i·f_i einmal, und der Verifier bildet Σ ρ^i·cm_i durch eine MSM über k Punkte. Eine falsche Behauptung überlebt mit Wahrscheinlichkeit höchstens (k − 1)/|Fr|.

Die Spalten im Batch stammen aus drei Quellen, in einer festen Reihenfolge, die Teil dessen ist, was bewiesen wird: die Commitments der Speicherspalten aus der Aussage, die der Witness-Spalten aus dem Shard-Beweis und die der Setup-Spalten aus dem Verifikationsschlüssel. Weil die Setup-Commitments dem Schlüssel entnommen werden, bindet die Öffnung die dekodierten Tabellen und das Image, die die Programmidentität committet.

Aufgeschobene Verifikation#

Eine aufgeschobene Prüfung führt alles außer dem Pairing aus und behält die Terme der Relation als zwölf Einträge (side, scalar, point). Genau das nutzt die Rekursion: Jeder Knoten des Baums berechnet die zwölf Skalare eines Shards in seiner eigenen Arithmetik, gewichtet sie mit frischen Challenges und addiert sie in ein einziges laufendes Punktepaar (A, B). Die Öffnung jedes Shards im gesamten Baum wird in eine einzige Behauptung e(A, [1]_2) = e(B, [x]_2) gefaltet, die erst der Contract auf Ethereum endgültig prüft. Rekursion und Abwicklung zeigt das Falten.

Gemessen#

Auf einem Apple M5 Pro mit 18 Kernen:

Operation Zeit
Eine 2^22-Spalte committen 1,30 s
Eine 2^22-Spalte öffnen 2,89 s
16 Spalten der Größe 2^20, als ein Batch geöffnet 1,01 s, verifiziert in 4,8 ms
Dieselben 16 einzeln geöffnet 9,79 s, verifiziert in 62 ms

Sicherheit#

Knowledge Soundness gilt im algebraischen Gruppenmodell unter q-DLOG, mit Fiat–Shamir über dem Poseidon2-Transkript im Random-Oracle-Modell und einem SRS, dessen τ niemand kennt. Die statistischen Terme, Schwartz–Zippel über die Challenges, bleiben für jede verwendete Instanz unter 2^−220; das Sicherheitsniveau ist also das von BN254, etwa 100 Bit. Nichts ist hiding, und nichts wird verblindet: Apogee v1.0.0 ist nicht zero-knowledge.

Als SRS dienen die Perpetual Powers of Tau der PSE, Beitrag 80, deren Soundness gegeben ist, solange ein Beitragender ehrlich war. Ihre Datei wird dekodiert, und für jeden Punkt wird geprüft, dass er auf seiner Kurve und in der richtigen Untergruppe liegt; nichts beweist, dass eine Datei die dieser Zeremonie ist, und deshalb vergleicht ein Verifier den SRS-Digest mit dem der Zeremonie aus seinem eigenen Kanal.

Die Spezifikation: Mercury, Der strukturierte Referenzstring.

Architektur

Speicher und Lookups

Zwei Argumente tragen alles, was über eine Zeile hinausreicht. Eine einzige Lese-/Schreib-Multimenge über die gesamte Ausführung, einmal abgeglichen, sorgt dafür, dass jeder Lesezugriff den letzten Schreibzugriff liefert, und ordnet jede Zeile; LogUp-Kanäle machen jeden Wert zu einem Byte, einem Wort oder einer Tabellenzeile.

Als Markdown anzeigen

Die Gates einer Familie stellen Bedingungen an jeweils eine Zeile. Alles, was über Zeilen, Shards oder Familien hinausreicht, etwa was ein Register enthält, was ein Load liefert, welchen Befehl eine Zeile ausführt und ob ein Wert in 32 Bit passt, tragen zwei Argumente, die in denselben GKR-Schaltkreisen leben.

Das Speicherargument#

Jeder Speicherzugriff wird zu einem Körperelement, einem Tupel:

T(AS, ADDR, TS, VAL) = γ + AS + α_addr·ADDR + α_ts·TS + α_val·VAL

über dem Adressraum, der Adresse, dem Zeitstempel und dem Wert, mit vier Challenges, die einmal pro Aussage gezogen werden. Eine Abfrage trägt ihr Lesetupel zu einer Seite bei und ihr Schreibtupel zur anderen. Der Schaltkreis jedes Shards gibt zwei Zahlen aus, das Produkt seiner Lesetupel und das Produkt seiner Schreibtupel. Der Verifier prüft dann eine einzige Gleichung über die gesamte Aussage:

∏ read roots · R_b  =  ∏ write roots · W_b          over every shard of every family

W_b und R_b sind die Anfangs- und Endtupel der 32 Register und des pc, die keine eigenen Zeilen haben: Der Verifier bildet sie aus 64 Randskalaren, die die Aussage mitführt, und multipliziert sie ein. Der RAM erhält seine Anfangs- und Endwerte aus den Shards der Fensterfamilien, die ein Wort pro Zeile enthalten: das Image des Programms in Fenster 0, Nullen in jedem anderen Fenster, das der Lauf berührt hat, die Eingabe der Aussage im öffentlichen Eingabefenster, die Bytes des Provers in den Hilfsdaten (Advice).

Gilt die Gleichung, sind die Multimengen mit überwältigender Wahrscheinlichkeit gleich. Gleiche Multimengen bedeuten, dass jeder Lesezugriff den letzten Schreibzugriff davor liefert: Jeder Lesezugriff wird genau einem Schreibzugriff zugeordnet, ein Lesezugriff muss strikt auf den Schreibzugriff folgen, den er konsumiert, und jede Adresse hat genau einen anfänglichen Schreibzugriff.

Reihenfolge ohne Zusatzkosten#

Der pc ist eine Speicherzelle wie jede andere, an Adresse 0 seines eigenen Adressraums. Jede Zeile liest den pc und schreibt den nächsten, und ihr Schaltkreis erzwingt, dass der Schreibzugriff mindestens vier Zeitstempel nach dem Lesezugriff liegt. Die Historie des pc ist also ein einziger Pfad durch jede aktive Zeile jeder Familie, vom Einsprungpunkt bis zur Exit-Zeile. Dieser eine Pfad liefert:

  • Programmreihenfolge, da Zeilen nach ihren pc-Schreibzugriffen geordnet sind;
  • Kontinuität über Shards und Familien hinweg, da der pc-Lesezugriff jeder Zeile den pc-Schreibzugriff irgendeiner Zeile konsumiert;
  • keinen doppelt bewiesenen Zyklus, da kein Schreibzugriff zweimal konsumiert werden kann.

Kein Shard ist mit seinem Nachbarn verkettet, und keiner muss es sein. Das behauptete Zeitfenster eines Shards bindet nichts; es wird nur auf seine Form geprüft.

Was zuerst feststehen muss#

Die Speicher-Challenges werden einmal pro Aussage gezogen, am Ende des globalen Transkripts, nachdem alles feststeht, was ein Tupel lesen kann: die Speicher-Commitments jedes Shards, die Programmidentität (die den Einsprung-pc und das Image festlegt), der Digest der Zeremonie, die Shard-Zahlen und die Fensterliste, der Digest von öffentlicher Eingabe und Journal und zuletzt die 64 Randskalare. Ein nach den Challenges gewählter Wert ließe sich passend errechnen; die Reihenfolge des Transkripts verhindert das. Aus demselben Grund darf ein Speichertupel nur Speicher-, Setup- und virtuelle Spalten lesen, nie eine Witness-Spalte, die im eigenen Transkript des Shards nach den Challenges committet wird. Die Schaltkreiskonstruktoren weisen jedes Artefakt zurück, das dagegen verstößt.

Lookups#

Ein Lookup besagt, dass ein Tupel aus Werten einer Zeile eine Zeile irgendeiner Tabelle ist. Apogee beweist jeden Lookup eines Shards mit LogUp: eine Identität pro Tabelle oder Kanal,

Σ_rows Σ_lookups 1/(E(y) + g)  −  Σ_rows mult(y)/(T(y) + g)  =  0

summiert über einen Bruchbaum im eigenen GKR-Schaltkreis der Familie und an dessen Wurzel geprüft: Zähler null, Nenner ungleich null. Die Multiplizitätsspalte braucht überhaupt keinen Constraint: Ein Tupel, das in keiner Tabellenzeile vorkommt, hinterlässt eine Polstelle, die die Multiplizitäten nicht aufheben können.

Kanal Tabelle Verwendet für
TIMESTAMP [0, 2^19), virtuell den Zeitstempelabstand jeder Abfrage, als zwei Chunks zu je 19 Bit
RANGE16 [0, 2^16), virtuell 32-Bit-Werte als zwei Halbwörter; Überträge; Frame-Schranken
XOR8 alle Bytepaare und ihr XOR, virtuell Keccak und SHA-256, Byte für Byte
GENERIC eine committete Tabelle aus Zeilen für AND, Vorzeichen und Shift-Potenzen bitweise Operationen, Vorzeichenbits, Shift-Beträge
DECODER die dekodierte Tabelle der Familie, committet durch die Identität Bindung jeder ausgeführten Zeile an das Programm

Drei der Tabellen sind virtuell: geschlossene Formen des Zeilenindex, die der Verifier selbst auswertet und die kein Commitment kosten. Die generische Tabelle wird einmal mit den Potenzen der Zeremonie committet und ist durch den SRS-Digest abgedeckt.

Der Decoder-Lookup#

Jede Ausführungsfamilie macht pro aktiver Zeile einen Lookup in ihre eigene dekodierte Tabelle, mit dem pc als Schlüssel, den die Zeile aus dem Speicher gelesen hat. Dieser eine Lookup bindet den Zyklus an das Programm: Operanden, Immediate und Befehlsart der Zeile sind die des Programms an diesem pc, und ihre Bits für die Befehlsart sind one-hot, weil jede aktive Zeile der Tabelle eine One-hot-Maske enthält und jede Padding-Zeile −1, was keine Summe solcher Bits erreicht. Eine Zeile an einem pc, an dem das Programm keinen Befehl hat, findet überhaupt keine Tabellenzeile.

Schlüssel müssen beschränkt sein#

Ein Kanal beweist die Zugehörigkeit zu einer Tabelle und nichts weiter. Mehrere Teiltabellen teilen sich die generische Tabelle unter disjunkten Schlüsselbereichen; ein unbeschränkter Schlüssel könnte also in der falschen Teiltabelle landen und ein falsches AND beweisen. Jede Familie beschränkt deshalb jeden Schlüssel, den sie nachschlägt, mit einem Bereichs-Lookup unter demselben Selektor, und die Schaltkreiskonstruktoren prüfen, dass eine über einen Skalierungsfaktor geschriebene Schranke zusätzlich eine direkte Schranke trägt. Die Spezifikation nennt für jede Familie den Angriff, den das verhindert.

Wie sich alles zusammenfügt#

Zusammen mit den Gates jeder Familie geben diese beiden Argumente der Aussage ihre Bedeutung: Jede Zeile befolgt ihren Befehl, der Befehl ist der des Programms, jeder Lesezugriff sieht den letzten Schreibzugriff, die Zeilen bilden einen einzigen Pfad vom Einsprung bis zum Exit, jeder Wert ist die ganze Zahl, die er zu sein behauptet, und die öffentlichen Fenster enthalten die Bytes der Aussage. Die Soundness-Karte führt jede Behauptung zu den Abschnitten, die sie beweisen.

Die Spezifikation: Das Speicherargument, Lookups.

Architektur

Delegationen

Wie eine teure Funktion einen eigenen Schaltkreis erhält, ohne dass die Befehlsschaltkreise wachsen. Wie eine Delegation aufgerufen wird, der Anker, der jede Anfrage mit genau einem Aufruf paart, die sechs Schaltkreise und ihre Wirtschaftlichkeit.

Als Markdown anzeigen

Hashing und Arithmetik mit großen Ganzzahlen dominieren reale Arbeitslasten: In einem Ethereum-Mainnet-Block machten allein die Körpermultiplikation und -quadrierung von secp256k1 44 % der Zyklen aus, bevor sie delegiert wurden. Sie Befehl für Befehl zu beweisen ist möglich und langsam. Eine Delegation gibt einer solchen Funktion eine eigene Schaltkreisfamilie, die aus dem Gastprogramm (Guest) aufgerufen wird; die Befehlsschaltkreise bleiben so klein, und ein Programm bezahlt nur für die Delegationen, die es aufruft.

Der Aufruf#

Eine Delegation wird aufgerufen, nie dekodiert. Das Gastprogramm schreibt einen Frame aus 32-Bit-Wörtern in den RAM und setzt einen ecall ab, mit der Nummer der Delegation in a7 und der Basisadresse des Frames in a0. Der ecall ist eine Zeile der Familie ADD_SUB_LUI_AUIPC, die Anfrage. Die Arbeit ist eine Zeile der eigenen Familie der Delegation, der Aufruf, der im anfragenden Zyklus jedes Frame-Wort liest und jedes Frame-Wort zurückschreibt, darunter die Ergebnisse. Ein Aufruf besitzt keinen Zyklus; er läuft in dem Zyklus mit, der ihn angefordert hat.

Frames sind wortausgerichtet und liegen vollständig im gewöhnlichen RAM; kein Frame überlappt also ein öffentliches Fenster oder die Hilfsdaten (Advice), und die Lese- und Schreibzugriffe des Frames sind gewöhnliche Speicherabfragen. Was eine Delegation berechnet hat, ist daher genau so gebunden wie jeder Store: über die eine Speicher-Multimenge.

Der Anker#

Anfragen und Aufrufe müssen sich eins zu eins paaren: Sonst könnten viele Anfragen gegen einen einzigen Aufruf aufgehen und Anfragen unausgeführt zurücklassen, oder ein nicht angefragter Aufruf könnte einen Frame überschreiben. Sie paaren sich über dieselbe Speicher-Multimenge, in einem Ankerraum, der allein dem Delegationstyp gehört und den kein Befehl erreichen kann:

Liest Schreibt
Anfrage, Zyklus c T(s, base, 0, 0) T(s, base, 4c + 3, v)
Aufruf T(s, base, 4c + 3, v′) T(s, base, 0, 0)

Drei Gates auf der Anfrageseite legen ihren Lesezugriff auf Zeitstempel 0 und Wert 0 fest und lassen sie 0 in a0 schreiben. Dann sind die mit 0 gestempelten Tupel genau die Lesezugriffe der Anfragen und die Antworten der Aufrufe; es gibt also ebenso viele Aufrufe wie Anfragen über denselben Basisadressen; und da keine zwei Anfragen einen Zyklus teilen, ist der Lesezugriff jedes Aufrufs genau der Schreibzugriff einer Anfrage. Jeder Aufruf sitzt an der Basisadresse und im Zyklus seiner Anfrage. Kein Gate in einem Delegationsschaltkreis musste überhaupt etwas von Anfragen wissen.

Viele Aufrufe, eine Operation#

Eine Operation, die für eine Zeile zu breit ist, besteht aus mehreren Aufrufen auf einem Frame, wobei ein Frame-Wort den Schritt benennt: Eine Permutation von keccak-f[1600] sind 24 Rundenaufrufe, eine SHA-256-Kompression 16 Aufrufe zu je vier Runden, eine vollständige Punktaddition drei Aufrufe. Kein Gate verbindet zwei Zeilen. Jeder Aufruf beweist seinen Schritt auf dem Frame so, wie er ihn vorfindet; seine Lesezugriffe liegen auf der einzigen Speicherhistorie jedes Worts, also liest er die Schreibzugriffe des vorherigen Schritts. Dass jeder Schritt läuft, und zwar in der richtigen Reihenfolge, muss der aufrufende Code sicherstellen, und der aufrufende Code ist Code des Gastprogramms, der als Befehle bewiesen wird. Das SDK setzt jede Operation aus mehreren Aufrufen aus einer einzigen Funktion ab; ein Gastprogramm ordnet die Schritte also nie von Hand.

Statisch deklariert#

Der Befehlsdurchlauf kann einen Aufruf nicht sehen, weil die Nummer ein Laufzeitwert von a7 ist. Deshalb hinterlässt jeder Shim im SDK einen 12-Byte-Deklarationsdatensatz in einer eigenen Linker-Sektion, der nur erhalten bleibt, wenn der Shim erreichbar ist. Die Ableitung des Programms durchsucht das Image nach solchen Datensätzen, und eine deklarierte Familie wird Teil der Konfiguration, gebunden durch die Identität über die Bytes des Images. Eine gelinkte, aber nie aufgerufene Familie beweist null Shards; eine aufgerufene Nummer, deren Familie das Programm nie deklariert hat, hat keinen Beweis.

Die sechs Schaltkreise#

Familie Ein Aufruf Aufgebaut aus
KECCAK_F eine Runde von keccak-f[1600] über einem Frame aus 51 Wörtern Bytes: 1.020 XOR8-Lookups pro Runde; Rotationen als Linearformen über Bytes und maskierten Kopien
SHA256_COMP vier Runden und vier Schedule-Wörter Bytes und XOR8: 52 Verpflichtungen pro Runde, 32 pro Schedule-Wort; Ch und Maj als Linearformen in XORs
POSEIDON2 eine Permutation der Breite 3 die Runden in den eigenen Schichten des Schaltkreises berechnet, drei Gate-Listen pro Runde, ohne Lookup; die einzige Delegation, die oberhalb ihrer ersten Schicht rechnet
FR_ARITH eine Fr-Addition, -Multiplikation oder -Inversion in Montgomery-Form Bitzerlegungen und Kanonizitätsketten gegen p
MOD_MUL ein 256-Bit-a·b mod m, vier Ethereum-Moduln 32-Bit-Limbs, ein Quotient, Überträge und eine Kanonizitätskette, die out < m beweist
EC_ADD ein Drittel einer vollständigen Punktaddition auf secp256k1 oder BN254 G1 die vollständige Formel von Renes–Costello–Batina als drei Reduktionen pro Zeile

Einige Konstruktionen kehren in ihnen wieder. Eine Ein-Code-Regel dekodiert ein Frame-Wort, das einen von k Fällen benennt, in boolesche Selektoren, von denen genau einer gesetzt ist, weil sich Codes addieren: Ohne sie beantworten die Selektoren 1 und 3 eine Anfrage nach 4. Eine Kanonizitätskette beweist mittels Borrows über 32-Bit-Limbs, dass ein 256-Bit-Wert unter einem Modul liegt. Und jedes geschriebene Wort ist unter 2^32 beschränkt, damit der RAM aus Wörtern besteht, worauf sich jede Befehlsfamilie verlässt.

Die Wirtschaftlichkeit#

Die Höhe einer Delegationsfamilie bestimmt, wie viele Aufrufe ein Shard aufnimmt, und ein Shard kostet seine Höhe, unabhängig von seiner Belegung:

Familie Höhe Einheiten pro Shard Shard-Beweis
KECCAK_F 2^18 10.922 Permutationen 381.100 B
SHA256_COMP 2^18 16.384 Kompressionen 189.988 B
EC_ADD 2^16 21.845 Additionen 434.916 B
MOD_MUL 2^16 65.536 Multiplikationen 135.220 B
POSEIDON2 2^8 256 Permutationen 664.780 B
FR_ARITH 2^8 256 Operationen 266.292 B

Für eine Familie mit vielen Aufrufen ist der größere Shard der günstigere: Ein KECCAK_F-Beweis wächst von 2^16 auf 2^18 kaum. Der Preis ist Speicher. Der Vorwärtsdurchlauf eines KECCAK_F-Shards der Höhe 2^18 umfasst 42 GiB, und zwei davon, gleichzeitig in Bearbeitung, bestimmten die Spitze des gemessenen Blocks.

Was delegiert wird und was nicht#

Bibliothekscode erreicht die Delegationen über gepatchte Kopien von k256, ark-ff und revm-precompile: Die secp256k1-Recovery wird zu k256-Code über MOD_MUL und EC_ADD, und ein BN254-Pairing wird zu ark-bn254-Code über MOD_MUL. MULMOD der EVM mit beliebigem Modul, MODEXP, BLS12-381 und jedes Signaturverfahren als Ganzes laufen als Befehle. Eigene Signaturunterstützung für Gastprogramme ist Teil der Flugbahn zu v2.0.0.

Die Spezifikation: Delegations-ABI, Delegationsschaltkreise.

Architektur

Der Streaming-Prover

Der Prover führt das Gastprogramm zweimal aus und hält nie den Ausführungs-Trace. Der Speicherbedarf folgt den gerade bearbeiteten Shards, nicht der Länge des Laufs, und der Beweis hängt nicht vom Scheduling ab.

Als Markdown anzeigen

Ein vollständiger Ethereum-Block umfasst etwa 200 Millionen Zyklen. Sein Trace, jede Zeile jeder Familie und jedes Speicherereignis, wäre etwa 300 Byte pro Zyklus groß: Dutzende Gigabyte, bevor eine einzige Spalte committet ist. Der Prover von Apogee baut ihn nie. Er führt das Gastprogramm (Guest) zweimal aus und hält nur die Shards, an denen er gerade arbeitet.

Zwei Durchläufe#

Die zwei Durchläufe Durchlauf 1 führt das Gastprogramm aus, committet Shard für Shard, sobald jeder gefüllt ist, dessen Speicherspalten und behält nur die Commitments; beim Exit baut er die Aussage und zieht die gemeinsamen Challenges. Durchlauf 2 führt erneut aus, beweist jeden Shard, sobald er gefüllt ist, und behält nur den Beweis. DURCHLAUF 1 · COMMIT Ausführen Shard gefülltM committen · Zeilen löschen Beim ExitFenster · Aussage · G1–G11 Challenges festgelegtSpeicher-Challenges · Digest DURCHLAUF 2 · BEWEIS Erneut ausführen Shard gefülltjede Spalte · GKR · Öffnung Beweis behaltenalles andere löschen BlockProofWurzeln in die Aussage jedes Shard-Transkript startet mit dem globalen Digest
Erst committen, dann beweisen. Die Speicher-Challenges müssen auf die Speicher-Commitments jedes Shards folgen; deshalb kommen die Commitments zuerst, aus einer Ausführung, und die Beweise aus einer zweiten.

Durchlauf 1 führt das Gastprogramm aus und committet, sobald ein Shard gefüllt ist, dessen Speicherspalten, behält die Commitments und verwirft die Zeilen. Beim Exit leitet er alles Weitere, was die Aussage braucht, aus dem finalen Speicherzustand ab: die Randwerte von Registern und pc, die Liste der Speicherfenster, die der Lauf berührt hat, und die Shards der Fensterfamilien. Dann führt er das globale Transkript aus, das die Aussage einschließlich jedes Speicher-Commitments absorbiert, und zieht die Speicher-Challenges sowie den Digest, mit dem jeder Shard initialisiert wird.

Durchlauf 2 führt erneut aus. Der Emulator ist eine reine Funktion seiner Eingabe und schneidet daher dieselben Shards, und Durchlauf 2 prüft per Assertion, dass sein Zyklenprofil, seine Fensterliste und seine Randwerte die von Durchlauf 1 sind. Jeder Shard erhält jede committete Spalte und wird bewiesen: sein Transkript, sein GKR-Durchlauf, seine Öffnung. Die Speicherspalten werden nicht erneut committet: Die Öffnung nimmt ihre Commitments aus der Aussage und ihre Werte aus Durchlauf 2; Spalten, die sich zwischen den Durchläufen unterschieden, ergäben also eine Öffnung, die der Verifier zurückweist.

Die Reihenfolge ist durch die Soundness erzwungen. Die Speicher-Challenges müssen auf jeden Wert folgen, den ein Speichertupel lesen kann; deshalb werden die Speicherspalten jedes Shards committet, bevor irgendein Shard bewiesen werden kann.

Die Pipeline#

Eine feste Zahl von Workern, max_in_flight, teilt sich ein Lock um den Executor. Unter dem Lock gibt ein Worker seinen fertigen Shard zurück und beansprucht den nächsten: einen gefüllten Shard, falls einer wartet, und andernfalls lässt er den Executor selbst Schritt für Schritt laufen, bis ein Puffer voll ist. Außerhalb des Locks baut er die Spalten des Shards, beweist ihn und verwirft ihn.

  • Der Executor läuft dem Bedarf nie voraus. Höchstens ein gefüllter, nicht beanspruchter Shard pro Familie wartet, in Form von Zeilen.
  • Innerhalb eines Shards ist die Arbeit datenparallel, über alle Kerne. Ein Worker, der in dieser Arbeit blockiert ist, nimmt keinen zweiten Shard.
  • Der Block hängt nicht vom Scheduling ab. Der Beweis eines Shards ist eine Funktion des globalen Zustands und seiner eigenen Spalten; Beweise werden nach ihrer Position in der Aussage abgelegt. Die Bytes sind bei 1 und bei 8 gleichzeitig bearbeiteten Shards identisch.
  • Fehlschläge sind deterministisch. Zurückgegeben wird der früheste Fehlschlag in der Befüllungsreihenfolge, bei jeder Zahl von Workern.

Was es kostet#

Der Speicher besteht aus einem Teilpuffer pro Familie, den Tabellen der letzten Zugriffe, den gerade bearbeiteten Shards und der Ausgabe. Den Speicherbedarf eines Shards bestimmt vor allem sein Vorwärtsdurchlauf, jede innere GKR-Schicht als Körperelemente: 8,4 GiB für einen SHIFT_BITWISE-Shard der Höhe 2^20, 42 GiB für einen KECCAK_F-Shard der Höhe 2^18. max_in_flight beschränkt also, wie viele davon zusammenfallen, und die Höhen bestimmen, wie groß jeder ist.

Gemessen an Block 257.510 (60 Transaktionen, 101,5 Mgas, 198 Mio. Zyklen, 207 Shards) auf 32 vCPUs und 247,7 GiB, mit zwölf gleichzeitig bearbeiteten Shards:

Durchlauf 1 191 s; Befüllungen auf einem einzigen Thread machen 81 % seiner Shard-Sekunden aus; per Stichprobe gemessener Speicher höchstens 15,9 GiB
Durchlauf 2 2.290 s, mit etwa 12 von 12 gehaltenen Shards und 30,4 ausgelasteten vCPUs bis zum Exit des Gastprogramms, danach eine Schlussphase von 460 s
Speicherspitze 173,92 GiB: die beiden KECCAK_F-Shards der Höhe 2^18, gemeinsam in der Schlussphase, ohne dass sonst etwas in Bearbeitung war

Die Spitze ergab sich aus der Höhe einer einzigen Delegationsfamilie, nicht aus den zwölf gleichzeitig bearbeiteten Shards. Das ist der Hebel: Ein Block mit weniger Keccak-Aufrufen oder mit Keccak bei geringerer Höhe erreicht eine niedrigere Spitze.

Die Spezifikation: Der Streaming-Prover.

Architektur

Rekursion und Abwicklung

Wie aus einem Basisbeweis aus Hunderten von Shards ein einziger Groth16-Beweis wird, den ein Contract prüft. Apogee beim Beweisen seines eigenen Verifiers, die Tapes, die das günstig machen, das über einen Baum verkettete Transkript und Pairings, die gefaltet werden, bis nur eines übrig bleibt.

Als Markdown anzeigen

Ein Basisbeweis eines Ethereum-Blocks besteht aus 207 Shard-Beweisen: 14,5 MB, jeder Shard ein GKR-Beweis mit seinen Commitments und einer Mercury-Öffnung, die in einer Pairing-Prüfung endet. Ein Contract kann nichts davon direkt prüfen. Die Rekursion komprimiert ihn, und die Art, wie sie das tut, ist nach der GKR-Engine selbst die folgenreichste Designentscheidung.

Der Basisbeweis bleibt unberührt#

Die erste Entscheidung betrifft das, was die Rekursion nicht tut. Kein Basisschlüssel, keine Basisaussage und kein Basisbeweis ändert sich, um die Rekursion möglich zu machen: Ein Blatt verifiziert Basis-Shards genau so, wie es ein nativer Verifier täte. Alles, was die Rekursion braucht, wird oberhalb des Basisbeweises hinzugefügt, nie in ihm. Ein Block lässt sich aus denselben Bytes nativ verifizieren, rekursiv verarbeiten oder beides.

Apogee beweist seinen eigenen Verifier#

Ein Knoten des Baums ist Apogee beim Beweisen eines Verifier-Programms. Ein Blatt verifiziert eine Reihe aufeinanderfolgender Basis-Shards, von from bis to; ein innerer Knoten verifiziert zwei bis vier Kinder, jedes ein vollständiger Beweis eines Blatt- oder Knotenprogramms; die Wurzel deckt jeden Basis-Shard ab. Jeder Knoten wird von demselben Streaming-Prover bewiesen wie ein Basisblock.

Den Rust-Verifier als RISC-V-Befehle auszuführen würde funktionieren und ergab gemessen 3,0 Milliarden Zyklen für die 207 Shards eines Blocks: das Fünfzehnfache des Blocks selbst. Deshalb laufen Knoten stattdessen in einem Rekursionsformat.

Das Rekursionsformat#

Eine Aussage ist genau dann im Rekursionsformat, wenn ihr Programm Körperfamilien deklariert. Das Format fügt einen Adressraum und vier Koprozessor-Familien darauf hinzu, die alle über das gewöhnliche Delegations-ABI aufgerufen werden, und ändert sonst nichts:

Familie Eine Zeile ist
FIELD_WINDOWS eine Körperzelle: eine Speicherzelle, die ein ganzes Fr-Element enthält, in derselben Speicher-Multimenge wie der RAM
FR_OP eine Körperoperation über Zellen: Multiplizieren, Addieren, Subtrahieren, Multiply-Accumulate, Invertieren, Assert-Equal und Schritte zum Aufbau von Konstanten
P2_FIELD ein Poseidon2-Duplex-Schritt über Zellen; ein Transkript läuft also mit einer Zeile pro Permutation
FIELD_IO acht RAM-Wörter in eine Zelle oder eine Zelle zurück in acht Wörter
FQ_OP eine Operation im Basiskörper von BN254, wobei ein Element aus vier Zellen mit 64-Bit-Limbs besteht; Kurvenarithmetik läuft also mit einer Zeile pro Körperoperation

Zwei weitere Änderungen machen einen Rekursions-Shard für seinen Elternknoten günstiger zu verifizieren. Seine Speicher- und Witness-Spalten werden als Stapel von bis zu 2^24 Auswertungen committet; ein Elternknoten faltet also eine Handvoll Punkte statt Hunderter. Und eine Rekursionsanfrage schreibt ihre Frame-Basis hinter das Ende des Frames vorgerückt zurück; direkt hintereinander liegende Frames werden also als direkt aufeinanderfolgende ecalls abgespielt, eine Zeile pro Aufruf.

Tapes#

Die Prüfungen eines Shards haben für seine Familie und Höhe eine feste Form. Deshalb kompiliert der Host sie einmal in ein Tape: eine lineare Liste von Koprozessor-Aufrufen über absoluten Zellen, in der nichts abhängig von einem Wert verzweigt. Das Tape eines Shards besteht aus den Schritten des nativen Verifiers für diesen Shard, Aufruf für Aufruf: das Shard-Transkript, der GKR-Rückwärtsdurchlauf, die Lookup- und Wurzelprüfungen und die zwölf Skalare der Mercury-Öffnung. Jede Prüfung ist ein Assert-Equal.

Die Tapes, Faltungsvorlagen und Konstanten jedes Rekursionsprogramms werden zur Compile-Zeit vom Verifier-Crate selbst gebaut und in den schreibgeschützten Daten des Programms abgelegt. Die Programmidentität bindet daher jedes Tape, das das Programm abspielt: Zu beweisen, dass ein Knoten sein Programm ausgeführt hat, heißt zu beweisen, dass er genau diese Prüfungen ausgeführt hat.

Ein über den Baum verkettetes Transkript#

Das globale Transkript der Basisaussage ist ein einziger Sponge über die gesamte Aussage. Der Baum teilt es auf, ohne es zu verändern. Der Knoten, der Shard 0 hält, führt das Präfix aus, bis einschließlich des Digests der öffentlichen Eingabe; jeder Knoten absorbiert die Speicher-Commitments seiner eigenen Shards und setzt dabei an dem Zustand an, den sein Vorgänger hinterlassen hat; der Knoten, der den letzten Shard hält, führt das Suffix aus und zieht die Speicher-Challenges, die jeder Knoten als Behauptungen übernommen hatte. Das Journal eines Knotens hält den Zustand der Kette an beiden Enden seines Bereichs fest, und ein Elternknoten verlangt, dass die Zustände seiner Kinder aneinander anschließen.

Ein Knoten gleicht seine Kinder außerdem miteinander ab: Exit-Status 0, eine einzige Basisaussage (ihre Form, ihr Digest, ihre Challenges, der Digest von Eingabe und Journal, ihr Exit-Status und ihre Shard-Zahl), aneinandergrenzende Shard-Bereiche, aneinander anschließende Kettenzustände, Zeitfenster in Reihenfolge über die Naht hinweg und die Identitäten der Rekursionsprogramme. Ein Knoten, der eine ganze Aussage hält, führt das Speicherargument aus.

Die Pairings falten#

Kein Knoten berechnet ein Pairing. Die Mercury-Prüfung jedes Shards wird als zwölf Einträge (side, scalar, point) aufgeschoben; nach dem Tape des Shards absorbiert das eigene Transkript des Knotens den Endzustand des Shard-Transkripts und zieht Gewichte, und jeder Eintrag wird gewichtet in ein einziges laufendes Punktepaar (A, B) addiert, das die Behauptung e(A, [1]_2) = e(B, [x]_2) darstellt. Die Batch-Prüfung, die das kombinierte Commitment eines Shards an seine Spalten bindet, wird daneben gefaltet. Punkte, die alle Shards einer Familie teilen, etwa [1]_1 und die Setup-Commitments, sammeln je einen Skalar an und gehen einmal ein. Das (A, B) eines Kindes geht unter einem Gewicht ein, das nach seinem gesamten Journal gezogen wird.

Jede Seite ist eine einzige Multi-Skalar-Multiplikation auf FQ_OP, ausgeführt als statische Vorlage: Pippenger mit 8-Bit-Ziffern über GLV-Hälften, für jeden Punkt erzwungen, dass er auf der Kurve liegt, jeder Schritt im Voraus festgelegt. Ein Punkt kostet etwa 400 FQ_OP-Aufrufe.

An der Wurzel ist der gesamte Inhalt des Baums zusammengefallen: jeder Basis-Shard verifiziert, das Transkript von Anfang bis Ende ausgeführt, das Speicherargument geführt und jede Öffnung in eine einzige Pairing-Behauptung gefaltet. Übrig bleiben diese Behauptung und zwei Programmidentitäten.

Der Decider#

Die Wurzel ist immer noch ein GKR-Beweis und einige Hundert Punkte, was ein Contract nicht prüfen kann. Der Decider ist ein Groth16-Schaltkreis, der das Knotenverfahren über einem einzigen Kind, der Wurzel, ausführt, und zwar über einen Treiber, der Rank-1-Constraints statt Koprozessor-Aufrufen schreibt, und der das Journal der Wurzel an den gesamten Bereich der Basis-Shards bindet. Er faltet nichts: Jeder Punkt, den die Wurzel dem finalen Pairing schuldet, wird mit seinem Skalar zu einem gebundenen Wire, einem Wert, den der Verifier besitzt, im Beweis unter einer fünften Falltür committet, statt als öffentliche Eingabe übergeben zu werden. Dasselbe gilt für die beiden Identitäten, den Exit-Status der Basisaussage sowie die öffentliche Eingabe und das Journal der Basisaussage, Byte für Byte.

Das Groth16 von Apogee weicht in drei Punkten vom Lehrbuch ab: das Commitment auf die gebundenen Wires, keine Verblindung und ein Beweisschlüssel über der Lagrange-Basis, die die Powers-of-Tau-Zeremonie bereits veröffentlicht. Sein Schlüssel stammt aus einer zweiphasigen Zeremonie: Phase 1 ist dieselbe Zeremoniedatei, auf der die Commitments beruhen; Phase 2 gehört allein dem Schaltkreis, wobei die Beiträge zu α und β abgeschlossen sind, bevor irgendein Beitrag zu γ, δ und η erfolgt, eine Reihenfolge, die selbst Teil der Soundness ist.

ApogeeVerifier.sol rekonstruiert die gebundenen Werte aus den Calldata, prüft die Groth16-Gleichung, faltet die Punkte beider Seiten mit ecMul und ecAdd, was zugleich erzwingt, dass jeder Punkt auf der Kurve liegt, und prüft das eine verbleibende Pairing. Sein Konstruktor legt den Schlüssel, die beiden G2-Punkte der Zeremonie und die Identitäten der beiden Rekursionsprogramme fest. Ein Deployment bedient ein einziges Basisprogramm, eine einzige Wurzelform und feste Längen der öffentlichen Werte.

Gemessen#

Block 257.510, der Baum auf einer Maschine mit 32 CPUs, Zeremonie und Decider auf einem Laptop mit 18 Kernen:

Basisbeweis 207 Shards, 14,5 MB, 2.481 s
Baum 4 Blätter aus höchstens 64 Basis-Shards und eine Wurzel: insgesamt 116 Shards
Blätter, vier gleichzeitig 21, 24, 23 und 27 Shards; 2.157 s; Spitze 92 GiB
Wurzel 21 Shards, 460 s, 1,03 MB
Decider 7.896.686 Constraints; Beweis in 18,5 s und 6,1 GB
Contract 358 Punkte; 3.620.026 gas; 34.980 Byte Calldata

Die Spezifikation: Rekursion und Decider. Selbst ausführen: On-Chain abwickeln.

Architektur

Ethereum-Blöcke

Die Referenz-Arbeitslast von Apogee. Ein revm-Gastprogramm, das Ethereum-Blöcke innerhalb der VM ausführt, ein zustandsloser Validator, dessen Ergebnisse mit jedem Fall des zkEVM-Test-Release übereinstimmen, und was ein Beweis eines Blocks aussagt.

Als Markdown anzeigen

Apogee beweist beliebige RV32IMAC-Programme. Seine Referenz-Arbeitslast, an der es gemessen und abgestimmt wird, ist die schwierigste gängige: die Validierung eines Ethereum-Blocks innerhalb der VM mit revm, der EVM in Rust. Eine einzige Bibliothek, revm_block, wird sowohl für das Gastprogramm (Guest) als auch für den Host kompiliert, und zwei Binaries beweisen zwei verschiedene Aussagen.

Zwei Binaries, zwei Aussagen#

Binary Hilfsdaten (Advice) Journal Sagt aus
revm-block ein BlockWitness: der Vorzustand, den die Transaktionen lesen pro Transaktion ihr Status, ihr Gas und ihre Rückgabedaten; ein Commitment auf die Logs; eine Zusammenfassung des Nachzustands irgendein kanonischer Witness bringt revm_block::run dazu, dieses Journal zu erzeugen: ein Beweis einer Ausführung, nicht der Gültigkeit eines Blocks
revm-block-stateless die zustandslose Eingabe des zkEVM-Benchmark-Formats 43 Byte: die Wurzel des Payloads, das Urteil, die Chain-ID, die Schema-ID der Payload mit dieser Wurzel ist auf dieser Chain unter diesem Fork ein gültiger Block oder nicht

Der zustandslose Validator ist derjenige, der Blöcke beweist. Er implementiert verify_stateless_new_payload aus den Execution Specifications von Ethereum: Er dekodiert die Anfrage, prüft die Vorfahren-Header und die Header-Regeln gegen den Elternblock, rekonstruiert jeden Absender, führt jede Transaktion gegen einen Vorzustand aus, der über Hashes an die Zustandswurzel des Elternblocks gebunden ist, wendet Withdrawals und Requests an und berechnet die Receipts-Wurzel, den Bloom-Filter, das Gas, den Requests-Hash, die Block Access List und die Wurzel des Nachzustands neu. Der Witness braucht keine eigene Bindung: Die veröffentlichte Wurzel legt den Payload fest, und der Witness ist über Hashes an sie gebunden; ein falscher Witness kann also keinen ungültigen Payload gültig machen. Ein Urteil false bedeutet nur, dass diese Eingabe die Validierung nicht bestanden hat.

Forks und Konformität#

Der Validator bestimmt den Fork anhand der Schema-ID der Eingabe, ohne einkompilierten Aktivierungsplan: Osaka, BPO1, BPO2 und Amsterdam. Alle 67.251 Paare des Release tests-zkevm v21.0.1 stimmen nativ überein, und die CI prüft die Bibliothek gegen eine eingecheckte Teilmenge von 34 Fällen, die jede Regel abdeckt, die das Release erreicht, in beiden Eingabe-Layouts.

Delegationen in der Praxis#

Beide Binaries deklarieren KECCAK_F, SHA256_COMP, MOD_MUL und EC_ADD. Keccak erreicht seinen Schaltkreis über den Native-Keccak-Hook von alloy-primitives. SHA-256, secp256k1 und BN254 erreichen ihre Schaltkreise über gepatchte Kopien von revm-precompile, k256 und ark-ff, jeweils mit dem Upstream-Code als Fallback-Pfad. Jeder Absender wird im Gastprogramm nach den Regeln von EIP-2 rekonstruiert, mit k256-Arithmetik, die die Patches zu MOD_MUL und EC_ADD leiten.

Die Binaries werden mit --release gebaut und mit jeder Familie, deren Höhe wählbar ist, bei 2^20 bewiesen: Der .text-Abschnitt des zustandslosen Binarys ist etwa 1,96 MB groß, 96,6 % dessen, was eine 2^20-Tabelle erreicht.

Der gemessene Block#

Block 257.510 von glamsterdam-devnet-8, über revm-block-stateless: 60 Transaktionen, 101,5 Mgas, 198 Millionen Zyklen, geschnitten in 207 Shards. Der Basisbeweis dauerte 2.481 s auf einer Maschine mit 32 vCPUs und 247,7 GiB und erreichte eine Spitze von 174 GiB, die Rekursion kam mit etwa 2.620 s hinzu, und der Contract akzeptierte das Ergebnis für 3.620.026 gas. Performance schlüsselt jede Stufe auf.

Was ein Blockbeweis nicht leistet#

  • Der Validator bezieht seine Eingabe von einem externen Witness-Erzeuger. eth_getProof liefert die Trie-Knoten auf dem Pfad jedes Schlüssels, aber eine Löschung, die einen Branch zusammenfallen lässt, braucht einen Geschwisterknoten, der auf dem Pfad keines geänderten Schlüssels liegt. Der Recorder des Repositorys kann deshalb keine zustandslosen Eingaben erzeugen; sie stammen aus einem Release von tests-zkevm oder aus den Datensätzen des zkEVM-Benchmarks.
  • Das Mini-Block-Binary beweist eine Ausführung über einem aufgezeichneten Vorzustand, meist die der ersten paar Transaktionen eines Blocks, und stellt keine Behauptung über die Zustandswurzel auf. Sein Journal wächst um einen Datensatz pro Transaktion und wird größer als das öffentliche Fenster; deshalb laufen vollständige Blöcke über den zustandslosen Validator.

Die Spezifikation: Ethereum-Blöcke.

Architektur

Sicherheitsmodell

Was ein Beweis begründet, was er voraussetzt, was ein Verifier selbst besitzen muss, auf welchem Code die Soundness beruht und die Grenzen von Version 1.0.0.

Als Markdown anzeigen

Was ein Beweis begründet#

Ein Beweis, der die Verifikation besteht, begründet, dass das Programm einer gegebenen Identität, gestartet an seinem Einsprung-pc auf seinem Image, mit der öffentlichen Eingabe in seinem Eingabefenster und irgendwelchen vom Prover gewählten Hilfsdaten (Advice), Befehl für Befehl bis zu EXIT mit einem gegebenen Status ausgeführt wird und dabei ein gegebenes Journal geschrieben hat. Über die Hilfsdaten wird nichts behauptet. Nichts wird verborgen.

Annahmen#

Annahme Wo sie eingeht
Knowledge Soundness von Mercury und KZG im algebraischen Gruppenmodell unter q-DLOG jede Öffnung eines Commitments
Poseidon2 als Random Oracle für Fiat–Shamir jede Challenge, im Basisbeweis und in der Rekursion
Die eigenen Annahmen von Groth16 der Decider, der letzte Schritt zum Contract
Ein ehrlicher Beitragender zu den Perpetual Powers of Tau der PSE der SRS, auf dem jedes Commitment beruht
Ein ehrlicher Beitragender pro Runde der Phase-2-Zeremonie des Deciders der Schlüssel des Deciders

BN254 bietet etwa 100 Bit Sicherheit. Der statistische Fehler jeder Protokollschicht, von den Sumchecks und der gebündelten Öffnung bis zu den Speicher- und Lookup-Argumenten, liegt weit darunter: unter 2^14/|Fr| für den gesamten GKR-Durchlauf eines Schaltkreises, unter 2^−190 für jeden Lookup-Kanal, unter 2^−220 für jede Mercury-Instanz.

Was ein Verifier selbst besitzen muss#

Zwei Werte, bezogen über einen Kanal, den der Prover nicht kontrolliert:

  • Die Programmidentität. Gegenüber einer vom Prover gelieferten Identität zeigt ein Beweis nur, dass irgendein Programm gelaufen ist.
  • Der SRS-Digest der Zeremonie. Ein Schlüssel wird unter jedem Digest geladen, den seine eigenen Punkte ergeben; ein Schlüssel, der über einem bekannten τ gebaut wurde, wird also nur durch diesen Vergleich zurückgewiesen.

Der Verifikationsschlüssel selbst darf von beliebiger Seite stammen. Beim Laden werden die Identität und der SRS-Digest aus seinem eigenen Inhalt neu berechnet, und seine Schaltkreise werden mit der Registry des Verifiers abgeglichen: Die Identität bindet das Programm, die Registry bindet die Schaltkreise. Das Kommandozeilenwerkzeug verifier vergleicht nur die Identität und entnimmt den SRS-Digest dem Schlüssel; host::verify vergleicht keines von beiden und überlässt es seinem Aufrufer, auch Eingabe, Journal und Exit-Status der Aussage zu prüfen.

Vertrauenswürdiger Code#

Die Soundness liegt allein beim Verifier. Sie beruht auf constants, field, curve, transcript, poly, sumcheck, pcs-verify, pcs, gkr-verify, verifier-core, verifier und constraints, weil die Schaltkreise Teil der Aussage sind und ein fehlendes Gate ein Soundness-Fehler ist. Das Berechnen einer Identität aus einem ELF vertraut außerdem loader, isa und program. Der letzte Schritt zur Chain fügt die Rekursionsprogramme, groth16, den Schaltkreis des Deciders und den Contract hinzu.

Dem Prover, dem Emulator, den Trace-Buildern und der beweisenden Hälfte des Host-SDK wird nicht vertraut. Der Prover validiert nichts; eine falsche Eingabe führt bei einem ehrlichen Prover zu einem Beweis, der fehlschlägt, und ein betrügerischer Prover führt diesen Code ohnehin nicht aus.

Nicht in konstanter Zeit, nicht zero-knowledge#

Nichts im Code läuft in konstanter Zeit: Reduktionen, Exponentiationen, Punktadditionen und Skalarleitern verzweigen abhängig von ihren Operanden. Das ist hier unschädlich, weil kein Beweis zero-knowledge ist und das Beweisen somit nichts geheim hält. Das einzige Geheimnis, mit dem der Code umgeht, ist der Faktor eines Beitragenden zur Zeremonie des Deciders, und er durchläuft dieselbe Leiter mit variabler Laufzeit; führen Sie Beiträge auf einer Maschine aus, die Sie kontrollieren.

Grenzen von v1.0.0#

Grenze Einzelheiten
Nicht zero-knowledge keine Verblindung in Mercury, GKR oder dem Decider
Hilfsdaten sind ungebunden ein Gastprogramm (Guest) prüft sie gegen etwas, das ein Beweis bindet
Öffentliche Werte höchstens je 16.380 Byte Eingabe und Journal
sc.w gelingt immer die einzige Abweichung von der Semantik von RV32IMAC; es gibt keinen Reservierungszustand
Traps sind nicht beweisbar ein nicht ausgerichteter Zugriff, ein Zugriff außerhalb des gemappten Speichers, ebreak oder ein pc ohne Befehl beendet eine Ausführung ohne Beweis
Code ist statisch der Befehlsstrom ist das beim Laden dekodierte Image; ein einziges nicht dekodierbares Wort im ausführbaren Code führt zur Zurückweisung des Programms
Codegröße .text innerhalb der Reichweite einer dekodierten Tabelle, 7,94 MiB bei 2^22; das Image standardmäßig innerhalb von 4 MiB
Ausführungslänge 2^36 − 1 Zyklen
Delegationen bilden eine feste Menge sechs im Basisformat; eine Delegation beweist einen Schritt ihrer Funktion, und das Zusammensetzen der Schritte, darunter die Validierung von Kurvenpunkten, ist Sache des aufrufenden Codes
Speicher des Provers bestimmt durch die gleichzeitig bearbeiteten Shards: Der gemessene Block erreichte eine Spitze von 174 GiB
Block-Witnesses der zustandslose Validator bezieht seine Eingabe von einem externen Witness-Erzeuger
Der Schlüssel des Deciders einer pro Wurzelform und nur so vertrauenswürdig wie seine Zeremonie; der Entwicklungsschlüssel ist fälschbar
On-Chain-Kosten etwa 3,6 Mio. gas für den gemessenen Block

Wohin die nächste Version das verschiebt#

Jede der obigen Annahmen, die BN254 betrifft, von q-DLOG und dem Pairing bis zu Groth16, hält einem hinreichend großen Quantencomputer nicht stand. Die Flugbahn zu v2.0.0 führt zu einem Beweiskern, dessen Soundness stattdessen auf Gitterproblemen beruht.

Die Sicht eines Auditors auf dasselbe Modell, Crate für Crate und Argument für Argument: Audit-Leitfaden, Soundness-Karte.

Architektur

Performance-Messwerte

Jeder Messwert für v1.0.0 mit seiner Quelle: der Basisbeweis eines vollständigen Ethereum-Blocks, der Rekursionsbaum, der Decider und der Contract, der Shard-Beweis jeder Familie und die Kosten von Mercury selbst.

Als Markdown anzeigen

Alle End-to-End-Zahlen beziehen sich auf Block 257.510 von glamsterdam-devnet-8, bewiesen über das Gastprogramm (Guest) des zustandslosen Validators: 60 Transaktionen, 101,5 Mgas, 198 Millionen Zyklen. Jede Zahl stammt aus den Messungen der Spezifikation.

Von Anfang bis Ende#

Stufe Maschine Ergebnis
Basisbeweis 32 vCPUs, 247,7 GiB, 12 gleichzeitig bearbeitete Shards 207 Shards, 14,5 MB, 2.481 s; Spitze 173,92 GiB
Rekursionsblätter 32 CPUs, vier Blätter gleichzeitig 21, 24, 23 und 27 Shards; 2.157 s; Spitze 92 GiB
Rekursionswurzel dieselbe Maschine, vier gleichzeitig bearbeitete Shards 21 Shards, 460 s, 1,03 MB
Zeremonie des Deciders Laptop mit 18 Kernen init 65 s; ein Beitrag 50–56 s; key 70 s und 12,7 GB; der Schlüssel 2,65 GB
Decider-Beweis Laptop mit 18 Kernen Schlüssel in 1 s eingelesen, Beweis 18,5 s, 6,1 GB; 7.896.686 Constraints über 2^23
On-Chain-Verifikation revm 3.620.026 gas; 34.980 Byte Calldata; 358 Punkte

Der Basisbeweis, Durchlauf für Durchlauf#

Durchlauf 1, Commit 191 s; im Mittel 25,7 ausgelastete vCPUs; Befüllungen auf einem einzigen Thread machen 81 % seiner Shard-Sekunden aus; per Stichprobe gemessener Speicher höchstens 15,9 GiB
Durchlauf 2, Beweis 2.290 s; im Mittel 11,95 von 12 Shards gehalten und 30,4 vCPUs ausgelastet bis zum Exit des Gastprogramms; danach eine Schlussphase von 460 s, deren längste Abschnitte die Befüllungen der beiden KECCAK_F-Shards auf einem einzigen Thread sind, 200 s und 279 s
Speicherspitze 173,92 GiB: die beiden KECCAK_F-Shards der Höhe 2^18, gemeinsam in der Schlussphase, ohne dass sonst etwas in Bearbeitung war

Die Spitze wurde durch die Höhe einer einzigen Delegationsfamilie bestimmt, nicht durch die Zahl der gleichzeitig bearbeiteten Shards.

Ein kleines Gastprogramm#

Das Gastprogramm des Schnellstarts, 114 Zyklen in vier Befehlsfamilien, bewiesen bei Befehlshöhen von 2^20 und Fensterhöhen von 2^16 mit zwei gleichzeitig bearbeiteten Shards: 7 Shards, 52 s und eine Spitze von 18 GB auf einem Laptop mit 18 Kernen und 48 GiB, fast alles davon die beiden 2^20-Shards in Bearbeitung. Die Untergrenze eines Beweises bestimmen seine Familien und Höhen, nicht seine Zyklenzahl.

Der Shard jeder Familie#

Bei den Standardhöhen. Die Beweisgröße eines Shards ist durch Form und Höhe seines Schaltkreises festgelegt; seine Beweiskosten folgen seiner Höhe mal der Breite seines Schaltkreises, unabhängig davon, wie viele Zeilen aktiv sind.

Familie Höhe Committet M/W/S Constraint-Gates Innere Spalten Shard-Beweis
ADD_SUB_LUI_AUIPC 2^22 27 / 35 / 7 63 314 64.764 B
JUMP_BRANCH_SLT 2^22 21 / 44 / 10 42 392 69.436 B
SHIFT_BITWISE 2^22 21 / 61 / 10 48 478 76.644 B
MUL_DIV 2^20 21 / 54 / 9 54 444 67.412 B
MEM_WORD 2^22 31 / 24 / 7 33 314 63.836 B
MEM_SUBWORD 2^22 31 / 55 / 10 53 472 76.196 B
ATOMICS 2^20 26 / 54 / 9 46 472 68.468 B
INIT_TEARDOWN 2^22 2 / 0 / 1 0 46 36.316 B
ZERO_WINDOWS 2^22 2 / 0 / 0 0 46 36.284 B
KECCAK_F 2^18 208 / 1.556 / 0 385 5.490 381.100 B
POSEIDON2 2^8 100 / 4.092 / 0 4.248 2.020 664.780 B
FR_ARITH 2^8 104 / 2.576 / 0 2.701 142 266.292 B
PUBLIC_INPUT, PUBLIC_OUTPUT 2^12 3 bzw. 2 / 0 / 0 0 26 12.556 B, 12.524 B
ADVICE_WINDOWS 2^22 3 / 0 / 0 0 46 36.316 B
MOD_MUL 2^16 104 / 221 / 0 125 2.244 135.220 B
SHA256_COMP 2^18 104 / 520 / 0 119 2.802 189.988 B
EC_ADD 2^16 392 / 1.028 / 0 637 8.772 434.916 B

Eine Höhe ändert nur die Zahl der halbierenden Listen und Sumcheck-Runden, nicht die Gates: Bei 2^20 hat ADD_SUB_LUI_AUIPC 298 innere Spalten und einen Beweis von 57.196 Byte, gegenüber 314 und 64.764 bei 2^22.

Speicher des Vorwärtsdurchlaufs#

Der GKR-Prover hält jede innere Schicht als Körperelemente, 32 Byte pro Zelle. Repräsentativer Speicherbedarf:

Shard Vorwärtsdurchlauf
SHIFT_BITWISE bei 2^20 8,4 GiB
MOD_MUL bei 2^16 4,6 GB
EC_ADD bei 2^16 18,3 GB
SHA256_COMP bei 2^18 22,6 GB
KECCAK_F bei 2^18 42 GiB

Mercury#

Auf einem Apple M5 Pro mit 18 Kernen:

Commit, n = 2^22 1,30 s
Öffnung, n = 2^22 2,89 s
16 Spalten der Größe 2^20 als ein Batch geöffnet in 1,01 s, verifiziert in 4,8 ms
Dieselben 16 einzeln geöffnet 9,79 s, verifiziert in 62 ms

Wie diese Zahlen zu lesen sind#

Die Beweiserzeugung ist speichergebunden, und ihr Speicherbedarf folgt den gleichzeitig bearbeiteten Shards und deren Höhen, nie der Länge der Ausführung. Die Zeit folgt der Zyklenzahl, Familie für Familie. Die On-Chain-Kosten folgen der Zahl der Punkte, die die Wurzel dem finalen Pairing schuldet, mit etwa 9.000 gas pro Punkt. Zyklenzahlen selbst sind exakt und maschinenunabhängig; der Zyklus-Profiler ist also das richtige erste Werkzeug, um alles Übrige abzuschätzen.

Quantensprung

Quantensprung

Wohin Apogee als Nächstes geht. Die Flugbahn zu v2.0.0 – ein Beweiskern auf Gittern, ein passend dazu gewählter Körper, Signaturen, die Gastprogramme verifizieren können, und ein Deployment-System, das eine Anwendung vom Repository bis zur laufenden Chain bringt.

Als Markdown anzeigen
BriefingProgramm: Apogee VMZiel: v2.0.0Status: aktive EntwicklungRelease-Fenster:

Version 1.0.0 klärt die Frage, ob die Architektur im vollen Maßstab trägt: ein ganzer Ethereum-Block, von einem Gastprogramm (Guest) in Rust bis zu einem Contract, der true sagt. Version 2.0.0 ändert, worauf diese Architektur ruht und wem sie dient. Der Beweiskern wechselt auf Mathematik, die ein Quantencomputer nicht bricht, und die Oberfläche für Entwickler wächst von einem Repository zu einem System, das Anwendungen deployt.

Wichtig

Dieser Abschnitt beschreibt laufende Arbeiten. Nichts hier ändert die Garantien von v1.0.0, die vollständig im Sicherheitsmodell dargelegt sind.

Die vier Initiativen#

Die Flugbahn#

v1.0.0, heute v2.0.0, die Flugbahn
Commitments Mercury über KZG: Pairings, q-DLOG Gitterbasiert, bindend unter Module-SIS
Körper Der Skalarkörper von BN254, 254 Bit Ein kleiner Primkörper mit Challenges aus einem Erweiterungskörper, passend zum Commitment
Gegen einen Quantenangreifer Jede Annahme ist ein diskreter Logarithmus Ein Beweiskern, der auf Gitterproblemen beruht
Signaturen in Gastprogrammen secp256k1 über delegierte Körper- und Kurvenarithmetik ZK-freundliche und Post-Quanten-Verfahren als Aufrufe im Gastprogramm
Für Entwickler Ein Repository, seine Werkzeuge und dieses Handbuch Das Deployment-System: Portal, kanonische Bridges, Telemetrie, ein KI-Gateway

Was erhalten bleibt#

Der Sprung betrifft die Fundamente, nicht das Modell, gegen das ein Entwickler programmiert:

  • Das Gastprogramm. Rust, RISC-V, drei Speicherbereiche, eine Identität, ein Journal. Für v1.0.0 geschriebene Programme behalten ihre Form.
  • Die GKR-Engine. Geschichtete Schaltkreise und Sumcheck sind über jedem Körper definiert. Die Engine, die einen Schaltkreis auf einen einzigen Punkt zurückführt, ist der Teil von Apogee, der am direktesten in den neuen Körper wechselt.
  • Die Argumente. Eine einzige Speicher-Multimenge über die gesamte Ausführung und die LogUp-Kanäle bleiben als Konstruktionen erhalten; ihre Tabellen und Bereichsargumente werden für die Größe des neuen Körpers neu hergeleitet.
  • Die Disziplin. Zuerst die Spezifikation, jede Schicht gegen ein unabhängiges Orakel geprüft, jede Fälschungsklasse durch einen Manipulations-Zwilling abgedeckt.

Warum jetzt#

Ein Gültigkeitsbeweis ist nur so quantensicher wie das System, das ihn erzeugt. Ein Beweis über BN254 beruht auf Pairings und diskreten Logarithmen; ein hinreichend großer Quantencomputer könnte ihn also fälschen, ohne irgendetwas anzutasten, was das Gastprogramm berechnet hat. Blockchain-native Anwendungen sollen über Jahrzehnte Werte halten. Das Fundament, auf dem ihre Abwicklung beruht, muss die Maschinen überdauern, die eines Tages die heutigen Kurven brechen werden, und der richtige Zeitpunkt für den Umzug liegt, bevor es diese Maschinen gibt.

Zu den Initiativen: Post-Quanten-Beweise · Signaturen für Gastprogramme · Das Deployment-System.

Quantensprung

Post-Quanten-Beweise

Initiativen QL-01 und QL-02. Ein Gitter-Commitment anstelle des pairing-basierten und ein passend dazu gewählter Körper, damit der Beweiskern nicht länger auf diskreten Logarithmen beruht.

Als Markdown anzeigen
QL-01 · QL-02Gitterbasierte CommitmentsKörperwechselStatus: aktive Entwicklung

Was bricht, und wo#

Jede kryptografische Annahme, auf der Apogee v1.0.0 ruht, betrifft BN254. Die Soundness von Mercury und KZG gilt unter q-DLOG im algebraischen Gruppenmodell; der Rekursionsbaum faltet Pairing-Prüfungen; der Decider ist Groth16. Shors Algorithmus löst auf einem hinreichend großen Quantencomputer diskrete Logarithmen und bricht damit jede einzelne dieser Annahmen. Die Berechnung des Gastprogramms (Guest) bliebe, was sie war; der Beweis, dass sie korrekt lief, hätte keinerlei Bedeutung mehr.

Im Commitment konzentriert sich die Abhängigkeit. Jede Spalte jedes Shards wird damit committet, jede Öffnung endet in seinem Pairing, und die Rekursion existiert, um diese Pairings zu falten. Ersetzt man das Commitment, bleibt im Rest des Beweiskerns nichts mehr, das von einem diskreten Logarithmus abhängt.

QL-01 · Gitterbasierte Commitments#

Ein Gitter-Commitment ist eine lineare Abbildung, t = A·s mod q, angewandt auf einen Vektor s mit kleinen Einträgen. Es ist bindend, solange niemand einen kurzen Vektor finden kann, den die Matrix auf null abbildet: Module-SIS, die Annahmenfamilie hinter ML-DSA und ML-KEM, den Post-Quanten-Standards des NIST, gestützt auf Worst-Case-Reduktionen und jahrzehntelange Kryptoanalyse.

Es bewahrt, was KZG für Apogee so nützlich gemacht hat und was hashbasierte Commitments aufgeben: Es ist homomorph. Commitments auf viele Teilstücke lassen sich mit Challenge-Koeffizienten kombinieren, und die Kombination wird durch einen einzigen Vektor geöffnet. Das Bündeln der Spalten eines Shards, das Aufschieben von Prüfungen und ihr Falten einen Baum hinauf sind lineare Operationen, und lineare Operationen überstehen den Wechsel. Merkle-Pfade lassen sich überhaupt nicht kombinieren.

Der Preis ist eine Einschränkung, die es anderswo nicht gibt: Das Commitment bindet nur kurze Vektoren, jede Kombination macht den Vektor länger, und der Prover muss zeigen, dass er noch kurz genug ist. Die Verfahren der letzten zwei Jahre unterscheiden sich vor allem darin, wie sie diesen Preis zahlen, und sie haben sich schnell entwickelt. Für Polynome mit 2^30 Koeffizienten liefern die veröffentlichten Module-SIS-Verfahren Auswertungsbeweise von 53 bis 72 KB, und die Verifikation sank von 2,8 Sekunden im Jahr 2024 auf 8 bis 16 Millisekunden im Jahr 2026. Die Abhandlung Lattice-Based Polynomial Commitment Schemes stellt sie Verfahren für Verfahren vor und setzt die Zahlen in Beziehung zur hashbasierten Seite.

QL-02 · Körperwechsel#

Die Gitterverfahren leben nicht in der Welt von BN254. Die führenden Konstruktionen arbeiten über kleinen Primzahlmoduln, mit Auswertungspunkten aus einem Erweiterungskörper, um die Soundness zu erhalten; das passt zu einem Beweissystem über einem kleinen Körper und nicht zu einem über einem 254-Bit-Körper. Der Wechsel des Commitments zieht also den Körper mit: v2.0.0 überträgt die Arithmetisierung auf einen neuen Körper und verlegt jeden Schaltkreis vom Skalarkörper von BN254 auf einen kleinen, zum Commitment passenden Körper.

Der Wechsel macht sich bezahlt:

  • Jede Schicht wird günstiger. Ein GKR-Prover verbringt seine Zeit mit Körperarithmetik, und eine Multiplikation in einem kleinen Körper kostet einen Bruchteil einer Multiplikation in einem 254-Bit-Körper. Die zentrale Ersparnis der Engine, dass Zwischenschichten nie committet werden, verstärkt sich mit billigerer Arithmetik auf jeder verbleibenden Schicht.
  • Commitments werden günstiger. Das Commitment auf eine Trace-Spalte ist eine lineare Abbildung über kleinen Ziffern, bezahlt pro Eintrag ungleich null, statt einer Multi-Skalar-Multiplikation über einer Kurve.
  • Die Engine bleibt erhalten. GKR und Sumcheck sind über jedem Körper definiert. Challenges wandern in einen Erweiterungskörper; der Rückwärtsdurchlauf, das Schichtenmodell und die darauf aufbauenden Argumente behalten ihre Struktur.

Neu gebaut werden muss alles, was einen großen Körper voraussetzte: Werte auf Wortebene, die mit reichlich Spielraum in ein BN254-Element passen, Bereichsargumente und Überträge, die auf einen 254-Bit-Modul ausgelegt sind, Kanonizitätsketten und die Körperzellen des Rekursionsformats. Jedes davon wird für den neuen Körper neu hergeleitet und so spezifiziert wie die von v1.0.0, mit eigenem Orakel und eigenen Manipulations-Zwillingen.

Abwicklung#

Die Verifikations-Precompiles von Ethereum sind heute pairing-basiert. Wie die Beweise eines Post-Quanten-Beweiskerns auf dieser Chain abgewickelt werden und welcher Teil des letzten Schritts auf Gittern ruhen kann, gehört zum selben Arbeitsprogramm und wird, bevor es ausgeliefert wird, mit derselben Sorgfalt spezifiziert wie der Kern.

Zurück zum Missionsbriefing.

Quantensprung

Signaturen für Gastprogramme

Initiative QL-03. ZK-freundliche und Post-Quanten-Signaturverifikation, für jedes Gastprogramm als Aufruf verfügbar, damit Autorisierung in einer Blockchain-nativen Anwendung eine einzige Codezeile ist.

Als Markdown anzeigen
QL-03Signaturen für GastprogrammeStatus: aktive Entwicklung

Fast jede Blockchain-native Anwendung stellt bei jeder Anfrage dieselbe Frage: Hat der richtige Schlüssel das autorisiert? In v1.0.0 beantwortet ein Gastprogramm (Guest) sie mit Code. Die secp256k1-Recovery ist k256, ausgeführt auf den delegierten Schaltkreisen MOD_MUL und EC_ADD, was Ethereums eigene Signaturen erschwinglich macht; alles andere besteht aus gewöhnlichen Befehlen. QL-03 macht die Signaturverifikation zu einer vollwertigen Operation des Gastprogramms.

Die Verfahren#

Ob sich ein Signaturverfahren günstig beweisen lässt, hängt fast ausschließlich von seinem Algorithmus zur Verifikation ab: welche Arithmetik er ausführt, in welchem Körper und welchen Hash er aufruft. Der Signierer läuft nie innerhalb des Beweises. Vier Entwürfe decken das Feld ab:

Verfahren Idee Warum es für ein Gastprogramm zählt
Schnorr über einer nativen Kurve Schnorrs Protokoll auf einer Kurve, deren Basiskörper der eigene Körper des Beweissystems ist, so wie Grumpkin zu BN254 konstruktionsbedingt das günstigste: Arithmetik und Hash sind beide nativ im Schaltkreis
ML-DSA (FIPS 204) Schnorr, übertragen auf Gitter, mit einer kurzen Antwort, die durch Rejection Sampling gleichverteilt bleibt die primäre Post-Quanten-Signatur des NIST; die Kosten werden von ihrem Hash und, sofern der Schlüssel nicht fest ist, von ihrer Matrixexpansion dominiert
FN-DSA (Falcon) Hash-and-Sign mit einer Gitter-Falltür, verborgen durch Gauß-Sampling die wenigste Arithmetik und das wenigste Hashing der drei Post-Quanten-Verfahren
SLH-DSA (FIPS 205) Signaturen allein aus einer Hashfunktion überhaupt keine Algebra und rund zweitausend Hash-Aufrufe; die konservativste Annahme

Alle vier Verifikationsalgorithmen folgen dem Muster „berechnen und vergleichen“, ohne Geheimnisse und ohne Verzweigungen auf Geheimnissen, und genau das macht jeden von ihnen beweisbar. Die Abhandlung ZK-Friendly Signature Schemes arbeitet jedes Verfahren an einem Spielzeugbeispiel durch und vergleicht ihre Kosten im Schaltkreis.

Was die Kosten bestimmt#

Zwei Hebel bewegen jede Zahl:

  • Der Körper, über dem der Prover arbeitet. Ein Verfahren ist nativ, wenn seine Arithmetik die Arithmetik des Schaltkreises ist. Welches Verfahren am günstigsten ist, folgt daher aus dem Körperwechsel von QL-02, und die Auswahl wird zusammen mit ihm getroffen.
  • Der Hash. Bei den Post-Quanten-Verifikationsalgorithmen dominiert ihr Standard-Hash, nicht ihre Algebra. Ihn durch einen arithmetischen Hash zu ersetzen, verlässt den Standard und ist die Variante, die die ZK-orientierten Konstruktionen wählen; die Interoperabilität mit bestehenden Schlüsseln verlangt, ihn beizubehalten. Beides hat seinen Platz, und das Gastprogramm entscheidet.

Für Entwickler#

Das Ziel ist ein Gastprogramm, das eine Autorisierung so verifiziert, wie es heute einen Hash berechnet: ein Aufruf, an einen Schaltkreis delegiert, keine Kryptografie im eigenen Code der Anwendung. Das erschließt die Muster, auf denen Blockchain-native Anwendungen aufbauen: Konten, deren Schlüssel nicht die der Chain sind, Mehrparteien-Freigaben innerhalb der Zustandsübergangsfunktion, Session-Keys und Identitäten, die gültig bleiben, nachdem die Kurven, auf denen sie entstanden sind, gefallen sind.

Zurück zum Missionsbriefing.

Quantensprung

Das Deployment-System

Initiative QL-04. Das Blockchain-native Deployment-System von Apogee, ein Portal und eine Toolchain, die eine Anwendung vom Quellcode bis zu einer laufenden Blockchain-nativen Umgebung bringen.

Als Markdown anzeigen
QL-04Blockchain-natives Deployment-System von ApogeeStatus: aktive Entwicklung

Version 1.0.0 gibt einem Entwickler ein Repository, seine Werkzeuge und dieses Handbuch. Alles zwischen einem bewiesenen Gastprogramm (Guest) und einer Anwendung im Live-Betrieb, von Schlüsseln und Zeremonien bis zu Verifier-Contracts, Bridges und Betrieb, muss der Entwickler noch selbst zusammensetzen. Das Blockchain-native Deployment-System von Apogee ist die zweite Hälfte des Sprungs: ein Portal und eine Toolchain, die ein Gastprogramm in eine laufende Blockchain-native Umgebung verwandeln und sie am Laufen halten.

Die Module#

KonsoleAktiv

Ein Ort für jede Deployment-Ressource

Programme und ihre Identitäten, Höhen und Parameter, Verifikationsschlüssel, Decider-Schlüssel und ihre Zeremonien, Verifier-Contracts und die Netzwerke, in denen sie leben, geordnet nach Anwendung und Release.

BridgesAktiv

Kanonische On-Chain-Contracts

Wiederverwendbare Vorlagen für das, was heute jede Anwendung neu baut: eine durch Beweise fortgeschriebene Registry für Zustandswurzeln, Bridges für Ein- und Auszahlungen, Upgrade-Pfade von einer Programmidentität zur nächsten.

TelemetrieAktiv

Die Vitalwerte einer Anwendung

Erzeugte und abgewickelte Beweise, Zyklen pro Anfrage, Latenz und Kosten der Beweiserzeugung, Profile nach Shard und Familie, für die Verifikation verbrauchtes Gas und die Historie der Zustandswurzeln der Anwendung.

GatewayAktiv

Eine Tür für Ihre eigene KI

Eine Schnittstelle, über die der lokale KI-Agent eines Entwicklers eine Anwendung untersuchen, ihre Telemetrie abfragen, Änderungen über die Deployment-Pipeline vorschlagen und ausführen und jedes Ergebnis zurücklesen kann, unter der Kontrolle des Entwicklers.

VorlagenGeplant

Anwendungen, die von einer funktionierenden Form ausgehen

Projekte für Gastprogramme außerhalb des Repositorys, mit bereits passenden Profilen, Linker-Einstellungen und mitgelieferten Crates und dem bereits eingerichteten KI-Begleiter.

ZeremonienGeplant

Zeremonien als Dienst, nicht als lästige Pflicht

Koordination der Phase-2-Beiträge zum Decider-Schlüssel, jeder einzelne gegen den Schaltkreis und die Zeremoniedatei verifizierbar, damit der Schlüssel eines Deployments die ehrlichen Beitragenden hat, die seine Soundness braucht.

Warum ein System und nicht mehr Werkzeuge#

Die These hinter Apogee lautet: eine optimierte Umgebung pro wirtschaftlicher Anwendung. Das vervielfacht die Zahl der Umgebungen und mit ihr die operative Oberfläche: Jede hat ihr Programm, ihre Schlüssel, ihre Zeremonie, ihre Contracts und ihre Metriken. Ein Ansatz, der jedes Team diese Oberfläche von Hand zusammensetzen lässt, skaliert nicht auf die vielen Umgebungen, die die These braucht. Das Deployment-System macht jede davon zur Routine, sodass beim Start einer Blockchain-nativen Anwendung die Anwendung selbst der schwierige Teil ist.

Das Gateway folgt aus derselben Überlegung. Entwickler arbeiten bereits Seite an Seite mit KI-Modellen, und der KI-Begleiter instruiert diese Modelle für das Schreiben von Gastprogrammen. Das Gateway gibt ihnen, unter der Kontrolle des Entwicklers, die Möglichkeit, auf Grundlage dessen zu handeln, was sie schreiben: deployen, beobachten und iterieren.

Zurück zum Missionsbriefing.

Auditoren

Audit-Leitfaden

Alles, was ein Auditor von Apogee VM v1.0.0 für den Einstieg braucht: der Umfang, die normative Spezifikation und ihr Aufbau, die Notation, die Vertrauensgrenze, eine Lesereihenfolge und die Eigenschaften, deren Prüfung sich zuerst am meisten lohnt.

Als Markdown anzeigen

Dieser Abschnitt ist die vollständige Konstruktion von Apogee VM v1.0.0, für die Begutachtung geordnet. Sein Kern ist die normative Spezifikation, wörtlich wiedergegeben: eine Seite pro Thema, mit jeder committeten Spalte nach Index und Name, jedem Gate als Polynom, jedem Lookup und seinem Kanal, jeder Transkript-Nachricht in ihrer Reihenfolge und jedem Byte jedes Serialisierungsformats. Darum herum bieten dieser Leitfaden und die Soundness-Karte einem Auditor einen Einstieg.

Umfang#

Im Umfang Wo
Die Aussage, die ein Beweis begründet, und was ein Verifier besitzen muss Das System von Anfang bis Ende, Der Beweis
Die Arithmetik: Fr, der Fq-Turm, G1 und G2, das Pairing, MSM, multilineare Polynome, der Sumcheck Primitive
Fiat–Shamir: Poseidon2, der Duplex-Sponge, jedes Tag Transkript
Das Setup und das Commitment-Verfahren Strukturierter Referenzstring, Mercury
Das Programm: Laden, Dekodieren, Tabellen, Konfiguration, Identität Programm und Identität, Gastprogramm-ABI
Das Ausführungsmodell und der Trace Ausführungs-Trace, Öffentliche Werte und Hilfsdaten
Das Beweissystem: GKR, die Schaltkreis-Registry, das Speicherargument, Lookups GKR-Engine, Schaltkreise, Speicherargument, Lookups
Jede Schaltkreisfamilie, Spalte für Spalte die sieben Befehlsfamilien und die sechs Delegationsschaltkreise
Der Aufbau des Provers Streaming-Prover
Rekursion, der Groth16-Decider und seine Zeremonie, der Contract Rekursion und Decider
Die Ethereum-Arbeitslast Ethereum-Blöcke

Die Seiten der Spezifikation sind aus dem Verzeichnis docs/ des Repositorys von Apogee VM in der Quellrevision 3571370 übernommen; relative Links wurden in Links auf dieser Website umgewandelt und jeder §-Verweis in einen Link auf seinen Abschnitt. Eine Wendung im Glossar ist umformuliert, damit sie zum Rest dieser Website passt; sonst ist nichts verändert. Wo eine Seite der Spezifikation und der Code voneinander abweichen, hat der Code recht, und die Abweichung ist ein Befund.

Wie die Spezifikation zu lesen ist#

Die Seiten sind dafür geschrieben, neben dem Code gelesen zu werden. Jede nennt das Crate und die Funktion, die implementieren, was sie festlegt, und Kommentare im Quellcode verweisen abschnittsgenau zurück auf die Spezifikation (docs/spec/memory.md §2.4). Einige wiederkehrende Konventionen:

Notation Bedeutung
M[i], W[i], S[i] committete Spalten eines Schaltkreises: Speicherspalten (im globalen Transkript gebunden, vor den Speicher-Challenges), Witness-Spalten (im eigenen Transkript des Shards gebunden), Setup-Spalten (durch die Programmidentität oder den SRS-Digest gebunden)
V[…] eine virtuelle Tabelle: eine geschlossene Form des Zeilenindex, nie committet
L{k}[j], C{k}[j], scratch[i] innere Spalte j der Schicht k; ein zwischengespeicherter Eintrag; ein Zwischenwert einer flachen Relation
W[8..14] ein halboffener Bereich von Spaltenindizes, W[8] bis W[13]
T(AS, ADDR, TS, VAL) ein Speichertupel, γ_M + AS + α_addr·ADDR + α_ts·TS + α_val·VAL
4c + Δ der Zeitstempel von Slot Δ in Zyklus c
G1–G11, S1–S6 die Schritte des globalen und des Shard-Transkripts
Schritte 1–12, B1–B6 die Prüfungen des Verifiers, in ihrer Reihenfolge, für einen Shard und einen Block
2^n eine Zweierpotenz; Höhen sind 2^8, 2^12, 2^16, 2^18, 2^20, 2^22

Ein als Ausdruck geschriebenes Gate wird auf 0 gehalten. Ein Lookup wird als sein Kanal, sein Selektor und sein Tupel geschrieben. „Beschränkt“ (bounded) bedeutet per Range-Check geprüft, und ein Wert, der Wort (word) heißt, ist eine ganze Zahl in [0, 2^32).

Lesereihenfolge#

Für einen ersten Durchgang, der das ganze Argument aufbaut, bevor er in die Schaltkreise hinabsteigt:

  1. Das System von Anfang bis Ende. Die Behauptung, die Kompositionstabelle, die Annahmen, die Grenzen.
  2. Der Beweis. Die Aussage, beide Transkripte, die Verifikationsreihenfolge, der Schlüssel und seine Laderegeln.
  3. GKR-Engine. Das Schichtenmodell, das Artefakt und seine Gesetze, der Rückwärtsdurchlauf und die Begründung seiner Soundness.
  4. Speicherargument und Lookups. Die beiden Argumente, auf denen alles beruht, was über eine Zeile hinausreicht, mit ihren Regeln zur Konstruktionszeit.
  5. Schaltkreise. Die Registry, die Formen und wie ein Familienschaltkreis zusammengesetzt wird.
  6. Die Befehlsfamilien, beginnend mit ADD_SUB_LUI_AUIPC, die jeden ecall und die Anfrageseite jeder Delegation trägt.
  7. Delegations-ABI und Delegationsschaltkreise.
  8. Programm und Identität, Öffentliche Werte, Ausführungs-Trace.
  9. Transkript, SRS, Mercury, Primitive.
  10. Rekursion und Decider, dann der Contract.

Die Vertrauensgrenze#

Die Soundness liegt allein beim Verifier, und der Code des Verifiers ist eine festgelegte Menge von Crates:

Vertrauen nötig für Crates
Verifizieren eines Blocks constants, field, curve, transcript, poly, sumcheck, pcs-verify, pcs, gkr-verify, verifier-core, verifier und constraints, weil die Schaltkreise Teil der Aussage sind und ein fehlendes Gate ein Soundness-Fehler ist
Berechnen einer Identität aus einem ELF loader, isa, program
Der letzte Schritt zur Chain guests/recursion, groth16, der Schaltkreis des Deciders, contracts/ApogeeVerifier.sol
Kein Vertrauen nötig prover, emulator, trace, die beweisende Hälfte von host: Der Prover validiert nichts

Die Annahmen sind die Knowledge Soundness von Mercury und KZG im algebraischen Gruppenmodell unter q-DLOG, Poseidon2 als Random Oracle, die eigenen Annahmen von Groth16 für den letzten Schritt und ein ehrlicher Beitragender zu jeder Zeremonie. Nichts ist in konstanter Laufzeit implementiert, und kein Beweis ist zero-knowledge. Ein Verifier muss die Programmidentität und den SRS-Digest der Zeremonie über einen Kanal beziehen, den der Prover nicht kontrolliert.

Wo zuerst hinsehen#

Dies sind die Eigenschaften, deren Verletzung eine Fälschung wäre, jeweils mit der Stelle, an der sie begründet werden. Die Soundness-Karte behandelt jede Behauptung der Aussage auf dieselbe Weise.

Eigenschaft Begründet in
Jede Challenge wird nach allem gezogen, was sie schützt: die Speicher-Challenges nach jedem M-Commitment, der Fensterliste, io_digest und den 64 Randskalaren; g und β nach den W-Commitments des Shards Beweis §2, §4; Speicher §6.1
Kein Speichertupel und keine Wurzel liest eine W-Spalte, die erst nach den Speicher-Challenges committet wird Speicher §8
Ein Frame erzwingt für seine Masken nur, dass sie boolesch sind; jede Familie legt jede Maske auf m_pc mal die Arten fest, die die Abfrage stellen Speicher §2.1; die Seite jeder Familie
Jeder Schlüssel, den ein Tabellenkanal nachschlägt, ist durch seine Familie beschränkt, und jede über eine Copower geschriebene Schranke trägt zusätzlich eine direkte Schranke Lookup §4, §11; shift-bitwise §3
Ein aus einem Frame-Wort dekodierter Selektor ist one-hot, da sich Codes addieren Delegationsschaltkreise §1
Jede Delegationsanfrage wird mit genau einem Aufruf gepaart Delegation §5
Nur die Exit-Zeile kann HALT_PC schreiben; next_pc wird überall dort, wo eine Familie es berechnet, als gerade erzwungen Speicher §5; jump-branch-slt §5
Ein Anfangswert pro Adresse: die Fensterregeln Speicher §3.5, §9
Die Eingabe- und Journal-Fenster enthalten die Bytes der Aussage; das Journal hat keine Anfangsspalte Öffentliche Werte §5
Die Öffnung entnimmt die Setup-Commitments dem Schlüssel und bindet so die Tabellen und das Image, die die Identität committet Beweis §5; Speicher §6.2
Die Schaltkreise eines Schlüssels sind die der Registry, und sein SRS-Digest wird mit dem der Zeremonie verglichen Beweis §3, §7
Rekursion: durch die Programmidentität gebundene Tapes, die Transkriptkette, Faltungsgewichte, die nach dem gezogen werden, was sie gewichten, die gebundenen Wires des Deciders und die Reihenfolge der Zeremonierunden Rekursion §7, §8, §9

Reproduzieren#

Nichts von dem, was die CI ausführt, braucht eine Zeremoniedatei. Die Suiten, die echte Shards beweisen, laufen über einem eigenen Spielzeug-SRS und brauchen Dutzende GiB; deshalb werden sie namentlich aufgerufen:

sh
cargo test --workspace                                   # every unit, law and row suite
cargo run -p kat-gen && git diff --exit-code             # fixtures regenerate identically
cargo test --release -p prover --test <suite> -- --include-ignored --test-threads=1
#   acceptance, control, alu, mem, fills, block, streaming, keccak, recursion, public_io, revm
cargo test --release -p checker --test tamper -- --include-ignored --test-threads=1   # every tamper twin
cargo run -p checker -- laws <artifact>                  # Laws 1–4 and the lookup rules, independently

Die Implementierung prüfen beschreibt, was jedes Orakel und jede Suite belegt und was keines von ihnen leistet.

Befunde melden#

Melden Sie Befunde an admin@gweb3networks.com, mit dem Abschnitt der Spezifikation oder dem Codepfad, der betroffenen Eigenschaft und nach Möglichkeit einem Manipulations-Zwilling: einem gefälschten Witness, bewiesen, wie ein ehrlicher Prover ihn beweisen würde, der verifiziert wird.

Auditoren

Soundness-Karte

Jede Behauptung, die ein verifizierter Beweis aufstellt, das Argument, das sie trägt, und die genauen Abschnitte der Spezifikation, in denen dieses Argument formuliert und begründet wird.

Als Markdown anzeigen

Ein verifizierter Block begründet einen einzigen Satz: Das Programm dieser Identität, gestartet an seinem Einsprung-pc auf seinem Image, mit dieser öffentlichen Eingabe und irgendwelchen Hilfsdaten (Advice), wird Befehl für Befehl bis zu EXIT mit diesem Status ausgeführt und hat dabei dieses Journal geschrieben. Diese Seite zerlegt diesen Satz in die Behauptungen, aus denen er besteht, und führt jede zu dem Argument, das sie beweist.

Das Programm#

Behauptung Argument Spezifikation
Der Schlüssel beschreibt das registrierte Programm beim Laden eines Schlüssels wird die Identität aus seiner eigenen Konfiguration, seinem Einsprung-pc und seinen Setup-Commitments neu berechnet; der Verifier vergleicht sie mit seiner eigenen Kopie Beweis §7.2, Programm §8
Die Tabellen, die ein Beweis liest, sind die, die die Identität committet die gebündelte Öffnung jedes Shards entnimmt ihre Setup-Commitments dem Schlüssel Beweis §5
Jede ausgeführte Zeile ist der Befehl des Programms an ihrem pc der Decoder-Lookup, mit dem von der Zeile selbst gelesenen pc als Schlüssel, in eine Tabelle, deren aktive Zeilen one-hot und deren Padding-Zeilen −1 sind Lookup §10, Programm §5, §6
Der Speicher beginnt mit dem Image des Programms die Init-Spalte von INIT_TEARDOWN ist eine Setup-Spalte, die von der Identität committet wird; kein aus der Datei stammendes Byte liegt außerhalb von Fenster 0 Speicher §6.2, §3.4
Die Ausführung beginnt am Einsprung-pc das Anfangstupel des pc verwendet den Einsprung-pc des Schlüssels, den die Identität bindet Speicher §4.2, §6.2
Die Schaltkreise sind die richtigen die Schaltkreise eines Schlüssels müssen bei seinen Höhen der Registry des Verifiers entsprechen und die Gesetze, die Speicherregeln und die Erfüllungsregel bestehen Beweis §7.2, Schaltkreise §1, GKR §4.2

Jede Zeile#

Behauptung Argument Spezifikation
Eine Zeile befolgt ihren Befehl die Constraint-Gates der Familie, null auf jeder Zeile, das Soundness-Argument jeder Familie add-sub §4, jump-branch-slt §5, shift-bitwise §5, mul-div §5, memory-ops §3–§6
Eine Zeile stellt genau die Abfragen ihres Befehls jede Maske festgelegt auf m_pc mal die Arten, die diese Abfrage stellen Speicher §2.1
x0 liest und schreibt 0 das x0-Gadget und die Rückschreibungen Speicher §2.4
Register- und RAM-Werte sind Wörter jeder Registerschreibzugriff auf seiner eigenen Zeile beschränkt; jeder RAM-Schreibzugriff einer Ausführungsfamilie ein Wort; Anfangswerte sind Wörter, außer Hilfsdaten, auf die sich keine Familie verlässt memory-ops §5
Eine Padding-Zeile fügt kein Speicherereignis hinzu jede Maske einer Padding-Zeile ist 0, also sind ihre Blätter 1 Speicher §2.3, GKR §4.3

Speicher und Reihenfolge#

Behauptung Argument Spezifikation
Jeder Lesezugriff liefert den letzten Schreibzugriff eine Lese-/Schreib-Multimenge über alle Shards, einmal gegen die Randwerte von Registern und pc abgeglichen Speicher §4, §9
Ein Lesezugriff folgt strikt auf den Schreibzugriff, den er konsumiert der Zeitstempelabstand jeder Abfrage besteht aus zwei TIMESTAMP-Chunks zu je 19 Bit Speicher §2.4, §7
Jede Adresse hat genau einen Anfangswert die Fensterregeln: eine Höhe, disjunkte Fenster, je ein Shard für jedes der festen Fenster Speicher §3.5, §9
Die Multimenge kann sich nicht über eine Schleife schließen Zeitstempel sind ganze Zahlen auf beschränkten Pfaden: Eine Schleife bräuchte mehr als 2^215 Kanten Speicher §4.2
Die Zeilen bilden einen einzigen Pfad vom Einsprung zum Exit, in Programmreihenfolge der pc ist eine Speicherzelle, die mindestens vier Zeitstempel nach ihrem Lesen geschrieben wird Speicher §5, §9
Die Ausführung endet an der Exit-Zeile HALT_PC = 1 ist ungerade; nur die Exit-Zeile schreibt es; JUMP_BRANCH_SLT prüft per Range-Check, dass sein next_pc gerade ist Speicher §5, jump-branch-slt §5, add-sub §4
Das Zeitfenster eines Shards trägt nichts bei Zeitfenster werden nur auf ihre Form geprüft; die Reihenfolge ergibt sich allein aus der Multimenge Beweis §8

Werte und Lookups#

Behauptung Argument Spezifikation
Jedes per Selektor geschaltete Tupel ist eine Zeile seiner Tabelle eine LogUp-Identität pro Kanal, summiert über einen Bruchbaum im GKR-Durchlauf; Zähler an der Wurzel 0, Nenner ungleich null Lookup §1, §6, §8
Selektoren sind boolesch der Selektor jedes Lookups wird durch ein Constraint-Gate der Liste 0 auf s − s² gehalten Lookup §2
Ein Lookup antwortet aus seiner eigenen Teiltabelle eine Breite pro Kanal, disjunkte Schlüsselbereiche mit dem Offset +1 und die Schranke jeder Familie für ihren Schlüssel Lookup §4, §9, §11
Die Tabellen sind die beabsichtigten virtuelle Tabellen sind die geschlossenen Formen des Verifiers; die generische Tabelle ist durch den SRS-Digest gebunden, die dekodierten Tabellen durch die Identität Lookup §3, §12
Jede deklarierte Verpflichtung wird erfüllt die Erfüllungsregel beim Zusammensetzen und bei jedem Laden eines Schlüssels Lookup §11

Öffentliche Werte#

Behauptung Argument Spezifikation
Das Eingabefenster enthielt die Eingabe der Aussage die Anfangsspalte von PUBLIC_INPUT stimmt an einem zufälligen Punkt mit den Wörtern der Eingabe überein, nachdem G7 die Bytes festgelegt hat Öffentliche Werte §5
Das Journal ist das, was die Store-Befehle des Gastprogramms (Guest) hinterlassen haben die Endspalte von PUBLIC_OUTPUT stimmt mit den Wörtern des Journals überein; das Fenster hat keine Anfangsspalte, die ein Prover füllen könnte Öffentliche Werte §5
Der Exit-Status ist der Endwert von x10 die Exit-Zeile schreibt das gelesene a0 zurück; der Verifier gleicht v_10 mit dem Status der Aussage ab add-sub §4, Speicher §4.1, Beweis §6

Delegationen#

Behauptung Argument Spezifikation
Jede Anfrage wird genau einmal ausgeführt der Anker: Anfragen und Aufrufe paaren sich eins zu eins über die Multimenge im eigenen Raum des Typs Delegation §5
Ein Aufruf berechnet seine Funktion das Soundness-Argument jedes Schaltkreises, einschließlich Frame und Kanonizitätsketten Delegationsschaltkreise §2–§7
Eine Operation aus mehreren Aufrufen ist die Komposition ihrer Schritte RAM-Glue: Jeder Schritt liest die Schreibzugriffe des vorherigen Schritts auf einer einzigen Speicherhistorie; die Reihenfolge bestimmt der aufrufende Code Delegationsschaltkreise §1

Das Beweissystem#

Behauptung Argument Spezifikation
Die Ausgaben eines Shards sind sein Schaltkreis, ausgewertet auf seinen committeten Spalten der GKR-Rückwärtsdurchlauf, jede Challenge gezogen nach dem, was sie schützt GKR §5.4
Die behaupteten Spaltenwerte sind die der committeten Polynome eine gebündelte Mercury-Öffnung am Punkt des Durchlaufs Mercury §5, §7
Eine Aussage wird nur durch alle ihre Shards verifiziert Exaktheit der Shard-Menge beim Dekodieren und in verify_block; der Abgleich liest die Wurzeln jedes Shards Beweis §1.3, §6
Challenges folgen auf jedes Commitment, das sie schützen das globale Transkript G1–G11 und das Shard-Transkript S1–S6 Beweis §2, §4, Transkript §3
Das Setup ist das der Zeremonie der SRS-Digest, vom Verifier mit dem der Zeremonie verglichen Beweis §3, SRS §3

Rekursion und der Contract#

Behauptung Argument Spezifikation
Ein Knoten hat genau die Prüfungen des Basis-Verifiers ausgeführt die Prüfungen sind in Tapes im Image des Knotenprogramms kompiliert, das durch die Identität des Knotenprogramms gebunden ist Rekursion §7, §8.1
Der Baum deckt eine einzige Basisaussage ab, jeden Shard, in Reihenfolge die Transkriptkette über die Knoten hinweg, benachbarte Shard-Bereiche und die Prüfungen jedes Knotens an seinen Kindern Rekursion §8.1, §8.2
Jede aufgeschobene Öffnung gilt jede unter Gewichten gefaltet, die nach allem gezogen werden, was sie gewichten, und durch ein einziges Pairing im Contract eingelöst Rekursion §8.3, Mercury §6
Der Decider bindet, was der Contract hält gebundene Wires, committet vor ihrer Challenge; der Schaltkreis bindet das Journal an den gesamten Basisbereich Rekursion §9
Der Decider-Schlüssel hat keine bekannte Falltür eine zweiphasige Zeremonie mit einem ehrlichen Beitragenden pro Runde, Runden in Reihenfolge Rekursion §9, SRS §7

Bewusst nicht behauptet#

  • Irgendetwas über Hilfsdaten. Hilfsdaten sind bewusst an nichts gebunden; ein Gastprogramm prüft sie.
  • Zero-Knowledge. Nichts wird verblindet.
  • Fehlschlagsemantik von sc.w. sc.w gelingt immer; ein Programm, das sich auf sein Fehlschlagen verlässt, liegt außerhalb der Behauptung.
  • Traps. Ein Lauf, der in einen Trap läuft, hat überhaupt keinen Beweis.
  • Dass der Prover korrekt ist. Dem Prover wird nicht vertraut; nur die Crates des Verifiers tragen die Soundness.
  • Dass eine Zeremoniedatei die der Zeremonie ist, oder dass das τ eines Schlüssels unbekannt ist, ohne den eigenen Vergleich des SRS-Digests durch den Verifier.

Auditoren

Die Implementierung prüfen

Wie der Code gegen etwas anderes als sich selbst geprüft wird. Das unabhängige Orakel jeder Schicht, die zweite Implementierung der Schaltkreisregeln, die Manipulations-Zwillinge, die belegen, dass Fälschungen zurückgewiesen werden, und was keine Prüfung abdeckt.

Als Markdown anzeigen

Keine Komponente von Apogee wird gegen eine zweite Implementierung des gesamten Systems geprüft. Stattdessen hat jede Schicht ein eigenes Orakel, so gewählt, dass die Prüfung so wenig Code wie möglich mit dem Geprüften teilt.

Jede Schicht und ihr Orakel#

Schicht Geprüft gegen
Körper, Kurve, Pairing, MSM Known-Answer-Vektoren, erzeugt mit arkworks, das die Tests auch live ausführen; die Tests leiten jede arithmetische Konstante neu her, die die Crates lesen
Poseidon2 und das Transkript tools/transcript-ref: Poseidon2 aus Plonky3 mit den Rundenkonstanten aus zkhash und eine Umsetzung der Spezifikation, die neben dem Duplex-Challenger von Plonky3 läuft und bei jedem Squeeze übereinstimmt
Der Decoder alle 2^30 32-Bit-Wörter mit den niedrigen Bits 11, gegen die aus den Tabellen der ISA abgeleitete Anzahl akzeptierter Wörter und gegen einen unabhängigen Encoder; llvm-objdump über die eingecheckten Gastprogramme (Guests)
RVC-Expansion der eigene Encoder von LLVM, über ein Gastprogramm, das sowohl komprimiert als auch unkomprimiert assembliert wurde
Schaltkreise als Daten checker: die vier Gesetze, die Lookup-Regeln und der Padding-Vertrag, neu implementiert ohne den Code von constraints; gemeinsam ist nur der Gate-Kernel
Die Gates jeder Familie Zeilen-Suiten, die Zeilen mit Rusts eigener Ganzzahlarithmetik bauen und über den Checker auswerten; die arithmetischen Kerne erschöpfend bei reduzierten Wortbreiten
Die Speicher- und Lookup-Argumente native Evaluatoren in checker, angewandt auf ausgeführte Traces
Der Executor die Selbstprüfung seines eigenen Trace und die obigen Argumente; es gibt keinen zweiten Executor
Das revm-Gastprogramm natives revm, gebaut aus ungepatchten Upstream-Crates
Der zustandslose Validator eine eingecheckte Teilmenge von tests-zkevm v21.0.1 nativ in der CI; das gesamte Release nativ und die Teilmenge über das Binary des Gastprogramms von Hand; tools/stateless-ref für die Kodierung der Eingabe
Der Decider der Beweis nativ geprüft und der Contract in revm ausgeführt

Erschöpfende Prüfungen bei kleinen Breiten#

Mehrere arithmetische Kerne sind mit ihrer Wortbreite als Parameter geschrieben, sodass sich die Kodierung über jede Eingabe bei einer Breite prüfen lässt, die klein genug zum Aufzählen ist:

  • das Vergleichs-Gadget bei 6 Bit: Für jedes Operandenpaar, mit und ohne Vorzeichen, findet die Prüfung genau ein (lt, gap), nämlich das der ISA;
  • die Arithmetik von MUL_DIV bei 4 Bit: Für jeden Dividenden und Divisor und jede Divisionsart ist genau ein (q, r) zulässig, nämlich das von RV32M;
  • das Einsetzen (Splice) von MEM_SUBWORD bei einem 4-Bit-Wort: Für jedes Wort, jeden Offset und jede Breite ist genau ein (high, sub, low) zulässig.

Die Schaltkreisregeln, zweimal#

CircuitArtifact::validate und die Konstruktionsregeln für Speicher und Lookups laufen überall dort, wo ein Artefakt gebaut oder ein Schlüssel geladen wird. crates/checker setzt dieselben Regeln ein zweites Mal mit eigenem Code durch, ruft validate nie auf und wertet Gates ausschließlich über gkr_verify::eval_gate aus, den einen Kernel, den beide Seiten als semantische Autorität behandeln. Seine Validatoren prüfen die Gesetze durch Auswertung an Stichprobenpunkten, während validate normalisierte Expansionen vergleicht, berechnen die Summe jedes Kanals Zeile für Zeile statt über einen Baum neu und benennen jedes Tupel, das in keiner Tabellenzeile vorkommt, und bauen die Speicherspalten einer Ausführungsfamilie aus dem Ereignislog statt aus den Zeilen eines Shards neu auf.

sh
cargo run -p checker -- laws <artifact>       # Laws 1–4, then the lookup rules
cargo run -p checker -- padding <artifact>    # the padding contract
cargo run -p checker -- dump <artifact>       # the circuit, readably

Manipulations-Zwillinge#

Ein Manipulations-Zwilling ist eine Fälschung, die genau so bewiesen wird, wie ein ehrlicher Prover sie beweisen würde. Die Manipulations-Suite (checker::TamperHarness) beweist eine Aussage mit veränderten Witness-Zellen oder Randskalaren erneut: Die Multiplizitäten jedes Kanals werden neu gezählt, veränderte Speicherspalten in einer frischen globalen Commit-Phase neu committet, jeder Shard neu bewiesen. Dann verifiziert sie einen Shard oder den Block und prüft per Assertion die Klasse der Zurückweisung, Constraint, Lookup mit seinem Kanal oder MemoryArgument, oder sie stellt sicher, dass eine Änderung, die nichts bricht, verifiziert wird.

Die Zwillinge beruhen darauf, dass der Prover nichts prüft, und das ist so gewollt: Ein gefälschter Witness erhält den besten Beweis, den ein ehrlicher Prover davon erstellen könnte, und der Verifier muss ihn in der erwarteten Klasse zurückweisen. Die Suite enthält auch die Fälschungen des Delegationsankers, und am Mainnet-Mini-Block zeigt sie die andere Seite der Regel für Hilfsdaten (Advice): Eine verfälschte Zelle der Hilfsdaten wird vom Speicherargument zurückgewiesen, eine konsistent verfälschte dagegen wird verifiziert, weil Hilfsdaten an nichts gebunden sind.

sh
cargo test --release -p checker --test tamper -- --include-ignored --test-threads=1

Fixtures, die sich neu erzeugen lassen#

kat-gen schreibt jeden eingecheckten Known-Answer-Vektor, jedes Listing, jedes Schaltkreis-Artefakt und jede Identität, jeweils mit zugehörigem SHA-256, den die einlesenden Tests fest verankern. Die CI erzeugt die Standardgruppen und beide Referenzorakel neu und schlägt bei jeder Abweichung in den Vektorverzeichnissen fehl:

sh
cargo run -p kat-gen && git diff --exit-code

Ein ELF eines Gastprogramms ist nicht maschinenübergreifend reproduzierbar, weil rustc absolute Pfade in die Strings für Panic-Locations einbettet; zwei saubere Builds auf derselben Maschine stimmen überein. Deshalb werden die ELFs der Gastprogramme von Hand auf einer Maschine neu erzeugt, und die CI erzeugt nur neu, was sich von ihnen ableitet.

Was keine Prüfung abdeckt#

  • Es gibt keinen zweiten Executor. Der Emulator wird an einer Neuformulierung seiner eigenen Frame-Tabelle und an den Speicher- und Lookup-Argumenten gemessen, nicht an einer unabhängigen RISC-V-Implementierung, und kein Executor hier nimmt den Software-Fallback eines Delegations-Shims.
  • Die Konstruktionsregeln, die der Checker nicht wiederholt: Die Konstruktionsregeln für den Speicher, die Copower-Regel und die übrigen Konstruktionsregeln von validate, darunter die Gradobergrenze, werden nur einmal durchgesetzt.
  • Der Prover wird nicht auf Korrektheit geprüft, nur auf Vollständigkeit, und zwar über die Suiten, die echte Shards beweisen und außerhalb der CI laufen, weil jede Dutzende GiB braucht.
  • Zustandslose Eingaben der Osaka-Familie haben kein End-to-End-Orakel. Das Release befüllt nur Amsterdam; das Electra/Fulu-Layout wird an eth-act/ere-guests gemessen, die Header-Regeln an zwei Mainnet-Blöcken.

Auditoren/System

Das System von Anfang bis Ende

Normative Spezifikationdocs/architecture.mdAls Markdown anzeigen

Zusammenfassung

Der Überblick der Spezifikation selbst über Apogee VM. Er legt genau fest, was ein verifizierter Beweis begründet und welche drei Werte ein Verifier unabhängig vom Prover besitzt; verfolgt ein Programm von seinem ELF bis zu seinem Beweis; stellt in einer Tabelle das Argument zusammen, das jeden Teil der Aussage trägt; führt die kryptografischen Annahmen sowie die Annahmen zu Setup und Code auf, einschließlich der Crates, auf denen die Soundness beruht; legt jede Grenze von v1.0.0 dar; und nennt das unabhängige Orakel, gegen das jede Schicht des Codes geprüft wird.

Der folgende normative Text wird auf Englisch gepflegt, der kanonischen Sprache der Spezifikation.

Apogee proves executions of RV32IMAC programs. This page is the system end to end: what a proof states, how one is made and checked, what it assumes and where it stops. Each paragraph names the page that specifies its subject; glossary.md indexes the vocabulary.

1 What a proof states#

A verifier holds three things it does not take from the prover's word, and two of them from a channel the prover does not control (proof.md §1, §3):

  • the program identity, one field element: a digest of the program's instruction tables, its initial memory image, its entry pc and its VmConfig — the circuit families it uses and their heights (program.md §8);
  • the SRS digest of the ceremony, which a verifying key must carry;
  • a verifying key: that config, each family's circuit and setup commitments, the SRS's verifier points and the generic lookup table's commitments. Loading it recomputes the identity and the SRS digest from its own contents and holds its circuits to the registry's bytes (proof.md §7).

The proof's statement, PublicInputs, carries the public input bytes, the public output bytes (the journal), the exit status, and the record of the execution's shape — shard counts, memory windows, the final registers and pc, every shard's memory commitments and roots (proof.md §1).

A proof that verifies establishes that the program of that identity, started at its entry pc over its image, with the public input in its input window and some advice of the prover's choosing, executes instruction by instruction to EXIT with that status, having written that journal. Nothing is claimed of the advice, and nothing is hidden: no commitment or proof is blinded.

2 From a binary to a proof#

  1. The program. loader reads the ELF into a ProgramImage, expanding compressed instructions in place; isa decodes; program routes each instruction to one of seven instruction families, builds every family's decoded table — a row per halfword of code — and commits to them as the identity (program.md).
  2. Execution. emulator runs the guest. A cycle is one row of the family that owns its instruction, recording its memory queries: timestamped reads and writes of the pc, registers and RAM (execution-trace.md). A guest issues no system call but EXIT: its input, journal and advice are regions of memory (public-values.md, ecall-abi.md). Hashing and big-integer arithmetic are delegated: an ecall names a frame in RAM, and a row of a delegation family does the work on it (delegation.md, delegation-circuits.md).
  3. Shards. A family's rows are cut into shards of the family's height, a power of two between 2^8 and 2^22. The memory an execution touches is covered by shards of the window families, which give each word its initial and final tuple (memory.md §3). A shard is the unit of proving; a block is hundreds (circuits.md §1).
  4. A shard's proof. Its columns are committed with Mercury (mercury.md). The family's GKR circuit is run backward from its outputs to those columns, a sumcheck a layer (gkr.md), and every column is opened at the one point that pass ends on, in one batched opening (proof.md §5).
  5. The block. A BlockProof is the statement and its shard proofs. verify_block runs the global transcript once, each shard's checks, and once the memory reconciliation over every shard's roots (proof.md §6).
  6. Recursion. Verifier programs, proved by this VM in a format of its own, verify runs of shards and fold their deferred pairings; a tree of them ends in a root, a Groth16 circuit re-verifies the root, and ApogeeVerifier.sol checks that proof and the folded pairing (recursion.md).

The prover executes the guest twice: once to commit every shard's memory columns, which fixes the statement and its challenges, and once to prove each shard as it fills. Its memory is bounded by the shards in flight, not by the execution (streaming.md).

3 How soundness composes#

Each shard's GKR pass and opening tie its circuit's outputs to committed columns. On top of that, these arguments span the execution:

claim carried by
every row obeys its instruction the family circuit's enforcing gates, zero on every row the family pages, circuits.md
a row's instruction is the program's at its pc a lookup of the row's pc and fields in the family's decoded table, which the identity commits lookup.md §10
every read returns the last write one multiset over all shards: an access reads a tuple (space, address, timestamp, value) and writes one with a later timestamp; the verifier multiplies every shard's read and write roots against boundary factors for the registers and the pc memory.md
the rows are one path from the entry pc to the exit, in program order the pc is a cell of that multiset: a row reads its pc and writes the next one at least four timestamps later, so shard order, cycle uniqueness and continuity need no other argument memory.md §5, §9
a value is a byte, a word, a sign, an XOR LogUp channels over range, byte and generic tables lookup.md
the public input and the journal are the claimed bytes the two public windows' initial and final columns, held to the bytes' multilinear extensions public-values.md §5
a delegated computation is the function's invocation rows that read and write the frame through the same multiset, paired one to one with their ecall by an anchor tuple delegation.md §5

Challenges come from a Poseidon2 duplex transcript (transcript.md). The global transcript absorbs the whole statement, every shard's memory commitments included, before the memory challenges exist; each shard's transcript is seeded from its final state (proof.md §2, §4).

4 What it assumes#

  • Cryptography. Mercury's and KZG's knowledge soundness in the algebraic group model under q-DLOG (mercury.md §7); Poseidon2 as a random oracle for Fiat–Shamir; for the last step, Groth16's own assumptions. BN254 gives about 100 bits.
  • Setup. The SRS is the PSE perpetual powers of tau, sound while one contributor was honest. The code checks a file's structure and decodes every point; nothing proves it is that ceremony's, and no proving path runs Srs::validate (srs.md §3). The decider's Groth16 key comes from a second, circuit-specific ceremony (recursion.md §9).
  • What a verifier must obtain itself. The program identity and the ceremony's SRS digest. A key loads under whatever digest its own points give, so a key built over a known τ is refused only by that comparison; the verifier CLI compares identity only, and host::verify neither (proof.md §1, §3).
  • Trusted code. Soundness is the verifier's alone: constants, field, curve, transcript, poly, sumcheck, pcs-verify, pcs, gkr-verify, verifier-core, verifier, and constraints — the circuits are part of the statement, and a missing gate is a soundness bug. Computing an identity from an ELF trusts loader, isa and program. The last step adds guests/recursion, groth16, the decider's circuit and the contract. prover, emulator, trace and the proving half of host are untrusted: the prover validates nothing, and a wrong input costs an honest prover a proof that fails.
  • Nothing is constant-time (primitives.md). No proof is zero-knowledge, so proving keeps nothing secret; the one secret the code handles is a Groth16 ceremony contributor's factor, which groth16::phase2 multiplies in with the same variable-time ladder.

5 Limits#

not zero-knowledge no blinding in Mercury, GKR or the Groth16 decider
advice is unbound a guest checks it against something a proof binds (public-values.md §6)
public values at most 16,380 bytes each of input and journal (public-values.md §9)
sc.w always succeeds the one deviation from RV32IMAC's semantics; there is no reservation state (memory-ops.md §6)
traps are not provable a misaligned access, an access outside mapped memory, ebreak or a pc with no instruction ends an execution with no proof (execution-trace.md §10)
code is static the instruction stream is the image decoded at load; one undecodable word in an executable segment refuses the program (program.md)
code size .text within a decoded table's reach of its load address, 7.94 MiB at 2^22, and the image within bytecode_size_words, 4 MiB by default (program.md §5, §7)
execution length timestamps are 38 bits: 2^36 − 1 cycles (execution-trace.md §1)
delegations are a fixed set six in the base format; an EVM MULMOD with an arbitrary modulus is not one; a delegation proves one step of its function, and composing steps, validating curve points among them, is the calling code's (delegation.md §11, delegation-circuits.md)
prover memory set by the shards in flight: the measured full block peaked at 174 GiB (streaming.md §1)
block witnesses the stateless validator takes its input from an external witness producer; the built-in recorder cannot record every block (ethereum.md §4, §6)
the decider's key one per root shape, and only as trustworthy as its ceremony; the development key is forgeable (recursion.md §9)
on-chain cost about 3.6M gas for the measured block (recursion.md §10)

6 How the code is checked#

No component is checked against a second implementation of the whole system; each layer has its own independent oracle.

layer checked against
fields, curve, pairing, MSM known-answer vectors generated from arkworks, which the tests also run live
Poseidon2 and the transcript tools/transcript-ref: Plonky3 and zkhash
the decoder every 32-bit word of the instruction space against counts from the ISA; llvm-objdump over the committed guests
circuits as data checker: the circuit laws, the lookup rules and the padding contract re-implemented without constraints' code, sharing only the gate kernel (circuits.md §3)
each family's gates row suites that build rows with Rust's own integer arithmetic and evaluate them through the checker; the arithmetic cores exhaustively at reduced word widths; tamper twins, a forged witness proved as an honest prover would and refused in the expected class
the memory and lookup arguments native evaluators in checker over executed traces
the executor its own trace's self-check and the arguments above; there is no second executor, and no executor here takes a delegation shim's software fallback
the revm guest native revm, built from unpatched upstream crates
the stateless validator a committed subset of tests-zkevm v21.0.1 natively in CI; the whole release natively and the subset through the guest binary by hand; tools/stateless-ref for the input encoding
the decider the proof checked natively, and the contract executed in revm

Committed fixtures are regenerated and compared in CI (tools.md §7). The suites that prove real shards, over a toy SRS, need tens of GiB and run outside CI (README).

7 Cost#

recursion.md §10 has the end-to-end measurements for one block, from the base proof to the contract call; streaming.md §1 breaks the base proof down; and circuits.md §1 gives every circuit's width and proof size, which a shard's cost follows.

References#

  • L. Eagen, A. Gabizon. MERCURY: A multilinear polynomial commitment scheme with constant proof size and linear field work. ePrint 2025/385. publication/2025-385.pdf
  • D. Boneh, J. Drake, B. Fisch, A. Gabizon. Efficient polynomial commitment schemes for multiple points and polynomials. ePrint 2020/081. publication/2020-081.pdf
  • J.-L. Beuchat et al. High-speed software implementation of the optimal ate pairing over Barreto–Naehrig curves. ePrint 2010/354. publication/2010-354.pdf

Auditoren/Grundlagen

Primitive: Körper, Kurve, Pairing, Polynome, Sumcheck

Normative Spezifikationdocs/spec/primitives.mdAls Markdown anzeigen

Zusammenfassung

Die Arithmetik, auf der alles andere aufbaut, vollständig im Repository implementiert, ohne Traits, unsafe-Code oder Assembler. Die Seite definiert Fr und seine drei Byte-Formen, den Fq-Turm bis Fq12, die Gruppen G1 und G2 mit ihren unkomprimierten Kodierungen und validierenden Decodern, das optimale Ate-Pairing und seine exakte finale Exponentiation, Pippenger-MSM mit Fenstern, die Indexkonvention multilinearer Polynome und den Zerocheck-Sumcheck, dessen Rundenformat die GKR-Engine wiederverwendet. Nichts davon läuft in konstanter Zeit.

Der folgende normative Text wird auf Englisch gepflegt, der kanonischen Sprache der Spezifikation.

BN254's scalar field Fr, its base field Fq and the tower to Fq12, the groups G1 and G2, the optimal ate pairing, multi-scalar multiplication, multilinear polynomials and the zerocheck. The byte encodings of field elements and points (§1–§3) and the polynomial index convention (§6) are defined here.

  • All of it is this repository's code: concrete types, no field trait, no unsafe, assembly or intrinsics. field, poly and sumcheck are #![no_std] and build for the guest target; curve is std, with rayon, and no guest links it.
  • arkworks is a test oracle only, for field, curve and poly, live and through vectors tools/kat-gen generates; the tests of field and curve re-derive every arithmetic constant those crates read.
  • Nothing is constant-time: reductions, exponentiations, point additions and scalar ladders branch on their operands. No proof is zero-knowledge, so no witness is secret; the one secret this code handles, a decider ceremony contributor's factor (recursion.md §9), goes through the same variable-time ladder (§3).

1 Fr#

p = 21888242871839275222246405745257275088548364400416034343698204186575808495617

field::Fr is the integers mod p (constants::FR_MODULUS), 254 bits; p is also the order of G1 and G2, r in §3–§4. In memory an element is four little-endian 64-bit limbs of x·R mod p, R = 2^256 mod p, always reduced below p, so equal limbs are equal values. Multiplication is CIOS Montgomery over u128 intermediates. Fr::inverse is x^(p−2), None at 0; field::batch_inverse is Montgomery's trick and leaves a 0 entry 0. p − 1 = 2^28·c with c odd, and every FFT domain is a subgroup of the one constants::FR_TWO_ADIC_ROOT_OF_UNITY generates.

byte form used in
wire the value, not x·R, as 32 little-endian bytes: Fr::to_bytes. Fr::from_bytes is None for a value ≥ p and never reduces; serde goes through both every proof, key and artifact
source literal 0x and exactly 64 lowercase hex digits, big-endian: Fr::from_hex, None for any other spelling or a value ≥ p constants in crates/constants
memory the four limbs: Fr::to_memory_bytes, Fr::from_memory_bytes, None at or above p the FR_ARITH delegation's frame alone

On the guest target (cfg(target_arch = "riscv32")) addition, Montgomery multiplication and inversion call the FR_ARITH delegation through guest_sdk::recursion::fr_arith (delegation.md §10). Its circuit proves these three functions of the memory form the frame carries (delegation-circuits.md §4), so a delegated result is the software result bit for bit.

2 The Fq tower#

q    = 21888242871839275222246405745257275088696311157297823662689037894645226208583
Fq2  = Fq[u]/(u^2 + 1)
Fq6  = Fq2[v]/(v^3 − ξ)       ξ = 9 + u
Fq12 = Fq6[w]/(w^2 − v)

curve::Fq is the field of coordinates (constants::FQ_MODULUS). Its limb arithmetic is Fr's, copied literally over q's constants, and so are its wire and source-literal forms. Fq2 encodes as c0 ‖ c1; nothing above it has a byte form.

Products are schoolbook and squarings above Fq2 are products. A Frobenius map multiplies coefficients by powers of ξ tabulated in constants (FQ6_FROBENIUS_C1, FQ6_FROBENIUS_C2, FQ12_FROBENIUS_C1); Fq12::conjugate is the q^6 one. Nothing in the tower or the pairing is sparse, cyclotomic or precomputed: a pairing is only ever computed to verify something, and the code is written to be read.

3 G1, G2 and their encodings#

G1 = E(Fq)                E:   y^2 = x^3 + 3        #E  = r             generator (1, 2)
G2 ⊂ E′(Fq2), order r     E′:  y^2 = x^3 + 3/ξ      #E′ = r·(2q − r)    generator EIP-197's

G1Affine { x, y, infinity } is a point and G1Projective its Jacobian form, Z = 0 the identity, under the EFD formulas dbl-2009-l, add-2007-bl and madd-2007-bl, with the identity, P = Q and P = −Q branched on explicitly. Scalar multiplication is a fixed 4-bit window. G2Affine and G2Projective are the same code over Fq2.

A point's wire form is uncompressed affine, and there is no compressed one:

G1Affine    64 bytes    x ‖ y
G2Affine   128 bytes    x.c0 ‖ x.c1 ‖ y.c0 ‖ y.c1        x = x.c0 + x.c1·u
infinity                every byte zero

Each coordinate is an Fq in wire form; (0, 0) is on neither curve, so zero is unambiguous. G1Affine::from_bytes and G2Affine::from_bytes return None unless the bytes are all zero, or every coordinate is below q, the point satisfies its curve's equation and, in G2, whose cofactor 2q − r is not 1, [r]P is the identity — by the window ladder, with no endomorphism.

Nothing else validates: the affine structs' fields are public, and the group law, msm and the pairing compute on whatever they are given. A point's transcript form is transcript.md §4's; a .ptau file (srs.md §2) and the contract (recursion.md §9) have encodings of their own.

4 The pairing#

e(P, Q) = f_{6x+2, Q}(P)^((q^12 − 1)/r)        x = 4965661367192848881

curve::pairing::miller_loop(pairs) is Algorithm 1 of Beuchat et al. (ePrint 2010/354) with homogeneous projective line formulas: the 66-digit NAF of 6x + 2 (constants::ATE_LOOP_NAF), then the two lines adding ψ(Q) and −ψ(ψ(Q)), ψ the untwist-Frobenius-twist map. For a Q of order r no step adds a point to itself, to its negative or to the identity, so the line formulas have no exceptional case.

final_exponentiation returns exactly f^((q^12 − 1)/r): the easy part (q^6 − 1)(q^2 + 1), then the hard exponent (q^4 − q^2 + 1)/r as its base-q expansion λ0 + λ1·q + λ2·q^2 + q^3, by three exponentiations (constants::FINAL_EXP_LAMBDA_0 to FINAL_EXP_LAMBDA_2, the two negative ones conjugated) and three Frobenius maps. The Fuentes-Castañeda hard part, which arkworks uses, returns this value raised to 2x(6x^2 + 3x + 1): the two libraries agree on every pairing check and on no pairing value but 1, and the test vectors are arkworks' Miller outputs raised to the literal exponent.

pairing_check(pairs) is Π e(P_i, Q_i) = 1 by one Miller loop, whose Fq12 squarings the pairs share, and one final exponentiation: the form of every pairing equation in the system. A pair holding a point at infinity contributes 1 and is skipped; an empty product is 1.

5 MSM#

curve::msm::msm(bases, scalars) is Σ scalars_i·bases_i in G1 by windowed Pippenger, and msm_small_u32 the same sum over u32 scalars, recoded from 32 bits instead of 254. Neither looks at a scalar's size: the caller chooses, and pcs::commit chooses by a column's backing (§6). G2 has no MSM here; crates/groth16 carries its own.

  • Width. w = 3 below 32 points, otherwise ⌊0.69·⌈log2 n⌉⌋ + 2: arkworks' rule.
  • Digits. A scalar is recoded into signed digits in [−2^(w−1), 2^(w−1)], one a window, over ⌈(bits + 1)/w⌉ windows; the sign costs a negated base and halves the buckets to 2^(w−1). At 2^20 points w = 15: 17 windows for an Fr, 3 for a u32.
  • Parallelism. Each (window, chunk of the input) is a rayon task that adds bases into buckets by mixed addition and reduces them by a running sum; the tasks' sums are combined serially, w doublings a window. Group sums are exact, so the point does not depend on the thread count.

6 Multilinear polynomials#

poly::MultilinearPoly is a table of 2^n evaluations over {0,1}^n, the type of every column.

Index convention. Variable j is bit j of the index: the evaluation at y = (y_0, …, y_{n−1}) is entry Σ_j y_j·2^j. bind(r) fixes variable 0, the low bit,

f′(i) = f(2i) + r·(f(2i + 1) − f(2i))

and the old variable 1 becomes variable 0. So binding r_0, r_1, … in order leaves evaluate(&[r_0, r_1, …]), whose point[j] is variable j. Sumcheck round i binds variable i (§7), so a claim's point lists its challenges in variable order, the order evaluate and a Mercury opening (mercury.md §1) take.

Backing. PolyBacking holds the table as a bitset (U1), u8, u16, u32 or Fr. A trace column is filled and committed at its integer width (§5). get and evaluate embed an entry in Fr as they read it and leave the table alone; the first bind folds the integer table straight into an Fr table of half the length, and the backing is Fr from then on.

eq. eq_table(r) tabulates eq(r, ·) over the cube in the same index order; eq_eval(r, y) is its closed form, for any r and y.

new, get, bind, evaluate and eq_eval panic on a table, index or point of the wrong size rather than return an error.

7 The sumcheck#

crates/sumcheck proves that a gate vanishes on the cube. A sumcheck::Gate is a sum of GateTerms coef·x_a·x_b, the second factor optional, over input columns of n variables: degree at most 2 in each variable, by construction.

G(y) = 0 on all of {0,1}^n is proved as the sumcheck 0 = Σ_y eq(r, y)·G(y) at a random r. eq·G has degree at most 3 in each variable, so a round polynomial is a cubic and a round message its four coefficients [c0, c1, c2, c3], ascending — four whatever the gate, so a proof's shape depends on n and the number of inputs alone.

prove_zerocheck and verify_zerocheck run one schedule, under the tags of transcript.md §5:

1   the caller binds the columns to the transcript
2   r_0 … r_{n−1}                                      n × SUMCHECK_CHALLENGE
3   for i in 0..n:   g_i, one message of four          SUMCHECK_ROUND
                     ρ_i, binding variable i           SUMCHECK_CHALLENGE
4   final_evals: each input column at ρ, one message   SUMCHECK_FINAL_EVALS

The verifier checks the proof's shape, then g_0(0) + g_0(1) = 0 and g_i(0) + g_i(1) = g_{i−1}(ρ_{i−1}), each before absorbing g_i, and, with final_evals absorbed, g_{n−1}(ρ_{n−1}) = eq(r, ρ)·G(final_evals). It returns SumcheckClaim { point: ρ, final_evals } or a SumcheckError, and does not panic on a proof.

That last check is one equation over all the claimed evaluations and ties none of them to its column: the caller owes an opening of each at ρ, as it owes step 1. The step 1 its callers use is witness_digest, a hash and not a commitment: a sponge of its own absorbs [column count, n] and each column's cells under WITNESS_DIGEST, and its raw squeeze enters the transcript under the same tag.

What uses it. No proof in the system is this zerocheck, and no circuit is made of its Gate (gkr.md §3). The proving stack takes one type from the crate, SumcheckProof { rounds: Vec<[Fr; 4]>, final_evals }, as each layer of a gkr_verify::GkrProof. The GKR layer sumcheck (gkr.md §5) repeats step 3's rounds and checks under the same two tags, from a batched claim instead of 0 and to a final check of its own, in gkr::prove_sumcheck and gkr_verify::verify_sumcheck. prove_zerocheck and verify_zerocheck are called only by tests and tools/bench.

Auditoren/Grundlagen

Das Transkript

Normative Spezifikationdocs/spec/transcript.mdAls Markdown anzeigen

Zusammenfassung

Wie jede Challenge im Protokoll gezogen wird. Die Seite spezifiziert die Poseidon2-Permutation der Breite 3 über Fr mit ihren festgelegten Rundenkonstanten, den Duplex-Sponge (Rate 2, Kapazität 1) mit seinen Regeln für Absorb und Squeeze, die typisierte Rahmung von Nachrichten, die jeden absorbierten Strom injektiv macht, wie ein G1-Punkt als vier Limbs absorbiert wird, und die Tabelle aller 45 Transkript-Tags mit Art und Stelle jedes einzelnen.

Der folgende normative Text wird auf Englisch gepflegt, der kanonischen Sprache der Spezifikation.

Every challenge in the protocol is drawn from a Poseidon2 duplex sponge over Fr through a typed message layer. This page specifies the permutation, the sponge, the framing, a G1 point's transcript form and every tag. Implementation: crates/transcript, #![no_std].

1 The Poseidon2 permutation#

Width 3 over Fr, S-box x^5, 4 full rounds, 56 partial rounds (S-box on lane 0 only), 4 full rounds. The round constants are RC3 of HorizenLabs/poseidon2, plain_implementations/src/poseidon2/poseidon2_instance_bn256.rs at commit 055bde3f4782731ba5f5ce5888a440a94327eaf3.

E(s) = s + (s₀+s₁+s₂)·(1,1,1)                 circ(2, 1, 1)
I(s) = s + (s₀+s₁+s₂)·(1,1,1) + (0,0,s₂)      1 + diag(1, 1, 2)

poseidon2_permute(s):
  s ← E(s)
  RC3 rows 0–3:    s_i ← (s_i + c_i)^5, every lane;   s ← E(s)
  RC3 rows 4–59:   s₀ ← (s₀ + c₀)^5;                  s ← I(s)
  RC3 rows 60–63:  s_i ← (s_i + c_i)^5, every lane;   s ← E(s)

poseidon2_permute([0, 1, 2])₀ = 0x0bb61d24daca55eebcb1929a82650f328134334da98ea4f847f760054f4a3033

constants::POSEIDON2_RC3_INITIAL, _INTERNAL and _TERMINAL hold the 80 entries read (upstream's partial rows are zero in lanes 1 and 2) as upstream's big-endian hex literals, character for character, decoded by Fr::from_hex on every call. They are pinned through the permutation, by the oracle's 128 vectors (§2), each of which reads every constant.

On riscv32, poseidon2_permute is one POSEIDON2 delegation call over the lanes' canonical bytes, falling back to these rounds when the executor answers -ENOSYS (delegation.md §10).

2 The duplex sponge#

state  [Fr; 3]   lanes 0, 1 the rate, lane 2 the capacity; zero in Transcript::new()
input  [Fr; 2]   absorbed, not yet permuted: 0 or 1 pending between operations
output [Fr; 2]   squeezed, not yet handed out: 0 to 2

observe(x):  output ← []; input.push(x); if |input| = 2: duplex()
sample():    if |input| > 0 or |output| = 0: duplex(); return output.pop()
duplex():    n ← |input|; state[0..n] ← input; input ← []
             if n > 0: state[n..2] ← 0; state[2] += n
             poseidon2_permute(state); output ← [state[0], state[1]]
  • Absorption overwrites the rate. A short absorb zero-fills the rest of it and adds its length to the capacity, so [a] and [a, 0] differ; with nothing pending, a duplex is a pure squeeze and does neither.
  • Squeezed lanes leave from the end: the first sample after an absorb is state[1], the second state[0], and a third permutes again.
  • observe drops unread output and lanes past a buffer's length stay zero, which moves no challenge and makes the state a function of the operation sequence alone.

This is Plonky3's DuplexChallenger at width 3 and rate 2. The vectors crates/transcript is tested against come from tools/transcript-ref, which shares no code with it: Plonky3's Poseidon2 keyed with zkhash's own RC3, and a transcription of this section and §3 run beside that type, agreeing with it on every squeeze. The recursion format replays the same sponge over field cells, one P2_FIELD row a duplex step (recursion.md §4).

3 Typed messages#

append_scalars(tag, xs):  observe(tag); observe(|xs|); observe(x) for x in xs
append_scalar(tag, x)  =  append_scalars(tag, [x])
append_bytes(tag, b):     observe(tag); observe(|b|); observe(c) for each 31-byte chunk c of b,
                          zero-padded to 32 bytes, read little-endian
challenge_scalar(tag):    observe(tag); return sample()

The length, the scalar count or for bytes the byte count, delimits a message: "abc" and "abc\0" are each one chunk, below 2^248 < p, and differ. A challenge absorbs its tag, so it always comes from a fresh permutation.

The framing carries no kind, so each tag names exactly one of scalars, bytes or a challenge (§5): a tag of two kinds would make append_bytes(T, b"") and append_scalars(T, []) the same T, 0. So every digest — program identity, the SRS digest, transcript::io_digest, sumcheck::witness_digest, pcs::accumulator_digest — is a fresh sponge of typed messages ended by a raw sample(), never by a challenge under one of its message tags.

snapshot() captures the state and both buffers, and Transcript::restore resumes the same challenge stream. Its postcard form is 226 bytes, state[3], input[2], input_len: u8, output[2], output_len: u8, each Fr canonical; decoding refuses input_len ≥ 2, output_len > 2 and a nonzero lane past either length. The archived path's phase files hold the global transcript, and each shard's after its GKR pass, in this form (streaming.md §6). A shard transcript is no restored global sponge but a fresh one whose first message carries the global state digest (proof.md §4).

Each typed operation appends Absorb { tag, n_scalars } (payload elements: scalars, or chunks) or Challenge { tag } to event_log(). Raw observe and sample are not logged, the log never feeds the sponge and a snapshot omits it; checker::tape holds the global transcript's log to the order G1–G11 (tools.md §4).

4 G1 points#

A point is absorbed as four Fr limbs of its 64-byte encoding x ‖ y (primitives.md §3), with no curve arithmetic (transcript::g1_limbs):

[ x[0..16], x[16..32], y[0..16], y[16..32] ]   each half read little-endian, below 2^128 < p
[ S, S, S, S ]                                 the 64 zero bytes of infinity; S = 2^128

A coordinate is an Fq element and q > p, hence the halves. S is constants::G1_INFINITY_SENTINEL: no 16-byte half reaches 2^128, so the limbs determine the 64 bytes whether or not they encode a point on the curve. The absorber never refuses; a point is validated where it is decoded, before a pairing reads it.

transcript::append_g1_points(tr, tag, points) absorbs k points as one message of 4k limbs, never k messages, so the framed length binds k. pcs::append_g1_list is it over G1Affine::to_bytes, and pcs::append_g1 a list of one.

5 Tags#

Tag = u64: constants::transcript_tags, 45 tags numbered from 1 and named by transcript_tags::NAMES[tag − 1]; 0 is not a tag. Kinds: S scalars, B bytes, C challenge. Where: G1–G11 and the shard transcript are proof.md §2, §4, the SRS digest §3 there; identity program.md §8; Mercury mercury.md; GKR gkr.md §5; io_digest public-values.md §5; stacks and nodes recursion.md §1.3, §8.3. † marks a tag on no proof path.

tag where
1 PROTOCOL_SUITE S G1: [PROTOCOL_VERSION]
2 PUBLIC_INPUTS B G7: io_digest's 32 canonical bytes
3 COMMITMENT S a commitment list: identity, G8, shard witness, Mercury
4 SUMCHECK_ROUND S a sumcheck round's coefficients
5 SUMCHECK_CHALLENGE C a round's challenge; first, a zerocheck's eq-randomizers
6 EVALUATION_CLAIM S Mercury: the point, then the claimed values
7 PCS_OPENING S Mercury: proof points and evaluations
8 WITNESS_DIGEST S sumcheck::witness_digest's sponge, and its result †
9 SUMCHECK_FINAL_EVALS S the zerocheck's final evaluations †
10 MERCURY_INSTANCE S Mercury: [n]
11 MERCURY_ALPHA C Mercury: α
12 MERCURY_GAMMA C Mercury: γ
13 MERCURY_Z C Mercury: z, redrawn while 0
14 BDFG_BATCH C Mercury: δ
15 BDFG_POINT C Mercury: z′
16 PAIRING_MERGE C Mercury: the pairing merge ρ
17 MERCURY_BATCH C Mercury: the column batch ρ
18 ACCUMULATOR_DIGEST S pcs::discharge: the entry words' sponge, and its result †
19 ACCUMULATOR_MERGE C pcs::discharge: the per-check weight †
20 PUBLIC_INPUT_STREAM B io_digest: the input
21 PUBLIC_OUTPUT_STREAM B io_digest: the output
22 PROGRAM_IDENTITY S identity: [code_version]; G6: [identity]
23 VM_CONFIG S identity; G3
24 SHARD_COUNTS S G4
25 GKR_OUTPUTS S GKR: the output tables
26 GKR_OUTPUT_POINT C GKR: the top point
27 GKR_BATCH C GKR: a transition's claim batch
28 GKR_LAYER_CLAIMS S GKR: a transition's claimed values
29 GKR_CHILD C GKR: a halving transition's line point
30 MEMORY_WINDOWS S G5
31 MEMORY_BOUNDARY S G9
32 PROGRAM_ENTRY S identity: [entry_pc]
33 LOOKUP_CHALLENGE C shard: g, then β (lookup.md §2)
34 SRS_DIGEST S G2
35 SRS_VERIFIER B the SRS digest: the 320-byte SrsVerifier
36 MEMORY_GROUP S G8: [family, shard count]
37 MEMORY_CHALLENGE C G10, four times
38 GLOBAL_STATE_DIGEST C G11
39 SHARD_SEED S shard: [digest, family, index]
40 SHARD_TS_WINDOW S shard: [start, end]
41 GENERIC_TABLE S the SRS digest: the generic table's 3 points, 12 limbs
42 STACK_CHALLENGE C a recursion-format shard: its σ stack challenges
43 FOLD_STATE S a node: a verified shard's final transcript state
44 FOLD_WEIGHT C a node: a shard's w, w′, or a child's weight
45 FOLD_CHILD S a node: a child's journal

Auditoren/Grundlagen

Der strukturierte Referenzstring

Normative Spezifikationdocs/spec/srs.mdAls Markdown anzeigen

Zusammenfassung

Das Setup des gesamten Systems. Die Seite nennt die Zeremonie, die Perpetual Powers of Tau der PSE, Beitrag 80, gibt deren [τ]_1 zur Identifikation an, spezifiziert, wie eine .ptau-Datei eingelesen wird und was das Dekodieren beweist, legt dar, was vorausgesetzt statt geprüft wird, und definiert das SRS-Archiv, die drei Punkte, die ein Verifier besitzt, KZG über den Potenzen und die Lagrange-Basen, aus denen der Groth16-Schlüssel des Deciders gebildet wird.

Der folgende normative Text wird auf Englisch gepflegt, der kanonischen Sprache der Spezifikation.

The powers of τ every commitment is made under: the ceremony they come from, how its file is read and what is checked, the archive an SRS is cached in, the three points a verifier holds, KZG over them, and the Groth16 first phase read from the same file. Implementation: crates/srs.

1 The ceremony#

The SRS is PSE's perpetual powers of tau, contribution 80: files ppot_0080_<p>.ptau, kept in assets/ptau/, which is gitignored. Hermez's powersOfTau28_hez_final_*.ptau is another ceremony with another τ; the reader ingests it as readily, and every commitment, key and identity over it differs. This ceremony's [τ]_1, as the hex of its canonical encoding x ‖ y:

9bbb31bedc304e081e2aada4b56c2217e0e94ee16874e3517d14bef5dcec3a16
317ff1589e53513fa333591b318e8f1e55ef7c37d92beb6d9a61d770a2f39506

PSE's files are cut from one ceremony: the first 2^k powers, and the Lagrange bases of domains up to 2^k, agree in every file of power k or more. One file, ppot_0080_24.ptau (19.3 GB), serves every use. A base key needs as many powers as its tallest family has rows, at most 2^22, the menu's top, and at least the generic table's 2^18 (lookup.md §9); bench prove reads 2^22. The recursion format reads 2^24, its largest stack (verifier_core::STACK_LOG, recursion.md §1.3), and the decider its domain's Lagrange bases (§7).

2 Ingesting a .ptau file#

Srs::from_ptau(path, k) reads snarkjs's .ptau container, the one ingestion format. Integers are little-endian.

0    4    "ptau"
4    4    version: 1
8    4    section count: at most 64
12   ..   sections: id u32 | size u64 | payload

id 1   header, 44 bytes: n8 = 32 | q (n8 bytes) = BN254's Fq modulus | power p | ceremonyPower
id 2   tauG1: 2^(p+1) − 1 G1 points, [τ^0]_1 first
id 3   tauG2: 2^p G2 points, [1]_2 then [τ]_2

Sections 1–3 must each occur once, at the sizes p implies; the others (alpha, beta, the contribution record, the Lagrange bases of §7) are not read here. from_ptau takes the first 2^k points of section 2 (k ≤ p) and the first two of section 3.

A point is uncompressed affine in little-endian Montgomery form: each 32-byte coordinate holds coord·R mod q, R = 2^256; G1 is x ‖ y, G2 x.c0 ‖ x.c1 ‖ y.c0 ‖ y.c1. It is the only non-canonical point encoding the code reads. A coordinate is read as a canonical Fq (refused at or above q), multiplied by R^−1 and re-encoded, and the canonical bytes go through G1Affine::from_bytes or G2Affine::from_bytes (primitives.md §3), the one validating decoder. All-zero bytes are infinity in both forms.

from_ptau never panics: every refusal is an SrsError — Io, Truncated, BadMagic, BadVersion, BadSection (over 64 sections, sections 1–3 not each present once, a header not 44 bytes, a power outside 1..=30, a section size p does not imply), WrongCurve, PowerTooLarge (k > p), and InvalidPoint { index }, a failing point but not necessarily the first.

3 What is validated, and what is presumed#

Decoding proves every point canonical, on its curve and in the order-r subgroup. Srs::validate adds that they are powers of one τ:

g1[0] = G1 generator      g2_gen = G2 generator      no point is infinity
e(Σ_i c_i·g1[i], g2_tau) = e(Σ_i c_i·g1[i+1], g2_gen)       i < n − 1

with each c_i 31 bytes from /dev/urandom, so that no file can be built to pass: a power that is not τ times the one before survives with probability at most 2^−248. The infinity check excludes τ = 0: pairing_check skips a pair at infinity (primitives.md §4), so such an SRS would pass vacuously and kzg_verify over it accept any opening. validate identifies nothing, and no proving path runs it.

Soundness needs nobody to know τ, which this code presumes of the ceremony. A statement binds the SRS only through the SRS digest (proof.md §3), which covers the SrsVerifier and the generic table's three commitments, not the powers, which only a prover reads. A key's loader recomputes the digest from the key's own points, so a key whose SrsVerifier has a known τ loads under its own digest: a verifier takes the ceremony's digest from a channel the prover does not control, or recomputes it from the ceremony. Program identity covers neither the SrsVerifier nor the table; in the recursion tree the digest is a constant of both programs' images, which their identities bind (recursion.md §8.1).

4 The SRS archive#

Srs::save and Srs::load keep an ingested SRS in a file of their own, integers little-endian and points canonical (primitives.md §3); bench recurse caches its 2^24 powers in one.

0     8          "APOGESRS"
8     4          version: 1
12    4          power k, at most 30
16    8          G1 count: 2^k
24    128        g2_gen
152   128        g2_tau
280   64·2^k     g1, [τ^0]_1 first

load requires exactly 280 + 64·2^k bytes before reading a point (the cap on k keeps the product from wrapping) and decodes every point through from_bytes, which catches a corrupted coordinate, not a substituted archive. It refuses with Truncated, BadMagic, BadVersion, BadSection and InvalidPoint.

5 SrsVerifier#

The only SRS material a verifier takes: g1_gen = [1]_1, g2_gen = [1]_2 and g2_tau = [τ]_2, what Mercury's pairings read. A verifier never commits; the generic table's commitments reach it as given points. The wire form is 320 bytes, g1_gen ‖ g2_gen ‖ g2_tau, canonical, unframed — its postcard form, VerifyingKey's srs_verifier and verifier::encode_srs_verifier alike — and every reader decodes it through the validating from_bytes.

6 KZG#

srs::kzg, over coefficients little-endian in the degree (coeffs[i] multiplies X^i, as g1[i] is [τ^i]_1):

kzg_commit(f)   = Σ_i f_i·[τ^i]_1                               one MSM
kzg_open(f, z)  = (f(z), [q(τ)]_1), q = (f − f(z))/(X − z)      one Horner pass gives both
kzg_verify(cm, z, v, w):  e(cm − v·[1]_1 + z·w, [1]_2) · e(−w, [τ]_2) = 1

More coefficients than powers is an error, never a truncation. The zero polynomial commits to infinity and opens to (0, infinity), which verifies. A Mercury commitment is exactly kzg_commit of the evaluation table read as coefficients (mercury.md §2). Mercury calls neither kzg_open nor kzg_verify, but its pairing relations take their shape, e(A, [1]_2) = e(B, [τ]_2) with both G2 arguments SRS constants, which is what lets recursion fold them instead of pairing (recursion.md §8.3).

7 Phase 1#

srs::Phase1::from_ptau(path, m), for m ≤ p and m ≤ 28, reads what a Groth16 key takes from the ceremony at a domain of n = 2^m: tau_g1, [τ^i]_1 for i < 2n − 1, from section 2; and lagrange_g1 and lagrange_g2, [L_j(τ)] in each group, L_j the Lagrange polynomial at ω^j and ω of order n squared down from constants::FR_TWO_ADIC_ROOT_OF_UNITY, from sections 12 and 13, which hold the bases of domains 1, 2, 4, … in turn, domain n from point n − 1. It refuses a basis that is not this domain's: tau_g1[0] and each basis's sum must be the generator, and Σ_j ω^j·[L_j(τ)]_1 = [τ]_1. The G2 basis is held to the curve, not the subgroup. The decider's key is made over it (recursion.md §9).

Auditoren/Grundlagen

Mercury

Normative Spezifikationdocs/spec/mercury.mdAls Markdown anzeigen

Zusammenfassung

Das polynomielle Commitment-Verfahren. Die Seite legt fest, was die Arbeiten zu Mercury und BDFG20 offenlassen: die Aufteilung der Variablen, das Commitment als KZG-Commitment der Auswertungstabelle, das vollständige Öffnungsprotokoll mit seinem Transkriptablauf in sechzehn Schritten, den 704-Byte-Beweis und die zusammengeführte Pairing-Prüfung des Verifiers mit zwei Paaren. Hinzu kommen ein Batch vieler Spalten an einem Punkt, den jeder Shard verwendet, und eine aufgeschobene Form aus zwölf Akkumulator-Einträgen, die der Rekursionsbaum faltet, statt Pairings zu berechnen. Kosten und Sicherheit schließen die Seite ab.

Der folgende normative Text wird auf Englisch gepflegt, der kanonischen Sprache der Spezifikation.

Every committed column is opened with Mercury (Eagen and Gabizon, ePrint 2025/385), finished by the batched KZG opening of BDFG20 (Boneh, Drake, Fisch and Gabizon, ePrint 2020/081). This page pins what the papers leave open, and adds a batch of k columns at one point and the deferred form the recursion tree folds. crates/pcs is the prover, the curve side and the pairings; crates/pcs-verify, no_std, is the verifier's field side, which the recursion guest links.

1 Parameters and the variable split#

n = 2^{2t} evaluations with 1 ≤ t ≤ 27, b = 2^t = √n, s = 2t variables; u ∈ Fr^s is the opening point and v the claimed value. pcs_verify::check_num_vars refuses every other variable count (PcsError::UnsupportedNumVars) and never pads, which is why every trace height is an even power of two (program.md §7). The ceiling, pcs_verify::MAX_NUM_VARS = 54, is where Fr's 2-adicity of 28 runs out of the 2b-th roots of unity §3.1 needs, and it keeps 2^{|u|} in range for a u the verifier is handed.

The evaluation table is read as coefficients, and variable m is bit m of an index (primitives.md §6). Write an index i + j·b with i the low t bits, as Mercury §3.1 does; its evaluation is the coefficient of X^{i+j·b}. The point splits the same way: u1 is its first half, u_0..u_{t−1}, and pairs with i; u2 is u_t..u_{2t−1} and pairs with j.

f(X) = Σ_{i<b} X^i·f_i(X^b),    f_i(X) = Σ_{j<b} f_{i+j·b}·X^j
f̂(u) = Σ_{i,j<b} eq(i, u1)·eq(j, u2)·f_{i+j·b}

pcs::open returns what poly::MultilinearPoly::evaluate gives at u, and a verifier handed the two halves swapped rejects.

2 Commitment#

pcs::commit returns [f(x)]_1 for §1's f(X), an MSM over the first n SRS powers: exactly the KZG commitment of the evaluation table read as coefficients (srs::kzg::kzg_commit), with no second scheme behind it. It refuses an SRS of fewer than n powers (SrsTooSmall). A column backed by U1, U8, U16 or U32 (poly::PolyBacking) is widened to u32 and committed through curve::msm::msm_small_u32, never lifted to Fr; an Fr backing goes through curve::msm::msm.

The map from a table to its commitment is Fr-linear, which §5 uses, and a zero coefficient adds nothing: a column extended by zero rows keeps its commitment. So the generic table's commitments serve every height that holds the table (lookup.md §9), and pcs::commit_stack commits a recursion stack without building it.

3 The opening protocol#

3.1 The polynomials#

definition coefficients sent as
h Σ_i eq(i, u1)·f_i(X); its X^j coefficient is f̂(u1, j) b h
q, g f = (X^b − α)·q + g, so g = Σ_i f_i(α)·X^i n − b, b q, g
S the symmetrized witness below b − 1 s
D X^{b−1}·g(1/X): g reversed b d
H (f − (z^b − α)·q − g_z)/(X − z) n − 1 pi_z
W, W′ §3.3 b − 1 each w, w_prime

P_u(X) = Σ_{i<b} eq(i, u)·X^i = Π_{m<t}(u_m·X^{2^m} + 1 − u_m), so ⟨P_u, g⟩ = ĝ(u) for g of fewer than b coefficients (Mercury §4.2). The prover uses its coefficients, poly::eq_table(u); the verifier evaluates the product in O(t).

The fold (Mercury §5) divides every f_i by X − α, b Horner divisions advanced together in one pass over the rows, with no transform. Then ĝ(u1) = h(α) and ĥ(u2) = f̂(u) = v, and one S proves both inner products (Mercury §4.1), the left side's constant coefficient being 2·(⟨g, P_u1⟩ + γ·⟨h, P_u2⟩):

g(X)·P_u1(1/X) + g(1/X)·P_u1(X) + γ·(h(X)·P_u2(1/X) + h(1/X)·P_u2(X))
    = 2·(h(α) + γ·v) + X·S(X) + S(1/X)/X

S is coefficients b..2b−2 of X^{b−1} times the left side, computed with four forward transforms of size 2b and one inverse; no transform in an opening is larger (crates/pcs/src/fft.rs, over constants::FR_TWO_ADIC_ROOT_OF_UNITY).

3.2 The transcript schedule#

pcs::open and pcs_verify::scalars run this Fiat–Shamir schedule step for step. A point or a list of points is one message (transcript.md §4).

# tag message
1 absorb MERCURY_INSTANCE n
2 absorb COMMITMENT cm, as passed: open never recommits it
3 absorb EVALUATION_CLAIM u_0..u_{s−1}, then v
4 absorb PCS_OPENING h
5 squeeze MERCURY_ALPHA α
6 absorb PCS_OPENING [q, g]
7 squeeze MERCURY_GAMMA γ
8 absorb PCS_OPENING [s, d]
9 squeeze MERCURY_Z z, by §3.4's rule
10 absorb PCS_OPENING g_z, g_{1/z}, h_z, h_{1/z}, s_z, s_{1/z}, one message
11 absorb PCS_OPENING pi_z, before δ although the batch does not read it
12 squeeze BDFG_BATCH δ
13 absorb PCS_OPENING w
14 squeeze BDFG_POINT z′
15 absorb PCS_OPENING w_prime
16 squeeze PAIRING_MERGE ρ, after all eight points and six values

The prover draws ρ too and discards it, so both sides leave the transcript in one state and an opening composes inside a larger transcript, the shard transcript (proof.md §4).

3.3 The BDFG20 batch#

Mercury §6 step 4(e) leaves the batched KZG opening to BDFG20 §4. The point set is T = {z, 1/z, α}, and the four polynomials are batched in this order, which fixes the power of δ each carries (pcs_verify::bdfg::items, which both sides read):

i f_i S_i Z_{T∖S_i} r_i interpolates
0 g {z, 1/z} X − α g_z, g_{1/z}
1 h {z, 1/z, α} 1 h_z, h_{1/z}, h_α
2 S {z, 1/z} X − α s_z, s_{1/z}
3 D {z} (X − 1/z)(X − α) D_z
F(X) = Σ_i δ^i·Z_{T∖S_i}(X)·(f_i(X) − r_i(X))                          W  = [(F/Z_T)(x)]_1
L(X) = Σ_i δ^i·Z_{T∖S_i}(z′)·(f_i(X) − r_i(z′)) − Z_T(z′)·(F/Z_T)(X)    W′ = [(L/(X − z′))(x)]_1

Both divisions are exact for an honest prover, and open asserts it (pcs_verify::bdfg::{quotient, linearization}).

3.4 Challenges and derived values#

Mercury draws z ∈ F*; here z is drawn again under MERCURY_Z while it is zero (pcs_verify::challenge_z). T needs three distinct points, so both sides refuse with PcsError::DegenerateChallenge when z² = 1, z = α or z·α = 1 (pcs_verify::degenerate): probability about 2^−252, and a loss of completeness only. The recursion tape draws z once and asserts all four conditions (verifier_core::tape::mercury_scalars).

The verifier is not sent h(α) or D(z): it derives them, as Mercury §6 step 4(c) does (pcs_verify::derive_h_alpha), and the prover builds the batch around the same derived values.

D_z = z^{b−1}·g_{1/z}
h_α = (g_z·P_u1(1/z) + g_{1/z}·P_u1(z) + γ·(h_z·P_u2(1/z) + h_{1/z}·P_u2(z) − 2v)
       − z·s_z − s_{1/z}/z) / 2

Opening D at z to D_z is the degree check on g (Mercury §4.3); opening h at α to h_α is §3.1's identity at z.

4 The proof and the verifier's checks#

pcs::MercuryProof is eight points and six values. Its field order is its byte order and its transcript order, and to_bytes writes pcs::PROOF_BYTES = 704 bytes for every n and k:

h  q  g  s  d  pi_z  w  w_prime                 8 × 64 bytes, G1 uncompressed (primitives.md §3)
g_z  g_inv_z  h_z  h_inv_z  s_z  s_inv_z        6 × 32 bytes, canonical Fr (primitives.md §1)

from_bytes returns None unless every point decodes through curve::G1Affine::from_bytes (canonical and on the curve; G1's cofactor is 1) and every value through field::Fr::from_bytes.

Two relations are checked, each written e(A, [1]_2) = e(B, [x]_2) so that both G2 arguments are SRS constants: the fold identity at z (Mercury §6 step 4(f), its z term moved into G1) and the BDFG20 batch (BDFG20 §4.1). They merge under ρ into one curve::pairing::pairing_check of two pairs:

A1 = cm − (z^b − α)·q − g_z·[1]_1 + z·pi_z                  B1 = pi_z
A2 = Σ_i c_i·cm_i − K·[1]_1 − Z_T(z′)·w + z′·w_prime        B2 = w_prime
     cm_i = g, h, s, d    c_i = δ^i·Z_{T∖S_i}(z′)    K = Σ_i c_i·r_i(z′)
     Z_T(z′) = (z′ − z)(z′ − 1/z)(z′ − α)
accept iff  e(A1 + ρ·A2, [1]_2)·e(−(B1 + ρ·B2), [x]_2) = 1

If either relation is false the merged one holds for at most one ρ, and ρ follows every proof element. The verifier reads three SRS points, srs::SrsVerifier's [1]_1, [1]_2 and [x]_2, and does no G2 arithmetic.

pcs::verify refuses, in order: a u whose length is not an instance's (§1), before anything is absorbed (UnsupportedNumVars); a proof point or cm off the curve (InvalidPoint), checked again because a proof built in memory has met no decoder; a degenerate T (DegenerateChallenge); a failed pairing check (VerificationFailed), which does not say which relation failed.

5 Batching k columns at one point#

Not in the papers. k commitments to columns of one size, opened at one point u, are one Mercury instance with one proof (pcs::batch_open, pcs::batch_verify); a shard proof's opening is one such batch (proof.md §5). Three steps precede §3.2's sixteen (pcs_verify::batch_preamble):

# tag message
B1 absorb COMMITMENT cm_0..cm_{k−1}, as passed, one message of 4k limbs
B2 absorb EVALUATION_CLAIM u_0..u_{s−1}, then v_0..v_{k−1}
B3 squeeze MERCURY_BATCH ρ

The opening then runs on (cm*, u, v*), with cm* = Σ_i ρ^i·cm_i and v* = Σ_i ρ^i·v_i.

  • ρ follows every commitment and every claimed value. Column i carries ρ^i, column 0 carrying 1, so a reordered or shortened list is a different statement.
  • The list is one message, so its length 4k fixes k, and then s from B2's s + k scalars: the absorbed stream is injective.
  • ρ = 0 is not redrawn: it checks column 0 alone, and is one of the roots the bound below counts.
  • A batch of one is a different transcript from a bare opening; their proofs do not interchange.

The batch is sound: by §2's linearity cm* commits to f* = Σ_i ρ^i·f_i, and evaluation at u is linear, so v* − f̂*(u) = Σ_i (v_i − f̂_i(u))·ρ^i, a polynomial in ρ of degree at most k − 1 fixed before ρ is drawn. A false claim survives with probability at most (k − 1)/|Fr|.

The prover builds f* as one Fr column and opens it once; mixed sizes are refused (MixedColumnSizes). The verifier refuses an empty list (EmptyBatch) or a value count that differs (BatchLengthMismatch), checks every cm_i on the curve before summing, derives cm* by a k-point MSM and runs §4 on it. pcs::batch_open_stacked opens recursion stacks at u ‖ r (recursion.md §1.3); batch_open is it at r = [], one column a stack.

6 Deferred verification and the accumulator#

6.1 The twelve entries#

Deferring a verification runs every check of §4 but the pairing and keeps the relation's terms: twelve pcs::AccumulatorEntry { side, scalar, point }, side a pcs::PairingSide, G2One for [1]_2 or G2X for [x]_2. The points are [cm, h, q, g, s, d, pi_z, w, w_prime, [1]_1], as pcs_verify::ENTRY_POINTS indexes them, and the scalars are pcs_verify::scalars's, in §4's notation:

# side point scalar # side point scalar
0 G2One cm 1 6 G2One pi_z z
1 G2One h ρ·c_1 7 G2One w −ρ·Z_T(z′)
2 G2One q −(z^b − α) 8 G2One w_prime ρ·z′
3 G2One g ρ·c_0 9 G2One [1]_1 −(g_z + ρ·K)
4 G2One s ρ·c_2 10 G2X pi_z 1
5 G2One d ρ·c_3 11 G2X w_prime ρ

The G2One terms sum to A1 + ρ·A2 and the G2X terms to B1 + ρ·B2; entry 9 carries both relations' [1]_1, and entry 2 is zero exactly when z^b = α, which is legal. A batch derives cm* first, so entry 0 is cm* and a check is ENTRIES_PER_CHECK = 12 entries whatever k. pcs::verify and pcs::batch_verify spend the entries at once; pcs::verify_deferred and pcs::batch_verify_deferred return them.

6.2 What uses it#

  • Base verification pairs: crates/verifier runs pcs::batch_verify for each shard, and no ShardProof or BlockProof carries an entry.
  • The recursion tree folds. A shard's tape computes the twelve scalars over field cells (verifier_core::tape::mercury_scalars), cm* being a hint; the node folds them with the batch check cm* = Σ_i ρ^i·cm_i (recursion.md §8.3), and one pairing check at the top discharges every shard's (recursion.md §9). Natively, host::recursion runs pcs::batch_verify_deferred on each shard for its cm*.
  • Nothing else: pcs::verify_deferred, §6.3's word form, pcs::accumulator_digest and pcs::discharge are called only by crates/pcs's tests and tools/kat-gen.

6.3 The word form and discharge#

A list is grouped into deferred checks, checks[j] being group j's entry count, and written as canonical Fr words (pcs::accumulator_words, inverse pcs::accumulator_from_words):

group:  count  entry_0 .. entry_{count−1}
entry:  side  scalar  x_lo  x_hi  y_lo  y_hi       side 0 = G2One, 1 = G2X; ENTRY_WORDS = 6

A word is 32 bytes, so an entry is 192, and the limbs are the point's transcript form (transcript.md §4). There is no header, so two lists concatenate into a list whose checks keep their groups. Decoding refuses a count of 2^64 or more or one that overruns, a side other than 0 or 1, a limb of 2^128 or more other than the sentinel, a partial sentinel, the all-zero quadruple (infinity has one spelling), and a point that is not canonical or not on the curve. The digest is the words as one ACCUMULATOR_DIGEST message in a fresh sponge, then a raw sample; covering the count words, it binds the grouping.

discharge(vsrs, entries, checks):
  every entry's point on the curve, before anything else
  ν = fresh sponge: absorb ACCUMULATOR_DIGEST [digest], challenge ACCUMULATOR_MERGE
  A = Σ_j ν^j·(group j's G2One terms)      B = Σ_j ν^j·(group j's G2X terms)
  accept iff e(A, [1]_2)·e(−B, [x]_2) = 1

An entry's point is a claim: absorption binds only its limbs, and an entry built in memory has met no decoder. The weight keeps the checks apart: at weight 1, two checks with equal and opposite errors pass together, and weighted, a false group passes only where ν is a root of a nonzero polynomial of degree below the group count. ν is a function of the words because discharge takes no transcript. An empty list discharges.

7 Cost and security#

prover, field O(n): a pass for h, the fold, H's division; S in O(b log b)
prover, MSMs 2n + 5b − 4 scalar multiplications: q n − b, pi_z n − 1, h, g, d b each, s, w, w_prime b − 1 each. A commitment is one more MSM of n
batch of k k multiply-adds a coefficient for f* and a k-point MSM for cm*, then one opening
verifier O(t) field operations, MSMs of ten points and of two (and of k), one two-pair pairing check
measured n = 2^22: commit 1.30 s, open 2.89 s. 16 columns of 2^20: a batch opens in 1.01 s and verifies in 4.8 ms, 16 single openings take 9.79 s and 62 ms. 18-core Apple M5 Pro; bench mercury, bench mercury-batch

Knowledge soundness holds in the algebraic group model under q-DLOG (Mercury §6, BDFG20 §4), with Fiat–Shamir over the Poseidon2 transcript in the random-oracle model and an SRS whose x nobody knows (srs.md §3). The statistical terms are Schwartz–Zippel over α, z and z′, of order a committed polynomial's degree over |Fr|, a few 1/|Fr| for γ, δ and the merge ρ, (k − 1)/|Fr| for a batch and the group count over |Fr| for ν: each is below 2^−220 for every instance in use, and the level is BN254's (architecture.md §4). Nothing is hiding and nothing is blinded.

Mercury's SRS has exactly n powers; here one SRS serves every size, so a prover can commit to a polynomial of degree n or more, and no degree bound is checked. None is needed: Mercury §6's argument goes through with its Schwartz–Zippel terms over that degree, and the opening at u is the multilinear extension of the polynomial's first n coefficients. A commitment binds that truncation, which is linear, so §5's argument holds for it too.

Auditoren/Programm und Ausführung

Das Programm: vom ELF zur Identität

Normative Spezifikationdocs/spec/program.mdAls Markdown anzeigen

Zusammenfassung

Wie ein Gastprogramm-Binary zur statischen Beschreibung wird, die ein Verifier kennt. Die Seite spezifiziert das Laden des ELF, den Durchlauf über die Halbwörter, der komprimierte Befehle expandiert, ohne eine Adresse zu verschieben, das Serialisierungsformat von ProgramImage, den Decoder und seine Zuordnung der 59 Befehle von RV32IMA zu sieben Familien, die dekodierten Tabellen mit einer Zeile pro Halbwort, die One-hot-Befehlsmaske, die VmConfig und ihre Höhen sowie die Programmidentität: genau, was sie bindet und was nicht.

Der folgende normative Text wird auf Englisch gepflegt, der kanonischen Sprache der Spezifikation.

How a guest binary becomes the static, verifier-known description of a program. crates/loader reads an ELF into a ProgramImage, crates/isa decodes its instructions, and crates/program routes them into per-family decoded tables, derives the VmConfig and commits to all of it as the program identity. Every step is a pure function of its input.

1 Loading#

loader::load_elf accepts a static executable — ELFCLASS32, little-endian, ET_EXEC, EM_RISCV — whose PT_LOAD segments lie inside guest RAM (constants::guest_memory) at even addresses, pairwise disjoint, with p_filesz ≤ p_memsz, at least one of them executable. Anything else is a named LoaderError: DynamicElf for ET_DYN, PT_DYNAMIC or PT_INTERP, EntryNotAnInstruction for an e_entry that is not the first halfword of an instruction, and the sweep's refusals (§2).

Of a program header it reads p_type, p_offset, p_vaddr, p_filesz, p_memsz and the PF_X bit, and nothing else: the VM has no pages, and all of RAM is addressable whatever the segments declare. The address map, and the segment layout a guest ELF keeps for host loaders, are ecall-abi.md §6.

2 RVC expansion and slots#

slots holds one Slot per halfword from slot_base, the lowest loaded address, to the end of the highest executable segment: the slot of pc is slots[(pc − slot_base)/2]. load_elf sweeps each executable segment's file bytes from its start, by the halfword at pc:

low bits 11   pc += 4   Instruction { word: the four bytes, compressed: false }, MidInstruction
0x0000        pc += 2   NonInstruction
otherwise     pc += 2   Instruction { word: rvc::expand(halfword), compressed: true }

Every other halfword is NonInstruction. An encoding longer than 32 bits (InstructionTooLong), one cut off by the end of the file bytes (TextTruncated) or a halfword rvc::expand refuses (RvcIllegal) refuses the image; 32-bit words are decoded in §5.

  • Addresses are never compacted. A c.addi at 0x1002 stays there and occupies two bytes, so linker-resolved addresses hold; compressed, the instruction's length, is the only record of whether the next pc is pc + 2 or pc + 4.
  • rvc::expand takes the base C extension in its RV32 form and refuses the floating-point forms, the RV64-only forms (c.addw, c.subw, a shift with shamt[5]), the reserved code points and the Zc* encodings. A HINT such as c.addi x0, 5 is expanded; its 32-bit form writes x0.
  • 0x0000, RVC's defined-illegal encoding, is not refused: LLVM pads unreachable blocks with it. Reaching it is fatal at run time.

A desynchronised sweep cannot make a wrong instruction provable. A slot is a function of the bytes at its own pc, so every Instruction slot is what a hart fetching there would decode; data that shifts the sweep off the true boundaries can only lose true instruction starts, whose pcs then have no table row (§5), or meet an unclaimed encoding and refuse the image. crates/loader/tests/differential.rs holds committed guests' slots to llvm-objdump's listing, and the expansion to LLVM's own encoder over guests/rvc-dense, one sequence assembled compressed and not: a wrong expansion would be a valid proof of another program.

3 ProgramImage and its wire form#

The wire form is postcard over ProgramImage's four fields in order, with no header; every integer but kind is a LEB128 varint:

ProgramImage = entry ‖ n ‖ n × Segment ‖ slot_base ‖ m ‖ m × Slot
Segment      = vaddr ‖ mem_len ‖ len ‖ bytes    mem_len is p_memsz; bytes, the p_filesz file bytes
Slot         = kind: u8 ‖ word                  kind 0 a four-byte instruction, 1 a two-byte one,
                                                2 MidInstruction, 3 NonInstruction; word 0 for 2, 3

The reader re-checks what load_elf establishes — segments at even addresses, sorted, disjoint and inside RAM; slot_base the lowest segment's address; each four-byte Instruction followed by its MidInstruction; entry an Instruction slot — but not slots against the bytes. artifact-dump writes this form (tools.md §5); no prover or verifier reads it, host::setup starting from the ELF.

ProgramImage::initial_word(addr) is the little-endian word at addr before the first cycle: file bytes where a segment has them, zero elsewhere. The image column (§8) and the trace's initial RAM values are read from it.

4 The instruction set and family routing#

isa::decode takes 32-bit words only and accepts exactly RV32IMA's 59 instructions — 40 of RV32I, 8 of M, 11 of A — with any value in an operand field, x0 destinations included, and the one legal value in every fixed field: funct7, jalr's funct3, all of ecall and ebreak, an atomic's .w width, lr.w's rs2 = 0. Everything else is a DecodeError: RV64 encodings, F, D, Zicsr, fence.i, privileged instructions. crates/isa/tests/sweep.rs holds it, over all 2^30 words with low bits 11, to accepted counts derived from the ISA's tables and to an independent encoder.

  • fence is every MISC-MEM word with funct3 = 000, 2^22 of them, whatever its rd, rs1, fm, pred and succ: the ISA has a base implementation treat a reserved setting as a normal fence (llvm-objdump prints those <unknown>), and on one hart a fence does nothing.
  • An immediate is the value the instruction uses: sign-extended for I, S, B and J, the shifted word for U, the amount for a shift immediate. B and J displacements are even by encoding; nothing asks for 4-byte alignment.

program::row_kind, a total function, routes an instruction to one family and one bit of that family's mask (§6):

id family mnemonics, from mask bit 0 up
0 ADD_SUB_LUI_AUIPC system (ecall ebreak fence), addi auipc add sub lui
1 JUMP_BRANCH_SLT slti sltiu slt sltu beq bne blt bge bltu bgeu jalr jal
2 SHIFT_BITWISE slli xori srli srai ori andi sll xor srl sra or and
3 MUL_DIV mul mulh mulhsu mulhu div divu rem remu
4 MEM_WORD lw sw
5 MEM_SUBWORD lb lh lbu lhu sb sh
6 ATOMICS amoadd amoswap lr sc amoxor amoor amoand amomin amomax amominu amomaxu, each .w

5 Decoded tables#

program::decode_program(image, params) decodes every Instruction slot — a word isa::decode refuses fails the program, reachable or not (NotAllOpcodesSupported) — and builds a table for each family of the config. An instruction family's columns are committed setup columns, which each cycle's decoder lookup reads (lookup.md §10); other families' tables have none.

  • One row per halfword, absolute. Row i is pc 2i, and a table has exactly its family's height h (§7).
  • A live row holds one of the family's instructions in the fields of its lookup tuple (program::lookup_tuple): pc, next_pc, rs1, rs2, rd, imm, extra_mask, without imm for MUL_DIV and ATOMICS; no tuple holds funct3, the mask saying more. next_pc is the fall-through, pc + 2 or pc + 4 by the slot's length, never a branch target. A register the form lacks is 0; imm is the two's complement of §4's value, 0 where the form has none, or a system code (§6).
  • Every other row is padding, Fr::MINUS_ONE in every field (FamilyTable::column_poly). An all-zero row would be a claimable instruction at pc 0 with an empty mask; a live field is below 2^32, so no live row is the padding row.
  • Reach. Derivation fails (TableTooShort) unless the family's own last instruction has pc ≤ 2h − 4; another family's code may lie beyond it. Code is linked from RAM_ORIGIN = 2^16, so a family reaches 1.9375 MiB of it at 2^20 and 7.9375 MiB at 2^22, the largest height.

Every Instruction slot is a live row of exactly one table, and no table has another (check_partition).

Code is static. A cycle's instruction comes from these tables, never from RAM: a store into .text changes what a load reads, not what executes, and a pc that is not an Instruction slot has no row, so reaching it is fatal and unprovable (execution-trace.md §10).

6 The extra mask#

A tuple's last field, family_extra_mask, is 1 << kind, the kind being the instruction's position in its row of §4's table (constants::extra_mask). A kind is a mnemonic, except family 0's bit 0, the system kind, whose three instructions are told apart by imm (constants::extra_mask::system_code): ecall 0, ebreak 1, fence 2. A fence's fm, pred and succ, and an atomic's aq and rl, are not recorded; on one hart they order nothing.

One-hotness is the table's, not a gate's: a circuit holds each bit it extracts boolean, and the decoder lookup, which admits only the table's rows, is what excludes an empty or many-bit mask (lookup.md §10).

7 VmConfig and heights#

verifier_core::VmConfig { families: Vec<(family, height)>, bytecode_size_words } is a program's static shape: its families, ascending by id (constants::family), each with its height, the row count of one of its shards; an execution's shard counts are not in it. decode_program derives the family set, and nothing selects it:

  1. an instruction family (0–6), whose rows are cycles, is present when the image holds one of its instructions;
  2. a window family, whose rows are memory locations — INIT_TEARDOWN (7), ZERO_WINDOWS (8), PUBLIC_INPUT (12), PUBLIC_OUTPUT (13), ADVICE_WINDOWS (14) — is always present, and FIELD_WINDOWS (18) when one of families 19–22 is, which puts the config in the recursion format (VmConfig::is_recursion), the one whose registry VmConfig::circuit reads for 18–22 (recursion.md §1.1, §1.2);
  3. a delegation family (9–11, 15–17, 19–22), whose rows are invocations, is present when the image declares it by a record among its file bytes (program::declared_delegations, delegation.md §7); a record naming a number no family answers is UnknownDelegation.

A height is a parameter (ProgramParams::heights, defaulting to constants::family::DEFAULT_HEIGHTS, circuits.md §1), except the public families', pinned at 2^12 because it places their windows (public-values.md §2). Four things constrain it:

  • the menu, constants::family::HEIGHT_MENU: 2^8, 2^12, 2^16, 2^18, 2^20, 2^22 (HeightNotOnMenu), even powers of two as a Mercury opening needs (mercury.md §1);
  • an instruction family's code (§5);
  • the window rules (verifier_core::window_height, WindowRule; memory.md §3.5), and RAM window 0, [0, 4·h_w), holding every file byte of the image (ImageOutsideWindow);
  • the floor of the family's lookup channels, below which the registry has no circuit and no key can be built: 2^20 for an instruction family (lookup.md §3), its own for a delegation family (delegation.md §9).

bytecode_size_words, 2^20 (4 MiB) by default, is a declared ceiling on the words from RAM_ORIGIN to the image's last file byte (ProgramTooLarge); no circuit reads it.

The wire form is 8k + 8 bytes; VmConfig::from_bytes refuses a wrong length, a family id above 22, ids not strictly ascending, a height off the menu, and what window_height refuses:

k: u32 LE ‖ k × (family: u32 LE ‖ height: u32 LE) ‖ bytecode_size_words: u32 LE

In a transcript a config is one VM_CONFIG message, [f_1 … f_k, h_1 … h_k, bytecode_size_words]: the second message of the identity (§8) and the first of a statement's descriptor (proof.md §2).

8 Program identity#

One Fr, ProgramIdentity, on the wire its canonical 32 bytes (primitives.md §1): the raw squeeze of a fresh transcript after these messages (verifier_core::identity_digest; tag values in transcript.md §5):

# tag message
1 PROGRAM_IDENTITY [code_version]: constants::family::CODE_VERSION, 0, the only one derivation builds (UnsupportedCodeVersion)
2 VM_CONFIG §7's
3 PROGRAM_ENTRY [entry_pc]
4 COMMITMENT, one per family of the config, ascending its setup commitments, four limbs a point (transcript.md §4)

A family's setup commitments (program::setup_commitments) are Mercury commitments (mercury.md §2): of an instruction family's decoded-table columns in tuple order, 7 or 6 points; of INIT_TEARDOWN's image column (program::image_init_column), row y being image.initial_word(4y) over RAM window 0, one point; none, an empty message, for every other family.

Committing needs the ceremony's SRS, with as many powers as the tallest table has rows (srs.md §1). The digest over given points (program::identity_from_commitments) needs no SRS and no curve arithmetic: a verifying key carries the lists, its load recomputes the identity from them, and every shard opens its setup columns against the same points (proof.md §5, §7), which is what ties the tables a proof reads to the identity.

It binds every instruction the sweep found, with its pc, length, operands and kind, and that no other pc holds one; every file byte of the image (.text, .rodata, .data, the delegation declarations among them); the entry pc; the family set, every height, bytecode_size_words and the code version. One ELF at two settings of the heights has two identities.

It does not bind:

  • the SRS its commitments are under, or the generic lookup table: those are the SRS digest's (proof.md §3);
  • the circuits, which a key's load holds to the registry (proof.md §7);
  • anything an execution chooses: its input, advice, shard counts, window list;
  • memory past a segment's file bytes (.bss, the heap and the stack): mem_len enters nothing, and such memory starts at zero whatever is declared;
  • the symbol table, which loader::function_symbols and loader::symbol_names read beside the image for the profiler and listings, or anything else of the ELF §1 does not read.

A verifier takes the identity from a channel the prover does not control and compares it with its key's; it never sees an ELF. Against a prover-supplied identity a proof shows only that some program ran. Whoever holds the ELF, the parameters and the ceremony file recomputes it.

Auditoren/Programm und Ausführung

Das ABI der Gastprogramme

Normative Spezifikationdocs/spec/ecall-abi.mdAls Markdown anzeigen

Zusammenfassung

Die Schnittstelle zwischen einem Gastprogramm (Guest) und der Maschine. Die Seite spezifiziert die Aufrufkonvention für ecall, die Nummernbereiche und jede implementierte Nummer (EXIT und zehn Delegationen), die stillgelegten Nummern und warum sie nie wiederverwendet werden, wie die Maschine auf jede andere Nummer antwortet, die gesamte 32-Bit-Adresskarte mit dem vom Linker-Skript erzwungenen Segmentlayout, und die öffentliche Oberfläche des Guest-SDK mit seinen Exit-Status.

Der folgende normative Text wird auf Englisch gepflegt, der kanonischen Sprache der Spezifikation.

The ecall convention, every ecall number, the guest's address space and the SDK over them. A guest has no file descriptors and no I/O syscall: its public input, journal and advice are memory (public-values.md), so an ecall only ends the execution or hands a frame to a circuit. crates/constants/tests/ecall_abi.rs holds this page's tables to constants::ecall and its MEMORY line to link.ld.

1 The calling convention#

Register Role
a7 the number
a0 in: the one argument, an exit status or a delegation's frame base (delegation.md §4)
a0 out: the result, 0 or a negated errno

a1–a5 are reserved for a call that needs more arguments; none does. A recursion-format delegation answers its frame base advanced past the frame (recursion.md §1.4). An ecall preserves every register but a0: its row writes no other (execution-trace.md §6), so the memory argument carries the rest across it.

2 The number ranges#

Constant Value What
ZKVM_IO_FIRST 0x0400 first host call
ZKVM_IO_LAST 0x04FF last host call
PRECOMPILE_FIRST 0x0500 first precompile
PRECOMPILE_LAST 0x05FF last precompile

Both ranges lie above 1023, the whole Linux number space, and are disjoint, so a number says its class: a host call would return a value the prover chose, a precompile is a deterministic function of guest memory that its circuit proves. The host-call range is reserved and empty: advice is a memory region the prover fills and the guest checks (public-values.md §6), inside the memory argument, where a value returned in a register would be bound to nothing.

3 Syscall numbers#

Every number this VM implements. All but EXIT are delegations, whose families, anchor spaces and frames are delegation.md §3's registry: the first six are the base format's, the last four the recursion format's (recursion.md §2).

Number Constant Class What
93 EXIT deterministic end the execution with status a0, the statement's exit status; nonzero is a failed execution, still provable
0x0500 PRECOMPILE_POSEIDON2 deterministic the width-3 Poseidon2 permutation over canonical Fr lanes
0x0502 PRECOMPILE_FR_ARITH deterministic one Fr add, multiply or inverse over Fr's in-memory form
0x0504 PRECOMPILE_MOD_MUL deterministic a·b mod m, m one of four Ethereum moduli a selector names
0x0506 PRECOMPILE_EC_ADD deterministic one third of a complete point addition, secp256k1 or BN254 G1
0x0507 PRECOMPILE_KECCAK_F deterministic one round of keccak-f[1600]; a permutation is 24 calls
0x0508 PRECOMPILE_SHA256_COMP deterministic four rounds of SHA-256's compression; a compression is 16 calls
0x0509 PRECOMPILE_FR_OP deterministic one operation over field cells
0x050A PRECOMPILE_P2_FIELD deterministic one transcript duplex step over field cells
0x050B PRECOMPILE_FIELD_IO deterministic eight RAM words into a field cell, or back
0x050C PRECOMPILE_FQ_OP deterministic one BN254 base-field operation over field cells

A class says who chooses the result. deterministic: a function of the guest's own state, which a circuit proves. advice: chosen by the prover; no number has it (§2). These are exactly the ecalls a proof admits: ADD_SUB_LUI_AUIPC holds every ecall row's a7 to 93 or to a registered delegation number, the base format's circuit knowing the first six (add-sub.md, recursion.md §1.2).

4 Retired numbers#

Number Constant Was
63 none POSIX read(fd, buf, len)
64 none POSIX write(fd, buf, len)
0x0501 RETIRED_KECCAK_F_WHOLE_PERMUTATION a whole keccak-f[1600] over a 200-byte frame, which 0x0507 replaces
0x0503 RETIRED_MOD_MUL_WITNESSED_MODULUS a·b mod m over a 128-byte frame carrying m, which 0x0504 replaces
0x0505 RETIRED_SHA256_COMP_WHOLE_COMPRESSION a whole compression over a 96-byte frame, which 0x0508 replaces

A number is assigned once. A retired one is never reassigned and answers -ENOSYS (§5): given a second meaning, it would run an old binary with its frame misread to a plausible wrong answer.

5 Every other number#

Constant Value What
ENOSYS 38 answered as -ENOSYS in a0

A number not in §3 answers -ENOSYS and falls through: the retired numbers, the host-call range, and every syscall a library might make for host data — getrandom, clock_gettime, the seeding of std's RandomState. Host data is prover advice, and a guest that needs it takes it from the advice region, where checking it is visibly the guest's job. Such a call executes and cannot be proved: the ADD_SUB_LUI_AUIPC fill refuses its row.

-ENOSYS is also the delegation ABI's "no circuit" answer, on which a base-format shim runs its software path (delegation.md §2); this executor never gives it to a §3 number. A registered number the image did not declare is the fatal DelegationFamilyAbsent on the tracing paths (delegation.md §7).

6 The memory map#

crates/guest-sdk/link.ld declares one region, constants::guest_memory's RAM_ORIGIN and RAM_LENGTH:

ld
MEMORY { RAM (rwx) : ORIGIN = 0x00010000, LENGTH = 0x7FFF0000 }

The whole 32-bit address space:

[0x0000_0000, 0x0000_8000)  hole: no family initializes it; an access is a fatal OutOfBounds
[0x0000_8000, 0x0000_C000)  public input window    PUBLIC_INPUT_ORIGIN    16 KiB
[0x0000_C000, 0x0001_0000)  journal                PUBLIC_OUTPUT_ORIGIN   16 KiB
[0x0001_0000, 0x8000_0000)  RAM                    RAM_ORIGIN, RAM_LENGTH
    0x0001_0000             .text, _start first; .rodata, .data, .bss, each page-aligned
    __heap_start            .bss's end rounded up to 16; the heap grows up from here
    0x7F80_0000             __stack_top − STACK_RESERVE (8 MiB): no heap block ends above it
    0x8000_0000             __stack_top, the initial sp; the stack grows down
[0x8000_0000, 2^32)         advice                 ADVICE_ORIGIN          up to 2^29 words
  • A load or store reaches the four regions alike, every word carrying the RAM tag; the windows' and the advice's layouts, families and binding are public-values.md §2–§6. None is in the ELF, so no linker symbol names them. Advice is addressable only up to the words the host supplied, and not at all when it supplied none.
  • The hole makes a null dereference a fatal error rather than a trace nothing could prove.
  • A delegation frame lies wholly in RAM (delegation.md §4). crates/loader refuses a PT_LOAD outside RAM (program.md §1), and a decoded table's height bounds how far .text reaches (program.md §5).
  • crt0's _start sets sp, zeroes [__bss_start, __bss_end) byte by byte, so that a zero .bss is the image's property and not the executor's, calls main, and exits 0 if it returns.
  • Nothing detects a stack that grows past its reserve after the heap has filled below it.

6.1 The segment layout#

A guest ELF loads under two loaders. crates/loader lays its PT_LOADs into a flat space the executor makes addressable whatever the headers say, with no pages and no permissions. A host loader maps exactly the PT_LOADs, page by page, at their permissions, and nothing else exists. The headers are the image's account of its own memory, read by every tool but this VM, so link.ld makes them true:

  • Every writable byte is declared. .bss runs to ORIGIN(RAM) + LENGTH(RAM), so the heap and the stack lie in one writable segment ending at __stack_top, whose file bytes stop at or before .bss, the 2 GiB reservation being NOBITS. Undeclared, the first stack push would fault.
  • No two segments share a page. .text, .rodata, .data and .bss are each 4096-aligned: a page two mappings share takes the second's permissions, stripping execute from .text's tail or putting zero fill on a read-only page, which a host loader refuses.

crates/loader/tests/layout.rs holds the committed guest ELFs to both by parsing their headers.

7 The guest-sdk surface#

crates/guest-sdk is the guest's runtime. Only exit and the delegation shims issue an ecall; the rest is loads and stores.

Item What
entry!(f) exports the main crt0 calls, a wrapper calling f
public_input() the public input payload, its length word clamped to the window
read_input(buf) copies min(buf.len(), public_input().len()) bytes and returns the count: it may return short
commit(bytes) appends to the journal and its length word; exits 70 rather than overflow the window
journal() what has been committed
advice() the advice payload, its length clamped to the region; bound by nothing, so the guest checks it
exit(code) EXIT; publishes nothing beyond what was committed
keccak256, sha256 over KECCAK_F and SHA256_COMP, with a software fallback on -ENOSYS from the first call
poseidon2_permute over POSEIDON2; false on -ENOSYS, for the caller's own permutation
ec_add, ec_mul, ec_identity homogeneous projective points over EC_ADD; None on -ENOSYS
recursion::* the raw shims over word-aligned frame types, false on -ENOSYS; the recursion format's (fr_op, p2_field, field_io, fq_op and the tape helpers import, import_run, replay) have no software path
allocator bumps up from __heap_start, never frees; exits 71 when a block would end above __stack_top − STACK_RESERVE or the live sp
panic handler exits 101 and writes nothing: a panicking guest is provable, having published what it committed

A delegation answer other than 0 or -ENOSYS exits 72, as do -ENOSYS after the first call of a multi-call operation and a recursion call that does not leave a0 past its frame. Each shim reads its number from its declaration record (delegation.md §7); which library code reaches which shim is delegation.md §10's.

Auditoren/Programm und Ausführung

Der Ausführungs-Trace

Normative Spezifikationdocs/spec/execution-trace.mdAls Markdown anzeigen

Zusammenfassung

Was jede Speicherabfrage einer Ausführung ist und wann sie stattfindet. Die Seite definiert den Zyklus aus vier Zeitstempeln und die 38-Bit-Uhr, die Adressräume, eine Abfrage als Lese- und Schreibzugriff an einer Adresse, den Frame jeder Befehlsklasse, die x0-Regel, die ecall-Zeile, die Reihenfolge des Ereignislogs, die Zuordnung von Zyklen zu Familien, die Selbstprüfung des Speichers auf Trace-Ebene, den Emulator mit seinen drei Abweichungen von einer gehosteten RV32IMAC-Umgebung und die Container, die einen Trace aufnehmen.

Der folgende normative Text wird auf Englisch gepflegt, der kanonischen Sprache der Spezifikation.

What every memory query of an execution is, when it happens and the order the trace records it in; the emulator that produces a trace and the containers that hold one. The memory argument (memory.md) and every family's frame are built on this convention.

1 The clock#

Cycle c occupies the four timestamps 4c + Δ, one per slot Δ ∈ {0, 1, 2, 3} (constants::memory::TS_STEP). Every instruction is one cycle and nothing else is: a delegation invocation rides the cycle that requested it, so the cycle count is the instruction count.

  • Cycles are numbered from 1. Timestamp 0 is every address's initial write, and a read must strictly precede its write, so a cycle-0 pc query could not follow the value it reads.
  • The clock is 38 bits (TS_BITS): every timestamp is below 2^38, the last cycle is 2^36 − 1, and the cycle that would pass it is the fatal ClockOverflow, raised before it is recorded.

2 Address spaces#

Tag Space Address At timestamp 0
1 REG a register index, 0..32 0, x0 included
2 RAM the byte address of a 4-aligned word in RAM, a public window or the advice region the image's bytes in RAM, 0 past them; the public input's and the advice's layouts; 0 in the journal
3 PC 0 the entry point
4–9 DELEGATION_KECCAK_F … DELEGATION_EC_ADD, a base-format delegation family's anchor each a frame base no initial write
10 FIELD a cell, any u32 0
11–14 DELEGATION_FR_OP … DELEGATION_FQ_OP, a recursion family's anchor each a frame base no initial write

The tags are nonzero so that no real tuple is all zeros, as x0's initial write would be. A byte or halfword access queries its word, and trace::InitialMemory is what RAM starts from. An anchor space, in delegation.md §3's order, is a delegation family's type, not memory: a query there reads the tuple stamped 0 with value 0, whatever came before; requests pair with invocations and nothing chains (delegation.md §5). Field cells hold whole Fr elements, reached only by the recursion families (recursion.md §2).

3 A query#

A memory query is one event at one address (trace::MemoryEvent): a read of read_value, last written at read_ts, and a write of write_value at ts = 4c + Δ.

  • A query that only reads writes back what it read: a register read or a load is one query.
  • read_ts < ts, strictly; the gap ts − read_ts − 1 is below 2^38.
  • Queries at distinct addresses may share a slot; two at one address never do. An address may be queried at two slots of a cycle — add a0, a0, a1 reads a0 at slot 1 and writes it at slot 3 — which is why the log is ordered by slot.

4 The frame of each instruction class#

Slot 0 is the pc query, every cycle: pc read, next_pc written. A register query exists for every register field of the decoded instruction, whatever register it names, x0 included.

Class Δ = 1 Δ = 2 Δ = 3
lui, auipc, jal rd
jalr, register-immediate rs1 rd
branches rs1 rs2
register-register, M rs1 rs2 rd
loads rs1 the word, read rd
stores rs1 rs2 the word, the stored bytes merged in
lr.w rs1 the word, written back; rd ← it
sc.w rs1 rs2 the word ← rs2; rd ← 0
AMOs rs1 rs2 the word ← op(old, rs2); rd ← old
fence
ecall a7 a0 (§6) a0 ← the result; a delegation's mirror query

ebreak has no row (§10). An atomic's row and a delegation request's carry two queries at slot 3, at distinct addresses. A family's frame is the union of its instructions' queries (memory.md §2). next_pc is the fall-through — pc + 2 after a compressed instruction, pc + 4 otherwise — except a jal's or taken branch's pc + imm, a jalr's (rs1 + imm) & !1, and the exit row's HALT_PC (memory.md §5).

An invocation's accesses ride its requesting cycle but belong to its own family's row: its frame words in RAM at slot 0 (constants::delegation::FRAME_DELTA), a FIELD_IO invocation's eight data words in RAM at slot 1 (constants::field_io::DATA_DELTA), and a recursion family's field cells at slots of its own (delegation.md §4, recursion.md §2.1).

5 The x0 rule#

x0 is an ordinary register in the trace and a constant in the machine: it starts at 0, a read of it is a REG query at address 0, and an instruction whose rd is x0 logs its slot-3 write with value 0, whatever it computed. So every query at x0 reads and writes 0, which the x0 gadget enforces (memory.md §2).

6 ecall#

An ecall's row is one cycle of ADD_SUB_LUI_AUIPC. It reads a7 at slot 1 and writes a0 at slot 3; the rest depends on the number (ecall-abi.md):

a7 Δ = 2 a0 written next_pc Besides
EXIT a0, the status the status HALT_PC the execution stops
a delegation number a0, the frame base 0, or for a recursion type the base past the frame (recursion.md §1.4) fall-through the mirror query at the frame base (§7); the invocation (§4)
any other none -ENOSYS fall-through no proof admits the row (ecall-abi.md §3)

7 The order of the log#

Events are recorded in cycle order and, within a cycle, by slot and then by role: the pc query; then an invocation riding the cycle, its frame words in frame order and a FIELD_IO invocation's data words after them; then one query per role the row has, in trace::ROLES order:

Role Slot Space What
rs1 1 REG rs1; an ecall's a7
rs2 2 REG rs2; an ecall's argument a0
load 2 RAM a load's word
ram 3 RAM a store's or an atomic's word
rd 3 REG rd; an ecall's result a0
delegate 3 the requested family's anchor space a delegation request's mirror query

ROLES is in slot order, so the log is in timestamp order, which MemoryState::record asserts; trace::Row::present holds one bit per role in a u8, and no two roles share a (space, slot) pair. The atomics family keeps its word at slot 3 for every instruction, lr.w included, so one frame serves the whole A extension.

8 Routing#

Every cycle goes to the one family whose decoded table claims its pc (program.md §4); a pc no table claims, or one claimed by a family not its instruction's, panics the tracer. An invocation goes by its type to its family's buffer.

9 The trace-level memory check#

MemoryEventLog::self_check(&InitialMemory) runs the memory argument natively over a whole log. First the timestamp rules: every address one its space has, every timestamp on the clock and in order, every read before its write, one query per address and timestamp. Then the balance: as multisets of (space, address, timestamp, value), an initial write at timestamp 0 of every touched address plus every query's write equals every query's read plus a teardown read of every address's last write. With one write per address and timestamp and no negative gap, this pairs each read with the last write before it: sequential consistency. What it cannot see:

  • Teardown is each address's last write, taken from the log, so everything after an address's last honest query balances by construction: a final value changed, a final query moved later or added, trailing cycles removed. In a proof the final values are the boundary scalars and the window families' teardown columns, fixed before any memory challenge, and the verifier fixes x0's and the pc's (memory.md §4, §5).
  • An anchor-space query is credited with its invocation's two tuples and balances alone; that requests and invocations pair 1:1 is the circuits' (delegation.md §5).
  • Field-cell accesses are not events (recursion.md §2.1).

10 The emulator#

crates/emulator runs RV32IMAC on one hart over a ProgramImage, with no interrupts and no privilege levels; aq/rl and fence order nothing. emulator::run returns an Execution: the registers, the exit status, the cycle count and the public values. emulator::trace_run returns the family buffers, the MemoryEventLog and the CycleProfile too, and emulator::StreamingRun, the prover's pull-based tracer, hands over a family's buffer as a ShardChunk the moment it reaches its height (streaming.md §2). The three differ only in what records a cycle, and a run is a pure function of (image, io), with no clock, randomness or threads, so two runs cut the same shards. A nonzero exit status is an execution, not an error.

Three points differ from a hosted RV32IMAC. sc.w always succeeds, storing and writing 0, as the circuits do (memory-ops.md §6). A misaligned halfword or word access is fatal, never split. The instruction stream is the image decoded at load, so a store into .text changes RAM and not what executes.

Every other stop is a fatal EmuError, and run and trace_run return no trace beside one: NotAnInstruction (the all-zero halfword included), IllegalInstruction, Ebreak, Misaligned (a frame base too), OutOfBounds (an access outside ecall-abi.md §6's regions, an advice word past what the host supplied, or a frame not wholly in RAM), ClockOverflow, PublicInputTooLong and JournalTooLong (the input, or the journal's length word at exit, above a window's payload), DelegationFamilyAbsent (on the tracing paths, a delegation number the image did not declare) and DelegationFrame (a frame its family has no witness for, delegation.md §6). An unassigned ecall number is not an error but -ENOSYS (§6).

There is no second executor: crates/emulator/tests/trace.rs restates §4's table and checks every traced row against it, and §9's check and the checker's multiset, memory and family-row suites hold the rest.

11 Trace containers#

crates/trace holds what an execution leaves; the emulator is its only producer.

  • Family buffers. trace::FamilyTraces holds one buffer per family of the VmConfig. A FamilyTrace, empty for a window family, is raw live rows, column-major, in small integer types: cycle, pc, next_pc, present, and per role addr, read_ts, read_value, write_value; no padding, no polynomial. A row stores everything its queries carry but a write timestamp, 4c + Δ, and the pc query's read timestamp, 4(c − 1). A delegation family's DelegationTrace has a row per invocation: the requesting cycle, the frame base, the frame words, and a recursion family's cell and data-word accesses.
  • RowSlice, FrameSlice. One shard's rows, [i·h, min((i + 1)·h, len)), borrowed: what the memory column builders read, never the log (memory.md §2). Row::delegation_space recovers a mirror query's space from the a7 the row read.
  • MemoryState, the last-access tables: each register's, the pc's, each RAM word's and each field cell's last (ts, value). O(touched addresses), and all the register and pc boundary, the RAM window list (trace::init_windows) and the window families' teardown need.
  • MemoryEventLog, the events and a MemoryState: O(cycles), kept only by trace_run, read by §9's check, the TraceArchive and checker::memory_columns_from_log, the independent reading the column builders are held to.
  • TraceArchive, the post-execution snapshot: buffers, log, profile, public values and advice. Its file is two postcard values, five phase sections and then their timings, so the deterministic payload is a byte prefix of it, and only a canonical encoding of self-consistent parts is read back. No proving path reads one; checker::TamperHarness and the retained archived path do (streaming.md §6).
  • CycleProfile, ShardPlan. The profile counts rows per family, cycles for a cycle-owning family (summing to the cycle count) and invocations for a delegation family. trace::plan_shards is ⌈count / height⌉ per family; a window family plans 0 there, its count being the prover's (streaming.md §4).

Auditoren/Programm und Ausführung

Öffentliche Werte und Hilfsdaten

Normative Spezifikationdocs/spec/public-values.mdAls Markdown anzeigen

Zusammenfassung

Wie die Eingaben und Ausgaben einer Ausführung Teil ihrer Aussage werden. Die Seite legt die öffentliche Eingabe, das Journal und die Hilfsdaten (Advice) in Speicherfenster, definiert ihr Layout mit Längenpräfix und die Fensterfamilien, die sie initialisieren, und legt die Bindung dar: Der Digest von Eingabe und Journal geht vor jeder Challenge in das Transkript ein, die Multimenge macht den ersten und letzten Wert der Fenster zu denen der Ausführung, und eine Prüfung an einem zufälligen Punkt bindet diese Fenster an die Bytes der Aussage. Hilfsdaten sind bewusst an nichts gebunden.

Der folgende normative Text wird auf Englisch gepflegt, der kanonischen Sprache der Spezifikation.

How an execution's public input and public output, the journal, are bound to its proof, and what the prover's advice is. All three are regions of guest memory, each initialized by a window family of its own and carried by the memory argument (memory.md).

1 Three regions, no I/O syscall#

region contents chosen by bound by
public input the statement step 10c, to the statement's input (§5)
journal the guest's stores step 10c, to the statement's output (§5)
advice the prover nothing (§6)

There is no I/O syscall: a guest uses ordinary loads and stores, and a provable guest's only ecalls are EXIT and delegation numbers (the retired POSIX numbers: ecall-abi.md §4). A host supplies emulator::GuestIo's input and advice and reads the journal from emulator::Execution::io. A byte-moving syscall would need cross-row constraints tying each transfer row to its buffer and length, which this arithmetization has no place for, while a window is bound by the multiset and one comparison (§5) and asks nothing of the guest. Nor does the guest hash its output: nothing rests on its honesty, and a panic loses nothing it committed.

2 The memory map#

region bytes family windows
public input [0x8000, 0xC000) PUBLIC_INPUT PUBLIC_INPUT_WINDOW = 2, at 2^12
journal [0xC000, 0x1_0000) PUBLIC_OUTPUT PUBLIC_OUTPUT_WINDOW = 3, at 2^12
advice [0x8000_0000, 2^32) ADVICE_WINDOWS from 2^29/h, at the window height h

The full address map is ecall-abi.md §6. The public windows take the upper half of [0, RAM_ORIGIN), 64 KiB that no RAM window family initializes (INIT_TEARDOWN masks window 0's rows below RAM_ORIGIN and no ZERO_WINDOWS id is 0, memory.md §3), so they cost RAM nothing, and [0, 0x8000) stays a hole in which a null dereference cannot balance.

A window's first address is 4·height·id, so the pinned height constants::family::PUBLIC_WINDOW_HEIGHT = 2^12 is what makes the origins windows 2 and 3, 16 KiB each, ending flush against RAM_ORIGIN. It is the ceiling: at the next menu height, 2^14, two windows need 128 KiB, and the one window in the hole is window 0, which would initialize address zero. Anything larger means moving RAM_ORIGIN, which moves every program's load address and shortens every decoded table's pc reach (program.md §5).

program::decode_program assigns that height whatever its caller asks, and the verifier refuses any other, and any RAM window height that would let a zero window reach the public windows (memory.md §3.5).

3 Layout and the length word#

word 0       the payload's byte length
words 1 …    the payload, little-endian, zero-padded to the end of the window

A public window is 2^12 words, so a payload is at most guest_memory::PUBLIC_PAYLOAD_BYTES = 16,380 bytes. verifier_core::public_io_words is the one spelling: the executor seeds the input window with it, the prover commits it and the verifier evaluates it. The length word makes the binding exact: without it [1, 2, 3] and [1, 2, 3, 0] fill the same window. verifier_core::derive_global_phase refuses an input or output longer than 16,380 bytes as Statement, and the executor refuses such an input before the first cycle.

4 The window families#

family id height shards init leaf step 10c holds
PUBLIC_INPUT 12 2^12 exactly 1 M[2] init_value M[2] to input
PUBLIC_OUTPUT 13 2^12 exactly 1 literal 0 M[1] teardown_value to output
ADVICE_WINDOWS 14 h k ≥ 0 M[2] init_value nothing

All three are in every VmConfig and own no cycles. Each public family proves exactly one shard in every statement (memory.md §3.5), so step 10c always runs: an unread input is still the window's initial contents, and an unwritten journal is empty.

The circuits are memory.md §3.3's. PUBLIC_OUTPUT's is ZERO_WINDOWS' byte for byte, whose init leaf writes the literal 0, so no column holds an initial journal (§5). PUBLIC_INPUT's and ADVICE_WINDOWS' initial values are M[2], one execution's values, committed before the memory challenges and bound by no program identity.

All three regions' tuples carry constants::address_space::RAM; which family initializes an address is what makes a word public, advice or heap. A space of their own would need an address-space column, and a gate pinning it, on the memory path of MEM_WORD, MEM_SUBWORD and ATOMICS; under RAM those circuits need nothing for them, their addressing already covering every 4-aligned address below 2^32 (memory-ops.md §2).

5 The binding#

io_digest absorbs the statement's two strings in a transcript of its own (transcript::io_digest):

t ← Transcript::new()
t.append_bytes(PUBLIC_INPUT_STREAM,  input)      tag, byte length, 31-byte limbs
t.append_bytes(PUBLIC_OUTPUT_STREAM, output)
io_digest ← t.sample()                           one raw squeeze

The framing (transcript.md §3) parses back to exactly one ordered pair, and the squeeze is raw, as every digest's is. The guest never computes it. G7 absorbs it before the memory commitments (G8) and challenges (G10) (proof.md §2), so both strings are fixed before any challenge exists.

The multiset. At a window address the init leaf is the only write at timestamp 0, every access consumes a write and produces a strictly later one, and the teardown balances only against the last (memory.md §9). So PUBLIC_INPUT's M[2] holds each word's value before its first access, and PUBLIC_OUTPUT's M[1] its value at the end.

Step 10c of verifier_core::verify_shard_local (proof.md §6). Of a public shard's base claims, which share one point u and each name a column, the verifier takes the one on M[2] (PUBLIC_INPUT) or M[1] (PUBLIC_OUTPUT), refusing its absence as Malformed, and compares it with its own evaluation at u of the multilinear extension of public_io_words(input) or public_io_words(output). A mismatch is MemoryArgument; the shard's opening then holds the claim to the committed column. Column and string are fixed before u is drawn, so a column other than the window passes with probability at most 12/p.

PUBLIC_INPUT's teardown is free: a guest may overwrite its input. PUBLIC_OUTPUT has no init column, and that is the point: with one, a prover could place the journal there at timestamp 0 and the teardown would match without the guest storing a byte.

5.1 The argument, stated plainly#

G7 fixes input and output, and G8 the window columns, before any challenge. Step 10c says the columns are those strings' windows; the multiset says they are the execution's first values in the input window and its last values in the journal window. So the guest found the statement's input in its input window, and the statement's output is what its stores left in the journal window. That rests on no cooperation, hash or register convention of the guest's, and says nothing about advice.

Recursion carries the binding unchanged: a node recomputes io_digest from the windows' words and repeats step 10c over them, and the decider binds the contract's input and output calldata to io_digest (recursion.md §8.1, §9).

6 Advice#

Advice is memory whose initial values the prover chose: ADVICE_WINDOWS initializes [ADVICE_ORIGIN, ADVICE_ORIGIN + 4hk) from an M[2] that nothing binds, not identity, not the statement, not a gate. A guest reads it with ordinary loads.

  • Layout. §3's framing over 1 + ⌈len/4⌉ words (trace::advice_region_words), spelled once by trace::advice_word for the executor and the prover; guest_sdk::advice reads it back.
  • Windows. At the window families' one height h, shard i is window verifier_core::advice_first_window(h) + i, and advice_first_window(h) = 2^29/h is the first window above RAM. Consecutive, they need no list: a statement carries only their count k = ⌈words/h⌉ (trace::advice_window_count), which covers what the host supplied, an untouched word's tuples cancelling. check_memory_windows asks only 2^29/h + k ≤ 2^30/h, the top of the address space, and ZERO_WINDOWS ids stay below 2^29/h (memory.md §3).
  • No advice, no region. Then k = 0 and there is no shard; guest_sdk::advice on such a run is a fatal emulator::EmuError::OutOfBounds.
  • Not read-only. A store there is an ordinary store. Refusing it would need a space selector and a gate on three families' memory path, and would buy nothing: advice is unbound either way.

What a guest owes. A proof says that some advice exists under which the program, given the public input, published the journal; advice that changes the journal unchecked is a value the prover chose. The check is against something the proof binds: a commitment in the public input (guests/public-io, at toy scale, with a position-weighted checksum standing in for a hash), or one the journal publishes. revm-block-stateless publishes the root of the payload it validated and holds its witness to that payload by hashes (ethereum.md §4).

7 The guest's view#

A guest reaches the regions with loads and stores at the constants::guest_memory constants, through guest_sdk::public_input, guest_sdk::commit and guest_sdk::advice, none of which issues an ecall; ecall-abi.md §7 is the API and the guest program manual the walkthrough. Nothing is published at exit, so a guest that panics has published what it committed, and its run is proved like any other.

8 Cost#

  • No address space, transcript message, tag, challenge or statement field; no gate elsewhere.
  • Two 2^12-row shards a statement, five committed columns between them; one h-row shard of three columns per advice window.
  • The native verifier: two 4,096-point multilinear evaluations, 4,095 multiplications each. A recursion node's cost follows the payload instead: it evaluates the payload's words alone, times 1 − r_j for each variable above them (verifier_core::chain::public_value).
  • The guest: nothing at exit; a byte store per journal byte and a word store per commit.

9 Limits#

  • 16,380 bytes each, and no larger window (§2). A journal that grows with the execution has no fixed bound: the mini-block binary's, a 13-byte record plus return data per transaction (ethereum.md §3), holds at most 1,255 transactions, and one record can exceed it. Large outputs belong behind a digest (the stateless binary's journal is 43 bytes), large inputs in advice.
  • The journal is the window's whole final contents. Anything but a length of at most 16,380, that many bytes, then zeros, matches no statement: the executor refuses an oversized length (EmuError::JournalTooLong), and a nonzero byte past it fails step 10c. commit keeps that form; a guest writing the window directly must.
  • Nothing orders the journal's writes, and nothing forces a guest to read its input. The proof binds a window's contents, not its accesses.
  • Read the exit status first. It is x10's final value (memory.md §4): a failed run, a panic included, has a verifying proof and a journal too (§7).
  • A deployed contract fixes both lengths, a decider key being per shape (recursion.md §9).

Auditoren/Beweissystem

Die GKR-Engine

Normative Spezifikationdocs/spec/gkr.mdAls Markdown anzeigen

Zusammenfassung

Die Engine, mit der jeder Familienschaltkreis bewiesen wird. Die Seite definiert das Schichtenmodell mit zeilenweisen und halbierenden Gate-Listen, die Adresse jedes Polynoms, die virtuellen Tabellen, die sieben Gate-Formen und die Gradobergrenze 2, das Schaltkreis-Artefakt mit seinem Serialisierungsformat, seinen vier Gesetzen und seinem Padding-Vertrag sowie den Rückwärtsdurchlauf: was der Aufrufer schuldet, den Transkriptablauf, den Schicht-Sumcheck, warum jede Challenge genau dann gezogen wird, wenn sie gezogen wird, und die Fehler, die ein Verifier zurückgibt.

Der folgende normative Text wird auf Englisch gepflegt, der kanonischen Sprache der Spezifikation.

The layered-circuit model every family circuit is written in, the artifact that carries one, its laws, and the backward pass reducing a circuit's outputs to claims on its committed columns at one point, which the shard's opening discharges (proof.md §5).

crates/constraints is §1–§4; crates/gkr-verify is §5's verifier half and the verifier's helpers for the memory argument (memory.md §3, §4) and LogUp (lookup.md §2, §8). Both are no_std, as verifier-core and the recursion guest build on them. crates/gkr, std and rayon, is the prover half and re-exports gkr-verify. crates/checker enforces §4.2–§4.3 again (circuits.md §3).

1 The layer model#

Layer k, 0 ≤ k ≤ N, N ≥ 1, is w_k columns of n_k variables, indexed as primitives.md §6 fixes. Layer 0 is the committed columns M, W, S in layout order at n_0 = trace_vars, beside the virtual tables the artifact lists (§2.1), which count in no width. Gate list k reads layer k and writes layer k + 1; the top, layer N, is exactly the outputs. A list is row-wise, n_{k+1} = n_k, or halving, n_{k+1} = n_k − 1.

A halving list halves each column of its layer: it writes w_k columns by halving shapes (§3) reading layer-k columns at both children — child 0 is rows [0, h), child 1 rows [h, 2h), h = 2^{n_k−1}, the child bit being the highest variable. An entry may read any column, as a fraction tree's numerator reads its denominator (lookup.md §6), but every column is read (§4.2). Only halving lists hold halving shapes; a halving list is never list 0, has no cached or enforcing entries and needs n_k ≥ 1. Every relation has the one template checker dump prints:

producing, row-wise   L{k+1}[j](x) = Σ_y eq(x, y)·G(layer k at y)
producing, halving    L{k+1}[j](x) = Σ_y eq(x, y)·G(layer k at (y, 0) and (y, 1))
enforcing             0 = G(layer k at y)   for every y ∈ {0,1}^{n_k}

2 Addresses#

constraints::PolyAddress names every polynomial; dumps use its Display notation:

variant notation read by
Memory(i), Witness(i), Setup(i) M[i], W[i], S[i] committed columns list 0, relations, lookups
Virtual(kind) V[row], … virtual tables, §2.1 the same, if virtuals lists it
Inner { layer, offset } L{k}[j] column j of layer k ≥ 1 list k
Cached { layer, offset } C{k}[j] cached entry j of list k, §3.1 list k
Scratch(i) scratch[i] an intermediate of the flat relation list, §4 relations

The scratch bijection maps each scratch[i] to one L{k}[j], covering every inner column once. A committed value needed above layer 1 is carried up by copy gates. M, W and S differ in when they are bound (memory.md §8).

2.1 Virtual tables#

A virtual table is a closed form, evaluated per row by gkr_verify::virtual_at_row and at a point by virtual_at_point, never materialized, committed or claimed. Each form is its table's multilinear extension, so the verifier evaluates what the prover sums (crates/gkr/tests/{lookup,ram_live}.rs check all but V[row]). Wire form: a u32, in table order from 0.

kind notation value at row y closed form at (y_0, …, y_{n−1})
RowIndex V[row] y Σ_{j<n} 2^j·y_j
RamLive V[ram_live] 1 if y ≥ 2^14, else 0 1 − Π_{14≤j<n} (1 − y_j); 0 if n ≤ 14
Range19 V[range19] y mod 2^19 Σ_{j<min(19,n)} 2^j·y_j
Range16 V[range16] y mod 2^16 Σ_{j<min(16,n)} 2^j·y_j
Xor8A V[xor8_a] a = y mod 2^8 Σ_{j<8} 2^j·y_j
Xor8B V[xor8_b] b = ⌊y/2^8⌋ mod 2^8 Σ_{j<8} 2^j·y_{j+8}
Xor8Out V[xor8_out] a ⊕ b Σ_{j<8} 2^j·(y_j + y_{j+8} − 2·y_j·y_{j+8})

14 is constants::memory::RAM_LIVE_BIT (memory.md §3); the range and XOR8 kinds are channel tables (lookup.md §3). Xor8Out's form is multilinear because y ⊕ z = y + z − 2yz is.

3 Gate shapes#

constraints::GateDef is a closed enum. A coefficient is Coeff::Literal(Fr) or Coeff::Challenge(slot), a constants::challenge_slot read from the pass's ExternalChallenges, of degree 0.

tag variant value
0 Linear { terms, constant } Σ c_i·x_i + c_0
1 Product { coeff, left, right } c·x·y
2 MaskIntoIdentity { input, mask } x·m + (1 − m)
3 AffineProduct { left, left_constant, right, right_constant } (Σ a_i·x_i + a_0)·(Σ b_j·y_j + b_0)
4 TreeProduct { input } x(·,0)·x(·,1)
5 Quadratic { constant, linear, products } c_0 + Σ a_i·x_i + Σ b_j·y_j·z_j
6 TreeCross { left, right } p(·,0)·q(·,1) + p(·,1)·q(·,0)

Quadratic spells degree-2 relations, such as a·b + c·d − e·f, that no product of affine forms does. The kernel, gkr_verify::eval_gate, takes one value per operand in GateDef::operands order, a halving shape's each at child 0 then child 1, and is the semantic authority. Both passes reach it through gkr_verify::ResolvedList, crates/checker calls it over the relations, and verifier_core::tape transcribes it for the recursion nodes (recursion.md §7).

3.1 Cached entries and the degree ceiling#

A cached entry C{k}[j] = H is a sub-expression of row-wise list k over its layer's columns, not another cached entry, substituted into the gates of its list naming it, with no table, claim or width. The prover evaluates H at every round node and never binds it: a bound table is the extension of H's values, which for a degree-2 H is not H of the extensions. No registered circuit has one. CircuitArtifact::inline_cached writes a Product with one Linear cached factor as an AffineProduct and refuses any other reference; both prove the same bytes.

Degree is read from the shape after substitution — a column or virtual table 1, a challenge 0, C{k}[j] its expression's, a halving shape 2, a Quadratic its widest term — and validate holds every gate, cached entry and relation to at most 2, so a higher relation is split across layers. With eq multilinear, every round polynomial is then a cubic (§5.3).

4 The circuit artifact#

constraints::CircuitArtifact holds a circuit twice: as layered gates, which the engine proves, and as a flat relation list over M, W, S, V and scratch, which the row-local checks read (circuits.md §3). Law 4 makes them one constraint set. In wire order:

CircuitArtifact = (format_version = 1, coefficient_encoding = 0, trace_vars ≤ 30,
                   memory, witness, setup: [name], virtuals: [(VirtualKind, name)],
                   layers: [LayerSpec], relations: [Relation], lookups: [LookupExpr],
                   scratch: [(name, L{k}[j])], outputs: [L{N}[j]],
                   padding: (row: [Fr], zero_row_valid: bool))
LayerSpec       = (halving, num_vars, width,
                   cached:    [(name, C{k}[j], GateDef)],
                   producing: [(relation, L{k+1}[j], GateDef)],
                   enforcing: [(relation, GateDef)])
Relation        = (name, output: Option<scratch index>, GateDef)
LookupExpr      = (name, channel, selector: PolyAddress, tuple: [GateDef])

validate holds the first three to those values and every name to non-empty [a-z0-9_], unique in the artifact; names mean nothing to the engine. Encoding 0, COEFFICIENT_ENCODING_CANONICAL_LE, is every Fr canonical 32-byte little-endian, and 30 is MAX_TRACE_VARS. outputs orders the top layer as OutputClaims lists it; a relation with an output defines that slot, one without is enforcing; lookups are lookup.md §1's.

4.1 Wire form#

postcard over §4's tuples, hand-written serde: a u32 is a varint, a u8 tag and a bool a byte, an Option a tag byte, a sequence a varint count then its elements, a name a str, an Fr its 32 canonical bytes.

PolyAddress  (tag u8, a u32, b u32): 0 M, 1 W, 2 S, 5 scratch (a = index); 3 V (a = kind);
             4 L, 6 C (a = layer, b = offset); unused fields 0
Coeff        (tag u8, slot u32, value Fr): 0 literal (slot 0), 1 challenge (value 0)
GateDef      (tag u8, split u32, coefficients [Coeff], operands [PolyAddress] in operands() order)
  0 Linear            split 0  c_1..c_t, c_0                 x_1..x_t
  1 Product           split 0  c                             x, y
  2 MaskIntoIdentity  split 0  —                             x, m
  3 AffineProduct     split t  a_1..a_t, a_0, b_1..b_u, b_0  x_1..x_t, y_1..y_u
  4 TreeProduct       split 0  —                             x
  5 Quadratic         split t  c_0, a_1..a_t, b_1..b_u       x_1..x_t, y_1, z_1, …, y_u, z_u
  6 TreeCross         split 0  —                             p, q

CircuitArtifact::from_bytes refuses a format_version other than 1 before decoding the rest, postcard not being self-describing; refuses an unknown tag, a nonzero unused field, a gate with counts its shape lacks and a non-canonical Fr; re-encodes and compares, as postcard admits overlong varints and trailing bytes; never panics or reserves what a declared length asks; and checks no law.

4.2 The laws#

CircuitArtifact::validate runs once where an artifact is built or loaded, never per proof: each constraints constructor panics on a refusal, and verifier_core::VerifyingKey::check applies it to a key's circuits, for prover and verifier (proof.md §7). checker::check_laws enforces Laws 1–4 and the lookup rules again, sharing no code with crates/constraints/src/laws.rs (circuits.md §3).

  1. Locality. Every operand of list k is in range and readable at layer k (§2): a V only if listed, a C{k}[j] only one of list k's own, from a producing or enforcing gate.
  2. Derived width. A list's stored width is its producing count, entry j writes L{k+1}[j], and its stored num_vars is n_k, or n_k − 1 if halving.
  3. Top layer. outputs is a permutation of L{N}[0..w_N).
  4. Single source of truth. Relations and gate entries correspond one to one, a producing entry's relation defining the slot the bijection maps to its output, an enforcing entry's none, and each pair is one polynomial, scratch read through the bijection and cached entries substituted: validate compares normalized expansions, checker evaluations at random points.

validate also refuses, each a ConstraintError naming what broke: §4's bounds, no gate list, padding.row not w_0 long, a virtual kind listed twice, §1's halving rules, degree above 2, a relation reading anything but M, W, S, listed V and existing scratch, a scratch list that is no bijection onto the inner columns or not defined once each, a slot outside constants::challenge_slot, and a relation constructed and then dropped — an inner column below the top the list above never reads, a cached entry no gate names, an enforcing gate whose expansion is zero. Reads are decided on normalized expansions: x − x and 0·x read nothing.

The lookup rules. A lookup's channel is in constants::lookup_channel; its tuple is one expression on a range channel, else 1 to lookup_channel::MAX_TUPLE (7), as wide as its channel's other lookups'; its selector is an in-range committed column some enforcing gate of list 0 holds to booleanity (x − x² up to normal form); and each expression is Linear over in-range committed columns and listed virtual tables, with literal coefficients, unit and constant-free above position 0 (lookup.md says what each protects).

4.3 The padding contract#

The engine gates nothing, an enforcing gate being a zerocheck over the whole cube, so a family switches relations off with its own columns (memory.md §2). On padding.row, a committed row, the row-local scratch values, those of producing relations not at or above a halving shape, make every row-local enforcing relation vanish at every challenge value and row index; zero_row_valid says whether the all-zero row does too. The product-tree clause: where shards have inactive rows, every column the first halving list reads is 1 on padding.row, so padding leaves each product unchanged; the RAM window families (memory.md §3) and the columns a TreeCross reads (lookup.md §6) are exempt. This is completeness, not soundness: a cheating prover's padding rows are its family's gates' business. Nor is padding.row the row a prover writes, multiplicities and setup columns differing; no prover or verifier reads it, and checker::check_padding and checker::check_padding_identity test it.

5 The backward pass#

gkr::forward materializes every layer from the committed columns; gkr::prove proves those values as they stand, one sumcheck::SumcheckProof per transition; gkr_verify::verify replays the schedule, checking, from OutputClaims, one table per output, to BaseClaims or a GkrError. gkr::self_check, naming the first failing gate, row and relation, and gkr::explain_self_check, listing that row's operands, are a debugging hook costing a second forward pass (tools.md §3). Rayon splits rows and row pairs, never lists or rounds: proofs do not depend on the thread count.

5.1 What the caller owes#

  • The base is bound into the transcript before prove or verify, which absorb none of it (proof.md §4 binds a shard's commitments).
  • Each challenge is drawn after every committed column its gates reach is bound, or is derived: a fixed function of such challenges and of statement data bound before them, computed by the verifier. That suffices for GKR; the memory argument needs more (memory.md §8).
  • The artifact has passed validate (§4.2) and is not checked again; on a lawless one the engine may panic, and verify may accept.
  • The prover's inputs have the artifact's shape; it checks none, nor that its values satisfy the gates. Soundness is verify's alone and a cheating prover runs none of this code, so a bad input costs the honest prover only a panic or a failing proof.

5.2 The transcript schedule#

prove and verify run these steps and end in one sponge state; the tags are transcript.md §5's. p is the claim point, v_j the claim on column j of the layer the next list writes.

step op tag message
O1 absorb GKR_OUTPUTS the output tables in output-map order, rows in index order: one message of w_N·2^{n_N} scalars
O2 squeeze ×n_N GKR_OUTPUT_POINT p = r, r_i binding variable i; v_j = tables[i](r) for outputs[i] = L{N}[j]
L1 squeeze GKR_BATCH λ; the claim is c = Σ_j λ^j·v_j
L2 ×n_{k+1}: absorb, squeeze SUMCHECK_ROUND, SUMCHECK_CHALLENGE a round's cubic, then ρ_i, binding variable i
L3 absorb GKR_LAYER_CLAIMS row-wise: L{k}[j](ρ) per j in offset order, layout order at k = 0; halving: L{k}[j](ρ,0), L{k}[j](ρ,1) per j
L4 squeeze, halving only GKR_CHILD τ; p = (ρ, τ); v_j = L{k}[j](ρ,0) + τ·(L{k}[j](ρ,1) − L{k}[j](ρ,0))

L1–L4 run for k = N − 1 down to 0; after a row-wise list p = ρ and v is L3's message. The base claims are layer 0's, in layout order at one point. Every registered circuit halves to a top with no variables (circuits.md §2), so O2 draws nothing and O1 fixes the roots before λ.

5.3 The layer sumcheck#

Transition k proves c = Σ_{y∈{0,1}^{n_{k+1}}} eq(p, y)·S_k(y), where

row-wise   S_k(y) = Σ_j λ^j·G_j(layer k at y) + Σ_e λ^{w_{k+1}+e}·E_e(layer k at y)
halving    S_k(y) = Σ_j λ^j·G_j(layer k at (y, 0) and (y, 1))

G_j writes L{k+1}[j] and E_e, the list's e-th enforcing gate, claims 0: enforcing gates are zerochecks sharing the descending point and its batch. The rounds are primitives.md §7's cubics, run from c, one per variable of layer k + 1, a halving list's two children being separate tables. After L3 the verifier checks claim = eq(p, ρ)·S_k(values), layer-k operands taking L3's values, virtual tables their closed form at ρ, cached entries their expression; with n_{k+1} = 0 there are no rounds and the check is c = S_k(values). A zero claim is legal. gkr::prove_sumcheck and gkr_verify::verify_sumcheck run L2.

5.4 Why it is sound#

Each challenge is drawn after what it protects:

  • r after the outputs, or a prover predicting r claims another table agreeing with the true one there.
  • λ after the claims and p. If some v_j is not the true v̂_j, or some E_e is nonzero on the cube, Σ_j λ^j·(v_j − v̂_j) − Σ_e λ^{w_{k+1}+e}·Ê_e(p) is a nonzero polynomial in λ of degree below w_{k+1} + |E_k|; Ê_e, the extension of E_e's values, is fixed before p is drawn and vanishes there with probability at most n_{k+1}/|Fr|.
  • ρ_i after round i: a wrong cubic agrees with the true one there with chance ≤ 3/|Fr|.
  • τ after both children: a wrong pair's line meets τ ↦ L{k}[j](ρ, τ) in at most one point.

Summed over a registered circuit's transitions at its default height, these stay under 2^14/|Fr|. The random-oracle assumption is architecture.md's.

5.5 Shapes and errors#

Transition k carries n_{k+1} rounds and w_k claims, 2·w_k if halving, so a proof's shape is the artifact's alone (wire form: proof.md §9). verify checks, in order and before touching the transcript, and on a validated artifact never panics on proof or claim data:

GkrError when
MissingChallenge { slot } a gate names a slot not supplied
OutputShape OutputClaims mismatches the output map in count or variables
ProofShape { layer } layer = N: a wrong transition count; else transition layer, lowest first, has a wrong round or claim count
LayerInconsistency { layer } a round or the final check of transition layer fails

One LayerInconsistency covers a wrong descending claim and a violated enforcing gate alike: a batched sum cannot tell them apart, and the proof spends nothing on it. proof.md §6 maps these errors to its classes.

Auditoren/Beweissystem

Schaltkreise

Normative Spezifikationdocs/spec/circuits.mdAls Markdown anzeigen

Zusammenfassung

Die Registry aller 23 Familienschaltkreise, mit Höhe, committeten Spalten, Gates, Lookups, inneren Spalten, Artefaktgröße und Größe des Shard-Beweises jedes einzelnen; wie ein Familienschaltkreis aus Speicherblättern, Lookup-Brüchen, Constraint-Gates, zeilenweisen Bäumen und halbierenden Listen zusammengesetzt wird; und wie der unabhängige Checker die Gesetze, den Padding-Vertrag und die Lookup-Regeln erneut validiert, zusammen mit der Manipulations-Suite, die belegt, dass Fälschungen zurückgewiesen werden.

Der folgende normative Text wird auf Englisch gepflegt, der kanonischen Sprache der Spezifikation.

Every shard is proved by its family's circuit, a constraints::CircuitArtifact in gkr.md's model, fixed by the format, the family and the height. This page lists the circuits and their shapes (§1), how one is assembled (§2) and how crates/checker checks one independently (§3); each family's own page specifies its columns, gates and lookups.

1 The registry#

constraints::family_circuit(family, trace_vars) is the base format's registry, constraints::recursion_circuit the recursion format's, and VmConfig::circuit picks one by format (recursion.md §1.1). Each returns a FamilyCircuit, the artifact and its channel specs (lookup.md §11). A verifying key loads only if its circuits are the registry's at its heights (proof.md §7), and the prover registers the same (§2).

Families 0–6 (constants::family) are the execution families, one executed instruction a row (add-sub.md, jump-branch-slt.md, shift-bitwise.md, mul-div.md, memory-ops.md §3, §4, §6); 7–8 and 12–14 the window families, one memory word a row (memory.md §3, public-values.md §4); 9–11 and 15–17 the delegation families, one invocation a row (delegation-circuits.md §2 to §7, by id); 18–22 the recursion format's (recursion.md §2 to §6).

Shapes at the default height 2^n (constants::family::DEFAULT_HEIGHTS): committed columns, enforcing gates, obligations per channel (TIMESTAMP/RANGE16/GENERIC/DECODER/XOR8), row-wise gate lists (the halving ones are n), inner columns, artifact bytes, and a base-format shard proof's bytes, proof.md §9's layout over the shape:

id family n M W S gates lookups row-wise inner bytes proof
0 ADD_SUB_LUI_AUIPC 22 27 35 7 63 10/4/0/1/0 5 314 72,064 64,764
recursion format 22 27 39 7 75 10/4/0/1/0 5 314 79,077 —
1 JUMP_BRANCH_SLT 22 21 44 10 42 8/11/2/1/0 5 392 76,980 69,436
2 SHIFT_BITWISE 22 21 61 10 48 8/24/6/1/0 6 478 102,837 76,644
3 MUL_DIV 20 21 54 9 54 8/16/2/1/0 6 444 92,640 67,412
4 MEM_WORD 22 31 24 7 33 12/5/0/1/0 5 314 60,383 63,836
5 MEM_SUBWORD 22 31 55 10 53 12/22/1/1/0 6 472 98,846 76,196
6 ATOMICS 20 26 54 9 46 10/19/6/1/0 6 472 101,593 68,468
7 INIT_TEARDOWN 22 2 0 1 0 — 1 46 3,907 36,316
8 ZERO_WINDOWS 22 2 0 0 0 — 1 46 3,418 36,284
9 KECCAK_F 18 208 1,556 0 385 0/210/0/0/1,020 11 5,490 1,900,468 381,100
10 POSEIDON2 8 100 4,092 0 4,248 — 193 2,020 2,056,361 664,780
11 FR_ARITH 8 104 2,576 0 2,701 — 6 142 1,063,214 266,292
12 PUBLIC_INPUT 12 3 0 0 0 — 1 26 2,455 12,556
13 PUBLIC_OUTPUT 12 2 0 0 0 — 1 26 2,338 12,524
14 ADVICE_WINDOWS 22 3 0 0 0 — 1 46 3,535 36,316
15 MOD_MUL 16 104 221 0 125 0/274/0/0/0 10 2,244 550,391 135,220
16 SHA256_COMP 18 104 520 0 119 0/114/0/0/336 10 2,802 845,456 189,988
17 EC_ADD 16 392 1,028 0 637 0/1,110/0/0/0 12 8,772 2,350,670 434,916
18 FIELD_WINDOWS 20 2 0 0 0 — 1 42 2,758 —
19 FR_OP 20 31 31 0 44 0/36/0/0/0 7 370 89,741 —
20 P2_FIELD 18 45 382 0 372 0/58/0/0/0 7 392 294,425 —
21 FIELD_IO 18 43 39 0 24 0/70/0/0/0 8 650 164,713 —
22 FQ_OP 20 48 73 0 38 30/50/0/0/0 7 630 158,326 —

Heights. Both registries return None above MAX_TRACE_VARS = 30, and below the floor lookup.md §3 derives from the family's channels: 19 with TIMESTAMP, else 16 with RANGE16 or XOR8, else 0. A height changes trace_vars, each list's variable count and the number of halving lists, one per variable and as wide as the outputs, and no gate below them: at 2^20 ADD_SUB_LUI_AUIPC has 298 inner columns, 70,974 bytes and a 57,196-byte proof.

Shared circuits. The registries agree on families 1–17; the recursion format's ADD_SUB_LUI_AUIPC is add_sub::recursion_artifact (add-sub.md §2). PUBLIC_OUTPUT's circuit is ZERO_WINDOWS' and ADVICE_WINDOWS' is PUBLIC_INPUT's, byte for byte at one height, and FIELD_WINDOWS' is the zero window at a stride of one cell, all constraints::memory constructors (memory.md §3). Every other family's is its own module's artifact.

2 How a family circuit is assembled#

layer 0        M ‖ W ‖ S in layout order, beside the V tables' closed forms
gate list 0    memory leaves: the read side, then the write side, each padded to a power of
                 two with the literal 1
               per channel, in spec order: (−mult, T + g), then (1, E_l + g) per lookup,
                 then (0, 1) up to a power of two                      (lookup.md §6)
               every enforcing gate
lists 1 … r    row-wise: each tree combines sibling nodes, a product by a·b, a fraction by
                 (n_a·d_b + n_b·d_a, d_a·d_b); a tree already at one node is copied up
lists r+1 …    halving, one per variable: TreeProduct on a product, TreeCross (num) and
                 TreeProduct (den) on a fraction
top            no variables: read_root, write_root, then (num, den) per channel

r is the largest tree's depth, so the circuit has r + 1 row-wise lists; every registered circuit, POSEIDON2 included, ends in a top with no variables. crates/constraints/src/build.rs assembles it, writing the flat relation list and an all-zero padding row, zero_row_valid read off the gates' constants, and validating (gkr.md §4). constraints::memory::assemble gives it the product trees and lookup::channel_trees' fraction trees (lookup.md §11), then runs memory::check_memory (memory.md §8) and lookup::check_discharge: a constructor panics on a refusal, so every circuit that exists has passed them. Its callers:

  • memory::frame_with_channels_artifact(queries, trace_vars, FamilySpec), the execution families: memory.md §2's frame over memory::frame_queries(family), then the family's witness columns after the frame's w + 3, setup columns from S[0], virtual tables, enforcing gates after the frame's, lookups after its 2w gap obligations, and a non-empty channel list;
  • the window constructors (memory.md §3);
  • the delegation and recursion families, every gate in list 0, from constraints::delegation's shared columns, leaves and gates (delegation-circuits.md §1) — but POSEIDON2, which builds its own lists (delegation::Assembly): 192 row-wise lists of rounds beside its product trees, the last holding three gates on the output lanes.

Beyond the frame, each execution family has m_pc as the row's liveness and every other mask held to m_pc times the kinds making that query (<q>_mask_rule, memory.md §2); its decoded row as W columns, bound by decode_row to its table at the row's pc, and decoded_mask_bits, the mask as boolean kind bits, one-hot by the table's domain (lookup.md §10, program.md §6); a next_pc_rule (memory.md §5); a bound on each register value it writes (memory-ops.md §5); and channels ordered TIMESTAMP, RANGE16, GENERIC if read, DECODER.

prover::family_fill(family) is the prover's side: a prover::Fill writes a shard's committed columns but the multiplicities, which trace::build_multiplicities counts. prover::register pairs fill and circuit for each family of a VmConfig (ProverError::Unregistered if either is missing).

3 Checking a circuit independently#

crates/checker's validators enforce the rules again in code sharing nothing with crates/constraints/src/laws.rs, never calling validate. They evaluate a gate only through the kernel gkr_verify::eval_gate (gkr.md §3), so they re-read the rules, not the gates' meaning. Sampled checks use eight pseudo-random points from fixed seeds.

checks
check_laws (check_law1 … check_law4) the four laws, then the lookup rules (gkr.md §4); Law 4 and selector booleanity by evaluation, where validate compares expansions
check_padding, check_padding_identity the padding contract and its product-tree clause, fraction trees exempt
check_lookup_discharge lookup.md §11's discharge rule, gating and compression re-derived
violated_relations, violated_lookups a witness row's row-local relations and range obligations
channel_sums, check_channel_roots each channel's sum and denominator product, folded row by row rather than by a tree, naming every tuple no table row holds; then the circuit's root pairs against them
memory_roots the two roots as products over the rows the halving phase reads
memory_columns_from_log, frame_witness_from_log an execution family's frame columns from the memory event log, where trace builds them from a shard's rows

They do not re-implement check_memory, the copower rule (lookup.md §11), or validate's other construction rules, the degree ceiling among them.

checker::TamperHarness re-proves a statement with witness cells or boundary scalars changed, as an honest prover would prove the changed witness — each channel's multiplicities recounted unless one is what changed or the changed tuple is in no table, changed M columns recommitted in a fresh global commit phase, every shard re-proved — then verifies a shard or the block and asserts the refusal's class (a Lookup's channel too), or that a change breaking nothing verifies. It relies on the prover checking nothing (gkr.md §5), runs on the archived path (streaming.md §6), and carries the delegation anchor's forgeries (checker::assert_anchor_twins_refused, delegation.md §5).

A dump (checker::dump, CLI in tools.md §4) prints the columns by address and name, each list's gates in gkr.md §1's template with their relations, the flat relations over scratch[i], the scratch bijection, outputs, lookups and padding row. Relations are numbered list by list, producing before enforcing; a producing one is define_<column>, an enforcing one bears its gate's name; a node is named for its tree and layer (range16_3_1_num, read_root), a leaf for what it holds (write_pad_0, rd_hi_range_den). A literal below 2^32 prints in decimal, p − k for such a k as -k, any other as 0x and 64 big-endian hex digits; a challenge as its constants::challenge_slot::NAMES entry.

Auditoren/Beweissystem

Das Speicherargument

Normative Spezifikationdocs/spec/memory.mdAls Markdown anzeigen

Zusammenfassung

Das Argument, dass jeder Lesezugriff den letzten Schreibzugriff liefert, über alle Shards einer Aussage hinweg. Die Seite definiert das Speichertupel, die Frame-Spalten, Blätter und Produkte jeder Ausführungsfamilie und die Gadgets, die jede Familie mitführt, die RAM-Fensterfamilien und die Regeln, an die der Verifier sie bindet, die Randwerte von Registern und pc und die eine Abgleichsgleichung, wie das Anhalten festgelegt wird, was den Speicher-Challenges vorausgehen muss, die Bereichsverpflichtungen, die Regeln zur Konstruktionszeit und worauf das Argument beruht, einschließlich der Begründung, warum der Pfad des pc die Programmreihenfolge ohne Zusatzkosten liefert.

Der folgende normative Text wird auf Englisch gepflegt, der kanonischen Sprache der Spezifikation.

Offline memory checking over a whole statement. Each shard's circuit outputs the product of its read tuples and of its write tuples; the verifier checks, once per statement, that all reads times the register and pc finals equal all writes times their initial values. RAM is initialized by window families over fixed address windows; registers and the pc have no rows. The section numbers are the ones the code cites.

1 The tuple#

T(AS, ADDR, TS, VAL) = γ_M + AS + α_addr·ADDR + α_ts·TS + α_val·VAL

The parts are in the order of constants::memory::{PART_AS, PART_ADDR, PART_TS, PART_VAL}. AS, an address-space tag (execution-trace.md §2), is unweighted; a RAM address is a 4-aligned word's byte address. γ_M, α_addr, α_ts, α_val are constants::challenge_slot slots 1–4, MEM_GAMMA to MEM_ALPHA_VAL, drawn once per statement after everything §6.1 lists; slot 5, MEM_WINDOW_CONSTANT, is derived per window shard by the verifier and never read from a proof (§3.3). A gate coefficient is one literal or one slot (gkr.md §3), so α_ts·4·cycle is the term (α_ts, cycle) four times.

Every memory artifact outputs its read product at outputs[READ_ROOT = 0] and its write product at outputs[WRITE_ROOT = 1] (constants::memory), before any channel's roots (lookup.md §6). All tuples of a statement form one multiset, over REG, RAM and PC, one anchor space per delegation type, where a request meets its invocation (delegation.md §5), and the recursion format's FIELD cells (recursion.md §2.1).

2 An execution family's memory subtree#

2.1 The frame columns#

A row of an execution family is one cycle; its accesses are queries, each a read and a write at one address, the write at 4·cycle + Δ (execution-trace.md §1). The query table is constraints::memory::{FRAME_NAMES, FRAME_SPACE, FRAME_DELTA}:

id query space Δ
0 pc PC 0 address 0; reads pc, writes next_pc
1 rs1 REG 1 read-only; an ecall's a7
2 rs2 REG 2 read-only; an ecall's a0
3 load RAM 2 read-only; a load's word
4 ram RAM 3 a store's or an atomic's word
5 rd REG 3 the x0 rule (§2.4)
6 deleg the row's 3 a delegation request's mirror (delegation.md §5)

A family's frame is exactly the queries its instructions make (execution-trace.md §4), in table order: constraints::memory::frame_queries, which crates/trace/tests/memory.rs holds to the union over all 59 instructions. A missing query would leave an instruction's written value unconstrained. No instruction routed to ADD_SUB_LUI_AUIPC touches RAM, and ATOMICS keeps every RAM access at Δ = 3, lr.w included. Window and delegation families have no frame (§3.3; delegation-circuits.md §1).

family queries, in slot order w leaves a side
ADD_SUB_LUI_AUIPC pc rs1 rs2 rd deleg 5 8
JUMP_BRANCH_SLT, SHIFT_BITWISE, MUL_DIV pc rs1 rs2 rd 4 4
MEM_WORD, MEM_SUBWORD pc rs1 rs2 load ram rd 6 8
ATOMICS pc rs1 rs2 ram rd 5 8

Columns are addressed by slot s, a query's position in its family's list:

M[0]                  cycle
M[1 + 5s + f]         slot s's <q>_mask, <q>_addr, <q>_read_ts, <q>_read_value, <q>_write_value
M[1 + 5w]             deleg_space, in the one frame holding deleg
W[s]                  <q>_gap_hi, for s < w
W[w], W[w+1], W[w+2]  rd_inv, rd_is_zero, rd_selected

That is 1 + 5w M columns, plus deleg_space, and w + 3 W columns, the family's own following (circuits.md §2). One deleg query serves every delegation type, so its space is the value of deleg_space, an M column the family pins to its type selectors: a leaf may read no W column (§8). The honest fill (trace::build_memory_columns, trace::build_frame_witness, over a shard's trace::RowSlice) sets a mask to 1 where the row is live and has the query, and every column of an absent query or a padding row to 0.

A frame holds a mask only to booleanity, so on the frame alone a padding row's rd query could rewrite x10, the exit status, after the exit row, and a live row could drop a query or carry one its instruction lacks. Every execution family makes m_pc the row's liveness and its decoder lookup's selector (lookup.md §10), and holds each other mask to m_q = m_pc·uses_q (its <q>_mask_rule gates), uses_q the sum of the row's kind and ecall-type selectors that make the query.

2.2 The leaves#

For the query at slot s with mask m, space AS and in-cycle slot Δ:

read_<q>    m·T(AS, addr, read_ts, read_value) + 1 − m
write_<q>   m·T(AS, addr, 4·cycle + Δ, write_value) + 1 − m

Each is one flat Quadratic of gate list 0, built from the unmasked tuple, a Linear whose AS and Δ terms sit on m (constraints::memory::read_tuple is the read one): constant 1; linear terms (γ_M, m), (−1, m), (AS, m) and, on the write side, (α_ts, m) Δ times; every other term multiplied by m, as is deleg's AS, the product (1, deleg_space, m). At m = 0 a leaf is 1 whatever its columns hold, at m = 1 the tuple, and it is one or the other only at a boolean m (§2.4).

2.3 The product#

Each side is padded to w rounded up to a power of two with read_pad_<i> and write_pad_<i>, the literal 1, reading no column. Row-wise Product lists reduce each side to one value a row, and trace_vars halving lists of TreeProduct multiply the rows (gkr.md §1), so the two roots are the products of the shard's read and write tuples. A padding row has every mask 0 and so every leaf 1, the padding contract's product-tree clause (gkr.md §4). The family's channel trees share the layers (circuits.md §2).

2.4 The gadgets every execution family carries#

Gate list 0's first enforcing gates, in this order, and the circuit's first 2w obligations:

<q>_mask_boolean       m − m·m = 0                       every query
<q>_writes_back        write_value − read_value = 0      rs1, rs2 and load, where held
rd_is_zero_inverse     addr·rd_inv + z − m = 0           on rd; z = rd_is_zero
rd_is_zero_at_nonzero  addr·z = 0
rd_is_zero_boolean     z − z·z = 0
rd_write_masked        write_value − sel + z·sel = 0     sel = rd_selected

gap_hi_<q>   TIMESTAMP, selector m:   hi                                   hi = <q>_gap_hi
gap_lo_<q>   TIMESTAMP, selector m:   4·cycle + δ_q − read_ts − 2^19·hi        δ_q = Δ − 1; δ_pc = −4
  • Booleanity. At m = −1 a pc query's leaves are each −T(REG, …): one sign flip a side, so the products balance and the pc access reads as a register access.
  • Write-back. Without it a read of x0 could write 5 there.
  • x0. The first two rd gates (constraints::gadgets::is_zero) make z = m·[addr = 0] and the last write_value = (1 − z)·sel: every write to x0 writes 0, whatever the family computed into sel, and with write-backs and x0's init 0 every read of it returns 0. The boundary's final x0 = 0 (§4.1) pins only its last write: a write of 5, a read of 5 and a write of 0 would otherwise balance.
  • Gap. Both chunks below 2^19 put gap = 4·cycle + δ_q − read_ts in [0, 2^38), so read_ts < 4·cycle + Δ as integers, every timestamp being a canonical integer by §4.2's count. The pc query's δ = −4 puts a row's pc write at least 4 after the one it reads, so consecutive rows' timestamps never interleave (§9). The frame's construction asserts two obligations per query.

3 RAM windows#

3.1 Geometry#

Window w at height h = 2^n covers the bytes [4h·w, 4h·(w + 1)), its row y being the word at 4h·w + 4y; the windows tile [0, 2^32) from 0. Ordinary RAM is [RAM_ORIGIN, ADVICE_ORIGIN) = [2^16, 2^31) (trace::in_ram), ending where window N = 2^29/h begins (verifier_core::advice_first_window). Window 0's rows y < 2^14 (constants::memory::RAM_LIVE_BIT) lie below RAM_ORIGIN at every height, and INIT_TEARDOWN masks them (§3.3).

3.2 The window families#

A window family's shard initializes and tears down one window; its rows are addresses.

region family id windows init value
[0, 0x8000) none: a hole
[0x8000, 0x10000) PUBLIC_INPUT, PUBLIC_OUTPUT 12, 13 2 and 3 at their pinned 2^12 the statement's input; 0 (public-values.md §4)
[RAM_ORIGIN, 4h) INIT_TEARDOWN 7 0, one shard S[0], the image column
[4h, 2^31) ZERO_WINDOWS 8 the listed w_1 < … < w_k in [1, N − 1] 0
[2^31, 2^32) ADVICE_WINDOWS 14 N … N + k_a − 1 M[2], bound to nothing (public-values.md §6)
FIELD cells FIELD_WINDOWS 18 0 … k_f − 1 0 (recursion.md §2.2)

INIT_TEARDOWN, ZERO_WINDOWS and ADVICE_WINDOWS share the window height h, 2^22 by default (§3.5). An unlisted RAM window is initialized by nothing. Every statement proves window 0 and, in practice, the stack's window N − 1, the initial sp being ADVICE_ORIGIN: two h-row shards however small the program, besides the public pair.

3.3 The artifacts#

INIT_TEARDOWN   image_window_artifact   M[0] teardown_ts, M[1] teardown_value, S[0] init_value
  read    live·(WC + α_addr·4·row + α_ts·M[0] + α_val·M[1]) + 1 − live     live = V[ram_live]
  write   live·(WC + α_addr·4·row + α_val·S[0]) + 1 − live
ZERO_WINDOWS, PUBLIC_OUTPUT    zero_window_artifact     M[0], M[1]
  read    WC + α_addr·4·row + α_ts·M[0] + α_val·M[1]        write   WC + α_addr·4·row
PUBLIC_INPUT, ADVICE_WINDOWS   value_window_artifact    M[0], M[1], M[2] init_value
  read    as above                                          write   WC + α_addr·4·row + α_val·M[2]
FIELD_WINDOWS   field_window_artifact: zero_window_artifact over α_addr·row, one cell a row

All are constraints::memory constructors: one leaf a side, then n halving lists; no witness column, enforcing gate or lookup; every init timestamp the literal 0; row is V[row]. V[ram_live] is [y ≥ 2^14], boolean on the cube by construction (gkr.md §2). The window enters only through WC, so one artifact serves every window:

WC = γ_M + RAM + α_addr·4h·w        gkr_verify::window_challenges
WC = γ_M + FIELD + α_addr·h·w       gkr_verify::field_window_challenges

w is verifier_core::shard_window's: 0 for INIT_TEARDOWN, the list's i-th id for ZERO_WINDOWS shard i, 2 and 3 for the public pair, N + i for advice shard i, i for field shard i. An init column is S where program identity binds it and M where it is one execution's, committed before the challenges (§6.1). A window shard has no inactive rows.

3.4 The columns a prover fills#

trace::init_windows(state, h) is ZERO_WINDOWS' list: the distinct ⌊a/4h⌋ over touched words a of ordinary RAM, ascending, without 0. A public or advice word is a RAM tuple too, and a zero window over it would be its second init row. trace::build_init_teardown_columns fills INIT_TEARDOWN and ZERO_WINDOWS, and trace::build_value_window_columns the value windows with their M[2], from the last-access tables (trace::MemoryState):

row y, a = 4h·w + 4y teardown_ts teardown_value
w = 0, y < 2^14 (masked) 0 0
a touched its last write's timestamp its last write's value
a untouched 0 its init value

An untouched row's two tuples are equal and cancel. The image column, program::image_init_column(image, h), has row y = ProgramImage::initial_word(4y): the word assembled byte by byte from file-backed bytes, 0 elsewhere, which is the trace's initial RAM value too. decode_program refuses an image with a file-backed byte at or above 4h (ProgramError::ImageOutsideWindow): it would sit in a zero window, read as 0, bound by nothing.

3.5 The verifier's window rules#

verifier_core::check_memory_windows, step 2 of derive_global_phase, before the global transcript (program::check_memory_windows wraps it):

rule why
INIT_TEARDOWN, ZERO_WINDOWS, ADVICE_WINDOWS at one height h a lower zero-window height would re-initialize image words; an advice height of its own is a grid advice_first_window(h) does not describe
PUBLIC_INPUT, PUBLIC_OUTPUT at PUBLIC_WINDOW_HEIGHT = 2^12 the height places their windows (public-values.md §2)
4h ≥ PUBLIC_OUTPUT_ORIGIN + PUBLIC_WINDOW_BYTES = 0x10000: h ≥ 2^16 on the menu the public windows lie in window 0's masked rows, out of every zero window's reach
one shard each of INIT_TEARDOWN, PUBLIC_INPUT, PUBLIC_OUTPUT (public-values.md §4 for the pair)
one id per ZERO_WINDOWS shard, strictly increasing, in [1, N − 1] disjoint windows; id 0 is unmasked over [0, RAM_ORIGIN); N up is advice
N + k_a ≤ 2^30/h, k_a the advice shard count advice ends by 2^32; it needs no list, starting where the zero ids stop
k_f·h ≤ 2^32 field cells recursion.md §2.2

The first three are verifier_core::window_height, which VmConfig::from_bytes runs too: a config breaking them does not decode.

4 The register and pc boundary#

Registers and the pc have no rows: the verifier multiplies in their initial and final tuples, once per statement. Rows for them would repeat the init tuples in every shard holding them, and a stale read would balance against the copy.

4.1 The boundary scalars#

The statement carries 64 scalars, gkr_verify::BoundaryFinals, absorbed as one MEMORY_BOUNDARY message in this order (verifier_core::boundary_scalars):

positions
0–31 t_0 … t_31 x_r's final timestamp: its last query's write, 0 if never queried
32 t_pc the pc's: the exit row's pc write
33–63 v_1 … v_31 x_r's final value, 0 if never queried

The final values of x0, 0, and of the pc, HALT_PC, are constants, not carried. PublicInputs::from_bytes refuses t ≥ 2^38 or v ≥ 2^32, and verify_global_memory re-checks the timestamps and holds v_10 to the exit status; no other register carries a public value. t_pc is not a cycle count: the pc's timestamps increase but need not be consecutive. trace::build_boundary_finals(state) is the fill.

4.2 The factors and the reconciliation#

W_b = ∏_{r=0}^{31} T(REG, r, 0, 0) · T(PC, 0, 0, entry_pc)
R_b = T(REG, 0, t_0, 0) · ∏_{r=1}^{31} T(REG, r, t_r, v_r) · T(PC, 0, t_pc, HALT_PC)

∏ read roots · R_b  =  ∏ write roots · W_b  ≠  0       over every shard of the statement

entry_pc is the verifying key's (§6.2). gkr_verify::boundary_factors evaluates each tuple through gkr_verify::eval_gate on the circuits' own tuple gate, read_tuple of pc or rs1, so the boundary and the circuits cannot disagree on the parts; gkr_verify::reconciles is the equation, which verifier_core::verify_global_memory runs once per statement (proof.md §6). A shard's roots are its GKR outputs, held to the statement's entry by its own verification.

The count. Read each query as an edge from its read tuple to its write tuple. Inits are only written and finals only read, so a balanced multiset is paths from inits to finals plus loops. An edge advances the timestamp by an integer in [1, 2^38 + 3] (the gap plus the query's least advance, 4 at the pc and 1 elsewhere), so a loop needs more than p/(2^38 + 3) > 2^215 edges, and a statement has fewer than 2^67 tuples: under 2^32 shards a family (a u32 count), 23 families, at most 2^22 rows (the menu's top), at most 196 tuples a row (EC_ADD's 97 frame words and its anchor, both sides), and 66 boundary tuples. So nothing loops: every path starts at an init at timestamp 0 and ends at a final, and every timestamp on it is an integer below 2^105.

5 Halting#

constants::memory::HALT_PC = 1. The exit row, ecall with a7 = 93, writes next_pc = HALT_PC instead of its fall-through (execution-trace.md §6), and R_b fixes the pc's final value to it. Nothing else writes it: HALT_PC is odd, every other next_pc even, and "odd" is a constraint only where a family makes it one.

  • A family copying the decoded fall-through, which is even, holds next_pc − decoded_next_pc = 0 and needs no bound.
  • JUMP_BRANCH_SLT, the one family computing a pc, range-checks every next_pc it writes even; otherwise a jalr whose rs1 + imm is 1 could write HALT_PC (jump-branch-slt.md).
  • ADD_SUB_LUI_AUIPC writes HALT_PC on its exit row alone (add-sub.md).

HALT_PC is below RAM_ORIGIN, so no decoded-table row claims it and no live row reads it (lookup.md §10). The pc's path therefore ends with the exit row's write, consumed by the final read. With a free final pc every prefix of an execution would balance.

6 Binding#

6.1 What precedes the memory challenges#

The four challenges are squeezed once per statement, at the end of the global transcript (proof.md §2 is the schedule), after everything a tuple or the reconciliation reads, because what is chosen after them can be solved for:

  • every shard's M commitments, every column a leaf may read but S and V (§8);
  • program identity, fixing entry_pc and the image column (§6.2), and the SRS digest, fixing the generic table's S columns (proof.md §3);
  • the shard counts and MEMORY_WINDOWS, the zero-window ids, fixing every window shard's addresses through WC: a list chosen afterwards is a union over up to 2^(N − 1) lists, 2^127 at h = 2^22 and no bound at all at 2^20;
  • io_digest, fixing the public windows' contents (public-values.md §5);
  • last, the 64 boundary scalars: a final value chosen afterwards reconciles any trace, v_r = (target − γ_M − REG − α_addr·r − α_ts·t_r)/α_val.

The roots are not absorbed: each shard's GKR proof binds its own.

6.2 The image column and the entry pc#

Program identity (program.md §8 is the recipe) binds INIT_TEARDOWN's one setup commitment, the image column's, and entry_pc, under PROGRAM_ENTRY. Recomputing identity binds a commitment, not the column a proof reads; the INIT_TEARDOWN shard's batched opening closes that by taking S[0]'s commitment from the verifying key, the list identity is recomputed over (proof.md §5, §7). Without it a statement over another image, with a trace consistent with that image, would verify. Without entry_pc in identity, a key carrying the registered identity beside another entry pc would verify an execution starting elsewhere. Identity binds nothing an execution chooses: no shard count, window list, public or advice word.

7 Range obligations#

A range obligation holds where its selector is 0 or its one expression is below its channel's bound (lookup.md §1, §3). Every circuit bounds a value one way:

  • a 32-bit value v: a witnessed high halfword h and RANGE16 obligations on h and on v − 2^16·h, under the row's selector, and no gate;
  • a result r = e mod 2^32 of an exact 0 ≤ e < 2^33: a witnessed wrap, the gates wrap − wrap·wrap = 0 and e − r − 2^32·wrap = 0, and r bounded as above; a wider carry is a family's own construction;
  • a timestamp gap: two 19-bit TIMESTAMP chunks, no wrap (§2.4). Delegation and recursion families decompose theirs their own way (delegation-circuits.md §1).

8 Construction-time rules#

constraints::memory::check_memory refuses, naming the gate, a memory artifact with:

  1. provenance: a gate or output whose cone both names a memory slot (1–5) and reads a W column, computed forward with two flags a column, so a tuple times a copy of a W column two layers up is refused too;
  2. a root over W: outputs[READ_ROOT] or outputs[WRITE_ROOT] whose cone reads a W column at all, slot or not, which rule 1 does not see;
  3. a memory slot over anything but M, S and V: a gate carrying one reads no W, inner or cached column;
  4. an unconstrained mask: a leaf — a producing Quadratic of gate list 0 with constant 1 and a slot-weighted linear term — whose mask, that term's operand, is an M, W or S column with no m − m·m enforcing gate in gate list 0, or a virtual column but V[ram_live].

It runs beside CircuitArtifact::validate, whose laws it assumes, wherever a memory artifact is built (constraints::memory's assembly panics on a refusal) and in VerifyingKey::check. A W column is committed in a shard's own transcript, after the memory challenges, so a tuple or root over one is chosen after them and balances any trace. M columns precede the challenges and V columns are closed forms; S columns are admitted because they precede them too, bound by identity or, for the generic table, by the SRS digest (§6.1).

9 What the argument rests on#

Both sides of §4.2 are products of linear forms in (γ_M, α_addr, α_ts, α_val), one per distinct tuple, every tuple fixed before those are drawn (§6.1, §8). By Schwartz–Zippel they agree on unequal multisets with probability at most N/p, N < 2^67 (§4.2), and on equal ones §4.2's count gives:

  • One init per address of REG, PC, RAM and FIELD: the 33 boundary inits once per statement, and §3.5's windows, disjoint and of one height. A second init would let a stale read balance. An anchor space has no init: each invocation's answer, stamped 0, starts a path one request long (delegation.md §5).
  • Coverage. Every query lies on a path from an init, so nothing reaches an address no family initializes: the hole [0, 0x8000), where a null dereference does not balance, an unlisted window, a register above x31, a pc address but 0. A query reading its own write would balance with no init; the gap forbids it.
  • Consistency per address: on its one path every read returns the write before it, and the final tuple holds the last.
  • Initial values: the image's, by §3.4's refusal and §6.2's opening of S[0]; 0 in every zero window and the journal; the statement's input in its window (public-values.md §5). Advice is bound to nothing by design (public-values.md §6).
  • Order across rows, shards and families. The pc's path runs from T(PC, 0, 0, entry_pc) through every live row of every execution family, each m_pc = 1 row one edge, to the exit row (§5). That is pc continuity; it orders the rows by their pc writes 4·cycle, which are therefore distinct, so no cycle is proved twice. Nothing else carries it: there is no per-shard pc chaining, and a shard's time window ties to no row (proof.md §8).

Per address, the order is timestamp order, and it is program order: the pc query's gap puts consecutive pc writes at least 4 apart (§2.4), so each cycle's four timestamps precede the next cycle's whatever value cycle takes, and a row never reads an address before its predecessor's write there. An invocation rides its requesting row's cycle (delegation.md §5) and is ordered with it.

Auditoren/Beweissystem

Lookups

Normative Spezifikationdocs/spec/lookup.mdAls Markdown anzeigen

Zusammenfassung

Wie Lookup-Verpflichtungen bewiesen werden: eine LogUp-Identität pro Tabelle und Shard, summiert über einen Bruchbaum im eigenen GKR-Durchlauf des Schaltkreises und an dessen Wurzel geprüft. Die Seite definiert, was ein Kanal behauptet, seine Challenges, die fünf Tabellen, wie eine abgeschaltete Zeile ein neutrales Tupel nachschlägt, den Nenner, den Bruchbaum, Multiplizitäten, die Wurzelprüfung, die gepackte generische Tabelle, den Decoder-Kanal, der jede Zeile an das Programm bindet, die Konstruktionsregeln und worauf das Argument beruht, einschließlich der Begründung, warum jeder Schlüssel durch seine Familie beschränkt sein muss.

Der folgende normative Text wird auf Englisch gepflegt, der kanonischen Sprache der Spezifikation.

How a circuit's lookup obligations are proved: per shard, by one LogUp channel per table, each summed by a fraction tree inside the circuit's own GKR pass and checked at its root.

1 What a channel claims#

A lookup is LookupExpr { name, channel, selector, tuple }: a channel of constants::lookup_channel, a committed M, W or S column as selector, and a tuple of Linear expressions with literal coefficients over committed columns and the circuit's virtual tables. It holds on a row where the selector is 0, or

  • on a range channel, where its one expression's canonical integer is below 2^BITS[channel];
  • on a table channel, where its tuple is a row of the channel's one table: 1 to MAX_TUPLE = 7 expressions, the same number for every lookup of the channel.

memory.md §7 is the convention range obligations follow. A channel discharges all of a shard's lookups on it as one identity over the shard's rows y:

Σ_y Σ_l 1/(E_l(y) + g)  −  Σ_y mult(y)/(T(y) + g)  =  0

E_l(y) is lookup l's gated tuple (§4) and T(y) the table's row y, both compressed by β (§5); mult is the channel's multiplicity column (§7). A range table is the one column [0, 2^BITS).

2 The challenges#

slot challenge_slot value
6 LOOKUP_G g, drawn
7 LOOKUP_BETA β, drawn
8–12 LOOKUP_BETA_2 … LOOKUP_BETA_6 β^2 … β^6, derived
13 LOOKUP_DECODER_NEUTRAL g − Σ_{j<W} β^j, derived; W the decoder tuple's width

g and β are shard-local: the shard's transcript draws them, in that order under the tag LOOKUP_CHALLENGE (33), right after absorbing its witness commitments, multiplicities included (proof.md §4). M columns are committed in the global transcript the shard is seeded from and S columns are bound by identity or the SRS digest, so every column a channel reads is fixed before either challenge exists.

β^0 is the literal 1, so a one-column tuple names no slot. A gate coefficient is one literal or one slot (gkr.md §3), so each higher power is a slot of its own, computed by the verifier and never read from a proof (gkr_verify::insert_lookup_challenges, which reads W off the artifact's decoder lookup).

Selectors are boolean: CircuitArtifact::validate refuses a lookup whose selector no enforcing gate of gate list 0 holds to s − s·s = 0. The selector multiplies the tuple inside the denominator (§5), so a channel proves the gated tuple s·(e + o) + n is a table row, which is the obligation only at s ∈ {0, 1}. At any other s a scaled tuple is looked up instead: on a range channel, s = t·e⁻¹ lands any nonzero e on any table value t.

3 Tables#

channel id kind table, at row y width table_vars
TIMESTAMP 0 range V[range19]: y mod 2^19 1 19
RANGE16 1 range V[range16]: y mod 2^16 1 16
GENERIC 2 table, committed the packed table (§9) 3 0
DECODER 3 table, committed the family's decoded table (§10) 7 or 6 0
XOR8 4 table, virtual V[xor8_a], V[xor8_b], V[xor8_out]: y's low two bytes and their XOR 3 16

A virtual table is a closed form of the row index, never committed: the verifier evaluates its multilinear extension where the GKR pass ends (gkr_verify::virtual_at_point, gkr.md §2). Each is a weighted sum of the row's bits but V[xor8_out], Σ_{j<8} 2^j·(y_j + y_{j+8} − 2·y_j·y_{j+8}), which is exact because y ^ z = y + z − 2yz is multilinear. So XOR8 costs no commitment and nothing in the SRS digest.

constraints::lookup::table_vars is the fewest variables at which a table is complete: BITS for a range channel, 16 for XOR8, 0 for a committed table, a setup column at the circuit's own height. Below it a virtual table holds only part of its range, which costs completeness, not soundness. family_circuit returns None below the largest table_vars of a family's channels, so a key naming such a height fails to load (proof.md §7). On the height menu (program.md §7) a family carrying TIMESTAMP is at 2^20 or more, and one carrying RANGE16 or XOR8 at 2^16 or more. The packed table needs 2^18 rows (§9), and every family that reads it carries TIMESTAMP. Above table_vars a table repeats, which §7 makes harmless.

XOR8's tuple is three wide so that membership bounds each entry to [0, 256) on its own; a packed key x + 256·y would bound neither, (x, y) and (x + 256, y − 1) compressing alike. Every other bitwise operation on bytes is a linear form over its results (delegation-circuits.md §1).

4 Gated keys#

A lookup expression is evaluated on every row, so the selector sends a row whose key means nothing to a neutral tuple, which is a real table row:

gating channels gated position j neutral tuple
NoOffset TIMESTAMP, RANGE16, XOR8 s·e_j all zero
ZeroEntry GENERIC s·(e_0 + 1), then s·e_j the all-zero ZeroEntry row
MinusOne DECODER s·(e_j + 1) − 1 −1 in every column, a padding row

The + 1 keeps every real key of the packed table at 1 or above, so no real entry is the all-zero tuple a switched-off row looks up. A range table needs no offset, 0 being in range, and one would push 2^BITS − 1 out of it; XOR8's (0, 0, 0) is a true entry. A decoded table has no all-zero row, pc 0 being a valid pc, and its MINUS_ONE padding rows (program.md §5) are the neutral entry.

Each key a table channel looks up is bounded by the family that looks it up, because a channel proves membership and nothing more. A selected row whose key expression is −1 gates to the ZeroEntry, and an unbounded key reaches any sub-table of the packed table: an AND key a + AND_BASE with a unbounded lands on a U16GetSign row and proves a false AND. Families bound their keys with RANGE16 obligations or build them from bounded columns (shift-bitwise.md §3 and the other family pages); the decoder's key is §10's.

5 The denominator#

With s the selector, e_j = Σ_i c_{j,i}·x_{j,i} + k_j, and §4's offset o_j (1 or 0) and neutral value n_j (−1 or 0):

E + g  =  Σ_j β^j·(s·(e_j + o_j) + n_j)  +  g
       =  Σ_j β^j·s·e_j  +  Σ_j β^j·o_j·s  +  (g + Σ_j β^j·n_j)
T + g  =  Σ_j β^j·t_j  +  g

E + g is one Quadratic (constraints::lookup::row_denominator): each term of e_j the product (β^j·c)·s·x, each offset the linear term β^j·o_j·s, and the bracket the slot LOOKUP_G or, for the decoder, LOOKUP_DECODER_NEUTRAL. β^j·c is one coefficient only where β^0 = 1 makes it a literal or c = 1 makes it the slot, so position 0 takes any literal coefficients and constant and every later position weights its columns by 1 with no constant. T + g is one Linear over the table's columns (table_denominator).

6 The fraction tree#

A channel's leaf level is (num, den) pairs of gate-list-0 columns, P = (L + 1).next_power_of_two() of them for L lookups:

leaf num den
the table, first −mult T + g
each lookup, in artifact order 1 E_l + g
padding, up to P 0 1

Row-wise gate lists add sibling pairs, (n_a·d_b + n_b·d_a, d_a·d_b), until each row holds one pair; a tree shallower than the circuit's deepest copies itself up. Then trace_vars halving lists add the rows' pairs, TreeCross writing the numerator and TreeProduct the denominator (gkr.md §3). The circuit's outputs are the memory argument's read and write roots, then each channel's (num, den) in the order of its channel specs (crates/constraints/src/build.rs).

A channel costs 4P − 2 inner columns to reduce a row, 2 more per copy-up layer and 2 per halving list, and one committed column. P doubles each time L reaches a power of two.

The padding clause (gkr.md §4) asks a padding row to feed 1 into every product tree. A fraction tree is exempt: its identity is (0, 1), and a padding row is not idle in a channel but looks up the neutral tuple, which the multiplicity counts. checker::check_padding_identity exempts every column a TreeCross reads.

7 Multiplicities#

Each channel has one multiplicity column, a committed W column; a circuit's are its last W columns, in channel order. Row t counts the (row, lookup) pairs of the shard whose gated tuple is table row t's, switched-off rows included. A tuple at several table rows is credited to the lowest; every other copy holds 0 and contributes 0/(T + g). The count is over raw gated tuples, the column being committed before g and β exist (trace::build_multiplicities, which refuses a tuple no table row holds: the honest prover cannot balance it).

No gate or range check constrains the column, and soundness needs none. If a gated tuple v is in no table row, the left side of §1's identity, as a rational function of g, has a pole at −v whose residue is the number of lookups producing v: a positive integer below p, whatever the column holds.

8 The root check#

accept  iff  num = 0  and  den ≠ 0

on each channel's root pair, at step 9 of proof.md §6 (gkr_verify::channel_holds); a failure is VerifyError::Lookup { channel }. The GKR pass absorbs the pair before its first challenge and proves it (gkr.md §5). den is the product of every leaf denominator, and num = 0 means the sum vanishes only where den ≠ 0: one leaf (0, 0) — a table row whose T + g vanishes, counted 0 — makes the root (0, 0) whatever the other leaves hold. With g drawn after the columns that has probability at most fractions/|Fr|, and den ≠ 0 makes it a refusal.

9 The generic table#

One committed table of constants::generic_table::WIDTH = 3 columns, a key and two values, packing three sub-tables under disjoint key ranges (program::lookup_tables::generic_table):

row 0                      ZeroEntry    (0, 0, 0)
rows 1 ..= 2^16            AND          (AND_BASE + a + 1,    b,        a & b)        a, b < 2^8
rows 2^16+1 ..= 2^17       U16GetSign   (SIGN_BASE + h + 1,   h >> 15,  0)            h < 2^16
rows 2^17+1 ..= 2^17+32    ShiftPowers  (SHIFT_BASE + s + 1,  2^s,      2^(31 − s))   s < 32
rows above                 zero

AND_BASE = 0, SIGN_BASE = 256 and SHIFT_BASE = 65,792 put the keys at 1..=256, 257..=65,792 and 65,793..=65,824; a lookup's key expression is x + BASE, and the gating adds the 1. U16GetSign serves every sign an execution family computes, AND the bitwise operations of SHIFT_BITWISE and ATOMICS, ShiftPowers the shifts. The copower 2^(32 − s) is stored halved (SHIFT_COPOWER_BITS = 31), 2^32 not fitting a u32 column, and the two gates that read it carry the factor 2 (shift-bitwise.md §4). 131,105 rows in all (GENERIC_ROWS).

Its commitments are a constant of the ceremony. A Mercury commitment reads the evaluation table as coefficients (mercury.md §2) and the table is zero past its entries, so over 2^n rows it commits to the same three points for every n ≥ 18; generic_commitments(srs) computes them at 2^18 (GENERIC_LOG_HEIGHT). Every verifying key carries them once, as VerifyingKey::generic_table, whether or not a family reads the channel, and its SRS digest covers them (proof.md §3); program identity does not. A circuit that reads GENERIC names the table as its three setup columns after identity's (FamilyCircuit::reads_generic_table), and a shard's opening checks them against the key's points (proof.md §5).

10 The decoder channel#

The DECODER table is the family's decoded table, program::lookup_tuple(family)'s columns, as its first setup columns at its height, row i holding pc 2i (program.md §5); program identity commits them (program.md §8). Each execution family makes one lookup on it, imm absent for MUL_DIV and ATOMICS:

decode_row    selector m_pc    tuple (pc read value, next_pc, rs1, rs2, rd, [imm], extra_mask)

The key is the frame's own pc read (memory.md §2), so the cycle itself is bound to the program; the rest are the row's decoded columns, which the family's other gates read. The selector is the row's liveness, so a padding row looks up the MINUS_ONE tuple, which every decoded table holds, being taller than its last instruction.

The family's decoded_mask_bits gate ties the packed mask to boolean kind bits. That the bits are one-hot, and that a live row is an instruction at all, is the table's domain: its live rows hold one-hot masks and its padding rows −1, which no sum of kind bits reaches. Boolean columns looked up one by one would lose this: booleanity admits any subset of bits, the empty one included, and an all-zero mask makes every gate a kind selects vacuous.

11 Construction rules#

CircuitArtifact::validate enforces §1's form and widths, §2's selector rule and §5's coefficients wherever an artifact is built or loaded (gkr.md §4). When a circuit is assembled, constraints::lookup asserts that a channel has a lookup, that its multiplicity is a W column, that every lookup has its table's width, and that a range channel's table is the one its bound names (range_table) with BITS ≤ trace_vars; constraints::memory::frame_with_channels_artifact refuses an empty channel list, which would leave a frame's gap obligations discharged by nothing.

The discharge rule, constraints::lookup::check_discharge, at assembly and at every key load (VerifyingKey::check): every lookup is the denominator of exactly one gate-list-0 column, its numerator 1 directly before it; no column is two lookups'; each channel's (−mult, T + g) appears once. It matches by normalized expansion inside the cone below the channel's own root pair, so an obligation or table fraction in another channel's tree is refused, and the two range channels, which gate alike, are not confused. Which output pair is whose root, which columns are a table and which counts it is not in the artifact but in its ChannelSpecs, which a key carries in FamilyCircuit::channels and must hold as the registry's (circuits.md §1). checker enforces this rule and the lookup rules a second time, with code of its own (circuits.md §3).

The copower rule, constraints::lookup::check_copowers, run by every constructor that bounds a column through a copower. A bound x < p written as x·p′ < 2^32, p·p′ = 2^32, bounds nothing alone: p′ is a unit of Fr, so x = s·p′⁻¹ ranges over a coset of 2^32 values. Each such x therefore also carries a direct RANGE16 bound, as a halfword or as a high chunk and a remainder, under the same selector.

12 What it rests on#

  • Every gated tuple is a table row: §1's identity over challenges drawn after every column it reads, boolean selectors (§2), both root conditions (§8) and the GKR pass. The error is at most fractions/|Fr| for g, plus looked-up tuples × table rows × (width − 1)/|Fr| for a β collision: below 2^−190 at every menu height.
  • A lookup answers from its own sub-table: one width per channel (§11), disjoint key ranges and the + 1 (§9), and its family's bound on the key (§4).
  • A switched-off row costs nothing: its neutral tuple is a table row the multiplicity counts (§4).
  • The table is the intended one: the verifier's own closed form (§3), or a table bound by identity or by the SRS digest, as trustworthy as the channel the verifier took that from (program.md §8, srs.md §3).
  • Every declared obligation is discharged: the discharge rule over the registry's specs (§11).

The channel does not check the multiplicity column (§7), a key's bound (§4), or that a committed table holds its neutral row, a property of its values that no artifact states: a table without one stops the honest prover at trace::build_multiplicities.

Auditoren/Beweissystem

Der Beweis

Normative Spezifikationdocs/spec/proof.mdAls Markdown anzeigen

Zusammenfassung

Wie sich die Argumente zu einer verifizierten Aussage zusammensetzen. Die Seite definiert PublicInputs und die Reihenfolge der Aussage, den Block und die Exaktheit seiner Shard-Menge, das globale Transkript G1–G11, den SRS-Digest, das Shard-Transkript S1–S6, die gebündelte Öffnung, die Prüfungen des Verifiers in ihrer Reihenfolge mit der Fehlerklasse jeder einzelnen, den Verifikationsschlüssel und seine Laderegeln, Zeitfenster und die Serialisierungsformate jedes Objekts auf Byte-Ebene, einschließlich des Beweisarchivs.

Der folgende normative Text wird auf Englisch gepflegt, der kanonischen Sprache der Spezifikation.

What a verifier checks and the formats it reads, from the statement to the bytes. The memory argument, LogUp, GKR and Mercury are their own pages; this one is how they compose. crates/verifier-core (#![no_std]) implements everything here but step 12, the opening, which crates/verifier runs.

1 The statement#

A statement is a PublicInputs under a verifying key (§7): one execution of the key's program. It is proved by one ShardProof per statement shard, a (family, index) with index below the family's shard count, each verified against the same PublicInputs. A shard's proof establishes its own circuit and opening, and the reconciliation it joins reads the roots the statement claims for every other shard, which only their own proofs establish: a statement is verified when its proofs are exactly its shards and all pass, never by a subset.

Every entry point is (&VerifyingKey, proof, &PublicInputs): verifier::verify_shard, verifier::verify_block, verifier_core::reduce_shard. A verifier holds two values from a channel the prover does not control, the program identity and the SRS digest (§3), and compares them with the key's; the key itself may come from anyone (§7). The verifier CLI compares identity only (tools.md §6); host::verify(vk, block), which is verify_block(vk, block, block.statement()), compares neither and leaves its caller to check the statement's input, output and exit status too.

1.1 PublicInputs#

field
input: Vec<u8> the public input window's payload, at most PUBLIC_PAYLOAD_BYTES = 16,380 (public-values.md §3)
output: Vec<u8> the journal, the public output window's payload, as long
exit_status: u32 x10's final value
shard_counts: Vec<u32> one per family of the VmConfig, in its order, possibly 0
windows: Vec<u32> ZERO_WINDOWS' window ids, one per shard (memory.md §3.5)
boundary: BoundaryFinals the 64 register and pc boundary scalars (memory.md §4.1)
memory_commitments: Vec<Vec<[u8; 64]>> per statement shard, its M columns' commitments in layout order
memory_roots: Vec<[Fr; 2]> per statement shard, [read_root, write_root]

The first three are the claim; the rest is the execution's record, which the prover chooses. All of it but the roots and the exit status is absorbed before any challenge (§2).

1.2 Statement order#

verifier_core::statement_shards(config, counts):

(INIT_TEARDOWN, 0)
(ZERO_WINDOWS, 0) … (ZERO_WINDOWS, k − 1)
every other family of the VmConfig, ascending by id, shards 0 … count − 1 each

It orders memory_commitments, memory_roots, G8's groups and a block's proofs. The two leading families are ids 7 and 8, so the order is not ascending by id. A family with count 0 has no entry.

1.3 The block#

verifier_core::BlockProof { config, statement, shards } is one execution closed: the static VmConfig, the statement and one proof per statement shard, in statement order; it adds no evidence to the proofs'. The config and counts are public data of the proof, absorbed at G3 and G4, so the block carries both and check B1 holds them to the key's and the verifier's.

Shard-set exactness, BlockProof::shape, at decode and again in verify_block: one count per config family; the counts' total, summed in u64 before any list is built from them, equal to the numbers of proofs, commitment lists and root pairs; the proofs naming statement_shards in order. No (family, index) is missing, repeated or extra.

BlockProof::reconciliation is the cross-shard record set, a BlockReconciliation of one ShardRecord { family, shard_index, ts_window, memory_commitments, roots } per statement shard, assembled from the shard's proof (the window) and the statement (the rest).

2 The global transcript#

verifier_core::global_commit(vk, statement), run by the verifier in derive_global_phase and by the prover once every shard's M columns are committed (streaming.md §2): a fresh transcript, tag values in transcript.md §5.

# op tag message
G1 absorb PROTOCOL_SUITE [PROTOCOL_VERSION], 0
G2 absorb SRS_DIGEST [vk.srs_digest] (§3)
G3 absorb VM_CONFIG the config (program.md §7)
G4 absorb SHARD_COUNTS shard_counts
G5 absorb MEMORY_WINDOWS windows
G6 absorb PROGRAM_IDENTITY [vk.identity]
G7 absorb PUBLIC_INPUTS, bytes the 32 bytes of io_digest(input, output) (public-values.md §5)
G8 per family MEMORY_GROUP, COMMITMENT below
G9 absorb MEMORY_BOUNDARY the 64 boundary scalars
G10 squeeze ×4 MEMORY_CHALLENGE γ_M, α_addr, α_ts, α_val, challenge slots 1–4
G11 squeeze GLOBAL_STATE_DIGEST the global state digest

G3–G5 are verifier_core::absorb_statement_descriptor. G8 is one group per family of the config, in statement order, a family with count 0 included:

MEMORY_GROUP   [family, shard count]
COMMITMENT     per shard, ascending: its memory_commitments, one message of 4k limbs

A statement's or proof's points are absorbed as limbs (transcript.md §4) and decoded only at step 12.

Everything a memory tuple or the reconciliation reads precedes G10 (memory.md §6.1 says why for each). Two fields are not absorbed: memory_roots, which depend on the challenges and are bound by each shard's own GKR proof (step 10a), and exit_status, which step 10b holds to v_10, absorbed at G9. The digest seeds every shard (§4); a proof carries the digest it was seeded with (ShardProof::global_digest) and step 5 compares it with the replay, so a shard proof is for one statement under one key.

3 The SRS digest#

t ← Transcript::new()
t.append_bytes(SRS_VERIFIER, srs_verifier)           320 bytes, srs.md §5
append_g1_points(t, GENERIC_TABLE, generic_table)    the table's 3 points, one 12-limb message
srs_digest ← t.sample()                              one raw squeeze

verifier_core::srs_digest; GENERIC_TABLE is absorbed in this sponge and nowhere else, the points key column first. G2 absorbs the digest, so a proof is bound to the three points its pairings read and the table its GENERIC lookups read. Both are constants of the ceremony (lookup.md §9), so one digest serves every key. It does not cover the powers, which only a prover reads: an opening is checked against g2_tau whatever powers made the commitment.

A key's load recomputes the digest from the key's own points (§7), which shows they agree, not that they are the ceremony's, and identity binds neither (program.md §8). So the verifier compares vk.srs_digest with the ceremony's (srs.md §3). Without that comparison, whoever built the key chose τ, so can open anything, and chose the table every GENERIC lookup is held to.

4 The shard transcript#

Shard (family, index) runs a fresh sponge (verifier_core::shard_transcript), not a restored global one:

# op tag message
S1 absorb SHARD_SEED [global state digest, family, index]
S2 absorb SHARD_TS_WINDOW [ts_start, ts_end] (§8)
S3 absorb COMMITMENT the shard's W commitments, multiplicities included, one message
S4 squeeze ×2 LOOKUP_CHALLENGE g, then β (lookup.md §2)
S5 the GKR backward pass (gkr.md §5.2)
S6 the batch opening (§5): B1–B3 of mercury.md §5, then the sixteen steps of its §3

S4 is drawn for every shard, whether or not its circuit has a channel. Every challenge follows every commitment the circuit reads: M at G8, S through identity at G6 or the SRS digest at G2, W at S3, as GKR requires of its caller (gkr.md §5.1).

The circuit's external challenges (verifier_core::shard_challenges) are slots 1–4 from G10; for a window family, slot 5 at the window verifier_core::shard_window gives the shard (memory.md §3.3); then the lookup slots from g and β. Its outputs, the top layer, are the two memory roots and then each channel's (num, den) (lookup.md §6), 2 + 2c of them for c channels.

In the recursion format a shard commits M and W as stacks of 2^σ columns, and S6 opens with σ STACK_CHALLENGE squeezes extending the opening point (recursion.md §1.3). At σ = 0, the base format, there are none.

5 The opening#

After S5 every committed column has one claim, layer 0's, all at one point u (gkr.md §5.2). So there is nothing for a claim-merging sumcheck to merge, and S6 opens every column as one batch (mercury.md §5), one 704-byte Mercury proof a shard:

columns      the circuit's committed layout: M[0..], W[0..], S[0..]
commitments  M  PublicInputs.memory_commitments[the shard's position]
             W  ShardProof.witness_commitments
             S  VerifyingKey.setup_commitments[family], then VerifyingKey.generic_table
                when the circuit reads GENERIC (FamilyCircuit::reads_generic_table)
point        u, variable j at index j
values       layer 0's claims, ShardProof.gkr.layers[0].final_evals

Column i carries ρ^i, so this order is part of what is proved. Virtual columns are neither claimed nor opened: the verifier evaluates their closed forms. Taking S from the key is what makes the opening bind the columns identity commits, the decoded tables and the image column (memory.md §6.2), and the generic table the SRS digest covers.

reduce_shard ends at an OpeningClaim: these commitments, the point, the values and the live shard transcript. verify_shard spends it with pcs::batch_verify (step 12); a recursion node defers it (recursion.md §8.3).

6 Verification#

verifier::verify_shard(vk, proof, public) returns the first failure, in this order, as a VerifyError:

step class check
1 Statement one shard count per config family; the key's circuits are its config's families, in order, with one setup list each
2 Statement check_memory_windows (memory.md §3.5); input and output each at most PUBLIC_PAYLOAD_BYTES
3 Statement one root pair and one commitment list per statement shard, each list its family's M width; the total summed in u64 first
G1–G11 (§2)
4 Statement ts_start ≤ ts_end ≤ 2^38
5 Statement the replayed global state digest is proof.global_digest
6 Malformed (family, index) is a statement shard; the witness commitments and outputs have the circuit's counts
7 Constraint { layer } gkr_verify::verify over the shard transcript: LayerInconsistency { layer }; its ProofShape, OutputShape and MissingChallenge are Malformed
8 Constraint { layer: 0 } every base claim at one point
9 Lookup { channel } gkr_verify::channel_holds on each channel's root pair, in channel order (lookup.md §8)
10a MemoryArgument the proof's two roots are the statement's for its position
10c MemoryArgument a PUBLIC_INPUT or PUBLIC_OUTPUT shard's value column is the statement's string (public-values.md §5)
10b MemoryArgument every boundary timestamp below 2^38; v_10 = exit_status; gkr_verify::reconciles over every shard's roots and boundary_factors(challenges, vk.entry_pc, boundary) (memory.md §4.2)
11 — the opening claim (§5)
12 Opening the SrsVerifier, every commitment and the Mercury proof through their validating decoders, then pcs::batch_verify; any failure

Step 8 cannot fail on verify's output, whose base claims share layer 0's point; it states what step 11 relies on. Steps 1–3 hold the statement to the key before the replay indexes by it, so nothing a proof or statement carries makes the core panic, for a loaded key.

The split, by what each part reads (verifier_core):

function reads steps runs
derive_global_phase(vk, public) → GlobalChallenges key, statement 1–3, G1–G11 once a statement
verify_shard_local(vk, global, proof, public) → OpeningClaim and one ShardProof 4–10a, 10c, 11 once a shard
verify_global_memory(vk, global, public) key, statement, challenges 10b once a statement

GlobalChallenges is the four memory challenges and the digest. reduce_shard is the three in that order, verify_shard that and step 12; step 11 cannot fail, so 10b after it is 10b in place. Step 10b reads only vk.entry_pc, the boundary, the roots and the challenges, so a block runs it once; step 10a puts each shard into the product by holding the roots its GKR proof outputs to the statement's entry, and shard-set exactness makes every root there a verified shard's. verify_shard_local alone verifies no memory argument: without verify_global_memory it accepts shards, each valid, whose multiset does not close.

verifier::verify_block(vk, block, public):

class check
B1 Statement block.config is vk.config, and block.statement is public
B2 Statement derive_global_phase, once
B3 Statement BlockProof::shape (§1.3)
B4 Statement check_ts_windows over the records (§8)
B5 MemoryArgument verify_global_memory, once
B6 as verify_shard per shard, in statement order: verify_shard_local, then step 12

B1–B5 read no GKR proof or opening, so a statement that cannot reconcile is refused before any circuit runs, and the class can differ from verify_shard's: a change to anything G1–G9 absorb that B1–B4 admit moves the challenges, so the honest roots stop reconciling and verify_block answers MemoryArgument where verify_shard names the seed at step 5; a forgery that unbalances the multiset is MemoryArgument even where it also breaks a gate. A dropped shard fails B5: the truncated statement, re-proved honestly with its counts, lists and roots adjusted, passes B1–B4 and misses that shard's memory events on one side of the product.

In verify_shard's order the class names the fault: a tampered witness proved honestly, its multiplicities recounted, columns recommitted and statement rebuilt, fails at the gate (Constraint), table membership (Lookup) or multiset (MemoryArgument) it broke, which checker::TamperHarness asserts (circuits.md §3).

7 The verifying key#

7.1 Fields#

field
code_version: u32 constants::family::CODE_VERSION, 0
config: VmConfig the static shape (program.md §7)
entry_pc: u32 the image's entry pc
identity: ProgramIdentity program.md §8
setup_commitments: Vec<Vec<[u8; 64]>> identity's commitment lists, one per config family, in its order
srs_verifier: [u8; 320] the SrsVerifier (srs.md §5)
generic_table: [[u8; 64]; 3] the generic table's commitments, key column first, in every key (lookup.md §9)
srs_digest: Fr §3
circuits: Vec<FamilyCircuit> one per config family, in its order: the family, its CircuitArtifact and its ChannelSpecs (lookup.md §11)

A key carries every family's artifact, so its size is mostly its delegation families' (circuits.md §1).

7.2 Loading#

VerifyingKey::from_bytes decodes (§9), refuses bytes that are not the key's canonical encoding, and runs VerifyingKey::check, which refuses, in order:

  1. a VmConfig no derivation produces (VmConfig::from_bytes of its own bytes), or a code_version other than CODE_VERSION;
  2. a setup list count other than the config's family count;
  3. an identity that identity_digest(code_version, config, entry_pc, setup_commitments) does not reproduce;
  4. an srs_digest that srs_digest(srs_verifier, generic_table) does not reproduce;
  5. a circuit count other than the family count; then, family by family: a circuit for another family; a height the registry has no circuit for; a circuit, artifact or channel specs, other than config.circuit(family, trace_vars), the registry of the config's format (circuits.md §1); an artifact failing CircuitArtifact::validate, constraints::memory::check_memory or constraints::lookup::check_discharge; a setup list whose length, plus 3 if the circuit reads GENERIC, is not the artifact's S count; and GENERIC specs naming anything but the 3 setup columns after identity's, §5's order, which no registry circuit fails.

verifier::load_verifying_key then decodes every curve point: the SrsVerifier's three, each setup commitment and each generic-table commitment, through the validating readers. The circuits are held to the registry because nothing else binds them: identity binds the program, not the circuit that proves it.

A key from an untrusted source. Every field is recomputed from or compared with one of the verifier's two trusted values (§1), or fixed by the code: config, entry_pc and the setup lists through identity; srs_verifier and generic_table through the SRS digest; code_version and the circuits by the verifier's own registry. So a key may come from the prover, provided both comparisons are made.

Validation runs once, at load; verify_shard and verify_block assume a loaded key. On one edited in memory a changed config or circuit list is still refused as Statement, but an edit inside a circuit may go unnoticed.

prover::ProverSetup::new(program, srs) builds the key: each family's circuit from VmConfig::circuit and fill from prover::family_fill (ProverError::Unregistered if either is missing), identity's and the generic table's commitments over srs, then check (ProverError::Key). srs needs as many powers as the tallest family has rows, and 2^18 for the generic table (program::lookup_tables::GENERIC_LOG_HEIGHT); fewer panics.

8 Time windows#

Each shard claims [ts_start, ts_end) (ShardProof::ts_window): the slice of the clock (execution-trace.md §1) its rows write in, their reads reaching back before it. S2 absorbs it before the witness commitments, so a proof made under one window fails under another; step 4 holds it to ts_start ≤ ts_end ≤ 2^38 and nothing more.

verifier_core::check_ts_windows (B4), over the records in statement order: within each cycle-owning family (constants::family::CYCLE_OWNING, the execution families 0–6), every window is non-empty and no shard's ts_end exceeds the next shard's ts_start; a family's records are consecutive and ascending, so neighbours suffice. It is per family because families interleave — ADD_SUB_LUI_AUIPC may own cycles 1 and 3 and JUMP_BRANCH_SLT cycle 2 — and every other family is exempt: a window family's rows are words, a delegation family's invocations at their requesting cycles (delegation.md §8).

The prover reads a window off the shard's committed M[0] cycle column (ts_window, crates/prover/src/lib.rs): [4·c_0, 4·c_max + 4), c_0 row 0's cycle and c_max the largest, padding rows carrying 0, for cycle-owning and delegation families alike. Window families claim verifier_core::TRIVIAL_TS_WINDOW = [0, 2^38).

A window binds nothing. No gate ties it to the rows committed under it, so a prover may claim any windows the rule admits; cross-shard order, cycle uniqueness and pc continuity are the memory multiset's alone (memory.md §9). B4 checks the shape of the shard plan and adds nothing to soundness.

9 Wire forms and the proof archive#

verifier_core::wire: integers little-endian; an Fr its 32 canonical bytes (primitives.md §1), refused at or above p; a G1 its 64 bytes (primitives.md §3), opaque to the core; bytes a u32 length then the bytes; list<T> a u32 count then the items; T[k] exactly k items, no count. Every decoder is total: it refuses a count the remaining bytes cannot hold, so it reserves nothing an untrusted length asks for, and refuses trailing bytes.

PublicInputs   input bytes, output bytes, exit_status u32,
               shard_counts list<u32>, windows list<u32>,
               boundary Fr[64]                     memory.md §4.1's order and ranges
               memory_commitments list<list<G1>>, memory_roots list<Fr[2]>

ShardProof     family u32, shard_index u32, ts_start u64, ts_end u64, global_digest Fr,
               witness_commitments list<G1>, outputs list<Fr>,
               gkr list<(rounds list<Fr[4]>, final_evals list<Fr>)>       transition 0 first
               opening u8[704]                     pcs::MercuryProof, mercury.md §4

BlockProof     config bytes                        VmConfig, program.md §7
               statement bytes                     PublicInputs
               shards list<bytes>                  each a ShardProof; then BlockProof::shape

VerifyingKey   code_version u32, config bytes, entry_pc u32, identity Fr,
               setup_commitments list<list<G1>>, srs_verifier u8[320], generic_table G1[3],
               srs_digest Fr,
               circuits list<(family u32, artifact bytes,          CircuitArtifact, gkr.md §4.1
                              channels list<(channel u32, table list<Address>,
                                             multiplicity Address)>)>
Address        tag u8 (0 M, 1 W, 2 S, 3 V), index u32; a V's index is its gkr.md §2.1 kind tag

BlockReconciliation   list<(family u32, shard_index u32, ts_start u64, ts_end u64,
                            memory_commitments list<G1>, read_root Fr, write_root Fr)>

A ShardProof's lengths are fixed by its key and family, and steps 6–7 hold them: transition k carries n_{k+1} rounds and w_k claims, twice that if halving (gkr.md §5.5). For a base-format circuit at 2^n with W witness, C committed and I inner columns, O outputs, and R row-wise lists before its n halving ones, that is

772 + 64·W + 32·O + 8·(R + n) + 128·(R·n + n(n − 1)/2) + 32·(C + I + O·(n − 1))   bytes

which circuits.md §1 tabulates per family.

The proof archive. verifier::proof_archive::write_proof(dir, stem, vk, block), re-exported as host::proof_archive, writes four files, each a bare to_bytes with no header of its own:

<stem>.vk         VerifyingKey
<stem>.identity   the key's identity: its 32 bytes in order, 64 lowercase hex digits, a newline
<stem>.public     PublicInputs: the block's own statement
<stem>.block      BlockProof

read_proof(dir, stem) is the inverse, each file through its type's decoder and the key through load_verifying_key. .identity records what the run claimed, and read_proof returns it unchecked: a verifier's identity comes from its own channel (§1). .public repeats the statement .block carries, for the CLI, which takes it as a file (tools.md §6).

Auditoren/Befehlsfamilien

Die Familie ADD_SUB_LUI_AUIPC

Normative Spezifikationdocs/spec/add-sub.mdAls Markdown anzeigen

Zusammenfassung

Der Schaltkreis der Familie 0, Spalte für Spalte und Gate für Gate. Neben add, sub, addi, lui, auipc und fence ist jeder ecall eine Zeile dieser Familie, sodass ihr Schaltkreis auch den Exit und die Anfrageseite jedes Delegationsaufrufs beweist. Die Seite führt ihre committeten Spalten, ihre 63 Constraint-Gates und ihre 15 Lookup-Verpflichtungen auf, begründet ihre Soundness und nennt ihre Grenzen.

Der folgende normative Text wird auf Englisch gepflegt, der kanonischen Sprache der Spezifikation.

add, sub, addi, lui, auipc, ecall, ebreak and fence, compressed forms included, are family 0, one executed instruction a row; constraints::add_sub::artifact is its circuit. Every ecall is a row of it, so the circuit also proves the exit and the request side of every delegation call. This page specifies what it adds beside the memory frame (memory.md §2).

1 Columns#

The decoded tuple is pc next_pc rs1 rs2 rd imm extra_mask (program.md §5), the mask one-hot over the kinds system addi auipc add sub lui in bit order (constants::extra_mask::add_sub_lui_auipc). ecall, ebreak and fence share the system kind, with imm 0, 1 and 2 (constants::extra_mask::system_code); elsewhere imm is what the instruction adds — addi's sign-extended immediate, lui's and auipc's shifted left by 12, 0 for add and sub — and a register field the instruction lacks is 0.

M[0..26] and W[0..8] are the frame of the five queries pc rs1 rs2 rd deleg, and M[26], deleg_space, is the requested delegation type's anchor address space, 0 on a row requesting none (memory.md §2). The family adds these columns, and reads V[range19] and V[range16]:

column name
W[8..14] decoded_next_pc, decoded_rs1, decoded_rs2, decoded_rd, decoded_imm, decoded_mask the claimed decoded row
W[14..20] kind_system … kind_lui b_k, the mask's bits
W[20], W[21] is_ecall, is_fence the system kind, split by its code
W[22..28] is_deleg_<f>, f = 9, 10, 11, 15, 16, 17 d_t: a request of delegation type t, family f
W[28], W[29] wrap, rd_hi the sum's carry or the difference's borrow; sel's high halfword
W[30], W[31] pc_wrap, next_pc_hi next_pc's wrap and high halfword
W[32..35] mult_timestamp, mult_range16, mult_decoder one multiplicity a channel
S[0..7] table_pc … table_extra_mask the decoded table

Below, m_q, a_q, ts_q and v_q are query q's mask, address, read timestamp and read value; pc and next_pc the pc query's read and write values; sel is rd_selected (W[7]), the value the frame writes to a nonzero rd. N_t and tag_t are type t's ecall number and anchor space, the first six rows of constants::delegation::TYPES in order (delegation.md §3); is_exit = is_ecall − Σ_t d_t; 93 is constants::ecall::EXIT and HALT_PC is 1 (memory.md §5).

The family's fill (prover::family_fill, crates/prover/src/fill.rs) writes sel as the computed value even where rd = x0.

2 Gates#

63 enforcing gates, all in gate list 0, each of degree at most 2 and 0 on the all-zero row: the frame's eleven (memory.md §2) and these 52, in artifact order, each held to 0:

gate expression
kind_<k>_boolean, six b_k − b_k²
decoded_mask_bits Σ_k 2^k·b_k − decoded_mask
is_ecall_boolean, is_fence_boolean y − y²
system_split is_ecall + is_fence − b_system
ecall_code is_ecall·decoded_imm
fence_code is_fence·(decoded_imm − 2)
per type: is_deleg_<f>_boolean, deleg_<f>_is_an_ecall, deleg_<f>_number d_t − d_t²; d_t·(1 − is_ecall); d_t·(v_rs1 − N_t)
ecall_is_exit is_exit·(v_rs1 − 93)
rs1_mask_rule m_rs1 − m_pc·(b_add + b_sub + b_addi + is_ecall)
rs2_mask_rule m_rs2 − m_pc·(b_add + b_sub + is_ecall)
rd_mask_rule m_rd − m_pc·(b_add + b_sub + b_addi + b_auipc + b_lui + is_ecall)
deleg_mask_rule m_deleg − m_pc·Σ_t d_t
rs1_addr_rule m_rs1·(a_rs1 − decoded_rs1 − 17·is_ecall)
rs2_addr_rule, rd_addr_rule m_q·(a_q − decoded_q − 10·is_ecall)
rs1_value_masked, rs2_value_masked v_q − m_q·v_q
add_addi_auipc (b_add + b_addi + b_auipc)·(v_rs1 + v_rs2 + decoded_imm − sel − 2^32·wrap) + b_auipc·pc
sub b_sub·(v_rs1 − v_rs2 − sel + 2^32·wrap)
lui b_lui·(decoded_imm − sel)
exit_status is_exit·(v_rd − sel)
deleg_writes_no_register m_deleg·sel
deleg_read_ts_zero, deleg_read_value_zero m_deleg·ts_deleg; m_deleg·v_deleg
deleg_addr_rule m_deleg·(a_deleg − v_rs2)
deleg_space_rule deleg_space − Σ_t tag_t·d_t
wrap_boolean, pc_wrap_boolean y − y²
next_pc_rule next_pc + 2^32·pc_wrap − (1 − is_exit)·decoded_next_pc − is_exit·HALT_PC

N_t and tag_t are literals read from constants::delegation::TYPES, so the base circuit depends on the registry's first BASE_TYPES = 6 rows and on no row appended after them. The recursion format's circuit, add_sub::recursion_artifact, carries a selector and its three gates for each of the ten types, the columns after them shifted by four, and in place of deleg_writes_no_register deleg_a0_rule, Σ_{t<6} d_t·sel + Σ_{t≥6} d_t·(sel − v_rs2 − 4·words_t) with words_t the type's frame length: a recursion type's request leaves a0 past its frame (recursion.md §1.4).

3 Lookups#

15 obligations on three channels, none of them GENERIC, so the setup columns are the decoded table alone: the frame's ten TIMESTAMP gaps, two a query under its mask (memory.md §2), and five under m_pc:

lookup channel tuple
rd_hi_range, rd_lo_range RANGE16 rd_hi; sel − 2^16·rd_hi
next_pc_hi_range, next_pc_lo_range RANGE16 next_pc_hi; next_pc − 2^16·next_pc_hi
decode_row DECODER pc, decoded_next_pc, decoded_rs1, decoded_rs2, decoded_rd, decoded_imm, decoded_mask

The channels, in output order (add_sub::channels), are TIMESTAMP over V[range19], RANGE16 over V[range16] and DECODER over S[0..7].

4 Why it is sound#

On a live row (m_pc = 1) decode_row makes the claimed row the table's at pc, so exactly one b_k is 1 (lookup.md §10). The mask rules make each query present exactly where the row's kind or request makes it (execution-trace.md §4, §6). The address rules make a register query the decoded register, or on an ecall row, whose decoded registers are 0, a7 (17) for rs1 and a0 (10) for rs2 and rd. The _value_masked gates make an absent operand read 0, which lets one gate serve three sums: an addi or auipc row's v_rs2, and an auipc row's v_rs1, would otherwise be free addends, and add's imm is the table's 0.

Read values are words (memory-ops.md §5), sel is a word by its range pair and wrap is boolean, so each arithmetic gate is an identity over ℤ with one solution: the sum mod 2^32 and its carry, the difference mod 2^32 and its borrow, or imm. Without the pair, a sum at or above 2^32 would satisfy the gate with wrap = 0 and reach a register. The frame's x0 rule then writes sel or discards it.

next_pc is a word by its range pair and is decoded_next_pc — the table's fall-through, so a compressed instruction advances by 2 (program.md §5) — or HALT_PC on the exit row, less 2^32·pc_wrap. Both are far below 2^32, so pc_wrap = 0 on every live row.

On a system row system_split sets exactly one of is_ecall and is_fence, and the code gates make it the one imm names; ebreak's code 1 satisfies neither, so an ebreak row is unprovable. A fence row makes no query but the pc's and falls through. Off a system row both bits are 0, and so, by deleg_<f>_is_an_ecall, is every d_t.

A set d_t forces is_ecall = 1 and a7 = N_t. The numbers are pairwise distinct and none is 93, const assertions beside the circuit, so at most one d_t is set, is_exit is 0 or 1, and an ecall row is the exit, with a7 = 93, or a request of exactly one type; no other a7 passes.

  • The exit row writes back the a0 it read (exit_status), so x10's final value is the status the statement carries (proof.md §6, step 10b), and writes HALT_PC, after which no row runs (memory.md §5).
  • A request row falls through, writes 0 to a0, and makes the mirror query at the frame base it read from a0, in the space deleg_space names, reading timestamp 0 and value 0. Those three zeroings pair it one-to-one with an invocation of its type (delegation.md §5); the mirror's write value is free here, and what the call computed is the invoked family's circuit (delegation-circuits.md). deleg_space is an M column because a memory leaf may read no W column (memory.md §8); deleg_space_rule ties it to the selectors.

A row with m_pc = 0 is bound to no table row and its kind bits are free; the arithmetic gates are gated by those bits alone, m_pc times a bit being degree 3. That is harmless: the mask rules zero the row's other four masks and every lookup is off, so it adds no memory tuple. On a live row, wrap outside the four sums and sel on a fence row are free, and nothing reads them.

5 Limits#

  • An ecall whose a7 is neither 93 nor a type the format's circuit knows has no proof. The emulator answers an unassigned number -ENOSYS and continues (ecall-abi.md §5); the fill refuses that trace, naming the cycle.
  • An ebreak has no proof; it is fatal in the emulator (execution-trace.md §10).
  • No row touches RAM: an ecall row reads a7 and a0 and writes a0, and a request's operands travel in the invoked family's frame (delegation.md §4).

Auditoren/Befehlsfamilien

Die Familie JUMP_BRANCH_SLT

Normative Spezifikationdocs/spec/jump-branch-slt.mdAls Markdown anzeigen

Zusammenfassung

Der Schaltkreis der Familie 1: die Set-less-than-Befehle, die sechs Verzweigungen, jalr und jal. Die Seite spezifiziert die Spalten, das Is-zero-Gadget und das Vergleichs-Gadget, in dem eine einzige per Range-Check geprüfte Differenz die Ordnung mit und ohne Vorzeichen ohne Vergleichstabelle entscheidet, die Gates einschließlich der einzigen next_pc-Regel mit ihrer Prüfung auf Geradheit, die Lookups und warum der Schaltkreis genau das Verhalten der ISA zulässt.

Der folgende normative Text wird auf Englisch gepflegt, der kanonischen Sprache der Spezifikation.

The circuit of slti, sltiu, slt, sltu, the six branches, jalr and jal: what it adds beside the memory frame every execution family carries (memory.md §2), and the two gadgets other families reuse (§3). One comparison settles signed and unsigned order for the branches and the slt kinds alike. The circuit is constraints::jump_branch_slt::artifact (crates/constraints/src/jump_branch_slt.rs); prover::family_fill writes its witness.

1 What the circuit reads from the decoded table#

The tuple is pc next_pc rs1 rs2 rd imm extra_mask (program.md §5). next_pc is the fall-through, seq below; imm is the two's-complement word of the value the instruction uses: the sign-extended immediate of slti and sltiu (which sltiu compares unsigned), a branch's or jal's displacement, jalr's offset. extra_mask is one-hot over constants::extra_mask::jump_branch_slt:

bit    0     1      2    3     4    5    6    7    8     9     10    11
kind   slti  sltiu  slt  sltu  beq  bne  blt  bge  bltu  bgeu  jalr  jal

The legal masks are these twelve one-bit values, jump_branch_slt::LEGAL_MASKS; rd = x0 is the table's rd, not a mask. The circuit commits the twelve bits b_k, and every signal it needs is a linear form over them: the signed-comparison flag sc = b_slti + b_slt + b_blt + b_bge, the compared immediate (b_slti + b_sltiu)·imm, so that a branch's displacement never reaches the comparison, and the branch weights of taken_rule (§4).

2 Columns#

The frame is M[0..21] and W[0..7], over the queries pc rs1 rs2 rd at slots 0–3 (memory.md §2). Its W[6], rd_selected (sel below), holds the value the instruction computes, which the x0 rule masks into the write. The circuit adds:

column name value
W[7..13] decoded_next_pc … decoded_mask the claimed decoded row after pc
W[13..25] kind_slti … kind_jal the bits b_k, in §1's order
W[25] cmp_rhs the right operand, rs2 + (b_slti + b_sltiu)·imm
W[26], W[27] rs1_hi, rs1_sign rs1 >> 16, rs1 >> 31
W[28], W[29] cmp_rhs_hi, cmp_rhs_sign the same of cmp_rhs
W[30] lt rs1 < cmp_rhs, signed where sc = 1
W[31], W[32] cmp_gap, cmp_gap_hi (rs1 − cmp_rhs) mod 2^32, and its high halfword
W[33], W[34] eq, eq_inv [rs1 = cmp_rhs] on a live row; the difference's inverse
W[35] taken a taken branch
W[36] jalr_drop bit 0 of rs1 + imm on a jalr row
W[37] pc_wrap the carry out of whichever sum next_pc is
W[38], W[39] next_pc_hi, rd_hi next_pc >> 16, sel >> 16
W[40..44] mult_timestamp … mult_decoder one multiplicity per channel, in channel order
S[0..7] table_pc … table_extra_mask the decoded table, which program identity binds
S[7..10] generic_key … generic_result the packed table (lookup.md §9)
V[range19], V[range16] the TIMESTAMP and RANGE16 tables

21 M, 44 W and 10 S columns, 75 committed; 42 enforcing gates, the frame's 10 and §4's 32; 22 lookups: 8 TIMESTAMP, 11 RANGE16, 2 GENERIC, 1 DECODER, counts artifact asserts.

3 The gadgets#

constraints::gadgets returns gates and lookups as data. is_zero also builds the frame's x0 rule (memory.md §2) and MUL_DIV's zero tests (mul-div.md); the comparison also orders ATOMICS' minimum and maximum (memory-ops.md §6).

3.1 is_zero(x, inv, z, enable)#

x·inv + z − enable = 0          x = Σ c_i·x_i, a linear form
z·x = 0

With enable boolean, which the caller establishes, these force z = enable·[x = 0]: at x ≠ 0 the second gives z = 0 and the first inv = enable/x; at x = 0 the first gives z = enable. So z is boolean with no gate of its own, and enable = 0 gives z = 0, which keeps the all-zero row valid.

3.2 The comparison#

Comparison names one comparison lhs < rhs by its columns, its lookups' selector, and the kind bits signed whose sum is sc, which the caller holds to 0 or 1 on a selected row. comparison returns, for x each of lhs, rhs and gap:

name kind expression
<p>_order gate lhs − rhs − 2^32·sc·lhs_sign + 2^32·sc·rhs_sign + 2^32·lt − gap
<p>_lt_boolean gate lt − lt²
<p>_<x>_hi_range, <p>_<x>_lo_range RANGE16 x_hi; x − 2^16·x_hi
<p>_lhs_get_sign, <p>_rhs_get_sign GENERIC (x_hi + SIGN_BASE, x_sign, 0)

The range pairs make lhs, rhs and gap words and each x_hi the true high halfword (memory.md §7), which keeps each sign key inside U16GetSign's range (lookup.md §4), so each sign is its operand's bit 31. Let D = lhs − rhs − 2^32·sc·(lhs_sign − rhs_sign): both operands read in two's complement where sc = 1, so mixed signs are no case split, and D ∈ (−2^32, 2^32). The gate says gap = D + 2^32·lt, and only lt = [D < 0] puts gap in [0, 2^32): at D ≥ 0, lt = 1 puts it at 2^32 or above; at D < 0, lt = 0 makes it a negative field element. So the range check on gap carries the order, and no comparison table exists; the honest gap is (lhs − rhs) mod 2^32 whatever sc is. Both gates are ungated, since a selector would make the order gate degree 3, and every row satisfies them with the gap its own values give. comparison_equation(c, word_bits) builds the order gate at any width to 32, and the row suite evaluates it at 6 bits over every operand pair, signed and unsigned, finding exactly one (lt, gap), the ISA's.

4 Gates#

After the frame's ten in gate list 0, with m_q, a_q, v_q query q's mask, address and read value, and pc, next_pc the pc query's read and write:

gate polynomial
kind_<k>_boolean ×12 b_k − b_k²
decoded_mask_bits Σ_k 2^k·b_k − decoded_mask
rs1_mask_rule m_rs1 − m_pc·(Σ_k b_k − b_jal)
rs2_mask_rule m_rs2 − m_pc·(b_slt + b_sltu + the six branch bits)
rd_mask_rule m_rd − m_pc·(b_slti + b_sltiu + b_slt + b_sltu + b_jalr + b_jal)
<q>_addr_rule, for rs1, rs2, rd m_q·(a_q − decoded_q)
<q>_value_masked, for rs1, rs2 v_q − m_q·v_q
cmp_rhs_rule cmp_rhs − v_rs2 − (b_slti + b_sltiu)·imm
cmp_order, cmp_lt_boolean §3.2: lhs = v_rs1, rhs = cmp_rhs, signed the bits of sc
eq_inverse, eq_at_nonzero §3.1: x = v_rs1 − cmp_rhs, z = eq, enable = m_pc
taken_rule taken − w_1 − w_eq·eq − w_lt·lt
taken_boolean, jalr_drop_boolean, pc_wrap_boolean x − x²
next_pc_rule §5's equation
rd_value_rule sel − (b_jal + b_jalr)·seq − (b_slti + b_sltiu + b_slt + b_sltu)·lt

The branch weights are w_1 = b_bne + b_bge + b_bgeu, w_eq = b_beq − b_bne and w_lt = b_blt + b_bltu − b_bge − b_bgeu. Every gate has degree at most 2 and is 0 on the all-zero row, which artifact asserts.

4.1 Lookups#

After the frame's 8 TIMESTAMP obligations, all under m_pc:

lookup channel expression
cmp_<x>_hi_range, cmp_<x>_lo_range ×3 RANGE16 §3.2 over v_rs1, cmp_rhs, cmp_gap
cmp_lhs_get_sign, cmp_rhs_get_sign GENERIC §3.2
rd_hi_range, rd_lo_range RANGE16 rd_hi; sel − 2^16·rd_hi
next_pc_hi_range, next_pc_lo_range RANGE16 next_pc_hi; next_pc − 2^16·next_pc_hi
next_pc_even RANGE16 2^−1·next_pc − 2^15·next_pc_hi
decode_row DECODER pc and W[7..13] (lookup.md §10)

The channels, in output order, are TIMESTAMP on V[range19], RANGE16 on V[range16], GENERIC on S[7..10] and DECODER on S[0..7]. next_pc_even is the low halfword lo halved, (lo + p)/2 and far above 2^16 when lo is odd. Because it scales next_pc, the constructor runs lookup::check_copowers (lookup.md §11) over (next_pc, m_pc).

5 Why it is sound#

On a live row, m_pc = 1, the decoder lookup makes the claimed tuple the table's row at pc, so pc is even, seq is below 2^24 and exactly one b_k is 1 (lookup.md §10). The mask and address rules make the frame's queries the instruction's (execution-trace.md §4): jal reads nothing, and a branch has no rd query, so nothing it computes is written. An absent operand reads 0, so cmp_rhs is rs2 or the immediate, never their sum, and the comparison's pairs make both operands words. So lt is the ISA's order (§3.2), eq its equality (§3.1), and taken its branch decision: eq on beq, 1 − eq on bne, lt on blt and bltu, 1 − lt on bge and bgeu, and 0 off the branches, every term of taken_rule carrying a branch bit. taken is a committed bit because, inlined, taken·(pc + imm) would be degree 3.

next_pc is held by one gate:

next_pc + 2^32·pc_wrap = (1 − taken − b_jal − b_jalr)·seq
                       + (taken + b_jal)·(pc + imm)
                       + b_jalr·(v_rs1 + imm − jalr_drop)

At most one of taken, b_jal, b_jalr is 1, so one sum is selected, and one wrap bit outside the selectors serves all three: imm is a two's-complement word, so every backward branch and jump wraps, not only jalr. With next_pc an even word and pc_wrap, jalr_drop boolean:

  • the default arm is seq, below 2^24, so pc_wrap = 0;
  • pc + imm and v_rs1 + imm are below 2^33, so one wrap bit holds the carry, uniquely;
  • on jalr, v_rs1 + imm − jalr_drop − 2^32·pc_wrap is a unique even word, (rs1 + imm) mod 2^32 with bit 0 cleared; a false jalr_drop makes next_pc odd or negative.

A branch's or jal's target is even unchecked, pc and imm both being even. Evenness is what keeps the family off HALT_PC = 1 (memory.md §5): without next_pc_even, a jalr whose rs1 + imm ≡ 1 keeps bit 0 and writes HALT_PC with every other gate and lookup holding, and a program that would crash by jumping to address 0 is proven to exit cleanly.

The link is seq, a table value and not a sum, so it has no wrap bit; the rd pair range-checks it and lt like every register write, and the x0 rule masks both at x0. A target needs no check of its own: at an address holding no instruction, the next row's decoder lookup fails whatever family claims the row, no table holding a live row there (lookup.md §10).

On a padding row, m_pc = 0, the mask rules zero every query mask, eq is 0 and every lookup is off, so the row reaches no memory event whatever its free bits hold.

Auditoren/Befehlsfamilien

Die Familie SHIFT_BITWISE

Normative Spezifikationdocs/spec/shift-bitwise.mdAls Markdown anzeigen

Zusammenfassung

Der Schaltkreis der Familie 2: die sechs Shifts und die bitweisen Operationen and, or und xor. Ein Shift in jede Richtung ist ein einziges Produkt mit einer Zweierpotenz, die in der generischen Tabelle nachgeschlagen wird; AND besteht aus vier Byte-Lookups, OR und XOR sind Linearformen über dem AND-Ergebnis. Die Seite spezifiziert die Spalten, die Tabellen, warum jeder Schlüssel seine eigene Schranke tragen muss, die Copower-Prüfung, die Gates und Lookups sowie das Soundness-Argument.

Der folgende normative Text wird auf Englisch gepflegt, der kanonischen Sprache der Spezifikation.

The circuit of the shifts sll, slli, srl, srli, sra, srai and the bitwise and, andi, or, ori, xor, xori, one family, beside the memory frame (memory.md §2). A shift either way is one product with a looked-up power of two; AND is four byte lookups, and OR and XOR are linear forms over it. The circuit is constraints::shift_bitwise::artifact (crates/constraints/src/shift_bitwise.rs); prover::family_fill writes its witness.

1 What the circuit reads from the decoded table#

The tuple is pc next_pc rs1 rs2 rd imm extra_mask (program.md §5), next_pc the fall-through, seq below. imm is the shamt of slli, srli and srai, below 32 because the decoder refuses shamt[5] on RV32; the sign-extended immediate, as a word, of andi, ori and xori; and 0 on a register form. extra_mask is one-hot over constants::extra_mask::shift_bitwise, the legal masks its twelve one-bit values (shift_bitwise::LEGAL_MASKS):

bit    0     1     2     3     4    5     6    7    8    9    10   11
kind   slli  xori  srli  srai  ori  andi  sll  xor  srl  sra  or   and

The second operand of all twelve is src2 = rs2 + imm: an immediate form has no rs2 query, so rs2 reads 0, and a register form's imm is 0. One addend is always zero, so the sum needs no wrap bit, and an immediate never enters the rs2 column the memory argument ties. The circuit commits the twelve bits b_k, and its flags are linear forms over them:

left    b_slli + b_sll
right   b_srli + b_srai + b_srl + b_sra
arith   b_srai + b_sra
t1      b_or + b_ori + b_xor + b_xori
t2      b_and + b_andi − b_or − b_ori − 2·(b_xor + b_xori)

Only the two halves' sums, f_shift and f_bitwise, are columns: each selects lookups, and a selector is a committed boolean (lookup.md §2).

2 Columns#

The frame is M[0..21] and W[0..7], over the queries pc rs1 rs2 rd at slots 0–3 (memory.md §2); its W[6], rd_selected (sel below), holds the value the instruction computes, which the x0 rule masks into the write. The circuit adds:

column name value
W[7..13] decoded_next_pc … decoded_mask the claimed decoded row after pc
W[13..25] kind_slli … kind_and the bits b_k, in §1's order
W[25], W[26] f_shift, f_bitwise the two halves
W[27], W[28] rs1_hi, rs1_sign rs1 >> 16, rs1 >> 31
W[29] src2_hi src2 >> 16
W[30] amount src2 & 31
W[31], W[32] pow, copow 2^amount, 2^(31 − amount) on a shift row
W[33], W[34] high, high_hi src2 >> 5, and its high halfword
W[35] se arith·rs1_sign
W[36], W[37] shift_in, shift_prod both directions' multiplicand, and shift_in·pow
W[38], W[39] ovf, ovf_hi a left shift's discarded high word, and its high halfword
W[40], W[41] residue, residue_hi a right shift's remainder, and its high halfword
W[42], W[43] scaled, scaled_hi residue·2^(32 − amount), and its high halfword
W[44..52] byte_a<j>, byte_b<j> the bytes of rs1, then of src2, low first
W[52..56] byte_and<j> their bytewise AND
W[56] rd_hi sel >> 16
W[57..61] mult_timestamp … mult_decoder one multiplicity per channel, in channel order
S[0..10], V[range19], V[range16] as in jump-branch-slt.md §2

21 M, 61 W and 10 S columns, 92 committed; 48 enforcing gates, the frame's 10 and §4's 38; 39 lookups: 8 TIMESTAMP, 24 RANGE16, 6 GENERIC, 1 DECODER, counts artifact asserts.

3 Tables, and the bound on every key#

3.1 ShiftPowers#

Row s of the packed table's top sub-table (lookup.md §9) is (SHIFT_BASE + s + 1, 2^s, 2^(31 − s)), one for each of the 32 shift amounts and for none other, so a key past its last row matches nothing. The second value is the copower a residue bound multiplies by, 2^(32 − s), stored halved (SHIFT_COPOWER_BITS = 31): at s = 0 it is 2^32, which the table's u32 columns cannot hold, so the two gates that read it carry the factor 2 (§4.2, §4.3).

3.2 The AND rows#

An AND row is (AND_BASE + a + 1, b, a & b) over bytes a and b, so a key inside their range matches a row that makes byte_b<j> a byte and byte_and<j> its AND with byte_a<j>: those two need no bound of their own.

3.3 Every key is bounded#

The channel proves membership of the packed table, not of a sub-table (lookup.md §4), so an out-of-range key lands on another sub-table's row. A bitwise row with byte_a0 = 65,823 gates to key 65,824, ShiftPowers' row (65,824, 2^31, 1); with rs1 = 65,823 and rs2 = 2^31 every gate holds, and and writes 1 where the answer is 0. So every key carries its own bound, as RANGE16 obligations under its lookup's selector:

key bound obligations selector
rs1_hi + SIGN_BASE rs1_hi < 2^16 rs1's 16+16 pair m_pc
amount + SHIFT_BASE amount < 2^5 amount; 2^11·amount f_shift
byte_a<j> + AND_BASE byte_a<j> < 2^8 byte_a<j>; 2^8·byte_a<j> f_bitwise

A bound below a halfword takes both obligations: the scaled one alone does not make the key an integer (lookup.md §11), and the direct one alone admits every halfword, byte_a0 = 256 landing on U16GetSign's row (257, 0, 0).

3.4 The copower check#

artifact runs lookup::check_copowers (lookup.md §11) over each column it bounds by scaling, which must carry its direct bound under its scaled obligation's own selector: residue, scaled by the looked-up copower (§4.3), under m_pc; amount under f_shift; each byte_a<j> under f_bitwise.

4 Gates#

Gate list 0 holds the frame's ten and these 38. m_q, a_q, v_q are query q's mask, address and read value, and a flag of §1 times (…) stands for each of its weighted bits times (…), so every term is of degree 2.

4.1 Presence and next_pc#

gate polynomial
kind_<k>_boolean ×12, decoded_mask_bits as in jump-branch-slt.md §4
f_shift_rule, f_bitwise_rule f − Σ its half's six bits
f_shift_boolean, f_bitwise_boolean f − f²
rs1_mask_rule, rd_mask_rule m_q − m_pc·Σ_k b_k
rs2_mask_rule m_rs2 − m_pc·(b_sll + b_srl + b_sra + b_and + b_or + b_xor)
<q>_addr_rule ×3, <q>_value_masked ×2 as in jump-branch-slt.md §4
next_pc_rule next_pc − seq

No kind computes a pc: next_pc is the decoder-bound fall-through, with no wrap bit and no bound of its own, and HALT_PC is beyond the family's reach (memory.md §5).

4.2 The shift amount#

amount_split    rs2 + imm − 32·high − amount
copower_rule    pow·copow − 2^31·f_shift

amount_split is ungated. copower_rule says pow·(2·copow) = 2^32 on a shift row, and pow·copow = 0 on a bitwise row.

4.3 The one product, both directions#

se_rule           se − arith·rs1_sign
rs1_sign_boolean  rs1_sign − rs1_sign²
se_boolean        se − se²
shift_in_rule     shift_in − left·v_rs1 − right·(sel − 2^32·se)
shift_prod_rule   shift_prod − shift_in·pow
shift_out_rule    left·(shift_prod − sel − 2^32·ovf)
                    + right·(shift_prod + residue − v_rs1 + 2^32·se)
scaled_rule       scaled − 2·residue·copow

shift_prod_rule, ungated, is the one multiplication by pow; shift_in_rule picks its multiplicand, which keeps shift_out_rule at degree 2 where left·(v_rs1·pow − …) would be 3, and se is committed for the same reason. A right shift is the floor division rs1 − 2^32·se = (sel − 2^32·se)·2^s + residue, which covers sra: the arithmetic shift of a negative word is the floor division of its signed value, and the result keeps the operand's sign. shift_in and shift_prod are the only columns that are not words: the multiplicand is negative where se = 1, and a left shift's product reaches 2^63.

4.4 The bitwise half#

rs1_bytes         v_rs1 − Σ_j 2^(8j)·byte_a<j>
src2_bytes        rs2 + imm − Σ_j 2^(8j)·byte_b<j>
bitwise_out_rule  f_bitwise·sel − t1·(v_rs1 + rs2 + imm) − t2·Σ_j 2^(8j)·byte_and<j>

Per byte, OR is a + b − (a & b) and XOR is a + b − 2·(a & b). Summed by weight through the two decompositions, sel is rs1 & src2 at (t1, t2) = (0, 1), their OR at (1, −1) and their XOR at (1, −2), exactly, no carry crossing a byte: there is no OR or XOR table, and the AND accumulator is a linear form, not a column. sel is gated by f_bitwise because t1 and t2 are 0 on a shift row, where a bare sel would force rd = 0. The decompositions are ungated: on a shift row the bytes carry no lookup, and a decomposition always exists.

4.5 Lookups#

After the frame's 8 TIMESTAMP obligations, in the channel order of jump-branch-slt.md §4.1:

RANGE16   <x>_hi_range, <x>_lo_range     under m_pc, x = rs1 src2 high ovf residue scaled rd
          amount_range, amount_scaled    under f_shift      §3.3
          byte_a<j>_range, _scaled ×4    under f_bitwise    §3.3
GENERIC   rs1_get_sign   (rs1_hi + SIGN_BASE, rs1_sign, 0)                 under m_pc
          shift_powers   (amount + SHIFT_BASE, pow, copow)                 under f_shift
          and_byte_<j>   (byte_a<j> + AND_BASE, byte_b<j>, byte_and<j>)    under f_bitwise
DECODER   decode_row     under m_pc (lookup.md §10)

5 Why it is sound#

On a live row the decoder lookup makes the claimed tuple the table's row at pc, so one kind bit is 1 and one of f_shift, f_bitwise (lookup.md §10); the mask and address rules make the queries the instruction's (execution-trace.md §4). rs1, src2 and sel are words by their pairs, rs1_hi is rs1's true high halfword and rs1_sign its bit 31. Every term of §4 that reads sel carries a shift bit or f_bitwise, so the inactive half never constrains it.

  • The amount is the ISA's. §3.3 bounds amount below 32 and high's pair bounds high below 2^32, so amount_split is an integer identity below 2^37, amount = src2 mod 32, and the ShiftPowers row it keys gives pow = 2^amount. Without high's pair, sll by rs2 = 4 can shift by 8, at high = −1/8.
  • A left shift: shift_prod = rs1·2^s < 2^63, and sel + 2^32·ovf, both words, is its unique split, so sel = (rs1·2^s) mod 2^32.
  • A right shift: se is rs1's bit 31 on sra and srai and 0 otherwise, so se_rule alone keeps an srai from carrying srli's answer. Every term of the floor division is below 2^64 in magnitude, so residue is the integer (rs1 − 2^32·se) − (sel − 2^32·se)·2^s, and scaled's pair puts it in [0, 2^s): sel − 2^32·se is the floor of (rs1 − 2^32·se)/2^s. residue's own pair, which check_copowers requires, bounds it without appeal to sel's, the scaled pair alone saying nothing of a non-integer: 2^−28 passes it at s = 3.
  • A bitwise result: each byte_a<j> is below 256, so its lookup matches an AND row (§3.2); with rs1 and src2 words, both decompositions are the unique byte splits and §4.4's identity holds.

copower_rule is implied by the bounded key and kept as the circuit's own reading of the table: a ShiftPowers row generated wrong stops the honest prover rather than license a residue bound that is not one. It also confines the key to ShiftPowers alone, no other row's two values having the product 2^31: an AND row's is at most 255·255, every other row's 0.

On a padding row, m_pc = 0, every query mask is 0 and every obligation under m_pc vacuous. f_shift and f_bitwise are free booleans there, so a padding row may look up ShiftPowers or the AND rows, which consumes a multiplicity and changes nothing.

Auditoren/Befehlsfamilien

Die Familie MUL_DIV

Normative Spezifikationdocs/spec/mul-div.mdAls Markdown anzeigen

Zusammenfassung

Der Schaltkreis der Familie 3, der M-Erweiterung. Eine einzige Produktidentität dient allen vier Multiplikationen und der Division; eine Vorzeichenregel und eine per Range-Check geprüfte Differenz machen die Division abschneidend statt abrundend; und ein einziges Gate legt die Division durch null fest. Die Seite spezifiziert die Spalten, die Vorzeichenanpassungen, die Gates und das Soundness-Argument, einschließlich des Falls des vorzeichenbehafteten Überlaufs und der ehrlichen Befüllung.

Der folgende normative Text wird auf Englisch gepflegt, der kanonischen Sprache der Spezifikation.

The M extension — mul, mulh, mulhsu, mulhu, div, divu, rem, remu — as one circuit beside the memory frame every execution family carries (memory.md §2). One product identity serves the four multiplies and the division; a sign rule and a range-checked gap make the division truncated, and one gate pins division by zero. constraints::mul_div builds it: 21 M, 54 W and 9 S columns, 54 enforcing gates, 27 lookups.

1 What the circuit reads from the decoded table#

Every M instruction is R-type, so the decoded tuple has no imm: pc next_pc rs1 rs2 rd extra_mask, six columns (program.md §5). extra_mask is one-hot over constants::extra_mask::mul_div, bits 0–7 in the order above; the legal masks are its eight single bits (mul_div::LEGAL_MASKS), which the table's domain enforces (lookup.md §10). The circuit commits the bits b_k and reads every signal as a linear form over them:

signal form
reads rs1 signed b_mul + b_mulh + b_mulhsu + b_div + b_rem
reads rs2 signed b_mul + b_mulh + b_div + b_rem
a multiply, Σ_mul b_mul + b_mulh + b_mulhsu + b_mulhu
a division, f_div b_div + b_divu + b_rem + b_remu, a column: the is-zero gadgets' enable

mul is read signed × signed: its low word is the same either way, which lets one product identity serve all four multiplies. mulhsu's asymmetry is the two lists, not a case split.

2 Columns#

The frame is pc rs1 rs2 rd, M[0..21] and W[0..7] (memory.md §2.1); below, m_q, a_q and v_q are query q's mask, address and read value, and rs1, rs2 the operands' read values. The family adds:

W[7..12]   decoded_next_pc decoded_rs1 decoded_rs2 decoded_rd decoded_mask    (no imm)
W[12..20]  kind_mul … kind_remu
W[20]      f_div
W[21..27]  rs1_hi rs1_top rs2_hi rs2_top s1 s2          high halfwords, bit 31, §3
W[27..34]  mx my p_low p_low_hi p_high p_high_hi p_sign  the product
W[34..40]  q q_hi q_sign r r_hi r_sign                   quotient and remainder
W[40..45]  r_inv rz d1 d_inv dz                          is_zero(r), f_div·s1, is_zero(rs2)
W[45..50]  abs_r abs_d gap gap_hi rd_hi
W[50..54]  mult_timestamp mult_range16 mult_generic mult_decoder
S[0..6]    the decoded table, bound by identity
S[6..9]    the packed generic table (lookup.md §9)
V          range19 range16

The fill keeps mx, my (signed) and r_inv, d_inv (inverses) in Fr, every other column in u32.

3 The sign adjustments#

rs1_adj = rs1 − 2^32·s1      s1 = (b_mul + b_mulh + b_mulhsu + b_div + b_rem)·rs1_top
rs2_adj = rs2 − 2^32·s2      s2 = (b_mul + b_mulh + b_div + b_rem)·rs2_top
q_adj   = q − 2^32·q_sign    r_adj = r − 2^32·r_sign

rs1_top is the U16GetSign lookup of rs1_hi, which rs1's 16+16 pair makes its true high halfword, so the key lies in that sub-table's range and the answer is bit 31 (lookup.md §4); rs2_top likewise. An unsigned position forces its adjustment to 0 whatever the top bit, which keeps the selection degree 2. q_sign and r_sign are not sign lookups (§5.3). f_div, rs1_top, rs2_top, s1, s2, p_sign, q_sign and r_sign carry booleanity gates; rz and dz are boolean by the is-zero gadget (jump-branch-slt.md §3), d1 as a product of booleans.

4 Gates#

The frame's ten enforcing gates (memory.md §2.4) and the family's 44, all in gate list 0, each formula = 0. The plumbing:

kind_<k>_boolean          b_k − b_k²                       eight
decoded_mask_bits         Σ_k 2^k·b_k − decoded_mask
<q>_mask_rule             m_q − m_pc·Σ_k b_k               rs1, rs2, rd: every kind uses all three
<q>_addr_rule             m_q·(a_q − decoded_q)            rs1, rs2, rd
<q>_value_masked          v_q − m_q·v_q                    rs1, rs2
next_pc_rule              next_pc − decoded_next_pc        the fall-through (memory.md §5)

The arithmetic, mul_div::arithmetic_gates(32), written with §3's abbreviations:

f_div_rule, s1_rule, s2_rule   §1's and §3's forms, and eight booleanity gates (§3)
mx_rule                mx − Σ_mul b·rs1_adj − f_div·rs2_adj
my_rule                my − Σ_mul b·rs2_adj − f_div·q_adj
product_rule           mx·my − p_low − 2^32·p_high + 2^64·p_sign
division_rule          f_div·(p_low + 2^32·p_high − 2^64·p_sign + r_adj − rs1_adj)
rz_inverse             r·r_inv + rz − f_div          rz_at_nonzero   rz·r
dz_inverse             rs2·d_inv + dz − f_div        dz_at_nonzero   dz·rs2
d1_rule                d1 − f_div·s1
r_sign_rule            r_sign − d1 + d1·rz           so r_sign = f_div·s1·(1 − [r = 0])
abs_r_rule             abs_r − r − 2^32·r_sign + 2·r·r_sign       abs_r = |r_adj|
abs_d_rule             abs_d − rs2 − 2^32·s2 + 2·rs2·s2           abs_d = |rs2_adj|
gap_rule               gap − f_div·(abs_d − abs_r − 1) − 2^32·dz
zero_divisor_quotient  dz·(q − (2^32 − 1))
rd_value_rule          rd_selected − b_mul·p_low − (b_mulh + b_mulhsu + b_mulhu)·p_high
                         − (b_div + b_divu)·q − (b_rem + b_remu)·r

The lookups: the frame's eight TIMESTAMP gap chunks, each under its query's mask; and under m_pc, 16+16 RANGE16 pairs on rs1, rs2, p_low, p_high, q, r, gap and rd_selected, rs1_get_sign, (rs1_hi + SIGN_BASE, rs1_top, 0) on GENERIC, and rs2_get_sign, and decode_row on DECODER.

The width is a parameter of arithmetic_gates so the encoding can be checked whole: crates/checker/tests/mul_div.rs evaluates arithmetic_gates(4) through gkr::eval_gate over every (dividend, divisor) pair of a 4-bit word and each division kind, and exactly one (q, r) survives, RV32M's.

5 Why it is sound#

On a live row the decoder lookup makes exactly one kind bit 1 (lookup.md §10), and rs1, rs2 are words whose _top is bit 31, so rs1_adj, rs2_adj ∈ [−2^31, 2^32) are the operands as the kind reads them.

5.1 The product#

On a multiply row mx·my = rs1_adj·rs2_adj; on a division row it is rs2_adj·q_adj, q's pair and q_sign's booleanity putting q_adj in [−2^32, 2^32). Either way |mx·my| < 2^64, and two words and a boolean cover [−2^64, 2^64) once, so product_rule holds over the integers with one solution: p_low, p_high are the words of the 64-bit two's-complement product, RV32M's for each multiply. product_rule is ungated and the circuit's only product of two row values, which is what lets both readings share it at degree 2.

5.2 The division#

With rs2_adj ≠ 0, division_rule is rs2_adj·q_adj + r_adj = rs1_adj over the integers. Truncated division is its one solution with |r_adj| < |rs2_adj| and r_adj zero or of the dividend's sign, and two gates state exactly that:

  • The sign. r_sign = f_div·s1·(1 − [r = 0]) makes r_adj the word r on an unsigned row or a non-negative dividend, and r − 2^32 < 0 on a negative one unless r = 0. It is what separates truncated division from floored: without it DIV(−7, 2) admits q = −4, r = 1 as readily as q = −3, r = −1. As a definition, through d1, it is degree 2.
  • The magnitude. gap = abs_d − abs_r − 1 is range-checked, and neither magnitude reaches 2^32: abs_d ≤ 2^31 where s2 = 1, abs_r ≤ 2^32 − 1 where r_sign = 1, which needs r ≠ 0, and each is a word elsewhere. So the difference lies in [−2^32, 2^32), in range exactly when |r_adj| < |rs2_adj|. The comparison gadget would repeat bounds that hold and has no place for the zero divisor's term.

So q_adj and r_adj are RV32M's, and q_sign is pinned only by q's range: one value puts q_adj + 2^32·q_sign in [0, 2^32).

A zero divisor makes dz = 1 and rs2_adj = 0: the identity leaves r_adj = rs1_adj, so r is the dividend's word; zero_divisor_quotient, the one pin, makes q all ones; the 2^32·dz term lifts gap to 2^32 − 1 − abs_r, so the divisor imposes no bound. q_sign is free and harmless: mx = 0, and rd reads the word q.

The identity is gated. On a multiply row r_sign = 0 and r is a word, so an ungated identity would demand rs1_adj − rs1_adj·rs2_adj ∈ [0, 2^32), false for nearly every multiply: 7 × 3, a negative rs1 times x0.

5.3 The signed overflow, and why q_sign is free#

DIV(−2^31, −1) needs no pin: |r_adj| < 1 forces r = 0, the identity q_adj = 2^31, and q's range q_sign = 0, q = 0x80000000, RV32M's answer; REM gives 0. This row is why q_sign is a free boolean: tied to bit 31 of q, as s1 and s2 are to their operands', it would force q_adj = −2^31 and make the row unprovable. r_sign likewise follows the dividend's sign, not the remainder's word.

rd_selected's pair is implied by its four sources' and kept, every family bounding what it writes to rd (memory-ops.md §5). Every gate is zero on the all-zero padding row, which mul_div::artifact asserts with each channel's obligation count.

5.4 The fill#

prover::family_fill(MUL_DIV) (crates/prover/src/fill.rs) computes the witness with Rust's integers: the product in i128, the division by wrapping_div and wrapping_rem, which give RV32M's overflow answer, with the zero divisor an arm of its own, and q_sign from the sign of q_adj. It writes the computed value to rd_selected, which the frame's x0 rule masks, and panics, on rows the emulator cannot produce, if the identity does not divide, a product exceeds two words, or the trace's rd write or next_pc is not what the instruction computes.

Auditoren/Befehlsfamilien

Die Familien für Speicheroperationen

Normative Spezifikationdocs/spec/memory-ops.mdAls Markdown anzeigen

Zusammenfassung

Die Schaltkreise der Familien 4, 5 und 6: Laden und Speichern von Wörtern, Laden und Speichern von Bytes und Halbwörtern sowie die atomaren Operationen. Die Seite spezifiziert ihre gemeinsame Adressierung, die Kopien von MEM_WORD, das Einsetzen eines Teilworts in sein Wort durch MEM_SUBWORD, die Induktion auf der Schreibseite, durch die jeder Register- und RAM-Wert ein 32-Bit-Wort bleibt, und ATOMICS, einschließlich der einzigen Abweichung von RV32IMAC: sc.w gelingt immer.

Der folgende normative Text wird auf Englisch gepflegt, der kanonischen Sprache der Spezifikation.

MEM_WORD (lw, sw), MEM_SUBWORD (lb, lh, lbu, lhu, sb, sh) and ATOMICS (lr.w, sc.w, the nine AMOs): the execution families whose rows touch RAM, each a circuit beside the memory frame (memory.md §2), sharing §2's addressing. They are constraints::{mem_word, mem_subword, atomics}, filled by prover::family_fill (crates/prover/src/fill.rs), which computes each witness with Rust's integer operations.

family M W S gates, frame + own TIMESTAMP, RANGE16, GENERIC, DECODER
MEM_WORD 31 24 7 13 + 20 12, 5, 0, 1
MEM_SUBWORD 31 55 10 13 + 40 12, 22, 1, 1
ATOMICS 26 54 9 11 + 35 10, 19, 6, 1

Each artifact asserts its gate and obligation counts and that the all-zero padding row satisfies every gate. Below, m_q, a_q and v_q are query q's mask, address and read value, rs1 and rs2 the operands' read values (memory.md §2.1), and b_k (b_lw, b_lr, …) the committed kind bits.

1 What the circuits read from the decoded table#

MEM_WORD's and MEM_SUBWORD's tuple is pc next_pc rs1 rs2 rd imm extra_mask, imm the offset's two's-complement u32; ATOMICS' has no imm, its address being rs1 (program.md §5). The tuple is the first setup columns, and the packed generic table follows it where a family reads one: S[7..10] in MEM_SUBWORD, S[6..9] in ATOMICS (lookup.md §9). extra_mask is one-hot over constants::extra_mask, bit k the k-th mnemonic below, and each module's LEGAL_MASKS is those single bits:

mem_word      lw sw
mem_subword   lb lh lbu lhu sb sh
atomics       amoadd amoswap lr sc amoxor amoor amoand amomin amomax amominu amomaxu

The atomics order is ascending funct5; aq and rl order nothing on one hart and are not recorded. MEM_SUBWORD's modifiers are linear forms over its bits:

LOADK = b_lb + b_lh + b_lbu + b_lhu     BYTE = b_lb + b_lbu + b_sb     SIGNEXT = b_lb + b_lh
STORE = b_sb + b_sh                     HALF = b_lh + b_lhu + b_sh

All three carry the same plumbing, each formula = 0:

kind_<k>_boolean    b_k − b_k²
decoded_mask_bits   Σ_k 2^k·b_k − decoded_mask
<q>_mask_rule       m_q − m_pc·uses_q              every query but pc
<q>_addr_rule       m_q·(a_q − decoded_q)          rs1, rs2, rd
                    m_q·(a_q − 4·word_index)       load, ram (§2)
<q>_value_masked    v_q − m_q·v_q                  rs1, rs2
next_pc_rule        next_pc − decoded_next_pc      the fall-through (memory.md §5)
uses_q rs1 rs2 load ram rd
MEM_WORD b_lw + b_sw b_sw b_lw b_sw b_lw
MEM_SUBWORD LOADK + STORE STORE LOADK STORE LOADK
ATOMICS every bit every bit but b_lr no query every bit every bit

m_rs2 is keyed on b_lr, the one kind without an rs2 field, and not on rs2 = x0: an amoadd.w whose rs2 is x0 still reads it.

2 Addressing#

The effective address is rs1 + imm mod 2^32, or rs1 for an atomic. One degree-1 gate splits it, with wrap, bit0 and bit1 boolean:

MEM_WORD      addr_split   rs1 + imm − 2^32·wrap − 4·word_index
MEM_SUBWORD   addr_split   rs1 + imm − 2^32·wrap − 4·word_index − 2·bit1 − bit0
ATOMICS       addr_word    rs1 − 4·word_index

Over Fr that says nothing, 4 being a unit. Three RANGE16 obligations under m_pc, on word_index_hi, word_index − 2^16·word_index_hi and 4·word_index_hi (word_index_hi_range, word_index_lo_range, word_index_hi_scaled), cap word_index at 2^30 − 1, the top word's. With rs1 a word (§5) and imm a table value the split is then one of integers: wrap is the true carry, bit1 and bit0 the true low bits, and every RAM address is a 4-aligned address below 2^32. Having no offset bits, a misaligned MEM_WORD or ATOMICS access needs a word_index that is not an integer, which its pair refuses; the emulator refuses it first (execution-trace.md §10). addr_word derives rs1 < 2^32 rather than assuming it. half_aligned, HALF·bit0 = 0, refuses a halfword at an odd address and keeps w·p a divisor of 2^32 (§4.3).

Every RAM query's address is 4·word_index, so byte, halfword, word and atomic accesses to one word name one cell; the byte position lives only in MEM_SUBWORD's splice. Confining an access to initialized memory is the multiset's (memory.md §9): an out-of-window access fails the statement's memory argument, not a gate.

3 MEM_WORD#

A load copies the word into rd, a store copies rs2 into the word; there is no splice, no generic lookup, and the decoded table is the only setup.

W[9..15]   decoded_next_pc decoded_rs1 decoded_rs2 decoded_rd decoded_imm decoded_mask
W[15..21]  kind_lw kind_sw wrap word_index word_index_hi rd_hi
W[21..24]  mult_timestamp mult_range16 mult_decoder

wrap_boolean, addr_split (§2)
rd_value_rule       rd_selected − b_lw·load_read_value
store_value_rule    ram_write_value − m_ram·rs2

Its RANGE16 obligations are §2's three and the 16+16 pair on rd_selected, and the two copies are its whole semantics. rd_selected is range-checked although it copies a RAM word, because a RAM word need not be a word, advice's initial values being bound to nothing (public-values.md §6): the pair keeps every register value a word without reference to RAM (§5). No gate reads ram_read_value, the word a store overwrites; the memory argument alone pins it.

4 MEM_SUBWORD#

4.1 The splice#

A sub-word's position in its word lives only in

word = high·(w·p) + sub·p + low      p = 2^(8·offset), offset = 2·bit1 + bit0
                                     w, the access width: 2^8 if BYTE, 2^16 if HALF

p and its copower are degree-2 forms in the offset bits, written as gates rather than looked up:

p_rule        p − m_pc − 255·bit0 − 65535·bit1 − K·bit0·bit1      K = 2^24 − 2^16 − 2^8 + 1
pcopow_rule   p·pcopow − 2^31·m_pc                                 pcopow = 2^31/p
wph_rule      wph − 32768·p + 32640·BYTE·p                         wph = w·p/2
p_ram_rule    p_ram − m_ram·p

p_rule takes the four offsets to 1, 2^8, 2^16, 2^24, m_pc standing for the constant so the all-zero row satisfies it. The copower and w·p are stored halved so that 2^32 fits a u32 column, the gates reading them carrying the factor 2, as ShiftPowers' do (lookup.md §9). p_ram keeps store_rule degree 2. A table keyed by the offset would pin nothing addr_split does not, and add a key to bound (lookup.md §4).

4.2 Columns and gates#

W[9..21]   the decoded row; kind_lb … kind_sh
W[21..31]  wrap word_index word_index_hi bit0 bit1 p pcopow wph p_ram word
W[31..42]  high high_hi high_scaled high_scaled_hi sub sub_scaled sub_scaled_hi
           low low_hi low_scaled low_scaled_hi
W[42..51]  src_sub src_sub_scaled src_sub_scaled_hi src_high src_high_hi sign_in sign se rd_hi
W[51..55]  the four multiplicities

Its gates, beside the plumbing: wrap_boolean, bit0_boolean, bit1_boolean, addr_split, half_aligned, §4.1's four, and

word_rule            word − LOADK·load_read_value − STORE·ram_read_value
splice_rule          word − high_scaled − sub·p − low
high_scaled_rule     high_scaled − 2·high·wph                             = high·w·p
sub_scaled_rule      sub_scaled − 2^16·sub − (2^24 − 2^16)·BYTE·sub       = sub·2^32/w
low_scaled_rule      low_scaled − 2·low·pcopow                            = low·2^32/p
src_sub_rule         rs2 − src_sub − 2^16·src_high + 65280·BYTE·src_high
src_sub_scaled_rule  src_sub_scaled − 2^16·src_sub − (2^24 − 2^16)·BYTE·src_sub
store_rule           ram_write_value − m_ram·word − (src_sub − sub)·p_ram
sign_in_rule         sign_in − sub − 255·BYTE·sub                         = 2^8·sub or sub
se_rule              se − SIGNEXT·sign
rd_value_rule        rd_selected − LOADK·sub − (2^32 − 2^16)·se − 65280·BYTE·se

mem_subword::splice_gates(byte_bits) builds the twelve whose literals depend on the byte width — §4.1's first three and these but word_rule and se_rule — and the circuit takes it at BYTE_BITS = 8. Its RANGE16 obligations, all under m_pc, are §2's three, 16+16 pairs on high, high_scaled, sub_scaled, low, low_scaled, src_sub_scaled, src_high and rd_selected, and one obligation each on sub, src_sub and sign_in; its GENERIC lookup is sub_get_sign, (sign_in + SIGN_BASE, sign, 0).

4.3 Why it is sound#

§2 fixes the offset bits and half_aligned clears bit0 at halfword width, so p and w are the access's. Each part has a direct bound and a scaled one: high < 2^32 makes high·w·p an integer, sub_scaled < 2^32 is sub < w and low_scaled < 2^32 is low < p. So splice_rule holds over ℤ with one solution, the base-(p, w) digits of the word, and a word not below 2^32 has none. A scaled bound alone admits non-integers, its scale being a unit of Fr (lookup.md §11); constraints::lookup::check_copowers holds word_index_hi, high, sub, low and src_sub to their direct bounds, one obligation being exact for sub and src_sub, both below w ≤ 2^16. src_high's pair makes rs2 = src_sub + w·src_high integral, so src_sub is rs2 mod w: without it sb could store a byte unrelated to rs2.

A load writes rd = sub + (2^32 − w)·se: the sub-word, or at se = 1 its two's-complement extension (lb of 0x88 is 0xffffff88). sign_in is 2^8·sub for a byte and sub for a halfword, so its bit 15 is the sign at either width and one U16GetSign lookup serves both; its own obligation bounds the key into that sub-table (lookup.md §4). se is a one-hot sum times a table bit, boolean without a gate.

A store writes word + (src_sub − sub)·p = high_scaled + src_sub·p + low, a word with no appeal to memory: high_scaled is a multiple of w·p below 2^32 and w·p divides 2^32 (a halfword at offset 3 would make it 2^40; half_aligned excludes it), so high_scaled ≤ 2^32 − w·p and src_sub·p + low ≤ w·p − 1. That is why high_scaled keeps its own pair.

crates/checker/tests/mem_subword.rs checks the splice whole at a 4-bit word: for every word, admissible offset and width, splice_gates(1) and the bounds admit exactly one (high, sub, low).

5 The write-side induction#

A circuit may use a register operand as a word without bounding it. That rests on two facts:

  • Every register write is a word on its own row. Every execution family's rd_selected carries a 16+16 pair under m_pc, but ATOMICS', which is the old word or 0, the old word bounded by its comparison's pair under m_pc (§6). The frame writes (1 − z)·rd_selected (memory.md §2.4), registers start at 0 and a read returns the last write (memory.md §9), so every register read is a word, with no appeal to RAM.
  • Every RAM write of an execution family is a word: MEM_WORD writes rs2, a register value; MEM_SUBWORD bounds its merged word itself (§4.3); each ATOMICS arm is bounded (§6); a read-only query writes back what it read.

RAM's initial values are words — the image's, 0, the public input's — but advice's, which nothing bounds. No execution family relies on a RAM word being one: each bounds the value it uses, by MEM_WORD's rd pair, MEM_SUBWORD's splice or ATOMICS' comparison, so a row using a non-word is unprovable. The register half is what every carry needs: a + b − 2^32·wrap is a reduction only for words (memory.md §7), and addr_split's integer argument needs rs1 < 2^32.

6 ATOMICS#

One row is one read-modify-write: the ram query reads old and writes new at Δ = 3, beside rd (execution-trace.md §4), lr.w included, which writes its word back.

W[8..13]   decoded_next_pc decoded_rs1 decoded_rs2 decoded_rd decoded_mask    (no imm)
W[13..24]  kind_amoadd … kind_amomaxu
W[24..30]  word_index word_index_hi sum sum_hi add_wrap f_bitwise
W[30..42]  byte_a0..3 byte_b0..3 byte_and0..3       old's bytes, rs2's, their AND
W[42..50]  old_hi old_sign src_hi src_sign lt cmp_gap cmp_gap_hi lo
W[50..54]  the four multiplicities

With A = Σ_j 2^(8j)·byte_and_j inlined, its gates beside the plumbing are:

ram_value_rule    new − b_lr·old − (b_sc + b_amoswap)·rs2 − b_amoadd·sum − b_amoand·A
                    − b_amoor·(old + rs2 − A) − b_amoxor·(old + rs2 − 2A)
                    − (b_amomin + b_amominu)·lo − (b_amomax + b_amomaxu)·(old + rs2 − lo)
rd_value_rule     rd_selected − Σ_{k ≠ sc} b_k·old
add_rule          old + rs2 − sum − 2^32·add_wrap
f_bitwise_rule    f_bitwise − b_amoand − b_amoor − b_amoxor
old_bytes_rule    old − Σ_j 2^(8j)·byte_a_j          src_bytes_rule   rs2 − Σ_j 2^(8j)·byte_b_j
lo_rule           lo − rs2 − lt·(old − rs2)
addr_word (§2); add_wrap_boolean, f_bitwise_boolean; cmp_order, cmp_lt_boolean (below)

Each takes a kind's bit through its constants::extra_mask constant, from which the table's masks are built too, so a transposed arm would pass the decoder lookup. Under m_pc the family looks up the comparison's pairs on old, rs2 and cmp_gap and its two signs, §2's three and sum's pair; under f_bitwise, for each j, byte_a_j and 2^8·byte_a_j on RANGE16 and and_byte_j, (byte_a_j + AND_BASE, byte_b_j, byte_and_j), on GENERIC.

The comparison is constraints::gadgets::comparison (jump-branch-slt.md §3) with selector m_pc, lhs = old, rhs = rs2 and signed = [b_amomin, b_amomax]. The family's assemble asserts all four, nothing else in the artifact determining them: signed widened to amominu orders it signed, lhs and rhs swapped turn amomin into a max, and a selector narrowed to the min/max kinds drops old's bound on the other seven, and with it the bound on their rd write (§5). lo is the smaller under the ordering lt settles, and old + rs2 − lo the larger.

Why new is a word. old and rs2 are bounded by the comparison, sum by its own pair (add_rule is ungated: sum = (old + rs2) mod 2^32 on every live row), lo and the larger by being old and rs2. On a bitwise row byte_a_j's pair puts the key in the AND sub-table, whose row bounds byte_b_j and fixes byte_and_j = byte_a_j & byte_b_j; the byte rules are then the operands' decompositions, and A, old + rs2 − A, old + rs2 − 2A are AND, OR and XOR, carry-free byte by byte. Without its pair byte_a0 = 65,823 reads ShiftPowers' (65,824, 2^31, 1) (lookup.md §4); check_copowers takes the four keys under f_bitwise, which covers all three bitwise kinds: under b_amoand alone amoor and amoxor would read free byte_and.

sc.w always succeeds: it stores rs2 and writes 0 to rd, and the machine holds no reservation. The emulator does the same (execution-trace.md §10), and a row claiming failure, a nonzero rd or an unchanged word, is refused by rd_value_rule or ram_value_rule. This is a conformance deviation, not a soundness one: the proof is of what the program did on this machine. A guest may not rely on an sc.w failing where the ISA requires it to: with no valid reservation (no earlier lr.w, or one an earlier sc.w consumed) or at an address outside the reservation set. The lr.w/sc.w retry loop compiled code uses is unaffected, first-pass success being legal on any hart.

Auditoren/Delegationen

Delegation

Normative Spezifikationdocs/spec/delegation.mdAls Markdown anzeigen

Zusammenfassung

Das Delegations-ABI: wie ein Gastprogramm per ecall einen Frame aus RAM-Wörtern an einen Schaltkreis übergibt, die Registry der Delegationstypen, die Frame-Regeln, der Anker, der jede Anfrage über die Speicher-Multimenge mit genau einem Aufruf paart, die Seite des Executors, die statische Deklaration der Familien, die ein Programm aufruft, Shards und Zeitfenster, Höhen und ihre Abwägungen, die Aufrufer auf Seiten der Gastprogramme und die Grenzen der delegierten Menge.

Der folgende normative Text wird auf Englisch gepflegt, der kanonischen Sprache der Spezifikation.

The delegation ABI: how a guest hands a frame of RAM words to a circuit with an ecall, how each request pairs with exactly one invocation, how a program declares the families it calls, and how each family is sized. Frame layouts and circuits are delegation-circuits.md's, the recursion format's four families recursion.md's.

1 What a delegation family is#

A delegation family proves a function of guest memory too costly to run as instructions. It is invoked, never decoded: its number is a run-time value of a7, so it claims no pc and has no decoded table. A row is one invocation, which rides the cycle that requested it and owns no cycle (execution-trace.md §1); its accesses join the one memory multiset; it is in a VmConfig exactly when the image declares it (§7). Otherwise it is an ordinary family, an arm in constraints::family_circuit and a fill in prover::family_fill. A call is one row, the anchor's two leaves being a row's (§5); an operation wider than a row is several calls on one frame, chained through RAM (delegation-circuits.md §1, RAM glue).

2 The calling convention#

A call is an ecall (ecall-abi.md §1): a7 the number, a0 the frame base. It writes 0 to a0 and falls through (execution-trace.md §6); a recursion-format type writes a0 + 4·words instead (recursion.md §1.4).

An executor without a family's circuit answers -ENOSYS, on which a base-format shim's caller computes the same function in software, so an executor may implement any subset of the families; any other nonzero answer is fatal (ecall-abi.md §7).

3 The registry#

constants::delegation::TYPES, also program::DELEGATIONS, is one table of (family, number, anchor space, frame words), ascending by family, which the emulator dispatches on and constraints::add_sub builds its request gates from. The first BASE_TYPES = 6 rows are the base format's (recursion.md §1.2). Why each family has its height is §9's.

family id number anchor space frame words
KECCAK_F 9 0x0507 4 51
POSEIDON2 10 0x0500 5 24
FR_ARITH 11 0x0502 6 25
MOD_MUL 15 0x0504 7 25
SHA256_COMP 16 0x0508 8 25
EC_ADD 17 0x0506 9 97
FR_OP 19 0x0509 11 4
P2_FIELD 20 0x050A 12 5
FIELD_IO 21 0x050B 13 3
FQ_OP 22 0x050C 14 4

constraints::add_sub asserts at compile time that every number is in the precompile range and not EXIT, and that numbers and spaces are pairwise distinct, so an ecall row is the exit or a request of one type; a type costs that circuit a selector is_deleg_<f>, three gates and a term in five shared ones (add-sub.md §2, §4). A type's anchor space is the type: only its requests and invocations touch it, so the anchor's address is the frame base alone. A reserved range of RAM would need an argument that no guest access reaches it.

4 The frame#

A frame is words 32-bit words at the base a0 names, word j at base + 4j, read and written in place. Its base is word-aligned and it lies in RAM, RAM_ORIGIN ≤ base and base + 4·words ≤ 2^31: the executor refuses any other (Misaligned, OutOfBounds, the sum taken in u64) and the circuit has no witness for one (delegation-circuits.md §1, frame chain). So no frame lies in a public window or in advice.

An invocation reads and writes every word, unchanged ones written back, each a RAM query of the requesting cycle at slot constants::delegation::FRAME_DELTA = 0, ahead of the request's own queries (execution-trace.md §4, §7).

5 The anchor#

Requests and invocations pair one to one through the memory multiset, in the requested type's anchor space s. Otherwise N requests could close against one invocation, N − 1 calls going unexecuted, or an unrequested invocation could rewrite a frame.

5.1 The two sides#

                       reads                               writes
request (deleg)        T(s, a0, 0, 0)                      T(s, a0, 4c + 3, v)
invocation (anchor)    T(s, base, 4c + 3, anchor_value)    T(s, base, 0, 0)

The request is the deleg query of an ADD_SUB_LUI_AUIPC ecall row at cycle c (memory.md §2.1): deleg_mask_rule makes its mask m_pc·Σ_t is_deleg_t and deleg_addr_rule its address the a0 the row read. One query serves every type, so its space is deleg_space, an M column deleg_space_rule pins to Σ_t tag_t·is_deleg_t: a memory leaf may read no W column, and the selectors are W (memory.md §8).

The invocation's two leaves are the anchor read (delegation-circuits.md §1): it writes the answer, stamped 0 with value 0, and reads back the request's write at 4c + 3, c its cycle column. v and anchor_value are free and cancel only when equal; an honest prover writes 0 on both. Each answer starts a path one request long (memory.md §9).

5.2 The three request-side zeroings#

gate, under the request's mask forces
deleg_writes_no_register 0 written to a0, so the result is not the prover's choice
deleg_read_ts_zero the mirror read stamped 0
deleg_read_value_zero the mirror read's value 0

With deleg_addr_rule the last two make the mirror read the answer tuple, so every request consumes an answer of its own; without the timestamp, requests at one base chain, each consuming the previous one's write. The gates are the request row's, the same for every family, so the pairing needs nothing from a family's frame, and a call that changes no memory value has nothing else to expose it. The recursion format's deleg_a0_rule replaces the first (recursion.md §1.4).

5.3 Why the pairing is one to one#

In s the only tuples are the requests' and the invocations': no instruction reaches it, no window initializes it, nothing chains there (trace::AddressSpace::chains).

  1. A live row's 4c + 3 is not 0: the request's pc write and the invocation's frame writes at 4c lie on memory paths, whose timestamps are integers below 2^105 (memory.md §4.2).
  2. So the tuples stamped 0 are the requests' reads and the invocations' answers: as many invocations as requests, with the same multiset of bases.
  3. The rest are the requests' writes and the invocations' reads. No two requests share a cycle (memory.md §9), so each invocation's read is exactly one request's write: every invocation sits at its request's base and cycle, its frame accesses at that point of each word's history.

The trace-level check credits each anchor-space query with its invocation's tuples and sees none of this (execution-trace.md §9).

6 The executor's side#

For a registered number, Machine::ecall and Machine::delegate (crates/emulator/src/lib.rs) read a7 and a0; on the tracing paths refuse a family the VmConfig lacks (§7); read the frame, refusing §4's rules; compute the function natively (emulator::keccak_round, transcript::poseidon2_permute, Fr's operators, schoolbook products with long division, emulator::sha256_call) and write the whole frame back, a recursion family leaving it unchanged and working on field cells; stage the mirror query, reading and writing 0; and write a0 (constants::delegation::a0_after).

EmuError::DelegationFrame refuses a frame the circuit has no witness for, which the arithmetic would answer — long division is right for an unreduced operand too — leaving a proof that fails inside the GKR pass with nothing named: a KECCAK_F round word above 23, a SHA256_COMP group word above 15, an FR_ARITH code other than 1, 2, 3 or operand at or above p in memory form, a MOD_MUL or EC_ADD selector naming nothing or operand its row reads at or above the modulus, a POSEIDON2 lane at or above p. The recursion families' refusals are recursion.md §3–§6's.

The tracer records each invocation in its family's trace::DelegationTrace (execution-trace.md §11), which a shard reads as a trace::FrameSlice, ⌈invocations / height⌉ shards a family. The fill (prover::family_fill) commits the recorded words and derives the circuit's intermediates from those read. It never recomputes a written word: what is committed is what the execution did, and the circuit says that is the function. The circuit's side — frame chain, anchor read, gap decomposition, RAM glue — is delegation-circuits.md §1's.

7 Static detachment#

The instruction sweep cannot see a call, so each shim declares its family with a declaration record (constants::delegation):

MARKER_MAGIC = "APOGDEL1" (8 bytes) ‖ ecall number (u32 LE)          MARKER_BYTES = 12

guest_sdk emits one per family, a static whose #[link_section] is its own allocated section, .rodata.apogee.delegations.<family>, which link.ld's *(.rodata*) absorbs.

  • Its own section, because the linker's garbage collection keeps or drops whole input sections: records sharing one would be kept together, and reaching one shim would declare all.
  • Kept by reachability, not #[used], which keeps every record in every guest. Only the family's shim references its record, reading its own number from it through core::hint::black_box: a linked shim has a record, calls the number it declares, and the optimizer cannot fold the read away.
  • Statically: a call linked but never executed declares its family, which proves zero shards.

program::declared_delegations scans the image's file-backed bytes at every byte offset, a static's address being the linker's; a duplicate is one declaration, and a number no family answers is ProgramError::UnknownDelegation. Identity binds a record through the image column (program.md §8).

A called number whose family the VmConfig lacks is the fatal DelegationFamilyAbsent on the tracing paths; emulator::run, having no VmConfig, executes it. No proof covers it: the statement has no shard of that family, so the mirror read has no answer to consume.

8 Shards, time windows and the block#

A delegation shard's window is proof.md §8's, taken over its invocations' requesting cycles, so it lies inside the span of the ADD_SUB_LUI_AUIPC windows that made the requests. A delegation family is not cycle-owning, so the block holds its windows to nothing beyond start ≤ end ≤ 2^38; the anchor, not the window, places an invocation in time (§5.3).

9 Heights and channels#

A height sets how many calls a shard holds and limits no program. It is a parameter (ProgramParams::heights, defaulting to constants::family::DEFAULT_HEIGHTS, circuits.md §1) in identity's VM_CONFIG: a program's, not an execution's (program.md §7).

  • Floor: constraints::family_circuit returns None below the most variables any of the family's channel tables needs (constraints::lookup::table_vars, lookup.md §3).
  • Trade: a shard costs its height, not its occupancy (streaming.md §1), but its proof grows with the height only by a sumcheck round a variable in each gate list, a height changing no gate, only the number of halving lists. For a family with many calls the fatter shard is the smaller proof.
family channels floor unit of work calls a unit units a shard
KECCAK_F RANGE16, XOR8 2^16 keccak-f[1600] 24 10,922
POSEIDON2 none none width-3 permutation 1 256
FR_ARITH none none Fr add, multiply or inverse 1 256
MOD_MUL RANGE16 2^16 a·b mod m 1 65,536
SHA256_COMP RANGE16, XOR8 2^16 compression 16 16,384
EC_ADD RANGE16 2^16 complete point addition 3 21,845
  • POSEIDON2 and FR_ARITH take 2^8, the menu's smallest shard, where no table fits: every bound is a boolean decomposition. MOD_MUL and EC_ADD take their floor.
  • KECCAK_F and SHA256_COMP take 2^18, two variables above it: four times the calls for 2% more proof (a KECCAK_F shard's is 381,100 bytes, against 373,276 at 2^16). The price is memory: two 2^18 KECCAK_F shards in flight set the measured block's peak (streaming.md §1).
  • No base family carries TIMESTAMP, whose table needs 2^19 rows. FR_OP, P2_FIELD and FIELD_IO carry RANGE16, and FQ_OP TIMESTAMP and RANGE16, flooring it at 2^20.

10 Guest-side callers#

delegation reached from
KECCAK_F guest_sdk::keccak256; in guests/revm-block every alloy-primitives keccak, through its native-keccak hook native_keccak256
SHA256_COMP guest_sdk::sha256; revm-precompile's Crypto::sha256, the 0x02 precompile and the stateless guest's SSZ hashing
POSEIDON2 transcript::poseidon2_permute; guest_sdk::poseidon2_permute
FR_ARITH field::Fr's addition, Montgomery multiplication (*, square, pow, the conversions in from_u64, from_bytes, to_bytes) and nonzero inverse
MOD_MUL k256's FieldElement10x26::{mul, square}, Scalar::mul; ark-ff's MontBackend::{mul_assign, square_in_place} for BN254's two fields, as the product and then ·R⁻¹
EC_ADD guest_sdk::{ec_add, ec_mul}; k256's ProjectivePoint::{add, add_mixed, double}; revm-precompile's Crypto::{bn254_g1_add, bn254_g1_mul}
  • The shims are guest_sdk::recursion's but KECCAK_F's, which only keccak256 reaches (ecall-abi.md §7). Their frame types are #[repr(C, align(4))], so §4's alignment is the type's and not where the code generator put a local.
  • A multi-call operation's order is the caller's, and nothing refuses a wrong one: it computes something else. So each is one SDK function, keccak256's permutation, guest_sdk::recursion::sha256_comp and guest_sdk::recursion::ec_add_complete.
  • The transparent backends: field and transcript call the shims under cfg(target_arch = "riscv32"), through a target dependency on guest-sdk that a host build never resolves, not a cargo feature. Cargo refusing the cycle, guest-sdk cannot name Fr, so the shims take frames of bytes. The software path is each crate's own code, one branch below the call. A guest declares what its library calls reach: Fr arithmetic FR_ARITH, poseidon2_permute both.
  • FR_ARITH's frame carries Fr's memory form (primitives.md §1): canonical values would cost a Montgomery multiplication per value, more than the one the call replaces. POSEIDON2's carries canonical values, six conversions against the permutation's 240 multiplications.
  • The vendored crates, k256 0.13.4, ark-ff 0.6.0 and revm-precompile 43.0.2, are what a guest compiles through guests/Cargo.toml's [patch.crates-io], each route under the same cfg with upstream's code as its software path; the root workspace is unpatched. A MOD_MUL or EC_ADD operand must be below its modulus, so k256 first reduces its lazily reduced field elements. Changed files: guests/vendor/README.md.

11 Limits#

  • The EVM's MULMOD and MODEXP, BLS12-381 and every primitive outside §10's table run as instructions. No signature or pairing is delegated: secp256k1 recovery is k256 code over MOD_MUL and EC_ADD, a BN254 pairing ark-bn254 code over MOD_MUL.
  • A delegation is an operation's core: padding, a sponge or block loop, a scalar multiplication's ladder and a multi-call operation's order are guest code, proven as instructions.
  • A call's result is bound to memory alone: the frame after it is the function of the frame before.
  • This executor implements every family, so no proof here runs a base shim's software path.
  • Retired numbers are ecall-abi.md §4's.

Auditoren/Delegationen

Die Delegationsschaltkreise

Normative Spezifikationdocs/spec/delegation-circuits.mdAls Markdown anzeigen

Zusammenfassung

Die sechs Delegationsschaltkreise im Basisformat. Nach den gemeinsamen Konstruktionen (der Frame-Kette, dem Lesen des Ankers, der Zerlegung der Differenz, der Kanonizitätskette, der Ein-Code-Regel, Byte-Operationen über XOR8 und RAM-Glue für Operationen aus mehreren Aufrufen) spezifiziert die Seite für jeden Schaltkreis Frame, Spalten, Gates und Lookups, warum er seine Funktion und keine andere zulässt, sowie seine Kosten und Aufrufer: KECCAK_F, POSEIDON2, FR_ARITH, MOD_MUL, SHA256_COMP und EC_ADD.

Der folgende normative Text wird auf Englisch gepflegt, der kanonischen Sprache der Spezifikation.

The circuits of the six delegation families the base format registers (recursion.md §1.2): KECCAK_F, POSEIDON2, FR_ARITH, MOD_MUL, SHA256_COMP, EC_ADD. A row is one invocation of a function of a frame of guest memory. For each circuit: its frame, columns, gates and lookups, and why it admits that function and no other. The call, the anchor's pairing, declaration and heights are delegation.md's.

1 Shared constructions#

Each circuit is constraints::delegation's frame over words frame words beside the family's function. None has a setup column; its only tables are its channels' virtual ones (lookup.md §3). live is the one mask, boolean by live_boolean and every lookup's selector. A padding row is all zero and satisfies every gate, a constant term riding live (gkr.md §4).

M[0..4]        cycle  live  base  anchor_value
M[4 + 4j ..]   word j: addr_j  read_ts_j  read_j  write_j        w{j}_addr … w{j}_write_value

Frame chain. Each word is read and written once at a pinned address, as two RAM leaves over memory.md §1's tuple T, from M columns because a leaf reads no W (memory.md §8):

read_w{j}        live·T(RAM, addr_j, read_ts_j, read_j) + 1 − live
write_w{j}       live·T(RAM, addr_j, 4·cycle, write_j) + 1 − live
addr_w{j}        live·(addr_j − base − 4j) = 0
base_aligned     live·(base − RAM_ORIGIN − 4·base_low) = 0          base_low  < 2^29
base_in_window   live·(2^31 − 4·words − base − base_room) = 0       base_room < 2^31

The bounds are delegation.md §4's frame rules, alignment a decomposition because 4 is a unit of Fr. A word the call leaves alone is held by writes_back_w{j}, write_j = read_j; every other written word is bounded below 2^32 by its circuit. A frame lies in RAM proper (delegation.md §4), which starts as the image's words or 0 and which every writer — an execution family (memory-ops.md §5), a frame, FIELD_IO's export (recursion.md §5) — leaves holding words, so a frame word a circuit reads is a word without a bound of its own.

Anchor read. Two leaves in the family's address space s (delegation.md §5) pair the row with its request: it writes the answer T(s, base, 0, 0) and reads T(s, base, 4·cycle + 3, anchor_value), what the request wrote back; anchor_value is free. That makes words + 1 leaves a side, padded with literal 1s to a power of two.

Gap decomposition. Each read precedes the row's write: gap_j = 4·cycle − 1 − read_ts_j is in [0, 2^38). TIMESTAMP would need a 2^20 shard (lookup.md §3), so the frame bounds its gaps, base_low and base_room itself, at the head of W:

  • bit form, at 2^8, where no table fits: 38 booleans a word, gap{j}_{i}, under gap_w{j}, live·(gap_j − Σ_i 2^i·g_i) = 0, and 29 and 31 for base_low and base_room: 38·words + 60 columns, each with its booleanity gate.
  • chunk form, at 2^16 and above, with no gate: a bound x ∈ [0, 2^{16q+r}), 0 < r < 16, is q committed chunks c_k of weight 2^{16(k+1)}, a RANGE16 obligation on each and on the remainder x − Σ_k 2^{16(k+1)}·c_k, and one on 2^{16−r}·c_top, which bounds only beside the chunk's direct one (lookup.md §11). A gap (r = 6) is gap{j}_c0 and gap{j}_c1; base_low and base_room (r = 13, 15) take base_low_hi and base_room_hi: 2·words + 4 columns and 4·words + 6 obligations.

The frame's gates are live_boolean, the addr_w{j}, base_aligned and base_in_window, words + 3, and in the bit form the gap_w{j} and each bit's booleanity besides.

Canonicity chain. A value X in limbs x_0 … x_7 < 2^32 is compared with a modulus m, limbs m_i < 2^32, through boolean borrows β_i and differences d_i ∈ [0, 2^32):

<v>_canonical{i}    x_i − m_i − β_{i−1} + 2^32·β_i − d_i = 0        i = 0 … 7, β_{−1} = 0

Every term is a small integer, so the eight sum over ℤ to X − m + 2^256·β_7 = D, 0 ≤ D < 2^256: β_7 = 1 exactly when X < m. Against Fr's p (§3, §4) the m_i are literals, x_i − p_i rides live and each d_i is 32 booleans; against a selected modulus (§5, §7) the m_i are columns, 0 on a padding row, and each d_i has a 32-bit bound (memory.md §7).

Gated conclusion. The chain's last gate, <v>_below_modulus, is live − β_7 = 0 where every live row reads X; where only rows with enable = 1 read it, it is the gated conclusion enable·(1 − β_7) = 0. β_7 = enable would demand X ≥ m wherever enable = 0, so a row holding a reduced X it does not read would have no witness.

One-code rule. A frame word naming one of k cases is decoded into boolean selectors s_c by word − Σ_c code_c·s_c = 0 and Σ_c s_c − live = 0. The second is not implied: a code 0 has no selector set and a code that is a sum of two has two (1 + 2 = 3), mixing cases. With both, the word and any column pinned to Σ_c lit_c·s_c are one entry of a table of literals, selected and bounded by a degree-1 gate.

Byte operations. Where the unit is the byte (§2, §6), each Boolean operation is one XOR8 obligation (e_0, e_1, e_2), e_2 = e_0 ^ e_1 with all three bytes (lookup.md §3): e_1 and e_2 columns, e_0 any literal-weighted form with a constant (lookup.md §5). The rest is linear in the results: a & b = (a + b − (a ^ b))/2, ¬a & b = (b − a + (a ^ b))/2; against a literal k, v & k = (v + k − (v ^ k))/2 splits a byte at any bit, so a rotation or shift of a word held as bytes is a literal-weighted form over its bytes and their masked copies; and (0, c, c) bounds c to a byte. On true bytes and true XORs each form is exact over ℤ, so its value is the integer it denotes.

RAM glue. An operation too wide for a row is several invocations on one frame, a frame word naming the step (§2, §6, §7). Each proves its step on the frame as it finds it: its reads lie on each word's one history (memory.md §9), so it reads the previous step's writes unless the guest wrote there between. No gate joins two rows, and a shard boundary may fall between them. That every step runs, in order, is the calling code's, which the execution families prove.

2 KECCAK_F#

One invocation is one round of keccak-f[1600]; a permutation is 24 on one frame, the sponge and padding being guest code. The circuit, constraints::keccak, is flat, every gate in gate list 0, and its unit is the byte (§1): no column is a bit but live and the 24 round selectors.

2.1 Frame and columns#

51 words (constants::keccak; M[0..208]), the state in SHA-3 byte order: lane A[x][y], i = 5y + x, at words 1 + 2i (low half) and 2 + 2i. A[i][b] is its byte b; lane coordinates are mod 5.

word read written
0 the round r ∈ [0, 24) yes unchanged
1–50 the state yes the round's output
W name
0..106 the frame's chunks (§1)
106..130 round_sel{r} s_r, one a round
130..134 rc_b{b} rc_t, byte b_t = 0, 1, 3, 7 of the round's constant
134..334 state_in_l{i}_b{b} A
334..494 parity_x{x}_b{b}_s{s} column x's lanes XORed in four steps, the last C[x]
494..574 c_mask_…, theta_d_… C ^ 0x80; D
574..774 theta_a_… A′ = A ^ D
774..950 rho_mask_… A′ ^ mask on the 22 lanes not rotated by whole bytes
950..1150 rho_out_… B, after ρ and π
1150..1550 chi_and_…, chi_out_… B1 ^ B2; χ's output
1550..1554 iota_out_b{b} lane 0's bytes b_t after ι
1554..1556 the multiplicities

2.2 Gates and obligations#

385 gates; O is chi_out, but iota_out at lane 0's bytes b_t; r_xy = ROTATIONS[y][x].

gate count expression
the frame's (§1) 54
round{r}_boolean 24 s_r − s_r²
round_rule 1 read_0 − Σ_r r·s_r
one_round_a_live_row 1 Σ_r s_r − live
rc{t}_rule 4 rc_t − Σ_r s_r·(byte b_t of ROUND_CONSTANTS[r])
writes_back_w0 1 write_0 − read_0
input_w{j}, j = 1 + 2i + h 50 read_j − Σ_{k<4} 2^{8k}·A[i][4h + k]
output_w{j} 50 write_j − Σ_{k<4} 2^{8k}·O[i][4h + k]
rho_pi_l{i}_b{j} 200 B[y][2x + 3y][j] − rot_j(A′[x][y], r_xy), its constant times live

A rotation by 8q + s is linear in a lane's bytes v and their copies μ = v ^ mask (§1), mask = 256 − 2^{8−s} being the top s bits; with u = j − q and w = u − 1 mod 8,

rot_j(v) = 2^{s−1}·(v_u + μ_u) + 2^{s−9}·(v_w − μ_w) + mask·(2^{s−9} − 2^{s−1})      s > 0
rot_j(v) = v_u                                                                    s = 0

v_u's low bits moved up and v_w's top bits down, (v + mask − μ)/2 being v & mask.

The obligations are the frame's 210 on RANGE16 (§1) and 1,020 on XOR8, one a byte:

step count obligation e_2 = e_0 ^ e_1
θ 160 parity_s = parity_{s−1} ^ A[x][s + 1], s < 4, parity_{−1} = A[x][0]
θ 40 c_mask = 0x80 ^ C[x]
θ 40 D[x] = rot(C[x + 1], 1) ^ C[x − 1], c_mask as μ
θ 200 A′[x][y] = D[x] ^ A[x][y]
ρ 176 rho_mask = mask ^ A′
χ 200 chi_and = B1 ^ B2, Bk = B[x + k][y]
χ 200 chi_out = ((B2 − B1 + chi_and)/2) ^ B[x][y]
ι 4 iota_out_t = rc_t ^ chi_out[0][b_t]

2.3 Why it is sound#

Every byte column is an entry of some obligation, so all are bytes, each obligation is the operation it names and each form the integer it denotes (§1): rot because μ is the true XOR, and (B2 − B1 + chi_and)/2 is ¬B1 & B2. The channel alone fixes parity, c_mask, theta_d and chi_and. B is committed, and pinned by rho_pi, because χ reads every lane at an entry only a column may fill.

input_w and output_w are each a word's byte decomposition and its 32-bit bound, so no state word has a range obligation; without output_w a row could write any state. Both are ungated and degree 1, a padding row's words and bytes being 0, which pins its state bytes to 0; a cell that only live-gated gates and obligations reach is free on a padding row, to no effect.

one_round_a_live_row is the one-code rule (§1) over codes 0 … 23: without it a live row could set no selector, claiming round 0, or two spelling a third, and ι would add no constant or a wrong one. The constant is a table of literals the selectors pick (rc{t}_rule), with no lookup or commitment. So a live row writes round read_0 of the state it read.

A permutation is RAM glue (§1) over guest_sdk::keccak256's loop, which stores r = 0 … 23 in word 0 before each call. crates/checker/tests/keccak.rs holds every gate and obligation over 24 such rows to a round written apart in u64 and, through emulator::keccak_round, to tiny-keccak.

2.4 Cost and callers#

1,764 committed columns and, at 2^18, 5,490 inner ones in 29 gate lists, 11 row-wise and 18 halving, all the two memory trees' and the two fraction trees'. The 1,020 obligations and the table's fraction fill 1,021 of the XOR8 tree's 1,024 leaves (lookup.md §6); four more would double it, 4,100 more inner columns. So ι is four obligations: a round constant is zero outside bytes 0, 1, 3 and 7 (constants::keccak::IOTA_BYTES_ARE_THE_ONLY_ONES, checked at compile time).

A 2^18 shard (delegation.md §9) holds 10,922 permutations; its proof is 381,100 bytes (proof.md §9), 34.9 a permutation, and its forward pass 45.2 GB of inner layers (streaming.md §1), which is what sets a block's peak. Caller: guest_sdk::keccak256 (delegation.md §10).

3 POSEIDON2#

One invocation is one transcript::poseidon2_permute (transcript.md §1). The circuit, constraints::poseidon2, is at 2^8 with no lookup, bounding in bits (§1), and is the one delegation circuit that computes above gate list 0.

3.1 Frame and columns#

24 words (constants::poseidon2):

words read written
8l … 8l + 7 lane l, l < 3 yes the permuted lane

A lane is its value's canonical encoding (Fr::to_bytes), not §4's Montgomery form, so the circuit is the permutation itself; the caller's six conversions are small beside the 240 S-box multiplications a call replaces.

M[0..100] and W[0..972] are the frame (§1). W[972..4092] holds 520 booleans for each of six values, the lanes read (in0 … in2) then written (out0 … out2): 256 word bits, then the canonicity chain's (§1) 256 difference bits and 8 borrows.

3.2 Gates#

Gate list 0 holds 4,245: the frame's 51 (§1), a booleanity gate on each W column, and 17 a value, over its read or written words: eight <v>_word{k}, word_k − Σ_t 2^t·bit_{k,t}, and its canonicity chain against p (§1), eight <v>_canonical{i} and <v>_below_modulus, live − β_7.

The permutation is computed, not witnessed: three gate lists a round r, S-boxing every lane of a full round and lane 0 of a partial one, whose other lanes the first two lists copy:

list 3r         q_i = (x_i + c_{r,i})²       t_i = x_i + c_{r,i}
list 3r + 1     q2_i = q_i²                  t_i copied
list 3r + 2     x′ = M_r·v                   v_i = q2_i·t_i, or x_i on a copied lane

M_r is E or I and the constants are literals of the gates; round 0's x is E applied to in_l = Σ_k 2^{32k}·read_{8l+k}. A committed column is read by gate list 0 only (gkr.md §2), so live and out_l = Σ_k 2^{32k}·write_{8l+k} are carried up to gate list 192, which holds the last three gates,

out_lane{l}     live·(x_l − out_l) = 0          x the state after round 63

gated because a padding row computes the permutation of the zero state, which is not zero.

3.3 Why it is sound#

A layer's column is forced by the gate that writes it, so x is the permutation of (in_0, in_1, in_2) as field elements. The word gates make each in_l and out_l the integer its words spell, and the chains put it below p: a lane at or above p has no witness, and out_lane fixes all 24 written words, where without the chains on out a row could write x_l + p. The forward pass accepts Plonky3's permutation vectors (crates/checker/tests/poseidon2.rs).

3.4 Cost and callers#

4,192 committed columns and 2,020 inner ones in 201 gate lists, 193 row-wise and 8 halving: 736 the rounds' (15 a full round, 11 a partial one), 768 the four carried columns', the rest the memory trees'. A 2^8 shard holds 256 permutations; its proof is 664,780 bytes, 2,597 a permutation. Caller: transcript::poseidon2_permute on the guest target (delegation.md §10).

4 FR_ARITH#

One invocation is one Fr addition, multiplication or inversion. The circuit, constraints::fr_arith, is flat, at 2^8 with no lookup, bounding in bits (§1).

4.1 Frame and encoding#

25 words (constants::fr_arith):

words read written
0 the code: 1 add, 2 multiply, 3 inverse (OPS) yes unchanged
1–8, 9–16 a, b yes unchanged
17–24 out yes, unconstrained the result

A value is Fr's in-memory form, Fr::to_memory_bytes: the canonical encoding of the Montgomery representative x·R, R = 2^256 mod p. The circuit computes what Fr's own operators compute on representatives,

add         out = a + b
multiply    out = a·b·R⁻¹
inverse     out = R²·a⁻¹, and 0 at a = 0

because a frame of values would cost the guest a Montgomery conversion per value, more than the multiplication a call replaces. Fr::inverse answers None at 0 itself and makes no call.

4.2 Columns and gates#

M[0..104] and W[0..1010] are the frame (§1); W[1010..2570] 520 booleans for each of a, b (read) and out (written), as §3.1; W[2570..2573] the selectors f_add, f_mul, f_inv (selector1 … selector3); W[2573..2576] the field columns prod, inv and z (is_zero). The 2,701 gates: the frame's 53 (§1); 2,573 booleanity gates, on every bit and selector; §3.2's 17 per value; writes_back_w{j} for j < 17; and, a, b and out being the forms Σ_k 2^{32k}·word_k,

gate expression
opcode_rule read_0 − f_add − 2·f_mul − 3·f_inv
one_op_a_live_row f_add + f_mul + f_inv − live
prod_rule prod − a·b
inv_is_an_inverse a·inv + z − f_inv
is_zero_at_nonzero a·z
inverse_of_zero_is_zero z·inv
out_rule out − f_add·(a + b) − R⁻¹·f_mul·prod − R²·f_inv·inv

R⁻¹ and R² are literals derived from constants::FR_R.

4.3 Why it is sound#

As in §3.3, each value is the integer below p its words spell, so out_rule fixes the eight written words. prod is committed, under an ungated gate, because a selector times a·b is degree 3. On an inverse row a ≠ 0 forces z = 0 and inv = a⁻¹, and a = 0 forces z = 1 and inv = 0; without is_zero_at_nonzero, z = 1 and inv = 0 pass at any a, and without inverse_of_zero_is_zero, inv is free at a = 0. one_op_a_live_row is the one-code rule (§1): 1 + 2 = 3, so opcode_rule alone lets f_add and f_mul answer an inversion with a + b + a·b·R⁻¹.

4.4 Cost and callers#

2,680 committed columns and 142 inner ones, all the memory trees', in 14 gate lists, 6 row-wise and 8 halving. A 2^8 shard holds 256 operations; its proof is 266,292 bytes, 1,040 an operation. Caller: field's addition, Montgomery multiplication and inverse on the guest target (delegation.md §10).

5 MOD_MUL#

One invocation is one multiplication out = a·b mod m of 256-bit integers, m one of four fixed primes a frame word selects. The circuit is constraints::mod_mul.

5.1 The frame and the columns#

25 words (constants::mod_mul). A value is a plain residue, not a Montgomery one, in eight 32-bit limbs, least significant first.

words
0 the selector: 1 secp256k1's base field p, 2 its order n, 3 BN254's base field q, 4 its scalar field r (CODES, MODULI) read, written back
1–8, 9–16 a, b, each below the selected modulus read, written back
17–24 out written; the value read is ignored

Codes start at 1, so a zero word names no field. The EVM's MULMOD, whose modulus is arbitrary, is not this call and runs as guest code.

M[0..104], W[0..54]   the frame (§1)
W[54..58]     selector1 … selector4        s_c, one a code
W[58..66]     m_limb{k}                    m_k, the selected modulus
W[66..162]    <v>{k}_hi, <v>_diff{i}, <v>_diff{i}_hi, <v>_borrow{i}     for v = a, b, out
W[162..178]   q_limb{k}, q_limb{k}_hi      the quotient and its halfwords
W[178..220]   carry{k}, carry{k}_c0, carry{k}_c1      c_k + 2^36 for k < 14, and two chunks
W[220]        range16_multiplicity

5.2 Gates and lookups#

read_j and write_j are word j's two values (§1), a_i and b_i read limbs, out_i written ones, and c_k = carry{k} − 2^36·live. Each expression is held to 0:

gate count expression
the frame's (§1) 28
writes_back_w{j}, j < 17 17 write_j − read_j
selector{c}_boolean; selector_rule; one_modulus_a_live_row 6 s_c − s_c²; read_0 − Σ_c c·s_c; Σ_c s_c − live
m_limb{k}_rule 8 m_k − Σ_c s_c·MODULI[c][k]
<v>_borrow{i}_boolean, <v>_canonical{i}, <v>_below_modulus 51 v's canonicity chain (§1) against the m_k columns, concluding live − β_7
limb{k}, k < 15 15 Σ_{i+j=k} (a_i·b_j − q_i·m_j) − out_k + c_{k−1} − 2^32·c_k; out_k past limb 7, c_{−1} and c_14 are 0

274 RANGE16 obligations, all under live: the frame's 106 (§1); a pair — the 32-bit bound of memory.md §7, two obligations over a committed high halfword — on every limb of a, b, out and q and on every diff_i (112); and each carry{k} in [0, 2^37), by two chunks and four obligations as a gap (§1) (56).

5.3 Why it is sound#

The field. By the one-code rule (§1), m is the modulus word 0 names. Codes add (1 + 3 = 4), so without one_modulus_a_live_row selectors 1 and 3 answer a request for r modulo p + q; with it each m_k is one literal, which is all that keeps m's limbs, bound by no obligation, below 2^32.

The product. Every limb of a, b, out, q and m being below 2^32, a position's products sum below 2^67 a side and the carries lie in [−2^36, 2^36), so no term nears Fr's modulus: the fifteen limb{k} equations hold over ℤ and, weighted by 2^{32k}, sum to a·b = q·m + out, position 14 having no carry out.

The reduction is out_below_modulus: without it (q − 1, out + m) satisfies every other relation wherever out + m fits eight limbs.

The operand bounds make the relation total, not out right: with a, b < m, q = (a·b − out)/m < m, so every frame the circuit admits has an eight-limb quotient. A caller holding a lazily reduced value therefore owes a reduction below m, not below 2^256. The emulator's mod_mul_frame refuses the frames no proof could cover, a selector that is no code and an operand at or above m (EmuError::DelegationFrame).

5.4 Cost and callers#

Shape: circuits.md §1. A 2^16 shard (delegation.md §9) is 65,536 multiplications at 2.1 proof bytes each; its forward pass, 2,180 row-wise inner columns × 2^16 rows × 32 bytes, is 4.6 GB.

guest_sdk::recursion::mod_mul makes the call over a ModMulFrame. The vendored k256 reaches it from its field and scalar multiplies (codes 1, 2), the vendored ark-ff from BN254's Montgomery multiply (codes 3, 4): delegation.md §10.

6 SHA256_COMP#

One invocation is four rounds of SHA-256's compression function and four words of its message schedule; a compression is sixteen invocations on one frame, joined by RAM glue (§1). Padding, the block loop and the final addition of the chaining value are the caller's. The circuit is constraints::sha256.

6.1 The frame#

25 words (constants::sha256):

words read written
0 the round group r < 16 unchanged
1–8 the working variables a … h a … h four rounds on
9–24 the schedule window W_{4r} … W_{4r+15} moved down four words, W_{4r+16} … W_{4r+19} last

Call 0 reads the chaining value as a … h and the block, decoded big-endian, as the window. Over a row the state is two sequences: A_0 … A_{−3} are a … d as read, A_4 … A_1 are a … d as written, and E_j is the same over e … h, so each of the sixteen is a frame column. For k < 4 and m < 4, every sum mod 2^32:

T1          = E_{k−3} + Σ1(E_k) + Ch(E_k, E_{k−1}, E_{k−2}) + K_{4r+k} + W_{4r+k}
A_{k+1}     = T1 + Σ0(A_k) + Maj(A_k, A_{k−1}, A_{k−2})
E_{k+1}     = A_{k−3} + T1
W_{4r+16+m} = σ1(W_{4r+14+m}) + W_{4r+9+m} + σ0(W_{4r+1+m}) + W_{4r+m}

Call r + 4's rounds read the words call r derives, so the guest computes no schedule; calls 12–15 derive words no round reads.

6.2 Bytes and their obligations#

No column is a bit but live and the group selectors g_r. A word that enters a Boolean operation has four byte columns, and each such operation is one XOR8 obligation (x, y, x ^ y) a byte (lookup.md §3), of which position 0 alone may be a literal-weighted form (lookup.md §5).

  • A rotation is linear. With μ = v ^ (2^s − 1) committed, a byte v splits into lo = (v + 2^s − 1 − μ)/2 and hi = (v − lo)/2^s. Byte j of ROTR_{8t+s}(V) is hi(v_{j+t}) + 2^{8−s}·lo(v_{j+t+1}), indices mod 4, and for s < 8 the word ROTR_s(V) is (V − lo(v_0))/2^s + 2^{32−s}·lo(v_0).
  • The big sigmas nest, Σ0(a) = ROTR2(a ^ ROTR11(a ^ ROTR9(a))) and Σ1(e) = ROTR6(e ^ ROTR5(e ^ ROTR14(e))), so each XOR has one rotated operand and the outer rotation is a word's: 17 obligations a sigma.
  • The small sigmas end in a shift, σ0(x) = ROTR7(x ^ ROTR11(x)) ^ SHR3(x) and σ1(x) = ROTR17(x ^ ROTR2(x)) ^ SHR10(x), so their outer XOR has two derived operands: the shifted bytes are committed and pinned by gates. 16 and 15 obligations, SHR10's top byte being 0.
  • Ch and Maj are linear in XORs, Ch(e, f, g) = (f + g − (e ^ f) + (e ^ g))/2 and Maj(a, b, c) = (a + b + c − (a ^ b ^ c))/2: 8 obligations each.
  • A carry c is a byte by (0, c, c).

That is 52 obligations a round and 32 a schedule word, 336 on XOR8. RANGE16 carries 114: the frame's 106 (§1) and a pair (§5.2) on each written word without bytes, A_4, E_4, W_{4r+18} and W_{4r+19}.

M[0..104], W[0..54]   the frame (§1)
W[54..70]     group{r}                     g_r, one a group
W[70..118]    a{j}_b{b}, e{j}_b{b}         bytes of A_{−2} … A_3 and E_{−2} … E_3 (j = m2 … 3)
W[118..150]   w{i}_b{b}, n{m}_b{b}         bytes of window words 1–4, 14, 15, derived words 0, 1
W[150..358]   r{k}_…                       52 a round: the big sigmas' masks and XORs (34),
                                           e^f, e^g, a^b, c^a^b (16), two carries
W[358..514]   s{m}_…                       39 a schedule word: the small sigmas' masks, XORs
                                           and shifted bytes (38), a carry
W[514..518]   w{j}_written_hi              high halfwords of A_4, E_4, W_{4r+18}, W_{4r+19}
W[518..520]   range16_multiplicity, xor8_multiplicity

6.3 Gates#

All of degree 1 but the frame's and the booleans:

gate count expression
the frame's (§1) 28
group{r}_boolean; group_rule; one_group_a_live_row 18 g_r − g_r²; read_0 − Σ_r r·g_r; Σ_r g_r − live
writes_back_w0 1 write_0 − read_0
a{j}_decode, a{j}_encode, e{j}_…, w{i}_decode, n{m}_encode 20 a word − Σ_b 2^{8b}·byte_b, for every word with bytes
w{i}_shift, i < 12 12 write_{9+i} − read_{13+i}
r{k}_a, r{k}_e 8 §6.1's A_{k+1} and E_{k+1}, as word + 2^32·carry − sum
s{m}_sum 4 §6.1's W_{4r+16+m}, likewise
s{m}_shr3_b{b}, s{m}_shr10_b{b} 28 a committed shifted byte − its form

K_{4r+k} is the form Σ_r K_{4r+k}·g_r.

6.4 Why it is sound#

A sum's operands are words: those with bytes by their obligations, and d, h, W_{4r} and W_{4r+9} … W_{4r+12}, which only sums read, because the frame lies in [RAM_ORIGIN, 2^31) (§1), below advice, where every initial value and every write is a word (memory-ops.md §5; §1 for these circuits). Its carry being a byte, a sum gate holds over ℤ, and its left word, bounded by its bytes or its pair, is the sum mod 2^32. Without the carry's range any word satisfies the gate; without the pair on A_4, a carry of 0 writes the unreduced sum. Every word a row writes is therefore a word: a copy, one with bytes, or one of the four with a pair.

Group 0's code being 0, group_rule alone admits a live row with no selector or with g_0 beside another; one_group_a_live_row refuses those and two selectors spelling a third group, each a round under a wrong constant. Sixteen rows are one compression by RAM glue (§1) and by guest_sdk::recursion::sha256_comp, which stores r = 0 … 15 in word 0 before each call; the emulator's sha256_frame refuses a group word of 16 or more. crates/checker/tests/sha256.rs evaluates every gate and obligation over sixteen chained rows built from FIPS 180-4 in u32 arithmetic and holds their output to the standard's abc digest.

6.5 Cost and callers#

Shape: circuits.md §1. A 2^18 shard (delegation.md §9) holds 16,384 compressions at 11.6 proof bytes each; its forward pass, 2,694 row-wise inner columns × 2^18 × 32 bytes, is 22.6 GB.

guest_sdk::sha256 pads, walks the blocks, and for each runs sha256_comp's sixteen calls and adds the result to the chaining value. The vendored revm-precompile routes Crypto::sha256 to it: precompile 0x02, and the stateless guest's SSZ hashing (delegation.md §10).

7 EC_ADD#

One invocation is a third of one complete point addition P1 + P2 on secp256k1 or BN254 G1, in homogeneous projective coordinates (x = X/Z, y = Y/Z). An addition is three invocations on one frame in group order, joined by RAM glue (§1); scalar multiplication is guest code over it. The circuit is constraints::ec_add.

7.1 The formula#

Renes–Costello–Batina 2015, Algorithm 7, for y² = x³ + b, with b3 = 3b: 21 and 9 (constants::ec_add::CURVE_B3).

group 0   xx = X1·X2            yy = Y1·Y2            zz = Z1·Z2
group 1   m4 = (X1+Y1)(X2+Y2)   m5 = (Y1+Z1)(Y2+Z2)   m6 = (X1+Z1)(X2+Z2)
group 2   X3 = xy·ym − byz3·xz  Y3 = yp·ym + bxx9·xz  Z3 = yz·yp + xx3·xy

xy = m4 − xx − yy   yz = m5 − yy − zz   xz = m6 − xx − zz   ym = yy − b3·zz
yp = yy + b3·zz     byz3 = b3·yz        xx3 = 3·xx          bxx9 = 3·b3·xx

Both groups have prime order, so the formula is complete: a doubling, P + (−P), the identity (0 : 1 : 0) and any Z take no special case, in the guest or in a row, and nothing is inverted. The formula is the caller's: the vendored k256's ProjectivePoint addition is this algorithm on these coordinates, so the delegated and the software path return the same representative.

The twelve multiplications are nine reductions, each of X3, Y3, Z3 being two products under one quotient. A row holds three, not nine, because a shard's memory grows with its row's width and its height cannot fall below 2^16 (§7.5).

7.2 The frame and the columns#

97 words (constants::ec_add), a value as in §5.1:

words read by group written by group
0 the selector, one of CODES: 1–3 secp256k1's groups 0–2, 4–6 BN254 G1's all none
1–24 X1, Y1, Z1 0, 1 2, as X3, Y3, Z3
25–48 X2, Y2, Z2 0, 1 none
49–72 xx, yy, zz 2 0
73–96 m4, m5, m6 2 1

A row has three slots, each one reduction of one shape:

A·B + C·D + 1024·m² = q·m + out,     out < m

Group 0's (A, B) are (X1, X2), (Y1, Y2), (Z1, Z2) and group 1's the three pairs of sums, both with C = D = 0. Group 2's (A, B, C, D) are (xy, ym, byz3, −xz), (yp, ym, bxx9, xz) and (yz, yp, xx3, xy): a product's sign rides its operand.

M[0..392], W[0..198]   the frame (§1)
W[198..295]   word{j}_hi                   the high halfword of every word's read value
W[295..301]   selector{c}                  s_c, one a code
W[301..310]   m_limb{k}, b3                the curve's modulus and 3b
W[310..334]   bzz3_{k}, byz3_{k}, bxx9_{k}     b3·zz_k, b3·(m5_k − yy_k − zz_k), 3·b3·xx_k
W[334..622]   <v>_diff{i}, <v>_diff{i}_hi, <v>_borrow{i}     chains of the twelve values x1 … m6
W[622..1027]  slot{r}_…, out{r}_…, q{r}_…, carry{r}_…    135 a slot: four operands (32), out and
              its halfwords (16), a nine-limb q and its halfwords (18), 15 carries c_k + 2^46
              with two chunks each (45), out's chain (24)
W[1027]       range16_multiplicity

7.3 Gates and lookups#

G_g is the sum of the two selectors naming group g, and c_k = carry − 2^46·live:

gate count expression
the frame's (§1) 100
selector{c}_boolean, selector_rule, one_code_a_live_row 8 §5.2's, over six codes
m_limb{k}_rule, b3_rule 9 the column − Σ_c s_c·(its literal for code c's curve)
bzz3_{k}_rule, byz3_{k}_rule, bxx9_{k}_rule 24 the column − its product above
<v>_borrow{i}_boolean, <v>_canonical{i} 240 canonicity chains (§1) of the twelve values and the three outs, against m_k
<v>_below_modulus 15 e·(1 − β_7): e is G_0 + G_1 for x1 … z2, G_2 for xx … m6, live for an out
operand{r}_{o}_{k}_rule 96 an operand limb − Σ_g G_g·(group g's expression at that limb)
slot{r}_limb{k}, k < 16 48 Σ_{i+j=k} (A_i·B_j + C_i·D_j + 1024·m_i·m_j − q_i·m_j) − out_k + c_{k−1} − 2^32·c_k, q having nine limbs
writes_back_w{j} 97 write_j − read_j − G_g·(out_k − read_j), for the group g and slot limb k that write word j, if any

1,110 RANGE16 obligations under live: the frame's 394 (§1); a pair (§5.2) on every word's read value (194), every chain difference (240), every out limb (48) and every q limb (54); and each carry in [0, 2^47), four apiece (180).

7.4 Why it is sound#

The curve and the group are §5.3's argument over six codes: codes add (1 + 3 = 4, 2 + 4 = 6), and the one-code rule is also all that keeps m's limbs and b3 literals.

The operands. An operand limb is its group's expression: a combination, with coefficients of at most 3, of frame limbs below 2^32 and of their products with b3. Its pin is therefore its bound, below 2^38 in magnitude, and it carries no obligation. It is a committed column because the expression depends on the group, and a selector times a product of limbs would be degree 3; b3 enters through the three helper columns for the same reason.

The identity. As in §5.3: a position stays below 2^78, the carries in [−2^46, 2^46) (CARRY_OFFSET_BITS), the sixteen equations hold over ℤ and close because position 15 has no carry out, and out < m makes out the residue of A·B + C·D. The 1024·m² (OFFSET_MULTIPLE) keeps the left side non-negative, a quotient's limbs being unsigned: it is lowest in group 2's Y3, at −673·m² by its operands' ceilings 22m, 22m, 63m and 3m. One literal serves every slot, a group-dependent offset being degree 3, and q < 1697·m fits nine limbs. ec_add::artifact checks both constants against the ceilings when it builds the circuit.

Canonicity. out < m is the reduction. A read value below m is what the ceilings assume, and so what gives every admitted frame a quotient; the emulator's ec_add_frame refuses a frame whose group reads a value at or above m, or whose selector is no code. Each such conclusion is gated (§1) on the groups that read the value: every lane is below m on every row a guest builds (EcAddFrame::of zeroes the intermediates), so β_7 = e would leave no row a witness.

What the guest owns. Each third is proved; their order is the guest's. guest_sdk::recursion::ec_add_complete writes the three codes in turn, and groups out of order are not refused but compute another point from stale lanes. Nor is a point held to its curve: what is proved is the formula's arithmetic.

7.5 Cost and callers#

Every bound is a RANGE16 obligation, so the family sits at 2^16, the channel's floor (lookup.md §3), and no other height is practical: as bits the 97 gaps alone would be 3,686 columns, and at 2^18 the forward pass below would be 73 GB. Shape: circuits.md §1. A shard holds 21,845 additions at 19.9 proof bytes each; its forward pass, 8,708 row-wise inner columns × 2^16 × 32 bytes, is 18.3 GB.

guest_sdk::ec_add makes the three calls over an EcAddFrame, and guest_sdk::ec_mul is double-and-add over it. The vendored k256 routes ProjectivePoint's addition, mixed addition and doubling here, and the vendored revm-precompile routes Crypto::bn254_g1_add and Crypto::bn254_g1_mul, precompiles 0x06 and 0x07: delegation.md §10.

Auditoren/Pipeline

Der Streaming-Prover

Normative Spezifikationdocs/spec/streaming.mdAls Markdown anzeigen

Zusammenfassung

Wie aus einer Ausführung ein Blockbeweis wird, ohne dass ihr Trace jemals vollständig gehalten wird. Die Seite nennt die gemessenen Kosten des Provers, spezifiziert die beiden Durchläufe und warum das Commitment dem Beweisen vorausgehen muss, was eine Ausführung überdauert, den Shard-Plan und wie Shards geschnitten werden, während der Executor läuft, die Worker-Pipeline und ihre Garantien, einschließlich der Garantie, dass der Beweis nicht vom Scheduling abhängt, und den archivierten Pfad, auf den sich die Manipulations-Suite stützt.

Der folgende normative Text wird auf Englisch gepflegt, der kanonischen Sprache der Spezifikation.

How one execution becomes a BlockProof: the guest runs twice, and a fixed number of workers commit, then prove, its shards as the executor fills them. The block is proof.md's, and its bytes do not depend on the schedule; this page fixes when each column exists, and so what a proof costs.

1 The prover, and what it costs#

prover::prove_block_streaming(setup, io, max_in_flight) proves every block: host::prove wraps it over the ProverSetup that host::setup builds from an ELF, bench prove drives it (tools.md §1), and recursion nodes are proved through it. Beside the block it returns a StreamingReport: each pass's wall clock and the executor's time inside it, the cycle and shard counts, and the most shards held at once.

Its memory follows the shards in flight, not the shard or cycle count: a partial buffer per family (§4), the last-access tables (§3), at most max_in_flight shards being worked and one filled shard's rows per family waiting (§5), and the output, 64 bytes a commitment and the ShardProofs. The executor's whole output, emulator::trace_run's buffers and event log at about 300 bytes a cycle, never exists; executing twice (§2) costs time instead.

A shard costs its height times its circuit's width (circuits.md §1), however few of its rows are live. gkr::forward holds every inner layer as field elements, 32·Σ_{k≥1} w_k·2^{n_k} bytes over layer k's width and variable count (42 GiB for a 2^18 KECCAK_F shard, 8.4 GiB for a 2^20 SHIFT_BITWISE one), and gkr::prove adds a copy of the layer it reduces and an eq table. The opening, after the forward pass is dropped, copies every committed column.

Measured on the base proof of recursion.md §10:

workload block 257,510 of glamsterdam-devnet-8, revm-block-stateless: 60 transactions, 101.5 Mgas, 198M cycles, 207 shards
machine 32 vCPUs, 247.7 GiB, --in-flight 12
pass 1 191 s; 25.7 vCPUs busy on average; one-thread fills 81% of its shard-seconds; sampled RSS at most 15.9 GiB
pass 2 2,290 s; on average 11.95 of 12 shards held and 30.4 vCPUs busy until the guest exits; then the exit's 460 s tail, whose longest stretches are the two KECCAK_F shards' one-thread fills, 200 s and 279 s
peak RSS 173.92 GiB: the two 2^18 KECCAK_F shards, together in the tail with nothing else in flight

Twelve shards in flight never reached that peak: a delegation family's height set it, and max_in_flight bounds only how many shards coincide.

2 The two passes#

pass 1                                       pass 2
  execute; for each shard as it fills:         execute again; for each shard as it fills:
    fill its M columns, commit them,             fill every column, prove the shard,
    keep the points, drop the columns            keep the proof, drop the rest
  at exit: the window families' shards         at exit: the window families' shards
  the statement, then G1–G11                   the statement's roots, the BlockProof

Each pass drives a fresh emulator::StreamingRun, which hands over a family's buffer as a ShardChunk the moment it reaches the family's height (§4); the workers of §5 take the chunks.

Pass 1 commits each shard's M columns, its family's fill with them moved out and no multiplicities, one commitment a column (in the recursion format one a stack of 2^σ, recursion.md §1.3). At exit it derives the window list, shard counts and boundary from the final state (§3), asserts the cut equals trace::plan_shards, commits the window families' shards, puts the commitments in statement order by (family, index) (proof.md §1) and runs G1–G11 (proof.md §2). A commitment reads no transcript, so only its place in the absorbed order matters, not when it was computed.

Pass 2 re-executes. The emulator is a pure function of (image, io), with no clock, randomness or threads, so it cuts the same shards; pass 2 asserts that its CycleProfile, Execution, window list and boundary are pass 1's. Each shard gets every committed column, multiplicities included, and prover::prove_shard_columns: shard transcript, GKR proof, opening (proof.md §4, §5). Proofs go to their statement positions, and prover::public_inputs copies each shard's two memory roots into the statement.

M is not recommitted: a shard's opening takes its M commitments from the statement, pass 1's, and its polynomials from pass 2's columns, so columns that differed would give an opening the verifier refuses.

3 What survives an execution#

A streaming run records no memory event. At exit StreamingRun::finish hands over each non-empty partial buffer as its family's last shard, and a StreamedExecution: the last-access tables (trace::MemoryState), the CycleProfile and the Execution (execution-trace.md §11). Beyond the shards' rows, the guest's inputs and its journal, everything the statement needs is a function of that final state: the boundary (trace::build_boundary_finals), ZERO_WINDOWS' list (trace::init_windows), the window families' teardowns (memory.md §3) and the field-window count. So a window family's shard exists only once the execution is over (§5).

A fill reads one shard through prover::ShardSource, its ShardRows a trace::RowSlice (a cycle-owning family), a trace::FrameSlice (a delegation family) or, for a window family alone, the final MemoryState. The streaming path builds it over a fresh chunk, ShardSource::archived over a slice of a TraceArchive (§6); nothing else differs. Memory columns come from a shard's rows alone (execution-trace.md §11), and checker::memory_columns_from_log rebuilds them from the event log, independently (circuits.md §3).

4 The shard plan#

A family's rows, in the order they are appended, are cut into shards of its VmConfig height h: shard i is rows [i·h, min((i + 1)·h, len)), the last padded to h with zero rows (memory.md §2). trace::plan_shards is ⌈rows/h⌉ per family over the CycleProfile, cycles for a cycle-owning family and invocations for a delegation family, so a family the execution never reached has no shard. A window family plans 0; its shards are windows (memory.md §3), counted by shard_counts in crates/prover/src/lib.rs: one INIT_TEARDOWN shard and one of each public window whatever the execution did, a ZERO_WINDOWS shard per entry of init_windows, one per advice window supplied (trace::advice_window_count), and field windows through the highest cell touched (MemoryState::field_windows). The counts are the statement's shard_counts (proof.md §1).

The flush. StreamingRun makes the cut as it runs. After a cycle is recorded, a buffer that has reached h rows is handed over as ShardChunk { family, index, rows }, index = rows/h − 1, and replaced by an empty one. A cycle appends at most one row to any buffer, the owning family's and, for a delegation request, one invocation to the delegation family's, so a buffer reaches h without passing it, a step fills at most two, and no chunk is split. At exit finish hands over the partial buffers. Chunks arrive in fill order, not statement order, and pass 1 asserts that each family's count is the plan's.

5 The pipeline#

pipeline (crates/prover/src/streaming.rs) runs both passes: max_in_flight workers under std::thread::scope and one std::sync::Mutex around a Source, which holds the executor, the filled shards no worker has claimed, and the counts. Under the lock a worker gives back its shard and claims the next: a waiting one, or else it steps the executor itself until a buffer fills (Source::claim, the only place the guest runs). Outside the lock it builds the shard's columns, works it and drops it. These are the prover's only threads and only lock; within a shard, parallelism is rayon over data.

held bound by
claimed shards, and all built from them max_in_flight one a worker; asserted in Source::claim
filled, unclaimed shards rows only, one per family the executor steps only for a claim with nothing waiting, a step fills at most two buffers and the exit one per family; asserted in Source::admit

The executor never runs ahead of demand, and there is no batch: a slow shard holds one worker. The workers are not rayon threads. A shard's MSMs, forward pass, sumcheck and opening run on rayon's global pool, so RAYON_NUM_THREADS sets the cores the shards share, and a worker blocked in that work cannot take a second shard as a rayon thread waiting in a nested join would. Fills run on the workers' own threads, one each, so up to RAYON_NUM_THREADS + max_in_flight threads are runnable. Fork-join cannot express this: below one shard per core, a batch waits for its slowest shard.

  • The block is independent of the schedule. A shard's proof is a function of the global state and its own columns, its transcript a fresh sponge seeded with the digest (proof.md §4); no proof depends on the thread count (gkr.md §5); proofs are placed by statement position. crates/prover/tests/streaming.rs compares the bytes at 1 and 8 in flight.
  • The failure returned is the earliest in fill order, at any worker count: claims follow fill order, a claimed shard is worked to its end, a failure stops later claims (Source::fail), and an executor failure ranks after every shard it filled.
  • No deadlock: the lock is never held while a shard is worked or taken twice by one worker, and nothing waits under it but the executor's step.
  • A panic stops the claims, through a drop guard (StopOnPanic) or, inside the executor, the poisoned lock; the shards in flight finish, and the panic is re-raised as itself.

The window families' shards follow the pipeline, built from the final state in rayon batches of at most max_in_flight, which are all that a ThreadPool::install around the call bounds.

The knob. max_in_flight, at least 1, is an argument because only the caller knows the machine; bench prove --in-flight defaults to 8. StreamingReport::peak_in_flight is the most shards claimed or batched at once in either pass.

6 The retained archived path#

emulator::trace_run keeps a whole execution, every buffer and the MemoryEventLog, and trace::TraceArchive::from_execution holds it (execution-trace.md §11). The per-shard component reads one through ShardSource::archived, with the same fills and shard proving: prover::statement_inputs (counts, windows, boundary, every shard's M columns), global_commit_phase, shard_columns, shard_memory_columns, prove_shard, prove_shard_columns and public_inputs. checker::TamperHarness is built on it (circuits.md §3): it writes changed cells into shards' columns, recommits changed M columns in a fresh global commit phase and re-proves, which needs an execution held still and read twice. Streaming has no such seam: pass 2 rebuilds, by re-executing, the columns pass 1 committed, so a cell changed in either pass would contradict the other.

prover::prove_block(setup, archive, plan), advance(setup, archive, until) and finish(archive) prove a block from an archive; nothing outside crates/prover/src/phases.rs calls them. prove_block refuses a plan that is not plan_shards of the archive's profile. advance fills the archive's four later phase sections in order, timing each, and decodes any it already holds, so an imported archive resumes; a stopped streaming run starts again. No column is stored: a phase rebuilds them from the archive. The sections, in proof.md §9's encodings, each refusing a byte too many or too few:

section content
PostCommit the statement's PublicInputs bytes, without roots; the global transcript after G11 as its 226-byte postcard snapshot (transcript.md §3); the four memory challenges; the digest
PostGkr per shard, in statement order: family u32, index u32, ts_start and ts_end u64, the witness commitments, the outputs, the GKR proof, the base claims' point, the shard transcript's snapshot after the GKR proof
PostOpening each shard's ShardProof bytes
Final the complete PublicInputs bytes, then the proofs

Auditoren/Pipeline

Rekursion

Normative Spezifikationdocs/spec/recursion.mdAls Markdown anzeigen

Zusammenfassung

Wie aus einem Basisbeweis ein einziger Groth16-Beweis wird, den ein Contract prüft, ohne dass der Basisbeweis verändert wird. Die Seite spezifiziert das Rekursionsformat und seine gestapelten Commitments, den Speicher für Körperelemente und die vier Koprozessor-Familien, die Tapes, die ein Knoten abspielt, die Blatt- und Knotenprogramme mit dem über den Baum verketteten Transkript, das Journal des Knotens, wie aufgeschobene Öffnungen zu einem einzigen Akkumulator gefaltet werden, den Scheduler, den Groth16-Decider mit gebundenen Wires und seine zweiphasige Zeremonie, den Contract und die gemessenen Kosten.

Der folgende normative Text wird auf Englisch gepflegt, der kanonischen Sprache der Spezifikation.

How one base block proof becomes one Groth16 proof a contract checks. The section numbers are the ones the code cites. Where this page and the code disagree, the code is right.

base proof ──► leaves ──────► internal nodes ──► root ───► decider ───► contract
N shards,      each a run     each 2–4           covers    the root in  folds the root's
base format    of base shards children           0..N      Groth16      points; two pairings
  • A node is this VM proving a verifier program. It verifies shards and folds every Mercury check they defer into one accumulator (A, B), the claim e(A, [1]_2) = e(B, [x]_2). Nothing pairs before the contract.
  • Base proving is untouched. No base key, statement or proof moved a byte: a leaf verifies base shards as they are.
  • Nodes are proved in a recursion format (§1) over a field memory (§2) with four coprocessor families on it (§3–§6), and they replay tapes (§7) rather than run verifier-core on RV32, which measured 3.0B cycles for block 257,510's 207 shards: fifteen times the block itself.

1 Two formats, one code path#

1.1 The rule#

A statement is in the recursion format exactly when its VmConfig holds FIELD_WINDOWS (VmConfig::is_recursion), which is exactly when its program declares a field family. No wire form says which format applies.

1.2 The delegation registry#

constants::delegation::TYPES is one append-only table, and its first BASE_TYPES = 6 rows are all the base format knows. constraints::family_circuit is the base registry; constraints::recursion_circuit differs from it in two ways only: its ADD_SUB knows every row and carries §1.4's rule, and the five families of §2–§6 exist. VmConfig::circuit picks the registry, for a key's load rule and the prover alike.

1.3 Stacked commitments#

Every commitment a shard opens is a point its parent folds (§8.3). So a recursion shard commits each of its two phases — its M columns, and its W columns with the multiplicities — as stacks of 2^σ columns. At height 2^n, with k_M and k_W columns (VmConfig::stack_vars):

σ = min(24 − n, the smallest even σ with 2^σ ≥ max(k_M, k_W))
  • Column i is slot i mod 2^σ of stack ⌊i / 2^σ⌋. A stack is the (n + σ)-variate multilinear whose evaluations [j·2^n, (j + 1)·2^n) are slot j's column, and its commitment is that polynomial's Mercury commitment. 24 is the ceremony's size.
  • The GKR pass leaves each column's value v at u. Then σ challenges r are drawn (STACK_CHALLENGE), a stack's value is Σ_j eq(r, j)·v_j, and a setup column is a stack of one, eq(r, 0)·v, its commitment unchanged. The shard's one batch opening is at u ‖ r, over the M stacks, the W stacks, then the setup columns.

σ = 0 is the base format exactly.

1.4 A recursion request leaves a0 past its frame#

A base delegation request writes 0 into a0. A request of a type past BASE_TYPES writes a0 + 4·words (constants::delegation::a0_after), which the recursion ADD_SUB's deleg_a0_rule holds it to. So frames laid back to back replay as back-to-back ecalls, one RISC-V row a call.

2 The field memory#

2.1 The space#

address_space::FIELD = 10: cells addressed by a u32, each a whole Fr. Its tuples (FIELD, cell, ts, value) join RAM's in the one memory multiset. No instruction reaches it. Only §3–§6's rows do, each access at its row's requesting cycle c and its own slot, 4c + Δ, with a read's usual gap check; a read-only access writes back what it read. A field access is not a MemoryEventLog event, a value not being a u32: trace::MemoryState keeps each cell's last (ts, value), and a recursion execution has no TraceArchive form. It streams.

2.2 FIELD_WINDOWS#

Family 18, 2^20 rows: ZERO_WINDOWS' circuit at a stride of one cell a row. Window w is cells [h·w, h·(w + 1)), initialized to 0. The windows are consecutive from cell 0 — shard i is window i — so a statement lists none, and a cell outside them has no tuple to balance a read against.

The four families on it are invoked, by the delegation ABI: an ecall whose a0 is a frame of words in RAM. A frame's words name cells.

§ family id ecall anchor space height frame a row is
3 FR_OP 19 0x0509 11 2^20 [op, d, a, b] one field operation
4 P2_FIELD 20 0x050A 12 2^18 [n, s, x, y, d] one transcript duplex step
5 FIELD_IO 21 0x050B 13 2^18 [op, cell, ptr] eight RAM words to a cell, or back
6 FQ_OP 22 0x050C 14 2^20 [op, d, a, b] one BN254 base-field operation

3 FR_OP — one field operation a row#

op op
1 MUL d ← a·b 6 EQ a = b, or the row has no witness
2 ADD d ← a + b 7 IMM d ← word b, as an integer
3 SUB d ← a − b 8 SHL d ← a·2^32 + word b
4 MAC d ← d + a·b 9 DIGIT d ← a's low byte, b ← (a − d)/2^8
5 INV d ← a⁻¹, and 0 at a = 0

a, b and d are accessed at slots of their own, so any two may name one cell. EQ is how a tape asserts. IMM and SHL are how it builds a constant with no field arithmetic of the guest's. A scalar's 32 DIGITs ending at 0 represent it mod p, which is all a scalar multiplication needs.

4 P2_FIELD — one duplex step a row#

With the state at cells s..s+3, the row absorbs n ∈ {0, 1, 2} of the cells x, y — the lanes are (n ≥ 1 ? x : s₀, n = 2 ? y : n = 1 ? 0 : s₁, s₂ + n) — and writes poseidon2_permute of them to d..d+3. A state is never overwritten, so a challenge is a cell of the triple that made it. The circuit is flat, every S-box's u² and u⁴ committed, so a parent verifies it as one gate list.

5 FIELD_IO — between RAM and a cell#

Over the eight RAM words w_k at ptr:

  • IMPORT (1): the cell takes Σ_k w_k·2^{32k} mod p. A non-canonical encoding is harmless.
  • EXPORT (2): the words take limbs below 2^32 congruent to the cell. Congruence, not canonicity: a guest that needs the canonical value compares the words with p itself.

Addressability is the multiset's. A word no window initializes cannot balance.

6 FQ_OP — one base-field operation a row#

An element of BN254's Fq is four consecutive cells of 64-bit limbs, congruent to its value mod q and not necessarily below it. Only this family writes one. The op word is a code, three flags and a digit cell (word >> 6): a flagged operand's element is its word plus 8·digit, a bucket chosen by a digit, which is what lets an MSM be a static tape (§8.3).

op
1 MUL d ← a·b
2 ADD d ← a + b
3 SUB d ← a − b
4 MULEQ asserts a·b ≡ d
5 FROM128 d ← a₀ + 2^128·a₁ from two cells below 2^128: a coordinate from its transcript limbs

One integer identity serves all five, a·y + z = q·K + d′, checked over 128-bit groups of limbs with a range-checked quotient and carries. Some of those ranges go through TIMESTAMP, which is why the family is at 2^20. b's and d's four cells share one read timestamp, so an element is only ever written whole; tape::run refuses a tape that reads one written apart, before a fill would.

7 Tapes#

A shard's checks have a fixed shape for its family and height. So the host compiles them, once, into a tape (verifier_core::tape): a straight-line list of coprocessor calls over absolute cells, in which nothing branches on a value. A tape reads three kinds of cell:

  • constants, which it builds itself from IMM and SHL, so they are bound with it;
  • slots, which its caller fills: the statement's digest and memory challenges, and the shard's index, window, roots and commitments;
  • inputs, the proof, IMPORTed from a blob laid out in the tape's Input order (tape::shard_blob), which the tape's own checks are what bind.

tape::shard_tape is verify_shard_local's steps 7–11 and Mercury's field side (pcs_verify), call for call: the shard transcript, the GKR backward pass, the lookup and root checks, the stack challenges and values, and the opening's twelve scalars. Every check is an EQ. It leaves three things to its caller (§8): the shard's time window against its neighbours', step 10c on the two public shards, and everything on the curve — to a tape a point is four transcript limbs, and the batch's cm* is a hint.

tape::schedule reorders a tape into runs of one family's calls, tape::encode is the form a guest replays — the cells its imports fill, then runs of frames — and tape::run is the native reading, over a Memory that models each access's timestamp as well as each cell's value.

8 Nodes and the tree#

8.1 Two programs, one procedure#

guests/recursion is two binaries. The leaf verifies shards from..to of one statement of the base program. The node verifies two to four whole statements of the two recursion programs — its children's proofs — reads each child's journal out of the output window step 10c binds, holds the children to one another, and folds their accumulators beside their shards'.

Both run verifier_core::node::node through a Driver. The host runs it natively (host::recursion), so it refuses whatever a guest would, first and by name, and it writes the advice the guest reads. The guest runs it by coprocessor calls. A binary's image — every shard tape, the fold's templates, the constants — is built by build.rs with verifier-core itself and sits in .rodata, so a program's identity binds every tape it replays.

  • The base program's identity is a constant of the leaf's image. The SRS digest and the generic table are constants of both images.
  • A node takes the two recursion programs' identities as claims and journals them, for the top to check once.
  • A program's setup commitments are advice, held to its identity by recomputing it.

The global transcript is a chain across the tree (verifier_core::chain). The node with shard 0 runs the prefix, G1–G7. Every node absorbs its own shards' memory commitments, G8, from the state its predecessor left. The node with the last shard runs the suffix, G9–G11, which settles the digest and the memory challenges every node took as claims. A node that holds a whole statement makes its memory argument, Π reads · R_b = Π writes · W_b.

A node holds its children to: exit status 0; one base statement — its shape, digest, challenges, io_digest, exit status and shard count; adjacent shards; chain states that meet; time windows in order across the seam; and, of a node child, the two identities it requires itself.

8.2 The journal#

47 cells, each a 32-byte word (node::journal):

cells
0 a digest of the base statement's shape: its shard counts and windows
1–7 its global digest, four memory challenges, io_digest, exit status
8–10 its shard count, and the shards this node covers, from..to
11–18 the chain's state at from and at to: three lanes and a pending input each
19–22 the covered shards' read and write root products; the boundary factors where to is the count
23–28 the first and last covered shard's family and time window
29–44 A and B, each x then y in four 64-bit limbs
45–46 the leaf program's and the node program's identities this node requires; 0 for a leaf

The root covers 0..count: every base shard verified, the transcript run end to end, the memory argument made. What is left is one pairing check and two identities.

8.3 Folding#

After each shard's tape the node's own transcript absorbs the shard transcript's final state (FOLD_STATE) and draws w and w′ (FOLD_WEIGHT), so a shard's weights follow everything they weight. Then, as scalars of points:

  • entry i of the shard's Mercury check gets w·e_i, on its side (pcs_verify::ENTRY_POINTS);
  • the batch check cm* = Σ ρ^i·cm_i is folded beside it: cm* gets w′ more, and each cm_i gets −w′·ρ^i;
  • [1]_1 and the setup commitments, which every shard of a family shares, accumulate one scalar each and enter once;
  • a child's A and B enter under a weight drawn after its whole journal (FOLD_CHILD).

Each side is one MSM on FQ_OP (verifier_core::fold): Pippenger with 8-bit digits over GLV halves, 16 windows of 256 buckets, every step a static template. A point is held to the curve and its scalar's split to the scalar, then added to one bucket a window through an indirect operand. Inversions are host witnesses held by a MULEQ, and buckets start at offsets so that no addition degenerates. A point costs about 400 FQ_OP calls.

8.4 The scheduler#

bench recurse <dir>/<stem> --out <out>, over a base proof archive (tools/bench/src/recurse.rs):

  • Keys. It writes base.key and programs.key into <out> and builds the two binaries with APOGEE_RECURSION_KEYS=<out>, where their build.rs reads them. The node is built twice: once with no image, for the two programs' keys, and once over them.
  • Plan, fixed before anything is proved (host::recursion::Tree::plan, <out>/tree.txt): leaves of at most --leaf (64) consecutive base shards, then levels of internal nodes over two to --fan-in (4) children, a lone leftover carried up. A program has sixteen families and each costs at least a shard, so a node is sixteen shards before any work and leaves are cut large.
  • A node is a process, bench recurse-node: it verifies its inputs natively, builds its advice only then, proves, verifies, holds the proved journal to the native one and writes <out>/<id>.block. At most --in-flight nodes run, with --in-flight × --shards-in-flight shards in flight across them: a node takes its share of what is spare when it starts, so a root alone has the whole budget. A proof already in <out> is kept, so a stopped run resumes; a run whose plan or programs differ is refused.
  • At the root it checks what a verifier owes beside the root's own proof — the journal covers 0..count and is the archive's statement, it requires the two programs' identities, and (A, B) discharges — and then runs §9.

9 The decider#

The root is still a GKR proof and some hundreds of points, and a contract can check neither. host::decider splits its verification in two.

The circuit is §8.1's node procedure over one child, the root, through a Driver that writes rank-1 constraints: an FR_OP is one constraint in the common case and none where it only copies, a duplex is 255, advice is a free wire. It verifies the root as a node would, and holds its journal to from = 0 and to = count. But it folds nothing: every MSM template is skipped, and each point's four limbs and its scalar are bound wires instead, after the two identities, the base statement's exit status, and its public input and output, a wire a byte, whose digest the circuit holds to the journal's io_digest.

crates/groth16 is Groth16 over this repository's BN254. A circuit streams its constraints into a sink, so no matrix is held. Three things are not the textbook's:

  • Bound wires are values the verifier holds, too many to be public inputs. The proof carries their commitment D = Σ w_j·[(β·A_j + α·B_j + C_j)/η]_1 under a fifth trapdoor η. A challenge c is SHA-256 of D and the verifier's values, and the circuit ends with acc ← (acc + wire)·c over the bound wires. The public inputs are c and that result, both of which the verifier computes from its own values, and the check is e(A, B) = e(α, β)·e(IC, γ)·e(C, δ)·e(D, η), IC being the public wires' points under 1, c and the result. D is fixed before c, so wires that differ from the values agree with them at c with probability len/r.
  • No blinding. A proof hides nothing and is a function of its witness.
  • A Lagrange basis. A and B are sums over the constraints, Σ_j (A·w)_j·[L_j(τ)], not over the wires. So the one element a key holds a wire is [(β·A_i + α·B_i + C_i)/x]_1, x being γ, η or δ — and a powers-of-tau ceremony already publishes [L_j(τ)].

The key is a ceremony's, in two phases:

  • Phase 1 is ppot_0080_24.ptau, the ceremony the tree's own commitments are under (srs::Phase1): the Lagrange basis at the circuit's domain in both groups, and the powers a quotient takes. Everything of the key that depends on τ is a combination of those points, and nothing derives τ.

  • Phase 2 is the circuit's own (groth16::phase2, bench ceremony), and makes α, β, γ, δ and η from 1 by contributions: each multiplies a trapdoor by a factor only its contributor knew, so a trapdoor is unknown while one contributor to it was honest.

    step
    init every trapdoor 1: a wire's [A_i(τ)]_1, [B_i(τ)]_1, [C_i(τ)]_1, and [τ^k·Z(τ)]_1. Deterministic from the circuit and the file
    round 1, contribute to α and β: [β·A_i]_1 and [α·B_i]_1, kept apart
    seal a wire's three terms summed
    round 2, contribute to γ, δ and η: the sum over the wire's trapdoor, and [τ^k·Z(τ)/δ]_1
    key the last state verified and, if every trapdoor has a contribution, written as the key

    The order of the rounds is the soundness. A prover may hold a wire's three terms only summed, over δ or η: apart, it could give A, B and C three witnesses. A contribution to α or β scales the terms apart, so those are finished before anything is divided.

    A state carries each contribution's record — its factor in G1 with a Schnorr proof of knowing it, bound to the records before it, and the trapdoor in G2 afterwards. Verifying a state checks that chain, then its elements against init's under those trapdoors, one pairing equation over a random combination: against the circuit and the file alone, with no earlier state. Every step lists the records by their factors' points, so a contributor finds its own under the state the key is made of. bench decide reads the key key wrote, and nothing else writes one.

  • setup_dev, bench decide --dev-key, derives all six trapdoors from a public seed. It is for development and tests: anyone forges under it.

The contract (contracts/ApogeeVerifier.sol) is verify(input, output, exitStatus, proof, points), a point being x, y, scalar, side [1]_2's points and then side [x]_2's. It rebuilds the bound values — a point's limbs are its coordinates' halves, or four sentinels at infinity, which is what the root's transcript absorbed — recomputes c and the result, checks the Groth16 pairing, folds each side with ecMul and ecAdd, which is also what holds a point to the curve, and checks e(A, [1]_2) = e(B, [x]_2). Its Groth16 key, the ceremony's two G2 points and the two identities are set at deployment.

bench decide <out> proves under the ceremony's key, checks the proof natively, deploys and calls the contract in revm, and writes decision.constructor and decision.calldata — under --dev-key, development.*.

What a deployment still owes. A key is as trustworthy as its ceremony: one honest contributor a round, which a ceremony run on one machine is not. The circuit depends on the root's shape — its program, its shard counts, the public values' lengths — so a key, and its ceremony, is per shape. And the contract pays about 9k gas a point, because the circuit folds none.

10 Running it#

bench prove --stateless <fixture> --out <dir>             the base proof
bench recurse <dir>/<stem> --out <out> --in-flight 4      the tree
bench ceremony <out> init                                 the decider's key: once a root shape,
bench ceremony <out> contribute                           each contributor in turn, to alpha and beta
bench ceremony <out> seal
bench ceremony <out> contribute                           and to gamma, delta and eta
bench ceremony <out> key
bench decide <out>                                        the Groth16 proof, and the contract

It needs assets/ptau/ppot_0080_24.ptau. Measured on block 257,510 — the tree on a 32-CPU, 247 GiB machine, the ceremony and the decider on an 18-core laptop:

base proof 207 shards, 14.5 MB, 2,481 s
tree 4 leaves of at most 64 base shards and a root: 116 shards
leaves, four at once 21, 24, 23 and 27 shards; 2,157 s; 92 GiB peak
root, four shards in flight 21 shards, 460 s, 1.03 MB
decider's circuit 7,896,686 constraints, a domain of 2^23
ceremony init 65 s; a contribution 50–56 s; key 70 s, 12.7 GB; the key 2.65 GB
decider the key read in 1 s, the proof 18.5 s, 6.1 GB
contract 358 points; 3,620,026 gas; 34,980 bytes of calldata

Auditoren/Pipeline

Ethereum-Blöcke

Normative Spezifikationdocs/spec/ethereum.mdAls Markdown anzeigen

Zusammenfassung

Die Referenz-Arbeitslast: Ethereum-Blöcke, ausgeführt auf revm innerhalb der VM. Die Seite spezifiziert das Crate des Gastprogramms und seine zwei Binaries, das Witness-Format und das Ausgabe-Commitment des Mini-Blocks, Eingabe und Ausgabe des zustandslosen Validators, die unterstützten Forks, was ein Ergebnis beweist, wie ein Block auf revm läuft, die Signatur-Recovery, die Konformität mit dem zkEVM-Test-Release, wo jede Validierungsregel geprüft wird und wie ein Block aufgezeichnet wird.

Der folgende normative Text wird auf Englisch gepflegt, der kanonischen Sprache der Spezifikation.

guests/revm-block runs Ethereum blocks on revm inside the VM. This page specifies its two binaries: the mini-block binary's input, BlockWitness, and its output commitment; the stateless validator's input, result and rules; how each runs a block on revm; and how a block is recorded.

1 The guest crate#

One library, revm_block (src/lib.rs and its modules), and two binaries, each its own program identity, the same for every block because the block is advice (public-values.md §6):

binary advice journal exit status
revm-block (src/main.rs) a BlockWitness (§2) the output commitment (§3) 0; 61 not a canonical witness; 62 not executable
revm-block-stateless (src/stateless_main.rs) statelessInputBytes (§4) the 43-byte result always 0

The mini-block binary runs transactions, usually a block's first few, over a pre-state recorded from a node (§6); it proves an execution, not a block's validity (§3). Full blocks are proved with the stateless validator: the mini journal grows by a record a transaction and outgrows the public window (public-values.md §9). On the host the library is the oracle crates/emulator/tests/revm.rs holds the mini binary's journal to, over unpatched upstream crates.

Dependencies. revm is the workload being proven: it and what revm-precompile brings (arkworks, k256, p256, sha2, ripemd) are reachable from no prover, verifier or other guest. It is built without default features, so without blst, c-kzg or libsecp256k1. The pin is exact, =43.0.1, because an identity is a digest of the image (program.md §8), and the set is the reference stateless guest's (paradigmxyz/stateless's lock): twelve crates, held in both lockfiles by crates/host/tests/revm_lock.rs, its revm-handler 43.0.1 carrying the EIP-8037 system-call state-gas reservoir tests-zkevm@v21.0.1 expects.

Delegations. Both binaries declare KECCAK_F, SHA256_COMP, MOD_MUL and EC_ADD: keccak through alloy-primitives' native-keccak hook (revm_block::native_keccak256), SHA-256, secp256k1 and BN254 through vendored crates (delegation.md §10, guests/vendor/README.md). revm-precompile is patched in its Crypto trait's default bodies, not given a second implementation by install_crypto: two types behind crypto()'s OnceLock<Box<dyn Crypto>> stop LLVM devirtualizing its calls, keeping code it otherwise strips, 870,828 bytes of .text on the mini binary, past its tables' reach.

Code size. No ELF is committed: host::fixture::build_revm_guest builds either binary at --release, proved at host::fixture::revm_params — 2^20 for every family whose height is a choice (revm_block::TRACE_HEIGHT_RELEASE), each delegation family's default, bytecode_size_words = 2^21. A 2^20 table reaches 1.9375 MiB of .text (program.md §5); the stateless binary's is about 1.96 MB, 96.6% of it. The debug image needs 2^22 and is only ever run.

2 BlockWitness#

The mini binary's advice: postcard of revm_block::BlockWitness, a format of this repository's, written by host::recorder (§6). Fields in declaration order; a word is 32 big-endian bytes; a u8, an Option tag (0 or 1) and a fixed array are raw bytes; every other integer and every length is a varint.

BlockWitness          env BlockEnvWitness; accounts Vec<AccountWitness>, by address;
                      txs Vec<TxWitness>, in execution order
BlockEnvWitness       chain_id u64; spec_id u8 (revm's SpecId); number word; beneficiary [20];
                      timestamp word; gas_limit u64; basefee u64; difficulty word;
                      prevrandao Option<word>; excess_blob_gas Option<u64>;
                      blob_gasprice Option<u128>; slot_num u64;
                      block_hashes Vec<(u64, word)>, by number
AccountWitness        address [20]; nonce u64; balance word; code Vec<u8>;
                      slots Vec<(word, word)>, by key, zero values included
TxWitness             caller [20]; to Option<[20]>, None a creation; value word; data Vec<u8>;
                      gas_limit u64; gas_price u128, the max fee from type 2;
                      gas_priority_fee Option<u128>; nonce u64; chain_id Option<u64>;
                      access_list Vec<([20], Vec<word>)>; blob_hashes Vec<word>;
                      max_fee_per_blob_gas Option<u128>; authorizations Vec<AuthorizationWitness>
AuthorizationWitness  chain_id word; address [20]; nonce u64; authority Option<[20]>, recovered

2.1 One state, one encoding#

BlockWitness::decode refuses, with exit 61:

rule WitnessError
spec_id is a SpecId UnknownSpec
excess_blob_gas and blob_gasprice both present or both absent BlobPairing
accounts, each account's slots, block_hashes strictly ascending AccountsNotSorted, SlotsNotSorted, BlockHashesNotSorted
the bytes are exactly BlockWitness::encode's for the value Malformed

The last closes postcard's two second encodings: postcard::from_bytes ignores trailing bytes, and its varints accept non-minimal forms (81 00 reads as 1). A code hash is computed, not carried, and a transaction's type is derived from its fields (TxEnv::derive_tx_type).

2.2 Execution#

revm_block::WitnessDb answers revm from the witness and refuses every miss (DbError, exit 62): the witness is unbound advice, so a default would be a value the prover chose.

  • Absence is recorded: WitnessDb::basic answers None for an account recorded with nonce 0, balance 0 and no code. A zero slot is recorded like any other.
  • BLOCKHASH reads env.block_hashes. revm answers 0 without asking for any block but the 256 before the current one, and serves those from the database, not EIP-2935's contract: at most 256 entries.
  • Code is Bytecode::new_raw_checked's: bytes beginning 0xef01 that are not a 23-byte EIP-7702 delegation, which a few pre-EIP-3541 accounts hold, are DbError::MalformedCode, where Bytecode::new_raw would panic, an exit 101 that names nothing (ecall-abi.md §7).
  • The block gas limit is a running bound. revm checks each transaction against the block's limit and keeps no total; revm_block::run, the block executor, refuses transaction i unless gas_limit_i ≤ env.gas_limit − Σ_{j<i} gas_used_j, the Yellow Paper's intrinsic validity.
  • The blob gas price is recorded (§6). It derives from the excess through the fork's update fraction, 3,338,477 at Cancun, 5,007,716 at Prague, raised by each BPO fork (§4.2), and revm 43 knows only the first two. revm holds each type-3 transaction's max_fee_per_blob_gas to it.

Not in the witness: signatures, caller and each authority being the producer's recovery, unchecked; a parent header, and the header rules against it; a state root (§3); a slot number, which the recorder writes as 0, no JSON-RPC method serving EIP-7843's.

3 The output commitment#

The mini binary's journal, revm_block::run's return:

per transaction, in order   status u8 (0 halt, 1 revert, 2 success) ‖ gas_used u64 LE
                            ‖ output_len u32 LE ‖ output: the return data, empty on a halt
logs commitment       32    keccak256 of  count u32 LE ‖ per log, in emission order:
                            address 20 ‖ topic_count u8 ‖ topics, 32 each ‖ data_len u32 LE ‖ data
post-state summary    32    keccak256 of  count u32 LE ‖ per account, by address:
                            address 20 ‖ nonce u64 LE ‖ balance 32 BE ‖ code_hash 32
                            ‖ slot_count u32 LE ‖ per slot, by key: key 32 BE ‖ value 32 BE

The record count is the witness's. The summary covers the state revm's finalize returns: every account the block loaded, read-only and nonexistent ones included, with every slot it loaded.

What a proof states: some canonical BlockWitness makes revm_block::run return this journal. Nothing ties the witness to a chain; a reader holding one recomputes the journal natively. And the journal tells witnesses apart only as far as the execution reads them: a slot read and then overwritten unconditionally reaches nothing, while every loaded account's final balance and nonce are in the summary.

4 The stateless validator#

revm-block-stateless maps tests-zkevm@v21.0.1's statelessInputBytes to its statelessOutputBytes, byte for byte. The formats and rules are ethereum/execution-specs' verify_stateless_new_payload at the release's commit (host::zkevm::RELEASE_COMMIT); revm_block::stateless::run is the guest's whole computation, and §5 lists its rules.

4.1 Input and output#

input    schema_id u16 BE ‖ SSZ(StatelessInput)                          ssz::decode
           new_payload_request   the schema's fork's NewPayloadRequest
           witness               state: trie-node preimages; codes; headers: RLP, oldest
                                 first, the parent last, at most 256
           chain_id              u64
           public_keys           eth-act/ere-guests v0.17.1's layout only: 65 bytes a transaction
output   new_payload_request_root 32 ‖ successful_validation 1 ‖ chain_id u64 LE ‖ schema_id u16 LE

The layouts' fixed parts are 16 and 20 bytes, so no input is both; the second is the zkEVM benchmark's. The root is hash_tree_root under EIP-7916's and EIP-7495's progressive forms as of 2026-01-15 (ssz::request_root), whatever the layout. Decoding is as strict as the spec's: every offset against the bytes it bounds, every bounded list against its limit, nothing after the end.

The guest exits 0 on every input. One that does not decode, or names a schema §4.2 does not list, publishes the sentinel, 43 zero bytes (ssz::SENTINEL); any other publishes its request's root, its verdict, its chain id and its schema id. The empty input is the one a run cannot be given, a run without advice having no advice region.

4.2 Forks#

The schema id, fork_index << 8 | 0x01, names the fork; no activation schedule is compiled in (block::fork).

schema fork request revm SpecId blob target, max update fraction
0x1201 Osaka Electra/Fulu OSAKA 6, 9 5,007,716
0x1301 BPO1 Electra/Fulu OSAKA 10, 15 8,346,193
0x1401 BPO2 Electra/Fulu OSAKA 14, 21 11,684,671
0x1501 Amsterdam Gloas: a block access list, a slot number, EIP-8282's two request types AMSTERDAM 14, 21 11,684,671

4.3 What a result proves#

true says the request whose root is published is a valid block on chain chain_id under the fork schema_id names. The witness needs no binding: the root fixes the payload, and the witness is held to it by hashes — the parent header to the payload's parent_hash, each ancestor to its child's, the state trie to the parent's state_root and each node to its parent's reference, each code to its account's code hash. A node a read needs and the witness lacks is an error, never an absence (mpt::get). So a wrong witness cannot make an invalid payload valid; but false says only that this input did not validate, which a prover can arrange for any payload.

4.4 How a block runs on revm#

The pre-state is witness::WitnessDb behind revm's State: the state trie under the parent's root, each storage trie parsed on its first read, codes by hash, and BLOCKHASH numbering each ancestor by its position below the block. stateless::execute is the spec's apply_body: the EIP-4788 and EIP-2935 system calls; each transaction; the withdrawals; the requests, from deposit logs and the checked system calls of EIP-7002, EIP-7251 and, from Amsterdam, EIP-8282. The calls before the transactions are block access list index 0, each transaction has its own, and what follows them shares the last. A transaction must fit what is left, Amsterdam metering regular and state gas apart (EIP-8037):

before Amsterdam   tx.gas_limit ≤ gas_limit − Σ gas_used
Amsterdam          tx.gas_limit ≤ 2^32 − 1,  min(tx.gas_limit, 2^24) ≤ gas_limit − Σ regular,
                   tx.gas_limit ≤ gas_limit − Σ state;  the block uses max(Σ regular, Σ state)
both               2^17·blobs ≤ 2^17·max − Σ blob gas

Three rules make the result the spec's where following reth would not:

  • Code loads when revm asks (witness::WitnessDb::code_by_hash), never with its account: the witness carries only the code the spec's execution read, and a coinbase may be a contract nothing calls.
  • Every write precedes every deletion in the post-state replay (stateless::post_state_root), in each trie, as the spec's mpt_set_storage_slots orders them. A deletion that leaves a branch one child needs that child's node, on no changed key's path; the witness carries those the spec's order needs, and writing first needs a subset.
  • One commit per index (stateless::commit_index). revm 43's access-list builder records a value that differs from its commit's baseline, and revm re-bases a value at each call, so committing call by call records a slot one call toggles and the next restores. An index's calls are committed once, each baseline reset to the committed state.

Also the spec's: a checked system contract must have code, deposit events are parsed to the byte, withdrawals precede requests; the TxEnv is built field by field (build_fill would put a dummy authorization in an empty type-4 list); the blob price is a checked fake_exponential (block::blob_gas_price); 0xef01 code that is not a delegation runs as legacy. Declared lengths are added checked and trie parsing is depth-bounded, a panic publishing nothing.

4.5 Signatures#

Every sender and EIP-7702 authority is recovered in the guest (tx::recover_key) under EIP-2's rules, 0 < r < n, 0 < s ≤ n/2, a parity bit, as Q = r⁻¹(s·R − z·G) with k256's arithmetic, which the vendored k256 routes to MOD_MUL and EC_ADD. The verification upstream's recover_from_prehash ends with cannot fail once recovery succeeds and costs about as much again, so it is not done. An authorization that does not recover is skipped, as EIP-7702 says. A key in ere-guests' layout is checked, never used: one a transaction, 0x04 ‖ x ‖ y, naming the recovered sender.

4.6 Conformance#

All 67,251 pairs of tests-zkevm@v21.0.1 match natively (crates/host/tests/conformance.rs, by hand); CI holds the library to a committed subset of 34 — a case for each rule the release reaches, the smallest valid one, every undecodable one — in both layouts, and the binary runs the subset by hand. The release fills only Amsterdam: tools/stateless-ref holds the Electra/Fulu layout to eth-act/ere-guests v0.17.1 and crates/host/tests/canonical.rs the encodings and header rules to two mainnet blocks, but no Osaka-family input has an end-to-end oracle.

5 Where each rule is checked#

The mini binary's rules, then the validator's step by step. A validator refusal is a stateless::Invalid variant, which host::zkevm::verdict names and the guest publishes as false. Paths are revm_block's.

rule refusal code
mini: a canonical witness exit 61 BlockWitness::decode
mini: every read recorded, code well formed exit 62 WitnessDb
mini: each transaction fits the gas left and executes exit 62 run_against
the input decodes under a listed schema sentinel ssz::decode, block::fork
the ancestors decode and chain Ancestors stateless::ancestors
no empty transaction EmptyTransaction stateless::verify
the base fee fits a u64 Unrepresentable stateless::payload_header
the header the payload implies hashes to block_hash BlockHash stateless::{verify, payload_header}
each transaction decodes (EIP-2718, types 0–4) Transaction(i) tx::decode
keyed layout: a key a transaction PublicKeys stateless::verify
the versioned hashes are the request's VersionedHashes stateless::verify
EIP-7934's block size BlockSize block::block_rlp_len
the header against its parent, twelve rules Header(_) block::validate_header
the blob gas price fits a u128 Unrepresentable block::blob_gas_price
the parent's state root is in the witness Witness(_) witness::WitnessDb::new
chain id; signature; keyed layout: the key names the sender ChainId(i), Signature(i), PublicKeys stateless::execute, tx::sender
the transaction fits what is left Capacity(i) stateless::execute
revm executes it, every read in the witness Execution(i) stateless::execute
the system calls SystemCall stateless::{execute, commit_index}
the deposit events Deposits block::deposit_requests
gas used, receipts root, bloom, blob gas used, requests hash GasUsed, ReceiptsRoot, Bloom, BlobGasUsed, RequestsHash stateless::verify, block
Amsterdam: the access list's item count and hash AccessList stateless::verify, alloy_eip7928
the post-state root StateRoot, Witness(_) stateless::post_state_root, mpt

6 Recording a block#

host::recorder::record(rpc, block_number, range) makes a BlockWitness for a block's first n transactions or all of them (recorder::TxRange) by running them once against a node: recorder::WitnessRecorder is a revm::Database over the parent block's state that records each answer, and the transactions run through revm_block::run_against, the guest's own executor, so the record is what the guest will read. The result is put through BlockWitness::decode.

  • Reads. An account is eth_getProof with no keys, absent when nonce, balance, code hash and storage hash are all empty or both hashes are zero, Geth's answer; code eth_getCode, checked against the hash; a slot eth_getStorageAt; a header eth_getBlockByNumber; the blob gas price eth_feeHistory's baseFeePerBlobGas, a receipt's blobGasPrice existing only for type 3.
  • Choices. The hardfork is mainnet's by number (recorder::mainnet_spec): before the Merge is refused, after Osaka runs as Osaka. caller is the node's from; authorities are recovered on the host.
  • The client, host::rpc::Rpc, files each response under the SHA-256 of its canonical request in the fixture's rpc-cache/, so a second recording is byte-identical and offline; a miss without ETH_RPC_URL is an error. A request goes through curl, the endpoint and its key on the command line, retried on a transport failure, a 5xx or a 429.
  • On disk (host::fixture): <stem>.json, a Pin naming the block and the length and SHA-256 of <stem>-witness.bin and <stem>-journal.bin, native revm's journal, beside rpc-cache/. crates/host/tests/vectors/mini-block* is block 26,057,509's first two transactions, refreshed by kat-gen -- block (tools.md §7).

Nothing here produces a stateless input. eth_getProof returns the nodes on a key's path, and a deletion that collapses a branch needs its surviving sibling's node, which is on no changed key's path (mpt::MptError::BlindedCollapse is the validator's refusal without it), so the proofs of a block's keys are not a witness. Stateless inputs come from an external producer, a tests-zkevm release or the zkEVM benchmark's datasets; host::zkevm reads every JSON object carrying both statelessInputBytes and statelessOutputBytes, and bench prove --stateless proves one as it is (tools.md §1).

Referenz

Glossar

Normative Spezifikationdocs/glossary.mdAls Markdown anzeigen

Zusammenfassung

Das für Apogee VM spezifische Vokabular, eine Zeile pro Begriff, jeder mit dem Abschnitt der Spezifikation verlinkt, der ihn definiert. Begriffe, die die Literatur festlegt, wie GKR, LogUp, KZG und RISC-V, sind nicht aufgeführt.

Der folgende normative Text wird auf Englisch gepflegt, der kanonischen Sprache der Spezifikation.

The project's own vocabulary, one line a term, each linked to the section that defines it. Terms the literature fixes (GKR, LogUp, KZG, RISC-V) are not listed.

term meaning defined in
accumulator, accumulator entry a deferred Mercury check as twelve (side, scalar, point) entries mercury.md §6
advice memory whose initial values the prover chose, bound by nothing public-values.md §6
anchor, anchor space tuples in a delegation type's own space pairing a request with one invocation delegation.md §5
archived path proving from a held TraceArchive; only the tamper suite (checker::TamperHarness) does streaming.md §6
artifact a circuit as data, CircuitArtifact; also an exported ProgramImage gkr.md §4, program.md §3
base claims each committed column's claimed value where the backward pass ends gkr.md §5
base format, recursion format recursion if a statement's VmConfig holds FIELD_WINDOWS, else base recursion.md §1
block BlockProof: config, statement and its shards' proofs proof.md §1
bound wire a decider value the verifier holds, committed instead of a public input recursion.md §9
boundary the registers' and pc's final timestamps and values; they have no rows memory.md §4
cached entry a sub-expression inlined into its list's gates, not a column gkr.md §3
canonical form an element as its value, 32 bytes little-endian, below the modulus primitives.md §1
challenge slot a gate coefficient's challenge: drawn, or derived by the verifier gkr.md §3, §5
channel one LogUp identity over a shard's lookups into one table lookup.md §1
copower x < p as x·2^32/p < 2^32, void without a direct bound lookup.md §11
cycle-owning the execution families 0–6, whose time windows are ordered proof.md §8
decider a Groth16 proof that the recursion root verifies, for the contract recursion.md §9
declaration record, static detachment 12 bytes a linked shim leaves in the image: how a delegation is declared delegation.md §7
decoded table an instruction family's setup columns: row i is pc 2i program.md §5
delegation a family proving a function of a RAM frame, invoked by ecall delegation.md §1
discharge spending an accumulator; the rule that each lookup is one leaf of its tree mercury.md §6, lookup.md §11
enforcing, producing a gate vanishing on every row; one writing the next layer gkr.md §1
extra mask, kind, kind bit family_extra_mask = 1 << kind, a kind being a mnemonic's index; b_k its bit program.md §6
family a circuit and the rows it proves: instructions (0–6), memory locations or invocations circuits.md §1
field memory address space FIELD: cells of one Fr, for the recursion families recursion.md §2
fold merging a node's deferred Mercury checks into one (A, B) recursion.md §8
frame an execution family's queries; a delegation's RAM words at a0 memory.md §2, delegation.md §4
gate list, row-wise, halving the gates from layer k to k + 1, keeping the height or halving it gkr.md §1
gated key, neutral tuple a lookup tuple under its selector; off, it reads a neutral table row lookup.md §4
generic table the committed table of ZeroEntry, AND, U16GetSign, ShiftPowers lookup.md §9
global transcript, global state digest G1–G11: the statement, M commitments, memory challenges; G11 seeds each shard proof.md §2
HALT_PC 1: the exit row's next_pc, the pc's final value memory.md §5
height a family's rows a shard: 2^8, 2^12, 2^16, 2^18, 2^20 or 2^22 program.md §7
identity, image column one Fr digest of the decoded tables, the image, the entry pc, VmConfig program.md §8
in flight shards worked at once, at most max_in_flight streaming.md §5
invocation, request a delegation's row doing one call; the ecall row asking for it delegation.md §1, §5
journal the public output: what the guest leaves in the output window public-values.md §1
laws Laws 1–4: locality, derived width, top layer, single source of truth gkr.md §4
layer layer 0 the committed columns, the top the outputs; L{k}[j] between gkr.md §1
leaf, node, root recursion programs: a leaf verifies base shards, a node 2–4 child proofs; the root, all recursion.md §8
live row, padding row m_pc = 1, or a zero row; in a decoded table, an instruction, or −1 throughout memory.md §2, program.md §5
M, W, S, V memory, witness and setup columns; virtual tables gkr.md §2
memory form an Fr's Montgomery limbs x·R; on the wire only in FR_ARITH's frame primitives.md §1
mini-block the revm-block binary: transactions over a recorded pre-state ethereum.md §1
multiplicity a channel's W column counting each table row's lookups lookup.md §7
padding contract padding.row makes row-local relations vanish and tree inputs 1 gkr.md §4
pairing side G2One or G2X: an entry's G2 argument, [1]_2 or [x]_2 mercury.md §6
pass 1, pass 2 executing to commit every shard's M columns; again to prove each streaming.md §2
phase 1, phase 2 the decider key's ceremonies: powers of tau, then the circuit's own recursion.md §9
public window windows 2 and 3 at 2^12: input at 0x8000, journal at 0xC000 public-values.md §2
query one read and one write at one address in one cycle execution-trace.md §3
RAM glue invocations chained through their frame's words in RAM delegation-circuits.md §1
reconciliation ∏ read roots · R_b = ∏ write roots · W_b, once a statement memory.md §4
registry family_circuit, recursion_circuit: each family's one circuit circuits.md §1
scratch scratch[i], a flat relation's intermediate, one per inner column gkr.md §2
shard h rows of one family, or one window, proved alone but for the memory argument streaming.md §4
slot Δ in a cycle's timestamps 4c + Δ; a ProgramImage halfword; a frame position execution-trace.md §1, program.md §2, memory.md §2
SRS digest a digest of the SrsVerifier and the generic table's commitments proof.md §3
stack 2^σ columns committed as one, in the recursion format recursion.md §1
statement PublicInputs: input, journal, exit status and the execution's record proof.md §1
statement shard, shard-set exactness a (family, index) below its count; a block proves each once, in order proof.md §1
tamper twin a forgery proved as an honest prover would, refused in its class circuits.md §3
tape straight-line coprocessor calls a node replays; checker tape's listing recursion.md §7, tools.md §4
time window a shard's claimed [ts_start, ts_end); it binds nothing proof.md §8
transcript form a G1 point as four 128-bit Fr limbs; infinity, four 2^128 transcript.md §4
tuple T(AS, ADDR, TS, VAL): a memory access as one field element memory.md §1
u1, u2 a Mercury opening point's halves, pairing with an index's low and high bits mercury.md §1
VmConfig a program's families, their heights, bytecode_size_words program.md §7
window h words from byte 4h·w, initialized and torn down by one shard memory.md §3
write-side induction an execution family writes only words, so operands need no bound memory-ops.md §5

Referenz

Werkzeuge

Normative Spezifikationdocs/tools.mdAls Markdown anzeigen

Zusammenfassung

Jedes Binary rund um Prover und Verifier, keines davon auf einem Beweispfad: bench, das misst und beweist; der Zyklus-Profiler mit seiner Klassifizierung und Bepreisung von Delegationskandidaten; das Debug-Log des Provers und seine Marker; checker; artifact-dump; das Kommandozeilenwerkzeug verifier; kat-gen und die eingecheckten Fixtures; und die beiden Referenzorakel außerhalb des Workspace.

Der folgende normative Text wird auf Englisch gepflegt, der kanonischen Sprache der Spezifikation.

The binaries around the prover and verifier, none on a proof path: bench measures and proves (§1), profiler counts a guest's cycles (§2), a debug-info build logs a proving run (§3), checker validates circuits and the global transcript (§4), artifact-dump exports a guest's ProgramImage (§5), verifier checks a proof from files (§6), kat-gen regenerates the committed fixtures (§7), and two generators outside the workspace are reference oracles (§8).

1 bench#

cargo run --release -p bench [-- <routine>...]   every routine, or those named; --list lists them
cargo run --release -p bench -- prove <stem> | --stateless <file> [--case <name>]
    [--in-flight <n>] [--out <dir>] [--json <path>] [--hourly-usd <price>] [--toy-srs]

The routines time one component each, over their own data: fr-arith, poly-bind, msm, mercury, mercury-batch, zerocheck-prove, zerocheck-verify, gkr-prove. msm, mercury and mercury-batch run over ceremony bases, assets/ptau/ppot_0080_24.ptau, and return without them.

prove proves a block through host::prove (streaming.md) and verifies it (host::verify).

  • <stem> names a recorded block under crates/host/tests/vectors: its pin <stem>.json, to which <stem>-witness.bin and <stem>-journal.bin are held, names the guest that proves it; mini-block is committed (ethereum.md §6).
  • --stateless <file> is one input to revm-block-stateless. A .json EEST fixture gives its statelessInputBytes as the advice, unchanged, and its statelessOutputBytes as the journal the proof must bind, checked by revm_block::stateless::run first and on the proof after; --case picks one input by part of its name. Any other file is the raw input.
  • The guest is built at --release (host::fixture::build_revm_guest), decoded at host::fixture::revm_params and keyed over 2^22 ceremony powers or, with --toy-srs, over τ = 0xc0ffee, cached as apogee-bench-toy-22.srs in the temporary directory: the same timings, another identity, which the report names.
  • --in-flight is max_in_flight, 8 by default. The verb asserts that the guest exits 0 and the block verifies; --out then writes the proof archive (proof.md §9) under the stem's or the input file's name.

The printed BenchReport (--json writes it too) holds the block, identity, SRS, cycles per gas, shards per family, proof and statement bytes, clocks, peak RSS, cost and hardware. commit and gkr are pass 1's and pass 2's wall clocks; execution, the executor's time, runs inside them and is left out of their total; opening and final are 0; unattributed is the rest of the proving wall clock; setup and verify are apart. Peak RSS is Linux's VmHWM, absent elsewhere, where /usr/bin/time -l gives it. --hourly-usd adds the cost, price · proving_ms / 3,600,000, and the cost per Mgas. Any failure exits 1, a wrong journal or a failed --out after the report prints; a usage error exits 2.

The verbs recurse, recurse-node, ceremony and decide are recursion.md §8.4–§10's.

2 The cycle profiler#

cargo run --release -p profiler -- elf <file> [--advice <f>] [--input <f>] [<common>]
cargo run --release -p profiler -- block <stem> [<common>]
cargo run --release -p profiler -- record <number|latest> [--txs <n>] [--cache <dir>] [<common>]
    <common>: [--top <n>] [--json <path>]

elf runs any guest over the given input and advice, at the smallest menu height its code fits; block runs the revm guest over a recorded fixture; record records a block from ETH_RPC_URL (latest is the finalized one; every transaction unless --txs; cached in target/profiler-cache) and runs revm-block over it, its gas the transactions' limits capped at the block's. A run prints a table, the --top (30) functions in it, and with --json writes a ProfileReport; any error exits 2. Its numbers are counts of executed cycles, the same on any machine.

2.1 One histogram over pc#

profiler::profile adds 1 to one u64 per halfword slot of the image for each executed cycle, reading each chunk's pc column off emulator::StreamingRun and dropping the chunk, so it holds the histogram and one partial buffer per family. Delegation rows add nothing: their requesting cycle is the ecall row's. A function's cycles are the sum over its [st_value, st_value + st_size) (loader::function_symbols), its own and not its callees'; its calls are the count at its first instruction, which runs once a call, so code entered only past its entry shows cycles and no calls. A mnemonic's cycles are the sum over its slots, a category's over its functions', and the unattributed ones are at slots no symbol covers.

2.2 Classification#

tools/profiler/src/categories.rs puts each function in one of 14 categories by RULES, ordered substring rules where the first match wins, then FALLBACK_RULES, the generic runtime paths, each matched against the demangled path and the raw symbol (categories::classify). The order is the meaning: revm_interpreter::instructions::system::keccak256 is hashing because its rule comes before revm_interpreter::'s. Legacy mangling is decoded whole, v0 to its identifiers.

A function's cycles include what was inlined into it: ruint's 256-bit operations count in the EVM opcode handlers, each a symbol of its own, revm dispatching through a table of function pointers. The unattributed share and the mnemonic mix, which no symbol table can misattribute, are the checks on attribution.

2.3 Pricing a candidate#

removable = max(0, cycles − calls·(4 + 2·frame_words))

cycles is the category's, calls the entry counts of the candidate's named entry symbols, and 4 + 2·frame_words (categories::shim_cycles) the shim a delegation leaves: the frame's stores, the ecall, the results' loads. CANDIDATES prices secp256k1, 256-bit arithmetic, BN254, SHA-256/RIPEMD-160 and the keccak sponge; one without entry symbols is charged no shim and flagged. It is a ceiling: it charges nothing for the new family's shards (delegation.md §9) or for marshalling operands into a frame.

3 The proving debug log#

crates/prover/src/debug.rs and the prover's log lines exist only with its debug-info feature, the workspace's one cargo feature: off by default, enabling no dependency, changing no proof byte (crates/prover/tests/debug_info.rs proves one statement with the log off and at deep and compares the blocks). Without it dlog! and debug_only! expand to nothing, so no scan is compiled into a proving run. gkr::explain_self_check is compiled always.

cargo run --release -p bench --features prover/debug-info -- prove ...
cargo test --release -p prover --features debug-info --test <suite> -- --include-ignored
APOGEE_DEBUG=off | phase | detail | deep [:FAMILY,FAMILY]

APOGEE_DEBUG, read at each log site, picks the level, case ignored: unset or empty is phase, none and 0 also mean off, 1 to 3 the other levels. :FAMILY,… (names as the log prints them, or ids) keeps those families at the level and lowers the others one step; lines naming no family stay. A bad level falls back to phase, an unknown family is dropped, and either is reported once as apogee ERROR. Lines go to the raw io::stderr() handle, one locked write each: libtest shows captured eprintln! output only for a failed test, and an OOM kill, a hang or a SIGINT loses it.

level adds
phase identity and SRS digest in full, in to_bytes order as the verifier CLI takes them; each claim's take and its committed or proved; the global digest and memory challenges, on the apogee commit line; each shard's begin h= … gkr done and open begin … open done
detail each family's circuit inventory; each shard's time window, g, β, roots and opening commitments; gkr::self_check; the scans
deep each GKR layer's shape and bytes; the top layer's all-zero columns

Where a run died. A begin without its done names the shard that died (FAMILY#index, [k/N] its statement position); a take without committed or proved, one in flight. fill# is fill order, which picks the failure returned, and in_flight= below the bound mid-pass means the workers wait on the executor. fill_ms is the one-thread fill, ms a wall clock shared with the shards in flight. Every shard forks from the apogee commit line's values, so two runs that should agree diverge there or inside a shard.

self_check recomputes every gate on every row before the backward pass, a second forward pass (gkr.md §5). gkr::explain_self_check turns a failure into the row's first disagreeing gate and every operand's value, a committed column by its artifact name and an inner one by the relation that wrote it, where a verifier says only LayerInconsistency { layer }.

The scans read each base delegation shard's live rows: invocations against the height, cycle and frame-base ranges, timestamp gaps, selector and round histograms, and canonicity, a tally for POSEIDON2 and FR_ARITH, whose < p conclusions are gated to the rows that read a value, and a verdict for MOD_MUL's operands and the values each EC_ADD row's group reads (debug::ec_add_reads). On ADD_SUB_LUI_AUIPC they count requests per type, which sum to each delegation family's invocations, and exit rows, one in all. The log's verdicts:

marker
self_check FAILED a gate fails on the prover's own values
NOT CANONICAL a frame value at or above its modulus where a gate needs it below
UNBALANCED an EC_ADD curve whose three groups' counts differ
OVER the a timestamp gap beyond 38 bits
NAMES NO MODULUS a MOD_MUL selector naming no modulus
DISAGREES a SHA256_COMP frame its rounds do not produce: the fill's refusal, in every build
NOT LOOPING 24 TIMES KECCAK_F round counts more than 1 apart
ABORTED a nonzero exit status: the block proves a failed execution
OUTPUT-LAYOUT-BREAK outputs other than 2 + 2·channels: reduce_shard and channel_cones index channel roots from opposite ends
ALL ZERO a top-layer column all zero: a root of 0
DECLARED BUT NEVER INVOKED a delegation shard with no live row
console
$ APOGEE_DEBUG=detail <a debug-info run> 2>&1 | tee run.log
$ grep -c 'begin h=' run.log; grep -c 'gkr done' run.log    # unequal: a shard died
$ grep 'begin h=' run.log | tail -1
$ grep -E 'FAIL|NOT CANONICAL|UNBALANCED|OVER the|NAMES NO|DISAGREES|NOT LOOPING|ABORTED' run.log
$ grep -E 'LAYOUT-BREAK|ALL ZERO' run.log

At detail the self-check doubles each shard's forward work and the scans cost O(live rows × frame words); deep reads no layer's cells but the top's.

4 checker#

cargo run -p checker -- laws <artifact>       Laws 1–4, then the lookup rules (check_laws)
cargo run -p checker -- padding <artifact>    the padding contract (check_padding)
cargo run -p checker -- dump <artifact>       the circuit, readably (checker::dump)
cargo run -p checker -- tape <verifying-key> <public-inputs>

An artifact is a CircuitArtifact file, decoded for encoding only so that a lawless one reaches the checks, such as crates/constraints/tests/vectors/*.bin. The validators are circuits.md §3's, independent of constraints; padding omits the product-tree clause; dump prints any decodable artifact.

tape loads a key (verifier::load_verifying_key) and a PublicInputs file, an archive's .vk and .public, refuses a statement the key does not describe (verifier_core::derive_global_phase), and runs checker::check_global_tape. That renders the global commit phase's event log a line a message, absorb <TAG> <n> (n scalars, or a bytes message's 31-byte chunks) or squeeze <TAG>, and holds it to expected_global_tape: G1–G11 (proof.md §2) written from the statement's shape, sharing nothing with verifier_core::global_commit but statement_shards. It prints the tape or the first line out of order, and checks the script, not the values, which the log does not carry. checker exits 0 when a check holds or a listing prints, 1 naming the failure, 2 on a usage error.

5 artifact-dump#

cargo run -p artifact-dump -- <guest.elf> [--out <dir>]
cargo run --release -p artifact-dump -- tables <guest.elf> [--ptau <file>]

The first writes <name>.img, the ELF's ProgramImage in its postcard wire form with nothing around it (program.md §3), and <name>.img.txt, a report rendered from the image read back off those bytes, which must equal the loaded one or nothing is written: segments, the listing (address, length, encoding, expanded word), .symtab names marked as outside the artifact, and the artifact's and the ELF's SHA-256, which pin bytes and are not the program identity. <name> is the ELF's stem; --out defaults to the working directory.

tables prints the VmConfig and each instruction's pc, next_pc, family, mnemonic and decoded fields at ProgramParams::defaults(); with --ptau it reads 2^22 powers, the largest default height, and prints the program identity (program.md §8).

6 The verifier CLI#

cargo run --release -p verifier -- <verifying-key> <identity-hex> <public-inputs> <proof>...
cargo run --release -p verifier -- block <verifying-key> <identity-hex> <public-inputs> <block>

The key is loaded by verifier::load_verifying_key (proof.md §7) and its identity must equal <identity-hex>, 64 lowercase hex digits of its canonical bytes from a channel the prover does not control: never the key, the proof or an archive's .identity. The first form verifies each ShardProof file with verify_shard and requires the files to be the statement's shards, each once, in any order; the second verifies a BlockProof with verify_block, as an archive's .vk, .public and .block (proof.md §9). It takes no SRS digest, using the key file's (srs.md §3). Exit 0 when all verifies, 1 naming the first file refused or a wrong shard set, 2 on usage or a malformed identity.

7 kat-gen and the committed fixtures#

cargo run -p kat-gen regenerates the default groups, cargo run -p kat-gen -- <group> one. Each file written prints its SHA-256, which the tests reading it pin.

group writes from
field, poly, curve, tower, pairing, msm, srs, moduli arithmetic, ceremony-point, KZG and MOD_MUL modulus vectors arkworks; srs's points through its own .ptau reader
pcs G1 absorption limbs; Mercury proofs arkworks; pcs
loader, isa listings of the committed guest ELFs, synthetic ELFs; an RV32IMA corpus, words that must not decode the pinned toolchain's llvm-objdump, llvm-nm
program, tape program identities, the generic table's commitments; guests/shards' global tape program; checker
gkr, memory, lookup, family, delegation CircuitArtifact files: toy circuits; the four frames, the two RAM-window circuits and the seven execution circuits, at 2^22; each base delegation circuit's shape and SHA-256 constraints
revm a synthetic block's witness, output commitment and delegated keccak-f frames native revm, held to the guest
opt-in: block, zkevm, guests mini-block, over ETH_RPC_URL (ethereum.md §6); zkevm-subset.json, cut from the tests-zkevm release at APOGEE_ZKEVM_FIXTURES only if every pair matches; the guest ELFs, each built twice and compared

srs, and program's identities and table commitments, need assets/ptau/ppot_0080_24.ptau and are skipped without it. CI runs the default groups and both oracles (§8) and fails on any git diff in the vector directories. A guest ELF is not reproducible across machines, since rustc embeds absolute paths in the panic-location strings of core and of crates outside the guest workspace, which the guest build does not remap; two clean builds on one machine agree. So guests is run by hand on one machine, and CI regenerates only what derives from the ELFs.

8 Reference oracles#

cargo run --manifest-path tools/transcript-ref/Cargo.toml
cargo run --manifest-path tools/stateless-ref/Cargo.toml

tools/transcript-ref implements transcript.md from its text over Plonky3's Poseidon2 and HorizenLabs zkhash's round constants, pinned by revision, and writes crates/transcript/tests/vectors/: permutation vectors, transcript scripts and io_digest cases. tools/stateless-ref encodes stateless inputs with eth-act/ere-guests v0.17.1's stateless-validator-common over libssz 0.3.0 and writes stateless_ref.txt under crates/host/tests/vectors/: per input, its request's hash_tree_root or reject. Each is its own workspace root because its dependencies enable features, serde/std among them, that cargo's feature unification would carry into the workspace's no_std crates; the one repository crate either links is tools/test-support, a seeded RNG, SHA-256 and hex with no dependencies.

Referenz

Repository-Übersicht

Wo im Repository von Apogee VM was liegt, was jedes Crate ist und welche Seite der Spezifikation es definiert.

Als Markdown anzeigen

Das Repository von Apogee VM besteht aus zwei Cargo-Workspaces: dem Root-Workspace für alles, was auf Ihrem Host läuft, und guests/ für alles, was in der VM läuft.

Crates#

Pfad Was es ist Spezifiziert in
crates/constants jede Protokollkonstante, jedes Tag und jeder Bezeichner; keine Logik die Seite, die sie jeweils verwendet
crates/field, curve, poly, sumcheck Fr; der Fq-Turm, G1, G2, das Pairing, MSM; multilineare Polynome; der Zerocheck Primitive
crates/transcript Poseidon2 und das Duplex-Transkript Transkript
crates/srs Einlesen der Zeremonie, das SRS-Archiv, KZG, Phase 1 von Groth16 SRS
crates/pcs, pcs-verify Mercury und aufgeschobene Verifikation; pcs-verify ist die Körperseite der Verifikation Mercury
crates/loader, isa, program ELF zu ProgramImage; der Decoder; dekodierte Tabellen, VmConfig, Programmidentität Programm und Identität
crates/emulator, trace der Executor und seine Tracer; Zeilen, Speicherzustand, Spalten-Builder Ausführungs-Trace
crates/constraints jeder Schaltkreis als Daten: Speicher-Frames, Lookup-Kanäle, die Familienschaltkreise, die Registries GKR-Engine, Speicherargument, Lookups, Schaltkreise und die Familienseiten
crates/gkr-verify, gkr der GKR-Verifier und der GKR-Prover GKR-Engine
crates/verifier-core Aussage, Transkripte, Verifikationsschlüssel, jede Prüfung eines Shards und eines Blocks außer der Öffnung; Tapes, Knoten und Falten der Rekursion Der Beweis, Rekursion
crates/verifier verify_shard, verify_block, das Beweisarchiv, die CLI verifier Der Beweis
crates/prover Schlüsselkonstruktion, Befüllen der Spalten, der Streaming-Prover, das Debug-Log Streaming-Prover
crates/groth16 Groth16 mit gebundenen Wires und einer zweiphasigen Zeremonie Rekursion §9
crates/host das Host-SDK: Setup, Beweisen, Verifizieren; der Recorder für Block-Witnesses; Rekursionsbaum und Decider Ethereum-Blöcke, Rekursion
crates/checker unabhängige Validatoren der Schaltkreisgesetze, native Evaluatoren für Lookups und Speicher, die Manipulations-Suite, die CLI checker Schaltkreise §3
crates/guest-sdk die Laufzeit der Gastprogramme (Guests): Einsprung, Linker-Skript, Allokator, Speicherbereiche, Delegations-Shims Gastprogramm-ABI, Delegations-ABI
guests/ Gastprogramme für Tests und Arbeitslasten, ein eigener Workspace; vendor/ enthält gepatchte Upstream-Crates Beispiel-Gastprogramme
contracts/ ApogeeVerifier.sol Rekursion §9
tools/ kat-gen, bench, profiler, artifact-dump, test-support; transcript-ref und stateless-ref, unabhängige Orakel außerhalb des Workspace Werkzeuge und CLIs
docs/ die Architekturübersicht, das Glossar, das Handbuch für Gastprogramme, die Werkzeugseite und spec/, eine Seite pro Thema diese Website

Voraussetzungen#

  • Die Toolchain, ihre Komponenten und das Target riscv32imac-unknown-none-elf sind in rust-toolchain.toml festgelegt; rustup installiert sie bei der ersten Verwendung.
  • Programmidentität, echte Schlüssel und Beweiserzeugung benötigen die Zeremoniedatei assets/ptau/ppot_0080_24.ptau. Die Workspace-Tests nicht.
  • Die Beweiserzeugung ist speichergebunden: Ein vollständiger Block erreichte eine Spitze von 174 GiB.

Befehle#

sh
# What CI runs
cargo fmt --all -- --check
cargo clippy --workspace --all-targets -- -D warnings
cargo test --workspace
cargo run -p kat-gen && git diff --exit-code     # committed fixtures regenerate identically

# Guests: their own workspace and target
(cd guests && cargo clippy --bins -- -D warnings)
(cd guests/fib && cargo build --target riscv32imac-unknown-none-elf)   # --release for proving

# Prove and verify a block, then recurse and decide
cargo run --release -p bench -- prove mini-block --out <dir>
cargo run --release -p verifier -- block <stem>.vk <identity-hex> <stem>.public <stem>.block
cargo run --release -p bench -- recurse <dir>/<stem> --out <out>

Die Suiten, die echte Shards beweisen, sind mit #[ignore] markiert, und die CI führt sie nicht aus: Jede beweist über einem eigenen Spielzeug-SRS und braucht Dutzende GiB.

sh
cargo test --release -p prover --test <suite> -- --include-ignored --test-threads=1
#   acceptance, control, alu, mem, fills, block, streaming, keccak, recursion, public_io, revm
cargo test --release -p host --test prove -- --include-ignored --test-threads=1     # a mainnet mini-block
cargo test --release -p checker --test tamper -- --include-ignored --test-threads=1 # every tamper twin

Referenz

Versionshinweise

Apogee VM v1.0.0, das erste Release. Was es beweist, was es enthält, wie es gemessen und geprüft wurde und wo seine bekannten Grenzen liegen.

Als Markdown anzeigen

v1.0.0#

Das erste Release von Apogee VM: eine RISC-V-zkVM, die RV32IMAC-Programme beweist und ihre Beweise auf Ethereum abwickelt. Die Spezifikation, die diese Website wiedergibt, ist das Verzeichnis docs/ des Repositorys in der Quellrevision 3571370.

Was es beweist#

Dass ein Programm, benannt durch einen Digest seines Images, mit einer öffentlichen Eingabe bis zu einem Exit-Status lief und ein Journal schrieb – über einen Rekursionsbaum bis zu einem einzigen Groth16-Beweis getragen, den ApogeeVerifier.sol prüft. Beweise sind succinct, aber nicht zero-knowledge.

Was enthalten ist#

  • Die Maschine. RV32IMAC auf einem Hart; die 59 Befehle von RV32IMA, komprimierte Befehle beim Laden expandiert; ein Guest-SDK mit drei Speicherbereichen für Eingabe, Hilfsdaten (Advice) und Ausgabe.
  • Das Beweissystem. 23 Schaltkreisfamilien über dem Skalarkörper von BN254, jede ein geschichteter GKR-Schaltkreis: sieben Befehlsfamilien, fünf Speicherfenster-Familien, sechs Delegationen und fünf Rekursionsfamilien. Eine einzige Lese-/Schreib-Speicher-Multimenge über die gesamte Ausführung; LogUp-Lookups über fünf Kanäle.
  • Delegationen. KECCAK_F, SHA256_COMP, POSEIDON2, FR_ARITH, MOD_MUL und EC_ADD, erreichbar aus dem SDK und aus gepatchten Versionen von k256, ark-ff und revm-precompile.
  • Commitments. Mercury über KZG auf den Perpetual Powers of Tau der PSE, eine 704-Byte-Öffnung pro Shard; aufgeschobene Verifikation für die Rekursion.
  • Der Prover. Ein Streaming-Prover mit zwei Durchläufen, dessen Speicherbedarf den gerade bearbeiteten Shards folgt.
  • Abwicklung. Ein Rekursionsbaum aus Blatt- und Knotenprogrammen in einem Rekursionsformat mit Speicher für Körperelemente und vier Koprozessoren; ein Groth16-Decider mit gebundenen Wires und einer zweiphasigen Zeremonie; ApogeeVerifier.sol.
  • Die Ethereum-Arbeitslast. Ein revm-Gastprogramm (Guest) mit einem Mini-Block-Binary und einem zustandslosen Validator für Osaka, BPO1, BPO2 und Amsterdam.
  • Werkzeuge. bench, der Zyklus-Profiler, das Debug-Log des Provers, checker, artifact-dump, die CLI verifier, kat-gen und zwei Referenzorakel.
  • Keine externe Kryptografie. Körper, Kurve, Pairing, MSM, Hash, PCS, GKR und Groth16 sind im Repository implementiert.

Gemessen#

Block 257.510 von glamsterdam-devnet-8 (60 Transaktionen, 101,5 Mgas, 198 Mio. Zyklen): ein Basisbeweis aus 207 Shards in 2.481 s auf 32 vCPUs mit einer Spitze von 174 GiB; ein Rekursionsbaum aus 116 Shards; ein Decider-Beweis in 18,5 s; On-Chain-Verifikation für 3.620.026 gas. Alle 67.251 Paare von tests-zkevm v21.0.1 stimmen nativ überein. Performance enthält jede Zahl.

Bekannte Grenzen#

Nicht zero-knowledge; Hilfsdaten bewusst ungebunden; höchstens je 16.380 Byte öffentliche Eingabe und Journal; sc.w gelingt immer; Traps sind nicht beweisbar; eine feste Menge von sechs Basis-Delegationen; Speicherbedarf des Provers bestimmt durch die gerade bearbeiteten Shards; der Decider-Schlüssel je Wurzelform und nur so vertrauenswürdig wie seine Zeremonie. Das Sicherheitsmodell nennt jede Grenze mit ihrer Begründung.

Dokumentation#

Diese Website, auf Englisch, Französisch (Kanada), vereinfachtem Chinesisch und Deutsch, mit der normativen Spezifikation in jeder Sprache auf Englisch. Der KI-Begleiter und llms.txt dienen KI-Agenten.

404

Diese Seite hat ihre Umlaufbahn verlassen

Unter dieser Adresse gibt es nichts. Durchsuchen Sie die Dokumentation oder beginnen Sie noch einmal beim ersten Schritt.