# Delegations

> How an expensive function gets a circuit of its own without growing the instruction circuits. The call, the anchor that pairs each request with exactly one invocation, the six circuits, and their economics.

Hashing and big-integer arithmetic dominate real workloads: in a mainnet Ethereum block, secp256k1's field multiplication and squaring alone were 44% of the cycles before they were delegated. Proving them instruction by instruction is possible and slow. A **delegation** gives such a function a circuit family of its own, invoked from the guest, so the instruction circuits stay small and a program pays only for the delegations it calls.

## The call

A delegation is invoked, never decoded. The guest writes a frame of 32-bit words in RAM and issues an `ecall` with the delegation's number in `a7` and the frame's base address in `a0`. The `ecall` is one row of the `ADD_SUB_LUI_AUIPC` family, the **request**. The work is one row of the delegation's own family, the **invocation**, which reads every frame word and writes every frame word back, the results among them, at the requesting cycle. An invocation owns no cycle; it rides the one that asked for it.

Frames are word-aligned and lie wholly in ordinary RAM, so no frame overlaps a public window or advice, and the frame's reads and writes are ordinary memory queries. What a delegation computed is therefore bound exactly as any store is: through the one memory multiset.

## The anchor

Requests and invocations must pair one to one: otherwise many requests could close against one invocation and leave calls unexecuted, or an unrequested invocation could rewrite a frame. They pair through the same memory multiset, in an **anchor space** that belongs to the delegation type alone and that no instruction can reach:

| | Reads | Writes |
| --- | --- | --- |
| Request, cycle `c` | `T(s, base, 0, 0)` | `T(s, base, 4c + 3, v)` |
| Invocation | `T(s, base, 4c + 3, v′)` | `T(s, base, 0, 0)` |

Three gates on the request side fix its read at timestamp 0 and value 0 and make it write 0 to `a0`. Then the tuples stamped 0 are exactly the requests' reads and the invocations' answers, so there are as many invocations as requests over the same bases; and since no two requests share a cycle, each invocation's read is exactly one request's write. Every invocation sits at its request's base and cycle. No gate in a delegation circuit had to know about requests at all.

## Many calls, one operation

An operation too wide for one row is several invocations on one frame, a frame word naming the step: a keccak-f[1600] permutation is 24 round calls, a SHA-256 compression 16 calls of four rounds, a complete point addition three calls. No gate joins two rows. Each call proves its step on the frame as it finds it, its reads lying on each word's single memory history, so it reads the previous step's writes. That every step runs, in order, is the calling code's to ensure, and the calling code is guest code proved as instructions. The SDK issues each multi-call operation from one function, so a guest never orders the steps by hand.

## Declared statically

The instruction sweep cannot see a call, because the number is a run-time value of `a7`. So each shim in the SDK leaves a 12-byte declaration record in its own linker section, kept only if the shim is reachable. The program derivation scans the image for records, and a declared family joins the configuration, bound by the identity through the image bytes. A family linked but never called proves zero shards; a called number whose family the program never declared has no proof.

## The six circuits

| Family | One invocation | Built from |
| --- | --- | --- |
| `KECCAK_F` | one round of keccak-f[1600] over a 51-word frame | bytes: 1,020 `XOR8` lookups a round; rotations as linear forms over bytes and masked copies |
| `SHA256_COMP` | four rounds and four schedule words | bytes and `XOR8`: 52 obligations a round, 32 a schedule word; `Ch` and `Maj` as linear forms in XORs |
| `POSEIDON2` | one width-3 permutation | the rounds computed in the circuit's own layers, three gate lists a round, with no lookup; the only delegation that computes above its first layer |
| `FR_ARITH` | one `Fr` add, multiply or inverse in Montgomery form | bit decompositions and canonicity chains against `p` |
| `MOD_MUL` | one 256-bit `a·b mod m`, four Ethereum moduli | 32-bit limbs, a quotient, carries, and a canonicity chain proving `out < m` |
| `EC_ADD` | one third of a complete point addition on secp256k1 or BN254 G1 | Renes–Costello–Batina's complete formula as three reductions a row |

A few constructions recur across them. A **one-code rule** decodes a frame word naming one of `k` cases into boolean selectors with exactly one set, because codes add: without it, selectors 1 and 3 answer a request for 4. A **canonicity chain** proves a 256-bit value is below a modulus through borrows over 32-bit limbs. And every written word is bounded below `2^32`, so that RAM stays words, which every instruction family relies on.

## The economics

A delegation family's height sets how many calls one shard holds, and a shard costs its height whatever its occupancy:

| Family | Height | Units a shard | Shard proof |
| --- | --- | --- | --- |
| `KECCAK_F` | `2^18` | 10,922 permutations | 381,100 B |
| `SHA256_COMP` | `2^18` | 16,384 compressions | 189,988 B |
| `EC_ADD` | `2^16` | 21,845 additions | 434,916 B |
| `MOD_MUL` | `2^16` | 65,536 multiplications | 135,220 B |
| `POSEIDON2` | `2^8` | 256 permutations | 664,780 B |
| `FR_ARITH` | `2^8` | 256 operations | 266,292 B |

For a family with many calls, the fatter shard is the cheaper one: a `KECCAK_F` proof barely grows from `2^16` to `2^18`. The price is memory. A `2^18` `KECCAK_F` shard's forward pass is 42 GiB, and two of them in flight set the measured block's peak.

## What is delegated, and what is not

Library code reaches the delegations through patched copies of `k256`, `ark-ff` and `revm-precompile`: secp256k1 recovery becomes `k256` code over `MOD_MUL` and `EC_ADD`, and a BN254 pairing becomes `ark-bn254` code over `MOD_MUL`. The EVM's `MULMOD` with an arbitrary modulus, `MODEXP`, BLS12-381 and every whole signature scheme run as instructions. Dedicated signature support for guests is part of the [v2.0.0 trajectory](https://apogee.gweb3networks.com/docs/quantum-leap/signatures).

The specification: [Delegation ABI](https://apogee.gweb3networks.com/docs/auditors/spec/delegation), [Delegation circuits](https://apogee.gweb3networks.com/docs/auditors/spec/delegation-circuits).
