From e9b3c564343b3812032df150d4fd6593f9f68d55 Mon Sep 17 00:00:00 2001 From: Tom Wambsgans Date: Mon, 14 Sep 2026 19:51:49 -0400 Subject: [PATCH 01/40] Memory words become 64 bits leanVM memory cells hold words of K = GF(2^64) instead of F192 elements. The ISA has eight instructions: XOR64, MUL64, SET, DEREF, JUMP, BLAKE2S, XOR192 and MUL192. The two 192-bit ones read three consecutive cells as an element of F192 = K[Y]/(Y^3+Y+1), limbs low first, and are appended as tables 6 and 7, so the existing table indices keep their meaning. - VM: tables, execution, witness layout and the memory bus carry one value column. BLAKE2S reads its block as four two-word chunks, its chaining value and digest as four words, and its metadata as two. The public input is four words; after the table sumcheck the verifier samples two challenges and adds one claim on the memory column whose value it computes, so nothing rides the stream for it. The Rust verifier, the Python verifier and the recursion guest implement the same protocol. - Compiler: `+` and `*` stay 64-bit. A 192-bit value is a three-cell run, computed by add192, mul192 and div192 (the quotient back-solved and pinned by one MUL192), with f192 constants, assert_eq192, assert_ne192, slice stores, runs through calls and match dispatch, and compile-time folding of 192-bit constants. blake2s operands are four-cell runs, and a computed chunk word is evaluated straight into its pair. - Recursion guest: transcript scalars and challenges are three-word runs, 16-byte hash values two words and digests four. The WHIR opening folds each query row at its fold point and writes challenges directly at their witness coordinate, and list hash tails run in power-of-two chunks, which keeps the bytecode at 2^18. Measured on this machine against main (b7a10725), best of three passes: - aggregate --xmss 900: 1,573,474 to 2,188,060 cycles, committed 2^26.195 to 2^26.384, 1.010 s to 1.103 s - aggregate --sphincs 220: 1,914,871 to 2,318,383 cycles, committed 2^26.301 to 2^26.160, 1.105 s to 1.109 s - recursion --n 2 --xmss-per-leaf 900 --log-inv-rate 2: 570,223 to 834,835 cycles, committed 2^24.086 to 2^24.681, 0.359 s to 0.487 s Also fixes three aggregation tamper tests that failed before reaching the guest check they target. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_017HVkj8EFExoGGa1okY9gvX --- crates/lean_compiler/src/ast.rs | 25 +- crates/lean_compiler/src/filler.rs | 22 +- crates/lean_compiler/src/ir.rs | 31 +- crates/lean_compiler/src/lib.rs | 46 +- crates/lean_compiler/src/lower.rs | 221 +- crates/lean_compiler/src/lower/builtins.rs | 233 +- crates/lean_compiler/src/lower/call.rs | 140 +- crates/lean_compiler/src/lower/eval.rs | 56 +- crates/lean_compiler/src/lower/mem.rs | 8 +- crates/lean_compiler/src/lower/run.rs | 331 ++ crates/lean_compiler/src/parser.rs | 102 +- crates/lean_compiler/src/parser/consts.rs | 20 +- crates/lean_compiler/src/parser/subst.rs | 1 + .../tests/programs/const_params.py | 27 +- .../lean_compiler/tests/programs/field192.py | 32 + .../tests/programs/hash_heap_chain.py | 25 +- .../tests/programs/hash_slices.py | 36 +- crates/lean_compiler/tests/programs/unroll.py | 16 +- .../lean_compiler/tests/programs/wots_walk.py | 34 +- crates/lean_compiler/tests/suite/assert_ne.rs | 50 +- .../lean_compiler/tests/suite/common/mod.rs | 19 +- .../tests/suite/const_placeholder.rs | 9 +- .../lean_compiler/tests/suite/determinism.rs | 31 +- .../lean_compiler/tests/suite/disassemble.rs | 19 +- crates/lean_compiler/tests/suite/field192.rs | 157 + crates/lean_compiler/tests/suite/field_div.rs | 27 +- crates/lean_compiler/tests/suite/filler.rs | 12 +- .../tests/suite/hint_log2_ceil.rs | 28 +- .../lean_compiler/tests/suite/inline_expr.rs | 8 +- .../lean_compiler/tests/suite/loop_frames.rs | 12 +- crates/lean_compiler/tests/suite/main.rs | 2 +- crates/lean_compiler/tests/suite/pack64x2.rs | 57 - .../lean_compiler/tests/suite/print_debug.rs | 4 +- crates/lean_compiler/tests/suite/py_source.rs | 33 +- .../lean_compiler/tests/suite/range_check.rs | 28 +- crates/lean_compiler/tests/suite/sharing.rs | 92 +- .../tests/suite/soundness/cases.rs | 241 +- .../tests/suite/soundness/mod.rs | 46 +- .../tests/suite/soundness/pairs.rs | 229 +- .../lean_compiler/tests/suite/stack_bits.rs | 28 +- crates/lean_compiler/tests/suite/stack_buf.rs | 351 +- .../lean_compiler/tests/suite/statements.rs | 176 +- .../tests/suite/transcript_helpers.rs | 71 +- crates/lean_compiler/tests/suite/vm_proofs.rs | 28 +- crates/lean_vm/src/cpu/execute.rs | 353 +- crates/lean_vm/src/cpu/filler.rs | 40 +- crates/lean_vm/src/cpu/hints.rs | 6 +- crates/lean_vm/src/cpu/isa.rs | 45 +- crates/lean_vm/src/cpu/layout.rs | 107 +- crates/lean_vm/src/cpu/mod.rs | 378 +- crates/lean_vm/src/cpu/trace.rs | 50 +- crates/lean_vm/src/hash_flock.rs | 42 +- crates/lean_vm/src/lib.rs | 16 +- crates/lean_vm/src/tables.rs | 759 ++-- .../tests/verifiers/python_verifier.rs | 43 +- .../rec_aggregation/guests/lean_ethereum.py | 3368 ++++++++--------- crates/rec_aggregation/src/aggregation.rs | 784 ++-- crates/rec_aggregation/src/fibonacci.rs | 8 +- crates/rec_aggregation/src/hash_chain.rs | 49 +- python-verifier/verifier.py | 170 +- 60 files changed, 4864 insertions(+), 4518 deletions(-) create mode 100644 crates/lean_compiler/src/lower/run.rs create mode 100644 crates/lean_compiler/tests/programs/field192.py create mode 100644 crates/lean_compiler/tests/suite/field192.rs delete mode 100644 crates/lean_compiler/tests/suite/pack64x2.rs diff --git a/crates/lean_compiler/src/ast.rs b/crates/lean_compiler/src/ast.rs index 73e75f3e2..003efdcb2 100644 --- a/crates/lean_compiler/src/ast.rs +++ b/crates/lean_compiler/src/ast.rs @@ -1,12 +1,13 @@ //! The surface AST produced by the parser: expressions, statements, functions. -use primitives::field::F192; +use primitives::field::F64; -/// An expression. Arithmetic is the field's own: `+` is `XOR`, `*` is `MUL`. +/// An expression. Arithmetic is the 64-bit field's own: `+` is `XOR64`, `*` is +/// `MUL64`; the 192-bit operations are the `add192`/`mul192`/`div192` builtins. #[derive(Clone, Debug, PartialEq)] pub enum Expr { - /// Integer / field literal: the source syntax provides a raw 128-bit value, - /// embedded into the low two limbs of the 192-bit tower element (`c2 = 0`). + /// Integer literal. Compile-time integer arithmetic reads all of it; as a + /// value it is a 64-bit word, so it must fit in one. Lit(u128), /// The generator `g`, written `GEN`. A logical index `i` rides the exponent /// as `gⁱ`, so `GEN` is the unit step. @@ -57,8 +58,7 @@ pub enum Expr { HeapBufDyn(Box), /// `StackBuf(n)`: allocate `n` *consecutive* frame (stack) cells, bound as a /// stack value. Its cells `sa[0..n]` are written/read directly (no heap deref), - /// and a size-2 `StackBuf` is a valid `blake2s` operand (the four 64-bit hash - /// words live as two lanes in each of two consecutive 128-bit cells). + /// and a size-4 `StackBuf` is a valid `blake2s` operand (four 64-bit words). StackBuf(u64), /// `arr[idx]`: read a cell. For a heap `arr` (a pointer), `m[arr·idx]` (idx a /// g-power). For a [`Expr::StackBuf`]: the frame cell `base + idx` (idx a @@ -67,8 +67,8 @@ pub enum Expr { /// `buf[lo:hi]`: a run of cells of a [`Expr::StackBuf`] (frame cells /// `base+lo..base+hi`) or of a [`Expr::HeapBuf`] (heap cells /// `ptr·g^lo..ptr·g^hi`), with compile-time integer bounds (`hi` - /// exclusive). Only meaningful as a `blake2s` operand, where it must span - /// exactly 2 cells (one 256-bit value). + /// exclusive), or a runtime heap start `buf[i:i + k]`. A run value: an + /// operand of `blake2s` or of a 192-bit builtin, a run binding or a run store. Slice(Box, Box, Box), /// `[a, b, …]`: an initialized [`Expr::StackBuf`], so `x = [a, b]` allocates /// a StackBuf of the element count and writes each element in place, sugar @@ -161,6 +161,9 @@ pub enum StmtKind { }, /// `arr[idx] = value`: store into a heap cell (write-once). Store(Expr, Expr, Expr), + /// `buf[lo:hi] = value`: store a run value (a 192-bit element, a digest) into + /// the slice's cells, one write each. + StoreRun(Expr, Expr), /// `for i in mul_range(GEN ** lo, stop)`: the counter rides the exponent as /// `gⁱ`, advancing by `×g` from `g^lo` until it reaches `stop`, which is not /// itself executed. There is no step knob, the bounds being field elements @@ -290,7 +293,7 @@ pub struct Ast { /// Indexed `NAME[i]` and measured `len(NAME)` at compile time only, `i` being /// a literal, a constant or an `unroll` variable. Unlike a scalar constant /// these are not textually substituted, but resolved at lowering. - pub const_arrays: Vec<(String, Vec)>, + pub const_arrays: Vec<(String, Vec)>, } // Free-variable analysis: pure AST, no lowering state. Its one consumer is the @@ -443,6 +446,10 @@ pub(crate) fn free_vars_stmt<'a>(s: &'a Stmt, refs: &mut Vec<&'a str>, bound: &m free_vars_expr(idx, refs); free_vars_expr(val, refs); } + StmtKind::StoreRun(target, val) => { + free_vars_expr(target, refs); + free_vars_expr(val, refs); + } StmtKind::Return(es) => es.iter().for_each(|e| free_vars_expr(e, refs)), StmtKind::For { hi, body, .. } => { if let ForBound::Runtime(b) = hi { diff --git a/crates/lean_compiler/src/filler.rs b/crates/lean_compiler/src/filler.rs index 5c4f1c66f..2d752fca3 100644 --- a/crates/lean_compiler/src/filler.rs +++ b/crates/lean_compiler/src/filler.rs @@ -25,14 +25,14 @@ /// What a table's dummy instruction is: the cheapest instruction of that opcode that can /// be executed any number of times in one frame, given write-once memory. All but -/// `Blake2s` name a single scratch cell as every operand, so the value they write there is +/// `Blake2s` name one scratch run as every operand, so the value they write there is /// the value already there (`FnLower::lower_filler_blocks` fixes the frame offsets). #[derive(Clone, Copy, Debug, PartialEq, Eq)] pub enum FillerOp { - /// `XOR s, s -> s`: pins the scratch cell to `m[s] + m[s] = 0`. - Xor, - /// `MUL s, s -> s`: pins it to `m[s]^2`, so to `0` given the above. - Mul, + /// `XOR64 s, s -> s`: pins the scratch cell to `m[s] + m[s] = 0`. + Xor64, + /// `MUL64 s, s -> s`: pins it to `m[s]^2`, so to `0` given the above. + Mul64, /// `SET s = 0`. Set, /// `DEREF` through the frame's pointer cell, which the interpreter sets to `g^0`, so @@ -44,15 +44,21 @@ pub enum FillerOp { /// One compression of message and chaining-value cells nothing ever writes, its /// digest placed clear of them, so every traversal compresses the same input. Blake2s, + /// `XOR192` over the three-cell scratch run, pinning it to zero. + Xor192, + /// `MUL192` over the same run. + Mul192, } /// The tables, in `lean_vm::cpu::Stats::TABLES` order, which is how the solver indexes /// them. -pub const TABLES: [(u8, FillerOp); 6] = [ - (0, FillerOp::Xor), - (1, FillerOp::Mul), +pub const TABLES: [(u8, FillerOp); 8] = [ + (0, FillerOp::Xor64), + (1, FillerOp::Mul64), (2, FillerOp::Set), (3, FillerOp::Deref), (4, FillerOp::Jump), (5, FillerOp::Blake2s), + (6, FillerOp::Xor192), + (7, FillerOp::Mul192), ]; diff --git a/crates/lean_compiler/src/ir.rs b/crates/lean_compiler/src/ir.rs index 29d5a8ec3..089b81221 100644 --- a/crates/lean_compiler/src/ir.rs +++ b/crates/lean_compiler/src/ir.rs @@ -10,9 +10,8 @@ pub(crate) type Off = u32; pub(crate) enum KVal { /// The stride to the next frame in a reserved loop run. FrameSize, - /// A 192-bit machine-word constant. Source literals fill only c0/c1, while - /// compiler-generated constants may use the full field. - Const(F192), + /// A 64-bit machine-word constant. + Const(F64), Entry(String), /// The halt sentinel pc `g^{B-1}` (last bytecode slot), fixed once the /// padded bytecode size `B` is known. `main` jumps here to terminate. @@ -47,12 +46,23 @@ pub(crate) enum LOp { o: Off, k: KVal, }, - Xor { + Xor64 { a: Off, b: Off, c: Off, }, - Mul { + Mul64 { + a: Off, + b: Off, + c: Off, + }, + /// Operands and result are the first cells of three-cell runs. + Xor192 { + a: Off, + b: Off, + c: Off, + }, + Mul192 { a: Off, b: Off, c: Off, @@ -68,10 +78,9 @@ pub(crate) enum LOp { od: Off, of: Off, }, - /// `BLAKE2s`: the four 128-bit input chunks `ins` are addressed independently, - /// one frame cell each. The 32-byte output occupies the two consecutive - /// 128-bit cells `c, c+1`; `md` is the cell holding the byte counter and the - /// two flags. + /// `BLAKE2s`: the four 16-byte input chunks `ins` are addressed independently, + /// two frame cells each. The chaining value `cv` and the digest `c` span four + /// consecutive cells, the metadata `md` two. Blake2s { ins: [Off; 4], cv: Off, @@ -117,8 +126,8 @@ pub(crate) struct Lowered { /// A resolved run of consecutive cells ([`crate::lower::FnLower::cell_run`]): a /// frame (stack) run, used in place, or a heap slice (the buffer pointer's cell -/// plus the first g-power offset), which a `blake2s` operand must bridge through -/// the stack since `BLAKE2s` addresses only frame cells. +/// plus the first g-power offset), which an instruction operand must bridge +/// through the stack since operands address only frame cells. pub(crate) enum CellRun { Stack { base: Off, len: u32 }, Heap { ptr: Off, lo: u32, len: u32 }, diff --git a/crates/lean_compiler/src/lib.rs b/crates/lean_compiler/src/lib.rs index 4390d137f..4aa6dc598 100644 --- a/crates/lean_compiler/src/lib.rs +++ b/crates/lean_compiler/src/lib.rs @@ -79,7 +79,7 @@ fn compile_inner(ast: &Ast, with_filler: bool) -> Program { // Definitions by name, for Const-parameter specialization at call sites. let defs: HashMap<&str, &Func> = ast.funcs.iter().map(|f| (f.name.as_str(), f)).collect(); // Constant arrays by name, resolved at lowering (`NAME[i]`, `len(NAME)`). - let const_arrays: HashMap<&str, &[F192]> = ast + let const_arrays: HashMap<&str, &[F64]> = ast .const_arrays .iter() .map(|(name, values)| (name.as_str(), values.as_slice())) @@ -191,7 +191,7 @@ fn compile_inner(ast: &Ast, with_filler: bool) -> Program { } // Pad the bytecode to `B` (the sentinel slot g^{B-1} must exist for execution). - prog.resize(bytecode_size, Op::Set { o: 0, k: F192::ZERO }); + prog.resize(bytecode_size, Op::Set { o: 0, k: F64::ZERO }); let mut program = Program::assemble(prog, hints, frame_size["main"]); program.src_lines = src_lines; program.fn_ranges = lowered @@ -215,20 +215,20 @@ pub fn disassemble(prog: &[Op]) -> String { gmap.entry(acc).or_insert(j); acc *= primitives::field::G; } - // A machine word is 192-bit; K-valued immediates (both high limbs zero) may be small - // g-powers (code addresses, indices), shown as `gʲ`. - let kfmt = |k: F192| match (k.c1 == 0 && k.c2 == 0).then(|| gmap.get(&F64(k.c0))).flatten() { + // An immediate may be a small g-power (a code address, an index), shown as `gʲ`. + let kfmt = |k: F64| match gmap.get(&k) { Some(j) => format!("g^{j}"), - None if k.c1 == 0 && k.c2 == 0 => format!("0x{:016x}", k.c0), - None => format!("0x{:016x}{:016x}{:016x}", k.c2, k.c1, k.c0), + None => format!("0x{:016x}", k.0), }; let mut out = String::new(); for (pc, op) in prog.iter().enumerate() { let line = match op { Op::Set { o, k } => format!("SET fp[{o}] = {}", kfmt(*k)), - Op::Xor { a, b, c } => format!("XOR fp[{c}] = fp[{a}] ^ fp[{b}]"), - Op::Mul { a, b, c } => format!("MUL fp[{c}] = fp[{a}] * fp[{b}]"), + Op::Xor64 { a, b, c } => format!("XOR64 fp[{c}] = fp[{a}] ^ fp[{b}]"), + Op::Mul64 { a, b, c } => format!("MUL64 fp[{c}] = fp[{a}] * fp[{b}]"), + Op::Xor192 { a, b, c } => format!("XOR192 fp[{c}..] = fp[{a}..] ^ fp[{b}..]"), + Op::Mul192 { a, b, c } => format!("MUL192 fp[{c}..] = fp[{a}..] * fp[{b}..]"), Op::Deref { o1, o2, o3, mode } => { let src = match mode { DerefMode::Cell => format!("fp[{o3}]"), @@ -242,7 +242,7 @@ pub fn disassemble(prog: &[Op]) -> String { } Op::Blake2s { ins, cv, out, md } => { format!( - "BLAKE2S fp[{out}..]= compress(cv=fp[{cv}..], m=fp[{}],fp[{}],fp[{}],fp[{}], meta=fp[{md}])", + "BLAKE2S fp[{out}..] = compress(cv=fp[{cv}..], m=fp[{}..],fp[{}..],fp[{}..],fp[{}..], md=fp[{md}..])", ins[0], ins[1], ins[2], ins[3] ) } @@ -252,9 +252,9 @@ pub fn disassemble(prog: &[Op]) -> String { out } -/// Embed a `u128` source literal into the low 128 bits of a 192-bit machine word. -pub(crate) fn lit_field(n: u128) -> F192 { - F192::new(n as u64, (n >> 64) as u64, 0) +/// A source literal as the machine word with its bits, if it fits in one. +pub(crate) fn lit_field(n: u128) -> Option { + u64::try_from(n).ok().map(F64) } /// `g^e` for a `u128` exponent (square-and-multiply). `field::g_pow` only takes @@ -274,19 +274,17 @@ fn g_pow_u128(mut e: u128) -> F64 { } fn resolve(op: &LOp, entry: &HashMap, sentinel: u32, base: u32, frame_size: u32) -> Op { - let resolve_kval = |kv: &KVal| -> F192 { + let resolve_kval = |kv: &KVal| -> F64 { match kv { - KVal::FrameSize => g_pow(frame_size as usize).into(), + KVal::FrameSize => g_pow(frame_size as usize), KVal::Const(c) => *c, - // Address / entry / sentinel constants are K-valued g-powers; - // embed them canonically as (c0, 0, 0). - KVal::Entry(name) => g_pow(entry[name] as usize).into(), - KVal::EndSentinel => g_pow(sentinel as usize).into(), - KVal::Local(i) => g_pow((base + i) as usize).into(), + KVal::Entry(name) => g_pow(entry[name] as usize), + KVal::EndSentinel => g_pow(sentinel as usize), + KVal::Local(i) => g_pow((base + i) as usize), } }; match op { - LOp::MulNextFrame { a, b, c } => Op::Mul { + LOp::MulNextFrame { a, b, c } => Op::Mul64 { a: *a, b: *b, c: frame_size.checked_add(*c).expect("frame offset overflow"), @@ -295,8 +293,10 @@ fn resolve(op: &LOp, entry: &HashMap, sentinel: u32, base: u32, fra o: *o, k: resolve_kval(kv), }, - LOp::Xor { a, b, c } => Op::Xor { a: *a, b: *b, c: *c }, - LOp::Mul { a, b, c } => Op::Mul { a: *a, b: *b, c: *c }, + LOp::Xor64 { a, b, c } => Op::Xor64 { a: *a, b: *b, c: *c }, + LOp::Mul64 { a, b, c } => Op::Mul64 { a: *a, b: *b, c: *c }, + LOp::Xor192 { a, b, c } => Op::Xor192 { a: *a, b: *b, c: *c }, + LOp::Mul192 { a, b, c } => Op::Mul192 { a: *a, b: *b, c: *c }, LOp::Deref { o1, o2, o3, mode } => Op::Deref { o1: *o1, o2: *o2, diff --git a/crates/lean_compiler/src/lower.rs b/crates/lean_compiler/src/lower.rs index b6ee3e442..28581525a 100644 --- a/crates/lean_compiler/src/lower.rs +++ b/crates/lean_compiler/src/lower.rs @@ -15,6 +15,7 @@ //! arity, because they place the return area from their own idea of it. //! - [`mod@builtins`] is the precompile and the hints: the two places a value //! arrives without an instruction computing it. +//! - [`mod@run`] is values spanning several cells: 192-bit elements and digests. //! //! [`Scope`] is what a name means HERE, and it reverts at a branch join, so a //! cell whose `SET` sits inside a branch is never trusted outside it. @@ -27,6 +28,7 @@ mod builtins; mod call; mod eval; mod mem; +mod run; use call::ret_binding; use eval::field_pow; @@ -84,12 +86,14 @@ impl Abi { /// that folds an exponent into `β` measures it against this. const FOLD_MAX: u128 = 1 << lean_vm::cpu::MIN_LOG_MEM; -/// The two pure operations worth interning. Both are commutative, so operands -/// are stored sorted. +/// The pure operations worth interning. All are commutative, so operands are +/// stored sorted. A 192-bit one names runs by their first cell. #[derive(Clone, Copy, PartialEq, Eq, Hash)] enum PureOp { Xor, Mul, + Xor192, + Mul192, } /// How an inlined `@inline` tail-return value binds into the caller @@ -111,7 +115,9 @@ enum Binding { Scalar(Off), Stack(Off, u32), Gaddr(GAddr), - FConst(F192), + FConst(F64), + /// A 192-bit constant, materialized as a pooled run only where a use needs cells. + Const192(F192), } /// What one name means: its value binding, plus an OPTIONAL compile-time integer @@ -150,16 +156,16 @@ struct Scope { /// Dominating memory equalities. Reads reuse the frame cell; stores still /// emit their equality checks. Branch joins restore this with the scope. load_cells: HashMap<(Off, u32), Off>, - /// Every lazily-`SET` constant cell: field value (as bits) → the frame cell - /// holding it. Cells are write-once and read-many, so one `SET` serves every - /// use in scope. A `SET` first emitted inside a branch must not be named from - /// outside it, where the other path leaves the cell unwritten and therefore - /// prover-chosen, which is why this reverts at a join with the bindings. - const_cells: HashMap<[u64; 3], Off>, - /// Two consecutive frame cells holding the standard BLAKE2s IV, emitted - /// lazily at the first dominating default-IV compression in this - /// control-flow scope. - blake2s_iv: Option, + /// The same for a heap run read into the frame: `(ptr, offset, len)` → its run. + run_loads: HashMap<(Off, u32, u32), Off>, + /// Every lazily-`SET` constant cell: its word → the frame cell holding it. + /// Cells are write-once and read-many, so one `SET` serves every use in scope. + /// A `SET` first emitted inside a branch must not be named from outside it, + /// where the other path leaves the cell unwritten and therefore prover-chosen, + /// which is why this reverts at a join with the bindings. + const_cells: HashMap, + /// The same for constant runs ([`FnLower::const_run`]): the words → the run. + const_runs: HashMap, Off>, } impl Scope { @@ -232,7 +238,7 @@ struct FnLower<'a> { defs: &'a HashMap<&'a str, &'a Func>, /// Top-level constant arrays, resolved at compile time: `NAME[i]` yields the /// element (a field value or an index), `len(NAME)` its length. - const_arrays: &'a HashMap<&'a str, &'a [F192]>, + const_arrays: &'a HashMap<&'a str, &'a [F64]>, } impl FnLower<'_> { @@ -278,7 +284,7 @@ impl FnLower<'_> { self.emit(LOp::Set { o, k }); } - fn set_const(&mut self, o: Off, v: F192) { + fn set_const(&mut self, o: Off, v: F64) { self.set(o, KVal::Const(v)); } @@ -302,7 +308,7 @@ impl FnLower<'_> { /// branch join, or on a path this hint does not belong to). fn anchor(&mut self) { let o = self.fresh(); - self.set_const(o, F192::ZERO); + self.set_const(o, F64::ZERO); } /// A top-level constant name is reserved (`zkDSL.md` §Global constants). A @@ -354,11 +360,16 @@ impl FnLower<'_> { if let Some(&o) = self.scope.pure_cells.get(&key) { return o; } - let o = self.fresh(); - match op { - PureOp::Xor => self.emit(LOp::Xor { a, b, c: o }), - PureOp::Mul => self.emit(LOp::Mul { a, b, c: o }), - } + let o = match op { + PureOp::Xor | PureOp::Mul => self.fresh(), + PureOp::Xor192 | PureOp::Mul192 => self.alloc_stack(3), + }; + self.emit(match op { + PureOp::Xor => LOp::Xor64 { a, b, c: o }, + PureOp::Mul => LOp::Mul64 { a, b, c: o }, + PureOp::Xor192 => LOp::Xor192 { a, b, c: o }, + PureOp::Mul192 => LOp::Mul192 { a, b, c: o }, + }); self.scope.pure_cells.insert(key, o); o } @@ -380,7 +391,7 @@ impl FnLower<'_> { /// A frame cell holding `1` (always-taken `JUMP` condition). fn one(&mut self) -> Off { - self.const_cell(F192::ONE) + self.const_cell(F64::ONE) } /// A frame cell holding `v`, shared by every dominated use in the current @@ -394,22 +405,20 @@ impl FnLower<'_> { /// the cell unwritten and so prover-chosen. Several call sites hoist /// [`Self::one`] above a branch on purpose; the revert is what makes that an /// optimization rather than the thing holding the invariant up. - fn const_cell(&mut self, v: F192) -> Off { - let key = [v.c0, v.c1, v.c2]; - if let Some(&o) = self.scope.const_cells.get(&key) { + fn const_cell(&mut self, v: F64) -> Off { + if let Some(&o) = self.scope.const_cells.get(&v.0) { return o; } let o = self.fresh(); self.set_const(o, v); - self.scope.const_cells.insert(key, o); + self.scope.const_cells.insert(v.0, o); o } - /// A frame cell holding `0`, set lazily once: the source for forwarded zero - /// words (a `BLAKE2s` padding half), and the destination every `assert a == b` - /// in this scope XORs into. + /// A frame cell holding `0`, set lazily once: the destination every + /// `assert a == b` in this scope XORs into. fn zero(&mut self) -> Off { - self.const_cell(F192::ZERO) + self.const_cell(F64::ZERO) } /// Terminate `main`: jump to the halt sentinel `g^{B-1}` with `fp = g^0`. @@ -459,20 +468,31 @@ impl FnLower<'_> { table, }); for _ in 0..size { + let scratch = (fr::SCRATCH, fr::SCRATCH, fr::SCRATCH); self.emit(match op { - FillerOp::Xor => LOp::Xor { - a: fr::SCRATCH, - b: fr::SCRATCH, - c: fr::SCRATCH, + FillerOp::Xor64 => LOp::Xor64 { + a: scratch.0, + b: scratch.1, + c: scratch.2, }, - FillerOp::Mul => LOp::Mul { - a: fr::SCRATCH, - b: fr::SCRATCH, - c: fr::SCRATCH, + FillerOp::Mul64 => LOp::Mul64 { + a: scratch.0, + b: scratch.1, + c: scratch.2, + }, + FillerOp::Xor192 => LOp::Xor192 { + a: scratch.0, + b: scratch.1, + c: scratch.2, + }, + FillerOp::Mul192 => LOp::Mul192 { + a: scratch.0, + b: scratch.1, + c: scratch.2, }, FillerOp::Set => LOp::Set { o: fr::SCRATCH, - k: KVal::Const(F192::ZERO), + k: KVal::Const(F64::ZERO), }, FillerOp::Deref => LOp::Deref { o1: fr::PTR, @@ -493,7 +513,7 @@ impl FnLower<'_> { // prover choosing otherwise only picks which compression the // dummy proves, which nothing reads (`lean_vm::cpu::filler`). FillerOp::Blake2s => LOp::Blake2s { - ins: [fr::DIGEST + 2, fr::DIGEST + 3, fr::DIGEST + 4, fr::DIGEST + 5], + ins: [fr::DIGEST + 4, fr::DIGEST + 6, fr::DIGEST + 8, fr::DIGEST + 10], cv: fr::SCRATCH, c: fr::DIGEST, md: fr::ZERO, @@ -516,7 +536,7 @@ impl FnLower<'_> { /// `dst = src` (no MOV: multiply by `1`). fn copy(&mut self, src: Off, dst: Off) { let one = self.one(); - self.emit(LOp::Mul { a: src, b: one, c: dst }); + self.emit(LOp::Mul64 { a: src, b: one, c: dst }); } /// A frame cell holding this function's own `fp` (the g-power element), @@ -613,16 +633,6 @@ impl FnLower<'_> { /// `i = j` into cells every arm shares. Write-once makes that sound, exactly /// one arm running, and [`Self::ret_targets`] says which cells those are. fn lower_match(&mut self, targets: &[Expr], x: &Expr, arms: &[Expr]) { - for arm in arms { - if let Expr::Call(f, _) = arm - && self - .defs - .get(f.as_str()) - .is_some_and(|d| !d.inline && d.return_shapes.iter().any(|s| matches!(s, Shape::StackBuf(_)))) - { - self.fail("a normal function's StackBuf return cannot cross a match join; bind it with `let`"); - } - } // Calls with identical runtime args share one callee frame and a // two-instruction trampoline per arm. Const args select specializations; // see `lower_dispatched_call` for the shared argument/return layout checks. @@ -643,6 +653,19 @@ impl FnLower<'_> { // Not uniform: fall through (the specializations queued above are // re-requested idempotently by `call_into`). } + for arm in arms { + if let Expr::Call(f, _) = arm + && self + .defs + .get(f.as_str()) + .is_some_and(|d| !d.inline && d.return_shapes.iter().any(|s| matches!(s, Shape::StackBuf(_)))) + { + self.fail( + "a normal function's StackBuf return crosses a match join only when every arm calls it \ + with the same runtime arguments; bind it with `let` otherwise", + ); + } + } let xo = self.expr(x); let (rcells, binds) = self.ret_targets(targets); self.lower_match_dispatch(xo, arms.len(), |s, j| { @@ -696,7 +719,7 @@ impl FnLower<'_> { self.set(kcell, KVal::Local(0)); // patched: table base T let x2 = self.pure(PureOp::Mul, xo, xo); let d = self.fresh(); - self.emit(LOp::Mul { a: kcell, b: x2, c: d }); + self.emit(LOp::Mul64 { a: kcell, b: x2, c: d }); self.emit(LOp::Jump { oc: one, od: d, of }); kset } @@ -770,10 +793,10 @@ impl FnLower<'_> { if let Some((n, f)) = known.diverging_readings() { self.fail(format!( "`{e:?}` reads as the integer {n} where a condition folds, and as the field \ - element {:#x}:{:#x} where a value is wanted, so this branch would be decided \ + element {:#x} where a value is wanted, so this branch would be decided \ by one reading and its body run under the other. Write `if const(...)` to \ decide it with integer arithmetic, or spell the operand so the two agree.", - f.c1, f.c0 + f.0 )) } } @@ -858,8 +881,8 @@ impl FnLower<'_> { let inv = self.fresh(); self.pending.push(Hint::Resolved(RHint::Inverse { value: x, dst: inv })); let p = self.fresh(); - self.emit(LOp::Mul { a: x, b: inv, c: p }); - self.set_const(p, F192::ONE); + self.emit(LOp::Mul64 { a: x, b: inv, c: p }); + self.set_const(p, F64::ONE); } /// The frame cell holding `g^{k-1}`, the range-check product target, shared @@ -871,7 +894,7 @@ impl FnLower<'_> { /// the `SET` precedes every use and each later write is the write-once /// equality, but a second WRITER on any of those paths would land on all. fn bound_cell(&mut self, k: u64) -> Off { - self.const_cell(g_pow_u128((k - 1) as u128).into()) + self.const_cell(g_pow_u128((k - 1) as u128)) } /// `assert log x < log GEN ** k`: the 3-cycle range check in the exponent @@ -920,9 +943,9 @@ impl FnLower<'_> { )) }; let bcell = self.expr(b); - let inv = self.const_cell(F192::new(primitives::field::G.inv().0, 0, 0)); + let inv = self.const_cell(primitives::field::G.inv()); let c = self.fresh(); - self.emit(LOp::Mul { a: bcell, b: inv, c }); + self.emit(LOp::Mul64 { a: bcell, b: inv, c }); c } }; @@ -931,7 +954,7 @@ impl FnLower<'_> { let t1 = self.fresh(); // DEREF targets: unconstrained touch cells let t2 = self.fresh(); self.deref(x, 0, t1, DerefMode::Cell); - self.emit(LOp::Mul { a: x, b: y, c: kcell }); + self.emit(LOp::Mul64 { a: x, b: y, c: kcell }); self.deref(y, 0, t2, DerefMode::Cell); } @@ -942,18 +965,20 @@ impl FnLower<'_> { return self.const_cell(v); } match e { - Expr::Lit(_) | Expr::Gen | Expr::GPow(_) => unreachable!("a literal folds above"), + Expr::Lit(n) => self.fail(format!("the literal {n} does not fit in a 64-bit word")), + Expr::Gen | Expr::GPow(_) => unreachable!("a g-power folds above"), // Not folded above, so its exponent is not a compile-time integer, // which `gpow_exp` reports. Expr::GenPow(e) => { let k = self.gpow_exp(e); - self.const_cell(g_pow_u128(k).into()) + self.const_cell(g_pow_u128(k)) } Expr::Pow(b, e) => self.pow_expr(b, e), Expr::Var(v) => match self.scope.bound(v).map(|b| b.val) { Some(Binding::Stack(..)) => { self.fail(format!("StackBuf `{v}` used as a scalar; index it (`{v}[k]`) or pass it to blake2s")); } + Some(Binding::Const192(_)) => self.fail(format!("the 192-bit constant `{v}` used as a scalar")), Some(Binding::Gaddr(ga)) => self.materialize(ga), Some(Binding::Scalar(o)) => o, _ => { @@ -997,11 +1022,13 @@ impl FnLower<'_> { // `a` must already be written, which `self.expr(a)` guarantees. let (la, lb) = (self.expr(a), self.expr(b)); let q = self.fresh(); - self.emit(LOp::Mul { a: q, b: lb, c: la }); + self.emit(LOp::Mul64 { a: q, b: lb, c: la }); q } - // A well-formed one folds above, so this is a malformed call. - Expr::Call(f, _) if f == "f192" => self.fail("f192 needs three literal u64 limbs"), + Expr::Call(f, _) if parser::RUN192_BUILTINS.contains(&f.as_str()) => self.fail(format!( + "`{f}(...)` is a 192-bit value, a run of three cells, not a scalar: bind it or pass it \ + where a 192-bit value is expected" + )), // Folded above when it is what it claims to be, so reaching here means // it is not: name that, rather than reporting an unknown function. Expr::Call(f, args) if f == "const" => { @@ -1083,7 +1110,7 @@ impl FnLower<'_> { "`-`, `//`, `%` are compile-time only (field subtraction is `+`); use them in an index, a bound, or a `Const` argument, got `{e:?}`" )) } - Expr::Slice(..) => self.fail("a slice is not a scalar; it is only a blake2s operand"), + Expr::Slice(..) => self.fail("a slice is a run of cells, not a scalar"), Expr::ListLit(..) => self.fail("a list literal must be bound to a name: `x = [a, b]`"), } } @@ -1100,7 +1127,7 @@ impl FnLower<'_> { } if k == 0 { let o = self.fresh(); - self.set_const(o, F192::ONE); + self.set_const(o, F64::ONE); return o; } // Runtime base: square-and-multiply over the compile-time exponent bits. @@ -1149,7 +1176,7 @@ impl FnLower<'_> { // Not `pure`: `dst` is the caller's, so it may be written // again and the second write be the assertion. See its doc. let (la, lb) = (self.expr(a), self.expr(b)); - self.emit(LOp::Xor { a: la, b: lb, c: dst }); + self.emit(LOp::Xor64 { a: la, b: lb, c: dst }); } } Expr::Mul(a, b) => { @@ -1157,7 +1184,7 @@ impl FnLower<'_> { self.expr_into(x, dst); } else { let (la, lb) = (self.expr(a), self.expr(b)); - self.emit(LOp::Mul { a: la, b: lb, c: dst }); + self.emit(LOp::Mul64 { a: la, b: lb, c: dst }); } } // A call writes its single return value straight into `dst` (an @@ -1189,16 +1216,32 @@ impl FnLower<'_> { let base = self.alloc_stack(*n as u32); self.rebind(name, Binding::Stack(base, *n as u32)); } + // `x = f192(..)`, or a 192-bit builtin over constants: folded, no cells. + _ if !matches!(e, Expr::ListLit(_)) && self.const192(e).is_some() => { + let c = self.const192(e).expect("guarded above"); + self.rebind(name, Binding::Const192(c)); + } // `x = [a, b, …]`: an initialized StackBuf. Allocate the run and - // write each element in place, through the ordinary stack-store - // path. Elements are lowered before `name` rebinds, so they may - // read its old binding (`fs = [fs[1], fs[0]]`). - Expr::ListLit(es) => { - let base = self.alloc_stack(es.len() as u32); - for (k, el) in es.iter().enumerate() { - self.expr_into(el, base + k as u32); - } - self.rebind(name, Binding::Stack(base, es.len() as u32)); + // write each element in place, a run element over its cells. + // Elements are lowered before `name` rebinds, so they may read its + // old binding (`fs = [fs[1], fs[0]]`). Fresh cells even for constant + // elements, so a later store into `x` stores rather than asserts. + _ if matches!(e, Expr::ListLit(_)) => { + let len = self.run_len(e).expect("a run"); + let base = self.alloc_stack(len); + self.run_into(e, base, len); + self.rebind(name, Binding::Stack(base, len)); + } + // `x = add192(a, b)` and the other 192-bit builtins, and `x = buf[lo:hi]`: + // the value's run, a heap slice copied onto the stack. + _ if matches!(e, Expr::Slice(..)) + || matches!(e, Expr::Call(f, _) if parser::RUN192_BUILTINS.contains(&f.as_str())) => + { + let len = self + .run_len(e) + .unwrap_or_else(|| self.fail(format!("a slice binding needs a known length, got `{e:?}`"))); + let base = self.run(e, len); + self.rebind(name, Binding::Stack(base, len)); } // `p = addr(sb)` binds the address itself, so the offset folds // into every later access; in any other position `expr` has to @@ -1232,9 +1275,12 @@ impl FnLower<'_> { self.rebind(name, Binding::FConst(c)); } else if let Some(k) = known.int { // Integer-only fold (`//`, `-`, `%` of constants): a - // compile-time value too, and as a scalar it is the field - // element with those 128 bits, materialized on demand. - self.rebind(name, Binding::FConst(lit_field(k))); + // compile-time value too, and as a scalar it is the word + // with those bits, materialized on demand. + let word = lit_field(k).unwrap_or_else(|| { + self.fail(format!("the compile-time integer {k} does not fit in a 64-bit word")) + }); + self.rebind(name, Binding::FConst(word)); } else if let Expr::Call(cf, cargs) = e && self.defs.contains_key(cf.as_str()) { @@ -1279,7 +1325,7 @@ impl FnLower<'_> { StmtKind::AssertEq(a, b) => { let (la, lb) = (self.expr(a), self.expr(b)); let z = self.zero(); - self.emit(LOp::Xor { a: la, b: lb, c: z }); + self.emit(LOp::Xor64 { a: la, b: lb, c: z }); } StmtKind::AssertNe(a, b) => self.lower_assert_ne(a, b), StmtKind::AssertLt(e, bound) => self.lower_assert_lt(e, bound), @@ -1335,6 +1381,7 @@ impl FnLower<'_> { self.deref(ptr, o2, v, DerefMode::Cell); } } + StmtKind::StoreRun(target, val) => self.lower_store_run(target, val), StmtKind::Return(es) => self.lower_return(es), StmtKind::CallIfNe(lhs, rhs, callee, args) => { // A conditional call: the frame setup runs either way, and the @@ -1575,7 +1622,7 @@ pub(crate) fn lower_func( queue: &mut Vec, loop_ctr: &mut usize, defs: &HashMap<&str, &Func>, - const_arrays: &HashMap<&str, &[F192]>, + const_arrays: &HashMap<&str, &[F64]>, with_filler: bool, loop_bounds: &mut HashMap)>, ) -> Lowered { @@ -1611,7 +1658,13 @@ pub(crate) fn lower_func( names, ..Default::default() }, - next: abi_end + u32::from(loop_frame), + // `main`'s frame is memory from cell 0, whose first four cells hold the + // public input. + next: if f.name == "main" { + abi_end.max(4) + } else { + abi_end + u32::from(loop_frame) + }, arg_cells, return_shapes: &f.return_shapes, is_main: f.name == "main", diff --git a/crates/lean_compiler/src/lower/builtins.rs b/crates/lean_compiler/src/lower/builtins.rs index 68e39b635..52c4ef1e0 100644 --- a/crates/lean_compiler/src/lower/builtins.rs +++ b/crates/lean_compiler/src/lower/builtins.rs @@ -2,7 +2,7 @@ //! instruction computing it. //! //! `blake2s` is a STATEMENT, not an expression: it writes its digest into a -//! two-cell run the caller names, so a pre-written destination checks the digest +//! four-cell run the caller names, so a pre-written destination checks the digest //! instead of computing it, by the same write-once rule as any store. //! //! A hint writes values the prover chose and the circuit did not, so **the @@ -13,15 +13,24 @@ use super::*; +/// One word of a `blake2s` message operand: a compile-time constant, a cell, or +/// the list element at this index, evaluated only once its chunk has cells. +#[derive(Clone, Copy)] +enum Word { + Const(F64), + Cell(Off), + Expr(usize), +} + impl FnLower<'_> { /// `blake2s(a, b, out)`: the digest of the two 256-bit operands lands in the - /// existing 2-cell run `out` (write-once: if `out` was already written, this + /// existing 4-cell run `out` (write-once: if `out` was already written, this /// asserts the digest equals it). A heap `out` slice takes the digest via a - /// fresh stack pair and two `DEREF`s after the hash, the store direction + /// fresh stack run and four `DEREF`s after the hash, the store direction /// being the same instruction as the load (write-once fills the unset side). /// Keyword arguments set the metadata: `counter=` / `final=` / `last_node=` - /// build it at compile time, `md=` takes the whole word from a value the - /// program computed. + /// build it at compile time, `md=` takes the two cells from values the program + /// computed. fn lower_blake2s(&mut self, args: &[Expr]) { let first_kw = args .iter() @@ -53,11 +62,11 @@ impl FnLower<'_> { self.fail(format!("unknown blake2s keyword {bad:?}; the keywords are {allowed:?}")) }; let customized = kwargs.keys().any(|k| matches!(*k, "counter" | "final" | "last_node")); - // `md=` hands over the whole metadata word as a runtime value, so it - // replaces the three keywords that would otherwise build it. + // `md=` hands over the whole metadata as runtime values, so it replaces the + // three keywords that would otherwise build it. let runtime_md = kwargs.get("md").copied(); if runtime_md.is_some() && customized { - self.fail("blake2s md= is the whole metadata word, so counter=, final= and last_node= cannot come with it") + self.fail("blake2s md= is the whole metadata, so counter=, final= and last_node= cannot come with it") }; if kwargs.contains_key("cv") && !customized && runtime_md.is_none() { self.fail( @@ -66,31 +75,32 @@ impl FnLower<'_> { ) }; - let a = self.blake2s_input(&args[0]); - let b = self.blake2s_input(&args[1]); - let (c, heap_out) = match self.blake2s_operand(&args[2]) { + let a = self.blake2s_chunks(&args[0]); + let b = self.blake2s_chunks(&args[1]); + let out = self.cell_run(&args[2]); + if out.cells() != 4 { + self.fail("a blake2s digest destination spans exactly 4 cells; slice a larger buffer: `buf[lo:lo + 4]`") + } + let (c, heap_out) = match out { CellRun::Stack { base, .. } => (base, None), - CellRun::Heap { ptr, lo, .. } => (self.alloc_stack(2), Some((ptr, lo))), + CellRun::Heap { ptr, lo, .. } => (self.alloc_stack(4), Some((ptr, lo))), }; - let cv = if let Some(value) = kwargs.get("cv") { - self.blake2s_cv(value) - } else { - self.default_blake2s_cv() + let cv = match kwargs.get("cv") { + Some(value) => self.run(value, 4), + None => self.const_run(&lean_vm::hash_flock::IV), }; let md = match runtime_md { - // A metadata word the program computes, which is what lets a hash whose - // block count is only known at run time carry the byte counter the - // standard asks for (doc §sec:prog-byte-counter). It owes the same - // canonical embedding as every other operand: the memory interaction - // carries a literal zero above its two low limbs. + // Metadata the program computes, which is what lets a hash whose block + // count is only known at run time carry the byte counter the standard + // asks for (doc §sec:prog-byte-counter). // // Aliasing the digest destination is the one case write-once does not // catch: the runner reads the metadata before storing the digest, while // the witness reads the finished memory image, so the two disagree and // the proof fails its opening rather than saying why. Some(expr) => { - let md = self.expr(expr); - if md == c || md == c + 1 { + let md = self.run(expr, 2); + if md + 2 > c && md < c + 4 { self.fail("blake2s md= must not name a cell of the digest destination") }; md @@ -103,7 +113,7 @@ impl FnLower<'_> { this.try_const_int(e).unwrap_or_else(|| { self.fail(format!( "BLAKE2s `{name}` must be a compile-time integer, got `{e:?}`; \ - a metadata word computed at run time goes through md=" + a metadata computed at run time goes through md=" )) }) }) @@ -127,15 +137,11 @@ impl FnLower<'_> { } else { 0 }; - // A compile-time metadata is a pooled `SET`: one per distinct value + // A compile-time metadata is a pooled run: one per distinct value // per frame, however many compressions read it. - self.const_cell(lean_vm::hash_flock::metadata(counter, f0, f1)) + self.const_run(&lean_vm::hash_flock::metadata(counter, f0, f1)) } }; - // Each operand is two 128-bit chunk cells; the flexible opcode addresses - // the four input cells independently (`blake2s_input` forwards the real - // chunk sources where it can). The digest occupies the two consecutive - // output cells `c, g·c`. self.emit(LOp::Blake2s { ins: [a[0], a[1], b[0], b[1]], cv, @@ -143,7 +149,7 @@ impl FnLower<'_> { md, }); if let Some((ptr, lo)) = heap_out { - for k in 0..2 { + for k in 0..4 { self.deref(ptr, lo + k, c + k, DerefMode::Cell); } } @@ -154,7 +160,7 @@ impl FnLower<'_> { /// advice, re-checked in-circuit by their caller: `hint_decompose_bits` /// writes a value's bits into a buffer, `hint_decompose_bits_exponent` the /// bits of `n` where the value is `g^n` (a bounded dlog at witness - /// generation), `hint_f192_limbs` a value's coordinate limbs. + /// generation). pub(super) fn lower_builtin(&mut self, f: &str, args: &[Expr]) -> bool { match f { "hint_decompose_bits" | "hint_decompose_bits_exponent" => { @@ -165,6 +171,11 @@ impl FnLower<'_> { )) }; let nbits = self.const_index(&args[2]); + if f == "hint_decompose_bits" && nbits > 64 { + self.fail(format!( + "hint_decompose_bits of a 64-bit word takes at most 64 bits, got {nbits}" + )) + } let bits = self.bits_dest(&args[0], nbits, f); let value = self.expr(&args[1]); self.pending.push(Hint::Resolved(if f == "hint_decompose_bits" { @@ -174,102 +185,77 @@ impl FnLower<'_> { })); } "blake2s" => self.lower_blake2s(args), - "assert_in_k" => { - if args.len() != 2 { - self.fail("assert_in_k(a, b) takes two scalar cells") - }; - let a = self.expr(&args[0]); - let b = self.expr(&args[1]); - let zero = self.zero(); - self.emit(LOp::Jump { oc: zero, od: a, of: b }); - } - "hint_f192_limbs" => { - if args.len() != 2 { - self.fail(format!( - "hint_f192_limbs takes two arguments, `(dest, value)`, got {}", - args.len() - )) - }; - let (base, len) = self.stack_of(&args[0]).unwrap_or_else(|| { - self.fail(format!( - "hint_f192_limbs writes 1..=3 frame cells, so its destination must be a \ - StackBuf, got `{:?}`", - args[0] - )) - }); - if !((1..=3).contains(&len)) { - self.fail("hint_f192_limbs destination must have 1..=3 cells") + "assert_eq192" | "assert_ne192" => { + let [a, b] = args else { + self.fail(format!("{f}(a, b) takes two 192-bit values")) }; - let value = self.expr(&args[1]); - // Names the physical cells, as the two consumers above do: whatever - // the program stores into them afterwards is a second write, and so - // the assertion that pins these limbs. - self.pending - .push(Hint::Resolved(RHint::FieldLimbs { value, base, len })); + if f == "assert_eq192" { + self.lower_assert_eq192(a, b); + } else { + self.lower_assert_ne192(a, b); + } } + _ if parser::RUN192_BUILTINS.contains(&f) => self.fail(format!( + "`{f}(...)` is a 192-bit value, not a statement: bind it (`x = {f}(...)`) or store it" + )), _ => return false, } true } - /// Resolve a `blake2s` operand: a [`Self::cell_run`] pinned to exactly 2 - /// cells, a 256-bit value being two 128-bit cells. Stack operands are used - /// in place; heap operands must be bridged through the stack, since - /// `BLAKE2s` addresses only frame cells (see [`Self::blake2s_input`]). - fn blake2s_operand(&mut self, e: &Expr) -> CellRun { - let run = self.cell_run(e); - if run.cells() != 2 { - self.fail("a blake2s operand must span exactly 2 cells (two 128-bit words); slice a larger buffer: `buf[lo:lo + 2]`") + /// A `blake2s` message operand as its two chunk bases, each two consecutive + /// cells. The opcode addresses the chunks independently, so a LIST LITERAL + /// operand gathers nothing it does not have to: a chunk whose two words already + /// sit side by side is used in place, a constant one is a pooled run, and any + /// other chunk is a fresh pair. A computed word is evaluated straight into its + /// pair, so only a word that already lives in a cell is copied. + pub(super) fn blake2s_chunks(&mut self, e: &Expr) -> [Off; 2] { + let Expr::ListLit(elems) = e else { + let base = self.run(e, 4); + return [base, base + 2]; }; - run - } - - /// A `blake2s` *input* operand as its two independently-addressed 128-bit - /// chunk bases (each chunk is ONE 128-bit cell): stack runs in place; a heap - /// slice is pulled into a fresh stack pair first, one `DEREF` per cell - /// (`m[ptr·g^{lo+k}] == m[fp+t+k]`, the `β` immediate doing the pointer - /// offset). The heap cells must already be written. - /// - /// A LIST LITERAL names its two words directly and allocates nothing. The - /// opcode addresses its four input chunks independently, so an operand - /// assembled out of values living elsewhere never has to be gathered into a - /// consecutive run: `blake2s([a, b], …)` is the spelling that says so. - pub(super) fn blake2s_input(&mut self, e: &Expr) -> [Off; 2] { - if let Expr::ListLit(words) = e { - if words.len() != 2 { - self.fail(format!( - "a blake2s operand written as a list needs exactly 2 words, got {}", - words.len() - )) - }; - return [self.expr(&words[0]), self.expr(&words[1])]; - } - match self.blake2s_operand(e) { - CellRun::Stack { base, .. } => [base, base + 1], - CellRun::Heap { ptr, lo, .. } => { - let t = self.alloc_stack(2); - for k in 0..2 { - self.deref(ptr, lo + k, t + k, DerefMode::Cell); + let mut words: Vec = Vec::new(); + for (i, el) in elems.iter().enumerate() { + match self.run_len(el) { + Some(w) => { + let base = self.run(el, w); + words.extend((base..base + w).map(Word::Cell)); } - [t, t + 1] + None => words.push(match (self.try_field_const(el), el) { + (Some(v), _) => Word::Const(v), + (None, Expr::Var(_)) => Word::Cell(self.expr(el)), + (None, Expr::Index(arr, idx)) if self.stack_of(arr).is_some() => { + Word::Cell(self.frame_cell(arr, idx).expect("a StackBuf element")) + } + (None, _) => Word::Expr(i), + }), } } - } - - fn default_blake2s_cv(&mut self) -> Off { - if let Some(o) = self.scope.blake2s_iv { - return o; + if words.len() != 4 { + self.fail(format!( + "a blake2s message operand is 4 words, this list spells {}", + words.len() + )) } - let o = self.alloc_stack(2); - for (k, value) in lean_vm::hash_flock::IV_CELLS.into_iter().enumerate() { - self.set_const(o + k as u32, value); - self.scope - .const_cells - .entry([value.c0, value.c1, value.c2]) - .or_insert(o + k as u32); + let mut chunks = [0; 2]; + for (i, chunk) in chunks.iter_mut().enumerate() { + *chunk = match (words[2 * i], words[2 * i + 1]) { + (Word::Const(x), Word::Const(y)) => self.const_run(&[x, y]), + (Word::Cell(x), Word::Cell(y)) if y == x + 1 => x, + (x, y) => { + let t = self.alloc_stack(2); + for (k, w) in [x, y].into_iter().enumerate() { + match w { + Word::Const(v) => self.set_const(t + k as u32, v), + Word::Cell(c) => self.copy(c, t + k as u32), + Word::Expr(i) => self.expr_into(&elems[i], t + k as u32), + } + } + t + } + }; } - self.scope.blake2s_iv = Some(o); - o + chunks } /// A computed-advice bit buffer's destination ([`BitsDest`]). Not @@ -283,7 +269,6 @@ impl FnLower<'_> { "{what} needs {nbits} cells, its StackBuf destination has {len}" )) }; - // The hint names the physical cells, as `hint_f192_limbs` does. BitsDest::Stack(base) } None => { @@ -309,20 +294,4 @@ impl FnLower<'_> { }; self.pending.push(Hint::Resolved(hint)); } - /// A BLAKE2s chaining value must occupy two consecutive frame cells because - /// the opcode carries one base offset for both words. Preserve a genuine - /// consecutive pair, including a heap pair already bridged by - /// [`Self::blake2s_input`]. A `cv` written as a two-word LIST exposes two - /// sources that need not be adjacent, so those are copied into a fresh - /// consecutive pair. - fn blake2s_cv(&mut self, e: &Expr) -> Off { - let pair = self.blake2s_input(e); - if pair[1] == pair[0] + 1 { - return pair[0]; - } - let cv = self.alloc_stack(2); - self.copy(pair[0], cv); - self.copy(pair[1], cv + 1); - cv - } } diff --git a/crates/lean_compiler/src/lower/call.rs b/crates/lean_compiler/src/lower/call.rs index 1ff6de773..b37288df4 100644 --- a/crates/lean_compiler/src/lower/call.rs +++ b/crates/lean_compiler/src/lower/call.rs @@ -43,6 +43,7 @@ fn stmt_inline_safe(s: &Stmt) -> bool { match &s.kind { StmtKind::Let(..) | StmtKind::Store(..) + | StmtKind::StoreRun(..) | StmtKind::HintWitness { .. } | StmtKind::LetHintWitness { .. } | StmtKind::Print { .. } @@ -74,12 +75,12 @@ impl FnLower<'_> { } // Nothing by that name is going to be lowered, so the entry pc it // needs will not exist. Caught here, where there is a line: a typo, a - // statement-only builtin used as a value (`x = assert_in_k(a, b)`), + // statement-only builtin used as a value (`x = assert_eq192(a, b)`), // or an `@inline` callee reached where inlining did not happen, all // used to die later in `resolve` as a bare `no entry found for key`. None => self.fail(format!( "no function named `{callee}`. A builtin that writes into a destination \ - (`blake2s`, `assert_in_k`, a `hint_*`) is a statement and returns nothing, so it \ + (`blake2s`, `assert_eq192`, a `hint_*`) is a statement and returns nothing, so it \ cannot be called for a value" )), _ => {} @@ -98,16 +99,16 @@ impl FnLower<'_> { let base = Abi::arg(shapes.iter().copied(), i); match shapes.get(i).copied().unwrap_or(Shape::Scalar) { Shape::StackBuf(n) => { - let (src, len) = self.stack_of(a).unwrap_or_else(|| { - self.fail(format!( - "`{callee}` parameter {i} is a StackBuf({n}); pass one, got `{a:?}`" - )) - }); - if len != n { - self.fail(format!( - "`{callee}` parameter {i} is a StackBuf({n}), got a StackBuf({len})" - )) + match self.run_len(a) { + Some(len) if len == n => {} + Some(len) => self.fail(format!( + "`{callee}` parameter {i} is a StackBuf({n}), got a {len}-cell value" + )), + None => self.fail(format!( + "`{callee}` parameter {i} is a StackBuf({n}); pass a {n}-cell value, got `{a:?}`" + )), } + let src = self.run(a, n); for k in 0..n { let cell = src + k; arg_offs.push((base + k, cell)); @@ -209,31 +210,27 @@ impl FnLower<'_> { /// Each taken arm is then just the trampoline's `SET entry; JUMP`: no /// per-arm frame setup, call, or return jump. pub(super) fn lower_dispatched_call(&mut self, targets: &[Expr], x: &Expr, callees: &[String], rt_args: &[&Expr]) { - // The arms share ONE frame, so they must share one argument layout too: - // a `StackBuf` parameter in one callee and a scalar in another at the - // same position would put the return area in two places. The arity check - // below is the count; this is the widths. + // The arms share ONE frame, so they must share one argument and return layout + // too: a `StackBuf` in one callee and a scalar in another at the same position + // would put the return area, or the value read back, in two places. The arity + // checks below are the counts; this is the widths. + let mut param_shapes = vec![Shape::Scalar; rt_args.len()]; + let mut ret_shapes = vec![Shape::Scalar; targets.len()]; if let Some(shared) = callees.iter().find_map(|c| self.callee_def(c)) { for c in callees { if let Some(def) = self.callee_def(c) - && !def.param_shapes().eq(shared.param_shapes()) + && (!def.param_shapes().eq(shared.param_shapes()) || def.return_shapes != shared.return_shapes) { self.fail(format!( - "`{c}` does not take the same parameter shapes as the other arms of this dispatch" + "`{c}` does not take and return the same shapes as the other arms of this dispatch" )) } } - // Fused dispatch writes one cell per argument; only scalar parameters fit. - if let Some(i) = shared.param_shapes().position(|s| s != Shape::Scalar) { - self.fail(format!( - "a `match` arm cannot pass a `StackBuf` parameter (parameter {i} of `{}`): the \ - fused dispatch writes one cell per argument. Give the arms `Const` arguments so each \ - specializes into its own call instead of fusing", - callees.first().map(String::as_str).unwrap_or("?") - )) + if shared.params.len() == rt_args.len() && shared.return_shapes.len() == targets.len() { + param_shapes = shared.param_shapes().collect(); + ret_shapes = shared.return_shapes.clone(); } } - let n_args = rt_args.len() as u32; // The join below reads one return cell per bound name, so every callee has // to declare exactly that many. Unchecked, a name past a callee's arity // `DEREF`s a frame offset nothing on that path writes, and since the shared @@ -277,17 +274,39 @@ impl FnLower<'_> { targets.len() )) }; - if shapes.iter().any(|s| *s != Shape::Scalar) { - self.fail(format!( - "`{callee}`: a multi-cell StackBuf return cannot cross a dispatched join" - )) - }; } - let (rcells, binds) = self.ret_targets(targets); + // A run return binds a name to fresh cells; a scalar one may also land in a + // frame element. + let mut rcells = Vec::with_capacity(targets.len()); + let mut binds = Vec::new(); + for (t, &shape) in targets.iter().zip(&ret_shapes) { + match (t, shape) { + (Expr::Var(n), Shape::StackBuf(w)) => { + let base = self.alloc_stack(w); + rcells.push(base); + binds.push((n.as_str(), Binding::Stack(base, w))); + } + (_, Shape::Scalar) => { + let (cells, names) = self.ret_targets(std::slice::from_ref(t)); + rcells.push(cells[0]); + binds.extend(names.into_iter().map(|(n, c)| (n, Binding::Scalar(c)))); + } + _ => self.fail(format!( + "a StackBuf return of a dispatched call binds a name, got `{t:?}`" + )), + } + } // Shared callee frame: args, retfp, and retpc = the join (so the callee // returns straight past the dispatch). Evaluated once. - let arg_offs: Vec = rt_args.iter().map(|a| self.expr(a)).collect(); + let arg_vals: Vec<(Off, u32)> = rt_args + .iter() + .zip(¶m_shapes) + .map(|(a, &shape)| match shape { + Shape::Scalar => (self.expr(a), 1), + Shape::StackBuf(n) => (self.run(a, n), n), + }) + .collect(); let xo = self.expr(x); let one = self.one(); let sfp = self.self_fp(); @@ -297,13 +316,11 @@ impl FnLower<'_> { ptr: nfp, callees: callees.to_vec(), }); - for (i, &ao) in arg_offs.iter().enumerate() { - self.deref( - nfp, - Abi::arg(std::iter::repeat_n(Shape::Scalar, rt_args.len()), i), - ao, - DerefMode::Cell, - ); + for (i, &(src, width)) in arg_vals.iter().enumerate() { + let base = Abi::arg(param_shapes.iter().copied(), i); + for k in 0..width { + self.deref(nfp, base + k, src + k, DerefMode::Cell); + } } self.deref(nfp, Abi::RET_FP, 0, DerefMode::Fp); let join_cell = self.fresh(); @@ -320,11 +337,17 @@ impl FnLower<'_> { // Join: read the return values (written by whichever callee ran). self.patch_local(join_set, self.code.len()); - for (i, &r) in rcells.iter().enumerate() { - self.deref(nfp, Abi::ret(n_args, i as u32), r, DerefMode::Cell); + let arg_cells = Abi::arg_cells(param_shapes.iter().copied()); + let mut offset = 0; + for (&r, shape) in rcells.iter().zip(&ret_shapes) { + for k in 0..shape.cells() { + self.deref(nfp, Abi::ret(arg_cells, offset + k), r + k, DerefMode::Cell); + } + offset += shape.cells(); + } + for (name, binding) in binds { + self.rebind(name, binding); } - - self.bind_targets(&binds); } /// Inline an `@inline` `callee(args)` into the current frame, binding its @@ -367,8 +390,8 @@ impl FnLower<'_> { if p.kind == ParamKind::Const { continue; } - let b = if let Some((base, size)) = self.stack_of(a) { - Binding::Stack(base, size) + let b = if let Some(len) = self.run_len(a) { + Binding::Stack(self.run(a, len), len) } else if let Some(ga) = self.gaddr_of(a) { Binding::Gaddr(ga) } else { @@ -416,8 +439,8 @@ impl FnLower<'_> { // advanced cursor together. let mut binds = Vec::with_capacity(dsts.len()); for (e, &d) in exprs.iter().zip(&dsts) { - binds.push(if let Some((base, size)) = self.stack_of(e) { - RetBind::Stack(base, size) + binds.push(if let Some(len) = self.run_len(e) { + RetBind::Stack(self.run(e, len), len) } else if let Some(ga) = self.gaddr_of(e) { RetBind::Gaddr(ga) } else { @@ -446,18 +469,11 @@ impl FnLower<'_> { for (e, &shape) in exprs.iter().zip(self.return_shapes) { match shape { Shape::Scalar => self.expr_into(e, ret), - Shape::StackBuf(size) => { - let (base, actual) = self - .stack_of(e) - .unwrap_or_else(|| self.fail(format!("expected a StackBuf({size}) return, got `{e:?}`"))); - if actual != size { - self.fail(format!("returned StackBuf has size {actual}, expected {size}")) - }; - for k in 0..size { - let src = base + k; - self.copy(src, ret + k); - } - } + Shape::StackBuf(size) => match self.run_len(e) { + Some(actual) if actual == size => self.run_into(e, ret, size), + Some(actual) => self.fail(format!("returned value has {actual} cells, expected {size}")), + None => self.fail(format!("expected a StackBuf({size}) return, got `{e:?}`")), + }, } ret += shape.cells(); } @@ -508,7 +524,7 @@ impl FnLower<'_> { /// describes those logical bindings to the surrounding let/tuple lowering. pub(super) fn call(&mut self, callee: &str, args: &[Expr], n_ret: usize) -> Vec { if callee == "blake2s" { - self.fail("blake2s is a statement: `blake2s(a, b, out)` writes the digest into the 2-cell stack run `out`") + self.fail("blake2s is a statement: `blake2s(a, b, out)` writes the digest into the 4-cell run `out`") }; self.inline_stack_ret = None; if self.defs.get(callee).is_some_and(|d| d.inline) { @@ -597,7 +613,7 @@ impl FnLower<'_> { } /// Original definitions take precedence over generated functions in the queue. - fn callee_def(&self, callee: &str) -> Option<&Func> { + pub(super) fn callee_def(&self, callee: &str) -> Option<&Func> { self.defs .get(callee) .copied() diff --git a/crates/lean_compiler/src/lower/eval.rs b/crates/lean_compiler/src/lower/eval.rs index bd9cd6222..967c83aab 100644 --- a/crates/lean_compiler/src/lower/eval.rs +++ b/crates/lean_compiler/src/lower/eval.rs @@ -17,16 +17,16 @@ pub(super) struct Known { /// The compile-time INTEGER, wanted by a size, an index, a bound, an exponent. pub(super) int: Option, /// The FIELD element a value position sees, where `+` is XOR. - pub(super) field: Option, + pub(super) field: Option, /// The ADDRESS the compiler tracks: a base cell times `g^exp`. pub(super) addr: Option, } impl Known { /// The integer and field readings when both exist and disagree. - pub(super) fn diverging_readings(self) -> Option<(u128, F192)> { + pub(super) fn diverging_readings(self) -> Option<(u128, F64)> { let (n, f) = (self.int?, self.field?); - (f != lit_field(n)).then_some((n, f)) + (Some(f) != lit_field(n)).then_some((n, f)) } } @@ -47,8 +47,8 @@ fn gmul(a: GAddr, b: GAddr) -> Option { } /// `b^k` by square-and-multiply, in logarithmically many field operations. -pub(super) fn field_pow(b: F192, mut k: u32) -> F192 { - let (mut acc, mut sq) = (F192::ONE, b); +pub(super) fn field_pow(b: F64, mut k: u32) -> F64 { + let (mut acc, mut sq) = (F64::ONE, b); while k > 0 { if k & 1 == 1 { acc *= sq; @@ -82,12 +82,12 @@ impl FnLower<'_> { // guard exists to stop. let int = |n: u128| Known { int: Some(n), - field: Some(lit_field(n)), + field: lit_field(n), addr: None, }; let gpow = |exp: u128| Known { int: None, - field: Some(g_pow_u128(exp).into()), + field: Some(g_pow_u128(exp)), addr: Some(GAddr { base: None, exp, @@ -96,10 +96,8 @@ impl FnLower<'_> { }; match e { // `g = x`, so the literal `2^k` IS `g^k`, but ONLY while `k < 64`: at - // and above it the modulus folds the monomial back into the low limb - // while the literal's bit `k` lands in the next limb, the tower - // coefficient of `y`. Without the guard the guest's own `Y_TOWER = - // 2^64` would read as `g^64` in a pointer position. + // and above it the modulus folds the monomial back into the word while + // the literal no longer fits in one. Expr::Lit(n) => Known { addr: (n.is_power_of_two() && *n < (1 << 64)).then(|| GAddr { base: None, @@ -128,7 +126,7 @@ impl FnLower<'_> { Binding::Gaddr(ga) => Known { int, // A constant g-power also reads as that field element. - field: (ga.base.is_none()).then(|| g_pow_u128(ga.exp).into()), + field: (ga.base.is_none()).then(|| g_pow_u128(ga.exp)), addr: Some(ga), }, // A plain scalar is its own base, unshifted. @@ -141,7 +139,7 @@ impl FnLower<'_> { run: None, }), }, - Binding::Stack(..) => Known { + Binding::Stack(..) | Binding::Const192(_) => Known { int, ..Known::default() }, @@ -195,22 +193,12 @@ impl FnLower<'_> { // bits spell. Expr::Index(..) => match self.const_array_elem(e) { Some(v) => Known { - int: (v.c2 == 0).then_some(v.c0 as u128 | ((v.c1 as u128) << 64)), + int: Some(u128::from(v.0)), field: Some(v), addr: None, }, None => Known::default(), }, - Expr::Call(f, args) if f == "f192" && args.len() == 3 => { - let limb = |i: usize| match &args[i] { - Expr::Lit(n) => u64::try_from(*n).ok(), - _ => None, - }; - Known { - field: (|| Some(F192::new(limb(0)?, limb(1)?, limb(2)?)))(), - ..Known::default() - } - } // `const(e)`: the one construct that asks for the INTEGER reading in a // position that would otherwise take the field one. It reinterprets the // OPERATORS, so its leaves must mean the same thing either way. @@ -243,7 +231,7 @@ impl FnLower<'_> { /// `e` as a compile-time FIELD constant, where `+` is XOR. `None` for a /// runtime value or for arithmetic the field has no meaning for (`-`, `//`, /// `%`). - pub(super) fn try_field_const(&self, e: &Expr) -> Option { + pub(super) fn try_field_const(&self, e: &Expr) -> Option { self.eval(e).field } @@ -272,8 +260,8 @@ impl FnLower<'_> { } /// If `e` is `NAME[i]` for a top-level constant array `NAME` with a - /// compile-time index `i`, its element (a raw `u128`). - fn const_array_elem(&self, e: &Expr) -> Option { + /// compile-time index `i`, its element. + fn const_array_elem(&self, e: &Expr) -> Option { if let Expr::Index(arr, idx) = e && let Expr::Var(v) = arr.as_ref() && let Some(a) = self.const_arrays.get(v.as_str()) @@ -305,20 +293,20 @@ impl FnLower<'_> { /// preserve. So `x + 0` lowers to just `x`: no cell, no XOR. Kills the /// `acc = 0; acc = acc + t` accumulator seed and similar. pub(super) fn add_identity<'e>(&self, a: &'e Expr, b: &'e Expr) -> Option<&'e Expr> { - if self.try_field_const(a) == Some(F192::ZERO) { + if self.try_field_const(a) == Some(F64::ZERO) { return Some(b); } - (self.try_field_const(b) == Some(F192::ZERO)).then_some(a) + (self.try_field_const(b) == Some(F64::ZERO)).then_some(a) } /// The surviving operand of `a * b` when the other is a compile-time one, a /// no-op multiply. Kills the `acc = GEN ** 0` (= 1) accumulator seed's first /// `1 * f` in every product loop. pub(super) fn mul_identity<'e>(&self, a: &'e Expr, b: &'e Expr) -> Option<&'e Expr> { - if self.try_field_const(a) == Some(F192::ONE) { + if self.try_field_const(a) == Some(F64::ONE) { return Some(b); } - (self.try_field_const(b) == Some(F192::ONE)).then_some(a) + (self.try_field_const(b) == Some(F64::ONE)).then_some(a) } /// The field value of `e` when it is a trivial compile-time constant (a @@ -360,9 +348,9 @@ impl FnLower<'_> { if let Some((n, f)) = self.eval(leaf).diverging_readings() { self.fail(format!( "const(...) reads its operators as integer arithmetic, but it cannot reinterpret \ - `{leaf:?}`, which is the integer {n} and the value {:#x}:{:#x}: two different \ + `{leaf:?}`, which is the integer {n} and the value {:#x}: two different \ numbers. Bind it in one regime and name that one", - f.c1, f.c0 + f.0 )) } } @@ -390,6 +378,6 @@ impl FnLower<'_> { return None; } let j = n.trailing_zeros(); - (u128::from(j) <= FOLD_MAX && k.field? == g_pow_u128(u128::from(j)).into()).then_some(j) + (u128::from(j) <= FOLD_MAX && k.field? == g_pow_u128(u128::from(j))).then_some(j) } } diff --git a/crates/lean_compiler/src/lower/mem.rs b/crates/lean_compiler/src/lower/mem.rs index a2f8875f3..fe0c7c35d 100644 --- a/crates/lean_compiler/src/lower/mem.rs +++ b/crates/lean_compiler/src/lower/mem.rs @@ -112,9 +112,9 @@ impl FnLower<'_> { same number as cell n. Write `buf[GEN ** {k}]` and say which you mean" ), None => format!( - "folds to the field constant {:#x}:{:#x}, which is not a g-power, so it names \ + "folds to the field constant {:#x}, which is not a g-power, so it names \ no heap cell (did an integer index leak in from a StackBuf conversion?)", - c.c1, c.c0 + c.0 ), }; self.fail(format!("heap index {why}")); @@ -234,7 +234,7 @@ impl FnLower<'_> { if extra == 0 { return (a, 0); } - let k = self.const_cell(g_pow_u128(extra).into()); + let k = self.const_cell(g_pow_u128(extra)); (self.pure(PureOp::Mul, a, k), 0) } @@ -269,7 +269,7 @@ impl FnLower<'_> { base: Some(c), exp: 0, .. } => c, GAddr { base, exp, .. } => { - let k = self.const_cell(g_pow_u128(exp).into()); + let k = self.const_cell(g_pow_u128(exp)); let Some(c) = base else { return k }; self.pure(PureOp::Mul, c, k) } diff --git a/crates/lean_compiler/src/lower/run.rs b/crates/lean_compiler/src/lower/run.rs new file mode 100644 index 000000000..af1fa2190 --- /dev/null +++ b/crates/lean_compiler/src/lower/run.rs @@ -0,0 +1,331 @@ +//! Run values: consecutive cells read and written as one value. A 192-bit element +//! is a run of three cells, its limbs low first; a BLAKE2s digest is a run of four. +//! +//! An instruction reads only frame cells, so a heap slice used as a value is copied +//! onto the stack first, one `DEREF` a cell. A copy between frame runs is an +//! `XOR192` against a pooled zero run, one instruction per three cells. + +use super::*; + +impl FnLower<'_> { + /// How many cells the value `e` spans, or `None` for a scalar. Emits nothing. + pub(super) fn run_len(&self, e: &Expr) -> Option { + match e { + Expr::Var(v) => match self.scope.bound(v)?.val { + Binding::Stack(_, n) => Some(n), + Binding::Const192(_) => Some(3), + _ => None, + }, + Expr::Slice(_, lo, hi) => match (self.try_const_index(lo), self.try_const_index(hi)) { + (Some(lo), Some(hi)) if lo < hi => Some(hi - lo), + _ => plus_k(lo, hi).and_then(|k| u32::try_from(k).ok()), + }, + Expr::ListLit(elems) => Some(elems.iter().map(|el| self.run_len(el).unwrap_or(1)).sum()), + Expr::Call(f, _) if parser::RUN192_BUILTINS.contains(&f.as_str()) => Some(3), + Expr::Call(f, _) => match self.callee_def(f)?.return_shapes.as_slice() { + [Shape::StackBuf(n)] => Some(*n), + _ => None, + }, + _ => None, + } + } + + fn expect_run(&self, e: &Expr, len: u32) { + match self.run_len(e) { + Some(n) if n == len => {} + Some(n) => self.fail(format!("expected a {len}-cell value, got a {n}-cell one: `{e:?}`")), + None => self.fail(format!("expected a {len}-cell value, got the scalar `{e:?}`")), + } + } + + /// The 192-bit constant `e` is, when it is one: an `f192` literal, a name bound + /// to one, a list of three constant words, or a 192-bit builtin over constants. + pub(super) fn const192(&self, e: &Expr) -> Option { + match e { + Expr::Var(v) => match self.scope.bound(v)?.val { + Binding::Const192(c) => Some(c), + _ => None, + }, + Expr::Call(f, args) if f == "f192" && args.len() == 3 => { + let limb = |a: &Expr| self.try_const_int(a).and_then(lit_field); + Some(F192::new(limb(&args[0])?.0, limb(&args[1])?.0, limb(&args[2])?.0)) + } + Expr::ListLit(elems) if elems.len() == 3 && elems.iter().all(|el| self.run_len(el).is_none()) => { + let w = |i: usize| self.try_field_const(&elems[i]); + Some(F192::new(w(0)?.0, w(1)?.0, w(2)?.0)) + } + Expr::Call(f, args) if matches!(f.as_str(), "add192" | "mul192" | "div192") && args.len() == 2 => { + // A product with a zero factor is zero whatever the other one is. + if f == "mul192" && args.iter().any(|a| self.const192(a) == Some(F192::ZERO)) { + return Some(F192::ZERO); + } + let (a, b) = (self.const192(&args[0])?, self.const192(&args[1])?); + match f.as_str() { + "add192" => Some(a + b), + "mul192" => Some(a * b), + _ => (!b.is_zero()).then(|| a * b.inv()), + } + } + _ => None, + } + } + + /// The operand a 192-bit builtin reduces to without an instruction: `a + 0`, + /// `a · 1` and `a / 1` are `a`. + fn identity192<'e>(&self, f: &str, args: &'e [Expr]) -> Option<&'e Expr> { + let [a, b] = args else { return None }; + let (ca, cb) = (self.const192(a), self.const192(b)); + match f { + "add192" if ca == Some(F192::ZERO) => Some(b), + "add192" if cb == Some(F192::ZERO) => Some(a), + "mul192" if ca == Some(F192::ONE) => Some(b), + "mul192" | "div192" if cb == Some(F192::ONE) => Some(a), + _ => None, + } + } + + /// The value `e` as `len` consecutive frame cells, returning the first. A run + /// already in the frame is used in place; anything else lands in fresh cells. + pub(super) fn run(&mut self, e: &Expr, len: u32) -> Off { + self.expect_run(e, len); + if let Some(c) = self.const192(e) { + return self.const_run(&[F64(c.c0), F64(c.c1), F64(c.c2)]); + } + if let Expr::Call(f, args) = e + && let Some(x) = self.identity192(f, args) + { + return self.run(x, len); + } + match e { + Expr::Var(_) => self.stack_of(e).expect("a run name").0, + Expr::Slice(..) => match self.cell_run(e) { + CellRun::Stack { base, .. } => base, + CellRun::Heap { ptr, lo, len } => { + if let Some(&t) = self.scope.run_loads.get(&(ptr, lo, len)) { + // The dominating DEREFs may have deferred equalities between + // unwritten cells. + for k in 0..len { + self.pending.push(Hint::Resolved(RHint::ResolveDeref { + ptr, + offset: lo + k, + dst: t + k, + })); + } + return t; + } + let t = self.alloc_stack(len); + for k in 0..len { + self.deref(ptr, lo + k, t + k, DerefMode::Cell); + } + self.scope.run_loads.insert((ptr, lo, len), t); + t + } + }, + Expr::ListLit(elems) => { + if let Some(words) = elems + .iter() + .map(|el| self.try_field_const(el)) + .collect::>>() + { + return self.const_run(&words); + } + let t = self.alloc_stack(len); + self.run_into(e, t, len); + t + } + Expr::Call(f, args) if f == "f192" => { + let words = self.f192_words(args); + self.const_run(&words) + } + Expr::Call(f, args) if f == "add192" || f == "mul192" => { + let (a, b) = self.run_operands(f, args); + let op = if f == "add192" { PureOp::Xor192 } else { PureOp::Mul192 }; + self.pure(op, a, b) + } + Expr::Call(f, _) if f == "div192" => { + let t = self.alloc_stack(3); + self.run_into(e, t, 3); + t + } + Expr::Call(f, args) => { + let dst = self.call(f, args, 1)[0]; + match self.inline_stack_ret.take().and_then(|b| b.into_iter().next()) { + Some(RetBind::Stack(base, _)) => base, + _ => dst, + } + } + _ => unreachable!("`run_len` names every run shape"), + } + } + + /// Evaluate the value `e` into the `len` cells at `dst`. A cell already written + /// makes its write the equality assertion, as for any store. + pub(super) fn run_into(&mut self, e: &Expr, dst: Off, len: u32) { + self.expect_run(e, len); + if let Expr::Call(f, args) = e + && self.const192(e).is_none() + && let Some(x) = self.identity192(f, args) + { + return self.run_into(x, dst, len); + } + match e { + Expr::Call(f, args) if f == "add192" || f == "mul192" => { + let (a, b) = self.run_operands(f, args); + if f == "add192" { + self.emit(LOp::Xor192 { a, b, c: dst }); + } else { + self.emit(LOp::Mul192 { a, b, c: dst }); + } + } + // The quotient is the product's unwritten operand: witness generation + // back-solves it and the `MUL192` pins `q·b == a`. + Expr::Call(f, args) if f == "div192" => { + let (a, b) = self.run_operands(f, args); + self.emit(LOp::Mul192 { a: dst, b, c: a }); + } + Expr::Call(f, args) if f == "f192" => { + for (k, w) in self.f192_words(args).into_iter().enumerate() { + self.set_const(dst + k as u32, w); + } + } + Expr::ListLit(elems) => { + let mut off = 0; + for el in elems { + match self.run_len(el) { + Some(w) => { + self.run_into(el, dst + off, w); + off += w; + } + None => { + self.expr_into(el, dst + off); + off += 1; + } + } + } + } + // A heap run `DEREF`s straight into `dst`: either side may be the unset one. + Expr::Slice(arr, ..) if self.stack_of(arr).is_none() => { + let CellRun::Heap { ptr, lo, .. } = self.cell_run(e) else { + unreachable!("a heap slice") + }; + for k in 0..len { + self.deref(ptr, lo + k, dst + k, DerefMode::Cell); + } + } + _ => { + let src = self.run(e, len); + self.copy_run(src, dst, len); + } + } + } + + /// `dst = src` over `len` cells: `XOR192` against the pooled zero run three + /// cells at a time, the rest cell by cell. + pub(super) fn copy_run(&mut self, src: Off, dst: Off, len: u32) { + if src == dst { + return; + } + let mut k = 0; + if len >= 3 { + let zero = self.const_run(&[F64::ZERO; 3]); + while k + 3 <= len { + self.emit(LOp::Xor192 { + a: src + k, + b: zero, + c: dst + k, + }); + k += 3; + } + } + for k in k..len { + self.copy(src + k, dst + k); + } + } + + /// A run holding `words`, shared by every dominated use in scope. Like + /// [`Self::const_cell`] its `SET`s precede every use and it reverts at a branch + /// join with the rest of [`Scope`]. + pub(super) fn const_run(&mut self, words: &[F64]) -> Off { + let key: Vec = words.iter().map(|w| w.0).collect(); + if let Some(&o) = self.scope.const_runs.get(&key) { + return o; + } + let o = self.alloc_stack(words.len() as u32); + for (k, &w) in words.iter().enumerate() { + self.set_const(o + k as u32, w); + self.scope.const_cells.entry(w.0).or_insert(o + k as u32); + } + self.scope.const_runs.insert(key, o); + o + } + + /// The three limbs of `f192(c0, c1, c2)`, each a compile-time integer. + fn f192_words(&self, args: &[Expr]) -> Vec { + if args.len() != 3 { + self.fail("f192(c0, c1, c2) takes three limbs") + } + args.iter() + .map(|a| { + self.try_const_int(a) + .and_then(lit_field) + .unwrap_or_else(|| self.fail(format!("an f192 limb is a compile-time 64-bit integer, got `{a:?}`"))) + }) + .collect() + } + + /// The two three-cell operands of a 192-bit builtin, left first. + fn run_operands(&mut self, f: &str, args: &[Expr]) -> (Off, Off) { + let [a, b] = args else { + self.fail(format!("{f}(a, b) takes two 192-bit values")) + }; + (self.run(a, 3), self.run(b, 3)) + } + + /// `assert_eq192(a, b)`: `XOR192` into the pooled zero run, whose double write is + /// the assertion, as for `assert a == b`. + pub(super) fn lower_assert_eq192(&mut self, a: &Expr, b: &Expr) { + let (la, lb) = (self.run(a, 3), self.run(b, 3)); + let zero = self.const_run(&[F64::ZERO; 3]); + self.emit(LOp::Xor192 { a: la, b: lb, c: zero }); + } + + /// `assert_ne192(a, b)`: `x = a + b`, a hinted `inv = x⁻¹`, and `MUL192 x·inv` + /// into the pooled one run. Sound as for `assert a != b`: `x = 0` makes the + /// product zero whatever the hint. + pub(super) fn lower_assert_ne192(&mut self, a: &Expr, b: &Expr) { + let (la, lb) = (self.run(a, 3), self.run(b, 3)); + let x = self.pure(PureOp::Xor192, la, lb); + let one = self.const_run(&[F64::ONE, F64::ZERO, F64::ZERO]); + let inv = self.alloc_stack(3); + self.pending + .push(Hint::Resolved(RHint::Inverse192 { value: x, dst: inv })); + self.emit(LOp::Mul192 { a: x, b: inv, c: one }); + } + + /// `buf[lo:hi] = value`: a frame run takes the value in place; a heap run takes + /// it through one `DEREF` a cell, the value lowered first as for a scalar store. + pub(super) fn lower_store_run(&mut self, target: &Expr, value: &Expr) { + let Expr::Slice(arr, ..) = target else { + unreachable!("the parser builds a run store from a slice") + }; + if self.stack_of(arr).is_some() { + let CellRun::Stack { base, len } = self.cell_run(target) else { + unreachable!("a frame slice") + }; + self.run_into(value, base, len); + return; + } + let len = self + .run_len(target) + .unwrap_or_else(|| self.fail(format!("a run store needs a slice of known length, got `{target:?}`"))); + let v = self.run(value, len); + let CellRun::Heap { ptr, lo, .. } = self.cell_run(target) else { + unreachable!("a heap slice") + }; + for k in 0..len { + self.deref(ptr, lo + k, v + k, DerefMode::Cell); + } + // The heap run now equals `v`, so a dominated read of it reuses `v`. + self.scope.run_loads.entry((ptr, lo, len)).or_insert(v); + } +} diff --git a/crates/lean_compiler/src/parser.rs b/crates/lean_compiler/src/parser.rs index 71eddc029..8a3bc8792 100644 --- a/crates/lean_compiler/src/parser.rs +++ b/crates/lean_compiler/src/parser.rs @@ -92,7 +92,7 @@ pub fn parse_with_replacements(src: &str, replacements: &BTreeMap = BTreeMap::new(); - let mut const_arrays: Vec<(String, Vec)> = Vec::new(); + let mut const_arrays: Vec<(String, Vec)> = Vec::new(); let mut start = 0; while start < lines.len() { let Line { @@ -145,38 +145,31 @@ pub fn parse_with_replacements(src: &str, replacements: &BTreeMap> 64) as u64, 0) - }; + let n = eval_const_int(p).map_err(|e| at(format!("global constant array `{name}`: {e}")))?; + let elem = lit_field(n).ok_or_else(|| { + at(format!( + "global constant array `{name}`: {n} does not fit in a 64-bit word" + )) + })?; elems.push(elem); } const_arrays.push((name, elems)); } else { - // A scalar constant: an `f192` literal, else a compile-time integer, - // else a field-valued expression. + // A scalar constant: an `f192` literal (a 192-bit run constant), else a + // compile-time integer, else a field-valued expression. if let Some(value) = parse_f192_const(rhs) { - let v = value.map_err(|e| at(format!("global constant `{name}`: {e}")))?; - consts.insert(name, format!("f192({},{},{})", v.c0, v.c1, v.c2)); + let [c0, c1, c2] = value.map_err(|e| at(format!("global constant `{name}`: {e}")))?; + consts.insert(name, format!("f192({c0},{c1},{c2})")); } else if let Ok(value) = eval_const_int(rhs) { consts.insert(name, value.to_string()); } else { // `GEN ** 2` and friends. The ISA is written in g-powers, so this // is the natural spelling for a constant one, and it is not an - // integer expression. Rendered as a decimal wherever the value - // fits the low two limbs, so the constant still works in the - // positions that demand a literal rather than only as a value. + // integer expression. Rendered as a decimal, so the constant still + // works in the positions that demand a literal rather than only as a + // value. let v = parse_const(rhs).map_err(|e| at(format!("global constant `{name}`: {e}")))?; - consts.insert( - name, - if v.c2 == 0 { - (v.c0 as u128 | ((v.c1 as u128) << 64)).to_string() - } else { - format!("f192({},{},{})", v.c0, v.c1, v.c2) - }, - ); + consts.insert(name, v.0.to_string()); } } start += 1; @@ -235,22 +228,44 @@ pub fn parse_with_replacements(src: &str, replacements: &BTreeMap Option { + let k = match (const_int_expr(lo), const_int_expr(hi)) { + (Some(lo), Some(hi)) => hi.checked_sub(lo)?, + _ => match hi { + Expr::Add(a, b) => match (a.as_ref(), b.as_ref()) { + (Expr::Lit(k), other) | (other, Expr::Lit(k)) if other == lo => *k, + _ => return None, + }, + _ => return None, + }, + }; + u32::try_from(k).ok() +} + /// Infer the compile-time representation of each tail-return value: a `StackBuf` /// carries its static size, a `HeapBuf` stays a one-cell pointer (its allocation /// hint ran in the creating function). Iterated to a fixed point, so a wrapper @@ -265,7 +280,18 @@ fn infer_return_shapes(funcs: &mut [Func]) -> Result<(), String> { Ok(match e { Expr::Var(v) => locals.get(v.as_str()).copied().unwrap_or(Shape::Scalar), Expr::StackBuf(n) => Shape::StackBuf(fits(*n)?), - Expr::ListLit(es) => Shape::StackBuf(fits(es.len() as u64)?), + Expr::ListLit(es) => { + let mut cells = 0u64; + for el in es { + cells += match expr_shape(el, locals, known)? { + Shape::StackBuf(n) => u64::from(n), + Shape::Scalar => 1, + }; + } + Shape::StackBuf(fits(cells)?) + } + Expr::Slice(_, lo, hi) => slice_len(lo, hi).map_or(Shape::Scalar, Shape::StackBuf), + Expr::Call(f, _) if RUN192_BUILTINS.contains(&f.as_str()) => Shape::StackBuf(3), Expr::Call(f, _) => known .get(f) .filter(|r| r.len() == 1) @@ -315,6 +341,19 @@ fn infer_return_shapes(funcs: &mut [Func]) -> Result<(), String> { StmtKind::LetHintWitness { name, .. } => { locals.insert(name.as_str(), Shape::Scalar); } + // Every arm returns the same shapes, so the first arm's callee names them. + StmtKind::Match { targets, arms, .. } => { + let shapes = match arms.first() { + Some(Expr::Call(f, _)) => known.get(f), + _ => None, + }; + for (i, target) in targets.iter().enumerate() { + if let Expr::Var(name) = target { + let shape = shapes.and_then(|s| s.get(i)).copied().unwrap_or(Shape::Scalar); + locals.insert(name.as_str(), shape); + } + } + } StmtKind::Return(es) => { returns = es .iter() @@ -739,13 +778,14 @@ impl Parser<'_> { }); } let rhs_expr = parse_expr(rhs)?; - // Indexed LHS `arr[idx] = value` is a heap store. + // Indexed LHS `arr[idx] = value` is a store, and a sliced one + // `arr[lo:hi] = value` a run store. if lhs.trim_end().ends_with(']') { - let lhs = lhs.trim(); - let open = lhs.find('[').ok_or("malformed store target")?; - let arr = parse_expr(&lhs[..open])?; - let idx = parse_expr(&lhs[open + 1..lhs.len() - 1])?; - return Ok(StmtKind::Store(arr, idx, rhs_expr)); + return match parse_expr(lhs)? { + Expr::Index(arr, idx) => Ok(StmtKind::Store(*arr, *idx, rhs_expr)), + target @ Expr::Slice(..) => Ok(StmtKind::StoreRun(target, rhs_expr)), + other => Err(format!("malformed store target `{other:?}`")), + }; } let targets = split_top(lhs, ','); if targets.len() == 1 { diff --git a/crates/lean_compiler/src/parser/consts.rs b/crates/lean_compiler/src/parser/consts.rs index ebaa3d046..b7d243ada 100644 --- a/crates/lean_compiler/src/parser/consts.rs +++ b/crates/lean_compiler/src/parser/consts.rs @@ -63,16 +63,15 @@ pub(super) fn const_int_expr(e: &Expr) -> Option { } /// Evaluate a compile-time constant expression (integer literals, `GEN`, -/// `GEN ** k`, and `+`/`*` combinations of those) to its field element. -/// Used for the `# public_input: , ` annotation of `.py` test +/// `GEN ** k`, and `+`/`*` combinations of those) to its 64-bit word. +/// Used for the `# public_input:` and `# witness` annotations of `.py` test /// programs (see `tests/py_source.rs`). -pub fn parse_const(s: &str) -> Result { - fn eval(e: &Expr) -> Result { +pub fn parse_const(s: &str) -> Result { + fn eval(e: &Expr) -> Result { match e { - // An integer literal is the raw 128-bit bit pattern of a machine word. - Expr::Lit(n) => Ok(F192::new(*n as u64, (*n >> 64) as u64, 0)), - Expr::Gen => Ok(g_pow(1).into()), - Expr::GPow(k) => Ok(g_pow_u128(*k).into()), + Expr::Lit(n) => lit_field(*n).ok_or_else(|| format!("literal {n} does not fit in a 64-bit word")), + Expr::Gen => Ok(g_pow(1)), + Expr::GPow(k) => Ok(g_pow_u128(*k)), Expr::Add(a, b) => Ok(eval(a)? + eval(b)?), Expr::Mul(a, b) => Ok(eval(a)? * eval(b)?), other => Err(format!("not a constant expression: `{other:?}`")), @@ -81,7 +80,8 @@ pub fn parse_const(s: &str) -> Result { eval(&parse_expr(s)?) } -pub(super) fn parse_f192_const(s: &str) -> Option> { +/// The three limbs of an `f192(c0, c1, c2)` literal, each a compile-time integer. +pub(super) fn parse_f192_const(s: &str) -> Option> { let inner = s.trim().strip_prefix("f192(")?.strip_suffix(')')?; let parts = split_top(inner, ','); Some((|| { @@ -93,7 +93,7 @@ pub(super) fn parse_f192_const(s: &str) -> Option> { limbs[i] = u64::try_from(eval_const_int(p.trim())?).map_err(|_| "an f192 limb does not fit in u64".to_string())?; } - Ok(F192::new(limbs[0], limbs[1], limbs[2])) + Ok(limbs) })()) } diff --git a/crates/lean_compiler/src/parser/subst.rs b/crates/lean_compiler/src/parser/subst.rs index 0affd3ed9..898d48a53 100644 --- a/crates/lean_compiler/src/parser/subst.rs +++ b/crates/lean_compiler/src/parser/subst.rs @@ -78,6 +78,7 @@ fn subst_kind(s: &StmtKind, name: &str, to: &Expr) -> (StmtKind, bool) { false, ), StmtKind::Store(a, i, v) => (StmtKind::Store(e(a), e(i), e(v)), false), + StmtKind::StoreRun(target, v) => (StmtKind::StoreRun(e(target), e(v)), false), StmtKind::Return(es) => (StmtKind::Return(es.iter().map(e).collect()), false), StmtKind::CallIfNe(a, b, f, args) => ( StmtKind::CallIfNe(e(a), e(b), f.clone(), args.iter().map(e).collect()), diff --git a/crates/lean_compiler/tests/programs/const_params.py b/crates/lean_compiler/tests/programs/const_params.py index 81a1e833c..959005f49 100644 --- a/crates/lean_compiler/tests/programs/const_params.py +++ b/crates/lean_compiler/tests/programs/const_params.py @@ -2,29 +2,28 @@ # site passes a compile-time constant and gets a monomorphized copy with `k` # substituted as the integer literal, usable in compile-time positions (the # slice bounds below). The direct call and match arm 0 share the k=0 -# specialization. A 256-bit BLAKE2s value occupies two canonical cells. -# Published: the two 128-bit digest cells of H(quad0, quad0) XOR H(quad1, quad1) -#: the direct k=0 digest XORed with the arm the runtime x = GEN selects (k=1). -# public_input: 252517949230448393340326710819579834691, 263897057969456650752475895236275386743 +# specialization. A 256-bit BLAKE2s value occupies four cells. +# Published: the four digest words of H(quad0, quad0) XOR H(quad1, quad1): the +# direct k=0 digest XORed with the arm the runtime x = GEN selects (k=1). +# public_input: 10463089001937318211, 13689025457361823030, 10489568288789412215, 14305888178150900108 from snark_lib import * def main(): - buf = HeapBuf(4) - buf[1] = 5 - buf[GEN] = 7 - buf[GEN ** 2] = 11 - buf[GEN ** 3] = 13 - a0, a1 = hash_pair(buf, 0) + buf = HeapBuf(8) + buf[0:8] = [5, 0, 7, 0, 11, 0, 13, 0] + a0, a1, a2, a3 = hash_pair(buf, 0) x = GEN - b0, b1 = match(log(x), range(0, 2), lambda i: hash_pair(buf, i)) + b0, b1, b2, b3 = match(log(x), range(0, 2), lambda i: hash_pair(buf, i)) p = GEN ** 0 p[1] = a0 + b0 p[GEN] = a1 + b1 + p[GEN ** 2] = a2 + b2 + p[GEN ** 3] = a3 + b3 return def hash_pair(buf, k: Const): - h = StackBuf(2) - blake2s(buf[k * 2:k * 2 + 2], buf[k * 2:k * 2 + 2], h) - return h[0], h[1] + h = StackBuf(4) + blake2s(buf[k * 4:k * 4 + 4], buf[k * 4:k * 4 + 4], h) + return h[0], h[1], h[2], h[3] diff --git a/crates/lean_compiler/tests/programs/field192.py b/crates/lean_compiler/tests/programs/field192.py new file mode 100644 index 000000000..360fa3fe0 --- /dev/null +++ b/crates/lean_compiler/tests/programs/field192.py @@ -0,0 +1,32 @@ +# 192-bit arithmetic on three-cell runs: `add192`, `mul192` and `div192` compute in +# GF(2^192) = K[y]/(y^3 + y + 1), limbs low first. A value lives in a StackBuf(3), +# a slice, a list literal or an `f192` constant, crosses a call as a StackBuf(3) +# parameter or return, and reaches the heap through a slice store. The quotient +# is back-solved at witness generation and pinned by the product. +# Published: the limbs of q = (a·b + c) / a, then the low limb of a·b. +# public_input: 5664748180115531082, 13, 4212248646752574772, 1544 +from snark_lib import * + + +def main(): + a = f192(3, 5, 7) + b = [GEN ** 9, 11, 13] + c = StackBuf(3) + c[0] = 17 + c[1] = 19 + c[2] = 23 + heap = HeapBuf(6) + heap[0:3] = mul192(a, b) + q = div192(add192(heap[0:3], c), a) + assert_eq192(mul192(q, a), add192(mul192(a, b), c)) + assert_ne192(q, c) + assert_eq192(square(q), mul192(q, q)) + heap[3:6] = q + p = GEN ** 0 + p[0:3] = heap[3:6] + p[GEN ** 3] = heap[GEN ** 0] + return + + +def square(x: StackBuf(3)): + return mul192(x, x) diff --git a/crates/lean_compiler/tests/programs/hash_heap_chain.py b/crates/lean_compiler/tests/programs/hash_heap_chain.py index 6a6950b16..e75eaf958 100644 --- a/crates/lean_compiler/tests/programs/hash_heap_chain.py +++ b/crates/lean_compiler/tests/programs/hash_heap_chain.py @@ -1,21 +1,18 @@ -# Runtime slices: `buf[i:i + 2]` with a runtime g-power index `i` names the -# heap cells `buf·i·g^k`, k < 2 (one MUL folds `i` into the pointer). A BLAKE2s -# chain over heap pairs (256-bit BLAKE2s value = two canonical cells), -# addressed by the loop counter: value k sits at cells g^{2k}..g^{2k+1}, and -# value k+1 = H(value k, value k). Published: the two 128-bit digest cells of -# H^3(5, 7). -# public_input: 64347157528245356000384183465036755063, 163839818445703091465558660402169004232 +# Runtime slices: `buf[i:i + 4]` with a runtime g-power index `i` names the +# heap cells `buf·i·g^k`, k < 4 (one MUL folds `i` into the pointer). A BLAKE2s +# chain over heap runs (a 256-bit BLAKE2s value is four cells), addressed by the +# loop counter: value k sits at cells g^{4k}..g^{4k+3}, and value k+1 = +# H(value k, value k). Published: the four digest words of H^3(5, 7). +# public_input: 248045690890498167, 3488266399269529831, 15161625676544491720, 8881774354923095707 from snark_lib import * def main(): - buf = HeapBuf(8) - buf[1] = 5 - buf[GEN] = 7 + buf = HeapBuf(16) + buf[0:4] = [5, 0, 7, 0] for i in mul_range(1, GEN ** 3): - b = i * i # value k at cells g^{2k}..g^{2k+1} - blake2s(buf[b:b + 2], buf[b:b + 2], buf[b * GEN ** 2:b * GEN ** 2 + 2]) + b = i ** 4 # value k at cells g^{4k}..g^{4k+3} + blake2s(buf[b:b + 4], buf[b:b + 4], buf[b * GEN ** 4:b * GEN ** 4 + 4]) p = GEN ** 0 - p[1] = buf[GEN ** 6] - p[GEN] = buf[GEN ** 7] + p[0:4] = buf[12:16] return diff --git a/crates/lean_compiler/tests/programs/hash_slices.py b/crates/lean_compiler/tests/programs/hash_slices.py index 5fdfd1184..8010ddf89 100644 --- a/crates/lean_compiler/tests/programs/hash_slices.py +++ b/crates/lean_compiler/tests/programs/hash_slices.py @@ -1,27 +1,31 @@ -# BLAKE2s over slices: `buf[lo:hi]` (2 cells) is a 256-bit operand under 128-bit -# machine words, with compile-time bounds: literals, literal-bound names, and -# their integer arithmetic (`x:x + 2`). Slices work on a large StackBuf (in -# place) and on a HeapBuf (bridged through the stack, one DEREF per cell), as -# inputs and as the output. Published: the two 128-bit digest cells of -# H(H(a[0:2], hb[0:2]), a[0:2]) read back from the heap. -# public_input: 249862442812096632729038560305983163980, 150628675827268462743577983046613573776 +# BLAKE2s over slices: `buf[lo:hi]` (4 cells) is a 256-bit operand, with +# compile-time bounds: literals, literal-bound names, and their integer +# arithmetic (`x:x + 4`). Slices work on a StackBuf (in place) and on a HeapBuf +# (bridged through the stack, one DEREF per cell), as inputs and as the output. +# Published: the four digest words of H(H(a[0:4], hb[0:4]), a[0:4]) read back +# from the heap. +# public_input: 14680300424583614028, 13545070165970514047, 15217079420999540880, 8165596878526962706 from snark_lib import * def main(): a = StackBuf(4) a[0] = 5 - a[1] = 7 - a[2] = 0 + a[1] = 0 + a[2] = 7 a[3] = 0 - hb = HeapBuf(4) + hb = HeapBuf(8) hb[1] = 11 # heap cell g^0 - hb[GEN] = 13 # heap cell g^1 + hb[GEN] = 0 + hb[GEN ** 2] = 13 + hb[GEN ** 3] = 0 x = 0 - h = StackBuf(2) - blake2s(a[x:x + 2], hb[0:2], h) # stack slice + heap input slice - blake2s(h, a[0:2], hb[2:4]) # digest lands in heap cells g^2, g^3 + h = StackBuf(4) + blake2s(a[x:x + 4], hb[0:4], h) # stack slice + heap input slice + blake2s(h, a[0:4], hb[4:8]) # digest lands in heap cells g^4..g^7 p = GEN ** 0 - p[1] = hb[GEN ** 2] - p[GEN] = hb[GEN ** 3] + p[1] = hb[GEN ** 4] + p[GEN] = hb[GEN ** 5] + p[GEN ** 2] = hb[GEN ** 6] + p[GEN ** 3] = hb[GEN ** 7] return diff --git a/crates/lean_compiler/tests/programs/unroll.py b/crates/lean_compiler/tests/programs/unroll.py index 84c622964..62eaf1c81 100644 --- a/crates/lean_compiler/tests/programs/unroll.py +++ b/crates/lean_compiler/tests/programs/unroll.py @@ -2,10 +2,10 @@ # as the integer literal of each iteration: zero loop overhead (no call, no # frame, no counter). Bounds are compile-time integers, including Const # parameters: `chain(buf, 3)` specializes and unrolls three BLAKE2s steps over -# heap slices indexed by `i` (a 256-bit BLAKE2s value is two canonical cells). -# Published: the two 128-bit digest cells of H^3(5, 7): same chain as +# heap slices indexed by `i` (a 256-bit BLAKE2s value is four cells). +# Published: the four digest words of H^3(5, 7): same chain as # hash_heap_chain.py, unrolled instead of looped. -# public_input: 64347157528245356000384183465036755063, 163839818445703091465558660402169004232 +# public_input: 248045690890498167, 3488266399269529831, 15161625676544491720, 8881774354923095707 from snark_lib import * @@ -15,17 +15,15 @@ def main(): for i in unroll(0, 7): sb[i + 1] = sb[i] * GEN # sb[k] = g^k assert sb[7] == GEN ** 7 - buf = HeapBuf(8) - buf[1] = 5 - buf[GEN] = 7 + buf = HeapBuf(16) + buf[0:4] = [5, 0, 7, 0] chain(buf, 3) p = GEN ** 0 - p[1] = buf[GEN ** 6] - p[GEN] = buf[GEN ** 7] + p[0:4] = buf[12:16] return def chain(buf, n: Const): for i in unroll(0, n): - blake2s(buf[i * 2:i * 2 + 2], buf[i * 2:i * 2 + 2], buf[i * 2 + 2:i * 2 + 4]) + blake2s(buf[i * 4:i * 4 + 4], buf[i * 4:i * 4 + 4], buf[i * 4 + 4:i * 4 + 8]) return diff --git a/crates/lean_compiler/tests/programs/wots_walk.py b/crates/lean_compiler/tests/programs/wots_walk.py index 5558b2aa7..c99c9357a 100644 --- a/crates/lean_compiler/tests/programs/wots_walk.py +++ b/crates/lean_compiler/tests/programs/wots_walk.py @@ -1,38 +1,34 @@ # A miniature WOTS-style chain walk bundling the DSL's moving parts: a # runtime digit is range-checked (dispatch soundness), then match # dispatches it to a Const-specialized walker whose BLAKE2s chain is unrolled -# over heap slices (a 256-bit BLAKE2s value occupies two canonical cells); -# the walker also builds g^{2n} at runtime (unrolled MULs) to read its final -# pair back through g-power indexing. The recomputation at the end lands on an -# already-written StackBuf pair, so write-once turns the hash into a digest -# assertion; the dead `if` branch holds an impossible assert that must never -# execute. Published: the two 128-bit digest cells of H^2(5, 7). -# public_input: 218111983282286173876109193675516368367, 307986319416510097496844621979881208676 +# over heap slices (a 256-bit BLAKE2s value occupies four cells); the walker +# reads its final digest back through a folded g^{4n} pointer. The +# recomputation at the end lands on an already-written run, so write-once +# turns the hash into a digest assertion; the dead `if` branch holds an +# impossible assert that must never execute. Published: the four digest words +# of H^2(5, 7). +# public_input: 14532901099560457711, 11823874305988834691, 15091288840866790244, 16695971830359737202 from snark_lib import * def main(): buf = HeapBuf(16) - buf[1] = 5 - buf[GEN] = 7 + buf[0:4] = [5, 0, 7, 0] d = GEN ** 2 # the runtime digit assert log(d) < 4 # bound the scrutinee before dispatching on it - t0, t1 = match(log(d), range(0, 4), lambda i: walk(buf, i)) + t0, t1, t2, t3 = match(log(d), range(0, 4), lambda i: walk(buf, i)) if d != GEN ** 2: assert 1 == 0 # dead branch: never executes - v = StackBuf(2) - v[0] = t0 - v[1] = t1 - blake2s(buf[2:4], buf[2:4], v) # recompute H(value1, value1): asserts v[0:2] == (t0, t1) + v = [t0, t1, t2, t3] + blake2s(buf[4:8], buf[4:8], v) # recompute H(value1, value1): asserts v == (t0, t1, t2, t3) p = GEN ** 0 - p[1] = t0 - p[GEN] = t1 + p[0:4] = v return def walk(buf, n: Const): p = 1 for i in unroll(0, n): - blake2s(buf[i * 2:i * 2 + 2], buf[i * 2:i * 2 + 2], buf[i * 2 + 2:i * 2 + 4]) - p = p * GEN * GEN - return buf[p], buf[p * GEN] + blake2s(buf[i * 4:i * 4 + 4], buf[i * 4:i * 4 + 4], buf[i * 4 + 4:i * 4 + 8]) + p = p * GEN ** 4 + return buf[p], buf[p * GEN], buf[p * GEN ** 2], buf[p * GEN ** 3] diff --git a/crates/lean_compiler/tests/suite/assert_ne.rs b/crates/lean_compiler/tests/suite/assert_ne.rs index ada182ee0..47ef7c002 100644 --- a/crates/lean_compiler/tests/suite/assert_ne.rs +++ b/crates/lean_compiler/tests/suite/assert_ne.rs @@ -1,11 +1,13 @@ -//! `assert a != b`: a proof-enforced inequality. It lowers to `XOR x = a + b`, -//! a hinted `inv = x⁻¹`, `MUL p = x·inv` and `SET p = 1`, the write-once +//! `assert a != b`: a proof-enforced inequality. It lowers to `XOR64 x = a + b`, +//! a hinted `inv = x⁻¹`, `MUL64 p = x·inv` and `SET p = 1`, the write-once //! conflict on `p` being the assertion. Sound whatever the hint: `x = 0` forces //! `p = 0`, which cannot then be set to `1`. Three rows and no `JUMP`. use lean_compiler::{compile, parse}; use lean_vm::cpu::{Op, prove, verify}; -use primitives::field::{F64, F192, g_pow}; +use primitives::field::{F64, g_pow}; + +use crate::common::pi; /// Honest inequality over runtime values: prove + verify pass, and corrupting /// the public output is still caught (the assert does not disturb the trace). @@ -24,11 +26,11 @@ def main(): return "; let program = compile(&parse(src).expect("parse")); - let want = [F192::from(g_pow(12)), F192::from(g_pow(5))]; + let want = pi(&[g_pow(12), g_pow(5)]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("inequality program verifies"); - let bad = [F192::from(g_pow(11)), F192::from(g_pow(5))]; + let bad = pi(&[g_pow(11), g_pow(5)]); assert!( verify(&program, &bad, &proof).is_err(), "wrong public input must be rejected" @@ -53,9 +55,9 @@ def main(): "; let run = |a: F64, b: F64| -> bool { let mut program = compile(&parse(src).expect("parse")); - program.set_witness("vals", vec![vec![F192::from(a), F192::from(b)]]); + program.set_witness("vals", vec![vec![a, b]]); std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| { - let pi = [F192::from(a), F192::from(b)]; + let pi = pi(&[a, b]); let (proof, _) = prove(&program, pi, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &pi, &proof).is_ok() })) @@ -81,7 +83,7 @@ def main(): return "; let program = compile(&parse(src).expect("parse")); - let want = [F192::from(F64(5)), F192::from(F64(7))]; + let want = pi(&[F64(5), F64(7)]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("loop inequality verifies"); } @@ -102,8 +104,8 @@ fn opcode_delta(with: &str, without: &str) -> (i64, i64, i64) { let (mut xor, mut mul, mut jump) = (0i64, 0i64, 0i64); for op in &p.prog { match op { - Op::Xor { .. } => xor += 1, - Op::Mul { .. } => mul += 1, + Op::Xor64 { .. } => xor += 1, + Op::Mul64 { .. } => mul += 1, Op::Jump { .. } => jump += 1, _ => {} } @@ -114,10 +116,10 @@ fn opcode_delta(with: &str, without: &str) -> (i64, i64, i64) { (a.0 - b.0, a.1 - b.1, a.2 - b.2) } -/// The check costs one `XOR`, one `MUL` and, above all, no `JUMP`: a -/// branch-based lowering would put an `E`-valued condition back on the one table -/// that carries constraints. The `SET` that closes the check is not counted, -/// constant materialisation elsewhere moving with the frame layout. +/// The check costs one `XOR64`, one `MUL64` and, above all, no `JUMP`: a +/// branch-based lowering would put the condition back on the one table that +/// carries constraints. The `SET` that closes the check is not counted, constant +/// materialisation elsewhere moving with the frame layout. #[test] fn assert_ne_emits_no_jump() { let body = |extra: &str| { @@ -134,11 +136,11 @@ def main(): ) }; let delta = opcode_delta(&body(" assert x != y\n"), &body("")); - assert_eq!(delta, (1, 1, 0), "one XOR, one MUL, no JUMP"); + assert_eq!(delta, (1, 1, 0), "one XOR64, one MUL64, no JUMP"); } /// The check survives cell sharing. `SET p = 1` writes a constant another cell -/// may already hold, and `MUL p = x·inv` a product that could otherwise be +/// may already hold, and `MUL64 p = x·inv` a product that could otherwise be /// shared; both are kept because `p` is written twice, which is what makes the /// write-once conflict the assertion. Dropping either would delete the check /// silently, so this pins it: the same product exists elsewhere in the frame, @@ -160,9 +162,9 @@ def main(): "; let run = |a: F64, b: F64| -> bool { let mut program = compile(&parse(src).expect("parse")); - program.set_witness("vals", vec![vec![F192::from(a), F192::from(b)]]); + program.set_witness("vals", vec![vec![a, b]]); std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| { - let pi = [F192::from(a) + F192::from(b), F192::ONE]; + let pi = pi(&[a + b, F64::ONE]); let (proof, _) = prove(&program, pi, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &pi, &proof).is_ok() })) @@ -192,18 +194,18 @@ def main(): p[GEN] = v[1] return "; - let run = |a: F192, b: F192, inv: F192| -> bool { + let run = |a: F64, b: F64, inv: F64| -> bool { let mut program = compile(&parse(src).expect("parse")); program.set_witness("vals", vec![vec![a, b, inv]]); std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| { - let (proof, _) = prove(&program, [a, b], lean_vm::pcs::TEST_LOG_INV_RATE); - verify(&program, &[a, b], &proof).is_ok() + let (proof, _) = prove(&program, pi(&[a, b]), lean_vm::pcs::TEST_LOG_INV_RATE); + verify(&program, &pi(&[a, b]), &proof).is_ok() })) .unwrap_or(false) }; - let (a, b) = (F192::from(g_pow(3)), F192::from(g_pow(5))); + let (a, b) = (g_pow(3), g_pow(5)); let d = a + b; assert!(run(a, b, d.inv()), "the true inverse verifies"); - assert!(!run(a, b, d.inv() + F192::ONE), "a wrong inverse must be rejected"); - assert!(!run(a, a, F192::ONE), "equal sides admit no inverse at all"); + assert!(!run(a, b, d.inv() + F64::ONE), "a wrong inverse must be rejected"); + assert!(!run(a, a, F64::ONE), "equal sides admit no inverse at all"); } diff --git a/crates/lean_compiler/tests/suite/common/mod.rs b/crates/lean_compiler/tests/suite/common/mod.rs index b54c2f88b..771fe840a 100644 --- a/crates/lean_compiler/tests/suite/common/mod.rs +++ b/crates/lean_compiler/tests/suite/common/mod.rs @@ -2,18 +2,33 @@ #![allow(dead_code)] use lean_compiler::{compile_without_filler, parse}; -use primitives::field::F192; +use primitives::field::F64; /// The program's own instruction mix: a build without the fill blocks, executed but not /// proven. Proving needs them, since a table's height has to be a power of two with no /// padding rows, but their dummy rows would drown out exactly what these counts are /// measuring. -pub fn mix(src: &str, pi: [F192; 2]) -> [usize; lean_vm::cpu::Stats::TABLES.len()] { +pub fn mix(src: &str, pi: [F64; 4]) -> [usize; lean_vm::cpu::Stats::TABLES.len()] { compile_without_filler(&parse(src).expect("parse")) .execute(pi) .base_counts } +/// A public input from up to four words, zero-padded. +pub fn pi(words: &[F64]) -> [F64; 4] { + let mut pi = [F64::ZERO; 4]; + pi[..words.len()].copy_from_slice(words); + pi +} + +/// The panic message of a caught unwind, for asserting on a diagnostic. +pub fn panic_message(err: &(dyn std::any::Any + Send)) -> String { + err.downcast_ref::() + .cloned() + .or_else(|| err.downcast_ref::<&str>().map(|s| s.to_string())) + .unwrap_or_default() +} + /// An AST's shape with source lines stripped. Two spellings of the same program /// (a constant against its substituted value, a placeholder against the filled /// text) are the same program but rarely occupy the same lines, so comparing diff --git a/crates/lean_compiler/tests/suite/const_placeholder.rs b/crates/lean_compiler/tests/suite/const_placeholder.rs index 8a7acf75e..88237c4a7 100644 --- a/crates/lean_compiler/tests/suite/const_placeholder.rs +++ b/crates/lean_compiler/tests/suite/const_placeholder.rs @@ -78,10 +78,9 @@ def main(): /// /// The scalar path tried an `f192` literal, then an integer expression, and /// stopped, so `GEN ** 2` was rejected as "not a compile-time integer constant -/// expression" while `f192(4, 0, 0)` naming the same element was accepted. It -/// now falls back to the field evaluator and renders the value as a decimal -/// wherever it fits the low two limbs, so the constant still works in the -/// positions that demand a literal rather than only as a value. +/// expression". It now falls back to the field evaluator and renders the word as +/// a decimal, so the constant still works in the positions that demand a literal +/// rather than only as a value. #[test] fn a_global_constant_may_be_a_g_power() { for (decl, exp) in [("GEN ** 2", 2usize), ("GEN * GEN", 2), ("GEN ** 70", 70)] { @@ -96,7 +95,7 @@ def main(): " ); let program = compile(&parse(&src).unwrap_or_else(|e| panic!("`{decl}`: {e}"))); - let want = [g_pow(exp).into(), g_pow(0).into()]; + let want = crate::common::pi(&[g_pow(exp), g_pow(0)]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).unwrap_or_else(|e| panic!("`{decl}` is not g^{exp}: {e:?}")); } diff --git a/crates/lean_compiler/tests/suite/determinism.rs b/crates/lean_compiler/tests/suite/determinism.rs index a68931849..ed92452f5 100644 --- a/crates/lean_compiler/tests/suite/determinism.rs +++ b/crates/lean_compiler/tests/suite/determinism.rs @@ -31,21 +31,22 @@ use lean_vm::cpu::Program; /// without a digest. #[rustfmt::skip] const GOLDEN: &[(&str, &str)] = &[ - ("conditionals", "e8df2b807d3eed366d83843e207f0a9441c1e2014ee0add958a399af562a761c"), - ("const_params", "f54f8df2c2259de3452565448aa2ae51b3dc3516ce21e04ba6d4a092848f9da3"), - ("fibonacci", "1419063250d0c54fa808499fdca691748dfb91f84f1007c6f0e939d5ef6779b6"), - ("hash_heap_chain", "9d8735c8393dd6cde7e07a71277a52435ca903186c6a984161b2b8821d6b552d"), - ("hash_slices", "99b3de0af47b797b51b8f38f37b5ccb7ff64d3f5acd9ba23bdd1ff4cfc47e9c3"), - ("heapbuf_dyn", "9e570732cf9258379eec6caf1efb2bd2d8ed484b5aa8c7b7da5360d2a79103ef"), - ("hint", "d9601df070184a3f1a5702d769e3897013c17a264889161df8ae7efe0daecc13"), - ("identities", "49a2bbd6bf785786f2ce8bf8f63a57a21bb6c547eae0c0f245f20a5222ac1c7a"), - ("match", "803e3a09e25825144dca768065f8acf69fa06a6ffb49c2b26e09f91849ae004b"), - ("match_arms", "00dcb9eff4ae060c20316f2567516c79d78fd0402abbefe35fbf3b62eee7bbb6"), - ("nested", "ae5e3556925b1ce14a0e7b6000fb5dc4a8909a7788e19f70c90dba418be56951"), - ("runtime_loop", "f33c6b5b82ed8ade198fb8978ba443554fc8714082a98dc784dbabdcee92e0fe"), - ("scoping", "c466babcc1af1deba56dda628e730d815d2fdffe065d695f9118167e0bb8669f"), - ("unroll", "08ebe1f4b51d862c6335b90694cf60d2fd2841d3a9913f6400cb53322117d309"), - ("wots_walk", "060605d64f78c7574598f40b0b060a58059a2122bfdcb79eb9ab3d317530d9d7"), + ("conditionals", "403f17d0ce5316049962b07eaa74c9f55c7c94add1d14978697a3487e3d54912"), + ("const_params", "f0816c2457b66413300192aabab590726e605bc551115c9252f26269d8ddd532"), + ("fibonacci", "ea71300abd23765da4f368175d4d0857b7b1175e121bccae42bd39d2dbb682cc"), + ("field192", "30d6d705f008a775a5e619a2ee3011bf8f39ea0244eaa10fc668d66c5a26ae02"), + ("hash_heap_chain", "7f19d09176d36956983a684b021c1b9c7897cafe3ec34fecfcc440c6203a60b5"), + ("hash_slices", "821c3206c7f5ce9cf37dff51bf8d2bd42c480e855d8ee120d465bc615abfc0fe"), + ("heapbuf_dyn", "cc6ac06781d7b71d7bbee5cb55d56dd4b7dee30ff45e433ae7fae5a26715690d"), + ("hint", "743fd1ba4dbe1bca619ef3b790681f14f3e3e5eb750c242be663373be8c266e9"), + ("identities", "b9e28816fcc5465c3d18801934d89d1cd0e39af798278ba29ddc313716f8d111"), + ("match", "f561aad705e79b8fa1ab304234aea570df0097d413451334b6d4b3cd22c3a3f7"), + ("match_arms", "bad6c31b086d81fb8d31e7e3207188f0331d598dab498ddbd1511fb7616d6149"), + ("nested", "b24339ff465c09e2f3efd0c8c349ee8b9d21ef74a7a5c64d01898b4c35bef48d"), + ("runtime_loop", "18435f31c81f25baaa40e72964746a4d95f441f1c4854599d89f1d641053d378"), + ("scoping", "0c37e2e4f7d5481b041752c2b0a5f737477cc57c8ec4d19acf4061ed572e0248"), + ("unroll", "28af13b92d23587254ddd72d1c66f5320417b84c280d07ba8787d79688360409"), + ("wots_walk", "1711401e58484c9158ae987e4c32342bdc0cf48cfbed2f85981ab1aad8f9b5f0"), ]; fn digest(p: &Program) -> String { diff --git a/crates/lean_compiler/tests/suite/disassemble.rs b/crates/lean_compiler/tests/suite/disassemble.rs index 462f55490..111169e44 100644 --- a/crates/lean_compiler/tests/suite/disassemble.rs +++ b/crates/lean_compiler/tests/suite/disassemble.rs @@ -1,4 +1,4 @@ -//! `disassemble` must render every one of the six opcodes without panicking, +//! `disassemble` must render every one of the eight opcodes without panicking, //! so it stays usable for the `DBG_DISASM` workflow (a failed guest `assert` //! surfaces as a write-once conflict, and the pc is all you get). @@ -8,25 +8,18 @@ use primitives::pretty_integer; #[test] fn disassemble_covers_every_opcode() { let src = "\ -@inline -def pack64x2(a, b): - assert_in_k(a, b) - return a + f192(0, 1, 0) * b - def main(): buff = HeapBuf(6) buff[1] = 1 buff[GEN] = GEN for i in mul_range(1, GEN ** 4): buff[i * GEN ** 2] = buff[i] * buff[i * GEN] - h = StackBuf(2) - h[0] = 5 - h[1] = 7 - d = StackBuf(2) + h = [5, 0, 7, 0] + d = StackBuf(4) blake2s(h, h, d) - packed = pack64x2(5, 7) + e = mul192(add192(d[0:3], f192(1, 2, 3)), d[1:4]) p = 1 - p[1] = buff[GEN ** 4] + packed + p[1] = buff[GEN ** 4] + e[0] p[GEN] = d[0] return "; @@ -41,7 +34,7 @@ def main(): let text = disassemble(&program.prog); print!("{text}"); - for mnemonic in ["SET", "XOR", "MUL", "DEREF", "JUMP", "BLAKE2S"] { + for mnemonic in ["SET", "XOR64", "MUL64", "DEREF", "JUMP", "BLAKE2S", "XOR192", "MUL192"] { assert!(text.contains(mnemonic), "disassembly is missing {mnemonic}"); } } diff --git a/crates/lean_compiler/tests/suite/field192.rs b/crates/lean_compiler/tests/suite/field192.rs new file mode 100644 index 000000000..6e2090b68 --- /dev/null +++ b/crates/lean_compiler/tests/suite/field192.rs @@ -0,0 +1,157 @@ +//! 192-bit values: runs of three cells, limbs low first, computed by `add192`, +//! `mul192` and `div192`. A run and a scalar never stand in for each other, runs +//! cross the heap in both directions, and the quotient back-solve refuses a zero +//! divisor. The soundness of the arithmetic itself is in `soundness`. + +use lean_compiler::{compile, compile_without_filler, parse}; +use lean_vm::cpu::{prove, verify}; +use primitives::field::{F64, F192, g_pow}; + +use crate::common::panic_message; + +/// Where a 192-bit value and a scalar meet, the program is rejected with a +/// diagnostic naming the widths, rather than one side reading the other's first +/// cell or a statement dropping a value on the floor. +#[test] +fn a_run_and_a_scalar_do_not_stand_in_for_each_other() { + for (body, want) in [ + ( + "y = add192(f192(1, 2, 3), f192(4, 5, 6)) + 1", + "is a 192-bit value, a run of three cells, not a scalar", + ), + ("x = f192(1, 2, 3)\n assert x == 5", "used as a scalar"), + ( + "y = mul192(GEN, f192(4, 5, 6))", + "expected a 3-cell value, got the scalar", + ), + ( + "h = StackBuf(4)\n y = mul192(h, f192(4, 5, 6))", + "expected a 3-cell value, got a 4-cell one", + ), + ( + "assert_eq192(f192(1, 2, 3), GEN)", + "expected a 3-cell value, got the scalar", + ), + ( + "add192(f192(1, 2, 3), f192(4, 5, 6))", + "is a 192-bit value, not a statement", + ), + ( + "hb = HeapBuf(4)\n hb[0:4] = f192(1, 2, 3)", + "expected a 4-cell value, got a 3-cell one", + ), + ( + "s = StackBuf(2)\n s[0:2] = mul192(f192(1, 2, 3), f192(4, 5, 6))", + "expected a 2-cell value, got a 3-cell one", + ), + ("r = sq(GEN)", "pass a 3-cell value"), + ("y = f192(1, 2)", "takes three limbs"), + ( + "y = f192(1, 2, 18446744073709551616)", + "an f192 limb is a compile-time 64-bit integer", + ), + ] { + let src = format!("def sq(x: StackBuf(3)):\n return mul192(x, x)\n\ndef main():\n {body}\n return\n"); + let msg = match parse(&src) { + Err(e) => e, + Ok(ast) => match std::panic::catch_unwind(|| compile(&ast)) { + Err(err) => panic_message(&*err), + Ok(_) => panic!("accepted: {body}"), + }, + }; + assert!(msg.contains(want), "`{body}`: wanted `{want}`, got `{msg}`"); + } +} + +/// A run read off the heap and stored back is the same value: runtime heap slices +/// written and read in a loop (a cached pair of loads per iteration), a stack copy +/// of a heap run, a list flattening two runs into one buffer, and a heap run +/// store publishing the result. The published scalar is the heap cell holding +/// the low limb, which pins the limb order. +#[test] +fn heap_runs_round_trip() { + let src = "\ +def main(): + x = StackBuf(3) + hint_witness(x, \"x\") + heap = HeapBuf(12) + heap[0:3] = x + for i in mul_range(1, GEN ** 3): + b = i ** 3 + heap[b * GEN ** 3:b * GEN ** 3 + 3] = mul192(heap[b:b + 3], heap[b:b + 3]) + y = heap[9:12] + both = [y, heap[0:3]] + assert_eq192(both[3:6], x) + p = GEN ** 0 + p[0:3] = both[0:3] + p[GEN ** 3] = heap[GEN ** 9] + return +"; + let x = F192::new(g_pow(7).0, 0x0123_4567_89ab_cdef, 42); + let y = (0..3).fold(x, |acc, _| acc * acc); + let mut program = compile(&parse(src).expect("parse")); + program.set_witness("x", vec![vec![F64(x.c0), F64(x.c1), F64(x.c2)]]); + let want = [F64(y.c0), F64(y.c1), F64(y.c2), F64(y.c0)]; + let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); + verify(&program, &want, &proof).expect("x^8 comes back off the heap"); + + let mut bad = want; + bad[2] += F64::ONE; + assert!(verify(&program, &bad, &proof).is_err(), "the top limb is bound"); +} + +/// `div192` by zero has no quotient, so the back-solve refuses it rather than +/// writing one the product cannot check. A divisor with zero low limbs is not +/// zero, so it divides. +#[test] +fn div192_by_zero_is_rejected() { + let src = "\ +def main(): + d = StackBuf(3) + hint_witness(d, \"d\") + q = div192(f192(1, 0, 0), d) + p = GEN ** 0 + p[0:3] = mul192(q, d) + return +"; + let program = compile_without_filler(&parse(src).expect("parse")); + let run = |d: [u64; 3]| { + let mut p = program.clone(); + p.set_witness("d", vec![d.map(F64).to_vec()]); + std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| { + p.execute([F64::ONE, F64::ZERO, F64::ZERO, F64::ZERO]) + })) + }; + assert!(run([0, 0, 1]).is_ok(), "y^2 is invertible"); + let Err(err) = run([0, 0, 0]) else { + panic!("a zero divisor has no quotient"); + }; + let msg = panic_message(&*err); + assert!(msg.contains("zero operand"), "got `{msg}`"); +} + +/// A `div192` stored into a run already holding some of the quotient's limbs +/// asserts those and fills the rest, as any store into written cells does: the +/// relation `q · b == a` determines every limb once `b` is nonzero. +#[test] +fn div192_into_a_partly_written_run() { + let src = "\ +def main(): + q = StackBuf(3) + q[0] = FIRST + q[0:3] = div192(mul192(f192(3, 5, 7), f192(11, 13, 17)), f192(11, 13, 17)) + p = GEN ** 0 + p[0:3] = q + return +"; + let program = compile(&parse(&src.replace("FIRST", "3")).expect("parse")); + let want = [F64(3), F64(5), F64(7), F64::ZERO]; + let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); + verify(&program, &want, &proof).expect("the written limb agrees with the quotient"); + + let wrong = compile_without_filler(&parse(&src.replace("FIRST", "4")).expect("parse")); + assert!( + std::panic::catch_unwind(|| wrong.execute(want)).is_err(), + "a written limb that disagrees with the quotient is rejected" + ); +} diff --git a/crates/lean_compiler/tests/suite/field_div.rs b/crates/lean_compiler/tests/suite/field_div.rs index e64646a1c..58ccbda7e 100644 --- a/crates/lean_compiler/tests/suite/field_div.rs +++ b/crates/lean_compiler/tests/suite/field_div.rs @@ -1,13 +1,15 @@ -//! Field division `a / b` (single slash): a runtime `a · b⁻¹`, distinct from -//! the compile-time floor-division `//`. It lowers to a single `MUL` whose +//! Field division `a / b` (single slash): a runtime `a · b⁻¹` in `K`, distinct +//! from the compile-time floor-division `//`. It lowers to a single `MUL64` whose //! quotient operand is left unset: the write-once back-solve fills it with -//! `a · b⁻¹` and the `MUL` constraint `quotient · b == a` binds it, with no +//! `a · b⁻¹` and the `MUL64` constraint `quotient · b == a` binds it, with no //! prover hint (the same back-solve the range-check gadget already uses). A //! zero divisor is rejected (the back-solve cannot invert 0). use lean_compiler::{compile, parse}; use lean_vm::cpu::{prove, verify}; -use primitives::field::{F64, F192, g_pow}; +use primitives::field::{F64, g_pow}; + +use crate::common::pi; /// `a / b` and `1 / b` over runtime values: the quotient satisfies `q·b == a`, /// checked by publishing it and reproducing the dividend. @@ -26,11 +28,11 @@ def main(): "; let program = compile(&parse(src).expect("parse")); // q·b must reproduce a = g^20; r·b must be 1. - let want = [F192::from(g_pow(20)), F192::from(F64::ONE)]; + let want = pi(&[g_pow(20), F64::ONE]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("division program verifies"); - let bad = [F192::from(g_pow(21)), F192::from(F64::ONE)]; + let bad = pi(&[g_pow(21), F64::ONE]); assert!( verify(&program, &bad, &proof).is_err(), "wrong quotient product rejected" @@ -54,7 +56,7 @@ def main(): "; let program = compile(&parse(src).expect("parse")); // q = g^6 / g^2 = g^4 (runtime `/`); z = g^(6//2) = g^3 (compile-time `//`). - let want = [F192::from(g_pow(4)), F192::from(g_pow(3))]; + let want = pi(&[g_pow(4), g_pow(3)]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("mixed //-and-/ program verifies"); } @@ -75,14 +77,11 @@ def main(): "; let run = |den: F64| -> bool { let mut program = compile(&parse(src).expect("parse")); - program.set_witness("den", vec![vec![F192::from(den)]]); + program.set_witness("den", vec![vec![den]]); std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| { - let (proof, _) = prove( - &program, - [F192::from(F64::ONE), F192::from(F64::ONE)], - lean_vm::pcs::TEST_LOG_INV_RATE, - ); - verify(&program, &[F192::from(F64::ONE), F192::from(F64::ONE)], &proof).is_ok() + let want = pi(&[F64::ONE, F64::ONE]); + let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); + verify(&program, &want, &proof).is_ok() })) .unwrap_or(false) }; diff --git a/crates/lean_compiler/tests/suite/filler.rs b/crates/lean_compiler/tests/suite/filler.rs index ef8d41212..64f410fcc 100644 --- a/crates/lean_compiler/tests/suite/filler.rs +++ b/crates/lean_compiler/tests/suite/filler.rs @@ -15,23 +15,25 @@ use lean_compiler::{compile, parse}; use lean_vm::cpu::filler; use lean_vm::cpu::{prove, verify}; -use primitives::field::F192; +use primitives::field::F64; -const PROGRAMS: [&str; 5] = [ +const PROGRAMS: [&str; 6] = [ // Folds to nothing, so the fill is all there is. "def main():\n x = GEN ** 5\n y = x * x\n return\n", "def main():\n b = HeapBuf(4)\n b[1] = GEN\n y = b[1] * b[1]\n return\n", "def main():\n for i in mul_range(1, GEN ** 20):\n z = i * i\n return\n", // A compression, so BLAKE2s is non-empty too. - "def main():\n a = StackBuf(2)\n a[0] = 5\n a[1] = 7\n c = StackBuf(2)\n blake2s(a, a, c)\n return\n", + "def main():\n a = [5, 0, 7, 0]\n c = StackBuf(4)\n blake2s(a, a, c)\n return\n", "def main():\n for i in mul_range(1, GEN ** 300):\n z = i + GEN\n return\n", + // The 192-bit tables, so their fill ops run too. + "def main():\n a = f192(3, 5, 7)\n b = mul192(a, a)\n assert_eq192(add192(b, a), add192(a, b))\n return\n", ]; #[test] fn every_table_lands_on_a_power_of_two() { for src in PROGRAMS { let program = compile(&parse(src).expect("parse")); - let pi = [F192::ZERO, F192::ZERO]; + let pi = [F64::ZERO; 4]; let (proof, stats) = prove(&program, pi, lean_vm::pcs::TEST_LOG_INV_RATE); assert!(filler::is_filled(stats.counts), "{:?} for {src:?}", stats.counts); verify(&program, &pi, &proof).expect("a filled program verifies"); @@ -45,7 +47,7 @@ fn every_table_lands_on_a_power_of_two() { fn the_cost_model_is_exact() { for src in PROGRAMS { let program = compile(&parse(src).expect("parse")); - let stats = prove(&program, [F192::ZERO, F192::ZERO], lean_vm::pcs::TEST_LOG_INV_RATE).1; + let stats = prove(&program, [F64::ZERO; 4], lean_vm::pcs::TEST_LOG_INV_RATE).1; let plan = filler::solve(stats.base_counts, filler::NO_FLOORS).expect("solvable"); assert_eq!( filler::filled(stats.base_counts, &plan), diff --git a/crates/lean_compiler/tests/suite/hint_log2_ceil.rs b/crates/lean_compiler/tests/suite/hint_log2_ceil.rs index e3bc0ea1d..8ed6d0853 100644 --- a/crates/lean_compiler/tests/suite/hint_log2_ceil.rs +++ b/crates/lean_compiler/tests/suite/hint_log2_ceil.rs @@ -6,16 +6,22 @@ use lean_compiler::{compile, parse}; use lean_vm::cpu::{prove, verify}; -use primitives::field::{F64, F192, g_pow}; +use primitives::field::{F64, g_pow}; -fn log2_ceil_of(v: u128) -> usize { +use crate::common::pi; + +fn log2_ceil_of(v: u64) -> usize { if v <= 1 { 0 } else { - (128 - (v - 1).leading_zeros()) as usize + (64 - (v - 1).leading_zeros()) as usize } } +fn bits_of(v: u64) -> Vec { + (0..8).map(|j| F64((v >> j) & 1)).collect() +} + #[test] fn log2_ceil_advice_computes_the_log() { let src = "\ @@ -28,14 +34,13 @@ def main(): p[GEN] = 1 return "; - for v in [1u128, 2, 3, 4, 5, 7, 8, 200] { + for v in [1u64, 2, 3, 4, 5, 7, 8, 200] { let mut program = compile(&parse(src).expect("parse")); - let bits: Vec = (0..8).map(|j| F192::from(F64(((v >> j) & 1) as u64))).collect(); - program.set_witness("bits", vec![bits]); - let want = [F192::from(g_pow(log2_ceil_of(v))), F192::from(F64::ONE)]; + program.set_witness("bits", vec![bits_of(v)]); + let want = pi(&[g_pow(log2_ceil_of(v)), F64::ONE]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).unwrap_or_else(|_| panic!("v={v}: log2_ceil advice must verify")); - let bad = [F192::from(g_pow(log2_ceil_of(v) + 1)), F192::from(F64::ONE)]; + let bad = pi(&[g_pow(log2_ceil_of(v) + 1), F64::ONE]); assert!( verify(&program, &bad, &proof).is_err(), "v={v}: wrong g_mu must be rejected" @@ -57,11 +62,10 @@ def main(): p[GEN] = 1 return "; - for (v, mu) in [(2u128, 5usize), (4, 5), (64, 6), (200, 8)] { + for (v, mu) in [(2u64, 5usize), (4, 5), (64, 6), (200, 8)] { let mut program = compile(&parse(src).expect("parse")); - let bits: Vec = (0..8).map(|j| F192::from(F64(((v >> j) & 1) as u64))).collect(); - program.set_witness("bits", vec![bits]); - let want = [F192::from(g_pow(mu)), F192::from(F64::ONE)]; + program.set_witness("bits", vec![bits_of(v)]); + let want = pi(&[g_pow(mu), F64::ONE]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).unwrap_or_else(|_| panic!("v={v}: floored log2_ceil must verify")); } diff --git a/crates/lean_compiler/tests/suite/inline_expr.rs b/crates/lean_compiler/tests/suite/inline_expr.rs index 196482f4e..420015a68 100644 --- a/crates/lean_compiler/tests/suite/inline_expr.rs +++ b/crates/lean_compiler/tests/suite/inline_expr.rs @@ -8,7 +8,9 @@ use lean_compiler::{compile, parse}; use lean_vm::cpu::{prove, verify}; -use primitives::field::{F64, F192}; +use primitives::field::F64; + +use crate::common::pi; #[test] fn inline_call_in_expression_positions() { @@ -48,13 +50,13 @@ def main(): let y = f7 * (f3 * (one + f5)); // heap-store RHS: idx 3 -> 3·5 let o = f3 * f5; - let want = [F192::from(x), F192::from(y + o)]; + let want = pi(&[x, y + o]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("expression-position inline calls compute correctly"); let mut bad = want; - bad[0] += F192::ONE; + bad[0] += F64::ONE; assert!( verify(&program, &bad, &proof).is_err(), "wrong published value must be rejected" diff --git a/crates/lean_compiler/tests/suite/loop_frames.rs b/crates/lean_compiler/tests/suite/loop_frames.rs index c1f975be8..e87b4e198 100644 --- a/crates/lean_compiler/tests/suite/loop_frames.rs +++ b/crates/lean_compiler/tests/suite/loop_frames.rs @@ -1,6 +1,8 @@ use lean_compiler::{compile, compile_without_filler, parse}; use lean_vm::cpu::{prove, verify}; -use primitives::field::{F64, F192, g_pow}; +use primitives::field::{F64, g_pow}; + +use crate::common::pi; #[test] fn loop_frames_preserve_escaped_cells_and_nested_allocations() { @@ -48,7 +50,7 @@ def main(): let program = compile(&parse(source).unwrap()); for end in [2, 3, 9] { let sum = (2..end).fold(F64::ZERO, |sum, i| sum + g_pow(i)); - let public = [F192::from(sum), F192::from(g_pow(end))]; + let public = pi(&[sum, g_pow(end)]); assert!(program.execute(public).unconstrained_reads.is_empty()); if end == 9 { let (proof, _) = prove(&program, public, lean_vm::pcs::TEST_LOG_INV_RATE); @@ -71,7 +73,7 @@ def main(): return "#; let program = compile_without_filler(&parse(source).unwrap()); - let execution = program.execute([F192::from(F64(7)), F192::ZERO]); + let execution = program.execute(pi(&[F64(7)])); assert!(execution.unconstrained_reads.is_empty()); } @@ -92,7 +94,7 @@ def main(): let program = compile_without_filler(&parse(source).unwrap()); assert!( program - .execute([g_pow(65536).into(), g_pow(65537).into()]) + .execute(pi(&[g_pow(65536), g_pow(65537)])) .unconstrained_reads .is_empty() ); @@ -119,7 +121,7 @@ def main(): assert seen[GEN] == GEN return "#; - let public = [F192::ZERO, g_pow(2).into()]; + let public = pi(&[F64::ZERO, g_pow(2)]); for bound in ["GEN ** 2", "public[GEN]"] { let program = compile(&parse(&source.replace("STOP", bound)).unwrap()); assert!(program.execute(public).unconstrained_reads.is_empty()); diff --git a/crates/lean_compiler/tests/suite/main.rs b/crates/lean_compiler/tests/suite/main.rs index f7d070318..54dec0b94 100644 --- a/crates/lean_compiler/tests/suite/main.rs +++ b/crates/lean_compiler/tests/suite/main.rs @@ -9,13 +9,13 @@ mod assert_ne; mod const_placeholder; mod determinism; mod disassemble; +mod field192; mod field_div; mod field_towers; mod filler; mod hint_log2_ceil; mod inline_expr; mod loop_frames; -mod pack64x2; mod print_debug; mod py_source; mod range_check; diff --git a/crates/lean_compiler/tests/suite/pack64x2.rs b/crates/lean_compiler/tests/suite/pack64x2.rs deleted file mode 100644 index 33cb0b95c..000000000 --- a/crates/lean_compiler/tests/suite/pack64x2.rs +++ /dev/null @@ -1,57 +0,0 @@ -use lean_compiler::{compile, parse}; -use lean_vm::cpu::{prove, verify}; -use primitives::field::{F64, F192}; - -use crate::common::mix; - -#[test] -fn pack64x2_proves_and_verifies() { - let src = "\ -@inline -def pack64x2(a, b): - assert_in_k(a, b) - return a + f192(0, 1, 0) * b - -def main(): - a = 5 - b = 7 - packed = pack64x2(a, b) - p = 1 - p[1] = packed - p[GEN] = packed - return -"; - let program = compile(&parse(src).expect("parse")); - let want = [F192::new(5, 7, 0), F192::new(5, 7, 0)]; - let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); - let counts = mix(src, want); - assert_eq!( - (counts[0], counts[1], counts[4]), - (1, 2, 2), - "XOR, MUL and JUMP lowering" - ); - verify(&program, &want, &proof).expect("pack64x2 program verifies"); -} - -#[test] -#[should_panic(expected = "JUMP target is not a K-valued word")] -fn pack64x2_rejects_extension_field_source() { - let src = "\ -@inline -def pack64x2(a, b): - assert_in_k(a, b) - return a + f192(0, 1, 0) * b - -def main(): - a = StackBuf(1) - hint_witness(a[0:1], \"a\") - packed = pack64x2(a[0], 7) - p = 1 - p[1] = packed - p[GEN] = packed - return -"; - let mut program = compile(&parse(src).expect("parse")); - program.set_witness("a", vec![vec![F192::new(5, 1, 0)]]); - let _ = program.execute([F192::from(F64::ONE), F192::from(F64::ONE)]); -} diff --git a/crates/lean_compiler/tests/suite/print_debug.rs b/crates/lean_compiler/tests/suite/print_debug.rs index c8e303526..15cfdc31c 100644 --- a/crates/lean_compiler/tests/suite/print_debug.rs +++ b/crates/lean_compiler/tests/suite/print_debug.rs @@ -3,7 +3,7 @@ use lean_compiler::{compile, parse}; use lean_vm::cpu::{prove, verify}; -use primitives::field::{F64, F192}; +use primitives::field::F64; #[test] fn print_is_constraint_free() { @@ -22,7 +22,7 @@ def main(): return "; let program = compile(&parse(src).expect("parse")); - let want = [F192::from(F64(5) * primitives::field::g_pow(1)), F192::from(F64(3))]; + let want = crate::common::pi(&[F64(5) * primitives::field::g_pow(1), F64(3)]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("prints must not disturb proving"); } diff --git a/crates/lean_compiler/tests/suite/py_source.rs b/crates/lean_compiler/tests/suite/py_source.rs index 1ff21d743..8902ceee3 100644 --- a/crates/lean_compiler/tests/suite/py_source.rs +++ b/crates/lean_compiler/tests/suite/py_source.rs @@ -4,42 +4,43 @@ //! //! The harness is generic: every `tests/programs/*.py` is parsed, compiled, //! proven, and verified. A program declares the public input it expects with a -//! top-of-file annotation of two constant field elements, +//! top-of-file annotation of up to four constant words, zero-padded, //! //! ```text -//! # public_input: GEN ** 89, 101229015297003380629709256178361811305 +//! # public_input: GEN ** 89, 101229015297003380 //! ``` //! -//! or omits it to run with the empty public input (two zeros). +//! or omits it to run with the empty public input (four zeros). use std::fs; use lean_compiler::{compile, parse, parse_const}; use lean_vm::cpu::{prove, verify}; -use primitives::field::F192; +use primitives::field::F64; -/// The `# public_input: , ` annotation, or `[0, 0]` if absent. -fn public_input(src: &str) -> [F192; 2] { +/// The `# public_input: , …` annotation, zero-padded to four words. +fn public_input(src: &str) -> [F64; 4] { + let mut words = [F64::ZERO; 4]; for line in src.lines() { if let Some(rest) = line.trim().strip_prefix("# public_input:") { let parts: Vec<&str> = rest.split(',').collect(); - assert_eq!( - parts.len(), - 2, - "`# public_input:` needs two field elements, got `{rest}`" + assert!( + parts.len() <= 4, + "`# public_input:` holds at most four words, got `{rest}`" ); - let elt = |s: &str| parse_const(s).unwrap_or_else(|e| panic!("bad public_input: {e}")); - return [elt(parts[0]), elt(parts[1])]; + for (word, s) in words.iter_mut().zip(parts) { + *word = parse_const(s).unwrap_or_else(|e| panic!("bad public_input: {e}")); + } } } - [F192::ZERO; 2] + words } -/// The `# witness : , …` annotations: one line per *entry* +/// The `# witness : , …` annotations: one line per *entry* /// (repeated lines with the same name are the stream's successive entries, /// popped by successive `hint_witness` calls). -fn witness(src: &str) -> std::collections::HashMap>> { - let mut streams: std::collections::HashMap>> = Default::default(); +fn witness(src: &str) -> std::collections::HashMap>> { + let mut streams: std::collections::HashMap>> = Default::default(); for rest in src.lines().filter_map(|l| l.trim().strip_prefix("# witness ")) { let (name, vals) = rest.split_once(':').expect("`# witness` needs `name: values`"); let entry = vals diff --git a/crates/lean_compiler/tests/suite/range_check.rs b/crates/lean_compiler/tests/suite/range_check.rs index 1263995ce..e15c4bb47 100644 --- a/crates/lean_compiler/tests/suite/range_check.rs +++ b/crates/lean_compiler/tests/suite/range_check.rs @@ -1,6 +1,6 @@ //! Range checks *in the exponent*: `assert log x < log GEN ** k` (or //! `assert log x < k`) proves `log_g(x) < k`, i.e. `x ∈ {g^0, g^1, …, g^{k-1}}`, -//! in 3 cycles: `DEREF x` bounds `log(x)` by the memory size, a `MUL` into the +//! in 3 cycles: `DEREF x` bounds `log(x)` by the memory size, a `MUL64` into the //! write-once constant cell `g^{k-1}` back-solves and binds the complement //! `y = g^{k-1-log(x)}`, and `DEREF y` bounds the complement. leanVM's DEREF //! range-check trick, transported to g-powers; the only nondeterminism is the @@ -8,9 +8,9 @@ use lean_compiler::{compile, parse}; use lean_vm::cpu::{prove, verify}; -use primitives::field::{F64, F192, g_pow}; +use primitives::field::{F64, g_pow}; -use crate::common::mix; +use crate::common::{mix, pi}; /// Both bound forms (`log GEN ** k` and a plain integer exponent) with the /// boundary elements (`g^{k-1}`, `1 = g^0`), end-to-end: prove + verify, and a @@ -33,13 +33,13 @@ def main(): return "; let program = compile(&parse(src).expect("parse")); - let want = [F192::from(g_pow(12)), F192::from(g_pow(5))]; + let want = pi(&[g_pow(12), g_pow(5)]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); // 2 DEREFs per range check (4 checks) + 2 publishing stores. assert_eq!(mix(src, want)[3], 10, "DEREF count"); verify(&program, &want, &proof).expect("range-checked program verifies"); - let bad = [F192::from(g_pow(12)), F192::from(g_pow(6))]; + let bad = pi(&[g_pow(12), g_pow(6)]); assert!( verify(&program, &bad, &proof).is_err(), "wrong public input must be rejected" @@ -61,7 +61,7 @@ def main(): return "; let program = compile(&parse(src).expect("parse")); - let want = [F192::from(g_pow(300)), F192::from(g_pow(300))]; + let want = pi(&[g_pow(300), g_pow(300)]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("deferred-fill program verifies"); } @@ -81,7 +81,7 @@ def main(): return "; let program = compile(&parse(src).expect("parse")); - let want = [F192::from(g_pow(5)); 2]; + let want = pi(&[g_pow(5), g_pow(5)]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("max-bound range check verifies"); } @@ -102,7 +102,7 @@ def main(): return "; let program = compile(&parse(src).expect("parse")); - let want = [F192::from(F64(5)), F192::from(F64(7))]; + let want = pi(&[F64(5), F64(7)]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); // 6 iterations × 2 range-check DEREFs, plus call/publish plumbing. assert!(mix(src, want)[3] >= 12, "at least the 12 range-check DEREFs"); @@ -117,7 +117,7 @@ def main(): fn range_check_at_bound_rejected() { let src = "def main():\n x = GEN ** 8\n assert log x < 8\n return\n"; let program = compile(&parse(src).expect("parse")); - program.execute([F192::ZERO, F192::ZERO]); + program.execute([F64::ZERO; 4]); } /// A value that is no small g-power at all (5 = x^2 + 1) fails at the first @@ -127,7 +127,7 @@ fn range_check_at_bound_rejected() { fn range_check_non_g_power_rejected() { let src = "def main():\n x = 5\n assert log x < 8\n return\n"; let program = compile(&parse(src).expect("parse")); - program.execute([F192::ZERO, F192::ZERO]); + program.execute([F64::ZERO; 4]); } /// Bound 0 names the empty set: rejected at compile time. @@ -177,8 +177,8 @@ def main(): return "; let mut program = compile(&parse(src).expect("parse")); - program.set_witness("n", vec![vec![F192::from(g_pow(6))]]); - let want = [F192::from(g_pow(5)), F192::from(g_pow(6))]; + program.set_witness("n", vec![vec![g_pow(6)]]); + let want = pi(&[g_pow(5), g_pow(6)]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("runtime-bound range check verifies"); } @@ -198,8 +198,8 @@ def main(): return "; let mut program = compile(&parse(src).expect("parse")); - program.set_witness("n", vec![vec![F192::from(g_pow(5))]]); - program.execute([F192::ZERO, F192::ZERO]); + program.set_witness("n", vec![vec![g_pow(5)]]); + program.execute([F64::ZERO; 4]); } /// A bound that folds at parse time but is not a power of `GEN` stays a parse diff --git a/crates/lean_compiler/tests/suite/sharing.rs b/crates/lean_compiler/tests/suite/sharing.rs index 79db8a49f..6ea98aa0a 100644 --- a/crates/lean_compiler/tests/suite/sharing.rs +++ b/crates/lean_compiler/tests/suite/sharing.rs @@ -6,7 +6,9 @@ use lean_compiler::{compile, parse}; use lean_vm::cpu::{prove, verify}; -use primitives::field::{F64, F192, g_pow}; +use primitives::field::{F64, g_pow}; + +use crate::common::pi; /// A returned value that repeats a constant computed earlier in the same /// function. The return slot lives in the callee frame and is read by the @@ -30,7 +32,7 @@ def main(): return "; let program = compile(&parse(src).expect("parse")); - let want = [F192::from(g_pow(3)) * F192::from(F64(7)), F192::from(F64(7))]; + let want = pi(&[g_pow(3) * F64(7), F64(7)]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("returned duplicate constant is preserved"); } @@ -55,7 +57,7 @@ def main(): "; let program = compile(&parse(src).expect("parse")); // (k + k) + (k + k) == 0 in characteristic two. - let want = [F192::ZERO, F192::ZERO]; + let want = [F64::ZERO; 4]; let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("duplicated call arguments are preserved"); } @@ -86,7 +88,7 @@ def main(): p[GEN] = x return "; - let run = |pi: [F192; 2]| -> bool { + let run = |pi: [F64; 4]| -> bool { let program = compile(&parse(src).expect("parse")); std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| { let (proof, _) = prove(&program, pi, lean_vm::pcs::TEST_LOG_INV_RATE); @@ -95,27 +97,27 @@ def main(): .unwrap_or(false) }; assert!( - run([F192::from(g_pow(4)), F192::from(g_pow(3))]), + run(pi(&[g_pow(4), g_pow(3)])), "the arm that runs keeps its own constant" ); // The wrong value is the half that bites: a cell only the untaken arm writes // is the prover's to choose, and the honest claim would verify anyway. assert!( - !run([F192::from(g_pow(7)), F192::from(g_pow(3))]), + !run(pi(&[g_pow(7), g_pow(3)])), "the stored constant is pinned by the arm that ran" ); } -/// The assert idiom is `XOR fp[t] = a ^ b` into the pooled zero cell, whose +/// The assert idiom is `XOR64 fp[t] = a ^ b` into the pooled zero cell, whose /// second write IS the assertion, so that cell must never be shared and the -/// `XOR` must never be skipped in favour of one computed earlier. +/// `XOR64` must never be skipped in favour of one computed earlier. /// /// The operands are HEAP READS on purpose. Written `a = GEN ** 9`, both sides -/// fold and `diff = a + b` emits no `XOR` at all, so the duplicate the test +/// fold and `diff = a + b` emits no `XOR64` at all, so the duplicate the test /// names does not exist and the test cannot fail: skipping the assert on a cache /// hit then passed the whole suite. Read from a `HeapBuf` the two values are -/// runtime cells, `diff` really does emit `XOR fp[t] = a ^ b`, and the assert's -/// own `XOR` has a genuine duplicate to be folded into. +/// runtime cells, `diff` really does emit `XOR64 fp[t] = a ^ b`, and the assert's +/// own `XOR64` has a genuine duplicate to be folded into. fn duplicated_comparison(second: u32) -> String { format!( "\ @@ -139,7 +141,7 @@ def main(): #[test] fn assert_survives_a_duplicated_comparison() { let program = compile(&parse(&duplicated_comparison(9)).expect("parse")); - let want = [F192::ZERO, F192::from(g_pow(9))]; + let want = pi(&[F64::ZERO, g_pow(9)]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("passing assert still verifies"); } @@ -150,7 +152,7 @@ fn assert_survives_a_duplicated_comparison() { #[should_panic(expected = "write-once conflict")] fn failing_assert_still_conflicts() { let program = compile(&parse(&duplicated_comparison(10)).expect("parse")); - let want = [F192::ZERO, F192::from(g_pow(9))]; + let want = pi(&[F64::ZERO, g_pow(9)]); let _ = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); } @@ -174,26 +176,26 @@ def main(): return "; let mut program = compile(&parse(src).expect("parse")); - program.set_witness("flag", vec![vec![F192::ZERO]]); - let want = [F192::ZERO, F192::ZERO]; + program.set_witness("flag", vec![vec![F64::ZERO]]); + let want = [F64::ZERO; 4]; let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("untaken branch must not consume its witness"); } -/// A BLAKE2s chaining value names a CONSECUTIVE PAIR, so neither half may be +/// A BLAKE2s chaining value names FOUR CONSECUTIVE cells, so none of them may be /// folded into a canonical elsewhere and the base may not be rewritten: a /// substitution speaks for one cell, and redirecting the base silently redirects -/// the second word too. `rewrite_reads` used to map `cv` like any single-cell -/// read, so when the first of the two assembling copies duplicated an earlier -/// copy of the same source, the compression absorbed the OTHER pair's second -/// word. Silent, and a soundness break in a transcript. +/// the other words too. `rewrite_reads` used to map `cv` like any single-cell +/// read, so when the first of the assembling copies duplicated an earlier copy of +/// the same source, the compression absorbed the OTHER run's later words. Silent, +/// and a soundness break in a transcript. /// /// The two compressions here differ in nothing but their chaining value, and -/// their two `cv` pairs share a first word, which is what made the first copy a -/// duplicate. If either pair is rewritten or dropped, the digests coincide and -/// the inequality fails at witness generation. +/// their two `cv` runs share their first three words, which is what made those +/// copies duplicates. If either run is rewritten or dropped, the digests coincide +/// and the inequality fails at witness generation. #[test] -fn a_chaining_value_pair_is_neither_rewritten_nor_dropped() { +fn a_chaining_value_run_is_neither_rewritten_nor_dropped() { let src = "\ def main(): hb = HeapBuf(4) @@ -210,16 +212,20 @@ def main(): msg[1] = y msg[2] = y msg[3] = y - t = StackBuf(2) + t = StackBuf(4) t[0] = x - t[1] = z - o1 = StackBuf(2) - blake2s(msg[0:2], msg[2:4], o1, cv=t, counter=64, final=1) - s = StackBuf(2) + t[1] = x + t[2] = x + t[3] = z + o1 = StackBuf(4) + blake2s(msg, msg, o1, cv=t, counter=64, final=1) + s = StackBuf(4) s[0] = x - s[1] = w - o2 = StackBuf(2) - blake2s(msg[0:2], msg[2:4], o2, cv=s, counter=64, final=1) + s[1] = x + s[2] = x + s[3] = w + o2 = StackBuf(4) + blake2s(msg, msg, o2, cv=s, counter=64, final=1) assert o1[0] != o2[0] p = 1 p[1] = x @@ -227,7 +233,7 @@ def main(): return "; let program = compile(&parse(src).expect("parse")); - let want = [F192::from(g_pow(11)), F192::from(g_pow(22))]; + let want = pi(&[g_pow(11), g_pow(22)]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("each compression absorbs its own chaining value"); } @@ -248,12 +254,12 @@ def main(): h[1] = public[GEN] return "#; - let value = F192::new(17, 31, 43); + let value = F64(0x2b_1f_11); let mut program = compile(&parse(source).unwrap()); - program.set_witness("values", vec![vec![value, F192::ONE, value]]); - for branch in [F192::ZERO, F192::ONE] { - assert!(program.execute([branch, value]).unconstrained_reads.is_empty()); - assert!(std::panic::catch_unwind(|| program.execute([branch, value + F192::ONE])).is_err()); + program.set_witness("values", vec![vec![value, F64::ONE, value]]); + for branch in [F64::ZERO, F64::ONE] { + assert!(program.execute(pi(&[branch, value])).unconstrained_reads.is_empty()); + assert!(std::panic::catch_unwind(|| program.execute(pi(&[branch, value + F64::ONE]))).is_err()); } } @@ -278,8 +284,8 @@ def main(): assert public[GEN] == out[1] return "#; - let value = F192::from(F64(7)); - let public = [value * value, value]; + let value = F64(7); + let public = pi(&[value * value, value]); for touch in ["assert log(h) < 1024", "early = StackBuf(1)\n early[0] = h[1]"] { for fill in ["hint_witness(h[0:1], \"value\")", "fill(h)"] { let source = source.replace("TOUCH", touch).replace("FILL", fill); @@ -288,7 +294,7 @@ def main(): assert!(program.execute(public).unconstrained_reads.is_empty()); let (proof, _) = prove(&program, public, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &public, &proof).unwrap(); - assert!(std::panic::catch_unwind(|| program.execute([public[0] + F192::ONE, value])).is_err()); + assert!(std::panic::catch_unwind(|| program.execute(pi(&[public[0] + F64::ONE, value]))).is_err()); } } } @@ -313,8 +319,8 @@ def main(): assert public[1] == result return "#; - let value = F192::from(F64(7)); - let public = [value * value, F192::ZERO]; + let value = F64(7); + let public = pi(&[value * value]); for (dest, fill) in [ ("early", "early[0] = value"), ("other", "other[0] = value"), diff --git a/crates/lean_compiler/tests/suite/soundness/cases.rs b/crates/lean_compiler/tests/suite/soundness/cases.rs index 62dc52f56..8abc40b03 100644 --- a/crates/lean_compiler/tests/suite/soundness/cases.rs +++ b/crates/lean_compiler/tests/suite/soundness/cases.rs @@ -6,14 +6,15 @@ //! every poke names one constraint. A poke that is accepted says which one is //! missing. //! -//! The pokes lean on witness streams rather than the public input, because two +//! The pokes lean on witness streams rather than the public input, because four //! public words is all there is and because the streams are where a real guest's //! untrusted data actually enters. -use super::{Case, Trial, check_case, g, k, pi, wit}; -use primitives::field::F192; +use super::{Case, Trial, check_case, g, k, limbs, pi, wit}; +use lean_vm::vmhash::compress; +use primitives::field::{F64, F192}; -/// `XOR`/`MUL` relations, both assert forms, and the division back-solve. The +/// `XOR64`/`MUL64` relations, both assert forms, and the division back-solve. The /// quotient cell is written by nothing but the back-solve, so this case also /// pins the one legitimate way a cell may be read before any instruction writes /// it. @@ -34,7 +35,7 @@ def main(): p[GEN] = v[0] + v[1] return ", - valid: Trial::new([g(8), g(3) + g(5)]).stream("w", vec![vec![g(3), g(5), g(8)]]), + valid: Trial::new(&[g(8), g(3) + g(5)]).stream("w", vec![vec![g(3), g(5), g(8)]]), pokes: vec![ // Each of the three hinted cells breaks the product relation. wit("w", 0, g(4)), @@ -50,6 +51,59 @@ def main(): }); } +/// The same relations over 192-bit runs: `MUL192`, both 192-bit assert forms, and +/// the `div192` back-solve, whose quotient run is written by nothing else. A +/// second quotient lands on a hinted run, so a claimed quotient is checked by the +/// product rather than trusted: the pokes on `q` are a forged quotient. +#[test] +fn arithmetic192_and_asserts() { + let a = F192::new(g(3).0, 11, 13); + let b = F192::new(g(5).0, 11, 13); + let c = a * b; + let w: Vec = [limbs(a), limbs(b), limbs(c)].concat(); + let sum = limbs(a + b); + let bump = |x: u64| F64(x) + F64::ONE; + check_case(&Case { + name: "arithmetic192_and_asserts", + src: "\ +def main(): + v = StackBuf(9) + hint_witness(v, \"w\") + assert_eq192(mul192(v[0:3], v[3:6]), v[6:9]) + assert_ne192(v[0:3], v[3:6]) + q = div192(v[6:9], v[0:3]) + assert_eq192(q, v[3:6]) + claimed = StackBuf(3) + hint_witness(claimed, \"q\") + claimed[0:3] = div192(v[6:9], v[3:6]) + p = GEN ** 0 + p[0:3] = add192(v[0:3], v[3:6]) + p[GEN ** 3] = v[6] + return +", + valid: Trial::new(&[sum[0], sum[1], sum[2], F64(c.c0)]) + .stream("w", vec![w]) + .stream("q", vec![limbs(a).to_vec()]), + pokes: vec![ + // Either factor, and the product's unpublished limbs. + wit("w", 1, bump(a.c1)), + wit("w", 5, bump(b.c2)), + wit("w", 7, bump(c.c1)), + wit("w", 8, bump(c.c2)), + // Equal operands, the poke `assert_ne192` exists for. + wit("w", 3, g(3)), + // A forged quotient, one limb at a time. + wit("q", 0, bump(a.c0)), + wit("q", 1, bump(a.c1)), + wit("q", 2, bump(a.c2)), + // The published sum and product limb. + pi(0, bump(sum[0].0)), + pi(2, bump(sum[2].0)), + pi(3, bump(c.c0)), + ], + }); +} + /// The exponent range check and `match` dispatch. The dispatch is only /// sound because the matched value was range-checked first (doc §Match /// statements), so a poke past the bound must be caught by the check rather than @@ -75,7 +129,7 @@ def sq(x): return x * x ", // Arm 3 runs: sq(3) = 3·3 in K = (x+1)^2 = x^2+1 = 5. - valid: Trial::new([g(3), k(5)]).stream("w", vec![vec![g(3), k(5)]]), + valid: Trial::new(&[g(3), k(5)]).stream("w", vec![vec![g(3), k(5)]]), pokes: vec![ // Past the bound: the range check's complement DEREF must catch it. wit("w", 0, g(8)), @@ -113,7 +167,7 @@ def main(): p[GEN] = v[0] return ", - valid: Trial::new([g(6), g(2)]).stream("w", vec![vec![g(2), g(5)]]), + valid: Trial::new(&[g(6), g(2)]).stream("w", vec![vec![g(2), g(5)]]), pokes: vec![ // Takes the else arm, which multiplies by g^3 instead of g. wit("w", 0, g(1)), @@ -150,7 +204,7 @@ def main(): return ", // n = g^5: five iterations, acc[j] = g^{2j}, so acc[5] = g^10. - valid: Trial::new([g(10), g(5)]).stream("n", vec![vec![g(5)]]), + valid: Trial::new(&[g(10), g(5)]).stream("n", vec![vec![g(5)]]), pokes: vec![ // Fewer and more iterations: acc[n] is then g^8 and g^12. wit("n", 0, g(4)), @@ -163,69 +217,42 @@ def main(): }); } -/// `pack64x2`'s range assertion: both sources must lie in K. Its untaken JUMP -/// puts them in the destination and frame slots, whose memory reads have -/// literal-zero upper limbs. +/// The digest-as-verification idiom: a hinted preimage, hashed, and the result +/// pinned against a hinted digest through a heap run store. This is the shape a +/// signature verifier has, so it is the one that most needs a regression test. #[test] -fn pack64x2_range_assertion() { +fn digest_pins_its_preimage() { + let d = digest_5_7(); check_case(&Case { - name: "pack64x2_range_assertion", + name: "digest_pins_its_preimage", src: "\ -@inline -def pack64x2(a, b): - assert_in_k(a, b) - return a + f192(0, 1, 0) * b - def main(): - v = StackBuf(2) - hint_witness(v, \"w\") - c = pack64x2(v[0], v[1]) + m = StackBuf(8) + hint_witness(m, \"msg\") + d = StackBuf(4) + blake2s(m[0:4], m[4:8], d) + e = HeapBuf(4) + hint_witness(e[0:4], \"dig\") + e[0:4] = d p = GEN ** 0 - p[1] = c - p[GEN] = v[0] + p[1] = m[0] + p[GEN] = m[1] return ", - valid: Trial::new([F192::new(5, 7, 0), k(5)]).stream("w", vec![vec![k(5), k(7)]]), - pokes: vec![ - // Either source outside K. - wit("w", 0, F192::new(5, 1, 0)), - wit("w", 0, F192::new(5, 0, 1)), - wit("w", 1, F192::new(7, 1, 0)), - // In K, but not the packing that was published. - wit("w", 0, k(6)), - wit("w", 1, k(8)), - pi(0, F192::new(5, 8, 0)), - pi(1, k(6)), - ], - }); -} - -/// The digest-as-verification idiom: a hinted preimage, hashed, and the result -/// pinned against a hinted digest through a heap store. This is the shape a -/// signature verifier has, so it is the one that most needs a regression test. -/// -/// The digest constant comes from [`print_blake2s_digest`], not from a hand -/// computation: what the case tests is that a *wrong* digest is rejected, and -/// for that the honest value only has to be honest. -#[test] -fn digest_pins_its_preimage() { - check_case(&Case { - name: "digest_pins_its_preimage", - src: BLAKE2S_PIN_SRC, - valid: Trial::new([k(5), k(7)]) - .stream("msg", vec![vec![k(5), k(7), F192::ZERO, F192::ZERO]]) - .stream("dig", vec![vec![DIGEST_5_7[0], DIGEST_5_7[1]]]), + valid: Trial::new(&[k(5), k(7)]) + .stream("msg", vec![vec![k(5), k(7), k(0), k(0), k(0), k(0), k(0), k(0)]]) + .stream("dig", vec![d.to_vec()]), pokes: vec![ - // A different preimage hashes to something else. + // A different preimage hashes to something else, published or not. wit("msg", 0, k(6)), wit("msg", 1, k(8)), wit("msg", 2, k(1)), - wit("msg", 3, k(1)), + wit("msg", 7, k(1)), // A wrong digest is what the write-once store has to catch. - wit("dig", 0, F192::ZERO), - wit("dig", 1, F192::ZERO), - wit("dig", 0, DIGEST_5_7[0] + F192::ONE), - wit("dig", 1, DIGEST_5_7[1] + F192::ONE), + wit("dig", 0, F64::ZERO), + wit("dig", 3, F64::ZERO), + wit("dig", 1, d[1] + F64::ONE), + wit("dig", 2, d[2] + F64::ONE), // The published preimage words. pi(0, k(6)), pi(1, k(8)), @@ -233,47 +260,11 @@ fn digest_pins_its_preimage() { }); } -const BLAKE2S_PIN_SRC: &str = "\ -def main(): - m = StackBuf(4) - hint_witness(m, \"msg\") - d = StackBuf(2) - blake2s(m[0:2], m[2:4], d) - e = HeapBuf(2) - hint_witness(e[0:2], \"dig\") - e[1] = d[0] - e[GEN] = d[1] - p = GEN ** 0 - p[1] = m[0] - p[GEN] = m[1] - return -"; - -/// BLAKE2s of the 64-byte block whose four canonical cells are `(5, 7, 0, 0)`. -pub const DIGEST_5_7: [F192; 2] = [ - F192::new(0xbbc8_c175_8cb7_7642, 0xf299_5d40_1fad_f4ff, 0), - F192::new(0x83ea_6ade_289a_53c8, 0x57e6_e523_12ec_734b, 0), -]; - -/// Regenerate [`DIGEST_5_7`]: `cargo test --release -p lean_compiler -/// print_blake2s_digest -- --ignored --nocapture`. Kept so the constant above is -/// reproducible rather than folklore. -#[test] -#[ignore = "prints a constant; not a check"] -fn print_blake2s_digest() { - let src = "\ -def main(): - m = StackBuf(4) - hint_witness(m, \"msg\") - d = StackBuf(2) - blake2s(m[0:2], m[2:4], d) - print(d[0]) - print(d[1]) - return -"; - let mut p = super::build(src); - p.set_witness("msg", vec![vec![k(5), k(7), F192::ZERO, F192::ZERO]]); - p.execute([F192::ZERO, F192::ZERO]); +/// BLAKE2s of the 64-byte block whose eight words are `(5, 7, 0, 0, 0, 0, 0, 0)`, +/// from the reference compression: what a case tests is that a *wrong* digest is +/// rejected, and for that the honest value only has to be honest. +pub fn digest_5_7() -> [F64; 4] { + compress([k(5), k(7), k(0), k(0)], [k(0); 4]) } /// The fused `match` path must reject a call that binds more names than @@ -284,9 +275,11 @@ def main(): /// /// Fusion needs every arm to be a call to the same function with identical /// runtime arguments, so the two programs below are the fused shape: one over -/// mixed-arity callees, one over a single over-bound callee. +/// mixed-arity callees, one over a single over-bound callee. The mixed arms are +/// caught by the shared-layout check, which compares the return shapes before any +/// count is read; the over-bound callee has one layout and reaches the count. #[test] -#[should_panic(expected = "dispatched call binds")] +#[should_panic(expected = "does not take and return the same shapes")] fn dispatched_call_rejects_a_mixed_arity_arm() { super::build( "\ @@ -416,37 +409,49 @@ def two(v, k: Const): /// words into a buffer, so hashing one way and the other must agree. /// /// Self-comparing on purpose: an equivalence pair cannot check this, because a -/// trial that must be ACCEPTED has to name the digest and no test should carry a -/// hash constant. Asserting the two digests equal needs no constant, and the two -/// operands are DIFFERENT words so that reordering within a list is visible: with -/// both operands equal the swap would cancel out. +/// trial that must be ACCEPTED has to name the digest, and asserting the two +/// digests equal needs no digest at all. The list spells every chunk shape the +/// lowering distinguishes: a run element whose two words are used in place, a +/// reversed pair that has to be copied, a constant pair, and a pair mixing a cell +/// and a constant. The operands are DIFFERENT words so that reordering within a +/// list is visible. #[test] fn a_blake2s_word_list_hashes_like_the_buffer_it_replaces() { check_case(&Case { name: "a_blake2s_word_list_hashes_like_the_buffer_it_replaces", src: "\ def main(): - v = StackBuf(2) + v = StackBuf(4) hint_witness(v, \"w\") - named = StackBuf(2) - blake2s([v[0], v[1]], [v[1], v[0]], named) - l = StackBuf(2) + named = StackBuf(4) + blake2s([v[0:2], v[3], v[2]], [5, 7, v[1], 9], named) + l = StackBuf(4) l[0] = v[0] l[1] = v[1] - r = StackBuf(2) - r[0] = v[1] - r[1] = v[0] - gathered = StackBuf(2) + l[2] = v[3] + l[3] = v[2] + r = StackBuf(4) + r[0] = 5 + r[1] = 7 + r[2] = v[1] + r[3] = 9 + gathered = StackBuf(4) blake2s(l, r, gathered) assert named[0] == gathered[0] assert named[1] == gathered[1] + assert named[2] == gathered[2] + assert named[3] == gathered[3] p = GEN ** 0 - p[1] = v[0] - p[GEN] = v[1] + p[0:4] = v return ", - valid: Trial::new([k(11), k(22)]).stream("w", vec![vec![k(11), k(22)]]), - pokes: vec![wit("w", 0, k(12)), wit("w", 1, k(23))], + valid: Trial::new(&[k(11), k(22), k(33), k(44)]).stream("w", vec![vec![k(11), k(22), k(33), k(44)]]), + pokes: vec![ + wit("w", 0, k(12)), + wit("w", 1, k(23)), + wit("w", 2, k(34)), + wit("w", 3, k(45)), + ], }); } diff --git a/crates/lean_compiler/tests/suite/soundness/mod.rs b/crates/lean_compiler/tests/suite/soundness/mod.rs index 43566c679..88edb669f 100644 --- a/crates/lean_compiler/tests/suite/soundness/mod.rs +++ b/crates/lean_compiler/tests/suite/soundness/mod.rs @@ -37,35 +37,43 @@ mod cases; mod pairs; /// `g^k` as a machine word, the way every index, address and counter is written. -pub fn g(k: usize) -> F192 { - F192::from(g_pow(k)) +pub fn g(k: usize) -> F64 { + g_pow(k) } -/// A K-valued literal in the low lane. -pub fn k(x: u64) -> F192 { - F192::from(F64(x)) +/// A word given by its bits. +pub fn k(x: u64) -> F64 { + F64(x) +} + +/// The three limb cells of a 192-bit element, low first. +pub fn limbs(x: F192) -> [F64; 3] { + [F64(x.c0), F64(x.c1), F64(x.c2)] } /// One `hint_witness` stream: the name, then one entry per call naming it. -pub type Stream = (&'static str, Vec>); +pub type Stream = (&'static str, Vec>); /// Everything a run consumes: the public statement and the prover's advice. #[derive(Clone)] pub struct Trial { - pub pi: [F192; 2], + pub pi: [F64; 4], pub streams: Vec, } impl Trial { - pub fn new(pi: [F192; 2]) -> Self { + /// A trial publishing `pi`, zero-padded to the four public words. + pub fn new(pi: &[F64]) -> Self { + let mut words = [F64::ZERO; 4]; + words[..pi.len()].copy_from_slice(pi); Self { - pi, + pi: words, streams: Vec::new(), } } /// Add a stream whose every call takes one entry of `cells`. - pub fn stream(mut self, name: &'static str, entries: Vec>) -> Self { + pub fn stream(mut self, name: &'static str, entries: Vec>) -> Self { self.streams.push((name, entries)); self } @@ -91,35 +99,33 @@ impl Trial { /// exactly the constraint that is missing. #[derive(Clone, Copy)] pub enum Poke { - /// Public-input word 0 or 1. - Pi { slot: usize, to: F192 }, + /// Public-input word 0 to 3. + Pi { slot: usize, to: F64 }, /// Cell `cell` of entry `entry` of witness stream `name`. Wit { name: &'static str, entry: usize, cell: usize, - to: F192, + to: F64, }, } impl Poke { fn label(&self) -> String { match self { - Poke::Pi { slot, to } => format!("pi[{slot}] := {:x}:{:x}:{:x}", to.c2, to.c1, to.c0), - Poke::Wit { name, entry, cell, to } => { - format!("{name}[{entry}][{cell}] := {:x}:{:x}:{:x}", to.c2, to.c1, to.c0) - } + Poke::Pi { slot, to } => format!("pi[{slot}] := {:#x}", to.0), + Poke::Wit { name, entry, cell, to } => format!("{name}[{entry}][{cell}] := {:#x}", to.0), } } } /// Poke a public-input word. -pub fn pi(slot: usize, to: F192) -> Poke { +pub fn pi(slot: usize, to: F64) -> Poke { Poke::Pi { slot, to } } /// Poke cell `cell` of the first entry of stream `name`. -pub fn wit(name: &'static str, cell: usize, to: F192) -> Poke { +pub fn wit(name: &'static str, cell: usize, to: F64) -> Poke { Poke::Wit { name, entry: 0, @@ -129,7 +135,7 @@ pub fn wit(name: &'static str, cell: usize, to: F192) -> Poke { } /// Poke cell `cell` of entry `entry` of stream `name`. -pub fn wit_at(name: &'static str, entry: usize, cell: usize, to: F192) -> Poke { +pub fn wit_at(name: &'static str, entry: usize, cell: usize, to: F64) -> Poke { Poke::Wit { name, entry, cell, to } } diff --git a/crates/lean_compiler/tests/suite/soundness/pairs.rs b/crates/lean_compiler/tests/suite/soundness/pairs.rs index 2f21cd9bb..eab2ce51c 100644 --- a/crates/lean_compiler/tests/suite/soundness/pairs.rs +++ b/crates/lean_compiler/tests/suite/soundness/pairs.rs @@ -11,12 +11,12 @@ //! Every pair here is a promise `zkDSL.md` makes. When one fails, quote the //! promise in the bug report; the `why` field is there to be quoted. -use super::{Pair, Trial, check_pair, g, k}; -use primitives::field::F192; +use super::{Pair, Trial, check_pair, g, k, limbs}; +use primitives::field::{F64, F192}; /// One hinted pair of cells, published so the trial's public input pins them. -fn two(a: F192, b: F192) -> Trial { - Trial::new([a, b]).stream("w", vec![vec![a, b]]) +fn two(a: F64, b: F64) -> Trial { + Trial::new(&[a, b]).stream("w", vec![vec![a, b]]) } /// `@inline` is documented as a pure call-site expansion: "the body is inlined at @@ -53,6 +53,45 @@ def shift(x): }); } +/// The same promise for a function taking and returning a 192-bit run: the plain +/// call passes the run through its argument cells and copies the result back out +/// of its return slots, while the inlined one binds both in place. +#[test] +fn inline_and_plain_run_calls_agree() { + let body = "\ +def main(): + v = StackBuf(6) + hint_witness(v, \"w\") + assert_eq192(square(v[0:3]), v[3:6]) + p = GEN ** 0 + p[0:3] = v[0:3] + p[GEN ** 3] = v[3] + return + + +@INLINE +def square(x: StackBuf(3)): + return mul192(x, x) +"; + let x = F192::new(g(3).0, 5, 7); + let trial = |x: F192, y: F192| { + let [x0, x1, x2] = limbs(x); + Trial::new(&[x0, x1, x2, F64(y.c0)]).stream("w", vec![[limbs(x), limbs(y)].concat()]) + }; + check_pair(&Pair { + name: "inline_and_plain_run_calls_agree", + why: "zkDSL.md §`@inline`: inlining is a call-site expansion, not a change of meaning.", + a: &body.replace("@INLINE\n", "@inline\n"), + b: &body.replace("@INLINE\n", ""), + trials: vec![ + trial(x, x * x), + trial(x, x * x + F192::Y), + trial(F192::ZERO, F192::ZERO), + trial(x, x), + ], + }); +} + /// Write-once memory is the assertion mechanism, so `assert a == b` and two /// stores of `a` and `b` into one heap cell are the same statement. `zkDSL.md` /// §Memory: "a second write of the same value is a no-op, of a different value a @@ -87,8 +126,8 @@ def main(): trials: vec![ two(g(3), g(3)), two(g(3), g(4)), - two(F192::ZERO, F192::ZERO), - two(F192::ZERO, k(1)), + two(F64::ZERO, F64::ZERO), + two(F64::ZERO, k(1)), ], }); } @@ -133,8 +172,54 @@ def main(): }); } -fn three(a: F192, b: F192, c: F192) -> Trial { - Trial::new([a, c]).stream("w", vec![vec![a, b, c]]) +fn three(a: F64, b: F64, c: F64) -> Trial { + Trial::new(&[a, c]).stream("w", vec![vec![a, b, c]]) +} + +/// `div192` is the same promise over 192-bit runs: the quotient run is left +/// unset and `MUL192` checks `quotient · b == a`. So dividing and comparing must +/// accept exactly what comparing the product accepts, for a nonzero divisor. +#[test] +fn division192_and_checked_product_agree() { + let trial = |divisor: F192, dividend: F192, claimed: F192| { + let [d0, d1, d2] = limbs(divisor); + Trial::new(&[d0, d1, d2, F64(claimed.c0)]) + .stream("w", vec![[limbs(divisor), limbs(dividend), limbs(claimed)].concat()]) + }; + let x = F192::new(g(3).0, 5, 7); + let y = F192::new(11, g(9).0, 13); + check_pair(&Pair { + name: "division192_and_checked_product_agree", + why: "`div192(a, b)` emits exactly the relation `quotient · b == a`.", + a: "\ +def main(): + v = StackBuf(9) + hint_witness(v, \"w\") + q = div192(v[3:6], v[0:3]) + assert_eq192(q, v[6:9]) + p = GEN ** 0 + p[0:3] = v[0:3] + p[GEN ** 3] = v[6] + return +", + b: "\ +def main(): + v = StackBuf(9) + hint_witness(v, \"w\") + assert_eq192(mul192(v[6:9], v[0:3]), v[3:6]) + p = GEN ** 0 + p[0:3] = v[0:3] + p[GEN ** 3] = v[6] + return +", + trials: vec![ + trial(x, x * y, y), + trial(x, x * y, y + F192::Y), + trial(F192::ONE, y, y), + trial(x, x, F192::ONE), + trial(x, x * y, x), + ], + }); } /// `unroll` is documented as compile-time unrolling, so a loop and its expansion @@ -183,8 +268,8 @@ def main(): } /// A published pair whose first word is the claim and whose second is the hint. -fn one(published: F192, hint: F192) -> Trial { - Trial::new([published, hint]).stream("w", vec![vec![hint]]) +fn one(published: F64, hint: F64) -> Trial { + Trial::new(&[published, hint]).stream("w", vec![vec![hint]]) } /// An `@inline` function returning a one-cell `StackBuf`, used in expression @@ -271,7 +356,7 @@ def main(): trials: vec![ pinned(g(3), g(9)), // the hint agrees with the pin pinned(g(4), g(9)), // it does not: both must reject - pinned(F192::ZERO, g(1)), + pinned(F64::ZERO, g(1)), pinned(g(2), g(3)), ], }); @@ -282,8 +367,68 @@ def main(): /// forwards through the very alias that dropped it and both spellings agree by /// accident. Publishing the pin makes a dropped pin visible as a program that /// accepts every hint. -fn pinned(hint0: F192, hint1: F192) -> Trial { - Trial::new([g(3), hint1]).stream("w", vec![vec![hint0, hint1]]) +fn pinned(hint0: F64, hint1: F64) -> Trial { + Trial::new(&[g(3), hint1]).stream("w", vec![vec![hint0, hint1]]) +} + +/// The run form of the pinning promise: a 192-bit value stored into a hinted run +/// asserts equality exactly as `assert_eq192` does, whether the value lands in +/// place (`s[0:3] = value`), through a bound copy (`t = value`, then +/// `s[0:3] = t`), or through a heap run already holding the hint. A constant +/// value and a computed one lower differently (limb `SET`s against an instruction +/// writing its result), so both are compared. +/// +/// As for [`pinned`], the trials publish the pin's input `x`, never the hint. +#[test] +fn a_run_store_pins_a_hint_like_assert_eq192() { + let body = "\ +def main(): + s = StackBuf(3) + hint_witness(s, \"s\") + x = StackBuf(3) + hint_witness(x, \"x\") + STORE + p = GEN ** 0 + p[0:3] = x + return +"; + let trial = |s: F192, x: F192| { + Trial::new(&limbs(x)) + .stream("s", vec![limbs(s).to_vec()]) + .stream("x", vec![limbs(x).to_vec()]) + }; + let c = F192::new(3, 5, 7); + let x = F192::new(g(4).0, 9, 1); + let computed = vec![ + trial(x * c, x), + trial(x * c + F192::ONE, x), + trial(F192::ZERO, F192::ZERO), + trial(x, x), + ]; + let constant = vec![ + trial(c, x), + trial(c + F192::Y, x), + trial(c, F192::ZERO), + trial(F192::ZERO, x), + ]; + for (value, trials) in [("mul192(x, f192(3, 5, 7))", computed), ("f192(3, 5, 7)", constant)] { + let spell = |store: &str| body.replace("STORE", &store.replace("VALUE", value)); + let asserted = spell("assert_eq192(s, VALUE)"); + for store in [ + "s[0:3] = VALUE", + "t = VALUE\n s[0:3] = t", + "h = HeapBuf(3)\n h[0:3] = s\n h[0:3] = VALUE", + ] { + check_pair(&Pair { + name: "a_run_store_pins_a_hint_like_assert_eq192", + why: "zkDSL.md §Memory: a store into already-written cells IS an equality assertion, \ + for a run as for one cell.", + a: &asserted, + b: &spell(store), + trials: trials.clone(), + }); + } + } } /// `zkDSL.md` §BLAKE2s: "If `out` was already written, the statement *asserts* @@ -298,52 +443,48 @@ fn pinned(hint0: F192, hint1: F192) -> Trial { /// prover could put any message under the hash. #[test] fn prewritten_blake2s_out_asserts_the_digest() { + let four = |w: [F64; 4]| Trial::new(&w).stream("w", vec![w.to_vec()]); + let d = super::cases::digest_5_7(); check_pair(&Pair { name: "prewritten_blake2s_out_asserts_the_digest", why: "zkDSL.md §BLAKE2s: a pre-written `out` turns the hash into a verification.", a: "\ def main(): - v = StackBuf(2) + v = StackBuf(4) hint_witness(v, \"w\") - m = StackBuf(4) - m[0] = 5 - m[1] = 7 - m[2] = 0 - m[3] = 0 - d = StackBuf(2) + m = [5, 7, 0, 0, 0, 0, 0, 0] + d = StackBuf(4) d[0] = v[0] d[1] = v[1] - blake2s(m[0:2], m[2:4], d) + d[2] = v[2] + d[3] = v[3] + blake2s(m[0:4], m[4:8], d) p = GEN ** 0 - p[1] = v[0] - p[GEN] = v[1] + p[0:4] = v return ", b: "\ def main(): - v = StackBuf(2) + v = StackBuf(4) hint_witness(v, \"w\") - m = StackBuf(4) - m[0] = 5 - m[1] = 7 - m[2] = 0 - m[3] = 0 - d = HeapBuf(2) + m = [5, 7, 0, 0, 0, 0, 0, 0] + d = HeapBuf(4) d[1] = v[0] d[GEN] = v[1] - blake2s(m[0:2], m[2:4], d[0:2]) + d[GEN ** 2] = v[2] + d[GEN ** 3] = v[3] + blake2s(m[0:4], m[4:8], d[0:4]) p = GEN ** 0 - p[1] = v[0] - p[GEN] = v[1] + p[0:4] = v return ", trials: vec![ - // The real digest of the block whose cells are (5, 7, 0, 0). - two(super::cases::DIGEST_5_7[0], super::cases::DIGEST_5_7[1]), + // The real digest of the block whose words are (5, 7, 0, ..., 0). + four(d), // Anything else must be rejected by both spellings. - two(F192::ZERO, F192::ZERO), - two(super::cases::DIGEST_5_7[0], F192::ZERO), - two(g(3), g(5)), + four([F64::ZERO; 4]), + four([d[0], d[1], d[2], F64::ZERO]), + four([g(3), g(5), g(7), g(9)]), ], }); } @@ -446,8 +587,8 @@ def main(): /// A trial for the branch pair: publishes `s[0]` and `v[0]`, which the assertion /// makes equal on every accepting path. -fn branch3(a: F192, b: F192, c: F192) -> Trial { - Trial::new([a, a]).stream("w", vec![vec![a, b, c]]) +fn branch3(a: F64, b: F64, c: F64) -> Trial { + Trial::new(&[a, a]).stream("w", vec![vec![a, b, c]]) } /// A multi-value target may be a `StackBuf` element, and `zkDSL.md` §match @@ -486,10 +627,10 @@ def main(): "t, e = match(log(v[0]), range(0, 2), lambda i: pick(v[1], i))\n sb[0] = t", ), trials: vec![ - Trial::new([g(4), g(0)]).stream("w", vec![vec![g(0), g(4)]]), - Trial::new([g(5), g(1)]).stream("w", vec![vec![g(1), g(4)]]), - Trial::new([g(4), g(1)]).stream("w", vec![vec![g(1), g(4)]]), - Trial::new([g(9), g(0)]).stream("w", vec![vec![g(0), g(4)]]), + Trial::new(&[g(4), g(0)]).stream("w", vec![vec![g(0), g(4)]]), + Trial::new(&[g(5), g(1)]).stream("w", vec![vec![g(1), g(4)]]), + Trial::new(&[g(4), g(1)]).stream("w", vec![vec![g(1), g(4)]]), + Trial::new(&[g(9), g(0)]).stream("w", vec![vec![g(0), g(4)]]), ], }); } diff --git a/crates/lean_compiler/tests/suite/stack_bits.rs b/crates/lean_compiler/tests/suite/stack_bits.rs index a7b1f73b5..d91221967 100644 --- a/crates/lean_compiler/tests/suite/stack_bits.rs +++ b/crates/lean_compiler/tests/suite/stack_bits.rs @@ -4,7 +4,9 @@ use lean_compiler::{compile, parse}; use lean_vm::cpu::{Stats, prove, verify}; -use primitives::field::{F64, F192}; +use primitives::field::{F64, g_pow}; + +use crate::common::pi; const V: u64 = 0b1011_0110; @@ -36,7 +38,7 @@ def main(): " ) }; - let want = [F192::from(F64(V)), F192::from(F64::ONE)]; + let want = pi(&[F64(V), F64::ONE]); let deref = |s: &str| crate::common::mix(s, want)[deref_index()]; assert_eq!( deref(&src("HeapBuf(GEN ** 8)", "GEN ** i")) - deref(&src("StackBuf(8)", "i")), @@ -48,8 +50,8 @@ def main(): /// A stack bit run is addressed by CONTIGUITY, so no cell of one may be given /// away to a duplicate elsewhere: `hint_log2_ceil` reads `fp+base+k` whatever the /// lowerer decided, so a dropped store would leave it holding nothing. The -/// duplicate `MUL` here comes FIRST, which is the order that would make the store -/// the one dropped. +/// duplicate `MUL64` here comes FIRST, which is the order that would make the +/// store the one dropped. #[test] fn a_stack_bit_run_survives_cell_sharing() { let src = "\ @@ -71,10 +73,8 @@ def main(): // store leaves its cell unwritten, and the advice then computed off the hole // collides with the published public input. let mut program = compile(&parse(src).expect("parse")); - let bits: Vec = [1u64, 1, 0, 1].iter().map(|&b| F192::from(F64(b))).collect(); - program.set_witness("bits", vec![bits]); - let want = [F192::from(primitives::field::g_pow(4)), F192::from(F64::ONE)]; - let exec = program.execute(want); + program.set_witness("bits", vec![[1u64, 1, 0, 1].map(F64).to_vec()]); + let exec = program.execute(pi(&[g_pow(4), F64::ONE])); assert!( exec.unconstrained_reads.is_empty(), "every cell of the run must still be written" @@ -117,11 +117,11 @@ def probe(v): return total "; let mut program = compile(&parse(src).expect("parse")); - program.set_witness("v", vec![vec![F192::from(F64(V))]]); - let want = [F192::from(F64(V)), F192::from(F64::ONE)]; + program.set_witness("v", vec![vec![F64(V)]]); + let want = pi(&[F64(V), F64::ONE]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("the pointer reads the frame run"); - let bad = [F192::from(F64(V + 1)), F192::from(F64::ONE)]; + let bad = pi(&[F64(V + 1), F64::ONE]); assert!(verify(&program, &bad, &proof).is_err(), "a wrong value is rejected"); } @@ -143,7 +143,7 @@ def main(): b[0] = GEN ** 9 return "; - compile(&parse(src).expect("parse")).execute([F192::ZERO; 2]); + compile(&parse(src).expect("parse")).execute([F64::ZERO; 4]); } /// The same hazard for a run declared AFTER the escape, which the test above @@ -168,7 +168,7 @@ def main(): assert b[0] == GEN ** 9 return "; - compile(&parse(src).expect("parse")).execute([F192::ZERO; 2]); + compile(&parse(src).expect("parse")).execute([F64::ZERO; 4]); } /// A frame pointer carries the same compile-time bound a `HeapBuf` pointer gets. @@ -209,5 +209,5 @@ def main(): assert b[0] == GEN ** 9 return "; - compile(&parse(src).expect("parse")).execute([F192::ZERO; 2]); + compile(&parse(src).expect("parse")).execute([F64::ZERO; 4]); } diff --git a/crates/lean_compiler/tests/suite/stack_buf.rs b/crates/lean_compiler/tests/suite/stack_buf.rs index 583adc894..a9cc68dbf 100644 --- a/crates/lean_compiler/tests/suite/stack_buf.rs +++ b/crates/lean_compiler/tests/suite/stack_buf.rs @@ -1,88 +1,79 @@ //! `StackBuf`: a run of consecutive frame (stack) cells in the zkDSL. Indexed -//! reads/writes go straight to `base+k` (no heap deref), and a size-2 `StackBuf` -//! is a `blake2s` operand: its two canonical 128-bit cells hold the 256-bit value, so +//! reads/writes go straight to `base+k` (no heap deref), and a size-4 `StackBuf` +//! is a `blake2s` operand: its four 64-bit cells hold the 256-bit value, so //! `blake2s(a, b, out)` reads them in place with no copies (a self-hash -//! `blake2s(h, h, out)` aliases one pair into both input operands) and writes -//! the digest into the pre-allocated pair `out`. -//! -//! Since these DSL scalars are K-embedded F192 cells, a `StackBuf(2)` written -//! cell-by-cell holds the flock words `[v0, 0, v1, 0]` -//!: the reference `compress` is fed that lane layout. +//! `blake2s(h, h, out)` aliases one run into both input operands) and writes +//! the digest into the pre-allocated run `out`. use lean_compiler::{compile, parse}; use lean_vm::cpu::{Op, prove, verify}; use lean_vm::hash_flock::{compression, digest, metadata, unpack_metadata}; use lean_vm::vmhash::compress; -use primitives::field::{F64, F192, g_pow}; +use primitives::field::{F64, g_pow}; -use crate::common::mix; +use crate::common::{mix, panic_message, pi}; -/// The two 128-bit digest cells of `compress(a, b)` as `F192`s (lo = word 0/2, -/// hi = word 1/3): what a `blake2s(...)` output `StackBuf(2)` holds cell-by-cell. -fn digest_cells(a: [F64; 4], b: [F64; 4]) -> [F192; 2] { - let d = compress(a, b); - [F192::new(d[0].0, d[1].0, 0), F192::new(d[2].0, d[3].0, 0)] -} - -/// A size-2 `StackBuf` fed to `blake2s` as a self-hash `blake2s(h, h)`, then the -/// digest's two 128-bit cells published to `m[0], m[1]`. Proves and verifies, and -/// a wrong published digest is rejected: so the whole path (StackBuf load → -/// aliased blake2s → stack read → publish) is exercised end-to-end. +/// A size-4 `StackBuf` fed to `blake2s` as a self-hash `blake2s(h, h)`, then the +/// digest's four words published to `m[0..4]`. Proves and verifies, and a wrong +/// published digest is rejected: so the whole path (StackBuf load → aliased +/// blake2s → heap run store → publish) is exercised end-to-end. #[test] fn stack_buf_blake2s_self_hash() { let src = "\ def main(): - a = StackBuf(2) + a = StackBuf(4) a[0] = 5 - a[1] = 7 - c = StackBuf(2) + a[1] = 0 + a[2] = 7 + a[3] = 0 + c = StackBuf(4) blake2s(a, a, c) - p = 1 - p[1] = c[0] - p[GEN] = c[1] + p = GEN ** 0 + p[0:4] = c return "; let program = compile(&parse(src).expect("parse")); - - // Each cell holds one scalar in its low lane, so the hashed words are [5,0,7,0]. - let h = [F64(5), F64(0), F64(7), F64(0)]; - let want = digest_cells(h, h); + let h = [5, 0, 7, 0].map(F64); + let want = compress(h, h); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); assert_eq!(mix(src, want)[5], 1, "one BLAKE2s instruction"); verify(&program, &want, &proof).expect("StackBuf self-hash verifies"); let mut bad = want; - bad[0] += F192::ONE; + bad[0] += F64::ONE; assert!(verify(&program, &bad, &proof).is_err(), "wrong digest must be rejected"); } +/// The standard BLAKE2s of the 80 bytes that the words `1, 0, 2, 0, 3, 0, 4, 0, 5, 0` +/// spell little-endian. +fn standard_80_byte_digest() -> [F64; 4] { + let input: Vec = [1u64, 0, 2, 0, 3, 0, 4, 0, 5, 0] + .iter() + .flat_map(|w| w.to_le_bytes()) + .collect(); + let d = primitives::hash::hash(&input); + std::array::from_fn(|i| F64(u64::from_le_bytes(d[8 * i..8 * i + 8].try_into().unwrap()))) +} + /// Optional BLAKE2s metadata and a memory-supplied chaining value reproduce a /// standard two-block (80-byte) BLAKE2s hash. #[test] fn blake2s_keywords_standard_multiblock() { let src = "\ def main(): - block0 = [1, 2, 3, 4] - tail = [5, 0, 0, 0] - cv = StackBuf(2) - blake2s(block0[0:2], block0[2:4], cv, counter=64, final=0) - out = StackBuf(2) - blake2s(tail[0:2], tail[2:4], out, cv=cv, counter=80, final=1) - p = 1 - p[1] = out[0] - p[GEN] = out[1] + block0 = [1, 0, 2, 0, 3, 0, 4, 0] + tail = [5, 0, 0, 0, 0, 0, 0, 0] + cv = StackBuf(4) + blake2s(block0[0:4], block0[4:8], cv, counter=64, final=0) + out = StackBuf(4) + blake2s(tail[0:4], tail[4:8], out, cv=cv, counter=80, final=1) + p = GEN ** 0 + p[0:4] = out return "; let program = compile(&parse(src).expect("parse")); - let mut input = Vec::new(); - for value in 1u64..=5 { - input.extend_from_slice(&value.to_le_bytes()); - input.extend_from_slice(&0u64.to_le_bytes()); - } - let d = primitives::hash::hash(&input); - let word = |o: usize| u64::from_le_bytes(d[o..o + 8].try_into().unwrap()); - let want = [F192::new(word(0), word(8), 0), F192::new(word(16), word(24), 0)]; + let want = standard_80_byte_digest(); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); assert_eq!(mix(src, want)[5], 2); verify(&program, &want, &proof).expect("standard two-block BLAKE2s verifies"); @@ -91,53 +82,45 @@ def main(): /// The same 80-byte hash with its second block's metadata computed at run time, /// the shape a hash of runtime length needs: the counter's high part is a word /// the program produced and its low part a compile-time constant, and their set -/// bits are disjoint, so one `XOR` is their integer sum (doc +/// bits are disjoint, so one `XOR64` is their integer sum (doc /// §sec:prog-byte-counter). The hint stands in for the high part a real absorb /// loop derives from its own counter. #[test] fn blake2s_runtime_metadata_matches_the_standard_hash() { let src = "\ def main(): - block0 = [1, 2, 3, 4] - tail = [5, 0, 0, 0] - cv = StackBuf(2) - blake2s(block0[0:2], block0[2:4], cv, counter=64, final=0) + block0 = [1, 0, 2, 0, 3, 0, 4, 0] + tail = [5, 0, 0, 0, 0, 0, 0, 0] + cv = StackBuf(4) + blake2s(block0[0:4], block0[4:8], cv, counter=64, final=0) high = hint_witness(\"high\") assert high == 64 - out = StackBuf(2) - blake2s(tail[0:2], tail[2:4], out, cv=cv, md=high + f192(16, 4294967295, 0)) - p = 1 - p[1] = out[0] - p[GEN] = out[1] + out = StackBuf(4) + blake2s(tail[0:4], tail[4:8], out, cv=cv, md=[high + 16, 4294967295]) + p = GEN ** 0 + p[0:4] = out return "; let mut program = compile(&parse(src).expect("parse")); - program.set_witness("high", vec![vec![F192::new(64, 0, 0)]]); - let mut input = Vec::new(); - for value in 1u64..=5 { - input.extend_from_slice(&value.to_le_bytes()); - input.extend_from_slice(&0u64.to_le_bytes()); - } - let d = primitives::hash::hash(&input); - let word = |o: usize| u64::from_le_bytes(d[o..o + 8].try_into().unwrap()); - let want = [F192::new(word(0), word(8), 0), F192::new(word(16), word(24), 0)]; + program.set_witness("high", vec![vec![F64(64)]]); + let want = standard_80_byte_digest(); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); - verify(&program, &want, &proof).expect("a runtime metadata word hashes to the standard digest"); + verify(&program, &want, &proof).expect("a runtime metadata hashes to the standard digest"); } #[test] fn blake2s_counter_accepts_full_u64_range() { let src = "\ def main(): - block = [1, 2, 3, 4] - out = StackBuf(2) + block = [1, 0, 2, 0, 3, 0, 4, 0] + out = StackBuf(4) counter = 18446744073709551615 // 1 - blake2s(block[0:2], block[2:4], out, counter=counter, final=1) + blake2s(block[0:4], block[4:8], out, counter=counter, final=1) return "; let program = compile(&parse(src).expect("parse")); // The metadata is a memory operand, so what carries the counter is the `SET` - // immediate that wrote the cell the instruction reads. + // immediates that wrote the two cells the instruction reads. let md = program .prog .iter() @@ -146,15 +129,17 @@ def main(): _ => None, }) .expect("BLAKE2s instruction"); - let metadata = program - .prog - .iter() - .find_map(|op| match op { - Op::Set { o, k } if *o == md => Some(*k), - _ => None, - }) - .expect("the metadata cell's SET"); - assert_eq!(unpack_metadata(metadata), (u64::MAX, u32::MAX, 0)); + let set_of = |cell: u32| { + program + .prog + .iter() + .find_map(|op| match op { + Op::Set { o, k } if *o == cell => Some(*k), + _ => None, + }) + .expect("the metadata cell's SET") + }; + assert_eq!(unpack_metadata([set_of(md), set_of(md + 1)]), (u64::MAX, u32::MAX, 0)); } #[test] @@ -162,9 +147,9 @@ def main(): fn blake2s_counter_rejects_values_above_u64() { let src = "\ def main(): - block = [1, 2, 3, 4] - out = StackBuf(2) - blake2s(block[0:2], block[2:4], out, counter=18446744073709551616, final=1) + block = [1, 0, 2, 0, 3, 0, 4, 0] + out = StackBuf(4) + blake2s(block[0:4], block[4:8], out, counter=18446744073709551616, final=1) return "; compile(&parse(src).expect("parse")); @@ -179,21 +164,20 @@ fn blake2s_default_iv_after_runtime_branch() { def main(): flag = StackBuf(1) hint_witness(flag, \"flag\") - a = [1, 2, 3, 4] + a = [1, 0, 2, 0, 3, 0, 4, 0] if flag[0] == 1: - ignored = StackBuf(2) - blake2s(a[0:2], a[2:4], ignored) - out = StackBuf(2) - blake2s(a[0:2], a[2:4], out) - p = 1 - p[1] = out[0] - p[GEN] = out[1] + ignored = StackBuf(4) + blake2s(a[0:4], a[4:8], ignored) + out = StackBuf(4) + blake2s(a[0:4], a[4:8], out) + p = GEN ** 0 + p[0:4] = out return "; - let want = digest_cells([F64(1), F64(0), F64(2), F64(0)], [F64(3), F64(0), F64(4), F64(0)]); + let want = compress([1, 0, 2, 0].map(F64), [3, 0, 4, 0].map(F64)); for flag in [0, 1] { let mut program = compile(&parse(src).expect("parse")); - program.set_witness("flag", vec![vec![F192::new(flag, 0, 0)]]); + program.set_witness("flag", vec![vec![F64(flag)]]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("post-join default IV is initialized on both paths"); } @@ -207,73 +191,70 @@ fn blake2s_default_iv_in_both_runtime_branches() { def main(): flag = StackBuf(1) hint_witness(flag, \"flag\") - a = [1, 2, 3, 4] - out = StackBuf(2) + a = [1, 0, 2, 0, 3, 0, 4, 0] + out = StackBuf(4) if flag[0] == 1: - blake2s(a[0:2], a[2:4], out) + blake2s(a[0:4], a[4:8], out) else: - blake2s(a[0:2], a[2:4], out) - p = 1 - p[1] = out[0] - p[GEN] = out[1] + blake2s(a[0:4], a[4:8], out) + p = GEN ** 0 + p[0:4] = out return "; - let want = digest_cells([F64(1), F64(0), F64(2), F64(0)], [F64(3), F64(0), F64(4), F64(0)]); + let want = compress([1, 0, 2, 0].map(F64), [3, 0, 4, 0].map(F64)); for flag in [0, 1] { let mut program = compile(&parse(src).expect("parse")); - program.set_witness("flag", vec![vec![F192::new(flag, 0, 0)]]); + program.set_witness("flag", vec![vec![F64(flag)]]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("each branch initializes its default IV"); } } /// Deferred aliases may expose non-adjacent source words for a syntactically -/// consecutive CV StackBuf. The compiler must materialize that pair because +/// consecutive CV StackBuf. The compiler must materialize that run because /// the BLAKE2s opcode carries only one CV base offset. #[test] -fn blake2s_materializes_aliased_cv_pair() { +fn blake2s_materializes_aliased_cv_run() { let src = "\ def main(): - msg = [1, 2, 3, 4] - sources = [5, 99, 6] - cv = [sources[0], sources[2]] - out = StackBuf(2) - blake2s(msg[0:2], msg[2:4], out, cv=cv, counter=128) - p = 1 - p[1] = out[0] - p[GEN] = out[1] + msg = [1, 0, 2, 0, 3, 0, 4, 0] + sources = [5, 99, 6, 99, 7, 99, 8] + cv = [sources[0], sources[2], sources[4], sources[6]] + out = StackBuf(4) + blake2s(msg[0:4], msg[4:8], out, cv=cv, counter=128) + p = GEN ** 0 + p[0:4] = out return "; let program = compile(&parse(src).expect("parse")); let block = compression( - [F64(1), F64(0), F64(2), F64(0)], - [F64(3), F64(0), F64(4), F64(0)], - [F64(5), F64(0), F64(6), F64(0)], + [1, 0, 2, 0].map(F64), + [3, 0, 4, 0].map(F64), + [5, 6, 7, 8].map(F64), metadata(128, 0, 0), ); - let d = digest(&block); - let want = [F192::new(d[0].0, d[1].0, 0), F192::new(d[2].0, d[3].0, 0)]; + let want = digest(&block); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("materialized custom CV verifies"); } -/// A custom CV with the default one-block metadata is not a chained block. /// A metadata cell inside the digest destination would be read before the digest /// is stored and re-read from the finished image by the witness, so the two would /// disagree and the proof would fail its opening with nothing to point at. Every -/// other overlap is a write-once conflict, which does say where it happened. +/// other overlap is a write-once conflict, which does say where it happened. A +/// partial overlap is enough. #[test] #[should_panic(expected = "md= must not name a cell of the digest destination")] fn blake2s_metadata_inside_the_destination_is_rejected() { let src = "\ def main(): - msg = [1, 2, 3, 4] - out = StackBuf(2) - out[0] = 7 - blake2s(msg[0:2], msg[2:4], out, md=out[0]) + msg = [1, 0, 2, 0, 3, 0, 4, 0] + out = StackBuf(4) + out[1] = 7 + blake2s(msg[0:4], msg[4:8], out, md=out[1:3]) return "; - let _ = compile(&parse(src).expect("parse")); + compile(&parse(src).expect("parse")); } /// Require the caller to state the byte counter explicitly. @@ -282,13 +263,13 @@ def main(): fn blake2s_cv_alone_is_rejected() { let src = "\ def main(): - msg = [1, 2, 3, 4] - cv = [5, 6] - out = StackBuf(2) - blake2s(msg[0:2], msg[2:4], out, cv=cv) + msg = [1, 0, 2, 0, 3, 0, 4, 0] + cv = [5, 6, 7, 8] + out = StackBuf(4) + blake2s(msg[0:4], msg[4:8], out, cv=cv) return "; - let _ = compile(&parse(src).expect("parse")); + compile(&parse(src).expect("parse")); } /// A general (non-blake2s) `StackBuf(3)`: indexed writes, an indexed read feeding @@ -309,7 +290,7 @@ def main(): "; let program = compile(&parse(src).expect("parse")); // `+` is XOR: 3 ^ 4 = 7. Published: (sa[2], sa[1]) = (7, 4). - let want = [F192::from(F64(7)), F192::from(F64(4))]; + let want = pi(&[F64(7), F64(4)]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); assert_eq!(mix(src, want)[5], 0, "no BLAKE2s here"); verify(&program, &want, &proof).expect("StackBuf indexing verifies"); @@ -341,7 +322,7 @@ def make(v): "; let program = compile(&parse(src).expect("parse")); // Field addition is XOR: 5 ^ (5 ^ 3) == 3. - program.execute([F192::from(F64(3)), F192::from(F64(11))]); + program.execute(pi(&[F64(3), F64(11)])); } /// Tuple returns retain their source-level arity even though a StackBuf member @@ -361,7 +342,7 @@ def make(v): return out, v + 1 "; let program = compile(&parse(src).expect("parse")); - program.execute([F192::from(F64(15)), F192::from(F64(8))]); + program.execute(pi(&[F64(15), F64(8)])); } /// HeapBuf already crosses a normal call as its one-cell pointer. Allocation @@ -384,7 +365,7 @@ def make(): return out "; let program = compile(&parse(src).expect("parse")); - program.execute([F192::from(F64(17)), F192::from(F64(23))]); + program.execute(pi(&[F64(17), F64(23)])); } /// A StackBuf index literal that does not fit `u32` is rejected at compile time, @@ -403,7 +384,7 @@ fn stack_buf_index_overflow_rejected() { fn stack_buf_rebind_to_scalar() { let src = "def main():\n x = StackBuf(2)\n x = 5\n p = 1\n p[1] = x\n p[GEN] = x\n return\n"; let program = compile(&parse(src).expect("parse")); - let want = [F192::from(F64(5)), F192::from(F64(5))]; + let want = pi(&[F64(5), F64(5)]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("rebound-scalar program verifies"); } @@ -427,40 +408,41 @@ fn stack_buf_loop_capture_rejected() { fn inline_returns_stackbuf_and_scalar() { let src = "\ def main(): - s = StackBuf(2) + s = StackBuf(4) s[0] = 5 - s[1] = 7 + s[1] = 0 + s[2] = 7 + s[3] = 0 s, x = step(s, 9) s, y = step(s, x) - p = 1 - p[1] = s[0] - p[GEN] = s[1] + p = GEN ** 0 + p[0:4] = s return @inline def step(state, v): - tg = StackBuf(2) + tg = StackBuf(4) tg[0] = v - tg[1] = 3 - nb = StackBuf(2) + tg[1] = 0 + tg[2] = 3 + tg[3] = 0 + nb = StackBuf(4) blake2s(state, tg, nb) return nb, v "; let program = compile(&parse(src).expect("parse")); - // Each cell = one scalar in its low lane, so a StackBuf(2) hashes words - // [c0, 0, c1, 0]. x == v == 9 (the scalar return), so both steps use tag 9. - let tag = [F64(9), F64(0), F64(3), F64(0)]; - let s1 = compress([F64(5), F64(0), F64(7), F64(0)], tag); - let s2 = compress(s1, tag); // the returned StackBuf (holding s1's words) fed back in - let want = [F192::new(s2[0].0, s2[1].0, 0), F192::new(s2[2].0, s2[3].0, 0)]; + // x == v == 9 (the scalar return), so both steps use tag 9. + let tag = [9, 0, 3, 0].map(F64); + let s1 = compress([5, 0, 7, 0].map(F64), tag); + let want = compress(s1, tag); // the returned StackBuf (holding s1's words) fed back in let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); assert_eq!(mix(src, want)[5], 2, "two BLAKE2s instructions (one per inlined step)"); verify(&program, &want, &proof).expect("inline StackBuf+scalar tuple return verifies"); let mut bad = want; - bad[1] += F192::ONE; + bad[1] += F64::ONE; assert!( verify(&program, &bad, &proof).is_err(), "wrong published state must be rejected" @@ -503,7 +485,7 @@ def select_pair(flag, a, b): return first, second "; let program = compile(&parse(src).expect("parse")); - program.execute([F192::ONE, F192::ZERO]); + program.execute(pi(&[F64::ONE])); } /// An `@inline` may also alias-return a folded **g-address** among its values: @@ -519,9 +501,11 @@ def main(): hb[1] = 10 hb[GEN] = 20 hb[GEN ** 2] = 30 - fs = StackBuf(2) + fs = StackBuf(4) fs[0] = 1 - fs[1] = 2 + fs[1] = 0 + fs[2] = 2 + fs[3] = 0 cur = hb fs, a, cur = step(fs, cur) fs, b, cur = step(fs, cur) @@ -534,17 +518,19 @@ def main(): @inline def step(state, cursor): x = cursor[GEN ** 0] - tg = StackBuf(2) + tg = StackBuf(4) tg[0] = x - tg[1] = 3 - nb = StackBuf(2) + tg[1] = 0 + tg[2] = 3 + tg[3] = 0 + nb = StackBuf(4) blake2s(state, tg, nb) return nb, x, cursor * GEN "; let program = compile(&parse(src).expect("parse")); // a = hb[0] = 10, b = hb[1] = 20, v = hb[2] = 30 read through the cursor // returned twice-advanced. a + b is XOR: 10 ^ 20 = 30. - let want = [F192::from(F64(30)), F192::from(F64(30))]; + let want = pi(&[F64(30), F64(30)]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("inline advanced-cursor return verifies"); } @@ -558,19 +544,18 @@ def step(state, cursor): fn stack_buf_list_literal() { let src = "\ def main(): - s = [5, 7] - s = [s[1], s[0]] - t = [s[0] + s[1], 3] - out = StackBuf(2) + s = [5, 7, 0, 0] + s = [s[1], s[0], s[2], s[3]] + t = [s[0] + s[1], 3, 0, 0] + out = StackBuf(4) blake2s(s, t, out) - p = 1 - p[1] = out[0] - p[GEN] = out[1] + p = GEN ** 0 + p[0:4] = out return "; let program = compile(&parse(src).expect("parse")); - // s = [7, 5] after the swap → words [7,0,5,0]; t = [7 ^ 5, 3] = [2, 3] → [2,0,3,0]. - let want = digest_cells([F64(7), F64(0), F64(5), F64(0)], [F64(2), F64(0), F64(3), F64(0)]); + // s = [7, 5, 0, 0] after the swap; t = [7 ^ 5, 3, 0, 0] = [2, 3, 0, 0]. + let want = compress([7, 5, 0, 0].map(F64), [2, 3, 0, 0].map(F64)); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); assert_eq!(mix(src, want)[5], 1, "one BLAKE2s instruction"); verify(&program, &want, &proof).expect("list-literal StackBuf verifies"); @@ -612,12 +597,12 @@ fn heap_hint_slice_oob_rejected() { } /// A blake2s heap slice straddling the buffer end is rejected. The 256-bit -/// operand `hb[7:9]` is two 128-bit cells, so the bound check trips at -/// `7 + 2 = 9 > 8`. +/// operand `hb[6:10]` is four 64-bit cells, so the bound check trips at +/// `6 + 4 = 10 > 8`. #[test] -#[should_panic(expected = "heap slice 7:9 out of bounds for `hb` (HeapBuf size 8)")] +#[should_panic(expected = "heap slice 6:10 out of bounds for `hb` (HeapBuf size 8)")] fn heap_blake2s_slice_oob_rejected() { - let src = "def main():\n hb = HeapBuf(8)\n hb[GEN ** 7] = 5\n out = StackBuf(2)\n blake2s(hb[7:9], hb[7:9], out)\n return\n"; + let src = "def main():\n hb = HeapBuf(8)\n hb[GEN ** 7] = 5\n out = StackBuf(4)\n blake2s(hb[6:10], hb[6:10], out)\n return\n"; let _ = compile(&parse(src).expect("parse")); } @@ -647,7 +632,7 @@ def main(): fn heap_index_boundary_ok() { let src = "def main():\n hb = HeapBuf(8)\n hb[GEN ** 7] = 5\n row = hb * GEN ** 4\n y = row[GEN ** 3]\n assert y == 5\n return\n"; let program = compile(&parse(src).expect("parse")); - let pi = [F192::from(F64(3)), F192::from(F64(4))]; + let pi = pi(&[F64(3), F64(4)]); let (proof, _) = prove(&program, pi, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &pi, &proof).expect("boundary access verifies"); } @@ -679,15 +664,15 @@ def main(): let ast = parse(pin).expect("parse"); // The honest prover hints what the callee asserts, and it verifies. let mut program = compile(&ast); - program.set_witness("adv", vec![vec![g_pow(5).into(), g_pow(6).into()]]); - let want = [g_pow(5).into(), g_pow(6).into()]; + program.set_witness("adv", vec![vec![g_pow(5), g_pow(6)]]); + let want = pi(&[g_pow(5), g_pow(6)]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("the honest hint matches the pin"); // A prover hinting anything else must be rejected: that is what the pin is. let mut bad = compile(&ast); - bad.set_witness("adv", vec![vec![g_pow(13).into(), g_pow(14).into()]]); - let dishonest = [g_pow(13).into(), g_pow(14).into()]; + bad.set_witness("adv", vec![vec![g_pow(13), g_pow(14)]]); + let dishonest = pi(&[g_pow(13), g_pow(14)]); assert!( std::panic::catch_unwind(|| bad.execute(dishonest)).is_err(), "the pin must reject a hint it does not match" @@ -710,7 +695,7 @@ def main(): return "; let program = compile(&parse(two).expect("parse")); - let want = [g_pow(3).into(), g_pow(0).into()]; + let want = pi(&[g_pow(3), g_pow(0)]); assert!( std::panic::catch_unwind(|| program.execute(want)).is_err(), "`s[0] = s[1]` asserts that they are equal" @@ -720,7 +705,7 @@ def main(): /// A multi-cell value can cross a call in BOTH directions. /// /// It could always be returned as a run of cells and never passed as one, so a -/// two-cell digest went in through a pointer or an `@inline` expansion while +/// multi-cell digest went in through a pointer or an `@inline` expansion while /// coming back out whole. A `s: StackBuf(n)` parameter takes the same n /// consecutive cells a `StackBuf(n)` return value occupies, placed by the same /// `Abi`, which is why the argument area is now a WIDTH rather than a count. @@ -744,7 +729,7 @@ def main(): return "; let program = compile(&parse(src).expect("parse")); - let want = [g_pow(2).into(), g_pow(1).into()]; + let want = pi(&[g_pow(2), g_pow(1)]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("the run went in and the swapped run came back"); @@ -752,9 +737,9 @@ def main(): for (arg, want) in [ ( "b = StackBuf(3)\n b[0] = GEN ** 1\n r = f(b)", - "got a StackBuf(3)", + "got a 3-cell value", ), - ("r = f(GEN ** 1)", "pass one"), + ("r = f(GEN ** 1)", "pass a 2-cell value"), ] { let src = format!( "def f(s: StackBuf(2)):\n return s[0]\n\ndef main():\n {arg}\n p = GEN ** 0\n p[1] = r\n p[GEN] = GEN ** 0\n return\n" @@ -763,7 +748,7 @@ def main(): let Err(err) = std::panic::catch_unwind(|| compile(&ast)) else { panic!("accepted: {arg}"); }; - let msg = err.downcast_ref::().map(String::as_str).unwrap_or(""); + let msg = panic_message(&*err); assert!(msg.contains(want), "got `{msg}`"); } } @@ -785,7 +770,7 @@ def main(): return "; let program = compile(&parse(src).expect("parse")); - let want = [F192::from(g_pow(5)); 2]; + let want = pi(&[g_pow(5), g_pow(5)]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("three spellings of cell 2 agree"); } diff --git a/crates/lean_compiler/tests/suite/statements.rs b/crates/lean_compiler/tests/suite/statements.rs index f3b00348e..2f70bae2b 100644 --- a/crates/lean_compiler/tests/suite/statements.rs +++ b/crates/lean_compiler/tests/suite/statements.rs @@ -4,7 +4,9 @@ use lean_compiler::{compile, parse}; use lean_vm::cpu::{prove, verify}; -use primitives::field::{F192, g_pow}; +use primitives::field::{F64, g_pow}; + +use crate::common::pi; /// A name is a plain identifier. Two different mistakes arrive here, and both /// used to compile: a mis-split statement (`x /= 2` became a binding named @@ -114,7 +116,7 @@ def main(): return " ); - compile(&parse(&src).expect("parse")).execute([F192::ZERO; 2]); + compile(&parse(&src).expect("parse")).execute([F64::ZERO; 4]); } // `s = [s[1], s[0]]` rebinds to a fresh run and must still swap. let swap = "\ @@ -128,11 +130,7 @@ def main(): p[GEN] = s[1] return "; - let want = [ - F192::from(primitives::field::F64(7)), - F192::from(primitives::field::F64(5)), - ]; - compile(&parse(swap).expect("parse")).execute(want); + compile(&parse(swap).expect("parse")).execute(pi(&[F64(7), F64(5)])); } /// A `mul_range` stop bound that is a compile-time value but not a power of GEN @@ -184,37 +182,33 @@ def main(): // Cell 2 is written on the taken arm (the branch-local `w = GEN`); cell 3 on // the other, from the outer `w`. let program = compile(&parse(src).expect("parse")); - let want = [F192::from(g_pow(1)), F192::from(g_pow(3))]; + let want = pi(&[g_pow(1), g_pow(3)]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("the else arm reads the outer binding"); } /// `g` is `x`, so a literal `2^k` is `g^k` ONLY while `k < 64`: at and above -/// that the modulus folds the monomial back into the low limb while the -/// literal's bit `k` lands in the next one, the tower coefficient of `y`. The -/// guest's own `Y_TOWER` is exactly `2^64`, machine-generated by `dsl_u128`, so -/// a g-power recognizer without that guard gives one literal two values -/// depending on whether it went through a binding. +/// that the literal no longer fits in a word, while the modulus would fold the +/// monomial back into one. A g-power recognizer without that guard reads `2^64` +/// as `g^64`, so a literal that is no machine word compiles as a value, and only +/// in the spelling that goes through the recognizer. Both spellings must be +/// rejected, and `2^63` must still be `g^63`. #[test] fn a_power_of_two_literal_is_a_g_power_only_below_two_to_the_64() { - let src = "\ -Y_TOWER = 18446744073709551616 - -def main(): - yt = Y_TOWER - a = yt * GEN - b = Y_TOWER * GEN - assert a == b - p = GEN ** 0 - p[1] = a - p[GEN] = b - return -"; - let program = compile(&parse(src).expect("parse")); - // y·x, i.e. the tower coefficient shifted, NOT g^65. - let want = [F192::new(0, 2, 0); 2]; - let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); - verify(&program, &want, &proof).expect("2^64 is the tower element y, not g^64"); + let prog = |body: &str| format!("Y_TOWER = 18446744073709551616\n\ndef main():\n {body}\n return\n"); + for body in [ + "yt = Y_TOWER\n a = yt * GEN\n assert a == a", + "a = Y_TOWER * GEN\n assert a == a", + ] { + let ast = parse(&prog(body)).expect("parse"); + let Err(err) = std::panic::catch_unwind(|| compile(&ast)) else { + panic!("`{body}` was accepted"); + }; + let msg = crate::common::panic_message(&*err); + assert!(msg.contains("does not fit in a 64-bit word"), "`{body}`: got `{msg}`"); + } + let top = "a = 9223372036854775808 * GEN\n assert a == GEN ** 64"; + compile(&parse(&prog(top)).expect("parse")).execute([F64::ZERO; 4]); } /// An operator missing a side says which operator and which side. @@ -257,8 +251,8 @@ fn one_hinted_value_needs_no_buffer() { let new = body(" m = hint_witness(\"m\")\n"); let run = |src: &str| { let mut program = compile(&parse(src).expect("parse")); - program.set_witness("m", vec![vec![g_pow(5).into()]]); - let want = [g_pow(5).into(), g_pow(0).into()]; + program.set_witness("m", vec![vec![g_pow(5)]]); + let want = pi(&[g_pow(5), g_pow(0)]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("verifies"); program.execute(want).base_counts.iter().sum::() @@ -278,11 +272,8 @@ def main(): return "; let mut program = compile(&parse(looped).expect("parse")); - program.set_witness( - "w", - vec![vec![g_pow(1).into()], vec![g_pow(2).into()], vec![g_pow(3).into()]], - ); - let want = [g_pow(2).into(), g_pow(0).into()]; + program.set_witness("w", vec![vec![g_pow(1)], vec![g_pow(2)], vec![g_pow(3)]]); + let want = pi(&[g_pow(2), g_pow(0)]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); // `mul_range(1, 8)` is g^0..g^3, so THREE iterations and all three entries // are popped. With `mul_range(1, 4)` the third would sit unread and mask an @@ -332,8 +323,8 @@ fn the_scalar_hint_binds_like_any_other_binder() { "@inline\ndef pick():\n m = hint_witness(\"m\")\n return m\n\ndef main():\n r = pick()\n{tail}" ); let mut program = compile(&parse(&inlined).expect("an @inline body may bind a hint")); - program.set_witness("m", vec![vec![g_pow(5).into()]]); - let want = [g_pow(5).into(), g_pow(0).into()]; + program.set_witness("m", vec![vec![g_pow(5)]]); + let want = pi(&[g_pow(5), g_pow(0)]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("the inlined binding fires"); @@ -355,8 +346,8 @@ def main(): return "; let mut program = compile(&parse(shadowed).expect("shadowing is not capturing")); - program.set_witness("w", vec![vec![g_pow(7).into()], vec![g_pow(8).into()]]); - let want = [g_pow(1).into(), g_pow(0).into()]; + program.set_witness("w", vec![vec![g_pow(7)], vec![g_pow(8)]]); + let want = pi(&[g_pow(1), g_pow(0)]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("the outer StackBuf is untouched"); @@ -368,8 +359,8 @@ def main(): "def pick():\n s = StackBuf(2)\n s[0] = GEN ** 1\n s[1] = GEN ** 2\n s = hint_witness(\"m\")\n return s\n\ndef main():\n r = pick()\n{tail}" ); let mut program = compile(&parse(&rebound).expect("a hint may rebind a StackBuf name")); - program.set_witness("m", vec![vec![g_pow(6).into()]]); - let want = [g_pow(6).into(), g_pow(0).into()]; + program.set_witness("m", vec![vec![g_pow(6)]]); + let want = pi(&[g_pow(6), g_pow(0)]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("the rebound name returns as a scalar"); } @@ -424,7 +415,7 @@ fn a_program_names_what_it_means() { // A call to something that will never be lowered died in the assembler as a // bare `no entry found for key`, with no line. - for callee in ["nosuchfn(1)", "assert_in_k(GEN ** 1, GEN ** 1)"] { + for callee in ["nosuchfn(1)", "assert_eq192(f192(1, 0, 0), f192(1, 0, 0))"] { let src = format!("def main():\n x = {callee}\n{tail}"); let ast = parse(&src).expect("parses"); let Err(err) = std::panic::catch_unwind(|| compile(&ast)) else { @@ -449,7 +440,7 @@ def main(): return "; let program = compile(&parse(captured).expect("a constant may be captured")); - let want = [F192::new(5, 0, 0), g_pow(0).into()]; + let want = pi(&[F64(5), g_pow(0)]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("the captured constant reaches the body"); } @@ -516,8 +507,8 @@ def main(): return "; let mut program = compile(&parse(ok).expect("parse")); - program.set_witness("w", vec![vec![F192::new(1234, 0, 0), F192::new(5678, 0, 0)]]); - let want = [F192::new(1234, 0, 0), g_pow(0).into()]; + program.set_witness("w", vec![vec![F64(1234), F64(5678)]]); + let want = pi(&[F64(1234), g_pow(0)]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("an in-bounds runtime-start slice still works"); } @@ -576,7 +567,7 @@ fn a_match_target_binds_and_every_field_is_walked() { let two = "def two(i: Const):\n return GEN ** i, GEN ** i\n\n"; let publish = " p = GEN ** 0\n p[1] = r\n p[GEN] = GEN ** 0\n return\n"; // Arm 1 returns (g, g), so the two returns XOR to zero wherever they are read. - let want = [F192::ZERO, g_pow(0).into()]; + let want = pi(&[F64::ZERO, g_pow(0)]); let verifies = |src: &str, why: &str| { let program = compile(&parse(src).unwrap_or_else(|e| panic!("{why}: {e}"))); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); @@ -634,7 +625,7 @@ fn a_value_may_ask_for_the_integer_regime() { ); let program = compile(&parse(&src).expect("parse")); std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| { - program.execute([F192::ZERO, F192::ZERO]); + program.execute([F64::ZERO; 4]); })) .is_ok() }; @@ -679,7 +670,17 @@ fn a_value_may_ask_for_the_integer_regime() { /// name. #[test] fn a_function_may_not_shadow_a_builtin() { - for name in ["const", "f192", "addr", "blake2s", "len", "hint_witness", "StackBuf"] { + for name in [ + "const", + "f192", + "add192", + "assert_ne192", + "addr", + "blake2s", + "len", + "hint_witness", + "StackBuf", + ] { let src = format!("def {name}(x):\n assert x == 99\n return x\n\ndef main():\n return\n"); let err = parse(&src).expect_err(&format!("`def {name}` must be rejected")); assert!(err.contains("is a builtin"), "got `{err}`"); @@ -731,7 +732,7 @@ fn an_ambiguous_compile_time_branch_must_be_declared() { }; let fold = |cond: &str| { let program = compile(&parse(&prog(cond)).unwrap_or_else(|e| panic!("{cond}: {e}"))); - let want = [F192::new(5, 0, 0), g_pow(0).into()]; + let want = pi(&[F64(5), g_pow(0)]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).unwrap_or_else(|e| panic!("`{cond}` did not take the then arm: {e:?}")); }; @@ -781,7 +782,7 @@ fn an_ambiguous_compile_time_branch_must_be_declared() { // A variable that merely starts with `const` is not a near miss. let plain = "def main():\n const = 4\n hb = HeapBuf(4)\n if const == 4:\n hb[GEN ** 0] = 5\n else:\n hb[GEN ** 0] = 7\n p = GEN ** 0\n p[1] = hb[GEN ** 0]\n p[GEN] = GEN ** 0\n return\n"; let program = compile(&parse(plain).expect("a name beginning with `const` is an ordinary name")); - let want = [F192::new(5, 0, 0), g_pow(0).into()]; + let want = pi(&[F64(5), g_pow(0)]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("`const == 4` is a comparison, not a wrapper"); @@ -815,8 +816,8 @@ def main(): return "; let mut program = compile(&parse(src).expect("a `,`, `#` or `]` inside a string is part of the name")); - program.set_witness("x,y#z]w", vec![vec![F192::new(7, 0, 0)]]); - let want = [F192::new(7, 0, 0), g_pow(0).into()]; + program.set_witness("x,y#z]w", vec![vec![F64(7)]]); + let want = pi(&[F64(7), g_pow(0)]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("the stream name survived parsing intact"); } @@ -850,11 +851,11 @@ def main(): return "; let mut program = compile(&parse(src).expect("parse")); - program.set_witness("r", vec![vec![g_pow(0).into()]]); - program.set_witness("vals", vec![(10u64..14).map(|v| F192::new(v, 0, 0)).collect()]); + program.set_witness("r", vec![vec![g_pow(0)]]); + program.set_witness("vals", vec![(10u64..14).map(F64).collect()]); // `r` is 1, so the index is `K` itself: cell 1, holding 11. Reading the // integer view instead would name cell 2, holding 12. - let want = [F192::new(11, 0, 0), g_pow(0).into()]; + let want = pi(&[F64(11), g_pow(0)]); let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &want, &proof).expect("the index means what the value means"); } @@ -881,7 +882,7 @@ def main(): assert sa[0] == GEN return "; - let exec = compile(&parse(src).expect("parse")).execute([F192::ZERO; 2]); + let exec = compile(&parse(src).expect("parse")).execute([F64::ZERO; 4]); assert!(exec.unconstrained_reads.is_empty(), "no prover-chosen read"); } @@ -904,7 +905,7 @@ def main(): assert hb[1] == GEN return "; - let exec = compile(&parse(src).expect("parse")).execute([F192::ZERO; 2]); + let exec = compile(&parse(src).expect("parse")).execute([F64::ZERO; 4]); assert!(exec.unconstrained_reads.is_empty(), "no prover-chosen read"); } @@ -997,7 +998,7 @@ fn a_failed_assert_names_its_source_line() { let src = "\ndef main():\n x = GEN ** 3\n\n assert x == GEN ** 4\n return\n"; let program = compile(&parse(src).expect("parse")); let err = std::panic::catch_unwind(|| { - program.execute([F192::ZERO; 2]); + program.execute([F64::ZERO; 4]); }) .expect_err("the assert cannot hold"); let msg = err @@ -1017,7 +1018,7 @@ fn an_inline_call_does_not_steal_the_call_site_line() { let src = "\n@inline\ndef idf(x):\n y = x * x\n return y\n\n\ndef main():\n hb = HeapBuf(4)\n hb[GEN] = GEN ** 3\n a = hb[GEN]\n assert idf(a) == GEN\n return\n"; let program = compile(&parse(src).expect("parse")); let err = std::panic::catch_unwind(|| { - program.execute([F192::ZERO; 2]); + program.execute([F64::ZERO; 4]); }) .expect_err("the assert cannot hold"); let msg = err @@ -1045,30 +1046,45 @@ fn fill_blocks_carry_no_source_line() { ); } -/// A dispatched `match` join reads one cell per bound name, but a callee -/// returning a `StackBuf` flattens it into several ABI cells, so the name would -/// silently bind the run's FIRST cell and the rest would be written where -/// nothing reads them. The guard against that is `all(is scalar)`; negating it -/// as `all(is not scalar)` rather than `any(is not scalar)` left it firing only -/// when EVERY return is a buffer, so a `(scalar, StackBuf)` pair walked through -/// and no existing test noticed. +/// A callee returning a `StackBuf` flattens it into several ABI cells, so a +/// `match` join that read one cell per bound name would bind the run's FIRST cell +/// and leave the rest written where nothing reads them. A mixed `(scalar, +/// StackBuf)` return is the shape that slipped past the old all-scalar guard, and +/// with runs allowed across the join it is the shape to pin: every cell of the +/// run must reach the caller, through an `@inline` arm and through the fused +/// dispatch of a plain one, with a run argument going in as well. #[test] -#[should_panic(expected = "StackBuf return cannot cross a dispatched join")] -fn a_mixed_stack_buf_return_cannot_cross_a_dispatched_join() { +fn a_mixed_stack_buf_return_crosses_a_dispatched_join_whole() { let src = "\ -@inline -def f(k: Const): +@INLINE +def f(v: StackBuf(2), k: Const): s = StackBuf(2) - s[0] = GEN ** 7 - s[1] = GEN ** 9 + s[0] = v[0] * GEN ** k + s[1] = v[1] * GEN ** k return GEN ** k, s def main(): - x = GEN - a, b = match(log(x), range(0, 2), lambda i: f(i)) - assert a == a - assert b == GEN ** 7 + x = hint_witness(\"x\") + assert log x < 2 + pair = [GEN ** 7, GEN ** 9] + a, b = match(log(x), range(0, 2), lambda i: f(pair, i)) + p = GEN ** 0 + p[1] = a + p[GEN] = b[0] + p[GEN ** 2] = b[1] return "; - compile(&parse(src).expect("parse")); + for decorator in ["@inline\n", ""] { + let mut program = compile(&parse(&src.replace("@INLINE\n", decorator)).expect("parse")); + program.set_witness("x", vec![vec![g_pow(1)]]); + let want = pi(&[g_pow(1), g_pow(8), g_pow(10)]); + let (proof, _) = prove(&program, want, lean_vm::pcs::TEST_LOG_INV_RATE); + verify(&program, &want, &proof).unwrap_or_else(|e| panic!("`{decorator}`: {e:?}")); + let mut bad = want; + bad[2] = g_pow(9); + assert!( + verify(&program, &bad, &proof).is_err(), + "`{decorator}`: the run's second cell is bound" + ); + } } diff --git a/crates/lean_compiler/tests/suite/transcript_helpers.rs b/crates/lean_compiler/tests/suite/transcript_helpers.rs index f3d892137..38924c6da 100644 --- a/crates/lean_compiler/tests/suite/transcript_helpers.rs +++ b/crates/lean_compiler/tests/suite/transcript_helpers.rs @@ -1,57 +1,52 @@ +//! The Fiat-Shamir helper shape: nested `@inline` functions passing a four-cell +//! state and a 192-bit scalar as runs, absorbing the scalar into a compression +//! through a list literal that flattens the run, and returning both the next +//! state and a 192-bit challenge sliced out of it. + use lean_compiler::{compile, parse}; -use primitives::field::F192; +use lean_vm::vmhash::compress; +use primitives::field::{F64, F192}; #[test] fn transcript_helpers_are_ordinary_nested_inline_zkdsl() { let src = r#" from snark_lib import * -Y = f192(0, 1, 0) - -@inline -def pack64x2(a, b): - assert_in_k(a, b) - return a + Y * b - @inline -def challenge_from_state(state): - lo = StackBuf(2) - hi = StackBuf(2) - hint_f192_limbs(lo, state[0]) - hint_f192_limbs(hi, state[1]) - state[0] = pack64x2(lo[0], lo[1]) - state[1] = pack64x2(hi[0], hi[1]) - return lo[0] + Y * (lo[1] + Y * hi[0]) +def absorb(state, scalar): + out = StackBuf(4) + blake2s(state, [scalar, 13], out) + return out @inline -def fs_compress(state, scalar, tail, out): - limbs = StackBuf(3) - hint_f192_limbs(limbs, scalar) - block = StackBuf(2) - block[0] = pack64x2(limbs[0], limbs[1]) - block[1] = pack64x2(limbs[2], tail) - assert scalar == limbs[0] + Y * (limbs[1] + Y * limbs[2]) - blake2s(state, block, out) - return +def squeeze(state): + return state[0:3] @inline def observe(state, scalar): - out = StackBuf(2) - fs_compress(state, scalar, 13, out) - return out + fresh = absorb(state, scalar) + challenge = squeeze(fresh) + return fresh, challenge def main(): - state = StackBuf(2) - state[0] = f192(1, 2, 0) - state[1] = f192(3, 4, 0) - out = observe(state, f192(5, 6, 7)) - challenge = challenge_from_state(out) - assert challenge == challenge + state = [1, 2, 3, 4] + s, c = observe(state, f192(5, 6, 7)) + p = GEN ** 0 + p[0:3] = mul192(c, c) + p[GEN ** 3] = s[3] return "#; - // `assert challenge == challenge` is the zkDSL keep-alive idiom: it forces - // the value to be materialized. The assertion under test is that `execute` - // runs the lowered helpers without a write-once memory conflict. let program = compile(&parse(src).expect("parse transcript helpers")); - program.execute([F192::ZERO; 2]); + let d = compress([1, 2, 3, 4].map(F64), [5, 6, 7, 13].map(F64)); + let c = F192::new(d[0].0, d[1].0, d[2].0); + let sq = c * c; + let want = [F64(sq.c0), F64(sq.c1), F64(sq.c2), d[3]]; + assert!(program.execute(want).unconstrained_reads.is_empty()); + + let mut bad = want; + bad[2] += F64::ONE; + assert!( + std::panic::catch_unwind(|| program.execute(bad)).is_err(), + "the challenge's top limb is bound" + ); } diff --git a/crates/lean_compiler/tests/suite/vm_proofs.rs b/crates/lean_compiler/tests/suite/vm_proofs.rs index 969af2a93..735e647e1 100644 --- a/crates/lean_compiler/tests/suite/vm_proofs.rs +++ b/crates/lean_compiler/tests/suite/vm_proofs.rs @@ -15,25 +15,25 @@ use primitives::field::{F64, F192}; /// flock's sub-proof over a real compression. const HASHING: &str = "\ def main(): - a = StackBuf(2) + a = StackBuf(4) a[0] = 5 - a[1] = 7 - c = StackBuf(2) + a[1] = 0 + a[2] = 7 + a[3] = 0 + c = StackBuf(4) blake2s(a, a, c) - p = 1 - p[1] = c[0] - p[GEN] = c[1] + p = GEN ** 0 + p[0:4] = c return "; /// The public input `HASHING` publishes. -fn hashing_pi() -> [F192; 2] { - let h = [F64(5), F64(0), F64(7), F64(0)]; - let d = compress(h, h); - [F192::new(d[0].0, d[1].0, 0), F192::new(d[2].0, d[3].0, 0)] +fn hashing_pi() -> [F64; 4] { + let h = [5, 0, 7, 0].map(F64); + compress(h, h) } -fn hashing_proof() -> (lean_vm::cpu::Program, [F192; 2], Proof) { +fn hashing_proof() -> (lean_vm::cpu::Program, [F64; 4], Proof) { let program = compile(&parse(HASHING).expect("parse")); let pi = hashing_pi(); let (proof, _) = prove(&program, pi, lean_vm::pcs::TEST_LOG_INV_RATE); @@ -84,13 +84,13 @@ fn a_proof_does_not_verify_against_another_program() { // job, and publishing nothing keeps the public input the same for both. let src = |k: u32| { format!( - "def main():\n a = StackBuf(2)\n a[0] = {k}\n a[1] = 7\n \ - c = StackBuf(2)\n blake2s(a, a, c)\n return\n" + "def main():\n a = StackBuf(4)\n a[0] = {k}\n a[1] = 0\n a[2] = 7\n a[3] = 0\n \ + c = StackBuf(4)\n blake2s(a, a, c)\n return\n" ) }; let program = compile(&parse(&src(5)).expect("parse")); let other = compile(&parse(&src(6)).expect("parse")); - let pi = [F192::ZERO, F192::ZERO]; + let pi = [F64::ZERO; 4]; let (proof, _) = prove(&program, pi, lean_vm::pcs::TEST_LOG_INV_RATE); verify(&program, &pi, &proof).expect("honest proof verifies"); assert!( diff --git a/crates/lean_vm/src/cpu/execute.rs b/crates/lean_vm/src/cpu/execute.rs index 5abd178b7..7382ca7a7 100644 --- a/crates/lean_vm/src/cpu/execute.rs +++ b/crates/lean_vm/src/cpu/execute.rs @@ -10,7 +10,7 @@ use primitives::{ }; pub struct Execution { - pub mem: Vec, // data memory after the run, write-once (size cells, power of two) + pub mem: Vec, // data memory after the run, write-once (size cells, power of two) pub cycles: usize, // number of instructions the run executed (trace length) pub mem_used: usize, // cells actually touched, before the power-of-two pad of `mem` /// Rows per table before the fill blocks ran: the work the program itself does, as @@ -33,18 +33,12 @@ pub struct Execution { pub(crate) trace: Trace, // rows + final access-count columns, emitted in the same walk } -/// A memory word interpreted as a K-valued address: valid only when both -/// extension limbs are zero (every g-power is a K-element). -fn as_addr(v: F192) -> Option { - (v.c1 == 0 && v.c2 == 0).then_some(F64(v.c0)) -} - fn pop_witness<'a>( - witness: &'a HashMap>>, + witness: &'a HashMap>>, positions: &mut HashMap<&'a str, usize>, name: &'a str, len: u32, -) -> &'a [F192] { +) -> &'a [F64] { let entries = witness .get(name) .unwrap_or_else(|| panic!("no witness stream `{name}` (Program::set_witness)")); @@ -70,9 +64,9 @@ fn pop_witness<'a>( impl Program { /// Run the program in write-once *fill* mode to produce its [`Execution`]: /// the final memory image and the step count. The public input seeds the - /// first two memory cells `m[0], m[1]` (§sec:e2e-pi). Compilation yields the - /// `Program`; executing it (here) and proving it are separate later phases. - pub fn execute(&self, public_input: [F192; 2]) -> Execution { + /// first four memory cells (§sec:e2e-pi). Compilation yields the `Program`; + /// executing it (here) and proving it are separate later phases. + pub fn execute(&self, public_input: [F64; 4]) -> Execution { self.execute_filled(public_input, super::filler::NO_FLOORS) } @@ -87,7 +81,7 @@ impl Program { /// executing anything; one re-run then realises it, and the loop only exists /// because the fill's own closing jumps and frames feed back into the size. /// Runs that already clear the floor (every one of consequence) execute once. - pub(crate) fn execute_to_floor(&self, public_input: [F192; 2]) -> Execution { + pub(crate) fn execute_to_floor(&self, public_input: [F64; 4]) -> Execution { let mut exec = self.execute_filled(public_input, super::filler::NO_FLOORS); if self.min_log_committed == 0 { return exec; @@ -143,7 +137,7 @@ impl Program { /// ([`Program::min_log_committed`]). pub(crate) fn execute_filled( &self, - public_input: [F192; 2], + public_input: [F64; 4], fill_floors: [usize; crate::tables::N_TABLES], ) -> Execution { // One interpretation of the program, then the fill. The blocks that bring every @@ -161,20 +155,20 @@ impl Program { // Dense write-once data memory (read path stays a vector for speed), the // per-cell access count (g^{count}, default g^0 = 1), and a written mask. - let n0 = self.main_frame.max(2) as usize; + let n0 = self.main_frame.max(4) as usize; let mut m = Mem { - cells: vec![F192::ZERO; n0], + cells: vec![F64::ZERO; n0], written: vec![false; n0], count: vec![F64::ONE; n0], dbg_pc: 0, dbg_line: 0, dbg_hint: None, }; - // Seed the public input into m[0], m[1] (addresses g^0, g^1, §sec:e2e-pi). - m.cells[0] = public_input[0]; - m.cells[1] = public_input[1]; - m.written[0] = true; - m.written[1] = true; + // Seed the public input into the first four cells (§sec:e2e-pi). + for (i, &word) in public_input.iter().enumerate() { + m.cells[i] = word; + m.written[i] = true; + } // Per-pc bytecode execution count (g^{count}). let mut bytecode_count: Vec = vec![F64::ONE; self.prog.len()]; @@ -215,8 +209,10 @@ impl Program { // Per-opcode trace rows, accumulated during the walk and assembled into the // `Trace` once the run finishes (alongside the final count columns). - let mut xor: Vec = Vec::new(); - let mut mul: Vec = Vec::new(); + let mut xor64: Vec = Vec::new(); + let mut mul64: Vec = Vec::new(); + let mut xor192: Vec = Vec::new(); + let mut mul192: Vec = Vec::new(); let mut set: Vec = Vec::new(); let mut deref: Vec = Vec::new(); let mut jump: Vec = Vec::new(); @@ -232,7 +228,7 @@ impl Program { // The three dense per-cell vectors, kept in lockstep. Every method is // `#[inline(always)]`: they sit in the interpreter's hot opcode loop. struct Mem { - cells: Vec, + cells: Vec, written: Vec, count: Vec, /// The pc of the currently executing instruction, and the name of the @@ -253,48 +249,58 @@ impl Program { fn ensure(&mut self, idx: usize) { if idx >= self.cells.len() { let n = idx + 1; - self.cells.resize(n, F192::ZERO); + self.cells.resize(n, F64::ZERO); self.written.resize(n, false); self.count.resize(n, F64::ONE); } } + #[inline(always)] + fn is_written(&self, cell: u32) -> bool { + (cell as usize) < self.written.len() && self.written[cell as usize] + } // Read a cell; an unwritten cell reads as ZERO. #[inline(always)] - fn get(&self, cell: u32) -> F192 { - let c = cell as usize; - if c < self.written.len() && self.written[c] { - self.cells[c] + fn get(&self, cell: u32) -> F64 { + if self.is_written(cell) { + self.cells[cell as usize] } else { - F192::ZERO + F64::ZERO } } + // The `E` element in three consecutive cells. + #[inline(always)] + fn get3(&self, cell: u32) -> F192 { + F192::new(self.get(cell).0, self.get(cell + 1).0, self.get(cell + 2).0) + } // Write-once store: writing a different value to an already-set cell panics. #[inline(always)] - fn put(&mut self, cell: u32, v: F192) { + fn put(&mut self, cell: u32, v: F64) { self.ensure(cell as usize); let c = cell as usize; if self.written[c] { assert!( self.cells[c] == v, - "write-once conflict at cell {cell} ({}, hint {:?}): had {:x}:{:x}:{:x}, new {:x}:{:x}:{:x}", + "write-once conflict at cell {cell} ({}, hint {:?}): had {:#x}, new {:#x}", if self.dbg_line == 0 { format!("pc {}", self.dbg_pc) } else { format!("line {}, pc {}", self.dbg_line, self.dbg_pc) }, self.dbg_hint, - self.cells[c].c2, - self.cells[c].c1, - self.cells[c].c0, - v.c2, - v.c1, - v.c0 + self.cells[c].0, + v.0 ); } else { self.cells[c] = v; self.written[c] = true; } } + #[inline(always)] + fn put3(&mut self, cell: u32, v: F192) { + self.put(cell, F64(v.c0)); + self.put(cell + 1, F64(v.c1)); + self.put(cell + 2, F64(v.c2)); + } // Read the running access count and advance it by ×g (the free increment). // ×g is ×x, i.e. `mul_by_g`, a shift+fold rather than a PMULL; this runs on every // memory access (several million per run), so the cheap form matters. @@ -306,6 +312,14 @@ impl Program { self.count[cell_idx] = mul_by_g(count); count } + #[inline(always)] + fn bump3(&mut self, cell: u32) -> [F64; 3] { + [ + self.bump_access_count(cell), + self.bump_access_count(cell + 1), + self.bump_access_count(cell + 2), + ] + } } // Bounded discrete log for `hint_decompose_bits_exponent`: find n < 2^nbits // with g^n = x, by baby-step giant-step (baby table g^j for j < 2^17, @@ -337,8 +351,8 @@ impl Program { // The cell a heap run starts at: read the pointer back out of memory and // invert it. Shared by every hint that writes through one. fn heap_base(m: &Mem, g: &mut GPow, cell: u32, what: &str) -> u32 { - let p = as_addr(m.get(cell)).unwrap_or_else(|| panic!("{what} pointer is not a K-valued g-power")); - g.log(p).unwrap_or_else(|| panic!("{what} pointer is not a g-power")) + g.log(m.get(cell)) + .unwrap_or_else(|| panic!("{what} pointer is not a g-power")) } // Where a computed-advice bit buffer starts: a frame run needs no lookup @@ -367,7 +381,16 @@ impl Program { if switch { if left.is_none() { assert_eq!((pc, fp), (ending_pc, 0), "main must halt at the sentinel pc g^{{B-1}}"); - let counts = [xor.len(), mul.len(), set.len(), deref.len(), jump.len(), blake2s.len()]; + let counts = [ + xor64.len(), + mul64.len(), + set.len(), + deref.len(), + jump.len(), + blake2s.len(), + xor192.len(), + mul192.len(), + ]; base_counts = Some(counts); fill_base = (1usize << crate::cpu::MIN_LOG_MEM).max(next_free as usize); // A frame per cycle, from the same bump allocator that serves `Alloc` @@ -386,9 +409,9 @@ impl Program { // What the closing jump reads: back to the block's own first // instruction, in this same frame. Then the pointer the `DEREF` // dummy follows, memory cell `0`. - m.put(frame + fr::DEST, F192::from(g.pow(block_pc as usize))); - m.put(frame + fr::NEXT_FP, F192::from(g.pow(frame as usize))); - m.put(frame + fr::PTR, F192::ONE); + m.put(frame + fr::DEST, g.pow(block_pc as usize)); + m.put(frame + fr::NEXT_FP, g.pow(frame as usize)); + m.put(frame + fr::PTR, F64::ONE); runs.push((block_pc, frame, n * (size as usize + 1))); frame += fr::CELLS; } @@ -430,8 +453,8 @@ impl Program { RHint::Log2Ceil { .. } => "Log2Ceil", RHint::BitDecompose { .. } => "BitDecompose", RHint::BitDecomposeExp { .. } => "BitDecomposeExp", - RHint::FieldLimbs { .. } => "FieldLimbs", RHint::Inverse { .. } => "Inverse", + RHint::Inverse192 { .. } => "Inverse192", RHint::Print { .. } => "Print", }); match h { @@ -460,8 +483,7 @@ impl Program { end, start_inverse, } => { - let span = - as_addr(m.get(fp + end)).expect("loop bound is not in K") * start_inverse; + let span = m.get(fp + end) * start_inverse; assert!(!span.is_zero(), "loop bound is zero"); let max_frames = ((1u64 << 28) - u64::from(next_free)) / u64::from(size); let max_span = max_frames.saturating_sub(1) as usize; @@ -477,7 +499,7 @@ impl Program { // the cell holds g^k, allocate k cells (reverse // g-power lookup, growing the index if needed). RHint::AllocDyn { ptr, size } => { - let sz = as_addr(m.get(fp + size)).expect("HeapBuf size is not a K-valued g-power"); + let sz = m.get(fp + size); let cells = g.log(sz).unwrap_or_else(|| { g.grow_to(1 << 20); g.log(sz) @@ -496,7 +518,7 @@ impl Program { // The base is about to become a pointer in memory. g.note(base as usize); m.ensure(next_free as usize); - m.cells[cell as usize] = F192::from(g.pow(base as usize)); + m.cells[cell as usize] = g.pow(base as usize); m.written[cell as usize] = true; } } @@ -506,25 +528,20 @@ impl Program { if m.written[c as usize] { let v = m.cells[c as usize]; // Small integers and small g-powers overlap (8 = x^3 - // = g^3): show every reading that applies. Only a - // K-valued word (extension limbs 0) can be a g-power. - let k = as_addr(v).and_then(|lo| g.log(lo)); - let small = v.c2 == 0 && v.c1 == 0 && v.c0 < 1 << 32; + // = g^3): show every reading that applies. + let k = g.log(v); + let small = v.0 < 1 << 32; match (k, small) { - (Some(k), true) => eprintln!( - "[print] {label} = {} (g^{})", - pretty_integer(v.c0), - pretty_integer(k) - ), + (Some(k), true) => { + eprintln!("[print] {label} = {} (g^{})", pretty_integer(v.0), pretty_integer(k)) + } (Some(k), false) => { eprintln!("[print] {label} = g^{}", pretty_integer(k)) } (None, true) => { - eprintln!("[print] {label} = {}", pretty_integer(v.c0)) - } - (None, false) => { - eprintln!("[print] {label} = {:#x}:{:#x}:{:#x}", v.c2, v.c1, v.c0) + eprintln!("[print] {label} = {}", pretty_integer(v.0)) } + (None, false) => eprintln!("[print] {label} = {:#x}", v.0), } } else { eprintln!("[print] {label} = "); @@ -562,39 +579,31 @@ impl Program { u128::BITS - (word - 1).leading_zeros() }; let mu = cl.max(*floor); - m.put(fp + dst, F192::from(primitives::field::g_pow(mu as usize))); + m.put(fp + dst, primitives::field::g_pow(mu as usize)); } RHint::BitDecompose { value, bits, nbits } => { - assert!(*nbits <= 192, "a machine word has 192 bits"); - let v = m.get(fp + value); - let limbs = [v.c0, v.c1, v.c2]; + assert!(*nbits <= 64, "a machine word has 64 bits"); + let v = m.get(fp + value).0; let bb = bits_base(&m, &mut g, fp, *bits, "decompose"); for j in 0..*nbits { - let bit = (limbs[j as usize / 64] >> (j % 64)) & 1; - m.put(bb + j, F192::new(bit, 0, 0)); + m.put(bb + j, F64((v >> j) & 1)); } } RHint::BitDecomposeExp { value, bits, nbits } => { - let x = as_addr(m.get(fp + value)) - .expect("hint_decompose_bits_exponent value is not a K-valued g-power"); + let x = m.get(fp + value); let n = bounded_dlog(&mut dlog_cache, x, *nbits); let bb = bits_base(&m, &mut g, fp, *bits, "hint_decompose_bits_exponent"); for j in 0..*nbits { - let bit = ((n >> j) & 1) as u64; - m.put(bb + j, F192::new(bit, 0, 0)); - } - } - RHint::FieldLimbs { value, base, len } => { - assert!((1..=3).contains(len), "an F192 value has three K limbs"); - let v = m.get(fp + value); - let limbs = [v.c0, v.c1, v.c2]; - for j in 0..*len { - m.put(fp + base + j, F192::new(limbs[j as usize], 0, 0)); + m.put(bb + j, F64(((n >> j) & 1) as u64)); } } RHint::Inverse { value, dst } => { let v = m.get(fp + value); - m.put(fp + dst, if v.is_zero() { F192::ZERO } else { v.inv() }); + m.put(fp + dst, if v.is_zero() { F64::ZERO } else { v.inv() }); + } + RHint::Inverse192 { value, dst } => { + let v = m.get3(fp + value); + m.put3(fp + dst, if v.is_zero() { F192::ZERO } else { v.inv() }); } } m.dbg_hint = None; @@ -613,12 +622,12 @@ impl Program { v }; - // Loaded once: the shared Xor/Mul arm needs the discriminant again, + // Loaded once: the shared arithmetic arms need the discriminant again, // and `Op` is wide enough that re-reading it costs a second load. let op = self.prog[pc as usize]; match op { - Op::Xor { a, b, c } | Op::Mul { a, b, c } => { - let is_xor = matches!(op, Op::Xor { .. }); + Op::Xor64 { a, b, c } | Op::Mul64 { a, b, c } => { + let is_xor = matches!(op, Op::Xor64 { .. }); let (aa, ab, ac) = (fp + a, fp + b, fp + c); // The row is the equality `m[c] = m[a] op m[b]` over write-once // memory. Normally the operands are known and the result is @@ -633,12 +642,11 @@ impl Program { // users are `MUL`), and an `XOR` into an already-written cell is // how `assert a == b` is spelled, so deducing there would define // the operand the assert exists to check instead of failing on it. - let is_set = |w: &[bool], cell: u32| (cell as usize) < w.len() && w[cell as usize]; - if !is_xor && is_set(&m.written, ac) { - let (ha, hb) = (is_set(&m.written, aa), is_set(&m.written, ab)); + if !is_xor && m.is_written(ac) { + let (ha, hb) = (m.is_written(aa), m.is_written(ab)); if ha ^ hb { let vk = m.get(if ha { aa } else { ab }); - assert!(!vk.is_zero(), "cannot back-solve MUL through a zero operand"); + assert!(!vk.is_zero(), "cannot back-solve MUL64 through a zero operand"); m.put(if ha { ab } else { aa }, m.get(ac) * vk.inv()); } } @@ -658,9 +666,41 @@ impl Program { bytecode_read, }; if is_xor { - xor.push(row); + xor64.push(row); + } else { + mul64.push(row); + } + pc += 1; + } + Op::Xor192 { a, b, c } | Op::Mul192 { a, b, c } => { + let is_xor = matches!(op, Op::Xor192 { .. }); + let (aa, ab, ac) = (fp + a, fp + b, fp + c); + let written = |m: &Mem, cell: u32| (0..3).filter(|&i| m.is_written(cell + i)).count(); + // The `MUL64` back-solve, `div192`'s quotient: an operand not fully + // written is solved for when the other one is, its written limbs + // checked like any store. + if !is_xor && written(&m, ac) == 3 { + let (wa, wb) = (written(&m, aa), written(&m, ab)); + if (wa == 3) != (wb == 3) { + let vk = m.get3(if wa == 3 { aa } else { ab }); + assert!(!vk.is_zero(), "cannot back-solve MUL192 through a zero operand"); + m.put3(if wa == 3 { ab } else { aa }, m.get3(ac) * vk.inv()); + } + } + let (va, vb) = (m.get3(aa), m.get3(ab)); + m.put3(ac, if is_xor { va + vb } else { va * vb }); + let row = X3row { + pc, + fp, + ra: m.bump3(aa), + rb: m.bump3(ab), + rc: m.bump3(ac), + bytecode_read, + }; + if is_xor { + xor192.push(row); } else { - mul.push(row); + mul192.push(row); } pc += 1; } @@ -679,15 +719,7 @@ impl Program { Op::Deref { o1, o2, o3, mode } => { let a1 = fp + o1; let p = m.get(a1); - let p_addr = as_addr(p).unwrap_or_else(|| { - panic!( - "DEREF pointer is not a K-valued g-power at pc {pc} (in {}): {:x}:{:x}", - self.site_at(pc), - p.c1, - p.c0 - ) - }); - let base = match g.log(p_addr) { + let base = match g.log(p) { Some(b) => b, None => { // Not indexed yet: grow the g-power index to the minimum @@ -697,13 +729,13 @@ impl Program { // pointer: a wild deref, or a failed range check // (`assert log _ < _`) surfacing honestly. g.grow_to(1 << MIN_LOG_MEM); - g.log(p_addr).unwrap_or_else(|| { + g.log(p).unwrap_or_else(|| { panic!( "DEREF pointer is not a small g-power at pc {pc} (in {}): a wild \ pointer, or a failed range check \ (value 0x{:016x})", self.site_at(pc), - p_addr.0 + p.0 ) }) } @@ -715,22 +747,15 @@ impl Program { // Equality m[a2] == m[a3]: fill the unset side. m.ensure(a2); let has2 = m.written[a2]; - let has3 = (a3 as usize) < m.written.len() && m.written[a3 as usize]; + let has3 = m.is_written(a3); match (has2, has3) { - (true, true) => { - assert!( - m.cells[a2] == m.get(a3), - "DEREF mismatch at pc {pc} (in {}): m[{a2}] = {:x}:{:x}:{:x} but \ - m[fp+{o3}] = {:x}:{:x}:{:x}", - self.site_at(pc), - m.cells[a2].c2, - m.cells[a2].c1, - m.cells[a2].c0, - m.get(a3).c2, - m.get(a3).c1, - m.get(a3).c0, - ) - } + (true, true) => assert!( + m.cells[a2] == m.get(a3), + "DEREF mismatch at pc {pc} (in {}): m[{a2}] = {:#x} but m[fp+{o3}] = {:#x}", + self.site_at(pc), + m.cells[a2].0, + m.get(a3).0, + ), (true, false) => { let v = m.cells[a2]; m.put(a3, v); @@ -755,12 +780,12 @@ impl Program { // The return target and the frame base are stored as // addresses, and JUMP reads them back. g.note(pc as usize + 2); - let v = F192::from(g.pow(pc as usize + 2)); + let v = g.pow(pc as usize + 2); m.put(a2 as u32, v); } DerefMode::Fp => { g.note(fp as usize); - let v = F192::from(g.pow(fp as usize)); + let v = g.pow(fp as usize); m.put(a2 as u32, v); } } @@ -779,15 +804,7 @@ impl Program { } Op::Jump { oc, od, of } => { let (ac, ad, af) = (fp + oc, fp + od, fp + of); - // All three cells are K-valued on EVERY row, taken or not: the - // table commits one lane each and their memory flushes carry - // literal zeros above it (§sec:tab-jump), which the bus would - // otherwise not balance. A guest branches on g-powers; the one - // idiom that once branched on a word, `assert a != b`, takes an - // inverse hint instead (§sec:prog-div-ne). - let c = as_addr(m.get(ac)).expect("JUMP condition is not a K-valued word"); - let d = as_addr(m.get(ad)).expect("JUMP target is not a K-valued word"); - let f = as_addr(m.get(af)).expect("JUMP fp is not a K-valued word"); + let (c, d, f) = (m.get(ac), m.get(ad), m.get(af)); // The is-nonzero witness `w = c⁻¹` is never used for control // flow, only recorded as a witness column, so it is not // computed here at all: `JumpTable::fill` batch-inverts every @@ -795,7 +812,6 @@ impl Program { let rc = m.bump_access_count(ac); let rd = m.bump_access_count(ad); let rf = m.bump_access_count(af); - let taken = !c.is_zero(); jump.push(Jrow { pc, fp, @@ -804,62 +820,35 @@ impl Program { rf, bytecode_read, }); - if taken { + if c.is_zero() { + pc += 1; + } else { pc = g.log(d).expect("JUMP target not a g-power"); fp = g.log(f).expect("JUMP fp not a g-power"); - } else { - pc += 1; } } - Op::Blake2s { ins, cv, out, md } => { - // Four independently-addressed 128-bit message chunks, each a - // single cell; the chaining value and the output each span two - // consecutive cells; the metadata is one more cell. - let (aa0, aa1, ab0, ab1) = (fp + ins[0], fp + ins[1], fp + ins[2], fp + ins[3]); - let acv = fp + cv; - let ac = fp + out; - let amd = fp + md; - let words = [aa0, aa1, ab0, ab1, acv, acv + 1, amd].map(|a| m.get(a)); - // Naming the operand and the line matters most for the metadata, - // the one a guest builds with field arithmetic rather than reads. - if let Some((i, w)) = words.iter().enumerate().find(|(_, w)| w.c2 != 0) { - const CELLS: [&str; 7] = ["m0", "m1", "m2", "m3", "cv0", "cv1", "md"]; - panic!( - "BLAKE2s {} cell is not a canonical 128-bit embedding at pc {pc} (in {}): \ - top limb 0x{:016x}", - CELLS[i], - self.site_at(pc), - w.c2 - ); + Op::Blake2s { .. } => { + // Every cell the row touches, in value-lane order: the message + // chunks, the digest, the chaining value, the metadata. + let cells = crate::tables::blake2s_cells(&self.prog, pc, fp); + let w = cells.map(|a| m.get(a)); + // Compress the 64 message bytes to the 32-byte result. No table + // constraint covers the digest (the relation is proven by flock, + // §hash_flock); the interpreter still computes the definite digest + // so the output cells are consistent for any later read. + let digest = blake2s_compress( + [w[0], w[1], w[2], w[3]], + [w[4], w[5], w[6], w[7]], + [w[12], w[13], w[14], w[15]], + [w[16], w[17]], + ); + for (k, &word) in digest.iter().enumerate() { + m.put(cells[8 + k], word); } - let va = [F64(words[0].c0), F64(words[0].c1), F64(words[1].c0), F64(words[1].c1)]; - let vb = [F64(words[2].c0), F64(words[2].c1), F64(words[3].c0), F64(words[3].c1)]; - let vcv = [F64(words[4].c0), F64(words[4].c1), F64(words[5].c0), F64(words[5].c1)]; - let metadata = words[6]; - // Compress the 64 message bytes to the 32-byte result, then - // write it to c's two cells. No table constraint covers the - // digest (the relation is proven by flock, §hash_flock); the - // interpreter still computes the definite digest so the output - // cells are consistent for any later read. - let vc = blake2s_compress(va, vb, vcv, metadata); - let outputs = [F192::new(vc[0].0, vc[1].0, 0), F192::new(vc[2].0, vc[3].0, 0)]; - m.put(ac, outputs[0]); - m.put(ac + 1, outputs[1]); - let ra = [m.bump_access_count(aa0), m.bump_access_count(aa1)]; - let rb = [m.bump_access_count(ab0), m.bump_access_count(ab1)]; - let rcv = [m.bump_access_count(acv), m.bump_access_count(acv + 1)]; - let rc = [m.bump_access_count(ac), m.bump_access_count(ac + 1)]; - // Last, matching the flush order, so an md cell aliasing another - // operand still pairs each read with its own count. - let rmd = m.bump_access_count(amd); blake2s.push(Brow { pc, fp, - ra, - rb, - rcv, - rc, - rmd, + r: cells.map(|a| m.bump_access_count(a)), bytecode_read, }); pc += 1; @@ -931,8 +920,8 @@ impl Program { } {} for (a2, a3) in deferred { // Never written: the cells are genuinely unconstrained; fix them to ZERO. - m.put(a2 as u32, F192::ZERO); - m.put(a3, F192::ZERO); + m.put(a2 as u32, F64::ZERO); + m.put(a3, F64::ZERO); } // Cells an instruction touched that nothing ever wrote. Read off the two @@ -951,15 +940,17 @@ impl Program { let mem_used = m.cells.len(); let cells = m.cells.len().next_power_of_two().max(1 << MIN_LOG_MEM); assert!(cells <= 1 << MAX_LOG_MEM, "data memory exceeds 2^{MAX_LOG_MEM} cells"); - m.cells.resize(cells, F192::ZERO); + m.cells.resize(cells, F64::ZERO); m.count.resize(cells, F64::ONE); let trace = Trace { - xor, - mul, + xor64, + mul64, set, deref, jump, blake2s, + xor192, + mul192, mem_count: m.count, bytecode_count, }; diff --git a/crates/lean_vm/src/cpu/filler.rs b/crates/lean_vm/src/cpu/filler.rs index 72301f03b..d8c2b8fd5 100644 --- a/crates/lean_vm/src/cpu/filler.rs +++ b/crates/lean_vm/src/cpu/filler.rs @@ -40,7 +40,7 @@ pub const SIZES: [usize; 8] = [128, 64, 32, 16, 8, 4, 2, 1]; /// Least rows a table can be proven over. Only `BLAKE2s` has one above `1`: flock sizes /// its argument to at least eight instances, so filling that table below the floor /// would leave it padded up to it, which is the padding this exists to avoid. -pub const MIN_ROWS: [usize; N_TABLES] = [1, 1, 1, 1, 1, 8]; +pub const MIN_ROWS: [usize; N_TABLES] = [1, 1, 1, 1, 1, 8, 1, 1]; /// The `JUMP` table's index in [`crate::cpu::Stats::TABLES`]. Every traversal of every /// block lands its closing jump here, so this table is solved last, absorbing the cost @@ -64,11 +64,12 @@ pub struct Block { /// follows. No instruction writes them, which is why a traversal costs its own rows and /// nothing more. /// -/// The rest is what the dummies use: a cell that is never written, so the `JUMP` table's -/// dummy reads a zero condition and falls through instead of leaving the block; the -/// scratch cell a dummy writes, which doubles as the `BLAKE2s` dummy's chaining value and -/// so spans `SCRATCH..SCRATCH+2`; and the digest, placed clear of it so that a digest -/// never becomes the next traversal's chaining value. `DIGEST+2..DIGEST+6` are the +/// The rest is what the dummies use: two cells that are never written, so the `JUMP` +/// table's dummy reads a zero condition and falls through instead of leaving the block, +/// and the `BLAKE2s` dummy reads a zero metadata pair; the scratch run a dummy writes +/// (three cells for a 192-bit one), which doubles as the `BLAKE2s` dummy's chaining value +/// and so spans `SCRATCH..SCRATCH+4`; and the digest, placed clear of it so that a digest +/// never becomes the next traversal's chaining value. `DIGEST+4..DIGEST+12` are the /// message cells, never written, so every traversal compresses the same input. pub mod frame { /// Where the closing jump goes, and in which frame. @@ -76,14 +77,14 @@ pub mod frame { pub const NEXT_FP: u32 = 1; /// The pointer a `DEREF` dummy follows: `g^0`, memory cell `0`. pub const PTR: u32 = 2; - /// Never written, so it reads as zero. + /// Never written, so it and its successor read as zero. pub const ZERO: u32 = 3; /// What a dummy writes. - pub const SCRATCH: u32 = 4; - /// The `BLAKE2s` dummy's output pair. - pub const DIGEST: u32 = 6; + pub const SCRATCH: u32 = 5; + /// The `BLAKE2s` dummy's digest. + pub const DIGEST: u32 = 9; /// Cells a block's frame occupies. - pub const CELLS: u32 = 12; + pub const CELLS: u32 = 21; } /// Traversals per block: `plan[t][k]` is how many times the size-`SIZES[k]` block of @@ -226,15 +227,14 @@ mod tests { #[test] fn solve_reaches_power_of_two_floors() { let cases: [[usize; N_TABLES]; 6] = [ - [0, 0, 0, 0, 0, 0], - [1, 1, 1, 1, 1, 1], - // Roughly the XMSS run's mix. - [125_000, 286_000, 341_000, 508_000, 114_000, 130_000], + [0; N_TABLES], + [1; N_TABLES], + [125_000, 286_000, 341_000, 508_000, 114_000, 130_000, 90_000, 70_000], // Tables already exactly on a power of two, the awkward case: the closing // jumps of every other table's traversals still have to fit somewhere. - [1 << 17, 1 << 12, 1000, 1 << 19, 1 << 16, 8], - [1, 2, 3, 4, 5, 6], - [0, 0, 0, 0, 1 << 20, 0], + [1 << 17, 1 << 12, 1000, 1 << 19, 1 << 16, 8, 1 << 10, 1 << 11], + [1, 2, 3, 4, 5, 6, 7, 8], + [0, 0, 0, 0, 1 << 20, 0, 0, 0], ]; for base in cases { let plan = solve(base, NO_FLOORS).unwrap_or_else(|| panic!("no plan for {base:?}")); @@ -250,7 +250,7 @@ mod tests { /// still lands on one: what a run too small for its consumer buys with. #[test] fn floor_grows_target_table() { - let base = [1_000, 2_000, 3_000, 4_000, 500, 8]; + let base = [1_000, 2_000, 3_000, 4_000, 500, 8, 300, 200]; let mut floors = NO_FLOORS; floors[PAD_TABLE] = 1 << 16; let plan = solve(base, floors).expect("solvable"); @@ -264,7 +264,7 @@ mod tests { /// remainders. #[test] fn fill_uses_bulk_blocks() { - let base = [125_000, 286_000, 341_000, 508_000, 114_000, 130_000]; + let base = [125_000, 286_000, 341_000, 508_000, 114_000, 130_000, 90_000, 70_000]; let plan = solve(base, NO_FLOORS).expect("solvable"); let fill: usize = delivered(&plan).iter().sum(); assert!( diff --git a/crates/lean_vm/src/cpu/hints.rs b/crates/lean_vm/src/cpu/hints.rs index cc97f8852..c5ddc04d6 100644 --- a/crates/lean_vm/src/cpu/hints.rs +++ b/crates/lean_vm/src/cpu/hints.rs @@ -214,13 +214,13 @@ pub enum RHint { /// Write the `nbits` bits of `n`, where `m[fp+value] = g^n` (a bounded /// discrete log at witness generation), into `bits`. BitDecomposeExp { value: Off, bits: BitsDest, nbits: u32 }, - /// Write the first `len` K-coordinate limbs of `m[fp+value]` to - /// `m[fp+base..]`. Computed advice; callers constrain the result. - FieldLimbs { value: Off, base: Off, len: u32 }, /// Write `m[fp+value]⁻¹` to `m[fp+dst]`, or `0` when the value is zero. /// Untrusted: `assert a != b` multiplies the two back together and asserts /// `1`, which a zero value cannot satisfy (`FnLower::lower_assert_ne`). Inverse { value: Off, dst: Off }, + /// [`RHint::Inverse`] in `E`: the three limb cells at `fp+dst` receive the + /// inverse of the element in the three at `fp+value`. + Inverse192 { value: Off, dst: Off }, /// Prover-side debug print (`print(...)` in the zkDSL): display the value /// of `m[fp+cell]` at this program point. Witness generation only. Print { label: String, cell: Off }, diff --git a/crates/lean_vm/src/cpu/isa.rs b/crates/lean_vm/src/cpu/isa.rs index 9b81cd90e..063e54850 100644 --- a/crates/lean_vm/src/cpu/isa.rs +++ b/crates/lean_vm/src/cpu/isa.rs @@ -1,25 +1,24 @@ //! The ISA and the `DEREF` store modes. -use primitives::field::{F64, F192}; +use primitives::field::F64; #[derive(Clone, Copy, Debug)] pub enum Op { - Xor { + /// `m[c] = m[a] + m[b]` in `K`. + Xor64 { a: u32, b: u32, c: u32, }, - Mul { + /// `m[c] = m[a] · m[b]` in `K`. + Mul64 { a: u32, b: u32, c: u32, }, Set { o: u32, - /// The immediate stored into `mem[fp·o]`. A full 192-bit machine word - /// (`E = F192`); K-valued constants (addresses, small ints) ride the - /// low lane with `c1 = c2 = 0`. - k: F192, + k: F64, }, Deref { o1: u32, @@ -32,27 +31,29 @@ pub enum Op { od: u32, of: u32, }, - /// `BLAKE2s`: one standard BLAKE2s compression. The four 16-byte - /// message chunks `ins` (each a canonical 128-bit chunk in ONE 192-bit cell, - /// top limb zero) form the 64-byte block; the digest lands in the TWO - /// consecutive cells `out, out+1`. Each message chunk is addressed - /// independently, with no forced contiguity, so the caller need not assemble - /// its operands into adjacent cells. Every operand is a memory operand, the - /// metadata included. The compression relation is proven by flock. + /// One standard BLAKE2s compression. Each message chunk `ins[i]` names two + /// consecutive 64-bit cells, so the 64-byte block is addressed as four + /// independent 128-bit chunks. The chaining value and the digest each span + /// four consecutive cells, the metadata `counter | f0 ‖ f1` two. The + /// compression relation is proven by flock. Blake2s { ins: [u32; 4], - /// Base of two consecutive cells holding the 256-bit chaining value - /// (canonical 128-bit chunks, top limbs zero). cv: u32, out: u32, - /// The cell holding the metadata `counter:u64 | f0:u32 | f1:u32`, - /// little-endian in its two low K-lanes (top lane zero, as for every - /// other cell this opcode reads). `f0` is the final-block flag and `f1` - /// the last-node flag. A memory operand like the rest, so a program can - /// hash any length: a compile-time counter is one pooled `SET` per - /// frame, a runtime one any cell the program computes. md: u32, }, + /// `m[c..c+3] = m[a..a+3] + m[b..b+3]` in `E`, each operand three consecutive cells. + Xor192 { + a: u32, + b: u32, + c: u32, + }, + /// `m[c..c+3] = m[a..a+3] · m[b..b+3]` in `E = K[y]/(y³+y+1)`. + Mul192 { + a: u32, + b: u32, + c: u32, + }, } /// The source `DEREF` stores at `mem[loc_o1·o2]`: a local cell, the return diff --git a/crates/lean_vm/src/cpu/layout.rs b/crates/lean_vm/src/cpu/layout.rs index 6d7c8811d..d20113029 100644 --- a/crates/lean_vm/src/cpu/layout.rs +++ b/crates/lean_vm/src/cpu/layout.rs @@ -9,12 +9,10 @@ use super::*; // Shared committed columns (indices `0..N_SHARED`). The program (opcode + // operands) is PUBLIC, not committed: it rides the bytecode seed/finalize blocks // as `Coord::Public`; only the witness-dependent finalize counts are committed. -// The data-memory image, a 192-bit word per cell committed as three K-lane columns. -pub const MEM_LO: usize = 0; -pub const MEM_HI: usize = 1; -pub const MEM_TOP: usize = 2; -pub const MFCNT: usize = 3; // per-cell memory access count, g^{A[i]} -pub const BFCNT: usize = 4; // per-pc bytecode execution count, g^{A[pc]} +// The data-memory image, one K word per cell. +pub const MEM: usize = 0; +pub const MFCNT: usize = 1; // per-cell memory access count, g^{A[i]} +pub const BFCNT: usize = 2; // per-pc bytecode execution count, g^{A[pc]} // flock's packed BLAKE2s witness `q_flock`, committed in the SAME stack as every // other column (single PCS). Size `2^(K_LOG+n_log-6)` F64 words, always ≥ 1 // instance (a no-BLAKE2s program commits one full padding instance). It is the @@ -22,8 +20,8 @@ pub const BFCNT: usize = 4; // per-pc bytecode execution count, g^{A[pc]} // virtual and their memory-bus claims route to `q_flock` slots (§hash_flock), so // nothing duplicates them. flock's R1CS validity is discharged by the single // stacked WHIR opening over this commitment. -pub const QFLOCK: usize = 5; -pub const N_SHARED: usize = 6; +pub const QFLOCK: usize = 3; +pub const N_SHARED: usize = 4; /// Global column indexing: the shared columns occupy `0..N_SHARED`, then each /// table `t` (in [`tables::tables`] order) owns the contiguous block `[base[t], @@ -78,9 +76,9 @@ pub struct Layout { /// The stacked witness's shape: its announced `2^mu` size, plus how many lane /// blocks of it the prover actually commits (see [`witness::StackShape`]). pub shape: witness::StackShape, - /// Public input: the first two memory cells `m[0], m[1]` (each a 192-bit - /// word), bound to the committed memory at verification (§sec:e2e-pi). - pub pi: [F192; 2], + /// Public input: the first four memory cells, bound to the committed memory at + /// verification (§sec:e2e-pi). + pub pi: [F64; 4], pub taus: [usize; tables::N_TABLES], } @@ -142,9 +140,7 @@ impl Witness { pub fn col_kappa_sources(log_bytecode: usize) -> Vec> { let sch = schema(); let mut k = vec![Some((0usize, 0usize)); sch.n]; - k[MEM_LO] = Some((1, 0)); - k[MEM_HI] = Some((1, 0)); - k[MEM_TOP] = Some((1, 0)); + k[MEM] = Some((1, 0)); k[MFCNT] = Some((1, 0)); k[BFCNT] = Some((0, log_bytecode)); // q_flock is `2^(K_LOG + n_blocks_log - LOG_PACKING)` F64 words, always ≥ 1 @@ -229,7 +225,9 @@ pub fn bytecode_columns(prog: &[Op]) -> [Vec; 8] { let max_op = prog .iter() .map(|op| match *op { - Op::Xor { a, b, c } | Op::Mul { a, b, c } => a.max(b).max(c), + Op::Xor64 { a, b, c } | Op::Mul64 { a, b, c } | Op::Xor192 { a, b, c } | Op::Mul192 { a, b, c } => { + a.max(b).max(c) + } Op::Set { o, .. } => o, Op::Deref { o1, o2, o3, .. } => o1.max(o2).max(o3), Op::Jump { oc, od, of } => oc.max(od).max(of), @@ -241,19 +239,22 @@ pub fn bytecode_columns(prog: &[Op]) -> [Vec; 8] { let g_at = |i: u32| gpow[i as usize]; // operand g-power let opcode = |op: &Op| match op { - Op::Xor { .. } => OP_XOR, - Op::Mul { .. } => OP_MUL, + Op::Xor64 { .. } => OP_XOR64, + Op::Mul64 { .. } => OP_MUL64, Op::Set { .. } => OP_SET, Op::Deref { .. } => OP_DEREF, Op::Jump { .. } => OP_JUMP, Op::Blake2s { .. } => OP_BLAKE2S, + Op::Xor192 { .. } => OP_XOR192, + Op::Mul192 { .. } => OP_MUL192, }; let operands = |op: &Op| -> (F64, F64, F64) { match *op { - Op::Xor { a, b, c } | Op::Mul { a, b, c } => (g_at(a), g_at(b), g_at(c)), - // The immediate's first two K-limbs ride operand slots o2/o3; c2 - // rides the fpc slot below. - Op::Set { o, k } => (g_at(o), F64(k.c0), F64(k.c1)), + Op::Xor64 { a, b, c } | Op::Mul64 { a, b, c } | Op::Xor192 { a, b, c } | Op::Mul192 { a, b, c } => { + (g_at(a), g_at(b), g_at(c)) + } + // The immediate rides the second operand slot. + Op::Set { o, k } => (g_at(o), k, F64::ZERO), Op::Deref { o1, o2, o3, .. } => (g_at(o1), g_at(o2), g_at(o3)), Op::Jump { oc, od, of } => (g_at(oc), g_at(od), g_at(of)), // BLAKE2s's first three input-word offsets; the last two ride the @@ -266,7 +267,6 @@ pub fn bytecode_columns(prog: &[Op]) -> [Vec; 8] { let fpc = |op: &Op| match op { Op::Deref { mode, .. } => mode.f_pc(), Op::Blake2s { ins, .. } => g_at(ins[3]), - Op::Set { k, .. } => F64(k.c2), _ => F64::ZERO, }; let ffp = |op: &Op| match op { @@ -324,7 +324,7 @@ pub fn bytecode_table(prog: &[Op]) -> Vec { crate::leaf::stacked_bytecode_table(std::slice::from_ref(&block)) } -pub fn layout(prog: &[Op], log_mem: usize, taus: [usize; tables::N_TABLES], pi: [F192; 2]) -> Layout { +pub fn layout(prog: &[Op], log_mem: usize, taus: [usize; tables::N_TABLES], pi: [F64; 4]) -> Layout { let bytecode_size = prog.len(); let log_bytecode = crate::log2_strict_usize(bytecode_size); @@ -355,30 +355,9 @@ pub fn layout(prog: &[Op], log_mem: usize, taus: [usize; tables::N_TABLES], pi: 0, vec![Const(SEP_STATE), Const(g_pow(final_pc as usize)), Const(one)], )); - // memory seed + finalize (every address real, no padding). The value is the - // full three-limb 192-bit word. - push.push(blk( - log_mem, - vec![ - Const(SEP_MEM), - Index, - Const(one), - Col(MEM_LO), - Col(MEM_HI), - Col(MEM_TOP), - ], - )); - pull.push(blk( - log_mem, - vec![ - Const(SEP_MEM), - Index, - Col(MFCNT), - Col(MEM_LO), - Col(MEM_HI), - Col(MEM_TOP), - ], - )); + // memory seed + finalize (every address real, no padding). + push.push(blk(log_mem, vec![Const(SEP_MEM), Index, Const(one), Col(MEM)])); + pull.push(blk(log_mem, vec![Const(SEP_MEM), Index, Col(MFCNT), Col(MEM)])); // bytecode seed + finalize (program columns are public; padding entries // self-cancel at count 1, so the whole 2^log_bytecode is "real"). let bytecode_block = |count: Coord| { @@ -467,7 +446,7 @@ impl Program { crate::hash_flock::n_blocks_log(row_counts[tables::BLAKE2S_TABLE]), "the BLAKE2s table must be filled to flock's instance floor" ); - let pi = [exec.mem[0], exec.mem[1]]; + let pi = [exec.mem[0], exec.mem[1], exec.mem[2], exec.mem[3]]; let l = layout(&self.prog, log_mem, taus, pi); // The stacked witness is written exactly ONCE: allocate it, carve one window @@ -508,14 +487,12 @@ impl Program { let ctx = FillCtx::new(tr, &exec.mem, &gpow, &self.prog, 1 << l.taus[t]); tables::fill_table(*table, &ctx, &mut windows[base..base + n]); } - // Shared columns. The 192-bit memory image splits into three K-limbs. - // These five plus `QFLOCK` below are every shared column, and each has to - // be written: the stack is uninitialized, so one left out would be read - // as indeterminate bytes rather than caught by a length mismatch. - const _: () = assert!(N_SHARED == 6, "a new shared column needs a fill here"); - parallel::fill(windows[MEM_LO], |i| F64(exec.mem[i].c0)); - parallel::fill(windows[MEM_HI], |i| F64(exec.mem[i].c1)); - parallel::fill(windows[MEM_TOP], |i| F64(exec.mem[i].c2)); + // Shared columns. These three plus `QFLOCK` below are every shared + // column, and each has to be written: the stack is uninitialized, so one + // left out would be read as indeterminate bytes rather than caught by a + // length mismatch. + const _: () = assert!(N_SHARED == 4, "a new shared column needs a fill here"); + windows[MEM].copy_from_slice(&exec.mem); parallel::fill(windows[MFCNT], |i| tr.mem_count[i]); // counts ended at g^{A[i]} parallel::fill(windows[BFCNT], |i| tr.bytecode_count[i]); // … at g^{A[pc]} }); @@ -525,20 +502,16 @@ impl Program { // program with no BLAKE2s still carries a single padding instance. let flock_reduction = crate::stage!("Build q_flock", || { // The rows carry only their access counts; the compression's input - // words are the nine cells they read, in the finished (write-once) - // memory image. + // words are the cells they read, in the finished (write-once) memory + // image. let blocks: Vec<_> = parallel::map_collect(tr.blake2s.len(), |i| { let r = &tr.blake2s[i]; - let a = tables::blake2s_addresses(&self.prog, r); - let chunk = |c0: u32, c1: u32| { - let (w0, w1) = (exec.mem[c0 as usize], exec.mem[c1 as usize]); - [F64(w0.c0), F64(w0.c1), F64(w1.c0), F64(w1.c1)] - }; + let w = tables::blake2s_cells(&self.prog, r.pc, r.fp).map(|a| exec.mem[a as usize]); crate::hash_flock::compression( - chunk(a[0], a[1]), - chunk(a[2], a[3]), - chunk(a[4], a[4] + 1), - exec.mem[a[6] as usize], + [w[0], w[1], w[2], w[3]], + [w[4], w[5], w[6], w[7]], + [w[12], w[13], w[14], w[15]], + [w[16], w[17]], ) }); crate::hash_flock::build_qflock_prepared(&blocks, windows[QFLOCK]) diff --git a/crates/lean_vm/src/cpu/mod.rs b/crates/lean_vm/src/cpu/mod.rs index 4d17ebca9..7cf311a44 100644 --- a/crates/lean_vm/src/cpu/mod.rs +++ b/crates/lean_vm/src/cpu/mod.rs @@ -1,12 +1,11 @@ //! Whole-program assembly over GF(2^64) (`doc/leanvm/main.tex`): the instruction tables //! sharing the state / memory / bytecode buses, bound to one field-valued //! commitment and verified oracle-free. Addresses, the program counter, and read -//! counts are g-powers, so every increment is a free ×g. Machine-word arithmetic -//! is over `E = F192 = K[y]/(y³+y+1)` (XOR degree 1, MUL_NATIVE degree 2), -//! with each word carried by three committed `K = F64` limbs. `BLAKE2s` -//! adds the memory/state/bytecode plumbing for a 64→32-byte compression -//! whose relation is discharged by flock (see [`crate::hash_flock`]). All -//! Challenges and transcript scalars live in the same tower E. +//! counts are g-powers, so every increment is a free ×g. A memory word is one +//! `K = F64` element; `XOR192`/`MUL192` compute in `E = F192 = K[y]/(y³+y+1)` over +//! three consecutive cells. `BLAKE2s` adds the memory/state/bytecode plumbing for a +//! 64→32-byte compression whose relation is discharged by flock (see +//! [`crate::hash_flock`]). Challenges and transcript scalars live in `E`. use std::collections::HashMap; @@ -15,8 +14,8 @@ use crate::constraints; use crate::leaf::{self, Block, ColumnClaim, Coord}; use crate::pcs; use crate::tables::{ - self, FillCtx, FlushBuilder, OP_BLAKE2S, OP_DEREF, OP_JUMP, OP_MUL, OP_SET, OP_XOR, SEP_BYTECODE, SEP_MEM, - SEP_STATE, + self, FillCtx, FlushBuilder, OP_BLAKE2S, OP_DEREF, OP_JUMP, OP_MUL64, OP_MUL192, OP_SET, OP_XOR64, OP_XOR192, + SEP_BYTECODE, SEP_MEM, SEP_STATE, }; use crate::transcript::{Challenger, ProverState, Receiver, Transmitter, VerifierState}; use crate::witness; @@ -31,14 +30,13 @@ mod trace; pub use execute::Execution; pub use isa::{DerefMode, Op}; pub use layout::*; -pub(crate) use trace::{Brow, Drow, Jrow, Srow, Trace, Xrow}; - -/// Witness-gen `BLAKE2s` compression: the four message cells' eight -/// words are laid out little-endian into 64 bytes, combined with the supplied -/// chaining value and metadata, and the 32-byte result is split back into the -/// four output words `c`. Flock proves this same compression relation -/// ([`crate::hash_flock`]). -fn blake2s_compress(va: [F64; 4], vb: [F64; 4], vcv: [F64; 4], metadata: F192) -> [F64; 4] { +pub(crate) use trace::{Brow, Drow, Jrow, Srow, Trace, X3row, Xrow}; + +/// Witness-gen `BLAKE2s` compression: the eight message words laid out +/// little-endian into 64 bytes, combined with the supplied chaining value and +/// metadata, and the 32-byte result split back into four words. Flock proves this +/// same compression relation ([`crate::hash_flock`]). +fn blake2s_compress(va: [F64; 4], vb: [F64; 4], vcv: [F64; 4], metadata: [F64; 2]) -> [F64; 4] { crate::hash_flock::digest(&crate::hash_flock::compression(va, vb, vcv, metadata)) } @@ -63,7 +61,7 @@ const MAX_LOG_ROWS: usize = 32; /// `2^32` instructions. const MAX_LOG_BYTECODE: usize = 32; -/// The Fiat-Shamir IV: ONE 32-byte digest, as two field words, committing to +/// The Fiat-Shamir IV: ONE 32-byte digest, as its four words, committing to /// everything fixed about the proving environment. /// /// Two things go in. [`flock::hash::R1CS_DIGEST`] names the flock BLAKE2s @@ -78,8 +76,8 @@ const MAX_LOG_BYTECODE: usize = 32; /// The IV IS the transcript's starting chaining value ([`fiat_shamir::FiatShamirState::new`]), /// so all challenges depend on the circuit version and the program before /// anything else; a recursion guest carries the INNER program's IV in its public -/// input, pinning both with one word pair. -pub fn fs_seed(program: &Program) -> [F192; 2] { +/// input, pinning both with one digest. +pub fn fs_seed(program: &Program) -> [F64; 4] { let mut h = primitives::hash::Hasher::new(); h.update(b"leanvm"); // Length-framed so the preimage parses one way: the domain and the bytecode @@ -87,22 +85,7 @@ pub fn fs_seed(program: &Program) -> [F192; 2] { h.update(&(flock::hash::R1CS_DIGEST.len() as u64).to_le_bytes()); h.update(&flock::hash::R1CS_DIGEST); h.update(&program.bytecode_hash); - let d = h.finalize(); - let word = |o: usize| u64::from_le_bytes(d[o..o + 8].try_into().unwrap()); - [F192::new(word(0), word(8), 0), F192::new(word(16), word(24), 0)] -} - -/// The two 128-bit halves a digest travels in, as the four words the -/// Fiat-Shamir chain runs on. Only defined for a real digest, whose halves have -/// no third limb; [`read_public`] rejects a public input that has one, so a -/// third limb can never be silently dropped from what the transcript binds. -fn digest_words(halves: &[F192; 2]) -> [F64; 4] { - [ - F64(halves[0].c0), - F64(halves[0].c1), - F64(halves[1].c0), - F64(halves[1].c1), - ] + fiat_shamir::digest_words(&h.finalize()) } /// Announce the prover's sizes (`log_mem`, every table's log height, the PCS rate) @@ -127,7 +110,7 @@ fn announce_public(ps: &mut ProverState, log_mem: usize, taus: [usize; tables::N /// rate from the stream, validate them, and reconstruct the public [`Layout`] /// from the program + sizes + public input. (The public input was already bound /// by seeding the transcript.) -fn read_public(vs: &mut VerifierState, prog: &Program, public_input: &[F192; 2]) -> Result<(Layout, usize), CpuError> { +fn read_public(vs: &mut VerifierState, prog: &Program, public_input: &[F64; 4]) -> Result<(Layout, usize), CpuError> { let read_size = |vs: &mut VerifierState| -> Result { let word = vs.next_scalar().map_err(CpuError::Transcript)?; if word.c1 != 0 || word.c2 != 0 { @@ -136,11 +119,6 @@ fn read_public(vs: &mut VerifierState, prog: &Program, public_input: &[F192; 2]) usize::try_from(word.c0).map_err(|_| CpuError::PublicInput) }; - // The transcript binds a public input as two 128-bit halves, so a third limb - // would be dropped and two statements would share a transcript. - if public_input.iter().any(|half| half.c2 != 0) { - return Err(CpuError::PublicInput); - } let log_mem = read_size(vs)?; let mut taus = [0usize; tables::N_TABLES]; for t in &mut taus { @@ -196,7 +174,7 @@ pub struct Program { /// slice of values per `hint_witness` call; the same symbol may be /// hinted many times); each call pops the next entry, whose length must /// match its destination. Prover-side only; verification ignores them. - pub(crate) witness: HashMap>>, + pub(crate) witness: HashMap>>, /// The fill blocks in the bytecode ([`filler`]): the cycles the interpreter /// traverses, after the program halts, to bring every table's row count to a power /// of two. Set by the compiler, prover-side only, and no program code reaches them, @@ -279,7 +257,7 @@ impl Program { /// `hint_witness(dest, "name")` call, popped in order (the same symbol /// may be hinted many times). Prover-side data: entirely unconstrained, /// invisible to verification. - pub fn set_witness(&mut self, name: impl Into, entries: Vec>) { + pub fn set_witness(&mut self, name: impl Into, entries: Vec>) { self.witness.insert(name.into(), entries); } @@ -434,8 +412,7 @@ fn blake2s_value_slot(col: usize) -> Option { } /// Run statistics returned alongside the proof: the cycle count (total executed -/// instructions), the per-opcode counts -/// `[XOR, MUL, SET, DEREF, JUMP, BLAKE2s]`, and the +/// instructions), the per-opcode counts in [`Stats::TABLES`] order, and the /// committed witness size, the sum of the column lengths, i.e. the real data /// before the stacked witness is zero-padded to a power of two `2^m`. pub struct Stats { @@ -456,7 +433,8 @@ pub struct Stats { impl Stats { /// Table names in `counts` order. - pub const TABLES: [&'static str; tables::N_TABLES] = ["XOR", "MUL", "SET", "DEREF", "JUMP", "BLAKE2S"]; + pub const TABLES: [&'static str; tables::N_TABLES] = + ["XOR64", "MUL64", "SET", "DEREF", "JUMP", "BLAKE2S", "XOR192", "MUL192"]; /// One line of per-table instruction counts and shares, largest first, followed by memory and committed-witness sizes. /// @@ -498,7 +476,7 @@ impl Stats { /// run [`Stats`]. `log_inv_rate` selects the PCS rate and is announced in the /// Fiat-Shamir transcript before the commitment. #[tracing::instrument(name = "Prove", skip_all, fields(log_inv_rate))] -pub fn prove(program: &Program, public_input: [F192; 2], log_inv_rate: usize) -> (Proof, Stats) { +pub fn prove(program: &Program, public_input: [F64; 4], log_inv_rate: usize) -> (Proof, Stats) { ::pcs::whir::validate_log_inv_rate(log_inv_rate).expect("valid log_inv_rate"); // One proof is one arena phase: every transient buffer below is bump-allocated // and reclaimed wholesale here, rather than faulted in and unmapped again per @@ -527,11 +505,7 @@ pub fn prove(program: &Program, public_input: [F192; 2], log_inv_rate: usize) -> let committed_size = w.committed_size(); // The public statement (program digest + input) seeds the transcript, so // every challenge depends on the exact program and public input. - debug_assert!( - public_input.iter().all(|h| h.c2 == 0), - "a public input is a 256-bit digest" - ); - let mut ps = ProverState::new(digest_words(&fs_seed(program)), digest_words(&public_input)); + let mut ps = ProverState::new(fs_seed(program), public_input); // Announce the prover's sizes, then commit, before sampling any challenge. announce_public(&mut ps, w.log_mem, w.layout.taus, log_inv_rate); @@ -558,7 +532,7 @@ pub fn prove(program: &Program, public_input: [F192; 2], log_inv_rate: usize) -> leaf::prove_balance(&l.push, &l.pull, &l.count, &cols, &owners, &spans, &mut ps) }); let table_claims = crate::stage!("Prove constraints", || { - // One sumcheck for all six tables (§constraints). + // One sumcheck for all eight tables (§constraints). let table_cols: Vec> = spans .iter() .map(|&(base, n)| (0..n).map(|c| cols[base + c]).collect()) @@ -581,23 +555,8 @@ pub fn prove(program: &Program, public_input: [F192; 2], log_inv_rate: usize) -> }; let l = &w.layout; - // The PI binding transmits the two LOW memory limbs' evaluations - // (§sec:e2e-pi); the verifier checks them against the public-input line at - // `r_pi`. The top limb of both public words is zero, so its evaluation is - // zero at every `r_pi` and rides no scalar. - let r_pi = ps.sample(); - let pi_limbs = [ - primitives::multilinear::interp_k(F64(l.pi[0].c0), F64(l.pi[1].c0), r_pi), - primitives::multilinear::interp_k(F64(l.pi[0].c1), F64(l.pi[1].c1), r_pi), - F192::ZERO, - ]; - for v in &pi_limbs[..2] { - ps.add_scalar(*v); - } - // Memory binds the message, chaining-value, and output words; bytecode binds - // the counter and flags. All corresponding value columns are virtual and route - // to q_flock through `slot_claims`. - let slots = finish_claims(l, bus.claims, &table_claims, r_pi, pi_limbs); + let r_pi = [ps.sample(), ps.sample()]; + let slots = finish_claims(l, bus.claims, &table_claims, r_pi); // Run flock's reduction (zerocheck + lincheck) over the prepared native // layouts retained from the fused q_flock build pass; it returns the @@ -626,14 +585,13 @@ pub fn prove(program: &Program, public_input: [F192; 2], log_inv_rate: usize) -> /// Everything the PCS has to open, in the ORDER that feeds the batch's weights: /// the bus's framework claims, then the zerocheck's per-table column claims, then -/// the three public-input limb claims, each located in its committed slot. Both -/// sides assemble it here, so a claim can never shift by one element. +/// the public-input claim, each located in its committed slot. Both sides assemble +/// it here, so a claim can never shift by one element. fn finish_claims( l: &Layout, bus_claims: Vec, table_claims: &[constraints::Claims], - r_pi: F192, - pi_limbs: [F192; 3], + r_pi: [F192; 2], ) -> Vec { let mut claims = bus_claims; let sch = schema(); @@ -647,23 +605,22 @@ fn finish_claims( }); } } - claims.extend(bind_pi_claim(r_pi, &l.placements, pi_limbs)); + claims.push(bind_pi_claim(r_pi, &l.placements, &l.pi)); slot_claims(l, claims) } -/// The public-input binding (§sec:e2e-pi): the committed `MEM` at `(r, 0,…,0)` must -/// equal `interp(pi[0], pi[1], r)`, one transmitted evaluation per physical `K` -/// limb. The caller has already checked the three against the line; here they -/// simply become the three claims the opening discharges. `placements` comes from -/// the prover's or verifier's layout, so both sides build byte-identical claims. -fn bind_pi_claim(r: F192, placements: &[witness::Placement], limbs: [F192; 3]) -> [ColumnClaim; 3] { - let mut point = vec![F192::ZERO; placements[MEM_LO].n_vars]; - point[0] = r; - [MEM_LO, MEM_HI, MEM_TOP].map(|col| ColumnClaim { - col, - point: point.clone(), - value: limbs[col - MEM_LO], - }) +/// The public-input binding (§sec:e2e-pi): the committed `MEM` at `(r_0, r_1, 0,…,0)` +/// must equal the multilinear extension of the four public words at `(r_0, r_1)`. +/// Both parties know those words, so the claim's value is computed rather than +/// transmitted, and the opening discharges it like any other. +fn bind_pi_claim(r: [F192; 2], placements: &[witness::Placement], pi: &[F64; 4]) -> ColumnClaim { + let mut point = vec![F192::ZERO; placements[MEM].n_vars]; + point[..2].copy_from_slice(&r); + ColumnClaim { + col: MEM, + point, + value: primitives::multilinear::mle_eval(pi, &r), + } } /// Everything a recursion harness needs from an accepting verify run, named @@ -692,8 +649,8 @@ pub struct VerifySummary { /// every scalar the prover wrote and pull the PCS hints, then assert the stream /// was fully consumed. Takes only public inputs, never the prover's witness. #[tracing::instrument(name = "Verify", skip_all)] -pub fn verify(program: &Program, public_input: &[F192; 2], proof: &Proof) -> Result { - let mut vs = VerifierState::new(digest_words(&fs_seed(program)), proof, digest_words(public_input)); +pub fn verify(program: &Program, public_input: &[F64; 4], proof: &Proof) -> Result { + let mut vs = VerifierState::new(fs_seed(program), proof, *public_input); let (l, log_inv_rate) = read_public(&mut vs, program, public_input)?; let root = pcs::read_commitment(&mut vs).map_err(CpuError::Transcript)?; @@ -726,18 +683,8 @@ pub fn verify(program: &Program, public_input: &[F192; 2], proof: &Proof) -> Res ) .map_err(CpuError::Constraint)?; - let r_pi = vs.sample(); - let mut pi_limbs = [F192::ZERO; 3]; - for v in &mut pi_limbs[..2] { - *v = vs.next_scalar().map_err(CpuError::Transcript)?; - } - // The two claimed evaluations must sit on the public-input line, the top - // limb's being zero (§sec:e2e-pi). - let want = primitives::multilinear::interp(l.pi[0], l.pi[1], r_pi); - if pi_limbs[0] + F192::Y * pi_limbs[1] != want { - return Err(CpuError::PublicInput); - } - let slots = finish_claims(&l, bus.claims, &table_claims, r_pi, pi_limbs); + let r_pi = [vs.sample(), vs.sample()]; + let slots = finish_claims(&l, bus.claims, &table_claims, r_pi); // Replay flock's reduction straight off the shared stream (each scalar bound // as it is read) to recover its validity claim on q_flock, then @@ -800,184 +747,93 @@ fn slot_claims(l: &Layout, claims: Vec) -> Vec { mod tests { use super::*; - /// A K-embedded immediate (both extension limbs zero). - fn w(x: u64) -> F192 { - F192::new(x, 0, 0) - } - - /// Pack two 64-bit flock words into the canonical BLAKE2s subspace of F192. - fn cell(lo: F64, hi: F64) -> F192 { - F192::new(lo.0, hi.0, 0) - } + const PI: [F64; 4] = [F64(7), F64(11), F64(13), F64(17)]; /// The default one-block-root metadata for a hand-built BLAKE2s op. - fn md() -> F192 { + fn md() -> [F64; 2] { crate::hash_flock::metadata(crate::hash_flock::PINNED_T, crate::hash_flock::FINAL_FLAG, 0) } - /// The four chaining-value lanes of the two cv cells. - fn cv_lanes(cv0: F192, cv1: F192) -> [F64; 4] { - [F64(cv0.c0), F64(cv0.c1), F64(cv1.c0), F64(cv1.c1)] - } - - /// A hand-built straight-line program with one BLAKE2s row: set up the two - /// 256-bit inputs (`a` at cells 2,3, `b` at cells 4,5, one 128-bit word per - /// cell) and the metadata (cell 8), hash them into the output `c` (cells 6,7), - /// pad with filler SETs so the last executed instruction lands one before the - /// sentinel, and halt there. The flock validity sub-proof plus the memory / - /// state / bytecode bus interactions are verified end-to-end (the proof - /// carries the WHIR opening they assert on). - fn blake2s_program(a: [F64; 4], b: [F64; 4]) -> Program { - // a → cells 2,3 and b → cells 4,5 (two flock lanes per BLAKE2s cell). - let mut prog = vec![ - Op::Set { - o: 2, - k: cell(a[0], a[1]), - }, - Op::Set { - o: 3, - k: cell(a[2], a[3]), - }, - Op::Set { - o: 4, - k: cell(b[0], b[1]), - }, - Op::Set { - o: 5, - k: cell(b[2], b[3]), - }, - Op::Set { o: 8, k: md() }, - // The chaining value reads cells 0,1 (the public input); any - // canonical cv is legal. - Op::Blake2s { - ins: [2, 3, 4, 5], - cv: 0, - out: 6, - md: 8, - }, - ]; // c → cells 6,7 - // 16 slots: 6 executed, then 9 filler SETs step the pc to 15, whose slot is - // the never-executed sentinel. - for k in 0..9u32 { - prog.push(Op::Set { - o: 16 + k, - k: F192::ONE, - }); + /// A hand-built straight-line program over the first `frame` cells, padded with + /// `SET`s so its last instruction lands just before the never-executed sentinel + /// of a `len`-slot bytecode. + fn padded(mut prog: Vec, len: usize, frame: u32) -> Program { + let mut next = frame; + while prog.len() < len - 1 { + prog.push(Op::Set { o: next, k: F64::ONE }); + next += 1; } - prog.push(Op::Xor { a: 0, b: 0, c: 0 }); // sentinel - assert_eq!(prog.len(), 16); - Program::from_bytecode(prog, 32) + prog.push(Op::Xor64 { a: 0, b: 0, c: 0 }); + Program::from_bytecode(prog, next) } - /// The opcode's execution semantics: the digest of the two message pairs under - /// the public input's chaining value lands in the output pair. Proving a program - /// is exercised from `lean_compiler`'s tests, which can compile one whose tables - /// come out powers of two. - #[test] - fn blake2s_computes_the_compression() { - let a: [F64; 4] = [ - F64(0x0123_4567_89ab_cdef), - F64(0xfedc_ba98_7654_3210), - F64(0x1111_2222_3333_4444), - F64(0x5555_6666_7777_8888), - ]; - let b: [F64; 4] = [ - F64(0xdead_beef_cafe_babe), - F64(0x0badf00d_0badf00d), - F64(0x9999_aaaa_bbbb_cccc), - F64(0xdddd_eeee_ffff_0000), - ]; - let program = blake2s_program(a, b); - - let pi = [w(7), w(11)]; - let exec = program.execute(pi); - - // The output cells hold the compression of the two inputs under the - // pi-supplied chaining value (two 128-bit chunks). - let d = blake2s_compress(a, b, cv_lanes(pi[0], pi[1]), md()); - assert_eq!(exec.mem[6], cell(d[0], d[1])); - assert_eq!(exec.mem[7], cell(d[2], d[3])); + /// One BLAKE2s row over hand-set cells: `a` at 4..8 and `b` at 8..12, the + /// metadata at 12..14, the public input as chaining value, the digest at 14..18. + fn blake2s_program(a: [F64; 4], b: [F64; 4], ins: [u32; 4]) -> Program { + let mut prog: Vec = a + .iter() + .chain(&b) + .chain(&md()) + .enumerate() + .map(|(i, &k)| Op::Set { o: 4 + i as u32, k }) + .collect(); + prog.push(Op::Blake2s { + ins, + cv: 0, + out: 14, + md: 12, + }); + padded(prog, 16, 18) } - /// BLAKE consumes the `(c0,c1,0)` embedding. This is not an extra AIR - /// constraint: the full three-limb memory bus makes a request carrying a - /// literal zero in limb 2 match only such a stored word. + const A: [F64; 4] = [ + F64(0x0123_4567_89ab_cdef), + F64(0xfedc_ba98_7654_3210), + F64(0x1111_2222_3333_4444), + F64(0x5555_6666_7777_8888), + ]; + const B: [F64; 4] = [ + F64(0xdead_beef_cafe_babe), + F64(0x0bad_f00d_0bad_f00d), + F64(0x9999_aaaa_bbbb_cccc), + F64(0xdddd_eeee_ffff_0000), + ]; + + /// The opcode's execution semantics: the digest of the four message chunks under + /// the public input's chaining value lands in the four output cells. Proving a + /// program is exercised from `lean_compiler`'s tests, which can compile one whose + /// tables come out powers of two. #[test] - #[should_panic(expected = "BLAKE2s m0 cell is not a canonical 128-bit embedding")] - fn blake2s_requires_zero_third_limb() { - let mut program = blake2s_program([F64::ZERO; 4], [F64::ZERO; 4]); - program.prog[0] = Op::Set { - o: 2, - k: F192::new(0, 0, 1), - }; - let _ = program.execute([w(7), w(11)]); + fn blake2s_computes_the_compression() { + let exec = blake2s_program(A, B, [4, 6, 8, 10]).execute(PI); + assert_eq!(exec.mem[14..18], blake2s_compress(A, B, PI, md())); } - /// A self-hash `BLAKE2s(h, h)` (the hash-chain step) passes the *same* input - /// chunks as both `a` and `b` (`ins[0..2] == ins[2..4]`), so one 256-bit quad - /// feeds both inputs with no copy. The row reads those cells twice; the - /// running access counts thread through and the bus still balances. This is - /// the aliasing the DSL's hash-chain lowering relies on. + /// A self-hash `BLAKE2s(h, h)` names the same chunks as both halves, the aliasing + /// the DSL's hash-chain lowering relies on. #[test] fn blake2s_self_hash_aliased_operands() { - let h: [F64; 4] = [ - F64(0xfeed_face_dead_beef), - F64(0x0123_4567_89ab_cdef), - F64(0xcafe_d00d_1337_c0de), - F64(0x8877_6655_4433_2211), - ]; - // a == b: hash h ‖ h into cells 4,5, both input operands aliasing one pair - let mut prog = vec![ - Op::Set { - o: 2, - k: cell(h[0], h[1]), - }, - Op::Set { - o: 3, - k: cell(h[2], h[3]), - }, - Op::Set { o: 6, k: md() }, - Op::Blake2s { - ins: [2, 3, 2, 3], - cv: 0, - out: 4, - md: 6, - }, - ]; - // 8 slots: 4 executed, 3 filler SETs stepping the pc, then the sentinel. - for k in 0..3u32 { - prog.push(Op::Set { - o: 12 + k, - k: F192::ONE, - }); - } - prog.push(Op::Xor { a: 0, b: 0, c: 0 }); // sentinel - assert_eq!(prog.len(), 8); - let program = Program::from_bytecode(prog, 16); - let pi = [w(3), w(5)]; - - let exec = program.execute(pi); - let d = blake2s_compress(h, h, cv_lanes(pi[0], pi[1]), md()); - assert_eq!(exec.mem[4], cell(d[0], d[1])); - assert_eq!(exec.mem[5], cell(d[2], d[3])); + let exec = blake2s_program(A, B, [4, 6, 4, 6]).execute(PI); + assert_eq!(exec.mem[14..18], blake2s_compress(A, A, PI, md())); } - /// A 192-bit-word MUL: the E-product of two full machine words. Full-limb - /// constants are why this one is hand-written bytecode: a source literal fills - /// only the low two limbs. + /// `MUL192` multiplies the elements its three-cell operands spell, in the tower. #[test] - fn mul_192bit_word() { + fn mul192_multiplies_in_the_tower() { let x = F192::new(0x0123_4567_89ab_cdef, 0xfeed_face_dead_beef, 0x1111_2222_3333_4444); let y = F192::new(0x9999_aaaa_bbbb_cccc, 0x1357_9bdf_2468_ace0, 0x5555_6666_7777_8888); - let prog = vec![ - Op::Set { o: 2, k: x }, - Op::Set { o: 3, k: y }, - Op::Mul { a: 2, b: 3, c: 4 }, - Op::Xor { a: 0, b: 0, c: 0 }, // sentinel (never executed) - ]; - let program = Program::from_bytecode(prog, 5); - let pi = [w(1), w(2)]; - let exec = program.execute(pi); - assert_eq!(exec.mem[4], x * y, "MUL computes the E product"); + let prog: Vec = [x, y] + .iter() + .flat_map(|v| [v.c0, v.c1, v.c2]) + .enumerate() + .map(|(i, limb)| Op::Set { + o: 4 + i as u32, + k: F64(limb), + }) + .chain([Op::Mul192 { a: 4, b: 7, c: 10 }]) + .collect(); + let exec = padded(prog, 8, 13).execute(PI); + let p = x * y; + assert_eq!(exec.mem[10..13], [F64(p.c0), F64(p.c1), F64(p.c2)]); } } diff --git a/crates/lean_vm/src/cpu/trace.rs b/crates/lean_vm/src/cpu/trace.rs index d83ca703f..e894dec26 100644 --- a/crates/lean_vm/src/cpu/trace.rs +++ b/crates/lean_vm/src/cpu/trace.rs @@ -1,17 +1,17 @@ //! Per-opcode trace rows, emitted during execution and assembled into a [`Trace`]. //! //! A row carries only what the witness fill cannot recover: the step's -//! `(pc, fp)`, the access counts read at access time, and `DEREF`'s resolved -//! store index. Everything else is a function of those: -//! operands and immediates come from `prog[pc]`, addresses from `fp` plus those -//! operands, and values from the final memory image, which is write-once and so -//! still holds what each accessed cell held at the time it was accessed. The -//! rows are the interpreter's largest write stream, so what they do not carry -//! they do not pay for, twice: once writing them and once reading them back. +//! `(pc, fp)` and the access counts read at access time. Everything else is a +//! function of those: operands and immediates come from `prog[pc]`, addresses +//! from `fp` plus those operands, and values from the final memory image, which +//! is write-once and so still holds what each accessed cell held at the time it +//! was accessed. The rows are the interpreter's largest write stream, so what +//! they do not carry they do not pay for, twice: once writing them and once +//! reading them back. use primitives::field::F64; -/// `XOR` or `MUL` row: the three cells are `fp·g^{a,b,c}`. +/// `XOR64` or `MUL64` row: the three cells are `fp·g^{a,b,c}`. pub(crate) struct Xrow { pub(crate) pc: u32, pub(crate) fp: u32, // frame base: address = fp + offset, operand = g^offset @@ -20,6 +20,15 @@ pub(crate) struct Xrow { pub(crate) rc: F64, pub(crate) bytecode_read: F64, } +/// `XOR192` or `MUL192` row: one count per limb cell of each operand. +pub(crate) struct X3row { + pub(crate) pc: u32, + pub(crate) fp: u32, + pub(crate) ra: [F64; 3], + pub(crate) rb: [F64; 3], + pub(crate) rc: [F64; 3], + pub(crate) bytecode_read: F64, +} pub(crate) struct Srow { pub(crate) pc: u32, pub(crate) fp: u32, @@ -43,29 +52,24 @@ pub(crate) struct Jrow { pub(crate) bytecode_read: F64, } -/// `BLAKE2s` row: the nine per-cell memory access counts of the four -/// message-chunk cells, the chaining value's two cells, the output's two and the -/// metadata cell. The addresses are `fp·g^{ins[i]}`, `fp·g^{cv}`, `fp·g^{out}`, -/// `fp·g^{md}` and the successors of the middle two; the eighteen flock words are -/// those cells' lanes. +/// `BLAKE2s` row: one access count per cell, in the value-lane order of +/// [`crate::tables::BLAKE2S_VALUE_COLS`]. pub(crate) struct Brow { pub(crate) pc: u32, pub(crate) fp: u32, - pub(crate) ra: [F64; 2], // per-cell counts for the two a input cells - pub(crate) rb: [F64; 2], // … the two b input cells - pub(crate) rcv: [F64; 2], // … the two cv input cells - pub(crate) rc: [F64; 2], // … the two c output cells - pub(crate) rmd: F64, // … and the metadata cell + pub(crate) r: [F64; 18], pub(crate) bytecode_read: F64, } pub(crate) struct Trace { - pub(crate) xor: Vec, - pub(crate) mul: Vec, + pub(crate) xor64: Vec, + pub(crate) mul64: Vec, pub(crate) set: Vec, pub(crate) deref: Vec, pub(crate) jump: Vec, pub(crate) blake2s: Vec, + pub(crate) xor192: Vec, + pub(crate) mul192: Vec, pub(crate) mem_count: Vec, // per-cell running access count g^{count}; final = g^{A[i]} pub(crate) bytecode_count: Vec, // per-pc running execution count g^{count}; final = g^{A[pc]} } @@ -74,12 +78,14 @@ impl Trace { /// Rows per instruction table, in [`crate::cpu::Stats::TABLES`] order. pub(crate) fn row_counts(&self) -> [usize; crate::tables::N_TABLES] { [ - self.xor.len(), - self.mul.len(), + self.xor64.len(), + self.mul64.len(), self.set.len(), self.deref.len(), self.jump.len(), self.blake2s.len(), + self.xor192.len(), + self.mul192.len(), ] } } diff --git a/crates/lean_vm/src/hash_flock.rs b/crates/lean_vm/src/hash_flock.rs index f56baed61..7524320f0 100644 --- a/crates/lean_vm/src/hash_flock.rs +++ b/crates/lean_vm/src/hash_flock.rs @@ -12,10 +12,10 @@ //! ## The mapping //! //! The VM's `BLAKE2s(a, b, cv, metadata) -> c` is one standard BLAKE2s -//! compression. `metadata` packs `counter:u64 | f0:u32 | f1:u32` in -//! little-endian order. All inputs are witness values in `q_flock`, and the -//! memory interaction binds every one of them: `a`, `b`, `cv` and `metadata` -//! are cells the instruction reads. +//! compression. `metadata` is two cells, the byte counter and `f0 | f1 << 32`. +//! All inputs are witness values in `q_flock`, and the memory interaction binds +//! every one of them: `a`, `b`, `cv` and `metadata` are cells the instruction +//! reads. //! //! ## The layout (aligned re-layout, `MSG_BASE = 640`, 64-bit words) //! @@ -37,7 +37,7 @@ use flock::hash::{ Blake2sSetup, Compression, K_LOG, ReductionReplay, blake2s_compress, generate_witness_with_ab_packed_and_lincheck, }; use flock::verifier::VerifyError; -use primitives::field::{F64, F192}; +use primitives::field::F64; use primitives::stream::Stream; use zk_alloc::ArenaVec; @@ -125,24 +125,21 @@ fn pack_words(w: [u32; 2]) -> F64 { F64((w[0] as u64) | ((w[1] as u64) << 32)) } -/// Pack BLAKE2s's compression metadata as one little-endian 128-bit value in -/// the two low K-lanes of a 192-bit word (top lane zero). -pub const fn metadata(counter: u64, f0: u32, f1: u32) -> F192 { - F192::new(counter, (f0 as u64) | ((f1 as u64) << 32), 0) +/// BLAKE2s's compression metadata as its two memory cells: the byte counter, then +/// `f0 | f1 << 32`. +pub const fn metadata(counter: u64, f0: u32, f1: u32) -> [F64; 2] { + [F64(counter), F64((f0 as u64) | ((f1 as u64) << 32))] } -/// Unpack `counter:u64 | f0:u32 | f1:u32` from a 192-bit word (the top lane -/// must be zero). -pub const fn unpack_metadata(x: F192) -> (u64, u32, u32) { - assert!(x.c2 == 0, "BLAKE2s metadata must have a zero top lane"); - (x.c0, x.c1 as u32, (x.c1 >> 32) as u32) +/// Unpack `counter:u64 | f0:u32 | f1:u32` from the two metadata cells. +pub const fn unpack_metadata(x: [F64; 2]) -> (u64, u32, u32) { + (x[0].0, x[1].0 as u32, (x[1].0 >> 32) as u32) } -/// BLAKE2s-256's initial chaining value as four flock words (the two -/// chaining-value cells' low lanes, in canonical lane order). This is the IV -/// with the parameter block (digest length 32, unkeyed, fanout and depth 1) -/// folded into word 0, which is what makes a 64-byte compression equal -/// `blake2s` of those 64 bytes. +/// BLAKE2s-256's initial chaining value as the four cells a chaining value +/// occupies. This is the IV with the parameter block (digest length 32, unkeyed, +/// fanout and depth 1) folded into word 0, which is what makes a 64-byte +/// compression equal `blake2s` of those 64 bytes. /// /// Derived from [`flock::hash::param_iv`] rather than written out: the /// parameter block touches three bytes of word 0, and a hand-copied constant @@ -155,12 +152,8 @@ pub const IV: [F64; 4] = { [w(h[0], h[1]), w(h[2], h[3]), w(h[4], h[5]), w(h[6], h[7])] }; -/// The initial chaining value as the two 192-bit VM memory cells a chaining -/// value occupies (canonical 128-bit chunks, top limbs zero). -pub const IV_CELLS: [F192; 2] = [F192::new(IV[0].0, IV[1].0, 0), F192::new(IV[2].0, IV[3].0, 0)]; - /// The flock [`Compression`] for one VM instruction. -pub fn compression(a: [F64; 4], b: [F64; 4], cv: [F64; 4], meta: F192) -> Compression { +pub fn compression(a: [F64; 4], b: [F64; 4], cv: [F64; 4], meta: [F64; 2]) -> Compression { let mut m = [0u32; 16]; for (i, &w) in a.iter().enumerate() { m[2 * i..2 * i + 2].copy_from_slice(&words_of(w)); @@ -259,6 +252,7 @@ pub fn verify_reduction(n_blocks: usize, vs: &mut VerifierState) -> Result F64 { F64(x) diff --git a/crates/lean_vm/src/lib.rs b/crates/lean_vm/src/lib.rs index 3103fc224..527f8d8f7 100644 --- a/crates/lean_vm/src/lib.rs +++ b/crates/lean_vm/src/lib.rs @@ -1,21 +1,21 @@ //! leanVM: arithmetization of a minimal zkVM (see `doc/leanvm/main.tex`). //! -//! Machine words are `c0 + c1*y + c2*y² ∈ E = K[y]/(y³ + y + 1)`. -//! Addresses, pc/fp, read counters, and logical indices live in +//! Machine words, addresses, pc/fp, read counters, and logical indices live in //! `K = GF(2^64)`; indices are powers of a fixed generator `g`, so incrementing -//! one is a multiplication by `g`, a free virtual operation. Every physical -//! witness column is K-valued (an E-valued word is three K-lane columns) and is -//! committed directly by a dense multilinear PCS. Challenges and transcript -//! scalars live in `E = GF(2^192)`, leaving ample margin for 128-bit soundness. +//! one is a multiplication by `g`, a free virtual operation. `XOR192`/`MUL192` +//! compute in `E = K[y]/(y³ + y + 1)` over three consecutive cells. Every physical +//! witness column is K-valued and is committed directly by a dense multilinear PCS. +//! Challenges and transcript scalars live in `E = GF(2^192)`, leaving ample margin +//! for 128-bit soundness. //! //! - [`transcript`]: the shared Fiat-Shamir transcript (re-exported from `fiat_shamir`). //! - [`pcs`]: `K`-committed witness, `E`-opened, via the stacked WHIR (§sec:stacking, §annex:pcs). //! - [`witness`]: `K`-valued columns stacked into one committed witness. //! - [`gkr`]: the grand product via GKR (§sec:gkr), balancing the bus. //! - [`leaf`]: the shared bus: grand-product balance, decomposed to per-column claims (§sec:gp through §sec:leafstack, §sec:omc). -//! - [`constraints`]: one table sumcheck over all six tables' +//! - [`constraints`]: one table sumcheck over all eight tables' //! degree-2 identities plus their three bus forms (§sec:air). -//! - [`tables`]: the six instruction tables (columns, flushes, constraints). +//! - [`tables`]: the eight instruction tables (columns, flushes, constraints). //! - [`cpu`]: whole-program assembly, control flow, and the prove/verify entry points. //! - [`hash_flock`]: the `BLAKE2s` glue: flock's R1CS validity proof over the same commitment. //! - [`vmhash`]: VM-native hashing (one-block compression and standard BLAKE2s slice hashing). diff --git a/crates/lean_vm/src/tables.rs b/crates/lean_vm/src/tables.rs index 8dd5ef14e..1b4f333dd 100644 --- a/crates/lean_vm/src/tables.rs +++ b/crates/lean_vm/src/tables.rs @@ -4,14 +4,12 @@ //! and its degree-2 constraint. Column indices here are *local* (`0..n_committed_columns`); //! `cpu`'s schema offsets them to global witness columns. //! -//! Columns are `K`-valued (`F64`). The pc/fp, operands, counts, opcodes and -//! separators are single `K`-columns; a **machine word** (memory value) is -//! 192-bit (`E = F192`), committed as THREE `K`-lane columns. Nothing a row -//! DERIVES is a column at all: an operand address `fp·o`, an `XOR`/`MUL` result, -//! the `DEREF` store, the `JUMP` successors are each written out as the degree-2 -//! bus coordinate that carries them (§sec:m3), which leaves `JUMP`'s is-nonzero -//! indicator as the one identity any table still has. Every identity is `K`-valued, -//! so a relation on machine words is written out lane by lane; after the round a +//! Every column is `K`-valued (`F64`), and so is every memory cell: an `E`-valued +//! operand of `XOR192`/`MUL192` is three consecutive cells, one read each. Nothing +//! a row DERIVES is a column at all: an operand address `fp·o`, an arithmetic +//! result, the `DEREF` store, the `JUMP` successors are each written out as the +//! degree-2 bus coordinate that carries them (§sec:m3), which leaves `JUMP`'s +//! is-nonzero indicator as the one identity any table still has. After the round a //! table joins the batch its columns are `E`-valued, which is what //! `eval_constraint` takes. @@ -25,23 +23,10 @@ use primitives::field::{F64, F192, mul_by_g}; // Each is written ONCE, generic over the column type: `F64` in the round a table // joins the batch, `F192` afterwards (see [`ColVal`]). Products of two `K` // columns stay 64-bit and an `η`-power multiplies through `mul_e`. -// -// EVERY identity is `K`-valued, like the columns it reads and like the bus -// coordinates (§sec:air). A relation between machine WORDS is therefore written out -// lane by lane, with the tower multiplication unrolled by hand ([`TOWER_LANES`]), -// rather than assembled into one `E` equation. A table's identity slice is then a -// mixed dot product against its `η`-range: in the round a table joins the batch (the -// batch's largest) that is one 64-bit product per identity plus three PMULL for its -// `η`-power, with ONE reduction for the whole slice ([`ColVal::dot`]). One `E`-valued -// identity would instead cost a full `E×E` product against its `η`-power, plus a -// reduction of its own, and the lanes it bundles are the same lane polynomials -// either way. /// The tower product `x·y` in `E = K[y]/(y³+y+1)` as three lane sums: lane `i` is -/// `Σ x_j·y_k` over `TOWER_LANES[i]`. The five partial sums of §sec:tab-mul fold -/// into `c0 = p0+p3`, `c1 = p1+p3+p4`, `c2 = p2+p4`; written once here because both -/// `MUL`'s result coordinate ([`arith_result`]) and `JUMP`'s inverse identity need -/// the same unrolling. +/// `Σ x_j·y_k` over `TOWER_LANES[i]`. The five partial sums of §sec:tab-mul192 fold +/// into `c0 = p0+p3`, `c1 = p1+p3+p4`, `c2 = p2+p4`. const TOWER_LANES: [&[(usize, usize)]; 3] = [ &[(0, 0), (1, 2), (2, 1)], &[(0, 1), (1, 0), (1, 2), (2, 1), (2, 2)], @@ -49,9 +34,8 @@ const TOWER_LANES: [&[(usize, usize)]; 3] = [ ]; /// One lane of the tower product, in the columns' own field. Only the test that -/// checks [`TOWER_LANES`] against the field needs it: `MUL`'s result rides the bus -/// as a coordinate ([`arith_result`]), and `JUMP`'s condition is `K`-valued, so no -/// identity assembles a tower product any more. +/// checks [`TOWER_LANES`] against the field needs it: `MUL192`'s result rides the +/// bus as a coordinate ([`arith192_result`]). #[cfg(test)] fn tower_lane(lane: usize, x: [T; 3], y: [T; 3]) -> T { TOWER_LANES[lane].iter().fold(T::ZERO, |acc, &(j, k)| acc + x[j] * y[k]) @@ -63,10 +47,6 @@ fn tower_lane(lane: usize, x: [T; 3], y: [T; 3]) -> T { /// gives `b = 1` (and the first `w = cond⁻¹`); when `cond = 0` the first gives /// `b = 0`. The two selections need no identity: the state push carries each as its /// own degree-2 coordinate (§sec:m3). -/// -/// The condition is `K`-valued, so both identities are single-lane. Its memory -/// flush carries literal zeros above the low limb (`memory_k`), so a word outside -/// `K` cannot balance the bus; the interpreter rejects one outright. fn jump_identity(pows: &[F192], cols: &[T], quadratic: bool) -> F192 { use jump::*; let (b, b1) = if quadratic { @@ -95,13 +75,15 @@ pub(crate) const SEP_STATE: F64 = g_pow(0); pub(crate) const SEP_MEM: F64 = g_pow(1); pub(crate) const SEP_BYTECODE: F64 = g_pow(2); -// Opcodes (coordinate 3 of a bytecode tuple). -pub(crate) const OP_XOR: F64 = g_pow(0); -pub(crate) const OP_MUL: F64 = g_pow(1); +// Opcodes (coordinate 3 of a bytecode tuple): `g^t` for table `t` of [`tables`]. +pub(crate) const OP_XOR64: F64 = g_pow(0); +pub(crate) const OP_MUL64: F64 = g_pow(1); pub(crate) const OP_SET: F64 = g_pow(2); pub(crate) const OP_DEREF: F64 = g_pow(3); pub(crate) const OP_JUMP: F64 = g_pow(4); pub(crate) const OP_BLAKE2S: F64 = g_pow(5); +pub(crate) const OP_XOR192: F64 = g_pow(6); +pub(crate) const OP_MUL192: F64 = g_pow(7); // ---- flush builder ----------------------------------------------------------- @@ -154,35 +136,16 @@ impl FlushBuilder { self.pair(push, pull); } - /// The shape every memory interaction shares: the word at `addr` carried as - /// three value coordinates, with the cell's access count advanced by ×g on the - /// push side. A value the row DERIVES rather than commits (an `XOR`/`MUL` - /// result, a `DEREF` store) is passed here as its form: the cell then holds - /// whatever the form says, which removes both the value columns and the - /// identity that used to tie them (§sec:m3). - pub(crate) fn memory_coords(&mut self, addr: Coord, count: usize, vals: [Coord; 3]) { - let mut push = vec![Const(SEP_MEM), addr.clone(), GCol(count, 1)]; - let mut pull = vec![Const(SEP_MEM), addr, Col(count)]; - push.extend_from_slice(&vals); - pull.extend_from_slice(&vals); - self.pair(push, pull); - } - - /// Memory access: read the three-limb word at `addr`. - pub(crate) fn memory(&mut self, addr: Coord, count: usize, val0: usize, val1: usize, val2: usize) { - self.memory_coords(addr, count, [Col(val0), Col(val1), Col(val2)]); - } - - /// Memory read of a K-valued word: both higher limbs are literal zero. Used where - /// the word is carried by a single K column (e.g. the DEREF pointer). Sound - /// because the bus balances only if the stored value's HI lane is likewise 0. - pub(crate) fn memory_k(&mut self, addr: Coord, count: usize, val: usize) { - self.memory_coords(addr, count, [Col(val), Const(F64::ZERO), Const(F64::ZERO)]); - } - - /// Memory access to a canonical 128-bit word `(lo, hi, 0)`. - pub(crate) fn memory_128(&mut self, addr: Coord, count: usize, lo: usize, hi: usize) { - self.memory_coords(addr, count, [Col(lo), Col(hi), Const(F64::ZERO)]); + /// Memory access: the word at `addr`, with the cell's access count advanced by + /// ×g on the push side. A value the row DERIVES rather than commits (an + /// arithmetic result, a `DEREF` store) is passed as its form: the cell then holds + /// whatever the form says, which removes both a value column and the identity + /// that would tie it (§sec:m3). + pub(crate) fn memory(&mut self, addr: Coord, count: usize, val: Coord) { + self.pair( + vec![Const(SEP_MEM), addr.clone(), GCol(count, 1), val.clone()], + vec![Const(SEP_MEM), addr, Col(count), val], + ); } } @@ -192,7 +155,7 @@ impl FlushBuilder { /// image (for read values), and `g^0..` for O(1) address/operand lookups. pub struct FillCtx<'a> { pub(crate) trace: &'a Trace, - pub(crate) mem: &'a [F192], + pub(crate) mem: &'a [F64], pub(crate) gpow: &'a [F64], pub(crate) prog: &'a [Op], /// This table's height `2^tau`, the length of every window in `out`, and its row @@ -209,7 +172,7 @@ pub struct FillCtx<'a> { pub type ColumnOut<'a> = &'a mut [F64]; impl<'a> FillCtx<'a> { - pub(crate) fn new(trace: &'a Trace, mem: &'a [F192], gpow: &'a [F64], prog: &'a [Op], rows: usize) -> Self { + pub(crate) fn new(trace: &'a Trace, mem: &'a [F64], gpow: &'a [F64], prog: &'a [Op], rows: usize) -> Self { Self { trace, mem, @@ -224,27 +187,29 @@ impl<'a> FillCtx<'a> { self.gpow[i as usize] } - /// The three frame offsets of an `XOR`/`MUL` row. A row records - /// only its `(pc, fp)`; the operands are the instruction's, so they are read - /// back from the bytecode rather than copied into every row (§the trace rows - /// in `cpu::trace`). + fn cell(&self, addr: u32) -> F64 { + self.mem[addr as usize] + } + + /// The three frame offsets of an arithmetic row, read back from the bytecode + /// rather than copied into every row (§the trace rows in `cpu::trace`). fn ternary_operands(&self, pc: u32) -> (u32, u32, u32) { match self.prog[pc as usize] { - Op::Xor { a, b, c } | Op::Mul { a, b, c } => (a, b, c), + Op::Xor64 { a, b, c } | Op::Mul64 { a, b, c } | Op::Xor192 { a, b, c } | Op::Mul192 { a, b, c } => { + (a, b, c) + } op => unreachable!("a three-operand row's pc {pc} holds {op:?}"), } } - /// Write local column `at`: `f` over the trace rows, then its pad value to the - /// end of the window. + /// Write local column `at`: `f` over the trace rows. fn col(&self, out: &mut [ColumnOut], rows: &[R], at: usize, f: impl Fn(&R) -> F64 + Sync) { self.cols(out, rows, at, |r| [f(r)]); } /// Write the `N` local columns at `at..at + N` from one closure per row. - /// Columns fed by the same read fill together: the three lanes of a 192-bit - /// memory word are one random access, and splitting them across `N` passes - /// pays for it `N` times. + /// Columns fed by the same bytecode decode fill together: splitting them + /// across `N` passes pays for the decode `N` times. fn cols( &self, out: &mut [ColumnOut], @@ -284,12 +249,6 @@ impl<'a> FillCtx<'a> { } }); } - - /// The three `K`-lanes of the 192-bit word in memory cell `addr`. - fn limbs(&self, addr: u32) -> [F64; 3] { - let w = self.mem[addr as usize]; - [F64(w.c0), F64(w.c1), F64(w.c2)] - } } /// Fill one table's columns and check that every window was written. The stack is @@ -314,7 +273,7 @@ pub trait Table: Sync { /// Local indices of this table's read-count columns: the `g^{count}` values /// recording how many times each accessed cell (and the pc) was read. The /// framework treats them specially: each gets its own single-column "count" - /// bus block, and padding rows fill them with `1` (= g^0) instead of `0`. + /// bus block. fn count_columns(&self) -> &'static [usize]; /// How many identities [`eval_constraint`](Table::eval_constraint) folds. /// Sizes this table's slice of the batch's disjoint `xi`-range (§constraints). @@ -326,7 +285,7 @@ pub trait Table: Sync { /// Evaluate the table's degree-2 constraint at one row, reading column values /// by local index from `cols` (e.g. `cols[jump::V_COND]`) and weighting identity /// `i` by `pows[i]`, this table's slice of the batch's `xi`-powers. The slice is - /// is exactly [`n_constraints`](Table::n_constraints) long: an identity indexed + /// exactly [`n_constraints`](Table::n_constraints) long: an identity indexed /// past its end panics rather than silently reaching into the next table's /// range. The table sumcheck carries every committed column of a table, in /// local order, so `cols` is indexed directly. With `quadratic=false` it returns @@ -347,177 +306,150 @@ pub trait Table: Sync { /// Declare the table's bus interactions. fn flushes(&self, f: &mut FlushBuilder); /// Fill this table's columns from the trace: `out[i]` is local column `i`'s - /// window, already at its padded length. Every window must be written in full; - /// use `FillCtx::col` / `FillCtx::cols`, which append the column's pad value - /// past the last trace row and record the coverage `fill_table` checks. + /// window, already at its final length. Every window must be written in full; + /// use `FillCtx::col` / `FillCtx::cols`, which record the coverage `fill_table` + /// checks. fn fill(&self, ctx: &FillCtx, out: &mut [ColumnOut]); } -/// The tables in fixed order `[XOR, MUL, SET, DEREF, JUMP, BLAKE2S]`, the -/// order of `row_counts` / `taus` throughout `cpu`. -pub const N_TABLES: usize = 6; +/// The tables in fixed order `[XOR64, MUL64, SET, DEREF, JUMP, BLAKE2S, XOR192, +/// MUL192]`, the order of `row_counts` / `taus` throughout `cpu`. Table `t`'s +/// opcode is `g^t`. +pub const N_TABLES: usize = 8; pub fn tables() -> [&'static dyn Table; N_TABLES] { [ - &Arith { is_xor: true }, - &Arith { is_xor: false }, + &Arith64 { is_xor: true }, + &Arith64 { is_xor: false }, &SetTable, &DerefTable, &JumpTable, &Blake2sTable, + &Arith192 { is_xor: true }, + &Arith192 { is_xor: false }, ] } /// Index of the BLAKE2s table in [`tables`]. pub(crate) const BLAKE2S_TABLE: usize = 5; -/// The seven base addresses a `BLAKE2s` row reads: the four message cells, the -/// chaining-value base, the output base (each of those two spans that cell and -/// its successor) and the metadata cell. Recovered from the instruction, not -/// stored per row. -pub(crate) fn blake2s_addresses(prog: &[Op], r: &Brow) -> [u32; 7] { - match prog[r.pc as usize] { - Op::Blake2s { ins, cv, out, md } => [ - r.fp + ins[0], - r.fp + ins[1], - r.fp + ins[2], - r.fp + ins[3], - r.fp + cv, - r.fp + out, - r.fp + md, - ], - op => unreachable!("a BLAKE2s row's pc {} holds {op:?}", r.pc), +/// The eighteen cells a `BLAKE2s` row reads, in value-lane order: the four message +/// chunks' two cells each, the digest's four, the chaining value's four and the +/// metadata's two. Recovered from the instruction, not stored per row. +pub(crate) fn blake2s_cells(prog: &[Op], pc: u32, fp: u32) -> [u32; 18] { + match prog[pc as usize] { + Op::Blake2s { ins, cv, out, md } => { + let [m0, m1, m2, m3] = ins.map(|o| fp + o); + let (cv, out, md) = (fp + cv, fp + out, fp + md); + [ + m0, + m0 + 1, + m1, + m1 + 1, + m2, + m2 + 1, + m3, + m3 + 1, + out, + out + 1, + out + 2, + out + 3, + cv, + cv + 1, + cv + 2, + cv + 3, + md, + md + 1, + ] + } + op => unreachable!("a BLAKE2s row's pc {pc} holds {op:?}"), } } /// BLAKE2s value-column LOCAL indices in canonical slot order /// `[a0..a3, b0..b3, c0..c3, cv0..cv3, md_lo, md_hi]` (matches -/// `hash_flock::SLOTS`). These columns are -/// VIRTUAL (never committed): `q_flock` already holds those words at fixed packed -/// slots, so `cpu` routes their memory-bus evaluation claims straight to `q_flock` -/// (`slot_claims`): the value the bus flushes IS the flock-proven word. -pub const BLAKE2S_VALUE_COLS: [usize; 18] = [ - blake2st::V_M0, - blake2st::V_M0 + 1, - blake2st::V_M0 + 2, - blake2st::V_M0 + 3, - blake2st::V_M2, - blake2st::V_M2 + 1, - blake2st::V_M2 + 2, - blake2st::V_M2 + 3, - blake2st::V_OUT0, - blake2st::V_OUT0 + 1, - blake2st::V_OUT0 + 2, - blake2st::V_OUT0 + 3, - blake2st::V_CV0, - blake2st::V_CV0 + 1, - blake2st::V_CV0 + 2, - blake2st::V_CV0 + 3, - blake2st::MD0, - blake2st::MD1, -]; -// The eighteen value lanes are laid out contiguously (V_M0..V_M0+17), so they map -// 1:1 onto `hash_flock::SLOTS`. -const _: () = assert!( - blake2st::V_M2 == blake2st::V_M0 + 4 - && blake2st::V_OUT0 == blake2st::V_M0 + 8 - && blake2st::V_CV0 == blake2st::V_M0 + 12 - && blake2st::MD0 == blake2st::V_M0 + 16 - && blake2st::MD1 == blake2st::V_M0 + 17 -); - -// ---- XOR / MUL --------------------------------------------------------------- - -/// `XOR` and `MUL_NATIVE` share their column layout, flushes, and fill; they -/// differ only in the opcode tag and in how the destination cell's value rides -/// the bus (`v_A + v_B` for `XOR`, `v_A·v_B` in `E = K[y]/(y³+y+1)` for `MUL`). -/// Neither commits that value and neither has an identity: the destination's -/// memory flush carries the result as a degree-≤2 coordinate over the operand -/// lanes, so bus balance IS the assertion (§sec:m3). -struct Arith { +/// `hash_flock::SLOTS`). These columns are VIRTUAL (never committed): `q_flock` +/// already holds those words at fixed packed slots, so `cpu` routes their +/// memory-bus evaluation claims straight to `q_flock` (`slot_claims`): the value the +/// bus flushes IS the flock-proven word. +pub const BLAKE2S_VALUE_COLS: [usize; 18] = { + let mut cols = [0; 18]; + let mut i = 0; + while i < 18 { + cols[i] = blake2st::V0 + i; + i += 1; + } + cols +}; + +// ---- XOR64 / MUL64 ----------------------------------------------------------- + +/// `XOR64` and `MUL64` share their column layout, flushes, and fill; they differ +/// only in the opcode tag and in how the destination cell's value rides the bus +/// (`v_A + v_B` or `v_A·v_B`). Neither commits that value and neither has an +/// identity: bus balance IS the assertion (§sec:m3). +struct Arith64 { is_xor: bool, } -mod arith { +mod arith64 { pub const PC: usize = 0; pub const FP: usize = 1; pub const OA: usize = 2; pub const OB: usize = 3; pub const OC: usize = 4; - // No absolute-address columns: the memory bus carries `fp·o` as a product - // coordinate (§sec:m3), which is why there is no address binding below. - // The two read words, each three K-limbs. The third (the result) is DERIVED. - pub const VA_LO: usize = 5; - pub const VA_HI: usize = 6; - pub const VA_TOP: usize = 7; - pub const VB_LO: usize = 8; - pub const VB_HI: usize = 9; - pub const VB_TOP: usize = 10; - pub const RA: usize = 11; - pub const RB: usize = 12; - pub const RC: usize = 13; - pub const RBC: usize = 14; - pub const N: usize = 15; + pub const VA: usize = 5; + pub const VB: usize = 6; + pub const RA: usize = 7; + pub const RB: usize = 8; + pub const RC: usize = 9; + pub const RBC: usize = 10; + pub const N: usize = 11; } -/// The result word's three K-lanes as forms over the operand lanes. For `XOR` -/// that is the lane-wise sum; for `MUL` it is the tower product, unrolled through -/// [`TOWER_LANES`]. -fn arith_result(is_xor: bool) -> [Coord; 3] { - use arith::*; - if is_xor { - return [ - Coord::Sum(vec![Col(VA_LO), Col(VB_LO)]), - Coord::Sum(vec![Col(VA_HI), Col(VB_HI)]), - Coord::Sum(vec![Col(VA_TOP), Col(VB_TOP)]), - ]; - } - let (a, b) = ([VA_LO, VA_HI, VA_TOP], [VB_LO, VB_HI, VB_TOP]); - let lane = |i: usize| Coord::Sum(TOWER_LANES[i].iter().map(|&(j, k)| Prod(a[j], b[k], 0)).collect()); - [lane(0), lane(1), lane(2)] -} - -impl Table for Arith { +impl Table for Arith64 { fn n_committed_columns(&self) -> usize { - arith::N + arith64::N } fn count_columns(&self) -> &'static [usize] { - use arith::*; + use arith64::*; &[RA, RB, RC, RBC] } fn flushes(&self, f: &mut FlushBuilder) { - use arith::*; + use arith64::*; f.state_step(PC, FP); f.bytecode( PC, RBC, - if self.is_xor { OP_XOR } else { OP_MUL }, + if self.is_xor { OP_XOR64 } else { OP_MUL64 }, &[Col(OA), Col(OB), Col(OC), Const(F64::ZERO), Const(F64::ZERO)], ); - f.memory(Prod(FP, OA, 0), RA, VA_LO, VA_HI, VA_TOP); - f.memory(Prod(FP, OB, 0), RB, VB_LO, VB_HI, VB_TOP); - f.memory_coords(Prod(FP, OC, 0), RC, arith_result(self.is_xor)); + f.memory(Prod(FP, OA, 0), RA, Col(VA)); + f.memory(Prod(FP, OB, 0), RB, Col(VB)); + let result = if self.is_xor { + Coord::Sum(vec![Col(VA), Col(VB)]) + } else { + Prod(VA, VB, 0) + }; + f.memory(Prod(FP, OC, 0), RC, result); } fn fill(&self, ctx: &FillCtx, out: &mut [ColumnOut]) { - use arith::*; - let rows = if self.is_xor { &ctx.trace.xor } else { &ctx.trace.mul }; + use arith64::*; + let rows = if self.is_xor { + &ctx.trace.xor64 + } else { + &ctx.trace.mul64 + }; ctx.col(out, rows, PC, |r| ctx.g_at(r.pc)); ctx.col(out, rows, FP, |r| ctx.g_at(r.fp)); - // The offsets and both operand words come out of ONE bytecode decode: split - // across passes, the row's instruction is fetched once per pass. ctx.cols(out, rows, OA, |r| { let (a, b, c) = ctx.ternary_operands(r.pc); - let (va, vb) = (ctx.limbs(r.fp + a), ctx.limbs(r.fp + b)); [ ctx.g_at(a), ctx.g_at(b), ctx.g_at(c), - va[0], - va[1], - va[2], - vb[0], - vb[1], - vb[2], + ctx.cell(r.fp + a), + ctx.cell(r.fp + b), ] }); ctx.cols(out, rows, RA, |r| [r.ra, r.rb, r.rc]); @@ -525,6 +457,104 @@ impl Table for Arith { } } +// ---- XOR192 / MUL192 --------------------------------------------------------- + +/// `XOR192` and `MUL192`: each operand is three consecutive cells `fp·o·g^i`, the +/// limbs of one `E` element, read one at a time. The result's three lanes ride the +/// bus as forms over the operand lanes: the lane-wise sum, or the tower product +/// unrolled through [`TOWER_LANES`]. +struct Arith192 { + is_xor: bool, +} + +mod arith192 { + pub const PC: usize = 0; + pub const FP: usize = 1; + pub const OA: usize = 2; + pub const OB: usize = 3; + pub const OC: usize = 4; + pub const VA: usize = 5; // three lanes + pub const VB: usize = 8; // three lanes + pub const RA: usize = 11; // one count per limb cell + pub const RB: usize = 14; + pub const RC: usize = 17; + pub const RBC: usize = 20; + pub const N: usize = 21; +} + +/// The result word's three K-lanes as forms over the operand lanes. +fn arith192_result(is_xor: bool) -> [Coord; 3] { + use arith192::*; + std::array::from_fn(|i| { + if is_xor { + Coord::Sum(vec![Col(VA + i), Col(VB + i)]) + } else { + Coord::Sum(TOWER_LANES[i].iter().map(|&(j, k)| Prod(VA + j, VB + k, 0)).collect()) + } + }) +} + +impl Table for Arith192 { + fn n_committed_columns(&self) -> usize { + arith192::N + } + fn count_columns(&self) -> &'static [usize] { + use arith192::*; + &[RA, RA + 1, RA + 2, RB, RB + 1, RB + 2, RC, RC + 1, RC + 2, RBC] + } + fn flushes(&self, f: &mut FlushBuilder) { + use arith192::*; + f.state_step(PC, FP); + f.bytecode( + PC, + RBC, + if self.is_xor { OP_XOR192 } else { OP_MUL192 }, + &[Col(OA), Col(OB), Col(OC), Const(F64::ZERO), Const(F64::ZERO)], + ); + // A limb's cell is a free ×g^i on the address product. + for i in 0..3 { + f.memory(Prod(FP, OA, i as u32), RA + i, Col(VA + i)); + } + for i in 0..3 { + f.memory(Prod(FP, OB, i as u32), RB + i, Col(VB + i)); + } + for (i, lane) in arith192_result(self.is_xor).into_iter().enumerate() { + f.memory(Prod(FP, OC, i as u32), RC + i, lane); + } + } + fn fill(&self, ctx: &FillCtx, out: &mut [ColumnOut]) { + use arith192::*; + let rows = if self.is_xor { + &ctx.trace.xor192 + } else { + &ctx.trace.mul192 + }; + ctx.col(out, rows, PC, |r| ctx.g_at(r.pc)); + ctx.col(out, rows, FP, |r| ctx.g_at(r.fp)); + ctx.cols(out, rows, OA, |r| { + let (a, b, c) = ctx.ternary_operands(r.pc); + let (va, vb) = (r.fp + a, r.fp + b); + [ + ctx.g_at(a), + ctx.g_at(b), + ctx.g_at(c), + ctx.cell(va), + ctx.cell(va + 1), + ctx.cell(va + 2), + ctx.cell(vb), + ctx.cell(vb + 1), + ctx.cell(vb + 2), + ] + }); + ctx.cols(out, rows, RA, |r| { + [ + r.ra[0], r.ra[1], r.ra[2], r.rb[0], r.rb[1], r.rb[2], r.rc[0], r.rc[1], r.rc[2], + ] + }); + ctx.col(out, rows, RBC, |r| r.bytecode_read); + } +} + // ---- SET --------------------------------------------------------------------- struct SetTable; @@ -533,13 +563,11 @@ mod set { pub const PC: usize = 0; pub const FP: usize = 1; pub const O: usize = 2; - // The stored immediate's three K-limbs ride the bytecode's spare slots. - pub const K_LO: usize = 3; - pub const K_HI: usize = 4; - pub const K_TOP: usize = 5; - pub const R: usize = 6; - pub const RBC: usize = 7; - pub const N: usize = 8; + // The stored immediate rides the bytecode's second operand slot. + pub const K: usize = 3; + pub const R: usize = 4; + pub const RBC: usize = 5; + pub const N: usize = 6; } impl Table for SetTable { @@ -553,31 +581,26 @@ impl Table for SetTable { fn flushes(&self, f: &mut FlushBuilder) { use set::*; f.state_step(PC, FP); - // The immediate's three limbs occupy bytecode operand slots o2..o4 - // (matching layout::operands for SET). f.bytecode( PC, RBC, OP_SET, - &[Col(O), Col(K_LO), Col(K_HI), Col(K_TOP), Const(F64::ZERO)], + &[Col(O), Col(K), Const(F64::ZERO), Const(F64::ZERO), Const(F64::ZERO)], ); - // The stored constant K is the cell's value. - f.memory(Prod(FP, O, 0), R, K_LO, K_HI, K_TOP); + f.memory(Prod(FP, O, 0), R, Col(K)); } fn fill(&self, ctx: &FillCtx, out: &mut [ColumnOut]) { use set::*; let rows = &ctx.trace.set; - // The offset and the stored immediate are the instruction's. let imm = |r: &Srow| match ctx.prog[r.pc as usize] { Op::Set { o, k } => (o, k), op => unreachable!("a SET row's pc {} holds {op:?}", r.pc), }; ctx.col(out, rows, PC, |r| ctx.g_at(r.pc)); ctx.col(out, rows, FP, |r| ctx.g_at(r.fp)); - ctx.col(out, rows, O, |r| ctx.g_at(imm(r).0)); - ctx.cols(out, rows, K_LO, |r| { - let k = imm(r).1; - [F64(k.c0), F64(k.c1), F64(k.c2)] + ctx.cols(out, rows, O, |r| { + let (o, k) = imm(r); + [ctx.g_at(o), k] }); ctx.col(out, rows, R, |r| r.r); ctx.col(out, rows, RBC, |r| r.bytecode_read); @@ -596,35 +619,30 @@ mod deref { pub const O3: usize = 4; pub const FPC: usize = 5; pub const FFP: usize = 6; - // The pointer word is a SINGLE K-lane, so its extension limbs are provably - // zero: they are NOT committed, and the memory read carries literal zeros - // there. Being a column is what puts it in K, and the pointer-relative - // address it forms on the bus, `p·obe`, is a K product for the same reason. + // The pointer, which forms the pointer-relative address `p·o2` on the bus. pub const P: usize = 7; - // The local cell, a full 192-bit word. The store target is DERIVED from it, - // the two flags, `pc` and `fp`, so it is no column. - pub const V3_LO: usize = 8; - pub const V3_HI: usize = 9; - pub const V3_TOP: usize = 10; - pub const R1: usize = 11; - pub const R2: usize = 12; - pub const R3: usize = 13; - pub const RBC: usize = 14; - pub const N: usize = 15; + // The local cell. The store target is DERIVED from it, the two flags, `pc` and + // `fp`, so it is no column. + pub const V3: usize = 8; + pub const R1: usize = 9; + pub const R2: usize = 10; + pub const R3: usize = 11; + pub const RBC: usize = 12; + pub const N: usize = 13; } -/// The stored word's three K-lanes as forms: -/// `v_2 = (1+f_pc+f_fp)·v_3 + f_pc·(g²·pc) + f_fp·fp`, the flag-selected source -/// of §sec:tab-deref. The `pc` source is the virtual return target `g²·pc`, a -/// free `×g²` on the product coordinate. Only the low lane takes the two K-valued -/// sources; the upper two are the gated local lanes alone. -fn deref_store() -> [Coord; 3] { +/// The stored word as a form: `v_2 = (1+f_pc+f_fp)·v_3 + f_pc·(g²·pc) + f_fp·fp`, the +/// flag-selected source of §sec:tab-deref. The `pc` source is the virtual return +/// target `g²·pc`, a free `×g²` on the product coordinate. +fn deref_store() -> Coord { use deref::*; - let gated = |v: usize| vec![Col(v), Prod(FPC, v, 0), Prod(FFP, v, 0)]; - let mut lo = gated(V3_LO); - lo.push(Prod(FPC, PC, 2)); - lo.push(Prod(FFP, FP, 0)); - [Coord::Sum(lo), Coord::Sum(gated(V3_HI)), Coord::Sum(gated(V3_TOP))] + Coord::Sum(vec![ + Col(V3), + Prod(FPC, V3, 0), + Prod(FFP, V3, 0), + Prod(FPC, PC, 2), + Prod(FFP, FP, 0), + ]) } impl Table for DerefTable { @@ -640,45 +658,33 @@ impl Table for DerefTable { f.state_step(PC, FP); f.bytecode(PC, RBC, OP_DEREF, &[Col(O1), Col(O2), Col(O3), Col(FPC), Col(FFP)]); // The pointer cell and the local cell are frame-relative; the store target - // is pointer-relative, so its address is `p·obe`, and its value is the + // is pointer-relative, so its address is `p·o2`, and its value is the // flag-selected source rather than a column. - f.memory_k(Prod(FP, O1, 0), R1, P); - f.memory_coords(Prod(P, O2, 0), R2, deref_store()); - f.memory(Prod(FP, O3, 0), R3, V3_LO, V3_HI, V3_TOP); + f.memory(Prod(FP, O1, 0), R1, Col(P)); + f.memory(Prod(P, O2, 0), R2, deref_store()); + f.memory(Prod(FP, O3, 0), R3, Col(V3)); } fn fill(&self, ctx: &FillCtx, out: &mut [ColumnOut]) { use deref::*; let rows = &ctx.trace.deref; - // Offsets and store mode are the instruction's. Neither the store target's - // address nor its value is committed. let ins = |r: &Drow| match ctx.prog[r.pc as usize] { Op::Deref { o1, o2, o3, mode } => (o1, o2, o3, mode), op => unreachable!("a DEREF row's pc {} holds {op:?}", r.pc), }; ctx.col(out, rows, PC, |r| ctx.g_at(r.pc)); ctx.col(out, rows, FP, |r| ctx.g_at(r.fp)); - debug_assert!( - rows.iter().all(|r| { - let p = ctx.mem[(r.fp + ins(r).0) as usize]; - p.c1 == 0 && p.c2 == 0 - }), - "deref pointer must be K-valued" - ); - // The three offsets, the two mode flags, the pointer lane and the local - // word all follow from ONE bytecode decode. + // The three offsets, the two mode flags, the pointer and the local cell all + // follow from ONE bytecode decode. ctx.cols(out, rows, O1, |r| { let (o1, o2, o3, mode) = ins(r); - let v3 = ctx.limbs(r.fp + o3); [ ctx.g_at(o1), ctx.g_at(o2), ctx.g_at(o3), mode.f_pc(), mode.f_fp(), - F64(ctx.mem[(r.fp + o1) as usize].c0), - v3[0], - v3[1], - v3[2], + ctx.cell(r.fp + o1), + ctx.cell(r.fp + o3), ] }); ctx.cols(out, rows, R1, |r| [r.r1, r.r2, r.r3]); @@ -696,11 +702,6 @@ mod jump { pub const OC: usize = 2; pub const OD: usize = 3; pub const OF: usize = 4; - // The condition, destination and frame words are all K-valued, so each is a - // SINGLE lane read through `memory_k`: bus balance forces the stored words - // into K, exactly as for the DEREF pointer. A guest branches on g-powers, - // never on an arbitrary word: `assert a != b` takes an inverse hint instead - // of a branch (§sec:prog-div-ne). pub const V_COND: usize = 5; pub const V_PC: usize = 6; pub const V_FP: usize = 7; @@ -709,8 +710,7 @@ mod jump { pub const RF: usize = 10; pub const RBC: usize = 11; // Local witness columns (committed, never flushed): the inverse hint `w = c⁻¹` - // and the taken indicator `b = [c ≠ 0]` it certifies (the `JUMP` table in - // `doc/leanvm/body/07-instruction-tables.tex`). Both are single K lanes. + // and the taken indicator `b = [c ≠ 0]` it certifies. pub const W: usize = 12; pub const B: usize = 13; pub const N: usize = 14; @@ -750,9 +750,9 @@ impl Table for JumpTable { OP_JUMP, &[Col(OC), Col(OD), Col(OF), Const(F64::ZERO), Const(F64::ZERO)], ); - f.memory_k(Prod(FP, OC, 0), RC, V_COND); - f.memory_k(Prod(FP, OD, 0), RD, V_PC); - f.memory_k(Prod(FP, OF, 0), RF, V_FP); + f.memory(Prod(FP, OC, 0), RC, Col(V_COND)); + f.memory(Prod(FP, OD, 0), RD, Col(V_PC)); + f.memory(Prod(FP, OF, 0), RF, Col(V_FP)); } fn fill(&self, ctx: &FillCtx, out: &mut [ColumnOut]) { use jump::*; @@ -761,33 +761,29 @@ impl Table for JumpTable { Op::Jump { oc, od, of } => (oc, od, of), op => unreachable!("a JUMP row's pc {} holds {op:?}", r.pc), }; - let cell = |r: &Jrow, o: u32| ctx.mem[(r.fp + o) as usize]; - let cond = |r: &Jrow| cell(r, ins(r).0); + let cond = |r: &Jrow| ctx.cell(r.fp + ins(r).0); ctx.col(out, rows, PC, |r| ctx.g_at(r.pc)); ctx.col(out, rows, FP, |r| ctx.g_at(r.fp)); // The three offsets and the three cells they name come out of ONE decode. - // Those cells are K-valued on every row, taken or not (`cpu::execute` - // rejects anything else), so each is one lane and the memory flush carries - // literal zeros above it. ctx.cols(out, rows, OC, |r| { let (oc, od, of) = ins(r); [ ctx.g_at(oc), ctx.g_at(od), ctx.g_at(of), - F64(cell(r, oc).c0), - F64(cell(r, od).c0), - F64(cell(r, of).c0), + ctx.cell(r.fp + oc), + ctx.cell(r.fp + od), + ctx.cell(r.fp + of), ] }); // The is-nonzero witness `w = c⁻¹` (0 where c = 0) for every row, in one - // batched Montgomery inversion. `prefix[i]` is the - // running product of the nonzero conditions before row `i`, so `acc` ends - // as their full product (nonzero, hence invertible). The taken indicator - // `b = [c ≠ 0]` falls out of the same pass, so it costs no extra decode. + // batched Montgomery inversion. `prefix[i]` is the running product of the + // nonzero conditions before row `i`, so `acc` ends as their full product + // (nonzero, hence invertible). The taken indicator `b = [c ≠ 0]` falls out + // of the same pass, so it costs no extra decode. let (w, b) = { - let mut acc = F192::ONE; - let mut prefix: Vec = Vec::with_capacity(rows.len()); + let mut acc = F64::ONE; + let mut prefix: Vec = Vec::with_capacity(rows.len()); let mut b = vec![F64::ZERO; rows.len()]; for (i, r) in rows.iter().enumerate() { prefix.push(acc); @@ -798,7 +794,7 @@ impl Table for JumpTable { } } let mut inv = acc.inv(); - let mut w = vec![F192::ZERO; rows.len()]; + let mut w = vec![F64::ZERO; rows.len()]; for (i, r) in rows.iter().enumerate().rev() { let c = cond(r); if !c.is_zero() { @@ -808,7 +804,7 @@ impl Table for JumpTable { } (w, b) }; - ctx.cols_at(out, rows.len(), W, |i| [F64(w[i].c0), b[i]]); + ctx.cols_at(out, rows.len(), W, |i| [w[i], b[i]]); ctx.cols(out, rows, RC, |r| [r.rc, r.rd, r.rf]); ctx.col(out, rows, RBC, |r| r.bytecode_read); } @@ -816,141 +812,95 @@ impl Table for JumpTable { // ---- BLAKE2s ------------------------------------------------------------------ -/// `BLAKE2s` (“BLAKE2s” in `doc/leanvm/body/07-instruction-tables.tex`): one standard compression. The four 128-bit message -/// chunks are addressed *independently* at `fp·o_i` (`o_i = g^{ins[i]}`), each a -/// single cell, with no forced contiguity between chunks, so a caller hashing e.g. -/// `(tweak, pp)` need not copy them into adjacent cells. The chaining value and the -/// 32-byte output each occupy two consecutive cells, based at `fp·o_cv` and -/// `fp·o_c`, and the metadata is one more cell at `fp·o_md`, so the row reads nine -/// cells in all. No address is committed: each rides the bus as the product `fp·o_X` -/// (§sec:m3). The compression relating output words to input words carries no -/// table constraint either: it is proven by flock's R1CS validity via `q_flock` +/// `BLAKE2s` (§sec:tab-blake2s): one standard compression. The four 128-bit message +/// chunks are addressed *independently* at `fp·o_i`, each spanning that cell and its +/// successor, so a caller hashing e.g. `(tweak, pp)` need not copy them into +/// adjacent cells. The digest and the chaining value span four consecutive cells, +/// the metadata two, so the row reads eighteen cells. No address is committed: each +/// rides the bus as the product `fp·o·g^k` (§sec:m3). The compression relating +/// output words to input words is proven by flock's R1CS via `q_flock` /// (§hash_flock), which leaves this table with no identity of its own. /// -/// A 128-bit chunk is two flock 64-bit words (lo, hi lanes), so the eighteen -/// memory-borne flock words are eighteen value LANE columns over the nine cells. -/// They are listed in `n_committed_columns` (they need a local index for the -/// flushes and are filled from the trace for the bus), but `cpu` treats them as -/// VIRTUAL (not committed) and routes their bus claims to `q_flock`, which already -/// holds those words (see [`BLAKE2S_VALUE_COLS`]). +/// The eighteen cells are eighteen value columns. They are listed in +/// `n_committed_columns` (they need a local index for the flushes and are filled +/// from the trace for the bus), but `cpu` treats them as VIRTUAL (not committed) and +/// routes their bus claims to `q_flock`, which already holds those words (see +/// [`BLAKE2S_VALUE_COLS`]). struct Blake2sTable; pub(crate) mod blake2st { pub const PC: usize = 0; pub const FP: usize = 1; - pub const O_M0: usize = 2; // operand g-powers (offsets) of the four message cells … - pub const O_M1: usize = 3; - pub const O_M2: usize = 4; - pub const O_M3: usize = 5; - pub const O_CV: usize = 6; // … the chaining-value base … - pub const O_OUT: usize = 7; // … the output base … - pub const O_MD: usize = 8; // … and the metadata cell - // The eighteen flock lanes: a (lo, hi) pair for each of the nine cells, - // the four message cells first, then the output pair, the chaining-value - // pair, and last the metadata cell's counter and flag lanes. - pub const V_M0: usize = 9; // m0.lo, m0.hi, m1.lo, m1.hi - pub const V_M2: usize = 13; // m2.lo, m2.hi, m3.lo, m3.hi - pub const V_OUT0: usize = 17; // out0.lo, out0.hi, out1.lo, out1.hi - pub const V_CV0: usize = 21; // cv0.lo, cv0.hi, cv1.lo, cv1.hi - pub const MD0: usize = 25; // metadata: the counter lane … - pub const MD1: usize = 26; // … and the final ‖ last_node lane - pub const R_M0: usize = 27; // one read count per cell: the four message cells … - pub const R_M1: usize = 28; - pub const R_M2: usize = 29; - pub const R_M3: usize = 30; - pub const R_CV0: usize = 31; // … the two chaining-value cells … - pub const R_CV1: usize = 32; - pub const R_OUT0: usize = 33; // … the two output cells … - pub const R_OUT1: usize = 34; - pub const R_MD: usize = 35; // … and the metadata cell. - pub const RBC: usize = 36; - pub const N: usize = 37; + pub const O_M0: usize = 2; // operand g-powers of the four message chunks … + pub const O_CV: usize = 6; // … the chaining value … + pub const O_OUT: usize = 7; // … the digest … + pub const O_MD: usize = 8; // … and the metadata. + // The eighteen value lanes, one per cell: the message chunks' eight, then the + // digest's four, the chaining value's four and the metadata's counter and flags. + pub const V0: usize = 9; + // One read count per value lane, in the same order. + pub const R0: usize = 27; + pub const RBC: usize = 45; + pub const N: usize = 46; } +/// The operand and the offset from it of each value lane's cell, in lane order. +const BLAKE2S_LANE_CELLS: [(usize, u32); 18] = { + use blake2st::*; + let mut cells = [(0, 0); 18]; + let mut i = 0; + while i < 18 { + cells[i] = match i { + 0..8 => (O_M0 + i / 2, (i % 2) as u32), + 8..12 => (O_OUT, (i - 8) as u32), + 12..16 => (O_CV, (i - 12) as u32), + _ => (O_MD, (i - 16) as u32), + }; + i += 1; + } + cells +}; + impl Table for Blake2sTable { fn n_committed_columns(&self) -> usize { blake2st::N } fn count_columns(&self) -> &'static [usize] { - use blake2st::*; - &[R_M0, R_M1, R_M2, R_M3, R_CV0, R_CV1, R_OUT0, R_OUT1, R_MD, RBC] + const COUNTS: [usize; 19] = { + let mut c = [blake2st::RBC; 19]; + let mut i = 0; + while i < 18 { + c[i] = blake2st::R0 + i; + i += 1; + } + c + }; + &COUNTS } fn flushes(&self, f: &mut FlushBuilder) { use blake2st::*; f.state_step(PC, FP); - f.bytecode( - PC, - RBC, - OP_BLAKE2S, - &[ - Col(O_M0), - Col(O_M1), - Col(O_M2), - Col(O_M3), - Col(O_CV), - Col(O_OUT), - Col(O_MD), - ], - ); - // Nine cell reads: four independent 128-bit message cells, the chaining - // value's two consecutive cells (ACV, g·ACV), the output's two - // consecutive cells (AC, g·AC), and the metadata cell. Each carries its - // chunk's two lanes with a literal-zero top limb (`memory_128`), so the - // canonical embedding is proof-enforced and the zero limbs are never - // committed. A consecutive cell is a free ×g on the product's g-power. - f.memory_128(Prod(FP, O_M0, 0), R_M0, V_M0, V_M0 + 1); - f.memory_128(Prod(FP, O_M1, 0), R_M1, V_M0 + 2, V_M0 + 3); - f.memory_128(Prod(FP, O_M2, 0), R_M2, V_M2, V_M2 + 1); - f.memory_128(Prod(FP, O_M3, 0), R_M3, V_M2 + 2, V_M2 + 3); - f.memory_128(Prod(FP, O_CV, 0), R_CV0, V_CV0, V_CV0 + 1); - f.memory_128(Prod(FP, O_CV, 1), R_CV1, V_CV0 + 2, V_CV0 + 3); - f.memory_128(Prod(FP, O_OUT, 0), R_OUT0, V_OUT0, V_OUT0 + 1); - f.memory_128(Prod(FP, O_OUT, 1), R_OUT1, V_OUT0 + 2, V_OUT0 + 3); - // The metadata rides the memory bus like every other operand: the read - // is what binds flock's counter and flag inputs, so a compile-time - // counter is pinned by the `SET` immediate that wrote the cell. - f.memory_128(Prod(FP, O_MD, 0), R_MD, MD0, MD1); + f.bytecode(PC, RBC, OP_BLAKE2S, &std::array::from_fn::<_, 7, _>(|i| Col(O_M0 + i))); + // A successor cell is a free ×g^k on the address product. The metadata rides + // the memory bus like every other operand: the read is what binds flock's + // counter and flag inputs. + for (lane, &(operand, k)) in BLAKE2S_LANE_CELLS.iter().enumerate() { + f.memory(Prod(FP, operand, k), R0 + lane, Col(V0 + lane)); + } } fn fill(&self, ctx: &FillCtx, out: &mut [ColumnOut]) { use blake2st::*; let rows = &ctx.trace.blake2s; - let ad = |r: &Brow| blake2s_addresses(ctx.prog, r); ctx.col(out, rows, PC, |r| ctx.g_at(r.pc)); ctx.col(out, rows, FP, |r| ctx.g_at(r.fp)); - // O_M0..O_MD are the seven base addresses' offsets, from the instruction decode. - ctx.cols(out, rows, O_M0, |r| ad(r).map(|a| ctx.g_at(a - r.fp))); - // The eighteen memory-borne flock words are the nine cells' lo/hi lanes: - // the four message cells, then the cv pair, the output pair and the - // metadata. A cell's two lanes are one read, so each group of four takes two. - let word_pair = |c0: u32, c1: u32| { - let (w0, w1) = (ctx.mem[c0 as usize], ctx.mem[c1 as usize]); - [F64(w0.c0), F64(w0.c1), F64(w1.c0), F64(w1.c1)] - }; - ctx.cols(out, rows, V_M0, |r| { - let a = ad(r); - word_pair(a[0], a[1]) - }); - ctx.cols(out, rows, V_M2, |r| { - let a = ad(r); - word_pair(a[2], a[3]) - }); - ctx.cols(out, rows, V_OUT0, |r| { - let a = ad(r); - word_pair(a[5], a[5] + 1) + ctx.cols(out, rows, O_M0, |r: &Brow| match ctx.prog[r.pc as usize] { + Op::Blake2s { ins, cv, out, md } => [ins[0], ins[1], ins[2], ins[3], cv, out, md].map(|o| ctx.g_at(o)), + op => unreachable!("a BLAKE2s row's pc {} holds {op:?}", r.pc), }); - ctx.cols(out, rows, V_CV0, |r| { - let a = ad(r); - word_pair(a[4], a[4] + 1) - }); - ctx.cols(out, rows, MD0, |r| { - let md = ctx.mem[ad(r)[6] as usize]; - [F64(md.c0), F64(md.c1)] - }); - ctx.cols(out, rows, R_M0, |r| { - [ - r.ra[0], r.ra[1], r.rb[0], r.rb[1], r.rcv[0], r.rcv[1], r.rc[0], r.rc[1], r.rmd, - ] + ctx.cols(out, rows, V0, |r| { + blake2s_cells(ctx.prog, r.pc, r.fp).map(|a| ctx.cell(a)) }); + ctx.cols(out, rows, R0, |r| r.r); ctx.col(out, rows, RBC, |r| r.bytecode_read); } } @@ -962,10 +912,9 @@ mod tests { const C: F192 = F192::new(0x0123_4567_89ab_cdef, 0xfeed_face_dead_beef, 0x1111_2222_3333_4444); - /// The hand-unrolled tower product IS `E`'s multiplication, lane by lane. Both - /// `MUL`'s result coordinate and `JUMP`'s inverse identity are written out from - /// [`TOWER_LANES`], and neither can be checked against the field at run time (one - /// is a bus coordinate, the other a `K` identity), so pin the unrolling here. + /// The hand-unrolled tower product IS `E`'s multiplication, lane by lane. + /// `MUL192`'s result coordinate is written out from [`TOWER_LANES`] and cannot + /// be checked against the field at run time, so pin the unrolling here. #[test] fn unrolled_tower_product_matches_field_product() { let lanes = |v: F192| [F64(v.c0), F64(v.c1), F64(v.c2)]; @@ -979,15 +928,9 @@ mod tests { #[test] fn jump_identities_bind_indicator_and_inverse() { let pows = powers(F192::new(0x9e37_79b9_7f4a_7c15, 0x1234_5678_9abc_def0, 7), 2); - // The condition is K-valued (`memory_k` on its read, §sec:tab-jump), so the - // pair is single-lane: `w = c⁻¹` in K too. for cond in [F64::ZERO, F64(0x9e37_79b9_7f4a_7c15)] { let mut row = vec![F64::ZERO; jump::N]; - let w = if cond.is_zero() { - F64::ZERO - } else { - F64(F192::from(cond).inv().c0) - }; + let w = if cond.is_zero() { F64::ZERO } else { cond.inv() }; row[jump::V_COND] = cond; row[jump::W] = w; row[jump::B] = if cond.is_zero() { F64::ZERO } else { F64::ONE }; diff --git a/crates/lean_vm/tests/verifiers/python_verifier.rs b/crates/lean_vm/tests/verifiers/python_verifier.rs index 452b39e2e..1bf2e0e93 100644 --- a/crates/lean_vm/tests/verifiers/python_verifier.rs +++ b/crates/lean_vm/tests/verifiers/python_verifier.rs @@ -6,7 +6,7 @@ use fiat_shamir::transcript::RawProof; use lean_compiler::{compile, parse_with_replacements}; use lean_vm::cpu::{prove, verify}; -use primitives::field::{F64, F192, g_pow}; +use primitives::field::{F64, g_pow}; use std::collections::BTreeMap; use std::path::Path; use std::process::{Command, Output}; @@ -18,14 +18,14 @@ from snark_lib import * LOOP_STEPS = LOOP_STEPS_PLACEHOLDER def mix(value, tag): - # A JUMP condition is K-valued (a g-power here), never an arbitrary word. + # The tag is a g-power, so the JUMP is always taken. if tag == 0: return value return value * GEN + value def main(): - seed = [5, 7] - digest = StackBuf(2) + seed = [5, 0, 7, 0] + digest = StackBuf(4) blake2s(seed, seed, digest) chain = HeapBuf(LOOP_STEPS + 1) @@ -36,6 +36,8 @@ def main(): public = GEN ** 0 public[1] = chain[GEN ** LOOP_STEPS] public[GEN] = mix(digest[1], GEN ** 0) + public[GEN ** 2] = digest[2] + public[GEN ** 3] = digest[3] return "#; @@ -76,30 +78,20 @@ fn python_verify(directory: &Path, bytecode: &Path, public_input: &Path, raw: &R .expect("run native Python verifier") } -fn public_input() -> [F192; 2] { +fn public_input() -> [F64; 4] { use lean_vm::hash_flock::{FINAL_FLAG, IV, PINNED_T, compression, digest, metadata}; let seed = [F64(5), F64::ZERO, F64(7), F64::ZERO]; - let metadata = metadata(PINNED_T, FINAL_FLAG, 0); - let digest = digest(&compression(seed, seed, IV, metadata)); - let digest = [ - F192::new(digest[0].0, digest[1].0, 0), - F192::new(digest[2].0, digest[3].0, 0), - ]; + let digest = digest(&compression(seed, seed, IV, metadata(PINNED_T, FINAL_FLAG, 0))); + let generator = g_pow(1); + let mix = |value: F64| value * generator + value; let mut value = digest[0]; - let mut index = F192::ONE; - let generator = F192::from(g_pow(1)); + let mut index = F64::ONE; for _ in 0..LOOP_STEPS { - let candidate = value + index; - let product = candidate * generator; - value = (if product == F192::ZERO { - candidate - } else { - product + candidate - }) + index; + value = mix(value + index) + index; index *= generator; } - [value, digest[1] * generator + digest[1]] + [value, mix(digest[1]), digest[2], digest[3]] } #[test] @@ -121,17 +113,14 @@ fn test_python_verifier() { let bytecode_path = directory.join("bytecode.bin"); let public_input_path = directory.join("public_input.bin"); let encoded = bincode::serialize(&proof).expect("serialize proof"); - // The statement the verifier takes is the bytecode multilinear plus 256 bits - // of public input, not a structured program. + // The statement the verifier takes is the bytecode multilinear plus four + // public words, not a structured program. let table: Vec = lean_vm::cpu::layout::bytecode_table(&program.prog) .iter() .flat_map(|w| w.0.to_le_bytes()) .collect(); std::fs::write(&bytecode_path, &table).expect("write bytecode"); - let pi: Vec = public_input - .iter() - .flat_map(|v| [v.c0.to_le_bytes(), v.c1.to_le_bytes()].concat()) - .collect(); + let pi: Vec = public_input.iter().flat_map(|w| w.0.to_le_bytes()).collect(); std::fs::write(&public_input_path, &pi).expect("write public input"); let verification_started = Instant::now(); diff --git a/crates/rec_aggregation/guests/lean_ethereum.py b/crates/rec_aggregation/guests/lean_ethereum.py index 1e7739d21..677b4d84b 100644 --- a/crates/rec_aggregation/guests/lean_ethereum.py +++ b/crates/rec_aggregation/guests/lean_ethereum.py @@ -13,24 +13,22 @@ from snark_lib import * # ---------------------------------------------------------------- proof stream -# The proof stream rides ONE padded witness hint (the guest walks only the prefix -# the shape dictates); binding always comes from the per-word absorbs. +# The proof stream rides ONE padded witness hint of 64-bit words, three to a +# transcript scalar (the guest walks only the prefix the shape dictates); binding +# always comes from the per-scalar absorbs. STREAM_CAP = STREAM_CAP_PLACEHOLDER MIN_LOG_MEM = MIN_LOG_MEM_PLACEHOLDER INV_GEN = INV_GEN_PLACEHOLDER # ------------------------------------------------------------- the field, GF(2^192) -# Tower F192 = F64[Y]/(Y^3+Y+1). Y_TOWER embeds Y for reassembling -# e192(lo,hi,top) = lo + hi*Y + top*Y², and Y_INV also deduces a top limb at the -# opening boundary once the low and high limbs are transmitted. COORD_BASIS is the -# coordinate basis e_i of F192 (which spans the WHOLE field, unlike the g-power -# basis GEN**i, which spans only F64): hint_decompose_bits emits a word's -# coordinate bits, so a value reconstructs as Σ b_i·COORD_BASIS[i]. +# A memory cell is one 64-bit word. An element of F192 = F64[Y]/(Y^3+Y+1) is a run +# of three cells, its limbs low first, computed with add192/mul192/div192; heap +# buffers of them have stride three. A product of a word and an element is +# limb-wise, three 64-bit multiplies and no reduction (`scale`). FIELD_BITS = 192 BASE_FIELD_BITS = 64 -Y_TOWER = Y_TOWER_PLACEHOLDER -Y_INV = Y_INV_PLACEHOLDER -COORD_BASIS = COORD_BASIS_PLACEHOLDER +ONE = f192(1, 0, 0) +ZERO = f192(0, 0, 0) # Six challenges compose the F2-linear map that batches the 192 transposed # ring-switch coordinates. RING_MAP_SHIFTS = [32, 16, 8, 4, 2, 1] @@ -46,46 +44,41 @@ LOG_WORD_BITS = 6 # ------------------------------------------------------------------- Fiat-Shamir -# Every absorbed block carries its domain tag in lane 3, which is exactly -# `fs_compress`'s `tail` argument, so a role is never smuggled through the data -# lanes. The seeding block has none, being fixed at the head of the chain. +# The state is four words. Every absorbed block carries a scalar's three limbs +# and its domain tag in word 3, so a role is never smuggled through the data +# words. The seeding block has none, being fixed at the head of the chain. DS_OBSERVE = 1 DS_SQ = 2 DS_POW_BASE = 3 DS_POW_NONCE = 4 # ----------------------------------------------------------- loop-carried chains -# A loop whose state is several fields keeps ONE heap run of N cells per iteration, -# so a step is one pointer multiply rather than one per field. -PAIR_SLOTS = 2 # the Fiat-Shamir state pair alone +# A loop whose state is several values keeps ONE heap run per iteration, so a +# step is one pointer multiply rather than one per value. Each run leads with the +# four-word Fiat-Shamir state. +FS_SLOTS = 4 # A Fiat-Shamir chain carrying one accumulator (the claim batching loops). -ACC_FS0 = 0 -ACC_FS1 = 1 -ACC_VALUE = 2 -ACC_SLOTS = 3 +ACC_VALUE = 4 +ACC_SLOTS = 7 # One sumcheck round of the bus GKR, the table batch, or flock's multilinear rounds. -ROUND_FS0 = 0 -ROUND_FS1 = 1 -ROUND_CURSOR = 2 -ROUND_CLAIM = 3 -ROUND_SLOTS = 4 +ROUND_CURSOR = 4 +ROUND_CLAIM = 5 +ROUND_SLOTS = 8 # One product layer of the bus GKR. -LAYER_FS0 = 0 -LAYER_FS1 = 1 -LAYER_CURSOR = 2 -LAYER_PUSH = 3 -LAYER_PULL = 4 -LAYER_COUNT = 5 -LAYER_LAMBDA = 6 -LAYER_ROW = 7 -LAYER_POS = 8 -LAYER_SLOTS = 9 +LAYER_CURSOR = 4 +LAYER_PUSH = 5 +LAYER_PULL = 8 +LAYER_COUNT = 11 +LAYER_LAMBDA = 14 +LAYER_ROW = 17 +LAYER_POS = 18 +LAYER_SLOTS = 19 # One OOD sample of a WHIR level. OOD_BETA = 0 -OOD_Y = 1 -OOD_C0 = 2 -OOD_C2 = 3 -OOD_SLOTS = 4 +OOD_Y = 3 +OOD_C0 = 6 +OOD_C2 = 9 +OOD_SLOTS = 12 # ---------------------------------------------------- the bus: sides and blocks # GKR sides. The layer counts mu_s are hinted and certified from the block kappas; @@ -160,7 +153,7 @@ BYTECODE_COLS = BYTECODE_COLS_PLACEHOLDER LOG2_BYTECODE_COLS = LOG2_BYTECODE_COLS_PLACEHOLDER -# --------------------------------------------------------------- the six tables +# ------------------------------------------------------------- the eight tables # The table sumcheck's batch carries EVERY committed column of a table, because its # bus forms read the flushed ones and its constraint the rest; TABLE_COLS_CAP caps # the evaluation frame. ETA_OFFSET[t] starts table t's disjoint range of zc_xi @@ -168,14 +161,16 @@ # every table, and that sharing is what makes the batch's target derivable from the # three leaf claims. FLOORS[t] is the table's tau floor (BLAKE2s is sized to # flock's instance count, >= 2^3). -TABLE_XOR = 0 -TABLE_MUL = 1 +TABLE_XOR64 = 0 +TABLE_MUL64 = 1 TABLE_SET = 2 TABLE_DEREF = 3 TABLE_JUMP = 4 TABLE_BLAKE2s = 5 +TABLE_XOR192 = 6 +TABLE_MUL192 = 7 N_TABLES = N_TABLES_PLACEHOLDER -FLOORS = [0, 0, 0, 0, 0, 3] +FLOORS = [0, 0, 0, 0, 0, 3, 0, 0] N_TABLE_COLS = N_TABLE_COLS_PLACEHOLDER TABLE_COLS_CAP = TABLE_COLS_CAP_PLACEHOLDER ETA_OFFSET = ETA_OFFSET_PLACEHOLDER @@ -189,7 +184,7 @@ # denominator per domain (combined, S). The zerocheck point/round buffers are sized # at runtime in the exponent (m = K_LOG + tau_5 and m - 6, both certified); # LINCHECK_ROUNDS = K_LOG - K_SKIP is protocol-fixed and PIN_COLUMN is the -# const-pin column. +# const-pin column. FIXED_CHALLENGES and PHI8_NODES hold three words an element. K_SKIP = K_SKIP_PLACEHOLDER N_FIXED_CHALLENGE_ROUNDS = N_FIXED_CHALLENGE_ROUNDS_PLACEHOLDER FIXED_CHALLENGES = FIXED_CHALLENGES_PLACEHOLDER @@ -243,7 +238,6 @@ LIG_FOLDS = LIG_FOLDS_PLACEHOLDER LIG_INTERLEAVE = LIG_INTERLEAVE_PLACEHOLDER LIG_LEAF_BLOCKS = LIG_LEAF_BLOCKS_PLACEHOLDER -LIG_PACKED_ROW_CAP = LIG_PACKED_ROW_CAP_PLACEHOLDER LIG_ROW_CAP = LIG_ROW_CAP_PLACEHOLDER LIG_PATH_CAP = LIG_PATH_CAP_PLACEHOLDER LIG_TREE_DEPTH = LIG_TREE_DEPTH_PLACEHOLDER @@ -254,7 +248,7 @@ LIG_RESIDUAL_PREFIX_LEN = LIG_RESIDUAL_PREFIX_LEN_PLACEHOLDER LIG_FOLDS_OFF = LIG_FOLDS_OFF_PLACEHOLDER LIG_VANISH_OFF = LIG_VANISH_OFF_PLACEHOLDER -LIG_VANISH_VALS = LIG_VANISH_VALS_PLACEHOLDER +LIG_VANISH_VALS = LIG_VANISH_VALS_PLACEHOLDER # words: the novel basis lives in F64 LIG_VANISH_INVS = LIG_VANISH_INVS_PLACEHOLDER LIG_N_CANDIDATES = LIG_N_CANDIDATES_PLACEHOLDER LIG_MIN_SHIFT_INV = LIG_MIN_SHIFT_INV_PLACEHOLDER @@ -275,10 +269,11 @@ # ------------------------------------------------------ statements and deferral # A node defers three claims on fixed polynomials. DEFER_SIZE is the region one -# sub-proof's verification exports (a bytecode point plus the flock lincheck data, -# see verify_sub's defer_out layout); DEFER_STMT_* index the batched claims a -# node's OWN statement carries. The Fiat-Shamir seed rides the public input rather -# than being baked, so one compiled guest verifies proofs of any inner program. +# sub-proof's verification exports, in field elements (a bytecode point plus the +# flock lincheck data, see verify_sub's defer_out layout); DEFER_STMT_* index the +# batched claims a node's OWN statement carries. The Fiat-Shamir seed rides the +# statement rather than being baked, so one compiled guest verifies proofs of any +# inner program. BYTECODE_LOG = BYTECODE_LOG_PLACEHOLDER # log rows of the bytecode blocks DEFER_SIZE = DEFER_SIZE_PLACEHOLDER BYTECODE_VARS = BYTECODE_VARS_PLACEHOLDER # = BYTECODE_LOG + LOG2_BYTECODE_COLS @@ -297,65 +292,65 @@ DEFER_STMT_B_VALUE = BYTECODE_VARS + 2 + 2 * K_LOG AGG_SEED_0 = AGG_SEED_0_PLACEHOLDER AGG_SEED_1 = AGG_SEED_1_PLACEHOLDER -# The statement digest's preimage: the STMT_HEADER header values as the 16-byte -# cells they already are (the seed and the signer-set digest, which itself binds -# the epoch groups and every count), then the deferred cells' tower limbs, two to -# a cell and four cells to a 64-byte block. No domain tag: the seed leads, and it -# binds this bytecode and flock's R1CS. +AGG_SEED_2 = AGG_SEED_2_PLACEHOLDER +AGG_SEED_3 = AGG_SEED_3_PLACEHOLDER +# The statement digest's preimage: the STMT_HEADER header words (the seed, the +# signer-set digest, which itself binds the epoch groups and every count, and the +# DA root-list digest), then the deferred elements' three limbs each, zero-filled +# to whole blocks. No domain tag: the seed leads, and it binds this bytecode and +# flock's R1CS. STMT_HEADER = STMT_HEADER_PLACEHOLDER -STMT_DEFER_OFF = STMT_HEADER -STMT_ODD = STMT_ODD_PLACEHOLDER -STMT_PAIRS = STMT_PAIRS_PLACEHOLDER -STMT_PAD_CELLS = STMT_PAD_CELLS_PLACEHOLDER +STMT_PAD = STMT_PAD_PLACEHOLDER STMT_BLOCKS = STMT_BLOCKS_PLACEHOLDER -# The declared lists are hashed with plain BLAKE2s over a flat run of cells, 64 +# The declared lists are hashed with plain BLAKE2s over a flat run of words, 64 # bytes a compression. A block's byte counter is a runtime value and the ISA has no # integer addition, so it splits as in doc §sec:prog-byte-counter: a window of # SIGNERS_WINDOW blocks shares one base 64·SIGNERS_WINDOW·q, whose set bits all sit -# above the window's own offsets 64(j+1), so a block's metadata cell is one XOR. The +# above the window's own offsets 64(j+1), so a block's counter word is one XOR. The # base comes from the window loop's own counter, and the one block whose offset # overlaps it takes the next window's base instead. SIGNERS_WINDOW = SIGNERS_WINDOW_PLACEHOLDER SIGNERS_WINDOW_LOG = SIGNERS_WINDOW_LOG_PLACEHOLDER SIGNERS_MAX_WINDOWS = SIGNERS_MAX_WINDOWS_PLACEHOLDER SIGNERS_COUNT_BITS = SIGNERS_COUNT_BITS_PLACEHOLDER +# The lists a hash tail walks (`tail_blocks`). +LIST_PLAIN = 0 +LIST_KEYS = 1 +LIST_SPHINCS = 2 +LIST_CHILD_KEYS = 3 +LIST_CHILD_SPHINCS = 4 # BLAKE2s's parameterized initial chaining value, which every hash here starts from, -# and the metadata of a final block with a zero counter, to add a length into. +# and the flag word of a final block's metadata. BLAKE2S_IV_0 = BLAKE2S_IV_0_PLACEHOLDER BLAKE2S_IV_1 = BLAKE2S_IV_1_PLACEHOLDER +BLAKE2S_IV_2 = BLAKE2S_IV_2_PLACEHOLDER +BLAKE2S_IV_3 = BLAKE2S_IV_3_PLACEHOLDER MD_FINAL = MD_FINAL_PLACEHOLDER # ---------------------------------------------------------- XMSS (host-supplied) # Every 16-byte native value (tweak, digest, chain tip, sibling, public parameter) -# is one canonical 128-bit cell. XM_* are the tweaks minus their index field, as -# the host read them out of the native `make_tweak`, so the guest holds no byte -# layout of its own; XM_INDEX_WEIGHT[b] is what bit b of an index weighs, an index -# being its set bits summed. +# is two words, a 32-byte one (a message, a key) four. A tweak's first word holds +# its type and sub-position, its second the index at bit 32 (`xmss::make_tweak`): +# XM_* are the first words, as the host read them out of `make_tweak`, and +# XM_INDEX_WEIGHT[b] is what bit b of an index weighs in the second, an index being +# its set bits summed. A sub-position weighs XM_P_MUL a unit. V = V_PLACEHOLDER W = W_PLACEHOLDER TARGET_SUM = TARGET_SUM_PLACEHOLDER LOG_LIFETIME = LOG_LIFETIME_PLACEHOLDER CHAIN_LENGTH = 2 ** W CHAIN_STEPS = CHAIN_LENGTH - 1 -WORDS_PER_VALUE = 1 -WORDS_PER_BLOCK = 2 -# Tweak table (one 1-cell tweak per index): encoding | V·CHAIN_STEPS chain | -# wots-pk | merkle. Derived in-circuit, once per epoch group the statement carries. -N_TWEAKS = 1 + V * CHAIN_STEPS + 1 + LOG_LIFETIME -N_TWEAK_CELLS = WORDS_PER_VALUE * N_TWEAKS -WOTS_PK_TWEAK_IDX = 1 + V * CHAIN_STEPS -MERKLE_TWEAK_IDX = WOTS_PK_TWEAK_IDX + 1 -MERKLE_BIT_CELLS = WORDS_PER_VALUE * LOG_LIFETIME # one 1-cell bit word per level XM_ENC_TWEAK = XM_ENC_TWEAK_PLACEHOLDER XM_PK_TWEAK = XM_PK_TWEAK_PLACEHOLDER -XM_CHAIN_TWEAKS = XM_CHAIN_TWEAKS_PLACEHOLDER # indexed CHAIN_STEPS·i + s -XM_MERKLE_TWEAKS = XM_MERKLE_TWEAKS_PLACEHOLDER # indexed by level +XM_CHAIN_TWEAK = XM_CHAIN_TWEAK_PLACEHOLDER +XM_MERKLE_TWEAK = XM_MERKLE_TWEAK_PLACEHOLDER XM_INDEX_WEIGHT = XM_INDEX_WEIGHT_PLACEHOLDER -# Digits packed per digest lane: W bits each in GF(2^64)'s monomial budget (the -# lane's leftover top bits are ground to zero by the signer). +XM_P_MUL = XM_P_MUL_PLACEHOLDER +# Digits packed per digest word: W bits each in GF(2^64)'s monomial budget (the +# word's leftover top bits are ground to zero by the signer). DIGITS_PER_WORD = V / 2 -TIP_CELLS = WORDS_PER_VALUE * V -WOTS_PK_BLOCKS = (2 + V) / 4 # prefix (tweak, pp) + V tips, four cells a block +TIP_WORDS = 2 * V +WOTS_PK_BLOCKS = (2 + V) / 4 # prefix (tweak, pp) + V tips, eight words a block # ------------------------------------------------------ SPHINCS+ (host-supplied) # The scheme's own letters, prefixed SP_ where XMSS has the same one. @@ -371,16 +366,15 @@ SP_CHAIN_LENGTH = 2 ** SP_W SP_CHAIN_STEPS = SP_CHAIN_LENGTH - 1 SP_DIGITS_PER_WORD = SP_V / 2 -SP_TIP_CELLS = SP_V -SP_LEAF_BLOCKS = (2 + SP_V) / 4 # prefix (tweak, pp) + V tips, four cells a block +SP_TIP_WORDS = 2 * SP_V +SP_LEAF_BLOCKS = (2 + SP_V) / 4 # prefix (tweak, pp) + V tips, eight words a block SP_N_FTS = SP_K - 1 # the forest drops the last index's tree SP_ROOT_BLOCKS = (2 + SP_N_FTS) / 4 -# The message digest is h + k*a bits of a BLAKE2s output: the whole low cell and -# the low 48 bits of the high one. Decomposing the high cell's low lane covers -# them, so the buffer holds three lanes and the top 16 are never read. +# The message digest is h + k*a bits of a BLAKE2s output, within its first three +# words, so the bit buffer holds three words' decompositions. SP_BIT_LANES = 3 SP_BIT_CELLS = SP_BIT_LANES * BASE_FIELD_BITS -# Native tweak prefixes, including the protocol domain separator and type. +# Native tweak first words, including the protocol domain separator and type. SP_TW_CHAIN = SP_TW_CHAIN_PLACEHOLDER SP_TW_LEAF = SP_TW_LEAF_PLACEHOLDER SP_TW_NODE = SP_TW_NODE_PLACEHOLDER @@ -389,14 +383,14 @@ SP_TW_FTS_NODE = SP_TW_FTS_NODE_PLACEHOLDER SP_TW_FTS_ROOTS = SP_TW_FTS_ROOTS_PLACEHOLDER SP_TW_MSG = SP_TW_MSG_PLACEHOLDER -# Tweak layout: protocol_domain_sep | type | layer | zero | p | tree | index. -# Each 32-bit field stays within one 64-bit lane. +# Tweak layout: protocol_domain_sep | type | layer | zero | p in the first word, +# tree | index in the second, each 32-bit field within one word. SP_LAY_MUL = 2 ** 16 SP_P_MUL = 2 ** 32 -SP_TAU_POS = BASE_FIELD_BITS -SP_J_POS = BASE_FIELD_BITS + 32 +SP_TAU_POS = 0 +SP_J_POS = 32 SP_CHAIN_MUL = SP_CHAIN_LENGTH * SP_P_MUL # chain i's tweaks start at p = 2^w * i -# The encoding counter, LE_32 in the low four bytes of its cell: bounded by +# The encoding counter, LE_32 in the low four bytes of its word: bounded by # decomposing exactly that many bits, so the guest accepts no preimage the native # verifier cannot parse. SP_COUNTER_BITS = 32 @@ -420,10 +414,8 @@ DA_LOG_CELL = DA_LOG_CELL_PLACEHOLDER DA_MAX_ROWS = DA_MAX_ROWS_PLACEHOLDER DA_LOG_MAX_ROWS = DA_LOG_MAX_ROWS_PLACEHOLDER -DA_PAD_CELL_0 = DA_PAD_CELL_0_PLACEHOLDER -DA_PAD_CELL_1 = DA_PAD_CELL_1_PLACEHOLDER -DA_PAD_ROW_0 = DA_PAD_ROW_0_PLACEHOLDER -DA_PAD_ROW_1 = DA_PAD_ROW_1_PLACEHOLDER +DA_PAD_CELL = DA_PAD_CELL_PLACEHOLDER +DA_PAD_ROW = DA_PAD_ROW_PLACEHOLDER DA_CELL = 2 ** DA_LOG_CELL # symbols in a cell DA_BLOCK_BITS = DA_LOG_K + 1 - DA_LOG_CELL # log of the cells per row @@ -433,110 +425,87 @@ DA_ROW_BLOCKS = DA_PREFIX_CELLS // 2 # BLAKE2s blocks in one row digest DA_TREE_ARMS = DA_LOG_MAX_ROWS + 1 # tree depths the row count can dispatch to -# =================================== field packing ================================== -# Serializing a word means exposing its three K limbs, and the tower representation -# is unique, so proving that the exposed lanes are in K and weight back to the word -# is the whole check: no equality to make, one less hint to distrust. +# =================================== field helpers ================================== @inline -def pack64x2(a, b): - assert_in_k(a, b) - return a + Y_TOWER * b - - -def assert_canonical(word): - # Hint the low limb and derive the quotient by Y. Requiring both in K proves - # `word = lo + Y*hi`, hence that its top limb is zero. - lo = StackBuf(1) - hint_f192_limbs(lo, word) - hi = (word + lo[0]) * Y_INV - assert_in_k(lo[0], hi) - return 0 +def scale(s, x): + # A word times an element, limb by limb. + return [s * x[0], s * x[1], s * x[2]] @inline -def challenge_from_state(state): - # Both words are BLAKE2s outputs with zero top limbs. - # Hint d2 and derive d3 = (state[1] + d2)/Y. - # Requiring both in K binds d2 by the tower representation. - # The challenge uses d2; d3 is checked and discarded. - d2 = StackBuf(1) - hint_f192_limbs(d2, state[1]) - d3 = (state[1] + d2[0]) * Y_INV - assert_in_k(d2[0], d3) - return state[0] + Y_TOWER * Y_TOWER * d2[0] +def add64(x, s): + # An element plus a word: only the low limb moves. + return [x[0] + s, x[1], x[2]] -# ==================================== Fiat-Shamir =================================== +@inline +def fixed_challenge(i: Const): + return f192(FIXED_CHALLENGES[3 * i], FIXED_CHALLENGES[3 * i + 1], FIXED_CHALLENGES[3 * i + 2]) @inline -def fs_compress(state, scalar, tail, out): - # BLAKE2s requires block[0] to have zero top limb, binding the hinted top. - limbs = StackBuf(3) - hint_f192_limbs(limbs, scalar) - assert_in_k(limbs[2], tail) - block = StackBuf(2) - block[0] = scalar + Y_TOWER * Y_TOWER * limbs[2] - block[1] = limbs[2] + Y_TOWER * tail - blake2s(state, block, out) - return +def phi8(i: Const): + return f192(PHI8_NODES[3 * i], PHI8_NODES[3 * i + 1], PHI8_NODES[3 * i + 2]) -@inline -def obs(state, x): - # Bind one scalar into the chain: state <- compress(state, (x, DS_OBSERVE)). - # Returns the successor StackBuf; the call site aliases it (zero copies). - nb = StackBuf(2) - fs_compress(state, x, DS_OBSERVE, nb) - return nb +# ==================================== Fiat-Shamir =================================== @inline def fs_next(state, cursor): - # Fetch, observe and advance in one act: read the word under `cursor`, fold it - # into the state, and hand back the successor state, the word, AND the cursor - # stepped one word on. Reading and absorbing are inseparable here, so no - # proof-stream word can enter the computation unbound: the soundness invariant - # the whole guest rests on. All three returns alias into the caller for free. - x = cursor[GEN ** 0] - nb = obs(state, x) - return nb, x, cursor * GEN + # Fetch, observe and advance in one act: read the scalar under `cursor`, fold + # it into the state, and hand back the successor state, the scalar, AND the + # cursor stepped one scalar on. Reading and absorbing are inseparable here, so + # no proof-stream word can enter the computation unbound: the soundness + # invariant the whole guest rests on. All three returns alias into the caller. + block = StackBuf(4) + block[0:3] = cursor[0:3] + block[3] = DS_OBSERVE + nb = StackBuf(4) + blake2s(state, block, nb) + return nb, block[0:3], cursor * GEN ** 3 @inline -def squeeze_state(state): - nb = StackBuf(2) - blake2s(state, [0, Y_TOWER * DS_SQ], nb) - return nb +def fs_next_half(state, cursor): + # `fs_next` of a 128-bit half of a Merkle root, which rides the stream as a + # scalar with a zero top limb (merkle.rs `scalars_to_hash`). + block = StackBuf(4) + block[0:3] = cursor[0:3] + block[3] = DS_OBSERVE + assert block[2] == 0 + nb = StackBuf(4) + blake2s(state, block, nb) + return nb, block[0:3], cursor * GEN ** 3 @inline -def squeeze(state): - # Ratchet: the canonical 128+128 digest is the new state; its first three K - # lanes are reassembled as the F192 challenge. - nb = squeeze_state(state) - challenge = challenge_from_state(nb) - return nb, challenge +def obs(state, x): + # Bind one scalar into the chain: state <- compress(state, (x, DS_OBSERVE)). + nb = StackBuf(4) + blake2s(state, [x, DS_OBSERVE], nb) + return nb @inline -def absorb_nonce(state, x): - # Full-field grinding nonce absorb: [x.c0, x.c1, x.c2, DS_POW_NONCE]. - nb = StackBuf(2) - fs_compress(state, x, DS_POW_NONCE, nb) +def obs_at(state, ptr): + # `obs` of the scalar at a heap pointer, read straight into the block. + block = StackBuf(4) + block[0:3] = ptr[0:3] + block[3] = DS_OBSERVE + nb = StackBuf(4) + blake2s(state, block, nb) return nb @inline -def squeeze_step(state_0, state_1): - # `squeeze` exposing BOTH output words, so a query-squeeze loop can chain the - # state through a heap buffer. Returns (challenge, next_state_0, next_state_1). - state = [state_0, state_1] - next_state = squeeze_state(state) - challenge = challenge_from_state(next_state) - return challenge, next_state[0], next_state[1] +def squeeze(state): + # Ratchet: the digest is the new state, its first three words the challenge. + nb = StackBuf(4) + blake2s(state, [0, 0, 0, DS_SQ], nb) + return nb, nb[0:3] # ============================ bits, logs, and the exponent ========================== @@ -544,14 +513,14 @@ def squeeze_step(state_0, state_1): @inline def bind_bits(bits_ptr, value, n: Const): - # Tie an advice bit run back to the value it decomposes. Booleanity is a + # Tie an advice bit run back to the word it decomposes. Booleanity is a # write-once pin: the cell already holds the bit, so storing its square IS the # assert, one instruction shorter than a separate equality. acc = 0 for i in unroll(0, n): b = bits_ptr[GEN ** i] bits_ptr[GEN ** i] = b * b - acc += b * COORD_BASIS[i] + acc += b * 2 ** i assert acc == value return @@ -638,86 +607,94 @@ def log2_ceil_in_the_exponent(g_N, g_logs_pow2, g_squares, floor: Const, nbits: return g_log -def decode_query_bits(squeezed_word, positions_out, bit_ptrs_out, depth: Const): - # The squeezed word's bits are advice-decomposed HERE, boolean-constrained and - # tied back by reconstruction; each depth-bit group also becomes a query - # position (little-endian), with a pointer to its bit run (the Merkle direction - # bits). Each field word packs FIELD_BITS // depth positions. The bits live in - # FRAME cells, every index into them being compile-time, and `addr` names the - # run so the direction-bit pointers still reach it. +def decode_query_bits(squeezed: StackBuf(3), positions_out, bit_ptrs_out, depth: Const): + # The squeezed challenge's 192 bits are advice-decomposed HERE, a limb at a time, + # boolean-constrained and tied back by reconstruction; each depth-bit group also + # becomes a query position (little-endian), with a pointer to its bit run (the + # Merkle direction bits). The bits live in FRAME cells, every index into them + # being compile-time, and `addr` names the run so the direction-bit pointers + # still reach it. per_word = FIELD_BITS // depth bits = StackBuf(FIELD_BITS) - hint_decompose_bits(bits, squeezed_word, FIELD_BITS) bits_ptr = addr(bits) - reconstructed = 0 + hint_decompose_bits(bits, squeezed[0], BASE_FIELD_BITS) + hint_decompose_bits(bits_ptr * GEN ** 64, squeezed[1], BASE_FIELD_BITS) + hint_decompose_bits(bits_ptr * GEN ** 128, squeezed[2], BASE_FIELD_BITS) + acc0 = 0 + acc1 = 0 + acc2 = 0 for j in unroll(0, per_word): base_bit = j * depth - # A group inside one 64-bit limb shifts as a WHOLE, the coordinate basis - # being the polynomial basis there: COORD_BASIS[base_bit + b] == - # COORD_BASIS[base_bit] * COORD_BASIS[b] below 64, so the group contributes - # one multiply rather than one per bit. A group straddling the boundary - # splits into the two runs that do stay inside a limb; `b // cut == 0` IS - # `b < cut`, the DSL's `if` comparing for equality only. - cut = 64 - base_bit % 64 # bits of this group below the next limb + lane = base_bit // 64 + shift = base_bit % 64 + # A group inside one limb is one run of bits; a group straddling a limb + # boundary splits into the two runs that do stay inside one. `b // cut == 0` + # IS `b < cut`, the DSL's `if` comparing for equality only. + cut = 64 - shift # bits of this group below the next limb p_lo = 0 p_hi = 0 for b in unroll(0, depth): t = bits[base_bit + b] bits[base_bit + b] = t * t # booleanity, as a write-once pin if b // cut == 0: - p_lo += t * COORD_BASIS[b] + p_lo += t * 2 ** b else: - p_hi += t * COORD_BASIS[b - cut] + p_hi += t * 2 ** (b - cut) + if const(lane == 0): + acc0 += p_lo * 2 ** shift + if const(lane == 1): + acc1 += p_lo * 2 ** shift + if const(lane == 2): + acc2 += p_lo * 2 ** shift # position = p_lo + 2^cut * p_hi: multiplying by X^cut concatenates the two - # runs, since both degrees stay below 64. + # runs, both degrees staying below 64. if cut // depth == 0: # `cut < depth`: this group straddles the boundary - positions_out[GEN ** j] = p_lo + COORD_BASIS[cut] * p_hi - reconstructed += COORD_BASIS[base_bit] * p_lo + COORD_BASIS[base_bit + cut] * p_hi + positions_out[GEN ** j] = p_lo + p_hi * 2 ** cut + if const(lane == 0): + acc1 += p_hi + if const(lane == 1): + acc2 += p_hi else: positions_out[GEN ** j] = p_lo - reconstructed += COORD_BASIS[base_bit] * p_lo bit_ptrs_out[GEN ** j] = bits_ptr * GEN ** base_bit for i in unroll(per_word * depth, FIELD_BITS): t = bits[i] bits[i] = t * t - reconstructed += t * COORD_BASIS[i] - assert reconstructed == squeezed_word + if const(i // 64 == 0): + acc0 += t * 2 ** (i % 64) + if const(i // 64 == 1): + acc1 += t * 2 ** (i % 64) + if const(i // 64 == 2): + acc2 += t * 2 ** (i % 64) + assert acc0 == squeezed[0] + assert acc1 == squeezed[1] + assert acc2 == squeezed[2] return -def grind_check(state_0, state_1, nonce, nbits_g): +def grind_check(state: StackBuf(4), nonce: StackBuf(3), nbits_g): # WHIR fold/query grinding: digest = H(H(state, POW_BASE), (nonce, POW_NONCE)), # whose low nbits (nbits_g = g^nbits) must be zero. The PoW window of # transcript::pow_bits_ok is `digest.0 & ((1 << bits) - 1)` with nbits < 64, so - # it lives entirely in the digest's FIRST 64-bit lane: only that lane is - # advice-decomposed and verified, not all FIELD_BITS of the cell. The caller - # absorbs the full field nonce afterwards. The honest prover searches the - # deterministic u64 subset while verification permits the full field: each - # candidate still costs one hash and succeeds with probability 2^-bits. + # it lives entirely in the digest's first word: only that word is + # advice-decomposed and verified. The caller absorbs the full field nonce + # afterwards. The honest prover searches the deterministic u64 subset while + # verification permits the full field: each candidate still costs one hash and + # succeeds with probability 2^-bits. if nbits_g == GEN ** 0: - assert nonce == 0 # native canonical zero-work nonce - st = [state_0, state_1] - base = StackBuf(2) - blake2s(st, [0, Y_TOWER * DS_POW_BASE], base) - out = StackBuf(2) - fs_compress(base, nonce, DS_POW_NONCE, out) - lanes = StackBuf(1) - hint_f192_limbs(lanes, out[0]) - high = (out[0] + lanes[0]) * Y_INV - assert_in_k(lanes[0], high) + assert_eq192(nonce, ZERO) # native canonical zero-work nonce + base = StackBuf(4) + blake2s(state, [0, 0, 0, DS_POW_BASE], base) + out = StackBuf(4) + blake2s(base, [nonce, DS_POW_NONCE], out) # Frame cells for the unrolled pass (no DEREF per bit), named by `addr` for the # zero-check walk, whose bound is runtime and so must index a pointer. - lane_bits = StackBuf(BASE_FIELD_BITS) - hint_decompose_bits(lane_bits, lanes[0], BASE_FIELD_BITS) - acc = 0 - for i in unroll(0, BASE_FIELD_BITS): - b = lane_bits[i] - lane_bits[i] = b * b # booleanity, as a write-once pin - acc += b * COORD_BASIS[i] - assert acc == lanes[0] # the bits ARE the lane's coordinates, so the pins below bind it - lane_ptr = addr(lane_bits) + word_bits = StackBuf(BASE_FIELD_BITS) + hint_decompose_bits(word_bits, out[0], BASE_FIELD_BITS) + word_ptr = addr(word_bits) + bind_bits(word_ptr, out[0], BASE_FIELD_BITS) for xb in mul_range(1, nbits_g): - assert lane_ptr[xb] == 0 + assert word_ptr[xb] == 0 return @@ -727,40 +704,43 @@ def grind_check(state_0, state_1, nonce, nbits_g): @inline def eq_weight(ch, count: Const, idx: Const, msb_span: Const): # The eq-tensor weight of compile-time index `idx` against the challenge run - # ch[0..count): prod_c eq(bit(idx), ch[c]), where the bit is bit c of idx - # (msb_span == 0) or bit (msb_span - 1 - c) (an MSB-first walk over an - # msb_span-bit index). - w = GEN ** 0 + # ch[0..count), three words an element: prod_c eq(bit(idx), ch[c]), where the + # bit is bit c of idx (msb_span == 0) or bit (msb_span - 1 - c) (an MSB-first + # walk over an msb_span-bit index). + w = ONE for c in unroll(0, count): - cv = ch[GEN ** c] + k = 3 * c + cv = ch[k:k + 3] if msb_span == 0: - if (idx // (2 ** c)) % 2 == 1: - w *= cv - else: - w *= (1 + cv) + bit = (idx // (2 ** c)) % 2 else: - if (idx // (2 ** (msb_span - 1 - c))) % 2 == 1: - w *= cv - else: - w *= (1 + cv) + bit = (idx // (2 ** (msb_span - 1 - c))) % 2 + if bit == 1: + w = mul192(w, cv) + else: + w = mul192(w, add192(ONE, cv)) return w @inline def eqtree(point_ptr, out, n_coords: Const): # The eq tensor of the n_coords challenges at point_ptr[0..n_coords), built by - # doubling into out (size 2^(n_coords+1) - 2); the final 2^n_coords values start - # at offset 2^n_coords - 2. - r0 = point_ptr[GEN ** 0] - out[GEN ** 0] = 1 + r0 - out[GEN ** 1] = r0 + # doubling into out (size 2^(n_coords+1) - 2 elements); the final 2^n_coords + # values start at element 2^n_coords - 2. + r0 = point_ptr[0:3] + out[0:3] = add192(ONE, r0) + out[3:6] = r0 for t in unroll(1, n_coords): - rt = point_ptr[GEN ** t] - one_plus_rt = 1 + rt + k = 3 * t + rt = point_ptr[k:k + 3] + one_plus_rt = add192(ONE, rt) for i in unroll(0, 2 ** t): - pw = out[GEN ** (2 ** t - 2 + i)] - out[GEN ** (2 ** (t + 1) - 2 + i)] = pw * one_plus_rt - out[GEN ** (2 ** (t + 1) - 2 + 2 ** t + i)] = pw * rt + src = 3 * (2 ** t - 2 + i) + lo = 3 * (2 ** (t + 1) - 2 + i) + hi = 3 * (2 ** (t + 1) - 2 + 2 ** t + i) + pw = out[src:src + 3] + out[lo:lo + 3] = mul192(pw, one_plus_rt) + out[hi:hi + 3] = mul192(pw, rt) return @@ -768,30 +748,36 @@ def eqtree(point_ptr, out, n_coords: Const): def lag64(z, out, node_base: Const): # The 64 phi8-domain Lagrange NUMERATORS at z over nodes # PHI8_NODES[node_base .. node_base + 64]: out[i] = prod_{j != i} (z + - # PHI8_NODES[node_base + j]). Every barycentric denominator over an aligned phi8 - # window is the same element, so callers scale the finished sum once by - # LAGRANGE_INV_S / LAGRANGE_INV_COMBINED instead of the numerators one by one. - pre = StackBuf(65) - pre[0] = 1 + # PHI8_NODES[node_base + j]), three cells an element. Every barycentric + # denominator over an aligned phi8 window is the same element, so callers scale + # the finished sum once by LAGRANGE_INV_S / LAGRANGE_INV_COMBINED instead of the + # numerators one by one. + pre = StackBuf(3 * 65) + pre[0:3] = ONE for i in unroll(0, 64): - pre[i + 1] = pre[i] * (z + PHI8_NODES[node_base + i]) - suf = StackBuf(65) - suf[64] = 1 + k = 3 * i + pre[k + 3:k + 6] = mul192(pre[k:k + 3], add192(z, phi8(node_base + i))) + suf = StackBuf(3 * 65) + suf[192:195] = ONE for i in unroll(0, 64): - suf[63 - i] = suf[64 - i] * (z + PHI8_NODES[node_base + 63 - i]) + k = 3 * (63 - i) + suf[k:k + 3] = mul192(suf[k + 3:k + 6], add192(z, phi8(node_base + 63 - i))) for i in unroll(0, 64): - out[i] = pre[i] * suf[i + 1] + k = 3 * i + out[k:k + 3] = mul192(pre[k:k + 3], suf[k + 3:k + 6]) return -def eq_prefix_chain(chain, seed, a, b, count_g): +def eq_prefix_chain(chain, seed: StackBuf(3), a, b, count_g): # Prefix products of eq(a_k, b_k) = 1 + a_k + b_k from `seed`, so a reader picks # the partial product up at its own certified length. Entry t is written from # inputs with index < t only, so a garbage tail past a buffer's written extent - # cannot corrupt any shorter prefix. - chain[GEN ** 0] = seed + # cannot corrupt any shorter prefix. Every buffer holds three words an element. + chain[0:3] = seed for xk in mul_range(1, count_g): - chain[xk * GEN] = chain[xk] * (1 + a[xk] + b[xk]) + x3 = xk ** 3 + nxt = x3 * GEN ** 3 + chain[nxt:nxt + 3] = mul192(chain[x3:x3 + 3], add192(ONE, add192(a[x3:x3 + 3], b[x3:x3 + 3]))) return @@ -800,126 +786,140 @@ def rs_eq_run(chain, z_vals, point, count_g): # (z_j^(2^k) + 1 + ris_j): coordinate x multiplies row k by (z^(2^k) + 1 + # point_x), z evolving by squaring per row. The runtime coordinates walk # OUTSIDE and the fixed Frobenius powers inside, so nothing stores a z-power - # table. + # table. A row is BASE_FIELD_BITS elements. for xk in mul_range(1, count_g): - zv = z_vals[xk] - one_plus = 1 + point[xk] - row = chain * xk ** BASE_FIELD_BITS - nxt = row * GEN ** BASE_FIELD_BITS + x3 = xk ** 3 + zv = z_vals[x3:x3 + 3] + one_plus = add192(ONE, point[x3:x3 + 3]) + row = chain * xk ** (3 * BASE_FIELD_BITS) + nxt = row * GEN ** (3 * BASE_FIELD_BITS) for k in unroll(0, BASE_FIELD_BITS): - nxt[GEN ** k] = row[GEN ** k] * (zv + one_plus) + c = 3 * k + nxt[c:c + 3] = mul192(row[c:c + 3], add192(zv, one_plus)) if k != BASE_FIELD_BITS - 1: - zv *= zv + zv = mul192(zv, zv) return def fold_final_msg(msg, point, log_len: Const): - weights = StackBuf(2 * YR_LOG_CAP) - for j in unroll(0, log_len): - weights[2 * j] = 1 + point[GEN ** j] - weights[2 * j + 1] = point[GEN ** j] # Weighted fold of the final_msg multilinear over 2^log_len values (log_len is # the candidate's yr_log_n; the frame buffers use the global max size). - l0 = StackBuf(2 ** YR_LOG_CAP) + l0 = StackBuf(3 * 2 ** YR_LOG_CAP) + p0 = point[0:3] for t in unroll(0, 2 ** log_len // 2): - l0[t] = weights[0] * msg[GEN ** (2 * t)] + weights[1] * msg[GEN ** (2 * t + 1)] + k = 3 * t + lo = 6 * t + hi = 6 * t + 3 + a = msg[lo:lo + 3] + l0[k:k + 3] = add192(a, mul192(p0, add192(a, msg[hi:hi + 3]))) cursor = l0 n = 2 ** log_len // 2 for j in unroll(1, log_len): - nxt = StackBuf(2 ** YR_LOG_CAP) + pj_off = 3 * j + pj = point[pj_off:pj_off + 3] + nxt = StackBuf(3 * 2 ** YR_LOG_CAP) for t in unroll(0, n // 2): - nxt[t] = weights[2 * j] * cursor[2 * t] + weights[2 * j + 1] * cursor[2 * t + 1] + k = 3 * t + lo = 6 * t + hi = 6 * t + 3 + a = cursor[lo:lo + 3] + nxt[k:k + 3] = add192(a, mul192(pj, add192(a, cursor[hi:hi + 3]))) cursor = nxt n = n // 2 - return cursor[0] - - -def sumcheck_round4(state_0, state_1, msg_cursor, claim): - # One PLAIN sumcheck round. The prover sends the round polynomial's coefficients - # bar the one the split h(0) + h(1) == claim fixes, so the verifier derives that - # one and reads h at the challenge by Horner. Nothing is reapplied: no eq - # factor, no separate term for the tables still waiting, and the eq point is not - # read here at all. - fs = [state_0, state_1] - fs, c0, msg_cursor = fs_next(fs, msg_cursor) - fs, c2, msg_cursor = fs_next(fs, msg_cursor) - fs, c3, msg_cursor = fs_next(fs, msg_cursor) - c1 = claim + c2 + c3 # the split identity fixes it, so it is neither sent nor bound + return cursor[0:3] + + +def sumcheck_round4(rd): + # One PLAIN sumcheck round off the round record at `rd`, whose successor it + # writes. The prover sends the round polynomial's coefficients bar the one the + # split h(0) + h(1) == claim fixes, so the verifier derives that one and reads h + # at the challenge by Horner. Nothing is reapplied: no eq factor, no separate + # term for the tables still waiting, and the eq point is not read here at all. + fs, c0, cursor = fs_next(rd[0:4], rd[GEN ** ROUND_CURSOR]) + fs, c2, cursor = fs_next(fs, cursor) + fs, c3, cursor = fs_next(fs, cursor) + c1 = add192(rd[ROUND_CLAIM:ROUND_CLAIM + 3], add192(c2, c3)) # the split fixes it, so it is neither sent nor bound fs, y = squeeze(fs) - return fs[0], fs[1], msg_cursor, c0 + y * (c1 + y * (c2 + y * c3)), y - - -def sumcheck_round5(state_0, state_1, msg_cursor, claim, prev_challenge): - # One GKR round. The prover sends every coefficient but c0, which the round's - # pulled-out eq factor leaves fixed: `c0 + prev_challenge * (c1 + ... + c4) == - # claim`. - fs = [state_0, state_1] - fs, c1, msg_cursor = fs_next(fs, msg_cursor) - fs, c2, msg_cursor = fs_next(fs, msg_cursor) - fs, c3, msg_cursor = fs_next(fs, msg_cursor) - fs, c4, msg_cursor = fs_next(fs, msg_cursor) + nxt = rd * GEN ** ROUND_SLOTS + nxt[0:4] = fs + nxt[GEN ** ROUND_CURSOR] = cursor + nxt[ROUND_CLAIM:ROUND_CLAIM + 3] = add192(c0, mul192(y, add192(c1, mul192(y, add192(c2, mul192(y, c3)))))) + return y + + +def sumcheck_round5(rd, prev, challenge_out): + # One GKR round off the round record at `rd`, whose successor it writes, its + # challenge landing at `challenge_out`. The prover sends every coefficient but + # c0, which the round's pulled-out eq factor (the challenge at `prev`) leaves + # fixed: `c0 + prev_challenge * (c1 + ... + c4) == claim`. + prev_challenge = prev[0:3] + fs, c1, cursor = fs_next(rd[0:4], rd[GEN ** ROUND_CURSOR]) + fs, c2, cursor = fs_next(fs, cursor) + fs, c3, cursor = fs_next(fs, cursor) + fs, c4, cursor = fs_next(fs, cursor) fs, y = squeeze(fs) - c0 = claim + prev_challenge * (c1 + c2 + c3 + c4) - return fs[0], fs[1], msg_cursor, c0 + y * (c1 + y * (c2 + y * (c3 + y * c4))), y + challenge_out[0:3] = y + c0 = add192(rd[ROUND_CLAIM:ROUND_CLAIM + 3], mul192(prev_challenge, add192(add192(c1, c2), add192(c3, c4)))) + nxt = rd * GEN ** ROUND_SLOTS + nxt[0:4] = fs + nxt[GEN ** ROUND_CURSOR] = cursor + nxt[ROUND_CLAIM:ROUND_CLAIM + 3] = add192(c0, mul192(y, add192(c1, mul192(y, add192(c2, mul192(y, add192(c3, mul192(y, c4)))))))) + return -def batch_sumcheck(fs0, fs1, msgs, running, point, n_rounds: Const): +def batch_sumcheck(fs: StackBuf(4), msgs, running: StackBuf(3), point, n_rounds: Const): # The rounds of a claim-batching sumcheck: two hinted values per round # (g(1) and g(inf)), the split fixing the third against the running claim, and # the challenges collected into `point`. - fs = [fs0, fs1] for rd in unroll(0, n_rounds): - fs, msg_g1, c = fs_next(fs, msgs * GEN ** (2 * rd)) + fs, msg_g1, c = fs_next(fs, msgs * GEN ** (6 * rd)) fs, msg_ginf, c = fs_next(fs, c) fs, rv = squeeze(fs) - point[GEN ** rd] = rv - g_zero = running + msg_g1 - c_one = g_zero + msg_g1 + msg_ginf - running = (msg_ginf * rv + c_one) * rv + g_zero # fold the degree-2 round at rv - return fs[0], fs[1], running + k = 3 * rd + point[k:k + 3] = rv + g_zero = add192(running, msg_g1) + c_one = add192(g_zero, add192(msg_g1, msg_ginf)) + running = add192(mul192(add192(mul192(msg_ginf, rv), c_one), rv), g_zero) # fold the degree-2 round at rv + return fs, running # ==================================== Merkle paths ================================== @inline -def order_children(node, sibling, bit): - # Branchless child ordering: `bit` is boolean-pinned wherever it comes from, so - # `m = bit*(node + sibling)` selects rather than branches, leaving (node, - # sibling) at bit 0 and (sibling, node) at bit 1. - m = bit * (node + sibling) - kids = [node + m, sibling + m] - return kids +def order_children(n0, n1, sibling, bit): + # Branchless child ordering over two-word values: `bit` is boolean-pinned + # wherever it comes from, so `m = bit*(node + sibling)` selects rather than + # branches, leaving (node, sibling) at bit 0 and (sibling, node) at bit 1. + m0 = bit * (n0 + sibling[0]) + m1 = bit * (n1 + sibling[1]) + return [n0 + m0, n1 + m1, sibling[0] + m0, sibling[1] + m1] @inline -def verify_merkle_path(leaf_0, leaf_1, direction_bits, depth: Const): +def verify_merkle_path(leaf, direction_bits, depth: Const): # Hinted child pairs are hashed in order; the boolean query bit selects the # child that must equal the running node, binding each link of the path. path = StackBuf(LIG_PATH_CAP) - hint_witness(path[0:4 * depth], "merkle_children") + hint_witness(path[0:8 * depth], "merkle_children") path_ptr = addr(path) - node_0 = leaf_0 - node_1 = leaf_1 + node = leaf for level in unroll(0, depth): dir_bit = direction_bits[GEN ** level] - selected = path_ptr * GEN ** (4 * level) * (1 + dir_bit * (1 + GEN ** 2)) - selected[1] = node_0 - selected[GEN] = node_1 - parent = StackBuf(2) - blake2s(path[4 * level:4 * level + 2], path[4 * level + 2:4 * level + 4], parent) - node_0 = parent[0] - node_1 = parent[1] - return node_0, node_1 + selected = path_ptr * GEN ** (8 * level) * (1 + dir_bit * (1 + GEN ** 4)) + selected[0:4] = node + parent = StackBuf(4) + blake2s(path[8 * level:8 * level + 4], path[8 * level + 4:8 * level + 8], parent) + node = parent + return node @inline def hash_cap_node(cap, index): - children = cap * index ** 4 - parent = StackBuf(2) - blake2s([children[1], children[GEN]], [children[GEN ** 2], children[GEN ** 3]], parent) - cap[index * index] = parent[0] - cap[GEN * index * index] = parent[1] + # Node n of a cap occupies words 4n..4n+4, so its children start at 8n. + children = cap * index ** 8 + parent = cap * index ** 4 + blake2s(children[0:4], children[4:8], parent[0:4]) return @@ -933,101 +933,99 @@ def verify_merkle_cap(cap, flags, depth: Const): if flags[child] != 0: flags[parent] = 1 # every active node forces its parent to be hashed hash_cap_node(cap, child) - return cap[GEN ** 2], cap[GEN ** 3] + return cap[4:8] # ============================== the stacked WHIR opening ============================ -def opening_fold(fs0, fs1, cursor, c0, c1, c2): - fs = [fs0, fs1] +def opening_fold(fs: StackBuf(4), cursor, c0: StackBuf(3), c1: StackBuf(3), c2: StackBuf(3)): fs, r = squeeze(fs) - claim = (c2 * r + c1) * r + c0 - fs, c0, cursor = fs_next(fs, cursor) - fs, c2, cursor = fs_next(fs, cursor) - return fs[0], fs[1], cursor, claim, c0, c2, r + claim = add192(mul192(add192(mul192(c2, r), c1), r), c0) + fs, n0, cursor = fs_next(fs, cursor) + fs, n2, cursor = fs_next(fs, cursor) + return fs, cursor, claim, n0, n2, r -def opening_ood_point(fs0, fs1, point, n_g): +def opening_ood_point(fs: StackBuf(4), point, n_g): # Share the squeeze loop across point dimensions. - states = HeapBuf((n_g * GEN) ** 2) - states[1] = fs0 - states[GEN] = fs1 + states = HeapBuf((n_g * GEN) ** FS_SLOTS) + states[0:4] = fs for x in mul_range(1, n_g): - state = states * x * x - r, f0, f1 = squeeze_step(state[1], state[GEN]) - point[x] = r - state[GEN ** 2] = f0 - state[GEN ** 3] = f1 - last = states * n_g * n_g - return last[1], last[GEN] + state = states * x ** FS_SLOTS + nb, r = squeeze(state[0:4]) + x3 = x ** 3 + point[x3:x3 + 3] = r + state[4:8] = nb + last = states * n_g ** FS_SLOTS + return last[0:4] -def opening_final_message(fs0, fs1, cursor, out, n: Const): - fs = [fs0, fs1] +def opening_final_message(fs: StackBuf(4), cursor, out, n: Const): for i in unroll(0, n): fs, value, cursor = fs_next(fs, cursor) - out[GEN ** i] = value - return fs[0], fs[1], cursor + k = 3 * i + out[k:k + 3] = value + return fs, cursor -def opening_row_weights(point, out, folds: Const, reverse: Const): - for i in unroll(0, 2 ** folds): - if reverse == 1: - slot = 2 ** folds - 1 - i - else: - slot = i - out[GEN ** i] = eq_weight(point, folds, slot, 0) - return - - -def opening_queries(cap, flags, query_weights, query_bit_ptrs, row_eq_weights, n_queries_g, base: Const, interleave: Const, blocks: Const, depth: Const, cap_depth: Const): +def opening_queries(cap, flags, query_weights, query_bit_ptrs, fold_point, n_queries_g, base: Const, folds: Const, blocks: Const, depth: Const, cap_depth: Const): # Specialize by row and path shape so opening configurations share query code. - query_sum_chain = HeapBuf(n_queries_g * GEN) - query_sum_chain[GEN ** 0] = 0 + # A row enters its level's claim through its multilinear extension at the + # level's fold point, folded one coordinate at a time. At level 0, slot i of a + # leaf image is interleaving index n-1-i (see open_stacked), which complements + # every index bit, so a fold there keeps the high half's term. + query_sum_chain = HeapBuf((n_queries_g * GEN) ** 3) + query_sum_chain[0:3] = ZERO for xe in mul_range(1, n_queries_g): - if base == 1: - row_len = interleave - else: - row_len = 3 * interleave + interleave = 2 ** folds # bound in the body, which captures names by value row = StackBuf(LIG_ROW_CAP) - hint_witness(row[0:row_len], "merkle_leaf_rows") - row_dot = 0 - packed_row = StackBuf(LIG_PACKED_ROW_CAP) + fold = StackBuf(3 * LIG_ROW_CAP) + r0 = fold_point[0:3] if base == 1: - # Packing proves each hinted lane is in K before hashing or folding it. - for jb in unroll(0, interleave // 4): - e0 = row[4 * jb] - e1 = row[4 * jb + 1] - e2 = row[4 * jb + 2] - e3 = row[4 * jb + 3] - packed_row[2 * jb] = pack64x2(e0, e1) - packed_row[2 * jb + 1] = pack64x2(e2, e3) - row_dot += e0 * row_eq_weights[GEN ** (4 * jb)] + e1 * row_eq_weights[GEN ** (4 * jb + 1)] + e2 * row_eq_weights[GEN ** (4 * jb + 2)] + e3 * row_eq_weights[GEN ** (4 * jb + 3)] + # Level 0's lanes are words, so its first fold runs limb by limb. + hint_witness(row[0:interleave], "merkle_leaf_rows") + for t in unroll(0, interleave // 2): + lo = row[2 * t] + hi = row[2 * t + 1] + d = lo + hi + k = 3 * t + fold[k:k + 3] = [hi + d * r0[0], d * r0[1], d * r0[2]] else: - # Pack the checked tower limbs into the leaf's contiguous byte image. - for jb in unroll(0, 3 * interleave // 4): - packed_row[2 * jb] = pack64x2(row[4 * jb], row[4 * jb + 1]) - packed_row[2 * jb + 1] = pack64x2(row[4 * jb + 2], row[4 * jb + 3]) - for jw in unroll(0, interleave): - if 3 * jw % 2 == 0: - # limbs (3w, 3w+1) are a pack; add Y^2 * limb(3w+2). - row_word = packed_row[3 * jw // 2] + Y_TOWER * Y_TOWER * row[3 * jw + 2] + # A deeper row is `interleave` elements, three limbs each. + hint_witness(row[0:3 * interleave], "merkle_leaf_rows") + for t in unroll(0, interleave // 2): + k = 3 * t + lo = 6 * t + a = row[lo:lo + 3] + fold[k:k + 3] = add192(a, mul192(r0, add192(a, row[lo + 3:lo + 6]))) + # Fold c reads fold c-1's outputs and writes the next run of the buffer. + for c in unroll(1, folds): + rc = fold_point[3 * c:3 * c + 3] + src = 3 * (interleave - interleave // 2 ** (c - 1)) + dst = 3 * (interleave - interleave // 2 ** c) + for t in unroll(0, interleave // 2 ** (c + 1)): + a = fold[src + 6 * t:src + 6 * t + 3] + b = fold[src + 6 * t + 3:src + 6 * t + 6] + if base == 1: + fold[dst + 3 * t:dst + 3 * t + 3] = add192(b, mul192(rc, add192(a, b))) else: - # limbs (3w+1, 3w+2) are a pack; shift it by Y and add limb(3w). - row_word = row[3 * jw] + Y_TOWER * packed_row[(3 * jw + 1) // 2] - row_dot += row_word * row_eq_weights[GEN ** jw] - # Hash the packed row as full BLAKE2s blocks. - leaf_hash_state = StackBuf(2) - blake2s(packed_row[0:2], packed_row[2:4], leaf_hash_state, counter=64, final=1 // blocks) + fold[dst + 3 * t:dst + 3 * t + 3] = add192(a, mul192(rc, add192(a, b))) + dot_at = 3 * (interleave - 2) + row_dot = fold[dot_at:dot_at + 3] + # The row's words are its leaf's byte image, hashed as full BLAKE2s blocks. + leaf_state = StackBuf(4) + blake2s(row[0:4], row[4:8], leaf_state, counter=64, final=1 // blocks) for jb in unroll(1, blocks): - leaf_digest = StackBuf(2) - blake2s(packed_row[4 * jb:4 * jb + 2], packed_row[4 * jb + 2:4 * jb + 4], leaf_digest, cv=leaf_hash_state, counter=64 * (jb + 1), final=(jb + 1) // blocks) - leaf_hash_state = leaf_digest - query_sum_chain[xe * GEN] = query_sum_chain[xe] + query_weights[xe] * row_dot + leaf_digest = StackBuf(4) + blake2s(row[8 * jb:8 * jb + 4], row[8 * jb + 4:8 * jb + 8], leaf_digest, cv=leaf_state, counter=64 * (jb + 1), final=(jb + 1) // blocks) + leaf_state = leaf_digest + xe3 = xe ** 3 + nxt = xe3 * GEN ** 3 + query_sum_chain[nxt:nxt + 3] = add192(query_sum_chain[xe3:xe3 + 3], mul192(query_weights[xe3:xe3 + 3], row_dot)) direction_bits = query_bit_ptrs[xe] path_depth = depth - cap_depth - node_0, node_1 = verify_merkle_path(leaf_hash_state[0], leaf_hash_state[1], direction_bits, path_depth) + node = verify_merkle_path(leaf_state, direction_bits, path_depth) if cap_depth != 0: parent = GEN ** (2 ** (cap_depth - 1)) for bit in unroll(0, cap_depth - 1): @@ -1036,18 +1034,36 @@ def opening_queries(cap, flags, query_weights, query_bit_ptrs, row_eq_weights, n cap_index = parent * parent * (1 + direction_bits[GEN ** path_depth] * (1 + GEN)) else: cap_index = GEN - cap[cap_index * cap_index] = node_0 - cap[GEN * cap_index * cap_index] = node_1 - return query_sum_chain[n_queries_g] + slot = cap * cap_index ** 4 + slot[0:4] = node + end = n_queries_g ** 3 + return query_sum_chain[end:end + 3] -def open_stacked(m_idx: Const, fs0, fs1, target, commit_root_0, commit_root_1, cursor): +@inline +def fold_at(point, f: Const, lane_folds: Const, tail_start: Const): + # Where fold challenge f (in round order) sits in the witness-order point: the + # lane folds after the tail, every later fold `lane_folds` places down. + if const(f // lane_folds == 0): + pos = tail_start + f + else: + pos = f - lane_folds + return point * GEN ** (3 * pos) + + +def open_stacked(m_idx: Const, fs: StackBuf(4), target: StackBuf(3), commit_root: StackBuf(4), cursor): # The stacked WHIR opening, one specialization per (rate, committed log-size) # candidate: every LIG_* table reads row m_idx, per level row `ml`, and all # opening proof data is hinted here, so only the executed arm pops its streams. # - # The returned point is in witness order. Round-order challenges stay local - # for the induced weights, OOD claims, and residual evaluation. + # The returned point is in witness order. The folds bind coordinates in ROUND + # order, and level 0's folds are the lane fold, binding the witness's TOP k + # coordinates: lane l of the commitment is the stack block q[l * 2^(mu-k) ...], + # which is what makes the witness's zero padding whole lanes for the committer + # to leave out of the encode. Every transparent weight downstream is written in + # witness coordinates, so each challenge is written straight to its witness + # coordinate as it arrives (`fold_at`), and the round-order readers below index + # it there. n_levels = LIG_N_LEVELS[m_idx] yr_level = LIG_YR_LEVEL[m_idx] yr_log = LIG_YR_LOG_LEN[m_idx] @@ -1055,105 +1071,107 @@ def open_stacked(m_idx: Const, fs0, fs1, target, commit_root_0, commit_root_1, c n_folds = LIG_TOTAL_FOLDS[m_idx] max_q = LIG_MAX_QUERIES[m_idx] ood_stride = LIG_MAX_OOD_SAMPLES * OOD_SLOTS + lane_folds = LIG_FOLDS[m_idx * LIG_MAX_LEVELS] + fold_head = n_folds - lane_folds + tail_start = fold_head + yr_log + # Padding is zero so shared prefix chains may compute unused entries above m + # without reading uninitialized cells. + point = HeapBuf(3 * (SIZE_BITS + SLOT_STRIDE_LOG)) + for j in unroll(n_folds + yr_log, SIZE_BITS + SLOT_STRIDE_LOG): + k = 3 * j + point[k:k + 3] = ZERO + tail = point * GEN ** (3 * fold_head) # The opening's scalars (sumcheck messages, level roots, nonces, final message) # ride the SHARED stream, walked on in protocol order. The K opener binds a - # Merkle root as its two F192 scalars, not as a byte string, and every digest - # uses ONE 128/128 encoding, so those scalars are exactly the root's two cells. - fs = [fs0, fs1] + # Merkle root as two scalars, one per 128-bit half. msg_cursor = cursor fs, round_quad_c, msg_cursor = fs_next(fs, msg_cursor) # the round polynomial in coefficients fs, round_quad_a, msg_cursor = fs_next(fs, msg_cursor) # bar the linear one, which the split fixes - round_quad_b = target + round_quad_a + round_quad_b = add192(target, round_quad_a) sumcheck_target = target - # Caps are shared across queries; rows and paths are hinted in each query's frame. - merkle_caps = HeapBuf(GEN ** (4 * LIG_CAP_LEN[m_idx])) - hint_witness(merkle_caps[0:4 * LIG_CAP_LEN[m_idx]], "merkle_caps") + # Caps are shared across queries, a node four words; rows and paths are hinted in + # each query's frame. + merkle_caps = HeapBuf(GEN ** (8 * LIG_CAP_LEN[m_idx])) + hint_witness(merkle_caps[0:8 * LIG_CAP_LEN[m_idx]], "merkle_caps") cap_flags = HeapBuf(GEN ** LIG_CAP_LEN[m_idx]) hint_witness(cap_flags[0:LIG_CAP_LEN[m_idx]], "merkle_cap_active") - final_msg = HeapBuf(GEN ** yr_len) # filled from the stream at the last level + final_msg = HeapBuf(GEN ** (3 * yr_len)) # filled from the stream at the last level # Each cap root is checked against its transcript-bound level root. - level_roots = HeapBuf(GEN ** (2 * n_levels)) - level_roots[GEN ** 0] = commit_root_0 - level_roots[GEN ** 1] = commit_root_1 - # ...and guest-filled accumulators (one slot per fold / per level / per query): - fold_challenges = HeapBuf(GEN ** n_folds) - level_betas = HeapBuf(GEN ** n_levels) - query_weights = HeapBuf(GEN ** (n_levels * max_q)) + level_roots = HeapBuf(GEN ** (4 * n_levels)) + level_roots[0:4] = commit_root + # ...and guest-filled accumulators (one slot per level / per query): + level_betas = HeapBuf(GEN ** (3 * n_levels)) + query_weights = HeapBuf(GEN ** (3 * n_levels * max_q)) query_positions = HeapBuf(GEN ** (LIG_POSITIONS_LEN[m_idx])) query_bit_ptrs = HeapBuf(GEN ** (LIG_POSITIONS_LEN[m_idx])) # Explicit OOD claims bind every recursive Johnson-list commitment. L0 needs # none: the opening claim itself is its post-commit binding value. An OOD claim # is read before its level's query positions but batched after them (one # challenge per level), so its value and intro message wait here. - ood_z = HeapBuf(GEN ** (n_levels * LIG_MAX_OOD_SAMPLES * LIG_LOG_MSG_COLS_CAP)) + ood_z = HeapBuf(GEN ** (3 * n_levels * LIG_MAX_OOD_SAMPLES * LIG_LOG_MSG_COLS_CAP)) ood = HeapBuf(GEN ** (n_levels * ood_stride)) for lvl in unroll(0, n_levels): ml = m_idx * LIG_MAX_LEVELS + lvl n_queries = LIG_QUERIES[ml] depth = LIG_TREE_DEPTH[ml] - interleave = LIG_INTERLEAVE[ml] folds_off = LIG_FOLDS_OFF[ml] pos_off = LIG_POSITIONS_OFF[ml] for j in unroll(0, LIG_FOLDS[ml]): - f0, f1, msg_cursor, sumcheck_target, round_quad_c, round_quad_a, fold_challenge = opening_fold(fs[0], fs[1], msg_cursor, round_quad_c, round_quad_b, round_quad_a) - fs = [f0, f1] - fold_challenges[GEN ** (folds_off + j)] = fold_challenge - round_quad_b = sumcheck_target + round_quad_a + fs, msg_cursor, sumcheck_target, round_quad_c, round_quad_a, fold_challenge = opening_fold(fs, msg_cursor, round_quad_c, round_quad_b, round_quad_a) + fc = fold_at(point, folds_off + j, lane_folds, tail_start) + fc[0:3] = fold_challenge + round_quad_b = add192(sumcheck_target, round_quad_a) if lvl == yr_level: - f0, f1, msg_cursor = opening_final_message(fs[0], fs[1], msg_cursor, final_msg, yr_len) - fs = [f0, f1] + fs, msg_cursor = opening_final_message(fs, msg_cursor, final_msg, yr_len) else: - fs, next_root_a, msg_cursor = fs_next(fs, msg_cursor) - fs, next_root_b, msg_cursor = fs_next(fs, msg_cursor) # A non-canonical half is rejected (merkle.rs `scalars_to_hash`), as # the commitment root was at its own read. - canon = StackBuf(2) - canon[0] = assert_canonical(next_root_a) - canon[1] = assert_canonical(next_root_b) - level_roots[GEN ** (2 * lvl + 2)] = next_root_a - level_roots[GEN ** (2 * lvl + 3)] = next_root_b + fs, root_a, msg_cursor = fs_next_half(fs, msg_cursor) + fs, root_b, msg_cursor = fs_next_half(fs, msg_cursor) + k = 4 * lvl + 4 + level_roots[k:k + 2] = root_a[0:2] + level_roots[k + 2:k + 4] = root_b[0:2] # OOD binding for the newly observed level-(lvl+1) commitment. The # random point has the just-folded witness dimension, namely this # level's message-column dimension. for os in unroll(0, LIG_OOD_SAMPLES[ml + 1]): - oz = ood_z * GEN ** (((lvl + 1) * LIG_MAX_OOD_SAMPLES + os) * LIG_LOG_MSG_COLS_CAP) - f0, f1 = opening_ood_point(fs[0], fs[1], oz, GEN ** LIG_LOG_MSG_COLS[ml]) - fs = [f0, f1] + oz = ood_z * GEN ** (3 * ((lvl + 1) * LIG_MAX_OOD_SAMPLES + os) * LIG_LOG_MSG_COLS_CAP) + fs = opening_ood_point(fs, oz, GEN ** LIG_LOG_MSG_COLS[ml]) sample = ood * GEN ** ((lvl + 1) * ood_stride + os * OOD_SLOTS) fs, ood_y, msg_cursor = fs_next(fs, msg_cursor) fs, ood_c0, msg_cursor = fs_next(fs, msg_cursor) fs, ood_c2, msg_cursor = fs_next(fs, msg_cursor) - sample[GEN ** OOD_Y] = ood_y - sample[GEN ** OOD_C0] = ood_c0 - sample[GEN ** OOD_C2] = ood_c2 # the split fixes c1 = y + c2 - q_nonce = msg_cursor[GEN ** 0] # raw transport word: bound by the DS_POW_NONCE absorb below - msg_cursor = msg_cursor * GEN + sample[OOD_Y:OOD_Y + 3] = ood_y + sample[OOD_C0:OOD_C0 + 3] = ood_c0 + sample[OOD_C2:OOD_C2 + 3] = ood_c2 # the split fixes c1 = y + c2 + q_nonce = msg_cursor[0:3] # raw transport scalar: bound by the DS_POW_NONCE absorb below + msg_cursor = msg_cursor * GEN ** 3 if LIG_QUERY_GRIND_BITS[ml] != 0: - grind_check(fs[0], fs[1], q_nonce, GEN ** LIG_QUERY_GRIND_BITS[ml]) + grind_check(fs, q_nonce, GEN ** LIG_QUERY_GRIND_BITS[ml]) else: - assert q_nonce == 0 - fs = absorb_nonce(fs, q_nonce) + assert_eq192(q_nonce, ZERO) + nonce_state = StackBuf(4) + blake2s(fs, [q_nonce, DS_POW_NONCE], nonce_state) + fs = nonce_state - sqz = HeapBuf((GEN ** (LIG_MAX_SQUEEZES[m_idx] + 1)) ** PAIR_SLOTS) - sqz[GEN ** 0] = fs[0] - sqz[GEN ** 1] = fs[1] + sqz = HeapBuf((GEN ** (LIG_MAX_SQUEEZES[m_idx] + 1)) ** FS_SLOTS) + sqz[0:4] = fs for xs in mul_range(1, GEN ** LIG_SQUEEZES[ml]): # a loop body captures free names BY VALUE, so the compile-time aliases # are rebound here (m_idx and lvl are substituted literals) depth = LIG_TREE_DEPTH[m_idx * LIG_MAX_LEVELS + lvl] pos_off = LIG_POSITIONS_OFF[m_idx * LIG_MAX_LEVELS + lvl] - row = sqz * xs ** PAIR_SLOTS - packed_word, next_c0, next_c1 = squeeze_step(row[GEN ** 0], row[GEN ** 1]) - row[GEN ** PAIR_SLOTS] = next_c0 - row[GEN ** (PAIR_SLOTS + 1)] = next_c1 + row = sqz * xs ** FS_SLOTS + nb, packed = squeeze(row[0:4]) + row[4:8] = nb query_ptr = xs ** (FIELD_BITS // depth) - decode_query_bits(packed_word, query_positions * GEN ** pos_off * query_ptr, query_bit_ptrs * GEN ** pos_off * query_ptr, depth) - sqz_end = sqz * (GEN ** LIG_SQUEEZES[ml]) ** PAIR_SLOTS - fs = [sqz_end[GEN ** 0], sqz_end[GEN ** 1]] + decode_query_bits(packed, query_positions * GEN ** pos_off * query_ptr, query_bit_ptrs * GEN ** pos_off * query_ptr, depth) + sqz_end = sqz * (GEN ** LIG_SQUEEZES[ml]) ** FS_SLOTS + fs = sqz_end[0:4] # One batching challenge for the level, drawn once every claim it batches is # fixed: its OOD claims above and these query positions. Claim tau of the @@ -1161,160 +1179,146 @@ def open_stacked(m_idx: Const, fs0, fs1, target, commit_root_0, commit_root_1, c # Protocol 1 step 1): query i is claim n_ood + 1 + i, so its weight splits # into lam^i here and the level scalar lam^(n_ood+1) below. fs, lam = squeeze(fs) - lam_pow = 1 + lam_pow = ONE for i in unroll(0, n_queries): - query_weights[GEN ** (lvl * max_q + i)] = lam_pow - lam_pow = lam_pow * lam - # At level 0, slot i of a leaf image is interleaving index n-1-i: the image - # reads its lanes from the top down, so the lanes a padding-free commitment - # leaves out are its LEADING words, whose whole blocks the committer hashes - # once for all leaves. The flip is a compile-time index and the guest still - # hashes the full image. Deeper levels commit every lane, ascending. - row_eq_weights = HeapBuf(GEN ** (LIG_MAX_INTERLEAVE[m_idx])) - opening_row_weights(fold_challenges * GEN ** folds_off, row_eq_weights, LIG_FOLDS[ml], 1 // (lvl + 1)) + k = 3 * (lvl * max_q + i) + query_weights[k:k + 3] = lam_pow + lam_pow = mul192(lam_pow, lam) cap_depth = LIG_CAP_DEPTH[ml] - cap = merkle_caps * GEN ** (4 * LIG_CAP_OFF[ml]) + cap = merkle_caps * GEN ** (8 * LIG_CAP_OFF[ml]) flags = cap_flags * GEN ** LIG_CAP_OFF[ml] - root_0, root_1 = verify_merkle_cap(cap, flags, cap_depth) - level_roots[GEN ** (2 * lvl)] = root_0 - level_roots[GEN ** (2 * lvl + 1)] = root_1 + level_root = verify_merkle_cap(cap, flags, cap_depth) + k = 4 * lvl + level_roots[k:k + 4] = level_root - level_query_sum = opening_queries(cap, flags, query_weights * GEN ** (lvl * max_q), query_bit_ptrs * GEN ** pos_off, row_eq_weights, GEN ** n_queries, 1 // (lvl + 1), interleave, LIG_LEAF_BLOCKS[ml], depth, cap_depth) + level_query_sum = opening_queries(cap, flags, query_weights * GEN ** (3 * lvl * max_q), query_bit_ptrs * GEN ** pos_off, fold_at(point, folds_off, lane_folds, tail_start), GEN ** n_queries, 1 // (lvl + 1), LIG_FOLDS[ml], LIG_LEAF_BLOCKS[ml], depth, cap_depth) # Every level, including the last, ties its commitment in through an intro # message. The level's claims then enter the running one with powers of # `lam`: the OOD claims held above first, then this query batch. fs, intro_c0, msg_cursor = fs_next(fs, msg_cursor) fs, intro_c2, msg_cursor = fs_next(fs, msg_cursor) - intro_c1 = level_query_sum + intro_c2 # the split fixes the linear coefficient + intro_c1 = add192(level_query_sum, intro_c2) # the split fixes the linear coefficient if lvl == yr_level: beta_lvl = lam # no OOD claim at the last level: no new oracle else: ood_scalar = lam for os in unroll(0, LIG_OOD_SAMPLES[ml + 1]): sample = ood * GEN ** ((lvl + 1) * ood_stride + os * OOD_SLOTS) - ood_y = sample[GEN ** OOD_Y] - ood_c2 = sample[GEN ** OOD_C2] - sample[GEN ** OOD_BETA] = ood_scalar - round_quad_c += ood_scalar * sample[GEN ** OOD_C0] - round_quad_b += ood_scalar * (ood_y + ood_c2) - round_quad_a += ood_scalar * ood_c2 - sumcheck_target += ood_scalar * ood_y - ood_scalar = ood_scalar * lam + ood_y = sample[OOD_Y:OOD_Y + 3] + ood_c2 = sample[OOD_C2:OOD_C2 + 3] + sample[OOD_BETA:OOD_BETA + 3] = ood_scalar + round_quad_c = add192(round_quad_c, mul192(ood_scalar, sample[OOD_C0:OOD_C0 + 3])) + round_quad_b = add192(round_quad_b, mul192(ood_scalar, add192(ood_y, ood_c2))) + round_quad_a = add192(round_quad_a, mul192(ood_scalar, ood_c2)) + sumcheck_target = add192(sumcheck_target, mul192(ood_scalar, ood_y)) + ood_scalar = mul192(ood_scalar, lam) beta_lvl = ood_scalar - level_betas[GEN ** lvl] = beta_lvl - round_quad_c += beta_lvl * intro_c0 - round_quad_b += beta_lvl * intro_c1 - round_quad_a += beta_lvl * intro_c2 - sumcheck_target += beta_lvl * level_query_sum + k = 3 * lvl + level_betas[k:k + 3] = beta_lvl + round_quad_c = add192(round_quad_c, mul192(beta_lvl, intro_c0)) + round_quad_b = add192(round_quad_b, mul192(beta_lvl, intro_c1)) + round_quad_a = add192(round_quad_a, mul192(beta_lvl, intro_c2)) + sumcheck_target = add192(sumcheck_target, mul192(beta_lvl, level_query_sum)) # ---- finish the sumcheck over the tail coordinates ---- - tail_challenges = HeapBuf(GEN ** YR_LOG_CAP) for j in unroll(0, yr_log - 1): - f0, f1, msg_cursor, sumcheck_target, round_quad_c, round_quad_a, tail_c = opening_fold(fs[0], fs[1], msg_cursor, round_quad_c, round_quad_b, round_quad_a) - fs = [f0, f1] - tail_challenges[GEN ** j] = tail_c - round_quad_b = sumcheck_target + round_quad_a + fs, msg_cursor, sumcheck_target, round_quad_c, round_quad_a, tail_c = opening_fold(fs, msg_cursor, round_quad_c, round_quad_b, round_quad_a) + k = 3 * j + tail[k:k + 3] = tail_c + round_quad_b = add192(sumcheck_target, round_quad_a) # The closing round sends no following message. fs, tail_last = squeeze(fs) - tail_challenges[GEN ** (yr_log - 1)] = tail_last - sumcheck_target = round_quad_c + tail_last * round_quad_b + tail_last * tail_last * round_quad_a - for j in unroll(yr_log, YR_LOG_CAP): - tail_challenges[GEN ** j] = 0 - - yr_at_tail = fold_final_msg(final_msg, tail_challenges, yr_log) - - # ---- the same point, indexed by committed-witness coordinate ---- - # The folds bind coordinates in ROUND order, and level 0's folds are the lane - # fold, binding the witness's TOP k coordinates: lane l of the commitment is the - # stack block q[l * 2^(mu-k) ...], which is what makes the witness's zero - # padding whole lanes for the committer to leave out of the encode. Every - # transparent weight downstream is written in witness coordinates, so rotate the - # point left by those k rounds here, while the level shape is still - # compile-time. Padding is zero so shared prefix chains may compute unused - # entries above m without reading uninitialized cells. - lane_folds = LIG_FOLDS[m_idx * LIG_MAX_LEVELS] - fold_head = n_folds - lane_folds - point = HeapBuf(SIZE_BITS + SLOT_STRIDE_LOG) - for j in unroll(0, fold_head): - point[GEN ** j] = fold_challenges[GEN ** (lane_folds + j)] - for j in unroll(0, yr_log): - point[GEN ** (fold_head + j)] = tail_challenges[GEN ** j] - for j in unroll(0, lane_folds): - point[GEN ** (fold_head + yr_log + j)] = fold_challenges[GEN ** j] - for j in unroll(n_folds + yr_log, SIZE_BITS + SLOT_STRIDE_LOG): - point[GEN ** j] = 0 + k = 3 * (yr_log - 1) + tail[k:k + 3] = tail_last + sumcheck_target = add192(round_quad_c, mul192(tail_last, add192(round_quad_b, mul192(tail_last, round_quad_a)))) + + yr_at_tail = fold_final_msg(final_msg, tail, yr_log) # ---- per-level induced bases at the single terminal point ---- # Every query of a level runs the SAME product shape over its message-column # coordinates; only the novel-basis chain (the query position's - # subspace-vanishing walk) differs. Since + # subspace-vanishing walk, in F64) differs. Since # 1 + c_t * (1 + chain_t * inv_t) == (1 + c_t) + (c_t * inv_t) * chain_t # a coordinate's two coefficients depend on the challenge and the baked # vanishing inverse alone, so they hoist out of the query loop (one row a level, # fold coords then tail coords) and each query is one multiply-add a # coordinate. - basis_a = HeapBuf(GEN ** (n_levels * LIG_LOG_MSG_COLS_CAP)) - basis_b = HeapBuf(GEN ** (n_levels * LIG_LOG_MSG_COLS_CAP)) + basis_a = HeapBuf(GEN ** (3 * n_levels * LIG_LOG_MSG_COLS_CAP)) + basis_b = HeapBuf(GEN ** (3 * n_levels * LIG_LOG_MSG_COLS_CAP)) for lvl in unroll(0, n_levels): ml = m_idx * LIG_MAX_LEVELS + lvl prefix_len = LIG_RESIDUAL_PREFIX_LEN[ml] vanish = m_idx * LIG_MAX_VANISH_LEN + LIG_VANISH_OFF[ml] for t in unroll(0, prefix_len): - fold_c = fold_challenges[GEN ** (LIG_RESIDUAL_FOLD_OFF[ml] + t)] - basis_a[GEN ** (lvl * LIG_LOG_MSG_COLS_CAP + t)] = 1 + fold_c - basis_b[GEN ** (lvl * LIG_LOG_MSG_COLS_CAP + t)] = fold_c * LIG_VANISH_INVS[vanish + t] + fold_c = fold_at(point, LIG_RESIDUAL_FOLD_OFF[ml] + t, lane_folds, tail_start) + dst = 3 * (lvl * LIG_LOG_MSG_COLS_CAP + t) + basis_a[dst:dst + 3] = add192(ONE, fold_c[0:3]) + basis_b[dst:dst + 3] = scale(LIG_VANISH_INVS[vanish + t], fold_c[0:3]) for j in unroll(0, yr_log): - tail_c = tail_challenges[GEN ** j] - basis_a[GEN ** (lvl * LIG_LOG_MSG_COLS_CAP + prefix_len + j)] = 1 + tail_c - basis_b[GEN ** (lvl * LIG_LOG_MSG_COLS_CAP + prefix_len + j)] = tail_c * LIG_VANISH_INVS[vanish + prefix_len + j] - inner_chain = HeapBuf(GEN ** (n_levels + 1)) - inner_chain[GEN ** 0] = 0 + src = 3 * j + tail_c = tail[src:src + 3] + dst = 3 * (lvl * LIG_LOG_MSG_COLS_CAP + prefix_len + j) + basis_a[dst:dst + 3] = add192(ONE, tail_c) + basis_b[dst:dst + 3] = scale(LIG_VANISH_INVS[vanish + prefix_len + j], tail_c) + inner_chain = HeapBuf(GEN ** (3 * (n_levels + 1))) + inner_chain[0:3] = ZERO for lvl in unroll(0, n_levels): ml = m_idx * LIG_MAX_LEVELS + lvl - vanish = m_idx * LIG_MAX_VANISH_LEN + LIG_VANISH_OFF[ml] - basis_row = lvl * LIG_LOG_MSG_COLS_CAP - residual_chain = HeapBuf(GEN ** (max_q + 1)) - residual_chain[GEN ** 0] = 0 + residual_chain = HeapBuf(GEN ** (3 * (max_q + 1))) + residual_chain[0:3] = ZERO for xr in mul_range(1, GEN ** LIG_QUERIES[ml]): ml = m_idx * LIG_MAX_LEVELS + lvl # rebound: the body captures by value vanish = m_idx * LIG_MAX_VANISH_LEN + LIG_VANISH_OFF[ml] basis_row = lvl * LIG_LOG_MSG_COLS_CAP max_q = LIG_MAX_QUERIES[m_idx] basis_chain = query_positions[GEN ** LIG_POSITIONS_OFF[ml] * xr] - prefix_eq = basis_a[GEN ** basis_row] + basis_b[GEN ** basis_row] * basis_chain + b0 = 3 * basis_row + prefix_eq = add192(basis_a[b0:b0 + 3], scale(basis_chain, basis_b[b0:b0 + 3])) for t in unroll(1, LIG_LOG_MSG_COLS[ml]): # subspace-vanishing recurrence for the novel-basis point basis_chain *= (basis_chain + LIG_VANISH_VALS[vanish + t - 1]) - prefix_eq *= basis_a[GEN ** (basis_row + t)] + basis_b[GEN ** (basis_row + t)] * basis_chain - residual_chain[xr * GEN] = residual_chain[xr] + query_weights[GEN ** (lvl * max_q)* xr] * prefix_eq + bt = 3 * (basis_row + t) + prefix_eq = mul192(prefix_eq, add192(basis_a[bt:bt + 3], scale(basis_chain, basis_b[bt:bt + 3]))) + xr3 = xr ** 3 + nxt = xr3 * GEN ** 3 + qw = GEN ** (3 * lvl * max_q) * xr3 + residual_chain[nxt:nxt + 3] = add192(residual_chain[xr3:xr3 + 3], mul192(query_weights[qw:qw + 3], prefix_eq)) # accumulate beta_lvl * (per-level residual sum) into the grand residual - inner_chain[GEN ** (lvl + 1)] = inner_chain[GEN ** lvl] + level_betas[GEN ** lvl] * residual_chain[GEN ** LIG_QUERIES[ml]] + end = GEN ** (3 * LIG_QUERIES[ml]) + k = 3 * lvl + inner_chain[k + 3:k + 6] = add192(inner_chain[k:k + 3], mul192(level_betas[k:k + 3], residual_chain[end:end + 3])) # Explicit OOD eq bases at the same terminal point. - ood_inner = 0 + ood_inner = ZERO for ood_lvl in unroll(1, n_levels): ml = m_idx * LIG_MAX_LEVELS + ood_lvl z_folded = LIG_LOG_MSG_COLS[ml - 1] - yr_log ris_start = LIG_FOLDS_OFF[ml] for os in unroll(0, LIG_OOD_SAMPLES[ml]): - oz = ood_z * GEN ** ((ood_lvl * LIG_MAX_OOD_SAMPLES + os) * LIG_LOG_MSG_COLS_CAP) - scalar = ood[GEN ** (ood_lvl * ood_stride + os * OOD_SLOTS + OOD_BETA)] + oz = ood_z * GEN ** (3 * (ood_lvl * LIG_MAX_OOD_SAMPLES + os) * LIG_LOG_MSG_COLS_CAP) + sample = ood * GEN ** (ood_lvl * ood_stride + os * OOD_SLOTS) + scalar = sample[OOD_BETA:OOD_BETA + 3] for t in unroll(0, z_folded): - scalar *= (1 + oz[GEN ** t] + fold_challenges[GEN ** (ris_start + t)]) + zk = 3 * t + fc = fold_at(point, ris_start + t, lane_folds, tail_start) + scalar = mul192(scalar, add192(ONE, add192(oz[zk:zk + 3], fc[0:3]))) for t in unroll(0, yr_log): - scalar *= (1 + oz[GEN ** (z_folded + t)] + tail_challenges[GEN ** t]) - ood_inner += scalar - return sumcheck_target, point, inner_chain[GEN ** n_levels] + ood_inner, yr_at_tail + zk = 3 * (z_folded + t) + tk = 3 * t + scalar = mul192(scalar, add192(ONE, add192(oz[zk:zk + 3], tail[tk:tk + 3]))) + ood_inner = add192(ood_inner, scalar) + inner_end = 3 * n_levels + return sumcheck_target, point, add192(inner_chain[inner_end:inner_end + 3], ood_inner), yr_at_tail # ============================== inner-proof verification ============================ # The phases of `verify_sub`, in the order it runs them. Each takes and returns the -# Fiat-Shamir state pair and the stream cursor, so the sequence is what binds them. +# Fiat-Shamir state and the stream cursor, so the sequence is what binds them. -def verify_bus_gkr(fs0, fs1, cursor, g_bus_mu, zeta): +def verify_bus_gkr(fs: StackBuf(4), cursor, g_bus_mu, zeta): # ONE GKR grand product over push, pull and count, RLC-batched. Push and pull # have equal depth (matched blocks) and the count tree is padded with identity # leaves up to it (product unchanged), so a single sumcheck serves all three. @@ -1324,21 +1328,18 @@ def verify_bus_gkr(fs0, fs1, cursor, g_bus_mu, zeta): # values with the walked Fiat-Shamir state and stream cursor. layers = HeapBuf((g_bus_mu * GEN ** 2) ** LAYER_SLOTS) # mu + 2 layers rounds = HeapBuf(GKR_ROUNDS_CAP * ROUND_SLOTS) - gkr_pts = HeapBuf(GKR_POINTS_CAP) + gkr_pts = HeapBuf(3 * GKR_POINTS_CAP) assert log(g_bus_mu) < COUNT_BITS - fs = [fs0, fs1] fs, root_push, cursor = fs_next(fs, cursor) - root_pull = root_push fs, root_count, cursor = fs_next(fs, cursor) - assert root_count != 0 # count-tree root nonzero: no read count self-cancels + assert_ne192(root_count, ZERO) # count-tree root nonzero: no read count self-cancels fs, root_lambda = squeeze(fs) - layers[GEN ** LAYER_FS0] = fs[0] - layers[GEN ** LAYER_FS1] = fs[1] + layers[0:4] = fs layers[GEN ** LAYER_CURSOR] = cursor - layers[GEN ** LAYER_PUSH] = root_push - layers[GEN ** LAYER_PULL] = root_pull - layers[GEN ** LAYER_COUNT] = root_count - layers[GEN ** LAYER_LAMBDA] = root_lambda # λ over the three roots + layers[LAYER_PUSH:LAYER_PUSH + 3] = root_push + layers[LAYER_PULL:LAYER_PULL + 3] = root_push + layers[LAYER_COUNT:LAYER_COUNT + 3] = root_count + layers[LAYER_LAMBDA:LAYER_LAMBDA + 3] = root_lambda # λ over the three roots layers[GEN ** LAYER_ROW] = gkr_pts layers[GEN ** LAYER_POS] = GEN ** 0 @@ -1354,94 +1355,95 @@ def verify_bus_gkr(fs0, fs1, cursor, g_bus_mu, zeta): if depth_shift[g_bus_mu] != 1: # The odd layer is layer 0, so its round state would be written and read # back at the same position: read it straight off the layer instead. - lam = layers[GEN ** LAYER_LAMBDA] - tail_fs = [layers[GEN ** LAYER_FS0], layers[GEN ** LAYER_FS1]] + lam = layers[LAYER_LAMBDA:LAYER_LAMBDA + 3] + tail_fs = layers[0:4] tcur = layers[GEN ** LAYER_CURSOR] - tclaim = layers[GEN ** LAYER_PUSH] + lam * (layers[GEN ** LAYER_PULL] + lam * layers[GEN ** LAYER_COUNT]) - nextrow = layers[GEN ** LAYER_ROW] * GEN ** MU_CAP - evals = StackBuf(2 * N_GKR_SIDES) # the two children of each side, in side order + tclaim = add192(layers[LAYER_PUSH:LAYER_PUSH + 3], mul192(lam, add192(layers[LAYER_PULL:LAYER_PULL + 3], mul192(lam, layers[LAYER_COUNT:LAYER_COUNT + 3])))) + nextrow = gkr_pts * GEN ** (3 * MU_CAP) + evals = StackBuf(3 * 2 * N_GKR_SIDES) # the two children of each side, in side order for i in unroll(0, 2 * N_GKR_SIDES): tail_fs, ev, tcur = fs_next(tail_fs, tcur) - evals[i] = ev - combined = 0 + k = 3 * i + evals[k:k + 3] = ev + combined = ZERO for i in unroll(0, N_GKR_SIDES): side = N_GKR_SIDES - 1 - i # Horner in lam, so the top side lands last - combined = evals[2 * side] * evals[2 * side + 1] + lam * combined - assert tclaim == combined + k = 6 * side + combined = add192(mul192(evals[k:k + 3], evals[k + 3:k + 6]), mul192(lam, combined)) + assert_eq192(tclaim, combined) tail_fs, c0 = squeeze(tail_fs) - nextrow[GEN ** 0] = c0 + nextrow[0:3] = c0 tail_fs, tail_lambda = squeeze(tail_fs) # fresh λ pins the tail individuals nxt = layers * GEN ** LAYER_SLOTS - nxt[GEN ** LAYER_FS0] = tail_fs[0] - nxt[GEN ** LAYER_FS1] = tail_fs[1] + nxt[0:4] = tail_fs nxt[GEN ** LAYER_CURSOR] = tcur for side in unroll(0, N_GKR_SIDES): - nxt[GEN ** (LAYER_PUSH + side)] = evals[2 * side] + c0 * (evals[2 * side] + evals[2 * side + 1]) - nxt[GEN ** LAYER_LAMBDA] = tail_lambda + k = 6 * side + out = LAYER_PUSH + 3 * side + ev_lo = evals[k:k + 3] + nxt[out:out + 3] = add192(ev_lo, mul192(c0, add192(ev_lo, evals[k + 3:k + 6]))) + nxt[LAYER_LAMBDA:LAYER_LAMBDA + 3] = tail_lambda nxt[GEN ** LAYER_ROW] = nextrow nxt[GEN ** LAYER_POS] = GEN for x_pair in mul_range(1, pair_bounds[g_bus_mu]): x_layer = x_pair * x_pair * depth_shift[g_bus_mu] layer = layers * x_layer ** LAYER_SLOTS - lam = layer[GEN ** LAYER_LAMBDA] + lam = layer[LAYER_LAMBDA:LAYER_LAMBDA + 3] point_row = layer[GEN ** LAYER_ROW] round_pos = layer[GEN ** LAYER_POS] - nextrow = point_row * GEN ** MU_CAP + nextrow = point_row * GEN ** (3 * MU_CAP) head = rounds * round_pos ** ROUND_SLOTS - head[GEN ** ROUND_FS0] = layer[GEN ** LAYER_FS0] - head[GEN ** ROUND_FS1] = layer[GEN ** LAYER_FS1] + head[0:4] = layer[0:4] head[GEN ** ROUND_CURSOR] = layer[GEN ** LAYER_CURSOR] - head[GEN ** ROUND_CLAIM] = layer[GEN ** LAYER_PUSH] + lam * (layer[GEN ** LAYER_PULL] + lam * layer[GEN ** LAYER_COUNT]) + head[ROUND_CLAIM:ROUND_CLAIM + 3] = add192(layer[LAYER_PUSH:LAYER_PUSH + 3], mul192(lam, add192(layer[LAYER_PULL:LAYER_PULL + 3], mul192(lam, layer[LAYER_COUNT:LAYER_COUNT + 3])))) for x_round in mul_range(1, x_layer): - rd = rounds * (round_pos * x_round) ** ROUND_SLOTS - nfs0, nfs1, ncur, nclaim, rk = sumcheck_round5(rd[GEN ** ROUND_FS0], rd[GEN ** ROUND_FS1], rd[GEN ** ROUND_CURSOR], rd[GEN ** ROUND_CLAIM], point_row[x_round]) - nextrow[x_round * GEN ** 2] = rk - rd_next = rd * GEN ** ROUND_SLOTS - rd_next[GEN ** ROUND_FS0] = nfs0 - rd_next[GEN ** ROUND_FS1] = nfs1 - rd_next[GEN ** ROUND_CURSOR] = ncur - rd_next[GEN ** ROUND_CLAIM] = nclaim + xr3 = x_round ** 3 + sumcheck_round5(rounds * (round_pos * x_round) ** ROUND_SLOTS, point_row * xr3, nextrow * xr3 * GEN ** 6) tail = rounds * (round_pos * x_layer) ** ROUND_SLOTS - tail_fs = [tail[GEN ** ROUND_FS0], tail[GEN ** ROUND_FS1]] + tail_fs = tail[0:4] tcur = tail[GEN ** ROUND_CURSOR] - tclaim = tail[GEN ** ROUND_CLAIM] - evals = StackBuf(4 * N_GKR_SIDES) # the four children of each side, in side order + tclaim = tail[ROUND_CLAIM:ROUND_CLAIM + 3] + evals = StackBuf(3 * 4 * N_GKR_SIDES) # the four children of each side, in side order for i in unroll(0, 4 * N_GKR_SIDES): tail_fs, ev, tcur = fs_next(tail_fs, tcur) - evals[i] = ev - combined = 0 + k = 3 * i + evals[k:k + 3] = ev + combined = ZERO for i in unroll(0, N_GKR_SIDES): side = N_GKR_SIDES - 1 - i # Horner in lam, so the top side lands last - combined = evals[4 * side] * evals[4 * side + 1] * evals[4 * side + 2] * evals[4 * side + 3] + lam * combined - assert tclaim == combined + k = 12 * side + combined = add192(mul192(mul192(evals[k:k + 3], evals[k + 3:k + 6]), mul192(evals[k + 6:k + 9], evals[k + 9:k + 12])), mul192(lam, combined)) + assert_eq192(tclaim, combined) tail_fs, c0 = squeeze(tail_fs) tail_fs, c1 = squeeze(tail_fs) - nextrow[GEN ** 0] = c0 - nextrow[GEN ** 1] = c1 + nextrow[0:3] = c0 + nextrow[3:6] = c1 tail_fs, tail_lambda = squeeze(tail_fs) nxt = layer * GEN ** (2 * LAYER_SLOTS) - nxt[GEN ** LAYER_FS0] = tail_fs[0] - nxt[GEN ** LAYER_FS1] = tail_fs[1] + nxt[0:4] = tail_fs nxt[GEN ** LAYER_CURSOR] = tcur for side in unroll(0, N_GKR_SIDES): - lo = evals[4 * side] + c0 * (evals[4 * side] + evals[4 * side + 1]) - hi = evals[4 * side + 2] + c0 * (evals[4 * side + 2] + evals[4 * side + 3]) - nxt[GEN ** (LAYER_PUSH + side)] = lo + c1 * (lo + hi) - nxt[GEN ** LAYER_LAMBDA] = tail_lambda + k = 12 * side + out = LAYER_PUSH + 3 * side + e0 = evals[k:k + 3] + e2 = evals[k + 6:k + 9] + lo = add192(e0, mul192(c0, add192(e0, evals[k + 3:k + 6]))) + hi = add192(e2, mul192(c0, add192(e2, evals[k + 9:k + 12]))) + nxt[out:out + 3] = add192(lo, mul192(c1, add192(lo, hi))) + nxt[LAYER_LAMBDA:LAYER_LAMBDA + 3] = tail_lambda nxt[GEN ** LAYER_ROW] = nextrow nxt[GEN ** LAYER_POS] = round_pos * x_layer * GEN last = layers * g_bus_mu ** LAYER_SLOTS - fs = [last[GEN ** LAYER_FS0], last[GEN ** LAYER_FS1]] - cursor = last[GEN ** LAYER_CURSOR] final_point_row = last[GEN ** LAYER_ROW] for xt in mul_range(1, g_bus_mu): - zeta[xt] = final_point_row[xt] # the ONE shared point - return fs[0], fs[1], cursor, last[GEN ** LAYER_PUSH], last[GEN ** LAYER_PULL], last[GEN ** LAYER_COUNT] + x3 = xt ** 3 + zeta[x3:x3 + 3] = final_point_row[x3:x3 + 3] # the ONE shared point + return last[0:4], last[GEN ** LAYER_CURSOR], last[LAYER_PUSH:LAYER_PUSH + 3], last[LAYER_PULL:LAYER_PULL + 3], last[LAYER_COUNT:LAYER_COUNT + 3] -def verify_tables(fs0, fs1, cursor, pi_0, pi_1, zeta, g_bus_mu, dims_g, block_kappa, g_squares, fp_w, beta, claim_pool, claim_cplen_g, chi, claim_push, claim_pull, claim_count): - # Settle the bus against the six tables, in four steps: certify each side's +def verify_tables(fs: StackBuf(4), cursor, pi: StackBuf(4), zeta, g_bus_mu, dims_g, block_kappa, g_squares, fp_w, beta: StackBuf(3), claim_pool, claim_cplen_g, chi, claim_push: StackBuf(3), claim_pull: StackBuf(3), claim_count: StackBuf(3)): + # Settle the bus against the eight tables, in four steps: certify each side's # leaf-cube tiling, decompose the three GKR leaf values over it, run the ONE # table sumcheck they all reduce to, and bind the public input. Pooled claims # land in claim_pool/claim_cplen_g and the reduced point in `chi`; the batch's @@ -1455,11 +1457,6 @@ def verify_tables(fs0, fs1, cursor, pi_0, pi_1, zeta, g_bus_mu, dims_g, block_ka # consecutive offsets force a valid tiling, and the grand product is # position-independent, so any tiling is sound. Pull's blocks mirror push's and # share zeta, so only push and count need offsets (pull's slots go unread). - fs = [fs0, fs1] - gkr_claims = StackBuf(N_GKR_SIDES) - gkr_claims[PUSH_SIDE] = claim_push - gkr_claims[PULL_SIDE] = claim_pull - gkr_claims[COUNT_SIDE] = claim_count sort_order = HeapBuf(N_BLOCKS) hint_witness(sort_order[0:N_BLOCKS], "sort_order") block_side_tab = HeapBuf(N_BLOCKS) # global block -> its side @@ -1492,22 +1489,24 @@ def verify_tables(fs0, fs1, cursor, pi_0, pi_1, zeta, g_bus_mu, dims_g, block_ka # Index-MLE value instead of recomputing them; its column values are mostly # deduped pool reads (COORD_FRESH). The identity check against pull's own GKR # claim still binds everything. - bc_share = hint_witness("bytecode_val") + bc_share = StackBuf(3) + hint_witness(bc_share, "bytecode_val") idxc_tab = HeapBuf(SIZE_BITS) # INDEX_MLE_FACTORS[t] = 1 + g^(2^t) for t in unroll(0, SIZE_BITS): idxc_tab[GEN ** t] = INDEX_MLE_FACTORS[t] - bus_table_total = StackBuf(N_GKR_SIDES) # per side, what its tables' blocks owe - block_eq_hi = StackBuf(N_BLOCKS) # every block's eq_hi, reused below - block_index_mle = HeapBuf(N_BLOCKS) # per push block with an Index coord + bus_table_total = StackBuf(3 * N_GKR_SIDES) # per side, what its tables' blocks owe + block_eq_hi = StackBuf(3 * N_BLOCKS) # every block's eq_hi, reused below + block_index_mle = HeapBuf(3 * N_BLOCKS) # per push block with an Index coord for s in unroll(0, N_GKR_SIDES): - acc = 0 - selector_sum = 0 + acc = ZERO + selector_sum = ZERO for b in unroll(SIDE_BLOCK_START[s], SIDE_BLOCK_START[s + 1]): block_has_public = 0 kappa_g = block_kappa[GEN ** b] assert log(kappa_g) < SIZE_BITS if s == PULL_SIDE: - eq_hi = block_eq_hi[b - SIDE_BLOCK_START[PULL_SIDE]] + twin = 3 * (b - SIDE_BLOCK_START[PULL_SIDE]) + eq_hi = block_eq_hi[twin:twin + 3] else: # eq_hi over the ζ coords above κ against the selector bits, whose # run is mu_s − κ = g^mu_s / g^κ long. Selector bits = offset >> κ: @@ -1517,84 +1516,102 @@ def verify_tables(fs0, fs1, cursor, pi_0, pi_1, zeta, g_bus_mu, dims_g, block_ka # in one shot; the low κ bit cells are written but never read. sel_len_g = g_bus_mu / kappa_g # g^(mu - κ) assert log(sel_len_g) < SIZE_BITS - zeta_hi = zeta * kappa_g + zeta_hi = zeta * kappa_g ** 3 offset_bits = HeapBuf(GEN ** SIZE_BITS) hint_decompose_bits_exponent(offset_bits, block_off_g[GEN ** b], SIZE_BITS) sel_bits = offset_bits * kappa_g - eq_chain = HeapBuf(MU_CAP + 2) + eq_chain = HeapBuf(3 * (MU_CAP + 2)) goff_chain = HeapBuf(MU_CAP + 2) # rebuild g^offset from the high bits - eq_chain[GEN ** 0] = 1 + eq_chain[0:3] = ONE goff_chain[GEN ** 0] = 1 for xk in mul_range(1, sel_len_g): sbit = sel_bits[xk] sel_bits[xk] = sbit * sbit # booleanity as a write-once pin - eq_chain[xk * GEN] = eq_chain[xk] * (1 + sbit + zeta_hi[xk]) # eq over GF(2) is 1 + b + z + x3 = xk ** 3 + nxt = x3 * GEN ** 3 + eq_chain[nxt:nxt + 3] = mul192(eq_chain[x3:x3 + 3], add64(zeta_hi[x3:x3 + 3], 1 + sbit)) # eq over GF(2) is 1 + b + z goff_chain[xk * GEN] = goff_chain[xk] * (1 + sbit * (g_squares[kappa_g * xk] + 1)) - eq_hi = eq_chain[sel_len_g] + sel_end = sel_len_g ** 3 + eq_hi = eq_chain[sel_end:sel_end + 3] assert goff_chain[sel_len_g] == block_off_g[GEN ** b] # bits == offset >> κ, κ-aligned - selector_sum += eq_hi - block_eq_hi[b] = eq_hi + selector_sum = add192(selector_sum, eq_hi) + bk = 3 * b + block_eq_hi[bk:bk + 3] = eq_hi # A TABLE's block streams no value here: the table sumcheck settles its # fingerprint from that table's column evaluations. Only the framework # blocks (boundary, memory, bytecode) still decompose. if BLOCK_TABLE[b] == NO_TABLE: # inner fingerprint Σ_i w_i · coord_i(ζ_lo); the count side weighs # slot 0 alone (α⃗ = 0), γ = 0. - inner_sum = 0 + inner_sum = ZERO for i in unroll(0, BLOCK_COORD_COUNT[b]): ci = BLOCK_COORD_OFF[b] + i # a compile-time index, so `ci` costs nothing + coord_val = ZERO + slot = 3 * COORD_CLAIM_SLOT[ci] if COORD_TYPE[ci] == COORD_KIND_CONST: - coord_val = COORD_CONST[ci] + coord_val = f192(COORD_CONST[ci], 0, 0) if COORD_TYPE[ci] == COORD_KIND_COL: if COORD_FRESH[ci] == 1: fs, coord_val, cursor = fs_next(fs, cursor) - claim_pool[GEN ** COORD_CLAIM_SLOT[ci]] = coord_val + claim_pool[slot:slot + 3] = coord_val claim_cplen_g[GEN ** COORD_CLAIM_SLOT[ci]] = kappa_g # cplen = block kappa else: - coord_val = claim_pool[GEN ** COORD_CLAIM_SLOT[ci]] + coord_val = claim_pool[slot:slot + 3] if COORD_TYPE[ci] == COORD_KIND_GCOL: if COORD_FRESH[ci] == 1: fs, rawv, cursor = fs_next(fs, cursor) - claim_pool[GEN ** COORD_CLAIM_SLOT[ci]] = rawv + claim_pool[slot:slot + 3] = rawv claim_cplen_g[GEN ** COORD_CLAIM_SLOT[ci]] = kappa_g else: - rawv = claim_pool[GEN ** COORD_CLAIM_SLOT[ci]] - coord_val = COORD_CONST[ci] * rawv + rawv = claim_pool[slot:slot + 3] + coord_val = scale(COORD_CONST[ci], rawv) if COORD_TYPE[ci] == COORD_KIND_INDEX: if s == PULL_SIDE: - coord_val = block_index_mle[GEN ** (b - SIDE_BLOCK_START[PULL_SIDE])] + twin = 3 * (b - SIDE_BLOCK_START[PULL_SIDE]) + coord_val = block_index_mle[twin:twin + 3] else: # Index-coord MLE: prod_t (1 + zeta_t * (1 + g^(2^t))) - idx_chain = HeapBuf(MU_CAP + 2) - idx_chain[GEN ** 0] = 1 + idx_chain = HeapBuf(3 * (MU_CAP + 2)) + idx_chain[0:3] = ONE for xt in mul_range(1, kappa_g): - idx_chain[xt * GEN] = idx_chain[xt] * (1 + zeta[xt] * idxc_tab[xt]) - coord_val = idx_chain[kappa_g] + x3 = xt ** 3 + nxt = x3 * GEN ** 3 + idx_chain[nxt:nxt + 3] = mul192(idx_chain[x3:x3 + 3], add64(scale(idxc_tab[xt], zeta[x3:x3 + 3]), 1)) + idx_end = kappa_g ** 3 + coord_val = idx_chain[idx_end:idx_end + 3] if s == PUSH_SIDE: - block_index_mle[GEN ** b] = coord_val + block_index_mle[bk:bk + 3] = coord_val if COORD_TYPE[ci] == COORD_KIND_PUBLIC: # The public slots carry no value of their own here: their # alpha-weighted sum IS bc_share, added once per block below # (push and pull share zeta, so both get the same one). - coord_val = 0 block_has_public = 1 if s == COUNT_SIDE: - inner_sum += coord_val + inner_sum = add192(inner_sum, coord_val) else: - inner_sum += fp_w[GEN ** i] * coord_val - inner_sum += block_has_public * bc_share # the bytecode blocks' public slots + wk = 3 * i + inner_sum = add192(inner_sum, mul192(fp_w[wk:wk + 3], coord_val)) + if block_has_public == 1: + inner_sum = add192(inner_sum, bc_share) # the bytecode blocks' public slots if s == COUNT_SIDE: - acc += eq_hi * inner_sum + acc = add192(acc, mul192(eq_hi, inner_sum)) else: - acc += eq_hi * (beta + inner_sum) - acc += 1 + selector_sum + acc = add192(acc, mul192(eq_hi, add192(beta, inner_sum))) + acc = add64(add192(acc, selector_sum), 1) # What the tables' blocks owe this side: its GKR leaf value less the # framework decomposition. DERIVED, not read: a transmitted total would be a # free value in its own check. The table sumcheck's target pins it below. - bus_table_total[s] = acc + gkr_claims[s] - claim_idx = N_BUS_CLAIMS # AIR/PI/pin claims pool after the deduped bus claims - - # ---- ONE table sumcheck for all six tables ---- + if s == PUSH_SIDE: + leaf = claim_push + if s == PULL_SIDE: + leaf = claim_pull + if s == COUNT_SIDE: + leaf = claim_count + sk = 3 * s + bus_table_total[sk:sk + 3] = add192(acc, leaf) + claim_idx = N_BUS_CLAIMS # table and PI claims pool after the deduped bus claims + + # ---- ONE table sumcheck for all eight tables ---- # Mirrors lean_vm::constraints::verify. zc_xi ONCE, each table folding its own # identities with a DISJOINT range of its powers (ETA_OFFSET[t]); one shared # point zeta (the bus GKR's); n = max_t tau_t rounds. Rounds bind the HIGHEST @@ -1616,135 +1633,133 @@ def verify_tables(fs0, fs1, cursor, pi_0, pi_1, zeta, g_bus_mu, dims_g, block_ka # zeta only holds mu coords: unwritten heap there is prover-chosen. assert log(g_bus_mu / g_zc_n) < COUNT_BITS fs, zc_xi = squeeze(fs) - zc_xi_pows = StackBuf(N_ETA_POWS) - zc_xi_pows[0] = 1 + zc_xi_pows = StackBuf(3 * N_ETA_POWS) + zc_xi_pows[0:3] = ONE for k in unroll(1, N_ETA_POWS): - zc_xi_pows[k] = zc_xi_pows[k - 1] * zc_xi + c = 3 * k + zc_xi_pows[c:c + 3] = mul192(zc_xi_pows[c - 3:c], zc_xi) # The eq point is the bus GKR's zeta, NOT a fresh one, which is what lets the # batch settle the bus forms alongside the constraints. It is also why no target # is read: what the three sides' tables owe, each in its own shared power of # zc_xi, IS the sum the batch must reach, and zc_xi is squeezed after those # totals are fixed, so hitting one number forces all three side equations. - bus_target = 0 + bus_target = ZERO for sd in unroll(0, N_GKR_SIDES): - bus_target += zc_xi_pows[ETA_FORM_BASE + sd] * bus_table_total[sd] + e = 3 * (ETA_FORM_BASE + sd) + k = 3 * sd + bus_target = add192(bus_target, mul192(zc_xi_pows[e:e + 3], bus_table_total[k:k + 3])) # n vanilla sumcheck rounds: the round polynomial arrives whole, so a round is # `h(0) + h(1) == claim` and a fold, with no eq to reapply. The tables still # waiting ride inside h, so nothing here is indexed by height; the heights enter # only the per-table weights below. zc_rounds = HeapBuf((g_zc_n * GEN) ** ROUND_SLOTS) - zc_cprod = HeapBuf(g_zc_n * GEN) # the challenges bound so far, multiplied - zc_rounds[GEN ** ROUND_FS0] = fs[0] - zc_rounds[GEN ** ROUND_FS1] = fs[1] + zc_cprod = HeapBuf((g_zc_n * GEN) ** 3) # the challenges bound so far, multiplied + zc_rounds[0:4] = fs zc_rounds[GEN ** ROUND_CURSOR] = cursor - zc_rounds[GEN ** ROUND_CLAIM] = bus_target - zc_cprod[GEN ** 0] = 1 + zc_rounds[ROUND_CLAIM:ROUND_CLAIM + 3] = bus_target + zc_cprod[0:3] = ONE for xk in mul_range(1, g_zc_n): - rd = zc_rounds * xk ** ROUND_SLOTS - nfs0, nfs1, ncur, nclaim, rk = sumcheck_round4(rd[GEN ** ROUND_FS0], rd[GEN ** ROUND_FS1], rd[GEN ** ROUND_CURSOR], rd[GEN ** ROUND_CLAIM]) - chi[g_zc_n * INV_GEN / xk] = rk # g^(n-1-j): the variable round j binds - nxt = rd * GEN ** ROUND_SLOTS - nxt[GEN ** ROUND_FS0] = nfs0 - nxt[GEN ** ROUND_FS1] = nfs1 - nxt[GEN ** ROUND_CURSOR] = ncur - nxt[GEN ** ROUND_CLAIM] = nclaim + rk = sumcheck_round4(zc_rounds * xk ** ROUND_SLOTS) + bound = (g_zc_n * INV_GEN / xk) ** 3 # g^(n-1-j): the variable round j binds + chi[bound:bound + 3] = rk # cprod is the weight of a table that joins here; peq below is the rest - zc_cprod[xk * GEN] = zc_cprod[xk] * rk + x3 = xk ** 3 + nxt = x3 * GEN ** 3 + zc_cprod[nxt:nxt + 3] = mul192(zc_cprod[x3:x3 + 3], rk) zc_last = zc_rounds * g_zc_n ** ROUND_SLOTS - fs = [zc_last[GEN ** ROUND_FS0], zc_last[GEN ** ROUND_FS1]] + fs = zc_last[0:4] cursor = zc_last[GEN ** ROUND_CURSOR] - claim = zc_last[GEN ** ROUND_CLAIM] + claim = zc_last[ROUND_CLAIM:ROUND_CLAIM + 3] # peq[g^tau] = eq(zeta[..tau], chi[..tau]), as a prefix chain. - zc_peq = HeapBuf(g_zc_n * GEN) - zc_peq[GEN ** 0] = 1 - for xi in mul_range(1, g_zc_n): - zc_peq[xi * GEN] = zc_peq[xi] * (1 + zeta[xi] + chi[xi]) + zc_peq = HeapBuf((g_zc_n * GEN) ** 3) + eq_prefix_chain(zc_peq, ONE, zeta, chi, g_zc_n) # Per table: every committed column's evaluation (pooled), its AIR constraint # at its own reduced point chi[..tau_t], weighted into the batch's final claim. - air_acc = 0 + air_acc = ZERO for t in unroll(0, N_TABLES): tau_g = dims_g[GEN ** (t + 1)] - col_evals = StackBuf(TABLE_COLS_CAP) + col_evals = StackBuf(3 * TABLE_COLS_CAP) for k in unroll(0, N_TABLE_COLS[t]): fs, e, cursor = fs_next(fs, cursor) - col_evals[k] = e - claim_pool[GEN ** claim_idx] = e + c = 3 * k + col_evals[c:c + 3] = e + p = 3 * claim_idx + claim_pool[p:p + 3] = e claim_cplen_g[GEN ** claim_idx] = tau_g # cplen = tau_t claim_idx += 1 # The table's AIR constraint at the final point (col_evals is indexed by # local column index; the formulas mirror tables.rs eval_constraint). Every - # value relation now rides the bus as a degree-2 coordinate, so only JUMP's + # value relation rides the bus as a degree-2 coordinate, so only JUMP's # is-nonzero indicator is left with an identity of its own. - constraint_eval = 0 + constraint_eval = ZERO if t == TABLE_JUMP: - # `b = c*w` and `c*(b+1) = 0`. The condition is K-valued, its memory read - # carrying literal zeros above the low limb, so both identities are - # single-lane (tables.rs jump_identity). Local columns: v_cond at 5, w at - # 12, the indicator b at 13. - c = col_evals[5] - b = col_evals[13] - constraint_eval = zc_xi_pows[ETA_OFFSET[t] + 0] * (b + c * col_evals[12]) - constraint_eval += zc_xi_pows[ETA_OFFSET[t] + 1] * (c * (b + 1)) + # `b = c*w` and `c*(b+1) = 0` (tables.rs jump_identity). Local columns: + # v_cond at 5, w at 12, the indicator b at 13. + cond = col_evals[15:18] + indicator = col_evals[39:42] + x0 = 3 * ETA_OFFSET[t] + x1 = x0 + 3 + constraint_eval = add192(mul192(zc_xi_pows[x0:x0 + 3], add192(indicator, mul192(cond, col_evals[36:39]))), mul192(zc_xi_pows[x1:x1 + 3], mul192(cond, add192(indicator, ONE)))) # The table's three bus forms, evaluated at the SAME column evaluations: # Σ_b eq_hi(b) · (γ + Σ_i α^i · coord_i), the coords read off col_evals at # their local index. This is what replaces opening those columns at ζ. for sd in unroll(0, N_GKR_SIDES): - form = 0 + form = ZERO for b in unroll(0, N_BLOCKS): if BLOCK_SIDE[b] == sd: if BLOCK_TABLE[b] == t: - inner = 0 + inner = ZERO for i in unroll(0, BLOCK_COORD_COUNT[b]): # Each coord is the sum of its terms, over this table's # column evaluations. A product term (an address, an # arithmetic result) is degree 2, which the batch's # round polynomial already allows. ci = BLOCK_COORD_OFF[b] + i # a compile-time index, so `ci` costs nothing - cv = 0 + cv = ZERO for j in unroll(0, COORD_TERM_COUNT[ci]): tj = COORD_TERM_OFF[ci] + j + ca = 3 * TERM_COL_A[tj] + cb = 3 * TERM_COL_B[tj] if TERM_TYPE[tj] == COORD_KIND_CONST: - cv += TERM_CONST[tj] + cv = add192(cv, f192(TERM_CONST[tj], 0, 0)) if TERM_TYPE[tj] == COORD_KIND_COL: - cv += col_evals[TERM_COL_A[tj]] + cv = add192(cv, col_evals[ca:ca + 3]) if TERM_TYPE[tj] == COORD_KIND_GCOL: - cv += TERM_CONST[tj] * col_evals[TERM_COL_A[tj]] + cv = add192(cv, scale(TERM_CONST[tj], col_evals[ca:ca + 3])) if TERM_TYPE[tj] == COORD_KIND_PROD: - cv += TERM_CONST[tj] * (col_evals[TERM_COL_A[tj]] * col_evals[TERM_COL_B[tj]]) + cv = add192(cv, scale(TERM_CONST[tj], mul192(col_evals[ca:ca + 3], col_evals[cb:cb + 3]))) if sd == COUNT_SIDE: - inner += cv + inner = add192(inner, cv) else: - inner += fp_w[GEN ** i] * cv + wk = 3 * i + inner = add192(inner, mul192(fp_w[wk:wk + 3], cv)) + bk = 3 * b if sd == COUNT_SIDE: - form += block_eq_hi[b] * inner + form = add192(form, mul192(block_eq_hi[bk:bk + 3], inner)) else: - form += block_eq_hi[b] * (beta + inner) - constraint_eval += zc_xi_pows[ETA_FORM_BASE + sd] * form - air_acc += zc_cprod[g_zc_n / tau_g] * zc_peq[tau_g] * constraint_eval # cprod[n - tau] * peq[tau] - assert air_acc == claim - - # ---- public-input binding claim: MEM as ONE logical E-column ---- - # The VM's bind_pi_claim makes a SINGLE E-claim at [rm, 0..]: - # MEM(rm) = interp(pi_0, pi_1, rm) = pi_0 + rm*(pi_0 + pi_1) - # over the E-valued public input (no lane splitting, no Frobenius). Both - # public words have a zero top limb, so that limb's evaluation is zero at - # every rm; only the two low ones ride the stream and must reassemble it: - # MEM = v_lo + Y*v_hi (doc sec:e2e-pi). - fs, rm = squeeze(fs) - mem = pi_0 + rm * (pi_0 + pi_1) - fs, mem_lo, cursor = fs_next(fs, cursor) - fs, mem_hi, cursor = fs_next(fs, cursor) - assert mem == mem_lo + mem_hi * Y_TOWER - claim_pool[GEN ** claim_idx] = mem_lo - claim_idx += 1 - claim_pool[GEN ** claim_idx] = mem_hi - claim_idx += 1 - claim_pool[GEN ** claim_idx] = 0 - claim_idx += 1 - return fs[0], fs[1], cursor, g_zc_n, bc_share, rm - - -def verify_flock(fs0, fs1, cursor, tau_blake2s_g, zerocheck_chis, lincheck_rs, z_partial): + form = add192(form, mul192(block_eq_hi[bk:bk + 3], add192(beta, inner))) + e = 3 * (ETA_FORM_BASE + sd) + constraint_eval = add192(constraint_eval, mul192(zc_xi_pows[e:e + 3], form)) + joins = (g_zc_n / tau_g) ** 3 + sits = tau_g ** 3 + air_acc = add192(air_acc, mul192(mul192(zc_cprod[joins:joins + 3], zc_peq[sits:sits + 3]), constraint_eval)) # cprod[n - tau] * peq[tau] + assert_eq192(air_acc, claim) + + # ---- public-input binding claim ---- + # The VM's bind_pi_claim: the committed MEM at (r0, r1, 0, ...) equals the + # multilinear extension of the four public words at (r0, r1), low variable + # first. Both sides know the words, so the value is computed and nothing rides + # the stream. + fs, r0 = squeeze(fs) + fs, r1 = squeeze(fs) + v_lo = add64(scale(pi[0] + pi[1], r0), pi[0]) + v_hi = add64(scale(pi[2] + pi[3], r0), pi[2]) + p = 3 * claim_idx + claim_pool[p:p + 3] = add192(v_lo, mul192(r1, add192(v_lo, v_hi))) + return fs, cursor, g_zc_n, bc_share, r0, r1 + + +def verify_flock(fs: StackBuf(4), cursor, tau_blake2s_g, zerocheck_chis, lincheck_rs, z_partial): # Flock's zerocheck (univariate skip, k_skip = 6) then its lincheck, whose # matrix evaluation is DEFERRED to the caller's statement. The three run buffers # come in pre-sized; the point z, lincheck's alpha and the deferred matrix part @@ -1758,129 +1773,129 @@ def verify_flock(fs0, fs1, cursor, tau_blake2s_g, zerocheck_chis, lincheck_rs, z # fixed inner values then sampled outer ones. The prover builds round 1 from # this equality tail, so its sampled part is squeezed before round 1 is fetched, # and round 1 before z, which evaluates it. - fs = [fs0, fs1] mr1cs_g = tau_blake2s_g * GEN ** K_LOG # runtime m = K_LOG + tau_5, in the exponent - zerocheck_r = HeapBuf(mr1cs_g) - for i in unroll(0, N_FIXED_CHALLENGE_ROUNDS): - zerocheck_r[GEN ** (K_SKIP + i)] = FIXED_CHALLENGES[i] - flock_pts = HeapBuf((mr1cs_g * GEN ** 2) ** PAIR_SLOTS) - seed = flock_pts * (GEN ** (K_SKIP + N_FIXED_CHALLENGE_ROUNDS)) ** PAIR_SLOTS - seed[GEN ** 0] = fs[0] - seed[GEN ** 1] = fs[1] + zerocheck_r = HeapBuf(mr1cs_g ** 3) + flock_pts = HeapBuf((mr1cs_g * GEN ** 2) ** FS_SLOTS) + seed = flock_pts * (GEN ** (K_SKIP + N_FIXED_CHALLENGE_ROUNDS)) ** FS_SLOTS + seed[0:4] = fs for xi in mul_range(GEN ** (K_SKIP + N_FIXED_CHALLENGE_ROUNDS), mr1cs_g): - row = flock_pts * xi ** PAIR_SLOTS - point_fs = [row[GEN ** 0], row[GEN ** 1]] - point_fs, zerocheck_challenge = squeeze(point_fs) - zerocheck_r[xi] = zerocheck_challenge - row[GEN ** PAIR_SLOTS] = point_fs[0] - row[GEN ** (PAIR_SLOTS + 1)] = point_fs[1] - pts_last = flock_pts * mr1cs_g ** PAIR_SLOTS - fs = [pts_last[GEN ** 0], pts_last[GEN ** 1]] - # round-1 message (P = P^AB + P^C on Lambda, 2^K_SKIP words): fetch + - # observe each word as it comes off the stream, then sample z. - zc_round1 = HeapBuf(2 ** K_SKIP) + row = flock_pts * xi ** FS_SLOTS + point_fs, zerocheck_challenge = squeeze(row[0:4]) + x3 = xi ** 3 + zerocheck_r[x3:x3 + 3] = zerocheck_challenge + row[4:8] = point_fs + pts_last = flock_pts * mr1cs_g ** FS_SLOTS + fs = pts_last[0:4] + # round-1 message (P = P^AB + P^C on Lambda, 2^K_SKIP scalars): fetch + + # observe each as it comes off the stream, then sample z. + zc_round1 = StackBuf(3 * 2 ** K_SKIP) for i in unroll(0, 2 ** K_SKIP): fs, w, cursor = fs_next(fs, cursor) - zc_round1[GEN ** i] = w + k = 3 * i + zc_round1[k:k + 3] = w fs, zerocheck_z = squeeze(fs) # cursor now sits at the multilinear round messages, walked below # P(z), interpolated at z over ALL 128 phi8 nodes: the transmitted Lambda values # (nodes 64..128) plus the S half, zero by the zerocheck identity. The finished # sum is scaled once by the domain's inverse denominator; the full-domain # product only adds the S-half factor to the Lambda numerators. - lagrange_nums = StackBuf(2 ** K_SKIP) + lagrange_nums = StackBuf(3 * 2 ** K_SKIP) lag64(zerocheck_z, lagrange_nums, 2 ** K_SKIP) - s_half_product = GEN ** 0 - zc_running = 0 # the zerocheck running claim entering the multilinear rounds + s_half_product = ONE + zc_running = ZERO # the zerocheck running claim entering the multilinear rounds for i in unroll(0, 2 ** K_SKIP): - s_half_product *= (zerocheck_z + PHI8_NODES[i]) - zc_running += lagrange_nums[i] * zc_round1[GEN ** i] - zc_running *= s_half_product * LAGRANGE_INV_COMBINED - mr1cs_rounds_g = mr1cs_g * INV_GEN ** 6 # the multilinear rounds: m - 6 + k = 3 * i + s_half_product = mul192(s_half_product, add192(zerocheck_z, phi8(i))) + zc_running = add192(zc_running, mul192(lagrange_nums[k:k + 3], zc_round1[k:k + 3])) + zc_running = mul192(zc_running, mul192(s_half_product, LAGRANGE_INV_COMBINED)) for i in unroll(0, N_FIXED_CHALLENGE_ROUNDS): - r_eq = zerocheck_r[GEN ** (K_SKIP + i)] fs, g_1, cursor = fs_next(fs, cursor) # G's coefficients, bar the constant one fs, g_2, cursor = fs_next(fs, cursor) - g_0 = zc_running + r_eq * (g_1 + g_2) # the eq-weighted split fixes it + g_0 = add192(zc_running, mul192(fixed_challenge(i), add192(g_1, g_2))) # the eq-weighted split fixes it fs, chi_v = squeeze(fs) - zerocheck_chis[GEN ** i] = chi_v - zc_running = g_0 + chi_v * (g_1 + chi_v * g_2) + k = 3 * i + zerocheck_chis[k:k + 3] = chi_v + zc_running = add192(g_0, mul192(chi_v, add192(g_1, mul192(chi_v, g_2)))) # the sampled rounds: K_LOG + tau_5 - K_SKIP in all, certified nmlv_g = tau_blake2s_g * GEN ** (K_LOG - K_SKIP) - flock_rounds = HeapBuf((mr1cs_rounds_g * GEN ** 2) ** ROUND_SLOTS) + flock_rounds = HeapBuf((nmlv_g * GEN ** 2) ** ROUND_SLOTS) seed = flock_rounds * (GEN ** N_FIXED_CHALLENGE_ROUNDS) ** ROUND_SLOTS - seed[GEN ** ROUND_FS0] = fs[0] - seed[GEN ** ROUND_FS1] = fs[1] + seed[0:4] = fs seed[GEN ** ROUND_CURSOR] = cursor - seed[GEN ** ROUND_CLAIM] = zc_running + seed[ROUND_CLAIM:ROUND_CLAIM + 3] = zc_running for xi in mul_range(GEN ** N_FIXED_CHALLENGE_ROUNDS, nmlv_g): rd = flock_rounds * xi ** ROUND_SLOTS - round_fs = [rd[GEN ** ROUND_FS0], rd[GEN ** ROUND_FS1]] - r_eq = zerocheck_r[GEN ** K_SKIP * xi] - cur_i = rd[GEN ** ROUND_CURSOR] - round_fs, g_1, cur_i = fs_next(round_fs, cur_i) # coefficients, bar the constant one + x3 = xi ** 3 + r_at = x3 * GEN ** (3 * K_SKIP) + round_fs, g_1, cur_i = fs_next(rd[0:4], rd[GEN ** ROUND_CURSOR]) # coefficients, bar the constant one round_fs, g_2, cur_i = fs_next(round_fs, cur_i) - g_0 = rd[GEN ** ROUND_CLAIM] + r_eq * (g_1 + g_2) # the eq-weighted split fixes it + g_0 = add192(rd[ROUND_CLAIM:ROUND_CLAIM + 3], mul192(zerocheck_r[r_at:r_at + 3], add192(g_1, g_2))) # the eq-weighted split fixes it round_fs, chi_v = squeeze(round_fs) - zerocheck_chis[xi] = chi_v + zerocheck_chis[x3:x3 + 3] = chi_v nxt = rd * GEN ** ROUND_SLOTS - nxt[GEN ** ROUND_FS0] = round_fs[0] - nxt[GEN ** ROUND_FS1] = round_fs[1] + nxt[0:4] = round_fs nxt[GEN ** ROUND_CURSOR] = cur_i - nxt[GEN ** ROUND_CLAIM] = g_0 + chi_v * (g_1 + chi_v * g_2) + nxt[ROUND_CLAIM:ROUND_CLAIM + 3] = add192(g_0, mul192(chi_v, add192(g_1, mul192(chi_v, g_2)))) fr_last = flock_rounds * nmlv_g ** ROUND_SLOTS - fs = [fr_last[GEN ** ROUND_FS0], fr_last[GEN ** ROUND_FS1]] - zc_running = fr_last[GEN ** ROUND_CLAIM] - cursor = fr_last[GEN ** ROUND_CURSOR] # walked past all 2*n_mlv round words, now at a_eval + fs = fr_last[0:4] + zc_running = fr_last[ROUND_CLAIM:ROUND_CLAIM + 3] + cursor = fr_last[GEN ** ROUND_CURSOR] # walked past all 2*n_mlv round scalars, now at a_eval # final: observe a_eval, b_eval; the terminal identity is what defines # c_eval, so nothing is checked here. C rode the rounds above, so all three # claims sit at the same point and lincheck pins all three at once. fs, a_eval, cursor = fs_next(fs, cursor) fs, b_eval, cursor = fs_next(fs, cursor) - c_eval = zc_running + a_eval * b_eval + c_eval = add192(zc_running, mul192(a_eval, b_eval)) # The phi8 Lagrange weights at z over the S nodes: the quirky extension's # own combination, which the lincheck terminal applies to the 64 slices. - claim_nums = StackBuf(2 ** K_SKIP) + claim_nums = StackBuf(3 * 2 ** K_SKIP) lag64(zerocheck_z, claim_nums, 0) # ---- flock lincheck (matrix evaluation DEFERRED) ---- - matrix_eval = hint_witness("matpart") + matrix_eval = StackBuf(3) + hint_witness(matrix_eval, "matpart") fs, lincheck_alpha = squeeze(fs) - lincheck_beta = lincheck_alpha * lincheck_alpha - lincheck_cube = lincheck_beta * lincheck_alpha - lc_running = a_eval + lincheck_alpha * b_eval + lincheck_beta * c_eval + lincheck_cube # seed: a + alpha*b + alpha^2*c + alpha^3 (the two matrix claims, C, and the pin) + lincheck_beta = mul192(lincheck_alpha, lincheck_alpha) + lincheck_cube = mul192(lincheck_beta, lincheck_alpha) + # seed: a + alpha*b + alpha^2*c + alpha^3 (the two matrix claims, C, and the pin) + lc_running = add192(add192(a_eval, mul192(lincheck_alpha, b_eval)), add192(mul192(lincheck_beta, c_eval), lincheck_cube)) for i in unroll(0, LINCHECK_ROUNDS): fs, c0, cursor = fs_next(fs, cursor) # q's coefficients, bar the linear one fs, c2, cursor = fs_next(fs, cursor) - c1 = lc_running + c2 # the split fixes it against the running claim + c1 = add192(lc_running, c2) # the split fixes it against the running claim fs, rv = squeeze(fs) - lincheck_rs[GEN ** i] = rv - lc_running = c0 + rv * (c1 + rv * c2) # fold the degree-2 round poly at the challenge rv - # post-sumcheck collapse: fetch + observe each word + k = 3 * i + lincheck_rs[k:k + 3] = rv + lc_running = add192(c0, mul192(rv, add192(c1, mul192(rv, c2)))) # fold the degree-2 round poly at the challenge rv + # post-sumcheck collapse: fetch + observe each scalar for i in unroll(0, 2 ** K_SKIP): fs, w, cursor = fs_next(fs, cursor) - z_partial[GEN ** i] = w + k = 3 * i + z_partial[k:k + 3] = w # final consistency: running == matpart (DEFERRED) + beta * pin term. The # const-pin column folds through the top-variable bindings: weight = # prod_j (bit_{klog-1-j}(PIN_COLUMN) ? r_j : 1+r_j), surviving z_partial index # = PIN_COLUMN low 6 bits. - pin_term = lincheck_cube * eq_weight(lincheck_rs, LINCHECK_ROUNDS, PIN_COLUMN, K_LOG) - pin_term *= z_partial[GEN ** (PIN_COLUMN % 2 ** K_SKIP)] + pin = 3 * (PIN_COLUMN % 2 ** K_SKIP) + pin_term = mul192(mul192(lincheck_cube, eq_weight(lincheck_rs, LINCHECK_ROUNDS, PIN_COLUMN, K_LOG)), z_partial[pin:pin + 3]) # The C term. Its column weight is the row weight itself (C = I), and both # sides are tensors, so it collapses to eq(chi_in, chi_in_prime) times the # phi8 Lagrange combination of the 64 slices: no second matrix walk, no # second family. - c_point_eq = GEN ** 0 + c_point_eq = ONE for t in unroll(0, LINCHECK_ROUNDS): - c_point_eq *= (1 + zerocheck_chis[GEN ** t] + lincheck_rs[GEN ** (LINCHECK_ROUNDS - 1 - t)]) - c_slice_value = 0 + zk = 3 * t + lk = 3 * (LINCHECK_ROUNDS - 1 - t) + c_point_eq = mul192(c_point_eq, add192(ONE, add192(zerocheck_chis[zk:zk + 3], lincheck_rs[lk:lk + 3]))) + c_slice_value = ZERO for i in unroll(0, 2 ** K_SKIP): - c_slice_value += claim_nums[i] * z_partial[GEN ** i] - c_slice_value *= LAGRANGE_INV_S + k = 3 * i + c_slice_value = add192(c_slice_value, mul192(claim_nums[k:k + 3], z_partial[k:k + 3])) + c_slice_value = mul192(c_slice_value, LAGRANGE_INV_S) # deferred matrix eval + pin + C - assert lc_running == matrix_eval + pin_term + lincheck_beta * c_point_eq * c_slice_value + assert_eq192(lc_running, add192(add192(matrix_eval, pin_term), mul192(mul192(lincheck_beta, c_point_eq), c_slice_value))) # z_partial IS the claim: the terminal identity above pins its 64 slices, and # ring switching binds every one of them. - return fs[0], fs[1], cursor, zerocheck_z, lincheck_alpha, matrix_eval + return fs, cursor, zerocheck_z, lincheck_alpha, matrix_eval def certify_placement(kappa_base, g_squares): @@ -1927,38 +1942,45 @@ def column_selector(offset, point, kappa: Const): bits = StackBuf(MAX_STACK_LOG) hint_decompose_bits_exponent(bits, offset, MAX_STACK_LOG) rebuilt = GEN ** 0 - selector = GEN ** 0 + selector = ONE for k in unroll(kappa, MAX_STACK_LOG): bit = bits[k] bits[k] = bit * bit rebuilt *= 1 + bit * (1 + GEN ** (2 ** k)) - selector *= 1 + bit + point[GEN ** k] + c = 3 * k + selector = mul192(selector, add64(point[c:c + 3], 1 + bit)) assert rebuilt == offset return selector -def check_opening_terminal(zeta, chi, rm, g_bus_mu, g_zc_n, g_log_mem, tau_blake2s_g, claim_cplen_g, lam_pool, col_offsets, col_kappas, z_vals, c_table, point, inner_total, yr_at_tail, sumcheck_target): +def check_opening_terminal(zeta, chi, r0: StackBuf(3), r1: StackBuf(3), g_bus_mu, g_zc_n, g_log_mem, tau_blake2s_g, claim_cplen_g, lam_pool, col_offsets, col_kappas, z_vals, c_table, point, inner_total: StackBuf(3), yr_at_tail: StackBuf(3), sumcheck_target: StackBuf(3)): # Evaluate each transparent weight at the complete point in witness order. # A claim is its low point followed by the certified column's selector bits; # q_flock slots prepend their fixed slot bits to the low point. - zeta_eq_chain = HeapBuf(SIZE_BITS + 1) - eq_prefix_chain(zeta_eq_chain, 1, zeta, point, g_bus_mu) - chi_eq_chain = HeapBuf(SIZE_BITS + 1) - eq_prefix_chain(chi_eq_chain, 1, chi, point, g_zc_n) - chi_slot_eq_chain = HeapBuf(SIZE_BITS + 1) - eq_prefix_chain(chi_slot_eq_chain, 1, chi, point * GEN ** SLOT_STRIDE_LOG, tau_blake2s_g) - pi_chain = HeapBuf(SIZE_BITS + 1) - pi_chain[GEN ** 0] = 1 - pi_chain[GEN ** 1] = 1 + rm + point[GEN ** 0] - for xk in mul_range(GEN, g_log_mem): - pi_chain[xk * GEN] = pi_chain[xk] * (1 + point[xk]) - pi_eq = pi_chain[g_log_mem] - - selectors = StackBuf(N_COMMITTED_COLS) + zeta_eq_chain = HeapBuf(3 * (SIZE_BITS + 1)) + eq_prefix_chain(zeta_eq_chain, ONE, zeta, point, g_bus_mu) + chi_eq_chain = HeapBuf(3 * (SIZE_BITS + 1)) + eq_prefix_chain(chi_eq_chain, ONE, chi, point, g_zc_n) + chi_slot_eq_chain = HeapBuf(3 * (SIZE_BITS + 1)) + eq_prefix_chain(chi_slot_eq_chain, ONE, chi, point * GEN ** (3 * SLOT_STRIDE_LOG), tau_blake2s_g) + # The PI claim's point is (r0, r1, 0, ..., 0) over log_mem coordinates. + pi_chain = HeapBuf(3 * (SIZE_BITS + 1)) + pi_chain[3:6] = add192(ONE, add192(r0, point[0:3])) + pi_chain[6:9] = mul192(pi_chain[3:6], add192(ONE, add192(r1, point[3:6]))) + for xk in mul_range(GEN ** 2, g_log_mem): + x3 = xk ** 3 + nxt = x3 * GEN ** 3 + pi_chain[nxt:nxt + 3] = mul192(pi_chain[x3:x3 + 3], add192(ONE, point[x3:x3 + 3])) + pi_end = g_log_mem ** 3 + pi_eq = pi_chain[pi_end:pi_end + 3] + + selectors = StackBuf(3 * N_COMMITTED_COLS) for c in unroll(0, N_COMMITTED_COLS): offset = col_offsets[GEN ** c] kappa_g = col_kappas[GEN ** c] - selectors[c] = match(log(kappa_g), range(0, N_COLUMN_LOGS), lambda kappa: column_selector(offset, point, kappa)) + selector = match(log(kappa_g), range(0, N_COLUMN_LOGS), lambda kappa: column_selector(offset, point, kappa)) + k = 3 * c + selectors[k:k + 3] = selector inner_sum = inner_total for j in unroll(0, N_CLAIMS): @@ -1968,66 +1990,79 @@ def check_opening_terminal(zeta, chi, rm, g_bus_mu, g_zc_n, g_log_mem, tau_blake else: cplen_g = claim_cplen_g[GEN ** j] nlow = cplen_g + low_at = cplen_g ** 3 if CLAIM_POINT_BUF[j] == POINT_BUF_ZETA: - low_eq = zeta_eq_chain[cplen_g] + low_eq = zeta_eq_chain[low_at:low_at + 3] if CLAIM_POINT_BUF[j] == POINT_BUF_RHO: - low_eq = chi_eq_chain[cplen_g] + low_eq = chi_eq_chain[low_at:low_at + 3] if CLAIM_POINT_BUF[j] == POINT_BUF_QFLOCK_RHO: - slot_eq = GEN ** 0 + slot_eq = ONE for k in unroll(0, SLOT_STRIDE_LOG): - slot_eq *= 1 + CLAIM_QFLOCK_SLOT_BITS[SLOT_STRIDE_LOG * j + k] + point[GEN ** k] - low_eq = slot_eq * chi_slot_eq_chain[cplen_g] + pk = 3 * k + if CLAIM_QFLOCK_SLOT_BITS[SLOT_STRIDE_LOG * j + k] == 1: + slot_eq = mul192(slot_eq, point[pk:pk + 3]) + else: + slot_eq = mul192(slot_eq, add192(ONE, point[pk:pk + 3])) + low_eq = mul192(slot_eq, chi_slot_eq_chain[low_at:low_at + 3]) nlow = cplen_g * GEN ** SLOT_STRIDE_LOG assert nlow == col_kappas[GEN ** CLAIM_COMMITTED_COL[j]] - inner_sum += lam_pool[GEN ** j] * low_eq * selectors[CLAIM_COMMITTED_COL[j]] + lk = 3 * j + sk = 3 * CLAIM_COMMITTED_COL[j] + inner_sum = add192(inner_sum, mul192(mul192(lam_pool[lk:lk + 3], low_eq), selectors[sk:sk + 3])) qflockv_g = tau_blake2s_g * GEN ** SLOT_STRIDE_LOG assert qflockv_g == col_kappas[GEN ** QFLOCK_COMMITTED_COL] - prod_chains = HeapBuf((qflockv_g * GEN) ** BASE_FIELD_BITS) + prod_chains = HeapBuf((qflockv_g * GEN) ** (3 * BASE_FIELD_BITS)) for k in unroll(0, BASE_FIELD_BITS): - prod_chains[GEN ** k] = 1 + c = 3 * k + prod_chains[c:c + 3] = ONE rs_eq_run(prod_chains, z_vals, point, qflockv_g) - prod_final = prod_chains * qflockv_g ** BASE_FIELD_BITS - rs_weight = 0 + prod_final = prod_chains * qflockv_g ** (3 * BASE_FIELD_BITS) + rs_weight = ZERO for k in unroll(0, BASE_FIELD_BITS): - rs_weight += c_table[GEN ** k] * prod_final[GEN ** k] - inner_sum += rs_weight * selectors[QFLOCK_COMMITTED_COL] - assert inner_sum * yr_at_tail == sumcheck_target + c = 3 * k + rs_weight = add192(rs_weight, mul192(c_table[c:c + 3], prod_final[c:c + 3])) + qk = 3 * QFLOCK_COMMITTED_COL + inner_sum = add192(inner_sum, mul192(rs_weight, selectors[qk:qk + 3])) + assert_eq192(mul192(inner_sum, yr_at_tail), sumcheck_target) + return -def verify_sub(pi_0, pi_1, seed_0, seed_1, g_logs_pow2, g_squares, defer_out): - # In-circuit verification of ONE inner proof for the statement (pi_0, pi_1), - # mirroring cpu::verify step for step; the `# ---- ... ----` headers below run - # in that order. All proof data is hinted HERE, so each call pops the next - # sub-proof's entry of every witness stream and the body lowers once. The - # exponent tables are shared read-only; the deferred claims go to `defer_out`. +def verify_sub(pi: StackBuf(4), seed, g_logs_pow2, g_squares, defer_out): + # In-circuit verification of ONE inner proof for the statement `pi`, mirroring + # cpu::verify step for step; the `# ---- ... ----` headers below run in that + # order. All proof data is hinted HERE, so each call pops the next sub-proof's + # entry of every witness stream and the body lowers once. The exponent tables + # are shared read-only; the deferred claims go to `defer_out`, three words an + # element. # # The pool holds every committed-coordinate claim's value in decompose order # (the points are the GKR zetas, resolvable from the baked block structure) and # its certified low dimension, which the terminal pins its lengths against. - claim_pool = HeapBuf(N_CLAIMS) + claim_pool = HeapBuf(3 * N_CLAIMS) claim_cplen_g = HeapBuf(N_CLAIMS) # ---- seed (statement pre-bound: hinted sub pi + baked program digest) ---- - fs = StackBuf(2) - blake2s([seed_0, seed_1], [pi_0, pi_1], fs) + fs = StackBuf(4) + blake2s(seed[0:4], pi, fs) stream = HeapBuf(STREAM_CAP) hint_witness(stream[0:STREAM_CAP], "stream") - cursor = stream # the proof stream, replayed word by word (advance = * g) + cursor = stream # the proof stream, replayed scalar by scalar (advance = * g^3) # ---- announced layout and PCS rate (observed, then certified) ---- - # The stream announces the sizes as integer WORDS, one log each; the - # shape-generic phases need them as G-POWERS (loop bounds, match scrutinees), so - # each is reassembled from its advice-decomposed bits, with no hint and no - # g^j -> j lookup. Every table's rows are real rows (the prover's fill blocks - # bring each count up to a power of two), so a height is all there is to - # announce. - sizes = StackBuf(N_TABLES + 1) - for i in unroll(0, N_TABLES + 1): + # The stream announces the sizes as integer scalars, one log each, whose upper + # limbs must be zero; the shape-generic phases need them as G-POWERS (loop + # bounds, match scrutinees), so each is reassembled from its advice-decomposed + # bits, with no hint and no g^j -> j lookup. Every table's rows are real rows + # (the prover's fill blocks bring each count up to a power of two), so a height + # is all there is to announce. + sizes = StackBuf(N_TABLES + 2) + for i in unroll(0, N_TABLES + 2): fs, x, cursor = fs_next(fs, cursor) - sizes[i] = x - fs, log_inv_rate, cursor = fs_next(fs, cursor) - rate_sel = g_power_of_word(log_inv_rate, g_squares, LOG_WORD_BITS) / GEN + assert x[1] == 0 + assert x[2] == 0 + sizes[i] = x[0] + rate_sel = g_power_of_word(sizes[N_TABLES + 1], g_squares, LOG_WORD_BITS) / GEN assert log(rate_sel) < LIG_N_RATES dims_g = HeapBuf(N_TABLES + 1) # [g^log_mem, g^tau_0 .. g^tau_{N_TABLES-1}] g_log_mem = g_power_of_word(sizes[0], g_squares, LOG_WORD_BITS) @@ -2058,46 +2093,43 @@ def verify_sub(pi_0, pi_1, seed_0, seed_1, g_logs_pow2, g_squares, defer_out): for b in unroll(SIDE_BLOCK_START[PUSH_SIDE], SIDE_BLOCK_START[PUSH_SIDE + 1]): push_total *= g_squares[block_kappa[GEN ** b]] # g^(sum of 2^kappa) g_bus_mu = log2_ceil_in_the_exponent(push_total, g_logs_pow2, g_squares, 0, SIZE_BITS) - zeta = HeapBuf(g_bus_mu) # the ONE shared GKR point: exactly mu coords + zeta = HeapBuf(g_bus_mu ** 3) # the ONE shared GKR point: exactly mu coords - # ---- commitment root (2 words), kept for the opening phase ---- - fs, commit_root_0, cursor = fs_next(fs, cursor) - fs, commit_root_1, cursor = fs_next(fs, cursor) + # ---- commitment root (two halves), kept for the opening phase ---- # A non-canonical half is rejected here (merkle.rs `scalars_to_hash`); the level # roots get the same treatment at their own read. - root_cells = StackBuf(2) - root_cells[0] = assert_canonical(commit_root_0) - root_cells[1] = assert_canonical(commit_root_1) + fs, root_a, cursor = fs_next_half(fs, cursor) + fs, root_b, cursor = fs_next_half(fs, cursor) + commit_root = [root_a[0], root_a[1], root_b[0], root_b[1]] # ---- bus challenges (F192 provides the soundness margin without grinding) ---- # A tuple is fingerprinted multilinearly: slot x weighs eq(alphas, x), so a leaf # factor has total degree N_TUPLE_BITS in the challenges and the aligned bytecode # polynomial is read off at the challenge vector itself (doc sec:gp, sec:e2e-bc). - bus_alpha = HeapBuf(N_TUPLE_BITS) + bus_alpha = HeapBuf(3 * N_TUPLE_BITS) for t in unroll(0, N_TUPLE_BITS): fs, av = squeeze(fs) - bus_alpha[GEN ** t] = av - fp_w = HeapBuf(N_TUPLE_SLOTS) + k = 3 * t + bus_alpha[k:k + 3] = av + fp_w = HeapBuf(3 * N_TUPLE_SLOTS) for x in unroll(0, N_TUPLE_SLOTS): - fp_w[GEN ** x] = eq_weight(bus_alpha, N_TUPLE_BITS, x, 0) + k = 3 * x + fp_w[k:k + 3] = eq_weight(bus_alpha, N_TUPLE_BITS, x, 0) fs, beta = squeeze(fs) # ---- ONE GKR grand product: push, pull, and count RLC-batched ---- - fs0, fs1, cursor, claim_push, claim_pull, claim_count = verify_bus_gkr(fs[0], fs[1], cursor, g_bus_mu, zeta) - fs = [fs0, fs1] + fs, cursor, claim_push, claim_pull, claim_count = verify_bus_gkr(fs, cursor, g_bus_mu, zeta) # ---- the bus leaves, the table sumcheck, and the public-input claim ---- - chi = HeapBuf(SIZE_BITS) # chi[i] = the challenge that bound variable i - fs0, fs1, cursor, g_zc_n, bc_share, rm = verify_tables(fs[0], fs[1], cursor, pi_0, pi_1, zeta, g_bus_mu, dims_g, block_kappa, g_squares, fp_w, beta, claim_pool, claim_cplen_g, chi, claim_push, claim_pull, claim_count) - fs = [fs0, fs1] + chi = HeapBuf(3 * SIZE_BITS) # chi[i] = the challenge that bound variable i + fs, cursor, g_zc_n, bc_share, r0, r1 = verify_tables(fs, cursor, pi, zeta, g_bus_mu, dims_g, block_kappa, g_squares, fp_w, beta, claim_pool, claim_cplen_g, chi, claim_push, claim_pull, claim_count) # ---- flock zerocheck and lincheck (the matrix evaluation is DEFERRED) ---- tau_blake2s_g = dims_g[GEN ** (TABLE_BLAKE2s + 1)] # the BLAKE2s table's certified tau - zerocheck_chis = HeapBuf(tau_blake2s_g * GEN ** (K_LOG - K_SKIP)) # m - 6 rounds - lincheck_rs = HeapBuf(LINCHECK_ROUNDS) - z_partial = HeapBuf(2 ** K_SKIP) - fs0, fs1, cursor, zerocheck_z, lincheck_alpha, matrix_eval = verify_flock(fs[0], fs[1], cursor, tau_blake2s_g, zerocheck_chis, lincheck_rs, z_partial) - fs = [fs0, fs1] + zerocheck_chis = HeapBuf((tau_blake2s_g * GEN ** (K_LOG - K_SKIP)) ** 3) # m - 6 rounds + lincheck_rs = HeapBuf(3 * LINCHECK_ROUNDS) + z_partial = HeapBuf(3 * 2 ** K_SKIP) + fs, cursor, zerocheck_z, lincheck_alpha, matrix_eval = verify_flock(fs, cursor, tau_blake2s_g, zerocheck_chis, lincheck_rs, z_partial) # ---- stacked mixed opening: ring-switch front + claim combination ---- # The ring-switch slices are z_partial, read and bound above; this block only @@ -2105,49 +2137,58 @@ def verify_sub(pi_0, pi_1, seed_0, seed_1, g_logs_pow2, g_squares, defer_out): # 32,16,8,4,2,1: their expansion has all 64 Frobenius terms soundness needs, # while direct application costs 63 squarings and only six general # multiplications. - map_challenges = HeapBuf(6) # len(RING_MAP_SHIFTS) - c_table = HeapBuf(BASE_FIELD_BITS) - z_vals = HeapBuf(QFLOCK_VARS_CAP) + map_challenges = HeapBuf(3 * 6) # len(RING_MAP_SHIFTS) + c_table = HeapBuf(3 * BASE_FIELD_BITS) + z_vals = HeapBuf(3 * QFLOCK_VARS_CAP) for stage in unroll(0, len(RING_MAP_SHIFTS)): fs, map_challenge = squeeze(fs) - map_challenges[GEN ** stage] = map_challenge + k = 3 * stage + map_challenges[k:k + 3] = map_challenge # Expand the same composition once for the later transparent-weight evaluation. # Before shift d, the populated coefficients are exactly at multiples of 2d; the # new branch fills the adjacent d-offset entries. - c_table[GEN ** 0] = 1 + c_table[0:3] = ONE for stage in unroll(0, len(RING_MAP_SHIFTS)): shift = RING_MAP_SHIFTS[stage] - map_challenge = map_challenges[GEN ** stage] + mk = 3 * stage + map_challenge = map_challenges[mk:mk + 3] for slot in unroll(0, BASE_FIELD_BITS // (2 * shift)): - coefficient = c_table[GEN ** (slot * 2 * shift)] + src = 3 * (slot * 2 * shift) + coefficient = c_table[src:src + 3] for k in unroll(0, shift): - coefficient *= coefficient - c_table[GEN ** (slot * 2 * shift + shift)] = map_challenge * coefficient - # Evaluate the claim and combine its 64 packing rows: the running x-power and - # the running sum ride one two-slot chain. - rs_chain = HeapBuf(((2 ** K_SKIP) + 1) * PAIR_SLOTS) + coefficient = mul192(coefficient, coefficient) + dst = 3 * (slot * 2 * shift + shift) + c_table[dst:dst + 3] = mul192(map_challenge, coefficient) + # Evaluate the claim and combine its 64 packing rows: the running x-power (a + # word) and the running sum ride one four-slot chain. + rs_chain = HeapBuf(4 * ((2 ** K_SKIP) + 1)) rs_chain[GEN ** 0] = GEN ** 0 # x^i - rs_chain[GEN ** 1] = 0 # the running sum + rs_chain[1:4] = ZERO # the running sum for x_round in mul_range(1, GEN ** (2 ** K_SKIP)): - lin_eval = z_partial[x_round] + x3 = x_round ** 3 + lin_eval = z_partial[x3:x3 + 3] for stage in unroll(0, len(RING_MAP_SHIFTS)): frobenius = lin_eval for k in unroll(0, RING_MAP_SHIFTS[stage]): - frobenius *= frobenius - lin_eval += map_challenges[GEN ** stage] * frobenius - row = rs_chain * x_round ** PAIR_SLOTS + frobenius = mul192(frobenius, frobenius) + mk = 3 * stage + lin_eval = add192(lin_eval, mul192(map_challenges[mk:mk + 3], frobenius)) + row = rs_chain * x_round ** 4 x_pow = row[GEN ** 0] - row[GEN ** PAIR_SLOTS] = x_pow * 2 - row[GEN ** (PAIR_SLOTS + 1)] = row[GEN ** 1] + x_pow * lin_eval - rs_end = rs_chain * (GEN ** (2 ** K_SKIP)) ** PAIR_SLOTS - transposed_claim = rs_end[GEN ** 1] + row[GEN ** 4] = x_pow * 2 + row[5:8] = add192(row[1:4], scale(x_pow, lin_eval)) + rs_end = rs_chain * (GEN ** (2 ** K_SKIP)) ** 4 + transposed_claim = rs_end[1:4] # Suffix point for the transparent weight. for t in unroll(0, LINCHECK_ROUNDS): - z_vals[GEN ** t] = lincheck_rs[GEN ** (LINCHECK_ROUNDS - 1 - t)] - zv_lo = z_vals * GEN ** LINCHECK_ROUNDS - zr_hi = zerocheck_chis * GEN ** LINCHECK_ROUNDS + dst = 3 * t + src = 3 * (LINCHECK_ROUNDS - 1 - t) + z_vals[dst:dst + 3] = lincheck_rs[src:src + 3] + zv_lo = z_vals * GEN ** (3 * LINCHECK_ROUNDS) + zr_hi = zerocheck_chis * GEN ** (3 * LINCHECK_ROUNDS) for xt in mul_range(1, tau_blake2s_g): - zv_lo[xt] = zr_hi[xt] + x3 = xt ** 3 + zv_lo[x3:x3 + 3] = zr_hi[x3:x3 + 3] # ONE batching challenge for the whole pool: N_CLAIMS - 1 fewer Fiat-Shamir # compressions than a challenge per claim, and none for the values themselves, # `fs_next` having bound every one of them as it read it, so `lam_cl` already @@ -2155,12 +2196,13 @@ def verify_sub(pi_0, pi_1, seed_0, seed_1, g_logs_pow2, g_squares, defer_out): # the ring-switch claim takes lam_cl^0, the pool lam_cl^1 onward. fs, lam_cl = squeeze(fs) target = transposed_claim - lam_pool = HeapBuf(N_CLAIMS) + lam_pool = HeapBuf(3 * N_CLAIMS) lam_pow = lam_cl for j in unroll(0, N_CLAIMS): - lam_pool[GEN ** j] = lam_pow - target += lam_pow * claim_pool[GEN ** j] - lam_pow *= lam_cl + k = 3 * j + lam_pool[k:k + 3] = lam_pow + target = add192(target, mul192(lam_pow, claim_pool[k:k + 3])) + lam_pow = mul192(lam_pow, lam_cl) g_total, col_offsets, col_kappas = certify_placement(kappa_base, g_squares) @@ -2174,51 +2216,61 @@ def verify_sub(pi_0, pi_1, seed_0, seed_1, g_logs_pow2, g_squares, defer_out): # dispatch independently for every inner proof in a mixed-rate batch. config_sel = size_sel * rate_sel ** LIG_N_LOG_SIZES assert log(config_sel) < LIG_N_CANDIDATES - sumcheck_target, point, inner_total, yr_at_tail = match(log(config_sel), range(0, LIG_N_CANDIDATES), lambda m_idx: open_stacked(m_idx, fs[0], fs[1], target, commit_root_0, commit_root_1, cursor)) + sumcheck_target, point, inner_total, yr_at_tail = match(log(config_sel), range(0, LIG_N_CANDIDATES), lambda m_idx: open_stacked(m_idx, fs, target, commit_root, cursor)) # `stream` is a fixed-capacity witness transport. The shape fixes the exact # consumed prefix, whose every word is transcript-bound; the unused suffix # is outside the recursively verified proof and intentionally unconstrained. # ---- generalized eval_b terminal (runtime claim shapes) ---- - check_opening_terminal(zeta, chi, rm, g_bus_mu, g_zc_n, g_log_mem, tau_blake2s_g, claim_cplen_g, lam_pool, col_offsets, col_kappas, z_vals, c_table, point, inner_total, yr_at_tail, sumcheck_target) + check_opening_terminal(zeta, chi, r0, r1, g_bus_mu, g_zc_n, g_log_mem, tau_blake2s_g, claim_cplen_g, lam_pool, col_offsets, col_kappas, z_vals, c_table, point, inner_total, yr_at_tail, sumcheck_target) # ---- export this sub-proof's deferred-claim data to the caller (FRESH_*) ---- for k in unroll(0, BYTECODE_LOG): - defer_out[GEN ** k] = zeta[GEN ** k] + c = 3 * k + defer_out[c:c + 3] = zeta[c:c + 3] for k in unroll(0, LOG2_BYTECODE_COLS): - defer_out[GEN ** (BYTECODE_LOG + k)] = bus_alpha[GEN ** k] - defer_out[GEN ** FRESH_BC_VALUE] = bc_share - defer_out[GEN ** FRESH_ALPHA] = lincheck_alpha - defer_out[GEN ** FRESH_Z_SKIP] = zerocheck_z + src = 3 * k + dst = 3 * (BYTECODE_LOG + k) + defer_out[dst:dst + 3] = bus_alpha[src:src + 3] + k = 3 * FRESH_BC_VALUE + defer_out[k:k + 3] = bc_share + k = 3 * FRESH_ALPHA + defer_out[k:k + 3] = lincheck_alpha + k = 3 * FRESH_Z_SKIP + defer_out[k:k + 3] = zerocheck_z for k in unroll(0, LINCHECK_ROUNDS): - defer_out[GEN ** (FRESH_ZCHI + k)] = zerocheck_chis[GEN ** k] - defer_out[GEN ** (FRESH_LINCHECK_RS + k)] = lincheck_rs[GEN ** k] + c = 3 * k + dst = 3 * (FRESH_ZCHI + k) + defer_out[dst:dst + 3] = zerocheck_chis[c:c + 3] + dst = 3 * (FRESH_LINCHECK_RS + k) + defer_out[dst:dst + 3] = lincheck_rs[c:c + 3] for k in unroll(0, 2 ** K_SKIP): - defer_out[GEN ** (FRESH_Z_PARTIAL + k)] = z_partial[GEN ** k] - defer_out[GEN ** FRESH_MATPART] = matrix_eval + c = 3 * k + dst = 3 * (FRESH_Z_PARTIAL + k) + defer_out[dst:dst + 3] = z_partial[c:c + 3] + k = 3 * FRESH_MATPART + defer_out[k:k + 3] = matrix_eval return # ============================ XMSS signature verification =========================== # One signature of its epoch group's (epoch, message), against the signer's public -# key at `pk_ptr[g^0..g^1]` = (merkle_root, public_param). Every 16-byte native -# value (tweak, digest, chain tip, sibling, pp) is one canonical 128-bit cell. -# Tweak table layout (tweak index t at cell g^t): -# 0 : encoding tweak -# 1 + CHAIN_STEPS·i + s : chain tweak, chain i < V, step s < CHAIN_STEPS -# WOTS_PK_TWEAK_IDX : wots-pk tweak -# MERKLE_TWEAK_IDX + l : merkle tweak, level l < LOG_LIFETIME +# key at `pk[0:4]` = (merkle_root, public_param), two words each. A tweak is two +# words: a constant first word, and the index word shared by every hash at one +# epoch or, for a Merkle node, the parent's. Index table layout (at cell g^t): +# 0 : the epoch's index word (encoding, chains, wots-pk) +# 1 + l : the parent index word of Merkle level l < LOG_LIFETIME -def fill_xmss_epoch_tables(epoch, merkle_bits, tweak_table): - # The tweak table and the Merkle direction bits at `epoch`, shared by every XMSS +def fill_xmss_epoch_tables(epoch, merkle_bits, index_table): + # The index words and the Merkle direction bits at `epoch`, shared by every XMSS # signature this node verifies at it. One bit decomposition gives all three uses: # a tweak's index field is the epoch (encoding, chain, wots-pk) or the parent - # index `epoch >> lvl` at Merkle level lvl - 1, and the direction bit at that - # level IS bit lvl - 1. Booleanity is a write-once pin and the reconstruction - # ties the bits back to the epoch, which also bounds it to LOG_LIFETIME bits. - # SPHINCS shares none of this, deriving every tweak from the index its own - # digest picks, which is neither public nor shared between signers. + # index `epoch >> (l + 1)` at Merkle level l, and the direction bit at that + # level IS bit l. Booleanity is a write-once pin and the reconstruction ties the + # bits back to the epoch, which also bounds it to LOG_LIFETIME bits. SPHINCS + # shares none of this, deriving every tweak from the index its own digest picks, + # which is neither public nor shared between signers. hint_decompose_bits(merkle_bits, epoch, LOG_LIFETIME) bits = StackBuf(LOG_LIFETIME) reconstructed = 0 @@ -2227,97 +2279,98 @@ def fill_xmss_epoch_tables(epoch, merkle_bits, tweak_table): bit = merkle_bits[GEN ** b] merkle_bits[GEN ** b] = bit * bit bits[b] = bit - reconstructed += bit * COORD_BASIS[b] + reconstructed += bit * 2 ** b index += bit * XM_INDEX_WEIGHT[b] assert reconstructed == epoch - tweak_table[1] = index + XM_ENC_TWEAK - for i in unroll(0, V): - for s in unroll(0, CHAIN_STEPS): - tweak_table[GEN ** (1 + CHAIN_STEPS * i + s)] = index + XM_CHAIN_TWEAKS[CHAIN_STEPS * i + s] - tweak_table[GEN ** WOTS_PK_TWEAK_IDX] = index + XM_PK_TWEAK + index_table[GEN ** 0] = index # Merkle level lvl - 1 hashes the parent at `epoch >> lvl`: the epoch's bits from # lvl up, each weighed lvl places down. The top level gets the empty sum. for lvl in unroll(1, LOG_LIFETIME + 1): parent = 0 for b in unroll(lvl, LOG_LIFETIME): parent += bits[b] * XM_INDEX_WEIGHT[b - lvl] - tweak_table[GEN ** (MERKLE_TWEAK_IDX + lvl - 1)] = parent + XM_MERKLE_TWEAKS[lvl - 1] + index_table[GEN ** lvl] = parent return -def verify_sig(message, tweak_table, merkle_bits, pk_ptr): - pp = pk_ptr[GEN] +def verify_sig(message, index_table, merkle_bits, pk): + pp = pk[2:4] + index = index_table[GEN ** 0] # Encoding digest D = BLAKE2s(tweak | pp | msg | randomness | zero-pad), 96 - # bytes: one full 64-byte block then a 32-byte final block (24 bytes of - # randomness and the specified 8-byte zero pad). A packing helper source is read - # as (lo, 0, 0) where BLAKE2s reads (lo, hi, 0), which is what pins that pad. - after_msg = StackBuf(WORDS_PER_BLOCK) - blake2s([tweak_table[1], pp], [message[1], message[GEN]], after_msg, counter=64, final=0) - rand_block = StackBuf(WORDS_PER_BLOCK) - hint_witness(rand_block, "rand") - assert_in_k(rand_block[1], 0) - digest = StackBuf(WORDS_PER_BLOCK) - blake2s(rand_block, [0, 0], digest, cv=after_msg, counter=96, final=1) + # bytes: one full 64-byte block then a 32-byte final block, 24 bytes of + # randomness and the specified 8-byte zero pad, which is a zero word. + after_msg = StackBuf(4) + blake2s([XM_ENC_TWEAK, index, pp], message[0:4], after_msg, counter=64, final=0) + rand = StackBuf(3) + hint_witness(rand, "rand") + digest = StackBuf(4) + blake2s([rand, 0], [0, 0, 0, 0], digest, cv=after_msg, counter=96, final=1) # V WOTS chains. Per chain the digit is hinted in the exponent (g^{e_i}), range # checked and dispatched once; arm k walks the remaining CHAIN_STEPS-k steps and - # returns the tip cell plus the digit literal. The product of the digits is the + # returns the tip plus the digit literal. The product of the digits is the # target sum (g^{Σe_i}); the digits, weighted by CHAIN_LENGTH^i inside their own - # 64-bit lane (DIGITS_PER_WORD digits a lane, GF(2^64)'s monomial budget, each - # lane's leftover top bits ground to zero by the signer), reconstruct D's first - # cell as `acc_lo + acc_hi·Y`. - tips = StackBuf(TIP_CELLS) - chain_tweaks = tweak_table * GEN ** WORDS_PER_VALUE # chain i at cell 1 + CHAIN_STEPS·i + # digest word (DIGITS_PER_WORD digits a word, GF(2^64)'s monomial budget, each + # word's leftover top bits ground to zero by the signer), reconstruct D's first + # two words. + tips = StackBuf(TIP_WORDS) digit_product = 1 acc_lo = 0 acc_hi = 0 for i in unroll(0, V): digit = hint_witness("digits") assert log(digit) < CHAIN_LENGTH - chain_start = hint_witness("chain_starts") - tips[i], e = match(log(digit), range(0, CHAIN_LENGTH), lambda k: walk(chain_start, chain_tweaks, pp, k)) + start = StackBuf(2) + hint_witness(start, "chain_starts") + tips[2 * i], tips[2 * i + 1], e = match(log(digit), range(0, CHAIN_LENGTH), lambda k: walk(start[0], start[1], XM_CHAIN_TWEAK + i * CHAIN_LENGTH * XM_P_MUL, index, pp[0], pp[1], k)) digit_product = digit_product * digit term = e * CHAIN_LENGTH ** (i % DIGITS_PER_WORD) # e_i in its monomial subspace if i // DIGITS_PER_WORD == 0: acc_lo = acc_lo + term else: acc_hi = acc_hi + term - chain_tweaks = chain_tweaks * GEN ** (WORDS_PER_VALUE * CHAIN_STEPS) assert digit_product == GEN ** TARGET_SUM - assert acc_lo + acc_hi * Y_TOWER == digest[0] + assert acc_lo == digest[0] + assert acc_hi == digest[1] # WOTS public-key leaf = standard BLAKE2s over prefix + V tips: WOTS_PK_BLOCKS # full blocks, carrying the chaining value between instructions. - leaf = StackBuf(WORDS_PER_BLOCK) - blake2s([tweak_table[GEN ** (WORDS_PER_VALUE * WOTS_PK_TWEAK_IDX)], pp], tips[0:2], leaf, counter=64, final=0) + leaf = StackBuf(4) + blake2s([XM_PK_TWEAK, index, pp], tips[0:4], leaf, counter=64, final=0) for q in unroll(1, WOTS_PK_BLOCKS): - next_leaf = StackBuf(WORDS_PER_BLOCK) - blake2s(tips[4 * q - 2:4 * q], tips[4 * q:4 * q + 2], next_leaf, cv=leaf, counter=64 * (q + 1), final=(q + 1) // WOTS_PK_BLOCKS) + next_leaf = StackBuf(4) + blake2s(tips[8 * q - 4:8 * q], tips[8 * q:8 * q + 4], next_leaf, cv=leaf, counter=64 * (q + 1), final=(q + 1) // WOTS_PK_BLOCKS) leaf = next_leaf # Merkle path from the leaf to the root: the epoch bit orders the two children at # each level, and the tweak carries that level's parent index. - node = leaf[0] + n0 = leaf[0] + n1 = leaf[1] for lvl in unroll(0, LOG_LIFETIME): - sibling = hint_witness("siblings") - children = order_children(node, sibling, merkle_bits[GEN ** (WORDS_PER_VALUE * lvl)]) - parent = StackBuf(WORDS_PER_BLOCK) - blake2s([tweak_table[GEN ** (WORDS_PER_VALUE * (MERKLE_TWEAK_IDX + lvl))], pp], children, parent) - node = parent[0] - assert node == pk_ptr[1] + sibling = StackBuf(2) + hint_witness(sibling, "siblings") + children = order_children(n0, n1, sibling, merkle_bits[GEN ** lvl]) + parent = StackBuf(4) + blake2s([XM_MERKLE_TWEAK + const((lvl + 1) * XM_P_MUL), index_table[GEN ** (lvl + 1)], pp], children, parent) + n0 = parent[0] + n1 = parent[1] + assert n0 == pk[1] + assert n1 == pk[GEN] return -def walk(value, chain_tweaks, pp, k: Const): - # Walk WOTS chain steps k..CHAIN_STEPS-1: value' = H(tweak|pp, value|0). Step s - # reads its tweak at cell s off the chain's subtable, a compile-time offset. - word = value +def walk(v0, v1, tw, index, pp0, pp1, k: Const): + # Walk WOTS chain steps k..CHAIN_STEPS-1: value' = H(tweak|pp, value|0). Step s's + # tweak is the chain's first word plus s sub-positions. + w0 = v0 + w1 = v1 for s in unroll(k, CHAIN_STEPS): - out = StackBuf(WORDS_PER_BLOCK) - blake2s([chain_tweaks[GEN ** (WORDS_PER_VALUE * s)], pp], [word, 0], out, counter=48, final=1) - word = out[0] - return word, k + out = StackBuf(4) + blake2s([tw + s * XM_P_MUL, index, pp0, pp1], [w0, w1, 0, 0], out, counter=48, final=1) + w0 = out[0] + w1 = out[1] + return w0, w1, k # ========================== SPHINCS+ signature verification ========================= @@ -2325,56 +2378,58 @@ def walk(value, chain_tweaks, pp, k: Const): @inline def sp_bit_field(bits_ptr, off: Const, n: Const, pos: Const): - # The integer held by bits [off, off+n) of the digest, weighed into the - # coordinate basis at `pos`: a tweak field placed where the tweak wants it, one - # fused multiply-add a bit, whatever lane the bits came from. + # The integer held by bits [off, off+n) of the digest, weighed into a tweak word + # at bit `pos`: a tweak field placed where the tweak wants it, one fused + # multiply-add a bit, whatever digest word the bits came from. acc = 0 for i in unroll(0, n): - acc += bits_ptr[GEN ** (off + i)] * COORD_BASIS[pos + i] + acc += bits_ptr[GEN ** (off + i)] * 2 ** (pos + i) return acc -def sp_walk(value, tw_base, pp, k: Const): +def sp_walk(v0, v1, tw, tw1, pp0, pp1, k: Const): # Walk chain steps k..SP_CHAIN_STEPS-1: value' = Th(P, tw_chain, value). - # `tw_base` already carries the type byte, the layer, 2^w*i and the position - # (tau, e), so step s's tweak is one addition of a compile-time literal. - word = value + # `tw` already carries the type byte, the layer and 2^w*i, and `tw1` the + # position (tau, e), so step s's tweak is one addition of a compile-time literal. + w0 = v0 + w1 = v1 for s in unroll(k, SP_CHAIN_STEPS): - out = StackBuf(WORDS_PER_BLOCK) - blake2s([tw_base + s * SP_P_MUL, pp], [word, 0], out, counter=48, final=1) - word = out[0] - return word, k + out = StackBuf(4) + blake2s([tw + s * SP_P_MUL, tw1, pp0, pp1], [w0, w1, 0, 0], out, counter=48, final=1) + w0 = out[0] + w1 = out[1] + return w0, w1, k -def sp_ots_leaf(tw_pos, pp, msg): - # One layer's one-time verification: the encoding of `msg` under the hinted +def sp_ots_leaf(tw_lay, tw1, pp0, pp1, m0, m1): + # One layer's one-time verification: the encoding of `m` under the hinted # counter, the V chains walked from the revealed values, and the leaf they hash - # to. `tw_pos` is the position's tweak base (layer, tau, e); this function is - # called once per layer, so the V dispatch tables are compiled once for the - # whole scheme. + # to. `tw_lay` is the layer's first-word field and `tw1` the position word (tau, + # e); this function is called once per layer, so the V dispatch tables are + # compiled once for the whole scheme. ctr = hint_witness("sp_counter") ctr_bits = HeapBuf(GEN ** SP_COUNTER_BITS) hint_decompose_bits(ctr_bits, ctr, SP_COUNTER_BITS) - bind_bits(ctr_bits, ctr, SP_COUNTER_BITS) # LE_32: four counter bytes, twelve of padding + bind_bits(ctr_bits, ctr, SP_COUNTER_BITS) # LE_32: four counter bytes, four of padding - # D = Th(P, tw_enc, msg | LE_32(c)), a 52-byte one-block hash. - digest = StackBuf(WORDS_PER_BLOCK) - blake2s([tw_pos + SP_TW_ENC, pp], [msg, ctr], digest, counter=52, final=1) + # D = Th(P, tw_enc, m | LE_32(c)), a 52-byte one-block hash. + digest = StackBuf(4) + blake2s([tw_lay + SP_TW_ENC, tw1, pp0, pp1], [m0, m1, ctr, 0], digest, counter=52, final=1) # The codeword, as in XMSS: each digit hinted in the exponent, range checked and # dispatched once, arm k walking the remaining steps; the product of the digits - # is the target sum, and the digits weighted by 2^w within each 64-bit lane - # reconstruct D, which pins each lane's leftover top bits to zero. - tips = StackBuf(SP_TIP_CELLS) + # is the target sum, and the digits weighted by 2^w within each digest word + # reconstruct D, which pins each word's leftover top bits to zero. + tips = StackBuf(SP_TIP_WORDS) digit_product = 1 acc_lo = 0 acc_hi = 0 for i in unroll(0, SP_V): digit = hint_witness("sp_digits") assert log(digit) < SP_CHAIN_LENGTH - chain_start = hint_witness("sp_chain_starts") - tw_chain = tw_pos + SP_TW_CHAIN + i * SP_CHAIN_MUL - tips[i], e = match(log(digit), range(0, SP_CHAIN_LENGTH), lambda k: sp_walk(chain_start, tw_chain, pp, k)) + start = StackBuf(2) + hint_witness(start, "sp_chain_starts") + tips[2 * i], tips[2 * i + 1], e = match(log(digit), range(0, SP_CHAIN_LENGTH), lambda k: sp_walk(start[0], start[1], tw_lay + SP_TW_CHAIN + i * SP_CHAIN_MUL, tw1, pp0, pp1, k)) digit_product = digit_product * digit term = e * SP_CHAIN_LENGTH ** (i % SP_DIGITS_PER_WORD) if i // SP_DIGITS_PER_WORD == 0: @@ -2382,51 +2437,42 @@ def sp_ots_leaf(tw_pos, pp, msg): else: acc_hi = acc_hi + term assert digit_product == GEN ** SP_TARGET_SUM - assert acc_lo + acc_hi * Y_TOWER == digest[0] + assert acc_lo == digest[0] + assert acc_hi == digest[1] - leaf = StackBuf(WORDS_PER_BLOCK) - blake2s([tw_pos + SP_TW_LEAF, pp], tips[0:2], leaf, counter=64, final=0) + leaf = StackBuf(4) + blake2s([tw_lay + SP_TW_LEAF, tw1, pp0, pp1], tips[0:4], leaf, counter=64, final=0) for q in unroll(1, SP_LEAF_BLOCKS): - next_leaf = StackBuf(WORDS_PER_BLOCK) - blake2s(tips[4 * q - 2:4 * q], tips[4 * q:4 * q + 2], next_leaf, cv=leaf, counter=64 * (q + 1), final=(q + 1) // SP_LEAF_BLOCKS) + next_leaf = StackBuf(4) + blake2s(tips[8 * q - 4:8 * q], tips[8 * q:8 * q + 4], next_leaf, cv=leaf, counter=64 * (q + 1), final=(q + 1) // SP_LEAF_BLOCKS) leaf = next_leaf - return leaf[0] + return leaf[0], leaf[1] def verify_sig_sphincs(signer): - # `signer` is one 4-cell entry of the SPHINCS coverage table: the key's root and + # `signer` is one 8-word entry of the SPHINCS coverage table: the key's root and # public parameter, then the message THAT signer signed. Where XMSS's message is # one statement field for the whole node, a SPHINCS message rides its own slot, # and the signer-set digest binds the two together. - pp = signer[GEN] + pp = signer[2:4] # ---- the message digest, which chooses the few-time key ---- # D = Truncate(H(tw_msg | P | rho | root | m)), 96 bytes in two blocks. - rho_root = StackBuf(WORDS_PER_BLOCK) - hint_witness(rho_root[0:1], "sp_rand") - rho_root[1] = signer[1] - prefix = StackBuf(WORDS_PER_BLOCK) - blake2s([SP_TW_MSG, pp], rho_root, prefix, counter=64, final=0) - digest = StackBuf(WORDS_PER_BLOCK) - blake2s([signer[GEN ** 2], signer[GEN ** 3]], [0, 0], digest, cv=prefix, counter=96, final=1) + rho = StackBuf(2) + hint_witness(rho, "sp_rand") + prefix = StackBuf(4) + blake2s([SP_TW_MSG, 0, pp], [rho, signer[0:2]], prefix, counter=64, final=0) + digest = StackBuf(4) + blake2s(signer[4:8], [0, 0, 0, 0], digest, cv=prefix, counter=96, final=1) # The index and the k leaf indices are bit fields of that digest, so its bits are - # advice-decomposed here and bound lane by lane. Nothing else derives them: every + # advice-decomposed here and bound word by word. Nothing else derives them: every # tweak below is built from these bits. bits = HeapBuf(GEN ** SP_BIT_CELLS) - lo = StackBuf(1) - hint_f192_limbs(lo, digest[0]) - hi = (digest[0] + lo[0]) * Y_INV - assert_in_k(lo[0], hi) - tail = StackBuf(1) - hint_f192_limbs(tail, digest[1]) - tail_hi = (digest[1] + tail[0]) * Y_INV - assert_in_k(tail[0], tail_hi) - lanes = [lo[0], hi, tail[0]] for lane in unroll(0, SP_BIT_LANES): - run = bits * GEN ** (lane * BASE_FIELD_BITS) - hint_decompose_bits(run, lanes[lane], BASE_FIELD_BITS) - bind_bits(run, lanes[lane], BASE_FIELD_BITS) + lane_bits = bits * GEN ** (lane * BASE_FIELD_BITS) + hint_decompose_bits(lane_bits, digest[lane], BASE_FIELD_BITS) + bind_bits(lane_bits, digest[lane], BASE_FIELD_BITS) # The digest is admissible only if its last leaf index is zero, which is what # lets the forest drop that tree. @@ -2435,34 +2481,39 @@ def verify_sig_sphincs(signer): # ---- the few-time signature: one opened leaf per tree of the forest ---- idx_tau = sp_bit_field(bits, 0, SP_H, SP_TAU_POS) - roots = StackBuf(SP_N_FTS) + roots = StackBuf(2 * SP_N_FTS) for kappa in unroll(0, SP_N_FTS): leaf_off = SP_H + kappa * SP_A - secret = StackBuf(WORDS_PER_BLOCK) - hint_witness(secret[0:1], "sp_fts_secrets") - fts_leaf = StackBuf(WORDS_PER_BLOCK) + secret = StackBuf(2) + hint_witness(secret, "sp_fts_secrets") + fts_leaf = StackBuf(4) node_index = sp_bit_field(bits, leaf_off, SP_A, SP_J_POS) - blake2s([SP_TW_FTS_LEAF + kappa * SP_LAY_MUL + idx_tau + node_index, pp], [secret[0], 0], fts_leaf, counter=48, final=1) - node = fts_leaf[0] + blake2s([SP_TW_FTS_LEAF + kappa * SP_LAY_MUL, idx_tau + node_index, pp], [secret, 0, 0], fts_leaf, counter=48, final=1) + n0 = fts_leaf[0] + n1 = fts_leaf[1] for level in unroll(0, SP_A): - sibling = hint_witness("sp_fts_paths") - children = order_children(node, sibling, bits[GEN ** (leaf_off + level)]) - parent = StackBuf(WORDS_PER_BLOCK) + sibling = StackBuf(2) + hint_witness(sibling, "sp_fts_paths") + children = order_children(n0, n1, sibling, bits[GEN ** (leaf_off + level)]) + parent = StackBuf(4) if const(level + 1 == SP_A): node_index = 0 else: - # The index fits in one lane; clearing its low bit makes division by GEN a right shift. - node_index = (node_index + bits[GEN ** (leaf_off + level)] * COORD_BASIS[SP_J_POS]) / GEN - blake2s([SP_TW_FTS_NODE + kappa * SP_LAY_MUL + const((level + 1) * SP_P_MUL) + idx_tau + node_index, pp], children, parent) - node = parent[0] - roots[kappa] = node - fts_key = StackBuf(WORDS_PER_BLOCK) - blake2s([SP_TW_FTS_ROOTS + idx_tau, pp], roots[0:2], fts_key, counter=64, final=0) + # The index fits in one word; clearing its low bit makes division by GEN a right shift. + node_index = (node_index + bits[GEN ** (leaf_off + level)] * 2 ** SP_J_POS) / GEN + blake2s([SP_TW_FTS_NODE + kappa * SP_LAY_MUL + const((level + 1) * SP_P_MUL), idx_tau + node_index, pp], children, parent) + n0 = parent[0] + n1 = parent[1] + roots[2 * kappa] = n0 + roots[2 * kappa + 1] = n1 + fts_key = StackBuf(4) + blake2s([SP_TW_FTS_ROOTS, idx_tau, pp], roots[0:4], fts_key, counter=64, final=0) for q in unroll(1, SP_ROOT_BLOCKS): - next_key = StackBuf(WORDS_PER_BLOCK) - blake2s(roots[4 * q - 2:4 * q], roots[4 * q:4 * q + 2], next_key, cv=fts_key, counter=64 * (q + 1), final=(q + 1) // SP_ROOT_BLOCKS) + next_key = StackBuf(4) + blake2s(roots[8 * q - 4:8 * q], roots[8 * q:8 * q + 4], next_key, cv=fts_key, counter=64 * (q + 1), final=(q + 1) // SP_ROOT_BLOCKS) fts_key = next_key - signed = fts_key[0] + s0 = fts_key[0] + s1 = fts_key[1] # ---- the hypertree, bottom layer first ---- # Layer lay signs what the layer below produced: the few-time key at the bottom, @@ -2472,103 +2523,144 @@ def verify_sig_sphincs(signer): leaf_index_off = SP_SUFFIX[lay + 1] tau_field = sp_bit_field(bits, SP_SUFFIX[lay], SP_H - SP_SUFFIX[lay], SP_TAU_POS) node_index = sp_bit_field(bits, leaf_index_off, SP_HEIGHTS[lay], SP_J_POS) - tw_pos = tau_field + node_index + lay * SP_LAY_MUL - node = sp_ots_leaf(tw_pos, pp, signed) + n0, n1 = sp_ots_leaf(lay * SP_LAY_MUL, tau_field + node_index, pp[0], pp[1], s0, s1) for level in unroll(0, SP_HEIGHTS[lay]): - sibling = hint_witness("sp_siblings") - children = order_children(node, sibling, bits[GEN ** (leaf_index_off + level)]) - parent = StackBuf(WORDS_PER_BLOCK) + sibling = StackBuf(2) + hint_witness(sibling, "sp_siblings") + children = order_children(n0, n1, sibling, bits[GEN ** (leaf_index_off + level)]) + parent = StackBuf(4) if const(level + 1 == SP_HEIGHTS[lay]): node_index = 0 else: - node_index = (node_index + bits[GEN ** (leaf_index_off + level)] * COORD_BASIS[SP_J_POS]) / GEN - blake2s([SP_TW_NODE + lay * SP_LAY_MUL + const((level + 1) * SP_P_MUL) + tau_field + node_index, pp], children, parent) - node = parent[0] - signed = node - assert signed == signer[1] + node_index = (node_index + bits[GEN ** (leaf_index_off + level)] * 2 ** SP_J_POS) / GEN + blake2s([SP_TW_NODE + lay * SP_LAY_MUL + const((level + 1) * SP_P_MUL), tau_field + node_index, pp], children, parent) + n0 = parent[0] + n1 = parent[1] + s0 = n0 + s1 = n1 + assert s0 == signer[1] + assert s1 == signer[GEN] return # =========================== statements and the signer set ========================== -def statement_digest(seed_0, seed_1, signers_hash, da_0, da_1, defer): - # A node's statement, hashed to the two words the VM publishes, over the - # proving environment's Fiat-Shamir seed (flock's R1CS and this bytecode), the - # two-cell signer-set digest, the digest of its possibly empty DA root list, - # and the DEFER_STMT_CELLS deferred-claim cells. A - # parent rebuilds a child's statement with this same call, over a signer-set - # digest it re-absorbed itself, which forces the child to be a proof of THIS - # bytecode over groups checked against the parent's own. +def statement_digest(seed, signers: StackBuf(4), da: StackBuf(4), defer): + # A node's statement, hashed to the four words the VM publishes, over the + # proving environment's Fiat-Shamir seed (at `seed`: flock's R1CS and this + # bytecode), the signer-set digest, the digest of its possibly empty DA root + # list, and the DEFER_STMT_CELLS deferred elements. A parent rebuilds a child's + # statement with this same call, over a signer-set digest it re-absorbed itself, + # which forces the child to be a proof of THIS bytecode over groups checked + # against the parent's own. # - # The preimage is fixed-length, so a plain BLAKE2s beats the Fiat-Shamir chain. - # A header value is a canonical cell and needs no check, the BLAKE2s table - # reading only cells whose top limb is zero. A deferred cell is a full field - # element, so two fill three cells as (s0,s1) (s2,t0) (t1,t2), each top limb - # derived from the two hinted below it and each pack proving its lanes in K. - cells = StackBuf(4 * STMT_BLOCKS) - cells[0] = seed_0 # the STMT_HEADER header cells - cells[1] = seed_1 - cells[2] = signers_hash[1] - cells[3] = signers_hash[GEN] - cells[4] = da_0 - cells[5] = da_1 - for p in unroll(0, STMT_PAIRS): - s = defer[GEN ** (2 * p)] - if const(2 * p + 1 == DEFER_STMT_CELLS): - t = 0 # an odd cell count pairs the last one with a zero partner - else: - t = defer[GEN ** (2 * p + 1)] - s_lo = StackBuf(2) - t_lo = StackBuf(2) - hint_f192_limbs(s_lo, s) - hint_f192_limbs(t_lo, t) - cells[STMT_DEFER_OFF + 3 * p] = pack64x2(s_lo[0], s_lo[1]) - cells[STMT_DEFER_OFF + 3 * p + 1] = pack64x2(((s + s_lo[0]) * Y_INV + s_lo[1]) * Y_INV, t_lo[0]) - cells[STMT_DEFER_OFF + 3 * p + 2] = pack64x2(t_lo[1], ((t + t_lo[0]) * Y_INV + t_lo[1]) * Y_INV) - for k in unroll(0, STMT_PAD_CELLS): - cells[STMT_DEFER_OFF + 3 * STMT_PAIRS + k] = 0 - st = StackBuf(2) - blake2s(cells[0:2], cells[2:4], st, counter=64, final=1 // STMT_BLOCKS) + # The preimage is fixed-length, so a plain BLAKE2s beats the Fiat-Shamir chain: + # the seed and signer-set digest are the first block, then the DA digest and + # the deferred elements' limbs, zero-filled to a whole block. + words = StackBuf(8 * (STMT_BLOCKS - 1)) + words[0:4] = da + for p in unroll(0, DEFER_STMT_CELLS): + k = 3 * p + words[4 + k:4 + k + 3] = defer[k:k + 3] + for k in unroll(0, STMT_PAD): + words[4 + 3 * DEFER_STMT_CELLS + k] = 0 + st = StackBuf(4) + blake2s(seed[0:4], signers, st, counter=64, final=0) for b in unroll(1, STMT_BLOCKS): - nxt = StackBuf(2) - blake2s(cells[4 * b:4 * b + 2], cells[4 * b + 2:4 * b + 4], nxt, cv=st, counter=64 * (b + 1), final=(b + 1) // STMT_BLOCKS) + lo = 8 * (b - 1) + nxt = StackBuf(4) + blake2s(words[lo:lo + 4], words[lo + 4:lo + 8], nxt, cv=st, counter=64 * (b + 1), final=(b + 1) // STMT_BLOCKS) st = nxt - return st[0], st[1] + return st -def keys_window(state_0, state_1, base, keys_ptr, x_q, g_squares): +def keys_window(state: StackBuf(4), base, keys_ptr, x_q, g_squares): # One window of an epoch group's key hash: SIGNERS_WINDOW blocks, two declared - # keys each (a key is two cells, a block four). Counters as in `sphincs_window`. + # keys each (a key is four words, a block eight). Counters as in `sphincs_window`. nxt = scaled_log(x_q * GEN, g_squares, const(6 + SIGNERS_WINDOW_LOG)) - st = StackBuf(2) - st[0] = state_0 - st[1] = state_1 + st = state for j in unroll(0, SIGNERS_WINDOW): - pair = keys_ptr * (GEN ** (4 * j)) - hint_witness(pair[0:4], "pubkeys") - out = StackBuf(2) + pair = keys_ptr * GEN ** (8 * j) + hint_witness(pair[0:8], "pubkeys") + out = StackBuf(4) if const(j + 1 == SIGNERS_WINDOW): - blake2s(pair[0:2], pair[2:4], out, cv=st, md=nxt) + blake2s(pair[0:4], pair[4:8], out, cv=st, md=[nxt, 0]) else: - blake2s(pair[0:2], pair[2:4], out, cv=st, md=base + const(64 * (j + 1))) + blake2s(pair[0:4], pair[4:8], out, cv=st, md=[base + const(64 * (j + 1)), 0]) st = out - return st[0], st[1], nxt - - -def keys_tail(state_0, state_1, base, keys_ptr, k: Const): - # The key pairs past the last whole window, all non-final, so every offset stays - # below the base's lowest set bit. - st = StackBuf(2) - st[0] = state_0 - st[1] = state_1 - for j in unroll(0, k): - pair = keys_ptr * (GEN ** (4 * j)) - hint_witness(pair[0:4], "pubkeys") - out = StackBuf(2) - blake2s(pair[0:2], pair[2:4], out, cv=st, md=base + const(64 * (j + 1))) + return st, nxt + + + + +def list_chunk(state: StackBuf(4), off, ptr, cover, marks, origin_g, limit_g, size: Const, kind: Const): + # `size` non-final blocks of a list hash, block j's counter `off + 64(j + 1)`, + # its contents by list kind (`tail_blocks`). + st = state + for j in unroll(0, size): + out = StackBuf(4) + md = [off + const(64 * (j + 1)), 0] + if const(kind == LIST_CHILD_KEYS): + two = StackBuf(2) + hint_witness(two, "child_index") + assert log(two[0]) < log(limit_g) + assert log(two[1]) < log(limit_g) + cover[origin_g * two[0]] = marks * GEN ** (2 * j) + cover[origin_g * two[1]] = marks * GEN ** (2 * j + 1) + key_a = ptr * two[0] ** 4 + key_b = ptr * two[1] ** 4 + blake2s(key_a[0:4], key_b[0:4], out, cv=st, md=md) + if const(kind == LIST_CHILD_SPHINCS): + off_hint = hint_witness("child_sphincs_index") + assert log(off_hint) < log(limit_g) + cover[origin_g * off_hint] = marks * GEN ** j + entry = ptr * off_hint ** 8 + blake2s(entry[0:4], entry[4:8], out, cv=st, md=md) + if const(kind // 3 == 0): + block = ptr * GEN ** (8 * j) + if const(kind == LIST_KEYS): + hint_witness(block[0:8], "pubkeys") + if const(kind == LIST_SPHINCS): + hint_witness(block[0:8], "sphincs_signers") + blake2s(block[0:4], block[4:8], out, cv=st, md=md) st = out - return st[0], st[1], keys_ptr * (GEN ** (4 * k)) + return st + + +def tail_blocks(state: StackBuf(4), base, ptr, cover, marks, origin_g, limit_g, tail_g, kind: Const): + # The blocks a list hash has left past its last whole window, fewer than a + # window: tail_g = g^k, absorbed in power-of-two chunks, largest first, one per + # set bit of k. A chunk starts at a multiple of twice its size, so every counter + # offset 64(start + j + 1) in it is still one XOR onto the running base, and one + # chunk per size serves every k. Returns the state, then the list pointer and the + # coverage marks stepped past the tail. + bits = StackBuf(SIGNERS_WINDOW_LOG) + hint_decompose_bits_exponent(bits, tail_g, SIGNERS_WINDOW_LOG) + states = StackBuf(4 * (SIGNERS_WINDOW_LOG + 1)) + states[0:4] = state + rebuilt = GEN ** 0 + off = base + for i in unroll(0, SIGNERS_WINDOW_LOG): + b = SIGNERS_WINDOW_LOG - 1 - i + bit = bits[b] + bits[b] = bit * bit # booleanity, as a write-once pin + rebuilt *= 1 + bit * (GEN ** (2 ** b) + 1) + cur = 4 * i + if bit != 0: + states[cur + 4:cur + 8] = list_chunk(states[cur:cur + 4], off, ptr, cover, marks, origin_g, limit_g, 2 ** b, kind) + else: + states[cur + 4:cur + 8] = states[cur:cur + 4] + off = off + bit * const(64 * 2 ** b) + if const(kind // 3 == 0): + ptr = ptr * (1 + bit * (GEN ** (8 * 2 ** b) + 1)) + if const(kind == LIST_CHILD_KEYS): + marks = marks * (1 + bit * (GEN ** (2 * 2 ** b) + 1)) + if const(kind == LIST_CHILD_SPHINCS): + marks = marks * (1 + bit * (GEN ** (2 ** b) + 1)) + assert rebuilt == tail_g + last = 4 * SIGNERS_WINDOW_LOG + return states[last:last + 4], ptr, marks def key_list_digest(keys_ptr, half_g, odd_g, n_keys_g, g_squares): @@ -2584,80 +2676,58 @@ def key_list_digest(keys_ptr, half_g, odd_g, n_keys_g, g_squares): assert log(tail) < SIGNERS_WINDOW assert log(windows) < SIGNERS_MAX_WINDOWS assert windows ** SIGNERS_WINDOW * tail == half_g * odd_g * INV_GEN - chain = HeapBuf((windows * GEN) ** 4) # state pair, base, first key of the window - chain[1] = BLAKE2S_IV_0 - chain[GEN] = BLAKE2S_IV_1 - chain[GEN ** 2] = 0 - chain[GEN ** 3] = keys_ptr + chain = HeapBuf((windows * GEN) ** 6) # state, base, first key of the window + chain[0:4] = [BLAKE2S_IV_0, BLAKE2S_IV_1, BLAKE2S_IV_2, BLAKE2S_IV_3] + chain[GEN ** 4] = 0 + chain[GEN ** 5] = keys_ptr for xq in mul_range(1, windows): - slot = chain * (xq ** 4) - s0, s1, nb = keys_window(slot[1], slot[GEN], slot[GEN ** 2], slot[GEN ** 3], xq, g_squares) - step = chain * ((xq * GEN) ** 4) - step[1] = s0 - step[GEN] = s1 - step[GEN ** 2] = nb - step[GEN ** 3] = slot[GEN ** 3] * (GEN ** (4 * SIGNERS_WINDOW)) - end = chain * (windows ** 4) - t0, t1, last = match(log(tail), range(0, SIGNERS_WINDOW), lambda k: keys_tail(end[1], end[GEN], end[GEN ** 2], end[GEN ** 3], k)) - final = scaled_log(n_keys_g, g_squares, 5) + MD_FINAL - digest = StackBuf(2) + slot = chain * xq ** 6 + st, nb = keys_window(slot[0:4], slot[GEN ** 4], slot[GEN ** 5], xq, g_squares) + step = slot * GEN ** 6 + step[0:4] = st + step[GEN ** 4] = nb + step[GEN ** 5] = slot[GEN ** 5] * GEN ** (8 * SIGNERS_WINDOW) + end = chain * windows ** 6 + t, last, unused = tail_blocks(end[0:4], end[GEN ** 4], end[GEN ** 5], 0, 0, 0, 0, tail, LIST_KEYS) + final = [scaled_log(n_keys_g, g_squares, 5), MD_FINAL] + digest = StackBuf(4) if odd_g == 1: - hint_witness(last[0:4], "pubkeys") - blake2s(last[0:2], last[2:4], digest, cv=[t0, t1], md=final) + hint_witness(last[0:8], "pubkeys") + blake2s(last[0:4], last[4:8], digest, cv=t, md=final) else: # The odd key out fills half its block, the rest being the zero bytes the # counter already accounts for. - hint_witness(last[0:2], "pubkeys") - blake2s(last[0:2], [0, 0], digest, cv=[t0, t1], md=final) - return digest[0], digest[1] + hint_witness(last[0:4], "pubkeys") + blake2s(last[0:4], [0, 0, 0, 0], digest, cv=t, md=final) + return digest -def child_keys_window(state_0, state_1, base, keys_ptr, cover, marks, origin_g, limit_g, x_q, g_squares): +def child_keys_window(state: StackBuf(4), base, keys_ptr, cover, marks, origin_g, limit_g, x_q, g_squares): # One window of a child's key hash, absorbed exactly as the child absorbed it, # but with both keys of a block read at hinted indices into THIS node's table and # marked in the coverage table. Each index is an offset into the parent group the # caller mapped this child group to, bounded by that group's size, so a child's # key can only ever land on an XMSS slot of the right epoch. nxt = scaled_log(x_q * GEN, g_squares, const(6 + SIGNERS_WINDOW_LOG)) - st = StackBuf(2) - st[0] = state_0 - st[1] = state_1 + st = state for j in unroll(0, SIGNERS_WINDOW): two = StackBuf(2) hint_witness(two, "child_index") assert log(two[0]) < log(limit_g) # precondition as in the raw loops assert log(two[1]) < log(limit_g) - cover[origin_g * two[0]] = marks * (GEN ** (2 * j)) - cover[origin_g * two[1]] = marks * (GEN ** (2 * j + 1)) - key_a = keys_ptr * (two[0] * two[0]) - key_b = keys_ptr * (two[1] * two[1]) - out = StackBuf(2) + cover[origin_g * two[0]] = marks * GEN ** (2 * j) + cover[origin_g * two[1]] = marks * GEN ** (2 * j + 1) + key_a = keys_ptr * two[0] ** 4 + key_b = keys_ptr * two[1] ** 4 + out = StackBuf(4) if const(j + 1 == SIGNERS_WINDOW): - blake2s(key_a[0:2], key_b[0:2], out, cv=st, md=nxt) + blake2s(key_a[0:4], key_b[0:4], out, cv=st, md=[nxt, 0]) else: - blake2s(key_a[0:2], key_b[0:2], out, cv=st, md=base + const(64 * (j + 1))) + blake2s(key_a[0:4], key_b[0:4], out, cv=st, md=[base + const(64 * (j + 1)), 0]) st = out - return st[0], st[1], nxt + return st, nxt -def child_keys_tail(state_0, state_1, base, keys_ptr, cover, marks, origin_g, limit_g, k: Const): - # The child's key pairs past its last whole window, all non-final. - st = StackBuf(2) - st[0] = state_0 - st[1] = state_1 - for j in unroll(0, k): - two = StackBuf(2) - hint_witness(two, "child_index") - assert log(two[0]) < log(limit_g) - assert log(two[1]) < log(limit_g) - cover[origin_g * two[0]] = marks * (GEN ** (2 * j)) - cover[origin_g * two[1]] = marks * (GEN ** (2 * j + 1)) - key_a = keys_ptr * (two[0] * two[0]) - key_b = keys_ptr * (two[1] * two[1]) - out = StackBuf(2) - blake2s(key_a[0:2], key_b[0:2], out, cv=st, md=base + const(64 * (j + 1))) - st = out - return st[0], st[1], marks * (GEN ** (2 * k)) def child_key_list_digest(keys_ptr, cover, base, origin_g, limit_g, half_g, odd_g, n_keys_g, g_squares): @@ -2672,23 +2742,21 @@ def child_key_list_digest(keys_ptr, cover, base, origin_g, limit_g, half_g, odd_ assert log(tail) < SIGNERS_WINDOW assert log(windows) < SIGNERS_MAX_WINDOWS assert windows ** SIGNERS_WINDOW * tail == half_g * odd_g * INV_GEN - chain = HeapBuf((windows * GEN) ** 4) - chain[1] = BLAKE2S_IV_0 - chain[GEN] = BLAKE2S_IV_1 - chain[GEN ** 2] = 0 - chain[GEN ** 3] = base + chain = HeapBuf((windows * GEN) ** 6) # state, base, coverage marks + chain[0:4] = [BLAKE2S_IV_0, BLAKE2S_IV_1, BLAKE2S_IV_2, BLAKE2S_IV_3] + chain[GEN ** 4] = 0 + chain[GEN ** 5] = base for xq in mul_range(1, windows): - slot = chain * (xq ** 4) - s0, s1, nb = child_keys_window(slot[1], slot[GEN], slot[GEN ** 2], keys_ptr, cover, slot[GEN ** 3], origin_g, limit_g, xq, g_squares) - step = chain * ((xq * GEN) ** 4) - step[1] = s0 - step[GEN] = s1 - step[GEN ** 2] = nb - step[GEN ** 3] = slot[GEN ** 3] * (GEN ** (2 * SIGNERS_WINDOW)) - end = chain * (windows ** 4) - t0, t1, marks = match(log(tail), range(0, SIGNERS_WINDOW), lambda k: child_keys_tail(end[1], end[GEN], end[GEN ** 2], keys_ptr, cover, end[GEN ** 3], origin_g, limit_g, k)) - final = scaled_log(n_keys_g, g_squares, 5) + MD_FINAL - digest = StackBuf(2) + slot = chain * xq ** 6 + st, nb = child_keys_window(slot[0:4], slot[GEN ** 4], keys_ptr, cover, slot[GEN ** 5], origin_g, limit_g, xq, g_squares) + step = slot * GEN ** 6 + step[0:4] = st + step[GEN ** 4] = nb + step[GEN ** 5] = slot[GEN ** 5] * GEN ** (2 * SIGNERS_WINDOW) + end = chain * windows ** 6 + t, unused, marks = tail_blocks(end[0:4], end[GEN ** 4], keys_ptr, cover, end[GEN ** 5], origin_g, limit_g, tail, LIST_CHILD_KEYS) + final = [scaled_log(n_keys_g, g_squares, 5), MD_FINAL] + digest = StackBuf(4) if odd_g == 1: two = StackBuf(2) hint_witness(two, "child_index") @@ -2696,24 +2764,24 @@ def child_key_list_digest(keys_ptr, cover, base, origin_g, limit_g, half_g, odd_ assert log(two[1]) < log(limit_g) cover[origin_g * two[0]] = marks cover[origin_g * two[1]] = marks * GEN - key_a = keys_ptr * (two[0] * two[0]) - key_b = keys_ptr * (two[1] * two[1]) - blake2s(key_a[0:2], key_b[0:2], digest, cv=[t0, t1], md=final) + key_a = keys_ptr * two[0] ** 4 + key_b = keys_ptr * two[1] ** 4 + blake2s(key_a[0:4], key_b[0:4], digest, cv=t, md=final) else: tail_idx = hint_witness("child_index") assert log(tail_idx) < log(limit_g) cover[origin_g * tail_idx] = marks - key_last = keys_ptr * (tail_idx * tail_idx) - blake2s(key_last[0:2], [0, 0], digest, cv=[t0, t1], md=final) - return digest[0], digest[1] + key_last = keys_ptr * tail_idx ** 4 + blake2s(key_last[0:4], [0, 0, 0, 0], digest, cv=t, md=final) + return digest def scaled_log(x, g_squares, shift: Const): # 2^shift times the exponent of `x`, as a bit pattern (doc §sec:prog-byte-counter). # The exponent's bits are advice, tied back by the g-power product; weighing them - # at COORD_BASIS[j] assembles the exponent itself and the final multiply is the - # shift, exact because nothing reduces below degree 64. Both sides of the product - # stay under the order of g, so the bits ARE that exponent, hence below + # at 2^j assembles the exponent itself and the final multiply is the shift, + # exact because nothing reduces below degree 64. Both sides of the product stay + # under the order of g, so the bits ARE that exponent, hence below # 2^SIGNERS_COUNT_BITS, which every count and window index here is. bits = StackBuf(SIGNERS_COUNT_BITS) hint_decompose_bits_exponent(bits, x, SIGNERS_COUNT_BITS) @@ -2722,47 +2790,32 @@ def scaled_log(x, g_squares, shift: Const): for j in unroll(0, SIGNERS_COUNT_BITS): b = bits[j] bits[j] = b * b # booleanity, as a write-once pin - value += b * COORD_BASIS[j] + value += b * 2 ** j rebuilt *= (1 + b * (g_squares[GEN ** j] + 1)) assert rebuilt == x - return value * COORD_BASIS[shift] + return value * 2 ** shift -def sphincs_window(state_0, state_1, base, entries_ptr, x_q, g_squares): +def sphincs_window(state: StackBuf(4), base, entries_ptr, x_q, g_squares): # One window of the SPHINCS list's hash: SIGNERS_WINDOW claims, one 64-byte block # each (the claimed key, then the message it signed). Block j's counter is # base + 64(j+1), one XOR, except the last, whose offset is the base's own lowest # bit and which therefore takes the NEXT window's base as its whole counter. That # base is derived here and carried out for the following window. nxt = scaled_log(x_q * GEN, g_squares, const(6 + SIGNERS_WINDOW_LOG)) - st = StackBuf(2) - st[0] = state_0 - st[1] = state_1 + st = state for j in unroll(0, SIGNERS_WINDOW): - entry = entries_ptr * (GEN ** (4 * j)) - hint_witness(entry[0:4], "sphincs_signers") - out = StackBuf(2) + entry = entries_ptr * GEN ** (8 * j) + hint_witness(entry[0:8], "sphincs_signers") + out = StackBuf(4) if const(j + 1 == SIGNERS_WINDOW): - blake2s(entry[0:2], entry[2:4], out, cv=st, md=nxt) + blake2s(entry[0:4], entry[4:8], out, cv=st, md=[nxt, 0]) else: - blake2s(entry[0:2], entry[2:4], out, cv=st, md=base + const(64 * (j + 1))) - st = out - return st[0], st[1], nxt - - -def sphincs_tail(state_0, state_1, base, entries_ptr, k: Const): - # The blocks the window loop leaves over, fewer than a window, so every offset - # 64(j+1) stays below the base's lowest set bit and needs no next base. - st = StackBuf(2) - st[0] = state_0 - st[1] = state_1 - for j in unroll(0, k): - entry = entries_ptr * (GEN ** (4 * j)) - hint_witness(entry[0:4], "sphincs_signers") - out = StackBuf(2) - blake2s(entry[0:2], entry[2:4], out, cv=st, md=base + const(64 * (j + 1))) + blake2s(entry[0:4], entry[4:8], out, cv=st, md=[base + const(64 * (j + 1)), 0]) st = out - return st[0], st[1], entries_ptr * (GEN ** (4 * k)) + return st, nxt + + def sphincs_list_digest(entries_ptr, n_g, g_squares): @@ -2770,10 +2823,10 @@ def sphincs_list_digest(entries_ptr, n_g, g_squares): # over exactly 64n bytes and no block is partial. The last block is absorbed # apart, carrying the total length as its counter and the final-block flag; the # n - 1 before it run in windows plus a tail (doc §sec:prog-byte-counter). - digest = StackBuf(2) + digest = StackBuf(4) if n_g == 1: # No claims: the hash of the empty string, one compression of a zero block. - blake2s([0, 0], [0, 0], digest, md=MD_FINAL) + blake2s([0, 0, 0, 0], [0, 0, 0, 0], digest, md=[0, MD_FINAL]) else: split = StackBuf(2) hint_witness(split, "signers_split") # g^windows, g^tail_blocks @@ -2782,67 +2835,48 @@ def sphincs_list_digest(entries_ptr, n_g, g_squares): assert log(tail) < SIGNERS_WINDOW assert log(windows) < SIGNERS_MAX_WINDOWS assert windows ** SIGNERS_WINDOW * tail == n_g * INV_GEN - # Four cells a window: the state pair, the window's base, its first entry. - chain = HeapBuf((windows * GEN) ** 4) - chain[1] = BLAKE2S_IV_0 - chain[GEN] = BLAKE2S_IV_1 - chain[GEN ** 2] = 0 - chain[GEN ** 3] = entries_ptr + # Six words a window: the state, the window's base, its first entry. + chain = HeapBuf((windows * GEN) ** 6) + chain[0:4] = [BLAKE2S_IV_0, BLAKE2S_IV_1, BLAKE2S_IV_2, BLAKE2S_IV_3] + chain[GEN ** 4] = 0 + chain[GEN ** 5] = entries_ptr for xq in mul_range(1, windows): - slot = chain * (xq ** 4) - s0, s1, nb = sphincs_window(slot[1], slot[GEN], slot[GEN ** 2], slot[GEN ** 3], xq, g_squares) - step = chain * ((xq * GEN) ** 4) - step[1] = s0 - step[GEN] = s1 - step[GEN ** 2] = nb - step[GEN ** 3] = slot[GEN ** 3] * (GEN ** (4 * SIGNERS_WINDOW)) - end = chain * (windows ** 4) - t0, t1, last = match(log(tail), range(0, SIGNERS_WINDOW), lambda k: sphincs_tail(end[1], end[GEN], end[GEN ** 2], end[GEN ** 3], k)) - hint_witness(last[0:4], "sphincs_signers") - final = scaled_log(n_g, g_squares, 6) + MD_FINAL - blake2s(last[0:2], last[2:4], digest, cv=[t0, t1], md=final) - return digest[0], digest[1] - - -def child_sphincs_window(state_0, state_1, base, entries_ptr, cover, marks, origin_g, limit_g, x_q, g_squares): + slot = chain * xq ** 6 + st, nb = sphincs_window(slot[0:4], slot[GEN ** 4], slot[GEN ** 5], xq, g_squares) + step = slot * GEN ** 6 + step[0:4] = st + step[GEN ** 4] = nb + step[GEN ** 5] = slot[GEN ** 5] * GEN ** (8 * SIGNERS_WINDOW) + end = chain * windows ** 6 + t, last, unused = tail_blocks(end[0:4], end[GEN ** 4], end[GEN ** 5], 0, 0, 0, 0, tail, LIST_SPHINCS) + hint_witness(last[0:8], "sphincs_signers") + final = [scaled_log(n_g, g_squares, 6), MD_FINAL] + blake2s(last[0:4], last[4:8], digest, cv=t, md=final) + return digest + + +def child_sphincs_window(state: StackBuf(4), base, entries_ptr, cover, marks, origin_g, limit_g, x_q, g_squares): # One window of a child's SPHINCS list, absorbed exactly as the child absorbed # it, but with each block's claim read at a hinted index into THIS node's table # and marked in the coverage table. The index is an offset into the SPHINCS # region and bounded by that region's size, so a child's claim can only ever # land on a SPHINCS slot. Counters as in `sphincs_window`. nxt = scaled_log(x_q * GEN, g_squares, const(6 + SIGNERS_WINDOW_LOG)) - st = StackBuf(2) - st[0] = state_0 - st[1] = state_1 + st = state for j in unroll(0, SIGNERS_WINDOW): off_hint = hint_witness("child_sphincs_index") assert log(off_hint) < log(limit_g) # precondition as in the raw loops - cover[origin_g * off_hint] = marks * (GEN ** j) - entry = entries_ptr * (off_hint ** 4) - out = StackBuf(2) + cover[origin_g * off_hint] = marks * GEN ** j + entry = entries_ptr * off_hint ** 8 + out = StackBuf(4) if const(j + 1 == SIGNERS_WINDOW): - blake2s(entry[0:2], entry[2:4], out, cv=st, md=nxt) + blake2s(entry[0:4], entry[4:8], out, cv=st, md=[nxt, 0]) else: - blake2s(entry[0:2], entry[2:4], out, cv=st, md=base + const(64 * (j + 1))) + blake2s(entry[0:4], entry[4:8], out, cv=st, md=[base + const(64 * (j + 1)), 0]) st = out - return st[0], st[1], nxt + return st, nxt -def child_sphincs_tail(state_0, state_1, base, entries_ptr, cover, marks, origin_g, limit_g, k: Const): - # The blocks past the child's last whole window, all of them non-final, so every - # offset stays below this base's lowest set bit. - st = StackBuf(2) - st[0] = state_0 - st[1] = state_1 - for j in unroll(0, k): - off_hint = hint_witness("child_sphincs_index") - assert log(off_hint) < log(limit_g) - cover[origin_g * off_hint] = marks * (GEN ** j) - entry = entries_ptr * (off_hint ** 4) - out = StackBuf(2) - blake2s(entry[0:2], entry[2:4], out, cv=st, md=base + const(64 * (j + 1))) - st = out - return st[0], st[1] def child_sphincs_list_digest(entries_ptr, cover, base, origin_g, limit_g, n_g, g_squares): @@ -2850,9 +2884,9 @@ def child_sphincs_list_digest(entries_ptr, cover, base, origin_g, limit_g, n_g, # child hashed (`sphincs_list_digest`), so the digest it rebuilds is the one the # child's statement carries. `base` prefixes the coverage write values, which # count the claims off as they are marked. - digest = StackBuf(2) + digest = StackBuf(4) if n_g == 1: - blake2s([0, 0], [0, 0], digest, md=MD_FINAL) + blake2s([0, 0, 0, 0], [0, 0, 0, 0], digest, md=[0, MD_FINAL]) else: split = StackBuf(2) hint_witness(split, "signers_split") @@ -2861,59 +2895,44 @@ def child_sphincs_list_digest(entries_ptr, cover, base, origin_g, limit_g, n_g, assert log(tail) < SIGNERS_WINDOW assert log(windows) < SIGNERS_MAX_WINDOWS assert windows ** SIGNERS_WINDOW * tail == n_g * INV_GEN - chain = HeapBuf((windows * GEN) ** 4) - chain[1] = BLAKE2S_IV_0 - chain[GEN] = BLAKE2S_IV_1 - chain[GEN ** 2] = 0 + chain = HeapBuf((windows * GEN) ** 5) # state, base + chain[0:4] = [BLAKE2S_IV_0, BLAKE2S_IV_1, BLAKE2S_IV_2, BLAKE2S_IV_3] + chain[GEN ** 4] = 0 for xq in mul_range(1, windows): - slot = chain * (xq ** 4) - marks = base * (xq ** SIGNERS_WINDOW) - s0, s1, nb = child_sphincs_window(slot[1], slot[GEN], slot[GEN ** 2], entries_ptr, cover, marks, origin_g, limit_g, xq, g_squares) - step = chain * ((xq * GEN) ** 4) - step[1] = s0 - step[GEN] = s1 - step[GEN ** 2] = nb - end = chain * (windows ** 4) - marks = base * (windows ** SIGNERS_WINDOW) - t0, t1 = match(log(tail), range(0, SIGNERS_WINDOW), lambda k: child_sphincs_tail(end[1], end[GEN], end[GEN ** 2], entries_ptr, cover, marks, origin_g, limit_g, k)) + slot = chain * xq ** 5 + marks = base * xq ** SIGNERS_WINDOW + st, nb = child_sphincs_window(slot[0:4], slot[GEN ** 4], entries_ptr, cover, marks, origin_g, limit_g, xq, g_squares) + step = slot * GEN ** 5 + step[0:4] = st + step[GEN ** 4] = nb + end = chain * windows ** 5 + marks = base * windows ** SIGNERS_WINDOW + t, unused, unused_marks = tail_blocks(end[0:4], end[GEN ** 4], entries_ptr, cover, marks, origin_g, limit_g, tail, LIST_CHILD_SPHINCS) off_hint = hint_witness("child_sphincs_index") assert log(off_hint) < log(limit_g) cover[origin_g * off_hint] = base * (n_g * INV_GEN) - entry = entries_ptr * (off_hint ** 4) - final = scaled_log(n_g, g_squares, 6) + MD_FINAL - blake2s(entry[0:2], entry[2:4], digest, cv=[t0, t1], md=final) - return digest[0], digest[1] + entry = entries_ptr * off_hint ** 8 + final = [scaled_log(n_g, g_squares, 6), MD_FINAL] + blake2s(entry[0:4], entry[4:8], digest, cv=t, md=final) + return digest -def plain_window(state_0, state_1, base, run_ptr, x_q, g_squares): - # One window over a run of cells already in memory, four to a block, hinting +def plain_window(state: StackBuf(4), base, run_ptr, x_q, g_squares): + # One window over a run of words already in memory, eight to a block, hinting # nothing. Counters as in `sphincs_window`. nxt = scaled_log(x_q * GEN, g_squares, const(6 + SIGNERS_WINDOW_LOG)) - st = StackBuf(2) - st[0] = state_0 - st[1] = state_1 + st = state for j in unroll(0, SIGNERS_WINDOW): - block = run_ptr * (GEN ** (4 * j)) - out = StackBuf(2) + block = run_ptr * GEN ** (8 * j) + out = StackBuf(4) if const(j + 1 == SIGNERS_WINDOW): - blake2s(block[0:2], block[2:4], out, cv=st, md=nxt) + blake2s(block[0:4], block[4:8], out, cv=st, md=[nxt, 0]) else: - blake2s(block[0:2], block[2:4], out, cv=st, md=base + const(64 * (j + 1))) + blake2s(block[0:4], block[4:8], out, cv=st, md=[base + const(64 * (j + 1)), 0]) st = out - return st[0], st[1], nxt - - -def plain_tail(state_0, state_1, base, run_ptr, k: Const): - # The blocks of the run past its last whole window, all non-final. - st = StackBuf(2) - st[0] = state_0 - st[1] = state_1 - for j in unroll(0, k): - block = run_ptr * (GEN ** (4 * j)) - out = StackBuf(2) - blake2s(block[0:2], block[2:4], out, cv=st, md=base + const(64 * (j + 1))) - st = out - return st[0], st[1], run_ptr * (GEN ** (4 * k)) + return st, nxt + + def signer_set_digest(run_ptr, n_epochs_g, g_squares): @@ -2930,25 +2949,23 @@ def signer_set_digest(run_ptr, n_epochs_g, g_squares): assert log(tail) < SIGNERS_WINDOW assert log(windows) < SIGNERS_MAX_WINDOWS assert windows ** SIGNERS_WINDOW * tail == blocks * INV_GEN - chain = HeapBuf((windows * GEN) ** 4) - chain[1] = BLAKE2S_IV_0 - chain[GEN] = BLAKE2S_IV_1 - chain[GEN ** 2] = 0 - chain[GEN ** 3] = run_ptr + chain = HeapBuf((windows * GEN) ** 6) + chain[0:4] = [BLAKE2S_IV_0, BLAKE2S_IV_1, BLAKE2S_IV_2, BLAKE2S_IV_3] + chain[GEN ** 4] = 0 + chain[GEN ** 5] = run_ptr for xq in mul_range(1, windows): - slot = chain * (xq ** 4) - s0, s1, nb = plain_window(slot[1], slot[GEN], slot[GEN ** 2], slot[GEN ** 3], xq, g_squares) - step = chain * ((xq * GEN) ** 4) - step[1] = s0 - step[GEN] = s1 - step[GEN ** 2] = nb - step[GEN ** 3] = slot[GEN ** 3] * (GEN ** (4 * SIGNERS_WINDOW)) - end = chain * (windows ** 4) - t0, t1, last = match(log(tail), range(0, SIGNERS_WINDOW), lambda k: plain_tail(end[1], end[GEN], end[GEN ** 2], end[GEN ** 3], k)) - final = scaled_log(blocks, g_squares, 6) + MD_FINAL - digest = StackBuf(2) - blake2s(last[0:2], last[2:4], digest, cv=[t0, t1], md=final) - return digest[0], digest[1] + slot = chain * xq ** 6 + st, nb = plain_window(slot[0:4], slot[GEN ** 4], slot[GEN ** 5], xq, g_squares) + step = slot * GEN ** 6 + step[0:4] = st + step[GEN ** 4] = nb + step[GEN ** 5] = slot[GEN ** 5] * GEN ** (8 * SIGNERS_WINDOW) + end = chain * windows ** 6 + t, last, unused = tail_blocks(end[0:4], end[GEN ** 4], end[GEN ** 5], 0, 0, 0, 0, tail, LIST_PLAIN) + final = [scaled_log(blocks, g_squares, 6), MD_FINAL] + digest = StackBuf(4) + blake2s(last[0:4], last[4:8], digest, cv=t, md=final) + return digest def rebuild_child_groups(nsub_e_g, run_ptr, base, epochs, msgs, group_base, group_slots, n_epochs_g, xmss_table, cover, g_squares): @@ -2963,16 +2980,15 @@ def rebuild_child_groups(nsub_e_g, run_ptr, base, epochs, msgs, group_base, grou counts = HeapBuf(nsub_e_g * GEN) counts[GEN ** 0] = 1 for xj in mul_range(1, nsub_e_g): - grp = StackBuf(4) - hint_witness(grp, "child_group") # epoch, msg_lo, msg_hi, count - n_keys = grp[3] + grp = StackBuf(6) + hint_witness(grp, "child_group") # epoch, message (four words), count + n_keys = grp[5] assert log(n_keys) < MAX_KEYS parent = hint_witness("child_group_map") assert log(parent) < log(n_epochs_g) assert epochs[parent] == grp[0] - parent_msg = msgs * (parent * parent) - assert parent_msg[1] == grp[1] - assert parent_msg[GEN] == grp[2] + parent_msg = msgs * parent ** 4 + parent_msg[0:4] = grp[1:5] halves = StackBuf(2) hint_witness(halves, "child_halves") assert log(halves[1]) < 2 @@ -2980,16 +2996,15 @@ def rebuild_child_groups(nsub_e_g, run_ptr, base, epochs, msgs, group_base, grou assert halves[0] * halves[0] * halves[1] == n_keys gb = group_base[parent] prefix = counts[xj] - kd_0, kd_1 = child_key_list_digest(xmss_table * (gb * gb), cover, base * prefix, gb, group_slots[parent], halves[0], halves[1], n_keys, g_squares) - slot = run_ptr * (xj ** 8) * (GEN ** 4) + kd = child_key_list_digest(xmss_table * gb ** 4, cover, base * prefix, gb, group_slots[parent], halves[0], halves[1], n_keys, g_squares) + slot = run_ptr * xj ** 16 * GEN ** 8 slot[1] = grp[0] - slot[GEN] = n_keys - slot[GEN ** 2] = grp[1] - slot[GEN ** 3] = grp[2] - slot[GEN ** 4] = kd_0 - slot[GEN ** 5] = kd_1 - slot[GEN ** 6] = 0 - slot[GEN ** 7] = 0 + slot[GEN] = 0 + slot[GEN ** 2] = n_keys + slot[GEN ** 3] = 0 + slot[4:8] = grp[1:5] + slot[8:12] = kd + slot[12:16] = [0, 0, 0, 0] counts[xj * GEN] = prefix * n_keys return counts[nsub_e_g] @@ -3000,7 +3015,7 @@ def rebuild_child_groups(nsub_e_g, run_ptr, base, epochs, msgs, group_base, grou def da_verify(g_squares): - # Returns the matrix root and membership-vector hash, two canonical cells each. + # Returns the matrix root and membership-vector hash, four words each. # The row count is a run-time parameter. The trees need a compile-time depth, # so the rows are padded to a power of two and the two `match`es below dispatch # on its log; everything else walks the real rows only. A padding row is the @@ -3012,30 +3027,29 @@ def da_verify(g_squares): g_log_pad = shape[1] n_pad_g, gap_g = da_row_shape(n_blob_g, g_log_pad, g_squares) - prefix_digests = HeapBuf(n_pad_g ** (2 * DA_PREFIX_CELLS)) + prefix_digests = HeapBuf(n_pad_g ** (4 * DA_PREFIX_CELLS)) prefix_bases = HeapBuf(n_blob_g) for xi in mul_range(1, n_blob_g): - prefix_bases[xi] = prefix_digests * (xi ** (2 * DA_PREFIX_CELLS)) + prefix_bases[xi] = prefix_digests * xi ** (4 * DA_PREFIX_CELLS) - hashes = HeapBuf(2 * (DA_CELLS + 1)) - hashes[1] = BLAKE2S_IV_0 - hashes[GEN] = BLAKE2S_IV_1 + hashes = HeapBuf(4 * (DA_CELLS + 1)) + hashes[0:4] = [BLAKE2S_IV_0, BLAKE2S_IV_1, BLAKE2S_IV_2, BLAKE2S_IV_3] counter_values = StackBuf(3 * DA_CELLS + 1) for w in unroll(0, 3 * DA_CELLS + 1): counter_values[w] = const(w * 2 ** (DA_LOG_CELL + 3)) counters = addr(counter_values) - # Each row's running sum uses a fresh write-once cell per column. - acc = HeapBuf(n_blob_g ** (DA_CELLS + 1)) + # Each row's running sum uses a fresh write-once element per column. + acc = HeapBuf(n_blob_g ** (3 * (DA_CELLS + 1))) rowbase = HeapBuf(n_blob_g) for xi in mul_range(1, n_blob_g): - chain = acc * (xi ** (DA_CELLS + 1)) + chain = acc * xi ** (3 * (DA_CELLS + 1)) rowbase[xi] = chain - chain[1] = 0 + chain[0:3] = ZERO # Prefix blocks first, so the digests the row branch needs are stored by a # loop that knows it is inside the prefix, with no per-block branch. - coltree = HeapBuf(4 * DA_CELLS) + coltree = HeapBuf(8 * DA_CELLS) for xb in mul_range(1, GEN ** DA_PREFIX_CELLS): da_verify_column(xb, n_blob_g, n_pad_g, gap_g, g_log_pad, prefix_bases, coltree, rowbase, hashes, counters, 1) for xb in mul_range(GEN ** DA_PREFIX_CELLS, GEN ** DA_CELLS): @@ -3043,31 +3057,31 @@ def da_verify(g_squares): for xi in mul_range(1, n_blob_g): chain = rowbase[xi] - assert chain[GEN ** DA_CELLS] == 0 - rc0, rc1 = da_levels(coltree, DA_CELLS, DA_BLOCK_BITS) + k = 3 * DA_CELLS + assert_eq192(chain[k:k + 3], ZERO) + col_root = da_levels(coltree, DA_CELLS, DA_BLOCK_BITS) # The row branch: each row's prefix digests, hashed as one string. - rowtree = HeapBuf(n_pad_g ** 4) + rowtree = HeapBuf(n_pad_g ** 8) for xi in mul_range(1, n_blob_g): - run = prefix_bases[xi] - st = StackBuf(2) - blake2s(run[0:2], run[2:4], st, counter=64, final=1 // DA_ROW_BLOCKS) + prefix = prefix_bases[xi] + st = StackBuf(4) + blake2s(prefix[0:4], prefix[4:8], st, counter=64, final=1 // DA_ROW_BLOCKS) for b in unroll(1, DA_ROW_BLOCKS): - nxt = StackBuf(2) - blake2s(run[4 * b:4 * b + 2], run[4 * b + 2:4 * b + 4], nxt, cv=st, counter=64 * (b + 1), final=(b + 1) // DA_ROW_BLOCKS) + nxt = StackBuf(4) + blake2s(prefix[8 * b:8 * b + 4], prefix[8 * b + 4:8 * b + 8], nxt, cv=st, counter=64 * (b + 1), final=(b + 1) // DA_ROW_BLOCKS) st = nxt - slot = rowtree * (xi ** 2) - slot[1] = st[0] - slot[GEN] = st[1] + slot = rowtree * xi ** 4 + slot[0:4] = st for xd in mul_range(1, gap_g): - pad = rowtree * (n_blob_g ** 2) * (xd ** 2) - pad[1] = DA_PAD_ROW_0 - pad[GEN] = DA_PAD_ROW_1 - rr0, rr1 = match(log(g_log_pad), range(0, DA_TREE_ARMS), lambda k: da_levels(rowtree, 2 ** k, k)) + pad = rowtree * n_blob_g ** 4 * xd ** 4 + pad[0:4] = [DA_PAD_ROW[0], DA_PAD_ROW[1], DA_PAD_ROW[2], DA_PAD_ROW[3]] + row_root = match(log(g_log_pad), range(0, DA_TREE_ARMS), lambda k: da_levels(rowtree, 2 ** k, k)) - root = StackBuf(2) - blake2s([rr0, rr1], [rc0, rc1], root) - return root[0], root[1], hashes[GEN ** (2 * DA_CELLS)], hashes[GEN ** (2 * DA_CELLS + 1)] + root = StackBuf(4) + blake2s(row_root, col_root, root) + vec = 4 * DA_CELLS + return root, hashes[vec:vec + 4] def da_row_shape(n_blob_g, g_log_pad, g_squares): @@ -3083,151 +3097,135 @@ def da_row_shape(n_blob_g, g_log_pad, g_squares): def da_verify_column(xb, n_blob_g, n_pad_g, gap_g, g_log_pad, prefix_bases, coltree, rowbase, hashes, counters, store: Const): - st = hashes * (xb ** 2) - h0, h1, lvals = da_vector_dispatch(xb, st[1], st[GEN], counters) - nxt = hashes * ((xb * GEN) ** 2) - nxt[1] = h0 - nxt[GEN] = h1 - node = HeapBuf(n_pad_g ** 4) + st = hashes * xb ** 4 + h, lvals = da_vector_dispatch(xb, st[0:4], counters) + nxt = st * GEN ** 4 + nxt[0:4] = h + node = HeapBuf(n_pad_g ** 8) + xb3 = xb ** 3 for xi in mul_range(1, n_blob_g): - d0, d1, s = da_verify_cell(lvals) - chain = rowbase[xi] * xb - chain[GEN] = chain[1] + s - leaf = node * (xi ** 2) - leaf[1] = d0 - leaf[GEN] = d1 + d, s = da_verify_cell(lvals) + chain = rowbase[xi] * xb3 + chain[3:6] = add192(chain[0:3], s) + leaf = node * xi ** 4 + leaf[0:4] = d if const(store == 1): - slot = prefix_bases[xi] * (xb * xb) - slot[1] = d0 - slot[GEN] = d1 + slot = prefix_bases[xi] * xb ** 4 + slot[0:4] = d for xd in mul_range(1, gap_g): - pad = node * (n_blob_g ** 2) * (xd ** 2) - pad[1] = DA_PAD_CELL_0 - pad[GEN] = DA_PAD_CELL_1 - c0, c1 = match(log(g_log_pad), range(0, DA_TREE_ARMS), lambda k: da_levels(node, 2 ** k, k)) - out = coltree * (xb * xb) - out[1] = c0 - out[GEN] = c1 + pad = node * n_blob_g ** 4 * xd ** 4 + pad[0:4] = [DA_PAD_CELL[0], DA_PAD_CELL[1], DA_PAD_CELL[2], DA_PAD_CELL[3]] + c = match(log(g_log_pad), range(0, DA_TREE_ARMS), lambda k: da_levels(node, 2 ** k, k)) + out = coltree * xb ** 4 + out[0:4] = c return def da_verify_cell(lvals): - # Packing checks that each symbol used by both the hash and dot product is in K. + # The symbols are words, so the hash and the dot product read them directly; the + # weights are elements, three words each, so the dot product runs limb by limb. sym = StackBuf(DA_CELL) hint_witness(sym, "da_symbols") - packed = StackBuf(DA_CELL // 2) - s = 0 - for e in unroll(0, DA_CELL // 2): - packed[e] = pack64x2(sym[2 * e], sym[2 * e + 1]) - s = s + lvals[GEN ** (2 * e)] * sym[2 * e] - s = s + lvals[GEN ** (2 * e + 1)] * sym[2 * e + 1] - st = StackBuf(2) - blake2s(packed[0:2], packed[2:4], st, counter=64, final=1 // DA_CELL_BLOCKS) + s0 = 0 + s1 = 0 + s2 = 0 + for e in unroll(0, DA_CELL): + k = 3 * e + sv = sym[e] + s0 += sv * lvals[GEN ** k] + s1 += sv * lvals[GEN ** (k + 1)] + s2 += sv * lvals[GEN ** (k + 2)] + st = StackBuf(4) + blake2s(sym[0:4], sym[4:8], st, counter=64, final=1 // DA_CELL_BLOCKS) for b in unroll(1, DA_CELL_BLOCKS): - nxt = StackBuf(2) - blake2s(packed[4 * b:4 * b + 2], packed[4 * b + 2:4 * b + 4], nxt, cv=st, counter=64 * (b + 1), final=(b + 1) // DA_CELL_BLOCKS) + nxt = StackBuf(4) + blake2s(sym[8 * b:8 * b + 4], sym[8 * b + 4:8 * b + 8], nxt, cv=st, counter=64 * (b + 1), final=(b + 1) // DA_CELL_BLOCKS) st = nxt - return st[0], st[1], s + return st, [s0, s1, s2] def da_levels(tree, n: Const, log_n: Const): # The internal levels of a Merkle tree whose leaves already sit at level 0 of - # `tree` (2 cells a node, level lvl at 4n - 4n//2**lvl). Returns the root pair. + # `tree` (4 words a node, level lvl at 8n - 8n//2**lvl). Returns the root. for lvl in unroll(0, log_n): for xp in mul_range(1, GEN ** (n // 2 ** (lvl + 1))): - a = tree * (GEN ** (4 * n - 4 * n // 2 ** lvl)) * (xp ** 4) - b = tree * (GEN ** (4 * n - 4 * n // 2 ** (lvl + 1))) * (xp * xp) - blake2s(a[0:2], a[2:4], b[0:2]) - return tree[GEN ** (4 * n - 4)], tree[GEN ** (4 * n - 3)] + a = tree * GEN ** (8 * n - 8 * n // 2 ** lvl) * xp ** 8 + b = tree * GEN ** (8 * n - 8 * n // 2 ** (lvl + 1)) * xp ** 4 + blake2s(a[0:4], a[4:8], b[0:4]) + root = 8 * n - 8 + return tree[root:root + 4] -def da_vector_dispatch(xb, h0, h1, counters): - out = StackBuf(3) +def da_vector_dispatch(xb, h: StackBuf(4), counters): + out = StackBuf(4) + weights = StackBuf(1) if xb == GEN ** (DA_CELLS - 1): - a, b, weights = da_vector_cell(xb, h0, h1, counters, 1) - out[0] = a - out[1] = b - out[2] = weights + a, w = da_vector_cell(xb, h, counters, 1) + out[0:4] = a + weights[0] = w else: - a, b, weights = da_vector_cell(xb, h0, h1, counters, 0) - out[0] = a - out[1] = b - out[2] = weights - return out[0], out[1], out[2] - - -def da_vector_cell(xb, h0, h1, counters, final: Const): - weights = StackBuf(DA_CELL) - hint_witness(weights[0:DA_CELL], "da_weights") - packed = StackBuf(3 * DA_CELL // 2) - # Two extension-field entries become three canonical 128-bit cells, without padding. - for p in unroll(0, DA_CELL // 2): - s = weights[2 * p] - t = weights[2 * p + 1] - slo = StackBuf(2) - tlo = StackBuf(2) - hint_f192_limbs(slo, s) - hint_f192_limbs(tlo, t) - packed[3 * p] = pack64x2(slo[0], slo[1]) - packed[3 * p + 1] = pack64x2(((s + slo[0]) * Y_INV + slo[1]) * Y_INV, tlo[0]) - packed[3 * p + 2] = pack64x2(tlo[1], ((t + tlo[0]) * Y_INV + tlo[1]) * Y_INV) - st = StackBuf(2) - st[0] = h0 - st[1] = h1 + a, w = da_vector_cell(xb, h, counters, 0) + out[0:4] = a + weights[0] = w + return out, weights[0] + + +def da_vector_cell(xb, h: StackBuf(4), counters, final: Const): + weights = StackBuf(3 * DA_CELL) + hint_witness(weights, "da_weights") + st = h # Three power-of-two byte windows per cell keep counter offsets disjoint from the base. base = counters[xb ** 3] for w in unroll(0, 3): window = xb ** 3 * GEN ** w end = counters[window * GEN] for b in unroll(0, DA_CELL_BLOCKS): - out = StackBuf(2) + out = StackBuf(4) if const(b + 1 == DA_CELL_BLOCKS): - md = end + counter = end else: - md = base + const(64 * (b + 1)) + counter = base + const(64 * (b + 1)) + flags = 0 if const(final == 1): if const(w == 2): if const(b + 1 == DA_CELL_BLOCKS): - md = md + MD_FINAL - blake2s(packed[w * (DA_CELL // 2) + 4 * b:w * (DA_CELL // 2) + 4 * b + 2], packed[w * (DA_CELL // 2) + 4 * b + 2:w * (DA_CELL // 2) + 4 * b + 4], out, cv=st, md=md) + flags = MD_FINAL + lo = w * DA_CELL + 8 * b + blake2s(weights[lo:lo + 4], weights[lo + 4:lo + 8], out, cv=st, md=[counter, flags]) st = out base = end weights_ptr = addr(weights) - return st[0], st[1], weights_ptr + return st, weights_ptr # ================================ the aggregation node ============================== def da_list_digest(roots, n_g): # At most 16 roots: every BLAKE2s byte counter and final flag is constant. - a, b = match(log(n_g), range(0, DA_ROOT_COUNTS), lambda n: da_hash_roots(roots, n)) - return a, b + digest = match(log(n_g), range(0, DA_ROOT_COUNTS), lambda n: da_hash_roots(roots, n)) + return digest def da_hash_roots(roots, n: Const): - digest = StackBuf(2) + digest = StackBuf(4) if const(n == 0): - blake2s([0, 0], [0, 0], digest, counter=0, final=1) + blake2s([0, 0, 0, 0], [0, 0, 0, 0], digest, counter=0, final=1) else: - st = StackBuf(2) - st[0] = BLAKE2S_IV_0 - st[1] = BLAKE2S_IV_1 - for i in unroll(0, n): - claim = roots * GEN ** (4 * i) - out = StackBuf(2) - blake2s(claim[0:2], claim[2:4], out, cv=st, counter=64 * (i + 1), final=(i + 1) // n) - st = out - digest[0] = st[0] - digest[1] = st[1] - return digest[0], digest[1] + blake2s(roots[0:4], roots[4:8], digest, counter=64, final=1 // n) + for i in unroll(1, n): + claim = roots * GEN ** (8 * i) + out = StackBuf(4) + blake2s(claim[0:4], claim[4:8], out, cv=digest, counter=64 * (i + 1), final=(i + 1) // n) + digest = out + return digest def cover_da_root(roots, cover, n_slots_g, mark): + # Marks one DA slot and returns its eight words: the root, then the vector hash. index = hint_witness("da_index") assert log(index) < log(n_slots_g) cover[index] = mark - root = roots * (index ** 4) - return root[1], root[GEN], root[GEN ** 2], root[GEN ** 3] + return roots * index ** 8 def main(): @@ -3249,8 +3247,8 @@ def main(): # # so one range check per write keeps each writer inside its own region: that is # what makes the statement's split mean which scheme verified which key against - # which (epoch, message). An XMSS slot is two cells, a SPHINCS slot four: a key - # and the message that key signed. A DA root occupies two cells. + # which (epoch, message). An XMSS slot is four words, a SPHINCS slot eight: a key + # and the message that key signed. A DA root occupies eight words. meta = StackBuf(7) hint_witness(meta, "meta") # every count in the exponent n_decl_g = meta[0] @@ -3280,14 +3278,14 @@ def main(): assert log(da_slots_g) < MAX_RECURSIONS * MAX_DA_ROOTS + 2 # ---- the epoch groups: geometry pass ---- - # Per group: its epoch, its two message cells, and its declared, duplicate and + # Per group: its epoch, its four message words, and its declared, duplicate and # raw-signature counts, each count bounded before it enters a product (up to # 2^16 factors of exponent < 2^17 stay far from the order 2^64 - 1, so nothing # wraps). Region bases and the three totals ride a stride-4 chain; the per-group # values land in heap buffers the later passes and the children's hinted group # maps read back at runtime. epochs = HeapBuf(n_epochs_g) - msgs = HeapBuf(n_epochs_g * n_epochs_g) + msgs = HeapBuf(n_epochs_g ** 4) group_n_keys = HeapBuf(n_epochs_g) group_n_dups = HeapBuf(n_epochs_g) group_n_raw = HeapBuf(n_epochs_g) @@ -3298,28 +3296,27 @@ def main(): geo[GEN] = 1 geo[GEN ** 2] = 1 for xe in mul_range(1, n_epochs_g): - grp = StackBuf(6) - hint_witness(grp, "group") # epoch, msg_lo, msg_hi, n, n_dup, n_raw - assert log(grp[3]) < MAX_KEYS - assert log(grp[4]) < MAX_KEYS + grp = StackBuf(8) + hint_witness(grp, "group") # epoch, message (four words), n, n_dup, n_raw assert log(grp[5]) < MAX_KEYS + assert log(grp[6]) < MAX_KEYS + assert log(grp[7]) < MAX_KEYS epochs[xe] = grp[0] - msg = msgs * (xe * xe) - msg[1] = grp[1] - msg[GEN] = grp[2] - group_n_keys[xe] = grp[3] - group_n_dups[xe] = grp[4] - group_n_raw[xe] = grp[5] - state = geo * (xe ** 4) + msg = msgs * xe ** 4 + msg[0:4] = grp[1:5] + group_n_keys[xe] = grp[5] + group_n_dups[xe] = grp[6] + group_n_raw[xe] = grp[7] + state = geo * xe ** 4 base = state[1] group_base[xe] = base - slots = grp[3] * grp[4] + slots = grp[5] * grp[6] group_slots[xe] = slots - nxt = geo * ((xe * GEN) ** 4) + nxt = geo * (xe * GEN) ** 4 nxt[1] = base * slots - nxt[GEN] = state[GEN] * grp[3] - nxt[GEN ** 2] = state[GEN ** 2] * grp[5] - geo_end = geo * (n_epochs_g ** 4) + nxt[GEN] = state[GEN] * grp[5] + nxt[GEN ** 2] = state[GEN ** 2] * grp[7] + geo_end = geo * n_epochs_g ** 4 xmss_slots_g = geo_end[1] n_raw_x_g = geo_end[GEN ** 2] sphincs_slots_g = n_sphincs_g * n_sdup_g @@ -3332,21 +3329,20 @@ def main(): # The proving environment (flock's R1CS and this bytecode) as one digest. It # rides the statement rather than the bytecode, so nothing here has to know its # own hash; the outer verifier pins it, and every child statement rebuilt below - # copies it, which is what keeps a whole tree on one bytecode. - fs_seed = StackBuf(2) - hint_witness(fs_seed, "fs_seed") - seed_0 = fs_seed[0] - seed_1 = fs_seed[1] + # copies it, which is what keeps a whole tree on one bytecode. It lives on the + # heap, where the loops over children can reach it. + seed = HeapBuf(4) + hint_witness(seed[0:4], "fs_seed") # ---- the signer set ---- g_logs_pow2, g_squares = exponent_tables() # ---- data availability ---- - da_roots = HeapBuf(da_slots_g ** 4) + da_roots = HeapBuf(da_slots_g ** 8) for xd in mul_range(1, da_slots_g): - root = da_roots * (xd ** 4) - hint_witness(root[0:4], "da_roots") - da_0, da_1 = da_list_digest(da_roots, n_da_g) + root = da_roots * xd ** 8 + hint_witness(root[0:8], "da_roots") + da_digest = da_list_digest(da_roots, n_da_g) # One table per scheme, the XMSS one an epoch group at a time: each group its # declared list (strictly sorted, checked by the outer verifier, which holds it) # followed by its own duplicate slots. The coverage indices below run over one @@ -3359,14 +3355,16 @@ def main(): # prefix of another's and the digest binds its own lengths. `half` and `odd` are # hinted per group and pinned by half*half*odd == n with odd in {0, 1}, which # leaves half = n // 2 and odd = n % 2 as the only solution. - xmss_table = HeapBuf(xmss_slots_g * xmss_slots_g) - sphincs_table = HeapBuf(sphincs_slots_g ** 4) + xmss_table = HeapBuf(xmss_slots_g ** 4) + sphincs_table = HeapBuf(sphincs_slots_g ** 8) # The run the set's hash covers: both lengths and the SPHINCS list's digest - # in one block, then two a group. Eight cells a group, so a group's + # in one block, then two a group. Sixteen words a group, so a group's # header and its key digest are one block each. - signers_run = HeapBuf(n_decl_g ** 8 * GEN ** 4) + signers_run = HeapBuf(n_decl_g ** 16 * GEN ** 8) signers_run[1] = n_decl_g - signers_run[GEN] = n_sphincs_g + signers_run[GEN] = 0 + signers_run[GEN ** 2] = n_sphincs_g + signers_run[GEN ** 3] = 0 decl_keys = HeapBuf(n_decl_g * GEN) decl_keys[GEN ** 0] = 1 for xe in mul_range(1, n_decl_g): @@ -3378,38 +3376,33 @@ def main(): assert log(halves[1]) < 2 assert log(halves[0]) < MAX_KEYS assert halves[0] * halves[0] * halves[1] == n_keys - kd_0, kd_1 = key_list_digest(xmss_table * (base * base), halves[0], halves[1], n_keys, g_squares) - group_msg = msgs * (xe * xe) - slot = signers_run * (xe ** 8) * (GEN ** 4) + kd = key_list_digest(xmss_table * base ** 4, halves[0], halves[1], n_keys, g_squares) + group_msg = msgs * xe ** 4 + slot = signers_run * xe ** 16 * GEN ** 8 slot[1] = epochs[xe] - slot[GEN] = n_keys - slot[GEN ** 2] = group_msg[1] - slot[GEN ** 3] = group_msg[GEN] - slot[GEN ** 4] = kd_0 - slot[GEN ** 5] = kd_1 - slot[GEN ** 6] = 0 - slot[GEN ** 7] = 0 + slot[GEN] = 0 + slot[GEN ** 2] = n_keys + slot[GEN ** 3] = 0 + slot[4:8] = group_msg[0:4] + slot[8:12] = kd + slot[12:16] = [0, 0, 0, 0] # The duplicate slots ride the same table but outside the hashed prefix, past # each group's declared keys. Over the whole table: a dropped group has only these. for xe in mul_range(1, n_epochs_g): n_keys = group_n_keys[xe] base = group_base[xe] - dup_ptr = xmss_table * (base * base * n_keys * n_keys) + dup_ptr = xmss_table * (base * n_keys) ** 4 for xd in mul_range(1, group_n_dups[xe]): - dup = dup_ptr * (xd * xd) - hint_witness(dup[0:2], "dup_pubkeys") - sp_0, sp_1 = sphincs_list_digest(sphincs_table, n_sphincs_g, g_squares) - signers_run[GEN ** 2] = sp_0 - signers_run[GEN ** 3] = sp_1 + dup = dup_ptr * xd ** 4 + hint_witness(dup[0:4], "dup_pubkeys") + sp_digest = sphincs_list_digest(sphincs_table, n_sphincs_g, g_squares) + signers_run[4:8] = sp_digest # At least one published signature claim or DA root also ensures a nonempty coverage table. assert decl_keys[n_decl_g] * n_sphincs_g * n_da_g != 1 - set_0, set_1 = signer_set_digest(signers_run, n_decl_g, g_squares) - signers_hash = HeapBuf(WORDS_PER_BLOCK) - signers_hash[1] = set_0 - signers_hash[GEN] = set_1 + signers_hash = signer_set_digest(signers_run, n_decl_g, g_squares) for xd in mul_range(1, n_sdup_g): - dup = sphincs_table * ((n_sphincs_g * xd) ** 4) - hint_witness(dup[0:4], "dup_sphincs") + dup = sphincs_table * (n_sphincs_g * xd) ** 8 + hint_witness(dup[0:8], "dup_sphincs") # ---- coverage ---- # Every one of the n_total slots is written exactly once: write-once memory @@ -3418,25 +3411,25 @@ def main(): # declared claim is covered by a direct check of its own kind or by a verified # child. Each writer is confined to its own region, including DA roots. # The raw XMSS walk runs one loop per epoch group, each signature verified - # against that group's tables (built here, only for a group that holds raw + # against that group's index words (built here, only for a group that holds raw # signatures); a stride-1 chain threads the running count across the groups. cover = HeapBuf(n_total_g) da_cover = cover * da_base_g - merkle_bits = HeapBuf(n_epochs_g ** MERKLE_BIT_CELLS) - tweak_tables = HeapBuf(n_epochs_g ** N_TWEAK_CELLS) + merkle_bits = HeapBuf(n_epochs_g ** LOG_LIFETIME) + index_tables = HeapBuf(n_epochs_g ** (LOG_LIFETIME + 1)) raw_count = HeapBuf(n_epochs_g * GEN) raw_count[GEN ** 0] = 1 for xe in mul_range(1, n_epochs_g): n_raw = group_n_raw[xe] prefix = raw_count[xe] if n_raw != 1: - tweak_table = tweak_tables * (xe ** N_TWEAK_CELLS) - group_bits = merkle_bits * (xe ** MERKLE_BIT_CELLS) - fill_xmss_epoch_tables(epochs[xe], group_bits, tweak_table) + index_table = index_tables * xe ** (LOG_LIFETIME + 1) + group_bits = merkle_bits * xe ** LOG_LIFETIME + fill_xmss_epoch_tables(epochs[xe], group_bits, index_table) slots = group_slots[xe] base = group_base[xe] - keys = xmss_table * (base * base) - group_msg = msgs * (xe * xe) + keys = xmss_table * base ** 4 + group_msg = msgs * xe ** 4 for xi in mul_range(1, n_raw): idx = hint_witness("raw_index") # A runtime bound, whose `n_total < 2^MIN_LOG_MEM` precondition is @@ -3447,7 +3440,7 @@ def main(): # message) covers no other group's declared key. assert log(idx) < log(slots) cover[base * idx] = prefix * xi - verify_sig(group_msg, tweak_table, group_bits, keys * (idx * idx)) + verify_sig(group_msg, index_table, group_bits, keys * idx ** 4) raw_count[xe * GEN] = prefix * n_raw else: raw_count[xe * GEN] = prefix @@ -3455,28 +3448,26 @@ def main(): off_hint = hint_witness("sp_raw_index") assert log(off_hint) < log(sphincs_slots_g) cover[xmss_slots_g * off_hint] = n_raw_x_g * xj - verify_sig_sphincs(sphincs_table * (off_hint ** 4)) + verify_sig_sphincs(sphincs_table * off_hint ** 8) if n_direct_da_g == GEN: - root_0, root_1, vector_0, vector_1 = da_verify(g_squares) - expected_0, expected_1, expected_v0, expected_v1 = cover_da_root(da_roots, da_cover, da_slots_g, n_raw_x_g * n_raw_s_g) - assert root_0 == expected_0 - assert root_1 == expected_1 - assert vector_0 == expected_v0 - assert vector_1 == expected_v1 + da_root, da_vector = da_verify(g_squares) + expected = cover_da_root(da_roots, da_cover, da_slots_g, n_raw_x_g * n_raw_s_g) + expected[0:4] = da_root + expected[4:8] = da_vector # ---- children ---- - child_pi = HeapBuf(n_children_g * n_children_g) - child_fresh = HeapBuf(n_children_g ** DEFER_SIZE) - child_carried = HeapBuf(n_children_g ** DEFER_STMT_CELLS) + child_pi = HeapBuf(n_children_g ** 4) + child_fresh = HeapBuf(n_children_g ** (3 * DEFER_SIZE)) + child_carried = HeapBuf(n_children_g ** (3 * DEFER_STMT_CELLS)) written = HeapBuf(n_children_g * GEN) # loop-carried write count, one per child written[GEN ** 0] = n_raw_x_g * n_raw_s_g * n_direct_da_g for xc in mul_range(1, n_children_g): base = written[xc] # The child's two list lengths, then its groups, rebuilt into its signer-set - # chain by rebuild_child_groups: everything hinted there is pinned by the - # chain, which the child's statement digest carries, so a lie about any of - # it changes the public input its proof has to satisfy. Nothing demands a + # run by rebuild_child_groups: everything hinted there is pinned by the + # run's hash, which the child's statement digest carries, so a lie about any + # of it changes the public input its proof has to satisfy. Nothing demands a # mid-tree statement be canonical (sorted, distinct groups); it still binds # every claim to its (epoch, message), which is all the group map relies on. child_meta = StackBuf(2) @@ -3485,65 +3476,63 @@ def main(): nsub_s_g = child_meta[1] assert log(nsub_e_g) < MAX_EPOCHS + 1 assert log(nsub_s_g) < MAX_KEYS - sub_run = HeapBuf(nsub_e_g ** 8 * GEN ** 4) + sub_run = HeapBuf(nsub_e_g ** 16 * GEN ** 8) sub_run[1] = nsub_e_g - sub_run[GEN] = nsub_s_g + sub_run[GEN] = 0 + sub_run[GEN ** 2] = nsub_s_g + sub_run[GEN ** 3] = 0 nsub_x_g = rebuild_child_groups(nsub_e_g, sub_run, base, epochs, msgs, group_base, group_slots, n_epochs_g, xmss_table, cover, g_squares) # Implied by the per-group bounds and the child's own n_total assert; stands # as documentation. assert log(nsub_x_g) < MAX_KEYS nsub_g = nsub_x_g * nsub_s_g - csp_0, csp_1 = child_sphincs_list_digest(sphincs_table, cover, base * nsub_x_g, xmss_slots_g, sphincs_slots_g, nsub_s_g, g_squares) - sub_run[GEN ** 2] = csp_0 - sub_run[GEN ** 3] = csp_1 - sub_set_0, sub_set_1 = signer_set_digest(sub_run, nsub_e_g, g_squares) - sub_hash = HeapBuf(WORDS_PER_BLOCK) - sub_hash[1] = sub_set_0 - sub_hash[GEN] = sub_set_1 - carried = child_carried * xc ** DEFER_STMT_CELLS - hint_witness(carried[0:DEFER_STMT_CELLS], "child_defer") + csp = child_sphincs_list_digest(sphincs_table, cover, base * nsub_x_g, xmss_slots_g, sphincs_slots_g, nsub_s_g, g_squares) + sub_run[4:8] = csp + sub_hash = signer_set_digest(sub_run, nsub_e_g, g_squares) + carried = child_carried * xc ** (3 * DEFER_STMT_CELLS) + hint_witness(carried[0:3 * DEFER_STMT_CELLS], "child_defer") nsub_da_g = hint_witness("child_da_count") assert log(nsub_da_g) < MAX_DA_ROOTS + 1 - child_da = HeapBuf(nsub_da_g ** 4) + child_da = HeapBuf(nsub_da_g ** 8) for xd in mul_range(1, nsub_da_g): - root = child_da * (xd ** 4) - root_0, root_1, vector_0, vector_1 = cover_da_root(da_roots, da_cover, da_slots_g, base * nsub_g * xd) - root[1] = root_0 - root[GEN] = root_1 - root[GEN ** 2] = vector_0 - root[GEN ** 3] = vector_1 - child_da_0, child_da_1 = da_list_digest(child_da, nsub_da_g) - pi_0, pi_1 = statement_digest(seed_0, seed_1, sub_hash, child_da_0, child_da_1, carried) - pi = xc * xc - child_pi[pi] = pi_0 - child_pi[pi * GEN] = pi_1 - verify_sub(pi_0, pi_1, seed_0, seed_1, g_logs_pow2, g_squares, child_fresh * xc ** DEFER_SIZE) + covered = cover_da_root(da_roots, da_cover, da_slots_g, base * nsub_g * xd) + slot = child_da * xd ** 8 + slot[0:8] = covered[0:8] + child_da_digest = da_list_digest(child_da, nsub_da_g) + pi = statement_digest(seed, sub_hash, child_da_digest, carried) + pi_slot = child_pi * xc ** 4 + pi_slot[0:4] = pi + verify_sub(pi, seed, g_logs_pow2, g_squares, child_fresh * xc ** (3 * DEFER_SIZE)) written[xc * GEN] = base * nsub_g * nsub_da_g assert written[n_children_g] == n_total_g # ---- this node's own deferred claims ---- - defer_stmt = HeapBuf(DEFER_STMT_CELLS) + defer_stmt = HeapBuf(3 * DEFER_STMT_CELLS) if n_children_g == 1: # A leaf has nothing to batch, so it defers the three fixed polynomials at # the all-zeros point. Their values ride a hint and are checked nowhere # here: the outer verifier recomputes them and rebuilds the statement, so a # lie changes the public input rather than the claim. - leaf_values = StackBuf(3) + leaf_values = StackBuf(9) hint_witness(leaf_values, "leaf_defer") for k in unroll(0, BYTECODE_VARS): - defer_stmt[GEN ** k] = 0 - defer_stmt[GEN ** DEFER_STMT_BC_VALUE] = leaf_values[0] - for k in unroll(0, 2 * K_LOG): - defer_stmt[GEN ** (DEFER_STMT_MAT_POINT + k)] = 0 - defer_stmt[GEN ** DEFER_STMT_A_VALUE] = leaf_values[1] - defer_stmt[GEN ** DEFER_STMT_B_VALUE] = leaf_values[2] + c = 3 * k + defer_stmt[c:c + 3] = ZERO + k = 3 * DEFER_STMT_BC_VALUE + defer_stmt[k:k + 3] = leaf_values[0:3] + for j in unroll(0, 2 * K_LOG): + c = 3 * (DEFER_STMT_MAT_POINT + j) + defer_stmt[c:c + 3] = ZERO + k = 3 * DEFER_STMT_A_VALUE + defer_stmt[k:k + 3] = leaf_values[3:6] + k = 3 * DEFER_STMT_B_VALUE + defer_stmt[k:k + 3] = leaf_values[6:9] else: aggregate_claims(n_children_g, child_pi, child_fresh, child_carried, defer_stmt) - own_0, own_1 = statement_digest(seed_0, seed_1, signers_hash, da_0, da_1, defer_stmt) - pub_ptr = GEN ** 0 - assert pub_ptr[1] == own_0 - assert pub_ptr[GEN] == own_1 + own = statement_digest(seed, signers_hash, da_digest, defer_stmt) + pub = GEN ** 0 + pub[0:4] = own return @@ -3558,160 +3547,171 @@ def aggregate_claims(n_children_g, child_pi, child_fresh, child_carried, defer_s # A carried claim is a plain point, so its weight is an eq product; a fresh one # carries flock's zerocheck/lincheck structure and keeps the succinct weight the # sub-verifier exported. That is the only asymmetry. - bc_msgs = HeapBuf(2 * BYTECODE_VARS) - hint_witness(bc_msgs[0:2 * BYTECODE_VARS], "bc_sumcheck_msgs") - mat_msgs = HeapBuf(4 * K_LOG) - hint_witness(mat_msgs[0:4 * K_LOG], "mat_sumcheck_msgs") - bytecode_star = hint_witness("bc_star_hint") - mat_stars = StackBuf(2) - hint_witness(mat_stars[0:2], "mat_stars_hint") + bc_msgs = HeapBuf(6 * BYTECODE_VARS) + hint_witness(bc_msgs[0:6 * BYTECODE_VARS], "bc_sumcheck_msgs") + mat_msgs = HeapBuf(12 * K_LOG) + hint_witness(mat_msgs[0:12 * K_LOG], "mat_sumcheck_msgs") + bytecode_star = StackBuf(3) + hint_witness(bytecode_star, "bc_star_hint") + mat_stars = StackBuf(6) + hint_witness(mat_stars, "mat_stars_hint") # ---- one transcript over every child's statement and both its claim sets ---- + # A child's public input is observed as two scalars, a 128-bit half each. fresh_row = HeapBuf(n_children_g) carried_row = HeapBuf(n_children_g) - agg_fs = [AGG_SEED_0, AGG_SEED_1] - agg_fs = obs(agg_fs, n_children_g) - absorb = HeapBuf((n_children_g * GEN) ** PAIR_SLOTS) - absorb[GEN ** 0] = agg_fs[0] - absorb[GEN ** 1] = agg_fs[1] + agg_fs = obs([AGG_SEED_0, AGG_SEED_1, AGG_SEED_2, AGG_SEED_3], [n_children_g, 0, 0]) + absorb = HeapBuf((n_children_g * GEN) ** FS_SLOTS) + absorb[0:4] = agg_fs for xc in mul_range(1, n_children_g): - row = absorb * xc ** PAIR_SLOTS - st = [row[GEN ** 0], row[GEN ** 1]] - pi = xc * xc - st = obs(st, child_pi[pi]) - st = obs(st, child_pi[pi * GEN]) - fresh = child_fresh * xc ** DEFER_SIZE + row = absorb * xc ** FS_SLOTS + pi = child_pi * xc ** 4 + st = obs(row[0:4], [pi[1], pi[GEN], 0]) + st = obs(st, [pi[GEN ** 2], pi[GEN ** 3], 0]) + fresh = child_fresh * xc ** (3 * DEFER_SIZE) fresh_row[xc] = fresh for k in unroll(0, DEFER_SIZE): - st = obs(st, fresh[GEN ** k]) - carried = child_carried * xc ** DEFER_STMT_CELLS + st = obs_at(st, fresh * GEN ** (3 * k)) + carried = child_carried * xc ** (3 * DEFER_STMT_CELLS) carried_row[xc] = carried for k in unroll(0, DEFER_STMT_CELLS): - st = obs(st, carried[GEN ** k]) - row[GEN ** PAIR_SLOTS] = st[0] - row[GEN ** (PAIR_SLOTS + 1)] = st[1] - absorbed = absorb * n_children_g ** PAIR_SLOTS + st = obs_at(st, carried * GEN ** (3 * k)) + row[4:8] = st + absorbed = absorb * n_children_g ** FS_SLOTS # ---- bytecode batching sumcheck (BYTECODE_VARS variables, 2 per child) ---- # Fresh and carried share the bytecode layout (point, then value), so the two # differ only in which buffer they come from. - lam_bc = HeapBuf(n_children_g * n_children_g) + lam_bc = HeapBuf(n_children_g ** 6) bc_chain = HeapBuf((n_children_g * GEN) ** ACC_SLOTS) - bc_chain[GEN ** ACC_FS0] = absorbed[GEN ** 0] - bc_chain[GEN ** ACC_FS1] = absorbed[GEN ** 1] - bc_chain[GEN ** ACC_VALUE] = 0 + bc_chain[0:4] = absorbed[0:4] + bc_chain[ACC_VALUE:ACC_VALUE + 3] = ZERO for xc in mul_range(1, n_children_g): row = bc_chain * xc ** ACC_SLOTS - st = [row[GEN ** ACC_FS0], row[GEN ** ACC_FS1]] - st, lam_fresh = squeeze(st) + st, lam_fresh = squeeze(row[0:4]) st, lam_carried = squeeze(st) - pair = xc * xc - lam_bc[pair] = lam_fresh - lam_bc[pair * GEN] = lam_carried + pair = lam_bc * xc ** 6 + pair[0:3] = lam_fresh + pair[3:6] = lam_carried fresh = fresh_row[xc] carried = carried_row[xc] nxt = row * GEN ** ACC_SLOTS - nxt[GEN ** ACC_FS0] = st[0] - nxt[GEN ** ACC_FS1] = st[1] - nxt[GEN ** ACC_VALUE] = row[GEN ** ACC_VALUE] + lam_fresh * fresh[GEN ** FRESH_BC_VALUE] + lam_carried * carried[GEN ** DEFER_STMT_BC_VALUE] + nxt[0:4] = st + fb = 3 * FRESH_BC_VALUE + cb = 3 * DEFER_STMT_BC_VALUE + nxt[ACC_VALUE:ACC_VALUE + 3] = add192(row[ACC_VALUE:ACC_VALUE + 3], add192(mul192(lam_fresh, fresh[fb:fb + 3]), mul192(lam_carried, carried[cb:cb + 3]))) bc_end = bc_chain * n_children_g ** ACC_SLOTS - agg_fs = [bc_end[GEN ** ACC_FS0], bc_end[GEN ** ACC_FS1]] - bc_running = bc_end[GEN ** ACC_VALUE] - bc_point = HeapBuf(BYTECODE_VARS) - fs0, fs1, bc_running = batch_sumcheck(agg_fs[0], agg_fs[1], bc_msgs, bc_running, bc_point, BYTECODE_VARS) - agg_fs = [fs0, fs1] - bc_wsum = HeapBuf(n_children_g * GEN) - bc_wsum[GEN ** 0] = 0 + bc_point = HeapBuf(3 * BYTECODE_VARS) + agg_fs, bc_running = batch_sumcheck(bc_end[0:4], bc_msgs, bc_end[ACC_VALUE:ACC_VALUE + 3], bc_point, BYTECODE_VARS) + bc_wsum = HeapBuf((n_children_g * GEN) ** 3) + bc_wsum[0:3] = ZERO for xc in mul_range(1, n_children_g): fresh = fresh_row[xc] carried = carried_row[xc] - eq_fresh = GEN ** 0 - eq_carried = GEN ** 0 + eq_fresh = ONE + eq_carried = ONE for k in unroll(0, BYTECODE_VARS): - rk = bc_point[GEN ** k] - eq_fresh *= (1 + fresh[GEN ** k] + rk) - eq_carried *= (1 + carried[GEN ** k] + rk) - pair = xc * xc - bc_wsum[xc * GEN] = bc_wsum[xc] + lam_bc[pair] * eq_fresh + lam_bc[pair * GEN] * eq_carried - assert bc_running == bytecode_star * bc_wsum[n_children_g] + c = 3 * k + rk = bc_point[c:c + 3] + eq_fresh = mul192(eq_fresh, add192(ONE, add192(fresh[c:c + 3], rk))) + eq_carried = mul192(eq_carried, add192(ONE, add192(carried[c:c + 3], rk))) + pair = lam_bc * xc ** 6 + x3 = xc ** 3 + nxt = x3 * GEN ** 3 + bc_wsum[nxt:nxt + 3] = add192(bc_wsum[x3:x3 + 3], add192(mul192(pair[0:3], eq_fresh), mul192(pair[3:6], eq_carried))) + ws_end = n_children_g ** 3 + assert_eq192(bc_running, mul192(bytecode_star, bc_wsum[ws_end:ws_end + 3])) # ---- matrix batching sumcheck (2*K_LOG variables, 3 claims per child) ---- # The fresh claim is one value against A0 weighted by lincheck's alpha plus B0; # a carried claim is one value per matrix at a shared point. - lam_mat = HeapBuf(n_children_g ** 3) + lam_mat = HeapBuf(n_children_g ** 9) mat_chain = HeapBuf((n_children_g * GEN) ** ACC_SLOTS) - mat_chain[GEN ** ACC_FS0] = agg_fs[0] - mat_chain[GEN ** ACC_FS1] = agg_fs[1] - mat_chain[GEN ** ACC_VALUE] = 0 + mat_chain[0:4] = agg_fs + mat_chain[ACC_VALUE:ACC_VALUE + 3] = ZERO for xc in mul_range(1, n_children_g): row = mat_chain * xc ** ACC_SLOTS - st = [row[GEN ** ACC_FS0], row[GEN ** ACC_FS1]] - st, lam_fresh = squeeze(st) + st, lam_fresh = squeeze(row[0:4]) st, lam_a = squeeze(st) st, lam_b = squeeze(st) - triple = xc ** 3 - lam_mat[triple] = lam_fresh - lam_mat[triple * GEN] = lam_a - lam_mat[triple * GEN ** 2] = lam_b + triple = lam_mat * xc ** 9 + triple[0:3] = lam_fresh + triple[3:6] = lam_a + triple[6:9] = lam_b fresh = fresh_row[xc] carried = carried_row[xc] nxt = row * GEN ** ACC_SLOTS - nxt[GEN ** ACC_FS0] = st[0] - nxt[GEN ** ACC_FS1] = st[1] - nxt[GEN ** ACC_VALUE] = row[GEN ** ACC_VALUE] + lam_fresh * fresh[GEN ** FRESH_MATPART] + lam_a * carried[GEN ** DEFER_STMT_A_VALUE] + lam_b * carried[GEN ** DEFER_STMT_B_VALUE] + nxt[0:4] = st + fm = 3 * FRESH_MATPART + ca = 3 * DEFER_STMT_A_VALUE + cb = 3 * DEFER_STMT_B_VALUE + nxt[ACC_VALUE:ACC_VALUE + 3] = add192(row[ACC_VALUE:ACC_VALUE + 3], add192(mul192(lam_fresh, fresh[fm:fm + 3]), add192(mul192(lam_a, carried[ca:ca + 3]), mul192(lam_b, carried[cb:cb + 3])))) mat_end = mat_chain * n_children_g ** ACC_SLOTS - agg_fs = [mat_end[GEN ** ACC_FS0], mat_end[GEN ** ACC_FS1]] - mat_running = mat_end[GEN ** ACC_VALUE] - mat_point = HeapBuf(2 * K_LOG) - fs0, fs1, mat_running = batch_sumcheck(agg_fs[0], agg_fs[1], mat_msgs, mat_running, mat_point, 2 * K_LOG) - agg_fs = [fs0, fs1] + mat_point = HeapBuf(3 * 2 * K_LOG) + agg_fs, mat_running = batch_sumcheck(mat_end[0:4], mat_msgs, mat_end[ACC_VALUE:ACC_VALUE + 3], mat_point, 2 * K_LOG) # Terminal weights. A fresh claim's is U_t(r*) = urow_t(r*_row) * wcol_t(r*_col), # with row_weight = (sum_i L_i(zz_t) eq(r*[0..6], i)) * eq(zchi_t, r*[6..K_LOG]) # and col_weight = (sum_i z_partial_t[i] eq(r*[K_LOG..K_LOG+6], i)) * prod_j (1 + # lrr_j + r*[2*K_LOG-1-j]) (the lincheck binds column variables top-down). A # carried claim's is a plain eq over all 2*K_LOG coordinates. - eq_rows = HeapBuf(2 ** (K_SKIP + 1) - 2) + eq_rows = HeapBuf(3 * (2 ** (K_SKIP + 1) - 2)) eqtree(mat_point, eq_rows, K_SKIP) - eq_cols = HeapBuf(2 ** (K_SKIP + 1) - 2) - eqtree(mat_point * GEN ** K_LOG, eq_cols, K_SKIP) - w_sums = HeapBuf((n_children_g * GEN) ** PAIR_SLOTS) # the A and B weight sums - w_sums[GEN ** 0] = 0 - w_sums[GEN ** 1] = 0 + eq_cols = HeapBuf(3 * (2 ** (K_SKIP + 1) - 2)) + eqtree(mat_point * GEN ** (3 * K_LOG), eq_cols, K_SKIP) + w_sums = HeapBuf((n_children_g * GEN) ** 6) # the A and B weight sums + w_sums[0:3] = ZERO + w_sums[3:6] = ZERO for xc in mul_range(1, n_children_g): fresh = fresh_row[xc] - row_nums = StackBuf(2 ** K_SKIP) - lag64(fresh[GEN ** FRESH_Z_SKIP], row_nums, 0) - row_weight = 0 + row_nums = StackBuf(3 * 2 ** K_SKIP) + zs = 3 * FRESH_Z_SKIP + lag64(fresh[zs:zs + 3], row_nums, 0) + row_weight = ZERO for i in unroll(0, 2 ** K_SKIP): - row_weight += row_nums[i] * eq_rows[GEN ** (2 ** K_SKIP - 2 + i)] - row_weight *= LAGRANGE_INV_S + k = 3 * i + e = 3 * (2 ** K_SKIP - 2 + i) + row_weight = add192(row_weight, mul192(row_nums[k:k + 3], eq_rows[e:e + 3])) + row_weight = mul192(row_weight, LAGRANGE_INV_S) for k in unroll(0, LINCHECK_ROUNDS): - row_weight *= (1 + fresh[GEN ** (FRESH_ZCHI + k)] + mat_point[GEN ** (K_SKIP + k)]) - col_weight = 0 + zc = 3 * (FRESH_ZCHI + k) + mp = 3 * (K_SKIP + k) + row_weight = mul192(row_weight, add192(ONE, add192(fresh[zc:zc + 3], mat_point[mp:mp + 3]))) + col_weight = ZERO for i in unroll(0, 2 ** K_SKIP): - col_weight += fresh[GEN ** (FRESH_Z_PARTIAL + i)] * eq_cols[GEN ** (2 ** K_SKIP - 2 + i)] + zp = 3 * (FRESH_Z_PARTIAL + i) + e = 3 * (2 ** K_SKIP - 2 + i) + col_weight = add192(col_weight, mul192(fresh[zp:zp + 3], eq_cols[e:e + 3])) for j in unroll(0, LINCHECK_ROUNDS): - col_weight *= (1 + fresh[GEN ** (FRESH_LINCHECK_RS + j)] + mat_point[GEN ** (2 * K_LOG - 1 - j)]) - weight_u = row_weight * col_weight + lr = 3 * (FRESH_LINCHECK_RS + j) + mp = 3 * (2 * K_LOG - 1 - j) + col_weight = mul192(col_weight, add192(ONE, add192(fresh[lr:lr + 3], mat_point[mp:mp + 3]))) + weight_u = mul192(row_weight, col_weight) carried = carried_row[xc] - eq_carried = GEN ** 0 + eq_carried = ONE for k in unroll(0, 2 * K_LOG): - eq_carried *= (1 + carried[GEN ** (DEFER_STMT_MAT_POINT + k)] + mat_point[GEN ** k]) - triple = xc ** 3 - lam_fresh = lam_mat[triple] - row = w_sums * xc ** PAIR_SLOTS - row[GEN ** PAIR_SLOTS] = row[GEN ** 0] + lam_fresh * weight_u + lam_mat[triple * GEN] * eq_carried - row[GEN ** (PAIR_SLOTS + 1)] = row[GEN ** 1] + lam_fresh * fresh[GEN ** FRESH_ALPHA] * weight_u + lam_mat[triple * GEN ** 2] * eq_carried - w_end = w_sums * n_children_g ** PAIR_SLOTS - a_star = mat_stars[0] - b_star = mat_stars[1] - assert mat_running == a_star * w_end[GEN ** 0] + b_star * w_end[GEN ** 1] + cp = 3 * (DEFER_STMT_MAT_POINT + k) + mp = 3 * k + eq_carried = mul192(eq_carried, add192(ONE, add192(carried[cp:cp + 3], mat_point[mp:mp + 3]))) + triple = lam_mat * xc ** 9 + lam_fresh = triple[0:3] + fa = 3 * FRESH_ALPHA + row = w_sums * xc ** 6 + row[6:9] = add192(row[0:3], add192(mul192(lam_fresh, weight_u), mul192(triple[3:6], eq_carried))) + row[9:12] = add192(row[3:6], add192(mul192(mul192(lam_fresh, fresh[fa:fa + 3]), weight_u), mul192(triple[6:9], eq_carried))) + w_end = w_sums * n_children_g ** 6 + assert_eq192(mat_running, add192(mul192(mat_stars[0:3], w_end[0:3]), mul192(mat_stars[3:6], w_end[3:6]))) for k in unroll(0, BYTECODE_VARS): - defer_stmt[GEN ** k] = bc_point[GEN ** k] - defer_stmt[GEN ** DEFER_STMT_BC_VALUE] = bytecode_star - for k in unroll(0, 2 * K_LOG): - defer_stmt[GEN ** (DEFER_STMT_MAT_POINT + k)] = mat_point[GEN ** k] - defer_stmt[GEN ** DEFER_STMT_A_VALUE] = a_star - defer_stmt[GEN ** DEFER_STMT_B_VALUE] = b_star + c = 3 * k + defer_stmt[c:c + 3] = bc_point[c:c + 3] + k = 3 * DEFER_STMT_BC_VALUE + defer_stmt[k:k + 3] = bytecode_star + for j in unroll(0, 2 * K_LOG): + c = 3 * j + d = 3 * (DEFER_STMT_MAT_POINT + j) + defer_stmt[d:d + 3] = mat_point[c:c + 3] + k = 3 * DEFER_STMT_A_VALUE + defer_stmt[k:k + 3] = mat_stars[0:3] + k = 3 * DEFER_STMT_B_VALUE + defer_stmt[k:k + 3] = mat_stars[3:6] return diff --git a/crates/rec_aggregation/src/aggregation.rs b/crates/rec_aggregation/src/aggregation.rs index fe3eef387..62e7612e1 100644 --- a/crates/rec_aggregation/src/aggregation.rs +++ b/crates/rec_aggregation/src/aggregation.rs @@ -155,62 +155,43 @@ const _: () = assert!(sphincs::CHAIN_LEN * sphincs::V < 1 << 32); const _: () = assert!(sphincs::A <= 32 && sphincs::HEIGHTS[0] <= 32); /// A count as the guest carries it: in the exponent, `g^n`. -fn count(n: usize) -> F192 { - F192::new(g_pow(n).0, 0, 0) -} - -/// A field element as the decimal `u128` literal the zkDSL parser accepts. -fn dsl_u128(value: F192) -> u128 { - assert_eq!(value.c2, 0, "u128 DSL literal cannot encode the top F192 limb"); - (value.c0 as u128) | ((value.c1 as u128) << 64) +fn count(n: usize) -> F64 { + g_pow(n) } fn f192_literal(f: F192) -> String { format!("f192({},{},{})", f.c0, f.c1, f.c2) } -/// Pack the Fiat-Shamir state's four K lanes as two canonical 128-bit VM cells. -fn pack_state(state: [F64; 4]) -> [F192; 2] { - [ - F192::new(state[0].0, state[1].0, 0), - F192::new(state[2].0, state[3].0, 0), - ] +/// Field elements as the words the guest holds them in, three limbs each. +fn limbs(values: &[F192]) -> Vec { + values.iter().flat_map(|v| [F64(v.c0), F64(v.c1), F64(v.c2)]).collect() } -/// Pack a 32-byte Merkle node as the same canonical 128+128 cell pair used by -/// the VM's sole BLAKE2s representation. -fn pack_hash_state(hash: &[u8; 32]) -> [F192; 2] { - let word_at = |offset: usize| u64::from_le_bytes(hash[offset..offset + 8].try_into().unwrap()); - [ - F192::new(word_at(0), word_at(8), 0), - F192::new(word_at(16), word_at(24), 0), - ] +/// A byte string of whole words as its little-endian words, the order a BLAKE2s +/// block or digest occupies memory in. +fn words(bytes: &[u8]) -> Vec { + let (chunks, rest) = bytes.as_chunks::<8>(); + assert!(rest.is_empty(), "a byte string of whole words"); + chunks.iter().map(|w| F64(u64::from_le_bytes(*w))).collect() } -/// A 16-byte native value as one canonical 128-bit cell. -fn pack_16_bytes(bytes: &[u8]) -> F192 { - let word_at = |offset: usize| u64::from_le_bytes(bytes[offset..offset + 8].try_into().unwrap()); - F192::new(word_at(0), word_at(8), 0) +/// A 32-byte digest as its four words. +fn digest_words(hash: &[u8; 32]) -> [F64; 4] { + words(hash).try_into().expect("four words") } -/// A public key as the two cells the guest hashes and `verify_sig` reads: the +/// A public key as the four words the guest hashes and `verify_sig` reads: the /// root then the public parameter. Both schemes lay a key out the same way, and /// the statement keeps them in separate lists rather than telling them apart by /// their bytes. -fn key_cells(pk: &XmssPublicKey) -> [F192; 2] { - [pack_16_bytes(&pk.merkle_root), pack_16_bytes(&pk.public_param)] +fn key_words(pk: &XmssPublicKey) -> Vec { + [words(&pk.merkle_root), words(&pk.public_param)].concat() } -/// A run of canonical 128-bit cells as the byte string BLAKE2s hashes: each cell -/// is its two low limbs, little-endian, which is the order the VM's compression -/// reads a memory cell in. -fn cell_bytes(cells: impl IntoIterator) -> Vec { - let mut bytes = Vec::new(); - for cell in cells { - bytes.extend_from_slice(&cell.c0.to_le_bytes()); - bytes.extend_from_slice(&cell.c1.to_le_bytes()); - } - bytes +/// A run of words as the byte string BLAKE2s hashes over it. +fn word_bytes(words: impl IntoIterator) -> Vec { + words.into_iter().flat_map(|w| w.0.to_le_bytes()).collect() } /// How the guest splits a list's `n - 1` non-final blocks: whole windows, then the @@ -218,7 +199,7 @@ fn cell_bytes(cells: impl IntoIterator) -> Vec { /// agree with them (`sphincs_list_digest` in the guest). A list is never empty /// here: a group holds at least one key, and both claim lists are guarded at the /// call, since the guest hashes an empty one without a split at all. -fn signers_split(blocks: usize) -> Vec { +fn signers_split(blocks: usize) -> Vec { assert!(blocks > 0, "an empty list has no window split"); let leading = blocks - 1; vec![count(leading / SIGNERS_WINDOW), count(leading % SIGNERS_WINDOW)] @@ -227,47 +208,43 @@ fn signers_split(blocks: usize) -> Vec { /// One epoch group's declared keys under plain BLAKE2s: 32 bytes a key, so the /// hashed string is `32n` bytes and only its last block is partial. The guest /// computes this same digest a window of blocks at a time (`key_list_digest`). -fn key_list_digest(keys: &[XmssPublicKey]) -> [F192; 2] { - let cells = keys.iter().flat_map(key_cells); - pack_hash_state(&primitives::hash::hash(&cell_bytes(cells))) +fn key_list_digest(keys: &[XmssPublicKey]) -> [F64; 4] { + digest_words(&primitives::hash::hash(&word_bytes(keys.iter().flat_map(key_words)))) } /// The declared SPHINCS claims under plain BLAKE2s: one 64-byte block per claim, /// its key then the message it signed, so the hashed string is exactly `64n` bytes /// and an empty list hashes the empty string. The guest computes this same digest a /// window of blocks at a time (`sphincs_list_digest`). -fn sphincs_list_digest(signers: &[SphincsClaim]) -> [F192; 2] { - let cells = signers.iter().flat_map(sphincs_signer_cells); - pack_hash_state(&primitives::hash::hash(&cell_bytes(cells))) +fn sphincs_list_digest(signers: &[SphincsClaim]) -> [F64; 4] { + digest_words(&primitives::hash::hash(&word_bytes( + signers.iter().flat_map(sphincs_signer_words), + ))) } -/// A SPHINCS signer as the four cells the guest hashes and `verify_sig_sphincs` +/// A SPHINCS signer as the eight words the guest hashes and `verify_sig_sphincs` /// reads: the key, then the message that key signed. -fn sphincs_signer_cells((pk, message): &SphincsClaim) -> [F192; 4] { - [ - pack_16_bytes(&pk.root), - pack_16_bytes(&pk.public_param), - pack_16_bytes(&message[..16]), - pack_16_bytes(&message[16..]), - ] +fn sphincs_signer_words((pk, message): &SphincsClaim) -> Vec { + [words(&pk.root), words(&pk.public_param), words(message)].concat() } -/// One XMSS tweak as the cell the guest adds into: `xmss::make_tweak`'s own -/// output, packed. The guest holds no byte layout of its own, building a tweak -/// as this constant half plus one [`tweak_index_weight`] per set epoch bit, so a -/// field that moves in `make_tweak` moves both halves together. -fn tweak_cell(tweak_type: u8, sub_position: u32) -> F192 { - pack_16_bytes(&xmss::make_tweak(tweak_type, sub_position, 0)) +/// The first word of an XMSS tweak, `xmss::make_tweak`'s own output: its type and +/// sub-position. The guest holds no byte layout of its own, building a tweak from +/// these words, a sub-position unit and one [`tweak_index_weight`] per set epoch +/// bit, so a field that moves in `make_tweak` moves both sides together. +fn tweak_word(tweak_type: u8, sub_position: u32) -> F64 { + words(&xmss::make_tweak(tweak_type, sub_position, 0))[0] } -/// What bit `b` of the epoch weighs in a tweak's index field, so an index is its +/// What bit `b` of the epoch weighs in a tweak's second word, so an index is its /// set bits summed. The one property of the layout this assumes is that the /// index field is linear in the index. Subtract the constant protocol prefix. -fn tweak_index_weight(b: usize) -> F192 { - pack_16_bytes(&xmss::make_tweak(0, 0, 1 << b)) + pack_16_bytes(&xmss::make_tweak(0, 0, 0)) +fn tweak_index_weight(b: usize) -> F64 { + let second = |index: u32| words(&xmss::make_tweak(0, 0, index))[1]; + second(1 << b) + second(0) } /// The signer-set digest: plain BLAKE2s of one byte string, laid out in whole -/// 64-byte blocks so the guest can absorb it four cells at a time +/// 64-byte blocks so the guest can absorb it eight words at a time /// (`signer_set_digest` there). The first block carries both list lengths and the /// SPHINCS list's own digest, followed by two blocks a group: its `(epoch, /// count, message)`, then its key list's digest. Leading with both lengths makes @@ -275,25 +252,17 @@ fn tweak_index_weight(b: usize) -> F192 { /// digest binds its own lengths, the groups' epochs and messages, and every split. /// The two list digests carry the bulk, each a stock hash of its own /// ([`key_list_digest`], [`sphincs_list_digest`]). -fn signers_hash(xmss_signers: &[XmssClaimGroup], sphincs_signers: &[SphincsClaim]) -> [F192; 2] { - let sphincs = sphincs_list_digest(sphincs_signers); - let mut cells = vec![ - count(xmss_signers.len()), - count(sphincs_signers.len()), - sphincs[0], - sphincs[1], - ]; +fn signers_hash(xmss_signers: &[XmssClaimGroup], sphincs_signers: &[SphincsClaim]) -> [F64; 4] { + let zero = F64::ZERO; + let mut run = vec![count(xmss_signers.len()), zero, count(sphincs_signers.len()), zero]; + run.extend(sphincs_list_digest(sphincs_signers)); for XmssClaimGroup { epoch, message, keys } in xmss_signers { - cells.extend([ - F192::new(*epoch as u64, 0, 0), - count(keys.len()), - pack_16_bytes(&message[..16]), - pack_16_bytes(&message[16..]), - ]); - let keys = key_list_digest(keys); - cells.extend([keys[0], keys[1], F192::ZERO, F192::ZERO]); + run.extend([F64(*epoch as u64), zero, count(keys.len()), zero]); + run.extend(words(message)); + run.extend(key_list_digest(keys)); + run.extend([zero; 4]); } - pack_hash_state(&primitives::hash::hash(&cell_bytes(cells))) + digest_words(&primitives::hash::hash(&word_bytes(run))) } /// The claims on the three fixed polynomials that a node defers rather than @@ -369,20 +338,13 @@ impl DeferredClaim { } } -/// The statement's fixed header, ahead of the deferred cells: the seed, the +/// The statement's fixed header, ahead of the deferred elements: the seed, the /// signer-set digest (which itself binds the epoch groups and every count), and -/// the DA root-list digest. Fed to the guest as `STMT_HEADER`, so the two cannot drift. -const STATEMENT_HEADER: usize = 6; - -/// A plain BLAKE2s over a lane stream, zero-filled to a whole 64-byte block: -/// what the guest gets by streaming four 128-bit cells a block. -fn lane_hash(lanes: impl Iterator) -> [F192; 2] { - let mut bytes: Vec = lanes.flat_map(u64::to_le_bytes).collect(); - bytes.resize(bytes.len().next_multiple_of(64), 0); - pack_hash_state(&primitives::hash::hash(&bytes)) -} +/// the DA root-list digest, four words each. Fed to the guest as `STMT_HEADER`, so +/// the two cannot drift. +const STATEMENT_HEADER: usize = 12; -/// A node's public statement, hashed to the two words the VM publishes. The +/// A node's public statement, hashed to the four words the VM publishes. The /// guest's `statement_digest` computes exactly this, both for itself and when it /// rebuilds a child's, which is what forces a whole tree onto one bytecode and /// each child's `(epoch, message)` groups, bound by the signer-set digest, @@ -391,37 +353,22 @@ fn lane_hash(lanes: impl Iterator) -> [F192; 2] { /// Fixed-length preimage, so a plain BLAKE2s, with no domain tag of its own: the /// header leads with the environment digest, which binds this bytecode and /// flock's R1CS and so already separates the preimage from every other use of -/// BLAKE2s here. The header is hashed as the canonical cells it already is (two -/// lanes each, whence the assert, the guest being unable to hash a third), then -/// all three lanes of each deferred cell. -fn statement_digest(signers_hash: [F192; 2], da_digest: [u8; 32], defer: &DeferredClaim) -> [F192; 2] { +/// BLAKE2s here. The header's words, then all three limbs of each deferred +/// element, zero-filled to a whole block. +fn statement_digest(signers_hash: [F64; 4], da_digest: [u8; 32], defer: &DeferredClaim) -> [F64; 4] { let seed = lean_vm::cpu::fs_seed(unified_guest()); - let da_digest = pack_hash_state(&da_digest); - let header = [ - seed[0], - seed[1], - signers_hash[0], - signers_hash[1], - da_digest[0], - da_digest[1], - ]; + let header = [seed, signers_hash, digest_words(&da_digest)].concat(); assert_eq!(header.len(), STATEMENT_HEADER); - let mut cells = defer.cells(); - if !cells.len().is_multiple_of(2) { - cells.push(F192::ZERO); // the guest pairs the odd cell with a zero scalar - } - let head = header.iter().flat_map(|x| { - assert_eq!(x.c2, 0, "a header value is a canonical cell"); - [x.c0, x.c1] - }); - lane_hash(head.chain(cells.iter().flat_map(|x| [x.c0, x.c1, x.c2]))) + let mut bytes = word_bytes(header.into_iter().chain(limbs(&defer.cells()))); + bytes.resize(bytes.len().next_multiple_of(64), 0); + digest_words(&primitives::hash::hash(&bytes)) } /// The deferred-claim data the guest binds to the outer public input: the outer /// verifier checks each claim natively (`doc/leanvm/main.tex` §Deferred evaluation claims; /// n_rec = 1 forwards fresh claims without batching). struct DeferredSubproof { - public_input: [F192; 2], + public_input: [F64; 4], bytecode_row_point: Vec, bytecode_selector_point: Vec, bytecode_value: F192, @@ -628,21 +575,20 @@ fn check_da_roots(roots: &[[u8; 32]]) -> Result<(), AggregateVerifyError> { Ok(()) } -fn da_claim_cells(root: &[u8; 32]) -> Vec { +fn da_claim_words(root: &[u8; 32]) -> Vec { let vector_hash = lean_da::vector_digest(&lean_da::membership_vector(root)); - [pack_hash_state(root), pack_hash_state(&vector_hash)].concat() + [words(root), words(&vector_hash)].concat() } fn da_list_digest(roots: &[[u8; 32]]) -> [u8; 32] { // Recompute vector hashes from the roots, including when verifying a received proof. // Trusting prover-supplied hashes would let zero weights certify any matrix. - let cells = roots.iter().flat_map(da_claim_cells); - primitives::hash::hash(&cell_bytes(cells)) + primitives::hash::hash(&word_bytes(roots.iter().flat_map(da_claim_words))) } impl EthereumProof { /// This aggregate's own public statement, as the VM publishes it. - fn public_input(&self) -> [F192; 2] { + fn public_input(&self) -> [F64; 4] { statement_digest( signers_hash(&self.xmss_signers, &self.sphincs_signers), self.da_commitments_digest(), @@ -1011,10 +957,12 @@ fn aggregate_deferred_claims( let klog = flock::hash::K_LOG; let mut transcript = FiatShamirState::from_label(RECURSION_AGG_LABEL); - transcript.observe(count(child_count)); + transcript.observe(count(child_count).into()); for (subproof, carried) in subproofs.iter().zip(carried_claims) { - transcript.observe(subproof.public_input[0]); - transcript.observe(subproof.public_input[1]); + // A public input is observed as two scalars, one 128-bit half each. + let pi = subproof.public_input; + transcript.observe(F192::new(pi[0].0, pi[1].0, 0)); + transcript.observe(F192::new(pi[2].0, pi[3].0, 0)); for &value in &subproof.bytecode_row_point { transcript.observe(value); } @@ -1272,10 +1220,10 @@ fn aggregate_deferred_claims( drop(_span); let hints = vec![ - ("bc_sumcheck_msgs", bscr), - ("mat_sumcheck_msgs", mscr), - ("bc_star_hint", vec![v_bc]), - ("mat_stars_hint", vec![v_a, v_b]), + ("bc_sumcheck_msgs", limbs(&bscr)), + ("mat_sumcheck_msgs", limbs(&mscr)), + ("bc_star_hint", limbs(&[v_bc])), + ("mat_stars_hint", limbs(&[v_a, v_b])), ]; ( hints, @@ -1332,8 +1280,8 @@ enum ClaimSite { /// A table column; `is_virtual` marks the q_flock-backed value /// columns, whose claim is a strided slot rather than a plain column. TableColumn { column: usize, is_virtual: bool }, - /// One of the three PI memory limbs (MEM_LO, MEM_HI, MEM_TOP). - MemoryLimb { column: usize }, + /// The public-input binding claim on `MEM`. + PublicInput { column: usize }, } /// The guest's `COORD_KIND_*` code for a coordinate (`guests/lean_ethereum.py`), @@ -1352,11 +1300,11 @@ fn coord_kind(c: &Coord) -> usize { /// The `K` scalar a coordinate carries beside its columns: the constant itself, /// or the `g^k` a `GCol`/`Prod` scales by. Zero for every other kind. -fn coord_scale(c: &Coord) -> F192 { +fn coord_scale(c: &Coord) -> F64 { match c { - Coord::Const(v) => F192::new(v.0, 0, 0), - Coord::GCol(_, k) | Coord::Prod(_, _, k) => F192::new(g_pow(*k as usize).0, 0, 0), - _ => F192::ZERO, + Coord::Const(v) => *v, + Coord::GCol(_, k) | Coord::Prod(_, _, k) => g_pow(*k as usize), + _ => F64::ZERO, } } @@ -1378,7 +1326,7 @@ fn push_coord_terms(c: &Coord, base: usize, terms: &mut Vec) { }; terms.push(Term { kind: coord_kind(c), - constant: dsl_u128(coord_scale(c)), + constant: coord_scale(c).0, column_a, column_b, }); @@ -1386,7 +1334,7 @@ fn push_coord_terms(c: &Coord, base: usize, terms: &mut Vec) { /// Visit the claim pool in the exact order the guest indexes it: the framework /// bus claims (deduped by `(column, kappa)`, as `leaf.rs` pools them), then every -/// table's committed columns, then the PI memory triple. The placeholder map's +/// table's committed columns, then the PI memory claim. The placeholder map's /// claim descriptors follow this order. fn walk_claims(layout: &lean_vm::cpu::Layout, kbc: usize, mut visit: impl FnMut(ClaimSite)) { let sides: [&[Block]; 3] = [&layout.push, &layout.pull, &layout.count]; @@ -1427,9 +1375,9 @@ fn walk_claims(layout: &lean_vm::cpu::Layout, kbc: usize, mut visit: impl FnMut( }); } } - for &column in &[lean_vm::cpu::MEM_LO, lean_vm::cpu::MEM_HI, lean_vm::cpu::MEM_TOP] { - visit(ClaimSite::MemoryLimb { column }); - } + visit(ClaimSite::PublicInput { + column: lean_vm::cpu::MEM, + }); } /// Config + hints for the recursion guest (`guests/lean_ethereum.py`), built @@ -1437,7 +1385,7 @@ fn walk_claims(layout: &lean_vm::cpu::Layout, kbc: usize, mut visit: impl FnMut( /// `cpu::verify` run (zero hand-mirroring drift). fn gen_verify( program: &Program, - public_input: [F192; 2], + public_input: [F64; 4], summary: lean_vm::cpu::VerifySummary, ) -> Result<(SubHints, DeferredSubproof), AggregationError> { let proof_stream = &summary.raw.stream; @@ -1451,7 +1399,7 @@ fn gen_verify( let side_layouts = sides.map(lean_vm::leaf::layout); // Fixed capacities: every buffer/stride placeholder is a global cap so // the placeholder map is SHAPE-INDEPENDENT (the definition of generic). - assert!(side_layouts.iter().all(|side| side.mu <= MU_CAP) && proof_stream.len() <= STREAM_CAP); + assert!(side_layouts.iter().all(|side| side.mu <= MU_CAP) && 3 * proof_stream.len() <= STREAM_CAP); // The guest holds one opening arm per candidate committed size, so a child // outside that window has no arm to dispatch to. `min_log_committed` keeps // every aggregate above the low end, leaving only the ceiling reachable. @@ -1521,20 +1469,19 @@ fn gen_verify( // polynomial at (ζ_lo, α⃗), the slot coordinates of the claim's own point being // the fingerprint challenges (§sec:e2e-bc). let bytecode_value = summary.bytecode_claim.value; - let bcv = vec![bytecode_value]; // ---- per-sub HINT data (the placeholder map is built once, elsewhere) ---- // Per side, the packing order read straight off `leaf::layout`'s offsets: // sort_order[side_base + rank] = g^{side-local index of the rank-r block}. // The guest only perm-checks it and derives offsets; any aligned tiling is // sound, so this canonical order just has to match the committed leaf. - let mut sort_order: Vec = Vec::new(); + let mut sort_order: Vec = Vec::new(); let mut gbase = 0usize; for (s, blocks) in sides.iter().enumerate() { let mut order: Vec = (0..blocks.len()).collect(); order.sort_by_key(|&i| side_layouts[s].offsets[i]); for &i in &order { - sort_order.push(F192::new(g_pow(gbase + i).0, 0, 0)); // g^{global block index} + sort_order.push(g_pow(gbase + i)); // g^{global block index} } gbase += blocks.len(); } @@ -1554,10 +1501,7 @@ fn gen_verify( } let mut col_order = committed_globals; col_order.sort_by_key(|&global| layout.placements[global].offset); - let col_sort_order: Vec = col_order - .iter() - .map(|&global| F192::new(g_pow(compact_col[global]).0, 0, 0)) - .collect(); + let col_sort_order: Vec = col_order.iter().map(|&global| g_pow(compact_col[global])).collect(); // ---- Phase E2 hints (the stacked WHIR opening) ---- // Share the upper Merkle tree across queries. Unknown, unused subtrees stay @@ -1570,12 +1514,9 @@ fn gen_verify( let n = 1 << cap_depth; let path_depth = depth - cap_depth; let mut nodes = vec![[0u8; 32]; 2 * n]; - let mut active = vec![F192::ZERO; n]; + let mut active = vec![F64::ZERO; n]; for opening in openings.by_ref().take(queries) { - query_hints.push(( - "merkle_leaf_rows", - opening.leaf_data.iter().map(|x| F192::from(*x)).collect(), - )); + query_hints.push(("merkle_leaf_rows", opening.leaf_data.clone())); let mut path_children = Vec::with_capacity(4 * path_depth); let bytes: Vec<_> = opening.leaf_data.iter().flat_map(|x| x.0.to_le_bytes()).collect(); let mut node = pcs::merkle::hash_leaf(&bytes); @@ -1584,7 +1525,7 @@ fn gen_verify( if height >= path_depth { nodes[index] = node; nodes[index ^ 1] = *sibling; - active[index >> 1] = F192::ONE; + active[index >> 1] = F64::ONE; } let (left, right) = if index & 1 == 0 { (&node, sibling) @@ -1592,8 +1533,8 @@ fn gen_verify( (sibling, &node) }; if height < path_depth { - path_children.extend(pack_hash_state(left)); - path_children.extend(pack_hash_state(right)); + path_children.extend(digest_words(left)); + path_children.extend(digest_words(right)); } node = pcs::merkle::hash_pair(left, right); index >>= 1; @@ -1601,7 +1542,7 @@ fn gen_verify( nodes[1] = node; query_hints.push(("merkle_children", path_children)); } - caps.extend(nodes.iter().flat_map(pack_hash_state)); + caps.extend(nodes.iter().flat_map(digest_words)); cap_active.extend(active); } let mut bytecode_row_point = summary.bytecode_claim.point; @@ -1628,25 +1569,22 @@ fn gen_verify( // picks it up at `msg_cursor = cursor`, which sits where the flock // reduction stopped; the ring-switch messages are struct-observed and // still do not advance that cursor. - let mut stream = summary.raw.stream; + let mut stream = limbs(&summary.raw.stream); assert!( stream.len() <= STREAM_CAP, - "stream {} exceeds cap {STREAM_CAP}", + "stream of {} words exceeds cap {STREAM_CAP}", stream.len() ); - stream.resize(STREAM_CAP, F192::ZERO); + stream.resize(STREAM_CAP, F64::ZERO); stream }), - ("bytecode_val", bcv), - ("matpart", vec![matpart]), + ("bytecode_val", limbs(&[bytecode_value])), + ("matpart", limbs(&[matpart])), ("merkle_caps", caps), ("merkle_cap_active", cap_active), // the table sumcheck's round count: max_t tau_t, certified in-guest as a // maximum (one of the taus, and dominating them all). - ( - "zc_tau_max", - vec![F192::new(g_pow(*taus.iter().max().unwrap()).0, 0, 0)], - ), + ("zc_tau_max", vec![g_pow(*taus.iter().max().unwrap())]), ("col_sort_order", col_sort_order), ("sort_order", sort_order), ]; @@ -1666,14 +1604,15 @@ const _: () = assert!(MU_MIN >= lean_vm::pcs::MIN_MU); /// `gen_verify` admits against: one definition, so a hinted shape can never /// outgrow the buffer the guest was compiled with. const MU_CAP: usize = 40; -const STREAM_CAP: usize = 8192; +/// In words, three to a transcript scalar. +const STREAM_CAP: usize = 3 * 8192; /// Named hint entries for a single sub-proof, ordered within each stream. -type SubHints = Vec<(&'static str, Vec)>; +type SubHints = Vec<(&'static str, Vec)>; /// One `hint_witness` stream: a name and its entries, in the order the guest /// pops them. #[derive(Default)] -pub(crate) struct Hints(Vec<(String, Vec>)>); +pub(crate) struct Hints(Vec<(String, Vec>)>); impl Hints { #[cfg(test)] @@ -1681,7 +1620,7 @@ impl Hints { self.0.is_empty() } - fn push(&mut self, name: &str, entry: Vec) { + fn push(&mut self, name: &str, entry: Vec) { match self.0.iter_mut().find(|(n, _)| n == name) { Some((_, entries)) => entries.push(entry), None => self.0.push((name.to_string(), vec![entry])), @@ -1690,7 +1629,7 @@ impl Hints { /// The entries of one stream, for the adversarial test to corrupt. #[cfg(test)] - fn entries(&mut self, name: &str) -> &mut Vec> { + fn entries(&mut self, name: &str) -> &mut Vec> { &mut self .0 .iter_mut() @@ -1950,20 +1889,15 @@ fn push_signature_hints( let wots = &sig.wots_signature; let encoding = xmss::wots_encode(message, xmss_epoch, &pk.public_param, &wots.randomness) .ok_or(AggregationError::MalformedRawSignature)?; - let mut randomness = [0u8; xmss::STATE_LEN]; - randomness[..xmss::RANDOMNESS_LEN].copy_from_slice(&wots.randomness); - hints.push( - "rand", - vec![pack_16_bytes(&randomness[..16]), pack_16_bytes(&randomness[16..])], - ); + hints.push("rand", words(&wots.randomness)); for &e in &encoding { hints.push("digits", vec![count(e as usize)]); } for tip in &wots.chain_tips { - hints.push("chain_starts", vec![pack_16_bytes(tip)]); + hints.push("chain_starts", words(tip)); } for sibling in &sig.merkle_proof { - hints.push("siblings", vec![pack_16_bytes(sibling)]); + hints.push("siblings", words(sibling)); } Ok(()) } @@ -1982,12 +1916,12 @@ fn push_sphincs_hints( sig: &SphincsSignature, ) -> Result<(), AggregationError> { let pp = &pk.public_param; - hints.push("sp_rand", vec![pack_16_bytes(&sig.randomizer)]); + hints.push("sp_rand", words(&sig.randomizer)); let (idx, u) = sphincs::message_digest(pp, &pk.root, &sig.randomizer, message); for kappa in 0..sphincs::NUM_FTS_TREES { - hints.push("sp_fts_secrets", vec![pack_16_bytes(&sig.fts.secrets[kappa])]); + hints.push("sp_fts_secrets", words(&sig.fts.secrets[kappa])); for sibling in &sig.fts.paths[kappa] { - hints.push("sp_fts_paths", vec![pack_16_bytes(sibling)]); + hints.push("sp_fts_paths", words(sibling)); } } let mut signed = sphincs::fts_recover(pp, idx, &u, &sig.fts); @@ -1995,14 +1929,14 @@ fn push_sphincs_hints( let pos = sphincs::Pos::new(lay, sphincs::tree_of(idx, lay), sphincs::leaf_of(idx, lay)); let counter = sig.counters[lay]; let codeword = sphincs::encode(pp, pos, &signed, counter).ok_or(AggregationError::MalformedRawSignature)?; - hints.push("sp_counter", vec![F192::new(u64::from(counter), 0, 0)]); + hints.push("sp_counter", vec![F64(u64::from(counter))]); for (&digit, opened) in codeword.iter().zip(&sig.ots[lay]) { hints.push("sp_digits", vec![count(digit as usize)]); - hints.push("sp_chain_starts", vec![pack_16_bytes(opened)]); + hints.push("sp_chain_starts", words(opened)); } let path = &sig.paths[sphincs::path_range(lay)]; for sibling in path { - hints.push("sp_siblings", vec![pack_16_bytes(sibling)]); + hints.push("sp_siblings", words(sibling)); } let leaf = sphincs::ots_leaf(pp, pos, &signed, counter, &sig.ots[lay]) .ok_or(AggregationError::MalformedRawSignature)?; @@ -2159,13 +2093,7 @@ pub(crate) fn aggregate_tampered( return Err(AggregationError::TooLarge); } let n_sphincs = cover.sphincs_signers.len(); - let group_cells = |group: &XmssClaimGroup| { - [ - F192::new(group.epoch as u64, 0, 0), - pack_16_bytes(&group.message[..16]), - pack_16_bytes(&group.message[16..]), - ] - }; + let group_words = |group: &XmssClaimGroup| [vec![F64(group.epoch as u64)], words(&group.message)].concat(); let mut hints = Hints::default(); hints.push( @@ -2181,14 +2109,14 @@ pub(crate) fn aggregate_tampered( ], ); let fs_seed = lean_vm::cpu::fs_seed(guest); - hints.push("fs_seed", vec![fs_seed[0], fs_seed[1]]); - // Per group: its epoch, its two message cells, and its declared, duplicate + hints.push("fs_seed", fs_seed.to_vec()); + // Per group: its epoch, its four message words, and its declared, duplicate // and raw-signature counts, in the guest's geometry-pass order. The keys // then ride two per `pubkeys` entry, so the guest can halve its loop // frames, the odd key out on a final one-key entry; each group's // duplicates follow its keys. for (j, group) in cover.xmss_groups.iter().enumerate() { - let mut entry = group_cells(group).to_vec(); + let mut entry = group_words(group); entry.extend([ count(group.keys.len()), count(cover.xmss_dups[j].len()), @@ -2200,16 +2128,12 @@ pub(crate) fn aggregate_tampered( hints.push("pk_halves", vec![count(keys.len() / 2), count(keys.len() % 2)]); hints.push("signers_split", signers_split(keys.len().div_ceil(2))); for pair in keys.chunks(2) { - let mut entry = key_cells(&pair[0]).to_vec(); - if let Some(second) = pair.get(1) { - entry.extend_from_slice(&key_cells(second)); - } - hints.push("pubkeys", entry); + hints.push("pubkeys", pair.iter().flat_map(key_words).collect()); } } for dups in &cover.xmss_dups { for pk in dups { - hints.push("dup_pubkeys", key_cells(pk).to_vec()); + hints.push("dup_pubkeys", key_words(pk)); } } if !cover.sphincs_signers.is_empty() { @@ -2217,10 +2141,10 @@ pub(crate) fn aggregate_tampered( } hints.push("signers_split", signers_split(1 + 2 * cover.n_declared)); for signer in &cover.sphincs_signers { - hints.push("sphincs_signers", sphincs_signer_cells(signer).to_vec()); + hints.push("sphincs_signers", sphincs_signer_words(signer)); } for signer in &cover.sphincs_dups { - hints.push("dup_sphincs", sphincs_signer_cells(signer).to_vec()); + hints.push("dup_sphincs", sphincs_signer_words(signer)); } // Group-major over the table, not over the epochs `raw_xmss` is sorted by: a // declaration puts the undeclared groups last. Each index is an offset within @@ -2245,7 +2169,7 @@ pub(crate) fn aggregate_tampered( vec![count(child.xmss_signers.len()), count(child.sphincs_signers.len())], ); for (group, (parent_group, offsets)) in child.xmss_signers.iter().zip(&cover.child_xmss[i]) { - let mut entry = group_cells(group).to_vec(); + let mut entry = group_words(group); entry.push(count(group.keys.len())); hints.push("child_group", entry); hints.push("child_group_map", vec![count(*parent_group)]); @@ -2262,7 +2186,7 @@ pub(crate) fn aggregate_tampered( hints.push("child_sphincs_index", vec![count(offset)]); } hints.push("signers_split", signers_split(1 + 2 * child.xmss_signers.len())); - hints.push("child_defer", child.defer.cells()); + hints.push("child_defer", limbs(&child.defer.cells())); hints.push("child_da_count", vec![count(child.da_roots.len())]); let (sub_hints, defer) = gen_verify(guest, pi, summary)?; for (name, entry) in sub_hints { @@ -2278,7 +2202,7 @@ pub(crate) fn aggregate_tampered( let leaf = DeferredClaim::leaf(); hints.push( "leaf_defer", - vec![leaf.bytecode_value, leaf.matrix_a_value, leaf.matrix_b_value], + limbs(&[leaf.bytecode_value, leaf.matrix_a_value, leaf.matrix_b_value]), ); leaf } else { @@ -2300,7 +2224,7 @@ pub(crate) fn aggregate_tampered( .as_chunks::() .0 { - hints.push("da_weights", block.to_vec()); + hints.push("da_weights", limbs(block)); } hints.push( "da_shape", @@ -2315,7 +2239,7 @@ pub(crate) fn aggregate_tampered( "da_symbols", witness.codewords[i * m + j * c..i * m + (j + 1) * c] .iter() - .map(|&w| F192::from(F64(w))) + .map(|&w| F64(w)) .collect(), ); } @@ -2348,7 +2272,7 @@ pub(crate) fn aggregate_tampered( } hints.push("da_meta", vec![count(da_roots.len()), count(da_dups.len())]); for root in da_roots.iter().chain(&da_dups) { - hints.push("da_roots", da_claim_cells(root)); + hints.push("da_roots", da_claim_words(root)); } let public_input = statement_digest( @@ -2377,7 +2301,7 @@ pub(crate) fn aggregate_tampered( } struct CoordinateDescriptor { kind: usize, - constant: u128, + constant: u64, fresh: usize, claim_slot: usize, terms: Range, @@ -2385,7 +2309,7 @@ struct CoordinateDescriptor { struct Term { kind: usize, - constant: u128, + constant: u64, column_a: usize, column_b: usize, } @@ -2421,8 +2345,8 @@ struct OpeningShape { vanish_offsets: Vec, fold_offsets: Vec, residual_fold_offsets: Vec, - vanish_values: Vec, - vanish_inverses: Vec, + vanish_values: Vec, + vanish_inverses: Vec, ood_samples: Vec, } @@ -2431,13 +2355,8 @@ struct OpeningShape { fn placeholder_map(kbc: usize) -> BTreeMap { // Only block and coordinate structure is used here; dummy instructions and // table sizes let us derive it before the guest's bytecode exists. - let stand_in = vec![lean_vm::cpu::Op::Xor { a: 0, b: 0, c: 0 }; 1 << kbc]; - let layout = lean_vm::cpu::layout( - &stand_in, - 20, - [1usize << 10; lean_vm::tables::N_TABLES], - [F192::ZERO, F192::ZERO], - ); + let stand_in = vec![lean_vm::cpu::Op::Xor64 { a: 0, b: 0, c: 0 }; 1 << kbc]; + let layout = lean_vm::cpu::layout(&stand_in, 20, [1usize << 10; lean_vm::tables::N_TABLES], [F64::ZERO; 4]); let sides: [&[Block]; 3] = [&layout.push, &layout.pull, &layout.count]; let lcrounds = flock::hash::K_LOG - 6; @@ -2488,7 +2407,7 @@ fn placeholder_map(kbc: usize) -> BTreeMap { nbcv += usize::from(matches!(c, Coord::Public(_))); coordinates.push(CoordinateDescriptor { kind: coord_kind(c), - constant: dsl_u128(coord_scale(c)), + constant: coord_scale(c).0, fresh, claim_slot: slot, terms: start..terms.len(), @@ -2498,7 +2417,7 @@ fn placeholder_map(kbc: usize) -> BTreeMap { sblk.push(nblocks); } let evtot: usize = lean_vm::tables::tables().iter().map(|t| t.n_committed_columns()).sum(); - let ncl = nclaims + evtot + 3; // bus + constraint + the three PI memory-limb claims + let ncl = nclaims + evtot + 1; // bus + constraint + the PI memory claim // ---- claim descriptor buffer ids (structural) ---- let valcols = blake2s_value_columns(); @@ -2539,7 +2458,7 @@ fn placeholder_map(kbc: usize) -> BTreeMap { 0 }, }, - ClaimSite::MemoryLimb { column } => ClaimDescriptor { + ClaimSite::PublicInput { column } => ClaimDescriptor { buffer: 2, column: compact_col_pm[column], qflock_slot: 0, @@ -2550,19 +2469,15 @@ fn placeholder_map(kbc: usize) -> BTreeMap { assert_eq!(claims.len(), ncl, "descriptor count == pool size"); // ---- the placeholder map ---- - let flds = |v: &[F192]| { - format!( - "[{}]", - v.iter().map(|&x| f192_literal(x)).collect::>().join(", ") - ) - }; + // A field-element array, flattened to three words an element. + let flds = |v: &[F192]| literals(limbs(v).iter().map(|w| w.0)); let mut rep = BTreeMap::new(); let mut ps = |k: &str, v: String| { rep.insert(format!("{k}_PLACEHOLDER"), v); }; ps("STREAM_CAP", STREAM_CAP.to_string()); ps("MIN_LOG_MEM", lean_vm::cpu::MIN_LOG_MEM.to_string()); - ps("INV_GEN", dsl_u128(F192::new(G.inv().0, 0, 0)).to_string()); + ps("INV_GEN", G.inv().0.to_string()); ps("MU_CAP", MU_CAP.to_string()); ps("NO_TABLE", layout.taus.len().to_string()); ps("GKR_ROUNDS_CAP", (MU_CAP * (MU_CAP + 1) / 2 + MU_CAP + 2).to_string()); @@ -2602,13 +2517,13 @@ fn placeholder_map(kbc: usize) -> BTreeMap { ps("TERM_COL_A", literals(terms.iter().map(|t| t.column_a))); ps("TERM_COL_B", literals(terms.iter().map(|t| t.column_b))); ps("N_BUS_CLAIMS", nclaims.to_string()); - let idxc: Vec = (0..34) + let idxc: Vec = (0..34) .map(|i| { - let mut g2k = F192::new(G.0, 0, 0); + let mut g2k = G; for _ in 0..i { g2k = g2k * g2k; } - dsl_u128(F192::ONE + g2k) + (F64::ONE + g2k).0 }) .collect(); ps("INDEX_MLE_FACTORS", literals(&idxc)); @@ -2641,23 +2556,6 @@ fn placeholder_map(kbc: usize) -> BTreeMap { ps("K_SKIP", "6".to_string()); ps("N_FIXED_CHALLENGE_ROUNDS", fixed_challenges.len().to_string()); ps("PHI8_NODES", flds(&primitives::field::PHI_8_TABLE_192[..128])); - // Tower F192 = F64[Y]/(Y^3+Y+1), Y = new(0,1,0). Y_TOWER embeds Y for - // AIR lane reassembly; Y_INV helps derive the top PI-memory limb. - let y_tower = F192::new(0, 1, 0); - ps("Y_TOWER", dsl_u128(y_tower).to_string()); - ps("Y_INV", f192_literal(y_tower.inv())); - // Coordinate basis e_i of F192 over F2 (spans the whole field): the 64 - // binary basis vectors in each of the three tower limbs. The guest uses - // these vectors to reconstruct a word from its 192 coordinate bits. - let coord_basis: Vec = (0..192) - .map(|i| match i / 64 { - 0 => F192::new(1u64 << i, 0, 0), - 1 => F192::new(0, 1u64 << (i - 64), 0), - 2 => F192::new(0, 0, 1u64 << (i - 128)), - _ => unreachable!(), - }) - .collect(); - ps("COORD_BASIS", flds(&coord_basis)); // One constant per domain, not one per node: every barycentric denominator over an aligned φ₈ // window is the same element (`primitives::multilinear::window_denominator`). ps( @@ -2717,12 +2615,8 @@ fn placeholder_map(kbc: usize) -> BTreeMap { let mut c_ivk = Vec::new(); for &cl_lv in cl.iter().take(cn) { for &v in &pcs::whir::eval_sk_at_vks(cl_lv) { - c_svk.push(F192::new(v.0, 0, 0)); - c_ivk.push(if v == F64::ZERO { - F192::ZERO - } else { - F192::new(v.inv().0, 0, 0) - }); + c_svk.push(v); + c_ivk.push(if v == F64::ZERO { F64::ZERO } else { v.inv() }); } } OpeningShape { @@ -2828,7 +2722,6 @@ fn placeholder_map(kbc: usize) -> BTreeMap { .max() .unwrap(); ps("LIG_ROW_CAP", row_cap.to_string()); - ps("LIG_PACKED_ROW_CAP", (row_cap / 2).to_string()); ps( "LIG_PATH_CAP", cands @@ -2837,7 +2730,7 @@ fn placeholder_map(kbc: usize) -> BTreeMap { c.tree_depths .iter() .zip(&c.cap_depths) - .map(|(depth, cap)| 4 * (depth - cap)) + .map(|(depth, cap)| 8 * (depth - cap)) }) .max() .unwrap() @@ -2887,21 +2780,18 @@ fn placeholder_map(kbc: usize) -> BTreeMap { let padded_len = svk2.len() + maxsvk; svk2.extend_from_slice(&candidate.vanish_values); ivk2.extend_from_slice(&candidate.vanish_inverses); - svk2.resize(padded_len, F192::ZERO); - ivk2.resize(padded_len, F192::ZERO); + svk2.resize(padded_len, F64::ZERO); + ivk2.resize(padded_len, F64::ZERO); } - ps("LIG_VANISH_VALS", flds(&svk2)); - ps("LIG_VANISH_INVS", flds(&ivk2)); + ps("LIG_VANISH_VALS", literals(svk2.iter().map(|v| v.0))); + ps("LIG_VANISH_INVS", literals(ivk2.iter().map(|v| v.0))); } let n_log_sizes = maxm - minm + 1; let n_rates = MAX_LOG_INV_RATE - MIN_LOG_INV_RATE + 1; ps("LIG_N_LOG_SIZES", n_log_sizes.to_string()); ps("LIG_N_RATES", n_rates.to_string()); ps("LIG_N_CANDIDATES", (n_log_sizes * n_rates).to_string()); - ps( - "LIG_MIN_SHIFT_INV", - dsl_u128(F192::new(g_pow(minm).inv().0, 0, 0)).to_string(), - ); + ps("LIG_MIN_SHIFT_INV", g_pow(minm).inv().0.to_string()); ps("CLAIM_POINT_BUF", literals(claims.iter().map(|c| c.buffer))); ps("CLAIM_COMMITTED_COL", literals(claims.iter().map(|c| c.column))); let slot_stride_log = lean_vm::hash_flock::SLOT_STRIDE_LOG; @@ -2923,28 +2813,24 @@ fn placeholder_map(kbc: usize) -> BTreeMap { ps("LOG2_BYTECODE_COLS", log2_bc_cols.to_string()); ps("DEFER_SIZE", (kbc + log2_bc_cols + 2 * lcrounds + 68).to_string()); ps("BYTECODE_VARS", (kbc + log2_bc_cols).to_string()); - let agg_state = pack_state(FiatShamirState::from_label(RECURSION_AGG_LABEL).state()); - ps("AGG_SEED_0", dsl_u128(agg_state[0]).to_string()); - ps("AGG_SEED_1", dsl_u128(agg_state[1]).to_string()); + let agg_state = FiatShamirState::from_label(RECURSION_AGG_LABEL).state(); + for (i, word) in agg_state.iter().enumerate() { + ps(&format!("AGG_SEED_{i}"), word.0.to_string()); + } // ---- LeanDA (`doc/leanvm` §sec:leanda) ---- let (pad_cell, pad_row) = lean_da::padding_digests(); - let (pad_cell, pad_row) = (pack_hash_state(&pad_cell), pack_hash_state(&pad_row)); ps("DA_LOG_K", DA_LOG_K.to_string()); ps("DA_LOG_CELL", DA_LOG_CELL.to_string()); ps("DA_MAX_ROWS", DA_MAX_ROWS.to_string()); ps("DA_LOG_MAX_ROWS", DA_MAX_ROWS.ilog2().to_string()); - ps("DA_PAD_CELL_0", f192_literal(pad_cell[0])); - ps("DA_PAD_CELL_1", f192_literal(pad_cell[1])); - ps("DA_PAD_ROW_0", f192_literal(pad_row[0])); - ps("DA_PAD_ROW_1", f192_literal(pad_row[1])); + ps("DA_PAD_CELL", literals(digest_words(&pad_cell).iter().map(|w| w.0))); + ps("DA_PAD_ROW", literals(digest_words(&pad_row).iter().map(|w| w.0))); let defer_cells = kbc + log2_bc_cols + 1 + 2 * flock::hash::K_LOG + 2; ps("STMT_HEADER", STATEMENT_HEADER.to_string()); - let (off, pairs) = (STATEMENT_HEADER, defer_cells.div_ceil(2)); - let blocks = (off + 3 * pairs).div_ceil(4); - ps("STMT_ODD", (defer_cells % 2).to_string()); - ps("STMT_PAIRS", pairs.to_string()); - ps("STMT_PAD_CELLS", (4 * blocks - off - 3 * pairs).to_string()); + let stmt_words = STATEMENT_HEADER + 3 * defer_cells; + let blocks = stmt_words.div_ceil(8); + ps("STMT_PAD", (8 * blocks - stmt_words).to_string()); ps("STMT_BLOCKS", blocks.to_string()); // A list is at most MAX_KEYS blocks (one a claim is the widest it gets), so it // holds fewer than that many windows; a declared count is below MAX_KEYS, hence @@ -2954,12 +2840,11 @@ fn placeholder_map(kbc: usize) -> BTreeMap { ps("SIGNERS_WINDOW_LOG", SIGNERS_WINDOW.ilog2().to_string()); ps("SIGNERS_MAX_WINDOWS", SIGNERS_MAX_WINDOWS.to_string()); ps("SIGNERS_COUNT_BITS", SIGNERS_COUNT_BITS.to_string()); - ps("BLAKE2S_IV_0", dsl_u128(lean_vm::hash_flock::IV_CELLS[0]).to_string()); - ps("BLAKE2S_IV_1", dsl_u128(lean_vm::hash_flock::IV_CELLS[1]).to_string()); - ps( - "MD_FINAL", - dsl_u128(lean_vm::hash_flock::metadata(0, lean_vm::hash_flock::FINAL_FLAG, 0)).to_string(), - ); + for (i, word) in lean_vm::hash_flock::IV.iter().enumerate() { + ps(&format!("BLAKE2S_IV_{i}"), word.0.to_string()); + } + let final_md = lean_vm::hash_flock::metadata(0, lean_vm::hash_flock::FINAL_FLAG, 0); + ps("MD_FINAL", final_md[1].0.to_string()); // The XMSS instance, from which the guest derives every table width by // compile-time integer arithmetic. @@ -2967,31 +2852,27 @@ fn placeholder_map(kbc: usize) -> BTreeMap { ps("W", xmss::W.to_string()); ps("TARGET_SUM", xmss::TARGET_SUM.to_string()); ps("LOG_LIFETIME", xmss::LOG_LIFETIME.to_string()); - // Every XMSS tweak the guest builds is one of these constants plus the - // epoch's weighed bits, so the byte layout lives in `xmss::make_tweak` and - // nowhere else. The chain table is indexed `CHAIN_STEPS * i + s` and the - // Merkle one by level, exactly as `verify_sig` walks them. - ps( - "XM_ENC_TWEAK", - dsl_u128(tweak_cell(xmss::TWEAK_TYPE_ENCODING, 0)).to_string(), + // Every XMSS tweak the guest builds is one of these first words plus a + // sub-position, next to the epoch's weighed bits, so the byte layout lives in + // `xmss::make_tweak` and nowhere else. A chain's tweaks start at sub-position + // `CHAIN_LENGTH * i` and a Merkle level's at `level + 1`, exactly as + // `verify_sig` walks them. + ps("XM_ENC_TWEAK", tweak_word(xmss::TWEAK_TYPE_ENCODING, 0).0.to_string()); + ps("XM_PK_TWEAK", tweak_word(xmss::TWEAK_TYPE_WOTS_PK, 0).0.to_string()); + ps("XM_CHAIN_TWEAK", tweak_word(xmss::TWEAK_TYPE_CHAIN, 0).0.to_string()); + ps("XM_MERKLE_TWEAK", tweak_word(xmss::TWEAK_TYPE_MERKLE, 0).0.to_string()); + let p_mul = tweak_word(0, 1) + tweak_word(0, 0); + assert!( + (0..(xmss::V * xmss::CHAIN_LENGTH + xmss::LOG_LIFETIME) as u32).all(|p| { + tweak_word(xmss::TWEAK_TYPE_CHAIN, p) == tweak_word(xmss::TWEAK_TYPE_CHAIN, 0) + F64(u64::from(p)) * p_mul + }), + "a tweak's sub-position is a whole number of units in its first word" ); + ps("XM_P_MUL", p_mul.0.to_string()); ps( - "XM_PK_TWEAK", - dsl_u128(tweak_cell(xmss::TWEAK_TYPE_WOTS_PK, 0)).to_string(), + "XM_INDEX_WEIGHT", + literals((0..xmss::LOG_LIFETIME).map(|b| tweak_index_weight(b).0)), ); - let chain_tweaks: Vec = (0..xmss::V) - .flat_map(|i| { - (0..xmss::CHAIN_LENGTH - 1) - .map(move |s| tweak_cell(xmss::TWEAK_TYPE_CHAIN, (i * xmss::CHAIN_LENGTH + s) as u32)) - }) - .collect(); - ps("XM_CHAIN_TWEAKS", flds(&chain_tweaks)); - let merkle_tweaks: Vec = (0..xmss::LOG_LIFETIME) - .map(|level| tweak_cell(xmss::TWEAK_TYPE_MERKLE, (level + 1) as u32)) - .collect(); - ps("XM_MERKLE_TWEAKS", flds(&merkle_tweaks)); - let index_weights: Vec = (0..xmss::LOG_LIFETIME).map(tweak_index_weight).collect(); - ps("XM_INDEX_WEIGHT", flds(&index_weights)); ps("MAX_KEYS", MAX_KEYS.to_string()); ps("MAX_DA_ROOTS", MAX_DA_ROOTS.to_string()); ps("DA_ROOT_COUNTS", (MAX_DA_ROOTS + 1).to_string()); @@ -3024,10 +2905,7 @@ fn placeholder_map(kbc: usize) -> BTreeMap { ("SP_TW_FTS_ROOTS", sphincs::TWEAK_FTS_ROOTS), ("SP_TW_MSG", sphincs::TWEAK_MSG), ] { - ps( - name, - dsl_u128(pack_16_bytes(&sphincs::tweak(tag, 0, 0, 0, 0))).to_string(), - ); + ps(name, words(&sphincs::tweak(tag, 0, 0, 0, 0))[0].0.to_string()); } rep } @@ -3287,13 +3165,13 @@ mod tests { let sphincs_tweak = sphincs::tweak(sphincs_tag, 0, 0, position, index); assert_eq!(&xmss_tweak[1..], &sphincs_tweak[1..]); assert_ne!(xmss_tweak[0], sphincs_tweak[0]); - let mut guest_tweak = tweak_cell(xmss_tag, position); + let mut guest_index = F64::ZERO; for bit in 0..32 { if index & (1 << bit) != 0 { - guest_tweak += tweak_index_weight(bit); + guest_index += tweak_index_weight(bit); } } - assert_eq!(guest_tweak, pack_16_bytes(&xmss_tweak)); + assert_eq!(words(&xmss_tweak), [tweak_word(xmss_tag, position), guest_index]); } } } @@ -3410,25 +3288,25 @@ mod tests { let source = format!( r#"{helpers} def main(): - point = HeapBuf(MAX_STACK_LOG) - hint_witness(point[0:MAX_STACK_LOG], "point") + point = HeapBuf(3 * MAX_STACK_LOG) + hint_witness(point[0:3 * MAX_STACK_LOG], "point") offset = hint_witness("offset") kappa_g = hint_witness("kappa") weight = match(log(kappa_g), range(0, N_COLUMN_LOGS), lambda kappa: column_selector(offset, point, kappa)) public = GEN ** 0 - assert public[1] == weight + public[0:3] = weight return "# ); let guest = compile(&parse_with_replacements(&source, &placeholder_map(18)).unwrap()); let run = |point: &[F192], offset: usize, kappa: usize, expected: F192| { let mut hints = Hints::default(); - hints.push("point", point.to_vec()); + hints.push("point", limbs(point)); hints.push("offset", vec![count(offset)]); hints.push("kappa", vec![count(kappa)]); let mut program = guest.clone(); hints.install(&mut program); - program.execute([expected, F192::ZERO]) + program.execute([F64(expected.c0), F64(expected.c1), F64(expected.c2), F64::ZERO]) }; let mut rng = StdRng::seed_from_u64(813); for mu in MU_MIN..=MU_MAX { @@ -3470,12 +3348,11 @@ def main(): for k in unroll(0, 3): bits[k] = bits[k] * bits[k] direction = addr(bits) - leaf = StackBuf(2) - hint_witness(leaf[0:2], "leaf") - a, b = verify_merkle_path(leaf[0], leaf[1], direction, 3) + leaf = StackBuf(4) + hint_witness(leaf[0:4], "leaf") + root = verify_merkle_path(leaf, direction, 3) public = GEN ** 0 - assert public[1] == a - assert public[GEN] == b + public[0:4] = root return "# ); @@ -3489,18 +3366,15 @@ def main(): } let run = |index: usize, leaf: [u8; 32], pairs: &[[[u8; 32]; 2]], root: [u8; 32]| { let mut hints = Hints::default(); - hints.push( - "bits", - (0..3).map(|k| F192::from(F64(((index >> k) & 1) as u64))).collect(), - ); - hints.push("leaf", pack_hash_state(&leaf).to_vec()); + hints.push("bits", (0..3).map(|k| F64(((index >> k) & 1) as u64)).collect()); + hints.push("leaf", digest_words(&leaf).to_vec()); hints.push( "merkle_children", - pairs.iter().flatten().flat_map(pack_hash_state).collect(), + pairs.iter().flatten().flat_map(digest_words).collect(), ); let mut program = guest.clone(); hints.install(&mut program); - program.execute(pack_hash_state(&root)) + program.execute(digest_words(&root)) }; for index in 0..8 { let pairs: Vec<_> = (0..3) @@ -3516,7 +3390,7 @@ def main(): ); for level in 0..3 { for side in 0..2 { - for byte in [0, 16] { + for byte in [0, 8, 16, 24] { let mut forged = pairs.clone(); forged[level][side][byte] ^= 1; assert!(std::panic::catch_unwind(|| run(index, tree[8 + index], &forged, tree[1])).is_err()); @@ -3524,7 +3398,7 @@ def main(): } // Rehash a forged running child all the way to a matching public root. // Root equality alone passes; the selected child must still bind to its predecessor. - for byte in [0, 16] { + for byte in [0, 8, 16, 24] { let mut forged = pairs.clone(); forged[level][(index >> level) & 1][byte] ^= 1; let mut root = pcs::merkle::hash_pair(&forged[level][0], &forged[level][1]); @@ -3807,7 +3681,7 @@ def main(): // The first child's omitted root occupies slot 1; it cannot cover slot 0 // (the second child's declared root) or write outside the DA region. - for index in [count(0), count(2), count(MAX_KEYS - 1), F192::ZERO, F192::new(0, 1, 0)] { + for index in [count(0), count(2), count(MAX_KEYS - 1), F64::ZERO] { let outcome = std::panic::catch_unwind(|| { aggregate_tampered( &children, @@ -3859,7 +3733,7 @@ def main(): LOG_INV_RATE, |h| { h.entries("da_meta")[0] = vec![count(declared), count(duplicates)]; - h.entries("da_roots").push(da_claim_cells(&second)); + h.entries("da_roots").push(da_claim_words(&second)); }, ) }); @@ -3881,9 +3755,9 @@ def main(): let root = h .entries("da_roots") .iter_mut() - .find(|r| **r == da_claim_cells(&second)) + .find(|r| **r == da_claim_words(&second)) .unwrap(); - *root = da_claim_cells(&first); + *root = da_claim_words(&first); }, ) }); @@ -3892,7 +3766,7 @@ def main(): "every child's complete DA list must be authenticated" ); } - for limb in [2, 3] { + for word in 4..8 { let outcome = std::panic::catch_unwind(|| { aggregate_tampered( &children, @@ -3908,9 +3782,9 @@ def main(): let omitted = h .entries("da_roots") .iter_mut() - .find(|claim| claim[..2] == pack_hash_state(&first)) + .find(|claim| claim[..4] == digest_words(&first)) .unwrap(); - omitted[limb] += F192::ONE; + omitted[word] += F64::ONE; }, ) }); @@ -3989,14 +3863,13 @@ def main(): def main(): n_g = hint_witness("n") assert log(n_g) < MAX_DA_ROOTS + 1 - roots = HeapBuf((n_g * GEN) ** 4) + roots = HeapBuf((n_g * GEN) ** 8) for x in mul_range(1, n_g): - root = roots * (x ** 4) - hint_witness(root[0:4], "root") - a, b = da_list_digest(roots, n_g) + root = roots * (x ** 8) + hint_witness(root[0:8], "root") + digest = da_list_digest(roots, n_g) public = GEN ** 0 - assert public[1] == a - assert public[GEN] == b + public[0:4] = digest return "# ); @@ -4008,11 +3881,11 @@ def main(): let mut hints = Hints::default(); hints.push("n", vec![count(n)]); for root in &roots { - hints.push("root", da_claim_cells(root)); + hints.push("root", da_claim_words(root)); } let mut program = guest.clone(); hints.install(&mut program); - let public = pack_hash_state(&da_list_digest(&roots)); + let public = digest_words(&da_list_digest(&roots)); if n <= MAX_DA_ROOTS { let execution = program.execute(public); assert!(execution.unconstrained_reads.is_empty(), "{n} roots"); @@ -4031,19 +3904,18 @@ def main(): let source = format!( r#"{helpers} def main(): - roots = HeapBuf(12) - hint_witness(roots[0:12], "roots") + roots = HeapBuf(24) + hint_witness(roots[0:24], "roots") n_slots = hint_witness("n_slots") assert log(n_slots) < 3 cover = HeapBuf(4) # Adjacent signature slots must stay outside the DA writer's range. cover[1] = 1 - a, b, v0, v1 = cover_da_root(roots, cover * GEN, n_slots, GEN) - digest = StackBuf(2) - blake2s([a, b], [v0, v1], digest) + claim = cover_da_root(roots, cover * GEN, n_slots, GEN) + digest = StackBuf(4) + blake2s(claim[0:4], claim[4:8], digest) public = GEN ** 0 - assert public[1] == digest[0] - assert public[GEN] == digest[1] + public[0:4] = digest return "# ); @@ -4051,14 +3923,14 @@ def main(): let first = [0x13; 32]; let second = [0x27; 32]; let outside = [0x39; 32]; - let run = |slots: usize, index: F192, claimed: [u8; 32]| { + let run = |slots: usize, index: F64, claimed: [u8; 32]| { let mut hints = Hints::default(); hints.push( "roots", [ - da_claim_cells(&first), - da_claim_cells(&second), - da_claim_cells(&outside), + da_claim_words(&first), + da_claim_words(&second), + da_claim_words(&outside), ] .concat(), ); @@ -4066,7 +3938,7 @@ def main(): hints.push("da_index", vec![index]); let mut program = guest.clone(); hints.install(&mut program); - program.execute(pack_hash_state(&da_list_digest(&[claimed]))) + program.execute(digest_words(&da_list_digest(&[claimed]))) }; assert!(run(2, count(0), first).unconstrained_reads.is_empty()); assert!(run(2, count(1), second).unconstrained_reads.is_empty()); @@ -4076,8 +3948,7 @@ def main(): (0, count(0), first), (2, count(2), outside), (2, count(1).inv(), first), - (2, F192::ZERO, first), - (2, F192::new(0, 1, 0), first), + (2, F64::ZERO, first), ] { assert!(std::panic::catch_unwind(|| run(slots, index, claimed)).is_err()); } @@ -4089,14 +3960,14 @@ def main(): let source = include_str!("../guests/lean_ethereum.py"); let (helpers, _) = source.split_once("\ndef main():").unwrap(); let source = format!( - "{helpers}\ndef main():\n _, squares = exponent_tables()\n a, b, v0, v1 = da_verify(squares)\n digest = StackBuf(2)\n blake2s([a, b], [v0, v1], digest)\n public = GEN ** 0\n assert public[1] == digest[0]\n assert public[GEN] == digest[1]\n return\n" + "{helpers}\ndef main():\n _, squares = exponent_tables()\n root, vector = da_verify(squares)\n digest = StackBuf(4)\n blake2s(root, vector, digest)\n public = GEN ** 0\n public[0:4] = digest\n return\n" ); let guest = compile(&parse_with_replacements(&source, &placeholder_map(20)).unwrap()); let n_rows = 3usize; let codewords = lean_da::encode_rows(&da_rows(3, 101)); - let run = |words: &[u64], tamper: &dyn Fn(&mut Hints, &mut [F192; 2])| { - let (commitment, _) = lean_da::commit_codewords(words.to_vec()); - let mut public = pack_hash_state(&da_list_digest(&[commitment.root])); + let run = |symbols: &[u64], tamper: &dyn Fn(&mut Hints, &mut [F64; 4])| { + let (commitment, _) = lean_da::commit_codewords(symbols.to_vec()); + let mut public = digest_words(&da_list_digest(&[commitment.root])); let mut hints = Hints::default(); hints.push( "da_shape", @@ -4107,10 +3978,7 @@ def main(): let start = i * CODEWORD_SYMBOLS + j * CELL_SYMBOLS; hints.push( "da_symbols", - words[start..start + CELL_SYMBOLS] - .iter() - .map(|&w| F192::from(F64(w))) - .collect(), + symbols[start..start + CELL_SYMBOLS].iter().map(|&w| F64(w)).collect(), ); } } @@ -4118,7 +3986,7 @@ def main(): .as_chunks::() .0 { - hints.push("da_weights", block.to_vec()); + hints.push("da_weights", limbs(block)); } tamper(&mut hints, &mut public); let mut program = guest.clone(); @@ -4130,10 +3998,10 @@ def main(): // Every vector is orthogonal to zero rows: only hashing can reject these changes. let zeros = vec![0; codewords.len()]; - for limb in [F192::ONE, F192::new(0, 1, 0), F192::new(0, 0, 1)] { + for limb in 0..3 { assert!( std::panic::catch_unwind(|| run(&zeros, &|h, _| { - h.entries("da_weights")[0][0] += limb; + h.entries("da_weights")[0][limb] += F64::ONE; })) .is_err(), "every limb of L must be bound by its hash" @@ -4142,7 +4010,7 @@ def main(): assert!( std::panic::catch_unwind(|| run(&zeros, &|h, _| { for block in h.entries("da_weights") { - block.fill(F192::ZERO); + block.fill(F64::ZERO); } })) .is_err(), @@ -4172,20 +4040,14 @@ def main(): assert_ne!(forged_digest, da_list_digest(&[bad_root])); let unchecked = run(&bad, &|h, public| { for block in h.entries("da_weights") { - block.fill(F192::ZERO); + block.fill(F64::ZERO); } - *public = pack_hash_state(&forged_digest); + *public = digest_words(&forged_digest); }); assert!(unchecked.unconstrained_reads.is_empty()); assert!( std::panic::catch_unwind(|| run(&codewords, &|_, public| { - public[0] += F192::ONE; - })) - .is_err() - ); - assert!( - std::panic::catch_unwind(|| run(&codewords, &|h, _| { - h.entries("da_symbols")[0][0] += F192::new(0, 1, 0); + public[0] += F64::ONE; })) .is_err() ); @@ -4197,7 +4059,7 @@ def main(): assert!( std::panic::catch_unwind(|| run(&codewords, &|h, public| { h.entries("da_shape")[0][1] = count(3); - *public = pack_hash_state(&da_list_digest(&[padded_commitment.root])); + *public = digest_words(&da_list_digest(&[padded_commitment.root])); })) .is_err(), "three rows must not use an eight-row tree, even with a matching root" @@ -4240,7 +4102,8 @@ def main(): counts.dedup(); for rows in counts { for depth in 0..=max_depth + 1 { - let result = std::panic::catch_unwind(|| guest.execute([count(rows), count(depth)])); + let result = + std::panic::catch_unwind(|| guest.execute([count(rows), count(depth), F64::ZERO, F64::ZERO])); let valid = (1..=DA_MAX_ROWS).contains(&rows) && rows.next_power_of_two() == 1 << depth; assert_eq!(result.is_ok(), valid, "rows={rows}, depth={depth}"); if let Ok(execution) = result { @@ -4248,8 +4111,11 @@ def main(): } } } - for bad in [F192::ZERO, F192::new(0, 1, 0), count(1).inv()] { - for public in [[bad, count(0)], [count(1), bad]] { + for bad in [F64::ZERO, count(1).inv()] { + for public in [ + [bad, count(0), F64::ZERO, F64::ZERO], + [count(1), bad, F64::ZERO, F64::ZERO], + ] { assert!(std::panic::catch_unwind(|| guest.execute(public)).is_err()); } } @@ -4276,12 +4142,11 @@ def main(): let expected = value.max(1).next_power_of_two().ilog2() as usize; for depth in 0..=9 { let mut program = guest.clone(); - program.set_witness( - "bits", - vec![(0..8).map(|j| F192::from(F64(((value >> j) & 1) as u64))).collect()], - ); + program.set_witness("bits", vec![(0..8).map(|j| F64(((value >> j) & 1) as u64)).collect()]); program.set_witness("ceil_log", vec![vec![count(depth)]]); - let result = std::panic::catch_unwind(|| program.execute([count(value), count(depth)])); + let result = std::panic::catch_unwind(|| { + program.execute([count(value), count(depth), F64::ZERO, F64::ZERO]) + }); assert_eq!( result.is_ok(), depth == expected.max(floor), @@ -4853,13 +4718,14 @@ def main(): h.entries("raw_index")[0] = vec![count(SMALL_LEAF_SIZE)]; }), ("group (n_xmss inflated)", &|h: &mut Hints| { - h.entries("group")[0][3] = count(SMALL_LEAF_SIZE + 1); + h.entries("group")[0][5] = count(SMALL_LEAF_SIZE + 1); }), ("group (n_raw_xmss understated)", &|h: &mut Hints| { - h.entries("group")[0][5] = count(SMALL_LEAF_SIZE - 1); + h.entries("group")[0][7] = count(SMALL_LEAF_SIZE - 1); }), ("group (a spurious duplicate slot)", &|h: &mut Hints| { - h.entries("group")[0][4] = count(1); + h.entries("group")[0][6] = count(1); + h.push("dup_pubkeys", vec![F64::ZERO; 4]); }), // One more group than the hint stream carries: witness generation // has nothing to pop for it. @@ -4867,7 +4733,7 @@ def main(): h.entries("meta")[0][0] = count(2); }), ("pubkeys (a key nobody signed for)", &|h: &mut Hints| { - h.entries("pubkeys")[0][0] += F192::ONE; + h.entries("pubkeys")[0][0] += F64::ONE; }), // The window split of a list hash is advice, so both halves are pinned: // the product identity ties them to the block count, and the tail's own @@ -4882,23 +4748,23 @@ def main(): h.entries("signers_split")[0][1] = count(SIGNERS_WINDOW); }), ("fs_seed", &|h: &mut Hints| { - h.entries("fs_seed")[0][0] += F192::ONE; + h.entries("fs_seed")[0][0] += F64::ONE; }), ("leaf_defer", &|h: &mut Hints| { - h.entries("leaf_defer")[0][0] += F192::ONE; + h.entries("leaf_defer")[0][0] += F64::ONE; }), // A leaf derives its group's tweak table from this, so a wrong // epoch is caught by the signatures long before the statement digest. ("group (another epoch's tweak table)", &|h: &mut Hints| { - h.entries("group")[0][0] += F192::ONE; + h.entries("group")[0][0] += F64::ONE; }), ("group (wider than the u32 the verifier holds)", &|h: &mut Hints| { - h.entries("group")[0][0] += F192::new(0, 1, 0); + h.entries("group")[0][0] += F64(1 << 32); }), // The group's signatures were made over another message, so the // encoding digests reject long before the statement digest. ("group (another message under the signatures)", &|h: &mut Hints| { - h.entries("group")[0][1] += F192::ONE; + h.entries("group")[0][1] += F64::ONE; }), ]; for (description, tamper) in leaf_cases { @@ -4926,8 +4792,8 @@ def main(): entries[0][0] = other; }), ("group (a group's count moved to the other)", &|h: &mut Hints| { - h.entries("group")[0][3] = count(1); - h.entries("group")[1][3] = count(2); + h.entries("group")[0][5] = count(1); + h.entries("group")[1][5] = count(2); }), // The declared count is advice: keep both table groups but hash only the // first, while the statement still publishes both. @@ -4963,38 +4829,38 @@ def main(): entries[1] = entries[0].clone(); }), ("meta (n_sphincs inflated)", &|h: &mut Hints| { - h.entries("meta")[0][1] = count(3); + h.entries("meta")[0][2] = count(3); }), ("meta (n_raw_sphincs understated)", &|h: &mut Hints| { - h.entries("meta")[0][3] = count(1); + h.entries("meta")[0][4] = count(1); }), ("sphincs_signers (a message nobody signed)", &|h: &mut Hints| { - h.entries("sphincs_signers")[0][2] += F192::ONE; + h.entries("sphincs_signers")[0][4] += F64::ONE; }), ("sphincs_signers (a key nobody signed for)", &|h: &mut Hints| { - h.entries("sphincs_signers")[0][0] += F192::ONE; + h.entries("sphincs_signers")[0][0] += F64::ONE; }), ("sp_rand (another randomizer, so another index)", &|h: &mut Hints| { - h.entries("sp_rand")[0][0] += F192::ONE; + h.entries("sp_rand")[0][0] += F64::ONE; }), ("sp_counter", &|h: &mut Hints| { - h.entries("sp_counter")[0][0] += F192::ONE; + h.entries("sp_counter")[0][0] += F64::ONE; }), ("sp_digits", &|h: &mut Hints| { let entries = h.entries("sp_digits"); - entries[0][0] *= F192::from(primitives::field::G); + entries[0][0] *= G; }), ("sp_chain_starts", &|h: &mut Hints| { - h.entries("sp_chain_starts")[0][0] += F192::ONE; + h.entries("sp_chain_starts")[0][0] += F64::ONE; }), ("sp_fts_secrets", &|h: &mut Hints| { - h.entries("sp_fts_secrets")[0][0] += F192::ONE; + h.entries("sp_fts_secrets")[0][0] += F64::ONE; }), ("sp_fts_paths", &|h: &mut Hints| { - h.entries("sp_fts_paths")[0][0] += F192::ONE; + h.entries("sp_fts_paths")[0][0] += F64::ONE; }), ("sp_siblings", &|h: &mut Hints| { - h.entries("sp_siblings")[0][0] += F192::ONE; + h.entries("sp_siblings")[0][0] += F64::ONE; }), ]; for (description, tamper) in mixed_cases { @@ -5018,75 +4884,79 @@ def main(): order[0] = count(order.len()); }), ("merkle cap (all hashes skipped)", &|h: &mut Hints| { - h.entries("merkle_cap_active")[0].fill(F192::ZERO); + h.entries("merkle_cap_active")[0].fill(F64::ZERO); }), ("merkle cap (one ancestor skipped)", &|h: &mut Hints| { let flags = &mut h.entries("merkle_cap_active")[0]; - let active = flags.iter_mut().skip(2).find(|x| **x == F192::ONE).unwrap(); - *active = F192::ZERO; + let active = flags.iter_mut().skip(2).find(|x| **x == F64::ONE).unwrap(); + *active = F64::ZERO; }), ("merkle cap (root)", &|h: &mut Hints| { - h.entries("merkle_caps")[0][2] += F192::ONE; + h.entries("merkle_caps")[0][4] += F64::ONE; }), ("merkle cap (subtree)", &|h: &mut Hints| { - h.entries("merkle_caps")[0][4] += F192::ONE; + h.entries("merkle_caps")[0][8] += F64::ONE; }), ("merkle children (reversed)", &|h: &mut Hints| { let children = &mut h.entries("merkle_children")[0]; - children.swap(0, 2); - children.swap(1, 3); + for word in 0..4 { + children.swap(word, word + 4); + } }), ("merkle path (below cap)", &|h: &mut Hints| { - h.entries("merkle_children")[0][0] += F192::ONE; + h.entries("merkle_children")[0][0] += F64::ONE; }), ("merkle leaf", &|h: &mut Hints| { - h.entries("merkle_leaf_rows")[0][0] += F192::ONE; - }), - ("merkle leaf (extension limb outside K)", &|h: &mut Hints| { - let rows = h.entries("merkle_leaf_rows"); - let row = rows - .iter_mut() - .find(|row| row.len() == 3 << pcs::whir_config::SUBSEQUENT_FOLDING_FACTOR) - .unwrap(); - row[0] += F192::new(0, 1, 0); + h.entries("merkle_leaf_rows")[0][0] += F64::ONE; }), + ( + "merkle leaf (an upper limb in a later level's row)", + &|h: &mut Hints| { + let rows = h.entries("merkle_leaf_rows"); + let row = rows + .iter_mut() + .find(|row| row.len() == 3 << pcs::whir_config::SUBSEQUENT_FOLDING_FACTOR) + .unwrap(); + row[1] += F64::ONE; + }, + ), ("merkle path (last query)", &|h: &mut Hints| { let paths = h.entries("merkle_children"); - *paths.last_mut().unwrap().last_mut().unwrap() += F192::ONE; + *paths.last_mut().unwrap().last_mut().unwrap() += F64::ONE; }), ("child_index (duplicate slot)", &|h: &mut Hints| { let entries = h.entries("child_index"); entries[1] = entries[0].clone(); }), ("child_index (out of range)", &|h: &mut Hints| { - h.entries("child_index")[0] = vec![count(2 * SMALL_LEAF_SIZE)]; + h.entries("child_index")[0][0] = count(2 * SMALL_LEAF_SIZE); }), ("child_group (count understated)", &|h: &mut Hints| { - h.entries("child_group")[0][3] = count(SMALL_LEAF_SIZE - 1); + h.entries("child_group")[0][5] = count(SMALL_LEAF_SIZE - 1); }), // A child group claimed at an epoch or under a message the child // never carried: the map equality or the rebuilt digest rejects. ("child_group (epoch)", &|h: &mut Hints| { - h.entries("child_group")[0][0] += F192::ONE; + h.entries("child_group")[0][0] += F64::ONE; }), ("child_group (message)", &|h: &mut Hints| { - h.entries("child_group")[0][1] += F192::ONE; + h.entries("child_group")[0][1] += F64::ONE; }), ("child_defer (a forged carried claim)", &|h: &mut Hints| { - h.entries("child_defer")[0][0] += F192::ONE; + h.entries("child_defer")[0][0] += F64::ONE; }), ("bc_star_hint", &|h: &mut Hints| { - h.entries("bc_star_hint")[0][0] += F192::ONE; + h.entries("bc_star_hint")[0][0] += F64::ONE; }), ("mat_stars_hint", &|h: &mut Hints| { - h.entries("mat_stars_hint")[0][0] += F192::ONE; + h.entries("mat_stars_hint")[0][0] += F64::ONE; }), // The one hint carrying flock's whole lincheck terminal. Pinned not // by the guest's own assert (which merely defines it) but by the // matrix batching, whose reduced claims the root discharges against // the real A_0/B_0. ("matpart", &|h: &mut Hints| { - h.entries("matpart")[0][0] += F192::ONE; + h.entries("matpart")[0][0] += F64::ONE; }), // A node holding no raw XMSS signature builds no tweak tables, so // the statement digest is all that pins its epochs. The children's @@ -5095,7 +4965,7 @@ def main(): ( "group (a node that derives nothing from the epoch)", &|h: &mut Hints| { - h.entries("group")[0][0] += F192::ONE; + h.entries("group")[0][0] += F64::ONE; }, ), ]; @@ -5249,7 +5119,7 @@ def main(): ); } - /// Every `BLAKE2s` the guest itself runs reads a metadata cell an earlier + /// Every `BLAKE2s` the guest itself runs reads two metadata cells an earlier /// instruction of its own function wrote: a `SET` for a compile-time counter, /// an `XOR` for a window's base plus its offset. An unwritten cell is /// prover-chosen (write-once memory constrains only what something writes), so @@ -5270,17 +5140,16 @@ def main(): // Which frame cell an instruction writes, if any. A `DEREF` in cell mode is // bidirectional under write-once, so its local operand counts as a write. let written = |op: &Op| match *op { - Op::Set { o, .. } => vec![o], - Op::Xor { c, .. } | Op::Mul { c, .. } => vec![c], - Op::Deref { o3, mode, .. } => { - if mode == DerefMode::Cell { - vec![o3] - } else { - vec![] - } - } - Op::Blake2s { out, .. } => vec![out, out + 1], - Op::Jump { .. } => vec![], + Op::Set { o, .. } => o..o + 1, + Op::Xor64 { c, .. } | Op::Mul64 { c, .. } => c..c + 1, + Op::Xor192 { c, .. } | Op::Mul192 { c, .. } => c..c + 3, + Op::Deref { + o3, + mode: DerefMode::Cell, + .. + } => o3..o3 + 1, + Op::Deref { .. } | Op::Jump { .. } => 0..0, + Op::Blake2s { out, .. } => out..out + 4, }; let program = unified_guest(); let fill: Vec> = program @@ -5298,8 +5167,13 @@ def main(): if fill.iter().any(|f| f.contains(&pc)) { continue; } - if !program.prog[range.start..pc].iter().any(|op| written(op).contains(&md)) { - unwritten.push(format!("{name} pc {pc} md fp[{md}]")); + for cell in md..md + 2 { + if !program.prog[range.start..pc] + .iter() + .any(|op| written(op).contains(&cell)) + { + unwritten.push(format!("{name} pc {pc} md fp[{cell}]")); + } } } } diff --git a/crates/rec_aggregation/src/fibonacci.rs b/crates/rec_aggregation/src/fibonacci.rs index 22180ca06..ba69ae862 100644 --- a/crates/rec_aggregation/src/fibonacci.rs +++ b/crates/rec_aggregation/src/fibonacci.rs @@ -5,7 +5,7 @@ use lean_compiler::{compile, parse}; use lean_vm::cpu::{prove, verify}; use primitives::{ bench::Plan, - field::{F64, F192, g_pow}, + field::{F64, g_pow}, pretty_f64, pretty_integer, }; @@ -55,8 +55,8 @@ pub fn run_fibonacci(n: usize, log_inv_rate: usize, plan: Plan) { /// Build the demo program: Fibonacci in the exponent over `fib_n` steps (an /// unrolled `mul_range` loop over a `HeapBuf`), with the result `g^{F(N)}` /// published into cell `m[0]`. Returns the zkDSL source and the public input -/// `[g^{F(N)}, 0]`. -fn fibonacci_program(fib_n: usize) -> (String, [F192; 2]) { +/// `[g^{F(N)}, 0, 0, 0]`. +fn fibonacci_program(fib_n: usize) -> (String, [F64; 4]) { const UNROLL: usize = 1000; assert!( fib_n >= UNROLL && fib_n.is_multiple_of(UNROLL), @@ -70,7 +70,7 @@ fn fibonacci_program(fib_n: usize) -> (String, [F192; 2]) { previous = current; current = next; } - let public_input = [F192::from(previous), F192::ZERO]; + let public_input = [previous, F64::ZERO, F64::ZERO, F64::ZERO]; // `K` blocks: each reads its boundary pair into locals, runs `UNROLL` // Fibonacci `MUL`s in registers, and writes the next pair (4 DEREFs per diff --git a/crates/rec_aggregation/src/hash_chain.rs b/crates/rec_aggregation/src/hash_chain.rs index dfa5aad83..5e4e382da 100644 --- a/crates/rec_aggregation/src/hash_chain.rs +++ b/crates/rec_aggregation/src/hash_chain.rs @@ -5,12 +5,9 @@ use std::time::Instant; use lean_compiler::{compile, compile_without_filler, parse}; use lean_vm::cpu::{prove, verify}; use lean_vm::vmhash::compress; -use primitives::{ - field::{F64, F192}, - pretty_f64, pretty_integer, -}; +use primitives::{field::F64, pretty_f64, pretty_integer}; -fn instruction_counts(source: &str, public_input: [F192; 2]) -> [usize; lean_vm::cpu::Stats::TABLES.len()] { +fn instruction_counts(source: &str, public_input: [F64; 4]) -> [usize; lean_vm::cpu::Stats::TABLES.len()] { compile_without_filler(&parse(source).expect("parse")) .execute(public_input) .base_counts @@ -22,37 +19,32 @@ fn chain_source(steps: usize, unroll: usize) -> String { "N must be a positive multiple of UNROLL" ); let blocks = steps / unroll; - let result_cell = 2 * blocks; + let result_cell = 4 * blocks; let mut body = String::new(); - body.push_str(" b = i * i\n"); - body.push_str(" h0 = StackBuf(2)\n"); - body.push_str(" h0[0] = buff[b]\n"); - body.push_str(" h0[1] = buff[b * GEN]\n"); + body.push_str(" b = i ** 4\n"); + body.push_str(" h0 = StackBuf(4)\n"); + body.push_str(" h0[0:4] = buff[b:b + 4]\n"); for step in 1..=unroll { - body.push_str(&format!(" h{step} = StackBuf(2)\n")); + body.push_str(&format!(" h{step} = StackBuf(4)\n")); body.push_str(&format!( " blake2s(h{previous}, h{previous}, h{step})\n", previous = step - 1 )); } - for word in 0..2 { - body.push_str(&format!(" buff[b * GEN ** {}] = h{unroll}[{word}]\n", 2 + word)); - } + body.push_str(" nxt = b * GEN ** 4\n"); + body.push_str(&format!(" buff[nxt:nxt + 4] = h{unroll}\n")); format!( "def main():\n\ \x20 buff = HeapBuf({size})\n\ - \x20 buff[1] = 0\n\ - \x20 buff[GEN] = 0\n\ + \x20 buff[0:4] = [0, 0, 0, 0]\n\ \x20 for i in mul_range(1, GEN ** {blocks}):\n\ {body}\ \x20 output = 1\n\ - \x20 output[1] = buff[GEN ** {result_cell}]\n\ - \x20 output[GEN] = buff[GEN ** {next_cell}]\n\ + \x20 output[0:4] = buff[{result_cell}:{result_cell} + 4]\n\ \x20 return\n", - size = result_cell + 2, - next_cell = result_cell + 1, + size = result_cell + 4, ) } @@ -71,14 +63,10 @@ fn blake2s_hash_chain() { "LEANVM_HASH_N must be a multiple of LEANVM_HASH_UNROLL" ); - let mut digest = [F64::ZERO; 4]; + let mut public_input = [F64::ZERO; 4]; for _ in 0..steps { - digest = compress(digest, digest); + public_input = compress(public_input, public_input); } - let public_input = [ - F192::new(digest[0].0, digest[1].0, 0), - F192::new(digest[2].0, digest[3].0, 0), - ]; let source = chain_source(steps, unroll); let program = compile(&parse(&source).expect("parse")); @@ -98,16 +86,13 @@ fn blake2s_hash_chain() { pretty_integer(unroll) ); println!(" cycles (VM steps) : {}", pretty_integer(stats.cycles)); - for (name, &c) in ["XOR", "MUL", "SET", "DEREF", "JUMP", "BLAKE2S"] - .iter() - .zip(&stats.counts) - { + for (name, &c) in lean_vm::cpu::Stats::TABLES.iter().zip(&stats.counts) { let pow = if c == 0 { "0".to_string() } else { format!("2^{}", pretty_f64((c as f64).log2())) }; - println!(" {name:<6} instructions : {pow}"); + println!(" {name:<7} instructions : {pow}"); } println!( " committed witness size : 2^{:.3}", @@ -127,6 +112,6 @@ fn blake2s_hash_chain() { ); let mut wrong_input = public_input; - wrong_input[0] += F192::ONE; + wrong_input[0] += F64::ONE; assert!(verify(&program, &wrong_input, &proof).is_err()); } diff --git a/python-verifier/verifier.py b/python-verifier/verifier.py index c2f3f7612..cef9b5a8d 100644 --- a/python-verifier/verifier.py +++ b/python-verifier/verifier.py @@ -180,7 +180,6 @@ def __repr__(self) -> str: ZERO = E(0) ONE = E(1) GEN = E(2) -Y = E(0, 1) # the tower generator, y^3 = y + 1 def powers(base: E, count: int) -> list[E]: @@ -213,13 +212,9 @@ def words(self) -> tuple[int, int, int, int]: """Its four 64-bit words, the form the compression chain runs in.""" return unpack("<4Q", self.value) - def halves(self) -> tuple[E, E]: - """Its two 128-bit halves, the one form a digest travels in.""" - w0, w1, w2, w3 = self.words() - return (E(w0, w1), E(w2, w3)) - @classmethod def from_halves(cls, low: E, high: E) -> Digest: + """A digest as it travels on the stream: two 128-bit halves.""" require(not (low.c2 or high.c2), "a digest half is 128-bit") return cls(pack("<4Q", low.c0, low.c1, high.c0, high.c1)) @@ -344,9 +339,9 @@ def compress(left: Sequence[K | int], right: Sequence[K | int]) -> tuple[int, in class Transcript: - def __init__(self, proof: Proof, fiat_shamir_IV: Digest, public_input: Digest) -> None: + def __init__(self, proof: Proof, fiat_shamir_IV: Digest, public_input: Sequence[K]) -> None: self.proof = proof - self.state = compress(fiat_shamir_IV.words(), public_input.words()) + self.state = compress(fiat_shamir_IV.words(), public_input) self.stream_offset = 0 # in E field elements self.opening_offset = 0 # in bytes @@ -549,12 +544,14 @@ def verify_bus_balance(layout: Layout, transcript: Transcript) -> BusResult: # name them: the memory image on push, then each array's final count on pull. memory_low = tuple(point[: layout.log_memory]) bytecode_low = tuple(point[: layout.log_bytecode]) - memory = transcript.next_scalars(3) + memory = transcript.next_scalar() memory_final = transcript.next_scalar() bytecode_final = transcript.next_scalar() - claims = [ColumnClaim(column, memory_low, memory[column]) for column in (MEMORY_0, MEMORY_1, MEMORY_2)] - claims.append(ColumnClaim(MEMORY_FINAL_COUNTERS, memory_low, memory_final)) - claims.append(ColumnClaim(BYTECODE_FINAL_COUNTERS, bytecode_low, bytecode_final)) + claims = [ + ColumnClaim(MEMORY, memory_low, memory), + ColumnClaim(MEMORY_FINAL_COUNTERS, memory_low, memory_final), + ColumnClaim(BYTECODE_FINAL_COUNTERS, bytecode_low, bytecode_final), + ] memory_index = index_mle(memory_low) bytecode_index = index_mle(bytecode_low) bytecode_value = multilinear_eval(layout.bytecode, (*bytecode_low, *alphas)) @@ -565,7 +562,7 @@ def fingerprints(pc: E, memory_count: E, bytecode_count: E) -> tuple[E, E, E]: with the committed final count and ends at the last pc. Both boundaries sit in frame 0.""" return ( dot(weights[:3], (SEP_STATE, pc, _gpow(0))), - dot(weights[:6], (SEP_MEM, memory_index, memory_count, *memory)), + dot(weights[:4], (SEP_MEM, memory_index, memory_count, memory)), dot(weights[:3], (SEP_BYTECODE, bytecode_index, bytecode_count)) + bytecode_value, ) @@ -633,8 +630,8 @@ def table_sumcheck( R1CS_DIGEST = bytes.fromhex("537ad20790308f8eb8c0e8bd3e6c58ee64573371e3d53c30613dd04d87c0b7ea") # The columns no instruction table owns. They come first in the global column numbering, the tables after. -NUM_GLOBAL_COLUMNS = 6 -MEMORY_0, MEMORY_1, MEMORY_2, MEMORY_FINAL_COUNTERS, BYTECODE_FINAL_COUNTERS, QFLOCK = range(NUM_GLOBAL_COLUMNS) +NUM_GLOBAL_COLUMNS = 4 +MEMORY, MEMORY_FINAL_COUNTERS, BYTECODE_FINAL_COUNTERS, QFLOCK = range(NUM_GLOBAL_COLUMNS) BLAKE2S_R1CS_LOG_SIZE = 14 K_BITS = 64 @@ -706,11 +703,8 @@ def _counted(self, prefix: Sequence[Form], count: int, suffix: Sequence[Form]) - def bytecode(self, pc: int, count: int, opcode: int, operands: Sequence[Form]) -> None: self._counted((_const(SEP_BYTECODE), _col(pc)), count, (_const(_gpow(opcode)), *operands)) - def memory(self, address: Form, count: int, values: Sequence[Form]) -> None: - self._counted((_const(SEP_MEM), address), count, values) - - def memory_cols(self, address: Form, count: int, *columns: int) -> None: - self.memory(address, count, [_col(column) for column in columns] + [_const(ZERO)] * (3 - len(columns))) + def memory(self, address: Form, count: int, value: Form) -> None: + self._counted((_const(SEP_MEM), address), count, (value,)) # The instruction tables ------------------------------------------------------ @@ -739,50 +733,60 @@ def count_columns(self) -> tuple[int, ...]: return tuple(i for i, name in enumerate(self.columns) if name.startswith("cnt")) -def _flushes_arith(opcode: int, multiply: bool) -> Flushes: - pc, fp, o_a, o_b, o_c, cnt_a, cnt_b, cnt_c, cnt_bc = _cols(ARITH_COLUMNS, "pc", "fp", "o_a", "o_b", "o_c", "cnt_a", "cnt_b", "cnt_c", "cnt_bc") - va, vb = _cols(ARITH_COLUMNS, "va_0", "va_1", "va_2"), _cols(ARITH_COLUMNS, "vb_0", "vb_1", "vb_2") +def _flushes_arith64(opcode: int, multiply: bool) -> Flushes: + pc, fp, o_a, o_b, o_c, va, vb = _cols(ARITH64_COLUMNS, "pc", "fp", "o_a", "o_b", "o_c", "va", "vb") + cnt_a, cnt_b, cnt_c, cnt_bc = _cols(ARITH64_COLUMNS, "cnt_a", "cnt_b", "cnt_c", "cnt_bc") flushes = Flushes() flushes.state_step(pc, fp) flushes.bytecode(pc, cnt_bc, opcode, (_col(o_a), _col(o_b), _col(o_c), _const(ZERO), _const(ZERO))) - flushes.memory_cols(_prod(fp, o_a), cnt_a, *va) - flushes.memory_cols(_prod(fp, o_b), cnt_b, *vb) + flushes.memory(_prod(fp, o_a), cnt_a, _col(va)) + flushes.memory(_prod(fp, o_b), cnt_b, _col(vb)) + flushes.memory(_prod(fp, o_c), cnt_c, _prod(va, vb) if multiply else _col(va) + _col(vb)) + return flushes + + +def _flushes_arith192(opcode: int, multiply: bool) -> Flushes: + """Each operand is three consecutive cells `fp*o*g^i`, the limbs of one E element, read one at a time.""" + pc, fp, cnt_bc = _cols(ARITH192_COLUMNS, "pc", "fp", "cnt_bc") + va, vb = _cols(ARITH192_COLUMNS, "va_0", "va_1", "va_2"), _cols(ARITH192_COLUMNS, "vb_0", "vb_1", "vb_2") + operands = _cols(ARITH192_COLUMNS, "o_a", "o_b", "o_c") + flushes = Flushes() + flushes.state_step(pc, fp) + flushes.bytecode(pc, cnt_bc, opcode, (*(_col(o) for o in operands), _const(ZERO), _const(ZERO))) + # The tower product's three lanes: lane i is the sum of va[j]*vb[k] over TOWER_LANES[i]. TOWER_LANES = (((0, 0), (1, 2), (2, 1)), ((0, 1), (1, 0), (1, 2), (2, 1), (2, 2)), ((0, 2), (1, 1), (2, 0), (2, 2))) result = ( tuple(Form.sum(_prod(va[j], vb[k]) for j, k in lane) for lane in TOWER_LANES) if multiply else tuple(_col(va[i]) + _col(vb[i]) for i in range(3)) ) - flushes.memory(_prod(fp, o_c), cnt_c, result) + lanes = ((_col(v) for v in va), (_col(v) for v in vb), result) + for name, operand, values in zip("abc", operands, lanes, strict=True): + for limb, value in enumerate(values): + flushes.memory(_prod(fp, operand, limb), _cols(ARITH192_COLUMNS, f"cnt_{name}_{limb}")[0], value) return flushes def _flushes_set() -> Flushes: - pc, fp, o, cnt, cnt_bc = _cols(SET_COLUMNS, "pc", "fp", "o", "cnt", "cnt_bc") - k = _cols(SET_COLUMNS, "k_0", "k_1", "k_2") + pc, fp, o, k, cnt, cnt_bc = _cols(SET_COLUMNS, "pc", "fp", "o", "k", "cnt", "cnt_bc") flushes = Flushes() flushes.state_step(pc, fp) - flushes.bytecode(pc, cnt_bc, OP_SET, (_col(o), *(_col(limb) for limb in k), _const(ZERO))) - flushes.memory_cols(_prod(fp, o), cnt, *k) + flushes.bytecode(pc, cnt_bc, OP_SET, (_col(o), _col(k), _const(ZERO), _const(ZERO), _const(ZERO))) + flushes.memory(_prod(fp, o), cnt, _col(k)) return flushes def _flushes_deref() -> Flushes: - pc, fp, o1, o2, o3, f_pc, f_fp, ptr = _cols(DEREF_COLUMNS, "pc", "fp", "o1", "o2", "o3", "f_pc", "f_fp", "ptr") + pc, fp, o1, o2, o3, f_pc, f_fp, ptr, v3 = _cols(DEREF_COLUMNS, "pc", "fp", "o1", "o2", "o3", "f_pc", "f_fp", "ptr", "v3") cnt_ptr, cnt_target, cnt_local, cnt_bc = _cols(DEREF_COLUMNS, "cnt_ptr", "cnt_target", "cnt_local", "cnt_bc") - v3 = _cols(DEREF_COLUMNS, "v3_0", "v3_1", "v3_2") - - def gated(lane: int) -> list[Form]: - return [_col(lane), _prod(f_pc, lane), _prod(f_fp, lane)] - - # v2 = (1 + f_pc + f_fp)*v3 + f_pc*(g^2*pc) + f_fp*fp, lane-wise: only the low lane takes the two K-valued sources. - store = (Form.sum((*gated(v3[0]), _prod(f_pc, pc, 2), _prod(f_fp, fp))), Form.sum(gated(v3[1])), Form.sum(gated(v3[2]))) + # v2 = (1 + f_pc + f_fp)*v3 + f_pc*(g^2*pc) + f_fp*fp + store = Form.sum((_col(v3), _prod(f_pc, v3), _prod(f_fp, v3), _prod(f_pc, pc, 2), _prod(f_fp, fp))) flushes = Flushes() flushes.state_step(pc, fp) flushes.bytecode(pc, cnt_bc, OP_DEREF, (_col(o1), _col(o2), _col(o3), _col(f_pc), _col(f_fp))) - flushes.memory_cols(_prod(fp, o1), cnt_ptr, ptr) + flushes.memory(_prod(fp, o1), cnt_ptr, _col(ptr)) flushes.memory(_prod(ptr, o2), cnt_target, store) - flushes.memory_cols(_prod(fp, o3), cnt_local, *v3) + flushes.memory(_prod(fp, o3), cnt_local, _col(v3)) return flushes @@ -793,9 +797,9 @@ def _flushes_jump() -> Flushes: # next_pc = b*dest + (b+1)*g*pc, next_fp = b*frame + (b+1)*fp, both derived. flushes.state_derived(pc, fp, _prod(b, dest) + _prod(b, pc, 1) + _col(pc, 1), _prod(b, frame) + _prod(b, fp) + _col(fp)) flushes.bytecode(pc, cnt_bc, OP_JUMP, (_col(o_c), _col(o_d), _col(o_f), _const(ZERO), _const(ZERO))) - flushes.memory_cols(_prod(fp, o_c), cnt_c, cond) - flushes.memory_cols(_prod(fp, o_d), cnt_d, dest) - flushes.memory_cols(_prod(fp, o_f), cnt_f, frame) + flushes.memory(_prod(fp, o_c), cnt_c, _col(cond)) + flushes.memory(_prod(fp, o_d), cnt_d, _col(dest)) + flushes.memory(_prod(fp, o_f), cnt_f, _col(frame)) return flushes @@ -806,50 +810,58 @@ def _jump_constraints(columns: Sequence[E]) -> tuple[E, ...]: def _flushes_blake2s() -> Flushes: pc, fp, cnt_bc = _cols(BLAKE2S_COLUMNS, "pc", "fp", "cnt_bc") - operands = _cols(BLAKE2S_COLUMNS, "o_0", "o_1", "o_2", "o_3", "o_v", "o_out", "o_md") + operands = _cols(BLAKE2S_COLUMNS, "o_0", "o_1", "o_2", "o_3", "o_cv", "o_out", "o_md") flushes = Flushes() flushes.state_step(pc, fp) flushes.bytecode(pc, cnt_bc, OP_BLAKE2S, tuple(_col(i) for i in operands)) - # The nine cells read, as (cell, operand, offset from it): four addressed message chunks, then - # the consecutive chaining-value and output pairs, then the metadata cell (the byte counter and - # the two flags). Each holds two q_flock limbs and a zero top. - cells = (("m0", "o_0", 0), ("m1", "o_1", 0), ("m2", "o_2", 0), ("m3", "o_3", 0), - ("cv0", "o_v", 0), ("cv1", "o_v", 1), ("out0", "o_out", 0), ("out1", "o_out", 1), - ("md", "o_md", 0)) # fmt: skip - for cell, operand, exponent in cells: - address, count, lo, hi = _cols(BLAKE2S_COLUMNS, operand, f"cnt_{cell}", f"{cell}_lo", f"{cell}_hi") - flushes.memory_cols(_prod(fp, address, exponent), count, lo, hi) + for lane, operand, exponent in BLAKE2S_LANES: + value, address, count = _cols(BLAKE2S_COLUMNS, lane, operand, f"cnt_{lane}") + flushes.memory(_prod(fp, address, exponent), count, _col(value)) return flushes -OP_XOR, OP_MUL, OP_SET, OP_DEREF, OP_JUMP, OP_BLAKE2S = range(6) +OP_XOR64, OP_MUL64, OP_SET, OP_DEREF, OP_JUMP, OP_BLAKE2S, OP_XOR192, OP_MUL192 = range(8) -ARITH_COLUMNS = ("pc", "fp", "o_a", "o_b", "o_c", "va_0", "va_1", "va_2", "vb_0", "vb_1", "vb_2", "cnt_a", "cnt_b", "cnt_c", "cnt_bc",) # fmt: skip -SET_COLUMNS = ("pc", "fp", "o", "k_0", "k_1", "k_2", "cnt", "cnt_bc") -DEREF_COLUMNS = ("pc", "fp", "o1", "o2", "o3", "f_pc", "f_fp", "ptr", "v3_0", "v3_1", "v3_2", "cnt_ptr", "cnt_target", "cnt_local", "cnt_bc",) # fmt: skip +ARITH64_COLUMNS = ("pc", "fp", "o_a", "o_b", "o_c", "va", "vb", "cnt_a", "cnt_b", "cnt_c", "cnt_bc") +ARITH192_COLUMNS = ( + "pc", "fp", "o_a", "o_b", "o_c", "va_0", "va_1", "va_2", "vb_0", "vb_1", "vb_2", + "cnt_a_0", "cnt_a_1", "cnt_a_2", "cnt_b_0", "cnt_b_1", "cnt_b_2", "cnt_c_0", "cnt_c_1", "cnt_c_2", "cnt_bc", +) # fmt: skip +SET_COLUMNS = ("pc", "fp", "o", "k", "cnt", "cnt_bc") +DEREF_COLUMNS = ("pc", "fp", "o1", "o2", "o3", "f_pc", "f_fp", "ptr", "v3", "cnt_ptr", "cnt_target", "cnt_local", "cnt_bc") JUMP_COLUMNS = ("pc", "fp", "o_c", "o_d", "o_f", "v_cond", "v_pc", "v_fp", "cnt_c", "cnt_d", "cnt_f", "cnt_bc", "w", "b",) # fmt: skip +# The eighteen cells a BLAKE2s row reads, as (value lane, operand, offset from it): each message chunk's two cells, +# then the digest's four, the chaining value's four and the metadata's two (the byte counter, then the flags). +BLAKE2S_LANES = ( + ("m0_lo", "o_0", 0), ("m0_hi", "o_0", 1), ("m1_lo", "o_1", 0), ("m1_hi", "o_1", 1), + ("m2_lo", "o_2", 0), ("m2_hi", "o_2", 1), ("m3_lo", "o_3", 0), ("m3_hi", "o_3", 1), + ("out0", "o_out", 0), ("out1", "o_out", 1), ("out2", "o_out", 2), ("out3", "o_out", 3), + ("cv0", "o_cv", 0), ("cv1", "o_cv", 1), ("cv2", "o_cv", 2), ("cv3", "o_cv", 3), + ("md_lo", "o_md", 0), ("md_hi", "o_md", 1), +) # fmt: skip BLAKE2S_COLUMNS = ( - "pc", "fp", "o_0", "o_1", "o_2", "o_3", "o_v", "o_out", "o_md", - # These eighteen value limbs live in q_flock, not here: each is already a flock witness slot. - "m0_lo", "m0_hi", "m1_lo", "m1_hi", "m2_lo", "m2_hi", "m3_lo", "m3_hi", - "out0_lo", "out0_hi", "out1_lo", "out1_hi", "cv0_lo", "cv0_hi", "cv1_lo", "cv1_hi", "md_lo", "md_hi", + "pc", "fp", "o_0", "o_1", "o_2", "o_3", "o_cv", "o_out", "o_md", + # The value lanes live in q_flock, not here: each is already a flock witness slot. + *(lane for lane, _, _ in BLAKE2S_LANES), # ...and the read counts, committed here like every other column. - "cnt_m0", "cnt_m1", "cnt_m2", "cnt_m3", "cnt_cv0", "cnt_cv1", "cnt_out0", "cnt_out1", "cnt_md", "cnt_bc", + *(f"cnt_{lane}" for lane, _, _ in BLAKE2S_LANES), "cnt_bc", ) # fmt: skip TABLES = ( - Table("xor", OP_XOR, ARITH_COLUMNS, _flushes_arith(OP_XOR, multiply=False)), - Table("mul", OP_MUL, ARITH_COLUMNS, _flushes_arith(OP_MUL, multiply=True)), + Table("xor64", OP_XOR64, ARITH64_COLUMNS, _flushes_arith64(OP_XOR64, multiply=False)), + Table("mul64", OP_MUL64, ARITH64_COLUMNS, _flushes_arith64(OP_MUL64, multiply=True)), Table("set", OP_SET, SET_COLUMNS, _flushes_set()), Table("deref", OP_DEREF, DEREF_COLUMNS, _flushes_deref()), Table("jump", OP_JUMP, JUMP_COLUMNS, _flushes_jump(), _jump_constraints), Table("blake2s", OP_BLAKE2S, BLAKE2S_COLUMNS, _flushes_blake2s()), + Table("xor192", OP_XOR192, ARITH192_COLUMNS, _flushes_arith192(OP_XOR192, multiply=False)), + Table("mul192", OP_MUL192, ARITH192_COLUMNS, _flushes_arith192(OP_MUL192, multiply=True)), ) -# Where in the flock witness each embedded BLAKE2s limb live: one 64-bit slot per limb, the chaining value first, then the +# Where in the flock witness each BLAKE2s value lane lives: one 64-bit slot per lane, the chaining value first, then the # digest, the message block and the metadata. Slots 8 and 9 are flock's constant wire and the padding up to its message base, which no memory cell carries. BLAKE2S_SLOTS = ( - "cv0_lo", "cv0_hi", "cv1_lo", "cv1_hi", "out0_lo", "out0_hi", "out1_lo", "out1_hi", None, None, + "cv0", "cv1", "cv2", "cv3", "out0", "out1", "out2", "out3", None, None, "m0_lo", "m0_hi", "m1_lo", "m1_hi", "m2_lo", "m2_hi", "m3_lo", "m3_hi", "md_lo", "md_hi", ) # fmt: skip @@ -881,11 +893,11 @@ def build_layout(bytecode: Sequence[K], log_memory: int, table_log_heights: Sequ # Every column's log size, in global order: the framework's, q_flock's, then each table's block. qflock_kappa = table_log_heights[OP_BLAKE2S] + QFLOCK_SLOT_BITS - kappas = [log_memory, log_memory, log_memory, log_memory, log_bytecode, qflock_kappa] + kappas = [log_memory, log_memory, log_bytecode, qflock_kappa] for table in TABLES: kappas += [table_log_heights[table.opcode]] * table.width - # A BLAKE2s value limb gets no block of its own: it is committed inside q_flock, whose slots + # A BLAKE2s value lane gets no block of its own: it is committed inside q_flock, whose slots # interleave, so it sits at q_flock's offset behind its own slot's bits. Same width either way. limbs = {GLOBAL_COLUMN_BASES[OP_BLAKE2S] + _cols(BLAKE2S_COLUMNS, name)[0]: slot for slot, name in enumerate(BLAKE2S_SLOTS) if name} blocks = {column: kappa for column, kappa in enumerate(kappas) if column not in limbs} @@ -1357,7 +1369,8 @@ def verify_stacked_opening(transcript: Transcript, root: Digest, stack_log: int, verify_whir(transcript, stack_log, log_inv_rate, dot(scales, values), root, lambda point: dot(scales, [weight(point) for weight in weights])) -def verify_execution(bytecode: Sequence[K], public_input: Digest, proof: Proof) -> None: +def verify_execution(bytecode: Sequence[K], public_input: Sequence[K], proof: Proof) -> None: + require(len(public_input) == 4, "the public input is four words") bytecode_hash = blake2s_hash(b"".join(word.to_bytes() for word in bytecode)) iv_preimage = b"leanvm" + pack(" int: parser = argparse.ArgumentParser(description="Verify a leanVM execution proof") parser.add_argument("bytecode", type=Path, help="stacked bytecode multilinear, little-endian 64-bit words") - parser.add_argument("public_input", type=Path, help="256-bit public input") + parser.add_argument("public_input", type=Path, help="public input, four little-endian 64-bit words") parser.add_argument("stream", type=Path, help="the proof's scalar stream, 24-byte little-endian field elements") parser.add_argument("merkle_openings", type=Path, help="every Merkle opening: its leaf's words, then its sibling digests") arguments = parser.parse_args(argv) @@ -1422,8 +1433,11 @@ def main(argv: Sequence[str] | None = None) -> int: encoded_bytecode = arguments.bytecode.read_bytes() require(len(encoded_bytecode) % 8 == 0, "bytecode is not a whole number of 64-bit words") bytecode = [K(int.from_bytes(encoded_bytecode[i : i + 8], "little")) for i in range(0, len(encoded_bytecode), 8)] + encoded_public_input = arguments.public_input.read_bytes() + require(len(encoded_public_input) == 32, "the public input is four 64-bit words") + public_input = [K(word) for word in unpack("<4Q", encoded_public_input)] proof = Proof.load(arguments.stream, arguments.merkle_openings) - verify_execution(bytecode, Digest(arguments.public_input.read_bytes()), proof) + verify_execution(bytecode, public_input, proof) except (OSError, ValueError, KeyError, VerificationError) as exc: parser.exit(1, f"verification failed: {exc}\n") print("verification succeeded") From 72f2864b8582ab9735f9b74b56d251438bbfc39d Mon Sep 17 00:00:00 2001 From: Tom Wambsgans Date: Mon, 14 Sep 2026 19:51:49 -0400 Subject: [PATCH 02/40] Document the 64-bit zkDSL zkDSL.md and the snark_lib stub describe 64-bit words, 192-bit runs and their builtins, four-cell blake2s operands with two-cell metadata, the four-word public input, and runs through calls and match dispatch. AGENTS.md counts eight opcodes and eight-word SPHINCS slots. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_017HVkj8EFExoGGa1okY9gvX --- AGENTS.md | 2 +- crates/lean_compiler/snark_lib.py | 114 +++++++++------ crates/lean_compiler/zkDSL.md | 226 ++++++++++++++++-------------- 3 files changed, 190 insertions(+), 152 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 1bf03c2c4..99867bfd6 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -98,7 +98,7 @@ The same verification algorithm is written out three times, in three languages. 2. **Python**, `python-verifier/verifier.py` (no dependencies), for readability and simplicity. Pinned by `lean_vm/tests/verifiers/python_verifier.rs`. 3. **Recursive verifier**, `crates/rec_aggregation/guests/lean_ethereum.py`. Its zkDSL compiles to the ISA; proving its execution gives a proof of child proofs. -Understand the third before changing the verifier. `guests/lean_ethereum.py` is zkDSL, not runnable Python. `lean_compiler` lowers it to the six-opcode, write-once-memory VM, so the prover proves every verifier step. Its size and instruction mix are what the recursion benchmark reports first. It verifies raw signatures of both schemes: a node's coverage table has one contiguous region per XMSS epoch group and separate regions for SPHINCS and DA roots, so the one range check a write already needs also keeps a signature off another group's declared keys, of either scheme, and the statement's signer lists say which scheme verified which key against which `(epoch, message)`. The XMSS signers are grouped by epoch, each group carrying its own message, a runtime number of groups (at most `MAX_EPOCHS`) bound through the signer-set digest, which is plain BLAKE2s of a byte string (each list's own digest folded into it, likewise plain BLAKE2s): a run-time length rides the byte counter because the counter is a memory operand, split as `doc/leanvm` §sec:prog-byte-counter describes. A child's groups need not equal its parent's, a hinted map tying each child group to a parent group with the same epoch and message. A SPHINCS signer's message rides its own four-cell slot, so that list is `(key, message)` pairs; both lists count claims rather than distinct signers, an XMSS key claiming once per epoch it signed at. Both schemes' tweaks are built in-circuit: XMSS's from the epochs the statement carries, derived once per group that verifies raw XMSS signatures and skipped by one that verifies none, SPHINCS's per signature from the index its message digest picks. Two consequences: +Understand the third before changing the verifier. `guests/lean_ethereum.py` is zkDSL, not runnable Python. `lean_compiler` lowers it to the eight-opcode, write-once-memory VM, so the prover proves every verifier step. Its size and instruction mix are what the recursion benchmark reports first. It verifies raw signatures of both schemes: a node's coverage table has one contiguous region per XMSS epoch group and separate regions for SPHINCS and DA roots, so the one range check a write already needs also keeps a signature off another group's declared keys, of either scheme, and the statement's signer lists say which scheme verified which key against which `(epoch, message)`. The XMSS signers are grouped by epoch, each group carrying its own message, a runtime number of groups (at most `MAX_EPOCHS`) bound through the signer-set digest, which is plain BLAKE2s of a byte string (each list's own digest folded into it, likewise plain BLAKE2s): a run-time length rides the byte counter because the counter is a memory operand, split as `doc/leanvm` §sec:prog-byte-counter describes. A child's groups need not equal its parent's, a hinted map tying each child group to a parent group with the same epoch and message. A SPHINCS signer's message rides its own slot, eight words holding the key then the message, so that list is `(key, message)` pairs; both lists count claims rather than distinct signers, an XMSS key claiming once per epoch it signed at. Both schemes' tweaks are built in-circuit: XMSS's from the epochs the statement carries, derived once per group that verifies raw XMSS signatures and skipped by one that verifies none, SPHINCS's per signature from the index its message digest picks. Two consequences: - The guest is **self-referential**: it verifies proofs of itself, so `unified_guest` compiles it to a fixed point on its own log size. The digest needs no fixed point, riding the statement instead of the code, which is also what lets one bytecode serve any inner size and PCS rate. - It does not verify *quite* everything in-circuit. Three claims on fixed polynomials (stacked bytecode, flock's A0/B0) are deferred. Each node batches its children's carried claims with the fresh ones its verifications raise, `2n` per polynomial down to one; only the root's are discharged natively, by `EthereumProof::verify` (explained in `doc/leanvm/`). diff --git a/crates/lean_compiler/snark_lib.py b/crates/lean_compiler/snark_lib.py index 9715f64a1..95ed8fd47 100644 --- a/crates/lean_compiler/snark_lib.py +++ b/crates/lean_compiler/snark_lib.py @@ -13,10 +13,11 @@ class _Elt: - """A 192-bit machine word in E = GF(2^192), represented as a cubic tower - over K = GF(2^64). Indices and addresses are K-valued powers of GEN, i.e. - "in the exponent": `GEN ** k` is the k-th index and `x * GEN` its successor. - A heap pointer is K-valued too; `buf[i]` is the write-once cell at `buf * i`.""" + """A 64-bit machine word in K = GF(2^64). Indices and addresses are powers of + GEN, i.e. "in the exponent": `GEN ** k` is the k-th index and `x * GEN` its + successor. A heap pointer is a word too; `buf[i]` is the write-once cell at + `buf * i`. A 192-bit element of E = K[y]/(y^3 + y + 1) is not a word but a + run of three cells, limbs low first, computed with `add192`/`mul192`/`div192`.""" def __add__(self, other): # field addition = XOR _ = other @@ -24,7 +25,7 @@ def __add__(self, other): # field addition = XOR __radd__ = __add__ - def __mul__(self, other): # tower-field product + def __mul__(self, other): # GF(2^64) product _ = other return _Elt() @@ -40,28 +41,61 @@ def __pow__(self, k: int): _ = k return _Elt() - def __getitem__(self, idx): # heap read m[self · idx] + def __getitem__(self, idx): # heap read m[self · idx], or a slice (a run) _ = idx return _Elt() - def __setitem__(self, idx, value): # heap store m[self · idx] (write-once) + def __setitem__(self, idx, value): # heap store m[self · idx], or a run store (write-once) _ = idx, value def f192(c0: int, c1: int, c2: int) -> _Elt: - """Construct a field constant from its three little-endian GF(2^64) limbs.""" + """A 192-bit constant `c0 + c1·y + c2·y^2`, a three-cell run. Each limb is a + compile-time unsigned 64-bit integer. Also valid as a global constant, + `ONE = f192(1, 0, 0)`.""" _ = c0, c1, c2 return _Elt() +def add192(a, b) -> _Elt: + """The 192-bit sum of two three-cell runs: one XOR192.""" + _ = a, b + return _Elt() + + +def mul192(a, b) -> _Elt: + """The 192-bit product of two three-cell runs: one MUL192.""" + _ = a, b + return _Elt() + + +def div192(a, b) -> _Elt: + """The 192-bit quotient `a · b⁻¹`: one MUL192 whose unwritten operand is the + quotient, back-solved at witness generation. A zero divisor is rejected there.""" + _ = a, b + return _Elt() + + +def assert_eq192(a, b) -> None: + """Assert two three-cell runs are equal: one XOR192 into a zero run.""" + _ = a, b + + +def assert_ne192(a, b) -> None: + """Assert two three-cell runs differ: their sum times a hinted inverse, one + MUL192 into a one run. Sound whatever the hint, as for `assert a != b`.""" + _ = a, b + + GEN = _Elt() """The fixed generator g = x of K^× = GF(2^64)^× (order 2^64 - 1).""" def hint_decompose_bits(bits, value, nbits: int) -> None: - """Computed advice: the prover writes the `nbits` bits of `value` into the - `bits` buffer. UNCONSTRAINED: the caller must check booleanity and that the - bits reconstruct `value` (a range check that `value < 2^nbits`).""" + """Computed advice: the prover writes the low `nbits` (at most 64) bits of the + word `value` into the `bits` buffer. UNCONSTRAINED: the caller must check + booleanity and that the bits reconstruct `value` (a range check that + `value < 2^nbits`).""" _ = bits, value, nbits @@ -150,8 +184,8 @@ def HeapBuf(n) -> _Elt: def StackBuf(n: int) -> _Elt: - """Allocate `n` consecutive frame (stack) cells. A size-2 StackBuf holds a - 256-bit value and is a valid `blake2s` operand.""" + """Allocate `n` consecutive frame (stack) cells. A size-3 StackBuf holds a + 192-bit value, and a size-4 one a 256-bit value, a valid `blake2s` operand.""" _ = n return _Elt() @@ -167,13 +201,13 @@ def addr(buf) -> _Elt: def hint_witness(dest, name: Optional[str] = None) -> Any: - """Take the next ENTRY (a slice of values) of the named prover witness + """Take the next ENTRY (a slice of words) of the named prover witness stream. Two forms. As a statement, `hint_witness(dest, "name")` fills `dest` (a StackBuf, or a StackBuf/HeapBuf slice of any length). As an expression, - `x = hint_witness("name")` binds ONE value and needs no destination, the - entry then having to hold exactly one value. + `x = hint_witness("name")` binds ONE word and needs no destination, the + entry then having to hold exactly one word. The same symbol may be hinted many times, each call popping the next entry (`Program::set_witness`; test programs declare one `# witness name: v1, …` @@ -184,19 +218,6 @@ def hint_witness(dest, name: Optional[str] = None) -> Any: return _Elt() -def assert_in_k(a, b) -> None: - """Prove that both machine words are GF(2^64)-valued. This is the sole - packing-related intrinsic and lowers to one untaken JUMP.""" - _ = a, b - - -def hint_f192_limbs(dest, value) -> None: - """Computed advice: write the first `len(dest)` GF(2^64) coordinate limbs - of `value` into a 1-to-3-cell StackBuf. UNCONSTRAINED; callers bind the - result with `assert_in_k` and field reconstruction.""" - _ = dest, value - - def blake2s( a, b, @@ -209,8 +230,8 @@ def blake2s( md=None, ) -> None: """One standard BLAKE2s compression of the two 256-bit message operands - `a`, `b`, written into the 2-cell run `out` (write-once: if `out` was - already written, this asserts it equals the chaining value). + `a`, `b`, written into the 4-cell run `out` (write-once: if `out` was + already written, this asserts it equals the digest). With no keywords this hashes exactly 64 bytes: the parameterized BLAKE2s-256 initial chaining value, byte counter 64, final-block flag set. That is @@ -218,19 +239,20 @@ def blake2s( For a longer message, drive the blocks yourself. `counter` is the CUMULATIVE byte count through this block (`64 * whole_blocks_before + bytes_in_this_block`) - and `final=1` marks the last block; `cv` carries the previous block's output - and requires one of `counter`, `final`, `last_node` or `md`. Setting any of - `counter`, `final` or `last_node` makes `final` default - to 0, so a single short block needs `counter=, final=1`. Bytes past the - block's real length must be zero-filled by the program. `last_node` is - BLAKE2s's tree-mode `f1` and is 0 everywhere here. - - `md` is the whole 128-bit metadata word as a value the program computed, for - a hash whose block count is only known at run time. It replaces `counter`, - `final` and `last_node` (giving both is an error), and it must not name a - cell of `out`. - - Message, chaining-value, and output operands are size-2 StackBufs or - 2-cell slices `buf[lo:hi]` of larger StackBufs or HeapBufs (heap inputs are - bridged through the stack, one DEREF per cell).""" + and `final=1` marks the last block; `cv` carries the previous block's 4-cell + output and requires one of `counter`, `final`, `last_node` or `md`. Setting any + of `counter`, `final` or `last_node` makes `final` default to 0, so a single + short block needs `counter=, final=1`. Bytes past the block's real length + must be zero-filled by the program. `last_node` is BLAKE2s's tree-mode `f1` + and is 0 everywhere here. + + `md` is the whole metadata as a 2-cell run the program computed, + `[counter, f0 | f1 << 32]`, for a hash whose block count is only known at run + time. It replaces `counter`, `final` and `last_node` (giving both is an + error), and it must not name a cell of `out`. + + Message, chaining-value, and output operands are 4-cell runs: a StackBuf(4), + or a slice `buf[lo:hi]` of a larger StackBuf or a HeapBuf (heap inputs are + bridged through the stack, one DEREF per cell). A message operand may also + be a list literal of four words, a run element flattening into its words.""" _ = a, b, out, cv, counter, final, last_node, md diff --git a/crates/lean_compiler/zkDSL.md b/crates/lean_compiler/zkDSL.md index a2c55ccd1..09dd30634 100644 --- a/crates/lean_compiler/zkDSL.md +++ b/crates/lean_compiler/zkDSL.md @@ -1,8 +1,8 @@ # zkDSL Language Reference (leanVM) -The zkDSL is a Python-syntax language that compiles to the leanVM ISA: six instructions (`XOR`, `MUL`, `SET`, `DEREF`, `JUMP`, `BLAKE2s`) over the binary field GF(2^192), with write-once memory and all indices carried "in the exponent" as powers of a fixed generator. For the underlying VM and proving system, see [`doc/leanvm/main.tex`](../../doc/leanvm/main.tex). +The zkDSL is a Python-syntax language that compiles to the leanVM ISA: eight instructions (`XOR64`, `MUL64`, `SET`, `DEREF`, `JUMP`, `BLAKE2S`, `XOR192`, `MUL192`) over 64-bit memory words in the binary field GF(2^64), the two 192-bit instructions acting on runs of three cells, with write-once memory and all indices carried "in the exponent" as powers of a fixed generator. For the underlying VM and proving system, see [`doc/leanvm/main.tex`](../../doc/leanvm/main.tex). -Source files use the `.py` extension and are **Python-shaped**: they import the [`snark_lib`](snark_lib.py) stub, which defines `GEN`, `log`, `mul_range`, `HeapBuf`, `StackBuf`, `assert_in_k`, and `blake2s`, so editors and linters resolve the intrinsic names. The compiler skips the import. Ordinary helpers such as `pack64x2` are defined in the single-file guest. A program that uses placeholders is not a runnable Python file: its `*_PLACEHOLDER` identifiers are undefined until the host fills them in, so importing it raises `NameError`. +Source files use the `.py` extension and are **Python-shaped**: they import the [`snark_lib`](snark_lib.py) stub, which defines `GEN`, `log`, `mul_range`, `HeapBuf`, `StackBuf`, `blake2s`, the 192-bit builtins (`f192`, `add192`, `mul192`, `div192`, `assert_eq192`, `assert_ne192`) and the other intrinsics, so editors and linters resolve the intrinsic names. The compiler skips the import. A program that uses placeholders is not a runnable Python file: its `*_PLACEHOLDER` identifiers are undefined until the host fills them in, so importing it raises `NameError`. Entry points: `lean_compiler::parse` / `parse_file_with_replacements` → `lean_compiler::compile` → `lean_vm::cpu::prove` / `verify`. @@ -16,20 +16,20 @@ The fields are `K = GF(2)[x]/(x^64 + x^4 + x^3 + x + 1)` and `E = K[y]/(y^3 + y + 1) = GF(2^192)`. -Machine **words** (the contents of a memory cell, an immediate, a hashed value, the `JUMP` condition) are elements of `E`. **Addresses**, the program counter, the frame pointer, read counters, operands, opcodes, and domain separators live in the 64-bit subfield `K = GF(2^64)`. There are no runtime integers. +Machine **words** (the contents of a memory cell, an immediate, one word of a hashed value, the `JUMP` condition) are elements of the 64-bit field `K = GF(2^64)`, and so are addresses, the program counter, the frame pointer, read counters, operands, opcodes, and domain separators. An element of `E` is not a word but a **run** of three consecutive cells (see "192-bit values"). There are no runtime integers. -- `+` is field addition = bitwise **XOR** (192-bit on words, so `x + x == 0`), -- `*` is multiplication in `E`; for g-powers and addresses it stays within `K`, -- `/` is runtime field division, `a / b = a · b⁻¹`. It costs one `MUL`: the compiler leaves the quotient cell unset and emits the checked relation `quotient · b == a`, which witness generation back-solves. Division by zero is undefined. This is distinct from `//`, compile-time integer floor division in sizes and indices, -- an integer literal `n` supplies up to 128 raw bits and is embedded as `F192(c0, c1, 0)`. This is a source-syntax limit, not the machine-word width: words have three 64-bit limbs. Thus `5` is `1 + x^2`, not the integer five, and the literal `18446744073709551616` is the tower element `y`. Written `2 ** 64` in a value position it is not: `**` there is a field power, so it is `x^64` reduced. `const(2 ** 64)` is the literal. Full-width constants use `f192(c0, c1, c2)`, with each limb an unsigned 64-bit compile-time integer, -- `GEN` is the fixed generator `g = x` of the 64-bit subfield `K^×` (multiplicative order `2^64 − 1`), +- `+` is field addition = bitwise **XOR** (64-bit on words, so `x + x == 0`), +- `*` is multiplication in `K`; arithmetic in `E` goes through the builtins `add192`, `mul192` and `div192`, +- `/` is runtime field division, `a / b = a · b⁻¹`. It costs one `MUL64`: the compiler leaves the quotient cell unset and emits the checked relation `quotient · b == a`, which witness generation back-solves. Division by zero is undefined. This is distinct from `//`, compile-time integer floor division in sizes and indices, +- an integer literal `n` is a 64-bit word whose bits are the coefficients of `x`: `5` is `1 + x^2`, not the integer five. A literal that does not fit in 64 bits is rejected in a value position, while compile-time integer arithmetic (a size, a bound, a keyword) reads it whole. Written `2 ** 64` in a value position it is not a literal: `**` there is a field power, so it is `x^64` reduced. A 192-bit constant is `f192(c0, c1, c2)`, with each limb an unsigned 64-bit compile-time integer, +- `GEN` is the fixed generator `g = x` of the 64-bit field `K^×` (multiplicative order `2^64 − 1`), - `GEN ** e` is the compile-time constant `g^e ∈ K` (`**` takes base `GEN` and a compile-time integer exponent: a literal, a constant, an `unroll` variable, `len(...)`, or index arithmetic of those). So `buf[GEN ** i]` names heap cell `i` directly inside an `unroll` loop, with no running-pointer cursor. -- constant arithmetic means different things in the two positions, and this is a silent trap: `a + b` on two constants is **integer** addition in an index, a bound or a keyword (`buf[GEN ** (i + 1)]`, `unroll(0, n + 1)`, `counter=64 * (q + 1)`), and **XOR** in a value, where `1 + 1` is `0`. So a literal built in a value position must not add overlapping integers: `tweak = base + (level + 1) * SHIFT` drops the whole term on odd levels. Products are safe (an integer times a power of two is that shift, as long as the top bit stays inside the limb); to add, name the regime with `const(...)`: `tweak = base + const((level + 1) * SHIFT)`. +- constant arithmetic means different things in the two positions, and this is a silent trap: `a + b` on two constants is **integer** addition in an index, a bound or a keyword (`buf[GEN ** (i + 1)]`, `unroll(0, n + 1)`, `counter=64 * (q + 1)`), and **XOR** in a value, where `1 + 1` is `0`. So a literal built in a value position must not add overlapping integers: `tweak = base + (level + 1) * SHIFT` drops the whole term on odd levels. Products are safe (an integer times a power of two is that shift, as long as the top bit stays inside the word); to add, name the regime with `const(...)`: `tweak = base + const((level + 1) * SHIFT)`. - a **global constant** and a **`Const` parameter** are integer-arithmetic throughout, which is deliberate (it is what makes a derived size right) but means the same text means different things in the two places: `STEP = 3 + 1` is the integer `4` everywhere it appears, while the identical `x = 3 + 1` written inside a function is the field element `2`. Neither is wrong; they are two regimes, and a name crossing between them is where the trap bites. - a **compile-time `if`** must say which regime it means when the two disagree. The fold decides on the integer reading while the runtime test of the same condition compares field values, so `if 3 + 1 == 4` is true one way and false the other. Such a condition is **rejected**; write `if const(3 + 1 == 4):` to decide it with integer arithmetic, or spell the operands so the two readings agree (a product or a shift rather than a sum of overlapping integers). A condition whose readings already agree needs no wrapper. - `base ** e` with a **non-`GEN`** base and a compile-time exponent `e` is square-and-multiply: integer arithmetic in an index/bound position (`2 ** c`), or field arithmetic in a value position (`x ** k`, e.g. a loop counter `g^i` raised to a stride to reach cell `i·stride`). The base may be runtime. -A logical **index** `i` is carried as `g^i` in the 64-bit subfield (order `2^64 − 1`): incrementing is one multiplication by `GEN`, and memory/bytecode addresses are g-powers. This is the design idiom of the whole VM: loops, heap addressing, and range checks below all live in the exponent, in `K`. +A logical **index** `i` is carried as `g^i` in `K` (order `2^64 − 1`): incrementing is one multiplication by `GEN`, and memory/bytecode addresses are g-powers. This is the design idiom of the whole VM: loops, heap addressing, and range checks below all live in the exponent, in `K`. ## Program shape @@ -50,22 +50,25 @@ def helper(a, b): # other functions `import snark_lib` / `from snark_lib import *` are the only imports accepted; anything else is a compile error (no multi-file programs yet). Comments (`#`) and blank lines are free. Indentation is block structure, as in Python. -Ordinary functions may return scalars, `HeapBuf` pointers, and `StackBuf` values, including mixtures in a tuple return. A returned `StackBuf(n)` has a compile-time-known size: its `n` cells are copied through `n` consecutive return slots and the caller binds the result as a new `StackBuf(n)`. A `HeapBuf` return is just its one-cell pointer; the allocation hint already ran where the buffer was created, so no size metadata needs to cross the call. +Ordinary functions may return scalars, `HeapBuf` pointers, and runs (a `StackBuf`, a slice, a list literal, a 192-bit value), including mixtures in a tuple return. A returned run has a compile-time-known size: its `n` cells are copied through `n` consecutive return slots and the caller binds the result as a new `StackBuf(n)`. A `HeapBuf` return is just its one-cell pointer; the allocation hint already ran where the buffer was created, so no size metadata needs to cross the call. ## Public input -Memory cells `m[0]` and `m[1]` hold the two public-input words, each an F192 machine word. A program *publishes* results by asserting them against those cells through the write-once heap store (the pointer `g^0` addresses absolute memory): +Memory cells `m[0]` to `m[3]` hold the four public-input words. A program *publishes* results by asserting them against those cells through the write-once heap store (the pointer `g^0` addresses absolute memory): ```python p = GEN ** 0 p[1] = result_a # m[p·1] = m[0], an equality assert against the public input p[GEN] = result_b # m[p·g] = m[1] +p[2:4] = pair # a slice store of a 2-cell run into m[2], m[3] ``` -Test programs under `tests/programs/` declare the public input they expect with a top-of-file annotation of two constant elements (or omit it to run with two zeros); the generic harness `tests/py_source.rs` proves and verifies every program in the directory: +A 4-cell digest publishes whole as `p[0:4] = digest`. + +Test programs under `tests/programs/` declare the public input they expect with a top-of-file annotation of up to four constant words, zero-padded (or omit it to run with four zeros); the generic harness `tests/suite/py_source.rs` proves and verifies every program in the directory: ```python -# public_input: GEN ** 89, 101229015297003380629709256178361811305 +# public_input: GEN ** 89, 101229015297003380 ``` ## Global constants and placeholders @@ -79,6 +82,7 @@ N = 8 # an integer size / value STEP = GEN ** 2 # a g-power constant (index carried in the exponent) WIDE = N + 1 # compile-time INTEGER arithmetic (`+ - * / **`); # references to *earlier* constants are allowed +ONE = f192(1, 0, 0) # a 192-bit constant, a three-cell run wherever it appears def main(): buf = StackBuf(N) # a constant is a plain literal: usable as a size, @@ -103,7 +107,7 @@ so one source template compiles at many sizes. An unfilled placeholder (no repla ### Constant arrays -A global constant may be a **list literal**, `NAME = [a, b, c]`, of compile-time values (integers or field values, each a ``). Unlike a scalar constant it is **not** textually substituted; it is carried to lowering and consumed at compile time: +A global constant may be a **list literal**, `NAME = [a, b, c]`, of compile-time values (integers or field values, each a `` that fits in a 64-bit word). Unlike a scalar constant it is **not** textually substituted; it is carried to lowering and consumed at compile time: ```python QUERIES = [290, 177, 145] # or QUERIES = QUERIES_PLACEHOLDER, filled "[290, 177, 145]" @@ -116,7 +120,7 @@ def main(): ... ``` -`NAME[i]` yields the element (as a field value in value position, or as an integer where an index / slice bound / `unroll` count / `**` exponent is expected), and `len(NAME)` its length. The index `i` must be compile-time (a literal, a constant, or an `unroll` variable). This is what lets one source file adapt to a per-level config vector (query counts, fold factors, sizes) without Rust-side code generation. Nested lists are not (yet) supported: flatten a 2-D table into one array plus an offsets array. +`NAME[i]` yields the element (as a field value in value position, or as an integer where an index / slice bound / `unroll` count / `**` exponent is expected), and `len(NAME)` its length. The index `i` must be compile-time (a literal, a constant, or an `unroll` variable). This is what lets one source file adapt to a per-level config vector (query counts, fold factors, sizes) without Rust-side code generation. Nested lists are not (yet) supported: flatten a 2-D table into one array plus an offsets array. A table of 192-bit constants is flattened the same way, three limbs an entry, and read back with a compile-time `i` (a `Const` parameter, say) as `f192(T[3 * i], T[3 * i + 1], T[3 * i + 2])`. ## Functions @@ -129,30 +133,30 @@ z = f(p, q) # expression position: first return f(p, q) # statement: returns discarded ``` -Functions may recurse. Each call gets a **fresh frame**: the frame pointer is prover-hinted (write-once memory makes an unconstrained cell prover-chosen), arguments and the return address/frame are stored with `DEREF`s, and control transfers with one `JUMP`. Cost: about `n_args + n_returns + 4` instructions per call. Every non-`main` function must end in an explicit `return`; in `main`, `return` is a no-op (main halts at a sentinel automatically). +Functions may recurse. Each call gets a **fresh frame**: the frame pointer is prover-hinted (write-once memory makes an unconstrained cell prover-chosen), arguments and the return address/frame are stored with `DEREF`s, and control transfers with one `JUMP`. Cost: about `n_args + n_returns + 4` instructions per call, counting a run as its cells. Every non-`main` function must end in an explicit `return`; in `main`, `return` is a no-op (main halts at a sentinel automatically). ### `StackBuf` parameters ```python -def compress(cv: StackBuf(2), block: StackBuf(2)): - out = StackBuf(2) - blake2s(cv, block, out) +def node(left: StackBuf(4), right: StackBuf(4)): + out = StackBuf(4) + blake2s(left, right, out) return out ``` -`s: StackBuf(n)` marks a parameter as a **run of n cells**, passed whole. The caller must pass a `StackBuf` of exactly that size (a whole named one: a slice is not yet accepted), and the run is copied into the callee's frame. +`s: StackBuf(n)` marks a parameter as a **run of n cells**, passed whole. The caller passes any run value of exactly that size (a `StackBuf`, a stack or heap slice, a list literal, a 192-bit value), and the run is copied into the callee's frame, one `DEREF` a cell. Those cells arrive **already written**, unlike a local `StackBuf`'s, so a store into one is the write-once equality *assertion* rather than a fresh store. That is what makes a callee able to pin its caller's values: `s[k] = ` inside the callee asserts that the caller's cell already held it. It also means a run parameter initializes nothing, so passing a partly-written buffer passes its unwritten cells, which the prover then chooses, exactly as anywhere else. -A `match` arm cannot pass one: the fused dispatch writes one cell per argument, so give such arms `Const` arguments and let each specialize instead. This is otherwise the same mechanism a `StackBuf` **return** already used, in the other direction: the argument area is a width rather than a count, and a run occupies the cells its size asks for. Without it a two-cell value could come out of a function whole but only go in through a `HeapBuf` pointer or an `@inline` expansion, which grows the caller's frame at every call site. +This is the same mechanism a `StackBuf` **return** uses, in the other direction: the argument area is a width rather than a count, and a run occupies the cells its size asks for. It is what lets a digest or a 192-bit value go into an ordinary function whole, rather than through a `HeapBuf` pointer or an `@inline` expansion, which grows the caller's frame at every call site. A `match` arm passes runs the same way. ### `Const` parameters ```python def hash_pair(buf, k: Const): - h = StackBuf(2) - blake2s(buf[k * 2:k * 2 + 2], buf[k * 2:k * 2 + 2], h) - return h[0], h[1] + h = StackBuf(4) + blake2s(buf[k * 4:k * 4 + 4], buf[k * 4:k * 4 + 4], h) + return h ``` `k: Const` marks a **compile-time parameter**: the call site must pass a constant (an integer literal, `GEN ** k`, or a literal-bound name), and the compiler *specializes* the function per distinct constant tuple, a monomorphized copy (`hash_pair__L1`) with the parameter substituted as its literal, shared by every call with the same constants; only the runtime arguments are passed. Inside the body the parameter *is* the literal, so it works in compile-time positions: stack indexes, slice bounds. A function with a `Const` parameter is a template: it is never lowered itself. The idiomatic pairing dispatches a runtime index to a const-indexed helper: @@ -177,9 +181,9 @@ def combine(a, b, k: Const): An `@inline` function is **expanded at each call site** instead of emitting a real call: no frame, no argument/return `DEREF`s, no call/return `JUMP`s. Its body must end with one top-level `return`. Builtins, ordinary calls, nested inline calls, `if`, and `unroll` are allowed; `mul_range` loops, `match`, tuple assignments within the body, and nested/early returns are rejected. Ordinary calls retain their own frames; nested inline calls expand recursively, with direct or indirect recursive expansion rejected. -An `@inline` function may also **return a `StackBuf`**: the caller's binding aliases the returned cell run (zero copies), and `StackBuf` arguments alias likewise. +An `@inline` function may also **return a run**: the caller's binding aliases the returned cell run (zero copies), and run arguments alias likewise. -An `@inline` call may also sit in **expression position**: embedded in arithmetic, as a store's RHS, or as a single-target `match` arm. An aliased return (a folded g-address) then materializes into a plain cell (free for a var; one `MUL` for a shifted pointer); a multi-cell `StackBuf` return still needs a `let` binding, since only a name can alias a cell run. +An `@inline` call may also sit in **expression position**: embedded in arithmetic, as a store's RHS, or as a single-target `match` arm. An aliased return (a folded g-address) then materializes into a plain cell (free for a var; one `MUL64` for a shifted pointer). A multi-cell run return is not a scalar: bind it with `let`, or use it where a run is expected (a 192-bit operand, a slice store, a `StackBuf` argument), as a real call's run return may be too. An `@inline` function that also takes a `Const` parameter and is used as a `match` arm is specialized rather than expanded: the fused dispatch enters one real function, so `@inline` is simply not honoured there. One without a `Const` parameter has no entry to dispatch to and is rejected. @@ -194,12 +198,12 @@ A name bound to an integer literal (`x = 2`) additionally acts as a **compile-ti Two families of binding are folded and carried **virtually**, costing no instruction until used as a value: - **g-powers and shifted pointers**: a cursor like `s = s * GEN` or a pointer view `p = buf * GEN ** k`. The offset folds into the `DEREF` address of each access; only a scalar use materializes it. -- **field constants**: a value built from literals / `GEN ** k` by field `+` and `*`, e.g. a running weight `w = w * CHAIN_LENGTH` in an unrolled loop. The arithmetic that advances it is compile-time (zero instructions); each use is one `SET` of the folded constant. +- **field constants**: a value built from literals / `GEN ** k` by field `+` and `*`, e.g. a running weight `w = w * CHAIN_LENGTH` in an unrolled loop. The arithmetic that advances it is compile-time (zero instructions); each use is one `SET` of the folded constant. A 192-bit constant (`a = f192(...)`, or a 192-bit builtin over constants) folds likewise, its three `SET`s emitted where a use needs cells. A store into a stack cell is NOT virtual: `sa[k] = other` always emits. If the cell already holds a value the store is the write-once equality *assertion* below, which is what makes `s[k] = ` pin a hint and a pre-written `blake2s` output verify a digest; if it does not, the store is what gives the cell its value. The compiler tracks nothing to tell those apart, the machine's write-once memory being what distinguishes them. ## Debugging -`print(expr)` / `print("label", expr)` displays a value at witness generation (prover side only, with no constraints and nothing entering the transcript). The label defaults to the argument's source text; output goes to stderr as `[print] label = ...`, showing the decimal reading for small integers, `g^k` when the value is a small g-power (both when they overlap: `8 (g^3)`), or `c2:c1:c0` hex otherwise, from the most significant limb to the least significant. Each print costs one anchor instruction, so the witness differs from a print-free build: strip prints before benchmarking. +`print(expr)` / `print("label", expr)` displays a word at witness generation (prover side only, with no constraints and nothing entering the transcript). The label defaults to the argument's source text; output goes to stderr as `[print] label = ...`, showing the decimal reading for small integers, `g^k` when the value is a small g-power (both when they overlap: `8 (g^3)`), or `0x…` hex otherwise. A run prints one cell at a time (`print(q[0])`). Each print costs one anchor instruction, so the witness differs from a print-free build: strip prints before benchmarking. ## Memory @@ -215,9 +219,9 @@ v = buf[i] # m[buf·i], i any runtime g-power (e.g. a loop counter) buf[i * GEN] = v # the next cell along ``` -The index is a field element; cell `k` of the buffer lives at address `buf · g^k`. A read or store is one `DEREF`. A **runtime** index costs one extra `MUL` for the `buf·i` pointer, but a **compile-time g-power** offset (`buf[1]`, `buf[GEN ** k]`, or a cursor advanced by `× GEN ** m`) folds into the `DEREF`'s address immediate for free: no `MUL`, no `SET`, and the cursor arithmetic itself vanishes (so a `× GEN` walk over consecutive cells is zero instructions). +The index is a field element; cell `k` of the buffer lives at address `buf · g^k`. A read or store is one `DEREF`. A **runtime** index costs one extra `MUL64` for the `buf·i` pointer, but a **compile-time g-power** offset (`buf[1]`, `buf[GEN ** k]`, or a cursor advanced by `× GEN ** m`) folds into the `DEREF`'s address immediate for free: no `MUL64`, no `SET`, and the cursor arithmetic itself vanishes (so a `× GEN` walk over consecutive cells is zero instructions). -**Compile-time indices are bounds-checked.** When the whole index is a compile-time exponent and the pointer resolves to a declared `HeapBuf` (directly, or through shifted aliases like `row = buf * GEN ** k`), the compiler rejects `index >= size`, and the same for the spans of `hint_witness` and `blake2s` slices. **Runtime** indices are not checked (their value is unknown at compile time): there the buffer remains a region convention, and a stray access surfaces at proving time as a write-once conflict or wild deref. +**Compile-time indices are bounds-checked.** When the whole index is a compile-time exponent and the pointer resolves to a declared `HeapBuf` (directly, or through shifted aliases like `row = buf * GEN ** k`), the compiler rejects `index >= size`, and the same for the span of every slice. **Runtime** indices are not checked (their value is unknown at compile time): there the buffer remains a region convention, and a stray access surfaces at proving time as a write-once conflict or wild deref. ### `StackBuf(n)`: frame-cell runs, indexed by compile-time integers @@ -230,7 +234,7 @@ v = sa[x + 1] # indexes: literals, literal-bound names, and + * // % of tg = [v, 7] # list literal: an initialized StackBuf, one cell per element ``` -A **list literal** `x = [a, b, …]` is an initialized `StackBuf`: it allocates one cell per element and writes each element in place, exactly the alloc-then-store idiom above, in one line. Elements are arbitrary runtime expressions; each write goes through the same stack-store path. It exists only as the RHS of a plain assignment inside a function; a *top-level* `NAME = [...]` is a constant array (see "Constant arrays"). The elements are lowered before the name rebinds, so `s = [s[1], s[0]]` swaps through the old binding. +A **list literal** `x = [a, b, …]` is an initialized `StackBuf`: it allocates one cell per scalar element and a run element's cells for each run element (so `[q, 7]` with `q` a 192-bit value is four cells), and writes each element in place, exactly the alloc-then-store idiom above, in one line. Elements are arbitrary runtime expressions; each write goes through the same stack-store path. Bound to a name it is the RHS of a plain assignment inside a function, and it is also a run value wherever a run is expected (a `blake2s` operand, a `StackBuf` argument, a 192-bit operand, a slice store); a *top-level* `NAME = [...]` is a constant array (see "Constant arrays"). The elements are lowered before the name rebinds, so `s = [s[1], s[0]]` swaps through the old binding. Stack indexes and slice bounds are **compile-time integers**, and index arithmetic (`+ * // %`) is *integer* arithmetic (`x + 1` above is 2, `k // 2` floor-divides, `k % 2` is a remainder: index space, not the field, where XOR is what `+` means and `//`/`%` have no meaning at all: using one as a runtime field value is a compile error). Bounds are checked at compile time. A `StackBuf` name is a run of cells, not a scalar: using it as one is an error, and it cannot be captured into a `for` loop body (carry state through a `HeapBuf` instead). @@ -240,13 +244,15 @@ A runtime index through such a pointer is unchecked, as on the heap, but it fail ### Slices: `buf[lo:hi]` -`buf[lo:hi]` names a run of cells (`hi` exclusive). BLAKE2s operands must span exactly two cells; `hint_witness` accepts any supported literal length. Two forms: +`buf[lo:hi]` names a run of cells (`hi` exclusive). A `blake2s` operand spans four cells, a 192-bit value three, and `hint_witness` accepts any length. Two forms: - **compile-time bounds** (integers, as for stack indexes): frame cells `base+lo .. base+hi` of a `StackBuf`, or heap cells `ptr·g^lo .. ptr·g^hi` of a `HeapBuf`, so `hb[2:4]` is the pair `g^2, g^3`; -- **runtime start, heap only**: `buf[i:i + k]` with a runtime g-power index `i` (e.g. a loop counter) and literal length `k` names the cells `buf·i`, `buf·i·g`, and so on; one `MUL` folds `i` into the pointer. The `hi` bound cannot be evaluated, only shape-checked: it must be syntactically `lo + k` (`buf[b * GEN ** 2 : b * GEN ** 2 + 2]` is fine). A `StackBuf` slice cannot have a runtime start: frame offsets are baked into the bytecode operands. +- **runtime start, heap only**: `buf[i:i + k]` with a runtime g-power index `i` (e.g. a loop counter) and literal length `k` names the cells `buf·i`, `buf·i·g`, and so on; one `MUL64` folds `i` into the pointer. The `hi` bound cannot be evaluated, only shape-checked: it must be syntactically `lo + k` (`buf[b * GEN ** 4:b * GEN ** 4 + 4]` is fine). A `StackBuf` slice cannot have a runtime start: frame offsets are baked into the bytecode operands. Note the two index spaces, consistent with plain indexing: compile-time bounds are integer exponents (`hb[2:4]` ≡ `hb[GEN ** 2 : GEN ** 2 + 2]`), runtime starts are g-power elements. +A slice is a value and a store target. Read as a value, a stack slice is used in place and a heap slice is copied onto the stack, one `DEREF` a cell. `buf[lo:hi] = value` stores a run of the same length: into a frame slice the value is written in place, into a heap slice through one `DEREF` a cell, and a cell already written makes its write the equality assertion, as for any store. + ## Control flow ### `for i in mul_range(start, stop)`: loops in the exponent @@ -280,7 +286,7 @@ for i in unroll(0, 7): def chain(buf, n: Const): for i in unroll(0, n): # a Const parameter as a bound - blake2s(buf[i * 2:i * 2 + 2], buf[i * 2:i * 2 + 2], buf[i * 2 + 2:i * 2 + 4]) + blake2s(buf[i * 4:i * 4 + 4], buf[i * 4:i * 4 + 4], buf[i * 4 + 4:i * 4 + 8]) return ``` @@ -297,9 +303,9 @@ else: r[1] = 9 ``` -Conditions are field-equality tests: `a == b` or `a != b` (there are no other predicates: order facts come from range-check asserts). The lowering is one `XOR` plus one conditional `JUMP` on it; the taken jump goes to whichever block the test doesn't fall into, so no negation gadget is needed. An `elif` is sugar for an `else` holding a nested `if`. +Conditions are field-equality tests on words: `a == b` or `a != b` (there are no other predicates: order facts come from range-check asserts). The lowering is one `XOR64` plus one conditional `JUMP` on it; the taken jump goes to whichever block the test doesn't fall into, so no negation gadget is needed. An `elif` is sugar for an `else` holding a nested `if`. -When **both sides are compile-time integers** (e.g. after a `Const` parameter is substituted, `if k % 2 == 0:`), the condition is known at compile time and the `if` **folds** to just the taken branch: no `XOR`, no `JUMP`, no `self-fp`. This is what lets an `@inline` function bake different straight-line code per `Const` value. A side whose integer reading and field reading disagree (`3 + 1` is the integer 4 and the field element 2) is **rejected** rather than folded either way, since the fold and a runtime test of the same condition would answer differently; write `if const(...)` below to decide it with integer arithmetic. Note that the rejection is per side, so it fires however the OTHER side is spelled. +When **both sides are compile-time integers** (e.g. after a `Const` parameter is substituted, `if k % 2 == 0:`), the condition is known at compile time and the `if` **folds** to just the taken branch: no `XOR64`, no `JUMP`, no `self-fp`. This is what lets an `@inline` function bake different straight-line code per `Const` value. A side whose integer reading and field reading disagree (`3 + 1` is the integer 4 and the field element 2) is **rejected** rather than folded either way, since the fold and a runtime test of the same condition would answer differently; write `if const(...)` below to decide it with integer arithmetic. Note that the rejection is per side, so it fires however the OTHER side is spelled. Two write-once-flavored rules: @@ -317,9 +323,9 @@ a, b = match(log(x), range(0, 2), lambda j: g(1), range(2, 6), lambda j: g(j)) The one dispatch construct. It matches the **log** of a g-power scrutinee against integer arms, which must cover consecutive integers from 0 (the dispatch table is dense; there is no default arm). Arm `j` is the lambda body with the parameter replaced by the **integer literal** `j`, usable as a field constant or a compile-time index, expanded at parse time over the contiguous `(range, lambda)` pairs. The whole call sits on one line, there being no line continuation. -Arms produce VALUES: every arm writes its results into the same cells, which is sound under write-once because exactly one arm runs. A target may be a name, bound after the join, or a **`StackBuf` element**, which the arms write into directly and which costs one instruction less than a name plus a store. The ABI returns into cells the CALLER picks, the same reason `sb[i] = f(x)` never needed a temporary, so reach for the element form wherever a returned value's home is a buffer slot. A target index must be a compile-time integer inside the buffer, both errors naming the line; a `HeapBuf` element is not a target, its cells not being frame cells. Multiple targets take a multi-return call as the arm body. +Arms produce VALUES: every arm writes its results into the same cells, which is sound under write-once because exactly one arm runs. A target may be a name, bound after the join, or a **`StackBuf` element**, which the arms write into directly and which costs one instruction less than a name plus a store. The ABI returns into cells the CALLER picks, the same reason `sb[i] = f(x)` never needed a temporary, so reach for the element form wherever a returned value's home is a buffer slot. A target index must be a compile-time integer inside the buffer, both errors naming the line; a `HeapBuf` element is not a target, its cells not being frame cells. Multiple targets take a multi-return call as the arm body. A run return binds a name, and crosses the join only through the fused dispatch below. -A branch body with statements in it goes in a function, and the arm calls it: that is the idiom the recursion guest uses throughout (`lambda k: walk(chain_start, tweaks, pp, k)`), and it names the body instead of inlining it. Where the arms only PRODUCE values, as there, this costs nothing. Where each arm's real work is a WRITE, it costs: the writer function needs a return value and the statement a target, both dead, and the natural translation measures about twice the instructions of a body inlined into the dispatching frame. An arm cannot take a `StackBuf` parameter either, so it cannot hash or hint into the caller's frame buffer. If that shape matters to a program, dispatch on a value and write after the join. +A branch body with statements in it goes in a function, and the arm calls it: that is the idiom the recursion guest uses throughout (`lambda k: walk(chain_start, tweaks, pp, k)`), and it names the body instead of inlining it. Where the arms only PRODUCE values, as there, this costs nothing. Where each arm's real work is a WRITE, it costs: the writer function needs a return value and the statement a target, both dead, so the natural translation costs more instructions than a body inlined into the dispatching frame. An arm may pass runs to its callee, but a run parameter is the callee's copy, so a store into it asserts against the caller's cells rather than filling them. If that shape matters to a program, dispatch on a value and write after the join. **Lowering** is two jumps through a *trampoline table* in the bytecode: the dispatch jumps to `g^T · x²`, the j-th two-instruction slot (`SET` the arm's address, `JUMP` to it) of a table at base `T`, and the slot jumps to the arm, which can sit anywhere, unaligned and of any length. Cost is about 7 cycles, independent of the arm count. @@ -327,7 +333,7 @@ A branch body with statements in it goes in a function, and the arm calls it: th **Soundness**: nothing in the dispatch bounds `x`, so a scrutinee outside `[0, n)` jumps to an arbitrary pc. A hinted value must be range-checked first (`assert log(x) < n`, 3 cycles), as in leanVM. -**Dispatched-call fusion.** When *every* arm is a call to the same function with identical runtime arguments (the common `lambda k: f(a, b, k)`, where only a `Const` argument varies), the compiler builds the callee frame **once** and the dispatch jumps straight into the selected specialization's entry, which returns past the join. Each taken arm is then just the trampoline's two instructions (`SET entry; JUMP`) instead of a full call: no per-arm frame setup, call jump, or return jump. (The `walk`-per-digit dispatch in the XMSS verifier is the motivating case.) +**Dispatched-call fusion.** When *every* arm is a call to the same function with identical runtime arguments (the common `lambda k: f(a, b, k)`, where only a `Const` argument varies), the compiler builds the callee frame **once** and the dispatch jumps straight into the selected specialization's entry, which returns past the join. Each taken arm is then just the trampoline's two instructions (`SET entry; JUMP`) instead of a full call: no per-arm frame setup, call jump, or return jump. The arms share that one frame, so every callee must take and return the same shapes, runs included, and this is the path that lets a run come back out of a `match` (`h = match(log(x), range(0, 4), lambda i: hash_pair(buf, i))`). (The `walk`-per-digit dispatch in the XMSS verifier is the motivating case.) Statements without effect are rejected. @@ -352,7 +358,7 @@ tweak = TW_NODE + const((level + 1) * P_MUL) + tau # (level+1)*P_MUL as intege The same wrapper, the same meaning: read this with **integer** arithmetic and emit the literal. It is needed because `+` in a value position is XOR, so `level + 1` with `level = 3` is 2 rather than 4, and silently: the value is well-formed, just not the one the arithmetic reads like. `-`, `//` and `%` have no field meaning at all, so `const(...)` is the only way to write them in a value position. -The inner expression must be a compile-time integer (a literal, a global constant, a `Const` parameter, an `unroll` counter, a name bound to one, a constant-array element, and `+ - * // % **` of those), and one that is not says so rather than falling back to a runtime computation. The result is one pooled `SET`, so a repeat costs nothing. +The inner expression must be a compile-time integer (a literal, a global constant, a `Const` parameter, an `unroll` counter, a name bound to one, a constant-array element, and `+ - * // % **` of those) that fits in a 64-bit word, and one that is not says so rather than falling back to a runtime computation. The result is one pooled `SET`, so a repeat costs nothing. In a position that is ALREADY integer arithmetic (a size, a count, an exponent, a bound, a stack index, a global constant) the wrapper is transparent: it asks for the only reading there is, so it changes nothing and is allowed rather than redundant. Where it earns its keep is a value, a condition, and anywhere `-`, `//` or `%` has to appear. @@ -362,11 +368,11 @@ The wrapper reinterprets the **operators**, not the leaves, and that is the whol ### `assert a == b` -A proof-enforced equality: 1 cycle (`XOR` into the frame's zero cell, whose write-once double write is the assert). +A proof-enforced equality of two words: 1 cycle (`XOR64` into the frame's zero cell, whose write-once double write is the assert). ### `assert a != b` -A proof-enforced inequality, in **3 instructions and no branch**: `XOR` for `x = a + b`, a prover-hinted `inv = x⁻¹`, then `MUL p = x·inv` and `SET p = 1`, where the write-once conflict is the assertion, exactly as for `assert a == b`. It is sound because `x = 0` forces `p = 0` whatever the prover hints, and `p` cannot then also be `1`; the hint needs no checking of its own, which is why an unconstrained value is safe here. Since there is no `JUMP` there is no self-frame or branch setup to amortize either. A compile-time assertion such as `assert 5 != 5` is rejected while compiling. +A proof-enforced inequality, in **3 instructions and no branch**: `XOR64` for `x = a + b`, a prover-hinted `inv = x⁻¹`, then `MUL64 p = x·inv` and `SET p = 1`, where the write-once conflict is the assertion, exactly as for `assert a == b`. It is sound because `x = 0` forces `p = 0` whatever the prover hints, and `p` cannot then also be `1`; the hint needs no checking of its own, which is why an unconstrained value is safe here. Since there is no `JUMP` there is no self-frame or branch setup to amortize either. A compile-time assertion such as `assert 5 != 5` is rejected while compiling. ### Range checks: `assert log x < log Y` and `assert log x < k` @@ -378,85 +384,89 @@ assert log(x) < 8 # the same check assert log(x) < log(n) # n = g^k runtime: same gadget, +1 cycle ``` -A **runtime** bound costs one extra `MUL` for `g^{k-1} = n·g⁻¹` and is otherwise identical, except that the `k ≤ 2^16` cap becomes the program's to enforce: range-check the bound itself first, with `assert log n < 2^16`. That check is not optional: without it the gadget is unsound. +A **runtime** bound costs one extra `MUL64` for `g^{k-1} = n·g⁻¹` and is otherwise identical, except that the `k ≤ 2^16` cap becomes the program's to enforce: range-check the bound itself first, with `assert log n < 2^16`. That check is not optional: without it the gadget is unsound. Cost: **3 cycles** (leanVM's DEREF range-check trick, in the exponent) plus one amortized `SET` per distinct bound per frame: 1. `DEREF` through `x`: the dereferenced address must be one of the memory's `2^h` g-power addresses, so the memory bus itself proves `x = g^e`, `e < 2^h`; -2. `MUL x·y` into the write-once cell holding `g^{k-1}`: the runner back-solves the complement `y = g^{k-1-e}` (the one unknown operand of a known product), and the double-write asserts `x·y = g^{k-1}`; +2. `MUL64 x·y` into the write-once cell holding `g^{k-1}`: the runner back-solves the complement `y = g^{k-1-e}` (the one unknown operand of a known product), and the double-write asserts `x·y = g^{k-1}`; 3. `DEREF` through `y`: bounds the complement; a "negative" `k-1-e` would wrap to `≈ 2^64`, far beyond any memory size, so together `e ≤ k-1`. The two `DEREF` target cells are unconstrained touches, back-filled at the end of execution. A failing check surfaces at witness generation as the complement's `DEREF` panic ("not a small g-power … a failed range check"). -## K membership and packing +## 192-bit values -```python -assert_in_k(lo, hi) -``` - -`assert_in_k(a, b)` is the sole packing-related compiler intrinsic. It proves that both source memory words are in the base field GF(2^64) with one untaken `JUMP`: its condition is a known zero, while its destination and frame operands are `a` and `b`. Although neither value affects the successor state, both memory reads carry literal-zero upper limbs, so a source outside GF(2^64) cannot balance the memory permutation. - -Packing is ordinary zkDSL built on that assertion: +A memory cell holds one 64-bit word, so an element `c0 + c1·y + c2·y^2` of `E = GF(2^192)` is a **run** of three consecutive cells, its limbs low first. `+` and `*` stay 64-bit; the 192-bit instructions `XOR192` and `MUL192`, each reading two three-cell operands and writing a three-cell result, are reached through builtins: ```python -@inline -def pack64x2(a, b): - assert_in_k(a, b) - return a + f192(0, 1, 0) * b -``` +ONE = f192(1, 0, 0) # a global 192-bit constant + -The inline helper takes three cycles, one `JUMP`, one `MUL` and one `XOR`, and returns the canonical 128-bit packing `(a.c0, b.c0, 0)`. Assignment-target lowering writes its return directly into the destination, including an already-written cell whose second write is an equality assertion. A caller needing only membership uses `assert_in_k` directly and pays no packing arithmetic. +def main(): + a = f192(3, 5, 7) # 3 + 5·y + 7·y^2: folded, no cells yet + b = [GEN ** 9, 11, 13] # three words written into a fresh run + x = StackBuf(3) + hint_witness(x, "x") # one entry of three words, limbs low first + heap = HeapBuf(6) + heap[0:3] = mul192(a, b) # a run store into heap cells g^0..g^2 + q = div192(add192(heap[0:3], x), a) # back-solved, pinned by one MUL192 + assert_eq192(add192(q, b), div192(x, a)) + assert_ne192(q, ONE) + heap[3:6] = square(q) # a run goes into a call and comes back out + return -The recursion transcript uses `challenge_from_state(state)` to reinterpret the first three 64-bit lanes of a canonical two-cell BLAKE2s digest as one extension field challenge. For `state = [s0, s1]`, it lowers exactly as follows (the limb hints cost no cycles, but are not trusted): -```python -d2 = StackBuf(1) -hint_f192_limbs(d2, state[1]) -d3 = (state[1] + d2[0]) * Y_INV -assert_in_k(d2[0], d3) -challenge = state[0] + d2[0] * f192(0, 0, 1) +def square(v: StackBuf(3)): + return mul192(v, v) ``` -Both state words are BLAKE2s outputs, so their top limbs are already zero. Only `d2` must be exposed separately; deriving `d3 = (s1+d2)/Y` and proving both values lie in GF(2^64) binds the one hinted limb by uniqueness of the tower representation. The challenge is `s0+d2·Y² = d0+d1·Y+d2·Y²`, while `d3` is checked but deliberately discarded. `challenge_from_state` is not a compiler intrinsic: this is the complete `@inline` helper used by the recursion guest. +- **A run is** a `StackBuf(3)`, a three-cell slice of a `StackBuf` or a `HeapBuf` (a runtime heap start `buf[i:i + 3]` included), a list literal of three words (a run element flattens into its cells), `f192(c0, c1, c2)`, a 192-bit builtin, or a call returning a `StackBuf(3)`. A name bound to a non-constant one names its cells, so `q[0]` is its low limb as an ordinary word. +- `add192(a, b)` is one `XOR192` and `mul192(a, b)` one `MUL192`. `div192(a, b)` is one `MUL192` whose unwritten operand is the quotient: witness generation back-solves `q = a·b⁻¹` and the instruction pins `q·b == a`. A zero divisor has no quotient, and witness generation refuses it. +- `assert_eq192(a, b)` is one `XOR192` into a pooled zero run, whose double write is the assertion. `assert_ne192(a, b)` is `x = a + b`, a hinted inverse, and one `MUL192 x·inv` into a pooled one run, sound for the reason `assert a != b` is. +- **Constants fold.** A builtin over constants is computed while compiling, `add192(a, 0)`, `mul192(a, 1)` and `div192(a, 1)` reduce to `a`, and `mul192(a, 0)` to zero, all with no instruction. A constant run costs three pooled `SET`s where a use first needs cells. +- **A run and a scalar never stand in for each other.** A word where a run is expected, a run where a word is expected (`+`, `*`, `assert`, `print`), a run of the wrong width, and a 192-bit builtin used as a statement are compile errors naming the widths. +- **Moving runs.** `buf[lo:hi] = value` stores one, an `x: StackBuf(3)` parameter takes one, a function returns one, and a fused `match` returns one. A copy between frame runs is one `XOR192` against the pooled zero run. + +A slice of a digest is a run like any other, so the recursion guest takes a transcript challenge as the first three words of its BLAKE2s state, `state[0:3]`. ## BLAKE2s ```python -h = StackBuf(2) -blake2s(a, b, h) # digest of (a, b) written into h -blake2s(t[0:2], t[x:x + 2], t[4:6]) # slices of one large StackBuf -blake2s(h, hb[0:2], hb[2:4]) # HeapBuf slices, input and output -blake2s(hb[i:i + 2], h, hb[j:j + 2]) # runtime-indexed heap slices (i, j g-powers) +h = StackBuf(4) +blake2s(a, b, h) # digest of (a, b) written into h +blake2s(t[0:4], t[x:x + 4], t[8:12]) # slices of one large StackBuf +blake2s(h, hb[0:4], hb[4:8]) # HeapBuf slices, input and output +blake2s(hb[i:i + 4], h, hb[j:j + 4]) # runtime-indexed heap slices (i, j g-powers) +blake2s([tag, 0, pair], h, out) # a list operand: four words, the run `pair` flattening # A standard 80-byte hash as two blocks. Keyword values are compile-time. -block0 = [1, 2, 3, 4] # 64 bytes -tail = [5, 0, 0, 0] # 16 more, the rest of the block zero-filled -blake2s(block0[0:2], block0[2:4], cv, counter=64, final=0) -blake2s(tail[0:2], tail[2:4], out, cv=cv, counter=80, final=1) +block0 = [1, 0, 2, 0, 3, 0, 4, 0] # 64 bytes, eight words +tail = [5, 0, 0, 0, 0, 0, 0, 0] # 16 more, the rest of the block zero-filled +blake2s(block0[0:4], block0[4:8], cv, counter=64, final=0) +blake2s(tail[0:4], tail[4:8], out, cv=cv, counter=80, final=1) # The same, with the second block's metadata computed at run time. -blake2s(tail[0:2], tail[2:4], out, cv=cv, md=high + f192(16, 4294967295, 0)) +blake2s(tail[0:4], tail[4:8], out, cv=cv, md=[high + 16, 4294967295]) ``` -The three positional arguments form a **statement**: one standard BLAKE2s compression consumes the two 256-bit message operands `a`, `b` (64 bytes) and writes its 32-byte result into the 2-cell run `out`. With no keywords it computes the standard hash of exactly 64 bytes: the parameterized BLAKE2s-256 initial chaining value (digest length 32, unkeyed, fanout and depth 1), byte counter 64, final-block flag `f0` set. That is `blake2s(a || b)`, the form every Fiat-Shamir step and Merkle node uses. +The three positional arguments form a **statement**: one standard BLAKE2s compression consumes the two 256-bit message operands `a`, `b` (64 bytes, four words each, little-endian) and writes its 32-byte result into the 4-cell run `out`. With no keywords it computes the standard hash of exactly 64 bytes: the parameterized BLAKE2s-256 initial chaining value (digest length 32, unkeyed, fanout and depth 1), byte counter 64, final-block flag `f0` set. That is `blake2s(a || b)`, the form every Fiat-Shamir step and Merkle node uses. -Every compression also has a 256-bit chaining value and a 128-bit metadata word. The optional keywords are: +Every compression also has a 256-bit chaining value and two metadata words. The optional keywords are: -- `cv=`: a consecutive 2-cell chaining value, the previous block's output; omitting it selects the parameterized IV above. On each runtime path, a function emits two `SET`s at its first such hash only and reuses those cells thereafter. Supplying `cv=` also requires one of the four below, since a chained block is never the default one-block hash; +- `cv=`: a 4-cell chaining value, the previous block's output; omitting it selects the parameterized IV above, four pooled `SET`s that every hash on the same runtime path of a function shares (a branch's copy does not survive its join). Supplying `cv=` also requires one of the four below, since a chained block is never the default one-block hash; - `counter=`: BLAKE2s's byte counter `t`, **cumulative** through this block, so `64 * whole_blocks_before + bytes_in_this_block`. Defaults to 64; - `final=<0|1>`: BLAKE2s's final-block flag `f0`. It defaults to 1 for the bare three-argument call, but to **0** as soon as `counter=` or `last_node=` appears, so a chained hash must set `final=1` on its last block and a single short block needs `counter=, final=1`. Any compile-time expression works, nonzero meaning set, which is what lets the guests write a predicate like `final=(q + 1) // BLOCKS_PER_HASH`; - `last_node=<0|1>`: BLAKE2s's tree-mode flag `f1`. Defaults to 0, and nothing here uses tree mode; -- `md=`: the whole 128-bit metadata word, as a value the program computed, for a hash whose block count is only known at run time. It replaces the three keywords above (giving both is an error) and it owes the same canonical embedding as every other operand, its top limb being read as a literal zero. The cheap way to build one is the disjoint-bit split of `doc/leanvm` §Byte counters for a hash of runtime length: XOR a runtime high part against a compile-time `metadata(64·j, f0, f1)` constant, one instruction per block. +- `md=`: the whole metadata as a 2-cell run the program computed, for a hash whose block count is only known at run time. It replaces the three keywords above (giving both is an error) and must not overlap `out`. The cheap way to build one is the disjoint-bit split of `doc/leanvm` §Byte counters for a hash of runtime length: XOR a runtime high part of the counter against its compile-time low part, `md=[high + 16, 4294967295]` above, one instruction per block. -The metadata is packed as `counter:u64 | f0:u32 | f1:u32`, little-endian, into one memory cell the instruction reads, like every other operand. With compile-time keywords that cell is one pooled `SET`: a frame emits it once per distinct metadata value, however many compressions read it, and the immediate that wrote it is public bytecode. There is no block-length field: the counter is what states how many of the 64 bytes are message, so only the last block may be partial and the program must zero-fill the bytes past its real length, which the compression circuit does not enforce. A multi-block hash therefore feeds each result back with `cv=`, advances `counter=` by the bytes actually absorbed, and sets `final=1` on the last block. +The metadata is two words, the counter `t` as a `u64` and then `f0 | f1 << 32` with each flag a `u32`, in two consecutive memory cells the instruction reads like every other operand. With compile-time keywords that pair is a pooled constant run: a frame emits its two `SET`s once per distinct metadata value, however many compressions read it, and the immediates that wrote it are public bytecode. There is no block-length field: the counter is what states how many of the 64 bytes are message, so only the last block may be partial and the program must zero-fill the bytes past its real length, which the compression circuit does not enforce. A multi-block hash therefore feeds each result back with `cv=`, advances `counter=` by the bytes actually absorbed, and sets `final=1` on the last block. -Operands are size-2 `StackBuf`s or 2-cell slices: +Operands are 4-cell runs: -- an **input operand written as a list**, `blake2s([a, b], [c, d], out)`, names its two words directly and allocates nothing: the opcode addresses its four input chunks independently, so an operand whose words live in different places never has to be gathered into a consecutive run. This is the spelling to reach for instead of `p = StackBuf(2); p[0] = a; p[1] = b`; -- **stack operands** are read in place, at zero copies; a self-hash `blake2s(h, h, out)` names one 2-cell pair as both inputs; -- the instruction addresses its **four canonical 128-bit message chunks independently** (each is a full F192 memory cell constrained at this use to the BLAKE2s subspace `c2 = 0`), so an operand gathered into a buffer (`p = StackBuf(2); p[0] = t0; p[1] = t1; blake2s(p, …)`) costs one instruction per assembling store, which the list form above avoids entirely; -- the chaining value has only one opcode offset and therefore must be consecutive. If a 2-cell `cv` was assembled from non-adjacent copied cells, the compiler materializes those two cells into a fresh consecutive run; -- **heap slices** are still bridged through the stack for the *input pull* (the operand's words come from the heap): +1 `DEREF` per heap cell, and the output, if a heap slice, is stored after: write-once memory fills whichever side is unset. +- an **input operand written as a list**, `blake2s([a, b, c, d], [e, f, g, k], out)`, names its four words directly. The opcode addresses its four 128-bit input chunks, two words each, independently, so a chunk whose two words already sit side by side is read in place, a constant chunk is a pooled pair of `SET`s, and only a chunk mixing words from different places is copied into a fresh pair. A run element flattens into its words. This is the spelling to reach for instead of gathering an operand into a `StackBuf` one store at a time; +- **stack operands** are read in place, at zero copies; a self-hash `blake2s(h, h, out)` names one 4-cell run as both inputs; +- the chaining value has only one opcode offset and therefore must be four consecutive cells: a `cv` written as a list of words from different places is gathered into a fresh run first; +- **heap slices** are bridged through the stack for the *input pull* (the operand's words come from the heap): +1 `DEREF` per heap cell, and the output, if a heap slice, is stored after: write-once memory fills whichever side is unset. If `out` was already written, the statement *asserts* the digest equals it, write-once turning the hash into a verification, which is exactly what a signature verifier wants. @@ -480,9 +490,9 @@ assert log m < 8 # still unconstrained: pin it which is the one-line form of allocating a `StackBuf(1)`, filling a slice of it, and reading the cell back out, and costs exactly the same (nothing). Everything below about a stream's entries applies to it: each such binding pops one entry, whose length must be 1. -Prover-supplied data (leanVM's `hint_witness`): a stream is a sequence of **entries**, one slice of values per `hint_witness` call, and the same symbol may be hinted many times. Each call pops the stream's next entry (whose length must match the destination run) and writes it into `dest` through the hint mechanism, at **zero cycles**. The values are completely unconstrained; the program must constrain them itself (asserts, range checks, hashes): an unconstrained hint consumed by anything security-relevant is a critical vulnerability. Runtime-start heap slices (`buf[i:i + k]`, `k` a literal) work too. +Prover-supplied data (leanVM's `hint_witness`): a stream is a sequence of **entries**, one slice of words per `hint_witness` call, and the same symbol may be hinted many times. Each call pops the stream's next entry (whose length must match the destination run) and writes it into `dest` through the hint mechanism, at **zero cycles**. The values are completely unconstrained; the program must constrain them itself (asserts, range checks, hashes): an unconstrained hint consumed by anything security-relevant is a critical vulnerability. Runtime-start heap slices (`buf[i:i + k]`, `k` a literal) work too. A 192-bit value is three words of an entry, limbs low first. -The prover supplies streams with `program.set_witness("name", entries)` (`Vec>`); test programs declare them as annotations, one line per entry, and repeated lines with the same name are its successive entries: +The prover supplies streams with `program.set_witness("name", entries)` (`Vec>`); test programs declare them as annotations, one line per entry, and repeated lines with the same name are its successive entries: ```python # witness r: GEN ** 5, 12 @@ -493,38 +503,44 @@ The prover supplies streams with `program.set_witness("name", entries)` (`Vec` / `GEN ** k` | 1 `SET` | -| `a + b` | 1 `XOR` | -| `a * b` | 1 `MUL` | -| `a / b` | 1 `MUL` (write-once back-solve; division by zero is undefined) | -| heap read / store `buf[i]` | 1 `DEREF`; +1 `MUL` for a *runtime* index (a compile-time g-power offset folds into the `DEREF`, for free) | +| `a + b` | 1 `XOR64` | +| `a * b` | 1 `MUL64` | +| `a / b` | 1 `MUL64` (write-once back-solve; division by zero is undefined) | +| heap read / store `buf[i]` | 1 `DEREF`; +1 `MUL64` for a *runtime* index (a compile-time g-power offset folds into the `DEREF`, for free) | | stack read `sa[k]` | 0 (direct cell addressing); a *store* is 1, like any other write | +| `add192(a, b)` / `mul192(a, b)` / `div192(a, b)` | 1 `XOR192` / 1 `MUL192` / 1 `MUL192`; **0 when the operands fold** | +| `f192(c0, c1, c2)` | 0 bound to a name; 3 pooled `SET`s where a use first needs cells | +| heap run read / store `buf[lo:hi]` | 1 `DEREF` per cell (+1 `MUL64` for a runtime start) | +| run copy between frame cells | 1 `XOR192` per three cells | | `assert a == b` | 1 (+ 1 `SET` amortized per frame for the zero cell) | -| `assert a != b` | 3 (`XOR`, `MUL`, `SET`), no branch, one hinted inverse | -| `assert log x < k` | 3 (+1 `SET` amortized per bound per frame; a runtime bound costs 1 `MUL` instead) | +| `assert a != b` | 3 (`XOR64`, `MUL64`, `SET`), no branch, one hinted inverse | +| `assert_eq192(a, b)` | 1 `XOR192` (+ 3 `SET`s amortized per frame for the zero run) | +| `assert_ne192(a, b)` | 2 (`XOR192`, `MUL192`), one hinted inverse (+ 3 `SET`s amortized per frame for the one run) | +| `assert log x < k` | 3 (+1 `SET` amortized per bound per frame; a runtime bound costs 1 `MUL64` instead) | | `if a == b: …` | 3 (+2 to skip a non-empty `else`; +2 amortized `self-fp` per branching function); **0 if the condition is compile-time** | | `… = match(log(x), …)` | ≈ 7 for the dispatch + the arm; results written into the targets directly. Uniform-call arms (`lambda k: f(a, b, k)`) **fuse**: one shared frame + dispatch to entry, each arm just `SET`+`JUMP` | -| function call | ≈ `n_args + n_returns + 4` (0 when the callee is `@inline`) | -| `mul_range` iteration | body + ≈ 1 `MUL` + 1 `XOR` + call overhead | +| function call | ≈ `n_args + n_returns + 4`, a run counting its cells (0 when the callee is `@inline`) | +| `mul_range` iteration | body + ≈ 1 `MUL64` + 1 `XOR64` + call overhead | | `unroll` iteration | body only (compile-time replication) | -| `blake2s(a, b, out, ...)` | 1; plus one `SET` once per frame per distinct metadata value (nothing with `md=`, which costs whatever building the word costs), and two more when `cv` is omitted; message/CV words are read in place, +1 `DEREF` per heap input or CV word, +1 `MUL` per runtime slice start | -| `hint_witness(dest, "name")` | 0 (+1 `MUL` for a runtime slice start) | +| `blake2s(a, b, out, ...)` | 1; plus two `SET`s once per frame per distinct metadata value (nothing with `md=`, which costs whatever building the pair costs), and four more when `cv` is omitted; message/CV words are read in place, +1 `DEREF` per heap input or CV word, +1 `MUL64` per runtime slice start, +1 per word of a list chunk that has to be gathered | +| `hint_witness(dest, "name")` | 0 (+1 `MUL64` for a runtime slice start) | -Every cost above is the FIRST occurrence. Two identical pure operations in one function share one cell and the second is free, so `hb[i]` twice, or `row[i]` where `row = hb * GEN ** 2`, costs one pointer `MUL` between them. The sharing stops at a branch: a cell whose instruction sits inside an `if` is not reused after the join, because the other path leaves it unwritten and therefore prover-chosen. +Every cost above is the FIRST occurrence. Two identical pure operations in one function share one cell and the second is free, so `hb[i]` twice, or `row[i]` where `row = hb * GEN ** 2`, costs one pointer `MUL64` between them. The sharing stops at a branch: a cell whose instruction sits inside an `if` is not reused after the join, because the other path leaves it unwritten and therefore prover-chosen. ## Example -Fibonacci in the exponent (`tests/programs/fibonacci.py`): `fib[g^k]` holds `GEN ** F_k`, so one field `MUL` is one Fibonacci step. +Fibonacci in the exponent (`tests/programs/fibonacci.py`): `fib[g^k]` holds `GEN ** F_k`, so one field `MUL64` is one Fibonacci step. ```python # public_input: GEN ** 89, GEN ** 89 @@ -548,4 +564,4 @@ def main(): ## Not (yet) supported -Mutable variables; conditions other than field (in)equality; `match` default and non-contiguous arms; multi-file imports; `Const` parameters as `mul_range` or range-check bounds (a substituted literal is a bit-pattern element, not the g-power a bound needs); runtime slice starts on a `StackBuf`; precompiles beyond `BLAKE2s`. +Mutable variables; conditions other than field (in)equality of words; `match` default and non-contiguous arms; multi-file imports; `Const` parameters as `mul_range` or range-check bounds (a substituted literal is a bit-pattern element, not the g-power a bound needs); runtime slice starts on a `StackBuf`; 192-bit arithmetic through `+` and `*`; precompiles beyond `BLAKE2s`. From eae9bfe603220611df0dd62e10d1c6c8e8d52203 Mon Sep 17 00:00:00 2001 From: Tom Wambsgans Date: Mon, 14 Sep 2026 20:51:55 -0400 Subject: [PATCH 03/40] Describe 64-bit memory in the leanVM document The specification, bus, instruction tables, bytecode encoding, public-input binding, programming and statement sections describe 64-bit memory words, the XOR192 and MUL192 tables, the eighteen-word BLAKE2S row and the two-challenge public-input claim. Co-Authored-By: Claude Opus 5 (1M context) Claude-Session: https://claude.ai/code/session_017HVkj8EFExoGGa1okY9gvX --- doc/leanvm/body/02-vm-specification.tex | 30 ++--- doc/leanvm/body/06-bus-interactions.tex | 12 +- doc/leanvm/body/07-instruction-tables.tex | 124 +++++++++++++-------- doc/leanvm/body/08-end-to-end-protocol.tex | 30 ++--- doc/leanvm/body/09-isa-programming.tex | 20 ++-- doc/leanvm/body/10-ethereum.tex | 6 +- doc/leanvm/preamble/macros.tex | 5 +- 7 files changed, 136 insertions(+), 91 deletions(-) diff --git a/doc/leanvm/body/02-vm-specification.tex b/doc/leanvm/body/02-vm-specification.tex index 6217b8ccc..9a2a051d0 100644 --- a/doc/leanvm/body/02-vm-specification.tex +++ b/doc/leanvm/body/02-vm-specification.tex @@ -21,16 +21,16 @@ \section{VM specification}\label{sec:vm} \item the frame pointer $\fp \xleftarrow{} 1_\K$ \end{itemize} -Then, initialize the write-once memory: a buffer, of size $\nmem = 2^{\kmem}$, containing $\E$-valued \textit{memory words} (192 bits). The memory is indexed \textit{in the exponent} of $\gen$ (64 bits). The first memory word is $\mem[\gen^0] \in \E$, the second is $\mem[\gen^1]$, the third is $\mem[\gen^2]$, etc. The bytecode addressing follows the same convention. Each memory access beyond $\gen^{\nmem - 1}$ (resp. $\gen^{\nprog - 1}$ for the bytecode) is invalid. +Then, initialize the write-once memory: a buffer, of size $\nmem = 2^{\kmem}$, containing $\K$-valued \textit{memory words} (64 bits). The memory is indexed \textit{in the exponent} of $\gen$ (64 bits). The first memory word is $\mem[\gen^0] \in \K$, the second is $\mem[\gen^1]$, the third is $\mem[\gen^2]$, etc. The bytecode addressing follows the same convention. Each memory access beyond $\gen^{\nmem - 1}$ (resp. $\gen^{\nprog - 1}$ for the bytecode) is invalid. -The public input fixes the first two memory words $\mem[\gen^0] = \textsf{input}_0 + \textsf{input}_1 \cdot y \in \E$ and $\mem[\gen^1] = \textsf{input}_2 + \textsf{input}_3 \cdot y \in \E$. +The public input fixes the first four memory words, $\mem[\gen^i] = \textsf{input}_i \in \K$ for $i<4$. The Virtual Machine execution loop is the following: \begin{enumerate} \item fetch the instruction $\mathsf{inst}$ at address $\pc$. \item execute $\mathsf{inst}$: \begin{itemize} - \item assert the memory constraints (e.g. $\loc{o_C}=\loc{o_A}+\loc{o_B}$ for XOR) + \item assert the memory constraints (e.g. $\loc{o_C}=\loc{o_A}+\loc{o_B}$ for \texttt{XOR64}) \item update $\pc$ and $\fp$ (all instructions except \texttt{JUMP} set $\pc \xleftarrow{} \pc \cdot \gen$, advancing to the next instruction, and leave $\fp \xleftarrow{} \fp$ unchanged) \end{itemize} \item if $\pc$ equals $\pc_{\textsf{final}} := \gen^{\nprog - 1}$, exit the loop (successful execution). Otherwise go back to $1.$ @@ -48,23 +48,29 @@ \section{VM specification}\label{sec:vm} In particular, the memory log-size $\kmem$ is a non deterministic parameter, freely chosen at the beginning of execution, with the constraint $16\le\kmem\le 32$. -\paragraph{Operands and addressing.} An instruction is an opcode followed by a list of $\K$-valued operands. An operand is either an \emph{fp-relative reference}, an offset $j$ into the current frame encoded as the $\gen$-power $o=\gen^{\,j}$, or one $\K$-lane of an \emph{immediate} ($k_0,k_1,k_2$ in \texttt{SET\_CONSTANT}). With the frame based at $\fp=\gen^{\mathrm{base}}$, a reference names the cell +\paragraph{Operands and addressing.} An instruction is an opcode followed by a list of $\K$-valued operands. An operand is either an \emph{fp-relative reference}, an offset $j$ into the current frame encoded as the $\gen$-power $o=\gen^{\,j}$, or the \emph{immediate} $k\in\K$ of \texttt{SET\_CONSTANT}. With the frame based at $\fp=\gen^{\mathrm{base}}$, a reference names the cell \[ \mem[\fp\cdot o],\qquad \fp\cdot o\;=\;\gen^{\mathrm{base}}\cdot\gen^{\,j}\;=\;\gen^{\,\mathrm{base}+j}. \] +A run of three consecutive cells spells one element of $\E$, its limbs low first: +\[ + \locE{o}\;=\;\loc{o}+\loc{\gen\cdot o}\,y+\loc{\gen^{2}\cdot o}\,y^{2}\;\in\;\E. +\] -\paragraph{Instruction set.} The machine implements six instructions. +\paragraph{Instruction set.} The machine implements eight instructions. \begin{center} \begin{tabularx}{\textwidth}{@{}llX@{}} \hline Instruction & Operands & Semantics\\ \hline -\texttt{XOR} & $[o_A,o_B,o_C]$ & $\loc{o_C}=\loc{o_A}+\loc{o_B}$ \quad(addition in $\E$)\\ -\texttt{MUL\_NATIVE} & $[o_A,o_B,o_C]$ & $\loc{o_C}=\loc{o_A}\cdot\loc{o_B}$ (multiplication in $\E$)\\ -\texttt{SET\_CONSTANT} & $[o,k_0,k_1,k_2]$ & $\loc{o}=k_0+k_1y+k_2y^{2}$\\ +\texttt{XOR64} & $[o_A,o_B,o_C]$ & $\loc{o_C}=\loc{o_A}+\loc{o_B}$ \quad(addition in $\K$)\\ +\texttt{MUL64} & $[o_A,o_B,o_C]$ & $\loc{o_C}=\loc{o_A}\cdot\loc{o_B}$ (multiplication in $\K$)\\ +\texttt{SET\_CONSTANT} & $[o,k]$ & $\loc{o}=k$\\ \texttt{DEREF} & $[o_1,o_2,o_3;\,\dmode]$ & $\mem[\loc{o_1}\cdot o_2]=\srcsel{\dmode}$ (see below)\\ \texttt{JUMP} & $[o_c,o_d,o_f]$ & conditional jump (see below)\\ \texttt{BLAKE2S} & $[o_{m_0},o_{m_1},o_{m_2},o_{m_3},o_{\mathit{cv}},o_{\mathit{out}},o_{\md}]$ & BLAKE2s compression (see below)\\ +\texttt{XOR192} & $[o_A,o_B,o_C]$ & $\locE{o_C}=\locE{o_A}+\locE{o_B}$ \quad(addition in $\E$)\\ +\texttt{MUL192} & $[o_A,o_B,o_C]$ & $\locE{o_C}=\locE{o_A}\cdot\locE{o_B}$ (multiplication in $\E$)\\ \hline \end{tabularx} \end{center} @@ -80,18 +86,16 @@ \section{VM specification}\label{sec:vm} \end{cases} \] -Additionally, it asserts $\loc{o_1} \in \K$. - \vspace{5mm} -\texttt{JUMP} reads a condition $c=\loc{o_c}$, and branches on whether it is zero. When $c\neq0$ it transfers control, $\pc\gets\loc{o_d}$ and $\fp\gets\loc{o_f}$; when $c=0$ it defaults to $\pc\gets\gen\cdot\pc$ and $\fp\gets\fp$. Additionally, it asserts $\loc{o_c},\loc{o_f}, \loc{o_d} \in \K$ (unconditionally). +\texttt{JUMP} reads a condition $c=\loc{o_c}$, and branches on whether it is zero. When $c\neq0$ it transfers control, $\pc\gets\loc{o_d}$ and $\fp\gets\loc{o_f}$; when $c=0$ it defaults to $\pc\gets\gen\cdot\pc$ and $\fp\gets\fp$. \vspace{5mm} -\texttt{BLAKE2S} is the standard BLAKE2s compression, consuming a $64$-byte message block and a $32$-byte chaining value and producing $32$ bytes. Each $128$-bit chunk occupies one full $\E$ memory cell through the canonical embedding $a_0+a_1y\mapsto a_0+a_1y+0y^2$; using a cell as a BLAKE2s operand therefore constrains its top limb to zero. The four message chunks are addressed \emph{independently} by references $o_{m_0},\dots,o_{m_3}$ (chunk $i$ at address $\fp\cdot o_{m_i}$), while $o_{\mathit{cv}}$ names two \emph{consecutive} chaining-value cells $\loc{o_{\mathit{cv}}},\loc{\gen\cdot o_{\mathit{cv}}}$. The $256$-bit result is written to the consecutive cells $\loc{o_{\mathit{out}}},\loc{\gen\cdot o_{\mathit{out}}}$, and $o_{\md}$ names one more cell, so each row accesses nine. That cell $\md=\loc{o_{\md}}=\md_0+\md_1y$ packs the standard compression metadata in little-endian order, +\texttt{BLAKE2S} is the standard BLAKE2s compression, consuming a $64$-byte message block and a $32$-byte chaining value and producing $32$ bytes. Memory holds these bytes as words: word $i$ of a byte string is its bytes $8i,\dots,8i+7$ read as a little-endian integer, so the block is eight words and the chaining value and the result are four each. The block splits into four $16$-byte message chunks, addressed \emph{independently} by references $o_{m_0},\dots,o_{m_3}$: chunk $i$ is the two consecutive cells $\loc{o_{m_i}},\loc{\gen\cdot o_{m_i}}$. The chaining value is read from the four \emph{consecutive} cells $\loc{\gen^{k}\cdot o_{\mathit{cv}}}$, $k<4$, and the result is written to the four consecutive cells $\loc{\gen^{k}\cdot o_{\mathit{out}}}$. Finally $o_{\md}$ names two more cells, so each row accesses eighteen. They hold the standard compression metadata, $\md_0=\loc{o_{\md}}$ and $\md_1=\loc{\gen\cdot o_{\md}}$, little-endian: \[ - \md=\mathit{counter}_{64}\;\Vert\;\mathit{final}_{32}\;\Vert\;\mathit{last\_node}_{32}. + \md_0=\mathit{counter}_{64},\qquad \md_1=\mathit{final}_{32}\;\Vert\;\mathit{last\_node}_{32}. \] diff --git a/doc/leanvm/body/06-bus-interactions.tex b/doc/leanvm/body/06-bus-interactions.tex index 4ba01488e..970ac2113 100644 --- a/doc/leanvm/body/06-bus-interactions.tex +++ b/doc/leanvm/body/06-bus-interactions.tex @@ -77,20 +77,22 @@ \subsection{The lookup protocol}\label{sec:lookup} \paragraph{The count product.} All counts are nonzero exactly when their product is. The prover therefore sends that product, one $\E$-element, and the verifier checks it is nonzero. It is proven like any other grand product, by GKR over the stacking of every count column (\S\ref{sec:leafstack}), batched with the two other GKRs of the bus (push / pull) (\S\ref{sec:gkr}). -\paragraph{Instance caps.} The verifier checks the announced sizes against public caps (\S\ref{sec:e2e-unrolled}): at most $2^{32}$ rows per table, at most $2^{32}$ memory cells (\S\ref{sec:memchan}) and at most $2^{32}$ bytecode instructions (\S\ref{sec:bytecode}). A row reads at most ten times (\texttt{BLAKE2S}: nine cells and its bytecode entry), so the six tables flush $\nread\le10\cdot6\cdot2^{32}<2^{38}$ reads, far under the $2^{64}-1$ that Lemma~\ref{lem:invcnt} requires. +\paragraph{Instance caps.} The verifier checks the announced sizes against public caps (\S\ref{sec:e2e-unrolled}): at most $2^{32}$ rows per table, at most $2^{32}$ memory cells (\S\ref{sec:memchan}) and at most $2^{32}$ bytecode instructions (\S\ref{sec:bytecode}). A row reads at most nineteen times (\texttt{BLAKE2S}: eighteen cells and its bytecode entry), so the eight tables flush $\nread\le19\cdot8\cdot2^{32}<2^{40}$ reads, far under the $2^{64}-1$ that Lemma~\ref{lem:invcnt} requires. \paragraph{Relaxing the hypotheses.} A bus entry may be any polynomial of degree at most $d = 2$ in the row's columns (\S\ref{sec:m3}), so none of the three (address, count, value) need be committed. Example 1: we don't need to commit memory addresses of the form $\fp\cdot o$ (\S\ref{sec:memchan}) when $\fp$ and $o$ are already committed. Example 2: once the pull counter $\cnt$ is committed, no need to commit to the corresponding push counter $\gen\cdot\cnt$. \subsection{Memory}\label{sec:memchan} -The memory is an array of $\nmem = 2^{\kmem}$ cells, $16\le\kmem\le32$ (\S\ref{sec:vm}). Bus interaction uses domain separation $\sep=\dsep{MEM} = \gen^{1}$. A memory word is $192$-bit (i.e. one $\E$-element), so $\ncomp=3$, memory is committed as 3 $\K$-valued multilinears in $\kmem$ variables $\mem_0,\mem_1,\mem_2$. +The memory is an array of $\nmem = 2^{\kmem}$ cells, $16\le\kmem\le32$ (\S\ref{sec:vm}). Bus interaction uses domain separation $\sep=\dsep{MEM} = \gen^{1}$. A memory word is $64$-bit (i.e. one $\K$-element), so $\ncomp=1$, and memory is committed as one $\K$-valued multilinear $\mem$ in $\kmem$ variables. \subsection{Bytecode}\label{sec:bytecode} The bytecode is an array of $\nprog = 2^{\kbc}$ instructions, $\kbc \le 32$ (\S\ref{sec:e2e-bc}). Bus interaction uses domain separation $\sep=\dsep{BC} = \gen^{2}$. An instruction is an opcode plus seven operands, so $\ncomp=8$: with the separator, the address and the count ahead of them, this is the widest entry on the bus at $11$ coordinates, which is what forces the four index bits, hence the $m=16$ slots of \S\ref{sec:m3}. Coordinate $3$ is the opcode, naming the instruction: -\[ - \opc{XOR}=\gen^{0},\quad \opc{MUL}=\gen^{1},\quad \opc{SET}=\gen^{2},\quad \opc{DRF}=\gen^{3},\quad \opc{JMP}=\gen^{4},\quad \opc{B2S}=\gen^{5}. -\] The program is public, i.e. not committed, the verifier can evaluate it directly. +\begin{gather*} + \opc{XOR64}=\gen^{0},\quad \opc{MUL64}=\gen^{1},\quad \opc{SET}=\gen^{2},\quad \opc{DRF}=\gen^{3},\\ + \opc{JMP}=\gen^{4},\quad \opc{B2S}=\gen^{5},\quad \opc{XOR192}=\gen^{6},\quad \opc{MUL192}=\gen^{7}. +\end{gather*} +The program is public, i.e. not committed, the verifier can evaluate it directly. \subsection{The index column}\label{sec:idxcol} diff --git a/doc/leanvm/body/07-instruction-tables.tex b/doc/leanvm/body/07-instruction-tables.tex index bcb5441f7..290d2d310 100644 --- a/doc/leanvm/body/07-instruction-tables.tex +++ b/doc/leanvm/body/07-instruction-tables.tex @@ -5,63 +5,53 @@ \section{The instruction tables}\label{sec:tables} Write $\rdf{\sep,\addr,\cnt,\val_0,\val_1, \dots}$ for the pair \pull\ $\tup{\sep,\addr,\cnt,\val_0,\val_1, \dots }$, \push\ $\tup{\sep,\addr,\gen\cdot\cnt,\val_0,\val_1, \dots}$: every flush below but the state's. -\subsection{\texttt{XOR}}\label{sec:tab-xor} +\subsection{\texttt{XOR64}}\label{sec:tab-xor} -\texttt{XOR}\ $[o_A,o_B,o_C]$ asserts $\loc{o_C}=\loc{o_A}+\loc{o_B}$, the $192$-bit field sum (in $\E$). +\texttt{XOR64}\ $[o_A,o_B,o_C]$ asserts $\loc{o_C}=\loc{o_A}+\loc{o_B}$, the $64$-bit field sum (in $\K$). \begin{itemize} -\item \textbf{Columns:} $\pc,\fp$; operands $o_A,o_B,o_C$; the two read values $(v_{A,i})_{i\in\{0,1,2\}}$ and $(v_{B,i})_{i\in\{0,1,2\}}$; memory counts $r_A,r_B,r_C$; bytecode count $r_{\mathrm{bc}}$. +\item \textbf{Columns:} $\pc,\fp$; operands $o_A,o_B,o_C$; the two read values $v_A,v_B$; memory counts $r_A,r_B,r_C$; bytecode count $r_{\mathrm{bc}}$. \item \textbf{Constraints:} none. \item \textbf{Flushes:} \begin{itemize} \item \textbf{State:} \pull\ $\tup{\dsep{ST},\pc,\fp}$, \push\ $\tup{\dsep{ST},\gen\cdot\pc,\fp}$. - \item \textbf{Bytecode:} $\rdf{\dsep{BC},\pc,r_{\mathrm{bc}},\opc{XOR},o_A,o_B,o_C,0,0}$. + \item \textbf{Bytecode:} $\rdf{\dsep{BC},\pc,r_{\mathrm{bc}},\opc{XOR64},o_A,o_B,o_C,0,0}$. \item \textbf{Memory:} \begin{align*} - &\rdf{\dsep{MEM},\fp\cdot o_X,r_X,v_{X,0},v_{X,1},v_{X,2}}\quad(X\in\{A,B\}),\\ - &\rdf{\dsep{MEM},\fp\cdot o_C,r_C,v_{A,0}+v_{B,0},v_{A,1}+v_{B,1},v_{A,2}+v_{B,2}}. + &\rdf{\dsep{MEM},\fp\cdot o_X,r_X,v_X}\quad(X\in\{A,B\}),\\ + &\rdf{\dsep{MEM},\fp\cdot o_C,r_C,v_A+v_B}. \end{align*} \end{itemize} \end{itemize} -\subsection{\texttt{MUL\_NATIVE}}\label{sec:tab-mul} +\subsection{\texttt{MUL64}}\label{sec:tab-mul} -\texttt{MUL\_NATIVE}\ $[o_A,o_B,o_C]$ asserts $\loc{o_C}=\loc{o_A} \cdot \loc{o_B}$, the $192$-bit field multiplication (in $\E$). +\texttt{MUL64}\ $[o_A,o_B,o_C]$ asserts $\loc{o_C}=\loc{o_A} \cdot \loc{o_B}$, the $64$-bit field multiplication (in $\K$). \begin{itemize} -\item \textbf{Columns:} $\pc,\fp$; operands $o_A,o_B,o_C$; the two read values $(v_{A,i})_{i\in\{0,1,2\}}$ and $(v_{B,i})_{i\in\{0,1,2\}}$; memory counts $r_A,r_B,r_C$; bytecode count $r_{\mathrm{bc}}$. +\item \textbf{Columns:} $\pc,\fp$; operands $o_A,o_B,o_C$; the two read values $v_A,v_B$; memory counts $r_A,r_B,r_C$; bytecode count $r_{\mathrm{bc}}$. \item \textbf{Constraints:} none. \item \textbf{Flushes:} \begin{itemize} \item \textbf{State:} \pull\ $\tup{\dsep{ST},\pc,\fp}$, \push\ $\tup{\dsep{ST},\gen\cdot\pc,\fp}$. - \item \textbf{Bytecode:} $\rdf{\dsep{BC},\pc,r_{\mathrm{bc}},\opc{MUL},o_A,o_B,o_C,0,0}$. - \item \textbf{Memory} (the destination carries $v_Av_B$ in $\E$, of degree two). With $y^3=y+1$, put - \[ - p_0=v_{A,0}v_{B,0},\quad - p_1=v_{A,0}v_{B,1}+v_{A,1}v_{B,0},\quad - p_2=v_{A,0}v_{B,2}+v_{A,1}v_{B,1}+v_{A,2}v_{B,0}, - \] - \[ - p_3=v_{A,1}v_{B,2}+v_{A,2}v_{B,1},\qquad - p_4=v_{A,2}v_{B,2}, - \] - twelve products over the nine pairs $v_{A,i}v_{B,j}$: + \item \textbf{Bytecode:} $\rdf{\dsep{BC},\pc,r_{\mathrm{bc}},\opc{MUL64},o_A,o_B,o_C,0,0}$. + \item \textbf{Memory} (the destination carries $v_Av_B$, of degree two): \begin{align*} - &\rdf{\dsep{MEM},\fp\cdot o_X,r_X,v_{X,0},v_{X,1},v_{X,2}}\quad(X\in\{A,B\}),\\ - &\rdf{\dsep{MEM},\fp\cdot o_C,r_C,p_0+p_3,p_1+p_3+p_4,p_2+p_4}. + &\rdf{\dsep{MEM},\fp\cdot o_X,r_X,v_X}\quad(X\in\{A,B\}),\\ + &\rdf{\dsep{MEM},\fp\cdot o_C,r_C,v_Av_B}. \end{align*} \end{itemize} \end{itemize} \subsection{\texttt{SET\_CONSTANT}}\label{sec:tab-set} -\texttt{SET\_CONSTANT}\ $[o,k_0,k_1,k_2]$ asserts $\loc{o}=k_0+k_1y+k_2y^{2}$. +\texttt{SET\_CONSTANT}\ $[o,k]$ asserts $\loc{o}=k$. \begin{itemize} -\item \textbf{Columns:} $\pc,\fp$; operand $o$; immediate $(k_i)_{i\in\{0,1,2\}}$; memory count $r$; bytecode count $r_{\mathrm{bc}}$. +\item \textbf{Columns:} $\pc,\fp$; operand $o$; immediate $k$; memory count $r$; bytecode count $r_{\mathrm{bc}}$. \item \textbf{Constraints:} none. \item \textbf{Flushes:} \begin{itemize} \item \textbf{State:} \pull\ $\tup{\dsep{ST},\pc,\fp}$, \push\ $\tup{\dsep{ST},\gen\cdot\pc,\fp}$. - \item \textbf{Bytecode:} $\rdf{\dsep{BC},\pc,r_{\mathrm{bc}},\opc{SET},o,k_0,k_1,k_2,0}$. - \item \textbf{Memory:} $\rdf{\dsep{MEM},\fp\cdot o,r,k_0,k_1,k_2}$. + \item \textbf{Bytecode:} $\rdf{\dsep{BC},\pc,r_{\mathrm{bc}},\opc{SET},o,k,0,0,0}$. + \item \textbf{Memory:} $\rdf{\dsep{MEM},\fp\cdot o,r,k}$. \end{itemize} \end{itemize} @@ -75,7 +65,7 @@ \subsection{\texttt{DEREF}}\label{sec:tab-deref} \end{itemize} The store asserts that $v_2$ equals the source chosen by the mode $\dmode$ (\S\ref{sec:vm}): the local cell $v_3$, the return address $\gen^{2}\cdot\pc$, or the frame pointer $\fp$. The mode is two boolean flags $(f_{\pc},f_{\fp})$: $(0,0)$ for $\texttt{deref\_cell}[o_3]$, $(1,0)$ for $\texttt{deref\_pc}$, $(0,1)$ for $\texttt{deref\_fp}$. \begin{itemize} -\item \textbf{Columns:} $\pc,\fp$; operands $o_1,o_2,o_3$; flags $f_{\pc},f_{\fp}$; pointer $p$ (a single $\K$-limb); the local value $(v_{3,i})_{i\in\{0,1,2\}}$; memory counts $r_1,r_2,r_3$; bytecode count $r_{\mathrm{bc}}$. +\item \textbf{Columns:} $\pc,\fp$; operands $o_1,o_2,o_3$; flags $f_{\pc},f_{\fp}$; pointer $p$; the local value $v_3$; memory counts $r_1,r_2,r_3$; bytecode count $r_{\mathrm{bc}}$. \item \textbf{Constraints:} none. \item \textbf{Flushes:} \begin{itemize} @@ -83,9 +73,9 @@ \subsection{\texttt{DEREF}}\label{sec:tab-deref} \item \textbf{Bytecode:} $\rdf{\dsep{BC},\pc,r_{\mathrm{bc}},\opc{DRF},o_1,o_2,o_3,f_{\pc},f_{\fp}}$. \item \textbf{Memory}, the target giving $v_3$, $\gen^{2}\cdot\pc$ or $\fp$ at the three flag settings: \begin{align*} - &\rdf{\dsep{MEM},\fp\cdot o_1,r_1,p,0,0},\\ - &\rdf{\dsep{MEM},\fp\cdot o_3,r_3,v_{3,0},v_{3,1},v_{3,2}},\\ - &\rdf{\dsep{MEM},p\cdot o_2,r_2,\bar f\,v_{3,0}+f_{\pc}\,(\gen^{2}\cdot\pc)+f_{\fp}\,\fp,\ \bar f\,v_{3,1},\ \bar f\,v_{3,2}}, + &\rdf{\dsep{MEM},\fp\cdot o_1,r_1,p},\\ + &\rdf{\dsep{MEM},\fp\cdot o_3,r_3,v_3},\\ + &\rdf{\dsep{MEM},p\cdot o_2,r_2,\bar f\,v_3+f_{\pc}\,(\gen^{2}\cdot\pc)+f_{\fp}\,\fp}, \qquad \bar f=1+f_{\pc}+f_{\fp}. \end{align*} \end{itemize} @@ -93,7 +83,7 @@ \subsection{\texttt{DEREF}}\label{sec:tab-deref} \subsection{\texttt{JUMP}}\label{sec:tab-jump} -\texttt{JUMP}\ $[o_c,o_d,o_f]$ performs a conditional jump on whether the condition $v_{\mathrm{cond}}=\loc{o_c}$ is nonzero. It transfers to $\pc\gets v_{\pc} = \loc{o_d}$, $\fp\gets v_{\fp} = \loc{o_f}$ when $v_{\mathrm{cond}}\neq0$, and defaults to $\pc\gets\gen\cdot\pc$, $\fp\gets\fp$ otherwise. $v_{\mathrm{cond}}, v_{\fp}, v_{\pc}$ must be $\K$-valued. The branch is taken according to a boolean indicator column $b=[\,v_{\mathrm{cond}}\neq0\,]$, certified by a prover-supplied inverse $w$ and 2 constraints: +\texttt{JUMP}\ $[o_c,o_d,o_f]$ performs a conditional jump on whether the condition $v_{\mathrm{cond}}=\loc{o_c}$ is nonzero. It transfers to $\pc\gets v_{\pc} = \loc{o_d}$, $\fp\gets v_{\fp} = \loc{o_f}$ when $v_{\mathrm{cond}}\neq0$, and defaults to $\pc\gets\gen\cdot\pc$, $\fp\gets\fp$ otherwise. The branch is taken according to a boolean indicator column $b=[\,v_{\mathrm{cond}}\neq0\,]$, certified by a prover-supplied inverse $w$ and 2 constraints: \begin{itemize} \item \textbf{Columns:} $\pc,\fp$; operands $o_c,o_d,o_f$; the condition $v_{\mathrm{cond}}$, the destination $v_{\pc}$ and the frame $v_{\fp}$; inverse $w$ (honest prover sets $w = v_{\mathrm{cond}}^{-1}$ if $v_{\mathrm{cond}}\neq0$, otherwise $w=0$); boolean indicator $b$; memory counts $r_c,r_d,r_f$; bytecode count $r_{\mathrm{bc}}$. @@ -104,22 +94,22 @@ \subsection{\texttt{JUMP}}\label{sec:tab-jump} \item \textbf{Bytecode:} $\rdf{\dsep{BC},\pc,r_{\mathrm{bc}},\opc{JMP},o_c,o_d,o_f,0,0}$. \item \textbf{Memory:} \begin{align*} - &\rdf{\dsep{MEM},\fp\cdot o_c,r_c,v_{\mathrm{cond}},0,0},\\ - &\rdf{\dsep{MEM},\fp\cdot o_d,r_d,v_{\pc},0,0},\\ - &\rdf{\dsep{MEM},\fp\cdot o_f,r_f,v_{\fp},0,0}. + &\rdf{\dsep{MEM},\fp\cdot o_c,r_c,v_{\mathrm{cond}}},\\ + &\rdf{\dsep{MEM},\fp\cdot o_d,r_d,v_{\pc}},\\ + &\rdf{\dsep{MEM},\fp\cdot o_f,r_f,v_{\fp}}. \end{align*} \end{itemize} \end{itemize} \subsection{\texttt{BLAKE2S}}\label{sec:tab-blake2s} -\texttt{BLAKE2S}\ $[o_{m_0},o_{m_1},o_{m_2},o_{m_3},o_{\mathit{cv}},o_{\mathit{out}},o_{\md}]$ compresses a $64$-byte message block under a $32$-byte chaining value into a $32$-byte result, with the BLAKE2s flags and byte counter in the metadata cell $\fp\cdot o_{\md}$. Each $128$-bit half occupies one cell, top limb zero: the four message chunks at $\fp\cdot o_{m_0},\dots,\fp\cdot o_{m_3}$, addressed independently, the chaining value at the consecutive pair $\fp\cdot o_{\mathit{cv}},\gen\cdot\fp\cdot o_{\mathit{cv}}$, and the result at $\fp\cdot o_{\mathit{out}},\gen\cdot\fp\cdot o_{\mathit{out}}$. +\texttt{BLAKE2S}\ $[o_{m_0},o_{m_1},o_{m_2},o_{m_3},o_{\mathit{cv}},o_{\mathit{out}},o_{\md}]$ compresses a $64$-byte message block under a $32$-byte chaining value into a $32$-byte result, with the BLAKE2s byte counter and flags in the two metadata cells from $\fp\cdot o_{\md}$ (\S\ref{sec:vm}). Every cell holds one $8$-byte word: message chunk $i$ spans the cells $\gen^{k}\cdot\fp\cdot o_{m_i}$, $k<2$, the four chunks addressed independently; the chaining value spans $\gen^{k}\cdot\fp\cdot o_{\mathit{cv}}$ and the result $\gen^{k}\cdot\fp\cdot o_{\mathit{out}}$, $k<4$; the metadata spans $\gen^{k}\cdot\fp\cdot o_{\md}$, $k<2$. Flock~\cite{Flock26} proves the compression relation with a batched R1CS over a Boolean witness (Annex~\ref{annex:flock}). The witness is packed into $\qflock$ via Ring-Switching (Annex~\ref{annex:rs}), with $64$ bits per $\K$-element. -The message, result, chaining value and metadata are all committed in $\qflock$: the two low limbs $z_0,z_1$ of each of the nine cells $z\in\{m_0,\dots,m_3,\mathit{cv}_0,\mathit{cv}_1,\mathit{out}_0,\mathit{out}_1,\md\}$, eighteen in all. Linking these values to the memory interaction implicitly forces the top limb of every memory cell accessed by BLAKE2s to be zero. The generalized R1CS matrices take the chaining value, byte counter, and the two finalization flags as free input: the memory interactions via the bus bind all four. +The message, result, chaining value and metadata are all committed in $\qflock$: the eighteen words $m_{i,k}$ ($i<4$, $k<2$), $\mathit{out}_k$ and $\mathit{cv}_k$ ($k<4$), and $\md_k$ ($k<2$), one $\qflock$ slot each. The generalized R1CS matrices take the chaining value, byte counter, and the two finalization flags as free input: the memory interactions via the bus bind all four. \begin{itemize} -\item \textbf{Columns:} $\pc,\fp$; operands $o_{m_0},o_{m_1},o_{m_2},o_{m_3},o_{\mathit{cv}},o_{\mathit{out}},o_{\md}$; one memory count $r_z$ per cell $z$; bytecode count $r_{\mathrm{bc}}$. The eighteen limbs are committed in $\qflock$, no duplicate columns needed. The nine addresses are the products $\gen^{k}\fp\cdot o$, $k\in\{0,1\}$. +\item \textbf{Columns:} $\pc,\fp$; operands $o_{m_0},o_{m_1},o_{m_2},o_{m_3},o_{\mathit{cv}},o_{\mathit{out}},o_{\md}$; one memory count per word read, $r_{m_i,k}$, $r_{\mathit{out},k}$, $r_{\mathit{cv},k}$, $r_{\md,k}$; bytecode count $r_{\mathrm{bc}}$. The eighteen words are committed in $\qflock$, no duplicate columns needed. The eighteen addresses are the products $\gen^{k}\cdot\fp\cdot o$. \item \textbf{Constraints:} none. The compression relation is handled by Flock. \item \textbf{Flushes:} \begin{itemize} @@ -127,12 +117,58 @@ \subsection{\texttt{BLAKE2S}}\label{sec:tab-blake2s} \item \textbf{Bytecode:} $\rdf{\dsep{BC},\pc,r_{\mathrm{bc}},\opc{B2S},o_{m_0},o_{m_1},o_{m_2},o_{m_3},o_{\mathit{cv}},o_{\mathit{out}},o_{\md}}$. \item \textbf{Memory:} \begin{align*} - &\rdf{\dsep{MEM},\fp\cdot o_{m_i},r_{m_i},m_{i,0},m_{i,1},0}\quad(i=0,\dots,3),\\ - &\rdf{\dsep{MEM},\fp\cdot o_{\mathit{cv}},r_{\mathit{cv}_0},\mathit{cv}_{0,0},\mathit{cv}_{0,1},0},\\ - &\rdf{\dsep{MEM},\gen\cdot\fp\cdot o_{\mathit{cv}},r_{\mathit{cv}_1},\mathit{cv}_{1,0},\mathit{cv}_{1,1},0},\\ - &\rdf{\dsep{MEM},\fp\cdot o_{\mathit{out}},r_{\mathit{out}_0},\mathit{out}_{0,0},\mathit{out}_{0,1},0},\\ - &\rdf{\dsep{MEM},\gen\cdot\fp\cdot o_{\mathit{out}},r_{\mathit{out}_1},\mathit{out}_{1,0},\mathit{out}_{1,1},0},\\ - &\rdf{\dsep{MEM},\fp\cdot o_{\md},r_{\md},\md_0,\md_1,0}. + &\rdf{\dsep{MEM},\gen^{k}\cdot\fp\cdot o_{m_i},r_{m_i,k},m_{i,k}}\quad(i<4,\ k<2),\\ + &\rdf{\dsep{MEM},\gen^{k}\cdot\fp\cdot o_{\mathit{out}},r_{\mathit{out},k},\mathit{out}_{k}}\quad(k<4),\\ + &\rdf{\dsep{MEM},\gen^{k}\cdot\fp\cdot o_{\mathit{cv}},r_{\mathit{cv},k},\mathit{cv}_{k}}\quad(k<4),\\ + &\rdf{\dsep{MEM},\gen^{k}\cdot\fp\cdot o_{\md},r_{\md,k},\md_{k}}\quad(k<2). + \end{align*} + \end{itemize} +\end{itemize} + +\subsection{\texttt{XOR192}}\label{sec:tab-xor192} + +\texttt{XOR192}\ $[o_A,o_B,o_C]$ asserts $\locE{o_C}=\locE{o_A}+\locE{o_B}$, the $192$-bit field sum (in $\E$). Each operand is three cells, read one limb at a time: limb $i$ of operand $X$ sits at $\gen^{i}\cdot\fp\cdot o_X$. +\begin{itemize} +\item \textbf{Columns:} $\pc,\fp$; operands $o_A,o_B,o_C$; the two read values $(v_{A,i})_{i\in\{0,1,2\}}$ and $(v_{B,i})_{i\in\{0,1,2\}}$; memory counts $(r_{X,i})_{i\in\{0,1,2\}}$ for $X\in\{A,B,C\}$; bytecode count $r_{\mathrm{bc}}$. +\item \textbf{Constraints:} none. +\item \textbf{Flushes:} + \begin{itemize} + \item \textbf{State:} \pull\ $\tup{\dsep{ST},\pc,\fp}$, \push\ $\tup{\dsep{ST},\gen\cdot\pc,\fp}$. + \item \textbf{Bytecode:} $\rdf{\dsep{BC},\pc,r_{\mathrm{bc}},\opc{XOR192},o_A,o_B,o_C,0,0}$. + \item \textbf{Memory}, for $i\in\{0,1,2\}$: + \begin{align*} + &\rdf{\dsep{MEM},\gen^{i}\cdot\fp\cdot o_X,r_{X,i},v_{X,i}}\quad(X\in\{A,B\}),\\ + &\rdf{\dsep{MEM},\gen^{i}\cdot\fp\cdot o_C,r_{C,i},v_{A,i}+v_{B,i}}. + \end{align*} + \end{itemize} +\end{itemize} + +\subsection{\texttt{MUL192}}\label{sec:tab-mul192} + +\texttt{MUL192}\ $[o_A,o_B,o_C]$ asserts $\locE{o_C}=\locE{o_A}\cdot\locE{o_B}$, the $192$-bit field multiplication (in $\E$), over the same three-cell operands as \texttt{XOR192}. +\begin{itemize} +\item \textbf{Columns:} $\pc,\fp$; operands $o_A,o_B,o_C$; the two read values $(v_{A,i})_{i\in\{0,1,2\}}$ and $(v_{B,i})_{i\in\{0,1,2\}}$; memory counts $(r_{X,i})_{i\in\{0,1,2\}}$ for $X\in\{A,B,C\}$; bytecode count $r_{\mathrm{bc}}$. +\item \textbf{Constraints:} none. +\item \textbf{Flushes:} + \begin{itemize} + \item \textbf{State:} \pull\ $\tup{\dsep{ST},\pc,\fp}$, \push\ $\tup{\dsep{ST},\gen\cdot\pc,\fp}$. + \item \textbf{Bytecode:} $\rdf{\dsep{BC},\pc,r_{\mathrm{bc}},\opc{MUL192},o_A,o_B,o_C,0,0}$. + \item \textbf{Memory} (the destination carries $v_Av_B$ in $\E$, each limb of degree two). With $y^3=y+1$, put + \[ + p_0=v_{A,0}v_{B,0},\quad + p_1=v_{A,0}v_{B,1}+v_{A,1}v_{B,0},\quad + p_2=v_{A,0}v_{B,2}+v_{A,1}v_{B,1}+v_{A,2}v_{B,0}, + \] + \[ + p_3=v_{A,1}v_{B,2}+v_{A,2}v_{B,1},\qquad + p_4=v_{A,2}v_{B,2}, + \] + twelve products over the nine pairs $v_{A,i}v_{B,j}$: + \begin{align*} + &\rdf{\dsep{MEM},\gen^{i}\cdot\fp\cdot o_X,r_{X,i},v_{X,i}}\quad(X\in\{A,B\},\ i\in\{0,1,2\}),\\ + &\rdf{\dsep{MEM},\fp\cdot o_C,r_{C,0},p_0+p_3},\\ + &\rdf{\dsep{MEM},\gen\cdot\fp\cdot o_C,r_{C,1},p_1+p_3+p_4},\\ + &\rdf{\dsep{MEM},\gen^{2}\cdot\fp\cdot o_C,r_{C,2},p_2+p_4}. \end{align*} \end{itemize} \end{itemize} diff --git a/doc/leanvm/body/08-end-to-end-protocol.tex b/doc/leanvm/body/08-end-to-end-protocol.tex index 682c7678e..b310686f4 100644 --- a/doc/leanvm/body/08-end-to-end-protocol.tex +++ b/doc/leanvm/body/08-end-to-end-protocol.tex @@ -5,7 +5,7 @@ \subsection{Bytecode encoding}\label{sec:e2e-bc} The bytecode (of $\nprog=2^{\kbc}$ instructions) is encoded as one $\K$-valued multilinear $P$ in $\kbc+4$ variables: \begin{enumerate}[itemsep=1pt,topsep=3pt] -\item lay instruction $z$ out as $16$ slots of $\K$: the opcode in slot $3$, the operands and immediate lanes in slots $4,\dots,10$, and $0$ everywhere else. +\item lay instruction $z$ out as $16$ slots of $\K$: the opcode in slot $3$, the operands and the immediate in slots $4,\dots,10$, and $0$ everywhere else. \item read the resulting $2^{\kbc}\times16$ table as $P$, the low $\kbc$ variables indexing the instruction and the high four its slot, bits low first. \end{enumerate} \begin{center} @@ -13,12 +13,14 @@ \subsection{Bytecode encoding}\label{sec:e2e-bc} \hline Instruction & $3$ & $4$ & $5$ & $6$ & $7$ & $8$ & $9$ & $10$\\ \hline -\texttt{XOR} & $\opc{XOR}$ & $o_A$ & $o_B$ & $o_C$ & $0$ & $0$ & $0$ & $0$\\ -\texttt{MUL\_NATIVE} & $\opc{MUL}$ & $o_A$ & $o_B$ & $o_C$ & $0$ & $0$ & $0$ & $0$\\ -\texttt{SET\_CONSTANT} & $\opc{SET}$ & $o$ & $k_0$ & $k_1$ & $k_2$ & $0$ & $0$ & $0$\\ +\texttt{XOR64} & $\opc{XOR64}$ & $o_A$ & $o_B$ & $o_C$ & $0$ & $0$ & $0$ & $0$\\ +\texttt{MUL64} & $\opc{MUL64}$ & $o_A$ & $o_B$ & $o_C$ & $0$ & $0$ & $0$ & $0$\\ +\texttt{SET\_CONSTANT} & $\opc{SET}$ & $o$ & $k$ & $0$ & $0$ & $0$ & $0$ & $0$\\ \texttt{DEREF} & $\opc{DRF}$ & $o_1$ & $o_2$ & $o_3$ & $f_{\pc}$ & $f_{\fp}$ & $0$ & $0$\\ \texttt{JUMP} & $\opc{JMP}$ & $o_c$ & $o_d$ & $o_f$ & $0$ & $0$ & $0$ & $0$\\ \texttt{BLAKE2S} & $\opc{B2S}$ & $o_{m_0}$ & $o_{m_1}$ & $o_{m_2}$ & $o_{m_3}$ & $o_{\mathit{cv}}$ & $o_{\mathit{out}}$ & $o_{\md}$\\ +\texttt{XOR192} & $\opc{XOR192}$ & $o_A$ & $o_B$ & $o_C$ & $0$ & $0$ & $0$ & $0$\\ +\texttt{MUL192} & $\opc{MUL192}$ & $o_A$ & $o_B$ & $o_C$ & $0$ & $0$ & $0$ & $0$\\ \hline \end{tabular} \end{center} @@ -26,11 +28,11 @@ \subsection{Bytecode encoding}\label{sec:e2e-bc} \subsection{Public input}\label{sec:e2e-pi} -The public input is the first two memory cells $\mem[\gen^{0}],\mem[\gen^{1}]$, two $192$-bit words (with top limb $0$) both parties know. The verifier samples $r_m\in\E$, the prover sends $c_0,c_1\in\E$ claiming $c_\ell=\mle{\mem_\ell}(r_m,0,\dots,0)$, and the verifier checks +The public input is the first four memory cells $\mem[\gen^{0}],\dots,\mem[\gen^{3}]$, four words both parties know. The verifier samples $r=(r_0,r_1)\in\E^{2}$ and, with nothing sent, pools the claim \[ - c_0+y c_1\;\qeq\;(1+r_m)\,\mem[\gen^{0}]+r_m\,\mem[\gen^{1}]. + \mle{\mem}(r_0,r_1,0,\dots,0)\;\qeq\;\sum_{b\in\cube{2}}\eq(r,b)\,\mem[\gen^{\,b_0+2b_1}], \] -The PCS binds $(c_0,c_1,0)$ to the three memory limbs at $(r_m,0,\dots,0)$. A mismatch gives a nonzero polynomial of degree at most one, so it passes with probability at most $1/|\E|$ (Lemma~\ref{lem:sz}). +whose right side, the multilinear extension of the four public words at $r$ ($r_0$ on the low bit), the verifier computes itself. By Fact~\ref{fact:eq} the left side is the same sum over the four committed cells, so a mismatch leaves a nonzero multilinear polynomial in $r$, of total degree at most two, and the claim holds with probability at most $2/|\E|$ (Lemma~\ref{lem:sz}). The PCS opening binds the claim to the committed memory. \subsection{Filling the tables}\label{sec:e2e-pad} @@ -52,12 +54,12 @@ \subsection{The unrolled protocol}\label{sec:e2e-unrolled} \paragraph{Setup.} -The Fiat-Shamir initial state (\S\ref{sec:e2e-fiat-shamir}) is seeded by the public input, the bytecode and Flock's BLAKE2s R1CS matrices. The prover sends the memory log-size $\kmem$ and the six table log-heights $\tau_j$; the verifier checks them against the caps (\S\ref{sec:lookup}) and derives the layout of every stack (PCS and GKR leaves). +The Fiat-Shamir initial state (\S\ref{sec:e2e-fiat-shamir}) is seeded by the public input, the bytecode and Flock's BLAKE2s R1CS matrices. The prover sends the memory log-size $\kmem$ and the eight table log-heights $\tau_j$; the verifier checks them against the caps (\S\ref{sec:lookup}) and derives the layout of every stack (PCS and GKR leaves). \paragraph{Commitment.} \begin{enumerate} -\item \textbf{Prover.} Stack (\S\ref{sec:stacking}) every column, the three memory limbs, the two finalize (memory / bytecode) counts, and Flock's packed witness $\qflock$ (\S\ref{sec:tab-blake2s}), into $q$ and sends its commitment (Annex~\ref{annex:pcs}). +\item \textbf{Prover.} Stack (\S\ref{sec:stacking}) every column, the memory $\mem$, the two finalize (memory / bytecode) counts, and Flock's packed witness $\qflock$ (\S\ref{sec:tab-blake2s}), into $q$ and sends its commitment (Annex~\ref{annex:pcs}). \item \textbf{Verifier.} Receives the commitment to $q$. \end{enumerate} @@ -70,18 +72,18 @@ \subsection{The unrolled protocol}\label{sec:e2e-unrolled} \item \textbf{Prover \& Verifier.} Decompose each leaf claim (\S\ref{sec:leafstack}). The verifier forms every public part itself: the block selectors, the constant entries, the index column (\S\ref{sec:idxcol}), and the program's whole share of the two bytecode blocks, one evaluation at $(\zeta,\alpha)$ (\S\ref{sec:e2e-bc}). Three blocks per side belong to no table: the state boundary, plus each array's (memory / bytecode) seed (push side) or finalization (pull side). For their committed columns the prover sends one evaluation each, entering the claim pool. Every other block belongs to a table, and is discharged by the table sumcheck (\S\ref{sec:air}). \end{enumerate} -\paragraph{Table sumcheck.} One run proves the six tables' constraints together with their share of the bus (\S\ref{sec:air}). +\paragraph{Table sumcheck.} One run proves the eight tables' constraints together with their share of the bus (\S\ref{sec:air}). \begin{enumerate} \item \textbf{Verifier.} Samples and sends the batch challenge $\xi\in\E$. Its first $\nconstot=\sum_j\ncons{j}$ powers weight the constraints, a disjoint range per table; the last $\nside$ weight the bus forms, one per side and shared by every table. It draws no point: the batch runs at the bus point, $T_j$ tested at $\zeta_{<\tau_j}$. \item \textbf{Prover \& Verifier.} Run the sumcheck, $\taumax=\max_j\tau_j$ rounds with $T_j$ joining at round $\taumax-\tau_j$, against the target the verifier derived from the three leaf claims. Its summand has individual degree three and the round check pins one of the four coefficients (\S\ref{sec:prelim}), so a round costs three $\E$-elements. -\item \textbf{Prover \& Verifier.} The prover sends one value per column of every table. The verifier rebuilds each $C_{j,i}$ and $B^{s}_j$ from them and checks that they reproduce the sumcheck's final value; the column values then enter the claim pool, $T_j$'s at $\fc_{<\tau_j}$. \texttt{BLAKE2S}'s eighteen value limbs are columns of its table but live in $\qflock$, not the stack, so their claims route there (\S\ref{sec:tab-blake2s}). +\item \textbf{Prover \& Verifier.} The prover sends one value per column of every table. The verifier rebuilds each $C_{j,i}$ and $B^{s}_j$ from them and checks that they reproduce the sumcheck's final value; the column values then enter the claim pool, $T_j$'s at $\fc_{<\tau_j}$. \texttt{BLAKE2S}'s eighteen value words are columns of its table but live in $\qflock$, not the stack, so their claims route there (\S\ref{sec:tab-blake2s}). \end{enumerate} \paragraph{Public input.} \begin{enumerate} -\item \textbf{Verifier.} Sends $r_m\in\E$ and forms the line $(1+r_m)\mem[\gen^0]+r_m\mem[\gen^1]$ from the public words. -\item \textbf{Prover \& Verifier.} The prover sends $c_0,c_1$; the verifier checks $c_0+y c_1$ against the line and pools $(c_0,c_1,0)$ as claims on $\mem_0,\mem_1,\mem_2$ at $(r_m,0,\dots,0)$ (\S\ref{sec:e2e-pi}). +\item \textbf{Verifier.} Sends $r=(r_0,r_1)\in\E^{2}$ and evaluates the multilinear extension of the four public words at $r$. +\item \textbf{Prover \& Verifier.} Pool that value as a claim on $\mem$ at $(r_0,r_1,0,\dots,0)$; the prover sends nothing (\S\ref{sec:e2e-pi}). \end{enumerate} \paragraph{BLAKE2s validity.} @@ -97,5 +99,5 @@ \subsection{The unrolled protocol}\label{sec:e2e-unrolled} \item \textbf{Prover \& Verifier.} Each pooled claim is one evaluation of a column at a point, its value already supplied; via the stacking selectors (\S\ref{sec:stacking}) it becomes a weighted sum $\sum_w W_j(w)\,\widetilde q(w)=c_j$ over the stacked witness. (The ring-switched claim arrives already in this form, its weight supported on the $\qflock$ region.) \item \textbf{Verifier.} Sends one batching challenge $\lambda\in\E$, drawn after every claim value is bound. Claim $j$ takes the weight $\lambda^{j-1}$, the ring-switched claim first and the pooled column claims after: the $J$ weights are distinct powers of the one challenge (Theorem~\ref{thm:rbr}). \item \textbf{Prover \& Verifier.} Fold the $J$ claims into $W_\lambda=\sum_j\lambda^{j-1} W_j$ and $C_\lambda=\sum_j\lambda^{j-1} c_j$, and run the inner-product PCS opening on $\sum_w W_\lambda(w)\,\widetilde q(w)=C_\lambda$ (Annex~\ref{annex:pcs}), the verifier evaluating $W_\lambda$ itself. This one opening discharges every pooled claim; there is no separate reduction sumcheck. -\item \textbf{Verifier.} Accepts if every check above passed: the caps, the one bus root for both sides, the nonzero count root, every sumcheck's rounds and final value, the public-input line, flock's reduction, and the PCS opening. +\item \textbf{Verifier.} Accepts if every check above passed: the caps, the one bus root for both sides, the nonzero count root, every sumcheck's rounds and final value, flock's reduction, and the PCS opening. \end{enumerate} diff --git a/doc/leanvm/body/09-isa-programming.tex b/doc/leanvm/body/09-isa-programming.tex index 2f3ab5b68..64669996e 100644 --- a/doc/leanvm/body/09-isa-programming.tex +++ b/doc/leanvm/body/09-isa-programming.tex @@ -1,7 +1,7 @@ % !TeX root = ../drafts/09-isa-programming.tex \section{ISA programming}\label{sec:isa-programming} -How the zkDSL compiler (\path{crates/lean_compiler}; surface language in \path{crates/lean_compiler/zkDSL.md}) programs the six-instruction ISA. Two themes run through every pattern: \emph{write-once memory is an assertion mechanism} (a second write of a cell is an equality check, an unwritten cell is prover-chosen), and \emph{indices live in the exponent} (a logical index $i$ is the element $\gen^{i}$, so incrementing, offsetting, and address formation are single field multiplications). +How the zkDSL compiler (\path{crates/lean_compiler}; surface language in \path{crates/lean_compiler/zkDSL.md}) programs the eight-instruction ISA. Two themes run through every pattern: \emph{write-once memory is an assertion mechanism} (a second write of a cell is an equality check, an unwritten cell is prover-chosen), and \emph{indices live in the exponent} (a logical index $i$ is the element $\gen^{i}$, so incrementing, offsetting, and address formation are single field multiplications). \subsection{Hints}\label{sec:prog-hints} @@ -9,9 +9,9 @@ \subsection{Hints}\label{sec:prog-hints} \subsection{Division and inequality}\label{sec:prog-div-ne} -The zkDSL's single-slash division $q=a/b$ uses write-once back-solving: it emits \texttt{MUL\_NATIVE} with $q$ as the one unset input and $a$ as the already-written output. Witness generation fills $q=a\,b^{-1}$, while the ordinary multiplication table proves $q b=a$. It costs one VM instruction. Division by zero is undefined. Double slash \texttt{//} remains compile-time integer floor division for sizes and indices; it is not a field operation. +The zkDSL's single-slash division $q=a/b$ uses write-once back-solving: it emits \texttt{MUL64} with $q$ as the one unset input and $a$ as the already-written output. Witness generation fills $q=a\,b^{-1}$, while the ordinary multiplication table proves $q b=a$. It costs one VM instruction. Division by zero is undefined. Double slash \texttt{//} remains compile-time integer floor division for sizes and indices; it is not a field operation. -The assertion \texttt{assert a != b} computes $x=a+b$ with \texttt{XOR}, takes a hinted $w=x^{-1}$, forms $p=x\,w$ with \texttt{MUL\_NATIVE}, and writes $1$ to $p$ with \texttt{SET\_CONSTANT}. Memory being write-once, the two writes to $p$ agree only if $p=1$, and $x=0$ forces $p=0$ whatever the hint. +The assertion \texttt{assert a != b} computes $x=a+b$ with \texttt{XOR64}, takes a hinted $w=x^{-1}$, forms $p=x\,w$ with \texttt{MUL64}, and writes $1$ to $p$ with \texttt{SET\_CONSTANT}. Memory being write-once, the two writes to $p$ agree only if $p=1$, and $x=0$ forces $p=0$ whatever the hint. \subsection{Functions}\label{sec:prog-functions} @@ -28,33 +28,33 @@ \subsection{Functions}\label{sec:prog-functions} \subsection{Loops}\label{sec:prog-loops} -A counted loop keeps its counter in the exponent: it starts at $\gen^{\mathrm{lo}}$, advances by one multiplication $i\gets\gen\cdot i$, and exits on reaching $\gen^{\mathrm{hi}}$. The body is lowered as a tail-recursive helper function whose exit test \emph{is} the recursive call's \texttt{JUMP} condition (the \texttt{XOR} $i+\gen^{\mathrm{hi}}$ is nonzero exactly while the loop must continue), so an iteration costs one call plus two instructions, with no is-zero gadget. Loop state is threaded through captured arguments or a heap buffer. +A counted loop keeps its counter in the exponent: it starts at $\gen^{\mathrm{lo}}$, advances by one multiplication $i\gets\gen\cdot i$, and exits on reaching $\gen^{\mathrm{hi}}$. The body is lowered as a tail-recursive helper function whose exit test \emph{is} the recursive call's \texttt{JUMP} condition (the \texttt{XOR64} $i+\gen^{\mathrm{hi}}$ is nonzero exactly while the loop must continue), so an iteration costs one call plus two instructions, with no is-zero gadget. Loop state is threaded through captured arguments or a heap buffer. \subsection{Conditionals}\label{sec:prog-conditionals} -\texttt{if}/\texttt{else} on a field equality is one \texttt{XOR} plus one conditional \texttt{JUMP}: the taken jump goes to whichever block the test should \emph{not} fall into ($=$ falls into \emph{then}, $\neq$ into \emph{else}), so no negation is ever computed. Intra-function jump targets are bytecode addresses $\gen^{\mathrm{entry}+i}$, backpatched at layout. +\texttt{if}/\texttt{else} on a field equality is one \texttt{XOR64} plus one conditional \texttt{JUMP}: the taken jump goes to whichever block the test should \emph{not} fall into ($=$ falls into \emph{then}, $\neq$ into \emph{else}), so no negation is ever computed. Intra-function jump targets are bytecode addresses $\gen^{\mathrm{entry}+i}$, backpatched at layout. A taken \texttt{JUMP} reloads $\fp$ from a cell, and the ISA has no $\fp$-read; a branching function therefore materializes its own $\fp$ once: a \texttt{DEREF} in mode \texttt{fp} writes $\fp$ into a fresh one-cell heap buffer, a \texttt{DEREF} in mode \texttt{cell} copies it back into the frame (2 instructions, amortized; in \texttt{main}, $\fp=\gen^{0}=1$ is already a constant). Bindings made inside a branch are local to it; branches communicate through write-once cells: only one branch executes, so both may write the \emph{same} cell. -\subsection{Checking that a value is in the subfield \texorpdfstring{$\K$}{K}}\label{sec:prog-subfield-checks} +\subsection{Values in \texorpdfstring{$\E$}{E}}\label{sec:prog-values-in-e} -To check that $\loc{o_1}$ and $\loc{o_2}$ are both in $\K$, it suffices to emit \texttt{JUMP} $[o_0,o_1,o_2]$ with $\loc{o_0}=0$. Indeed, the jump falls through, while \texttt{JUMP} unconditionally requires all three values it reads to be in $\K$. +Every memory word is in $\K$, so a value in $\E$ (a transcript challenge, say) is a run of three consecutive cells, its limbs low first (\S\ref{sec:vm}), and each run spells exactly one element. The compiler lowers arithmetic on such runs to \texttt{XOR192} and \texttt{MUL192}, with the patterns of \S\ref{sec:prog-div-ne}: a division back-solves the unset operand of a \texttt{MUL192}; an equality is an \texttt{XOR192} into a pooled run holding zero, whose second write is the check; an inequality takes a hinted inverse and writes the \texttt{MUL192} product into a pooled run holding one. A copy between two runs is one \texttt{XOR192} against the zero run. \subsection{Range checks}\label{sec:prog-range-checks} The range check \emph{in the exponent} proves $x\in\{\gen^{0},\dots,\gen^{k-1}\}$, i.e.\ $\log_{\gen}x Date: Mon, 14 Sep 2026 20:52:12 -0400 Subject: [PATCH 04/40] update readme with new benchmarks --- README.md | 56 +++++++++++++++++++++++++++---------------------------- 1 file changed, 28 insertions(+), 28 deletions(-) diff --git a/README.md b/README.md index e1bd5a0bb..ed96ba70b 100644 --- a/README.md +++ b/README.md @@ -77,12 +77,12 @@ cargo run --release -- aggregate --xmss 900 --log-inv-rate 1 --repeat 3 ``` aggregation, 900 XMSS signatures - cycles (VM steps) : 1,573,849 = 2^20.586 - details : DEREF 2^18.947 (32.1%) SET 2^18.528 (24.0%) MUL 2^18.259 (19.9%) BLAKE2S 2^16.989 (8.3%) XOR 2^16.979 (8.2%) JUMP 2^16.839 (7.4%) MEMORY 2^21.305 TOTAL_COMMITTED 2^26.195 - proof size : 295.4 KiB - proving time : 0.749 s ± 2.1% peak memory 8.769 GiB - per signature : 1,201.795 signatures/s - verifying : 3.799 ms + cycles (VM steps) : 2,188,060 = 2^21.061 + details : SET 2^19.265 (28.8%) DEREF 2^19.126 (26.2%) MUL64 2^18.734 (19.9%)XOR64 2^18.21 (13.9%) BLAKE2S 2^16.989 (5.9%) JUMP 2^16.827 (5.3%) XOR192 2^5.209 (0.0%) MEMORY 2^22.014 TOTAL_COMMITTED 2^26.384 + proof size : 306.6 KiB + proving time : 0.868 s ± 0.8% peak memory 9.837 GiB + per signature : 1,036.452 signatures/s + verifying : 4.237 ms ``` ### SPHINCS aggregation @@ -95,12 +95,12 @@ cargo run --release -- aggregate --sphincs 245 --log-inv-rate 1 --repeat 3 ``` aggregation, 245 SPHINCS signatures - cycles (VM steps) : 2,132,425 = 2^21.024 - details : XOR 2^18.951 (23.8%) MUL 2^18.93 (23.4%) SET 2^18.847 (22.1%) DEREF 2^18.711 (20.1%) BLAKE2S 2^16.992 (6.1%) JUMP 2^16.543 (4.5%) MEMORY 2^21.564 TOTAL_COMMITTED 2^26.301 - proof size : 300.1 KiB - proving time : 0.861 s ± 1.5% peak memory 9.348 GiB - per signature : 284.603 signatures/s - verifying : 3.841 ms + cycles (VM steps) : 2,581,573 = 2^21.3 + details : MUL64 2^19.304 (25.1%) XOR64 2^19.139 (22.4%) SET 2^19.13 (22.2%) DEREF 2^19.089 (21.6%) BLAKE2S 2^16.992 (5.0%) JUMP 2^16.534 (3.7%) XOR192 2^4.858 (0.0%) MEMORY 2^22.056 TOTAL_COMMITTED 2^26.562 + proof size : 316.4 KiB + proving time : 1.058 s ± 4.7% peak memory 12.524 GiB + per signature : 231.676 signatures/s + verifying : 4.348 ms ``` ### data availability @@ -111,12 +111,12 @@ cargo run --release -- aggregate --blobs 16 --log-inv-rate 1 --repeat 3 ``` aggregation, 16 blobs - cycles (VM steps) : 2,989,506 = 2^21.511 - details : MUL 2^19.899 (32.7%) XOR 2^19.809 (30.7%) DEREF 2^19.138 (19.3%) JUMP 2^18.299 (10.8%) SET 2^16.787 (3.8%) BLAKE2S 2^16.295 (2.7%) MEMORY 2^21.696 TOTAL_COMMITTED 2^26.695 - proof size : 322.4 KiB - proving time : 0.998 s ± 0.9% peak memory 12.646 GiB - blob throughput : 16.032 blobs/s, 2.004 MiB/s - verifying : 6.659 ms + cycles (VM steps) : 5,220,524 = 2^22.316 + details : MUL64 2^20.672 (32.0%) DEREF 2^20.667 (31.9%) XOR64 2^20.595 (30.3%) SET 2^17.643 (3.9%) BLAKE2S 2^16.295 (1.5%) JUMP 2^13.448 (0.2%) XOR192 2^12.139 (0.1%) MEMORY 2^22.545 TOTAL_COMMITTED 2^26.952 + proof size : 342.6 KiB + proving time : 1.741 s ± 82.6% peak memory 16.34 GiB + blob throughput : 9.189 blobs/s, 1.149 MiB/s + verifying : 6.892 ms ``` ### recursion @@ -127,11 +127,11 @@ cargo run --release -- recursion --n 2 --xmss-per-leaf 900 --log-inv-rate 2 --re ``` recursion 2→1, over leaves of 900 XMSS signatures - cycles (VM steps) : 570,113 = 2^19.121 - details : MUL 2^17.838 (41.1%) DEREF 2^16.988 (22.8%) XOR 2^16.747 (19.3%) SET 2^15.79 (9.9%) JUMP 2^14.488 (4.0%) BLAKE2S 2^13.978 (2.8%) MEMORY 2^19.507 TOTAL_COMMITTED 2^24.086 - proof size : 191.3 KiB - proving time : 0.287 s ± 15.9% peak memory 10.124 GiB - verifying : 4.121 ms + cycles (VM steps) : 834,835 = 2^19.671 + details : MUL64 2^18.093 (33.5%) DEREF 2^18.064 (32.8%) SET 2^16.321 (9.8%) XOR64 2^16.299 (9.7%) MUL192 2^15.582 (5.9%) XOR192 2^15.424 (5.3%) BLAKE2S 2^13.997 (2.0%) JUMP 2^13.206 (1.1%) MEMORY 2^20.434 TOTAL_COMMITTED 2^24.681 + proof size : 209.0 KiB + proving time : 0.394 s ± 6.0% peak memory 11.354 GiB + verifying : 4.346 ms ``` ### hashing @@ -164,11 +164,11 @@ cargo run --release -- fibonacci --n 2000000 --log-inv-rate 1 --repeat 3 ``` Fibonacci (in the exponent, i.e. modulo 2^64 - 1), N = 2,000,000 - cycles (VM steps) : 2,127,880 - details : MUL 2^20.944 (98.9%) SET 2^13.288 (0.5%) DEREF 2^12.967 (0.4%) JUMP 2^10.968 (0.1%) XOR2^10.966 (0.1%) MEMORY 2^20.96 TOTAL_COMMITTED 2^25.26 - proof size : 285.4 KiB - proving : 0.391 s ± 1.1% 5,442,734 cycles/s peak memory 5.203 GiB - verifying : 2.092 ms + cycles (VM steps) : 2,127,882 + details : MUL64 2^20.944 (98.9%) SET 2^13.288 (0.5%) DEREF 2^12.967 (0.4%) JUMP 2^10.968 (0.1%) XOR64 2^10.966 (0.1%) MEMORY 2^20.96 TOTAL_COMMITTED 2^24.716 + proof size : 299.3 KiB + proving : 0.272 s ± 10.7% 7,813,917 cycles/s peak memory 3.418 GiB + verifying : 2.3 ms ``` ## SNARK machinery From 208a71c7964586f1299694e9fe088b131869b3e8 Mon Sep 17 00:00:00 2001 From: Tom Wambsgans Date: Thu, 17 Sep 2026 13:11:29 -0400 Subject: [PATCH 05/40] Remove the signature schemes, LeanDA and recursive aggregation This branch explores a plain zkVM, so the XMSS and SPHINCS crates, the LeanDA encoder and the aggregation crate with its recursion guest all go. The Fibonacci benchmark moves to the root package, the BLAKE2s hash chain becomes the no-arena test, and the public API is now compile, prove and verify. Co-Authored-By: Claude Fable 5.1 --- Cargo.lock | 2485 +------- Cargo.toml | 13 +- crates/lean_da/Cargo.toml | 19 - crates/lean_da/src/commit.rs | 223 - crates/lean_da/src/encode.rs | 84 - crates/lean_da/src/lib.rs | 63 - crates/lean_da/src/membership.rs | 206 - crates/rec_aggregation/Cargo.toml | 26 - .../rec_aggregation/guests/lean_ethereum.py | 3717 ------------ crates/rec_aggregation/src/aggregation.rs | 5182 ----------------- crates/rec_aggregation/src/benchmark.rs | 252 - crates/rec_aggregation/src/hash_chain.rs | 117 - crates/rec_aggregation/src/lib.rs | 45 - crates/rec_aggregation/src/signers_cache.rs | 322 - crates/rec_aggregation/tests/arena_prove.rs | 27 - crates/sphincs/Cargo.toml | 14 - crates/sphincs/src/fts.rs | 86 - crates/sphincs/src/hash.rs | 60 - crates/sphincs/src/lib.rs | 102 - crates/sphincs/src/ots.rs | 103 - crates/sphincs/src/sphincs.rs | 468 -- crates/sphincs/tests/sphincs_tests.rs | 223 - crates/xmss/Cargo.toml | 18 - crates/xmss/src/hash.rs | 56 - crates/xmss/src/lib.rs | 106 - crates/xmss/src/ssz_serialization.rs | 158 - crates/xmss/src/wots.rs | 156 - crates/xmss/src/xmss.rs | 383 -- crates/xmss/tests/xmss_tests.rs | 259 - .../rec_aggregation/src => src}/fibonacci.rs | 18 +- src/lib.rs | 53 +- src/main.rs | 67 +- tests/api.rs | 166 +- tests/no_arena.rs | 75 +- 34 files changed, 335 insertions(+), 15017 deletions(-) delete mode 100644 crates/lean_da/Cargo.toml delete mode 100644 crates/lean_da/src/commit.rs delete mode 100644 crates/lean_da/src/encode.rs delete mode 100644 crates/lean_da/src/lib.rs delete mode 100644 crates/lean_da/src/membership.rs delete mode 100644 crates/rec_aggregation/Cargo.toml delete mode 100644 crates/rec_aggregation/guests/lean_ethereum.py delete mode 100644 crates/rec_aggregation/src/aggregation.rs delete mode 100644 crates/rec_aggregation/src/benchmark.rs delete mode 100644 crates/rec_aggregation/src/hash_chain.rs delete mode 100644 crates/rec_aggregation/src/lib.rs delete mode 100644 crates/rec_aggregation/src/signers_cache.rs delete mode 100644 crates/rec_aggregation/tests/arena_prove.rs delete mode 100644 crates/sphincs/Cargo.toml delete mode 100644 crates/sphincs/src/fts.rs delete mode 100644 crates/sphincs/src/hash.rs delete mode 100644 crates/sphincs/src/lib.rs delete mode 100644 crates/sphincs/src/ots.rs delete mode 100644 crates/sphincs/src/sphincs.rs delete mode 100644 crates/sphincs/tests/sphincs_tests.rs delete mode 100644 crates/xmss/Cargo.toml delete mode 100644 crates/xmss/src/hash.rs delete mode 100644 crates/xmss/src/lib.rs delete mode 100644 crates/xmss/src/ssz_serialization.rs delete mode 100644 crates/xmss/src/wots.rs delete mode 100644 crates/xmss/src/xmss.rs delete mode 100644 crates/xmss/tests/xmss_tests.rs rename {crates/rec_aggregation/src => src}/fibonacci.rs (87%) diff --git a/Cargo.lock b/Cargo.lock index 2e1dd0e70..0f1e69996 100644 --- a/Cargo.lock +++ b/Cargo.lock @@ -11,54 +11,6 @@ dependencies = [ "memchr", ] -[[package]] -name = "alloy-primitives" -version = "1.7.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ce7b00f0cb42c66ec353076ded1dff1fbf818f6e0e26c40c8a8456c04483fca4" -dependencies = [ - "alloy-rlp", - "bytes", - "cfg-if", - "const-hex", - "derive_more", - "fixed-cache", - "foldhash", - "hashbrown 0.17.1", - "indexmap 2.14.1", - "itoa", - "k256", - "keccak-asm", - "paste", - "proptest", - "rand 0.9.4", - "rapidhash", - "ruint", - "rustc-hash", - "secp256k1", - "serde", - "sha3", -] - -[[package]] -name = "alloy-rlp" -version = "0.3.16" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "24671b1f62edcf0f9b62994c7bf72cd621a04a4b99f5020ece1a647b40e2f103" -dependencies = [ - "arrayvec", - "bytes", -] - -[[package]] -name = "android_system_properties" -version = "0.1.6" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ae221649c9976a6f6c56ae1facf410f3ddb33cc661c4b7b61020a912d4237fbc" -dependencies = [ - "libc", -] - [[package]] name = "ansi_term" version = "0.12.1" @@ -119,2088 +71,366 @@ dependencies = [ ] [[package]] -name = "ark-ff" -version = "0.3.0" +name = "bincode" +version = "1.3.3" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "6b3235cc41ee7a12aaaf2c575a2ad7b46713a8a50bda2fc3b003a04845c05dd6" +checksum = "b1f45e9417d87227c7a56d22e471c6206462cba514c7590c09aff4cf6d1ddcad" dependencies = [ - "ark-ff-asm 0.3.0", - "ark-ff-macros 0.3.0", - "ark-serialize 0.3.0", - "ark-std 0.3.0", - "derivative", - "num-bigint", - "num-traits", - "paste", - "rustc_version 0.3.3", - "zeroize", + "serde", ] [[package]] -name = "ark-ff" -version = "0.4.2" +name = "cfg-if" +version = "1.0.4" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ec847af850f44ad29048935519032c33da8aa03340876d351dfab5660d2966ba" -dependencies = [ - "ark-ff-asm 0.4.2", - "ark-ff-macros 0.4.2", - "ark-serialize 0.4.2", - "ark-std 0.4.0", - "derivative", - "digest 0.10.7", - "itertools 0.10.5", - "num-bigint", - "num-traits", - "paste", - "rustc_version 0.4.1", - "zeroize", -] +checksum = "9330f8b2ff13f34540b44e946ef35111825727b38d33286ef986142615121801" [[package]] -name = "ark-ff" -version = "0.5.0" +name = "clap" +version = "4.6.1" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "a177aba0ed1e0fbb62aa9f6d0502e9b46dad8c2eab04c14258a1212d2557ea70" +checksum = "1ddb117e43bbf7dacf0a4190fef4d345b9bad68dfc649cb349e7d17d28428e51" dependencies = [ - "ark-ff-asm 0.5.0", - "ark-ff-macros 0.5.0", - "ark-serialize 0.5.0", - "ark-std 0.5.0", - "arrayvec", - "digest 0.10.7", - "educe", - "itertools 0.13.0", - "num-bigint", - "num-traits", - "paste", - "zeroize", + "clap_builder", + "clap_derive", ] [[package]] -name = "ark-ff" -version = "0.6.0" +name = "clap_builder" +version = "4.6.0" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f7a806ac6c8307b929df4645776290a50ee2aac754ad09d8bdf73391309e43af" +checksum = "714a53001bf66416adb0e2ef5ac857140e7dc3a0c48fb28b2f10762fc4b5069f" dependencies = [ - "ark-ff-asm 0.6.0", - "ark-ff-macros 0.6.0", - "ark-serialize 0.6.0", - "ark-std 0.6.0", - "digest 0.10.7", - "educe", - "num-bigint", - "num-traits", - "zeroize", + "anstream", + "anstyle", + "clap_lex", + "strsim", ] [[package]] -name = "ark-ff-asm" -version = "0.3.0" +name = "clap_derive" +version = "4.6.1" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "db02d390bf6643fb404d3d22d31aee1c4bc4459600aef9113833d17e786c6e44" +checksum = "f2ce8604710f6733aa641a2b3731eaa1e8b3d9973d5e3565da11800813f997a9" dependencies = [ + "heck", + "proc-macro2", "quote", - "syn 1.0.109", + "syn", ] [[package]] -name = "ark-ff-asm" -version = "0.4.2" +name = "clap_lex" +version = "1.1.0" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "3ed4aa4fe255d0bc6d79373f7e31d2ea147bcf486cba1be5ba7ea85abdb92348" -dependencies = [ - "quote", - "syn 1.0.109", -] +checksum = "c8d4a3bb8b1e0c1050499d1815f5ab16d04f0959b233085fb31653fbfc9d98f9" [[package]] -name = "ark-ff-asm" -version = "0.5.0" +name = "colorchoice" +version = "1.0.5" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "62945a2f7e6de02a31fe400aa489f0e0f5b2502e69f95f853adb82a96c7a6b60" -dependencies = [ - "quote", - "syn 2.0.118", -] +checksum = "1d07550c9036bf2ae0c684c4297d503f838287c83c53686d05370d0e139ae570" [[package]] -name = "ark-ff-asm" -version = "0.6.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "1479009684adc073dff49a1025d3a7065b317a9ead25aaaca38cdc70058ba8a2" +name = "fiat_shamir" +version = "0.1.0" dependencies = [ - "quote", - "syn 2.0.118", + "parallel", + "primitives", + "serde", ] [[package]] -name = "ark-ff-macros" -version = "0.3.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "db2fd794a08ccb318058009eefdf15bcaaaaf6f8161eb3345f907222bac38b20" +name = "flock" +version = "0.1.0" dependencies = [ - "num-bigint", - "num-traits", - "quote", - "syn 1.0.109", + "fiat_shamir", + "parallel", + "pcs", + "primitives", + "zk_alloc", ] [[package]] -name = "ark-ff-macros" -version = "0.4.2" +name = "getrandom" +version = "0.3.4" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7abe79b0e4288889c4574159ab790824d0033b9fdcb2a112a3182fac2e514565" +checksum = "899def5c37c4fd7b2664648c28120ecec138e4d395b459e5ca34f9cce2dd77fd" dependencies = [ - "num-bigint", - "num-traits", - "proc-macro2", - "quote", - "syn 1.0.109", + "cfg-if", + "libc", + "r-efi", + "wasip2", ] [[package]] -name = "ark-ff-macros" +name = "heck" version = "0.5.0" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "09be120733ee33f7693ceaa202ca41accd5653b779563608f1234f78ae07c4b3" -dependencies = [ - "num-bigint", - "num-traits", - "proc-macro2", - "quote", - "syn 2.0.118", -] - -[[package]] -name = "ark-ff-macros" -version = "0.6.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "4a0691ed21ef00ef89c1e9bda832eba493dda3ec2f8d892fb25b705f73f06bb8" -dependencies = [ - "num-bigint", - "num-traits", - "proc-macro2", - "quote", - "syn 2.0.118", -] - -[[package]] -name = "ark-serialize" -version = "0.3.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "1d6c2b318ee6e10f8c2853e73a83adc0ccb88995aa978d8a3408d492ab2ee671" -dependencies = [ - "ark-std 0.3.0", - "digest 0.9.0", -] +checksum = "2304e00983f87ffb38b55b444b5e3b60a884b5d30c0fca7d82fe33449bbe55ea" [[package]] -name = "ark-serialize" -version = "0.4.2" +name = "is_terminal_polyfill" +version = "1.70.2" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "adb7b85a02b83d2f22f89bd5cac66c9c89474240cb6207cb1efc16d098e822a5" -dependencies = [ - "ark-std 0.4.0", - "digest 0.10.7", - "num-bigint", -] +checksum = "a6cb138bb79a146c1bd460005623e142ef0181e3d0219cb493e02f7d08a35695" [[package]] -name = "ark-serialize" -version = "0.5.0" +name = "lazy_static" +version = "1.5.0" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "3f4d068aaf107ebcd7dfb52bc748f8030e0fc930ac8e360146ca54c1203088f7" -dependencies = [ - "ark-std 0.5.0", - "arrayvec", - "digest 0.10.7", - "num-bigint", -] +checksum = "bbd2bcb4c963f2ddae06a2efc7e9f3591312473c50c6685e1f298068316e66fe" [[package]] -name = "ark-serialize" -version = "0.6.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "a74dd304fd536fb95d0a328e72be759209cc496a9da094c5bc56e5fea4f9e86b" +name = "lean_compiler" +version = "0.1.0" dependencies = [ - "ark-serialize-derive", - "ark-std 0.6.0", - "digest 0.10.7", - "num-bigint", - "serde_with", + "bincode", + "lean_vm", + "primitives", + "rand", ] [[package]] -name = "ark-serialize-derive" -version = "0.6.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "4f153690697a2b91e5e1251ff98411ee5371500a111a0fd317a70e588eb300f9" +name = "lean_vm" +version = "0.1.0" dependencies = [ - "proc-macro2", - "quote", - "syn 2.0.118", + "bincode", + "fiat_shamir", + "flock", + "lean_compiler", + "parallel", + "pcs", + "primitives", + "tracing", + "zk_alloc", ] [[package]] -name = "ark-std" -version = "0.3.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "1df2c09229cbc5a028b1d70e00fdb2acee28b1055dfb5ca73eea49c5a25c4e7c" +name = "leanvm" +version = "0.1.0" dependencies = [ - "num-traits", - "rand 0.8.8", + "bincode", + "clap", + "lean_compiler", + "lean_vm", + "primitives", + "tracing", + "zk_alloc", ] [[package]] -name = "ark-std" -version = "0.4.0" +name = "libc" +version = "0.2.186" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "94893f1e0c6eeab764ade8dc4c0db24caf4fe7cbbaafc0eba0a9030f447b5185" -dependencies = [ - "num-traits", - "rand 0.8.8", -] +checksum = "68ab91017fe16c622486840e4c83c9a37afeff978bd239b5293d61ece587de66" [[package]] -name = "ark-std" -version = "0.5.0" +name = "log" +version = "0.4.33" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "246a225cc6131e9ee4f24619af0f19d67761fff15d7ccc22e42b80846e69449a" -dependencies = [ - "num-traits", - "rand 0.8.8", -] +checksum = "0ceec5bc11778974d1bcb055b18002eba7f4b3518b6a0081b3af5f21666da9ad" [[package]] -name = "ark-std" -version = "0.6.0" +name = "matchers" +version = "0.2.0" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "367c9c827ed431bff6868b7aa926e05b16eb46603cc8b6e768e4a5553fa1d155" +checksum = "d1525a2a28c7f4fa0fc98bb91ae755d1e2d1505079e05539e35bc876b5d65ae9" dependencies = [ - "num-traits", - "rand 0.8.8", + "regex-automata", ] [[package]] -name = "arrayvec" -version = "0.7.8" +name = "memchr" +version = "2.8.3" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d3fb67a6e08acf24fdeccbac2cb6ac4305825bd1f117462e0e6f2f193345ad56" +checksum = "cf8baf1c55e62ffcace7a9f06f4bd9cd3f0c4beb022d3b367256b91b87513d98" [[package]] -name = "auto_impl" -version = "1.3.0" +name = "nu-ansi-term" +version = "0.50.3" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ffdcb70bdbc4d478427380519163274ac86e52916e10f0a8889adf0f96d3fee7" +checksum = "7957b9740744892f114936ab4a57b3f487491bbeafaf8083688b16841a4240e5" dependencies = [ - "proc-macro2", - "quote", - "syn 2.0.118", + "windows-sys", ] [[package]] -name = "autocfg" -version = "1.5.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f2032f911046de80f0a198e0901378627c33f59ea0ac00e363d481118bd70a53" - -[[package]] -name = "base16ct" -version = "0.2.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "4c7f02d4ea65f2c1853089ffd8d2787bdbc63de2f0d29dedbcf8ccdfa0ccd4cf" - -[[package]] -name = "base64" -version = "0.22.1" +name = "once_cell" +version = "1.21.4" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "72b3254f16251a8381aa12e40e3c4d2f0199f8c6508fbecb9d91f575e0fbb8c6" +checksum = "9f7c3e4beb33f85d45ae3e3a1792185706c8e16d043238c593331cc7cd313b50" [[package]] -name = "base64ct" -version = "1.8.3" +name = "once_cell_polyfill" +version = "1.70.2" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "2af50177e190e07a26ab74f8b1efbfe2ef87da2116221318cb1c2e82baf7de06" +checksum = "384b8ab6d37215f3c5301a95a4accb5d64aa607f1fcb26a11b5303878451b4fe" [[package]] -name = "bincode" -version = "1.3.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b1f45e9417d87227c7a56d22e471c6206462cba514c7590c09aff4cf6d1ddcad" +name = "parallel" +version = "0.1.0" dependencies = [ - "serde", + "libc", ] [[package]] -name = "bitcoin-consensus-encoding" -version = "1.2.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "6712f9c6fd6785b3b270884e57c441c403dc5d7e19ca45368c97c7a1de3000ec" +name = "pcs" +version = "0.1.0" dependencies = [ - "bitcoin-internals", - "hex-conservative 1.2.0", + "bincode", + "fiat_shamir", + "parallel", + "primitives", "serde", + "tracing", + "zk_alloc", ] [[package]] -name = "bitcoin-internals" -version = "0.6.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d573f4cf32996a8dce612e4348cece65a241f1882ed594047c9ba348e8869fa5" - -[[package]] -name = "bitcoin-io" -version = "0.1.101" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "bb5de036369d1ac59d3c1819ebc4d850f89466f5401c571a285b6ed564a4cb78" -dependencies = [ - "bitcoin-consensus-encoding", -] - -[[package]] -name = "bitcoin_hashes" -version = "0.14.101" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "bca4c7abb40c8817d77403c880988cfd484f23ab2365726afb2f798363e2c4a2" -dependencies = [ - "bitcoin-io", - "hex-conservative 0.2.3", -] - -[[package]] -name = "bitflags" -version = "1.3.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "bef38d45163c2f1dde094a7dfd33ccf595c92905c8f8f4fdc18d06fb1037718a" - -[[package]] -name = "bitflags" -version = "2.13.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b588b76d00fde79687d7646a9b5bdf3cc0f655e0bbd080335a95d7e96f3587da" - -[[package]] -name = "bitvec" -version = "1.1.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ddcec3d12c579d40898fe0a9a358a803c23e9c52ca3c425707f81c9436211837" -dependencies = [ - "funty", - "radium", - "tap", - "wyz", -] - -[[package]] -name = "block-buffer" -version = "0.10.4" +name = "pin-project-lite" +version = "0.2.17" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "3078c7629b62d3f0439517fa394996acacc5cbc91c5a20d8c658e77abd503a71" -dependencies = [ - "generic-array", -] +checksum = "a89322df9ebe1c1578d689c92318e070967d1042b512afbe49518723f4e6d5cd" [[package]] -name = "block-buffer" -version = "0.12.1" +name = "ppv-lite86" +version = "0.2.21" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d2f6c7dbe95a6ed67ad9f18e57daf93a2f034c524b99fd2b76d18fdfeb6660aa" +checksum = "85eae3c4ed2f50dcfe72643da4befc30deadb458a9b590d720cde2f2b1e97da9" dependencies = [ - "hybrid-array", + "zerocopy", ] [[package]] -name = "bs58" -version = "0.5.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "bf88ba1141d185c399bee5288d850d63b8369520c1eafc32a0430b5b6c287bf4" +name = "primitives" +version = "0.1.0" dependencies = [ - "tinyvec", + "bincode", + "libc", + "parallel", + "primitives", + "rand", + "serde", + "tracing-forest", + "tracing-subscriber", + "zk_alloc", ] [[package]] -name = "bumpalo" -version = "3.20.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "72f5acc6cb2ba439de613abc23857ec3d78374d8ed5ac84e9d11336e87da8649" - -[[package]] -name = "byte-slice-cast" -version = "1.2.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7575182f7272186991736b70173b0ea045398f984bf5ebbb3804736ce1330c9d" - -[[package]] -name = "byteorder" -version = "1.5.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "1fd0f2584146f6f2ef48085050886acf353beff7305ebd1ae69500e27c67f64b" - -[[package]] -name = "bytes" -version = "1.12.1" +name = "proc-macro2" +version = "1.0.106" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "fc652a48c352aef3ea3aed32080501cf3ef6ed5da78602a020c991775b0aff04" -dependencies = [ - "serde", -] - -[[package]] -name = "cc" -version = "1.4.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0ad534f4357a5264cce5019c989cf66a4f0dc4e0d1b1d15f8aacec0ff7360273" -dependencies = [ - "find-msvc-tools", - "shlex", -] - -[[package]] -name = "cfg-if" -version = "1.0.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9330f8b2ff13f34540b44e946ef35111825727b38d33286ef986142615121801" - -[[package]] -name = "chrono" -version = "0.4.45" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "1aa79e62e7697b8e29b513a68abacf485adcd1fe8284a4316c5ae868e6633327" -dependencies = [ - "iana-time-zone", - "num-traits", - "serde", - "windows-link", -] - -[[package]] -name = "clap" -version = "4.6.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "1ddb117e43bbf7dacf0a4190fef4d345b9bad68dfc649cb349e7d17d28428e51" -dependencies = [ - "clap_builder", - "clap_derive", -] - -[[package]] -name = "clap_builder" -version = "4.6.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "714a53001bf66416adb0e2ef5ac857140e7dc3a0c48fb28b2f10762fc4b5069f" -dependencies = [ - "anstream", - "anstyle", - "clap_lex", - "strsim", -] - -[[package]] -name = "clap_derive" -version = "4.6.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f2ce8604710f6733aa641a2b3731eaa1e8b3d9973d5e3565da11800813f997a9" -dependencies = [ - "heck", - "proc-macro2", - "quote", - "syn 2.0.118", -] - -[[package]] -name = "clap_lex" -version = "1.1.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c8d4a3bb8b1e0c1050499d1815f5ab16d04f0959b233085fb31653fbfc9d98f9" - -[[package]] -name = "colorchoice" -version = "1.0.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "1d07550c9036bf2ae0c684c4297d503f838287c83c53686d05370d0e139ae570" - -[[package]] -name = "const-hex" -version = "1.19.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "33e2a781ebdf4467d1428dc4593067825fb646f6871475098d8577421af73558" -dependencies = [ - "cfg-if", - "cpufeatures 0.2.17", - "proptest", - "serde_core", -] - -[[package]] -name = "const-oid" -version = "0.9.6" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c2459377285ad874054d797f3ccebf984978aa39129f6eafde5cdc8315b612f8" - -[[package]] -name = "const_format" -version = "0.2.36" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "4481a617ad9a412be3b97c5d403fef8ed023103368908b9c50af598ff467cc1e" -dependencies = [ - "const_format_proc_macros", - "konst", -] - -[[package]] -name = "const_format_proc_macros" -version = "0.2.34" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "1d57c2eccfb16dbac1f4e61e206105db5820c9d26c3c472bc17c774259ef7744" -dependencies = [ - "proc-macro2", - "quote", - "unicode-xid", -] - -[[package]] -name = "convert_case" -version = "0.10.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "633458d4ef8c78b72454de2d54fd6ab2e60f9e02be22f3c6104cdc8a4e0fceb9" -dependencies = [ - "unicode-segmentation", -] - -[[package]] -name = "core-foundation-sys" -version = "0.8.7" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "773648b94d0e5d620f64f280777445740e61fe701025087ec8b57f45c791888b" - -[[package]] -name = "cpufeatures" -version = "0.2.17" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "59ed5838eebb26a2bb2e58f6d5b5316989ae9d08bab10e0e6d103e656d1b0280" -dependencies = [ - "libc", -] - -[[package]] -name = "cpufeatures" -version = "0.3.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "5ca28b0ae3115b884660db4118d803791fd6756b6e88f39c0f3f7859060d7566" -dependencies = [ - "libc", -] - -[[package]] -name = "crunchy" -version = "0.2.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "460fbee9c2c2f33933d720630a6a0bac33ba7053db5344fac858d4b8952d77d5" - -[[package]] -name = "crypto-bigint" -version = "0.5.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0dc92fb57ca44df6db8059111ab3af99a63d5d0f8375d9972e319a379c6bab76" -dependencies = [ - "generic-array", - "rand_core 0.6.4", - "subtle", - "zeroize", -] - -[[package]] -name = "crypto-common" -version = "0.1.6" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "1bfb12502f3fc46cca1bb51ac28df9d618d813cdc3d2f25b9fe775a34af26bb3" -dependencies = [ - "generic-array", - "typenum", -] - -[[package]] -name = "crypto-common" -version = "0.2.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ce6e4c961d6cd6c9a86db418387425e8bdeaf05b3c8bc1411e6dca4c252f1453" -dependencies = [ - "hybrid-array", -] - -[[package]] -name = "defmt" -version = "1.1.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e2953bfe4f93bbd20cc71198842756f77d161884c99ebbabc41d80231ded88d1" -dependencies = [ - "bitflags 1.3.2", - "defmt-macros", -] - -[[package]] -name = "defmt-macros" -version = "1.1.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "bad9c72e7ca2137e0dc3813245a0d282fd6daad32fd800af018306a9169b5fe8" -dependencies = [ - "defmt-parser", - "proc-macro2", - "quote", - "syn 2.0.118", -] - -[[package]] -name = "defmt-parser" -version = "1.0.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "10d60334b3b2e7c9d91ef8150abfb6fa4c1c39ebbcf4a81c2e346aad939fee3e" -dependencies = [ - "thiserror", -] - -[[package]] -name = "der" -version = "0.7.10" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e7c1832837b905bbfb5101e07cc24c8deddf52f93225eee6ead5f4d63d53ddcb" -dependencies = [ - "const-oid", - "zeroize", -] - -[[package]] -name = "deranged" -version = "0.5.8" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7cd812cc2bc1d69d4764bd80df88b4317eaef9e773c75226407d9bc0876b211c" -dependencies = [ - "serde_core", -] - -[[package]] -name = "derivative" -version = "2.2.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "fcc3dd5e9e9c0b295d6e1e4d811fb6f157d5ffd784b8d202fc62eac8035a770b" -dependencies = [ - "proc-macro2", - "quote", - "syn 1.0.109", -] - -[[package]] -name = "derive_more" -version = "2.1.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d751e9e49156b02b44f9c1815bcb94b984cdcc4396ecc32521c739452808b134" -dependencies = [ - "derive_more-impl", -] - -[[package]] -name = "derive_more-impl" -version = "2.1.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "799a97264921d8623a957f6c3b9011f3b5492f557bbb7a5a19b7fa6d06ba8dcb" -dependencies = [ - "convert_case", - "proc-macro2", - "quote", - "rustc_version 0.4.1", - "syn 2.0.118", - "unicode-xid", -] - -[[package]] -name = "digest" -version = "0.9.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d3dd60d1080a57a05ab032377049e0591415d2b31afd7028356dbf3cc6dcb066" -dependencies = [ - "generic-array", -] - -[[package]] -name = "digest" -version = "0.10.7" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9ed9a281f7bc9b7576e61468ba615a66a5c8cfdff42420a70aa82701a3b1e292" -dependencies = [ - "block-buffer 0.10.4", - "const-oid", - "crypto-common 0.1.6", - "subtle", -] - -[[package]] -name = "digest" -version = "0.11.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f1dd6dbb5841937940781866fa1281a1ff7bd3bf827091440879f9994983d5c2" -dependencies = [ - "block-buffer 0.12.1", - "crypto-common 0.2.2", -] - -[[package]] -name = "dyn-clone" -version = "1.0.20" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d0881ea181b1df73ff77ffaaf9c7544ecc11e82fba9b5f27b262a3c73a332555" - -[[package]] -name = "ecdsa" -version = "0.16.9" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ee27f32b5c5292967d2d4a9d7f1e0b0aed2c15daded5a60300e4abb9d8020bca" -dependencies = [ - "der", - "digest 0.10.7", - "elliptic-curve", - "rfc6979", - "signature", - "spki", -] - -[[package]] -name = "educe" -version = "0.6.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "1d7bc049e1bd8cdeb31b68bbd586a9464ecf9f3944af3958a7a9d0f8b9799417" -dependencies = [ - "enum-ordinalize", - "proc-macro2", - "quote", - "syn 2.0.118", -] - -[[package]] -name = "either" -version = "1.18.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "252afb9ae5eaa683babdc6a068b3f5726eb19e05070c731f9b2a23a7c3e8ed34" - -[[package]] -name = "elliptic-curve" -version = "0.13.8" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b5e6043086bf7973472e0c7dff2142ea0b680d30e18d9cc40f267efbf222bd47" -dependencies = [ - "base16ct", - "crypto-bigint", - "digest 0.10.7", - "ff", - "generic-array", - "group", - "pkcs8", - "rand_core 0.6.4", - "sec1", - "subtle", - "zeroize", -] - -[[package]] -name = "enum-ordinalize" -version = "4.4.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "89dd01549b09589510cf0647475075d12071456586d70f5c75c98ae2a5537677" -dependencies = [ - "enum-ordinalize-derive", -] - -[[package]] -name = "enum-ordinalize-derive" -version = "4.4.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "a65863d15a4ce2888bd2f0f543cc963d3879c3a022c8ee43f6141d479a3ac815" -dependencies = [ - "proc-macro2", - "quote", - "syn 3.0.4", -] - -[[package]] -name = "equivalent" -version = "1.0.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "877a4ace8713b0bcf2a4e7eec82529c029f1d0619886d18145fea96c3ffe5c0f" - -[[package]] -name = "ethereum_serde_utils" -version = "0.8.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "38df44a7a271ab43835678f9215b53cc2523e4714a215da6643d83dc110245da" -dependencies = [ - "alloy-primitives", - "hex", - "serde", - "serde_derive", - "serde_json", -] - -[[package]] -name = "ethereum_ssz" -version = "0.10.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e462875ad8693755ea8913d6e905715c76ea4836e2254e18c9cf0f7a8f8c2a13" -dependencies = [ - "alloy-primitives", - "ethereum_serde_utils", - "itertools 0.14.0", - "serde", - "serde_derive", - "smallvec", - "typenum", -] - -[[package]] -name = "fastrlp" -version = "0.3.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "139834ddba373bbdd213dffe02c8d110508dcf1726c2be27e8d1f7d7e1856418" -dependencies = [ - "arrayvec", - "auto_impl", - "bytes", -] - -[[package]] -name = "fastrlp" -version = "0.4.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ce8dba4714ef14b8274c371879b175aa55b16b30f269663f19d576f380018dc4" -dependencies = [ - "arrayvec", - "auto_impl", - "bytes", -] - -[[package]] -name = "ff" -version = "0.13.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c0b50bfb653653f9ca9095b427bed08ab8d75a137839d9ad64eb11810d5b6393" -dependencies = [ - "rand_core 0.6.4", - "subtle", -] - -[[package]] -name = "fiat_shamir" -version = "0.1.0" -dependencies = [ - "parallel", - "primitives", - "serde", -] - -[[package]] -name = "find-msvc-tools" -version = "0.1.11" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d45db016d36b838f563236e9193d0ee6ce38f3f68b6c94e914b4929c96bbb890" - -[[package]] -name = "fixed-cache" -version = "0.1.10" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "2fe63500644ef0269fe6b744e7e5dc5c20b5eebf3d881bc2be53f194636f6583" -dependencies = [ - "equivalent", - "rapidhash", -] - -[[package]] -name = "fixed-hash" -version = "0.8.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "835c052cb0c08c1acf6ffd71c022172e18723949c8282f2b9f27efbc51e64534" -dependencies = [ - "byteorder", - "rand 0.8.8", - "rustc-hex", - "static_assertions", -] - -[[package]] -name = "flock" -version = "0.1.0" -dependencies = [ - "fiat_shamir", - "parallel", - "pcs", - "primitives", - "zk_alloc", -] - -[[package]] -name = "foldhash" -version = "0.2.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "77ce24cb58228fbb8aa041425bb1050850ac19177686ea6e0f41a70416f56fdb" - -[[package]] -name = "funty" -version = "2.0.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e6d5a32815ae3f33302d95fdcb2ce17862f8c65363dcfd29360480ba1001fc9c" - -[[package]] -name = "futures-core" -version = "0.3.34" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "92d699e522242e69e3003b94ecc1f960f3a5e015aa7c5d7486e65ad01dd94f5e" - -[[package]] -name = "futures-task" -version = "0.3.34" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "cd417de3d1d015fc3bfd2b1ea46dfc7bab72ef86f1cc7cc9c78e728b34a6d1fd" - -[[package]] -name = "futures-util" -version = "0.3.34" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0d50a92467f8ba5dd6e3ee5d4bd04d73ab2e4e1c44474a0674821dfce14b79bc" -dependencies = [ - "futures-core", - "futures-task", - "pin-project-lite", - "slab", -] - -[[package]] -name = "generic-array" -version = "0.14.9" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "4bb6743198531e02858aeaea5398fcc883e71851fcbcb5a2f773e2fb6cb1edf2" -dependencies = [ - "typenum", - "version_check", - "zeroize", -] - -[[package]] -name = "getrandom" -version = "0.2.17" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ff2abc00be7fca6ebc474524697ae276ad847ad0a6b3faa4bcb027e9a4614ad0" -dependencies = [ - "cfg-if", - "libc", - "wasi", -] - -[[package]] -name = "getrandom" -version = "0.3.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "899def5c37c4fd7b2664648c28120ecec138e4d395b459e5ca34f9cce2dd77fd" -dependencies = [ - "cfg-if", - "libc", - "r-efi", - "wasip2", -] - -[[package]] -name = "group" -version = "0.13.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f0f9ef7462f7c099f518d754361858f86d8a07af53ba9af0fe635bbccb151a63" -dependencies = [ - "ff", - "rand_core 0.6.4", - "subtle", -] - -[[package]] -name = "hashbrown" -version = "0.12.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "8a9ee70c43aaf417c914396645a0fa852624801b24ebb7ae78fe8272889ac888" - -[[package]] -name = "hashbrown" -version = "0.17.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ed5909b6e89a2db4456e54cd5f673791d7eca6732202bbf2a9cc504fe2f9b84a" -dependencies = [ - "foldhash", - "serde", - "serde_core", -] - -[[package]] -name = "heck" -version = "0.5.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "2304e00983f87ffb38b55b444b5e3b60a884b5d30c0fca7d82fe33449bbe55ea" - -[[package]] -name = "hex" -version = "0.4.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7f24254aa9a54b5c858eaee2f5bccdb46aaf0e486a595ed5fd8f86ba55232a70" - -[[package]] -name = "hex-conservative" -version = "0.2.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "db3fef046dca3ca91ee1408a8c1b80ab777e80a4d308d1bf4e7adb3fcb047e08" -dependencies = [ - "arrayvec", -] - -[[package]] -name = "hex-conservative" -version = "1.2.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "35431185f361ccf3ffc58254628af5f1f5d5f28531da2e02e5d6c82bbc282a10" -dependencies = [ - "arrayvec", -] - -[[package]] -name = "hmac" -version = "0.12.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "6c49c37c09c17a53d937dfbb742eb3a961d65a994e6bcdcf37e7399d0cc8ab5e" -dependencies = [ - "digest 0.10.7", -] - -[[package]] -name = "hybrid-array" -version = "0.4.14" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "707114b52a152fa7bdb290cd7cd5912d9467273b6d74e21b8d81aca1f8533f6b" -dependencies = [ - "typenum", -] - -[[package]] -name = "iana-time-zone" -version = "0.1.65" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e31bc9ad994ba00e440a8aa5c9ef0ec67d5cb5e5cb0cc7f8b744a35b389cc470" -dependencies = [ - "android_system_properties", - "core-foundation-sys", - "iana-time-zone-haiku", - "js-sys", - "log", - "wasm-bindgen", - "windows-core", -] - -[[package]] -name = "iana-time-zone-haiku" -version = "0.1.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f31827a206f56af32e590ba56d5d2d085f558508192593743f16b2306495269f" -dependencies = [ - "cc", -] - -[[package]] -name = "impl-codec" -version = "0.6.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ba6a270039626615617f3f36d15fc827041df3b78c439da2cadfa47455a77f2f" -dependencies = [ - "parity-scale-codec", -] - -[[package]] -name = "impl-trait-for-tuples" -version = "0.2.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "a0eb5a3343abf848c0984fe4604b2b105da9539376e24fc0a3b0007411ae4fd9" -dependencies = [ - "proc-macro2", - "quote", - "syn 2.0.118", -] - -[[package]] -name = "indexmap" -version = "1.9.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "bd070e393353796e801d209ad339e89596eb4c8d430d18ede6a1cced8fafbd99" -dependencies = [ - "autocfg", - "hashbrown 0.12.3", - "serde", -] - -[[package]] -name = "indexmap" -version = "2.14.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "07aa2048142242915a31d35844fb311e0e53fcca590c3a0a40dcf1b841fa09eb" -dependencies = [ - "equivalent", - "hashbrown 0.17.1", - "serde", - "serde_core", -] - -[[package]] -name = "is_terminal_polyfill" -version = "1.70.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "a6cb138bb79a146c1bd460005623e142ef0181e3d0219cb493e02f7d08a35695" - -[[package]] -name = "itertools" -version = "0.10.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b0fd2260e829bddf4cb6ea802289de2f86d6a7a690192fbe91b3f46e0f2c8473" -dependencies = [ - "either", -] - -[[package]] -name = "itertools" -version = "0.13.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "413ee7dfc52ee1a4949ceeb7dbc8a33f2d6c088194d9f922fb8318faf1f01186" -dependencies = [ - "either", -] - -[[package]] -name = "itertools" -version = "0.14.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "2b192c782037fadd9cfa75548310488aabdbf3d2da73885b31bd0abd03351285" -dependencies = [ - "either", -] - -[[package]] -name = "itoa" -version = "1.0.18" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "8f42a60cbdf9a97f5d2305f08a87dc4e09308d1276d28c869c684d7777685682" - -[[package]] -name = "jiff" -version = "0.2.35" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "668b7183bd07af9a4885f5c35b0cc5c83c4607a913c16b7e17291832910d2dcc" -dependencies = [ - "defmt", - "jiff-core", - "jiff-static", - "jiff-tzdb-platform", - "log", - "portable-atomic", - "portable-atomic-util", - "serde_core", - "windows-link", -] - -[[package]] -name = "jiff-core" -version = "0.1.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7feca88439efe53da3754500c1851dedf3cb36c524dd5cf8225cc0794de95d09" -dependencies = [ - "defmt", -] - -[[package]] -name = "jiff-static" -version = "0.2.35" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "3a69dcb3a21cfb32ce1cd056169337ca284af0766dd766e7878819b251a49204" -dependencies = [ - "jiff-core", - "proc-macro2", - "quote", - "syn 2.0.118", -] - -[[package]] -name = "jiff-tzdb" -version = "0.1.8" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "142bd39932ad231f10513df9ab62661fead8719872150b7ad02a2df79f4e141e" - -[[package]] -name = "jiff-tzdb-platform" -version = "0.1.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "875a5a69ac2bab1a891711cf5eccbec1ce0341ea805560dcd90b7a2e925132e8" -dependencies = [ - "jiff-tzdb", -] - -[[package]] -name = "js-sys" -version = "0.3.104" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0e0c1080212aad755ea003d18543e8768dd432c48819efd73a7bf1e39b7a5a3a" -dependencies = [ - "cfg-if", - "futures-util", - "wasm-bindgen", -] - -[[package]] -name = "k256" -version = "0.13.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f6e3919bbaa2945715f0bb6d3934a173d1e9a59ac23767fbaaef277265a7411b" -dependencies = [ - "cfg-if", - "ecdsa", - "elliptic-curve", - "once_cell", - "sha2", -] - -[[package]] -name = "keccak" -version = "0.2.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d8f198d1db720e4940b5a493201d199d9f24f568f8f746bd13706243a2f71598" -dependencies = [ - "cfg-if", - "cpufeatures 0.3.1", -] - -[[package]] -name = "keccak-asm" -version = "0.1.8" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "dd5dc2c0d691cbf7595cde551ced329cca99c2387c2cbc97754c5d0cd045d3ee" -dependencies = [ - "digest 0.10.7", - "sha3-asm", -] - -[[package]] -name = "konst" -version = "0.2.20" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "128133ed7824fcd73d6e7b17957c5eb7bacb885649bd8c69708b2331a10bcefb" -dependencies = [ - "konst_macro_rules", -] - -[[package]] -name = "konst_macro_rules" -version = "0.2.19" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "a4933f3f57a8e9d9da04db23fb153356ecaf00cbd14aee46279c33dc80925c37" - -[[package]] -name = "lazy_static" -version = "1.5.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "bbd2bcb4c963f2ddae06a2efc7e9f3591312473c50c6685e1f298068316e66fe" - -[[package]] -name = "lean_compiler" -version = "0.1.0" -dependencies = [ - "bincode", - "lean_vm", - "primitives", - "rand 0.9.4", -] - -[[package]] -name = "lean_da" -version = "0.1.0" -dependencies = [ - "fiat_shamir", - "parallel", - "pcs", - "primitives", - "rand 0.9.4", - "serde", - "tracing", -] - -[[package]] -name = "lean_vm" -version = "0.1.0" -dependencies = [ - "bincode", - "fiat_shamir", - "flock", - "lean_compiler", - "parallel", - "pcs", - "primitives", - "tracing", - "zk_alloc", -] - -[[package]] -name = "leanvm" -version = "0.1.0" -dependencies = [ - "clap", - "lean_da", - "lean_vm", - "primitives", - "rand 0.9.4", - "rec_aggregation", - "sphincs", - "xmss", - "zk_alloc", -] - -[[package]] -name = "libc" -version = "0.2.186" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "68ab91017fe16c622486840e4c83c9a37afeff978bd239b5293d61ece587de66" - -[[package]] -name = "libm" -version = "0.2.16" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b6d2cec3eae94f9f509c767b45932f1ada8350c4bdb85af2fcab4a3c14807981" - -[[package]] -name = "log" -version = "0.4.33" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0ceec5bc11778974d1bcb055b18002eba7f4b3518b6a0081b3af5f21666da9ad" - -[[package]] -name = "matchers" -version = "0.2.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d1525a2a28c7f4fa0fc98bb91ae755d1e2d1505079e05539e35bc876b5d65ae9" -dependencies = [ - "regex-automata", -] - -[[package]] -name = "memchr" -version = "2.8.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "cf8baf1c55e62ffcace7a9f06f4bd9cd3f0c4beb022d3b367256b91b87513d98" - -[[package]] -name = "nu-ansi-term" -version = "0.50.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7957b9740744892f114936ab4a57b3f487491bbeafaf8083688b16841a4240e5" -dependencies = [ - "windows-sys", -] - -[[package]] -name = "num-bigint" -version = "0.4.8" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c89e69e7e0f03bea5ef08013795c25018e101932225a656383bd384495ecc367" -dependencies = [ - "num-integer", - "num-traits", -] - -[[package]] -name = "num-conv" -version = "0.2.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "521739c6d2bac4aa25192232afe6841231376b2b26d4d9fae5ecf8ca5772e441" - -[[package]] -name = "num-integer" -version = "0.1.47" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7ce2d95d4b3734dc35aa2f45e1aa22cd416814592a4f9d9205e11affd5b8e10b" -dependencies = [ - "num-traits", -] - -[[package]] -name = "num-traits" -version = "0.2.19" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "071dfc062690e90b734c0b2273ce72ad0ffa95f0c74596bc250dcfd960262841" -dependencies = [ - "autocfg", - "libm", -] - -[[package]] -name = "once_cell" -version = "1.21.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9f7c3e4beb33f85d45ae3e3a1792185706c8e16d043238c593331cc7cd313b50" - -[[package]] -name = "once_cell_polyfill" -version = "1.70.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "384b8ab6d37215f3c5301a95a4accb5d64aa607f1fcb26a11b5303878451b4fe" - -[[package]] -name = "parallel" -version = "0.1.0" -dependencies = [ - "libc", -] - -[[package]] -name = "parity-scale-codec" -version = "3.7.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "799781ae679d79a948e13d4824a40970bfa500058d245760dd857301059810fa" -dependencies = [ - "arrayvec", - "bitvec", - "byte-slice-cast", - "const_format", - "impl-trait-for-tuples", - "parity-scale-codec-derive", - "rustversion", - "serde", -] - -[[package]] -name = "parity-scale-codec-derive" -version = "3.7.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "34b4653168b563151153c9e4c08ebed57fb8262bebfa79711552fa983c623e7a" -dependencies = [ - "proc-macro-crate", - "proc-macro2", - "quote", - "syn 2.0.118", -] - -[[package]] -name = "paste" -version = "1.0.15" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "57c0d7b74b563b49d38dae00a0c37d4d6de9b432382b2892f0574ddcae73fd0a" - -[[package]] -name = "pcs" -version = "0.1.0" -dependencies = [ - "bincode", - "fiat_shamir", - "parallel", - "primitives", - "serde", - "tracing", - "zk_alloc", -] - -[[package]] -name = "pest" -version = "2.9.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "5a07a60cc7a4d00c91f95c685609d1d2f79050e6804b70ebedd7650f0b839bcf" -dependencies = [ - "memchr", - "ucd-trie", -] - -[[package]] -name = "pin-project-lite" -version = "0.2.17" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "a89322df9ebe1c1578d689c92318e070967d1042b512afbe49518723f4e6d5cd" - -[[package]] -name = "pkcs8" -version = "0.10.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f950b2377845cebe5cf8b5165cb3cc1a5e0fa5cfa3e1f7f55707d8fd82e0a7b7" -dependencies = [ - "der", - "spki", -] - -[[package]] -name = "portable-atomic" -version = "1.15.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "05c8b63e8d9609db387f0324918f81d68fe27748f084ef092fb35954d0539a85" - -[[package]] -name = "portable-atomic-util" -version = "0.2.7" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c2a106d1259c23fac8e543272398ae0e3c0b8d33c88ed73d0cc71b0f1d902618" -dependencies = [ - "portable-atomic", -] - -[[package]] -name = "powerfmt" -version = "0.2.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "439ee305def115ba05938db6eb1644ff94165c5ab5e9420d1c1bcedbba909391" - -[[package]] -name = "ppv-lite86" -version = "0.2.21" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "85eae3c4ed2f50dcfe72643da4befc30deadb458a9b590d720cde2f2b1e97da9" -dependencies = [ - "zerocopy", -] - -[[package]] -name = "primitive-types" -version = "0.12.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0b34d9fd68ae0b74a41b21c03c2f62847aa0ffea044eee893b4c140b37e244e2" -dependencies = [ - "fixed-hash", - "impl-codec", - "uint", -] - -[[package]] -name = "primitives" -version = "0.1.0" -dependencies = [ - "bincode", - "libc", - "parallel", - "primitives", - "rand 0.9.4", - "serde", - "tracing-forest", - "tracing-subscriber", - "zk_alloc", -] - -[[package]] -name = "proc-macro-crate" -version = "3.5.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e67ba7e9b2b56446f1d419b1d807906278ffa1a658a8a5d8a39dcb1f5a78614f" -dependencies = [ - "toml_edit", -] - -[[package]] -name = "proc-macro2" -version = "1.0.106" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "8fd00f0bb2e90d81d1044c2b32617f68fcb9fa3bb7640c23e9c748e53fb30934" +checksum = "8fd00f0bb2e90d81d1044c2b32617f68fcb9fa3bb7640c23e9c748e53fb30934" dependencies = [ "unicode-ident", ] [[package]] -name = "proptest" -version = "1.11.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "4b45fcc2344c680f5025fe57779faef368840d0bd1f42f216291f0dc4ace4744" -dependencies = [ - "bitflags 2.13.1", - "num-traits", - "rand 0.9.4", - "rand_chacha 0.9.0", - "rand_xorshift", - "regex-syntax", - "unarray", -] - -[[package]] -name = "quote" -version = "1.0.46" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "dfbc457d0c7a0759a614551b11a6409e5951f6c7537be1f1b7682b9ae9230368" -dependencies = [ - "proc-macro2", -] - -[[package]] -name = "r-efi" -version = "5.3.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "69cdb34c158ceb288df11e18b4bd39de994f6657d83847bdffdbd7f346754b0f" - -[[package]] -name = "radium" -version = "0.7.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "dc33ff2d4973d518d823d61aa239014831e521c75da58e3df4840d3f47749d09" - -[[package]] -name = "rand" -version = "0.8.8" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e058c7de0b26af77780c769414d6257830bb240f3c38477dbc2c16e5f54d6d4c" -dependencies = [ - "libc", - "rand_chacha 0.3.1", - "rand_core 0.6.4", -] - -[[package]] -name = "rand" -version = "0.9.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "44c5af06bb1b7d3216d91932aed5265164bf384dc89cd6ba05cf59a35f5f76ea" -dependencies = [ - "rand_chacha 0.9.0", - "rand_core 0.9.5", - "serde", -] - -[[package]] -name = "rand_chacha" -version = "0.3.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e6c10a63a0fa32252be49d21e7709d4d4baf8d231c2dbce1eaa8141b9b127d88" -dependencies = [ - "ppv-lite86", - "rand_core 0.6.4", -] - -[[package]] -name = "rand_chacha" -version = "0.9.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d3022b5f1df60f26e1ffddd6c66e8aa15de382ae63b3a0c1bfc0e4d3e3f325cb" -dependencies = [ - "ppv-lite86", - "rand_core 0.9.5", -] - -[[package]] -name = "rand_core" -version = "0.6.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ec0be4795e2f6a28069bec0b5ff3e2ac9bafc99e6a9a7dc3547996c5c816922c" -dependencies = [ - "getrandom 0.2.17", -] - -[[package]] -name = "rand_core" -version = "0.9.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "76afc826de14238e6e8c374ddcc1fa19e374fd8dd986b0d2af0d02377261d83c" -dependencies = [ - "getrandom 0.3.4", - "serde", -] - -[[package]] -name = "rand_xorshift" -version = "0.4.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "513962919efc330f829edb2535844d1b912b0fbe2ca165d613e4e8788bb05a5a" -dependencies = [ - "rand_core 0.9.5", -] - -[[package]] -name = "rapidhash" -version = "4.5.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "5da7e78a036ce858e8d55b7e7dc8ba3a88b78350fd2155d3591bbd966b58589e" -dependencies = [ - "rustversion", -] - -[[package]] -name = "rec_aggregation" -version = "0.1.0" -dependencies = [ - "bincode", - "flock", - "lean_compiler", - "lean_da", - "lean_vm", - "parallel", - "pcs", - "primitives", - "rand 0.9.4", - "serde", - "sphincs", - "tracing", - "xmss", - "zk_alloc", -] - -[[package]] -name = "ref-cast" -version = "1.0.27" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7e440fb4e4b4147295338efb76001ab9e4efc0e5839df2c47fc5ac2381d365c3" -dependencies = [ - "ref-cast-impl", -] - -[[package]] -name = "ref-cast-impl" -version = "1.0.27" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "92ecd8964f8453721699a1ed72037b0db49ce2f5a5138486ee89bed6f67cdf3a" -dependencies = [ - "proc-macro2", - "quote", - "syn 3.0.4", -] - -[[package]] -name = "regex-automata" -version = "0.4.16" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "8fcfdb36bda0c880c5931cdc7a2bcdc8ba4556847b9d912bca70bc94708711ad" -dependencies = [ - "aho-corasick", - "memchr", - "regex-syntax", -] - -[[package]] -name = "regex-syntax" -version = "0.8.11" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d6f6ff9a378485b298a5286656da665ba74413d36db0979633275d2e708145d4" - -[[package]] -name = "rfc6979" -version = "0.4.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f8dd2a808d456c4a54e300a23e9f5a67e122c3024119acbfd73e3bf664491cb2" -dependencies = [ - "hmac", - "subtle", -] - -[[package]] -name = "rlp" -version = "0.5.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "bb919243f34364b6bd2fc10ef797edbfa75f33c252e7998527479c6d6b47e1ec" -dependencies = [ - "bytes", - "rustc-hex", -] - -[[package]] -name = "ruint" -version = "1.20.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f5e99bff0393163bb25029a6af25d3d8d202ba5b5438a74d1bd8789f5c822970" -dependencies = [ - "alloy-rlp", - "ark-ff 0.3.0", - "ark-ff 0.4.2", - "ark-ff 0.5.0", - "ark-ff 0.6.0", - "bytes", - "fastrlp 0.3.1", - "fastrlp 0.4.0", - "num-bigint", - "num-integer", - "num-traits", - "parity-scale-codec", - "primitive-types", - "proptest", - "rand 0.8.8", - "rand 0.9.4", - "rlp", - "ruint-macro", - "serde_core", - "valuable", - "zeroize", -] - -[[package]] -name = "ruint-macro" -version = "1.2.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "48fd7bd8a6377e15ad9d42a8ec25371b94ddc67abe7c8b9127bec79bebaaae18" - -[[package]] -name = "rustc-hash" -version = "2.1.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "6b1e7f9a428571be2dc5bc0505c13fb6bf936822b894ec87abf8a08a4e51742d" - -[[package]] -name = "rustc-hex" -version = "2.1.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "3e75f6a532d0fd9f7f13144f392b6ad56a32696bfcd9c78f797f16bbb6f072d6" - -[[package]] -name = "rustc_version" -version = "0.3.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f0dfe2087c51c460008730de8b57e6a320782fbfb312e1f4d520e6c6fae155ee" -dependencies = [ - "semver 0.11.0", -] - -[[package]] -name = "rustc_version" -version = "0.4.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "cfcb3a22ef46e85b45de6ee7e79d063319ebb6594faafcf1c225ea92ab6e9b92" -dependencies = [ - "semver 1.0.28", -] - -[[package]] -name = "rustversion" -version = "1.0.23" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "cf54715a573b99ac80df0bc206da022bcd442c974952c7b9720069370852e21f" - -[[package]] -name = "schemars" -version = "0.9.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "4cd191f9397d57d581cddd31014772520aa448f65ef991055d7f61582c65165f" -dependencies = [ - "dyn-clone", - "ref-cast", - "serde", - "serde_json", -] - -[[package]] -name = "schemars" -version = "1.2.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "687274d293b6cdc6e73e0fee520bf2049650090d7164f87672d212a3c530cf4a" -dependencies = [ - "dyn-clone", - "ref-cast", - "serde", - "serde_json", -] - -[[package]] -name = "sec1" -version = "0.7.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d3e97a565f76233a6003f9f5c54be1d9c5bdfa3eccfb189469f11ec4901c47dc" -dependencies = [ - "base16ct", - "der", - "generic-array", - "pkcs8", - "subtle", - "zeroize", -] - -[[package]] -name = "secp256k1" -version = "0.31.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "2c3c81b43dc2d8877c216a3fccf76677ee1ebccd429566d3e67447290d0c42b2" -dependencies = [ - "bitcoin_hashes", - "rand 0.9.4", - "secp256k1-sys", -] - -[[package]] -name = "secp256k1-sys" -version = "0.11.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "dcb913707158fadaf0d8702c2db0e857de66eb003ccfdda5924b5f5ac98efb38" -dependencies = [ - "cc", -] - -[[package]] -name = "semver" -version = "0.11.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f301af10236f6df4160f7c3f04eec6dbc70ace82d23326abad5edee88801c6b6" -dependencies = [ - "semver-parser", -] - -[[package]] -name = "semver" -version = "1.0.28" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "8a7852d02fc848982e0c167ef163aaff9cd91dc640ba85e263cb1ce46fae51cd" - -[[package]] -name = "semver-parser" -version = "0.10.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9900206b54a3527fdc7b8a938bffd94a568bac4f4aa8113b209df75a09c0dec2" -dependencies = [ - "pest", -] - -[[package]] -name = "serde" -version = "1.0.228" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9a8e94ea7f378bd32cbbd37198a4a91436180c5bb472411e48b5ec2e2124ae9e" -dependencies = [ - "serde_core", - "serde_derive", -] - -[[package]] -name = "serde_core" -version = "1.0.228" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "41d385c7d4ca58e59fc732af25c3983b67ac852c1a25000afe1175de458b67ad" -dependencies = [ - "serde_derive", -] - -[[package]] -name = "serde_derive" -version = "1.0.228" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d540f220d3187173da220f885ab66608367b6574e925011a9353e4badda91d79" -dependencies = [ - "proc-macro2", - "quote", - "syn 2.0.118", -] - -[[package]] -name = "serde_json" -version = "1.0.151" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c841b55ecdae098c80dcae9cf767f6f8a0c2cdb3416bbef72181df4d0fe73f14" -dependencies = [ - "itoa", - "memchr", - "serde", - "serde_core", - "zmij", -] - -[[package]] -name = "serde_with" -version = "3.22.0" +name = "quote" +version = "1.0.46" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ee78f1fbe43ac4a0e47aadb3dbd357b69eb0d3793e948624cd03dd2750ab1c0a" +checksum = "dfbc457d0c7a0759a614551b11a6409e5951f6c7537be1f1b7682b9ae9230368" dependencies = [ - "base64", - "bs58", - "chrono", - "hex", - "indexmap 1.9.3", - "indexmap 2.14.1", - "jiff", - "schemars 0.9.0", - "schemars 1.2.2", - "serde_core", - "serde_json", - "time", + "proc-macro2", ] [[package]] -name = "sha2" -version = "0.10.9" +name = "r-efi" +version = "5.3.0" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "69cdb34c158ceb288df11e18b4bd39de994f6657d83847bdffdbd7f346754b0f" + +[[package]] +name = "rand" +version = "0.9.4" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "a7507d819769d01a365ab707794a4084392c824f54a7a6a7862f8c3d0892b283" +checksum = "44c5af06bb1b7d3216d91932aed5265164bf384dc89cd6ba05cf59a35f5f76ea" dependencies = [ - "cfg-if", - "cpufeatures 0.2.17", - "digest 0.10.7", + "rand_chacha", + "rand_core", ] [[package]] -name = "sha3" -version = "0.11.0" +name = "rand_chacha" +version = "0.9.0" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "be176f1a57ce4e3d31c1a166222d9768de5954f811601fb7ca06fc8203905ce1" +checksum = "d3022b5f1df60f26e1ffddd6c66e8aa15de382ae63b3a0c1bfc0e4d3e3f325cb" dependencies = [ - "digest 0.11.3", - "keccak", + "ppv-lite86", + "rand_core", ] [[package]] -name = "sha3-asm" -version = "0.1.8" +name = "rand_core" +version = "0.9.5" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "a6287fd675f713484342a89cbf0a386abef5f15919cfad607e5e1f19e1e15331" +checksum = "76afc826de14238e6e8c374ddcc1fa19e374fd8dd986b0d2af0d02377261d83c" dependencies = [ - "cc", - "cfg-if", + "getrandom", ] [[package]] -name = "sharded-slab" -version = "0.1.7" +name = "regex-automata" +version = "0.4.16" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f40ca3c46823713e0d4209592e8d6e826aa57e928f09752619fc696c499637f6" +checksum = "8fcfdb36bda0c880c5931cdc7a2bcdc8ba4556847b9d912bca70bc94708711ad" dependencies = [ - "lazy_static", + "aho-corasick", + "memchr", + "regex-syntax", ] [[package]] -name = "shlex" -version = "2.0.1" +name = "regex-syntax" +version = "0.8.11" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "f8fadd59c855ef2080decdef8ff161eb6661b86933c9d82e5ba29dc602a55aba" +checksum = "d6f6ff9a378485b298a5286656da665ba74413d36db0979633275d2e708145d4" [[package]] -name = "signature" -version = "2.2.0" +name = "serde" +version = "1.0.228" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "77549399552de45a898a580c1b41d445bf730df867cc44e6c0233bbc4b8329de" +checksum = "9a8e94ea7f378bd32cbbd37198a4a91436180c5bb472411e48b5ec2e2124ae9e" dependencies = [ - "digest 0.10.7", - "rand_core 0.6.4", + "serde_core", + "serde_derive", ] [[package]] -name = "slab" -version = "0.4.12" +name = "serde_core" +version = "1.0.228" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0c790de23124f9ab44544d7ac05d60440adc586479ce501c1d6d7da3cd8c9cf5" +checksum = "41d385c7d4ca58e59fc732af25c3983b67ac852c1a25000afe1175de458b67ad" +dependencies = [ + "serde_derive", +] [[package]] -name = "smallvec" -version = "1.15.2" +name = "serde_derive" +version = "1.0.228" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "8ed6a63f02c8539c91a8685a86f4099661ba3da017932f6ebbea6de3f0fa7c90" - -[[package]] -name = "sphincs" -version = "0.1.0" +checksum = "d540f220d3187173da220f885ab66608367b6574e925011a9353e4badda91d79" dependencies = [ - "parallel", - "primitives", - "rand 0.9.4", - "serde", + "proc-macro2", + "quote", + "syn", ] [[package]] -name = "spki" -version = "0.7.3" +name = "sharded-slab" +version = "0.1.7" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "d91ed6c858b01f942cd56b37a94b3e0a1798290327d1236e4d9cf4eaca44d29d" +checksum = "f40ca3c46823713e0d4209592e8d6e826aa57e928f09752619fc696c499637f6" dependencies = [ - "base64ct", - "der", + "lazy_static", ] [[package]] -name = "static_assertions" -version = "1.1.0" +name = "smallvec" +version = "1.15.2" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "a2eb9349b6444b326872e140eb1cf5e7c522154d69e7a0ffb0fb81c06b37543f" +checksum = "8ed6a63f02c8539c91a8685a86f4099661ba3da017932f6ebbea6de3f0fa7c90" [[package]] name = "strsim" @@ -2208,23 +438,6 @@ version = "0.11.1" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "7da8b5736845d9f2fcb837ea5d9e2628564b3b043a70948a3f0b778838c5fb4f" -[[package]] -name = "subtle" -version = "2.6.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "13c2bddecc57b384dee18652358fb23172facb8a2c51ccc10d74c157bdea3292" - -[[package]] -name = "syn" -version = "1.0.109" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "72b64191b275b66ffe2469e8af2c1cfe3bafa67b529ead792a6d0160888b4237" -dependencies = [ - "proc-macro2", - "quote", - "unicode-ident", -] - [[package]] name = "syn" version = "2.0.118" @@ -2236,23 +449,6 @@ dependencies = [ "unicode-ident", ] -[[package]] -name = "syn" -version = "3.0.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e6275cddf4610d1775e6d1fe9469b2e77d0f39fd98fb7450901b821e0c53649f" -dependencies = [ - "proc-macro2", - "quote", - "unicode-ident", -] - -[[package]] -name = "tap" -version = "1.0.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "55937e1799185b12863d447f42597ed69d9928686b8d88a1df17376a097d8369" - [[package]] name = "thiserror" version = "2.0.18" @@ -2270,7 +466,7 @@ checksum = "ebc4ee7f67670e9b64d05fa4253e753e016c6c95ff35b89b7941d6b856dec1d5" dependencies = [ "proc-macro2", "quote", - "syn 2.0.118", + "syn", ] [[package]] @@ -2282,81 +478,6 @@ dependencies = [ "cfg-if", ] -[[package]] -name = "time" -version = "0.3.55" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "cdb87b95ec50ddfa440816d227a17b2ccbdda963a316a727fda0fc4334f7d134" -dependencies = [ - "deranged", - "num-conv", - "powerfmt", - "serde_core", - "time-core", - "time-macros", -] - -[[package]] -name = "time-core" -version = "0.1.9" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9e1c906769ad99c88eaa54e728060edef082f8e358ff32030cb7c7d315e81109" - -[[package]] -name = "time-macros" -version = "0.2.32" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7e689342a48d2ea927c87ea50cabf8594854bf940e9310208848d680d668ed85" -dependencies = [ - "num-conv", - "time-core", -] - -[[package]] -name = "tinyvec" -version = "1.12.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "bb4ebadaa0af04fab11ae01eb5f9fdb5f9c5b875506e210e71c07873528baa7f" -dependencies = [ - "tinyvec_macros", -] - -[[package]] -name = "tinyvec_macros" -version = "0.1.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "1f3ccbac311fea05f86f61904b462b55fb3df8837a366dfc601a0161d0532f20" - -[[package]] -name = "toml_datetime" -version = "1.1.1+spec-1.1.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "3165f65f62e28e0115a00b2ebdd37eb6f3b641855f9d636d3cd4103767159ad7" -dependencies = [ - "serde_core", -] - -[[package]] -name = "toml_edit" -version = "0.25.13+spec-1.1.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "6975367e4d2ef766d86af01ffad14b622fecc8d4357a998fbc4deb6e9bacaf9b" -dependencies = [ - "indexmap 2.14.1", - "toml_datetime", - "toml_parser", - "winnow", -] - -[[package]] -name = "toml_parser" -version = "1.1.3+spec-1.1.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "1d38ac1cf9b95face32296c0a3ede1fdc270627c9d9c02a7274dd6d960dc4d56" -dependencies = [ - "winnow", -] - [[package]] name = "tracing" version = "0.1.44" @@ -2376,7 +497,7 @@ checksum = "7490cfa5ec963746568740651ac6781f701c9c5ea257c58e057f3ba8cf69e8da" dependencies = [ "proc-macro2", "quote", - "syn 2.0.118", + "syn", ] [[package]] @@ -2431,54 +552,12 @@ dependencies = [ "tracing-log", ] -[[package]] -name = "typenum" -version = "1.20.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b6f5e870be6c3b371b77fe0ee0bafb859fa4964b4404c27de1d380043c4dda20" - -[[package]] -name = "ucd-trie" -version = "0.1.7" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "2896d95c02a80c6d6a5d6e953d479f5ddf2dfdb6a244441010e373ac0fb88971" - -[[package]] -name = "uint" -version = "0.9.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "76f64bba2c53b04fcab63c01a7d7427eadc821e3bc48c34dc9ba29c501164b52" -dependencies = [ - "byteorder", - "crunchy", - "hex", - "static_assertions", -] - -[[package]] -name = "unarray" -version = "0.1.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "eaea85b334db583fe3274d12b4cd1880032beab409c0d774be044d4480ab9a94" - [[package]] name = "unicode-ident" version = "1.0.24" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "e6e4313cd5fcd3dad5cafa179702e2b244f760991f45397d14d4ebf38247da75" -[[package]] -name = "unicode-segmentation" -version = "1.13.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c6f5d3c3b1bf09027a88a6bc961fc00497d651009560b5463668dc81b0fa87a8" - -[[package]] -name = "unicode-xid" -version = "0.2.6" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ebc1c04c71510c7f702b52b7c350734c9ff1295c464a03335b00bb84fc54f853" - [[package]] name = "utf8parse" version = "0.2.2" @@ -2491,18 +570,6 @@ version = "0.1.1" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "ba73ea9cf16a25df0c8caa16c51acb937d5712a8429db78a3ee29d5dcacd3a65" -[[package]] -name = "version_check" -version = "0.9.5" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0b928f33d975fc6ad9f86c8f283853ad26bdd5b10b7f1542aa2fa15e2289105a" - -[[package]] -name = "wasi" -version = "0.11.1+wasi-snapshot-preview1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ccf3ec651a847eb01de73ccad15eb7d99f80485de043efb2f370cd654f4ea44b" - [[package]] name = "wasip2" version = "1.0.4+wasi-0.2.12" @@ -2512,51 +579,6 @@ dependencies = [ "wit-bindgen", ] -[[package]] -name = "wasm-bindgen" -version = "0.2.127" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "1b70935747edd64d89de3efa29d73789b806c15798f8e7dca4d8ac356b50ce70" -dependencies = [ - "cfg-if", - "once_cell", - "rustversion", - "wasm-bindgen-macro", - "wasm-bindgen-shared", -] - -[[package]] -name = "wasm-bindgen-macro" -version = "0.2.127" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "77775f8f3f7217702089053b94958f8f54061a3f663417df76e19cbdcca29bc1" -dependencies = [ - "quote", - "wasm-bindgen-macro-support", -] - -[[package]] -name = "wasm-bindgen-macro-support" -version = "0.2.127" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e11d33f857dc2fb11b8bc75aee111aa9cbeb12cd9f25efd3d4c2a3dd4e235284" -dependencies = [ - "bumpalo", - "proc-macro2", - "quote", - "syn 2.0.118", - "wasm-bindgen-shared", -] - -[[package]] -name = "wasm-bindgen-shared" -version = "0.2.127" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7ef64dbcc55df09c7e5a46182d181c2cfa3e925f3da937ea764728b4bbb9dcbf" -dependencies = [ - "unicode-ident", -] - [[package]] name = "winapi" version = "0.3.9" @@ -2579,65 +601,12 @@ version = "0.4.0" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "712e227841d057c1ee1cd2fb22fa7e5a5461ae8e48fa2ca79ec42cfc1931183f" -[[package]] -name = "windows-core" -version = "0.62.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b8e83a14d34d0623b51dce9581199302a221863196a1dde71a7663a4c2be9deb" -dependencies = [ - "windows-implement", - "windows-interface", - "windows-link", - "windows-result", - "windows-strings", -] - -[[package]] -name = "windows-implement" -version = "0.60.2" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "053e2e040ab57b9dc951b72c264860db7eb3b0200ba345b4e4c3b14f67855ddf" -dependencies = [ - "proc-macro2", - "quote", - "syn 2.0.118", -] - -[[package]] -name = "windows-interface" -version = "0.59.3" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "3f316c4a2570ba26bbec722032c4099d8c8bc095efccdc15688708623367e358" -dependencies = [ - "proc-macro2", - "quote", - "syn 2.0.118", -] - [[package]] name = "windows-link" version = "0.2.1" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "f0805222e57f7521d6a62e36fa9163bc891acd422f971defe97d64e70d0a4fe5" -[[package]] -name = "windows-result" -version = "0.4.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7781fa89eaf60850ac3d2da7af8e5242a5ea78d1a11c49bf2910bb5a73853eb5" -dependencies = [ - "windows-link", -] - -[[package]] -name = "windows-strings" -version = "0.5.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7837d08f69c77cf6b07689544538e017c1bfcf57e34b4c0ff58e6c2cd3b37091" -dependencies = [ - "windows-link", -] - [[package]] name = "windows-sys" version = "0.61.2" @@ -2647,42 +616,12 @@ dependencies = [ "windows-link", ] -[[package]] -name = "winnow" -version = "1.0.4" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "23b97319f7b8343df12cc98938e5c3eb436064524c8d2b4e30a1d3a36eecdf81" -dependencies = [ - "memchr", -] - [[package]] name = "wit-bindgen" version = "0.57.1" source = "registry+https://github.com/rust-lang/crates.io-index" checksum = "1ebf944e87a7c253233ad6766e082e3cd714b5d03812acc24c318f549614536e" -[[package]] -name = "wyz" -version = "0.5.1" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "05f360fc0b24296329c78fda852a1e9ae82de9cf7b27dae4b7f62f118f77b9ed" -dependencies = [ - "tap", -] - -[[package]] -name = "xmss" -version = "0.1.0" -dependencies = [ - "bincode", - "ethereum_ssz", - "parallel", - "primitives", - "rand 0.9.4", - "serde", -] - [[package]] name = "zerocopy" version = "0.8.52" @@ -2700,27 +639,7 @@ checksum = "1ae7f38b72ec2a254e2b87ef277cf2cd4fb97cbebf944faa6f33354da0867930" dependencies = [ "proc-macro2", "quote", - "syn 2.0.118", -] - -[[package]] -name = "zeroize" -version = "1.9.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e13c156562582aa81c60cb29407084cdb54c4164760106ab78e6c5b0858cf64e" -dependencies = [ - "zeroize_derive", -] - -[[package]] -name = "zeroize_derive" -version = "1.5.0" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "3c50655cbb0fe3fc43170059e702f1ce5e19b84cec58dc87b037a09935c2f328" -dependencies = [ - "proc-macro2", - "quote", - "syn 2.0.118", + "syn", ] [[package]] @@ -2729,9 +648,3 @@ version = "0.1.0" dependencies = [ "libc", ] - -[[package]] -name = "zmij" -version = "1.0.23" -source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "29666d0abbfad1e3dc4dcf6144730dd3a3ab225bbbdac83319345b1b44ccfc1b" diff --git a/Cargo.toml b/Cargo.toml index cb60b7e46..a20d6941a 100644 --- a/Cargo.toml +++ b/Cargo.toml @@ -12,13 +12,11 @@ workspace = true [dependencies] primitives.workspace = true -rec_aggregation.workspace = true lean_vm.workspace = true -lean_da.workspace = true -xmss.workspace = true -sphincs.workspace = true +lean_compiler.workspace = true zk_alloc.workspace = true -rand.workspace = true +bincode.workspace = true +tracing.workspace = true clap = { version = "4", features = ["derive"] } [workspace.package] @@ -40,16 +38,11 @@ pcs = { path = "crates/pcs" } flock = { path = "crates/flock" } lean_vm = { path = "crates/lean_vm" } lean_compiler = { path = "crates/lean_compiler" } -lean_da = { path = "crates/lean_da" } -rec_aggregation = { path = "crates/rec_aggregation" } -xmss = { path = "crates/xmss" } -sphincs = { path = "crates/sphincs" } zk_alloc = { path = "crates/zk_alloc" } parallel = { path = "crates/parallel" } libc = "0.2" serde = { version = "1", features = ["derive"] } bincode = "1" -ethereum_ssz = "0.10" rand = "0.9" tracing = "0.1.26" tracing-forest = { version = "0.3.0", features = ["ansi", "smallvec"] } diff --git a/crates/lean_da/Cargo.toml b/crates/lean_da/Cargo.toml deleted file mode 100644 index 79c118ad7..000000000 --- a/crates/lean_da/Cargo.toml +++ /dev/null @@ -1,19 +0,0 @@ -[package] -name = "lean_da" -version.workspace = true -edition.workspace = true -publish = false - -[lints] -workspace = true - -[dependencies] -primitives.workspace = true -parallel.workspace = true -pcs.workspace = true -fiat_shamir.workspace = true -serde.workspace = true -tracing.workspace = true - -[dev-dependencies] -rand.workspace = true diff --git a/crates/lean_da/src/commit.rs b/crates/lean_da/src/commit.rs deleted file mode 100644 index 8e8d594f8..000000000 --- a/crates/lean_da/src/commit.rs +++ /dev/null @@ -1,223 +0,0 @@ -//! Two commitment branches over the same cell digests. -//! -//! A row digest hashes its first `k/c` cell digests, covering the systematic -//! payload. A Merkle tree over these digests gives `root_row`. -//! A second tree has all cell digests as leaves, in column-major order; its -//! intermediate column roots authenticate samples, and its root is `root_col`. -//! The final commitment is `H(root_row, root_col)`. - -use fiat_shamir::merkle::{Hash, hash_pair}; -use primitives::hash::{OUT_LEN, hash, hash_many_dyn}; - -use crate::{CELL_SYMBOLS, CELLS_PER_ROW, CODEWORD_SYMBOLS, PAYLOAD_CELLS, encode_rows, row_count}; - -/// What the builder publishes. -#[derive(Clone, Copy, Debug, PartialEq, Eq, serde::Serialize, serde::Deserialize)] -pub struct DaCommitment { - /// `H(root_row, root_col)`, binding the matrix including its zero padding. - pub root: Hash, - pub root_row: Hash, - pub root_col: Hash, -} - -/// What the builder keeps, to answer samples and to feed the proof. -pub struct DaWitness { - /// `n_rows · m` symbols, row-major. Padding rows are zero and not stored. - pub codewords: Vec, - /// `n_rows · ℓ` cell digests, row-major. - pub cell_digests: Vec, - /// The flat tree over the column-major cell digests; see the module docs. - pub column_tree: Vec, - /// The flat tree over the `n_padded` row digests. - pub row_tree: Vec, -} - -impl DaWitness { - /// The `ℓ` column roots `C_j`, the level at height `log n_padded`. - pub fn column_roots(&self) -> &[Hash] { - let (n_pad, cells) = ( - row_count(self.codewords.len(), CODEWORD_SYMBOLS).next_power_of_two(), - CELLS_PER_ROW, - ); - let offset = 2 * n_pad * cells - 2 * cells; - &self.column_tree[offset..offset + cells] - } -} - -/// Commit to payload rows of little-endian 64-bit symbols and their systematic encoding. -#[tracing::instrument(name = "Commit", skip_all)] -pub fn commit(rows: &[u64]) -> (DaCommitment, DaWitness) { - let codewords = encode_rows(rows); - commit_codewords(codewords) -} - -/// [`commit`] over rows that are already encoded. -pub fn commit_codewords(codewords: Vec) -> (DaCommitment, DaWitness) { - let n_rows = row_count(codewords.len(), CODEWORD_SYMBOLS); - let (cells, n_pad, t) = (CELLS_PER_ROW, n_rows.next_power_of_two(), PAYLOAD_CELLS); - let cell_digests = hash_cells(&codewords); - let (padding_cell, padding_row) = padding_digests(); - - // Row branch: hash each payload's cell digests. - let mut prefixes = vec![Hash::default(); n_rows * t]; - parallel::chunks_mut(&mut prefixes, t, |i, prefix| { - prefix.copy_from_slice(&cell_digests[i * cells..i * cells + t]); - }); - let mut row_digests = vec![padding_row; n_pad]; - hash_many_dyn( - prefixes.as_flattened(), - t * OUT_LEN, - row_digests[..n_rows].as_flattened_mut(), - ); - drop(prefixes); - let row_tree = tree_from_leaves(row_digests); - - // Column branch: the same digests, column-major, one tree carrying both levels. - let mut column_major = vec![Hash::default(); cells * n_pad]; - parallel::chunks_mut(&mut column_major, n_pad, |j, column| { - for (i, slot) in column.iter_mut().enumerate() { - *slot = if i < n_rows { - cell_digests[i * cells + j] - } else { - padding_cell - }; - } - }); - let column_tree = tree_from_leaves(column_major); - - let (root_row, root_col) = (*row_tree.last().unwrap(), *column_tree.last().unwrap()); - let commitment = DaCommitment { - root: hash_pair(&root_row, &root_col), - root_row, - root_col, - }; - let witness = DaWitness { - codewords, - cell_digests, - column_tree, - row_tree, - }; - (commitment, witness) -} - -/// Shape-dependent padding digests: the zero cell and its repeated digest for a row. -pub fn padding_digests() -> (Hash, Hash) { - let cell = hash(&[0u8; CELL_SYMBOLS * size_of::()]); - (cell, hash([cell; PAYLOAD_CELLS].as_flattened())) -} - -/// `e_{i,j} = H(W_{i,j})` for every cell of every real row, row-major. -#[tracing::instrument(name = "Hashing cells", skip_all)] -fn hash_cells(codewords: &[u64]) -> Vec { - let (cells, m) = (CELLS_PER_ROW, CODEWORD_SYMBOLS); - let n_rows = codewords.len() / m; - let mut digests = vec![Hash::default(); n_rows * cells]; - parallel::chunks_mut(&mut digests, cells, |i, row| { - hash_many_dyn( - as_bytes(&codewords[i * m..(i + 1) * m]), - CELL_SYMBOLS * size_of::(), - row.as_flattened_mut(), - ); - }); - digests -} - -/// The flat Merkle tree over leaves that are already digests: `tree[..n]` is the -/// leaves, then each level in turn, the root last. Unlike [`pcs::merkle`] the -/// leaves are not re-hashed, since a cell digest is already the leaf. -fn tree_from_leaves(mut tree: Vec) -> Vec { - let n = tree.len(); - assert!(n.is_power_of_two(), "leaf count must be a power of two"); - tree.resize(2 * n - 1, Hash::default()); - - let (mut base, mut width) = (0, n); - while width > 1 { - let (read, write) = tree.split_at_mut(base + width); - hash_many_dyn( - read[base..].as_flattened(), - 2 * OUT_LEN, - write[..width / 2].as_flattened_mut(), - ); - base += width; - width /= 2; - } - tree -} - -fn as_bytes(data: &[u64]) -> &[u8] { - const { assert!(cfg!(target_endian = "little"), "digests use little-endian symbols") }; - // SAFETY: u64 has no padding; the endian check fixes the byte representation. - unsafe { core::slice::from_raw_parts(data.as_ptr().cast::(), size_of_val(data)) } -} - -#[cfg(test)] -mod tests { - use super::*; - use crate::BLOB_SYMBOLS; - use rand::{Rng, SeedableRng, rngs::StdRng}; - - fn payload(n_rows: usize, seed: u64) -> Vec { - let mut rng = StdRng::seed_from_u64(seed); - (0..n_rows * BLOB_SYMBOLS).map(|_| rng.random()).collect() - } - - /// `column_roots` indexes into the flat tree by an offset derived from the - /// shape, which is the easiest thing here to get wrong by a level. Each one has - /// to be the root of its own column's subtree, rebuilt independently. - #[test] - fn column_roots_are_their_own_subtrees() { - let n_rows = 5usize; - let (_, witness) = commit(&payload(n_rows, 1)); - let (n_pad, cells) = (n_rows.next_power_of_two(), CELLS_PER_ROW); - let padding_cell = hash(&[0u8; CELL_SYMBOLS * size_of::()]); - - for (j, &got) in witness.column_roots().iter().enumerate() { - let column: Vec = (0..n_pad) - .map(|i| { - if i < n_rows { - witness.cell_digests[i * cells + j] - } else { - padding_cell - } - }) - .collect(); - assert_eq!(got, *tree_from_leaves(column).last().unwrap(), "column {j}"); - } - } - - /// Both branches have to be binding, and independently so: a change inside the - /// first half must move `root_row`, and a change in the second half must - /// still move `root_col` even though no row digest covers it. - #[test] - fn both_branches_bind() { - let n_rows = 3usize; - let codewords = crate::encode_rows(&payload(n_rows, 2)); - let (base, _) = commit_codewords(codewords.clone()); - - for &position in &[0usize, BLOB_SYMBOLS - 1, BLOB_SYMBOLS, CODEWORD_SYMBOLS - 1] { - let mut corrupted = codewords.clone(); - corrupted[position] ^= 1; - let (moved, _) = commit_codewords(corrupted); - assert_ne!(moved.root, base.root, "root ignored a flip at {position}"); - assert_ne!(moved.root_col, base.root_col, "root_col ignored a flip at {position}"); - if position < BLOB_SYMBOLS { - assert_ne!(moved.root_row, base.root_row, "root_row ignored a payload flip"); - } - } - } - - /// Padding rows are zero codewords sharing one cell digest. A payload whose - /// real rows are unchanged must commit identically whether or not the row count - /// happens to be a power of two, up to the padding the shape declares. - #[test] - fn padding_rows_are_the_zero_codeword() { - let n_rows = 3usize; - let rows = payload(n_rows, 3); - let mut extended = rows.clone(); - extended.extend(std::iter::repeat_n(0, BLOB_SYMBOLS)); - - let (short, _) = commit(&rows); - let (long, _) = commit(&extended); - assert_eq!(short.root, long.root); - } -} diff --git a/crates/lean_da/src/encode.rs b/crates/lean_da/src/encode.rs deleted file mode 100644 index 1f3f27376..000000000 --- a/crates/lean_da/src/encode.rs +++ /dev/null @@ -1,84 +0,0 @@ -//! Systematic Reed-Solomon encoding at rate 1/2. -//! -//! An inverse additive NTT interpolates each payload row into novel-basis -//! coefficients. Encoding these on the full domain preserves the payload in -//! the first half of the codeword and adds redundancy in the second half. - -use pcs::ntt::AdditiveNttF64; -use primitives::field::F64; - -use crate::{BLOB_SYMBOLS, CODEWORD_SYMBOLS, DA_LOG_K, LOG_M, row_count}; - -/// Encode `n` rows of `k` symbols into `n` codewords of `m` symbols, row-major. -/// -/// `rows` is the payload, `n_rows · k` long. The returned buffer is `n_rows · m` -/// long, with each payload row unchanged in its codeword's first half. -/// Symbols are little-endian 64-bit words; every bit pattern is valid. -/// Padding rows are not materialized, being zero codewords. -#[tracing::instrument(name = "Encoding", skip_all)] -pub fn encode_rows(rows: &[u64]) -> Vec { - let n_rows = row_count(rows.len(), BLOB_SYMBOLS); - let (k, m) = (BLOB_SYMBOLS, CODEWORD_SYMBOLS); - let interpolation = AdditiveNttF64::standard(DA_LOG_K); - let ntt = AdditiveNttF64::standard(LOG_M); - - // One row at a time: the transform dispatches internally at any row worth - // encoding, and the pool panics on a nested dispatch, so the row loop must stay - // sequential and let the NTT own the parallelism. - let mut codewords = vec![0; n_rows * m]; - for (i, codeword) in codewords.chunks_exact_mut(m).enumerate() { - codeword[..k].copy_from_slice(&rows[i * k..(i + 1) * k]); - let codeword = as_field_mut(codeword); - interpolation.inverse_transform(&mut codeword[..k]); - ntt.encode_interleaved_in_place(codeword, 1, 1); - } - codewords -} - -/// Borrow native symbols as field elements without allocating or copying. -pub(crate) fn as_field_mut(data: &mut [u64]) -> &mut [F64] { - // SAFETY: F64 is repr(transparent) over u64, with every bit pattern valid. - unsafe { core::slice::from_raw_parts_mut(data.as_mut_ptr().cast::(), data.len()) } -} - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn encoding_preserves_payload_and_degree_bound() { - let payload: Vec = (0..3 * BLOB_SYMBOLS) - .map(|i| 0x9E37_79B9_7F4A_7C15u64.wrapping_mul(i as u64 + 1)) - .collect(); - let mut codewords = encode_rows(&payload); - let ntt = AdditiveNttF64::standard(LOG_M); - for (row, codeword) in payload - .as_chunks::() - .0 - .iter() - .zip(codewords.as_chunks_mut::().0.iter_mut()) - { - assert_eq!(&codeword[..BLOB_SYMBOLS], row); - ntt.inverse_transform(as_field_mut(codeword)); - assert!(codeword[BLOB_SYMBOLS..].iter().all(|&x| x == 0)); - } - } - - /// Encoding commutes with XOR of payloads. - #[test] - fn encoding_is_linear() { - let n = 2 * BLOB_SYMBOLS; - let a: Vec = (0..n) - .map(|i| 0x9E37_79B9_7F4A_7C15u64.wrapping_mul(i as u64 + 1)) - .collect(); - let b: Vec = (0..n) - .map(|i| 0xBF58_476D_1CE4_E5B9u64.wrapping_mul(i as u64 + 3)) - .collect(); - let sum: Vec = a.iter().zip(&b).map(|(&x, &y)| x ^ y).collect(); - - let (ca, cb, cs) = (encode_rows(&a), encode_rows(&b), encode_rows(&sum)); - for ((&x, &y), &z) in ca.iter().zip(&cb).zip(&cs) { - assert_eq!(x ^ y, z); - } - } -} diff --git a/crates/lean_da/src/lib.rs b/crates/lean_da/src/lib.rs deleted file mode 100644 index d8a66ce53..000000000 --- a/crates/lean_da/src/lib.rs +++ /dev/null @@ -1,63 +0,0 @@ -//! LeanDA over `K = GF(2^64)` with BLAKE2s, following -//! . -//! -//! Each payload row contains `k` symbols, systematically encoded to `m = 2k` -//! evaluations on an additive domain. The row branch commits to the first `k` -//! evaluations, which are the original payload; the column branch authenticates -//! sampled cells. [`check_membership`] checks that every row belongs to the code. -//! The aggregation guest proves this check together with both commitment branches. -//! -//! Row and cell widths are powers of two. Trees pad the row count with zero rows. -//! A root binds this padded matrix, not the original row count: a trailing zero -//! row is indistinguishable from padding within the same padded height. - -mod commit; -mod encode; -mod membership; - -pub use commit::{DaCommitment, DaWitness, commit, commit_codewords, padding_digests}; -pub use encode::encode_rows; -pub use membership::{ - check_membership, dual_codeword, membership_challenges, membership_vector, row_residuals, vector_digest, -}; - -/// Payload symbols per blob, as a logarithm (128 KiB). -pub const DA_LOG_K: usize = 14; -/// Symbols per sampling cell, as a logarithm (2 KiB). -pub const DA_LOG_CELL: usize = 8; -/// Maximum blobs in one commitment. -pub const DA_MAX_ROWS: usize = 1024; - -/// 64-bit symbols in one payload blob. -pub const BLOB_SYMBOLS: usize = 1 << DA_LOG_K; -/// 64-bit symbols in one encoded blob. -pub const CODEWORD_SYMBOLS: usize = 2 * BLOB_SYMBOLS; -/// 64-bit symbols in one sampling cell. -pub const CELL_SYMBOLS: usize = 1 << DA_LOG_CELL; -/// Sampling cells in one encoded blob. -pub const CELLS_PER_ROW: usize = CODEWORD_SYMBOLS / CELL_SYMBOLS; -const PAYLOAD_CELLS: usize = BLOB_SYMBOLS / CELL_SYMBOLS; -const LOG_M: usize = DA_LOG_K + 1; - -fn row_count(symbols: usize, width: usize) -> usize { - assert!(symbols.is_multiple_of(width), "partial blob row"); - let rows = symbols / width; - assert!((1..=DA_MAX_ROWS).contains(&rows), "expected 1..=DA_MAX_ROWS blobs"); - rows -} - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn row_counts_reject_empty_partial_and_oversized_inputs() { - for width in [BLOB_SYMBOLS, CODEWORD_SYMBOLS] { - assert_eq!(row_count(width, width), 1); - assert_eq!(row_count(DA_MAX_ROWS * width, width), DA_MAX_ROWS); - for symbols in [0, width - 1, width + 1, (DA_MAX_ROWS + 1) * width, usize::MAX] { - assert!(std::panic::catch_unwind(|| row_count(symbols, width)).is_err()); - } - } - } -} diff --git a/crates/lean_da/src/membership.rs b/crates/lean_da/src/membership.rs deleted file mode 100644 index fe82af744..000000000 --- a/crates/lean_da/src/membership.rs +++ /dev/null @@ -1,206 +0,0 @@ -//! Reed-Solomon membership on an additive domain. -//! -//! For an additive domain `D` of size `m`, `Σ_{x∈D} h(x) = 0` when -//! `deg(h) <= m-2`. Pairing codewords and comparing dimensions therefore gives -//! `RS_k(D)^⊥ = RS_{m-k}(D)`. At rate 1/2 the code is self-dual. -//! -//! The check uses `L(X) = ∏_{j < log2(k)} (1 + z_j · Vhat_j(X))`, where -//! `Vhat_j = X_{2^j}` is the normalized subspace polynomial of degree `2^j`. -//! Expanding this product gives every novel-basis coefficient below `k` with -//! tensor weights `⊗_j(1, z_j)`, so `L` is a codeword. -//! -//! For a fixed invalid row `w`, `⟨L,w⟩` is a nonzero multilinear polynomial in -//! `z`. Fresh independent uniform challenges accept it with probability at most -//! `log2(k)/2^192`. The external verifier derives the vector from the commitment. -//! The guest binds its hinted vector by hashing it and checks the inner products. - -use fiat_shamir::FiatShamirState; -use fiat_shamir::merkle::hash_to_scalars; -use pcs::ntt::AdditiveNttF64; -use primitives::field::{F64, F192, F192BaseUnreduced}; - -use crate::{CODEWORD_SYMBOLS, DA_LOG_K, DaCommitment, LOG_M, row_count}; - -/// Transcript label, so a membership challenge can never be replayed as any other -/// challenge in the stack. -const LABEL: &[u8] = b"leanDA/rs-membership/v1"; - -/// `z_0 … z_{log k - 1}`, bound to the commitment. -/// -/// The matrix root binds both commitment branches. The vector hash is not observed, -/// avoiding a circular dependency between the challenges and the vector. -pub fn membership_challenges(root: &[u8; 32]) -> Vec { - let mut fs = FiatShamirState::from_label(LABEL); - for scalar in hash_to_scalars(root) { - fs.observe(scalar); - } - fs.sample_vec(DA_LOG_K) -} - -/// The test vector deterministically derived from a matrix root. -pub fn membership_vector(root: &[u8; 32]) -> Vec { - dual_codeword(&membership_challenges(root)) -} - -/// BLAKE2s of the vector's three little-endian 64-bit limbs per entry, in domain order. -pub fn vector_digest(vector: &[F192]) -> [u8; 32] { - assert_eq!(vector.len(), CODEWORD_SYMBOLS); - let bytes: Vec = vector - .iter() - .flat_map(|v| [v.c0, v.c1, v.c2].into_iter().flat_map(u64::to_le_bytes)) - .collect(); - primitives::hash::hash(&bytes) -} - -/// `L` over the whole domain: the tensor `⊗_j (1, z_j)` encoded as a codeword. -/// -/// The additive NTT's twiddles are `K`-valued and the transform is `K`-linear, so -/// an `E`-valued transform is three `K`-transforms on the limbs. `F192` is -/// `repr(C)` over three `u64`, which is exactly the interleaved layout -/// [`AdditiveNttF64::encode_interleaved_in_place`] wants for `num_ntts = 3`. -pub fn dual_codeword(z: &[F192]) -> Vec { - assert_eq!(z.len(), DA_LOG_K, "one challenge per novel-basis variable"); - let m = CODEWORD_SYMBOLS; - - let mut buffer = vec![F192::ZERO; m]; - buffer[0] = F192::ONE; - for (j, &zj) in z.iter().enumerate() { - let (low, high) = buffer[..1 << (j + 1)].split_at_mut(1 << j); - for (h, &l) in high.iter_mut().zip(low.iter()) { - *h = l * zj; - } - } - - // SAFETY: `F192` is `repr(C)` over three `u64` and `F64` is `repr(transparent)` - // over `u64`, so the buffer is `3m` limb-interleaved `F64` words with no padding. - let limbs = unsafe { core::slice::from_raw_parts_mut(buffer.as_mut_ptr().cast::(), 3 * m) }; - AdditiveNttF64::standard(LOG_M).encode_interleaved_in_place(limbs, 3, 1); - buffer -} - -/// `⟨L, w_i⟩` for every row, which the caller checks against zero. Rows are `K` -/// valued and `L` is `E` valued, so a term is one `mul_base` (three PMULL), and the -/// whole row accumulates unreduced. -pub fn row_residuals(codewords: &[u64], dual: &[F192]) -> Vec { - let n_rows = row_count(codewords.len(), CODEWORD_SYMBOLS); - assert_eq!(dual.len(), CODEWORD_SYMBOLS, "the dual codeword spans the domain"); - let m = CODEWORD_SYMBOLS; - parallel::map_collect(n_rows, |i| { - let row = &codewords[i * m..(i + 1) * m]; - let mut acc = F192BaseUnreduced::ZERO; - for (&l, &w) in dual.iter().zip(row) { - acc ^= l.mul_base_unreduced(F64(w)); - } - acc.reduce() - }) -} - -/// Draw challenges from the commitment and test every row for membership. -/// The caller must separately bind `codewords` to the commitment. -#[tracing::instrument(name = "RS membership", skip_all)] -pub fn check_membership(commitment: &DaCommitment, codewords: &[u64]) -> bool { - let dual = membership_vector(&commitment.root); - row_residuals(codewords, &dual).iter().all(|&r| r == F192::ZERO) -} - -#[cfg(test)] -mod tests { - use super::*; - use crate::BLOB_SYMBOLS; - use crate::encode_rows; - use rand::{Rng, SeedableRng, rngs::StdRng}; - - fn random_rows(rng: &mut StdRng, n: usize) -> Vec { - (0..n).map(|_| rng.random()).collect() - } - - /// The claim the whole check rests on: over an F_2-subspace domain at rate - /// 1/2 the code is self-dual, so any two codewords are orthogonal. If this - /// fails, `check_membership` is testing the wrong relation. - #[test] - fn code_is_self_dual() { - let mut rng = StdRng::seed_from_u64(3); - for _ in 0..4 { - let a = encode_rows(&random_rows(&mut rng, BLOB_SYMBOLS)); - let b = encode_rows(&random_rows(&mut rng, BLOB_SYMBOLS)); - let dot = a.iter().zip(&b).fold(F64::ZERO, |acc, (&x, &y)| acc + F64(x) * F64(y)); - assert_eq!(dot, F64::ZERO); - } - } - - /// `dual_codeword` runs one interleaved transform over the three `F192` limbs - /// in place, which leans on `F192` being `repr(C)` over three `u64`. A scalar - /// transform of each limb must give the same codeword. - #[test] - fn dual_codeword_matches_limbwise_transform() { - let mut rng = StdRng::seed_from_u64(5); - let z: Vec = (0..DA_LOG_K) - .map(|_| F192::new(rng.random(), rng.random(), rng.random())) - .collect(); - let dual = dual_codeword(&z); - - let mut tensor = vec![F192::ZERO; BLOB_SYMBOLS]; - tensor[0] = F192::ONE; - for (j, &zj) in z.iter().enumerate() { - let (low, high) = tensor[..1 << (j + 1)].split_at_mut(1 << j); - for (h, &l) in high.iter_mut().zip(low.iter()) { - *h = l * zj; - } - } - let ntt = AdditiveNttF64::standard(LOG_M); - for limb in 0..3 { - let mut encoded = vec![F64::ZERO; CODEWORD_SYMBOLS]; - for (out, c) in encoded.iter_mut().zip(&tensor) { - *out = F64([c.c0, c.c1, c.c2][limb]); - } - ntt.forward_transform_scalar(&mut encoded); - for (x, (&got, &want)) in dual.iter().zip(&encoded).enumerate() { - assert_eq!(F64([got.c0, got.c1, got.c2][limb]), want, "limb {limb} at {x}"); - } - } - } - - /// An honestly encoded payload passes, and flipping a single symbol of a single - /// row fails. That single flip is the case the check exists for: it is exactly - /// what a builder would do to make a row unrecoverable while still answering - /// every sample consistently. - #[test] - fn membership_accepts_codewords_and_rejects_a_flipped_symbol() { - let n_rows = 5; - let mut rng = StdRng::seed_from_u64(11); - let rows = random_rows(&mut rng, n_rows * BLOB_SYMBOLS); - let codewords = encode_rows(&rows); - let (commitment, _) = crate::commit_codewords(codewords.clone()); - - assert!(check_membership(&commitment, &codewords)); - - for &position in &[0usize, 1, BLOB_SYMBOLS - 1, BLOB_SYMBOLS, CODEWORD_SYMBOLS - 1] { - let mut corrupted = codewords.clone(); - let index = 3 * CODEWORD_SYMBOLS + position; - corrupted[index] ^= 1; - let (corrupted_commitment, _) = crate::commit_codewords(corrupted.clone()); - assert!( - !check_membership(&corrupted_commitment, &corrupted), - "a flip at {position} passed the membership check" - ); - } - } - - /// A row that is a codeword of the wrong degree (the full domain rather than - /// the degree bound) must fail: rate is what the check enforces, and a - /// rate-1 "codeword" is what an unrecoverable payload looks like. - #[test] - fn membership_rejects_a_full_degree_row() { - let n_rows = 2; - let mut rng = StdRng::seed_from_u64(13); - let rows = random_rows(&mut rng, n_rows * BLOB_SYMBOLS); - let mut codewords = encode_rows(&rows); - let ntt = AdditiveNttF64::standard(LOG_M); - let mut full = random_rows(&mut rng, CODEWORD_SYMBOLS); - ntt.forward_transform_scalar(crate::encode::as_field_mut(&mut full)); - codewords[..CODEWORD_SYMBOLS].copy_from_slice(&full); - let (commitment, _) = crate::commit_codewords(codewords.clone()); - - assert!(!check_membership(&commitment, &codewords)); - } -} diff --git a/crates/rec_aggregation/Cargo.toml b/crates/rec_aggregation/Cargo.toml deleted file mode 100644 index 963b74665..000000000 --- a/crates/rec_aggregation/Cargo.toml +++ /dev/null @@ -1,26 +0,0 @@ -[package] -name = "rec_aggregation" -version.workspace = true -edition.workspace = true -publish = false - -[lints] -workspace = true - -[dependencies] -primitives.workspace = true -parallel.workspace = true -pcs.workspace = true -flock.workspace = true -lean_vm.workspace = true -lean_compiler.workspace = true -xmss.workspace = true -sphincs.workspace = true -lean_da.workspace = true -rand.workspace = true -bincode.workspace = true -serde.workspace = true -tracing.workspace = true - -[dev-dependencies] -zk_alloc.workspace = true diff --git a/crates/rec_aggregation/guests/lean_ethereum.py b/crates/rec_aggregation/guests/lean_ethereum.py deleted file mode 100644 index 677b4d84b..000000000 --- a/crates/rec_aggregation/guests/lean_ethereum.py +++ /dev/null @@ -1,3717 +0,0 @@ -# The recursive aggregation guest: zkDSL, not runnable Python (see -# crates/lean_compiler/zkDSL.md). One node of an aggregation tree verifies raw XMSS -# and SPHINCS+ signatures and sub-proofs OF THIS SAME BYTECODE, and publishes the -# statement digest binding its signature claims and a possibly empty list of LeanDA roots. -# It may check a blob matrix directly and retain any roots proved by its children. -# Reading order: `main` is the -# node, `verify_sub` the in-circuit copy of `lean_vm::cpu::verify`, and -# `open_stacked` the WHIR opening it dispatches into. -# -# Every `*_PLACEHOLDER` below is filled by the host at compile time -# (rec_aggregation::aggregation::placeholder_map), so one source serves every inner -# shape. Names follow doc/leanvm/preamble/macros.tex. -from snark_lib import * - -# ---------------------------------------------------------------- proof stream -# The proof stream rides ONE padded witness hint of 64-bit words, three to a -# transcript scalar (the guest walks only the prefix the shape dictates); binding -# always comes from the per-scalar absorbs. -STREAM_CAP = STREAM_CAP_PLACEHOLDER -MIN_LOG_MEM = MIN_LOG_MEM_PLACEHOLDER -INV_GEN = INV_GEN_PLACEHOLDER - -# ------------------------------------------------------------- the field, GF(2^192) -# A memory cell is one 64-bit word. An element of F192 = F64[Y]/(Y^3+Y+1) is a run -# of three cells, its limbs low first, computed with add192/mul192/div192; heap -# buffers of them have stride three. A product of a word and an element is -# limb-wise, three 64-bit multiplies and no reduction (`scale`). -FIELD_BITS = 192 -BASE_FIELD_BITS = 64 -ONE = f192(1, 0, 0) -ZERO = f192(0, 0, 0) -# Six challenges compose the F2-linear map that batches the 192 transposed -# ring-switch coordinates. -RING_MAP_SHIFTS = [32, 16, 8, 4, 2, 1] -# Exponent bit-widths: an announced 32-bit count decomposes into COUNT_BITS bits, -# its top bit constrained to zero so the native strict 32-bit bound holds; any -# structural size (sums of 2^kappa, packing offsets) fits SIZE_BITS bits; and a -# structural LOG (log_mem, tau_t, log_inv_rate), announced as an integer word and -# raised to a g-power by g_power_of_word, is below SIZE_BITS, so LOG_WORD_BITS bits -# are enough and the reconstruction IS the bound (a larger announced log cannot -# reproduce itself from this many bits). -COUNT_BITS = 33 -SIZE_BITS = 34 -LOG_WORD_BITS = 6 - -# ------------------------------------------------------------------- Fiat-Shamir -# The state is four words. Every absorbed block carries a scalar's three limbs -# and its domain tag in word 3, so a role is never smuggled through the data -# words. The seeding block has none, being fixed at the head of the chain. -DS_OBSERVE = 1 -DS_SQ = 2 -DS_POW_BASE = 3 -DS_POW_NONCE = 4 - -# ----------------------------------------------------------- loop-carried chains -# A loop whose state is several values keeps ONE heap run per iteration, so a -# step is one pointer multiply rather than one per value. Each run leads with the -# four-word Fiat-Shamir state. -FS_SLOTS = 4 -# A Fiat-Shamir chain carrying one accumulator (the claim batching loops). -ACC_VALUE = 4 -ACC_SLOTS = 7 -# One sumcheck round of the bus GKR, the table batch, or flock's multilinear rounds. -ROUND_CURSOR = 4 -ROUND_CLAIM = 5 -ROUND_SLOTS = 8 -# One product layer of the bus GKR. -LAYER_CURSOR = 4 -LAYER_PUSH = 5 -LAYER_PULL = 8 -LAYER_COUNT = 11 -LAYER_LAMBDA = 14 -LAYER_ROW = 17 -LAYER_POS = 18 -LAYER_SLOTS = 19 -# One OOD sample of a WHIR level. -OOD_BETA = 0 -OOD_Y = 3 -OOD_C0 = 6 -OOD_C2 = 9 -OOD_SLOTS = 12 - -# ---------------------------------------------------- the bus: sides and blocks -# GKR sides. The layer counts mu_s are hinted and certified from the block kappas; -# GKR_ROUNDS_CAP caps the per-tree round positions (triangle rounds plus one slot -# per layer) and GKR_POINTS_CAP the point triangle (rows x MU_CAP). -PUSH_SIDE = 0 -PULL_SIDE = 1 -COUNT_SIDE = 2 -N_GKR_SIDES = 3 -GKR_ROUNDS_CAP = GKR_ROUNDS_CAP_PLACEHOLDER -MU_CAP = MU_CAP_PLACEHOLDER -GKR_POINTS_CAP = GKR_POINTS_CAP_PLACEHOLDER -# Bus blocks, flattened across the 3 sides (side s covers blocks -# [SIDE_BLOCK_START[s], SIDE_BLOCK_START[s+1])). The block STRUCTURE is -# protocol-fixed and baked: each block's coord range [BLOCK_COORD_OFF, -# +BLOCK_COORD_COUNT), per coord its COORD_TYPE kind (mirroring leaf.rs::Coord), -# COORD_CONST (the const value, a product's or gcol's g^k, else 0), and the kappa -# SOURCE map (BLOCK_KAPPA_SRC/ADJ: 0 = const adj, 1 = log_mem, 2+t = tau_t). The -# block SHAPES are all reconstructed at runtime from the certified logs: kappa -# directly, the selector bits by pinned advice-decompositions. BLOCK_TABLE names -# the table a block's flush belongs to, or NO_TABLE for the framework blocks -# (boundary, memory seed/finalize, bytecode seed/finalize); it is also what marks a -# block as owned, an owned block's fingerprint being settled by the table sumcheck -# off its table's column evaluations. -COORD_KIND_CONST = 0 -COORD_KIND_COL = 1 -COORD_KIND_GCOL = 2 -COORD_KIND_INDEX = 3 -COORD_KIND_PUBLIC = 4 -COORD_KIND_PROD = 5 -NO_TABLE = NO_TABLE_PLACEHOLDER -SIDE_BLOCK_START = SIDE_BLOCK_START_PLACEHOLDER -N_BLOCKS = N_BLOCKS_PLACEHOLDER -BLOCK_KAPPA_SRC = BLOCK_KAPPA_SRC_PLACEHOLDER -BLOCK_KAPPA_ADJ = BLOCK_KAPPA_ADJ_PLACEHOLDER -BLOCK_TABLE = BLOCK_TABLE_PLACEHOLDER -BLOCK_SIDE = BLOCK_SIDE_PLACEHOLDER -BLOCK_COORD_OFF = BLOCK_COORD_OFF_PLACEHOLDER -BLOCK_COORD_COUNT = BLOCK_COORD_COUNT_PLACEHOLDER -COORD_TYPE = COORD_TYPE_PLACEHOLDER -COORD_CONST = COORD_CONST_PLACEHOLDER -# Claim dedup: push/pull share their GKR point, so a column read by two blocks with -# the same kappa (across OR within the sides) is streamed and opened ONCE. -# COORD_FRESH = 1 on the first occurrence (read the stream, fill pool slot -# COORD_CLAIM_SLOT), 0 on a duplicate (reuse that slot). The count side has its own -# point, so its claims never dedup against the pair's. -COORD_FRESH = COORD_FRESH_PLACEHOLDER -COORD_CLAIM_SLOT = COORD_CLAIM_SLOT_PLACEHOLDER -# A TABLE block's coordinates, flattened into TERMS: coord c is -# Σ_{j < COORD_TERM_COUNT[c]} term(COORD_TERM_OFF[c] + j), each term a TERM_TYPE -# kind over that table's LOCAL column indices TERM_COL_A/TERM_COL_B, scaled by -# TERM_CONST. Those coords raise no claim (the table sumcheck settles them), which -# is what lets one carry a value the row DERIVES from its columns: an XOR/MUL -# result, a DEREF store, a JUMP successor. A framework coord has no terms. -COORD_TERM_OFF = COORD_TERM_OFF_PLACEHOLDER -COORD_TERM_COUNT = COORD_TERM_COUNT_PLACEHOLDER -TERM_TYPE = TERM_TYPE_PLACEHOLDER -TERM_CONST = TERM_CONST_PLACEHOLDER -TERM_COL_A = TERM_COL_A_PLACEHOLDER -TERM_COL_B = TERM_COL_B_PLACEHOLDER -N_BUS_CLAIMS = N_BUS_CLAIMS_PLACEHOLDER -INDEX_MLE_FACTORS = INDEX_MLE_FACTORS_PLACEHOLDER # 1 + g^(2^i) -# Committed-coordinate claims (Col/GCol coords across all sides) and the deferred -# bytecode values (Public coords). -N_CLAIMS = N_CLAIMS_PLACEHOLDER -# A bus tuple's coordinates index the 2^N_TUPLE_BITS fingerprint slots (doc -# sec:gp). The stacked bytecode has BYTECODE_COLS encoding columns, stacked along -# LOG2_BYTECODE_COLS selector bits into ONE multilinear; push and pull share their -# GKR point, so the columns are opened ONCE. -N_TUPLE_BITS = 4 -N_TUPLE_SLOTS = 16 -BYTECODE_COLS = BYTECODE_COLS_PLACEHOLDER -LOG2_BYTECODE_COLS = LOG2_BYTECODE_COLS_PLACEHOLDER - -# ------------------------------------------------------------- the eight tables -# The table sumcheck's batch carries EVERY committed column of a table, because its -# bus forms read the flushed ones and its constraint the rest; TABLE_COLS_CAP caps -# the evaluation frame. ETA_OFFSET[t] starts table t's disjoint range of zc_xi -# powers; the three bus forms take ETA_FORM_BASE + side, the SAME three powers for -# every table, and that sharing is what makes the batch's target derivable from the -# three leaf claims. FLOORS[t] is the table's tau floor (BLAKE2s is sized to -# flock's instance count, >= 2^3). -TABLE_XOR64 = 0 -TABLE_MUL64 = 1 -TABLE_SET = 2 -TABLE_DEREF = 3 -TABLE_JUMP = 4 -TABLE_BLAKE2s = 5 -TABLE_XOR192 = 6 -TABLE_MUL192 = 7 -N_TABLES = N_TABLES_PLACEHOLDER -FLOORS = [0, 0, 0, 0, 0, 3, 0, 0] -N_TABLE_COLS = N_TABLE_COLS_PLACEHOLDER -TABLE_COLS_CAP = TABLE_COLS_CAP_PLACEHOLDER -ETA_OFFSET = ETA_OFFSET_PLACEHOLDER -ETA_FORM_BASE = ETA_FORM_BASE_PLACEHOLDER -N_ETA_POWS = N_ETA_POWS_PLACEHOLDER - -# ------------------------------------------------------------ flock (the R1CS) -# Univariate skip: K_SKIP variables fold in one skip round (half-domain 2^K_SKIP -# phi8 nodes), then N_FIXED_CHALLENGE_ROUNDS fixed inner rounds (FIXED_CHALLENGES), -# then sampled outer rounds. LAGRANGE_INV_* are the one baked inverse barycentric -# denominator per domain (combined, S). The zerocheck point/round buffers are sized -# at runtime in the exponent (m = K_LOG + tau_5 and m - 6, both certified); -# LINCHECK_ROUNDS = K_LOG - K_SKIP is protocol-fixed and PIN_COLUMN is the -# const-pin column. FIXED_CHALLENGES and PHI8_NODES hold three words an element. -K_SKIP = K_SKIP_PLACEHOLDER -N_FIXED_CHALLENGE_ROUNDS = N_FIXED_CHALLENGE_ROUNDS_PLACEHOLDER -FIXED_CHALLENGES = FIXED_CHALLENGES_PLACEHOLDER -PHI8_NODES = PHI8_NODES_PLACEHOLDER -LAGRANGE_INV_COMBINED = LAGRANGE_INV_COMBINED_PLACEHOLDER -LAGRANGE_INV_S = LAGRANGE_INV_S_PLACEHOLDER -LINCHECK_ROUNDS = LINCHECK_ROUNDS_PLACEHOLDER -PIN_COLUMN = PIN_COLUMN_PLACEHOLDER -K_LOG = K_LOG_PLACEHOLDER -SLOT_STRIDE_LOG = SLOT_STRIDE_LOG_PLACEHOLDER # = K_LOG - LOG_PACKING (=8); the q_flock slot stride - -# ------------------------------------------------- the stacked WHIR opening -# The opening is dispatched by the certified committed log-size m through `match`. -# The LIG_* tables carry one row per (rate, m), emitted from the same -# derive_profile/level_shapes the prover uses: scalars index as TBL[m_idx], -# per-level values as TBL[m_idx * LIG_MAX_LEVELS + lvl] where m_idx is the -# flattened rate-major configuration index, and the subspace vanishing constants -# with the LIG_MAX_VANISH_LEN stride. -LIG_MIN_LOG_SIZE = LIG_MIN_LOG_SIZE_PLACEHOLDER -LIG_N_LOG_SIZES = LIG_N_LOG_SIZES_PLACEHOLDER -LIG_N_RATES = LIG_N_RATES_PLACEHOLDER -# Committed-column kappa sources (0 = const COL_KAPPA_ADJ, 1 = log_mem, 2+t = tau_t) -# and the PCS floor for the stacked size. -N_COMMITTED_COLS = N_COMMITTED_COLS_PLACEHOLDER -N_COLUMN_LOGS = N_COLUMN_LOGS_PLACEHOLDER -COL_KAPPA_SRC = COL_KAPPA_SRC_PLACEHOLDER -COL_KAPPA_ADJ = COL_KAPPA_ADJ_PLACEHOLDER -PCS_MIN_MU = PCS_MIN_MU_PLACEHOLDER -# Global maxima; StackBuf frame sizes are parse-time, so they must be baked. -LIG_MAX_LEVELS = LIG_MAX_LEVELS_PLACEHOLDER -LIG_MAX_VANISH_LEN = LIG_MAX_VANISH_LEN_PLACEHOLDER -LIG_MAX_OOD_SAMPLES = LIG_MAX_OOD_SAMPLES_PLACEHOLDER -LIG_LOG_MSG_COLS_CAP = LIG_LOG_MSG_COLS_CAP_PLACEHOLDER -YR_LOG_CAP = YR_LOG_CAP_PLACEHOLDER -MAX_STACK_LOG = LIG_MIN_LOG_SIZE + LIG_N_LOG_SIZES - 1 -LIG_N_LEVELS = LIG_N_LEVELS_PLACEHOLDER -LIG_YR_LEVEL = LIG_YR_LEVEL_PLACEHOLDER -LIG_YR_LOG_LEN = LIG_YR_LOG_LEN_PLACEHOLDER -LIG_YR_LEN = LIG_YR_LEN_PLACEHOLDER -LIG_TOTAL_FOLDS = LIG_TOTAL_FOLDS_PLACEHOLDER -LIG_MAX_QUERIES = LIG_MAX_QUERIES_PLACEHOLDER -LIG_MAX_SQUEEZES = LIG_MAX_SQUEEZES_PLACEHOLDER -LIG_MAX_INTERLEAVE = LIG_MAX_INTERLEAVE_PLACEHOLDER -LIG_POSITIONS_LEN = LIG_POSITIONS_LEN_PLACEHOLDER -LIG_CAP_DEPTH = LIG_CAP_DEPTH_PLACEHOLDER -LIG_CAP_OFF = LIG_CAP_OFF_PLACEHOLDER -LIG_CAP_LEN = LIG_CAP_LEN_PLACEHOLDER -LIG_QUERY_GRIND_BITS = LIG_QUERY_GRIND_BITS_PLACEHOLDER -LIG_OOD_SAMPLES = LIG_OOD_SAMPLES_PLACEHOLDER -LIG_QUERIES = LIG_QUERIES_PLACEHOLDER -LIG_FOLDS = LIG_FOLDS_PLACEHOLDER -LIG_INTERLEAVE = LIG_INTERLEAVE_PLACEHOLDER -LIG_LEAF_BLOCKS = LIG_LEAF_BLOCKS_PLACEHOLDER -LIG_ROW_CAP = LIG_ROW_CAP_PLACEHOLDER -LIG_PATH_CAP = LIG_PATH_CAP_PLACEHOLDER -LIG_TREE_DEPTH = LIG_TREE_DEPTH_PLACEHOLDER -LIG_SQUEEZES = LIG_SQUEEZES_PLACEHOLDER -LIG_POSITIONS_OFF = LIG_POSITIONS_OFF_PLACEHOLDER -LIG_LOG_MSG_COLS = LIG_LOG_MSG_COLS_PLACEHOLDER -LIG_RESIDUAL_FOLD_OFF = LIG_RESIDUAL_FOLD_OFF_PLACEHOLDER -LIG_RESIDUAL_PREFIX_LEN = LIG_RESIDUAL_PREFIX_LEN_PLACEHOLDER -LIG_FOLDS_OFF = LIG_FOLDS_OFF_PLACEHOLDER -LIG_VANISH_OFF = LIG_VANISH_OFF_PLACEHOLDER -LIG_VANISH_VALS = LIG_VANISH_VALS_PLACEHOLDER # words: the novel basis lives in F64 -LIG_VANISH_INVS = LIG_VANISH_INVS_PLACEHOLDER -LIG_N_CANDIDATES = LIG_N_CANDIDATES_PLACEHOLDER -LIG_MIN_SHIFT_INV = LIG_MIN_SHIFT_INV_PLACEHOLDER -# eval_b claim descriptors. CLAIM_POINT_BUF says which point buffer a pooled -# claim's x-part lives in, CLAIM_COMMITTED_COL maps it to the compact index of the -# committed column it must open (a virtual BLAKE2s value claim maps to QFLOCK), -# CLAIM_QFLOCK_SLOT_BITS holds the fixed packed-slot bits of every logical claim -# (zero for a non-virtual one), and QFLOCK_COMMITTED_COL is the ring-switch target. -POINT_BUF_ZETA = 0 -POINT_BUF_RHO = 1 -POINT_BUF_PI = 2 -POINT_BUF_QFLOCK_RHO = 3 -CLAIM_POINT_BUF = CLAIM_POINT_BUF_PLACEHOLDER -CLAIM_COMMITTED_COL = CLAIM_COMMITTED_COL_PLACEHOLDER -CLAIM_QFLOCK_SLOT_BITS = CLAIM_QFLOCK_SLOT_BITS_PLACEHOLDER -QFLOCK_COMMITTED_COL = QFLOCK_COMMITTED_COL_PLACEHOLDER -QFLOCK_VARS_CAP = QFLOCK_VARS_CAP_PLACEHOLDER - -# ------------------------------------------------------ statements and deferral -# A node defers three claims on fixed polynomials. DEFER_SIZE is the region one -# sub-proof's verification exports, in field elements (a bytecode point plus the -# flock lincheck data, see verify_sub's defer_out layout); DEFER_STMT_* index the -# batched claims a node's OWN statement carries. The Fiat-Shamir seed rides the -# statement rather than being baked, so one compiled guest verifies proofs of any -# inner program. -BYTECODE_LOG = BYTECODE_LOG_PLACEHOLDER # log rows of the bytecode blocks -DEFER_SIZE = DEFER_SIZE_PLACEHOLDER -BYTECODE_VARS = BYTECODE_VARS_PLACEHOLDER # = BYTECODE_LOG + LOG2_BYTECODE_COLS -# The exported record's layout: the shared bytecode point, then the flock data. -FRESH_BC_VALUE = BYTECODE_VARS -FRESH_ALPHA = BYTECODE_VARS + 1 -FRESH_Z_SKIP = BYTECODE_VARS + 2 -FRESH_ZCHI = BYTECODE_VARS + 3 -FRESH_LINCHECK_RS = FRESH_ZCHI + LINCHECK_ROUNDS -FRESH_Z_PARTIAL = FRESH_LINCHECK_RS + LINCHECK_ROUNDS -FRESH_MATPART = FRESH_Z_PARTIAL + 2 ** K_SKIP -DEFER_STMT_CELLS = BYTECODE_VARS + 1 + 2 * K_LOG + 2 -DEFER_STMT_BC_VALUE = BYTECODE_VARS -DEFER_STMT_MAT_POINT = BYTECODE_VARS + 1 -DEFER_STMT_A_VALUE = BYTECODE_VARS + 1 + 2 * K_LOG -DEFER_STMT_B_VALUE = BYTECODE_VARS + 2 + 2 * K_LOG -AGG_SEED_0 = AGG_SEED_0_PLACEHOLDER -AGG_SEED_1 = AGG_SEED_1_PLACEHOLDER -AGG_SEED_2 = AGG_SEED_2_PLACEHOLDER -AGG_SEED_3 = AGG_SEED_3_PLACEHOLDER -# The statement digest's preimage: the STMT_HEADER header words (the seed, the -# signer-set digest, which itself binds the epoch groups and every count, and the -# DA root-list digest), then the deferred elements' three limbs each, zero-filled -# to whole blocks. No domain tag: the seed leads, and it binds this bytecode and -# flock's R1CS. -STMT_HEADER = STMT_HEADER_PLACEHOLDER -STMT_PAD = STMT_PAD_PLACEHOLDER -STMT_BLOCKS = STMT_BLOCKS_PLACEHOLDER -# The declared lists are hashed with plain BLAKE2s over a flat run of words, 64 -# bytes a compression. A block's byte counter is a runtime value and the ISA has no -# integer addition, so it splits as in doc §sec:prog-byte-counter: a window of -# SIGNERS_WINDOW blocks shares one base 64·SIGNERS_WINDOW·q, whose set bits all sit -# above the window's own offsets 64(j+1), so a block's counter word is one XOR. The -# base comes from the window loop's own counter, and the one block whose offset -# overlaps it takes the next window's base instead. -SIGNERS_WINDOW = SIGNERS_WINDOW_PLACEHOLDER -SIGNERS_WINDOW_LOG = SIGNERS_WINDOW_LOG_PLACEHOLDER -SIGNERS_MAX_WINDOWS = SIGNERS_MAX_WINDOWS_PLACEHOLDER -SIGNERS_COUNT_BITS = SIGNERS_COUNT_BITS_PLACEHOLDER -# The lists a hash tail walks (`tail_blocks`). -LIST_PLAIN = 0 -LIST_KEYS = 1 -LIST_SPHINCS = 2 -LIST_CHILD_KEYS = 3 -LIST_CHILD_SPHINCS = 4 -# BLAKE2s's parameterized initial chaining value, which every hash here starts from, -# and the flag word of a final block's metadata. -BLAKE2S_IV_0 = BLAKE2S_IV_0_PLACEHOLDER -BLAKE2S_IV_1 = BLAKE2S_IV_1_PLACEHOLDER -BLAKE2S_IV_2 = BLAKE2S_IV_2_PLACEHOLDER -BLAKE2S_IV_3 = BLAKE2S_IV_3_PLACEHOLDER -MD_FINAL = MD_FINAL_PLACEHOLDER - -# ---------------------------------------------------------- XMSS (host-supplied) -# Every 16-byte native value (tweak, digest, chain tip, sibling, public parameter) -# is two words, a 32-byte one (a message, a key) four. A tweak's first word holds -# its type and sub-position, its second the index at bit 32 (`xmss::make_tweak`): -# XM_* are the first words, as the host read them out of `make_tweak`, and -# XM_INDEX_WEIGHT[b] is what bit b of an index weighs in the second, an index being -# its set bits summed. A sub-position weighs XM_P_MUL a unit. -V = V_PLACEHOLDER -W = W_PLACEHOLDER -TARGET_SUM = TARGET_SUM_PLACEHOLDER -LOG_LIFETIME = LOG_LIFETIME_PLACEHOLDER -CHAIN_LENGTH = 2 ** W -CHAIN_STEPS = CHAIN_LENGTH - 1 -XM_ENC_TWEAK = XM_ENC_TWEAK_PLACEHOLDER -XM_PK_TWEAK = XM_PK_TWEAK_PLACEHOLDER -XM_CHAIN_TWEAK = XM_CHAIN_TWEAK_PLACEHOLDER -XM_MERKLE_TWEAK = XM_MERKLE_TWEAK_PLACEHOLDER -XM_INDEX_WEIGHT = XM_INDEX_WEIGHT_PLACEHOLDER -XM_P_MUL = XM_P_MUL_PLACEHOLDER -# Digits packed per digest word: W bits each in GF(2^64)'s monomial budget (the -# word's leftover top bits are ground to zero by the signer). -DIGITS_PER_WORD = V / 2 -TIP_WORDS = 2 * V -WOTS_PK_BLOCKS = (2 + V) / 4 # prefix (tweak, pp) + V tips, eight words a block - -# ------------------------------------------------------ SPHINCS+ (host-supplied) -# The scheme's own letters, prefixed SP_ where XMSS has the same one. -SP_V = SP_V_PLACEHOLDER -SP_W = SP_W_PLACEHOLDER -SP_TARGET_SUM = SP_TARGET_SUM_PLACEHOLDER -SP_D = SP_D_PLACEHOLDER -SP_HEIGHTS = SP_HEIGHTS_PLACEHOLDER # h_lay, one per hypertree layer, top first -SP_SUFFIX = SP_SUFFIX_PLACEHOLDER # SP_SUFFIX[lay] = sum of h_j for j >= lay -SP_A = SP_A_PLACEHOLDER -SP_K = SP_K_PLACEHOLDER -SP_H = SP_H_PLACEHOLDER # the total hypertree height, SP_SUFFIX[0] -SP_CHAIN_LENGTH = 2 ** SP_W -SP_CHAIN_STEPS = SP_CHAIN_LENGTH - 1 -SP_DIGITS_PER_WORD = SP_V / 2 -SP_TIP_WORDS = 2 * SP_V -SP_LEAF_BLOCKS = (2 + SP_V) / 4 # prefix (tweak, pp) + V tips, eight words a block -SP_N_FTS = SP_K - 1 # the forest drops the last index's tree -SP_ROOT_BLOCKS = (2 + SP_N_FTS) / 4 -# The message digest is h + k*a bits of a BLAKE2s output, within its first three -# words, so the bit buffer holds three words' decompositions. -SP_BIT_LANES = 3 -SP_BIT_CELLS = SP_BIT_LANES * BASE_FIELD_BITS -# Native tweak first words, including the protocol domain separator and type. -SP_TW_CHAIN = SP_TW_CHAIN_PLACEHOLDER -SP_TW_LEAF = SP_TW_LEAF_PLACEHOLDER -SP_TW_NODE = SP_TW_NODE_PLACEHOLDER -SP_TW_ENC = SP_TW_ENC_PLACEHOLDER -SP_TW_FTS_LEAF = SP_TW_FTS_LEAF_PLACEHOLDER -SP_TW_FTS_NODE = SP_TW_FTS_NODE_PLACEHOLDER -SP_TW_FTS_ROOTS = SP_TW_FTS_ROOTS_PLACEHOLDER -SP_TW_MSG = SP_TW_MSG_PLACEHOLDER -# Tweak layout: protocol_domain_sep | type | layer | zero | p in the first word, -# tree | index in the second, each 32-bit field within one word. -SP_LAY_MUL = 2 ** 16 -SP_P_MUL = 2 ** 32 -SP_TAU_POS = 0 -SP_J_POS = 32 -SP_CHAIN_MUL = SP_CHAIN_LENGTH * SP_P_MUL # chain i's tweaks start at p = 2^w * i -# The encoding counter, LE_32 in the low four bytes of its word: bounded by -# decomposing exactly that many bits, so the guest accepts no preimage the native -# verifier cannot parse. -SP_COUNTER_BITS = 32 - -# --------------------------------------------------------------- node capacities -# MAX_KEYS caps the coverage table's slots, both schemes' declared keys and their -# duplicates, which is what the coverage range check needs below 2^MIN_LOG_MEM; -# MAX_RECURSIONS is the arity of an aggregation tree; MAX_EPOCHS caps the runtime -# number of XMSS epoch groups. -MAX_KEYS = MAX_KEYS_PLACEHOLDER -MAX_DA_ROOTS = MAX_DA_ROOTS_PLACEHOLDER -DA_ROOT_COUNTS = DA_ROOT_COUNTS_PLACEHOLDER -MAX_RECURSIONS = MAX_RECURSIONS_PLACEHOLDER -MAX_EPOCHS = MAX_EPOCHS_PLACEHOLDER - - -# ---------------------------------- LeanDA ------------------------------------------ -# Blob and cell widths are fixed. The row count is hinted and bounded by DA_MAX_ROWS; -# the Merkle trees dispatch on the checked log of its padded value. -DA_LOG_K = DA_LOG_K_PLACEHOLDER -DA_LOG_CELL = DA_LOG_CELL_PLACEHOLDER -DA_MAX_ROWS = DA_MAX_ROWS_PLACEHOLDER -DA_LOG_MAX_ROWS = DA_LOG_MAX_ROWS_PLACEHOLDER -DA_PAD_CELL = DA_PAD_CELL_PLACEHOLDER -DA_PAD_ROW = DA_PAD_ROW_PLACEHOLDER - -DA_CELL = 2 ** DA_LOG_CELL # symbols in a cell -DA_BLOCK_BITS = DA_LOG_K + 1 - DA_LOG_CELL # log of the cells per row -DA_CELLS = 2 ** DA_BLOCK_BITS # cells per row -DA_PREFIX_CELLS = 2 ** (DA_LOG_K - DA_LOG_CELL) # cells in the first half -DA_CELL_BLOCKS = DA_CELL // 8 # BLAKE2s blocks in one cell -DA_ROW_BLOCKS = DA_PREFIX_CELLS // 2 # BLAKE2s blocks in one row digest -DA_TREE_ARMS = DA_LOG_MAX_ROWS + 1 # tree depths the row count can dispatch to - -# =================================== field helpers ================================== - - -@inline -def scale(s, x): - # A word times an element, limb by limb. - return [s * x[0], s * x[1], s * x[2]] - - -@inline -def add64(x, s): - # An element plus a word: only the low limb moves. - return [x[0] + s, x[1], x[2]] - - -@inline -def fixed_challenge(i: Const): - return f192(FIXED_CHALLENGES[3 * i], FIXED_CHALLENGES[3 * i + 1], FIXED_CHALLENGES[3 * i + 2]) - - -@inline -def phi8(i: Const): - return f192(PHI8_NODES[3 * i], PHI8_NODES[3 * i + 1], PHI8_NODES[3 * i + 2]) - - -# ==================================== Fiat-Shamir =================================== - - -@inline -def fs_next(state, cursor): - # Fetch, observe and advance in one act: read the scalar under `cursor`, fold - # it into the state, and hand back the successor state, the scalar, AND the - # cursor stepped one scalar on. Reading and absorbing are inseparable here, so - # no proof-stream word can enter the computation unbound: the soundness - # invariant the whole guest rests on. All three returns alias into the caller. - block = StackBuf(4) - block[0:3] = cursor[0:3] - block[3] = DS_OBSERVE - nb = StackBuf(4) - blake2s(state, block, nb) - return nb, block[0:3], cursor * GEN ** 3 - - -@inline -def fs_next_half(state, cursor): - # `fs_next` of a 128-bit half of a Merkle root, which rides the stream as a - # scalar with a zero top limb (merkle.rs `scalars_to_hash`). - block = StackBuf(4) - block[0:3] = cursor[0:3] - block[3] = DS_OBSERVE - assert block[2] == 0 - nb = StackBuf(4) - blake2s(state, block, nb) - return nb, block[0:3], cursor * GEN ** 3 - - -@inline -def obs(state, x): - # Bind one scalar into the chain: state <- compress(state, (x, DS_OBSERVE)). - nb = StackBuf(4) - blake2s(state, [x, DS_OBSERVE], nb) - return nb - - -@inline -def obs_at(state, ptr): - # `obs` of the scalar at a heap pointer, read straight into the block. - block = StackBuf(4) - block[0:3] = ptr[0:3] - block[3] = DS_OBSERVE - nb = StackBuf(4) - blake2s(state, block, nb) - return nb - - -@inline -def squeeze(state): - # Ratchet: the digest is the new state, its first three words the challenge. - nb = StackBuf(4) - blake2s(state, [0, 0, 0, DS_SQ], nb) - return nb, nb[0:3] - - -# ============================ bits, logs, and the exponent ========================== - - -@inline -def bind_bits(bits_ptr, value, n: Const): - # Tie an advice bit run back to the word it decomposes. Booleanity is a - # write-once pin: the cell already holds the bit, so storing its square IS the - # assert, one instruction shorter than a separate equality. - acc = 0 - for i in unroll(0, n): - b = bits_ptr[GEN ** i] - bits_ptr[GEN ** i] = b * b - acc += b * 2 ** i - assert acc == value - return - - -def exponent_tables(): - # Read-only lookup tables over the exponent domain, indexed at runtime g-powers - # (so they must be heap, not stack): g_logs_pow2[g^j] = 2^j raises a g-power's - # log, and g_squares[g^j] = g^(2^j) turns integer sums of powers of two into - # field products. Both span SIZE_BITS because verify_log2_ceil bounds its result - # there, so g_log reaches g^(SIZE_BITS-1) and indexes g_logs_pow2 at it; sizing - # to COUNT_BITS would leave that lookup reading a prover-chosen cell. - g_logs_pow2 = HeapBuf(SIZE_BITS) - for j in unroll(0, SIZE_BITS): - g_logs_pow2[GEN ** j] = 2 ** j - g_squares = HeapBuf(SIZE_BITS) - sq_run = GEN - for j in unroll(0, SIZE_BITS): - g_squares[GEN ** j] = sq_run - sq_run *= sq_run - return g_logs_pow2, g_squares - - -def g_power_of_word(value, g_squares, nbits: Const): - # g^value for a concrete integer `value` < 2^nbits: advice-decompose its bits, - # tie them back to the word, and assemble Π g^(bit_j·2^j). - bits = HeapBuf(GEN ** nbits) - hint_decompose_bits(bits, value, nbits) - word = 0 - g_value = GEN ** 0 - for j in unroll(0, nbits): - bit = bits[GEN ** j] - assert bit * bit == bit - word += bit * (2 ** j) - g_value *= (1 + bit * (g_squares[GEN ** j] + 1)) - assert word == value - return g_value - - -def verify_log2_ceil(bits_buf, g_logs_pow2, g_squares, floor: Const, nbits: Const): - # Given `nbits` bits already in bits_buf, return (g_log, exp_prod) for - # word = Σ bit_j 2^j: exp_prod = g^word and g_log = g^max(log2_ceil(word), - # floor). g_log is prover advice, pinned to log2_ceil(word) by psum[g_log] == - # word (word < 2^log, the == 2^log case via g_logs_pow2) and word > 2^(log-1) - # (waived at floor). Callers fill the bits and tie word or exp_prod to their - # value. NB: log2 here is the base-2 log of the integer word, not the discrete - # log base g that `log(...)` means. - psum_buf = HeapBuf(SIZE_BITS + 1) # psum_buf[g^j] = value of bits [0, j) - psum_buf[GEN ** 0] = 0 - word = 0 - exp_prod = GEN ** 0 - for j in unroll(0, nbits): - bit = bits_buf[GEN ** j] - assert bit * bit == bit - exp_prod *= (1 + bit * (g_squares[GEN ** j] + 1)) - word += bit * (2 ** j) - psum_buf[GEN ** (j + 1)] = word - for j in unroll(nbits + 1, SIZE_BITS + 1): - psum_buf[GEN ** j] = word - g_log = hint_log2_ceil(bits_buf, nbits, floor) # prover advice; verified below - assert log(g_log) < SIZE_BITS - assert log(g_log / (GEN ** floor)) < SIZE_BITS - low_bits = psum_buf[g_log] # value of bits [0, log) - high_bits = low_bits + word # value of bits [log, nbits) - assert high_bits * low_bits == 0 # word < 2^log (high bits clear) OR ... - assert high_bits * (word + g_logs_pow2[g_log]) == 0 # ... word == 2^log - if g_log != GEN ** floor: - # minimality (word > 2^(log-1)); skip at g_log == g^0, where word is in - # {0,1}, its ceil-log 0 already minimal and psum_buf[g^-1] out of range. - if g_log != GEN ** 0: - low_bits_prev = psum_buf[g_log * INV_GEN] # bits [0, log-1) - word_vs_2logprev = word + g_logs_pow2[g_log * INV_GEN] # 0 iff word == 2^(log-1) - assert (low_bits_prev + word) * word_vs_2logprev != 0 - return g_log, exp_prod - - -def log2_ceil_in_the_exponent(g_N, g_logs_pow2, g_squares, floor: Const, nbits: Const): - # g^log2_ceil(N) given g_N = g^N (N < 2^nbits). There is no in-circuit log, so - # the prover hints N's bits; they are verified and tied back, the value they - # decode to having to equal g_N. - bits = HeapBuf(GEN ** nbits) - hint_decompose_bits_exponent(bits, g_N, nbits) - g_log, g_bits_value = verify_log2_ceil(bits, g_logs_pow2, g_squares, floor, nbits) - assert g_bits_value == g_N - return g_log - - -def decode_query_bits(squeezed: StackBuf(3), positions_out, bit_ptrs_out, depth: Const): - # The squeezed challenge's 192 bits are advice-decomposed HERE, a limb at a time, - # boolean-constrained and tied back by reconstruction; each depth-bit group also - # becomes a query position (little-endian), with a pointer to its bit run (the - # Merkle direction bits). The bits live in FRAME cells, every index into them - # being compile-time, and `addr` names the run so the direction-bit pointers - # still reach it. - per_word = FIELD_BITS // depth - bits = StackBuf(FIELD_BITS) - bits_ptr = addr(bits) - hint_decompose_bits(bits, squeezed[0], BASE_FIELD_BITS) - hint_decompose_bits(bits_ptr * GEN ** 64, squeezed[1], BASE_FIELD_BITS) - hint_decompose_bits(bits_ptr * GEN ** 128, squeezed[2], BASE_FIELD_BITS) - acc0 = 0 - acc1 = 0 - acc2 = 0 - for j in unroll(0, per_word): - base_bit = j * depth - lane = base_bit // 64 - shift = base_bit % 64 - # A group inside one limb is one run of bits; a group straddling a limb - # boundary splits into the two runs that do stay inside one. `b // cut == 0` - # IS `b < cut`, the DSL's `if` comparing for equality only. - cut = 64 - shift # bits of this group below the next limb - p_lo = 0 - p_hi = 0 - for b in unroll(0, depth): - t = bits[base_bit + b] - bits[base_bit + b] = t * t # booleanity, as a write-once pin - if b // cut == 0: - p_lo += t * 2 ** b - else: - p_hi += t * 2 ** (b - cut) - if const(lane == 0): - acc0 += p_lo * 2 ** shift - if const(lane == 1): - acc1 += p_lo * 2 ** shift - if const(lane == 2): - acc2 += p_lo * 2 ** shift - # position = p_lo + 2^cut * p_hi: multiplying by X^cut concatenates the two - # runs, both degrees staying below 64. - if cut // depth == 0: # `cut < depth`: this group straddles the boundary - positions_out[GEN ** j] = p_lo + p_hi * 2 ** cut - if const(lane == 0): - acc1 += p_hi - if const(lane == 1): - acc2 += p_hi - else: - positions_out[GEN ** j] = p_lo - bit_ptrs_out[GEN ** j] = bits_ptr * GEN ** base_bit - for i in unroll(per_word * depth, FIELD_BITS): - t = bits[i] - bits[i] = t * t - if const(i // 64 == 0): - acc0 += t * 2 ** (i % 64) - if const(i // 64 == 1): - acc1 += t * 2 ** (i % 64) - if const(i // 64 == 2): - acc2 += t * 2 ** (i % 64) - assert acc0 == squeezed[0] - assert acc1 == squeezed[1] - assert acc2 == squeezed[2] - return - - -def grind_check(state: StackBuf(4), nonce: StackBuf(3), nbits_g): - # WHIR fold/query grinding: digest = H(H(state, POW_BASE), (nonce, POW_NONCE)), - # whose low nbits (nbits_g = g^nbits) must be zero. The PoW window of - # transcript::pow_bits_ok is `digest.0 & ((1 << bits) - 1)` with nbits < 64, so - # it lives entirely in the digest's first word: only that word is - # advice-decomposed and verified. The caller absorbs the full field nonce - # afterwards. The honest prover searches the deterministic u64 subset while - # verification permits the full field: each candidate still costs one hash and - # succeeds with probability 2^-bits. - if nbits_g == GEN ** 0: - assert_eq192(nonce, ZERO) # native canonical zero-work nonce - base = StackBuf(4) - blake2s(state, [0, 0, 0, DS_POW_BASE], base) - out = StackBuf(4) - blake2s(base, [nonce, DS_POW_NONCE], out) - # Frame cells for the unrolled pass (no DEREF per bit), named by `addr` for the - # zero-check walk, whose bound is runtime and so must index a pointer. - word_bits = StackBuf(BASE_FIELD_BITS) - hint_decompose_bits(word_bits, out[0], BASE_FIELD_BITS) - word_ptr = addr(word_bits) - bind_bits(word_ptr, out[0], BASE_FIELD_BITS) - for xb in mul_range(1, nbits_g): - assert word_ptr[xb] == 0 - return - - -# =============================== multilinear primitives ============================= - - -@inline -def eq_weight(ch, count: Const, idx: Const, msb_span: Const): - # The eq-tensor weight of compile-time index `idx` against the challenge run - # ch[0..count), three words an element: prod_c eq(bit(idx), ch[c]), where the - # bit is bit c of idx (msb_span == 0) or bit (msb_span - 1 - c) (an MSB-first - # walk over an msb_span-bit index). - w = ONE - for c in unroll(0, count): - k = 3 * c - cv = ch[k:k + 3] - if msb_span == 0: - bit = (idx // (2 ** c)) % 2 - else: - bit = (idx // (2 ** (msb_span - 1 - c))) % 2 - if bit == 1: - w = mul192(w, cv) - else: - w = mul192(w, add192(ONE, cv)) - return w - - -@inline -def eqtree(point_ptr, out, n_coords: Const): - # The eq tensor of the n_coords challenges at point_ptr[0..n_coords), built by - # doubling into out (size 2^(n_coords+1) - 2 elements); the final 2^n_coords - # values start at element 2^n_coords - 2. - r0 = point_ptr[0:3] - out[0:3] = add192(ONE, r0) - out[3:6] = r0 - for t in unroll(1, n_coords): - k = 3 * t - rt = point_ptr[k:k + 3] - one_plus_rt = add192(ONE, rt) - for i in unroll(0, 2 ** t): - src = 3 * (2 ** t - 2 + i) - lo = 3 * (2 ** (t + 1) - 2 + i) - hi = 3 * (2 ** (t + 1) - 2 + 2 ** t + i) - pw = out[src:src + 3] - out[lo:lo + 3] = mul192(pw, one_plus_rt) - out[hi:hi + 3] = mul192(pw, rt) - return - - -@inline -def lag64(z, out, node_base: Const): - # The 64 phi8-domain Lagrange NUMERATORS at z over nodes - # PHI8_NODES[node_base .. node_base + 64]: out[i] = prod_{j != i} (z + - # PHI8_NODES[node_base + j]), three cells an element. Every barycentric - # denominator over an aligned phi8 window is the same element, so callers scale - # the finished sum once by LAGRANGE_INV_S / LAGRANGE_INV_COMBINED instead of the - # numerators one by one. - pre = StackBuf(3 * 65) - pre[0:3] = ONE - for i in unroll(0, 64): - k = 3 * i - pre[k + 3:k + 6] = mul192(pre[k:k + 3], add192(z, phi8(node_base + i))) - suf = StackBuf(3 * 65) - suf[192:195] = ONE - for i in unroll(0, 64): - k = 3 * (63 - i) - suf[k:k + 3] = mul192(suf[k + 3:k + 6], add192(z, phi8(node_base + 63 - i))) - for i in unroll(0, 64): - k = 3 * i - out[k:k + 3] = mul192(pre[k:k + 3], suf[k + 3:k + 6]) - return - - -def eq_prefix_chain(chain, seed: StackBuf(3), a, b, count_g): - # Prefix products of eq(a_k, b_k) = 1 + a_k + b_k from `seed`, so a reader picks - # the partial product up at its own certified length. Entry t is written from - # inputs with index < t only, so a garbage tail past a buffer's written extent - # cannot corrupt any shorter prefix. Every buffer holds three words an element. - chain[0:3] = seed - for xk in mul_range(1, count_g): - x3 = xk ** 3 - nxt = x3 * GEN ** 3 - chain[nxt:nxt + 3] = mul192(chain[x3:x3 + 3], add192(ONE, add192(a[x3:x3 + 3], b[x3:x3 + 3]))) - return - - -def rs_eq_run(chain, z_vals, point, count_g): - # One run of the telescoped ring-switch product E = sum_k c_k * prod_j - # (z_j^(2^k) + 1 + ris_j): coordinate x multiplies row k by (z^(2^k) + 1 + - # point_x), z evolving by squaring per row. The runtime coordinates walk - # OUTSIDE and the fixed Frobenius powers inside, so nothing stores a z-power - # table. A row is BASE_FIELD_BITS elements. - for xk in mul_range(1, count_g): - x3 = xk ** 3 - zv = z_vals[x3:x3 + 3] - one_plus = add192(ONE, point[x3:x3 + 3]) - row = chain * xk ** (3 * BASE_FIELD_BITS) - nxt = row * GEN ** (3 * BASE_FIELD_BITS) - for k in unroll(0, BASE_FIELD_BITS): - c = 3 * k - nxt[c:c + 3] = mul192(row[c:c + 3], add192(zv, one_plus)) - if k != BASE_FIELD_BITS - 1: - zv = mul192(zv, zv) - return - - -def fold_final_msg(msg, point, log_len: Const): - # Weighted fold of the final_msg multilinear over 2^log_len values (log_len is - # the candidate's yr_log_n; the frame buffers use the global max size). - l0 = StackBuf(3 * 2 ** YR_LOG_CAP) - p0 = point[0:3] - for t in unroll(0, 2 ** log_len // 2): - k = 3 * t - lo = 6 * t - hi = 6 * t + 3 - a = msg[lo:lo + 3] - l0[k:k + 3] = add192(a, mul192(p0, add192(a, msg[hi:hi + 3]))) - cursor = l0 - n = 2 ** log_len // 2 - for j in unroll(1, log_len): - pj_off = 3 * j - pj = point[pj_off:pj_off + 3] - nxt = StackBuf(3 * 2 ** YR_LOG_CAP) - for t in unroll(0, n // 2): - k = 3 * t - lo = 6 * t - hi = 6 * t + 3 - a = cursor[lo:lo + 3] - nxt[k:k + 3] = add192(a, mul192(pj, add192(a, cursor[hi:hi + 3]))) - cursor = nxt - n = n // 2 - return cursor[0:3] - - -def sumcheck_round4(rd): - # One PLAIN sumcheck round off the round record at `rd`, whose successor it - # writes. The prover sends the round polynomial's coefficients bar the one the - # split h(0) + h(1) == claim fixes, so the verifier derives that one and reads h - # at the challenge by Horner. Nothing is reapplied: no eq factor, no separate - # term for the tables still waiting, and the eq point is not read here at all. - fs, c0, cursor = fs_next(rd[0:4], rd[GEN ** ROUND_CURSOR]) - fs, c2, cursor = fs_next(fs, cursor) - fs, c3, cursor = fs_next(fs, cursor) - c1 = add192(rd[ROUND_CLAIM:ROUND_CLAIM + 3], add192(c2, c3)) # the split fixes it, so it is neither sent nor bound - fs, y = squeeze(fs) - nxt = rd * GEN ** ROUND_SLOTS - nxt[0:4] = fs - nxt[GEN ** ROUND_CURSOR] = cursor - nxt[ROUND_CLAIM:ROUND_CLAIM + 3] = add192(c0, mul192(y, add192(c1, mul192(y, add192(c2, mul192(y, c3)))))) - return y - - -def sumcheck_round5(rd, prev, challenge_out): - # One GKR round off the round record at `rd`, whose successor it writes, its - # challenge landing at `challenge_out`. The prover sends every coefficient but - # c0, which the round's pulled-out eq factor (the challenge at `prev`) leaves - # fixed: `c0 + prev_challenge * (c1 + ... + c4) == claim`. - prev_challenge = prev[0:3] - fs, c1, cursor = fs_next(rd[0:4], rd[GEN ** ROUND_CURSOR]) - fs, c2, cursor = fs_next(fs, cursor) - fs, c3, cursor = fs_next(fs, cursor) - fs, c4, cursor = fs_next(fs, cursor) - fs, y = squeeze(fs) - challenge_out[0:3] = y - c0 = add192(rd[ROUND_CLAIM:ROUND_CLAIM + 3], mul192(prev_challenge, add192(add192(c1, c2), add192(c3, c4)))) - nxt = rd * GEN ** ROUND_SLOTS - nxt[0:4] = fs - nxt[GEN ** ROUND_CURSOR] = cursor - nxt[ROUND_CLAIM:ROUND_CLAIM + 3] = add192(c0, mul192(y, add192(c1, mul192(y, add192(c2, mul192(y, add192(c3, mul192(y, c4)))))))) - return - - -def batch_sumcheck(fs: StackBuf(4), msgs, running: StackBuf(3), point, n_rounds: Const): - # The rounds of a claim-batching sumcheck: two hinted values per round - # (g(1) and g(inf)), the split fixing the third against the running claim, and - # the challenges collected into `point`. - for rd in unroll(0, n_rounds): - fs, msg_g1, c = fs_next(fs, msgs * GEN ** (6 * rd)) - fs, msg_ginf, c = fs_next(fs, c) - fs, rv = squeeze(fs) - k = 3 * rd - point[k:k + 3] = rv - g_zero = add192(running, msg_g1) - c_one = add192(g_zero, add192(msg_g1, msg_ginf)) - running = add192(mul192(add192(mul192(msg_ginf, rv), c_one), rv), g_zero) # fold the degree-2 round at rv - return fs, running - - -# ==================================== Merkle paths ================================== - - -@inline -def order_children(n0, n1, sibling, bit): - # Branchless child ordering over two-word values: `bit` is boolean-pinned - # wherever it comes from, so `m = bit*(node + sibling)` selects rather than - # branches, leaving (node, sibling) at bit 0 and (sibling, node) at bit 1. - m0 = bit * (n0 + sibling[0]) - m1 = bit * (n1 + sibling[1]) - return [n0 + m0, n1 + m1, sibling[0] + m0, sibling[1] + m1] - - -@inline -def verify_merkle_path(leaf, direction_bits, depth: Const): - # Hinted child pairs are hashed in order; the boolean query bit selects the - # child that must equal the running node, binding each link of the path. - path = StackBuf(LIG_PATH_CAP) - hint_witness(path[0:8 * depth], "merkle_children") - path_ptr = addr(path) - node = leaf - for level in unroll(0, depth): - dir_bit = direction_bits[GEN ** level] - selected = path_ptr * GEN ** (8 * level) * (1 + dir_bit * (1 + GEN ** 4)) - selected[0:4] = node - parent = StackBuf(4) - blake2s(path[8 * level:8 * level + 4], path[8 * level + 4:8 * level + 8], parent) - node = parent - return node - - -@inline -def hash_cap_node(cap, index): - # Node n of a cap occupies words 4n..4n+4, so its children start at 8n. - children = cap * index ** 8 - parent = cap * index ** 4 - blake2s(children[0:4], children[4:8], parent[0:4]) - return - - -def verify_merkle_cap(cap, flags, depth: Const): - if depth != 0: - hash_cap_node(cap, GEN) - flags[GEN] = 1 - for parent in mul_range(GEN, GEN ** (2 ** (depth - 1))): - for side in unroll(0, 2): - child = parent * parent * GEN ** side - if flags[child] != 0: - flags[parent] = 1 # every active node forces its parent to be hashed - hash_cap_node(cap, child) - return cap[4:8] - - -# ============================== the stacked WHIR opening ============================ - - -def opening_fold(fs: StackBuf(4), cursor, c0: StackBuf(3), c1: StackBuf(3), c2: StackBuf(3)): - fs, r = squeeze(fs) - claim = add192(mul192(add192(mul192(c2, r), c1), r), c0) - fs, n0, cursor = fs_next(fs, cursor) - fs, n2, cursor = fs_next(fs, cursor) - return fs, cursor, claim, n0, n2, r - - -def opening_ood_point(fs: StackBuf(4), point, n_g): - # Share the squeeze loop across point dimensions. - states = HeapBuf((n_g * GEN) ** FS_SLOTS) - states[0:4] = fs - for x in mul_range(1, n_g): - state = states * x ** FS_SLOTS - nb, r = squeeze(state[0:4]) - x3 = x ** 3 - point[x3:x3 + 3] = r - state[4:8] = nb - last = states * n_g ** FS_SLOTS - return last[0:4] - - -def opening_final_message(fs: StackBuf(4), cursor, out, n: Const): - for i in unroll(0, n): - fs, value, cursor = fs_next(fs, cursor) - k = 3 * i - out[k:k + 3] = value - return fs, cursor - - -def opening_queries(cap, flags, query_weights, query_bit_ptrs, fold_point, n_queries_g, base: Const, folds: Const, blocks: Const, depth: Const, cap_depth: Const): - # Specialize by row and path shape so opening configurations share query code. - # A row enters its level's claim through its multilinear extension at the - # level's fold point, folded one coordinate at a time. At level 0, slot i of a - # leaf image is interleaving index n-1-i (see open_stacked), which complements - # every index bit, so a fold there keeps the high half's term. - query_sum_chain = HeapBuf((n_queries_g * GEN) ** 3) - query_sum_chain[0:3] = ZERO - for xe in mul_range(1, n_queries_g): - interleave = 2 ** folds # bound in the body, which captures names by value - row = StackBuf(LIG_ROW_CAP) - fold = StackBuf(3 * LIG_ROW_CAP) - r0 = fold_point[0:3] - if base == 1: - # Level 0's lanes are words, so its first fold runs limb by limb. - hint_witness(row[0:interleave], "merkle_leaf_rows") - for t in unroll(0, interleave // 2): - lo = row[2 * t] - hi = row[2 * t + 1] - d = lo + hi - k = 3 * t - fold[k:k + 3] = [hi + d * r0[0], d * r0[1], d * r0[2]] - else: - # A deeper row is `interleave` elements, three limbs each. - hint_witness(row[0:3 * interleave], "merkle_leaf_rows") - for t in unroll(0, interleave // 2): - k = 3 * t - lo = 6 * t - a = row[lo:lo + 3] - fold[k:k + 3] = add192(a, mul192(r0, add192(a, row[lo + 3:lo + 6]))) - # Fold c reads fold c-1's outputs and writes the next run of the buffer. - for c in unroll(1, folds): - rc = fold_point[3 * c:3 * c + 3] - src = 3 * (interleave - interleave // 2 ** (c - 1)) - dst = 3 * (interleave - interleave // 2 ** c) - for t in unroll(0, interleave // 2 ** (c + 1)): - a = fold[src + 6 * t:src + 6 * t + 3] - b = fold[src + 6 * t + 3:src + 6 * t + 6] - if base == 1: - fold[dst + 3 * t:dst + 3 * t + 3] = add192(b, mul192(rc, add192(a, b))) - else: - fold[dst + 3 * t:dst + 3 * t + 3] = add192(a, mul192(rc, add192(a, b))) - dot_at = 3 * (interleave - 2) - row_dot = fold[dot_at:dot_at + 3] - # The row's words are its leaf's byte image, hashed as full BLAKE2s blocks. - leaf_state = StackBuf(4) - blake2s(row[0:4], row[4:8], leaf_state, counter=64, final=1 // blocks) - for jb in unroll(1, blocks): - leaf_digest = StackBuf(4) - blake2s(row[8 * jb:8 * jb + 4], row[8 * jb + 4:8 * jb + 8], leaf_digest, cv=leaf_state, counter=64 * (jb + 1), final=(jb + 1) // blocks) - leaf_state = leaf_digest - xe3 = xe ** 3 - nxt = xe3 * GEN ** 3 - query_sum_chain[nxt:nxt + 3] = add192(query_sum_chain[xe3:xe3 + 3], mul192(query_weights[xe3:xe3 + 3], row_dot)) - direction_bits = query_bit_ptrs[xe] - path_depth = depth - cap_depth - node = verify_merkle_path(leaf_state, direction_bits, path_depth) - if cap_depth != 0: - parent = GEN ** (2 ** (cap_depth - 1)) - for bit in unroll(0, cap_depth - 1): - parent *= 1 + direction_bits[GEN ** (path_depth + 1 + bit)] * (1 + GEN ** (2 ** bit)) - flags[parent] = 1 # the cap check propagates this obligation to the root - cap_index = parent * parent * (1 + direction_bits[GEN ** path_depth] * (1 + GEN)) - else: - cap_index = GEN - slot = cap * cap_index ** 4 - slot[0:4] = node - end = n_queries_g ** 3 - return query_sum_chain[end:end + 3] - - -@inline -def fold_at(point, f: Const, lane_folds: Const, tail_start: Const): - # Where fold challenge f (in round order) sits in the witness-order point: the - # lane folds after the tail, every later fold `lane_folds` places down. - if const(f // lane_folds == 0): - pos = tail_start + f - else: - pos = f - lane_folds - return point * GEN ** (3 * pos) - - -def open_stacked(m_idx: Const, fs: StackBuf(4), target: StackBuf(3), commit_root: StackBuf(4), cursor): - # The stacked WHIR opening, one specialization per (rate, committed log-size) - # candidate: every LIG_* table reads row m_idx, per level row `ml`, and all - # opening proof data is hinted here, so only the executed arm pops its streams. - # - # The returned point is in witness order. The folds bind coordinates in ROUND - # order, and level 0's folds are the lane fold, binding the witness's TOP k - # coordinates: lane l of the commitment is the stack block q[l * 2^(mu-k) ...], - # which is what makes the witness's zero padding whole lanes for the committer - # to leave out of the encode. Every transparent weight downstream is written in - # witness coordinates, so each challenge is written straight to its witness - # coordinate as it arrives (`fold_at`), and the round-order readers below index - # it there. - n_levels = LIG_N_LEVELS[m_idx] - yr_level = LIG_YR_LEVEL[m_idx] - yr_log = LIG_YR_LOG_LEN[m_idx] - yr_len = LIG_YR_LEN[m_idx] - n_folds = LIG_TOTAL_FOLDS[m_idx] - max_q = LIG_MAX_QUERIES[m_idx] - ood_stride = LIG_MAX_OOD_SAMPLES * OOD_SLOTS - lane_folds = LIG_FOLDS[m_idx * LIG_MAX_LEVELS] - fold_head = n_folds - lane_folds - tail_start = fold_head + yr_log - # Padding is zero so shared prefix chains may compute unused entries above m - # without reading uninitialized cells. - point = HeapBuf(3 * (SIZE_BITS + SLOT_STRIDE_LOG)) - for j in unroll(n_folds + yr_log, SIZE_BITS + SLOT_STRIDE_LOG): - k = 3 * j - point[k:k + 3] = ZERO - tail = point * GEN ** (3 * fold_head) - - # The opening's scalars (sumcheck messages, level roots, nonces, final message) - # ride the SHARED stream, walked on in protocol order. The K opener binds a - # Merkle root as two scalars, one per 128-bit half. - msg_cursor = cursor - fs, round_quad_c, msg_cursor = fs_next(fs, msg_cursor) # the round polynomial in coefficients - fs, round_quad_a, msg_cursor = fs_next(fs, msg_cursor) # bar the linear one, which the split fixes - round_quad_b = add192(target, round_quad_a) - sumcheck_target = target - - # Caps are shared across queries, a node four words; rows and paths are hinted in - # each query's frame. - merkle_caps = HeapBuf(GEN ** (8 * LIG_CAP_LEN[m_idx])) - hint_witness(merkle_caps[0:8 * LIG_CAP_LEN[m_idx]], "merkle_caps") - cap_flags = HeapBuf(GEN ** LIG_CAP_LEN[m_idx]) - hint_witness(cap_flags[0:LIG_CAP_LEN[m_idx]], "merkle_cap_active") - final_msg = HeapBuf(GEN ** (3 * yr_len)) # filled from the stream at the last level - # Each cap root is checked against its transcript-bound level root. - level_roots = HeapBuf(GEN ** (4 * n_levels)) - level_roots[0:4] = commit_root - # ...and guest-filled accumulators (one slot per level / per query): - level_betas = HeapBuf(GEN ** (3 * n_levels)) - query_weights = HeapBuf(GEN ** (3 * n_levels * max_q)) - query_positions = HeapBuf(GEN ** (LIG_POSITIONS_LEN[m_idx])) - query_bit_ptrs = HeapBuf(GEN ** (LIG_POSITIONS_LEN[m_idx])) - # Explicit OOD claims bind every recursive Johnson-list commitment. L0 needs - # none: the opening claim itself is its post-commit binding value. An OOD claim - # is read before its level's query positions but batched after them (one - # challenge per level), so its value and intro message wait here. - ood_z = HeapBuf(GEN ** (3 * n_levels * LIG_MAX_OOD_SAMPLES * LIG_LOG_MSG_COLS_CAP)) - ood = HeapBuf(GEN ** (n_levels * ood_stride)) - - for lvl in unroll(0, n_levels): - ml = m_idx * LIG_MAX_LEVELS + lvl - n_queries = LIG_QUERIES[ml] - depth = LIG_TREE_DEPTH[ml] - folds_off = LIG_FOLDS_OFF[ml] - pos_off = LIG_POSITIONS_OFF[ml] - for j in unroll(0, LIG_FOLDS[ml]): - fs, msg_cursor, sumcheck_target, round_quad_c, round_quad_a, fold_challenge = opening_fold(fs, msg_cursor, round_quad_c, round_quad_b, round_quad_a) - fc = fold_at(point, folds_off + j, lane_folds, tail_start) - fc[0:3] = fold_challenge - round_quad_b = add192(sumcheck_target, round_quad_a) - - if lvl == yr_level: - fs, msg_cursor = opening_final_message(fs, msg_cursor, final_msg, yr_len) - else: - # A non-canonical half is rejected (merkle.rs `scalars_to_hash`), as - # the commitment root was at its own read. - fs, root_a, msg_cursor = fs_next_half(fs, msg_cursor) - fs, root_b, msg_cursor = fs_next_half(fs, msg_cursor) - k = 4 * lvl + 4 - level_roots[k:k + 2] = root_a[0:2] - level_roots[k + 2:k + 4] = root_b[0:2] - # OOD binding for the newly observed level-(lvl+1) commitment. The - # random point has the just-folded witness dimension, namely this - # level's message-column dimension. - for os in unroll(0, LIG_OOD_SAMPLES[ml + 1]): - oz = ood_z * GEN ** (3 * ((lvl + 1) * LIG_MAX_OOD_SAMPLES + os) * LIG_LOG_MSG_COLS_CAP) - fs = opening_ood_point(fs, oz, GEN ** LIG_LOG_MSG_COLS[ml]) - sample = ood * GEN ** ((lvl + 1) * ood_stride + os * OOD_SLOTS) - fs, ood_y, msg_cursor = fs_next(fs, msg_cursor) - fs, ood_c0, msg_cursor = fs_next(fs, msg_cursor) - fs, ood_c2, msg_cursor = fs_next(fs, msg_cursor) - sample[OOD_Y:OOD_Y + 3] = ood_y - sample[OOD_C0:OOD_C0 + 3] = ood_c0 - sample[OOD_C2:OOD_C2 + 3] = ood_c2 # the split fixes c1 = y + c2 - q_nonce = msg_cursor[0:3] # raw transport scalar: bound by the DS_POW_NONCE absorb below - msg_cursor = msg_cursor * GEN ** 3 - if LIG_QUERY_GRIND_BITS[ml] != 0: - grind_check(fs, q_nonce, GEN ** LIG_QUERY_GRIND_BITS[ml]) - else: - assert_eq192(q_nonce, ZERO) - nonce_state = StackBuf(4) - blake2s(fs, [q_nonce, DS_POW_NONCE], nonce_state) - fs = nonce_state - - sqz = HeapBuf((GEN ** (LIG_MAX_SQUEEZES[m_idx] + 1)) ** FS_SLOTS) - sqz[0:4] = fs - for xs in mul_range(1, GEN ** LIG_SQUEEZES[ml]): - # a loop body captures free names BY VALUE, so the compile-time aliases - # are rebound here (m_idx and lvl are substituted literals) - depth = LIG_TREE_DEPTH[m_idx * LIG_MAX_LEVELS + lvl] - pos_off = LIG_POSITIONS_OFF[m_idx * LIG_MAX_LEVELS + lvl] - row = sqz * xs ** FS_SLOTS - nb, packed = squeeze(row[0:4]) - row[4:8] = nb - query_ptr = xs ** (FIELD_BITS // depth) - decode_query_bits(packed, query_positions * GEN ** pos_off * query_ptr, query_bit_ptrs * GEN ** pos_off * query_ptr, depth) - sqz_end = sqz * (GEN ** LIG_SQUEEZES[ml]) ** FS_SLOTS - fs = sqz_end[0:4] - - # One batching challenge for the level, drawn once every claim it batches is - # fixed: its OOD claims above and these query positions. Claim tau of the - # level is weighted lam^tau, the running claim keeping lam^0 (Annex B, - # Protocol 1 step 1): query i is claim n_ood + 1 + i, so its weight splits - # into lam^i here and the level scalar lam^(n_ood+1) below. - fs, lam = squeeze(fs) - lam_pow = ONE - for i in unroll(0, n_queries): - k = 3 * (lvl * max_q + i) - query_weights[k:k + 3] = lam_pow - lam_pow = mul192(lam_pow, lam) - - cap_depth = LIG_CAP_DEPTH[ml] - cap = merkle_caps * GEN ** (8 * LIG_CAP_OFF[ml]) - flags = cap_flags * GEN ** LIG_CAP_OFF[ml] - level_root = verify_merkle_cap(cap, flags, cap_depth) - k = 4 * lvl - level_roots[k:k + 4] = level_root - - level_query_sum = opening_queries(cap, flags, query_weights * GEN ** (3 * lvl * max_q), query_bit_ptrs * GEN ** pos_off, fold_at(point, folds_off, lane_folds, tail_start), GEN ** n_queries, 1 // (lvl + 1), LIG_FOLDS[ml], LIG_LEAF_BLOCKS[ml], depth, cap_depth) - - # Every level, including the last, ties its commitment in through an intro - # message. The level's claims then enter the running one with powers of - # `lam`: the OOD claims held above first, then this query batch. - fs, intro_c0, msg_cursor = fs_next(fs, msg_cursor) - fs, intro_c2, msg_cursor = fs_next(fs, msg_cursor) - intro_c1 = add192(level_query_sum, intro_c2) # the split fixes the linear coefficient - if lvl == yr_level: - beta_lvl = lam # no OOD claim at the last level: no new oracle - else: - ood_scalar = lam - for os in unroll(0, LIG_OOD_SAMPLES[ml + 1]): - sample = ood * GEN ** ((lvl + 1) * ood_stride + os * OOD_SLOTS) - ood_y = sample[OOD_Y:OOD_Y + 3] - ood_c2 = sample[OOD_C2:OOD_C2 + 3] - sample[OOD_BETA:OOD_BETA + 3] = ood_scalar - round_quad_c = add192(round_quad_c, mul192(ood_scalar, sample[OOD_C0:OOD_C0 + 3])) - round_quad_b = add192(round_quad_b, mul192(ood_scalar, add192(ood_y, ood_c2))) - round_quad_a = add192(round_quad_a, mul192(ood_scalar, ood_c2)) - sumcheck_target = add192(sumcheck_target, mul192(ood_scalar, ood_y)) - ood_scalar = mul192(ood_scalar, lam) - beta_lvl = ood_scalar - k = 3 * lvl - level_betas[k:k + 3] = beta_lvl - round_quad_c = add192(round_quad_c, mul192(beta_lvl, intro_c0)) - round_quad_b = add192(round_quad_b, mul192(beta_lvl, intro_c1)) - round_quad_a = add192(round_quad_a, mul192(beta_lvl, intro_c2)) - sumcheck_target = add192(sumcheck_target, mul192(beta_lvl, level_query_sum)) - - # ---- finish the sumcheck over the tail coordinates ---- - for j in unroll(0, yr_log - 1): - fs, msg_cursor, sumcheck_target, round_quad_c, round_quad_a, tail_c = opening_fold(fs, msg_cursor, round_quad_c, round_quad_b, round_quad_a) - k = 3 * j - tail[k:k + 3] = tail_c - round_quad_b = add192(sumcheck_target, round_quad_a) - # The closing round sends no following message. - fs, tail_last = squeeze(fs) - k = 3 * (yr_log - 1) - tail[k:k + 3] = tail_last - sumcheck_target = add192(round_quad_c, mul192(tail_last, add192(round_quad_b, mul192(tail_last, round_quad_a)))) - - yr_at_tail = fold_final_msg(final_msg, tail, yr_log) - - # ---- per-level induced bases at the single terminal point ---- - # Every query of a level runs the SAME product shape over its message-column - # coordinates; only the novel-basis chain (the query position's - # subspace-vanishing walk, in F64) differs. Since - # 1 + c_t * (1 + chain_t * inv_t) == (1 + c_t) + (c_t * inv_t) * chain_t - # a coordinate's two coefficients depend on the challenge and the baked - # vanishing inverse alone, so they hoist out of the query loop (one row a level, - # fold coords then tail coords) and each query is one multiply-add a - # coordinate. - basis_a = HeapBuf(GEN ** (3 * n_levels * LIG_LOG_MSG_COLS_CAP)) - basis_b = HeapBuf(GEN ** (3 * n_levels * LIG_LOG_MSG_COLS_CAP)) - for lvl in unroll(0, n_levels): - ml = m_idx * LIG_MAX_LEVELS + lvl - prefix_len = LIG_RESIDUAL_PREFIX_LEN[ml] - vanish = m_idx * LIG_MAX_VANISH_LEN + LIG_VANISH_OFF[ml] - for t in unroll(0, prefix_len): - fold_c = fold_at(point, LIG_RESIDUAL_FOLD_OFF[ml] + t, lane_folds, tail_start) - dst = 3 * (lvl * LIG_LOG_MSG_COLS_CAP + t) - basis_a[dst:dst + 3] = add192(ONE, fold_c[0:3]) - basis_b[dst:dst + 3] = scale(LIG_VANISH_INVS[vanish + t], fold_c[0:3]) - for j in unroll(0, yr_log): - src = 3 * j - tail_c = tail[src:src + 3] - dst = 3 * (lvl * LIG_LOG_MSG_COLS_CAP + prefix_len + j) - basis_a[dst:dst + 3] = add192(ONE, tail_c) - basis_b[dst:dst + 3] = scale(LIG_VANISH_INVS[vanish + prefix_len + j], tail_c) - inner_chain = HeapBuf(GEN ** (3 * (n_levels + 1))) - inner_chain[0:3] = ZERO - for lvl in unroll(0, n_levels): - ml = m_idx * LIG_MAX_LEVELS + lvl - residual_chain = HeapBuf(GEN ** (3 * (max_q + 1))) - residual_chain[0:3] = ZERO - for xr in mul_range(1, GEN ** LIG_QUERIES[ml]): - ml = m_idx * LIG_MAX_LEVELS + lvl # rebound: the body captures by value - vanish = m_idx * LIG_MAX_VANISH_LEN + LIG_VANISH_OFF[ml] - basis_row = lvl * LIG_LOG_MSG_COLS_CAP - max_q = LIG_MAX_QUERIES[m_idx] - basis_chain = query_positions[GEN ** LIG_POSITIONS_OFF[ml] * xr] - b0 = 3 * basis_row - prefix_eq = add192(basis_a[b0:b0 + 3], scale(basis_chain, basis_b[b0:b0 + 3])) - for t in unroll(1, LIG_LOG_MSG_COLS[ml]): - # subspace-vanishing recurrence for the novel-basis point - basis_chain *= (basis_chain + LIG_VANISH_VALS[vanish + t - 1]) - bt = 3 * (basis_row + t) - prefix_eq = mul192(prefix_eq, add192(basis_a[bt:bt + 3], scale(basis_chain, basis_b[bt:bt + 3]))) - xr3 = xr ** 3 - nxt = xr3 * GEN ** 3 - qw = GEN ** (3 * lvl * max_q) * xr3 - residual_chain[nxt:nxt + 3] = add192(residual_chain[xr3:xr3 + 3], mul192(query_weights[qw:qw + 3], prefix_eq)) - # accumulate beta_lvl * (per-level residual sum) into the grand residual - end = GEN ** (3 * LIG_QUERIES[ml]) - k = 3 * lvl - inner_chain[k + 3:k + 6] = add192(inner_chain[k:k + 3], mul192(level_betas[k:k + 3], residual_chain[end:end + 3])) - - # Explicit OOD eq bases at the same terminal point. - ood_inner = ZERO - for ood_lvl in unroll(1, n_levels): - ml = m_idx * LIG_MAX_LEVELS + ood_lvl - z_folded = LIG_LOG_MSG_COLS[ml - 1] - yr_log - ris_start = LIG_FOLDS_OFF[ml] - for os in unroll(0, LIG_OOD_SAMPLES[ml]): - oz = ood_z * GEN ** (3 * (ood_lvl * LIG_MAX_OOD_SAMPLES + os) * LIG_LOG_MSG_COLS_CAP) - sample = ood * GEN ** (ood_lvl * ood_stride + os * OOD_SLOTS) - scalar = sample[OOD_BETA:OOD_BETA + 3] - for t in unroll(0, z_folded): - zk = 3 * t - fc = fold_at(point, ris_start + t, lane_folds, tail_start) - scalar = mul192(scalar, add192(ONE, add192(oz[zk:zk + 3], fc[0:3]))) - for t in unroll(0, yr_log): - zk = 3 * (z_folded + t) - tk = 3 * t - scalar = mul192(scalar, add192(ONE, add192(oz[zk:zk + 3], tail[tk:tk + 3]))) - ood_inner = add192(ood_inner, scalar) - inner_end = 3 * n_levels - return sumcheck_target, point, add192(inner_chain[inner_end:inner_end + 3], ood_inner), yr_at_tail - - -# ============================== inner-proof verification ============================ -# The phases of `verify_sub`, in the order it runs them. Each takes and returns the -# Fiat-Shamir state and the stream cursor, so the sequence is what binds them. - - -def verify_bus_gkr(fs: StackBuf(4), cursor, g_bus_mu, zeta): - # ONE GKR grand product over push, pull and count, RLC-batched. Push and pull - # have equal depth (matched blocks) and the count tree is padded with identity - # leaves up to it (product unchanged), so a single sumcheck serves all three. - # Radix four contracts two binary levels per layer; after checking the combined - # product identity, a fresh λ pins the individual values. All three trees reduce - # to the one shared point `zeta`, which this fills, returning their three leaf - # values with the walked Fiat-Shamir state and stream cursor. - layers = HeapBuf((g_bus_mu * GEN ** 2) ** LAYER_SLOTS) # mu + 2 layers - rounds = HeapBuf(GKR_ROUNDS_CAP * ROUND_SLOTS) - gkr_pts = HeapBuf(3 * GKR_POINTS_CAP) - assert log(g_bus_mu) < COUNT_BITS - fs, root_push, cursor = fs_next(fs, cursor) - fs, root_count, cursor = fs_next(fs, cursor) - assert_ne192(root_count, ZERO) # count-tree root nonzero: no read count self-cancels - fs, root_lambda = squeeze(fs) - layers[0:4] = fs - layers[GEN ** LAYER_CURSOR] = cursor - layers[LAYER_PUSH:LAYER_PUSH + 3] = root_push - layers[LAYER_PULL:LAYER_PULL + 3] = root_push - layers[LAYER_COUNT:LAYER_COUNT + 3] = root_count - layers[LAYER_LAMBDA:LAYER_LAMBDA + 3] = root_lambda # λ over the three roots - layers[GEN ** LAYER_ROW] = gkr_pts - layers[GEN ** LAYER_POS] = GEN ** 0 - - # Contract two binary product levels at a time. pair_bounds[g^d] = g^(d//2) is - # the radix-four layer count, and shift is g exactly when the depth is odd, in - # which case the root-most BINARY layer runs first. - pair_bounds = HeapBuf(COUNT_BITS) - depth_shift = HeapBuf(COUNT_BITS) - for depth in unroll(0, COUNT_BITS): - pair_bounds[GEN ** depth] = GEN ** (depth // 2) - depth_shift[GEN ** depth] = GEN ** (depth % 2) - - if depth_shift[g_bus_mu] != 1: - # The odd layer is layer 0, so its round state would be written and read - # back at the same position: read it straight off the layer instead. - lam = layers[LAYER_LAMBDA:LAYER_LAMBDA + 3] - tail_fs = layers[0:4] - tcur = layers[GEN ** LAYER_CURSOR] - tclaim = add192(layers[LAYER_PUSH:LAYER_PUSH + 3], mul192(lam, add192(layers[LAYER_PULL:LAYER_PULL + 3], mul192(lam, layers[LAYER_COUNT:LAYER_COUNT + 3])))) - nextrow = gkr_pts * GEN ** (3 * MU_CAP) - evals = StackBuf(3 * 2 * N_GKR_SIDES) # the two children of each side, in side order - for i in unroll(0, 2 * N_GKR_SIDES): - tail_fs, ev, tcur = fs_next(tail_fs, tcur) - k = 3 * i - evals[k:k + 3] = ev - combined = ZERO - for i in unroll(0, N_GKR_SIDES): - side = N_GKR_SIDES - 1 - i # Horner in lam, so the top side lands last - k = 6 * side - combined = add192(mul192(evals[k:k + 3], evals[k + 3:k + 6]), mul192(lam, combined)) - assert_eq192(tclaim, combined) - tail_fs, c0 = squeeze(tail_fs) - nextrow[0:3] = c0 - tail_fs, tail_lambda = squeeze(tail_fs) # fresh λ pins the tail individuals - nxt = layers * GEN ** LAYER_SLOTS - nxt[0:4] = tail_fs - nxt[GEN ** LAYER_CURSOR] = tcur - for side in unroll(0, N_GKR_SIDES): - k = 6 * side - out = LAYER_PUSH + 3 * side - ev_lo = evals[k:k + 3] - nxt[out:out + 3] = add192(ev_lo, mul192(c0, add192(ev_lo, evals[k + 3:k + 6]))) - nxt[LAYER_LAMBDA:LAYER_LAMBDA + 3] = tail_lambda - nxt[GEN ** LAYER_ROW] = nextrow - nxt[GEN ** LAYER_POS] = GEN - - for x_pair in mul_range(1, pair_bounds[g_bus_mu]): - x_layer = x_pair * x_pair * depth_shift[g_bus_mu] - layer = layers * x_layer ** LAYER_SLOTS - lam = layer[LAYER_LAMBDA:LAYER_LAMBDA + 3] - point_row = layer[GEN ** LAYER_ROW] - round_pos = layer[GEN ** LAYER_POS] - nextrow = point_row * GEN ** (3 * MU_CAP) - head = rounds * round_pos ** ROUND_SLOTS - head[0:4] = layer[0:4] - head[GEN ** ROUND_CURSOR] = layer[GEN ** LAYER_CURSOR] - head[ROUND_CLAIM:ROUND_CLAIM + 3] = add192(layer[LAYER_PUSH:LAYER_PUSH + 3], mul192(lam, add192(layer[LAYER_PULL:LAYER_PULL + 3], mul192(lam, layer[LAYER_COUNT:LAYER_COUNT + 3])))) - for x_round in mul_range(1, x_layer): - xr3 = x_round ** 3 - sumcheck_round5(rounds * (round_pos * x_round) ** ROUND_SLOTS, point_row * xr3, nextrow * xr3 * GEN ** 6) - tail = rounds * (round_pos * x_layer) ** ROUND_SLOTS - tail_fs = tail[0:4] - tcur = tail[GEN ** ROUND_CURSOR] - tclaim = tail[ROUND_CLAIM:ROUND_CLAIM + 3] - evals = StackBuf(3 * 4 * N_GKR_SIDES) # the four children of each side, in side order - for i in unroll(0, 4 * N_GKR_SIDES): - tail_fs, ev, tcur = fs_next(tail_fs, tcur) - k = 3 * i - evals[k:k + 3] = ev - combined = ZERO - for i in unroll(0, N_GKR_SIDES): - side = N_GKR_SIDES - 1 - i # Horner in lam, so the top side lands last - k = 12 * side - combined = add192(mul192(mul192(evals[k:k + 3], evals[k + 3:k + 6]), mul192(evals[k + 6:k + 9], evals[k + 9:k + 12])), mul192(lam, combined)) - assert_eq192(tclaim, combined) - tail_fs, c0 = squeeze(tail_fs) - tail_fs, c1 = squeeze(tail_fs) - nextrow[0:3] = c0 - nextrow[3:6] = c1 - tail_fs, tail_lambda = squeeze(tail_fs) - nxt = layer * GEN ** (2 * LAYER_SLOTS) - nxt[0:4] = tail_fs - nxt[GEN ** LAYER_CURSOR] = tcur - for side in unroll(0, N_GKR_SIDES): - k = 12 * side - out = LAYER_PUSH + 3 * side - e0 = evals[k:k + 3] - e2 = evals[k + 6:k + 9] - lo = add192(e0, mul192(c0, add192(e0, evals[k + 3:k + 6]))) - hi = add192(e2, mul192(c0, add192(e2, evals[k + 9:k + 12]))) - nxt[out:out + 3] = add192(lo, mul192(c1, add192(lo, hi))) - nxt[LAYER_LAMBDA:LAYER_LAMBDA + 3] = tail_lambda - nxt[GEN ** LAYER_ROW] = nextrow - nxt[GEN ** LAYER_POS] = round_pos * x_layer * GEN - last = layers * g_bus_mu ** LAYER_SLOTS - final_point_row = last[GEN ** LAYER_ROW] - for xt in mul_range(1, g_bus_mu): - x3 = xt ** 3 - zeta[x3:x3 + 3] = final_point_row[x3:x3 + 3] # the ONE shared point - return last[0:4], last[GEN ** LAYER_CURSOR], last[LAYER_PUSH:LAYER_PUSH + 3], last[LAYER_PULL:LAYER_PULL + 3], last[LAYER_COUNT:LAYER_COUNT + 3] - - -def verify_tables(fs: StackBuf(4), cursor, pi: StackBuf(4), zeta, g_bus_mu, dims_g, block_kappa, g_squares, fp_w, beta: StackBuf(3), claim_pool, claim_cplen_g, chi, claim_push: StackBuf(3), claim_pull: StackBuf(3), claim_count: StackBuf(3)): - # Settle the bus against the eight tables, in four steps: certify each side's - # leaf-cube tiling, decompose the three GKR leaf values over it, run the ONE - # table sumcheck they all reduce to, and bind the public input. Pooled claims - # land in claim_pool/claim_cplen_g and the reduced point in `chi`; the batch's - # round count g^n, the deferred bytecode share and the PI point come back. - # - # Bus-leaf packing offsets, for the selector certification. Each side's blocks - # tile its leaf cube, block b at offset_b; the hinted order is only - # PERMUTATION-checked and offsets accumulate as g^offset = Π_{earlier} g^(2^κ). - # The decompose below pins each block's selector bits against this offset, - # forcing κ-alignment, and no sort or tie-break check is needed: alignment plus - # consecutive offsets force a valid tiling, and the grand product is - # position-independent, so any tiling is sound. Pull's blocks mirror push's and - # share zeta, so only push and count need offsets (pull's slots go unread). - sort_order = HeapBuf(N_BLOCKS) - hint_witness(sort_order[0:N_BLOCKS], "sort_order") - block_side_tab = HeapBuf(N_BLOCKS) # global block -> its side - for b in unroll(0, N_BLOCKS): - block_side_tab[GEN ** b] = BLOCK_SIDE[b] - block_off_g = HeapBuf(N_BLOCKS) # g^offset per block, keyed by global index - for cert in unroll(0, 2): - s = COUNT_SIDE * cert # PUSH_SIDE (0), then COUNT_SIDE (2) - g_off = GEN ** 0 - for r in unroll(SIDE_BLOCK_START[s], SIDE_BLOCK_START[s + 1]): - global_g = sort_order[GEN ** r] # g^{global block index at this rank} - assert log(global_g) < N_BLOCKS # a valid block index - assert block_side_tab[global_g] == s # ...belonging to THIS side - # write-once: a repeat collides, and an omission fails the decompose's - # offset read below - block_off_g[global_g] = g_off - g_off *= g_squares[block_kappa[global_g]] - - # ---- 3x leaf decomposition (claims pooled; bytecode Public DEFERRED) ---- - # Reconstruct Ṽ₀(ζ) per side and assert it equals the GKR leaf value. The - # committed-coordinate values ride the stream (observed, pooled); Index - # coordinates use the factored index MLE; and the program's whole share of a - # bytecode leaf is ONE evaluation of the stacked polynomial, since its slots are - # aligned with the tuple and the weights are eq(α⃗, ·), so the share IS that - # polynomial at (ζ_lo, α⃗) (doc sec:e2e-bc): one hinted value, exported as a - # deferred claim, with no per-coordinate values and no selector challenge. - # - # Pull's blocks mirror push's (same kappas, same offsets, generator-asserted - # pairing) and share zeta, so each pull block REUSES its push twin's eq_hi and - # Index-MLE value instead of recomputing them; its column values are mostly - # deduped pool reads (COORD_FRESH). The identity check against pull's own GKR - # claim still binds everything. - bc_share = StackBuf(3) - hint_witness(bc_share, "bytecode_val") - idxc_tab = HeapBuf(SIZE_BITS) # INDEX_MLE_FACTORS[t] = 1 + g^(2^t) - for t in unroll(0, SIZE_BITS): - idxc_tab[GEN ** t] = INDEX_MLE_FACTORS[t] - bus_table_total = StackBuf(3 * N_GKR_SIDES) # per side, what its tables' blocks owe - block_eq_hi = StackBuf(3 * N_BLOCKS) # every block's eq_hi, reused below - block_index_mle = HeapBuf(3 * N_BLOCKS) # per push block with an Index coord - for s in unroll(0, N_GKR_SIDES): - acc = ZERO - selector_sum = ZERO - for b in unroll(SIDE_BLOCK_START[s], SIDE_BLOCK_START[s + 1]): - block_has_public = 0 - kappa_g = block_kappa[GEN ** b] - assert log(kappa_g) < SIZE_BITS - if s == PULL_SIDE: - twin = 3 * (b - SIDE_BLOCK_START[PULL_SIDE]) - eq_hi = block_eq_hi[twin:twin + 3] - else: - # eq_hi over the ζ coords above κ against the selector bits, whose - # run is mu_s − κ = g^mu_s / g^κ long. Selector bits = offset >> κ: - # advice-decompose the offset and read it shifted by κ. Rebuilding - # g^offset from those high bits alone (weights g^(2^(κ+k))) and - # asserting it equals block_off_g pins the bits AND the κ-alignment - # in one shot; the low κ bit cells are written but never read. - sel_len_g = g_bus_mu / kappa_g # g^(mu - κ) - assert log(sel_len_g) < SIZE_BITS - zeta_hi = zeta * kappa_g ** 3 - offset_bits = HeapBuf(GEN ** SIZE_BITS) - hint_decompose_bits_exponent(offset_bits, block_off_g[GEN ** b], SIZE_BITS) - sel_bits = offset_bits * kappa_g - eq_chain = HeapBuf(3 * (MU_CAP + 2)) - goff_chain = HeapBuf(MU_CAP + 2) # rebuild g^offset from the high bits - eq_chain[0:3] = ONE - goff_chain[GEN ** 0] = 1 - for xk in mul_range(1, sel_len_g): - sbit = sel_bits[xk] - sel_bits[xk] = sbit * sbit # booleanity as a write-once pin - x3 = xk ** 3 - nxt = x3 * GEN ** 3 - eq_chain[nxt:nxt + 3] = mul192(eq_chain[x3:x3 + 3], add64(zeta_hi[x3:x3 + 3], 1 + sbit)) # eq over GF(2) is 1 + b + z - goff_chain[xk * GEN] = goff_chain[xk] * (1 + sbit * (g_squares[kappa_g * xk] + 1)) - sel_end = sel_len_g ** 3 - eq_hi = eq_chain[sel_end:sel_end + 3] - assert goff_chain[sel_len_g] == block_off_g[GEN ** b] # bits == offset >> κ, κ-aligned - selector_sum = add192(selector_sum, eq_hi) - bk = 3 * b - block_eq_hi[bk:bk + 3] = eq_hi - # A TABLE's block streams no value here: the table sumcheck settles its - # fingerprint from that table's column evaluations. Only the framework - # blocks (boundary, memory, bytecode) still decompose. - if BLOCK_TABLE[b] == NO_TABLE: - # inner fingerprint Σ_i w_i · coord_i(ζ_lo); the count side weighs - # slot 0 alone (α⃗ = 0), γ = 0. - inner_sum = ZERO - for i in unroll(0, BLOCK_COORD_COUNT[b]): - ci = BLOCK_COORD_OFF[b] + i # a compile-time index, so `ci` costs nothing - coord_val = ZERO - slot = 3 * COORD_CLAIM_SLOT[ci] - if COORD_TYPE[ci] == COORD_KIND_CONST: - coord_val = f192(COORD_CONST[ci], 0, 0) - if COORD_TYPE[ci] == COORD_KIND_COL: - if COORD_FRESH[ci] == 1: - fs, coord_val, cursor = fs_next(fs, cursor) - claim_pool[slot:slot + 3] = coord_val - claim_cplen_g[GEN ** COORD_CLAIM_SLOT[ci]] = kappa_g # cplen = block kappa - else: - coord_val = claim_pool[slot:slot + 3] - if COORD_TYPE[ci] == COORD_KIND_GCOL: - if COORD_FRESH[ci] == 1: - fs, rawv, cursor = fs_next(fs, cursor) - claim_pool[slot:slot + 3] = rawv - claim_cplen_g[GEN ** COORD_CLAIM_SLOT[ci]] = kappa_g - else: - rawv = claim_pool[slot:slot + 3] - coord_val = scale(COORD_CONST[ci], rawv) - if COORD_TYPE[ci] == COORD_KIND_INDEX: - if s == PULL_SIDE: - twin = 3 * (b - SIDE_BLOCK_START[PULL_SIDE]) - coord_val = block_index_mle[twin:twin + 3] - else: - # Index-coord MLE: prod_t (1 + zeta_t * (1 + g^(2^t))) - idx_chain = HeapBuf(3 * (MU_CAP + 2)) - idx_chain[0:3] = ONE - for xt in mul_range(1, kappa_g): - x3 = xt ** 3 - nxt = x3 * GEN ** 3 - idx_chain[nxt:nxt + 3] = mul192(idx_chain[x3:x3 + 3], add64(scale(idxc_tab[xt], zeta[x3:x3 + 3]), 1)) - idx_end = kappa_g ** 3 - coord_val = idx_chain[idx_end:idx_end + 3] - if s == PUSH_SIDE: - block_index_mle[bk:bk + 3] = coord_val - if COORD_TYPE[ci] == COORD_KIND_PUBLIC: - # The public slots carry no value of their own here: their - # alpha-weighted sum IS bc_share, added once per block below - # (push and pull share zeta, so both get the same one). - block_has_public = 1 - if s == COUNT_SIDE: - inner_sum = add192(inner_sum, coord_val) - else: - wk = 3 * i - inner_sum = add192(inner_sum, mul192(fp_w[wk:wk + 3], coord_val)) - if block_has_public == 1: - inner_sum = add192(inner_sum, bc_share) # the bytecode blocks' public slots - if s == COUNT_SIDE: - acc = add192(acc, mul192(eq_hi, inner_sum)) - else: - acc = add192(acc, mul192(eq_hi, add192(beta, inner_sum))) - acc = add64(add192(acc, selector_sum), 1) - # What the tables' blocks owe this side: its GKR leaf value less the - # framework decomposition. DERIVED, not read: a transmitted total would be a - # free value in its own check. The table sumcheck's target pins it below. - if s == PUSH_SIDE: - leaf = claim_push - if s == PULL_SIDE: - leaf = claim_pull - if s == COUNT_SIDE: - leaf = claim_count - sk = 3 * s - bus_table_total[sk:sk + 3] = add192(acc, leaf) - claim_idx = N_BUS_CLAIMS # table and PI claims pool after the deduped bus claims - - # ---- ONE table sumcheck for all eight tables ---- - # Mirrors lean_vm::constraints::verify. zc_xi ONCE, each table folding its own - # identities with a DISJOINT range of its powers (ETA_OFFSET[t]); one shared - # point zeta (the bus GKR's); n = max_t tau_t rounds. Rounds bind the HIGHEST - # variable first, so a 2^tau table sits out the first n - tau and joins carrying - # the challenges it sat out, weighing cprod[n - tau] * peq[tau] where peq[tau] = - # eq(zeta[..tau], chi[..tau]); the zc_xi-powers are inside constraint_eval. - # - # g^n is hinted, then pinned exactly: the range-checked division slacks force it - # to dominate every certified tau and the product identity forces it to BE one. - g_zc_n = hint_witness("zc_tau_max") - zc_is_a_tau = 1 - for t in unroll(0, N_TABLES): - tau_g = dims_g[GEN ** (t + 1)] - assert log(g_zc_n / tau_g) < COUNT_BITS - zc_is_a_tau *= g_zc_n + tau_g - assert zc_is_a_tau == 0 - # n <= mu, the `Error::Truncated` of constraints.rs. Every table pushes at - # kappa = tau, so it holds structurally, but zc_peq below reads zeta[..n] and - # zeta only holds mu coords: unwritten heap there is prover-chosen. - assert log(g_bus_mu / g_zc_n) < COUNT_BITS - fs, zc_xi = squeeze(fs) - zc_xi_pows = StackBuf(3 * N_ETA_POWS) - zc_xi_pows[0:3] = ONE - for k in unroll(1, N_ETA_POWS): - c = 3 * k - zc_xi_pows[c:c + 3] = mul192(zc_xi_pows[c - 3:c], zc_xi) - # The eq point is the bus GKR's zeta, NOT a fresh one, which is what lets the - # batch settle the bus forms alongside the constraints. It is also why no target - # is read: what the three sides' tables owe, each in its own shared power of - # zc_xi, IS the sum the batch must reach, and zc_xi is squeezed after those - # totals are fixed, so hitting one number forces all three side equations. - bus_target = ZERO - for sd in unroll(0, N_GKR_SIDES): - e = 3 * (ETA_FORM_BASE + sd) - k = 3 * sd - bus_target = add192(bus_target, mul192(zc_xi_pows[e:e + 3], bus_table_total[k:k + 3])) - # n vanilla sumcheck rounds: the round polynomial arrives whole, so a round is - # `h(0) + h(1) == claim` and a fold, with no eq to reapply. The tables still - # waiting ride inside h, so nothing here is indexed by height; the heights enter - # only the per-table weights below. - zc_rounds = HeapBuf((g_zc_n * GEN) ** ROUND_SLOTS) - zc_cprod = HeapBuf((g_zc_n * GEN) ** 3) # the challenges bound so far, multiplied - zc_rounds[0:4] = fs - zc_rounds[GEN ** ROUND_CURSOR] = cursor - zc_rounds[ROUND_CLAIM:ROUND_CLAIM + 3] = bus_target - zc_cprod[0:3] = ONE - for xk in mul_range(1, g_zc_n): - rk = sumcheck_round4(zc_rounds * xk ** ROUND_SLOTS) - bound = (g_zc_n * INV_GEN / xk) ** 3 # g^(n-1-j): the variable round j binds - chi[bound:bound + 3] = rk - # cprod is the weight of a table that joins here; peq below is the rest - x3 = xk ** 3 - nxt = x3 * GEN ** 3 - zc_cprod[nxt:nxt + 3] = mul192(zc_cprod[x3:x3 + 3], rk) - zc_last = zc_rounds * g_zc_n ** ROUND_SLOTS - fs = zc_last[0:4] - cursor = zc_last[GEN ** ROUND_CURSOR] - claim = zc_last[ROUND_CLAIM:ROUND_CLAIM + 3] - # peq[g^tau] = eq(zeta[..tau], chi[..tau]), as a prefix chain. - zc_peq = HeapBuf((g_zc_n * GEN) ** 3) - eq_prefix_chain(zc_peq, ONE, zeta, chi, g_zc_n) - # Per table: every committed column's evaluation (pooled), its AIR constraint - # at its own reduced point chi[..tau_t], weighted into the batch's final claim. - air_acc = ZERO - for t in unroll(0, N_TABLES): - tau_g = dims_g[GEN ** (t + 1)] - col_evals = StackBuf(3 * TABLE_COLS_CAP) - for k in unroll(0, N_TABLE_COLS[t]): - fs, e, cursor = fs_next(fs, cursor) - c = 3 * k - col_evals[c:c + 3] = e - p = 3 * claim_idx - claim_pool[p:p + 3] = e - claim_cplen_g[GEN ** claim_idx] = tau_g # cplen = tau_t - claim_idx += 1 - # The table's AIR constraint at the final point (col_evals is indexed by - # local column index; the formulas mirror tables.rs eval_constraint). Every - # value relation rides the bus as a degree-2 coordinate, so only JUMP's - # is-nonzero indicator is left with an identity of its own. - constraint_eval = ZERO - if t == TABLE_JUMP: - # `b = c*w` and `c*(b+1) = 0` (tables.rs jump_identity). Local columns: - # v_cond at 5, w at 12, the indicator b at 13. - cond = col_evals[15:18] - indicator = col_evals[39:42] - x0 = 3 * ETA_OFFSET[t] - x1 = x0 + 3 - constraint_eval = add192(mul192(zc_xi_pows[x0:x0 + 3], add192(indicator, mul192(cond, col_evals[36:39]))), mul192(zc_xi_pows[x1:x1 + 3], mul192(cond, add192(indicator, ONE)))) - # The table's three bus forms, evaluated at the SAME column evaluations: - # Σ_b eq_hi(b) · (γ + Σ_i α^i · coord_i), the coords read off col_evals at - # their local index. This is what replaces opening those columns at ζ. - for sd in unroll(0, N_GKR_SIDES): - form = ZERO - for b in unroll(0, N_BLOCKS): - if BLOCK_SIDE[b] == sd: - if BLOCK_TABLE[b] == t: - inner = ZERO - for i in unroll(0, BLOCK_COORD_COUNT[b]): - # Each coord is the sum of its terms, over this table's - # column evaluations. A product term (an address, an - # arithmetic result) is degree 2, which the batch's - # round polynomial already allows. - ci = BLOCK_COORD_OFF[b] + i # a compile-time index, so `ci` costs nothing - cv = ZERO - for j in unroll(0, COORD_TERM_COUNT[ci]): - tj = COORD_TERM_OFF[ci] + j - ca = 3 * TERM_COL_A[tj] - cb = 3 * TERM_COL_B[tj] - if TERM_TYPE[tj] == COORD_KIND_CONST: - cv = add192(cv, f192(TERM_CONST[tj], 0, 0)) - if TERM_TYPE[tj] == COORD_KIND_COL: - cv = add192(cv, col_evals[ca:ca + 3]) - if TERM_TYPE[tj] == COORD_KIND_GCOL: - cv = add192(cv, scale(TERM_CONST[tj], col_evals[ca:ca + 3])) - if TERM_TYPE[tj] == COORD_KIND_PROD: - cv = add192(cv, scale(TERM_CONST[tj], mul192(col_evals[ca:ca + 3], col_evals[cb:cb + 3]))) - if sd == COUNT_SIDE: - inner = add192(inner, cv) - else: - wk = 3 * i - inner = add192(inner, mul192(fp_w[wk:wk + 3], cv)) - bk = 3 * b - if sd == COUNT_SIDE: - form = add192(form, mul192(block_eq_hi[bk:bk + 3], inner)) - else: - form = add192(form, mul192(block_eq_hi[bk:bk + 3], add192(beta, inner))) - e = 3 * (ETA_FORM_BASE + sd) - constraint_eval = add192(constraint_eval, mul192(zc_xi_pows[e:e + 3], form)) - joins = (g_zc_n / tau_g) ** 3 - sits = tau_g ** 3 - air_acc = add192(air_acc, mul192(mul192(zc_cprod[joins:joins + 3], zc_peq[sits:sits + 3]), constraint_eval)) # cprod[n - tau] * peq[tau] - assert_eq192(air_acc, claim) - - # ---- public-input binding claim ---- - # The VM's bind_pi_claim: the committed MEM at (r0, r1, 0, ...) equals the - # multilinear extension of the four public words at (r0, r1), low variable - # first. Both sides know the words, so the value is computed and nothing rides - # the stream. - fs, r0 = squeeze(fs) - fs, r1 = squeeze(fs) - v_lo = add64(scale(pi[0] + pi[1], r0), pi[0]) - v_hi = add64(scale(pi[2] + pi[3], r0), pi[2]) - p = 3 * claim_idx - claim_pool[p:p + 3] = add192(v_lo, mul192(r1, add192(v_lo, v_hi))) - return fs, cursor, g_zc_n, bc_share, r0, r1 - - -def verify_flock(fs: StackBuf(4), cursor, tau_blake2s_g, zerocheck_chis, lincheck_rs, z_partial): - # Flock's zerocheck (univariate skip, k_skip = 6) then its lincheck, whose - # matrix evaluation is DEFERRED to the caller's statement. The three run buffers - # come in pre-sized; the point z, lincheck's alpha and the deferred matrix part - # come back with the walked Fiat-Shamir state and stream cursor. - # - # tau's reach is bounded: the count gadget gives tau < 34 (every flock buffer is - # sized for that), and q_flock's committed kappa = K_LOG + tau feeds the - # certified size m, whose dispatch bound caps tau below any baked structure. The - # first K_SKIP Boolean rounds are replaced by the univariate skip and consume no - # equality challenges; the remaining r coordinates are N_FIXED_CHALLENGE_ROUNDS - # fixed inner values then sampled outer ones. The prover builds round 1 from - # this equality tail, so its sampled part is squeezed before round 1 is fetched, - # and round 1 before z, which evaluates it. - mr1cs_g = tau_blake2s_g * GEN ** K_LOG # runtime m = K_LOG + tau_5, in the exponent - zerocheck_r = HeapBuf(mr1cs_g ** 3) - flock_pts = HeapBuf((mr1cs_g * GEN ** 2) ** FS_SLOTS) - seed = flock_pts * (GEN ** (K_SKIP + N_FIXED_CHALLENGE_ROUNDS)) ** FS_SLOTS - seed[0:4] = fs - for xi in mul_range(GEN ** (K_SKIP + N_FIXED_CHALLENGE_ROUNDS), mr1cs_g): - row = flock_pts * xi ** FS_SLOTS - point_fs, zerocheck_challenge = squeeze(row[0:4]) - x3 = xi ** 3 - zerocheck_r[x3:x3 + 3] = zerocheck_challenge - row[4:8] = point_fs - pts_last = flock_pts * mr1cs_g ** FS_SLOTS - fs = pts_last[0:4] - # round-1 message (P = P^AB + P^C on Lambda, 2^K_SKIP scalars): fetch + - # observe each as it comes off the stream, then sample z. - zc_round1 = StackBuf(3 * 2 ** K_SKIP) - for i in unroll(0, 2 ** K_SKIP): - fs, w, cursor = fs_next(fs, cursor) - k = 3 * i - zc_round1[k:k + 3] = w - fs, zerocheck_z = squeeze(fs) # cursor now sits at the multilinear round messages, walked below - # P(z), interpolated at z over ALL 128 phi8 nodes: the transmitted Lambda values - # (nodes 64..128) plus the S half, zero by the zerocheck identity. The finished - # sum is scaled once by the domain's inverse denominator; the full-domain - # product only adds the S-half factor to the Lambda numerators. - lagrange_nums = StackBuf(3 * 2 ** K_SKIP) - lag64(zerocheck_z, lagrange_nums, 2 ** K_SKIP) - s_half_product = ONE - zc_running = ZERO # the zerocheck running claim entering the multilinear rounds - for i in unroll(0, 2 ** K_SKIP): - k = 3 * i - s_half_product = mul192(s_half_product, add192(zerocheck_z, phi8(i))) - zc_running = add192(zc_running, mul192(lagrange_nums[k:k + 3], zc_round1[k:k + 3])) - zc_running = mul192(zc_running, mul192(s_half_product, LAGRANGE_INV_COMBINED)) - for i in unroll(0, N_FIXED_CHALLENGE_ROUNDS): - fs, g_1, cursor = fs_next(fs, cursor) # G's coefficients, bar the constant one - fs, g_2, cursor = fs_next(fs, cursor) - g_0 = add192(zc_running, mul192(fixed_challenge(i), add192(g_1, g_2))) # the eq-weighted split fixes it - fs, chi_v = squeeze(fs) - k = 3 * i - zerocheck_chis[k:k + 3] = chi_v - zc_running = add192(g_0, mul192(chi_v, add192(g_1, mul192(chi_v, g_2)))) - # the sampled rounds: K_LOG + tau_5 - K_SKIP in all, certified - nmlv_g = tau_blake2s_g * GEN ** (K_LOG - K_SKIP) - flock_rounds = HeapBuf((nmlv_g * GEN ** 2) ** ROUND_SLOTS) - seed = flock_rounds * (GEN ** N_FIXED_CHALLENGE_ROUNDS) ** ROUND_SLOTS - seed[0:4] = fs - seed[GEN ** ROUND_CURSOR] = cursor - seed[ROUND_CLAIM:ROUND_CLAIM + 3] = zc_running - for xi in mul_range(GEN ** N_FIXED_CHALLENGE_ROUNDS, nmlv_g): - rd = flock_rounds * xi ** ROUND_SLOTS - x3 = xi ** 3 - r_at = x3 * GEN ** (3 * K_SKIP) - round_fs, g_1, cur_i = fs_next(rd[0:4], rd[GEN ** ROUND_CURSOR]) # coefficients, bar the constant one - round_fs, g_2, cur_i = fs_next(round_fs, cur_i) - g_0 = add192(rd[ROUND_CLAIM:ROUND_CLAIM + 3], mul192(zerocheck_r[r_at:r_at + 3], add192(g_1, g_2))) # the eq-weighted split fixes it - round_fs, chi_v = squeeze(round_fs) - zerocheck_chis[x3:x3 + 3] = chi_v - nxt = rd * GEN ** ROUND_SLOTS - nxt[0:4] = round_fs - nxt[GEN ** ROUND_CURSOR] = cur_i - nxt[ROUND_CLAIM:ROUND_CLAIM + 3] = add192(g_0, mul192(chi_v, add192(g_1, mul192(chi_v, g_2)))) - fr_last = flock_rounds * nmlv_g ** ROUND_SLOTS - fs = fr_last[0:4] - zc_running = fr_last[ROUND_CLAIM:ROUND_CLAIM + 3] - cursor = fr_last[GEN ** ROUND_CURSOR] # walked past all 2*n_mlv round scalars, now at a_eval - # final: observe a_eval, b_eval; the terminal identity is what defines - # c_eval, so nothing is checked here. C rode the rounds above, so all three - # claims sit at the same point and lincheck pins all three at once. - fs, a_eval, cursor = fs_next(fs, cursor) - fs, b_eval, cursor = fs_next(fs, cursor) - c_eval = add192(zc_running, mul192(a_eval, b_eval)) - # The phi8 Lagrange weights at z over the S nodes: the quirky extension's - # own combination, which the lincheck terminal applies to the 64 slices. - claim_nums = StackBuf(3 * 2 ** K_SKIP) - lag64(zerocheck_z, claim_nums, 0) - - # ---- flock lincheck (matrix evaluation DEFERRED) ---- - matrix_eval = StackBuf(3) - hint_witness(matrix_eval, "matpart") - fs, lincheck_alpha = squeeze(fs) - lincheck_beta = mul192(lincheck_alpha, lincheck_alpha) - lincheck_cube = mul192(lincheck_beta, lincheck_alpha) - # seed: a + alpha*b + alpha^2*c + alpha^3 (the two matrix claims, C, and the pin) - lc_running = add192(add192(a_eval, mul192(lincheck_alpha, b_eval)), add192(mul192(lincheck_beta, c_eval), lincheck_cube)) - for i in unroll(0, LINCHECK_ROUNDS): - fs, c0, cursor = fs_next(fs, cursor) # q's coefficients, bar the linear one - fs, c2, cursor = fs_next(fs, cursor) - c1 = add192(lc_running, c2) # the split fixes it against the running claim - fs, rv = squeeze(fs) - k = 3 * i - lincheck_rs[k:k + 3] = rv - lc_running = add192(c0, mul192(rv, add192(c1, mul192(rv, c2)))) # fold the degree-2 round poly at the challenge rv - # post-sumcheck collapse: fetch + observe each scalar - for i in unroll(0, 2 ** K_SKIP): - fs, w, cursor = fs_next(fs, cursor) - k = 3 * i - z_partial[k:k + 3] = w - # final consistency: running == matpart (DEFERRED) + beta * pin term. The - # const-pin column folds through the top-variable bindings: weight = - # prod_j (bit_{klog-1-j}(PIN_COLUMN) ? r_j : 1+r_j), surviving z_partial index - # = PIN_COLUMN low 6 bits. - pin = 3 * (PIN_COLUMN % 2 ** K_SKIP) - pin_term = mul192(mul192(lincheck_cube, eq_weight(lincheck_rs, LINCHECK_ROUNDS, PIN_COLUMN, K_LOG)), z_partial[pin:pin + 3]) - # The C term. Its column weight is the row weight itself (C = I), and both - # sides are tensors, so it collapses to eq(chi_in, chi_in_prime) times the - # phi8 Lagrange combination of the 64 slices: no second matrix walk, no - # second family. - c_point_eq = ONE - for t in unroll(0, LINCHECK_ROUNDS): - zk = 3 * t - lk = 3 * (LINCHECK_ROUNDS - 1 - t) - c_point_eq = mul192(c_point_eq, add192(ONE, add192(zerocheck_chis[zk:zk + 3], lincheck_rs[lk:lk + 3]))) - c_slice_value = ZERO - for i in unroll(0, 2 ** K_SKIP): - k = 3 * i - c_slice_value = add192(c_slice_value, mul192(claim_nums[k:k + 3], z_partial[k:k + 3])) - c_slice_value = mul192(c_slice_value, LAGRANGE_INV_S) - # deferred matrix eval + pin + C - assert_eq192(lc_running, add192(add192(matrix_eval, pin_term), mul192(mul192(lincheck_beta, c_point_eq), c_slice_value))) - # z_partial IS the claim: the terminal identity above pins its 64 slices, and - # ring switching binds every one of them. - return fs, cursor, zerocheck_z, lincheck_alpha, matrix_eval - - -def certify_placement(kappa_base, g_squares): - # Certify the native order: descending kappa, then ascending column index. - # Accumulate g^offset, with each column advancing it by g^(2^kappa). - col_kappa_g = HeapBuf(N_COMMITTED_COLS) - for c in unroll(0, N_COMMITTED_COLS): - col_kappa_g[GEN ** c] = kappa_base[GEN ** COL_KAPPA_SRC[c]] * GEN ** COL_KAPPA_ADJ[c] - col_sort_order = HeapBuf(N_COMMITTED_COLS) - hint_witness(col_sort_order[0:N_COMMITTED_COLS], "col_sort_order") - col_off_g = HeapBuf(N_COMMITTED_COLS) - g_total = GEN ** 0 - prev_col = GEN ** 0 - prev_kappa = GEN ** 0 - for rank in unroll(0, N_COMMITTED_COLS): - col = col_sort_order[GEN ** rank] - assert log(col) < N_COMMITTED_COLS - kappa_g = col_kappa_g[col] - if rank != 0: - # A negative exponent wraps around the order of GEN and cannot pass this - # small range check, hence prev_kappa >= kappa. - assert log(prev_kappa / kappa_g) < SIZE_BITS - if prev_kappa == kappa_g: - # Equal-sized columns use their compact (native column-order) index - # as the deterministic ascending tie-break. - assert log(col / prev_col) < N_COMMITTED_COLS - col_off_g[col] = g_total # write-once: a duplicate permutation entry collides - # g_squares spans SIZE_BITS, and every kappa is under it: a certified log - # <= 32 (log_mem or a tau), the baked bytecode log, or q_flock's tau_5 + 8. - # Bound it before the lookup, independently of m derived from this product. - assert log(kappa_g) < SIZE_BITS - g_total *= g_squares[kappa_g] - prev_col = col - prev_kappa = kappa_g - - return g_total, col_off_g, col_kappa_g - - -def column_selector(offset, point, kappa: Const): - # Rebuild the offset using only bits above the column's kappa, certifying its - # alignment. The same bits select the column at the complete opening point. - # Both offset and point are zero above m, so extending to MAX_STACK_LOG adds - # only factors eq(0, 0) = 1. - bits = StackBuf(MAX_STACK_LOG) - hint_decompose_bits_exponent(bits, offset, MAX_STACK_LOG) - rebuilt = GEN ** 0 - selector = ONE - for k in unroll(kappa, MAX_STACK_LOG): - bit = bits[k] - bits[k] = bit * bit - rebuilt *= 1 + bit * (1 + GEN ** (2 ** k)) - c = 3 * k - selector = mul192(selector, add64(point[c:c + 3], 1 + bit)) - assert rebuilt == offset - return selector - - -def check_opening_terminal(zeta, chi, r0: StackBuf(3), r1: StackBuf(3), g_bus_mu, g_zc_n, g_log_mem, tau_blake2s_g, claim_cplen_g, lam_pool, col_offsets, col_kappas, z_vals, c_table, point, inner_total: StackBuf(3), yr_at_tail: StackBuf(3), sumcheck_target: StackBuf(3)): - # Evaluate each transparent weight at the complete point in witness order. - # A claim is its low point followed by the certified column's selector bits; - # q_flock slots prepend their fixed slot bits to the low point. - zeta_eq_chain = HeapBuf(3 * (SIZE_BITS + 1)) - eq_prefix_chain(zeta_eq_chain, ONE, zeta, point, g_bus_mu) - chi_eq_chain = HeapBuf(3 * (SIZE_BITS + 1)) - eq_prefix_chain(chi_eq_chain, ONE, chi, point, g_zc_n) - chi_slot_eq_chain = HeapBuf(3 * (SIZE_BITS + 1)) - eq_prefix_chain(chi_slot_eq_chain, ONE, chi, point * GEN ** (3 * SLOT_STRIDE_LOG), tau_blake2s_g) - # The PI claim's point is (r0, r1, 0, ..., 0) over log_mem coordinates. - pi_chain = HeapBuf(3 * (SIZE_BITS + 1)) - pi_chain[3:6] = add192(ONE, add192(r0, point[0:3])) - pi_chain[6:9] = mul192(pi_chain[3:6], add192(ONE, add192(r1, point[3:6]))) - for xk in mul_range(GEN ** 2, g_log_mem): - x3 = xk ** 3 - nxt = x3 * GEN ** 3 - pi_chain[nxt:nxt + 3] = mul192(pi_chain[x3:x3 + 3], add192(ONE, point[x3:x3 + 3])) - pi_end = g_log_mem ** 3 - pi_eq = pi_chain[pi_end:pi_end + 3] - - selectors = StackBuf(3 * N_COMMITTED_COLS) - for c in unroll(0, N_COMMITTED_COLS): - offset = col_offsets[GEN ** c] - kappa_g = col_kappas[GEN ** c] - selector = match(log(kappa_g), range(0, N_COLUMN_LOGS), lambda kappa: column_selector(offset, point, kappa)) - k = 3 * c - selectors[k:k + 3] = selector - - inner_sum = inner_total - for j in unroll(0, N_CLAIMS): - if CLAIM_POINT_BUF[j] == POINT_BUF_PI: - cplen_g = g_log_mem - low_eq = pi_eq - else: - cplen_g = claim_cplen_g[GEN ** j] - nlow = cplen_g - low_at = cplen_g ** 3 - if CLAIM_POINT_BUF[j] == POINT_BUF_ZETA: - low_eq = zeta_eq_chain[low_at:low_at + 3] - if CLAIM_POINT_BUF[j] == POINT_BUF_RHO: - low_eq = chi_eq_chain[low_at:low_at + 3] - if CLAIM_POINT_BUF[j] == POINT_BUF_QFLOCK_RHO: - slot_eq = ONE - for k in unroll(0, SLOT_STRIDE_LOG): - pk = 3 * k - if CLAIM_QFLOCK_SLOT_BITS[SLOT_STRIDE_LOG * j + k] == 1: - slot_eq = mul192(slot_eq, point[pk:pk + 3]) - else: - slot_eq = mul192(slot_eq, add192(ONE, point[pk:pk + 3])) - low_eq = mul192(slot_eq, chi_slot_eq_chain[low_at:low_at + 3]) - nlow = cplen_g * GEN ** SLOT_STRIDE_LOG - assert nlow == col_kappas[GEN ** CLAIM_COMMITTED_COL[j]] - lk = 3 * j - sk = 3 * CLAIM_COMMITTED_COL[j] - inner_sum = add192(inner_sum, mul192(mul192(lam_pool[lk:lk + 3], low_eq), selectors[sk:sk + 3])) - - qflockv_g = tau_blake2s_g * GEN ** SLOT_STRIDE_LOG - assert qflockv_g == col_kappas[GEN ** QFLOCK_COMMITTED_COL] - prod_chains = HeapBuf((qflockv_g * GEN) ** (3 * BASE_FIELD_BITS)) - for k in unroll(0, BASE_FIELD_BITS): - c = 3 * k - prod_chains[c:c + 3] = ONE - rs_eq_run(prod_chains, z_vals, point, qflockv_g) - prod_final = prod_chains * qflockv_g ** (3 * BASE_FIELD_BITS) - rs_weight = ZERO - for k in unroll(0, BASE_FIELD_BITS): - c = 3 * k - rs_weight = add192(rs_weight, mul192(c_table[c:c + 3], prod_final[c:c + 3])) - qk = 3 * QFLOCK_COMMITTED_COL - inner_sum = add192(inner_sum, mul192(rs_weight, selectors[qk:qk + 3])) - assert_eq192(mul192(inner_sum, yr_at_tail), sumcheck_target) - return - - -def verify_sub(pi: StackBuf(4), seed, g_logs_pow2, g_squares, defer_out): - # In-circuit verification of ONE inner proof for the statement `pi`, mirroring - # cpu::verify step for step; the `# ---- ... ----` headers below run in that - # order. All proof data is hinted HERE, so each call pops the next sub-proof's - # entry of every witness stream and the body lowers once. The exponent tables - # are shared read-only; the deferred claims go to `defer_out`, three words an - # element. - # - # The pool holds every committed-coordinate claim's value in decompose order - # (the points are the GKR zetas, resolvable from the baked block structure) and - # its certified low dimension, which the terminal pins its lengths against. - claim_pool = HeapBuf(3 * N_CLAIMS) - claim_cplen_g = HeapBuf(N_CLAIMS) - - # ---- seed (statement pre-bound: hinted sub pi + baked program digest) ---- - fs = StackBuf(4) - blake2s(seed[0:4], pi, fs) - stream = HeapBuf(STREAM_CAP) - hint_witness(stream[0:STREAM_CAP], "stream") - cursor = stream # the proof stream, replayed scalar by scalar (advance = * g^3) - - # ---- announced layout and PCS rate (observed, then certified) ---- - # The stream announces the sizes as integer scalars, one log each, whose upper - # limbs must be zero; the shape-generic phases need them as G-POWERS (loop - # bounds, match scrutinees), so each is reassembled from its advice-decomposed - # bits, with no hint and no g^j -> j lookup. Every table's rows are real rows - # (the prover's fill blocks bring each count up to a power of two), so a height - # is all there is to announce. - sizes = StackBuf(N_TABLES + 2) - for i in unroll(0, N_TABLES + 2): - fs, x, cursor = fs_next(fs, cursor) - assert x[1] == 0 - assert x[2] == 0 - sizes[i] = x[0] - rate_sel = g_power_of_word(sizes[N_TABLES + 1], g_squares, LOG_WORD_BITS) / GEN - assert log(rate_sel) < LIG_N_RATES - dims_g = HeapBuf(N_TABLES + 1) # [g^log_mem, g^tau_0 .. g^tau_{N_TABLES-1}] - g_log_mem = g_power_of_word(sizes[0], g_squares, LOG_WORD_BITS) - assert log(g_log_mem) < COUNT_BITS - assert log(g_log_mem / GEN ** MIN_LOG_MEM) < COUNT_BITS # native MIN_LOG_MEM <= log_mem - dims_g[GEN ** 0] = g_log_mem - for t in unroll(0, N_TABLES): - g_tau = g_power_of_word(sizes[t + 1], g_squares, LOG_WORD_BITS) - assert log(g_tau) < COUNT_BITS - # A table's floor: flock sizes its BLAKE2s argument to at least 2^3 instances. - assert log(g_tau / GEN ** FLOORS[t]) < COUNT_BITS - dims_g[GEN ** (t + 1)] = g_tau - # kappa_base maps a kappa source index to its certified announced log (source 0 - # = const via the baked adj). Each block's kappa then DERIVES from its - # structural source as a compile-time offset off a certified log: no hint, and - # nothing left free. - kappa_base = HeapBuf(N_TABLES + 2) - kappa_base[GEN ** 0] = 1 - kappa_base[GEN ** 1] = g_log_mem - for t in unroll(0, N_TABLES): - kappa_base[GEN ** (2 + t)] = dims_g[GEN ** (t + 1)] - block_kappa = HeapBuf(N_BLOCKS) - for b in unroll(0, N_BLOCKS): - block_kappa[GEN ** b] = kappa_base[GEN ** BLOCK_KAPPA_SRC[b]] * GEN ** BLOCK_KAPPA_ADJ[b] - # The ONE bus depth, COMPUTED (not hinted): mu = log2_ceil(Σ_b 2^κ_b) over - # PUSH's blocks; pull matches by pairing, the count tree is padded to it. - push_total = GEN ** 0 - for b in unroll(SIDE_BLOCK_START[PUSH_SIDE], SIDE_BLOCK_START[PUSH_SIDE + 1]): - push_total *= g_squares[block_kappa[GEN ** b]] # g^(sum of 2^kappa) - g_bus_mu = log2_ceil_in_the_exponent(push_total, g_logs_pow2, g_squares, 0, SIZE_BITS) - zeta = HeapBuf(g_bus_mu ** 3) # the ONE shared GKR point: exactly mu coords - - # ---- commitment root (two halves), kept for the opening phase ---- - # A non-canonical half is rejected here (merkle.rs `scalars_to_hash`); the level - # roots get the same treatment at their own read. - fs, root_a, cursor = fs_next_half(fs, cursor) - fs, root_b, cursor = fs_next_half(fs, cursor) - commit_root = [root_a[0], root_a[1], root_b[0], root_b[1]] - - # ---- bus challenges (F192 provides the soundness margin without grinding) ---- - # A tuple is fingerprinted multilinearly: slot x weighs eq(alphas, x), so a leaf - # factor has total degree N_TUPLE_BITS in the challenges and the aligned bytecode - # polynomial is read off at the challenge vector itself (doc sec:gp, sec:e2e-bc). - bus_alpha = HeapBuf(3 * N_TUPLE_BITS) - for t in unroll(0, N_TUPLE_BITS): - fs, av = squeeze(fs) - k = 3 * t - bus_alpha[k:k + 3] = av - fp_w = HeapBuf(3 * N_TUPLE_SLOTS) - for x in unroll(0, N_TUPLE_SLOTS): - k = 3 * x - fp_w[k:k + 3] = eq_weight(bus_alpha, N_TUPLE_BITS, x, 0) - fs, beta = squeeze(fs) - - # ---- ONE GKR grand product: push, pull, and count RLC-batched ---- - fs, cursor, claim_push, claim_pull, claim_count = verify_bus_gkr(fs, cursor, g_bus_mu, zeta) - - # ---- the bus leaves, the table sumcheck, and the public-input claim ---- - chi = HeapBuf(3 * SIZE_BITS) # chi[i] = the challenge that bound variable i - fs, cursor, g_zc_n, bc_share, r0, r1 = verify_tables(fs, cursor, pi, zeta, g_bus_mu, dims_g, block_kappa, g_squares, fp_w, beta, claim_pool, claim_cplen_g, chi, claim_push, claim_pull, claim_count) - - # ---- flock zerocheck and lincheck (the matrix evaluation is DEFERRED) ---- - tau_blake2s_g = dims_g[GEN ** (TABLE_BLAKE2s + 1)] # the BLAKE2s table's certified tau - zerocheck_chis = HeapBuf((tau_blake2s_g * GEN ** (K_LOG - K_SKIP)) ** 3) # m - 6 rounds - lincheck_rs = HeapBuf(3 * LINCHECK_ROUNDS) - z_partial = HeapBuf(3 * 2 ** K_SKIP) - fs, cursor, zerocheck_z, lincheck_alpha, matrix_eval = verify_flock(fs, cursor, tau_blake2s_g, zerocheck_chis, lincheck_rs, z_partial) - - # ---- stacked mixed opening: ring-switch front + claim combination ---- - # The ring-switch slices are z_partial, read and bound above; this block only - # binds them to the commitment. Compose six two-term F2-linear maps with shifts - # 32,16,8,4,2,1: their expansion has all 64 Frobenius terms soundness needs, - # while direct application costs 63 squarings and only six general - # multiplications. - map_challenges = HeapBuf(3 * 6) # len(RING_MAP_SHIFTS) - c_table = HeapBuf(3 * BASE_FIELD_BITS) - z_vals = HeapBuf(3 * QFLOCK_VARS_CAP) - for stage in unroll(0, len(RING_MAP_SHIFTS)): - fs, map_challenge = squeeze(fs) - k = 3 * stage - map_challenges[k:k + 3] = map_challenge - # Expand the same composition once for the later transparent-weight evaluation. - # Before shift d, the populated coefficients are exactly at multiples of 2d; the - # new branch fills the adjacent d-offset entries. - c_table[0:3] = ONE - for stage in unroll(0, len(RING_MAP_SHIFTS)): - shift = RING_MAP_SHIFTS[stage] - mk = 3 * stage - map_challenge = map_challenges[mk:mk + 3] - for slot in unroll(0, BASE_FIELD_BITS // (2 * shift)): - src = 3 * (slot * 2 * shift) - coefficient = c_table[src:src + 3] - for k in unroll(0, shift): - coefficient = mul192(coefficient, coefficient) - dst = 3 * (slot * 2 * shift + shift) - c_table[dst:dst + 3] = mul192(map_challenge, coefficient) - # Evaluate the claim and combine its 64 packing rows: the running x-power (a - # word) and the running sum ride one four-slot chain. - rs_chain = HeapBuf(4 * ((2 ** K_SKIP) + 1)) - rs_chain[GEN ** 0] = GEN ** 0 # x^i - rs_chain[1:4] = ZERO # the running sum - for x_round in mul_range(1, GEN ** (2 ** K_SKIP)): - x3 = x_round ** 3 - lin_eval = z_partial[x3:x3 + 3] - for stage in unroll(0, len(RING_MAP_SHIFTS)): - frobenius = lin_eval - for k in unroll(0, RING_MAP_SHIFTS[stage]): - frobenius = mul192(frobenius, frobenius) - mk = 3 * stage - lin_eval = add192(lin_eval, mul192(map_challenges[mk:mk + 3], frobenius)) - row = rs_chain * x_round ** 4 - x_pow = row[GEN ** 0] - row[GEN ** 4] = x_pow * 2 - row[5:8] = add192(row[1:4], scale(x_pow, lin_eval)) - rs_end = rs_chain * (GEN ** (2 ** K_SKIP)) ** 4 - transposed_claim = rs_end[1:4] - # Suffix point for the transparent weight. - for t in unroll(0, LINCHECK_ROUNDS): - dst = 3 * t - src = 3 * (LINCHECK_ROUNDS - 1 - t) - z_vals[dst:dst + 3] = lincheck_rs[src:src + 3] - zv_lo = z_vals * GEN ** (3 * LINCHECK_ROUNDS) - zr_hi = zerocheck_chis * GEN ** (3 * LINCHECK_ROUNDS) - for xt in mul_range(1, tau_blake2s_g): - x3 = xt ** 3 - zv_lo[x3:x3 + 3] = zr_hi[x3:x3 + 3] - # ONE batching challenge for the whole pool: N_CLAIMS - 1 fewer Fiat-Shamir - # compressions than a challenge per claim, and none for the values themselves, - # `fs_next` having bound every one of them as it read it, so `lam_cl` already - # depends on all of them. Disjoint power ranges, as for the zc_xi-powers above: - # the ring-switch claim takes lam_cl^0, the pool lam_cl^1 onward. - fs, lam_cl = squeeze(fs) - target = transposed_claim - lam_pool = HeapBuf(3 * N_CLAIMS) - lam_pow = lam_cl - for j in unroll(0, N_CLAIMS): - k = 3 * j - lam_pool[k:k + 3] = lam_pow - target = add192(target, mul192(lam_pow, claim_pool[k:k + 3])) - lam_pow = mul192(lam_pow, lam_cl) - - g_total, col_offsets, col_kappas = certify_placement(kappa_base, g_squares) - - # ---- certify g^m: m = max(log2_ceil(sum_cols 2^kappa), PCS_MIN_MU) ---- - # g_total is g^(sum 2^kappa) from the certified placement walk above. - gmv = log2_ceil_in_the_exponent(g_total, g_logs_pow2, g_squares, PCS_MIN_MU, SIZE_BITS) # g^m - size_sel = gmv * LIG_MIN_SHIFT_INV # g^(m - MIN) - assert log(size_sel) < LIG_N_LOG_SIZES - # Flatten (rate-1, m-MIN) in rate-major order. Both coordinates are - # transcript-bound and range-checked above, so a single compiled guest can - # dispatch independently for every inner proof in a mixed-rate batch. - config_sel = size_sel * rate_sel ** LIG_N_LOG_SIZES - assert log(config_sel) < LIG_N_CANDIDATES - sumcheck_target, point, inner_total, yr_at_tail = match(log(config_sel), range(0, LIG_N_CANDIDATES), lambda m_idx: open_stacked(m_idx, fs, target, commit_root, cursor)) - # `stream` is a fixed-capacity witness transport. The shape fixes the exact - # consumed prefix, whose every word is transcript-bound; the unused suffix - # is outside the recursively verified proof and intentionally unconstrained. - - # ---- generalized eval_b terminal (runtime claim shapes) ---- - check_opening_terminal(zeta, chi, r0, r1, g_bus_mu, g_zc_n, g_log_mem, tau_blake2s_g, claim_cplen_g, lam_pool, col_offsets, col_kappas, z_vals, c_table, point, inner_total, yr_at_tail, sumcheck_target) - - # ---- export this sub-proof's deferred-claim data to the caller (FRESH_*) ---- - for k in unroll(0, BYTECODE_LOG): - c = 3 * k - defer_out[c:c + 3] = zeta[c:c + 3] - for k in unroll(0, LOG2_BYTECODE_COLS): - src = 3 * k - dst = 3 * (BYTECODE_LOG + k) - defer_out[dst:dst + 3] = bus_alpha[src:src + 3] - k = 3 * FRESH_BC_VALUE - defer_out[k:k + 3] = bc_share - k = 3 * FRESH_ALPHA - defer_out[k:k + 3] = lincheck_alpha - k = 3 * FRESH_Z_SKIP - defer_out[k:k + 3] = zerocheck_z - for k in unroll(0, LINCHECK_ROUNDS): - c = 3 * k - dst = 3 * (FRESH_ZCHI + k) - defer_out[dst:dst + 3] = zerocheck_chis[c:c + 3] - dst = 3 * (FRESH_LINCHECK_RS + k) - defer_out[dst:dst + 3] = lincheck_rs[c:c + 3] - for k in unroll(0, 2 ** K_SKIP): - c = 3 * k - dst = 3 * (FRESH_Z_PARTIAL + k) - defer_out[dst:dst + 3] = z_partial[c:c + 3] - k = 3 * FRESH_MATPART - defer_out[k:k + 3] = matrix_eval - return - - -# ============================ XMSS signature verification =========================== -# One signature of its epoch group's (epoch, message), against the signer's public -# key at `pk[0:4]` = (merkle_root, public_param), two words each. A tweak is two -# words: a constant first word, and the index word shared by every hash at one -# epoch or, for a Merkle node, the parent's. Index table layout (at cell g^t): -# 0 : the epoch's index word (encoding, chains, wots-pk) -# 1 + l : the parent index word of Merkle level l < LOG_LIFETIME - - -def fill_xmss_epoch_tables(epoch, merkle_bits, index_table): - # The index words and the Merkle direction bits at `epoch`, shared by every XMSS - # signature this node verifies at it. One bit decomposition gives all three uses: - # a tweak's index field is the epoch (encoding, chain, wots-pk) or the parent - # index `epoch >> (l + 1)` at Merkle level l, and the direction bit at that - # level IS bit l. Booleanity is a write-once pin and the reconstruction ties the - # bits back to the epoch, which also bounds it to LOG_LIFETIME bits. SPHINCS - # shares none of this, deriving every tweak from the index its own digest picks, - # which is neither public nor shared between signers. - hint_decompose_bits(merkle_bits, epoch, LOG_LIFETIME) - bits = StackBuf(LOG_LIFETIME) - reconstructed = 0 - index = 0 - for b in unroll(0, LOG_LIFETIME): - bit = merkle_bits[GEN ** b] - merkle_bits[GEN ** b] = bit * bit - bits[b] = bit - reconstructed += bit * 2 ** b - index += bit * XM_INDEX_WEIGHT[b] - assert reconstructed == epoch - index_table[GEN ** 0] = index - # Merkle level lvl - 1 hashes the parent at `epoch >> lvl`: the epoch's bits from - # lvl up, each weighed lvl places down. The top level gets the empty sum. - for lvl in unroll(1, LOG_LIFETIME + 1): - parent = 0 - for b in unroll(lvl, LOG_LIFETIME): - parent += bits[b] * XM_INDEX_WEIGHT[b - lvl] - index_table[GEN ** lvl] = parent - return - - -def verify_sig(message, index_table, merkle_bits, pk): - pp = pk[2:4] - index = index_table[GEN ** 0] - - # Encoding digest D = BLAKE2s(tweak | pp | msg | randomness | zero-pad), 96 - # bytes: one full 64-byte block then a 32-byte final block, 24 bytes of - # randomness and the specified 8-byte zero pad, which is a zero word. - after_msg = StackBuf(4) - blake2s([XM_ENC_TWEAK, index, pp], message[0:4], after_msg, counter=64, final=0) - rand = StackBuf(3) - hint_witness(rand, "rand") - digest = StackBuf(4) - blake2s([rand, 0], [0, 0, 0, 0], digest, cv=after_msg, counter=96, final=1) - - # V WOTS chains. Per chain the digit is hinted in the exponent (g^{e_i}), range - # checked and dispatched once; arm k walks the remaining CHAIN_STEPS-k steps and - # returns the tip plus the digit literal. The product of the digits is the - # target sum (g^{Σe_i}); the digits, weighted by CHAIN_LENGTH^i inside their own - # digest word (DIGITS_PER_WORD digits a word, GF(2^64)'s monomial budget, each - # word's leftover top bits ground to zero by the signer), reconstruct D's first - # two words. - tips = StackBuf(TIP_WORDS) - digit_product = 1 - acc_lo = 0 - acc_hi = 0 - for i in unroll(0, V): - digit = hint_witness("digits") - assert log(digit) < CHAIN_LENGTH - start = StackBuf(2) - hint_witness(start, "chain_starts") - tips[2 * i], tips[2 * i + 1], e = match(log(digit), range(0, CHAIN_LENGTH), lambda k: walk(start[0], start[1], XM_CHAIN_TWEAK + i * CHAIN_LENGTH * XM_P_MUL, index, pp[0], pp[1], k)) - digit_product = digit_product * digit - term = e * CHAIN_LENGTH ** (i % DIGITS_PER_WORD) # e_i in its monomial subspace - if i // DIGITS_PER_WORD == 0: - acc_lo = acc_lo + term - else: - acc_hi = acc_hi + term - assert digit_product == GEN ** TARGET_SUM - assert acc_lo == digest[0] - assert acc_hi == digest[1] - - # WOTS public-key leaf = standard BLAKE2s over prefix + V tips: WOTS_PK_BLOCKS - # full blocks, carrying the chaining value between instructions. - leaf = StackBuf(4) - blake2s([XM_PK_TWEAK, index, pp], tips[0:4], leaf, counter=64, final=0) - for q in unroll(1, WOTS_PK_BLOCKS): - next_leaf = StackBuf(4) - blake2s(tips[8 * q - 4:8 * q], tips[8 * q:8 * q + 4], next_leaf, cv=leaf, counter=64 * (q + 1), final=(q + 1) // WOTS_PK_BLOCKS) - leaf = next_leaf - - # Merkle path from the leaf to the root: the epoch bit orders the two children at - # each level, and the tweak carries that level's parent index. - n0 = leaf[0] - n1 = leaf[1] - for lvl in unroll(0, LOG_LIFETIME): - sibling = StackBuf(2) - hint_witness(sibling, "siblings") - children = order_children(n0, n1, sibling, merkle_bits[GEN ** lvl]) - parent = StackBuf(4) - blake2s([XM_MERKLE_TWEAK + const((lvl + 1) * XM_P_MUL), index_table[GEN ** (lvl + 1)], pp], children, parent) - n0 = parent[0] - n1 = parent[1] - assert n0 == pk[1] - assert n1 == pk[GEN] - return - - -def walk(v0, v1, tw, index, pp0, pp1, k: Const): - # Walk WOTS chain steps k..CHAIN_STEPS-1: value' = H(tweak|pp, value|0). Step s's - # tweak is the chain's first word plus s sub-positions. - w0 = v0 - w1 = v1 - for s in unroll(k, CHAIN_STEPS): - out = StackBuf(4) - blake2s([tw + s * XM_P_MUL, index, pp0, pp1], [w0, w1, 0, 0], out, counter=48, final=1) - w0 = out[0] - w1 = out[1] - return w0, w1, k - - -# ========================== SPHINCS+ signature verification ========================= - - -@inline -def sp_bit_field(bits_ptr, off: Const, n: Const, pos: Const): - # The integer held by bits [off, off+n) of the digest, weighed into a tweak word - # at bit `pos`: a tweak field placed where the tweak wants it, one fused - # multiply-add a bit, whatever digest word the bits came from. - acc = 0 - for i in unroll(0, n): - acc += bits_ptr[GEN ** (off + i)] * 2 ** (pos + i) - return acc - - -def sp_walk(v0, v1, tw, tw1, pp0, pp1, k: Const): - # Walk chain steps k..SP_CHAIN_STEPS-1: value' = Th(P, tw_chain, value). - # `tw` already carries the type byte, the layer and 2^w*i, and `tw1` the - # position (tau, e), so step s's tweak is one addition of a compile-time literal. - w0 = v0 - w1 = v1 - for s in unroll(k, SP_CHAIN_STEPS): - out = StackBuf(4) - blake2s([tw + s * SP_P_MUL, tw1, pp0, pp1], [w0, w1, 0, 0], out, counter=48, final=1) - w0 = out[0] - w1 = out[1] - return w0, w1, k - - -def sp_ots_leaf(tw_lay, tw1, pp0, pp1, m0, m1): - # One layer's one-time verification: the encoding of `m` under the hinted - # counter, the V chains walked from the revealed values, and the leaf they hash - # to. `tw_lay` is the layer's first-word field and `tw1` the position word (tau, - # e); this function is called once per layer, so the V dispatch tables are - # compiled once for the whole scheme. - ctr = hint_witness("sp_counter") - ctr_bits = HeapBuf(GEN ** SP_COUNTER_BITS) - hint_decompose_bits(ctr_bits, ctr, SP_COUNTER_BITS) - bind_bits(ctr_bits, ctr, SP_COUNTER_BITS) # LE_32: four counter bytes, four of padding - - # D = Th(P, tw_enc, m | LE_32(c)), a 52-byte one-block hash. - digest = StackBuf(4) - blake2s([tw_lay + SP_TW_ENC, tw1, pp0, pp1], [m0, m1, ctr, 0], digest, counter=52, final=1) - - # The codeword, as in XMSS: each digit hinted in the exponent, range checked and - # dispatched once, arm k walking the remaining steps; the product of the digits - # is the target sum, and the digits weighted by 2^w within each digest word - # reconstruct D, which pins each word's leftover top bits to zero. - tips = StackBuf(SP_TIP_WORDS) - digit_product = 1 - acc_lo = 0 - acc_hi = 0 - for i in unroll(0, SP_V): - digit = hint_witness("sp_digits") - assert log(digit) < SP_CHAIN_LENGTH - start = StackBuf(2) - hint_witness(start, "sp_chain_starts") - tips[2 * i], tips[2 * i + 1], e = match(log(digit), range(0, SP_CHAIN_LENGTH), lambda k: sp_walk(start[0], start[1], tw_lay + SP_TW_CHAIN + i * SP_CHAIN_MUL, tw1, pp0, pp1, k)) - digit_product = digit_product * digit - term = e * SP_CHAIN_LENGTH ** (i % SP_DIGITS_PER_WORD) - if i // SP_DIGITS_PER_WORD == 0: - acc_lo = acc_lo + term - else: - acc_hi = acc_hi + term - assert digit_product == GEN ** SP_TARGET_SUM - assert acc_lo == digest[0] - assert acc_hi == digest[1] - - leaf = StackBuf(4) - blake2s([tw_lay + SP_TW_LEAF, tw1, pp0, pp1], tips[0:4], leaf, counter=64, final=0) - for q in unroll(1, SP_LEAF_BLOCKS): - next_leaf = StackBuf(4) - blake2s(tips[8 * q - 4:8 * q], tips[8 * q:8 * q + 4], next_leaf, cv=leaf, counter=64 * (q + 1), final=(q + 1) // SP_LEAF_BLOCKS) - leaf = next_leaf - return leaf[0], leaf[1] - - -def verify_sig_sphincs(signer): - # `signer` is one 8-word entry of the SPHINCS coverage table: the key's root and - # public parameter, then the message THAT signer signed. Where XMSS's message is - # one statement field for the whole node, a SPHINCS message rides its own slot, - # and the signer-set digest binds the two together. - pp = signer[2:4] - - # ---- the message digest, which chooses the few-time key ---- - # D = Truncate(H(tw_msg | P | rho | root | m)), 96 bytes in two blocks. - rho = StackBuf(2) - hint_witness(rho, "sp_rand") - prefix = StackBuf(4) - blake2s([SP_TW_MSG, 0, pp], [rho, signer[0:2]], prefix, counter=64, final=0) - digest = StackBuf(4) - blake2s(signer[4:8], [0, 0, 0, 0], digest, cv=prefix, counter=96, final=1) - - # The index and the k leaf indices are bit fields of that digest, so its bits are - # advice-decomposed here and bound word by word. Nothing else derives them: every - # tweak below is built from these bits. - bits = HeapBuf(GEN ** SP_BIT_CELLS) - for lane in unroll(0, SP_BIT_LANES): - lane_bits = bits * GEN ** (lane * BASE_FIELD_BITS) - hint_decompose_bits(lane_bits, digest[lane], BASE_FIELD_BITS) - bind_bits(lane_bits, digest[lane], BASE_FIELD_BITS) - - # The digest is admissible only if its last leaf index is zero, which is what - # lets the forest drop that tree. - for b in unroll(0, SP_A): - assert bits[GEN ** (SP_H + (SP_K - 1) * SP_A + b)] == 0 - - # ---- the few-time signature: one opened leaf per tree of the forest ---- - idx_tau = sp_bit_field(bits, 0, SP_H, SP_TAU_POS) - roots = StackBuf(2 * SP_N_FTS) - for kappa in unroll(0, SP_N_FTS): - leaf_off = SP_H + kappa * SP_A - secret = StackBuf(2) - hint_witness(secret, "sp_fts_secrets") - fts_leaf = StackBuf(4) - node_index = sp_bit_field(bits, leaf_off, SP_A, SP_J_POS) - blake2s([SP_TW_FTS_LEAF + kappa * SP_LAY_MUL, idx_tau + node_index, pp], [secret, 0, 0], fts_leaf, counter=48, final=1) - n0 = fts_leaf[0] - n1 = fts_leaf[1] - for level in unroll(0, SP_A): - sibling = StackBuf(2) - hint_witness(sibling, "sp_fts_paths") - children = order_children(n0, n1, sibling, bits[GEN ** (leaf_off + level)]) - parent = StackBuf(4) - if const(level + 1 == SP_A): - node_index = 0 - else: - # The index fits in one word; clearing its low bit makes division by GEN a right shift. - node_index = (node_index + bits[GEN ** (leaf_off + level)] * 2 ** SP_J_POS) / GEN - blake2s([SP_TW_FTS_NODE + kappa * SP_LAY_MUL + const((level + 1) * SP_P_MUL), idx_tau + node_index, pp], children, parent) - n0 = parent[0] - n1 = parent[1] - roots[2 * kappa] = n0 - roots[2 * kappa + 1] = n1 - fts_key = StackBuf(4) - blake2s([SP_TW_FTS_ROOTS, idx_tau, pp], roots[0:4], fts_key, counter=64, final=0) - for q in unroll(1, SP_ROOT_BLOCKS): - next_key = StackBuf(4) - blake2s(roots[8 * q - 4:8 * q], roots[8 * q:8 * q + 4], next_key, cv=fts_key, counter=64 * (q + 1), final=(q + 1) // SP_ROOT_BLOCKS) - fts_key = next_key - s0 = fts_key[0] - s1 = fts_key[1] - - # ---- the hypertree, bottom layer first ---- - # Layer lay signs what the layer below produced: the few-time key at the bottom, - # that layer's root above it, and the public key's root at the top. - for step in unroll(0, SP_D): - lay = SP_D - 1 - step - leaf_index_off = SP_SUFFIX[lay + 1] - tau_field = sp_bit_field(bits, SP_SUFFIX[lay], SP_H - SP_SUFFIX[lay], SP_TAU_POS) - node_index = sp_bit_field(bits, leaf_index_off, SP_HEIGHTS[lay], SP_J_POS) - n0, n1 = sp_ots_leaf(lay * SP_LAY_MUL, tau_field + node_index, pp[0], pp[1], s0, s1) - for level in unroll(0, SP_HEIGHTS[lay]): - sibling = StackBuf(2) - hint_witness(sibling, "sp_siblings") - children = order_children(n0, n1, sibling, bits[GEN ** (leaf_index_off + level)]) - parent = StackBuf(4) - if const(level + 1 == SP_HEIGHTS[lay]): - node_index = 0 - else: - node_index = (node_index + bits[GEN ** (leaf_index_off + level)] * 2 ** SP_J_POS) / GEN - blake2s([SP_TW_NODE + lay * SP_LAY_MUL + const((level + 1) * SP_P_MUL), tau_field + node_index, pp], children, parent) - n0 = parent[0] - n1 = parent[1] - s0 = n0 - s1 = n1 - assert s0 == signer[1] - assert s1 == signer[GEN] - return - - -# =========================== statements and the signer set ========================== - - -def statement_digest(seed, signers: StackBuf(4), da: StackBuf(4), defer): - # A node's statement, hashed to the four words the VM publishes, over the - # proving environment's Fiat-Shamir seed (at `seed`: flock's R1CS and this - # bytecode), the signer-set digest, the digest of its possibly empty DA root - # list, and the DEFER_STMT_CELLS deferred elements. A parent rebuilds a child's - # statement with this same call, over a signer-set digest it re-absorbed itself, - # which forces the child to be a proof of THIS bytecode over groups checked - # against the parent's own. - # - # The preimage is fixed-length, so a plain BLAKE2s beats the Fiat-Shamir chain: - # the seed and signer-set digest are the first block, then the DA digest and - # the deferred elements' limbs, zero-filled to a whole block. - words = StackBuf(8 * (STMT_BLOCKS - 1)) - words[0:4] = da - for p in unroll(0, DEFER_STMT_CELLS): - k = 3 * p - words[4 + k:4 + k + 3] = defer[k:k + 3] - for k in unroll(0, STMT_PAD): - words[4 + 3 * DEFER_STMT_CELLS + k] = 0 - st = StackBuf(4) - blake2s(seed[0:4], signers, st, counter=64, final=0) - for b in unroll(1, STMT_BLOCKS): - lo = 8 * (b - 1) - nxt = StackBuf(4) - blake2s(words[lo:lo + 4], words[lo + 4:lo + 8], nxt, cv=st, counter=64 * (b + 1), final=(b + 1) // STMT_BLOCKS) - st = nxt - return st - - -def keys_window(state: StackBuf(4), base, keys_ptr, x_q, g_squares): - # One window of an epoch group's key hash: SIGNERS_WINDOW blocks, two declared - # keys each (a key is four words, a block eight). Counters as in `sphincs_window`. - nxt = scaled_log(x_q * GEN, g_squares, const(6 + SIGNERS_WINDOW_LOG)) - st = state - for j in unroll(0, SIGNERS_WINDOW): - pair = keys_ptr * GEN ** (8 * j) - hint_witness(pair[0:8], "pubkeys") - out = StackBuf(4) - if const(j + 1 == SIGNERS_WINDOW): - blake2s(pair[0:4], pair[4:8], out, cv=st, md=[nxt, 0]) - else: - blake2s(pair[0:4], pair[4:8], out, cv=st, md=[base + const(64 * (j + 1)), 0]) - st = out - return st, nxt - - - - -def list_chunk(state: StackBuf(4), off, ptr, cover, marks, origin_g, limit_g, size: Const, kind: Const): - # `size` non-final blocks of a list hash, block j's counter `off + 64(j + 1)`, - # its contents by list kind (`tail_blocks`). - st = state - for j in unroll(0, size): - out = StackBuf(4) - md = [off + const(64 * (j + 1)), 0] - if const(kind == LIST_CHILD_KEYS): - two = StackBuf(2) - hint_witness(two, "child_index") - assert log(two[0]) < log(limit_g) - assert log(two[1]) < log(limit_g) - cover[origin_g * two[0]] = marks * GEN ** (2 * j) - cover[origin_g * two[1]] = marks * GEN ** (2 * j + 1) - key_a = ptr * two[0] ** 4 - key_b = ptr * two[1] ** 4 - blake2s(key_a[0:4], key_b[0:4], out, cv=st, md=md) - if const(kind == LIST_CHILD_SPHINCS): - off_hint = hint_witness("child_sphincs_index") - assert log(off_hint) < log(limit_g) - cover[origin_g * off_hint] = marks * GEN ** j - entry = ptr * off_hint ** 8 - blake2s(entry[0:4], entry[4:8], out, cv=st, md=md) - if const(kind // 3 == 0): - block = ptr * GEN ** (8 * j) - if const(kind == LIST_KEYS): - hint_witness(block[0:8], "pubkeys") - if const(kind == LIST_SPHINCS): - hint_witness(block[0:8], "sphincs_signers") - blake2s(block[0:4], block[4:8], out, cv=st, md=md) - st = out - return st - - -def tail_blocks(state: StackBuf(4), base, ptr, cover, marks, origin_g, limit_g, tail_g, kind: Const): - # The blocks a list hash has left past its last whole window, fewer than a - # window: tail_g = g^k, absorbed in power-of-two chunks, largest first, one per - # set bit of k. A chunk starts at a multiple of twice its size, so every counter - # offset 64(start + j + 1) in it is still one XOR onto the running base, and one - # chunk per size serves every k. Returns the state, then the list pointer and the - # coverage marks stepped past the tail. - bits = StackBuf(SIGNERS_WINDOW_LOG) - hint_decompose_bits_exponent(bits, tail_g, SIGNERS_WINDOW_LOG) - states = StackBuf(4 * (SIGNERS_WINDOW_LOG + 1)) - states[0:4] = state - rebuilt = GEN ** 0 - off = base - for i in unroll(0, SIGNERS_WINDOW_LOG): - b = SIGNERS_WINDOW_LOG - 1 - i - bit = bits[b] - bits[b] = bit * bit # booleanity, as a write-once pin - rebuilt *= 1 + bit * (GEN ** (2 ** b) + 1) - cur = 4 * i - if bit != 0: - states[cur + 4:cur + 8] = list_chunk(states[cur:cur + 4], off, ptr, cover, marks, origin_g, limit_g, 2 ** b, kind) - else: - states[cur + 4:cur + 8] = states[cur:cur + 4] - off = off + bit * const(64 * 2 ** b) - if const(kind // 3 == 0): - ptr = ptr * (1 + bit * (GEN ** (8 * 2 ** b) + 1)) - if const(kind == LIST_CHILD_KEYS): - marks = marks * (1 + bit * (GEN ** (2 * 2 ** b) + 1)) - if const(kind == LIST_CHILD_SPHINCS): - marks = marks * (1 + bit * (GEN ** (2 ** b) + 1)) - assert rebuilt == tail_g - last = 4 * SIGNERS_WINDOW_LOG - return states[last:last + 4], ptr, marks - - -def key_list_digest(keys_ptr, half_g, odd_g, n_keys_g, g_squares): - # BLAKE2s of one epoch group's declared key list: 32 bytes a key, so the hashed - # string is 32·n bytes and its last block is the only partial one. The n // 2 - # pairs and the odd key out make half + odd blocks; all but the last run in - # windows plus a tail (doc §sec:prog-byte-counter), and the last carries the - # total length as its counter and the final-block flag. - split = StackBuf(2) - hint_witness(split, "signers_split") # g^windows, g^tail_blocks - windows = split[0] - tail = split[1] - assert log(tail) < SIGNERS_WINDOW - assert log(windows) < SIGNERS_MAX_WINDOWS - assert windows ** SIGNERS_WINDOW * tail == half_g * odd_g * INV_GEN - chain = HeapBuf((windows * GEN) ** 6) # state, base, first key of the window - chain[0:4] = [BLAKE2S_IV_0, BLAKE2S_IV_1, BLAKE2S_IV_2, BLAKE2S_IV_3] - chain[GEN ** 4] = 0 - chain[GEN ** 5] = keys_ptr - for xq in mul_range(1, windows): - slot = chain * xq ** 6 - st, nb = keys_window(slot[0:4], slot[GEN ** 4], slot[GEN ** 5], xq, g_squares) - step = slot * GEN ** 6 - step[0:4] = st - step[GEN ** 4] = nb - step[GEN ** 5] = slot[GEN ** 5] * GEN ** (8 * SIGNERS_WINDOW) - end = chain * windows ** 6 - t, last, unused = tail_blocks(end[0:4], end[GEN ** 4], end[GEN ** 5], 0, 0, 0, 0, tail, LIST_KEYS) - final = [scaled_log(n_keys_g, g_squares, 5), MD_FINAL] - digest = StackBuf(4) - if odd_g == 1: - hint_witness(last[0:8], "pubkeys") - blake2s(last[0:4], last[4:8], digest, cv=t, md=final) - else: - # The odd key out fills half its block, the rest being the zero bytes the - # counter already accounts for. - hint_witness(last[0:4], "pubkeys") - blake2s(last[0:4], [0, 0, 0, 0], digest, cv=t, md=final) - return digest - - -def child_keys_window(state: StackBuf(4), base, keys_ptr, cover, marks, origin_g, limit_g, x_q, g_squares): - # One window of a child's key hash, absorbed exactly as the child absorbed it, - # but with both keys of a block read at hinted indices into THIS node's table and - # marked in the coverage table. Each index is an offset into the parent group the - # caller mapped this child group to, bounded by that group's size, so a child's - # key can only ever land on an XMSS slot of the right epoch. - nxt = scaled_log(x_q * GEN, g_squares, const(6 + SIGNERS_WINDOW_LOG)) - st = state - for j in unroll(0, SIGNERS_WINDOW): - two = StackBuf(2) - hint_witness(two, "child_index") - assert log(two[0]) < log(limit_g) # precondition as in the raw loops - assert log(two[1]) < log(limit_g) - cover[origin_g * two[0]] = marks * GEN ** (2 * j) - cover[origin_g * two[1]] = marks * GEN ** (2 * j + 1) - key_a = keys_ptr * two[0] ** 4 - key_b = keys_ptr * two[1] ** 4 - out = StackBuf(4) - if const(j + 1 == SIGNERS_WINDOW): - blake2s(key_a[0:4], key_b[0:4], out, cv=st, md=[nxt, 0]) - else: - blake2s(key_a[0:4], key_b[0:4], out, cv=st, md=[base + const(64 * (j + 1)), 0]) - st = out - return st, nxt - - - - -def child_key_list_digest(keys_ptr, cover, base, origin_g, limit_g, half_g, odd_g, n_keys_g, g_squares): - # BLAKE2s of one epoch group of a child's keys, over the same 32·n bytes the - # child hashed (`key_list_digest`), so the digest it rebuilds is the one the - # child's statement carries. `base` prefixes the coverage write values, which - # count the keys off as they are marked. - split = StackBuf(2) - hint_witness(split, "signers_split") - windows = split[0] - tail = split[1] - assert log(tail) < SIGNERS_WINDOW - assert log(windows) < SIGNERS_MAX_WINDOWS - assert windows ** SIGNERS_WINDOW * tail == half_g * odd_g * INV_GEN - chain = HeapBuf((windows * GEN) ** 6) # state, base, coverage marks - chain[0:4] = [BLAKE2S_IV_0, BLAKE2S_IV_1, BLAKE2S_IV_2, BLAKE2S_IV_3] - chain[GEN ** 4] = 0 - chain[GEN ** 5] = base - for xq in mul_range(1, windows): - slot = chain * xq ** 6 - st, nb = child_keys_window(slot[0:4], slot[GEN ** 4], keys_ptr, cover, slot[GEN ** 5], origin_g, limit_g, xq, g_squares) - step = slot * GEN ** 6 - step[0:4] = st - step[GEN ** 4] = nb - step[GEN ** 5] = slot[GEN ** 5] * GEN ** (2 * SIGNERS_WINDOW) - end = chain * windows ** 6 - t, unused, marks = tail_blocks(end[0:4], end[GEN ** 4], keys_ptr, cover, end[GEN ** 5], origin_g, limit_g, tail, LIST_CHILD_KEYS) - final = [scaled_log(n_keys_g, g_squares, 5), MD_FINAL] - digest = StackBuf(4) - if odd_g == 1: - two = StackBuf(2) - hint_witness(two, "child_index") - assert log(two[0]) < log(limit_g) - assert log(two[1]) < log(limit_g) - cover[origin_g * two[0]] = marks - cover[origin_g * two[1]] = marks * GEN - key_a = keys_ptr * two[0] ** 4 - key_b = keys_ptr * two[1] ** 4 - blake2s(key_a[0:4], key_b[0:4], digest, cv=t, md=final) - else: - tail_idx = hint_witness("child_index") - assert log(tail_idx) < log(limit_g) - cover[origin_g * tail_idx] = marks - key_last = keys_ptr * tail_idx ** 4 - blake2s(key_last[0:4], [0, 0, 0, 0], digest, cv=t, md=final) - return digest - - -def scaled_log(x, g_squares, shift: Const): - # 2^shift times the exponent of `x`, as a bit pattern (doc §sec:prog-byte-counter). - # The exponent's bits are advice, tied back by the g-power product; weighing them - # at 2^j assembles the exponent itself and the final multiply is the shift, - # exact because nothing reduces below degree 64. Both sides of the product stay - # under the order of g, so the bits ARE that exponent, hence below - # 2^SIGNERS_COUNT_BITS, which every count and window index here is. - bits = StackBuf(SIGNERS_COUNT_BITS) - hint_decompose_bits_exponent(bits, x, SIGNERS_COUNT_BITS) - value = 0 - rebuilt = GEN ** 0 - for j in unroll(0, SIGNERS_COUNT_BITS): - b = bits[j] - bits[j] = b * b # booleanity, as a write-once pin - value += b * 2 ** j - rebuilt *= (1 + b * (g_squares[GEN ** j] + 1)) - assert rebuilt == x - return value * 2 ** shift - - -def sphincs_window(state: StackBuf(4), base, entries_ptr, x_q, g_squares): - # One window of the SPHINCS list's hash: SIGNERS_WINDOW claims, one 64-byte block - # each (the claimed key, then the message it signed). Block j's counter is - # base + 64(j+1), one XOR, except the last, whose offset is the base's own lowest - # bit and which therefore takes the NEXT window's base as its whole counter. That - # base is derived here and carried out for the following window. - nxt = scaled_log(x_q * GEN, g_squares, const(6 + SIGNERS_WINDOW_LOG)) - st = state - for j in unroll(0, SIGNERS_WINDOW): - entry = entries_ptr * GEN ** (8 * j) - hint_witness(entry[0:8], "sphincs_signers") - out = StackBuf(4) - if const(j + 1 == SIGNERS_WINDOW): - blake2s(entry[0:4], entry[4:8], out, cv=st, md=[nxt, 0]) - else: - blake2s(entry[0:4], entry[4:8], out, cv=st, md=[base + const(64 * (j + 1)), 0]) - st = out - return st, nxt - - - - -def sphincs_list_digest(entries_ptr, n_g, g_squares): - # BLAKE2s of the declared SPHINCS claims: n blocks of 64 bytes, so the hash is - # over exactly 64n bytes and no block is partial. The last block is absorbed - # apart, carrying the total length as its counter and the final-block flag; the - # n - 1 before it run in windows plus a tail (doc §sec:prog-byte-counter). - digest = StackBuf(4) - if n_g == 1: - # No claims: the hash of the empty string, one compression of a zero block. - blake2s([0, 0, 0, 0], [0, 0, 0, 0], digest, md=[0, MD_FINAL]) - else: - split = StackBuf(2) - hint_witness(split, "signers_split") # g^windows, g^tail_blocks - windows = split[0] - tail = split[1] - assert log(tail) < SIGNERS_WINDOW - assert log(windows) < SIGNERS_MAX_WINDOWS - assert windows ** SIGNERS_WINDOW * tail == n_g * INV_GEN - # Six words a window: the state, the window's base, its first entry. - chain = HeapBuf((windows * GEN) ** 6) - chain[0:4] = [BLAKE2S_IV_0, BLAKE2S_IV_1, BLAKE2S_IV_2, BLAKE2S_IV_3] - chain[GEN ** 4] = 0 - chain[GEN ** 5] = entries_ptr - for xq in mul_range(1, windows): - slot = chain * xq ** 6 - st, nb = sphincs_window(slot[0:4], slot[GEN ** 4], slot[GEN ** 5], xq, g_squares) - step = slot * GEN ** 6 - step[0:4] = st - step[GEN ** 4] = nb - step[GEN ** 5] = slot[GEN ** 5] * GEN ** (8 * SIGNERS_WINDOW) - end = chain * windows ** 6 - t, last, unused = tail_blocks(end[0:4], end[GEN ** 4], end[GEN ** 5], 0, 0, 0, 0, tail, LIST_SPHINCS) - hint_witness(last[0:8], "sphincs_signers") - final = [scaled_log(n_g, g_squares, 6), MD_FINAL] - blake2s(last[0:4], last[4:8], digest, cv=t, md=final) - return digest - - -def child_sphincs_window(state: StackBuf(4), base, entries_ptr, cover, marks, origin_g, limit_g, x_q, g_squares): - # One window of a child's SPHINCS list, absorbed exactly as the child absorbed - # it, but with each block's claim read at a hinted index into THIS node's table - # and marked in the coverage table. The index is an offset into the SPHINCS - # region and bounded by that region's size, so a child's claim can only ever - # land on a SPHINCS slot. Counters as in `sphincs_window`. - nxt = scaled_log(x_q * GEN, g_squares, const(6 + SIGNERS_WINDOW_LOG)) - st = state - for j in unroll(0, SIGNERS_WINDOW): - off_hint = hint_witness("child_sphincs_index") - assert log(off_hint) < log(limit_g) # precondition as in the raw loops - cover[origin_g * off_hint] = marks * GEN ** j - entry = entries_ptr * off_hint ** 8 - out = StackBuf(4) - if const(j + 1 == SIGNERS_WINDOW): - blake2s(entry[0:4], entry[4:8], out, cv=st, md=[nxt, 0]) - else: - blake2s(entry[0:4], entry[4:8], out, cv=st, md=[base + const(64 * (j + 1)), 0]) - st = out - return st, nxt - - - - -def child_sphincs_list_digest(entries_ptr, cover, base, origin_g, limit_g, n_g, g_squares): - # BLAKE2s of a child's declared SPHINCS claims, over the same 64n bytes the - # child hashed (`sphincs_list_digest`), so the digest it rebuilds is the one the - # child's statement carries. `base` prefixes the coverage write values, which - # count the claims off as they are marked. - digest = StackBuf(4) - if n_g == 1: - blake2s([0, 0, 0, 0], [0, 0, 0, 0], digest, md=[0, MD_FINAL]) - else: - split = StackBuf(2) - hint_witness(split, "signers_split") - windows = split[0] - tail = split[1] - assert log(tail) < SIGNERS_WINDOW - assert log(windows) < SIGNERS_MAX_WINDOWS - assert windows ** SIGNERS_WINDOW * tail == n_g * INV_GEN - chain = HeapBuf((windows * GEN) ** 5) # state, base - chain[0:4] = [BLAKE2S_IV_0, BLAKE2S_IV_1, BLAKE2S_IV_2, BLAKE2S_IV_3] - chain[GEN ** 4] = 0 - for xq in mul_range(1, windows): - slot = chain * xq ** 5 - marks = base * xq ** SIGNERS_WINDOW - st, nb = child_sphincs_window(slot[0:4], slot[GEN ** 4], entries_ptr, cover, marks, origin_g, limit_g, xq, g_squares) - step = slot * GEN ** 5 - step[0:4] = st - step[GEN ** 4] = nb - end = chain * windows ** 5 - marks = base * windows ** SIGNERS_WINDOW - t, unused, unused_marks = tail_blocks(end[0:4], end[GEN ** 4], entries_ptr, cover, marks, origin_g, limit_g, tail, LIST_CHILD_SPHINCS) - off_hint = hint_witness("child_sphincs_index") - assert log(off_hint) < log(limit_g) - cover[origin_g * off_hint] = base * (n_g * INV_GEN) - entry = entries_ptr * off_hint ** 8 - final = [scaled_log(n_g, g_squares, 6), MD_FINAL] - blake2s(entry[0:4], entry[4:8], digest, cv=t, md=final) - return digest - - -def plain_window(state: StackBuf(4), base, run_ptr, x_q, g_squares): - # One window over a run of words already in memory, eight to a block, hinting - # nothing. Counters as in `sphincs_window`. - nxt = scaled_log(x_q * GEN, g_squares, const(6 + SIGNERS_WINDOW_LOG)) - st = state - for j in unroll(0, SIGNERS_WINDOW): - block = run_ptr * GEN ** (8 * j) - out = StackBuf(4) - if const(j + 1 == SIGNERS_WINDOW): - blake2s(block[0:4], block[4:8], out, cv=st, md=[nxt, 0]) - else: - blake2s(block[0:4], block[4:8], out, cv=st, md=[base + const(64 * (j + 1)), 0]) - st = out - return st, nxt - - - - -def signer_set_digest(run_ptr, n_epochs_g, g_squares): - # BLAKE2s of the signer set: both list lengths and the SPHINCS list's digest - # in the first block, then two blocks a group, its (epoch, count, message) - # and its key list's digest. Every block is full, so the hash is over exactly - # 64·(1 + 2·epochs) bytes, and leading with both lengths makes the encoding - # prefix-free: no set's string is a prefix of another's. - blocks = n_epochs_g * n_epochs_g * GEN # g^(1 + 2·epochs) - split = StackBuf(2) - hint_witness(split, "signers_split") - windows = split[0] - tail = split[1] - assert log(tail) < SIGNERS_WINDOW - assert log(windows) < SIGNERS_MAX_WINDOWS - assert windows ** SIGNERS_WINDOW * tail == blocks * INV_GEN - chain = HeapBuf((windows * GEN) ** 6) - chain[0:4] = [BLAKE2S_IV_0, BLAKE2S_IV_1, BLAKE2S_IV_2, BLAKE2S_IV_3] - chain[GEN ** 4] = 0 - chain[GEN ** 5] = run_ptr - for xq in mul_range(1, windows): - slot = chain * xq ** 6 - st, nb = plain_window(slot[0:4], slot[GEN ** 4], slot[GEN ** 5], xq, g_squares) - step = slot * GEN ** 6 - step[0:4] = st - step[GEN ** 4] = nb - step[GEN ** 5] = slot[GEN ** 5] * GEN ** (8 * SIGNERS_WINDOW) - end = chain * windows ** 6 - t, last, unused = tail_blocks(end[0:4], end[GEN ** 4], end[GEN ** 5], 0, 0, 0, 0, tail, LIST_PLAIN) - final = [scaled_log(blocks, g_squares, 6), MD_FINAL] - digest = StackBuf(4) - blake2s(last[0:4], last[4:8], digest, cv=t, md=final) - return digest - - -def rebuild_child_groups(nsub_e_g, run_ptr, base, epochs, msgs, group_base, group_slots, n_epochs_g, xmss_table, cover, g_squares): - # The child's epoch groups, written into the run its own signer-set hash covers, - # two blocks a group exactly as the child laid them out: its (epoch, count, - # message), then the digest of its keys, read from THIS node's table through - # hinted indices. A hinted map ties each group to the parent group holding the - # same epoch AND message, whose region its keys land in. Everything hinted here - # is pinned by the digest, which the child's statement carries. The running - # product of the group counts prefixes each group's coverage writes and ends as - # the child's XMSS claim count. - counts = HeapBuf(nsub_e_g * GEN) - counts[GEN ** 0] = 1 - for xj in mul_range(1, nsub_e_g): - grp = StackBuf(6) - hint_witness(grp, "child_group") # epoch, message (four words), count - n_keys = grp[5] - assert log(n_keys) < MAX_KEYS - parent = hint_witness("child_group_map") - assert log(parent) < log(n_epochs_g) - assert epochs[parent] == grp[0] - parent_msg = msgs * parent ** 4 - parent_msg[0:4] = grp[1:5] - halves = StackBuf(2) - hint_witness(halves, "child_halves") - assert log(halves[1]) < 2 - assert log(halves[0]) < MAX_KEYS - assert halves[0] * halves[0] * halves[1] == n_keys - gb = group_base[parent] - prefix = counts[xj] - kd = child_key_list_digest(xmss_table * gb ** 4, cover, base * prefix, gb, group_slots[parent], halves[0], halves[1], n_keys, g_squares) - slot = run_ptr * xj ** 16 * GEN ** 8 - slot[1] = grp[0] - slot[GEN] = 0 - slot[GEN ** 2] = n_keys - slot[GEN ** 3] = 0 - slot[4:8] = grp[1:5] - slot[8:12] = kd - slot[12:16] = [0, 0, 0, 0] - counts[xj * GEN] = prefix * n_keys - return counts[nsub_e_g] - - -# ================================ LeanDA =============================== -# Hash the encoded rows and the hinted membership vector, then check their inner products. -# The external verifier derives the vector hash from the root; recursion preserves both. - - -def da_verify(g_squares): - # Returns the matrix root and membership-vector hash, four words each. - # The row count is a run-time parameter. The trees need a compile-time depth, - # so the rows are padded to a power of two and the two `match`es below dispatch - # on its log; everything else walks the real rows only. A padding row is the - # zero codeword, so its cell digest and its row digest are constants, and the - # gap loop writes them without hashing anything. - shape = StackBuf(2) - hint_witness(shape, "da_shape") # n_blob, then log2 of the padded count - n_blob_g = shape[0] - g_log_pad = shape[1] - n_pad_g, gap_g = da_row_shape(n_blob_g, g_log_pad, g_squares) - - prefix_digests = HeapBuf(n_pad_g ** (4 * DA_PREFIX_CELLS)) - prefix_bases = HeapBuf(n_blob_g) - for xi in mul_range(1, n_blob_g): - prefix_bases[xi] = prefix_digests * xi ** (4 * DA_PREFIX_CELLS) - - hashes = HeapBuf(4 * (DA_CELLS + 1)) - hashes[0:4] = [BLAKE2S_IV_0, BLAKE2S_IV_1, BLAKE2S_IV_2, BLAKE2S_IV_3] - counter_values = StackBuf(3 * DA_CELLS + 1) - for w in unroll(0, 3 * DA_CELLS + 1): - counter_values[w] = const(w * 2 ** (DA_LOG_CELL + 3)) - counters = addr(counter_values) - - # Each row's running sum uses a fresh write-once element per column. - acc = HeapBuf(n_blob_g ** (3 * (DA_CELLS + 1))) - rowbase = HeapBuf(n_blob_g) - for xi in mul_range(1, n_blob_g): - chain = acc * xi ** (3 * (DA_CELLS + 1)) - rowbase[xi] = chain - chain[0:3] = ZERO - - # Prefix blocks first, so the digests the row branch needs are stored by a - # loop that knows it is inside the prefix, with no per-block branch. - coltree = HeapBuf(8 * DA_CELLS) - for xb in mul_range(1, GEN ** DA_PREFIX_CELLS): - da_verify_column(xb, n_blob_g, n_pad_g, gap_g, g_log_pad, prefix_bases, coltree, rowbase, hashes, counters, 1) - for xb in mul_range(GEN ** DA_PREFIX_CELLS, GEN ** DA_CELLS): - da_verify_column(xb, n_blob_g, n_pad_g, gap_g, g_log_pad, prefix_bases, coltree, rowbase, hashes, counters, 0) - - for xi in mul_range(1, n_blob_g): - chain = rowbase[xi] - k = 3 * DA_CELLS - assert_eq192(chain[k:k + 3], ZERO) - col_root = da_levels(coltree, DA_CELLS, DA_BLOCK_BITS) - - # The row branch: each row's prefix digests, hashed as one string. - rowtree = HeapBuf(n_pad_g ** 8) - for xi in mul_range(1, n_blob_g): - prefix = prefix_bases[xi] - st = StackBuf(4) - blake2s(prefix[0:4], prefix[4:8], st, counter=64, final=1 // DA_ROW_BLOCKS) - for b in unroll(1, DA_ROW_BLOCKS): - nxt = StackBuf(4) - blake2s(prefix[8 * b:8 * b + 4], prefix[8 * b + 4:8 * b + 8], nxt, cv=st, counter=64 * (b + 1), final=(b + 1) // DA_ROW_BLOCKS) - st = nxt - slot = rowtree * xi ** 4 - slot[0:4] = st - for xd in mul_range(1, gap_g): - pad = rowtree * n_blob_g ** 4 * xd ** 4 - pad[0:4] = [DA_PAD_ROW[0], DA_PAD_ROW[1], DA_PAD_ROW[2], DA_PAD_ROW[3]] - row_root = match(log(g_log_pad), range(0, DA_TREE_ARMS), lambda k: da_levels(rowtree, 2 ** k, k)) - - root = StackBuf(4) - blake2s(row_root, col_root, root) - vec = 4 * DA_CELLS - return root, hashes[vec:vec + 4] - - -def da_row_shape(n_blob_g, g_log_pad, g_squares): - assert n_blob_g != 1 - assert log(n_blob_g) < DA_MAX_ROWS + 1 - assert log(g_log_pad) < DA_LOG_MAX_ROWS + 1 - n_pad_g = g_squares[g_log_pad] - # n_blob <= n_pad < 2*n_blob: both differences must be nonnegative. - gap_g = n_pad_g / n_blob_g - assert log(gap_g) < DA_MAX_ROWS + 1 - assert log(n_blob_g * n_blob_g / (n_pad_g * GEN)) < DA_MAX_ROWS - return n_pad_g, gap_g - - -def da_verify_column(xb, n_blob_g, n_pad_g, gap_g, g_log_pad, prefix_bases, coltree, rowbase, hashes, counters, store: Const): - st = hashes * xb ** 4 - h, lvals = da_vector_dispatch(xb, st[0:4], counters) - nxt = st * GEN ** 4 - nxt[0:4] = h - node = HeapBuf(n_pad_g ** 8) - xb3 = xb ** 3 - for xi in mul_range(1, n_blob_g): - d, s = da_verify_cell(lvals) - chain = rowbase[xi] * xb3 - chain[3:6] = add192(chain[0:3], s) - leaf = node * xi ** 4 - leaf[0:4] = d - if const(store == 1): - slot = prefix_bases[xi] * xb ** 4 - slot[0:4] = d - for xd in mul_range(1, gap_g): - pad = node * n_blob_g ** 4 * xd ** 4 - pad[0:4] = [DA_PAD_CELL[0], DA_PAD_CELL[1], DA_PAD_CELL[2], DA_PAD_CELL[3]] - c = match(log(g_log_pad), range(0, DA_TREE_ARMS), lambda k: da_levels(node, 2 ** k, k)) - out = coltree * xb ** 4 - out[0:4] = c - return - - -def da_verify_cell(lvals): - # The symbols are words, so the hash and the dot product read them directly; the - # weights are elements, three words each, so the dot product runs limb by limb. - sym = StackBuf(DA_CELL) - hint_witness(sym, "da_symbols") - s0 = 0 - s1 = 0 - s2 = 0 - for e in unroll(0, DA_CELL): - k = 3 * e - sv = sym[e] - s0 += sv * lvals[GEN ** k] - s1 += sv * lvals[GEN ** (k + 1)] - s2 += sv * lvals[GEN ** (k + 2)] - st = StackBuf(4) - blake2s(sym[0:4], sym[4:8], st, counter=64, final=1 // DA_CELL_BLOCKS) - for b in unroll(1, DA_CELL_BLOCKS): - nxt = StackBuf(4) - blake2s(sym[8 * b:8 * b + 4], sym[8 * b + 4:8 * b + 8], nxt, cv=st, counter=64 * (b + 1), final=(b + 1) // DA_CELL_BLOCKS) - st = nxt - return st, [s0, s1, s2] - - -def da_levels(tree, n: Const, log_n: Const): - # The internal levels of a Merkle tree whose leaves already sit at level 0 of - # `tree` (4 words a node, level lvl at 8n - 8n//2**lvl). Returns the root. - for lvl in unroll(0, log_n): - for xp in mul_range(1, GEN ** (n // 2 ** (lvl + 1))): - a = tree * GEN ** (8 * n - 8 * n // 2 ** lvl) * xp ** 8 - b = tree * GEN ** (8 * n - 8 * n // 2 ** (lvl + 1)) * xp ** 4 - blake2s(a[0:4], a[4:8], b[0:4]) - root = 8 * n - 8 - return tree[root:root + 4] - - -def da_vector_dispatch(xb, h: StackBuf(4), counters): - out = StackBuf(4) - weights = StackBuf(1) - if xb == GEN ** (DA_CELLS - 1): - a, w = da_vector_cell(xb, h, counters, 1) - out[0:4] = a - weights[0] = w - else: - a, w = da_vector_cell(xb, h, counters, 0) - out[0:4] = a - weights[0] = w - return out, weights[0] - - -def da_vector_cell(xb, h: StackBuf(4), counters, final: Const): - weights = StackBuf(3 * DA_CELL) - hint_witness(weights, "da_weights") - st = h - # Three power-of-two byte windows per cell keep counter offsets disjoint from the base. - base = counters[xb ** 3] - for w in unroll(0, 3): - window = xb ** 3 * GEN ** w - end = counters[window * GEN] - for b in unroll(0, DA_CELL_BLOCKS): - out = StackBuf(4) - if const(b + 1 == DA_CELL_BLOCKS): - counter = end - else: - counter = base + const(64 * (b + 1)) - flags = 0 - if const(final == 1): - if const(w == 2): - if const(b + 1 == DA_CELL_BLOCKS): - flags = MD_FINAL - lo = w * DA_CELL + 8 * b - blake2s(weights[lo:lo + 4], weights[lo + 4:lo + 8], out, cv=st, md=[counter, flags]) - st = out - base = end - weights_ptr = addr(weights) - return st, weights_ptr - -# ================================ the aggregation node ============================== - - -def da_list_digest(roots, n_g): - # At most 16 roots: every BLAKE2s byte counter and final flag is constant. - digest = match(log(n_g), range(0, DA_ROOT_COUNTS), lambda n: da_hash_roots(roots, n)) - return digest - - -def da_hash_roots(roots, n: Const): - digest = StackBuf(4) - if const(n == 0): - blake2s([0, 0, 0, 0], [0, 0, 0, 0], digest, counter=0, final=1) - else: - blake2s(roots[0:4], roots[4:8], digest, counter=64, final=1 // n) - for i in unroll(1, n): - claim = roots * GEN ** (8 * i) - out = StackBuf(4) - blake2s(claim[0:4], claim[4:8], out, cv=digest, counter=64 * (i + 1), final=(i + 1) // n) - digest = out - return digest - - -def cover_da_root(roots, cover, n_slots_g, mark): - # Marks one DA slot and returns its eight words: the root, then the vector hash. - index = hint_witness("da_index") - assert log(index) < log(n_slots_g) - cover[index] = mark - return roots * index ** 8 - - -def main(): - # One node of an aggregation tree: raw XMSS signatures grouped by the epoch they - # were made at (a RUNTIME number of groups), n_raw_sphincs SPHINCS signatures - # and n_children sub-proofs OF THIS SAME BYTECODE. Each XMSS group carries its - # own (epoch, message) pair, and each SPHINCS signature is against the message - # in its own coverage slot. DA roots occupy a separate region of the same table. - # - # The declared lists are the signer set; the duplicate slots absorb keys a child - # covers that the set does not declare. The coverage table is one region per - # epoch group, each its declared keys then its own duplicates, then SPHINCS - # and DA regions shaped the same way: - # - # [group 0: declared | dup]...[group n_epochs-1: declared | dup][SPHINCS: declared | dup][DA: declared | dup] - # - # The first n_decl groups are the signer set's; the n_drop after them declare - # nothing, holding a child group's (epoch, message) without publishing it. - # - # so one range check per write keeps each writer inside its own region: that is - # what makes the statement's split mean which scheme verified which key against - # which (epoch, message). An XMSS slot is four words, a SPHINCS slot eight: a key - # and the message that key signed. A DA root occupies eight words. - meta = StackBuf(7) - hint_witness(meta, "meta") # every count in the exponent - n_decl_g = meta[0] - n_drop_g = meta[1] - n_sphincs_g = meta[2] - n_sdup_g = meta[3] - n_raw_s_g = meta[4] - n_children_g = meta[5] - n_direct_da_g = meta[6] - # Declared plus dropped, so bounding each side pins n_decl <= n_epochs. - assert log(n_decl_g) < MAX_EPOCHS + 1 - assert log(n_drop_g) < MAX_EPOCHS + 1 - n_epochs_g = n_decl_g * n_drop_g - assert log(n_epochs_g) < MAX_EPOCHS + 1 - assert log(n_sphincs_g) < MAX_KEYS - assert log(n_sdup_g) < MAX_KEYS - assert log(n_raw_s_g) < MAX_KEYS - assert log(n_children_g) < MAX_RECURSIONS + 1 - assert log(n_direct_da_g) < 2 - da_meta = StackBuf(2) - hint_witness(da_meta, "da_meta") - n_da_g = da_meta[0] - n_da_dup_g = da_meta[1] - assert log(n_da_g) < MAX_DA_ROOTS + 1 - assert log(n_da_dup_g) < MAX_RECURSIONS * MAX_DA_ROOTS + 2 - da_slots_g = n_da_g * n_da_dup_g - assert log(da_slots_g) < MAX_RECURSIONS * MAX_DA_ROOTS + 2 - - # ---- the epoch groups: geometry pass ---- - # Per group: its epoch, its four message words, and its declared, duplicate and - # raw-signature counts, each count bounded before it enters a product (up to - # 2^16 factors of exponent < 2^17 stay far from the order 2^64 - 1, so nothing - # wraps). Region bases and the three totals ride a stride-4 chain; the per-group - # values land in heap buffers the later passes and the children's hinted group - # maps read back at runtime. - epochs = HeapBuf(n_epochs_g) - msgs = HeapBuf(n_epochs_g ** 4) - group_n_keys = HeapBuf(n_epochs_g) - group_n_dups = HeapBuf(n_epochs_g) - group_n_raw = HeapBuf(n_epochs_g) - group_base = HeapBuf(n_epochs_g) - group_slots = HeapBuf(n_epochs_g) - geo = HeapBuf((n_epochs_g * GEN) ** 4) # [base, n_xmss product, n_raw product] - geo[1] = 1 - geo[GEN] = 1 - geo[GEN ** 2] = 1 - for xe in mul_range(1, n_epochs_g): - grp = StackBuf(8) - hint_witness(grp, "group") # epoch, message (four words), n, n_dup, n_raw - assert log(grp[5]) < MAX_KEYS - assert log(grp[6]) < MAX_KEYS - assert log(grp[7]) < MAX_KEYS - epochs[xe] = grp[0] - msg = msgs * xe ** 4 - msg[0:4] = grp[1:5] - group_n_keys[xe] = grp[5] - group_n_dups[xe] = grp[6] - group_n_raw[xe] = grp[7] - state = geo * xe ** 4 - base = state[1] - group_base[xe] = base - slots = grp[5] * grp[6] - group_slots[xe] = slots - nxt = geo * (xe * GEN) ** 4 - nxt[1] = base * slots - nxt[GEN] = state[GEN] * grp[5] - nxt[GEN ** 2] = state[GEN ** 2] * grp[7] - geo_end = geo * n_epochs_g ** 4 - xmss_slots_g = geo_end[1] - n_raw_x_g = geo_end[GEN ** 2] - sphincs_slots_g = n_sphincs_g * n_sdup_g - # The sum of every region bounds the coverage indices, so it is what has to sit - # below the minimum memory size. - da_base_g = xmss_slots_g * sphincs_slots_g - n_total_g = da_base_g * da_slots_g - assert log(n_total_g) < MAX_KEYS - - # The proving environment (flock's R1CS and this bytecode) as one digest. It - # rides the statement rather than the bytecode, so nothing here has to know its - # own hash; the outer verifier pins it, and every child statement rebuilt below - # copies it, which is what keeps a whole tree on one bytecode. It lives on the - # heap, where the loops over children can reach it. - seed = HeapBuf(4) - hint_witness(seed[0:4], "fs_seed") - - # ---- the signer set ---- - g_logs_pow2, g_squares = exponent_tables() - - # ---- data availability ---- - da_roots = HeapBuf(da_slots_g ** 8) - for xd in mul_range(1, da_slots_g): - root = da_roots * xd ** 8 - hint_witness(root[0:8], "da_roots") - da_digest = da_list_digest(da_roots, n_da_g) - # One table per scheme, the XMSS one an epoch group at a time: each group its - # declared list (strictly sorted, checked by the outer verifier, which holds it) - # followed by its own duplicate slots. The coverage indices below run over one - # space: the group regions in order, then SPHINCS, then DA. - # - # The digest is a plain BLAKE2s of one string, in whole blocks: both lengths - # and the SPHINCS list's digest, then per group its (epoch, count, - # message) and its key list's digest, each list hashed plainly in turn. Leading - # with both lengths makes the encoding prefix-free, so no set's string is a - # prefix of another's and the digest binds its own lengths. `half` and `odd` are - # hinted per group and pinned by half*half*odd == n with odd in {0, 1}, which - # leaves half = n // 2 and odd = n % 2 as the only solution. - xmss_table = HeapBuf(xmss_slots_g ** 4) - sphincs_table = HeapBuf(sphincs_slots_g ** 8) - # The run the set's hash covers: both lengths and the SPHINCS list's digest - # in one block, then two a group. Sixteen words a group, so a group's - # header and its key digest are one block each. - signers_run = HeapBuf(n_decl_g ** 16 * GEN ** 8) - signers_run[1] = n_decl_g - signers_run[GEN] = 0 - signers_run[GEN ** 2] = n_sphincs_g - signers_run[GEN ** 3] = 0 - decl_keys = HeapBuf(n_decl_g * GEN) - decl_keys[GEN ** 0] = 1 - for xe in mul_range(1, n_decl_g): - n_keys = group_n_keys[xe] - base = group_base[xe] - decl_keys[xe * GEN] = decl_keys[xe] * n_keys - halves = StackBuf(2) - hint_witness(halves, "pk_halves") - assert log(halves[1]) < 2 - assert log(halves[0]) < MAX_KEYS - assert halves[0] * halves[0] * halves[1] == n_keys - kd = key_list_digest(xmss_table * base ** 4, halves[0], halves[1], n_keys, g_squares) - group_msg = msgs * xe ** 4 - slot = signers_run * xe ** 16 * GEN ** 8 - slot[1] = epochs[xe] - slot[GEN] = 0 - slot[GEN ** 2] = n_keys - slot[GEN ** 3] = 0 - slot[4:8] = group_msg[0:4] - slot[8:12] = kd - slot[12:16] = [0, 0, 0, 0] - # The duplicate slots ride the same table but outside the hashed prefix, past - # each group's declared keys. Over the whole table: a dropped group has only these. - for xe in mul_range(1, n_epochs_g): - n_keys = group_n_keys[xe] - base = group_base[xe] - dup_ptr = xmss_table * (base * n_keys) ** 4 - for xd in mul_range(1, group_n_dups[xe]): - dup = dup_ptr * xd ** 4 - hint_witness(dup[0:4], "dup_pubkeys") - sp_digest = sphincs_list_digest(sphincs_table, n_sphincs_g, g_squares) - signers_run[4:8] = sp_digest - # At least one published signature claim or DA root also ensures a nonempty coverage table. - assert decl_keys[n_decl_g] * n_sphincs_g * n_da_g != 1 - signers_hash = signer_set_digest(signers_run, n_decl_g, g_squares) - for xd in mul_range(1, n_sdup_g): - dup = sphincs_table * (n_sphincs_g * xd) ** 8 - hint_witness(dup[0:8], "dup_sphincs") - - # ---- coverage ---- - # Every one of the n_total slots is written exactly once: write-once memory - # rejects a second write (the value written is the running count, so two writes - # to one slot disagree), and the count below rejects a missed one. So every - # declared claim is covered by a direct check of its own kind or by a verified - # child. Each writer is confined to its own region, including DA roots. - # The raw XMSS walk runs one loop per epoch group, each signature verified - # against that group's index words (built here, only for a group that holds raw - # signatures); a stride-1 chain threads the running count across the groups. - cover = HeapBuf(n_total_g) - da_cover = cover * da_base_g - merkle_bits = HeapBuf(n_epochs_g ** LOG_LIFETIME) - index_tables = HeapBuf(n_epochs_g ** (LOG_LIFETIME + 1)) - raw_count = HeapBuf(n_epochs_g * GEN) - raw_count[GEN ** 0] = 1 - for xe in mul_range(1, n_epochs_g): - n_raw = group_n_raw[xe] - prefix = raw_count[xe] - if n_raw != 1: - index_table = index_tables * xe ** (LOG_LIFETIME + 1) - group_bits = merkle_bits * xe ** LOG_LIFETIME - fill_xmss_epoch_tables(epochs[xe], group_bits, index_table) - slots = group_slots[xe] - base = group_base[xe] - keys = xmss_table * base ** 4 - group_msg = msgs * xe ** 4 - for xi in mul_range(1, n_raw): - idx = hint_witness("raw_index") - # A runtime bound, whose `n_total < 2^MIN_LOG_MEM` precondition is - # discharged by `assert log(n_total_g) < MAX_KEYS` above. Without it - # this degenerates to what DEREF alone gives and an index could - # reach past `cover`, which is the whole bijection. The bound is - # this GROUP's region, so a signature verified at this (epoch, - # message) covers no other group's declared key. - assert log(idx) < log(slots) - cover[base * idx] = prefix * xi - verify_sig(group_msg, index_table, group_bits, keys * idx ** 4) - raw_count[xe * GEN] = prefix * n_raw - else: - raw_count[xe * GEN] = prefix - for xj in mul_range(1, n_raw_s_g): - off_hint = hint_witness("sp_raw_index") - assert log(off_hint) < log(sphincs_slots_g) - cover[xmss_slots_g * off_hint] = n_raw_x_g * xj - verify_sig_sphincs(sphincs_table * off_hint ** 8) - - if n_direct_da_g == GEN: - da_root, da_vector = da_verify(g_squares) - expected = cover_da_root(da_roots, da_cover, da_slots_g, n_raw_x_g * n_raw_s_g) - expected[0:4] = da_root - expected[4:8] = da_vector - - # ---- children ---- - child_pi = HeapBuf(n_children_g ** 4) - child_fresh = HeapBuf(n_children_g ** (3 * DEFER_SIZE)) - child_carried = HeapBuf(n_children_g ** (3 * DEFER_STMT_CELLS)) - written = HeapBuf(n_children_g * GEN) # loop-carried write count, one per child - written[GEN ** 0] = n_raw_x_g * n_raw_s_g * n_direct_da_g - for xc in mul_range(1, n_children_g): - base = written[xc] - # The child's two list lengths, then its groups, rebuilt into its signer-set - # run by rebuild_child_groups: everything hinted there is pinned by the - # run's hash, which the child's statement digest carries, so a lie about any - # of it changes the public input its proof has to satisfy. Nothing demands a - # mid-tree statement be canonical (sorted, distinct groups); it still binds - # every claim to its (epoch, message), which is all the group map relies on. - child_meta = StackBuf(2) - hint_witness(child_meta, "child_meta") # n_epochs, n_sphincs - nsub_e_g = child_meta[0] - nsub_s_g = child_meta[1] - assert log(nsub_e_g) < MAX_EPOCHS + 1 - assert log(nsub_s_g) < MAX_KEYS - sub_run = HeapBuf(nsub_e_g ** 16 * GEN ** 8) - sub_run[1] = nsub_e_g - sub_run[GEN] = 0 - sub_run[GEN ** 2] = nsub_s_g - sub_run[GEN ** 3] = 0 - nsub_x_g = rebuild_child_groups(nsub_e_g, sub_run, base, epochs, msgs, group_base, group_slots, n_epochs_g, xmss_table, cover, g_squares) - # Implied by the per-group bounds and the child's own n_total assert; stands - # as documentation. - assert log(nsub_x_g) < MAX_KEYS - nsub_g = nsub_x_g * nsub_s_g - csp = child_sphincs_list_digest(sphincs_table, cover, base * nsub_x_g, xmss_slots_g, sphincs_slots_g, nsub_s_g, g_squares) - sub_run[4:8] = csp - sub_hash = signer_set_digest(sub_run, nsub_e_g, g_squares) - carried = child_carried * xc ** (3 * DEFER_STMT_CELLS) - hint_witness(carried[0:3 * DEFER_STMT_CELLS], "child_defer") - nsub_da_g = hint_witness("child_da_count") - assert log(nsub_da_g) < MAX_DA_ROOTS + 1 - child_da = HeapBuf(nsub_da_g ** 8) - for xd in mul_range(1, nsub_da_g): - covered = cover_da_root(da_roots, da_cover, da_slots_g, base * nsub_g * xd) - slot = child_da * xd ** 8 - slot[0:8] = covered[0:8] - child_da_digest = da_list_digest(child_da, nsub_da_g) - pi = statement_digest(seed, sub_hash, child_da_digest, carried) - pi_slot = child_pi * xc ** 4 - pi_slot[0:4] = pi - verify_sub(pi, seed, g_logs_pow2, g_squares, child_fresh * xc ** (3 * DEFER_SIZE)) - written[xc * GEN] = base * nsub_g * nsub_da_g - assert written[n_children_g] == n_total_g - - # ---- this node's own deferred claims ---- - defer_stmt = HeapBuf(3 * DEFER_STMT_CELLS) - if n_children_g == 1: - # A leaf has nothing to batch, so it defers the three fixed polynomials at - # the all-zeros point. Their values ride a hint and are checked nowhere - # here: the outer verifier recomputes them and rebuilds the statement, so a - # lie changes the public input rather than the claim. - leaf_values = StackBuf(9) - hint_witness(leaf_values, "leaf_defer") - for k in unroll(0, BYTECODE_VARS): - c = 3 * k - defer_stmt[c:c + 3] = ZERO - k = 3 * DEFER_STMT_BC_VALUE - defer_stmt[k:k + 3] = leaf_values[0:3] - for j in unroll(0, 2 * K_LOG): - c = 3 * (DEFER_STMT_MAT_POINT + j) - defer_stmt[c:c + 3] = ZERO - k = 3 * DEFER_STMT_A_VALUE - defer_stmt[k:k + 3] = leaf_values[3:6] - k = 3 * DEFER_STMT_B_VALUE - defer_stmt[k:k + 3] = leaf_values[6:9] - else: - aggregate_claims(n_children_g, child_pi, child_fresh, child_carried, defer_stmt) - - own = statement_digest(seed, signers_hash, da_digest, defer_stmt) - pub = GEN ** 0 - pub[0:4] = own - return - - -def aggregate_claims(n_children_g, child_pi, child_fresh, child_carried, defer_stmt): - # Every child contributes two claims per fixed polynomial: the one IT deferred - # (carried in its statement) and the fresh one raised by verifying its proof. A - # fresh transcript binds all of them, samples the batching coefficients, and two - # sumchecks (one for the bytecode, one shared by the two matrices) reduce the - # lot to one claim each, which is what the node then defers in its own - # statement. - # - # A carried claim is a plain point, so its weight is an eq product; a fresh one - # carries flock's zerocheck/lincheck structure and keeps the succinct weight the - # sub-verifier exported. That is the only asymmetry. - bc_msgs = HeapBuf(6 * BYTECODE_VARS) - hint_witness(bc_msgs[0:6 * BYTECODE_VARS], "bc_sumcheck_msgs") - mat_msgs = HeapBuf(12 * K_LOG) - hint_witness(mat_msgs[0:12 * K_LOG], "mat_sumcheck_msgs") - bytecode_star = StackBuf(3) - hint_witness(bytecode_star, "bc_star_hint") - mat_stars = StackBuf(6) - hint_witness(mat_stars, "mat_stars_hint") - - # ---- one transcript over every child's statement and both its claim sets ---- - # A child's public input is observed as two scalars, a 128-bit half each. - fresh_row = HeapBuf(n_children_g) - carried_row = HeapBuf(n_children_g) - agg_fs = obs([AGG_SEED_0, AGG_SEED_1, AGG_SEED_2, AGG_SEED_3], [n_children_g, 0, 0]) - absorb = HeapBuf((n_children_g * GEN) ** FS_SLOTS) - absorb[0:4] = agg_fs - for xc in mul_range(1, n_children_g): - row = absorb * xc ** FS_SLOTS - pi = child_pi * xc ** 4 - st = obs(row[0:4], [pi[1], pi[GEN], 0]) - st = obs(st, [pi[GEN ** 2], pi[GEN ** 3], 0]) - fresh = child_fresh * xc ** (3 * DEFER_SIZE) - fresh_row[xc] = fresh - for k in unroll(0, DEFER_SIZE): - st = obs_at(st, fresh * GEN ** (3 * k)) - carried = child_carried * xc ** (3 * DEFER_STMT_CELLS) - carried_row[xc] = carried - for k in unroll(0, DEFER_STMT_CELLS): - st = obs_at(st, carried * GEN ** (3 * k)) - row[4:8] = st - absorbed = absorb * n_children_g ** FS_SLOTS - - # ---- bytecode batching sumcheck (BYTECODE_VARS variables, 2 per child) ---- - # Fresh and carried share the bytecode layout (point, then value), so the two - # differ only in which buffer they come from. - lam_bc = HeapBuf(n_children_g ** 6) - bc_chain = HeapBuf((n_children_g * GEN) ** ACC_SLOTS) - bc_chain[0:4] = absorbed[0:4] - bc_chain[ACC_VALUE:ACC_VALUE + 3] = ZERO - for xc in mul_range(1, n_children_g): - row = bc_chain * xc ** ACC_SLOTS - st, lam_fresh = squeeze(row[0:4]) - st, lam_carried = squeeze(st) - pair = lam_bc * xc ** 6 - pair[0:3] = lam_fresh - pair[3:6] = lam_carried - fresh = fresh_row[xc] - carried = carried_row[xc] - nxt = row * GEN ** ACC_SLOTS - nxt[0:4] = st - fb = 3 * FRESH_BC_VALUE - cb = 3 * DEFER_STMT_BC_VALUE - nxt[ACC_VALUE:ACC_VALUE + 3] = add192(row[ACC_VALUE:ACC_VALUE + 3], add192(mul192(lam_fresh, fresh[fb:fb + 3]), mul192(lam_carried, carried[cb:cb + 3]))) - bc_end = bc_chain * n_children_g ** ACC_SLOTS - bc_point = HeapBuf(3 * BYTECODE_VARS) - agg_fs, bc_running = batch_sumcheck(bc_end[0:4], bc_msgs, bc_end[ACC_VALUE:ACC_VALUE + 3], bc_point, BYTECODE_VARS) - bc_wsum = HeapBuf((n_children_g * GEN) ** 3) - bc_wsum[0:3] = ZERO - for xc in mul_range(1, n_children_g): - fresh = fresh_row[xc] - carried = carried_row[xc] - eq_fresh = ONE - eq_carried = ONE - for k in unroll(0, BYTECODE_VARS): - c = 3 * k - rk = bc_point[c:c + 3] - eq_fresh = mul192(eq_fresh, add192(ONE, add192(fresh[c:c + 3], rk))) - eq_carried = mul192(eq_carried, add192(ONE, add192(carried[c:c + 3], rk))) - pair = lam_bc * xc ** 6 - x3 = xc ** 3 - nxt = x3 * GEN ** 3 - bc_wsum[nxt:nxt + 3] = add192(bc_wsum[x3:x3 + 3], add192(mul192(pair[0:3], eq_fresh), mul192(pair[3:6], eq_carried))) - ws_end = n_children_g ** 3 - assert_eq192(bc_running, mul192(bytecode_star, bc_wsum[ws_end:ws_end + 3])) - - # ---- matrix batching sumcheck (2*K_LOG variables, 3 claims per child) ---- - # The fresh claim is one value against A0 weighted by lincheck's alpha plus B0; - # a carried claim is one value per matrix at a shared point. - lam_mat = HeapBuf(n_children_g ** 9) - mat_chain = HeapBuf((n_children_g * GEN) ** ACC_SLOTS) - mat_chain[0:4] = agg_fs - mat_chain[ACC_VALUE:ACC_VALUE + 3] = ZERO - for xc in mul_range(1, n_children_g): - row = mat_chain * xc ** ACC_SLOTS - st, lam_fresh = squeeze(row[0:4]) - st, lam_a = squeeze(st) - st, lam_b = squeeze(st) - triple = lam_mat * xc ** 9 - triple[0:3] = lam_fresh - triple[3:6] = lam_a - triple[6:9] = lam_b - fresh = fresh_row[xc] - carried = carried_row[xc] - nxt = row * GEN ** ACC_SLOTS - nxt[0:4] = st - fm = 3 * FRESH_MATPART - ca = 3 * DEFER_STMT_A_VALUE - cb = 3 * DEFER_STMT_B_VALUE - nxt[ACC_VALUE:ACC_VALUE + 3] = add192(row[ACC_VALUE:ACC_VALUE + 3], add192(mul192(lam_fresh, fresh[fm:fm + 3]), add192(mul192(lam_a, carried[ca:ca + 3]), mul192(lam_b, carried[cb:cb + 3])))) - mat_end = mat_chain * n_children_g ** ACC_SLOTS - mat_point = HeapBuf(3 * 2 * K_LOG) - agg_fs, mat_running = batch_sumcheck(mat_end[0:4], mat_msgs, mat_end[ACC_VALUE:ACC_VALUE + 3], mat_point, 2 * K_LOG) - # Terminal weights. A fresh claim's is U_t(r*) = urow_t(r*_row) * wcol_t(r*_col), - # with row_weight = (sum_i L_i(zz_t) eq(r*[0..6], i)) * eq(zchi_t, r*[6..K_LOG]) - # and col_weight = (sum_i z_partial_t[i] eq(r*[K_LOG..K_LOG+6], i)) * prod_j (1 + - # lrr_j + r*[2*K_LOG-1-j]) (the lincheck binds column variables top-down). A - # carried claim's is a plain eq over all 2*K_LOG coordinates. - eq_rows = HeapBuf(3 * (2 ** (K_SKIP + 1) - 2)) - eqtree(mat_point, eq_rows, K_SKIP) - eq_cols = HeapBuf(3 * (2 ** (K_SKIP + 1) - 2)) - eqtree(mat_point * GEN ** (3 * K_LOG), eq_cols, K_SKIP) - w_sums = HeapBuf((n_children_g * GEN) ** 6) # the A and B weight sums - w_sums[0:3] = ZERO - w_sums[3:6] = ZERO - for xc in mul_range(1, n_children_g): - fresh = fresh_row[xc] - row_nums = StackBuf(3 * 2 ** K_SKIP) - zs = 3 * FRESH_Z_SKIP - lag64(fresh[zs:zs + 3], row_nums, 0) - row_weight = ZERO - for i in unroll(0, 2 ** K_SKIP): - k = 3 * i - e = 3 * (2 ** K_SKIP - 2 + i) - row_weight = add192(row_weight, mul192(row_nums[k:k + 3], eq_rows[e:e + 3])) - row_weight = mul192(row_weight, LAGRANGE_INV_S) - for k in unroll(0, LINCHECK_ROUNDS): - zc = 3 * (FRESH_ZCHI + k) - mp = 3 * (K_SKIP + k) - row_weight = mul192(row_weight, add192(ONE, add192(fresh[zc:zc + 3], mat_point[mp:mp + 3]))) - col_weight = ZERO - for i in unroll(0, 2 ** K_SKIP): - zp = 3 * (FRESH_Z_PARTIAL + i) - e = 3 * (2 ** K_SKIP - 2 + i) - col_weight = add192(col_weight, mul192(fresh[zp:zp + 3], eq_cols[e:e + 3])) - for j in unroll(0, LINCHECK_ROUNDS): - lr = 3 * (FRESH_LINCHECK_RS + j) - mp = 3 * (2 * K_LOG - 1 - j) - col_weight = mul192(col_weight, add192(ONE, add192(fresh[lr:lr + 3], mat_point[mp:mp + 3]))) - weight_u = mul192(row_weight, col_weight) - carried = carried_row[xc] - eq_carried = ONE - for k in unroll(0, 2 * K_LOG): - cp = 3 * (DEFER_STMT_MAT_POINT + k) - mp = 3 * k - eq_carried = mul192(eq_carried, add192(ONE, add192(carried[cp:cp + 3], mat_point[mp:mp + 3]))) - triple = lam_mat * xc ** 9 - lam_fresh = triple[0:3] - fa = 3 * FRESH_ALPHA - row = w_sums * xc ** 6 - row[6:9] = add192(row[0:3], add192(mul192(lam_fresh, weight_u), mul192(triple[3:6], eq_carried))) - row[9:12] = add192(row[3:6], add192(mul192(mul192(lam_fresh, fresh[fa:fa + 3]), weight_u), mul192(triple[6:9], eq_carried))) - w_end = w_sums * n_children_g ** 6 - assert_eq192(mat_running, add192(mul192(mat_stars[0:3], w_end[0:3]), mul192(mat_stars[3:6], w_end[3:6]))) - - for k in unroll(0, BYTECODE_VARS): - c = 3 * k - defer_stmt[c:c + 3] = bc_point[c:c + 3] - k = 3 * DEFER_STMT_BC_VALUE - defer_stmt[k:k + 3] = bytecode_star - for j in unroll(0, 2 * K_LOG): - c = 3 * j - d = 3 * (DEFER_STMT_MAT_POINT + j) - defer_stmt[d:d + 3] = mat_point[c:c + 3] - k = 3 * DEFER_STMT_A_VALUE - defer_stmt[k:k + 3] = mat_stars[0:3] - k = 3 * DEFER_STMT_B_VALUE - defer_stmt[k:k + 3] = mat_stars[3:6] - return diff --git a/crates/rec_aggregation/src/aggregation.rs b/crates/rec_aggregation/src/aggregation.rs deleted file mode 100644 index 62e7612e1..000000000 --- a/crates/rec_aggregation/src/aggregation.rs +++ /dev/null @@ -1,5182 +0,0 @@ -//! Recursive proofs of XMSS and SPHINCS signature claims and LeanDA blob well-formedness. -//! One bytecode (`guests/lean_ethereum.py`) serves every node of an aggregation tree. -//! -//! A node verifies `n_raw_xmss` XMSS signatures, `n_raw_sphincs` SPHINCS -//! signatures and `n_children` sub-proofs **of this same bytecode**, and -//! by default publishes the sorted deduplicated union of their signer sets. The XMSS -//! signers are grouped by epoch, each group -//! carrying its own message. A SPHINCS -//! signer carries its own message, so that half of the statement is a list of -//! `(key, message)` pairs. Coverage is what carries the security claim: a write-once slot per -//! declared signer, written once by each raw signature and each child key, plus -//! a final count, so every declared signer is backed by a real signature or a -//! verified child. -//! -//! Those slots are one contiguous region per XMSS epoch group and one for -//! SPHINCS, so the one range check a write already needs also keeps a -//! signature off another group's declared keys, of either scheme: that is what -//! makes the published split mean which scheme verified which key against -//! which `(epoch, message)`, at every level of the tree. -//! -//! A duplicate slot sits outside the prefix the digest hashes, so a key in one is -//! covered and not claimed. `aggregate`'s `declare` rests on that, the table also -//! holding undeclared groups so a whole `(epoch, message)` can go unpublished. -//! A child's groups need -//! not equal its parent's: a hinted map, checked by the guest, ties each -//! non-empty child group to a parent group with the same epoch and message. -//! An XMSS slot holds the -//! key's two cells and a SPHINCS slot four, its key and its message, so the -//! guest reads each SPHINCS signature's message out of the slot it verifies. -//! -//! The bytecode is compiled to a fixed point on its own size -//! ([`unified_guest`]): the recursion placeholders depend on the inner bytecode -//! size, and here the inner bytecode is this one. Its digest does not need a -//! fixed point, riding the statement instead of the code. -//! -//! Three fixed polynomials (the stacked bytecode and flock's `A0`/`B0`) are too -//! big to evaluate in-circuit, so each node exports one deferred claim on each -//! and batches its children's carried claims with the fresh ones its -//! verifications raise (`doc/leanvm/main.tex` §Deferred evaluation claims). Only -//! the root's are discharged natively, by [`EthereumProof::verify`]. -//! -//! `gen_verify` derives the guest's whole witness for a child from the real -//! `cpu::layout` of the inner program and the summary of a real `cpu::verify` -//! run, so there is no hand-mirrored copy of the protocol to drift. - -use bincode::Options as _; -use pcs::whir::{MAX_LOG_INV_RATE, MIN_LOG_INV_RATE}; -use std::collections::{BTreeMap, BTreeSet}; -use std::ops::Range; - -use lean_compiler::{compile, parse_with_replacements}; -use lean_da::{BLOB_SYMBOLS, CELL_SYMBOLS, CELLS_PER_ROW, CODEWORD_SYMBOLS}; -pub use lean_da::{DA_LOG_CELL, DA_LOG_K, DA_MAX_ROWS}; -use lean_vm::cpu::{Program, prove, verify}; -use lean_vm::leaf::{Block, Coord}; -use lean_vm::transcript::FiatShamirState; -use primitives::field::{F64, F192, G, g_pow}; -use primitives::multilinear::mle_eval_par; -use xmss::{XmssPublicKey, XmssSignature}; - -use sphincs::{SphincsPublicKey, SphincsSignature}; - -/// One SPHINCS claim: a key, and the message it signed. Where an XMSS group -/// shares one message, every SPHINCS signer carries its own. -pub type SphincsClaim = (SphincsPublicKey, sphincs::Message); - -/// The XMSS signers sharing one epoch: the epoch, the message they all signed -/// at it, and their strictly sorted keys. -#[derive(Clone, Debug, PartialEq, Eq, serde::Serialize, serde::Deserialize)] -pub struct XmssClaimGroup { - pub epoch: xmss::Epoch, - pub message: xmss::Message, - pub keys: Vec, -} - -/// Why the guest reads every `q_flock` slot claim's instance point off `chi`: a -/// virtual value column is referenced only by its own table's bus blocks, which -/// the table sumcheck settles, so no framework block can raise one at `zeta`. -const VALCOL_FRAMEWORK: &str = "a framework block must not reference a virtual value column"; -const RECURSION_AGG_LABEL: &[u8] = b"leanvm/recursion-aggregation/v1"; - -/// The most earlier aggregates one [`aggregate`] call can take, so the arity of -/// an aggregation tree. -pub const MAX_RECURSIONS: usize = 16; - -/// Exclusive cap on coverage slots: signatures of both schemes and DA roots, -/// including duplicates and omitted claims. Exactly this many is already too many. -/// -/// It is where the coverage indices' range check has to sit to stay below -/// `2^MIN_LOG_MEM`, so the bound means the same at every announced memory size. -/// The guest checks it against a COMPILE-TIME bound, so raising it past -/// `2^MIN_LOG_MEM` fails the build rather than weakening the index bound. -pub const MAX_KEYS: usize = 1 << 16; - -/// Maximum number of LeanDA roots in one statement, including inherited and direct roots. -pub const MAX_DA_ROOTS: usize = 16; -const _: () = assert!(MAX_DA_ROOTS < MAX_KEYS); - -/// The most [`XmssClaimGroup`]s one aggregate can carry: headroom over the few epochs -/// expected in practice. An empty group is in-circuit unprovable, its list hash -/// having no valid window split, so this is a bound on cost, not on soundness. -pub const MAX_EPOCHS: usize = 1024; - -/// Blocks the guest absorbs per loop frame when it hashes a declared list, and so -/// how many share one byte-counter base (`doc/leanvm` §Byte counters for a hash of -/// runtime length). Larger amortizes the base's bit decomposition over more blocks -/// and costs bytecode in the tail's dispatch arms; the split it induces is what -/// [`signers_split`] hands the guest. -const SIGNERS_WINDOW: usize = 32; -// The counter split is a bit split: a window's base is `64·SIGNERS_WINDOW·q`, which -// has to be a power of two for its bits to sit clear of the window's own offsets. -const _: () = assert!(SIGNERS_WINDOW.is_power_of_two()); - -/// The window bound the guest range-checks its hinted window count against. A -/// range check takes a COUNT (`assert log(x) < k` bounds the exponent by `k`), so -/// this has to cover the most windows a list can hold, which a bit width would not: -/// under the bound every list still hashes, so the mistake shows up only past it. -const SIGNERS_MAX_WINDOWS: usize = MAX_KEYS / SIGNERS_WINDOW; -// The widest list is one block a claim, so at most MAX_KEYS blocks, of which the -// last is absorbed apart. The set's own string is 2 + 2·MAX_EPOCHS blocks, hashed -// the same way, so it needs the bound too. -const _: () = assert!((MAX_KEYS - 1) / SIGNERS_WINDOW < SIGNERS_MAX_WINDOWS); -const _: () = assert!((2 * MAX_EPOCHS + 1) / SIGNERS_WINDOW < SIGNERS_MAX_WINDOWS); -// The guest decomposes a count into SIGNERS_COUNT_BITS bits and shifts the result -// left to make a byte counter, so a count has to fit and the shift must not reduce. -const SIGNERS_COUNT_BITS: u32 = MAX_KEYS.ilog2(); -const _: () = assert!(MAX_KEYS.is_power_of_two() && 2 * MAX_EPOCHS + 2 < 1 << SIGNERS_COUNT_BITS); -const _: () = assert!(SIGNERS_COUNT_BITS + 6 + SIGNERS_WINDOW.ilog2() <= 64); - -// The guest bakes a bytecode claim's width from `N_TUPLE_BITS` while `bytecode_vars` -// reads it off the stacked table, which is `N_BYTECODE_SELECTORS` wide. Two constants -// that happen to agree: were they to drift, a leaf's claim point would be one length in -// the guest and another in the statement, and nothing else would notice. -const _: () = assert!(lean_vm::leaf::N_TUPLE_BITS == lean_vm::leaf::N_BYTECODE_SELECTORS); -// The epoch fills a tweak's four-byte index field, so a longer lifetime would -// need a weight per bit that `xmss::make_tweak` cannot express. -const _: () = assert!(xmss::LOG_LIFETIME <= 32); -// The guest's `WOTS_PK_BLOCKS = (2 + V) / 4` truncates, so a bad `V` would drop -// the last tips. -const _: () = assert!((2 + xmss::V).is_multiple_of(4)); -// The SPHINCS side of the same shape. `SP_LEAF_BLOCKS = (2 + V) / 4` and -// `SP_ROOT_BLOCKS = (2 + NUM_FTS_TREES) / 4` truncate, and a truncated loop -// would leave the last tips or roots out of the hash while the signature still -// carries them: revealed values no longer bound by the leaf they belong to. -const _: () = assert!((2 + sphincs::V).is_multiple_of(4)); -const _: () = assert!((2 + sphincs::NUM_FTS_TREES).is_multiple_of(4)); -// The guest reads the message digest's bits out of three 64-bit lanes, and a -// dynamically sized `HeapBuf` gets no compile-time index check, so a wider -// digest would read leaf indices from cells nothing writes. -const _: () = assert!(sphincs::DIGEST_BITS <= 3 * 64); -// The guest packs each tweak field into its own 32-bit word: p at bit 32, -// tau at bit 64, and j at bit 96. -const _: () = assert!(sphincs::H <= 32); -const _: () = assert!(sphincs::CHAIN_LEN * sphincs::V < 1 << 32); -const _: () = assert!(sphincs::A <= 32 && sphincs::HEIGHTS[0] <= 32); - -/// A count as the guest carries it: in the exponent, `g^n`. -fn count(n: usize) -> F64 { - g_pow(n) -} - -fn f192_literal(f: F192) -> String { - format!("f192({},{},{})", f.c0, f.c1, f.c2) -} - -/// Field elements as the words the guest holds them in, three limbs each. -fn limbs(values: &[F192]) -> Vec { - values.iter().flat_map(|v| [F64(v.c0), F64(v.c1), F64(v.c2)]).collect() -} - -/// A byte string of whole words as its little-endian words, the order a BLAKE2s -/// block or digest occupies memory in. -fn words(bytes: &[u8]) -> Vec { - let (chunks, rest) = bytes.as_chunks::<8>(); - assert!(rest.is_empty(), "a byte string of whole words"); - chunks.iter().map(|w| F64(u64::from_le_bytes(*w))).collect() -} - -/// A 32-byte digest as its four words. -fn digest_words(hash: &[u8; 32]) -> [F64; 4] { - words(hash).try_into().expect("four words") -} - -/// A public key as the four words the guest hashes and `verify_sig` reads: the -/// root then the public parameter. Both schemes lay a key out the same way, and -/// the statement keeps them in separate lists rather than telling them apart by -/// their bytes. -fn key_words(pk: &XmssPublicKey) -> Vec { - [words(&pk.merkle_root), words(&pk.public_param)].concat() -} - -/// A run of words as the byte string BLAKE2s hashes over it. -fn word_bytes(words: impl IntoIterator) -> Vec { - words.into_iter().flat_map(|w| w.0.to_le_bytes()).collect() -} - -/// How the guest splits a list's `n - 1` non-final blocks: whole windows, then the -/// tail. Its own product identity and range check pin both, so this only has to -/// agree with them (`sphincs_list_digest` in the guest). A list is never empty -/// here: a group holds at least one key, and both claim lists are guarded at the -/// call, since the guest hashes an empty one without a split at all. -fn signers_split(blocks: usize) -> Vec { - assert!(blocks > 0, "an empty list has no window split"); - let leading = blocks - 1; - vec![count(leading / SIGNERS_WINDOW), count(leading % SIGNERS_WINDOW)] -} - -/// One epoch group's declared keys under plain BLAKE2s: 32 bytes a key, so the -/// hashed string is `32n` bytes and only its last block is partial. The guest -/// computes this same digest a window of blocks at a time (`key_list_digest`). -fn key_list_digest(keys: &[XmssPublicKey]) -> [F64; 4] { - digest_words(&primitives::hash::hash(&word_bytes(keys.iter().flat_map(key_words)))) -} - -/// The declared SPHINCS claims under plain BLAKE2s: one 64-byte block per claim, -/// its key then the message it signed, so the hashed string is exactly `64n` bytes -/// and an empty list hashes the empty string. The guest computes this same digest a -/// window of blocks at a time (`sphincs_list_digest`). -fn sphincs_list_digest(signers: &[SphincsClaim]) -> [F64; 4] { - digest_words(&primitives::hash::hash(&word_bytes( - signers.iter().flat_map(sphincs_signer_words), - ))) -} - -/// A SPHINCS signer as the eight words the guest hashes and `verify_sig_sphincs` -/// reads: the key, then the message that key signed. -fn sphincs_signer_words((pk, message): &SphincsClaim) -> Vec { - [words(&pk.root), words(&pk.public_param), words(message)].concat() -} - -/// The first word of an XMSS tweak, `xmss::make_tweak`'s own output: its type and -/// sub-position. The guest holds no byte layout of its own, building a tweak from -/// these words, a sub-position unit and one [`tweak_index_weight`] per set epoch -/// bit, so a field that moves in `make_tweak` moves both sides together. -fn tweak_word(tweak_type: u8, sub_position: u32) -> F64 { - words(&xmss::make_tweak(tweak_type, sub_position, 0))[0] -} - -/// What bit `b` of the epoch weighs in a tweak's second word, so an index is its -/// set bits summed. The one property of the layout this assumes is that the -/// index field is linear in the index. Subtract the constant protocol prefix. -fn tweak_index_weight(b: usize) -> F64 { - let second = |index: u32| words(&xmss::make_tweak(0, 0, index))[1]; - second(1 << b) + second(0) -} -/// The signer-set digest: plain BLAKE2s of one byte string, laid out in whole -/// 64-byte blocks so the guest can absorb it eight words at a time -/// (`signer_set_digest` there). The first block carries both list lengths and the -/// SPHINCS list's own digest, followed by two blocks a group: its `(epoch, -/// count, message)`, then its key list's digest. Leading with both lengths makes -/// the encoding prefix-free, so no set's string is a prefix of another's, and the -/// digest binds its own lengths, the groups' epochs and messages, and every split. -/// The two list digests carry the bulk, each a stock hash of its own -/// ([`key_list_digest`], [`sphincs_list_digest`]). -fn signers_hash(xmss_signers: &[XmssClaimGroup], sphincs_signers: &[SphincsClaim]) -> [F64; 4] { - let zero = F64::ZERO; - let mut run = vec![count(xmss_signers.len()), zero, count(sphincs_signers.len()), zero]; - run.extend(sphincs_list_digest(sphincs_signers)); - for XmssClaimGroup { epoch, message, keys } in xmss_signers { - run.extend([F64(*epoch as u64), zero, count(keys.len()), zero]); - run.extend(words(message)); - run.extend(key_list_digest(keys)); - run.extend([zero; 4]); - } - digest_words(&primitives::hash::hash(&word_bytes(run))) -} - -/// The claims on the three fixed polynomials that a node defers rather than -/// evaluating in-circuit: one point and value on the stacked bytecode, one point -/// and two values on flock's `A0`/`B0` (`doc/leanvm/main.tex` §Deferred -/// evaluation claims). -/// -/// Only the points are transmitted; the values are derived from them on receipt, -/// so a prover that lies about a value changes the statement its proof has to -/// satisfy rather than the claim anyone checks. -#[derive(Clone, Debug, PartialEq, Eq)] -pub struct DeferredClaim { - bytecode_point: Vec, - bytecode_value: F192, - matrix_point: Vec, - matrix_a_value: F192, - matrix_b_value: F192, -} - -impl DeferredClaim { - /// What a leaf defers: the all-zeros point on each polynomial, where the - /// value is just the table's first entry. - fn leaf() -> Self { - Self::recompute( - vec![F192::ZERO; bytecode_vars()], - vec![F192::ZERO; 2 * flock::hash::K_LOG], - ) - .expect("the all-zeros point has the right shape") - } - - /// Evaluate the three fixed polynomials at `bytecode_point` / `matrix_point`. - fn recompute(bytecode_point: Vec, matrix_point: Vec) -> Result { - let klog = flock::hash::K_LOG; - if bytecode_point.len() != bytecode_vars() || matrix_point.len() != 2 * klog { - return Err(AggregateVerifyError::MalformedClaim); - } - // Every leaf defers the all-zeros point, where the bytecode polynomial is - // just its table's first entry. Still worth special-casing that half: the - // general path is a pass over 2^23 entries, and a leaf is the aggregate - // people verify most. The matrix half needs no special case, its walk - // being O(circuit) either way. - let zero_point = bytecode_point.iter().chain(&matrix_point).all(|x| *x == F192::ZERO); - let bytecode_value = if zero_point { - F192::from(stacked_bytecode()[0]) - } else { - let sp = tracing::info_span!("bytecode mle").entered(); - let value = mle_eval_par(stacked_bytecode(), &bytecode_point); - drop(sp); - value - }; - let sp = tracing::info_span!("matrix walk").entered(); - let eq_r = pcs::whir::build_eq_table_ext(&matrix_point[..klog]); - let eq_c = pcs::whir::build_eq_table_ext(&matrix_point[klog..]); - let (matrix_a_value, matrix_b_value) = flock::hash::bilinear_walk_pair(&eq_r, &eq_c); - drop(sp); - Ok(Self { - bytecode_point, - bytecode_value, - matrix_point, - matrix_a_value, - matrix_b_value, - }) - } - - /// The cells the statement digest absorbs, in the guest's `defer_stmt` order. - fn cells(&self) -> Vec { - let mut cells = self.bytecode_point.clone(); - cells.push(self.bytecode_value); - cells.extend_from_slice(&self.matrix_point); - cells.push(self.matrix_a_value); - cells.push(self.matrix_b_value); - cells - } -} - -/// The statement's fixed header, ahead of the deferred elements: the seed, the -/// signer-set digest (which itself binds the epoch groups and every count), and -/// the DA root-list digest, four words each. Fed to the guest as `STMT_HEADER`, so -/// the two cannot drift. -const STATEMENT_HEADER: usize = 12; - -/// A node's public statement, hashed to the four words the VM publishes. The -/// guest's `statement_digest` computes exactly this, both for itself and when it -/// rebuilds a child's, which is what forces a whole tree onto one bytecode and -/// each child's `(epoch, message)` groups, bound by the signer-set digest, -/// onto its parent's list. -/// -/// Fixed-length preimage, so a plain BLAKE2s, with no domain tag of its own: the -/// header leads with the environment digest, which binds this bytecode and -/// flock's R1CS and so already separates the preimage from every other use of -/// BLAKE2s here. The header's words, then all three limbs of each deferred -/// element, zero-filled to a whole block. -fn statement_digest(signers_hash: [F64; 4], da_digest: [u8; 32], defer: &DeferredClaim) -> [F64; 4] { - let seed = lean_vm::cpu::fs_seed(unified_guest()); - let header = [seed, signers_hash, digest_words(&da_digest)].concat(); - assert_eq!(header.len(), STATEMENT_HEADER); - let mut bytes = word_bytes(header.into_iter().chain(limbs(&defer.cells()))); - bytes.resize(bytes.len().next_multiple_of(64), 0); - digest_words(&primitives::hash::hash(&bytes)) -} - -/// The deferred-claim data the guest binds to the outer public input: the outer -/// verifier checks each claim natively (`doc/leanvm/main.tex` §Deferred evaluation claims; -/// n_rec = 1 forwards fresh claims without batching). -struct DeferredSubproof { - public_input: [F64; 4], - bytecode_row_point: Vec, - bytecode_selector_point: Vec, - bytecode_value: F192, - matrix_a_coefficient: F192, - skip_point: F192, - zerocheck_row_point: Vec, - lincheck_round_point: Vec, - lincheck_terminal_values: Vec, - matrix_claim: F192, -} - -/// A proof that every key in [`Self::xmss_signers`] signed its group's message -/// at its group's epoch under XMSS, and that every `(key, message)` in -/// [`Self::sphincs_signers`] is backed by a valid SPHINCS signature. Each root in -/// [`Self::da_commitments`] also attests that the committed blob rows are valid Reed-Solomon codewords. -/// These claims can be established directly or carried from verified child proofs. -/// -/// The two lists describe the signature claims and remain separate because the proof says which scheme -/// verified which key (the module docs give the coverage argument). Until -/// [`Self::verify`] returns `Ok` they are claims, not attestation, and even then -/// the epochs and messages are the prover's, so a caller that reads either list -/// as attestation of something must compare it against what it expected. -/// -/// **Neither length counts signers, only claims.** One key may appear in several -/// XMSS groups or under several SPHINCS messages, so a committee threshold has -/// to count distinct keys itself. -#[derive(Clone, Debug)] -pub struct EthereumProof { - /// The XMSS signers: strictly increasing epochs, - /// each group non-empty and strictly sorted. May be empty. The claims of both schemes together are - /// strictly fewer than [`MAX_KEYS`]. - xmss_signers: Vec, - /// Strictly sorted and deduplicated on the whole `(key, message)` pair. - sphincs_signers: Vec, - /// Strictly sorted LeanDA roots. Their list digest rides the public statement. - da_roots: Vec<[u8; 32]>, - /// What this aggregate defers to whoever discharges it: its parent, in - /// circuit, or [`Self::verify`], natively. - defer: DeferredClaim, - proof: lean_vm::cpu::Proof, -} - -#[derive(Clone, Debug, PartialEq, Eq)] -pub enum AggregateVerifyError { - /// The signer set is unsorted, holds a duplicate, or has [`MAX_KEYS`] keys or more. - MalformedSignerSet, - /// The DA root list is unsorted, holds duplicates, or exceeds [`MAX_DA_ROOTS`]. - MalformedDaCommitments, - /// A deferred claim's point has the wrong number of coordinates. - MalformedClaim, - /// The bytes are not a valid encoding of an aggregate. - MalformedEncoding, - /// The snark itself did not verify. - Snark(lean_vm::cpu::CpuError), -} - -#[derive(Clone, Debug, PartialEq, Eq)] -pub enum AggregationError { - /// Two claims at one epoch, raw or through children, carry different - /// messages: within an aggregate the message is a function of the epoch. - ConflictingMessages, - /// The union of the epochs, over the raw signatures and the children, - /// exceeds [`MAX_EPOCHS`]. - TooManyEpochs, - /// A child aggregate does not verify. - InvalidChild(AggregateVerifyError), - /// No signature claims or DA roots to publish. - Empty, - /// A declared claim is not one the contributions cover. A claim is a key, an - /// epoch and a message, so another epoch or another message is another claim. - NotCovered, - /// A raw signature's randomness does not decode to a target-sum encoding, - /// so there is no witness to build for it. - MalformedRawSignature, - /// More than [`MAX_RECURSIONS`] children, or [`MAX_KEYS`] signers or more - /// once the duplicate slots are counted, or more than [`MAX_DA_ROOTS`] DA roots. - TooLarge, - /// The payload has a partial row or exceeds [`DA_MAX_ROWS`]. - InvalidBlobSize { symbols: usize }, - /// A requested DA commitment is absent from the children and the direct blob check. - BlobNotCovered, - /// `log_inv_rate` is outside the range the WHIR configuration accepts. - InvalidRate { log_inv_rate: usize }, - /// A child's committed witness falls outside the opening arms the guest was - /// compiled with. Small aggregates are padded up to the floor, so this means - /// a child too big: more signatures than one node can hold. - ChildOutOfRange { log_committed: usize }, -} - -impl std::fmt::Display for AggregateVerifyError { - fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { - match self { - Self::MalformedSignerSet => write!(f, "malformed signer set"), - Self::MalformedDaCommitments => write!(f, "malformed DA commitment list"), - Self::MalformedClaim => write!(f, "malformed deferred claim"), - Self::MalformedEncoding => write!(f, "not a valid aggregate encoding"), - Self::Snark(e) => write!(f, "the snark did not verify: {e:?}"), - } - } -} - -impl std::error::Error for AggregateVerifyError {} - -impl std::fmt::Display for AggregationError { - fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { - match self { - Self::ConflictingMessages => write!(f, "two claims at one epoch carry different messages"), - Self::TooManyEpochs => write!(f, "more than MAX_EPOCHS ({MAX_EPOCHS}) epochs"), - Self::InvalidChild(_) => write!(f, "invalid child aggregate"), - Self::Empty => write!(f, "no signature claims or DA roots to publish"), - Self::NotCovered => write!(f, "a declared claim is not covered by the children and raw signatures"), - Self::MalformedRawSignature => write!(f, "a raw signature does not decode to a target-sum encoding"), - Self::TooLarge => { - write!(f, "too many children, signature claims, or DA roots") - } - Self::InvalidBlobSize { symbols } => { - write!( - f, - "{symbols} blob symbols do not form 1..={DA_MAX_ROWS} rows of {} symbols", - 1 << DA_LOG_K - ) - } - Self::BlobNotCovered => write!(f, "the requested blob commitment is not carried by any child"), - Self::InvalidRate { log_inv_rate } => { - write!( - f, - "log_inv_rate {log_inv_rate} is outside {}..={}", - MIN_LOG_INV_RATE, MAX_LOG_INV_RATE - ) - } - Self::ChildOutOfRange { log_committed } => { - write!( - f, - "a child commits 2^{log_committed} words, outside the guest's opening arms" - ) - } - } - } -} - -impl std::error::Error for AggregationError { - fn source(&self) -> Option<&(dyn std::error::Error + 'static)> { - match self { - Self::InvalidChild(e) => Some(e), - _ => None, - } - } -} - -/// Everything but the signer set, which a receiver may already hold. -type WireCore = (Vec<[u8; 32]>, Vec, Vec, lean_vm::cpu::Proof); - -/// Signature claims grouped by scheme: XMSS epoch/message groups, then SPHINCS key/message pairs. -#[derive(Clone, Debug, Default, PartialEq, Eq, serde::Serialize, serde::Deserialize)] -pub struct SignatureClaims { - pub xmss: Vec, - pub sphincs: Vec, -} - -/// Exactly the claims to publish. Empty lists publish no claims of that kind. -#[derive(Clone, Copy)] -pub struct ClaimSelection<'a> { - pub signatures: &'a SignatureClaims, - pub da_commitments: &'a [[u8; 32]], -} - -/// The wire encoding: bincode's fixed-width integers, as the free functions use, -/// but rejecting trailing bytes, which they do not. Without that an accepted -/// aggregate has unboundedly many encodings, so anything downstream that dedupes -/// or indexes on the serialized bytes can be made to see one aggregate as many. -fn wire() -> impl bincode::Options { - bincode::DefaultOptions::new().with_fixint_encoding() -} - -/// Reject a signer set that the coverage argument does not cover: strict sorting -/// within each list is what stops one signer being counted many times: the XMSS -/// list's length is a count of distinct `(epoch, key)` claims, the SPHINCS -/// list's of distinct `(key, message)` claims. The epoch groups are strictly -/// increasing, non-empty (an absent epoch is an absent group, the one -/// encoding of each set) and at most [`MAX_EPOCHS`]. Either list may be empty; -/// both may be empty for a blob proof. [`MAX_KEYS`] is exclusive here, as in the guest. -fn check_signer_set( - xmss_signers: &[XmssClaimGroup], - sphincs_signers: &[SphincsClaim], -) -> Result<(), AggregateVerifyError> { - let total = xmss_signers.iter().map(|group| group.keys.len()).sum::() + sphincs_signers.len(); - if total >= MAX_KEYS - || xmss_signers.len() > MAX_EPOCHS - || !xmss_signers.windows(2).all(|w| w[0].epoch < w[1].epoch) - || xmss_signers - .iter() - .any(|group| group.keys.is_empty() || !group.keys.windows(2).all(|w| w[0] < w[1])) - || !sphincs_signers.windows(2).all(|w| w[0] < w[1]) - { - return Err(AggregateVerifyError::MalformedSignerSet); - } - Ok(()) -} - -fn check_da_roots(roots: &[[u8; 32]]) -> Result<(), AggregateVerifyError> { - if roots.len() > MAX_DA_ROOTS || roots.windows(2).any(|pair| pair[0] >= pair[1]) { - return Err(AggregateVerifyError::MalformedDaCommitments); - } - Ok(()) -} - -fn da_claim_words(root: &[u8; 32]) -> Vec { - let vector_hash = lean_da::vector_digest(&lean_da::membership_vector(root)); - [words(root), words(&vector_hash)].concat() -} - -fn da_list_digest(roots: &[[u8; 32]]) -> [u8; 32] { - // Recompute vector hashes from the roots, including when verifying a received proof. - // Trusting prover-supplied hashes would let zero weights certify any matrix. - primitives::hash::hash(&word_bytes(roots.iter().flat_map(da_claim_words))) -} - -impl EthereumProof { - /// This aggregate's own public statement, as the VM publishes it. - fn public_input(&self) -> [F64; 4] { - statement_digest( - signers_hash(&self.xmss_signers, &self.sphincs_signers), - self.da_commitments_digest(), - &self.defer, - ) - } - - /// Strictly increasing epochs, each group non-empty and strictly sorted. - /// May be empty, including in a blob-only proof. - pub fn xmss_signers(&self) -> &[XmssClaimGroup] { - &self.xmss_signers - } - - /// Strictly sorted and deduplicated on the whole `(key, message)` pair. - pub fn sphincs_signers(&self) -> &[SphincsClaim] { - &self.sphincs_signers - } - - /// Strictly sorted, distinct LeanDA roots proved directly or inherited from children. - /// An empty list makes no blob claim. - pub fn da_commitments(&self) -> &[[u8; 32]] { - &self.da_roots - } - - /// BLAKE2s of the concatenated `(root, vector hash)` pairs, or of empty input. - /// Each vector hash is derived from its root outside the SNARK. - /// This digest is bound into the public statement. - pub fn da_commitments_digest(&self) -> [u8; 32] { - da_list_digest(&self.da_roots) - } - - /// The declared claims, as many as the coverage table's declared slots. - /// - /// NOT a count of distinct signers: a SPHINCS key may hold several claims, - /// one per message it signed, and an XMSS key one per epoch it signed at - /// (see the notes on the two lists). A caller that wants signers has to - /// deduplicate by key itself. - pub fn num_signature_claims(&self) -> usize { - self.xmss_signers.iter().map(|group| group.keys.len()).sum::() + self.sphincs_signers.len() - } - - /// The wire format: the signer set (each group with its epoch and - /// message), the DA root list, the two deferred points, and the VM proof. The claim *values* are not transmitted; - /// [`Self::from_bytes`] recomputes them, so there is nothing to lie about. - pub fn to_bytes(&self) -> Vec { - wire() - .serialize(&((&self.xmss_signers, &self.sphincs_signers), self.core())) - .expect("an aggregate serializes") - } - - /// Parsing does NOT verify: the proof is untouched, only shapes are checked. - /// Call [`Self::verify`] before believing any of it. - pub fn from_bytes(bytes: &[u8]) -> Result { - let (keys, core): (SignatureClaims, WireCore) = wire() - .deserialize(bytes) - .map_err(|_| AggregateVerifyError::MalformedEncoding)?; - Self::from_parts(keys, core) - } - - /// Without the signer set, for a receiver that already knows it. A set other - /// than the one aggregated fails verification. - pub fn to_bytes_without_pubkeys(&self) -> Vec { - wire().serialize(&self.core()).expect("an aggregate serializes") - } - - pub(crate) fn proof(&self) -> &lean_vm::cpu::Proof { - &self.proof - } - - /// Inverse of [`Self::to_bytes_without_pubkeys`], verifying nothing either. - pub fn from_bytes_without_pubkeys(bytes: &[u8], keys: SignatureClaims) -> Result { - Self::from_parts( - keys, - wire() - .deserialize(bytes) - .map_err(|_| AggregateVerifyError::MalformedEncoding)?, - ) - } - - fn core(&self) -> (&[[u8; 32]], &[F192], &[F192], &lean_vm::cpu::Proof) { - ( - &self.da_roots, - &self.defer.bytecode_point, - &self.defer.matrix_point, - &self.proof, - ) - } - - fn from_parts(keys: SignatureClaims, core: WireCore) -> Result { - let SignatureClaims { - xmss: xmss_signers, - sphincs: sphincs_signers, - } = keys; - let (da_roots, bytecode_point, matrix_point, proof) = core; - // Cheap rejections first. `recompute` below is a pass over the whole stacked - // bytecode plus a walk of the BLAKE2s circuit, on points a peer chose, so - // anything decidable without it has to be decided before it. - check_signer_set(&xmss_signers, &sphincs_signers)?; - check_da_roots(&da_roots)?; - Ok(Self { - xmss_signers, - sphincs_signers, - da_roots, - // The wire carries no value, only the points to derive it from. - defer: DeferredClaim::recompute(bytecode_point, matrix_point)?, - proof, - }) - } - - /// Verify the aggregate's internal consistency: the signer set and DA root list are well - /// formed, the three deferred fixed-polynomial claims hold at their - /// transmitted points, and the VM proof satisfies the statement built from - /// all of it. - /// - /// This says "every key in `xmss_signers` signed its group's message at its - /// group's epoch, and every `(key, message)` in `sphincs_signers` is a valid - /// SPHINCS signature", with the epochs and messages chosen by whoever - /// produced the aggregate: an aggregate over the same keys at different - /// epochs, or under different messages, verifies just as well. A caller that - /// expects particular pairs has to check the two lists against them. Every root in - /// [`Self::da_commitments`] also attests to a well-formed blob matrix; callers check - /// that this list contains the commitments they require. - pub fn verify(&self) -> Result<(), AggregateVerifyError> { - check_signer_set(&self.xmss_signers, &self.sphincs_signers)?; - check_da_roots(&self.da_roots)?; - // Discharging the deferred claims IS recomputing them: their values have - // to be the true evaluations, or the recursion below proves nothing about - // the fixed polynomials. - let _s = tracing::info_span!("Recompute deferred claims").entered(); - let defer = DeferredClaim::recompute(self.defer.bytecode_point.clone(), self.defer.matrix_point.clone())?; - drop(_s); - if defer != self.defer { - return Err(AggregateVerifyError::MalformedClaim); - } - let _s = tracing::info_span!("Statement digest").entered(); - let pi = self.public_input(); - drop(_s); - verify(unified_guest(), &pi, &self.proof).map_err(AggregateVerifyError::Snark)?; - Ok(()) - } -} -/// The stacked bytecode polynomial of the aggregation guest: the one fixed -/// table every node's bytecode claims are about. Cached, because verification -/// evaluates it and building it walks the whole program. -fn stacked_bytecode() -> &'static [F64] { - static TABLE: std::sync::OnceLock> = std::sync::OnceLock::new(); - TABLE.get_or_init(|| lean_vm::cpu::layout::bytecode_table(&unified_guest().prog)) -} - -/// The slots of the stacked bytecode that are not structurally zero. -/// -/// Eight encoding columns sit inside sixteen stacking slots, so half the -/// table is zero and contributes nothing to any round of the batching sumcheck. -/// Read off the table rather than from the column count, so an all-zero column -/// at the edge only ever shrinks the window. -fn bytecode_window() -> Range { - static WINDOW: std::sync::OnceLock> = std::sync::OnceLock::new(); - WINDOW - .get_or_init(|| { - let table = stacked_bytecode(); - let kbc = bytecode_vars() - lean_vm::leaf::N_BYTECODE_SELECTORS; - let live = |s: usize| table[s << kbc..(s + 1) << kbc].iter().any(|v| *v != F64::ZERO); - let slots = 1 << lean_vm::leaf::N_BYTECODE_SELECTORS; - let start = (0..slots).find(|&s| live(s)).expect("the bytecode is not all zero"); - let end = (0..slots).rfind(|&s| live(s)).expect("the bytecode is not all zero") + 1; - start..end - }) - .clone() -} - -/// One round message against a K-valued table: `bt` is the bytecode itself, so -/// the products are base-by-extension. -fn round_msg_base(bytecode: &[F64], weights: &[F192]) -> (F192, F192) { - let half = bytecode.len() / 2; - let term = |i: usize| { - let (bytecode_0, bytecode_1) = (bytecode[2 * i], bytecode[2 * i + 1]); - let (weight_0, weight_1) = (weights[2 * i], weights[2 * i + 1]); - ( - weight_1.mul_base(bytecode_1), - (weight_0 + weight_1).mul_base(bytecode_0 + bytecode_1), - ) - }; - if half >= PAR_MIN { - parallel::fold_reduce( - half, - || (F192::ZERO, F192::ZERO), - |acc: &mut (F192, F192), i| { - let (x, y) = term(i); - acc.0 += x; - acc.1 += y; - }, - |a: (F192, F192), b: (F192, F192)| (a.0 + b.0, a.1 + b.1), - ) - } else { - (0..half).fold((F192::ZERO, F192::ZERO), |acc, i| { - let (x, y) = term(i); - (acc.0 + x, acc.1 + y) - }) - } -} - -/// The first fold of a K-valued table, which is where it becomes extension-valued. -fn fold_lsb_base(table: &[F64], challenge: F192) -> Vec { - parallel::map_collect(table.len() / 2, |i| { - let (left, right) = (table[2 * i], table[2 * i + 1]); - F192::from(left) + challenge.mul_base(left + right) - }) -} - -/// `Σ_t γ_t · eq(points[t], (r_row, ·))` over the stacking slots, once the row -/// variables are bound. A closed form, so the row rounds never have to carry the -/// slot half of a `2^kbcv` weight table. -fn slot_weights(points: &[Vec], lambdas: &[F192], r_row: &[F192], kbc: usize) -> Vec { - let slots = lean_vm::leaf::N_BYTECODE_SELECTORS; - let mut weights = vec![F192::ZERO; 1 << slots]; - for (point, &lambda) in points.iter().zip(lambdas) { - let row_weight: F192 = (0..kbc).fold(lambda, |acc, k| acc * (F192::ONE + point[k] + r_row[k])); - for (slot, weight) in weights.iter_mut().enumerate() { - let slot_weight = (0..slots).fold(row_weight, |acc, bit| { - let coordinate = point[kbc + bit]; - acc * if (slot >> bit) & 1 == 1 { - coordinate - } else { - F192::ONE + coordinate - } - }); - *weight += slot_weight; - } - } - weights -} - -/// Variables of a bytecode claim's point: the log row count plus the stacking -/// selectors. -fn bytecode_vars() -> usize { - stacked_bytecode().len().trailing_zeros() as usize -} - -/// Below this a parallel dispatch costs more than the loop it replaces. -const PAR_MIN: usize = 1 << 16; - -fn fold_lsb(table: &mut Vec, challenge: F192) { - let half = table.len() / 2; - if half >= PAR_MIN { - let source: &[F192] = table; - let folded = parallel::map_collect(half, |i| { - let (left, right) = (source[2 * i], source[2 * i + 1]); - left + challenge * (left + right) - }); - *table = folded; - return; - } - for i in 0..half { - table[i] = table[2 * i] + challenge * (table[2 * i] + table[2 * i + 1]); - } - table.truncate(half); -} - -/// Compressed product-sumcheck round message over γ-weighted table pairs: -/// (g1, g∞) with g0 recovered from the running claim. -fn round_msg(pairs: &[(&[F192], &[F192], F192)]) -> (F192, F192) { - let (mut g1, mut gi) = (F192::ZERO, F192::ZERO); - for &(u, m, lambda) in pairs { - let half = u.len() / 2; - let term = |i: usize| { - let (u0, u1) = (u[2 * i], u[2 * i + 1]); - let (m0, m1) = (m[2 * i], m[2 * i + 1]); - (u1 * m1, (u0 + u1) * (m0 + m1)) - }; - let (a1, ai) = if half >= PAR_MIN { - parallel::fold_reduce( - half, - || (F192::ZERO, F192::ZERO), - |acc: &mut (F192, F192), i| { - let (x, y) = term(i); - acc.0 += x; - acc.1 += y; - }, - |a: (F192, F192), b: (F192, F192)| (a.0 + b.0, a.1 + b.1), - ) - } else { - (0..half).fold((F192::ZERO, F192::ZERO), |acc, i| { - let (x, y) = term(i); - (acc.0 + x, acc.1 + y) - }) - }; - g1 += lambda * a1; - gi += lambda * ai; - } - (g1, gi) -} - -/// One round of a batching sumcheck, in the transcript order the guest mirrors -/// word for word: observe `(g1, g_inf)`, squeeze the fold challenge, record -/// both, and advance the running claim through the compressed round polynomial. -/// Returns the challenge, which the caller folds its tables with. -fn absorb_round( - transcript: &mut FiatShamirState, - messages: &mut Vec, - challenges: &mut Vec, - claim: &mut F192, - (g1, gi): (F192, F192), -) -> F192 { - transcript.observe(g1); - transcript.observe(gi); - let challenge = transcript.sample(); - messages.extend([g1, gi]); - challenges.push(challenge); - let g0 = *claim + g1; - let c1 = g0 + g1 + gi; - *claim = (gi * challenge + c1) * challenge + g0; - challenge -} - -/// `Σ_t γ_t · eq(points[t], ·)`, over the `active` window of `2^vars` entries. -/// -/// Each point splits in half; the halves' eq tables are small and serial, and -/// the full table is their outer product, one fused multiply-add per entry, -/// parallel over the high index. Materializing a `2^vars` eq table per claim and -/// summing them is the same arithmetic done twice, serially. -fn weighted_eq_table(points: &[Vec], lambdas: &[F192], vars: usize, active: Range) -> Vec { - let lo_vars = vars / 2; - let lo_len = 1usize << lo_vars; - debug_assert!(active.start.is_multiple_of(lo_len) && active.end.is_multiple_of(lo_len)); - // `build_eq_table_ext` is LSB-first, so the low variables index the low bits - // and entry `hi * lo_len + lo` is `eq_lo[lo] * eq_hi[hi]`. - let halves: Vec<(Vec, Vec)> = points - .iter() - .zip(lambdas) - .map(|(point, &lambda)| { - let low = pcs::whir::build_eq_table_ext(&point[..lo_vars]); - let mut high = pcs::whir::build_eq_table_ext(&point[lo_vars..]); - high.iter_mut().for_each(|weight| *weight *= lambda); - (low, high) - }) - .collect(); - let mut weights = vec![F192::ZERO; active.len()]; - let first = active.start / lo_len; - parallel::chunks_mut(&mut weights, lo_len, |high_index, chunk| { - for (low, high) in &halves { - let scale = high[first + high_index]; - for (output, &low_weight) in chunk.iter_mut().zip(low) { - *output += scale * low_weight; - } - } - }); - weights -} - -/// Mirror the guest's `aggregate_claims` transcript and prove the two batching -/// sumchecks: dense for the bytecode, two-phase sparse for the matrices. -/// -/// Each child brings two claims per fixed polynomial: the one it deferred and -/// the fresh one its verification raised. They differ -/// only in the weight they enter with. A fresh matrix claim carries flock's -/// zerocheck/lincheck structure; a carried one is a plain point, so its weight -/// is an eq table on each side. Returns the guest hints and the single claim -/// per polynomial they reduce to. -fn aggregate_deferred_claims( - subproofs: &[DeferredSubproof], - carried_claims: &[&DeferredClaim], -) -> (SubHints, DeferredClaim) { - let child_count = subproofs.len(); - assert_eq!(child_count, carried_claims.len(), "one carried claim per child"); - let kbcv = bytecode_vars(); - let klog = flock::hash::K_LOG; - - let mut transcript = FiatShamirState::from_label(RECURSION_AGG_LABEL); - transcript.observe(count(child_count).into()); - for (subproof, carried) in subproofs.iter().zip(carried_claims) { - // A public input is observed as two scalars, one 128-bit half each. - let pi = subproof.public_input; - transcript.observe(F192::new(pi[0].0, pi[1].0, 0)); - transcript.observe(F192::new(pi[2].0, pi[3].0, 0)); - for &value in &subproof.bytecode_row_point { - transcript.observe(value); - } - for &value in &subproof.bytecode_selector_point { - transcript.observe(value); - } - transcript.observe(subproof.bytecode_value); - transcript.observe(subproof.matrix_a_coefficient); - transcript.observe(subproof.skip_point); - for &value in &subproof.zerocheck_row_point { - transcript.observe(value); - } - for &value in &subproof.lincheck_round_point { - transcript.observe(value); - } - for &value in &subproof.lincheck_terminal_values { - transcript.observe(value); - } - transcript.observe(subproof.matrix_claim); - for value in carried.cells() { - transcript.observe(value); - } - } - - let _span = tracing::info_span!("Bytecode batch", vars = kbcv).entered(); - let gbc: Vec = (0..2 * child_count).map(|_| transcript.sample()).collect(); - let points: Vec> = subproofs - .iter() - .zip(carried_claims) - .flat_map(|(subproof, carried)| { - [ - subproof - .bytecode_row_point - .iter() - .chain(&subproof.bytecode_selector_point) - .copied() - .collect::>(), - carried.bytecode_point.clone(), - ] - }) - .collect(); - let values: Vec = subproofs - .iter() - .zip(carried_claims) - .flat_map(|(subproof, carried)| [subproof.bytecode_value, carried.bytecode_value]) - .collect(); - let mut brun: F192 = (0..2 * child_count) - .map(|index| gbc[index] * values[index]) - .fold(F192::ZERO, |acc, value| acc + value); - let mut bscr = Vec::new(); - let mut r_bc = Vec::new(); - // The row variables, over the populated slot window only: the rest of the - // stacked table is structurally zero and contributes nothing to any round - // message, and folding LSB-first pairs entries within a slot, so the window's - // blocks stay aligned all the way down. - let n_slots = 1 << lean_vm::leaf::N_BYTECODE_SELECTORS; - let slot_window = bytecode_window(); - let kbc = kbcv - lean_vm::leaf::N_BYTECODE_SELECTORS; - let mut wt = weighted_eq_table( - &points, - &gbc, - kbcv, - (slot_window.start << kbc)..(slot_window.end << kbc), - ); - // Round zero runs against the K-valued table itself: the bytecode never needs - // to exist as `2^kbcv` extension elements (three times the memory traffic), - // and base-by-extension is cheaper than extension-by-extension. Every later - // round is extension-valued anyway. - let bc = &stacked_bytecode()[(slot_window.start << kbc)..(slot_window.end << kbc)]; - let mut bt = { - let msg = round_msg_base(bc, &wt); - let r = absorb_round(&mut transcript, &mut bscr, &mut r_bc, &mut brun, msg); - let folded = fold_lsb_base(bc, r); - fold_lsb(&mut wt, r); - folded - }; - for _ in 1..kbc { - let msg = round_msg(&[(&bt, &wt, F192::ONE)]); - let r = absorb_round(&mut transcript, &mut bscr, &mut r_bc, &mut brun, msg); - fold_lsb(&mut bt, r); - fold_lsb(&mut wt, r); - } - // The slot variables. The window has folded to one entry per populated slot; - // put those back among the zeros. The weights come from a closed form rather - // than from having carried a `2^kbcv` table through the rows. - let mut bt_slots = vec![F192::ZERO; n_slots]; - bt_slots[slot_window.clone()].copy_from_slice(&bt); - let wt_slots = slot_weights(&points, &gbc, &r_bc, kbc); - let (mut bt, mut wt) = (bt_slots, wt_slots); - for _ in 0..lean_vm::leaf::N_BYTECODE_SELECTORS { - let msg = round_msg(&[(&bt, &wt, F192::ONE)]); - let r = absorb_round(&mut transcript, &mut bscr, &mut r_bc, &mut brun, msg); - fold_lsb(&mut bt, r); - fold_lsb(&mut wt, r); - } - let v_bc = bt[0]; - assert_eq!(brun, v_bc * wt[0], "bytecode sumcheck terminal"); - - drop(_span); - let _span = tracing::info_span!("Matrix batch").entered(); - // One group per claim: the row weights, the column weights, and the - // coefficient each matrix enters with. Downstream is shape-blind. - let mut us: Vec> = Vec::with_capacity(2 * child_count); - let mut ws: Vec> = Vec::with_capacity(2 * child_count); - let mut ga: Vec = Vec::with_capacity(2 * child_count); - let mut gb: Vec = Vec::with_capacity(2 * child_count); - let mut mrun = F192::ZERO; - for (subproof, carried) in subproofs.iter().zip(carried_claims) { - let (gf, cga, cgb) = (transcript.sample(), transcript.sample(), transcript.sample()); - us.push(flock::lincheck::build_quirky_eq_table( - subproof.skip_point, - &subproof.zerocheck_row_point, - 6, - )); - ws.push( - (0..1usize << klog) - .map(|col| { - let mut w = subproof.lincheck_terminal_values[col & 63]; - for (j, &rj) in subproof.lincheck_round_point.iter().enumerate() { - let bit = (col >> (klog - 1 - j)) & 1; - w *= if bit == 1 { rj } else { F192::ONE + rj }; - } - w - }) - .collect(), - ); - ga.push(gf); - gb.push(gf * subproof.matrix_a_coefficient); - us.push(pcs::whir::build_eq_table_ext(&carried.matrix_point[..klog])); - ws.push(pcs::whir::build_eq_table_ext(&carried.matrix_point[klog..])); - ga.push(cga); - gb.push(cgb); - mrun += gf * subproof.matrix_claim + cga * carried.matrix_a_value + cgb * carried.matrix_b_value; - } - let _cols = tracing::info_span!("Contract columns").entered(); - // One forward walk of the circuit per claim yields that claim's two row - // tables `(A_0 w, B_0 w)` directly, in O(circuit): no matrix, and no pass - // over the ~89M nonzeros. A before B, the order `ga`/`gb` index. - let mut ms: Vec> = Vec::with_capacity(2 * ws.len()); - for w in &ws { - let (ra, rb) = flock::hash::row_values_walk(w); - ms.push(ra); - ms.push(rb); - } - drop(_cols); - // sanity: every claim really is the bilinear form over the matrices. - #[cfg(debug_assertions)] - for t in 0..2 * child_count { - let form = |m: &[F192]| { - m.iter() - .zip(&us[t]) - .map(|(&m, &u)| m * u) - .fold(F192::ZERO, |a, x| a + x) - }; - let (fa, fb) = (form(&ms[2 * t]), form(&ms[2 * t + 1])); - if t % 2 == 0 { - let subproof = &subproofs[t / 2]; - assert_eq!( - fa + subproof.matrix_a_coefficient * fb, - subproof.matrix_claim, - "fresh matrix claim, child {}", - t / 2 - ); - } else { - let carried = &carried_claims[t / 2]; - assert_eq!( - (fa, fb), - (carried.matrix_a_value, carried.matrix_b_value), - "carried matrix claim" - ); - } - } - let mut mscr = Vec::new(); - let mut r_row = Vec::new(); - let _rounds = tracing::info_span!("Row rounds").entered(); - for _ in 0..klog { - let pairs: Vec<(&[F192], &[F192], F192)> = (0..2 * child_count) - .flat_map(|t| { - [ - (&us[t][..], &ms[2 * t][..], ga[t]), - (&us[t][..], &ms[2 * t + 1][..], gb[t]), - ] - }) - .collect(); - let msg = round_msg(&pairs); - let r = absorb_round(&mut transcript, &mut mscr, &mut r_row, &mut mrun, msg); - for u in us.iter_mut() { - fold_lsb(u, r); - } - for m in ms.iter_mut() { - fold_lsb(m, r); - } - } - drop(_rounds); - let eq_rstar = pcs::whir::build_eq_table_ext(&r_row); - let _rows = tracing::info_span!("Contract rows").entered(); - // `A_0ᵀ eq` and `B_0ᵀ eq` are the column marginals, which the circuit walks - // backwards (`gf2`'s `back_*`) in O(circuit). - let (mut acol, mut bcol) = flock::hash::marginal_walk_pair(&eq_rstar); - drop(_rows); - let mut wa = vec![F192::ZERO; 1 << klog]; - let mut wb = vec![F192::ZERO; 1 << klog]; - for t in 0..2 * child_count { - let (sa, sb) = (ga[t] * us[t][0], gb[t] * us[t][0]); - for j in 0..1 << klog { - wa[j] += sa * ws[t][j]; - wb[j] += sb * ws[t][j]; - } - } - let mut r_col = Vec::new(); - for _ in 0..klog { - let pairs: Vec<(&[F192], &[F192], F192)> = vec![(&acol, &wa, F192::ONE), (&bcol, &wb, F192::ONE)]; - let msg = round_msg(&pairs); - let r = absorb_round(&mut transcript, &mut mscr, &mut r_col, &mut mrun, msg); - for tb in [&mut acol, &mut bcol, &mut wa, &mut wb] { - fold_lsb(tb, r); - } - } - let (v_a, v_b) = (acol[0], bcol[0]); - assert_eq!(mrun, v_a * wa[0] + v_b * wb[0], "matrix sumcheck terminal"); - // The guest reaches the same two weights by a succinct formula rather than by - // folding these tables, and nothing else compares the two: the aggregation - // layer has no third implementation the way `cpu::verify` does. So this runs - // unconditionally (it is O(n) against a 2^28 sumcheck), and it compares the - // weights COMPONENT-WISE. Checking only the combination `v_a·wa + v_b·wb` - // would let two correlated errors through. - { - let eqr = pcs::whir::build_eq_table_ext(&r_row[..6]); - let eqc = pcs::whir::build_eq_table_ext(&r_col[..6]); - let (mut wam, mut wbm) = (F192::ZERO, F192::ZERO); - for (t, (subproof, carried)) in subproofs.iter().zip(carried_claims).enumerate() { - let lam = primitives::multilinear::lagrange_weights_naive(6, subproof.skip_point); - let mut urow: F192 = (0..64).map(|i| lam[i] * eqr[i]).fold(F192::ZERO, |a, x| a + x); - for (k, &z) in subproof.zerocheck_row_point.iter().enumerate() { - urow *= F192::ONE + z + r_row[6 + k]; - } - let mut wcol: F192 = (0..64) - .map(|i| subproof.lincheck_terminal_values[i] * eqc[i]) - .fold(F192::ZERO, |a, x| a + x); - for (j, &rj) in subproof.lincheck_round_point.iter().enumerate() { - wcol *= F192::ONE + rj + r_col[klog - 1 - j]; - } - let fresh = urow * wcol; - let mut plain = F192::ONE; - for (k, &p) in carried.matrix_point.iter().enumerate() { - let r = if k < klog { r_row[k] } else { r_col[k - klog] }; - plain *= F192::ONE + p + r; - } - wam += ga[2 * t] * fresh + ga[2 * t + 1] * plain; - wbm += gb[2 * t] * fresh + gb[2 * t + 1] * plain; - } - assert_eq!(wa[0], wam, "guest row-weight formula for A0"); - assert_eq!(wb[0], wbm, "guest row-weight formula for B0"); - } - - drop(_span); - let hints = vec![ - ("bc_sumcheck_msgs", limbs(&bscr)), - ("mat_sumcheck_msgs", limbs(&mscr)), - ("bc_star_hint", limbs(&[v_bc])), - ("mat_stars_hint", limbs(&[v_a, v_b])), - ]; - ( - hints, - DeferredClaim { - bytecode_point: r_bc, - bytecode_value: v_bc, - matrix_point: r_row.iter().chain(&r_col).copied().collect(), - matrix_a_value: v_a, - matrix_b_value: v_b, - }, - ) -} -/// The verifier-side WHIR config for one committed size and rate, plus the -/// query packing derived from it. The hint builder needs it for the real -/// opening and the placeholder map for every candidate size, so it lives here: -/// a candidate whose shape differed from the real one would compile a guest -/// that cannot open the proof it is handed. -struct WhirShape { - config: pcs::whir::VerifierConfig, - levels: pcs::whir::LevelShapes, - /// Merkle tree depth per level. - depth: Vec, - /// Query positions carried by one squeezed F192 word, per level. - per_squeeze: Vec, -} - -fn whir_shape(mu: usize, log_inv_rate: usize) -> WhirShape { - let config = pcs::whir::config_for_rate(mu, log_inv_rate).expect("stacked whir config"); - let levels = config.level_shapes(mu); - let depth: Vec = levels.block_len.iter().map(|b| b.trailing_zeros() as usize).collect(); - let per_squeeze = depth.iter().map(|&d| 192 / d).collect(); - WhirShape { - config, - levels, - depth, - per_squeeze, - } -} - -fn merkle_cap_depth(queries: usize, depth: usize) -> usize { - (queries.next_power_of_two().ilog2() as usize).min(depth) -} - -/// The BLAKE2s table's virtual value columns, in `hash_flock::SLOTS` order. -fn blake2s_value_columns() -> Vec { - let base = lean_vm::cpu::schema().base[5]; - lean_vm::tables::BLAKE2S_VALUE_COLS.iter().map(|&c| base + c).collect() -} - -/// One entry of the guest's claim pool. -enum ClaimSite { - /// A committed column read by a framework bus block. - Framework { column: usize }, - /// A table column; `is_virtual` marks the q_flock-backed value - /// columns, whose claim is a strided slot rather than a plain column. - TableColumn { column: usize, is_virtual: bool }, - /// The public-input binding claim on `MEM`. - PublicInput { column: usize }, -} - -/// The guest's `COORD_KIND_*` code for a coordinate (`guests/lean_ethereum.py`), -/// shared by its `COORD_TYPE` and `TERM_TYPE` arrays. -fn coord_kind(c: &Coord) -> usize { - match c { - Coord::Const(_) => 0, - Coord::Col(_) => 1, - Coord::GCol(..) => 2, - Coord::Index => 3, - Coord::Public(_) => 4, - Coord::Prod(..) => 5, - Coord::Sum(..) => 6, - } -} - -/// The `K` scalar a coordinate carries beside its columns: the constant itself, -/// or the `g^k` a `GCol`/`Prod` scales by. Zero for every other kind. -fn coord_scale(c: &Coord) -> F64 { - match c { - Coord::Const(v) => *v, - Coord::GCol(_, k) | Coord::Prod(_, _, k) => g_pow(*k as usize), - _ => F64::ZERO, - } -} - -/// Flatten one table-block coordinate into the guest's term arrays, in local -/// column indices. A [`Coord::Sum`]'s children are its terms; every other kind is -/// one term. `Index`/`Public` never reach a table block. -fn push_coord_terms(c: &Coord, base: usize, terms: &mut Vec) { - let (column_a, column_b) = match c { - Coord::Const(_) => (0, 0), - Coord::Col(i) | Coord::GCol(i, _) => (*i - base, 0), - Coord::Prod(i, j, _) => (*i - base, *j - base), - Coord::Sum(cs) => { - for c in cs { - push_coord_terms(c, base, terms); - } - return; - } - Coord::Index | Coord::Public(_) => unreachable!("a table's bus block carries no virtual coordinate"), - }; - terms.push(Term { - kind: coord_kind(c), - constant: coord_scale(c).0, - column_a, - column_b, - }); -} - -/// Visit the claim pool in the exact order the guest indexes it: the framework -/// bus claims (deduped by `(column, kappa)`, as `leaf.rs` pools them), then every -/// table's committed columns, then the PI memory claim. The placeholder map's -/// claim descriptors follow this order. -fn walk_claims(layout: &lean_vm::cpu::Layout, kbc: usize, mut visit: impl FnMut(ClaimSite)) { - let sides: [&[Block]; 3] = [&layout.push, &layout.pull, &layout.count]; - let valcols = blake2s_value_columns(); - // Only the framework blocks raise claims: a table's coords are settled inside - // the table sumcheck. - let is_framework: Vec = lean_vm::cpu::block_kappa_sources(kbc) - .into_iter() - .map(|(src, _)| src < 2) - .collect(); - let mut seen: std::collections::HashSet<(usize, usize)> = Default::default(); - let mut bi = 0usize; - for blocks in sides.iter() { - for blk in blocks.iter() { - let framework = is_framework[bi]; - bi += 1; - if !framework { - continue; - } - for c in &blk.coords { - if let Coord::Col(i) | Coord::GCol(i, _) = c { - if !seen.insert((*i, blk.kappa)) { - continue; // deduped: pooled once at its first occurrence - } - assert!(!valcols.contains(i), "{VALCOL_FRAMEWORK}"); - visit(ClaimSite::Framework { column: *i }); - } - } - } - } - let sch = lean_vm::cpu::schema(); - for (t, table) in lean_vm::tables::tables().iter().enumerate() { - for c in 0..table.n_committed_columns() { - let column = sch.base[t] + c; - visit(ClaimSite::TableColumn { - column, - is_virtual: layout.placements[column].is_virtual(), - }); - } - } - visit(ClaimSite::PublicInput { - column: lean_vm::cpu::MEM, - }); -} - -/// Config + hints for the recursion guest (`guests/lean_ethereum.py`), built -/// from the REAL `cpu::layout` of the inner program and the summary of a real -/// `cpu::verify` run (zero hand-mirroring drift). -fn gen_verify( - program: &Program, - public_input: [F64; 4], - summary: lean_vm::cpu::VerifySummary, -) -> Result<(SubHints, DeferredSubproof), AggregationError> { - let proof_stream = &summary.raw.stream; - let layout = lean_vm::cpu::layout( - &program.prog, - proof_stream[0].c0 as usize, - std::array::from_fn(|i| proof_stream[1 + i].c0 as usize), - public_input, - ); - let sides: [&[Block]; 3] = [&layout.push, &layout.pull, &layout.count]; - let side_layouts = sides.map(lean_vm::leaf::layout); - // Fixed capacities: every buffer/stride placeholder is a global cap so - // the placeholder map is SHAPE-INDEPENDENT (the definition of generic). - assert!(side_layouts.iter().all(|side| side.mu <= MU_CAP) && 3 * proof_stream.len() <= STREAM_CAP); - // The guest holds one opening arm per candidate committed size, so a child - // outside that window has no arm to dispatch to. `min_log_committed` keeps - // every aggregate above the low end, leaving only the ceiling reachable. - if !(MU_MIN..=MU_MAX).contains(&layout.shape.mu) { - return Err(AggregationError::ChildOutOfRange { - log_committed: layout.shape.mu, - }); - } - - // ---- typed extraction: proof structs + the verifier's summary ---- - // Push and pull share the bytecode point. - let kbc = summary.bytecode_claim.point.len() - lean_vm::leaf::N_BYTECODE_SELECTORS; - - let taus = layout.taus; - // Flock replay data, all named struct fields. - let lcrounds = flock::hash::K_LOG - 6; - let zcf = [summary.zc_claim.a_eval, summary.zc_claim.b_eval]; - let zc_z = summary.zc_claim.z; - let zchi = &summary.zc_claim.mlv_challenges; - let lc_alpha = summary.lc_claim.alpha; - let lc_beta = summary.lc_claim.beta; - let lrr = &summary.lc_claim.r_rounds; - - // ---- the stacked opening: config + the opening summary ---- - let stack = whir_shape(layout.shape.mu, summary.log_inv_rate); - - // flock's reduction ends at `flock_stream_end`, where the WHIR opening's own - // scalars start: its last 64 scalars are lincheck's `z_partial` (which the - // summary already carries as `s_hat_v`), immediately preceded by the - // coefficient PAIRS of the `lcrounds` lincheck rounds: the linear one is not - // sent, the running claim fixing it. - let ns = summary.flock_stream_end; - let lcr = &proof_stream[ns - 64 - 2 * lcrounds..ns - 64]; - let lcz = &summary.lc_claim.s_hat_v; - - // matpart = the deferred weighted matrix evaluation: the lincheck running - // claim minus (= plus, char 2) the const-pin and c-claim contributions. - // α² from α, not from β: `LincheckClaim::beta` (the pin, at α³) is zero for - // a circuit with no const-pin column, while every verifier draws the - // c-claim's coefficient unconditionally. - let lc_sq = lc_alpha.square(); - let mut lrun = zcf[0] + lc_alpha * zcf[1] + lc_sq * summary.zc_claim.c_eval + lc_beta; - for i in 0..lcrounds { - let (c0, c2) = (lcr[2 * i], lcr[2 * i + 1]); - lrun = primitives::multilinear::poly_eval(&[c0, lrun + c2, c2], lrr[i]); - } - let mut pinw = lc_beta; - for (j, &rv) in lrr.iter().enumerate() { - let bit = (flock::hash::Z_CONST_POS >> (flock::hash::K_LOG - 1 - j)) & 1; - pinw *= if bit == 1 { rv } else { F192::ONE + rv }; - } - pinw *= lcz[flock::hash::Z_CONST_POS % 64]; - // The c term: eq(ρ_in, ρ'_in) times the φ8-Lagrange combination of the 64 - // slices, ρ'_in being the lincheck challenges read back in coordinate order. - let mut c_point_eq = F192::ONE; - for (t, &rin) in zchi[..lcrounds].iter().enumerate() { - c_point_eq *= F192::ONE + rin + lrr[lcrounds - 1 - t]; - } - let c_slice_value = primitives::multilinear::lagrange_weights_naive(6, zc_z) - .iter() - .zip(lcz) - .fold(F192::ZERO, |acc, (&w, &s)| acc + w * s); - let matpart = lrun + pinw + lc_sq * c_point_eq * c_slice_value; - - // ---- hints ---- - // The program's whole share of a bytecode leaf: ONE value, the stacked - // polynomial at (ζ_lo, α⃗), the slot coordinates of the claim's own point being - // the fingerprint challenges (§sec:e2e-bc). - let bytecode_value = summary.bytecode_claim.value; - - // ---- per-sub HINT data (the placeholder map is built once, elsewhere) ---- - // Per side, the packing order read straight off `leaf::layout`'s offsets: - // sort_order[side_base + rank] = g^{side-local index of the rank-r block}. - // The guest only perm-checks it and derives offsets; any aligned tiling is - // sound, so this canonical order just has to match the committed leaf. - let mut sort_order: Vec = Vec::new(); - let mut gbase = 0usize; - for (s, blocks) in sides.iter().enumerate() { - let mut order: Vec = (0..blocks.len()).collect(); - order.sort_by_key(|&i| side_layouts[s].offsets[i]); - for &i in &order { - sort_order.push(g_pow(gbase + i)); // g^{global block index} - } - gbase += blocks.len(); - } - // The stacked commitment uses witness::placements_of: committed columns - // sorted by descending kappa, then by their native column index. Transport - // compact committed-column indices; the guest certifies the permutation, - // ordering, and accumulated offsets before using them as claim selectors. - let col_sources = lean_vm::cpu::col_kappa_sources(kbc); - let committed_globals: Vec = col_sources - .iter() - .enumerate() - .filter_map(|(i, source)| source.map(|_| i)) - .collect(); - let mut compact_col = vec![usize::MAX; col_sources.len()]; - for (compact, &global) in committed_globals.iter().enumerate() { - compact_col[global] = compact; - } - let mut col_order = committed_globals; - col_order.sort_by_key(|&global| layout.placements[global].offset); - let col_sort_order: Vec = col_order.iter().map(|&global| g_pow(compact_col[global])).collect(); - - // ---- Phase E2 hints (the stacked WHIR opening) ---- - // Share the upper Merkle tree across queries. Unknown, unused subtrees stay - // opaque; the guest authenticates every cap leaf a query reaches. - let mut query_hints = Vec::new(); - let (mut caps, mut cap_active) = (Vec::new(), Vec::new()); - let mut openings = summary.raw.merkle.iter(); - for (&queries, &depth) in stack.config.queries.iter().zip(&stack.depth) { - let cap_depth = merkle_cap_depth(queries, depth); - let n = 1 << cap_depth; - let path_depth = depth - cap_depth; - let mut nodes = vec![[0u8; 32]; 2 * n]; - let mut active = vec![F64::ZERO; n]; - for opening in openings.by_ref().take(queries) { - query_hints.push(("merkle_leaf_rows", opening.leaf_data.clone())); - let mut path_children = Vec::with_capacity(4 * path_depth); - let bytes: Vec<_> = opening.leaf_data.iter().flat_map(|x| x.0.to_le_bytes()).collect(); - let mut node = pcs::merkle::hash_leaf(&bytes); - let mut index = (1 << depth) + opening.leaf_index; - for (height, sibling) in opening.path.iter().enumerate() { - if height >= path_depth { - nodes[index] = node; - nodes[index ^ 1] = *sibling; - active[index >> 1] = F64::ONE; - } - let (left, right) = if index & 1 == 0 { - (&node, sibling) - } else { - (sibling, &node) - }; - if height < path_depth { - path_children.extend(digest_words(left)); - path_children.extend(digest_words(right)); - } - node = pcs::merkle::hash_pair(left, right); - index >>= 1; - } - nodes[1] = node; - query_hints.push(("merkle_children", path_children)); - } - caps.extend(nodes.iter().flat_map(digest_words)); - cap_active.extend(active); - } - let mut bytecode_row_point = summary.bytecode_claim.point; - let bytecode_selector_point = bytecode_row_point.split_off(kbc); - let deferred = DeferredSubproof { - public_input, - bytecode_row_point, - bytecode_selector_point, - bytecode_value, - matrix_a_coefficient: lc_alpha, - skip_point: zc_z, - zerocheck_row_point: zchi[..lcrounds].to_vec(), - lincheck_round_point: summary.lc_claim.r_rounds, - lincheck_terminal_values: summary.lc_claim.s_hat_v, - matrix_claim: matpart, - }; - - let mut hints = vec![ - ("stream", { - // The guest replays the WHIR opening off the same stream the native - // verifier reads: every transmitted scalar (sumcheck messages, level - // roots, OOD claims, grind nonces, `yr`) is already there in protocol - // order, so there is nothing to reassemble. The guest's `open_stacked` - // picks it up at `msg_cursor = cursor`, which sits where the flock - // reduction stopped; the ring-switch messages are struct-observed and - // still do not advance that cursor. - let mut stream = limbs(&summary.raw.stream); - assert!( - stream.len() <= STREAM_CAP, - "stream of {} words exceeds cap {STREAM_CAP}", - stream.len() - ); - stream.resize(STREAM_CAP, F64::ZERO); - stream - }), - ("bytecode_val", limbs(&[bytecode_value])), - ("matpart", limbs(&[matpart])), - ("merkle_caps", caps), - ("merkle_cap_active", cap_active), - // the table sumcheck's round count: max_t tau_t, certified in-guest as a - // maximum (one of the taus, and dominating them all). - ("zc_tau_max", vec![g_pow(*taus.iter().max().unwrap())]), - ("col_sort_order", col_sort_order), - ("sort_order", sort_order), - ]; - hints.extend(query_hints); - Ok((hints, deferred)) -} - -/// The guest's stacked-size dispatch range: one `match_range` opening arm per -/// candidate `mu` in `MU_MIN..=MU_MAX` (mirrored by the soundness test's -/// residual-log cap). -const MU_MIN: usize = 22; -const MU_MAX: usize = lean_vm::pcs::MAX_MU; - -const _: () = assert!(MU_MIN >= lean_vm::pcs::MIN_MU); - -/// The guest's baked buffer caps, which `placeholder_map` compiles in and -/// `gen_verify` admits against: one definition, so a hinted shape can never -/// outgrow the buffer the guest was compiled with. -const MU_CAP: usize = 40; -/// In words, three to a transcript scalar. -const STREAM_CAP: usize = 3 * 8192; -/// Named hint entries for a single sub-proof, ordered within each stream. -type SubHints = Vec<(&'static str, Vec)>; - -/// One `hint_witness` stream: a name and its entries, in the order the guest -/// pops them. -#[derive(Default)] -pub(crate) struct Hints(Vec<(String, Vec>)>); - -impl Hints { - #[cfg(test)] - fn is_empty(&self) -> bool { - self.0.is_empty() - } - - fn push(&mut self, name: &str, entry: Vec) { - match self.0.iter_mut().find(|(n, _)| n == name) { - Some((_, entries)) => entries.push(entry), - None => self.0.push((name.to_string(), vec![entry])), - } - } - - /// The entries of one stream, for the adversarial test to corrupt. - #[cfg(test)] - fn entries(&mut self, name: &str) -> &mut Vec> { - &mut self - .0 - .iter_mut() - .find(|(n, _)| n == name) - .unwrap_or_else(|| panic!("no hint stream `{name}`")) - .1 - } - - fn install(self, guest: &mut Program) { - for (name, entries) in self.0 { - guest.set_witness(name, entries); - } - } -} - -/// The coverage slot each write in the guest's coverage walk targets, in walk -/// order: the raw signatures first (grouped by epoch, as the guest loops over -/// them), then each child's key lists. -/// -/// The table is one contiguous region per XMSS epoch group, then the SPHINCS -/// pair, each region its declared keys followed by its own duplicate slots: -/// -/// | slots | holds | -/// | --- | --- | -/// | `[0, n_0 + d_0)` | epoch group 0: declared keys, then duplicates | -/// | ... | one such region per declared group, in epoch order | -/// | ... | then one per covered-but-undeclared group, all duplicates | -/// | `[X, X + n_sphincs)` | the declared SPHINCS keys | -/// | `[X + n_sphincs, n_total)` | SPHINCS duplicate slots | -/// -/// with `X` the sum of every group's slots. A key the declared set holds takes its -/// slot there, one it does not takes a fresh duplicate slot in the same region, so -/// the walk hits every one of the `n_total` slots exactly once. -/// That bijection, enforced in-circuit by write-once memory plus the final -/// count, is what makes every declared key covered by a real signature or a -/// verified child. -/// -/// Keeping each region contiguous is what binds the scheme and the epoch: the -/// guest addresses every writer as an offset into one region, bounded by that -/// region's size, one range check per write, so no signature can reach a key -/// declared under another epoch or the other scheme. -struct Coverage { - /// Declared groups in statement order, then undeclared ones, all duplicates. - xmss_groups: Vec, - /// How many of `xmss_groups` the signer set declares. - n_declared: usize, - /// One duplicate list per group, aligned with `xmss_groups`. - xmss_dups: Vec>, - sphincs_signers: Vec, - sphincs_dups: Vec, - /// Offsets within each raw signature's own group region, indexed as `raw_xmss`. - raw_xmss: Vec, - /// `raw_xmss` indices in the guest's walk order: it walks the table, whose - /// groups are declared-first rather than epoch-sorted. - raw_walk: Vec, - /// Offsets past `X`, in the SPHINCS region. - raw_sphincs: Vec, - /// Per child, per child group: the parent group it maps to, and the offset - /// of each child key within that parent group's region. - child_xmss: Vec)>>, - child_sphincs: Vec>, -} - -impl Coverage { - fn declared(&self) -> &[XmssClaimGroup] { - &self.xmss_groups[..self.n_declared] - } - - fn n_keys(&self) -> usize { - self.xmss_groups.iter().map(|group| group.keys.len()).sum::() + self.sphincs_signers.len() - } - - fn n_total(&self) -> usize { - self.n_keys() + self.xmss_dups.iter().map(Vec::len).sum::() + self.sphincs_dups.len() - } -} - -/// A claim takes its declared slot once; subsequent or omitted occurrences take -/// fresh slots outside the hashed prefix. -fn take_slot(claims: &[K], claimed: &mut [bool], duplicates: &mut Vec, claim: &K) -> usize { - match claims.binary_search(claim) { - Ok(pos) if !claimed[pos] => { - claimed[pos] = true; - pos - } - _ => { - duplicates.push(claim.clone()); - claims.len() + duplicates.len() - 1 - } - } -} - -/// Record one epoch's message, rejecting a second, different one: within an -/// aggregate the message is a function of the epoch. -fn bind_message( - messages: &mut BTreeMap, - epoch: xmss::Epoch, - message: xmss::Message, -) -> Result<(), AggregationError> { - match messages.insert(epoch, message) { - Some(previous) if previous != message => Err(AggregationError::ConflictingMessages), - _ => Ok(()), - } -} - -fn plan_coverage( - raw_xmss: &[(XmssPublicKey, xmss::Epoch, xmss::Message)], - raw_sphincs: &[SphincsClaim], - children: &[EthereumProof], - declare: Option<&SignatureClaims>, -) -> Result { - // The union, as `(epoch, key)` claims plus the epoch-to-message function - // every contributor must agree on, then grouped: consecutive equal epochs - // of the sorted deduplicated list are one group. - let mut messages: BTreeMap = BTreeMap::new(); - let mut claims: Vec<(xmss::Epoch, XmssPublicKey)> = Vec::with_capacity(raw_xmss.len()); - for (pk, epoch, message) in raw_xmss { - bind_message(&mut messages, *epoch, *message)?; - claims.push((*epoch, pk.clone())); - } - let mut sphincs_signers = raw_sphincs.to_vec(); - for child in children { - for XmssClaimGroup { epoch, message, keys } in &child.xmss_signers { - bind_message(&mut messages, *epoch, *message)?; - claims.extend(keys.iter().map(|pk| (*epoch, pk.clone()))); - } - sphincs_signers.extend_from_slice(&child.sphincs_signers); - } - claims.sort(); - claims.dedup(); - // On the whole pair, so one key signing two messages is two claims. - sphincs_signers.sort(); - sphincs_signers.dedup(); - let mut union_groups: Vec = Vec::new(); - for (epoch, pk) in claims { - match union_groups.last_mut() { - Some(group) if group.epoch == epoch => group.keys.push(pk), - _ => union_groups.push(XmssClaimGroup { - epoch, - message: messages[&epoch], - keys: vec![pk], - }), - } - } - // Groups the declaration holds nothing of go last, so the declared ones are the - // prefix the digest hashes. Claims are struck off, so leftovers are uncovered. - let mut wanted: BTreeSet<(xmss::Epoch, XmssPublicKey)> = BTreeSet::new(); - let mut wanted_sphincs: BTreeSet = BTreeSet::new(); - if let Some(SignatureClaims { xmss, sphincs }) = declare { - for XmssClaimGroup { epoch, message, keys } in xmss { - if messages.get(epoch) != Some(message) { - return Err(AggregationError::NotCovered); - } - wanted.extend(keys.iter().map(|key| (*epoch, key.clone()))); - } - wanted_sphincs.extend(sphincs.iter().copied()); - } - let mut xmss_groups: Vec = Vec::new(); - let mut covered_only: Vec = Vec::new(); - for mut group in union_groups { - group - .keys - .retain(|key| declare.is_none() || wanted.remove(&(group.epoch, key.clone()))); - if group.keys.is_empty() { - covered_only.push(group); - } else { - xmss_groups.push(group); - } - } - let n_declared = xmss_groups.len(); - xmss_groups.append(&mut covered_only); - if xmss_groups.len() > MAX_EPOCHS { - return Err(AggregationError::TooManyEpochs); - } - if declare.is_some() { - sphincs_signers.retain(|signer| wanted_sphincs.remove(signer)); - if !wanted.is_empty() || !wanted_sphincs.is_empty() { - return Err(AggregationError::NotCovered); - } - } - // The table is no longer sorted by epoch. - let region_of: BTreeMap = xmss_groups - .iter() - .enumerate() - .map(|(j, group)| (group.epoch, j)) - .collect(); - let mut xmss_claimed: Vec> = xmss_groups.iter().map(|group| vec![false; group.keys.len()]).collect(); - let mut xmss_dups: Vec> = vec![Vec::new(); xmss_groups.len()]; - let mut sphincs_claimed = vec![false; sphincs_signers.len()]; - let mut sphincs_dups = Vec::new(); - let raw_xmss_slots: Vec = raw_xmss - .iter() - .map(|(pk, epoch, _)| { - let g = region_of[epoch]; - take_slot(&xmss_groups[g].keys, &mut xmss_claimed[g], &mut xmss_dups[g], pk) - }) - .collect(); - let raw_sphincs_slots: Vec = raw_sphincs - .iter() - .map(|signer| take_slot(&sphincs_signers, &mut sphincs_claimed, &mut sphincs_dups, signer)) - .collect(); - let mut child_xmss = Vec::with_capacity(children.len()); - let mut child_sphincs = Vec::with_capacity(children.len()); - for child in children { - child_xmss.push( - child - .xmss_signers - .iter() - .map(|group| { - let g = region_of[&group.epoch]; - let offsets = group - .keys - .iter() - .map(|pk| take_slot(&xmss_groups[g].keys, &mut xmss_claimed[g], &mut xmss_dups[g], pk)) - .collect(); - (g, offsets) - }) - .collect::)>>(), - ); - child_sphincs.push( - child - .sphincs_signers - .iter() - .map(|signer| take_slot(&sphincs_signers, &mut sphincs_claimed, &mut sphincs_dups, signer)) - .collect(), - ); - } - // Stable, so signatures within a group keep the order their slots were taken in. - let mut raw_walk: Vec = (0..raw_xmss.len()).collect(); - raw_walk.sort_by_key(|&i| region_of[&raw_xmss[i].1]); - let cover = Coverage { - xmss_groups, - n_declared, - xmss_dups, - sphincs_signers, - sphincs_dups, - raw_xmss: raw_xmss_slots, - raw_walk, - raw_sphincs: raw_sphincs_slots, - child_xmss, - child_sphincs, - }; - if cover.n_total() >= MAX_KEYS { - return Err(AggregationError::TooLarge); - } - Ok(cover) -} -/// One signature's witness: the WOTS randomness, the encoding digits (in the -/// exponent), the chain tips they start from, and the Merkle siblings. -fn push_signature_hints( - hints: &mut Hints, - pk: &XmssPublicKey, - sig: &XmssSignature, - message: &xmss::Message, - xmss_epoch: xmss::Epoch, -) -> Result<(), AggregationError> { - let wots = &sig.wots_signature; - let encoding = xmss::wots_encode(message, xmss_epoch, &pk.public_param, &wots.randomness) - .ok_or(AggregationError::MalformedRawSignature)?; - hints.push("rand", words(&wots.randomness)); - for &e in &encoding { - hints.push("digits", vec![count(e as usize)]); - } - for tip in &wots.chain_tips { - hints.push("chain_starts", words(tip)); - } - for sibling in &sig.merkle_proof { - hints.push("siblings", words(sibling)); - } - Ok(()) -} - -/// One SPHINCS signature's witness: the randomizer, the few-time opening, and -/// per layer the encoding counter, the codeword digits (in the exponent), the -/// chain values they start from, and the Merkle siblings. -/// -/// The guest derives the index and the leaf indices from the digest itself, so -/// nothing here carries them; what it does carry is the per-layer message, which -/// this walk recomputes exactly as the guest will. The signer's own message is -/// not hinted either: it rides its slot in the coverage table. -fn push_sphincs_hints( - hints: &mut Hints, - (pk, message): &SphincsClaim, - sig: &SphincsSignature, -) -> Result<(), AggregationError> { - let pp = &pk.public_param; - hints.push("sp_rand", words(&sig.randomizer)); - let (idx, u) = sphincs::message_digest(pp, &pk.root, &sig.randomizer, message); - for kappa in 0..sphincs::NUM_FTS_TREES { - hints.push("sp_fts_secrets", words(&sig.fts.secrets[kappa])); - for sibling in &sig.fts.paths[kappa] { - hints.push("sp_fts_paths", words(sibling)); - } - } - let mut signed = sphincs::fts_recover(pp, idx, &u, &sig.fts); - for lay in (0..sphincs::D).rev() { - let pos = sphincs::Pos::new(lay, sphincs::tree_of(idx, lay), sphincs::leaf_of(idx, lay)); - let counter = sig.counters[lay]; - let codeword = sphincs::encode(pp, pos, &signed, counter).ok_or(AggregationError::MalformedRawSignature)?; - hints.push("sp_counter", vec![F64(u64::from(counter))]); - for (&digit, opened) in codeword.iter().zip(&sig.ots[lay]) { - hints.push("sp_digits", vec![count(digit as usize)]); - hints.push("sp_chain_starts", words(opened)); - } - let path = &sig.paths[sphincs::path_range(lay)]; - for sibling in path { - hints.push("sp_siblings", words(sibling)); - } - let leaf = sphincs::ots_leaf(pp, pos, &signed, counter, &sig.ots[lay]) - .ok_or(AggregationError::MalformedRawSignature)?; - signed = sphincs::tree_fold(pp, pos, leaf, path); - } - debug_assert_eq!(signed, pk.root, "the hinted walk reaches the public key"); - Ok(()) -} -#[derive(Clone, Copy, Default)] -pub(crate) struct DaInput<'a> { - pub rows: &'a [u64], - pub roots: Option<&'a [[u8; 32]]>, -} - -/// Prove existence of signatures and valid encoding of PQ, potentially using recursive children. -/// -/// - `children`: child proofs; at most [`MAX_RECURSIONS`]. -/// - `raw_xmss`: list of `(public_key, epoch, message, signature)`, any order; one message per -/// epoch across the whole result, at most [`MAX_EPOCHS`] epochs. -/// - `raw_sphincs`: list of `(public_key, message, signature)`, any order. -/// - `blobs`: concatenated blobs, each containing [`BLOB_SYMBOLS`] little-endian `u64` symbols; -/// at most [`DA_MAX_ROWS`] blobs. -/// - `declare`: `None` keeps all claims; `Some` specifies exactly the signatures and DA roots -/// to publish. Every declared claim must be covered by the inputs above. -/// - `log_inv_rate`: PCS code rate `2^-log_inv_rate`, higher `log_inv_rate` means a smaller proof -/// but slower proving; in [`MIN_LOG_INV_RATE`]..=[`MAX_LOG_INV_RATE`]. -/// -/// The combined XMSS and SPHINCS claim count, including duplicate coverage slots, must be strictly below [`MAX_KEYS`]. -/// -/// IMPORTANT: -/// - `aggregate` should not be called more than once at a time in parallel per process. -/// - Raw signatures are assumed valid (otherwise `aggregate` will panic). -/// -/// XMSS Performance: it is optimized for a small set of different (epoch, message), and many XMSS -/// sharing each such pair. -pub fn aggregate( - children: &[EthereumProof], - raw_xmss: Vec<(XmssPublicKey, xmss::Epoch, xmss::Message, XmssSignature)>, - raw_sphincs: Vec<(SphincsPublicKey, sphincs::Message, SphincsSignature)>, - blobs: &[u64], - declare: Option>, - log_inv_rate: usize, -) -> Result { - let da_input = DaInput { - rows: blobs, - roots: declare.map(|claims| claims.da_commitments), - }; - aggregate_with_stats( - children, - raw_xmss, - raw_sphincs, - declare.map(|claims| claims.signatures), - da_input, - log_inv_rate, - ) - .map(|(sig, _)| sig) -} - -/// [`aggregate`], keeping the prover statistics the benchmark reports. -pub(crate) fn aggregate_with_stats( - children: &[EthereumProof], - raw_xmss: Vec<(XmssPublicKey, xmss::Epoch, xmss::Message, XmssSignature)>, - raw_sphincs: Vec<(SphincsPublicKey, sphincs::Message, SphincsSignature)>, - declare: Option<&SignatureClaims>, - da_input: DaInput<'_>, - log_inv_rate: usize, -) -> Result<(EthereumProof, lean_vm::cpu::Stats), AggregationError> { - aggregate_tampered(children, raw_xmss, raw_sphincs, declare, da_input, log_inv_rate, |_| {}) -} - -/// [`aggregate`], with a hook to corrupt the witness before proving. -/// -/// The coverage argument and the claim batching are enforced entirely by guest -/// asserts over prover advice, so the only way to test them is to lie in a hint -/// and require the guest to notice. That is what `tamper` is for -/// (`aggregate_hints_bind`); with an empty hook this is the production path. -pub(crate) fn aggregate_tampered( - children: &[EthereumProof], - raw_xmss: Vec<(XmssPublicKey, xmss::Epoch, xmss::Message, XmssSignature)>, - raw_sphincs: Vec<(SphincsPublicKey, sphincs::Message, SphincsSignature)>, - declare: Option<&SignatureClaims>, - da_input: DaInput<'_>, - log_inv_rate: usize, - tamper: impl FnOnce(&mut Hints), -) -> Result<(EthereumProof, lean_vm::cpu::Stats), AggregationError> { - // Otherwise this reaches `cpu::prove`, which asserts rather than reporting. - if !(lean_vm::pcs::MIN_LOG_INV_RATE..=lean_vm::pcs::MAX_LOG_INV_RATE).contains(&log_inv_rate) { - return Err(AggregationError::InvalidRate { log_inv_rate }); - } - if children.len() > MAX_RECURSIONS { - return Err(AggregationError::TooLarge); - } - let rows = da_input.rows; - if !rows.len().is_multiple_of(BLOB_SYMBOLS) || rows.len() / BLOB_SYMBOLS > DA_MAX_ROWS { - return Err(AggregationError::InvalidBlobSize { symbols: rows.len() }); - } - let mut available_roots = BTreeSet::new(); - for child in children { - check_da_roots(&child.da_roots).map_err(AggregationError::InvalidChild)?; - available_roots.extend(child.da_roots.iter().copied()); - } - let mut da_roots = match da_input.roots { - Some(roots) => roots.iter().copied().collect::>().into_iter().collect(), - None => available_roots.iter().copied().collect::>(), - }; - if da_roots.len() > MAX_DA_ROOTS { - return Err(AggregationError::TooLarge); - } - if rows.is_empty() && da_roots.iter().any(|root| !available_roots.contains(root)) { - return Err(AggregationError::BlobNotCovered); - } - - let guest = unified_guest(); - // Sorted by `(epoch, key)` to group them; `Coverage::raw_walk` then puts the - // groups in the guest's order. Dedup is on the whole triple, so one - // `(epoch, key)` under two messages reaches `plan_coverage`, whose message - // binding rejects it. - let mut raw_xmss = raw_xmss; - raw_xmss.sort_by(|(a, ae, _, _), (b, be, _, _)| (ae, a).cmp(&(be, b))); - raw_xmss.dedup_by(|(a, ae, am, _), (b, be, bm, _)| ae == be && a == b && am == bm); - // On the whole (key, message) pair, so a signer may appear once per message. - let mut raw_sphincs = raw_sphincs; - raw_sphincs.sort_by_key(|(pk, message, _)| (*pk, *message)); - raw_sphincs.dedup_by(|(a, am, _), (b, bm, _)| (a, am) == (b, bm)); - - // Verifying a child here is not a courtesy: `gen_verify` derives the guest's - // whole witness for it from a real verification's summary. Its deferred - // claim is deliberately NOT recomputed, which would cost a full pass over - // each fixed polynomial per child: the batching sumcheck below already - // forces every batched value to be the true evaluation, and the root - // discharges the one claim they reduce to. - let mut verified = Vec::with_capacity(children.len()); - let _span = tracing::info_span!("Verify children").entered(); - for child in children { - check_signer_set(&child.xmss_signers, &child.sphincs_signers).map_err(AggregationError::InvalidChild)?; - let pi = child.public_input(); - let summary = verify(guest, &pi, &child.proof) - .map_err(|e| AggregationError::InvalidChild(AggregateVerifyError::Snark(e)))?; - verified.push((pi, summary)); - } - - drop(_span); - - let _span = tracing::info_span!("Build witness").entered(); - let raw_xmss_claims: Vec<(XmssPublicKey, xmss::Epoch, xmss::Message)> = raw_xmss - .iter() - .map(|(pk, epoch, message, _)| (pk.clone(), *epoch, *message)) - .collect(); - let raw_sphincs_keys: Vec = raw_sphincs.iter().map(|(pk, message, _)| (*pk, *message)).collect(); - let cover = plan_coverage(&raw_xmss_claims, &raw_sphincs_keys, children, declare)?; - let da_contributions = - usize::from(!rows.is_empty()) + children.iter().map(|child| child.da_roots.len()).sum::(); - if cover.n_total() + da_contributions >= MAX_KEYS { - return Err(AggregationError::TooLarge); - } - let n_sphincs = cover.sphincs_signers.len(); - let group_words = |group: &XmssClaimGroup| [vec![F64(group.epoch as u64)], words(&group.message)].concat(); - - let mut hints = Hints::default(); - hints.push( - "meta", - vec![ - count(cover.n_declared), - count(cover.xmss_groups.len() - cover.n_declared), - count(n_sphincs), - count(cover.sphincs_dups.len()), - count(raw_sphincs.len()), - count(children.len()), - count(usize::from(!rows.is_empty())), - ], - ); - let fs_seed = lean_vm::cpu::fs_seed(guest); - hints.push("fs_seed", fs_seed.to_vec()); - // Per group: its epoch, its four message words, and its declared, duplicate - // and raw-signature counts, in the guest's geometry-pass order. The keys - // then ride two per `pubkeys` entry, so the guest can halve its loop - // frames, the odd key out on a final one-key entry; each group's - // duplicates follow its keys. - for (j, group) in cover.xmss_groups.iter().enumerate() { - let mut entry = group_words(group); - entry.extend([ - count(group.keys.len()), - count(cover.xmss_dups[j].len()), - count(raw_xmss.iter().filter(|(_, e, _, _)| *e == group.epoch).count()), - ]); - hints.push("group", entry); - } - for XmssClaimGroup { keys, .. } in cover.declared() { - hints.push("pk_halves", vec![count(keys.len() / 2), count(keys.len() % 2)]); - hints.push("signers_split", signers_split(keys.len().div_ceil(2))); - for pair in keys.chunks(2) { - hints.push("pubkeys", pair.iter().flat_map(key_words).collect()); - } - } - for dups in &cover.xmss_dups { - for pk in dups { - hints.push("dup_pubkeys", key_words(pk)); - } - } - if !cover.sphincs_signers.is_empty() { - hints.push("signers_split", signers_split(cover.sphincs_signers.len())); - } - hints.push("signers_split", signers_split(1 + 2 * cover.n_declared)); - for signer in &cover.sphincs_signers { - hints.push("sphincs_signers", sphincs_signer_words(signer)); - } - for signer in &cover.sphincs_dups { - hints.push("dup_sphincs", sphincs_signer_words(signer)); - } - // Group-major over the table, not over the epochs `raw_xmss` is sorted by: a - // declaration puts the undeclared groups last. Each index is an offset within - // the signature's own group region. - for &i in &cover.raw_walk { - let (pk, epoch, message, sig) = &raw_xmss[i]; - hints.push("raw_index", vec![count(cover.raw_xmss[i])]); - push_signature_hints(&mut hints, pk, sig, message, *epoch)?; - } - // A SPHINCS slot is hinted as an offset into the SPHINCS region, which is - // how one range check keeps the scheme's writers off the other's keys. - for (&offset, (pk, message, sig)) in cover.raw_sphincs.iter().zip(&raw_sphincs) { - hints.push("sp_raw_index", vec![count(offset)]); - push_sphincs_hints(&mut hints, &(*pk, *message), sig)?; - } - - let mut subs = Vec::with_capacity(children.len()); - let mut carried = Vec::with_capacity(children.len()); - for (i, (child, (pi, summary))) in children.iter().zip(verified).enumerate() { - hints.push( - "child_meta", - vec![count(child.xmss_signers.len()), count(child.sphincs_signers.len())], - ); - for (group, (parent_group, offsets)) in child.xmss_signers.iter().zip(&cover.child_xmss[i]) { - let mut entry = group_words(group); - entry.push(count(group.keys.len())); - hints.push("child_group", entry); - hints.push("child_group_map", vec![count(*parent_group)]); - hints.push("child_halves", vec![count(offsets.len() / 2), count(offsets.len() % 2)]); - hints.push("signers_split", signers_split(offsets.len().div_ceil(2))); - for pair in offsets.chunks(2) { - hints.push("child_index", pair.iter().map(|&idx| count(idx)).collect()); - } - } - if !cover.child_sphincs[i].is_empty() { - hints.push("signers_split", signers_split(cover.child_sphincs[i].len())); - } - for &offset in &cover.child_sphincs[i] { - hints.push("child_sphincs_index", vec![count(offset)]); - } - hints.push("signers_split", signers_split(1 + 2 * child.xmss_signers.len())); - hints.push("child_defer", limbs(&child.defer.cells())); - hints.push("child_da_count", vec![count(child.da_roots.len())]); - let (sub_hints, defer) = gen_verify(guest, pi, summary)?; - for (name, entry) in sub_hints { - hints.push(name, entry); - } - subs.push(defer); - carried.push(&child.defer); - } - - drop(_span); - - let defer = if children.is_empty() { - let leaf = DeferredClaim::leaf(); - hints.push( - "leaf_defer", - limbs(&[leaf.bytecode_value, leaf.matrix_a_value, leaf.matrix_b_value]), - ); - leaf - } else { - let _span = tracing::info_span!("Batch deferred claims").entered(); - let (agg_hints, reduced) = aggregate_deferred_claims(&subs, &carried); - for (name, entry) in agg_hints { - hints.push(name, entry); - } - reduced - }; - - let direct_root = if rows.is_empty() { - None - } else { - let _span = tracing::info_span!("LeanDA commit").entered(); - let n_rows = rows.len() / BLOB_SYMBOLS; - let (commitment, witness) = lean_da::commit(rows); - for block in lean_da::membership_vector(&commitment.root) - .as_chunks::() - .0 - { - hints.push("da_weights", limbs(block)); - } - hints.push( - "da_shape", - vec![count(n_rows), count(n_rows.next_power_of_two().ilog2() as usize)], - ); - // Padding rows are constants the guest bakes, so only the real rows' - // symbols ride the stream. - let (c, m) = (CELL_SYMBOLS, CODEWORD_SYMBOLS); - for j in 0..CELLS_PER_ROW { - for i in 0..n_rows { - hints.push( - "da_symbols", - witness.codewords[i * m + j * c..i * m + (j + 1) * c] - .iter() - .map(|&w| F64(w)) - .collect(), - ); - } - } - Some(commitment.root) - }; - if let Some(root) = direct_root { - available_roots.insert(root); - if da_input.roots.is_none() { - da_roots = available_roots.iter().copied().collect(); - } - if da_roots.len() > MAX_DA_ROOTS { - return Err(AggregationError::TooLarge); - } - } - if cover.n_declared == 0 && cover.sphincs_signers.is_empty() && da_roots.is_empty() { - return Err(AggregationError::Empty); - } - let mut claimed = vec![false; da_roots.len()]; - let mut da_dups = Vec::new(); - for root in direct_root - .iter() - .chain(children.iter().flat_map(|child| &child.da_roots)) - { - let slot = take_slot(&da_roots, &mut claimed, &mut da_dups, root); - hints.push("da_index", vec![count(slot)]); - } - if claimed.contains(&false) { - return Err(AggregationError::BlobNotCovered); - } - hints.push("da_meta", vec![count(da_roots.len()), count(da_dups.len())]); - for root in da_roots.iter().chain(&da_dups) { - hints.push("da_roots", da_claim_words(root)); - } - - let public_input = statement_digest( - signers_hash(cover.declared(), &cover.sphincs_signers), - da_list_digest(&da_roots), - &defer, - ); - let mut program = guest.clone(); - // Every aggregate is a potential child, and the guest has no opening arm below - // `2^MU_MIN`. A run smaller than that (a leaf of a few dozen signatures) grows - // its SET table until it clears the floor. - program.min_log_committed = MU_MIN; - tamper(&mut hints); - hints.install(&mut program); - let (proof, stats) = prove(&program, public_input, log_inv_rate); - Ok(( - EthereumProof { - xmss_signers: cover.declared().to_vec(), - sphincs_signers: cover.sphincs_signers, - da_roots, - defer, - proof, - }, - stats, - )) -} -struct CoordinateDescriptor { - kind: usize, - constant: u64, - fresh: usize, - claim_slot: usize, - terms: Range, -} - -struct Term { - kind: usize, - constant: u64, - column_a: usize, - column_b: usize, -} - -struct ClaimDescriptor { - buffer: usize, - column: usize, - qflock_slot: usize, -} - -fn literals(values: impl IntoIterator) -> String { - format!( - "[{}]", - values.into_iter().map(|v| v.to_string()).collect::>().join(", ") - ) -} - -struct OpeningShape { - n_levels: usize, - yr_level: usize, - yr_log_len: usize, - folds: Vec, - log_message_columns: Vec, - queries: Vec, - tree_depths: Vec, - positions_per_squeeze: Vec, - squeezes: Vec, - interleaving: Vec, - query_grinding_bits: Vec, - cap_depths: Vec, - cap_offsets: Vec, - positions_offsets: Vec, - vanish_offsets: Vec, - fold_offsets: Vec, - residual_fold_offsets: Vec, - vanish_values: Vec, - vanish_inverses: Vec, - ood_samples: Vec, -} - -/// The recursion program's placeholder map depends on table structure and bytecode -/// size, so one compiled guest serves every proof layout. -fn placeholder_map(kbc: usize) -> BTreeMap { - // Only block and coordinate structure is used here; dummy instructions and - // table sizes let us derive it before the guest's bytecode exists. - let stand_in = vec![lean_vm::cpu::Op::Xor64 { a: 0, b: 0, c: 0 }; 1 << kbc]; - let layout = lean_vm::cpu::layout(&stand_in, 20, [1usize << 10; lean_vm::tables::N_TABLES], [F64::ZERO; 4]); - let sides: [&[Block]; 3] = [&layout.push, &layout.pull, &layout.count]; - let lcrounds = flock::hash::K_LOG - 6; - - // ---- flattened block/coord descriptors (structural) ---- - let mut sblk = vec![0usize]; - let mut block_coords = Vec::new(); - let mut coordinates = Vec::new(); - let mut terms = Vec::new(); - let (mut nclaims, mut nbcv, mut nblocks) = (0usize, 0usize, 0usize); - // Claim dedup (mirrors leaf.rs): per coord, fresh = first (group, col, - // kappa) occurrence gets the next pool slot; duplicates point at it. - let mut slot_of: std::collections::HashMap<(usize, usize), usize> = Default::default(); - // A TABLE block's coordinates, flattened into terms: the guest rebuilds each as - // `Σ_terms`, so a derived value (an XOR/MUL result, a DEREF store, a JUMP - // successor) costs terms rather than columns. A framework coordinate has none: - // it decomposes into pooled claims instead. - // The table sumcheck settles table claims; only framework blocks stream column values. - let sch_pm = lean_vm::cpu::schema(); - let owner_pm: Vec> = lean_vm::cpu::block_kappa_sources(kbc) - .into_iter() - .map(|(src, _)| src.checked_sub(2)) - .collect(); - for blocks in sides.iter() { - for blk in blocks.iter() { - block_coords.push(coordinates.len()..coordinates.len() + blk.coords.len()); - let owner = owner_pm[nblocks]; - nblocks += 1; - for c in &blk.coords { - // One COORD_FRESH/COORD_CLAIM_SLOT entry PER coord (the guest - // indexes them by global coord offset); only a framework block's - // Col/GCol raises a claim. - let (mut fresh, mut slot) = (0usize, 0usize); - if let (Coord::Col(i) | Coord::GCol(i, _), None) = (c, owner) { - let key = (*i, blk.kappa); - if let Some(&known) = slot_of.get(&key) { - slot = known; - } else { - slot_of.insert(key, nclaims); - fresh = 1; - slot = nclaims; - nclaims += 1; - } - } - let start = terms.len(); - if let Some(t) = owner { - push_coord_terms(c, sch_pm.base[t], &mut terms); - } - nbcv += usize::from(matches!(c, Coord::Public(_))); - coordinates.push(CoordinateDescriptor { - kind: coord_kind(c), - constant: coord_scale(c).0, - fresh, - claim_slot: slot, - terms: start..terms.len(), - }); - } - } - sblk.push(nblocks); - } - let evtot: usize = lean_vm::tables::tables().iter().map(|t| t.n_committed_columns()).sum(); - let ncl = nclaims + evtot + 1; // bus + constraint + the PI memory claim - - // ---- claim descriptor buffer ids (structural) ---- - let valcols = blake2s_value_columns(); - let col_sources_pm = lean_vm::cpu::col_kappa_sources(kbc); - let mut compact_col_pm = vec![usize::MAX; col_sources_pm.len()]; - let mut n_committed = 0usize; - for (global, source) in col_sources_pm.iter().enumerate() { - if source.is_some() { - compact_col_pm[global] = n_committed; - n_committed += 1; - } - } - let qflock_compact = compact_col_pm[lean_vm::cpu::QFLOCK]; - assert_ne!(qflock_compact, usize::MAX, "QFLOCK must be committed"); - // Buffer codes are the guest's POINT_BUF_*: zeta, chi, pi, qflock-chi. - let mut claims = Vec::new(); - walk_claims(&layout, kbc, |site| { - let descriptor = match site { - ClaimSite::Framework { column, .. } => { - let column = compact_col_pm[column]; - assert_ne!(column, usize::MAX, "framework claim must target a committed column"); - ClaimDescriptor { - buffer: 0, - column, - qflock_slot: 0, - } - } - ClaimSite::TableColumn { column, is_virtual, .. } => ClaimDescriptor { - buffer: if is_virtual { 3 } else { 1 }, - column: if is_virtual { - qflock_compact - } else { - compact_col_pm[column] - }, - qflock_slot: if is_virtual { - lean_vm::hash_flock::SLOTS[valcols.iter().position(|&v| v == column).unwrap()] - } else { - 0 - }, - }, - ClaimSite::PublicInput { column } => ClaimDescriptor { - buffer: 2, - column: compact_col_pm[column], - qflock_slot: 0, - }, - }; - claims.push(descriptor); - }); - assert_eq!(claims.len(), ncl, "descriptor count == pool size"); - - // ---- the placeholder map ---- - // A field-element array, flattened to three words an element. - let flds = |v: &[F192]| literals(limbs(v).iter().map(|w| w.0)); - let mut rep = BTreeMap::new(); - let mut ps = |k: &str, v: String| { - rep.insert(format!("{k}_PLACEHOLDER"), v); - }; - ps("STREAM_CAP", STREAM_CAP.to_string()); - ps("MIN_LOG_MEM", lean_vm::cpu::MIN_LOG_MEM.to_string()); - ps("INV_GEN", G.inv().0.to_string()); - ps("MU_CAP", MU_CAP.to_string()); - ps("NO_TABLE", layout.taus.len().to_string()); - ps("GKR_ROUNDS_CAP", (MU_CAP * (MU_CAP + 1) / 2 + MU_CAP + 2).to_string()); - ps("GKR_POINTS_CAP", ((MU_CAP + 1) * MU_CAP).to_string()); - ps("SIDE_BLOCK_START", literals(&sblk)); - ps("N_BLOCKS", nblocks.to_string()); - let bks = lean_vm::cpu::block_kappa_sources(kbc); - // Push and pull emit bus blocks in matched pairs, so their baked kappa-source - // segments are identical; the guest computes only push's side total and - // aliases pull's mu to push's on this basis. - assert_eq!( - bks[sblk[0]..sblk[1]], - bks[sblk[1]..sblk[2]], - "push/pull kappa sources must match" - ); - ps("BLOCK_KAPPA_SRC", literals(bks.iter().map(|&(s, _)| s))); - ps("BLOCK_KAPPA_ADJ", literals(bks.iter().map(|&(_, a)| a))); - ps( - "BLOCK_TABLE", - literals(bks.iter().map(|&(s, _)| if s >= 2 { s - 2 } else { layout.taus.len() })), - ); - let mut block_side = Vec::new(); - for (s, blocks) in sides.iter().enumerate() { - block_side.extend(std::iter::repeat_n(s, blocks.len())); - } - ps("BLOCK_SIDE", literals(&block_side)); - ps("BLOCK_COORD_OFF", literals(block_coords.iter().map(|r| r.start))); - ps("BLOCK_COORD_COUNT", literals(block_coords.iter().map(|r| r.len()))); - ps("COORD_TYPE", literals(coordinates.iter().map(|c| c.kind))); - ps("COORD_CONST", literals(coordinates.iter().map(|c| c.constant))); - ps("COORD_FRESH", literals(coordinates.iter().map(|c| c.fresh))); - ps("COORD_CLAIM_SLOT", literals(coordinates.iter().map(|c| c.claim_slot))); - ps("COORD_TERM_OFF", literals(coordinates.iter().map(|c| c.terms.start))); - ps("COORD_TERM_COUNT", literals(coordinates.iter().map(|c| c.terms.len()))); - ps("TERM_TYPE", literals(terms.iter().map(|t| t.kind))); - ps("TERM_CONST", literals(terms.iter().map(|t| t.constant))); - ps("TERM_COL_A", literals(terms.iter().map(|t| t.column_a))); - ps("TERM_COL_B", literals(terms.iter().map(|t| t.column_b))); - ps("N_BUS_CLAIMS", nclaims.to_string()); - let idxc: Vec = (0..34) - .map(|i| { - let mut g2k = G; - for _ in 0..i { - g2k = g2k * g2k; - } - (F64::ONE + g2k).0 - }) - .collect(); - ps("INDEX_MLE_FACTORS", literals(&idxc)); - ps("N_CLAIMS", ncl.to_string()); - ps("N_TABLES", layout.taus.len().to_string()); - // The table sumcheck's xi layout, from the native verifier's own numbers: - // a disjoint range of identities per table, then THREE powers shared by every - // table, one per bus side. Sharing is what lets the target be derived from the - // three leaf claims instead of trusted (lean_vm::cpu::xi_form_base). - let n_id: Vec = lean_vm::tables::tables().iter().map(|t| t.n_constraints()).collect(); - let form_base = lean_vm::cpu::xi_form_base(); - ps( - "ETA_OFFSET", - literals(lean_vm::constraints::xi_offsets(n_id.iter().copied())), - ); - ps("ETA_FORM_BASE", form_base.to_string()); - ps("N_ETA_POWS", (form_base + 3).to_string()); - let committed: Vec = lean_vm::tables::tables() - .iter() - .map(|t| t.n_committed_columns()) - .collect(); - ps("N_TABLE_COLS", literals(&committed)); - ps("TABLE_COLS_CAP", (committed.iter().max().unwrap() + 1).to_string()); - let fixed_challenges: Vec = flock::zerocheck::univariate_skip_optimized::small_challenges() - .into_iter() - .chain(flock::zerocheck::univariate_skip_optimized::medium_challenges()) - .collect(); - ps("FIXED_CHALLENGES", flds(&fixed_challenges)); - // Flock univariate skip: 6 skipped variables, then the fixed inner rounds. - ps("K_SKIP", "6".to_string()); - ps("N_FIXED_CHALLENGE_ROUNDS", fixed_challenges.len().to_string()); - ps("PHI8_NODES", flds(&primitives::field::PHI_8_TABLE_192[..128])); - // One constant per domain, not one per node: every barycentric denominator over an aligned φ₈ - // window is the same element (`primitives::multilinear::window_denominator`). - ps( - "LAGRANGE_INV_COMBINED", - f192_literal(primitives::multilinear::window_denominator(128)), - ); - ps( - "LAGRANGE_INV_S", - f192_literal(primitives::multilinear::window_denominator(64)), - ); - ps("LINCHECK_ROUNDS", lcrounds.to_string()); - ps("PIN_COLUMN", flock::hash::Z_CONST_POS.to_string()); - ps("K_LOG", flock::hash::K_LOG.to_string()); - // The q_flock Strided-claim slot stride is K_LOG - LOG_PACKING (= 8), so the - // qflock point-claim slot must use THIS, not LOG2_FIELD_BITS. - ps("SLOT_STRIDE_LOG", lean_vm::hash_flock::SLOT_STRIDE_LOG.to_string()); - - // ---- LIG candidate tables (fixed [minm, maxm] range; open_stacked config) ---- - let oshape = |m: usize, log_inv_rate: usize| { - let shape = whir_shape(m, log_inv_rate); - let (vc, sh) = (&shape.config, &shape.levels); - let (cn, cr) = (sh.levels, vc.level_steps); - // Every cap root must match a transcript-bound root, including the final level. - assert_eq!(cr, cn - 1, "the yr level must be the last one"); - let (ck, cl, cyr) = (&sh.ks, &sh.log_msg_cols, sh.yr_log_n); - let cq = &vc.queries; - let (cd, cp) = (&shape.depth, &shape.per_squeeze); - let cs: Vec = (0..cn).map(|i| cq[i].div_ceil(cp[i])).collect(); - let cni: Vec = ck.iter().map(|&k| 1usize << k).collect(); - assert!( - cni.iter().enumerate().all(|(lv, &n)| { - let (bytes, whole_blocks) = if lv == 0 { - (8 * n, n % 8 == 0) - } else { - (24 * n, (3 * n) % 8 == 0) - }; - bytes <= 1024 && whole_blocks - }), - "recursive WHIR guest supports whole-block Merkle rows of at most one 1024-byte BLAKE2s chunk" - ); - let psum = |f: &dyn Fn(usize) -> usize| -> Vec { - let mut offsets = Vec::with_capacity(cn); - let mut acc = 0; - for lv in 0..cn { - offsets.push(acc); - acc += f(lv); - } - offsets - }; - let cap_depths: Vec<_> = (0..cn).map(|lv| merkle_cap_depth(cq[lv], cd[lv])).collect(); - let cap_offsets = psum(&|lv| 1 << cap_depths[lv]); - let c_qpoff = psum(&|lv| cs[lv] * cp[lv]); - let c_svkoff = psum(&|lv| cl[lv] + 1); - let c_foldbase = psum(&|lv| ck[lv]); - let c_risstart: Vec = (0..cn).map(|k| c_foldbase[k] + ck[k]).collect(); - let mut c_svk = Vec::new(); - let mut c_ivk = Vec::new(); - for &cl_lv in cl.iter().take(cn) { - for &v in &pcs::whir::eval_sk_at_vks(cl_lv) { - c_svk.push(v); - c_ivk.push(if v == F64::ZERO { F64::ZERO } else { v.inv() }); - } - } - OpeningShape { - n_levels: cn, - yr_level: cr, - yr_log_len: cyr, - folds: shape.levels.ks, - log_message_columns: shape.levels.log_msg_cols, - queries: shape.config.queries, - tree_depths: shape.depth, - positions_per_squeeze: shape.per_squeeze, - squeezes: cs, - interleaving: cni, - query_grinding_bits: shape.config.grinding_bits, - cap_depths, - cap_offsets, - positions_offsets: c_qpoff, - vanish_offsets: c_svkoff, - fold_offsets: c_foldbase, - residual_fold_offsets: c_risstart, - vanish_values: c_svk, - vanish_inverses: c_ivk, - ood_samples: shape.config.ood_samples, - } - }; - let (minm, maxm) = (MU_MIN, MU_MAX); - let rates = pcs::whir::MIN_LOG_INV_RATE..=pcs::whir::MAX_LOG_INV_RATE; - let cands: Vec<_> = rates - .clone() - .flat_map(|r| (minm..=maxm).map(move |m| oshape(m, r))) - .collect(); - let maxlev = cands.iter().map(|c| c.n_levels).max().unwrap(); - let maxsvk = cands.iter().map(|c| c.vanish_values.len()).max().unwrap(); - let maxood = cands.iter().flat_map(|c| &c.ood_samples).copied().max().unwrap_or(0); - ps("LIG_MAX_LEVELS", maxlev.to_string()); - ps("LIG_MAX_VANISH_LEN", maxsvk.to_string()); - ps("LIG_MAX_OOD_SAMPLES", maxood.to_string()); - ps("LIG_MIN_LOG_SIZE", minm.to_string()); - let cks: Vec<(usize, usize)> = lean_vm::cpu::col_kappa_sources(kbc).into_iter().flatten().collect(); - ps("N_COMMITTED_COLS", cks.len().to_string()); - ps("N_COLUMN_LOGS", (MU_MAX + 1).to_string()); - ps("COL_KAPPA_SRC", literals(cks.iter().map(|&(s, _)| s))); - ps("COL_KAPPA_ADJ", literals(cks.iter().map(|&(_, a)| a))); - ps("PCS_MIN_MU", lean_vm::pcs::MIN_MU.to_string()); - ps( - "LIG_LOG_MSG_COLS_CAP", - cands - .iter() - .map(|c| *c.log_message_columns.iter().max().unwrap()) - .max() - .unwrap() - .to_string(), - ); - ps( - "YR_LOG_CAP", - cands.iter().map(|c| c.yr_log_len).max().unwrap().to_string(), - ); - { - let flat = |f: &dyn Fn(&OpeningShape) -> Vec| { - let rows: Vec = cands - .iter() - .flat_map(|c| { - let mut row = f(c); - row.resize(maxlev, 0); - row - }) - .collect(); - literals(rows) - }; - let scal = |f: &dyn Fn(&OpeningShape) -> usize| literals(cands.iter().map(f)); - ps("LIG_N_LEVELS", scal(&|c| c.n_levels)); - ps("LIG_YR_LEVEL", scal(&|c| c.yr_level)); - // The guest rotates the terminal point by the lane-fold count to index it by - // witness coordinate, and the residual segment is what it rotates the last - // lane challenges past, so the residual may never be longer than that fold - // (`RESIDUAL_MAX_LOG` < `INITIAL_FOLDING_FACTOR` keeps this true by a margin). - assert!( - cands.iter().all(|c| c.yr_log_len <= c.folds[0]), - "residual longer than the lane fold: the guest's point rotation has no room" - ); - ps("LIG_YR_LOG_LEN", scal(&|c| c.yr_log_len)); - ps("LIG_YR_LEN", scal(&|c| 1usize << c.yr_log_len)); - ps("LIG_TOTAL_FOLDS", scal(&|c| c.folds.iter().sum())); - ps("LIG_MAX_QUERIES", scal(&|c| *c.queries.iter().max().unwrap())); - ps("LIG_MAX_SQUEEZES", scal(&|c| *c.squeezes.iter().max().unwrap())); - ps("LIG_MAX_INTERLEAVE", scal(&|c| *c.interleaving.iter().max().unwrap())); - ps( - "LIG_POSITIONS_LEN", - scal(&|c| { - (0..c.n_levels) - .map(|level| c.squeezes[level] * c.positions_per_squeeze[level]) - .sum() - }), - ); - let row_cap = cands - .iter() - .flat_map(|c| { - c.interleaving - .iter() - .enumerate() - .map(|(level, &n)| n * if level == 0 { 1 } else { 3 }) - }) - .max() - .unwrap(); - ps("LIG_ROW_CAP", row_cap.to_string()); - ps( - "LIG_PATH_CAP", - cands - .iter() - .flat_map(|c| { - c.tree_depths - .iter() - .zip(&c.cap_depths) - .map(|(depth, cap)| 8 * (depth - cap)) - }) - .max() - .unwrap() - .to_string(), - ); - ps("LIG_QUERY_GRIND_BITS", flat(&|c| c.query_grinding_bits.clone())); - ps("LIG_OOD_SAMPLES", flat(&|c| c.ood_samples.clone())); - ps("LIG_QUERIES", flat(&|c| c.queries.clone())); - ps("LIG_FOLDS", flat(&|c| c.folds.clone())); - ps("LIG_INTERLEAVE", flat(&|c| c.interleaving.clone())); - // 64-byte BLAKE2s blocks per leaf row: level 0's committed rows are - // base-field F64 (8 bytes/lane); deeper levels are native F192 - // (24 bytes/word, received as three embedded K limbs each). Rows are - // whole blocks only (asserted at candidate construction). - ps( - "LIG_LEAF_BLOCKS", - flat(&|c| { - c.interleaving - .iter() - .enumerate() - .map(|(level, &n)| if level == 0 { n / 8 } else { 3 * n / 8 }) - .collect() - }), - ); - ps("LIG_TREE_DEPTH", flat(&|c| c.tree_depths.clone())); - ps("LIG_CAP_DEPTH", flat(&|c| c.cap_depths.clone())); - ps("LIG_CAP_OFF", flat(&|c| c.cap_offsets.clone())); - ps("LIG_CAP_LEN", scal(&|c| c.cap_depths.iter().map(|&d| 1 << d).sum())); - ps("LIG_SQUEEZES", flat(&|c| c.squeezes.clone())); - ps("LIG_POSITIONS_OFF", flat(&|c| c.positions_offsets.clone())); - ps("LIG_LOG_MSG_COLS", flat(&|c| c.log_message_columns.clone())); - ps("LIG_RESIDUAL_FOLD_OFF", flat(&|c| c.residual_fold_offsets.clone())); - ps( - "LIG_RESIDUAL_PREFIX_LEN", - flat(&|c| { - c.log_message_columns - .iter() - .map(|&columns| columns - c.yr_log_len) - .collect() - }), - ); - ps("LIG_FOLDS_OFF", flat(&|c| c.fold_offsets.clone())); - ps("LIG_VANISH_OFF", flat(&|c| c.vanish_offsets.clone())); - let mut svk2 = Vec::with_capacity(cands.len() * maxsvk); - let mut ivk2 = Vec::with_capacity(cands.len() * maxsvk); - for candidate in &cands { - let padded_len = svk2.len() + maxsvk; - svk2.extend_from_slice(&candidate.vanish_values); - ivk2.extend_from_slice(&candidate.vanish_inverses); - svk2.resize(padded_len, F64::ZERO); - ivk2.resize(padded_len, F64::ZERO); - } - ps("LIG_VANISH_VALS", literals(svk2.iter().map(|v| v.0))); - ps("LIG_VANISH_INVS", literals(ivk2.iter().map(|v| v.0))); - } - let n_log_sizes = maxm - minm + 1; - let n_rates = MAX_LOG_INV_RATE - MIN_LOG_INV_RATE + 1; - ps("LIG_N_LOG_SIZES", n_log_sizes.to_string()); - ps("LIG_N_RATES", n_rates.to_string()); - ps("LIG_N_CANDIDATES", (n_log_sizes * n_rates).to_string()); - ps("LIG_MIN_SHIFT_INV", g_pow(minm).inv().0.to_string()); - ps("CLAIM_POINT_BUF", literals(claims.iter().map(|c| c.buffer))); - ps("CLAIM_COMMITTED_COL", literals(claims.iter().map(|c| c.column))); - let slot_stride_log = lean_vm::hash_flock::SLOT_STRIDE_LOG; - let cpqbits: Vec = claims - .iter() - .flat_map(|c| (0..slot_stride_log).map(move |k| (c.qflock_slot >> k) & 1)) - .collect(); - ps("CLAIM_QFLOCK_SLOT_BITS", literals(&cpqbits)); - ps("QFLOCK_COMMITTED_COL", qflock_compact.to_string()); - ps("QFLOCK_VARS_CAP", (33 + slot_stride_log).to_string()); - ps("BYTECODE_LOG", kbc.to_string()); - // The stacked bytecode: nbcv/2 encoding columns per side, aligned with the bus - // tuple, so their slots span the fingerprint's own bits. The defer region is - // 2*kbc points + sel bits + 2 reduced + alpha + z_skip + 2*lcrounds rounds - // + 64 z_partial + 1 matpart. - let bc_cols = nbcv / 2; - let log2_bc_cols = lean_vm::leaf::N_TUPLE_BITS; - ps("BYTECODE_COLS", bc_cols.to_string()); - ps("LOG2_BYTECODE_COLS", log2_bc_cols.to_string()); - ps("DEFER_SIZE", (kbc + log2_bc_cols + 2 * lcrounds + 68).to_string()); - ps("BYTECODE_VARS", (kbc + log2_bc_cols).to_string()); - let agg_state = FiatShamirState::from_label(RECURSION_AGG_LABEL).state(); - for (i, word) in agg_state.iter().enumerate() { - ps(&format!("AGG_SEED_{i}"), word.0.to_string()); - } - - // ---- LeanDA (`doc/leanvm` §sec:leanda) ---- - let (pad_cell, pad_row) = lean_da::padding_digests(); - ps("DA_LOG_K", DA_LOG_K.to_string()); - ps("DA_LOG_CELL", DA_LOG_CELL.to_string()); - ps("DA_MAX_ROWS", DA_MAX_ROWS.to_string()); - ps("DA_LOG_MAX_ROWS", DA_MAX_ROWS.ilog2().to_string()); - ps("DA_PAD_CELL", literals(digest_words(&pad_cell).iter().map(|w| w.0))); - ps("DA_PAD_ROW", literals(digest_words(&pad_row).iter().map(|w| w.0))); - let defer_cells = kbc + log2_bc_cols + 1 + 2 * flock::hash::K_LOG + 2; - ps("STMT_HEADER", STATEMENT_HEADER.to_string()); - let stmt_words = STATEMENT_HEADER + 3 * defer_cells; - let blocks = stmt_words.div_ceil(8); - ps("STMT_PAD", (8 * blocks - stmt_words).to_string()); - ps("STMT_BLOCKS", blocks.to_string()); - // A list is at most MAX_KEYS blocks (one a claim is the widest it gets), so it - // holds fewer than that many windows; a declared count is below MAX_KEYS, hence - // decomposes into that many bits. The first two bound a range check, which takes - // a COUNT, the third a bit decomposition. - ps("SIGNERS_WINDOW", SIGNERS_WINDOW.to_string()); - ps("SIGNERS_WINDOW_LOG", SIGNERS_WINDOW.ilog2().to_string()); - ps("SIGNERS_MAX_WINDOWS", SIGNERS_MAX_WINDOWS.to_string()); - ps("SIGNERS_COUNT_BITS", SIGNERS_COUNT_BITS.to_string()); - for (i, word) in lean_vm::hash_flock::IV.iter().enumerate() { - ps(&format!("BLAKE2S_IV_{i}"), word.0.to_string()); - } - let final_md = lean_vm::hash_flock::metadata(0, lean_vm::hash_flock::FINAL_FLAG, 0); - ps("MD_FINAL", final_md[1].0.to_string()); - - // The XMSS instance, from which the guest derives every table width by - // compile-time integer arithmetic. - ps("V", xmss::V.to_string()); - ps("W", xmss::W.to_string()); - ps("TARGET_SUM", xmss::TARGET_SUM.to_string()); - ps("LOG_LIFETIME", xmss::LOG_LIFETIME.to_string()); - // Every XMSS tweak the guest builds is one of these first words plus a - // sub-position, next to the epoch's weighed bits, so the byte layout lives in - // `xmss::make_tweak` and nowhere else. A chain's tweaks start at sub-position - // `CHAIN_LENGTH * i` and a Merkle level's at `level + 1`, exactly as - // `verify_sig` walks them. - ps("XM_ENC_TWEAK", tweak_word(xmss::TWEAK_TYPE_ENCODING, 0).0.to_string()); - ps("XM_PK_TWEAK", tweak_word(xmss::TWEAK_TYPE_WOTS_PK, 0).0.to_string()); - ps("XM_CHAIN_TWEAK", tweak_word(xmss::TWEAK_TYPE_CHAIN, 0).0.to_string()); - ps("XM_MERKLE_TWEAK", tweak_word(xmss::TWEAK_TYPE_MERKLE, 0).0.to_string()); - let p_mul = tweak_word(0, 1) + tweak_word(0, 0); - assert!( - (0..(xmss::V * xmss::CHAIN_LENGTH + xmss::LOG_LIFETIME) as u32).all(|p| { - tweak_word(xmss::TWEAK_TYPE_CHAIN, p) == tweak_word(xmss::TWEAK_TYPE_CHAIN, 0) + F64(u64::from(p)) * p_mul - }), - "a tweak's sub-position is a whole number of units in its first word" - ); - ps("XM_P_MUL", p_mul.0.to_string()); - ps( - "XM_INDEX_WEIGHT", - literals((0..xmss::LOG_LIFETIME).map(|b| tweak_index_weight(b).0)), - ); - ps("MAX_KEYS", MAX_KEYS.to_string()); - ps("MAX_DA_ROOTS", MAX_DA_ROOTS.to_string()); - ps("DA_ROOT_COUNTS", (MAX_DA_ROOTS + 1).to_string()); - ps("MAX_RECURSIONS", MAX_RECURSIONS.to_string()); - ps("MAX_EPOCHS", MAX_EPOCHS.to_string()); - - // The SPHINCS instance. Its tweaks are derived per signature from the index - // the message digest picks, where XMSS's come from one public epoch, so the - // guest receives the shape and the native tweak prefixes. - let dsl_list = |values: &[usize]| { - let inner: Vec = values.iter().map(usize::to_string).collect(); - format!("[{}]", inner.join(", ")) - }; - ps("SP_V", sphincs::V.to_string()); - ps("SP_W", sphincs::W.to_string()); - ps("SP_TARGET_SUM", sphincs::TARGET_SUM.to_string()); - ps("SP_D", sphincs::D.to_string()); - ps("SP_A", sphincs::A.to_string()); - ps("SP_K", sphincs::K.to_string()); - ps("SP_H", sphincs::H.to_string()); - ps("SP_HEIGHTS", dsl_list(&sphincs::HEIGHTS)); - ps("SP_SUFFIX", dsl_list(&sphincs::SUFFIX)); - for (name, tag) in [ - ("SP_TW_CHAIN", sphincs::TWEAK_CHAIN), - ("SP_TW_LEAF", sphincs::TWEAK_LEAF), - ("SP_TW_NODE", sphincs::TWEAK_NODE), - ("SP_TW_ENC", sphincs::TWEAK_ENC), - ("SP_TW_FTS_LEAF", sphincs::TWEAK_FTS_LEAF), - ("SP_TW_FTS_NODE", sphincs::TWEAK_FTS_NODE), - ("SP_TW_FTS_ROOTS", sphincs::TWEAK_FTS_ROOTS), - ("SP_TW_MSG", sphincs::TWEAK_MSG), - ] { - ps(name, words(&sphincs::tweak(tag, 0, 0, 0, 0))[0].0.to_string()); - } - rep -} - -/// Build the guest bytecode and the stacked table its claims are about. Both are -/// cached, so this only moves the cost out of the first prove or verify. -pub fn warm_up() { - unified_guest(); - stacked_bytecode(); -} - -/// The aggregation bytecode, compiled to a fixed point on its own size. -/// -/// The recursion placeholders are a function of the inner bytecode's log size, -/// and here the inner bytecode is this one, so the size has to agree with -/// itself. Its *digest* needs no such loop: it rides the statement rather than -/// the code. The guess converges in one or two rounds because the map's only -/// size-dependent part is a handful of unrolled sumcheck rounds. -pub fn unified_guest() -> &'static Program { - static GUEST: std::sync::OnceLock = std::sync::OnceLock::new(); - GUEST.get_or_init(|| { - let mut guess = 20; - for _ in 0..8 { - let guest = compile_guest(guess); - let actual = guest.prog.len().trailing_zeros() as usize; - if actual == guess { - return guest; - } - guess = actual; - } - panic!("the aggregation bytecode's self-referential compile did not converge"); - }) -} - -fn compile_guest(kbc: usize) -> Program { - let replacements = placeholder_map(kbc); - // `DBG_PLACEHOLDERS=path`: dump the baked guest constants, to read alongside - // a `DBG_PROF_DUMP` profile (the guest's shape is entirely in these). - if let Ok(path) = std::env::var("DBG_PLACEHOLDERS") { - let dump: String = replacements.iter().map(|(k, v)| format!("{k} = {v}\n")).collect(); - std::fs::write(&path, dump).expect("write DBG_PLACEHOLDERS"); - } - let guest = compile( - &parse_with_replacements(include_str!("../guests/lean_ethereum.py"), &replacements) - .expect("the repository aggregation guest must parse"), - ); - // `DBG_DISASM=path`: dump the guest's disassembly, to read alongside a - // `DBG_PROF_DUMP` per-pc profile. Function boundaries lead the dump, so the - // pc a failed guest check reports can be resolved to a source function - // without re-deriving the layout by hand. - if let Ok(path) = std::env::var("DBG_DISASM") { - let mut ranges: Vec<_> = guest.fn_ranges.iter().collect(); - ranges.sort_by_key(|(_, entry, _)| *entry); - let mut dump = String::new(); - for (name, entry, len) in ranges { - dump += &format!("# fn {entry:>7}..{:<7} {name}\n", entry + len); - } - dump += &lean_compiler::disassemble(&guest.prog); - std::fs::write(&path, dump).expect("write DBG_DISASM"); - } - guest -} - -#[cfg(test)] -mod tests { - use super::*; - use rand::SeedableRng; - use rand::rngs::StdRng; - - use crate::signers_cache::{ - KEY_START, XMSS_EPOCH_A, XMSS_EPOCH_B, get_signers, get_signers_at, get_sphincs_signers, message, message_for, - }; - - const SMALL_LEAF_SIZE: usize = 6; - const LOG_INV_RATE: usize = lean_vm::pcs::TEST_LOG_INV_RATE; - - /// Cached `(key, signature)` pairs as the API takes them, every one at - /// `epoch` over the cache's message for it. - fn at_epoch( - signers: &[(XmssPublicKey, XmssSignature)], - epoch: xmss::Epoch, - ) -> Vec<(XmssPublicKey, xmss::Epoch, xmss::Message, XmssSignature)> { - signers - .iter() - .map(|(pk, sig)| (pk.clone(), epoch, message_for(epoch), sig.clone())) - .collect() - } - - fn xmss_claims(sig: &EthereumProof) -> usize { - sig.xmss_signers.iter().map(|group| group.keys.len()).sum() - } - - /// Distinct keys, strictly increasing, without generating any. - fn signer_set(len: usize) -> Vec { - (0..len) - .map(|i| XmssPublicKey { - merkle_root: (i as u128).to_be_bytes(), - public_param: [0; xmss::PUBLIC_PARAM_LEN], - }) - .collect() - } - - #[test] - fn signature_claims_keep_the_wire_layout() { - let claims = SignatureClaims { - xmss: vec![XmssClaimGroup { - epoch: XMSS_EPOCH_A, - message: message(), - keys: signer_set(2), - }], - sphincs: vec![( - SphincsPublicKey::from_bytes(&[0xa5; sphincs::PUB_KEY_SIZE]), - [0x3c; sphincs::MESSAGE_LEN], - )], - }; - let groups: Vec<_> = claims - .xmss - .iter() - .map(|group| (group.epoch, &group.message, &group.keys)) - .collect(); - let bytes = wire().serialize(&(groups, &claims.sphincs)).unwrap(); - assert_eq!(wire().serialize(&claims).unwrap(), bytes); - assert_eq!(wire().deserialize::(&bytes).unwrap(), claims); - } - - /// The guest compiles to one program, always. - /// - /// `unified_guest` finds a fixed point by compiling repeatedly and comparing - /// the result's log size, so a compiler that read a hash seed would not just - /// produce two incompatible transcripts, it could fail to converge at all. - /// The small programs in `lean_compiler`'s `determinism` suite pin the - /// digests; this pins the one program large enough to hit every path that - /// walks a map. A fixed `kbc` is enough: reproducibility does not depend on - /// the size being the fixed point. - #[test] - fn guest_compiles_reproducibly() { - let (one, two) = (compile_guest(20), compile_guest(20)); - assert_eq!( - format!("{:?}", one.prog), - format!("{:?}", two.prog), - "two compilations of the guest produced different bytecode, \ - so the compiler is reading a hash seed" - ); - } - - /// `MAX_KEYS` is exclusive at both host checks: one key short of it passes, - /// the cap itself is the documented error. The cap counts both schemes, so - /// one XMSS key short of it plus one SPHINCS claim is already over. No proof - /// involved, and the epoch cap has its own error alongside. - #[test] - fn max_keys_bound_is_exclusive() { - let full = signer_set(MAX_KEYS); - let group = |keys: &[XmssPublicKey]| { - vec![XmssClaimGroup { - epoch: XMSS_EPOCH_A, - message: message(), - keys: keys.to_vec(), - }] - }; - let claims = |keys: &[XmssPublicKey]| -> Vec<(XmssPublicKey, xmss::Epoch, xmss::Message)> { - keys.iter().map(|pk| (pk.clone(), XMSS_EPOCH_A, message())).collect() - }; - let claim = [( - SphincsPublicKey::from_bytes(&[0; sphincs::PUB_KEY_SIZE]), - [0; sphincs::MESSAGE_LEN], - )]; - check_signer_set(&group(&full[..MAX_KEYS - 1]), &[]).expect("one short of the cap"); - assert_eq!( - check_signer_set(&group(&full), &[]), - Err(AggregateVerifyError::MalformedSignerSet) - ); - assert_eq!( - check_signer_set(&group(&full[..MAX_KEYS - 1]), &claim), - Err(AggregateVerifyError::MalformedSignerSet) - ); - // One group per epoch: MAX_EPOCHS groups pass, one more is malformed. - let spread = |n: usize| -> Vec { - (0..n) - .map(|e| XmssClaimGroup { - epoch: e as u32, - message: message(), - keys: vec![full[e].clone()], - }) - .collect() - }; - check_signer_set(&spread(MAX_EPOCHS), &[]).expect("at the epoch cap"); - assert_eq!( - check_signer_set(&spread(MAX_EPOCHS + 1), &[]), - Err(AggregateVerifyError::MalformedSignerSet) - ); - plan_coverage(&claims(&full[..MAX_KEYS - 1]), &[], &[], None).expect("one short of the cap"); - assert_eq!( - plan_coverage(&claims(&full), &[], &[], None).err(), - Some(AggregationError::TooLarge) - ); - assert_eq!( - plan_coverage(&claims(&full[..MAX_KEYS - 1]), &claim, &[], None).err(), - Some(AggregationError::TooLarge) - ); - let spread_claims = |n: usize| -> Vec<(XmssPublicKey, xmss::Epoch, xmss::Message)> { - (0..n).map(|e| (full[e].clone(), e as u32, message())).collect() - }; - plan_coverage(&spread_claims(MAX_EPOCHS), &[], &[], None).expect("at the epoch cap"); - assert_eq!( - plan_coverage(&spread_claims(MAX_EPOCHS + 1), &[], &[], None).err(), - Some(AggregationError::TooManyEpochs) - ); - } - - fn prove_leaf(signers: &[(XmssPublicKey, XmssSignature)]) -> EthereumProof { - aggregate(&[], at_epoch(signers, XMSS_EPOCH_A), vec![], &[], None, LOG_INV_RATE).expect("leaf aggregates") - } - - #[test] - fn keygen_and_verification_hash_domains_are_disjoint() { - let xmss_tags = [ - xmss::TWEAK_TYPE_PRF, - xmss::TWEAK_TYPE_CHAIN, - xmss::TWEAK_TYPE_WOTS_PK, - xmss::TWEAK_TYPE_MERKLE, - xmss::TWEAK_TYPE_ENCODING, - xmss::TWEAK_TYPE_PARAMETER, - xmss::TWEAK_TYPE_FILLER, - ]; - let sphincs_tags = [ - sphincs::TWEAK_PRF, - sphincs::TWEAK_CHAIN, - sphincs::TWEAK_LEAF, - sphincs::TWEAK_NODE, - sphincs::TWEAK_ENC, - sphincs::TWEAK_FTS_PRF, - sphincs::TWEAK_FTS_LEAF, - sphincs::TWEAK_FTS_NODE, - sphincs::TWEAK_FTS_ROOTS, - sphincs::TWEAK_MSG, - sphincs::TWEAK_PARAMETER, - ]; - let domains: BTreeSet<_> = xmss_tags - .into_iter() - .map(|tag| xmss::make_tweak(tag, 0, 0)) - .chain(sphincs_tags.into_iter().map(|tag| sphincs::tweak(tag, 0, 0, 0, 0))) - .collect(); - assert_eq!(domains.len(), xmss_tags.len() + sphincs_tags.len()); - } - - #[test] - fn signature_tweaks_align_with_distinct_domains() { - for (xmss_tag, sphincs_tag) in [ - (xmss::TWEAK_TYPE_CHAIN, sphincs::TWEAK_CHAIN), - (xmss::TWEAK_TYPE_WOTS_PK, sphincs::TWEAK_LEAF), - (xmss::TWEAK_TYPE_MERKLE, sphincs::TWEAK_NODE), - (xmss::TWEAK_TYPE_ENCODING, sphincs::TWEAK_ENC), - ] { - for position in [0, 1, u32::MAX] { - for index in [0, 1, 3, 0xa0b0_c0d0, u32::MAX] { - let xmss_tweak = xmss::make_tweak(xmss_tag, position, index); - let sphincs_tweak = sphincs::tweak(sphincs_tag, 0, 0, position, index); - assert_eq!(&xmss_tweak[1..], &sphincs_tweak[1..]); - assert_ne!(xmss_tweak[0], sphincs_tweak[0]); - let mut guest_index = F64::ZERO; - for bit in 0..32 { - if index & (1 << bit) != 0 { - guest_index += tweak_index_weight(bit); - } - } - assert_eq!(words(&xmss_tweak), [tweak_word(xmss_tag, position), guest_index]); - } - } - } - } - - type RawSphincs = (SphincsPublicKey, sphincs::Message, SphincsSignature); - - fn prove_sphincs_leaf(signers: &[RawSphincs]) -> EthereumProof { - aggregate(&[], vec![], signers.to_vec(), &[], None, LOG_INV_RATE).expect("leaf aggregates") - } - - #[test] - fn aggregate_one_sphincs_signer() { - lean_vm::init_prover_pool(); - let aggregate = prove_sphincs_leaf(&get_sphincs_signers(1)); - aggregate.verify().expect("verifies"); - assert!(aggregate.xmss_signers.is_empty()); - assert_eq!(aggregate.sphincs_signers.len(), 1); - } - - #[test] - fn aggregate_one_signer() { - lean_vm::init_prover_pool(); - let aggregate = prove_leaf(&get_signers(1)); - aggregate.verify().expect("verifies"); - assert_eq!(aggregate.xmss_signers[0].epoch, XMSS_EPOCH_A); - assert_eq!(aggregate.xmss_signers[0].message, message()); - } - - /// An odd XMSS count, so its digest chain takes its odd-key-out branch and - /// the `pubkeys` stream ends in a short entry; the SPHINCS list has no parity - /// case, absorbing one entry a frame. - #[test] - fn aggregate_mixed_leaf() { - lean_vm::init_prover_pool(); - let aggregate = aggregate( - &[], - at_epoch(&get_signers(3), XMSS_EPOCH_A), - get_sphincs_signers(3), - &[], - None, - LOG_INV_RATE, - ) - .expect("leaf aggregates"); - aggregate.verify().expect("verifies"); - assert_eq!((xmss_claims(&aggregate), aggregate.sphincs_signers.len()), (3, 3)); - } - - /// A node over children of both schemes, overlapping in one signer of each: - /// the coverage table then needs a duplicate slot in both regions, and each - /// child's two key lists have to land in their own. - #[test] - fn aggregate_mixed_two_to_one() { - lean_vm::init_prover_pool(); - let xmss = get_signers(6); - let sphincs = get_sphincs_signers(4); - let leaf = |x: &[(XmssPublicKey, XmssSignature)], s: &[RawSphincs]| { - aggregate(&[], at_epoch(x, XMSS_EPOCH_A), s.to_vec(), &[], None, LOG_INV_RATE).expect("leaf aggregates") - }; - let left = leaf(&xmss[..4], &sphincs[..3]); - let right = leaf(&xmss[3..], &sphincs[2..]); - let node = aggregate(&[left, right], vec![], vec![], &[], None, LOG_INV_RATE).expect("node aggregates"); - node.verify().expect("node verifies"); - assert_eq!((xmss_claims(&node), node.sphincs_signers.len()), (6, 4)); - assert!(node.xmss_signers[0].keys.windows(2).all(|w| w[0] < w[1])); - assert!(node.sphincs_signers.windows(2).all(|w| w[0] < w[1])); - } - - /// A node whose children are each of one scheme only: every key list it - /// rebuilds is empty on one side, which is the only way the guest's - /// key-absorbing loops run over an empty range and its bound `log(x) < - /// log(g^0)` (unsatisfiable, so nothing may be written there) is reached. - #[test] - fn aggregate_one_scheme_per_child() { - lean_vm::init_prover_pool(); - let xmss_child = prove_leaf(&get_signers(3)); - let sphincs_child = prove_sphincs_leaf(&get_sphincs_signers(2)); - let node = - aggregate(&[xmss_child, sphincs_child], vec![], vec![], &[], None, LOG_INV_RATE).expect("node aggregates"); - node.verify().expect("node verifies"); - assert_eq!((xmss_claims(&node), node.sphincs_signers.len()), (3, 2)); - } - - /// The repeat the statement allows: one key signing two messages is two - /// claims, ordered by the pair, each needing its own signature. Generated - /// here rather than cached, the cache holding one message per key. - #[test] - fn aggregate_one_key_two_messages() { - lean_vm::init_prover_pool(); - let mut rng = StdRng::seed_from_u64(77); - let (secret_key, public_key) = sphincs::key_gen(&mut rng); - let raw: Vec = [3u8, 9] - .into_iter() - .map(|tag| { - let signed: sphincs::Message = std::array::from_fn(|i| tag.wrapping_mul(i as u8 + 1)); - let signature = sphincs::sign(&secret_key, &signed).expect("signs"); - (public_key, signed, signature) - }) - .collect(); - let aggregate = prove_sphincs_leaf(&raw); - aggregate.verify().expect("verifies"); - assert_eq!(aggregate.sphincs_signers.len(), 2); - let (first, second) = (aggregate.sphincs_signers[0], aggregate.sphincs_signers[1]); - assert_eq!(first.0, second.0, "the same key, twice"); - assert!(first.1 < second.1, "ordered by the message"); - } - - #[test] - fn guest_column_selectors_match_native_eq() { - lean_vm::init_prover_pool(); - let (helpers, _) = include_str!("../guests/lean_ethereum.py") - .split_once("\ndef main():") - .unwrap(); - let source = format!( - r#"{helpers} -def main(): - point = HeapBuf(3 * MAX_STACK_LOG) - hint_witness(point[0:3 * MAX_STACK_LOG], "point") - offset = hint_witness("offset") - kappa_g = hint_witness("kappa") - weight = match(log(kappa_g), range(0, N_COLUMN_LOGS), lambda kappa: column_selector(offset, point, kappa)) - public = GEN ** 0 - public[0:3] = weight - return -"# - ); - let guest = compile(&parse_with_replacements(&source, &placeholder_map(18)).unwrap()); - let run = |point: &[F192], offset: usize, kappa: usize, expected: F192| { - let mut hints = Hints::default(); - hints.push("point", limbs(point)); - hints.push("offset", vec![count(offset)]); - hints.push("kappa", vec![count(kappa)]); - let mut program = guest.clone(); - hints.install(&mut program); - program.execute([F64(expected.c0), F64(expected.c1), F64(expected.c2), F64::ZERO]) - }; - let mut rng = StdRng::seed_from_u64(813); - for mu in MU_MIN..=MU_MAX { - let mut point: Vec = (0..mu) - .map(|_| { - let (c0, c1, c2) = rand::Rng::random(&mut rng); - F192::new(c0, c1, c2) - }) - .collect(); - point.resize(MU_MAX, F192::ZERO); - for kappa in 0..=mu { - let mask = ((1usize << mu) - 1) & !((1usize << kappa) - 1); - for offset in [0, mask, 0x0555_5555 & mask] { - let selector: Vec<_> = (kappa..mu) - .map(|k| F192::from(F64(((offset >> k) & 1) as u64))) - .collect(); - let expected = primitives::multilinear::eq_eval(&selector, &point[kappa..mu]); - assert!(run(&point, offset, kappa, expected).unconstrained_reads.is_empty()); - if kappa != 0 { - assert!(std::panic::catch_unwind(|| run(&point, offset + 1, kappa, expected)).is_err()); - } - } - } - assert!(std::panic::catch_unwind(|| run(&point, 1usize << MU_MAX, 0, F192::ZERO)).is_err()); - } - } - - #[test] - fn guest_merkle_children_bind_every_link() { - lean_vm::init_prover_pool(); - let (helpers, _) = include_str!("../guests/lean_ethereum.py") - .split_once("\ndef main():") - .unwrap(); - let source = format!( - r#"{helpers} -def main(): - bits = StackBuf(3) - hint_witness(bits[0:3], "bits") - for k in unroll(0, 3): - bits[k] = bits[k] * bits[k] - direction = addr(bits) - leaf = StackBuf(4) - hint_witness(leaf[0:4], "leaf") - root = verify_merkle_path(leaf, direction, 3) - public = GEN ** 0 - public[0:4] = root - return -"# - ); - let guest = compile(&parse_with_replacements(&source, &placeholder_map(18)).unwrap()); - let mut tree = vec![[0u8; 32]; 16]; - for (i, leaf) in tree[8..].iter_mut().enumerate() { - *leaf = pcs::merkle::hash_leaf(&[i as u8; 64]); - } - for i in (1..8).rev() { - tree[i] = pcs::merkle::hash_pair(&tree[2 * i], &tree[2 * i + 1]); - } - let run = |index: usize, leaf: [u8; 32], pairs: &[[[u8; 32]; 2]], root: [u8; 32]| { - let mut hints = Hints::default(); - hints.push("bits", (0..3).map(|k| F64(((index >> k) & 1) as u64)).collect()); - hints.push("leaf", digest_words(&leaf).to_vec()); - hints.push( - "merkle_children", - pairs.iter().flatten().flat_map(digest_words).collect(), - ); - let mut program = guest.clone(); - hints.install(&mut program); - program.execute(digest_words(&root)) - }; - for index in 0..8 { - let pairs: Vec<_> = (0..3) - .map(|level| { - let left = ((8 + index) >> level) & !1; - [tree[left], tree[left + 1]] - }) - .collect(); - assert!( - run(index, tree[8 + index], &pairs, tree[1]) - .unconstrained_reads - .is_empty() - ); - for level in 0..3 { - for side in 0..2 { - for byte in [0, 8, 16, 24] { - let mut forged = pairs.clone(); - forged[level][side][byte] ^= 1; - assert!(std::panic::catch_unwind(|| run(index, tree[8 + index], &forged, tree[1])).is_err()); - } - } - // Rehash a forged running child all the way to a matching public root. - // Root equality alone passes; the selected child must still bind to its predecessor. - for byte in [0, 8, 16, 24] { - let mut forged = pairs.clone(); - forged[level][(index >> level) & 1][byte] ^= 1; - let mut root = pcs::merkle::hash_pair(&forged[level][0], &forged[level][1]); - for (height, pair) in forged.iter_mut().enumerate().skip(level + 1) { - pair[(index >> height) & 1] = root; - root = pcs::merkle::hash_pair(&pair[0], &pair[1]); - } - assert!(std::panic::catch_unwind(|| run(index, tree[8 + index], &forged, root)).is_err()); - } - } - assert!(std::panic::catch_unwind(|| run(index ^ 1, tree[8 + index], &pairs, tree[1])).is_err()); - } - } - - /// The right leaf holds more keys than one absorb window of its list hash - /// (`SIGNERS_WINDOW` blocks of two keys), so both that leaf and the parent - /// rebuilding it run the window loop and then a non-empty tail, while the left - /// leaf's list is a tail alone. Under a window everywhere, the loop would never - /// execute and neither would the byte counter's base. - #[test] - fn aggregate_two_to_one() { - lean_vm::init_prover_pool(); - let big = 2 * SIGNERS_WINDOW + 6; - let signers = get_signers(SMALL_LEAF_SIZE + big); - let left = prove_leaf(&signers[..SMALL_LEAF_SIZE]); - let right = prove_leaf(&signers[SMALL_LEAF_SIZE..]); - let node = aggregate(&[left, right], vec![], vec![], &[], None, LOG_INV_RATE).expect("node aggregates"); - node.verify().expect("node verifies"); - assert_eq!(xmss_claims(&node), SMALL_LEAF_SIZE + big); - } - - fn da_rows(n_rows: usize, seed: u64) -> Vec { - let mut rng = ::seed_from_u64(seed); - (0..n_rows * (1 << DA_LOG_K)) - .map(|_| rand::Rng::random(&mut rng)) - .collect() - } - - /// A LeanDA payload, proven without signatures and published in the - /// statement. The node has to reach the native committer's root, and the - /// aggregate has to verify against a statement that carries it. - #[test] - fn aggregate_with_a_da_payload() { - lean_vm::init_prover_pool(); - let rows = da_rows(3, 97); - - let node = aggregate(&[], vec![], vec![], &rows, None, LOG_INV_RATE).expect("node aggregates"); - node.verify().expect("node verifies"); - assert_eq!(node.num_signature_claims(), 0); - - let (commitment, _) = lean_da::commit(&rows); - assert_eq!( - node.da_roots, - vec![commitment.root], - "the guest committed to something else" - ); - let received = EthereumProof::from_bytes(&node.to_bytes()).unwrap(); - assert_eq!(received.da_commitments(), &[commitment.root]); - received.verify().unwrap(); - let keys = SignatureClaims { - xmss: node.xmss_signers.clone(), - sphincs: node.sphincs_signers.clone(), - }; - let mut core: WireCore = wire().deserialize(&node.to_bytes_without_pubkeys()).unwrap(); - core.0[0][0] ^= 1; - let bad = EthereumProof::from_bytes_without_pubkeys(&wire().serialize(&core).unwrap(), keys).unwrap(); - assert!(bad.verify().is_err(), "the DA root must bind the VM proof"); - } - - #[test] - fn invalid_da_selections_are_rejected_before_building_the_proof() { - let unknown = [[0xa5; 32]]; - assert_eq!( - aggregate( - &[], - vec![], - vec![], - &[], - Some(ClaimSelection { - signatures: &SignatureClaims::default(), - da_commitments: &unknown - }), - LOG_INV_RATE - ) - .unwrap_err(), - AggregationError::BlobNotCovered - ); - let too_many: Vec<_> = (0..=MAX_DA_ROOTS).map(|i| [i as u8; 32]).collect(); - assert_eq!( - aggregate( - &[], - vec![], - vec![], - &[], - Some(ClaimSelection { - signatures: &SignatureClaims::default(), - da_commitments: &too_many - }), - LOG_INV_RATE - ) - .unwrap_err(), - AggregationError::TooLarge - ); - } - - #[test] - fn da_roots_accumulate_and_can_be_selected_or_omitted() { - lean_vm::init_prover_pool(); - let signers = get_signers(SMALL_LEAF_SIZE); - let mut children = Vec::new(); - for seed in [509, 510] { - let rows = da_rows(1, seed); - children.push(aggregate(&[], at_epoch(&signers, XMSS_EPOCH_A), vec![], &rows, None, LOG_INV_RATE).unwrap()); - } - let signatures = SignatureClaims { - xmss: children[0].xmss_signers.clone(), - sphincs: children[0].sphincs_signers.clone(), - }; - let first = children[0].da_roots[0]; - let second = children[1].da_roots[0]; - assert_ne!(first, second); - let mut both = vec![first, second]; - both.sort(); - for selected in [ - None, - Some(vec![]), - Some(vec![first]), - Some(vec![second]), - Some(vec![second, first, second]), - ] { - let node = aggregate( - &children, - vec![], - vec![], - &[], - selected.as_deref().map(|roots| ClaimSelection { - signatures: &signatures, - da_commitments: roots, - }), - LOG_INV_RATE, - ) - .unwrap(); - node.verify().unwrap(); - let mut expected = selected.unwrap_or_else(|| both.clone()); - expected.sort(); - expected.dedup(); - assert_eq!(node.da_roots, expected); - assert_eq!(node.da_commitments_digest(), da_list_digest(&expected)); - let mut tampered = node.clone(); - tampered.da_roots = vec![[0xa5; 32]]; - assert!( - tampered.verify().is_err(), - "changing the root list requires a new proof" - ); - } - let repeated = aggregate( - &[children[0].clone(), children[0].clone()], - vec![], - vec![], - &[], - None, - LOG_INV_RATE, - ) - .unwrap(); - repeated.verify().unwrap(); - assert_eq!(repeated.da_roots, vec![first]); - let same_rows = da_rows(1, 509); - let repeated_direct = aggregate( - std::slice::from_ref(&repeated), - vec![], - vec![], - &same_rows, - None, - LOG_INV_RATE, - ) - .unwrap(); - repeated_direct.verify().unwrap(); - assert_eq!(repeated_direct.da_roots, vec![first]); - - let rows = da_rows(1, 511); - let new_root = lean_da::commit(&rows).0.root; - for selected in [vec![new_root], vec![first], vec![]] { - let node = aggregate( - &children, - vec![], - vec![], - &rows, - Some(ClaimSelection { - signatures: &signatures, - da_commitments: &selected, - }), - LOG_INV_RATE, - ) - .unwrap(); - node.verify().unwrap(); - assert_eq!(node.da_roots, selected); - } - assert!(matches!( - aggregate( - &children, - vec![], - vec![], - &rows, - Some(ClaimSelection { - signatures: &signatures, - da_commitments: &[[0xa5; 32]] - }), - LOG_INV_RATE - ), - Err(AggregationError::BlobNotCovered) - )); - let direct = aggregate(&children, vec![], vec![], &rows, None, LOG_INV_RATE).unwrap(); - direct.verify().unwrap(); - let mut three = both.clone(); - three.push(new_root); - three.sort(); - assert_eq!(direct.da_roots, three); - let received = EthereumProof::from_bytes(&direct.to_bytes()).unwrap(); - received.verify().unwrap(); - let nested = aggregate(&[received, repeated], vec![], vec![], &[], None, LOG_INV_RATE).unwrap(); - nested.verify().unwrap(); - assert_eq!(nested.da_roots, three); - let narrowed = aggregate( - &[nested], - vec![], - vec![], - &[], - Some(ClaimSelection { - signatures: &signatures, - da_commitments: &[second], - }), - LOG_INV_RATE, - ) - .unwrap(); - narrowed.verify().unwrap(); - assert_eq!(narrowed.da_roots, vec![second]); - let dropped = aggregate( - &[narrowed], - vec![], - vec![], - &[], - Some(ClaimSelection { - signatures: &signatures, - da_commitments: &[], - }), - LOG_INV_RATE, - ) - .unwrap(); - dropped.verify().unwrap(); - assert!(dropped.da_roots.is_empty()); - assert_eq!(dropped.da_commitments_digest(), primitives::hash::hash(&[])); - assert!(matches!( - aggregate( - &[dropped], - vec![], - vec![], - &[], - Some(ClaimSelection { - signatures: &signatures, - da_commitments: &[second] - }), - LOG_INV_RATE - ), - Err(AggregationError::BlobNotCovered) - )); - assert!(matches!( - aggregate( - &children, - vec![], - vec![], - &[], - Some(ClaimSelection { - signatures: &signatures, - da_commitments: &[[0xa5; 32]] - }), - LOG_INV_RATE - ), - Err(AggregationError::BlobNotCovered) - )); - - // The first child's omitted root occupies slot 1; it cannot cover slot 0 - // (the second child's declared root) or write outside the DA region. - for index in [count(0), count(2), count(MAX_KEYS - 1), F64::ZERO] { - let outcome = std::panic::catch_unwind(|| { - aggregate_tampered( - &children, - vec![], - vec![], - None, - DaInput { - rows: &[], - roots: Some(&[second]), - }, - LOG_INV_RATE, - |h| { - h.entries("da_index")[0] = vec![index]; - }, - ) - }); - assert!(!matches!(outcome, Ok(Ok(_))), "accepted false DA slot {index:?}"); - } - let outcome = std::panic::catch_unwind(|| { - aggregate_tampered( - &[children[0].clone(), children[0].clone()], - vec![], - vec![], - None, - DaInput::default(), - LOG_INV_RATE, - |h| { - h.entries("da_index")[1] = h.entries("da_index")[0].clone(); - }, - ) - }); - assert!( - !matches!(outcome, Ok(Ok(_))), - "identical roots still require distinct coverage writes" - ); - // An extra, unclaimed slot leaves the published digest and both child - // statements intact; only the final coverage count rejects it. - for (declared, duplicates) in [(1, 2), (MAX_DA_ROOTS + 1, 1), (1, MAX_RECURSIONS * MAX_DA_ROOTS + 2)] { - let outcome = std::panic::catch_unwind(|| { - aggregate_tampered( - &children, - vec![], - vec![], - None, - DaInput { - rows: &[], - roots: Some(&[second]), - }, - LOG_INV_RATE, - |h| { - h.entries("da_meta")[0] = vec![count(declared), count(duplicates)]; - h.entries("da_roots").push(da_claim_words(&second)); - }, - ) - }); - assert!(!matches!(outcome, Ok(Ok(_))), "accepted invalid DA coverage shape"); - } - for selected in [None, Some([].as_slice()), Some([second].as_slice())] { - let outcome = std::panic::catch_unwind(|| { - aggregate_tampered( - &children, - vec![], - vec![], - None, - DaInput { - rows: &[], - roots: selected, - }, - LOG_INV_RATE, - |h| { - let root = h - .entries("da_roots") - .iter_mut() - .find(|r| **r == da_claim_words(&second)) - .unwrap(); - *root = da_claim_words(&first); - }, - ) - }); - assert!( - !matches!(outcome, Ok(Ok(_))), - "every child's complete DA list must be authenticated" - ); - } - for word in 4..8 { - let outcome = std::panic::catch_unwind(|| { - aggregate_tampered( - &children, - vec![], - vec![], - None, - DaInput { - rows: &[], - roots: Some(&[second]), - }, - LOG_INV_RATE, - |h| { - let omitted = h - .entries("da_roots") - .iter_mut() - .find(|claim| claim[..4] == digest_words(&first)) - .unwrap(); - omitted[word] += F64::ONE; - }, - ) - }); - assert!( - !matches!(outcome, Ok(Ok(_))), - "a child's vector hash cannot change, even for an omitted root" - ); - } - for roots in [ - vec![second, second], - vec![both[1], both[0]], - vec![[0; 32]; MAX_DA_ROOTS], - ] { - let mut bad = direct.clone(); - bad.da_roots = roots; - assert_eq!(bad.verify(), Err(AggregateVerifyError::MalformedDaCommitments)); - assert_eq!( - EthereumProof::from_bytes(&bad.to_bytes()).unwrap_err(), - AggregateVerifyError::MalformedDaCommitments - ); - } - } - - #[test] - fn da_root_lists_merge_two_and_three() { - lean_vm::init_prover_pool(); - let mut leaves = Vec::new(); - let mut expected = Vec::new(); - for seed in 600..605 { - let rows = da_rows(1, seed); - let leaf = aggregate(&[], vec![], vec![], &rows, None, LOG_INV_RATE).unwrap(); - expected.extend_from_slice(leaf.da_commitments()); - leaves.push(leaf); - } - let left = aggregate(&leaves[..2], vec![], vec![], &[], None, LOG_INV_RATE).unwrap(); - let right = aggregate(&leaves[2..], vec![], vec![], &[], None, LOG_INV_RATE).unwrap(); - assert_eq!(left.da_commitments().len(), 2); - assert_eq!(right.da_commitments().len(), 3); - let root = aggregate(&[left, right], vec![], vec![], &[], None, LOG_INV_RATE).unwrap(); - let root = EthereumProof::from_bytes(&root.to_bytes()).unwrap(); - root.verify().unwrap(); - expected.sort(); - assert_eq!(root.num_signature_claims(), 0); - assert_eq!(root.da_commitments().len(), 5); - assert_eq!(root.da_commitments(), expected); - assert_eq!(root.da_commitments_digest(), da_list_digest(&expected)); - - let boundary: Vec<_> = (0..=MAX_DA_ROOTS).map(|i| [i as u8; 32]).collect(); - check_da_roots(&boundary[..MAX_DA_ROOTS]).unwrap(); - let mut full = root.clone(); - full.da_roots = boundary[..MAX_DA_ROOTS].to_vec(); - let mut extra = root.clone(); - extra.da_roots = boundary[MAX_DA_ROOTS..].to_vec(); - assert_eq!( - aggregate(&[full, extra], vec![], vec![], &[], None, LOG_INV_RATE).unwrap_err(), - AggregationError::TooLarge, - "reject the oversized union before verifying the modified child statements" - ); - let mut too_many = root; - too_many.da_roots = boundary; - assert_eq!(too_many.verify(), Err(AggregateVerifyError::MalformedDaCommitments)); - assert_eq!( - EthereumProof::from_bytes(&too_many.to_bytes()).unwrap_err(), - AggregateVerifyError::MalformedDaCommitments - ); - } - - #[test] - fn da_guest_hashes_root_lists() { - lean_vm::init_prover_pool(); - let (helpers, _) = include_str!("../guests/lean_ethereum.py") - .split_once("\ndef main():") - .unwrap(); - let source = format!( - r#"{helpers} -def main(): - n_g = hint_witness("n") - assert log(n_g) < MAX_DA_ROOTS + 1 - roots = HeapBuf((n_g * GEN) ** 8) - for x in mul_range(1, n_g): - root = roots * (x ** 8) - hint_witness(root[0:8], "root") - digest = da_list_digest(roots, n_g) - public = GEN ** 0 - public[0:4] = digest - return -"# - ); - let guest = compile(&parse_with_replacements(&source, &placeholder_map(20)).unwrap()); - for n in 0..=MAX_DA_ROOTS + 1 { - let roots: Vec<_> = (0..n) - .map(|i| primitives::hash::hash(&(i as u64).to_le_bytes())) - .collect(); - let mut hints = Hints::default(); - hints.push("n", vec![count(n)]); - for root in &roots { - hints.push("root", da_claim_words(root)); - } - let mut program = guest.clone(); - hints.install(&mut program); - let public = digest_words(&da_list_digest(&roots)); - if n <= MAX_DA_ROOTS { - let execution = program.execute(public); - assert!(execution.unconstrained_reads.is_empty(), "{n} roots"); - } else { - assert!(std::panic::catch_unwind(|| program.execute(public)).is_err()); - } - } - } - - #[test] - fn da_guest_bounds_coverage_slots() { - lean_vm::init_prover_pool(); - let (helpers, _) = include_str!("../guests/lean_ethereum.py") - .split_once("\ndef main():") - .unwrap(); - let source = format!( - r#"{helpers} -def main(): - roots = HeapBuf(24) - hint_witness(roots[0:24], "roots") - n_slots = hint_witness("n_slots") - assert log(n_slots) < 3 - cover = HeapBuf(4) - # Adjacent signature slots must stay outside the DA writer's range. - cover[1] = 1 - claim = cover_da_root(roots, cover * GEN, n_slots, GEN) - digest = StackBuf(4) - blake2s(claim[0:4], claim[4:8], digest) - public = GEN ** 0 - public[0:4] = digest - return -"# - ); - let guest = compile(&parse_with_replacements(&source, &placeholder_map(20)).unwrap()); - let first = [0x13; 32]; - let second = [0x27; 32]; - let outside = [0x39; 32]; - let run = |slots: usize, index: F64, claimed: [u8; 32]| { - let mut hints = Hints::default(); - hints.push( - "roots", - [ - da_claim_words(&first), - da_claim_words(&second), - da_claim_words(&outside), - ] - .concat(), - ); - hints.push("n_slots", vec![count(slots)]); - hints.push("da_index", vec![index]); - let mut program = guest.clone(); - hints.install(&mut program); - program.execute(digest_words(&da_list_digest(&[claimed]))) - }; - assert!(run(2, count(0), first).unconstrained_reads.is_empty()); - assert!(run(2, count(1), second).unconstrained_reads.is_empty()); - // Matching public roots must not bypass the region bound, even when the - // out-of-range root is present in memory. - for (slots, index, claimed) in [ - (0, count(0), first), - (2, count(2), outside), - (2, count(1).inv(), first), - (2, F64::ZERO, first), - ] { - assert!(std::panic::catch_unwind(|| run(slots, index, claimed)).is_err()); - } - } - - #[test] - fn da_guest_checks_commitment_and_codewords() { - lean_vm::init_prover_pool(); - let source = include_str!("../guests/lean_ethereum.py"); - let (helpers, _) = source.split_once("\ndef main():").unwrap(); - let source = format!( - "{helpers}\ndef main():\n _, squares = exponent_tables()\n root, vector = da_verify(squares)\n digest = StackBuf(4)\n blake2s(root, vector, digest)\n public = GEN ** 0\n public[0:4] = digest\n return\n" - ); - let guest = compile(&parse_with_replacements(&source, &placeholder_map(20)).unwrap()); - let n_rows = 3usize; - let codewords = lean_da::encode_rows(&da_rows(3, 101)); - let run = |symbols: &[u64], tamper: &dyn Fn(&mut Hints, &mut [F64; 4])| { - let (commitment, _) = lean_da::commit_codewords(symbols.to_vec()); - let mut public = digest_words(&da_list_digest(&[commitment.root])); - let mut hints = Hints::default(); - hints.push( - "da_shape", - vec![count(n_rows), count(n_rows.next_power_of_two().ilog2() as usize)], - ); - for j in 0..CELLS_PER_ROW { - for i in 0..n_rows { - let start = i * CODEWORD_SYMBOLS + j * CELL_SYMBOLS; - hints.push( - "da_symbols", - symbols[start..start + CELL_SYMBOLS].iter().map(|&w| F64(w)).collect(), - ); - } - } - for block in lean_da::membership_vector(&commitment.root) - .as_chunks::() - .0 - { - hints.push("da_weights", limbs(block)); - } - tamper(&mut hints, &mut public); - let mut program = guest.clone(); - hints.install(&mut program); - program.execute(public) - }; - let honest = run(&codewords, &|_, _| {}); - assert!(honest.unconstrained_reads.is_empty()); - - // Every vector is orthogonal to zero rows: only hashing can reject these changes. - let zeros = vec![0; codewords.len()]; - for limb in 0..3 { - assert!( - std::panic::catch_unwind(|| run(&zeros, &|h, _| { - h.entries("da_weights")[0][limb] += F64::ONE; - })) - .is_err(), - "every limb of L must be bound by its hash" - ); - } - assert!( - std::panic::catch_unwind(|| run(&zeros, &|h, _| { - for block in h.entries("da_weights") { - block.fill(F64::ZERO); - } - })) - .is_err(), - "zero weights must not bypass the external vector hash" - ); - - // Recommit the corrupted matrix and use that root as the public input. - // Hashing and statement binding now pass; only membership can reject it. - for position in [ - 0, - BLOB_SYMBOLS, - CODEWORD_SYMBOLS - 1, - CODEWORD_SYMBOLS, - codewords.len() - 1, - ] { - let mut bad = codewords.clone(); - bad[position] ^= 1; - assert!(std::panic::catch_unwind(|| run(&bad, &|_, _| {})).is_err()); - } - // A zero vector with its own hash passes the guest even for bad data. - // The verifier must derive the expected hash from the root, never trust this hash. - let mut bad = codewords.clone(); - bad[0] ^= 1; - let bad_root = lean_da::commit_codewords(bad.clone()).0.root; - let zero_hash = lean_da::vector_digest(&vec![F192::ZERO; CODEWORD_SYMBOLS]); - let forged_digest = primitives::hash::hash([bad_root, zero_hash].as_flattened()); - assert_ne!(forged_digest, da_list_digest(&[bad_root])); - let unchecked = run(&bad, &|h, public| { - for block in h.entries("da_weights") { - block.fill(F64::ZERO); - } - *public = digest_words(&forged_digest); - }); - assert!(unchecked.unconstrained_reads.is_empty()); - assert!( - std::panic::catch_unwind(|| run(&codewords, &|_, public| { - public[0] += F64::ONE; - })) - .is_err() - ); - // Give the inflated tree its own matching root, so rejection must come - // from the shape check rather than a mismatched public commitment. - let mut padded_words = codewords.clone(); - padded_words.resize(8 * CODEWORD_SYMBOLS, 0); - let (padded_commitment, _) = lean_da::commit_codewords(padded_words); - assert!( - std::panic::catch_unwind(|| run(&codewords, &|h, public| { - h.entries("da_shape")[0][1] = count(3); - *public = digest_words(&da_list_digest(&[padded_commitment.root])); - })) - .is_err(), - "three rows must not use an eight-row tree, even with a matching root" - ); - assert!( - std::panic::catch_unwind(|| run(&zeros, &|h, _| { - h.entries("da_shape")[0][0] = count(0); - })) - .is_err(), - "an empty payload with a matching zero-padded root must be rejected" - ); - for (rows, log_pad) in [(DA_MAX_ROWS + 1, 10), (3, 1), (3, DA_MAX_ROWS.ilog2() as usize + 1)] { - assert!( - std::panic::catch_unwind(|| run(&codewords, &|h, _| { - h.entries("da_shape")[0] = vec![count(rows), count(log_pad)]; - })) - .is_err(), - "accepted shape ({rows}, {log_pad})" - ); - } - } - - #[test] - fn da_row_shape_checks_power_of_two_boundaries() { - lean_vm::init_prover_pool(); - let (helpers, _) = include_str!("../guests/lean_ethereum.py") - .split_once("\ndef main():") - .unwrap(); - let source = format!( - "{helpers}\ndef main():\n _, squares = exponent_tables()\n public = GEN ** 0\n _, _ = da_row_shape(public[1], public[GEN], squares)\n return\n" - ); - let guest = compile(&parse_with_replacements(&source, &placeholder_map(20)).unwrap()); - let max_depth = DA_MAX_ROWS.ilog2() as usize; - let mut counts = vec![0, DA_MAX_ROWS + 1]; - for depth in 0..=max_depth { - let power = 1usize << depth; - counts.extend([power - 1, power, power + 1]); - } - counts.sort_unstable(); - counts.dedup(); - for rows in counts { - for depth in 0..=max_depth + 1 { - let result = - std::panic::catch_unwind(|| guest.execute([count(rows), count(depth), F64::ZERO, F64::ZERO])); - let valid = (1..=DA_MAX_ROWS).contains(&rows) && rows.next_power_of_two() == 1 << depth; - assert_eq!(result.is_ok(), valid, "rows={rows}, depth={depth}"); - if let Ok(execution) = result { - assert!(execution.unconstrained_reads.is_empty()); - } - } - } - for bad in [F64::ZERO, count(1).inv()] { - for public in [ - [bad, count(0), F64::ZERO, F64::ZERO], - [count(1), bad, F64::ZERO, F64::ZERO], - ] { - assert!(std::panic::catch_unwind(|| guest.execute(public)).is_err()); - } - } - } - - #[test] - fn ceil_log_hints_enforce_rounding_and_floor() { - lean_vm::init_prover_pool(); - let (helpers, _) = include_str!("../guests/lean_ethereum.py") - .split_once("\ndef main():") - .unwrap(); - // Replace only advice generation, leaving every guest constraint intact. - assert!(helpers.contains("g_log = hint_log2_ceil(bits_buf, nbits, floor)")); - let helpers = helpers.replace( - "g_log = hint_log2_ceil(bits_buf, nbits, floor)", - "g_log = hint_witness(\"ceil_log\")", - ); - for floor in [0usize, 3] { - let source = format!( - "{helpers}\ndef main():\n powers, squares = exponent_tables()\n bits = HeapBuf(8)\n hint_witness(bits[0:8], \"bits\")\n depth, value = verify_log2_ceil(bits, powers, squares, {floor}, 8)\n public = GEN ** 0\n assert public[1] == value\n assert public[GEN] == depth\n return\n" - ); - let guest = compile(&parse_with_replacements(&source, &placeholder_map(20)).unwrap()); - for value in [0usize, 1, 2, 3, 4, 7, 8, 9, 15, 16, 17, 127, 128, 129, 255] { - let expected = value.max(1).next_power_of_two().ilog2() as usize; - for depth in 0..=9 { - let mut program = guest.clone(); - program.set_witness("bits", vec![(0..8).map(|j| F64(((value >> j) & 1) as u64)).collect()]); - program.set_witness("ceil_log", vec![vec![count(depth)]]); - let result = std::panic::catch_unwind(|| { - program.execute([count(value), count(depth), F64::ZERO, F64::ZERO]) - }); - assert_eq!( - result.is_ok(), - depth == expected.max(floor), - "value={value}, floor={floor}, depth={depth}" - ); - if let Ok(execution) = result { - assert!(execution.unconstrained_reads.is_empty()); - } - } - } - } - } - - #[test] - fn invalid_blob_sizes_are_rejected() { - for symbols in [1, BLOB_SYMBOLS - 1, BLOB_SYMBOLS + 1, (DA_MAX_ROWS + 1) * BLOB_SYMBOLS] { - let rows = vec![0; symbols]; - assert!(matches!( - aggregate(&[], vec![], vec![], &rows, None, LOG_INV_RATE), - Err(AggregationError::InvalidBlobSize { symbols: n }) if n == symbols - )); - } - } - - /// The row count is a run-time parameter, so payloads of different heights - /// must prove against the *same* bytecode and each reach its own committer's - /// root. Powers of two and the counts between them alike: 3 pads to 4 and 5 to - /// 8, exercising two different arms of the tree dispatch and a non-empty gap. - #[test] - fn da_row_count_is_a_run_time_parameter() { - lean_vm::init_prover_pool(); - let signers = get_signers(SMALL_LEAF_SIZE); - for n_rows in [1usize, 3, 4, 5] { - let rows = da_rows(n_rows, 200 + n_rows as u64); - let node = aggregate(&[], at_epoch(&signers, XMSS_EPOCH_A), vec![], &rows, None, LOG_INV_RATE) - .expect("node aggregates"); - node.verify().expect("node verifies"); - let (commitment, _) = lean_da::commit(&rows); - assert_eq!(node.da_roots, vec![commitment.root], "{n_rows} rows"); - } - } - - /// What one blob-carrying proof costs, at the shape EIP-4844 and EIP-7594 fix. - /// Reported, not asserted: run it by name when the shape or the sweep changes. - #[test] - #[ignore] - fn da_blob_proof() { - lean_vm::init_prover_pool(); - let signers = get_signers(SMALL_LEAF_SIZE); - // One discarded proof: the first pays the flock circuit build and the - // arena's page faults, which would otherwise land entirely on the first row - // count reported. - warm_up(); - println!( - "bytecode: {} instructions, DA_MAX_ROWS = {DA_MAX_ROWS}", - unified_guest().prog.len() - ); - let _ = aggregate_with_stats( - &[], - at_epoch(&signers, XMSS_EPOCH_A), - vec![], - None, - DaInput { - rows: &da_rows(1, 7), - roots: None, - }, - LOG_INV_RATE, - ); - for n_rows in [1usize, 6, 14, 32] { - let rows = da_rows(n_rows, 300 + n_rows as u64); - let started = std::time::Instant::now(); - let (node, stats) = aggregate_with_stats( - &[], - at_epoch(&signers, XMSS_EPOCH_A), - vec![], - None, - DaInput { - rows: &rows, - roots: None, - }, - LOG_INV_RATE, - ) - .expect("node aggregates"); - let elapsed = started.elapsed(); - let payload = n_rows * (1 << DA_LOG_K) * 8; - println!( - "{n_rows:>3} blobs ({:>5} KiB): {:>8.2?} {:>6.0} KiB/s cycles 2^{:.1} mem 2^{:.1} proof {:.0} KiB", - payload / 1024, - elapsed, - payload as f64 / 1024.0 / elapsed.as_secs_f64(), - (stats.cycles as f64).log2(), - (stats.mem_used as f64).log2(), - node.to_bytes().len() as f64 / 1024.0, - ); - } - } - - /// A leaf carrying no payload publishes the digest of an empty root list. - #[test] - fn no_payload_publishes_the_empty_root_list() { - lean_vm::init_prover_pool(); - let signers = get_signers(SMALL_LEAF_SIZE); - let node = prove_leaf(&signers); - assert!(node.da_roots.is_empty()); - assert_eq!(node.da_commitments_digest(), primitives::hash::hash(&[])); - } - - /// Two epochs in one tree. Signer `i` holds the same key at both epochs, so - /// group B repeats keys of group A as distinct claims; the left leaf holds - /// one epoch, the right both, and the node maps each child group onto its - /// own region, with a duplicate slot for the key both leaves cover at A. - /// Enough epoch groups that the set's own hash runs its window loop: its string - /// is two blocks a group plus a leading one, so it takes sixteen groups to fill - /// one window of SIGNERS_WINDOW blocks. Every other test stays inside the tail, - /// where `plain_window` never executes and neither does the byte counter's base. - #[test] - fn aggregate_many_epoch_groups() { - lean_vm::init_prover_pool(); - // Two blocks a group plus a leading one, so SIGNERS_WINDOW / 2 groups make - // SIGNERS_WINDOW + 1 blocks: one whole window and the final block. The cached - // keys are activated over exactly that many epochs, and one key may claim - // once per epoch, so a single signer covers them all. - let groups = SIGNERS_WINDOW / 2; - let raw: Vec<_> = (0..groups) - .map(|i| { - let epoch = KEY_START + i as xmss::Epoch; - let (public_key, signature) = get_signers_at(1, epoch).remove(0); - (public_key, epoch, message_for(epoch), signature) - }) - .collect(); - let leaf = aggregate(&[], raw, vec![], &[], None, LOG_INV_RATE).expect("many-group leaf aggregates"); - leaf.verify().expect("it verifies"); - assert_eq!(leaf.xmss_signers.len(), groups); - assert!(leaf.xmss_signers.iter().all(|group| group.keys.len() == 1)); - } - - /// A whole `(epoch, message)` needs the table's undeclared groups; keys of a group - /// that stays need only its duplicate slots. The first narrowing does one, the - /// second both. - #[test] - fn a_node_may_publish_less_than_it_covers() { - lean_vm::init_prover_pool(); - let at_a = get_signers(3); - let at_b = get_signers_at(2, XMSS_EPOCH_B); - let mut raw = at_epoch(&at_a, XMSS_EPOCH_A); - raw.extend(at_epoch(&at_b, XMSS_EPOCH_B)); - let wide = aggregate(&[], raw, vec![], &[], None, LOG_INV_RATE).expect("the wide leaf aggregates"); - wide.verify().expect("the wide leaf verifies"); - assert_eq!(wide.xmss_signers.len(), 2); - assert_eq!(xmss_claims(&wide), 5); - - let narrowed = |wide: &EthereumProof, declare: &SignatureClaims| { - aggregate( - std::slice::from_ref(wide), - vec![], - vec![], - &[], - Some(ClaimSelection { - signatures: declare, - da_commitments: &[], - }), - LOG_INV_RATE, - ) - }; - let (group_a, group_b) = (wide.xmss_signers[0].clone(), wide.xmss_signers[1].clone()); - - // One group declared: the other's epoch and message go with it. - let narrow = narrowed( - &wide, - &SignatureClaims { - xmss: vec![group_b.clone()], - sphincs: vec![], - }, - ) - .expect("narrows to one group"); - narrow.verify().expect("the one-group narrowing verifies"); - assert_eq!(narrow.xmss_signers, vec![group_b.clone()]); - - // One key of one group: B goes whole, A keeps one of three. - let one_of_a = XmssClaimGroup { - epoch: XMSS_EPOCH_A, - message: message(), - keys: vec![group_a.keys[0].clone()], - }; - let part = narrowed( - &wide, - &SignatureClaims { - xmss: vec![one_of_a.clone()], - sphincs: vec![], - }, - ) - .expect("narrows to one key"); - part.verify().expect("the one-key narrowing verifies"); - assert_eq!(part.xmss_signers, vec![one_of_a]); - - assert_eq!( - narrowed(&wide, &SignatureClaims::default()).err(), - Some(AggregationError::Empty), - "a declaration has to publish something" - ); - // A key A holds and B does not, declared at B: the cache reuses keys. - let only_at_a = group_a - .keys - .iter() - .find(|key| !group_b.keys.contains(key)) - .expect("A holds a key B does not") - .clone(); - assert_eq!( - narrowed( - &wide, - &SignatureClaims { - xmss: vec![XmssClaimGroup { - epoch: XMSS_EPOCH_B, - message: message_for(XMSS_EPOCH_B), - keys: vec![only_at_a], - }], - sphincs: vec![], - } - ) - .err(), - Some(AggregationError::NotCovered) - ); - // A covered key, against another message. - assert_eq!( - narrowed( - &wide, - &SignatureClaims { - xmss: vec![XmssClaimGroup { - epoch: XMSS_EPOCH_A, - message: message_for(XMSS_EPOCH_B), - keys: group_a.keys.clone(), - }], - sphincs: vec![], - } - ) - .err(), - Some(AggregationError::NotCovered) - ); - - // Putting the group back is a different signer set, whatever it covered. - let mut rewidened = narrow.clone(); - rewidened.xmss_signers.insert(0, wide.xmss_signers[0].clone()); - assert!( - rewidened.verify().is_err(), - "a split may not be re-widened after the fact" - ); - } - - /// The guest walks raw signatures group by group over the table, and a - /// declaration puts the undeclared groups last, so the table stops agreeing with - /// the epoch order `raw_xmss` is sorted by. Declaring only the HIGHER epoch is - /// what separates the two: the table becomes [B, A] while the raw stream starts - /// with A, and a signature verified against the wrong group's tweaks fails. - #[test] - fn raw_signatures_follow_the_table_not_the_epochs() { - lean_vm::init_prover_pool(); - const _: () = assert!(XMSS_EPOCH_A < XMSS_EPOCH_B, "A must sort first for this to bite"); - let a = get_signers(1); - let b = get_signers_at(1, XMSS_EPOCH_B); - let mut raw = at_epoch(&a, XMSS_EPOCH_A); - raw.extend(at_epoch(&b, XMSS_EPOCH_B)); - let group_b = XmssClaimGroup { - epoch: XMSS_EPOCH_B, - message: message_for(XMSS_EPOCH_B), - keys: vec![b[0].0.clone()], - }; - let sig = aggregate( - &[], - raw, - vec![], - &[], - Some(ClaimSelection { - signatures: &SignatureClaims { - xmss: vec![group_b.clone()], - sphincs: vec![], - }, - da_commitments: &[], - }), - LOG_INV_RATE, - ) - .expect("the narrowing leaf aggregates"); - sig.verify().expect("it verifies"); - assert_eq!(sig.xmss_signers, vec![group_b]); - } - - #[test] - fn aggregate_two_epochs() { - lean_vm::init_prover_pool(); - let at_a = get_signers(4); - let at_b = get_signers_at(2, XMSS_EPOCH_B); - assert_eq!(at_a[0].0, at_b[0].0, "the cache reuses keys across epochs"); - let left = aggregate(&[], at_epoch(&at_a[..3], XMSS_EPOCH_A), vec![], &[], None, LOG_INV_RATE).expect("left"); - let mut right_raw = at_epoch(&at_a[2..], XMSS_EPOCH_A); - right_raw.extend(at_epoch(&at_b, XMSS_EPOCH_B)); - let right = aggregate(&[], right_raw, vec![], &[], None, LOG_INV_RATE).expect("right"); - right.verify().expect("the two-epoch leaf verifies"); - // A claim at epoch A under B's message conflicts with `left`'s group: - // within an aggregate the message is a function of the epoch. - let (pk, _, _, sig) = at_epoch(&at_a[3..], XMSS_EPOCH_A).remove(0); - assert_eq!( - aggregate( - std::slice::from_ref(&left), - vec![(pk, XMSS_EPOCH_A, message_for(XMSS_EPOCH_B), sig)], - vec![], - &[], - None, - LOG_INV_RATE - ) - .err(), - Some(AggregationError::ConflictingMessages) - ); - let node = aggregate(&[left, right], vec![], vec![], &[], None, LOG_INV_RATE).expect("node"); - node.verify().expect("the two-epoch node verifies"); - let messages: Vec = node.xmss_signers.iter().map(|group| group.message).collect(); - assert_eq!(messages, vec![message(), message_for(XMSS_EPOCH_B)]); - let epochs: Vec = node.xmss_signers.iter().map(|group| group.epoch).collect(); - assert_eq!(epochs, vec![XMSS_EPOCH_A, XMSS_EPOCH_B]); - assert_eq!(node.xmss_signers[0].keys.len(), 4); - let mut b_keys: Vec = at_b.iter().map(|(pk, _)| pk.clone()).collect(); - b_keys.sort(); - assert_eq!( - node.xmss_signers[1].keys, b_keys, - "the same keys, at B, are their own claims" - ); - // Statement tampers: no proving, the mutated aggregate just has to fail. - let tampered = |mutate: &dyn Fn(&mut EthereumProof)| { - let mut bad = node.clone(); - mutate(&mut bad); - assert!(bad.verify().is_err(), "a tampered aggregate must not verify"); - }; - tampered(&|s| s.xmss_signers.swap(0, 1)); - tampered(&|s| s.xmss_signers[1].epoch = XMSS_EPOCH_B + 1); - tampered(&|s| { - let moved = s.xmss_signers[1].keys.pop().expect("a key to move"); - s.xmss_signers[0].keys.push(moved); - s.xmss_signers[0].keys.sort(); - s.xmss_signers[0].keys.dedup(); - }); - tampered(&|s| { - // Relabel group B's claims as group A's: B's keys are already among - // A's, so this folds the two groups into one. - let XmssClaimGroup { keys, .. } = s.xmss_signers.remove(1); - s.xmss_signers[0].keys.extend(keys); - s.xmss_signers[0].keys.sort(); - s.xmss_signers[0].keys.dedup(); - }); - } - - #[test] - fn aggregate_overlapping_signers() { - lean_vm::init_prover_pool(); - let signers = get_signers(40); - let left = prove_leaf(&signers[..25]); - let right = prove_leaf(&signers[15..]); - let node = aggregate(&[left, right], vec![], vec![], &[], None, LOG_INV_RATE).expect("node aggregates"); - node.verify().expect("node verifies"); - assert_eq!(xmss_claims(&node), 40); - assert!(node.xmss_signers[0].keys.windows(2).all(|w| w[0] < w[1])); - } - - /// Three levels, both schemes. The SPHINCS claims are rebuilt twice over, once - /// into each node and again into the root, and the two nodes share one claim, - /// so the root needs a SPHINCS duplicate slot for a claim it never saw - /// directly. The root also adds a raw signature of each scheme alongside its - /// children. - #[test] - #[ignore] - fn aggregate_three_levels() { - lean_vm::init_prover_pool(); - let signers = get_signers(4 * SMALL_LEAF_SIZE); - let claims = get_sphincs_signers(5); - let leaf = |index: usize, sphincs: &[RawSphincs]| { - aggregate( - &[], - at_epoch( - &signers[index * SMALL_LEAF_SIZE..(index + 1) * SMALL_LEAF_SIZE], - XMSS_EPOCH_A, - ), - sphincs.to_vec(), - &[], - None, - LOG_INV_RATE, - ) - .expect("leaf aggregates") - }; - let node = |children: &[EthereumProof]| { - aggregate(children, vec![], vec![], &[], None, LOG_INV_RATE).expect("node aggregates") - }; - // Claim 1 is under both nodes; claim 4 arrives raw at the root, and so - // do two XMSS signatures at a second epoch, so the root holds a group - // its children never carried. - let left = node(&[leaf(0, &claims[..2]), leaf(1, &[])]); - let right = node(&[leaf(2, &claims[1..3]), leaf(3, &[])]); - let root = aggregate( - &[left, right], - at_epoch(&get_signers_at(2, XMSS_EPOCH_B), XMSS_EPOCH_B), - claims[4..].to_vec(), - &[], - None, - LOG_INV_RATE, - ) - .expect("root aggregates"); - root.verify().expect("root verifies"); - assert_eq!(xmss_claims(&root), 4 * SMALL_LEAF_SIZE + 2); - assert_eq!(root.xmss_signers.len(), 2, "the raw epoch-B group joins the children's"); - assert_eq!(root.sphincs_signers.len(), 4, "claims 0, 1, 2 and 4, the repeat merged"); - assert!( - root.xmss_signers - .iter() - .all(|group| group.keys.windows(2).all(|w| w[0] < w[1])) - ); - assert!(root.sphincs_signers.windows(2).all(|w| w[0] < w[1])); - } - - #[test] - #[ignore] - fn aggregate_statement_binds() { - lean_vm::init_prover_pool(); - let signers = get_signers(2 * SMALL_LEAF_SIZE); - let left = prove_leaf(&signers[..SMALL_LEAF_SIZE]); - let right = prove_leaf(&signers[SMALL_LEAF_SIZE..]); - // Mixed, so both published lists are non-empty and every tampering - // below has a SPHINCS counterpart. - let node = aggregate(&[left, right], vec![], get_sphincs_signers(3), &[], None, LOG_INV_RATE).expect("node"); - node.verify().expect("the honest node verifies"); - - assert_eq!( - EthereumProof::from_bytes(&node.to_bytes()) - .expect("round trip") - .to_bytes(), - node.to_bytes(), - "the wire format round-trips, recomputed claim values included" - ); - let without = EthereumProof::from_bytes_without_pubkeys( - &node.to_bytes_without_pubkeys(), - SignatureClaims { - xmss: node.xmss_signers.clone(), - sphincs: node.sphincs_signers.clone(), - }, - ) - .expect("round trip"); - without.verify().expect("a caller-supplied signer set verifies"); - - let tampered = |mutate: &dyn Fn(&mut EthereumProof)| { - let mut bad = node.clone(); - mutate(&mut bad); - assert!(bad.verify().is_err(), "a tampered aggregate must not verify"); - }; - tampered(&|s| s.xmss_signers[0].keys[0] = s.xmss_signers[0].keys[1].clone()); - tampered(&|s| { - s.xmss_signers[0].keys.swap(0, 1); - }); - tampered(&|s| { - s.xmss_signers[0].keys.pop(); - }); - tampered(&|s| s.sphincs_signers[0] = s.sphincs_signers[1]); - tampered(&|s| { - s.sphincs_signers.swap(0, 1); - }); - tampered(&|s| { - s.sphincs_signers.pop(); - }); - // Relabelling a signer's scheme: the same 32 bytes moved to the other - // list. Every count and the splits between them are in the statement, and - // the guest holds each region's writers to that region, so this is - // not a free relabelling of what the aggregate claims. - tampered(&|s| { - let moved = s.xmss_signers[0].keys.remove(0); - let claimed = ( - SphincsPublicKey::from_bytes(&moved.flatten()), - s.xmss_signers[0].message, - ); - s.sphincs_signers.push(claimed); - s.sphincs_signers.sort(); - }); - tampered(&|s| s.xmss_signers[0].epoch += 1); - tampered(&|s| s.xmss_signers[0].message[0] ^= 1); - // A signer's own message is in the statement too, so editing it is not a - // free re-attribution of that signature to another message. - tampered(&|s| s.sphincs_signers[0].1[0] ^= 1); - tampered(&|s| s.defer.bytecode_point[0] += F192::ONE); - tampered(&|s| s.defer.matrix_point[0] += F192::ONE); - tampered(&|s| s.xmss_signers[0].keys[0] = get_signers(2 * SMALL_LEAF_SIZE + 1)[2 * SMALL_LEAF_SIZE].0.clone()); - // Splitting one group's keys across two epochs: the same claims cannot - // be re-attributed to an epoch nothing signed at. - tampered(&|s| { - let moved = s.xmss_signers[0].keys.pop().expect("a key to move"); - let epoch = s.xmss_signers[0].epoch; - let message = s.xmss_signers[0].message; - s.xmss_signers.push(XmssClaimGroup { - epoch: epoch + 1, - message, - keys: vec![moved], - }); - }); - - // A claim off the wire carries only its points. Tampering with either - // half must be caught: a point by the recomputation, a value by the - // statement the proof is checked against. - let from_wire = |mutate: &dyn Fn(&mut EthereumProof)| { - let mut bad = EthereumProof::from_bytes(&node.to_bytes()).expect("round trip"); - mutate(&mut bad); - assert!(bad.verify().is_err(), "a tampered wire aggregate must not verify"); - }; - from_wire(&|s| s.defer.bytecode_point[0] += F192::ONE); - from_wire(&|s| s.defer.matrix_point[0] += F192::ONE); - from_wire(&|s| s.defer.bytecode_value += F192::ONE); - from_wire(&|s| s.defer.matrix_a_value += F192::ONE); - from_wire(&|s| s.defer.matrix_b_value += F192::ONE); - } - - /// The all-zeros fast path in `DeferredClaim::recompute` must agree with the - /// two full passes it replaces, or every leaf would verify against the wrong - /// statement. - #[test] - fn leaf_claim_matches_the_general_path() { - let klog = flock::hash::K_LOG; - let leaf = DeferredClaim::leaf(); - let general = { - let bytecode_value = mle_eval_par(stacked_bytecode(), &leaf.bytecode_point); - let eq_r = pcs::whir::build_eq_table_ext(&leaf.matrix_point[..klog]); - let eq_c = pcs::whir::build_eq_table_ext(&leaf.matrix_point[klog..]); - let (matrix_a_value, matrix_b_value) = flock::hash::bilinear_walk_pair(&eq_r, &eq_c); - (bytecode_value, matrix_a_value, matrix_b_value) - }; - assert_eq!((leaf.bytecode_value, leaf.matrix_a_value, leaf.matrix_b_value), general); - } - - type Tamper<'a> = (&'a str, &'a dyn Fn(&mut Hints)); - - /// Corrupt each security-critical hint and require rejection. - #[test] - #[ignore] - fn aggregate_hints_bind() { - lean_vm::init_prover_pool(); - let signers = get_signers(2 * SMALL_LEAF_SIZE); - - let rejects = |children: &[EthereumProof], - raw_signatures: Vec<(XmssPublicKey, xmss::Epoch, xmss::Message, XmssSignature)>, - raw_sphincs: Vec, - description: &str, - tamper: &dyn Fn(&mut Hints)| { - let outcome = std::panic::catch_unwind(std::panic::AssertUnwindSafe(|| { - aggregate_tampered( - children, - raw_signatures, - raw_sphincs, - None, - DaInput::default(), - LOG_INV_RATE, - |hints| tamper(hints), - ) - .map(|(signature, _)| signature.verify().is_ok()) - })); - assert!( - !matches!(outcome, Ok(Ok(true))), - "tampering {description} must be rejected" - ); - }; - - let raw_signatures = at_epoch(&signers[..SMALL_LEAF_SIZE], XMSS_EPOCH_A); - prove_leaf(&signers[..SMALL_LEAF_SIZE]); - let leaf_cases: &[Tamper] = &[ - ("raw_index (duplicate slot)", &|h: &mut Hints| { - let entries = h.entries("raw_index"); - entries[1] = entries[0].clone(); - }), - ("raw_index (out of range)", &|h: &mut Hints| { - h.entries("raw_index")[0] = vec![count(SMALL_LEAF_SIZE)]; - }), - ("group (n_xmss inflated)", &|h: &mut Hints| { - h.entries("group")[0][5] = count(SMALL_LEAF_SIZE + 1); - }), - ("group (n_raw_xmss understated)", &|h: &mut Hints| { - h.entries("group")[0][7] = count(SMALL_LEAF_SIZE - 1); - }), - ("group (a spurious duplicate slot)", &|h: &mut Hints| { - h.entries("group")[0][6] = count(1); - h.push("dup_pubkeys", vec![F64::ZERO; 4]); - }), - // One more group than the hint stream carries: witness generation - // has nothing to pop for it. - ("meta (n_epochs inflated)", &|h: &mut Hints| { - h.entries("meta")[0][0] = count(2); - }), - ("pubkeys (a key nobody signed for)", &|h: &mut Hints| { - h.entries("pubkeys")[0][0] += F64::ONE; - }), - // The window split of a list hash is advice, so both halves are pinned: - // the product identity ties them to the block count, and the tail's own - // range check keeps its `match` dispatch on a real arm. - ( - "signers_split (a window count the list does not have)", - &|h: &mut Hints| { - h.entries("signers_split")[0][0] = count(1); - }, - ), - ("signers_split (a tail past a whole window)", &|h: &mut Hints| { - h.entries("signers_split")[0][1] = count(SIGNERS_WINDOW); - }), - ("fs_seed", &|h: &mut Hints| { - h.entries("fs_seed")[0][0] += F64::ONE; - }), - ("leaf_defer", &|h: &mut Hints| { - h.entries("leaf_defer")[0][0] += F64::ONE; - }), - // A leaf derives its group's tweak table from this, so a wrong - // epoch is caught by the signatures long before the statement digest. - ("group (another epoch's tweak table)", &|h: &mut Hints| { - h.entries("group")[0][0] += F64::ONE; - }), - ("group (wider than the u32 the verifier holds)", &|h: &mut Hints| { - h.entries("group")[0][0] += F64(1 << 32); - }), - // The group's signatures were made over another message, so the - // encoding digests reject long before the statement digest. - ("group (another message under the signatures)", &|h: &mut Hints| { - h.entries("group")[0][1] += F64::ONE; - }), - ]; - for (description, tamper) in leaf_cases { - rejects(&[], raw_signatures.clone(), vec![], description, *tamper); - } - - // A leaf holding two epoch groups: the second group's region is one - // slot, so its writer cannot reach the first group's keys, and the two - // groups' tweak tables cannot be swapped. - let mut two_epoch_raw = at_epoch(&signers[..2], XMSS_EPOCH_A); - two_epoch_raw.extend(at_epoch(&get_signers_at(1, XMSS_EPOCH_B), XMSS_EPOCH_B)); - aggregate(&[], two_epoch_raw.clone(), vec![], &[], None, LOG_INV_RATE) - .expect("the honest two-epoch leaf aggregates"); - let two_epoch_cases: &[Tamper] = &[ - ( - "raw_index (an XMSS signature crossing into another epoch's region)", - &|h: &mut Hints| { - h.entries("raw_index")[2] = vec![count(1)]; - }, - ), - ("group (two epochs swapped)", &|h: &mut Hints| { - let entries = h.entries("group"); - let other = entries[1][0]; - entries[1][0] = entries[0][0]; - entries[0][0] = other; - }), - ("group (a group's count moved to the other)", &|h: &mut Hints| { - h.entries("group")[0][5] = count(1); - h.entries("group")[1][5] = count(2); - }), - // The declared count is advice: keep both table groups but hash only the - // first, while the statement still publishes both. - ("meta (a published group left out of the digest)", &|h: &mut Hints| { - h.entries("meta")[0][0] = count(1); - h.entries("meta")[0][1] = count(1); - }), - ]; - for (description, tamper) in two_epoch_cases { - rejects(&[], two_epoch_raw.clone(), vec![], description, *tamper); - } - - // A mixed leaf: three XMSS signers then two SPHINCS ones, so the XMSS - // region is slots 0..3 and the SPHINCS region 3..5. Each scheme's - // witness has to bind, and neither scheme's signature may cover the - // other's declared key, which is what the statement's split claims. - let mixed_xmss = at_epoch(&signers[..3], XMSS_EPOCH_A); - let mixed_sphincs = get_sphincs_signers(2); - aggregate(&[], mixed_xmss.clone(), mixed_sphincs.clone(), &[], None, LOG_INV_RATE) - .expect("the honest mixed leaf aggregates"); - let mixed_cases: &[Tamper] = &[ - ( - "raw_index (an XMSS signature reaching the SPHINCS region)", - &|h: &mut Hints| { - h.entries("raw_index")[0] = vec![count(3)]; - }, - ), - ("sp_raw_index (out of range)", &|h: &mut Hints| { - h.entries("sp_raw_index")[0] = vec![count(2)]; - }), - ("sp_raw_index (duplicate slot)", &|h: &mut Hints| { - let entries = h.entries("sp_raw_index"); - entries[1] = entries[0].clone(); - }), - ("meta (n_sphincs inflated)", &|h: &mut Hints| { - h.entries("meta")[0][2] = count(3); - }), - ("meta (n_raw_sphincs understated)", &|h: &mut Hints| { - h.entries("meta")[0][4] = count(1); - }), - ("sphincs_signers (a message nobody signed)", &|h: &mut Hints| { - h.entries("sphincs_signers")[0][4] += F64::ONE; - }), - ("sphincs_signers (a key nobody signed for)", &|h: &mut Hints| { - h.entries("sphincs_signers")[0][0] += F64::ONE; - }), - ("sp_rand (another randomizer, so another index)", &|h: &mut Hints| { - h.entries("sp_rand")[0][0] += F64::ONE; - }), - ("sp_counter", &|h: &mut Hints| { - h.entries("sp_counter")[0][0] += F64::ONE; - }), - ("sp_digits", &|h: &mut Hints| { - let entries = h.entries("sp_digits"); - entries[0][0] *= G; - }), - ("sp_chain_starts", &|h: &mut Hints| { - h.entries("sp_chain_starts")[0][0] += F64::ONE; - }), - ("sp_fts_secrets", &|h: &mut Hints| { - h.entries("sp_fts_secrets")[0][0] += F64::ONE; - }), - ("sp_fts_paths", &|h: &mut Hints| { - h.entries("sp_fts_paths")[0][0] += F64::ONE; - }), - ("sp_siblings", &|h: &mut Hints| { - h.entries("sp_siblings")[0][0] += F64::ONE; - }), - ]; - for (description, tamper) in mixed_cases { - rejects(&[], mixed_xmss.clone(), mixed_sphincs.clone(), description, *tamper); - } - - let left = prove_leaf(&signers[..SMALL_LEAF_SIZE]); - let right = prove_leaf(&signers[SMALL_LEAF_SIZE..]); - let children = vec![left, right]; - aggregate(&children, vec![], vec![], &[], None, LOG_INV_RATE).expect("the honest node aggregates"); - let node_cases: &[Tamper] = &[ - ("column placement (swapped)", &|h: &mut Hints| { - h.entries("col_sort_order")[0].swap(0, 1); - }), - ("column placement (duplicate)", &|h: &mut Hints| { - let order = &mut h.entries("col_sort_order")[0]; - order[1] = order[0]; - }), - ("column placement (out of range)", &|h: &mut Hints| { - let order = &mut h.entries("col_sort_order")[0]; - order[0] = count(order.len()); - }), - ("merkle cap (all hashes skipped)", &|h: &mut Hints| { - h.entries("merkle_cap_active")[0].fill(F64::ZERO); - }), - ("merkle cap (one ancestor skipped)", &|h: &mut Hints| { - let flags = &mut h.entries("merkle_cap_active")[0]; - let active = flags.iter_mut().skip(2).find(|x| **x == F64::ONE).unwrap(); - *active = F64::ZERO; - }), - ("merkle cap (root)", &|h: &mut Hints| { - h.entries("merkle_caps")[0][4] += F64::ONE; - }), - ("merkle cap (subtree)", &|h: &mut Hints| { - h.entries("merkle_caps")[0][8] += F64::ONE; - }), - ("merkle children (reversed)", &|h: &mut Hints| { - let children = &mut h.entries("merkle_children")[0]; - for word in 0..4 { - children.swap(word, word + 4); - } - }), - ("merkle path (below cap)", &|h: &mut Hints| { - h.entries("merkle_children")[0][0] += F64::ONE; - }), - ("merkle leaf", &|h: &mut Hints| { - h.entries("merkle_leaf_rows")[0][0] += F64::ONE; - }), - ( - "merkle leaf (an upper limb in a later level's row)", - &|h: &mut Hints| { - let rows = h.entries("merkle_leaf_rows"); - let row = rows - .iter_mut() - .find(|row| row.len() == 3 << pcs::whir_config::SUBSEQUENT_FOLDING_FACTOR) - .unwrap(); - row[1] += F64::ONE; - }, - ), - ("merkle path (last query)", &|h: &mut Hints| { - let paths = h.entries("merkle_children"); - *paths.last_mut().unwrap().last_mut().unwrap() += F64::ONE; - }), - ("child_index (duplicate slot)", &|h: &mut Hints| { - let entries = h.entries("child_index"); - entries[1] = entries[0].clone(); - }), - ("child_index (out of range)", &|h: &mut Hints| { - h.entries("child_index")[0][0] = count(2 * SMALL_LEAF_SIZE); - }), - ("child_group (count understated)", &|h: &mut Hints| { - h.entries("child_group")[0][5] = count(SMALL_LEAF_SIZE - 1); - }), - // A child group claimed at an epoch or under a message the child - // never carried: the map equality or the rebuilt digest rejects. - ("child_group (epoch)", &|h: &mut Hints| { - h.entries("child_group")[0][0] += F64::ONE; - }), - ("child_group (message)", &|h: &mut Hints| { - h.entries("child_group")[0][1] += F64::ONE; - }), - ("child_defer (a forged carried claim)", &|h: &mut Hints| { - h.entries("child_defer")[0][0] += F64::ONE; - }), - ("bc_star_hint", &|h: &mut Hints| { - h.entries("bc_star_hint")[0][0] += F64::ONE; - }), - ("mat_stars_hint", &|h: &mut Hints| { - h.entries("mat_stars_hint")[0][0] += F64::ONE; - }), - // The one hint carrying flock's whole lincheck terminal. Pinned not - // by the guest's own assert (which merely defines it) but by the - // matrix batching, whose reduced claims the root discharges against - // the real A_0/B_0. - ("matpart", &|h: &mut Hints| { - h.entries("matpart")[0][0] += F64::ONE; - }), - // A node holding no raw XMSS signature builds no tweak tables, so - // the statement digest is all that pins its epochs. The children's - // epochs must then land on slots of this altered list, and the map - // equality has no target, so the node cannot be proven. - ( - "group (a node that derives nothing from the epoch)", - &|h: &mut Hints| { - h.entries("group")[0][0] += F64::ONE; - }, - ), - ]; - for (description, tamper) in node_cases { - rejects(&children, vec![], vec![], description, *tamper); - } - - // A node over children of two different epochs: the hinted group map is - // what ties each child's group to the parent region of the same epoch. - let epoch_children = vec![ - prove_leaf(&signers[..2]), - aggregate( - &[], - at_epoch(&get_signers_at(2, XMSS_EPOCH_B), XMSS_EPOCH_B), - vec![], - &[], - None, - LOG_INV_RATE, - ) - .expect("the honest epoch-B leaf aggregates"), - ]; - aggregate(&epoch_children, vec![], vec![], &[], None, LOG_INV_RATE) - .expect("the honest two-epoch node aggregates"); - let epoch_node_cases: &[Tamper] = &[ - // Pointing the second child's group at the parent's epoch-A region: - // the epochs disagree, so the map equality fails. - ("child_group_map (a group mapped across epochs)", &|h: &mut Hints| { - h.entries("child_group_map")[1][0] = count(0); - }), - ("child_group_map (out of range)", &|h: &mut Hints| { - h.entries("child_group_map")[0][0] = count(2); - }), - ]; - for (description, tamper) in epoch_node_cases { - rejects(&epoch_children, vec![], vec![], description, *tamper); - } - - // The same discipline over a child's SPHINCS claims, which are rebuilt by - // their own loop (`hash_child_sphincs`) rather than the XMSS helper, so - // the cases above do not reach them: these children carry claims. - let sphincs = get_sphincs_signers(4); - let mixed_child = |x: &[(XmssPublicKey, XmssSignature)], s: &[RawSphincs]| { - aggregate(&[], at_epoch(x, XMSS_EPOCH_A), s.to_vec(), &[], None, LOG_INV_RATE) - .expect("the honest mixed child aggregates") - }; - let mixed_children = vec![ - mixed_child(&signers[..2], &sphincs[..2]), - mixed_child(&signers[2..4], &sphincs[2..]), - ]; - let mixed_node_cases: &[Tamper] = &[ - ("child_sphincs_index (duplicate slot)", &|h: &mut Hints| { - let entries = h.entries("child_sphincs_index"); - entries[1] = entries[0].clone(); - }), - ("child_sphincs_index (out of range)", &|h: &mut Hints| { - h.entries("child_sphincs_index")[0] = vec![count(4)]; - }), - ("child_meta (a child's SPHINCS count understated)", &|h: &mut Hints| { - h.entries("child_meta")[0][1] = count(1); - }), - ]; - for (description, tamper) in mixed_node_cases { - rejects(&mixed_children, vec![], vec![], description, *tamper); - } - } - - #[test] - #[ignore] - fn aggregate_rejects_a_bad_signature() { - lean_vm::init_prover_pool(); - let mut raw_signatures = at_epoch(&get_signers(3), XMSS_EPOCH_A); - raw_signatures[1].3.wots_signature.chain_tips[0][0] ^= 1; - let built = std::panic::catch_unwind(|| { - aggregate(&[], raw_signatures, vec![], &[], None, LOG_INV_RATE).map(|signature| signature.verify().is_ok()) - }); - assert!( - !matches!(built, Ok(Ok(true))), - "a forged signature must not produce a verifying aggregate" - ); - - let mut raw_sphincs = get_sphincs_signers(2); - raw_sphincs[1].2.ots[2][0][0] ^= 1; - let built = std::panic::catch_unwind(|| { - aggregate(&[], vec![], raw_sphincs, &[], None, LOG_INV_RATE).map(|signature| signature.verify().is_ok()) - }); - assert!( - !matches!(built, Ok(Ok(true))), - "a forged SPHINCS signature must not produce a verifying aggregate" - ); - } - - /// Randomness that does not decode to a target-sum encoding used to panic - /// the hint builder; it is a typed error now, and the abandoned builder - /// leaves nothing behind. - #[test] - fn malformed_raw_signature_is_an_error() { - let message = message(); - let pk = XmssPublicKey { - merkle_root: [0; xmss::DIGEST_LEN], - public_param: [0; xmss::PUBLIC_PARAM_LEN], - }; - let randomness = (0..=u8::MAX) - .find_map(|byte| { - let mut randomness = [0; xmss::RANDOMNESS_LEN]; - randomness[0] = byte; - xmss::wots_encode(&message, XMSS_EPOCH_A, &pk.public_param, &randomness) - .is_none() - .then_some(randomness) - }) - .expect("some randomness fails the target sum"); - let sig = XmssSignature { - wots_signature: xmss::WotsSignature { - chain_tips: [[0; xmss::DIGEST_LEN]; xmss::V], - randomness, - }, - merkle_proof: [[0; xmss::DIGEST_LEN]; xmss::LOG_LIFETIME], - }; - - let mut hints = Hints::default(); - assert_eq!( - push_signature_hints(&mut hints, &pk, &sig, &message, XMSS_EPOCH_A), - Err(AggregationError::MalformedRawSignature) - ); - assert!(hints.is_empty()); - assert_eq!( - aggregate( - &[], - vec![(pk, XMSS_EPOCH_A, message, sig)], - vec![], - &[], - None, - LOG_INV_RATE - ) - .err(), - Some(AggregationError::MalformedRawSignature) - ); - } - - /// The same for a SPHINCS claim, whose witness walk is equally fallible: a - /// counter that does not encode has no witness, and that is an error rather - /// than a panic inside the prover. - #[test] - fn malformed_raw_sphincs_signature_is_an_error() { - let (public_key, signed, mut signature) = get_sphincs_signers(1).pop().expect("one signer"); - signature.counters[sphincs::D - 1] ^= 1; - assert!(sphincs::verify(&public_key, &signed, &signature).is_err()); - let raw = vec![(public_key, signed, signature)]; - assert_eq!( - aggregate(&[], vec![], raw, &[], None, LOG_INV_RATE).err(), - Some(AggregationError::MalformedRawSignature) - ); - } - - /// Every `BLAKE2s` the guest itself runs reads two metadata cells an earlier - /// instruction of its own function wrote: a `SET` for a compile-time counter, - /// an `XOR` for a window's base plus its offset. An unwritten cell is - /// prover-chosen (write-once memory constrains only what something writes), so - /// a compression whose metadata nothing writes would hand the prover that - /// hash's byte counter and both flags, and every guest digest rests on those - /// being the ones the scheme specifies. The fill blocks are the deliberate - /// exception: their dummy reads a cell nothing writes, and nothing reads what - /// they compress (`lean_vm::cpu::filler`). - /// - /// This is a scan by pc, not a dominance check: a writer sitting in a branch - /// nobody took would satisfy it. What makes naming such a cell impossible is - /// `FnLower::scoped` reverting the constant pool at every join, and the - /// `blake2s_default_iv_*` tests are what guard that, by proving both paths. - #[test] - fn every_guest_blake2s_metadata_cell_is_written_first() { - use lean_vm::cpu::{DerefMode, Op}; - - // Which frame cell an instruction writes, if any. A `DEREF` in cell mode is - // bidirectional under write-once, so its local operand counts as a write. - let written = |op: &Op| match *op { - Op::Set { o, .. } => o..o + 1, - Op::Xor64 { c, .. } | Op::Mul64 { c, .. } => c..c + 1, - Op::Xor192 { c, .. } | Op::Mul192 { c, .. } => c..c + 3, - Op::Deref { - o3, - mode: DerefMode::Cell, - .. - } => o3..o3 + 1, - Op::Deref { .. } | Op::Jump { .. } => 0..0, - Op::Blake2s { out, .. } => out..out + 4, - }; - let program = unified_guest(); - let fill: Vec> = program - .filler - .iter() - .map(|b| b.pc as usize..(b.pc + b.size) as usize) - .collect(); - let mut unwritten = Vec::new(); - for (name, entry, len) in &program.fn_ranges { - let range = *entry as usize..(*entry + *len) as usize; - for pc in range.clone() { - let Op::Blake2s { md, .. } = program.prog[pc] else { - continue; - }; - if fill.iter().any(|f| f.contains(&pc)) { - continue; - } - for cell in md..md + 2 { - if !program.prog[range.start..pc] - .iter() - .any(|op| written(op).contains(&cell)) - { - unwritten.push(format!("{name} pc {pc} md fp[{cell}]")); - } - } - } - } - assert!(unwritten.is_empty(), "metadata cell never written: {unwritten:?}"); - } -} diff --git a/crates/rec_aggregation/src/benchmark.rs b/crates/rec_aggregation/src/benchmark.rs deleted file mode 100644 index c7f2392e2..000000000 --- a/crates/rec_aggregation/src/benchmark.rs +++ /dev/null @@ -1,252 +0,0 @@ -//! The two benchmarks: one leaf of the aggregation tree (`aggregate`), and an -//! n→1 recursion step over leaves of that size (`recursion`). Aggregation accepts -//! signature counts and a blob count; recursion accepts these counts per leaf. - -use primitives::bench::Plan; -use primitives::{pretty_f64, pretty_integer}; -use rand::{Rng, SeedableRng, rngs::StdRng}; -use xmss::{XmssPublicKey, XmssSignature}; - -use crate::aggregation::{DaInput, EthereumProof, aggregate, aggregate_with_stats}; -use crate::signers_cache; - -fn blobs(n: usize, seed: u64) -> Vec { - let mut rng = StdRng::seed_from_u64(seed); - (0..n * lean_da::BLOB_SYMBOLS).map(|_| rng.random()).collect() -} - -/// Cached signers `[from, to)`, as the aggregation API takes them: each raw -/// signature carries its epoch and message, the benchmarks using one pair for -/// all. -fn signers(from: usize, to: usize) -> Vec<(XmssPublicKey, xmss::Epoch, xmss::Message, XmssSignature)> { - if to == 0 { - return Vec::new(); - } - signers_cache::get_signers(to)[from..to] - .iter() - .map(|(pk, sig)| { - ( - pk.clone(), - signers_cache::XMSS_EPOCH_A, - signers_cache::message(), - sig.clone(), - ) - }) - .collect() -} - -/// Each SPHINCS signer comes with the message it signed, as the XMSS ones do. -fn sphincs_signers( - from: usize, - to: usize, -) -> Vec<(sphincs::SphincsPublicKey, sphincs::Message, sphincs::SphincsSignature)> { - if to == 0 { - return Vec::new(); - } - signers_cache::get_sphincs_signers(to)[from..to].to_vec() -} - -/// Report the shape and cost of one aggregation node. -fn report(label: &str, stats: &lean_vm::cpu::Stats, sig: &EthereumProof, prove_time: &primitives::bench::Timing) { - let base_cycles: usize = stats.base_counts.iter().sum(); - println!("{label}"); - // The program's own work, then what gets proven: the fill blocks bring each table - // to a power of two so that none needs padding rows, so the proven total is the sum - // of those powers. - println!( - " cycles (VM steps) : {} = {}", - pretty_integer(base_cycles), - crate::report::pow(base_cycles) - ); - println!(" details : {}", stats.details()); - crate::report::print_proof_size(sig.proof()); - // The whole `aggregate` call, not just `cpu::prove`: for a node that also - // covers verifying each child and batching the deferred claims, which are - // real per-node costs. `--tracing` breaks it down. - println!( - " proving time : {} s{} peak memory {} GiB", - pretty_f64(prove_time.mean()), - prove_time.spread(), - crate::report::peak_gib() - ); -} - -/// How a benchmark names a leaf of either scheme or of both. -fn describe(n_xmss: usize, n_sphincs: usize) -> String { - match (n_xmss, n_sphincs) { - (x, 0) => format!("{} XMSS", pretty_integer(x)), - (0, s) => format!("{} SPHINCS", pretty_integer(s)), - (x, s) => format!("{} XMSS and {} SPHINCS", pretty_integer(x), pretty_integer(s)), - } -} - -/// Prove signatures and `n_blobs` blobs in one leaf, then verify it. -/// -/// Proving runs one discarded warmup pass followed by `plan.repeat` measured -/// passes; see [`primitives::bench`] for why the first pass is not -/// representative and why the cooldown matters. -pub fn run_aggregation(n_xmss: usize, n_sphincs: usize, n_blobs: usize, log_inv_rate: usize, plan: Plan) { - assert!( - n_xmss > 0 || n_sphincs > 0 || n_blobs > 0, - "a leaf needs signatures or blobs" - ); - assert!(n_blobs <= lean_da::DA_MAX_ROWS, "too many blobs in one commitment"); - let trace_span = tracing::info_span!("aggregation", n_xmss, n_sphincs, n_blobs, log_inv_rate).entered(); - // Spawn the worker pool before any timed work, so no kernel pays the spawn - // cost. Opting into the arena is the calling *process's* decision (one region, - // one proof at a time), so it stays in `main`, not here. - lean_vm::init_prover_pool(); - let raw_xmss = signers(0, n_xmss); - let raw_sphincs = sphincs_signers(0, n_sphincs); - let blobs = blobs(n_blobs, 0); - // Only the final measured pass of each stage is traced: the tree describes the - // proof the reported timings are about, instead of repeating itself per pass. - let ((sig, stats), prove_time) = plan.warm_then_measure(|last| { - let _quiet = (!last).then(primitives::suppress_tracing); - aggregate_with_stats( - &[], - raw_xmss.clone(), - raw_sphincs.clone(), - None, - DaInput { - rows: &blobs, - roots: None, - }, - log_inv_rate, - ) - .expect("leaf aggregates") - }); - let (_, verify_time) = Plan::new(plan.repeat, 0).measure_quiet(|last| { - let _quiet = (!last).then(primitives::suppress_tracing); - sig.verify().expect("the leaf aggregate verifies"); - }); - drop(trace_span); - - let mut inputs = Vec::new(); - if n_xmss > 0 || n_sphincs > 0 { - inputs.push(format!("{} signatures", describe(n_xmss, n_sphincs))); - } - if n_blobs > 0 { - inputs.push(format!("{} blobs", pretty_integer(n_blobs))); - } - report( - &format!("\naggregation, {}", inputs.join(", ")), - &stats, - &sig, - &prove_time, - ); - if n_xmss > 0 || n_sphincs > 0 { - println!( - " per signature : {} signatures/s", - pretty_f64((n_xmss + n_sphincs) as f64 / prove_time.mean()) - ); - } - if n_blobs > 0 { - let payload_mib = (blobs.len() * size_of::()) as f64 / (1 << 20) as f64; - println!( - " blob throughput : {} blobs/s, {} MiB/s", - pretty_f64(n_blobs as f64 / prove_time.mean()), - pretty_f64(payload_mib / prove_time.mean()) - ); - } - println!( - " verifying : {} ms", - pretty_f64(verify_time.mean() * 1000.0) - ); -} - -/// Prove `n` leaves with signatures and distinct blob payloads, then aggregate them in one -/// recursion step and verify the result. The leaves are built once; only the -/// recursion step is measured. -pub fn run_recursion( - n: usize, - per_leaf: usize, - sphincs_per_leaf: usize, - blobs_per_leaf: usize, - log_inv_rate: usize, - enable_tracing: bool, - plan: Plan, -) { - assert!((1..=crate::MAX_RECURSIONS).contains(&n), "invalid child count"); - assert!( - per_leaf > 0 || sphincs_per_leaf > 0 || blobs_per_leaf > 0, - "a leaf needs signatures or blobs" - ); - assert!( - blobs_per_leaf <= lean_da::DA_MAX_ROWS, - "too many blobs in one commitment" - ); - assert!( - blobs_per_leaf == 0 || n <= crate::MAX_DA_ROOTS, - "too many distinct DA roots" - ); - lean_vm::init_prover_pool(); - let all = signers(0, n * per_leaf); - let all_sphincs = sphincs_signers(0, n * sphincs_per_leaf); - let started = std::time::Instant::now(); - let guest_instructions: usize = crate::aggregation::unified_guest() - .fn_ranges - .iter() - .map(|(_, _, len)| *len as usize) - .sum(); - let compile_time = started.elapsed(); - - let children: Vec = (0..n) - .map(|k| { - aggregate( - &[], - all[k * per_leaf..(k + 1) * per_leaf].to_vec(), - all_sphincs[k * sphincs_per_leaf..(k + 1) * sphincs_per_leaf].to_vec(), - &blobs(blobs_per_leaf, k as u64), - None, - log_inv_rate, - ) - .expect("leaf aggregates") - }) - .collect(); - - if enable_tracing { - primitives::init_tracing(); - } - let ((sig, stats), prove_time) = plan.warm_then_measure(|last| { - let _quiet = (!last).then(primitives::suppress_tracing); - aggregate_with_stats(&children, vec![], vec![], None, DaInput::default(), log_inv_rate) - .expect("node aggregates") - }); - let (_, verify_time) = Plan::new(plan.repeat, 0).measure_quiet(|last| { - let _quiet = (!last).then(primitives::suppress_tracing); - sig.verify().expect("the recursive aggregate verifies"); - }); - - println!( - "aggregation bytecode: {} instructions (2^{} padded), compiled in {} s", - pretty_integer(guest_instructions), - crate::aggregation::unified_guest().prog.len().trailing_zeros(), - pretty_f64(compile_time.as_secs_f64()) - ); - let mut inputs = Vec::new(); - if per_leaf > 0 || sphincs_per_leaf > 0 { - inputs.push(format!("{} signatures", describe(per_leaf, sphincs_per_leaf))); - } - if blobs_per_leaf > 0 { - inputs.push(format!("{blobs_per_leaf} blobs")); - } - report( - &format!("\nrecursion {n}\u{2192}1, over leaves of {}", inputs.join(", ")), - &stats, - &sig, - &prove_time, - ); - if blobs_per_leaf > 0 { - assert_eq!( - sig.da_commitments().len(), - n, - "every child's distinct root must survive recursion" - ); - println!(" retained DA roots : {n}"); - } - println!( - " verifying : {} ms", - pretty_f64(verify_time.mean() * 1000.0) - ); -} diff --git a/crates/rec_aggregation/src/hash_chain.rs b/crates/rec_aggregation/src/hash_chain.rs deleted file mode 100644 index 5e4e382da..000000000 --- a/crates/rec_aggregation/src/hash_chain.rs +++ /dev/null @@ -1,117 +0,0 @@ -//! End-to-end proof of an unrolled zkDSL BLAKE2s hash chain. - -use std::time::Instant; - -use lean_compiler::{compile, compile_without_filler, parse}; -use lean_vm::cpu::{prove, verify}; -use lean_vm::vmhash::compress; -use primitives::{field::F64, pretty_f64, pretty_integer}; - -fn instruction_counts(source: &str, public_input: [F64; 4]) -> [usize; lean_vm::cpu::Stats::TABLES.len()] { - compile_without_filler(&parse(source).expect("parse")) - .execute(public_input) - .base_counts -} - -fn chain_source(steps: usize, unroll: usize) -> String { - assert!( - unroll >= 1 && steps.is_multiple_of(unroll), - "N must be a positive multiple of UNROLL" - ); - let blocks = steps / unroll; - let result_cell = 4 * blocks; - - let mut body = String::new(); - body.push_str(" b = i ** 4\n"); - body.push_str(" h0 = StackBuf(4)\n"); - body.push_str(" h0[0:4] = buff[b:b + 4]\n"); - for step in 1..=unroll { - body.push_str(&format!(" h{step} = StackBuf(4)\n")); - body.push_str(&format!( - " blake2s(h{previous}, h{previous}, h{step})\n", - previous = step - 1 - )); - } - body.push_str(" nxt = b * GEN ** 4\n"); - body.push_str(&format!(" buff[nxt:nxt + 4] = h{unroll}\n")); - - format!( - "def main():\n\ - \x20 buff = HeapBuf({size})\n\ - \x20 buff[0:4] = [0, 0, 0, 0]\n\ - \x20 for i in mul_range(1, GEN ** {blocks}):\n\ - {body}\ - \x20 output = 1\n\ - \x20 output[0:4] = buff[{result_cell}:{result_cell} + 4]\n\ - \x20 return\n", - size = result_cell + 4, - ) -} - -#[test] -fn blake2s_hash_chain() { - let env_usize = |key: &str, default: usize| { - std::env::var(key) - .ok() - .and_then(|value| value.parse().ok()) - .unwrap_or(default) - }; - let unroll = env_usize("LEANVM_HASH_UNROLL", 4); - let steps = env_usize("LEANVM_HASH_N", 8); - assert!( - steps.is_multiple_of(unroll), - "LEANVM_HASH_N must be a multiple of LEANVM_HASH_UNROLL" - ); - - let mut public_input = [F64::ZERO; 4]; - for _ in 0..steps { - public_input = compress(public_input, public_input); - } - - let source = chain_source(steps, unroll); - let program = compile(&parse(&source).expect("parse")); - - let started = Instant::now(); - let (proof, stats) = prove(&program, public_input, lean_vm::pcs::TEST_LOG_INV_RATE); - let prove_time = started.elapsed(); - let started = Instant::now(); - verify(&program, &public_input, &proof).expect("hash-chain proof verifies"); - let verify_time = started.elapsed(); - - assert_eq!(instruction_counts(&source, public_input)[5], steps); - - println!( - "\nBLAKE2s hash chain, N = {}, unroll = {}", - pretty_integer(steps), - pretty_integer(unroll) - ); - println!(" cycles (VM steps) : {}", pretty_integer(stats.cycles)); - for (name, &c) in lean_vm::cpu::Stats::TABLES.iter().zip(&stats.counts) { - let pow = if c == 0 { - "0".to_string() - } else { - format!("2^{}", pretty_f64((c as f64).log2())) - }; - println!(" {name:<7} instructions : {pow}"); - } - println!( - " committed witness size : 2^{:.3}", - (stats.committed as f64).log2() - ); - let proof_bytes = bincode::serialized_size(&proof).expect("proof is serializable"); - println!(" proof size : {:.1} KiB", proof_bytes as f64 / 1024.0); - println!(" proving : {prove_time:?}"); - println!( - " verifying : {} ms", - pretty_f64(verify_time.as_secs_f64() * 1000.0) - ); - let hashes_per_second = (steps as f64 / prove_time.as_secs_f64()).round() as u64; - println!( - " throughput : {} hashes/s", - pretty_integer(hashes_per_second) - ); - - let mut wrong_input = public_input; - wrong_input[0] += F64::ONE; - assert!(verify(&program, &wrong_input, &proof).is_err()); -} diff --git a/crates/rec_aggregation/src/lib.rs b/crates/rec_aggregation/src/lib.rs deleted file mode 100644 index b1221e4bc..000000000 --- a/crates/rec_aggregation/src/lib.rs +++ /dev/null @@ -1,45 +0,0 @@ -pub mod aggregation; -pub mod benchmark; -pub mod fibonacci; -/// The BLAKE2s hash chain, proven end to end. A `src` module rather than its own -/// test binary so it shares the process, and so the slow flock circuit build, -/// with the other workloads. -#[cfg(test)] -mod hash_chain; -pub mod signers_cache; - -pub use aggregation::{ - AggregateVerifyError, AggregationError, ClaimSelection, DA_LOG_CELL, DA_LOG_K, DA_MAX_ROWS, EthereumProof, - MAX_DA_ROOTS, MAX_EPOCHS, MAX_KEYS, MAX_RECURSIONS, SignatureClaims, SphincsClaim, XmssClaimGroup, aggregate, - warm_up, -}; -pub use benchmark::{run_aggregation, run_recursion}; -pub use fibonacci::run_fibonacci; - -/// The pieces every workload's benchmark report ends with. -/// -/// Each caller drops its root tracing span before printing: tracing-forest -/// renders its tree only when that span closes, so the complete trace has to be -/// flushed above the report. -mod report { - use primitives::pretty_f64; - - /// A count as a power of two, or a dash when the opcode never ran. - pub fn pow(x: usize) -> String { - if x == 0 { - " -".into() - } else { - format!("2^{}", pretty_f64((x as f64).log2())) - } - } - - /// Peak resident set size, in GiB. - pub fn peak_gib() -> String { - pretty_f64(primitives::bench::peak_rss_bytes() as f64 / (1u64 << 30) as f64) - } - - pub fn print_proof_size(proof: &T) { - let bytes = bincode::serialized_size(proof).expect("proof is serializable"); - println!(" proof size : {:.1} KiB", bytes as f64 / 1024.0); - } -} diff --git a/crates/rec_aggregation/src/signers_cache.rs b/crates/rec_aggregation/src/signers_cache.rs deleted file mode 100644 index 1515bd0aa..000000000 --- a/crates/rec_aggregation/src/signers_cache.rs +++ /dev/null @@ -1,322 +0,0 @@ -//! Persistent cache for deterministic benchmark signatures, one file per -//! scheme. -//! -//! The cache grows as needed and is memoized in-process. Its filename binds the -//! parameters, hash construction, and encoding predicate. Loaded signatures are -//! also verified, so stale entries are regenerated from the first invalid one. - -use std::collections::BTreeMap; -use std::collections::hash_map::DefaultHasher; -use std::fs; -use std::hash::{Hash, Hasher}; -use std::io::Write; -use std::path::PathBuf; -use std::sync::Mutex; -use std::time::Instant; - -use primitives::{pretty_f64, pretty_integer}; -use rand::SeedableRng; -use rand::rngs::StdRng; -use xmss::*; - -type CachedSignature = (XmssPublicKey, XmssSignature); - -const SCHEMA_VERSION: u32 = 3; - -/// The epoch `get_signers` signs at. SPHINCS has none. -pub const XMSS_EPOCH_A: Epoch = 3_000_000_007; -/// A second epoch for multi-epoch tests; signer `i` holds the same key at both. -pub const XMSS_EPOCH_B: Epoch = 3_000_000_009; -/// First epoch the cached keys are activated at, so a test that wants many epoch -/// groups can walk the whole window (`key_gen`'s range is inclusive). -pub const KEY_START: Epoch = 3_000_000_000; -const KEY_END: Epoch = 3_000_000_015; - -pub fn message() -> Message { - std::array::from_fn(|i| (i * 5 + 1) as u8) -} - -/// The message signed at `epoch`: distinct per epoch, and exactly [`message`] -/// at [`XMSS_EPOCH_A`], so pre-existing cache files stay valid. -pub fn message_for(epoch: Epoch) -> Message { - let mut msg = message(); - for (byte, delta) in msg.iter_mut().zip((epoch ^ XMSS_EPOCH_A).to_le_bytes()) { - *byte ^= delta; - } - msg -} - -fn compute_signer(index: usize, epoch: Epoch) -> CachedSignature { - // The index over its full width: a one-byte seed repeats every 256 signers, - // and a repeated signer is invisible until something deduplicates the set, - // at which point a batch of 900 quietly becomes one of 256. - let mut seed = [10u8; 32]; - seed[..8].copy_from_slice(&(index as u64).to_le_bytes()); - let (sk, pk) = xmss::key_gen_from_seed(seed, KEY_START, KEY_END).expect("keygen"); - let sig = xmss::sign(&sk, &message_for(epoch), epoch).expect("sign"); - (pk, sig) -} - -fn hash_fingerprint() -> [Digest; 2] { - let pp = [0xA5u8; PUBLIC_PARAM_LEN]; - [ - tweak_hash(&pp, TWEAK_TYPE_CHAIN, 1, 2, &[0x5Au8; DIGEST_LEN]), - tweak_hash(&pp, TWEAK_TYPE_ENCODING, 3, 4, &[0x3Cu8; 2 * STATE_LEN]), - ] -} - -fn encoding_fingerprint(epoch: Epoch) -> (u64, [u8; V]) { - let pp = [0xA5u8; PUBLIC_PARAM_LEN]; - let msg = message_for(epoch); - for counter in 0u64.. { - let mut randomness = [0u8; RANDOMNESS_LEN]; - randomness[..8].copy_from_slice(&counter.to_le_bytes()); - if let Some(digits) = wots_encode(&msg, epoch, &pp, &randomness) { - return (counter, digits); - } - } - unreachable!("some counter randomness encodes") -} - -fn footprint(epoch: Epoch) -> u64 { - let mut hasher = DefaultHasher::new(); - SCHEMA_VERSION.hash(&mut hasher); - epoch.hash(&mut hasher); - KEY_START.hash(&mut hasher); - KEY_END.hash(&mut hasher); - message_for(epoch).hash(&mut hasher); - (V, W, CHAIN_LENGTH, LOG_LIFETIME, TARGET_SUM, RANDOMNESS_LEN).hash(&mut hasher); - hash_fingerprint().hash(&mut hasher); - encoding_fingerprint(epoch).hash(&mut hasher); - hasher.finish() -} - -fn cache_dir() -> PathBuf { - PathBuf::from(env!("CARGO_MANIFEST_DIR")).join("../../target/signers-cache") -} - -fn cache_path(epoch: Epoch) -> PathBuf { - cache_dir().join(format!("xmss_signers_{:016x}.bin", footprint(epoch))) -} - -fn try_load_cache(epoch: Epoch) -> Option> { - let bytes = fs::read(cache_path(epoch)).ok()?; - let (version, mut signers): (u32, Vec) = bincode::deserialize(&bytes).ok()?; - if version != SCHEMA_VERSION { - return None; - } - let msg = message_for(epoch); - let valid = signers - .iter() - .take_while(|(pk, sig)| xmss::verify(pk, &msg, sig, epoch).is_ok()) - .count(); - if valid < signers.len() { - eprintln!( - "warning: signers cache {} is stale (signer {valid} of {} no longer verifies); regenerating from there", - cache_path(epoch).display(), - signers.len() - ); - signers.truncate(valid); - } - Some(signers) -} - -fn save_cache(signers: &[CachedSignature], epoch: Epoch) { - let path = cache_path(epoch); - if let Some(parent) = path.parent() { - let _ = fs::create_dir_all(parent); - } - let bytes = bincode::serialize(&(SCHEMA_VERSION, signers)).expect("serialize signers cache"); - if let Err(error) = fs::write(&path, &bytes) { - eprintln!("warning: could not write signers cache to {}: {error}", path.display()); - } -} - -fn generate_range(start: usize, end: usize, epoch: Epoch) -> Vec { - let total = end - start; - let started = Instant::now(); - let mut signers = Vec::with_capacity(total); - for (done, index) in (start..end).enumerate() { - signers.push(compute_signer(index, epoch)); - print!( - "\r generating XMSS signers (one-time, then cached): {}/{}", - pretty_integer(done + 1), - pretty_integer(total) - ); - let _ = std::io::stdout().flush(); - } - println!( - "\r generated {} XMSS in {} s (cached to disk) ", - pretty_integer(total), - pretty_f64(started.elapsed().as_secs_f64()) - ); - signers -} - -static POOLS: Mutex>> = Mutex::new(BTreeMap::new()); - -pub fn get_signers(n: usize) -> Vec { - get_signers_at(n, XMSS_EPOCH_A) -} - -/// The first `n` cached signers, signing [`message_for`]`(epoch)` at `epoch`: -/// one cache file per epoch, the keys shared across them. -pub fn get_signers_at(n: usize, epoch: Epoch) -> Vec { - let mut pools = POOLS.lock().unwrap(); - let pool = pools.entry(epoch).or_default(); - if pool.len() < n { - if let Some(disk) = try_load_cache(epoch) - && disk.len() > pool.len() - { - *pool = disk; - } - if pool.len() < n { - let mut fresh = generate_range(pool.len(), n, epoch); - pool.append(&mut fresh); - save_cache(pool, epoch); - } - } - pool[..n].to_vec() -} - -/// A SPHINCS signer, generated the same way, with the message it signed. -/// Signing is stateless, so unlike XMSS there is no epoch and no key range: -/// one key answers for every index. -type CachedSphincsSignature = (sphincs::SphincsPublicKey, sphincs::Message, sphincs::SphincsSignature); - -/// Signer `index`'s own message, distinct from every other's and from the -/// XMSS ones, so a test that mixed them up would fail rather than pass. -pub fn sphincs_message(index: usize) -> sphincs::Message { - let mut msg = [0u8; sphincs::MESSAGE_LEN]; - msg[..8].copy_from_slice(&(index as u64).to_le_bytes()); - msg[8..].copy_from_slice(&[0xC5; sphincs::MESSAGE_LEN - 8]); - msg -} - -/// One cached SPHINCS signer, as fixed-size bytes: the scheme's own -/// serializations, so nothing here has to agree with a derived one. -const SPHINCS_RECORD: usize = sphincs::PUB_KEY_SIZE + sphincs::MESSAGE_LEN + sphincs::SIG_SIZE; - -fn compute_sphincs_signer(index: usize) -> CachedSphincsSignature { - let mut rng = StdRng::seed_from_u64(0x5F1A_C500 ^ index as u64); - let (secret_key, public_key) = sphincs::key_gen(&mut rng); - let message = sphincs_message(index); - let signature = sphincs::sign(&secret_key, &message).expect("sign"); - (public_key, message, signature) -} - -fn sphincs_footprint() -> u64 { - let mut hasher = DefaultHasher::new(); - SCHEMA_VERSION.hash(&mut hasher); - // The record layout and the per-signer messages, so a change to either - // invalidates the file rather than being read back as another scheme's. - SPHINCS_RECORD.hash(&mut hasher); - sphincs_message(0).hash(&mut hasher); - sphincs_message(1).hash(&mut hasher); - ( - sphincs::MASTER_SECRET_LEN, - sphincs::V, - sphincs::W, - sphincs::TARGET_SUM, - sphincs::D, - sphincs::HEIGHTS, - sphincs::A, - sphincs::K, - ) - .hash(&mut hasher); - // The tweakable hash itself, so a change to it invalidates the file. - sphincs::th( - &[0xA5; sphincs::PUBLIC_PARAM_LEN], - &sphincs::tweak(1, 2, 3, 4, 5), - &[0x3C; 16], - ) - .hash(&mut hasher); - hasher.finish() -} - -fn sphincs_cache_path() -> PathBuf { - cache_dir().join(format!("sphincs_signers_{:016x}.bin", sphincs_footprint())) -} - -fn try_load_sphincs_cache() -> Option> { - let bytes = fs::read(sphincs_cache_path()).ok()?; - let mut signers = Vec::with_capacity(bytes.len() / SPHINCS_RECORD); - for record in bytes.as_chunks::().0 { - let (key_bytes, rest) = record.split_at(sphincs::PUB_KEY_SIZE); - let (message_bytes, signature_bytes) = rest.split_at(sphincs::MESSAGE_LEN); - let public_key = sphincs::SphincsPublicKey::from_bytes(key_bytes.try_into().unwrap()); - let message: sphincs::Message = message_bytes.try_into().unwrap(); - let signature = sphincs::SphincsSignature::from_bytes(signature_bytes.try_into().unwrap()); - if sphincs::verify(&public_key, &message, &signature).is_err() { - eprintln!( - "warning: signers cache {} is stale (signer {} no longer verifies); regenerating from there", - sphincs_cache_path().display(), - signers.len() - ); - break; - } - signers.push((public_key, message, signature)); - } - Some(signers) -} - -fn save_sphincs_cache(signers: &[CachedSphincsSignature]) { - let path = sphincs_cache_path(); - if let Some(parent) = path.parent() { - let _ = fs::create_dir_all(parent); - } - let mut bytes = Vec::with_capacity(signers.len() * SPHINCS_RECORD); - for (public_key, message, signature) in signers { - bytes.extend_from_slice(&public_key.flatten()); - bytes.extend_from_slice(message); - bytes.extend_from_slice(&signature.to_bytes()); - } - if let Err(error) = fs::write(&path, &bytes) { - eprintln!("warning: could not write signers cache to {}: {error}", path.display()); - } -} - -static SPHINCS_POOL: Mutex> = Mutex::new(Vec::new()); - -pub fn get_sphincs_signers(n: usize) -> Vec { - let mut pool = SPHINCS_POOL.lock().unwrap(); - if pool.len() < n { - if let Some(disk) = try_load_sphincs_cache() - && disk.len() > pool.len() - { - *pool = disk; - } - // Key generation is one whole 2^12-leaf tree, which is the expensive - // part; it fans out internally, so this loop stays sequential. - let started = Instant::now(); - let missing = n.saturating_sub(pool.len()); - for index in pool.len()..n { - pool.push(compute_sphincs_signer(index)); - print!( - "\r generating SPHINCS signers (one-time, then cached): {}/{}", - pretty_integer(index + 1 - (n - missing)), - pretty_integer(missing) - ); - let _ = std::io::stdout().flush(); - } - if missing > 0 { - println!( - "\r generated {} SPHINCS in {} s (cached to disk) ", - pretty_integer(missing), - pretty_f64(started.elapsed().as_secs_f64()) - ); - save_sphincs_cache(&pool); - } - } - pool[..n].to_vec() -} - -#[cfg(test)] -mod tests { - use super::*; - - #[test] - fn signer_seed_uses_more_than_one_byte() { - assert_ne!(compute_signer(0, XMSS_EPOCH_A).0, compute_signer(256, XMSS_EPOCH_A).0); - } -} diff --git a/crates/rec_aggregation/tests/arena_prove.rs b/crates/rec_aggregation/tests/arena_prove.rs deleted file mode 100644 index 0097859d8..000000000 --- a/crates/rec_aggregation/tests/arena_prove.rs +++ /dev/null @@ -1,27 +0,0 @@ -//! End-to-end proof verification across arena phase resets. -//! -//! This has its own binary because arena phases are process-global and cannot nest. - -use primitives::bench::Plan; - -#[test] -fn repeated_proofs_survive_phase_resets() { - lean_vm::init_prover(); - assert!( - zk_alloc::is_enabled(), - "this test is meaningless unless the arena is engaged" - ); - - rec_aggregation::run_aggregation(3, 1, 1, lean_vm::pcs::TEST_LOG_INV_RATE, Plan::new(2, 0)); - - let stats = zk_alloc::stats(); - assert!(stats.phases >= 3, "expected one phase per proof, got {stats:?}"); - assert!( - stats.peak_bytes > 0, - "no buffer reached the arena, so nothing was actually exercised: {stats:?}" - ); - assert_eq!( - stats.overflow, 0, - "a slab overflowed into the system allocator, so SLAB_SIZE is undersized: {stats:?}" - ); -} diff --git a/crates/sphincs/Cargo.toml b/crates/sphincs/Cargo.toml deleted file mode 100644 index 7ecb4abe6..000000000 --- a/crates/sphincs/Cargo.toml +++ /dev/null @@ -1,14 +0,0 @@ -[package] -name = "sphincs" -version.workspace = true -edition.workspace = true -publish = false - -[lints] -workspace = true - -[dependencies] -primitives.workspace = true -parallel.workspace = true -rand.workspace = true -serde.workspace = true diff --git a/crates/sphincs/src/fts.rs b/crates/sphincs/src/fts.rs deleted file mode 100644 index f139f754f..000000000 --- a/crates/sphincs/src/fts.rs +++ /dev/null @@ -1,86 +0,0 @@ -//! The few-time signature: a forest of `k-1` Merkle trees of `2^a` secret -//! leaves, one leaf opened per tree at an index the message digest picks -//! (FORS+C). -//! -//! Reuse leaks rather than breaks: after `r` signatures on one instance an -//! adversary holds `r` leaves per tree, and can sign a message only if it lands -//! on that instance, has last index zero, and has every other index on a leaf it -//! already holds. - -use crate::*; - -/// What a signature carries for the few-time key: the opened secret and the -/// Merkle path of each of the `k-1` trees. -#[derive(Clone, Debug, PartialEq, Eq)] -pub struct FtsOpening { - pub secrets: [Digest; NUM_FTS_TREES], - pub paths: [[Digest; A]; NUM_FTS_TREES], -} - -/// `s_{idx,kappa,j} = Th(P, tw_ftsprf(idx,kappa,j), S)`. -fn fts_secret(pp: &PublicParam, master: &MasterSecret, idx: u64, kappa: usize, j: usize) -> Digest { - th(pp, &tweak(TWEAK_FTS_PRF, kappa, idx as u32, 0, j as u32), master) -} - -fn fts_leaf(pp: &PublicParam, idx: u64, kappa: usize, j: usize, secret: &Digest) -> Digest { - th(pp, &tweak(TWEAK_FTS_LEAF, kappa, idx as u32, 0, j as u32), secret) -} - -fn fts_node(pp: &PublicParam, idx: u64, kappa: usize, level: usize, j: usize, left: &Digest, right: &Digest) -> Digest { - let tw = tweak(TWEAK_FTS_NODE, kappa, idx as u32, level as u32, j as u32); - th_digests(pp, &tw, &[*left, *right]) -} - -/// `Fts.key`: the few-time public key, `Th` over the `k-1` roots. -fn fts_key_of_roots(pp: &PublicParam, idx: u64, roots: &[Digest; NUM_FTS_TREES]) -> Digest { - th_digests(pp, &tweak(TWEAK_FTS_ROOTS, 0, idx as u32, 0, 0), roots) -} - -/// `Fts.key` and `Fts.open` together, the forest being built once. `u[k-1]` is -/// ignored: its tree is the dropped one. -pub fn fts_open(pp: &PublicParam, master: &MasterSecret, idx: u64, u: &[u32; K]) -> (Digest, FtsOpening) { - let mut opening = FtsOpening { - secrets: [[0; N]; NUM_FTS_TREES], - paths: [[[0; N]; A]; NUM_FTS_TREES], - }; - let mut roots = [[0; N]; NUM_FTS_TREES]; - for kappa in 0..NUM_FTS_TREES { - let opened = u[kappa] as usize; - let mut nodes = Vec::with_capacity(1 << A); - for j in 0..1 << A { - let secret = fts_secret(pp, master, idx, kappa, j); - if j == opened { - opening.secrets[kappa] = secret; - } - nodes.push(fts_leaf(pp, idx, kappa, j, &secret)); - } - for level in 0..A { - opening.paths[kappa][level] = nodes[(opened >> level) ^ 1]; - nodes = (0..nodes.len() / 2) - .map(|j| fts_node(pp, idx, kappa, level + 1, j, &nodes[2 * j], &nodes[2 * j + 1])) - .collect(); - } - roots[kappa] = nodes[0]; - } - (fts_key_of_roots(pp, idx, &roots), opening) -} - -/// `Fts.recover`: the few-time key an opening reaches, which is `Fts.key` on an -/// opening of the leaves `u` of that instance and nothing else short of a -/// collision. -pub fn fts_recover(pp: &PublicParam, idx: u64, u: &[u32; K], opening: &FtsOpening) -> Digest { - let roots = std::array::from_fn(|kappa| { - let opened = u[kappa] as usize; - let leaf = fts_leaf(pp, idx, kappa, opened, &opening.secrets[kappa]); - (0..A).fold(leaf, |node, level| { - let sibling = &opening.paths[kappa][level]; - let (left, right) = if (opened >> level) & 1 == 0 { - (node, *sibling) - } else { - (*sibling, node) - }; - fts_node(pp, idx, kappa, level + 1, opened >> (level + 1), &left, &right) - }) - }); - fts_key_of_roots(pp, idx, &roots) -} diff --git a/crates/sphincs/src/hash.rs b/crates/sphincs/src/hash.rs deleted file mode 100644 index 3bda543a2..000000000 --- a/crates/sphincs/src/hash.rs +++ /dev/null @@ -1,60 +0,0 @@ -//! The tweakable hash `Th(P, tw, M) = Truncate_n(BLAKE2s(tw | P | M))`, and the -//! 16-byte tweak that names one hash call in the whole structure. -//! -//! Compressions per call, the input including the 32 bytes of tweak and public -//! parameter: 1 for a chain step, a Merkle node, a derived secret and an -//! encoding, 2 for the message digest, 4 for the few-time roots, and 11 for a -//! one-time leaf. - -use crate::*; - -pub const TWEAK_LEN: usize = 16; -pub type Tweak = [u8; TWEAK_LEN]; - -pub const PROTOCOL_DOMAIN_SEP: u8 = 1; - -// Tweak types (byte 1). -pub const TWEAK_PRF: u8 = 0; -pub const TWEAK_CHAIN: u8 = 1; -pub const TWEAK_LEAF: u8 = 2; -pub const TWEAK_NODE: u8 = 3; -pub const TWEAK_ENC: u8 = 4; -pub const TWEAK_FTS_PRF: u8 = 5; -pub const TWEAK_FTS_LEAF: u8 = 6; -pub const TWEAK_FTS_NODE: u8 = 7; -pub const TWEAK_FTS_ROOTS: u8 = 8; -pub const TWEAK_MSG: u8 = 9; -pub const TWEAK_PARAMETER: u8 = 10; -pub const TWEAK_RANDOMIZER: u8 = 12; - -/// `[protocol_domain_sep:1 | type:1 | layer:1 | zero:1 | p:4 | tree:4 | index:4]`, little endian. -/// `lay` identifies a hypertree layer or a tree of a few-time forest. -pub fn tweak(t: u8, lay: usize, tau: u32, p: u32, j: u32) -> Tweak { - debug_assert!(lay < 256); - let mut tw = [0u8; TWEAK_LEN]; - tw[0] = PROTOCOL_DOMAIN_SEP; - tw[1] = t; - tw[2] = lay as u8; - tw[4..8].copy_from_slice(&p.to_le_bytes()); - tw[8..12].copy_from_slice(&tau.to_le_bytes()); - tw[12..16].copy_from_slice(&j.to_le_bytes()); - tw -} - -/// `Th` over a byte payload: an encoding input or a derived secret. -pub fn th(pp: &PublicParam, tw: &Tweak, payload: &[u8]) -> Digest { - let mut hasher = primitives::hash::Hasher::new(); - hasher.update(tw).update(pp).update(payload); - hasher.finalize()[..N].try_into().unwrap() -} - -/// `Th` over a concatenation of digests: a Merkle node, a one-time leaf, or the -/// few-time roots. -pub fn th_digests(pp: &PublicParam, tw: &Tweak, values: &[Digest]) -> Digest { - let mut hasher = primitives::hash::Hasher::new(); - hasher.update(tw).update(pp); - for value in values { - hasher.update(value); - } - hasher.finalize()[..N].try_into().unwrap() -} diff --git a/crates/sphincs/src/lib.rs b/crates/sphincs/src/lib.rs deleted file mode 100644 index 5186138b5..000000000 --- a/crates/sphincs/src/lib.rs +++ /dev/null @@ -1,102 +0,0 @@ -//! SPHINCS+ over BLAKE2s: the stateless scheme specified in -//! `doc/sphincs/main.tex`, with WOTS+C and FORS+C, at `2^24` signatures per key -//! pair. A public key is 32 bytes, a signature 4924, and a verification 497 hash -//! calls. -//! -//! That specification is the reference and every symbol here carries its name: -//! `n`, `w`, `v`, `T`, `d`, `h_lay`, `a`, `k`. `Th` is standard BLAKE2s of -//! the exact byte string `tweak | P | payload` truncated to `n = 128` bits (the -//! `hash` module), and the tweak names one hash call in the whole structure. -//! -//! One 32-byte master seed derives the public parameter and all signing secrets. -//! The signer caches public nodes of the top tree. - -#![cfg_attr(not(test), warn(unused_crate_dependencies))] - -mod hash; -pub use hash::*; -mod ots; -pub use ots::*; -mod fts; -pub use fts::*; -mod sphincs; -pub use sphincs::*; - -/// `n`: hash value and Merkle node length, in bytes. -pub const N: usize = 16; -pub type Digest = [u8; N]; - -/// The public parameter derived from the master seed. -pub const PUBLIC_PARAM_LEN: usize = 16; -pub type PublicParam = [u8; PUBLIC_PARAM_LEN]; - -/// The master secret used to derive all WOTS and FORS secrets. -pub const MASTER_SECRET_LEN: usize = 32; -pub type MasterSecret = [u8; MASTER_SECRET_LEN]; - -/// The per-signature randomizer the message digest is computed under. -pub const RANDOMIZER_LEN: usize = 16; -pub type Randomizer = [u8; RANDOMIZER_LEN]; - -/// The message to sign (a 256-bit message hash). -pub const MESSAGE_LEN: usize = 32; -pub type Message = [u8; MESSAGE_LEN]; - -/// The serialized width of an encoding counter. -pub const COUNTER_LEN: usize = 4; - -// The one-time signature. -/// `w`: chunk size in bits. -pub const W: usize = 3; -/// `2^w`: one more than the steps of a hash chain. -pub const CHAIN_LEN: usize = 1 << W; -/// `v`: code length, one hash chain per chunk. -pub const V: usize = 42; -/// `T`: the sum every codeword has. Above the mean `v(2^w-1)/2 = 147`, so -/// verification walks fewer chain steps and the signer grinds a counter for it. -pub const TARGET_SUM: usize = 191; - -// The hypertree. -/// `d`: hypertree layers, numbered from the top. -pub const D: usize = 3; -/// `h_lay`: the Merkle tree height of each layer. -pub const HEIGHTS: [usize; D] = [12, 7, 7]; -/// `h`: total height, so `2^h` few-time keys. -pub const H: usize = 26; - -// The few-time signature. -/// `a`: log2 of the leaves in one few-time tree. -pub const A: usize = 10; -/// `k`: digest index groups. -pub const K: usize = 15; -/// The forest holds `k-1` trees: the tree of the last digest index carries no -/// information, that index being ground to zero (FORS$^+$C). -pub const NUM_FTS_TREES: usize = K - 1; - -/// `A_max`: digest attempts per signature. -pub const MAX_DIGEST_ATTEMPTS: u64 = 1 << 32; -/// `C_max`: encoding attempts per layer. -pub const MAX_ENCODING_ATTEMPTS: u64 = 1 << 32; - -/// `h + ka`: the message digest's width, all of it consumed by the index and the -/// `k` leaf indices. -pub const DIGEST_BITS: usize = H + K * A; -pub const DIGEST_BYTES: usize = DIGEST_BITS / 8; - -pub const PUB_KEY_SIZE: usize = N + PUBLIC_PARAM_LEN; -/// A secret key is its public parameter and its master secret; the rest is derived. -pub const SECRET_KEY_SIZE: usize = PUBLIC_PARAM_LEN + MASTER_SECRET_LEN; -pub const SIG_SIZE: usize = RANDOMIZER_LEN + NUM_FTS_TREES * (1 + A) * N + D * (COUNTER_LEN + V * N) + H * N; - -/// Calls to the hash function one verification makes: the digest, `Fts.recover`, -/// `d` times `Ots.leaf`, and `Tree.fold`. -pub const VERIFY_HASHES: usize = 1 + (NUM_FTS_TREES * (1 + A) + 1) + D * (V * (CHAIN_LEN - 1) - TARGET_SUM + 2) + H; - -const _: () = assert!(H == HEIGHTS[0] + HEIGHTS[1] + HEIGHTS[2]); -// Each half of an encoding digest holds `v/2` chunks and one pinned bit. -const _: () = assert!(W * V / 2 + 1 == 64); -const _: () = assert!(DIGEST_BITS == DIGEST_BYTES * 8); -const _: () = assert!(TARGET_SUM < V * (CHAIN_LEN - 1)); -const _: () = assert!(PUB_KEY_SIZE == 32); -const _: () = assert!(SIG_SIZE == 4924); -const _: () = assert!(VERIFY_HASHES == 497); diff --git a/crates/sphincs/src/ots.rs b/crates/sphincs/src/ots.rs deleted file mode 100644 index b9c345c48..000000000 --- a/crates/sphincs/src/ots.rs +++ /dev/null @@ -1,103 +0,0 @@ -//! The one-time signature: `v` hash chains of `2^w - 1` steps, and the -//! target-sum code that replaces the Winternitz checksum (WOTS+C). -//! -//! A codeword is `v` chunks summing to `T`. Two distinct words of equal sum -//! cannot be ordered componentwise, so revealing chain position `x_i` on every -//! chain gives a forger nothing: any other codeword needs a value above one of -//! the revealed ones. The price is that most messages do not encode into the -//! code at all, hence the counter the signer searches for and the signature -//! carries. - -use crate::*; - -/// One one-time key's position: the layer, the tree within it, the leaf within -/// that tree. -#[derive(Clone, Copy, Debug, PartialEq, Eq)] -pub struct Pos { - pub lay: usize, - pub tau: u32, - pub e: u32, -} - -impl Pos { - pub const fn new(lay: usize, tau: u32, e: u32) -> Self { - Self { lay, tau, e } - } -} - -/// `sk_{lay,tau,e,i} = Th(P, tw_prf(lay,tau,i,e), S)`. -pub fn ots_secret(pp: &PublicParam, master: &MasterSecret, pos: Pos, i: usize) -> Digest { - th(pp, &tweak(TWEAK_PRF, pos.lay, pos.tau, i as u32, pos.e), master) -} - -/// `Chain_{lay,tau,e,i}(P, start, steps, value)`: the step onto position `s` is -/// hashed under the tweak of the edge into it. -pub fn chain(pp: &PublicParam, pos: Pos, i: usize, start: usize, steps: usize, value: Digest) -> Digest { - debug_assert!(start + steps < CHAIN_LEN); - (1..=steps).fold(value, |current, step| { - let p = (CHAIN_LEN * i + start + step - 1) as u32; - th(pp, &tweak(TWEAK_CHAIN, pos.lay, pos.tau, p, pos.e), ¤t) - }) -} - -/// The Merkle leaf of a one-time key: `Th` over its `v` chain tips. -pub fn ots_leaf_hash(pp: &PublicParam, pos: Pos, tips: &[Digest; V]) -> Digest { - th_digests(pp, &tweak(TWEAK_LEAF, pos.lay, pos.tau, 0, pos.e), tips) -} - -/// `Enc(P, lay, tau, e, M, c)`: the codeword, or `None` if the digest of that -/// counter is not admissible. -pub fn encode(pp: &PublicParam, pos: Pos, m: &Digest, c: u32) -> Option<[u8; V]> { - let mut payload = [0u8; N + COUNTER_LEN]; - payload[..N].copy_from_slice(m); - payload[N..].copy_from_slice(&c.to_le_bytes()); - codeword(&th(pp, &tweak(TWEAK_ENC, pos.lay, pos.tau, 0, pos.e), &payload)) -} - -/// Each 64-bit half of the digest holds `v/2` chunks of `w` bits and one pinned -/// top bit; pinning it is what makes the codeword determine the digest. -fn codeword(digest: &Digest) -> Option<[u8; V]> { - let mut x = [0u8; V]; - let mut sum = 0; - for (q, half) in digest.as_chunks::<{ N / 2 }>().0.iter().enumerate() { - let d = u64::from_le_bytes(*half); - if d >> (W * V / 2) != 0 { - return None; - } - for r in 0..V / 2 { - let chunk = ((d >> (W * r)) & (CHAIN_LEN as u64 - 1)) as u8; - x[q * (V / 2) + r] = chunk; - sum += chunk as usize; - } - } - (sum == TARGET_SUM).then_some(x) -} - -/// `Ots.sign`: the LEAST admissible counter, and the chain value each chunk -/// opens. Deterministic in its inputs, which is what keeps one key to one -/// codeword: a resumed or randomized search would leak two incomparable -/// codewords and drop forgery to about `2^53`. -pub fn ots_sign(pp: &PublicParam, master: &MasterSecret, pos: Pos, m: &Digest) -> Option<(u32, [Digest; V])> { - let (c, x) = (0..MAX_ENCODING_ATTEMPTS).find_map(|c| encode(pp, pos, m, c as u32).map(|x| (c as u32, x)))?; - let signature = std::array::from_fn(|i| chain(pp, pos, i, 0, x[i] as usize, ots_secret(pp, master, pos, i))); - Some((c, signature)) -} - -/// `Ots.leaf`: the leaf a claimed signature recovers, or `None` if its counter -/// is not admissible for `m`. Does not touch the secrets, which is why it is the -/// verifier's half. -pub fn ots_leaf(pp: &PublicParam, pos: Pos, m: &Digest, c: u32, signature: &[Digest; V]) -> Option { - let x = encode(pp, pos, m, c)?; - let tips = std::array::from_fn(|i| { - let start = x[i] as usize; - chain(pp, pos, i, start, CHAIN_LEN - 1 - start, signature[i]) - }); - Some(ots_leaf_hash(pp, pos, &tips)) -} - -/// The leaf of the one-time key at `pos`, from the master secret: what key -/// generation and every tree rebuild spend their hashes on. -pub fn ots_public_leaf(pp: &PublicParam, master: &MasterSecret, pos: Pos) -> Digest { - let tips = std::array::from_fn(|i| chain(pp, pos, i, 0, CHAIN_LEN - 1, ots_secret(pp, master, pos, i))); - ots_leaf_hash(pp, pos, &tips) -} diff --git a/crates/sphincs/src/sphincs.rs b/crates/sphincs/src/sphincs.rs deleted file mode 100644 index 3b999e524..000000000 --- a/crates/sphincs/src/sphincs.rs +++ /dev/null @@ -1,468 +0,0 @@ -//! The hypertree and the three algorithms: `d` layers of Merkle trees over -//! one-time leaves, the bottom layer signing few-time keys, layer 0's root being -//! the public key. -//! -//! An index derived from the message digest says which few-time key signs, and -//! with it which tree and which leaf are used on every layer. Nothing is -//! reserved and nothing is spent: a key answers for all `2^h` indices, which is -//! what makes the scheme stateless. - -use rand::{CryptoRng, Rng}; -use serde::{Deserialize, Serialize}; - -use crate::*; - -/// `SUFFIX[lay] = sum_{j >= lay} h_j`, the height of everything at or below -/// layer `lay`: the divisors of the index decomposition. -const fn suffix_heights() -> [usize; D + 1] { - let mut suffix = [0; D + 1]; - let mut lay = D; - while lay > 0 { - lay -= 1; - suffix[lay] = suffix[lay + 1] + HEIGHTS[lay]; - } - suffix -} -pub const SUFFIX: [usize; D + 1] = suffix_heights(); -const _: () = assert!(SUFFIX[0] == H); - -/// The layer-0 depth whose nodes a signer caches, halfway up so that the subtree -/// to rebuild and the nodes to refold are both `2^(h_0/2)`. -pub const SPLIT_LEVEL: usize = HEIGHTS[0].div_ceil(2); -pub const CACHE_LEN: usize = 1 << (HEIGHTS[0] - SPLIT_LEVEL); -const _: () = assert!(CACHE_LEN * N == 1024); - -/// `tau_lay(idx)`: the tree used on layer `lay`. -pub fn tree_of(idx: u64, lay: usize) -> u32 { - (idx >> SUFFIX[lay]) as u32 -} - -/// `e_lay(idx)`: the leaf used within that tree. -pub fn leaf_of(idx: u64, lay: usize) -> u32 { - ((idx >> SUFFIX[lay + 1]) & ((1 << HEIGHTS[lay]) - 1)) as u32 -} - -/// Where layer `lay`'s siblings sit in a signature's flat path. -pub fn path_range(lay: usize) -> std::ops::Range { - let start: usize = HEIGHTS[..lay].iter().sum(); - start..start + HEIGHTS[lay] -} - -/// Ordered lexicographically on [`Self::flatten`], which is what an aggregate's -/// signer list is sorted and deduplicated by. -#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash, PartialOrd, Ord, Serialize, Deserialize)] -pub struct SphincsPublicKey { - pub root: Digest, - pub public_param: PublicParam, -} - -impl SphincsPublicKey { - pub fn flatten(&self) -> [u8; PUB_KEY_SIZE] { - let mut out = [0; PUB_KEY_SIZE]; - out[..N].copy_from_slice(&self.root); - out[N..].copy_from_slice(&self.public_param); - out - } - - pub fn from_bytes(bytes: &[u8; PUB_KEY_SIZE]) -> Self { - Self { - root: bytes[..N].try_into().unwrap(), - public_param: bytes[N..].try_into().unwrap(), - } - } -} - -/// `P`, the root, and the master secret every secret is derived from, plus -/// layer 0's nodes at [`SPLIT_LEVEL`]. Those nodes are a cache and not state: a -/// deterministic function of the master secret, so losing them costs -/// recomputation and nothing else. -#[derive(Clone, Debug)] -pub struct SphincsSecretKey { - pub public_param: PublicParam, - pub root: Digest, - master: MasterSecret, - cache: [Digest; CACHE_LEN], -} - -impl SphincsSecretKey { - /// SECRET KEY MATERIAL: public parameter and master secret. - pub fn to_bytes(&self) -> [u8; SECRET_KEY_SIZE] { - let mut out = [0; SECRET_KEY_SIZE]; - out[..PUBLIC_PARAM_LEN].copy_from_slice(&self.public_param); - out[PUBLIC_PARAM_LEN..].copy_from_slice(&self.master); - out - } - - /// Inverse of [`Self::to_bytes`], costing what [`key_gen_from`] costs: the - /// layer-0 tree is rebuilt rather than stored. - pub fn from_bytes(bytes: &[u8; SECRET_KEY_SIZE]) -> Self { - let (public_param, master) = bytes.split_at(PUBLIC_PARAM_LEN); - key_gen_from(public_param.try_into().unwrap(), master.try_into().unwrap()).0 - } -} - -#[derive(Clone, Debug, PartialEq, Eq)] -pub struct SphincsSignature { - pub randomizer: Randomizer, - pub fts: FtsOpening, - pub counters: [u32; D], - pub ots: [[Digest; V]; D], - /// Layer 0's `h_0` siblings, then layer 1's, then layer 2's. - pub paths: [Digest; H], -} - -impl SphincsSignature { - /// The specification's serialization, exactly [`SIG_SIZE`] bytes. - pub fn to_bytes(&self) -> [u8; SIG_SIZE] { - let mut out = [0; SIG_SIZE]; - let mut at = 0; - let mut put = |bytes: &[u8]| { - out[at..at + bytes.len()].copy_from_slice(bytes); - at += bytes.len(); - }; - put(&self.randomizer); - for kappa in 0..NUM_FTS_TREES { - put(&self.fts.secrets[kappa]); - for sibling in &self.fts.paths[kappa] { - put(sibling); - } - } - for lay in 0..D { - put(&self.counters[lay].to_le_bytes()); - for value in &self.ots[lay] { - put(value); - } - for sibling in &self.paths[path_range(lay)] { - put(sibling); - } - } - debug_assert_eq!(at, SIG_SIZE); - out - } - - pub fn from_bytes(bytes: &[u8; SIG_SIZE]) -> Self { - let mut at = 0; - let mut take = |len: usize| { - at += len; - &bytes[at - len..at] - }; - let randomizer = take(RANDOMIZER_LEN).try_into().unwrap(); - let mut fts = FtsOpening { - secrets: [[0; N]; NUM_FTS_TREES], - paths: [[[0; N]; A]; NUM_FTS_TREES], - }; - for kappa in 0..NUM_FTS_TREES { - fts.secrets[kappa] = take(N).try_into().unwrap(); - for level in 0..A { - fts.paths[kappa][level] = take(N).try_into().unwrap(); - } - } - let mut counters = [0; D]; - let mut ots = [[[0; N]; V]; D]; - let mut paths = [[0; N]; H]; - for lay in 0..D { - counters[lay] = u32::from_le_bytes(take(COUNTER_LEN).try_into().unwrap()); - for i in 0..V { - ots[lay][i] = take(N).try_into().unwrap(); - } - for level in path_range(lay) { - paths[level] = take(N).try_into().unwrap(); - } - } - debug_assert_eq!(at, SIG_SIZE); - Self { - randomizer, - fts, - counters, - ots, - paths, - } - } -} - -#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash)] -pub enum SphincsSignError { - /// `A_max` digests in a row had a nonzero last index. - NoAdmissibleDigest, - /// `C_max` counters in a row failed to encode. - NoAdmissibleEncoding, -} - -#[derive(Clone, Copy, Debug, PartialEq, Eq, Hash)] -pub enum SphincsVerifyError { - /// The digest's last index is not zero. - InadmissibleDigest, - /// A layer's counter does not encode the message it signs. - InadmissibleEncoding, - RootMismatch, -} - -impl std::fmt::Display for SphincsSignError { - fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { - match self { - Self::NoAdmissibleDigest => write!(f, "no admissible message digest within A_max attempts"), - Self::NoAdmissibleEncoding => write!(f, "no admissible encoding within C_max attempts"), - } - } -} - -impl std::error::Error for SphincsSignError {} - -impl std::fmt::Display for SphincsVerifyError { - fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { - match self { - Self::InadmissibleDigest => write!(f, "the digest's last index is not zero"), - Self::InadmissibleEncoding => write!(f, "a layer's counter does not encode the message it signs"), - Self::RootMismatch => write!(f, "the hypertree walk does not reach the key's root"), - } - } -} - -impl std::error::Error for SphincsVerifyError {} - -/// The message digest, read as the index and the `k` leaf indices. `h + ka` bits -/// of a random oracle output, so the index and the last leaf index are disjoint -/// and grinding one does not bias the other. -pub fn message_digest(pp: &PublicParam, root: &Digest, rho: &Randomizer, m: &Message) -> (u64, [u32; K]) { - let mut hasher = primitives::hash::Hasher::new(); - hasher - .update(&tweak(TWEAK_MSG, 0, 0, 0, 0)) - .update(pp) - .update(rho) - .update(root) - .update(m); - let digest = &hasher.finalize()[..DIGEST_BYTES]; - let field = |offset: usize, len: usize| { - (0..len).fold(0u64, |value, bit| { - let position = offset + bit; - value | (u64::from(digest[position / 8] >> (position % 8) & 1) << bit) - }) - }; - (field(0, H), std::array::from_fn(|kappa| field(H + kappa * A, A) as u32)) -} - -fn node(pp: &PublicParam, lay: usize, tau: u32, level: usize, j: u64, left: &Digest, right: &Digest) -> Digest { - let tw = tweak(TWEAK_NODE, lay, tau, level as u32, j as u32); - th_digests(pp, &tw, &[*left, *right]) -} - -/// Merkle levels `from_level..=to_level` of layer `lay`'s tree `tau`, given -/// `bottom`, the complete band of level-`from_level` nodes starting at index -/// `first`. Level `l` is `layers[l - from_level]`. -fn build_up( - pp: &PublicParam, - lay: usize, - tau: u32, - bottom: Vec, - from_level: usize, - to_level: usize, - first: u64, -) -> Vec> { - let mut layers = vec![bottom]; - for level in from_level + 1..=to_level { - let base = first >> (level - from_level); - let children = layers.last().unwrap(); - layers.push( - (0..children.len() / 2) - .map(|j| { - node( - pp, - lay, - tau, - level, - base + j as u64, - &children[2 * j], - &children[2 * j + 1], - ) - }) - .collect(), - ); - } - layers -} - -/// `Gen`, on given `P` and master secret. Only layer 0 is built; the trees below -/// it are built when a signature needs them. -pub fn key_gen_from(public_param: PublicParam, master: MasterSecret) -> (SphincsSecretKey, SphincsPublicKey) { - let leaves = parallel::map_collect(1 << HEIGHTS[0], |e| { - ots_public_leaf(&public_param, &master, Pos::new(0, 0, e as u32)) - }); - let layers = build_up(&public_param, 0, 0, leaves, 0, HEIGHTS[0], 0); - let root = layers[HEIGHTS[0]][0]; - let cache = std::array::from_fn(|i| layers[SPLIT_LEVEL][i]); - ( - SphincsSecretKey { - public_param, - root, - master, - cache, - }, - SphincsPublicKey { root, public_param }, - ) -} - -/// `Gen`, on a fresh key: the seed comes from `rng`, so nothing can regenerate -/// the key. -pub fn key_gen(rng: &mut impl CryptoRng) -> (SphincsSecretKey, SphincsPublicKey) { - key_gen_from_seed(rng.random()) -} - -/// Deterministic [`key_gen`]: the seed is the master secret, and a dedicated -/// tweak derives the public parameter from it. -pub fn key_gen_from_seed(seed: MasterSecret) -> (SphincsSecretKey, SphincsPublicKey) { - let parameter = th(&[0; PUBLIC_PARAM_LEN], &tweak(TWEAK_PARAMETER, 0, 0, 0, 0), &seed); - key_gen_from(parameter, seed) -} - -impl SphincsSecretKey { - pub fn public_key(&self) -> SphincsPublicKey { - SphincsPublicKey { - root: self.root, - public_param: self.public_param, - } - } - - /// Layer `lay`'s tree `tau` rebuilt whole: the siblings at `e` into `path`, - /// and the root. - fn tree_path_and_root(&self, lay: usize, tau: u32, e: u32, path: &mut [Digest]) -> Digest { - debug_assert_eq!(path.len(), HEIGHTS[lay]); - let leaves = (0..1 << HEIGHTS[lay]) - .map(|leaf| ots_public_leaf(&self.public_param, &self.master, Pos::new(lay, tau, leaf))) - .collect(); - let layers = build_up(&self.public_param, lay, tau, leaves, 0, HEIGHTS[lay], 0); - for (level, sibling) in path.iter_mut().enumerate() { - *sibling = layers[level][((e >> level) ^ 1) as usize]; - } - layers[HEIGHTS[lay]][0] - } - - /// Layer 0's siblings at `e`, from the cache: one `2^SPLIT_LEVEL`-leaf - /// subtree rebuilt below it, the cached nodes refolded above it. Returns the - /// root, which the cache reproduces. - fn cached_path_and_root(&self, e: u32, path: &mut [Digest]) -> Digest { - debug_assert_eq!(path.len(), HEIGHTS[0]); - let first = u64::from(e >> SPLIT_LEVEL) << SPLIT_LEVEL; - let leaves = (first..first + (1 << SPLIT_LEVEL)) - .map(|leaf| ots_public_leaf(&self.public_param, &self.master, Pos::new(0, 0, leaf as u32))) - .collect(); - let below = build_up(&self.public_param, 0, 0, leaves, 0, SPLIT_LEVEL, first); - let above = build_up( - &self.public_param, - 0, - 0, - self.cache.to_vec(), - SPLIT_LEVEL, - HEIGHTS[0], - 0, - ); - debug_assert_eq!(below[SPLIT_LEVEL][0], self.cache[(first >> SPLIT_LEVEL) as usize]); - for (level, sibling) in path.iter_mut().enumerate() { - let index = u64::from(e >> level) ^ 1; - *sibling = if level < SPLIT_LEVEL { - below[level][(index - (first >> level)) as usize] - } else { - above[level - SPLIT_LEVEL][index as usize] - }; - } - above[HEIGHTS[0] - SPLIT_LEVEL][0] - } -} - -/// Sign at most `2^24` messages per key. -/// Signing is deterministic and stateless. -pub fn sign(sk: &SphincsSecretKey, message: &Message) -> Result { - // The digest is admissible when its last leaf index is zero, which is what - // drops that tree from the forest; it takes 2^a attempts on average. - let (randomizer, idx, u) = (0..MAX_DIGEST_ATTEMPTS) - .find_map(|trial| { - let mut hasher = primitives::hash::Hasher::new(); - hasher.update(&tweak(TWEAK_RANDOMIZER, 0, 0, trial as u32, 0)); - hasher.update(&sk.public_param).update(&sk.master).update(message); - let randomizer = hasher.finalize()[..RANDOMIZER_LEN].try_into().unwrap(); - let (idx, u) = message_digest(&sk.public_param, &sk.root, &randomizer, message); - (u[K - 1] == 0).then_some((randomizer, idx, u)) - }) - .ok_or(SphincsSignError::NoAdmissibleDigest)?; - - let (fts_key, fts) = fts_open(&sk.public_param, &sk.master, idx, &u); - - let mut message_of_layer = fts_key; - let mut counters = [0; D]; - let mut ots = [[[0; N]; V]; D]; - let mut paths = [[0; N]; H]; - for lay in (0..D).rev() { - let (tau, e) = (tree_of(idx, lay), leaf_of(idx, lay)); - let pos = Pos::new(lay, tau, e); - let (c, signature) = ots_sign(&sk.public_param, &sk.master, pos, &message_of_layer) - .ok_or(SphincsSignError::NoAdmissibleEncoding)?; - counters[lay] = c; - ots[lay] = signature; - let path = &mut paths[path_range(lay)]; - message_of_layer = if lay == 0 { - sk.cached_path_and_root(e, path) - } else { - sk.tree_path_and_root(lay, tau, e, path) - }; - } - // Layer 0's root is discarded: it is the public key's whenever the signer is - // honest, which is also the only check the cache gets. - debug_assert_eq!(message_of_layer, sk.root); - - Ok(SphincsSignature { - randomizer, - fts, - counters, - ots, - paths, - }) -} - -/// `Tree.fold`: the other half of a Merkle opening. -pub fn tree_fold(pp: &PublicParam, pos: Pos, leaf: Digest, path: &[Digest]) -> Digest { - path.iter().enumerate().fold(leaf, |current, (level, sibling)| { - let (left, right) = if (pos.e >> level) & 1 == 0 { - (current, *sibling) - } else { - (*sibling, current) - }; - node( - pp, - pos.lay, - pos.tau, - level + 1, - u64::from(pos.e >> (level + 1)), - &left, - &right, - ) - }) -} - -/// `Ver`. -pub fn verify( - pk: &SphincsPublicKey, - message: &Message, - signature: &SphincsSignature, -) -> Result<(), SphincsVerifyError> { - let (idx, u) = message_digest(&pk.public_param, &pk.root, &signature.randomizer, message); - if u[K - 1] != 0 { - return Err(SphincsVerifyError::InadmissibleDigest); - } - let mut message_of_layer = fts_recover(&pk.public_param, idx, &u, &signature.fts); - for lay in (0..D).rev() { - let pos = Pos::new(lay, tree_of(idx, lay), leaf_of(idx, lay)); - let leaf = ots_leaf( - &pk.public_param, - pos, - &message_of_layer, - signature.counters[lay], - &signature.ots[lay], - ) - .ok_or(SphincsVerifyError::InadmissibleEncoding)?; - message_of_layer = tree_fold(&pk.public_param, pos, leaf, &signature.paths[path_range(lay)]); - } - if message_of_layer == pk.root { - Ok(()) - } else { - Err(SphincsVerifyError::RootMismatch) - } -} diff --git a/crates/sphincs/tests/sphincs_tests.rs b/crates/sphincs/tests/sphincs_tests.rs deleted file mode 100644 index 4317c96b4..000000000 --- a/crates/sphincs/tests/sphincs_tests.rs +++ /dev/null @@ -1,223 +0,0 @@ -use rand::{Rng, SeedableRng, rngs::StdRng}; -use sphincs::*; - -fn test_message() -> Message { - std::array::from_fn(|i| (i * 5 + 3) as u8) -} - -fn test_key(seed: u64) -> (SphincsSecretKey, SphincsPublicKey) { - key_gen(&mut StdRng::seed_from_u64(seed)) -} - -#[test] -fn keygen_sign_verify() { - let (sk, pk) = test_key(0); - assert_eq!(sk.public_key(), pk); - let message = test_message(); - let signature = sign(&sk, &message).unwrap(); - verify(&pk, &message, &signature).unwrap(); - assert_eq!(sign(&sk, &message).unwrap(), signature); -} - -#[test] -fn serialized_sizes_and_roundtrip() { - let (sk, pk) = test_key(1); - let message = test_message(); - let signature = sign(&sk, &message).unwrap(); - - let public_key_bytes = pk.flatten(); - assert_eq!(public_key_bytes.len(), 32); - assert_eq!(SphincsPublicKey::from_bytes(&public_key_bytes), pk); - - let signature_bytes = signature.to_bytes(); - assert_eq!(signature_bytes.len(), 4924); - let decoded = SphincsSignature::from_bytes(&signature_bytes); - assert_eq!(decoded, signature); - verify(&pk, &message, &decoded).unwrap(); -} - -#[test] -fn tampered_signatures_rejected() { - let (sk, pk) = test_key(2); - let message = test_message(); - let signature = sign(&sk, &message).unwrap(); - verify(&pk, &message, &signature).unwrap(); - - let mut other_message = message; - other_message[0] ^= 1; - assert!(verify(&pk, &other_message, &signature).is_err()); - - let mut other_key = pk; - other_key.root[0] ^= 1; - assert!(verify(&other_key, &message, &signature).is_err()); - - // Verification recomputes the digest, so a tampered randomizer asks for - // another index, and asks it of a digest that is admissible only one time in - // 2^a. - let mut tampered = signature.clone(); - tampered.randomizer[0] ^= 1; - assert_eq!( - verify(&pk, &message, &tampered), - Err(SphincsVerifyError::InadmissibleDigest) - ); - - // Everything the bottom layers carry feeds the message a layer above signs, - // and a counter is admissible for one message in 2^13.6, so tampering - // surfaces as an inadmissible encoding rather than as a wrong root. - for tamper in [ - (|s: &mut SphincsSignature| s.fts.secrets[5][0] ^= 1) as fn(&mut SphincsSignature), - |s: &mut SphincsSignature| s.fts.paths[9][4][0] ^= 1, - |s: &mut SphincsSignature| s.counters[2] ^= 1, - |s: &mut SphincsSignature| s.ots[1][17][0] ^= 1, - |s: &mut SphincsSignature| s.paths[H - 1][0] ^= 1, - ] { - let mut tampered = signature.clone(); - tamper(&mut tampered); - assert_eq!( - verify(&pk, &message, &tampered), - Err(SphincsVerifyError::InadmissibleEncoding) - ); - } - - // Layer 0's path is the exception: nothing is signed above it, so it can - // only fail the root comparison. - let mut tampered = signature.clone(); - tampered.paths[0][0] ^= 1; - assert_eq!(verify(&pk, &message, &tampered), Err(SphincsVerifyError::RootMismatch)); - - let mut tampered = signature.clone(); - tampered.ots[0][17][0] ^= 1; - assert_eq!(verify(&pk, &message, &tampered), Err(SphincsVerifyError::RootMismatch)); -} - -/// One key signs one codeword, on which the whole one-time argument rests: the -/// counter is the least admissible one, not any admissible one. -#[test] -fn ots_counter_is_the_least_admissible() { - let mut rng = StdRng::seed_from_u64(4); - let public_param: PublicParam = rng.random(); - let master: MasterSecret = rng.random(); - let pos = Pos::new(2, 1234, 56); - let message: Digest = rng.random(); - - let (counter, signature) = ots_sign(&public_param, &master, pos, &message).unwrap(); - assert!((0..counter).all(|c| encode(&public_param, pos, &message, c).is_none())); - assert_eq!( - ots_leaf(&public_param, pos, &message, counter, &signature), - Some(ots_public_leaf(&public_param, &master, pos)) - ); -} - -#[test] -fn index_decomposition_is_a_bijection_onto_the_bottom_layer() { - let mut rng = StdRng::seed_from_u64(5); - for _ in 0..1000 { - let idx = rng.random::() % (1 << H); - // Every layer's tree is the one whose root sits at the leaf its parent - // layer uses. - for lay in 1..D { - let expected = - u64::from(tree_of(idx, lay - 1)) * (1 << HEIGHTS[lay - 1]) + u64::from(leaf_of(idx, lay - 1)); - assert_eq!(u64::from(tree_of(idx, lay)), expected); - } - assert_eq!(tree_of(idx, 0), 0); - assert_eq!( - u64::from(tree_of(idx, D - 1)) * (1 << HEIGHTS[D - 1]) + u64::from(leaf_of(idx, D - 1)), - idx - ); - } -} - -/// The counter search and the digest resampling are the signer's two grinding -/// loops; both costs are a property of the predicates, so a drift here is a -/// change of scheme. -#[test] -#[ignore] -fn grinding_bits() { - let mut rng = StdRng::seed_from_u64(6); - let public_param: PublicParam = rng.random(); - let master: MasterSecret = rng.random(); - - let samples = 200; - let counters: u64 = (0..samples) - .map(|i| { - let message: Digest = rng.random(); - let pos = Pos::new(i % D, i as u32, i as u32); - u64::from(ots_sign(&public_param, &master, pos, &message).unwrap().0) - }) - .sum(); - // A codeword is one admissible digest, so 1/p is the number of them over - // 2^128: 2^13.60 for T = 191. - let encoding_bits = ((counters as f64 / samples as f64) + 1.0).log2(); - println!("counter search: 2^{encoding_bits:.2} attempts"); - assert!( - (12.6..14.6).contains(&encoding_bits), - "encoding cost moved: {encoding_bits:.2} bits" - ); - - let root: Digest = rng.random(); - let message = test_message(); - let mut attempts = 0u64; - for _ in 0..samples { - loop { - attempts += 1; - let randomizer: Randomizer = rng.random(); - if message_digest(&public_param, &root, &randomizer, &message).1[K - 1] == 0 { - break; - } - } - } - let digest_bits = (attempts as f64 / samples as f64).log2(); - println!("digest resampling: 2^{digest_bits:.2} attempts"); - assert!( - ((A as f64 - 1.0)..(A as f64 + 1.0)).contains(&digest_bits), - "digest cost moved: {digest_bits:.2} bits" - ); -} - -/// A reloaded secret key must sign exactly as the original does: `key_gen` -/// samples the master secret itself, so these bytes are the only way back to a -/// key it produced. -#[test] -fn secret_key_survives_a_round_trip() { - let (sk, pk) = test_key(7); - let reloaded = SphincsSecretKey::from_bytes(&sk.to_bytes()); - - assert_eq!(reloaded.public_key(), pk); - let message = test_message(); - let sig = sign(&reloaded, &message).unwrap(); - verify(&pk, &message, &sig).unwrap(); -} - -#[test] -fn secret_derivation_uses_full_master() { - let pp = [3; PUBLIC_PARAM_LEN]; - let master = [7; MASTER_SECRET_LEN]; - let pos = Pos::new(2, 5, 6); - let ots = ots_secret(&pp, &master, pos, 4); - let (fts, _) = fts_open(&pp, &master, 5, &[0; K]); - for byte in 0..MASTER_SECRET_LEN { - let mut changed = master; - changed[byte] ^= 1; - assert_ne!(ots_secret(&pp, &changed, pos, 4), ots); - assert_ne!(fts_open(&pp, &changed, 5, &[0; K]).0, fts); - } -} - -/// The split between the two entry points: the seed alone determines the key, -/// and the rng one draws a fresh seed per call rather than a fixed one. -#[test] -fn key_gen_entry_points() { - let seed = [17u8; 32]; - let (_, from_seed) = key_gen_from_seed(seed); - assert_eq!(key_gen_from_seed(seed).1, from_seed); - - let mut rng = StdRng::seed_from_u64(31); - let (_, first) = key_gen(&mut rng); - let (_, second) = key_gen(&mut rng); - assert_ne!(first, second); - - // The rng path is exactly a drawn seed, so it reproduces under a fixed rng. - let drawn: [u8; 32] = StdRng::seed_from_u64(31).random(); - assert_eq!(key_gen_from_seed(drawn).1, first); -} diff --git a/crates/xmss/Cargo.toml b/crates/xmss/Cargo.toml deleted file mode 100644 index d0d15a71c..000000000 --- a/crates/xmss/Cargo.toml +++ /dev/null @@ -1,18 +0,0 @@ -[package] -name = "xmss" -version.workspace = true -edition.workspace = true -publish = false - -[lints] -workspace = true - -[dependencies] -primitives.workspace = true -parallel.workspace = true -ethereum_ssz.workspace = true -rand.workspace = true -serde.workspace = true - -[dev-dependencies] -bincode.workspace = true diff --git a/crates/xmss/src/hash.rs b/crates/xmss/src/hash.rs deleted file mode 100644 index ffaeae847..000000000 --- a/crates/xmss/src/hash.rs +++ /dev/null @@ -1,56 +0,0 @@ -//! The XMSS hash layer: [`tweak_hash`] is standard BLAKE2s of the exact byte -//! string `tweak | pp | payload`, for chain steps, Merkle nodes, WOTS public -//! keys, and message encodings alike. -//! -//! The 16-byte tweak makes every call site a distinct hash function -//! (multi-target separation, as in leanVM) and the public parameter separates -//! users. Standard BLAKE2s binds the exact payload length. -//! -//! Compression counts per call: chain step 1, Merkle node 1, message encoding -//! 2, WOTS public key 11. A full XMSS verification is a constant 144 -//! compressions: 2 (encoding) + 99 (chains, fixed by the target sum) + 11 -//! (tips) + 32 (Merkle path). - -use crate::*; - -pub const PROTOCOL_DOMAIN_SEP: u8 = 0; - -// Tweak types (byte 1). -pub const TWEAK_TYPE_PRF: u8 = 0; -pub const TWEAK_TYPE_CHAIN: u8 = 1; -pub const TWEAK_TYPE_WOTS_PK: u8 = 2; -pub const TWEAK_TYPE_MERKLE: u8 = 3; -pub const TWEAK_TYPE_ENCODING: u8 = 4; -pub const TWEAK_TYPE_PARAMETER: u8 = 10; -pub const TWEAK_TYPE_FILLER: u8 = 11; -pub const TWEAK_TYPE_RANDOMIZER: u8 = 12; - -pub const TWEAK_LEN: usize = 16; -pub type Tweak = [u8; TWEAK_LEN]; - -/// A full 32-byte BLAKE2s chaining value/output. -pub const STATE_LEN: usize = 32; - -/// `[protocol_domain_sep:1 | type:1 | layer:1 | zero:1 | p:4 | tree:4 | index:4]`, little endian. -/// XMSS sets `layer` and `tree` to zero. -/// `index` is the epoch (chain / wots_pk / encoding) or the Merkle node index; -/// `sub_position` is the chain position or the Merkle level. -pub fn make_tweak(tweak_type: u8, sub_position: u32, index: u32) -> Tweak { - let mut tweak = [0u8; TWEAK_LEN]; - tweak[0] = PROTOCOL_DOMAIN_SEP; - tweak[1] = tweak_type; - tweak[4..8].copy_from_slice(&sub_position.to_le_bytes()); - tweak[12..16].copy_from_slice(&index.to_le_bytes()); - tweak -} - -/// Standard BLAKE2s of the exact-length `tweak | pp | payload` byte string. One -/// compression for chain steps (48 bytes total) and Merkle nodes (64 bytes -/// total), more for the multi-block WOTS public-key and encoding inputs. -pub fn tweak_hash(pp: &PublicParam, tweak_type: u8, sub_position: u32, index: u32, payload: &[u8]) -> Digest { - let mut hasher = primitives::hash::Hasher::new(); - hasher.update(&make_tweak(tweak_type, sub_position, index)); - hasher.update(pp); - hasher.update(payload); - hasher.finalize()[..DIGEST_LEN].try_into().unwrap() -} diff --git a/crates/xmss/src/lib.rs b/crates/xmss/src/lib.rs deleted file mode 100644 index 4ca40dbb5..000000000 --- a/crates/xmss/src/lib.rs +++ /dev/null @@ -1,106 +0,0 @@ -//! XMSS over BLAKE2s, with byte-oriented keys and signatures. -//! The concrete scheme is defined in the [XMSS specification]. -//! -//! Every hash is standard BLAKE2s of the exact byte string -//! `tweak | pp | payload`. Randomizer derivation retains 192 bits; other calls -//! retain 128 bits. See the `hash` module for the constructions. -//! -//! [XMSS specification]: https://github.com/leanEthereum/leanVM/releases/download/doc-latest/XMSS.pdf - -#![cfg_attr(not(test), warn(unused_crate_dependencies))] - -mod hash; -pub use hash::*; -mod ssz_serialization; -pub use ssz_serialization::*; -mod wots; -pub use wots::*; -mod xmss; -pub use xmss::*; - -/// n = 128 bits. -pub const DIGEST_LEN: usize = 16; -pub type Digest = [u8; DIGEST_LEN]; -pub type PublicParam = [u8; PUBLIC_PARAM_LEN]; -pub type Randomness = [u8; RANDOMNESS_LEN]; -/// The message to sign (a 256-bit message hash). -pub type Message = [u8; MESSAGE_LEN]; - -// WOTS -pub const V: usize = 42; // number of hash chains -pub const W: usize = 3; -pub const CHAIN_LENGTH: usize = 1 << W; // 8 -/// Chain hashes the VERIFIER walks, summed over all chains: `sum(chain_length - -/// 1 - e_i)`. Constant because the encoding sum is fixed to [`TARGET_SUM`]. -pub const NUM_CHAIN_HASHES: usize = 99; -/// A WOTS encoding `(e_0, .., e_{v-1})` is valid iff every `e_i < CHAIN_LENGTH`, -/// `sum(e_i) = TARGET_SUM`, and the 2 leftover digest bits are zero (see -/// [`wots_encode`]). The signer grinds the randomness until the encoding is -/// valid (no checksum chains). 195 sits above the mean (147) so verification -/// walks fewer chain steps; grinding takes fewer than 2^15 encode attempts on -/// average. -pub const TARGET_SUM: usize = V * (CHAIN_LENGTH - 1) - NUM_CHAIN_HASHES; // 195 -/// Maximum randomizer trials per signature. -pub const MAX_RANDOMIZER_TRIALS: u64 = 1 << 32; -pub const RANDOMNESS_LEN: usize = 24; -pub const MESSAGE_LEN: usize = 32; -pub const PUBLIC_PARAM_LEN: usize = 16; - -// XMSS -/// Merkle tree height: a key is valid for up to `2^32` epochs. -pub const LOG_LIFETIME: usize = 32; - -/// When a signature was made. Each epoch in the key's range may sign only one message. -pub type Epoch = u32; - -/// Serialized sizes (exact under bincode: fixed arrays, no length prefixes). -pub const WOTS_SIG_SIZE: usize = RANDOMNESS_LEN + V * DIGEST_LEN; // 696 -pub const SIG_SIZE: usize = WOTS_SIG_SIZE + LOG_LIFETIME * DIGEST_LEN; // 1208 -pub const PUB_KEY_SIZE: usize = DIGEST_LEN + PUBLIC_PARAM_LEN; // 32 - -// The encoding uses v*w = 126 of the digest's 128 bits; the 2 leftover top -// bits are ground to zero, so the digest decomposes exactly into the chunks. -const _: () = assert!(V * W + 2 == DIGEST_LEN * 8); - -/// Serde for `[T; N]` with N > 32 (serde only derives arrays up to 32): -/// serialized as a fixed-length tuple, exactly like the native array impls. -pub mod array_serialization { - use serde::de::{Error, SeqAccess, Visitor}; - use serde::ser::SerializeTuple; - use serde::{Deserialize, Deserializer, Serialize, Serializer}; - use std::marker::PhantomData; - - pub fn serialize(data: &[T; N], ser: S) -> Result { - let mut tup = ser.serialize_tuple(N)?; - for elem in data { - tup.serialize_element(elem)?; - } - tup.end() - } - - struct ArrayVisitor(PhantomData); - - impl<'de, T: Deserialize<'de> + Copy + Default, const N: usize> Visitor<'de> for ArrayVisitor { - type Value = [T; N]; - - fn expecting(&self, f: &mut std::fmt::Formatter) -> std::fmt::Result { - write!(f, "an array of length {N}") - } - - fn visit_seq>(self, mut seq: A) -> Result<[T; N], A::Error> { - let mut out = [T::default(); N]; - for (i, slot) in out.iter_mut().enumerate() { - *slot = seq.next_element()?.ok_or_else(|| Error::invalid_length(i, &self))?; - } - Ok(out) - } - } - - pub fn deserialize<'de, D, T, const N: usize>(de: D) -> Result<[T; N], D::Error> - where - D: Deserializer<'de>, - T: Deserialize<'de> + Copy + Default, - { - de.deserialize_tuple(N, ArrayVisitor::(PhantomData)) - } -} diff --git a/crates/xmss/src/ssz_serialization.rs b/crates/xmss/src/ssz_serialization.rs deleted file mode 100644 index a51affa29..000000000 --- a/crates/xmss/src/ssz_serialization.rs +++ /dev/null @@ -1,158 +0,0 @@ -//! SSZ encodings (Ethereum consensus-layer compatibility) for the public key and -//! the signature, the two objects that appear in consensus data. The secret key -//! never does, so it keeps only its serde persistence. -//! -//! Both are SSZ containers of fixed-size byte vectors, so each encoding is the -//! plain concatenation of its fields and every byte string of the right length -//! decodes: -//! -//! - [`XmssPublicKey`]: `merkle_root | public_param`, [`PUB_KEY_SSZ_LEN`] bytes. -//! - [`XmssSignature`]: `chain_tips | randomness | merkle_proof`, -//! [`SIGNATURE_SSZ_LEN`] bytes. - -pub use ssz::{Decode, DecodeError, Encode}; - -use crate::*; - -/// SSZ length of an encoded public key. Equal to [`PUB_KEY_SIZE`]: both encodings -/// concatenate the same fixed-size fields. -pub const PUB_KEY_SSZ_LEN: usize = PUB_KEY_SIZE; -/// SSZ length of an encoded signature. Equal to [`SIG_SIZE`], for the same reason. -pub const SIGNATURE_SSZ_LEN: usize = SIG_SIZE; - -/// Peels fixed-size fields off a buffer whose length was checked up front. -struct Reader<'a>(&'a [u8]); - -impl Reader<'_> { - fn take(&mut self) -> [u8; N] { - let (head, tail) = self.0.split_at(N); - self.0 = tail; - head.try_into().unwrap() - } -} - -fn check_len(bytes: &[u8], expected: usize) -> Result, DecodeError> { - if bytes.len() == expected { - Ok(Reader(bytes)) - } else { - Err(DecodeError::InvalidByteLength { - len: bytes.len(), - expected, - }) - } -} - -impl Encode for WotsSignature { - fn is_ssz_fixed_len() -> bool { - true - } - - fn ssz_fixed_len() -> usize { - WOTS_SIG_SIZE - } - - fn ssz_bytes_len(&self) -> usize { - WOTS_SIG_SIZE - } - - fn ssz_append(&self, buf: &mut Vec) { - for chain_tip in &self.chain_tips { - buf.extend_from_slice(chain_tip); - } - buf.extend_from_slice(&self.randomness); - } -} - -impl Decode for WotsSignature { - fn is_ssz_fixed_len() -> bool { - true - } - - fn ssz_fixed_len() -> usize { - WOTS_SIG_SIZE - } - - fn from_ssz_bytes(bytes: &[u8]) -> Result { - let mut reader = check_len(bytes, WOTS_SIG_SIZE)?; - Ok(Self { - chain_tips: std::array::from_fn(|_| reader.take()), - randomness: reader.take(), - }) - } -} - -impl Encode for XmssPublicKey { - fn is_ssz_fixed_len() -> bool { - true - } - - fn ssz_fixed_len() -> usize { - PUB_KEY_SSZ_LEN - } - - fn ssz_bytes_len(&self) -> usize { - PUB_KEY_SSZ_LEN - } - - fn ssz_append(&self, buf: &mut Vec) { - buf.extend_from_slice(&self.merkle_root); - buf.extend_from_slice(&self.public_param); - } -} - -impl Decode for XmssPublicKey { - fn is_ssz_fixed_len() -> bool { - true - } - - fn ssz_fixed_len() -> usize { - PUB_KEY_SSZ_LEN - } - - fn from_ssz_bytes(bytes: &[u8]) -> Result { - let mut reader = check_len(bytes, PUB_KEY_SSZ_LEN)?; - Ok(Self { - merkle_root: reader.take(), - public_param: reader.take(), - }) - } -} - -impl Encode for XmssSignature { - fn is_ssz_fixed_len() -> bool { - true - } - - fn ssz_fixed_len() -> usize { - SIGNATURE_SSZ_LEN - } - - fn ssz_bytes_len(&self) -> usize { - SIGNATURE_SSZ_LEN - } - - fn ssz_append(&self, buf: &mut Vec) { - self.wots_signature.ssz_append(buf); - for neighbour in &self.merkle_proof { - buf.extend_from_slice(neighbour); - } - } -} - -impl Decode for XmssSignature { - fn is_ssz_fixed_len() -> bool { - true - } - - fn ssz_fixed_len() -> usize { - SIGNATURE_SSZ_LEN - } - - fn from_ssz_bytes(bytes: &[u8]) -> Result { - let mut reader = check_len(bytes, SIGNATURE_SSZ_LEN)?; - Ok(Self { - wots_signature: WotsSignature::from_ssz_bytes(&reader.take::())?, - merkle_proof: std::array::from_fn(|_| reader.take()), - }) - } -} diff --git a/crates/xmss/src/wots.rs b/crates/xmss/src/wots.rs deleted file mode 100644 index a3d8cc6af..000000000 --- a/crates/xmss/src/wots.rs +++ /dev/null @@ -1,156 +0,0 @@ -//! WOTS (Winternitz one-time signature) with target-sum encoding. - -use serde::{Deserialize, Serialize}; - -use crate::*; - -#[derive(Debug)] -pub struct WotsSecretKey { - pre_images: [Digest; V], -} - -#[derive(Clone, Copy, Debug, PartialEq, Eq)] -pub struct WotsPublicKey(pub [Digest; V]); - -#[derive(Debug, Clone, Serialize, Deserialize, PartialEq, Eq, Hash)] -pub struct WotsSignature { - #[serde(with = "crate::array_serialization")] - pub chain_tips: [Digest; V], - pub randomness: Randomness, -} - -impl WotsSecretKey { - pub const fn new(pre_images: [Digest; V]) -> Self { - Self { pre_images } - } - - /// Walk every chain to its tip. Only key generation needs this; signing - /// stops each chain at its encoding digit. - pub fn public_key(&self, public_param: &PublicParam, epoch: Epoch) -> WotsPublicKey { - WotsPublicKey(std::array::from_fn(|i| { - iterate_hash(&self.pre_images[i], CHAIN_LENGTH - 1, public_param, epoch, i, 0) - })) - } - - /// `encoding` and `randomness` come as a pair from - /// [`find_randomness_for_wots_encoding`], which only returns a randomness - /// whose encoding is valid. - pub fn sign( - &self, - encoding: &[u8; V], - randomness: Randomness, - epoch: Epoch, - public_param: &PublicParam, - ) -> WotsSignature { - WotsSignature { - chain_tips: std::array::from_fn(|i| { - iterate_hash(&self.pre_images[i], encoding[i] as usize, public_param, epoch, i, 0) - }), - randomness, - } - } -} - -impl WotsSignature { - pub fn recover_public_key( - &self, - message: &Message, - epoch: Epoch, - public_param: &PublicParam, - ) -> Option { - let encoding = wots_encode(message, epoch, public_param, &self.randomness)?; - Some(WotsPublicKey(std::array::from_fn(|i| { - iterate_hash( - &self.chain_tips[i], - CHAIN_LENGTH - 1 - encoding[i] as usize, - public_param, - epoch, - i, - encoding[i] as usize, - ) - }))) - } -} - -impl WotsPublicKey { - /// The Merkle leaf: standard BLAKE2s over the tweak, public parameter, and - /// 42 concatenated chain tips (704 bytes, 11 compressions). - pub fn hash(&self, public_param: &PublicParam, epoch: Epoch) -> Digest { - tweak_hash(public_param, TWEAK_TYPE_WOTS_PK, 0, epoch, self.0.as_flattened()) - } -} - -/// One chain step (1 compression). The position `chain_index * CHAIN_LENGTH + -/// step` identifies the edge from chain value `step` to `step + 1`. -fn chain_step(public_param: &PublicParam, epoch: Epoch, chain_index: usize, step: usize, value: &Digest) -> Digest { - let position = (chain_index * CHAIN_LENGTH + step) as u32; - tweak_hash(public_param, TWEAK_TYPE_CHAIN, position, epoch, value) -} - -/// Walk chain `chain_index` for `n` steps starting at chain value `start_step`. -fn iterate_hash( - start: &Digest, - steps: usize, - public_param: &PublicParam, - epoch: Epoch, - chain_index: usize, - start_step: usize, -) -> Digest { - (0..steps).fold(*start, |value, offset| { - chain_step(public_param, epoch, chain_index, start_step + offset, &value) - }) -} - -pub fn find_randomness_for_wots_encoding( - message: &Message, - epoch: Epoch, - public_param: &PublicParam, - seed: &[u8; 32], -) -> Option<(Randomness, [u8; V], u64)> { - (0..MAX_RANDOMIZER_TRIALS).find_map(|trial| { - let mut hasher = primitives::hash::Hasher::new(); - hasher.update(&make_tweak(TWEAK_TYPE_RANDOMIZER, trial as u32, epoch)); - hasher.update(public_param).update(seed).update(message); - let randomness = hasher.finalize()[..RANDOMNESS_LEN].try_into().unwrap(); - wots_encode(message, epoch, public_param, &randomness).map(|encoding| (randomness, encoding, trial + 1)) - }) -} - -/// The target-sum encoding. `D = MD(msg | randomness | zeros)` under the -/// encoding tweak, truncated to 16 bytes: 2 standard BLAKE2s compressions over -/// the 96-byte exact input. `D`'s two little-endian 64-bit words each hold 21 -/// chunks of 3 bits (the VM's word width budgets the monomial encoding at 64 -/// bits per word: `g^k = x^k` only for `k < 64`): digit `i < 21` sits at bits -/// `3i` of word 0, digit `i >= 21` at bits `3(i-21)` of word 1. The encoding is -/// valid iff the leftover top bit of EACH word (bits 63 and 127) is zero AND -/// the chunks sum to [`TARGET_SUM`]. Grinding the top bits to zero makes each -/// digest word exactly `sum(e_i * 2^{3i})` of its 21 digits, so both words -/// decompose into the chunks with no slack term. In-circuit this is checked -/// over GF(2^64) per word by accumulating the dispatched digit literals against -/// `8^i` monomial weights (see `verify_sig` in `rec_aggregation`'s `guests/lean_ethereum.py`). -pub fn wots_encode( - message: &Message, - epoch: Epoch, - public_param: &PublicParam, - randomness: &Randomness, -) -> Option<[u8; V]> { - let mut data = [0u8; 2 * STATE_LEN]; - data[..MESSAGE_LEN].copy_from_slice(message); - data[MESSAGE_LEN..][..RANDOMNESS_LEN].copy_from_slice(randomness); - let digest = tweak_hash(public_param, TWEAK_TYPE_ENCODING, 0, epoch, &data); - - let mut encoding = [0u8; V]; - let mut sum = 0; - for (half, bytes) in digest.as_chunks::<8>().0.iter().enumerate() { - let word = u64::from_le_bytes(*bytes); - if word >> (W * V / 2) != 0 { - return None; - } - for i in 0..V / 2 { - let digit = ((word >> (W * i)) & (CHAIN_LENGTH as u64 - 1)) as u8; - encoding[half * (V / 2) + i] = digit; - sum += digit as usize; - } - } - (sum == TARGET_SUM).then_some(encoding) -} diff --git a/crates/xmss/src/xmss.rs b/crates/xmss/src/xmss.rs deleted file mode 100644 index a75e4e90d..000000000 --- a/crates/xmss/src/xmss.rs +++ /dev/null @@ -1,383 +0,0 @@ -//! XMSS: a Merkle tree of `2^LOG_LIFETIME` WOTS public-key hashes. -//! -//! For a range of R = epoch_end - -//! epoch_start + 1 epochs, storage is O(sqrt(R) + LOG_LIFETIME) instead of O(R). -//! The key stores the top tree (in-range band plus a thin spine) and one cached -//! bottom subtree, cut at `split_level = ceil(log2(R)) / 2`. Out-of-range nodes -//! are deterministic `gen_random_node` fillers. - -use std::sync::{Mutex, MutexGuard}; - -use rand::{CryptoRng, Rng}; -use serde::{Deserialize, Serialize}; - -use crate::*; - -/// The encoding is SECRET KEY MATERIAL: it carries the seed. -#[derive(Debug, Serialize, Deserialize)] -pub struct XmssSecretKey { - pub(crate) epoch_start: Epoch, - pub(crate) epoch_end: Epoch, - pub(crate) public_param: PublicParam, - pub(crate) seed: [u8; 32], - pub(crate) split_level: usize, - pub(crate) top: Vec>, - #[serde(skip)] - pub(crate) cache: Mutex>, -} - -#[derive(Debug)] -pub(crate) struct BottomSubtree { - subtree_index: u64, - layers: Vec>, -} - -#[derive(Debug, Clone, Serialize, Deserialize, PartialEq, Eq, Hash)] -pub struct XmssSignature { - pub wots_signature: WotsSignature, - pub merkle_proof: [Digest; LOG_LIFETIME], -} - -/// Ordered lexicographically on `flatten()`, which is what an aggregate's signer -/// set is sorted and deduplicated by. -#[derive(Debug, Clone, Serialize, Deserialize, PartialEq, Eq, Hash, PartialOrd, Ord)] -pub struct XmssPublicKey { - pub merkle_root: Digest, - pub public_param: PublicParam, -} - -impl XmssPublicKey { - pub fn flatten(&self) -> [u8; PUB_KEY_SIZE] { - let mut out = [0u8; PUB_KEY_SIZE]; - out[..DIGEST_LEN].copy_from_slice(&self.merkle_root); - out[DIGEST_LEN..].copy_from_slice(&self.public_param); - out - } -} - -fn gen_wots_secret_key(seed: &[u8; 32], public_param: &PublicParam, epoch: Epoch) -> WotsSecretKey { - let pre_images = std::array::from_fn(|i| tweak_hash(public_param, TWEAK_TYPE_PRF, i as u32, epoch, seed)); - WotsSecretKey::new(pre_images) -} - -fn gen_public_param(seed: &[u8; 32]) -> PublicParam { - tweak_hash(&[0; PUBLIC_PARAM_LEN], TWEAK_TYPE_PARAMETER, 0, 0, seed) -} - -fn gen_random_node(seed: &[u8; 32], public_param: &PublicParam, level: usize, index: u64) -> Digest { - tweak_hash(public_param, TWEAK_TYPE_FILLER, level as u32, index as u32, seed) -} - -/// Merkle parent at `level` (1 compression: both children fill one block). -fn merkle_node(public_param: &PublicParam, level: usize, index: u64, left: &Digest, right: &Digest) -> Digest { - let mut data = [0u8; 2 * DIGEST_LEN]; - data[..DIGEST_LEN].copy_from_slice(left); - data[DIGEST_LEN..].copy_from_slice(right); - tweak_hash(public_param, TWEAK_TYPE_MERKLE, level as u32, index as u32, &data) -} - -/// Level-0 layer: WOTS public-key hashes for the in-range leaves `[lo, hi]`. -/// -/// Sequential: this runs once per bottom subtree, and the subtrees are what -/// [`key_gen`] fans out over. -fn leaf_layer(seed: &[u8; 32], public_param: &PublicParam, first_epoch: u64, last_epoch: u64) -> Vec { - (first_epoch..=last_epoch) - .map(|epoch| { - gen_wots_secret_key(seed, public_param, epoch as Epoch) - .public_key(public_param, epoch as Epoch) - .hash(public_param, epoch as Epoch) - }) - .collect() -} - -/// Build levels `(from_level+1)..=to_level` onto `layers`; out-of-range children -/// use `gen_random_node`. -/// -/// Sequential: each level depends on the one below it, and every level here is at -/// most `O(sqrt(R))` wide, whether it is a bottom subtree's levels or the top part -/// built from the subtree roots. -fn build_up( - seed: &[u8; 32], - public_param: &PublicParam, - layers: &mut Vec>, - first_epoch: u64, - last_epoch: u64, - from_level: usize, - to_level: usize, -) { - for level in (from_level + 1)..=to_level { - let (first_node, last_node) = (first_epoch >> level, last_epoch >> level); - let (first_child, last_child) = (first_epoch >> (level - 1), last_epoch >> (level - 1)); - let children = layers.last().unwrap(); - let nodes: Vec = (first_node..=last_node) - .map(|index| { - let child = |child_index: u64| { - if child_index >= first_child && child_index <= last_child { - children[(child_index - first_child) as usize] - } else { - gen_random_node(seed, public_param, level - 1, child_index) - } - }; - merkle_node(public_param, level, index, &child(2 * index), &child(2 * index + 1)) - }) - .collect(); - layers.push(nodes); - } -} - -fn subtree_bounds(epoch_start: u64, epoch_end: u64, split_level: usize, subtree_index: u64) -> (u64, u64) { - ( - epoch_start.max(subtree_index << split_level), - epoch_end.min(((subtree_index + 1) << split_level) - 1), - ) -} - -fn build_subtree_layers( - seed: &[u8; 32], - public_param: &PublicParam, - first_epoch: u64, - last_epoch: u64, - to_level: usize, -) -> Vec> { - let mut layers = vec![leaf_layer(seed, public_param, first_epoch, last_epoch)]; - build_up(seed, public_param, &mut layers, first_epoch, last_epoch, 0, to_level); - layers -} - -#[derive(Debug, PartialEq, Eq, Clone, Copy, Hash)] -pub enum XmssKeyGenError { - InvalidRange, -} - -impl std::fmt::Display for XmssKeyGenError { - fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { - match self { - Self::InvalidRange => write!(f, "epoch_start is past epoch_end"), - } - } -} - -impl std::error::Error for XmssKeyGenError {} - -/// A fresh key pair for `epoch_start..=epoch_end`, with its seed sampled from `rng`. -pub fn key_gen( - rng: &mut impl CryptoRng, - epoch_start: Epoch, - epoch_end: Epoch, -) -> Result<(XmssSecretKey, XmssPublicKey), XmssKeyGenError> { - key_gen_from_seed(rng.random(), epoch_start, epoch_end) -} - -/// Deterministic [`key_gen`]: one `(seed, epoch range)` always regenerates the -/// same key pair, the seed being the key's entire secret material. -pub fn key_gen_from_seed( - seed: [u8; 32], - epoch_start: Epoch, - epoch_end: Epoch, -) -> Result<(XmssSecretKey, XmssPublicKey), XmssKeyGenError> { - if epoch_start > epoch_end { - return Err(XmssKeyGenError::InvalidRange); - } - let public_param = gen_public_param(&seed); - let (first_epoch, last_epoch) = (epoch_start as u64, epoch_end as u64); - - let split_level = primitives::log2_ceil_usize((last_epoch - first_epoch + 1) as usize).div_ceil(2); - - // The bottom subtrees are the only fan-out in key generation: `O(sqrt(R))` of - // them, each independent, each holding only its own `O(sqrt(R))` layers, so - // peak memory stays `O(sqrt(R))` and one flat dispatch covers the whole tree. - // Everything below this is sequential by construction. - let first_subtree = first_epoch >> split_level; - let last_subtree = last_epoch >> split_level; - let root_layer: Vec = parallel::map_collect((last_subtree - first_subtree + 1) as usize, |i| { - let subtree = first_subtree + i as u64; - let (first_leaf, last_leaf) = subtree_bounds(first_epoch, last_epoch, split_level, subtree); - build_subtree_layers(&seed, &public_param, first_leaf, last_leaf, split_level)[split_level][0] - }); - - let mut top = vec![root_layer]; - build_up( - &seed, - &public_param, - &mut top, - first_epoch, - last_epoch, - split_level, - LOG_LIFETIME, - ); - - let pub_key = XmssPublicKey { - merkle_root: top.last().unwrap()[0], - public_param, - }; - let secret_key = XmssSecretKey { - epoch_start, - epoch_end, - public_param, - seed, - split_level, - top, - cache: Mutex::new(None), - }; - Ok((secret_key, pub_key)) -} - -#[derive(Debug, PartialEq, Eq, Clone, Copy, Hash)] -pub enum XmssSignError { - EpochOutOfRange, - NoAdmissibleEncoding, -} - -impl std::fmt::Display for XmssSignError { - fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { - match self { - Self::EpochOutOfRange => write!(f, "the epoch is outside the key's range"), - Self::NoAdmissibleEncoding => write!(f, "no admissible encoding within the randomizer trial limit"), - } - } -} - -impl std::error::Error for XmssSignError {} - -/// Never use the same key and epoch to sign two different messages. -/// Signing is deterministic. -pub fn sign(secret_key: &XmssSecretKey, message: &Message, epoch: Epoch) -> Result { - if epoch < secret_key.epoch_start || epoch > secret_key.epoch_end { - return Err(XmssSignError::EpochOutOfRange); - } - let (randomness, encoding, _) = - find_randomness_for_wots_encoding(message, epoch, &secret_key.public_param, &secret_key.seed) - .ok_or(XmssSignError::NoAdmissibleEncoding)?; - let wots_secret_key = gen_wots_secret_key(&secret_key.seed, &secret_key.public_param, epoch); - let wots_signature = wots_secret_key.sign(&encoding, randomness, epoch, &secret_key.public_param); - - let cache = secret_key.cached_bottom_subtree(epoch); - let subtree = cache.as_ref().unwrap(); - let merkle_proof = std::array::from_fn(|level| { - let neighbour_index = ((epoch as u64) >> level) ^ 1; - secret_key.merkle_sibling(level, neighbour_index, subtree) - }); - drop(cache); - Ok(XmssSignature { - wots_signature, - merkle_proof, - }) -} - -impl XmssSecretKey { - /// The epochs this key can sign at. The caller must ensure each epoch signs only one message. - pub fn epoch_range(&self) -> std::ops::RangeInclusive { - self.epoch_start..=self.epoch_end - } - - pub fn public_key(&self) -> XmssPublicKey { - XmssPublicKey { - merkle_root: self.top.last().unwrap()[0], - public_param: self.public_param, - } - } - - /// Warms the signing cache for `epoch`: when the next epoch to sign at is - /// known in advance, calling this ahead of time makes the [`sign`] faster. - pub fn prepare(&self, epoch: Epoch) -> Result<(), XmssSignError> { - if epoch < self.epoch_start || epoch > self.epoch_end { - return Err(XmssSignError::EpochOutOfRange); - } - drop(self.cached_bottom_subtree(epoch)); - Ok(()) - } - - /// The bottom subtree covering `epoch`, rebuilt only on a miss: one subtree - /// serves all `2^split_level` epochs under it. - fn cached_bottom_subtree(&self, epoch: Epoch) -> MutexGuard<'_, Option> { - let subtree_index = (epoch as u64) >> self.split_level; - let mut cache = self.cache.lock().unwrap(); - if cache.as_ref().is_none_or(|s| s.subtree_index != subtree_index) { - *cache = Some(self.build_bottom_subtree(subtree_index)); - } - cache - } - - fn build_bottom_subtree(&self, subtree_index: u64) -> BottomSubtree { - let (lo, hi) = subtree_bounds( - self.epoch_start as u64, - self.epoch_end as u64, - self.split_level, - subtree_index, - ); - let layers = build_subtree_layers(&self.seed, &self.public_param, lo, hi, self.split_level); - BottomSubtree { subtree_index, layers } - } - - /// Authentication-path sibling at `level`: from the top part, the cached - /// subtree, or `gen_random_node`. - fn merkle_sibling(&self, level: usize, neighbour_index: u64, subtree: &BottomSubtree) -> Digest { - let (first_epoch, last_epoch, level_base, layers) = if level >= self.split_level { - ( - self.epoch_start as u64, - self.epoch_end as u64, - self.split_level, - &self.top, - ) - } else { - let (first_epoch, last_epoch) = subtree_bounds( - self.epoch_start as u64, - self.epoch_end as u64, - self.split_level, - subtree.subtree_index, - ); - (first_epoch, last_epoch, 0, &subtree.layers) - }; - let first_node = first_epoch >> level; - if neighbour_index >= first_node && neighbour_index <= (last_epoch >> level) { - layers[level - level_base][(neighbour_index - first_node) as usize] - } else { - gen_random_node(&self.seed, &self.public_param, level, neighbour_index) - } - } -} - -#[derive(Debug, PartialEq, Eq, Clone, Copy, Hash)] -pub enum XmssVerifyError { - InvalidWots, - InvalidMerklePath, -} - -impl std::fmt::Display for XmssVerifyError { - fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result { - match self { - Self::InvalidWots => write!(f, "the WOTS signature does not recover a public key"), - Self::InvalidMerklePath => write!(f, "the authentication path does not reach the key's root"), - } - } -} - -impl std::error::Error for XmssVerifyError {} - -pub fn verify( - pub_key: &XmssPublicKey, - message: &Message, - signature: &XmssSignature, - epoch: Epoch, -) -> Result<(), XmssVerifyError> { - let wots_public_key = signature - .wots_signature - .recover_public_key(message, epoch, &pub_key.public_param) - .ok_or(XmssVerifyError::InvalidWots)?; - let mut current = wots_public_key.hash(&pub_key.public_param, epoch); - for (level, neighbour) in signature.merkle_proof.iter().enumerate() { - let is_left = ((epoch as u64 >> level) & 1) == 0; - let parent_index = (epoch as u64) >> (level + 1); - let (left, right) = if is_left { - (current, *neighbour) - } else { - (*neighbour, current) - }; - current = merkle_node(&pub_key.public_param, level + 1, parent_index, &left, &right); - } - if current == pub_key.merkle_root { - Ok(()) - } else { - Err(XmssVerifyError::InvalidMerklePath) - } -} diff --git a/crates/xmss/tests/xmss_tests.rs b/crates/xmss/tests/xmss_tests.rs deleted file mode 100644 index da12fbcc4..000000000 --- a/crates/xmss/tests/xmss_tests.rs +++ /dev/null @@ -1,259 +0,0 @@ -use rand::{Rng, SeedableRng, rngs::StdRng}; -use xmss::*; - -fn test_message() -> Message { - std::array::from_fn(|i| (i * 3 + 7) as u8) -} - -#[test] -fn keygen_sign_verify() { - let seed: [u8; 32] = std::array::from_fn(|i| i as u8); - let message = test_message(); - - for epoch in [0u32, 1234, u32::MAX] { - let (sk, pk) = key_gen_from_seed(seed, epoch.saturating_sub(1), epoch.saturating_add(2)).unwrap(); - let sig = sign(&sk, &message, epoch).unwrap(); - verify(&pk, &message, &sig, epoch).unwrap(); - assert_eq!(sign(&sk, &message, epoch).unwrap(), sig); - } -} - -#[test] -fn serialize_deserialize_and_size() { - let seed: [u8; 32] = std::array::from_fn(|i| i as u8); - let message = test_message(); - let epoch = 110; - - let (sk, pk) = key_gen_from_seed(seed, 100, 115).unwrap(); - let sig = sign(&sk, &message, epoch).unwrap(); - - let public_key_bytes = bincode::serialize(&pk).unwrap(); - assert_eq!(public_key_bytes.len(), PUB_KEY_SIZE); - let decoded_public_key: XmssPublicKey = bincode::deserialize(&public_key_bytes).unwrap(); - assert_eq!(pk, decoded_public_key); - - let signature_bytes = bincode::serialize(&sig).unwrap(); - assert_eq!(signature_bytes.len(), SIG_SIZE); - let decoded_signature: XmssSignature = bincode::deserialize(&signature_bytes).unwrap(); - assert_eq!(sig, decoded_signature); - - verify(&decoded_public_key, &message, &decoded_signature, epoch).unwrap(); -} - -#[test] -fn deterministic_keygen_and_range_separation() { - let seed = [3u8; 32]; - let (_, pk) = key_gen_from_seed(seed, 50, 60).unwrap(); - let (_, same_seed_and_range) = key_gen_from_seed(seed, 50, 60).unwrap(); - assert_eq!(pk, same_seed_and_range); - // A different range changes the filler/real split, hence the root. - let (_, longer_range) = key_gen_from_seed(seed, 50, 61).unwrap(); - assert_ne!(pk.merkle_root, longer_range.merkle_root); -} - -#[test] -fn tweak_separates_hash_domains() { - let pp = [7u8; PUBLIC_PARAM_LEN]; - let x = [1u8; DIGEST_LEN]; - let base = tweak_hash(&pp, TWEAK_TYPE_CHAIN, 3, 5, &x); - // Different type, position, index, or public parameter: different hash. - assert_ne!(base, tweak_hash(&pp, TWEAK_TYPE_MERKLE, 3, 5, &x)); - assert_ne!(base, tweak_hash(&pp, TWEAK_TYPE_CHAIN, 4, 5, &x)); - assert_ne!(base, tweak_hash(&pp, TWEAK_TYPE_CHAIN, 3, 6, &x)); - assert_ne!(base, tweak_hash(&[8u8; PUBLIC_PARAM_LEN], TWEAK_TYPE_CHAIN, 3, 5, &x)); - // Standard BLAKE2s binds the exact payload length. - let mut extended = [0u8; STATE_LEN]; - extended[..DIGEST_LEN].copy_from_slice(&x); - assert_ne!(base, tweak_hash(&pp, TWEAK_TYPE_CHAIN, 3, 5, &extended)); -} - -/// The multi-block WOTS public-key hash is standard BLAKE2s of `tweak | pp | -/// payload`, assembled independently here and streamed in unrelated chunks. -#[test] -fn multi_block_tweak_hash_is_standard_blake2s() { - let pp = [9u8; PUBLIC_PARAM_LEN]; - let payload = [5u8; V * DIGEST_LEN]; - let mut input = Vec::new(); - input.extend_from_slice(&make_tweak(TWEAK_TYPE_WOTS_PK, 0, 42)); - input.extend_from_slice(&pp); - input.extend_from_slice(&payload); - - let mut hasher = primitives::hash::Hasher::new(); - for chunk in input.chunks(37) { - hasher.update(chunk); - } - let expected = hasher.finalize(); - assert_eq!(expected, primitives::hash::hash(&input)); - assert_eq!( - tweak_hash(&pp, TWEAK_TYPE_WOTS_PK, 0, 42, &payload), - expected[..DIGEST_LEN] - ); -} - -#[test] -fn tampered_signatures_rejected() { - let seed = [9u8; 32]; - let message = test_message(); - let epoch = 7; - let (sk, pk) = key_gen_from_seed(seed, 0, 15).unwrap(); - let sig = sign(&sk, &message, epoch).unwrap(); - verify(&pk, &message, &sig, epoch).unwrap(); - - let mut bad_message = message; - bad_message[0] ^= 1; - assert!(verify(&pk, &bad_message, &sig, epoch).is_err()); - - assert!(verify(&pk, &message, &sig, epoch + 1).is_err()); - - let mut bad_chain_tip = sig.clone(); - bad_chain_tip.wots_signature.chain_tips[5][0] ^= 1; - assert!(verify(&pk, &message, &bad_chain_tip, epoch).is_err()); - - let mut bad_randomness = sig.clone(); - bad_randomness.wots_signature.randomness[0] ^= 1; - assert!(verify(&pk, &message, &bad_randomness, epoch).is_err()); - - let mut bad_merkle_path = sig.clone(); - bad_merkle_path.merkle_proof[10][3] ^= 1; - assert_eq!( - verify(&pk, &message, &bad_merkle_path, epoch), - Err(XmssVerifyError::InvalidMerklePath) - ); - - assert_eq!(sign(&sk, &message, 16), Err(XmssSignError::EpochOutOfRange)); -} - -/// Detect changes to the encoding predicate through its grinding cost. -#[test] -#[ignore] -fn encoding_grinding_bits() { - let n = 200; - let pp = [0u8; PUBLIC_PARAM_LEN]; - let mut total_iters = 0u64; - for i in 0..n { - let mut rng = StdRng::seed_from_u64(i as u64); - let message: Message = rng.random(); - let (_, _, num_iters) = find_randomness_for_wots_encoding(&message, i as u32, &pp, &rng.random()).unwrap(); - total_iters += num_iters; - } - let bits = (total_iters as f64 / n as f64).log2(); - println!("Average grinding bits: {bits:.1}"); - assert!((13.5..15.5).contains(&bits), "grinding cost moved: {bits:.2} bits"); -} - -/// A reloaded secret key must sign exactly as the original does, so that -/// persisting one is a real alternative to regenerating it. -#[test] -fn secret_key_survives_a_round_trip() { - let seed: [u8; 32] = std::array::from_fn(|i| (i * 11) as u8); - let (sk, pk) = key_gen_from_seed(seed, 40, 45).unwrap(); - let reloaded: XmssSecretKey = bincode::deserialize(&bincode::serialize(&sk).unwrap()).unwrap(); - - assert_eq!(reloaded.public_key(), pk); - assert_eq!(reloaded.epoch_range(), 40..=45); - let message = test_message(); - for epoch in [40, 43, 45] { - let sig = sign(&reloaded, &message, epoch as u32).unwrap(); - verify(&pk, &message, &sig, epoch as u32).unwrap(); - } - assert_eq!(sign(&reloaded, &message, 46), Err(XmssSignError::EpochOutOfRange)); -} - -/// The SSZ encoding is the container's fields concatenated, in declaration -/// order. Assembled independently here, so a reordered or resized field cannot -/// be mirrored by the impl. -#[test] -fn ssz_layout_is_exact() { - let seed: [u8; 32] = std::array::from_fn(|i| (i * 5 + 1) as u8); - let message = test_message(); - let epoch = 300; - let (sk, pk) = key_gen_from_seed(seed, 290, 310).unwrap(); - let sig = sign(&sk, &message, epoch).unwrap(); - - let mut expected_pk = Vec::new(); - expected_pk.extend_from_slice(&pk.merkle_root); - expected_pk.extend_from_slice(&pk.public_param); - assert_eq!(expected_pk.len(), PUB_KEY_SSZ_LEN); - assert_eq!(pk.as_ssz_bytes(), expected_pk); - assert_eq!(expected_pk, pk.flatten()); - - let mut expected_sig = Vec::new(); - for chain_tip in &sig.wots_signature.chain_tips { - expected_sig.extend_from_slice(chain_tip); - } - expected_sig.extend_from_slice(&sig.wots_signature.randomness); - for neighbour in &sig.merkle_proof { - expected_sig.extend_from_slice(neighbour); - } - assert_eq!(expected_sig.len(), SIGNATURE_SSZ_LEN); - assert_eq!(sig.as_ssz_bytes(), expected_sig); - - // Round trip, and the decoded pair still verifies. - let decoded_pk = XmssPublicKey::from_ssz_bytes(&expected_pk).unwrap(); - let decoded_sig = XmssSignature::from_ssz_bytes(&expected_sig).unwrap(); - assert_eq!(decoded_pk, pk); - assert_eq!(decoded_sig, sig); - verify(&decoded_pk, &message, &decoded_sig, epoch).unwrap(); -} - -/// A fixed-length container rejects any other length, and only that. -#[test] -fn ssz_rejects_wrong_lengths() { - for len in [0, PUB_KEY_SSZ_LEN - 1, PUB_KEY_SSZ_LEN + 1] { - assert!(matches!( - XmssPublicKey::from_ssz_bytes(&vec![0u8; len]), - Err(DecodeError::InvalidByteLength { .. }) - )); - } - for len in [0, SIGNATURE_SSZ_LEN - 1, SIGNATURE_SSZ_LEN + 1] { - assert!(matches!( - XmssSignature::from_ssz_bytes(&vec![0u8; len]), - Err(DecodeError::InvalidByteLength { .. }) - )); - } - // Every byte string of the right length is a well-formed encoding. - XmssPublicKey::from_ssz_bytes(&[0xab; PUB_KEY_SSZ_LEN]).unwrap(); - XmssSignature::from_ssz_bytes(&[0xcd; SIGNATURE_SSZ_LEN]).unwrap(); -} - -/// `prepare` only warms a cache, so it must change no signature, and a later -/// epoch in a different bottom subtree must still evict what it left behind. -#[test] -fn prepare_warms_without_changing_signatures() { - let seed = [21u8; 32]; - let message = test_message(); - let (sk, pk) = key_gen_from_seed(seed, 0, 255).unwrap(); - - // 0 and 200 are far enough apart to land in different bottom subtrees. - sk.prepare(200).unwrap(); - let after_prepare = sign(&sk, &message, 200).unwrap(); - verify(&pk, &message, &after_prepare, 200).unwrap(); - - let fresh = sign(&sk, &message, 200).unwrap(); - assert_eq!(after_prepare, fresh); - - // A miss on the warmed subtree rebuilds rather than reusing it. - sk.prepare(0).unwrap(); - let other = sign(&sk, &message, 0).unwrap(); - verify(&pk, &message, &other, 0).unwrap(); - - assert_eq!(sk.prepare(256), Err(XmssSignError::EpochOutOfRange)); -} - -/// The rng entry point must forward `(seed, epoch_start, epoch_end)` in that -/// order, and draw a fresh seed on every call. -#[test] -fn key_gen_draws_a_usable_seed() { - let mut rng = StdRng::seed_from_u64(99); - let message = test_message(); - let (sk, pk) = key_gen(&mut rng, 70, 80).unwrap(); - assert_eq!(sk.epoch_range(), 70..=80); - let sig = sign(&sk, &message, 75).unwrap(); - verify(&pk, &message, &sig, 75).unwrap(); - - // A fresh draw is a different key. - let (_, other) = key_gen(&mut rng, 70, 80).unwrap(); - assert_ne!(pk, other); - - assert_eq!(key_gen(&mut rng, 80, 70).unwrap_err(), XmssKeyGenError::InvalidRange); -} diff --git a/crates/rec_aggregation/src/fibonacci.rs b/src/fibonacci.rs similarity index 87% rename from crates/rec_aggregation/src/fibonacci.rs rename to src/fibonacci.rs index ba69ae862..878ea3934 100644 --- a/crates/rec_aggregation/src/fibonacci.rs +++ b/src/fibonacci.rs @@ -1,13 +1,8 @@ //! Fibonacci in the exponent: the demo benchmark (`buff[g^k] = g^{F(k)}`, //! recurrence `buff[i·g²] = buff[i·g] · buff[i]`). -use lean_compiler::{compile, parse}; -use lean_vm::cpu::{prove, verify}; -use primitives::{ - bench::Plan, - field::{F64, g_pow}, - pretty_f64, pretty_integer, -}; +use leanvm::{F64, compile, g_pow, parse, prove, verify}; +use primitives::{bench::Plan, pretty_f64, pretty_integer}; /// Prove and verify Fibonacci-in-the-exponent over a `HeapBuf` (an unrolled /// `mul_range` recurrence), binding `g^{F(n)}` as the public input. Prints the @@ -19,7 +14,7 @@ pub fn run_fibonacci(n: usize, log_inv_rate: usize, plan: Plan) { let (src, pi) = fibonacci_program(n); let program = compile(&parse(&src).unwrap()); - // Only the final measured pass of each stage is traced (see `run_recursion`). + // Only the final measured pass of each stage is traced. let ((proof, stats), prove_time) = plan.warm_then_measure(|last| { let _quiet = (!last).then(primitives::suppress_tracing); prove(&program, pi, log_inv_rate) @@ -29,6 +24,8 @@ pub fn run_fibonacci(n: usize, log_inv_rate: usize, plan: Plan) { verify(&program, &pi, &proof).unwrap() }); + // tracing-forest renders its tree only when the root span closes, so the + // complete trace has to be flushed above the report. drop(trace_span); println!( @@ -37,14 +34,15 @@ pub fn run_fibonacci(n: usize, log_inv_rate: usize, plan: Plan) { ); println!(" cycles (VM steps) : {}", pretty_integer(stats.cycles)); println!(" details : {}", stats.details()); - crate::report::print_proof_size(&proof); + let proof_bytes = bincode::serialized_size(&proof).expect("proof is serializable"); + println!(" proof size : {:.1} KiB", proof_bytes as f64 / 1024.0); let cycles_per_second = (stats.cycles as f64 / prove_time.mean()).round() as u64; println!( " proving : {} s{} {} cycles/s peak memory {} GiB", pretty_f64(prove_time.mean()), prove_time.spread(), pretty_integer(cycles_per_second), - crate::report::peak_gib() + pretty_f64(primitives::bench::peak_rss_bytes() as f64 / (1u64 << 30) as f64) ); println!( " verifying : {} ms", diff --git a/src/lib.rs b/src/lib.rs index 4d8bad3d9..bcb130363 100644 --- a/src/lib.rs +++ b/src/lib.rs @@ -1,59 +1,26 @@ -//! leanVM proves XMSS and SPHINCS signature claims and LeanDA blob well-formedness. +//! leanVM: a minimal zkVM. A zkDSL program compiles to the ISA, [`prove`] runs it and +//! proves the run, [`verify`] checks the proof against the program and its public input. //! -//! Release only: the zkDSL compiler [`setup_verifier`] runs overflows the debug stack. +//! Release only: the zkDSL compiler overflows the debug stack. //! //! End to end in [`tests/api.rs`](https://github.com/leanEthereum/leanVM/blob/main/tests/api.rs). -pub use rec_aggregation::{ - AggregateVerifyError, AggregationError, ClaimSelection, DA_LOG_CELL, DA_LOG_K, DA_MAX_ROWS, EthereumProof, - MAX_DA_ROOTS, MAX_EPOCHS, MAX_KEYS, MAX_RECURSIONS, SignatureClaims, SphincsClaim, XmssClaimGroup, aggregate, -}; - +pub use lean_compiler::{compile, parse}; pub use lean_vm::{ - cpu::CpuError, + cpu::{CpuError, Program, Proof, Stats, prove, verify}, pcs::{MAX_LOG_INV_RATE, MIN_LOG_INV_RATE}, }; +pub use primitives::field::{F64, g_pow}; -/// LeanDA commitments. Blob symbols are `u64` words read in little-endian byte order. -pub mod lean_da { - pub use ::lean_da::{ - BLOB_SYMBOLS, CELL_SYMBOLS, CELLS_PER_ROW, CODEWORD_SYMBOLS, DA_LOG_CELL, DA_LOG_K, DA_MAX_ROWS, DaCommitment, - DaWitness, commit, - }; -} - -pub mod xmss { - /// The SSZ traits [`XmssPublicKey`] and [`XmssSignature`] implement: import - /// them to call `as_ssz_bytes` and `from_ssz_bytes`. - pub use ::xmss::{Decode, DecodeError, Encode}; - pub use ::xmss::{ - Digest, Epoch, LOG_LIFETIME, MESSAGE_LEN, Message, PUB_KEY_SIZE, PUB_KEY_SSZ_LEN, PublicParam, SIG_SIZE, - SIGNATURE_SSZ_LEN, WotsSignature, XmssKeyGenError, XmssPublicKey, XmssSecretKey, XmssSignError, XmssSignature, - XmssVerifyError, key_gen, key_gen_from_seed, sign, verify, - }; -} - -pub mod sphincs { - pub use ::sphincs::{ - Digest, FtsOpening, MASTER_SECRET_LEN, MESSAGE_LEN, MasterSecret, Message, PUB_KEY_SIZE, PublicParam, - SECRET_KEY_SIZE, SIG_SIZE, SphincsPublicKey, SphincsSecretKey, SphincsSignError, SphincsSignature, - SphincsVerifyError, key_gen, key_gen_from, key_gen_from_seed, sign, verify, - }; -} - -pub use rand; - -/// Call once before verifying an [`EthereumProof`]. Idempotent, and -/// [`setup_prover`] does it for you. +/// Call once before [`verify`]. Idempotent, and [`setup_prover`] does it for you. pub fn setup_verifier() { lean_vm::init_prover_pool(); - rec_aggregation::warm_up(); } -/// Call once before [`aggregate`]. +/// Call once before [`prove`]. /// -/// There is one arena per process, so only one [`aggregate`] call may run at a -/// time in a process: to aggregate in parallel, use separate processes. +/// There is one arena per process, so only one [`prove`] call may run at a +/// time in a process: to prove in parallel, use separate processes. pub fn setup_prover() { zk_alloc::enable_arena(); setup_prover_without_arena(); diff --git a/src/main.rs b/src/main.rs index ebf99ce77..ef488f700 100644 --- a/src/main.rs +++ b/src/main.rs @@ -1,7 +1,9 @@ -//! Benchmark CLI for signature and blob proofs, recursion, and the Fibonacci demo. +//! Benchmark CLI. use clap::{Parser, Subcommand}; +mod fibonacci; + #[derive(Parser)] struct Cli { /// WHIR inverse-rate logarithm (1 through 4). @@ -36,43 +38,7 @@ struct Cli { #[derive(Subcommand)] enum Command { - /// Prove signatures and blobs, then verify the proof. At least one count must be nonzero. - Aggregate { - /// XMSS signatures to aggregate. - #[arg(long, default_value = "0")] - xmss: usize, - /// SPHINCS signatures to aggregate. - #[arg(long, default_value = "0")] - sphincs: usize, - /// Blobs in one LeanDA commitment (128 KiB each). - #[arg( - long, - default_value_t = 0, - value_parser = clap::builder::RangedU64ValueParser::::new().range(0..=leanvm::lean_da::DA_MAX_ROWS as u64) - )] - blobs: usize, - }, - /// Aggregate n child proofs into one proof. - Recursion { - /// Number of child aggregates. - #[arg(long, default_value = "2")] - n: usize, - /// XMSS signatures in each child. Sets the child proof's committed size, - /// which is what the recursion cost should be quoted against. - #[arg(long, default_value = "900")] - xmss_per_leaf: usize, - /// SPHINCS signatures in each child, on top of the XMSS ones. - #[arg(long, default_value = "0")] - sphincs_per_leaf: usize, - /// Blobs in each child's LeanDA commitment. Use --xmss-per-leaf 0 for blobs alone. - #[arg( - long, - default_value_t = 0, - value_parser = clap::builder::RangedU64ValueParser::::new().range(0..=leanvm::lean_da::DA_MAX_ROWS as u64) - )] - blobs_per_leaf: usize, - }, - /// Prove and verify Fibonacci in the exponent (demo). + /// Prove and verify Fibonacci in the exponent. Fibonacci { /// Number of recurrence steps. #[arg(long, default_value = "2000000")] @@ -84,32 +50,11 @@ fn main() { let cli = Cli::parse(); lean_vm::init_prover(); let plan = primitives::bench::Plan::new(cli.repeat, cli.cooldown); - if cli.tracing && !matches!(&cli.command, Command::Recursion { .. }) { + if cli.tracing { primitives::init_tracing(); } match cli.command { - Command::Aggregate { xmss, sphincs, blobs } => { - rec_aggregation::run_aggregation(xmss, sphincs, blobs, cli.log_inv_rate, plan); - } - Command::Recursion { - n, - xmss_per_leaf, - sphincs_per_leaf, - blobs_per_leaf, - } => { - rec_aggregation::run_recursion( - n, - xmss_per_leaf, - sphincs_per_leaf, - blobs_per_leaf, - cli.log_inv_rate, - cli.tracing, - plan, - ); - } - Command::Fibonacci { n } => { - rec_aggregation::run_fibonacci(n, cli.log_inv_rate, plan); - } + Command::Fibonacci { n } => fibonacci::run_fibonacci(n, cli.log_inv_rate, plan), } if std::env::var_os("ZK_ALLOC_STATS").is_some() { eprintln!("{}", zk_alloc::stats()); diff --git a/tests/api.rs b/tests/api.rs index f15e8b429..dc772ed0e 100644 --- a/tests/api.rs +++ b/tests/api.rs @@ -1,125 +1,57 @@ use leanvm::*; -const EPOCH_0: xmss::Epoch = 7; -const EPOCH_1: xmss::Epoch = 9; -const EPOCH_2: xmss::Epoch = 11; -const MSG_0: xmss::Message = [0; xmss::MESSAGE_LEN]; -const MSG_1: xmss::Message = [1; xmss::MESSAGE_LEN]; -const MSG_2: xmss::Message = [2; xmss::MESSAGE_LEN]; +const STEPS: usize = 1000; + +/// Fibonacci in the exponent: `fib[g^k] = g^{F_k}`, the result published into `m[0]`. +fn fibonacci() -> (Program, [F64; 4]) { + let source = format!( + "def main():\n\ + \x20 fib = HeapBuf({size})\n\ + \x20 fib[1] = 1\n\ + \x20 fib[GEN] = GEN\n\ + \x20 for i in mul_range(1, GEN ** {STEPS}):\n\ + \x20 fib[i * GEN * GEN] = fib[i] * fib[i * GEN]\n\ + \x20 p = 1\n\ + \x20 p[1] = fib[GEN ** {last}]\n\ + \x20 return\n", + size = STEPS + 2, + last = STEPS + 1, + ); + let (mut previous, mut current) = (F64::ONE, g_pow(1)); + for _ in 0..STEPS { + (previous, current) = (current, previous * current); + } + ( + compile(&parse(&source).unwrap()), + [current, F64::ZERO, F64::ZERO, F64::ZERO], + ) +} #[test] fn public_api_end_to_end() { setup_prover(); - let rng = &mut rand::rng(); - - // 1. Eight XMSS signatures: three at the first (epoch, message), four at the second, one at the third. - let mut xmss_input = Vec::new(); - for (epoch, message, count) in [(EPOCH_0, MSG_0, 3), (EPOCH_1, MSG_1, 4), (EPOCH_2, MSG_2, 1)] { - for _ in 0..count { - let (secret_key, pub_key) = xmss::key_gen(rng, epoch, epoch).unwrap(); - let signature = xmss::sign(&secret_key, &message, epoch).unwrap(); - xmss_input.push((pub_key, epoch, message, signature)); - } - } - - // 2. Three SPHINCS signatures, each on its own message - let mut sphincs_input = Vec::new(); - for signer in 0..3u8 { - let (secret_key, pub_key) = sphincs::key_gen(rng); - let message = [signer; sphincs::MESSAGE_LEN]; - let signature = sphincs::sign(&secret_key, &message).unwrap(); - sphincs_input.push((pub_key, message, signature)); - } - - // 3. Two leaves, then a root over both. The leaves carry different epochs and the root's groups are their union. - let blobs: Vec<_> = (0..lean_da::BLOB_SYMBOLS).map(|i| i as u64).collect(); - let (commitment, _) = lean_da::commit(&blobs); - let left = aggregate( - &[], - xmss_input[..4].to_vec(), - sphincs_input[..1].to_vec(), - &blobs, - None, - 2, - ) - .unwrap(); - let left = EthereumProof::from_bytes(&left.to_bytes()).unwrap(); - left.verify().unwrap(); - assert_eq!(left.da_commitments(), &[commitment.root]); - let other_blobs: Vec<_> = blobs.iter().map(|x| x ^ 42).collect(); - let (other_commitment, _) = lean_da::commit(&other_blobs); - let right = aggregate( - &[], - xmss_input[4..7].to_vec(), - sphincs_input[1..].to_vec(), - &other_blobs, - None, - 2, - ) - .unwrap(); - let mut roots = vec![commitment.root, other_commitment.root]; - roots.sort(); - let root = aggregate(&[left, right], xmss_input[7..].to_vec(), vec![], &[], None, 2).unwrap(); - assert_eq!(root.num_signature_claims(), 11); - assert_eq!(root.da_commitments(), roots); - - // 4. Onto the wire, and back to a receiver, which checks the statement itself: - // verifying says these keys signed, the epochs and messages being the prover's. - let bytes = root.to_bytes(); - let received = EthereumProof::from_bytes(&bytes).unwrap(); - received.verify().unwrap(); - assert_eq!(received.da_commitments(), roots); - let pairs: Vec<_> = received - .xmss_signers() - .iter() - .map(|group| (group.epoch, group.message)) - .collect(); - assert_eq!(pairs, vec![(EPOCH_0, MSG_0), (EPOCH_1, MSG_1), (EPOCH_2, MSG_2)]); - - // 5. Removing some signatures from the aggregate: `declare` is what we keep. Here the first epoch group goes whole. - let mut groups = received.xmss_signers().to_vec(); - let mut sphincs_signers = received.sphincs_signers().to_vec(); - let dropped_group = groups.remove(0); - let dropped_signer = sphincs_signers.remove(0); - let retained_signatures = SignatureClaims { - xmss: groups, - sphincs: sphincs_signers, - }; - let narrowed = aggregate( - &[received], - vec![], - vec![], - &[], - Some(ClaimSelection { - signatures: &retained_signatures, - da_commitments: &[commitment.root], - }), - 2, - ) - .unwrap(); - narrowed.verify().unwrap(); - assert_eq!(narrowed.da_commitments(), &[commitment.root]); - assert_eq!(narrowed.num_signature_claims(), 11 - dropped_group.keys.len() - 1); - assert!( - !narrowed.xmss_signers().contains(&dropped_group), - "unpublished, epoch and message included" + let (program, public_input) = fibonacci(); + + // 1. Prove, then onto the wire and back to a receiver. + let (proof, _) = prove(&program, public_input, MIN_LOG_INV_RATE); + let bytes = bincode::serialize(&proof).unwrap(); + let received: Proof = bincode::deserialize(&bytes).unwrap(); + verify(&program, &public_input, &received).unwrap(); + + // 2. The proof is about this public input and no other. + let mut wrong_input = public_input; + wrong_input[0] += F64::ONE; + assert!(verify(&program, &wrong_input, &received).is_err()); + + // 3. One proof is one arena phase: the first proof outlives the second's phase. + let (second, _) = prove(&program, public_input, MIN_LOG_INV_RATE); + verify(&program, &public_input, &second).unwrap(); + verify(&program, &public_input, &received).unwrap(); + let stats = zk_alloc::stats(); + assert!(stats.phases >= 2, "expected one phase per proof, got {stats:?}"); + assert!(stats.peak_bytes > 0, "no buffer reached the arena: {stats:?}"); + assert_eq!( + stats.overflow, 0, + "a slab overflowed into the system allocator: {stats:?}" ); - assert!(!narrowed.sphincs_signers().contains(&dropped_signer)); - - let dropped = aggregate( - &[narrowed], - vec![], - vec![], - &[], - Some(ClaimSelection { - signatures: &retained_signatures, - da_commitments: &[], - }), - 2, - ) - .unwrap(); - dropped.verify().unwrap(); - assert!(dropped.da_commitments().is_empty()); - assert_eq!(dropped.xmss_signers(), retained_signatures.xmss); - assert_eq!(dropped.sphincs_signers(), retained_signatures.sphincs); } diff --git a/tests/no_arena.rs b/tests/no_arena.rs index 42356fad0..df993bd0c 100644 --- a/tests/no_arena.rs +++ b/tests/no_arena.rs @@ -2,25 +2,70 @@ //! one-way opt-in, so a test that must not have it cannot share a process with //! `tests/api.rs`. +use lean_vm::vmhash::compress; use leanvm::*; -const EPOCH: xmss::Epoch = 5; +/// An unrolled BLAKE2s hash chain from the zero block, the digest published into `m[0..4]`. +fn chain_source(steps: usize, unroll: usize) -> String { + assert!( + unroll >= 1 && steps.is_multiple_of(unroll), + "N must be a positive multiple of UNROLL" + ); + let blocks = steps / unroll; + let result_cell = 4 * blocks; + + let mut body = String::new(); + body.push_str(" b = i ** 4\n"); + body.push_str(" h0 = StackBuf(4)\n"); + body.push_str(" h0[0:4] = buff[b:b + 4]\n"); + for step in 1..=unroll { + body.push_str(&format!(" h{step} = StackBuf(4)\n")); + body.push_str(&format!( + " blake2s(h{previous}, h{previous}, h{step})\n", + previous = step - 1 + )); + } + body.push_str(" nxt = b * GEN ** 4\n"); + body.push_str(&format!(" buff[nxt:nxt + 4] = h{unroll}\n")); + + format!( + "def main():\n\ + \x20 buff = HeapBuf({size})\n\ + \x20 buff[0:4] = [0, 0, 0, 0]\n\ + \x20 for i in mul_range(1, GEN ** {blocks}):\n\ + {body}\ + \x20 output = 1\n\ + \x20 output[0:4] = buff[{result_cell}:{result_cell} + 4]\n\ + \x20 return\n", + size = result_cell + 4, + ) +} #[test] -fn aggregate_without_the_arena() { +fn blake2s_hash_chain_without_the_arena() { setup_prover_without_arena(); assert!(!zk_alloc::is_enabled(), "this path must leave the arena disengaged"); - let rng = &mut rand::rng(); - - let message = [3; xmss::MESSAGE_LEN]; - let signers = (0..2) - .map(|_| { - let (secret_key, pub_key) = xmss::key_gen(rng, EPOCH, EPOCH).unwrap(); - let signature = xmss::sign(&secret_key, &message, EPOCH).unwrap(); - (pub_key, EPOCH, message, signature) - }) - .collect(); - - let aggregated = aggregate(&[], signers, vec![], &[], None, 2).unwrap(); - aggregated.verify().unwrap(); + + let env_usize = |key: &str, default: usize| { + std::env::var(key) + .ok() + .and_then(|value| value.parse().ok()) + .unwrap_or(default) + }; + let unroll = env_usize("LEANVM_HASH_UNROLL", 4); + let steps = env_usize("LEANVM_HASH_N", 8); + + let mut public_input = [F64::ZERO; 4]; + for _ in 0..steps { + public_input = compress(public_input, public_input); + } + + let program = compile(&parse(&chain_source(steps, unroll)).expect("parse")); + let (proof, stats) = prove(&program, public_input, MIN_LOG_INV_RATE); + verify(&program, &public_input, &proof).expect("hash-chain proof verifies"); + assert_eq!(stats.base_counts[5], steps, "one BLAKE2S row per compression"); + + let mut wrong_input = public_input; + wrong_input[0] += F64::ONE; + assert!(verify(&program, &wrong_input, &proof).is_err()); } From 03fcc4e4a57c109c23d62e8d22b03448f9d720be Mon Sep 17 00:00:00 2001 From: Tom Wambsgans Date: Thu, 17 Sep 2026 13:11:30 -0400 Subject: [PATCH 06/40] Remove the recursion hooks from the kept crates The committed-size floor (min_log_committed and the filler floors), the verifier's recursion summary with its deferred bytecode claim, and the transcript's stream cursor existed only for the recursion harness. verify now returns unit, and verify_to_raw returns the expanded proof the Python verifier reads. Comments that pointed at the deleted guests are reworded. Co-Authored-By: Claude Fable 5.1 --- crates/fiat_shamir/src/lib.rs | 4 +- crates/fiat_shamir/src/merkle.rs | 12 +-- crates/fiat_shamir/src/transcript.rs | 19 ++--- crates/flock/src/hash.rs | 3 +- crates/flock/src/lincheck.rs | 3 +- crates/flock/tests/batch_proving_hashes.rs | 5 -- crates/lean_compiler/snark_lib.py | 3 +- .../tests/suite/const_placeholder.rs | 1 - crates/lean_compiler/tests/suite/filler.rs | 2 +- crates/lean_compiler/tests/suite/main.rs | 2 +- .../lean_compiler/tests/suite/range_check.rs | 3 +- crates/lean_compiler/tests/suite/sharing.rs | 2 +- .../lean_compiler/tests/suite/statements.rs | 4 +- crates/lean_compiler/zkDSL.md | 4 +- crates/lean_vm/src/constraints.rs | 5 +- crates/lean_vm/src/cpu/execute.rs | 77 +------------------ crates/lean_vm/src/cpu/filler.rs | 42 +++------- crates/lean_vm/src/cpu/hints.rs | 2 +- crates/lean_vm/src/cpu/layout.rs | 15 +--- crates/lean_vm/src/cpu/mod.rs | 66 ++++------------ crates/lean_vm/src/leaf.rs | 76 +----------------- crates/lean_vm/src/pcs.rs | 2 +- crates/lean_vm/src/vmhash.rs | 2 +- .../tests/verifiers/python_verifier.rs | 6 +- crates/pcs/src/ntt/additive_ntt_f64.rs | 3 +- crates/pcs/src/ring_switch.rs | 12 ++- crates/pcs/src/whir.rs | 2 +- crates/pcs/src/whir_config.rs | 9 +-- crates/pcs/src/whir_induce.rs | 4 +- crates/primitives/src/field/mod.rs | 2 +- crates/primitives/src/test_rng.rs | 4 +- 31 files changed, 79 insertions(+), 317 deletions(-) diff --git a/crates/fiat_shamir/src/lib.rs b/crates/fiat_shamir/src/lib.rs index 5dc63ec2c..e80f54e15 100644 --- a/crates/fiat_shamir/src/lib.rs +++ b/crates/fiat_shamir/src/lib.rs @@ -40,9 +40,7 @@ const DS_POW_NONCE: F64 = F64(4); /// `compress(base, (nonce.c0, nonce.c1, nonce.c2, DS_POW_NONCE))` has its low `bits` /// bits zero: the grinding predicate over the VM compression. A CONTIGUOUS -/// low-bit window (rather than byte-wise leading zeros) so a recursive verifier -/// re-checks it with a single loop over the bit decomposition of the digest word -/// (`grind_check` in `guests/lean_ethereum.py`). `bits` is always `< 64`. +/// low-bit window rather than byte-wise leading zeros. `bits` is always `< 64`. #[inline] fn pow_bits_ok(base: [F64; 4], nonce: F192, bits: u32) -> bool { debug_assert!(bits < 64, "grinding deficit fits the digest's low word"); diff --git a/crates/fiat_shamir/src/merkle.rs b/crates/fiat_shamir/src/merkle.rs index bdab343e7..bdd9a111e 100644 --- a/crates/fiat_shamir/src/merkle.rs +++ b/crates/fiat_shamir/src/merkle.rs @@ -8,8 +8,8 @@ pub type Hash = [u8; 32]; /// Encode a Merkle hash as the two field words transcripts carry it in: two /// 128-bit halves, each a K pair with a spare top lane. Every digest in the -/// protocol uses this one split (the commitment root, the public input, the -/// guest's MD state), so the VM sees one shape everywhere. +/// protocol uses this one split (the commitment root, the public input), so the +/// VM sees one shape everywhere. #[inline] pub fn hash_to_scalars(hash: &Hash) -> [F192; 2] { let word_at = |offset: usize| u64::from_le_bytes(hash[offset..offset + 8].try_into().unwrap()); @@ -241,12 +241,12 @@ impl PrunedMerklePaths { /// /// The redundant form. Several queries of one phase repeat whatever siblings /// they share, which is exactly what makes it simple to consume: recomputing -/// the root is a walk up one path, with no dedup bookkeeping. Recursive witness -/// construction and the Python verifier consume this; the wire format -/// ([`PrunedMerklePaths`]) sends each shared sibling once. +/// the root is a walk up one path, with no dedup bookkeeping. The Python +/// verifier consumes this; the wire format ([`PrunedMerklePaths`]) sends each +/// shared sibling once. #[derive(Clone, Debug, PartialEq, Eq, serde::Serialize, serde::Deserialize)] pub struct RawMerklePath { - /// Transcript-derived position, retained for recursive witness construction. + /// Transcript-derived position. pub leaf_index: usize, pub leaf_data: Vec, pub path: Vec, diff --git a/crates/fiat_shamir/src/transcript.rs b/crates/fiat_shamir/src/transcript.rs index fefba8478..8fc900c5c 100644 --- a/crates/fiat_shamir/src/transcript.rs +++ b/crates/fiat_shamir/src/transcript.rs @@ -11,11 +11,10 @@ pub struct Proof { pub merkle: Vec, } -/// The proof the recursion guest and the Python verifier consume: [`Proof`] with -/// every query's Merkle path written out, which is the one thing they would -/// otherwise have to reconstruct. A verifier run yields it as a by-product -/// ([`VerifierState::into_raw_proof`]), so that expansion is written once, in -/// Rust, instead of three times in three languages. +/// The proof the Python verifier consumes: [`Proof`] with every query's Merkle +/// path written out, which is the one thing it would otherwise have to +/// reconstruct. A verifier run yields it as a by-product +/// ([`VerifierState::into_raw_proof`]), so that expansion is written once, in Rust. pub type RawProof = Proof; #[derive(Clone, Copy, Debug, PartialEq, Eq)] @@ -44,8 +43,7 @@ pub trait Transmitter: Challenger { fn add_scalars(&mut self, xs: &[F192]); fn grind(&mut self, bits: u32); - /// Transmit a root as its two scalars, not as a byte string, so the recursion guest replays - /// one shape for every digest. Sending it is what binds it: no verifier absorbs a root + /// Transmit a root as its two scalars, not as a byte string. Sending it is what binds it: no verifier absorbs a root /// separately (see [`Receiver::next_root`], its mirror). fn add_root(&mut self, root: &Hash) { self.add_scalars(&hash_to_scalars(root)); @@ -200,13 +198,6 @@ impl<'a> VerifierState<'a> { } } - /// How many scalars have been read so far: the cursor into the stream a - /// caller needs to locate a sub-protocol's scalars without counting back - /// from the tail. - pub fn stream_offset(&self) -> usize { - self.offset - } - /// Assert the whole proof was consumed (no trailing/extra data). pub fn finish(&self) -> Result<(), Error> { if self.offset == self.stream.len() && self.phase == self.merkle.len() { diff --git a/crates/flock/src/hash.rs b/crates/flock/src/hash.rs index effca3f4a..801fd44c1 100644 --- a/crates/flock/src/hash.rs +++ b/crates/flock/src/hash.rs @@ -271,8 +271,7 @@ pub fn padding_block() -> Compression { /// commit that still had `build_matrices` and run `r1cs_digest_matches_baked`. /// /// The value is mirrored in `python-verifier/verifier.py`, which never could -/// rebuild the matrices, and in the recursion guest, so a deliberate circuit -/// change means bumping all three by hand. +/// rebuild the matrices, so a deliberate circuit change means bumping both by hand. pub const R1CS_DIGEST: [u8; 32] = [ 0x53, 0x7a, 0xd2, 0x07, 0x90, 0x30, 0x8f, 0x8e, 0xb8, 0xc0, 0xe8, 0xbd, 0x3e, 0x6c, 0x58, 0xee, 0x64, 0x57, 0x33, 0x71, 0xe3, 0xd5, 0x3c, 0x30, 0x61, 0x3d, 0xd0, 0x4d, 0x87, 0xc0, 0xb7, 0xea, diff --git a/crates/flock/src/lincheck.rs b/crates/flock/src/lincheck.rs index 72821d6b1..ab64d7079 100644 --- a/crates/flock/src/lincheck.rs +++ b/crates/flock/src/lincheck.rs @@ -1208,8 +1208,7 @@ pub fn verify( // The c term's `⟨eq_inner, w_col⟩`, by the tensor structure of both sides: // `eq_inner = eq(x_inner_rest) ⊗ λ(z_skip)` and `w_col = eq(r_inner_rest) ⊗ // z_partial`, so it is 8 eq factors times a 64-term Lagrange combination - // instead of a length-k inner product. That is the form the recursive - // verifier can afford. + // instead of a length-k inner product. let lambda_skip = lagrange_weights_naive(k_skip, x_ab.z_skip); let c_slice_value = lambda_skip .iter() diff --git a/crates/flock/tests/batch_proving_hashes.rs b/crates/flock/tests/batch_proving_hashes.rs index 96838c3b1..204a950e3 100644 --- a/crates/flock/tests/batch_proving_hashes.rs +++ b/crates/flock/tests/batch_proving_hashes.rs @@ -21,7 +21,6 @@ use primitives::{field::F64, pretty_integer, test_rng::Rng}; #[test] #[ignore = "manual release benchmark; needs a large-stack worker and substantial memory"] fn hash_batch_prove_verify() { - // The XMSS n=820 workload executes about 2^17 BLAKE2s compressions. let requested_n_log: usize = std::env::var("FLOCK_N_LOG") .ok() .map(|s| s.parse().expect("FLOCK_N_LOG must be an integer")) @@ -174,8 +173,4 @@ fn hash_batch_prove_verify() { pretty_integer(compressions_per_second), prove.spread() ); - println!( - " (~{:.1} XMSS/s equivalent at 146 compressions/signature)", - n as f64 / prove_s / 146.0 - ); } diff --git a/crates/lean_compiler/snark_lib.py b/crates/lean_compiler/snark_lib.py index 95ed8fd47..715685152 100644 --- a/crates/lean_compiler/snark_lib.py +++ b/crates/lean_compiler/snark_lib.py @@ -109,8 +109,7 @@ def hint_decompose_bits_exponent(bits, x, nbits: int) -> None: def hint_log2_ceil(bits, nbits: int, floor: int) -> _Elt: """Computed advice: returns `g^max(log2_ceil(v), floor)`, where `v` is the integer the `nbits`-cell `bits` buffer decodes to. The prover fills it at - witness-generation; it is UNCONSTRAINED, so the caller must verify it (see the - log2_ceil_in_the_exponent wrapper in the recursion guest). log2 = base-2 log of the integer, NOT the + witness-generation; it is UNCONSTRAINED, so the caller must verify it. log2 = base-2 log of the integer, NOT the discrete log base g that `log(...)` means.""" _ = bits, nbits, floor return _Elt() diff --git a/crates/lean_compiler/tests/suite/const_placeholder.rs b/crates/lean_compiler/tests/suite/const_placeholder.rs index 88237c4a7..9682dc916 100644 --- a/crates/lean_compiler/tests/suite/const_placeholder.rs +++ b/crates/lean_compiler/tests/suite/const_placeholder.rs @@ -151,7 +151,6 @@ def main(): assert a[0] == W return "; - // V = 42, W = 8, LOG_LIFETIME = 32 → the standard XMSS instance. let mut repl = BTreeMap::new(); repl.insert("V_PLACEHOLDER".to_string(), "42".to_string()); repl.insert("W_PLACEHOLDER".to_string(), "8".to_string()); diff --git a/crates/lean_compiler/tests/suite/filler.rs b/crates/lean_compiler/tests/suite/filler.rs index 64f410fcc..8763f5f4a 100644 --- a/crates/lean_compiler/tests/suite/filler.rs +++ b/crates/lean_compiler/tests/suite/filler.rs @@ -48,7 +48,7 @@ fn the_cost_model_is_exact() { for src in PROGRAMS { let program = compile(&parse(src).expect("parse")); let stats = prove(&program, [F64::ZERO; 4], lean_vm::pcs::TEST_LOG_INV_RATE).1; - let plan = filler::solve(stats.base_counts, filler::NO_FLOORS).expect("solvable"); + let plan = filler::solve(stats.base_counts).expect("solvable"); assert_eq!( filler::filled(stats.base_counts, &plan), stats.counts, diff --git a/crates/lean_compiler/tests/suite/main.rs b/crates/lean_compiler/tests/suite/main.rs index 54dec0b94..58a4e4251 100644 --- a/crates/lean_compiler/tests/suite/main.rs +++ b/crates/lean_compiler/tests/suite/main.rs @@ -1,7 +1,7 @@ //! Compiler integration tests share one binary and its initialization caches. //! //! These tests leave the proving arena disabled. A test that enables it needs -//! its own process (see `rec_aggregation`'s `arena_prove`). +//! its own process (see the root `tests/api.rs`). mod common; diff --git a/crates/lean_compiler/tests/suite/range_check.rs b/crates/lean_compiler/tests/suite/range_check.rs index e15c4bb47..6c8ce3727 100644 --- a/crates/lean_compiler/tests/suite/range_check.rs +++ b/crates/lean_compiler/tests/suite/range_check.rs @@ -158,8 +158,7 @@ fn range_check_without_log_rejected() { /// A *runtime* bound, `assert log x < log n`: the same gadget with `g^{k-1}` /// derived as `n·g^{-1}` instead of pooled from a constant. The bound rides a -/// hint here, as it does in the aggregation guest, where the signer count is -/// prover-announced. +/// hint here, as a prover-announced count would. #[test] fn range_check_runtime_bound() { let src = "\ diff --git a/crates/lean_compiler/tests/suite/sharing.rs b/crates/lean_compiler/tests/suite/sharing.rs index 6ea98aa0a..c38fc6058 100644 --- a/crates/lean_compiler/tests/suite/sharing.rs +++ b/crates/lean_compiler/tests/suite/sharing.rs @@ -13,7 +13,7 @@ use crate::common::pi; /// A returned value that repeats a constant computed earlier in the same /// function. The return slot lives in the callee frame and is read by the /// CALLER, so eliminating that write leaves the caller reading an unwritten -/// (prover-chosen) cell: `walk` in the XMSS guest returned a flag exactly this +/// (prover-chosen) cell: a chain-walk helper once returned a flag exactly this /// way, and folding it produced a proof whose caller-side assert failed. #[test] fn duplicate_constant_in_a_return_slot_survives() { diff --git a/crates/lean_compiler/tests/suite/statements.rs b/crates/lean_compiler/tests/suite/statements.rs index 2f70bae2b..bdef57093 100644 --- a/crates/lean_compiler/tests/suite/statements.rs +++ b/crates/lean_compiler/tests/suite/statements.rs @@ -608,8 +608,8 @@ fn a_match_target_binds_and_every_field_is_walked() { } /// The same construct in a VALUE position, for the same reason. `+` there is -/// XOR, so `lvl + 1` with `lvl = 3` is 2 and not 4, silently: the SPHINCS guest -/// could not write a Merkle level into a tweak and carried a generated table of +/// XOR, so `lvl + 1` with `lvl = 3` is 2 and not 4, silently: a program could +/// not write a Merkle level into a tweak and had to carry a generated table of /// one literal per level to get the integer reading instead. /// /// `const(e)` reads `e` with integer arithmetic and emits the literal, so one diff --git a/crates/lean_compiler/zkDSL.md b/crates/lean_compiler/zkDSL.md index 09dd30634..55e622cc2 100644 --- a/crates/lean_compiler/zkDSL.md +++ b/crates/lean_compiler/zkDSL.md @@ -325,7 +325,7 @@ The one dispatch construct. It matches the **log** of a g-power scrutinee agains Arms produce VALUES: every arm writes its results into the same cells, which is sound under write-once because exactly one arm runs. A target may be a name, bound after the join, or a **`StackBuf` element**, which the arms write into directly and which costs one instruction less than a name plus a store. The ABI returns into cells the CALLER picks, the same reason `sb[i] = f(x)` never needed a temporary, so reach for the element form wherever a returned value's home is a buffer slot. A target index must be a compile-time integer inside the buffer, both errors naming the line; a `HeapBuf` element is not a target, its cells not being frame cells. Multiple targets take a multi-return call as the arm body. A run return binds a name, and crosses the join only through the fused dispatch below. -A branch body with statements in it goes in a function, and the arm calls it: that is the idiom the recursion guest uses throughout (`lambda k: walk(chain_start, tweaks, pp, k)`), and it names the body instead of inlining it. Where the arms only PRODUCE values, as there, this costs nothing. Where each arm's real work is a WRITE, it costs: the writer function needs a return value and the statement a target, both dead, so the natural translation costs more instructions than a body inlined into the dispatching frame. An arm may pass runs to its callee, but a run parameter is the callee's copy, so a store into it asserts against the caller's cells rather than filling them. If that shape matters to a program, dispatch on a value and write after the join. +A branch body with statements in it goes in a function, and the arm calls it: that is the idiom to use throughout (`lambda k: walk(chain_start, tweaks, pp, k)`), and it names the body instead of inlining it. Where the arms only PRODUCE values, as there, this costs nothing. Where each arm's real work is a WRITE, it costs: the writer function needs a return value and the statement a target, both dead, so the natural translation costs more instructions than a body inlined into the dispatching frame. An arm may pass runs to its callee, but a run parameter is the callee's copy, so a store into it asserts against the caller's cells rather than filling them. If that shape matters to a program, dispatch on a value and write after the join. **Lowering** is two jumps through a *trampoline table* in the bytecode: the dispatch jumps to `g^T · x²`, the j-th two-instruction slot (`SET` the arm's address, `JUMP` to it) of a table at base `T`, and the slot jumps to the arm, which can sit anywhere, unaligned and of any length. Cost is about 7 cycles, independent of the arm count. @@ -427,7 +427,7 @@ def square(v: StackBuf(3)): - **A run and a scalar never stand in for each other.** A word where a run is expected, a run where a word is expected (`+`, `*`, `assert`, `print`), a run of the wrong width, and a 192-bit builtin used as a statement are compile errors naming the widths. - **Moving runs.** `buf[lo:hi] = value` stores one, an `x: StackBuf(3)` parameter takes one, a function returns one, and a fused `match` returns one. A copy between frame runs is one `XOR192` against the pooled zero run. -A slice of a digest is a run like any other, so the recursion guest takes a transcript challenge as the first three words of its BLAKE2s state, `state[0:3]`. +A slice of a digest is a run like any other, so a program can take the first three words of a BLAKE2s state as `state[0:3]`. ## BLAKE2s diff --git a/crates/lean_vm/src/constraints.rs b/crates/lean_vm/src/constraints.rs index 53772b270..ab0f993f1 100644 --- a/crates/lean_vm/src/constraints.rs +++ b/crates/lean_vm/src/constraints.rs @@ -24,12 +24,11 @@ //! `eq(ζ_m, Y)·p(Y) + Y·u`. Its constant, quadratic and cubic coefficients are //! sent; the running claim fixes the linear coefficient. The verifier evaluates //! it at the challenge. Heights and `ζ` enter only the per-table `weights`, which -//! may be accumulated along the way or deferred to the end, as the recursion guest does. +//! may be accumulated along the way or deferred to the end. //! //! The eq point is the caller's, not a fresh one (the bus's GKR point `ζ`), which //! is what lets the forms' sums settle the bus. Batching derived in `doc/leanvm/main.tex` -//! §sec:air. Both sides take `n = max τ_t` from the announced heights; a recursive -//! verifier certifies that maximum with one hinted `g`-power (§recursion), so there +//! §sec:air. Both sides take `n = max τ_t` from the announced heights, so there //! are no rounds in which no table has joined. use crate::PAR_THRESHOLD; diff --git a/crates/lean_vm/src/cpu/execute.rs b/crates/lean_vm/src/cpu/execute.rs index 7382ca7a7..6d444fe7b 100644 --- a/crates/lean_vm/src/cpu/execute.rs +++ b/crates/lean_vm/src/cpu/execute.rs @@ -67,79 +67,6 @@ impl Program { /// first four memory cells (§sec:e2e-pi). Compilation yields the `Program`; /// executing it (here) and proving it are separate later phases. pub fn execute(&self, public_input: [F64; 4]) -> Execution { - self.execute_filled(public_input, super::filler::NO_FLOORS) - } - - /// [`Self::execute`], then grow one table until the stacked witness reaches - /// [`Program::min_log_committed`]. - /// - /// The height is solved for, not approached: the padded table is a small - /// share of the committed total, so doubling it moves `m` by much less than - /// a bit, and a geometric search would neither converge in a predictable - /// number of rounds nor be monotone in the request. `committed_log` is pure - /// and cheap, so the smallest height that reaches the floor is found without - /// executing anything; one re-run then realises it, and the loop only exists - /// because the fill's own closing jumps and frames feed back into the size. - /// Runs that already clear the floor (every one of consequence) execute once. - pub(crate) fn execute_to_floor(&self, public_input: [F64; 4]) -> Execution { - let mut exec = self.execute_filled(public_input, super::filler::NO_FLOORS); - if self.min_log_committed == 0 { - return exec; - } - let mut floors = super::filler::NO_FLOORS; - for _ in 0..3 { - if self.committed_log(&exec) >= self.min_log_committed { - return exec; - } - floors[super::filler::PAD_TABLE] = 1 << self.padded_height(&exec); - exec = self.execute_filled(public_input, floors); - } - assert!( - self.committed_log(&exec) >= self.min_log_committed, - "no fill reaches the requested 2^{} committed witness", - self.min_log_committed - ); - exec - } - - /// The smallest height for the padded table that takes this run's committed - /// witness to [`Program::min_log_committed`]. `m` is nondecreasing in every - /// table height, so the first one that reaches the floor is the smallest. - fn padded_height(&self, exec: &Execution) -> usize { - let mut taus = exec.trace.row_counts().map(crate::log2_strict_usize); - let log_bytecode = crate::log2_strict_usize(self.prog.len()); - let log_mem = crate::log2_strict_usize(exec.mem.len()); - let pad = super::filler::PAD_TABLE; - for height in taus[pad] + 1..=crate::cpu::MAX_LOG_ROWS { - taus[pad] = height; - if super::layout::committed_log(log_mem, log_bytecode, taus) >= self.min_log_committed { - return height; - } - } - panic!( - "even a full {} table cannot reach 2^{}", - crate::cpu::MAX_LOG_ROWS, - self.min_log_committed - ) - } - - /// The stacked witness this run commits to, as a log2. - fn committed_log(&self, exec: &Execution) -> usize { - super::layout::committed_log( - crate::log2_strict_usize(exec.mem.len()), - crate::log2_strict_usize(self.prog.len()), - exec.trace.row_counts().map(crate::log2_strict_usize), - ) - } - - /// [`Self::execute`] with a floor on some tables' row counts, which the fill - /// blocks buy on top of their natural powers of two - /// ([`Program::min_log_committed`]). - pub(crate) fn execute_filled( - &self, - public_input: [F64; 4], - fill_floors: [usize; crate::tables::N_TABLES], - ) -> Execution { // One interpretation of the program, then the fill. The blocks that bring every // table to a power of two are cycles no program code enters (`cpu::filler`), so // they run after the chain has halted, by which point the program's own row @@ -324,7 +251,7 @@ impl Program { // Bounded discrete log for `hint_decompose_bits_exponent`: find n < 2^nbits // with g^n = x, by baby-step giant-step (baby table g^j for j < 2^17, // built once per run; giant step ×g^(-2^17)). Prover-side only: the - // guest re-verifies the hinted bits in-circuit. + // program re-verifies the hinted bits itself. fn bounded_dlog(cache: &mut Option<(GPow, F64)>, x: F64, nbits: u32) -> u128 { const LOG_BABY: u32 = 17; let (baby, giant) = cache.get_or_insert_with(|| { @@ -403,7 +330,7 @@ impl Program { // powers of two by itself, which `Layout` checks. if !self.filler.is_empty() { let mut frame = (1u32 << crate::cpu::MIN_LOG_MEM).max(next_free); - for (block_pc, size, n) in super::filler::cycles(&self.filler, counts, fill_floors) { + for (block_pc, size, n) in super::filler::cycles(&self.filler, counts) { g.grow_to(frame as usize); g.note(frame as usize); // What the closing jump reads: back to the block's own first diff --git a/crates/lean_vm/src/cpu/filler.rs b/crates/lean_vm/src/cpu/filler.rs index d8c2b8fd5..d9e98108d 100644 --- a/crates/lean_vm/src/cpu/filler.rs +++ b/crates/lean_vm/src/cpu/filler.rs @@ -25,14 +25,6 @@ use crate::tables::N_TABLES; -/// No table asked to grow past its natural power of two. -pub const NO_FLOORS: [usize; N_TABLES] = [0; N_TABLES]; - -/// The table a committed-size floor grows ([`crate::cpu::Program::min_log_committed`]). -/// `SET` writes one cell from an immediate: no operand reads, no second memory -/// touch, and no precompile behind it, so its rows are the cheapest to prove. -pub const PAD_TABLE: usize = 2; - /// Block sizes, largest first: a fill of `f` rows takes `f / 128` traversals of the /// largest block and then one per set bit of the remainder. pub const SIZES: [usize; 8] = [128, 64, 32, 16, 8, 4, 2, 1]; @@ -147,26 +139,23 @@ fn decompose_jump(gap: usize) -> Option<[usize; SIZES.len()]> { (left == 0).then_some(out) } -/// A plan taking every table from `base` to an exact power of two, at or above -/// `floors[t]`, or `None` if some table cannot be filled at all. +/// A plan taking every table from `base` to an exact power of two, or `None` if some +/// table cannot be filled at all. /// /// Every table but `JUMP` is independent: its fill is the distance to its next power of /// two, decomposed into traversals. `JUMP` is not, because every traversal of the whole /// fill lands a row there, its own traversals included. Counting those first makes it a /// single decomposition rather than a fixpoint. -/// -/// A floor above a table's natural target is how a run too small to be provable at the -/// size its consumer needs buys the difference: see [`crate::cpu::Program::min_log_committed`]. -pub fn solve(base: [usize; N_TABLES], floors: [usize; N_TABLES]) -> Option { +pub fn solve(base: [usize; N_TABLES]) -> Option { let mut plan: Plan = [[0; SIZES.len()]; N_TABLES]; for t in 0..N_TABLES { if t != JUMP { - plan[t] = decompose(ceil_pow2(base[t].max(MIN_ROWS[t]).max(floors[t])) - base[t]); + plan[t] = decompose(ceil_pow2(base[t].max(MIN_ROWS[t])) - base[t]); } } // What `JUMP` already owes: its own rows, plus one per traversal so far. let owed = base[JUMP] + traversals(&plan); - let mut target = ceil_pow2(owed.max(MIN_ROWS[JUMP]).max(floors[JUMP])); + let mut target = ceil_pow2(owed.max(MIN_ROWS[JUMP])); loop { if let Some(jump_steps) = decompose_jump(target - owed) { plan[JUMP] = jump_steps; @@ -180,8 +169,8 @@ pub fn solve(base: [usize; N_TABLES], floors: [usize; N_TABLES]) -> Option /// The cycles a run needs, in the order the interpreter should walk them: for each, the /// block's first pc, its size, and how many times to traverse it. Panics if `blocks` is /// missing one the plan calls for, which can only mean bytecode the compiler did not emit. -pub fn cycles(blocks: &[Block], base: [usize; N_TABLES], floors: [usize; N_TABLES]) -> Vec<(u32, u32, usize)> { - let plan = solve(base, floors).unwrap_or_else(|| panic!("no fill plan from {base:?}")); +pub fn cycles(blocks: &[Block], base: [usize; N_TABLES]) -> Vec<(u32, u32, usize)> { + let plan = solve(base).unwrap_or_else(|| panic!("no fill plan from {base:?}")); let mut out = Vec::new(); for (t, row) in plan.iter().enumerate() { for (k, &n) in row.iter().enumerate() { @@ -237,7 +226,7 @@ mod tests { [0, 0, 0, 0, 1 << 20, 0, 0, 0], ]; for base in cases { - let plan = solve(base, NO_FLOORS).unwrap_or_else(|| panic!("no plan for {base:?}")); + let plan = solve(base).unwrap_or_else(|| panic!("no plan for {base:?}")); let got = filled(base, &plan); assert!(is_filled(got), "{base:?} filled to {got:?}"); for t in 0..N_TABLES { @@ -246,26 +235,13 @@ mod tests { } } - /// A floor takes a table past its natural power of two, and every other table - /// still lands on one: what a run too small for its consumer buys with. - #[test] - fn floor_grows_target_table() { - let base = [1_000, 2_000, 3_000, 4_000, 500, 8, 300, 200]; - let mut floors = NO_FLOORS; - floors[PAD_TABLE] = 1 << 16; - let plan = solve(base, floors).expect("solvable"); - let got = filled(base, &plan); - assert!(is_filled(got), "{got:?}"); - assert_eq!(got[PAD_TABLE], 1 << 16); - } - /// The bulk of a fill rides the largest block, so the fill stays cheap: one closing /// jump per 128 rows, plus at most one traversal per size per table for the /// remainders. #[test] fn fill_uses_bulk_blocks() { let base = [125_000, 286_000, 341_000, 508_000, 114_000, 130_000, 90_000, 70_000]; - let plan = solve(base, NO_FLOORS).expect("solvable"); + let plan = solve(base).expect("solvable"); let fill: usize = delivered(&plan).iter().sum(); assert!( traversals(&plan) <= fill / SIZES[0] + SIZES.len() * N_TABLES, diff --git a/crates/lean_vm/src/cpu/hints.rs b/crates/lean_vm/src/cpu/hints.rs index c5ddc04d6..c05eb1fbb 100644 --- a/crates/lean_vm/src/cpu/hints.rs +++ b/crates/lean_vm/src/cpu/hints.rs @@ -8,7 +8,7 @@ use primitives::field::F64; pub type Off = u32; /// The `g^k` table paired with a reverse index `g^k ↦ k`, both grown on demand -/// (recursion depth, and so the address range, is unbounded). +/// (call depth, and so the address range, is unbounded). /// /// The interpreter needs both directions: a cell index becomes the address `g^k`, /// and a pointer word read back out of memory must be inverted to the index it diff --git a/crates/lean_vm/src/cpu/layout.rs b/crates/lean_vm/src/cpu/layout.rs index d20113029..9fcf5c767 100644 --- a/crates/lean_vm/src/cpu/layout.rs +++ b/crates/lean_vm/src/cpu/layout.rs @@ -129,9 +129,7 @@ impl Witness { } } -/// The committed columns' kappa SOURCES, for the recursion guest's -/// in-circuit certification of the stacked size m = max(log2_ceil(sum of -/// 2^kappa), MIN_MU). Per committed column: `Some((source, adj))` with +/// The committed columns' kappa SOURCES. Per committed column: `Some((source, adj))` with /// kappa = value(source) + adj, where source 0 is the constant 0 (kappa = /// adj; used for the fixed-size columns and the program bytecode length, /// which the caller passes as `log_bytecode`), source 1 is log_mem, and @@ -167,8 +165,7 @@ pub fn col_kappa_sources(log_bytecode: usize) -> Vec> { /// The bus flush blocks' kappa SOURCES, flattened in side order (push, pull, /// count) exactly as the blocks are constructed below: per block /// `(source, adj)` with kappa = value(source) + adj, source 0 = the constant -/// 0, 1 = log_mem, 2 + t = tau_t. For the recursion guest's in-circuit pin -/// of every hinted block kappa. Keep in lockstep with the block +/// 0, 1 = log_mem, 2 + t = tau_t. Keep in lockstep with the block /// construction in [`fn@layout`]. pub fn block_kappa_sources(log_bytecode: usize) -> Vec<(usize, usize)> { let mut push = vec![(0, 0), (1, 0), (0, log_bytecode)]; @@ -200,14 +197,6 @@ fn col_kappas(log_mem: usize, log_bytecode: usize, taus: [usize; tables::N_TABLE .collect() } -/// `log2` of the stacked witness the announced sizes imply. The one part of -/// [`layout`] a size floor needs, and far cheaper than the rest of it. -pub(crate) fn committed_log(log_mem: usize, log_bytecode: usize, taus: [usize; tables::N_TABLES]) -> usize { - crate::witness::placements_of(&col_kappas(log_mem, log_bytecode, taus)) - .1 - .mu -} - /// Build the public [`Layout`] from the program, the memory log-size `log_mem`, the /// instruction tables' log heights `taus`, and the public input `pi`. The flush blocks /// reference columns only by INDEX and the program only through its public columns, so diff --git a/crates/lean_vm/src/cpu/mod.rs b/crates/lean_vm/src/cpu/mod.rs index 7cf311a44..f1c4e6d81 100644 --- a/crates/lean_vm/src/cpu/mod.rs +++ b/crates/lean_vm/src/cpu/mod.rs @@ -75,8 +75,7 @@ const MAX_LOG_BYTECODE: usize = 32; /// /// The IV IS the transcript's starting chaining value ([`fiat_shamir::FiatShamirState::new`]), /// so all challenges depend on the circuit version and the program before -/// anything else; a recursion guest carries the INNER program's IV in its public -/// input, pinning both with one digest. +/// anything else. pub fn fs_seed(program: &Program) -> [F64; 4] { let mut h = primitives::hash::Hasher::new(); h.update(b"leanvm"); @@ -139,8 +138,8 @@ fn read_public(vs: &mut VerifierState, prog: &Program, public_input: &[F64; 4]) || taus.iter().any(|&t| t > MAX_LOG_ROWS) // flock sizes its argument to at least `n_blocks_log(1)` instances, and the // BLAKE2s table's value columns share that instance cube, so a height below the - // floor describes a layout the arithmetization cannot express. The other two - // verifiers reject it here too (`python-verifier`, `guests/lean_ethereum.py`). + // floor describes a layout the arithmetization cannot express. `python-verifier` + // rejects it here too. || taus[tables::BLAKE2S_TABLE] < crate::hash_flock::n_blocks_log(1) || ::pcs::whir::validate_log_inv_rate(log_inv_rate).is_err() { @@ -186,21 +185,10 @@ pub struct Program { /// diagnostic; empty for hand-assembled programs. pub fn_ranges: Vec<(String, u32, u32)>, /// Source line of the statement that emitted each pc. Prover-side only, and - /// outside `bytecode_hash`, so it costs nothing in the proof: a failed guest + /// outside `bytecode_hash`, so it costs nothing in the proof: a failed /// check reports a line rather than a pc to disassemble around. Empty for a /// hand-assembled program, and shorter than `prog`, which is padded. pub src_lines: Vec, - /// The smallest stacked witness this program's proofs may commit to, as a - /// log2. Zero (the default) asks for nothing. - /// - /// A consumer can need a proof to be no smaller than some size even when the - /// run is: the recursion guest holds one WHIR opening arm per committed size - /// it was compiled for, and has none below the first. A run that falls short - /// buys the difference in fill rows ([`crate::cpu::filler`]) rather than in a - /// padded commitment, which keeps the committed size a function of the - /// announced table heights, so neither the verifier nor the guest needs a new - /// parameter to certify. Prover-side only. - pub min_log_committed: usize, } /// The bytecode digest reinterprets the stacked table as bytes, which is its @@ -229,7 +217,6 @@ impl Program { filler: Vec::new(), fn_ranges: Vec::new(), src_lines: Vec::new(), - min_log_committed: 0, } } @@ -484,7 +471,7 @@ pub fn prove(program: &Program, public_input: [F64; 4], log_inv_rate: usize) -> // The returned `Proof` is system-allocated (`ps.into_proof()` builds `Vec`s), // so it survives the next phase. let _phase = zk_alloc::enter_phase(); - let exec = crate::stage!("Execute program", || program.execute_to_floor(public_input)); + let exec = crate::stage!("Execute program", || program.execute(public_input)); // A live value that came from outside the constraint system means the emitted // bytecode asserts less than its source asked for, so the proof would be about a // weaker statement than the program text. That is a compiler bug and never a @@ -623,33 +610,22 @@ fn bind_pi_claim(r: [F192; 2], placements: &[witness::Placement], pi: &[F64; 4]) } } -/// Everything a recursion harness needs from an accepting verify run, named -/// and typed: the deferred bytecode claim, flock's -/// reduction claims, and the stacked-opening summary (ring-switch challenges + -/// WHIR fold/query data). The sub-proof scalars themselves live on -/// `proof.stream`, ending at `flock_stream_end`. Ordinary callers just -/// `?`-discard it. -pub struct VerifySummary { - /// Transcript-bound inverse-rate logarithm used by this proof's PCS. - pub log_inv_rate: usize, - pub bytecode_claim: leaf::BytecodeClaim, - pub zc_claim: flock::zerocheck::ZerocheckClaim, - pub lc_claim: flock::lincheck::LincheckClaim, - /// Stream cursor just after flock's reduction, i.e. where the PCS opening's - /// own scalars start. The recursion harness reads flock's lincheck tail - /// from here rather than counting back from the end of the stream. - pub flock_stream_end: usize, - /// The proof this run just verified, in the unpruned form the recursion - /// guest and the Python verifier consume. - pub raw: fiat_shamir::transcript::RawProof, -} - /// Verify a proof against the public statement (program + public input): replay /// the transcript, reconstruct the public layout from the announced sizes, read /// every scalar the prover wrote and pull the PCS hints, then assert the stream /// was fully consumed. Takes only public inputs, never the prover's witness. +pub fn verify(program: &Program, public_input: &[F64; 4], proof: &Proof) -> Result<(), CpuError> { + verify_to_raw(program, public_input, proof).map(|_| ()) +} + +/// [`verify`], returning the proof it accepted with every query's Merkle path +/// written out, the form `python-verifier` reads. #[tracing::instrument(name = "Verify", skip_all)] -pub fn verify(program: &Program, public_input: &[F64; 4], proof: &Proof) -> Result { +pub fn verify_to_raw( + program: &Program, + public_input: &[F64; 4], + proof: &Proof, +) -> Result { let mut vs = VerifierState::new(fs_seed(program), proof, *public_input); let (l, log_inv_rate) = read_public(&mut vs, program, public_input)?; let root = pcs::read_commitment(&mut vs).map_err(CpuError::Transcript)?; @@ -694,18 +670,10 @@ pub fn verify(program: &Program, public_input: &[F64; 4], proof: &Proof) -> Resu let n_blocks = n_blake2s.max(1); let offset = l.placements[QFLOCK].offset; let replay = crate::hash_flock::verify_reduction(n_blocks, &mut vs).map_err(CpuError::Blake2s)?; - let flock_stream_end = vs.stream_offset(); let ring = crate::hash_flock::ring_switch_verify(n_blocks, offset, &replay.claim); pcs::verify(&mut vs, &slots, &ring, l.shape, log_inv_rate, &root).map_err(CpuError::Open)?; vs.finish().map_err(CpuError::Transcript)?; - Ok(VerifySummary { - bytecode_claim: bus.bytecode_claim, - zc_claim: replay.zc_claim, - lc_claim: replay.lc_claim, - log_inv_rate, - flock_stream_end, - raw: vs.into_raw_proof(), - }) + Ok(vs.into_raw_proof()) } /// Lift `ColumnClaim`s to located PCS claims: a claim on column `c` lives in diff --git a/crates/lean_vm/src/leaf.rs b/crates/lean_vm/src/leaf.rs index d9cfa91f2..7e603431e 100644 --- a/crates/lean_vm/src/leaf.rs +++ b/crates/lean_vm/src/leaf.rs @@ -576,20 +576,6 @@ fn decompose_verify( }) } -/// One reduced claim on the bytecode polynomial. The eight public encoding -/// columns (opcode plus seven operand/immediate slots), padded to sixteen slots -/// along four selector bits, form one multilinear polynomial B̃ in `κ_bc + 4` -/// variables. The native verifier combines its column evaluations at ζ with -/// the bus weights `eq(α⃗, ·)`, giving `B̃(ζ_lo, α⃗)`. The recursive verifier -/// defers this claim to its public input. -#[derive(Clone, Debug)] -pub struct BytecodeClaim { - /// `ζ_side_lo ++ s`, a point in `κ_bc + 4` variables. - pub point: Vec, - /// `B̃(point)`. - pub value: F192, -} - /// Selector bits of the stacked bytecode polynomial: the public encoding /// columns (opcode + seven operand/immediate slots = eight) stack along /// `2^N_BYTECODE_SELECTORS` slots. A column's slot is its bus tuple coordinate, @@ -602,9 +588,8 @@ pub const N_BYTECODE_SELECTORS: usize = 4; pub const BYTECODE_PUBLIC_SLOT: usize = 3; /// The stacked bytecode polynomial as a dense table: eight public encoding -/// columns at their tuple coordinates, padded to sixteen selector slots. This is -/// the polynomial [`BytecodeClaim`]s are claims about; the outermost verifier -/// evaluates it. +/// columns at their tuple coordinates, padded to sixteen selector slots. The +/// program's digest is taken over it ([`crate::cpu::Program`]). pub fn stacked_bytecode_table(blocks: &[Block]) -> Vec { let mut kbc = 0; let mut cols: Vec<&[F64]> = Vec::new(); @@ -644,33 +629,6 @@ fn sides<'a>( ] } -/// The program's whole share of a bus leaf, in ONE evaluation: a public column's -/// slot is its tuple coordinate and the weights are `eq(α⃗, ·)`, so the weighted -/// sum over the columns IS the stacked polynomial at `(ζ, α⃗)` (§sec:e2e-bc). -fn bytecode_claim(blocks: &[Block], point: &[F192], alphas: &[F192], public: &mut PublicEvals) -> BytecodeClaim { - let weights = fingerprint_weights(alphas); - let mut kbc = 0; - let mut slot = BYTECODE_PUBLIC_SLOT; - let mut value = F192::ZERO; - for blk in blocks { - for c in &blk.coords { - if let Coord::Public(vals) = c { - if slot == BYTECODE_PUBLIC_SLOT { - kbc = blk.kappa; - } - assert_eq!(vals.len(), 1 << kbc); - value += weights[slot] * public_eval(vals, &point[..kbc], public); - slot += 1; - } - } - } - let claim_point = [&point[..kbc], alphas].concat(); - BytecodeClaim { - value, - point: claim_point, - } -} - /// Prove the bus balances; returns the per-column claims to open (§sec:leafstack). `alpha`/ /// `beta` follow the witness commitment (the only ordering the grand product /// needs), and the block structure is public, so no shape is observed. @@ -873,12 +831,10 @@ fn tables_and_prods_at( .unzip() } -/// What [`verify_balance`] establishes: the per-column claims to open, the -/// reduced bytecode claim (push and pull share ζ), -/// and the table forms with their claimed sums. +/// What [`verify_balance`] establishes: the per-column claims to open and the +/// table forms with their claimed sums. pub struct BusVerify { pub claims: Vec, - pub bytecode_claim: BytecodeClaim, /// The GKR point ζ, reused as the table sumcheck's eq point. pub point: Vec, /// `forms[side][table]`, for the zerocheck to settle. @@ -957,7 +913,6 @@ pub fn verify_balance( Ok(BusVerify { claims, - bytecode_claim: bytecode_claim(push, &bus_gkr.point, &alphas, &mut public), point: bus_gkr.point, forms, totals, @@ -968,29 +923,6 @@ pub fn verify_balance( mod tests { use super::soundness_bits; - #[test] - fn bytecode_claim_matches_dense_stacking() { - use super::*; - - let columns: Vec<_> = (0..8) - .map(|col| Arc::new((0..32).map(|i| F64((i + 1) * (col + 1))).collect())) - .collect(); - let blocks = [Block { - kappa: 5, - coords: columns.iter().cloned().map(Coord::Public).collect(), - }]; - let point: Vec<_> = (0..5).map(|i| F192::new(i + 2, i + 17, i + 23)).collect(); - let alphas: Vec<_> = (0..N_TUPLE_BITS).map(|i| F192::new(i as u64 + 5, 3, 7)).collect(); - let table = stacked_bytecode_table(&blocks); - let mut public = PublicEvals::new(); - for col in columns.iter().rev() { - public_eval(col, &point, &mut public); - } - let claim = bytecode_claim(&blocks, &point, &alphas, &mut public); - assert_eq!(claim.point, [point, alphas].concat()); - assert_eq!(claim.value, mle_eval(&table, &claim.point)); - } - /// The bound is `(N_TUPLE_BITS + 1)·2^mu` plus the GKR terms: only the bus /// DEPTH costs bits now, the multilinear fingerprint having fixed each factor's /// degree at four in `α⃗` and one in `β`, whatever the tuple's width. diff --git a/crates/lean_vm/src/pcs.rs b/crates/lean_vm/src/pcs.rs index a13f5516b..6c407af13 100644 --- a/crates/lean_vm/src/pcs.rs +++ b/crates/lean_vm/src/pcs.rs @@ -47,7 +47,7 @@ pub use ::pcs::whir::{MAX_LOG_INV_RATE, MIN_LOG_INV_RATE}; const _: () = assert!(::pcs::whir::SECURITY_BITS == crate::SECURITY_BITS as usize); /// Minimum committed-witness log-size accepted by the WHIR level ladder, with one level of margin. pub const MIN_MU: usize = 15; -/// Largest committed size accepted by all verifiers and compiled into the recursion guest. +/// Largest committed size accepted by all verifiers. pub const MAX_MU: usize = 28; /// The shared WHIR config for a `2^μ`-word witness, memoized per `(μ, log_inv_rate)`. diff --git a/crates/lean_vm/src/vmhash.rs b/crates/lean_vm/src/vmhash.rs index 800007085..babd8aa64 100644 --- a/crates/lean_vm/src/vmhash.rs +++ b/crates/lean_vm/src/vmhash.rs @@ -1,7 +1,7 @@ //! VM-provable standard BLAKE2s hashing. //! //! The VM instruction exposes one BLAKE2s compression whose chaining value, byte -//! counter and finalization flags are all memory-supplied. A guest can therefore +//! counter and finalization flags are all memory-supplied. A program can therefore //! implement the standard sequential compression chain for any input length, //! including a zero-padded partial final block, and the length need not be known //! when the program is compiled. diff --git a/crates/lean_vm/tests/verifiers/python_verifier.rs b/crates/lean_vm/tests/verifiers/python_verifier.rs index 1bf2e0e93..55d150b12 100644 --- a/crates/lean_vm/tests/verifiers/python_verifier.rs +++ b/crates/lean_vm/tests/verifiers/python_verifier.rs @@ -5,7 +5,7 @@ use fiat_shamir::transcript::RawProof; use lean_compiler::{compile, parse_with_replacements}; -use lean_vm::cpu::{prove, verify}; +use lean_vm::cpu::{prove, verify, verify_to_raw}; use primitives::field::{F64, g_pow}; use std::collections::BTreeMap; use std::path::Path; @@ -104,9 +104,7 @@ fn test_python_verifier() { // Python reads the RAW proof: same protocol, each query carrying its own // full Merkle path instead of one octopus over the batch. A Rust verify // expands the wire form, so the pruning is written once. - let raw = verify(&program, &public_input, &proof) - .expect("honest proof verifies") - .raw; + let raw = verify_to_raw(&program, &public_input, &proof).expect("honest proof verifies"); let directory = std::env::temp_dir().join(format!("leanvm-python-verifier-test-{}", std::process::id())); std::fs::create_dir_all(&directory).expect("create test directory"); diff --git a/crates/pcs/src/ntt/additive_ntt_f64.rs b/crates/pcs/src/ntt/additive_ntt_f64.rs index 108eca73e..0202ecbcc 100644 --- a/crates/pcs/src/ntt/additive_ntt_f64.rs +++ b/crates/pcs/src/ntt/additive_ntt_f64.rs @@ -188,8 +188,7 @@ impl AdditiveNttF64 { /// `log_inv_rate` IS one replica, so a block's eight participating rows are eight /// message rows, and the pass can gather them itself instead of reading back a /// codeword someone else just filled. That turns three sweeps of the whole - /// codeword (fill it, read it, write it) into one gather and one write: at the - /// XMSS scale, three gigabytes moved instead of seven. Replica 0 IS the message, + /// codeword (fill it, read it, write it) into one gather and one write. Replica 0 IS the message, /// so the blocks run in descending order and it is transformed in place last, /// once every other replica has read it. /// diff --git a/crates/pcs/src/ring_switch.rs b/crates/pcs/src/ring_switch.rs index f03c23872..4a754f99f 100644 --- a/crates/pcs/src/ring_switch.rs +++ b/crates/pcs/src/ring_switch.rs @@ -96,11 +96,10 @@ pub const COMPOSITION_SHIFTS: [usize; 6] = [32, 16, 8, 4, 2, 1]; /// `sum_w weights[w]·t_w = sum_j x^j·Phi(y_j)` for the row /// view `y` and the column view `t` of the same tensor-algebra element. That /// identity is what lets a verifier evaluate the batched claim from `Phi`'s -/// six challenges instead of the 192 weights; the recursion guest does exactly -/// that. Expanding the composition puts a distinct monomial at every Frobenius +/// six challenges instead of the 192 weights. Expanding the composition puts a distinct monomial at every Frobenius /// exponent `0..64`: writing `k = sum_p k_p·2^(5-p)` for the binary digits of /// `k`, the coefficient is `C_k = prod_{p : k_p = 1} f_p^(2^(k mod 2^(5-p)))`, -/// which is what the guest's coefficient table builds. Applying the composed +/// which is what `python-verifier`'s coefficient table builds. Applying the composed /// form directly costs only 63 squarings and six multiplications. /// /// ## Soundness @@ -667,10 +666,9 @@ mod tests { } /// The contract between the native opener and every verifier that batches - /// from `Phi`'s coefficients (the recursion guest, the Python reference - /// verifier): weighting the COLUMN view by `build_coordinate_weights` must + /// from `Phi`'s coefficients (the Python reference verifier): weighting the COLUMN view by `build_coordinate_weights` must /// equal applying `Phi` to the ROW view and combining with `x^j`. If this - /// drifts, the guest computes a different opening target than the prover. + /// drifts, that verifier computes a different opening target than the prover. #[test] fn column_weights_match_the_row_side_linearized_map() { let mut rng = Rng::new(0xF00D_BEEF_1234_5678); @@ -682,7 +680,7 @@ mod tests { let columns = transpose_s_hat(&s_hat_v); let lhs = inner_product_base_ext(&columns, &build_coordinate_weights(&challenges)); - // Row side: sum_j x^j * Phi(y_j), the guest's loop. + // Row side: sum_j x^j * Phi(y_j). let x = F192::new(2, 0, 0); let mut rhs = F192::ZERO; let mut x_pow = F192::ONE; diff --git a/crates/pcs/src/whir.rs b/crates/pcs/src/whir.rs index c806faa8c..4757d2c60 100644 --- a/crates/pcs/src/whir.rs +++ b/crates/pcs/src/whir.rs @@ -1082,7 +1082,7 @@ impl<'a> SumcheckProver<'a> { /// Sample `count` query positions in transcript order: no dedup, no sort. /// `block_len = 2^d`; each squeezed field element yields `⌊192/d⌋` positions as -/// its disjoint d-bit chunks (low bits first), matching the scheme the recursive guest derives (fixed `192/d` per +/// its disjoint d-bit chunks (low bits first) (fixed `192/d` per /// squeeze, dup-tolerant: soundness matches the deployed PCS with the same /// `config.queries`). Duplicates are harmless, a repeated position re-opens the /// same Merkle-authenticated row. diff --git a/crates/pcs/src/whir_config.rs b/crates/pcs/src/whir_config.rs index 0058561d5..3fc2a441d 100644 --- a/crates/pcs/src/whir_config.rs +++ b/crates/pcs/src/whir_config.rs @@ -85,10 +85,9 @@ const _: () = assert!(RS_DOMAIN_SUBSEQUENT_REDUCTION_FACTOR <= SUBSEQUENT_FOLDIN /// clear instead of committed and folded further. pub const RESIDUAL_MAX_LOG: usize = 5; -// The recursion guest rotates the terminal point left by the lane fold to index it -// by witness coordinate, and the residual segment is what the last lane challenges -// rotate past, so the residual may never be longer than that fold -// (`rec_aggregation`'s placeholder emitter re-checks this per derived candidate). +// A verifier rotates the terminal point left by the lane fold to index it by +// witness coordinate, and the residual segment is what the last lane challenges +// rotate past, so the residual may never be longer than that fold. const _: () = assert!(RESIDUAL_MAX_LOG <= INITIAL_FOLDING_FACTOR); /// Shape plus per-level soundness parameters for one WHIR opening. Prover @@ -119,7 +118,7 @@ pub type VerifierConfig = ProverConfig; /// The per-level shape table a [`VerifierConfig`] implies for a /// `log_n`-variable opening: the numbers every consumer of the multilevel -/// protocol (the verifier itself, recursion harnesses) otherwise re-derives. +/// protocol otherwise re-derives. #[derive(Clone, Debug)] pub struct LevelShapes { /// Level count (`level_steps + 1`). diff --git a/crates/pcs/src/whir_induce.rs b/crates/pcs/src/whir_induce.rs index 271a7fcf0..c34b18982 100644 --- a/crates/pcs/src/whir_induce.rs +++ b/crates/pcs/src/whir_induce.rs @@ -26,9 +26,7 @@ fn next_s(s: F64, s_at_root: F64) -> F64 { } /// `sks_vks[k] = s_k(v_k)` for `k = 0..=log_n`, over K. Mirror of -/// `whir::eval_sk_at_vks`. Public for the recursion harness, which dumps -/// these vanishing-polynomial values as guest hints (base-field, embedded into -/// the tower with both extension limbs zero). +/// `whir::eval_sk_at_vks`. pub fn eval_sk_at_vks(log_n: usize) -> Vec { let mut sks_vks = vec![F64::ZERO; log_n + 1]; sks_vks[0] = F64::ONE; diff --git a/crates/primitives/src/field/mod.rs b/crates/primitives/src/field/mod.rs index b4aa6b6db..f77cbdb65 100644 --- a/crates/primitives/src/field/mod.rs +++ b/crates/primitives/src/field/mod.rs @@ -96,7 +96,7 @@ pub fn g_pow(i: usize) -> F64 { /// The fixed generator `g = x ∈ K`, with `ord(g) = 2^64 - 1` (pinned by a /// field test), larger than every index any admissible /// instance uses (the verifier's instance caps, §cpu). For `k < 64`, `g^k` is -/// the monomial `x^k` (bit `k`), which the XMSS encoding check relies on. +/// the monomial `x^k` (bit `k`). pub const G: F64 = F64::G; /// MLE of the index column `[g^0, …, g^{2^n−1}]` over the `n`-variable cube, diff --git a/crates/primitives/src/test_rng.rs b/crates/primitives/src/test_rng.rs index 1531e018a..1f62c008c 100644 --- a/crates/primitives/src/test_rng.rs +++ b/crates/primitives/src/test_rng.rs @@ -2,8 +2,8 @@ //! kernel and its scalar oracle see the same stream on every host. //! //! Exported normally rather than under `#[cfg(test)]`: integration tests in -//! `pcs`, `flock`, `primitives` and `rec_aggregation` are separate crates and -//! cannot reach a test-only module. +//! `pcs`, `flock` and `primitives` are separate crates and cannot reach a +//! test-only module. use rand::{RngCore, SeedableRng, rngs::StdRng}; From 01e6fbecad6eafcc19f81cee714db3f9425a4e84 Mon Sep 17 00:00:00 2001 From: Tom Wambsgans Date: Thu, 17 Sep 2026 13:11:47 -0400 Subject: [PATCH 07/40] Remove the signature specifications and their Lean proofs Drops doc/xmss, doc/sphincs, formal/, the Lean CI workflow, the XMSS and SPHINCS steps of the documentation workflow, and the cross-branch hash benchmark script that drove the aggregate and recursion commands. formal/ is ignored so the local Lean build cache stays out of git status. Co-Authored-By: Claude Fable 5.1 --- .github/workflows/doc.yml | 36 - .github/workflows/lean.yml | 73 - .gitignore | 1 + doc/bench_blake3_vs_sha2_vs_sha3.sh | 161 -- doc/sphincs/latexmkrc | 4 - doc/sphincs/main.tex | 471 --- doc/sphincs/refs.bib | 100 - doc/xmss/latexmkrc | 4 - doc/xmss/main.tex | 343 --- doc/xmss/refs.bib | 124 - formal/sphincs/.gitignore | 3 - formal/sphincs/PROOF.md | 130 - formal/sphincs/README.md | 43 - formal/sphincs/SphincsSecurity.lean | 34 - .../sphincs/SphincsSecurity/Completeness.lean | 54 - .../Completeness/Assembly.lean | 81 - .../SphincsSecurity/Completeness/Code.lean | 271 -- .../SphincsSecurity/Completeness/Counter.lean | 49 - .../SphincsSecurity/Completeness/Decay.lean | 130 - .../SphincsSecurity/Completeness/Digest.lean | 201 -- .../Completeness/Encoding.lean | 59 - .../SphincsSecurity/Completeness/Fresh.lean | 576 ---- .../SphincsSecurity/Completeness/Game.lean | 122 - .../SphincsSecurity/Completeness/Keygen.lean | 92 - .../Completeness/Recovery.lean | 645 ----- .../SphincsSecurity/Completeness/Search.lean | 209 -- .../SphincsSecurity/Completeness/Signing.lean | 128 - formal/sphincs/SphincsSecurity/Proof.lean | 1 - .../Proof/Adversary/Embedding.lean | 37 - .../Proof/Adversary/Security.lean | 19 - .../Proof/Base/BernoulliExcessMoments.lean | 31 - .../Proof/Base/BinomialMoments.lean | 315 -- .../Base/FinitePmfProductObservation.lean | 92 - .../Proof/Base/FirstSuccessFamily.lean | 81 - .../Proof/Base/FirstSuccessPrefix.lean | 130 - .../Proof/Base/FirstSuccessTable.lean | 212 -- .../Base/FourthMomentExceptionBound.lean | 34 - .../SphincsSecurity/Proof/Base/Prelude.lean | 34 - .../SphincsSecurity/Proof/Base/QueryCap.lean | 143 - .../Proof/Base/QueryCapAccounting.lean | 80 - .../Proof/Base/QueryCapBalance.lean | 56 - .../Proof/Base/QueryCapErasure.lean | 62 - .../Proof/Base/QueryCapState.lean | 56 - .../Proof/Base/QueryPause.lean | 44 - .../Proof/Base/QueryPauseInvariant.lean | 74 - .../Proof/Base/QueryPauseTrace.lean | 56 - .../Proof/Base/QueryTraceInvariant.lean | 52 - .../Proof/Base/QueryTracePotential.lean | 105 - .../Proof/Base/RomQueryCharge.lean | 150 - .../Proof/Base/RomQueryChargeBind.lean | 26 - .../Proof/Base/StirlingMomentBounds.lean | 40 - .../SphincsSecurity/Proof/Base/TraceSum.lean | 60 - .../Proof/Base/UniformTableCompletion.lean | 184 -- .../Proof/Base/UniformTableConditioning.lean | 41 - .../Proof/Base/UniformTableDisclosure.lean | 42 - .../Proof/Base/UniformTableJoin.lean | 104 - .../Proof/Base/UniformTableObservation.lean | 90 - .../Base/UniformTableObservationErasure.lean | 72 - .../Proof/Base/UniformTableOverwrite.lean | 41 - .../Proof/Base/UniformTableProducts.lean | 52 - .../Proof/Base/UniformTableRestriction.lean | 98 - .../Proof/Base/UniformTableSplit.lean | 111 - .../Proof/Base/WeightedQuery.lean | 66 - .../Chains/AdaptiveChainActualBudget.lean | 60 - .../Proof/Chains/AdaptiveChainAuxiliary.lean | 47 - .../Chains/AdaptiveChainBudgetPotential.lean | 47 - .../Proof/Chains/AdaptiveChainCap.lean | 92 - .../Proof/Chains/AdaptiveChainCapContact.lean | 105 - .../Proof/Chains/AdaptiveChainCapCost.lean | 68 - .../Chains/AdaptiveChainCapObservation.lean | 137 - .../Proof/Chains/AdaptiveChainCapTwoEdge.lean | 66 - .../Proof/Chains/AdaptiveChainCheckpoint.lean | 152 - .../AdaptiveChainCheckpointContact.lean | 90 - .../AdaptiveChainCheckpointProjection.lean | 52 - .../Chains/AdaptiveChainCompensation.lean | 48 - .../Proof/Chains/AdaptiveChainContact.lean | 56 - .../Chains/AdaptiveChainContactMoments.lean | 108 - .../Chains/AdaptiveChainCountedRows.lean | 45 - .../Proof/Chains/AdaptiveChainEndpoint.lean | 98 - .../Proof/Chains/AdaptiveChainErasure.lean | 52 - .../Proof/Chains/AdaptiveChainLikelihood.lean | 61 - .../Chains/AdaptiveChainObservation.lean | 91 - .../Proof/Chains/AdaptiveChainPotential.lean | 60 - .../Proof/Chains/AdaptiveChainQueryBound.lean | 75 - .../Proof/Chains/AdaptiveChainRestart.lean | 56 - .../Proof/Chains/AdaptiveChainSupport.lean | 35 - .../Proof/Chains/AdaptiveChainTwoEdge.lean | 126 - .../Proof/Chains/EndpointPreimageDensity.lean | 95 - .../Proof/Chains/PartialChainCompletion.lean | 158 - .../Chains/PartialChainContactCount.lean | 111 - .../Chains/PartialChainContactPotential.lean | 88 - .../Proof/Chains/PartialChainEndpoint.lean | 110 - .../Proof/Chains/PartialChainLastRow.lean | 117 - .../Chains/PartialChainLikelihoodLower.lean | 82 - .../Chains/PartialChainLongCompletion.lean | 146 - .../Proof/Chains/PartialChainObservation.lean | 76 - .../Proof/Chains/PartialChainPreparation.lean | 110 - .../Proof/Chains/PartialChainRowCharge.lean | 151 - .../Proof/Chains/PartialChainSuffix.lean | 128 - .../Proof/Chains/PartialChainTwoEdge.lean | 129 - .../Chains/PartialChainTwoEdgeCharge.lean | 103 - .../PartialChainTwoEdgeCompensation.lean | 110 - .../Proof/ClosingParameters.lean | 22 - .../Proof/Deterministic/CostState.lean | 53 - .../Proof/Deterministic/DerivationTable.lean | 84 - .../Proof/Deterministic/FreshRequests.lean | 219 -- .../Proof/Deterministic/GameComparison.lean | 56 - .../Proof/Deterministic/GameExpansion.lean | 114 - .../Proof/Deterministic/Inputs.lean | 69 - .../Proof/Deterministic/KeygenBudget.lean | 95 - .../Proof/Deterministic/LoggedSigning.lean | 46 - .../Proof/Deterministic/MemoErasure.lean | 83 - .../Proof/Deterministic/MemoGame.lean | 50 - .../Proof/Deterministic/MemoLog.lean | 118 - .../Proof/Deterministic/MemoTable.lean | 57 - .../Proof/Deterministic/Memoize.lean | 114 - .../Proof/Deterministic/Preparation.lean | 86 - .../Deterministic/ReferenceDistribution.lean | 77 - .../Proof/Deterministic/ReferenceSource.lean | 72 - .../Proof/Deterministic/Replay.lean | 90 - .../Proof/Deterministic/RequestSampling.lean | 150 - .../Proof/Deterministic/Security.lean | 34 - .../Proof/Deterministic/SignerSampling.lean | 68 - .../Proof/Deterministic/TableSigner.lean | 94 - .../Proof/Deterministic/TableToReference.lean | 75 - .../Deterministic/TranscriptReduction.lean | 81 - .../Proof/Deterministic/TrialLoop.lean | 73 - .../Proof/Deterministic/TrialSampling.lean | 133 - .../Forced/FtsGuessAuxiliaryProgram.lean | 149 - .../Proof/Forced/FtsGuessBudget.lean | 158 - .../Proof/Forced/FtsGuessCachedForced.lean | 228 -- .../Proof/Forced/FtsGuessCachedMessage.lean | 80 - .../Proof/Forced/FtsGuessCachedSigning.lean | 354 --- .../Proof/Forced/FtsGuessDeferredMessage.lean | 67 - .../Proof/Forced/FtsGuessDeferredSeed.lean | 96 - .../Proof/Forced/FtsGuessEventTransfer.lean | 117 - .../Proof/Forced/FtsGuessExceptionBounds.lean | 291 -- .../FtsGuessExceptionClassification.lean | 298 -- .../Forced/FtsGuessExceptionHistory.lean | 341 --- .../Forced/FtsGuessExceptionWeights.lean | 479 ---- .../Proof/Forced/FtsGuessHash.lean | 148 - .../Proof/Forced/FtsGuessInputCoverage.lean | 461 --- .../Forced/FtsGuessMonitoredAccounting.lean | 420 --- .../Forced/FtsGuessMonitoredPayment.lean | 356 --- .../Forced/FtsGuessMonitoredPotential.lean | 403 --- .../Proof/Forced/FtsGuessMonitoredStep.lean | 270 -- .../Proof/Forced/FtsGuessMonitoredValid.lean | 358 --- .../Proof/Forced/FtsGuessNearAssembly.lean | 386 --- .../Proof/Forced/FtsGuessNearDeferred.lean | 75 - .../Proof/Forced/FtsGuessNearEvent.lean | 69 - .../Proof/Forced/FtsGuessNearSource.lean | 164 -- .../Proof/Forced/FtsGuessNearWitness.lean | 125 - .../Proof/Forced/FtsGuessPairBound.lean | 60 - .../Proof/Forced/FtsGuessPairSource.lean | 117 - .../Proof/Forced/FtsGuessProgram.lean | 187 -- .../Forced/FtsGuessProposalCompleted.lean | 266 -- .../Forced/FtsGuessProposalInvariant.lean | 233 -- .../Proof/Forced/FtsGuessProposalPayment.lean | 263 -- .../Proof/Forced/FtsGuessProposalStep.lean | 413 --- .../Proof/Forced/FtsGuessReference.lean | 122 - .../Forced/FtsGuessReferenceWitness.lean | 110 - .../Proof/Forced/FtsGuessRemaining.lean | 65 - .../Proof/Forced/FtsGuessSeedCache.lean | 78 - .../Proof/Forced/FtsGuessSigning.lean | 117 - .../Proof/Forced/FtsGuessTracking.lean | 213 -- .../Proof/Forced/FtsGuessWitness.lean | 147 - .../Proof/Forced/FtsGuessWork.lean | 183 -- .../Proof/Forced/SecretGuessErasure.lean | 71 - .../Proof/Forced/SecretGuessForceBound.lean | 115 - .../Forced/SecretGuessForceLikelihood.lean | 122 - .../Proof/Forced/SecretGuessForced.lean | 79 - .../Forced/SecretGuessForcedProgram.lean | 105 - .../Proof/Forced/SecretGuessHitPayoff.lean | 67 - .../Proof/Forced/SecretGuessHitRun.lean | 104 - .../Proof/Forced/SecretGuessHitSum.lean | 148 - .../Proof/Forced/SecretGuessObservation.lean | 479 ---- .../Proof/Forced/SecretGuessPairBound.lean | 145 - .../Security127SmallBudgetArithmetic.lean | 80 - .../Proof/Fts/AdaptiveHiddenHazard.lean | 70 - .../Proof/Fts/AdaptiveProposalWords.lean | 44 - .../Proof/Fts/BankedCacheWeight.lean | 171 -- .../Proof/Fts/BankedProposalStep.lean | 162 -- .../Proof/Fts/BankedTargetEnvelope.lean | 220 -- .../Proof/Fts/BoundaryCertificateCache.lean | 125 - .../Proof/Fts/CacheIndexMultiplicity.lean | 21 - .../Proof/Fts/CacheMessageSignerWeight.lean | 51 - .../Proof/Fts/CacheMessageWeight.lean | 158 - .../SphincsSecurity/Proof/Fts/CacheSize.lean | 53 - .../Proof/Fts/CachedDigestRate.lean | 68 - .../Proof/Fts/CachedDigestSelection.lean | 87 - .../Fts/CachedIndexExcessConcentration.lean | 54 - .../Proof/Fts/CachedIndexExcessScore.lean | 122 - .../Proof/Fts/CachedIndexHashMoments.lean | 111 - .../Proof/Fts/CachedSigningViews.lean | 48 - .../Proof/Fts/CachedTargetIncrement.lean | 35 - .../Proof/Fts/CachedTargetSubsetMatch.lean | 32 - .../Fts/CanonicalCoordinateSampling.lean | 124 - .../Proof/Fts/CanonicalHiddenCoordinates.lean | 131 - .../Fts/CertificateBankCompleteness.lean | 101 - .../Fts/CertificateBoundaryInvariants.lean | 107 - .../Fts/CertificateCacheExceptionGame.lean | 36 - .../Fts/CertificateCacheExceptionGrowth.lean | 66 - .../Fts/CertificateCacheExceptionKernels.lean | 81 - .../CertificateCacheExceptionPotential.lean | 52 - .../Proof/Fts/CertificateCacheMonitor.lean | 100 - .../Fts/CertificateCachePersistence.lean | 69 - .../Proof/Fts/CertificateCleanExecution.lean | 163 -- .../Proof/Fts/CertificateCleanGame.lean | 83 - .../Proof/Fts/CertificateGame.lean | 125 - .../Proof/Fts/CertificateJointExceptions.lean | 24 - .../Proof/Fts/CertificateMessagePayment.lean | 144 - .../Proof/Fts/CertificateMonitor.lean | 247 -- .../Proof/Fts/CertificateMonitorStep.lean | 221 -- .../Fts/CertificateOriginalMessageCost.lean | 78 - .../Proof/Fts/CertificatePathBudget.lean | 89 - .../Fts/CertificateProposalInvariant.lean | 249 -- .../CertificateProposalPrefixException.lean | 23 - .../Proof/Fts/CertificateStoppedState.lean | 47 - .../Proof/Fts/CertificateTerminalGame.lean | 98 - .../Proof/Fts/ConcreteTargetShapeQuery.lean | 80 - .../Proof/Fts/ConcreteTargetShapeSigning.lean | 86 - .../Proof/Fts/DigestAttemptExpectation.lean | 74 - .../Proof/Fts/DigestCompletionBank.lean | 93 - .../Fts/DigestCompletionCacheGrowth.lean | 209 -- .../Proof/Fts/DigestCompletionLogGrowth.lean | 104 - .../Proof/Fts/DigestCompletionNewTarget.lean | 124 - .../Proof/Fts/DigestLoopRecord.lean | 13 - .../Proof/Fts/DigestMessageCost.lean | 127 - .../Proof/Fts/DigestSelectionIndex.lean | 172 -- .../Proof/Fts/DigestSelectionMass.lean | 60 - .../Proof/Fts/DigestSelectionWeight.lean | 251 -- .../Proof/Fts/DigestSigningCompletion.lean | 83 - .../Proof/Fts/ExactTargetShapeSigning.lean | 187 -- .../SphincsSecurity/Proof/Fts/ExtractFts.lean | 140 - .../Proof/Fts/FewTimeConditionalCoverage.lean | 15 - .../Proof/Fts/FewTimeFixedPrehit.lean | 55 - .../Proof/Fts/FewTimeFresh.lean | 231 -- .../Proof/Fts/FewTimeFreshMass.lean | 125 - .../Proof/Fts/FewTimeLoop.lean | 42 - .../Proof/Fts/FewTimePadding.lean | 49 - .../Proof/Fts/FewTimePrehit.lean | 126 - .../Proof/Fts/FewTimeProbability.lean | 25 - .../Proof/Fts/FewTimeRace.lean | 111 - .../Proof/Fts/FewTimeSignerView.lean | 88 - .../Proof/Fts/FewTimeSource.lean | 32 - .../Proof/Fts/FewTimeTargetCompletion.lean | 182 -- .../Proof/Fts/FewTimeUniform.lean | 352 --- .../Proof/Fts/FewTimeWeightedOriginRace.lean | 42 - .../Proof/Fts/FewTimeWitness.lean | 30 - .../Proof/Fts/FixedCertificateCoverage.lean | 92 - .../Proof/Fts/FixedProposalMoments.lean | 59 - .../Proof/Fts/FreshDigestHazard.lean | 50 - .../Proof/Fts/FreshTargetEnvelope.lean | 53 - .../Proof/Fts/FreshTargetPayload.lean | 37 - .../Proof/Fts/FreshTargetShapeAverage.lean | 37 - .../Proof/Fts/FtsProbeAdversary.lean | 61 - .../Proof/Fts/FtsProbeGame.lean | 34 - .../Proof/Fts/FtsProbeOrigin.lean | 27 - .../Proof/Fts/FtsProbeProbability.lean | 36 - .../Proof/Fts/FtsProbeSampling.lean | 63 - .../Proof/Fts/FtsProbeSimulation.lean | 65 - .../Proof/Fts/FtsProbeVerifierSource.lean | 64 - .../Proof/Fts/FtsVerifierWitness.lean | 111 - .../Proof/Fts/FutureCoverageBound.lean | 35 - .../Proof/Fts/HiddenLabelObservation.lean | 106 - .../Proof/Fts/HiddenLabelProbe.lean | 98 - .../Proof/Fts/InterleavedCoverStep.lean | 50 - .../Fts/InterleavedResidualDisclosure.lean | 82 - .../Proof/Fts/JointProbeMessageAnswers.lean | 11 - .../Proof/Fts/JointProbeMessageReserve.lean | 13 - .../Proof/Fts/MessageAdmissibleDeficit.lean | 79 - .../Proof/Fts/MessageByteTrace.lean | 110 - .../Proof/Fts/MessageCacheCountGrowth.lean | 51 - .../Proof/Fts/MessageCacheProjection.lean | 74 - .../Fts/MessageCertificateProjection.lean | 74 - .../Proof/Fts/MessageDeficitCacheGrowth.lean | 93 - .../Fts/MessageDeficitConcentration.lean | 42 - .../Proof/Fts/MessageDeficitHashMoments.lean | 114 - .../Proof/Fts/MessageDeficitMomentGrowth.lean | 49 - .../Proof/Fts/MessageDeficitScore.lean | 51 - .../Proof/Fts/MessageDigestHazard.lean | 124 - .../Proof/Fts/MessageInputMiss.lean | 67 - .../Proof/Fts/MessageNormalizedReuse.lean | 60 - .../Proof/Fts/MessagePrehit.lean | 126 - .../Proof/Fts/NearCertificateBound.lean | 45 - .../Proof/Fts/NewTargetEnvelopeCharge.lean | 36 - .../Proof/Fts/NormalizedTargetCacheQuery.lean | 89 - .../Proof/Fts/NormalizedTargetLogSigning.lean | 70 - .../Proof/Fts/NormalizedTargetMatches.lean | 66 - .../Proof/Fts/ObservedAdaptiveCoverBound.lean | 20 - .../Proof/Fts/ObservedFreshCoverBound.lean | 35 - .../Proof/Fts/ObservedOccupancy.lean | 49 - .../Fts/OriginalCacheExceptionBound.lean | 297 -- .../Proof/Fts/OriginalCertificateBound.lean | 209 -- .../Proof/Fts/OriginalCertificateTrace.lean | 105 - .../Proof/Fts/OriginalMessageAllocation.lean | 98 - .../Proof/Fts/OriginalProposalBudget.lean | 88 - .../Proof/Fts/OriginalProposalExecution.lean | 195 -- .../Fts/OriginalProposalPrefixBound.lean | 155 - .../Proof/Fts/OriginalTerminalProposal.lean | 71 - .../Proof/Fts/PairedHiddenMiss.lean | 61 - .../SphincsSecurity/Proof/Fts/Parameters.lean | 43 - .../Proof/Fts/PrimitiveMessagePotential.lean | 147 - .../Proof/Fts/ProposalBridgeKernel.lean | 200 -- .../Proof/Fts/ProposalLengthProjection.lean | 89 - .../Proof/Fts/ProposalPrefixExponential.lean | 127 - .../Proof/Fts/ProposalPrefixStop.lean | 14 - .../Proof/Fts/ProposalQueryProjection.lean | 174 -- .../Proof/Fts/ProposalWordDistribution.lean | 125 - .../Proof/Fts/PublicSigningRecord.lean | 85 - .../Proof/Fts/RawProposalMomentBound.lean | 143 - .../Proof/Fts/RawSigningMomentBound.lean | 193 -- .../Proof/Fts/ReferenceFtsCoverage.lean | 110 - .../Proof/Fts/ReuseCachedTargets.lean | 150 - .../Proof/Fts/ReuseRawEnvelope.lean | 30 - .../Proof/Fts/ReuseTargetEnvelope.lean | 92 - .../Proof/Fts/SelectedDigestCache.lean | 42 - .../Proof/Fts/SignerAdmissibleMessage.lean | 91 - .../Proof/Fts/SignerDigestSource.lean | 58 - .../Proof/Fts/SigningProposalRecord.lean | 168 -- .../Proof/Fts/SingleMessageCacheGrowth.lean | 58 - .../Proof/Fts/SourceGroupExpectation.lean | 117 - .../Proof/Fts/StoppedSigningLog.lean | 26 - .../Proof/Fts/SubsetTargetAssignment.lean | 34 - .../Proof/Fts/SubsetTargetExpectation.lean | 65 - .../Proof/Fts/TargetAssignmentCount.lean | 22 - .../Proof/Fts/TargetIndexEnvelope.lean | 71 - .../Fts/TargetMixedGrowthPolynomial.lean | 91 - .../Proof/Fts/TargetMomentShapes.lean | 47 - .../Proof/Fts/TargetShapeBlocks.lean | 83 - .../Proof/Fts/TargetShapeCardinality.lean | 64 - .../Proof/Fts/TargetShapeContinuation.lean | 54 - .../Proof/Fts/TargetShapeEnvelope.lean | 85 - .../Proof/Fts/TargetShapeExpectation.lean | 115 - .../Proof/Fts/TargetShapeOperators.lean | 106 - .../Proof/Fts/TargetShapeReindex.lean | 66 - .../Proof/Fts/TargetSigningMatchFactors.lean | 86 - .../Proof/Fts/TargetSourceMultiplicity.lean | 73 - .../Proof/Fts/TerminalCache.lean | 28 - .../Proof/Fts/TerminalCertificateCharge.lean | 227 -- .../Proof/Fts/TerminalProposalEnvelope.lean | 108 - .../Proof/Fts/TerminalProposalWord.lean | 79 - .../Fts/UniformProposalMixedMoments.lean | 223 -- .../Proof/Fts/UniformProposalMoments.lean | 149 - .../Proof/Fts/UniformProposalVariance.lean | 72 - .../Proof/Fts/UniformPublicCoordinates.lean | 89 - .../Proof/Fts/UnitCertificateCoverage.lean | 137 - .../Proof/Fts/UnrestrictedRowSwap.lean | 125 - .../Proof/Fts/UpperDigestSelection.lean | 68 - .../Proof/Fts/ValidInterleavedCover.lean | 12 - .../Proof/Fts/WeightedTargetGroups.lean | 51 - .../Proof/Hypertree/CanonicalGraph.lean | 142 - .../Proof/Hypertree/CanonicalGraphGame.lean | 124 - .../Proof/Hypertree/CanonicalGraphHonest.lean | 143 - .../Hypertree/CanonicalGraphResidual.lean | 113 - .../Hypertree/CanonicalGraphSampling.lean | 118 - .../Proof/Hypertree/CanonicalProbeCache.lean | 86 - .../Hypertree/CanonicalProbeRouting.lean | 266 -- .../Proof/Hypertree/CanonicalPublicPrior.lean | 78 - .../Hypertree/CanonicalSigningFrontier.lean | 100 - .../Proof/Hypertree/Descent.lean | 208 -- .../Proof/Hypertree/Extract.lean | 130 - .../Proof/Hypertree/FiniteGraphReplay.lean | 105 - .../Proof/Hypertree/FiniteGraphSampling.lean | 95 - .../Hypertree/FrontierEncodingCongruence.lean | 124 - .../Hypertree/FrontierGameProjection.lean | 138 - .../Hypertree/FrontierOracleCongruence.lean | 112 - .../Proof/Hypertree/FrontierOracleMask.lean | 75 - .../Proof/Hypertree/FrontierRandomOracle.lean | 102 - .../Hypertree/FrontierSignerErasure.lean | 153 - .../Hypertree/FrontierSigningEvaluation.lean | 82 - .../FrontierSigningOracleCongruence.lean | 74 - .../Hypertree/FrontierSigningOrigin.lean | 140 - .../Hypertree/FrontierTreeEvaluation.lean | 127 - .../Proof/Hypertree/GraphPayloadInputs.lean | 32 - .../Proof/Hypertree/Honest.lean | 248 -- .../Proof/Hypertree/Hypertree.lean | 35 - .../Proof/Hypertree/Parameters.lean | 32 - .../Proof/Hypertree/Position.lean | 160 -- .../Proof/Hypertree/PublicGraphOpenings.lean | 145 - .../Proof/Hypertree/PublicGraphSigner.lean | 94 - .../Hypertree/ReferenceGraphContext.lean | 119 - .../Hypertree/ReferenceGraphProgram.lean | 114 - .../Hypertree/ReferenceHypertreeWitness.lean | 131 - .../Proof/Hypertree/RootCache.lean | 168 -- .../Proof/Hypertree/Settled.lean | 34 - .../StructuralMatchAccumulation.lean | 139 - .../Proof/Hypertree/StructuralMatchBound.lean | 140 - .../Hypertree/StructuralMatchKernel.lean | 101 - .../StructuralOracleObservation.lean | 115 - .../Hypertree/StructuralOracleSplit.lean | 107 - .../Proof/Hypertree/StructuralTraceCache.lean | 148 - .../Proof/Hypertree/TreeFoldBound.lean | 13 - .../SphincsSecurity/Proof/IdealStatement.lean | 278 -- .../SphincsSecurity/Proof/LayerAssembly.lean | 51 - .../Proof/Ots/CanonicalEncodingSampling.lean | 110 - .../SphincsSecurity/Proof/Ots/Chain.lean | 30 - .../SphincsSecurity/Proof/Ots/Code.lean | 345 --- .../Proof/Ots/EncodingAdaptiveMarker.lean | 86 - .../Proof/Ots/EncodingBackwardWitness.lean | 94 - .../Proof/Ots/EncodingCached.lean | 33 - .../Proof/Ots/EncodingCharge.lean | 50 - .../Ots/EncodingConditionalObservation.lean | 62 - .../EncodingContactMarkerAccumulation.lean | 55 - .../Ots/EncodingContactMarkerSource.lean | 81 - .../Proof/Ots/EncodingContactMarkerStep.lean | 59 - .../Proof/Ots/EncodingFamilyObservation.lean | 34 - .../Proof/Ots/EncodingFamilyOracleSplit.lean | 63 - .../Proof/Ots/EncodingFreshRow.lean | 70 - .../Proof/Ots/EncodingInputs.lean | 22 - .../Proof/Ots/EncodingMarkerAccumulation.lean | 107 - .../Proof/Ots/EncodingMarkerAllocation.lean | 71 - .../Proof/Ots/EncodingMarkerBound.lean | 76 - .../Proof/Ots/EncodingMarkerKernel.lean | 55 - .../Proof/Ots/EncodingMarkerStep.lean | 82 - .../Proof/Ots/EncodingMatchAccumulation.lean | 151 - .../Proof/Ots/EncodingMatchBound.lean | 82 - .../Proof/Ots/EncodingMatchKernel.lean | 82 - .../Ots/EncodingNeighborProbability.lean | 34 - .../Proof/Ots/EncodingOracleObservation.lean | 72 - .../Proof/Ots/EncodingOracleSplit.lean | 92 - .../Proof/Ots/EncodingProbability.lean | 45 - .../Proof/Ots/EncodingSelectionCache.lean | 37 - .../Proof/Ots/EncodingTablePrior.lean | 59 - .../Proof/Ots/EncodingTarget.lean | 80 - .../Proof/Ots/EncodingTraceCache.lean | 154 - .../Proof/Ots/ExtractChain.lean | 75 - .../SphincsSecurity/Proof/Ots/ExtractOts.lean | 110 - .../Proof/Ots/LayerCompare.lean | 32 - .../Proof/Ots/LayerVerifierWitness.lean | 57 - .../SphincsSecurity/Proof/Ots/OneTime.lean | 55 - .../Proof/Ots/OtsChainBackward.lean | 157 - .../Proof/Ots/OtsContactAllocation.lean | 75 - .../Proof/Ots/OtsContactCheckpoint.lean | 53 - .../Proof/Ots/OtsContactEvents.lean | 94 - .../Proof/Ots/OtsContactFirstLaw.lean | 87 - .../Proof/Ots/OtsContactFirstProbability.lean | 116 - .../Proof/Ots/OtsContactMarkerBound.lean | 90 - .../Proof/Ots/OtsContactMarkerTrace.lean | 81 - .../Proof/Ots/OtsContactPause.lean | 41 - .../Proof/Ots/OtsContactProbability.lean | 61 - .../Proof/Ots/OtsContactRestart.lean | 63 - .../Proof/Ots/OtsContactSourceAllocation.lean | 57 - .../Proof/Ots/OtsContactSplit.lean | 59 - .../Proof/Ots/OtsContactTrace.lean | 94 - .../Proof/Ots/OtsDistinctContactBound.lean | 43 - .../Ots/OtsDistinctContactProbability.lean | 66 - .../Proof/Ots/OtsEncodingMarker.lean | 87 - .../Proof/Ots/OtsEndpointLikelihood.lean | 47 - .../Proof/Ots/OtsMarkerContactBound.lean | 46 - .../Proof/Ots/OtsMarkerContactPartition.lean | 109 - .../Ots/OtsMarkerContactProbability.lean | 97 - .../Proof/Ots/OtsMarkerContactSource.lean | 75 - .../Proof/Ots/OtsPrefixAccounting.lean | 102 - .../Proof/Ots/OtsPrefixAllocation.lean | 109 - .../Ots/OtsPrefixCheckpointCompletion.lean | 56 - .../Ots/OtsPrefixContactProbability.lean | 82 - .../Proof/Ots/OtsPrefixFrontier.lean | 101 - .../Proof/Ots/OtsPrefixIdealAllocation.lean | 76 - .../Proof/Ots/OtsPrefixInstrumentedSeed.lean | 56 - .../Ots/OtsPrefixInstrumentedSource.lean | 120 - .../Ots/OtsPrefixInstrumentedVisible.lean | 39 - .../Ots/OtsPrefixObservedAllocation.lean | 67 - .../Proof/Ots/OtsPrefixObservedBudget.lean | 55 - .../Proof/Ots/OtsPrefixObservedRun.lean | 52 - .../Proof/Ots/OtsPrefixObservedSource.lean | 96 - .../Proof/Ots/OtsPrefixOracle.lean | 142 - .../Proof/Ots/OtsPrefixRawOracle.lean | 94 - .../Proof/Ots/OtsPrefixRawSampling.lean | 95 - .../Proof/Ots/OtsPrefixReferenceSeed.lean | 67 - .../Proof/Ots/OtsPrefixSecretSampling.lean | 59 - .../Proof/Ots/OtsPrefixSeedAllocation.lean | 91 - .../Proof/Ots/OtsPrefixSeedGame.lean | 47 - .../Ots/OtsPrefixSeedReconstruction.lean | 134 - .../Proof/Ots/OtsPrefixSimulation.lean | 208 -- .../Proof/Ots/OtsPrefixSourceGame.lean | 79 - .../Ots/OtsPrefixTwoEdgeProbability.lean | 86 - .../Proof/Ots/OtsPrefixVisible.lean | 73 - .../Proof/Ots/OtsPrefixVisibleAccounting.lean | 69 - .../Proof/Ots/OtsPrefixVisibleCompletion.lean | 46 - .../Proof/Ots/OtsPrefixVisibleContact.lean | 128 - .../Proof/Ots/OtsPrefixVisibleTrace.lean | 128 - .../Ots/OtsProbeCanonicalChargeGame.lean | 49 - .../Proof/Ots/OtsProbeCompletionSampling.lean | 25 - .../Proof/Ots/OtsProbeOrigin.lean | 28 - .../Proof/Ots/OtsProbeSimulation.lean | 66 - .../Proof/Ots/OtsTraceCheckpointBudget.lean | 50 - .../Proof/Ots/OtsTraceCheckpointLaw.lean | 89 - .../Ots/OtsTraceCheckpointObservation.lean | 37 - .../Proof/Ots/OtsTraceEvents.lean | 76 - .../Proof/Ots/OtsTraceProbability.lean | 68 - .../Proof/Ots/OtsTraceRestart.lean | 138 - .../Proof/Ots/OtsTraceRowSource.lean | 45 - .../Proof/Ots/OtsTraceRows.lean | 111 - .../Proof/Ots/OtsTwoEdgeProbability.lean | 107 - .../Proof/Ots/OtsTwoEdgeSource.lean | 63 - .../Proof/Ots/OtsTwoEdgeTrace.lean | 54 - .../Proof/Ots/OtsVerifierWitness.lean | 115 - .../Proof/Ots/PrefixByteAction.lean | 83 - .../Proof/Ots/PrefixByteRun.lean | 19 - .../Proof/Ots/PrefixEncodingRisk.lean | 208 -- .../Proof/Ots/PublicEncodingMatch.lean | 95 - .../Proof/Ots/ReferenceEncodingContext.lean | 91 - .../Proof/Ots/ReferenceEncodingErasure.lean | 107 - .../Ots/ReferenceEncodingLazySource.lean | 65 - .../Ots/ReferenceEncodingPriorSource.lean | 45 - .../Proof/Ots/ReferenceEncodingProgram.lean | 105 - .../Proof/Ots/ReferenceEncodingSource.lean | 116 - .../Proof/Ots/ReferenceEncodingTable.lean | 61 - .../Proof/Ots/ReferenceFamilyAllocation.lean | 65 - .../Ots/ReferenceFamilyConditioning.lean | 77 - .../Proof/Ots/ReferenceFamilyGame.lean | 136 - .../Proof/Ots/ReferenceFamilySeed.lean | 66 - .../Proof/Ots/ReferenceLayerWitness.lean | 69 - .../Proof/Ots/ReferencePrefixGame.lean | 128 - .../Proof/Ots/ReferencePrefixResidual.lean | 71 - .../Proof/Ots/ReferencePrefixSigning.lean | 102 - .../Proof/Ots/SecretProbe.lean | 47 - .../Proof/Ots/TraceCheckpointObserver.lean | 46 - .../Proof/Ots/VerifierContactWitness.lean | 36 - .../Proof/RandomizedStatement.lean | 84 - .../Reference/BoundaryChargePartition.lean | 56 - .../Proof/Reference/BoundaryHashCost.lean | 101 - .../Reference/BoundaryHashEvaluation.lean | 262 -- .../Proof/Reference/BoundaryMessageCost.lean | 109 - .../Proof/Reference/BoundarySimulation.lean | 63 - .../Reference/CausalFrontierAllocation.lean | 131 - .../Proof/Reference/CausalFrontierGame.lean | 82 - .../Reference/CausalFrontierProgram.lean | 129 - .../Proof/Reference/CausalPublicSigning.lean | 150 - .../Proof/Reference/DirectQueryBudget.lean | 67 - .../Proof/Reference/FiniteHashWorld.lean | 151 - .../Proof/Reference/FixedHashBoundary.lean | 140 - .../Proof/Reference/FixedQueryBound.lean | 122 - .../Proof/Reference/QueryAllocation.lean | 99 - .../Proof/Reference/QueryBound.lean | 121 - .../Proof/Reference/QueryClassAllocation.lean | 91 - .../Reference/ReferenceAuxiliarySigning.lean | 101 - .../ReferenceCertificateCoverage.lean | 231 -- .../Reference/ReferenceCertificateTrace.lean | 168 -- .../Proof/Reference/ReferenceContactGame.lean | 98 - .../Reference/ReferenceForgeryAuxiliary.lean | 60 - .../Reference/ReferenceForgeryCoverage.lean | 45 - .../Reference/ReferenceForgerySource.lean | 252 -- .../Reference/ReferenceInstrumentedGame.lean | 63 - .../Proof/Reference/ReferenceJointPrior.lean | 31 - .../ReferenceOracleConditioning.lean | 49 - .../Reference/ReferencePrimitiveBound.lean | 180 -- .../Reference/ReferencePrimitiveWitness.lean | 66 - .../Reference/ReferenceQueryAllocation.lean | 62 - .../Reference/ReferenceResidualGame.lean | 67 - .../Reference/ReferenceResidualSampling.lean | 110 - .../Reference/ReferenceResidualSeeds.lean | 65 - .../Reference/ReferenceSigningReplay.lean | 102 - .../ReferenceVerifierInstantiation.lean | 103 - .../Reference/SigningBoundaryHashCost.lean | 104 - .../Proof/Reference/SigningTrace.lean | 76 - .../Proof/Reference/VerifierTraceDescent.lean | 22 - .../Proof/Reference/VerifierTraceSource.lean | 72 - .../VerifierWitnessClassification.lean | 42 - .../Residual/AdaptiveResidualErasure.lean | 148 - .../Residual/AdaptiveResidualLabels.lean | 216 -- .../Residual/CanonicalResidualQuery.lean | 97 - .../Residual/CanonicalResidualRouting.lean | 101 - .../Proof/Residual/CheckedByteExecution.lean | 49 - .../Residual/PublicReferenceResidual.lean | 101 - .../Proof/Residual/PublicResidualLookup.lean | 55 - .../Proof/Residual/ResidualByteAction.lean | 26 - .../Residual/ResidualByteCandidates.lean | 111 - .../Residual/ResidualByteCheckedHazard.lean | 76 - .../Residual/ResidualByteCorrespondence.lean | 91 - .../Proof/Residual/ResidualByteExecution.lean | 180 -- .../Proof/Residual/ResidualByteFrontend.lean | 100 - .../Proof/Residual/ResidualByteHazard.lean | 120 - .../Proof/Residual/ResidualByteRun.lean | 45 - .../Proof/Residual/ResidualGraphGame.lean | 49 - .../Residual/ResidualProbeCompletion.lean | 57 - .../Residual/ResidualSigningDisclosure.lean | 121 - .../Residual/ResidualSigningProgram.lean | 18 - .../Residual/ResidualTableCompletion.lean | 77 - .../Proof/Residual/RetainedObservation.lean | 143 - .../Residual/RetainedResidualAccounting.lean | 164 -- .../RetainedResidualBankCompleteness.lean | 181 -- .../RetainedResidualBoundaryCost.lean | 238 -- .../Residual/RetainedResidualBudget.lean | 109 - .../Residual/RetainedResidualByteRun.lean | 142 - .../RetainedResidualCacheAccounting.lean | 117 - .../RetainedResidualCacheHistory.lean | 107 - .../RetainedResidualCacheKernels.lean | 193 -- .../Residual/RetainedResidualCacheTail.lean | 90 - .../RetainedResidualCandidateHistory.lean | 39 - .../Residual/RetainedResidualCandidates.lean | 95 - .../RetainedResidualCertificateTransfer.lean | 138 - .../RetainedResidualCheckedTrace.lean | 108 - .../Residual/RetainedResidualCompletion.lean | 117 - .../Residual/RetainedResidualComposition.lean | 54 - .../Residual/RetainedResidualContext.lean | 116 - .../RetainedResidualCoverageStep.lean | 175 -- .../Residual/RetainedResidualDigestLaw.lean | 130 - .../RetainedResidualEncodingHistory.lean | 195 -- .../Residual/RetainedResidualEnvelope.lean | 67 - ...tainedResidualExceptionClassification.lean | 79 - .../RetainedResidualExceptionGame.lean | 149 - .../RetainedResidualExceptionHistory.lean | 118 - .../Residual/RetainedResidualExecution.lean | 173 -- .../RetainedResidualGameTransfer.lean | 172 -- .../Residual/RetainedResidualHashTrace.lean | 128 - .../Residual/RetainedResidualHazard.lean | 38 - .../Residual/RetainedResidualHistory.lean | 98 - .../Residual/RetainedResidualInitial.lean | 82 - .../RetainedResidualMessageKernel.lean | 141 - .../RetainedResidualMessagePayment.lean | 255 -- .../RetainedResidualMessageTrace.lean | 82 - .../RetainedResidualMonitorReadiness.lean | 131 - .../RetainedResidualMonitorStops.lean | 144 - .../RetainedResidualMonitoredCoverage.lean | 99 - .../RetainedResidualMonitoredErasure.lean | 149 - .../RetainedResidualMonitoredGame.lean | 122 - .../RetainedResidualMonitoredPayment.lean | 213 -- .../RetainedResidualMonitoredStep.lean | 157 - .../RetainedResidualNativePayment.lean | 116 - .../RetainedResidualOriginalBudget.lean | 267 -- .../RetainedResidualPaymentBudget.lean | 93 - .../RetainedResidualPrimitivePotential.lean | 554 ---- .../Residual/RetainedResidualProbeBudget.lean | 184 -- .../Residual/RetainedResidualProgram.lean | 81 - .../RetainedResidualProposalIndex.lean | 140 - .../RetainedResidualProposalInvariant.lean | 134 - .../RetainedResidualProposalPayment.lean | 170 -- .../Residual/RetainedResidualProposalRun.lean | 97 - .../RetainedResidualProposalStep.lean | 141 - .../RetainedResidualProposalSupport.lean | 159 - .../RetainedResidualProposalTail.lean | 126 - .../RetainedResidualProposalTailStep.lean | 126 - .../RetainedResidualProposalWord.lean | 101 - .../RetainedResidualQueryPotential.lean | 88 - .../Residual/RetainedResidualRecovery.lean | 282 -- .../Residual/RetainedResidualReplay.lean | 97 - .../Residual/RetainedResidualResources.lean | 143 - .../Proof/Residual/RetainedResidualRows.lean | 106 - .../RetainedResidualSigningCandidates.lean | 170 -- .../RetainedResidualSigningHistory.lean | 207 -- .../RetainedResidualSigningKernel.lean | 97 - .../Residual/RetainedResidualSigningLaw.lean | 127 - .../RetainedResidualSigningOrigin.lean | 203 -- .../RetainedResidualSigningTrace.lean | 174 -- .../Residual/RetainedResidualSource.lean | 169 -- .../RetainedResidualStrongCoverage.lean | 131 - .../RetainedResidualSuccessTransfer.lean | 141 - .../RetainedResidualTerminalCoverage.lean | 140 - .../RetainedResidualTraceValidity.lean | 107 - .../Residual/RetainedResidualVerify.lean | 127 - .../RetainedResidualVerifySupport.lean | 154 - .../Residual/RetainedResidualWorkCost.lean | 91 - .../RetainedResidualWorldCoverage.lean | 182 -- .../Residual/RetainedResidualWorldKernel.lean | 104 - .../Proof/Residual/RetainedSigningTrace.lean | 69 - .../Residual/RetainedWorldCoverBudget.lean | 33 - .../Residual/Security127LargeBudget.lean | 58 - .../SphincsSecurity/Proof/Scheme/Arith.lean | 50 - .../SphincsSecurity/Proof/Scheme/Bytes.lean | 180 -- .../SphincsSecurity/Proof/Scheme/Cached.lean | 56 - .../SphincsSecurity/Proof/Scheme/Charge.lean | 116 - .../SphincsSecurity/Proof/Scheme/Eval.lean | 46 - .../Proof/Scheme/Execution.lean | 37 - .../Proof/Scheme/FirstBad.lean | 63 - .../Proof/Scheme/ForgeryClassify.lean | 120 - .../SphincsSecurity/Proof/Scheme/Guess.lean | 36 - .../Proof/Scheme/NoMessage.lean | 273 -- .../SphincsSecurity/Proof/Scheme/Queried.lean | 187 -- .../Proof/Scheme/RawQueryMomentBound.lean | 178 -- .../SphincsSecurity/Proof/Scheme/Replay.lean | 25 - .../Proof/Scheme/ReplayWorld.lean | 120 - .../SphincsSecurity/Proof/Scheme/Secrets.lean | 83 - .../Proof/Scheme/SignSupport.lean | 158 - .../SphincsSecurity/Proof/Scheme/Slot.lean | 27 - .../Proof/Scheme/StatementLemmas.lean | 238 -- .../SphincsSecurity/Proof/Scheme/Support.lean | 289 -- .../Proof/Security127Completion.lean | 32 - .../Proof/Seeded/AdaptiveSeedGuessing.lean | 58 - .../Proof/Seeded/AlgorithmErasure.lean | 275 -- .../Proof/Seeded/BudgetTransfer.lean | 156 - .../Proof/Seeded/CacheCoupling.lean | 79 - .../Proof/Seeded/DerivationTable.lean | 90 - .../SphincsSecurity/Proof/Seeded/Erasure.lean | 167 -- .../Proof/Seeded/FiniteTable.lean | 62 - .../Proof/Seeded/FreshTable.lean | 166 -- .../Proof/Seeded/GameComparison.lean | 99 - .../Proof/Seeded/GameErasure.lean | 70 - .../Proof/Seeded/GameExpansion.lean | 107 - .../Proof/Seeded/HashTrace.lean | 92 - .../Proof/Seeded/KeyDerivation.lean | 46 - .../Proof/Seeded/KeygenBudget.lean | 106 - .../Proof/Seeded/Presampling.lean | 135 - .../Proof/Seeded/QueryBoundExtras.lean | 57 - .../Proof/Seeded/Security.lean | 58 - .../Proof/Seeded/SeedGuessing.lean | 34 - .../Proof/Seeded/StoppedRun.lean | 72 - .../Proof/Seeded/TableSampling.lean | 119 - .../Proof/SignatureLayout.lean | 32 - formal/sphincs/SphincsSecurity/Scheme.lean | 610 ---- formal/sphincs/SphincsSecurity/Statement.lean | 90 - .../SphincsSecurity/Tests/QueryBudget.lean | 58 - formal/sphincs/lake-manifest.json | 126 - formal/sphincs/lakefile.toml | 11 - formal/sphincs/lean-toolchain | 1 - formal/sphincs/scripts/Reach.lean | 45 - formal/sphincs/scripts/Taint.lean | 107 - formal/xmss/.gitignore | 1 - formal/xmss/README.md | 21 - formal/xmss/XmssSecurity.lean | 16 - formal/xmss/XmssSecurity/Proof.lean | 10 - .../Proof/AdaptiveEpochCollision.lean | 800 ------ .../Proof/AdaptiveFreshTarget.lean | 46 - .../Proof/AdaptiveRevealMonitor.lean | 85 - .../Proof/Adversary/Embedding.lean | 37 - .../Proof/Adversary/Security.lean | 19 - .../Proof/BoundedFirstLaneCoupling.lean | 215 -- .../Proof/BoundedSignProbability.lean | 85 - .../XmssSecurity/Proof/CacheAgreement.lean | 84 - .../XmssSecurity/Proof/CacheQuerySupport.lean | 117 - .../XmssSecurity/Proof/CacheReplayEval.lean | 421 --- .../xmss/XmssSecurity/Proof/CacheVerify.lean | 82 - .../CausalSigningKeygenCoupling.lean | 266 -- .../CappedChain/ChainEventDecomposition.lean | 195 -- .../Proof/CappedChain/ChainHiddenTable.lean | 112 - .../Proof/CappedChain/ChainInputTrace.lean | 91 - .../CappedChain/ChainOriginProbability.lean | 14 - .../CappedChain/ChainRevealFiltering.lean | 80 - .../Proof/CappedChain/ChainTracedGame.lean | 22 - .../CappedChain/DirectQueryAccounting.lean | 43 - .../Proof/CappedChain/EncodingQueryBound.lean | 105 - .../CappedChain/KeygenUnaddressedCache.lean | 129 - .../ReturnedChainValueCoverage.lean | 66 - .../CappedChain/SignatureChainValue.lean | 45 - .../Proof/CappedChain/SourceDirectTrace.lean | 121 - .../Proof/CappedChain/TreeCacheStability.lean | 139 - .../Proof/CappedConcreteExecution.lean | 139 - .../Proof/CappedDetailedQueryPresence.lean | 89 - .../Proof/CappedEncodingActionTrace.lean | 1033 ------- .../Proof/CappedEncodingEventProbability.lean | 44 - .../Proof/CappedEncodingExpectedBound.lean | 54 - .../Proof/CappedEncodingGameTrace.lean | 116 - .../Proof/CappedEncodingMonitor.lean | 853 ------ .../CappedEncodingPrehitExpectedBound.lean | 638 ----- .../CappedEncodingPrehitProbability.lean | 128 - .../Proof/CappedEncodingQueryBound.lean | 75 - .../Proof/CappedEncodingRejection.lean | 50 - .../Proof/CappedExactFirstLane.lean | 97 - .../Proof/CappedExactFirstLaneBound.lean | 1863 ------------ .../CappedExactFirstLaneSourceCoupling.lean | 763 ----- ...appedExactFirstLaneTransportReduction.lean | 2546 ----------------- .../Proof/CappedGlobalCausalSetup.lean | 642 ----- .../Proof/CappedGlobalCausalUniformTrace.lean | 25 - .../Proof/CappedGlobalChainHighCoverage.lean | 1831 ------------ .../CappedGlobalChainHighLocalCoupling.lean | 1177 -------- .../Proof/CappedGlobalChainHighReduction.lean | 730 ----- .../Proof/CappedGlobalChainHighReplay.lean | 1293 --------- .../Proof/CappedGlobalChainHighSetup.lean | 1340 --------- .../Proof/CappedGlobalChainOrigin.lean | 13 - .../CappedGlobalChainOutputSimulation.lean | 1317 --------- .../CappedGlobalChainOutputUniformity.lean | 103 - .../CappedGlobalCollisionProbability.lean | 76 - .../CappedGlobalFirstLaneExperiment.lean | 1510 ---------- .../Proof/CappedGlobalKeygen.lean | 2508 ---------------- .../CappedGlobalTreeCacheCorrespondence.lean | 1245 -------- .../Proof/CappedJointExactLoss.lean | 126 - .../Proof/CappedLeafEventProbability.lean | 442 --- .../Proof/CappedMerkleEventProbability.lean | 578 ---- .../Proof/CappedSigningCacheTrace.lean | 605 ---- .../Proof/CappedSigningLogReplay.lean | 156 - .../Proof/CappedSuffixEventProbability.lean | 530 ---- .../Proof/CappedUnifiedExpectedDigest.lean | 251 -- .../CappedUnifiedStructuralCollision.lean | 308 -- .../CausalAuthenticationPathIndependence.lean | 49 - .../Proof/CausalKeygenCoupling.lean | 394 --- .../Proof/CausalTreeCoupling.lean | 1029 ------- .../XmssSecurity/Proof/ChainHiddenTable.lean | 136 - .../XmssSecurity/Proof/ChainInputTrace.lean | 12 - .../Proof/ChainOraclePresampling.lean | 1238 -------- .../Proof/ChainTableUniformity.lean | 78 - .../XmssSecurity/Proof/ChainTargetInput.lean | 46 - .../Proof/ChainTrajectoryComposition.lean | 21 - .../Proof/ChainTrajectoryUniformity.lean | 130 - .../XmssSecurity/Proof/ChainWalkCache.lean | 45 - .../Proof/ConcreteCorrectness.lean | 249 -- .../XmssSecurity/Proof/ConcreteForgery.lean | 144 - .../Proof/ConcreteQueryBound.lean | 350 --- .../Proof/ConsistentQueryBound.lean | 171 -- .../XmssSecurity/Proof/DetailedExecution.lean | 74 - .../Proof/Deterministic/CostState.lean | 53 - .../Proof/Deterministic/DerivationTable.lean | 92 - .../Proof/Deterministic/FreshRequests.lean | 219 -- .../Proof/Deterministic/GameComparison.lean | 56 - .../Proof/Deterministic/GameExpansion.lean | 126 - .../Proof/Deterministic/Inputs.lean | 74 - .../Proof/Deterministic/KeygenBudget.lean | 96 - .../Proof/Deterministic/LoggedSigning.lean | 46 - .../Proof/Deterministic/MemoErasure.lean | 83 - .../Proof/Deterministic/MemoGame.lean | 50 - .../Proof/Deterministic/MemoLog.lean | 118 - .../Proof/Deterministic/MemoTable.lean | 57 - .../Proof/Deterministic/Memoize.lean | 114 - .../Proof/Deterministic/Preparation.lean | 86 - .../Deterministic/ReferenceDistribution.lean | 78 - .../Proof/Deterministic/ReferenceSource.lean | 72 - .../Proof/Deterministic/Replay.lean | 90 - .../Proof/Deterministic/RequestSampling.lean | 166 -- .../Proof/Deterministic/Security.lean | 34 - .../Proof/Deterministic/SignerSampling.lean | 19 - .../Proof/Deterministic/TableSigner.lean | 61 - .../Proof/Deterministic/TableToReference.lean | 75 - .../Deterministic/TranscriptReduction.lean | 81 - .../Proof/Deterministic/TrialLoop.lean | 73 - .../Proof/Deterministic/TrialSampling.lean | 145 - .../Proof/Deterministic/World.lean | 27 - .../Proof/EagerTraceInvariant.lean | 113 - .../Proof/EncodingActionTrace.lean | 265 -- .../Proof/EncodingEventProbability.lean | 147 - .../Proof/EncodingHashCacheReplay.lean | 39 - .../XmssSecurity/Proof/EncodingLemmas.lean | 238 -- .../Proof/EncodingOracleSimulation.lean | 30 - .../XmssSecurity/Proof/EncodingPrehit.lean | 15 - .../Proof/EncodingQueryAccounting.lean | 27 - .../XmssSecurity/Proof/EncodingTargetMap.lean | 89 - .../Proof/ExactKeygenQueryCount.lean | 211 -- .../XmssSecurity/Proof/ExactQueryCount.lean | 236 -- formal/xmss/XmssSecurity/Proof/Execution.lean | 75 - .../Proof/ExpectedAdaptiveFreshTarget.lean | 408 --- .../Proof/ExpectedQueryCount.lean | 312 -- .../Proof/FirstLaneEagerBound.lean | 932 ------ .../Proof/FirstLaneEagerSimulation.lean | 195 -- .../Proof/FirstLaneHazardEnforcement.lean | 427 --- .../Proof/FirstLaneOracleSimulation.lean | 76 - .../xmss/XmssSecurity/Proof/ForgeryCases.lean | 225 -- .../XmssSecurity/Proof/GlobalBadEvent.lean | 53 - .../GlobalWinningChainValueRevealed.lean | 14 - .../xmss/XmssSecurity/Proof/HashAddress.lean | 63 - .../XmssSecurity/Proof/HashInputLemmas.lean | 300 -- .../XmssSecurity/Proof/HashOutputHigh.lean | 8 - .../xmss/XmssSecurity/Proof/HiddenValue.lean | 19 - .../XmssSecurity/Proof/IdealStatement.lean | 85 - .../Proof/IndexedHiddenValue.lean | 58 - .../xmss/XmssSecurity/Proof/KeygenCache.lean | 139 - .../xmss/XmssSecurity/Proof/LazyScheme.lean | 143 - .../XmssSecurity/Proof/LeafTargetInput.lean | 43 - .../XmssSecurity/Proof/LossDecomposition.lean | 63 - .../XmssSecurity/Proof/MarginalCoupling.lean | 578 ---- formal/xmss/XmssSecurity/Proof/Merkle.lean | 87 - .../XmssSecurity/Proof/MerkleQueryBound.lean | 174 -- .../Proof/MerkleQueryPresence.lean | 262 -- .../MerkleVerificationQueryPresence.lean | 365 --- .../Proof/ObservedTraceMonitor.lean | 263 -- .../Proof/OutcomeClassification.lean | 156 - .../Proof/PrecomputedBoundedSign.lean | 121 - .../Proof/PrecomputedBoundedSignCache.lean | 345 --- .../PrecomputedBoundedSignProbability.lean | 162 -- .../Proof/PrecomputedKeyConsistency.lean | 275 -- .../Proof/PrecomputedKeygenCache.lean | 160 -- .../Proof/PrecomputedSignQueryBound.lean | 24 - .../XmssSecurity/Proof/QueryBoundSupport.lean | 41 - .../XmssSecurity/Proof/QueryCounting.lean | 62 - .../XmssSecurity/Proof/QueryPresence.lean | 409 --- .../xmss/XmssSecurity/Proof/RandomOracle.lean | 196 -- .../Proof/RandomOraclePresampling.lean | 173 -- .../Proof/RandomizedStatement.lean | 83 - .../Proof/RevealProbeOracleSimulation.lean | 457 --- .../XmssSecurity/Proof/RevealTraceReplay.lean | 63 - .../XmssSecurity/Proof/RunObservedAppend.lean | 91 - .../XmssSecurity/Proof/SecurityBudget.lean | 15 - .../Proof/Seeded/AdaptiveSeedGuessing.lean | 57 - .../Proof/Seeded/BudgetTransfer.lean | 170 -- .../Proof/Seeded/CacheCoupling.lean | 86 - .../XmssSecurity/Proof/Seeded/Erasure.lean | 165 -- .../Proof/Seeded/FiniteTable.lean | 62 - .../XmssSecurity/Proof/Seeded/FreshTable.lean | 179 -- .../Proof/Seeded/GameComparison.lean | 90 - .../Proof/Seeded/GameExpansion.lean | 78 - .../XmssSecurity/Proof/Seeded/HashTrace.lean | 92 - .../Proof/Seeded/KeyDerivation.lean | 50 - .../Proof/Seeded/KeygenBudget.lean | 95 - .../Proof/Seeded/KeygenExpansion.lean | 133 - .../Proof/Seeded/KeygenSampling.lean | 127 - .../Proof/Seeded/Presampling.lean | 135 - .../XmssSecurity/Proof/Seeded/Security.lean | 58 - .../Proof/Seeded/SeedGuessing.lean | 34 - .../XmssSecurity/Proof/Seeded/StoppedRun.lean | 72 - .../Proof/SignCacheHitProbability.lean | 232 -- .../XmssSecurity/Proof/SigningCacheTrace.lean | 197 -- .../Proof/SigningLogConsistency.lean | 17 - .../Proof/SigningRandomnessUniformity.lean | 14 - formal/xmss/XmssSecurity/Proof/StateLens.lean | 48 - .../XmssSecurity/Proof/StatementLemmas.lean | 173 -- .../Proof/TreeCacheVocabulary.lean | 241 -- .../XmssSecurity/Proof/TreeQueryBound.lean | 251 -- .../Proof/TreeValueTraversal.lean | 187 -- .../XmssSecurity/Proof/UnifiedBadEvent.lean | 38 - .../Proof/UniformFiniteTable.lean | 329 --- .../Proof/VerificationChainQuery.lean | 77 - .../Proof/WinningEventReduction.lean | 66 - formal/xmss/XmssSecurity/Proof/Wots.lean | 12 - .../XmssSecurity/Proof/WotsExtraction.lean | 271 -- formal/xmss/XmssSecurity/Scheme.lean | 448 --- formal/xmss/XmssSecurity/Statement.lean | 102 - .../xmss/XmssSecurity/Tests/QueryBudget.lean | 58 - formal/xmss/lake-manifest.json | 126 - formal/xmss/lakefile.toml | 11 - formal/xmss/lean-toolchain | 1 - 908 files changed, 1 insertion(+), 125776 deletions(-) delete mode 100644 .github/workflows/lean.yml delete mode 100755 doc/bench_blake3_vs_sha2_vs_sha3.sh delete mode 100644 doc/sphincs/latexmkrc delete mode 100644 doc/sphincs/main.tex delete mode 100644 doc/sphincs/refs.bib delete mode 100644 doc/xmss/latexmkrc delete mode 100644 doc/xmss/main.tex delete mode 100644 doc/xmss/refs.bib delete mode 100644 formal/sphincs/.gitignore delete mode 100644 formal/sphincs/PROOF.md delete mode 100644 formal/sphincs/README.md delete mode 100644 formal/sphincs/SphincsSecurity.lean delete mode 100644 formal/sphincs/SphincsSecurity/Completeness.lean delete mode 100644 formal/sphincs/SphincsSecurity/Completeness/Assembly.lean delete mode 100644 formal/sphincs/SphincsSecurity/Completeness/Code.lean delete mode 100644 formal/sphincs/SphincsSecurity/Completeness/Counter.lean delete mode 100644 formal/sphincs/SphincsSecurity/Completeness/Decay.lean delete mode 100644 formal/sphincs/SphincsSecurity/Completeness/Digest.lean delete mode 100644 formal/sphincs/SphincsSecurity/Completeness/Encoding.lean delete mode 100644 formal/sphincs/SphincsSecurity/Completeness/Fresh.lean delete mode 100644 formal/sphincs/SphincsSecurity/Completeness/Game.lean delete mode 100644 formal/sphincs/SphincsSecurity/Completeness/Keygen.lean delete mode 100644 formal/sphincs/SphincsSecurity/Completeness/Recovery.lean delete mode 100644 formal/sphincs/SphincsSecurity/Completeness/Search.lean delete mode 100644 formal/sphincs/SphincsSecurity/Completeness/Signing.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Adversary/Embedding.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Adversary/Security.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/BernoulliExcessMoments.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/BinomialMoments.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/FinitePmfProductObservation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/FirstSuccessFamily.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/FirstSuccessPrefix.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/FirstSuccessTable.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/FourthMomentExceptionBound.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/Prelude.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/QueryCap.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/QueryCapAccounting.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/QueryCapBalance.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/QueryCapErasure.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/QueryCapState.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/QueryPause.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/QueryPauseInvariant.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/QueryPauseTrace.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/QueryTraceInvariant.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/QueryTracePotential.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/RomQueryCharge.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/RomQueryChargeBind.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/StirlingMomentBounds.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/TraceSum.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/UniformTableCompletion.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/UniformTableConditioning.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/UniformTableDisclosure.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/UniformTableJoin.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/UniformTableObservation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/UniformTableObservationErasure.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/UniformTableOverwrite.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/UniformTableProducts.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/UniformTableRestriction.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/UniformTableSplit.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Base/WeightedQuery.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/AdaptiveChainActualBudget.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/AdaptiveChainAuxiliary.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/AdaptiveChainBudgetPotential.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/AdaptiveChainCap.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/AdaptiveChainCapContact.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/AdaptiveChainCapCost.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/AdaptiveChainCapObservation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/AdaptiveChainCapTwoEdge.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/AdaptiveChainCheckpoint.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/AdaptiveChainCheckpointContact.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/AdaptiveChainCheckpointProjection.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/AdaptiveChainCompensation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/AdaptiveChainContact.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/AdaptiveChainContactMoments.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/AdaptiveChainCountedRows.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/AdaptiveChainEndpoint.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/AdaptiveChainErasure.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/AdaptiveChainLikelihood.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/AdaptiveChainObservation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/AdaptiveChainPotential.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/AdaptiveChainQueryBound.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/AdaptiveChainRestart.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/AdaptiveChainSupport.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/AdaptiveChainTwoEdge.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/EndpointPreimageDensity.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/PartialChainCompletion.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/PartialChainContactCount.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/PartialChainContactPotential.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/PartialChainEndpoint.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/PartialChainLastRow.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/PartialChainLikelihoodLower.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/PartialChainLongCompletion.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/PartialChainObservation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/PartialChainPreparation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/PartialChainRowCharge.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/PartialChainSuffix.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/PartialChainTwoEdge.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/PartialChainTwoEdgeCharge.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Chains/PartialChainTwoEdgeCompensation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/ClosingParameters.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Deterministic/CostState.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Deterministic/DerivationTable.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Deterministic/FreshRequests.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Deterministic/GameComparison.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Deterministic/GameExpansion.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Deterministic/Inputs.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Deterministic/KeygenBudget.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Deterministic/LoggedSigning.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Deterministic/MemoErasure.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Deterministic/MemoGame.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Deterministic/MemoLog.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Deterministic/MemoTable.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Deterministic/Memoize.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Deterministic/Preparation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Deterministic/ReferenceDistribution.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Deterministic/ReferenceSource.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Deterministic/Replay.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Deterministic/RequestSampling.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Deterministic/Security.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Deterministic/SignerSampling.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Deterministic/TableSigner.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Deterministic/TableToReference.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Deterministic/TranscriptReduction.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Deterministic/TrialLoop.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Deterministic/TrialSampling.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessAuxiliaryProgram.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessBudget.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessCachedForced.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessCachedMessage.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessCachedSigning.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessDeferredMessage.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessDeferredSeed.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessEventTransfer.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessExceptionBounds.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessExceptionClassification.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessExceptionHistory.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessExceptionWeights.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessHash.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessInputCoverage.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessMonitoredAccounting.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessMonitoredPayment.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessMonitoredPotential.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessMonitoredStep.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessMonitoredValid.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessNearAssembly.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessNearDeferred.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessNearEvent.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessNearSource.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessNearWitness.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessPairBound.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessPairSource.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessProgram.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessProposalCompleted.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessProposalInvariant.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessProposalPayment.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessProposalStep.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessReference.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessReferenceWitness.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessRemaining.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessSeedCache.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessSigning.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessTracking.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessWitness.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/FtsGuessWork.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/SecretGuessErasure.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/SecretGuessForceBound.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/SecretGuessForceLikelihood.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/SecretGuessForced.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/SecretGuessForcedProgram.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/SecretGuessHitPayoff.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/SecretGuessHitRun.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/SecretGuessHitSum.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/SecretGuessObservation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/SecretGuessPairBound.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Forced/Security127SmallBudgetArithmetic.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/AdaptiveHiddenHazard.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/AdaptiveProposalWords.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/BankedCacheWeight.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/BankedProposalStep.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/BankedTargetEnvelope.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/BoundaryCertificateCache.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CacheIndexMultiplicity.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CacheMessageSignerWeight.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CacheMessageWeight.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CacheSize.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CachedDigestRate.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CachedDigestSelection.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CachedIndexExcessConcentration.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CachedIndexExcessScore.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CachedIndexHashMoments.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CachedSigningViews.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CachedTargetIncrement.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CachedTargetSubsetMatch.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CanonicalCoordinateSampling.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CanonicalHiddenCoordinates.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CertificateBankCompleteness.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CertificateBoundaryInvariants.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CertificateCacheExceptionGame.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CertificateCacheExceptionGrowth.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CertificateCacheExceptionKernels.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CertificateCacheExceptionPotential.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CertificateCacheMonitor.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CertificateCachePersistence.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CertificateCleanExecution.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CertificateCleanGame.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CertificateGame.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CertificateJointExceptions.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CertificateMessagePayment.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CertificateMonitor.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CertificateMonitorStep.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CertificateOriginalMessageCost.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CertificatePathBudget.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CertificateProposalInvariant.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CertificateProposalPrefixException.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CertificateStoppedState.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/CertificateTerminalGame.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/ConcreteTargetShapeQuery.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/ConcreteTargetShapeSigning.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/DigestAttemptExpectation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/DigestCompletionBank.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/DigestCompletionCacheGrowth.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/DigestCompletionLogGrowth.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/DigestCompletionNewTarget.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/DigestLoopRecord.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/DigestMessageCost.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/DigestSelectionIndex.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/DigestSelectionMass.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/DigestSelectionWeight.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/DigestSigningCompletion.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/ExactTargetShapeSigning.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/ExtractFts.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/FewTimeConditionalCoverage.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/FewTimeFixedPrehit.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/FewTimeFresh.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/FewTimeFreshMass.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/FewTimeLoop.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/FewTimePadding.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/FewTimePrehit.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/FewTimeProbability.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/FewTimeRace.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/FewTimeSignerView.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/FewTimeSource.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/FewTimeTargetCompletion.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/FewTimeUniform.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/FewTimeWeightedOriginRace.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/FewTimeWitness.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/FixedCertificateCoverage.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/FixedProposalMoments.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/FreshDigestHazard.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/FreshTargetEnvelope.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/FreshTargetPayload.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/FreshTargetShapeAverage.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/FtsProbeAdversary.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/FtsProbeGame.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/FtsProbeOrigin.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/FtsProbeProbability.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/FtsProbeSampling.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/FtsProbeSimulation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/FtsProbeVerifierSource.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/FtsVerifierWitness.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/FutureCoverageBound.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/HiddenLabelObservation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/HiddenLabelProbe.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/InterleavedCoverStep.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/InterleavedResidualDisclosure.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/JointProbeMessageAnswers.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/JointProbeMessageReserve.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/MessageAdmissibleDeficit.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/MessageByteTrace.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/MessageCacheCountGrowth.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/MessageCacheProjection.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/MessageCertificateProjection.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/MessageDeficitCacheGrowth.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/MessageDeficitConcentration.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/MessageDeficitHashMoments.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/MessageDeficitMomentGrowth.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/MessageDeficitScore.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/MessageDigestHazard.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/MessageInputMiss.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/MessageNormalizedReuse.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/MessagePrehit.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/NearCertificateBound.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/NewTargetEnvelopeCharge.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/NormalizedTargetCacheQuery.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/NormalizedTargetLogSigning.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/NormalizedTargetMatches.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/ObservedAdaptiveCoverBound.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/ObservedFreshCoverBound.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/ObservedOccupancy.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/OriginalCacheExceptionBound.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/OriginalCertificateBound.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/OriginalCertificateTrace.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/OriginalMessageAllocation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/OriginalProposalBudget.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/OriginalProposalExecution.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/OriginalProposalPrefixBound.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/OriginalTerminalProposal.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/PairedHiddenMiss.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/Parameters.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/PrimitiveMessagePotential.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/ProposalBridgeKernel.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/ProposalLengthProjection.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/ProposalPrefixExponential.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/ProposalPrefixStop.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/ProposalQueryProjection.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/ProposalWordDistribution.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/PublicSigningRecord.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/RawProposalMomentBound.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/RawSigningMomentBound.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/ReferenceFtsCoverage.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/ReuseCachedTargets.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/ReuseRawEnvelope.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/ReuseTargetEnvelope.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/SelectedDigestCache.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/SignerAdmissibleMessage.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/SignerDigestSource.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/SigningProposalRecord.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/SingleMessageCacheGrowth.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/SourceGroupExpectation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/StoppedSigningLog.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/SubsetTargetAssignment.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/SubsetTargetExpectation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/TargetAssignmentCount.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/TargetIndexEnvelope.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/TargetMixedGrowthPolynomial.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/TargetMomentShapes.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/TargetShapeBlocks.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/TargetShapeCardinality.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/TargetShapeContinuation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/TargetShapeEnvelope.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/TargetShapeExpectation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/TargetShapeOperators.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/TargetShapeReindex.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/TargetSigningMatchFactors.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/TargetSourceMultiplicity.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/TerminalCache.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/TerminalCertificateCharge.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/TerminalProposalEnvelope.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/TerminalProposalWord.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/UniformProposalMixedMoments.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/UniformProposalMoments.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/UniformProposalVariance.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/UniformPublicCoordinates.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/UnitCertificateCoverage.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/UnrestrictedRowSwap.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/UpperDigestSelection.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/ValidInterleavedCover.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Fts/WeightedTargetGroups.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/CanonicalGraph.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/CanonicalGraphGame.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/CanonicalGraphHonest.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/CanonicalGraphResidual.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/CanonicalGraphSampling.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/CanonicalProbeCache.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/CanonicalProbeRouting.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/CanonicalPublicPrior.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/CanonicalSigningFrontier.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/Descent.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/Extract.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/FiniteGraphReplay.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/FiniteGraphSampling.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/FrontierEncodingCongruence.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/FrontierGameProjection.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/FrontierOracleCongruence.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/FrontierOracleMask.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/FrontierRandomOracle.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/FrontierSignerErasure.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/FrontierSigningEvaluation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/FrontierSigningOracleCongruence.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/FrontierSigningOrigin.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/FrontierTreeEvaluation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/GraphPayloadInputs.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/Honest.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/Hypertree.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/Parameters.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/Position.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/PublicGraphOpenings.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/PublicGraphSigner.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/ReferenceGraphContext.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/ReferenceGraphProgram.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/ReferenceHypertreeWitness.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/RootCache.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/Settled.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/StructuralMatchAccumulation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/StructuralMatchBound.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/StructuralMatchKernel.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/StructuralOracleObservation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/StructuralOracleSplit.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/StructuralTraceCache.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Hypertree/TreeFoldBound.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/IdealStatement.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/LayerAssembly.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/CanonicalEncodingSampling.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/Chain.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/Code.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/EncodingAdaptiveMarker.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/EncodingBackwardWitness.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/EncodingCached.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/EncodingCharge.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/EncodingConditionalObservation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/EncodingContactMarkerAccumulation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/EncodingContactMarkerSource.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/EncodingContactMarkerStep.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/EncodingFamilyObservation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/EncodingFamilyOracleSplit.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/EncodingFreshRow.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/EncodingInputs.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/EncodingMarkerAccumulation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/EncodingMarkerAllocation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/EncodingMarkerBound.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/EncodingMarkerKernel.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/EncodingMarkerStep.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/EncodingMatchAccumulation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/EncodingMatchBound.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/EncodingMatchKernel.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/EncodingNeighborProbability.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/EncodingOracleObservation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/EncodingOracleSplit.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/EncodingProbability.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/EncodingSelectionCache.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/EncodingTablePrior.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/EncodingTarget.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/EncodingTraceCache.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/ExtractChain.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/ExtractOts.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/LayerCompare.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/LayerVerifierWitness.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OneTime.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsChainBackward.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsContactAllocation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsContactCheckpoint.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsContactEvents.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsContactFirstLaw.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsContactFirstProbability.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsContactMarkerBound.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsContactMarkerTrace.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsContactPause.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsContactProbability.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsContactRestart.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsContactSourceAllocation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsContactSplit.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsContactTrace.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsDistinctContactBound.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsDistinctContactProbability.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsEncodingMarker.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsEndpointLikelihood.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsMarkerContactBound.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsMarkerContactPartition.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsMarkerContactProbability.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsMarkerContactSource.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsPrefixAccounting.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsPrefixAllocation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsPrefixCheckpointCompletion.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsPrefixContactProbability.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsPrefixFrontier.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsPrefixIdealAllocation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsPrefixInstrumentedSeed.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsPrefixInstrumentedSource.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsPrefixInstrumentedVisible.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsPrefixObservedAllocation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsPrefixObservedBudget.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsPrefixObservedRun.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsPrefixObservedSource.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsPrefixOracle.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsPrefixRawOracle.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsPrefixRawSampling.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsPrefixReferenceSeed.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsPrefixSecretSampling.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsPrefixSeedAllocation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsPrefixSeedGame.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsPrefixSeedReconstruction.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsPrefixSimulation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsPrefixSourceGame.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsPrefixTwoEdgeProbability.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsPrefixVisible.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsPrefixVisibleAccounting.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsPrefixVisibleCompletion.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsPrefixVisibleContact.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsPrefixVisibleTrace.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsProbeCanonicalChargeGame.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsProbeCompletionSampling.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsProbeOrigin.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsProbeSimulation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsTraceCheckpointBudget.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsTraceCheckpointLaw.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsTraceCheckpointObservation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsTraceEvents.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsTraceProbability.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsTraceRestart.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsTraceRowSource.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsTraceRows.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsTwoEdgeProbability.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsTwoEdgeSource.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsTwoEdgeTrace.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/OtsVerifierWitness.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/PrefixByteAction.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/PrefixByteRun.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/PrefixEncodingRisk.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/PublicEncodingMatch.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/ReferenceEncodingContext.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/ReferenceEncodingErasure.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/ReferenceEncodingLazySource.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/ReferenceEncodingPriorSource.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/ReferenceEncodingProgram.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/ReferenceEncodingSource.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/ReferenceEncodingTable.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/ReferenceFamilyAllocation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/ReferenceFamilyConditioning.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/ReferenceFamilyGame.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/ReferenceFamilySeed.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/ReferenceLayerWitness.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/ReferencePrefixGame.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/ReferencePrefixResidual.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/ReferencePrefixSigning.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/SecretProbe.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/TraceCheckpointObserver.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Ots/VerifierContactWitness.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/RandomizedStatement.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/BoundaryChargePartition.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/BoundaryHashCost.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/BoundaryHashEvaluation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/BoundaryMessageCost.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/BoundarySimulation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/CausalFrontierAllocation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/CausalFrontierGame.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/CausalFrontierProgram.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/CausalPublicSigning.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/DirectQueryBudget.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/FiniteHashWorld.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/FixedHashBoundary.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/FixedQueryBound.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/QueryAllocation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/QueryBound.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/QueryClassAllocation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/ReferenceAuxiliarySigning.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/ReferenceCertificateCoverage.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/ReferenceCertificateTrace.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/ReferenceContactGame.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/ReferenceForgeryAuxiliary.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/ReferenceForgeryCoverage.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/ReferenceForgerySource.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/ReferenceInstrumentedGame.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/ReferenceJointPrior.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/ReferenceOracleConditioning.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/ReferencePrimitiveBound.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/ReferencePrimitiveWitness.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/ReferenceQueryAllocation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/ReferenceResidualGame.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/ReferenceResidualSampling.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/ReferenceResidualSeeds.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/ReferenceSigningReplay.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/ReferenceVerifierInstantiation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/SigningBoundaryHashCost.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/SigningTrace.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/VerifierTraceDescent.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/VerifierTraceSource.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Reference/VerifierWitnessClassification.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/AdaptiveResidualErasure.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/AdaptiveResidualLabels.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/CanonicalResidualQuery.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/CanonicalResidualRouting.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/CheckedByteExecution.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/PublicReferenceResidual.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/PublicResidualLookup.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/ResidualByteAction.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/ResidualByteCandidates.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/ResidualByteCheckedHazard.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/ResidualByteCorrespondence.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/ResidualByteExecution.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/ResidualByteFrontend.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/ResidualByteHazard.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/ResidualByteRun.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/ResidualGraphGame.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/ResidualProbeCompletion.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/ResidualSigningDisclosure.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/ResidualSigningProgram.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/ResidualTableCompletion.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedObservation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualAccounting.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualBankCompleteness.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualBoundaryCost.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualBudget.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualByteRun.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualCacheAccounting.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualCacheHistory.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualCacheKernels.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualCacheTail.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualCandidateHistory.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualCandidates.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualCertificateTransfer.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualCheckedTrace.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualCompletion.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualComposition.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualContext.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualCoverageStep.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualDigestLaw.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualEncodingHistory.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualEnvelope.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualExceptionClassification.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualExceptionGame.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualExceptionHistory.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualExecution.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualGameTransfer.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualHashTrace.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualHazard.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualHistory.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualInitial.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualMessageKernel.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualMessagePayment.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualMessageTrace.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualMonitorReadiness.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualMonitorStops.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualMonitoredCoverage.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualMonitoredErasure.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualMonitoredGame.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualMonitoredPayment.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualMonitoredStep.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualNativePayment.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualOriginalBudget.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualPaymentBudget.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualPrimitivePotential.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualProbeBudget.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualProgram.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualProposalIndex.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualProposalInvariant.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualProposalPayment.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualProposalRun.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualProposalStep.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualProposalSupport.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualProposalTail.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualProposalTailStep.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualProposalWord.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualQueryPotential.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualRecovery.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualReplay.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualResources.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualRows.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualSigningCandidates.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualSigningHistory.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualSigningKernel.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualSigningLaw.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualSigningOrigin.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualSigningTrace.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualSource.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualStrongCoverage.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualSuccessTransfer.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualTerminalCoverage.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualTraceValidity.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualVerify.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualVerifySupport.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualWorkCost.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualWorldCoverage.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedResidualWorldKernel.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedSigningTrace.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/RetainedWorldCoverBudget.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Residual/Security127LargeBudget.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Scheme/Arith.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Scheme/Bytes.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Scheme/Cached.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Scheme/Charge.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Scheme/Eval.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Scheme/Execution.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Scheme/FirstBad.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Scheme/ForgeryClassify.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Scheme/Guess.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Scheme/NoMessage.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Scheme/Queried.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Scheme/RawQueryMomentBound.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Scheme/Replay.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Scheme/ReplayWorld.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Scheme/Secrets.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Scheme/SignSupport.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Scheme/Slot.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Scheme/StatementLemmas.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Scheme/Support.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Security127Completion.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Seeded/AdaptiveSeedGuessing.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Seeded/AlgorithmErasure.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Seeded/BudgetTransfer.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Seeded/CacheCoupling.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Seeded/DerivationTable.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Seeded/Erasure.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Seeded/FiniteTable.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Seeded/FreshTable.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Seeded/GameComparison.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Seeded/GameErasure.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Seeded/GameExpansion.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Seeded/HashTrace.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Seeded/KeyDerivation.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Seeded/KeygenBudget.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Seeded/Presampling.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Seeded/QueryBoundExtras.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Seeded/Security.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Seeded/SeedGuessing.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Seeded/StoppedRun.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/Seeded/TableSampling.lean delete mode 100644 formal/sphincs/SphincsSecurity/Proof/SignatureLayout.lean delete mode 100644 formal/sphincs/SphincsSecurity/Scheme.lean delete mode 100644 formal/sphincs/SphincsSecurity/Statement.lean delete mode 100644 formal/sphincs/SphincsSecurity/Tests/QueryBudget.lean delete mode 100644 formal/sphincs/lake-manifest.json delete mode 100644 formal/sphincs/lakefile.toml delete mode 100644 formal/sphincs/lean-toolchain delete mode 100644 formal/sphincs/scripts/Reach.lean delete mode 100644 formal/sphincs/scripts/Taint.lean delete mode 100644 formal/xmss/.gitignore delete mode 100644 formal/xmss/README.md delete mode 100644 formal/xmss/XmssSecurity.lean delete mode 100644 formal/xmss/XmssSecurity/Proof.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/AdaptiveEpochCollision.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/AdaptiveFreshTarget.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/AdaptiveRevealMonitor.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Adversary/Embedding.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Adversary/Security.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/BoundedFirstLaneCoupling.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/BoundedSignProbability.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CacheAgreement.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CacheQuerySupport.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CacheReplayEval.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CacheVerify.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedChain/CausalSigningKeygenCoupling.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedChain/ChainEventDecomposition.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedChain/ChainHiddenTable.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedChain/ChainInputTrace.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedChain/ChainOriginProbability.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedChain/ChainRevealFiltering.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedChain/ChainTracedGame.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedChain/DirectQueryAccounting.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedChain/EncodingQueryBound.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedChain/KeygenUnaddressedCache.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedChain/ReturnedChainValueCoverage.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedChain/SignatureChainValue.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedChain/SourceDirectTrace.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedChain/TreeCacheStability.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedConcreteExecution.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedDetailedQueryPresence.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedEncodingActionTrace.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedEncodingEventProbability.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedEncodingExpectedBound.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedEncodingGameTrace.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedEncodingMonitor.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedEncodingPrehitExpectedBound.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedEncodingPrehitProbability.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedEncodingQueryBound.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedEncodingRejection.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedExactFirstLane.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedExactFirstLaneBound.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedExactFirstLaneSourceCoupling.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedExactFirstLaneTransportReduction.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedGlobalCausalSetup.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedGlobalCausalUniformTrace.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedGlobalChainHighCoverage.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedGlobalChainHighLocalCoupling.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedGlobalChainHighReduction.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedGlobalChainHighReplay.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedGlobalChainHighSetup.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedGlobalChainOrigin.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedGlobalChainOutputSimulation.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedGlobalChainOutputUniformity.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedGlobalCollisionProbability.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedGlobalFirstLaneExperiment.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedGlobalKeygen.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedGlobalTreeCacheCorrespondence.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedJointExactLoss.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedLeafEventProbability.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedMerkleEventProbability.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedSigningCacheTrace.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedSigningLogReplay.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedSuffixEventProbability.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedUnifiedExpectedDigest.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CappedUnifiedStructuralCollision.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CausalAuthenticationPathIndependence.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CausalKeygenCoupling.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/CausalTreeCoupling.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/ChainHiddenTable.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/ChainInputTrace.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/ChainOraclePresampling.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/ChainTableUniformity.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/ChainTargetInput.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/ChainTrajectoryComposition.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/ChainTrajectoryUniformity.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/ChainWalkCache.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/ConcreteCorrectness.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/ConcreteForgery.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/ConcreteQueryBound.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/ConsistentQueryBound.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/DetailedExecution.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Deterministic/CostState.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Deterministic/DerivationTable.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Deterministic/FreshRequests.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Deterministic/GameComparison.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Deterministic/GameExpansion.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Deterministic/Inputs.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Deterministic/KeygenBudget.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Deterministic/LoggedSigning.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Deterministic/MemoErasure.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Deterministic/MemoGame.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Deterministic/MemoLog.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Deterministic/MemoTable.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Deterministic/Memoize.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Deterministic/Preparation.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Deterministic/ReferenceDistribution.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Deterministic/ReferenceSource.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Deterministic/Replay.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Deterministic/RequestSampling.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Deterministic/Security.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Deterministic/SignerSampling.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Deterministic/TableSigner.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Deterministic/TableToReference.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Deterministic/TranscriptReduction.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Deterministic/TrialLoop.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Deterministic/TrialSampling.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Deterministic/World.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/EagerTraceInvariant.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/EncodingActionTrace.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/EncodingEventProbability.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/EncodingHashCacheReplay.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/EncodingLemmas.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/EncodingOracleSimulation.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/EncodingPrehit.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/EncodingQueryAccounting.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/EncodingTargetMap.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/ExactKeygenQueryCount.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/ExactQueryCount.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Execution.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/ExpectedAdaptiveFreshTarget.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/ExpectedQueryCount.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/FirstLaneEagerBound.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/FirstLaneEagerSimulation.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/FirstLaneHazardEnforcement.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/FirstLaneOracleSimulation.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/ForgeryCases.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/GlobalBadEvent.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/GlobalWinningChainValueRevealed.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/HashAddress.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/HashInputLemmas.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/HashOutputHigh.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/HiddenValue.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/IdealStatement.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/IndexedHiddenValue.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/KeygenCache.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/LazyScheme.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/LeafTargetInput.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/LossDecomposition.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/MarginalCoupling.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Merkle.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/MerkleQueryBound.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/MerkleQueryPresence.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/MerkleVerificationQueryPresence.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/ObservedTraceMonitor.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/OutcomeClassification.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/PrecomputedBoundedSign.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/PrecomputedBoundedSignCache.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/PrecomputedBoundedSignProbability.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/PrecomputedKeyConsistency.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/PrecomputedKeygenCache.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/PrecomputedSignQueryBound.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/QueryBoundSupport.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/QueryCounting.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/QueryPresence.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/RandomOracle.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/RandomOraclePresampling.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/RandomizedStatement.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/RevealProbeOracleSimulation.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/RevealTraceReplay.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/RunObservedAppend.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/SecurityBudget.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Seeded/AdaptiveSeedGuessing.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Seeded/BudgetTransfer.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Seeded/CacheCoupling.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Seeded/Erasure.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Seeded/FiniteTable.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Seeded/FreshTable.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Seeded/GameComparison.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Seeded/GameExpansion.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Seeded/HashTrace.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Seeded/KeyDerivation.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Seeded/KeygenBudget.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Seeded/KeygenExpansion.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Seeded/KeygenSampling.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Seeded/Presampling.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Seeded/Security.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Seeded/SeedGuessing.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Seeded/StoppedRun.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/SignCacheHitProbability.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/SigningCacheTrace.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/SigningLogConsistency.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/SigningRandomnessUniformity.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/StateLens.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/StatementLemmas.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/TreeCacheVocabulary.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/TreeQueryBound.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/TreeValueTraversal.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/UnifiedBadEvent.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/UniformFiniteTable.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/VerificationChainQuery.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/WinningEventReduction.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/Wots.lean delete mode 100644 formal/xmss/XmssSecurity/Proof/WotsExtraction.lean delete mode 100644 formal/xmss/XmssSecurity/Scheme.lean delete mode 100644 formal/xmss/XmssSecurity/Statement.lean delete mode 100644 formal/xmss/XmssSecurity/Tests/QueryBudget.lean delete mode 100644 formal/xmss/lake-manifest.json delete mode 100644 formal/xmss/lakefile.toml delete mode 100644 formal/xmss/lean-toolchain diff --git a/.github/workflows/doc.yml b/.github/workflows/doc.yml index 300b36174..1da1f73ed 100644 --- a/.github/workflows/doc.yml +++ b/.github/workflows/doc.yml @@ -5,15 +5,11 @@ on: branches: [ "main" ] paths: - 'doc/leanvm/**' - - 'doc/xmss/**' - - 'doc/sphincs/**' - 'doc/images/**' - '.github/workflows/doc.yml' pull_request: paths: - 'doc/leanvm/**' - - 'doc/xmss/**' - - 'doc/sphincs/**' - 'doc/images/**' - '.github/workflows/doc.yml' workflow_dispatch: @@ -38,22 +34,6 @@ jobs: - name: Fail on an undefined reference or citation run: | ! grep -qE 'Reference .* undefined|Citation .* undefined|multiply defined' doc/leanvm/.build/main.log - - name: Compile XMSS specification - uses: xu-cheng/latex-action@v3 - with: - working_directory: doc/xmss - root_file: main.tex - - name: Fail on an undefined XMSS reference or citation - run: | - ! grep -qE 'Reference .* undefined|Citation .* undefined|multiply defined' doc/xmss/.build/main.log - - name: Compile SPHINCS specification - uses: xu-cheng/latex-action@v3 - with: - working_directory: doc/sphincs - root_file: main.tex - - name: Fail on an undefined SPHINCS reference or citation - run: | - ! grep -qE 'Reference .* undefined|Citation .* undefined|multiply defined' doc/sphincs/.build/main.log build-pdf: if: github.event_name != 'pull_request' && github.ref == 'refs/heads/main' @@ -67,21 +47,9 @@ jobs: with: working_directory: doc/leanvm root_file: main.tex - - name: Compile XMSS specification - uses: xu-cheng/latex-action@v3 - with: - working_directory: doc/xmss - root_file: main.tex - - name: Compile SPHINCS specification - uses: xu-cheng/latex-action@v3 - with: - working_directory: doc/sphincs - root_file: main.tex - name: Name the artifacts run: | cp doc/leanvm/.build/main.pdf leanVM.pdf - cp doc/xmss/.build/main.pdf XMSS.pdf - cp doc/sphincs/.build/main.pdf SPHINCS.pdf - name: Publish PDFs as release assets uses: softprops/action-gh-release@v2 with: @@ -91,10 +59,6 @@ jobs: Auto-built on every push to `main`. `leanVM.pdf` contains the leanVM specification. - `XMSS.pdf` contains the XMSS specification. - `SPHINCS.pdf` contains the SPHINCS specification. make_latest: false files: | leanVM.pdf - XMSS.pdf - SPHINCS.pdf diff --git a/.github/workflows/lean.yml b/.github/workflows/lean.yml deleted file mode 100644 index a7a1e8556..000000000 --- a/.github/workflows/lean.yml +++ /dev/null @@ -1,73 +0,0 @@ -name: Lean - -on: - push: - branches: [ "main" ] - pull_request: - workflow_dispatch: - -permissions: - contents: read - -concurrency: - group: lean-${{ github.ref }} - cancel-in-progress: true - -jobs: - xmss-formalization: - runs-on: ubuntu-latest - steps: - - uses: actions/checkout@v4 - # Cheap first pass: the axiom guard in `XmssSecurity.lean` catches a `sorry` - # or a `native_decide` reaching the root theorem, this catches one parked - # anywhere in the project. - - name: Forbid proof escapes - run: | - ! grep -rnE '\b(sorry|sorryAx|admit|native_decide|unsafe|implemented_by)\b|#exit' \ - --include='*.lean' --exclude-dir=.lake formal/xmss | grep -vE ':[0-9]+: *(--|/-)' - # Installs the toolchain from `formal/xmss/lean-toolchain`, fetches the - # mathlib cache, and runs `lake build` on the default target, which - # elaborates the root module and with it the `#guard_msgs` check. The - # checked-in manifest is used as is: no `lake update`. - - uses: leanprover/lean-action@v1 - with: - lake-package-directory: formal/xmss - # The root module guards its own footprint with `#guard_msgs`, which an - # edit to the expected message would silence. This asks again from - # outside, against the list written here. - - name: Check the axiom footprint - working-directory: formal/xmss - run: | - printf 'import XmssSecurity\n#print axioms XmssSecurity.xmss_has_127_bits_of_classical_security\n' \ - > "$RUNNER_TEMP/axioms.lean" - lake env lean "$RUNNER_TEMP/axioms.lean" | tee "$RUNNER_TEMP/axioms.txt" - grep -qF "'XmssSecurity.xmss_has_127_bits_of_classical_security' depends on axioms: [propext, Classical.choice, Quot.sound]" \ - "$RUNNER_TEMP/axioms.txt" - - uses: actions/upload-artifact@v4 - with: - name: xmss-axioms - path: ${{ runner.temp }}/axioms.txt - - sphincs-formalization: - runs-on: ubuntu-latest - steps: - - uses: actions/checkout@v4 - - name: Forbid proof escapes - run: | - ! grep -rnE '\b(sorry|sorryAx|admit|native_decide|unsafe|implemented_by)\b|#exit' \ - --include='*.lean' --exclude-dir=.lake formal/sphincs | grep -vE ':[0-9]+: *(--|/-)' - - uses: leanprover/lean-action@v1 - with: - lake-package-directory: formal/sphincs - - name: Check the axiom footprint - working-directory: formal/sphincs - run: | - printf 'import SphincsSecurity\n#print axioms SphincsSecurity.sphincs_has_127_bits_of_classical_security\n' \ - > "$RUNNER_TEMP/axioms.lean" - lake env lean "$RUNNER_TEMP/axioms.lean" | tee "$RUNNER_TEMP/axioms.txt" - grep -qF "'SphincsSecurity.sphincs_has_127_bits_of_classical_security' depends on axioms: [propext, Classical.choice, Quot.sound]" \ - "$RUNNER_TEMP/axioms.txt" - - uses: actions/upload-artifact@v4 - with: - name: sphincs-axioms - path: ${{ runner.temp }}/axioms.txt diff --git a/.gitignore b/.gitignore index 60e0034dd..70c4e41c6 100644 --- a/.gitignore +++ b/.gitignore @@ -1,3 +1,4 @@ .build target __pycache__/ +formal/ diff --git a/doc/bench_blake3_vs_sha2_vs_sha3.sh b/doc/bench_blake3_vs_sha2_vs_sha3.sh deleted file mode 100755 index 498c8b950..000000000 --- a/doc/bench_blake3_vs_sha2_vs_sha3.sh +++ /dev/null @@ -1,161 +0,0 @@ -#!/usr/bin/env bash -# -# ./doc/bench_blake3_vs_sha2_vs_sha3.sh -# -# Compares the three candidate hashes on the branch that implements each: -# BLAKE2s (main), SHA-256 (sha2), Keccak (sha3). Every branch is fetched and -# fast-forwarded first, then measured four ways: -# -# raw hashing, from `multithreaded_throughput`, which reports one 64-byte -# input per compression (per permutation, for Keccak). Run HASH_RUNS times, -# best kept: the test already medians its own passes, so a low run is the -# machine being busy rather than the hash being slow. From it comes the -# compression rate, and the time to generate one XMSS key over a -# 2^LOG_LIFETIME lifetime at COMPRESSIONS compressions per epoch. Comparing -# the three rates, mind that a 64-byte hash is one compression for BLAKE2s -# and one permutation for Keccak, but two compressions for SHA-256, whose -# padding spills a 64-byte input into a second block. -# -# proving that hashing, from flock's `hash_batch_prove_verify`, whose -# throughput line counts compressions (permutations, for Keccak) proven per -# second, with none of the VM around it. -# -# XMSS aggregation, SPHINCS aggregation, and REC_N-to-1 recursion over leaves -# of XMSS aggregates. None of the three counts is the same on every branch: -# each is whatever fills a proof of the same proven size, so a costlier hash -# fits fewer signatures, as does flock's batch, and the counts below are the -# ones each branch's -# README quotes. Comparing them means comparing signatures per second, not -# proving times, the recursion rate counting every signature the node covers -# (REC_N leaves of XMSS_PER_LEAF). -# -# Wrapped in braces so bash parses the whole file before running it: checking -# out a branch where this script does not exist would otherwise pull the rest of -# it out from under the interpreter. -{ -set -euo pipefail - -HASH_RUNS=5 # hash-throughput runs per branch, best kept -HASH_COOLDOWN=10 # seconds between them, to let the machine settle -LOG_LIFETIME=30 # XMSS lifetime, as log2 of the number of epochs -COMPRESSIONS=390 # compressions per epoch of an XMSS keygen - -REPEAT=5 # measured proving passes per benchmark, after a warmup -COOLDOWN=5 # idle seconds before each of them -LOG_INV_RATE=1 # for the two aggregations -REC_LOG_INV_RATE=2 # for recursion -REC_N=2 # child aggregates per recursion node - -BRANCHES=(main sha2 sha3) -HASHES=( BLAKE2s SHA-256 Keccak) -XMSS=( 900 450 205) # signatures per leaf, per that branch's README -SPHINCS=( 245 122 55) # likewise, a leaf of SPHINCS signatures -PER_LEAF=(900 900 450) # likewise, the XMSS leaves a recursion node covers -FLOCK_LOG=(18 17 16) # likewise, log2 of flock's batch of compressions - -TEST=multithreaded_throughput -PACKAGE=primitives -BIN=hash_bench - -cd "$(dirname "$0")/.." - -if [ -n "$(git status --porcelain --untracked-files=no)" ]; then - echo "working tree has changes; commit or stash them first" >&2 - exit 1 -fi -start=$(git symbolic-ref --quiet --short HEAD || git rev-parse HEAD) - -# Two of these at once check out branches under each other, so the runs measure -# whatever branch the other one left in the tree. Take the lock before the first -# checkout, and before arming the trap that undoes it. -LOCK="$(git rev-parse --git-dir)/bench.lock" -if ! mkdir "$LOCK" 2>/dev/null; then - echo "another ./doc/bench.sh is running; wait for it, or remove $LOCK" >&2 - exit 1 -fi -trap 'rmdir "$LOCK"; git checkout --quiet "$start"' EXIT INT TERM - -git fetch --quiet --all - -# The number following $1 in $2, commas stripped. -num() { - local line - line=$(grep -m1 "$1" <<<"$2") || { echo "no \"$1\" line in the output" >&2; exit 1; } - sed -E "s#.*$1[^0-9]*([0-9,]+(\.[0-9]+)?).*#\1#" <<<"$line" | tr -d ',' -} - -# Best of HASH_RUNS `multithreaded_throughput` runs, in Mhash/s. -hash_rate() { - local r out line - local results=() - for ((r = 1; r <= HASH_RUNS; r++)); do - out=$(cargo test --release -p "$PACKAGE" --test "$BIN" "$TEST" -- --exact --ignored --nocapture) - line=$(grep -m1 'Mhash/s' <<<"$out") || { echo "no measurement in the output" >&2; exit 1; } - results+=("$(sed -E 's#.*[^0-9.]([0-9]+(\.[0-9]+)?) Mhash/s.*#\1#' <<<"$line")") - printf 'run %d/%d: %s\n' "$r" "$HASH_RUNS" "$line" >&2 - if ((r < HASH_RUNS)); then sleep "$HASH_COOLDOWN"; fi - done - printf 'samples: %s\n' "${results[*]}" >&2 - printf '%s\n' "${results[@]}" | sort -g | tail -1 -} - -# `aggregate` of $1 signatures of scheme $2, in signatures per second. -aggregate() { - local out - out=$(cargo run --release -- aggregate "--$2" "$1" --log-inv-rate "$LOG_INV_RATE" \ - --repeat "$REPEAT" --cooldown "$COOLDOWN") - echo "$out" >&2 - num 'per signature' "$out" -} - -# Flock's batch proving of 2^$1 compressions, in compressions per second. -flock_rate() { - local out - out=$(BENCH_REPEAT="$REPEAT" BENCH_COOLDOWN="$COOLDOWN" FLOCK_N_LOG="$1" \ - cargo test --release --package flock --test batch_proving_hashes -- \ - hash_batch_prove_verify --exact --nocapture --include-ignored) - echo "$out" >&2 - num 'throughput' "$out" -} - -# `recursion` over REC_N leaves of $1 XMSS signatures, in seconds. -recursion() { - local out - out=$(cargo run --release -- recursion --n "$REC_N" --xmss-per-leaf "$1" \ - --log-inv-rate "$REC_LOG_INV_RATE" --repeat "$REPEAT" --cooldown "$COOLDOWN") - echo "$out" >&2 - num 'proving time' "$out" -} - -rows="" -for i in "${!BRANCHES[@]}"; do - b=${BRANCHES[$i]} - git checkout --quiet "$b" - git merge --quiet --ff-only "origin/$b" - echo "=== $b ===" - rows+=$(printf '%s\t%s\t%s\t%s\t%s\t%s\t%s %s\n' \ - "$b" "${HASHES[$i]}" "$(hash_rate)" "$(flock_rate "${FLOCK_LOG[$i]}")" \ - "$(aggregate "${XMSS[$i]}" xmss)" "$(aggregate "${SPHINCS[$i]}" sphincs)" \ - "${PER_LEAF[$i]}" "$(recursion "${PER_LEAF[$i]}")")$'\n' - echo -done - -awk -F'\t' -v l="$LOG_LIFETIME" -v c="$COMPRESSIONS" -v n="$REC_N" ' -function dur(t) { - if (t < 60) return sprintf("%.1f s", t) - else if (t < 3600) return sprintf("%.1f min", t / 60) - else if (t < 86400) return sprintf("%.1f h", t / 3600) - else return sprintf("%.1f days", t / 86400) -} -BEGIN { printf "%-6s %-8s %10s %10s %20s %9s %11s\n", "branch", "hash", "compr/s", "proven/s", - sprintf("keygen (2^%d)", l), "XMSS/s", "SPHINCS/s" } -{ - split($7, r, " ") - t = 2 ^ l * c / ($3 * 1e6) - printf "%-6s %-8s %8.0f M %8.0f K %20s %9.0f %11.0f\n", $1, $2, $3, $4 / 1e3, - sprintf("%.0f s (%s)", t, dur(t)), $5, $6 - recs = recs sprintf("%-6s %-8s %10d %9.3f s\n", $1, $2, r[1], r[2]) -} -END { printf "\n%-6s %-8s %10s %11s\n%s", "branch", "hash", "XMSS/leaf", sprintf("%d->1", n), recs } -' <<<"${rows%$'\n'}" -} diff --git a/doc/sphincs/latexmkrc b/doc/sphincs/latexmkrc deleted file mode 100644 index ec7c0e475..000000000 --- a/doc/sphincs/latexmkrc +++ /dev/null @@ -1,4 +0,0 @@ -$pdf_mode = 1; -$out_dir = '.build'; -$bibtex_use = 2; -$clean_ext = 'bbl synctex.gz'; diff --git a/doc/sphincs/main.tex b/doc/sphincs/main.tex deleted file mode 100644 index cd755fe1b..000000000 --- a/doc/sphincs/main.tex +++ /dev/null @@ -1,471 +0,0 @@ -% Build with: latexmk -pdf main.tex -\documentclass[11pt]{article} - -\usepackage[T1]{fontenc} -\usepackage{lmodern} -\usepackage[margin=1in]{geometry} -\usepackage{microtype} -\usepackage{amsmath,amssymb,amsthm,mathtools} -\usepackage{booktabs} -\usepackage{enumitem} -\usepackage{xcolor} -\usepackage[colorlinks=true,linkcolor=blue!50!black,citecolor=blue!50!black,urlcolor=blue!50!black]{hyperref} - -\theoremstyle{plain} -\newtheorem{theorem}{Theorem}[section] -\theoremstyle{definition} -\newtheorem{definition}[theorem]{Definition} - -\newcommand{\bits}[1]{\{0,1\}^{#1}} -\newcommand{\getsr}{\stackrel{\$}{\gets}} -\newcommand{\Th}{\mathsf{Th}} -\newcommand{\Enc}{\mathsf{Enc}} -\newcommand{\Digest}{\mathsf{Digest}} -\newcommand{\Gen}{\mathsf{Gen}} -\newcommand{\Sig}{\mathsf{Sig}} -\newcommand{\Ver}{\mathsf{Ver}} -\newcommand{\hash}{\mathsf{H}} -\newcommand{\LE}{\mathsf{LE}} -\newcommand{\Truncate}{\mathsf{Truncate}} -\newcommand{\concat}{\mathbin\Vert} -\newcommand{\sk}{\mathit{sk}} -\newcommand{\pk}{\mathit{pk}} -\newcommand{\rootnode}{\mathit{root}} -\newcommand{\tw}{\mathit{tw}} -\newcommand{\qs}{q_{\mathrm{s}}} -\newcommand{\amax}{A_{\max}} -\newcommand{\cmax}{C_{\max}} -\newcommand{\idx}{\mathit{idx}} -\newcommand{\lay}{\mathit{lay}} -\newcommand{\OtsSign}{\mathsf{Ots.sign}} -\newcommand{\OtsLeaf}{\mathsf{Ots.leaf}} -\newcommand{\TreeRoot}{\mathsf{Tree.root}} -\newcommand{\TreePath}{\mathsf{Tree.path}} -\newcommand{\TreeFold}{\mathsf{Tree.fold}} -\newcommand{\FtsKey}{\mathsf{Fts.key}} -\newcommand{\FtsOpen}{\mathsf{Fts.open}} -\newcommand{\FtsRec}{\mathsf{Fts.recover}} - -\emergencystretch=1.5em -\setlist[enumerate]{leftmargin=2em, itemsep=4pt} - -\title{A SPHINCS$^+$ variant} -\author{} -\date{} - -\begin{document} -\maketitle - -\begin{abstract} - -This specification defines a SPHINCS$^+$ variant with a lifetime of $2^{24}$ signatures per key pair: - -\begin{itemize} - \item \textbf{NIST security level~1}~\cite{NISTPQC}, $\approx 128$-bit (resp. $\approx 64$-bit) of classical (resp. quantum) security, in the Random Oracle model, ROM (resp. Quantum Random Oracle model, QROM).\footnote{The classical bound is proved (Section~\ref{sec:security}); the quantum one is a target, not a proved statement (Section~\ref{sec:quantum}).} - \item \textbf{Public key: 32 bytes.} - \item \textbf{Signature: 4924 bytes.} - \item \textbf{Verification: 497 hashes.} - \item \textbf{Key generation: 1.38M hashes.} - \item \textbf{Signing: approximately 190K hashes on average}, with a 1024-byte public cache. -\end{itemize} -\end{abstract} - -The construction uses compressed variants of the Winternitz one-time signature (WOTS) and forest of random subsets (FORS), called WOTS$^+$C and FORS$^+$C~\cite{HK22C,KN25}. They are combined within the SPHINCS$^+$ framework~\cite{SPHINCSPLUS}. The definitions below specify the concrete variant completely. - -\section{Hashing and parameters} -\label{sec:parameters} - -Write $\bits r$ for $r$-bit strings, $\concat$ for concatenation, and $\LE_r(z)$ for the unsigned $r$-bit little-endian encoding of $z$. Write $x\getsr A$ for a uniform sample from $A$. Indices and bit positions start at zero; $\bot$ denotes failure. - -Let $\hash:\bits{*}\to\bits{256}$ be the hash function. Most operations use its first $n=128$ output bits: -\[ - \Th(P,\tw,M)=\Truncate_n\!\left(\hash(\tw\concat P\concat M)\right). -\] -Here $P$ is a 128-bit public parameter and $\tw$ is a 128-bit \emph{tweak}: an address identifying the operation and its position in the construction. Appendix~\ref{sec:tweaks} gives every tweak's exact bytes. $\Truncate_r$ always keeps the first $r$ bits, with bits read least significant first within each byte. - -The message $m$ and master secret $S$ are each 256 bits. WOTS signs 128-bit values $M$. The signature contains a 128-bit randomizer $\rho$ and one 32-bit counter $c$ per WOTS signature. Derived signing secrets, chain values and Merkle nodes are 128 bits; write $\mathcal H=\bits n$. - -\begin{center} -\begin{tabular}{@{}lll@{}} -\toprule -Symbol & Value & Meaning\\ -\midrule -$w$ & $3$ & bit-size of WOTS chain positions\\ -$v$ & $42$ & chains per WOTS key\\ -$T$ & $191$ & sum of the signed chain positions\\ -$d$ & $3$ & layers, numbered from the top\\ -$(h_0,h_1,h_2)$ & $(12,7,7)$ & tree heights at those layers\\ -$h$ & $26$ & total height, $h=h_0+h_1+h_2$\\ -$a$ & $10$ & height of each FORS tree\\ -$k$ & $15$ & digest indices, of which $k-1$ open trees\\ -$\qs$ & $2^{24}$ & signing requests allowed per key pair\\ -$\amax$ & $2^{32}$ & maximum randomizer trials per signature\\ -$\cmax$ & $2^{32}$ & maximum encoding attempts per WOTS signature\\ -\bottomrule -\end{tabular} -\end{center} - -The public parameter and all signing secrets are derived from $S$. In particular, -\[ - P=\Th\!\left(0^{128},\mathsf{tw}_{\mathrm{parameter}},S\right). -\] -The component definitions below derive signing secrets as needed from this fixed $S$. Functions that read secrets use $S$ implicitly. - -\section{WOTS: signing with hash chains} -\label{sec:ots} - -A WOTS key is identified by a layer $\lay$, a tree $\tau$ and a leaf $e$. It contains $v=42$ chains, each with eight positions, numbered $0$ through $7$. Fix such a key. For $0\leq i\lay}h_j}\right\rfloor\bmod 2^{h_\lay}. -\] -Concretely, $e_0$ is the high 12 bits of $\idx$, $e_1$ the next 7 bits, and $e_2$ the low 7 bits. Then -\[ - \tau_0=0,\qquad \tau_1=e_0,\qquad \tau_2=2^7e_0+e_1. -\] -Thus the key at $(0,0,e_0)$ signs tree $(1,\tau_1)$'s root, the key at $(1,\tau_1,e_1)$ signs tree $(2,\tau_2)$'s root, and the key at $(2,\tau_2,e_2)$ signs $\FtsKey(P,\idx)$. - -\paragraph{Selecting the route and FORS leaves.} -\label{rem:rho} -For public root $\rootnode$, message $m$ and randomizer $\rho$, define -\[ - \Digest(P,\rootnode,m,\rho)=\Truncate_{h+ka}\!\left( - \hash(\mathsf{tw}_{\mathrm{msg}}\concat P\concat\rho\concat\rootnode\concat m)\right). -\] -Interpret these $h+ka=176$ bits as a little-endian integer $N$ and extract -\[ - \idx=N\bmod2^h,\qquad - u_\kappa=\left\lfloor N/2^{h+\kappa a}\right\rfloor\bmod2^a, - \quad 0\leq\kappa