feat(rewards): always-on prover loop (DHT, mirror-coin gate, capsule challenge, entry writes) - #593
Conversation
Recovered from an uncommitted working tree after a session cap. NOT YET WIRED into dig-node-core (no `mod rewards;`, no mod.rs) and NEVER COMPILED -- committed as a checkpoint so the next pass cannot lose it again, per CLAUDE.md invariant 3. Contents, each transcribed against dig-rewards-coin SPEC.md v0.1.1: - spec_constants.rs -- every normative bound, tagged with its clause - port.rs -- the RewardsChainPort seam (dig-rewards-coin is SPEC-only, #3249) plus UnavailableChainPort, which reports Unavailable rather than silently no-op'ing - state.rs -- the SPEC 2.3 status record, deliberately with NO health boolean and NO precomputed staleness (2.4), plus an injectable Clock a test can refuse to advance - admission.rs -- THE single admission point (5.3); self-exclusion on BOTH coordinates, the puzzle-hash coordinate checked before the gate's eligibility is trusted - gate.rs -- the mirror-coin gate: advertises + declares_peer -> owner_puzzle_hash, fail-closed on every absence Known defects carried, fixed under review in this PR, not in this commit: - gate.rs passes the census ordinal `n` to `advertises`, but SPEC 4.6 clause 1 requires `n-1` exactly; as written it admits coins the census excludes - gate.rs evaluate() hardcodes in_grace_window=false, so the SPEC 4.6.3 grace window is unreachable in production and MIRROR_EPOCH_GRACE_SECONDS is dead - the test named absent_declaration_is_ineligible actually asserts PeerNotDeclared - no NewEpoch spend (SPEC 2.1 clause 2) Refs #3250
d76a0d2 to
21e52ed
Compare
Parent-side checkpoint taken while the implementer lane was blocked on a cold workspace compile. NOT YET VERIFIED -- this commit has never completed a build. It exists because this ticket has now lost uncommitted work to a dying process twice (five files in dig-node, two in dig-rpc-protocol) and a third loss was one crash away. - mod.rs + `pub mod rewards;` in lib.rs: the module is finally part of the build - cycle.rs: the always-on loop (period, heartbeat, cycle deadline) - writes.rs: the four SPEC 6.3 entry-write bounds - staleness.rs: the SPEC 12.4 chain-derived EntrySetStale computation - challenge.rs: SPEC 3.2 window selection and the 3.5 pass/fail decision - admission.rs / gate.rs / port.rs: in-progress fixes for the four defects the salvage commit recorded, plus D5 (an absent mirror-collateral epoch ordinal is a prover fault and must not strike a peer) - .gitkeep removed now that mod.rs exists Refs #3250
The five defects in this branch, with their fixes — written down so they survive a capThis diagnosis currently exists only in an orchestrator's context. This epic has already lost Every clause reference is to the merged D1 — the census ordinal is off by one, and it pays the wrong parties in BOTH directions
§4.6 clause 1: "A coin qualifies for the census of mirror-collateral epoch This is not a conservative error. Passing D2 — the §4.6.3 grace window is unreachable, so the constant is dead
That clause exists because "without the grace window every mirror in the network becomes Fix: General tell worth keeping: a named constant with no reader is a spec clause that silently did D3 — a test's name contradicts its assertion
D4 — no
|
…ike test The two errors that kept dig-node-core from compiling. OwnIdentity derived Copy while carrying `controlled_puzzle_hashes: Vec<[u8; 32]>`, which Vec cannot satisfy (E0204). Copy dropped, Clone kept. The field stays a growable collection deliberately: SPEC 5.2 excludes a candidate on MEMBERSHIP in the set of puzzle hashes this node's wallet controls, not on equality to one distinguished value, because a wallet may hold more than one payout address. `prover_fault_never_increments_any_peer_strike` bound its tracker `mut`, which clippy rejects under -D warnings. It does not need mut, and that is the point worth noticing rather than papering over: `record_prover_fault` takes `&self` because a prover fault must not be able to touch a peer's strike counter at all (SPEC 3.6 clause 4). The borrow checker now carries that invariant, so the absent `mut` is evidence of the design rather than an oversight. Committed from the parent side: the lane had both edits correct in its working tree but uncommitted, and this ticket has already lost or nearly lost in-flight work four times to a dying process. Refs #3250
…x a paused-time race The last three CI failures. No production behaviour changes. - challenge.rs: `HashMap<(Bytes32, Bytes32), Vec<(u32, Bytes32, u64)>>` tripped clippy::type_complexity. Factored into `ChallengeSubject` (the `(peer_id, launcher_id)` a no-repeat rule is keyed on) and `IssuedWindow` (`(cycle_index, resource_id, offset)`). The lint was right: those tuples were unreadable at the use site. - cycle.rs: passed `|| std::future::pending::<()>()` where the function itself does (clippy::redundant_closure). - cycle.rs: `heartbeat_loop_fires_on_its_own_timer` failed against CORRECT production code. Under `start_paused`, `tokio::spawn` does not poll the task, so `tokio::time::advance` jumped over a timer the loop had not registered yet; the loop then slept from the far side of the jump and never ticked. Yielding once before the advance lets it register the timer first. The assertion in that test is deliberately UNCHANGED. It is the only evidence that the heartbeat fires on its own timer rather than merely when called directly, and an always-on loop is trivially easy to keep green while it never runs -- so the fix had to be the sequencing, never the bound. Committed from the parent side: the lane had all three edits correct but uncommitted, the sixth time in-flight work on this ticket needed rescuing. Refs #3250
- Introduce type aliases ChallengeSubject and IssuedWindow in challenge.rs to address clippy::type_complexity lint (line 49) - Remove redundant closure wrapper from pending future in cycle.rs test (line 134) to address clippy::redundant_closure - Fix heartbeat_loop_fires_on_its_own_timer test by adding yield_now() before virtual time advance to ensure timer is registered Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
MichaelTaylor3d
left a comment
There was a problem hiding this comment.
Independent correctness review — PR #593, head efe57f1154fb0ab5a518d710ff88263c8aa989b6
Verdict: CHANGES-REQUIRED — one blocking finding (below), rest non-blocking. This PR's own D1-D5 diagnosis (comment above) is verified as correctly fixed in this diff, with tests placed at the decision level, not inside the fake.
D1-D5 verification
- D1 (census ordinal n vs n-1): FIXED.
gate.rs::evaluatecomputescensus_epoch = current_epoch.checked_sub(1), guardsn == 0(returnsDoesNotAdvertise, not a panic/wrap), and callsadvertises(.., census_epoch).census_epoch_is_n_minus_1_not_nproves the negative (declaringndirectly is rejected) andcensus_epoch_n_minus_1_is_admittedproves the positive. Both assert atevaluate(), not insideFakeReader. Confirmed. - D2 (grace window unreachable): FIXED.
EpochContext::in_grace_window()is a real reader overepoch_rolled_over_at/now, wired intoevaluate()viaadvertises_current_or_previous, andgrace_window_makes_previous_census_ordinal_admissible_via_evaluateexercises it through the publicevaluate()entrypoint (not the private helper) — this is the reachability D2 asked for.MIRROR_EPOCH_GRACE_SECONDSnow has a real reader. Confirmed. - D3 (test name vs assertion): FIXED.
absent_declaration_is_ineligiblenow setsdeclaringto the correct peer but leavesownersempty, soowner_puzzle_hash()resolvesNoneand the assertion is genuinelyAbsentDeclaration, distinct fromPeerNotDeclared. Confirmed. - D4 (no
NewEpochspend): FIXED at the seam.RewardsChainPort::spend_new_epochexists (port.rs), documented against §2.1 clause 3 (two willing spenders is correct, not a conflict; a not-yet-rolled epoch is not an error). The real spend logic is out of scope (#3249) — the seam is what this PR owns. Confirmed. - D5 (absent epoch ordinal as prover fault, not peer strike): FIXED, and well.
GateError::EpochOrdinalUnavailableis a distinct type fromGateIneligibleReason— not constructible as one — andAdmissionDecision::ChainSourceUnavailableis likewise a distinct variant fromGateIneligible.StrikeTracker::record_prover_faulttakes&self, not&mut self— it is compile-time incapable of mutating the strike map, which is a stronger guarantee than a runtime no-op. Confirmed structurally incapable, not just tested incapable.
Numbered items 1-4
- §2.4 honesty (no health boolean, no precomputed staleness): partially verified — see blocking finding below. No
healthy/ok/up/runningfield exists onRewardProverStatus, and the negative test checks JSON keys (not substrings, correctly avoiding thecontains("running")-on-Running-value trap). But the check is not recursive — see inline comment onstate.rs:178. - Wedged-loop tests: intact, confirmed the best possible outcome.
wedged_cycle_is_abandoned_at_the_deadline_and_does_not_fake_completionassertsconsecutive_cycle_failures == 1andlast_cycle_completed_at == Noneafter an abandoned cycle — not just the returned bool.heartbeat_loop_fires_on_its_own_timerretainstokio::task::yield_now().awaitbeforetokio::time::advance(the sequencing fix), and the assertion is exactlyobserved_at >= 1_000 + PROVER_HEARTBEAT_SECONDS— not weakened, not toleranced, not#[ignore]d. This is the single most important thing in this PR and it is correct. - §3.6 cl.4 prover-fault-never-strikes: verified structurally. See D5 above —
record_prover_fault(&self, ...)cannot touch the map by construction, and the test asserts the strike counters directly (consecutive_failures(peer, LAUNCHER) == 0for every peer), not just the reported state. Note: there is no top-level "run one full cycle across N candidates" function in this PR that would wire aGateError/AdmissionDecision::ChainSourceUnavailableintochallenge::StrikeTrackerat all — the two live in separate modules with no call path between them today. That is the strongest form of "cannot strike" (no path exists), but it also means the end-to-end wiring is unverified because it doesn't exist yet; that's consistent with "RPC handler wiring" being out of scope for this PR, so not blocking on it, but flagging for whoever writes that orchestrator next. - §6.3 cl.4 cooldown keyed on (payout_puzzle_hash, launcher_id), never peer_id: verified, and well — structurally, not just by test.
EntryWriteScheduler::is_in_reentry_cooldowndoesn't take apeer_idparameter at all, so there is no code path by which peer identity could select a different key.reentry_cooldown_survives_a_fresh_peer_id_for_the_same_payout_hashis the direct test.
Modulo-bias assessment (§3.2 cl.4, csprng_u64_below)
u64::from_le_bytes(buf) % bound is textbook modulo-biased: outcomes in [0, u64::MAX % bound) are drawn with slightly higher probability than the rest. Not exploitable for window prediction here: for any realistic bound (resource-length sums up to petabyte scale, or window-length-scale offsets, i.e. bound << 2^64), the bias per outcome is at most bound/2^64, roughly 2^-40 or smaller for any resource under ~16 TiB — far below what an adversary could detect or exploit to predict which offset/resource will be chosen. Non-blocking, but worth a one-line comment at the call site (or a rejection-sampling loop) so a future reader doesn't have to re-derive this analysis — left inline as a non-blocking suggestion.
Other checks (§5.2/§5.3 self-exclusion, §3.5/§3.6 bounds, §12.4/§12.6, unreferenced constants)
- §5.3 discovery paths:
DiscoveryPathhas exactly 4 variants; SPEC §5.3 cl.2 names a 5th (the off-chain hint), but that path is explicitly deferred (§13.2, dig_ecosystem#3252) and does not exist as code anywhere in this repo yet — there is nothing to bypass. When #3252 ships, that PR must route throughadmission::admittoo; flagging as a forward note, not a defect in this PR. - §5.3.4 control test: present and correct (
control_a_non_self_candidate_is_admitted_on_every_path) — distinguishes "excluded self" from "dropped everything". - §6.3 bounds (batch cap, rate, fee budget, withheld-not-dropped): all present with direct tests (
batch_cap_leaves_the_rest_pending,rate_limit_withholds_rather_than_drops,fee_budget_exhaustion_keeps_decisions_and_stops_writing). - §12.4/§12.6:
is_entry_set_stalecorrectly treatsNoneas maximally stale once the distributor is old enough (not "unknown"/false-default) and is a free function separate fromstate.rs, so it cannot leak onto the prover status record by accident.is_unfunded/is_entry_set_fullare pure predicates that don't touch the entry set — correctly leave eviction-on-Unfunded unimplemented rather than wrongly implemented. - Unreferenced constants (non-blocking, flagged because D2's own writeup calls this exact pattern out):
CHALLENGE_WINDOWS_PER_CYCLE,CHALLENGE_DEADLINE_SECONDS,CHALLENGE_PEER_DEADLINE_SECONDS,CHALLENGE_MIN_INTERVAL_SECONDS,CHALLENGE_MAX_PEERS_PER_CYCLE,MAX_MIRROR_URL_TERMShave no reader anywhere in this diff. UnlikeMIRROR_EPOCH_GRACE_SECONDSbefore this PR's fix, these are legitimately out of scope right now — they belong to the concretedig.fetchRangetransport and the per-cycle peer-selection/URL-parsing logic this PR explicitly defers — butCHALLENGE_WINDOWS_PER_CYCLE = 4(select 4 windows per peer per cycle) is core §3.2 soundness logic, andselect_windowonly ever produces one window per call with no caller in this diff that invokes it 4 times. Non-blocking since the loop presumably lives in the not-yet-written cycle orchestrator, but leaving a TODO, or at minimum not letting it silently stay unreferenced past the next PR, would help the next reader trust it shipped. - Attacker-controlled strings (§3.7 cl.4):
ChallengeFailure::RpcError(String)andChainPortError::Other(String)are the only free-form strings carried from a fault; neither is logged, interpolated, or displayed anywhere in this diff, so there is nothing to bound yet — correctly deferred to whichever module actually logs them.
Blocking finding
crates/dig-node-core/src/rewards/state.rs:178-190 — the §2.4 honesty negative test (serialized_status_has_no_health_or_staleness_key) checks json.as_object() at the top level only. It does not recurse into nested objects (e.g. counters). Today nothing nested carries a forbidden key, so the test currently passes correctly — but this is exactly the shape of test the brief for this review flagged as the highest-risk item in the whole PR, and a flat top-level check gives no protection against a future field added one level down (e.g. inside ProverCounters, or a new nested sub-record) carrying running/healthy/stale. Please make this walk the full serde_json::Value tree (objects and arrays) checking every key, not just the top-level map's keys, so the guarantee holds regardless of where a smuggled boolean gets added later.
Findings posted as inline threads below. KG: NONE (gate review, no new dig-pattern beyond what the PR's own D1-D5 comment already recorded).
Adversarial gate (third leg) — RATIFY-WITH-CONDITIONSHead SHA read: Three blocking conditions (C1–C3), two follow-up tickets (C4–C5). C1 is the one I would not merge without. 1. Is it safe to merge an autonomous spender whose chain adapter always returns
|
| Quantity | Bound | Value |
|---|---|---|
| Bundles/day (rate, §6.3 cl. 2) | 86,400 / 3,600 | 24 |
| Fee budget/day (§6.3 cl. 3) | 24 × F (FeeBudget::new) |
24 F |
| Entry actions/day | 24 × MAX_ENTRY_WRITES_PER_BUNDLE |
192 |
| Max XCH/day at F = 0.000005 XCH | 0.00012 XCH (~$0.002) | |
| Max XCH/day at a congestion F = 0.01 XCH | 0.24 XCH/day ≈ 88 XCH/yr | |
| Evict+re-add churn per entry/day | 6 h cooldown + 3 hourly strikes ≈ 9 h | ~2.4 |
| Max evictions/day (192 actions, 2 per churn) | ≤ 96, over a ≤ 250-entry set | |
| Max $DIG settled out of reserve/day by eviction | the entry set's whole accrued-but-unclaimed balance; ~1.3 days to flush all 250 | bounded by scheduled emission, plus all sub-payout_threshold dust (§6.4) |
Two findings fall out of that table.
(a) Bound 3 is not an independent bound. 24 bundles/day is simultaneously the rate ceiling and the fee ceiling. Under nominal fees the budget never binds — it bites only when the fee_mojos passed to decide exceeds the configured standard fee. Fine as a fee-escalation catch, but it means the only thing between the funder and >24 bundles/day is EntryWriteScheduler::last_bundle_sent_at.
(b) last_bundle_sent_at, cooldown_until and spent_mojos_today are in-memory with no persistence anywhere in the diff (no serde, no store, nothing). SPEC §12.1 cl. 2 is normative and explicit:
challenge strikes MUST reset to zero … cooldowns and fee budgets MUST persist … Lose the evidence, keep the bounds.
This PR implements the first half (StrikeTracker::reset_all, correctly) and drops the second half entirely. The doc comment in writes.rs says "One instance per running prover (not per cycle) so both bounds persist across cycles" — across cycles, which is true and is not what the SPEC asked for. Consequence: a crash-looping or frequently-restarted node writes one bundle per restart with no interval, no daily cap, and an empty re-entry cooldown map — exactly the churn drain §6.3 exists to prevent, arriving "at exactly the moment a restart loop would hit it hardest", in the SPEC's own words. The bound is a habit, not a bound.
On the number itself: 24 F/day/distributor is a number a funder would accept if they could see it. It is stated nowhere — not in the PR body, not in mod.rs, not in spec_constants.rs (which tags clauses but never composes them). A bound nobody can compute is not a bound anyone consented to. C3.
3. Do the wedged-loop tests bind? — Partly. Better than expected on the heartbeat; nothing at all on the cycle loop.
heartbeat_loop_fires_on_its_own_timer does bind: observed_at starts at 1,000 and is only ever written by heartbeat_tick, so observed_at >= 1_000 + 60 is unreachable unless the spawned loop actually ticked — clock.advance moves the fake clock, not observed_at. The yield_now fix is legitimate and did not weaken it. It would still pass under an interval shorter than 60 s (harmless direction) and fails under a longer one or a loop that never registers a timer. Keep it.
The real hole is one level up: there is no cycle loop to test. run_cycle_with_deadline is single-shot; nothing repeats it every PROVER_CYCLE_PERIOD_SECONDS, so there is nothing that could "silently stop running in production" and nothing that tests it. The honesty story is instead carried by is_wedged — a reader-side derivation against the reader's own clock that a wedged writer cannot flatter (§2.4) — plus §12.4 staleness from chain spend history. The design answers the question even though the tests cannot yet.
Decision: the surviving tests are adequate for what this PR contains; the missing test is missing because the code under test does not exist. It becomes mandatory in the wiring PR (C4): the periodic loop re-fires after a cycle fails, and is_wedged goes true when the spawned task is aborted. Does not block.
4. Is self-exclusion enforced by the type system or by convention? — By convention, and that should change here, not later.
admit() is correct on both coordinates (peer_id and controlled_puzzle_hashes, §5.2) and returns AdmissionDecision::Admit { payout_puzzle_hash }. But EntryAction::Add { payout_puzzle_hash, launcher_id } has public fields and no relationship to AdmissionDecision. Any future path can construct an Add and hand it to EntryWriteScheduler::decide with admit never running. DiscoveryPath has four variants; §5.3 cl. 2 names five paths; the fifth (#3252 off-chain hint) will be added by a lane that has not read §5.3. dig-node#261's lesson applies verbatim.
Decision: make it a compile error, in this PR. It blocks. Reason: ~40 lines and no caller exists yet to update; after #3249 wires callers it becomes a refactor across a live money path, which is when nobody does it. Failure direction of deferring: the node writes its own payout puzzle hash into a distributor it funds and pays itself the funder's $DIG — self-dealing, invisible in the chain view, and §5.1 calls the rule absolute. Shape (C2): pub struct AdmittedPeer { payout_puzzle_hash, launcher_id } with private fields, returned only by admit, and EntryAction::Add(AdmittedPeer) — so Add is unconstructible outside admission. Remove stays freely constructible; evicting yourself is not the hazard.
5. What is missing that nobody notices until it costs money?
The operator's journey against the SPEC's own §12 list:
| SPEC | State in this diff |
|---|---|
| §12.1 restart — strikes reset | Implemented (StrikeTracker::reset_all). The doc cites "§12.1 clause 3"; it is clause 2 — clause 3 is the status-record rule. |
| §12.1 restart — cooldowns/fee budget persist | ABSENT, and normatively required. C1. |
§12.1 restart — status record published before the first cycle, Running with last_cycle_completed_at absent |
Shape implemented in state.rs; nothing publishes it (no RPC wiring — out of scope, C4). |
§12.2 reorg — 32-block finality, re-derive not replay, never a strike, surface as ChainSourceUnavailable |
ABSENT entirely. Zero references to reorg or finality; CENSUS_FINALITY_DEPTH_BLOCKS is not imported; submit_entry_writes has no notion of unconfirmed. C5. |
| §12.3 funded, no prover | Chain-observable; not this crate's surface. Correctly absent. |
| §12.4 stale entry set | Implemented (staleness.rs, chain-derived, kept off the prover status record per §2.4). The strongest part of this PR. |
| §12.5 claim after eviction | #3251's claim loop. Correctly absent. |
| §12.6 reserve exhausted | Implemented as a predicate (is_unfunded, entry set kept). No caller yet. |
Two further blind spots, neither blocking:
decidecharges the fee and records the removal cooldown before the bundle is submitted. For the fee that is the safe direction. For the cooldown it is not: ifsubmit_entry_writesfails, an honest mirror is held out for 6 h for a removal that never reached the chain. Harms a peer, not the funder — fix in the wiring PR.FeeBudget's day rolls only insidetry_spend, and there is no read path, so no operator surface can answer "how much of today's budget is left?". The number exists and is unobservable. Fold into C3.
Conditions
C1 — BLOCKS. Persist the write bounds, or fail closed without them. Do not build a storage layer here. Introduce a CooldownStore-style trait seam that EntryWriteScheduler must be constructed with (delete the no-arg new()/Default so an unpersisted scheduler is not expressible), carrying last_bundle_sent_at, cooldown_until and the FeeBudget day counters. Ship a NoPersistence impl whose presence makes decide return WriteOutcome::Pending — a prover with no durable bound refuses to write rather than writing unbounded. Tests: a scheduler reloaded from a store still reports is_rate_limited and is_in_reentry_cooldown true; a NoPersistence scheduler never returns Bundle. Failure direction if skipped: a restart loop drains the funder's XCH while the distributor looks healthy.
C2 — BLOCKS. AdmittedPeer newtype with private fields, minted only by admit, required by EntryAction::Add. As in Q4, with the existing admission tests threaded through the new type.
C3 — BLOCKS (doc-only, ~15 lines). State the worst case in money where a funder will read it. Put the Q2 rows (24 bundles/day, 192 actions/day, 24 × standard_fee XCH/day, eviction settles the accrued balance including sub-threshold dust per §6.4) in rewards/mod.rs's module doc and in this PR's body. Say plainly that the cap scales linearly with the configured standard fee.
C4 — follow-up ticket, must exist before merge and be linked as blocking the wiring PR.
dig-node: gate the rewards prover before it can spend — wiring PR requirements. The PR that firstspawns the prover loop (not #3249's driver, and not in the same PR as it) MUST land: a config key defaulting to off; an explicit operator opt-in that names the daily XCH maximum from §6.3 cl. 3; a dry-run mode that runs a full cycle and logs the bundle it would submit without callingsubmit_entry_writes; a kill switch that stops the loop without stopping the node; a periodic-loop test proving the cycle re-fires after a failed cycle and thatis_wedgedgoes true when the task is aborted; and a full triple gate on the composed system. Refs dig_ecosystem#3250, dig-node#593.
C5 — follow-up ticket.
dig-node: implement SPEC §12.2 reorg handling in the rewards prover. Treat a submitted bundle as unconfirmed until buried bydig_mirror_collateral::CENSUS_FINALITY_DEPTH_BLOCKS = 32(reuse the constant; do not introduce a second finality number); on an unwound entry write re-derive the entry set from the new chain view and decide again, never replay the bundle; a reorg MUST NOT produce a challenge strike or an eviction, and MUST surface asChainSourceUnavailableor aconsecutive_cycle_failuresincrement. Refs dig_ecosystem#3250, dig-node#593.
What I would not change
The RewardsChainPort + UnavailableChainPort seam is the right call, and Unavailable-not-a-silent-no-op is the honest one. staleness.rs deriving §12.4 from chain spend history rather than a self-report; is_wedged comparing against the reader's own clock; the cooldown keyed on (payout_puzzle_hash, launcher_id) with peer_id not even a parameter; and GateError::EpochOrdinalUnavailable kept distinct from ineligibility so a chain outage cannot strike a peer — all four are the failure-direction-correct choice, and the provenance did not damage them.
Given that provenance (two caps, committed pre-compile, tests written after), I weighted "what would these tests still pass under" throughout. The answer that mattered: they would pass under a prover that loses every write bound on restart. That is C1.
Security audit — loop-security — PR #593Verdict: PASS Scope note that changes the calibrationThis diff adds The five explicit answers
Other findings
What I did not coverThe #3249 driver, the Posted by loop-security, read-only, no edits made. |
Parent-side checkpoint, UNVERIFIED -- this commit has not completed a build and may be mid-edit. Taken because in-flight work on this ticket has now needed rescuing seven times, and 486 uncommitted lines were one crash from gone. Work toward the four blocking conditions the triple gate attached to #593: 1. state.rs -- make the SPEC 2.4 no-health-key test RECURSIVE. It asserted over JSON object keys (correctly, not substrings) but only at the top level, so a smuggled isRunning inside `counters` or any future nested struct would pass. 2. writes.rs, port.rs -- persist the entry-write bounds. SPEC 12.1 clause 2 requires cooldowns and fee budgets to survive a restart; they lived only in memory, so a restart loop wrote one bundle per restart with no interval, no daily cap and an empty cooldown map -- unbounded XCH spend plus repeated reserve settlements via re-eviction, presenting as nothing being wrong. The shape is a WriteBoundStore seam with a NoPersistence impl that REFUSES to write rather than writing unbounded. 3. admission.rs -- make a self-exclusion bypass a COMPILE error. EntryAction::Add carried public fields unrelated to AdmissionDecision, so any future discovery path could mint an Add without calling admit. dig-node#261's lesson is the rule: an invariant enforced on some paths is not an invariant, it is a habit. 4. mod.rs -- state the worst-case spend where a human reads it: 24 bundles/day, 192 entry actions/day, a fee ceiling of 24x the configured standard fee, up to 96 evictions/day, and eviction can flush the entry set's whole accrued balance in ~1.3 days including sub-threshold dust. Plus the non-blocking items: the csprng modulo bias documented, NoRepeatMemory pruned past its 8-cycle horizon, and the removal cooldown recorded after a successful submit rather than before (a failed submit had been holding an honest mirror out for six hours). Refs #3250
Exactly the five sites Rustfmt flagged in writes.rs (215, 222, 523, 558, 585). Salvaged from an uncommitted working tree after the third session cap on this ticket; formatting only, no logic touched. Refs #3250
…it from a fee PersistedEntryWriter::decide took a per-bundle standard fee and formed the daily cap in place with saturating_mul(24). A caller passing the day's whole ceiling therefore got a bound 24x looser than the operator configured, and the restart regression test's post-restart bundle (100 mojos already spent, plus a 1,000,000 fee, against a 1,000,000/day ceiling) was admitted instead of refused. The cap predicate and the persistence seam were both already correct: spent_mojos_today does round-trip the store, which is why the rate and cooldown assertions in that test passed. The defect was units. decide now takes daily_limit_mojos, and FeeBudget::daily_limit_for is the single place that product is formed, so the in-memory and persisted write paths cannot bound the same spend differently. Mistaking a ceiling for a fee there now fails CLOSED -- the prover refuses to write, which is what SPEC 6.3 clause 3 asks of it -- rather than open. The two 86_400 literals and the 24 become named constants derived from ENTRY_WRITE_MIN_INTERVAL_SECONDS so the ceiling cannot drift from the rate bound it comes from. Adds decide_then_commit_persists_every_write_bound_field: all four WriteBoundState fields must round-trip with distinct non-zero values. That is the general form of this bug class -- a seam that carries most fields and silently drops the one that bounds spending. The failing test's assertions, numbers and step order are unchanged (byte-identical). Refs #593
loop-security re-gate — #593 DELTAVerdict: PASS The units-confusion fix, verified
Five explicit answers(a) Any (b) Can any saturating operation loosen the cap? One theoretical instance, not live: (c) Any route to (d) Can a successful submit with a failed persist cause overspend? (e) Do Other three conditions, as new surface
Structural invariants reconfirmed unweakened at new head
What I did not coverDid not re-audit |
Adversarial gate, re-gate leg — RATIFY-WITH-CONDITIONSHead SHA read: C1 satisfied on the path it was written for, not on all paths. C2 satisfied. C3 satisfied in shape, wrong in one number. Two new blocking conditions, both small: C6 (a failed 1. C1 — persistence-or-refusal: fail-closed on load, fail-OPEN on save. Not fully satisfied.The load path is right and the restart hole is closed. The hole the brief asked about is real. Worst case, measured honestly: C6 — BLOCKS. Make a failed save structurally fatal, not a doc obligation. Cheapest shape: give 2. The 24× defect and the bounds —
|
| Quantity | Derivation | Value | mod.rs says |
|---|---|---|---|
| Bundles/day | 86,400 / 3,600 | 24 | 24 — correct |
| Entry actions/day | 24 × 8 | 192 | 192 — correct |
| Daily fee ceiling | daily_limit_for = fee × 24 |
24 × standard fee | correct, and correctly labelled NOT independent of the rate bound |
| XCH/day at fee 0.000005 | 24 × 5e-6 | 0.00012 XCH/day | correct |
| XCH/day at a congested fee 0.01 | 24 × 0.01 | 0.24 XCH/day; × 365 = 87.6 ≈ 88 XCH/yr | correct |
| Removals/day, pure-eviction worst case | all 192 actions are Remove |
192/day | stated as 96 — wrong |
| Removals/day sustained with re-add | 192 / 2 actions per churn | 96/day | conflated with the above |
| Entry set flush time | 250 / 192 | 1.30 days (2.60 days at 96/day) | 1.3 days — correct only at 192/day |
| Per-entry churn rate | 6 h cooldown + 3 × 900 s strikes ≈ 8.25 h | ~2.9/day | not stated; fine |
C7 — BLOCKS (one-line doc fix). mod.rs writes "up to 96 Remove actions/day (half of 192, if every bundle is all removals)" and then computes the 1.3-day flush from 192. The parenthetical is false: if every bundle is all removals the figure is 192/day, and 96/day is the evict-plus-re-add churn ceiling, at which the flush takes 2.6 days. As it stands the section understates the eviction rate by 2× while quoting a flush time that only holds at the un-understated rate. This is the funder-facing paragraph C3 exists to produce, so a wrong number in it is exactly the money lie the brief names. Correct to: 192 Remove actions/day worst case (96/day if each evicted entry is re-added, two actions per churn); a 250-entry set flushable in ~1.3 days at 192/day, ~2.6 days at 96/day. Every other figure in the block is right and I verified each against the constants above. C3 is otherwise satisfied — the fee ceiling, its linear scaling with the operator's configured fee, and the §6.4 sub-threshold-dust consequence are all stated plainly where a human reads them.
4. C2 — satisfied, and better than I specified.
AdmittedPeer has private fields, no public constructor, accessors only, and EntryAction::Add(AdmittedPeer) — so an Add is unconstructible outside admission, exactly as asked. The for_test mint is #[cfg(test)], which is the right escape hatch and cannot ship. Threading launcher_id into admit so the admission decision names the distributor it was decided for is more than I asked for and is the correct extra coordinate: without it an AdmittedPeer admitted for one distributor could be written into another. All five self-exclusion tests were rethreaded rather than weakened.
5. Still safe to merge — confirmed at this head, and the persistence seam did not acquire a backend.
The PR is still 11 files: ten under crates/dig-node-core/src/rewards/ and lib.rs containing exactly pub mod rewards; (line 63) and nothing else. The delta adds no spawn, no call site, no config key, no default-on path, and no real store implementation — PersistedEntryWriter is constructed only in #[cfg(test)], FakeStore is test-only, and NoPersistence (the one production impl) refuses both operations. No file outside rewards/ was touched in the delta. The concern the brief raises is the right one to raise about a persistence seam and the answer here is clean: the seam arrived without a backend, and the backend is #3265's, under #3265's gate.
6. The rest of the delta — one over-claim, otherwise it strengthens what it touches.
state.rsforbidden-health-keys: strictly stronger than what it replaced — recurses through nested objects and arrays, and the doc correctly explains why it must assert on object keys and not on substrings of the serialized string (ProverState::Runninglegitimately serializes the value"running"). AddingisRunning,uptime,alive,livewidens it. Does not encode the defect.NoRepeatMemoryprune +no_repeat_memory_does_not_grow_without_bound: binds.peer_idis peer-supplied, the test drives 2,000 distinct identities against a horizon of 8, and assertsrecent.len() <= CHALLENGE_NO_REPEAT_CYCLESon every iteration — it fails without both the inner and the outer prune. Legitimate, and it closes a memory-growth primitive I did not catch in round one.decide_alone_does_not_record_a_cooldown: binds — the assertion is false the instantdeciderecords, since the cooldown would be inserted atnow = 0.- One over-claim, non-blocking.
challenge.rs's new modulo-bias comment says the bias was "judged inert by both the security and decider gates on this ticket." My round-one verdict never considered modulo bias; that sentence asserts a consent that did not exist when it was written. Adjudicating it now so the claim becomes true: for aboundof a few GiB the bias factor is ~2^33/2^64 ≈ 2^-31, inert, and the rejection-sampling note marks the right place to tighten ifboundever approaches 2^64. Keep the comment; re-word the attribution to cite this comment rather than a round it predates. Not blocking, but it is the class of claim that ships false in the commit that writes it.
Conditions at this head
- C1 — partially satisfied. Load path and restart scenario: closed. Save-failure path: open. Superseded by C6.
- C2 — SATISFIED.
- C3 — SATISFIED except one number. Superseded by C7.
- C6 — BLOCKS. A failed
commitmust forcePersistenceUnavailableon subsequentdecidecalls structurally, with a load-ok/save-err test. ~10 lines, no callers to update. - C7 — BLOCKS. Correct the
mod.rseviction figures: 192 removals/day worst case, 96/day with re-add, ~1.3 days at 192 and ~2.6 days at 96. - Non-blocking, added to #3265 (no new ticket):
Mojos/DailyCeilingnewtypes at thedecideboundary before its first real caller; usePersistedEntryWriterexclusively and makeEntryWriteScheduler::decidenon-pub; re-word the modulo-bias attribution. - #3265 and #3266 confirmed filed — C4 and C5 discharged.
Clear C6 and C7 and this is a RATIFY with no further adversarial leg needed; neither touches a shape I would want to re-audit.
MichaelTaylor3d
left a comment
There was a problem hiding this comment.
Re-gate verdict: PASS
Head reviewed: 9e8e1c51d1697f324804e9e5fa3638ffe9999d9f. Delta reviewed: efe57f1154fb0ab5a518d710ff88263c8aa989b6..9e8e1c51 (admission.rs, port.rs, state.rs, writes.rs, mod.rs, challenge.rs) — not the whole 2,000-line module. D1–D5, settled at efe57f11, are unchanged and not relitigated.
The five items, explicitly
1. §2.4 honesty test now recurses — YES. state.rs: assert_no_forbidden_health_keys walks serde_json::Value depth-first over Object and Array, checks map.contains_key (keys, never substrings) at every depth, and the forbidden-key list grew (isRunning, isStale, uptime, alive, live, …). This is exactly what the prior blocking finding asked for; that thread is fixed.
2. Write-bound persistence — YES. writes.rs: WriteBoundStore trait (load/save), NoPersistence fails closed (Err on both calls — refuses to write rather than running bounds unbounded), PersistedEntryWriter::decide loads before deciding and only commit (called by the caller post-confirm) persists. No tenth ProverState: PersistedWriteOutcome::PersistenceUnavailable reuses the existing ChainSourceUnavailable reporting path rather than inventing a new state (§2.3's nine-state set is untouched). restart_still_enforces_rate_daily_cap_and_cooldown_across_the_store genuinely proves rate bound, daily cap AND cooldown survive a fresh PersistedEntryWriter over the same FakeStore. decide_then_commit_persists_every_write_bound_field additionally proves no field silently fails to round-trip.
3. AdmittedPeer self-exclusion as a compile error — YES. Private fields, no public constructor except a #[cfg(test)]-gated for_test escape hatch. AdmissionDecision::Admit(AdmittedPeer) and EntryAction::Add(AdmittedPeer) — a discovery path cannot build an Add without going through admit. The §5.3.4 control test (admits_non_self_candidate — an identical non-self candidate IS admitted) still passes, using AdmittedPeer::for_test, so it still distinguishes "excluded self" from "dropped everything." EntryAction::Remove was left as loose fields, untouched — correct, a removal is not an admission.
4. The money bound doc in mod.rs — numbers match code. MAX_ENTRY_WRITES_PER_BUNDLE = 8 × MAX_BUNDLES_PER_DAY (= 86_400 / 3_600 = 24, exact) = 192 entry actions/day, matching the doc. Fee ceiling stated as 24× standard fee — matches FeeBudget::daily_limit_for. 96 evictions/day (half of 192) and the ~1.3-day full-set eviction-drain figure (250 entries / 192 per day) check out arithmetically. The doc explicitly states the rate bound and fee ceiling are ONE spend control, not two — as required.
5. THE BUG FIX — reviewed hardest.
- Every caller now passes an already-derived ceiling:
PersistedEntryWriter::decidehas no other call site anywhere in the codebase outside this module's own tests (grepconfirms — the chain port this seam feeds isn't wired to a live caller yet;UnavailableChainPortis still the only adapter). Nothing reinstates the hole because nothing else calls it. MAX_BUNDLES_PER_DAY = SECONDS_PER_DAY / ENTRY_WRITE_MIN_INTERVAL_SECONDS = 86_400 / 3_600 = 24exactly (integer division, no remainder).grepfor24\b|86_400|86400inwrites.rs/mod.rsturns up only the two named constants and prose — nothing else hardcodes either literal.- In-memory (
FeeBudget::new→daily_limit_for) and persisted (PersistedEntryWriter::decidetakesdaily_limit_mojos) paths both consumeFeeBudget::daily_limit_for's output — one formula, not two. The defect (two paths, two units) is gone, not relocated. - The originally-failing test (
restart_still_enforces_rate_daily_cap_and_cooldown_across_the_store) is confirmed byte-identical between the pre-fix commit (556c01a) and the fix commit (9e8e1c51) —diffof the extracted function body is empty. Only the production code (parameter rename,daily_limit_forextraction, named constants) and a new, additional test changed. - Class coverage: the new
decide_then_commit_persists_every_write_bound_fieldtest proves the general form (everyWriteBoundStatefield round-trips with a distinct non-zero value, not just the fields the scenario test happens to inspect) — this catches a broader class than "this exact 24× instance," though it still wouldn't catch a third call site independently re-deriving the product with its own arithmetic mistake (there is no such call site today, so this is not live risk).
Non-blocking, carried from the prior round (both still open, neither newly regressed)
crates/dig-node-core/src/rewards/spec_constants.rs:27—CHALLENGE_WINDOWS_PER_CYCLE = 4still has no reader in this diff; not addressed this round, still non-blocking (pre-existing, orchestrator-ticket-shaped, not part of the five conditions asked for here).crates/dig-node-core/src/rewards/challenge.rs:50— modulo-bias: addressed with a documentation comment oncsprng_u64_belowexplaining the bias is bounded bybound/2^64and marking exactly where rejection sampling would go ifboundever grows close to2^64. Satisfies the ask ("consider a one-line comment"); leaving open per brief (not authorized to resolve threads this round).
Scope note (non-blocking)
challenge.rs's NoRepeatMemory::record now prunes stale (peer_id, launcher_id) entries and empty window lists on every call, with a new no_repeat_memory_does_not_grow_without_bound regression test. This wasn't one of the five items in the re-gate brief, but it's in-scope of the same commit (72d5b01, "checkpoint the four gate conditions in progress") that carries the four applied conditions, is additive, tested, and fixes a real unbounded-growth primitive keyed on peer-supplied peer_id. No objection.
What I did not run
Did not run the full dig-node-core test suite locally (long-running local build hit an MSBuild/cmake path-length failure in a scratch clone unrelated to this diff — libz-sys/vcpkg toolchain issue, not a code defect). Relying on CI (Analyze (rust), Test + coverage — both pending at dispatch per the brief, everything else green) plus targeted manual trace of every touched call site.
Verdict: PASS at 9e8e1c51d1697f324804e9e5fa3638ffe9999d9f.
…tion arithmetic The adversarial gate's last two blocking conditions on #593. C6 -- the save-failure hole. The load path was already fail-closed, but `commit` returned a Result and the obligation "treat the next cycle as unavailable if the save failed" lived only in a doc comment. A store whose load succeeds and whose save fails -- a read-only filesystem, a full disk -- kept handing back pre-submit state forever, so `spent_mojos_today` never accumulated and the daily ceiling silently became 24x whatever fee the caller supplied instead of 24x the standard fee. The rate bound survived only by the coincidence that PROVER_CYCLE_PERIOD_SECONDS equals ENTRY_WRITE_MIN_INTERVAL_SECONDS. A doc comment is not a bound, so `commit`'s Err now poisons the writer for its whole lifetime and `decide` checks that before anything else. There is deliberately no unpoison method: recovery is a fresh writer once the operator fixes the store, because a reset path is how a poison flag becomes decorative. C7 -- the eviction arithmetic was wrong in the direction that understates it. The doc claimed 96 Remove actions/day "(half of 192)" and then derived a 1.3-day flush from 192; both halves cannot be true at once. If every bundle is all removals the ceiling is 192/day; 96/day is the separate evict-plus-re-add churn ceiling, and at that rate flushing a 250-entry set takes ~2.6 days. For a figure a funder is meant to rely on, understating the eviction rate 2x while quoting a flush time only valid at the un-understated rate is a money misstatement, not a doc nit. Also re-worded the modulo-bias comment, which credited a judgement the decider gate says it did not make in that round. Committed from the parent side; the lane had all three edits correct but uncommitted, the eighth time in-flight work on this ticket needed rescuing. Refs #3250
Resolving the three open review threads — each is addressed by a landed commit
Refs #3250 |
MichaelTaylor3d
left a comment
There was a problem hiding this comment.
Verdict: PASS
Head: c74a16bbaf23ccb184e99263d6fc2b71f2cf4563
Reviewed the incremental delta 9e8e1c51..c74a16bb (3 files, +116/-12: writes.rs, mod.rs, challenge.rs) per the round-3 brief. Did not re-audit the ~2,000-line base already passed in rounds 1–2.
C6 — save-failure poison flag (writes.rs)
decidecheckspoisonedbeforeload: confirmed,writes.rs:328-330— the check is the first statement in the method body, beforeself.store.load(...)at line 332. Returns(PersistenceUnavailable, None). ✅- No unpoison/reset method: confirmed —
poisonedis only read indecide(.get()) and only written incommit(.set(true)); no method clears it. Recovery is a freshPersistedEntryWriter. ✅ Cell<bool>soundness: sound as written.decide/commitboth take&self, soCell(notAtomicBool) is the right choice for single-threaded/no-concurrent-access use.PersistedEntryWriter<'a>holding aCell<bool>makes it!Syncby auto-trait inference (theWriteBoundStore: Send + Syncbound on the trait object doesn't propagateSyncto the writer itself) — so the compiler already refuses to let two threads share a&PersistedEntryWriterconcurrently. That means the "future hazard" is self-defending: a later attempt to share this across threads is a compile error, not a silent race. Recommend a one-line comment on the field noting whyCellremains correct if someone ever wraps the writer inArc(they'd hit a compile error and should reach forAtomicBoolthen, not before) — non-blocking, doc-only.- Test spans ≥2 subsequent cycles: confirmed,
save_failure_poisons_the_writer_for_every_subsequent_cycle(writes.rs, tests module) asserts cycle 2 (PersistenceUnavailable/None) and cycle 3 at a distinct, laternow(2 * ENTRY_WRITE_MIN_INTERVAL_SECONDS + 2), each with a differentEntryAction. Two independent follow-up calls, not one. ✅
C7 — eviction arithmetic (mod.rs)
Checked against spec_constants.rs at this head: MAX_ENTRY_WRITES_PER_BUNDLE = 8, ENTRY_WRITE_MIN_INTERVAL_SECONDS = 3_600 (→ 24 bundles/day), MAX_ENTRIES_PER_DISTRIBUTOR = 250.
- 192
Removeactions/day = 24 bundles × 8 actions — matches the doc's own stated 192-action/day cap (no longer presented as "half of" anything). ✅ - 96/day churn ceiling = 192 actions / 2 (each churn = 1 Remove + 1 Add) — arithmetic checks out and is now correctly labeled a separate number from the 192 raw-removal ceiling. ✅
- ~2.6 days to flush 250 entries at 96 churns/day = 250/96 = 2.604 — checks out. ✅
No remaining internal contradiction between the two halves of the old doc (the "half of 192" self-contradiction is gone).
Attribution fix (challenge.rs)
Confirmed: the comment no longer claims the decider judged the modulo bias inert in the first round. It now states the decider "did not judge this in its first round on this ticket, adjudicating it only afterward as inert (bias ≈ 2⁻³¹)." Substance (bias is inert at realistic bound sizes, two-line rejection-sampling fix noted if it ever needs tightening) is retained. ✅
Wedged-loop assertion
Not touched by this delta (outside the 3 changed files) — re-confirmed byte-identical by omission; no further action needed.
Merge-readiness / inertness
git diff 9e8e1c51..c74a16bb --stat touches exactly writes.rs, mod.rs, challenge.rs — no new spawner, call site, config key, default-on path, or production WriteBoundStore impl anywhere in the delta or the rest of the tree at this head (repo-wide code search for impl WriteBoundStore for outside the test module in writes.rs returns nothing). This PR remains inert library code; the composed money-moving system stays deferred to #3265. ✅
Non-blocking note
writes.rs: theCellcomment could preempt the future-hazard question by naming the compile-time guarantee explicitly (see C6 above). Doc-only, not gating.
No security-critical custody/replay/fund-safety re-derivation was needed — the four gate conditions and the fee-ceiling fix were already verified applied in round 2 and are untouched by this delta.
🤖 Generated with Claude Code
loop-security — final gate, dig-node#593CHANGES-REQUIRED ScopeDelta Finding 1 — LIVE (in the design this PR ships, not yet reachable in production): C6's poison flag does not survive the object-lifetime pattern this module's own tests establish as correct usage
The poison flag is the one piece of state that pattern silently discards: construct a new writer Exploit / failure path: operator's persistence backend goes read-only or the disk fills. Verdict on the carried question: C6 as written is a caller-discipline dependency Explicit answers(a) Any path to a (b) Does a per-cycle-constructed writer defeat C6, and what must #3265 gate on? Yes, it (c) Is a poisoned writer observable to an operator? Not silently swallowed by design: the (d) Do (e) Third item — modulo-bias comment
What I did not coverDid not re-audit |
Adversarial gate, final leg — RATIFYHead SHA read: C6 and C7 are satisfied. The adversarial leg is CLOSED. No third round, no new blocking condition. Two requirements move to #3265, one of which I consider the most important thing on that ticket. Answers to all five questions below, each with its reason. 1. C6 — satisfied here, with the residual relocated to #3265 deliberately, not by omissionWhat was built is what I asked for and slightly more: The hole the brief names is real and I am relocating it on purpose. A flag on the instance is a guarantee about a lifetime, and Why that is acceptable at this head rather than something that must change here:
#3265 requirement (blocking that PR, not this merge), and it is the top one: the persisted write path MUST advance and durably save the write bounds before 2. C7 — satisfied; one figure I asked for was dropped, and it moves to #3265The correction is right and the reasoning is now explicit: 192 What was dropped: my correction asked for both flush times — "~1.3 days at 192/day, ~2.6 days at 96/day". The new text states only 2.6 days, tied to churn. The faster figure is the funder-relevant one (pure eviction, no re-adds, 250 / 192 = 1.30 days), and stating only the slower one understates worst-case flush speed by 2x — the same class of error as C7, inverted. Not blocking, and the reason is narrow: no operator and no caller reads this paragraph today, because nothing runs. #3265 must restore the pure-eviction flush figure (~1.3 days at 192 removals/day) alongside the churn figure before the prover can be enabled. If it were reachable by a funder today it would block. 3. Terminal answer — the adversarial leg is CLOSEDRATIFY at 4. Merge order — merging #593 first is correct; press it as hard as you like, it holdsNothing in this delta changes it and nothing in it makes first-merge riskier. The delta touches only 5. The "inert" argument — we have passed the point, and this PR is the last one that may use itChallenged as asked, and the challenge lands. "Inert" is a valid claim about this merge and I have verified it three times at three heads: no It is no longer a valid claim about the module. Across three rounds this argument has admitted a persistence seam, a money-ceiling arithmetic fix, a poison flag and a type-level self-exclusion guard — four changes whose entire purpose is to behave correctly under failure, none of which has ever executed against a real chain, a real store, or a real clock outside a fake. The tests are good and several of them caught real defects (my own C1 regression test found the 24x coupling bug). But every one of them runs against Decision: #593 is the last PR that may land in
If #3265 lands the wiring without (1) and (2), the triple gate on it should reject regardless of how clean the diff is — the missing evidence is the finding. Also in the delta
Verdict: RATIFY at |
Orchestrator adjudication: the gates split, and I am merging. Reason below.Final gate round on
Both dissenting gates found the same defect and disagree only about whether it blocks. Security The findingC6's poison flag lives on Why I am merging anyway — and it is not "it's only test code"Security's preferred remedy does not work. Moving the poison bit into The decider's remedy is the real one: reserve-then-spend. Persist the advanced bounds before And the hole is unreachable at this head, which I verified rather than took on faith: So the choice is not "fix it now or ship a hole". It is "write the wrong fix now against an absent What this costs, made explicit#3265 now carries reserve-then-spend as its top blocking requirement — persist advanced bounds The decider's structural point, which I am adopting"'Inert' has expired." Three gate rounds have all reasoned that merging is safe because nothing #593 is the last PR that may land in Merging on: reviewer PASS, adversarial RATIFY with the leg closed, CI fully green (zero failing, Refs #3250 |
…p, prover-status RPC (#602) * feat(mirror): persist mirror-bond coin ids (#575) * chore: open lane for #574 * feat(mirror): persist mirror-bond coin ids so a restart cannot double-create Bond identity was reconstructed from a live chain scan on every read (`mirror/observe.rs`), with no persistence of its own. A restart, a cold replica, or a lagging/flaky chain source all rendered a real, unspent, confirmed bond as "no bonds" -- and because the in-flight suppression is keyed on pending/submitted audit entries, a bond whose create had already CONFIRMED was not suppressed either, so the same short scan that emptied the read surface also cleared the one thing that would have stopped a second coin being paid for collateral that already exists (dig-node#574). Persist the (store, root, epoch) -> coin_id mapping in the EXISTING spend audit record (spend-audit.jsonl) rather than a new store: a mirror-coin create already writes store_id + AuditedBond{root, epoch} + amount there, and the coin id itself becomes durable the moment resolve_landed_spends confirms it. This adds the one missing piece -- the advertised URL a create carries -- and a read-side query, confirmed_mirror_bond, that returns the newest CONFIRMED record naming a triple. Chain stays authoritative. mirror::local_bond::recheck_missing_bonds never trusts the record: for a held bond the live scan did not cover, it asks the record for a candidate coin id, then re-verifies that SPECIFIC coin against chain via the same independent check (chain_bond_verdict) that verifies an untrusted peer's claimed bond. Only a fresh `Bonded` verdict is folded back in, as covered; `Unbonded`/`Unverified` fall through to an ordinary create, exactly as if no record existed. Version: 0.254.86 (patch -- per #522 the MSI ProductVersion minor field is exhausted and the counter lives in patch). Co-Authored-By: Claude <noreply@anthropic.com> * test(mirror): prove the recovery wiring end to end through PassRunner::run Adds two integration-level tests over the REAL pass pipeline, not just the isolated recheck_missing_bonds unit tests: a bond missing from the live scan with a chain-reverified durable record is recovered (no double create, correct Bonded state reported), and the control -- the same record but chain disproves it -- correctly falls through to an ordinary create. Together these are the concrete regression test for the cold-start/lagging-chain-source double-create scenario the ticket asked to have measured. Also refactors in_flight_creates to take the already-folded SpendLedger instead of re-reading the log itself, so PassRunner::run reads the audit file once per pass and shares it with the new recovery step, and fixes a doc comment on in_flight_creates that the recovery step would otherwise have made stale on landing ("a Confirmed create has a coin the chain observation already sees" is no longer unconditionally true). Co-Authored-By: Claude <noreply@anthropic.com> * chore(fmt): wrap long test signatures to satisfy rustfmt Co-Authored-By: Claude <noreply@anthropic.com> * chore(clippy): use slice::from_ref instead of cloning for a single-element slice Co-Authored-By: Claude <noreply@anthropic.com> * chore(release): bump to v0.254.89 Base branch moved to develop after PR #576 merged there at v0.254.88 (main and develop are currently identical), leaving this branch's carried-forward .88 as a zero-increment against the new base. Bumped to the next free integer after fetching and verifying both origin/main and origin/develop tip at .88. Co-Authored-By: Claude <noreply@anthropic.com> --------- Co-authored-by: Claude <noreply@anthropic.com> * fix(peer): count accepted relayed circuits in the connected pool (#579) serve_accepted_relay_conn served every accepted relayed circuit (full mTLS auth, full L7 peer RPC) while registering it nowhere, so connected_peers under-reported every relayed inbound peer -- the relay-leg twin of the direct-inbound defect #402/#523 already fixed. adopt_inbound_peer_in_pool now dispatches by TraversalKind: Relayed routes to dig-gossip's already-published adopt_relayed_inbound_handle (v0.32.0, the rev this repo already pins), every other tier keeps the unchanged adopt_direct_inbound_handle path. serve_accepted_relay_conn adopts before serving and releases after, mirroring the direct listener exactly. Refs: https://github.com/DIG-Network/dig_ecosystem/issues/3124 * fix(cli): guard the exit-code namespace shared with diga against collisions (#582) * chore: open lane for #3189 * fix(cli): guard the exit-code namespace shared with diga against collisions dign and diga deliberately share one process exit-code numbering (dig-app's outcome.rs says so in its own doc comment), so a number is free only if it is unoccupied ecosystem-wide. dig-node#407 assigned exit 7 to NODE_UNREACHABLE by checking only this repo's own table, where 7 genuinely was free -- and collided with diga's NOT_CONNECTED. A reviewer caught it by hand; nothing failed automatically. Adds scripts/check-exit-code-collisions.sh: parses both enums' code()/name() match arms straight from their own source -- this repo's ExitCode, and a live fetch of dig-app's outcome.rs at its default branch -- and fails if a number carries two different names, or if either side draws a number from the reserved shell signal range (126, 127, 128+N). Ships with an 18-case hermetic test harness (scripts/tests/check-exit-code-collisions.test.sh) covering the actual #407 collision shape, arm-order independence, arm-count mismatch, the reserved-range boundary from both sides, the live-fetch path itself, and fail-closed behaviour on an empty/missing/unreachable table. Wires a real (unstubbed) invocation into ci.yml's existing "Release-script tests" job so a collision introduced by a future PR, on either side, is a red required check on that PR -- not a note a reviewer has to catch. The fetch retries twice (2s backoff) since this becomes a required, network- dependent check; a fetch failure still fails closed after retrying, never silently passing as "diga has no codes". Updates SPEC.md 8.4 to point at the mechanical guard instead of leaving "re-check both tables" as unenforced prose, and records that the extension's WALLET_WS_ERR.NOT_CONNECTED = -33001 is a separate JSON-RPC error-code space, not a rival of this one. Adds a doc-comment to the existing transcribed collision test pointing future readers at the live script as the authoritative check; the transcription remains as a narrower, hermetic regression pin for the #407 shape specifically. No renumbering: every currently-assigned code is unchanged. Refs #3189 Co-Authored-By: Claude <noreply@anthropic.com> --------- Co-authored-by: Claude <noreply@anthropic.com> * fix(hygiene): port the lost-continuation guard to 4 crates, fix 48 corrupted strings (#3190) (#583) * chore: open lane for #3190 * fix(hygiene): port the lost-continuation guard to 4 crates, fix 48 corrupted strings Replicates dig-node-service::continuation_guard (dig-node#526/#501) into dig-node-core, dig-wallet, dig-runtime and dig-chat-protocol, line-for-line apart from crate-specific constants -- ported rather than reinvented, per dig_ecosystem#3190. Wiring the guard in surfaced 48 pre-existing lost-continuation defects the ticket's own "no measured corruption in these four crates" note did not anticipate: 36 in dig-node-core, 12 in dig-wallet, mostly test-assertion prose where a multi-line message lost its `\` continuation and shipped the source's own indentation as a mid-sentence space run (one as the worse `\n`-plus-indentation variant). All 48 are collapsed to the single space the sentence always meant, with surrounding indentation and wording otherwise untouched. Two lines are real column-alignment, not defects, and get a targeted EXCLUDED_LINE_RANGES entry on dig-node-core instead of a rewrite: download.rs's `claimed(...)` fixture-table trailing comments, and net.rs's `label : value` debug-print alignment. Refs https://github.com/DIG-Network/dig_ecosystem/issues/3190 Refs https://github.com/DIG-Network/dig_ecosystem/issues/3130 Co-Authored-By: Claude <noreply@anthropic.com> --------- Co-authored-by: Claude <noreply@anthropic.com> * feat(mirror): detect an IP change daily and reconcile mirror coins to the current advertise URL Automatic half (D1-D4) of the daily mirror-URL reconcile: derived personal-day offset, two-observation hysteresis, nine ordered gates with K sized as a self-funding prefix before any reclaim, `submitted` never `completed`, audit lines gain `reclaim_reason` + `trigger`. Gates at b4c09866: loop-reviewer PASS (review 5129272285), loop-security PASS @ 175304e1 (tree byte-identical, `git diff 175304e1 b4c09866` empty), adversarial loop-decider PASS (comment 5567083644; SHOULD-FIX findings ticketed separately). Refs DIG-Network/dig-node#570 Refs DIG-Network/dig_ecosystem#3203 * feat(serve): content hosting + serve path batch, v0.255.0 (dig_ecosystem#3212) Nine commits from the #3212 serve-path lane, gated at 899cc68f (reviewer review 5130425808, security comment 5568662836), plus the single semver bump to 0.255.0 for the develop -> main batch. - store_id/root case normalised at the CapsuleKey boundary; cache delete targets the matched entry - tier-0 occupancy reads the eviction-aware ledger - profile-sync outbound budget in bytes; announcer asked first - melt confirmation depth on the terminal spend, fail-closed - EngineWarming (-32002) while the peer tier attaches, never -32004 - window completeness derived from the bytes read - deps: dig-stun 0.2, chia-query 0.24.3, dig-nat 0.21.2, dig-logging 0.2.2 Refs DIG-Network/dig_ecosystem#3212 * chore: untrack gitnexus-generated agent files (#590) * chore: untrack gitnexus-generated agent files These files were generated by `gitnexus analyze` as a side effect of indexing this repository. They are development-loop private tooling output, not product code, and carry no secrets. They are removed from tracking going forward via .gitignore; history is deliberately NOT rewritten. Refs #3177 * chore: drop private-repo reference from gitignore comment The ignore comment named a private repository and an internal issue number in a public file, which is the same disclosure class this change set exists to remove; the reference is dropped and the guidance kept. * feat(rewards): always-on prover loop engine -- honest liveness, type-enforced self-exclusion, bounded spend (#593) The always-on reward-prover engine: ~2,000 lines under `crates/dig-node-core/src/rewards/`, built against the merged `dig-rewards-coin` SPEC. Library only -- nothing spawns it, and the sole production `RewardsChainPort` refuses every call, so it cannot spend. Wiring the composed system is #3265, which carries its own gate. The epic's premise -- "anytime the process isn't running, rewards are not being distributed" -- is half wrong, and the false half is the dangerous one. `Sync`, `NewEpoch` and `InitiatePayout` need no manager authority, so funder downtime does not stop rewards: it FREEZES THE ENTRY SET while accrual and payouts continue. Peers that stopped mirroring keep earning; peers that started cannot begin. That shaped the whole design. Liveness honesty (SPEC 2.4). The status record carries no `healthy`/`ok`/`up` boolean and no precomputed staleness, because a wedged loop cannot report its own wedging -- whatever it last wrote stays there, so a writer-set flag reads true forever after the failure it exists to reveal. The reader derives staleness from `last_cycle_completed_at` against `observed_at` and its own clock. A recursive JSON-key test enforces the absence at every nesting depth; asserting on keys and never substrings, since `ProverState::Running` legitimately serializes the VALUE "running". The one legal staleness signal is chain-derived (SPEC 12.4: 48 hours AND a non-zero reserve, from the singleton's own spend history) and lives on the distributor read, where a wedged prover cannot fake it. Self-exclusion is a compile error, not a habit. dig-node#261's lesson is that an invariant enforced on some paths is not an invariant. `admit` is the single admission point, checks both SPEC 5.2 coordinates (own peer_id OR a payout puzzle hash this wallet controls), and mints an `AdmittedPeer` with private fields and no public constructor -- so `EntryAction::Add` cannot be built by a path that skipped admission. A prover's own fault can never strike a peer. `GateError` is a distinct type from `GateIneligibleReason` and `record_prover_fault` takes `&self`, so SPEC 3.6 clause 4 is enforced by the borrow checker rather than by comment. Without that, a misconfigured operator -- one missing mirror-collateral epoch ordinal -- would strike every peer at once and evict its entire 250-entry set in three hours, each eviction a fee it pays plus a settlement out of its own reserve. The money bounds are stated where a human reads them (`rewards/mod.rs`): 24 bundles/day, 192 entry actions/day, a fee ceiling of 24x the configured standard fee, 192 removals/day worst case with 96/day sustained churn. Recorded honestly: SPEC 6.3's rate bound and fee ceiling are ONE control, not two. Three gate rounds, every leg fresh-context. Round 3 at this head: reviewer PASS, adversarial decider RATIFY (leg closed), security CHANGES-REQUIRED on a finding the decider ratified deliberately -- adjudicated in https://github.com/DIG-Network/dig-node/pull/593#issuecomment-5601784815 and carried to #3265 with the remedy corrected, because the proposed fix would have persisted a poison flag to the very store whose writes were failing. Found and fixed under gate: a census ordinal off by one in both directions (SPEC 4.6 requires n-1 exactly); an unreachable grace window leaving a named constant with no reader; a missing `NewEpoch` spend; an absent-ordinal path attributing a prover fault to peers; and a daily fee ceiling 24x too high because a per-bundle fee was consumed as a daily ceiling. Refs DIG-Network/dig_ecosystem#3250 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: serve dig.getRewardProverStatus at Tier::Control (#595) * chore: open lane for #3269 Bump dig-rpc-protocol to 0.11.0 (adds dig.getRewardProverStatus and the other reward RPC methods to the wire). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(rewards): fail-closed Reward-tier guard + dig-rpc-protocol 0.11 line assertion - dependency_tree.rs: assert the resolved dig-rpc-protocol line is 0.11 (was 0.10); documents the known-red two-version state pending the dig-peer 0.14.0 / dig-download 0.23.0 cascade (#3269). - reward_methods_tier_guard.rs: fail-closed guard over the live Method::ALL catalogue -- every Reward-named method must be Tier::Control and not peer-reachable, so a fifth reward method added later is caught at the wrong tier automatically rather than inheriting a wrong default (binds #3261's rule node-side). - peer.rs: sibling unit test exercising the real (pub(crate)) is_peer_reachable_method, since an external integration test cannot see it -- same guard, executed against this node's own allowlist rather than only the shared crate's. Refs #3269 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * style: remove trailing blank line in reward_methods_tier_guard.rs * feat(rpc): serve dig.getRewardProverStatus at Tier::Control Adds the missing handler for PR#595: a new reward_prover_statuses registry + accessors on Node (empty until #3265 spawns a prover loop, so the registry read is real, not a stub), a dispatch.rs arm inside the Method enum match (never the string pre-match), and a field-for-field mapping from dig-node-core's internal rewards::state::RewardProverStatus (camelCase-tagged) onto dig-rpc-protocol 0.11's wire RewardProverStatus (snake_case-tagged struct, camelCase-tagged ProverState value), widening entry_count u32 -> u64 explicitly and hex-encoding the three [u8; 32] identity fields. An all-zero launcher_id (what an uninitialised registry slot hex-encodes to) is omitted at this boundary rather than rendered as a real distributor with a plausible-looking id -- the money-hole class the dig-rewards-coin driver's adversarial gates found three times. Tests (in dig-node-core::lib.rs's existing test module, where the pub(crate) registry accessors are visible) drive the real dispatch entry point (handle_rpc -> RpcDispatch::dispatch -> the Method arm) and assert field-for-field on the serialized JSON body: populated registry, empty registry (-> {"statuses": []}), zero-id omission, tier/peer-reachability, enum-match-not-string-prematch, and launcher_id filtering. The no-health-boolean / no-staleness assertion is by key set, not substring. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * style: rustfmt the reward-prover-status registry + tests Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(deps): bump dig-download 0.23, dig-peer 0.14, dig-rpc-protocol 0.11 Closes the two-versions-of-dig-rpc-protocol split (#3269): dig-download 0.23.0 and dig-peer 0.14.0 both now resolve dig-rpc-protocol ^0.11, matching dig-node-core's own dig-rpc-protocol = "0.11.0" line, and dig-node-service's two 0.10 lines (main dep + dev-dependency restatement for openrpc_drift_guard.rs) move to 0.11 to match. Manifests only. cargo update / Cargo.lock intentionally NOT run yet: a fourth capper, dig-peer-selector ("0.11" in dig-node-core/Cargo.toml), still requires dig-peer = "^0.13" in every published version through 0.11.1, so the tree cannot fully resolve until a dig-peer-selector release picks up dig-peer 0.14. CI will stay red on this commit for that reason, which is expected. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * feat(rpc): close dig-rpc-protocol 0.11 cascade + attribute payout figures Bump dig-peer-selector 0.11 -> 0.12 (the release that moves onto dig-peer ^0.14) and run cargo update, closing the two-versions-of-dig-rpc-protocol split: Cargo.lock now resolves exactly one dig-rpc-protocol, at 0.11.0, alongside dig-peer 0.14.0, dig-download 0.23.0, dig-peer-selector 0.12.0. Add a subject-attribution test and doc comments to reward_prover_status_to_wire: total_paid_out_base_units and reserve_base_units are per-distributor totals (this distributor's payout to ALL its mirrors, and this distributor's own reserve), never the querying node's own earnings and never summed/cross-attributed across distributors. This is the defect class a sibling adversarial gate found in dig-app#403's rewards pane, which rendered a distributor total as one mirror operator's personal earnings and overstated by up to 250x. Checked: rewards/state.rs, rewards/port.rs and rewards/mod.rs contain no Eligible/verdict/payout_hash symbol, so nothing in this mapping surfaces a persisted EligiblePayoutHash verdict. Refs #3269 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(rpc): silence dead_code on register_reward_prover_status pending #3265 Clippy's non-test lib target has no production caller for register_reward_prover_status yet, because #3265 (the always-on prover loop that would call it from bring-up) has not landed -- only tests call it today. cfg_attr(not(test), allow(dead_code)) stands in for that missing caller until #3265 wires a real one. Refs #3269 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(rpc): make the all-zero identity guard non-silent and cover all three fields Security (blocking) and the adversarial leg both found the same defect in the zero-launcher_id filter: it checked only launcher_id, so a registration bug that zeroed store_id or root beside a valid launcher_id would pass through as a plausible record, and dropping the bad record silently destroyed the evidence a registration bug happened at all -- SPEC Sec2.4 clause 1's exact prohibition. zeroed_identity_fields() now checks launcher_id, store_id AND root. The dispatch filter still excludes a record with any zeroed field (never renders an uninitialised slot as a real distributor), but first fires a tracing::warn! naming which field(s) were zero, so a bad registration is observable rather than swallowed. Kept isolated in dispatch.rs rather than woven into the wire mapping, since this belongs at #3265's writer once that lands. Replaced get_reward_prover_status_omits_an_all_zero_launcher_id (which proved the omission but not the observability, and never exercised a zeroed store_id/root beside a valid launcher_id) with get_reward_prover_status_logs_and_excludes_a_zeroed_identity_field, covering both a zeroed launcher_id and a zeroed store_id beside a valid launcher_id, and asserting the tracing::warn! output via the crate's existing capture_sync_logs test utility. Fixed a now-false "Known-red" doc comment on tests/dependency_tree.rs::the_workspace_carries_exactly_one_module_wire_crate: the dig-peer 0.14.0 / dig-download 0.23.0 / dig-peer-selector 0.12.0 cascade already landed in this PR's Cargo.toml/Cargo.lock, so the assertion is green, not red. Assertion itself untouched -- still exact-version. Refs #3269 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(rpc): correct a born-false "shipped dig-app 15.5.0" doc claim dig-app's latest release is v15.4.0 -- there is no v15.5.0 tag -- and dig-app#403 (the pane that would consume dig.getRewardProverStatus) is OPEN and unmerged. Point the doc comment at the real, unmerged consumer instead so a future reader doesn't take this as evidence a shipped consumer depends on the guard, which would wrongly discourage relocating it to #3265's writer. Refs #3269 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(rpc): correct doc placement, assert the zeroed field by name, treat root as an observation Three findings from the correctness gate on PR#595 at 134864a9. 1. The zeroed-identity helper's doc block was spliced onto the end of reward_prover_status_to_wire's block with no separator, so the wire-mapping rationale documented a boolean predicate and the mapping function was left with no doc at all. Each doc block now sits above the item it describes. 2. The log assertion `logs.contains("launcher_id")` was a tautology: the warn emits launcher_id as a structured field on every fire, so the property the guard exists to add -- naming which field was zeroed -- was unasserted. Deleting `zeroed_fields = ?zeroed` from the warn left every assertion green. The test now asserts the zeroed_fields value itself, which the fixture makes exact and disjoint across cases. 3. `root` is an observation, not an identity. A registered prover that has not completed its first cycle plausibly has no root, and a writer that zero-inits it would have made a healthy prover invisible. A zeroed launcher_id or store_id still excludes the record; a zeroed root alone warns and returns. Refs #3269 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(rpc): restore zeroed_fields structured field dropped from the pushed warn The previous commit (080be3df) landed with `zeroed_fields = ?zeroed` missing from the tracing::warn! call in the GetRewardProverStatus filter -- a one-line regression introduced while proving the new log assertion goes red without it, never restored before the commit was made. Without this field the log line never names WHICH field was zero, so an operator sees only that something was excluded, and the test asserting `zeroed_fields=[...]` per case would fail. Restored; all 7 reward-prover-status tests green. Refs #3269 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(rpc): split zeroed-field logging by level -- WARN for a missing identity, DEBUG for a zeroed root A zeroed launcher_id or store_id is a real registration bug: the record is excluded and now logs at WARN, naming the exact field(s) via `zeroed_fields=[...]`. A zeroed root alone is an ordinary pre-first-cycle state, not a fault: the record is still returned, and now logs at DEBUG instead of WARN, so an operator polling this endpoint sees warn-level volume proportional to real registration bugs, not to every not-yet-cycled prover on every poll. Updated the doc comments on `zeroed_fields`, the dispatch filter and the test to describe the level split, and extended the regression test to assert on level (WARN vs DEBUG) as well as the `zeroed_fields` value. Proved both directions: flipping the DEBUG branch back to WARN turns the test red on the level assertion; flipping the field-name assertion back to a bare `contains("launcher_id")` would have passed unconditionally (the prior tautology) and is no longer possible since the assertions now pin `zeroed_fields=[...]` plus the level string. Refs #3269 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> * feat(rewards): peer-side claim loop -- watch distributors, claim on cadence (#3251) (#594) * feat(rewards): peer claim loop skeleton -- discovery, cadence, config, claim port * test(rewards): write all twelve acceptance tests for the peer claim loop * feat(rewards): wire the seven rewards_claim submodules into the crate mod.rs declared no submodules, so types.rs/port.rs/config.rs/cadence.rs/ parser.rs/engine.rs/hints.rs (1,435 lines, 25 tests) were never part of the crate and never compiled. Declare them and re-export the public surface. * style(rewards): cargo fmt the rewards_claim submodules * chore(deps): bump dig-node-control-interface 0.33->0.35, dig-rpc-protocol 0.10->0.11.0 dig-rpc-protocol 0.11.0 is merged and tagged upstream; the other dig-*/chia-* deps of dig-node-service were already at the latest permitted-by-caret version in Cargo.lock. crates/dig-node-core/Cargo.toml is untouched (#3250's file set). * chore(deps): revert dig-node-control-interface and dig-rpc-protocol bumps Both create a duplicate-version split in this PR's scope and neither can be closed without editing a sibling crate's manifest this lane does not own: - dig-rpc-protocol 0.11.0 duplicates against dig-node-core/Cargo.toml:194 ("0.10.2"), which is #3250's live file set (dig-node#593). - dig-node-control-interface 0.35.0 duplicates against dig-wallet/Cargo.toml:81 ("0.33"), a sibling crate this lane does not own; the observed Clippy break (BalanceAsset/Asset type-identity mismatch, missing url_reconcile/url_current/urls fields) came from THIS duplicate, not from dig-rpc-protocol. Both belong to their own sequenced dep-bump unit of work, not this ticket. * fix(rewards-claim): fault laundering, permanent no-entry blacklist, fee ceiling magnitude Three independent gates on dig-node#594 (51516e62) found four logic defects; this addresses A, B and C per the corrected fix brief (D is documented only, not fixed here per the brief's own instruction). Defect A -- the anti-silence surface laundered every real fault into `Nominal`: - A1: `fault_reported` had no fault-bearing ClaimLoopState to fall through to, so a chain adapter erroring every cycle read `Nominal` forever. Added `ClaimLoopState::Faulted { cycles }`, outranking Nominal/ClaimableButNotClaiming, under ChainSourceUnavailable. - A2: inverted the test that asserted A1's bug as correct behaviour. - A3: `ClaimableButNotClaiming` compared a per-cycle snapshot (`distributors_claimable`) against a lifetime-cumulative counter (`claims_submitted`), so it latched healthy forever after one lifetime success. Added `claims_submitted_this_cycle` (per-cycle) as the correct comparand; kept `claims_submitted` as a cumulative counter. - A4: `last_discovery_at`/`last_cycle_at` were stamped even on a failed discovery or an all-faulted cycle, destroying the staleness signal a reader depends on. Now only stamped on success; added `last_attempt_at` to prove liveness separately. `fault_reported` and `distributors_faulted` now reset per cycle instead of latching for the process's lifetime. Defect B -- "terminal, stop retrying" was implemented as a process-lifetime blacklist (`terminal_no_entry: HashSet<Bytes32>`, never cleared). That blocked SPEC 12.5 clause 2's re-entry path (evicted, re-challenged, re-admitted never claims again) and permanently punished a peer that discovered a distributor before the funder's AddEntry landed. Removed the blacklist entirely -- `own_entry` is a cheap chain read, re-issued every cycle for every candidate, matching clause 3's "never cache across cycles". `NoEntrySlot` is now a per-cycle observation, not a lifetime sentence. Defect C -- the fee ceiling didn't bind anything and there was no aggregate cap: - C1: default `CLAIM_FEE_CEILING_MOJOS_DEFAULT` lowered from 1_000_000_000 (transplanted from `MIRROR_SPEND_FEE_CEILING_MOJOS`, sized for a mirror-coin spend) to 200_000 -- 2x the observed routine Chia fee range (5,000-100,000 mojos), so it actually binds instead of leaving 4-5 orders of magnitude of slack. - C2: added a per-cycle aggregate fee budget (`max_cycle_fee_budget_mojos`, default 10x the per-claim ceiling) checked across all claims in a cycle, closing the attacker-cost gap where funding K distributors could force a victim to spend K x the per-claim ceiling per cycle. New `ClaimOutcome::SkippedCycleBudgetExhausted`. Tests: rewards_claim test count 26 -> 37 (11 new: repeated_discovery_faults_never_ read_as_nominal, failed_discovery_leaves_last_discovery_at_unchanged, a_reported_ fault_surfaces_as_faulted_not_nominal, a_lifetime_submission_does_not_mask_a_ later_cycle_that_submits_nothing, no_entry_slot_then_re_admitted_produces_a_claim_ on_the_later_cycle, distributors_each_under_ceiling_do_not_collectively_exceed_ the_cycle_budget, the_default_per_claim_ceiling_actually_binds_a_routine_fee, plus renamed/rewritten no_entry_slot_is_non_terminal_and_re_checked_every_cycle). Refs #3251 * fix(rewards-claim): refuse a claim entry for the wrong payout puzzle hash CI fix: cadence.rs's RewardsClaimConfig literal was missing the max_cycle_fee_budget_mojos field added in the previous commit (E0063, caught by CI's Clippy/Test jobs -- the local cargo check for this workspace is too slow to use as the compiler here). Defect E (security-gate finding, folded in before this pass closes): submit_initiate_payout was called with entry.payout_puzzle_hash -- whatever the chain port handed back -- with no check against this node's own own_payout_puzzle_hash. UnavailableClaimChainPort is the only production adapter today so nothing can exploit this yet, but the whole point of the ClaimChainPort seam is that #3249 swaps in a real adapter with nothing above it changing, so deferring this would ship the landmine live with no review pass watching for it. Added an equality guard before the spend: a mismatch refuses to submit, counts (ClaimStatus::claims_refused_payout_mismatch), surfaces its own named outcome (ClaimOutcome::PayoutPuzzleHashMismatch), and is reported as a fault (a divergent entry means the port is confused or hostile, not that there is nothing to claim) -- never corrected by substituting our own hash and proceeding. Defect D: documented, not wired, per instruction -- added the "not yet wired into node startup" paragraph to mod.rs's module doc (the PR body carries the same paragraph) so the next reader arrives at the caveat in the code, not only in a merged PR description. Refs #3251 * fix(rewards-claim): B1 -- ClaimableButNotClaiming is a magnitude comparison, not a zero-test submitted_this_cycle < claimable_this_cycle now fires the anti-silence state, carrying the shortfall as ClaimableButNotClaiming { claimable, submitted }. The previous submitted_this_cycle == 0 zero-test let one submission mask any number of same-cycle skips (claimable=10, submitted=1 read Nominal). Also folds in B3's precedence fix (ChainSourceUnavailable > Faulted > ClaimableButNotClaiming > Idle > Nominal) and the per-distributor payout_hash_mismatches_this_cycle counter so a per-distributor fault can no longer pin the cycle-wide Faulted state, plus R2's rename of terminal_no_entry_slot to no_entry_slot_this_cycle now that it is no longer terminal. * fix(rewards-claim): B2/B3 -- value-ordered budget with rotation, per-distributor fault isolation B2: run_cycle now splits into a pre-budget phase (asset/entry/hash/threshold checks, producing the claimable set) and a budget phase, ordering the claimable set by accrued value descending before applying the fee ceiling and cycle budget. Dust distributors (low accrued value regardless of attacker-controlled fee) now sort last and are the ones the budget drops, closing the claim-suppression attack where ten high-fee dust distributors could consume the whole cycle budget ahead of a victim's real earnings. A rotation_cursor tie-breaks only WITHIN equal-accrued-value tiers so a genuinely tied honest tail that exceeds one cycle's budget every cycle still rotates through and is eventually served, rather than dropping the same tail forever. B3: the payout-hash mismatch check in evaluate_pre_budget now increments the per-distributor payout_hash_mismatches_this_cycle counter instead of setting fault_reported, so one hostile or buggy entry can no longer pin the cycle-wide Faulted state and bury ClaimableButNotClaiming for every other healthy distributor. R2: terminal_no_entry_slot -> no_entry_slot_this_cycle throughout. * fix(rewards-claim): R5 -- document not-yet-wired loop; persist B2 rotation cursor An operator reading their own rewards-claim.json and seeing enabled: true has no way to know from that file alone that no startup path constructs a ClaimEngine yet (#3268) -- mod.rs said so, but a config file reader does not arrive at a module doc. Also gives RewardsClaimConfig a rotation_cursor: Option<Bytes32> field so B2's tie-break cursor survives a save/load round-trip -- an in-memory-only cursor resets on every restart, which would starve a legitimately tied honest tail forever on any node that restarts daily. * fix(rewards-claim): clippy collapsible-match + SPEC v0.1.3 wording refresh Collapses the nested if into the outer match arm in run_cycle (clippy::collapsible_match). Also refreshes NoEntrySlot / no_entry_slot_this_cycle doc comments now that dig-rewards-coin v0.1.3's SPEC §12.5 amendment is merged and tagged: absence is terminal for one claim attempt only, never for the distributor, must not be cached, and must not accumulate into a permanent exclusion set -- confirming rather than diverging from the re-read-every-cycle behaviour already implemented. * fix(rewards-claim): add missing rotation_cursor field in cadence.rs test literal Struct literal in the cadence test module was not updated when RewardsClaimConfig gained rotation_cursor (R5 commit) -- CI's Clippy/Test jobs caught the missing field (E0063) that a local cargo check could not (killed by memory pressure before this workspace-wide build completed). * fix(rewards-claim): F1 -- ChainSourceUnavailable is per-cycle, never a latch compute_state() compared against self.state -- last cycle's OWN computed output -- so once any cycle took an Unavailable port path, every later cycle re-asserted ChainSourceUnavailable forever, even after the chain came back and real claims were submitting. A node still syncing, or one dropped connection, was enough to trip this permanently. Add ClaimStatus::chain_unavailable_this_cycle, reset to false at the top of every run_cycle and set true only on a cycle that actually took the Unavailable path; compute_state now reads that flag instead of self.state, so the reading is live again. Test: a_transient_unavailable_cycle_does_not_latch_state_for_the_rest_of_the_process (engine.rs) drives cycle 1 unavailable, cycle 2 healthy with a submission, and asserts cycle 2 reads Nominal. Plus a compute_state-level regression in types.rs. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(rewards-claim): F3/F4/F5 -- fix stale counters, dedup candidates, correct version doc F3: reset EVERY per-cycle counter (distributors_known/with_own_entry/ claimable/faulted, claims_submitted_this_cycle, no_entry_slot_this_cycle) at the TOP of run_cycle, before any early return. The three ChainUnavailable early-return paths skip the end-of-function assignment block entirely, so a cycle that hit one used to leave the PRIOR cycle's counts sitting on self.status while last_attempt_at stamped fresh for THIS cycle -- a stale count under a fresh timestamp, exactly what SPEC §2.4's staleness reasoning forbids. types.rs's doc sentence for no_entry_slot_this_cycle now correctly says it is dated by last_attempt_at (the field stamped unconditionally every cycle), not last_cycle_at. F4: dedup `candidates` by launcher id before phase 2. A real adapter scanning §1.3 launch comments across every (store_id, root) this node mirrors can plausibly return the same launcher id twice; without dedup phase 2 would evaluate it twice and submit InitiatePayout twice against one entry slot in one cycle -- the second spend is invalid but the fee is paid anyway. F5: dig-rewards-coin is v0.1.3, published on crates.io -- correct the stale "v0.1.1" module-doc claim. Tests: a_chain_unavailable_cycle_does_not_leave_prior_cycles_counters_stale (F3), a_duplicated_launcher_id_submits_exactly_once (F4). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(rewards_claim): fold payout-hash mismatches into the shortfall predicate (F2) A payout-hash mismatch never enters the eligible set, so it was counted in NEITHER distributors_claimable NOR claims_submitted_this_cycle -- the shortfall lived in neither term of compute_state's magnitude comparison. All-K-distributors mismatching therefore read Nominal (falsely healthy). Fold payout_hash_mismatches_this_cycle into the comparison's denominator: submitted < claimable + mismatches. The result is ClaimableButNotClaiming (a shortfall), never the cycle-wide Faulted -- Defect B3 stays fixed. Inverts the assertion at what was engine.rs:1305 (a_payout_mismatch_never_sets_the_cycle_wide_fault_or_masks_other_distributors): it previously asserted ClaimLoopState::Nominal across three cycles of an ongoing mismatch, which pinned the defect as intended behaviour (an A2-class test). It now asserts ClaimableButNotClaiming { claimable: 1, submitted: 1 }. Adds all_distributors_mismatching_is_a_shortfall_not_nominal, covering the brief's exact "what if every distributor refuses for the same reason" case. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(rewards-claim): F1/F3 -- AtomicU32 in fake ports, not Mutex, keeps discovery Send CI's Clippy job (the compiler for this crate, per brief) caught it: holding a std::sync::MutexGuard across the .await in FlakyThenHealthyPort and HealthyThenUnavailablePort's discover_distributors made the returned future not Send, which #[async_trait]'s generated trait signature requires. Neither fake needs a lock -- each holds one call counter, incremented once per call, never read-modify-written across an await point. AtomicU32's fetch_add removes the guard (and the Send bound violation) entirely. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(rewards-claim): F7 -- persist the fee window and cadence gate across restart The per-cycle aggregate fee budget and the 24h cadence clock both lived only in memory: `spent_this_cycle_mojos` was a `run_cycle` local and nothing on disk recorded a completed cycle. Every fresh process got a full `max_cycle_fee_budget_mojos` and an empty cadence clock, so a node stuck in a crash-restart loop could spend unbounded XCH on fees, one full budget per restart. Adds three `#[serde(default)]` fields to `RewardsClaimConfig` (`fee_window_start_unix`, `fee_spent_in_window_mojos`, `last_cycle_completed_at`) and a new opt-in `ClaimEngine::with_persisted_fee_ window(dir, cadence_seconds)` that: - restores the window/cadence state from `dir` at construction, - refuses to start a cycle until the cadence has elapsed since the last completed one, - rolls a fresh budget window only once the cadence has elapsed since it opened, otherwise keeps enforcing the budget against the persisted spend, - persists the spend BEFORE every chain submission (write-then-spend), never batched to cycle end, and persists the completed-cycle timestamp when a cycle finishes. Engines that never call `with_persisted_fee_window` (every pre-F7 test) are unaffected -- this is additive, opt-in state beside the existing rotation cursor, not a change to B2's value-ordering or rotation mechanism. `ClaimStatus`'s own counters stay in-memory on purpose (observability, meant to reset on restart); only the spend bound and the cadence gate persist. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(rewards-claim): F7 -- update cadence.rs test literal for new persisted fields The three new persisted RewardsClaimConfig fields (fee_window_start_unix, fee_spent_in_window_mojos, last_cycle_completed_at) broke this crate's only remaining full struct literal outside config.rs/engine.rs's own test modules -- E0063 missing fields, caught by CI's Clippy job. Switched to ..RewardsClaimConfig::default() so the next added field cannot break this literal again, the same fix already applied once before for rotation_cursor. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(rewards-claim): F8/F14 -- atomic state write, fail closed on a corrupt window Salvaged from a lane killed by a weekly cap before it could commit. Uncompiled at commit time; CI is the compile signal. Covers the fourth gate pass findings on the F7 persisted spend bound: - F8: RewardsClaimConfig::save_to now writes atomically (temp file + rename in the same directory), reusing the pattern already used by mirror/reconcile_state.rs for the same class of state. load_from distinguishes an ABSENT file (clean first run, defaults are correct) from a PRESENT but unparsable one, which fails CLOSED: the window is treated as fully spent and nothing is submitted. Never Default, and never a silent clamp downward, which would hand back the budget the corruption was hiding. - F14: the budget comparison uses saturating arithmetic so a corrupt disk-seeded fee_spent_in_window_mojos cannot panic under the release profile's overflow-checks. - F9/F10/F12/F13 in progress in the same files. Refs #3251 * fix(rewards-claim): negate with ! rather than the unimported Not trait Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com> * fix(rewards-claim): F13 -- correct stale test literals to the folded shortfall compute_state (types.rs) already reported the folded shortfall denominator (distributors_claimable + payout_hash_mismatches_this_cycle) as `claimable` -- that part of F13 landed in f478516a. The two engine.rs tests asserting this state were written against the pre-fold, un-folded numbers and never updated, so CI showed the implementation producing the correct folded value (`claimable: 2`, `claimable: 1`) while the test literals still expected the stale un-folded one (`claimable: 1`, `claimable: 0`). Update both literals -- and the comments describing them -- to the folded values the F13 fix actually produces. No production code change; compute_state's predicate and payload were already correct. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * feat(rewards-claim): add ClaimOutcome::Faulted variant Add the seventh ClaimOutcome variant: the type could only say a peer was legitimately not paid, never that a chain call failed. Carries the launcher id, a bounded (200 char) copy of the chain port's error text, and whether a pre-committed fee was reversed, so a reader can tell no money moved. Engine wiring at the two fault arms (engine.rs:332, :377) follows in the next commit. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(rewards-claim): both fault arms now push ClaimOutcome::Faulted engine.rs:332 and :377 used to increment `faulted` and discard the outcome, leaving a definitively-failed claim absent from the outcome stream -- indistinguishable from a cycle that never touched that distributor. Both PreBudgetResult::Fault and BudgetPhaseResult::Fault now carry the chain port's (bounded) error text, and the submit_initiate_payout failure path also carries the fee it reversed, so a reader can tell no money moved. The counter stays; it is not a substitute for the outcome. 7 call sites needed updating: 3 PreBudgetResult::Fault constructions (reserve_asset_id, own_entry, payout_threshold), 2 BudgetPhaseResult::Fault constructions (required_fee_mojos, submit_initiate_payout), and the 2 consuming match arms -- exactly the set that was silently discarding a failure before this change. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(rewards-claim): a failed submission produces a Faulted outcome Regression for the rework: reuses F12's fixture (a submission that definitely never broadcast) to prove both facts from one cycle -- the outcome exists and carries the reversed fee, and the persisted window still reflects zero net spend. Also fixes a rustfmt diff on the PreBudgetResult::Fault variant Clippy's Rustfmt job flagged. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(rewards-claim): delete fee_window_poisoned, stop latching self-healing state Finding 1 (dig-node#594 pass 6): a future-dated clock is self-healing by construction (`t > now` goes false the moment real time passes it), but the engine ORed it into `self.fee_window_poisoned` and set that field `true` permanently -- an RTC glitch or VM resume froze the claim loop forever instead of until the skew passed. This is the third instance of one mechanism (pass 3 latched ChainSourceUnavailable, pass 4 left a stale cadence-gate `state`), so the fix removes the FIELD, not just the bug: with no `fee_window_poisoned` on `ClaimEngine`, `self.fee_window_poisoned = true` is a compile error, not a convention to remember. Per-cycle conditions (corrupt + future-dated-clock) now live in a `CycleConditions` value built fresh at the top of every `run_cycle` from `now` plus a freshly reloaded `RewardsClaimConfig`, used, and dropped -- never stored on the engine. `corrupt` is now re-read from disk every cycle too (it previously latched at construction only), matching what `ClaimLoopState::PersistedStateCorrupt`'s doc already claimed but the code never did. Rewrites the single-cycle f10 regression into a two-cycle test: cycle 1 with a future-dated clock refuses; cycle 2, after the clock catches up and the cadence elapses, MUST claim. The old one-cycle version was green whether the latch bug was present or not. Refs #594 * fix(rewards-claim): satisfy clippy doc-list indent and rustfmt Clippy failed with 3x doc_lazy_continuation on the PersistedStateCorrupt doc comment (types.rs:165-167): continuation lines of a `-` bullet must be indented under the marker, not left flush. Indent them. Rustfmt failed on the new fail_reserve_asset_for early-return in FakeChainPort::reserve_asset_id (engine.rs:888): the Err(...) call exceeded the line-length limit unwrapped. Let rustfmt wrap it. Refs #594 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(rewards-claim): red proof for corrupt-then-repaired stale read Cycle 1 refuses a corrupt fee-window file; the file is then repaired to valid values with a fully-spent window and a recent completed-cycle time. Cycle 2 must neither grant a fresh budget nor skip the cadence gate. Fails against current `with_persisted_fee_window`, which loads the three fee-window fields once at construction and never refreshes them from the per-cycle `cfg` -- see engine.rs:149-157, #594. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(rewards-claim): resync fee-window fields from disk every cycle `with_persisted_fee_window` only loaded fee_window_start_unix, fee_spent_in_window_mojos and last_cycle_completed_at once, at construction. Once the now-deleted fee_window_poisoned latch stopped masking it, a file corrupt at construction and repaired later left those three fields stuck on poisoned()'s None/0/None placeholders -- a fresh budget and a skipped cadence gate, and persist_fee_window then overwrote the repaired disk values with them. CycleConditions now carries the three fields from the SAME freshly reloaded cfg it already used for the corrupt/future-dated check, and run_cycle copies them onto self before the cadence gate or window-roll logic runs, but only on a read that is neither corrupt nor future- dated. This also fixes Finding 2b: future_dated_clock now reads cfg's own clocks instead of self's stale ones. Corrects the doc claim at the old lines 236-238 to describe what the code now does for both halves. Closes #594. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * refactor(rewards-claim): make disk the sole store for the fee window Delete `fee_window_start_unix`, `fee_spent_in_window_mojos` and `last_cycle_completed_at` from `ClaimEngine`. `run_cycle` already re-reads `RewardsClaimConfig` fresh every cycle for the corrupt/future-dated check, so caching a copy on the engine bought nothing and cost exactly the stale-read defect class F16 just fixed. A local `FeeWindowState`, scoped to one `run_cycle` call, now threads the in-flight values through `evaluate_budget_phase`/`uncommit_fee`/`persist_fee_window` instead. With no field left to cache into, a future `self.fee_window_start_unix = ...` outside this file is an E0609 compile error, the same enforcement `fee_window_poisoned`'s removal already has. No behaviour change: every early return, the corrupt/future-dated fail- closed path, the cadence gate, the window roll, write-then-spend pre-commit/uncommit and the per-claim ceiling are unchanged -- only where the three values live changed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> * chore(release): v0.256.0 Bump dig-node-service to v0.256.0 for release. This release includes: - Reward distributor prover loop (#593) - Peer reward claim loop (#594) - Reward prover status RPC (#595) * ci: scope commitlint to PR-introduced commits, fix title suffix check A develop -> main release-cut PR was linting main..develop, the full inherited commit range, instead of just the commits it introduces. Every commit in that range was already linted at its own PR while it was still mutable; re-linting it at cut time adds no information and cannot be satisfied once merged (gitlinks and rev-pinned deps make history immutable). Use commitDepth: 1 on a main-base PR; keep the full-range lint unchanged for develop-base PRs, where authors can still fix the commits. Also fix the PR-title lint's blind spot: GitHub's squash merge lands "$PR_TITLE (#$PR_NUMBER)" as the commit subject, about eight characters longer than the title alone, so a title that passes header-max-length can still produce an over-limit commit subject that nothing checks. Lint the exact string that will land. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: Claude <noreply@anthropic.com>
…unning (#607) * feat(mirror): persist mirror-bond coin ids (#575) * chore: open lane for #574 * feat(mirror): persist mirror-bond coin ids so a restart cannot double-create Bond identity was reconstructed from a live chain scan on every read (`mirror/observe.rs`), with no persistence of its own. A restart, a cold replica, or a lagging/flaky chain source all rendered a real, unspent, confirmed bond as "no bonds" -- and because the in-flight suppression is keyed on pending/submitted audit entries, a bond whose create had already CONFIRMED was not suppressed either, so the same short scan that emptied the read surface also cleared the one thing that would have stopped a second coin being paid for collateral that already exists (dig-node#574). Persist the (store, root, epoch) -> coin_id mapping in the EXISTING spend audit record (spend-audit.jsonl) rather than a new store: a mirror-coin create already writes store_id + AuditedBond{root, epoch} + amount there, and the coin id itself becomes durable the moment resolve_landed_spends confirms it. This adds the one missing piece -- the advertised URL a create carries -- and a read-side query, confirmed_mirror_bond, that returns the newest CONFIRMED record naming a triple. Chain stays authoritative. mirror::local_bond::recheck_missing_bonds never trusts the record: for a held bond the live scan did not cover, it asks the record for a candidate coin id, then re-verifies that SPECIFIC coin against chain via the same independent check (chain_bond_verdict) that verifies an untrusted peer's claimed bond. Only a fresh `Bonded` verdict is folded back in, as covered; `Unbonded`/`Unverified` fall through to an ordinary create, exactly as if no record existed. Version: 0.254.86 (patch -- per #522 the MSI ProductVersion minor field is exhausted and the counter lives in patch). Co-Authored-By: Claude <noreply@anthropic.com> * test(mirror): prove the recovery wiring end to end through PassRunner::run Adds two integration-level tests over the REAL pass pipeline, not just the isolated recheck_missing_bonds unit tests: a bond missing from the live scan with a chain-reverified durable record is recovered (no double create, correct Bonded state reported), and the control -- the same record but chain disproves it -- correctly falls through to an ordinary create. Together these are the concrete regression test for the cold-start/lagging-chain-source double-create scenario the ticket asked to have measured. Also refactors in_flight_creates to take the already-folded SpendLedger instead of re-reading the log itself, so PassRunner::run reads the audit file once per pass and shares it with the new recovery step, and fixes a doc comment on in_flight_creates that the recovery step would otherwise have made stale on landing ("a Confirmed create has a coin the chain observation already sees" is no longer unconditionally true). Co-Authored-By: Claude <noreply@anthropic.com> * chore(fmt): wrap long test signatures to satisfy rustfmt Co-Authored-By: Claude <noreply@anthropic.com> * chore(clippy): use slice::from_ref instead of cloning for a single-element slice Co-Authored-By: Claude <noreply@anthropic.com> * chore(release): bump to v0.254.89 Base branch moved to develop after PR #576 merged there at v0.254.88 (main and develop are currently identical), leaving this branch's carried-forward .88 as a zero-increment against the new base. Bumped to the next free integer after fetching and verifying both origin/main and origin/develop tip at .88. Co-Authored-By: Claude <noreply@anthropic.com> --------- Co-authored-by: Claude <noreply@anthropic.com> * fix(peer): count accepted relayed circuits in the connected pool (#579) serve_accepted_relay_conn served every accepted relayed circuit (full mTLS auth, full L7 peer RPC) while registering it nowhere, so connected_peers under-reported every relayed inbound peer -- the relay-leg twin of the direct-inbound defect #402/#523 already fixed. adopt_inbound_peer_in_pool now dispatches by TraversalKind: Relayed routes to dig-gossip's already-published adopt_relayed_inbound_handle (v0.32.0, the rev this repo already pins), every other tier keeps the unchanged adopt_direct_inbound_handle path. serve_accepted_relay_conn adopts before serving and releases after, mirroring the direct listener exactly. Refs: https://github.com/DIG-Network/dig_ecosystem/issues/3124 * fix(cli): guard the exit-code namespace shared with diga against collisions (#582) * chore: open lane for #3189 * fix(cli): guard the exit-code namespace shared with diga against collisions dign and diga deliberately share one process exit-code numbering (dig-app's outcome.rs says so in its own doc comment), so a number is free only if it is unoccupied ecosystem-wide. dig-node#407 assigned exit 7 to NODE_UNREACHABLE by checking only this repo's own table, where 7 genuinely was free -- and collided with diga's NOT_CONNECTED. A reviewer caught it by hand; nothing failed automatically. Adds scripts/check-exit-code-collisions.sh: parses both enums' code()/name() match arms straight from their own source -- this repo's ExitCode, and a live fetch of dig-app's outcome.rs at its default branch -- and fails if a number carries two different names, or if either side draws a number from the reserved shell signal range (126, 127, 128+N). Ships with an 18-case hermetic test harness (scripts/tests/check-exit-code-collisions.test.sh) covering the actual #407 collision shape, arm-order independence, arm-count mismatch, the reserved-range boundary from both sides, the live-fetch path itself, and fail-closed behaviour on an empty/missing/unreachable table. Wires a real (unstubbed) invocation into ci.yml's existing "Release-script tests" job so a collision introduced by a future PR, on either side, is a red required check on that PR -- not a note a reviewer has to catch. The fetch retries twice (2s backoff) since this becomes a required, network- dependent check; a fetch failure still fails closed after retrying, never silently passing as "diga has no codes". Updates SPEC.md 8.4 to point at the mechanical guard instead of leaving "re-check both tables" as unenforced prose, and records that the extension's WALLET_WS_ERR.NOT_CONNECTED = -33001 is a separate JSON-RPC error-code space, not a rival of this one. Adds a doc-comment to the existing transcribed collision test pointing future readers at the live script as the authoritative check; the transcription remains as a narrower, hermetic regression pin for the #407 shape specifically. No renumbering: every currently-assigned code is unchanged. Refs #3189 Co-Authored-By: Claude <noreply@anthropic.com> --------- Co-authored-by: Claude <noreply@anthropic.com> * fix(hygiene): port the lost-continuation guard to 4 crates, fix 48 corrupted strings (#3190) (#583) * chore: open lane for #3190 * fix(hygiene): port the lost-continuation guard to 4 crates, fix 48 corrupted strings Replicates dig-node-service::continuation_guard (dig-node#526/#501) into dig-node-core, dig-wallet, dig-runtime and dig-chat-protocol, line-for-line apart from crate-specific constants -- ported rather than reinvented, per dig_ecosystem#3190. Wiring the guard in surfaced 48 pre-existing lost-continuation defects the ticket's own "no measured corruption in these four crates" note did not anticipate: 36 in dig-node-core, 12 in dig-wallet, mostly test-assertion prose where a multi-line message lost its `\` continuation and shipped the source's own indentation as a mid-sentence space run (one as the worse `\n`-plus-indentation variant). All 48 are collapsed to the single space the sentence always meant, with surrounding indentation and wording otherwise untouched. Two lines are real column-alignment, not defects, and get a targeted EXCLUDED_LINE_RANGES entry on dig-node-core instead of a rewrite: download.rs's `claimed(...)` fixture-table trailing comments, and net.rs's `label : value` debug-print alignment. Refs https://github.com/DIG-Network/dig_ecosystem/issues/3190 Refs https://github.com/DIG-Network/dig_ecosystem/issues/3130 Co-Authored-By: Claude <noreply@anthropic.com> --------- Co-authored-by: Claude <noreply@anthropic.com> * feat(mirror): detect an IP change daily and reconcile mirror coins to the current advertise URL Automatic half (D1-D4) of the daily mirror-URL reconcile: derived personal-day offset, two-observation hysteresis, nine ordered gates with K sized as a self-funding prefix before any reclaim, `submitted` never `completed`, audit lines gain `reclaim_reason` + `trigger`. Gates at b4c09866: loop-reviewer PASS (review 5129272285), loop-security PASS @ 175304e1 (tree byte-identical, `git diff 175304e1 b4c09866` empty), adversarial loop-decider PASS (comment 5567083644; SHOULD-FIX findings ticketed separately). Refs DIG-Network/dig-node#570 Refs DIG-Network/dig_ecosystem#3203 * feat(serve): content hosting + serve path batch, v0.255.0 (dig_ecosystem#3212) Nine commits from the #3212 serve-path lane, gated at 899cc68f (reviewer review 5130425808, security comment 5568662836), plus the single semver bump to 0.255.0 for the develop -> main batch. - store_id/root case normalised at the CapsuleKey boundary; cache delete targets the matched entry - tier-0 occupancy reads the eviction-aware ledger - profile-sync outbound budget in bytes; announcer asked first - melt confirmation depth on the terminal spend, fail-closed - EngineWarming (-32002) while the peer tier attaches, never -32004 - window completeness derived from the bytes read - deps: dig-stun 0.2, chia-query 0.24.3, dig-nat 0.21.2, dig-logging 0.2.2 Refs DIG-Network/dig_ecosystem#3212 * chore: untrack gitnexus-generated agent files (#590) * chore: untrack gitnexus-generated agent files These files were generated by `gitnexus analyze` as a side effect of indexing this repository. They are development-loop private tooling output, not product code, and carry no secrets. They are removed from tracking going forward via .gitignore; history is deliberately NOT rewritten. Refs #3177 * chore: drop private-repo reference from gitignore comment The ignore comment named a private repository and an internal issue number in a public file, which is the same disclosure class this change set exists to remove; the reference is dropped and the guidance kept. * feat(rewards): always-on prover loop engine -- honest liveness, type-enforced self-exclusion, bounded spend (#593) The always-on reward-prover engine: ~2,000 lines under `crates/dig-node-core/src/rewards/`, built against the merged `dig-rewards-coin` SPEC. Library only -- nothing spawns it, and the sole production `RewardsChainPort` refuses every call, so it cannot spend. Wiring the composed system is #3265, which carries its own gate. The epic's premise -- "anytime the process isn't running, rewards are not being distributed" -- is half wrong, and the false half is the dangerous one. `Sync`, `NewEpoch` and `InitiatePayout` need no manager authority, so funder downtime does not stop rewards: it FREEZES THE ENTRY SET while accrual and payouts continue. Peers that stopped mirroring keep earning; peers that started cannot begin. That shaped the whole design. Liveness honesty (SPEC 2.4). The status record carries no `healthy`/`ok`/`up` boolean and no precomputed staleness, because a wedged loop cannot report its own wedging -- whatever it last wrote stays there, so a writer-set flag reads true forever after the failure it exists to reveal. The reader derives staleness from `last_cycle_completed_at` against `observed_at` and its own clock. A recursive JSON-key test enforces the absence at every nesting depth; asserting on keys and never substrings, since `ProverState::Running` legitimately serializes the VALUE "running". The one legal staleness signal is chain-derived (SPEC 12.4: 48 hours AND a non-zero reserve, from the singleton's own spend history) and lives on the distributor read, where a wedged prover cannot fake it. Self-exclusion is a compile error, not a habit. dig-node#261's lesson is that an invariant enforced on some paths is not an invariant. `admit` is the single admission point, checks both SPEC 5.2 coordinates (own peer_id OR a payout puzzle hash this wallet controls), and mints an `AdmittedPeer` with private fields and no public constructor -- so `EntryAction::Add` cannot be built by a path that skipped admission. A prover's own fault can never strike a peer. `GateError` is a distinct type from `GateIneligibleReason` and `record_prover_fault` takes `&self`, so SPEC 3.6 clause 4 is enforced by the borrow checker rather than by comment. Without that, a misconfigured operator -- one missing mirror-collateral epoch ordinal -- would strike every peer at once and evict its entire 250-entry set in three hours, each eviction a fee it pays plus a settlement out of its own reserve. The money bounds are stated where a human reads them (`rewards/mod.rs`): 24 bundles/day, 192 entry actions/day, a fee ceiling of 24x the configured standard fee, 192 removals/day worst case with 96/day sustained churn. Recorded honestly: SPEC 6.3's rate bound and fee ceiling are ONE control, not two. Three gate rounds, every leg fresh-context. Round 3 at this head: reviewer PASS, adversarial decider RATIFY (leg closed), security CHANGES-REQUIRED on a finding the decider ratified deliberately -- adjudicated in https://github.com/DIG-Network/dig-node/pull/593#issuecomment-5601784815 and carried to #3265 with the remedy corrected, because the proposed fix would have persisted a poison flag to the very store whose writes were failing. Found and fixed under gate: a census ordinal off by one in both directions (SPEC 4.6 requires n-1 exactly); an unreachable grace window leaving a named constant with no reader; a missing `NewEpoch` spend; an absent-ordinal path attributing a prover fault to peers; and a daily fee ceiling 24x too high because a per-bundle fee was consumed as a daily ceiling. Refs DIG-Network/dig_ecosystem#3250 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat: serve dig.getRewardProverStatus at Tier::Control (#595) * chore: open lane for #3269 Bump dig-rpc-protocol to 0.11.0 (adds dig.getRewardProverStatus and the other reward RPC methods to the wire). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(rewards): fail-closed Reward-tier guard + dig-rpc-protocol 0.11 line assertion - dependency_tree.rs: assert the resolved dig-rpc-protocol line is 0.11 (was 0.10); documents the known-red two-version state pending the dig-peer 0.14.0 / dig-download 0.23.0 cascade (#3269). - reward_methods_tier_guard.rs: fail-closed guard over the live Method::ALL catalogue -- every Reward-named method must be Tier::Control and not peer-reachable, so a fifth reward method added later is caught at the wrong tier automatically rather than inheriting a wrong default (binds #3261's rule node-side). - peer.rs: sibling unit test exercising the real (pub(crate)) is_peer_reachable_method, since an external integration test cannot see it -- same guard, executed against this node's own allowlist rather than only the shared crate's. Refs #3269 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * style: remove trailing blank line in reward_methods_tier_guard.rs * feat(rpc): serve dig.getRewardProverStatus at Tier::Control Adds the missing handler for PR#595: a new reward_prover_statuses registry + accessors on Node (empty until #3265 spawns a prover loop, so the registry read is real, not a stub), a dispatch.rs arm inside the Method enum match (never the string pre-match), and a field-for-field mapping from dig-node-core's internal rewards::state::RewardProverStatus (camelCase-tagged) onto dig-rpc-protocol 0.11's wire RewardProverStatus (snake_case-tagged struct, camelCase-tagged ProverState value), widening entry_count u32 -> u64 explicitly and hex-encoding the three [u8; 32] identity fields. An all-zero launcher_id (what an uninitialised registry slot hex-encodes to) is omitted at this boundary rather than rendered as a real distributor with a plausible-looking id -- the money-hole class the dig-rewards-coin driver's adversarial gates found three times. Tests (in dig-node-core::lib.rs's existing test module, where the pub(crate) registry accessors are visible) drive the real dispatch entry point (handle_rpc -> RpcDispatch::dispatch -> the Method arm) and assert field-for-field on the serialized JSON body: populated registry, empty registry (-> {"statuses": []}), zero-id omission, tier/peer-reachability, enum-match-not-string-prematch, and launcher_id filtering. The no-health-boolean / no-staleness assertion is by key set, not substring. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * style: rustfmt the reward-prover-status registry + tests Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * chore(deps): bump dig-download 0.23, dig-peer 0.14, dig-rpc-protocol 0.11 Closes the two-versions-of-dig-rpc-protocol split (#3269): dig-download 0.23.0 and dig-peer 0.14.0 both now resolve dig-rpc-protocol ^0.11, matching dig-node-core's own dig-rpc-protocol = "0.11.0" line, and dig-node-service's two 0.10 lines (main dep + dev-dependency restatement for openrpc_drift_guard.rs) move to 0.11 to match. Manifests only. cargo update / Cargo.lock intentionally NOT run yet: a fourth capper, dig-peer-selector ("0.11" in dig-node-core/Cargo.toml), still requires dig-peer = "^0.13" in every published version through 0.11.1, so the tree cannot fully resolve until a dig-peer-selector release picks up dig-peer 0.14. CI will stay red on this commit for that reason, which is expected. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * feat(rpc): close dig-rpc-protocol 0.11 cascade + attribute payout figures Bump dig-peer-selector 0.11 -> 0.12 (the release that moves onto dig-peer ^0.14) and run cargo update, closing the two-versions-of-dig-rpc-protocol split: Cargo.lock now resolves exactly one dig-rpc-protocol, at 0.11.0, alongside dig-peer 0.14.0, dig-download 0.23.0, dig-peer-selector 0.12.0. Add a subject-attribution test and doc comments to reward_prover_status_to_wire: total_paid_out_base_units and reserve_base_units are per-distributor totals (this distributor's payout to ALL its mirrors, and this distributor's own reserve), never the querying node's own earnings and never summed/cross-attributed across distributors. This is the defect class a sibling adversarial gate found in dig-app#403's rewards pane, which rendered a distributor total as one mirror operator's personal earnings and overstated by up to 250x. Checked: rewards/state.rs, rewards/port.rs and rewards/mod.rs contain no Eligible/verdict/payout_hash symbol, so nothing in this mapping surfaces a persisted EligiblePayoutHash verdict. Refs #3269 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(rpc): silence dead_code on register_reward_prover_status pending #3265 Clippy's non-test lib target has no production caller for register_reward_prover_status yet, because #3265 (the always-on prover loop that would call it from bring-up) has not landed -- only tests call it today. cfg_attr(not(test), allow(dead_code)) stands in for that missing caller until #3265 wires a real one. Refs #3269 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(rpc): make the all-zero identity guard non-silent and cover all three fields Security (blocking) and the adversarial leg both found the same defect in the zero-launcher_id filter: it checked only launcher_id, so a registration bug that zeroed store_id or root beside a valid launcher_id would pass through as a plausible record, and dropping the bad record silently destroyed the evidence a registration bug happened at all -- SPEC Sec2.4 clause 1's exact prohibition. zeroed_identity_fields() now checks launcher_id, store_id AND root. The dispatch filter still excludes a record with any zeroed field (never renders an uninitialised slot as a real distributor), but first fires a tracing::warn! naming which field(s) were zero, so a bad registration is observable rather than swallowed. Kept isolated in dispatch.rs rather than woven into the wire mapping, since this belongs at #3265's writer once that lands. Replaced get_reward_prover_status_omits_an_all_zero_launcher_id (which proved the omission but not the observability, and never exercised a zeroed store_id/root beside a valid launcher_id) with get_reward_prover_status_logs_and_excludes_a_zeroed_identity_field, covering both a zeroed launcher_id and a zeroed store_id beside a valid launcher_id, and asserting the tracing::warn! output via the crate's existing capture_sync_logs test utility. Fixed a now-false "Known-red" doc comment on tests/dependency_tree.rs::the_workspace_carries_exactly_one_module_wire_crate: the dig-peer 0.14.0 / dig-download 0.23.0 / dig-peer-selector 0.12.0 cascade already landed in this PR's Cargo.toml/Cargo.lock, so the assertion is green, not red. Assertion itself untouched -- still exact-version. Refs #3269 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(rpc): correct a born-false "shipped dig-app 15.5.0" doc claim dig-app's latest release is v15.4.0 -- there is no v15.5.0 tag -- and dig-app#403 (the pane that would consume dig.getRewardProverStatus) is OPEN and unmerged. Point the doc comment at the real, unmerged consumer instead so a future reader doesn't take this as evidence a shipped consumer depends on the guard, which would wrongly discourage relocating it to #3265's writer. Refs #3269 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(rpc): correct doc placement, assert the zeroed field by name, treat root as an observation Three findings from the correctness gate on PR#595 at 134864a9. 1. The zeroed-identity helper's doc block was spliced onto the end of reward_prover_status_to_wire's block with no separator, so the wire-mapping rationale documented a boolean predicate and the mapping function was left with no doc at all. Each doc block now sits above the item it describes. 2. The log assertion `logs.contains("launcher_id")` was a tautology: the warn emits launcher_id as a structured field on every fire, so the property the guard exists to add -- naming which field was zeroed -- was unasserted. Deleting `zeroed_fields = ?zeroed` from the warn left every assertion green. The test now asserts the zeroed_fields value itself, which the fixture makes exact and disjoint across cases. 3. `root` is an observation, not an identity. A registered prover that has not completed its first cycle plausibly has no root, and a writer that zero-inits it would have made a healthy prover invisible. A zeroed launcher_id or store_id still excludes the record; a zeroed root alone warns and returns. Refs #3269 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(rpc): restore zeroed_fields structured field dropped from the pushed warn The previous commit (080be3df) landed with `zeroed_fields = ?zeroed` missing from the tracing::warn! call in the GetRewardProverStatus filter -- a one-line regression introduced while proving the new log assertion goes red without it, never restored before the commit was made. Without this field the log line never names WHICH field was zero, so an operator sees only that something was excluded, and the test asserting `zeroed_fields=[...]` per case would fail. Restored; all 7 reward-prover-status tests green. Refs #3269 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(rpc): split zeroed-field logging by level -- WARN for a missing identity, DEBUG for a zeroed root A zeroed launcher_id or store_id is a real registration bug: the record is excluded and now logs at WARN, naming the exact field(s) via `zeroed_fields=[...]`. A zeroed root alone is an ordinary pre-first-cycle state, not a fault: the record is still returned, and now logs at DEBUG instead of WARN, so an operator polling this endpoint sees warn-level volume proportional to real registration bugs, not to every not-yet-cycled prover on every poll. Updated the doc comments on `zeroed_fields`, the dispatch filter and the test to describe the level split, and extended the regression test to assert on level (WARN vs DEBUG) as well as the `zeroed_fields` value. Proved both directions: flipping the DEBUG branch back to WARN turns the test red on the level assertion; flipping the field-name assertion back to a bare `contains("launcher_id")` would have passed unconditionally (the prior tautology) and is no longer possible since the assertions now pin `zeroed_fields=[...]` plus the level string. Refs #3269 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> * feat(rewards): peer-side claim loop -- watch distributors, claim on cadence (#3251) (#594) * feat(rewards): peer claim loop skeleton -- discovery, cadence, config, claim port * test(rewards): write all twelve acceptance tests for the peer claim loop * feat(rewards): wire the seven rewards_claim submodules into the crate mod.rs declared no submodules, so types.rs/port.rs/config.rs/cadence.rs/ parser.rs/engine.rs/hints.rs (1,435 lines, 25 tests) were never part of the crate and never compiled. Declare them and re-export the public surface. * style(rewards): cargo fmt the rewards_claim submodules * chore(deps): bump dig-node-control-interface 0.33->0.35, dig-rpc-protocol 0.10->0.11.0 dig-rpc-protocol 0.11.0 is merged and tagged upstream; the other dig-*/chia-* deps of dig-node-service were already at the latest permitted-by-caret version in Cargo.lock. crates/dig-node-core/Cargo.toml is untouched (#3250's file set). * chore(deps): revert dig-node-control-interface and dig-rpc-protocol bumps Both create a duplicate-version split in this PR's scope and neither can be closed without editing a sibling crate's manifest this lane does not own: - dig-rpc-protocol 0.11.0 duplicates against dig-node-core/Cargo.toml:194 ("0.10.2"), which is #3250's live file set (dig-node#593). - dig-node-control-interface 0.35.0 duplicates against dig-wallet/Cargo.toml:81 ("0.33"), a sibling crate this lane does not own; the observed Clippy break (BalanceAsset/Asset type-identity mismatch, missing url_reconcile/url_current/urls fields) came from THIS duplicate, not from dig-rpc-protocol. Both belong to their own sequenced dep-bump unit of work, not this ticket. * fix(rewards-claim): fault laundering, permanent no-entry blacklist, fee ceiling magnitude Three independent gates on dig-node#594 (51516e62) found four logic defects; this addresses A, B and C per the corrected fix brief (D is documented only, not fixed here per the brief's own instruction). Defect A -- the anti-silence surface laundered every real fault into `Nominal`: - A1: `fault_reported` had no fault-bearing ClaimLoopState to fall through to, so a chain adapter erroring every cycle read `Nominal` forever. Added `ClaimLoopState::Faulted { cycles }`, outranking Nominal/ClaimableButNotClaiming, under ChainSourceUnavailable. - A2: inverted the test that asserted A1's bug as correct behaviour. - A3: `ClaimableButNotClaiming` compared a per-cycle snapshot (`distributors_claimable`) against a lifetime-cumulative counter (`claims_submitted`), so it latched healthy forever after one lifetime success. Added `claims_submitted_this_cycle` (per-cycle) as the correct comparand; kept `claims_submitted` as a cumulative counter. - A4: `last_discovery_at`/`last_cycle_at` were stamped even on a failed discovery or an all-faulted cycle, destroying the staleness signal a reader depends on. Now only stamped on success; added `last_attempt_at` to prove liveness separately. `fault_reported` and `distributors_faulted` now reset per cycle instead of latching for the process's lifetime. Defect B -- "terminal, stop retrying" was implemented as a process-lifetime blacklist (`terminal_no_entry: HashSet<Bytes32>`, never cleared). That blocked SPEC 12.5 clause 2's re-entry path (evicted, re-challenged, re-admitted never claims again) and permanently punished a peer that discovered a distributor before the funder's AddEntry landed. Removed the blacklist entirely -- `own_entry` is a cheap chain read, re-issued every cycle for every candidate, matching clause 3's "never cache across cycles". `NoEntrySlot` is now a per-cycle observation, not a lifetime sentence. Defect C -- the fee ceiling didn't bind anything and there was no aggregate cap: - C1: default `CLAIM_FEE_CEILING_MOJOS_DEFAULT` lowered from 1_000_000_000 (transplanted from `MIRROR_SPEND_FEE_CEILING_MOJOS`, sized for a mirror-coin spend) to 200_000 -- 2x the observed routine Chia fee range (5,000-100,000 mojos), so it actually binds instead of leaving 4-5 orders of magnitude of slack. - C2: added a per-cycle aggregate fee budget (`max_cycle_fee_budget_mojos`, default 10x the per-claim ceiling) checked across all claims in a cycle, closing the attacker-cost gap where funding K distributors could force a victim to spend K x the per-claim ceiling per cycle. New `ClaimOutcome::SkippedCycleBudgetExhausted`. Tests: rewards_claim test count 26 -> 37 (11 new: repeated_discovery_faults_never_ read_as_nominal, failed_discovery_leaves_last_discovery_at_unchanged, a_reported_ fault_surfaces_as_faulted_not_nominal, a_lifetime_submission_does_not_mask_a_ later_cycle_that_submits_nothing, no_entry_slot_then_re_admitted_produces_a_claim_ on_the_later_cycle, distributors_each_under_ceiling_do_not_collectively_exceed_ the_cycle_budget, the_default_per_claim_ceiling_actually_binds_a_routine_fee, plus renamed/rewritten no_entry_slot_is_non_terminal_and_re_checked_every_cycle). Refs #3251 * fix(rewards-claim): refuse a claim entry for the wrong payout puzzle hash CI fix: cadence.rs's RewardsClaimConfig literal was missing the max_cycle_fee_budget_mojos field added in the previous commit (E0063, caught by CI's Clippy/Test jobs -- the local cargo check for this workspace is too slow to use as the compiler here). Defect E (security-gate finding, folded in before this pass closes): submit_initiate_payout was called with entry.payout_puzzle_hash -- whatever the chain port handed back -- with no check against this node's own own_payout_puzzle_hash. UnavailableClaimChainPort is the only production adapter today so nothing can exploit this yet, but the whole point of the ClaimChainPort seam is that #3249 swaps in a real adapter with nothing above it changing, so deferring this would ship the landmine live with no review pass watching for it. Added an equality guard before the spend: a mismatch refuses to submit, counts (ClaimStatus::claims_refused_payout_mismatch), surfaces its own named outcome (ClaimOutcome::PayoutPuzzleHashMismatch), and is reported as a fault (a divergent entry means the port is confused or hostile, not that there is nothing to claim) -- never corrected by substituting our own hash and proceeding. Defect D: documented, not wired, per instruction -- added the "not yet wired into node startup" paragraph to mod.rs's module doc (the PR body carries the same paragraph) so the next reader arrives at the caveat in the code, not only in a merged PR description. Refs #3251 * fix(rewards-claim): B1 -- ClaimableButNotClaiming is a magnitude comparison, not a zero-test submitted_this_cycle < claimable_this_cycle now fires the anti-silence state, carrying the shortfall as ClaimableButNotClaiming { claimable, submitted }. The previous submitted_this_cycle == 0 zero-test let one submission mask any number of same-cycle skips (claimable=10, submitted=1 read Nominal). Also folds in B3's precedence fix (ChainSourceUnavailable > Faulted > ClaimableButNotClaiming > Idle > Nominal) and the per-distributor payout_hash_mismatches_this_cycle counter so a per-distributor fault can no longer pin the cycle-wide Faulted state, plus R2's rename of terminal_no_entry_slot to no_entry_slot_this_cycle now that it is no longer terminal. * fix(rewards-claim): B2/B3 -- value-ordered budget with rotation, per-distributor fault isolation B2: run_cycle now splits into a pre-budget phase (asset/entry/hash/threshold checks, producing the claimable set) and a budget phase, ordering the claimable set by accrued value descending before applying the fee ceiling and cycle budget. Dust distributors (low accrued value regardless of attacker-controlled fee) now sort last and are the ones the budget drops, closing the claim-suppression attack where ten high-fee dust distributors could consume the whole cycle budget ahead of a victim's real earnings. A rotation_cursor tie-breaks only WITHIN equal-accrued-value tiers so a genuinely tied honest tail that exceeds one cycle's budget every cycle still rotates through and is eventually served, rather than dropping the same tail forever. B3: the payout-hash mismatch check in evaluate_pre_budget now increments the per-distributor payout_hash_mismatches_this_cycle counter instead of setting fault_reported, so one hostile or buggy entry can no longer pin the cycle-wide Faulted state and bury ClaimableButNotClaiming for every other healthy distributor. R2: terminal_no_entry_slot -> no_entry_slot_this_cycle throughout. * fix(rewards-claim): R5 -- document not-yet-wired loop; persist B2 rotation cursor An operator reading their own rewards-claim.json and seeing enabled: true has no way to know from that file alone that no startup path constructs a ClaimEngine yet (#3268) -- mod.rs said so, but a config file reader does not arrive at a module doc. Also gives RewardsClaimConfig a rotation_cursor: Option<Bytes32> field so B2's tie-break cursor survives a save/load round-trip -- an in-memory-only cursor resets on every restart, which would starve a legitimately tied honest tail forever on any node that restarts daily. * fix(rewards-claim): clippy collapsible-match + SPEC v0.1.3 wording refresh Collapses the nested if into the outer match arm in run_cycle (clippy::collapsible_match). Also refreshes NoEntrySlot / no_entry_slot_this_cycle doc comments now that dig-rewards-coin v0.1.3's SPEC §12.5 amendment is merged and tagged: absence is terminal for one claim attempt only, never for the distributor, must not be cached, and must not accumulate into a permanent exclusion set -- confirming rather than diverging from the re-read-every-cycle behaviour already implemented. * fix(rewards-claim): add missing rotation_cursor field in cadence.rs test literal Struct literal in the cadence test module was not updated when RewardsClaimConfig gained rotation_cursor (R5 commit) -- CI's Clippy/Test jobs caught the missing field (E0063) that a local cargo check could not (killed by memory pressure before this workspace-wide build completed). * fix(rewards-claim): F1 -- ChainSourceUnavailable is per-cycle, never a latch compute_state() compared against self.state -- last cycle's OWN computed output -- so once any cycle took an Unavailable port path, every later cycle re-asserted ChainSourceUnavailable forever, even after the chain came back and real claims were submitting. A node still syncing, or one dropped connection, was enough to trip this permanently. Add ClaimStatus::chain_unavailable_this_cycle, reset to false at the top of every run_cycle and set true only on a cycle that actually took the Unavailable path; compute_state now reads that flag instead of self.state, so the reading is live again. Test: a_transient_unavailable_cycle_does_not_latch_state_for_the_rest_of_the_process (engine.rs) drives cycle 1 unavailable, cycle 2 healthy with a submission, and asserts cycle 2 reads Nominal. Plus a compute_state-level regression in types.rs. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(rewards-claim): F3/F4/F5 -- fix stale counters, dedup candidates, correct version doc F3: reset EVERY per-cycle counter (distributors_known/with_own_entry/ claimable/faulted, claims_submitted_this_cycle, no_entry_slot_this_cycle) at the TOP of run_cycle, before any early return. The three ChainUnavailable early-return paths skip the end-of-function assignment block entirely, so a cycle that hit one used to leave the PRIOR cycle's counts sitting on self.status while last_attempt_at stamped fresh for THIS cycle -- a stale count under a fresh timestamp, exactly what SPEC §2.4's staleness reasoning forbids. types.rs's doc sentence for no_entry_slot_this_cycle now correctly says it is dated by last_attempt_at (the field stamped unconditionally every cycle), not last_cycle_at. F4: dedup `candidates` by launcher id before phase 2. A real adapter scanning §1.3 launch comments across every (store_id, root) this node mirrors can plausibly return the same launcher id twice; without dedup phase 2 would evaluate it twice and submit InitiatePayout twice against one entry slot in one cycle -- the second spend is invalid but the fee is paid anyway. F5: dig-rewards-coin is v0.1.3, published on crates.io -- correct the stale "v0.1.1" module-doc claim. Tests: a_chain_unavailable_cycle_does_not_leave_prior_cycles_counters_stale (F3), a_duplicated_launcher_id_submits_exactly_once (F4). Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(rewards_claim): fold payout-hash mismatches into the shortfall predicate (F2) A payout-hash mismatch never enters the eligible set, so it was counted in NEITHER distributors_claimable NOR claims_submitted_this_cycle -- the shortfall lived in neither term of compute_state's magnitude comparison. All-K-distributors mismatching therefore read Nominal (falsely healthy). Fold payout_hash_mismatches_this_cycle into the comparison's denominator: submitted < claimable + mismatches. The result is ClaimableButNotClaiming (a shortfall), never the cycle-wide Faulted -- Defect B3 stays fixed. Inverts the assertion at what was engine.rs:1305 (a_payout_mismatch_never_sets_the_cycle_wide_fault_or_masks_other_distributors): it previously asserted ClaimLoopState::Nominal across three cycles of an ongoing mismatch, which pinned the defect as intended behaviour (an A2-class test). It now asserts ClaimableButNotClaiming { claimable: 1, submitted: 1 }. Adds all_distributors_mismatching_is_a_shortfall_not_nominal, covering the brief's exact "what if every distributor refuses for the same reason" case. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(rewards-claim): F1/F3 -- AtomicU32 in fake ports, not Mutex, keeps discovery Send CI's Clippy job (the compiler for this crate, per brief) caught it: holding a std::sync::MutexGuard across the .await in FlakyThenHealthyPort and HealthyThenUnavailablePort's discover_distributors made the returned future not Send, which #[async_trait]'s generated trait signature requires. Neither fake needs a lock -- each holds one call counter, incremented once per call, never read-modify-written across an await point. AtomicU32's fetch_add removes the guard (and the Send bound violation) entirely. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(rewards-claim): F7 -- persist the fee window and cadence gate across restart The per-cycle aggregate fee budget and the 24h cadence clock both lived only in memory: `spent_this_cycle_mojos` was a `run_cycle` local and nothing on disk recorded a completed cycle. Every fresh process got a full `max_cycle_fee_budget_mojos` and an empty cadence clock, so a node stuck in a crash-restart loop could spend unbounded XCH on fees, one full budget per restart. Adds three `#[serde(default)]` fields to `RewardsClaimConfig` (`fee_window_start_unix`, `fee_spent_in_window_mojos`, `last_cycle_completed_at`) and a new opt-in `ClaimEngine::with_persisted_fee_ window(dir, cadence_seconds)` that: - restores the window/cadence state from `dir` at construction, - refuses to start a cycle until the cadence has elapsed since the last completed one, - rolls a fresh budget window only once the cadence has elapsed since it opened, otherwise keeps enforcing the budget against the persisted spend, - persists the spend BEFORE every chain submission (write-then-spend), never batched to cycle end, and persists the completed-cycle timestamp when a cycle finishes. Engines that never call `with_persisted_fee_window` (every pre-F7 test) are unaffected -- this is additive, opt-in state beside the existing rotation cursor, not a change to B2's value-ordering or rotation mechanism. `ClaimStatus`'s own counters stay in-memory on purpose (observability, meant to reset on restart); only the spend bound and the cadence gate persist. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(rewards-claim): F7 -- update cadence.rs test literal for new persisted fields The three new persisted RewardsClaimConfig fields (fee_window_start_unix, fee_spent_in_window_mojos, last_cycle_completed_at) broke this crate's only remaining full struct literal outside config.rs/engine.rs's own test modules -- E0063 missing fields, caught by CI's Clippy job. Switched to ..RewardsClaimConfig::default() so the next added field cannot break this literal again, the same fix already applied once before for rotation_cursor. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(rewards-claim): F8/F14 -- atomic state write, fail closed on a corrupt window Salvaged from a lane killed by a weekly cap before it could commit. Uncompiled at commit time; CI is the compile signal. Covers the fourth gate pass findings on the F7 persisted spend bound: - F8: RewardsClaimConfig::save_to now writes atomically (temp file + rename in the same directory), reusing the pattern already used by mirror/reconcile_state.rs for the same class of state. load_from distinguishes an ABSENT file (clean first run, defaults are correct) from a PRESENT but unparsable one, which fails CLOSED: the window is treated as fully spent and nothing is submitted. Never Default, and never a silent clamp downward, which would hand back the budget the corruption was hiding. - F14: the budget comparison uses saturating arithmetic so a corrupt disk-seeded fee_spent_in_window_mojos cannot panic under the release profile's overflow-checks. - F9/F10/F12/F13 in progress in the same files. Refs #3251 * fix(rewards-claim): negate with ! rather than the unimported Not trait Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com> * fix(rewards-claim): F13 -- correct stale test literals to the folded shortfall compute_state (types.rs) already reported the folded shortfall denominator (distributors_claimable + payout_hash_mismatches_this_cycle) as `claimable` -- that part of F13 landed in f478516a. The two engine.rs tests asserting this state were written against the pre-fold, un-folded numbers and never updated, so CI showed the implementation producing the correct folded value (`claimable: 2`, `claimable: 1`) while the test literals still expected the stale un-folded one (`claimable: 1`, `claimable: 0`). Update both literals -- and the comments describing them -- to the folded values the F13 fix actually produces. No production code change; compute_state's predicate and payload were already correct. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * feat(rewards-claim): add ClaimOutcome::Faulted variant Add the seventh ClaimOutcome variant: the type could only say a peer was legitimately not paid, never that a chain call failed. Carries the launcher id, a bounded (200 char) copy of the chain port's error text, and whether a pre-committed fee was reversed, so a reader can tell no money moved. Engine wiring at the two fault arms (engine.rs:332, :377) follows in the next commit. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(rewards-claim): both fault arms now push ClaimOutcome::Faulted engine.rs:332 and :377 used to increment `faulted` and discard the outcome, leaving a definitively-failed claim absent from the outcome stream -- indistinguishable from a cycle that never touched that distributor. Both PreBudgetResult::Fault and BudgetPhaseResult::Fault now carry the chain port's (bounded) error text, and the submit_initiate_payout failure path also carries the fee it reversed, so a reader can tell no money moved. The counter stays; it is not a substitute for the outcome. 7 call sites needed updating: 3 PreBudgetResult::Fault constructions (reserve_asset_id, own_entry, payout_threshold), 2 BudgetPhaseResult::Fault constructions (required_fee_mojos, submit_initiate_payout), and the 2 consuming match arms -- exactly the set that was silently discarding a failure before this change. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(rewards-claim): a failed submission produces a Faulted outcome Regression for the rework: reuses F12's fixture (a submission that definitely never broadcast) to prove both facts from one cycle -- the outcome exists and carries the reversed fee, and the persisted window still reflects zero net spend. Also fixes a rustfmt diff on the PreBudgetResult::Fault variant Clippy's Rustfmt job flagged. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(rewards-claim): delete fee_window_poisoned, stop latching self-healing state Finding 1 (dig-node#594 pass 6): a future-dated clock is self-healing by construction (`t > now` goes false the moment real time passes it), but the engine ORed it into `self.fee_window_poisoned` and set that field `true` permanently -- an RTC glitch or VM resume froze the claim loop forever instead of until the skew passed. This is the third instance of one mechanism (pass 3 latched ChainSourceUnavailable, pass 4 left a stale cadence-gate `state`), so the fix removes the FIELD, not just the bug: with no `fee_window_poisoned` on `ClaimEngine`, `self.fee_window_poisoned = true` is a compile error, not a convention to remember. Per-cycle conditions (corrupt + future-dated-clock) now live in a `CycleConditions` value built fresh at the top of every `run_cycle` from `now` plus a freshly reloaded `RewardsClaimConfig`, used, and dropped -- never stored on the engine. `corrupt` is now re-read from disk every cycle too (it previously latched at construction only), matching what `ClaimLoopState::PersistedStateCorrupt`'s doc already claimed but the code never did. Rewrites the single-cycle f10 regression into a two-cycle test: cycle 1 with a future-dated clock refuses; cycle 2, after the clock catches up and the cadence elapses, MUST claim. The old one-cycle version was green whether the latch bug was present or not. Refs #594 * fix(rewards-claim): satisfy clippy doc-list indent and rustfmt Clippy failed with 3x doc_lazy_continuation on the PersistedStateCorrupt doc comment (types.rs:165-167): continuation lines of a `-` bullet must be indented under the marker, not left flush. Indent them. Rustfmt failed on the new fail_reserve_asset_for early-return in FakeChainPort::reserve_asset_id (engine.rs:888): the Err(...) call exceeded the line-length limit unwrapped. Let rustfmt wrap it. Refs #594 Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * test(rewards-claim): red proof for corrupt-then-repaired stale read Cycle 1 refuses a corrupt fee-window file; the file is then repaired to valid values with a fully-spent window and a recent completed-cycle time. Cycle 2 must neither grant a fresh budget nor skip the cadence gate. Fails against current `with_persisted_fee_window`, which loads the three fee-window fields once at construction and never refreshes them from the per-cycle `cfg` -- see engine.rs:149-157, #594. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * fix(rewards-claim): resync fee-window fields from disk every cycle `with_persisted_fee_window` only loaded fee_window_start_unix, fee_spent_in_window_mojos and last_cycle_completed_at once, at construction. Once the now-deleted fee_window_poisoned latch stopped masking it, a file corrupt at construction and repaired later left those three fields stuck on poisoned()'s None/0/None placeholders -- a fresh budget and a skipped cadence gate, and persist_fee_window then overwrote the repaired disk values with them. CycleConditions now carries the three fields from the SAME freshly reloaded cfg it already used for the corrupt/future-dated check, and run_cycle copies them onto self before the cadence gate or window-roll logic runs, but only on a read that is neither corrupt nor future- dated. This also fixes Finding 2b: future_dated_clock now reads cfg's own clocks instead of self's stale ones. Corrects the doc claim at the old lines 236-238 to describe what the code now does for both halves. Closes #594. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> * refactor(rewards-claim): make disk the sole store for the fee window Delete `fee_window_start_unix`, `fee_spent_in_window_mojos` and `last_cycle_completed_at` from `ClaimEngine`. `run_cycle` already re-reads `RewardsClaimConfig` fresh every cycle for the corrupt/future-dated check, so caching a copy on the engine bought nothing and cost exactly the stale-read defect class F16 just fixed. A local `FeeWindowState`, scoped to one `run_cycle` call, now threads the in-flight values through `evaluate_budget_phase`/`uncommit_fee`/`persist_fee_window` instead. With no field left to cache into, a future `self.fee_window_start_unix = ...` outside this file is an E0609 compile error, the same enforcement `fee_window_poisoned`'s removal already has. No behaviour change: every early return, the corrupt/future-dated fail- closed path, the cadence gate, the window roll, write-then-spend pre-commit/uncommit and the per-claim ceiling are unchanged -- only where the three values live changed. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com> * feat(rewards): chain port + listRewardDistributors (unit 2) (#604) * chore: open lane for #3269 (unit 2 -- rewards chain port + listRewardDistributors) * chore(rewards): add dig-rewards-coin 0.2 dep; record blocked-reader finding dig-rewards-coin 0.2.0 is published but ships no chain reader (its own state.rs module doc: SPEC 12.1's read_distributor is withheld pending DIG-Network/dig_ecosystem#3267). Separately, no registry in this codebase records which distributors this node funds. A "real" RewardsChainPort adapter over 0.2.0 therefore has no honest way to answer any of the four trait methods with live data yet -- reimplementing read_distributor or inventing a funded-distributor registry would be exactly the unreviewed money-shape guess kernel invariant 6 says to escalate instead of build. UnavailableChainPort remains the only production adapter; port.rs records the finding for the next unit. Refs #3269 * docs(rewards): revert dep add, name both blockers with evidence in port.rs Per L1 direction: an unused dig-rewards-coin dep with no consumer is inert weight and would want whichever version ships the reader (0.3.0+, PR#6 open against DIG-Network/dig-rewards-coin), not 0.2 -- so it's reverted here and belongs in the unit that actually consumes it. Expanded the port.rs module doc to name both blockers explicitly with what was read (state.rs:1-31, #3267, the open reader PR) and the negative grep that found no funder-ownership registry anywhere in the tree, plus why serving dig.listRewardDistributors through UnavailableChainPort was considered and rejected (false capability signal; the exact "dispatch surface with no function behind it" pattern dig-node#593 was the last PR allowed to land on). No RewardsChainPort adapter, no Node wiring, no dispatch arm -- all three reward methods stay -32601 pending #3267 and a funder-ownership registry (parallel tickets, both required). Refs #3269 * feat(rewards): durable funder-ownership registry (identity only) (#606) * feat(rewards): durable funder-ownership registry (identity only) Records WHICH reward distributors this node funds -- launcher id plus the store id when the funding act knew it -- and nothing else. No amount can be recorded: every money figure here is chain-derived and goes stale, and dig_ecosystem#3286's wrapping u64 share multiply means a figure crossing this boundary can already be wrong. Durable storage would make it permanent. Persistence mirrors rewards_claim::engine::ClaimEngine: an optional state directory (absent = inert, so tests and default builds need no disk), atomic write, and a corrupt record is never overwritten. The set is never cached on the registry -- every read re-reads the file -- so no transient state lives on the struct across calls (the engine's F16/F18 discipline). The read outcome is closed and distinguishes funds-nothing from every unknown: NotConfigured (no state dir / dir missing / nothing written yet), PersistedStateCorrupt and IoFailed. A corrupt record is quarantined by COPY and left in place, so the next read is corrupt too rather than decaying into an empty list -- SPEC 2.4 clause 1 in the place it costs most, since an empty dig.listRewardDistributors tells an operator it funds no distributors. Node carries it in a OnceLock slot with pub(crate) accessors, mirroring mirror_pointers and reward_prover_statuses. Nothing installs it in production yet: no dig-node code funds a distributor, and the startup wiring belongs to dig_ecosystem#3268, so the slot is marked the same way register_reward_prover_status is. Refs #3285 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(rewards): drop a duplicated funded_distributors initializer Two test-only `Node` literals got the slot twice (E0062), because the inserted line's own indentation made the wider-indented site match twice. Refs #3285 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(rewards): wire the peer claim loop onto a cadence driver from real startup (#3268) (#605) * feat(rewards): wire the peer claim loop onto a cadence driver from real startup (#3268) The peer reward-claim engine shipped complete and tested in #594 but INERT: nothing constructed it, so the 86400s cadence never fired while `rewards_claim.enabled` defaulted to `true` -- a config asserting a subsystem is on while nothing runs. `rewards_claim/driver.rs` is a SCHEDULER, not a chain adapter: it derives this node's own payout puzzle hash, loads `RewardsClaimConfig`, builds a `ClaimEngine` against the only production port that exists (`UnavailableClaimChainPort`, until #3249 lands a real one) and drives `run_cycle` every `cadence_seconds + jitter`, jitter drawn from the OS CSPRNG. `server.rs`'s `serve_with_shutdown` makes exactly one call into it, beside `self_heal::spawn_driver_if_service()`. `enabled = true` now means: a background task exists, drives a counted cycle per interval, and its outcome is readable in-process as a NAMED state. With `UnavailableClaimChainPort` every cycle honestly reports `ChainSourceUnavailable` -- the gap is loud instead of silent. Anti-silence: `ClaimLoopHandle` carries a monotonic `cycles_driven` counter alongside the status, because `Idle` before the first cycle is correct and honest, so status alone cannot tell "scheduler never fired" from "nothing was claimable". The gate takes an INJECTED handle rather than reading the process-wide singleton, so `ClaimDriverRefusal::{Disabled, ChainSyncDisabled, NoOperatorWallet}` and "spawned but never ticked" are four pairwise-distinct readings a test asserts in-process. Nothing goes on the wire: no RPC method, dispatch row, handler or OpenRPC entry. `ClaimStatus` stays off the wire until #3249's real adapter lets the status surface be re-derived against it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(rewards): silence the deliberately-ignored fake-port argument (#3268) `OneDistributorPort::own_entry` ignores the puzzle hash the engine passes in on purpose -- the fake always returns the entry keyed to `entry_keyed_to` so the ENGINE's own comparison is what decides claimable vs. refused. Named it `_payout_puzzle_hash` (clippy `-D unused-variables`) and moved the rationale onto the parameter, where the next reader meets it. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * test(rewards): close the untested joint between the claim gate and the drive loop (#3268) `decide_claim_driver` was tested and `drive` was tested, but the production body joining them -- load the config from the state dir, derive the engine, reach `drive` -- was exercised by nothing. That is the exact shape of #594, which shipped a complete, fully-tested and entirely inert claim engine: had this body returned early, built the engine wrong, or never reached `drive`, every test on this change would still have passed and a real node would still never claim. Split `run_claim_driver` on the same `load` / `load_from` pattern the config itself uses: `run_claim_driver_in(state_dir, own_payout_puzzle_hash, port, handle)` holds the whole body and is generic over the port, and `run_claim_driver` is reduced to the wallet-derivation adapter that cannot be reached from a test. Adds two tests through the real body: counted cycles from a written config (zero before the interval, exactly one per interval after), and `UnavailableClaimChainPort` reporting `ChainSourceUnavailable` by name on a driven cycle -- proving the production adapter path is reached, not only a fake. No behaviour change: same config, same engine construction, same port. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(rewards): settle before advancing, and keep the wrapped assertion out of rustfmt's reach (#3268) Two repairs to the new composition tests: - The `ChainSourceUnavailable` test advanced the paused clock before the spawned body had reached its first `sleep`, so the timer was not yet registered and the advance bought no cycle at all -- it read zero cycles, not a driven one. A `settle()` first, mirroring the counted-cycles test. - rustfmt rejoined a `\`-continued assertion message into one line, leaving 14 literal spaces mid-sentence and tripping the repo's own `continuation_guard`. `concat!` states the wrap explicitly, so no formatter pass can reintroduce the run. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * feat(rewards): emit a per-cycle event so the claim loop has a reader (#3268) The adversarial gate blocked #605 on this: the PR justified itself by making an inert subsystem loud, but nothing in the shipped binary could hear it. ClaimLoopHandle had no caller outside driver.rs tests, drive() emitted no event, and all three tracing calls fired only on paths where the loop does NOT run -- so on the default path (enabled=true, chain sync on) the observable output was identical to before the PR: silence. Today that silence covers a permanent ChainSourceUnavailable; after #3249 it would also cover Faulted, PersistedStateCorrupt and ClaimableButNotClaiming. log_cycle() now names the state and the cycle count after every cycle -- info for Nominal, warn for everything else, because "this peer is earning nothing and here is why" is a warning, not routine chatter. Tested by capturing the subscriber output rather than asserting the call site exists, since this ticket exists because a guarantee that cannot be observed in a running node is not a guarantee. Refs DIG-Network/dig_ecosystem#3268 Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(rewards): stop a u64::MAX jitter bound panicking the claim driver `OsJitter::jitter_seconds` computed `bound + 1` for its modulus. `jitter_seconds` comes from the node's persisted `rewards_claim` config and is not clamped, so a config carrying `u64::MAX` overflow-panicked inside the detached claim-driver task -- which has no restart and emits no further log output, so the claim loop would die silently for the rest of the process lifetime. `saturating_add(1)` keeps the draw within `0..=bound` for every input; the composed `next_interval_seconds` range is unchanged. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> * fix(rewards): sanitize the claim schedule so no config value silently disables the loop `next_interval_seconds` saturates instead of panicking, so a persisted `jitter_seconds = u64::MAX` no longer crashes the driver -- it schedules the next cycle ~585 billion years out. The claim loop then never fires again: no cycle, no `log_cycle` line, and a permanent, reassuring `0` cycle count. That is #594's inert-but-green shape reopened one level up, in the config file. `run_claim_driver_in` now sanitizes both schedule fields where it reads them, before either reaches the engine's fee window or `drive`: - `CLAIM_SCHEDULE_SECONDS_MAX = 31 * 24 * 60 * 60` (31 days) -- above every documented default (86,400s cadence, 3,600s jitter) and above "claim monthly", while excluding everything that means never. - out of range (or a zero cadence, which would busy-loop) substitutes the published default and emits `tracing::warn!` naming the field, the rejected value and the substituted one. Nothing is accepted silently. `config.rs` is untouched: it keeps reporting what is on disk. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com> * chore(release): v0.257.0 -- the reward distributor lifecycle starts running Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude <noreply@anthropic.com>
* chore(release): v0.256.0 -- reward distributor prover, peer claim loop, prover-status RPC (#602)
* feat(mirror): persist mirror-bond coin ids (#575)
* chore: open lane for #574
* feat(mirror): persist mirror-bond coin ids so a restart cannot double-create
Bond identity was reconstructed from a live chain scan on every read
(`mirror/observe.rs`), with no persistence of its own. A restart, a cold
replica, or a lagging/flaky chain source all rendered a real, unspent,
confirmed bond as "no bonds" -- and because the in-flight suppression is
keyed on pending/submitted audit entries, a bond whose create had already
CONFIRMED was not suppressed either, so the same short scan that emptied
the read surface also cleared the one thing that would have stopped a
second coin being paid for collateral that already exists (dig-node#574).
Persist the (store, root, epoch) -> coin_id mapping in the EXISTING spend
audit record (spend-audit.jsonl) rather than a new store: a mirror-coin
create already writes store_id + AuditedBond{root, epoch} + amount there,
and the coin id itself becomes durable the moment resolve_landed_spends
confirms it. This adds the one missing piece -- the advertised URL a
create carries -- and a read-side query, confirmed_mirror_bond, that
returns the newest CONFIRMED record naming a triple.
Chain stays authoritative. mirror::local_bond::recheck_missing_bonds
never trusts the record: for a held bond the live scan did not cover, it
asks the record for a candidate coin id, then re-verifies that SPECIFIC
coin against chain via the same independent check (chain_bond_verdict)
that verifies an untrusted peer's claimed bond. Only a fresh `Bonded`
verdict is folded back in, as covered; `Unbonded`/`Unverified` fall
through to an ordinary create, exactly as if no record existed.
Version: 0.254.86 (patch -- per #522 the MSI ProductVersion minor field
is exhausted and the counter lives in patch).
Co-Authored-By: Claude <noreply@anthropic.com>
* test(mirror): prove the recovery wiring end to end through PassRunner::run
Adds two integration-level tests over the REAL pass pipeline, not just the
isolated recheck_missing_bonds unit tests: a bond missing from the live
scan with a chain-reverified durable record is recovered (no double
create, correct Bonded state reported), and the control -- the same
record but chain disproves it -- correctly falls through to an ordinary
create. Together these are the concrete regression test for the
cold-start/lagging-chain-source double-create scenario the ticket asked
to have measured.
Also refactors in_flight_creates to take the already-folded SpendLedger
instead of re-reading the log itself, so PassRunner::run reads the audit
file once per pass and shares it with the new recovery step, and fixes a
doc comment on in_flight_creates that the recovery step would otherwise
have made stale on landing ("a Confirmed create has a coin the chain
observation already sees" is no longer unconditionally true).
Co-Authored-By: Claude <noreply@anthropic.com>
* chore(fmt): wrap long test signatures to satisfy rustfmt
Co-Authored-By: Claude <noreply@anthropic.com>
* chore(clippy): use slice::from_ref instead of cloning for a single-element slice
Co-Authored-By: Claude <noreply@anthropic.com>
* chore(release): bump to v0.254.89
Base branch moved to develop after PR #576 merged there at v0.254.88
(main and develop are currently identical), leaving this branch's
carried-forward .88 as a zero-increment against the new base. Bumped
to the next free integer after fetching and verifying both origin/main
and origin/develop tip at .88.
Co-Authored-By: Claude <noreply@anthropic.com>
---------
Co-authored-by: Claude <noreply@anthropic.com>
* fix(peer): count accepted relayed circuits in the connected pool (#579)
serve_accepted_relay_conn served every accepted relayed circuit (full mTLS
auth, full L7 peer RPC) while registering it nowhere, so connected_peers
under-reported every relayed inbound peer -- the relay-leg twin of the
direct-inbound defect #402/#523 already fixed.
adopt_inbound_peer_in_pool now dispatches by TraversalKind: Relayed routes to
dig-gossip's already-published adopt_relayed_inbound_handle (v0.32.0, the rev
this repo already pins), every other tier keeps the unchanged
adopt_direct_inbound_handle path. serve_accepted_relay_conn adopts before
serving and releases after, mirroring the direct listener exactly.
Refs: https://github.com/DIG-Network/dig_ecosystem/issues/3124
* fix(cli): guard the exit-code namespace shared with diga against collisions (#582)
* chore: open lane for #3189
* fix(cli): guard the exit-code namespace shared with diga against collisions
dign and diga deliberately share one process exit-code numbering (dig-app's
outcome.rs says so in its own doc comment), so a number is free only if it
is unoccupied ecosystem-wide. dig-node#407 assigned exit 7 to
NODE_UNREACHABLE by checking only this repo's own table, where 7 genuinely
was free -- and collided with diga's NOT_CONNECTED. A reviewer caught it by
hand; nothing failed automatically.
Adds scripts/check-exit-code-collisions.sh: parses both enums' code()/name()
match arms straight from their own source -- this repo's ExitCode, and a
live fetch of dig-app's outcome.rs at its default branch -- and fails if a
number carries two different names, or if either side draws a number from
the reserved shell signal range (126, 127, 128+N). Ships with an 18-case
hermetic test harness (scripts/tests/check-exit-code-collisions.test.sh)
covering the actual #407 collision shape, arm-order independence, arm-count
mismatch, the reserved-range boundary from both sides, the live-fetch path
itself, and fail-closed behaviour on an empty/missing/unreachable table.
Wires a real (unstubbed) invocation into ci.yml's existing "Release-script
tests" job so a collision introduced by a future PR, on either side, is a
red required check on that PR -- not a note a reviewer has to catch. The
fetch retries twice (2s backoff) since this becomes a required, network-
dependent check; a fetch failure still fails closed after retrying, never
silently passing as "diga has no codes".
Updates SPEC.md 8.4 to point at the mechanical guard instead of leaving
"re-check both tables" as unenforced prose, and records that the
extension's WALLET_WS_ERR.NOT_CONNECTED = -33001 is a separate JSON-RPC
error-code space, not a rival of this one. Adds a doc-comment to the
existing transcribed collision test pointing future readers at the live
script as the authoritative check; the transcription remains as a narrower,
hermetic regression pin for the #407 shape specifically.
No renumbering: every currently-assigned code is unchanged.
Refs #3189
Co-Authored-By: Claude <noreply@anthropic.com>
---------
Co-authored-by: Claude <noreply@anthropic.com>
* fix(hygiene): port the lost-continuation guard to 4 crates, fix 48 corrupted strings (#3190) (#583)
* chore: open lane for #3190
* fix(hygiene): port the lost-continuation guard to 4 crates, fix 48 corrupted strings
Replicates dig-node-service::continuation_guard (dig-node#526/#501) into dig-node-core,
dig-wallet, dig-runtime and dig-chat-protocol, line-for-line apart from crate-specific
constants -- ported rather than reinvented, per dig_ecosystem#3190.
Wiring the guard in surfaced 48 pre-existing lost-continuation defects the ticket's own
"no measured corruption in these four crates" note did not anticipate: 36 in dig-node-core,
12 in dig-wallet, mostly test-assertion prose where a multi-line message lost its `\`
continuation and shipped the source's own indentation as a mid-sentence space run (one
as the worse `\n`-plus-indentation variant). All 48 are collapsed to the single space the
sentence always meant, with surrounding indentation and wording otherwise untouched.
Two lines are real column-alignment, not defects, and get a targeted EXCLUDED_LINE_RANGES
entry on dig-node-core instead of a rewrite: download.rs's `claimed(...)` fixture-table
trailing comments, and net.rs's `label : value` debug-print alignment.
Refs https://github.com/DIG-Network/dig_ecosystem/issues/3190
Refs https://github.com/DIG-Network/dig_ecosystem/issues/3130
Co-Authored-By: Claude <noreply@anthropic.com>
---------
Co-authored-by: Claude <noreply@anthropic.com>
* feat(mirror): detect an IP change daily and reconcile mirror coins to the current advertise URL
Automatic half (D1-D4) of the daily mirror-URL reconcile: derived personal-day offset, two-observation hysteresis, nine ordered gates with K sized as a self-funding prefix before any reclaim, `submitted` never `completed`, audit lines gain `reclaim_reason` + `trigger`.
Gates at b4c09866: loop-reviewer PASS (review 5129272285), loop-security PASS @ 175304e1 (tree byte-identical, `git diff 175304e1 b4c09866` empty), adversarial loop-decider PASS (comment 5567083644; SHOULD-FIX findings ticketed separately).
Refs DIG-Network/dig-node#570
Refs DIG-Network/dig_ecosystem#3203
* feat(serve): content hosting + serve path batch, v0.255.0 (dig_ecosystem#3212)
Nine commits from the #3212 serve-path lane, gated at 899cc68f (reviewer review 5130425808,
security comment 5568662836), plus the single semver bump to 0.255.0 for the develop -> main batch.
- store_id/root case normalised at the CapsuleKey boundary; cache delete targets the matched entry
- tier-0 occupancy reads the eviction-aware ledger
- profile-sync outbound budget in bytes; announcer asked first
- melt confirmation depth on the terminal spend, fail-closed
- EngineWarming (-32002) while the peer tier attaches, never -32004
- window completeness derived from the bytes read
- deps: dig-stun 0.2, chia-query 0.24.3, dig-nat 0.21.2, dig-logging 0.2.2
Refs DIG-Network/dig_ecosystem#3212
* chore: untrack gitnexus-generated agent files (#590)
* chore: untrack gitnexus-generated agent files
These files were generated by `gitnexus analyze` as a side effect of
indexing this repository. They are development-loop private tooling
output, not product code, and carry no secrets. They are removed from
tracking going forward via .gitignore; history is deliberately NOT
rewritten.
Refs #3177
* chore: drop private-repo reference from gitignore comment
The ignore comment named a private repository and an internal issue
number in a public file, which is the same disclosure class this
change set exists to remove; the reference is dropped and the
guidance kept.
* feat(rewards): always-on prover loop engine -- honest liveness, type-enforced self-exclusion, bounded spend (#593)
The always-on reward-prover engine: ~2,000 lines under
`crates/dig-node-core/src/rewards/`, built against the merged `dig-rewards-coin`
SPEC. Library only -- nothing spawns it, and the sole production
`RewardsChainPort` refuses every call, so it cannot spend. Wiring the composed
system is #3265, which carries its own gate.
The epic's premise -- "anytime the process isn't running, rewards are not being
distributed" -- is half wrong, and the false half is the dangerous one. `Sync`,
`NewEpoch` and `InitiatePayout` need no manager authority, so funder downtime does
not stop rewards: it FREEZES THE ENTRY SET while accrual and payouts continue.
Peers that stopped mirroring keep earning; peers that started cannot begin. That
shaped the whole design.
Liveness honesty (SPEC 2.4). The status record carries no `healthy`/`ok`/`up`
boolean and no precomputed staleness, because a wedged loop cannot report its own
wedging -- whatever it last wrote stays there, so a writer-set flag reads true
forever after the failure it exists to reveal. The reader derives staleness from
`last_cycle_completed_at` against `observed_at` and its own clock. A recursive
JSON-key test enforces the absence at every nesting depth; asserting on keys and
never substrings, since `ProverState::Running` legitimately serializes the VALUE
"running". The one legal staleness signal is chain-derived (SPEC 12.4: 48 hours
AND a non-zero reserve, from the singleton's own spend history) and lives on the
distributor read, where a wedged prover cannot fake it.
Self-exclusion is a compile error, not a habit. dig-node#261's lesson is that an
invariant enforced on some paths is not an invariant. `admit` is the single
admission point, checks both SPEC 5.2 coordinates (own peer_id OR a payout puzzle
hash this wallet controls), and mints an `AdmittedPeer` with private fields and no
public constructor -- so `EntryAction::Add` cannot be built by a path that skipped
admission.
A prover's own fault can never strike a peer. `GateError` is a distinct type from
`GateIneligibleReason` and `record_prover_fault` takes `&self`, so SPEC 3.6 clause
4 is enforced by the borrow checker rather than by comment. Without that, a
misconfigured operator -- one missing mirror-collateral epoch ordinal -- would
strike every peer at once and evict its entire 250-entry set in three hours, each
eviction a fee it pays plus a settlement out of its own reserve.
The money bounds are stated where a human reads them (`rewards/mod.rs`): 24
bundles/day, 192 entry actions/day, a fee ceiling of 24x the configured standard
fee, 192 removals/day worst case with 96/day sustained churn. Recorded honestly:
SPEC 6.3's rate bound and fee ceiling are ONE control, not two.
Three gate rounds, every leg fresh-context. Round 3 at this head: reviewer PASS,
adversarial decider RATIFY (leg closed), security CHANGES-REQUIRED on a finding
the decider ratified deliberately -- adjudicated in
https://github.com/DIG-Network/dig-node/pull/593#issuecomment-5601784815 and
carried to #3265 with the remedy corrected, because the proposed fix would have
persisted a poison flag to the very store whose writes were failing.
Found and fixed under gate: a census ordinal off by one in both directions (SPEC
4.6 requires n-1 exactly); an unreachable grace window leaving a named constant
with no reader; a missing `NewEpoch` spend; an absent-ordinal path attributing a
prover fault to peers; and a daily fee ceiling 24x too high because a per-bundle
fee was consumed as a daily ceiling.
Refs DIG-Network/dig_ecosystem#3250
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat: serve dig.getRewardProverStatus at Tier::Control (#595)
* chore: open lane for #3269
Bump dig-rpc-protocol to 0.11.0 (adds dig.getRewardProverStatus and
the other reward RPC methods to the wire).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* test(rewards): fail-closed Reward-tier guard + dig-rpc-protocol 0.11 line assertion
- dependency_tree.rs: assert the resolved dig-rpc-protocol line is 0.11 (was 0.10);
documents the known-red two-version state pending the dig-peer 0.14.0 /
dig-download 0.23.0 cascade (#3269).
- reward_methods_tier_guard.rs: fail-closed guard over the live Method::ALL
catalogue -- every Reward-named method must be Tier::Control and not
peer-reachable, so a fifth reward method added later is caught at the wrong
tier automatically rather than inheriting a wrong default (binds #3261's rule
node-side).
- peer.rs: sibling unit test exercising the real (pub(crate)) is_peer_reachable_method,
since an external integration test cannot see it -- same guard, executed against
this node's own allowlist rather than only the shared crate's.
Refs #3269
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* style: remove trailing blank line in reward_methods_tier_guard.rs
* feat(rpc): serve dig.getRewardProverStatus at Tier::Control
Adds the missing handler for PR#595: a new reward_prover_statuses
registry + accessors on Node (empty until #3265 spawns a prover loop,
so the registry read is real, not a stub), a dispatch.rs arm inside
the Method enum match (never the string pre-match), and a
field-for-field mapping from dig-node-core's internal
rewards::state::RewardProverStatus (camelCase-tagged) onto
dig-rpc-protocol 0.11's wire RewardProverStatus (snake_case-tagged
struct, camelCase-tagged ProverState value), widening entry_count
u32 -> u64 explicitly and hex-encoding the three [u8; 32] identity
fields.
An all-zero launcher_id (what an uninitialised registry slot
hex-encodes to) is omitted at this boundary rather than rendered as
a real distributor with a plausible-looking id -- the money-hole
class the dig-rewards-coin driver's adversarial gates found three
times.
Tests (in dig-node-core::lib.rs's existing test module, where the
pub(crate) registry accessors are visible) drive the real dispatch
entry point (handle_rpc -> RpcDispatch::dispatch -> the Method arm)
and assert field-for-field on the serialized JSON body: populated
registry, empty registry (-> {"statuses": []}), zero-id omission,
tier/peer-reachability, enum-match-not-string-prematch, and
launcher_id filtering. The no-health-boolean / no-staleness
assertion is by key set, not substring.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* style: rustfmt the reward-prover-status registry + tests
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* chore(deps): bump dig-download 0.23, dig-peer 0.14, dig-rpc-protocol 0.11
Closes the two-versions-of-dig-rpc-protocol split (#3269): dig-download 0.23.0
and dig-peer 0.14.0 both now resolve dig-rpc-protocol ^0.11, matching
dig-node-core's own dig-rpc-protocol = "0.11.0" line, and dig-node-service's
two 0.10 lines (main dep + dev-dependency restatement for
openrpc_drift_guard.rs) move to 0.11 to match.
Manifests only. cargo update / Cargo.lock intentionally NOT run yet: a fourth
capper, dig-peer-selector ("0.11" in dig-node-core/Cargo.toml), still requires
dig-peer = "^0.13" in every published version through 0.11.1, so the tree
cannot fully resolve until a dig-peer-selector release picks up dig-peer 0.14.
CI will stay red on this commit for that reason, which is expected.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* feat(rpc): close dig-rpc-protocol 0.11 cascade + attribute payout figures
Bump dig-peer-selector 0.11 -> 0.12 (the release that moves onto dig-peer
^0.14) and run cargo update, closing the two-versions-of-dig-rpc-protocol
split: Cargo.lock now resolves exactly one dig-rpc-protocol, at 0.11.0,
alongside dig-peer 0.14.0, dig-download 0.23.0, dig-peer-selector 0.12.0.
Add a subject-attribution test and doc comments to
reward_prover_status_to_wire: total_paid_out_base_units and
reserve_base_units are per-distributor totals (this distributor's payout to
ALL its mirrors, and this distributor's own reserve), never the querying
node's own earnings and never summed/cross-attributed across distributors.
This is the defect class a sibling adversarial gate found in dig-app#403's
rewards pane, which rendered a distributor total as one mirror operator's
personal earnings and overstated by up to 250x.
Checked: rewards/state.rs, rewards/port.rs and rewards/mod.rs contain no
Eligible/verdict/payout_hash symbol, so nothing in this mapping surfaces a
persisted EligiblePayoutHash verdict.
Refs #3269
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* fix(rpc): silence dead_code on register_reward_prover_status pending #3265
Clippy's non-test lib target has no production caller for
register_reward_prover_status yet, because #3265 (the always-on prover loop
that would call it from bring-up) has not landed -- only tests call it today.
cfg_attr(not(test), allow(dead_code)) stands in for that missing caller
until #3265 wires a real one.
Refs #3269
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* fix(rpc): make the all-zero identity guard non-silent and cover all three fields
Security (blocking) and the adversarial leg both found the same defect in the
zero-launcher_id filter: it checked only launcher_id, so a registration bug
that zeroed store_id or root beside a valid launcher_id would pass through as
a plausible record, and dropping the bad record silently destroyed the
evidence a registration bug happened at all -- SPEC Sec2.4 clause 1's exact
prohibition.
zeroed_identity_fields() now checks launcher_id, store_id AND root. The
dispatch filter still excludes a record with any zeroed field (never renders
an uninitialised slot as a real distributor), but first fires a
tracing::warn! naming which field(s) were zero, so a bad registration is
observable rather than swallowed. Kept isolated in dispatch.rs rather than
woven into the wire mapping, since this belongs at #3265's writer once that
lands.
Replaced get_reward_prover_status_omits_an_all_zero_launcher_id (which
proved the omission but not the observability, and never exercised a zeroed
store_id/root beside a valid launcher_id) with
get_reward_prover_status_logs_and_excludes_a_zeroed_identity_field, covering
both a zeroed launcher_id and a zeroed store_id beside a valid launcher_id,
and asserting the tracing::warn! output via the crate's existing
capture_sync_logs test utility.
Fixed a now-false "Known-red" doc comment on
tests/dependency_tree.rs::the_workspace_carries_exactly_one_module_wire_crate:
the dig-peer 0.14.0 / dig-download 0.23.0 / dig-peer-selector 0.12.0 cascade
already landed in this PR's Cargo.toml/Cargo.lock, so the assertion is green,
not red. Assertion itself untouched -- still exact-version.
Refs #3269
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* fix(rpc): correct a born-false "shipped dig-app 15.5.0" doc claim
dig-app's latest release is v15.4.0 -- there is no v15.5.0 tag -- and
dig-app#403 (the pane that would consume dig.getRewardProverStatus) is OPEN
and unmerged. Point the doc comment at the real, unmerged consumer instead
so a future reader doesn't take this as evidence a shipped consumer depends
on the guard, which would wrongly discourage relocating it to #3265's writer.
Refs #3269
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* fix(rpc): correct doc placement, assert the zeroed field by name, treat root as an observation
Three findings from the correctness gate on PR#595 at 134864a9.
1. The zeroed-identity helper's doc block was spliced onto the end of
reward_prover_status_to_wire's block with no separator, so the wire-mapping
rationale documented a boolean predicate and the mapping function was left
with no doc at all. Each doc block now sits above the item it describes.
2. The log assertion `logs.contains("launcher_id")` was a tautology: the warn
emits launcher_id as a structured field on every fire, so the property the
guard exists to add -- naming which field was zeroed -- was unasserted.
Deleting `zeroed_fields = ?zeroed` from the warn left every assertion green.
The test now asserts the zeroed_fields value itself, which the fixture makes
exact and disjoint across cases.
3. `root` is an observation, not an identity. A registered prover that has not
completed its first cycle plausibly has no root, and a writer that zero-inits
it would have made a healthy prover invisible. A zeroed launcher_id or
store_id still excludes the record; a zeroed root alone warns and returns.
Refs #3269
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* fix(rpc): restore zeroed_fields structured field dropped from the pushed warn
The previous commit (080be3df) landed with `zeroed_fields = ?zeroed` missing
from the tracing::warn! call in the GetRewardProverStatus filter -- a
one-line regression introduced while proving the new log assertion goes red
without it, never restored before the commit was made. Without this field
the log line never names WHICH field was zero, so an operator sees only
that something was excluded, and the test asserting `zeroed_fields=[...]`
per case would fail. Restored; all 7 reward-prover-status tests green.
Refs #3269
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* fix(rpc): split zeroed-field logging by level -- WARN for a missing
identity, DEBUG for a zeroed root
A zeroed launcher_id or store_id is a real registration bug: the record is
excluded and now logs at WARN, naming the exact field(s) via
`zeroed_fields=[...]`. A zeroed root alone is an ordinary pre-first-cycle
state, not a fault: the record is still returned, and now logs at DEBUG
instead of WARN, so an operator polling this endpoint sees warn-level
volume proportional to real registration bugs, not to every
not-yet-cycled prover on every poll.
Updated the doc comments on `zeroed_fields`, the dispatch filter and the
test to describe the level split, and extended the regression test to
assert on level (WARN vs DEBUG) as well as the `zeroed_fields` value.
Proved both directions: flipping the DEBUG branch back to WARN turns the
test red on the level assertion; flipping the field-name assertion back to
a bare `contains("launcher_id")` would have passed unconditionally (the
prior tautology) and is no longer possible since the assertions now pin
`zeroed_fields=[...]` plus the level string.
Refs #3269
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* feat(rewards): peer-side claim loop -- watch distributors, claim on cadence (#3251) (#594)
* feat(rewards): peer claim loop skeleton -- discovery, cadence, config, claim port
* test(rewards): write all twelve acceptance tests for the peer claim loop
* feat(rewards): wire the seven rewards_claim submodules into the crate
mod.rs declared no submodules, so types.rs/port.rs/config.rs/cadence.rs/
parser.rs/engine.rs/hints.rs (1,435 lines, 25 tests) were never part of the
crate and never compiled. Declare them and re-export the public surface.
* style(rewards): cargo fmt the rewards_claim submodules
* chore(deps): bump dig-node-control-interface 0.33->0.35, dig-rpc-protocol 0.10->0.11.0
dig-rpc-protocol 0.11.0 is merged and tagged upstream; the other dig-*/chia-*
deps of dig-node-service were already at the latest permitted-by-caret version
in Cargo.lock. crates/dig-node-core/Cargo.toml is untouched (#3250's file set).
* chore(deps): revert dig-node-control-interface and dig-rpc-protocol bumps
Both create a duplicate-version split in this PR's scope and neither can be
closed without editing a sibling crate's manifest this lane does not own:
- dig-rpc-protocol 0.11.0 duplicates against dig-node-core/Cargo.toml:194
("0.10.2"), which is #3250's live file set (dig-node#593).
- dig-node-control-interface 0.35.0 duplicates against
dig-wallet/Cargo.toml:81 ("0.33"), a sibling crate this lane does not own;
the observed Clippy break (BalanceAsset/Asset type-identity mismatch,
missing url_reconcile/url_current/urls fields) came from THIS duplicate,
not from dig-rpc-protocol.
Both belong to their own sequenced dep-bump unit of work, not this ticket.
* fix(rewards-claim): fault laundering, permanent no-entry blacklist, fee ceiling magnitude
Three independent gates on dig-node#594 (51516e62) found four logic defects; this
addresses A, B and C per the corrected fix brief (D is documented only, not fixed
here per the brief's own instruction).
Defect A -- the anti-silence surface laundered every real fault into `Nominal`:
- A1: `fault_reported` had no fault-bearing ClaimLoopState to fall through to, so a
chain adapter erroring every cycle read `Nominal` forever. Added
`ClaimLoopState::Faulted { cycles }`, outranking Nominal/ClaimableButNotClaiming,
under ChainSourceUnavailable.
- A2: inverted the test that asserted A1's bug as correct behaviour.
- A3: `ClaimableButNotClaiming` compared a per-cycle snapshot
(`distributors_claimable`) against a lifetime-cumulative counter
(`claims_submitted`), so it latched healthy forever after one lifetime success.
Added `claims_submitted_this_cycle` (per-cycle) as the correct comparand; kept
`claims_submitted` as a cumulative counter.
- A4: `last_discovery_at`/`last_cycle_at` were stamped even on a failed discovery
or an all-faulted cycle, destroying the staleness signal a reader depends on.
Now only stamped on success; added `last_attempt_at` to prove liveness
separately. `fault_reported` and `distributors_faulted` now reset per cycle
instead of latching for the process's lifetime.
Defect B -- "terminal, stop retrying" was implemented as a process-lifetime
blacklist (`terminal_no_entry: HashSet<Bytes32>`, never cleared). That blocked
SPEC 12.5 clause 2's re-entry path (evicted, re-challenged, re-admitted never
claims again) and permanently punished a peer that discovered a distributor
before the funder's AddEntry landed. Removed the blacklist entirely -- `own_entry`
is a cheap chain read, re-issued every cycle for every candidate, matching clause
3's "never cache across cycles". `NoEntrySlot` is now a per-cycle observation, not
a lifetime sentence.
Defect C -- the fee ceiling didn't bind anything and there was no aggregate cap:
- C1: default `CLAIM_FEE_CEILING_MOJOS_DEFAULT` lowered from 1_000_000_000
(transplanted from `MIRROR_SPEND_FEE_CEILING_MOJOS`, sized for a mirror-coin
spend) to 200_000 -- 2x the observed routine Chia fee range (5,000-100,000
mojos), so it actually binds instead of leaving 4-5 orders of magnitude of
slack.
- C2: added a per-cycle aggregate fee budget
(`max_cycle_fee_budget_mojos`, default 10x the per-claim ceiling) checked
across all claims in a cycle, closing the attacker-cost gap where funding K
distributors could force a victim to spend K x the per-claim ceiling per cycle.
New `ClaimOutcome::SkippedCycleBudgetExhausted`.
Tests: rewards_claim test count 26 -> 37 (11 new: repeated_discovery_faults_never_
read_as_nominal, failed_discovery_leaves_last_discovery_at_unchanged, a_reported_
fault_surfaces_as_faulted_not_nominal, a_lifetime_submission_does_not_mask_a_
later_cycle_that_submits_nothing, no_entry_slot_then_re_admitted_produces_a_claim_
on_the_later_cycle, distributors_each_under_ceiling_do_not_collectively_exceed_
the_cycle_budget, the_default_per_claim_ceiling_actually_binds_a_routine_fee, plus
renamed/rewritten no_entry_slot_is_non_terminal_and_re_checked_every_cycle).
Refs #3251
* fix(rewards-claim): refuse a claim entry for the wrong payout puzzle hash
CI fix: cadence.rs's RewardsClaimConfig literal was missing the
max_cycle_fee_budget_mojos field added in the previous commit (E0063,
caught by CI's Clippy/Test jobs -- the local cargo check for this
workspace is too slow to use as the compiler here).
Defect E (security-gate finding, folded in before this pass closes):
submit_initiate_payout was called with entry.payout_puzzle_hash -- whatever
the chain port handed back -- with no check against this node's own
own_payout_puzzle_hash. UnavailableClaimChainPort is the only production
adapter today so nothing can exploit this yet, but the whole point of the
ClaimChainPort seam is that #3249 swaps in a real adapter with nothing
above it changing, so deferring this would ship the landmine live with no
review pass watching for it. Added an equality guard before the spend:
a mismatch refuses to submit, counts
(ClaimStatus::claims_refused_payout_mismatch), surfaces its own named
outcome (ClaimOutcome::PayoutPuzzleHashMismatch), and is reported as a
fault (a divergent entry means the port is confused or hostile, not that
there is nothing to claim) -- never corrected by substituting our own
hash and proceeding.
Defect D: documented, not wired, per instruction -- added the "not yet
wired into node startup" paragraph to mod.rs's module doc (the PR body
carries the same paragraph) so the next reader arrives at the caveat in
the code, not only in a merged PR description.
Refs #3251
* fix(rewards-claim): B1 -- ClaimableButNotClaiming is a magnitude comparison, not a zero-test
submitted_this_cycle < claimable_this_cycle now fires the anti-silence state, carrying
the shortfall as ClaimableButNotClaiming { claimable, submitted }. The previous
submitted_this_cycle == 0 zero-test let one submission mask any number of same-cycle
skips (claimable=10, submitted=1 read Nominal).
Also folds in B3's precedence fix (ChainSourceUnavailable > Faulted >
ClaimableButNotClaiming > Idle > Nominal) and the per-distributor
payout_hash_mismatches_this_cycle counter so a per-distributor fault can no longer
pin the cycle-wide Faulted state, plus R2's rename of terminal_no_entry_slot to
no_entry_slot_this_cycle now that it is no longer terminal.
* fix(rewards-claim): B2/B3 -- value-ordered budget with rotation, per-distributor fault isolation
B2: run_cycle now splits into a pre-budget phase (asset/entry/hash/threshold checks,
producing the claimable set) and a budget phase, ordering the claimable set by accrued
value descending before applying the fee ceiling and cycle budget. Dust distributors
(low accrued value regardless of attacker-controlled fee) now sort last and are the
ones the budget drops, closing the claim-suppression attack where ten high-fee dust
distributors could consume the whole cycle budget ahead of a victim's real earnings.
A rotation_cursor tie-breaks only WITHIN equal-accrued-value tiers so a genuinely
tied honest tail that exceeds one cycle's budget every cycle still rotates through
and is eventually served, rather than dropping the same tail forever.
B3: the payout-hash mismatch check in evaluate_pre_budget now increments the
per-distributor payout_hash_mismatches_this_cycle counter instead of setting
fault_reported, so one hostile or buggy entry can no longer pin the cycle-wide
Faulted state and bury ClaimableButNotClaiming for every other healthy distributor.
R2: terminal_no_entry_slot -> no_entry_slot_this_cycle throughout.
* fix(rewards-claim): R5 -- document not-yet-wired loop; persist B2 rotation cursor
An operator reading their own rewards-claim.json and seeing enabled: true has no way
to know from that file alone that no startup path constructs a ClaimEngine yet
(#3268) -- mod.rs said so, but a config file reader does not arrive at a module doc.
Also gives RewardsClaimConfig a rotation_cursor: Option<Bytes32> field so B2's
tie-break cursor survives a save/load round-trip -- an in-memory-only cursor resets
on every restart, which would starve a legitimately tied honest tail forever on any
node that restarts daily.
* fix(rewards-claim): clippy collapsible-match + SPEC v0.1.3 wording refresh
Collapses the nested if into the outer match arm in run_cycle (clippy::collapsible_match).
Also refreshes NoEntrySlot / no_entry_slot_this_cycle doc comments now that
dig-rewards-coin v0.1.3's SPEC §12.5 amendment is merged and tagged: absence is
terminal for one claim attempt only, never for the distributor, must not be cached,
and must not accumulate into a permanent exclusion set -- confirming rather than
diverging from the re-read-every-cycle behaviour already implemented.
* fix(rewards-claim): add missing rotation_cursor field in cadence.rs test literal
Struct literal in the cadence test module was not updated when RewardsClaimConfig
gained rotation_cursor (R5 commit) -- CI's Clippy/Test jobs caught the missing field
(E0063) that a local cargo check could not (killed by memory pressure before this
workspace-wide build completed).
* fix(rewards-claim): F1 -- ChainSourceUnavailable is per-cycle, never a latch
compute_state() compared against self.state -- last cycle's OWN computed
output -- so once any cycle took an Unavailable port path, every later
cycle re-asserted ChainSourceUnavailable forever, even after the chain
came back and real claims were submitting. A node still syncing, or one
dropped connection, was enough to trip this permanently.
Add ClaimStatus::chain_unavailable_this_cycle, reset to false at the top
of every run_cycle and set true only on a cycle that actually took the
Unavailable path; compute_state now reads that flag instead of
self.state, so the reading is live again.
Test: a_transient_unavailable_cycle_does_not_latch_state_for_the_rest_of_the_process
(engine.rs) drives cycle 1 unavailable, cycle 2 healthy with a
submission, and asserts cycle 2 reads Nominal. Plus a compute_state-level
regression in types.rs.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* fix(rewards-claim): F3/F4/F5 -- fix stale counters, dedup candidates, correct version doc
F3: reset EVERY per-cycle counter (distributors_known/with_own_entry/
claimable/faulted, claims_submitted_this_cycle, no_entry_slot_this_cycle)
at the TOP of run_cycle, before any early return. The three
ChainUnavailable early-return paths skip the end-of-function assignment
block entirely, so a cycle that hit one used to leave the PRIOR cycle's
counts sitting on self.status while last_attempt_at stamped fresh for
THIS cycle -- a stale count under a fresh timestamp, exactly what SPEC
§2.4's staleness reasoning forbids. types.rs's doc sentence for
no_entry_slot_this_cycle now correctly says it is dated by
last_attempt_at (the field stamped unconditionally every cycle), not
last_cycle_at.
F4: dedup `candidates` by launcher id before phase 2. A real adapter
scanning §1.3 launch comments across every (store_id, root) this node
mirrors can plausibly return the same launcher id twice; without dedup
phase 2 would evaluate it twice and submit InitiatePayout twice against
one entry slot in one cycle -- the second spend is invalid but the fee
is paid anyway.
F5: dig-rewards-coin is v0.1.3, published on crates.io -- correct the
stale "v0.1.1" module-doc claim.
Tests: a_chain_unavailable_cycle_does_not_leave_prior_cycles_counters_stale
(F3), a_duplicated_launcher_id_submits_exactly_once (F4).
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* fix(rewards_claim): fold payout-hash mismatches into the shortfall predicate (F2)
A payout-hash mismatch never enters the eligible set, so it was counted in
NEITHER distributors_claimable NOR claims_submitted_this_cycle -- the
shortfall lived in neither term of compute_state's magnitude comparison.
All-K-distributors mismatching therefore read Nominal (falsely healthy).
Fold payout_hash_mismatches_this_cycle into the comparison's denominator:
submitted < claimable + mismatches. The result is ClaimableButNotClaiming
(a shortfall), never the cycle-wide Faulted -- Defect B3 stays fixed.
Inverts the assertion at what was engine.rs:1305
(a_payout_mismatch_never_sets_the_cycle_wide_fault_or_masks_other_distributors):
it previously asserted ClaimLoopState::Nominal across three cycles of an
ongoing mismatch, which pinned the defect as intended behaviour (an
A2-class test). It now asserts ClaimableButNotClaiming { claimable: 1,
submitted: 1 }.
Adds all_distributors_mismatching_is_a_shortfall_not_nominal, covering the
brief's exact "what if every distributor refuses for the same reason" case.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* fix(rewards-claim): F1/F3 -- AtomicU32 in fake ports, not Mutex, keeps discovery Send
CI's Clippy job (the compiler for this crate, per brief) caught it: holding
a std::sync::MutexGuard across the .await in FlakyThenHealthyPort and
HealthyThenUnavailablePort's discover_distributors made the returned future
not Send, which #[async_trait]'s generated trait signature requires.
Neither fake needs a lock -- each holds one call counter, incremented once
per call, never read-modify-written across an await point. AtomicU32's
fetch_add removes the guard (and the Send bound violation) entirely.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* fix(rewards-claim): F7 -- persist the fee window and cadence gate across restart
The per-cycle aggregate fee budget and the 24h cadence clock both lived only
in memory: `spent_this_cycle_mojos` was a `run_cycle` local and nothing on
disk recorded a completed cycle. Every fresh process got a full
`max_cycle_fee_budget_mojos` and an empty cadence clock, so a node stuck in
a crash-restart loop could spend unbounded XCH on fees, one full budget per
restart.
Adds three `#[serde(default)]` fields to `RewardsClaimConfig`
(`fee_window_start_unix`, `fee_spent_in_window_mojos`,
`last_cycle_completed_at`) and a new opt-in `ClaimEngine::with_persisted_fee_
window(dir, cadence_seconds)` that:
- restores the window/cadence state from `dir` at construction,
- refuses to start a cycle until the cadence has elapsed since the last
completed one,
- rolls a fresh budget window only once the cadence has elapsed since it
opened, otherwise keeps enforcing the budget against the persisted spend,
- persists the spend BEFORE every chain submission (write-then-spend), never
batched to cycle end, and persists the completed-cycle timestamp when a
cycle finishes.
Engines that never call `with_persisted_fee_window` (every pre-F7 test) are
unaffected -- this is additive, opt-in state beside the existing rotation
cursor, not a change to B2's value-ordering or rotation mechanism.
`ClaimStatus`'s own counters stay in-memory on purpose (observability, meant
to reset on restart); only the spend bound and the cadence gate persist.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* fix(rewards-claim): F7 -- update cadence.rs test literal for new persisted fields
The three new persisted RewardsClaimConfig fields (fee_window_start_unix,
fee_spent_in_window_mojos, last_cycle_completed_at) broke this crate's only
remaining full struct literal outside config.rs/engine.rs's own test
modules -- E0063 missing fields, caught by CI's Clippy job. Switched to
..RewardsClaimConfig::default() so the next added field cannot break this
literal again, the same fix already applied once before for rotation_cursor.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* fix(rewards-claim): F8/F14 -- atomic state write, fail closed on a corrupt window
Salvaged from a lane killed by a weekly cap before it could commit. Uncompiled at
commit time; CI is the compile signal.
Covers the fourth gate pass findings on the F7 persisted spend bound:
- F8: RewardsClaimConfig::save_to now writes atomically (temp file + rename in the
same directory), reusing the pattern already used by mirror/reconcile_state.rs
for the same class of state. load_from distinguishes an ABSENT file (clean first
run, defaults are correct) from a PRESENT but unparsable one, which fails CLOSED:
the window is treated as fully spent and nothing is submitted. Never Default, and
never a silent clamp downward, which would hand back the budget the corruption
was hiding.
- F14: the budget comparison uses saturating arithmetic so a corrupt disk-seeded
fee_spent_in_window_mojos cannot panic under the release profile's
overflow-checks.
- F9/F10/F12/F13 in progress in the same files.
Refs #3251
* fix(rewards-claim): negate with ! rather than the unimported Not trait
Co-Authored-By: Claude Haiku 4.5 <noreply@anthropic.com>
* fix(rewards-claim): F13 -- correct stale test literals to the folded shortfall
compute_state (types.rs) already reported the folded shortfall
denominator (distributors_claimable + payout_hash_mismatches_this_cycle)
as `claimable` -- that part of F13 landed in f478516a. The two engine.rs
tests asserting this state were written against the pre-fold, un-folded
numbers and never updated, so CI showed the implementation producing the
correct folded value (`claimable: 2`, `claimable: 1`) while the test
literals still expected the stale un-folded one (`claimable: 1`,
`claimable: 0`).
Update both literals -- and the comments describing them -- to the
folded values the F13 fix actually produces. No production code change;
compute_state's predicate and payload were already correct.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* feat(rewards-claim): add ClaimOutcome::Faulted variant
Add the seventh ClaimOutcome variant: the type could only say a peer was
legitimately not paid, never that a chain call failed. Carries the launcher
id, a bounded (200 char) copy of the chain port's error text, and whether a
pre-committed fee was reversed, so a reader can tell no money moved.
Engine wiring at the two fault arms (engine.rs:332, :377) follows in the
next commit.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* fix(rewards-claim): both fault arms now push ClaimOutcome::Faulted
engine.rs:332 and :377 used to increment `faulted` and discard the
outcome, leaving a definitively-failed claim absent from the outcome
stream -- indistinguishable from a cycle that never touched that
distributor. Both PreBudgetResult::Fault and BudgetPhaseResult::Fault
now carry the chain port's (bounded) error text, and the
submit_initiate_payout failure path also carries the fee it reversed,
so a reader can tell no money moved. The counter stays; it is not a
substitute for the outcome.
7 call sites needed updating: 3 PreBudgetResult::Fault constructions
(reserve_asset_id, own_entry, payout_threshold), 2 BudgetPhaseResult::Fault
constructions (required_fee_mojos, submit_initiate_payout), and the 2
consuming match arms -- exactly the set that was silently discarding a
failure before this change.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* test(rewards-claim): a failed submission produces a Faulted outcome
Regression for the rework: reuses F12's fixture (a submission that
definitely never broadcast) to prove both facts from one cycle -- the
outcome exists and carries the reversed fee, and the persisted window
still reflects zero net spend. Also fixes a rustfmt diff on the
PreBudgetResult::Fault variant Clippy's Rustfmt job flagged.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* fix(rewards-claim): delete fee_window_poisoned, stop latching self-healing state
Finding 1 (dig-node#594 pass 6): a future-dated clock is self-healing by
construction (`t > now` goes false the moment real time passes it), but the
engine ORed it into `self.fee_window_poisoned` and set that field `true`
permanently -- an RTC glitch or VM resume froze the claim loop forever instead
of until the skew passed. This is the third instance of one mechanism (pass 3
latched ChainSourceUnavailable, pass 4 left a stale cadence-gate `state`), so
the fix removes the FIELD, not just the bug: with no `fee_window_poisoned` on
`ClaimEngine`, `self.fee_window_poisoned = true` is a compile error, not a
convention to remember.
Per-cycle conditions (corrupt + future-dated-clock) now live in a
`CycleConditions` value built fresh at the top of every `run_cycle` from `now`
plus a freshly reloaded `RewardsClaimConfig`, used, and dropped -- never
stored on the engine. `corrupt` is now re-read from disk every cycle too (it
previously latched at construction only), matching what
`ClaimLoopState::PersistedStateCorrupt`'s doc already claimed but the code
never did.
Rewrites the single-cycle f10 regression into a two-cycle test: cycle 1 with a
future-dated clock refuses; cycle 2, after the clock catches up and the
cadence elapses, MUST claim. The old one-cycle version was green whether the
latch bug was present or not.
Refs #594
* fix(rewards-claim): satisfy clippy doc-list indent and rustfmt
Clippy failed with 3x doc_lazy_continuation on the PersistedStateCorrupt
doc comment (types.rs:165-167): continuation lines of a `-` bullet must
be indented under the marker, not left flush. Indent them.
Rustfmt failed on the new fail_reserve_asset_for early-return in
FakeChainPort::reserve_asset_id (engine.rs:888): the Err(...) call
exceeded the line-length limit unwrapped. Let rustfmt wrap it.
Refs #594
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* test(rewards-claim): red proof for corrupt-then-repaired stale read
Cycle 1 refuses a corrupt fee-window file; the file is then repaired to
valid values with a fully-spent window and a recent completed-cycle
time. Cycle 2 must neither grant a fresh budget nor skip the cadence
gate. Fails against current `with_persisted_fee_window`, which loads
the three fee-window fields once at construction and never refreshes
them from the per-cycle `cfg` -- see engine.rs:149-157, #594.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* fix(rewards-claim): resync fee-window fields from disk every cycle
`with_persisted_fee_window` only loaded fee_window_start_unix,
fee_spent_in_window_mojos and last_cycle_completed_at once, at
construction. Once the now-deleted fee_window_poisoned latch stopped
masking it, a file corrupt at construction and repaired later left
those three fields stuck on poisoned()'s None/0/None placeholders --
a fresh budget and a skipped cadence gate, and persist_fee_window then
overwrote the repaired disk values with them.
CycleConditions now carries the three fields from the SAME freshly
reloaded cfg it already used for the corrupt/future-dated check, and
run_cycle copies them onto self before the cadence gate or window-roll
logic runs, but only on a read that is neither corrupt nor future-
dated. This also fixes Finding 2b: future_dated_clock now reads cfg's
own clocks instead of self's stale ones. Corrects the doc claim at the
old lines 236-238 to describe what the code now does for both halves.
Closes #594.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
* refactor(rewards-claim): make disk the sole store for the fee window
Delete `fee_window_start_unix`, `fee_spent_in_window_mojos` and
`last_cycle_completed_at` from `ClaimEngine`. `run_cycle` already re-reads
`RewardsClaimConfig` fresh every cycle for the corrupt/future-dated check,
so caching a copy on the engine bought nothing and cost exactly the
stale-read defect class F16 just fixed. A local `FeeWindowState`, scoped to
one `run_cycle` call, now threads the in-flight values through
`evaluate_budget_phase`/`uncommit_fee`/`persist_fee_window` instead. With
no field left to cache into, a future `self.fee_window_start_unix = ...`
outside this file is an E0609 compile error, the same enforcement
`fee_window_poisoned`'s removal already has.
No behaviour change: every early return, the corrupt/future-dated fail-
closed path, the cadence gate, the window roll, write-then-spend
pre-commit/uncommit and the per-claim ceiling are unchanged -- only where
the three values live changed.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
* chore(release): v0.256.0
Bump dig-node-service to v0.256.0 for release.
This release includes:
- Reward distributor prover loop (#593)
- Peer reward claim loop (#594)
- Reward prover status RPC (#595)
* ci: scope commitlint to PR-introduced commits, fix title suffix check
A develop -> main release-cut PR was linting main..develop, the full
inherited commit range, instead of just the commits it introduces.
Every commit in that range was already linted at its own PR while it
was still mutable; re-linting it at cut time adds no information and
cannot be satisfied once merged (gitlinks and rev-pinned deps make
history immutable). Use commitDepth: 1 on a main-base PR; keep the
full-range lint unchanged for develop-base PRs, where authors can
still fix the commits.
Also fix the PR-title lint's blind spot: GitHub's squash merge lands
"$PR_TITLE (#$PR_NUMBER)" as the commit subject, about eight
characters longer than the title alone, so a title that passes
header-max-length can still produce an over-limit commit subject that
nothing checks. Lint the exact string that will land.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
---------
Co-authored-by: Claude <noreply@anthropic.com>
* chore(release): v0.257.0 -- the reward distributor lifecycle starts running (#607)
* feat(mirror): persist mirror-bond coin ids (#575)
* chore: open lane for #574
* feat(mirror): persist mirror-bond coin ids so a restart cannot double-create
Bond identity was reconstructed from a live chain scan on every read
(`mirror/observe.rs`), with no persistence of its own. A restart, a cold
replica, or a lagging/flaky chain source all rendered a real, unspent,
confirmed bond as "no bonds" -- and because the in-flight suppression is
keyed on pending/submitted audit entries, a bond whose create had already
CONFIRMED was not suppressed either, so the same short scan that emptied
the read surface also cleared the one thing that would have stopped a
second coin being paid for collateral that already exists (dig-node#574).
Persist the (store, root, epoch) -> coin_id mapping in the EXISTING spend
audit record (spend-audit.jsonl) rather than a new store: a mirror-coin
create already writes store_id + AuditedBond{root, epoch} + amount there,
and the coin id itself becomes durable the moment resolve_landed_spends
confirms it. This adds the one missing piece -- the advertised URL a
create carries -- and a read-side query, confirmed_mirror_bond, that
returns the newest CONFIRMED record naming a triple.
Chain stays authoritative. mirror::local_bond::recheck_missing_bonds
never trusts the record: for a held bond the live scan did not cover, it
asks the record for a candidate coin id, then re-verifies that SPECIFIC
coin against chain via the same independent check (chain_bond_verdict)
that verifies an untrusted peer's claimed bond. Only a fresh `Bonded`
verdict is folded back in, as covered; `Unbonded`/`Unverified` fall
through to an ordinary create, exactly as if no record existed.
Version: 0.254.86 (patch -- per #522 the MSI ProductVersion minor field
is exhausted and the counter lives in patch).
Co-Authored-By: Claude <noreply@anthropic.com>
* test(mirror): prove the recovery wiring end to end through PassRunner::run
Adds two integration-level tests over the REAL pass pipeline, not just the
isolated recheck_missing_bonds unit tests: a bond missing from the live
scan with a chain-reverified durable record is recovered (no double
create, correct Bonded state reported), and the control -- the same
record but chain disproves it -- correctly falls through to an ordinary
create. Together these are the concrete regression test for the
cold-start/lagging-chain-source double-create scenario the ticket asked
to have measured.
Also refactors in_flight_creates to take the already-folded SpendLedger
instead of re-reading the log itself, so PassRunner::run reads the audit
file once per pass and shares it with the new recovery step, and fixes a
doc comment on in_flight_creates that the recovery step would otherwise
have made stale on landing ("a Confirmed create has a coin the chain
observation already sees" is no longer unconditionally true).
Co-Authored-By: Claude <noreply@anthropic.com>
* chore(fmt): wrap long test signatures to satisfy rustfmt
Co-Authored-By: Claude <noreply@anthropic.com>
* chore(clippy): use slice::from_ref instead of cloning for a single-element slice
Co-Authored-By: Claude <noreply@anthropic.com>
* chore(release): bump to v0.254.89
Base branch moved to develop after PR #576 merged there at v0.254.88
(main and develop are currently identical), leaving this branch's
carried-forward .88 as a zero-increment against the new base. Bumped
to the next free integer after fetching and verifying both origin/main
and origin/develop tip at .88.
Co-Authored-By: Claude <noreply@anthropic.com>
---------
Co-authored-by: Claude <noreply@anthropic.com>
* fix(peer): count accepted relayed circuits in the connected pool (#579)
serve_accepted_relay_conn served every accepted relayed circuit (full mTLS
auth, full L7 peer RPC) while registering it nowhere, so connected_peers
under-reported every relayed inbound peer -- the relay-leg twin of the
direct-inbound defect #402/#523 already fixed.
adopt_inbound_peer_in_pool now dispatches by TraversalKind: Relayed routes to
dig-gossip's already-published adopt_relayed_inbound_handle (v0.32.0, the rev
this repo already pins), every other tier keeps the unchanged
adopt_direct_inbound_handle path. serve_accepted_relay_conn adopts before
serving and releases after, mirroring the direct listener exactly.
Refs: https://github.com/DIG-Network/dig_ecosystem/issues/3124
* fix(cli): guard the exit-code namespace shared with diga against collisions (#582)
* chore: open lane for #3189
* fix(cli): guard the exit-code namespace shared with diga against collisions
dign and diga deliberately share one process exit-code numbering (dig-app's
outcome.rs says so in its own doc comment), so a number is free only if it
is unoccupied ecosystem-wide. dig-node#407 assigned exit 7 to
NODE_UNREACHABLE by checking only this repo's own table, where 7 genuinely
was free -- and collided with diga's NOT_CONNECTED. A reviewer caught it by
hand; nothing failed automatically.
Adds scripts/check-exit-code-collisions.sh: parses both enums' code()/name()
match arms straight from their own source -- this repo's ExitCode, and a
live fetch of dig-app's outcome.rs at its default branch -- and fails if a
number carries two different names, or if either side draws a number from
the reserved shell signal range (126, 127, 128+N). Ships with an 18-case
hermetic test harness (scripts/tests/check-exit-code-collisions.test.sh)
covering the actual #407 collision shape, arm-order independence, arm-count
mismatch, the reserved-range boundary from both sides, the live-fetch path
itself, and fail-closed behaviour on an empty/missing/unreachable table.
Wires a real (unstubbed) invocation into ci.yml's existing "Release-script
tests" job so a collision introduced by a future PR, on either side, is a
red required check on that PR -- not a note a reviewer has to catch. The
fetch retries twice (2s backoff) since this becomes a required, network-
dependent check; a fetch failure still fails closed after retrying, never
silently passing as "diga has no codes".
Updates SPEC.md 8.4 to point at the mechanical guard instead of leaving
"re-check both tables" as unenforced prose, and records that the
extension's WALLET_WS_ERR.NOT_CONNECTED = -33001 is a separate JSON-RPC
error-code space, not a rival of this one. Adds a doc-comment to the
existing transcribed collision test pointing future readers at the live
script as the authoritative check; the transcription remains as a narrower,
hermetic regression pin for the #407 shape specifically.
No renumbering: every currently-assigned code is unchanged.
Refs #3189
Co-Authored-By: Claude <noreply@anthropic.com>
---------
Co-authored-by: Claude <noreply@anthropic.com>
* fix(hygiene): port the lost-continuation guard to 4 crates, fix 48 corrupted strings (#3190) (#583)
* chore: open lane for #3190
* fix(hygiene): port the lost-continuation guard to 4 crates, fix 48 corrupted strings
Replicates dig-node-service::continuation_guard (dig-node#526/#501) into dig-node-core,
dig-wallet, dig-runtime and dig-chat-protocol, line-for-line apart from crate-specific
constants -- ported rather than reinvented, per dig_ecosystem#3190.
Wiring the guard in surfaced 48 pre-existing lost-continuation defects the ticket's own
"no measured corruption in these four crates" note did not anticipate: 36 in dig-node-core,
12 in dig-wallet, mostly test-assertion prose where a multi-line message lost its `\`
continuation and shipped the source's own indentation as a mid-sentence space run (one
as the worse `\n`-plus-indentation variant). All 48 are collapsed to the single space the
sentence always meant, with surrounding indentation and wording otherwise untouched.
Two lines are real column-alignment, not defects, and get a targeted EXCLUDED_LINE_RANGES
entry on dig-node-core instead of a rewrite: download.rs's `claimed(...)` fixture-table
trailing comments, and net.rs's `label : value` debug-print alignment.
Refs https://github.com/DIG-Network/dig_ecosystem/issues/3190
Refs https://github.com/DIG-Network/dig_ecosystem/issues/3130
Co-Authored-By: Claude <noreply@anthropic.com>
---------
Co-authored-by: Claude <noreply@anthropic.com>
* feat(mirror): detect an IP change daily and reconcile mirror coins to the current advertise URL
Automatic half (D1-D4) of the daily mirror-URL reconcile: derived personal-day offset, two-observation hysteresis, nine ordered gates with K sized as a self-funding prefix before any reclaim, `submitted` never `completed`, audit lines gain `reclaim_reason` + `trigger`.
Gates at b4c09866: loop-reviewer PASS (review 5129272285), loop-security PASS @ 175304e1 (tree byte-identical, `git diff 175304e1 b4c09866` empty), adversarial loop-decider PASS (comment 5567083644; SHOULD-FIX findings ticketed separately).
Refs DIG-Network/dig-node#570
Refs DIG-Network/dig_ecosystem#3203
* feat(serve): content hosting + serve path batch, v0.255.0 (dig_ecosystem#3212)
Nine commits from the #3212 serve-path lane, gated at 899cc68f (reviewer review 5130425808,
security comment 5568662836), plus the single semver bump to 0.255.0 for the develop -> main batch.
- store_id/root case normalised at the CapsuleKey boundary; cache delete targets the matched entry
- tier-0 occupancy reads the eviction-aware ledger
- profile-sync outbound budget in bytes; announcer asked first
- melt confirmation depth on the terminal spend, fail-closed
- EngineWarming (-32002) while the peer tier attaches, never -32004
- window completeness derived from the bytes read
- deps: dig-stun 0.2, chia-query 0.24.3, dig-nat 0.21.2, dig-logging 0.2.2
Refs DIG-Network/dig_ecosystem#3212
* chore: untrack gitnexus-generated agent files (#590)
* chore: untrack gitnexus-generated agent files
These files were generated by `gitnexus analyze` as a side effect of
indexing this repository. They are development-loop private tooling
output, not product code, and carry no secrets. They are removed from
tracking going forward via .gitignore; history is deliberately NOT
rewritten.
Refs #3177
* chore: drop private-repo reference from gitignore comment
The ignore comment named a private repository and an internal issue
number in a public file, which is the same disclosure class this
change set exists to remove; the reference is dropped and the
guidance kept.
* feat(rewards): always-on prover loop engine -- honest liveness, type-enforced self-exclusion, bounded spend (#593)
The always-on reward-prover engine: ~2,000 lines under
`crates/dig-node-core/src/rewards/`, built against the merged `dig-rewards-coin`
SPEC. Library only -- nothing spawns it, and the sole production
`RewardsChainPort` refuses every call, so it cannot spend. Wiring the composed
system is #3265, which carries its own gate.
The epic's premise -- "anytime the process isn't running, rewards are not being
distributed" -- is half wrong, and the false half is the dangerous one. `Sync`,
`NewEpoch` and `InitiatePayout` need no manager authority, so funder downtime does
not stop rewards: it FREEZES THE ENTRY SET while accrual and payouts continue.
Peers that stopped mirroring keep earning; peers that started cannot begin. That
shaped the whole design.
Liveness honesty (SPEC 2.4). The status record carries no `healthy`/`ok`/`up`
boolean and no precomputed staleness, because a wedged loop cannot report its own
wedging -- whatever it last wrote stays there, so a writer-set flag reads true
forever after the failure it exists to reveal. The reader derives staleness from
`last_cycle_completed_at` against `observed_at` and its own clock. A recursive
JSON-key test enforces the absence at every nesting depth; asserting on keys and
never substrings, since `ProverState::Running` legitimately serializes the VALUE
"running". The one legal staleness signal is chain-derived (SPEC 12.4: 48 hours
AND a non-zero reserve, from the singleton's own spend history) and lives on the
distributor read, where a wedged prover cannot fake it.
Self-exclusion is a compile error, not a habit. dig-node#261's lesson is that an
invariant enforced on some paths is not an invariant. `admit` is the single
admission point, checks both SPEC 5.2 coordinates (own peer_id OR a payout puzzle
hash this wallet controls), and mints an `AdmittedPeer` with private fields and no
public constructor -- so `EntryAction::Add` cannot be built by a path that skipped
admission.
A prover's own fault can never strike a peer. `GateError` is a distinct type from
`GateIneligibleReason` and `record_prover_fault` takes `&self`, so SPEC 3.6 clause
4 is enforced by the borrow checker rather than by comment. Without that, a
misconfigured operator -- one missing mirror-collateral epoch ordinal -- would
strike every peer at once and evict its entire 250-entry set in three hours, each
eviction a fee it pays plus a settlement out of its own reserve.
The money bounds are stated where a human reads them (`rewards/mod.rs`): 24
bundles/day, 192 entry actions/day, a fee ceiling of 24x the configured standard
fee, 192 removals/day worst case with 96/day sustained churn. Recorded honestly:
SPEC 6.3's rate bound and fee ceiling are ONE control, not two.
Three gate rounds, every leg fresh-context. Round 3 at this head: reviewer PASS,
adversarial decider RATIFY (leg closed), security CHANGES-REQUIRED on a finding
the decider ratified deliberately -- adjudicated in
https://github.com/DIG-Network/dig-node/pull/593#issuecomment-5601784815 and
carried to #3265 with the remedy corrected, because the proposed fix would have
persisted a poison flag to the very store whose writes were failing.
Found and fixed under gate: a census ordinal off by one in both directions (SPEC
4.6 requires n-1 exactly); an unreachable grace window leaving a named constant
with no reader; a missing `NewEpoch` spend; an absent-ordinal path attributing a
prover fault to peers; and a daily fee ceiling 24x too high because a per-bundle
fee was consumed as a daily ceiling.
Refs DIG-Network/dig_ecosystem#3250
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
* feat: serve dig.getRewardProverStatus at Tier::Control (#5…
Refs #3250
DO NOT MERGE — gate round in progress