ADR-0024: a pending admission survives a restart

On 2026-09-26 at 14:32 IST a log line that could not be written ended 8902's block producer (see the say!

Status: proposed (Design only). Dated 2026-09-26.

Status: proposed, 2026-09-26. Design only; no code yet. The owner chose this order on 2026-09-26: the rollout's pending guard first (509d154), this next, as node work for a cut after the one staged that day.

The problem, as found

On 2026-09-26 at 14:32 IST a log line that could not be written ended 8902's block producer (see the say! macros and turn() in the node, prepared the same day). When the node was restarted at 15:57 it replayed its 19 blocks and read 0 pending: the 6 transactions it had admitted and not yet sealed were gone with the process.

Losing them is the smaller half. Each of those transactions had been admitted under a signed receipt naming its batch and position (Node::admit: sequencer.accept signs Receipt { order_hash, batch, position, l2_slot, chain }, and the receipt is what the client gets back). The batch counter the node persists, sequencer.json's nextBatch, is written at seal as current_batch(): the batch the node then issues receipts INTO. On open, Sequencer::resume(id, cfg, next_batch, 0) reopens that same batch at position 0. The sequencer's own doc says the caller "is responsible for passing a batch number strictly greater than any batch it has already issued receipts in; this cannot check that, because the evidence lives in the receipts other people hold."

So after a restart with N pending admissions, the next N admissions get receipts for the same (batch, 0..N-1) slots with different order hashes. Two receipts under one key for one slot is the definition of equivocation this protocol slashes (bond::slash_equivocation), provable by anyone holding a lost receipt against the new block's logged receipts. On a bonded chain — gamma — an honest restart at the wrong moment costs the bond. 8902 (chain 9002) has no bond, so its contradiction stands with nothing at stake.

The stale-copy guard (ADR-0020, stale_copy_check) refuses a datadir that is behind L1. It does not see this case: L1 and the datadir agree on the next batch; what is missing is the node's own memory of what it already signed into it.

What holds today

  • Admissions are memory-only: Node::pending: Vec<wire::Admitted>, Node::receipts (issued, unsealed), seen / seen_order, forced_ids (the inbox entry a drained admission came from), and the explorer's admission timing. Blocks are persisted at seal (persist::append_block, blocks.jsonl), receipts verbatim with the block (LoggedReceipt: "a node can copy what it signed; it must not sign again").
  • Sequencer keeps batch, l2_slot, pending: Vec<OrderHash>, assigned: BTreeMap<order, position>, last_time. resume sets the first two; seal takes pending, clears assigned, advances batch.
  • The cut (rollout-fleet.ps1, since 509d154) refuses to stop a node whose getStats.pending stays above 0 after waiting. That is the guard for planned restarts. A crash, an out-of-memory kill, a machine sleep or a supervisor's exit are not planned.

Decision

Journal every admission in the datadir the moment it is made, and restore the journal on open, verbatim, so that the same (batch, position) slots carry the same order hashes after a restart and the same signed receipts are the ones this node holds. Nothing is re-signed.

The journal

<datadir>/pending.jsonl, append-only, one line per admission:

  • Written in Node::admit AFTER sequencer.accept returns the signed receipt and BEFORE the admission is pushed to pending and before admit returns — so a client never holds a receipt the journal does not. A crash between accept and the append loses a receipt no one was given; that is harmless.
  • Written with append + flush, as append_block is. A power loss can lose the tail of either file alike; that is the existing durability, not a new promise.
  • Cleared at seal, AFTER append_block and save_sequencer succeed: the sealed block now carries those admissions and their receipts. Clearing is a truncate (pending.jsonl becomes empty), not a delete, so a reader never sees a missing file mean something.
  • No datadir (an ephemeral node), no journal: as today.

Restore, in Node::open

After the log has replayed and the sequencer has resumed at next_batch:

  1. Read pending.jsonl. A malformed line refuses to open, by name: a journal the node cannot trust is a journal it must not guess at.
  2. Skip any entry whose id is already in seen — an admission the last logged block sealed, left in the journal by a crash between append_block and the truncate. Say how many were skipped.
  3. Every remaining entry must name batch == sequencer.current_batch() and position == its index among the kept entries, and its receipt must verify under this node's key. Anything else refuses to open, like stale_copy_check does: the journal describes signatures this key made, and a journal that disagrees with the sequencer's counter means the two files come from different histories; signing anything from that state is the equivocation this ADR exists to prevent.
  4. Rebuild the sequencer's open batch without signing: Sequencer::resume_with(id, cfg, batch, l2_slot, orders, last_time) in the sequencer crate, which pushes each order hash and its position into pending / assigned in order and refuses a duplicate or a position out of sequence. last_time is the last receipt's l2_slot-time as journaled, so time cannot go backwards on the next accept.
  5. For each entry, in order: receipts.push the journaled SignedReceipt (verbatim), seen.insert(id), seen_order.push_back((id, head + 1)), pending.push(Admitted { id, fee_payer, instructions, published }), and forced_ids.insert when forced is set. The admission gate is not re-run: these were admitted, under receipts, before the restart, and pending is what the gate's projection rebuilds from.
  6. Say what was restored: "restored 6 pending admission(s) into batch 303 (positions 0–5) from pending.jsonl".

What does not change

  • Replay of the log, the block format, the receipt format, the published bytes: untouched. The journal is a node-local file.
  • stale_copy_check: untouched, and still runs.
  • The rollout's pending guard stays: a planned restart with nothing pending is still the cleanest restart, and the guard costs nothing.

Consequences

  • An honest restart — planned or not — no longer manufactures a contradiction against the node's own key. The lost-transaction half of the problem goes with it: restored admissions seal in the next block, in the order and at the positions their receipts name.
  • A journal that disagrees with the counter stops the node instead of letting it sign. That is a new way for a node to refuse to start, and it is the right one; the message names both numbers.
  • Two crates change: solieum-sequencer gains resume_with; the node gains the journal in persist.rs, the append in admit, the truncate at seal, the restore in open. About 250 lines with tests.
  • Tests to write: an admission survives Node::new + open with the same receipt bytes and the next admission takes the next position; a journal entry the log already sealed is skipped; a journal naming another batch refuses to open; a malformed line refuses to open; a restored forced admission keeps its inbox index.

Alternatives considered

  • Bump the batch counter on every open, so a reopened batch is never one receipts were issued in. Breaks "batch k carries block k+1", which replay, the DA feed's next_seq and stale_copy_check all rely on, and still loses the transactions.
  • Journal only the receipts and re-derive the admissions from the published bytes on restore. Re-derivation re-runs validation (blockhash freshness, signatures) on a transaction that was already admitted under a receipt; freshness holds identically since no block advanced, but nothing is gained by re-deciding, and a forced entry's origin (forced_ids) would be lost. Journal the whole admission.
  • Accept the loss and rely on the rollout guard. Covers planned restarts only; today's case was not one.