(This is a big one, get ready - I told Claude to write the commit message cause I couldn't be bothered)
Root cause: reorg was broken at every layer and the failures compounded. Verified with the node's own
SKALACOIN_FORCE_ORPHAN_REORG debug mode, which stalled permanently at height 2 on 114 consecutive
coinbase-validation failures. A binary built at HEAD behaves identically, so none of this is a regression —
the reorg path had simply never worked.
Rollback (the keystone):
- Chain_RollbackToHeight always returned false on any node that had ever saved or loaded its chain, and
only after it had already truncated the chain and destroyed the balance sheet. Chain_RecomputeRuntimeState
bails on any header-only block, and Chain_SaveToFile nulls transactions on every in-memory block once
persisted, so the failure was universal in practice. Both callers treated the false as "nothing happened"
- Added Chain_BorrowBlockTransactions / Chain_ReturnBlockTransactions, which fall back to the on-disk copy
when the in-memory block has been compacted to headers
- Supply is now accumulated in the rollback's existing balance-sheet replay pass instead of a second
Chain_RecomputeRuntimeState pass
- Gave Chain_RecomputeRuntimeState the same disk fallback: it had been failing on every restart with an
existing chain, silently leaving currentSupply/currentReward at whatever came out of chain.meta
Fork choice is now cumulative work, not height:
- Added Chain_ComputeBlockWork / Chain_ComputeWorkRange / Chain_ComputeBranchWork, computing
2^256 / (target + 1) per block and summing over a range. Derived on demand from headers — no header,
chain.meta or wire-format change
- uint256 only had add/sub/cmp, so added uint256_divide (restoring binary long division),
uint256_from_be_bytes, uint256_bitwise_not and uint256_is_zero
- Comparison is strictly greater, so tied tips do not cause two nodes to keep swapping
- Height-based choice was wrong now that difficulty actually varies: a long low-difficulty branch beat a
short high-difficulty one
Atomic branch replacement:
- Added Chain_ReplaceBranch: validate linkage, apply the reorg penalty, compare work, snapshot the outgoing
blocks, roll back, apply. On any failure the original chain, balance sheet, supply, reward and difficulty
target are restored. The caller keeps ownership of its blocks in every case — the chain applies copies
- Split Chain_AddBlock and Chain_RollbackToHeight into locked public wrappers over unlocked internals, so a
whole branch swap happens under one lock acquisition and Chain_OnTipAdvanced runs once per reorg rather
than once per block
- Chain_AddBlock now validates header.prevHash against the tip. It never did — that check lived only in
Chain_IsValid and the network path, which is exactly why a rollback-then-reapply could splice blocks from
two different forks into a chain that no longer links up
- Moved the currentSupply/currentReward update into Chain_AddBlock. Each caller used to do it separately, so
the orphan-attach and maintenance-thread paths never did, and the next block's coinbase was then validated
against a stale currentReward and rejected forever. This was the height-2 stall
Reorg penalty (Horizen-style delayed block submission):
- The penalty was only ever reachable from the manual sync command. The P2P broadcast -> orphan pool ->
branch adoption path, which is the path an attacker actually uses, had none at all and picked the winner
by raw height. It is now enforced inside Chain_ReplaceBranch, the single choke point every adoption
passes through
- Removed its application to the sync fetch window. The height gap to a peer is not a reorg depth;
penalising it only throttled honest catch-up, and for gaps of 4-50 it collapsed the window to one block
per pass, defeating MAX_PARALLEL_FETCHES
- The depth is stamped once when a branch is first observed (orphan_entry_t.observedAtTipHeight) and never
recomputed. Re-deriving it from a moving tip never converges: depth and elapsed both grow by one per block
while penalty(depth) grows faster, so a penalized branch could never be adopted at all
- The initial-sync exemption now comes from Chain_IsInitialBlockDownload, which uses the local median time
past over MEDIAN_TIME_SPAN blocks. It used to key off the peer's advertised height, so any peer claiming
localHeight + INITIAL_SYNC_HEIGHT_DIFF could switch reorg handling off for the whole session. A median
rather than the tip alone means one backdated block cannot fake it either
Orphan pool (largely rewritten):
- Added a pool mutex. It had no synchronisation whatsoever while being mutated from the 1 Hz maintenance
thread, every per-peer TCP thread and the REPL thread; a concurrent insert could realloc the array while a
scan held a raw element pointer. The lock is never held across a call into chain.c
- Dedup by block hash, a MAX_ORPHAN_BLOCKS cap with oldest-first eviction, and pruning of entries that can
no longer apply. Nothing was ever reaped before, and orphans are reachable before the chain-derived
difficulty check, so this is also the memory-exhaustion fix
- Candidate branches are now assembled by following prevHash from the fork point. Taking the first orphan
found at each successive height could interleave blocks from two competing forks into one incoherent branch
- Fixed rollbackHeight = forkHeight - 1. Chain_RollbackToHeight is exclusive, so every non-genesis adoption
amputated one block too many and then failed Chain_AddBlock's index check
- Fixed a block_t wrapper leak on every successful attach (free the wrapper, not Block_Destroy — the chain
owns the transactions after a shallow copy)
- Permanently invalid orphans are dropped instead of being retried on every maintenance tick forever
Forks below the tip are now discoverable:
- A block at blockNumber < chainSize was rejected and freed, so the fork point and the lower half of any
competing branch were always thrown away and a sub-tip fork could never be learned. Now the hash is
compared: identical means a duplicate and is ignored, different means it goes to the orphan pool
- The sync loop probes downwards (RequestForkWindow, bounded by REORG_FETCH_DEPTH and MAX_FORK_PROBE_ROUNDS)
when it makes no progress while the peer is ahead. That is the only trigger that fires for a genuine
sub-tip fork, because the old divergence check could only see blocks that had already entered our chain.
FETCH_BLOCK already answers from the peer's own chain, so no protocol change was needed
- Removed the rollback-to-height-0 path. "Could not find the parent" used to wipe the entire local chain,
genesis included, and any peer could trigger it with a single unlinked block
Floating point removed from consensus and network math:
- Chain_ComputeTargetAtHeight (the difficulty retarget) used double ratio arithmetic, and
FetchScheduler_ComputeReorgPenaltyBlocks used double/pow/ceil. Both are consensus-critical and are now
integer only; float results are not reproducible across platforms and compilers, and a single last-digit
difference in a target or a penalty splits the network
- The penalty constants became integer rationals (REORG_PENALTY_FACTOR_NUM/DEN, integer EXPONENT and
REF_BLOCK_TIME) with saturating exponentiation and explicit ceiling division. Output is unchanged:
penalty(4)=10, penalty(8)=39, penalty(10)=60, penalty(50)=1500, penalty(100)=6000
- Removed the unused float macros DAG_MAX_UP/DOWN_SWING_PERCENTAGE and the now-dead math.h include from
constants.h. Both were latent: DAG size feeds PoW verification. Replaced with integer numerator/denominator
constants and used them at both clamp sites (values verified identical)
Other fixes that were blocking fork propagation:
- madeProgressOverall was set but never reset, so after one productive pass the "no progress -> stop" guard
could never fire again and the sync loop could spin forever holding the REPL
- seenBlocks was inserted before/regardless of a successful send, so a block broadcast while no peer was
connected was never offered again. It is now recorded only after the block actually goes out
- Broadcasts relayed to outbound connections only, so in a two-node setup the dialled node never pushed
anything back and the dialer learned of new blocks only via a manual sync. Inbound peers are now relayed to
Verified with two-node harnesses at a shortened adjustment interval:
- forced-orphan regression: was height 2 with 114 coinbase rejections, now reaches the peer's height with
zero rejections and both nodes report Chain OK
- sub-tip fork at depth 3: the node discovers the fork below its own tip, discards its three blocks, adopts
the heavier five, and both nodes converge on an identical tip hash
- deep fork at depth 8: the strictly heavier branch is correctly refused with depth=8 penalty=39 elapsed=0
- uint256 work arithmetic covered by a standalone test (division, big-endian conversion, monotonicity,
halved target doubles work)
One issue found along the way: the chain in build/chain_data does not pass fullverify. Block 7683 reverts to
INITIAL_DIFFICULTY where it should carry 0x1f06df14, i.e. it contains blocks mined before 1288a64 landed. A
binary built at HEAD rejects it too, so this is stale data rather than a regression — it needs a wipechain
and a re-mine.
Root cause: peers were keyed only by (IP, listen port). An IPv6 host holds several addresses at once, so one node appeared as several peers and could not recognise its own addresses.
Protocol:
- Added a random per-run node identity (localNodeId), advertised as a length-guarded trailing field in HELLO and ACK_HELLO — backwards compatible with peers that omit it
- Added peerNodeId to tcp_connection_t; peers are now identified by this rather than by an endpoint
- ACK_HELLO is now sent before any decision to drop the connection, so a rejected dialer learns whose address it reached instead of retrying forever
Self-connection:
- Identity match → close the connection and record that endpoint permanently as our own
- Self-endpoint set preemptively seeded from getifaddrs() at startup, so a node knows its own addresses before ever dialing one
- Node_ConnectPeer refuses self endpoints, which also kills the echo-back chain that could exhaust connection slots
Duplicate connections and churn:
- Dedup by identity per direction — one inbound and one outbound per physical peer, regardless of how many addresses it has
- Moved dial history out of the peer table, so striking a peer no longer resets its connect-retry cooldown (this was the loop engine: strike → re-learn via gossip → redial next tick)
- Node_HasOtherInboundFrom / Node_HasLiveConnectionTo now ignore connections that are already tearing down, and match on identity as well as endpoint
Gossip hygiene:
- Never hand a peer its own other addresses in a PEERS reply (matched on identity, not just the socket address)
- Reject unusable endpoints: link-local without scope id, unspecified, multicast, site-local (loopback stays allowed for local testing)
- Normalise IPv4-mapped IPv6 so one host cannot occupy two entries
Two bugs found along the way:
- All nodes drew the same identity — random_eight_byte() comes from srand(time(NULL)), so processes started in the same second produced identical values. Added random_secure_eight_byte() (/dev/urandom) for the identity, and mixed the pid into the seed so connection IDs stop colliding too
- Identity dedup initially left inbound-only nodes mute — broadcasts traverse outbound connections only, so suppressing a dial-back because an inbound existed would have silenced such a node. Corrected to per-direction
Other:
- peers output now shows each entry's node identity and the node's own endpoints
- Added _DEFAULT_SOURCE to the build so getifaddrs() stays visible on glibc
Reclaim outbound slots on peer disconnect:
- TcpClient_ThreadProc fired on_disconnect but never cleared the outbound
slot, closed the fd, joined the io thread, or freed the connection, so a
dead peer permanently held its outboundClients[] slot. After MAX_CONS (32)
churned connections the node could make no new outbound connections, and
leaked fds/threads/memory. (The inbound side already self-reclaimed.)
- Add a reaper (Node_ReapDeadOutbound) on the maintenance thread: under
outboundLock it detaches dead (disconnect-notified) slots, then joins the
io thread and destroys/frees each connection outside the lock.
- Guard against use-after-free with a pin count on tcp_connection_t
(TcpConnection_Pin/Unpin). The only cross-thread consumer holding a raw
connection pointer across a blocking op is the `sync` command (via
Node_GetBestOutboundPeer); it now pins the peer and unpins when done, and
the reaper skips pinned connections. Discovery's snapshots run on the
reaper's own thread, so they need no pin.
- Node_GetBestOutboundPeer/GetClientList/GetPeerEndpoints skip
disconnect-notified connections so a dead peer is never handed out.
- Node_Destroy stops+joins the maintenance thread before tearing down
outbound clients, so the reaper can't race shutdown.
Strike disconnected peers from the discovery peer list:
- Add NodeDiscovery_RemovePeer + Node_HandlePeerDisconnect: on disconnect,
remove the peer from the known-peer table once no live connection (inbound
or outbound) to its listen endpoint remains (Node_HasLiveConnectionTo;
disconnect-notified conns don't count, so both directions dropping at once
is handled). Wired into Node_Server_OnDisconnect and Node_Client_OnDisconnect.
- Implement NodeDiscovery engine: known-peer table (DynArr) with a
per-tick seed/ping/query/connect state machine driven by the node
maintenance thread; bounded multi-hop crawl (FANOUT peers per node,
hop-capped) that connects to reachable peers lowest-ping-first
- Add GET_PEERS/PEERS TCP opcodes for peer-list exchange, handled on
both inbound and outbound connections
- Measure UDP round-trip time and pass it to the on_pong callback
(previously the send timestamp was only used for retries)
- Advertise each node's listen port in HELLO/ACK_HELLO and store it
per-connection, so inbound-only peers and non-default ports are
discoverable (length-guarded parse; wire-compatible with old peers)
- Wire a udp_node_t + node_discovery_t into net_node_t: init/start in
Node_Create, tick in the maintenance loop, teardown in Node_Destroy
(stop UDP before destroying discovery to avoid callback races)
- Add Node_ConnListenEndpoint / Node_GetPeerEndpoints helpers to derive
peers' listen endpoints (outbound: dialed port; inbound: advertised),
with IPv4-mapped-IPv6 normalization and IP+port dedup
- Match ping pong/timeout callbacks by peer address (the UDP layer owns
the nonce), fixing discovered peers stuck UNREACHABLE
- Ping the peer's listen port instead of the ephemeral TCP source port
(fixes the original stub so pongs actually return)
- Add `peers` CLI command to dump the discovery table (endpoint/hop/
state/ping)
- Add discovery tunables to constants.h (fanout, max hops, target
connections, timeouts, caps)
Mining: blocks now include mempool txs, select spendable txs by fee, and pay coinbase as base reward + fees in main.c.
- Consensus: block validation now enforces coinbase accounting and rejects invalid coinbase placement, including coinbase on amount2, in block.c and transaction.c.
- Chain state: rollback now rebuilds currentSupply/currentReward, and block addition preflights spendability before mutating balances in chain.c.
- Orphans/reorgs: orphan retry is safer, rollback-triggered sync reattaches orphans immediately, and transient orphan failures no longer drop blocks in orphan_pool.c and main.c.
- Networking/mempool: node lifecycle now initializes the mempool, broadcasts can exclude one peer, and mempool snapshotting supports mining selection in net_node.c and txmempool.c.
- Ledger simulation: added non-mutating spendable-transaction selection for block assembly in balance_sheet.c.