The Deterministic Executor
- Why not tokio?
- Randomized AND deterministic
- Where is the reactor?
- The drive loop
- Coming from tokio
- Kill-on-drop
- Exploration process safety
- Verifying it yourself
Moonpool simulations do not run on tokio. They run on a purpose-built, single-threaded executor whose every scheduling decision derives from the iteration seed. This chapter explains why it exists, how it works, and what changes if you are used to tokio.
Why not tokio?
Two reasons, one practical and one fundamental.
The practical one: seeding tokio’s runtime randomness requires the unstable
RngSeed API, which forces --cfg tokio_unstable onto every crate in every
consumer’s workspace. For stable-pinned projects that is a hard adoption
barrier (issue #151).
The fundamental one: tokio’s current-thread scheduler is a FIFO queue, and that is the wrong policy for a bug-finding simulator. When one event makes two tasks runnable, FIFO runs them in registration order, on every seed, forever. Any race between those tasks is structurally unexplorable: no amount of seeds will ever try the other order. A simulation executor should treat “which runnable task goes next” as a decision to explore, the same way it explores message latencies and fault timings.
Randomized AND deterministic
Those two properties sound contradictory, but they are not. Deterministic
means “same seed, same execution, bit for bit”, not “one fixed schedule”.
The executor keeps its ready tasks in a Vec and picks the next one with a
swap_remove at a seeded-random index. madsim uses the same distribution,
and FoundationDB’s sim2 randomizes the same way:
seed 42: task-3, task-0, task-2, task-1, ...
seed 42: task-3, task-0, task-2, task-1, ... (always)
seed 43: task-1, task-3, task-0, task-2, ...
Each seed replays exactly. The population of seeds covers the ordering space. Ordering bugs, the kind that survive code review because both orders look fine locally, live exactly there.
The scheduling index is a draw on the simulation’s one random stream, the
same sim_random_range a process reaches through ctx.random(). There is no
private scheduling RNG: which task runs next is counted like every other
decision, so a fork-explorer count@seed recipe replays the schedule along
with the faults and the workload. The only time the executor does not draw is
when exactly one task is runnable, which is a forced choice, not a decision.
Where is the reactor?
There is none, and that is the deep reason this executor is small. A
production runtime pairs its executor with a reactor (epoll, timers) that
turns OS events into waker calls. In the simulation, SimWorld and its global
Scheduler<Event> play that role. The scheduler owns logical time, stable
same-time ordering, and cancellation. It dispatches targeted events to
NetworkSimulation and StorageEngine, which own their resource state,
operation results, and waker registries. The returned wake batches are invoked
after the world lock is released. The executor only answers one question: of
the tasks that are runnable right now, which runs next?
The drive loop
Executor::block_on(main) pins the orchestrator future on the stack as the
driver. It never enters the ready queue and is never scheduled randomly:
loop {
poll driver ── Ready? ─────────────► return
run_until_stalled() // poll ready tasks in seeded-random order
// until none is runnable
(driver not woken AND scheduler empty) ► panic: deadlock (with the seed)
}
The orchestrator’s step loops interleave the simulation with the task pool:
sim.step(); // process one virtual event
crate::executor::until_stalled().await; // let every woken task run
Because the driver is only re-polled after a complete drain, resuming from
until_stalled().await guarantees that every task that was runnable has
been polled to Pending or completion. Under tokio this guarantee was
approximated by yield_now plus FIFO ordering. Under randomized scheduling
“the back of the queue” does not exist, so the executor provides the
contract directly.
Inside a task, use executor::yield_now() instead: it reschedules the task
at a seeded-random position. until_stalled() is driver-only and asserts
it, in every build profile, because a task awaiting “run until nothing is
runnable” would be waiting for itself.
One consequence deserves a warning sign: a task must not busy-wait on
progress only the driver can make. A loop { yield_now().await } polling
for a flag that a simulation event sets keeps the ready queue non-empty
forever, so the drain never finishes and sim.step() never runs. The
executor kills such a loop with a poll-bound panic naming the task and the
seed. Park on a real wake instead: a sleep, a Notify, a channel.
Coming from tokio
The task API mirrors what the orchestrator (and your process/workload code,
through TaskProvider) already expected from tokio:
| tokio | moonpool executor |
|---|---|
tokio::spawn(fut) | executor::spawn("name", fut) (named, for seed debugging) |
JoinHandle::is_finished() | same |
JoinHandle::abort() | same (await yields Err(JoinError::Cancelled)) |
dropping a JoinHandle | same: detaches, task keeps running |
| fire-and-forget | JoinHandle::detach() — explicit form of the above |
| task panics | caught, handle yields Err(JoinError::Panicked), siblings unaffected |
dropping the Runtime | dropping the Executor cancels every live task |
tokio::select! | moonpool::select! (tokio’s own expansion, seeded start offset) |
Spawned futures stay Send + 'static, exactly as before: execution is one
OS thread, but the bounds let application code use Arc<RwLock<...>>,
DashMap, and friends without contortion.
The raw executor keeps panic isolation deliberately narrow. A detached task
spawned with executor::spawn records its panic in its own JoinHandle, so a
low-level driver can decide how to recover. Tasks spawned through an actor’s
ctx.task().spawn_task() carry a different contract: the simulation runner
records an unobserved child panic with its process or workload identity and
marks that seed as failed. If the caller awaits the handle and handles
Err(JoinError::Panicked), the panic is observed and the seed can recover.
Dropping or detaching the handle leaves the panic unobserved, even if the task
already finished. Dropping a task during process shutdown or executor teardown
remains cancellation, not a panic.
The runner also records each unobserved panic as an unobserved_task_panic
event with actor, task, and panic fields. It runs invariants once more
after task teardown, so those diagnostics remain visible even when setup or
check stalls and orchestration returns early.
time.sleep(Duration::ZERO) still schedules a real same-time timer, preserving
the scheduler’s FIFO ordering for a burst of immediate work. The runner allows
that fan-out, then requests shutdown after a generous fixed count of events
that keep rearming at the same logical instant. If the tasks ignore shutdown,
the next stagnant burst fails the seed. Time advancing or a task completing
resets the count, so ordinary cleanup and long timer-driven runs keep working.
Two behaviors are deliberately better than tokio’s:
- A genuine deadlock (driver not woken, nothing runnable) panics with the seed in the message instead of parking the thread forever.
- A task that stays runnable forever, whether a busy-yield loop or a select over a source it synchronously re-readies (something tokio’s cooperative budget used to interrupt), trips a poll-bound panic naming the task and the seed, in every build profile, instead of hanging the drain loop.
Kill-on-drop
Simulation iterations must be hermetic: a task leaked by seed N must never
run during seed N+1. Dropping the per-iteration tokio runtime used to
guarantee that, and the executor replicates it with a waker registry (the
same pattern async-executor uses). Every live task registers one waker.
Drop wakes them all, which schedules every parked task, then pops and
drops Runnables until the queue is empty. Dropping a Runnable cancels
the task and drops its future, and any tasks that drop wakes land in the
same queue and are consumed by the same loop.
Exploration process safety
The explorer only forks between complete timelines, after the deterministic
executor and all of its tasks have been dropped. Assertion macros record
discovery coordinates; they are not fork points. Forked workers are still an
explicit performance mode because unrelated host threads and resources may
exist outside the simulation. The default workers: 0 mode runs continuations
sequentially in-process and is the conservative choice for deterministic CI.
Verifying it yourself
The executor’s contracts are enforced by crates/moonpool-sim/tests/executor.rs
(join/abort/detach/panic/kill-on-drop, same-seed replay, cross-seed
diversity) and crates/moonpool-sim/tests/determinism.rs, an end-to-end tripwire:
two full SimulationBuilder runs of racing workloads on the same seed must
produce byte-identical execution traces. If any component, executor,
select!, providers, ever consults an unseeded randomness source, that test
fails.