Problem
ClientManager.pre_run_callback currently calls pre_run_callback on every registered client, then combines the results with all(...). If any client returns False, run_block skips kernel execution and does not call post_run_callback for any client.
This creates an inconsistent lifecycle contract for multi-client runs: a client that returned True may have already mutated launch/block state expecting a matching post_run_callback, but a later client can veto the block and prevent that cleanup from ever running.
Example
One concrete case is a symbolic client used together with another client that can veto blocks, such as profiler block sampling. If the symbolic client marks a block as active during pre_run_callback, and the profiler later returns False, the block never executes and no post_run_callback happens. That active count or related launch-scoped state can leak into the next launch.
This can lead to stale tensor/address state persisting longer than intended, and in sanitizer-like clients stale ranges could make later accesses appear in-bounds incorrectly.
Expected Behavior
The client lifecycle should make it unambiguous which callbacks are paired:
pre_run_callback should either be treated as a pure veto check, with no state that requires post_run_callback, or
- the manager should provide a separate callback after the aggregate pre-run decision is known, e.g.
block_start_callback, and clients should mutate paired block state there instead.
Notes
This is a broader multi-client lifecycle issue rather than a sanitizer-only problem. It is worth fixing at the ClientManager contract level so all clients have the same guarantees when another client vetoes a block.
Problem
ClientManager.pre_run_callbackcurrently callspre_run_callbackon every registered client, then combines the results withall(...). If any client returnsFalse,run_blockskips kernel execution and does not callpost_run_callbackfor any client.This creates an inconsistent lifecycle contract for multi-client runs: a client that returned
Truemay have already mutated launch/block state expecting a matchingpost_run_callback, but a later client can veto the block and prevent that cleanup from ever running.Example
One concrete case is a symbolic client used together with another client that can veto blocks, such as profiler block sampling. If the symbolic client marks a block as active during
pre_run_callback, and the profiler later returnsFalse, the block never executes and nopost_run_callbackhappens. That active count or related launch-scoped state can leak into the next launch.This can lead to stale tensor/address state persisting longer than intended, and in sanitizer-like clients stale ranges could make later accesses appear in-bounds incorrectly.
Expected Behavior
The client lifecycle should make it unambiguous which callbacks are paired:
pre_run_callbackshould either be treated as a pure veto check, with no state that requirespost_run_callback, orblock_start_callback, and clients should mutate paired block state there instead.Notes
This is a broader multi-client lifecycle issue rather than a sanitizer-only problem. It is worth fixing at the
ClientManagercontract level so all clients have the same guarantees when another client vetoes a block.