Conversation
Introduce an experimental runner that executes circuits as gate batches over cache-sized state blocks. Each batch selects causally executable raw gates, maps their logical qubits into block-local physical positions, fuses the mapped gates, and executes them across all state blocks in parallel. Bound planning cost by comparing greedy, resident, and upcoming-gate- seeded candidates. Score candidates by circuit progress while accounting for the number of required state-vector swap passes. Apply each set of disjoint qubit-position swaps in one state-vector pass. Keep architecture-specific SIMD lane positions pinned so remapping moves whole chunks on NEON, SSE, AVX2, and AVX-512 layouts. Add QubitMappedState to retain the logical-to-physical qubit layout across execution. The mapped-state API returns without restoring canonical qubit order, while the raw-state compatibility API restores identity before returning. Add a command-line runner for exercising the gate-batch implementation and register the runner and supporting headers with the Bazel build. Register a Makefile target that builds both the gate-batch runner and the existing base runner. Signed-off-by: Xin Xie <xinpin@google.com>
There was a problem hiding this comment.
Code Review
This pull request introduces a gate-batched cache-local simulation runner (QSimGateBatchRunner) along with qubit remapping and layout utilities (QubitLayout, QubitMappedState, and ApplyBitPairSwaps) to optimize simulation performance. It also adds a benchmark application (qsim_gate_batch) and updates the build configurations. Feedback was provided to relax an assertion in run_qsim_gate_batch.h that would otherwise cause failures for small circuits where the number of qubits is less than the SIMD chunk size.
| layout_(layout), | ||
| gate_batch_planner_(partition_.num_state_qubits, | ||
| partition_.block_qubits, chunk_qubits_) { | ||
| assert(partition_.block_qubits >= chunk_qubits_); |
There was a problem hiding this comment.
The assertion assert(partition_.block_qubits >= chunk_qubits_); will fail for small circuits where the number of qubits is less than the SIMD chunk size (e.g., a 2-qubit circuit on an AVX-512 backend where chunk_qubits_ is 4). Since cache blocking is not required for circuits that fit entirely within a single SIMD chunk, we should relax this assertion to allow block_qubits to be smaller than chunk_qubits_ when the total number of state qubits is also smaller than chunk_qubits_.
| assert(partition_.block_qubits >= chunk_qubits_); | |
| assert(partition_.block_qubits >= std::min(chunk_qubits_, partition_.num_state_qubits)); |
|
@xxie24 Thank you for this work. I know it's still a draft PR, but I wanted to mention that a recent failure in the CI workflow has now been fixed. I took the liberty of updating your branch to have the latest changes (I hope that's okay), and now the PR passes all the Ci tests. |
cool. thanks. |
Dynamic scheduling with small chunk sizes caused excessive atomic compare-and-swap synchronization contention on the shared work-sharing counter across high core counts. Switching to schedule(static) eliminates lock contention and preserves sequential hardware stream prefetch performance.
Allow logical CPUs that share a physical core to cooperate on one cache-resident state block. Add the -i/inner_threads option while preserving the existing one-thread-per-block behavior by default. Design decisions: - Split each simulator kernel's iteration range between team lanes with CooperativeFor, and synchronize the lanes after every gate because the next gate consumes the preceding gate's output. - Keep independent-block and SMT-team scheduling explicit, while sharing the block-local gate loop through a compile-time cooperative helper. The non-SMT path therefore pays no barrier or runtime branch overhead. - Discover Linux CPU topology from the process affinity and sysfs package/core identifiers. Order logical CPUs by physical core, pin every team lane explicitly, and reject configurations where the requested siblings are unavailable instead of silently pairing unrelated cores. - Keep non-OpenMP builds working and report a clear error if SMT teams are requested without OpenMP. - Register the cooperative loop and CPU topology headers in the Bazel library. On a Google C4 Intel Xeon Platinum 8581C with 8 physical cores and 16 vCPUs, q30 at d=30, L=18, and f=3 averaged 8.52 seconds with two-way SMT teams. This is about 9% faster than the 8-thread gate-batch run (9.30 seconds) and 2.11x faster than the base runner (17.98 seconds). Team execution reduced the gate phase from 6.32 to 5.93 seconds; remapping averaged 2.59 seconds. Team and non-team q24 runs produced identical amplitudes. Remaining team-SMT work: - Base adaptive block sizing on the number of teams rather than total logical threads. - Separate physical-core remap workers from the larger gate-thread pool. - Add pause/backoff behavior to the per-team spin barrier. - Add topology mocks and OpenMP CI coverage, and extend discovery beyond Linux. - Benchmark more CPU generations, circuit sizes, and affinity configurations. Signed-off-by: Xin Xie <xinpin@google.com>
On multi-node systems, allocating large state vectors on a single NUMA node can create memory controller and channel bottlenecks while remote memory channels remain underutilized. Apply page-level NUMA interleaving (MPOL_INTERLEAVE) via sysfs compute node discovery during state vector allocation on Linux. This balances memory traffic across available memory controllers, noticeably reducing gate execution times on multi-node architectures while gracefully falling back to standard allocation when NUMA support or multiple nodes are unavailable. Tested: bazel test //tests:all
In quantum circuits, phase-shift gates (diagonal unitaries such as Z, S, T, CZ, and Rz) commute with each other. In previous gate-batch planning, a skipped phase gate placed a hard causal barrier on all its qubits, blocking subsequent independent phase gates from packing into the current cache-resident block. Introduce soft barriers for phase gates so that skipped diagonal gates only bar subsequent non-diagonal operations, allowing compatible phase gates to commute past them. Additionally, expose a tunable eviction floor (-e) to enable streaming-optimized spans on wider vector extensions (e.g., floor 5 for AVX-512) and increase lookahead candidate exploration to 64 seeds by default. Tested: bazel test //tests:all
Run: local_script/evaluate_gate_seeds.sh circuits/circuit_q30 12 The harness compares lookahead seeds 0, 1, 8, and 64 with -f 3, -l 19, -e 5, and -x 1. Every run includes the empty and resident candidates. It reports total runtime, batches, executable gates, swaps, planning time, remapping time, and gate time.
The candidate-plan search accumulated a vector of fully materialized
GateBatchPlan objects and then picked the maximum, so peak planner memory
scaled with the number of seeds times the gates admitted by each. Buffer
the lightweight qubit sets instead and evaluate them as they are produced,
keeping only the incumbent best plan alive.
That removes both per-batch heap allocations and makes several helpers
redundant. CollectSeedGateIndices and PlanGateBatchFromQubitSet each had a
single caller and are folded in. GateBatchPlan::operator< became dead when
std::max_element went away.
The eviction floor depends only on constructor inputs, but was recomputed
inside AdditionalRemapSlotsFor, which runs per qubit of every gate of every
candidate seed. Precompute it once; chunk_qubits_ and min_eviction_floor_
were read by nothing else and are dropped as members.
PendingGate exposes `applied` directly rather than through an IsPending /
MarkApplied pair, which also removes a double negative from the three
planner scans.
Selection behavior is unchanged. Ties still resolve to the first candidate
(greedy, then resident, then seeded in circuit order) because the
comparison is strict; the added HasGates() guard fixes an edge case where
the default score of 0.0 would have suppressed a negative-scoring plan.
Net -22 lines. Planning time on q30 at -g 64 improved 7.7% (min of 5,
randomized interleaved A/B), attributable to the hoisted floor computation
and the two removed allocations.
Tested: 32-case equivalence vs parent across {q20,q24} x {-g 0,1,8,64} x
{-x 0,1} x {-l 19,12}, comparing batch/gate/swap counts and amplitudes.
All bit-identical. Built clean at -O3 -fopenmp -march=native -Wall -Wextra.
The file had grown so that its section banners no longer described their
contents. "Gate-batch diagnostics" spanned 433 lines and held 14 logic
members against 4 logging ones, making it the de facto home of remapping,
hot-qubit placement, fusion, and the entire block-execution engine. Nothing
signalled that ExecuteGateBatchOnBlocks lived under a banner saying
diagnostics.
Reorder into four tiers: plain records, then the two driver loops, then
phase implementations in the order the drivers call them, then diagnostics.
SimulateCircuit and ExecuteNextGateBatch are now adjacent, so the algorithm
reads top to bottom instead of being split across 394 lines.
Hoist GateBatchPlanner out of QSimGateBatchRunner. It referenced no
enclosing state and took every input through its constructor, so the
nesting only obscured that it is a self-contained unit; as a namespace-scope
template it is also testable without instantiating the runner. This shrinks
the runner class from 1044 to 795 lines.
Collect all eight Log* members into a trailing section, removing the nine
alternations between logging and logic that previously interrupted the
pipeline. Move IsDiagonalMatrix under the gate-utilities banner it belongs
to, and order SmtTeamBarrier last in the internal namespace, next to the
execution code that uses it.
No behavior change. Every construct moved verbatim; the only content edits
are the template header and <FP> parameterization the planner hoist
requires, plus its alias in the runner.
Tested: verified as a pure move by an indentation-independent line multiset
comparison against the parent, which reports 0 lines removed and exactly 3
added (template <typename FP>, and the two-line planner alias). 32-case
equivalence vs parent across {q20,q24} x {-g 0,1,8,64} x {-x 0,1} x
{-l 19,12}: all batch/gate/swap counts and amplitudes bit-identical. q30
smoke at 32 threads unchanged at 21 batches / 364 gates / 158 swaps. Built
clean at -O3 -fopenmp -march=native -Wall -Wextra.
ExecuteSmtBlockTeams was the longest function in the header at 57 lines,
and its non-OpenMP fallback needed seven (void) casts to silence unused
parameters for a body that could never run.
The fallback casts disappear by guarding the whole function with
#ifdef _OPENMP and moving the error path into its only caller, where
every parameter but team_thread_cpus is still consumed by the
independent-block path.
The team block loop walked a base index and recomputed a per-team offset,
burning empty trailing iterations whenever the base was still in range but
the derived block was not. Striding directly from team_id enumerates the
same block set with no dead iterations and drops the loop-invariant
active check out of the body. This is safe because the barrier inside
ExecuteGatesOnBlock is per-team, so a team's members always iterate
identically and inactive threads belong to wholly inactive teams.
The remaining index arithmetic moves into SmtTeamRole, a pure record with
no OpenMP in it, so the parallel region reads as four steps: assign role,
pin, barrier, run blocks.
Tested: New exhaustive check confirms AssignSmtTeamRole reproduces the
previous inline arithmetic across all 35360 combinations of
inner_threads <= 16, num_threads <= 64, and thread_id, and that active
teams tile blocks 0..n-1 exactly once for every inner_threads <= 8,
num_threads <= 64, num_blocks <= 40.
Tested: 72/72 bit-identical amplitudes and batch/gate/swap counts against
the parent commit over {circuit_q20, circuit_q24} x -g {0,8,64} x
-x {0,1} x (-t/-i) {4/1, 4/2, 8/2, 6/2, 6/3, 12/4}, covering one, two and
three teams and block counts not divisible by the team count.
Tested: 4/4 bit-identical on a build without -fopenmp; -i 2 still reports
that SMT teams require OpenMP.
Tested: Clean build at -O3 -fopenmp -march=native -Wall -Wextra.
Collect the OpenMP-dependent code into one shim block and reject SMT teams without OpenMP during parameter validation, before the state is touched. ExecuteSmtBlockTeams no longer needs an #ifdef.
The planner already computes the eviction floor once; derive the remap slot capacity from it instead of copying it into every candidate set.
Batch remapping and hot-qubit placement now share one helper that moves wanted qubits below a limit, taking victims from the top down. Hot qubits therefore fill the fixed zone in reverse order.
Drop the StartPhaseTimer/AccumulatePhaseSeconds helpers. Timings are always collected (a few calls per batch) and reported at verbosity > 1. Rename the Phase N section banners to Step N, since "phase" has a specific meaning in quantum computing.
They read the runner's state, partition, parameters and executable gates directly instead of taking them as arguments. ExecuteGatesOnBlock takes a block index, and a null barrier replaces the kCooperative template.
One AppendExecutableGates loop with a generic append lambda serves both modes. Qubit order is asserted rather than re-normalized, since raw batch gates and fused gates are already ascending.
std::swap_ranges measured consistently 1-3% slower swap passes.
Drop LogPlannedGateBatch, LogAppliedRemap and LogGateBatchSwapSummary, and keep the plan's swap count local to its score.
Behavior change for -f 0 and -f 1: - Before: -f 0 skipped the fuser and ran every raw gate as-is (no-fusion control mode); -f 1 ran the fuser, which still merges 1-qubit gates into neighboring 2-qubit gates (same result as -f 2 on q24). - After: both are rejected up front with "max_fused_size must be at least 2." Every batch now goes through the fuser. To restore a no-fusion mode, bring back the max_fused_size == 0 branch in FuseBatchGates that copied raw batch gates directly.
Drop the CpuCore struct, whose package and core ids were only map keys, and remove an unused variable in BuildTeamCpuOrder.
Treat a lone node "n" as the range "n-n" so one loop sets the mask. Same results as before on 20k generated node lists.
The detected interleave never applied: posix_memalign(64) returns a pointer 64 bytes past a page boundary with glibc, and mbind rejects a non-page-aligned start with EINVAL, so states kept the default policy. Allocate page-aligned and interleave across all nodes the process may use; the kernel limits the mask to the allowed memory nodes, so a single node is a no-op. This drops the sysfs node-list detection, including its preference for CPU nodes: memory-only nodes (CXL, HBM flat mode) are now included. Verified on WSL (one node): numa_maps shows interleave:0 for the state, where it previously showed default.
CooperativeFor only works together with the runner, which sets each SMT team member's lane before it runs a gate. As a separate header it let an app build the per-block simulator with any other For; with SMT teams enabled, both team members would then apply every gate to the whole block, giving wrong results without an error. Move it into run_qsim_gate_batch.h and have the runner derive its per-block simulator from Factory::Simulator by swapping in CooperativeFor (ReplaceFor). The runner drops its SeqSimulator template parameter and now matches QSimRunner<IO, Fuser, Factory>. Verified: identical output on the Mac and on WSL, including SMT teams; ReplaceFor checked for the Basic, SSE, AVX, AVX512 and NEON simulators.
qubit_remap.h, run_qsim_gate_batch.h and qsim_gate_batch.cc were missing the license block that every other qsim source file carries.
"Chunk" is not a standard SIMD term. Name the low amplitude-index bits that select a SIMD lane kLaneQubits/lane_qubits (log2 of the lane count), and the [re x N, im x N] unit a lane group. In qubit_remap.h, ChunkBitSwap, PartnerChunk, SwapChunkSpans and floats_per_chunk become GroupBitSwap, PartnerGroup, SwapGroupSpans and floats_per_group. No functional change: output is identical on the Mac, and every x86 state space compiles on WSL.
Drop the unused GateBatchQubitSet alias, replace the lane_qubits_ member with a static constexpr kLaneQubits (it always equals the state space's compile-time constant), and fix the ExecutableGate comment: gates are always fused now.
std::stable_sort keeps equal scores in ascending qubit order, so the comparator no longer needs an explicit tie-breaker.
Its only caller turned the returned list into a mask; return the mask directly and build the candidate list with std::iota.
|
/gemini review |
There was a problem hiding this comment.
Code Review
This pull request introduces a proposal-faithful gate-batched cache-local simulation runner (qsim_gate_batch) along with CPU topology discovery, qubit remapping, and NUMA interleaving support. Feedback on the changes highlights a critical buffer over-read vulnerability in the NUMA interleaving syscall, a performance degradation risk in the SMT barrier spin loop due to a lack of pause/yield instructions, and a missing validation check for num_threads when inner_threads is not active.
| bool ValidateParameters() const { | ||
| if (param_.max_fused_size < 2) { | ||
| IO::errorf("qsim_gate_batch: max_fused_size must be at least 2.\n"); | ||
| return false; | ||
| } | ||
| if (param_.inner_threads <= 1) return true; |
There was a problem hiding this comment.
If param_.inner_threads <= 1, ValidateParameters returns true early without validating param_.num_threads. If param_.num_threads is 0, this can cause undefined behavior or crashes in OpenMP parallel regions (e.g., #pragma omp parallel num_threads(0)). We should validate that param_.num_threads > 0 before the early return.
| bool ValidateParameters() const { | |
| if (param_.max_fused_size < 2) { | |
| IO::errorf("qsim_gate_batch: max_fused_size must be at least 2.\n"); | |
| return false; | |
| } | |
| if (param_.inner_threads <= 1) return true; | |
| bool ValidateParameters() const { | |
| if (param_.max_fused_size < 2) { | |
| IO::errorf("qsim_gate_batch: max_fused_size must be at least 2.\n"); | |
| return false; | |
| } | |
| if (param_.num_threads == 0) { | |
| IO::errorf("qsim_gate_batch: num_threads must be greater than 0.\n"); | |
| return false; | |
| } | |
| if (param_.inner_threads <= 1) return true; |
ValidateParameters returned early for inner_threads <= 1 without checking num_threads, so 0 could reach num_threads(0) in the OpenMP pragmas. Reject it up front, which also covers the SMT-team check. Pass mbind's maxnode as the mask's bit count, as the mbind(2) man page and current libnuma do, instead of the bit count plus one.
Remove casts that the surrounding arithmetic, comparison or parameter type already implies: the swap-loop bound, the plan score, the MatrixShuffle qubit count, the swap counter and the block address. The remaining casts feed printf-style arguments or are otherwise required. Use auto for locals whose initializer already states the type (GetTime timers, element references, shifts of typed constants, arithmetic on typed operands). Counters initialized from literals, loop indices and class members keep their explicit types.
Spell out that both arguments are physical positions, the layout's term for amplitude-index bits.
Rename PhysicalIndexOf to PhysicalAmplitudeIndex, so it no longer reads like PhysicalPositionOf: one maps an amplitude index, the other a qubit. Also remove the unused LogicalToPhysical accessor.
It holds one thread's place in the SMT team grid; AssignSmtTeam builds it. The local at the call site is now assignment instead of role.
The header is mostly QubitLayout; name it after that class.
gate_batch_runner.h holds the planner and pipeline as GateBatchRunner<IO, Fuser, Backend>. A backend owns the state and provides block sizing, swap passes, batch execution and synchronization; its Parameter and defaults (block_qubits, min_eviction_floor) extend the runner's. run_qsim_gate_batch.h keeps the CPU pieces (OpenMP blocks, SMT teams, CooperativeFor) as CpuGateBatchBackend, and QSimGateBatchRunner keeps its Run(param, factory, circuit, state) API. QubitSwap moves to qubit_layout.h with the layout that produces the swaps; qubit_remap.h is only the CPU swap kernel. No behavior change: same batches, swaps and amplitudes.
7cacbfb to
036ec95
Compare
block_qubits becomes tile_qubits, BlockPartition becomes TilePartition, and so on. "Tile" is the usual name for a cache-sized piece of work and leaves "block" free for CUDA thread blocks. The dependency sense of "blocked" (is_qubit_blocked_ and friends) is unchanged. API: Parameter::block_qubits is now tile_qubits. The -l flag is unchanged. No behavior change.
CudaGateBatchBackend plugs into GateBatchRunner: each batch runs as one kernel in which every CUDA thread block loads a shared-memory tile of 2^tile_qubits amplitudes, applies every fused gate, and writes it back once. tile_qubits (default 13) is an upper bound, lowered to fit the device's shared memory per block. qubit_remap_cuda.h is the CUDA counterpart of qubit_remap.h. Its PermuteLaneGroupsKernel also swaps lane bits while keeping global memory access coalesced, so lane_qubits and min_eviction_floor default to 0. qsim_gate_batch_cuda.cu is the benchmark app. qsim_gate_batch_cuda_test.cu compares full final states with qsim's SimulatorCUDA runner. util_cuda.h now includes <complex>, which it uses. Simulation time, best max_fused_size on each side (baseline f=2..6, gate batch f=2..5 with default options). RTX 4070 Laptop (8 GiB) under WSL2; q30 states spill about 1 GiB to system memory. circuit qsim CUDA gate batch CUDA speedup circuit_q24 0.215 s (f=4) 0.153 s (f=5) 1.40x q30_qaoa_sparse 9.18 s (f=3) 2.95 s (f=3) 3.11x q30_qaoa_dense 34.1 s (f=6) 15.3 s (f=3) 2.23x q30_qft 25.4 s (f=6) 9.15 s (f=3) 2.77x circuit_q30 43.1 s (f=4) 22.3 s (f=3) 1.93x Gate batch is within about 5% of its best at any f from 2 to 5. On circuit_q30, the 48 batch kernels take 64% of the time and the swap passes 36%.
build_gate_batch.sh
Summary
This PR adds an experimental gate-batched, cache-local state-vector runner for the NEON, SSE4.2, AVX2, and AVX-512 backends.
The runner groups causally executable gates, remaps their logical qubits into a cache-sized physical-qubit block, fuses the selected gates once, and applies them to every state block. It also includes adaptive block sizing for small circuits, optional topology-aware SMT teams, and Linux NUMA interleaving for large state-vector allocations.
Algorithm
The state vector is divided into contiguous blocks of
2^Lamplitudes. Each block represents the state space ofLphysical qubits, while the remainingn-Lqubits have a fixed bit assignment.While circuit gates remain:
Lblock-local physical positions.The default execution path follows the proposal's three-level loop:
One worker normally handles one contiguous block and applies all executable gates using a sequential SIMD simulator. The simulator's existing architecture-specific kernel retains the innermost local-index loop.
On Linux, optional team-SMT execution assigns the two logical CPUs belonging to one physical core to the same block. Each sibling processes a disjoint portion of the simulator kernel and the pair synchronizes between gates. CPU topology is discovered at runtime from the process affinity and sysfs; workers are pinned explicitly rather than assuming adjacent OpenMP thread IDs are siblings.
Lis treated as an upper bound and may be reduced for small circuits to expose enough blocks for parallel execution. Frequently used logical qubits are initially placed in a fixed block-local region. SIMD lane positions remain pinned because remapping moves whole SIMD chunks rather than individual lanes.QubitMappedStateretains the final logical-to-physical mapping so callers can avoid a final identity-restoration pass.On multi-node Linux systems, large state-vector allocations use
MPOL_INTERLEAVEwhen multiple compute nodes are available. Unsupported platforms, single-node systems, and failed NUMA-policy calls retain the normal aligned-allocation behavior.Performance
Times are in seconds and lower is better. Speedup is base-runner time divided by gate-batch-runner time. Results should only be compared within a row/platform.
Full
circuit_q30measurements — August 19, 2026These measurements use the full 1617-raw-gate circuit and compare the best configuration measured for each runner. They predate the team-SMT and NUMA-interleaving commits, so they should not be treated as measurements of the final PR head.
f=4, SMT on (192 threads)f=4, L=18, SMT off (96 threads)f=4, SMT on (192 threads)f=3, L=18, ordinary SMT (192 threads)f=4, SMT on (96 threads)f=3, L=17, SMT off (48 threads)f=3, SMT on (96 threads)f=3, L=17, SMT off (48 threads)f=4, 72 physical coresf=3, L=18, 72 physical coresThe results are architecture- and configuration-dependent. The largest measured gain is on AMD Turin with AVX-512. Axion shows a modest gain, while the best large Intel C4 base configuration is slightly faster than the best gate-batch configuration. Remapping remains a material part of total runtime, so improved gate locality does not automatically translate into an end-to-end speedup.
Intel C4 team-SMT experiment — August 24, 2026
This focused experiment used an 8-physical-core/16-vCPU
c4-standard-16Intel Xeon Platinum 8581C VM. Each core has a private 2 MiB L2, so the gate-batch runs usedL=18. The circuit was limited withd=30(502 raw gates), and each result is the average of two runs withf=3.Runtime topology discovery produced the sibling pairs
(0,8),(1,9), ...,(7,15). Team SMT was about 9% faster end-to-end than the 8-worker gate-batch run. Ordinary SMT made gate application slower because two siblings processed separate L2-sized blocks; assigning both siblings to one shared block recovered that loss and improved the gate phase.This short
d=30experiment is not directly comparable to the full-circuit table above. A full latest-head sweep, including NUMA interleaving and team SMT on the large C4/C4D machines, remains future work.Validation
bazel test //tests:allon Linux.1e-9in regression testing.Remaining work