Skip to content

Add gate-batched cache-local runner across SIMD backends - #1092

Draft
xxie24 wants to merge 51 commits into
quantumlib:mainfrom
xxie24:blocked-v2
Draft

xxie24 wants to merge 51 commits into
quantumlib:mainfrom
xxie24:blocked-v2

Conversation

@xxie24

@xxie24 xxie24 commented Jul 23, 2026 •

Copy link
Copy Markdown
Contributor

build_gate_batch.sh

Summary

This PR adds an experimental gate-batched, cache-local state-vector runner for the NEON, SSE4.2, AVX2, and AVX-512 backends.

The runner groups causally executable gates, remaps their logical qubits into a cache-sized physical-qubit block, fuses the selected gates once, and applies them to every state block. It also includes adaptive block sizing for small circuits, optional topology-aware SMT teams, and Linux NUMA interleaving for large state-vector allocations.

Algorithm

The state vector is divided into contiguous blocks of 2^L amplitudes. Each block represents the state space of L physical qubits, while the remaining n-L qubits have a fixed bit assignment.

While circuit gates remain:

  1. Select a causally executable gate batch whose logical qubits fit in the L block-local physical positions.
  2. Remap the required logical qubits into those positions. All disjoint qubit-position swaps are applied together in one state-vector pass.
  3. Fuse the selected gates once to produce the executable gates for the batch.
  4. Apply every executable gate to every state block while that block is cache-resident.

The default execution path follows the proposal's three-level loop:

while (gates_remain) {
  const auto plan = PlanNextGateBatch();
  ApplySwapsToState(plan);
  const auto executable_gates = FuseBatchGates(plan);

  #pragma omp parallel for schedule(dynamic, 1)
  for (int64_t block_index = 0;
       block_index < num_state_blocks;
       ++block_index) {
    auto block = GetStateBlock(block_index);

    for (const auto& gate : executable_gates) {
      sequential_simulator.ApplyGate(gate, block);
    }
  }
}

One worker normally handles one contiguous block and applies all executable gates using a sequential SIMD simulator. The simulator's existing architecture-specific kernel retains the innermost local-index loop.

On Linux, optional team-SMT execution assigns the two logical CPUs belonging to one physical core to the same block. Each sibling processes a disjoint portion of the simulator kernel and the pair synchronizes between gates. CPU topology is discovered at runtime from the process affinity and sysfs; workers are pinned explicitly rather than assuming adjacent OpenMP thread IDs are siblings.

L is treated as an upper bound and may be reduced for small circuits to expose enough blocks for parallel execution. Frequently used logical qubits are initially placed in a fixed block-local region. SIMD lane positions remain pinned because remapping moves whole SIMD chunks rather than individual lanes.

QubitMappedState retains the final logical-to-physical mapping so callers can avoid a final identity-restoration pass.

On multi-node Linux systems, large state-vector allocations use MPOL_INTERLEAVE when multiple compute nodes are available. Unsupported platforms, single-node systems, and failed NUMA-policy calls retain the normal aligned-allocation behavior.

Performance

Times are in seconds and lower is better. Speedup is base-runner time divided by gate-batch-runner time. Results should only be compared within a row/platform.

Full circuit_q30 measurements — August 19, 2026

These measurements use the full 1617-raw-gate circuit and compare the best configuration measured for each runner. They predate the team-SMT and NUMA-interleaving commits, so they should not be treated as measurements of the final PR head.

Platform SIMD backend Base best configuration Base (s) Gate-batch best configuration Gate-batch total (s) Gate application (s) Remap (s) Speedup
Google C4 / Intel AVX-512 f=4, SMT on (192 threads) 7.96 f=4, L=18, SMT off (96 threads) 8.18 4.12 4.05 0.97x
Google C4 / Intel SSE4.2 f=4, SMT on (192 threads) 10.67 f=3, L=18, ordinary SMT (192 threads) 11.97 7.52 4.45 0.89x
Google C4D / AMD Turin AVX-512 f=4, SMT on (96 threads) 10.53 f=3, L=17, SMT off (48 threads) 6.13 3.40 2.73 1.72x
Google C4D / AMD Turin SSE4.2 f=3, SMT on (96 threads) 17.39 f=3, L=17, SMT off (48 threads) 16.14 12.67 3.46 1.08x
Google C4A / Axion NEON f=4, 72 physical cores 11.00 f=3, L=18, 72 physical cores 9.99 6.82 3.17 1.10x

The results are architecture- and configuration-dependent. The largest measured gain is on AMD Turin with AVX-512. Axion shows a modest gain, while the best large Intel C4 base configuration is slightly faster than the best gate-batch configuration. Remapping remains a material part of total runtime, so improved gate locality does not automatically translate into an end-to-end speedup.

Intel C4 team-SMT experiment — August 24, 2026

This focused experiment used an 8-physical-core/16-vCPU c4-standard-16 Intel Xeon Platinum 8581C VM. Each core has a private 2 MiB L2, so the gate-batch runs used L=18. The circuit was limited with d=30 (502 raw gates), and each result is the average of two runs with f=3.

Configuration CPU allocation Total (s) Gate application (s) Remap (s) Speedup vs 8-core base
Base runner 8 physical cores 17.980 — — 1.00x
Gate-batch, physical only 8 independent workers 9.297 6.319 2.976 1.93x
Gate-batch, ordinary SMT 16 independent workers 9.257 6.762 2.493 1.94x
Gate-batch, team SMT 8 two-sibling teams 8.523 5.930 2.592 2.11x

Runtime topology discovery produced the sibling pairs (0,8), (1,9), ..., (7,15). Team SMT was about 9% faster end-to-end than the 8-worker gate-batch run. Ordinary SMT made gate application slower because two siblings processed separate L2-sized blocks; assigning both siblings to one shared block recovered that loss and improved the gate phase.

This short d=30 experiment is not directly comparable to the full-circuit table above. A full latest-head sweep, including NUMA interleaving and team SMT on the large C4/C4D machines, remains future work.

Validation

  • Compared logical amplitudes/full state vectors with the existing qsim runner, accounting for the retained qubit mapping.
  • Tested multiple block sizes and fusion sizes.
  • Tested NEON, SSE4.2, AVX2, and AVX-512 builds.
  • Verified generated NEON and AVX-512 SIMD instructions.
  • Verified identical q24 amplitudes between physical-only and team-SMT execution on Intel C4.
  • Verified runtime Intel sibling discovery and explicit worker pinning.
  • Verified that unavailable/restricted SMT siblings are rejected rather than paired incorrectly.
  • Tested the NUMA allocation change with bazel test //tests:all on Linux.
  • Observed numerical differences around 1e-9 in regression testing.

Remaining work

  • Re-run the full q30 performance matrix at the latest PR head.
  • Base adaptive block sizing on the number of block teams rather than total logical threads.
  • Separate physical-core remap workers from the larger gate-thread pool.
  • Add pause/backoff behavior to the team synchronization barrier.
  • Add mocked topology tests, OpenMP CI coverage, and topology discovery beyond Linux.

Introduce an experimental runner that executes circuits as gate batches
over cache-sized state blocks. Each batch selects causally executable
raw gates, maps their logical qubits into block-local physical positions,
fuses the mapped gates, and executes them across all state blocks in
parallel.

Bound planning cost by comparing greedy, resident, and upcoming-gate-
seeded candidates. Score candidates by circuit progress while accounting
for the number of required state-vector swap passes.

Apply each set of disjoint qubit-position swaps in one state-vector pass.
Keep architecture-specific SIMD lane positions pinned so remapping moves
whole chunks on NEON, SSE, AVX2, and AVX-512 layouts.

Add QubitMappedState to retain the logical-to-physical qubit layout
across execution. The mapped-state API returns without restoring
canonical qubit order, while the raw-state compatibility API restores
identity before returning.

Add a command-line runner for exercising the gate-batch implementation
and register the runner and supporting headers with the Bazel build.
Register a Makefile target that builds both the gate-batch runner and
the existing base runner.

Signed-off-by: Xin Xie <xinpin@google.com>
@github-actions github-actions Bot added the size: XL lines changed >1000 label Jul 23, 2026

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a gate-batched cache-local simulation runner (QSimGateBatchRunner) along with qubit remapping and layout utilities (QubitLayout, QubitMappedState, and ApplyBitPairSwaps) to optimize simulation performance. It also adds a benchmark application (qsim_gate_batch) and updates the build configurations. Feedback was provided to relax an assertion in run_qsim_gate_batch.h that would otherwise cause failures for small circuits where the number of qubits is less than the SIMD chunk size.

Comment thread lib/run_qsim_gate_batch.h Outdated
layout_(layout),
gate_batch_planner_(partition_.num_state_qubits,
partition_.block_qubits, chunk_qubits_) {
assert(partition_.block_qubits >= chunk_qubits_);

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

The assertion assert(partition_.block_qubits >= chunk_qubits_); will fail for small circuits where the number of qubits is less than the SIMD chunk size (e.g., a 2-qubit circuit on an AVX-512 backend where chunk_qubits_ is 4). Since cache blocking is not required for circuits that fit entirely within a single SIMD chunk, we should relax this assertion to allow block_qubits to be smaller than chunk_qubits_ when the total number of state qubits is also smaller than chunk_qubits_.

Suggested change
assert(partition_.block_qubits >= chunk_qubits_);
assert(partition_.block_qubits >= std::min(chunk_qubits_, partition_.num_state_qubits));

@xxie24
xxie24 marked this pull request as draft July 23, 2026 21:12
@mhucka

mhucka commented Aug 15, 2026

Copy link
Copy Markdown
Collaborator

@xxie24 Thank you for this work. I know it's still a draft PR, but I wanted to mention that a recent failure in the CI workflow has now been fixed. I took the liberty of updating your branch to have the latest changes (I hope that's okay), and now the PR passes all the Ci tests.

@xxie24

xxie24 commented Aug 17, 2026

Copy link
Copy Markdown
Contributor Author

@xxie24 Thank you for this work. I know it's still a draft PR, but I wanted to mention that a recent failure in the CI workflow has now been fixed. I took the liberty of updating your branch to have the latest changes (I hope that's okay), and now the PR passes all the Ci tests.

cool. thanks.

Xin Xie added 4 commits August 16, 2026 21:57
Dynamic scheduling with small chunk sizes caused excessive atomic compare-and-swap
synchronization contention on the shared work-sharing counter across high core counts.
Switching to schedule(static) eliminates lock contention and preserves sequential
hardware stream prefetch performance.
Xin Xie added 2 commits August 24, 2026 07:15
Allow logical CPUs that share a physical core to cooperate on one cache-resident state block. Add the -i/inner_threads option while preserving the existing one-thread-per-block behavior by default.

Design decisions:

- Split each simulator kernel's iteration range between team lanes with CooperativeFor, and synchronize the lanes after every gate because the next gate consumes the preceding gate's output.
- Keep independent-block and SMT-team scheduling explicit, while sharing the block-local gate loop through a compile-time cooperative helper. The non-SMT path therefore pays no barrier or runtime branch overhead.
- Discover Linux CPU topology from the process affinity and sysfs package/core identifiers. Order logical CPUs by physical core, pin every team lane explicitly, and reject configurations where the requested siblings are unavailable instead of silently pairing unrelated cores.
- Keep non-OpenMP builds working and report a clear error if SMT teams are requested without OpenMP.
- Register the cooperative loop and CPU topology headers in the Bazel library.

On a Google C4 Intel Xeon Platinum 8581C with 8 physical cores and 16 vCPUs, q30 at d=30, L=18, and f=3 averaged 8.52 seconds with two-way SMT teams. This is about 9% faster than the 8-thread gate-batch run (9.30 seconds) and 2.11x faster than the base runner (17.98 seconds). Team execution reduced the gate phase from 6.32 to 5.93 seconds; remapping averaged 2.59 seconds. Team and non-team q24 runs produced identical amplitudes.

Remaining team-SMT work:

- Base adaptive block sizing on the number of teams rather than total logical threads.
- Separate physical-core remap workers from the larger gate-thread pool.
- Add pause/backoff behavior to the per-team spin barrier.
- Add topology mocks and OpenMP CI coverage, and extend discovery beyond Linux.
- Benchmark more CPU generations, circuit sizes, and affinity configurations.

Signed-off-by: Xin Xie <xinpin@google.com>
On multi-node systems, allocating large state vectors on a single NUMA
node can create memory controller and channel bottlenecks while remote
memory channels remain underutilized.

Apply page-level NUMA interleaving (MPOL_INTERLEAVE) via sysfs compute
node discovery during state vector allocation on Linux. This balances
memory traffic across available memory controllers, noticeably reducing
gate execution times on multi-node architectures while gracefully falling
back to standard allocation when NUMA support or multiple nodes are
unavailable.

Tested: bazel test //tests:all
xinpin added 15 commits September 15, 2026 15:00
In quantum circuits, phase-shift gates (diagonal unitaries such as Z,
S, T, CZ, and Rz) commute with each other. In previous gate-batch
planning, a skipped phase gate placed a hard causal barrier on all its
qubits, blocking subsequent independent phase gates from packing into the
current cache-resident block.

Introduce soft barriers for phase gates so that skipped diagonal gates
only bar subsequent non-diagonal operations, allowing compatible phase
gates to commute past them. Additionally, expose a tunable eviction
floor (-e) to enable streaming-optimized spans on wider vector extensions
(e.g., floor 5 for AVX-512) and increase lookahead candidate exploration
to 64 seeds by default.

Tested: bazel test //tests:all
Run: local_script/evaluate_gate_seeds.sh circuits/circuit_q30 12

The harness compares lookahead seeds 0, 1, 8, and 64 with -f 3, -l 19, -e 5, and -x 1. Every run includes the empty and resident candidates. It reports total runtime, batches, executable gates, swaps, planning time, remapping time, and gate time.
The candidate-plan search accumulated a vector of fully materialized
GateBatchPlan objects and then picked the maximum, so peak planner memory
scaled with the number of seeds times the gates admitted by each. Buffer
the lightweight qubit sets instead and evaluate them as they are produced,
keeping only the incumbent best plan alive.

That removes both per-batch heap allocations and makes several helpers
redundant. CollectSeedGateIndices and PlanGateBatchFromQubitSet each had a
single caller and are folded in. GateBatchPlan::operator< became dead when
std::max_element went away.

The eviction floor depends only on constructor inputs, but was recomputed
inside AdditionalRemapSlotsFor, which runs per qubit of every gate of every
candidate seed. Precompute it once; chunk_qubits_ and min_eviction_floor_
were read by nothing else and are dropped as members.

PendingGate exposes `applied` directly rather than through an IsPending /
MarkApplied pair, which also removes a double negative from the three
planner scans.

Selection behavior is unchanged. Ties still resolve to the first candidate
(greedy, then resident, then seeded in circuit order) because the
comparison is strict; the added HasGates() guard fixes an edge case where
the default score of 0.0 would have suppressed a negative-scoring plan.

Net -22 lines. Planning time on q30 at -g 64 improved 7.7% (min of 5,
randomized interleaved A/B), attributable to the hoisted floor computation
and the two removed allocations.

Tested: 32-case equivalence vs parent across {q20,q24} x {-g 0,1,8,64} x
{-x 0,1} x {-l 19,12}, comparing batch/gate/swap counts and amplitudes.
All bit-identical. Built clean at -O3 -fopenmp -march=native -Wall -Wextra.
The file had grown so that its section banners no longer described their
contents. "Gate-batch diagnostics" spanned 433 lines and held 14 logic
members against 4 logging ones, making it the de facto home of remapping,
hot-qubit placement, fusion, and the entire block-execution engine. Nothing
signalled that ExecuteGateBatchOnBlocks lived under a banner saying
diagnostics.

Reorder into four tiers: plain records, then the two driver loops, then
phase implementations in the order the drivers call them, then diagnostics.
SimulateCircuit and ExecuteNextGateBatch are now adjacent, so the algorithm
reads top to bottom instead of being split across 394 lines.

Hoist GateBatchPlanner out of QSimGateBatchRunner. It referenced no
enclosing state and took every input through its constructor, so the
nesting only obscured that it is a self-contained unit; as a namespace-scope
template it is also testable without instantiating the runner. This shrinks
the runner class from 1044 to 795 lines.

Collect all eight Log* members into a trailing section, removing the nine
alternations between logging and logic that previously interrupted the
pipeline. Move IsDiagonalMatrix under the gate-utilities banner it belongs
to, and order SmtTeamBarrier last in the internal namespace, next to the
execution code that uses it.

No behavior change. Every construct moved verbatim; the only content edits
are the template header and <FP> parameterization the planner hoist
requires, plus its alias in the runner.

Tested: verified as a pure move by an indentation-independent line multiset
comparison against the parent, which reports 0 lines removed and exactly 3
added (template <typename FP>, and the two-line planner alias). 32-case
equivalence vs parent across {q20,q24} x {-g 0,1,8,64} x {-x 0,1} x
{-l 19,12}: all batch/gate/swap counts and amplitudes bit-identical. q30
smoke at 32 threads unchanged at 21 batches / 364 gates / 158 swaps. Built
clean at -O3 -fopenmp -march=native -Wall -Wextra.
ExecuteSmtBlockTeams was the longest function in the header at 57 lines,
and its non-OpenMP fallback needed seven (void) casts to silence unused
parameters for a body that could never run.

The fallback casts disappear by guarding the whole function with
#ifdef _OPENMP and moving the error path into its only caller, where
every parameter but team_thread_cpus is still consumed by the
independent-block path.

The team block loop walked a base index and recomputed a per-team offset,
burning empty trailing iterations whenever the base was still in range but
the derived block was not. Striding directly from team_id enumerates the
same block set with no dead iterations and drops the loop-invariant
active check out of the body. This is safe because the barrier inside
ExecuteGatesOnBlock is per-team, so a team's members always iterate
identically and inactive threads belong to wholly inactive teams.

The remaining index arithmetic moves into SmtTeamRole, a pure record with
no OpenMP in it, so the parallel region reads as four steps: assign role,
pin, barrier, run blocks.

Tested: New exhaustive check confirms AssignSmtTeamRole reproduces the
previous inline arithmetic across all 35360 combinations of
inner_threads <= 16, num_threads <= 64, and thread_id, and that active
teams tile blocks 0..n-1 exactly once for every inner_threads <= 8,
num_threads <= 64, num_blocks <= 40.
Tested: 72/72 bit-identical amplitudes and batch/gate/swap counts against
the parent commit over {circuit_q20, circuit_q24} x -g {0,8,64} x
-x {0,1} x (-t/-i) {4/1, 4/2, 8/2, 6/2, 6/3, 12/4}, covering one, two and
three teams and block counts not divisible by the team count.
Tested: 4/4 bit-identical on a build without -fopenmp; -i 2 still reports
that SMT teams require OpenMP.
Tested: Clean build at -O3 -fopenmp -march=native -Wall -Wextra.
Collect the OpenMP-dependent code into one shim block and reject SMT
teams without OpenMP during parameter validation, before the state is
touched. ExecuteSmtBlockTeams no longer needs an #ifdef.
The planner already computes the eviction floor once; derive the remap
slot capacity from it instead of copying it into every candidate set.
Batch remapping and hot-qubit placement now share one helper that moves
wanted qubits below a limit, taking victims from the top down. Hot
qubits therefore fill the fixed zone in reverse order.
Drop the StartPhaseTimer/AccumulatePhaseSeconds helpers. Timings are
always collected (a few calls per batch) and reported at verbosity > 1.
Rename the Phase N section banners to Step N, since "phase" has a
specific meaning in quantum computing.
They read the runner's state, partition, parameters and executable gates
directly instead of taking them as arguments. ExecuteGatesOnBlock takes a
block index, and a null barrier replaces the kCooperative template.
One AppendExecutableGates loop with a generic append lambda serves both
modes. Qubit order is asserted rather than re-normalized, since raw
batch gates and fused gates are already ascending.
Xin Xie added 5 commits September 26, 2026 20:28
std::swap_ranges measured consistently 1-3% slower swap passes.
Drop LogPlannedGateBatch, LogAppliedRemap and LogGateBatchSwapSummary,
and keep the plan's swap count local to its score.
Behavior change for -f 0 and -f 1:
- Before: -f 0 skipped the fuser and ran every raw gate as-is (no-fusion
  control mode); -f 1 ran the fuser, which still merges 1-qubit gates
  into neighboring 2-qubit gates (same result as -f 2 on q24).
- After: both are rejected up front with "max_fused_size must be at
  least 2." Every batch now goes through the fuser.

To restore a no-fusion mode, bring back the max_fused_size == 0 branch
in FuseBatchGates that copied raw batch gates directly.
Drop the CpuCore struct, whose package and core ids were only map keys,
and remove an unused variable in BuildTeamCpuOrder.
Treat a lone node "n" as the range "n-n" so one loop sets the mask.
Same results as before on 20k generated node lists.
Xin Xie added 5 commits September 26, 2026 21:48
The detected interleave never applied: posix_memalign(64) returns a
pointer 64 bytes past a page boundary with glibc, and mbind rejects a
non-page-aligned start with EINVAL, so states kept the default policy.

Allocate page-aligned and interleave across all nodes the process may
use; the kernel limits the mask to the allowed memory nodes, so a single
node is a no-op. This drops the sysfs node-list detection, including its
preference for CPU nodes: memory-only nodes (CXL, HBM flat mode) are now
included.

Verified on WSL (one node): numa_maps shows interleave:0 for the state,
where it previously showed default.
CooperativeFor only works together with the runner, which sets each SMT
team member's lane before it runs a gate. As a separate header it let an
app build the per-block simulator with any other For; with SMT teams
enabled, both team members would then apply every gate to the whole
block, giving wrong results without an error.

Move it into run_qsim_gate_batch.h and have the runner derive its
per-block simulator from Factory::Simulator by swapping in CooperativeFor
(ReplaceFor). The runner drops its SeqSimulator template parameter and
now matches QSimRunner<IO, Fuser, Factory>.

Verified: identical output on the Mac and on WSL, including SMT teams;
ReplaceFor checked for the Basic, SSE, AVX, AVX512 and NEON simulators.
qubit_remap.h, run_qsim_gate_batch.h and qsim_gate_batch.cc were missing
the license block that every other qsim source file carries.
"Chunk" is not a standard SIMD term. Name the low amplitude-index bits
that select a SIMD lane kLaneQubits/lane_qubits (log2 of the lane
count), and the [re x N, im x N] unit a lane group. In qubit_remap.h,
ChunkBitSwap, PartnerChunk, SwapChunkSpans and floats_per_chunk become
GroupBitSwap, PartnerGroup, SwapGroupSpans and floats_per_group.

No functional change: output is identical on the Mac, and every x86
state space compiles on WSL.
Drop the unused GateBatchQubitSet alias, replace the lane_qubits_ member
with a static constexpr kLaneQubits (it always equals the state space's
compile-time constant), and fix the ExecutableGate comment: gates are
always fused now.
Xin Xie added 3 commits September 26, 2026 22:30
std::stable_sort keeps equal scores in ascending qubit order, so the
comparator no longer needs an explicit tie-breaker.
Its only caller turned the returned list into a mask; return the mask
directly and build the candidate list with std::iota.
@xxie24

xxie24 commented Sep 27, 2026

Copy link
Copy Markdown
Contributor Author

/gemini review

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a proposal-faithful gate-batched cache-local simulation runner (qsim_gate_batch) along with CPU topology discovery, qubit remapping, and NUMA interleaving support. Feedback on the changes highlights a critical buffer over-read vulnerability in the NUMA interleaving syscall, a performance degradation risk in the SMT barrier spin loop due to a lack of pause/yield instructions, and a missing validation check for num_threads when inner_threads is not active.

Comment thread lib/vectorspace.h
Comment thread lib/run_qsim_gate_batch.h
Comment thread lib/run_qsim_gate_batch.h Outdated
Comment on lines +1054 to +1059
bool ValidateParameters() const {
if (param_.max_fused_size < 2) {
IO::errorf("qsim_gate_batch: max_fused_size must be at least 2.\n");
return false;
}
if (param_.inner_threads <= 1) return true;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

high

If param_.inner_threads <= 1, ValidateParameters returns true early without validating param_.num_threads. If param_.num_threads is 0, this can cause undefined behavior or crashes in OpenMP parallel regions (e.g., #pragma omp parallel num_threads(0)). We should validate that param_.num_threads > 0 before the early return.

Suggested change
bool ValidateParameters() const {
if (param_.max_fused_size < 2) {
IO::errorf("qsim_gate_batch: max_fused_size must be at least 2.\n");
return false;
}
if (param_.inner_threads <= 1) return true;
bool ValidateParameters() const {
if (param_.max_fused_size < 2) {
IO::errorf("qsim_gate_batch: max_fused_size must be at least 2.\n");
return false;
}
if (param_.num_threads == 0) {
IO::errorf("qsim_gate_batch: num_threads must be greater than 0.\n");
return false;
}
if (param_.inner_threads <= 1) return true;

Xin Xie added 8 commits September 27, 2026 09:08
ValidateParameters returned early for inner_threads <= 1 without checking
num_threads, so 0 could reach num_threads(0) in the OpenMP pragmas.
Reject it up front, which also covers the SMT-team check.

Pass mbind's maxnode as the mask's bit count, as the mbind(2) man page
and current libnuma do, instead of the bit count plus one.
Remove casts that the surrounding arithmetic, comparison or parameter
type already implies: the swap-loop bound, the plan score, the
MatrixShuffle qubit count, the swap counter and the block address. The
remaining casts feed printf-style arguments or are otherwise required.

Use auto for locals whose initializer already states the type (GetTime
timers, element references, shifts of typed constants, arithmetic on
typed operands). Counters initialized from literals, loop indices and
class members keep their explicit types.
Spell out that both arguments are physical positions, the layout's term
for amplitude-index bits.
Rename PhysicalIndexOf to PhysicalAmplitudeIndex, so it no longer reads
like PhysicalPositionOf: one maps an amplitude index, the other a qubit.
Also remove the unused LogicalToPhysical accessor.
It holds one thread's place in the SMT team grid; AssignSmtTeam builds
it. The local at the call site is now assignment instead of role.
The header is mostly QubitLayout; name it after that class.
gate_batch_runner.h holds the planner and pipeline as
GateBatchRunner<IO, Fuser, Backend>. A backend owns the state and provides
block sizing, swap passes, batch execution and synchronization; its
Parameter and defaults (block_qubits, min_eviction_floor) extend the
runner's.

run_qsim_gate_batch.h keeps the CPU pieces (OpenMP blocks, SMT teams,
CooperativeFor) as CpuGateBatchBackend, and QSimGateBatchRunner keeps its
Run(param, factory, circuit, state) API. QubitSwap moves to qubit_layout.h
with the layout that produces the swaps; qubit_remap.h is only the CPU
swap kernel.

No behavior change: same batches, swaps and amplitudes.
@xxie24
xxie24 force-pushed the blocked-v2 branch 2 times, most recently from 7cacbfb to 036ec95 Compare September 28, 2026 00:20
Xin Xie added 2 commits September 27, 2026 17:38
block_qubits becomes tile_qubits, BlockPartition becomes TilePartition,
and so on. "Tile" is the usual name for a cache-sized piece of work and
leaves "block" free for CUDA thread blocks. The dependency sense of
"blocked" (is_qubit_blocked_ and friends) is unchanged.

API: Parameter::block_qubits is now tile_qubits. The -l flag is
unchanged. No behavior change.
CudaGateBatchBackend plugs into GateBatchRunner: each batch runs as one
kernel in which every CUDA thread block loads a shared-memory tile of
2^tile_qubits amplitudes, applies every fused gate, and writes it back
once. tile_qubits (default 13) is an upper bound, lowered to fit the
device's shared memory per block. qubit_remap_cuda.h is the CUDA
counterpart of qubit_remap.h. Its PermuteLaneGroupsKernel also swaps lane
bits while keeping global memory access coalesced, so lane_qubits and
min_eviction_floor default to 0.

qsim_gate_batch_cuda.cu is the benchmark app. qsim_gate_batch_cuda_test.cu
compares full final states with qsim's SimulatorCUDA runner.
util_cuda.h now includes <complex>, which it uses.

Simulation time, best max_fused_size on each side (baseline f=2..6,
gate batch f=2..5 with default options). RTX 4070 Laptop (8 GiB) under
WSL2; q30 states spill about 1 GiB to system memory.

  circuit            qsim CUDA       gate batch CUDA   speedup
  circuit_q24         0.215 s (f=4)   0.153 s (f=5)    1.40x
  q30_qaoa_sparse     9.18 s  (f=3)   2.95 s  (f=3)    3.11x
  q30_qaoa_dense     34.1 s   (f=6)  15.3 s   (f=3)    2.23x
  q30_qft            25.4 s   (f=6)   9.15 s  (f=3)    2.77x
  circuit_q30        43.1 s   (f=4)  22.3 s   (f=3)    1.93x

Gate batch is within about 5% of its best at any f from 2 to 5. On
circuit_q30, the 48 batch kernels take 64% of the time and the swap
passes 36%.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size: XL lines changed >1000

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants