Skip to content

feat(bp): BP decoder memory-layout, scheduling, and OSD truncation optimizations (2.2x-5.25x) - #294

Merged
aria-googler merged 14 commits into
feat/bp-decoderfrom
feat/bp-decoder-optimizations
Sep 22, 2026
Merged

aria-googler merged 14 commits into
feat/bp-decoderfrom
feat/bp-decoder-optimizations

Conversation

@aria-googler

@aria-googler aria-googler commented Aug 8, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Performance and robustness work on the Tesseract-BP decoder. By re-architecting BP state into a contiguous 1D interleaved memory layout, switching to horizontal layered serial scheduling, adding stochastic check-node schedule permutations, and introducing OSD matrix column truncation, the decoder achieves a 2.2x–5.25x end-to-end throughput speedup with no degradation in logical error rate.

This branch has also been synced with main and carries the fixes required to keep both the Bazel and CMake builds green.


1. BP engine optimizations

  1. Flat interleaved 1D array layout (bab6616)

    • BP state moved into a single 64-byte-aligned array indexed posteriors_flat[v * 16 + b], interleaving BP_BATCH_SIZE = 16 shot streams side by side.
    • Previously the 16 streams lived in 16 separate heap allocations, so reading variable $V_i$ across the batch touched 16 disjoint cache lines (1,024 bytes transferred to consume 64). Now it is a single aligned SIMD load.
  2. Horizontal layered serial scheduling (b71c8fa)

    • Reimplemented the solver around $L'v = Q{c,v} + R'{c,v}$ with immediate posterior updates, eliminating the explicit 2D check-to-variable message table $R{c,v}$ and roughly halving BP memory traffic.
  3. Stochastic check-node permutation (eac413a)

    • Fisher–Yates shuffle of the check-node visit order driven by a Xorshift64 PRNG, breaking the static LDPC trapping sets that stall symmetric topologies (Bivariate Bicycle, Color codes) without disturbing SIMD alignment.
    • Exposed as --random-schedule.
  4. OSD matrix column truncation (3e01a32)

    • The dense OSD parity-check matrix is now bounded to min(num_errors, num_detectors * osd_truncation_factor) columns, sorted by posterior reliability, instead of materializing all $N$ error mechanisms.
    • New --osd-truncation-factor CLI flag and BPParams.osd_truncation_factor binding; default 1.2, 0.0 disables.
    • On a $d=11$ surface code this shrinks the dense matrix from 50,460 to 2,904 columns (15.3 MB → 877 KB).
  5. CLI telemetry + benchmark harness (565ae1f)

    • --threads N previously reported summed CPU thread time as wall-clock; now reports true elapsed duration.
    • Added devtools/benchmark.sh covering Surface ($d=3..9$), Color ($d=5,7$), and Bivariate Bicycle ($[[72,12,6]]$, $[[90,8,10]]$, $[[108,8,10]]$, $[[144,12,12]]$) codes.

Reverted during review

  • AVX-512 intrinsics (c900f95, reverted in 85f8c44). Explicit _mm512_* intrinsics were benchmarked at only ~2–4% over #pragma GCC ivdep auto-vectorization in optimized builds — not worth the readability cost.

2. Build and integration fixes

  • fix(build): removed leftover global AVX-512 copts from .bazelrc.
    c900f95 added -mavx512f/-mavx512bw/-mavx512dq to build:linux; the revert in 85f8c44 did not undo it. This applied AVX-512 codegen to every Linux target repo-wide, so all binaries — including unrelated ones like tesseract_tests, common_tests, and the Python tests — aborted with SIGILL on any CPU without AVX-512, which is the case for GitHub Actions runners. Vectorization remains portable via //src:OPT_COPTS (-march=native, or -march=x86-64 for portable wheels).

  • fix(bp): qualified symbols moved into the tesseract_decoder namespace.
    After syncing main, the BP sources (top-level namespace bp) no longer resolved common::merge_indistinguishable_errors, parallel_for_shots_in_order, or parse_py_object. Call sites are now fully qualified, and bp_main.cc gained using namespace tesseract_decoder; to match tesseract_main.cc / simplex_main.cc.

  • fix(cmake): linked dem_decomposition into error_correlations_test.
    error_correlations.test.cc calls prepare_two_component_dem(), defined in multi_pass/dem_decomposition.cc, but the CMake target linked only error_correlations — an undefined reference that broke cmake --build . outright. The Bazel target already declared the dep, which is why Bazel-only CI never caught it.

  • style: clang-format applied to the BP sources.


3. Benchmark results

48-vCPU Intel Xeon, 100,000 Monte Carlo shots per configuration.

QEC code family Decoder config Baseline (shots/s) Optimized (shots/s) Gain
Surface Code ($d=3$, $p=0.001$) serial-batched 306,867.6 922,330.5 3.01x
Surface Code ($d=5$, $p=0.001$) serial-batched 30,728.0 161,361.1 5.25x
Surface Code ($d=7$, $p=0.001$) serial-batched 6,777.6 32,631.1 4.81x
Surface Code ($d=9$, $p=0.001$) serial-batched 2,258.9 8,660.9 3.83x
Color Code Superdense ($d=5$) serial-batched 14,815.3 65,647.4 4.43x
Color Code Superdense ($d=7$) serial-batched 3,062.9 9,399.2 3.07x
Bivariate Bicycle $[[72,12,6]]$ serial-batched + OSD-0 1,384.9 4,142.6 2.99x
Bivariate Bicycle $[[72,12,6]]$ serial-batched + OSD-1 978.5 2,165.6 2.21x
Bivariate Bicycle $[[90,8,10]]$ serial-batched + OSD-0 539.3 1,466.1 2.72x
Bivariate Bicycle $[[108,8,10]]$ serial-batched + OSD-0 414.3 1,096.1 2.65x
Bivariate Bicycle $[[144,12,12]]$ serial-batched + OSD-0 208.9 462.0 2.21x

OSD truncation, isolated

$d=11$ / $r=11$ unrotated surface code, heterogeneous 1–7% noise, 1,000 shots, 50,460 error mechanisms vs 2,420 detectors.

Truncation factor Matrix width LER End-to-end Speedup
0.00 (off) 50,460 0.16000 124.5 s 1.00x
1.05 2,541 0.16300 43.6 s 2.85x
1.10 2,662 0.16300 43.8 s 2.84x
1.20 (default) 2,904 0.15600 44.9 s 2.77x

LER is unaffected; at 1.20 it is marginally better, since excluding extremely low-probability variables keeps OSD out of obscure high-weight equivalence classes.


4. Verification

Gate Before After
bazel test --jobs=1 src/... 11/14 FAILED (SIGILL) 21/21 PASSED
cmake --build . --parallel 1 && ctest link error 8/8 PASSED
clang-format --dry-run --Werror 7 violations clean
Staleness vs main 33 commits behind synced

All GitHub CI checks green, including build (ubuntu-latest) — the job that was previously crashing with SIGILL.


5. Out of scope / follow-ups

  1. OSD-0 early exit on residual syndrome — stop Gauss-Jordan elimination as soon as the residual syndrome is spanned by the pivots found so far, rather than driving to full rank.
  2. Top-$K$ partial sort — replace the full std::sort over all $N$ columns with std::nth_element + prefix sort, since truncation discards the tail anyway (~10x fewer comparisons at $d=21$).
  3. Tiered adaptive truncation — attempt OSD on the top ~500 columns first and expand only when the residual syndrome is not spanned.
  4. Batched OSD Gaussian elimination — OSD currently runs single-threaded per shot.
  5. OSD-2 perturbation pruning — candidate pair search is currently exhaustive.
  6. BP-guided Tesseract A* — feed BP posteriors into the search heuristic instead of using OSD as the fallback.

@aria-googler
aria-googler requested a review from a team as a code owner August 8, 2026 06:08
@aria-googler
aria-googler requested review from arshpreetmaan and viathor and removed request for a team and viathor August 8, 2026 06:08
…nsics"

This reverts commit c900f95, returning to GCC ivdep pragmas for auto-vectorization
in order to maximize code readability and simplicity. The marginal performance gain
in release builds (~2-4%) was deemed not worth the architectural complexity of
explicit intrinsics.

Also forces -c opt in benchmark.sh to prevent unintentional unoptimized profiling.
- Implements matrix width truncation in OSD Gaussian Elimination (bounded by factor * num_detectors).
- Substantially improves e2e BP-OSD CPU decoding speed (up to 32% faster) without degrading LER.
- Exposes `osd_truncation_factor` configuration in BPParams with a safe, optimized default of 1.2.
- Wires CLI parameter `--osd-truncation-factor` to bp_main.cc and adds PyBind11 bindings.
- Updates benchmark scripts to measure and test the optimization.
The -mavx512f/-mavx512bw/-mavx512dq flags were introduced alongside the
explicit AVX-512 intrinsics in c900f95. That commit was reverted in
85f8c44, but the revert did not touch .bazelrc, leaving the flags applied
to every Linux build target repo-wide.

This caused all binaries (including non-BP targets such as tesseract_tests,
common_tests and the Python tests) to abort with SIGILL / 'Illegal
instruction' on any CPU without AVX-512, which is the case for the GitHub
Actions ubuntu-latest runners.

Vectorization is still handled portably: //src:OPT_COPTS selects
-march=native for normal builds and -march=x86-64 for portable wheels, and
the serial min-sum kernel relies on '#pragma GCC ivdep' auto-vectorization.
Fixes clang-format --dry-run --Werror violations in bp_params.h,
bp.pybind.h, bp_serial_min_sum.test.cc, osd_post_processor.{h,cc},
tesseract_bp_decoder.cc and bp_main.cc. Formatting-only, no behavior change.
main wrapped the shared helpers in an outer 'tesseract_decoder' namespace.
The BP sources live in a top-level 'namespace bp' and therefore no longer
resolved these names after merging main:

  - tesseract_bp_decoder.cc: common::merge_indistinguishable_errors
  - bp_main.cc: parallel_for_shots_in_order
  - bp.pybind.h / bp_sinter_compat.pybind.h: parse_py_object

Fully qualify the call sites (and add 'using namespace tesseract_decoder;'
to bp_main.cc, matching tesseract_main.cc and simplex_main.cc).
error_correlations.test.cc calls prepare_two_component_dem(), which is
defined in multi_pass/dem_decomposition.cc, but the CMake target only
linked the 'error_correlations' library. This produced an undefined
reference at link time and broke 'cmake --build .' entirely.

The Bazel target //src:error_correlations_tests already declares the
dependency correctly, which is why CI (Bazel-only) did not catch it.
@aria-googler aria-googler changed the title Feat/bp decoder optimizations feat(bp): BP decoder memory-layout, scheduling, and OSD truncation optimizations (2.2x-5.25x) Sep 22, 2026
@aria-googler
aria-googler merged commit b1ea085 into feat/bp-decoder Sep 22, 2026
13 checks passed
@aria-googler
aria-googler deleted the feat/bp-decoder-optimizations branch September 22, 2026 05:33
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants