ci: track memory usage for benchmarks - #9640
Conversation
Signed-off-by: Mikhail Kot <mikhail@spiraldb.com>
Merging this PR will regress 2 benchmarks
|
| Mode | Benchmark | BASE |
HEAD |
Efficiency | |
|---|---|---|---|---|---|
| ❌ | WallTime | arrow_checked_add_u32_avx2[16384] |
17.7 µs | 21.3 µs | -16.84% |
| ❌ | Simulation | cold_misaligned[(16, 64)] |
345.7 µs | 390.5 µs | -11.46% |
| ⚡ | WallTime | arrow_checked_add_u32_avx512[16384] |
21.2 µs | 17.6 µs | +20.59% |
| ⚡ | Simulation | take[duplicates/repeated/primitive/nonnull/chunks=16/indices=1000] |
239.4 µs | 206.1 µs | +16.13% |
| ⚡ | WallTime | words_gather_scalar_avx2[65536] |
9.5 µs | 8.3 µs | +14.52% |
Tip
Investigate this regression by commenting @codspeedbot fix this regression on this PR, or directly use the CodSpeed MCP with your agent.
Comparing myrrc/bench-track-memory (12dd044) with develop (34a6912)
Footnotes
-
106 benchmarks were skipped, so the baseline results were used instead. If they were deleted from the codebase, click here and archive them to remove them from the performance reports. ↩
-
4 benchmarks were run, but are now archived. If they were deleted in another branch, consider rebasing to remove them from the report. Instead if they were added back, click here to restore them. ↩
Polar Signals Profiling ResultsLatest Run
Powered by Polar Signals Cloud |
Benchmarks: PolarSignals Profiling 📖Commits: PR datafusion / vortex-file-compressed / ns (1.012x ➖, 1↑ 1↓)
datafusion / vortex-file-compressed / unit (no group data, 0↑ 0↓)
No file size changes detected. |
Benchmarks: TPC-H SF=1 on NVME 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (0.995x ➖, 0↑ 0↓)
datafusion / vortex-file-compressed / unit (no group data, 0↑ 0↓)
datafusion / parquet / ns (0.987x ➖, 2↑ 0↓)
datafusion / parquet / unit (no group data, 0↑ 0↓)
duckdb / vortex-file-compressed / ns (0.991x ➖, 1↑ 0↓)
duckdb / vortex-file-compressed / unit (no group data, 0↑ 0↓)
duckdb / parquet / ns (0.999x ➖, 0↑ 0↓)
duckdb / parquet / unit (no group data, 0↑ 0↓)
No file size changes detected. |
Benchmarks: FineWeb NVMe 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (0.990x ➖, 1↑ 0↓)
datafusion / vortex-file-compressed / unit (no group data, 0↑ 0↓)
datafusion / parquet / ns (0.999x ➖, 0↑ 0↓)
datafusion / parquet / unit (no group data, 0↑ 0↓)
duckdb / vortex-file-compressed / ns (0.968x ➖, 0↑ 0↓)
duckdb / vortex-file-compressed / unit (no group data, 0↑ 0↓)
duckdb / parquet / ns (1.008x ➖, 0↑ 0↓)
duckdb / parquet / unit (no group data, 0↑ 0↓)
No file size changes detected. |
Benchmarks: Clickbench Sorted on NVME 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (1.002x ➖, 0↑ 0↓)
datafusion / vortex-file-compressed / unit (no group data, 0↑ 0↓)
datafusion / parquet / ns (1.018x ➖, 0↑ 1↓)
datafusion / parquet / unit (no group data, 0↑ 0↓)
duckdb / vortex-file-compressed / ns (0.974x ➖, 0↑ 0↓)
duckdb / vortex-file-compressed / unit (no group data, 0↑ 0↓)
duckdb / parquet / ns (1.022x ➖, 0↑ 1↓)
duckdb / parquet / unit (no group data, 0↑ 0↓)
File Size Changes (100 files changed, +0.0% overall, 48↑ 52↓)
Totals:
|
Benchmarks: TPC-DS SF=1 on NVME 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (0.998x ➖, 1↑ 0↓)
datafusion / vortex-file-compressed / unit (no group data, 0↑ 0↓)
datafusion / parquet / ns (1.003x ➖, 0↑ 1↓)
datafusion / parquet / unit (no group data, 0↑ 0↓)
duckdb / vortex-file-compressed / ns (0.986x ➖, 4↑ 2↓)
duckdb / vortex-file-compressed / unit (no group data, 0↑ 0↓)
duckdb / parquet / ns (1.001x ➖, 3↑ 4↓)
duckdb / parquet / unit (no group data, 0↑ 0↓)
No file size changes detected. |
Benchmarks: TPC-H SF=10 on NVME 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (1.006x ➖, 0↑ 0↓)
datafusion / vortex-file-compressed / unit (no group data, 0↑ 0↓)
datafusion / parquet / ns (1.001x ➖, 0↑ 0↓)
datafusion / parquet / unit (no group data, 0↑ 0↓)
duckdb / vortex-file-compressed / ns (1.012x ➖, 0↑ 0↓)
duckdb / vortex-file-compressed / unit (no group data, 0↑ 0↓)
duckdb / parquet / ns (1.015x ➖, 0↑ 1↓)
duckdb / parquet / unit (no group data, 0↑ 0↓)
No file size changes detected. |
Benchmarks: Statistical and Population Genetics 📖Commits: PR How to read Verdict and Engines
duckdb / vortex-file-compressed / ns (1.019x ➖, 1↑ 2↓)
duckdb / vortex-file-compressed / unit (no group data, 0↑ 0↓)
duckdb / parquet / ns (0.997x ➖, 0↑ 0↓)
duckdb / parquet / unit (no group data, 0↑ 0↓)
No file size changes detected. |
Benchmarks: FineWeb S3 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (0.861x ➖, 1↑ 0↓)
datafusion / vortex-file-compressed / unit (no group data, 0↑ 0↓)
datafusion / parquet / ns (0.977x ➖, 0↑ 0↓)
datafusion / parquet / unit (no group data, 0↑ 0↓)
duckdb / vortex-file-compressed / ns (0.939x ➖, 0↑ 0↓)
duckdb / vortex-file-compressed / unit (no group data, 0↑ 0↓)
duckdb / parquet / ns (0.857x ➖, 4↑ 3↓)
duckdb / parquet / unit (no group data, 0↑ 0↓)
|
Benchmarks: Clickbench on NVME 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (1.006x ➖, 0↑ 0↓)
datafusion / vortex-file-compressed / unit (no group data, 0↑ 0↓)
datafusion / parquet / ns (1.004x ➖, 1↑ 0↓)
datafusion / parquet / unit (no group data, 0↑ 0↓)
duckdb / vortex-file-compressed / ns (1.006x ➖, 0↑ 1↓)
duckdb / vortex-file-compressed / unit (no group data, 0↑ 0↓)
duckdb / parquet / ns (1.011x ➖, 0↑ 1↓)
duckdb / parquet / unit (no group data, 0↑ 0↓)
No file size changes detected. |
Benchmarks: TPC-H SF=1 on S3 📖Commits: PR How to read Verdict and Engines
datafusion / vortex-file-compressed / ns (0.968x ➖, 2↑ 3↓)
datafusion / vortex-file-compressed / unit (no group data, 0↑ 0↓)
datafusion / parquet / ns (1.005x ➖, 0↑ 0↓)
datafusion / parquet / unit (no group data, 0↑ 0↓)
duckdb / vortex-file-compressed / ns (1.055x ➖, 0↑ 3↓)
duckdb / vortex-file-compressed / unit (no group data, 0↑ 0↓)
duckdb / parquet / ns (1.007x ➖, 0↑ 0↓)
duckdb / parquet / unit (no group data, 0↑ 0↓)
|
There was a problem hiding this comment.
claude had this to say:
The .diff() sign fix looks right, and end_query was the only caller with the arguments backwards.
The peak change is the problem. VmHWM is a process-lifetime high water mark, but it lands in a per-query, per-format MemoryMeasurement that gets emitted as q{i}_peak and ingested as peak_physical. It never decreases, so every query reports the running max of everything the process did before it. runner.rs:341 loops formats on the outside within one process, so the second format's peaks inherit the first format's peak, which breaks the comparison this metric exists for.
Measured under mimalloc, which vortex-bench/src/lib.rs:78 installs:
start VmHWM= 6.1 MB VmRSS= 6.2 MB VmPeak= 1029.7 MB
q0: 512 MiB live VmHWM= 518.3 MB VmRSS= 518.3 MB VmPeak= 1029.7 MB
q0: freed VmHWM= 518.3 MB VmRSS= 518.3 MB VmPeak= 1029.7 MB
q1: 8 MiB VmHWM= 526.3 MB VmRSS= 526.3 MB VmPeak= 1029.7 MB
A query touching 8 MiB reports a 526 MB peak. VmPeak is also already 1029.7 MB before any query runs, because mimalloc reserves the address space up front, so peak_virtual_memory is a constant for every query, format and commit. Worth dropping the virtual columns rather than ingesting a constant.
Writing 5 to /proc/self/clear_refs in start_query does reset VmHWM (513.9 MB down to 2.1 MB in a standalone check, with the following 8 MiB query then reading 10.1 MB). It has no effect on VmPeak, and because mimalloc holds onto freed pages even the reset measures growth above the retained arena rather than a working set. Open question: is a per-query peak the right shape here, or should this be one process-level peak per run?
Separately, --track-memory in CI adds a permanently empty section to every benchmark PR comment. MemoryMeasurementJson carries no unit or value, and tpch_q00_memory/... does not match the compare script's _q(\d+)/ pattern, so the row can never resolve a value.
scripts/compare-benchmark-jsons.py output with memory rows in results.json
datafusion / vortex-file-compressed / unit (no group data, 0↑ 0↓)
| tpch_q00_memory/datafusion:vortex | — / — / no baseline |
Adding matching memory rows to base.json does not change this, because hot_value is NaN on both sides. Two extra sections per engine/format, on every bench job. The rows also get appended to data.json.gz on develop.
Real numbers still reach Postgres through results.ingest.jsonl, so the fix is either a unit and value on the JSON row or keeping memory rows out of the GhJson output.
On the tests: query_end_peak asserts peak > 0 and peak <= VmHWM read afterwards, and the second assertion is a tautology given monotonicity. Both new tests pass against the old global_tracker.peak_memory() code and with the sign bug still in place, and nothing asserts physical_memory_delta > 0, which is the bug being fixed. vec![7u8; 8 * 1024 * 1024] is also elided under optimization, so it needs black_box if it is meant to be load-bearing.
Nit: MemoryMeasurementResult::peak_physical_memory still documents itself as coming from the global tracker, and Some((peak_rss?, peak_vsz?)) discards one field's successful parse when the other is missing.
I think this makes sense? And also have we looked into other harnesses that can track memory at a finer granularity?
negative memory on memory usage increase