Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
22 commits
Select commit Hold shift + click to select a range
5a50b69
add generated upgrade test harness
aagbsn Aug 11, 2026
cdc4180
Fix ClickHouse schema divergence from ooni/backend verification
aagbsn Aug 17, 2026
bb8df48
Add per-step CI workflow for the upgrade test; fix obs_web from prod …
aagbsn Aug 17, 2026
d47d05c
Verify remaining tables against live SHOW CREATE TABLE output
aagbsn Aug 17, 2026
785ac2d
Here are the two commit messages, one per patch delivered above.
aagbsn Aug 17, 2026
52df4e9
Recommend stopping upgrade at 25.8.29.51 for now; record audit findings
aagbsn Aug 17, 2026
8364487
Bisect the 25.3.14.14 -> 25.8.29.51 mark-file incompatibility
aagbsn Aug 17, 2026
4ed89b4
Confirm the 25.8.29.51 mark-format break via bisection, document find…
aagbsn Aug 18, 2026
d9000d7
Don't abort staged-upgrade job on first node-upgrade failure within a…
aagbsn Aug 18, 2026
0275ae4
Promote production recommendation to latest LTS with a 4-hop runbook
aagbsn Aug 18, 2026
a820098
Add real-data end-to-end upgrade scenario (ooni/data#160)
aagbsn Aug 18, 2026
ad9b79a
Fix NameError in load_schema_and_seed() from the apply_schema() refactor
aagbsn Aug 18, 2026
e83178c
Add real-data-upgrade job to the CI workflow
aagbsn Aug 18, 2026
b8e4fbc
Add zero-downtime, zero-corruption upgrade proof (harness/availabilit…
aagbsn Aug 19, 2026
7101d09
add missing availability.py
aagbsn Sep 21, 2026
f9d73b6
decrease LTS upgrade hops
aagbsn Sep 21, 2026
b683a9f
update documented results
aagbsn Sep 21, 2026
2e17cd0
validate content integrity at each round of upgrades
aagbsn Sep 21, 2026
4e3eef4
skip ddl verify marker columns in integrity check
aagbsn Sep 24, 2026
c41055f
fix checksum verification over other tables
aagbsn Sep 24, 2026
1583a95
avoid nullable arguments to cityHash64
aagbsn Sep 24, 2026
4907c94
check integrity after each version upgrade
aagbsn Sep 24, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
532 changes: 532 additions & 0 deletions .github/workflows/clickhouse_upgrade_test.yml

Large diffs are not rendered by default.

8 changes: 8 additions & 0 deletions scripts/clickhouse-upgrade-test/.gitignore
Original file line number Diff line number Diff line change
@@ -0,0 +1,8 @@
__pycache__/
*.pyc

# Generated test output -- keep the directory (see results/.gitkeep) but not
# whatever a local run drops into it.
results/report.md
results/report.json
results/steps/
21 changes: 21 additions & 0 deletions scripts/clickhouse-upgrade-test/Makefile
Original file line number Diff line number Diff line change
@@ -0,0 +1,21 @@
.PHONY: test test-staged test-direct config down clean

# Full test: staged (recommended) upgrade path, then the naive direct-jump path
test:
python3 run_test.py --scenario both

test-staged:
python3 run_test.py --scenario staged

test-direct:
python3 run_test.py --scenario direct

# Validate docker-compose.yml without needing registry access
config:
docker compose config

down:
docker compose down -v

clean: down
rm -rf results/report.md results/report.json
1,046 changes: 1,046 additions & 0 deletions scripts/clickhouse-upgrade-test/README.md

Large diffs are not rendered by default.

316 changes: 316 additions & 0 deletions scripts/clickhouse-upgrade-test/ci_step.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,316 @@
#!/usr/bin/env python3
"""
Per-step CLI for running the upgrade test as discrete, individually
pass/fail-able steps -- built for .github/workflows/clickhouse_upgrade_test.yml,
where each hop / node-upgrade gets its own GitHub Actions step (own
checkmark, own timing, own expandable log), rather than one opaque job that
only reports pass/fail for the entire upgrade path at once.

Each invocation is a fresh process. State that would normally be threaded
through function arguments (which image tag each node is currently running)
is instead recovered by inspecting the already-running containers via
`docker inspect` (see harness/compose.py:current_env()) -- so steps are
just plain sequential shell commands in a workflow, no shared state file to
keep in sync, other than the results/steps/*.json this script writes after
every step (used by the final `report` step to assemble a combined summary
from whichever steps actually ran).

For local one-shot runs (not CI), use run_test.py / `make test` instead --
this script is intentionally low-level.

Usage:
python3 ci_step.py setup --base-version 24.8.6.70 --label setup
python3 ci_step.py upgrade-node --node ch1 --version 25.3.14.14 --label hop-25.3-ch1
python3 ci_step.py verify-ddl --version 25.3.14.14 --label hop-25.3-verify-ddl
python3 ci_step.py content-integrity --label content-integrity-25.3.14.14 # run once per hop, right after that hop's verify-ddl
python3 ci_step.py report
python3 ci_step.py teardown

`report`'s pass/fail now reflects whether the rollout ended with no data
loss or corruption, not whether every single step was individually clean
-- a node-upgrade step that logs a hard-looking mixed-version error but
self-heals by the end of its own hop (its hop's verify-ddl step settles
cleanly) no longer fails the job on its own; see harness/report.py's
self_healed()/effective_ok()/overall_ok() for the exact rule, and
harness/scenarios.py's content_integrity_step() for the "did anything
ALREADY there get lost or corrupted" check `report` also requires to pass
at every hop it ran at (not just the last one -- see .github/workflows/
clickhouse_upgrade_test.yml, which now invokes this once per hop with a
distinct --label so a corrupting release is pinpointed to its own hop
instead of only surfacing as "something in the rollout broke it").

Real-data scenario (harness/real_data.py) -- separate CLI verbs, since it's
a different flow (load real data once, then hop):
python3 ci_step.py setup-real-data --base-version 24.8.6.70 --label setup-real-data
python3 ci_step.py load-real-data --label load-real-data
python3 ci_step.py golden-snapshot --label golden-snapshot
python3 ci_step.py verify-e2e --label base-verify-e2e
python3 ci_step.py real-data-hop --version 25.3.14.14 --label rd-hop1
python3 ci_step.py zero-downtime-upgrade --label zero-downtime-upgrade
python3 ci_step.py report
python3 ci_step.py teardown-real-data

See .github/workflows/clickhouse_upgrade_test.yml for the full sequence.
"""
from __future__ import annotations

import argparse
import json
import os
import sys
from pathlib import Path

sys.path.insert(0, str(Path(__file__).resolve().parent))

from harness import compose, real_data, report
from harness import availability
from harness.scenarios import content_integrity_step, setup_step, step_ok, upgrade_node_step, verify_ddl_step
from harness.versions import RECOMMENDED_LTS_HOPS

PROJECT_DIR = Path(__file__).resolve().parent
RESULTS_DIR = PROJECT_DIR / "results"
STEPS_DIR = RESULTS_DIR / "steps"


def _save_step(label: str, result: dict) -> None:
STEPS_DIR.mkdir(parents=True, exist_ok=True)
safe = label.replace("/", "_")
(STEPS_DIR / f"{safe}.json").write_text(json.dumps(result, indent=2, default=str))


def cmd_setup(args) -> int:
result = setup_step(args.base_version, label=args.label)
_save_step(args.label, result)
ok = step_ok(result)
print(f"[{args.label}] {'OK' if ok else 'FAILED: ' + str(result.get('error'))}")
return 0 if ok else 1


def cmd_upgrade_node(args) -> int:
step = upgrade_node_step(args.node, args.version, label=args.label)
_save_step(args.label, step)
ok = step_ok(step)
print(f"[{args.label}] {'OK' if ok else 'PROBLEM DETECTED'}")
print(json.dumps(step, indent=2, default=str))
return 0 if ok else 1


def cmd_verify_ddl(args) -> int:
result = verify_ddl_step(args.version, label=args.label)
_save_step(result["label"], result)
ok = step_ok(result)
print(f"[{result['label']}] {'OK' if ok else 'FAILED: ' + str(result.get('error'))}")
return 0 if ok else 1


def cmd_report(args) -> int:
RESULTS_DIR.mkdir(exist_ok=True)
steps = []
if STEPS_DIR.exists():
# Chronological order (steps run strictly sequentially within a CI
# job), not alphabetical -- so mtime, not filename, decides order.
for f in sorted(STEPS_DIR.glob("*.json"), key=lambda p: p.stat().st_mtime):
steps.append(json.loads(f.read_text()))

md = report.render_ci_steps_report(steps)
(RESULTS_DIR / "report.md").write_text(md)
(RESULTS_DIR / "report.json").write_text(json.dumps(steps, indent=2, default=str))
print(md)

gh_summary = os.environ.get("GITHUB_STEP_SUMMARY")
if gh_summary:
with open(gh_summary, "a") as f:
f.write(md)
f.write("\n")

# Individual upgrade-node/verify-ddl steps now run with
# `continue-on-error: true` (see .github/workflows/clickhouse_upgrade_test.yml)
# so that a failure partway through a hop doesn't abort the job before
# the remaining nodes in that hop get a chance to upgrade too -- e.g. so
# we can observe whether a lagging node's stuck replication queue clears
# once it also reaches the new version. That means this step -- which
# has no continue-on-error and runs with `if: always()` -- is now the
# thing that actually has to fail the job when something stayed broken.
#
# report.overall_ok(), not a plain step_ok() scan: a node-upgrade step
# that logged a hard-looking mixed-version error but self-healed by the
# end of its own hop (its hop's verify-ddl step settled cleanly) no
# longer fails the job on its own -- see report.py's self_healed() /
# effective_ok(). The per-step results above still show its literal
# FAIL, annotated, so nothing is hidden; only the job-level gate changes.
any_fail = not report.overall_ok(steps)
return 1 if any_fail else 0


def cmd_content_integrity(args) -> int:
result = content_integrity_step(label=args.label)
_save_step(args.label, result)
ok = step_ok(result)
print(f"[{args.label}] {'OK -- no data loss or corruption in pre-existing seed data' if ok else 'FAILED -- see diffs'}")
print(json.dumps(result, indent=2, default=str))
return 0 if ok else 1


def cmd_teardown(args) -> int:
try:
compose.down(volumes=True)
except Exception as e:
print(f"teardown warning (non-fatal): {e}")
return 0


# ---------------------------------------------------------------------------
# Real-data scenario (harness/real_data.py) -- see the `real-data-upgrade`
# job in .github/workflows/clickhouse_upgrade_test.yml for the full step
# sequence these are wired into, and README.md for the design writeup.
# ---------------------------------------------------------------------------


def cmd_setup_real_data(args) -> int:
result = real_data.setup_real_data_step(args.base_version, label=args.label)
_save_step(args.label, result)
ok = step_ok(result)
print(f"[{args.label}] {'OK' if ok else 'FAILED: ' + str(result.get('error'))}")
return 0 if ok else 1


def cmd_load_real_data(args) -> int:
result = real_data.load_real_data_step(label=args.label)
_save_step(args.label, result)
ok = step_ok(result)
print(f"[{args.label}] {'OK' if ok else 'FAILED'}")
print(json.dumps(result, indent=2, default=str))
return 0 if ok else 1


def cmd_golden_snapshot(args) -> int:
result = real_data.take_golden_snapshot_step(label=args.label)
_save_step(args.label, result)
ok = step_ok(result)
print(f"[{args.label}] {'OK' if ok else 'FAILED: nodes disagree before any upgrade -- ' + str(result.get('mismatched_tables'))}")
return 0 if ok else 1


def cmd_verify_e2e(args) -> int:
result = real_data.run_e2e_verify_step(args.label)
_save_step(args.label, result)
ok = step_ok(result)
print(f"[{args.label}] {'OK' if ok else 'FAILED (exit ' + str(result.get('exit_code')) + ')'}")
return 0 if ok else 1


def cmd_real_data_hop(args) -> int:
result = real_data.real_data_hop_step(args.version, args.label)
_save_step(args.label, result)
ok = step_ok(result)
print(f"[{args.label}] {'OK' if ok else 'PROBLEM DETECTED'}")
print(json.dumps(result, indent=2, default=str))
return 0 if ok else 1


def cmd_teardown_real_data(args) -> int:
try:
compose.down(volumes=True, files=real_data.COMPOSE_FILES)
except Exception as e:
print(f"teardown warning (non-fatal): {e}")
return 0


def cmd_zero_downtime_upgrade(args) -> int:
# RECOMMENDED_LTS_HOPS[1:] -- skip the BASE_VERSION entry (24.8.6.70, None),
# since that's the version setup-real-data already brought the cluster
# up on; the canary starts running from the current (base) version and
# the first real hop upgrades away from it, exactly like real-data-hop
# is invoked once per remaining RECOMMENDED_LTS_HOPS entry today.
result = availability.run_zero_downtime_upgrade(RECOMMENDED_LTS_HOPS[1:], label=args.label)
_save_step(args.label, result)
ok = step_ok(result)
print(f"[{args.label}] {'OK' if ok else 'PROBLEM DETECTED'}")
print(json.dumps({k: v for k, v in result.items() if k != "hops"}, indent=2, default=str))
return 0 if ok else 1


def main() -> int:
parser = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter)
sub = parser.add_subparsers(dest="command", required=True)

p_setup = sub.add_parser("setup", help="Tear down any previous state, bring up a fresh cluster, load schema + seed data")
p_setup.add_argument("--base-version", required=True, help="ClickHouse image tag for all 3 nodes at startup")
p_setup.add_argument("--label", default="setup", help="Step label, used as the results/steps/<label>.json filename")
p_setup.set_defaults(func=cmd_setup)

p_upgrade = sub.add_parser("upgrade-node", help="Recreate one node on a new image tag; other nodes keep running")
p_upgrade.add_argument("--node", required=True, choices=["ch1", "ch2", "ch3"])
p_upgrade.add_argument("--version", required=True, help="ClickHouse image tag to upgrade this node to")
p_upgrade.add_argument("--label", required=True, help="Step label, e.g. hop-25.3-ch1")
p_upgrade.set_defaults(func=cmd_upgrade_node)

p_ddl = sub.add_parser("verify-ddl", help="Confirm an ON CLUSTER ALTER still propagates at the current version(s)")
p_ddl.add_argument("--version", required=True, help="Version label for this checkpoint (used in the test column name)")
p_ddl.add_argument("--label", default=None, help="Step label; defaults to verify-ddl-<version>")
p_ddl.set_defaults(func=cmd_verify_ddl)

p_ci = sub.add_parser(
"content-integrity",
help=(
"Does the pre-existing seed data (probe rows from mid-rollout writes "
"excluded) still checksum-match the golden snapshot taken at setup, on "
"all 3 nodes, right now? This is the 'no data loss or corruption' "
"signal, independent of whether any mid-rollout step self-healed. Run "
"once per hop (see .github/workflows/clickhouse_upgrade_test.yml), each "
"time with a distinct --label, so a mismatch pinpoints the hop/release "
"that caused it rather than only 'somewhere in the rollout'."
),
)
p_ci.add_argument("--label", default="content-integrity")
p_ci.set_defaults(func=cmd_content_integrity)

p_report = sub.add_parser("report", help="Aggregate every results/steps/*.json into results/report.md + report.json")
p_report.set_defaults(func=cmd_report)

p_teardown = sub.add_parser("teardown", help="docker compose down -v")
p_teardown.set_defaults(func=cmd_teardown)

p_setup_rd = sub.add_parser("setup-real-data", help="Fresh cluster + schema only (no synthetic seed) -- real-data scenario")
p_setup_rd.add_argument("--base-version", required=True, help="ClickHouse image tag for all 3 nodes at startup")
p_setup_rd.add_argument("--label", default="setup-real-data")
p_setup_rd.set_defaults(func=cmd_setup_real_data)

p_load_rd = sub.add_parser("load-real-data", help="Bring up postgres/valkey/api/downloader/fastpath, wait for the one-shot containers -- run exactly once per CI run")
p_load_rd.add_argument("--label", default="load-real-data")
p_load_rd.set_defaults(func=cmd_load_real_data)

p_golden = sub.add_parser("golden-snapshot", help="Snapshot real-data tables on all 3 nodes; persist as the baseline every later hop is diffed against")
p_golden.add_argument("--label", default="golden-snapshot")
p_golden.set_defaults(func=cmd_golden_snapshot)

p_verify_e2e = sub.add_parser("verify-e2e", help="Run ooni/data#160's real pytest suite against the cluster's current state (no upgrade)")
p_verify_e2e.add_argument("--label", default="verify-e2e")
p_verify_e2e.set_defaults(func=cmd_verify_e2e)

p_rd_hop = sub.add_parser("real-data-hop", help="Upgrade all 3 nodes to --version, then re-check integrity against the golden snapshot and re-run ooni/data's pytest suite")
p_rd_hop.add_argument("--version", required=True, help="ClickHouse image tag to upgrade all 3 nodes to")
p_rd_hop.add_argument("--label", required=True, help="Step label, e.g. rd-hop1")
p_rd_hop.set_defaults(func=cmd_real_data_hop)

p_zdt = sub.add_parser(
"zero-downtime-upgrade",
help=(
"Run every remaining RECOMMENDED_LTS_HOPS hop back-to-back in one process, with a "
"continuous read/write canary (round-robin + failover across all 3 nodes) running "
"the whole time -- proves no full-cluster downtime, no blocked writes, and no lost "
"or corrupted data across the entire rollout, not just at per-hop checkpoints"
),
)
p_zdt.add_argument("--label", default="zero-downtime-upgrade")
p_zdt.set_defaults(func=cmd_zero_downtime_upgrade)

p_teardown_rd = sub.add_parser("teardown-real-data", help="docker compose -f docker-compose.yml -f docker-compose.real-data.yml down -v")
p_teardown_rd.set_defaults(func=cmd_teardown_real_data)

args = parser.parse_args()
return args.func(args)


if __name__ == "__main__":
raise SystemExit(main())
40 changes: 40 additions & 0 deletions scripts/clickhouse-upgrade-test/config/ch1/node.xml
Original file line number Diff line number Diff line change
@@ -0,0 +1,40 @@
<?xml version="1.0"?>
<clickhouse>
<macros>
<shard>01</shard>
<replica>01</replica>
<cluster>oonidata_cluster</cluster>
</macros>

<!-- Embedded Keeper: this node is Keeper server id=1 -->
<keeper_server>
<tcp_port>9181</tcp_port>
<server_id>1</server_id>
<log_storage_path>/var/lib/clickhouse/coordination/log</log_storage_path>
<snapshot_storage_path>/var/lib/clickhouse/coordination/snapshots</snapshot_storage_path>

<coordination_settings>
<operation_timeout_ms>10000</operation_timeout_ms>
<session_timeout_ms>30000</session_timeout_ms>
<raft_logs_level>information</raft_logs_level>
</coordination_settings>

<raft_configuration>
<server>
<id>1</id>
<hostname>ch1</hostname>
<port>9234</port>
</server>
<server>
<id>2</id>
<hostname>ch2</hostname>
<port>9234</port>
</server>
<server>
<id>3</id>
<hostname>ch3</hostname>
<port>9234</port>
</server>
</raft_configuration>
</keeper_server>
</clickhouse>
Loading
Loading