Local LGTM observability stack for engine telemetry - #751
Open
tony wants to merge 24 commits into
Open
Conversation
tony
force-pushed
the
lgtm-observability
branch
from
August 23, 2026 01:21
a4d3724 to
eb4ede0
Compare
tony
force-pushed
the
lgtm-observability
branch
3 times, most recently
from
August 23, 2026 01:41
b07244b to
a1a468b
Compare
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## engine-ops-orchestration-bench #751 +/- ##
==================================================================
+ Coverage 73.94% 73.96% +0.02%
==================================================================
Files 173 173
Lines 11913 11913
Branches 1914 1914
==================================================================
+ Hits 8809 8812 +3
+ Misses 2439 2436 -3
Partials 665 665 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
tony
force-pushed
the
lgtm-observability
branch
2 times, most recently
from
August 23, 2026 02:21
a0d0564 to
3db13bf
Compare
tony
force-pushed
the
lgtm-observability
branch
2 times, most recently
from
August 23, 2026 14:15
1495fce to
8b35148
Compare
tony
force-pushed
the
lgtm-observability
branch
from
August 23, 2026 15:36
8b35148 to
aec2d99
Compare
tony
force-pushed
the
lgtm-observability
branch
from
August 23, 2026 15:58
aec2d99 to
734a31a
Compare
why: Every rung inherited the benchmark's flat 120s no-progress watchdog, whatever topology it built. The cost per pass is superlinear on this ladder's own shapes -- one control/async pass measures 14.3s at 400 panes, 24.0s at 800, 52.5s at 1200 and 177.5s at 1600 -- so the allowance that is generous at the base is short at the top, and a slower rung is reported as a stuck one. That stops the axis before it reaches any real limit, which is the opposite of what an escalating harness is for. what: - Give each rung an allowance derived from the panes it builds, and pass it through to the benchmark - Cap it at the hard per-rung limit, so a genuinely stuck rung dies no later than it would have - Cover the scaling, the cap, and that the value actually reaches the child
why: The engine instrumentation seam could emit telemetry but had nowhere to send it, so nobody could see what a tmux workload actually costs. Standing up that stack by hand is unreproducible, and a dashboard nobody queries rots into panels that render an empty grid and look healthy doing it. The workload is short-lived, so every panel has to read the window the viewer selected rather than a counter at an instant: Prometheus marks a finished run's series stale within minutes, and an instant query then returns nothing beside a timeseries still drawing the same run. what: - Add scripts/lgtm/up.sh, pinning the otel-lgtm image and bind-mounting the datasource and dashboard provisioning, with a config label so changing a mount recreates the container instead of restarting it with stale state - Publish Grafana on 3900 and Prometheus on 9099: a host process already bound to their defaults still lets Docker publish, then answers first, so queries reach the wrong server and return plausible data - Add scripts/lgtm/telemetry.py, an OTelSink emitting spans and metrics together so the duration histogram records while its span is current and carries an exemplar, giving the metric-to-trace pivot - Add scripts/otel_smoke.py driving all four transports, with grouped commands so the inlining panels are not flat zero and rejected commands so the failure panels have real data - Generate the three dashboards from scripts/lgtm/generate_dashboards.py rather than hand-editing JSON, and commit the output - Query each panel's own window with increase(...[$__range]) rather than reading counters at a point, so a panel reports what happened in the range the viewer selected - Add scripts/otel_acceptance.py, which reads those dashboards, runs every panel's own query the way Grafana runs it -- range or instant -- and fails naming any that returned nothing; it re-checks until ingestion catches up, since "no data yet" and "no data ever" are indistinguishable at any single instant - Reject any Prometheus target with no range selector in a test, which catches a stale-series panel offline; a live check only sees it once the series has gone stale, which is precisely when nobody is watching - Add just otel-up/down/dashboards/smoke/acceptance/verify and an otel dependency group, kept out of dev so the ordinary gates stay lean
why: Telemetry that cannot say which branch, worktree, or experiment produced it can be observed but not compared, which is most of what a benchmark is for. The trap is treating that as one decision: copying every fact onto every signal makes each commit mint a fresh set of Prometheus series, and the cost lands on whoever runs the stack next month. what: - Add scripts/lgtm/identity.py resolving repository, ref, revision, and worktree from one git rev-parse at startup, using the OpenTelemetry vcs.* conventions, overridable by environment for CI's detached checkouts - Split the facts by signal: metrics carry only branch, run id, and spike, the dimensions worth grouping by; revision and worktree ride on traces and profiles where the drill-down happens. A SHA is never a comparison axis -- each run has one, so grouping by it is grouping by run - Copy baggage onto spans with a processor, so a value that changes mid-run reaches spans created inside the engines without becoming a parameter on calls that have no business knowing about telemetry - Give every board Transport and Branch selectors, and add a Compare board grouping the same measurements by run and by branch - Add tests pinning the metric label set closed and rejecting a panel that filters on an undefined variable
…rd guidance why: A fixed-worker loop cannot find saturation. As latency rises the workers slow with it, offered load falls, and the graph bends instead of breaking, so the steady workload could say how fast a transport is but never where it stops keeping up. Separately, five boards with no entry point is how dashboard sprawl starts: without directed browsing, finding the right one is guesswork and the fix people reach for is duplicating it. what: - Add scripts/lgtm/load_tmux.py and just otel-load, driving the engines under rampa's ramping-arrival-rate; the same control-mode engine measures p99 around 2 ms steady and around 16 ms at the top of the ramp - Keep telemetry on this project's sink rather than a rampa output backend: rampa's own backend would export under its service name and metric vocabulary, so a load-shaped run would arrive as a second account of the same work rather than a comparable one - Add a Home board naming what each board answers, and link Overview's panels down to the board that explains them - Reorganize Overview as rate, errors, and duration, since an engine is a service and RED is the frame its caller thinks in - Drop stacking and relax refresh to a minute: the data arrives in bursts from short runs, so a thirty-second poll buys nothing - Teach the dashboard tests that a text panel is documentation and is not expected to query anything
why: The README explained at length that a host process already bound to a port shadows the container's, so a query succeeds against the wrong backend and returns plausible data -- then shipped no way to detect it. Explaining a hazard without supplying the check is the worst of both: the reader knows to worry and has nothing to act on. what: - Add scripts/lgtm/verify.sh, comparing each service's build info as seen from inside the container against the same URL from the host, since liveness cannot tell the two apart but identity can - Run it from up.sh before reporting success, so a shadowed port fails the start rather than surfacing later as a confusing query result - Expose it as just otel-ports and document it - Cover loki, which the ad-hoc checks had been skipping - Drop the LANE alias in the dashboard generator; SCOPE was the only name it needed, and the regenerated JSON is byte-identical - Show just otel-load as a runnable command rather than only naming it in prose
why: A closed port is not always refused promptly -- under WSL and inside containers the packets are dropped and the socket waits out its full timeout. With one request per panel target, forgetting `just otel-up` produced a twenty-minute silence and then a report blaming empty panels, which points at the dashboards rather than at the stack. what: - Probe each backend once before checking any panel, and exit naming the fix; the same mistake now costs eighteen seconds and says what to do - Say so in the README, so the guard is discoverable before it fires
why: `scripts` is not in pytest's testpaths, so the doctests under scripts/lgtm were never executed. The examples there carry the load-bearing decision about which facts a metric may keep and which belong on a profile, so they were comments that looked like tests. what: - Run the identity and telemetry doctests from the test suite, skipping the latter when the otel dependency group is absent - Put scripts/lgtm on sys.path from telemetry.py itself, so importing it does not depend on the caller having arranged that first
why: The README's commands and board names are only correct until someone renames one, and the failure is silent -- the prose still reads well and the command simply does not work. Both sides live in this repo, so the check is cheap and there is no reason to rely on noticing. what: - Add tests asserting every `just` command the README demonstrates is a real recipe, and every generated board is named somewhere in it - Point the instrumentation page at the stack that consumes what it describes; it explained how to export to OTLP without mentioning that an exporter, a Grafana stack, and dashboards already ship here - Drop a fragile test count from that page, which was already stale - Remove a dead assignment in scope() and the unused parameter on the rampa teardown, which is called with none
why: The stack advertised that a log line links to the trace it came from, and the datasource was wired for it, but no log record carried a trace id. A record only picks up trace context while a span is current, and the lane summaries were logged between lanes, with no span open. The claim was false and nothing checked it. what: - Log each lane's result from inside a short span, so the record carries a trace id and the Loki-to-Tempo jump works; verified by taking a trace id from a Loki line and resolving it to its span in Tempo - Put the lane's totals on that span too, so the trace answers the same question the log line does
why: The resolver's edge cases carried the load and none were tested. A detached checkout answers the literal string HEAD, so used directly every detached run on every branch would share one dimension value and could not be told apart; CI needs its explicit ref to win over the checkout's; and a directory outside git has to yield nothing rather than raise. All three were verified by hand and nothing kept them true. what: - Test the ref override, the detached-HEAD fallback to tag then short revision, and the empty result outside a repository, each against a real git repository rather than a fixture that happens to exist - Correct the latency bucket comment, which claimed control mode costs tens of microseconds; the measured median is a couple of hundred. The buckets were already right and still resolve every lane's median into its own band
why: A single unknown lane name saturated the machine. Building the scenario's state costs a tmux server and three exporter threads, and it was built before the step that failed, so every retry left another set behind: 3,539 tmux servers, 10,619 threads, a load average near 2,000, and a test suite that then failed 24 tests for reasons that had nothing to do with the code. rampa is not at fault. Isolating a failing iteration and running the next one is what a load tool should do, and a minimal scenario that raises on every iteration holds a flat three threads. Reproducing it took twenty lines: build three threads per iteration, then raise, and a six-second run reaches a thousand threads. The defect is a scenario that rebuilds expensive state on a path that cannot succeed. what: - Latch a setup failure and re-raise it, so an impossible setup is attempted once however many iterations follow - Validate the lane at import, before a tmux server exists, so a typo costs three seconds and names the valid lanes - Test both: that fifty iterations produce one setup attempt, and that the lane check precedes anything that creates
why: Both Grafana and Tempo expose MCP servers, and neither was usable here. Tempo's was off, and the config the image hands out assumes Grafana's default port, which this stack deliberately moves. Following the shipped copy connects to nothing, and the failure is quiet in the worst way: the agent authenticates against nothing, finds no data, and reports an empty stack rather than a misconfigured one. what: - Enable Tempo's MCP server, which offers traceql-search, get-trace, and attribute discovery, so an agent can find the engine attributes without being told the schema; verified by driving the handshake and searching for this repo's own spans - Add just otel-mcp, printing a client config with this stack's real ports and the running container's token, which is never written to the repository - Test that the config and the stack agree on the Grafana port, since drift there fails silently
why: Only CPU sampling was ever collected, so "where did the time go" had an answer and "what did it allocate" did not. Pyroscope lists the other profile types it knows, which made the gap look like a wiring fault rather than a choice. what: - Add --memory-profile to the smoke workload, collecting alloc_space, alloc_objects, and inuse_space; left off by default because allocation profiling costs more than sampling - Record that these are per run rather than per transport: the lane tag scopes the CPU sampler and the allocation profiler does not consult it, so claiming otherwise would be wrong - Record that the goroutine, mutex, and block types Pyroscope advertises stay empty for a Python process, so an empty panel there is expected
why: The check reported that a published port reaches something other than this container, which is the hard half of the problem, and then stopped. The reader is left knowing they have a collision and not that the port is overridable, so the obvious next move is to start deleting containers. what: - Print the environment variable that republishes the affected service, and only for the two this stack actually moves; naming one for a service that keeps its upstream default would send the reader somewhere useless - Have up.sh show the same remedy when it refuses to report success
why: Notifications arrive out of band, so the instrumentation seam never sees them -- a sink wraps run(), and nothing routes a %output through run(). That left the streaming half of control mode unmeasured, including the engine's own count of notifications dropped when a subscriber falls behind, which is the number that says the stream is unhealthy. what: - Subscribe to notifications while commands keep flowing, recording those received and those dropped; 240 arrive in a three second lane with none dropped, and the commands are unaffected - Establish the subscription before generating output and pace the sends, because a tight command loop starves the consumer and it then waits for output that has already gone by - Give the lane its own shell window: the workload's panes run sleep, which ignores keystrokes and emits nothing - Read the counters with max_over_time rather than increase, since a value written once per run is flat and increase over it is zero
…inst why: The telemetry work had only ever run against one tmux, and control mode is the part most likely to drift between releases -- its notification set and client flags have both changed over the versions libtmux supports. A build that emitted nothing would leave the streaming panels blank with no clue why. Exercised on 3.2a, 3.3a, 3.4, 3.5, 3.6, 3.7, 3.7a and 3.7b: identical command and inlining counts on both transports, notifications delivered on every one, none dropped. what: - State the range in the README, since "works with tmux" is not a claim a reader can act on
why: A cold start failed with grafana UNREACHABLE, and the same check passed seconds later. The container reports healthy before every service inside it has finished binding, so verifying immediately catches Grafana mid-startup and refuses a stack that is fine. Shadowing is deliberately not retried: two different services answering the same port is a settled fact, not a timing question, and retrying it would only delay a report that will not change. what: - Retry a port that is not answering yet, up to LIBTMUX_LGTM_WAIT seconds; two consecutive cold starts now pass where they previously failed, and a healthy stack still verifies in under two seconds - Record what has been verified in scripts/lgtm/VERIFICATION.md: transport against signal, tmux 3.2a through 3.7b, load shapes, profile types, identity resolution, failure paths, and what is deliberately not covered - Recompute the panel total from the boards in a test, so the record cannot claim a number the dashboards do not define
why: Cleanup lived only in rampa's teardown hook. SIGINT reaches it, but SIGKILL runs no user code, so a killed run stranded its tmux server, that server's pane, and its scratch root, each holding a pty, with nothing left to reclaim them. Repeated interrupted runs walk toward pty exhaustion. Separately, the pane's holding command is a run-duration ceiling rather than a cleanup mechanism: when it exits the window closes, the last window closing ends the session, and the server goes with it. destroy-unattached off only survives detach. At sleep 600 any run past ten minutes lost its server mid-flight. what: - Reap roots whose owner is proven absent, at the start of the next run -- the one path that survives an exit running no user code - Stamp each root with its owner's pid and start time, so a reused pid cannot be mistaken for a live run - Leave a root alone while its owner runs, and leave a pre-owner-file root alone while a tmux still answers on its socket - Raise the holding command above any plausible --duration - Cover both directions, including the reused-pid case
why: The shipped scenarios run 10 and 16 seconds, so nothing in the recipes approaches a limit and a reader has no reason to suspect one. A long run used to climb in memory until it ended in swap rather than in a clean error, which presents as a CPU spike that never recovers. Bisecting it showed removing the OTel sink left the growth unchanged while removing rampa's per-iteration recording cut it by 70%: the cause was rampa's runner buffering every sample for outputs that were never configured, and it is fixed upstream. what: - State that duration is safe, and name the version the growth belongs to, so a climbing run has an obvious first thing to check - Explain why the pane's holding command outlasts any plausible duration
why: CI runs `mypy .`, which covers `scripts/`; the configured file list that plain `mypy` uses does not, so these went unchecked on developer machines and failed the gate the first time this branch had a pull request. Two of them were real: telemetry and the load scenario still reached command_count through control_mode, which re-exports it rather than defining it, so the import resolved at runtime but not under strict checking. The rest were missing type arguments and a literal inferred as list[object]. what: - Take command_count from libtmux.engines.base, its definition site - Give the acceptance script's dict annotations their arguments, and name the decoded JSON body rather than returning Any - Annotate the error thresholds so the generated panel keeps its shape - Narrow the git-dir probe explicitly, since bool() does not narrow Optional - Pass rampa a timedelta rather than a duration string: its field coerces the string through a validator, so the value is identical and the annotation stops lying - Name the async lane factory separately from the sync one, and mark the optional profiler import untyped
why: CI runs `mypy .`, which reaches `scripts/`, but installs only the default and dev groups -- the `otel` group is deliberately excluded so the ordinary gates stay lean. Every OpenTelemetry, pyroscope, and rampa import was therefore unresolvable in CI while resolving fine on a machine that had the group installed. what: - Allow those three imports to be missing, so the scripts type-check in both environments rather than only the one the developer happens to have - Relax subclassing-Any and untyped-decorator for the two modules that build on them, scoped per module: an inline ignore would itself be unused wherever the group is installed
why: Both tests guard on `opentelemetry`, but that namespace package arrives transitively, so the guard passes wherever libtmux is installed. What the modules under test actually import is `opentelemetry.sdk`, the OTLP exporters, and rampa, all of which ship only in the optional `otel` group -- which CI does not install. The guard therefore never fired and the tests failed on import instead of skipping, which was the stated intent. what: - Guard on `opentelemetry.sdk` rather than the namespace package - Guard the load scenario on rampa as well, since it imports both Verified in both directions: without the group the two skip, and with it installed they still run and pass rather than being silently disabled.
…ng to
`scripts/lgtm/` holds the observability stack -- its compose files, dashboard
generator, telemetry wiring, and load driver. The smoke run and the acceptance
check sat outside it under an `otel_` prefix, though they are the two entry
points that stack exists to serve: `acceptance.py` already documented itself in
terms of `scripts.lgtm.generate_dashboards`, and `generate_dashboards` returns
the reference.
scripts/otel_smoke.py -> scripts/lgtm/smoke.py
scripts/otel_acceptance.py -> scripts/lgtm/acceptance.py
Inside the directory the prefix was saying what the directory already says, so
it goes, matching `identity.py`, `telemetry.py`, and `load_tmux.py` beside
them.
Moving one level deeper broke two paths that had been counting directories,
both silently:
- `acceptance.py` derived the repository root with `parent.parent`, which now
lands on `scripts/`, pointing every dashboard and helper lookup one level
wrong.
- `smoke.py` put `__file__.parent / "lgtm"` on `sys.path` to import
`telemetry`; from inside `lgtm/` that names a directory that does not exist.
It now adds its own directory, the idiom `load_tmux.py` already uses.
`--help` used to say `otel_smoke.py`, which told a reader where to look. Bare
`smoke.py` does not, so both parsers name their path, as the orchestration
scripts do.
Verified: both run under the invocation their `just` recipes use, every
derived path resolves, and the lgtm tests pass.
Both cover `scripts/lgtm/`, so they sit beside it. Their repository-root lookups count directories, and the counts move with them. The dashboard generator and the stack's README each pointed at the dashboard test by its old name, which is the pairing this layout now makes checkable from the path. With these, no branch in the stack carries a `scripts/bench_*` or a `tests/test_bench_*`: one convention, everywhere, rather than one that arrives partway up.
tony
force-pushed
the
lgtm-observability
branch
from
August 23, 2026 18:09
734a31a to
3faee62
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
just otel-*, so the engine instrumentation seam has somewhere to send telemetry and a tmux workload's real cost becomes visible.run()-shaped seam cannot see, including the engine's own count of notifications dropped when a subscriber falls behind.Changes by area
The stack
scripts/lgtm/up.sh,verify.sh: pin the image, bind-mount provisioning, and refuse to report success when a published port answers from something other than this container. Liveness cannot tell a shadowed port from a healthy one; identity can.scripts/lgtm/telemetry.py: an OTel sink emitting spans and metrics together, so the duration histogram records while its span is current and carries an exemplar — that is the metric-to-trace pivot.scripts/lgtm/identity.py: resolves repository, ref, revision, and worktree from onegit rev-parseat startup, using OpenTelemetryvcs.*conventions, overridable by environment for detached CI checkouts.scripts/lgtm/generate_dashboards.pyanddashboards/*.json: Home, Overview, Commands, Transports, and Compare boards, with Overview organised as rate/errors/duration because an engine is a service and RED is the frame its caller thinks in.scripts/lgtm/mcp-config.sh: prints a client config carrying this stack's real ports and the running container's token, so an agent can query traces over MCP.Workloads
scripts/otel_smoke.py: drives all four transports, with grouped commands so the inlining panels are not flat zero and rejected commands so the failure panels have real data.--memory-profileadds allocation profiles, off by default because they cost more than sampling.scripts/lgtm/load_tmux.py: the rampa scenario, reaping roots whose owner is proven absent at the start of the next run — the one path that survives an exit running no user code.scripts/otel_acceptance.py: reads the committed dashboards, runs each panel's own query the way Grafana runs it, and probes the backends first so a forgottenjust otel-upcosts seconds rather than a socket timeout per panel.Benchmark-harness fixes
Teardown no longer reads tmux's "no server" exit as a failure, the stress ladder scales its watchdog with the rung it is building, a stalled test worker can no longer outlive the test that parked it, and two tests stop depending on machine load or the caller's environment size.
Design decisions
Load telemetry goes through this project's sink, not a rampa output backend. rampa's own backend would export under its service name and metric vocabulary, so a load-shaped run would arrive as a second account of the same work rather than a comparable one.
Identity is split by signal. Metrics carry only branch, run id, and spike — the dimensions worth grouping by. Revision and worktree ride on traces and profiles. A revision is never a comparison axis: each run has exactly one, so grouping by it is grouping by run, and copying every fact onto every signal mints a fresh set of series per commit.
Panels query the viewer's selected window, not a counter at an instant. The workload is short-lived, so Prometheus marks its series stale within minutes and an instant query then returns nothing beside a timeseries still drawing the same run. A test rejects any Prometheus target with no range selector, which catches this offline — a live check only sees it once the series has gone stale, which is precisely when nobody is watching.
Port collisions fail the start. Grafana and Prometheus are published on non-default ports because a host process already bound to their defaults still lets Docker publish, then answers first, so queries reach the wrong server and return plausible data.
Verification
Confirm every
justrecipe the README demonstrates actually exists:$ uv run pytest tests/test_lgtm_dashboards.py -k readmeConfirm no Prometheus panel target reads a counter at an instant:
Test plan
uv run ruff check .— lint cleanuv run ruff format .— formatting clean, tree unchangeduv run mypy— types cleanuv run pytest --reruns 0— suite clean apart from the load-sensitive stall family noted belowjust build-docs— docs build cleantest_lgtm_load.py— a stranded root is reclaimed, a live owner's root is left alone, and a reused pid is not mistaken for a live runtests/test_bench_orchestration_script.pycarries a stall/cancellation family that is load-sensitive above roughly load 20; those tests pass in isolation and their failing subset varies between runs on identical trees.