feat(index): add worker resource watchdogs - #1724
Conversation
|
Thanks for opening this — it has been seen, and it is queued. This note is automated, but it is not a brush-off: it exists so you know where your PR stands instead of having to guess from silence. Current review status: working through a backlog. What that means for this PR, concretely:
Things that will genuinely speed it up whenever review does happen:
If this fixes a bug, a reproduction we can run is worth more than a description of the symptom. Thanks for contributing, and sorry in advance for the wait. |
ece9042 to
eaab28c
Compare
eaab28c to
a1339ec
Compare
Indexing accepts whatever a repository contains. A tree carrying a vendored monorepo, a generated dump, or a runaway build directory is discovered in full, and the first sign of trouble is a host under memory pressure with nothing that attributes it to indexing. Add two opt-in limits evaluated during discovery against accepted source files only: index_max_files and index_max_source_mb. Both default to off, so nothing changes until an operator sets one. Crossing a limit fails the whole attempt with a structured resource_limit_exceeded result naming the resource, the observed value and the limit; no partial graph is published, and an existing serving index keeps answering. Limits are read from the CLI-managed _config.db and are not MCP request arguments. A supervised parent replaces any caller-supplied policy before spawning its worker, and the worker rejects a missing or incomplete contract, so the CLI, the daemon and the supervised worker all enforce the same decision. Both keys reach an operator through the existing config get/set/list/reset with no new subcommand. `set` suppresses its own generic message for them because the policy writer names the precise reason -- so that writer speaks on every failure it can return, including a validated value whose write then fails on a database that cannot be written. Exiting non-zero in silence is not an acceptable answer from a CLI. The two shell regressions that hand-roll the supervisor's worker argv carry that contract as well. Without it the worker exits before either guard can observe anything, and the guard would go quietly vacuous. Signed-off-by: 刘冲 <mail@liuchong.dev>
Discovery limits bound what indexing accepts, not what it then costs. A repository well inside those bounds can still exhaust the host through parser memory, or simply never finish, and a supervised worker that hangs leaves the parent waiting with nothing to report. Add index_max_rss_mb and index_max_duration_seconds, enforced by the parent against the worker process tree rather than the worker process alone, so a runaway child cannot hide behind a small parent. Resident memory is sampled through the platform interface on macOS, Linux and Windows. Crossing a limit terminates the tree and yields one trusted, structured terminal result that attributes the failure to the resource that caused it. Both limits default to off. A measurement that cannot be taken fails the attempt instead of passing it: a watchdog that quietly stops watching is worse than no watchdog at all. The shell fixture that stands in for the supervisor names the two new keys. The worker accepts only a policy that spells out every key it knows, which is what keeps a stale supervisor from starting a worker it cannot bound. Signed-off-by: 刘冲 <mail@liuchong.dev>
a1339ec to
38ec2e4
Compare
|
Reviewed slice 2. The direction is already accepted (see #1723); this is the merit read, and the design holds up well. The three-valued status is the right call and it is the thing most people get wrong here. CBM_PROC_TREE_RSS_ERROR = -1,
CBM_PROC_TREE_RSS_EMPTY = 0,
CBM_PROC_TREE_RSS_OK = 1,"No trustworthy measurement" and "the tree has no observable members" are different facts, and collapsing them into "A watchdog that quietly stops watching is worse than no watchdog" is the right principle and you built to it. Using a Windows Job Object is the detail I want to single out. The error paths around it are thorough too: allocation failure, size overflow past Measuring the tree rather than the root process is the correct scope, and your reason for it — "a runaway child hiding behind a small parent is exactly the case that goes unattributed today" — is the sentence that justifies the extra complexity. Ignoring processes that exit mid-enumeration is the right way to handle the inherent race. Exposing Process notesClearance: this shows Rebase when #1723 lands. This diff currently carries #1723's commit; once that merges it collapses to Slice 3 onward is where the configuration surface really grows, so expect those to get a closer look than this one — but this slice is in good shape. |
Related to #1347.
Problem
Discovery limits bound what indexing accepts, not what it then costs. A repository well inside those bounds can still exhaust the host through parser memory, or simply never finish. A supervised worker that hangs leaves the parent waiting with nothing to report and no way to attribute the stall.
What this changes
Two opt-in limits, enforced by the parent against the worker process tree:
index_max_rss_mbindex_max_duration_secondsBoth default to
off.Measuring the tree rather than the worker process alone matters: indexing spawns children, and a runaway child hiding behind a small parent is exactly the case that goes unattributed today. Resident memory is sampled through the platform interface on macOS, Linux and Windows.
Crossing a limit terminates the process tree and produces one trusted, structured terminal result attributing the failure to the resource that caused it, so the caller learns which limit ended the attempt rather than seeing a generic worker failure.
A measurement that cannot be taken fails the attempt instead of passing it. A watchdog that quietly stops watching is worse than no watchdog, so the probe is fail-closed and says so in the result.
Testing
make -f Makefile.cbm testandmake -f Makefile.cbm lint-cion macOS. New coverage: process-tree RSS measurement, watchdog termination on each limit, probe-failure fail-closed behaviour, terminal-result attribution, and policy validation for the two new keys.Stack