Skip to content

fix(foundation): resume spawn backoff after EINTR (#1616) - #1799

Merged
DeusData merged 1 commit into
DeusData:mainfrom
rarepops:fix/nanosleep-eintr-1616
Sep 2, 2026
Merged

fix(foundation): resume spawn backoff after EINTR (#1616)#1799
DeusData merged 1 commit into
DeusData:mainfrom
rarepops:fix/nanosleep-eintr-1616

Conversation

@rarepops

Copy link
Copy Markdown
Contributor

What does this PR do?

Adds cbm_nanosleep_full() to resume POSIX sleeps after EINTR, and uses it for subprocess spawn backoff so transient signals cannot silently shorten the documented retry budget. Other sleep call sites retain their existing interruptible behavior.

Adds a POSIX regression test that repeatedly interrupts the deterministic spawn backoff with SIGALRM and verifies the full cumulative delay is preserved.

Fixes #1616

Checklist

  • Every commit is signed off (git commit -s) - required, CI rejects unsigned commits (DCO, see CONTRIBUTING.md)
  • Tests pass locally (subprocess: 32 passed)
  • Lint passes (clang-format-20, cppcheck 2.20.0)
  • New behavior is covered by a test (reproduce-first for bug fixes)

Signed-off-by: Rares Popa <2606875+rarepops@users.noreply.github.com>
@rarepops
rarepops requested a review from DeusData as a code owner August 22, 2026 13:56
@github-actions

Copy link
Copy Markdown

Thanks for opening this — it has been seen, and it is queued.

This note is automated, but it is not a brush-off: it exists so you know where your PR stands instead of having to guess from silence.

Current review status: working through a backlog. 0.9.1-rc.1 is out, so the release freeze that held reviews is over — but it left a large queue of open pull requests behind it, and we are reading through them oldest-first. The background is in discussion #1144.

What that means for this PR, concretely:

  • It will not be closed for inactivity. No stale bot touches pull requests here.
  • It may still sit a while before a human reads it. That is on us, not on you.
  • Older PRs are read first, so a recent one is not being skipped — it is behind a queue.

Things that will genuinely speed it up whenever review does happen:

  • Keep it rebased on main — the tree is moving quickly right now, and a conflicting branch cannot be reviewed as the diff you intended.
  • Get CI green, or say which failures you believe are pre-existing.
  • Keep the change to one claim. Bundled features and refactors get split before they get merged, which costs you a round trip.
  • Every commit needs a sign-off (git commit -s) — CI enforces DCO.

If this fixes a bug, a reproduction we can run is worth more than a description of the symptom.

Thanks for contributing, and sorry in advance for the wait.

@DeusData DeusData added bug Something isn't working stability/performance Server crashes, OOM, hangs, high CPU/memory priority/high Needs near-term maintainer attention; high-impact bug, regression, safety issue, or release blocker. labels Aug 24, 2026
@DeusData

Copy link
Copy Markdown
Owner

Thank you for the focused EINTR fix and deterministic regression test. I checked current main: POSIX maps cbm_nanosleep directly to nanosleep, and the subprocess spawn backoff discards its return value while passing no remainder. That confirms a signal can shorten the documented retry budget.

I have labeled this as a high-priority stability bug and routed it for review. CI is green. Our review queue is full, so detailed review may take a little time, but the change is now classified and visible. Thank you for limiting the new full-sleep behavior to the spawn-backoff path.

@DeusData
DeusData merged commit 1424934 into DeusData:main Sep 2, 2026
35 checks passed
@DeusData

DeusData commented Sep 2, 2026

Copy link
Copy Markdown
Owner

Merged as 14249349. That is your third landing today — thank you.

Two things about how this one got verified are worth telling you, because the usual route was not available.

Your CI was 35/35, but the branch was 205 commits behind main, so that green no longer described the code that would actually land. Normally I would rebase and re-run — except our Actions pool has been servicing roughly one job at a time against a 35-deep queue all day, so a rebase would have cost you hours and everyone else a slot.

So instead I sized the risk with evidence rather than guessing: exactly one commit on main had touched any of your four files since your base, and it touched only tests/test_subprocess.c, adding an unrelated Windows-only Job Object test. compat.c, compat.h and subprocess.c had no changes at all.

Then I built the actual merge result locally — current main plus your branch — and ran the affected suites. Clean build, 281 passed, 5 skipped, 0 failed. When #1779 landed and invalidated that, I redid it against the new main and included the mcp suite too, since that merge touched mcp.c. Still clean. That is stronger evidence than a fresh remote green would have been for the platforms it covers, and it cost the queue nothing.

On the fix itself: the detail that made it easy to accept is that the old call passed NULL for the remainder, so the unslept time was discarded rather than merely shortened. With the backoff reaching 2.56 s at shift 8, a signal could collapse the wait to almost nothing — tightening the retry loop under precisely the pressure it exists to ride out. Confining the change to the one call site and leaving other sleeps interruptible is what kept the blast radius honest.

The fixing SHA is recorded on #1616.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bug Something isn't working priority/high Needs near-term maintainer attention; high-impact bug, regression, safety issue, or release blocker. stability/performance Server crashes, OOM, hangs, high CPU/memory

Projects

None yet

Development

Successfully merging this pull request may close these issues.

cbm_nanosleep never resumes after EINTR: every timed wait can be silently truncated

2 participants