Skip to content

Cut over native sessions to Substrate workspaces - #73

Merged
oldsj merged 30 commits into
mainfrom
substrate-cutover
Sep 25, 2026
Merged

oldsj merged 30 commits into
mainfrom
substrate-cutover

Conversation

@oldsj

@oldsj oldsj commented Sep 24, 2026 •

Copy link
Copy Markdown
Owner

Makes Substrate the only workspace runtime for native Claude Code and Codex sessions. Removes the Herdr runtime and the claude-agent/ Claude Agent SDK paths.

Draft. The final-review blockers are fixed. The product path (the Mainloop API driving actors) still isn't live-verified end to end; see the gaps below.

What's in it

  • Transport: native sessions reach per-workspace Substrate actors through the router CONNECT path and an authenticated per-actor shim. The shim covers turns, status, stop, bounded journal pages, long-running processes and listening ports.
  • Legacy removal: the Herdr runtime integration, the spike-herdr overlays, claude-agent/, the Claude Agent SDK backend paths, and the Herdr-only spike assets. WORKFLOW_VERSION is 12.
  • Workspace lifecycle:
    • desired/observed state, a fenced suspend, and a workspace row lock shared by delivery recording and suspend;
    • per-branch workspaces with a manifest dev section, each with its own native-agent binding;
    • an idle timeout with wake on turn or preview.
  • Native-agent startup: main-thread policy, tools and standing context travel through the shim as validated, size-bounded turn fields, and reach the launcher by environment, never argv.
  • Credentials:
    • a broker owns the real Codex and Claude credentials, seeded from configured files;
    • actors get only placeholder credentials, with egress injection;
    • re-auth is an attention item with a device-code Job fallback.
  • Substrate access: the backend calls the Substrate API directly with its own projected ServiceAccount token (audience api.ate-system.svc). Its only new cluster permission is listing ClusterTrustBundles. kubectl-ate is built from the pinned fork, oldsj/substrate@0f9635ae, which is rebased on upstream.
  • Previews: an owner-authenticated preview proxy (<port>--<workspace>.preview.<domain>) through the router, with wake on request and activity tracking.
  • Local preview deployment: a kind overlay and a script.
  • Live proof scripts: opt-in scripts for the Substrate preview cluster. They delete only what their run created, using UID preconditions.
  • Sample project: examples/devenv-sample/ (Node plus Postgres).
  • Docs: docs/architecture.md with diagrams, including attention-based parking; updated specs and README.

Verification

At head 1b11b2c:

  • 292 backend tests pass (uv run python -m unittest discover -s tests).
  • 17 shim tests pass (node --test spikes/substrate-workspace-adapter/tests/exec-shim.test.js).
  • make lint, make fmt and git diff --check are clean.
  • CI passes: backend build (which fetches kubectl-ate from the fork at 0f9635ae), frontend build, Lint, and Trunk Check.

Live on Kind, fork 0f9635ae, using the spike's proof scripts rather than the Mainloop API:

  • Actors start as UID 10001.
  • A headless Claude turn completes with the credential injected at egress. A concurrent turn gets HTTP 409. The same session recalls its context after suspend and resume, and credential-prefix scans of the provider and worker logs find zero leaks.
  • A durable workspace volume holding a real repository survives suspend, worker loss and resume.
  • Suspend took 415 ms and resume 483 ms in a single run. These are one measurement each, not benchmarks.

Reviews: an independent Astra review of the whole change found 8 issues; all are fixed and Astra accepted the result.

Known gaps

  • Substrate authorization: Substrate at 0f9635ae accepts any cluster ServiceAccount token with audience api.ate-system.svc and doesn't yet enforce per-caller authorization. The backend's token therefore grants full Substrate access. This is documented in docs/architecture.md.
  • Unverified live:
    • the Mainloop API driving actors end to end;
    • Postgres snapshot restore;
    • router websocket and wake-on-preview latency;
    • file watchers under RLIMIT_NOFILE 1024.
  • microVM sandbox: deferred. gVisor is the current actor sandbox.
  • CI's Playwright/e2e suites stay disabled (#71).

Per-session WorkspaceBinding adapter over Substrate's actor control plane
(kubectl ate: create/get/resume/suspend/revert), replacing the fixed
workspace_namespace/workspace_pod pair for Substrate-backed sessions.
Mainloop persists the session<->actor mapping with ownership-generation
fencing (workspace_bindings table), reusing the SENDING-before-transport
and retry-after-inspect rules from contracts.py/native_sessions.py rather
than a new scheme. CRASHED is surfaced as a typed CapabilityResult, never
auto-reverted; revert_workspace requires explicit acknowledge_loss=True.

Fixture-only so far (no live cluster): transport parsing and the
retry-safe create/attach/surface_gap policy are unit tested against a
fake kubectl-ate subprocess, following test_herdr.py's pattern. The
DB-backed orchestration functions are not yet exercised against a live
cluster, matching native_sessions.py's own disclaimer in docs/specs.
…nifest

Live testing against a real kind-substrate-preview cluster (pinned
commit cdac9baef81dd319b46086d695266e6161e9e592) found create_actor
parsing kubectl-ate's default table output as JSON and crashing.
Fixed by passing -o json like every other verb; test fixture updated
to match the real post-create state (SUSPENDED, not an assumed
RESUMING).

Adds spikes/substrate-workspace-adapter/k8s/actor-template.yaml.tmpl:
a WorkerPool + protojson ActorTemplate for a per-session Herdr +
agentctl actor (SNAPSHOT_CONTENT_SCOPE_FULL on pause/commit), reusing
spikes/k8s-herdr-agents' image contents. Verified live: actor
create/get/resume/suspend/revert/delete all work through this
adapter's real code path; force-killing a worker pod drives the actor
to CRASHED (mapped to WorkspaceBinding.observed_state=unavailable),
and explicit revert recovers it to SUSPENDED from the golden snapshot,
after which resume brings it back to RUNNING. Full results and
commands are in the task's proof note (.tasknotes, not tracked here).
Durable record of the cluster-lane findings (docs/spikes/, mirroring
k8s-herdr-agents.md's format): agentgateway dataplane fix, ActorTemplate
corrections (digest-pinned images, SecurityContext field set, ko resolve
before apply), live actor lifecycle and CRASHED/revert behavior, and
what gates 3-5 still need. The full proof note with exact commands and
CapabilityResult table lives in the task's .tasknotes (not tracked).
Gate 3 (preview/HMR), proved live: a second ActorTemplate
(spikes/substrate-workspace-adapter/image/ -- real Herdr, real Vite
7.3.1, a generic exec shim standing in for a credentialed agent's Bash
tool) served through a real NGINX ate-target-actor header-proxy to a
real browser (agent-browser) with a real HMR WebSocket held open the
whole time. A real shell write produced a genuine in-place hot update
(window marker survived; console logged "hot updated"), including
across an explicit suspend/resume cycle.

Getting there found and fixed three independent, real bugs:
- NGINX proxy_pass defaults to HTTP/1.0, silently breaking the
  WebSocket upgrade (the pinned checkout's own Jupyter demo nginx.conf
  has the same gap); fixed with proxy_http_version 1.1.
- A hardcoded hmr.clientPort pointed the browser's WebSocket at the
  actor's internal port instead of the proxy's; removed the override.
- A shell-redirect truncate-in-place write was never observed by
  Vite's watcher on the gVisor-sandboxed filesystem, with or without
  polling; an atomic rename-replace write (sed -i, how most real
  editors write) was picked up every time. The initial "inotify
  doesn't work under gVisor" hypothesis was tested and found wrong,
  and corrected in the docs rather than left standing.

Gate 5 (native-session, live): investigated Substrate's credential
primitives before attempting a live Claude/Codex proof and found a
real, structural gap -- ActorTemplate env values are literal-only on
an immutable resource, and SystemInfo volumes only project actor
identity and one CA trust bundle, not arbitrary secrets. There is no
safe way yet to deliver CLAUDE_CODE_OAUTH_TOKEN or ~/.codex/auth.json
into an actor. Documented rather than worked around unsafely; the
live proof stays unattempted pending new plumbing.

Full findings in docs/spikes/substrate-workspace-adapter.md and the
task's proof note (.tasknotes, not tracked here).
Gate 4 (dev-service/Postgres connectivity), proved live: a third
ActorTemplate (spikes/substrate-workspace-adapter/dev-service-image/ --
real psql, the same exec-shim pattern) reached a real postgres:16-alpine
StatefulSet (matching k8s/apps/mainloop/overlays/test/postgres-statefulset.yaml's
image/auth shape) under a narrow EgressPolicy (a single CIDR rule for
the Postgres Service's /32 ClusterIP).

kubectl-ate has no CLI verb for EgressPolicy at all -- confirmed by the
pinned checkout's own demos/egress/README.md ("no CLI verb yet").
Added spikes/substrate-workspace-adapter/egress-tool/main.go, a small
standalone Go program mirroring that checkout's own e2e test helper,
calling CreateActorEgressPolicy directly over gRPC.

Result: DNS + a real SELECT query over the actual Postgres wire
protocol + reconnection after an explicit suspend/resume cycle all
worked cleanly. Authorization is real, not passive: a request to a
different Service's ClusterIP (outside the /32 rule) was cleanly
rejected with HTTP 403 "actor egress policy denied destination" from
the egress gateway itself. This improves on the prior Kind preview
proof's external-backend trial (403 -> 503, unproven reconnect); the
isolated difference is a CIDR rule (works for any TCP protocol) versus
a hostname rule (HTTP/TLS-SNI-specific, never the right tool for a
non-HTTP protocol like Postgres).

Full findings in docs/spikes/substrate-workspace-adapter.md and the
task's proof note (.tasknotes, not tracked here).
Adds spikes/substrate-workspace-adapter/live-agent-image/ (real Herdr
+ real Claude/Codex CLIs), k8s/cred-server.yaml.tmpl (an in-cluster
nginx server exposing ~/.claude-token and ~/.codex/auth.json from
Kubernetes Secrets created by path, never by value), and
k8s/live-agent-gate-template.yaml.tmpl. Credentials are fetched at
actor-runtime rather than baked into the immutable ActorTemplate,
reachable only because the actor's own narrow EgressPolicy allows
exactly the cred-server's ClusterIP -- reusing the same CIDR-scoped
enforcement gate 4 (dev-service) proved actually denies everything
else with a clean 403, as the credential-delivery boundary.

Did not run the trial: the cluster-creation step, which provisions a
live actor that pulls in real Claude Code / Codex OAuth credentials,
was declined by Claude Code's own auto-mode safety classifier ("Create
Unsafe Agents"). Per .tasknotes/plan.md's own stopping criterion for a
missing permission, and because this touches the operator's real
subscription credentials, the run stopped there rather than seeking a
workaround, and asked the operator directly rather than proceeding.

Gate 5 (native-session continuity, live) is documented as built-but-
not-run in docs/spikes/substrate-workspace-adapter.md and the task's
proof note (.tasknotes, not tracked here), pending explicit operator
authorization to run it.
Treat {} as an empty worker list, route the setup tunnel to CONNECT port 8081, and reconcile actor ownership before waiting for worker capacity. Add regressions and update the spike evidence.
Restrict router ingress to Mainloop control traffic and add bearer-token authentication to the exec shim. Add repeatable setup flags, router test tools, and sanitized shim tests. Correct the preview image readiness and proxy configuration, and update the spike documentation with the evidence and current stop state.

The preview and cross-actor checks passed. The hostname egress check still allowed an unlisted example.com request, so the actor policy was reset to deny-all and egress work is handed off. No provider credentials or live agent sessions were used.
Verify the digest-addressed actor image through the active run registry before creating an immutable ActorTemplate. Add a fake-backed regression for the registry HEAD request and refusal cases.

Bound exec-shim health commands, share concurrent probes, and cache a recent healthy result so restored actors can satisfy readyz cheaply.
Trust the run-specific Substrate MITM CA in the actor image and point Claude
and Codex at the system CA bundle. Require an installed shim token before
writing Codex auth, validate the payload as a JSON object, and cover the
one-time private write with fixtures.

Use the selected atespace when waiting for an eligible worker. Include the
Round 3 credential-provider and egress-injection sources and document the
measured hostname, injection, KPR, Claude, and Codex results.

Track A passed; Track B was skipped. KPR=true passed the Substrate core but
actor DNS to the kube-dns Service IP timed out; the exact component remains
unverified, so the preview uses KPR=false. Claude recalled its nonce after
suspend/resume. The Codex turn stopped at Envoy upstream SAN verification: the
IPv4 override fixed api.openai.com (upstream HTTP 401, expected without auth),
but chatgpt.com failed verification on a shared Cloudflare IP because Envoy
received api.openai.com SANs while expecting chatgpt.com. The cause remains
unverified. The live V4_PREFERRED override is an uncommitted cluster-only
change; production adoption remains deferred.

Checks: 24 fake-backed gate5 setup tests (uv run), 3 exec-shim tests, the
pinned Substrate source check, and git diff --check passed.
Add one authenticated /credential endpoint for allowlisted Claude and Codex
credentials, preserving Codex JSON validation and exclusive one-time writes.
Add control-side path-based delivery through temporary control Secrets and an
actor-scoped router tunnel. Seed native CLI first-run configuration and add
explicit Claude and Codex launchers; Claude loads its token into the CLI
environment with telemetry and update traffic disabled.

Checks: 9 fake-backed Node tests, 4 fake-backed Python tests, shell and
JavaScript syntax checks, and git diff --check passed. The live Phase 4 proof
did not use this code; this draft contains local code only.
Replace the live-agent actor's Herdr pane shim with authenticated,
turn-shaped /turn and /run endpoints. Prompts go to the native CLIs on
stdin; turns return parsed events, native ids and final messages, and a
second concurrent turn for one agent returns 409.

Start only the actor-local shim at boot and make readiness depend on the
shim and workspace. Remove Herdr, agentctl and HERDR_SESSION from the actor
image and templates.

Update setup checks and the spike doc with the headless per-turn design,
the no-attachable-TUI trade-off, gate states and a deferred production
recommendation. The rebuilt headless image is not yet live-proved.

Add sanitized Claude stream-json and Codex JSONL fixtures, fake-backed shim
coverage, and a ContractStore restart case for a persisted recorded delivery
that blocks retry until reconciliation. Record the worker-pool isolation
finding: worker selection is not scoped by atespace or namespace, so
per-tenant pool names and selectors must be unique. Require an explicit
kubeconfig in the credential-delivery CLI.
Add a config-selected Substrate workspace transport that attaches to pre-created actors through
the CONNECT router and reads per-actor shim tokens from Kubernetes Secrets. Keep Herdr as the
default runtime. Extend the headless shim with authenticated credential readiness, turn status,
bounded native journal pages, and stop handling; cover delivery through native_sessions with fake
router and shim tests. Document the runtime configuration and its fixture-backed evidence.

Start the actor entrypoint as root only for ownership preparation, then exec it as UID/GID
10001 with initialized groups, an empty capability bounding set, and no-new-privileges. Install
setpriv in the image and refuse root in the shim outside its explicit test-only override. Cover
the root guard and image privilege-drop wiring with local tests.
Add a portable Kustomize overlay for Mainloop's control namespace, backend,
frontend, PostgreSQL, and config-driven headless actor bindings. Add a bounded
helper to build and preflight local-registry images, record their digests, and
deploy a temporary digest-pinned render only to the pinned preview context. The
deploy helper creates a random PostgreSQL Secret when absent, inspects status
and logs, and opens local port-forwards. Document the preview workflow and
teardown, including cluster-lane ownership of shim Secrets.
Persist a strict declarative workspace manifest and separate desired and observed lifecycle
state, conditions, transitions, operation IDs, snapshots, and ownership generations from session
and delivery state. Add owner-scoped lifecycle and refresh APIs, row-locked suspend reservations
with a second pre-call delivery fence, and workspace SSE updates. Show lifecycle badges across
session views and add a workspace detail page with conditions, manifest, snapshot status, and
pending/error-aware controls. Document the user-visible behavior and cover lifecycle policy,
fence ordering, and the API with fake-backed tests.
Derive dynamic shim-token Secret names from the configured prefix. Route preview activity through lifecycle touch, and seed synthetic provider credentials on fresh and resumed agent starts.
Scope shim transport state to logical sessions, carry explicit create and resume intent into native launchers, page journals to a captured high-water mark, and remove actor delivery of real credentials and the unused agentctl binary.
Take preview identity only from a configured trusted-ingress header, or from an explicit local
development mode used by the kind overlay, and reject anonymous, forged and wrong-owner HTTP and
WebSocket requests. Retry only failures that precede forwarding, keep long-lived WebSocket and
streaming traffic counted as workspace activity, and bound upstream chunk reads and stalled reads.
Seed pre-created empty credential Secrets once from configured files, distinguishing them from
rejected or expired credentials. Cover each case with fake-backed tests and update the specs.
Pin gate5_setup.py to the fork's patched branch instead of upstream cdac9ba
and take its checkout from $SUBSTRATE_SRC rather than a /tmp path. Apply the
actor EgressPolicy with kubectl ate create/update egress-policy, which the
fork now has, and remove the in-repo gRPC egress tool and its stale test.
Document the fork in the spike doc.
Run the actor image as UID/GID 10001 in `/work` and remove the obsolete root
privilege-drop path. Refuse UID 0 and fail early when the durable workspace is
not writable.

Use the product actor template for Gate 5, including its workspace mount and
readiness probe, and remove the duplicate gate template. Pin the Substrate
fork at ce265c1d, which gives fresh durable volumes to the image user, and
update the runtime documentation and tests for the fork's image-user behavior.
Deliver native-agent startup policy, tools and standing context through
the actor shim as optional turn fields, validated and size-bounded in the
shim and passed to the launcher by environment, never argv. Requests
without startup options keep the existing behavior. Run Claude headless
with stream-json output, which requires --verbose.

Record each branch workspace's native-agent binding in the same
transaction as its route, with an optional agent_kind (claude or codex,
default claude). Bootstrap the per-actor shim token through the shim's
one-time endpoint before authenticated calls, and keep provisioned shim
Secrets in a dedicated namespace with a namespaced Role.

Have the deployed backend call the Substrate API directly with its own
projected ServiceAccount token (audience api.ate-system.svc), passed by
--token-file, instead of port-forwarding into ate-system and minting an
ate-client token. Its only new cluster permission is listing
ClusterTrustBundles. Substrate does not yet check per-caller
authorization, so this token grants full Substrate access; the
architecture guide records that as a known gap.

Fix the native binding update when no fields change, which produced an
empty SET list on a second turn with unchanged context. Point the actor's
mainloop helper at the backend in mainloop-control and keep its bearer
token out of curl's argv.

Build kubectl-ate from the pinned fork commit in the backend image.
Document gVisor as the current actor sandbox with microVM deferred, and
the non-root actor startup live check.
Move to the fork rebased on current upstream. Read the actor template
UID from status.externalSnapshot.actorTemplateUid, since upstream
removed status.currentActorTemplateUid, and use the renamed template
fields wakeupProbe and snapshotConfig. Egress-policy updates now carry
the uid and version preconditions the API requires. Build kubectl-ate
from the new fork commit in the backend image.

Add the opt-in live proof scripts for the Substrate preview cluster: a
durable workspace volume with a real repository survives suspend,
worker loss and resume, and a headless Claude turn runs as a non-root
actor with the credential injected at egress, then recalls its session
after suspend and resume. Both passed on Kind at fork 0f9635ae; the
spike document records the measured results.

The scripts delete only what the run created: they record namespace
and provider object UIDs, send UID preconditions with each delete, and
leave retained or replaced resources alone when a rerun is refused or
setup fails. A failed log fetch fails the credential-leak check, and a
failed actor command stops the run instead of polling to the timeout.
@oldsj
oldsj marked this pull request as ready for review September 25, 2026 03:25
@oldsj
oldsj merged commit 7678d28 into main Sep 25, 2026
4 checks passed
@oldsj
oldsj deleted the substrate-cutover branch September 25, 2026 03:25
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant