feat(fastapi): attribute per-request energy by CPU time and exclude idle power - #1430
Draft
davidberenstein1957 wants to merge 18 commits into
Draft
davidberenstein1957 wants to merge 18 commits into
davidberenstein1957 wants to merge 18 commits into
Conversation
…ervers Observers now receive cumulative CPU, GPU and RAM energy, the source quality of each component, and the CPU idle power when the power model fixes it (load and constant modes).
Each window is split per component into idle and dynamic energy. Idle power is the lower of a rolling minimum of window power and a good regression intercept against CPU utilisation, or the analytic value in load and constant modes. Idle and RAM energy go to a new idle_kwh bucket.
The share is the process's CPU time over the machine's busy CPU time from psutil.cpu_times(), clamped to [0, 1]. The rest goes to a new other_processes_kwh bucket. Load mode in process tracking is already per-process and keeps all of it.
The middleware drives the app coroutine through an awaitable that sums time.thread_time_ns() around every resumption. Each window, a request gets our CPU dynamic energy in proportion to the CPU time its meter saw; CPU time no meter claimed goes to process_unattributed_kwh. RequestEnergy gains cpu_seconds.
…request RequestEnergy.energy_kwh is now the CPU energy above idle charged by CPU time; GPU energy above idle is split by wall-clock overlap into a separate gpu_kwh. Each request reports its attribution_method and the weakest source quality it was charged from. Document the buckets and the limits of the method.
CI checks conservation over random fake windows, cpu_seconds of a known async CPU burn, and a sleeping endpoint's near-zero CPU time. A sync burn is marked xfail: worker-thread CPU time is not metered yet. fastapi_accuracy.py holds the opt-in RAPL checks (-m accuracy) and skips without readable RAPL; fastapi_overhead.py times 100k ASGI calls with and without the middleware.
A window skipped for a backwards counter left the meters' CPU time to be charged in the next window, against a process time that did not include it.
Wrap anyio.to_thread.run_sync process-wide while a tracker is attached. The wrapper only acts when a request meter is set in the current context, so sync endpoints and dependencies, sync streaming iterators and plain Starlette sync endpoints are metered, including routes from include_router. The patch is reference-counted, idempotent and restored on detach; meter_threadpool=False turns it off. On Linux a running worker call is read through its thread CPU clock, so long sync calls are charged window by window.
A request is charged cpu_seconds times the cost of a CPU-second: the power model's in load mode, the slope of a good power-on-load fit when measured, else the process average as before. Charges are capped at the process's share of dynamic CPU energy; what the cost leaves over goes to other_processes_kwh. The conservation invariant still holds exactly. This stops per-request energy from following other processes' load on a convex power curve.
A binding cap used to hand all of our share to the requests, leaving nothing for process CPU time no meter claimed. One cap now scales both.
ASGI only promises an awaitable, so drive its __await__ iterator instead of calling send/throw/close on the object itself.
Build the observer sample inside the same guard as the observers, so a bug in an integration's sample can't stop the core measure and flush.
Otherwise a concurrent total_ns() could see more of the running call than is then added, and the meter would go backwards.
A window with no width now advances the meters like a skipped one, and counts as skipped when energy arrived in it.
Each observer is bound to its tracker, and a sample from a tracker that is no longer attached is dropped instead of mixing into the new anchor.
A window without machine CPU times kept charging by the previous fit.
State idle for each CPU mode, the lower-of idle estimate, where unclaimed process CPU goes, what unattributed_kwh holds, the anchor sample race, nested middleware, the run_sync chain drop and child processes in process-tracking load mode.
Codecov Report❌ Patch coverage is
Additional details and impacted files@@ Coverage Diff @@
## feat/add-fastapi-middleware #1430 +/- ##
===============================================================
+ Coverage 91.74% 91.92% +0.18%
===============================================================
Files 52 52
Lines 5352 5646 +294
===============================================================
+ Hits 4910 5190 +280
- Misses 442 456 +14 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
13 tasks
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Makes the FastAPI middleware's per-request energy an estimate of what each request actually caused, instead of an even split of the whole machine's energy. Stacked on #1203; review that one first.
Today (#1203) each sampling window's energy, idle power included, is split across the requests in flight by how long they overlapped the window. A lone 1 ms request in an idle window gets the whole window; a request waiting on a database counts the same as one using the CPU; other processes on the host are charged to requests.
With this PR, per window:
idle_kwh). Idle is 10% of TDP in CPU load mode, 0 W in process-tracking load mode, the whole CPU power in constant mode, and otherwise the lower of the lowest window power seen and a fitted intercept.time.process_time()against busy CPU time; the rest goes toother_processes_kwh.anyio.to_thread.run_syncthat does nothing outside a metered request.meter_threadpool=Falseturns it off.gpu_kwh.attributed + idle + other_processes + process_unattributed + unattributed == settled, up to float rounding.New per-request fields:
cpu_seconds,gpu_kwh,attribution_method(cpu_time,wall,mixed) andquality(measuredfor RAPL/powermetrics/NVML,modeledfor load mode,none).energy_kwhnow means the request's CPU energy above idle; the field has not been released yet.Merge plan: merge together with #1203, after one validation run of
pytest -m accuracy tests/integrations/fastapi_accuracy.pyon Linux with readable RAPL.Related Issue
Fixes #1428. Stacked on #1203.
Motivation and Context
Users read
energy_kwhper request as "what this request cost". The overlap split can be off by orders of magnitude for short or I/O-bound requests. This keeps the number conservative and says how it was obtained.How Has This Been Tested?
uv run pytest tests/integrations/test_fastapi.py: 46 passed, 1 skipped (a Linux-only test for long sync calls).uv run task test-package: 702 passed.include_router, plain Starlette endpoints;dependency_overridesstill work; anasyncio.sleependpoint gets almost no CPU.Not yet run: the accuracy checks against RAPL (
pytest -m accuracy tests/integrations/fastapi_accuracy.py) need a Linux machine with readable RAPL; they skip elsewhere. In CPU load mode the check "per-request energy barely moves when another process loads the machine" fails (+289%) because the cubic power model has no single per-CPU-second cost; the docs say so and the check only runs against a measured source.Known limits
create_task,gather) is not metered and lands inprocess_unattributed_kwh.run_syncover ours, threadpool metering stops until restart.CodeCarbonMiddlewareinstances share one meter; the outer one loses threadpool CPU.Screenshots (if appropriate):
N/A
Types of changes
AI Usage Disclosure
Please refer to docs/how-to/ai-policy.md for detailed guidelines on how to disclose AI usage in your PR. Accurately completing this section is mandatory.
Checklist:
Commits