Skip to content

feat(fastapi): attribute per-request energy by CPU time and exclude idle power - #1430

Draft
davidberenstein1957 wants to merge 18 commits into
feat/add-fastapi-middlewarefrom
feat/fastapi-cpu-time-attribution
Draft

davidberenstein1957 wants to merge 18 commits into
feat/add-fastapi-middlewarefrom
feat/fastapi-cpu-time-attribution

Conversation

@davidberenstein1957

@davidberenstein1957 davidberenstein1957 commented Sep 28, 2026 •

Copy link
Copy Markdown
Collaborator

Description

Makes the FastAPI middleware's per-request energy an estimate of what each request actually caused, instead of an even split of the whole machine's energy. Stacked on #1203; review that one first.

Today (#1203) each sampling window's energy, idle power included, is split across the requests in flight by how long they overlapped the window. A lone 1 ms request in an idle window gets the whole window; a request waiting on a database counts the same as one using the CPU; other processes on the host are charged to requests.

With this PR, per window:

  • Idle power is removed first and reported separately (idle_kwh). Idle is 10% of TDP in CPU load mode, 0 W in process-tracking load mode, the whole CPU power in constant mode, and otherwise the lower of the lowest window power seen and a fitted intercept.
  • Only this process's share of the CPU energy above idle is kept, from time.process_time() against busy CPU time; the rest goes to other_processes_kwh.
  • Each request is charged its own CPU time at one cost per CPU-second (fitted from measured windows, or the power model's slope in load mode). CPU time is metered around every resumption of the request's coroutine and for work it sends to the threadpool (sync endpoints and dependencies, sync streaming iterators), through a patch of anyio.to_thread.run_sync that does nothing outside a metered request. meter_threadpool=False turns it off.
  • One cap applies to requests and to the process's unclaimed CPU time together, so requests are never charged for the server's own background work.
  • GPU energy has no per-request signal; it is split by overlap and reported separately as gpu_kwh.
  • Every window's energy is accounted for: attributed + idle + other_processes + process_unattributed + unattributed == settled, up to float rounding.

New per-request fields: cpu_seconds, gpu_kwh, attribution_method (cpu_time, wall, mixed) and quality (measured for RAPL/powermetrics/NVML, modeled for load mode, none). energy_kwh now means the request's CPU energy above idle; the field has not been released yet.

Merge plan: merge together with #1203, after one validation run of pytest -m accuracy tests/integrations/fastapi_accuracy.py on Linux with readable RAPL.

Related Issue

Fixes #1428. Stacked on #1203.

Motivation and Context

Users read energy_kwh per request as "what this request cost". The overlap split can be off by orders of magnitude for short or I/O-bound requests. This keeps the number conservative and says how it was obtained.

How Has This Been Tested?

  • uv run pytest tests/integrations/test_fastapi.py: 46 passed, 1 skipped (a Linux-only test for long sync calls).
  • uv run task test-package: 702 passed.
  • CPU-time metering against known burns (±10%): async and sync endpoints, sync dependencies, sync streaming, routes added with include_router, plain Starlette endpoints; dependency_overrides still work; an asyncio.sleep endpoint gets almost no CPU.
  • Conservation of the buckets over 500 random windows.
  • Overhead: about 1.7 µs added per request at p50 (100k in-process calls on macOS), against a 50 µs budget.
  • In CPU load mode on macOS, endpoints burning 1:2:4 ms of CPU are charged 1 : 1.96 : 3.88, and a sleep endpoint 1.2% of the 4 ms burn.

Not yet run: the accuracy checks against RAPL (pytest -m accuracy tests/integrations/fastapi_accuracy.py) need a Linux machine with readable RAPL; they skip elsewhere. In CPU load mode the check "per-request energy barely moves when another process loads the machine" fails (+289%) because the cubic power model has no single per-CPU-second cost; the docs say so and the check only runs against a measured source.

Known limits

  • CPU time is a proxy for energy; frequency scaling, SMT and wide vector instructions add error that has not been measured.
  • CPU work in child tasks (create_task, gather) is not metered and lands in process_unattributed_kwh.
  • If another library later restores its own run_sync over ours, threadpool metering stops until restart.
  • Nested CodeCarbonMiddleware instances share one meter; the outer one loses threadpool CPU.

Screenshots (if appropriate):

N/A

Types of changes

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to change)

AI Usage Disclosure

Please refer to docs/how-to/ai-policy.md for detailed guidelines on how to disclose AI usage in your PR. Accurately completing this section is mandatory.

  • 🟥 AI-vibecoded: You cannot explain the logic. Car analogy : the car drive by itself, you are outside it and just tell it where to go.
  • 🟠 AI-generated: Car analogy : the car drive by itself, you are inside and give instructions.
  • ⭐ AI-assisted. Car analogy : you drive the car, AI help you find your way.
  • ♻️ No AI used. Car analogy : you drive the car.

Checklist:

  • My code follows the code style of this project.
  • My change requires a change to the documentation.
  • I have updated the documentation accordingly.
  • I have read the docs/how-to/contributing.md document.
  • I have added tests to cover my changes.
  • All new and existing tests passed.

Commits

  • ae74a9a feat(tracker): pass a per-component WindowSample to energy window observers
  • 3d58a7a feat(fastapi): charge requests only energy above idle
  • 0b17fdb feat(fastapi): keep only this process's share of CPU dynamic energy
  • 26afaa4 feat(fastapi): split CPU energy by each request's metered CPU time
  • 6b75abb feat(fastapi): report GPU energy, attribution method and quality per request
  • c052bcc test(fastapi): add accuracy checks and an overhead benchmark
  • c89a986 fix(fastapi): drop CPU time metered in a skipped window
  • 01bd37f feat(fastapi): meter CPU time of work sent to the threadpool
  • 5ce9a8f feat(fastapi): price each request's CPU time at a per-CPU-second cost
  • 72043fb docs(fastapi): say load-mode charges rise under other processes' load
  • e0506e7 fix(fastapi): cap requests and unclaimed process CPU together
  • f646fea fix(fastapi): meter any awaitable the ASGI app returns
  • 678bb56 fix(tracker): never let a window sample error break measurement
  • d1dc599 fix(fastapi): read a worker call's CPU time under the meter lock
  • 57491e4 fix(fastapi): drop CPU time metered in a zero-width window
  • 57fee8f fix(fastapi): ignore late windows from a detached tracker
  • 989447b fix(fastapi): drop the fitted CPU cost when a window has no load reading
  • 16d4aed docs(fastapi): document idle per mode, the shared cap and known limits

…ervers

Observers now receive cumulative CPU, GPU and RAM energy, the source
quality of each component, and the CPU idle power when the power model
fixes it (load and constant modes).
Each window is split per component into idle and dynamic energy. Idle
power is the lower of a rolling minimum of window power and a good
regression intercept against CPU utilisation, or the analytic value in
load and constant modes. Idle and RAM energy go to a new idle_kwh bucket.
The share is the process's CPU time over the machine's busy CPU time
from psutil.cpu_times(), clamped to [0, 1]. The rest goes to a new
other_processes_kwh bucket. Load mode in process tracking is already
per-process and keeps all of it.
The middleware drives the app coroutine through an awaitable that sums
time.thread_time_ns() around every resumption. Each window, a request
gets our CPU dynamic energy in proportion to the CPU time its meter saw;
CPU time no meter claimed goes to process_unattributed_kwh. RequestEnergy
gains cpu_seconds.
…request

RequestEnergy.energy_kwh is now the CPU energy above idle charged by CPU
time; GPU energy above idle is split by wall-clock overlap into a
separate gpu_kwh. Each request reports its attribution_method and the
weakest source quality it was charged from. Document the buckets and
the limits of the method.
CI checks conservation over random fake windows, cpu_seconds of a known
async CPU burn, and a sleeping endpoint's near-zero CPU time. A sync
burn is marked xfail: worker-thread CPU time is not metered yet.
fastapi_accuracy.py holds the opt-in RAPL checks (-m accuracy) and skips
without readable RAPL; fastapi_overhead.py times 100k ASGI calls with
and without the middleware.
A window skipped for a backwards counter left the meters' CPU time to be
charged in the next window, against a process time that did not include
it.
Wrap anyio.to_thread.run_sync process-wide while a tracker is attached.
The wrapper only acts when a request meter is set in the current context,
so sync endpoints and dependencies, sync streaming iterators and plain
Starlette sync endpoints are metered, including routes from
include_router. The patch is reference-counted, idempotent and restored
on detach; meter_threadpool=False turns it off. On Linux a running
worker call is read through its thread CPU clock, so long sync calls are
charged window by window.
A request is charged cpu_seconds times the cost of a CPU-second: the
power model's in load mode, the slope of a good power-on-load fit when
measured, else the process average as before. Charges are capped at the
process's share of dynamic CPU energy; what the cost leaves over goes to
other_processes_kwh. The conservation invariant still holds exactly.
This stops per-request energy from following other processes' load on a
convex power curve.
A binding cap used to hand all of our share to the requests, leaving
nothing for process CPU time no meter claimed. One cap now scales both.
ASGI only promises an awaitable, so drive its __await__ iterator
instead of calling send/throw/close on the object itself.
Build the observer sample inside the same guard as the observers, so a
bug in an integration's sample can't stop the core measure and flush.
Otherwise a concurrent total_ns() could see more of the running call
than is then added, and the meter would go backwards.
A window with no width now advances the meters like a skipped one, and
counts as skipped when energy arrived in it.
Each observer is bound to its tracker, and a sample from a tracker that
is no longer attached is dropped instead of mixing into the new anchor.
A window without machine CPU times kept charging by the previous fit.
State idle for each CPU mode, the lower-of idle estimate, where unclaimed
process CPU goes, what unattributed_kwh holds, the anchor sample race,
nested middleware, the run_sync chain drop and child processes in
process-tracking load mode.
@codecov

codecov Bot commented Sep 28, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 94.16910% with 20 lines in your changes missing coverage. Please review.
✅ Project coverage is 91.92%. Comparing base (4732135) to head (16d4aed).

Files with missing lines Patch % Lines
codecarbon/emissions_tracker.py 78.94% 12 Missing ⚠️
codecarbon/integrations/fastapi/attribution.py 96.55% 8 Missing ⚠️
Additional details and impacted files
@@                       Coverage Diff                       @@
##           feat/add-fastapi-middleware    #1430      +/-   ##
===============================================================
+ Coverage                        91.74%   91.92%   +0.18%     
===============================================================
  Files                               52       52              
  Lines                             5352     5646     +294     
===============================================================
+ Hits                              4910     5190     +280     
- Misses                             442      456      +14     

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.
  • 📦 JS Bundle Analysis: Save yourself from yourself by tracking and limiting bundle sizes in JS merges.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant