You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Follow-up from the weekly dependency update in #805. Auditing the release notes and wheel diffs for every bumped provider SDK turned up new behavior our integrations could instrument. The fixes needed for CI (the ADK 2.10 tool-span regression, the dspy 3.4 VCR engine pin, and the litellm 1.103 Windows tiktoken fetch) are handled in #805. This issue tracks optional feature work only.
High priority
dspy ≥ 3.4: token metrics missing on the default engine
dspy.LM now defaults to engine="auto", which sends OpenAI (and other supported) routes through the vendored lm15 transport instead of LiteLLM. On that path our LiteLLM patching never runs, so dspy.lm spans have no token metrics.
Fix: in on_lm_end (py/src/braintrust/integrations/dspy/tracing.py), read usage from instance.history[-1]["usage"] and log prompt, completion, and total tokens. Also update the docstring that says to patch LiteLLM for metrics.
Testing note: lm15 writes to raw sockets, which VCR can't intercept. test_dspy.py pins engine="litellm" as of chore(deps): weekly dependency update #805, so this needs a different testing approach.
pipecat ≥ 1.12: classifier tokens may be counted in LLM span metrics (reasoned from code, not yet reproduced)
The new classifiers (pipecat/classifiers/, used by VoicemailDetector and UIWorker) emit LLMUsageMetricsData in MetricsFrames.
_capture_metrics in py/src/braintrust/integrations/pipecat/tracing.py merges every LLMUsageMetricsData into the current LLM span without checking processor, so classifier tokens can inflate LLM span metrics.
Fix: filter by processor. First add a test that reproduces it.
Nice to have
pipecat ≥ 1.12: unbounded _seen_frame_ids. Our observer keeps every frame id for the whole session, the same leak pipecat fixed in its own observers in 1.12. On 1.12+, gated by version, pass observe_every_push=False to BaseObserver.__init__ or use FramePushed.first_push.
livekit-agents 1.8.3: new STT usage fields.STTMetrics gained total_tokens and input_audio_tokens, but our stt_metrics allowlist in integrations/livekit_agents/tracing.py only keeps audio_duration and streamed.
pydantic-ai 2.50: RequestUsage.audio_seconds. A new duration-based billing field for voice models. It could become a metric next to input_audio_tokens in integrations/pydantic_ai/tracing.py.
openai 3.19.1: tools passed as iterables.chat.completions.parse now accepts generator tools. _filter_metadata in integrations/openai/tracing.py stores the raw value, so span metadata shows a generator repr. Normalize non-list iterables with list(), but only after the SDK has read them, or only when it's safe to.
google-adk 2.10: streamed partial function-call arguments. Check that LLM spans don't log partial chunks once there's a cassette with tool-call streaming. The ADK OpenAI adapter now also reports usage and reasoning tokens; confirm they reach span metrics.
anthropic 1.8.0: beta inline MCP tool definitions. These add new mcp_tool_listing and tool-change content blocks. They should pass through our generic handling; add a cassette test with mcp_toolset to confirm.
openrouter 1.2.32: new client.batch resource (create_batches, list, get_batches, delete). Only worth a thin span on create_batches if we want parity with the anthropic batches wrapper.
agno 3.0.11: add the new cancellation_stage on RunOutput/TeamRunOutput to metadata.
agentscope 2.0.9:TeamPipeline.reply_stream (pipeline/_team_pipeline.py) is untraced.
claude-agent-sdk 0.2.158: consider recording verbatim_prompts as span metadata.
Gaps that predate these bumps
openai: client.beta.responses.* is not patched.
openrouter: client.beta.responses (BetaResponses.send) is not patched. The patcher named openrouter.beta.responses actually targets client.responses.
pydantic-ai: realtime sessions (pydantic_ai.realtime.*) aren't traced. The gap is growing: 2.51 adds OpenAILiveModel and RealtimeSession.wait_for_reply().
strands: Tracer.start_multiagent_span/end_swarm_span and the memory spans are unpatched.
livekit-agents: the new core livekit.agents.inference.realtime.openai.RealtimeModel is untraced.
No action needed
Checked with no gaps found: openai 3.19.2 (apart from the item above), anthropic 1.8.0, google-genai 2.25.0 (the new client.voices is resource CRUD), litellm 1.103.0, strands 1.57.1, claude-agent-sdk 0.2.160, huggingface-hub 2.0.0 (every patched method still exists), cursor-sdk, langchain-core, deepagents, llama-index-core, typesafe-sdk.
Follow-up from the weekly dependency update in #805. Auditing the release notes and wheel diffs for every bumped provider SDK turned up new behavior our integrations could instrument. The fixes needed for CI (the ADK 2.10 tool-span regression, the dspy 3.4 VCR engine pin, and the litellm 1.103 Windows tiktoken fetch) are handled in #805. This issue tracks optional feature work only.
High priority
dspy ≥ 3.4: token metrics missing on the default engine
dspy.LMnow defaults toengine="auto", which sends OpenAI (and other supported) routes through the vendored lm15 transport instead of LiteLLM. On that path our LiteLLM patching never runs, sodspy.lmspans have no token metrics.on_lm_end(py/src/braintrust/integrations/dspy/tracing.py), read usage frominstance.history[-1]["usage"]and log prompt, completion, and total tokens. Also update the docstring that says to patch LiteLLM for metrics.test_dspy.pypinsengine="litellm"as of chore(deps): weekly dependency update #805, so this needs a different testing approach.pipecat ≥ 1.12: classifier tokens may be counted in LLM span metrics (reasoned from code, not yet reproduced)
pipecat/classifiers/, used byVoicemailDetectorandUIWorker) emitLLMUsageMetricsDatainMetricsFrames._capture_metricsinpy/src/braintrust/integrations/pipecat/tracing.pymerges everyLLMUsageMetricsDatainto the current LLM span without checkingprocessor, so classifier tokens can inflate LLM span metrics.Nice to have
_seen_frame_ids. Our observer keeps every frame id for the whole session, the same leak pipecat fixed in its own observers in 1.12. On 1.12+, gated by version, passobserve_every_push=FalsetoBaseObserver.__init__or useFramePushed.first_push.STTMetricsgainedtotal_tokensandinput_audio_tokens, but ourstt_metricsallowlist inintegrations/livekit_agents/tracing.pyonly keepsaudio_durationandstreamed.RequestUsage.audio_seconds. A new duration-based billing field for voice models. It could become a metric next toinput_audio_tokensinintegrations/pydantic_ai/tracing.py.chat.completions.parsenow accepts generatortools._filter_metadatainintegrations/openai/tracing.pystores the raw value, so span metadata shows a generator repr. Normalize non-list iterables withlist(), but only after the SDK has read them, or only when it's safe to.mcp_tool_listingand tool-change content blocks. They should pass through our generic handling; add a cassette test withmcp_toolsetto confirm.client.batchresource (create_batches,list,get_batches,delete). Only worth a thin span oncreate_batchesif we want parity with the anthropic batches wrapper.cancellation_stageonRunOutput/TeamRunOutputto metadata.TeamPipeline.reply_stream(pipeline/_team_pipeline.py) is untraced.verbatim_promptsas span metadata.Gaps that predate these bumps
openai:client.beta.responses.*is not patched.openrouter:client.beta.responses(BetaResponses.send) is not patched. The patcher namedopenrouter.beta.responsesactually targetsclient.responses.pydantic-ai: realtime sessions (pydantic_ai.realtime.*) aren't traced. The gap is growing: 2.51 addsOpenAILiveModelandRealtimeSession.wait_for_reply().strands:Tracer.start_multiagent_span/end_swarm_spanand the memory spans are unpatched.livekit-agents: the new corelivekit.agents.inference.realtime.openai.RealtimeModelis untraced.No action needed
Checked with no gaps found: openai 3.19.2 (apart from the item above), anthropic 1.8.0, google-genai 2.25.0 (the new
client.voicesis resource CRUD), litellm 1.103.0, strands 1.57.1, claude-agent-sdk 0.2.160, huggingface-hub 2.0.0 (every patched method still exists), cursor-sdk, langchain-core, deepagents, llama-index-core, typesafe-sdk.🤖 Generated with Claude Code