Skip to content
Judgment Labs
Esc
↑↓navigate↵open⌘Jpreview
On this page

Claude Managed Agents Tracing

Trace Claude Managed Agents sessions, model requests, tool calls, and sub-agent delegation with Judgment.

Claude Managed Agents integration captures traces from agents that run on Anthropic’s managed infrastructure. Judgment builds spans from the session’s event stream, so you see each turn’s model requests, token usage and cost, tool calls, and, for multi-agent sessions, each sub-agent’s run. This integration is designed for TypeScript applications.

Quickstart

Install Dependencies

@anthropic-ai/sdk version 0.129.0 or later is required.

npm install judgeval @anthropic-ai/sdk
yarn add judgeval @anthropic-ai/sdk
pnpm add judgeval @anthropic-ai/sdk
bun add judgeval @anthropic-ai/sdk

Wrap Your Anthropic Client

Wrap the client you use to stream session events, and initialize the tracer.

import Anthropic from "@anthropic-ai/sdk";
import { Tracer, wrapAnthropicManagedAgents } from "judgeval"; 

const client = wrapAnthropicManagedAgents(new Anthropic()); 

await Tracer.init({ projectName: "managed_agents_project" }); 

Use the wrapped client for everything else. Judgment traces sessions through client.beta.sessions.events.stream.

Run a Session

Run your session as usual. Spans are created as you consume the event stream, so iterate it until the session goes idle.

const agent = await client.beta.agents.create({
  name: "assistant",
  model: "claude-haiku-4-5",
  system: "You are a concise assistant. Use bash when asked to run commands.",
  tools: [{ type: "agent_toolset_20260401" }],
});
const environment = await client.beta.environments.create({
  name: "judgment-example",
  config: { type: "cloud", networking: { type: "limited" } },
});
const session = await client.beta.sessions.create({
  agent: agent.id,
  environment_id: environment.id,
});

const stream = await client.beta.sessions.events.stream(session.id);
await client.beta.sessions.events.send(session.id, {
  events: [
    {
      type: "user.message",
      content: [{ type: "text", text: "Run `echo hello` in bash and tell me the output." }],
    },
  ],
});

for await (const event of stream) {
  if (event.type === "agent.message") {
    for (const block of event.content) {
      if (block.type === "text") console.log(block.text);
    }
  }
  if (event.type === "session.status_idle" && event.stop_reason.type !== "requires_action") {
    break;
  }
}

await Tracer.shutdown(); 

What gets traced

Each turn of a session produces these spans:

Span Created for Includes
invoke_agent The turn The user’s messages and the agent’s final reply
generate_content Each model request Model, input and output messages, token usage, and cost
execute_tool <name> Each tool call The tool’s input and result; errors are marked on the span

A turn starts with the first message in the session and ends when the session goes idle. If it goes idle waiting for you to return a custom tool result (requires_action), the turn continues after you send the result. The turn’s invoke_agent span is the root of a new trace, or a child of your application’s active span when there is one.

Every span carries the Managed Agents session ID as its Judgment session ID, so the turns of one conversation are grouped together.

Multi-agent sessions

When a coordinator agent delegates to another agent, the delegation appears as an execute_tool transfer_to_agent span. Its input is the target agent and message, and its output is the sub-agent’s reply. The sub-agent’s own run (invoke_agent, model requests, and tool calls) is nested beneath that span, so the whole delegation tree is in one trace.

Verify

Run a session, then open your project in Judgment and check that the trace shows invoke_agent with a generate_content span for each model request and an execute_tool span for each tool call. See verifying a trace for the full checklist.

Next Steps

Was this page helpful?