Skip to main content
Working with an agent? Give them a link to this page as markdown.

Gateway deployment reference

Gateway hooks connect an LLM proxy to the Watcher API without a Watcher client on each developer machine. The proxy routes model traffic; Watcher reviews the conversation; the coding agent executes its own tools.

See the self-hosted deployment reference for Watcher's components, authentication paths, and internal network connections.

IntegrationSupported trafficEffect on client toolsSetup
Tailscale ApertureCaptures from the supported coding agents.Checks the conversation after the provider responds; cannot stop tools.Configure Aperture.
LiteLLM, or another proxy with a synchronous callbackAnthropic Messages, OpenAI Responses and OpenAI Chat Completions, including streaming.In enforce mode, holds each response until Watcher has decided on its tool calls.A callback you install in the proxy that calls the raw pre-tool-use hook.

Tailscale Aperture: review without blocking​

Aperture passes model requests and responses between the agent and the provider. It sends Watcher a copy over HTTPS with Watcher credentials. Watcher checks the copy in the background and cannot stop tools in the agent.

  1. The agent sends a model request through Aperture. Aperture forwards the provider's response to the agent without waiting for Watcher review.
  2. Aperture sends an entire_request event containing the request, response, and metadata to POST /api/v1/hooks/aperture. Configure both request_body and response_body send types.
  3. Watcher returns HTTP 202 Accepted. In the background, it checks new tool calls and saves the conversation and results. Text-only turns are saved too.

Scope the Aperture grant to the intended users or device tags and provider keys. Developer attribution comes from metadata.login_name. Check the capture requirements for your client and provider configuration.

LiteLLM: decisions before client tool delivery​

This path needs a callback installed in the proxy; installing the Watcher SDK does not install one. The LiteLLM example README documents setup, configuration and limits for Apollo's experimental callback. This example is not an officially supported integration.

LiteLLM passes requests to the model provider and buffers each response. A Watcher callback sends the raw request and provider response to Watcher over HTTPS with Watcher credentials. In enforce mode, LiteLLM releases the response only when every tool call is allowed; otherwise it withholds every tool call. The agent runs the tools it receives.

  1. The agent sends an HTTP request to LiteLLM using a configured model alias. LiteLLM calls the model provider and buffers the complete response, including a streamed one.
  2. A callback inside LiteLLM sends the client's request body and the complete provider response to POST /api/v1/hooks/pre-tool-use with format: "raw". Watcher parses the provider format, applies your policy and returns a decision for each new client tool call.
  3. In enforce mode, the callback releases the response only when every client tool call in it is allowed. If any call is denied, escalated, or has no decision, the callback withholds every client tool call in that response and returns an explanation instead. A streamed response is re-encoded from the body Watcher evaluated, so the agent can only run calls Watcher saw.

The client receives nothing until Watcher decides. Escalations follow the denial path; this integration has no human-approval interface. The coding agent's own permissions still apply after release. If Watcher is unreachable or returns an error, the callback should withhold the whole response rather than release it. A grading failure is different: Watcher returns a normal decision chosen by the organization's policy_monitors.on_grading_failure setting, so with allow the callback releases tools that were never fully graded. Set it to deny if those must be withheld; see delivery limits. The raw request section lists the checks a callback must make before it releases a response.

For enforcement, allow developer inference traffic only through the checked /v1/messages, /v1/responses, and /v1/chat/completions paths. Restrict direct provider access, and apply a separate access policy to administrative routes. Realtime, WebSockets, batches, files, and LiteLLM's provider pass-through routes bypass the callback and are not checked. Tools that run on the provider's servers, such as hosted web search, have already run by the time the response reaches the callback.

Apollo's experimental LiteLLM callback also rejects requests it cannot evaluate before the provider runs, with HTTP 400: more than one choice (n other than 1), Chat Completions audio output, and Responses requests that rely on history stored by the provider (previous_response_id, conversation, or background). Clients must send the full conversation in each request.

Credentials and the Watcher connection​

ConnectionCredentials
Developer → proxyAperture's access grants or LiteLLM's client credentials.
Proxy → model providerThe provider authentication configured in Aperture or LiteLLM.
Proxy → WatcherSeparate Watcher credentials, sent to the deployment's HTTPS endpoint on TCP 443. In SSO mode, use an organization API key as x-api-key. In proxy mode, authenticate through Watcher's reverse proxy instead.

Aperture requires sessions:write:any to attribute captures to developers. A LiteLLM callback can send the proxy key's user_email as the developer, which needs the same permission. Otherwise, recording uses the Watcher credential's write scope and can produce unattributed sessions. See hook authentication.

Watcher uses its own configured model providers for grading. Its deployment, storage, and retention are described in the self-hosted deployment reference and security reference. The proxy's logging and retention are configured separately.

Modes and verification​

The administrator's enforcement mode, supplied through managed client settings, controls hook behavior even when no Watcher client is installed.

ModeApertureLiteLLM
enforceChecks tools but cannot stop them.Holds each response and releases its tool calls only when all are allowed.
observeChecks tools but cannot stop them.Waits for checks, then releases the response unchanged. A response Watcher cannot parse, or a Watcher outage, can still prevent delivery.
pausedTranscript ingestion without hook judgments or grades.Releases responses and records the transcript without hook judgments or grades. A Watcher outage can still prevent delivery.

Verify a scoped coding task in the Analyzer: find its session, check attribution, and inspect decisions and available grades. Follow the Aperture verification steps.

Prefer one monitoring integration per conversation. Combining a local Watcher client with gateway hooks can repeat grading. Watcher recognizes supported coding agents in raw requests and uses their own session IDs. For other applications, send a stable conversation ID in external_ids; see raw request fields.