Gateway deployment reference
Gateway hooks connect an LLM proxy to the Watcher API without a Watcher client on each developer machine. The proxy routes model traffic; Watcher reviews the conversation; the coding agent executes its own tools.
See the self-hosted deployment reference for Watcher's components, authentication paths, and internal network connections.
| Integration | Supported traffic | Effect on client tools | Setup |
|---|---|---|---|
| Tailscale Aperture | Captures from the supported coding agents. | Checks the conversation after the provider responds; cannot stop tools. | Configure Aperture. |
| LiteLLM, or another proxy with a synchronous callback | Anthropic Messages, OpenAI Responses and OpenAI Chat Completions, including streaming. | In enforce mode, holds each response until Watcher has decided on its tool calls. | A callback you install in the proxy that calls the raw pre-tool-use hook. |
Tailscale Aperture: review without blocking
- The agent sends a model request through Aperture. Aperture forwards the provider's response to the agent without waiting for Watcher review.
- Aperture sends an
entire_requestevent containing the request, response, and metadata toPOST /api/v1/hooks/aperture. Configure bothrequest_bodyandresponse_bodysend types. - Watcher returns HTTP
202 Accepted. In the background, it checks new tool calls and saves the conversation and results. Text-only turns are saved too.
Scope the Aperture grant to the intended users or device tags and provider
keys. Developer attribution comes from metadata.login_name. Check the
capture requirements for your client
and provider configuration.
LiteLLM: decisions before client tool delivery
This path needs a callback installed in the proxy; installing the Watcher SDK does not install one. The LiteLLM example README documents setup, configuration and limits for Apollo's experimental callback. This example is not an officially supported integration.
- The agent sends an HTTP request to LiteLLM using a configured model alias. LiteLLM calls the model provider and buffers the complete response, including a streamed one.
- A callback inside LiteLLM sends the client's request body and the complete
provider response to
POST /api/v1/hooks/pre-tool-usewithformat: "raw". Watcher parses the provider format, applies your policy and returns a decision for each new client tool call. - In
enforcemode, the callback releases the response only when every client tool call in it is allowed. If any call is denied, escalated, or has no decision, the callback withholds every client tool call in that response and returns an explanation instead. A streamed response is re-encoded from the body Watcher evaluated, so the agent can only run calls Watcher saw.
The client receives nothing until Watcher decides. Escalations follow the
denial path; this integration has no human-approval interface. The coding
agent's own permissions still apply after release. If Watcher is unreachable
or returns an error, the callback should withhold the whole response rather
than release it. A grading failure is different: Watcher returns a normal
decision chosen by the organization's policy_monitors.on_grading_failure
setting, so with allow the callback releases tools that were never fully
graded. Set it to deny if those must be withheld; see
delivery limits. The
raw request section lists
the checks a callback must make before it releases a response.
For enforcement, allow developer inference traffic only through the checked
/v1/messages, /v1/responses, and /v1/chat/completions paths. Restrict
direct provider access, and apply a separate access policy to administrative
routes. Realtime, WebSockets, batches, files, and LiteLLM's provider
pass-through routes bypass the callback and are not checked. Tools that run
on the provider's servers, such as hosted web search, have already run by the
time the response reaches the callback.
Apollo's experimental LiteLLM callback also rejects requests it cannot evaluate
before the provider runs, with HTTP 400: more than one choice (n other than
1), Chat Completions audio output, and Responses requests that rely on history
stored by the provider (previous_response_id, conversation, or
background). Clients must send the full conversation in each request.
Credentials and the Watcher connection
| Connection | Credentials |
|---|---|
| Developer → proxy | Aperture's access grants or LiteLLM's client credentials. |
| Proxy → model provider | The provider authentication configured in Aperture or LiteLLM. |
| Proxy → Watcher | Separate Watcher credentials, sent to the deployment's HTTPS endpoint on TCP 443. In SSO mode, use an organization API key as x-api-key. In proxy mode, authenticate through Watcher's reverse proxy instead. |
Aperture requires sessions:write:any to attribute captures to developers.
A LiteLLM callback can send the proxy key's user_email as the developer, which
needs the same permission. Otherwise, recording uses the Watcher
credential's write scope and can produce unattributed
sessions. See
hook authentication.
Watcher uses its own configured model providers for grading. Its deployment, storage, and retention are described in the self-hosted deployment reference and security reference. The proxy's logging and retention are configured separately.
Modes and verification
The administrator's enforcement mode, supplied through managed client settings, controls hook behavior even when no Watcher client is installed.
| Mode | Aperture | LiteLLM |
|---|---|---|
enforce | Checks tools but cannot stop them. | Holds each response and releases its tool calls only when all are allowed. |
observe | Checks tools but cannot stop them. | Waits for checks, then releases the response unchanged. A response Watcher cannot parse, or a Watcher outage, can still prevent delivery. |
paused | Transcript ingestion without hook judgments or grades. | Releases responses and records the transcript without hook judgments or grades. A Watcher outage can still prevent delivery. |
Verify a scoped coding task in the Analyzer: find its session, check attribution, and inspect decisions and available grades. Follow the Aperture verification steps.
Prefer one monitoring integration per conversation. Combining a local
Watcher client with gateway hooks can repeat grading. Watcher recognizes
supported coding agents in raw requests and uses their own session IDs. For
other applications, send a stable conversation ID in external_ids; see
raw request fields.