Guide · Claude Enterprise

How to Set Up Inference Hooks: Connect the Endpoint, Return a Verdict, Control the Rollout

Last updated: AI-generated

Claude shows the signing secret exactly once, on the first save; after that it can only be rotated, never retrieved. Your server gets 5,000 milliseconds to return a verdict, adjustable between 1 and 10,000.

In short: Setting up Inference hooks is two jobs. An Owner turns the feature on in claude.ai under Organization settings › Data and privacy, saves an HTTPS address, and sets the timeout, failure handling, and rollout. In parallel, somebody builds the AI security server that answers every signed POST with {"action": "allow"} or {"action": "deny"}. It has to be HTTP 200, or the response counts as a failure rather than a denial.

What you need first

RequirementWhere it stands
PlanClaude Enterprise. Platform organizations, meaning API access through the Claude Platform, are out of scope. The feature does not exist on Amazon Bedrock or Google Cloud.
RoleThe organization:manage permission, held only by Owner and Primary owner. The Admin role does not have it.
EndpointAn https:// URL on port 443, on a publicly routable host, reachable without redirects, with a certificate that validates against the public CA trust store.
MaturityBeta. Field names, request shapes, and headers may still change.

Private, loopback, and carrier-grade NAT addresses are refused at connect time. Reverse tunnels such as ngrok are blocked by Anthropic’s network policy, so don’t test through a tunnel. Host the server on a domain you control.

The three states

The settings read more easily as three states than as a pile of switches:

StateWhat is setWhat happens
OffEnforce verdicts offYour server is never contacted, nothing is inspected
ShadowEnforce verdicts on, Mode set to Shadow modeYour server receives prompts and answers, nothing is blocked
EnforcingEnforce verdicts on, Mode set to Allow the request or Block the requestA deny blocks the request

Setting it up, step by step

1. Allow the feature

claude.ai › Organization settings › Data and privacy › the Inference hooks section › turn on Allow for your organization.

This only unlocks the settings page. It always forces Enforce verdicts off, even for a configuration that was enforcing before, so allowing the feature never starts inspection by itself.

2. Open the settings page

The page sits under Data and privacy and has no entry of its own in the settings nav; its breadcrumb reads Data and privacy / Inference hooks. Until you save an endpoint, the page warns that nothing is being inspected, and Enforce verdicts carries a Requires endpoint badge.

3. Save the endpoint

Configure opens the Set up endpoint dialog. It asks for the Endpoint URL and accepts only https://. Nothing else at this point. Next saves it, and the button then reads Edit.

The whole URL you enter is the endpoint. There is no fixed path suffix, so pick any path that suits your server.

4. Lock away the signing secret

The first save generates the webhook signing secret and reveals it once. Copy it and store it securely before you click Next. It cannot be retrieved later, only rotated.

5. Add headers and test the connection

Next on the secret dialog reopens the endpoint dialog with two more controls.

Custom request headers: up to 16 static headers sent with every request so your server can authenticate the caller. Values are stored encrypted and never shown again; only the names stay visible. Because they are write-only, every change means re-entering all of them. Changing the URL clears every stored value so your credentials never reach a new destination, so re-enter them after a URL change.

Header names use standard HTTP token characters with - rather than _. Reserved and therefore rejected: framing headers such as Content-* and Host, proxy and cookie headers, client-address headers such as X-Forwarded-*, the webhook-* signature headers, and anything prefixed X-Anthropic-. Values must be printable ASCII.

Test connection sends a synthetic test prompt to the URL and headers currently in the form, not to the saved ones, so re-enter stored header values before testing. The result says whether your server answered allow or deny, which catches a deny-everything default before you start enforcing.

Failure resultWhat to check
URL rejectedThe URL failed a structural check. Use https:// on port 443.
Private or internal IPThe host resolves to a private or internal address.
TimeoutNo verdict within the timeout.
Transport errorDNS, the TLS handshake, or the connection failed.
Non-200 statusAnything other than HTTP 200. Redirects are not followed and count as failures.
Unparseable responseA response arrived, but it is not a valid verdict.
Signing secret requiredNo secret exists, so the test would go out unsigned. Click Generate secret under Request signing.

6. Failure handling and timeout

Under Failure handling, Mode decides what applies while your server is unreachable or too slow:

  • Block the request: inference stops when no verdict arrives (fail closed).
  • Allow the request: the request reaches the model uninspected (fail open).

The third option in the same dropdown, Shadow mode, is a rollout tool, not a failure policy.

Prompt verdict timeout (ms) ranges from 1 to 10,000, defaulting to 5,000. The budget covers the whole exchange: connection, TLS handshake, request, and response. A slow verdict counts as an unreachable server, so use the lowest value your server reliably meets.

Changes in this section save as you type them. On first save the defaults are Allow the request and 5,000ms.

7. Rollout percentage

Under Rollout, Requests inspected (%) runs from 0 to 100 and decides what share gets inspected. The roll happens once per conversation turn, so a single conversation can be inspected on some turns and not others. Anything outside the sample proceeds uninspected, even with failure handling set to Block the request.

8. Start enforcing

Turn on Enforce verdicts and confirm in the dialog, which restates your failure handling choice. Allow about a minute for the change to reach every Anthropic server; requests already in flight finish under the old setting. Turning it off works just as quickly and keeps your configuration.

To watch before you block: set Mode to Shadow mode in step 6 first, then turn on Enforce verdicts. Your server receives real prompts and answers exactly as it would when enforcing, and nothing is blocked, not a deny and not an outage, with the end user seeing nothing. The settings page then shows a Shadow mode — not blocking badge. To leave, set Mode back to Allow the request or Block the request.

The server: the smallest thing that works

The smallest useful server reads the request and lets it through. This example is the one in the docs:

# Run with: python server.py
from http.server import BaseHTTPRequestHandler, ThreadingHTTPServer


class VerdictHandler(BaseHTTPRequestHandler):
    protocol_version = "HTTP/1.1"  # keep the connection open between verdicts

    def do_POST(self):
        # Drain the body; transcripts can be megabytes.
        self.rfile.read(int(self.headers.get("Content-Length", 0)))
        verdict = b'{"action": "allow"}'
        self.send_response(200)
        self.send_header("Content-Type", "application/json")
        self.send_header("Content-Length", str(len(verdict)))
        self.end_headers()
        self.wfile.write(verdict)


ThreadingHTTPServer(("", 8000), VerdictHandler).serve_forever()

It accepts everything, unsigned requests included. Signature verification goes in before you start enforcing.

What arrives

Every request carries these fixed headers, plus your own and the webhook-* signature headers:

HeaderValue
Content-Typeapplication/json
User-Agentanthropic-dlp/1
Accept-Encodingidentity

The body is a JSON object. There is one event today, the prompt frame, sent once per governed inference request and before inference begins:

{
  "type": "prompt",
  "request_id": "req_abc123",
  "tenant_id": "11111111-1111-1111-1111-111111111111",
  "actor": {
    "type": "user",
    "id": "user_01AbCdEfGhIjKlMnOpQrStUv",
    "email_address": "alice@example.com"
  },
  "source": { "application": "claude-ai" },
  "session_id": "22222222-2222-2222-2222-222222222222",
  "model": "claude-sonnet-4-5",
  "messages": [],
  "metadata": {}
}

request_id is the same value as the webhook-id header. source.application is an open string, not a closed enum: common values are claude-ai, claude-code, and cowork, while connection tests and the circuit breaker’s recovery checks arrive as config-test. Treat the field as routing metadata, not as a trust boundary.

messages holds the transcript as the user sees it: text, tool calls and their results, extracted attachment text, earlier turns. There are four block types, text, tool_use, tool_result, and attachment. What never appears: system prompts, tool definitions, Anthropic-internal context, Claude’s hidden reasoning, and raw file bytes.

Two things trip people up in practice:

  • A turn whose blocks are all excluded is dropped entirely. Don’t assume user and assistant alternate cleanly.
  • Transcripts arrive untruncated. In practice the context window keeps bodies under about 10 MB; the protocol allows 64 MiB. Common defaults sit far below that: nginx client_max_body_size is 1 MB, express.json() is 100 kB. A rejected body is a webhook failure, and with failure handling on Allow the request, the oversized prompt then reaches the model uninspected.

What goes back

HTTP 200 and a JSON verdict, for both outcomes. To allow:

{ "action": "allow" }

To deny:

{
  "action": "deny",
  "deny_reason": "This prompt appears to contain customer payment card data, which your organization's policy does not allow.",
  "reference_id": "scan_01HXPT4R9V"
}
FieldConstraintsMeaning
actionallow or deny, requiredAny other value is a webhook failure, not a denial.
deny_reasonString or null, at most 500 charactersWhat the user sees. Longer values are truncated. Ignored on allow.
reference_idString or null, at most 50 characters from [A-Za-z0-9._:/-]Your own identifier for this evaluation. It lands on the inference_hooks_request_denied activity and is never shown to the user. No content, no personal data.

A deny is never lost over a formatting problem: an oversize deny_reason is truncated, a malformed reference_id is silently dropped, and the denial still stands. The reverse does not hold. Never signal a denial with an error status. Anything other than HTTP 200 with a parseable verdict is a webhook failure, and then your failure handling applies instead of your judgment.

Anthropic reads at most 64 KiB of your response, uncompressed. Redirects are not followed, cookies are ignored, and unknown fields in the verdict are skipped.

Verifying the signature

Requests are signed per the Standard Webhooks specification, using three headers. Anthropic sends the names in lowercase and proxies are free to re-case them, so look them up case-insensitively.

HeaderContents
webhook-idUnique identifier for this delivery, identical to request_id in the body. Works as an idempotency key and is the first component of the signed payload.
webhook-timestampUnix time in seconds as a decimal string. Reject anything more than five minutes from your clock, in either direction.
webhook-signatureOne or more space-separated v1,<base64> values, each an HMAC-SHA256 over {webhook-id}.{webhook-timestamp}.{raw body bytes}. Accept if any one matches, using a constant-time comparison.

Two mistakes account for most verification bugs:

  1. Compute over the raw bytes, the body exactly as it arrived, before any parsing or re-encoding.
  2. Decode the secret with a standard base64 decoder. The key is the part after the whsec_ prefix, in the standard alphabet with + and /. A URL-safe decoder derives the wrong key bytes whenever the secret contains + or /, which is most of the time.

The check in Python, as the docs give it:

import base64, hashlib, hmac, time

TOLERANCE_SECONDS = 300


def verify(secret: str, headers: dict[str, str], body: bytes) -> bool:
    lowercased = {name.lower(): value for name, value in headers.items()}
    try:
        message_id = lowercased["webhook-id"]
        timestamp = lowercased["webhook-timestamp"]
        signatures = lowercased["webhook-signature"]
    except KeyError:
        return False  # unsigned: not from Anthropic

    try:
        signed_at = int(timestamp)
    except ValueError:
        return False
    if abs(time.time() - signed_at) > TOLERANCE_SECONDS:
        return False  # replayed, or the clocks disagree

    try:
        key = base64.b64decode(secret.removeprefix("whsec_"), validate=True)
    except ValueError:
        return False

    payload = f"{message_id}.{timestamp}.".encode() + body
    expected = b"v1," + base64.b64encode(
        hmac.new(key, payload, hashlib.sha256).digest()
    )
    return any(
        hmac.compare_digest(expected, candidate.encode())
        for candidate in signatures.split()
    )

Once your organization has a secret, every request is signed, the connection test included, because the setup flow generates the secret before the first test. So reject unsigned requests. One exception: organizations that turned Inference hooks on before the secret was required keep sending unsigned requests until an administrator generates one.

Rotation is an immediate cutover with no overlap. Requests signed with the old secret still arrive for about a minute afterwards, plus whatever was already in flight, so let your server accept both signatures during the switchover.

Running it

  • Retry: exactly once, after 100ms, and only when the connection attempt itself fails. The retry shares the same timeout budget and carries the same webhook-id and signature. Once your server has answered, the exchange is never retried.
  • Deduplication: on webhook-id. If you record verdicts, key the records on it.
  • Circuit breaker: sustained failures attributable to your server trip it. Anthropic then stops contacting your server, and your failure handling applies to every request. Under Block the request, that means your people are blocked until it resets. Starting 10 minutes after the trip, Anthropic checks at most about once a minute with the same synthetic test request; a valid verdict, allow or deny, resets it. Change any setting after a trip, even just rotating the secret, and those checks stop. Then you turn Enforce verdicts back on by hand.
  • Source IPs: requests come from 160.79.106.0/24, part of Anthropic’s published outbound ranges. The inbound ranges on the same page don’t cover it. An allowlist narrows your exposure but does not replace signature verification, since other Anthropic egress traffic uses the same block.
  • Latency: your server’s round trip sits on every governed request in the organization. Load-test it before rolling out to a large one.

Staying forward compatible

The protocol grows. Your server has to skip unknown top-level fields, unknown keys in metadata, new source.application values, new actor.type values, and blocks with an unrecognized type, rather than rejecting the request.

One special case: more event types are announced. If the top-level type is one you don’t recognize, skipping won’t do, because the request still needs a verdict. Return allow, not an error status. An error is a webhook failure, and sustained webhook failures trip the circuit breaker.

What the user sees on a deny

The message has two parts: your server’s deny_reason, then a blank line, then a standing text administrators set under Custom blocked prompt message, up to 500 characters, usually who to contact or where to request an exception. With no custom text, a built-in default points at the administrators; the appended part can also be switched off entirely, leaving only the deny_reason.

So write the deny_reason for the user rather than for your own team: what to change, which kind of content has to go, and not a scanner code only you can read.

Every denial is also recorded in the organization’s Activity Feed. A blocked message stays in the conversation on claude.ai, and the Inference hooks system keeps no copy of prompts or responses of its own, only configuration and metadata such as verdicts, timestamps, and request identifiers.

Who can be exempted

Under Exclusions you pick roles whose members are never covered: their prompts never reach your server. Only custom roles your organization created are offered, not the built-in ones, and the list starts empty. The exemption covers a person’s interactive sessions; traffic authenticated by machine credentials is always inspected. Changes to the list land in the audit trail, and making them requires identity management permission.

Keeping an eye on it

The endpoint health area of the settings page shows status (Healthy, Tripped, Not enforcing, Not configured), failures per minute averaged over the last two minutes, the block rate while the rollout percentage is below 100, and when the circuit breaker last tripped. Recent errors appear as a timestamp, an error type, and a one-line reason, with no content and no endpoint URL.

The docs name two limits themselves. The panel is best-effort and shows zero failures when it cannot read the counters, so a healthy-looking panel is no proof of a healthy server. And failures per minute counts network and DNS errors that never trip the breaker, so the number can run high while Circuit breaker tripped stays empty.

For alerting, the Activity Feed is the better source: each trip is recorded as inference_hooks_circuit_breaker_tripped, one activity per trip rather than one per affected request. That requires the Compliance API to be enabled for the organization. While the breaker is tripped, no per-request activities are written, so the trip activity is the only record of that window.

Turning it off again

Two levels:

  1. Enforce verdicts off on the settings page. Within about a minute, prompts stop going to your server, and requests already in flight finish under the old setting. The page stays available, which makes this the pause to use while you work on the server.
  2. Allow for your organization off in Data and privacy settings. The settings page goes away too. Either way, your endpoint, headers, and secret are kept; turning it back on forces Enforce verdicts off and clears a tripped circuit breaker.

Limits

  • Allow or deny only. Rewriting or redacting a prompt is not supported: it goes through whole or not at all.
  • One event, prompt. It fires once per governed inference request, before inference. Response-side enforcement is announced as a later event and does not exist yet.
  • Attachments arrive as metadata and extracted text. Raw file and image bytes are never sent, so image-only content, a screenshot of a document for instance, is not inspected.
  • Not covered: voice mode, ancillary requests such as conversation title generation, system prompts, and tool definitions.
  • Not available: to Platform organizations, on Amazon Bedrock, or on Google Cloud.
  • Beta. Field names, request shapes, and headers may change. This guide gets checked against the docs on the next run.

If you would rather look afterwards than intervene beforehand, that is what the Compliance API is for. Inference hooks acts inline and decides in real time, with Anthropic calling your server. The Compliance API works after the fact, and you call Anthropic.

Changes

  • 2026-09-22: First version. Every statement checked against the docs as of 2026-09-22.

Sources