Skip to content

feat(guardrails): add fire_and_forget dispatch to generic_guardrail_api - #43681

Open
caduri wants to merge 5 commits into
BerriAI:mainfrom
caduri:feature/generic-guardrail-fire-and-forget
Open

caduri wants to merge 5 commits into
BerriAI:mainfrom
caduri:feature/generic-guardrail-fire-and-forget

Conversation

@caduri

@caduri caduri commented Sep 29, 2026 •

Copy link
Copy Markdown

TLDR

Problem this solves:

  • Every request waits for the generic guardrail endpoint to answer
  • Audit-only setups pay that latency and can still be blocked by it

How it solves it:

  • New fire_and_forget option sends the guardrail call in the background
  • The endpoint gets the same URL, headers and JSON body as before
  • fire_and_forget_max_inflight caps background calls and drops the extra ones
  • Guardrail logs tell dispatched, dropped and unbuildable calls apart
  • A stream is checked once when it closes, even if the client hangs up

Relevant issues

Split out of #37055. Docs: BerriAI/litellm-docs#2047

User Flow

Before: an admin who only wants to audit traffic makes every request wait on the audit endpoint, and its answer can still block

  1. The admin points a generic_guardrail_api guardrail at an audit endpoint that takes 5 seconds to answer
  2. A developer sends POST https://litellm-domain/v1/chat/completions, and the same happens on /v1/responses and /v1/messages
  3. The reply only starts after 5 seconds and comes back 400 "audit says no" when the endpoint answers BLOCKED
  4. When the audit endpoint is down, the request fails with 500 "Generic Guardrail API failed"
  5. https://litellm-domain/ui/?page=logs shows the request as Failure with 5017ms of guardrail overhead

Before: the Generic Guardrail API form ends at streaming_transform_mode

Before: the request is a Failure and the guardrail added 5017ms

After: the same request goes straight to the model while the audit endpoint still receives every call

  1. The admin adds fire_and_forget: true to the guardrail, or sets it in the Add Guardrail form at https://litellm-domain/ui/?page=guardrails, and the boot log warns it is observe-only
  2. A developer sends the same POST https://litellm-domain/v1/chat/completions
  3. The reply comes back 200 in model time, and the endpoint's BLOCKED answer is only logged
  4. When the audit endpoint is down, the request still returns 200 and the proxy logs a warning with the call id
  5. https://litellm-domain/ui/?page=logs shows the guardrail as PASSED, or NOT RUN when the in-flight cap dropped the call

After: the Generic Guardrail API form gains fire_and_forget and fire_and_forget_max_inflight

After: a dispatched call is a Success with the guardrail PASSED and 0ms overhead

After: a call dropped by the in-flight cap is a Success with the guardrail NOT RUN

To reproduce the screenshots, open http://localhost:4000/ui/?page=guardrails, click Add New Guardrail, then Add Provider Guardrail, type any name, pick Generic Guardrail API as the provider, click Next and scroll to the end of the form. For the logs, run case 3 below, open http://localhost:4000/ui/?page=logs and click each of the three requests

Pre-Submission checklist

Please complete all items before asking a LiteLLM maintainer to review your PR

  • I have added meaningful tests
    • The new tests fail at the merge base 6f123b7, where neither option exists
    • Each of these 30 mutations fails a test: observing a stream no handler understands, not making a fire_and_forget stream observe-only, the streaming hook ignoring that flag, not closing the observed stream with the hook, observing only a stream that ends normally, dropping the cancellation shield, letting the check rewrite chunks already sent, observing an empty stream, letting a failing check break the stream, dropping the background call's headers, URL or configured timeout, swapping success and not_run, returning empty inputs, removing the observe-only warning, not forwarding either option from the config or the max_inflight to the dispatcher, letting an unreadable value turn fire_and_forget on, rejecting "true", raising on or ignoring max_inflight, accepting true as a max_inflight, an off-by-one cap, never freeing a slot, letting a background failure escape, running the task in the request's context, warning on every drop, and building the payload before the cap check
  • The handful of test files covering my change pass locally, e.g. uv run pytest tests/unit/<your_test_file>.py -v. Leave the suites (make test-unit-*, make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more
  • My PR passes all required CI/CD checks (e.g., lint, schema.d.ts sync check, etc.)
    • Of the non-required jobs, budget-ratchet fails because main dropped the LIT002 budget in chore(lint): remove the LIT002 mutable-construction rule #43971 after this branch's merge base, and osv-scan flags a braces advisory in the dashboard's package-lock.json. This PR touches neither budgets nor dependencies
    • make check passes locally with BASE_REF=upstream/main. A fork whose main lags upstream fails the budget gates against that stale base, so set BASE_REF when checking a fork
    • codecov/patch passes, and Codecov reports every modified line as covered
  • My PR's scope is as isolated as possible; it only solves 1 specific problem
  • I have received a Greptile Confidence Score of at least 4/5 before requesting a maintainer review (Greptile reviews automatically once the PR is opened; only comment @greptileai to re-request a review after pushing changes)
    • Greptile scored 5/5 on de597a5. Its abandoned-stream and payload-serialization findings are fixed in de06532 and 09922b7, the shutdown-drain finding is a listed caveat, and it accepted the reply on the stream copy
    • Veria reported no security issues on de597a5
    • Bugbot has not run on this PR

Delays in PR merge?

If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).

Screenshots / Proof of Fix

All runs are against a live proxy on localhost:4000 started with python litellm/proxy/proxy_cli.py --config proof_config.yaml --detailed_debug --use_v2_migration_resolver, backed by a real Postgres, and every completion is a real, billed call to Bedrock us.anthropic.claude-haiku-4-5-20251001-v1:0. $PROOF_KEY is a virtual key from POST /key/generate with access to bedrock-haiku-4-5. Non-streaming requests set metadata.include_guardrail_response, so the response lists the guardrail entries recorded for that call. Each request picks its guardrail with "guardrails": [...]

model_list:
  - model_name: bedrock-haiku-4-5
    litellm_params:
      model: bedrock/us.anthropic.claude-haiku-4-5-20251001-v1:0
      aws_region_name: us-east-1

guardrails:
  - guardrail_name: audit
    litellm_params:
      guardrail: generic_guardrail_api
      mode: [pre_call, post_call]
      api_base: http://127.0.0.1:8787
      api_key: audit-key
      default_on: false
      fire_and_forget: true
  - guardrail_name: audit-capped
    litellm_params:
      guardrail: generic_guardrail_api
      mode: pre_call
      api_base: http://127.0.0.1:8787
      api_key: audit-key
      default_on: false
      fire_and_forget: true
      fire_and_forget_max_inflight: 1
  - guardrail_name: audit-down
    litellm_params:
      guardrail: generic_guardrail_api
      mode: pre_call
      api_base: http://127.0.0.1:8799
      default_on: false
      fire_and_forget: true
  - guardrail_name: audit-post
    litellm_params:
      guardrail: generic_guardrail_api
      mode: post_call
      api_base: http://127.0.0.1:8787
      api_key: audit-key
      default_on: false
      fire_and_forget: true
  - guardrail_name: audit-diff
    litellm_params:
      guardrail: generic_guardrail_api
      mode: post_call
      api_base: http://127.0.0.1:8787
      api_key: audit-key
      default_on: false
      fire_and_forget: true
      streaming_transform_mode: incremental_diff

general_settings:
  master_key: os.environ/LITELLM_MASTER_KEY
  database_url: os.environ/DATABASE_URL
  proxy_batch_write_at: 1

At the merge base fire_and_forget and fire_and_forget_max_inflight do not exist, so the proxy ignores them and awaits every guardrail call. The audit endpoint is a small FastAPI app on 127.0.0.1:8787 that waits 5 seconds, logs what it received, and answers BLOCKED for everything. audit-down points at a port where nothing listens, and audit-post only checks responses. Nothing else is stubbed

import asyncio
import json
import time
from pathlib import Path

from fastapi import FastAPI, Request

LOG = Path(__file__).with_name("audit.jsonl")
app = FastAPI()


@app.post("/beta/litellm_basic_guardrail_api")
async def guardrail(request: Request) -> dict[str, str]:
    payload = await request.json()
    received = time.time()
    await asyncio.sleep(5)
    with LOG.open("a") as log:
        log.write(
            json.dumps(
                {
                    "received": received,
                    "input_type": payload.get("input_type"),
                    "litellm_call_id": payload.get("litellm_call_id"),
                    "x_api_key": request.headers.get("x-api-key"),
                    "texts": payload.get("texts"),
                }
            )
            + "\n"
        )
    return {"action": "BLOCKED", "blocked_reason": "audit says no"}


@app.post("/_reset")
async def reset() -> dict[str, int]:
    LOG.write_text("")
    return {"records": 0}

The Admin UI screenshots for the same runs are in User Flow above

Before (6f123b7)

Case 1: the audit endpoint takes 5 s and answers BLOCKED

  1. Run
    $ curl -s -o chat.json -w '%{http_code} %{time_total}s\n' http://localhost:4000/v1/chat/completions -H "Authorization: Bearer $PROOF_KEY" -H 'Content-Type: application/json' -d '{"model": "bedrock-haiku-4-5", "max_tokens": 10, "guardrails": ["audit"], "metadata": {"include_guardrail_response": true}, "messages": [{"role": "user", "content": "Reply with one word: hi"}]}'; jq -c 'if .error then {error: .error.message} else [.guardrail_information[] | {guardrail_status, guardrail_response}] end' chat.json
    
  2. Observe
    400 5.055909s
    {"error":"audit says no"}
    
  3. Run
    $ curl -s -o responses.json -w '%{http_code} %{time_total}s\n' http://localhost:4000/v1/responses -H "Authorization: Bearer $PROOF_KEY" -H 'Content-Type: application/json' -d '{"model": "bedrock-haiku-4-5", "max_output_tokens": 16, "guardrails": ["audit"], "metadata": {"include_guardrail_response": true}, "input": "Reply with one word: hi"}'; jq -c 'if .error then {error: .error.message} else [.guardrail_information[] | {guardrail_status, guardrail_response}] end' responses.json
    
  4. Observe
    400 5.028657s
    {"error":"audit says no"}
    
  5. Run
    $ curl -s -o messages.json -w '%{http_code} %{time_total}s\n' http://localhost:4000/v1/messages -H "Authorization: Bearer $PROOF_KEY" -H 'Content-Type: application/json' -d '{"model": "bedrock-haiku-4-5", "max_tokens": 10, "guardrails": ["audit"], "metadata": {"include_guardrail_response": true}, "messages": [{"role": "user", "content": "Reply with one word: hi"}]}'; jq -c 'if .error then {error: .error.message} else [.guardrail_information[] | {guardrail_status, guardrail_response}] end' messages.json
    
  6. Observe
    400 5.027310s
    {"error":"audit says no"}
    

Case 2: the audit endpoint still receives every call with its api_key

  1. Run
    $ sleep 6; jq -c '{input_type, x_api_key, chars: ((.texts // [""])[0] | length), text: ((.texts // [""])[0] | .[0:40])}' audit.jsonl
    
  2. Observe
    {"input_type":"request","x_api_key":"audit-key","chars":23,"text":"Reply with one word: hi"}
    {"input_type":"request","x_api_key":"audit-key","chars":23,"text":"Reply with one word: hi"}
    {"input_type":"request","x_api_key":"audit-key","chars":23,"text":"Reply with one word: hi"}
    

Case 3: fire_and_forget_max_inflight=1 with three requests at once

  1. Run
    $ for i in 1 2 3; do curl -s -o capped_$i.json http://localhost:4000/v1/chat/completions -H "Authorization: Bearer $PROOF_KEY" -H 'Content-Type: application/json' -d '{"model": "bedrock-haiku-4-5", "max_tokens": 10, "guardrails": ["audit-capped"], "metadata": {"include_guardrail_response": true}, "messages": [{"role": "user", "content": "Reply with one word: hi"}]}' & done; wait; for i in 1 2 3; do jq -c 'if .error then {error: .error.message} else [.guardrail_information[] | {guardrail_status, guardrail_response}] end' capped_$i.json; done
    
  2. Observe
    {"error":"audit says no"}
    {"error":"audit says no"}
    {"error":"audit says no"}
    
  3. Run
    $ sleep 6; jq -c '{input_type, x_api_key, chars: ((.texts // [""])[0] | length), text: ((.texts // [""])[0] | .[0:40])}' audit.jsonl
    
  4. Observe
    {"input_type":"request","x_api_key":"audit-key","chars":23,"text":"Reply with one word: hi"}
    {"input_type":"request","x_api_key":"audit-key","chars":23,"text":"Reply with one word: hi"}
    {"input_type":"request","x_api_key":"audit-key","chars":23,"text":"Reply with one word: hi"}
    
  5. Run
    $ grep -o 'dropped [0-9]* call(s) so far.*max_inflight=[0-9]*)' litellm.log
    
  6. Observe: no output

Case 4: the audit endpoint is down

  1. Run
    $ curl -s -o down.json -w '%{http_code} %{time_total}s\n' http://localhost:4000/v1/chat/completions -H "Authorization: Bearer $PROOF_KEY" -H 'Content-Type: application/json' -d '{"model": "bedrock-haiku-4-5", "max_tokens": 10, "guardrails": ["audit-down"], "metadata": {"include_guardrail_response": true}, "messages": [{"role": "user", "content": "Reply with one word: hi"}]}'; jq -c 'if .error then {error: .error.message} else [.guardrail_information[] | {guardrail_status, guardrail_response}] end' down.json
    
  2. Observe
    500 0.033690s
    {"error":"Generic Guardrail API failed: Cannot connect to host REDACTED:8799 ssl:<ssl.SSLContext object at 0x110b91bf0> [Connect call failed ('REDACTED', 8799)]"}
    
  3. Run
    $ grep -o 'Generic Guardrail API (audit-down, fire_and_forget) call failed[^:]*' litellm.log
    
  4. Observe: no output

Case 5: streaming, time to first byte and total time

  1. Run
    $ curl -sN -o stream.txt -w 'http %{http_code}, first byte %{time_starttransfer}s, total %{time_total}s\n' http://localhost:4000/v1/chat/completions -H "Authorization: Bearer $PROOF_KEY" -H 'Content-Type: application/json' -d '{"model": "bedrock-haiku-4-5", "max_tokens": 400, "stream": true, "guardrails": ["audit"], "messages": [{"role": "user", "content": "Write a 250 word story about a lighthouse"}]}'; grep -c '^data: {' stream.txt
    
  2. Observe
    http 400, first byte 5.017685s, total 5.017781s
    0
    
  3. Run
    $ sleep 6; jq -c '{input_type, x_api_key, chars: ((.texts // [""])[0] | length), text: ((.texts // [""])[0] | .[0:40])}' audit.jsonl
    
  4. Observe
    {"input_type":"request","x_api_key":"audit-key","chars":41,"text":"Write a 250 word story about a lighthous"}
    

Case 6: streaming with streaming_transform_mode: incremental_diff

  1. Run
    $ curl -sN -o diff.txt -w 'http %{http_code}, first byte %{time_starttransfer}s, total %{time_total}s\n' http://localhost:4000/v1/chat/completions -H "Authorization: Bearer $PROOF_KEY" -H 'Content-Type: application/json' -d '{"model": "bedrock-haiku-4-5", "max_tokens": 400, "stream": true, "guardrails": ["audit-diff"], "messages": [{"role": "user", "content": "Write a 250 word story about a lighthouse"}]}'; grep -c '^data: {' diff.txt
    
  2. Observe
    http 400, first byte 7.458337s, total 7.458581s
    0
    
  3. Run
    $ sleep 6; jq -c '{input_type, x_api_key, chars: ((.texts // [""])[0] | length), text: ((.texts // [""])[0] | .[0:40])}' audit.jsonl
    
  4. Observe
    {"input_type":"response","x_api_key":"audit-key","chars":49,"text":"# The Lighthouse Keeper's Secret\n\nFor fo"}
    

Case 7: the client hangs up mid-stream

  1. Run
    $ curl -sN --max-time 2.5 -o hangup.txt -w 'http %{http_code}, client gave up after %{time_total}s\n' http://localhost:4000/v1/chat/completions -H "Authorization: Bearer $PROOF_KEY" -H 'Content-Type: application/json' -d '{"model": "bedrock-haiku-4-5", "max_tokens": 400, "stream": true, "guardrails": ["audit-post"], "messages": [{"role": "user", "content": "Write a 250 word story about a lighthouse"}]}'; sed -n 's/^data: //p' hangup.txt | jq -Rrj 'fromjson? | .choices[0].delta.content // empty' | wc -m
    
  2. Observe
    http 200, client gave up after 2.507538s
    41
    
  3. Run
    $ sleep 6; jq -c '{input_type, x_api_key, chars: ((.texts // [""])[0] | length), text: ((.texts // [""])[0] | .[0:40])}' audit.jsonl
    
  4. Observe
    {"input_type":"response","x_api_key":"audit-key","chars":47,"text":"# The Lighthouse Keeper's Last Light\n\nFo"}
    

Case 8: startup warning

  1. Run
    $ grep -o 'Generic Guardrail API (audit): fire_and_forget=True.*' litellm.log
    
  2. Observe: no output

After (de597a5)

Case 1: the audit endpoint takes 5 s and answers BLOCKED

  1. Run
    $ curl -s -o chat.json -w '%{http_code} %{time_total}s\n' http://localhost:4000/v1/chat/completions -H "Authorization: Bearer $PROOF_KEY" -H 'Content-Type: application/json' -d '{"model": "bedrock-haiku-4-5", "max_tokens": 10, "guardrails": ["audit"], "metadata": {"include_guardrail_response": true}, "messages": [{"role": "user", "content": "Reply with one word: hi"}]}'; jq -c 'if .error then {error: .error.message} else [.guardrail_information[] | {guardrail_status, guardrail_response}] end' chat.json
    
  2. Observe
    200 2.613779s
    [{"guardrail_status":"success","guardrail_response":"fire_and_forget dispatched, verdict not read"}]
    
  3. Run
    $ curl -s -o responses.json -w '%{http_code} %{time_total}s\n' http://localhost:4000/v1/responses -H "Authorization: Bearer $PROOF_KEY" -H 'Content-Type: application/json' -d '{"model": "bedrock-haiku-4-5", "max_output_tokens": 16, "guardrails": ["audit"], "metadata": {"include_guardrail_response": true}, "input": "Reply with one word: hi"}'; jq -c 'if .error then {error: .error.message} else [.guardrail_information[] | {guardrail_status, guardrail_response}] end' responses.json
    
  4. Observe
    200 2.106694s
    [{"guardrail_status":"success","guardrail_response":"fire_and_forget dispatched, verdict not read"}]
    
  5. Run
    $ curl -s -o messages.json -w '%{http_code} %{time_total}s\n' http://localhost:4000/v1/messages -H "Authorization: Bearer $PROOF_KEY" -H 'Content-Type: application/json' -d '{"model": "bedrock-haiku-4-5", "max_tokens": 10, "guardrails": ["audit"], "metadata": {"include_guardrail_response": true}, "messages": [{"role": "user", "content": "Reply with one word: hi"}]}'; jq -c 'if .error then {error: .error.message} else [.guardrail_information[] | {guardrail_status, guardrail_response}] end' messages.json
    
  6. Observe
    200 2.206762s
    [{"guardrail_status":"success","guardrail_response":"fire_and_forget dispatched, verdict not read"}]
    

Case 2: the audit endpoint still receives every call with its api_key

  1. Run
    $ sleep 6; jq -c '{input_type, x_api_key, chars: ((.texts // [""])[0] | length), text: ((.texts // [""])[0] | .[0:40])}' audit.jsonl
    
  2. Observe
    {"input_type":"request","x_api_key":"audit-key","chars":23,"text":"Reply with one word: hi"}
    {"input_type":"response","x_api_key":"audit-key","chars":5,"text":"Hello"}
    {"input_type":"request","x_api_key":"audit-key","chars":23,"text":"Reply with one word: hi"}
    {"input_type":"response","x_api_key":"audit-key","chars":5,"text":"Hello"}
    {"input_type":"request","x_api_key":"audit-key","chars":23,"text":"Reply with one word: hi"}
    {"input_type":"response","x_api_key":"audit-key","chars":5,"text":"Hello"}
    

Case 3: fire_and_forget_max_inflight=1 with three requests at once

  1. Run
    $ for i in 1 2 3; do curl -s -o capped_$i.json http://localhost:4000/v1/chat/completions -H "Authorization: Bearer $PROOF_KEY" -H 'Content-Type: application/json' -d '{"model": "bedrock-haiku-4-5", "max_tokens": 10, "guardrails": ["audit-capped"], "metadata": {"include_guardrail_response": true}, "messages": [{"role": "user", "content": "Reply with one word: hi"}]}' & done; wait; for i in 1 2 3; do jq -c 'if .error then {error: .error.message} else [.guardrail_information[] | {guardrail_status, guardrail_response}] end' capped_$i.json; done
    
  2. Observe
    [{"guardrail_status":"success","guardrail_response":"fire_and_forget dispatched, verdict not read"}]
    [{"guardrail_status":"not_run","guardrail_response":"fire_and_forget_max_inflight reached, call dropped"}]
    [{"guardrail_status":"not_run","guardrail_response":"fire_and_forget_max_inflight reached, call dropped"}]
    
  3. Run
    $ sleep 6; jq -c '{input_type, x_api_key, chars: ((.texts // [""])[0] | length), text: ((.texts // [""])[0] | .[0:40])}' audit.jsonl
    
  4. Observe
    {"input_type":"request","x_api_key":"audit-key","chars":23,"text":"Reply with one word: hi"}
    
  5. Run
    $ grep -o 'dropped [0-9]* call(s) so far.*max_inflight=[0-9]*)' litellm.log
    
  6. Observe
    dropped 1 call(s) so far, 1 already in flight (fire_and_forget_max_inflight=1)
    

Case 4: the audit endpoint is down

  1. Run
    $ curl -s -o down.json -w '%{http_code} %{time_total}s\n' http://localhost:4000/v1/chat/completions -H "Authorization: Bearer $PROOF_KEY" -H 'Content-Type: application/json' -d '{"model": "bedrock-haiku-4-5", "max_tokens": 10, "guardrails": ["audit-down"], "metadata": {"include_guardrail_response": true}, "messages": [{"role": "user", "content": "Reply with one word: hi"}]}'; jq -c 'if .error then {error: .error.message} else [.guardrail_information[] | {guardrail_status, guardrail_response}] end' down.json
    
  2. Observe
    200 1.741006s
    [{"guardrail_status":"success","guardrail_response":"fire_and_forget dispatched, verdict not read"}]
    
  3. Run
    $ grep -o 'Generic Guardrail API (audit-down, fire_and_forget) call failed[^:]*' litellm.log
    
  4. Observe
    Generic Guardrail API (audit-down, fire_and_forget) call failed. input_type=request litellm_call_id=cde4052e-adcc-4051-ba7b-a0922bbb0e8d
    

Case 5: streaming, time to first byte and total time

  1. Run
    $ curl -sN -o stream.txt -w 'http %{http_code}, first byte %{time_starttransfer}s, total %{time_total}s\n' http://localhost:4000/v1/chat/completions -H "Authorization: Bearer $PROOF_KEY" -H 'Content-Type: application/json' -d '{"model": "bedrock-haiku-4-5", "max_tokens": 400, "stream": true, "guardrails": ["audit"], "messages": [{"role": "user", "content": "Write a 250 word story about a lighthouse"}]}'; grep -c '^data: {' stream.txt
    
  2. Observe
    http 200, first byte 2.267969s, total 5.715412s
    146
    
  3. Run
    $ sleep 6; jq -c '{input_type, x_api_key, chars: ((.texts // [""])[0] | length), text: ((.texts // [""])[0] | .[0:40])}' audit.jsonl
    
  4. Observe
    {"input_type":"request","x_api_key":"audit-key","chars":41,"text":"Write a 250 word story about a lighthous"}
    {"input_type":"response","x_api_key":"audit-key","chars":1429,"text":"# The Last Light\n\nThe lighthouse stood a"}
    

Case 6: streaming with streaming_transform_mode: incremental_diff

  1. Run
    $ curl -sN -o diff.txt -w 'http %{http_code}, first byte %{time_starttransfer}s, total %{time_total}s\n' http://localhost:4000/v1/chat/completions -H "Authorization: Bearer $PROOF_KEY" -H 'Content-Type: application/json' -d '{"model": "bedrock-haiku-4-5", "max_tokens": 400, "stream": true, "guardrails": ["audit-diff"], "messages": [{"role": "user", "content": "Write a 250 word story about a lighthouse"}]}'; grep -c '^data: {' diff.txt
    
  2. Observe
    http 200, first byte 1.778948s, total 5.172660s
    145
    
  3. Run
    $ sleep 6; jq -c '{input_type, x_api_key, chars: ((.texts // [""])[0] | length), text: ((.texts // [""])[0] | .[0:40])}' audit.jsonl
    
  4. Observe
    {"input_type":"response","x_api_key":"audit-key","chars":1506,"text":"# The Last Light\n\nThe lighthouse stood a"}
    

Case 7: the client hangs up mid-stream

  1. Run
    $ curl -sN --max-time 2.5 -o hangup.txt -w 'http %{http_code}, client gave up after %{time_total}s\n' http://localhost:4000/v1/chat/completions -H "Authorization: Bearer $PROOF_KEY" -H 'Content-Type: application/json' -d '{"model": "bedrock-haiku-4-5", "max_tokens": 400, "stream": true, "guardrails": ["audit-post"], "messages": [{"role": "user", "content": "Write a 250 word story about a lighthouse"}]}'; sed -n 's/^data: //p' hangup.txt | jq -Rrj 'fromjson? | .choices[0].delta.content // empty' | wc -m
    
  2. Observe
    http 200, client gave up after 2.503534s
    351
    
  3. Run
    $ sleep 6; jq -c '{input_type, x_api_key, chars: ((.texts // [""])[0] | length), text: ((.texts // [""])[0] | .[0:40])}' audit.jsonl
    
  4. Observe
    {"input_type":"response","x_api_key":"audit-key","chars":351,"text":"# The Last Light\n\nThe old lighthouse sto"}
    

Case 8: startup warning

  1. Run
    $ grep -o 'Generic Guardrail API (audit): fire_and_forget=True.*' litellm.log
    
  2. Observe
    Generic Guardrail API (audit): fire_and_forget=True makes this guardrail observe-only. action=BLOCKED and action=GUARDRAIL_INTERVENED are ignored, fail_on_error=True and unreachable_fallback=fail_closed cannot block the request, and a stream is checked once when it closes, with whatever reached the client.
    

Type

🆕 New Feature

Caveats (if any)

Severe

  • With fire_and_forget on, the guardrail never blocks or rewrites anything
    • BLOCKED and GUARDRAIL_INTERVENED answers are ignored
    • fail_on_error and unreachable_fallback can no longer fail the request

Medium

  • A dispatched call is recorded as success, so it counts as a pass
    • The logs page shows it PASSED with Risk 0/10, and the usage dashboard counts it
    • Same as akto's fire_and_forget and the generic guardrail's fail-open path
  • A background call that fails only logs a warning
  • The in-flight cap is per guardrail per worker process
    • N workers allow N times fire_and_forget_max_inflight
  • Each in-flight payload, images included, stays in memory until its call ends
  • Background calls still in flight are lost on proxy shutdown

Low

  • A stream is checked once when it closes, ignoring the streaming_* settings
    • This adds an opt-in streaming_observe_only path to the shared streaming hook
    • Guardrails that don't set it behave exactly as before
    • The check works on a deep copy, so chunks already sent stay untouched
  • On a hang-up while the proxy waits to send, the check runs when Python finalizes the stream
    • That can be a few event loop turns later, and case 7 shows the audit still arrives
  • Partial streams on other routes are checked by their existing handlers
    • /v1/responses cut after a tool call event only sends the tool calls
    • /v1/messages cut before any text sends an empty text
  • Background calls use timeout, or 30 seconds when it is unset
    • The 30 seconds apply per phase, so a slot can stay busy longer
  • An unreadable fire_and_forget is ignored with a warning and keeps enforcing
  • An invalid fire_and_forget_max_inflight behaves differently by config path
    • In YAML it falls back to 100 with a warning
    • The Admin UI form rejects it
  • Dropped calls log a warning only once every 100 drops
    • Each dropped request still gets its own not_run entry
  • Background call warnings lose trace_id and session_id correlation
    • They still carry litellm_call_id, input type and guardrail name
  • Dispatched calls show 0ms guardrail overhead because nobody waits
  • The applied-guardrails response header still lists a dropped call's guardrail
  • Re-creating the guardrail from the Admin UI starts a fresh in-flight count
  • The UI "test guardrail" call always passes for this guardrail
  • The "Ignoring ..." config warnings don't name the guardrail, same as feat(guardrails): add guardrail_information_scope to generic_guardrail_api #43680
  • Merge conflicts with feat(guardrails): add guardrail_information_scope to generic_guardrail_api #43680 and fix(guardrails): keep masked text when a guardrail echoes multipart rows #43679 are expected

Final Attestation

  • The tests check the right things, including the edge cases, and regressions in the respective real-world customer use-cases are not possible after this PR

@CLAassistant

CLAassistant commented Sep 29, 2026 •

Copy link
Copy Markdown

CLA assistant check
All committers have signed the CLA.

@codspeed

codspeed Bot commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Merging this PR will not alter performance

✅ 31 untouched benchmarks


Comparing caduri:feature/generic-guardrail-fire-and-forget (de597a5) with main (e768ad5)

Open in CodSpeed

@codecov

codecov Bot commented Sep 29, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

With fire_and_forget enabled the guardrail POST runs as a detached task and
apply_guardrail returns the inputs unchanged right away, so the guardrail
becomes observe-only. A dispatched call is recorded as success with a response
saying the verdict was not read. At most fire_and_forget_max_inflight calls
(default 100) are in flight at once per guardrail and worker. Extra calls are
dropped, counted with a rate-limited warning, and recorded as guardrail_status
not_run. The background POST has a fixed 30 second timeout and starts from an
empty context, so it does not keep request-scoped context vars alive

Failures in the background call are logged and never raised. A failure while
building the payload before dispatch is logged, passes the request through
regardless of fail_on_error, and is recorded as not_run with its own reason. A
non-bool fire_and_forget or a non-int fire_and_forget_max_inflight is rejected
at startup. Enabling it also forces streaming_end_of_stream_only so a stream
sends one call
Parse fire_and_forget and fire_and_forget_max_inflight the way pydantic parses config values, so "true" and "5" work. A value that cannot be read is ignored with a warning: fire_and_forget falls back to false, so the guardrail keeps enforcing, and fire_and_forget_max_inflight falls back to 100

Force streaming_transform_mode to block_only under fire_and_forget. With incremental_diff the unified guardrail buffered the stream and dispatched one call per chunk, so a stream now sends a single end-of-stream call while the chunks reach the client live

The background POST now honors the guardrail's timeout and falls back to 30 seconds when it is unset. It is built from the same URL, headers and payload as the awaited call. A payload that cannot be serialized is recorded as not_run before anything is dispatched

Tests that build GenericGuardrailAPI move to a mirror test_generic_guardrail_api.py, the dispatcher and parser tests stay in test_background_dispatch.py, and a shared conftest.py captures proxy warnings
@caduri

caduri commented Oct 3, 2026

Copy link
Copy Markdown
Author

@greptileai please re-review the latest commit, which hardens config parsing, forces block_only streaming and honors the guardrail timeout for background calls

@greptile-apps

greptile-apps Bot commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[Medium risk] Adds fire-and-forget dispatch mode to guardrail API.

The PR appears safe to merge based on the changes since the previous review and the current state of the prior findings.

Summary

The PR adds observe-only background dispatch for the Generic Guardrail API, with an in-flight cap and a close-time audit for streams. Since the previous review, the only changes broaden a test helper’s input type and add a test for an unrecognized stream.

Reviews (3) · Last reviewed commit: "test(guardrails): cover an observe-only ..."

caduri added 2 commits October 3, 2026 15:48
…free

BackgroundDispatcher.dispatch now takes a prepare callable and calls it only after the in-flight cap check passes. A call dropped by fire_and_forget_max_inflight no longer serializes its payload on the request path, which matters most for large or image-heavy requests while the endpoint is slow

A payload that cannot be serialized is still recorded as not_run before anything is dispatched
…ails

With fire_and_forget the response audit ran only after a stream finished, so a stream the client hung up on, or one that failed upstream, never reached the audit endpoint

Guardrails can now set streaming_observe_only. UnifiedLLMGuardrails then forwards every chunk untouched and, when the stream closes for any reason, runs one check over a copy of what reached the client. The check is shielded from cancellation, and a failure in it is logged and never breaks the stream. GenericGuardrailAPI sets the flag from fire_and_forget instead of forcing end-of-stream and block_only, which the new path makes unnecessary. Other guardrails are unchanged
@caduri

caduri commented Oct 3, 2026

Copy link
Copy Markdown
Author

@greptileai please re-review: abandoned streams are now observed, dropped calls skip serialization, and the PR body has a fresh live proof

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants