Skip to content

fix(langsmith): set AUTH_ENDPOINT in the shared ConfigMap [v16-stable] - #846

Merged
Mukil Loganathan (langchain-infra) merged 3 commits into
v16-stablefrom
fix/gateway-auth-endpoint-v16
Jul 28, 2026
Merged

Mukil Loganathan (langchain-infra) merged 3 commits into
v16-stablefrom
fix/gateway-auth-endpoint-v16

Conversation

@langchain-infra

@langchain-infra Mukil Loganathan (langchain-infra) commented Jul 27, 2026 •

Copy link
Copy Markdown
Contributor

v16-stable port of #845. Same change; only the chart version differs.

Problem

The agent gateway resolves frontend bearer tokens against the auth service. It reads the AUTH_ENDPOINT env var, which this chart never set — not in the gateway Deployment, and not in the shared ConfigMap (which defines LANGSMITH_AUTH_ENDPOINT, a different key the gateway does not read).

Whenever the gateway cannot construct a local auth backend, it panics at startup:

panic: gateway frontend bearer authentication requires AUTH_ENDPOINT for AUTH_TYPE "mixed"
  smith-go/gateway/middleware.go

The pod CrashLoopBackOffs, the rollout never completes, and the Helm upgrade fails with context deadline exceeded. This reproduced on a live mixed + OAuth cluster.

The trigger is narrow. In newFrontendBearerAuthWithConfig, the "use local auth" early return requires a non-nil local handler:

authType local handler Result
none NewNone returns before the check
oauth (AuthTypeOAuthPkce) NewOAuth early return, no panic
mixed + basic auth on NewBasicAuth early return, no panic
mixed + basic auth off nil — switch falls through panics

The panic is only a guard; the real consumer is auth.NewAuthBackend → NewAuthClientImpl(config.Env.AuthEndpoint). Removing the check would just produce silent HTTP calls to an empty base URL, so the var genuinely has to be set. The gateway cannot self-serve here: sessioned mixed auth is initialized by the main service, so the gateway must delegate over HTTP rather than duplicate goth/session state.

Fix

Set AUTH_ENDPOINT in the shared ConfigMap, pointing at the platform-backend service — the same target GO_ENDPOINT and LANGSMITH_AUTH_ENDPOINT already resolve to. It respects platformBackend.name, platformBackend.service.port, namespace, and clusterDomain.

The gateway Deployment already consumes the shared ConfigMap via envFrom, so no inline env entry is needed. Setting it centrally covers every consumer of AUTH_ENDPOINT rather than the single Deployment we happened to patch.

Setting it cluster-wide is inert for the other services that read the key:

  • Python — ServiceName.AUTH switches from its GO_ENDPOINT fallback to AUTH_ENDPOINT. Same URL in this chart, and the timeouts are identical (AUTH_TIMEOUT_SECS and GO_TIMEOUT_SECS both default to 2).
  • Go — nothing branches on AuthEndpoint != "". The auth decision tree keys off isDataPlane / authServiceSeparationEnabled / whether a local handler exists, and NewAuthBackend is only constructed on paths those already gate. The one other emptiness check is in testutil.

Chart version

Bumped to 0.16.0-rc.21. This branch previously carried 0.16.0-rc.19, which is now already tagged, and v16-stable has since moved to 0.16.0-rc.20. chart-releaser runs with CR_SKIP_EXISTING: true, so reusing a published version would be silently dropped.

Test Plan

  • New tests/auth_endpoint_test.yaml (6 cases): default render, each authType branch (mixed+OAuth, oauth, mixed+basic), parity with LANGSMITH_AUTH_ENDPOINT, custom platformBackend.name / service.port / clusterDomain, and the gateway's envFrom reference
  • Full chart suite green — 138 tests, 12 suites, no regressions
  • helm template on a mixed + OAuth config renders AUTH_ENDPOINT, LANGSMITH_AUTH_ENDPOINT, and GO_ENDPOINT all as http://<release>-platform-backend.<ns>.svc.cluster.local:1986, with no inline AUTH_ENDPOINT on the gateway container
  • Post-merge: confirm the gateway pod reaches Ready and the Helm upgrade completes

Follow-ups (not in this PR)

  • The durable fix belongs in smith-go: fall back to GO_ENDPOINT / LANGSMITH_AUTH_ENDPOINT when AUTH_ENDPOINT is unset and auth-service separation is off. In a single-cluster install the auth service is the platform backend. That covers self-hosted topologies a chart patch cannot reach, and the sibling supabase branch, which builds NewAuthClientImpl("") and degrades silently instead of panicking.
  • The platform-backend URL template is now written 3× in config-map.yaml. A langsmith.platformBackendURL helper would stop the copies drifting.
  • main and v16-stable have previously shared a chart version with different appVersions; worth tracking separately given CR_SKIP_EXISTING.

The gateway resolves frontend bearer tokens against the auth service and
reads AUTH_ENDPOINT, which the chart never set. Whenever the gateway
cannot build a local auth backend it panics at startup with:

  panic: gateway frontend bearer authentication requires AUTH_ENDPOINT
         for AUTH_TYPE "mixed"

Point AUTH_ENDPOINT at the platform-backend service, the same target
LANGSMITH_AUTH_ENDPOINT already uses in the shared ConfigMap. Scoped to
the gateway Deployment rather than the ConfigMap so the other services'
resolution of AUTH_ENDPOINT is unchanged.

Chart version goes to rc.19 rather than rc.18 so this branch and main do
not publish different charts under the same version.
AUTH_ENDPOINT was set as an inline env var on the agent gateway Deployment
only. Move it to the shared ConfigMap alongside GO_ENDPOINT and
LANGSMITH_AUTH_ENDPOINT, which already resolve to the same platform-backend
service.

The gateway already consumes the shared ConfigMap via envFrom, so it still
receives the value. Setting it centrally also covers every other consumer of
AUTH_ENDPOINT instead of the single Deployment we happened to patch.

Replace the gateway-scoped test with a ConfigMap suite covering the key's
presence across each authType branch, that it matches LANGSMITH_AUTH_ENDPOINT,
custom platformBackend name/port/clusterDomain, and the gateway's envFrom
reference that delivers it.
…h-endpoint-v16

# Conflicts:
#	charts/langsmith/Chart.yaml
@langchain-infra Mukil Loganathan (langchain-infra) changed the title fix(langsmith): set AUTH_ENDPOINT on the agent gateway [v16-stable] fix(langsmith): set AUTH_ENDPOINT in the shared ConfigMap [v16-stable] Jul 28, 2026
@langchain-infra
Mukil Loganathan (langchain-infra) merged commit b9c243b into v16-stable Jul 28, 2026
6 checks passed
@langchain-infra
Mukil Loganathan (langchain-infra) deleted the fix/gateway-auth-endpoint-v16 branch July 28, 2026 22:45
Bagatur (baskaryan) added a commit that referenced this pull request Jul 30, 2026
* Bump langsmith appVersion to 0.16.22rc1 (#842)

* Bump langsmith appVersion 0.16.21rc1 -> 0.16.22rc1

* Bump langsmith chart version to 0.16.0-rc.17

* Bump langsmith appVersion to 0.16.23rc1 (#853)

* fix(langsmith): normalize SmithDB OTLP endpoint scheme (#843)

* fix(langsmith): normalize SmithDB OTLP endpoint scheme

Derive the URI scheme required by SmithDB from the chart's TLS setting while preserving the bare endpoint consumed by existing LangSmith services.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(langsmith): allow disabling SmithDB tracing

Keep shared TLS tracing available to LangSmith services by adding a SmithDB-specific opt-out and validating TLS only when SmithDB export remains enabled.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(langsmith): disable SmithDB tracing with TLS

Keep TLS tracing available to supported services while SmithDB falls back to console logging until its OTLP exporter supports TLS.

Co-authored-by: Cursor <cursoragent@cursor.com>

* refactor(langsmith): simplify SmithDB tracing condition

Use the chart's explicit TLS setting as the source of truth for SmithDB exporter compatibility.

Co-authored-by: Cursor <cursoragent@cursor.com>

* docs(langsmith): clarify SmithDB TLS fallback

Distinguish disabled OTLP tracing from console-formatted stdout logging.

Co-authored-by: Cursor <cursoragent@cursor.com>

* refactor(langsmith): simplify SmithDB endpoint normalization

Make plaintext normalization explicit now that TLS compatibility is handled before rendering the endpoint.

Co-authored-by: Cursor <cursoragent@cursor.com>

* refactor(langsmith): prepend SmithDB OTLP scheme directly

Rely on the chart's bare endpoint contract instead of accepting protocol-bearing values.

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(langsmith): enable SmithDB OTLP TLS

Derive SmithDB's required endpoint URI scheme from the shared TLS setting while preserving the bare endpoint used by other services.

---------

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat: backport agent IAM providers to v16-stable (#852)

* feat: configure agent database IAM providers

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: validate agent IAM providers

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix: nest agent IAM providers under external databases

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: bump v16 chart version to 0.16.0-rc.18

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: joaquin-borggio-lc <213688804+joaquin-borggio-lc@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix(langsmith): template TRIGGER_SERVER_ENDPOINT in shared ConfigMap [v16-stable] (#856)

* fix(langsmith): template TRIGGER_SERVER_ENDPOINT in shared ConfigMap

v16-stable port. Fleet trigger proxy/client in platform-backend and the
tool-server need TRIGGER_SERVER_ENDPOINT. Wire the in-cluster
trigger-server URL into the shared ConfigMap when fleet +
fleetTriggerServer are enabled so customers do not need to set commonEnv.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore: drop TRIGGER_SERVER_ENDPOINT unittest

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore: drop TRIGGER_SERVER_ENDPOINT config-map comments

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore: bump chart version by one (rc.19)

Only increment from v16-stable's rc.18 instead of jumping to rc.22.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>

* feat(langsmith): bake stable-smithdb flag defaults into config-map (#848) (#860)

Extends the main -config ConfigMap with the LangSmith app feature-flag envs
that differ from code defaults for a stable SmithDB-backed deployment. New
SMITHDB_* / RUN_RULES_SMITHDB_* keys are gated on smithdb.enabled (same block
as SMITHDB_CLUSTER_MANAGER_ENABLED), so ClickHouse-only clusters keep app
code defaults. RUN_RULES_TWO_POINTER_BACKFILL_TENANTS is non-SmithDB-specific
and emitted unconditionally.

Values.yaml gains typed knobs under smithdb.langsmith.* and
config.runRules.twoPointerBackfillTenants for operator overrides.

* fix(langsmith): configure plaintext OTLP exporters (#859)

Set the standard OpenTelemetry insecure flag from useTls so Python exporters connect correctly to plaintext collectors while preserving TLS behavior.

* Bump langsmith appVersion to 0.16.24rc1 (#863)

Advance v16-stable chart defaults to the newly published 0.16.24rc1 images and bump the chart package to 0.16.0-rc.20.

Co-authored-by: Cursor <cursoragent@cursor.com>

* chore(langsmith): remove deprecated bundled agent-bootstrap; validate on use (#841)

Remove the bundled agent-bootstrap Job (backend.agentBootstrap) and its glue:
the bootstrap-job/rbac templates, its test, its values, the three
agentBootstrap.* helpers, and the -agent-builder-config configMap mounts it
created at runtime. Standalone agents (top-level fleet/insights/polly
deployments) replace it, so config.agentBuilder is retained (fleet uses it as
fallback defaults).

Add a validate.yaml gate: if backend.agentBootstrap is still set, fail with a
message directing users to contact LangChain support for migration guidance.

* fix(langsmith): set AUTH_ENDPOINT in the shared ConfigMap [v16-stable] (#846)

* fix(langsmith): set AUTH_ENDPOINT on the agent gateway

The gateway resolves frontend bearer tokens against the auth service and
reads AUTH_ENDPOINT, which the chart never set. Whenever the gateway
cannot build a local auth backend it panics at startup with:

  panic: gateway frontend bearer authentication requires AUTH_ENDPOINT
         for AUTH_TYPE "mixed"

Point AUTH_ENDPOINT at the platform-backend service, the same target
LANGSMITH_AUTH_ENDPOINT already uses in the shared ConfigMap. Scoped to
the gateway Deployment rather than the ConfigMap so the other services'
resolution of AUTH_ENDPOINT is unchanged.

Chart version goes to rc.19 rather than rc.18 so this branch and main do
not publish different charts under the same version.

* refactor(langsmith): move AUTH_ENDPOINT to the shared ConfigMap

AUTH_ENDPOINT was set as an inline env var on the agent gateway Deployment
only. Move it to the shared ConfigMap alongside GO_ENDPOINT and
LANGSMITH_AUTH_ENDPOINT, which already resolve to the same platform-backend
service.

The gateway already consumes the shared ConfigMap via envFrom, so it still
receives the value. Setting it centrally also covers every other consumer of
AUTH_ENDPOINT instead of the single Deployment we happened to patch.

Replace the gateway-scoped test with a ConfigMap suite covering the key's
presence across each authType branch, that it matches LANGSMITH_AUTH_ENDPOINT,
custom platformBackend name/port/clusterDomain, and the gateway's envFrom
reference that delivers it.

* Bump langsmith appVersion to 0.16.25rc1 (#865)

Advance v16-stable chart defaults to the newly published 0.16.25rc1
images and bump the chart package to 0.16.0-rc.22.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: align ace backend resources with SaaS defaults (#866) (#867)

Co-authored-by: Alex Kira <43946+akira@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat(langsmith): enable SMITHDB_UPGRADE_CRON_ENABLED by default (#868) (#869)

Set SMITHDB_UPGRADE_CRON_ENABLED=true into the shared LangSmith
ConfigMap when smithdb.enabled is true

(cherry picked from commit e17caa5)

* feat(langsmith): enable RULES_REDIS_TRACKING for SmithDB ingestion (#871)

Bakes RULES_REDIS_TRACKING_ENABLED=true and
RULES_REDIS_TRACKING_MULTIPART_ENABLED=true into the shared LangSmith
ConfigMap when smithdb.langsmith.ingestion.enabled is true — the "dual"
(ClickHouse + SmithDB) and "smithdb only" ingestion modes.

The Run Rules scheduler already reads from Redis (gated in #848 by
SMITHDB_RULES_REDIS_READ_ENABLED / _THREAD_ variants), but the ingestion
pipeline was not writing tracking entries there. These app-side flags
default to false, so we opt in from the chart whenever SmithDB is doing
the ingesting.

Placed in the existing `{{- if .Values.smithdb.langsmith.ingestion.enabled }}`
block alongside SMITHDB_MUTATIONS_SERVICE_URL / SMITHDB_INGESTION_SERVICE_URL,
so ClickHouse-only clusters remain on app code defaults (both false).

Unittest coverage: positive assertions in the dual-mode and smithdb-only
ingestion cases; notExists assertions in the disabled case.

Co-authored-by: Anirudh Satish <asatish@langchain.dev>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* fix(langsmith): set TaskDB resource defaults (#872)

Give the chart-managed migration TaskDB guaranteed CPU and memory so it does not run with BestEffort QoS.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix: cap Run Rules backfills at one day (#875)

Co-authored-by: ben <1986620+ben11211@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: add priorityClassName to all workloads (#876)

* feat: add priorityClassName to StatefulSets

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat: add priorityClassName to deployments

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* chore: align priorityClassName placement

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* ci: rerun transient helm repository check

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

---------

Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>

* feat(langsmith): set RUN_RULES_QUERY_LIMIT for SmithDB installs (#879)

Set RUN_RULES_QUERY_LIMIT=250 in the shared LangSmith ConfigMap when
smithdb.enabled is true, covering the "dual" (ClickHouse + SmithDB) and
"SmithDB only" self-hosted modes. Pure-ClickHouse installs keep the
application default.

---------

Co-authored-by: ben <nagengast@langchain.dev>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: joaquin-borggio-lc <joaquin@langchain.dev>
Co-authored-by: joaquin-borggio-lc <213688804+joaquin-borggio-lc@users.noreply.github.com>
Co-authored-by: open-swe[bot] <open-swe@users.noreply.github.com>
Co-authored-by: Anirudh Satish <asatish@langchain.dev>
Co-authored-by: Mukil Loganathan <mukil@langchain.dev>
Co-authored-by: Alex Kira <alex.kira@gmail.com>
Co-authored-by: Alex Kira <43946+akira@users.noreply.github.com>
Co-authored-by: ben <1986620+ben11211@users.noreply.github.com>
Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants