Skip to content

feat(setup): Anthropic OAuth onboarding with setup-token support - #384

Merged
ilblackdragon merged 18 commits into
mainfrom
feat/oauth-onboarding
Mar 7, 2026
Merged

ilblackdragon merged 18 commits into
mainfrom
feat/oauth-onboarding

Conversation

@serrrfirat

Copy link
Copy Markdown
Collaborator

Summary

  • Adds Anthropic OAuth support to the onboarding wizard (API key or OAuth token from claude login / claude setup-token)
  • New AnthropicOAuthProvider sends Authorization: Bearer + anthropic-beta: oauth-2025-04-20 (rig-core hardcodes x-api-key, so OAuth needs a custom HTTP provider)
  • Auto-detects existing OAuth tokens from OS credential stores (macOS Keychain, Linux ~/.claude/.credentials.json)
  • Deferred config resolution: providers return None instead of hard-erroring when credentials aren't available at startup, then re-resolve after secrets injection
  • All 4 provider backends (OpenAI, Anthropic, OpenAI-compatible, Tinfoil) defer gracefully

Motivation

Supersedes #113 and #143 which are stale (merge conflicts, based on old config.rs before module split). Key differences:

Feature #113 #143 This PR
Custom OAuth provider (Bearer auth) Yes (947 lines) No (sends as x-api-key → 401) Yes (580 lines)
anthropic-beta: oauth-2025-04-20 header Yes Wizard model-fetch only Yes (runtime)
OS credential store extraction No No Yes (Keychain + credentials.json)
Deferred config resolution No (hard-error) No (hard-error) Yes
Works without secrets DB / master key No No Yes
Model selection respects wizard choice No No Yes
Based on current codebase No (stale) No (stale) Yes
Provider resolution tests No No 20 tests

Credit to @LucaDeLeo (#113) for the native provider approach and AnthropicAuth enum pattern, and @bigguybobby (#143) for the wizard auth selector UX and token validation ideas.

Changes

  • src/llm/anthropic_oauth.rs (new) — Direct HTTP provider: Authorization: Bearer + anthropic-beta: oauth-2025-04-20 + anthropic-version: 2023-06-01. Follows the NearAiChatProvider pattern.
  • src/setup/wizard.rs — Anthropic auth offers "API Key" or "OAuth Token (from claude login)". Auto-detects tokens, retry flow, fallback to API key.
  • src/config/llm.rs — oauth_token field on AnthropicDirectConfig, deferred resolution for all providers, selected_model respected in resolution chain.
  • src/config/mod.rs — inject_os_credentials() loads Anthropic OAuth from OS stores even without secrets DB. Split from inject_llm_keys_from_secrets().
  • src/app.rs — OS credential injection + config re-resolution in no-master-key startup path.
  • src/llm/mod.rs — Routes to AnthropicOAuthProvider when oauth_token is present.
  • .env.example — Documents ANTHROPIC_OAUTH_TOKEN.

VPS / Headless usage

Set ANTHROPIC_OAUTH_TOKEN=<token> from claude setup-token (1-year validity). No browser needed on the server.

Test plan

  • cargo fmt — clean
  • cargo clippy --all --benches --tests --examples --all-features — zero warnings
  • cargo test — all pass
  • 20 provider resolution tests (all backends: deferred, API key, OAuth, selected_model)
  • 6 Anthropic OAuth provider tests (message conversion, system extraction, tool calls, response parsing)
  • Manual: wizard OAuth flow with existing claude login token → auto-detected, model selection works
  • Manual: claude-opus-4-6 via OAuth → successful chat
  • Manual: REPL, web gateway, Telegram all functional

Closes #113
Closes #143

Co-Authored-By: LucaDeLeo 46697542+LucaDeLeo@users.noreply.github.com
Co-Authored-By: bigguybobby 203908353+bigguybobby@users.noreply.github.com

🤖 Generated with Claude Code

serrrfirat and others added 11 commits February 26, 2026 12:38
Add OAuth token authentication as an alternative to API keys during
onboarding for both Anthropic (via `claude login`) and OpenAI/Codex
(via `~/.codex/auth.json`).

Key changes:
- New `AnthropicOAuthProvider` using `Authorization: Bearer` header
  (rig-core hardcodes `x-api-key` which rejects OAuth tokens)
- Wizard auth method selector: "Direct API Key" vs "OAuth Token"
  for both Anthropic and OpenAI providers
- Codex token extraction from `$CODEX_HOME/auth.json` / `~/.codex/auth.json`
- Claude Code sandbox sub-step in Docker setup (checks for credentials)
- Secret injection mappings for `ANTHROPIC_OAUTH_TOKEN` and `CODEX_OAUTH_TOKEN`
- `CODEX_OAUTH_TOKEN` falls back to `OPENAI_API_KEY` (same Bearer auth)

Supersedes #143 which had a broken auth flow (OAuth token sent as
x-api-key → 401). Credit to @bigguybobby for the original approach.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
OAuth tokens stored only in the secrets DB were invisible to
Config::from_env() which runs before the DB connects (chicken-and-egg).

Two fixes:
1. write_bootstrap_env() now persists ANTHROPIC_OAUTH_TOKEN and
   CODEX_OAUTH_TOKEN to ~/.ironclaw/.env (same pattern as NEARAI_API_KEY)
2. main.rs re-extracts a fresh token from the OS credential store
   (macOS Keychain / ~/.claude/.credentials.json) before config resolution,
   handling token expiry (8-12h) gracefully

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
All providers had the same chicken-and-egg issue: API keys stored in the
secrets DB were invisible to Config::from_env() which runs before DB
connects. Only NEARAI_API_KEY was written to bootstrap .env.

Now write_bootstrap_env() persists all credential env vars:
NEARAI_API_KEY, ANTHROPIC_API_KEY, ANTHROPIC_OAUTH_TOKEN, OPENAI_API_KEY,
CODEX_OAUTH_TOKEN, LLM_API_KEY, TINFOIL_API_KEY.

Also: setup_api_key_provider() now sets the env var during the wizard
session so write_bootstrap_env() can pick it up.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- Extract "oauth-placeholder" to named OAUTH_PLACEHOLDER constant shared
  across config and wizard to prevent silent drift
- Document plaintext credential tradeoff in write_bootstrap_env (API keys
  stored with 0o600 permissions, recommend full-disk encryption)
- Add blocking "Press Enter" wait in Anthropic OAuth retry flow so user
  has time to run `claude login` in another terminal
- Add escape hatch from manual OAuth paste back to API key flow (empty
  input switches to setup_api_key_provider)
- Fix Retry-After header: parse u64 seconds into Duration before passing
  to LlmError::RateLimited
- Make config::llm module pub(crate) for constant visibility
- Use .bearer_auth() instead of manual format!("Bearer {}")
- Remove response body from debug log (may contain PII)
- Update Anthropic API version to 2024-10-22

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Credentials (API keys, OAuth tokens) were being written in plaintext to
~/.ironclaw/.env to work around a chicken-and-egg problem: Config::from_env()
runs before the encrypted secrets DB is connected.

Instead of storing secrets on disk, LlmConfig::resolve() now defers
gracefully when credentials are missing — it returns None for the provider
config instead of hard-erroring with MissingRequired. After the DB connects,
AppBuilder::build_all() loads secrets from encrypted storage via
inject_llm_keys_from_secrets() and re-resolves the config.

For Anthropic OAuth tokens (which expire in 8-12h), the secret injection
step also tries the OS credential store (macOS Keychain / Linux
credentials.json) for a fresh token, overriding the potentially stale
copy in the DB.

Changes:
- LlmConfig::resolve(): OpenAI, Anthropic, OpenAI-compatible, and Tinfoil
  all return None instead of MissingRequired when credentials are absent
- write_bootstrap_env(): no longer writes any credential env vars
- inject_llm_keys_from_secrets(): refreshes Anthropic OAuth from OS
  credential store before overlay is finalized
- main.rs: removed OAuth re-extraction hack (no longer needed)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The OAuth token extraction from macOS Keychain / Linux credentials files
was only running inside inject_llm_keys_from_secrets(), which requires
the encrypted secrets DB. When no master key is configured, init_secrets()
returned early — skipping both DB secret loading AND OS credential store
extraction, leaving the Anthropic OAuth token unavailable.

Split into two paths:
- inject_llm_keys_from_secrets(): loads from encrypted DB + OS stores
- inject_os_credentials(): loads from OS stores only (no DB needed)

init_secrets() now calls inject_os_credentials() and re-resolves config
even in the no-master-key early-return path, so `claude login` tokens
are always available regardless of secrets DB state.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Anthropic's api.anthropic.com requires the `anthropic-beta: oauth-2025-04-20`
header to accept OAuth Bearer tokens. Without it, the API returns 401
"OAuth authentication is currently not supported."

Also reverts API version to 2023-06-01 since the OAuth beta flag does
not support the 2024-10-22 version (returns 400 "not a valid version").

This was the same bug that caused PR #143's 401 errors — the beta header
was missing entirely.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
The Anthropic and OpenAI config resolution ignored settings.selected_model
entirely, only checking the provider-specific env var (ANTHROPIC_MODEL,
OPENAI_MODEL) and falling back to a hardcoded default. This meant the
model chosen during onboarding wizard was silently overridden.

Now follows the same pattern as NearAI and OpenAI-compatible:
env var > settings.selected_model > hardcoded default.

Also deduplicated the Anthropic config construction (two identical
branches for API key vs OAuth now share model/base_url resolution).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Covers deferred resolution (no credentials → None instead of error),
credential presence, model selection fallback chain, and OAuth token
routing for Anthropic, OpenAI, Tinfoil, Ollama, and NearAI.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Codex CLI stores OAuth tokens in a nested format under
tokens.access_token (ChatGPT OAuth flow), not at the top level.
Also adds ENV_MUTEX to Codex token tests for thread safety.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Codex CLI OAuth tokens use a different endpoint
(chatgpt.com/backend-api/codex) and the Responses API wire format,
not api.openai.com with Chat Completions. The tokens lack the
model.request scope needed for the platform API, so they can't be
used as drop-in OPENAI_API_KEY replacements.

Removes: extract_codex_oauth_token(), wizard Codex OAuth flow,
CODEX_OAUTH_TOKEN env var support, and related tests.

OpenAI onboarding now uses direct API key only.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@github-actions github-actions Bot added scope: llm LLM integration scope: config Configuration scope: setup Onboarding / setup size: XL 500+ changed lines risk: high Safety, secrets, auth, or critical infrastructure contributor: experienced 6-19 merged PRs labels Feb 26, 2026
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello @serrrfirat, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly enhances the authentication and configuration mechanisms for Large Language Model (LLM) providers, particularly for Anthropic. It introduces a more flexible and robust way to handle Anthropic credentials by supporting OAuth tokens and integrating with OS-level credential stores. Furthermore, it refines the application's startup process by implementing deferred configuration resolution, ensuring that LLM providers can be configured even when secrets are loaded asynchronously. These changes aim to streamline the setup experience and improve the reliability of LLM integrations.

Highlights

  • Anthropic OAuth Onboarding: Added comprehensive support for Anthropic OAuth tokens in the onboarding wizard, allowing users to authenticate via API key or OAuth token obtained from claude login or claude setup-token.
  • Dedicated Anthropic OAuth Provider: Introduced a new AnthropicOAuthProvider to handle Authorization: Bearer authentication and the anthropic-beta: oauth-2025-04-20 header, as the existing rig-core client only supports x-api-key.
  • OS Credential Store Integration: Implemented auto-detection and extraction of existing Anthropic OAuth tokens from OS credential stores (macOS Keychain, Linux ~/.claude/.credentials.json), improving user experience and token freshness.
  • Deferred LLM Config Resolution: Enhanced LLM provider configuration to gracefully defer resolution when credentials are not immediately available at startup, allowing re-resolution after secrets injection and preventing hard errors.
  • Claude Code Sandbox Support: Integrated Claude Code sandbox mode, enabling the agent to delegate tasks to Claude CLI within sandboxed Docker containers, leveraging Anthropic credentials.
Changelog
  • .env.example
    • Documented ANTHROPIC_OAUTH_TOKEN for OAuth authentication.
  • src/app.rs
    • Added logic to inject OS credentials and re-resolve the configuration when no master key is available, ensuring tokens from OS stores are utilized.
  • src/config/llm.rs
    • Introduced OAUTH_PLACEHOLDER sentinel value for Anthropic API keys when only an OAuth token is present.
    • Added oauth_token field to AnthropicDirectConfig to store OAuth tokens.
    • Modified LlmConfig::resolve to implement deferred resolution for OpenAI, Anthropic, OpenAI-compatible, and Tinfoil providers, returning None if credentials are not yet available.
    • Added new test cases for Anthropic, OpenAI, and Tinfoil provider resolution, including scenarios for deferred resolution and OAuth token handling.
  • src/config/mod.rs
    • Made the llm module public to allow external access.
    • Included llm_anthropic_oauth_token in the secrets mapping for inject_llm_keys_from_secrets.
    • Added inject_os_credentials and inject_os_credential_store_tokens functions to load tokens from OS credential stores independently of the secrets database.
  • src/llm/anthropic_oauth.rs
    • Added a new module defining AnthropicOAuthProvider for direct HTTP communication with Anthropic using Bearer tokens and specific beta headers.
    • Implemented LlmProvider trait for AnthropicOAuthProvider, including methods for complete and complete_with_tools.
    • Included message conversion logic to adapt internal ChatMessage format to Anthropic's API requirements, handling system messages and tool calls/results.
    • Added unit tests for message conversion and response content extraction.
  • src/llm/mod.rs
    • Imported the new anthropic_oauth module.
    • Modified create_anthropic_provider to route to AnthropicOAuthProvider if an OAuth token is present in the configuration.
  • src/main.rs
    • Updated comments to clarify that initial config loading can occur before the DB is available and that LlmConfig::resolve() defers gracefully.
  • src/settings.rs
    • Added claude_code_enabled boolean field to SandboxSettings to control Claude Code sandbox mode.
  • src/setup/wizard.rs
    • Modified setup_anthropic to present users with a choice between API Key and OAuth Token authentication.
    • Introduced setup_anthropic_oauth to guide users through OAuth token setup, including auto-detection from OS stores, retry mechanisms, and manual input fallback.
    • Added save_anthropic_oauth_token to store OAuth tokens securely and set environment variables.
    • Updated setup_api_key_provider to set environment variables for immediate use during the wizard.
    • Added step_claude_code_sandbox to enable and configure Claude Code sandbox mode, checking for Anthropic credentials.
    • Updated write_bootstrap_env to include CLAUDE_CODE_ENABLED if enabled and clarified that credentials are not written to the bootstrap environment.
Activity
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces significant enhancements for Anthropic integration, including OAuth support, deferred configuration resolution, and an improved setup wizard. My review focuses on improving error handling consistency by propagating potential ConfigErrors, increasing robustness against unexpected inputs by explicitly handling None cases, and refactoring for better maintainability by consolidating repeated logic into reusable methods, all in line with repository guidelines.

Comment thread src/config/llm.rs Outdated
Comment on lines +305 to +310
.flatten()
.or_else(|| settings.selected_model.clone())
.unwrap_or_else(|| "claude-sonnet-4-20250514".to_string());
let base_url = optional_env("ANTHROPIC_BASE_URL").ok().flatten();
AnthropicDirectConfig {
api_key,

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The error handling for ANTHROPIC_MODEL and ANTHROPIC_BASE_URL is inconsistent with other providers like OpenAI. Using .ok().flatten() silently ignores potential ConfigErrors (e.g., for invalid UTF-8 in environment variables). It would be more robust and consistent to propagate these errors using the ? operator, similar to how OPENAI_MODEL is handled. This aligns with the practice of using specialized helper functions for environment variable resolution to ensure consistency and proper error handling.

Suggested change
.flatten()
.or_else(|| settings.selected_model.clone())
.unwrap_or_else(|| "claude-sonnet-4-20250514".to_string());
let base_url = optional_env("ANTHROPIC_BASE_URL").ok().flatten();
AnthropicDirectConfig {
api_key,
let model = optional_env("ANTHROPIC_MODEL")?
.or_else(|| settings.selected_model.clone())
.unwrap_or_else(|| "claude-sonnet-4-20250514".to_string());
let base_url = optional_env("ANTHROPIC_BASE_URL")?;
References
  1. Use specialized helper functions like parse_option_env<T> when resolving environment variables into Option<T> fields to minimize boilerplate and maintain consistency across configuration files.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Already resolved — the code now uses the registry-based resolve_registry_provider() which calls optional_env() with ? propagation throughout.

Comment on lines +427 to +448
Role::Tool => {
// Tool results go into a user message with tool_result blocks
let block = AnthropicContentBlock::ToolResult {
tool_use_id: msg.tool_call_id.unwrap_or_default(),
content: msg.content,
};
// If the last message is already a user message with blocks,
// append to it (Anthropic requires consecutive tool results
// in one user message).
if let Some(last) = anthropic_msgs.last_mut()
&& last.role == "user"
&& let AnthropicContent::Blocks(ref mut blocks) = last.content
{
blocks.push(block);
continue;
}
anthropic_msgs.push(AnthropicMessage {
role: "user".to_string(),
content: AnthropicContent::Blocks(vec![block]),
});
}
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Using unwrap_or_default() on tool_call_id could result in an empty string for tool_use_id if a ChatMessage with Role::Tool is ever constructed without a tool_call_id. This would likely cause an API error from Anthropic. It's safer to handle the None case explicitly, for example by skipping the message and logging a warning.

            Role::Tool => {
                let Some(tool_call_id) = msg.tool_call_id else {
                    tracing::warn!("Skipping Tool message without tool_call_id");
                    continue;
                };
                // Tool results go into a user message with tool_result blocks
                let block = AnthropicContentBlock::ToolResult {
                    tool_use_id: tool_call_id,
                    content: msg.content,
                };
                // If the last message is already a user message with blocks,
                // append to it (Anthropic requires consecutive tool results
                // in one user message).
                if let Some(last) = anthropic_msgs.last_mut()
                    && last.role == "user"
                    && let AnthropicContent::Blocks(ref mut blocks) = last.content
                {
                    blocks.push(block);
                    continue;
                }
                anthropic_msgs.push(AnthropicMessage {
                    role: "user".to_string(),
                    content: AnthropicContent::Blocks(vec![block]),
                });
            }

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Already resolved — the code already uses let Some(tool_call_id) = msg.tool_call_id else { continue } with a warning log, exactly as suggested.

Comment thread src/setup/wizard.rs Outdated
Comment on lines +1992 to +2025
let has_api_key = std::env::var("ANTHROPIC_API_KEY")
.is_ok_and(|v| !v.is_empty() && v != OAUTH_PLACEHOLDER);
let has_oauth = crate::config::ClaudeCodeConfig::extract_oauth_token().is_some()
|| std::env::var("ANTHROPIC_OAUTH_TOKEN").is_ok_and(|v| !v.is_empty());

if has_api_key || has_oauth {
self.settings.sandbox.claude_code_enabled = true;
print_success("Claude Code sandbox enabled");
} else {
print_error("No Anthropic credentials found.");
print_info(
"Claude Code needs ANTHROPIC_API_KEY or an OAuth token from `claude login`.",
);
println!();

if confirm("Retry after setting up credentials?", false).map_err(SetupError::Io)? {
let has_oauth_retry =
crate::config::ClaudeCodeConfig::extract_oauth_token().is_some();
let has_key_retry = std::env::var("ANTHROPIC_API_KEY")
.is_ok_and(|v| !v.is_empty() && v != OAUTH_PLACEHOLDER);

if has_key_retry || has_oauth_retry {
self.settings.sandbox.claude_code_enabled = true;
print_success("Claude Code sandbox enabled");
} else {
self.settings.sandbox.claude_code_enabled = false;
print_info("No credentials found. Claude Code disabled for now.");
print_info("Set ANTHROPIC_API_KEY or run `claude login` and enable later.");
}
} else {
self.settings.sandbox.claude_code_enabled = false;
print_info("Claude Code disabled. Enable with CLAUDE_CODE_ENABLED=true later.");
}
}

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The logic for checking for Anthropic credentials is repeated. To improve maintainability and reduce code duplication, you can define a closure to encapsulate this check, aligning with the practice of consolidating related operations into reusable methods.

        let has_credentials = || {
            let has_api_key = std::env::var("ANTHROPIC_API_KEY")
                .is_ok_and(|v| !v.is_empty() && v != OAUTH_PLACEHOLDER);
            let has_oauth = crate::config::ClaudeCodeConfig::extract_oauth_token().is_some()
                || std::env::var("ANTHROPIC_OAUTH_TOKEN").is_ok_and(|v| !v.is_empty());
            has_api_key || has_oauth
        };

        if has_credentials() {
            self.settings.sandbox.claude_code_enabled = true;
            print_success("Claude Code sandbox enabled");
        } else {
            print_error("No Anthropic credentials found.");
            print_info(
                "Claude Code needs ANTHROPIC_API_KEY or an OAuth token from `claude login`.",
            );
            println!();

            if confirm("Retry after setting up credentials?", false).map_err(SetupError::Io)? {
                if has_credentials() {
                    self.settings.sandbox.claude_code_enabled = true;
                    print_success("Claude Code sandbox enabled");
                } else {
                    self.settings.sandbox.claude_code_enabled = false;
                    print_info("No credentials found. Claude Code disabled for now.");
                    print_info("Set ANTHROPIC_API_KEY or run `claude login` and enable later.");
                }
            } else {
                self.settings.sandbox.claude_code_enabled = false;
                print_info("Claude Code disabled. Enable with CLAUDE_CODE_ENABLED=true later.");
            }
        }
References
  1. Consolidate related sequences of operations, such as creating, persisting, and scheduling a job, into a single reusable method to improve code consistency and maintainability.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Already resolved — the code already uses a has_credentials closure as suggested. Additionally, in 898e841 the closure was updated to use optional_env() instead of raw std::env::var.

- Use ? operator for ANTHROPIC_MODEL/BASE_URL env resolution instead of
  .ok().flatten() to propagate ConfigErrors consistently
- Skip Tool messages without tool_call_id with a warning instead of
  using unwrap_or_default() which would send empty string to Anthropic
- Extract credential check into closure to reduce duplication in
  Claude Code sandbox setup

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
zmanian
zmanian previously requested changes Mar 1, 2026

@zmanian zmanian left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: Anthropic OAuth Onboarding with Setup-Token Support

Good feature addition -- the AnthropicOAuthProvider is well-implemented, the OS credential extraction is sensible, and the deferred resolution approach is sound. Test coverage for config resolution is solid. However, there are several issues.

Blocker: OnceLock Race Between Credential Injection Paths

Both inject_os_credentials() and inject_llm_keys_from_secrets() call INJECTED_VARS.set(injected), but INJECTED_VARS is a OnceLock<HashMap<...>> -- whichever runs first wins; the second call silently drops its data.

In AppBuilder::init_secrets(), when master_key() returns None, inject_os_credentials() is called (setting INJECTED_VARS), then it returns early. If the code paths ever change ordering, credentials will silently vanish.

Fix: Replace OnceLock with a Mutex<HashMap> that supports merge semantics, or ensure at the type level that only one injection path can be called.

High

  1. No token expiry/refresh: OAuth tokens from claude login expire in ~8-12 hours. The AnthropicOAuthProvider stores the token once and never refreshes it. For IronClaw (a daemon), this means guaranteed auth failures after ~8 hours with no recovery path. At minimum, on 401 errors attempt to re-extract from the OS credential store. Ideally, use the refresh token. Can be a follow-up, but document the limitation.

  2. Comment/code path mismatch: Comment says ~/.codex/auth.json but code reads ~/.claude/.credentials.json. Fix the comment.

  3. OAUTH_PLACEHOLDER sentinel value: Storing "oauth-placeholder" in api_key when only an OAuth token is present is fragile. Any new code that reads config.anthropic.api_key without being aware of the sentinel will silently use an invalid key. Consider using Option<SecretString> for api_key instead.

Medium

  1. Deferred resolution changes error behavior for ALL providers: Previously, LLM_BACKEND=openai without OPENAI_API_KEY produced a clear startup error. Now it silently defers and fails at runtime with a less informative error. Add a post-init validation step that checks "if backend == X and config.X.is_none(), emit a clear error."

  2. unsafe { std::env::set_var } in wizard code: The comment says "single-threaded wizard context" but the wizard is async and Tokio's runtime is multi-threaded. Use the INJECTED_VARS overlay mechanism instead of set_var.

  3. No validation of extracted OAuth tokens: parse_oauth_access_token accepts any string. Add at least a prefix check (sk-ant-oat01-) and non-empty validation. Optionally check expiresAt.

Low

  1. claude_code_enabled in SandboxSettings: Somewhat tangential to OAuth onboarding -- increases review surface but there's a connection.

What's Good

  • No new dependencies added
  • Token stored in memory only (never written to .env)
  • Uses mask_api_key() for display -- no token leakage in logs
  • Clean AnthropicOAuthProvider with proper Authorization: Bearer header
  • 10+ unit tests for config resolution scenarios
  • No database schema changes needed -- both backends unaffected

Recommendation

Fix the OnceLock race condition before merge. Document the token expiry limitation and file an issue for follow-up refresh support. The other items can be addressed in follow-up PRs.

…try)

Adapt OAuth onboarding features to work with the new RegistryProviderConfig
architecture from main. Key changes:
- Add oauth_token field to RegistryProviderConfig
- Update AnthropicOAuthProvider to accept RegistryProviderConfig
- Route Anthropic OAuth in create_anthropic_from_registry
- Add setup_anthropic intercept in wizard's run_provider_setup
- Include ANTHROPIC_OAUTH_TOKEN in dynamic secret injection
- Fix pre-existing type error in wizard (CLAUDE_CODE_ENABLED)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@ilblackdragon
ilblackdragon force-pushed the feat/oauth-onboarding branch from 6652c20 to 0760268 Compare March 7, 2026 09:01
- Gate ANTHROPIC_OAUTH_TOKEN resolution to Anthropic provider only
  (was needlessly checked for all registry providers)
- Add 3 regression tests for OAuth config resolution:
  - oauth_token sets placeholder api_key
  - real api_key takes priority over oauth
  - non-Anthropic providers don't pick up oauth_token
- Validate OAuth token prefix (sk-ant-oat) in wizard to catch
  accidentally pasted API keys
- Improve error body read handling in AnthropicOAuthProvider
  (was silently swallowing read errors with unwrap_or_default)
- Remove extra blank line in write_bootstrap_env
- Remove stale blank line in RegistryProviderConfig doc comment

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@ilblackdragon
ilblackdragon force-pushed the feat/oauth-onboarding branch from 0760268 to 14ff926 Compare March 7, 2026 09:02

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds first-class Anthropic OAuth support (Claude CLI tokens) across config resolution, onboarding, and runtime provider creation, including OS credential-store token extraction for headless/no-secrets-DB startup paths.

Changes:

  • Introduces a direct HTTP Anthropic OAuth provider using Authorization: Bearer + required anthropic-beta header.
  • Extends the setup wizard to offer Anthropic API key vs OAuth token onboarding, plus an optional “Claude Code sandbox” toggle.
  • Adds OAuth token plumbing to config resolution and startup credential injection (secrets DB + OS credential stores), with updated env docs.

Reviewed changes

Copilot reviewed 9 out of 9 changed files in this pull request and generated 6 comments.

Show a summary per file
File Description
src/setup/wizard.rs Adds Anthropic OAuth onboarding flow, env/secrets persistence, and Claude Code sandbox sub-step.
src/settings.rs Persists a claude_code_enabled toggle in sandbox settings.
src/main.rs Clarifies that LLM config may defer until secrets are injected and config is re-resolved.
src/llm/mod.rs Routes Anthropic provider creation to the OAuth provider when an OAuth token is present.
src/llm/anthropic_oauth.rs New direct HTTP provider implementing Anthropic Messages API with Bearer auth.
src/config/mod.rs Injects Anthropic OAuth from secrets + OS credential stores into the config overlay.
src/config/llm.rs Adds oauth_token to RegistryProviderConfig and an OAUTH_PLACEHOLDER sentinel.
src/app.rs Injects OS credentials and re-resolves config when no secrets master key is available.
.env.example Documents ANTHROPIC_OAUTH_TOKEN usage.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread src/setup/wizard.rs
Comment on lines +1100 to +1102
// Cache for model fetching
self.llm_api_key = Some(SecretString::from(token.to_string()));

Copilot AI Mar 7, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

self.llm_api_key is reused here to store an OAuth token, but later fetch_anthropic_models() sends the cached value as x-api-key (not Authorization: Bearer + anthropic-beta), so model listing will always fail for OAuth and fall back to static defaults (after a 401). Consider either (a) teaching the model-fetch path to use the OAuth headers when ANTHROPIC_OAUTH_TOKEN is configured, or (b) storing OAuth separately from llm_api_key so the callsite can choose the right auth scheme.

Suggested change
// Cache for model fetching
self.llm_api_key = Some(SecretString::from(token.to_string()));

Copilot uses AI. Check for mistakes.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 898e841. fetch_anthropic_models() now detects OAuth tokens (via optional_env) and uses Bearer auth + anthropic-beta header. Also filters out OAUTH_PLACEHOLDER from the API key lookup.

Comment thread src/llm/mod.rs Outdated
Comment on lines +181 to +191
// Route to OAuth provider when an OAuth token is present
if config.oauth_token.is_some() {
tracing::info!(
provider = %config.provider_id,
model = %config.model,
base_url = if config.base_url.is_empty() { "default" } else { &config.base_url },
"Using Anthropic OAuth API"
);
let provider = anthropic_oauth::AnthropicOAuthProvider::new(config)?;
return Ok(Arc::new(provider));
}

Copilot AI Mar 7, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

create_anthropic_from_registry() routes to AnthropicOAuthProvider whenever oauth_token is set, which will also trigger when an API key is present (your LlmConfig::resolve test keeps oauth_token populated even when ANTHROPIC_API_KEY is set). This contradicts the intended “API key takes priority” behavior and will cause API-key users with ANTHROPIC_OAUTH_TOKEN set to unexpectedly use Bearer auth. Consider routing to OAuth only when api_key is missing or equals OAUTH_PLACEHOLDER, or clearing oauth_token when a real API key is present.

Copilot uses AI. Check for mistakes.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 898e841. OAuth routing now only triggers when api_key is None or equals OAUTH_PLACEHOLDER. When both an API key and OAuth token are set, the API key takes priority.

Comment thread src/setup/wizard.rs Outdated
Comment on lines +1094 to +1100
// Set env var for immediate use (model selection step)
// SAFETY: Single-threaded wizard context.
unsafe {
std::env::set_var("ANTHROPIC_OAUTH_TOKEN", token);
}

// Cache for model fetching

Copilot AI Mar 7, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

unsafe { std::env::set_var(...) } is called while the program is running on a multi-threaded Tokio runtime (see main.rs uses new_multi_thread()), so the “single-threaded wizard context” safety comment doesn’t hold and this can be UB if any other thread reads env concurrently. Instead of mutating process env, consider passing the key via self.llm_api_key (already used by model listing) or using the INJECTED_VARS overlay (or a wizard-local overlay) that optional_env() reads.

Suggested change
// Set env var for immediate use (model selection step)
// SAFETY: Single-threaded wizard context.
unsafe {
std::env::set_var("ANTHROPIC_OAUTH_TOKEN", token);
}
// Cache for model fetching
// Cache for model fetching / immediate use during model selection

Copilot uses AI. Check for mistakes.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 898e841. Replaced unsafe { std::env::set_var } with crate::config::inject_single_var(), which merges into the thread-safe INJECTED_VARS overlay. No more UB risk on multi-threaded runtimes.

Comment thread src/setup/wizard.rs Outdated
Comment on lines 1168 to 1174
// Set env var so subsequent wizard steps (model selection) can use it
// immediately. This is ephemeral — credentials are NOT written to the
// bootstrap .env; they live only in the encrypted secrets DB.
// SAFETY: Single-threaded wizard context.
unsafe { std::env::set_var(env_var, key_str) };

// Cache key in memory for model fetching later in the wizard

Copilot AI Mar 7, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Same concern as above: setting env vars from inside the async wizard requires unsafe and is not sound under the multi-threaded runtime. Since later wizard steps already consult self.llm_api_key (and could consult optional_env() for the injected overlay), it would be safer to remove this env mutation and plumb credentials through wizard state instead.

Suggested change
// Set env var so subsequent wizard steps (model selection) can use it
// immediately. This is ephemeral — credentials are NOT written to the
// bootstrap .env; they live only in the encrypted secrets DB.
// SAFETY: Single-threaded wizard context.
unsafe { std::env::set_var(env_var, key_str) };
// Cache key in memory for model fetching later in the wizard
// Cache key in memory so subsequent wizard steps (e.g. model selection)
// can use it without mutating process-wide environment variables.

Copilot uses AI. Check for mistakes.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 898e841. Same fix — both save_anthropic_oauth_token and setup_api_key_provider now use inject_single_var() instead of unsafe { std::env::set_var }.

Comment thread src/setup/wizard.rs
Comment on lines +2131 to +2137
let has_credentials = || {
let has_api_key = std::env::var("ANTHROPIC_API_KEY")
.is_ok_and(|v| !v.is_empty() && v != OAUTH_PLACEHOLDER);
let has_oauth = crate::config::ClaudeCodeConfig::extract_oauth_token().is_some()
|| std::env::var("ANTHROPIC_OAUTH_TOKEN").is_ok_and(|v| !v.is_empty());
has_api_key || has_oauth
};

Copilot AI Mar 7, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The Anthropic credential check here only looks at raw process env vars and ClaudeCodeConfig::extract_oauth_token(). If credentials are present only via the injected secrets overlay (inject_llm_keys_from_secrets + optional_env) or via wizard state (self.llm_api_key), this can incorrectly report “No Anthropic credentials found.” Consider using crate::config::helpers::optional_env() for env lookups and/or checking the wizard’s cached key/token state.

Copilot uses AI. Check for mistakes.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 898e841. Credential check now uses optional_env() which reads both real env vars and the injected overlay.

Comment thread src/app.rs Outdated
Some(k) => k,
None => {
// No secrets DB available, but we can still load tokens from
// OS credential stores (macOS Keychain, ~/.codex/auth.json).

Copilot AI Mar 7, 2026

Copy link

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This comment mentions loading tokens from ~/.codex/auth.json, but inject_os_credentials() currently only injects Anthropic OAuth from Claude Code’s credential store. Either expand OS credential injection to include Codex (if intended) or update the comment to match current behavior.

Suggested change
// OS credential stores (macOS Keychain, ~/.codex/auth.json).
// OS/IDE credential stores (e.g., Anthropic OAuth via Claude Code's credential store).

Copilot uses AI. Check for mistakes.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 898e841. Comment now accurately says "macOS Keychain / Linux ~/.claude/.credentials.json".

ilblackdragon and others added 2 commits March 7, 2026 01:30
…infra)

Resolve conflicts between OAuth onboarding and cache retention features:
- Keep both OAuth routing and CacheRetention imports in create_anthropic_from_registry
- Add cache_creation_input_tokens/cache_read_input_tokens to AnthropicOAuthProvider responses
- Keep both OAuth and CacheRetention test suites in config/llm.rs

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Blocker:
- Replace OnceLock<HashMap> with LazyLock<Mutex<HashMap>> for INJECTED_VARS
  so both inject_os_credentials() and inject_llm_keys_from_secrets() merge
  data instead of the second caller silently dropping its entries.

High:
- Add 401 retry with OS credential store re-extraction in
  AnthropicOAuthProvider, recovering from expired OAuth tokens (~8-12h)
  without manual intervention.
- Fix comment in app.rs: ~/.codex/auth.json → ~/.claude/.credentials.json.

Medium:
- Remove unsafe { std::env::set_var } from wizard; use thread-safe
  inject_single_var() overlay instead (safe on multi-threaded Tokio).
- Add post-init validation in AppBuilder: fail early with clear error when
  LLM_BACKEND is set but no credentials were resolved after secret injection.
- Add sk-ant-oat prefix validation in parse_oauth_access_token().
- Only route to AnthropicOAuthProvider when api_key is missing or equals
  OAUTH_PLACEHOLDER (API key takes priority over OAuth token).
- Teach fetch_anthropic_models() to use Bearer auth when only OAuth token
  is available (model listing no longer fails for OAuth-only users).

Low:
- Use optional_env() in wizard credential checks to read from injected
  overlay, not just raw env vars.

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@ilblackdragon

Copy link
Copy Markdown
Member

Review Fixes (898e841)

All items from @zmanian's review and bot comments have been addressed:

Blocker ✅

  • OnceLock race: Replaced OnceLock<HashMap> with LazyLock<Mutex<HashMap>>. Both inject_os_credentials() and inject_llm_keys_from_secrets() now merge data via merge_injected_vars() instead of the second caller silently dropping its entries.

High ✅

  1. Token expiry/refresh: AnthropicOAuthProvider now attempts to re-extract a fresh token from the OS credential store (macOS Keychain / Linux credentials file) on 401 errors and retries once. Full refresh token support can be a follow-up.
  2. Comment mismatch: Fixed ~/.codex/auth.json → ~/.claude/.credentials.json in app.rs.
  3. OAUTH_PLACEHOLDER: Kept by design (documented in code). The sentinel only exists so RegistryProviderConfig is Some; the factory always checks for it before use. Added routing guard so API key always takes priority when both are set.

Medium ✅

  1. Deferred resolution error: Added post-init validation in AppBuilder::build_all() — fails early with a clear error when LLM_BACKEND is set but no credentials were resolved.
  2. unsafe set_var: Replaced both instances with inject_single_var() which merges into the thread-safe INJECTED_VARS overlay. No more UB on multi-threaded runtimes.
  3. Token validation: Added sk-ant-oat prefix check in parse_oauth_access_token() (was already validated in save_anthropic_oauth_token).
  4. OAuth routing priority: Now only routes to AnthropicOAuthProvider when api_key is None or equals OAUTH_PLACEHOLDER.
  5. Model listing with OAuth: fetch_anthropic_models() now uses Bearer auth + anthropic-beta header when only an OAuth token is available.

Low ✅

  1. Credential check overlay: step_claude_code_sandbox() now uses optional_env() instead of raw std::env::var.

Already resolved (bot comments)

  • Gemini: .ok().flatten() inconsistency — code uses registry-based resolution with ? propagation
  • Gemini: unwrap_or_default() on tool_call_id — already uses let Some(...) else { continue }
  • Gemini: credential check duplication — already using has_credentials closure

CI: 2476 tests pass, zero clippy warnings.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@ilblackdragon
ilblackdragon merged commit 11c5e25 into main Mar 7, 2026
22 checks passed
@ilblackdragon
ilblackdragon deleted the feat/oauth-onboarding branch March 7, 2026 20:59
@github-actions github-actions Bot mentioned this pull request Mar 7, 2026
ALuhning pushed a commit to VitalPointAI/Bastion that referenced this pull request Mar 8, 2026
Switch ironclaw from API key billing to Anthropic OAuth (subscription-based)
by building from source to include PR nearai/ironclaw#384. Add global chat
mode so ironclaw is usable anywhere in the app, not just inside problem sets.

- Dockerfile: multi-stage build from source (main branch) for OAuth support
- docker-compose: add ANTHROPIC_OAUTH_TOKEN env var (dev + prod)
- Backend: add POST /global/message and GET /global/history routes
  scoped per-user via DID-derived thread IDs
- Frontend: enable chat input in global mode with mode indicator,
  dynamic WebSocket channel subscription from API response
- .env.example: document ironclaw config section

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
ALuhning pushed a commit to VitalPointAI/Bastion that referenced this pull request Mar 8, 2026
* fix(validation): build complete BastionState in executeAgentCall

The validation runner was constructing an incomplete BastionState missing
required fields (objectives, taskType, invocationCount, etc.) and passing
plain {role, content} objects instead of LangChain BaseMessage instances.
This caused buildAgentSystemPrompt to throw on state.objectives.join(),
silently falling back to [Simulated] responses that score near-zero on
every metric — triggering circuit breaker disables for ALL agents.

The bug was exposed by the fixture-loading fallback (cb0a447) which made
validation tests actually run for the first time.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat(ironclaw): Anthropic OAuth auth and global chat mode

Switch ironclaw from API key billing to Anthropic OAuth (subscription-based)
by building from source to include PR nearai/ironclaw#384. Add global chat
mode so ironclaw is usable anywhere in the app, not just inside problem sets.

- Dockerfile: multi-stage build from source (main branch) for OAuth support
- docker-compose: add ANTHROPIC_OAUTH_TOKEN env var (dev + prod)
- Backend: add POST /global/message and GET /global/history routes
  scoped per-user via DID-derived thread IDs
- Frontend: enable chat input in global mode with mode indicator,
  dynamic WebSocket channel subscription from API response
- .env.example: document ironclaw config section

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: jim-agent <259877510+jim-agent@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
@github-actions github-actions Bot mentioned this pull request Mar 10, 2026
bkutasi pushed a commit to bkutasi/ironclaw that referenced this pull request Mar 28, 2026
…rai#384)

* feat(setup): add Anthropic OAuth and Codex OAuth onboarding flows

Add OAuth token authentication as an alternative to API keys during
onboarding for both Anthropic (via `claude login`) and OpenAI/Codex
(via `~/.codex/auth.json`).

Key changes:
- New `AnthropicOAuthProvider` using `Authorization: Bearer` header
  (rig-core hardcodes `x-api-key` which rejects OAuth tokens)
- Wizard auth method selector: "Direct API Key" vs "OAuth Token"
  for both Anthropic and OpenAI providers
- Codex token extraction from `$CODEX_HOME/auth.json` / `~/.codex/auth.json`
- Claude Code sandbox sub-step in Docker setup (checks for credentials)
- Secret injection mappings for `ANTHROPIC_OAUTH_TOKEN` and `CODEX_OAUTH_TOKEN`
- `CODEX_OAUTH_TOKEN` falls back to `OPENAI_API_KEY` (same Bearer auth)

Supersedes nearai#143 which had a broken auth flow (OAuth token sent as
x-api-key → 401). Credit to @bigguybobby for the original approach.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: persist OAuth tokens in bootstrap .env and re-extract at startup

OAuth tokens stored only in the secrets DB were invisible to
Config::from_env() which runs before the DB connects (chicken-and-egg).

Two fixes:
1. write_bootstrap_env() now persists ANTHROPIC_OAUTH_TOKEN and
   CODEX_OAUTH_TOKEN to ~/.ironclaw/.env (same pattern as NEARAI_API_KEY)
2. main.rs re-extracts a fresh token from the OS credential store
   (macOS Keychain / ~/.claude/.credentials.json) before config resolution,
   handling token expiry (8-12h) gracefully

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: persist all LLM credentials in bootstrap .env, not just NEAR AI

All providers had the same chicken-and-egg issue: API keys stored in the
secrets DB were invisible to Config::from_env() which runs before DB
connects. Only NEARAI_API_KEY was written to bootstrap .env.

Now write_bootstrap_env() persists all credential env vars:
NEARAI_API_KEY, ANTHROPIC_API_KEY, ANTHROPIC_OAUTH_TOKEN, OPENAI_API_KEY,
CODEX_OAUTH_TOKEN, LLM_API_KEY, TINFOIL_API_KEY.

Also: setup_api_key_provider() now sets the env var during the wizard
session so write_bootstrap_env() can pick it up.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: address security review findings for OAuth onboarding

- Extract "oauth-placeholder" to named OAUTH_PLACEHOLDER constant shared
  across config and wizard to prevent silent drift
- Document plaintext credential tradeoff in write_bootstrap_env (API keys
  stored with 0o600 permissions, recommend full-disk encryption)
- Add blocking "Press Enter" wait in Anthropic OAuth retry flow so user
  has time to run `claude login` in another terminal
- Add escape hatch from manual OAuth paste back to API key flow (empty
  input switches to setup_api_key_provider)
- Fix Retry-After header: parse u64 seconds into Duration before passing
  to LlmError::RateLimited
- Make config::llm module pub(crate) for constant visibility
- Use .bearer_auth() instead of manual format!("Bearer {}")
- Remove response body from debug log (may contain PII)
- Update Anthropic API version to 2024-10-22

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* security: remove plaintext credentials from bootstrap .env

Credentials (API keys, OAuth tokens) were being written in plaintext to
~/.ironclaw/.env to work around a chicken-and-egg problem: Config::from_env()
runs before the encrypted secrets DB is connected.

Instead of storing secrets on disk, LlmConfig::resolve() now defers
gracefully when credentials are missing — it returns None for the provider
config instead of hard-erroring with MissingRequired. After the DB connects,
AppBuilder::build_all() loads secrets from encrypted storage via
inject_llm_keys_from_secrets() and re-resolves the config.

For Anthropic OAuth tokens (which expire in 8-12h), the secret injection
step also tries the OS credential store (macOS Keychain / Linux
credentials.json) for a fresh token, overriding the potentially stale
copy in the DB.

Changes:
- LlmConfig::resolve(): OpenAI, Anthropic, OpenAI-compatible, and Tinfoil
  all return None instead of MissingRequired when credentials are absent
- write_bootstrap_env(): no longer writes any credential env vars
- inject_llm_keys_from_secrets(): refreshes Anthropic OAuth from OS
  credential store before overlay is finalized
- main.rs: removed OAuth re-extraction hack (no longer needed)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: load OS credential store tokens even without secrets DB

The OAuth token extraction from macOS Keychain / Linux credentials files
was only running inside inject_llm_keys_from_secrets(), which requires
the encrypted secrets DB. When no master key is configured, init_secrets()
returned early — skipping both DB secret loading AND OS credential store
extraction, leaving the Anthropic OAuth token unavailable.

Split into two paths:
- inject_llm_keys_from_secrets(): loads from encrypted DB + OS stores
- inject_os_credentials(): loads from OS stores only (no DB needed)

init_secrets() now calls inject_os_credentials() and re-resolves config
even in the no-master-key early-return path, so `claude login` tokens
are always available regardless of secrets DB state.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: add anthropic-beta header required for OAuth authentication

Anthropic's api.anthropic.com requires the `anthropic-beta: oauth-2025-04-20`
header to accept OAuth Bearer tokens. Without it, the API returns 401
"OAuth authentication is currently not supported."

Also reverts API version to 2023-06-01 since the OAuth beta flag does
not support the 2024-10-22 version (returns 400 "not a valid version").

This was the same bug that caused PR nearai#143's 401 errors — the beta header
was missing entirely.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: Anthropic and OpenAI model resolution respects selected_model

The Anthropic and OpenAI config resolution ignored settings.selected_model
entirely, only checking the provider-specific env var (ANTHROPIC_MODEL,
OPENAI_MODEL) and falling back to a hardcoded default. This meant the
model chosen during onboarding wizard was silently overridden.

Now follows the same pattern as NearAI and OpenAI-compatible:
env var > settings.selected_model > hardcoded default.

Also deduplicated the Anthropic config construction (two identical
branches for API key vs OAuth now share model/base_url resolution).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* test: add provider resolution tests for all LLM backends

Covers deferred resolution (no credentials → None instead of error),
credential presence, model selection fallback chain, and OAuth token
routing for Anthropic, OpenAI, Tinfoil, Ollama, and NearAI.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: handle nested tokens.access_token format in Codex auth.json

Codex CLI stores OAuth tokens in a nested format under
tokens.access_token (ChatGPT OAuth flow), not at the top level.
Also adds ENV_MUTEX to Codex token tests for thread safety.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor: remove Codex OAuth onboarding (incompatible with OpenAI API)

Codex CLI OAuth tokens use a different endpoint
(chatgpt.com/backend-api/codex) and the Responses API wire format,
not api.openai.com with Chat Completions. The tokens lack the
model.request scope needed for the platform API, so they can't be
used as drop-in OPENAI_API_KEY replacements.

Removes: extract_codex_oauth_token(), wizard Codex OAuth flow,
CODEX_OAUTH_TOKEN env var support, and related tests.

OpenAI onboarding now uses direct API key only.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* style: fix formatting for CI (cargo fmt)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: address Gemini review feedback

- Use ? operator for ANTHROPIC_MODEL/BASE_URL env resolution instead of
  .ok().flatten() to propagate ConfigErrors consistently
- Skip Tool messages without tool_call_id with a warning instead of
  using unwrap_or_default() which would send empty string to Anthropic
- Extract credential check into closure to reduce duplication in
  Claude Code sandbox setup

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor(review): address PR review feedback for OAuth onboarding

- Gate ANTHROPIC_OAUTH_TOKEN resolution to Anthropic provider only
  (was needlessly checked for all registry providers)
- Add 3 regression tests for OAuth config resolution:
  - oauth_token sets placeholder api_key
  - real api_key takes priority over oauth
  - non-Anthropic providers don't pick up oauth_token
- Validate OAuth token prefix (sk-ant-oat) in wizard to catch
  accidentally pasted API keys
- Improve error body read handling in AnthropicOAuthProvider
  (was silently swallowing read errors with unwrap_or_default)
- Remove extra blank line in write_bootstrap_env
- Remove stale blank line in RegistryProviderConfig doc comment

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: address PR nearai#384 review comments

Blocker:
- Replace OnceLock<HashMap> with LazyLock<Mutex<HashMap>> for INJECTED_VARS
  so both inject_os_credentials() and inject_llm_keys_from_secrets() merge
  data instead of the second caller silently dropping its entries.

High:
- Add 401 retry with OS credential store re-extraction in
  AnthropicOAuthProvider, recovering from expired OAuth tokens (~8-12h)
  without manual intervention.
- Fix comment in app.rs: ~/.codex/auth.json → ~/.claude/.credentials.json.

Medium:
- Remove unsafe { std::env::set_var } from wizard; use thread-safe
  inject_single_var() overlay instead (safe on multi-threaded Tokio).
- Add post-init validation in AppBuilder: fail early with clear error when
  LLM_BACKEND is set but no credentials were resolved after secret injection.
- Add sk-ant-oat prefix validation in parse_oauth_access_token().
- Only route to AnthropicOAuthProvider when api_key is missing or equals
  OAUTH_PLACEHOLDER (API key takes priority over OAuth token).
- Teach fetch_anthropic_models() to use Bearer auth when only OAuth token
  is available (model listing no longer fails for OAuth-only users).

Low:
- Use optional_env() in wizard credential checks to read from injected
  overlay, not just raw env vars.

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* style: cargo fmt

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: ilblackdragon@gmail.com <ilblackdragon@gmail.com>
drchirag1991 pushed a commit to drchirag1991/ironclaw that referenced this pull request Apr 8, 2026
…rai#384)

* feat(setup): add Anthropic OAuth and Codex OAuth onboarding flows

Add OAuth token authentication as an alternative to API keys during
onboarding for both Anthropic (via `claude login`) and OpenAI/Codex
(via `~/.codex/auth.json`).

Key changes:
- New `AnthropicOAuthProvider` using `Authorization: Bearer` header
  (rig-core hardcodes `x-api-key` which rejects OAuth tokens)
- Wizard auth method selector: "Direct API Key" vs "OAuth Token"
  for both Anthropic and OpenAI providers
- Codex token extraction from `$CODEX_HOME/auth.json` / `~/.codex/auth.json`
- Claude Code sandbox sub-step in Docker setup (checks for credentials)
- Secret injection mappings for `ANTHROPIC_OAUTH_TOKEN` and `CODEX_OAUTH_TOKEN`
- `CODEX_OAUTH_TOKEN` falls back to `OPENAI_API_KEY` (same Bearer auth)

Supersedes nearai#143 which had a broken auth flow (OAuth token sent as
x-api-key → 401). Credit to @bigguybobby for the original approach.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: persist OAuth tokens in bootstrap .env and re-extract at startup

OAuth tokens stored only in the secrets DB were invisible to
Config::from_env() which runs before the DB connects (chicken-and-egg).

Two fixes:
1. write_bootstrap_env() now persists ANTHROPIC_OAUTH_TOKEN and
   CODEX_OAUTH_TOKEN to ~/.ironclaw/.env (same pattern as NEARAI_API_KEY)
2. main.rs re-extracts a fresh token from the OS credential store
   (macOS Keychain / ~/.claude/.credentials.json) before config resolution,
   handling token expiry (8-12h) gracefully

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: persist all LLM credentials in bootstrap .env, not just NEAR AI

All providers had the same chicken-and-egg issue: API keys stored in the
secrets DB were invisible to Config::from_env() which runs before DB
connects. Only NEARAI_API_KEY was written to bootstrap .env.

Now write_bootstrap_env() persists all credential env vars:
NEARAI_API_KEY, ANTHROPIC_API_KEY, ANTHROPIC_OAUTH_TOKEN, OPENAI_API_KEY,
CODEX_OAUTH_TOKEN, LLM_API_KEY, TINFOIL_API_KEY.

Also: setup_api_key_provider() now sets the env var during the wizard
session so write_bootstrap_env() can pick it up.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: address security review findings for OAuth onboarding

- Extract "oauth-placeholder" to named OAUTH_PLACEHOLDER constant shared
  across config and wizard to prevent silent drift
- Document plaintext credential tradeoff in write_bootstrap_env (API keys
  stored with 0o600 permissions, recommend full-disk encryption)
- Add blocking "Press Enter" wait in Anthropic OAuth retry flow so user
  has time to run `claude login` in another terminal
- Add escape hatch from manual OAuth paste back to API key flow (empty
  input switches to setup_api_key_provider)
- Fix Retry-After header: parse u64 seconds into Duration before passing
  to LlmError::RateLimited
- Make config::llm module pub(crate) for constant visibility
- Use .bearer_auth() instead of manual format!("Bearer {}")
- Remove response body from debug log (may contain PII)
- Update Anthropic API version to 2024-10-22

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* security: remove plaintext credentials from bootstrap .env

Credentials (API keys, OAuth tokens) were being written in plaintext to
~/.ironclaw/.env to work around a chicken-and-egg problem: Config::from_env()
runs before the encrypted secrets DB is connected.

Instead of storing secrets on disk, LlmConfig::resolve() now defers
gracefully when credentials are missing — it returns None for the provider
config instead of hard-erroring with MissingRequired. After the DB connects,
AppBuilder::build_all() loads secrets from encrypted storage via
inject_llm_keys_from_secrets() and re-resolves the config.

For Anthropic OAuth tokens (which expire in 8-12h), the secret injection
step also tries the OS credential store (macOS Keychain / Linux
credentials.json) for a fresh token, overriding the potentially stale
copy in the DB.

Changes:
- LlmConfig::resolve(): OpenAI, Anthropic, OpenAI-compatible, and Tinfoil
  all return None instead of MissingRequired when credentials are absent
- write_bootstrap_env(): no longer writes any credential env vars
- inject_llm_keys_from_secrets(): refreshes Anthropic OAuth from OS
  credential store before overlay is finalized
- main.rs: removed OAuth re-extraction hack (no longer needed)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: load OS credential store tokens even without secrets DB

The OAuth token extraction from macOS Keychain / Linux credentials files
was only running inside inject_llm_keys_from_secrets(), which requires
the encrypted secrets DB. When no master key is configured, init_secrets()
returned early — skipping both DB secret loading AND OS credential store
extraction, leaving the Anthropic OAuth token unavailable.

Split into two paths:
- inject_llm_keys_from_secrets(): loads from encrypted DB + OS stores
- inject_os_credentials(): loads from OS stores only (no DB needed)

init_secrets() now calls inject_os_credentials() and re-resolves config
even in the no-master-key early-return path, so `claude login` tokens
are always available regardless of secrets DB state.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: add anthropic-beta header required for OAuth authentication

Anthropic's api.anthropic.com requires the `anthropic-beta: oauth-2025-04-20`
header to accept OAuth Bearer tokens. Without it, the API returns 401
"OAuth authentication is currently not supported."

Also reverts API version to 2023-06-01 since the OAuth beta flag does
not support the 2024-10-22 version (returns 400 "not a valid version").

This was the same bug that caused PR nearai#143's 401 errors — the beta header
was missing entirely.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: Anthropic and OpenAI model resolution respects selected_model

The Anthropic and OpenAI config resolution ignored settings.selected_model
entirely, only checking the provider-specific env var (ANTHROPIC_MODEL,
OPENAI_MODEL) and falling back to a hardcoded default. This meant the
model chosen during onboarding wizard was silently overridden.

Now follows the same pattern as NearAI and OpenAI-compatible:
env var > settings.selected_model > hardcoded default.

Also deduplicated the Anthropic config construction (two identical
branches for API key vs OAuth now share model/base_url resolution).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* test: add provider resolution tests for all LLM backends

Covers deferred resolution (no credentials → None instead of error),
credential presence, model selection fallback chain, and OAuth token
routing for Anthropic, OpenAI, Tinfoil, Ollama, and NearAI.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: handle nested tokens.access_token format in Codex auth.json

Codex CLI stores OAuth tokens in a nested format under
tokens.access_token (ChatGPT OAuth flow), not at the top level.
Also adds ENV_MUTEX to Codex token tests for thread safety.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor: remove Codex OAuth onboarding (incompatible with OpenAI API)

Codex CLI OAuth tokens use a different endpoint
(chatgpt.com/backend-api/codex) and the Responses API wire format,
not api.openai.com with Chat Completions. The tokens lack the
model.request scope needed for the platform API, so they can't be
used as drop-in OPENAI_API_KEY replacements.

Removes: extract_codex_oauth_token(), wizard Codex OAuth flow,
CODEX_OAUTH_TOKEN env var support, and related tests.

OpenAI onboarding now uses direct API key only.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* style: fix formatting for CI (cargo fmt)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: address Gemini review feedback

- Use ? operator for ANTHROPIC_MODEL/BASE_URL env resolution instead of
  .ok().flatten() to propagate ConfigErrors consistently
- Skip Tool messages without tool_call_id with a warning instead of
  using unwrap_or_default() which would send empty string to Anthropic
- Extract credential check into closure to reduce duplication in
  Claude Code sandbox setup

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor(review): address PR review feedback for OAuth onboarding

- Gate ANTHROPIC_OAUTH_TOKEN resolution to Anthropic provider only
  (was needlessly checked for all registry providers)
- Add 3 regression tests for OAuth config resolution:
  - oauth_token sets placeholder api_key
  - real api_key takes priority over oauth
  - non-Anthropic providers don't pick up oauth_token
- Validate OAuth token prefix (sk-ant-oat) in wizard to catch
  accidentally pasted API keys
- Improve error body read handling in AnthropicOAuthProvider
  (was silently swallowing read errors with unwrap_or_default)
- Remove extra blank line in write_bootstrap_env
- Remove stale blank line in RegistryProviderConfig doc comment

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: address PR nearai#384 review comments

Blocker:
- Replace OnceLock<HashMap> with LazyLock<Mutex<HashMap>> for INJECTED_VARS
  so both inject_os_credentials() and inject_llm_keys_from_secrets() merge
  data instead of the second caller silently dropping its entries.

High:
- Add 401 retry with OS credential store re-extraction in
  AnthropicOAuthProvider, recovering from expired OAuth tokens (~8-12h)
  without manual intervention.
- Fix comment in app.rs: ~/.codex/auth.json → ~/.claude/.credentials.json.

Medium:
- Remove unsafe { std::env::set_var } from wizard; use thread-safe
  inject_single_var() overlay instead (safe on multi-threaded Tokio).
- Add post-init validation in AppBuilder: fail early with clear error when
  LLM_BACKEND is set but no credentials were resolved after secret injection.
- Add sk-ant-oat prefix validation in parse_oauth_access_token().
- Only route to AnthropicOAuthProvider when api_key is missing or equals
  OAUTH_PLACEHOLDER (API key takes priority over OAuth token).
- Teach fetch_anthropic_models() to use Bearer auth when only OAuth token
  is available (model listing no longer fails for OAuth-only users).

Low:
- Use optional_env() in wizard credential checks to read from injected
  overlay, not just raw env vars.

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* style: cargo fmt

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: ilblackdragon@gmail.com <ilblackdragon@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: experienced 6-19 merged PRs risk: high Safety, secrets, auth, or critical infrastructure scope: config Configuration scope: llm LLM integration scope: setup Onboarding / setup size: XL 500+ changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants