Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
157 commits
Select commit Hold shift + click to select a range
06bba33
docs: design NVIDIA-backed OpenCode development loop
seonghobae Aug 7, 2026
d32e2f1
docs: plan bounded OpenCode commercial development loop
seonghobae Aug 7, 2026
d90167f
feat(agent): define bounded OpenCode development policy
seonghobae Aug 7, 2026
9dc82ec
feat(agent): add commercial development package manifest
seonghobae Aug 7, 2026
780d941
test(agent): enforce complete commercial agent coverage
seonghobae Aug 7, 2026
4ba60d2
test(agent): define commercial development contracts
seonghobae Aug 7, 2026
6647172
feat(agent): implement bounded development contracts
seonghobae Aug 7, 2026
046c70e
test(agent): define deterministic issue selection
seonghobae Aug 7, 2026
395e760
feat(agent): implement bounded issue selection
seonghobae Aug 7, 2026
697ed99
test(agent): define prompt isolation contract
seonghobae Aug 7, 2026
5bd7876
feat(agent): build policy-isolated OpenCode prompts
seonghobae Aug 7, 2026
be2dd87
test(agent): define bounded diff policy
seonghobae Aug 7, 2026
a6ce477
feat(agent): enforce deterministic working-tree policy
seonghobae Aug 7, 2026
57ec20c
test(agent): define credential-free receipt composition
seonghobae Aug 7, 2026
223e864
feat(agent): compose credential-free execution receipts
seonghobae Aug 7, 2026
6b07f11
test(agent): define deterministic CLI boundaries
seonghobae Aug 7, 2026
ea8dd07
feat(agent): implement deterministic JSON CLI core
seonghobae Aug 7, 2026
d182ddc
feat(agent): expose commercial development CLI
seonghobae Aug 7, 2026
e9d8222
feat(agent): export commercial development contracts
seonghobae Aug 7, 2026
e4748ac
test(agent): add realistic autonomous loop dry run
seonghobae Aug 7, 2026
e75388f
test(agent): define hourly workflow security contract
seonghobae Aug 7, 2026
39d8d25
feat(agent): add hourly bounded OpenCode development loop
seonghobae Aug 7, 2026
89a081a
test(agent): verify public commercial agent surface
seonghobae Aug 7, 2026
983061f
ci(agent): bootstrap exact OpenCode package and verify contracts
seonghobae Aug 7, 2026
cc859ed
ci(agent): harden and verify OpenCode development contracts
seonghobae Aug 7, 2026
4a97056
docs: add OpenCode commercial loop runbook
seonghobae Aug 7, 2026
51805de
docs: record OpenCode loop standards and research
seonghobae Aug 7, 2026
227822f
ci(agent): stage final secure OpenCode loop
seonghobae Aug 7, 2026
58eb667
ci(agent): finalize secure OpenCode commercial loop
seonghobae Aug 7, 2026
b31fef8
ci(agent): autonomously resolve OpenCode loop review findings
seonghobae Aug 7, 2026
ba25565
ci(agent): merge OpenCode loop only after clean exact-head review
seonghobae Aug 7, 2026
85ee06f
fix(agent): parse issue references without dynamic regex
seonghobae Aug 7, 2026
3e8ae4d
test(agent): synthesize private-key detector fixture
seonghobae Aug 7, 2026
c56601e
ci(agent): finalize OpenCode development loop safely
seonghobae Aug 7, 2026
09187c3
fix(ci): allow the exact OpenCode lifecycle installation
seonghobae Aug 7, 2026
566eb98
fix(ci): remove the nonexistent commercial-agent fixture glob
seonghobae Aug 7, 2026
4e1f72b
test(agent): stage bounded contract corrections
seonghobae Aug 7, 2026
795bcaa
fix(ci): apply bounded commercial-agent contract repairs
seonghobae Aug 7, 2026
4b1850c
fix(agent): reject secret-shaped receipt version labels
seonghobae Aug 7, 2026
3697fb3
fix(agent): repair scheduled branch and process bounds
seonghobae Aug 7, 2026
c6ffe13
ci(agent): run bounded contract finalizer
seonghobae Aug 7, 2026
decd4a5
ci(agent): trigger bounded final verification
seonghobae Aug 7, 2026
5e1ebf0
test(agent): cover fail-closed commercial-agent boundaries
seonghobae Aug 7, 2026
c636afc
test(ci): rerun exhaustive commercial-agent coverage
seonghobae Aug 7, 2026
1507820
fix(ci): match the current exact-base recheck contract
seonghobae Aug 7, 2026
c420fa2
test(ci): rerun final OpenCode contract verification
seonghobae Aug 7, 2026
587199e
test(ci): diagnose exact OpenCode coverage gaps
seonghobae Aug 7, 2026
8f745a2
fix(agent): isolate untrusted model execution
seonghobae Aug 7, 2026
6f44a68
ci(agent): verify isolated OpenCode authority boundary
seonghobae Aug 7, 2026
cc0e540
fix(agent): remove stale exhaustive import
seonghobae Aug 7, 2026
8fa8150
ci(agent): include remaining quality repair
seonghobae Aug 7, 2026
80d99fb
fix(ci): normalize the OpenCode isolation patch boundary
seonghobae Aug 7, 2026
fd3ffab
test(agent): harden NIM credential and path boundaries
seonghobae Aug 7, 2026
056d981
test(ci): verify brokered OpenCode hardening test-first
seonghobae Aug 7, 2026
8495dd4
ci(agent): expose exact finalizer evidence
seonghobae Aug 7, 2026
1c01eb5
ci(agent): normalize finalizer Python source
seonghobae Aug 7, 2026
1b02e5c
chore(ci): remove temporary self-merging workflow
seonghobae Aug 7, 2026
f370f3a
chore(ci): remove temporary branch-writing review workflow
seonghobae Aug 7, 2026
a3573f6
chore(ci): remove temporary bootstrap workflow
seonghobae Aug 7, 2026
a96970d
chore(ci): remove temporary repair workflow
seonghobae Aug 7, 2026
7cc366b
chore(ci): remove temporary finalizer workflow
seonghobae Aug 7, 2026
dd1b142
chore(ci): remove self-removing verification workflow
seonghobae Aug 7, 2026
f954981
chore(ci): remove self-removing normalizer workflow
seonghobae Aug 7, 2026
c7d4b0f
chore(ci): remove bounded diagnostic workflow
seonghobae Aug 7, 2026
9aec803
chore(agent): remove one-shot finalizer script
seonghobae Aug 7, 2026
032c952
chore(agent): remove one-shot contract finalizer
seonghobae Aug 7, 2026
41b6787
chore(agent): remove one-shot isolation finalizer
seonghobae Aug 7, 2026
58ec424
chore(agent): remove encoded one-shot quality finalizer
seonghobae Aug 7, 2026
dfdd60d
fix(agent): remove unused coverage import
seonghobae Aug 7, 2026
877451a
ci(agent): generate lockfile evidence read-only
seonghobae Aug 7, 2026
2270d94
ci(agent): expose lockfile diagnostic on PR
seonghobae Aug 7, 2026
21a1609
ci(agent): remove temporary lockfile diagnostic
seonghobae Aug 7, 2026
82fbc7a
ci(agent): generate canonical lockfile evidence read-only
seonghobae Aug 7, 2026
2363c38
fix(agent): pin reviewed OpenCode CLI
seonghobae Aug 7, 2026
2a071e6
test(agent): enforce isolated NIM bridge trust boundary
seonghobae Aug 7, 2026
848c7a6
test(agent): reject model authority mutations
seonghobae Aug 7, 2026
8e7adbf
fix(agent): make verifier authority immutable to model
seonghobae Aug 7, 2026
638dc15
fix(agent): isolate model authority and NIM credential
seonghobae Aug 7, 2026
ad1316e
chore(ci): remove bounded lockfile diagnostic
seonghobae Aug 7, 2026
a8caf7e
ci: add bounded read-only lockfile diagnostic
seonghobae Aug 8, 2026
e92a11b
ci(agent): expose generated lockfile diagnostic artifact
seonghobae Aug 8, 2026
4fb1e3b
test(agent): cover nested protected policy paths
seonghobae Aug 8, 2026
d7342fa
test(agent): exercise policy regression in diagnostic
seonghobae Aug 8, 2026
8a47696
fix(agent): validate nested protected repository paths
seonghobae Aug 8, 2026
d4bddc8
ci(agent): retain bounded source diagnostic
seonghobae Aug 8, 2026
1c2851a
test(agent): lock trusted scheduler boundaries
seonghobae Aug 8, 2026
c3b1b31
test(agent): require evidence-based RCA before escalation
seonghobae Aug 8, 2026
f563cd4
ci(agent): exercise RCA prompt contract read-only
seonghobae Aug 8, 2026
9006a97
ci(agent): apply verified RCA contract repair
seonghobae Aug 8, 2026
3886a92
fix(agent): require evidence-based RCA and feasibility probes
seonghobae Aug 8, 2026
a75333a
ci(agent): verify RCA fix on current-main merge candidate
seonghobae Aug 8, 2026
51eb7a9
ci(agent): run verified RCA repair on pull request
seonghobae Aug 8, 2026
7bc87c8
ci(agent): finalize RCA contract through the live PR path
seonghobae Aug 8, 2026
2727ac3
ci(agent): execute verified RCA finalization through CI
seonghobae Aug 8, 2026
59b6d3a
ci(agent): run RCA finalizer with indentation-safe patch
seonghobae Aug 8, 2026
337baae
ci(agent): make lockfile diagnosis read-only
seonghobae Aug 8, 2026
1bf3f6b
chore(ci): remove prohibited branch finalizer
seonghobae Aug 8, 2026
e26b11c
chore(ci): remove prohibited RCA writer
seonghobae Aug 8, 2026
b989093
fix(agent): lint only maintained package files
seonghobae Aug 8, 2026
f541240
test(agent): preserve realistic multiline issue bodies
seonghobae Aug 8, 2026
ddce3a2
test(agent): capture focused red-green evidence read-only
seonghobae Aug 8, 2026
86ac80b
fix(agent): preserve bounded multiline issue evidence
seonghobae Aug 8, 2026
96d46f0
ci(agent): expand bounded read-only verification evidence
seonghobae Aug 8, 2026
10f6e09
ci: bound canonical lockfile repair
seonghobae Aug 8, 2026
316bb6e
fix(ci): synchronize commercial agent lockfile
github-actions[bot] Aug 8, 2026
b465d8d
chore(ci): remove completed lockfile repair workflow
seonghobae Aug 8, 2026
eca4ec1
ci(agent): add bounded one-shot formatter
seonghobae Aug 8, 2026
60fb185
style(agent): apply canonical Prettier output
github-actions[bot] Aug 8, 2026
d5941b5
chore(ci): remove completed formatting repair workflow
seonghobae Aug 8, 2026
cbd69d5
ci(agent): apply reviewed workflow security contract
seonghobae Aug 9, 2026
6ca0084
fix(ci): make workflow security repair deterministic
seonghobae Aug 9, 2026
1b373d6
ci(agent): prepare hardened workflow commit object
seonghobae Aug 9, 2026
c905c4e
fix(ci): expose validated workflow blob identity
seonghobae Aug 9, 2026
5614917
fix(agent): harden trusted development workflow boundary
seonghobae Aug 9, 2026
9f310c0
chore(ci): remove workflow commit preparation scaffold
seonghobae Aug 9, 2026
979425b
chore(ci): remove completed workflow security repair scaffold
seonghobae Aug 9, 2026
340ca59
fix(agent): reject tab control injection
seonghobae Aug 9, 2026
2d53b6f
fix(agent): classify unsafe paths as policy rejections
seonghobae Aug 9, 2026
84c10f0
fix(agent): accept validated diff decision receipts
seonghobae Aug 9, 2026
dd6807d
test(agent): align multiline fixture with tab rejection contract
seonghobae Aug 9, 2026
75f4052
refactor(agent): remove unreachable CLI coverage branches
seonghobae Aug 9, 2026
13b0605
refactor(agent): remove unreachable receipt serializer branch
seonghobae Aug 9, 2026
675d3a3
test(agent): cover remaining contract invariants
seonghobae Aug 9, 2026
cf9a22c
test(agent): cover diff path and root allowlist boundaries
seonghobae Aug 9, 2026
f7766a8
test(agent): cover CLI control and identifier boundaries
seonghobae Aug 9, 2026
7fe47c9
style(agent): format contract boundary regressions
seonghobae Aug 9, 2026
7ab3495
refactor(agent): remove unreachable CLI option branch
seonghobae Aug 9, 2026
74b6ca6
test(agent): cover malformed policy collection boundaries
seonghobae Aug 9, 2026
a60e26c
test(agent): cover valid deletion branch evidence
seonghobae Aug 9, 2026
7211e37
test(agent): cover fail-closed contract boundaries
seonghobae Aug 9, 2026
30eb4d0
test(agent): accept realistic multiline pull request evidence
seonghobae Aug 9, 2026
fd5ebd2
Merge branch 'main' into feat/opencode-commercial-development-loop
github-actions[bot] Aug 9, 2026
e057be9
fix(agent): preserve bounded multiline pull request evidence
seonghobae Aug 9, 2026
c0287df
test(agent): reject runner context in workflow-level env
seonghobae Aug 9, 2026
02b3d6c
fix(ci): scope runner temp to development job
seonghobae Aug 9, 2026
f9a3578
test(agent): require runtime runner temp initialization
seonghobae Aug 9, 2026
fa3ed7b
fix(ci): initialize receipt dir after runner starts
seonghobae Aug 9, 2026
002bf9c
test(agent): preserve complete invalid argv rows
seonghobae Aug 9, 2026
a302931
test(agent): expose contract boundary regressions
seonghobae Aug 9, 2026
2a51e9d
test(agent): reject case-variant dependency manifests
seonghobae Aug 9, 2026
efffe2a
test(agent): define untrusted text control boundaries
seonghobae Aug 9, 2026
70474c7
test(agent): isolate invalid receipt composition paths
seonghobae Aug 9, 2026
4d18dfe
style(agent): format control-character regression
seonghobae Aug 9, 2026
3ab3366
style(agent): apply exact Prettier output
seonghobae Aug 9, 2026
9285856
fix(agent): harden receipt and issue contracts
seonghobae Aug 9, 2026
bed75c6
fix(agent): preserve issue text whitespace
seonghobae Aug 9, 2026
1d04dca
fix(agent): reject mixed-case dependency manifests
seonghobae Aug 9, 2026
3db7b62
test(agent): align exhaustive control boundaries
seonghobae Aug 9, 2026
c86a22f
test(agent): cover workflow boundary regressions
seonghobae Aug 9, 2026
017b62e
fix(agent): harden isolated workflow boundaries
seonghobae Aug 9, 2026
e20f015
test(agent): match paginated jq aggregation
seonghobae Aug 9, 2026
0868909
docs(agent): correct credential and receipt operations
seonghobae Aug 9, 2026
73259e7
docs(agent): align provider research boundary
seonghobae Aug 9, 2026
f7bef9c
docs(agent): document loopback trust boundary
seonghobae Aug 9, 2026
165f9dd
docs(agent): align implementation plan with workflow
seonghobae Aug 9, 2026
f2ce6f7
test(agent): require trusted compose verification
seonghobae Aug 9, 2026
f6be470
test(agent): remove invalid Docker daemon assumption
seonghobae Aug 9, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1,026 changes: 1,026 additions & 0 deletions .github/workflows/opencode-commercial-development.yml

Large diffs are not rendered by default.

165 changes: 165 additions & 0 deletions docs/operations/opencode-commercial-development-loop.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,165 @@
# OpenCode commercial development loop runbook

## Purpose

The `OpenCode Commercial Development` workflow may implement one explicitly eligible LifeOS buyer-gap issue each hour. It does not merge, deploy, release, modify repository settings, or bypass the existing review loop. The deterministic Commercial Readiness workflow remains the authoritative audit and exact-head merge path.

## Enablement prerequisites

1. Store the NVIDIA provider credential as the repository secret `NVIDIA_NIM_API_KEY`.
2. Optionally configure repository variable `OPENCODE_NVIDIA_MODEL` with one NVIDIA NIM chat model identifier. When the value omits the provider prefix, the workflow prefixes `nvidia/`.
3. Keep `product/opencode-commercial-development-policy.json` under normal pull-request review.
4. Add an issue title to `eligible_issue_titles` only through a reviewed pull request after confirming the issue fits the initial non-destructive write boundary.
5. Verify the exact OpenCode package and `pnpm-lock.yaml` pin after every OpenCode update.

The workflow is disabled functionally when the NVIDIA secret is absent: it emits a credential-free `provider_credential_missing` receipt and performs no remote branch mutation.

## Hourly sequence

```mermaid
sequenceDiagram
participant Schedule as GitHub schedule
participant Audit as Deterministic audit
participant Selector as Policy selector
participant OpenCode
participant Validator as Diff validator
participant GitHub as GitHub mutation step
participant Review as Existing review loop

Schedule->>Audit: Snapshot repository and open PRs
Audit->>Selector: Bounded issue and PR projections
alt open PR or no eligible issue
Selector-->>Schedule: no_eligible_issue receipt
else eligible issue
Selector->>OpenCode: Fixed prompt + untrusted issue JSON
OpenCode->>OpenCode: Edit temporary UUIDv4 branch worktree
OpenCode-->>Validator: Working tree only
Validator->>Validator: Path, object, byte, line, content, base checks
alt rejected or verification failed
Validator-->>Schedule: credential-free rejected/failed receipt
else accepted and exact base unchanged
Validator->>GitHub: One commit, one branch, one draft PR
GitHub->>Review: Normal CI and review gates
end
end
```

## Credentials and trust boundaries

### NVIDIA credential

`NVIDIA_NIM_API_KEY` is present in exactly one workflow step. That step starts a loopback HTTP bridge as the separate system user `opencode_bridge`; only that bridge receives the credential and forwards bounded requests to NVIDIA NIM. OpenCode runs as `opencode_model` with `NVIDIA_API_KEY=local-loopback-placeholder` and a provider base URL pointing to the bridge.

During model execution, UID-based `iptables` rules reject other IPv4 and all IPv6 egress from `opencode_model`, permitting only the configured loopback bridge port. The workflow terminates the bridge before repository verification and removes the model and bridge processes, firewall rules, private homes, configuration, prompt, bridge code, and log in the `always()` cleanup step.

The credential must never appear in:

- Git configuration;
- issue or pull-request bodies;
- retained receipts or artifacts;
- model prompts or source files;
- the OpenCode/model process;
- repository tests or scripts;
- test logs;
- the `@life-os/commercial-development-agent` process;
- the GitHub mutation step;
- the existing review-agent credential scheme.

A suspected credential disclosure requires immediate secret rotation, cancellation of active runs, deletion of unreferenced automation branches, and review of workflow logs and draft pull requests. Do not retain or upload the raw OpenCode log while investigating. Treat `opencode_bridge`, rather than the OpenCode/model process, as the credential-bearing process during exposure analysis.

### GitHub credential

The checkout disables persisted credentials. OpenCode receives no `GITHUB_TOKEN` or `GH_TOKEN`. A later deterministic step receives `github.token` only after the diff, repository tests, and exact base SHA pass. That step may create one commit, push one same-repository UUIDv4 branch, and open one draft pull request. It cannot merge or release.

## Policy changes

Changes to allowed paths, issue titles, limits, model profiles, credentials, permissions, or mutation authority are security-sensitive architecture changes. They require:

- updated design and plan;
- realistic prompt-injection and policy tests;
- AppGuardrail, Semgrep, Security Scan, Commercial Readiness, CodeRabbit, and human review;
- exact-head success before merge.

The agent is prohibited from modifying its own workflow or policy in the initial slice.

## Receipts

The retained artifact contains only `receipt.json` using schema:

```text
life-os.opencode-commercial-development-receipt.v1
```

Retention is seven days. The receipt records counts, stable classifications, exact base SHA, external GitHub references, UUIDv4 run/branch identity, OpenCode version, model label, and deterministic validation outcomes. It excludes source paths, source diff, issue body, prompt, model output, hidden reasoning, credentials, provider bodies, raw logs, and stack traces.

## Failure handling

| Reason code | Operator interpretation | Remote mutation |
| ----------------------------- | ----------------------------------------------------------------------------- | -------------------------------------------------------------- |
| `no_eligible_issue` | Open PRs remain or no allowlisted issue is available | None |
| `provider_credential_missing` | NVIDIA secret is absent | None |
| `provider_unavailable` | Provider or OpenCode run failed | None |
| `opencode_unavailable` | Exact OpenCode CLI cannot execute | None |
| `invalid_configuration` | Policy, model, or workflow configuration is invalid | None |
| `prompt_rejected` | Prompt exceeds or violates the fixed contract | None |
| `diff_rejected` | Working-tree output violates path, object, size, content, or no-change policy | None |
| `base_changed` | `main` advanced after the run began | None |
| `verification_failed` | Repository tests or build failed | None |
| `draft_pull_request_failed` | A validated commit could not become a draft PR | Possible unreferenced automation branch; reconcile immediately |
| `completed` | One draft PR was created; normal review is still required | One branch and one draft PR |

The receipt contract reserves `opencode_unavailable`, `invalid_configuration`, and `prompt_rejected`; the current workflow receipt composer does not emit those codes.

## Branch reconciliation

Automation branches use:

```text
automation/opencode-commercial-<uuidv4>
```

An automation branch may be deleted when all conditions hold:

- no open or closed pull request references it;
- it is not the current head of an active workflow;
- the deterministic receipt does not report a pending draft-PR operation;
- an operator has verified that no unique reviewed work would be lost.

Never force-push an automation branch. If `main` advances, abandon the branch and rerun from the new exact base rather than silently rebasing model output.

## OpenCode update procedure

1. Create a feature branch.
2. Resolve the current official `opencode-ai` version once.
3. Add it with an exact version and update `pnpm-lock.yaml`.
4. Verify `opencode --version` and `opencode run --help`.
5. Run package and workflow-contract tests.
6. Inspect the lockfile and transitive dependency change.
7. Remove the temporary write-capable bootstrap workflow.
8. Obtain normal exact-head security and review evidence.

Never use a floating version, mutable installer script, or `curl | sh` in the persistent workflow.

## Central `.github` migration

The organization-central `.github` repository may later host a reusable wrapper that performs common checkout, OpenCode installation, secret scoping, and receipt upload. LifeOS must continue to own:

- issue-selection policy;
- product prompt policy;
- allowed/prohibited paths;
- realistic fixtures;
- receipt validation;
- package tests;
- branch and merge policy.

A central migration must pin the reusable workflow by exact commit SHA and preserve the existing review-agent credentials unchanged.

## Disablement

To stop model-assisted development immediately without affecting deterministic audit or merge behavior:

1. remove or rotate `NVIDIA_NIM_API_KEY`; or
2. remove every title from the eligible backlog through a reviewed policy change; or
3. disable the `OpenCode Commercial Development` workflow in GitHub Actions.

Do not disable the independent Commercial Readiness workflow when responding to a model-provider incident.
Original file line number Diff line number Diff line change
@@ -0,0 +1,95 @@
# OpenCode commercial development loop: standards and research basis

**Reviewed:** 2026-08-07
**Scope:** Hourly NVIDIA-backed OpenCode development in `ContextualWisdomLab/life-os`

## Evidence-status rule

LifeOS distinguishes normative standards and final vendor documentation from conference papers, research releases, and preprints. Research results motivate hypotheses and ablations; they do not grant repository authority or replace deterministic tests, security review, and exact-head merge gates.

## Normative security and governance basis

### NIST AI RMF and Generative AI Profile

The AI Risk Management Framework organizes governed AI risk work around Govern, Map, Measure, and Manage. The Generative AI Profile extends that framework with risks such as confabulation, information integrity, privacy, human-AI configuration, and value-chain integration. LifeOS maps these concepts into explicit policy, bounded authority, retained measurements, credential-free receipts, and independent deterministic review (National Institute of Standards and Technology, 2023, 2024).

### GitHub Actions hardening

GitHub recommends least-privilege `GITHUB_TOKEN` permissions, immutable third-party action references, protected branches, careful treatment of untrusted input, and separation of trusted and untrusted execution. The OpenCode step therefore receives no GitHub credential, runs against an exact main snapshot, and cannot push. A later deterministic step receives narrowly scoped repository authority only after the diff and base SHA have passed validation (GitHub, 2026a).

A future organization-central wrapper should use a reusable workflow pinned by exact commit SHA. Repository-specific issue, path, prompt, receipt, and merge policy remains inside LifeOS so shared automation cannot silently broaden product authority (GitHub, 2026b).

### OWASP risks for model-assisted software changes

The OWASP Top 10 for LLM Applications identifies prompt injection, sensitive-information disclosure, excessive agency, improper output handling, supply-chain risk, and unbounded consumption as material concerns. LifeOS treats issue text and model output as untrusted, prevents the model from receiving GitHub credentials, validates source output before execution or push, pins OpenCode and GitHub actions, scopes the NVIDIA key to one process, and enforces file, byte, line, time, recursion, decomposition, and concurrency limits (OWASP Foundation, 2025).

## OpenCode and NVIDIA provider boundary

OpenCode exposes a non-interactive `run` command and provider/model configuration. LifeOS uses one exact reviewed `opencode-ai` package version and verifies both the installed version and command contract. Auto-update and sharing are disabled. The model receives a private configuration and a source archive without `.git`; Bash is denied by default except for reviewed `pnpm`, `node`, `python3`, `grep`, `rg`, `find`, `ls`, and `cat` command patterns, while web-fetch, web-search, and external-directory access are denied. The prompt is attached from a private file instead of carrying issue text in process arguments (Anomaly, 2026).

NVIDIA NIM exposes hosted OpenAI-compatible inference authenticated with an API key. `NVIDIA_NIM_API_KEY` is mapped only to a loopback bridge running as `opencode_bridge`. OpenCode runs separately as `opencode_model` with a placeholder API-key value and an allowlisted minimal environment; UID-based `iptables` rules restrict its model-phase egress to the bridge. GitHub, review-agent, deployment, and unrelated repository credentials are absent. Provider availability is evidence, not a deterministic merge prerequisite (NVIDIA Corporation, 2026).

## Test-time compute allocation

### Strong single-agent baseline

Xu et al. report that a multi-turn single agent can match homogeneous multi-agent workflows in several evaluated settings and can benefit from KV-cache reuse. Because broader orchestration adds coordination and attack surface, LifeOS requires a strong single-model route as the mandatory baseline and does not assume that more agents are better (Xu et al., 2026b).

### Fugu

Sakana AI reports a system that dynamically selects between direct answering and an expert team. This supports an explicit routing decision rather than always-on multi-agent execution. LifeOS treats the result as a final research release and measures whether repository-specific issue fixtures justify deeper orchestration (Sakana AI, 2026).

### Conductor

Conductor learns natural-language communication topologies and targeted instructions, including recursive self-selection for dynamic test-time scaling. The final ICLR 2026 conference paper motivates explicit topology, access-list, recursive-depth, and decomposition fields in the LifeOS ablation contract (Nielsen et al., 2026).

### TRINITY

TRINITY reports a lightweight evolved coordinator assigning Thinker, Worker, and Verifier roles across multiple turns. The final ICLR 2026 conference paper motivates role-specific reasoning effort and explicit verification rather than an undifferentiated agent pool (Xu et al., 2026a).

## LifeOS design conclusions

The initial workflow intentionally runs one high-effort OpenCode model with recursion depth one. It records a versioned contract for planner, worker, verifier, and synthesizer roles but does not enable hidden multi-agent delegation. A contextual-orchestrator profile may be introduced only when:

1. the same realistic issue fixtures are used for route and orchestrated cells;
2. issue selection, prompt policy, source authority, and diff validation remain deterministic;
3. prompt-injection and sensitive-information tests do not regress;
4. quality or heterogeneous capability improves materially;
5. unsupported profile fields remain explicit rather than simulated;
6. the exact contextual-orchestrator commit and dependency hashes are reviewed;
7. retained artifacts exclude prompt, response, hidden reasoning, source diff, and credentials.

Latency and token use are recorded for cost and capacity review but are not the primary optimization objective. Product correctness, security, auditability, and buyer-visible quality determine the routing decision.

## Limitations

- Vendor documentation describes interfaces, not independent security assurance.
- The initial dry-run fixtures cannot establish general autonomous-development reliability.
- The workflow applies UID-based `iptables` restrictions to `opencode_model`—allowing only the loopback bridge during model execution and denying IPv6—but does not provide a general-purpose operating-system sandbox; the model also receives a source archive without `.git` or GitHub credentials.
- Provider-side retention and processing remain subject to the deployment operator's NVIDIA agreement and data-governance assessment.
- The model cannot discover or authorize new backlog work in this slice; issue eligibility is an explicit reviewed policy.
- A draft pull request is evidence for review, not proof of correctness or permission to merge.

## References

Anomaly. (2026). _OpenCode documentation_. https://opencode.ai/docs/

GitHub. (2026a). _Security hardening for GitHub Actions_. https://docs.github.com/en/actions/security-for-github-actions/security-guides/security-hardening-for-github-actions

GitHub. (2026b). _Reusing workflows_. https://docs.github.com/en/actions/using-workflows/reusing-workflows

National Institute of Standards and Technology. (2023). _Artificial intelligence risk management framework (AI RMF 1.0)_ (NIST AI 100-1). https://doi.org/10.6028/NIST.AI.100-1

National Institute of Standards and Technology. (2024). _Artificial intelligence risk management framework: Generative artificial intelligence profile_ (NIST AI 600-1). https://doi.org/10.6028/NIST.AI.600-1

Nielsen, S., Cetin, E., Schwendeman, P., Sun, Q., Xu, J., & Tang, Y. (2026). _Learning to orchestrate agents in natural language with the Conductor_ [Conference paper]. International Conference on Learning Representations. https://openreview.net/forum?id=U23A2BUKYt

NVIDIA Corporation. (2026). _API reference—NVIDIA NIM for large language models_. https://docs.nvidia.com/nim/large-language-models/latest/api-reference.html

OWASP Foundation. (2025). _OWASP Top 10 for large language model applications 2025_. https://genai.owasp.org/llm-top-10/

Sakana AI. (2026, June 22). _Sakana Fugu: One model to command them all_ [Final research release]. https://sakana.ai/fugu-release/

Xu, J., Sun, Q., Schwendeman, P., Nielsen, S., Cetin, E., & Tang, Y. (2026a). _TRINITY: An evolved LLM coordinator_ [Conference paper]. International Conference on Learning Representations. https://doi.org/10.48550/arXiv.2512.04695

Xu, J., Koesdwiady, A., Bei, S., Han, Y., Huang, B., Wang, D., Chen, Y., Wang, Z., Wang, P., Li, P., & Ding, Y. (2026b). _Rethinking the value of multi-agent workflow: A strong single agent baseline_ [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2601.12307
Loading
Loading