Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion MODULE.bazel
Original file line number Diff line number Diff line change
Expand Up @@ -92,7 +92,7 @@ go_deps.from_file(go_work = "//:go.work.bazel")
# use_repo entries below are auto-managed by `bazel mod tidy`. Re-run after
# adding new imports in nvcf-cli, src/libraries/go/lib, or
# src/libraries/go/worker.
use_repo(go_deps, "cat_dario_mergo", "com_github_aws_aws_sdk_go_v2", "com_github_aws_aws_sdk_go_v2_config", "com_github_aws_aws_sdk_go_v2_credentials", "com_github_aws_smithy_go", "com_github_awslabs_amazon_ecr_credential_helper_ecr_login", "com_github_carlmjohnson_versioninfo", "com_github_cenkalti_backoff_v4", "com_github_charmbracelet_bubbles", "com_github_charmbracelet_bubbletea", "com_github_charmbracelet_lipgloss", "com_github_dustin_go_humanize", "com_github_felixge_httpsnoop", "com_github_fsnotify_fsnotify", "com_github_go_co_op_gocron", "com_github_go_jose_go_jose_v4", "com_github_go_kit_kit", "com_github_go_viper_mapstructure_v2", "com_github_goccy_go_json", "com_github_golang_jwt_jwt_v5", "com_github_google_go_containerregistry", "com_github_google_shlex", "com_github_google_uuid", "com_github_gorilla_mux", "com_github_grpc_ecosystem_go_grpc_middleware", "com_github_grpc_ecosystem_grpc_gateway_v2", "com_github_hashicorp_go_cleanhttp", "com_github_hashicorp_go_retryablehttp", "com_github_hellofresh_health_go_v5", "com_github_jarcoal_httpmock", "com_github_kimmachinegun_automemlimit", "com_github_klauspost_compress", "com_github_madappgang_httplog", "com_github_madappgang_httplog_zap", "com_github_masterminds_semver_v3", "com_github_mattn_go_isatty", "com_github_mitchellh_go_homedir", "com_github_muesli_termenv", "com_github_nats_io_nats_go", "com_github_nats_io_nats_server_v2", "com_github_nats_io_nkeys", "com_github_nvidia_nvcf_src_libraries_go_lib", "com_github_oklog_oklog", "com_github_openzipkin_zipkin_go", "com_github_prometheus_client_golang", "com_github_prometheus_client_model", "com_github_quic_go_quic_go", "com_github_samber_lo", "com_github_senseyeio_duration", "com_github_sirupsen_logrus", "com_github_spf13_cobra", "com_github_spf13_pflag", "com_github_spf13_viper", "com_github_stretchr_testify", "com_github_urfave_cli_v2", "com_github_valyala_bytebufferpool", "com_github_volcengine_volcengine_go_sdk", "in_gopkg_yaml_v3", "in_yaml_go_yaml_v3", "io_k8s_api", "io_k8s_apiextensions_apiserver", "io_k8s_apimachinery", "io_k8s_client_go", "io_k8s_sigs_yaml", "io_opentelemetry_go_contrib_instrumentation_github.com_gorilla_mux_otelmux", "io_opentelemetry_go_contrib_instrumentation_google_golang_org_grpc_otelgrpc", "io_opentelemetry_go_contrib_instrumentation_host", "io_opentelemetry_go_contrib_instrumentation_net_http_httptrace_otelhttptrace", "io_opentelemetry_go_contrib_instrumentation_net_http_otelhttp", "io_opentelemetry_go_otel", "io_opentelemetry_go_otel_exporters_otlp_otlptrace_otlptracegrpc", "io_opentelemetry_go_otel_exporters_prometheus", "io_opentelemetry_go_otel_sdk", "io_opentelemetry_go_otel_sdk_metric", "io_opentelemetry_go_otel_trace", "io_opentelemetry_go_proto_otlp", "land_oras_oras_go_v2", "org_golang_google_genproto_googleapis_rpc", "org_golang_google_grpc", "org_golang_google_grpc_cmd_protoc_gen_go_grpc", "org_golang_google_protobuf", "org_golang_x_net", "org_golang_x_oauth2", "org_golang_x_sync", "org_golang_x_term", "org_uber_go_automaxprocs", "org_uber_go_zap")
use_repo(go_deps, "cat_dario_mergo", "com_github_aws_aws_sdk_go_v2", "com_github_aws_aws_sdk_go_v2_config", "com_github_aws_aws_sdk_go_v2_credentials", "com_github_aws_smithy_go", "com_github_awslabs_amazon_ecr_credential_helper_ecr_login", "com_github_carlmjohnson_versioninfo", "com_github_cenkalti_backoff_v4", "com_github_charmbracelet_bubbles", "com_github_charmbracelet_bubbletea", "com_github_charmbracelet_lipgloss", "com_github_dustin_go_humanize", "com_github_felixge_httpsnoop", "com_github_fsnotify_fsnotify", "com_github_go_co_op_gocron", "com_github_go_jose_go_jose_v4", "com_github_go_kit_kit", "com_github_go_viper_mapstructure_v2", "com_github_goccy_go_json", "com_github_golang_jwt_jwt_v5", "com_github_google_go_containerregistry", "com_github_google_shlex", "com_github_google_uuid", "com_github_gorilla_mux", "com_github_grpc_ecosystem_go_grpc_middleware", "com_github_grpc_ecosystem_grpc_gateway_v2", "com_github_hashicorp_go_cleanhttp", "com_github_hashicorp_go_retryablehttp", "com_github_hellofresh_health_go_v5", "com_github_jarcoal_httpmock", "com_github_kimmachinegun_automemlimit", "com_github_klauspost_compress", "com_github_madappgang_httplog", "com_github_madappgang_httplog_zap", "com_github_masterminds_semver_v3", "com_github_mattn_go_isatty", "com_github_mitchellh_go_homedir", "com_github_muesli_termenv", "com_github_nats_io_nats_go", "com_github_nats_io_nats_server_v2", "com_github_nats_io_nkeys", "com_github_nvidia_nvcf_src_libraries_go_lib", "com_github_oklog_oklog", "com_github_openzipkin_zipkin_go", "com_github_prometheus_client_golang", "com_github_prometheus_client_model", "com_github_quic_go_quic_go", "com_github_samber_lo", "com_github_senseyeio_duration", "com_github_sirupsen_logrus", "com_github_spf13_cobra", "com_github_spf13_pflag", "com_github_spf13_viper", "com_github_stretchr_testify", "com_github_urfave_cli_v2", "com_github_valyala_bytebufferpool", "com_github_volcengine_volcengine_go_sdk", "in_gopkg_yaml_v3", "in_yaml_go_yaml_v3", "io_k8s_api", "io_k8s_apiextensions_apiserver", "io_k8s_apimachinery", "io_k8s_client_go", "io_k8s_sigs_yaml", "io_opentelemetry_go_contrib_instrumentation_github.com_gorilla_mux_otelmux", "io_opentelemetry_go_contrib_instrumentation_google_golang_org_grpc_otelgrpc", "io_opentelemetry_go_contrib_instrumentation_host", "io_opentelemetry_go_contrib_instrumentation_net_http_httptrace_otelhttptrace", "io_opentelemetry_go_contrib_instrumentation_net_http_otelhttp", "io_opentelemetry_go_otel", "io_opentelemetry_go_otel_exporters_otlp_otlptrace_otlptracegrpc", "io_opentelemetry_go_otel_exporters_prometheus", "io_opentelemetry_go_otel_metric", "io_opentelemetry_go_otel_sdk", "io_opentelemetry_go_otel_sdk_metric", "io_opentelemetry_go_otel_trace", "io_opentelemetry_go_proto_otlp", "land_oras_oras_go_v2", "org_golang_google_genproto_googleapis_rpc", "org_golang_google_grpc", "org_golang_google_grpc_cmd_protoc_gen_go_grpc", "org_golang_google_protobuf", "org_golang_x_net", "org_golang_x_oauth2", "org_golang_x_sync", "org_golang_x_term", "org_uber_go_automaxprocs", "org_uber_go_zap")

# ============================================================================
# C++ cross-compilation (multi-arch container builds).
Expand Down
3 changes: 3 additions & 0 deletions NOTICE
Original file line number Diff line number Diff line change
Expand Up @@ -219,6 +219,7 @@ The following third-party licenses are included in this repository:
src/compute-plane-services/nvca/vendor/github.com/prometheus/client_model/NOTICE
src/compute-plane-services/nvca/vendor/github.com/prometheus/common/LICENSE
src/compute-plane-services/nvca/vendor/github.com/prometheus/common/NOTICE
src/compute-plane-services/nvca/vendor/github.com/prometheus/otlptranslator/LICENSE
src/compute-plane-services/nvca/vendor/github.com/prometheus/procfs/LICENSE
src/compute-plane-services/nvca/vendor/github.com/prometheus/procfs/NOTICE
src/compute-plane-services/nvca/vendor/github.com/run-ai/karta/LICENSE
Expand Down Expand Up @@ -251,9 +252,11 @@ The following third-party licenses are included in this repository:
src/compute-plane-services/nvca/vendor/go.opentelemetry.io/contrib/instrumentation/net/http/otelhttp/LICENSE
src/compute-plane-services/nvca/vendor/go.opentelemetry.io/otel/exporters/otlp/otlptrace/LICENSE
src/compute-plane-services/nvca/vendor/go.opentelemetry.io/otel/exporters/otlp/otlptrace/otlptracegrpc/LICENSE
src/compute-plane-services/nvca/vendor/go.opentelemetry.io/otel/exporters/prometheus/LICENSE
src/compute-plane-services/nvca/vendor/go.opentelemetry.io/otel/LICENSE
src/compute-plane-services/nvca/vendor/go.opentelemetry.io/otel/metric/LICENSE
src/compute-plane-services/nvca/vendor/go.opentelemetry.io/otel/sdk/LICENSE
src/compute-plane-services/nvca/vendor/go.opentelemetry.io/otel/sdk/metric/LICENSE
src/compute-plane-services/nvca/vendor/go.opentelemetry.io/otel/trace/LICENSE
src/compute-plane-services/nvca/vendor/go.opentelemetry.io/proto/otlp/LICENSE
src/compute-plane-services/nvca/vendor/go.uber.org/mock/LICENSE
Expand Down
46 changes: 46 additions & 0 deletions docs/user/cluster-management/nvcf-ui.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,46 @@
(enabling-nvcf-ui)=

# Enabling NVCF UI

The NVCF UI is an optional web interface for managing NVCF deployments. It is
disabled by default. Enable it only when the `nvcf-ui` addon is installed in
your cluster.

The NVCF UI addon runs as a Service named `nvcf-ui` in the `nvcf-ui` namespace
on port 8300. When enabled, the gateway-routes chart creates an HTTPRoute and a
ReferenceGrant that forward requests from `nvcf-ui.<domain>` to that Service.

## Prerequisites

- The `nvcf-ui` addon must be deployed in the `nvcf-ui` namespace before
enabling the gateway route.
- Gateway API ingress must be configured. See [Gateway Routing](../gateway-routing.md).

## Enable the gateway route

In your Helmfile environment values file (for example
`environments/<environment-name>.yaml`), set:

```yaml
ingress:
gatewayApi:
routes:
nvcfUi:
enabled: true
```

Then sync the ingress release to apply:

```bash
HELMFILE_ENV=<environment-name> helmfile --selector release-group=ingress sync
```

The UI is available at `http://nvcf-ui.<domain>` after the HTTPRoute is ready.

## Verify

```bash
kubectl get httproute nvcf-ui -n envoy-gateway
```

The route should show `Accepted` status and the hostname `nvcf-ui.<domain>`.
2 changes: 2 additions & 0 deletions fern/versions/dev.yml
Original file line number Diff line number Diff line change
Expand Up @@ -65,6 +65,8 @@ navigation:
path: ../../docs/user/cluster-management/configuration.md
- page: Multi-Tenancy
path: ../../docs/user/cluster-management/multi-tenancy.md
- page: NVCF UI
path: ../../docs/user/cluster-management/nvcf-ui.md
- page: KAI Scheduler
path: ../../docs/user/cluster-management/kai-scheduler.md
- section: Low Latency Streaming
Expand Down
21 changes: 21 additions & 0 deletions src/compute-plane-services/nvca/AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -396,6 +396,27 @@ Counter metrics are pre-initialized to zero in `internal/metrics/metrics.go` so
- `rate()` calculations give unexpected results if metrics appear mid-scrape
- Dashboards show gaps instead of zeros

### OTel client metrics (outbound dependencies)

Outbound dependency clients are instrumented with OpenTelemetry metrics that
follow the OpenTelemetry Semantic Conventions, separate from the client_golang
metrics above. This path does not use the manual zero-init loops; instruments are
created from an OTel MeterProvider and exported through the OTel to Prometheus
bridge onto the same `/metrics` endpoint.

- Pipeline: `internal/otel/meter.go` builds the MeterProvider and Prometheus
exporter. It is gated by the `ClientMetrics` feature flag; when off, a no-op
meter provider is installed and nothing is emitted.
- Recording: `internal/metrics/clientmetrics/` holds the shared `Recorder` and
the HTTP metrics `RoundTripper`. The wrapper is attached to a client through the
shared HTTP factory's `WithTransportWrapper` option.
- Labels: use the semconv helpers in `internal/metrics/semconv/` (`httpsemconv`,
`msgsemconv`, `rpcsemconv`). Keep label values bounded.
- Adding a dependency: for a new HTTP client, pass the metrics transport wrapper
with a `peer.service` name (add the constant in
`internal/metrics/clientmetrics`). For a new client type, add a thin decorator
over the shared `Recorder`. See `internal/metrics/METRICS.md` for the recipe.

Comment on lines +399 to +419

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

Keep this AGENTS.md below 400 lines.

The file now reaches line 442. Move this detailed metrics recipe to internal/metrics/METRICS.md and retain only a short pointer here.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/compute-plane-services/nvca/AGENTS.md` around lines 399 - 419, Reduce
AGENTS.md to below 400 lines by removing the detailed “OTel client metrics
(outbound dependencies)” recipe and moving its guidance to
internal/metrics/METRICS.md. Retain only a concise pointer in AGENTS.md
directing contributors to that document for OTel client metrics instructions.

Source: Coding guidelines

## Key Environment Variables

Useful local test variables:
Expand Down
3 changes: 3 additions & 0 deletions src/compute-plane-services/nvca/MODULE.bazel
Original file line number Diff line number Diff line change
Expand Up @@ -128,7 +128,10 @@ use_repo(
"io_opentelemetry_go_contrib_instrumentation_net_http_httptrace_otelhttptrace",
"io_opentelemetry_go_contrib_instrumentation_net_http_otelhttp",
"io_opentelemetry_go_otel",
"io_opentelemetry_go_otel_exporters_prometheus",
"io_opentelemetry_go_otel_metric",
"io_opentelemetry_go_otel_sdk",
"io_opentelemetry_go_otel_sdk_metric",
"io_opentelemetry_go_otel_trace",
"org_golang_x_time",
"xyz_gomodules_jsonpatch_v2",
Expand Down
83 changes: 83 additions & 0 deletions src/compute-plane-services/nvca/docs/users/byoc/featureflags.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,83 @@
# NVCA Attributes and Feature Flags

After installing the nvca-operator, edit its `NVCFBackend` object to add feature flags.
Default-enabled feature flags can be _disabled_ by prepending `-` to its name in the `values` list,
ex. `-CachingSupport`.

**Note**: make sure to copy over existing spec feature flag values into the equivalent override values,
since that list overwritten not merged.
Comment on lines +7 to +8

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

Remove Markdown bold emphasis.

Repository documentation rules prohibit Markdown bold.

  • src/compute-plane-services/nvca/docs/users/byoc/featureflags.md#L7-L8: replace the bold “Note” label with plain text.
  • src/compute-plane-services/nvca/docs/users/byoc/featureflags.md#L83-L83: replace the bold “Note” label with plain text.
📍 Affects 1 file
  • src/compute-plane-services/nvca/docs/users/byoc/featureflags.md#L7-L8 (this comment)
  • src/compute-plane-services/nvca/docs/users/byoc/featureflags.md#L83-L83
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/compute-plane-services/nvca/docs/users/byoc/featureflags.md` around lines
7 - 8, Remove Markdown bold emphasis from both “Note” labels in
src/compute-plane-services/nvca/docs/users/byoc/featureflags.md at lines 7-8 and
83, leaving each label as plain text while preserving the surrounding
documentation.

Source: Coding guidelines


Example:

```

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win

Declare the code-block language.

Use a yaml fence so Markdown tooling and renderers classify the configuration example correctly.

-```
+```yaml
🧰 Tools
🪛 markdownlint-cli2 (0.23.1)

[warning] 12-12: Fenced code blocks should have a language specified

(MD040, fenced-code-language)

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@src/compute-plane-services/nvca/docs/users/byoc/featureflags.md` at line 12,
Update the configuration example code fence in featureflags.md to declare the
yaml language, using a yaml fence for the existing code block while preserving
its contents.

Source: Linters/SAST tools

$ kubectl edit nvcfbackend -n nvca-operator
...
spec:
featureGate:
values:
- LogPosting # Existing feature flag
overrides:
featureGate:
values:
- LogPosting # Existing feature flag copied over
- -CachingSupport # Caching support disabled
- LowLatencyStreaming
- BYOObservability
...
```

## Attributes

| Name | Default | Description |
| --- | --- | --- |
| KataRuntimeIsolation | false | Forces NVCF workload pods to get a Kata runtime class, and allows "nvidia.com/pgpu" resource types on nodes when parsing node resources |
| HostIsolation | false | Prevents pods from more than one function from running on a given node at once |
| AccountIsolation | false | Prevents pods from different functions belonging to more than one Nvidia Cloud Account from running on a given node at once |
| TimeSlicingGPUEnabled | false | Forces NVCA's node feature handler to permit time-sliced GPUs when parsing node resources |
| PassthroughGPUEnabled | false | Allows "nvidia.com/pgpu" resource types on nodes when parsing node resources |
| OVCSecurityEnforcements | false | Turns on OVC security enforcements outlined by the "SensorRTX Risk Mitigation" SDD |
| NVLinkOptimized | false | Turns on NVLink Optimization on Clusters and related validations on Agent startup |

## Feature Flags

| Name | Default | Description |
| --- | --- | --- |
| LogPosting | false | Post instance logs to ICMS directly |
| CachingSupport | false | Enable NVMesh caching support for Container functions and tasks |
| HelmCachingSupport | false | Enable NVMesh caching support for Helm functions and tasks |
| NVMeshEncryption | false | Enable NVMesh encryption on cache data |
| PeriodicInstanceStatusUpdate | true | Enable periodic syncs with ICMS to reconcile instance state differences |
| HelmRBACEnforcement | true | Enforce RBAC constraints on Helm charts specified by functions |
| DynamicGPUDiscovery | true | Dynamically discover GPUs and instance types on this cluster |
| MultipleGPUTypesAllowed | true | Permit a heterogeneous set of GPUs across nodes in this cluster, ex. L40 and A100 |
| AutoPurgeDegradedWorkers | true | Automatically delete function instances and tasks that have degraded Pods |
| HelmSharedStorage | true | Configure Helm functions and tasks with shared read-only storage for ESS secrets |
| ClusterTargeting | true | Enable targeted cluster queues |
| HelmResourceConstraints | true | Enforce GPU quota adherence on Helm functions and tasks |
| BinPackTenantWorkloads | false | Prefer that pods from the same function or task are scheduled on the same node |
| GXCache | false | Enable GXCache support in NVCA |
| LowLatencyStreaming | true | Enable LLS support in NVCA |
| UseFunctionDeploymentStages | false | Enable container stage transition event logging to Function Deployment Stages service |
| PVCRebind | false | Force cache PVC's to rebind on failure |
| MultiNodeWorkloads | true | Instruct NVCA to send multi-node instance types to ICMS during registration |
| UseFunctionTranslator | true | Use the nvcf-icms-translate translator to generate function manifests instead of ICMS-generated artifacts |
| BYOObservability | false | Enable Bring-your-own observability support in NVCA |
| BYOOFluentBit | false | Enable Bring-your-own observability FluentBit logging sidecar in workload pods |
| ClientMetrics | false | Emit OpenTelemetry semantic-convention metrics for NVCA's outbound dependency clients |
| MaxSQSBatchPull | true | Increase the pull batch size from the SQS queue from 1 to 10 |
| InfraResourceOverhead | false | InfraResourceOverhead enables subtraction of infrastructure resource overhead from instance type resources, potentially removing any instance type that cannot satisfy infrastructure resources |
| EnforceHelmFunctionResourceLimits | false | Enforces resource limits on helm functions via ResourceQuota's. Sets `podSpec.{initContainers,containers}[*].resource.requests = limits` |
| EnforceContainerFunctionResourceLimits | false | Enforces resource limits on container functions via container resource limits. Sets `podSpec.{initContainers,containers}[*].resource.requests = limits` |
| EnforceHelmTaskResourceLimits | false | Enforces resource limits on helm tasks via ResourceQuota's. Sets `podSpec.{initContainers,containers}[*].resource.requests = limits` |
| EnforceContainerTaskResourceLimits | false | Enforces resource limits on container tasks via container resource limits. Sets `podSpec.{initContainers,containers}[*].resource.requests = limits` |
| CordonMaintenance | false | Sets the mode for NVCA to maintenance and only pauses new workloads on the cluster backend |
| CordonAndDrainMaintenance | false | Sets the mode for NVCA to maintenance and evicts existing workloads disruptively on the cluster backend |
| AckTaskRequestAfterPodsScheduled | false | Instructs the agent to only acknowledge ICMS requests with ICMS and delete queue messages after all NVCT task pods have been accepted by the cluster's scheduler |
| SelfHosted | false | Enable Self-Hosted mode |
| GracefulNoGPU | false | Allow NVCA to start and operate without GPUs, pausing queue processing until GPUs become available |
| HelmCustomAnnotations | false | Enable Custom Annotations for Helm workloads |
| KAIScheduler | false | Enables bin-packing support for efficient resource utilization using KAI scheduler |
| HelmAllowCPUNodes | false | Allows CPU-only pods in Helm functions to be scheduled on non-GPU nodes. GPU pods retain required instance-type affinity while CPU-only pods get anti-preference for GPU nodes. Mutually exclusive with HelmResourceConstraints |
| MiniServiceRevisionHistory | true | Enables saving prior helm values as ConfigMaps on each MiniService's values update. |
| AllowWorkloadKubernetesAPIAccess | false | Allows workload pods to access the Kubernetes API. Required for First Class Operator support |
| DynamoOperatorSupport | false | Enables First Class Operator support. The operator must be installed in the cluster and NVCA's validation policy configured with its CRD types. **Note:** enabling this flag automatically enables `AllowWorkloadKubernetesAPIAccess` |
Loading