Skip to content

Refactor Coralogix tests to cloud-only with simplified utilities - #1440

Closed
aantn wants to merge 3 commits into
masterfrom
claude/coralogix-label-settings-ji5Y5
Closed

aantn wants to merge 3 commits into
masterfrom
claude/coralogix-label-settings-ji5Y5

Conversation

@aantn

@aantn aantn commented Jan 28, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Refactored Coralogix integration tests from Kubernetes-based deployments to lightweight cloud-only tests that directly interact with Coralogix APIs. This simplifies test setup, reduces infrastructure requirements, and improves test reliability.

Key Changes

Test Infrastructure Simplification

  • Removed Kubernetes deployments: Eliminated payment-service.yaml, traffic-generator.yaml, and related Kubernetes manifests that required cluster setup and pod orchestration
  • Cloud-only approach: Tests now send data directly to Coralogix via REST/OTLP APIs without needing a running Kubernetes cluster
  • Reduced setup time: Setup timeout reduced from 900s to 300s by eliminating pod scheduling delays

New Shared Utilities

  • coralogix_test_utils.sh: Bash utility library with functions for:

    • Environment validation (cx_validate_env)
    • Ingestion and query endpoint resolution
    • Log sending via REST API (cx_send_logs)
    • DataPrime query execution (cx_query)
    • Log availability polling (cx_wait_for_logs)
    • Timestamp utilities for Coralogix format (milliseconds since epoch)
  • send_traces.py: Python script to send test traces via OTLP gRPC with:

    • Realistic multi-span trace generation
    • Error injection for anti-hallucination testing
    • Configurable trace IDs and error codes
  • send_metrics.py: Python script to send test metrics via OTLP gRPC with:

    • Multiple metric types (counters, histograms)
    • Realistic endpoint and status code attributes
    • Configurable metric prefixes

Test Case Updates

Test 173 (Coralogix Logs):

  • Simplified to direct REST API log ingestion
  • Injects 3 specific error codes (ERR-7291, ERR-4058, ERR-9463) for anti-hallucination verification
  • Validates logs are queryable before test execution
  • User prompt focuses on finding specific error codes via DataPrime queries

Test 174 (Coralogix Traces) (new):

  • Sends traces via OTLP gRPC using send_traces.py
  • Injects specific error code (ERR-5847) and trace ID
  • Tests DataPrime span queries
  • Validates failed span detection

Code Cleanup

  • Removed unused classes from holmes/plugins/toolsets/coralogix/utils.py:
    • FlattenedLog, CoralogixQueryResult, CoralogixLabelsConfig (no longer needed)
    • extract_field, flatten_structured_log_entries, stringify_flattened_logs, parse_json_objects (unused helper functions)
  • Simplified CoralogixConfig: Removed labels field (not used by DataPrime toolset)
  • Updated documentation: Clarified that labels field is not supported in DataPrime toolset config

Test Passthrough Configuration

  • Updated conftest.py to allow all Coralogix API calls across regions:
    • Changed from single hardcoded endpoint to regex patterns
    • Now supports .coralogix.com, .coralogix.us, and .coralogix.in domains
    • Covers both query and ingestion endpoints

Benefits

  • Faster test execution: No Kubernetes pod scheduling delays
  • Simpler maintenance: No complex YAML manifests to maintain
  • Better isolation: Each test is independent with unique app names
  • Improved reliability: Direct API calls are more predictable than pod-based tests
  • Reusable utilities: Shared bash/Python scripts can be used across multiple Coralogix tests
  • Anti-hallucination: Specific error codes and trace IDs ensure LLM must actually query Coralogix

https://claude.ai/code/session_017MjrWmKKG2sPBNQL4QcVV2

Summary by CodeRabbit

  • New Features

    • Added support for Coralogix endpoints across multiple regions (US, EU, India).
  • Bug Fixes

    • Simplified Coralogix configuration by removing the optional labels field.
  • Tests

    • Replaced Kubernetes-based test infrastructure with cloud-native test automation for improved reliability and faster execution.

✏️ Tip: You can customize this high-level summary in your review settings.

The CoralogixLabelsConfig and related flatten functions were never
actually used - the API returns raw JSON that is cleaned up directly
by _cleanup_coralogix_results in api.py without any label-based
flattening.

Removed:
- FlattenedLog, CoralogixQueryResult, CoralogixLabelsConfig classes
- labels field from CoralogixConfig
- extract_field, flatten_structured_log_entries, stringify_flattened_logs,
  parse_json_objects functions
- Related tests for extract_field
- Documentation reference to optional labels config

Signed-off-by: Claude <noreply@anthropic.com>
Replace the slow Kubernetes-based Coralogix evals (15+ min setup) with
a fast cloud-only approach that writes logs directly via REST API.

Changes:
- Replace 173_coralogix_logs: Now sends logs via REST API to Coralogix
  ingestion endpoint, no K8s required (~5 min setup including indexing)
- Delete 174_coralogix_traces_ad and 175_coralogix_metrics_frontend:
  These required complex K8s+OTLP setup. Can be re-added later using
  OTLP HTTP endpoints if needed.
- Add shared/coralogix/coralogix_test_utils.sh: Utilities for sending
  logs and querying Coralogix DataPrime API
- Update conftest.py: Allow all Coralogix domains (not just eu2)
- Delete unused coralogix_app.py and traffic_generator.py

The new 173 test uses anti-hallucination error codes (ERR-7291, ERR-4058,
ERR-9463) that the LLM can only find by actually querying Coralogix.

Signed-off-by: Claude <noreply@anthropic.com>
Add Python-based scripts and eval tests for Coralogix traces and metrics
that send data directly via OTLP gRPC (no K8s required):

- 174_coralogix_traces: Sends traces via OTLP, queries via DataPrime
  - Uses send_traces.py with OpenTelemetry SDK
  - Injects ERR-5847 error code for anti-hallucination

- 175_coralogix_metrics: Sends metrics via OTLP, queries via Prometheus API
  - Uses send_metrics.py with OpenTelemetry SDK
  - Creates eval175_* prefixed metrics for anti-hallucination

Both tests complete in ~5 minutes (vs 15+ min for old K8s-based tests).

Signed-off-by: Claude <noreply@anthropic.com>
@netlify

netlify Bot commented Jan 28, 2026 •

Copy link
Copy Markdown

✅ Deploy Preview for holmes-docs ready!

Name Link
🔨 Latest commit 2d06ddd
🔍 Latest deploy log https://app.netlify.com/projects/holmes-docs/deploys/6979bd3355c817000830284c
😎 Deploy Preview https://deploy-preview-1440--holmes-docs.netlify.app
📱 Preview on mobile
Toggle QR Code...

QR Code

Use your smartphone camera to open QR code link.

To edit notification comments on pull requests, go to your Netlify project configuration.

@linux-foundation-easycla

Copy link
Copy Markdown

CLA Not Signed

@coderabbitai

coderabbitai Bot commented Jan 28, 2026 •

Copy link
Copy Markdown
Contributor

Walkthrough

The changes migrate Coralogix test infrastructure from Kubernetes-based deployments to a cloud-only approach. This includes removing in-cluster services (payment, traffic generators), updating the Coralogix utils module by removing labels-related types and helper functions, and introducing shell utilities and Python scripts for direct Coralogix API interaction via REST and OTLP protocols.

Changes

Cohort / File(s) Summary
Configuration & API Changes
conftest.py, holmes/plugins/toolsets/coralogix/utils.py, docs/data-sources/builtin-toolsets/coralogix-logs.md
Expanded Coralogix allowlist from single static URL to three regex-based regional entries (coralogix.com, coralogix.us, coralogix.in); removed public types FlattenedLog, CoralogixQueryResult, CoralogixLabelsConfig and labels field from CoralogixConfig; removed helper functions extract_field, flatten_structured_log_entries, stringify_flattened_logs, parse_json_objects; updated documentation to remove optional labels field from configuration
Kubernetes Resources Removed
tests/llm/fixtures/test_ask_holmes/173_coralogix_logs/{payment-service.yaml, traffic-generator.yaml}, tests/llm/fixtures/test_ask_holmes/174_coralogix_traces_ad/{ad-service.yaml, ad-traffic-generator.yaml}, tests/llm/fixtures/test_ask_holmes/175_coralogix_metrics_frontend/{frontend-service.yaml, frontend-traffic-generator.yaml}
Removed Deployment and Service manifests for in-cluster payment, traffic generator, ad service, and frontend services; eliminates Kubernetes-based test infrastructure
Shared Test Infrastructure Removed
tests/llm/fixtures/shared/coralogix/{coralogix_app.py, traffic_generator.py}
Deleted Flask payment service application and traffic generator script that were previously deployed in Kubernetes
Cloud-based Test Utilities Added
tests/llm/fixtures/shared/coralogix/{coralogix_test_utils.sh, send_metrics.py, send_traces.py}
Introduced shell utility functions for Coralogix interaction (env validation, log sending, DataPrime queries, readiness checks); added Python scripts to send test metrics and traces via OTLP gRPC to Coralogix with proper Resource configuration and export handling
Test Case Reorganization
tests/llm/fixtures/test_ask_holmes/173_coralogix_logs/test_case.yaml, tests/llm/fixtures/test_ask_holmes/173_coralogix_logs/toolsets.yaml, tests/llm/fixtures/test_ask_holmes/174_coralogix_traces/{test_case.yaml, toolsets.yaml} (new), tests/llm/fixtures/test_ask_holmes/175_coralogix_metrics/{test_case.yaml, toolsets.yaml} (new), tests/llm/fixtures/test_ask_holmes/174_coralogix_traces_ad/test_case.yaml (removed), tests/llm/fixtures/test_ask_holmes/175_coralogix_metrics_frontend/test_case.yaml (removed)
Converted test cases from Kubernetes-based multi-service deployments to cloud-only Coralogix REST/OTLP flows with direct log, trace, and metric ingestion; replaced pod lifecycle management with Coralogix API polling for data availability; removed Kubernetes toolsets from configurations
Plugin Tests
tests/plugins/toolsets/coralogix/test_coralogix.py
Removed usage of extract_field function and CoralogixLabelsConfig; removed labels field from coralogix_config fixture and deleted corresponding test case

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

Suggested reviewers

  • moshemorad
  • arikalon1
🚥 Pre-merge checks | ✅ 2 | ❌ 1
❌ Failed checks (1 warning)
Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 76.47% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and specifically summarizes the main refactoring effort: migrating Coralogix tests from Kubernetes-based to cloud-only infrastructure with simplified utility functions.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.


Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 5

Caution

Some comments are outside the diff and can’t be posted inline due to platform limitations.

⚠️ Outside diff range comments (1)
holmes/plugins/toolsets/coralogix/utils.py (1)

18-35: Avoid logging full log lines on JSON decode errors.

line may contain user data from Coralogix. Logging it verbatim risks PII leakage. Prefer logging the line number and length instead.

🔒 Suggested update
-    for line in raw_text.strip().split("\n"):  # Split by newlines
+    for idx, line in enumerate(raw_text.strip().split("\n"), start=1):  # Split by newlines
         try:
             obj = json.loads(line)
             if isinstance(obj, dict):
                 # Remove userData from top level
                 obj.pop("userData", None)
                 # Remove userData from direct child dicts (one level deep, no recursion)
                 for key, value in list(obj.items()):
                     if isinstance(value, dict):
                         value.pop("userData", None)
                     elif isinstance(value, list):
                         for item in value:
                             if isinstance(item, dict):
                                 item.pop("userData", None)
             json_objects.append(obj)
         except json.JSONDecodeError:
-            logging.error(f"Failed to decode JSON from line: {line}")
+            logging.error("Failed to decode JSON on line %d (len=%d)", idx, len(line))
🤖 Fix all issues with AI agents
In `@tests/llm/fixtures/shared/coralogix/send_metrics.py`:
- Around line 30-33: The Coralogix integration tests (e.g., test
175_coralogix_metrics referenced in
tests/llm/fixtures/shared/coralogix/send_metrics.py) require OpenTelemetry
packages but they are missing from poetry dev deps; add opentelemetry-api,
opentelemetry-sdk, and opentelemetry-exporter-otlp-proto-grpc to the
[tool.poetry.group.dev.dependencies] section of pyproject.toml (specify
appropriate compatible versions or leave unpinned per project convention) so the
ImportError branch in send_metrics.py no longer triggers and Coralogix
metric/trace tests can run.

In `@tests/llm/fixtures/test_ask_holmes/174_coralogix_traces/test_case.yaml`:
- Around line 18-21: The test's expected_output entry uses an ambiguous phrase
"external_api_call or payment gateway" which makes the assertion
non-deterministic; update the expected_output list to assert the exact operation
name emitted by the trace fixture (replace the item "Must mention that the
external_api_call or payment gateway operation failed" with a specific string
containing the exact operation identifier/noun used in the trace), ensuring the
check references the concrete operation name found in the fixture so the test
deterministically matches the trace.
- Around line 16-67: The test hard-codes APP_NAME as "holmes-eval-174" and
embeds that literal in the DataPrime query; change APP_NAME to follow the
required namespace convention (e.g., "app-174" or "app-${TEST_ID}") and update
the query construction (the curl -d "{\"query\": \"source spans | lucene
'holmes-eval-174 AND $ERROR_CODE' ...}") to interpolate the APP_NAME variable
instead of the literal; ensure TRACE_ID/ERROR_CODE usage remains unchanged and
use a neutral app name format to avoid leaking test intent and to prevent
parallel-test collisions.

In `@tests/llm/fixtures/test_ask_holmes/175_coralogix_metrics/test_case.yaml`:
- Around line 18-21: The expected_output YAML uses generic patterns that allow
hallucination; update the expected_output entries in the test_case.yaml to
assert the exact, deterministic metric names and values produced by the test
setup (replace the three generic bullets under expected_output with concrete
strings that include the actual metric names like the exact eval175_* metric
identifiers and exact endpoint names and counts/error-rates from your test
data), ensuring the assertions match the setup fixture values so the model must
query Coralogix rather than invent results; specifically edit the
expected_output block referenced in the test_case.yaml to use those explicit,
discoverable values.
- Around line 16-40: Update the test to use the required namespace and neutral
prefix: replace APP_NAME ("holmes-eval-175") with the app-<testid> format (e.g.,
APP_NAME="app-175") and change METRIC_PREFIX ("eval175") to a neutral
business-context name (e.g., METRIC_PREFIX="service175" or "app175"); then
update user_prompt and expected_output to reference the new METRIC_PREFIX and
APP_NAME (ensure strings like "eval175_" and metric names
"eval175_http_requests_total" are replaced accordingly) so the prompt and
assertions remain consistent with the renamed symbols.
🧹 Nitpick comments (8)
tests/llm/fixtures/shared/coralogix/coralogix_test_utils.sh (3)

49-57: Separate declaration from assignment to avoid masking errors.

The local declaration combined with command substitution masks the return value. If cx_ingress_url or cat fails, the failure is hidden.

♻️ Proposed fix
-  local ingress_url=$(cx_ingress_url)
-  local payload=$(cat <<EOF
+  local ingress_url
+  ingress_url=$(cx_ingress_url)
+  local payload
+  payload=$(cat <<EOF

78-89: Same pattern: separate local declaration from assignment.

Line 83 has the same issue where local query_url=$(cx_query_url) masks potential failures.

♻️ Proposed fix
-  local query_url=$(cx_query_url)
+  local query_url
+  query_url=$(cx_query_url)

100-110: Separate declaration in the loop for proper error detection.

♻️ Proposed fix
   for i in $(seq 1 $max_attempts); do
-    local result=$(cx_query "source logs | lucene '$search_term' | limit 1")
+    local result
+    result=$(cx_query "source logs | lucene '$search_term' | limit 1")
tests/llm/fixtures/test_ask_holmes/173_coralogix_logs/test_case.yaml (1)

55-56: Consider extracting log entries to a separate JSON file for maintainability.

The inline JSON string is difficult to read and modify. While the comment explains YAML heredoc compatibility, an alternative would be storing the log entries template in a separate .json file and using envsubst or sed to inject the error codes and timestamps.

tests/llm/fixtures/test_ask_holmes/175_coralogix_metrics/test_case.yaml (1)

34-53: set -e prevents the custom failure message from running.

With set -e, a failing python3 exits before the $? check. Wrap the command in if ! ...; then to keep the explicit error output.

♻️ Suggested update
-  # Send metrics using Python OTLP exporter
-  python3 ../../shared/coralogix/send_metrics.py \
-    --app-name "$APP_NAME" \
-    --subsystem "$SUBSYSTEM" \
-    --metric-prefix "$METRIC_PREFIX"
-
-  if [ $? -ne 0 ]; then
+  # Send metrics using Python OTLP exporter
+  if ! python3 ../../shared/coralogix/send_metrics.py \
+    --app-name "$APP_NAME" \
+    --subsystem "$SUBSYSTEM" \
+    --metric-prefix "$METRIC_PREFIX"; then
     echo "❌ Failed to send metrics"
     exit 1
   fi
tests/llm/fixtures/test_ask_holmes/174_coralogix_traces/test_case.yaml (1)

33-54: set -e prevents the custom failure message from running.

With set -e, a failing python3 exits before the $? check. Wrap the command in if ! ...; then to keep the explicit error output.

♻️ Suggested update
-  # Send traces using Python OTLP exporter
-  python3 ../../shared/coralogix/send_traces.py \
-    --app-name "$APP_NAME" \
-    --subsystem "$SUBSYSTEM" \
-    --trace-id "$TRACE_ID" \
-    --error-code "$ERROR_CODE"
-
-  if [ $? -ne 0 ]; then
+  # Send traces using Python OTLP exporter
+  if ! python3 ../../shared/coralogix/send_traces.py \
+    --app-name "$APP_NAME" \
+    --subsystem "$SUBSYSTEM" \
+    --trace-id "$TRACE_ID" \
+    --error-code "$ERROR_CODE"; then
     echo "❌ Failed to send traces"
     exit 1
   fi
holmes/plugins/toolsets/coralogix/utils.py (2)

6-13: Allow legacy config fields for backward compatibility.

With labels removed, older configs may still pass that field. Setting extra="allow" keeps those configs accepted without reintroducing deprecated fields.

♻️ Suggested update
-from pydantic import BaseModel
+from pydantic import BaseModel, ConfigDict

 class CoralogixConfig(BaseModel):
+    model_config = ConfigDict(extra="allow")
     team_hostname: str
     domain: str
     api_key: str
As per coding guidelines: holmes/plugins/toolsets/**/*.py : Maintain backwards compatibility for renamed toolset config fields using Pydantic's 'extra="allow"' instead of defining deprecated fields in schema.

15-15: Add the missing type hint for raw_text.

♻️ Suggested update
-def parse_json_lines(raw_text) -> List[Dict[str, Any]]:
+def parse_json_lines(raw_text: str) -> List[Dict[str, Any]]:
As per coding guidelines: **/*.py : Use type hints throughout Python code and run 'mypy' for type checking.

Comment on lines +30 to +33
except ImportError:
print("ERROR: OpenTelemetry packages not installed. Run:")
print(" pip install opentelemetry-api opentelemetry-sdk opentelemetry-exporter-otlp-proto-grpc")
return False

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🔴 Critical

🧩 Analysis chain

🏁 Script executed:

# Check pyproject.toml for OpenTelemetry dependencies
echo "=== Checking pyproject.toml for OpenTelemetry ==="
rg -n "opentelemetry" pyproject.toml --color never || echo "No OpenTelemetry found in pyproject.toml"

# Check for requirements files
echo ""
echo "=== Checking for requirements files ==="
fd -t f -e txt -e in | grep -i req | head -20

# Check the actual test file to understand context
echo ""
echo "=== Examining send_metrics.py ==="
cat -n tests/llm/fixtures/shared/coralogix/send_metrics.py 2>/dev/null | head -50

# Check pyproject.toml structure for optional dependencies
echo ""
echo "=== Checking pyproject.toml structure ==="
rg -A 10 "optional-dependencies|dev|test" pyproject.toml | head -60

Repository: HolmesGPT/holmesgpt

Length of output: 5038


🏁 Script executed:

# Search for usage of send_metrics.py in the test suite
echo "=== Searching for usage of send_metrics.py ==="
rg "send_metrics" --type py | head -20

# Check if there's any documentation for this fixture/utility
echo ""
echo "=== Checking for documentation in coralogix directory ==="
find tests/llm/fixtures/shared/coralogix -type f \( -name "*.md" -o -name "*.txt" -o -name "README*" \)

# Check if there's a conftest or __init__ that might run this
echo ""
echo "=== Checking for conftest.py or __init__ that might import send_metrics ==="
rg -l "conftest|__init__" tests/llm/fixtures/shared/coralogix/

# See all files in the coralogix directory
echo ""
echo "=== Files in coralogix fixture directory ==="
find tests/llm/fixtures/shared/coralogix -type f

Repository: HolmesGPT/holmesgpt

Length of output: 785


🏁 Script executed:

# Check coralogix_test_utils.sh to see if it calls send_metrics
echo "=== Content of coralogix_test_utils.sh ==="
cat tests/llm/fixtures/shared/coralogix/coralogix_test_utils.sh

# Check if send_metrics.py has a main block
echo ""
echo "=== Checking for __main__ block in send_metrics.py ==="
grep -n "__main__" tests/llm/fixtures/shared/coralogix/send_metrics.py

# Check if this utility is referenced in any test files or CI configs
echo ""
echo "=== Searching for coralogix references in test files ==="
rg "coralogix|Coralogix" tests/ --type py | grep -v "Binary" | head -20

# Check CI/workflow files
echo ""
echo "=== Checking workflow/CI files for coralogix ==="
fd -t f -e yml -e yaml | xargs rg "coralogix|send_metrics" 2>/dev/null || echo "No coralogix references in CI configs"

Repository: HolmesGPT/holmesgpt

Length of output: 7908


🏁 Script executed:

# Check how test cases are structured and if there are conditional test runs
echo "=== Checking test case structure for Coralogix tests ==="
cat tests/llm/fixtures/test_ask_holmes/175_coralogix_metrics/test_case.yaml

# Check if there's any test runner documentation
echo ""
echo "=== Checking for test runner or pytest configuration ==="
find . -maxdepth 2 -name "pytest.ini" -o -name "conftest.py" | xargs cat 2>/dev/null | head -100

# Check if there's documentation about optional dependencies
echo ""
echo "=== Checking for documentation about optional/conditional tests ==="
fd -t f -e md | xargs rg -l "coralogix|optional.*test|test.*dependency" 2>/dev/null | head -10

Repository: HolmesGPT/holmesgpt

Length of output: 6748


Add OpenTelemetry dependencies to pyproject.toml for Coralogix integration tests.

The test case 175_coralogix_metrics (and similar Coralogix tests) document that "OpenTelemetry Python packages (installed via poetry)" are required. These dependencies are currently missing from pyproject.toml:

  • opentelemetry-api
  • opentelemetry-sdk
  • opentelemetry-exporter-otlp-proto-grpc

Add these to the [tool.poetry.group.dev.dependencies] section so the Coralogix metric/trace tests can run properly.

🤖 Prompt for AI Agents
In `@tests/llm/fixtures/shared/coralogix/send_metrics.py` around lines 30 - 33,
The Coralogix integration tests (e.g., test 175_coralogix_metrics referenced in
tests/llm/fixtures/shared/coralogix/send_metrics.py) require OpenTelemetry
packages but they are missing from poetry dev deps; add opentelemetry-api,
opentelemetry-sdk, and opentelemetry-exporter-otlp-proto-grpc to the
[tool.poetry.group.dev.dependencies] section of pyproject.toml (specify
appropriate compatible versions or leave unpinned per project convention) so the
ImportError branch in send_metrics.py no longer triggers and Coralogix
metric/trace tests can run.

Comment on lines +16 to +67
user_prompt: "Search the Coralogix traces for the 'holmes-eval-174' application. Look for failed spans and report what error codes you find. Also identify which operation failed."

expected_output:
- "Must identify that there are error/failed spans"
- "Must find the error code ERR-5847"
- "Must mention that the external_api_call or payment gateway operation failed"

tags:
- coralogix
- traces
- medium

setup_timeout: 300

before_test: |
source ../../shared/coralogix/coralogix_test_utils.sh
cx_setup
set -e

echo "🚀 Setting up Coralogix traces test 174"

APP_NAME="holmes-eval-174"
SUBSYSTEM="checkout-service"
TRACE_ID="TRACE-174-$(date +%s)"
ERROR_CODE="ERR-5847"

echo "⏳ Sending traces to Coralogix via OTLP..."

# Send traces using Python OTLP exporter
python3 ../../shared/coralogix/send_traces.py \
--app-name "$APP_NAME" \
--subsystem "$SUBSYSTEM" \
--trace-id "$TRACE_ID" \
--error-code "$ERROR_CODE"

if [ $? -ne 0 ]; then
echo "❌ Failed to send traces"
exit 1
fi

echo "✅ Traces sent to $APP_NAME/$SUBSYSTEM"

# Wait for traces to be queryable
echo "⏳ Waiting for traces to be queryable in Coralogix..."
QUERY_URL="https://ng-api-http.${CORALOGIX_DOMAIN}/api/v1/dataprime/query"

TRACES_READY=false
for i in {1..60}; do
RESULT=$(curl -sf -X POST "$QUERY_URL" \
-H "Authorization: Bearer ${CORALOGIX_API_KEY}" \
-H "Content-Type: application/json" \
-d "{\"query\": \"source spans | lucene 'holmes-eval-174 AND $ERROR_CODE' | limit 1\", \"metadata\": {\"syntax\": \"QUERY_SYNTAX_DATAPRIME\"}}" 2>/dev/null)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Align APP_NAME with namespace convention and avoid hard-coding it in the query.

Use the required app-<testid> format and interpolate the app name in the DataPrime query so it stays consistent when renaming.

♻️ Suggested update
-user_prompt: "Search the Coralogix traces for the 'holmes-eval-174' application. Look for failed spans and report what error codes you find. Also identify which operation failed."
+user_prompt: "Search the Coralogix traces for the 'app-174' application. Look for failed spans and report what error codes you find. Also identify which operation failed."

-  APP_NAME="holmes-eval-174"
+  APP_NAME="app-174"

-    -d "{\"query\": \"source spans | lucene 'holmes-eval-174 AND $ERROR_CODE' | limit 1\", \"metadata\": {\"syntax\": \"QUERY_SYNTAX_DATAPRIME\"}}" 2>/dev/null)
+    -d "{\"query\": \"source spans | lucene '${APP_NAME} AND $ERROR_CODE' | limit 1\", \"metadata\": {\"syntax\": \"QUERY_SYNTAX_DATAPRIME\"}}" 2>/dev/null)
Based on learnings: Applies to tests/llm/**/*.yaml : Use dedicated namespace per test with format 'app-' (e.g., 'app-177') to prevent conflicts in parallel execution; Never use obvious or hint-giving resource names in tests - use neutral, business-context names instead of technical indicators.
🤖 Prompt for AI Agents
In `@tests/llm/fixtures/test_ask_holmes/174_coralogix_traces/test_case.yaml`
around lines 16 - 67, The test hard-codes APP_NAME as "holmes-eval-174" and
embeds that literal in the DataPrime query; change APP_NAME to follow the
required namespace convention (e.g., "app-174" or "app-${TEST_ID}") and update
the query construction (the curl -d "{\"query\": \"source spans | lucene
'holmes-eval-174 AND $ERROR_CODE' ...}") to interpolate the APP_NAME variable
instead of the literal; ensure TRACE_ID/ERROR_CODE usage remains unchanged and
use a neutral app name format to avoid leaking test intent and to prevent
parallel-test collisions.

Comment on lines +18 to +21
expected_output:
- "Must identify that there are error/failed spans"
- "Must find the error code ERR-5847"
- "Must mention that the external_api_call or payment gateway operation failed"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Make the failed operation expectation specific.

“external_api_call or payment gateway” is ambiguous. Please assert the exact operation name emitted by the trace fixture so the check is deterministic.

Based on learnings: Applies to tests/llm/fixtures/test_ask_holmes/**/test_case.yaml : Use specific, discoverable values in expected_output instead of generic patterns to rule out hallucinations.

🤖 Prompt for AI Agents
In `@tests/llm/fixtures/test_ask_holmes/174_coralogix_traces/test_case.yaml`
around lines 18 - 21, The test's expected_output entry uses an ambiguous phrase
"external_api_call or payment gateway" which makes the assertion
non-deterministic; update the expected_output list to assert the exact operation
name emitted by the trace fixture (replace the item "Must mention that the
external_api_call or payment gateway operation failed" with a specific string
containing the exact operation identifier/noun used in the trace), ensuring the
check references the concrete operation name found in the fixture so the test
deterministically matches the trace.

Comment on lines +16 to +40
user_prompt: "Query the Coralogix metrics for the 'holmes-eval-175' application. Look for HTTP request metrics with prefix 'eval175_' and tell me: what metrics are available, how many total requests were made, and what endpoints have the most errors."

expected_output:
- "Must identify metrics with the eval175_ prefix"
- "Must mention eval175_http_requests_total, eval175_http_errors_total, or eval175_request_latency_seconds"
- "Must provide information about request counts or error rates by endpoint"

tags:
- coralogix
- metrics
- prometheus
- medium

setup_timeout: 300

before_test: |
source ../../shared/coralogix/coralogix_test_utils.sh
cx_setup
set -e

echo "🚀 Setting up Coralogix metrics test 175"

APP_NAME="holmes-eval-175"
SUBSYSTEM="api-gateway"
METRIC_PREFIX="eval175"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Align APP_NAME/METRIC_PREFIX with namespace + neutral naming.

holmes-eval-175 / eval175 are explicit test identifiers and don’t follow the app-<testid> convention. Please switch to the required namespace format and a neutral business-context prefix, then update the prompt/expectations to match.

♻️ Suggested update
-user_prompt: "Query the Coralogix metrics for the 'holmes-eval-175' application. Look for HTTP request metrics with prefix 'eval175_' and tell me: what metrics are available, how many total requests were made, and what endpoints have the most errors."
+user_prompt: "Query the Coralogix metrics for the 'app-175' application. Look for HTTP request metrics with prefix 'billing175_' and tell me: what metrics are available, how many total requests were made, and what endpoints have the most errors."

 expected_output:
-  - "Must identify metrics with the eval175_ prefix"
-  - "Must mention eval175_http_requests_total, eval175_http_errors_total, or eval175_request_latency_seconds"
+  - "Must identify metrics with the billing175_ prefix"
+  - "Must mention billing175_http_requests_total, billing175_http_errors_total, or billing175_request_latency_seconds"

   APP_NAME="holmes-eval-175"
   SUBSYSTEM="api-gateway"
-  METRIC_PREFIX="eval175"
+  APP_NAME="app-175"
+  METRIC_PREFIX="billing175"
Based on learnings: Applies to tests/llm/**/*.yaml : Use dedicated namespace per test with format 'app-' (e.g., 'app-177') to prevent conflicts in parallel execution; Never use obvious or hint-giving resource names in tests - use neutral, business-context names instead of technical indicators.
🤖 Prompt for AI Agents
In `@tests/llm/fixtures/test_ask_holmes/175_coralogix_metrics/test_case.yaml`
around lines 16 - 40, Update the test to use the required namespace and neutral
prefix: replace APP_NAME ("holmes-eval-175") with the app-<testid> format (e.g.,
APP_NAME="app-175") and change METRIC_PREFIX ("eval175") to a neutral
business-context name (e.g., METRIC_PREFIX="service175" or "app175"); then
update user_prompt and expected_output to reference the new METRIC_PREFIX and
APP_NAME (ensure strings like "eval175_" and metric names
"eval175_http_requests_total" are replaced accordingly) so the prompt and
assertions remain consistent with the renamed symbols.

Comment on lines +18 to +21
expected_output:
- "Must identify metrics with the eval175_ prefix"
- "Must mention eval175_http_requests_total, eval175_http_errors_total, or eval175_request_latency_seconds"
- "Must provide information about request counts or error rates by endpoint"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

⚠️ Potential issue | 🟡 Minor

Make expected_output assertions deterministic.

The current bullets allow plausible fabrication. Please pin the assertions to concrete values produced during setup (e.g., exact endpoint names and counts), so the model must actually query Coralogix.

Based on learnings: Applies to tests/llm/fixtures/test_ask_holmes/**/test_case.yaml : Use specific, discoverable values in expected_output instead of generic patterns to rule out hallucinations.

🤖 Prompt for AI Agents
In `@tests/llm/fixtures/test_ask_holmes/175_coralogix_metrics/test_case.yaml`
around lines 18 - 21, The expected_output YAML uses generic patterns that allow
hallucination; update the expected_output entries in the test_case.yaml to
assert the exact, deterministic metric names and values produced by the test
setup (replace the three generic bullets under expected_output with concrete
strings that include the actual metric names like the exact eval175_* metric
identifiers and exact endpoint names and counts/error-rates from your test
data), ensuring the assertions match the setup fixture values so the model must
query Coralogix rather than invent results; specifically edit the
expected_output block referenced in the test_case.yaml to use those explicit,
discoverable values.

@aantn aantn closed this Jan 31, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants