Skip to content

feat(ci): add BFCL function calling accuracy tests to nightly and PR CI - #520

Closed
vschandramourya wants to merge 1 commit into
mainfrom
tool-calling-tests
Closed

vschandramourya wants to merge 1 commit into
mainfrom
tool-calling-tests

Conversation

@vschandramourya

@vschandramourya vschandramourya commented Feb 24, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Add end-to-end BFCL v3 (Berkeley Function Calling Leaderboard) test infrastructure and integrate it into both nightly and PR CI workflows for tool-calling accuracy validation.

What's included

  • BFCL test infrastructure (e2e_test/bfcl/): data loader, evaluator, per-test JSON logging, and HuggingFace data downloader for 1,240 test cases across 5 categories (simple, multiple, parallel, parallel_multiple, irrelevance)
  • Baseline comparison framework (e2e_test/baselines/): stores accuracy baselines per model/backend, detects regressions with configurable tolerance (default ±3.0pp)
  • Test file (e2e_test/chat_completions/test_bfcl.py): pytest classes for each category with fixture-based backend setup + standalone mode

CI integration

CI Mode Scope Cases Runtimes
Nightly (nightly-benchmark.yml) Full BFCL suite 1,240 (all categories) SGLang
PR (pr-test-rust.yml) Lighter BFCL 100 (20 per category via BFCL_LIMIT=20) SGLang + vLLM

Nightly additions

  • New bfcl-accuracy job on 4×H100 running all 1,240 BFCL test cases with Qwen2.5-7B-Instruct
  • Results (summary.json, comparison.json, detailed per-test logs) uploaded as artifacts
  • BFCL section added to nightly_summarize.py with per-category accuracy table and baseline regression tracking
  • summarize-benchmarks job updated to depend on bfcl-accuracy

PR additions

  • BFCL_LIMIT=20 added to chat-completions-sglang and chat-completions-vllm env vars (20 cases × 5 categories = 100 tests per runtime)
  • New download_bfcl matrix flag + "Download BFCL test data" conditional step
  • BFCL runs as part of existing chat-completions-* jobs — no additional GPU allocation needed

Nightly summary rendering

  • Runtime comparison tables (SGLang vs vLLM) for performance benchmarks
  • BFCL accuracy section with per-category breakdown and baseline comparison (regressions/improvements)

Test plan

  • Verify BFCL_LIMIT=20 produces exactly 20 tests per category (100 total) in PR runs
  • Verify nightly bfcl-accuracy job runs all 1,240 cases on SGLang
  • Verify BFCL data download step runs only for chat-completions-sglang and chat-completions-vllm
  • Verify BFCL results appear in nightly summary markdown
  • Verify TRT-LLM correctly skips BFCL tests via skip_for_runtime marker

Made with Cursor

Summary by CodeRabbit

New Features

  • Added Berkeley Function Calling Leaderboard (BFCL) support to nightly benchmarks with automated accuracy testing across multiple categories (simple, multiple, parallel, irrelevance)
  • Introduced baseline comparison and regression detection for benchmark results to track performance changes over time
  • Extended CI/CD workflows to include BFCL test data downloads and evaluation

Tests

  • Added comprehensive BFCL end-to-end test suite with detailed result logging and summary reporting

Add end-to-end BFCL v3 (Berkeley Function Calling Leaderboard) test
infrastructure with 1,240 test cases across 5 categories (simple,
multiple, parallel, parallel_multiple, irrelevance) and integrate it
into both nightly and PR CI workflows.

Nightly CI: full BFCL suite (all 1,240 cases) runs on SGLang with
Qwen2.5-7B, results flow into the nightly summary with per-category
accuracy tables and baseline regression detection.

PR CI: lighter BFCL (20 cases per category = 100 total) runs on both
SGLang and vLLM via BFCL_LIMIT=20, providing fast function-calling
smoke coverage on every pull request.

Co-authored-by: Cursor <cursoragent@cursor.com>
@github-actions github-actions Bot added ci CI/CD configuration changes tests Test changes labels Feb 24, 2026
@coderabbitai

coderabbitai Bot commented Feb 24, 2026 •

Copy link
Copy Markdown

Caution

Review failed

The pull request is closed.

ℹ️ Recent review info

Configuration used: Organization UI

Review profile: ASSERTIVE

Plan: Pro

📥 Commits

Reviewing files that changed from the base of the PR and between fbbb99b and a0dd21e.

📒 Files selected for processing (19)
  • .github/workflows/nightly-benchmark.yml
  • .github/workflows/pr-test-rust.yml
  • e2e_test/baselines/__init__.py
  • e2e_test/baselines/compare.py
  • e2e_test/benchmarks/nightly_summarize.py
  • e2e_test/bfcl/__init__.py
  • e2e_test/bfcl/data/BFCL_v3_irrelevance.json
  • e2e_test/bfcl/data/BFCL_v3_multiple.json
  • e2e_test/bfcl/data/BFCL_v3_multiple_answer.json
  • e2e_test/bfcl/data/BFCL_v3_parallel.json
  • e2e_test/bfcl/data/BFCL_v3_parallel_answer.json
  • e2e_test/bfcl/data/BFCL_v3_parallel_multiple.json
  • e2e_test/bfcl/data/BFCL_v3_parallel_multiple_answer.json
  • e2e_test/bfcl/data/BFCL_v3_simple.json
  • e2e_test/bfcl/data/BFCL_v3_simple_answer.json
  • e2e_test/bfcl/download_data.py
  • e2e_test/bfcl/evaluator.py
  • e2e_test/bfcl/loader.py
  • e2e_test/chat_completions/test_bfcl.py

📝 Walkthrough

Walkthrough

This PR introduces comprehensive BFCL (Berkeley Function Calling Leaderboard) support to the project. It adds a nightly GPU-enabled benchmark workflow, BFCL data loader and evaluator modules, baseline comparison infrastructure for regression detection, a summary integration layer, and an end-to-end test suite covering simple/multiple/parallel/irrelevance function-calling scenarios.

Changes

Cohort / File(s) Summary
Workflow Configuration
.github/workflows/nightly-benchmark.yml, .github/workflows/pr-test-rust.yml
Added new bfcl-accuracy nightly job with GPU matrix, BFCL data download, test execution, and result aggregation; extended pr-test-rust.yml to download BFCL test data with BFCL_LIMIT environment variable in e2e matrix jobs.
Baseline Comparison Infrastructure
e2e_test/baselines/__init__.py, e2e_test/baselines/compare.py
New baseline comparison engine with deterministic file layout, tolerance-based regression detection via BASELINE_ACCURACY_TOLERANCE, and per-category/overall accuracy comparison producing Regression and ComparisonResult objects.
BFCL Core Infrastructure
e2e_test/bfcl/__init__.py, e2e_test/bfcl/download_data.py, e2e_test/bfcl/loader.py, e2e_test/bfcl/evaluator.py
BFCL data downloader from HuggingFace, category loader with OpenAI tool conversion, evaluator with per-test JSON logging and summary generation including latency statistics and failure tracking.
BFCL Test Data
e2e_test/bfcl/data/BFCL_v3_simple_answer.json, e2e_test/bfcl/data/BFCL_v3_multiple_answer.json, e2e_test/bfcl/data/BFCL_v3_parallel_answer.json, e2e_test/bfcl/data/BFCL_v3_parallel_multiple_answer.json
Added 200 entries each (800 total) with structured ground-truth function calls across diverse domains (math, finance, physics, geography, etc.) for benchmark evaluation scenarios.
Summary Integration
e2e_test/benchmarks/nightly_summarize.py
Extended SummaryResult with bfcl_results field; added BfclResult and BfclCategoryResult data classes; implemented discover_bfcl_results() and _section_bfcl() to locate, parse, and render BFCL results in nightly summary with per-category accuracy and optional baseline comparison.
BFCL Test Suite
e2e_test/chat_completions/test_bfcl.py
Comprehensive parametrized e2e test suite with session-scoped logging, five test classes (Simple, Multiple, Parallel, ParallelMultiple, Irrelevance variants for Qwen) plus standalone mode, capturing request/response payloads, latency, and optional baseline comparison.

Sequence Diagram

sequenceDiagram
    participant Workflow as GitHub Workflow
    participant DataMgr as Data Manager
    participant TestRunner as Test Runner
    participant Evaluator as Evaluator
    participant Logger as Logger/Summary
    participant Baseline as Baseline Comparison
    participant Output as Output Artifacts

    Workflow->>DataMgr: Download BFCL data
    DataMgr->>DataMgr: Fetch from HuggingFace
    Workflow->>TestRunner: Execute BFCL tests
    TestRunner->>TestRunner: Load category test cases
    TestRunner->>TestRunner: Convert to OpenAI tools
    TestRunner->>Evaluator: Evaluate tool calls
    Evaluator->>Logger: Save per-test JSON log
    Evaluator->>Logger: Compute results & latencies
    Logger->>Baseline: Load baseline (if exists)
    Baseline->>Baseline: Compare accuracy vs baseline
    Logger->>Logger: Generate summary with stats
    Logger->>Output: Upload summary & logs
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~60 minutes

Possibly related PRs

  • #327: Modifies .github/workflows/nightly-benchmark.yml and updates summarize-benchmarks job dependencies similarly to integrate new benchmark results.
  • #392: Modifies e2e_test/benchmarks/nightly_summarize.py and SummaryResult structure, directly overlapping with BFCL summary integration changes.
  • #373: Updates nightly-benchmark workflow's benchmark runner and summary wiring, related to BFCL job integration patterns.

Suggested reviewers

  • key4ng
  • CatherineSue
  • slin1237

Poem

🐰 Hop along, the functions call,
BFCL tests to measure all,
Baselines tracked, regressions caught,
Accuracy's battles fought! ✨

✨ Finishing Touches
  • 📝 Generate docstrings (stacked PR)
  • 📝 Generate docstrings (commit on current branch)
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch tool-calling-tests

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello @vschandramourya, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly enhances the project's testing capabilities by incorporating the Berkeley Function Calling Leaderboard (BFCL) v3. The primary goal is to establish robust, automated validation for tool-calling accuracy across different models and backends. By integrating these tests into both nightly and PR CI, the system gains continuous monitoring for potential regressions in function calling performance, ensuring the reliability and correctness of tool-use features.

Highlights

  • BFCL v3 Test Infrastructure: Added comprehensive end-to-end test infrastructure for Berkeley Function Calling Leaderboard (BFCL) v3, including data loaders, evaluators, per-test JSON logging, and a HuggingFace data downloader for 1,240 test cases across five categories (simple, multiple, parallel, parallel_multiple, irrelevance).
  • Baseline Comparison Framework: Introduced a new framework to store accuracy baselines per model/backend, enabling automatic detection of regressions with a configurable tolerance (default ±3.0 percentage points).
  • CI Integration for Nightly and PR Workflows: Integrated BFCL tests into both nightly and pull request CI workflows. Nightly runs the full BFCL suite (1,240 cases) on SGLang, while PR runs a lighter version (100 cases, 20 per category) on both SGLang and vLLM. This ensures continuous validation of tool-calling accuracy.
  • Nightly Summary Enhancements: Updated the nightly benchmark summarization script to include a dedicated BFCL section, displaying per-category accuracy tables and tracking baseline regressions/improvements, with results uploaded as artifacts.
  • New BFCL Test Suite: Added a new pytest file (e2e_test/chat_completions/test_bfcl.py) containing classes for each BFCL category, supporting fixture-based backend setup and a standalone execution mode.
Changelog
  • e2e_test/baselines/init.py
    • Added imports for baseline comparison utilities.
  • e2e_test/baselines/compare.py
    • Added core logic for comparing benchmark summaries against stored baselines.
    • Introduced data types for Regression and ComparisonResult to structure comparison outcomes.
    • Implemented functions to load, save, and compare baseline JSON files, including handling for missing baselines and configurable accuracy tolerances.
    • Added utility functions for normalizing model names and determining baseline file paths.
  • e2e_test/benchmarks/nightly_summarize.py
    • Added BfclCategoryResult and BfclResult data classes to represent BFCL accuracy outcomes.
    • Updated SummaryResult to include a list of bfcl_results.
    • Implemented discover_bfcl_results to locate and parse BFCL accuracy logs from nightly runs.
    • Added _section_bfcl to format and render BFCL accuracy results, including baseline comparisons, into the nightly summary markdown.
    • Modified generate_summary to discover BFCL results and incorporate them into the overall report, ensuring the summary is generated even if only BFCL results are present.
  • e2e_test/bfcl/init.py
    • Added imports for BFCL evaluator and loader functions.
  • e2e_test/bfcl/data/BFCL_v3_multiple_answer.json
    • Added BFCL v3 multiple category ground truth answer data.
  • e2e_test/bfcl/data/BFCL_v3_parallel_answer.json
    • Added BFCL v3 parallel category ground truth answer data.
  • e2e_test/bfcl/data/BFCL_v3_parallel_multiple_answer.json
    • Added BFCL v3 parallel multiple category ground truth answer data.
  • e2e_test/bfcl/data/BFCL_v3_simple_answer.json
    • Added BFCL v3 simple category ground truth answer data.
  • e2e_test/bfcl/download_data.py
    • Added a Python script to download BFCL v3 open-source test data and answer files from HuggingFace.
  • e2e_test/bfcl/evaluator.py
    • Added functions to create a timestamped run directory for BFCL logs.
    • Implemented evaluate_tool_calls to compare model output against BFCL ground truth, handling irrelevance tests and flexible argument matching.
    • Added save_test_log to write detailed JSON log files for individual BFCL test cases.
    • Implemented save_summary to generate a comprehensive JSON summary of all BFCL test results, including per-category statistics and latency distributions.
  • e2e_test/bfcl/loader.py
    • Added load_bfcl_category to load BFCL test cases and merge them with ground truth answers.
    • Implemented _fix_parameter_type to convert BFCL's non-standard JSON schema types to OpenAI-compatible formats.
    • Added bfcl_to_openai_tools to transform BFCL function definitions into the OpenAI tools format.
    • Added bfcl_messages to ensure messages are correctly formatted for the OpenAI API.
  • e2e_test/chat_completions/test_bfcl.py
    • Added BFCL E2E test classes for simple, multiple, parallel, parallel_multiple, and irrelevance categories.
    • Implemented _extract_tool_calls to parse structured tool calls from OpenAI ChatCompletion responses.
    • Introduced session-scoped fixtures (bfcl_run_dir, _write_summary_on_exit) for managing log directories and generating final summaries and baseline comparisons.
    • Created _run_bfcl_case as a core function to execute individual BFCL test cases, log results, and handle assertions.
    • Added a TestBFCLStandalone class for running tests against an externally managed gateway, configurable via environment variables.
Ignored Files
  • Ignored by pattern: .github/workflows/** (2)
    • .github/workflows/nightly-benchmark.yml
    • .github/workflows/pr-test-rust.yml
Activity
  • No human activity has been recorded on this pull request yet.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@mergify mergify Bot closed this Feb 24, 2026
@mergify

mergify Bot commented Feb 24, 2026

Copy link
Copy Markdown
Contributor

Closing this PR because the branch name tool-calling-tests does not follow the naming convention.

To fix this, please rename your branch locally, push the new branch, and open a new PR:

git branch -m <new-branch-name>
git push origin -u <new-branch-name>

@mergify

mergify Bot commented Feb 24, 2026

Copy link
Copy Markdown
Contributor

Hi @vschandramourya, the branch tool-calling-tests does not follow our naming convention.

Please use one of the following formats:

  • <type>/<description> — e.g. feat/add-auth, fix/null-pointer, dependabot/cargo/pyo3-0.28.1
  • <username>/<description> — e.g. changsu/fix-routing

Allowed types: feat, fix, chore, docs, refactor, test, ci, perf

Note: PRs with non-conforming branch names will be auto-closed. Please follow the naming convention for all branches.

@mergify

mergify Bot commented Feb 24, 2026

Copy link
Copy Markdown
Contributor

Hi @vschandramourya, the DCO sign-off check has failed. All commits must include a Signed-off-by line.

To fix existing commits:

# Sign off the last N commits (replace N with the number of unsigned commits)
git rebase HEAD~N --signoff
git push --force-with-lease

To sign off future commits automatically:

  • Use git commit -s every time, or
  • VSCode: enable Git: Always Sign Off in Settings
  • PyCharm: enable Sign-off commit in the Commit tool window

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces a comprehensive testing framework for function calling accuracy using the BFCL benchmark, which is a significant enhancement to the project's testing capabilities. The new infrastructure for baselining, evaluation, and logging is well-structured, and the integration into CI for both nightly and PR workflows is a great addition for regression testing. My feedback focuses on improving the robustness of the reporting scripts by adding more explicit error logging and ensuring consistency in the test result data collection.

Comment on lines +1243 to +1244
except Exception:
pass

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Silently ignoring exceptions when parsing metadata.json can mask problems. If the file is corrupt, it will be treated as if it doesn't exist, leading to 'unknown' values in the summary report. This can be misleading and make debugging harder. It's better to log a warning to stderr to make these issues visible.

Suggested change
except Exception:
pass
except Exception as e:
print(f"Warning: Failed to parse BFCL metadata {meta_path}: {e}", file=sys.stderr)

Comment on lines +1264 to +1265
except Exception:
pass

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

Similar to the metadata parsing, silently ignoring exceptions when parsing comparison.json can hide issues. If the file is corrupt, the baseline comparison section will be missing from the summary without any indication of why. Logging a warning would make this failure explicit and easier to debug.

Suggested change
except Exception:
pass
except Exception as e:
print(f"Warning: Failed to parse BFCL comparison data {comp_path}: {e}", file=sys.stderr)

Comment on lines +217 to +244
save_test_log(
run_dir,
test_id=test_id,
category=category,
model=model,
parser=parser,
backend=backend,
request_payload=request_payload,
response_payload=None,
ground_truth=case.get("ground_truth", []),
actual_tool_calls=[],
passed=False,
errors=[f"API error: {exc}"],
latency_ms=latency,
)
_all_results.append({
"test_id": test_id,
"category": category,
"passed": False,
"errors": [f"API error: {exc}"],
"latency_ms": latency,
"finish_reason": None,
"completion_tokens": None,
"had_reasoning": False,
"log_file": f"{category}/{test_id.replace('/', '_')}_FAIL.json",
"model": model,
"backend": backend,
})

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The log_file path stored in _all_results for failed API calls is constructed manually and is inconsistent with the success path. In the success path, log_path.name is used, which is just the filename. In this error path, a relative path including the category is constructed. To improve consistency and reduce brittleness, you should capture the returned path from save_test_log and use its .name attribute, just like in the success case.

        log_path = save_test_log(
            run_dir,
            test_id=test_id,
            category=category,
            model=model,
            parser=parser,
            backend=backend,
            request_payload=request_payload,
            response_payload=None,
            ground_truth=case.get("ground_truth", []),
            actual_tool_calls=[],
            passed=False,
            errors=[f"API error: {exc}"],
            latency_ms=latency,
        )
        _all_results.append({
            "test_id": test_id,
            "category": category,
            "passed": False,
            "errors": [f"API error: {exc}"],
            "latency_ms": latency,
            "finish_reason": None,
            "completion_tokens": None,
            "had_reasoning": False,
            "log_file": log_path.name,
            "model": model,
            "backend": backend,
        })

@CatherineSue CatherineSue left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can you split the PR into smaller ones? It is hard to review with 4k lines of changes. Thing like .yml doesn't have to be in the same PR.

@vschandramourya
vschandramourya deleted the tool-calling-tests branch February 24, 2026 00:36

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: a0dd21ef20

ℹ️ About Codex in GitHub

Codex has been enabled to automatically review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

When you sign up for Codex through ChatGPT, Codex can also answer questions or update the PR, like "@codex address that feedback".

Comment on lines +238 to +241
if isinstance(expected, list) and isinstance(actual, list):
if len(expected) == len(actual):
return all(_values_match(a, e) for a, e in zip(actual, expected))
return False

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Add dict-aware matching for nested BFCL arguments

_values_match never handles dictionary structures recursively, so BFCL cases whose expected values are object-shaped alternatives (for example budget/gradeDict entries in the answer files) are treated as mismatches even when the model returns correct scalar fields. Because only list recursion is implemented here, valid tool calls with nested object arguments are systematically marked as failures, which underreports BFCL accuracy.

Useful? React with 👍 / 👎.

Comment on lines +72 to +74
for i in range(check_count):
gt_entry = ground_truth[i]
actual = actual_tool_calls[i]

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Match parallel tool calls without positional coupling

The evaluator aligns calls strictly by index (ground_truth[i] vs actual_tool_calls[i]), so two equivalent sets of calls fail whenever the model emits them in a different order. This is especially problematic for the parallel and parallel_multiple categories, where call ordering is non-semantic and backend generation order can vary, leading to false negatives in accuracy reporting.

Useful? React with 👍 / 👎.

Comment on lines +14 to +16
"https://huggingface.co/datasets/gorilla-llm/"
"Berkeley-Function-Calling-Leaderboard/resolve/main"
)

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Pin BFCL download source to an immutable dataset revision

The download script pulls from resolve/main, so each CI run can overwrite the checked-in BFCL fixtures with whatever is currently on the upstream default branch. Since both nightly and PR workflows invoke this script before BFCL tests, benchmark coverage and baseline comparisons become non-reproducible and can regress due to upstream dataset changes unrelated to this repository.

Useful? React with 👍 / 👎.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

ci CI/CD configuration changes tests Test changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants