Skip to content

e2e: rewrite nightly summary with gRPC vs HTTP comparison - #381

Merged
slin1237 merged 1 commit into
mainfrom
nigh-fix
Feb 9, 2026
Merged

slin1237 merged 1 commit into
mainfrom
nigh-fix

Conversation

@slin1237

@slin1237 slin1237 commented Feb 9, 2026 •

Copy link
Copy Markdown
Member

Summary

  • Rewrites nightly_summarize.py to produce a comprehensive gRPC vs HTTP comparison report for the GitHub Actions summary page, replacing the previous raw-data-only tables
  • Fixes _TIMEOUT_SEC in test_nightly_perf.py from 3h to match workflow's 24h timeout

What changed

e2e_test/benchmarks/nightly_summarize.py — full rewrite:

  • Aggregate comparison table: avg/median % difference per metric (TTFT, E2E, TPOT, throughput, RPS) with winner labels
  • Performance by concurrency: collapsible TTFT/E2E/throughput tables broken down by concurrency level (1–256)
  • Win/loss scorecard: how often gRPC beats HTTP at 1%/2%/5%/10% thresholds for E2E, TTFT, and throughput
  • Top gRPC wins: 15 largest individual improvements (>10%)
  • Per-model summary: compact one-row-per-model table, plus collapsible per-config detail tables showing every scenario×concurrency % diff
  • Dynamic overview columns: auto-discovered from experiment data instead of hardcoded — adding new runtimes (e.g. trtllm) or protocols (e.g. HTTP vLLM) automatically produces new columns and comparisons
  • _KNOWN_RUNTIMES set: single place to register new runtimes (sglang, vllm, trtllm)
  • Comparison logic matches gRPC/HTTP pairs by (model, runtime, worker_type, scenario, concurrency), filtering out error-rate > 0

e2e_test/benchmarks/test_nightly_perf.py:

  • Changed _TIMEOUT_SEC from 10800 (3h) to 1440 * 60 (24h) to match timeout-minutes: 1440 in the workflow. The 3h limit was causing premature SIGKILL of genai-bench subprocesses.

Why

The previous summary page only showed raw data tables per model/config — no comparison between gRPC and HTTP, making it hard to draw conclusions without downloading artifacts and analyzing offline. The new summary surfaces the same analysis that was previously done manually.

The timeout fix prevents false failures where genai-bench is killed after 3h even though the workflow allows 24h.

Test plan

  • Ran nightly_summarize.py against downloaded artifacts from run #21807599943 — produces 495 matched comparison points, all sections render correctly
  • Verified output matches hand-analyzed report data (same top wins, same scorecard counts, same aggregate percentages)
  • Clean run with no stderr warnings
  • Verified dynamic column discovery works (no hardcoded runtime/protocol lists in overview table)

Summary by CodeRabbit

Release Notes

  • Tests
    • Enhanced nightly performance benchmark reporting with comprehensive gRPC vs HTTP protocol comparisons, including aggregate metrics, per-concurrency breakdowns, and detailed performance analysis
    • Increased timeout duration for nightly test runs to 24 hours to accommodate extended benchmark scenarios

…eout

Rewrite nightly_summarize.py to produce a comprehensive gRPC vs HTTP
comparison report instead of raw per-config data tables. The new summary
is written to GITHUB_STEP_SUMMARY and includes:

- Aggregate comparison table (avg/median % diff per metric with winner)
- Performance by concurrency level (TTFT, E2E, throughput breakdowns)
- Win/loss scorecard at 1%/2%/5%/10% thresholds
- Top 15 largest gRPC wins (>10% improvement)
- Per-model summary table plus collapsible detail tables

The overview table columns are now dynamically discovered from experiment
data instead of hardcoded, so adding new runtimes (trtllm) or protocols
(http for vllm) will automatically produce new columns and comparisons.

Also fix _TIMEOUT_SEC in test_nightly_perf.py from 10800 (3h) to match
the workflow timeout-minutes of 1440 (24h). The 3h timeout was causing
premature SIGKILL of genai-bench processes.

Files changed:
- e2e_test/benchmarks/nightly_summarize.py: full rewrite with comparison
  logic, dynamic column discovery, _KNOWN_RUNTIMES set for extensibility
- e2e_test/benchmarks/test_nightly_perf.py: _TIMEOUT_SEC = 1440 * 60
@github-actions github-actions Bot added the tests Test changes label Feb 9, 2026
@coderabbitai

coderabbitai Bot commented Feb 9, 2026 •

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

This pull request introduces a comprehensive gRPC vs HTTP comparison reporting system for nightly benchmarks. The primary file is refactored to parse benchmark results, build cross-protocol comparisons, and generate detailed reports with aggregated metrics, per-concurrency breakdowns, and win/loss scorecards. A test timeout is increased to 24 hours to accommodate extended benchmark runs.

Changes

Cohort / File(s) Summary
Benchmark Report Enhancement
e2e_test/benchmarks/nightly_summarize.py
Comprehensive refactor replacing simple per-model reports with gRPC vs HTTP comparison reporting. Introduces RunResult enhancements (total_throughput, error_rate), ExperimentInfo.group_key property, and new ComparisonPoint dataclass. Adds parsing logic with constants, folder-name extraction, and metrics aggregation. Implements comparison building by aligning gRPC/HTTP runs and new formatting sections including overview, aggregated tables, per-concurrency breakdowns, win/loss scorecard, top wins, and per-model details.
Test Configuration
e2e_test/benchmarks/test_nightly_perf.py
Increases per-model timeout from 10800 seconds (3 hours) to 86400 seconds (24 hours) in the _TIMEOUT_SEC constant.

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Suggested labels

tests

Suggested reviewers

  • CatherineSue
  • key4ng
  • XinyueZhang369

Poem

🐰 Benchmarks now compare with gRPC pride,
HTTP measured right there by its side,
Throughput metrics dance, win-loss cards arranged,
A comprehensive report, beautifully exchanged!
Twenty-four hours for testing to thrive, 🚀

🚥 Pre-merge checks | ✅ 2 | ❌ 1
❌ Failed checks (1 warning)
Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 68.18% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (2 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title accurately and concisely describes the main change: rewriting the nightly summary to focus on gRPC vs HTTP comparison, which is the primary objective of the pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing touches
  • 📝 Generate docstrings
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Post copyable unit tests in a comment
  • Commit unit tests in branch nigh-fix

Comment @coderabbitai help to get the list of available commands and usage tips.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Summary of Changes

Hello @slin1237, I'm Gemini Code Assist1! I'm currently reviewing this pull request and will post my feedback shortly. In the meantime, here's a summary to help you and other reviewers quickly get up to speed!

This pull request significantly upgrades the nightly benchmark reporting system by transforming raw data into insightful comparative analyses between gRPC and HTTP protocols. It introduces several new, detailed report sections that highlight performance differences, trends across concurrency levels, and specific wins, making it much easier to interpret benchmark results. Additionally, a critical timeout setting for benchmark runs has been corrected to prevent false failures and ensure complete data collection.

Highlights

  • Enhanced Nightly Benchmark Reporting: The nightly_summarize.py script has been completely rewritten to provide a comprehensive gRPC vs HTTP comparison report directly on the GitHub Actions summary page, replacing previous raw-data-only tables.
  • Detailed Performance Analysis Sections: New report sections include an aggregate comparison table, performance breakdown by concurrency, a win/loss scorecard, a list of top gRPC performance improvements, and a compact per-model summary with detailed collapsible tables.
  • Dynamic Report Generation: Overview columns are now dynamically discovered from experiment data, ensuring that new runtimes (e.g., trtllm) or protocols are automatically included without requiring hardcoded updates.
  • Improved Runtime Detection and Extensibility: A new _KNOWN_RUNTIMES set simplifies the registration and detection of new runtimes like sglang, vllm, and trtllm within the parsing logic.
  • Timeout Configuration Fix: The _TIMEOUT_SEC in test_nightly_perf.py has been adjusted from 3 hours to 24 hours to align with the GitHub Actions workflow's timeout, preventing premature termination of benchmark processes and false failures.
Changelog
  • e2e_test/benchmarks/nightly_summarize.py
    • Rewrote the script to generate a detailed gRPC vs HTTP comparison report, including aggregate stats, per-concurrency breakdown, win/loss scorecard, top wins, and per-model detail tables.
    • Introduced ComparisonPoint dataclass to encapsulate matched gRPC and HTTP run data for easier comparison.
    • Added total_throughput and error_rate fields to the RunResult dataclass to capture more comprehensive metrics.
    • Enhanced parse_folder_name and parse_experiment functions for more robust and extensible parsing of experiment metadata, utilizing _KNOWN_PROTOCOLS, _KNOWN_WORKER_TYPES, and a new _KNOWN_RUNTIMES set.
    • Implemented new helper functions for percentage calculation (_pct), formatting (_fmt_pct, _fmt_latency_s, _fmt_throughput), and determining performance winners (_winner).
    • Developed new section generators (_section_overview, _section_aggregate, _section_by_concurrency, _section_scorecard, _section_top_wins, _section_per_model) to structure the new report.
    • Removed previous generic table generation and formatting functions, replacing them with comparison-specific logic.
  • e2e_test/benchmarks/test_nightly_perf.py
    • Modified _TIMEOUT_SEC from 10800 seconds (3 hours) to 1440 * 60 seconds (24 hours) to match the GitHub Actions workflow's timeout-minutes setting, preventing premature benchmark termination.
Activity
  • The author executed nightly_summarize.py against existing benchmark artifacts from run #21807599943 to validate the new reporting logic and ensure all sections render correctly.
  • The generated output was cross-referenced with manually analyzed report data to confirm accuracy in top wins, scorecard counts, and aggregate percentages.
  • Confirmed that the script runs cleanly with no stderr warnings, indicating stable execution.
  • Verified the dynamic column discovery feature functions as intended, ensuring flexibility for new runtimes and protocols without hardcoding.
Using Gemini Code Assist

The full guide for Gemini Code Assist can be found on our documentation page, here are some quick tips.

Invoking Gemini

You can request assistance from Gemini at any point by creating a comment using either /gemini <command> or @gemini-code-assist <command>. Below is a summary of the supported commands on the current page.

Feature Command Description
Code Review /gemini review Performs a code review for the current pull request in its current state.
Pull Request Summary /gemini summary Provides a summary of the current pull request in its current state.
Comment @gemini-code-assist Responds in comments when explicitly tagged, both in pull request comments and review comments.
Help /gemini help Displays a list of available commands.

Customization

To customize Gemini Code Assist for GitHub experience, repository maintainers can create a configuration file and/or provide a custom code review style guide (such as PEP-8 for Python) by creating and adding files to a .gemini/ folder in the base of the repository. Detailed instructions can be found here.

Limitations & Feedback

Gemini Code Assist may make mistakes. Please leave feedback on any instances where its feedback is incorrect or counter productive. You can react with 👍 and 👎 on @gemini-code-assist comments. If you're interested in giving your feedback about your experience with Gemini Code Assist for Github and other Google products, sign up here.

You can also get AI-powered code generation, chat, as well as code reviews directly in the IDE at no cost with the Gemini Code Assist IDE Extension.

Footnotes

  1. Review the Privacy Notices, Generative AI Prohibited Use Policy, Terms of Service, and learn how to configure Gemini Code Assist in GitHub here. Gemini can make mistakes, so double check it and use code with caution. ↩

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request significantly enhances the nightly benchmark summary by rewriting nightly_summarize.py to produce a detailed gRPC vs. HTTP performance comparison. The new report is well-structured with aggregate stats, per-concurrency breakdowns, and win/loss scorecards, which is a massive improvement for performance analysis. The code is well-organized and modular. I've added a few suggestions to further improve robustness and maintainability. The timeout fix in test_nightly_perf.py is also a welcome correction.

runtime = "vllm"
# Fallback: detect runtime from folder name
folder_lower = folder.name.lower()
for rt in _KNOWN_RUNTIMES:

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The current logic for detecting the runtime from the folder name could be brittle. Since the iteration order over a set is not guaranteed, this could lead to incorrect parsing if _KNOWN_RUNTIMES ever contains a name that is a substring of another (e.g., "foo" and "foobar"). To make this more robust, you can sort the known runtimes by length in descending order before iterating.

Suggested change
for rt in _KNOWN_RUNTIMES:
for rt in sorted(list(_KNOWN_RUNTIMES), key=len, reverse=True):

Comment on lines +424 to +486
# TTFT mean
lines.extend(
[
"<details>",
"<summary><b>TTFT Mean by Concurrency</b></summary>",
"",
"| Concurrency | gRPC avg | HTTP avg | Diff % | Winner |",
"|---:|---:|---:|---:|:---|",
]
)
for conc in conc_levels:
cps = by_conc[conc]
g_avg = sum(cp.grpc.ttft_mean * 1000 for cp in cps) / len(cps)
h_avg = sum(cp.http.ttft_mean * 1000 for cp in cps) / len(cps)
pct = _pct(g_avg, h_avg)
lines.append(
f"| {conc} | {g_avg:.0f}ms | {h_avg:.0f}ms | {_fmt_pct(pct)} | "
f"{_winner(pct, lower_is_better=True)} |"
)
lines.extend(["", "</details>", ""])

# E2E mean
lines.extend(
[
"<details>",
"<summary><b>E2E Latency Mean by Concurrency</b></summary>",
"",
"| Concurrency | gRPC avg | HTTP avg | Diff % | Winner |",
"|---:|---:|---:|---:|:---|",
]
)
for conc in conc_levels:
cps = by_conc[conc]
g_avg = sum(cp.grpc.e2e_mean * 1000 for cp in cps) / len(cps)
h_avg = sum(cp.http.e2e_mean * 1000 for cp in cps) / len(cps)
pct = _pct(g_avg, h_avg)
lines.append(
f"| {conc} | {g_avg:.0f}ms | {h_avg:.0f}ms | {_fmt_pct(pct)} | "
f"{_winner(pct, lower_is_better=True)} |"
)
lines.extend(["", "</details>", ""])

# Output throughput
lines.extend(
[
"<details>",
"<summary><b>Output Throughput by Concurrency</b></summary>",
"",
"| Concurrency | gRPC avg | HTTP avg | Diff % | Winner |",
"|---:|---:|---:|---:|:---|",
]
)
for conc in conc_levels:
cps = by_conc[conc]
g_avg = sum(cp.grpc.output_throughput for cp in cps) / len(cps)
h_avg = sum(cp.http.output_throughput for cp in cps) / len(cps)
pct = _pct(g_avg, h_avg)
lines.append(
f"| {conc} | {_fmt_throughput(g_avg)} tok/s | "
f"{_fmt_throughput(h_avg)} tok/s | {_fmt_pct(pct)} | "
f"{_winner(pct, lower_is_better=False)} |"
)
lines.extend(["", "</details>", ""])

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

There's significant code duplication in this function for generating the tables for TTFT, E2E latency, and Output Throughput. The structure of each table generation block is nearly identical. Consider refactoring this into a helper function to improve maintainability and reduce code size. The helper could take parameters like the table title, metric field name, and whether it's a latency or throughput metric.

Comment on lines +526 to +539
if lower_better:
if pct < -thresh:
grpc_w += 1
elif pct > thresh:
http_w += 1
else:
within += 1
else:
if pct > thresh:
grpc_w += 1
elif pct < -thresh:
http_w += 1
else:
within += 1

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

The nested if/else logic to count gRPC vs. HTTP wins can be simplified. You can first check if the percentage difference is within the threshold, and if not, use a single conditional expression to determine the winner based on the lower_better flag. This makes the logic more concise and easier to read.

                if abs(pct) <= thresh:
                    within += 1
                else:
                    # gRPC wins if pct is negative for latency, or positive for throughput
                    is_grpc_win = (pct < 0) if lower_better else (pct > 0)
                    if is_grpc_win:
                        grpc_w += 1
                    else:
                        http_w += 1

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 1

🤖 Fix all issues with AI agents
In `@e2e_test/benchmarks/nightly_summarize.py`:
- Around line 45-65: The grouping keys on ExperimentInfo currently omit hardware
details; update the ExperimentInfo.group_key and ExperimentInfo.table_key
properties to include gpu_type and gpu_count so experiments are not compared
across different GPUs or counts — modify the f-strings in the ExperimentInfo
class (properties group_key and table_key) to append gpu_type and gpu_count
(e.g., include them separated with '|' in group_key and with '_' in table_key)
so grouping and table columns incorporate both GPU type and GPU count.
🧹 Nitpick comments (1)
e2e_test/benchmarks/nightly_summarize.py (1)

307-364: Consider flagging error_rate in the overview status.
A run with non-zero error_rate but non-zero throughput would still show ✅; including error_rate improves the signal.

♻️ Suggested tweak
-                has_errors = any(r.rps == 0 or r.output_throughput == 0 for r in exp.runs)
+                has_errors = any(
+                    r.error_rate > 0 or r.rps == 0 or r.output_throughput == 0
+                    for r in exp.runs
+                )

Comment thread e2e_test/benchmarks/nightly_summarize.py
@slin1237
slin1237 merged commit c4836aa into main Feb 9, 2026
19 checks passed
@slin1237
slin1237 deleted the nigh-fix branch February 9, 2026 17:59
ppraneth pushed a commit that referenced this pull request Feb 18, 2026
Signed-off-by: ppraneth <pranethparuchuri@gmail.com>
yetone added a commit that referenced this pull request Apr 24, 2026
…red`` + debug dump

The xfail helper was written against an old TokenSpeed main where
``create_grammar_backend()`` was defined but never called and
``Request.grammar`` was never assigned. That's no longer the case —
TokenSpeed's grammar pipeline is fully wired as of #236 + #381 + #395:

- ``request_handler.py`` instantiates ``GrammarManager``
- ``event_loop.py`` routes new requests through
  ``process_req_with_grammar`` + ``get_ready_grammar_requests``, with an
  async compile queue and abort support
- ``greedy.py`` / ``flashinfer*.py`` both apply ``sampling_info.vocab_mask``
  to logits before sampling when a grammar is active
- speculative (``eagle_utils.py:773``) and disagg decode paths already
  call ``grammar.accept_token``

So ``tokenspeed#361``'s surface (``sampling_params.json_schema`` →
``req.grammar`` → vocab mask) now works end-to-end. The three
``tool_choice=required/specific`` tests that were xfailed should run on
tokenspeed just like they do on sglang/vllm/trtllm; CI will confirm.

Also drops the short-lived debug dumps in ``servicer.py`` and the
``TOKENSPEED_DEBUG_OUTPUT=1`` toggle in ``e2e-gpu-job.yml`` — those were
staged to diagnose why meta-llama tool-call tests diverged from unsloth.
The diagnosis is done: output trimming is correct, chunks are
incremental, the remaining unconstrained (tool_choice=auto) drift is a
sampling-variance quality issue orthogonal to this PR.

Signed-off-by: yetone <yetoneful@gmail.com>
yetone added a commit that referenced this pull request Apr 27, 2026
…red`` + debug dump

The xfail helper was written against an old TokenSpeed main where
``create_grammar_backend()`` was defined but never called and
``Request.grammar`` was never assigned. That's no longer the case —
TokenSpeed's grammar pipeline is fully wired as of #236 + #381 + #395:

- ``request_handler.py`` instantiates ``GrammarManager``
- ``event_loop.py`` routes new requests through
  ``process_req_with_grammar`` + ``get_ready_grammar_requests``, with an
  async compile queue and abort support
- ``greedy.py`` / ``flashinfer*.py`` both apply ``sampling_info.vocab_mask``
  to logits before sampling when a grammar is active
- speculative (``eagle_utils.py:773``) and disagg decode paths already
  call ``grammar.accept_token``

So ``tokenspeed#361``'s surface (``sampling_params.json_schema`` →
``req.grammar`` → vocab mask) now works end-to-end. The three
``tool_choice=required/specific`` tests that were xfailed should run on
tokenspeed just like they do on sglang/vllm/trtllm; CI will confirm.

Also drops the short-lived debug dumps in ``servicer.py`` and the
``TOKENSPEED_DEBUG_OUTPUT=1`` toggle in ``e2e-gpu-job.yml`` — those were
staged to diagnose why meta-llama tool-call tests diverged from unsloth.
The diagnosis is done: output trimming is correct, chunks are
incremental, the remaining unconstrained (tool_choice=auto) drift is a
sampling-variance quality issue orthogonal to this PR.

Signed-off-by: yetone <yetoneful@gmail.com>
CatherineSue pushed a commit that referenced this pull request Apr 30, 2026
…red`` + debug dump

The xfail helper was written against an old TokenSpeed main where
``create_grammar_backend()`` was defined but never called and
``Request.grammar`` was never assigned. That's no longer the case —
TokenSpeed's grammar pipeline is fully wired as of #236 + #381 + #395:

- ``request_handler.py`` instantiates ``GrammarManager``
- ``event_loop.py`` routes new requests through
  ``process_req_with_grammar`` + ``get_ready_grammar_requests``, with an
  async compile queue and abort support
- ``greedy.py`` / ``flashinfer*.py`` both apply ``sampling_info.vocab_mask``
  to logits before sampling when a grammar is active
- speculative (``eagle_utils.py:773``) and disagg decode paths already
  call ``grammar.accept_token``

So ``tokenspeed#361``'s surface (``sampling_params.json_schema`` →
``req.grammar`` → vocab mask) now works end-to-end. The three
``tool_choice=required/specific`` tests that were xfailed should run on
tokenspeed just like they do on sglang/vllm/trtllm; CI will confirm.

Also drops the short-lived debug dumps in ``servicer.py`` and the
``TOKENSPEED_DEBUG_OUTPUT=1`` toggle in ``e2e-gpu-job.yml`` — those were
staged to diagnose why meta-llama tool-call tests diverged from unsloth.
The diagnosis is done: output trimming is correct, chunks are
incremental, the remaining unconstrained (tool_choice=auto) drift is a
sampling-variance quality issue orthogonal to this PR.

Signed-off-by: yetone <yetoneful@gmail.com>
key4ng pushed a commit that referenced this pull request Apr 30, 2026
…red`` + debug dump

The xfail helper was written against an old TokenSpeed main where
``create_grammar_backend()`` was defined but never called and
``Request.grammar`` was never assigned. That's no longer the case —
TokenSpeed's grammar pipeline is fully wired as of #236 + #381 + #395:

- ``request_handler.py`` instantiates ``GrammarManager``
- ``event_loop.py`` routes new requests through
  ``process_req_with_grammar`` + ``get_ready_grammar_requests``, with an
  async compile queue and abort support
- ``greedy.py`` / ``flashinfer*.py`` both apply ``sampling_info.vocab_mask``
  to logits before sampling when a grammar is active
- speculative (``eagle_utils.py:773``) and disagg decode paths already
  call ``grammar.accept_token``

So ``tokenspeed#361``'s surface (``sampling_params.json_schema`` →
``req.grammar`` → vocab mask) now works end-to-end. The three
``tool_choice=required/specific`` tests that were xfailed should run on
tokenspeed just like they do on sglang/vllm/trtllm; CI will confirm.

Also drops the short-lived debug dumps in ``servicer.py`` and the
``TOKENSPEED_DEBUG_OUTPUT=1`` toggle in ``e2e-gpu-job.yml`` — those were
staged to diagnose why meta-llama tool-call tests diverged from unsloth.
The diagnosis is done: output trimming is correct, chunks are
incremental, the remaining unconstrained (tool_choice=auto) drift is a
sampling-variance quality issue orthogonal to this PR.

Signed-off-by: yetone <yetoneful@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

tests Test changes

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant