Skip to content

Llane/release notes 26.02 audit fixes - #1519

Closed
lbliii wants to merge 103 commits into
NVIDIA-NeMo:r1.1.0from
lbliii:llane/release-notes-26.02-audit-fixes
Closed

Llane/release notes 26.02 audit fixes#1519
lbliii wants to merge 103 commits into
NVIDIA-NeMo:r1.1.0from
lbliii:llane/release-notes-26.02-audit-fixes

Conversation

@lbliii

@lbliii lbliii commented Feb 18, 2026

Copy link
Copy Markdown
Contributor

No description provided.

lbliii and others added 30 commits September 22, 2025 13:56
Signed-off-by: Lawrence Lane <llane@nvidia.com>
Signed-off-by: Lawrence Lane <llane@nvidia.com>
Signed-off-by: Lawrence Lane <llane@nvidia.com>
Signed-off-by: Lawrence Lane <llane@nvidia.com>
Signed-off-by: Lawrence Lane <llane@nvidia.com>
Signed-off-by: Lawrence Lane <llane@nvidia.com>
Signed-off-by: Lawrence Lane <llane@nvidia.com>
Signed-off-by: Lawrence Lane <llane@nvidia.com>
Signed-off-by: Lawrence Lane <llane@nvidia.com>
Signed-off-by: Lawrence Lane <llane@nvidia.com>
Signed-off-by: Lawrence Lane <llane@nvidia.com>
Signed-off-by: Lawrence Lane <llane@nvidia.com>
Co-authored-by: Praateek Mahajan <praateekmahajan@users.noreply.github.com>
Signed-off-by: L.B. <llane@nvidia.com>
Signed-off-by: Lawrence Lane <llane@nvidia.com>
Signed-off-by: Lawrence Lane <llane@nvidia.com>
Signed-off-by: Lawrence Lane <llane@nvidia.com>
Signed-off-by: Lawrence Lane <llane@nvidia.com>
Signed-off-by: Lawrence Lane <llane@nvidia.com>
Co-authored-by: Sarah Yurick <53962159+sarahyurick@users.noreply.github.com>
Signed-off-by: L.B. <llane@nvidia.com>
Signed-off-by: Lawrence Lane <llane@nvidia.com>
Signed-off-by: Lawrence Lane <llane@nvidia.com>
Signed-off-by: Lawrence Lane <llane@nvidia.com>
Signed-off-by: Lawrence Lane <llane@nvidia.com>
Signed-off-by: Lawrence Lane <llane@nvidia.com>
Signed-off-by: Lawrence Lane <llane@nvidia.com>
SwekeR-463 and others added 9 commits February 12, 2026 10:59
…NVIDIA-NeMo#1470)

* Refactor stage names and update paths in configuration files

Signed-off-by: SwekeR-463 <swekerswasti@gmail.com>

* Fix processing logic to handle None results in ProcessingStage

Signed-off-by: SwekeR-463 <swekerswasti@gmail.com>

* Refactor stage classes to propagate names from constructors and remove hardcoded names

Signed-off-by: SwekeR-463 <swekerswasti@gmail.com>

* Fuse document iterate and extract stages (NVIDIA-NeMo#1458)

* Fuse document iterate and extract stages

Signed-off-by: Sarah Yurick <sarahyurick@gmail.com>

* ruff

Signed-off-by: Sarah Yurick <sarahyurick@gmail.com>

* fix bug

Signed-off-by: Sarah Yurick <sarahyurick@gmail.com>

* update docs and tutorial

Signed-off-by: Sarah Yurick <sarahyurick@gmail.com>

* save progress

Signed-off-by: Sarah Yurick <sarahyurick@gmail.com>

* update more tests

Signed-off-by: Sarah Yurick <sarahyurick@gmail.com>

* ruff

Signed-off-by: Sarah Yurick <sarahyurick@gmail.com>

* fix tests

Signed-off-by: Sarah Yurick <sarahyurick@gmail.com>

* ruff

Signed-off-by: Sarah Yurick <sarahyurick@gmail.com>

* update benchmark

Signed-off-by: Sarah Yurick <sarahyurick@gmail.com>

* move class

Signed-off-by: Sarah Yurick <sarahyurick@gmail.com>

* add missing import

Signed-off-by: Sarah Yurick <sarahyurick@gmail.com>

* update comment

Signed-off-by: Sarah Yurick <sarahyurick@gmail.com>

---------

Signed-off-by: Sarah Yurick <sarahyurick@gmail.com>
Co-authored-by: Ayush Dattagupta <ayushdg95@gmail.com>
Signed-off-by: SwekeR-463 <swekerswasti@gmail.com>

* Llane/sdg ray docs (NVIDIA-NeMo#1347)

* sdg ray docs init

Signed-off-by: Lawrence Lane <llane@nvidia.com>

* header, tab fixes

Signed-off-by: Lawrence Lane <llane@nvidia.com>

* style guide

Signed-off-by: Lawrence Lane <llane@nvidia.com>

* release notes change, bump version

Signed-off-by: Lawrence Lane <llane@nvidia.com>

* feedback

Signed-off-by: Lawrence Lane <llane@nvidia.com>

* readme

Signed-off-by: Lawrence Lane <llane@nvidia.com>

* updates

Signed-off-by: Lawrence Lane <llane@nvidia.com>

* updates

Signed-off-by: Lawrence Lane <llane@nvidia.com>

* updates

Signed-off-by: Lawrence Lane <llane@nvidia.com>

* updates

Signed-off-by: Lawrence Lane <llane@nvidia.com>

* Update tutorials/synthetic/README.md

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
Signed-off-by: Lawrence Lane <llane@nvidia.com>

---------

Signed-off-by: Lawrence Lane <llane@nvidia.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
Signed-off-by: SwekeR-463 <swekerswasti@gmail.com>

* ci: Remove thirdparty aiohttp file from ray (NVIDIA-NeMo#1469)

Signed-off-by: Dong Hyuk Chang <donghyukc@nvidia.com>
Signed-off-by: SwekeR-463 <swekerswasti@gmail.com>

* refactor: Enhance stage classes to propagate performance metrics and set stage names

Signed-off-by: SwekeR-463 <swekerswasti@gmail.com>

* fixes as per greptile

Signed-off-by: SwekeR-463 <swekerswasti@gmail.com>

* fixes as per greptile

Signed-off-by: SwekeR-463 <swekerswasti@gmail.com>

* remove continue part

Signed-off-by: SwekeR-463 <swekerswasti@gmail.com>

* Update nemo_curator/stages/audio/common.py

Signed-off-by: Sarah Yurick <53962159+sarahyurick@users.noreply.github.com>

---------

Signed-off-by: SwekeR-463 <swekerswasti@gmail.com>
Signed-off-by: Sarah Yurick <sarahyurick@gmail.com>
Signed-off-by: Lawrence Lane <llane@nvidia.com>
Signed-off-by: Dong Hyuk Chang <donghyukc@nvidia.com>
Signed-off-by: Sarah Yurick <53962159+sarahyurick@users.noreply.github.com>
Co-authored-by: Sarah Yurick <53962159+sarahyurick@users.noreply.github.com>
Co-authored-by: Ayush Dattagupta <ayushdg95@gmail.com>
Co-authored-by: Lawrence Lane <llane@nvidia.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
Co-authored-by: Dong Hyuk Chang <thomaschang26@tutanota.com>
* first commit

* initial commit - workable

* fix lint

* fix lint

* addressing comments, tested

* fix lint

* update uv lock

* fix greptile suggestion

* addressing comments

* lint

* greptile fix

* lint

* addressing comments

* fix lint

* update uv lock with putest_server

* resolve uv lock

* restore uv.lock

* addressing comment

* ndd benchmark

* update  uv.lock

* udpate benchmarking

* adding num_input_chars, num_output_chars to metrics

* lint

* addressing comments

* fix bug

* lint

* lin

* lint

* fix import

* fix .github/workflows/cicd-main.yml

---------

Co-authored-by: Huy Vu2 <huvu@login-eos02.eos.clusters.nvidia.com>
NVIDIA-NeMo#1452)

* [benchmark] Add FastText filter benchmarking script (NVIDIA-NeMo#1411)

- Add fasttext_filter_benchmark.py script following the pattern from
  score_filter_benchmark.py
- Add fasttext_filter_raydata and fasttext_filter_xenna entries to
  nightly-benchmark.yaml
- Supports FastText language ID and quality filters with model setup
  requirements

Fixes NVIDIA-NeMo#1411

Signed-off-by: Kunal Sachdev <kunalmgsachdev@gmail.com>

* [benchmark] Wire FastText model paths explicitly and update nightly config (NVIDIA-NeMo#1411)

- Add separate dataset entries for FastText langid and quality models
- Pass FastText model paths as explicit CLI arguments to benchmarks
- Remove hardcoded model paths from Hydra overrides
- Update FastText filter benchmarks to use model_weights_path
- Align arxiv E2E benchmark arg naming with FastText langid usage

Signed-off-by: Kunal Sachdev <kunalmgsachdev@gmail.com>

* Updated fasttext_filter_raydata benchmark timeout in benchmarking/nightly-benchmark.yaml basis Sarah Yurick's test run (NVIDIA-NeMo#1411)

Signed-off-by: Kunal Sachdev <kunalmgsachdev@gmail.com>

* Updated fasttext_filter_xenna benchmark timeout in benchmarking/nightly-benchmark.yaml basis Sarah Yurick's test run (NVIDIA-NeMo#1411)

Co-authored-by: Sarah Yurick <53962159+sarahyurick@users.noreply.github.com>
Signed-off-by: Kunal Sachdev <kunalmgsachdev@gmail.com>

* Updated fasttext_quality_model dataset entry's model file name to model.bin in benchmarking/nightly-benchmark.yaml (NVIDIA-NeMo#1411)

Co-authored-by: Sarah Yurick <53962159+sarahyurick@users.noreply.github.com>
Signed-off-by: Kunal Sachdev <kunalmgsachdev@gmail.com>

* Adding ftz file option for fasttext_langid_model datasets entry in benchmarking/nightly-benchmark.yaml (NVIDIA-NeMo#1411)

Co-authored-by: Sarah Yurick <53962159+sarahyurick@users.noreply.github.com>
Signed-off-by: Kunal Sachdev <kunalmgsachdev@gmail.com>

* Moving fasttext_filter_raydata and fasttext_filter_xenna to run right after ScoreFilter benchmarks in benchmarking/nightly-benchmark.yaml (NVIDIA-NeMo#1411)

Signed-off-by: Kunal Sachdev <kunalmgsachdev@gmail.com>

---------

Signed-off-by: Kunal Sachdev <kunalmgsachdev@gmail.com>
Co-authored-by: Sarah Yurick <53962159+sarahyurick@users.noreply.github.com>
) (NVIDIA-NeMo#1507)

Signed-off-by: Arivunidhi A <arivunidhi.a@gmail.com>
Co-authored-by: Arivunidhi A <arivunidhi.a@gmail.com>
Co-authored-by: Sarah Yurick <53962159+sarahyurick@users.noreply.github.com>
…ck (NVIDIA-NeMo#1511)

* Refactor video frame extraction to improve PyNvCodec availability check

- Removed the try-except block for importing PyNvcFrameExtractor, simplifying the import logic.
- Updated the condition for initializing the PyNvcFrameExtractor in the VideoFrameExtractionStage to rely solely on the _PYNVC_AVAILABLE flag.
- Adjusted the handling of pixel format conversion in NvVideoDecoder to prepare for future updates to cvcuda.

Signed-off-by: Abhinav Garg <abhinavg@stanford.edu>

* Refactor NvVideoDecoder to replace deprecated nvcv_image with cvcuda tensor

- Updated NvVideoDecoder to remove the use of nvcv_image, which is deprecated, and replaced it with cvcuda tensor.
- Adjusted related tensor operations and tests to ensure compatibility with the new cvcuda implementation.

Signed-off-by: Abhinav Garg <abhinavg@stanford.edu>

* Update import statements in test_nvcodec_utils.py to include ruff linting rule

- Modified import statements in the test file to include the RUF100 linting rule, ensuring better adherence to coding standards.
- This change enhances the clarity of the import handling tests.

Signed-off-by: [Your Name] <your.email@example.com>
Signed-off-by: Abhinav Garg <abhinavg@stanford.edu>

* Update tests/utils/test_nvcodec_utils.py

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
Signed-off-by: Abhinav Garg <abhinavg@stanford.edu>

* Update tests/utils/test_nvcodec_utils.py

Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
Signed-off-by: Abhinav Garg <abhinavg@stanford.edu>

---------

Signed-off-by: Abhinav Garg <abhinavg@stanford.edu>
Signed-off-by: [Your Name] <your.email@example.com>
Co-authored-by: greptile-apps[bot] <165735046+greptile-apps[bot]@users.noreply.github.com>
Co-authored-by: Ayush Dattagupta <ayushdg95@gmail.com>
* add model weights

Signed-off-by: Vibhu Jawa <vjawa@nvidia.com>

* address validation feedback

Signed-off-by: Vibhu Jawa <vjawa@nvidia.com>

---------

Signed-off-by: Vibhu Jawa <vjawa@nvidia.com>
Co-authored-by: Sarah Yurick <53962159+sarahyurick@users.noreply.github.com>
… are specified (NVIDIA-NeMo#1508)

* Warn and resolve conflict when both blocksize and files_per_partition are specified (NVIDIA-NeMo#1401)

Signed-off-by: Arivunidhi A <arivunidhi.a@gmail.com>

* Add warning assertion to test_both_blocksize_and_files_per_partition_warns

Address review feedback: verify that the warning message is actually
logged when both blocksize and files_per_partition are specified,
using caplog fixture consistent with existing test patterns.

---------

Signed-off-by: Arivunidhi A <arivunidhi.a@gmail.com>
Co-authored-by: Arivunidhi A <arivunidhi.a@gmail.com>
@lbliii lbliii self-assigned this Feb 18, 2026
@copy-pr-bot

copy-pr-bot Bot commented Feb 18, 2026

Copy link
Copy Markdown

This pull request requires additional validation before any workflows can run on NVIDIA's runners.

Pull request vetters can view their responsibilities here.

Contributors can view more details about this message here.

- Fix YAML config example to use correct Hydra CLI (--config-path, --config-name)
- Correct image curation batch sizes to match source (dali_batch_size=100, num_threads=8)
- Update code linting claim to Ruff (markdownlint not in pre-commit)
- Remove unverifiable Ray Actor Pool progress bars and small cluster warnings

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Lawrence Lane <llane@nvidia.com>
@lbliii
lbliii force-pushed the llane/release-notes-26.02-audit-fixes branch from 1a65eba to 488c3d3 Compare February 18, 2026 18:05
@greptile-apps

greptile-apps Bot commented Feb 18, 2026

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR consolidates the 26.02 release documentation by adding a concise CHANGELOG entry and streamlining the verbose release notes. The changes include:

  • Added CHANGELOG.md section for version 1.1.0 with organized categories (New Features, Improvements, Dependency Updates, Bug Fixes, Infrastructure, Breaking Changes, Documentation)
  • Simplified docs/about/release-notes/index.md by consolidating sections like "Benchmarking Infrastructure" → "Stage and Pipeline Benchmarking" and "Workflow Results API" → "Pipeline Performance and Metric Logging"
  • Updated YAML configuration command example to match the actual syntax in nemo_curator/config/README.md
  • Removed several detailed subsections (Ray Actor Pool Executor Improvements, Enhanced Embedding Generation, specific text curation bullet points)

The edits follow a consistent editorial direction toward more concise documentation. All formatting is correct, and the command examples are consistent with existing documentation.

Confidence Score: 5/5

  • Documentation-only changes with no code modifications; safe to merge
  • This is a documentation-only PR that adds a properly formatted CHANGELOG entry and streamlines release notes. No code changes, no security concerns, and all formatting is correct. The command examples are verified against existing documentation.
  • No files require special attention

Important Files Changed

Filename Overview
CHANGELOG.md Added comprehensive 1.1.0 release section with proper formatting and complete feature coverage
docs/about/release-notes/index.md Streamlined 26.02 release notes by consolidating sections and removing verbose details; command example updated to match README

Last reviewed commit: 3c11c07

@greptile-apps greptile-apps Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1 file reviewed, no comments

Edit Code Review Agent Settings | Greptile

lbliii and others added 3 commits February 18, 2026 13:08
- Fix YAML config example to use correct Hydra CLI (--config-path, --config-name)
- Correct image curation batch sizes to match source (dali_batch_size=100, num_threads=8)
- Update code linting claim to Ruff (markdownlint not in pre-commit)
- Remove unverifiable Ray Actor Pool progress bars and small cluster warnings

Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Lawrence Lane <llane@nvidia.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Signed-off-by: Lawrence Lane <llane@nvidia.com>

@greptile-apps greptile-apps Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

2 files reviewed, 2 comments

Edit Code Review Agent Settings | Greptile

Comment thread docs/about/release-notes/index.md Outdated
Comment thread docs/about/release-notes/index.md Outdated

@sarahyurick sarahyurick left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Please update the PR to target r1.1.0 instead of main.

Comment thread docs/about/release-notes/index.md Outdated
Comment thread docs/about/release-notes/index.md Outdated
Comment thread CHANGELOG.md Outdated
Comment thread CHANGELOG.md Outdated
Comment thread CHANGELOG.md Outdated
Comment thread CHANGELOG.md Outdated
Comment thread docs/about/release-notes/index.md Outdated
Comment thread docs/about/release-notes/index.md Outdated
Comment thread docs/about/release-notes/index.md Outdated
Comment thread docs/about/release-notes/index.md Outdated
Comment thread docs/about/release-notes/index.md Outdated
Comment thread docs/about/release-notes/index.md Outdated
Comment thread CHANGELOG.md Outdated
Comment thread CHANGELOG.md Outdated
Signed-off-by: Lawrence Lane <llane@nvidia.com>
Signed-off-by: Lawrence Lane <llane@nvidia.com>

@greptile-apps greptile-apps Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

2 files reviewed, no comments

Edit Code Review Agent Settings | Greptile

@sarahyurick
sarahyurick changed the base branch from main to r1.1.0 February 19, 2026 19:05
@sarahyurick

Copy link
Copy Markdown
Contributor

Closing in favor of #1529.

- **End-to-End Pipeline Benchmarking**: Automated benchmarks for all curation modalities (text, image, video, audio)
- **Performance Tracking**: Integration with MLflow for metrics tracking and Slack for notifications
- **Nightly Benchmarks**: Continuous performance monitoring across:
- **Stage and Pipeline Benchmarking**: Automated benchmarks for curation modalities (text, image, video, audio)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Do we have for Audio ?

- Text pipelines: exact deduplication, fuzzy deduplication, semantic deduplication, score filters, modifiers
- Image curation workflows with DALI-based processing
- Video processing pipelines with scene detection and semantic deduplication
- Audio ASR inference and quality assessment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Again audio.

- **Performance Tracking**: Metrics tracking across:
- Text pipelines: exact deduplication, fuzzy deduplication, semantic deduplication, score filters, modifiers
- Image curation workflows with DALI-based processing
- Video processing pipelines with scene detection and semantic deduplication

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we add captioning here?

Video processing pipelines with splitting, scene detection, and captioning.

### Image Curation

- **Optimized Batch Sizes**: Reduced default batch sizes for better CPU memory usage (batch_size=50, num_threads=4)
- **Optimized Batch Sizes**: Configurable batch sizes for better CPU/GPU memory usage (batch_size=100, num_threads=8)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

num_threds is 16 I think

Comment thread CHANGELOG.md

- **Video**: Removed InternVideo2; vLLM 0.14.1, FFmpeg 8.0.1
- **Audio**: Enhanced ASR/WER docs, robust manifest handling
- **Image**: Optimized batch sizes (batch_size=100, num_threads=8), memory guidance

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

num_threads=16

Comment thread CHANGELOG.md

### Improvements

- **Video**: Removed InternVideo2; vLLM 0.14.1, FFmpeg 8.0.1

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

vLLM version is 0.15.1.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.