chore: update ty to 0.0.44 - #608
Conversation
Signed-off-by: Aaron Gonzales <aagonzales@nvidia.com>
|
No actionable comments were generated in the recent review. 🎉 ℹ️ Recent review info⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: CHILL Plan: Enterprise Run ID: 📒 Files selected for processing (1)
🚧 Files skipped from review as they are similar to previous changes (1)
📜 Recent review details⏰ Context from checks skipped due to timeout. (5)
WalkthroughThe PR updates typing guidance and runtime validation across datasets, data processing, privacy, telemetry, and evaluation code, with matching tests and a tool version bump. ChangesType Safety Hardening
Estimated code review effort🎯 3 (Moderate) | ⏱️ ~25 minutes Possibly related PRs
Suggested reviewers
🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✏️ Tip: You can configure your own custom pre-merge checks in the settings. ✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Comment |
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
Signed-off-by: Aaron Gonzales <aagonzales@nvidia.com>
Greptile SummaryThis PR bumps the
Confidence Score: 5/5Safe to merge; the type-checker upgrade uncovered and fixed real bugs rather than introducing new ones. Every source-code change is either a type-annotation correction that matches the runtime behavior already present, or an actual bug fix (infinite-loop risk in distributions, off-by-one and skipped attack iteration in AIA). The new unit tests directly cover the fixed paths. No logic was removed or re-architected in a way that could regress existing behavior. No files require special attention. Important Files Changed
|
There was a problem hiding this comment.
Actionable comments posted: 6
Caution
Some comments are outside the diff and can’t be posted inline due to platform limitations.
⚠️ Outside diff range comments (1)
src/nemo_safe_synthesizer/data_processing/actions/dates.py (1)
418-431: 🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick winPreserve the original column label in
date_min_dict.Casting
object_coltostrchanges the metadata key that later reversal code needs to map back onto the DataFrame. For non-string labels like0or tuples, this records"0"here while the transformed frame still carries0, so the reverse transform will miss the column or fail. Keep the original label as the key and widen the annotation instead of stringifying it.As per path instructions, review
src/nemo_safe_synthesizer/data_processing/**/*.pyfor data-contract regressions.Source: Path instructions
ℹ️ Review info
⚙️ Run configuration
Configuration used: Path: .coderabbit.yaml
Review profile: CHILL
Plan: Enterprise
Run ID: 3070e10e-c84b-4a12-9c48-d92f2d5c395b
⛔ Files ignored due to path filters (1)
mise.lockis excluded by!**/*.lock,!mise.lock
📒 Files selected for processing (19)
.mise.tomlSTYLE_GUIDE.mdsrc/nemo_safe_synthesizer/cli/datasets.pysrc/nemo_safe_synthesizer/data_processing/actions/dates.pysrc/nemo_safe_synthesizer/data_processing/actions/distributions.pysrc/nemo_safe_synthesizer/evaluation/components/attribute_inference_protection.pysrc/nemo_safe_synthesizer/evaluation/components/membership_inference_protection.pysrc/nemo_safe_synthesizer/evaluation/components/multi_modal_figures.pysrc/nemo_safe_synthesizer/evaluation/reports/multimodal/multimodal_report.pysrc/nemo_safe_synthesizer/privacy/dp_transformers/linear.pysrc/nemo_safe_synthesizer/telemetry.pysrc/nemo_safe_synthesizer/utils.pytests/cli/test_datasets.pytests/config/test_autoconfig.pytests/data_processing/test_dates.pytests/data_processing/test_distributions.pytests/evaluation/components/test_attribute_inference_protection.pytests/evaluation/components/test_membership_inference_protection.pytests/training/test_dp_linear.py
📜 Review details
⏰ Context from checks skipped due to timeout. (3)
- GitHub Check: Unit Tests (3.13)
- GitHub Check: Unit Tests (3.11)
- GitHub Check: Unit Tests (3.12)
🧰 Additional context used
📓 Path-based instructions (17)
.mise.toml
⚙️ CodeRabbit configuration file
Treat .mise.toml as toolchain supply-chain configuration. Check pinned tool choices, install cadence, platform coverage, environment settings, and whether changes require regenerating mise.lock.
Files:
.mise.toml
**/*.{md,markdown,py}
📄 CodeRabbit inference engine (.cursor/rules/agent-markdown-style.mdc)
**/*.{md,markdown,py}: Avoid decorative bold (**text**) in list items, body text, and docstrings; use structural cues (headers, list markers, colons, backticks) for emphasis instead
Use backticks for code identifiers, paths, and CLI commands in markdown and docstrings
Files:
src/nemo_safe_synthesizer/privacy/dp_transformers/linear.pytests/training/test_dp_linear.pySTYLE_GUIDE.mdtests/data_processing/test_distributions.pytests/data_processing/test_dates.pytests/config/test_autoconfig.pytests/cli/test_datasets.pytests/evaluation/components/test_attribute_inference_protection.pytests/evaluation/components/test_membership_inference_protection.pysrc/nemo_safe_synthesizer/evaluation/reports/multimodal/multimodal_report.pysrc/nemo_safe_synthesizer/evaluation/components/multi_modal_figures.pysrc/nemo_safe_synthesizer/cli/datasets.pysrc/nemo_safe_synthesizer/data_processing/actions/distributions.pysrc/nemo_safe_synthesizer/utils.pysrc/nemo_safe_synthesizer/evaluation/components/membership_inference_protection.pysrc/nemo_safe_synthesizer/data_processing/actions/dates.pysrc/nemo_safe_synthesizer/evaluation/components/attribute_inference_protection.pysrc/nemo_safe_synthesizer/telemetry.py
**/*.py
📄 CodeRabbit inference engine (AGENTS.md)
**/*.py: Place durable implementation guidance in function and class docstrings for public contracts and source comments for local invariants
Target Python 3.11–3.13 with modern syntax (X | Y,list[str],Self). Python 3.14+ is not supported
**/*.py: Source code must remain Python 3.11 syntax-compatible; do not use Python 3.12-only syntax such as PEP 695 type statements or bracketed generic class/function parameters in shared package code
Use ruff for formatting and linting via mise run format and mise run check tasks; run formatting before committing
Use ty for type checking via mise run check task; ensure type hints are present and valid
New features must include tests; bug fixes must include regression tests
**/*.py: In Pythonconfig/models, useNSSBaseModelfor user-facing configuration/parameter models; use rawBaseModelor module-specific bases for DTOs and internal structures.
UseBaseSettingsfor environment/CLI settings, preferAliasChoicesfor per-field aliases when a setting must respond to both its Python name and an environment variable name, and useenv_prefixonly for simple shared-prefix settings.
For Pydantic models, useField(description=...)as the canonical field documentation and always include it for model fields.
Prefer assignment-style Pydantic fields (name: T = Field(...)) overAnnotated[...]unless extra metadata beyondField()is needed; use bare assignments for defaults withAnnotated, except thatdefault_factorystill belongs in assignment-styleField(default_factory=...).
Use@dataclass(frozen=True)for immutable value objects and validators; mutable dataclasses are acceptable for builders, accumulators, and pipeline state.
Usefield(default_factory=list)(or another factory) for mutable dataclass defaults; never use a mutable literal default such as= [].
UseStrEnumfor string-valued enums used in configs or serialization, and plainEnumfor internal-only named constants.
In Python cod...
Files:
src/nemo_safe_synthesizer/privacy/dp_transformers/linear.pytests/training/test_dp_linear.pytests/data_processing/test_distributions.pytests/data_processing/test_dates.pytests/config/test_autoconfig.pytests/cli/test_datasets.pytests/evaluation/components/test_attribute_inference_protection.pytests/evaluation/components/test_membership_inference_protection.pysrc/nemo_safe_synthesizer/evaluation/reports/multimodal/multimodal_report.pysrc/nemo_safe_synthesizer/evaluation/components/multi_modal_figures.pysrc/nemo_safe_synthesizer/cli/datasets.pysrc/nemo_safe_synthesizer/data_processing/actions/distributions.pysrc/nemo_safe_synthesizer/utils.pysrc/nemo_safe_synthesizer/evaluation/components/membership_inference_protection.pysrc/nemo_safe_synthesizer/data_processing/actions/dates.pysrc/nemo_safe_synthesizer/evaluation/components/attribute_inference_protection.pysrc/nemo_safe_synthesizer/telemetry.py
**/*.{py,sh,yaml,yml,md}
📄 CodeRabbit inference engine (CONTRIBUTING.md)
All source files (.py, .sh, .yaml, .yml, .md) require SPDX copyright headers; mise run format adds them automatically
Files:
src/nemo_safe_synthesizer/privacy/dp_transformers/linear.pytests/training/test_dp_linear.pySTYLE_GUIDE.mdtests/data_processing/test_distributions.pytests/data_processing/test_dates.pytests/config/test_autoconfig.pytests/cli/test_datasets.pytests/evaluation/components/test_attribute_inference_protection.pytests/evaluation/components/test_membership_inference_protection.pysrc/nemo_safe_synthesizer/evaluation/reports/multimodal/multimodal_report.pysrc/nemo_safe_synthesizer/evaluation/components/multi_modal_figures.pysrc/nemo_safe_synthesizer/cli/datasets.pysrc/nemo_safe_synthesizer/data_processing/actions/distributions.pysrc/nemo_safe_synthesizer/utils.pysrc/nemo_safe_synthesizer/evaluation/components/membership_inference_protection.pysrc/nemo_safe_synthesizer/data_processing/actions/dates.pysrc/nemo_safe_synthesizer/evaluation/components/attribute_inference_protection.pysrc/nemo_safe_synthesizer/telemetry.py
src/nemo_safe_synthesizer/**/*.py
📄 CodeRabbit inference engine (CONTRIBUTING.md)
API reference pages are auto-generated from Python docstrings using Google-style format; write docstrings in src/nemo_safe_synthesizer/ and they will appear in the reference/
Files:
src/nemo_safe_synthesizer/privacy/dp_transformers/linear.pysrc/nemo_safe_synthesizer/evaluation/reports/multimodal/multimodal_report.pysrc/nemo_safe_synthesizer/evaluation/components/multi_modal_figures.pysrc/nemo_safe_synthesizer/cli/datasets.pysrc/nemo_safe_synthesizer/data_processing/actions/distributions.pysrc/nemo_safe_synthesizer/utils.pysrc/nemo_safe_synthesizer/evaluation/components/membership_inference_protection.pysrc/nemo_safe_synthesizer/data_processing/actions/dates.pysrc/nemo_safe_synthesizer/evaluation/components/attribute_inference_protection.pysrc/nemo_safe_synthesizer/telemetry.py
src/**/*.py
📄 CodeRabbit inference engine (STYLE_GUIDE.md)
Every directory under
src/that contains Python files must include an__init__.pyfile, even if empty.
Files:
src/nemo_safe_synthesizer/privacy/dp_transformers/linear.pysrc/nemo_safe_synthesizer/evaluation/reports/multimodal/multimodal_report.pysrc/nemo_safe_synthesizer/evaluation/components/multi_modal_figures.pysrc/nemo_safe_synthesizer/cli/datasets.pysrc/nemo_safe_synthesizer/data_processing/actions/distributions.pysrc/nemo_safe_synthesizer/utils.pysrc/nemo_safe_synthesizer/evaluation/components/membership_inference_protection.pysrc/nemo_safe_synthesizer/data_processing/actions/dates.pysrc/nemo_safe_synthesizer/evaluation/components/attribute_inference_protection.pysrc/nemo_safe_synthesizer/telemetry.py
⚙️ CodeRabbit configuration file
Review library code against STYLE_GUIDE.md. Focus on behavior, API contracts, error handling, resource cleanup, typing, logging, and user-facing failures. Public APIs and nontrivial functions need Google-style docstrings.
Files:
src/nemo_safe_synthesizer/privacy/dp_transformers/linear.pysrc/nemo_safe_synthesizer/evaluation/reports/multimodal/multimodal_report.pysrc/nemo_safe_synthesizer/evaluation/components/multi_modal_figures.pysrc/nemo_safe_synthesizer/cli/datasets.pysrc/nemo_safe_synthesizer/data_processing/actions/distributions.pysrc/nemo_safe_synthesizer/utils.pysrc/nemo_safe_synthesizer/evaluation/components/membership_inference_protection.pysrc/nemo_safe_synthesizer/data_processing/actions/dates.pysrc/nemo_safe_synthesizer/evaluation/components/attribute_inference_protection.pysrc/nemo_safe_synthesizer/telemetry.py
**/*
📄 CodeRabbit inference engine (STYLE_GUIDE.md)
**/*: Every source file must include an SPDX copyright header, with HTML comments for Markdown and hash comments for.py,.sh,.yaml, and.yml; Markdown files with YAML frontmatter must place the hash-comment SPDX header inside the frontmatter block.
Ensure files end with a newline and contain no trailing whitespace; keep single spacing between sentences.
Files:
src/nemo_safe_synthesizer/privacy/dp_transformers/linear.pytests/training/test_dp_linear.pySTYLE_GUIDE.mdtests/data_processing/test_distributions.pytests/data_processing/test_dates.pytests/config/test_autoconfig.pytests/cli/test_datasets.pytests/evaluation/components/test_attribute_inference_protection.pytests/evaluation/components/test_membership_inference_protection.pysrc/nemo_safe_synthesizer/evaluation/reports/multimodal/multimodal_report.pysrc/nemo_safe_synthesizer/evaluation/components/multi_modal_figures.pysrc/nemo_safe_synthesizer/cli/datasets.pysrc/nemo_safe_synthesizer/data_processing/actions/distributions.pysrc/nemo_safe_synthesizer/utils.pysrc/nemo_safe_synthesizer/evaluation/components/membership_inference_protection.pysrc/nemo_safe_synthesizer/data_processing/actions/dates.pysrc/nemo_safe_synthesizer/evaluation/components/attribute_inference_protection.pysrc/nemo_safe_synthesizer/telemetry.py
⚙️ CodeRabbit configuration file
**/*: Review as a senior maintainer for NeMo Safe Synthesizer. Prioritize issues that can change behavior, break user workflows, weaken privacy guarantees, hide failures, make tests unreliable, or create maintenance risk. Avoid generic style commentary unless it points to a concrete project convention that automated tools will not catch.
Comment only when the finding is actionable and tied to changed code. For each finding, state the impact, the condition that triggers it, and the smallest practical fix. Prefer one precise comment over broad advice. Do not ask for refactors outside the PR scope unless the changed code creates the problem.
Review type guidance: - Potential issue: use for correctness bugs, data loss, privacy leaks,
security risks, broken public APIs, invalid config behavior, missing
validation, hidden failures, nondeterministic tests, or CI breakage.
- Refactor suggestion: use for local maintainability problems introduced
by the diff when they have clear future cost, such as duplicated setup,
unclear boundaries, over-mocking, avoidable complexity, or opaque test
helpers.- Nitpick: avoid in chill mode. Do not emit formatting, import-order,
wording, or style-only comments unless automated tools cannot catch the
issue and it affects maintainability.Severity guidance: - Critical: security/privacy leaks, data loss, training/test/holdout
contamination, or broken release/package/core pipeline execution.
- Major: incorrect generation/training/evaluation behavior, broken
CLI/SDK public API, invalid config defaults or validators, or GPU/vLLM
cleanup and process-isolation bugs likely to fail CI or production
runs.- Minor: localized bugs, missing focused tests for changed behavior, or
bad test patterns that weaken regression coverage.- Trivial: small cleanup with no behavior impact. Usually suppress in
chill mode.- Info: context only. Avoid unless it helps reviewers understand risk.
Safe-Synthesizer-specific review focus: - Data ...
Files:
src/nemo_safe_synthesizer/privacy/dp_transformers/linear.pytests/training/test_dp_linear.pySTYLE_GUIDE.mdtests/data_processing/test_distributions.pytests/data_processing/test_dates.pytests/config/test_autoconfig.pytests/cli/test_datasets.pytests/evaluation/components/test_attribute_inference_protection.pytests/evaluation/components/test_membership_inference_protection.pysrc/nemo_safe_synthesizer/evaluation/reports/multimodal/multimodal_report.pysrc/nemo_safe_synthesizer/evaluation/components/multi_modal_figures.pysrc/nemo_safe_synthesizer/cli/datasets.pysrc/nemo_safe_synthesizer/data_processing/actions/distributions.pysrc/nemo_safe_synthesizer/utils.pysrc/nemo_safe_synthesizer/evaluation/components/membership_inference_protection.pysrc/nemo_safe_synthesizer/data_processing/actions/dates.pysrc/nemo_safe_synthesizer/evaluation/components/attribute_inference_protection.pysrc/nemo_safe_synthesizer/telemetry.py
**
⚙️ CodeRabbit configuration file
**:AGENTS.md
Guide for AI agents (Cursor, Windsurf, Claude Code, etc.) working in the Safe-Synthesizer repo.
This project loads local developer preferences from
@AGENTS.local.md. You MUST read this file if it exists and give its instructions top priority.Skills
Repo-specific skills live in
.agents/skills/; see.agents/README.mdfor the catalog. Read a skill when the task matches its scope instead of copying workflow details into this file.Durable implementation guidance belongs with the code it describes: function and class docstrings for public contracts and source comments for local invariants. Test-suite guidance belongs in
tests/TESTING.md.Repo Conventions
See STYLE_GUIDE.md for detailed code style conventions (Python, markdown, Dockerfiles, shell scripts, testing, config files, docstrings).
Use
uvfor everything -- neverpipor rawpython. Python 3.11–3.13 with modern syntax (X | Y,list[str],Self). Python 3.14+ is not supported.Common commands:
mise run test(unit tests),mise run format(auto-fix formatting + lint + copyright),mise run check(read-only local quality checks),mise run validate(pre-PR quality, lock, and CI unit checks),mise run typecheck(ty only). Always use mise tasks or the wrapper scripts intools/instead of runningruffortydirectly. Useuv runfor Python execution. When in doubt, inspectmise tasksandpytest --markers.The canonical
uv synccommand for a full GPU/dev environment is:uv sync --frozen --extra cu129 --extra engine --group devBare
uv sync --frozen(without extras) installs an incomplete environment --ty, import checks, and GPU tests will fail.Feature branches off
main. Branch names often include an issue number prefix (e.g.,<author>/123-short-name).Do ...
Files:
src/nemo_safe_synthesizer/privacy/dp_transformers/linear.pytests/training/test_dp_linear.pySTYLE_GUIDE.mdtests/data_processing/test_distributions.pytests/data_processing/test_dates.pytests/config/test_autoconfig.pytests/cli/test_datasets.pytests/evaluation/components/test_attribute_inference_protection.pytests/evaluation/components/test_membership_inference_protection.pysrc/nemo_safe_synthesizer/evaluation/reports/multimodal/multimodal_report.pysrc/nemo_safe_synthesizer/evaluation/components/multi_modal_figures.pysrc/nemo_safe_synthesizer/cli/datasets.pysrc/nemo_safe_synthesizer/data_processing/actions/distributions.pysrc/nemo_safe_synthesizer/utils.pysrc/nemo_safe_synthesizer/evaluation/components/membership_inference_protection.pysrc/nemo_safe_synthesizer/data_processing/actions/dates.pysrc/nemo_safe_synthesizer/evaluation/components/attribute_inference_protection.pysrc/nemo_safe_synthesizer/telemetry.py
src/nemo_safe_synthesizer/privacy/**/*.py
⚙️ CodeRabbit configuration file
Treat privacy changes as high-risk. Check DP accounting, parameter validation, data leakage, seed handling, model state persistence, and whether privacy guarantees are documented accurately.
Files:
src/nemo_safe_synthesizer/privacy/dp_transformers/linear.py
**/test_*.py
📄 CodeRabbit inference engine (AGENTS.md)
Use the
unitmarker instead of the deprecatedunit_testmarker for test identification
Files:
tests/training/test_dp_linear.pytests/data_processing/test_distributions.pytests/data_processing/test_dates.pytests/config/test_autoconfig.pytests/cli/test_datasets.pytests/evaluation/components/test_attribute_inference_protection.pytests/evaluation/components/test_membership_inference_protection.py
tests/**
📄 CodeRabbit inference engine (.cursor/rules/repo-navigation.mdc)
tests/**: Mirrorsrc/directory structure intests/directory for test organization
Auto-mark tests by directory:tests/e2e/→e2e,tests/smoke/→smoke, otherwise default tounitMirror source code directory structure in tests directory (e.g.,
tests/training/,tests/generation/parallel to source structure)
Files:
tests/training/test_dp_linear.pytests/data_processing/test_distributions.pytests/data_processing/test_dates.pytests/config/test_autoconfig.pytests/cli/test_datasets.pytests/evaluation/components/test_attribute_inference_protection.pytests/evaluation/components/test_membership_inference_protection.py
⚙️ CodeRabbit configuration file
tests/**:Testing Guide
Comprehensive testing reference for Safe-Synthesizer developers. Covers commands, markers, test data, fixtures, and gotchas.
Read First
tests/conftest.py-- auto-marking,load_test_dataset/load_test_dataframe,fixture_mock_processorpatternpytest.ini-- markers, asyncio, timeouttests/evaluation/conftest.py-- most complex: Faker-basedmake_df, nullable dtype conversiontests/generation/conftest.py-- JSONL/schema fixtures,fixture_valid_iris_dataset_jsonl_and_schemaRunning Tests
All mise test tasks, grouped by scope:
mise run test # Unit (excludes slow, e2e, and smoke) mise run test:unit-slow # Unit tests including slow (excludes e2e and smoke) mise run test:smoke # CPU smoke tests (~few min, no GPU required) mise run test:smoke:gpu # All staged GPU smoke tests (requires CUDA) mise run test:smoke:gpu:train-only mise run test:smoke:gpu:generation mise run test:smoke:gpu:resume mise run test:smoke:gpu:structured-generation mise run test:smoke:gpu:timeseries mise run test:smoke:gpu:smollm2 mise run test:e2e # All e2e (requires CUDA) -- runs default + dp mise run test:e2e:default # e2e default (no-DP) tests only mise run test:e2e:dp # e2e DP tests only mise run test:ci # CI unit tests with coverage (excludes slow, e2e, gpu, smoke) mise run test:ci-slow # CI slow tests with coverage mise run test:ci-container # CI tests in a Linux container (Docker/Podman)Run a single test:
uv run --frozen pytest tests/path/test_file.py::test_name -vvs -n0Test runner:
uv run --frozen pytest -n auto --dist loadscope -vv...
Files:
tests/training/test_dp_linear.pytests/data_processing/test_distributions.pytests/data_processing/test_dates.pytests/config/test_autoconfig.pytests/cli/test_datasets.pytests/evaluation/components/test_attribute_inference_protection.pytests/evaluation/components/test_membership_inference_protection.py
tests/**/*.py
📄 CodeRabbit inference engine (tests/TESTING.md)
tests/**/*.py: Auto-mark tests based on file path: tests under/e2e/gete2emarker, tests under/smoke/getsmokemarker, all others getunitmarker (only if no category marker already present)
Every test should have exactly one category marker:unit,smoke, ore2e
Usepytest.mark.requires_gpumodifier on tests that need CUDA hardware
Usepytest.mark.vllmon tests using vLLM generation backend and ensure each vLLM test file runs in its own process for GPU memory isolation
Usepytest.mark.slowon long-running tests
Usepytest.mark.smollm2for SmolLM2 Hub download tests to enable process isolation
Usepytest.mark.noautouseto skip autouse fixtures for specific tests
Useload_test_dataset(filename)helper to load test datasets fromtests/stub_datasets/as HuggingFaceDatasetobjects
Useload_test_dataframe(filename)helper to load test data files fromtests/stub_datasets/as pandas DataFrames
Convert pandas columns to nullable dtypes (pd.Int64Dtype(),pd.BooleanDtype()) before assigningnp.nanvalues
Usefake.seed_instance(seed)andrandom.seed(seed)together for Faker-based test data reproducibility
When sharing methods across multiple test files, define them inconftest.pyand import them using relative imports (e.g.,from .conftest import train_with_sdk); note that importing from other test files liketests/cli/helpers.pydoes not work
Usefixture_mock_processororfixture_mock_processor_without_valid_recordsfor mocking ParsedResponse objects withvalid_records,invalid_records,errors, andprompt_numberfields
Usepytest.importorskipto gate tests on optional dependencies that require specific extras (e.g.,sentence_transformers,vllm)
Run vLLM tests with separate pytest invocations (one per file) using-n 0(single process) for GPU memory isolation, or use staged mise tasks for CI visibility
Print statements are allowed in tests (ruffT201is suppressed fortests/directory) and should...
Files:
tests/training/test_dp_linear.pytests/data_processing/test_distributions.pytests/data_processing/test_dates.pytests/config/test_autoconfig.pytests/cli/test_datasets.pytests/evaluation/components/test_attribute_inference_protection.pytests/evaluation/components/test_membership_inference_protection.py
⚙️ CodeRabbit configuration file
Review tests against tests/TESTING.md. Check marker usage, fixture naming, tmp_path usage, determinism, and GPU/vLLM process-isolation requirements. Flag slop tests that only check that code runs, assert result is not None when stronger invariants exist, over-mock internal implementation details, patch around the bug instead of reproducing it, or add broad snapshot/golden churn without a clear contract. Flag change detector tests that fail on harmless refactors, formatting, record ordering, incidental wording, or private implementation details without demonstrating a behavior regression. Prefer existing fixtures or focused new fixtures for repeated setup; keep tests DRY when reasonable without making the behavior under test opaque. print() is allowed in tests.
Files:
tests/training/test_dp_linear.pytests/data_processing/test_distributions.pytests/data_processing/test_dates.pytests/config/test_autoconfig.pytests/cli/test_datasets.pytests/evaluation/components/test_attribute_inference_protection.pytests/evaluation/components/test_membership_inference_protection.py
**/*.{md,markdown}
📄 CodeRabbit inference engine (.cursor/rules/agent-markdown-style.mdc)
**/*.{md,markdown}: Bold is acceptable only in markdown tables where it's the conventional way to mark header-like cells in the body
Use##headers to segment markdown sections instead of bold text
Use--(em-dash) instead of-(hyphen) for asides in markdown
Files:
STYLE_GUIDE.md
**/*.md
📄 CodeRabbit inference engine (STYLE_GUIDE.md)
**/*.md: Do not use decorative bold text in Markdown body text, list items, or docstrings; use headers, list markers, colons, and backticks for structure instead.
Use--for asides in Markdown and single backticks for code identifiers, paths, and CLI commands.
For Mermaid diagrams in Markdown, use node IDs without spaces, quote labels that contain special characters, and do not use explicit colors or styles.
Files:
STYLE_GUIDE.md
tests/cli/**/*.py
📄 CodeRabbit inference engine (tests/TESTING.md)
Use
mock_workdir(tmp_path)helper in CLI tests to create temporary Workdir instances
Files:
tests/cli/test_datasets.py
src/nemo_safe_synthesizer/evaluation/**/*.py
⚙️ CodeRabbit configuration file
Treat evaluation changes as correctness-sensitive. Check metric inputs, holdout usage, privacy metric semantics, report data shape, missing-data handling, and whether unavailable metrics fail or degrade intentionally.
Files:
src/nemo_safe_synthesizer/evaluation/reports/multimodal/multimodal_report.pysrc/nemo_safe_synthesizer/evaluation/components/multi_modal_figures.pysrc/nemo_safe_synthesizer/evaluation/components/membership_inference_protection.pysrc/nemo_safe_synthesizer/evaluation/components/attribute_inference_protection.py
src/nemo_safe_synthesizer/data_processing/**/*.py
⚙️ CodeRabbit configuration file
Review for data-contract regressions. Check input/training/test/synthetic naming, group boundaries, token-budget math, record ordering, schema and column validation, nullable dtypes, and deterministic behavior.
Files:
src/nemo_safe_synthesizer/data_processing/actions/distributions.pysrc/nemo_safe_synthesizer/data_processing/actions/dates.py
🧠 Learnings (2)
📓 Common learnings
Learnt from: CR
Repo: NVIDIA-NeMo/Safe-Synthesizer
Timestamp: 2026-06-24T17:38:02.057Z
Learning: Use American English spelling in code and documentation (for example, "initialize" not "initialise", "recognize" not "recognise", "color" not "colour").
Learnt from: CR
Repo: NVIDIA-NeMo/Safe-Synthesizer
Timestamp: 2026-06-24T17:38:02.057Z
Learning: Treat `__all__` as the public API surface, and treat identifiers with a leading `_` as private and changeable without notice.
Learnt from: CR
Repo: NVIDIA-NeMo/Safe-Synthesizer
Timestamp: 2026-06-24T17:38:02.057Z
Learning: Follow local consistency in the surrounding code even when it differs from the style guide, and migrate legacy code toward the conventions here when practical.
📚 Learning: 2026-05-27T22:20:37.354Z
Learnt from: kendrickb-nvidia
Repo: NVIDIA-NeMo/Safe-Synthesizer PR: 520
File: tests/generation/test_vllm_backend.py:556-587
Timestamp: 2026-05-27T22:20:37.354Z
Learning: In NVIDIA-NeMo/Safe-Synthesizer, `tests/conftest.py`’s `pytest_collection_modifyitems` hook applies pytest category markers automatically based on each test file’s path: tests under `/e2e/` get `pytest.mark.e2e`, tests under `/smoke/` get `pytest.mark.smoke`, and all other tests get `pytest.mark.unit`. Therefore, when reviewing pytest tests outside `tests/e2e/` and `tests/smoke/`, do not flag missing explicit `pytest.mark.unit` decorators on test classes/functions as an issue (the hook will add them during collection). If a new test directory/category is introduced, ensure the hook is updated so it’s categorized correctly.
Applied to files:
tests/training/test_dp_linear.pytests/data_processing/test_distributions.pytests/data_processing/test_dates.pytests/config/test_autoconfig.pytests/cli/test_datasets.pytests/evaluation/components/test_attribute_inference_protection.pytests/evaluation/components/test_membership_inference_protection.py
🪛 LanguageTool
STYLE_GUIDE.md
[style] ~144-~144: Consider using “incompatible” to avoid wordiness.
Context: ...ot express because the narrowed type is not compatible with the input type. ```python from ty...
(NOT_ABLE_PREMIUM)
🪛 Ruff (0.15.18)
tests/training/test_dp_linear.py
[warning] 41-41: Pattern passed to match= contains metacharacters but is neither escaped nor raw
(RUF043)
[warning] 46-46: Pattern passed to match= contains metacharacters but is neither escaped nor raw
(RUF043)
src/nemo_safe_synthesizer/utils.py
[warning] 137-137: zip() without an explicit strict= parameter
Add explicit value for parameter strict=
(B905)
🔇 Additional comments (3)
src/nemo_safe_synthesizer/telemetry.py (1)
26-29: LGTM!Also applies to: 74-81, 170-265
src/nemo_safe_synthesizer/data_processing/actions/distributions.py (1)
79-92: LGTM!tests/data_processing/test_distributions.py (1)
9-27: LGTM!
Signed-off-by: Aaron Gonzales <aagonzales@nvidia.com>
Signed-off-by: Aaron Gonzales <aagonzales@nvidia.com>
Signed-off-by: Aaron Gonzales <aagonzales@nvidia.com>
|
|
||
| # As we process the attack dataset, we'll accumulate for each column the number of | ||
| # correct and incorrect predictions | ||
| training_columns = [str(column) for column in training_df.columns] |
There was a problem hiding this comment.
did this just need to be defined earlier?
There was a problem hiding this comment.
Yes, the intent was to define training_columns earlier so the quasi-identifier combinations and the prediction counters use the same column identity instead of mixing original column objects with stringified names.
Summary
Test plan
Related issue: #614
Summary by CodeRabbit
pandas.DataFrameand fails fast with a clear error when they don’t.TypeIstype narrowing.tytool version.