-
Notifications
You must be signed in to change notification settings - Fork 354
Add exp-test-smell-detection skill and evaluation #491
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Merged
Merged
Changes from all commits
Commits
Show all changes
7 commits
Select commit
Hold shift + click to select a range
715a4e8
Add exp-test-smell-detection skill and evaluation
Evangelink 6641114
Improve integration test fixture and non-activation timeout
Evangelink c80799e
Fix test projects
Evangelink 20d79b4
Merge branch 'main' into dev/amauryleve/test-smells
Evangelink c1856de
Add domains for testsmells
Evangelink 40e9c6e
Merge branch 'main' into dev/amauryleve/test-smells
Evangelink de50b2d
Remove extra domains
Evangelink File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
196 changes: 196 additions & 0 deletions
196
plugins/dotnet-experimental/skills/exp-test-smell-detection/SKILL.md
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,196 @@ | ||
| --- | ||
| name: exp-test-smell-detection | ||
| description: "Detects test smells — bad programming practices in test code that indicate design problems and reduce test effectiveness. Use when the user asks to find test smells, review test quality, audit test health, identify problematic test patterns, or detect anti-patterns in test suites. Produces a categorized report with severity, locations, and concrete fix suggestions. Works with any test framework and language. DO NOT USE FOR: writing new tests (use writing-mstest-tests), evaluating assertion quality specifically (use exp-assertion-quality), or detecting boilerplate duplication (use exp-test-boilerplate-detection)." | ||
| --- | ||
|
|
||
| # Test Smell Detection | ||
|
|
||
| Analyze test code to detect test smells — symptoms of bad design or implementation decisions that make tests harder to understand, more fragile, less effective at catching bugs, or more expensive to maintain. Produce a severity-ranked report of findings with specific locations and actionable fixes. | ||
|
|
||
| ## Why Test Smells Matter | ||
|
|
||
| Test smells erode confidence in a test suite and inflate maintenance costs: | ||
|
|
||
| | Problem | Consequence | | ||
| |---------|-------------| | ||
| | Tests with conditional logic | Some paths never execute — hidden testing gaps | | ||
| | Tests that depend on external resources | Flaky failures, slow execution, environment coupling | | ||
| | Tests that sleep to wait for results | Non-deterministic timing, slow suites, false failures | | ||
| | Tests without assertions | False confidence — coverage looks good but nothing is verified | | ||
| | Tests that call many production methods | Hard to diagnose failures, unclear what's being tested | | ||
| | Tests with magic numbers | Unreadable intent, unclear boundary conditions | | ||
| | Tests relying on ToString for comparison | Brittle to formatting changes, obscure failure messages | | ||
| | Tests with exception handling logic | Swallowed failures, tests that pass when they shouldn't | | ||
|
|
||
| ## When to Use | ||
|
|
||
| - User asks to find test smells or anti-patterns in test code | ||
| - User asks "are my tests well-written?" or "what's wrong with my tests?" | ||
| - User wants a test quality audit or health check | ||
| - User asks for a review of test design or structure | ||
| - User suspects tests are fragile, flaky, or giving false confidence | ||
|
|
||
| ## When Not to Use | ||
|
|
||
| - User wants to evaluate assertion diversity specifically (use `exp-assertion-quality`) | ||
| - User wants to find duplicated boilerplate across tests (use `exp-test-boilerplate-detection`) | ||
| - User wants to write new tests from scratch (help them directly) | ||
| - User wants to fix a specific failing test (diagnose and fix directly) | ||
|
|
||
| ## Inputs | ||
|
|
||
| | Input | Required | Description | | ||
| |-------|----------|-------------| | ||
| | Test code | Yes | One or more test files or a test project directory to analyze | | ||
| | Production code | No | The code under test, for context on whether patterns are justified | | ||
|
|
||
| ## Workflow | ||
|
|
||
| ### Step 1: Gather the test code | ||
|
|
||
| Read all test files the user provides. If the user points to a directory or project, scan for all test files by looking for test framework markers — see [extensions/dotnet.md](extensions/dotnet.md) for .NET-specific markers. | ||
|
|
||
| For a thorough audit, also consult the [extended smell catalog](references/test-smell-catalog.md) which covers 9 additional smell types beyond the core 10 below. | ||
|
|
||
| ### Step 2: Scan for test smells | ||
|
|
||
| For each test method and class, check for the following smell categories: | ||
|
|
||
| #### Smell 1: Conditional Test Logic | ||
|
|
||
| Test methods containing `if`, `else`, `switch`, ternary (`? :`), `for`, `foreach`, or `while` statements. Control flow in tests means some paths may never execute, hiding gaps. | ||
|
|
||
| **Severity:** High | ||
| **Detection:** Any control flow statement inside a test method body. | ||
| **Exception:** `foreach` used solely to assert every item in a known collection is acceptable when the assertion is the loop body. | ||
|
|
||
| #### Smell 2: Mystery Guest | ||
|
|
||
| Tests that depend on external resources — files on disk, databases, network endpoints, environment variables — without making the dependency explicit or using test doubles. | ||
|
|
||
| **Severity:** High | ||
| **Detection:** Test methods that read files, open database connections, make HTTP requests (without a test handler), read environment variables, or use hard-coded file paths. | ||
| **Exception:** In-memory fakes or test-specific handlers are fine. | ||
|
|
||
| #### Smell 3: Sleepy Test | ||
|
|
||
| Tests that call sleep or delay functions to wait for a condition. These introduce non-deterministic timing and slow down the suite. | ||
|
|
||
| **Severity:** High | ||
| **Detection:** Calls to sleep/delay functions inside test methods. See [extensions/dotnet.md](extensions/dotnet.md) for .NET-specific patterns. | ||
|
|
||
| #### Smell 4: Assertion-Free Test (Unknown Test) | ||
|
|
||
| Tests that execute code but never assert anything. Test frameworks report these as passing even if the code is completely broken, as long as no exception is thrown. | ||
|
|
||
| **Severity:** High | ||
| **Detection:** A test method with no assertion calls (framework-specific: `Assert.*`, `expect()`, `assert`, `Should*`, etc.) and no expected-exception annotation. | ||
| **Calibration:** A method named `*_DoesNotThrow` or `*_NoException` is implicitly asserting no exception — still flag it but note it may be intentional. | ||
|
|
||
| #### Smell 5: Eager Test | ||
|
|
||
| A test method that calls many different production methods, making it unclear what behavior is being tested. When it fails, diagnosis is difficult because the failure could stem from any of the calls. | ||
|
|
||
| **Severity:** Medium | ||
| **Detection:** A test method that calls 4+ distinct methods on the production object (excluding setup/construction). Count unique method names, not call count. | ||
| **Calibration:** Integration tests or workflow tests may legitimately call multiple methods — note this as a possible exception for end-to-end scenarios. | ||
|
|
||
| #### Smell 6: Magic Number Test | ||
|
|
||
| Assertions that contain unexplained numeric literals. The intent of `Assert.AreEqual(42, result)` is unclear without context — what does 42 represent? | ||
|
|
||
| **Severity:** Medium | ||
| **Detection:** Numeric literals (other than 0, 1, -1, and the literal used in the test name) appearing as `expected` parameters in assertion methods. | ||
| **Calibration:** Small integers in context (like count checks `Assert.AreEqual(3, list.Count)` where 3 items were just added) are acceptable — only flag when the number's meaning is genuinely unclear. | ||
|
|
||
| #### Smell 7: Sensitive Equality | ||
|
|
||
| Tests that use `ToString()` for comparison or assertion. If the `ToString()` implementation changes, the test breaks even though the actual behavior is correct. | ||
|
|
||
| **Severity:** Medium | ||
| **Detection:** `Assert.AreEqual(expected, obj.ToString())`, or `.ToString()` appearing inside an assertion parameter. | ||
|
|
||
| #### Smell 8: Exception Handling in Tests | ||
|
|
||
| Tests that contain `try`/`catch` blocks or `throw` statements. This typically means the test is manually managing exceptions rather than using the framework's built-in exception assertion facilities. | ||
|
|
||
| **Severity:** Medium | ||
| **Detection:** `try`/`catch` or `throw`/`raise` statements inside a test method. | ||
| **Exception:** `catch` blocks that capture an exception for further assertion are a lesser concern — note but don't flag as high severity. | ||
|
|
||
| #### Smell 9: General Fixture (Over-broad Setup) | ||
|
|
||
| The test setup method or constructor initializes fields that are not used by every test method. This means each test pays the cost of setting up objects it doesn't need. | ||
|
|
||
| **Severity:** Low | ||
| **Detection:** Fields initialized in setup that are referenced by fewer than half the test methods in the class. | ||
|
|
||
| #### Smell 10: Ignored/Disabled Test | ||
|
|
||
| Tests marked as skipped or disabled. These add overhead and clutter, and the underlying issue they were disabled for may never be addressed. | ||
|
|
||
| **Severity:** Low | ||
| **Detection:** Skip/ignore annotations or conditional compilation that disables a test. See [extensions/dotnet.md](extensions/dotnet.md) for framework-specific skip attributes. | ||
|
|
||
| ### Step 3: Apply calibration rules | ||
|
|
||
| Before reporting, calibrate findings to avoid false positives: | ||
|
|
||
| - **Integration tests have different norms.** A test class clearly marked as integration (by name, annotation, or category) legitimately uses external resources, calls multiple methods, and may use delays for async coordination. Downgrade Mystery Guest, Eager Test, and Sleepy Test severity for integration tests — note them but don't flag as problems. | ||
| - **Simple loop-assert patterns are fine.** Iterating a collection to assert on every item is readable and correct. Only flag loops with complex branching logic. | ||
| - **Context matters for magic numbers.** A count assertion right after adding a known number of items is self-documenting. Only flag numbers whose meaning requires looking at production code to understand. | ||
| - **Inconclusive/pending markers are not assertion-free.** Tests explicitly marked as incomplete should be flagged as Ignored Test, not Assertion-Free. | ||
| - **Capture-and-assert exception patterns are borderline.** Try/catch patterns that capture an exception then assert on its properties are ugly but functional. Note as a smell and suggest the framework's built-in exception assertion instead of calling it broken. | ||
| - **If the test suite is clean, say so.** A report finding few or no smells is perfectly valid. | ||
|
|
||
| ### Step 4: Report findings | ||
|
|
||
| Present the analysis in this structure: | ||
|
|
||
| 1. **Summary Dashboard** — Quick overview: | ||
| ``` | ||
| | Severity | Smell Count | Affected Tests | | ||
| |----------|-------------|----------------| | ||
| | High | 3 | 7 | | ||
| | Medium | 2 | 4 | | ||
| | Low | 1 | 2 | | ||
| | Total | 6 | 13 | | ||
| ``` | ||
|
|
||
| 2. **Findings by Severity** — For each smell found: | ||
| - Smell name and category | ||
| - Severity level with rationale | ||
| - Affected test methods (file and method name) | ||
| - Code snippet showing the smell | ||
| - Concrete fix: show what the code should look like after remediation | ||
| - Risk if left unfixed | ||
|
|
||
| 3. **Smell-Free Patterns** — If any test methods are well-written, briefly acknowledge this. Highlighting what's good helps the user understand the contrast. | ||
|
|
||
| 4. **Prioritized Remediation Plan** — Rank fixes by: | ||
| - Impact (high-severity smells affecting many tests first) | ||
| - Effort (quick fixes before refactoring) | ||
| - Risk (fixes that prevent false-passes before cosmetic improvements) | ||
|
|
||
| ## Validation | ||
|
|
||
| - [ ] Every finding includes the specific test method name and file location | ||
| - [ ] Every finding includes a code snippet showing the smell in context | ||
| - [ ] Every finding includes a concrete fix example (not just "fix this") | ||
| - [ ] Integration tests are not penalized for patterns that are appropriate for their scope | ||
| - [ ] Simple foreach-assert loops are not flagged as conditional test logic | ||
| - [ ] Contextually obvious numbers are not flagged as magic numbers | ||
| - [ ] If the test suite is clean, the report says so upfront | ||
| - [ ] Severity levels are justified, not arbitrary | ||
|
|
||
| ## Common Pitfalls | ||
|
|
||
| | Pitfall | Solution | | ||
| |---------|----------| | ||
| | Flagging integration tests for using real resources | Check for integration test markers and adjust severity accordingly | | ||
| | Flagging loop-over-collection-assert as conditional logic | Only flag loops with branching or complex logic, not assertion iterations | | ||
| | Flagging obvious count assertions after adding N items | Consider the immediate context — self-documenting numbers are fine | | ||
| | Missing framework-specific assertion syntax | Consult [extensions/dotnet.md](extensions/dotnet.md) for .NET framework assertion and skip APIs | | ||
| | Over-flagging try/catch that captures for assertion | Distinguish swallowed exceptions from capture-and-assert patterns | | ||
| | Treating skip annotations with reasons same as bare skips | Note that reasoned skips are less concerning than unexplained ones | | ||
| | Flagging `DoesNotThrow`-style tests as assertion-free | These implicitly assert no exception — note but acknowledge the intent | | ||
111 changes: 111 additions & 0 deletions
111
plugins/dotnet-experimental/skills/exp-test-smell-detection/extensions/dotnet.md
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,111 @@ | ||
| # .NET Extension | ||
|
|
||
| Language-specific detection patterns for .NET test frameworks (MSTest, xUnit, NUnit, TUnit). | ||
|
|
||
| ## Test File Identification | ||
|
|
||
| | Framework | Test class markers | Test method markers | | ||
| | --------- | ------------------ | ------------------- | | ||
| | MSTest | `[TestClass]` | `[TestMethod]`, `[DataTestMethod]` | | ||
| | xUnit | _(none — convention-based)_ | `[Fact]`, `[Theory]` | | ||
| | NUnit | `[TestFixture]` | `[Test]`, `[TestCase]`, `[TestCaseSource]` | | ||
| | TUnit | `[ClassDataSource]` | `[Test]` | | ||
|
|
||
| ## Assertion APIs by Framework | ||
|
|
||
| | Category | MSTest | xUnit | NUnit | | ||
| | -------- | ------ | ----- | ----- | | ||
| | Equality | `Assert.AreEqual` | `Assert.Equal` | `Assert.That(x, Is.EqualTo(y))` | | ||
| | Boolean | `Assert.IsTrue` / `Assert.IsFalse` | `Assert.True` / `Assert.False` | `Assert.That(x, Is.True)` | | ||
| | Null | `Assert.IsNull` / `Assert.IsNotNull` | `Assert.Null` / `Assert.NotNull` | `Assert.That(x, Is.Null)` | | ||
| | Exception | `Assert.Throws<T>()` / `Assert.ThrowsExactly<T>()` | `Assert.Throws<T>()` | `Assert.That(() => ..., Throws.TypeOf<T>())` | | ||
|
Evangelink marked this conversation as resolved.
|
||
| | Collection | `CollectionAssert.Contains` | `Assert.Contains` | `Assert.That(col, Has.Member(x))` | | ||
| | String | `StringAssert.Contains` | `Assert.Contains(str, sub)` | `Assert.That(str, Does.Contain(sub))` | | ||
| | Type | `Assert.IsInstanceOfType` | `Assert.IsAssignableFrom` | `Assert.That(x, Is.InstanceOf<T>())` | | ||
| | Inconclusive | `Assert.Inconclusive()` | _skip via `[Fact(Skip)]`_ | `Assert.Inconclusive()` | | ||
| | Fail | `Assert.Fail()` | `Assert.Fail()` (.NET 10+) | `Assert.Fail()` | | ||
|
Evangelink marked this conversation as resolved.
|
||
|
|
||
| Third-party assertion libraries: `Should*` (Shouldly), `.Should()` (FluentAssertions / AwesomeAssertions), `Verify()` (Verify). | ||
|
|
||
| ## Sleep/Delay Patterns | ||
|
|
||
| | Pattern | Example | | ||
| | ------- | ------- | | ||
| | Thread sleep | `Thread.Sleep(2000)` | | ||
| | Task delay | `await Task.Delay(1000)` | | ||
| | SpinWait | `SpinWait.SpinUntil(() => condition, timeout)` | | ||
|
|
||
| ## Skip/Ignore Annotations | ||
|
|
||
| | Framework | Annotation | With reason | | ||
| | --------- | ---------- | ----------- | | ||
| | MSTest | `[Ignore]` | `[Ignore("reason")]` | | ||
| | xUnit | `[Fact(Skip = "reason")]` | _(reason is required)_ | | ||
| | NUnit | `[Ignore("reason")]` | _(reason is required)_ | | ||
| | TUnit | `[Skip("reason")]` | _(reason is required)_ | | ||
| | Conditional | `#if false` / `#if NEVER` | _(no reason possible)_ | | ||
|
|
||
| ## Exception Handling — Idiomatic Alternatives | ||
|
|
||
| When a test uses `try`/`catch` to verify exceptions, suggest the framework-native alternative: | ||
|
|
||
| **MSTest:** | ||
|
|
||
| ```csharp | ||
| // Instead of try/catch (matches exact type): | ||
| var ex = Assert.ThrowsExactly<InvalidOperationException>( | ||
| () => processor.ProcessOrder(emptyOrder)); | ||
| Assert.AreEqual("Order must contain at least one item", ex.Message); | ||
|
|
||
| // Or (also matches derived types): | ||
| var ex = Assert.Throws<InvalidOperationException>( | ||
| () => processor.ProcessOrder(emptyOrder)); | ||
| Assert.AreEqual("Order must contain at least one item", ex.Message); | ||
| ``` | ||
|
|
||
| **xUnit:** | ||
|
|
||
| ```csharp | ||
| var ex = Assert.Throws<InvalidOperationException>( | ||
| () => processor.ProcessOrder(emptyOrder)); | ||
| Assert.Equal("Order must contain at least one item", ex.Message); | ||
| ``` | ||
|
|
||
| **NUnit:** | ||
|
|
||
| ```csharp | ||
| var ex = Assert.Throws<InvalidOperationException>( | ||
| () => processor.ProcessOrder(emptyOrder)); | ||
| Assert.That(ex.Message, Is.EqualTo("Order must contain at least one item")); | ||
| ``` | ||
|
|
||
| ## Mystery Guest — Common .NET Patterns | ||
|
|
||
| | Smell indicator | What to look for | | ||
| | --------------- | ---------------- | | ||
| | File system | `File.ReadAllText`, `File.Exists`, `File.WriteAllBytes`, `Directory.GetFiles`, `Path.Combine` with hard-coded paths | | ||
| | Database | `SqlConnection`, `DbContext` (without in-memory provider), `SqlCommand` | | ||
| | Network | `HttpClient` without `HttpMessageHandler` override, `WebRequest`, `TcpClient` | | ||
| | Environment | `Environment.GetEnvironmentVariable`, `Environment.CurrentDirectory` | | ||
| | Acceptable | `MemoryStream`, `StringReader`, `InMemory` database providers, custom `DelegatingHandler` | | ||
|
|
||
| ## Integration Test Markers | ||
|
|
||
| Recognize these as integration tests (adjust smell severity accordingly): | ||
|
|
||
| - Class name contains `Integration`, `E2E`, `EndToEnd`, or `Acceptance` | ||
| - `[TestCategory("Integration")]` (MSTest) | ||
| - `[Trait("Category", "Integration")]` (xUnit) | ||
| - `[Category("Integration")]` (NUnit) | ||
| - Project name ending in `.IntegrationTests` or `.E2ETests` | ||
|
|
||
| ## Setup/Teardown Methods | ||
|
|
||
| | Framework | Setup | Teardown | | ||
| | --------- | ----- | -------- | | ||
| | MSTest | `[TestInitialize]` or constructor | `[TestCleanup]` or `IDisposable.Dispose` / `IAsyncDisposable.DisposeAsync` | | ||
| | xUnit | constructor | `IDisposable.Dispose` / `IAsyncDisposable.DisposeAsync` | | ||
| | NUnit | `[SetUp]` | `[TearDown]` | | ||
| | MSTest (class) | `[ClassInitialize]` | `[ClassCleanup]` | | ||
| | NUnit (class) | `[OneTimeSetUp]` | `[OneTimeTearDown]` | | ||
| | xUnit (class) | `IClassFixture<T>` | fixture's `Dispose` | | ||
Oops, something went wrong.
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
Uh oh!
There was an error while loading. Please reload this page.