[inflight/Candidate] Fix Android Shell Flyout UI tests failure in Candidate PR #36411 - #36672
Conversation
|
Azure Pipelines: Successfully started running 1 pipeline(s). There may be pipelines that require an authorized user to comment /azp run to run. |
|
/azp run maui-pr-uitests , maui-pr-devicetests |
|
Azure Pipelines: Successfully started running 1 pipeline(s). There may be pipelines that require an authorized user to comment /azp run to run. |
|
Azure Pipelines: Successfully started running 2 pipeline(s). |
Tests Failure Analysis
Test Failure Review: Needs human investigation - click to expandOverall verdict: Needs human investigation The deterministic gate caps this PR at Coverage: 175 checks · 96 passing · 79 failing · 0 pending · 0 inaccessible · 2 unmapped · 133 unexplained build legs · 0 unaccounted failing checks · 1 aborted failing checks · 0 canceled-build checks · 3 device-test unverified · 252 unattributed · 0 regressed-vs-base. Deterministic ceiling: Needs human investigation — 2 failing checks have no inspectable AzDO build evidence (
Recommended actionDo not merge on this CI run. Manually inspect the unmapped red checks and the 133 unexplained failed build legs first, then rerun or re-gather after confirming the Android Shell UI-test failures and device-test Helix results; only the Evidence details
Visual failure comparisonsFull-resolution CI baseline, actual, and diff images are embedded below. They supplement the failure classification and do not change the deterministic verdict ceiling.
|
| CI baseline | Fresh PR actual | CI diff |
|---|---|---|
![]() |
![]() |
![]() |
VerifyBorderWithNullStrokeDashArray - android - Needs human investigation - visual comparison
CI reported 11.03% difference in build 1517762.
Relationship to PR: Needs human investigation - No decisive exact test-and-platform baseline attribution was available.
| CI baseline | Fresh PR actual | CI diff |
|---|---|---|
![]() |
![]() |
![]() |
VerifyBorderWithStrokeDashArrayValue - windows - Needs human investigation - visual comparison
CI reported 2.99% difference in build 1517762.
Relationship to PR: Needs human investigation - No decisive exact test-and-platform baseline attribution was available.
| CI baseline | Fresh PR actual | CI diff |
|---|---|---|
![]() |
![]() |
![]() |
VerifyBorderWithNullStrokeDashArray - windows - Needs human investigation - visual comparison
CI reported 3.17% difference in build 1517762.
Relationship to PR: Needs human investigation - No decisive exact test-and-platform baseline attribution was available.
| CI baseline | Fresh PR actual | CI diff |
|---|---|---|
![]() |
![]() |
![]() |
ShouldFlyoutTextWrapsInLandscape - android - Needs human investigation - visual comparison
CI reported 7.96% difference in build 1517762.
Relationship to PR: Needs human investigation - No decisive exact test-and-platform baseline attribution was available.
| CI baseline | Fresh PR actual | CI diff |
|---|---|---|
| Baseline unavailable: ambiguous across multiple snapshot environments | ![]() |
![]() |
VerifySwipeViewWithLabelContentAndThreshold - android - Needs human investigation - visual comparison
CI reported 0.68% difference in build 1517762.
Relationship to PR: Needs human investigation - No decisive exact test-and-platform baseline attribution was available.
| CI baseline | Fresh PR actual | CI diff |
|---|---|---|
![]() |
![]() |
![]() |
VerifySwipeViewWithImageContentAndThreshold - android - Needs human investigation - visual comparison
CI reported 0.68% difference in build 1517762.
Relationship to PR: Needs human investigation - No decisive exact test-and-platform baseline attribution was available.
| CI baseline | Fresh PR actual | CI diff |
|---|---|---|
![]() |
![]() |
![]() |
Material3Editor_VerifyzEditorTextWhenAutoSizeTextChangesSet - android - Needs human investigation - visual comparison
CI reported 0.91% difference in build 1517762.
Relationship to PR: Needs human investigation - No decisive exact test-and-platform baseline attribution was available.
| CI baseline | Fresh PR actual | CI diff |
|---|---|---|
![]() |
![]() |
![]() |
VerifySwipeViewWithCollectionViewContentAndThreshold - android - Needs human investigation - visual comparison
CI reported 0.68% difference in build 1517762.
Relationship to PR: Needs human investigation - No decisive exact test-and-platform baseline attribution was available.
| CI baseline | Fresh PR actual | CI diff |
|---|---|---|
![]() |
![]() |
![]() |
Material3Editor_VerifyzEditorPlaceholderWithAutoSizeTextChanges - android - Needs human investigation - visual comparison
CI reported 1.38% difference in build 1517762.
Relationship to PR: Needs human investigation - No decisive exact test-and-platform baseline attribution was available.
| CI baseline | Fresh PR actual | CI diff |
|---|---|---|
![]() |
![]() |
![]() |
Material3Editor_VerifyzEditorPlaceholderWithAutoSizeDiabled - android - Needs human investigation - visual comparison
CI reported 0.75% difference in build 1517762.
Relationship to PR: Needs human investigation - No decisive exact test-and-platform baseline attribution was available.
| CI baseline | Fresh PR actual | CI diff |
|---|---|---|
![]() |
![]() |
![]() |
ItemImageSourceShouldBeVisible - ios - Needs human investigation - visual comparison
CI reported 1.71% difference in build 1517762.
Relationship to PR: Needs human investigation - No decisive exact test-and-platform baseline attribution was available.
| CI baseline | Fresh PR actual | CI diff |
|---|---|---|
| Baseline unavailable: ambiguous across multiple snapshot environments | ![]() |
![]() |
VerifyEditorTextWhenFontAttributesSet - windows - Needs human investigation - visual comparison
CI reported 0.60% difference in build 1517762.
Relationship to PR: Needs human investigation - No decisive exact test-and-platform baseline attribution was available.
| CI baseline | Fresh PR actual | CI diff |
|---|---|---|
![]() |
![]() |
![]() |
Material3Editor_VerifyEditorPlaceholderWithShadow - android - Needs human investigation - visual comparison
CI reported 0.58% difference in build 1517762.
Relationship to PR: Needs human investigation - No decisive exact test-and-platform baseline attribution was available.
| CI baseline | Fresh PR actual | CI diff |
|---|---|---|
![]() |
![]() |
![]() |
kubaflo
left a comment
There was a problem hiding this comment.
Could you please check if test failures are related?
@kubaflo , I have ensured the test locally, failures are not related to my fix. |
<!-- Please let the below note in for people that find this PR --> > [!NOTE] > Are you waiting for the changes in this PR to be merged? > It would be very helpful if you could [test the resulting artifacts](https://github.com/dotnet/maui/wiki/Testing-PR-Builds) from this PR and let us know in a comment if this change resolves your issue. Thank you! ### Description of Change Automates visual snapshot evidence in `/review tests`. When public AzDO results contain failed screenshot comparisons, the command now emits exactly one test-failure analysis comment containing bounded, expandable baseline/actual/diff panels. Visual evidence remains supplementary: it does not change `gate.verdictCeiling`, deterministic attribution, or the merge-readiness verdict. #### One-comment flow 1. Trusted pre-activation code discovers failed visual results through the public AzDO `resultsbybuild` API, including retry-suffixed attachments such as `Snapshot[1].png` and `Snapshot-diff[1].png`. 2. It resolves baselines from the exact source version tested by AzDO and maps runtime evidence to the correct snapshot directory (`ios-26`, `android-notch-36`, `mac`, or `windows`). 3. It streams and validates bounded PNG assets, then stores them on `review-tests-assets` using immutable commit-pinned `raw.githubusercontent.com` URLs. 4. The Copilot agent emits the normal single `add_comment` analysis payload with a trusted insertion marker. 5. A sealed post-step validates the published asset manifest and injects as many expandable comparison panels as fit into that same comment. Each collapsed panel shows a conservative relationship label: - `Likely PR-caused` for an exact test/platform base regression or directly changed snapshot/test; - `Likely unrelated` for an exact base/known-issue match without direct visual scope; - `Needs human investigation` for unmatched or mixed evidence. 6. Excess comparisons are summarized as omitted instead of creating a companion comment. The local `.github/scripts/Review-Tests.ps1 -PostComment` path uses the same merger. It also recognizes complete reports returned in Copilot's final response, preserving nested evidence code fences without wrapping a second title or badge section. ### Security and Failure Safety - PR text, logs, test names, attachment metadata, changed files, and visual labels remain untrusted input. - The merger script and visual context are copied to a root-owned location before the workflow checks out the untrusted PR branch. - The post-step runs without `COPILOT_GITHUB_TOKEN`, `GH_TOKEN`, or `GITHUB_TOKEN`. - AzDO attachment URLs must match the expected public project and attachment route. - Published assets are size-bounded, signature-checked PNGs with validated dimensions and repository paths. - Raw image URLs must match the exact repository, asset commit, PR directory, and safe filename. - Labels are HTML-escaped and `@` is neutralized before insertion. - Relationship labels use fixed trusted text. Untrusted attribution values are never rendered, and same-named snapshots changed on another platform do not count as PR scope. - The final body is checked panel-by-panel against conservative limits of 45 URLs, 10 mentions, and 60,000 UTF-16 characters, below gh-aw's throwing limits. - The analysis JSON update is atomic (written to a temp file, then renamed over the original). Invalid context, malformed output, missing analysis payloads, limit failures, and dry-run/noop output leave the original analysis unchanged. - The publisher never creates or patches PR comments; only the existing gh-aw `add_comment` payload is mutated. ### What NOT to Do - Do not use the ordinary anonymous AzDO test-runs listing for discovery; it redirects to sign-in. Use the public failed-results endpoint. - Do not resolve baselines from the current PR head; use the source version actually tested by the selected AzDO build. - Do not let the agent construct or trust visual asset URLs. - Do not publish a second companion comment; merge bounded panels into the single analysis payload. ### Validation - 98 focused Pester tests pass. - Changed PowerShell scripts parse successfully. - `gh aw compile copilot-review-tests --approve` completes without errors or warnings. - A real `agent_output.json` from gh-aw run [29674953402](https://github.com/dotnet/maui/actions/runs/29674953402) was replayed through the post-step: - one `add_comment` item remained one item; - five visual panels were inserted; - the final body contained 26 URLs, one mention, and 9,585 characters. ### Live Single-Comment Examples The exact local `/review tests` path from this branch posted or repaired these merged comments after the PRs' `/azp run` pipelines completed: | PR | Single merged result | Included evidence | Relationship labels | Final limits | | --- | --- | --- | --- | --- | | #36413 | [Test-failure analysis with visual panels](#36413 (comment)) | 5 panels / 15 images | 1 PR-caused, 4 investigate | 23 URLs, 13,156 chars | | #36631 | [Test-failure analysis with visual panels](#36631 (comment)) | 6 panels / 18 images | 2 PR-caused, 4 investigate | 31 URLs, 21,367 chars | | #36395 | [Test-failure analysis with visual panels](#36395 (comment)) | 11 panels / 33 images; 19 omitted | 11 investigate | 43 URLs, 19,858 chars | | #36404 | [Test-failure analysis with visual panels](#36404 (comment)) | 14 panels / 40 images; 81 omitted | 14 investigate | 45 URLs, 23,779 chars | | #35846 | [Test-failure analysis with visual panels](#35846 (comment)) | 10 panels / 30 images; 9 omitted | 10 investigate | 43 URLs, 22,496 chars | | #36277 | [Test-failure analysis with visual panels](#36277 (comment)) | 7 panels / 19 images | 3 PR-caused, 4 investigate | 31 URLs, 18,355 chars | | #36170 | [Test-failure analysis with visual panels](#36170 (comment)) | 11 panels / 33 images; 8 omitted | 11 investigate | 44 URLs, 23,180 chars | | #35578 | [Test-failure analysis with visual panels](#35578 (comment)) | 12 panels / 36 images; 50 omitted | 12 investigate | 44 URLs, 25,443 chars | | #36672 | [Test-failure analysis with visual panels](#36672 (comment)) | 14 panels / 40 images; 9 omitted | 14 investigate | 45 URLs, 25,915 chars | | #31755 | [Test-failure analysis with visual panels](#31755 (comment)) | 12 panels / 36 images; 3 omitted | 12 investigate | 44 URLs, 22,728 chars | | #34637 | [Test-failure analysis with visual panels](#34637 (comment)) | 9 panels / 27 images; 81 omitted | 9 investigate | 43 URLs, 22,325 chars | | #35156 | [Test-failure analysis with visual panels](#35156 (comment)) | 2 panels / 6 images | 1 PR-caused, 1 investigate | 19 URLs, 11,911 chars | | #35885 | [Test-failure analysis with no visual failures](#35885 (comment)) | 0 panels / 0 images | No visual failures detected | 8 URLs, 4,239 chars | | #36577 | [Test-failure analysis with visual panels](#36577 (comment)) | 1 panel / 3 images | 1 investigate | 44 URLs, 21,768 chars | | #36212 | [Test-failure analysis with no visual failures](#36212 (comment)) | 0 panels / 0 images | No visual failures detected | 5 URLs, 5,982 chars | Each result contains one `Tests Failure Analysis` title and one merged review marker. Across 114 rendered panels, all 336 embedded image URLs returned HTTP 200. Seven panels were safely classified as likely PR-caused; no panel in this sample had enough exact evidence to be safely classified as likely unrelated, so the remaining 107 stayed at `Needs human investigation`. Another 260 comparisons were omitted safely by the comment limits. The latest eight-example batch was regenerated concurrently, and #36672, #31755, #34637, #35156, #35885, #36577, and #36212 were added afterward. The current [`review-tests-assets` head](f937993) retains the full asset history. The protected `copilot-pat-pool` environment rejects feature-branch `workflow_dispatch` runs before job execution. The live local-runner examples validate comment generation and asset publication, while the real gh-aw output replay validates the workflow post-step mutation without weakening that branch protection. ### Issues Fixed N/A - reviewer workflow enhancement. --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com> Co-authored-by: kubaflo <kubaflo@users.noreply.github.com> Copilot-Session: a280b482-e102-4ca0-9ff9-1cfe1946e21f Copilot-Session: d00747b7-96f3-4e7a-8dfb-e3a48db04b2d Copilot-Session: 478d195b-20f3-4bc6-aeed-f6aa88b55fda








































Issue Details
Android Shell Flyout UI Test Failures in Candidate PR #36411 due to PR #34510
Description of Change
On Android, the Shell flyout was incorrectly applying bottom padding even when the flyout didn't reach the bottom safe area. The fix checks if the view actually overlaps the bottom safe area before applying the bottom inset padding.
Issues Fixed
Fixes #36625
Test failure details