Skip to content

feat(reliability): yt-dlp 3× exponential backoff + snapshot restore timeout - #850

Merged
LucasSantana-Dev merged 13 commits into
release/v2.11.0from
feat/reliability-improvements
May 14, 2026
Merged

LucasSantana-Dev merged 13 commits into
release/v2.11.0from
feat/reliability-improvements

Conversation

@LucasSantana-Dev

@LucasSantana-Dev LucasSantana-Dev commented May 14, 2026 •

Copy link
Copy Markdown
Owner

Summary

  • yt-dlp 3× exponential backoff: spawnYtDlp now retries up to 3 times with increasing timeouts (30 s → 45 s → 60 s). Network blips or CDN hiccups no longer cause permanent playback failures on the first transient error.
  • Snapshot restore timeout: restoreSnapshot() in the connection lifecycle handler is now wrapped in Promise.race(2 s) with a warn-and-continue fallback. A slow or hung DB call can no longer block the entire queue restore path indefinitely.
  • Jest test coverage: retry behavior (success on first attempt, retry-and-succeed, exhaust-all-retries, timeout-kill-and-retry) fully covered in ytdlpExtractor.test.ts.
  • ADR: docs/decisions/2026-05-14-discord-integration-testing-strategy.md documents the DEFER decision for in-CI voice integration testing and the rationale for these higher-ROI reliability fixes instead.

Test plan

  • pnpm test --filter=bot — all 24 ytdlp tests pass
  • Manual playback smoke: queue a YouTube track; confirm it plays after a simulated network hiccup (or trust the retry unit tests)
  • Verify no regression in lifecycleHandlers — session restore still works within 2 s under normal conditions

Summary by CodeRabbit

Release Notes

  • Chores

    • Updated Docker infrastructure configuration with port adjustments and unprivileged user execution
    • Added resource limits via Docker Compose anchors
    • Pinned webhook service image version
  • Bug Fixes

    • Enhanced music extraction resilience with automatic retry logic
    • Added timeout protection for session restoration to prevent indefinite hangs
  • Documentation

    • Added Architecture Decision Records documenting Docker configuration updates and testing strategy

Review Change Stack

LucasSantana-Dev and others added 12 commits May 13, 2026 20:32
Build context was including `.worktrees/` (3.4GB), `worktrees/` (707MB),
`.wt-specs/` (41MB), `.claude/` (115MB), and `.agents/` (183MB) — totalling
~4.5GB of duplicated trees and AI agent state shipped to the Docker daemon
on every build. None of this is needed inside any image.

Also exclude `archive/` and `downloads/` (host-only state).
The previous HEALTHCHECK was `node -e "console.log('Service is running')"`,
which always exits 0 regardless of bot state — orchestrators could never
detect a wedged or disconnected bot.

New check opens a raw TCP socket to Redis ($REDIS_HOST / $REDIS_PORT) and
sends the RESP PING command. Pass = +PONG within 3s; anything else fails.
This confirms (1) node can execute inside the container and (2) the bot's
critical Redis dependency is reachable from this container. No new deps.

Start period bumped 5s → 30s to account for `prisma migrate deploy` running
before the bot process starts.
`docker-compose.dev.yml` referenced `target: development` but no such
stage existed in `Dockerfile` — `docker compose -f docker-compose.dev.yml
up --build` would fail with 'failed to find target development'.

New `development` stage derives from `base-runtime` (already has ffmpeg /
opus / yt-dlp), adds native build tools, and runs `tsx watch` via
`npm run dev --workspace=packages/bot`. Compose still bind-mounts host
source over `/app`; node_modules installed on first run to populate the
anonymous volume.
Switch `Dockerfile.nginx` and `Dockerfile.frontend` from `nginx:alpine`
(runs as root to bind port 80) to `nginxinc/nginx-unprivileged:1.27-alpine`
(UID 101, listens on 8080 by default — no NET_BIND_SERVICE capability
needed).

Changes:
- nginx confs (frontend + reverse proxy): `listen 80` → `listen 8080`
- nginx reverse-proxy upstream: `http://frontend:80` → `http://frontend:8080`
- Dockerfile.frontend + Dockerfile.nginx: pinned image, EXPOSE 8080, added
  HEALTHCHECK via `wget --spider` (busybox wget ships in the base image).
- docker-compose.yml: port mapping `${NGINX_PORT:-8080}:80` → `:8080`.

DEPLOY ACTION REQUIRED: update `cloudflared/config-lucky.yml` on the
homelab so the tunnel ingress points at `http://nginx:8080` instead of
`http://nginx:80` before merging this PR. Host-side `NGINX_PORT`
default is unchanged (8080).
Compose previously had no memory or CPU limits on any service — a runaway
bot or backend process could OOM-kill the homelab host. Adds tiered
defaults via YAML anchors:

  small-svc  (frontend, nginx, webhook, cloudflared)   128m  / 0.25 cpu
  medium-svc (redis, backend)                          512m  / 0.5  cpu
  large-svc  (postgres, bot)                           1g    / 1.0  cpu

Bot + backend also get `env_file: .env` so any var missing from the
explicit `environment:` block falls back to .env at startup. Explicit
entries still take precedence, so behavior is unchanged for vars already
listed.

Cloudflared `user: root` removed — the official image's nonroot default
is sufficient for `tunnel run` with a mounted config dir.

`docker-compose config -q` passes.
The deploy webhook container has `/var/run/docker.sock` bind-mounted in
`docker-compose.yml`, which gives anything inside it effective root on the
host. Pulling `almir/webhook:latest` (last published 2026-02-12) every
rebuild made that surface vulnerable to silent upstream changes.

Pin to `2.8.3` (current latest, identical digest as `latest` at time of
this change). Also fix the inline-comment placement on `USER root` — it
was on the same line which is parsed differently across Docker versions.
… node 22

Three small but durable cleanups:

1. Drop the no-op `base-runtime-backend` stage. `production-backend` now
   derives directly from `node:${NODE_VERSION}` — the intermediate stage
   only set WORKDIR, which the production stage already does.

2. Replace `pip3 install --break-system-packages` for yt-dlp with a
   proper PEP-668-compliant venv at `/opt/ytdlp`, symlinked into
   `/usr/local/bin`. Same behavior, no warning suppression, easier to
   audit and upgrade. /root/.cache is cleaned in the same layer.

3. `packages/frontend/Dockerfile.dev` was using `node:24-alpine` while
   production frontend is pinned to `node:22-alpine` (PR #846). Aligned
   to 22 to avoid silent native-module drift. Also removed
   `npm cache clean --force` which defeated the BuildKit cache mount.
Captures the 8-commit Docker surface overhaul: motivation, decisions per
commit, consequences (including the cloudflared config-lucky.yml port
edit required on the homelab before merge), out-of-scope items, and
revisit triggers.
- deploy/Dockerfile: pin almir/webhook:2.8.3 by manifest digest
  sha256:f77cc91c91d1527b48052280af38b50791e842ad8dd291e9b360a0c13c9ca991.
  Docker Hub tags are mutable by default; the digest makes the supply-chain
  story (this container holds docker.sock) actually defensible.
- docs/decisions/2026-05-13-docker-overhaul.md: remove dangling reference
  to a separate audit doc that was never committed; ADR contains audit
  findings inline.
…imeout

- Replace single-shot yt-dlp execution with 3-attempt retry loop using
  exponential backoff (30s → 45s → 60s). Each attempt has its own
  independent timeout; a settled flag prevents double-resolve when the
  process closes after a timeout kill.
- Wrap restoreSnapshot() in Promise.race() with a 2-second deadline so a
  slow or hung DB call cannot block the entire queue restore path. Logs
  a warning and falls through to an empty queue on timeout.
…rategy

- Add 4 retry-behavior tests: first-attempt success, retry-on-failure,
  all-attempts-exhausted (expects throw), timeout-kills-and-retries.
  Moved discord-player mock to file scope to eliminate jest.resetModules()
  worker crashes.
- Add ADR 2026-05-14: defer in-CI voice integration testing; yt-dlp retry
  + lifecycle snapshot timeout + backend scaffolding are higher ROI.
@vercel

vercel Bot commented May 14, 2026 •

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
lucky Ready Ready Preview, Comment May 14, 2026 5:33pm

Request Review

@greptile-apps greptile-apps Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LucasSantana-Dev has reached the 50-review limit for trial accounts. To continue receiving code reviews, upgrade your plan.

@github-actions

Copy link
Copy Markdown

This PR exceeds the recommended size of 1000 lines. Please make sure you are NOT addressing multiple issues with one PR. Note this PR might be rejected due to its size.

@github-actions

Copy link
Copy Markdown

Failed to generate code suggestions for PR

@github-actions

Copy link
Copy Markdown

Size Change: 0 B

Total Size: 369 kB

ℹ️ View Unchanged
Filename Size
packages/frontend/dist/assets/ActionPanel-BHyuNdpr.js 402 B
packages/frontend/dist/assets/Admin-B6MFsHen.js 1.99 kB
packages/frontend/dist/assets/api-AUPUM2Qa.js 3.37 kB
packages/frontend/dist/assets/authStore-Cu1BkXvq.js 556 B
packages/frontend/dist/assets/AutoMessages-B4sp3ygQ.js 2.68 kB
packages/frontend/dist/assets/AutoMod-DrBCoACw.js 4.11 kB
packages/frontend/dist/assets/badge-vPK-qUOM.js 501 B
packages/frontend/dist/assets/Card-C_uHUkFW.js 509 B
packages/frontend/dist/assets/CommandsConfig-DrTRW2xr.js 1.47 kB
packages/frontend/dist/assets/Config-ijWHAfPt.js 1.73 kB
packages/frontend/dist/assets/CustomCommands-CU1vBphA.js 2.16 kB
packages/frontend/dist/assets/DashboardOverview-DB0sQtH5.js 3.39 kB
packages/frontend/dist/assets/dialog-DSHSwnxX.js 951 B
packages/frontend/dist/assets/EmbedBuilder-DdolUBZ4.js 3.34 kB
packages/frontend/dist/assets/Features-mfnCuLdA.js 754 B
packages/frontend/dist/assets/GuildAutomation-CZNrZEuH.js 2.92 kB
packages/frontend/dist/assets/guildStore-il-poOxg.js 791 B
packages/frontend/dist/assets/index-CdlbTRdI.css 17.1 kB
packages/frontend/dist/assets/index-DRmb6K-r.js 52.1 kB
packages/frontend/dist/assets/input-Yz475ZAJ.js 469 B
packages/frontend/dist/assets/label-BzaZE6x2.js 491 B
packages/frontend/dist/assets/Landing-DiU_pB7d.js 3.74 kB
packages/frontend/dist/assets/LastFm-CWNGd_qk.js 1.74 kB
packages/frontend/dist/assets/Levels-CPoenlFQ.js 2.63 kB
packages/frontend/dist/assets/Login-uP19mdGs.js 2.5 kB
packages/frontend/dist/assets/Lyrics-BBCUe7JA.js 1.33 kB
packages/frontend/dist/assets/Moderation-BxnKnHIA.js 3.85 kB
packages/frontend/dist/assets/Music-CvpKxleP.js 7.12 kB
packages/frontend/dist/assets/MusicConfig-CfomaIrN.js 1.61 kB
packages/frontend/dist/assets/PreferredArtists-DOHVMZJO.js 3.7 kB
packages/frontend/dist/assets/PrivacyPolicy-D3x5lQf7.js 1.38 kB
packages/frontend/dist/assets/ReactionRoles-CibF-0KT.js 1.89 kB
packages/frontend/dist/assets/rolldown-runtime-BYbx6iT9.js 471 B
packages/frontend/dist/assets/SectionHeader-B-f_yjhM.js 895 B
packages/frontend/dist/assets/select-DM_YBjDr.js 1.23 kB
packages/frontend/dist/assets/ServerLogs-Dkdtmhsl.js 2.91 kB
packages/frontend/dist/assets/ServerSettings-B3l_5dU3.js 4.22 kB
packages/frontend/dist/assets/ServersPage-V6JR3l_J.js 2.92 kB
packages/frontend/dist/assets/shim-Qmczh487.js 510 B
packages/frontend/dist/assets/Skeleton-WB3VdUho.js 238 B
packages/frontend/dist/assets/Spotify-CjGactC2.js 1.75 kB
packages/frontend/dist/assets/Starboard-w6YFyfV2.js 2.08 kB
packages/frontend/dist/assets/StatTile-B79rBPkY.js 639 B
packages/frontend/dist/assets/switch-HLSMi6pZ.js 546 B
packages/frontend/dist/assets/TermsOfService-B_oLkpxV.js 1.37 kB
packages/frontend/dist/assets/TrackHistory-ClprovMB.js 2.16 kB
packages/frontend/dist/assets/TwitchNotifications-C2jAQojn.js 2.44 kB
packages/frontend/dist/assets/useFeatures-CQZj2LNh.js 2.03 kB
packages/frontend/dist/assets/usePageMetadata-DVjH87aq.js 329 B
packages/frontend/dist/assets/utils-TE4_8V5I.js 149 B
packages/frontend/dist/assets/vendor-forms-VJTTJemd.js 25.6 kB
packages/frontend/dist/assets/vendor-radix-DKbKEiWx.js 38.9 kB
packages/frontend/dist/assets/vendor-react-DNaNL-Gj.js 55.6 kB
packages/frontend/dist/assets/vendor-state-B_fYO9zl.js 24.1 kB
packages/frontend/dist/assets/vendor-ui-D9aFDhli.js 65 kB

compressed-size-action

@coderabbitai

coderabbitai Bot commented May 14, 2026

Copy link
Copy Markdown
📝 Walkthrough

Walkthrough

This PR standardizes Docker infrastructure across the application by migrating services to unprivileged images, moving from port 80 to 8080, isolating yt-dlp in a Python venv, and implementing yt-dlp retry logic with timeout guards alongside session restore resilience in the bot.

Changes

Infrastructure Overhaul & Bot Reliability Improvements

Layer / File(s) Summary
Docker Build Context & Image Standardization
.dockerignore, Dockerfile, Dockerfile.frontend, Dockerfile.nginx, deploy/Dockerfile, packages/frontend/Dockerfile.dev, docs/decisions/2026-05-13-docker-overhaul.md
Docker build context excludes worktree and AI agent directories. Base-runtime refactored to install yt-dlp in a Python venv at /opt/ytdlp with ffmpeg/opus support, avoiding Alpine PEP 668 constraints. Frontend and nginx containers switched to nginxinc/nginx-unprivileged:1.27-alpine running on port 8080 as non-root. Webhook image pinned to digest 2.8.3. Frontend dev image aligned to node:22-alpine. ADR documents the Docker surface overhaul decision and commits.
Port Migration & Service Orchestration
nginx/frontend.conf, nginx/nginx.conf, docker-compose.yml
All nginx and frontend services migrated from port 80 to port 8080. Nginx configuration updated to route frontend upstream traffic to port 8080. Docker Compose introduces YAML anchors (x-small-svc, x-medium-svc, x-large-svc) for shared resource limits and restart policies. Bot and backend services load environment variables from shared .env file. Service port mappings and health checks updated throughout.
Bot Healthcheck & yt-dlp Resilience
Dockerfile, packages/bot/src/handlers/player/lifecycleHandlers.ts, packages/bot/src/utils/music/ytdlpExtractor/service.ts, packages/bot/tests/utils/music/ytdlpExtractor.test.ts, docs/decisions/2026-05-14-discord-integration-testing-strategy.md
Bot production healthcheck replaced with Redis TCP PING liveness probe instead of log-based check. Session snapshot restoration wrapped in 2-second timeout via Promise.race to prevent blocking on restore hangs. YtDlpExtractorService refactored with spawnYtDlp helper that enforces per-invocation timeout and kills process on deadline. executeYtDlp implements retry loop with exponential timeout backoff across multiple attempts. Test suite expanded with retry behavior validation covering first-success, retry-on-failure, exhausted retries, and timeout-triggered process kill. ADR documents deferred Discord voice integration testing, prioritizing reliability fixes for yt-dlp retry, session restore timeout guard, and backend route test scaffolding.

Sequence Diagram(s)

sequenceDiagram
  participant executeYtDlp
  participant spawnYtDlp
  participant Process as child_process
  participant Test as Test Framework
  
  executeYtDlp->>spawnYtDlp: retry attempt 1 (timeout1)
  spawnYtDlp->>Process: spawn yt-dlp
  Process-->>spawnYtDlp: stdout/stderr + exit code
  alt timeout elapsed
    spawnYtDlp->>Process: kill process
    spawnYtDlp-->>executeYtDlp: timeout error
    executeYtDlp->>spawnYtDlp: retry attempt 2 (timeout2)
  else success on close
    spawnYtDlp-->>executeYtDlp: success result
  end
  executeYtDlp-->>Test: final result or exhausted retries
Loading

Estimated code review effort

🎯 4 (Complex) | ⏱️ ~45 minutes

Possibly related PRs

  • LucasSantana-Dev/Lucky#427: Both PRs modify the bot container startup flow in Dockerfile to preserve or ensure prisma migrate deploy runs before bot launch.
  • LucasSantana-Dev/Lucky#550: Both PRs modify packages/bot/src/handlers/player/lifecycleHandlers.ts session-restore path—this PR adds timeout guards while the related PR changes queue metadata typing used by restore logic.

Suggested labels

bot, infra, size/l

🚥 Pre-merge checks | ✅ 4 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage ⚠️ Warning Docstring coverage is 0.00% which is insufficient. The required threshold is 80.00%. Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (4 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The PR title directly and clearly summarizes the two main reliability improvements: yt-dlp exponential backoff retry logic (3×) and snapshot restore timeout wrapping.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch feat/reliability-improvements

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (4)
Dockerfile (1)

116-120: ⚡ Quick win

Extract the healthcheck to a standalone script for maintainability.

The Redis TCP PING implementation is correct and handles all expected cases appropriately (RESP protocol is properly formatted, socket cleanup is implicit via process.exit, timeout margins are adequate, and partial reads are negligible for this use case). However, the one-liner is difficult to debug and test locally. Extract to scripts/healthcheck-bot.js to enable standalone testing and improve readability without changing logic.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@Dockerfile` around lines 116 - 120, Replace the inline HEALTHCHECK node
one-liner with a call to a new executable script (scripts/healthcheck-bot.js)
that contains the same Redis TCP PING logic (uses net.createConnection, writes
the RESP PING frame, listens for 'data' to check for "+PONG", handles 'error' by
exiting non-zero, and enforces a 3s timeout before exiting 1); add the script to
the image via COPY and ensure it is executable, then update the Dockerfile
HEALTHCHECK line to run node /app/scripts/healthcheck-bot.js (or the container
path you COPY to) so behavior and exit codes remain identical while making the
logic maintainable and testable locally.
packages/bot/src/handlers/player/lifecycleHandlers.ts (1)

52-59: 💤 Low value

Consider clarifying the timeout fallback behavior in the log message.

The log states "continuing with empty queue", but the timeout doesn't explicitly clear the queue—it simply continues with whatever queue state exists at that moment. If tracks were already present or partially restored, the queue might not actually be empty.

Consider rephrasing to: "Snapshot restore timed out, continuing with current queue state" for accuracy.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@packages/bot/src/handlers/player/lifecycleHandlers.ts` around lines 52 - 59,
The timeout fallback log is misleading; update the message emitted in the
restoreDeadline handler so it accurately reflects that the process continues
with the current queue state rather than an empty queue — locate the
Promise.race block using restoreDeadline and
musicSessionSnapshotService.restoreSnapshot and change the infoLog call
(currently using queue.guild.name) to something like "Snapshot restore timed
out, continuing with current queue state" or equivalent.
packages/bot/src/utils/music/ytdlpExtractor/service.ts (1)

141-161: 💤 Low value

Unreachable fallback at line 160.

The loop at lines 145-158 always returns on the last iteration (line 155: if (isLastAttempt) return result), making the fallback at line 160 unreachable. While this doesn't affect runtime behavior, it adds dead code.

Consider removing line 160 or adding a TypeScript assertion to document that the loop guarantees a return:

🧹 Remove unreachable fallback
             debugLog({ message: `yt-dlp attempt ${attempt + 1} failed (${result.error}), retrying…` })
         }
-
-        return { success: false, error: 'yt-dlp exhausted all retries' }
     }
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@packages/bot/src/utils/music/ytdlpExtractor/service.ts` around lines 141 -
161, The final fallback return is dead code because executeYtDlp always returns
inside the retry loop when isLastAttempt is true; remove the unreachable `return
{ success: false, error: 'yt-dlp exhausted all retries' }` at the end of
executeYtDlp (or, if you prefer to be explicit, replace it with a strong
assertion/throw like `throw new Error('unreachable')`) and keep the existing
loop logic that returns `result` when `isLastAttempt` is true.
packages/bot/tests/utils/music/ytdlpExtractor.test.ts (1)

54-127: 💤 Low value

Add a test to verify the exponential timeout progression (30s → 45s → 60s).

The test suite comprehensively validates the retry behavior for success on first attempt, retry-and-recover after initial failure, exhausting all three attempts, and timeout triggering process kill and retry. Consider adding a test to explicitly verify the timeout values scale according to the implementation progression to improve test coverage clarity.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@packages/bot/tests/utils/music/ytdlpExtractor.test.ts` around lines 54 - 127,
Add a unit test in ytdlpExtractor.test.ts that constructs a
YtDlpExtractorService with initial timeout 30000 and uses
mockSpawn/createMockProcess to return three processes; advance timers and flush
microtasks between attempts and assert that jest.advanceTimersByTime was called
(or timers progressed) with 30000, then 45000, then 60000 before each retry and
that proc.kill was invoked on each timed-out process; reference the existing
test patterns (retry behavior block, createMockProcess, mockSpawn,
extractor.handle, jest.advanceTimersByTime, and flushMicrotasks) to mirror the
other retry tests and verify the exponential timeout progression (30s → 45s →
60s).
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In `@Dockerfile`:
- Around line 116-120: Replace the inline HEALTHCHECK node one-liner with a call
to a new executable script (scripts/healthcheck-bot.js) that contains the same
Redis TCP PING logic (uses net.createConnection, writes the RESP PING frame,
listens for 'data' to check for "+PONG", handles 'error' by exiting non-zero,
and enforces a 3s timeout before exiting 1); add the script to the image via
COPY and ensure it is executable, then update the Dockerfile HEALTHCHECK line to
run node /app/scripts/healthcheck-bot.js (or the container path you COPY to) so
behavior and exit codes remain identical while making the logic maintainable and
testable locally.

In `@packages/bot/src/handlers/player/lifecycleHandlers.ts`:
- Around line 52-59: The timeout fallback log is misleading; update the message
emitted in the restoreDeadline handler so it accurately reflects that the
process continues with the current queue state rather than an empty queue —
locate the Promise.race block using restoreDeadline and
musicSessionSnapshotService.restoreSnapshot and change the infoLog call
(currently using queue.guild.name) to something like "Snapshot restore timed
out, continuing with current queue state" or equivalent.

In `@packages/bot/src/utils/music/ytdlpExtractor/service.ts`:
- Around line 141-161: The final fallback return is dead code because
executeYtDlp always returns inside the retry loop when isLastAttempt is true;
remove the unreachable `return { success: false, error: 'yt-dlp exhausted all
retries' }` at the end of executeYtDlp (or, if you prefer to be explicit,
replace it with a strong assertion/throw like `throw new Error('unreachable')`)
and keep the existing loop logic that returns `result` when `isLastAttempt` is
true.

In `@packages/bot/tests/utils/music/ytdlpExtractor.test.ts`:
- Around line 54-127: Add a unit test in ytdlpExtractor.test.ts that constructs
a YtDlpExtractorService with initial timeout 30000 and uses
mockSpawn/createMockProcess to return three processes; advance timers and flush
microtasks between attempts and assert that jest.advanceTimersByTime was called
(or timers progressed) with 30000, then 45000, then 60000 before each retry and
that proc.kill was invoked on each timed-out process; reference the existing
test patterns (retry behavior block, createMockProcess, mockSpawn,
extractor.handle, jest.advanceTimersByTime, and flushMicrotasks) to mirror the
other retry tests and verify the exponential timeout progression (30s → 45s →
60s).

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

Run ID: bb010c5b-b21c-4bab-b2a3-b69c3ed55a96

📥 Commits

Reviewing files that changed from the base of the PR and between d033162 and 1695153.

📒 Files selected for processing (14)
  • .dockerignore
  • Dockerfile
  • Dockerfile.frontend
  • Dockerfile.nginx
  • deploy/Dockerfile
  • docker-compose.yml
  • docs/decisions/2026-05-13-docker-overhaul.md
  • docs/decisions/2026-05-14-discord-integration-testing-strategy.md
  • nginx/frontend.conf
  • nginx/nginx.conf
  • packages/bot/src/handlers/player/lifecycleHandlers.ts
  • packages/bot/src/utils/music/ytdlpExtractor/service.ts
  • packages/bot/tests/utils/music/ytdlpExtractor.test.ts
  • packages/frontend/Dockerfile.dev

@greptile-apps greptile-apps Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LucasSantana-Dev has reached the 50-review limit for trial accounts. To continue receiving code reviews, upgrade your plan.

@sonarqubecloud

Copy link
Copy Markdown

@LucasSantana-Dev
LucasSantana-Dev merged commit a44c3ef into release/v2.11.0 May 14, 2026
17 checks passed
@LucasSantana-Dev
LucasSantana-Dev deleted the feat/reliability-improvements branch May 23, 2026 02:21

This branch was successfully deployed

1 active deployment
Preview — a4d8e325 Deployed May 14, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant