Skip to content

fix(OMN-13469): durabilize dev redpanda partition cap via .bootstrap.yaml - #2064

Merged
jonahgabriel merged 1 commit into
devfrom
jonah/omn-13469-durabilize-partition-cap
Jun 22, 2026
Merged

jonahgabriel merged 1 commit into
devfrom
jonah/omn-13469-durabilize-partition-cap

Conversation

@jonahgabriel

@jonahgabriel jonahgabriel commented Jun 22, 2026 •

Copy link
Copy Markdown
Collaborator

fix(OMN-13469): durabilize dev redpanda partition cap via .bootstrap.yaml

Root Cause (verified live)

Fresh Redpanda volume boots with DEFAULT cluster cap topic_partitions_per_shard=1000 (single shard). With ~1392 contract topics, the broker jams at ~995 with BROKER_NOT_AVAILABLE, leaving the runtime stuck runtime_pending.

Why existing mechanisms fail on fresh volume / partial recreate:

  • --set topic_partitions_per_shard=7000 passed to redpanda start is a node-startup hint, NOT a persistent cluster config write — it does not survive fresh volume formation.
  • redpanda-partition-cap init service (restart:"no") runs once on initial stack creation but does NOT re-run after a volume reset or partial recreate.

Fix

Added docker/redpanda/.bootstrap.yaml:

topic_partitions_per_shard: 7000
topic_memory_per_partition: 1048576

Mounted read-only into the redpanda container at /etc/redpanda/.bootstrap.yaml. Redpanda reads this file on first cluster formation (fresh volume), applying the config before any topics are created — making the partition cap volume-reset-proof.

Belt-and-suspenders (all three layers agree on 7000)

Layer Mechanism Applies when
1 (primary, reset-proof) .bootstrap.yaml ← this PR Fresh volume / first cluster formation
2 (warm restart) --set topic_partitions_per_shard=7000 in compose command Node startup (existing)
3 (post-boot) redpanda-partition-cap init service via rpk cluster config set After redpanda healthy (existing)

Belts #2 and #3 are NOT removed — they are harmless and help on warm restarts. Comment added referencing OMN-13469 to explain the layering.

Files Changed

  • docker/redpanda/.bootstrap.yaml (new) — cluster bootstrap config, 7000 partitions/shard
  • docker/docker-compose.infra.yml — added bootstrap.yaml mount + clarifying comment
  • docker/catalog/services/redpanda.yaml — added bootstrap.yaml mount to catalog manifest

Scope

Dev compose only (docker-compose.infra.yml + docker/catalog/services/redpanda.yaml). Prod/stability/judge lanes have independent rpk-based cap management and are NOT touched by this PR. Follow-up: extend to prod/stability for full parity (tracked under OMN-13469).

Validation

  • docker-compose -f docker/docker-compose.infra.yml config -q passes (syntax valid; env-var interpolation warning is pre-existing)
  • pre-commit run --files docker/docker-compose.infra.yml docker/catalog/services/redpanda.yaml docker/redpanda/.bootstrap.yaml — all hooks pass

Evidence-Source: 815310265ba79c283c718c1524e1ebab6315f04e
Evidence-Ticket: OMN-13469

…yaml

Root cause: fresh Redpanda volume boots with DEFAULT cluster cap
topic_partitions_per_shard=1000 (single shard). With ~1392 contract
topics, the broker jams at ~995 with BROKER_NOT_AVAILABLE, leaving the
runtime stuck runtime_pending. The --set flag passed to `redpanda start`
is a node-startup flag that does not persist cluster config on a fresh
volume; the redpanda-partition-cap init service (restart:"no") does not
re-run after a volume reset.

Fix: add docker/redpanda/.bootstrap.yaml containing
  topic_partitions_per_shard: 7000
  topic_memory_per_partition: 1048576

Redpanda reads .bootstrap.yaml on FIRST cluster formation (fresh volume),
applying the config before any topics are created — making the cap
volume-reset-proof regardless of init-service execution order.

Three-layer belt-and-suspenders (all must agree on 7000):
  1. .bootstrap.yaml (reset-proof, primary)       ← this PR
  2. --set topic_partitions_per_shard=7000 (node startup flag, warm)
  3. redpanda-partition-cap init service (rpk cluster config set, post-boot)

Scope: dev compose (docker-compose.infra.yml + catalog/services/redpanda.yaml)
only. Prod/stability/judge lanes have independent rpk-based cap management
and are NOT touched (follow-up: OMN-13469 prod/stability parity).

Evidence-Ticket: OMN-13469
@jonahgabriel
jonahgabriel enabled auto-merge June 22, 2026 10:03
@coderabbitai

coderabbitai Bot commented Jun 22, 2026 •

Copy link
Copy Markdown

Review Change Stack

📝 Walkthrough

Walkthrough

Adds docker/redpanda/.bootstrap.yaml with topic_partitions_per_shard: 7000 and topic_memory_per_partition: 1048576, then mounts this file read-only into the Redpanda container at /etc/redpanda/.bootstrap.yaml in both docker/docker-compose.infra.yml and docker/catalog/services/redpanda.yaml.

Changes

Redpanda Partition Cap Bootstrap

Layer / File(s) Summary
Redpanda bootstrap config file
docker/redpanda/.bootstrap.yaml
New bootstrap YAML sets topic_partitions_per_shard to 7000 and topic_memory_per_partition to 1048576, and documents the three-layer reset-proof configuration strategy (bootstrap file, warm-restart --set flag, post-start rpk init).
Bootstrap file volume mount wiring
docker/docker-compose.infra.yml, docker/catalog/services/redpanda.yaml
Both compose definitions add a read-only bind-mount of ./redpanda/.bootstrap.yaml into the container at /etc/redpanda/.bootstrap.yaml:ro. The infra compose file also adds inline comments explaining that the --set startup flag may not persist on fresh volumes and that the bootstrap mount plus the redpanda-partition-cap init service are the agreed reset-proof mechanisms.

Estimated code review effort

🎯 1 (Trivial) | ⏱️ ~3 minutes

Poem

A bootstrap YAML, small but bold,
Sets partition caps that always hold.
Fresh volumes come and volumes go,
But /etc/redpanda keeps the flow.
Read-only mounted, reset-proof too —
🐇 One config file, the clusters knew! ✨

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed The title directly and accurately summarizes the main change: introducing a .bootstrap.yaml file to durabilize (make persistent) the Redpanda partition cap configuration in the dev environment, addressing the specific issue OMN-13469.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.

✏️ Tip: You can configure your own custom pre-merge checks in the settings.

✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch jonah/omn-13469-durabilize-partition-cap

Comment @coderabbitai help to get the list of available commands and usage tips.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
docker/docker-compose.infra.yml (1)

376-376: 🧹 Nitpick | 🔵 Trivial | 💤 Low value

Consider adding topic_memory_per_partition to the --set flag for consistency.

The .bootstrap.yaml (Layer 1) sets both topic_partitions_per_shard: 7000 and topic_memory_per_partition: 1048576, and the redpanda-partition-cap service (Layer 3, lines 423-424) also sets both values via rpk cluster config set. However, Layer 2 (the --set flag) only sets topic_partitions_per_shard=7000.

The PR objectives state: "All three layers must agree. Edit all three if the target value changes."

🔧 Add topic_memory_per_partition to --set flag
       # 7000. (OMN-13469)
       - --set topic_partitions_per_shard=7000
+      - --set topic_memory_per_partition=1048576
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@docker/docker-compose.infra.yml` at line 376, The --set flag in the
docker-compose.infra.yml file is missing the topic_memory_per_partition
parameter that is present in both the .bootstrap.yaml and redpanda-partition-cap
service configuration layers. Add topic_memory_per_partition=1048576 to the
--set flag alongside the existing topic_partitions_per_shard=7000 setting to
ensure consistency across all three configuration layers as specified in the PR
objectives.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In `@docker/docker-compose.infra.yml`:
- Line 376: The --set flag in the docker-compose.infra.yml file is missing the
topic_memory_per_partition parameter that is present in both the .bootstrap.yaml
and redpanda-partition-cap service configuration layers. Add
topic_memory_per_partition=1048576 to the --set flag alongside the existing
topic_partitions_per_shard=7000 setting to ensure consistency across all three
configuration layers as specified in the PR objectives.

ℹ️ Review info
⚙️ Run configuration

Configuration used: defaults

Review profile: CHILL

Plan: Pro

Run ID: da93c76d-f3e4-40b7-9a27-53bbc6e23464

📥 Commits

Reviewing files that changed from the base of the PR and between ec2d922 and 1f1e7d5.

📒 Files selected for processing (3)
  • docker/catalog/services/redpanda.yaml
  • docker/docker-compose.infra.yml
  • docker/redpanda/.bootstrap.yaml

@jonahgabriel
jonahgabriel added this pull request to the merge queue Jun 22, 2026
Merged via the queue into dev with commit 38e3d3b Jun 22, 2026
87 of 92 checks passed
@jonahgabriel
jonahgabriel deleted the jonah/omn-13469-durabilize-partition-cap branch June 22, 2026 10:29
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant