Skip to content

adding profiling context - #477

Merged
sunway513 merged 24 commits into
ROCm:mainfrom
mohbasit:enhance_profiling
Jul 21, 2026
Merged

sunway513 merged 24 commits into
ROCm:mainfrom
mohbasit:enhance_profiling

Conversation

@mohbasit

@mohbasit mohbasit commented Apr 2, 2026 •

Copy link
Copy Markdown
Contributor

Motivation

This PR aims to make the profile traces collected during ATOM run more detailed. It introduces 2 broad changes:

  1. enables annotations in the execution step with detailed information like KV lengths which could be used to calculate the roofline metrics
  2. Stores the capture phase traces as one trace file per batch size as well as omitting the warmup phase so that only meaningful information is retained and trace files are smaller in size

Technical Details

  1. The capture phase trace split uses the existing --mark-trace annotation.
  2. Setting the flag ATOM_ENABLE_DETAILED_ANNOTATION=1 gives you annotations in the trace file needed for calculating roofline metrics.

Test Plan

Test Result

Submission Checklist

@devalshahamd
devalshahamd requested a review from ChuanLi1101 May 28, 2026 16:10
@valarLip

Copy link
Copy Markdown
Collaborator

Review: overlap with ATOM's existing annotations + a few correctness items

Thanks for the PR. The capture-trace-per-bs split is a clear win, but the roofline annotation overlaps heavily with annotations ATOM already emits, and the way it's layered hits a known GPU-trace pitfall. Details below.

1. The roofline annotation duplicates the existing prefill[]/decode[] labels

run_model() already emits an unconditional record_function label on every forward (model_runner.py:1938-2000):

prefill[bs={bs} tok={total_tokens_num} ctx={context_lens}]
eager_decode[bs=...]
decode[bs={bs} tok={total_tokens_num} p={prefill_seqs} d={decode_seqs} spec={num_spec_step}]

Mapping the new execute_..._context_...(...)_generation_...(...) fields against it:

roofline field (new) already present in existing label
bs (actually total tokens) tok=total_tokens_num
context_{n} / generation_{n} p= (prefill seqs) / d= (decode seqs)
sq (sum N_Q) tok=
sk (sum N_KV) ctx=context_lens (existing is per-request; new is the sum)
sqsq (sum N_Q^2) not present
sqsk (sum N_Q*N_KV) not present

So the only genuinely new information is sqsq and sqsk (the quadratic terms needed for the roofline FLOP estimate). Everything else is a restatement or sub-set of what's already in the trace.

2. Layering it on forward() nests it above the existing label and likely loses the GPU annotation

The new annotation wraps forward(), while prefill[]/decode[] wraps run_model() (called by forward), so the trace ends up nested:

execute_..._context_...(roofline)   <- forward level (CPU annotation)
  \-- prefill[bs= tok= ctx=]         <- run_model level

A nested outer record_function does not produce a GPU-side annotation (only a CPU-timeline one) — we hit this before in this codebase. That means the roofline label probably won't carry GPU kernel time, which is exactly what a roofline analysis needs.

Suggested fix: drop the new forward()-level record_function and append sqsq/sqsk (and the sk sum if wanted) directly to the existing prefill[]/decode[] labels in run_model(). That layer is already a GPU-recognized annotation, it removes the duplication and nesting, and tools/parse_trace.py already keys off the prefill[/decode[ prefixes — so extending those fields is a smaller change than introducing a separate execute_... format.

Also note bs in the new annotation means "total tokens", which collides in name with the existing bs (effective batch size) while differing in value (they diverge significantly for prefill). Please rename to avoid confusion.

3. Smaller correctness / convention items

  • build_profile_annotation docstring is wrong (scheduler.py): it says "Return a context manager that annotates..." but the function returns None and mutates scheduled_batch.profile_annotation. The docstring also lists R_C/R_G fields that aren't in the emitted string.
  • New env var not registered: ATOM_ENABLE_ROOFLINE_ANNOTATION is read raw via os.environ.get(...) in two places. Per convention all ATOM_* vars belong in atom/utils/envs.py, and docs/environment_variables.md should be updated.
  • record_function without exception safety (model_runner.py forward): manual ctx.__enter__() / ctx.__exit__(None,None,None) leaks the annotation if run_model raises. Use a with block (this disappears entirely if you adopt the suggestion in section 2).
  • prof.step() alignment in start_capture_profiler is fragile: correctness relies on exactly 2 step() calls per bs against schedule(wait=1, warmup=0, active=1). Add assert self._profile_bs_idx < len(self.graph_bs) in on_trace_ready and a comment documenting the "2 steps per bs" invariant, otherwise a future edit silently mislabels traces or raises IndexError.
  • No tests added: at minimum a pure-Python unit test for build_profile_annotation (fake seqs, assert the prefill/decode aggregates) would cover the new logic without needing a GPU.

The capture-trace-per-bs change (replacing the old single-blob capture profiler in engine_core) looks good and doesn't overlap with anything existing.

@mohbasit

mohbasit commented Jul 4, 2026

Copy link
Copy Markdown
Contributor Author

1 & 2 — Duplication + the nested forward()-level annotation

Done as suggested. I dropped the separate forward()-level record_function entirely and now append only the genuinely-new quadratic terms to the existing prefill[]/decode[] labels in run_model() (the GPU-recognized layer). Labels now look like:

prefill[bs=1 tok=1024 ctx=1024 sqsq=1048576 sqsk=1048576 sk=1024]
decode[bs=64 tok=64 d=64 sqsq=64 sqsk=65600 sk=65600]
  • Only sqsq (Σ N_Q²), sqsk (Σ N_Q·N_KV), and the sk (Σ N_KV) sum are added; bs/tok/p=/d=/ctx= are left to the existing label — no duplication.
  • The separate execute_... string (and its conflicting bs field) is gone, so the naming collision is resolved.
  • tools/parse_trace.py keeps working since the prefill[/decode[ prefixes are unchanged.
  • Because it's now on the run_model() record_function instead of nested above it, the annotation carries GPU kernel time.

Design note: I aggregate at the batch level (summed over all reqs in that forward) rather than splitting context/generation, since each label corresponds to a single forward and that sum is the per-iteration roofline quantity. Happy to emit separate c_/g_ fields for mixed decode batches if you'd prefer.

3 — Correctness / convention items

  • Docstring: renamed build_profile_annotation → compute_roofline_aggregates; docstring now correctly says it mutates the batch in place (roofline_sqsq/roofline_sqsk/roofline_sk) and no longer claims to return a context manager or list R_C/R_G.
  • Env var registered: ATOM_ENABLE_ROOFLINE_ANNOTATION is now in atom/utils/envs.py and documented in docs/environment_variables.md. Both raw os.environ.get(...) reads are removed.
  • record_function exception safety: the manual __enter__()/__exit__() is gone entirely (dropped with the section-2 fix, as you predicted).
  • prof.step() invariant: added assert self._profile_bs_idx < len(self.graph_bs) in on_trace_ready with a comment documenting the "2 steps per bs" schedule.
  • Tests: added tests/test_scheduler.py::TestComputeRooflineAggregates — pure-Python, fake seqs asserting the prefill/decode aggregates plus no-op cases (env off / profiling inactive).
  • Extra bug fix: /start_profile (_handle_start_profile) never set scheduler.profile_active, so the aggregates would have silently never emitted at runtime — now set/cleared in the start/stop handlers.

Test results

Unit tests:

  • tests/test_scheduler.py — 52 passed (incl. 3 new roofline tests)
  • tests/test_envs.py — 31 passed

End-to-end (DeepSeek-R1-0528 FP8, TP8, ATOM_ENABLE_ROOFLINE_ANNOTATION=1, profiled run):

Successful requests:        64
Output token throughput:    1551 tok/s
Total token throughput:     3102 tok/s
Mean TPOT:                   26.6 ms

Roofline fields verified in the trace, e.g. prefill[... sqsq=1048576 sqsk=1048576 sk=1024] (1024² and 1024×1024 ✓) and decode[bs=64 ... sqsq=64 sqsk=65600 sk=65600] (64×1; KV sum 64×1025 ✓).

Capture-trace-per-bs left unchanged per your note.

@valarLip

valarLip commented Jul 5, 2026

Copy link
Copy Markdown
Collaborator

Review: adding profiling context

Adds opt-in profiling detail: per-batch-size CUDA-graph capture traces (capture_traces/bs_<bs>_rank<rank>.json.gz, replacing the old single mark-trace capture block in engine_core.py), and roofline FLOP aggregates (sqsq/sqsk/sk) appended to prefill[]/decode[] trace labels behind ATOM_ENABLE_ROOFLINE_ANNOTATION. All new behavior is gated behind profile_active + the env flag, so the production inference path is untouched. The core mechanics are sound — the zip(seqs.values(), num_scheduled_tokens) ordering matches the existing ScheduledBatch.__init__ invariant, the capture-profiler schedule(wait=1,active=1,repeat=0) fires exactly once per bs, and ScheduledBatch initializes the roofline fields so the label path never AttributeErrors.

Findings (most severe first)

1. scheduler.py compute_roofline_aggregates — int32 overflow silently corrupts the values for long prefills.
scheduled_batch.num_scheduled_tokens is stored as np.int32 (ScheduledBatch.__init__, np.asarray(..., dtype=np.int32)). In the prefill branch nq = num_tokens is an np.int32, so nq*nq (and nq*nkv) overflow int32 once a prefill/chunk reaches ~46341 tokens — e.g. np.int32(65536)**2 == 0. A long-context prefill then emits sqsq=0/garbage in the trace, defeating the exact estimate the feature exists to produce. Fix: nq = int(num_tokens) (or cast to int64) before multiplying.

2. engine_core.py (~439) — EngineCore.start_profiler / stop_profiler / print_mtp_statistics are dead code duplicating the live utility handlers.
Runtime profiler control flows through EngineUtilityHandler._UTILITY_HANDLERS → _handle_start_profile/_handle_stop_profile/_handle_get_mtp_stats (a different class; runner_mgr.call_func("start_profiler") targets ModelRunner, not EngineCore). These three new methods have zero callers and re-implement the same profile_active flip / spec_stats._log() logic, creating a second "authoritative" writer for profile_active that can drift if later wired up. They also entrench the pre-existing profile_enbaled typo. Recommend deleting all three — the _handle_* edits in this PR are the correct, complete path.

3. Naming: "roofline" is misleading, and sqsq/sqsk/sk are cryptic (violates the repo's Name-matches-function rule).
What is actually computed are attention quadratic-work terms, not a roofline:

  • A roofline needs FLOPs and bytes moved (for arithmetic intensity) and peak compute/bandwidth. This PR provides none of the last two, so it can't produce a roofline — it's the raw material for an attention FLOP estimate only.
  • It captures only attention's quadratic terms (QKᵀ / softmax·V). The linear GEMM FLOPs (MLP, projections) that dominate most shapes are absent, so it isn't even a full FLOP count.
  • The abbreviations are inconsistent: sk reads as "sum of N_KV", but by the same scheme sqsq looks like "sum-Q sum-Q" while it's actually Σ N_Q². A reader can't decode sqsq=... in a trace label without the docstring.

Suggest dropping "roofline" and naming things for what they are, e.g. ATOM_ANNOTATE_ATTN_FLOPS; fields attn_sum_nq2 (Σ N_Q²), attn_sum_nq_nkv (Σ N_Q·N_KV), attn_sum_nkv (Σ N_KV); label keys nq2= / nqnkv= / nkv=. If "roofline" is intentional (a downstream script builds roofline points from these), the docstring/docs should state that these are attention-quadratic terms only and the caller must add GEMM FLOPs + bytes.

4. model_runner.py on_trace_ready — assert on the bs-index invariant can abort CUDA-graph capture (server startup).
Capture runs at bring-up; an AssertionError inside on_trace_ready kills startup rather than degrading trace output, and the assert is stripped under python -O so the invariant it documents goes unchecked. A profiling-only diagnostic shouldn't crash capture — prefer log-and-skip (if idx >= len(graph_bs): logger.warning(...); return).

5. scheduler.py compute_roofline_aggregates — decode hardcodes nq=1, undercounting MTP/spec-decode.
For MTP a decode step schedules mtp_k+1 query tokens, but the DECODE branch ignores the zipped num_tokens and uses nq=1, so the sums are undercounted up to ~(mtp_k+1)× on spec-decode batches. Documented as intentional in the docstring, but for a FLOP-estimate feature this silently misleads MTP analysis — either use the scheduled token count or call out the limitation in the label.

6. docs/ — docs-match-code gap for the new capture output.
Only ATOM_ENABLE_ROOFLINE_ANNOTATION got a doc entry. The new per-bs capture_traces/bs_<bs>_rank<rank>.json.gz artifact and its --mark-trace gating (and the removal of the old single capture-graph trace) aren't documented in serving_benchmarking_guide.md, so users won't know where capture traces land or how to enable them.

7. scheduler.py — envs.ATOM_ENABLE_ROOFLINE_ANNOTATION is re-read via os.getenv every engine step while profiling is active (minor).
envs.__getattr__ re-invokes the lambda (no caching), so during a profiling run every _process_engine_step_inner pays an os.getenv. Cheap and profiling-only, but trivially avoidable by reading it once when profiling starts.

#1 is the only correctness bug (confined to the profiling path); #2 and #3 are the clearest cleanup/naming items; #4–7 are robustness/accuracy/docs.

@poznano-amd

Copy link
Copy Markdown
Contributor

Just chiming in here -- this capture+replay method is important for the work in my org.

@ChuanLi1101 ChuanLi1101 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the updates — the capture-trace-per-bs split looks good. Two things before I can approve:

1. Correctness bug (blocking): int32 overflow in compute_roofline_aggregates.
num_scheduled_tokens is stored as np.int32, so in the prefill branch nq = num_tokens is an np.int32 and nq*nq / nq*nkv overflow once a prefill/chunk exceeds ~46341 tokens (e.g. np.int32(65536)**2 == 0). Long-context prefills would emit sqsq=0/garbage. Please cast nq = int(num_tokens) (or int64). The new unit test uses plain Python ints, so it won't catch this — please add a case with a large N_Q.

2. Merge conflicts: the branch is currently conflicting with main; please rebase/resolve.

The remaining items from my Jul 5 review (naming "roofline"/sqsq/sqsk/sk, the assert in on_trace_ready that can abort capture at startup, the dead EngineCore start/stop_profiler, nq=1 undercounting MTP, and the capture_traces doc gap) are non-blocking — a quick reply on each would help. Once #1 and the conflicts are handled I'll re-review.

@mohbasit

Copy link
Copy Markdown
Contributor Author

Thanks @ChuanLi1101 — addressed the blocking items and went through the rest of the Jul 5 list too.

1. int32 overflow (blocking) — fixed.
compute_roofline_aggregates now casts to Python int before multiplying, so nq*nq/nq*nkv can't wrap:

nq = int(num_tokens)
nkv = int(seq.num_tokens) if decode else int(seq.num_cached_tokens) + nq

Added a regression test that feeds num_scheduled_tokens as an np.int32 array with N_Q=65536 and asserts sqsq == 65536**2 (4294967296, not 0): tests/test_scheduler.py::TestComputeRooflineAggregates::test_no_int32_overflow_large_prefill.

2. Merge conflicts — resolving. Rebasing onto latest main and pushing shortly.


Non-blocking items from the Jul 5 review:

Naming (roofline / sqsq/sqsk/sk). Renamed the env flag ATOM_ENABLE_ROOFLINE_ANNOTATION → ATOM_ENABLE_DETAILED_ANNOTATION to drop the "roofline" framing, and the docs now explicitly state these are attention-quadratic terms only (a full roofline still needs GEMM FLOPs + bytes). I kept the sqsq/sqsk/sk label keys as-is intentionally, since tools/parse_trace.py / analyze_trace_summary.py and existing downstream trace parsing key off the current label format .

assert in on_trace_ready. Changed to log-and-skip so a cadence mismatch can never abort CUDA-graph capture at startup (and isn't stripped under python -O):

if self._profile_bs_idx >= len(self.graph_bs):
    logger.warning(...); return

Dead EngineCore.start_profiler/stop_profiler/print_mtp_statistics. Deleted all three — the live path is ModelRunner via the EngineUtilityHandler._handle_* methods. Also removed the now-unused self.profile_enbaled attribute (kills the typo).

nq=1 undercounting MTP. Fixed — the decode branch now uses the scheduled query-token count (mtp_k+1 for spec-decode) instead of a hardcoded 1. Non-spec decode is unchanged (scheduled count is 1). Added test_decode_counts_scheduled_query_tokens.

capture_traces doc gap. Added section 5.6 CUDA-Graph Capture Traces to serving_benchmarking_guide.md documenting the capture_traces/bs_<bs>_rank<rank>.json.gz output, its --mark-trace gating, and the ATOM_ENABLE_DETAILED_ANNOTATION fields.

Minor: per-step os.getenv. The env flag is now read once and cached on the scheduler (_detailed_annotation_enabled), so profiling no longer re-reads it every engine step.

Unit tests updated accordingly.

ChuanLi1101
ChuanLi1101 previously approved these changes Jul 14, 2026

@ChuanLi1101 ChuanLi1101 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the thorough follow-up — all my blocking items and the Jul 5 findings are addressed. Verified on ec41d3e:

  • int32 overflow (blocking) fixed: nq = int(num_tokens) and nkv = int(...) in compute_roofline_aggregates, plus a test_no_int32_overflow_large_prefill regression test (N_Q=65536 → sqsq == nq*nq, not 0).
  • Merge conflicts resolved — branch is now MERGEABLE.
  • Dead EngineCore start/stop_profiler/print_mtp_statistics removed; profiler control stays on the _handle_* utility path.
  • on_trace_ready now log-and-skips instead of assert, so a cadence mismatch can't abort capture at startup.
  • Decode branch now uses the scheduled query-token count (no longer hardcoded nq=1), so MTP/spec-decode isn't undercounted (+ test).
  • Env flag cached once as self._detailed_annotation_enabled — no per-iteration os.getenv.
  • Naming/docs: renamed to ATOM_ENABLE_DETAILED_ANNOTATION; docs now state these are attention-quadratic terms only (a full roofline still needs GEMM FLOPs + bytes). capture_traces/bs_<bs>_rank<rank>.json.gz documented in the serving guide.

All new behavior stays gated behind profile_active + the env flag, so the production path is untouched. LGTM.

@ChuanLi1101

Copy link
Copy Markdown
Collaborator

Approving on the condition that CI passes and the benchmark shows no perf degradation before merge. The new logic is gated behind profile_active + ATOM_ENABLE_DETAILED_ANNOTATION, so the production path should be unaffected — but please run the full CI + perf benchmark and confirm no regression, then it's good to merge.

@ChuanLi1101
ChuanLi1101 self-requested a review July 14, 2026 17:19
@ChuanLi1101

Copy link
Copy Markdown
Collaborator

Re-ran the failed jobs — still red, but all failures are infra, none related to this PR:

  • aiter wheel artifact digest-mismatch on download → load model runner failed (Kimi-K2.7, DeepSeek-V4-Pro)
  • Model download timeout — 97GB models hitting the 30m MODEL_DOWNLOAD_TIMEOUT, plus an empty HF Bearer token (DeepSeek-R1-0528-FP4 MTP)
  • Stale GLM-5.2-FP8 server holding port 8000 → model-mismatch 400s (DeepSeek-R1 FP8/FP4, Kimi-K2.5)
  • HIP/NCCL init errors (hipErrorIllegalState) on a bad GPU/runner

Re-running just lands on the same broken environment. This needs infra to fix the aiter artifact digest, bump/pre-cache the model download, and clean the runners. @LingPeng — could you help? I'll re-approve once CI is green.

@zufayu
zufayu requested a review from valarLip July 16, 2026 06:10
ChuanLi1101
ChuanLi1101 previously approved these changes Jul 16, 2026

@ChuanLi1101 ChuanLi1101 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re-approving. All my blocking items and the Jul 5 findings have been addressed and verified (int32 overflow fixed + regression test, merge conflicts resolved, dead EngineCore start/stop_profiler removed, on_trace_ready log-and-skips instead of assert, decode branch uses the scheduled query-token count for MTP, env flag cached, renamed to ATOM_ENABLE_DETAILED_ANNOTATION + docs updated, capture_traces artifact documented). All new behavior stays gated behind profile_active + ATOM_ENABLE_DETAILED_ANNOTATION, so the production inference path is untouched.

Note on CI — the lightweight gates (black, ruff, non-GPU unit tests, docs build) are green. The remaining red jobs are all infrastructure, not code:

  • Required check Accuracy (DeepSeek-R1-0528-FP4 MTP): empty HF Bearer token → Illegal header value b'Bearer ', so the 97GB model can't resolve/download and hits MODEL_DOWNLOAD_TIMEOUT=30m → aborts. Not PR-related.
  • Non-required jobs (DeepSeek-V4-Pro TP8 SGLang, DeepSeek-V4-Flash, GLM-5.2, MiniMax-M3): GPU memory unbalanced / GPUs occupied by stale processes, plus the same empty-HF-token issue. The same DeepSeek-V4-Pro model passes in the ATOM Test workflow on a clean runner.

Once infra restores the HF token / pre-caches the models and the required Accuracy (DeepSeek-R1-0528-FP4 MTP) check re-runs green, this is good to merge.

@valarLip

Copy link
Copy Markdown
Collaborator

Following further review of the annotation portion of this PR, I would like to revise my earlier assessment. My previous comment characterized the emitted fields as "attention-quadratic terms only." That understated the problem: the computation in compute_roofline_aggregates assumes dense full attention (N_KV = full sequence length), and this assumption does not hold for the majority of the models in our benchmark roster. For those models the emitted values are not merely incomplete — they are incorrect.

Affected models:

  • DeepSeek-V4 / Kimi-K2.5: Compressed Sparse Attention selects index_topk (default 1024) KV entries per query (sparse_attn_v4_paged_*, top_k_per_row_*). Each query attends to a bounded budget rather than the full sequence, so sqsk/sk overcount by roughly two orders of magnitude at long context.
  • DeepSeek-V3.2: uses the same sparse indexer path in deepseek_v2.py (index_topk).
  • GLM-5: routed through the fused indexer (ATOM_ENABLE_GLM_FUSED_INDEXER).
  • gpt-oss: alternating sliding-window attention with sinks (gpt_oss.py:124); the window layers are bounded, not N_KV.
  • Qwen3-Next / Qwen3.5: Gated DeltaNet linear attention (mamba_type = "gdn_attention"). These layers have no N_Q·N_KV score matrix, so sqsq/sqsk are undefined for them rather than overcounted.

In addition, the aggregate is a single batch-level scalar with no per-layer distinction. It therefore cannot represent the mixed layer types within a single model (DSv4 interleaves dense-MLA and CSA layers; Qwen3-Next interleaves GDN and full-attention layers). The fields are only meaningful for genuinely dense-attention models (R1 / V3 / V3-0324), which are a minority of what we serve.

Separately, on naming. Only the environment variable was renamed (ATOM_ENABLE_ROOFLINE_ANNOTATION → ATOM_ENABLE_DETAILED_ANNOTATION). The term "roofline" remains throughout the code it controls: Scheduler.compute_roofline_aggregates, ScheduledBatch.roofline_sqsq/sqsk/sk, ModelRunner._roofline_label_suffix, build_run_label(roofline_suffix=...), and the associated docstrings. As a result the flag name no longer matches the symbols it drives, and "roofline" remains inaccurate — the feature provides no bytes-moved, peak compute, or GEMM FLOPs. This is inconsistent with the repository's name-matches-function guideline.

Regarding the label keys: I reviewed tools/parse_trace.py, which matches only the prefill[ / decode[ prefix and extracts bs= (lines 75, 85). No in-tree consumer reads sqsq/sqsk/sk, so the backward-compatibility rationale for retaining the abbreviated keys does not apply; renaming them to nq2= / nqnkv= / nkv= would not break any existing tooling.

Recommendation: I suggest splitting this PR. The per-batch-size capture-trace change is self-contained and useful, and can be merged on its own. For the annotation portion, please either (a) compute the terms per layer against the actual index_topk / sliding-window budget and gate by attention type so the values are correct on our sparse and linear-attention models, or (b) defer it and explicitly scope it to dense-attention models in both the label and the documentation. In either case, please complete the rename away from "roofline" (symbols and docstrings, not only the environment variable) and adopt readable label keys.

@mohbasit

Copy link
Copy Markdown
Contributor Author

Implemented the naming changes from roofline->detailed

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

7 participants