Skip to content

Add Hy3 IFP2 runtime, adaptive MoE, and SSD MTP cache - #30

Draft
ciru-ai wants to merge 4 commits into
charlie12345:mainfrom
ciru-ai:agent/hy3-ifp2-runtime-update
Draft

ciru-ai wants to merge 4 commits into
charlie12345:mainfrom
ciru-ai:agent/hy3-ifp2-runtime-update

Conversation

@ciru-ai

@ciru-ai ciru-ai commented Jul 13, 2026

Copy link
Copy Markdown
Contributor
  • feat(rocmfpx): add optimized 2.5 bpw ROCmFP2 kernels
  • perf(rocmfpx): add opt-in HY3 adaptive MoE launch
  • server: add SSD prompt cache for MTP
  • fix(rocmfpx): classify IFP2 in CPU clamp dispatch

@github-actions github-actions Bot added documentation Improvements or additions to documentation ggml CUDA testing server conversion labels Jul 13, 2026
charlie12345 added a commit that referenced this pull request Jul 29, 2026
vulkan: add ROCmFP2 Q8_1 decode kernels

Credits: ROCmFP2's core format/runtime and frozen codebook originated with
@ciru-ai in commit 9d1090e (PR #30) and landed through PR #32 with the
original authorship preserved.

Tested head: efcfb45
Validation: 34/35 applicable checks successful, zero failures.
Accepted exception: Windows CUDA 12.4 was still running and is not claimed as
passed.

Rollback by reverting this outer merge commit with `git revert -m 1`.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

conversion CUDA documentation Improvements or additions to documentation ggml server testing

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant