Skip to content

missing VMI features for high-perf simd kernels#997

Draft
learning-chip wants to merge 9 commits into
hw-native-sys:mainfrom
learning-chip:zjw/vmi-dsl-feature-gaps-rebase
Draft

missing VMI features for high-perf simd kernels#997
learning-chip wants to merge 9 commits into
hw-native-sys:mainfrom
learning-chip:zjw/vmi-dsl-feature-gaps-rebase

Conversation

@learning-chip

@learning-chip learning-chip commented Jul 24, 2026

Copy link
Copy Markdown
Contributor

See standalone docs & reproducers in docs/feature-gaps/vmi.

This is a bug report & feature request, not a mergeable PR.

Suggested priority (for highest perf gain): 03, 05, 06; then 07, 01, 02

@mouliangyu @Zhendong404 @WenboCodes

…, UE8M0, Persistent

Co-authored-by: Cursor <cursoragent@cursor.com>
…hecks

Align gaps with remaining VMI vs ASC BW blockers; add small-L ui8 widen;
check in lowered VPTO for working workarounds.
Name the compiling-today path for what it is: correct but slow.
Cover MHC-style continuous bf16 strips (EVEN-only UNPK today;
dist_mode=unpack illegal). Note cast-back as a 05 consumer.
Add minimal reference_asc_cce.asc kernels that compile on the Ascend CCE
path while target_mi HIVM still crashes on bisheng.
Express dual-buffer ping-pong and pipe-event sync in target_mi for gap 05
(still fails HIVM/bisheng). Retarget gap 06 MI to PAT_VL8; that shape now
builds with bisheng while VMI L=8 legalization remains the gap.
Pair current_slow_vmi.py / desired_vmi.py with each gap IR so the VMI
authoring surface is documented alongside the .pto reproducers.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant