Fmha pr 2 (#26)
* support hdim=64/128 in same example code * support v transpose * revert gemm.cpp, not intent to modify it * remove useless code * fix a bug for swizzle C encoding, no perf change * optimize LDS encoding * update LDS layout * clean up code
Showing
Please register or sign in to comment