[Bugfix] Remove hardcoded head_size=256
for Deepseek v2 and v3 (#12…
#31
Job | Run time |
---|---|
9s | |
9s |
head_size=256
for Deepseek v2 and v3 (#12…
#31
Job | Run time |
---|---|
9s | |
9s |