DSA or HISA have their selection baked into their attention kernel and, wiring a new kernel for every gpu arch is…
DSA or HISA have their selection baked into their attention kernel and, wiring a new kernel for every gpu arch is a cumbersome process, we avoid this and make selection out of attention.
We have a selection kernel which does bitonic topk m
How it ranks: Backlist reads my Twitter/X timeline, scores every tweet for substance with an LLM rubric (not engagement), and publishes the daily top picks with a one-line takeaway. Curated by Surya Dantuluri.