The allowed pattern may use local windows, global tokens, blocks, or learned selection. Reducing the number of query-key comparisons can make longer contexts practical, but the pattern may omit useful dependencies.
Sparse attention limits each query to attending to selected positions instead of every position in a sequence.
The allowed pattern may use local windows, global tokens, blocks, or learned selection. Reducing the number of query-key comparisons can make longer contexts practical, but the pattern may omit useful dependencies.