RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models
事件
arXiv:2609.20971v1 Announce Type: new Abstract: Long-context large language model inference is increasingly limited by prefill, where dense self-attention processes the entire prompt before generation begins. Sparse block selection can reduce this cost, but a block centroid may hide a highly relevan
来源
本条目由采集管线自动抓取并发布,完整内容见下方来源链接。
*采集源:arXiv cs.AI*