UltraAttn: Efficiently Parallelizing Attention through Hierarchical Context-Tiling
Published in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC), *CCF-A*, 2025
Long-context comprehension is critical for large language models. Context parallelism and irregular block-sparse attention are keyss to accelerating long-context training and inference. Existing context parallelism suffers from poor scalability due to the striped-like partition pattern, which causes high communication traffic, and the ring-based communication pattern, which limits kernel granularity, reduces device utilization, and incurs redundant communication. We present UltraAttn, a novel context parallelism solution for irregular attention. UltraAttn hierarchically tiles the context at the node and device levels to reduce communication cost, and applies kernel-level tiling to balance kernel overlap and device utilization. An ILP-based runtime further optimizes distributed attention latency. On 64 GPUs, UltraAttn achieves an average 5.5× speedup over state-of-the-art context parallelism methods across various irregular attention types.
