TraceFlow: Efficient Trace Analysis for Large-Scale Parallel Applications via Interaction Pattern-Aware Trace Distribution
Published in Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC), *CCF-A*, 2025
Trace analysis of large-scale parallel applications is crucial for understanding and optimizing performance. It primarily focuses on the interaction behaviors between different parallel processes, such as synchronization waits and asynchronous overlaps. The trace size explodes as the parallel scale of applications, thus current methods analyze traces in parallel to ensure analysis speed. However, due to the interaction pattern-agnostic trace distribution, they often introduce inter-process communications to fetch non-local event data during interaction analysis, leading to excessively long trace analysis time.
To address this issue, we propose TraceFlow, a trace analysis tool for large-scale parallel applications, which achieves a nearly communication-free analysis through an interaction pattern-aware trace distribution strategy. We first leverage static program structures and interaction patterns of applications to provide a global perspective. Guided by this perspective, we distribute events with interaction relationships to the same replay processes to avoid inter-process communications in the replay period. We evaluate the efficiency of TraceFlow on widely used benchmarks and several real-world applications with up to 8,192 processes. Experimental results show that TraceFlow achieves an average speedup of 13.49× in the analysis time compared to the state-of-the-art approaches.
