Distributed TensorFlow with MPI.pdf

时间:2023-01-29 08:46:24
【文件属性】:

文件名称:Distributed TensorFlow with MPI.pdf

文件大小:377KB

文件格式:PDF

更新时间:2023-01-29 08:46:24

TensorFlow AI 分布式

Machine Learning and Data Mining (MLDM) algorithms are becoming increasingly important in analyzing large volume of data generated by simulations, experiments and mobile devices. With increasing data volume, distributed memory systems (such as tightly connected supercomputers or cloud computing systems) are becoming important in designing in-memory and massively parallel MLDM algorithms. Yet, the majority of open source MLDM software is limited to sequential execution with a few supporting multi-core/manycore execution. In this paper, we extend recently proposed Google TensorFlow for execution on large scale clusters using Message Passing Interface (MPI). Our approach requires minimal changes to the TensorFlow runtime – making the proposed implementation generic and readily usable to increasingly large users of TensorFlow. We evaluate our implementation using an InfiniBand cluster and several well known datasets. Our evaluation indicates the efficiency of our proposed implementation.


网友评论