Safe Data Parallelism for General Streaming

Streaming applications process possibly infinite streams of data and often have both high throughput and low latency requirements. They are comprised of operator graphs that produce and consume data tuples. General streaming applications use stateful, selective, and user-defined operators. The strea...

Ausführliche Beschreibung

Gespeichert in:

Bibliographische Detailangaben
Veröffentlicht in:	IEEE transactions on computers Jg. 64; H. 2; S. 504 - 517
Hauptverfasser:	Schneider, Scott, Hirzel, Martin, Gedik, Bugra, Wu, Kun-Lung
Format:	Journal Article
Sprache:	Englisch
Veröffentlicht:	New York IEEE 01.02.2015 The Institute of Electrical and Electronics Engineers, Inc. (IEEE)
Schlagworte:	Aggregates Compilers Computer simulation Data mining Data processing distributed computing Graphs Operators Optimization Parallel processing parallel programming Partitioning Pipelines Run time (computers) Runtime Safety Semantics Streams
ISSN:	0018-9340, 1557-9956
Online-Zugang:	Volltext
Tags:	Tag hinzufügen Keine Tags, Fügen Sie den ersten Tag hinzu!

Beschreibung
Zusammenfassung:	Streaming applications process possibly infinite streams of data and often have both high throughput and low latency requirements. They are comprised of operator graphs that produce and consume data tuples. General streaming applications use stateful, selective, and user-defined operators. The stream programming model naturally exposes task and pipeline parallelism, enabling it to exploit parallel systems of all kinds, including large clusters. However, data parallelism must either be manually introduced by programmers, or extracted as an optimization by compilers. Previous data parallel optimizations did not apply to selective, stateful and user-defined operators. This article presents a compiler and runtime system that automatically extracts data parallelism for general stream processing. Data-parallelization is safe if the transformed program has the same semantics as the original sequential version. The compiler forms parallel regions while considering operator selectivity, state, partitioning, and graph dependencies. The distributed runtime system ensures that tuples always exit parallel regions in the same order they would without data parallelism, using the most efficient strategy as identified by the compiler. Our experiments using 100 cores across 14 machines show linear scalability for parallel regions that are computation-bound, and near linear scalability when tuples are shuffled across parallel regions.
Bibliographie:	ObjectType-Article-1 SourceType-Scholarly Journals-1 ObjectType-Feature-2 content type line 14 content type line 23
ISSN:	0018-9340 1557-9956
DOI:	10.1109/TC.2013.221