A Distributed Path Query Engine for Temporal Property Graphs

Ramesh, Shriram; Baranawal, Animesh; Simmhan, Yogesh

doi:10.1109/CCGrid49817.2020.00-43

Full-text links:

Download:

Current browse context:

cs.DC

< prev | next >

new | recent | 2002

Computer Science > Distributed, Parallel, and Cluster Computing

Title: A Distributed Path Query Engine for Temporal Property Graphs

Authors: Shriram Ramesh, Animesh Baranawal, Yogesh Simmhan

(Submitted on 9 Feb 2020 (v1), last revised 14 Jun 2020 (this version, v2))

Abstract: Property graphs are a common form of linked data, with path queries used to traverse and explore them for enterprise transactions and mining. Temporal property graphs are a recent variant where time is a first-class entity to be queried over, and their properties and structure vary over time. These are seen in social, telecom, transit and epidemic networks. However, current graph databases and query engines have limited support for temporal relations among graph entities, no support for time-varying entities and/or do not scale on distributed resources. We address this gap by extending a linear path query model over property graphs to include intuitive temporal predicates and aggregation operators over temporal graphs. We design a distributed execution model for these temporal path queries using the interval-centric computing model, and develop a novel cost model to select an efficient execution plan from several. We perform detailed experiments of our Granite distributed query engine using both static and dynamic temporal property graphs as large as 52M vertices, 218M edges and 325M properties, and a 1600-query workload, derived from the LDBC benchmark. We often offer sub-second query latencies on a commodity cluster, which is 149x-1140x faster compared to industry-leading Neo4J shared-memory graph database and the JanusGraph / Spark distributed graph query engine. Granite also completes 100% of the queries for all graphs, compared to only 32-92% workload completion by the baseline systems. Further, our cost model selects a query plan that is within 10% of the optimal execution time in 90% of the cases. Despite the irregular nature of graph processing, we exhibit a weak-scaling efficiency >= 60% on 8 nodes and >= 40% on 16 nodes, for most query workloads.

Comments:	An extended version of the paper that appears in IEEE/ACM International Symposium on Cluster, Cloud and Internet Computing (CCGrid), 2020
Subjects:	Distributed, Parallel, and Cluster Computing (cs.DC); Databases (cs.DB)
Journal reference:	IEEE/ACM International Symposium on Cluster, Cloud and Internet Computing (CCGrid), 2020, 499-508
DOI:	10.1109/CCGrid49817.2020.00-43
Cite as:	arXiv:2002.03274 [cs.DC]
	(or arXiv:2002.03274v2 [cs.DC] for this version)

Submission history

From: Shriram Ramesh [view email]
[v1] Sun, 9 Feb 2020 03:41:25 GMT (421kb,D)
[v2] Sun, 14 Jun 2020 14:35:23 GMT (1270kb,D)

Which authors of this paper are endorsers? | Disable MathJax (What is MathJax?)

Link back to: arXiv, form interface, contact.

> cs > arXiv:2002.03274

Download:

Current browse context:

Change to browse by:

References & Citations

DBLP - CS Bibliography

Bookmark

Computer Science > Distributed, Parallel, and Cluster Computing

Title: A Distributed Path Query Engine for Temporal Property Graphs

Submission history