Loading...
SPARQL Query Federation over Polystores
Khan, Yasar
Khan, Yasar
Citations
Altmetric
item.page.identifiers
Date
2026-07-24
item.page.type
doctoral thesis
item.page.downloads
Citation
Abstract
Data retrieval systems are facing a paradigm shift due to the proliferation of specialised data storage engines (SQL, NoSQL, Column Stores, MapReduce, Data Stream, Graph) supported by varied data models (CSV, JSON, RDB, RDF, XML). One immediate consequence of this paradigm shift results into data bottleneck over the Web; which means, Web applications are unable to retrieve data with the intensity at which data is being generated from different facilities. Especially in the genomics and healthcare verticals, data is growing from petascale to exascale and biomedical stakeholders are expecting seamless retrieval of these data over the Web.
This thesis is focused on two aspects of SPARQL query federation over heterogeneous data sources, i.e. enabling and optimising SPARQL query federation over heterogeneous data sources and a comprehensive benchmark to evaluate such SPARQL query federation systems.
The typical linked data approach to query independent data silos is to convert all the underlying native data models into the RDF data model and devise querying mechanism through which these independent data silos can be queried in unison. While this approach can be practical for simpler verticals, in case of the HCLS domain it is already predicted that 2–40 exabytes of storage capacity will be needed by 2025 just for the human genomes which will continue to grow. Nevertheless, raw storage is not the main concern, but the amount of variant data (text, relational, stream, graph, etc.) being queried and analysed is already seen as a major hurdle in the meaningful use of this vast amount of data. The bottleneck over the Web can be reduced by minimising the costly data conversion process and delegating query performance and processing loads to the specialised data storage engines over their native data models. In this thesis, we present a Web-based query federation mechanism– called PolyWeb– that unifies query answering over multiple native data models (CSV, RDB, and RDF). We emphasise two main challenges of query federation over native data models: (i) devise a method to select prospective data sources– with different underlying data models– that can satisfy a given query; and (ii) query optimisation, join and execution over different data models. We also present SAFE, an efficient source selection approach for SPARQL query federation, which additionally provides policy aware access to sensitive information represented as RDF data cubes.
The Web of Data presents a significant challenge for data integration due to its distributed and diverse nature. Linked data has been proposed as a solution and consequently SPARQL query federation systems were developed, which in turn has led to the development of benchmarks to assess the performance of these systems. However, all these benchmarks support a single data model. With the emergence of various data models and systems designed for them, it is not practical to assume that all data will be materialized as RDF. Hence SPARQL query federations systems have been developed to support multiple data models. However, there is no such benchmark which can be used to assess the fitness of these systems. In this thesis, we also present a benchmark for evaluating the performance of SPARQL query federation approaches which can query data sources complying with different data models, such as RDB, CSV, JSON and RDF. The benchmark is composed of 12 real world datasets with billions of records represented using multiple data models and 17 queries of varying characteristics and complexity. The benchmark is also evaluated against similar state-of-the-art benchmarks based on the diversity of data sources, data models and queries.
We evaluated state-of-the-art SPARQL query federation systems based on different metrics such as, number of relevant sources selected, source selection time, query execution time and index generation times and compression ratio. Our results show that SAFE enabled granular graph-level access control over dis tributed clinical RDF data cubes and efficiently reduced the source selection and overall query execution time when compared with general-purpose SPARQL query federation engines. In case of PolyWeb, our results show that PolyWeb outperformed the other systems in complex queries execution times. In the benchmark, our results suggests that all of the systems could not handle complex queries and timed out where either the intermediate results or end result sets were large which means there is a need to improve these systems, especially the cases of complex queries with large result set size or large intermediate results size or involving many data sources.
item.page.funder
Research Ireland
Publisher
University of Galway
item.page.publisherdoi
License
CC-BY-NC-ND