Features of HeFQUIN
HeFQUIN is a query federation engine for heterogeneous federations of graph data sources (e.g., federated knowledge graphs) that is designed to support querying across different data models and different data access interfaces.
This page provides an overview of the features of HeFQUIN, organized under the following headlines.
- Query Language Support
- Data Source Integration
- Authentication
- Response Caching
- Deployment Options
- Internals (Operators, Execution Models, and Optimization)
- Monitoring, Inspection, and Debugging
- Extension Points and Configurability
- Testing and Reliability
- Current Limitations
Query Language Support
HeFQUIN supports all features of SPARQL 1.1, with queries that contain SERVICE clauses to indicate which federation member should process which subpattern. Most features of the language are supported natively by the engine, including basic graph patterns, group graph patterns, UNION, OPTIONAL, MINUS, FILTER, BIND, and DISTINCT. The remaining features, such as subqueries, grouping, and aggregation, are implemented through integration of Apache Jena's query functionality into HeFQUIN.
In addition, HeFQUIN supports query constructs that are designed specifically for queries over (heterogeneous) federations, such as:
VALUESclauses to specify multiple federation members for the sameSERVICEclause in a more concise formPARAMSclauses for federation members that require request parametersUNFOLDclauses and dedicated functions for composite values (lists and maps, in particular)
For details, see the documentation of queries and query features supported by HeFQUIN.
Data Source Integration
HeFQUIN is designed for querying federations of heterogeneous types of data sources, including the following types of RDF-based data sources:
- SPARQL endpoints
- Triple Pattern Fragments (TPF)
- Bindings-Restricted Triple Pattern Fragments (brTPF)
For queries across RDF-based data sources that capture equivalent concepts and relationships using different vocabularies (ontologies), HeFQUIN can apply simple vocabulary mappings, which makes it possible to express such queries in terms of only one of the vocabularies.
HeFQUIN can also work with non-RDF data sources by giving them an RDF view via mappings described declaratively in RML, which is currently supported for:
- JSON-based Web APIs
To provide this functionality, HeFQUIN comes with its own RML processor, which can also be used as an independent component (i.e., separate from the query federation engine).
Ongoing development is adding support for further source types, including:
- Property Graphs accessible via openCypher
- GraphQL APIs
Authentication
HeFQUIN supports authentication for federation members, which can be configured per member and covers common Web authentication patterns such as:
- Basic authentication (username and password)
- Bearer token authentication
- Generic token-based authentication using custom HTTP Authorization schemes
Credentials are supplied through environment variables, which keeps them separate from the federation description.
Response Caching
HeFQUIN includes a two-level caching mechanism to cache responses retrieved for requests to federation members, which helps to minimize network overhead and to improve query performance. The main features of this caching mechanism are:
- Combination of a fast, in-memory L1 cache with a persistent L2 cache that retains responses across restarts of HeFQUIN
- Two backends for the L2 cache: one based on MapDB (suited for CLI use), the other based on Chronical Map (suited for the service deployment)
- TTL-based cache invalidation with the option to configure the time to live (TTL) per cache level
- Configurable cache capacity (per cache level), with LRU policy to handle capacity limits
- Options to bypass the cache for any given query, with explicit distinction between cardinality requests and data retrieval requests
Deployment Options
HeFQUIN can be deployed in several ways to fit different usage scenarios and infrastructure setups:
- It can be used as a Java library to interact with the engine directly from within a Java application.
- It can run as a service in a Docker container or via an existing Java servlet container.
- It can be used through command-line programs that execute queries directly via the engine or indirectly via a HeFQUIN service.
When deployed as a service, HeFQUIN exposes an HTTP API for interacting with the engine. This API supports:
- GET and POST requests for issuing queries
- Multiple result formats, including JSON, XML, CSV, TSV, and RDF serializations
- Retrieval of query processing information, including logical plans, physical plans, and execution statistics
The command-line programs support the same result formats and can provide the same query processing information.
Internals (Operators, Execution Models, and Optimization)
HeFQUIN implements a range of physical operators, including several join algorithms designed for federated query processing:
- Hash join
- Symmetric hash join (SHJ)
- Request-based nested-loops join (NLJ)
- Several variations of batching-based bind joins, including a FILTER-based variant, a VALUES-based variant, three UNION-based variants (including the bound join of FedX), and a version for brTPF, all implemented both using a sequential approach and a parallel approach to process the batches of input solutions
- Parallel multi-source bind join
It also implements two execution models:
- Push-based execution with multi-threaded intra-operator parallelism (default)
- Pull-based execution using non-blocking iterators
For query optimization, HeFQUIN uses both logical and physical optimization stages to make query execution efficient.
- Heuristics-based logical optimization uses plan rewriting rules such as filter push down, distinct push down, projection push down, filter redistribution, and selective union pull up
- Cardinality-based logical optimization determines a join order by leveraging cardinality requests
- Cost-based physical optimization can be set up using one of several strategies, including a greedy approach, the dynamic programming approach, simulated annealing, and randomized iterative improvement (notice that the physical optimizer is currently being redesigned)
Monitoring, Inspection, and Debugging
Several features provide visibility into the internal processes within the engine. Information that can be inspected includes:
- Logical and physical plans
- Query execution statistics, collected at the level of individual operators and data structures
- Exceptions that occurred during query execution
All of this information can be:
- printed to separate files (or to the terminal) when using the command-line programs
- accessed programmatically when using HeFQUIN as a Java library
- requested via the HTTP API of the HeFQUIN service
Extension Points and Configurability
HeFQUIN is developed using a clear separation between the various components, which makes it possible to swap out individual components and replace them by alternative implementations. To avoid recompiling the engine for each configuration, HeFQUIN uses an RDF-based configuration file that specifies which implementation and setup to use for which component. At start-up time, the engine reads this file and initializes all the components as specified in the file. This makes it possible to document and share the specific configurations used for experiments, but also to run the engine with external implementations of particular components (as long as the classes of these implementations are in the Java classpath).
Testing and Reliability
HeFQUIN is backed by a large automated test suite. The project currently includes more than 1000 unit tests, which helps ensure correctness and stability across features. Additionally, we have a testing environment set up in our lab to run regression tests based on federation benchmarks (LargeRDFBench and FedShop).
Current Limitations
HeFQUIN does not yet have a source selection component. All subpatterns of the queries given to HeFQUIN need to be wrapped in SERVICE clauses.