iAccount based inNorway
About this account
- Account based in
- Norway
- Connected via
- Web
Account-level information from X, not a live location or the device used for a specific post.
New Databases papers from https://nitter.cf/t.co/nAsAct7zgt: database management, datamining, and data processing. Thank you to arXiv for use of its open access interoperability.
- Tweets12.1K
- Following1
- Followers258
- Likes2
ALT The semiring framework and its extensions form the basis of a rich collection of theoretical results and implementations for provenance tracking of database queries. Many real-world queries use aggregation and conditions on the aggregate values. Support for such queries has been proposed by introducing semimodule elements as aggregate values and formal comparisons between aggregate values as tuple annotations, which takes the approach outside the standard semiring framework. In this work, we show how to introduce a semantics for the provenance of such queries in arbitrary commutative semirings with monus (or m-semirings), without the need for additional operators. This semantics is shown to agree with the standard provenance of the aggregation-free self-join rewriting of HAVING COUNT(*) queries in semirings that are absorptive and where times distributes over monus. We derive algorithms for this semantics and implement them within the ProvSQL system, with viable performance on a real-w
ALT Query containment is a fundamental decision problem in database theory: given two queries, determine whether, over all database instances, every answer produced by the first is also produced by the second. For conjunctive queries under set semantics, the problem is understood through the classical homomorphism-based characterisation. Under bag semantics, the interpretation underlying real relational databases, containment becomes a quantitative comparison of answer multiplicities. Despite decades of work, the decidability of bag containment for conjunctive queries remains open. This frontier is fragile: for slightly more expressive classes, bag containment is undecidable, with negative results relying on reductions from variants of Hilbert's 10th problem. This work develops a unified framework for bag containment of conjunctive queries that subsumes two previously studied decidable cases: projection-free and join-on-free containee queries. The framework yields decidability for a broade
ALT Navigable graphs for nearest-neighbor search are built either incrementally, each inserted point searching a graph that mutates as construction proceeds, or in batch over a fixed substrate, which buys determinism and parallelism at a price in build time. We measure that price and find where it comes from. Instrumenting a tuned Vamana and PiPNN and building every system repeatedly in a paired design on one 64-thread machine, we separate build time into work (distance evaluations) and cost per evaluation. Letting the batch builder's substrate mutate in B synchronous blocks recovers the feedback loop of incremental construction while the graph stays a function of (data, seed, parameters, B): eight trees with batched feedback match thirty-two frozen trees, and the build does 0.88× Vamana's distance work. Yet it takes 1.55× Vamana's wall-clock, because the mutating substrate costs more per evaluation. Pushing further, we find a wall that no search-based builder crosses. Every one of them, i
ALT We design, implement, and evaluate KathDB-FAO, a new query evaluation subsystem for our KathDB multimodal DBMS. KathDB-FAO takes as input a query in natural language (NL) and converts it into a query execution plan where each operator is a function whose body is synthesized during query evaluation, which allows powerful query-specific optimizations. To generate accurate and efficient plans from NL, KathDB-FAO first extracts fine-grained atomic actions for correctness, then establishes contracts on the inputs and outputs of those actions and groups them for efficiency, and finally synthesizes the function for each group on the fly. On SemBench, KathDB-FAO cuts execution cost by 58.8% on average across scenarios compared with the next best system, at comparable or better quality.
ALT Knowledge Graphs (KGs) play an increasingly important role in numerous applications ranging from traditional knowledge representation to serving as memory for LLMs to support downstream tasks. However, their construction is labor-intensive; thus, in recent years, numerous approaches for automizing this process have been proposed. Compared to the popularity of construction approaches that focus on textual input data, methods for semi-structured inputs remain underrepresented and as a result, no comprehensive benchmark and evaluation suite exists to judge the quality of mapping predictions and generated KGs. This is problematic, as a KG's quality has direct influence on the downstream applications it supports and thus, strong evaluation mechanisms for their construction are urgently needed. In this work, we thus focus on the evaluation of KG construction from semi-structured data and present a benchmark and evaluation pipeline for KG construction that combines the quality dimensions (1)
ALT We address the problems of giving a semantics to a relational database (RDB) that has missing values (MVs). The causes for the latter are governed by a Missingness Mechanism that is modelled as a Bayesian Network (BN) that involves the DB attributes as variables. The BN is called a Missingness Graph (MG). Our approach considerable departs from the treatment of RDBs with NULL (values). The combination of the MG and the observed DB allows us to build a block-independent probabilistic DB. We identify two optimal classes of its possible worlds on which QA can be performed. Those classes jointly capture probabilistic uncertainty and statistical plausibility of the implicit imputation of MVs. We obtain tractability results for the computation of some optimal classes; and we also obtain complexity results that characterize the computational feasibility of our approach.
ALT The rapid adoption of AI across industries for (e.g.) fine-tuning and analytics has accelerated the need for high-quality data. To satisfy this demand, public and private entities will buy and sell data with other organizations. But such data sharing can and does violate privacy norms and laws. EU and American lawmakers have endeavored to control data sharing, restricting what data may be shared with whom, and under what circumstances. Unfortunately, accurately assessing compliance with new regulatory regimes remains a challenge, as data provenance, that is, metadata about data sharing, is rarely preserved. And even when provenance information is retained, unilateral changes to one party's database can render the provenance stale. To fill this gap, we present Proof-of-Retention a novel framework and interactive protocol that enforces auditable data sharing. Our framework requires each party involved in data sharing to retain a subset of information associated with each data exchange, a
ALT Agent memory enables enterprise agents to retain knowledge acquired during work and reuse it across tasks and agents, turning execution experience into persistent organizational knowledge. Realizing this potential requires both source–memory integration, through which enterprise sources and accumulated memory can be utilized together, and memory governance, through which shared memory remains subject to organizational policies throughout its lifecycle. These requirements interact when information from enterprise sources persists in memory. As this information is repeatedly derived and reused under changing principals and policies, source restrictions may be bypassed, resulting in information leakage. Preventing such leakage requires authorization continuity, under which source restrictions remain effective throughout source-to-memory and memory-to-memory derivation and reuse. Existing approaches address these concerns individually, but do not treat source–memory integration, memory gov
ALT Assigning a standard commodity code to a free-text purchase line underpins enterprise spend analytics, and its reported accuracy cannot be checked. Published results use proprietary data or private samples under undocumented protocols; no two compare and there is no public benchmark. We build a harness with four protocols over 1.26 million labelled purchase lines from two US state governments. Enterprises rebuy continuously, so random splits put matching item text on both sides: 60.7% and 61.2% of test rows, 54.7% and 60.6% byte for byte. On California orders the best classical baseline scores 56.0% on repeated text but 31.9% on novel text, a 24-point gap widening with taxonomy depth. Embedding retrieval shrinks the gap from 23.8 to 19.1 points; a fine-tuned transformer does not escape it. On one corpus leaky evaluation cannot separate retrieval from the transformer, while three leak-free protocols put the transformer 2.3 to 3.6 points ahead, so leaky splits hide real differences, not
ALT The OpenAlex Colors Working Group proposes replacing the familiar Gold, Green, Hybrid, Bronze and Closed taxonomy with a work-level framework based on three dimensions: access, location and license. The proposal addresses genuine weaknesses in the legacy color system, particularly its dependence on journal-level characteristics, its difficulty in describing non-article outputs, and its conflation of access conditions with venue business models. This commentary supports that general direction but argues that the proposed labels should not become the primary data model. The central technical issue is that access, location and license are not best understood as three independent attributes of a work: scholarly works commonly have several versions or copies at different locations, and access status, license and version may differ from one location to another. The analysis therefore recommends treating the location/version observation as the atomic unit of OA metadata and deriving work-leve
ALT Many emerging applications, from computational biology to digital twins and urban planning, rely heavily on three-dimensional spatial joins over polyhedral meshes. These joins comprise computationally-intensive triangle–triangle intersection tests that pairwise compare the faces of polyhedral meshes. Since each mesh may contain thousands of faces, the resulting cost challenges the responsiveness of spatial data management techniques. Existing techniques follow the filter-and-refine paradigm, accelerating either the filtering step through indexing or the refinement step through progressive mesh compression combined with GPU parallelization of triangle–triangle tests. However, the former neglects the high cost of intra-geometry refinements, whereas the latter lowers this cost but still relies on the same pairwise triangle–triangle tests. In this paper, we introduce Pierce, an approach that reformulates three-dimensional spatial joins over complex polyhedral meshes as ray-tracing operatio
ALT Knowledge graphs are a core component of today's knowledge infrastructure, supporting reasoning and anchoring knowledge systems to verifiable facts. RDF stores and SPARQL engines fulfill this function, enabling a range of retrieval and inference tasks on structured knowledge. Coupling them with Language Models (LMs) extends RAG toward neurosymbolic reasoning, where structured queries gate or re-rank generative outputs. This line of reasoning requires that SPARQL evaluation natively support tensor operations on dense embeddings, enabling multimodal querying and learned similarity-based ranking to be expressed together with graph-structural constraints. This approach is feasible only if the engine can efficiently perform dense vector search. We present QLever-Unified Indexed Vector Embedding Retrieval (QUIVER), an extension to QLever that adds native support for dense vector retrieval within RDF knowledge graphs. It implements three optimizations: engine-level registration of tensor func
ALT Traditional data systems face profound limitations in the AI era, relying on human-crafted pipelines, lacking semantic understanding of heterogeneous data, and operating through rigid, reactive processing. To address these challenges, we propose a new paradigm called the Data Agent, designed to manage, process, and analyze data with minimal human intervention. Data agents autonomously execute a wide range of data-related tasks, transforming traditional data systems by shifting from manual design to autonomous orchestration, from literal manipulation to semantic interpretation, and from reactive to proactive processing. Our Data Agent system includes six components: semantic data organization, semantic operators, agentic pipeline orchestration and optimization, feedback-driven refinement, memory management, and proactive adaptation. Building on this foundation, we also develop two specialized agents: the data analytics agent and the data science agent. Experiments on real benchmarks dem
ALT Open-vocabulary object retrieval in videos requires answering free-form object queries under bounded query-time cost. Existing index-based systems typically store independent frame-level regions and retrieve them with vision-language similarity, which is effective for appearance queries but mismatched with predicates whose evidence is temporal or relational, such as stopped state, scene-region occupancy, persistence, and object interactions. We identify this gap as an evidence-unit mismatch: the query is expressed over tracklets or object tuples, while the index stores isolated boxes. To address it, we propose STEG-OVR, a structured spatio-temporal evidence graph for open-vocabulary object retrieval. STEG-OVR represents persistent objects as tracklet nodes and temporally compatible object pairs as relation edges, storing appearance, motion, scene occupancy, relative geometry, velocity compatibility, and symbolic relation evidence. A query is decomposed into entity, state, scene, tempor
ALT Graph databases are frequently positioned as categorically necessary for connected-data workloads, yet the systems dimension along which they actually differ - query planning, indexing, and data-readiness cost - is rarely isolated from vendor framing. We construct a synthetic, biomedical-shaped property graph (1.02 million nodes, 5.34 million total node and edge rows) and a twenty-query workload spanning neighborhood lookups, bounded paths, set intersections, anti-joins, grouped aggregation, top-k ranking, temporal filters, full scans, and relational joins. We benchmark Corvic AI - a purpose-built columnar query engine underlying Corvic's ontology management layer ("memories")- against seven purpose-built or graph-extension database systems (LoraDB, Ladybug, DuckPGQ, Memgraph, Neo4j, HugeGraph, and FalkorDB) at three graph scales spanning three orders of magnitude. We report query latency geomeans, bulk-ingest throughput, point-update latency, and answer correctness for each system, an
ALT Approximate nearest neighbor search (ANNS) underpins large-scale vector retrieval in search, recommendation, and retrieval-augmented generation. Graph-based indexes have demonstrated state-of-the-art search performance for ANNS. They connect each corpus vector to a small set of nearby or navigationally useful vertices and answer queries by traversing the resulting graph. Because these edges are selected using construction-time distances, the graph index is tied to the embedding model. Re-encoding a corpus with a new model may change distances and neighborhoods of the vectors. Reconstructing the graph for the new embedding vectors incurs substantial construction cost and delays deployment. When the embedding model changes, we observe a phenomenon in the old graph index that we call residual reachability. Specifically, although derived from different models, the vectors describe the same underlying objects and often retain part of their similarity structure. These shared relations are re
ALT GQL (ISO/IEC 39075:2024) is the first international standard for a graph query language, and SQL/PGQ (ISO/IEC 9075-16:2023) embeds the same pattern-matching core, GPML, in SQL. Both fix a precise semantics for path patterns: four path modes, four selectors, quantified segments. What engines compute for those patterns has never been measured. Work on GQL is theoretical and runs no engine; cross-engine work measures performance and normalises the semantics away. We build an executable reference semantics for the GPML path core and gate it against the six worked answers printed in the standard's own reference exposition. We then run a 17-construct suite against six releases of five products - Kuzu, DuckPGQ, Neo4j at 5.26 and 2026.04, Memgraph, Apache AGE - scoring each cell as conforming, diverging, rejected, or inexpressible in that dialect. Nine of seventeen constructs draw more than one answer across the engines that accepted them. Of 26 disagreements, 15 are silent - the query runs, r
ALT We describe the integration of ioᵤring into Oracle Database's storage layer and the architectural decisions required to deploy it in a production multi-process RDBMS. Our design uses per-process ring contexts that eliminate inter-process synchronization, a shared buffer registration mechanism now part of the mainline Linux kernel, and a transparent fallback to libaio on error. Evaluation on an internal development build of Oracle Database 26ai shows that ioᵤring's benefits concentrate on asynchronous batched I/O paths: on a mixed OLTP workload (TPC-C), ioᵤring delivers identical throughput while reducing server CPU utilization by 1.2 percentage points through more efficient background writes; on analytical queries (TPC-H), CPU per query drops by 8.5% (geometric mean). Isolating the write path alone shows 29% lower CPU per write and 34% higher throughput. Synchronous read paths – the dominant I/O in OLTP – show no improvement and even increased CPU usage for very large I/O sizes. These
ALT Collaborative research projects in life sciences increasingly need to integrate private, embargoed consortium data with public reference databases in order to reach statistically meaningful interpretations. The Resource Description Framework (RDF) is well suited to this task: it facilitates the integration of heterogeneous data sources, and allows researchers to keep data and their documentation as metadata in the same place, provided the knowledge graph itself remains private during the time course of the project. Nevertheless, the development and long-term maintenance of a scientific knowledge graph remains a challenging, labour-intensive endeavour owing to the state of constant flux of most public resources. To tackle this challenge, we present kgsteward, a Python command-line tool that builds and maintains knowledge graphs inside RDF stores from a single, version-controlled configuration file. kgsteward supports multiple triplestores, keeps the local graph up-to-date with its exter
ALT SQL has been augmented with AI operators, enabling modern data analytics platforms to derive insights from both structured and unstructured data. We observe that while current Text-to-SQL systems can successfully generate these AI-augmented queries, reliably evaluating their correctness remains a critical open challenge. Current metrics, which rely on exact query results and deterministic execution, systematically fail against the flexible, non-deterministic outputs of AI operators. In this paper, we formalize these unique evaluation failure modes and introduce a Multilayered Evaluation Framework that decouples deterministic database logic from flexible AI semantics. We test our approach across both industry (BigQuery) and academic (ThalamusDB) systems. We demonstrate that traditional Execution Accuracy severely penalizes valid queries, achieving as low as a 25% detection rate for correct translations. Furthermore, even a state-of-the-art LLM-based autorater falsely rejects 32% of accu