Physical Sciences › Computer Science › Signal Processing
Data Management and Algorithms
56 papers indexed
This topic and its hierarchy come from the OpenAlex classification, the open catalogue of the world's scientific research.
Monthly volume - last 12 months
Latest papers
- Tsubame: Tree Replay for Diffusion-Based Speculative Decoding
Yepeng Weng, Qiao Hu, Takehisa Yairi · 29 September 2026
Context-aware dynamic trees allocate the speculative decoding budget according to draft path probabilities, adapting their depth and branching to the current context. Under stochastic decoding, however, we find that this structural advantage does not always compensate for the acceptance gains of ran…
- Cartograph: Federated Tool Discovery with Operator-Attested Retrieval for AI Agents
Justice Owusu Agyemang, Michael Agyare, Kwame Opuni-Boachie Obour Agyekum, Kwame Agyeman-Prempeh Agyekum, Francisca Adoma Acheampong, Jerry John Kponyo · 28 September 2026
The Model Context Protocol (MCP) enables AI agents to discover and call tools, but loading every definition becomes expensive as connected catalogs grow. We present Cartograph, a federated MCP proxy that changes agent-visible tool discovery from $O(n)$ catalog traversal to $O(k)$ progressive disclos…
- Non-Commutative State Tracking with Input-Dependent Low-Rank Updates in Mamba-3
Hiroki Fujii, Masaki Yamakita · 24 September 2026
State tracking from sequential observations can require both retaining information and updating it by composing observed operations. We extend Mamba-3's diagonal transition with an input-dependent low-rank reflection term to support noncommutative state tracking, in which the order of operations mat…
- Active Spatial Inspection for Effective and Efficient Embodied Exploration
Wenbin Wang, Xiang Bai, Yizhao Wang, Hang Sun, Dong Ren, Jie Qin, Qingquan Li, Bing Wang · 22 September 2026
Achieving high task success and efficiency remains a central pursuit in embodied exploration. Existing frameworks typically guide agent behavior through a spatially coarse and indirect assessment of suggestive cues and directions, yet such designs may struggle to judge cue sufficiency and the need f…
- TripScore: Aligning LLMs for Real-World Travel Planning via Expert-Calibrated Reward
Yincen Qu, Huan Xiao, Feng Li, Gregory Li, Hui Zhou, Xiangying Dai, Xiaoru Dai, Xuan Huang · 18 September 2026
In our deployed travel-planning service, most users give minimal inputs or free-form requests rather than the structured constraint checklists assumed by existing benchmarks. We therefore present TripScore, a behavior-grounded benchmark and evaluation framework built from real user logs and calibrat…
- A Functional Pilot for Certified Freshness-Aware Semantic--Spatial Range Retrieval
Taimoor Ahmad · 18 September 2026
Geographic applications need every object inside a radius that satisfies a semantic threshold, yet embedding indexes return approximate top-ranked lists and may omit qualifying records silently. We present FRESH-GEORANGE, a semantic- spatial range design that separates source-watermark freshness fro…
- MiCRo: Mixture Modeling and Context-aware Routing for Personalized Preference Learning
Jingyan Shen, Jiarui Yao, Rui Yang, Yifan Sun, Feng Luo, Rui Pan, Tong Zhang, Han Zhao · 16 September 2026
Reward modeling is a key step in building safe foundation models when applying reinforcement learning from human feedback (RLHF) to align Large Language Models (LLMs). However, reward modeling based on the Bradley-Terry (BT) model assumes a global reward function, failing to capture the inherently d…
- GEOSTEER: Geodesic Optimization for Activation Steering in Large Language Models
Xuan Cuong Ngo, Hao Vo, Ngan Le · 11 September 2026
Activation steering provides a lightweight way to control large language models (LLMs) by modifying their hidden activations at inference time. Among these approaches, norm-preserving steering aims to change model behavior without altering the activation norm, reducing the risk of representation col…
- REFINE: Trajectory Representation Learning via Closed-Loop Transcription -- Extended Version
Sean Bin Yang, Ying Sun, Jilin Hu, Zongyi Xu, Kristian Torp, Hua Lu, Bin Yang, Christian S. Jensen · 9 September 2026
Trajectory representation learning underpins a wide range of trajectory analytics tasks; however, most existing self-supervised approaches, whether discriminative or generative, adopt an open-loop paradigm, relying on fixed data augmentations or random masking without feedback, which limits their ab…
- TraveL: Transformer-based Multi-view Path Distributional Representation Learning
Fang He, Tao-yang Fu, Wang-chien Lee · 4 September 2026
Path representation learning (PRL) for road networks has received increasing research attention, due to various path-related applications. Existing works on PRL typically exploit the co-occurrence relationship among road segments and paths to learn a vector as the path representation, without explor…
- Evidence for Shared Routing Geometry and Dynamics in Sparse Mixture-of-Experts
Kirill Labzin, Stepan Kulibaba, Artem Dzhalilov, Artem Gorokhov · 3 September 2026
Sparse mixture-of-experts (MoE) models use an independently parameterized router at each sparse layer to select experts for every token. Prior work has shown that routing decisions across depth can often be predicted from earlier routing signals, suggesting that routing is not fully independent acro…
- GeoSPRINT: Geometric Redundancy-Aware Step Pruning for Inference in Diffusion Trajectories
Arpita Joshi · 3 September 2026
Diffusion models achieve high sample quality but remain expensive at inference time because sampling requires many sequential neural function evaluations (NFEs). Existing acceleration methods either use fixed step-skipping schedules, adapt step sizes based on local numerical error, or require additi…
- A Power Law in Logarithm's Clothing: On the Scalability of Graph-Based Vector Search
Sajad Faghfoor Maghrebi, Navid Eslami, Niv Dayan · 3 September 2026
Most vector databases rely on graph-based indexes, notably HNSW and Vamana, for approximate nearest neighbor search. With embedding models widely adopted, the datasets these databases store grow rapidly. At a fixed accuracy, how does search cost scale with dataset size? The prevailing answer is poly…
- A Feature-Major Codebook for Memory-Efficient Sparse-Binary Self-Organizing Maps: Scaling a MEDLINE Atlas to 1.05 Million Neurons on a Single Consumer GPU
Andrew James Amos · 26 August 2026
A self-organising map turns a large corpus into a browsable two-dimensional atlas, but building one at MEDLINE scale has been impractical: the best-matching-unit (BMU) search that dominates training is bound by the bandwidth needed to read the codebook every epoch. I show that this bottleneck is lar…
- Mahalanobis-Based Multi-Head Attention for Complex State Propagation
Xiaohe Li · 26 August 2026
In this paper, we propose \textbf{Mahalanobis-Based Multi-Head Attention} (MHA-CSP), a novel attention mechanism that replaces the standard dot-product with a \textbf{Mahalanobis distance-based RBF kernel}, which effectively computes attention in an infinite-dimensional feature space without increas…
- Length-Controlled Margin-Based Preference Optimization without Reference Model
Gengxu Li, Tingyu Xia, Yi Chang, Yuan Wu · 25 August 2026
Direct Preference Optimization (DPO) is a widely adopted offline algorithm for preference-based reinforcement learning from human feedback (RLHF), designed to improve training simplicity and stability by redefining reward functions. However, DPO is hindered by several limitations, including length b…
- SceneGTMM: A Conformal Mapping-based Scene-Aware Transferable GNN-Transformer Dual-Graph Interaction Framework for Map Matching
Yongliang Zhang, Feng Song, Ji Chen, Lishuai Guo, Yong Deng, Yue Zheng, Tianyi Liu, Zhixiong Chen, Qixin Zhang · 21 August 2026
Map matching is a key technology connecting positioning data with high precision road networks, but it faces challenges in noise robustness, cross regional transfer, and interpretability. To addr ess the limitations of existing methods in local global fusion, dynamic road network adaptation, and rel…
- Beyond Retrieval: Query-Conditioned Reuse of Long-Horizon Agent Trajectories
Yifei Li, Heng Wang, Lingling Zhang, Muye Huang, Xinyu Zhang, Jiashuai Liu, Hang Yan, Rongman Xu · 14 August 2026
Retrieval can identify a past trajectory that may matter, yet it does not specify how an acting agent should use that trajectory after users, entities, constraints, or environment state have changed. We identify this post-retrieval reuse step as a distinct bottleneck for long-horizon trajectory memo…
- Trajectory inference via Acceleration Matching
Bartolo Dazzini, Giovanni Conforti, Alain Durmus, Aram-Alexandre Pooladian · 5 August 2026
Trajectory inference is a fundamental problem in many scientific domains: given a collection of unpaired snapshots of observations at discrete time points, the goal is to generate smooth trajectories that best resemble and interpolate the data. Existing algorithms exhibit computational challenges: t…
- Using Lower-Bound Representations for Trajectory Similarity Learning
Liwei Deng, Haotian Meng, Yupu Zhang, Yan Zhao, Torben Bach Pedersen, Kai Zheng, Christian S. Jensen · 4 August 2026
Trajectory similarity learning is fundamental to efficient trajectory retrieval under complex distance measures. Existing learning-based methods typically rely on embeddings trained to approximate trajectory distances or rankings, but they often lack guarantees with respect to the original distances…
- Journey Operators for Structured Multi-Axis Composition
Mahesh Godavarti · 30 July 2026
Many kinds of data have structure along one or more axes: words in a sentence, pixels in an image, nodes in a tree, frames in audio, or cells in a 3D volume. Along one axis, order matters: "the dog bit the man" is different from "the man bit the dog." Across independent axes, however, neither compos…
- Semantic Space Search Trajectory Networks
Julian Agudelo, Alberto Tonda, Gabriela Ochoa, Vincent Guigue, Cristina Manfredotti, Evelyne Lutton · 29 July 2026
Search Trajectory Networks (STNs) are a graph-based tool for visualizing and characterizing the behavior of optimization algorithms. STNs' reliance on discretization of the search space has largely confined them to low-dimensional or combinatorial settings. We introduce a methodology for constructin…
- Staypoint Detection from Noisy Trajectory Data [Experiment Paper]
Lance Kennedy, Hossein Amiri, Yueyang Liu, Riyang Bao, Hanqi Chen, Mohammad Hashemi, Ruochen Kong, Xiaotong Liu, Joon-Seok Kim, Shengpu Tang, Liang Zhao, Andreas Z\"ufle · 22 July 2026
Detecting staypoints from raw trajectory data is fundamental to numerous spatial computing applications. This process transforms raw numeric sequences of geolocations into semantically meaningful locations, such as homes, workplaces, or restaurants. Despite its importance for semantic trajectory ana…
- ANNLib: A Development Framework for Efficient Approximate Nearest Neighbor Search
Zheqi Shen, Jingbo Su, Zijin Wan, Yan Gu, Yihan Sun · 21 July 2026
Approximate Nearest Neighbor Search (ANNS) plays a pivotal role in modern deep learning pipelines. Recently, many ANNS systems have been proposed to either provide broad functionality or reach high performance. However, it is yet difficult to achieve both with minimal programming efforts. We propose…
- Lightning Fast Matching Dependency Discovery with Desbordante
Alexey Shlyonskikh, Michael Sinelnikov, Daniil Nikolaev, Yurii Litvinov, George Chernishev · 14 July 2026
Matching dependency is a generalization of the functional dependency concept, which allows users to apply custom similarity functions for matching individual attributes. Matching dependencies have a wide range of applications for solving various data quality problems, such as entity resolution, data…
