Adaptive Approximate Nearest Neighbor Search under Dynamic Data Distributions in Large-Scale Information Platforms
- Authors
-
-
Sanjay Khadka
Pokhara University, Ratna Marg 18, Department of Information Systems Engineering, Pokhara, Gandaki Province, Nepal
Author
-
Prakash Adhikari
Mid-West University, Birendranagar Road 42, Department of Software Technology and Analytics, Surkhet, Karnali Province, Nepal
Author
-
Ariful Islam
Dhaka, Bangladesh
Author
-
- Abstract
-
Large-scale information platforms routinely rely on embedding-based retrieval to connect users to relevant items under strict latency and throughput constraints. Approximate nearest neighbor (ANN) search has become a practical backbone for this retrieval, yet real deployments rarely satisfy the stationary assumptions that make static ANN indexes effective. Item embeddings evolve due to continual model updates, content churn introduces new vectors while removing others, and query distributions shift with seasonality, UI changes, and feedback loops. These dynamics create a moving target in which index structures, quantization codebooks, routing policies, and search hyperparameters can become miscalibrated, degrading recall or inflating tail latency. This paper studies adaptive ANN search under dynamic data distributions with an emphasis on large-scale platforms that operate continuously and must reconcile update throughput with predictable quality of service. We formalize a time-indexed retrieval process in which both data and query distributions drift, and we connect drift to measurable failure modes such as graph fragmentation, centroid staleness, and widening quantization error. We then develop a system-level approach that couples drift detection, adaptive index maintenance, and online control of search effort. The approach treats ANN as a closed-loop service, using observed difficulty and outcomes to adjust routing, candidate generation, and refinement. Analytical models characterize trade-offs among recall, latency, and update costs under nonstationarity, and design principles are proposed for robustness when drift is abrupt, heterogeneous across shards, or induced by model rollouts.
- References
- Downloads
- Published
- 2026-01-04
- Section
- Articles