logo final

Adaptive Approximate Nearest Neighbor Search under Dynamic Data Distributions in Large-Scale Information Platforms

Authors
  • Sanjay Khadka

    Pokhara University, Ratna Marg 18, Department of Information Systems Engineering, Pokhara, Gandaki Province, Nepal

    Author

  • Prakash Adhikari

    Mid-West University, Birendranagar Road 42, Department of Software Technology and Analytics, Surkhet, Karnali Province, Nepal

    Author

  • Ariful Islam

    Dhaka, Bangladesh

    Author

Abstract

Large-scale information platforms routinely rely on embedding-based retrieval to connect users to relevant items under strict latency and throughput constraints. Approximate nearest neighbor (ANN) search has become a practical backbone for this retrieval, yet real deployments rarely satisfy the stationary assumptions that make static ANN indexes effective. Item embeddings evolve due to continual model updates, content churn introduces new vectors while removing others, and query distributions shift with seasonality, UI changes, and feedback loops. These dynamics create a moving target in which index structures, quantization codebooks, routing policies, and search hyperparameters can become miscalibrated, degrading recall or inflating tail latency. This paper studies adaptive ANN search under dynamic data distributions with an emphasis on large-scale platforms that operate continuously and must reconcile update throughput with predictable quality of service. We formalize a time-indexed retrieval process in which both data and query distributions drift, and we connect drift to measurable failure modes such as graph fragmentation, centroid staleness, and widening quantization error. We then develop a system-level approach that couples drift detection, adaptive index maintenance, and online control of search effort. The approach treats ANN as a closed-loop service, using observed difficulty and outcomes to adjust routing, candidate generation, and refinement. Analytical models characterize trade-offs among recall, latency, and update costs under nonstationarity, and design principles are proposed for robustness when drift is abrupt, heterogeneous across shards, or induced by model rollouts.

References
Downloads
Published
2026-01-04
Section
Articles

How to Cite

[1]
S. Khadka, P. Adhikari, and A. Islam, “Adaptive Approximate Nearest Neighbor Search under Dynamic Data Distributions in Large-Scale Information Platforms”, JASCAR, vol. 16, no. 1, pp. 1–17, Jan. 2026, Accessed: Sep. 18, 2026. [Online]. Available: https://scichronicle.com/index.php/JASCAR/article/view/AdaptiveApproximateNearest