Autonomous RAG-based Enterprise Search & Analytics: A Hybrid Retrieval Framework for Grounded Question Answering
DOI:
https://doi.org/10.55544/sjmars.5.5.3Keywords:
retrieval-augmented generation, enterprise search, hybrid retrieval, BM25, reciprocal rank fusion, dense vector search, autonomous agents, business analytics, large language models, grounded generationAbstract
Enterprise organizations increasingly hold their institutional knowledge in scattered, unstructured document repositories that traditional keyword search cannot interpret reliably. This paper presents an Autonomous RAG-based Enterprise Search & Analytics system, a prototype platform that couples hybrid information retrieval with large language model (LLM) generation and a live operational analytics layer. The retrieval stage runs a custom BM25 keyword ranker and a dense vector ranker in parallel, and a configurable fusion layer — Reciprocal Rank Fusion (RRF) or weighted score fusion, switchable at run time — merges the two ranked lists before the top-scoring passages are handed, with citations, to Claude 3.5 Sonnet for grounded answer generation. A companion dashboard exposes real-time key performance indicators so that operators can observe and tune retrieval behaviour without redeploying the system. We report a preliminary pilot evaluation of the prototype over a corpus of ten enterprise policy documents and roughly 30 to 40 test queries: a mean end-to-end latency of 2.0 seconds against a 3-second target, top-1 retrieval scores in the 75 to 95% range, approximately 70% of queries classified as high-confidence, and zero crashes across the test session. The pilot is explicitly small in scale and its numbers are treated as indicative rather than generalizable; the paper is correspondingly transparent about this limitation and about the prototype's current use of placeholder embeddings and an in-memory vector store, both slated for replacement in future iterations.
References
[1] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, (2017). Attention is all you need. Advances in Neural Information Processing Systems, 30, 5998–6008.
[2] P. Lewis, E. Perez, A. Piktus, F. Petroni, V. Karpukhin, N. Goyal, H. Küttler, M. Lewis, W. Yih, T. Rocktäschel, S. Riedel, and D. Kiela, (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459–9474.
[3] V. Karpukhin, B. Oğuz, S. Min, P. Lewis, L. Wu, S. Edunov, D. Chen, and W. Yih, (2020). Dense passage retrieval for open-domain question answering. Proc. 2020 Conf. on Empirical Methods in Natural Language Processing (EMNLP), 6769–6781.
[4] S. Robertson and H. Zaragoza, (2009). The probabilistic relevance framework: BM25 and beyond. Foundations and Trends in Information Retrieval, 3(4), 333–389.
[5] G. V. Cormack, C. L. A. Clarke, and S. Büttcher, (2009). Reciprocal rank fusion outperforms Condorcet and individual rank learning methods. Proc. 32nd Int. ACM SIGIR Conf. on Research and Development in Information Retrieval, 758–759.
[6] J. Johnson, M. Douze, and H. Jégou, (2019). Billion-scale similarity search with GPUs. IEEE Transactions on Big Data, 7(3), 535–547.
Downloads
Published
How to Cite
Issue
Section
License

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.

