Agentic AISuketa Technology SolutionsSep 2025 to Dec 2025
Hybrid retrieval for large, mixed-format document sets
Finding the right passage in a large corpus that mixes medical papers, threat reports and law, across PDF, Word, CSV and JSON. We built a retrieval pipeline that ingests it all and returns the best five passages.
My roleIntern; two-person build, I owned chunking and the pipeline documentation
Raw files in, five ranked passages out per query, built with one other engineer.
- Every chunkPDF, Word, CSV, JSON; type-aware chunking
- 200 candidatesBM25 + dense search, fused with RRF
- Top 5Cross-encoder rerank
The problem
A large document set mixing medical research, threat-intelligence reports and legal text, in many file formats. One retrieval method can’t do justice to all of it.
What we built
Ingestion streams files through type-aware chunking (1,024-token chunks with 256 overlap), validates every chunk and bulk-loads it into Milvus. Queries run a hybrid search, keyword (BM25) and dense (all-mpnet-base-v2), fuse the results with reciprocal rank fusion over 200 candidates, and rerank them with a cross-encoder to the final five.
My part
I reworked the chunking strategy and wrote the full pipeline documentation, building it with one other engineer.
- Python
- Milvus
- BM25
- Sentence Transformers
- Docker