Ayush Patravali

Search the site

Jump to a page, project, paper or journey step

Agentic AISuketa Technology SolutionsSep 2025 to Dec 2025

Hybrid retrieval for large, mixed-format document sets

Finding the right passage in a large corpus that mixes medical papers, threat reports and law, across PDF, Word, CSV and JSON. We built a retrieval pipeline that ingests it all and returns the best five passages.

My roleIntern; two-person build, I owned chunking and the pipeline documentation

Raw files in, five ranked passages out per query, built with one other engineer.

  1. Every chunkPDF, Word, CSV, JSON; type-aware chunking
  2. 200 candidatesBM25 + dense search, fused with RRF
  3. Top 5Cross-encoder rerank
200 fused candidates, reranked to 5.

The problem

A large document set mixing medical research, threat-intelligence reports and legal text, in many file formats. One retrieval method can’t do justice to all of it.

What we built

Ingestion streams files through type-aware chunking (1,024-token chunks with 256 overlap), validates every chunk and bulk-loads it into Milvus. Queries run a hybrid search, keyword (BM25) and dense (all-mpnet-base-v2), fuse the results with reciprocal rank fusion over 200 candidates, and rerank them with a cross-encoder to the final five.

My part

I reworked the chunking strategy and wrote the full pipeline documentation, building it with one other engineer.