Job Requirements
Responsibilities
Lead the design and development of scalable vectorization pipelines to convert text, PDFs, and multimodal data into high-quality embeddings using models like BGE, Ada, or Cohere.
Manage and optimize vector databases (Pinecone, Weaviate, Milvus, pgvector), including building efficient indexing strategies (HNSW, DiskANN).
Develop advanced chunking and parsing techniques to maintain context and improve retrieval accuracy.
Build and fine-tune hybrid search workflows combining dense vector search with sparse keyword methods (e.g., BM25).
Create evaluation frameworks and gold-standard datasets to measure retrieval performance (Hit Rate, MRR, Faithfulness).
Implement re-ranking pipelines using cross-encoders to refine results before sending them to LLMs.
Work closely with engineering and product teams to ensure scalable, high-performance retrieval syste
Work Experience
2-5Years