Onkar Salvi

Senior Research Engineer · AI Agents / LLMs / ML / NLP

Research Engineer with 7+ years building, training, and scaling ML/NLP systems hand-in-hand with research scientists. Depth in LLM finetuning (LoRA/PEFT), contrastive & representation learning, agentic systems, and model optimization — turning research ideas into production systems that process 600M+ messages daily. Published in AAAI; 1st place at the DSTC7 dialogue challenge.

About Me

I'm a Research Engineer with deep expertise in AI agents, Large Language Models, and Natural Language Processing. Currently a Senior Research Engineer at Dataminr, I work hand-in-hand with research scientists to turn ideas into production systems — from multi-agent workflows and LLM finetuning to model optimization and serving infrastructure.

My work spans the entire AI lifecycle: LoRA/PEFT finetuning, contrastive & representation learning, agentic (semantic + keyword) retrieval, and cost-efficient inference at scale. I have a strong bias toward clean, reusable research infrastructure and fast, ablation-driven experimentation.

I've published in AAAI, placed 1st at the DSTC7 dialogue challenge, and served as a reviewer for the AIWILD Workshop at ICML and ICLR 2025.

7+
Years building & scaling ML/NLP systems
600M+
Daily messages processed by systems I've built
$800K
Annual savings from inference cost optimization
AAAI
Published researcher & DSTC7 challenge winner

Experience

Senior Research Engineer
Dataminr | New York, NY
April 2021 – Present
  • Architected multi-agent workflows for the Knowledge Accumulation pipeline powering Tailored Intelligence, Dataminr's flagship offering — agents autonomously gather client-specific context to enable per-client impact analysis of every alert
  • Re-architected app search from keyword lookup to agentic hybrid (semantic + keyword) retrieval — a Query Analyzer for intent/multi-hop decomposition over an OpenSearch vector store scaled to 100M+ vectors
  • LoRA fine-tuned LLaMA-3.1-8B in a two-stage architecture (contrastive-learning pre-filter + LLM deep analysis) for real-time threat detection across 600M+ daily messages; ablations lifted F1 by 23% (65% → 88%) at sub-second latency
  • Built an agent-assisted annotation pipeline supplying cultural context for threat classification, cutting annotation cost ~66% vs. human-only labeling
  • Shipped a Sales Enablement app combining a web-search agent with a RAG workflow — 3× email open rate and 60% lift in response rate
  • Built the company's first Model Context Protocol (MCP) server (web search, document parsing, internal API access); integrated Arize Phoenix and Langfuse for LLM/agent observability
  • Built and maintained the in-house model-serving framework (gRPC + REST proxy) deploying NLP/CV/LLM models on Kubernetes (AWS EKS), with metric-aware autoscaling on queue length, latency, and request volume
  • Led migration of 40+ AI services from ECS to Kubernetes and cut inference cost 70% across 14+ models via AWS Inferentia, graph optimization, and dynamic batching — $800K annual savings
AI Engineer
OneConnect Financial Technology (GammaLab) | New York, NY
Sept 2018 – April 2021
  • Researched and developed ML/DL models for natural language understanding, chatbots, machine reading comprehension, and NLG
  • 1st place, DSTC7 Dialog System Technology Challenge (knowledge-grounded conversation) — work published at AAAI
  • Built a zero-shot classification model reaching 98% accuracy on seen and 90% on unseen classes, eliminating retraining on new data
  • Applied BERT + K-Means clustering over large call-transcript corpora to group semantically similar utterances; used Celery for asynchronous model retraining
  • Owned a ranking model that ordered related news drawn from multiple sources

Skills & Technologies

🤖 Deep Learning & Research

PyTorch Transformers LoRA/PEFT Finetuning Contrastive & Representation Learning RAG Agentic Workflows Evaluation & Ablation Prompt Engineering Zero-/Few-Shot Learning

⚡ Model Optimization & Serving

vLLM ONNX TorchServe Dynamic Batching AWS Inferentia gRPC Quantization

☁️ Infrastructure & MLOps

Docker Kubernetes (AWS EKS) AWS OpenSearch MLflow Databricks Celery LiteLLM Langfuse Arize Phoenix

💻 Languages

Python Distributed / Multi-GPU Training

Education

M.S., Electrical Engineering
New York University
Sept 2016 – May 2018
B.E., Electronics Engineering
University of Mumbai
Sept 2011 – Jul 2015

Publications & Service

  • J. Zheng, O. Salvi, J. Chan. "Candidate Attended Dialogue State Tracking using BERT." AAAI, 2020.
  • J. Zheng, S. Kasturi, M. Lin, X. Chen, O. Salvi, H. J. Wang. "OneConn-MemNN System for Knowledge-Grounded Conversation Modeling." AAAI, 2019.
  • Reviewer, AIWILD Workshop, ICML 2025
  • Reviewer, AIWILD Workshop, ICLR 2025

Let's Connect

📧

Email

onkar.salvi7@gmail.com

📍

Location

Ramsey, NJ