Software & ML Engineer specializing in end-to-end intelligent systems — from distributed data pipelines and backend infrastructure to LLM evaluation, GPU inference optimization, and low-resource language model research.
Software & ML Engineer with 2+ years of experience across backend infrastructure, LLM evaluation, and inference optimization.
M.S. in Data Science from Montclair State University (GPA 3.87/4.0). Researching low-resource language modeling and LLM fairness alongside engineering work.
Full-time AI Engineer, ML Engineer, or Data Engineer roles where I can own systems end-to-end — from the pipeline to the model to the API serving it.
RAG evaluation platform benchmarking 3 retrievers against 3 LLM providers — faithfulness, accuracy, latency, cost, and pass rate, with GitHub Actions gates that catch regressions before release.
4-agent async financial research system covering US, NSE, and BSE markets with plurality-vote signals, LLM-as-judge scoring, FinBERT sentiment, and Markowitz portfolio optimization.
Voice-to-SQL app converting natural-language queries into guarded SQL and Plotly charts over 3.4M orders and 30M+ line items, with dynamic allowlists and a 29-test eval suite.
Python/CLI library for adapting, extending, pruning, and initializing Hugging Face tokenizers and embedding layers, plus an end-to-end CPT → SFT → DPO pipeline for cross-lingual LLM fine-tuning.
Exactly-once change data capture pipeline: PostgreSQL WAL → Debezium → Kafka → Flink → Apache Iceberg, chaos-tested against TaskManager kills and network partitions.
Delta Lake medallion architecture (bronze/silver/gold) with a compaction engine cutting per-partition file count 95% (87 → 4 files) in 14.5s, zero reader downtime.
Real-time data quality and LLM incident-triage platform detecting 6 failure modes, with automated anomaly alerting integrated into CI/CD data pipelines.
Compiled ResNet-50 to an FP16 TensorRT 11 engine on CUDA 12.4 with zero accuracy loss across 10K CIFAR-10 images, served via a C++ HTTP server.
XGBoost classifier on 7,043 records with 50-trial Optuna tuning, gated promotion via a two-tier quality gate against the Champion model in the MLflow Registry.
Feature pipeline on NYC Taxi data enforcing train-serve parity via PySpark → Feast offline store → Redis. Skew caught at inference time using SHA-256 hash comparison.
Visual workflow automation platform: React Flow editor, Spring Boot DAG executor, JWT auth, OAuth credentials encrypted at rest (Jasypt AES-256), live SSE execution logs.
High-throughput inventory reservations for flash-sale spikes: Redisson Redlock holds, Kafka-backed checkout, live seat availability over WebSocket with polling fallback.
Real-time collaborative document editor with live multi-cursor editing via Yjs CRDTs, a Tiptap-based React frontend, and a Spring Boot WebSocket hub.
Built a native C++ inference engine on ONNX Runtime for RLHF reward-model scoring and benchmarked it against PyTorch eager mode, torch.compile, and FastAPI serving on CPU and GPU.
Trained 5 GPT-2-architecture causal language models from scratch for Telugu, Tamil, Kannada, and Malayalam (4 monolingual + 1 multilingual), benchmarked against XLM-R, mBERT, and mGPT.
A fairness-regression audit pipeline for detecting whether a new model version introduces fairness regressions, applied to 12 model versions from 5 providers across 7 public bias benchmarks (~200K examples).
Investigates whether machine translation (IndicTrans2) can generate BabyLM-scale training corpora for low-resource languages, translating the 100M-word English BabyLM corpus into Hindi and Telugu.
A federated learning framework combining capsule endoscopy imaging with gene expression data for privacy-preserving Crohn's disease diagnosis.
Evaluates subword tokenization across Telugu, Hindi, and English to measure token inflation, computational cost, and accuracy-efficiency trade-offs for low-resource, morphologically rich languages.
Looking for full-time AI Engineer, ML Engineer, or Data Engineer work. If you're hiring, or just want to talk about any of this, reach out.