Agents & retrieval
Multi-agent orchestration, tool-calling and routing — with RAG, GraphRAG and memory infrastructure owned end to end.
Two years shipping production LLM systems — eval harnesses, red-teaming, retrieval and memory infrastructure.
scrollLegal AI, health assistants, retrieval and memory infrastructure — owned end to end, measured before they ship.
Multi-agent orchestration, tool-calling and routing — with RAG, GraphRAG and memory infrastructure owned end to end.
LoRA / QLoRA and DeepSpeed ZeRO distributed runs — embeddings, rerankers and low-resource translation.
Eval harnesses, golden sets and LLM-as-judge scoring — red-teaming, prompt-injection testing, OWASP LLM Top 10.
Parse corpora, OCR and structure into DB/CSV-grounded sources.
$ ingest --ocr --ground db,csv ✓ 8.0M judgmentsLoRA / QLoRA adaptation, DeepSpeed ZeRO runs and dataset curation.
$ train --lora r=16 --zero 3 ✓ loss 0.41 → 0.18Wire agents, tools and routing into a reliable workflow.
$ route --agents 42 --tools mcp ✓ p95 1.2sEval harnesses, red-teaming and regression gates before every deploy.
$ gate --golden 4M --redteam ✓ 96% zero-fail · PASSAn evidence-grounded health assistant built with Dr. Mohan Tanirru. I fine-tuned open-source LLMs with LoRA/QLoRA, own the evaluation pipeline — groundedness scoring, hallucination detection, regression gates — and red-team it against OWASP LLM Top 10, with automatic clinician-referral escalation on high-risk queries.
Full experience →Multi-agent orchestration routing each request to the right specialist — 42 agents, config-driven.
Compliance agent with hybrid RAG, grounding guardrails and LLM-as-judge evaluation.
Swap LLM providers inside Claude Code — a published npm package, global install.
Full fine-tune of Qwen for Hindi → Gujarati on a curated parallel corpus.
Agent crews in Rust — typed tool schemas, async runtime, zero-copy message passing.
Hybrid function-calling router on Gemma — local-first tool selection with a cloud fallback.