VJ/
loading the lab
NPU · LIVE
INIT
booting inference core
LAYER ACT.0 tok/s
VJ/
AI / ML engineer — LLM evaluation · agentic systems

Vatsa Joshi

Two years shipping production LLM systems — eval harnesses, red-teaming, retrieval and memory infrastructure.

scroll
01 — what I do

Legal AI, health assistants, retrieval and memory infrastructure — owned end to end, measured before they ship.

0+
years shipping LLM systems
0×
cheaper legal-NER extraction
0%
document-level zero-failure
0M
row ground-truth eval set
02 — focus

routing01

Agents & retrieval

Multi-agent orchestration, tool-calling and routing — with RAG, GraphRAG and memory infrastructure owned end to end.

LangGraph · CrewAI · GraphRAG · Qdrant
training02

Fine-tuning & training

LoRA / QLoRA and DeepSpeed ZeRO distributed runs — embeddings, rerankers and low-resource translation.

LoRA / QLoRA · Unsloth · DeepSpeed
gated03

Evaluation & AI security

Eval harnesses, golden sets and LLM-as-judge scoring — red-teaming, prompt-injection testing, OWASP LLM Top 10.

LLM-as-judge · red-team · OWASP
03 — how the work gets built

01 — ingest

Ingest & ground

Parse corpora, OCR and structure into DB/CSV-grounded sources.

$ ingest --ocr --ground db,csv ✓ 8.0M judgments
02 — train

Train & fine-tune

LoRA / QLoRA adaptation, DeepSpeed ZeRO runs and dataset curation.

$ train --lora r=16 --zero 3 ✓ loss 0.41 → 0.18
03 — orchestrate

Orchestrate

Wire agents, tools and routing into a reliable workflow.

$ route --agents 42 --tools mcp ✓ p95 1.2s
04 — ship

Ship & harden

Eval harnesses, red-teaming and regression gates before every deploy.

$ gate --golden 4M --redteam ✓ 96% zero-fail · PASS
04 — most recent

live on AWS · Arizona State University · Jun 2024 — present

An evidence-grounded health assistant built with Dr. Mohan Tanirru. I fine-tuned open-source LLMs with LoRA/QLoRA, own the evaluation pipeline — groundedness scoring, hallucination detection, regression gates — and red-team it against OWASP LLM Top 10, with automatic clinician-referral escalation on high-risk queries.

LoRA / QLoRA fine-tunesgroundedness evals · regression gatesred-teaming · OWASP LLM Top 10retrieval & memory architecturequantized GPU inference · AWS
Full experience →
06 — let's build