Robust agents, built through extreme vibe coding.
I build agent systems where evaluation is not the ending. It is the feedback signal for the next version.
I'm Pratik Bhavsar. I turn agent behavior into better agents.
Eval Engineer
Claude Code and Codex skills that turn Galileo evidence into a diagnosis, a bounded fix plan, and a verification plan.
BRAG open-source model release
Fine-tuned open-source LLMs for RAG, built to improve grounded answers, not just retrieval.
Books for building production AI systems.
Five field guides in the Mastering GenAI series.
Eval Engineering
The emerging discipline of building trust in production AI systems.
Mastering Multi-Agent Systems
Patterns for agents that coordinate and hand off work.
Agent systemsMastering AI Agents
Tools, memory, planning, and evaluation.
240+ pagesMastering RAG 2.0
Retrieval, generation, and evaluation patterns.
Judge pipelinesMastering LLM-as-a-Judge
Designing and calibrating model-based evals.
I build the loop where agent traces become better agents.
The interesting work starts after an agent runs: tool paths, judge disagreements, and failed actions become the raw material for the next version. Evals are the control system that lets agents improve run to run.
Instrument the run
Capture tool calls, retrievals, memory, and judge signals.
Select what survives
Separate durable behavior from lucky demos.
Patch the agent
Feed accepted traces into prompts, tools, and regression suites.
Publish the loop
Share lessons as books, leaderboards, and patterns.
Measuring agent behavior in the open.
Benchmarks should expose how systems behave under real pressure: multi-turn tasks, factuality, retrieval, and tool use.
Agent Leaderboard v2
30+ models on 500 multi-turn support scenarios across banking, healthcare, investment, telecom, and insurance.
Full-stack AI, from models to communities.
I work on tokenomics and agent evaluations at Cisco, which acquired Galileo in May 2026.
Before that: founding engineer at Enterpret, principal data scientist at TaskHuman, and first quantitative research hire at Morningstar, where I launched end-to-end AI initiatives.
My RAG evaluation research was featured in Andrew Ng's newsletter. I've been a guest on the Latent Space Podcast and was named one of the Top AI Developers to Watch in 2023.
IIT Bombay
M.Tech.
Agent evals + tokenomics
Galileo employee #26, Series A.
Employee #6
Founding engineer.
Top AI Developer
AI developers to watch, 2023.
Learning in public sharpens the work.
Speaking and writing force ideas to survive contact with practitioners.
Ranking Agentic LLMs
The Agent Leaderboard and measuring real agent behavior.
02 / DAIR.AI / 1.3K viewsAI Agent Evaluation
Testing agents beyond one-shot answers.
03 / Galileo Live / 213 viewsBehind Agent Leaderboard 2.0
Building a public agent benchmark for real workflows.
04 / DAIR.AI / 3.8K views101 Ways to Solve Search
Search systems and retrieval quality.
05 / WiMLDS / 2.6K viewsModeling Fallacies in NLP
Traps in NLP modeling and data.
06 / PyData Mumbai / 538 viewsAutomated Machine Learning
Earlier work on AutoML.
Applying AI at the edge of products.
Cisco
Tokenomics and agent evaluations.
Galileo
Led open-source evaluations and developer relations. Built the Agent Leaderboard, Hallucination Index, and BRAG; wrote five GenAI books.
Maxpool
Founded and grew a community of AI professionals.
Enterpret
Pre-seed. Semantic search, reranking, text generation, MLOps, and NLP pipelines for customer feedback.
Jina AI
Employee #6, seed. Contributed to its open-source multi-modal neural search framework.
TaskHuman and Morningstar
TaskHuman: semantic search with transformers, and recommendations. Morningstar: NLP extraction, sentiment, and quantitative ML.
The loop gets better when more builders can inspect it.
I founded Maxpool for AI engineers building production systems.
Let's build the next agent loop.
Reach out about agents, evals, writing, or speaking.