Recursive agent engineering

Robust agents, built through extreme vibe coding.

I build agent systems where evaluation is not the ending. It is the feedback signal for the next version.

I'm Pratik Bhavsar. I turn agent behavior into better agents.

01 / BuildAgent systems
02 / MeasureOpen benchmarks
03 / ShareBooks & field guides
04 / ImproveFeedback loops
5
Books
25K+
Followers
100+
Articles
11
Talks
6
Startups

I build the loop where agent traces become better agents.

The interesting work starts after an agent runs: tool paths, judge disagreements, and failed actions become the raw material for the next version. Evals are the control system that lets agents improve run to run.

01

Instrument the run

Capture tool calls, retrievals, memory, and judge signals.

02

Select what survives

Separate durable behavior from lucky demos.

03

Patch the agent

Feed accepted traces into prompts, tools, and regression suites.

04

Publish the loop

Share lessons as books, leaderboards, and patterns.

Measuring agent behavior in the open.

Benchmarks should expose how systems behave under real pressure: multi-turn tasks, factuality, retrieval, and tool use.

Featured benchmark

Agent Leaderboard v2

30+ models on 500 multi-turn support scenarios across banking, healthcare, investment, telecom, and insurance.

Agent Leaderboard v2 blog graphic
Agent Leaderboard v2 launch graphic
Hallucination Index
Model reliability
Tracking hallucination rates across foundation models
Reliability benchmark

Hallucination Index

Full-stack AI, from models to communities.

I work on tokenomics and agent evaluations at Cisco, which acquired Galileo in May 2026.

Before that: founding engineer at Enterpret, principal data scientist at TaskHuman, and first quantitative research hire at Morningstar, where I launched end-to-end AI initiatives.

My RAG evaluation research was featured in Andrew Ng's newsletter. I've been a guest on the Latent Space Podcast and was named one of the Top AI Developers to Watch in 2023.

Education

IIT Bombay

M.Tech.

Cisco / Galileo

Agent evals + tokenomics

Galileo employee #26, Series A.

Enterpret

Employee #6

Founding engineer.

Recognition

Top AI Developer

AI developers to watch, 2023.

Applying AI at the edge of products.

May 2026 - present

Cisco

Tokenomics and agent evaluations.

June 2023 - May 2026

Galileo

Led open-source evaluations and developer relations. Built the Agent Leaderboard, Hallucination Index, and BRAG; wrote five GenAI books.

2019 - present

Maxpool

Founded and grew a community of AI professionals.

Feb 2021 - May 2023

Enterpret

Pre-seed. Semantic search, reranking, text generation, MLOps, and NLP pipelines for customer feedback.

Sep 2020 - Dec 2020

Jina AI

Employee #6, seed. Contributed to its open-source multi-modal neural search framework.

2017 - 2020

TaskHuman and Morningstar

TaskHuman: semantic search with transformers, and recommendations. Morningstar: NLP extraction, sentiment, and quantitative ML.

Let's build the next agent loop.

Reach out about agents, evals, writing, or speaking.