A little about me

About

I’m Vaibhav, an AGI safety researcher interested in making advanced AI systems easier to understand, evaluate, and guide. I work across mechanistic interpretability, scalable oversight, and neurosymbolic AI, exploring how we can build systems whose capabilities remain legible and accountable.

My projects and writing examine questions from model unlearning and red teaming to how we evaluate social reasoning in AI. I’m especially drawn to research that connects a clear picture of how models work with practical ways to make them more reliable.