Dipesh Tharu Mahato

I am an AI safety and robotics researcher working on reliable evaluation, runtime safeguards, interpretability, and trustworthy machine learning systems. My research studies how intelligent systems fail under real-world conditions and how those failures can be measured, monitored, and controlled. I am particularly interested in agent safety, vision-language-action models, distribution shift, and statistical methods for dependable AI evaluation. I completed my master’s degree in Data Science at New York University’s Center for Data Science. Before founding and directing LatentOps, I was Co-founder and CTO at Offtrack.

Portrait of Dipesh Tharu Mahato
For an up-to-date list, visit my Google Scholar profile.

Stable Encodings, Changing Downstream Sensitivity: Measuring SAE Feature Identity Across Fine-Tuning August 2026

Dipesh Tharu Mahato*

I separate what an SAE coordinate encodes from how the adapted model uses it, showing that semantic selectivity can persist across fine-tuning even as downstream sensitivity and intervention effects change.

ProxyGuard: Direct Reliability Inference for Randomized Data Release Mechanisms with Shared Targets August 2026

Dipesh Tharu Mahato*, Pramod Dhungana

We develop finite-sample methods for telling whether a randomized data-release mechanism is genuinely reliable, rather than judging it by a single fortunate release.

Safeguard-Conditioned Uplift: Measuring Utility-Risk Frontiers for Dual-Use Biology Assistants June 2026

Dipesh Tharu Mahato*

I study how the access conditions around dual-use assistants change both legitimate utility and harmful actionability, treating safeguard evaluation as a deployment-level measurement problem.

Early Warning Signals for OpenVLA Failure under Visual Distribution Shift June 2026

Dipesh Tharu Mahato*, Rachel Ren

We test whether OpenVLA’s internal activations contain useful warning signals before task failure when visual conditions shift, and examine where retrospective discrimination falls short of deployable monitoring.

TriGuard: Testing Model Safety with Attribution Entropy, Verification, and Drift June 2025

Dipesh Tharu Mahato*, Rohan Poudel, Pramod Dhungana

We combine formal robustness verification, attribution entropy, and explanation drift to expose model fragilities that accuracy and adversarial robustness alone can miss.

Enabling Explicit Memory in Pretrained LLMs for Enhanced Reasoning and Efficiency May 2025

Junzhi Chen, Dipesh Tharu Mahato*, Xinran Lin, Akshay Nuthanapati

We explore sparse explicit key-value memory inside pretrained language models as an alternative to repeatedly placing retrieved documents in context, with the goal of improving reasoning efficiency.

Candy Clouds: Sharp Minima Can Generalize for Deep Nets April 2025

Candice Yao, Carina Sun, Dipesh Tharu Mahato*, Tanishq Sardana

We examine why common flatness and sharpness measures can be misleading in deep networks, where reparameterization and model symmetries can change loss geometry without changing predictions.