Hi, I'm Megha.
I'm a research engineer at Google DeepMind, where I work on evals and post-training: figuring out what a model can actually do, and turning the gaps into better models.
- model
- megha-joshi
- role
- Research Engineer · Google DeepMind
- objective
- evals + post-training
- prior ckpts
- Gemini Omni (video) · Google Maps
- training data
- Yale B.S. CS & Econ · Penn M.S. AI
- side quests
- violin · running
- status
- ● 11 / 11 suites passing
I measure what models can actually do
I'm a Research Engineer at Google DeepMind working on evals and post-training. When a new model candidate shows up, my job is to find out how it stacks up, both against our earlier models and against the rest of the field.
Then I turn what I find into signal: new evals for capabilities we weren't measuring yet, and insights that shape the next round of post-training.
Climbing the hill on Gemini Omni
I was a core contributor to Gemini Omni, Google's model that creates video from any input and edits it through conversation.
I built the "hillclimb" eval suite for video-to-video. It's the set of hard benchmarks each checkpoint has to climb, and it surfaced pre-training gaps that shaped our post-training plan. I also built stylization and reference datasets for DPO and tuned the recipes.
An eval for your evals
EvalSquared is a diagnostic tool I built to make writing evals painless. It cut eval creation time from days to hours, and it's now widely used across Google DeepMind and YouTube.
Because it makes datasets easy to inspect, teams also use it to find and fix problems in large-scale training data. Behind it are auto-rater pipelines that mine diverse sources for hard cases.
Teaching Gemini to spot broken maps
On Google Maps I led a video-based Gemini analysis that watches the app the way a person would and flags visual issues. It took a vague problem all the way to production and caught 10.5× more regressions for 25M daily users.
Along the way I worked on LoRA fine-tuning and prompting strategies like task decomposition. I also built an AI bug-deduplication system that cut report volume by about 80%.
Yale, then Penn
I studied Computer Science & Economics at Yale with a minor in Statistics & Data Science, then did an M.S. in Artificial Intelligence at the University of Pennsylvania.
The economics half still shows up in how I think about evals. A benchmark is a set of incentives, and models optimize for whatever you measure.
Running the Presidio
I run. One highlight was the Presidio 10K in San Francisco, with its hills, its trails, and the Golden Gate Bridge in view.
Training for a race is a lot like hillclimbing a model: small, measurable steps, one checkpoint at a time.
Playing with the South Bay Philharmonic
I'm a proud member of the South Bay Philharmonic, where I play violin.
An orchestra is a good lesson in collaboration: dozens of people, one tempo, and everyone listening as hard as they play.
New Haven to Bangalore, and back
College in New Haven, grad school in Philadelphia, work in Mountain View. I've flown to Bangalore to help onboard new Googlers and speak at the Geo EngProd summit.
Conferences took me to Vancouver for ICML 2025 and Orlando for Grace Hopper.
A few things I'm proud of
Sweety Silver Award: Google-wide recognition for advancing UI regression detection with Gemini vision models. Geo Tech Impact Award: for the video-based Maps analysis framework.
Lead contributor on two defensive publications: UI anomaly detection with prompt engineering and regression detection with generative video models.
Explaining things in rooms full of people
I've spoken at Google's LLM Applications Conference and the Geo EngProd Bangalore Summit, and I co-wrote my old org's 2025 AI strategy.
Recent rolls: graduation at Yale, ICML in Vancouver, Grace Hopper, a Philharmonic concert, and the Presidio 10K.
Summary: 11 / 11 suites passed
| # | Suite | Finding | Result |
|---|---|---|---|
| 01 | Eval Harness | Evals + post-training at Google DeepMind | ✓ running |
| 02 | Hillclimb | Gemini Omni V2V eval suite | ✓ shipped May 2026 |
| 03 | EvalSquared | Eval creation: days → hours | ✓ adopted org-wide |
| 04 | Map Tile QA | 10.5× more regressions caught | ✓ 25M daily users |
| 05 | Bookshelf | Yale B.S. CS & Econ · Penn M.S. AI | ✓ 2 degrees |
| 06 | Presidio 10K | 10.0 km | ✓ finished |
| 07 | Metronome | South Bay Philharmonic · violin | ✓ in tempo |
| 08 | Globe | New Haven → Bangalore → Vancouver | ✓ 7 places |
| 09 | Badge Wall | 2 awards · 2 defensive pubs | ✓ 8 badges |
| 10 | Camera Roll | Talks, conferences, finish lines | ✓ 7 frames |
| 11 | Mailbox | m3gha.joshi@gmail.com | ✓ open |
Let's talk evals
I love talking about measuring models: what makes a benchmark hard to game, how to read an auto-rater, and when a number really means "ship it."
Email m3gha.joshi@gmail.com, grab 30 minutes, or find me on LinkedIn. For my résumé, just ask by email.