megha-joshi · eval report v2026.10
suite 00/11
Eval report · subject

Hi, I'm Megha.

I'm a research engineer at Google DeepMind, where I work on evals and post-training: figuring out what a model can actually do, and turning the gaps into better models.

Model cardv2026.10
model
megha-joshi
role
Research Engineer · Google DeepMind
objective
evals + post-training
prior ckpts
Gemini Omni (video) · Google Maps
training data
Yale B.S. CS & Econ · Penn M.S. AI
side quests
violin · running
status
● 11 / 11 suites passing
SCROLL ↓ · 11 SUITES · EVERY FIGURE IS INTERACTIVE
Megha Joshi on a balcony overlooking Lucerne and the Alps
Suite 00● subject
Click to re-runevaluating subject…
Suite 01Eval Harness
Click to run the suiteawaiting candidate
Now · Google DeepMind

I measure what models can actually do

I'm a Research Engineer at Google DeepMind working on evals and post-training. When a new model candidate shows up, my job is to find out how it stacks up, both against our earlier models and against the rest of the field.

Then I turn what I find into signal: new evals for capabilities we weren't measuring yet, and insights that shape the next round of post-training.

Research EngineerEvalsPost-trainingsince 2026
Suite 02Hillclimb
Click to train a checkpointckpt 0 · base
2025 – 2026 · YouTube × DeepMind

Climbing the hill on Gemini Omni

I was a core contributor to Gemini Omni, Google's model that creates video from any input and edits it through conversation.

I built the "hillclimb" eval suite for video-to-video. It's the set of hard benchmarks each checkpoint has to climb, and it surfaced pre-training gaps that shaped our post-training plan. I also built stylization and reference datasets for DPO and tuned the recipes.

V2V evalsDPOFlume pipelineslaunched May 2026
Suite 03EvalSquared
Click to scan the dataset0 / 48 rows · 0 flagged
Tooling

An eval for your evals

EvalSquared is a diagnostic tool I built to make writing evals painless. It cut eval creation time from days to hours, and it's now widely used across Google DeepMind and YouTube.

Because it makes datasets easy to inspect, teams also use it to find and fix problems in large-scale training data. Behind it are auto-rater pipelines that mine diverse sources for hard cases.

days → hoursauto-ratersadopted org-wide
Suite 04Map Tile QA
Click to catch a regressionwatching 25M users' maps
2023 – 2025 · Google Geo

Teaching Gemini to spot broken maps

On Google Maps I led a video-based Gemini analysis that watches the app the way a person would and flags visual issues. It took a vague problem all the way to production and caught 10.5× more regressions for 25M daily users.

Along the way I worked on LoRA fine-tuning and prompting strategies like task decomposition. I also built an AI bug-deduplication system that cut report volume by about 80%.

10.5× regressions caught−80% duplicate bugsflakes 50% → <1%
Suite 05Bookshelf
Click to pull a book4 volumes
Education

Yale, then Penn

I studied Computer Science & Economics at Yale with a minor in Statistics & Data Science, then did an M.S. in Artificial Intelligence at the University of Pennsylvania.

The economics half still shows up in how I think about evals. A benchmark is a set of incentives, and models optimize for whatever you measure.

B.S. YaleM.S. Penn
Suite 06Presidio 10K
Click to run a kmkm 0.0 / 10.0
Weekends

Running the Presidio

I run. One highlight was the Presidio 10K in San Francisco, with its hills, its trails, and the Golden Gate Bridge in view.

Training for a race is a lot like hillclimbing a model: small, measurable steps, one checkpoint at a time.

Suite 07Stand & Metronome
Click to keep time♩ = — · tacet
Orchestra

Playing with the South Bay Philharmonic

I'm a proud member of the South Bay Philharmonic, where I play violin.

An orchestra is a good lesson in collaboration: dozens of people, one tempo, and everyone listening as hard as they play.

Suite 08Globe
Click · drag to spin7 places
On the map

New Haven to Bangalore, and back

College in New Haven, grad school in Philadelphia, work in Mountain View. I've flown to Bangalore to help onboard new Googlers and speak at the Geo EngProd summit.

Conferences took me to Vancouver for ICML 2025 and Orlando for Grace Hopper.

Suite 09Badge Wall
Click for the next badge8 badges
Recognition

A few things I'm proud of

Sweety Silver Award: Google-wide recognition for advancing UI regression detection with Gemini vision models. Geo Tech Impact Award: for the video-based Maps analysis framework.

Lead contributor on two defensive publications: UI anomaly detection with prompt engineering and regression detection with generative video models.

Suite 10Camera Roll
Click to advanceframe 01 / 07
Out loud

Explaining things in rooms full of people

I've spoken at Google's LLM Applications Conference and the Geo EngProd Bangalore Summit, and I co-wrote my old org's 2025 AI strategy.

Recent rolls: graduation at Yale, ICML in Vancouver, Grace Hopper, a Philharmonic concert, and the Presidio 10K.

Results

Summary: 11 / 11 suites passed

Click any row to jump back to that suite.
#SuiteFindingResult
01Eval HarnessEvals + post-training at Google DeepMind✓ running
02HillclimbGemini Omni V2V eval suite✓ shipped May 2026
03EvalSquaredEval creation: days → hours✓ adopted org-wide
04Map Tile QA10.5× more regressions caught✓ 25M daily users
05BookshelfYale B.S. CS & Econ · Penn M.S. AI✓ 2 degrees
06Presidio 10K10.0 km✓ finished
07MetronomeSouth Bay Philharmonic · violin✓ in tempo
08GlobeNew Haven → Bangalore → Vancouver✓ 7 places
09Badge Wall2 awards · 2 defensive pubs✓ 8 badges
10Camera RollTalks, conferences, finish lines✓ 7 frames
11Mailboxm3gha.joshi@gmail.com✓ open
Suite 11Mailbox
Click to raise the flagflag down
Say hi

Let's talk evals

I love talking about measuring models: what makes a benchmark hard to game, how to read an auto-rater, and when a number really means "ship it."

Email m3gha.joshi@gmail.com, grab 30 minutes, or find me on LinkedIn. For my résumé, just ask by email.