Senior Director, Google AI Models

Why This Matters: An impressive demo doesn’t prove your AI works. Reliable evaluation requires representative datasets, precise rubrics, calibrated judges, and failure analysis—so your team can detect regressions and make defensible release decisions.
Why Learn From Reah: Reah Miyara brings experience as Senior Director, Google Models; Product Lead for RL & Model Evaluation at OpenAI; and product leadership at Aporia and Arize AI—spanning frontier models, evaluation, observability, and guardrails.
What You’ll Build: Across 6 modules and 7.5 hours of live instruction, you’ll build an end-to-end evaluation system for a real AI feature: a representative dataset, scoring rubrics, calibrated judge, baseline results, failure diagnosis, and release gate.
Beyond Basic Benchmarks: Learn to resolve reviewer disagreement, track judge bias and drift, and uncover critical failures hidden behind average scores.
Built for Every Release: Establish dashboards, ownership, monitoring, and thresholds for launch, rollback, and human review... making evaluation part of your team’s operating cadence.
Master the complete AI evaluation lifecycle: from defining quality to diagnosing failures and making release decisions your team can trust.
Define observable criteria for usefulness, correctness, safety, tone, and task success.
Map failure severity across user segments and critical workflows.
Establish decision thresholds for deployment, rollback, and escalation to human review.
Design sampling strategies spanning common cases, long-tail behavior, and adversarial inputs.
Build golden sets grounded in trusted reference judgments.
Evolve datasets with production evidence while avoiding overfitting to easy examples.
Design annotation and adjudication protocols that resolve conflicting interpretations.
Translate subjective quality dimensions into explicit scoring anchors.
Measure reviewer agreement and refine criteria where ambiguity weakens the signal.
Develop reference-aware LLM judge prompts for dimension-specific scoring.
Calibrate automated judgments against expert human annotations.
Assess evaluator bias, variance, agreement, and drift to identify unreliable scoring.
Disaggregate results by user, task, language, risk, and behavior.
Identify recurring failure clusters concealed by aggregate performance.
Investigate regression causes across prompts, retrieval, models, tools, and UX.
Integrate offline evaluation with online monitoring.
Define dashboards, release gates, and accountability for quality decisions.
Establish an evaluation cadence that turns failure diagnosis into prioritized product improvements.

Head of Gemini, NanoBanana, Veo3, Live, Lyria, AlphaGenome, and all other models
For PMs & Leaders: Define AI quality, set release thresholds, and turn evaluation results into confident product decisions.
For Engineers & Tech Leads: Build rigorous eval systems, calibrate automated judges, and catch regressions before they reach users.
For AI Operators & Builders: Measure the quality of your prompts, agents, and workflows.

Live sessions
Learn directly from Reah Miyara in a real-time, interactive format.
Lifetime access
Go back to course content and recordings whenever you need to.
Community of peers
Stay accountable and share insights with like-minded professionals.
Certificate of completion
Share your new skills with your employer or on LinkedIn.
Maven Guarantee
Your purchase is backed by the Maven Guarantee.
Live sessions
2 hrs / week
Projects
1 hr / week
Async content
1 hr / week
Maven for Teams
Reimbursement
Get your company to pay
Everything L&D needs: email template, receipts, and certificate of completion.
Get reimbursedTeam discount
Learn with your teammates
Save 20%+ when 2 or more teammates enroll in the same cohort.
Save 20%+ with a teamPrivate cohort
Run a cohort for your org
A dedicated cohort with a custom schedule and curriculum, tailored to your team.
Book a private cohort$2,750
USD