Data & AI Architect | Author | Educator

Enroll by Sep 18th with code ERLYBRD and save $100.
As AI applications become more capable, evaluating them has become a major engineering challenge. Outputs are often open-ended, multiple answers may be valid, quality can depend on subjective or context-dependent criteria, and human evaluation is expensive and difficult to scale.
To address these challenges, many teams are turning to LLM-as-a-Judge systems to automate evaluation. However, LLM judges are not inherently objective or reliable, as their judgments can reflect biases of the underlying models and be highly sensitive to prompt wording, response ordering, evaluation criteria, model choice, and other design decisions. Without careful design and validation, they can introduce bias and inconsistency into evaluation pipelines, leading to misplaced confidence and poor decisions.
In this hands-on workshop, you'll learn a systematic engineering approach to designing, evaluating, and improving LLM judges. Through guided exercises, you'll build and validate LLM judges for different tasks and scenarios, uncover common sources of bias and instability, measure their impact, and apply practical techniques to improve reliability.
Move from ad hoc LLM judging to engineering reliable evaluation systems you can confidently use in production.
Design clear rubrics by breaking evaluation goals into observable criteria and minimizing ambiguity.
Write effective judgment prompts that translate rubrics into clear instructions and consistent outputs.
Choose the optimal LLM model based on the evaluation task, requirements, and model capabilities.
Measure agreement, consistency, and variability across judgments.
Identify and measure common sources of bias, variance, and instability.
Stress-test judges for sensitivity to prompts, response ordering, models, and other design choices.
Refine prompts, rubrics, models, and configurations based on evaluation evidence.
Combine multiple LLM judges through ensembles, aggregation, or debate to improve reliability.
Combine LLM judges with human evaluation and other methods to address their limitations.
Explore how LLM judges work, compare common judging approaches and use cases, and determine when LLM-as-a-Judge is the right evaluation method.
Turn evaluation goals into clear criteria and rubrics, choose appropriate models, write effective judgment prompts, and build judges for different evaluation tasks.
Validate LLM judges againstt human judgments, measure consistency, and variability. Identify biases and failure modes, and stress-test for sensitivity to different design choices.
Use evaluation evidence to improve judges, experiment with multiple-judge strategies, and combine LLM judges with human and other evaluation methods.

Data and AI architect, author, and educator
AI/ML Engineers & Data Scientists who build LLM applications and want to scale evaluation while ensuring reliable results.
AI Product Managers who need reliable evaluation to measure AI quality, compare alternatives, and make confident product decisions.
Technical Leads & AI Architects responsible for evaluation strategy and quality across LLM-powered systems and applications.
Lifetime access
Go back to course content and recordings whenever you need to.
Community of peers
Stay accountable and share insights with like-minded professionals.
Certificate of completion
Share your new skills with your employer or on LinkedIn.
Complimentary copy of "Evaluating AI Systems" book
Receive a complimentary MEAP copy of my forthcoming Manning book, Evaluating AI Systems, for a broader, systematic approach to evaluating modern AI systems.
Maven Guarantee
Your purchase is backed by the Maven Guarantee.
Maven for Teams
Reimbursement
Get your company to pay
Everything L&D needs: email template, receipts, and certificate of completion.
Get reimbursedPrivate cohort
Run a cohort for your org
A dedicated cohort with a custom schedule and curriculum, tailored to your team.
Book a private cohortGet course updates