Generative AI consultant and trainer.

Jev arrived on OpenRouter on September 18, 2026. It is not a chat model. It takes a JSON state and typed questions and returns calibrated probabilities in about 100 milliseconds, for $0.042 per million input tokens with output free. TypeSafe claims it is 193x faster and 445x cheaper than frontier models on routing, scoring, and verification. Nobody outside TypeSafe has reproduced those numbers.
Test both models on the same examples, compare their errors, latency, and cost, and decide when switching makes sense. You’ll learn how to recognize a meaningful improvement and when the evidence is too weak to justify a change.
In four hours you will learn Jev from the first call to production thresholds, build a paired bakeoff harness, and run it against Claude Opus 5.5 on routing, scoring, judging, and agent gating. Every result carries a confidence interval, a latency, and a cost per million decisions. You leave with the harness, the numbers, and an eleven-question decision guide that names the right model for any task you bring.
Master Jev end to end, the pros and the cons.
Recognize classification, routing, and scoring tasks that are candidates for Jev.
Separate decisions that need structured answers from tasks that need generated text.
Define the quality requirements a cheaper model must meet.
Use Choice, Score, and Noul to turn application data into structured decisions
Ask multiple questions in one request and inspect the returned probabilities.
Adapt runnable notebooks to your application without needing a GPU.
Run both models against a shared dataset with consistent scoring.
Compare errors, latency, and cost in a reusable evaluation harness.
Inspect disagreement cases to understand where each model succeeds or fails.
Calculate confidence intervals around your evaluation metrics.
Use paired comparisons to assess whether an apparent improvement is credible.
Recognize when the evidence is too weak to justify switching models.
Use Jev to check Claude outputs against clear, task-specific criteria.
Tune thresholds to balance missed problems against unnecessary escalations.
Measure the combined workflow’s quality, latency, and cost on held-out examples.
Use a decision checklist to recommend Jev, Claude, or a combination for a defined task.
Adapt the evaluation template to your own routing, scoring, or agent-gating problem.
Pin model versions and define when to rerun evaluations as your application changes.

PhD physicist, AI architect, author, and corporate trainer.


AI engineers and developers looking to reduce the cost and latency of classification, routing, and verification workflows.
Technical leads who need evidence to choose models and justify tradeoffs between cost, quality, and reliability.
Data scientists comfortable with Python who want practical methods to evaluate models and build reliable LLM applications.

Live sessions
Learn directly from Bruno Gonçalves in a real-time, interactive format.
Four hours of guided, hands-on Python work
Work through practical examples with instructor guidance. Ask questions as you explore model behavior, interpret results, and make implementation decisions.
Runnable notebooks you can reuse
Leave with the full repository and notebooks you can adapt to your own application. Reuse the code for model calls, evaluation, and reporting after the workshop.
Compare Jev and Claude on identical examples
Run both models against the same dataset with consistent scoring. Inspect their disagreements to understand which errors matter for your application
Understand the cost of your actual workload
Measure cost per decision and latency alongside task performance\. Assess whether potential savings justify a change and whether the cheaper option meets your quality requirements.
Statistical rigor, explained through code
Build confidence intervals and paired comparisons in Python. Understand what your results support, where uncertainty remains, and when you need more data.
Combine Claude’s generation with Jev’s verification
Explore a workflow in which Claude produces an answer and Jev checks it against defined criteria. Evaluate the combined system’s quality, latency, and cost.
A practical checklist for choosing your model
Leave with a decision guide for evaluating Jev, Claude, or a combination\. Connect your recommendation to a specific task, quality target, and operating budget.
Put agent decision gates to the test
Evaluate when an agent action should proceed, be blocked, or require review. Measure missed unsafe actions and unnecessary blocks, and understand the limits of your test results.
A straightforward setup with no GPU required
Use a laptop, Python, and hosted APIs. No prior Jev experience or statistics background is required; you should be comfortable reading Python and making API calls.
Lifetime access to recordings and materials
Revisit the explanations and rerun the labs at your own pace. Keep access to the workshop recordings, repository, and course materials as you apply the methods to new problems.
Maven Guarantee
Your purchase is backed by the Maven Guarantee.
Live sessions
4 hrs
Projects
1 hr
Maven for Teams
Reimbursement
Get your company to pay
Everything L&D needs: email template, receipts, and certificate of completion.
Get reimbursedTeam discount
Learn with your teammates
Save 20%+ when 2 or more teammates enroll in the same cohort.
Save 20%+ with a teamPrivate cohort
Run a cohort for your org
A dedicated cohort with a custom schedule and curriculum, tailored to your team.
Book a private cohort$300
USD