Conversational observability: evaluating and improving your GenAI chatbot

Maaike Groenewege, Convocat BV

Conversation designer, linguist, GenAI

What's the point of having an AI chatbot if you don't know how it's performing?

Note: you need your own Claude Code Pro subscription for this course!

So you've built your AI chatbot, fed it your website or knowledge base, created a beautiful system prompt or two, and put it live on your website. Only to find out you have no idea how it's doing.

Back in the olden days of NLU-based bots, life was pretty straightforward: you measured your F1 scores, walked through your convo logs, and optimised your model and flow accordingly.

Today, with chatbots generating answers of their own accord, no fixed flows, and no basic conversation logs anymore, it's a little bit harder. But not impossible.

In this course, I'll show you how to set up a basic RAG bot in Claude Code, connect it to Langfuse for observability, and work with the conversational observability lifecycle: collect, analyse, measure, experiment, and improve.

And as a bonus, you'll build a custom dashboard in Claude Code that turns technical Langfuse data into simple views your product owners and content team can actually use

What you’ll learn

You'll confidently design and implement the conversational observability lifecycle: collect, analyse, measure, experiment, improve.

  • Learn how to set up a Claude Code project in Claude Desktop

  • Keep track of your work with session backlogs and handover skills

  • Connect Claude Code to Langfuse

  • Learn a new way to evaluate GenAI-chatbots, building on what you already know as a conversation designer

  • Discover evaluation as an ongoing cycle, so your process keeps working as the bot and its answers evolve.

  • Master a step-by-step method (collect, analyse, measure, experiment, improve) that makes your evaluation repeatable and easy to explain.

  • Set up logging to capture the conversational data that you need for insights.

  • Trace analysis and labeling: learn what to track in a generative chatbot

  • Apply the transcript-reading and pattern-spotting skills you already use in flow-based design.

  • Learn failure modes and metrics that are unique to generative bots, like hallucinations, over-confidence, and off-brand tone of voice.

  • Turn the patterns you spot in Analyse into repeatable, automated checks you can run anytime.

  • Create experiments to compare prompt variants against the same test conversations, so you know what works before it goes live.

  • Turn your outcomes into concrete prompt or retrieval changes, backed by the numbers.

  • Turn overwhelming Langfuse data into simple, custom views that product owners and content people can actually use.

  • Use Claude Code to build a custom dashboard on top of Langfuse and your chatbot, no analyst skills required.

  • Give your whole team a shared, at a glance view of how the chatbot is performing, so insights reach beyond the data team.

Learn directly from Maaike

Maaike Groenewege, Convocat BV

Maaike Groenewege, Convocat BV

Conversation designer, linguist, fascinated by AI. Turning logs into insight.

Etsy
ABN AMRO
Independer
OneReach.ai
See all products from Convocat

Who this course is for

  • Conversation designers ready to grow their GenAI bot and career. Basic RAG knowledge and a Claude Pro subscription required.

  • Data analysts ready to learn the language-focused metrics crucial to evaluating GenAI chatbots. Claude Pro subscription required.

  • Product owners ready to oversee GenAI chatbot quality with confidence. Comfortable with chatbots and Claude Pro required.

What's included

Maaike Groenewege, Convocat BV

Live sessions

Learn directly from Maaike Groenewege, Convocat BV in a real-time, interactive format.

Lifetime access

Go back to course content and recordings whenever you need to.

Community of peers

Stay accountable and share insights with like-minded professionals.

Certificate of completion

Share your new skills with your employer or on LinkedIn.

Maven Guarantee

Your purchase is backed by the Maven Guarantee.

Course syllabus

Week 1

Aug 1—Aug 2

    Session 1 - Kick off: Why observability, and building your RAG chatbot -part 1

    4 items

    Session 2 - Building a basic RAG chatbot with Claude Code - part 2

    3 items

Week 2

Aug 3—Aug 9

    Session 3 - The conversational observability lifecycle

    3 items

    Session 4 - Analyse: reading conversations at scale

    7 items

Schedule

Live sessions

4 hrs / week

Two 2-hour sessions for 3 weeks

Frequently asked questions

Maven for Teams

Reimbursement

Get your company to pay

Everything L&D needs: email template, receipts, and certificate of completion.

Get reimbursed

Private cohort

Run a cohort for your org

A dedicated cohort with a custom schedule and curriculum, tailored to your team.

Book a private cohort