Skip to content

Build Evals with Mizan — series index

The Build Evals with Mizan series takes you from the judgment you already make by eye to writing, sharing, and extending evals with Mizan. The blog index lists every post newest-first; the list below is the reading order — follow it top to bottom to walk the arc.

  1. Part 1Why evals, and what is an LLM-as-a-judge, really?

    You already judge generative output by eye. An eval turns that private judgment into something explicit, repeatable, and shareable.

  2. Part 2One spectrum, four users: from brand alignment to technical metrics

    Four Mizan users sit on one spectrum, from "is this on-brand?" to "is this grounded and accurate?" The same machinery answers both. Only the criteria change.

  3. Part 3.1Using Mizan in your role: per-persona playbooks, part 1 (Asset creator, Asset manager)

    Two roles, two playbooks. How an asset creator finds the shortest path from one eval to a scorecard, and how an asset manager curates a set that reports each concern on its own.

  4. Part 3.2Using Mizan in your role: per-persona playbooks, part 2 (Genmedia configurator, Brand Lab)

    The second pair of playbooks: the genmedia configurator embedding an eval-set as a calling shape, and the Brand Lab user generating a rubric from a brand book and freezing it into a reusable metric.

  5. Part 4Contribute a template pack: extend Mizan without touching the core

    You do not have to fork Mizan to add value to it. This piece walks authoring, validating, and sharing a template pack that others can import.

  6. Part 4.5Turn a writing skill into an eval: grading this series with Mizan

    The capstone closes the loop: take the editorial standard behind these posts and encode it as a Mizan rubric that grades the series itself.

  7. Part 5Developing for Mizan: extending the core

    For contributors who do need to touch the core: an orientation to Mizan's architecture and where new metric kinds and behaviors plug in.

Prefer a feed reader? Subscribe to the RSS feed.