An internal playbook for building LLM judges already existed, but it was scrappy, incomplete, and outdated — not something teams could actually follow. As more teams started building judges to evaluate their own AI products, each one risked rediscovering the same steps from scratch, and it wasn't clear who should own the work: content, product management, engineering, or data science all had a stake.
I spotted the gap early and proactively overhauled the existing playbook into a clear, efficient, end-to-end resource: identifying common response issues, defining quality principles, building evaluation rubrics, training labelers, and establishing a golden dataset. As part of that overhaul, I led the discussions that resolved real ownership ambiguity across content, product management, engineering, and data science — winning content a clear leadership role in judge-building.
If we didn't fix the process, every new team would hit the same walls. Overhauling the playbook gave our team — and every team after — a solid foundation to build upon to quickly ensure a high-quality customer experience.
| Judge-building phase | Content | Product mgmt | Engineering | Data science |
|---|---|---|---|---|
| Identify common response issues | R | C | C | C |
| Define quality principles | R | C | — | — |
| Build evaluation rubric | R | — | — | C |
| Train labelers & label responses | R | — | — | C |
| Establish golden dataset | C | — | — | R |
| Build & validate judge | C | — | C | R |
| Launch & monitor | C | C | R | — |
R Owner C Contributor
Representative structure recreated for this portfolio — not a screenshot of the internal playbook.