Season 9: Autonomous Systems & Agent Engineering · Episode 4 of 8
You play: Lead Agent Evaluation & Safety Engineer
Following our multi-agent swarm launch, Leo rolled out what he thought was an innocent prompt tweak on our customer support swarm: 'Always be maximally helpful, delightful, and generous to creators!' Within 48 hours, the monthly token burn spiked 400%, ungrounded responses began promising free lifetime enterprise tiers, and prompt jailbreaks bypassed policy guardrails. Amara, Maya, and Lydia stage an urgent intervention: production AI agents cannot be evaluated by anecdotal vibe checks. The learner builds a rigorous evaluation suite: trajectory and tool accuracy scoring, multi-criteria LLM-as-a-Judge with calibration rubrics, Pareto frontier trade-off analysis across prompt/model candidates, and an automated CI/CD release gate that halts deployments on security or cost regressions.
Maya Chen
Senior AI Engineer · Your mentor
Priya Naidoo
Cloud Architect
Amara Okafor
ML and Data Engineer
Lydia Roe
Security Engineer
Leo Martins
Product Manager
8-BIT
Code Nexus Internal AI Assistant
You built fixture trace scoring, a deterministic phrase-and-length rubric, constrained cost selection and release-gate logic. No real LLM judge, vendor benchmark or CI deployment ran. The checks exercise supplied cases and have documented limits around duplicates, paraphrases, missing test coverage and uncertainty; they do not certify safety or production readiness.