How this lab works
Operate, in plain terms
Day two begins the moment you ship: the dashboards stay green while the answers quietly go stale. This lab plays out one authored incident, twelve weeks of signals, a silent-drift emergency, and the retrain / reindex / rollback / rescope call that loops the program back to Frame.
Read the four signal families
System SLOs, the model-quality canary, RAG freshness and staleness, and agent and cost signals on one 12 week time axis.
Infra health and answer quality are different things, and only one of them is on the ops dashboard by default.
Spot the silent drift
Around week 5 the canary pass rate starts decaying below the Build baseline while availability and p95 stay green.
This is the trap the stage exists to teach: SLOs tell you the system is up, canary evals tell you it is still right.
Work the week 7 incident
An index staleness incident lands, with blast radius, value at risk, and a projected breach of the quality floor.
Value at risk turns a quality metric into money, which is what gets a remediation funded.
Make the call
Choose reindex, retrain, rollback, or rescope, each with its cost, time to effect, and what it does (and does not) fix.
The decision is the deliverable: the right fix depends on which signal actually moved, not on which button is nearest.
Loop it back
Your decision becomes a typed feedback contract routed upstream to Frame, Build, Deploy, Realize, and Govern.
This is what makes the lifecycle a loop instead of a line, day two findings become the next cycle's framing.
Take the artifacts
Download the weekly ops review and the incident report, generated from the exact state you just produced.
Ops evidence you can hand to a steering meeting beats a screenshot of a green dashboard.