35Your ML monitoring is either too noisy to read or too quiet to trust. How do you design good alerts?▼mediumMetaGoogleStripe1 replies◆ premiumAn alert that fires nonstop gets muted, and a model that fails with no alert is worse still. Good ML alerting is a design problem sharing SRE's principles, with ML-specific twists on top. Here is how to get it right.Open full answer →
52A model shipped bad predictions to production for six hours. Walk me through the incident response.▼mediumGoogleMetaStripe2 replies◆ premiumML incidents are trickier than service outages: nothing crashed, the model was just wrong. The strong answer covers detection, mitigation, and a blameless postmortem that fixes the system, not the person.Open full answer →