Skip to content

Topic

Interpretability-Based Monitoring

Oversight that reads a model's internal states rather than only its outputs, on the assumption those internal signals stay meaningful as the model is trained further.

Current clusters