The Mythos of Model Interpretability.¶
Lipton, Z. C. (2018). The Mythos of Model Interpretability. Communications of the ACM, 61(10), 36-43.
Cited by¶
3 citations across 2 artifacts.
Each citation links to the sentence it supports in the citing article.
Primes¶
- Black Box vs. White Box Distinction
- Interpretability research (Lipton 2018, "The Mythos of Model Interpretability"
This sourceDecomposes the under-specified notion of interpretability into transparency (simulatability, decomposability, algorithmic transparency) and post-hoc explanation, arguing it is not a single binary property.
- Modern ML directly illustrates this (Lipton 2018
This sourceArgues the interpretability/accuracy trade-off is fundamental: nonlinear deep models gain accuracy at the cost of human-inspectable structure.
- Interpretability research (Lipton 2018, "The Mythos of Model Interpretability"
- Transparency
- Non-formal, structurally faithful: A large AI company, facing growing external pressure for AI transparency — pressure shaped by Lipton's (2018) influential critique of the "mythos of model interpretability," which decomposed interpretability into transparency (simulatability, decomposability, algorithmic-transparency) and post-hoc explanations (visualization, examples, text rationales) — and internal concerns about reputation for opacity, designs a comprehensive transparency framework: (a) model cards — for each deployed model, published documentation covering training data characteristics, evaluation results across benchmarks, known limitations, intended-use descriptions, and safety-evaluation findings; (b) system cards — product-level documentation of how models are integrated, user-facing controls, content-policy enforcement, and operational safeguards; © transparency reports — quarterly publication of content-policy enforcement actions, law-enforcement requests, safety-incident counts, model-update summaries; (d) researcher access — credentialed research program giving external researchers access to models and evaluation infrastructure under vetted-access terms; (e) red-team reports — published summaries of major red-teaming engagements with findings and remediation, redacted only where disclosure would increase harm; (f) integrity guarantees — external auditor review of transparency-report accuracy, cryptographically-signed commit histories for training-data documentation, immutable logging of enforcement-action metadata.
This sourceDecomposes the under-specified concept of model interpretability into transparency (simulatability, decomposability, algorithmic transparency) and post-hoc explanation (visualization, examples, text rationales); influential conceptual framework for AI-system transparency.
- Non-formal, structurally faithful: A large AI company, facing growing external pressure for AI transparency — pressure shaped by Lipton's (2018) influential critique of the "mythos of model interpretability," which decomposed interpretability into transparency (simulatability, decomposability, algorithmic-transparency) and post-hoc explanations (visualization, examples, text rationales) — and internal concerns about reputation for opacity, designs a comprehensive transparency framework: (a) model cards — for each deployed model, published documentation covering training data characteristics, evaluation results across benchmarks, known limitations, intended-use descriptions, and safety-evaluation findings; (b) system cards — product-level documentation of how models are integrated, user-facing controls, content-policy enforcement, and operational safeguards; © transparency reports — quarterly publication of content-policy enforcement actions, law-enforcement requests, safety-incident counts, model-update summaries; (d) researcher access — credentialed research program giving external researchers access to models and evaluation infrastructure under vetted-access terms; (e) red-team reports — published summaries of major red-teaming engagements with findings and remediation, redacted only where disclosure would increase harm; (f) integrity guarantees — external auditor review of transparency-report accuracy, cryptographically-signed commit histories for training-data documentation, immutable logging of enforcement-action metadata.
Verification¶
This reference passed the adversarial substantiation pipeline: it was checked to exist and to support the claim it is attached to. See how references were verified.
Registry ID ref:9fd1370f8e98 · see in the full table