“Why Should I Trust You?”: Explaining the Predictions of Any Classifier (2016)
Open in webOpen in zoteroOpen pdf
1 Abstract
Despite widespread adoption, machine learning models remain mostly black boxes. Understanding the reasons behind predictions is, however, quite important in assessing trust, which is fundamental if one plans to take action based on a prediction, or when choosing whether to deploy a new model. Such understanding also provides insights into the model, which can be used to transform an untrustworthy model or prediction into a trustworthy one. In this work, we propose LIME, a novel explanation technique that explains the predictions of any classifier in an interpretable and faithful manner, by learning an interpretable model locally around the prediction. We also propose a method to explain models by presenting representative individual predictions and their explanations in a non-redundant way, framing the task as a submodular optimization problem. We demonstrate the flexibility of these methods by explaining different models for text (e.g. random forests) and image classification (e.g. neural networks). We show the utility of explanations via novel experiments, both simulated and with human subjects, on various scenarios that require trust: deciding if one should trust a prediction, choosing between models, improving an untrustworthy classifier, and identifying why a classifier should not be trusted.
2 NOTES
This paper proposes the LIME method for explanation AI, which is pretty simple overall, but has some details that are confusing, especially due to the examples on the paper. Also, I don’t think the paper helps understanding the method, at all, especially for images. The main ideia in here starts with having a trained model on the dataset. After this, another model, a much simpler one, whose results are explainable is used. For this second one, the inputs are gonna be a single original sample from the dataset (whose understanding is desired) and multiple samples generated by varying the values of the original one. Therefore the models will have to worry about a local region of the dataset, referring only to the portion around that single original sample. By doing so the problem reduces from a much higher complexity to a simple one (hopefully). With this it is possible to investigate which parameter are important. For example, if trying to determine a subject probability to have a stroke, the parameters are gonna be something like age, weight, if is active, etc., with the model each is gonna have an value corresponding to its impact on the final decision. However, for EEG the only application I found is if separating the signal into multiple frequencies and them using LIME to see which are important. Other uses are harder to think, I don’t know exactly how to apply to covariance matrices. Maybe LIME can be adapted to the Riemannian manifold. Overall, it is a nice method, but is over 10 years old, so I bet there are much betters ones out there.