Tartalmi kivonat
UNCOVERING AND INDUCING INTERPRETABLE CAUSAL STRUCTURE IN DEEP LEARNING MODELS A DISSERTATION SUBMITTED TO THE DEPARTMENT OF LINGUISTICS AND THE COMMITTEE ON GRADUATE STUDIES OF STANFORD UNIVERSITY IN PARTIAL FULFILLMENT OF THE REQUIREMENTS FOR THE DEGREE OF DOCTOR OF PHILOSOPHY Atticus Reed Geiger November 2023 2023 by Atticus Geiger. All Rights Reserved Re-distributed by Stanford University under license with the author. This work is licensed under a Creative Commons Attribution3.0 United States License http://creativecommons.org/licenses/by/30/us/ This dissertation is online at: https://purl.stanfordedu/th321qf7186 ii I certify that I have read this dissertation and that, in my opinion, it is fully adequate in scope and quality as a dissertation for the degree of Doctor of Philosophy. Christopher Potts, Primary Adviser I certify that I have read this dissertation and that, in my opinion, it is fully adequate in scope and quality as a dissertation for the degree of Doctor
of Philosophy. Thomas Icard, III, Co-Adviser I certify that I have read this dissertation and that, in my opinion, it is fully adequate in scope and quality as a dissertation for the degree of Doctor of Philosophy. Michael Frank I certify that I have read this dissertation and that, in my opinion, it is fully adequate in scope and quality as a dissertation for the degree of Doctor of Philosophy. Noah Goodman Approved for the Stanford University Committee on Graduate Studies. Stacey F. Bent, Vice Provost for Graduate Education This signature page was generated electronically upon submission of this dissertation in electronic format. iii Abstract A faithful and interpretable explanation of an AI model’s behavior and internal structure is a high-level explanation that is human-intelligible but also consistent with the known, but often opaque low-level causal details of the model. We argue that the theory of causal abstraction provides the mathematical foundations for the desired
kinds of model explanations. In the analysis mode, we uncover causal structure using interventions on model-internal states to assess whether an interpretable high-level causal model is a faithful description of a deep learning model. In the training mode, we induce interpretable causal structure using interventions during model training to simulate counterfactuals in the deep learning model’s activation space. We show how to uncover and induce causal structures in a variety of case studies on deep learning models that reason over language and/or images. iv Preface This thesis is largely an assembly of existing publications that form a coherent research program centered around uncovering and inducing interpretable causal structure in deep learning models. 1. The first chapter situates the field interpretability in a larger context and is content original to this thesis. 2. The second chapter lays out the general theoretical framework for interpretability that is grounded in causal
abstraction. This chapter is the publication: A. Geiger, C Potts, and T Icard Causal abstraction for faithful model interpretation Ms, Stanford University, 2023a. URL https://arxivorg/abs/230104709 3. Chapters 3-6 contain experimental work where interpretable causal structure is uncovered in deep learning models. These chapters are the publications: A. Geiger, K Richardson, and C Potts Neural natural language inference models partially embed theories of lexical entailment and negation. In Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pages 163–173, Online, Nov. 2020b. Association for Computational Linguistics doi: 1018653/v1/2020blackboxnlp-116 URL https://www.aclweborg/anthology/2020blackboxnlp-116 A. Geiger, H Lu, T Icard, and C Potts Causal abstractions of neural networks In Advances in Neural Information Processing Systems, volume 34, pages 9574–9586, 2021b. URL https://
papers.nipscc/paper/2021/hash/4f5c422f4d49a5a807eda27434231040-Abstracthtml Z. Wu, A Geiger, C Potts, and N D Goodman Interpretability at scale: Identifying causal mechanisms in Alpaca. Ms, Stanford University, 2023b URL https://arxivorg/abs/2305 08809 A. Geiger, Z Wu, C Potts, T Icard, and N D Goodman Finding alignments between interpretable causal variables and distributed neural representations. Ms, Stanford University, 2023d. URL https://arxivorg/abs/230302536 4. Chapters 7-10 contain experimental work where interpretable causal structure is induced in deep learning models. These chapters are the publications: v A. Geiger, Z Wu, H Lu, J Rozner, E Kreiss, T Icard, N Goodman, and C Potts Inducing causal structure for interpretable neural networks. In K Chaudhuri, S Jegelka, L Song, C. Szepesvari, G Niu, and S Sabato, editors, Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pages 7324–7338. PMLR, 17–23
Jul 2022c. URL https://proceedingsmlrpress/v162/geiger22ahtml Z. Wu, A Geiger, J Rozner, E Kreiss, H Lu, T Icard, C Potts, and N D Goodman Causal distillation for language models. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4288–4295, Seattle, United States, July 2022e. Association for Computational Linguistics doi: 10.18653/v1/2022naacl-main318 URL https://aclanthologyorg/2022naacl-main318 Z. Wu, K D’Oosterlinck, A Geiger, A Zur, and C Potts Causal proxy models for concept-based model explanations. In International Conference on Machine Learning, pages 37313–37334 PMLR, 2023a A. Zur, E Kreiss, and C P A Geiger When interpretability enhances accessibility: Updating clip to prefer descriptions over captions. 2023 vi Acknowledgments I stand on the shoulders of the giants who came before me. Most of these giants are long dead figures of history, but some have been mentors
and role models to me throughout my academic career. Words won’t properly express my gratitude for these individuals, but I’m happy to try. Christopher Potts has been essential to my success at Stanford by being a constant source of guidance and support. Despite being an undergraduate student when we first met, Chris saw my potential, encouraged my autonomy, and engaged with me as an equal collaborator from day one. Most importantly, every time I shared a personal hardship with Chris, his response was to ask how he could help. I aspire to Chris’s integrity and singular dedication to his students and his work The first class I took at Stanford in the fall of 2015 was taught by Thomas Icard, and his deeply interdisciplinary way of thinking that has shaped me ever since. When I began working closely with Thomas during the first year of my PhD, I realized that he was not just a brilliant thinker who could feed my hungry mind, but also a kindred spirit who shares some of my deepest
beliefs and values. I can’t imagine where I’d be without the countless delightful conversations we’ve had over the years. I connected with Noah Goodman later into my PhD career, which is a shame, because we hit it off so well! His contributions to my research were invaluable; I would leave conversations with more promising ideas than I knew what to do with. Also, we both wear flip flips to work, and that says something good about both of us. My research with Mike Frank doesn’t appear in this thesis, but our project was a formative experience very early into my research career. Mike’s humility, kindness, and devotion to good science inspire me. I hope we get the chance to work together again To all of my peers and coauthors over the years: I appreciate you all so much, and I would not be where I am today without you. Special shout out to Elisa Kreiss, Zen Wu, and Karel D’oosterlinck for being amazing collaborators, coworkers, and friends. I also need to thank the late, great
Lauri Karttunen for graciously introducing me to academic research. I look back fondly on the summer I spent drinking espresso, eating chocolate, and thinking about implicative verbs. Finally, thank you to all of my friends and family who supported me over the years, because I needed it. More than anyone else, thank you Sarah Hartman You are my sunshine, without you I would wilt. vii Contents Abstract iv Preface v Acknowledgments vii 1 Interpretability as a Subfield of Explainable AI 1 1.1 Setting the Scene . 2 1.2 Black Box AI . 4 1.21 Modular Features . 4 1.3 Transparent Algorithms 4 1.4 1.31 Simulatability . 4 1.32 Conceptual Variables . 5 1.33 Natural Language Texts as Transparent Algorithms . 5 Interpretability .
6 1.41 Faithfulness . 6 1.42 Verification and Generalization . 6 1.43 Behavioral and Mechanistic Interpretability . 6 1.5 Explanation . 7 1.6 Data Generating Process . 7 1.61 Black Box AIs as Scientific Models . 8 Human Decision Makers . 8 1.71 Accidents and AI Safety . 8 1.72 Algorithmic Fairness and Recourse . 9 1.73 Plausibility and Human Evaluations of Interpretations . 9 1.8 The Contributions of this Thesis 10 1.7 2 Causal Abstraction for Faithful Model Interpretation 2.1 Introduction . viii 11 11 2.2 Related Work . 13 2.21 Causal Abstraction .
13 2.22 Faithful and Interpretable Causal Explanations of AI . 14 2.23 Methods for Explaining AI Behavior . 15 2.24 Methods for Explaining the Internal Structure of AI . 16 2.3 Causal Models . 18 2.4 Example of Causal Models: A Symbolic Algorithm and Neural Network . 20 2.41 Hierarchical Equality Task . 20 2.42 A Tree-Structured Algorithm for Hierarchical Equality . 21 2.43 A Fully Connected Neural Network for Hierarchical Equality . 22 Causal Abstraction and Interchange Intervention Analysis . 22 2.51 Alignments Between Causal Models . 22 2.52 Causal Consistency and Constructive Abstraction . 25 2.53 Interchange Intervention Analysis . 26 2.54 Explanation and Generalization . 28 2.6
Decomposing Constructive Causal Abstraction . 29 2.7 Example of Causal Abstraction: Tree-Structure in Neural Computation . 32 2.71 An Alignment Between the Algorithm and the Neural Network . 32 2.72 The Algorithm Abstracts the Neural Network . 33 2.73 The Algorithm can be Constructed from the Neural Network . 33 2.8 Approximate Abstraction and Interchange Intervention Accuracy 33 2.9 XAI Methods Grounded in Causal Abstraction 35 2.5 2.91 LIME: Behavioral Fidelity as Approximate Abstraction by a Two-Variable Chain 35 2.92 Causal Effect Estimation as Abstraction by a Two-Variable Chain . 37 2.93 Causal Mediation As Abstraction by a Three-Variable Chain . 37 2.94 Iterative Nullspace Projection As Abstraction by a Three Variable Chain . 39 2.95 Operationalizing Circuit-Based Explanations with Causal Abstraction . 40 2.96 Interchange Interventions
from Integrated Gradients . 41 2.10 Future Applications: Types, Infinite Variables, and Cycles 42 2.11 Coda: Abstraction for Probabilistic Models 45 2.12 Conclusion 48 . 3 Neural Natural Language Inference Models Partially Embed Theories of Lexical Entailment and Negation 49 3.1 Introduction . 49 3.2 Related work . 51 3.3 Monotonicity NLI dataset . 52 3.4 Models . 53 ix 3.5 3.6 3.7 Behavioral Evaluations . 54 3.51 MoNLI as a Challenge Test Set . 54 3.52 A Systematic Generalization Task . 55 Structural Evaluations . 56 3.61 Probes . 58 3.62
Interventions . 59 Conclusion . 4 Causal Abstractions of Neural Networks 61 63 4.1 Introduction . 63 4.2 Related Work . 65 4.3 Causal Abstraction Analysis of Neural Networks . 67 4.4 The Natural Language Inference Task and Models 69 4.5 A Case Study in Structural Neural Network Analysis 70 4.6 4.51 Causal Abstractions of Neural NLI models . 70 4.52 Comparison with Other Structural Analysis Methods . 74 Conclusion . 75 5 Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations 79 5.1 Introduction . 79 5.2 Related Work . 81 5.3 Methods .
82 5.4 Hierarchical Equality Experiment . 89 5.5 Monotonicity NLI Experiment . 90 5.6 Conclusion 92 . 6 Interpretability at Scale: Identifying Causal Mechanisms in Alpaca 93 6.1 Introduction . 93 6.2 Related Work . 94 6.3 Methods . 96 6.31 Background on Causal Abstraction . 96 6.32 Boundless Distributed Alignment Search . 97 Experiment . 98 6.41 Price Tagging . 98 6.42 Hypothesized Causal Models . 99 6.43 Boundless DAS Results . 101 6.44 Interchange Interventions with (In-)Correct Inputs . 103 6.4 x 6.45 Do Alignments
Robustly Generalize to Unseen Instructions and Inputs? . 103 6.46 Boundary Learning Dynamics . 105 6.5 Analytic Strengths and Limitations 106 6.6 Conclusion . 106 7 Inducing Causal Structure for Interpretable Neural Networks 107 7.1 Introduction . 107 7.2 Related Work . 108 7.3 Interchange Intervention Training . 109 7.4 MNIST Pointer-Value Retrieval . 114 7.5 Navigation and Language (ReaSCAN) . 118 7.6 Natural Language Inference (MQNLI) . 120 7.7 Conclusion . 122 8 Causal Distillation for Language Models 123 8.1 Introduction . 123 8.2 Related Work . 124 8.3
Causal Distillation . 126 8.4 Experimental Set-up . 127 8.5 Results . 129 8.6 Conclusion . 131 9 Causal Proxy Models for Concept-based Model Explanations 132 9.1 Introduction . 132 9.2 Related Work . 133 9.3 Causal Proxy Model (CPM) . 135 9.4 Experiment Setup . 138 9.5 9.6 9.41 Causal Estimation-Based Benchmark (CEBaB) . 138 9.42 Evaluation Metrics . 139 9.43 Baseline Methods . 139 9.44 Causal Proxy Models . 140 Results . 141 9.51 CEBaB Performance . 142 9.52 Self-Explanation
with CPM . 144 9.53 Concept-Aware Feature Attribution with Causal Proxy Models . 144 Conclusion . 145 xi 10 When Interpretability Enhances Accessibility: Updating CLIP to Prefer Descriptions Over Captions 146 10.1 Introduction 146 10.2 Related Work 148 10.3 Methods 150 10.31 CLIP 150 10.32 Causal Models and Interventions 150 10.33 Contrastive Training Objectives 151 10.4 Fine-Tuning CLIP on Concadia 153 10.5 Transfer Learning Evaluations 154 10.6 BLV and Sighted Human Evaluations 155 10.7 Integrated Gradients Through the Description–Caption Representation 156 10.8 Conclusion .
158 xii List of Tables 3.1 The results of our behavioral analysis The columns labeled No MoNLI fine-tuning display the challenge test set results (Section 3.51), and the columns labeled With MoNLI fine-tuning display systematic generalization task results (Section 3.52) The numbers are accuracy values; all the datasets have balanced label distributions. Dashes mark experiments that would involve untrained NLI parameters due to training/finetuning set-up. 4.1 53 N Largest subsets of examples on which specific models CNatLog are abstractions of an LSTM and BERT model trained on MQNLI. We record the size of such subsets as a percentage of the total 1000 examples. On this subset, we know that the neural models compute a representation of the relation between the aligned subphrases under N and use this information to make a final prediction. [-125ex] 5.1 . 72 Hierarchical equality alignment learning results. The
table can be read as follows: Layer 1, Layer 2, and Layer 3 indicate which layer of neurons is targeted, ∣N∣ is the number of neurons in a layer, k is the number of neurons aligned with each intermediate variable (red) where our subspace model occupies k2 with rounding up to the closest integer, and the values in each cell are interchange intervention accuracies for the learned alignment on training data. We report the best results from three runs with distinct random seeds. 5.2 88 Monotonicity NLI results. The table can be read as follows: Layer 7, Layer 9, and Layer 11 indicate which layer of neurons is targeted, ∣N∣ is the number of neurons in a layer, k is the number of neurons aligned with each intermediate variable (red) where our subspace model occupies k2 , and the values in each cell are interchange intervention accuracies for the learned alignment on training data. We report the best results from three runs with distinct random
seeds. xiii 91 6.1 Summary results for all experiments with task performance as accuracy (range [0, 1]), maximal interchange intervention accuracy (IIA) (range [0, 1]) across all positions and layers, Pearson correlations of IIA between two distributions (compared to ♣ or ♥; range [−1, 1]), and variance of IIA within a single experiment across all positions and † layers. This is empirical task performance on the evaluation dataset for this experiment104 7.1 θ Results for NPVR (ResNet18) trained on the PVR-MNIST dataset. Behavioral accuracy θ is the percentage of inputs that NPVR agrees with CPVR on. Interchange intervention accuracy quantifies the extent to which the interpretable causal model is a proxy for the network (Section 7.3) IIT delivers the best results, especially when combined with multi-task objectives. 116 8.1 Performance on the development sets of the WikiText, GLUE benchmark, CoNLL2003 corpus for
the name-entity recognition task, and SQuAD v1.1 for the question answering task. The score is the averaged performance scores with standard deviation † (SD) for all tasks across 15 distinct runs. Numbers are imputed from released models on Hugging Face [Wolf et al., 2020] 128 9.1 CEBaB scores measured in three different metrics on the test set for four different model architectures as a five-class sentiment classification task. Lower is better Results averaged over three distinct seeds, standard deviations in parentheses. The metrics are described in Section 9.4 Best averaged result is bolded (including ties) per approximate counterfactual creation strategy. 140 9.2 Task performance measured as Macro-F1 score on the test set (average of 3 distinct seeds; standard deviations in parentheses). 141 9.3 Self-explanation CEBaB scores measured in three different metrics on the test set for four different model
architectures as a five-class sentiment classification task. Lower is better. Average of 3 distinct seeds; standard deviations in parentheses 142 9.4 Visualizations of word importance scores using Integrated Gradient (IG) by restricting gradient flow through the corresponding intervention site of the targeted concept. Our target class pools positive and very positive. Individual word importance is the sum of neuron-level importance scores for each input, normalized to [ −1 , +1 ]. −1 means the word contributes the most negatively to predicting the target class (red); +1 means the word contributes the most positively (green). 143 10.1 Transfer learning results for CLIP models fine-tuned on Concadia The error bounds are 95% confidence intervals from runs with 5 random seeds. The number of classes for each dataset is stated in parentheses. 152 xiv 10.2 The results of fine-tuning the CLIP model on the Concadia dataset The error bounds
are 95% confidence intervals from runs with five random seeds. 153 10.3 Correlation between model similarity scores and human preference, across imaginability, relevance, irrelevance, and overall value of text description Kreiss et al. [2022a] Scores ∗ reported with an asterisk ( ) are statistically significant with p < 0.10 154 10.4 Correlation between integrated gradient attributions and per-token human labels for concreteness and imageability. The error bounds are 95% confidence intervals from runs with five random seeds. 158 xv List of Figures 1.1 A visualization of interpretability as a subfield of explainable AI Interpretability is about faithfully simplifying a black box AI into a transparent algorithm, while the larger field of expainable AI also concerns the real world processes that generate the data that is input to a black box AIs and the needs of humans that make decisions using transparent algorithms. .
3 2.1 A tree-structured algorithm that perfectly solves the hierarchical equality task with a compositional solution. 21 2.2 A fully-connected feed-forward neural network that labels inputs for the hierarchical equality task. We define weights for a network that was created with interchange intervention training to implement the tree-structured solution to the task. 23 2.3 An alignment between the causal graphs of a low-level fully-connected neural network (bottom) and a high-level tree structured algorithm (top). . 24 2.4 The result of aligned interchange intervention on the low-level fully-connected neural network and a high-level tree structured algorithm under the alignment in Figure 2.3 Observe the equivalent counterfactual behavior across the two levels. 28 2.5 An illustration of a fully-connected neural network being transformed into a tree structured algorithm by (1) marginalizing away
neurons aligned with no high-level variable, (2) merging sets of variables aligned with high level variables, and (3) merging the continuous values of neural activity into the symbolic values of the algorithm. . 31 2.6 A causal model representing the bubble sort algorithm (top) and abstractions of that model (bottom). 43 3.1 An algorithm able to solve the MoNLI dataset that provides a theoretically motivated learning target for neural models at an algorithmic level of analysis [Marr, 1982]. Infer takes in an example from MoNLI and outputs the relation between the premise and hypothesis. It uses three predefined functions get-lex-rel returns the relation (one of {⊐, ⊏}) between the substituted words in the premise and hypothesis. contains-not returns true iff negation is present. reverse maps ⊏ to ⊐ and vice-versa xvi 56 3.2 Results where classifier probes are trained on BERT representations to predict the value of lexrel and
the output of Infer (Figure 3.1) Selectivity is probe accuracy minus control probe accuracy [Hewitt and Liang, 2019]. The grey dotted line provides a soft ceiling for selectivity values, because we expect control probes trained on a binary task to at least achieve chance accuracy. . 57 3.3 An illustrative interchange intervention: The solid arrows represent a hypothesis about where the model stores and uses information about lexical entailment. The dotted arrow is an interchange intervention, where the green vector (top) we think stores reverse entailment, trees ⊐ elms, is interchanged with the red vector (middle) we think stores forward entailment, pugs ⊏ dogs, leading to a modified network (bottom). If our hypothesis is correct, then the output should change from entailment to neutral, because the negation in the green example reverses the relationship between lexical entailment and sentence-level entailment. If this label reversal is not observed, crucial
entailment information must lie elsewhere in the network. 4.1 60 Our motivating example where we hypothesis that a symbolic computation C+ is a causal abstraction of a neural network N+ under a particular alignment (top). We can experimentally confirm this hypothesis by conducting an interchange intervention on both the network and the computation with every pair of inputs and evaluating whether the intervened network and intervened computation have the same counterfactual output behavior. We schematically depict an interchange intervention on the network N+ (bottom left) and the computation C+ (bottom right) with the base input (1, 2, 3) and the source input (4, 5, 6). Observe that the output of the intervened neural network matches the output of the intervened symbolic computation, so we have success for this pair of inputs. 76 4.2 The natural logic causal model (top), MQNLI examples (left) and MQNLI results (right).
77 NPObj 4.3 A BERT-based NLI model (left) aligned with the natural logic causal model CNatLog P (right), where the fourth vector representation above the AdjObj token in the network is aligned with NPObj , the variable representing the relation between the object noun phrases. When analyzing a sample of 1000 examples, we found a subset of 383 where NPObj CNatLog is an abstraction of NNLI under this alignment. 4.4 77 Interchange intervention and probing results for the NPObj position. Vertical axes denote layers of BERT and horizontal axes denote the token position of hidden representations. The intervention success rates reported here are calculated based on intervention experiments with a change in the output label. Clique sizes are reported as % of 1000 examples. . xvii 78 5.1 A generic multi-source distributed interchange intervention The base input and two source inputs create three total
settings of a model. The top left (green) and right (blue) total model settings are determined by two source inputs and the middle total model setting (red) is determined by the base input. Three hidden units from each total setting are rotated with an orthogonal matrix R ∶ X Y. Then we intervene on the rotated representation for the base input and fix two dimensions to be the value they take on for each source input, respectively. Then we unrotate the representation −1 with R and compute a counterfactual total model setting for the base input. In DAS, the orthogonal matrix is found with SGD using a high-level causal model to guide the search process. . 85 5.2 High-level models. 90 6.1 Our pipeline for scaling causal explainability to LLMs with billions of parameters. 95 6.2 Four proposed high-level causal models for how Alpaca solves the price tagging task. Intermediate variables are in red. All
these models perfectly solve the task 99 6.3 Aligned distributed interchange interventions performed on the Alpaca model that is instructed to solve our Price Tagging game. It aligns the boolean variable representing whether the input amount is higher than the lower bound in the causal model. To train Boundless DAS, we sample two training examples and then swap the intermediate boolean value between them to produce a counterfactual output using our causal model. In parallel, we swap the aligned dimensions of the neural representations in rotated space. Lastly, we update our rotation matrix such that our neural network has a more similar counterfactual behavior to the causal model. 100 6.4 Interchange Intervention Accuracy (IIA) for four different alignment proposals. The Alpaca model achieves 85% task accuracy. The higher the number is, the more faithful the alignment is. We color each cell by scaling IIA using the model’s task performance as the upper bound and a
dummy classifier (predicting the most frequent label) as the lower bound. These results indicate that the top two are highly accurate hypotheses about how Alpaca solves the task, whereas the bottom two are inaccurate in this sense. Analyzing tokens includes special tokens (e.g, ‘<0x0A>’ for linebreaks) required by Alpaca’s instruct-tuning template. 102 6.5 Learned boundary width for intervention site and in-training evaluation interchange intervention accuracy (IIA) for two groups of data: (1) aligned group where the boundary does not shrink to 0 at the end of the training; (2) unaligned group where the boundary does shrink to 0 at the end of the training. 1 on the y-axis means either 100% accuracy for IIA, or the variable is occupying half of the hidden representation for the boundary width. 105 xviii 7.1 . 113 7.2 An illustration of an IIT update where a
neural network (right) is trained to realize a causal model (left) that solves the PVR-MNIST task. Solid lines are feed-forward connections, dashed lines are interchange interventions, red lines are the flow of backpropagation. Observe that when backpropagation reaches the interchange intervention, it flows into both the source input’s computation graph and the base input’s graph, updating the weights below the interchange intervention twice. 114 7.3 . 117 7.4 Performance of a pretrained BERT natural language inference model fine-tuned on the QPObj MQNLI dataset with the causal model CNatLog from Geiger et al. [2020a] We report the results on the evaluation set. While data augmentation leads to consistently excellent behavior accuracy (left) panel, it has very low interchange intervention accuracy. In other words, IIT is necessary for an interpretable model with high-performance. 122 8.1 An IIT update in the context
of masked language modelling (MLM) The teacher network (top) has 6 layers and the student (bottom) has 3 layers, and we align layer 2 in the student with layers 3–4 in the teacher. Solid lines are feed-forward connections, red lines show the flow of backpropagation, and dashed lines indicate interchange interventions. In this case, the student originally predicted the token “salad” under the interchange intervention, while the teacher predicted the token “pizza” under an aligned interchange intervention. DIITO trains the student to minimize the divergence between the student logits and the teacher logits under the interchange intervention. This updates the student to conform to causal dynamics of the teacher. 125 8.2 Perplexity score distribution for the development set of WikiText of models trained in a low-resource setting. The best model is the one with the richest alignment structure130 8.3 GLUE score distribution across 15 distinct runs of students in different
sizes. Following the evaluation for BERT Devlin et al. [2019b] we exclude WNLI for evaluation 130 9.1 Causal Proxy Model (CPM) summary. Every CPM for model N is trained to mimic the factual behavior of N (LMimic ). For CPMIN , the counterfactual objective is LIN For CPMHI , the counterfactual objective is LHI . . 136 10.1 A visual depiction of a training update during distributed interchange intervention training. The core idea is that a description of an image in Concadia should be assigned a higher CLIPScore than the caption of the same image, so we induce a representation of the description–caption distinction that increases the CLIPScore when the text is a description. The final label (in the diagram, < or >) is determined by whether the intervention is caption description or vice versa. 147 xix 10.2 Integrated gradient attributions for the IIT-DAS model run on a sample image and its corresponding description and caption from
the Concadia dataset. A positive token attribution means that the token contributed positively to the outputted CLIPScore (green), and negative token attribution means that it contributed negatively (red). The overall score is the sum of the token attributions within the sentence. 157 xx Chapter 1 Interpretability as a Subfield of Explainable AI We take the fundamental question of explainable artificial intelligence (XAI) to be why an AI model makes the predictions it does; the gold standard for explaining model behavior and the internal mechanisms that underlie it should be a causal explanation [Woodward, 2003, Pearl, 2019a]. However, not just any causal explanation will be a satisfying answer. After all, the parameters of deep learning models give us perfect ground-truth knowledge of the causal relationships between all components, so low-level causal explanations of behavior and internal reasoning can be easily provided in terms of real-valued vectors, activation
functions, and weight tensors. Ironically, ‘black box’ models are in some sense easily explainable. The problem is, these explanations are not transparent to humansthey fail to yield insights about the high-level reasoning that underlies AI behavior [Lipton, 2018a, Creel, 2020]. Simple algorithms that operate on human-intelligible concepts are easy to construct and understand, but under what conditions could we trust that such a transparent algorithm is a faithful interpretation [Jacovi and Goldberg, 2020a] of the known, but opaque, low-level details of a black box AI? We take this to be the fundamental question of the XAI subfield known as interpretability. It is crucial for human decision makers that interpretability methods not tell ‘just-so’ stories that have nothing to do with how AIs actually behave and reason. To achieve this, we need a common language for explicating and comparing methodologies and precisely defining faithful interpretation. We believe the theory causal
abstraction will provide this common language. Modern deep learning models are akin to the weather or a brain in the following sense: they contain a large number of densely connected microvariables that have complex, non-linear dynamics. In order to meet the challenges of aggregating microvariables into macrovariables [Jonas and Kording, 2017], Chalupka et al. [2017], Rubenstein et al [2017b], Beckers and Halpern [2019a] and others 1 CHAPTER 1. INTERPRETABILITY AS A SUBFIELD OF EXPLAINABLE AI 2 pioneered the theory of causal abstraction which has been used to analyze weather patterns [Chalupka et al., 2016a] and human brains [Dubois et al, 2020a,b] Causal abstraction provides a mathematical framework for analyzing a system at multiple levels of detail simultaneously by defining when a high-level causal model is causally consistent aggregation of a (typically larger) low-level model. Chalupka et al. [2015], Geiger et al [2020b, 2021d], Hu and Tian [2022], Geiger et al [2023d], Wu
et al. [2023b] already use causal abstraction to define when a transparent algorithm is a faithful interpretation of a black box AI. However, we want to argue that causal abstraction provides the mathematical foundations for faithful interpretation by unifying a wide range of existing methodologies in a common language. Furthermore, approximate causal abstraction [Beckers et al., 2019] supports a graded notion of faithfulness that can be flexibly adapted for a wide range of settings in which a black box AI is unlikely to be perfectly summarized by transparent algorithm. We show that a wide range of existing behavioral and mechanistic interpretability methods can be understood as (approximately) uncovering abstract causal structure. We are optimistic about productive interplay between theoretical work on causal abstraction and applied work on interpretability. In stark contrast to weather and brains, we can measure and manipulate the microvariables of deep learning models with perfect
precision and accuracy, and thus empirical claims about their structure can be held to the highest standard of rigorous falsification through experimentation. 1.1 Setting the Scene Our goal here is to situate interpretability within the larger field of explainable artificial intelligence (XAI). We present a narrative at the heart XAI, which is used to establish basic vocabulary for the relevant entities and relationships. Then we fill in the details of the story with pointers to relevant literature. See Figure 11 for a summary graphic Explainable AI (XAI) is an interdisciplinary field concerned with the following sort of situation. A real world process generates data which is input to a black-box AI that is interpreted as a transparent algorithm in order to provide an explanation to human decision makers. Interpretability is a subfield of XAI concerned with interpreting black-box AI using transparent algorithms. In particular, a black box AI is decomposed into modular features that
are mapped to conceptual variables in the transparent algorithm. Mechanistic interpretability is a subfield of interpretability concerned with cases where some modular features and conceptual variables are internal, mediating the causal effect of an input on the produced output. CHAPTER 1. INTERPRETABILITY AS A SUBFIELD OF EXPLAINABLE AI 3 Figure 1.1: A visualization of interpretability as a subfield of explainable AI Interpretability is about faithfully simplifying a black box AI into a transparent algorithm, while the larger field of expainable AI also concerns the real world processes that generate the data that is input to a black box AIs and the needs of humans that make decisions using transparent algorithms. CHAPTER 1. INTERPRETABILITY AS A SUBFIELD OF EXPLAINABLE AI 1.2 4 Black Box AI Black boxes are objects whose inner contents are obscured from observers. We have complete observational and interventional access to the innards of deep learning systems, but even the
engineers that designed the AI cannot make sense of the overwhelming amounts of detailed information. Using the vocabulary of Creel [2020], the typical deep learning model is transparent in two manners. We have system knowledge, meaning we know the code used to run the AI, and process knowledge, meaning we know the hardware the AI is implemented on and the input data provided. However, we have poor algorithmic knowledge; we don’t know the the “high-level logical rules according to which the system will transform a given input into an output” [Creel, 2020]. This is what motivates the need to interpret black box AI with transparent algorithms. 1.21 Modular Features A vexed question when analyzing black box AI is how to decompose a deep learning system into constituent parts. Should the units of analysis be individual real-valued neurons, directions in the activation space of real-valued vectors of neurons, or entire attention heads? The ideal theoretical framework won’t make
us choose, but rather support any and all decompositions of a deep learning system into modular features that each have separate mechanisms from one another. We should have the flexibility to choose the units of analysis, and not bake in restrictive assumptions that have the potential to rule out meaningful structures in deep learning models. Whether a particular decomposition of a deep learning system into modular features is useful for interpretability should be understood as an empirical hypothesis that can be falsified through experimentation. 1.3 Transparent Algorithms What exactly do we mean when we say we want the high-level logical rules that govern black box AI? Practically speaking, we need to be able to explain black box AI with algorithms that are simulatable by humans and have variables corresponding to intuitive concepts. The details of what exactly makes an algorithm transparent will depend on the downstream use by human decision makers. 1.31 Simulatability A
transparent algorithm should be simulatable, or simple enough for the human decision makers to “in reasonable time, step through every calculation required to produce a prediction” [Lipton, 2018a]. This subjective definition has the flexibility to account for a variety of situation based on a contextual definition of “reasonable time”. If we need to provide algorithmic recourse for automated loan approval decisions for the general member of society, we might require that all or most humans be able to contemplate the entire algorithm at once. However, if we instead need to provide a CHAPTER 1. INTERPRETABILITY AS A SUBFIELD OF EXPLAINABLE AI 5 algorithm to a domain expert in healthcare to inspect for safety purposes, a more complex algorithm might suffice. In general, fewer variables and sparser connections between variables will result in a more simulatable model. 1.32 Conceptual Variables The variables in transparent algorithms will represent intuitive concepts [Goyal
et al., 2019a, Feder et al., 2021, Abraham et al, 2022a] that are easily understood by human decision makers (this aligns with Lipton [2018a]’s notion of decomposability). These concepts can be abstract and mathematical, such as truth-valued propositional content, natural numbers, or real valued quantities like height or weight. These concepts can also be grounded and concrete, such as the breed of a dog, the ethnicity of a job applicant, or the pitch of a singer’s voice. In general, the kind and number of concepts that need to appear in a transparent algorithm will depend on the needs and capabilities of decision makers. 1.33 Natural Language Texts as Transparent Algorithms Explanations of black box AIs that come in the form of language text have the obvious benefit of being easy to understand and produce (Hendricks et al. [2016], Ling et al [2017], Wiegreffe and Marasovic [2021], Camburu et al. [2018], Kayser et al [2021], Do et al [2020], Kayser et al [2022]; see Rajani et
al. [2019] for a review) However, the downside is that text explanations are ambiguous and difficult to systematically evaluate for faithfulness Huang et al. [2023] While much existing work concerns producing plausible or convincing arguments, text explanations also have the potential to make claims about which transparent algorithms a black box AI can be faithfully interpreted as [Wiegreffe et al., 2021, Majumder et al, 2022, Atanasova et al, 2023] For example, Atanasova et al. [2023] have an example of a natural language inference model with text explanations being unfaithful. Initially, the text explanation Just because people are talking does not mean they are having a chat is provided for some input. However, the same inference model predicts that the sentence People are talking entails They are having a chat with the explanation People are talking is a rephrasing of they are having a chat. First, the model behavior of predicting People are talking entails They are having a chat
is unfaithful to the initial explanation. But furthermore, the two explanations provided make contradictory claims about model behavior, a common issue with these methods [Camburu et al., 2020, Jang et al, 2023] Text explanations can also make claims about the internal structure of a black box AI by describing a reasoning process. Suppose that a binary image classifier outputs true with the explanation I counted the number of dogs, counted the number of cats, then output true because there are more cats. This can be understood as a claim about the AI’s internal reasoning. Namely, there should be a decomposition of AI internals into modular features that represent the number of cats, the number of dogs, and the boolean value produced by comparing the two. However, while text explanations have CHAPTER 1. INTERPRETABILITY AS A SUBFIELD OF EXPLAINABLE AI 6 the expressive power to describe internal reasoning processes, they are also vague and ambiguous. For this reason, precise
formalisms such as causal models or computer programs have advantages over natural language as a medium for explanation [Huang et al., 2023] 1.4 Interpretability Interpretability is fundamentally about understanding whether (or to what degree) the conceptual variables in a transparent algorithm are faithful interpretations of modular features that form a decomposition of a black box AI. A mathematical framework for interpretability should provide the ability to define graded faithfulness metrics that allow for apples-to-apples comparisons between existing (and future) methods. 1.41 Faithfulness Informally, faithfulness has been defined as the degree to which an explanation accurately represents the ‘true reasoning process behind a model’s behavior’ [Wiegreffe and Pinter, 2019, Jacovi and Goldberg, 2020a, Lyu et al., 2022, Chan et al, 2022a] Crucially, faithfulness should be a graded notion, but precisely which metric of faithfulness is correct will depend on the situation.
It could be that for safety reasons there are some domains of inputs that we need a perfectly faithful interpretation of a black box AI, but others that matter less. Preferably, we can defer the exact details to be filled in based on the use of the iterpretation. 1.42 Verification and Generalization Verification is the process of empirically determining whether an interpretation is faithful [Camburu et al., 2019] The ultimate verification would be a proof that a transparent model is a faithful interpretation of a black box AI for every possible input. However, it is far easier to verify an interpretation to be faithful on a large, but finite dataset generated from a real world process. If we want to develop interpretable explanation methods that are faithful when applied to an unseen input, we should seek interpretations that are likely to generalize to test examples despite only being verified on training inputs. This problem of generalizations is in no way unique to XAI methods;
generalizing from training to testing data is a central question of machine learning as a field [Hinton, 1989]. It is crucial that we develop interpretability methods that are robust to input distribution shift and generalize to real-world data. 1.43 Behavioral and Mechanistic Interpretability Behavioral interpretability methods concern the input-output behaviors of a black box AI and transparent algorithm. Such methods do not depend on the inner workings of the AI in any way, CHAPTER 1. INTERPRETABILITY AS A SUBFIELD OF EXPLAINABLE AI 7 and will provide the same transparent model for two AIs with identical behaviors. On the other hand, mechanistic interpretability requires that we decompose the internals of a black box AI into modular features that are interpreted as internal conceptual variables in the transparent algorithm. We believe this to be consistent with existing definitions of mechanistic interpretability as reverse engineering deep learning models or artificial
neuroscience. 1.5 Explanation The philosophical literature on scientific explanation recognizes a variety of modes in which phenomena can be explained. In the specific instance that we are explaining how an artifact works, there is wide consensus that explanations based in causal mechanisms enjoy a privileged status [Machamer et al., 2000]. Particularly when the components of mechanisms are understood in counterfactual terms, such explanations have the virtue of providing (at least in principle) manipulation and control of the system or artifact in question [Woodward, 2002, 2003]. In a similar vein, they allow a researcher to answer a range of ‘What if?’ questions about the topic. A prominent theme in this literature is the idea that causal/mechanistic explanations should be pitched at an appropriate level of abstraction. The operative notion of abstraction adopted here relates closely to desiderata identified in this literature concerning what makes for an appropriate level of
causal analysis [Woodward, 2021]. The fundamental operation underlying causal explanations is the intervention, which sets some number of variables in a causal model to fixed values. The nature of this intervention in the context of interpretability is left open. The intervention may be on variables for model inputs, internal model components, real-world data generation processes [Feder et al., 2021, Abraham et al, 2022a], or model training data statistics [Elazar et al., 2022] When intervention is on internal neural vector representations, the values of the intervention might be zero vectors, a perturbed or jittered version of the original vector, a learned binary mask applied to the original vector [Csordás et al., 2021, De Cao et al., 2021b], the projection of the original vector onto the null space of a linear probe [Ravfogel et al., 2020a, Elazar et al, 2020], a tensor product representation [Soulos et al, 2020a], or values realized by the vector on some other actual input
[Geiger et al., 2019a, 2020b, Vig et al, 2020b, Meng et al., 2022a] 1.6 Data Generating Process Each piece of data that is input to an AI model is first produced as the result of a complex, real-world sociological, economic, and political systems. Before an AI model trained to assign a star rating to a restaurant review written by a customer, there was a dining experience where food is eaten, service is given, and ambiance is taken in. These events culminate to a human writing a piece of text that is CHAPTER 1. INTERPRETABILITY AS A SUBFIELD OF EXPLAINABLE AI 8 eventually put into an AI model, and we can seek to understand the relationships between these events, the input data, and the output provided by a black box AI [Abraham et al., 2022a, Wu et al, 2022b]. How model predictions are affected by real world events and concepts is a question central to explaining AI. 1.61 Black Box AIs as Scientific Models Crucially, using black box AI to scientifically model real-world
processes is not the same as the XAI task of understanding how AI are impacted by real-world processes. Using deep learning as a tool for science requires empirical evidence that showing the AI is an abstraction of the real-world process [Sullivan, 2022, Cao and Yamins, 2021]. On the other hand, the XAI problem requires empirical evidence revealing how AI behavior is a downstream effect of real-world data generating processes. 1.7 Human Decision Makers Explanations of AI are ultimately for human decision makers, sometimes in highly consequential settings such as justice, healthcare, or finance. In such high stakes settings, trust and understanding of AI technology is a must. 1.71 Accidents and AI Safety Amodei et al. [2016b] understand AI safety to be concerned with the mitigation of accidents in machine learning systems, where accident is defined as “unintended and harmful behavior that may emerge when we specify the wrong objective function, are not careful about the learning
process, or commit other machine learning-related implementation errors”. Many AI accidents can be understood as some form of reward misspecification; the AI overfit its objective in a manner that is misaligned with intentions of the human designer [Pan et al., 2022, Hadfield-Menell et al, 2017] However, even if we specify rewards perfectly and observe safe behavior on training data, distributional shift may result in anomalous behaviors and accidents. We need reliable robust AI that generalize from their training data to the real world situation in which they are deployed [Dathathri et al., 2020, Kiela et al., 2021, Hupkes et al, 2022] Any methodology for preventing AI accidents will face the challenge of scalable oversight [Bowman et al., 2022] How can we provide reliable supervision on increasingly intelligent AI with limited resources? A potential answer is through human feedback [Stiennon et al., 2020] or feedback from other AI [Bai et al., 2022] Among those concerned with
existential risk from powerful AI, there are some (see, for example, Nanda [2022] or Hubinger [2019]) who argue that mechanistic interpretability has a crucial role to play. CHAPTER 1. INTERPRETABILITY AS A SUBFIELD OF EXPLAINABLE AI 1.72 9 Algorithmic Fairness and Recourse Algorithmic fairness (Dwork et al. [2012], Zemel et al [2013], Hardt et al [2016a], Chouldechova [2017]; see Mitchell et al. [2021], Pessach and Shmueli [2020] for an overview) is concerned with uncovering whether an automated decision making system is discriminatory and correcting for discrimination when discovered. There are statistical notions of fairness [Zafar et al, 2017a,b], but discrimination has been argued by many to be the causal influence of a protected attribute on the prediction of an automated decision making system. On the other hand, algorithmic recourse (Venkatasubramanian and Alfano [2020], Upadhyay et al. [2021], Dominguez-Olmedo et al [2022]; see Karimi et al [2023] for an overview)
provides a recommendation on how to make consequential changes that result in a more favourable outcome. The individual is ascribed agency and provided a contrastive explanation that suggests a plan of action. Kügelgen et al [2022] proposes that a system for algorithmic recourse can itself be judged for its fairness based on whether the cost of the recourse suggested to an individual is causally influenced by a protected attribute. We believe there are situations where mechanistic interpretability may play a crucial role in algorithmic fairness and recourse. Suppose that an automated healthcare system requires photographs of a poorly healed wound to determine whether a surgery is considered cosmetic or medically necessary. Information about the ethnicity of the individual will be revealed by the photo, so there is potential for discrimination. However, creating a counterfactual input with the same wound, but different ethnicity, seems to be near impossible. If we could instead
simulate this counterfactual by performing an intervention on the internals of the automated system, we could evaluate the system for fairness. When counterfactuals that change the protected attribute of an individual are difficult to construct, mechanistic interpretability may offer a solution. 1.73 Plausibility and Human Evaluations of Interpretations There might be certain contexts where unfaithful, but plausible, interpretations of black box AIs are useful to humans. Suppose, for instance, an AI is trained to generate organic molecules by predicting a sequence of edits and a natural language text that provides reasoning for why the prediction is made. This natural language text may be informative to material science researchers, even if the text is unfaithful to the details of the AI. However, while faithful interpretation is not required for an informative explanation, in practice we believe they will coincide. We should not expect that evaluations based on human judgments
produce interpretations faithful to the true internal reasoning process of a model. Human evaluations should only be used when unfaithful interpretations are acceptable. CHAPTER 1. INTERPRETABILITY AS A SUBFIELD OF EXPLAINABLE AI 1.8 10 The Contributions of this Thesis The next chapter of this thesis present a theoretical framework for mechanistic interpretability. In short, black box AIs and transparent algorithms can both be represented as causal models and a transparent algorithm is defined to be a faithful interpretation of a black box AI when the variables in the algorithm are a causal abstraction of the modular features in the AI. This provides a precise and rigorous mathematical foundations for interpretability, which, in the remaining chapters, we leverage to conduct empirical investigations that both uncover and induce interpretable causal structure in deep learning models. Chapter 2 Causal Abstraction for Faithful Model Interpretation Abstract A faithful and
interpretable explanation of an AI model’s behavior and internal structure is a high-level explanation that is human-intelligible but also consistent with the known, but often opaque low-level causal details of the model. We argue that the theory of causal abstraction provides the mathematical foundations for the desired kinds of model explanations. In causal abstraction analysis, we use interventions on model-internal states to rigorously assess whether an interpretable high-level causal model is a faithful description of an AI model. Our contributions in this area are: (1) We generalize causal abstraction to cyclic causal structures and typed high-level variables. (2) We show how multi-source interchange interventions can be used to conduct causal abstraction analyses. (3) We define a notion of approximate causal abstraction that allows us to assess the degree to which a high-level causal model is a causal abstraction of a lower-level one. (4) We prove constructive causal
abstraction can be decomposed into three operations we refer to as marginalization, variable-merge, and value-merge. (5) We formalize the XAI methods of LIME, causal effect estimation, causal mediation analysis, iterated nullspace projection, and circuit-based explanations as special cases of causal abstraction analysis. 2.1 Introduction We take the fundamental question in explainable artificial intelligence (XAI) to be why a deep learning model makes the predictions it does. The gold standard for explaining model behavior and the internal reasoning that underlies it should be a causal analysis [Woodward, 2003, Pearl, 2019a]. However, not just any causal explanation will be a satisfying answer to this question. After all, the parameters of deep learning models give us perfect ground-truth knowledge of the causal relationships 11 CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION 12 between all their components, so low-level causal explanations of behavior and
internal reasoning can be easily provided in terms of neurons, activation functions, and weight tensors. Ironically, ‘black box’ models are in some sense easily explainable. The problem is that these explanations are not interpretable to humansthey fail to instill a sense of understanding [Lipton, 2018a]. Interpretable causal explanations in terms of human-intelligible concepts are easy to come by, but difficult to trust. How do we know a high-level explanation isn’t a ‘just-so’ story that is unfaithful [Jacovi and Goldberg, 2020a] to the known low-level details of the model? XAI needs a theory for when a high-level causal explanation is harmonious with a low-level causal explanation. This need is not unique to XAI. Modern deep learning models are akin to the weather or a brain: they contain a large number of densely connected microvariables that have complex, non-linear dynamics. In response to these challenges, Chalupka et al [2017], Rubenstein et al [2017b], Beckers and
Halpern [2019a] and others pioneered the theory of causal abstraction. Causal abstraction provides a mathematical framework for analyzing a system at multiple levels of detail simultaneously by defining when a human-intelligible, high-level causal explanation is a faithful interpretation of opaque low-level details. Specifically, a high-level (possibly symbolic) model is a faithful proxy for a low-lever (in our setting, usually neural) model when we can align high-level variables with sets of low-level variables that play the same causal role. Causal abstraction has been applied to the study of deep learning AI models [Chalupka et al., 2015, Geiger et al, 2020b, 2021d], weather patterns [Chalupka et al., 2016a], and human brains [Dubois et al, 2020a,b] Furthermore, approximate causal abstraction [Beckers et al., 2019] supports a graded notion of faithfulness appropriate for a wide range of settings in which a low-level model is unlikely to be perfectly aligned with a high-level one. We
are optimistic about the success of causal abstraction as a theoretical framework for XAI. In stark contrast to weather and brains, we can measure and manipulate the microvariables of deep learning models with perfect precision and accuracy, and thus empirical claims about their structure can be held to the highest standard of rigorous falsification through experimentation. This Paper We develop the theory of causal abstraction as a mathematical framework for XAI. 1. We generalize a familiar notion of causal abstraction from the literature [Beckers and Halpern, 2019a] to cyclic causal models and typed high-level variables. 2. We flesh out the full theory of interchange interventions, a method used to conduct causal abstraction analysis of deep learning models [Geiger et al., 2020b, 2021d] An interchange intervention sets the variables of a model to be the value they would take for a different input. The current theory only supports high-level causal explanations with a single
intermediate variable and has no connection to approximate causal abstraction [Beckers et al., 2019] We define a general theory of interchange interventions for XAI in which high-level explanations can involve several variables (Section 2.5), and we define a notion of approximate causal abstraction CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION 13 that corresponds and this gives rise to a graded faithfulness metric of interchange intervention accuracy (Section 2.8) 3. We prove that the relation of constructive abstraction holds exactly when the high-level model can be constructed from the low-level model by marginalizing away variables, merging sets of variables, and merging values of variables (Section 2.6) This constructive characterization decomposes abstraction into constituent parts and makes explicit the relationship between marginalization and abstraction. Variable merge, value merge, and marginalization are basic and ubiquitous operations that demystify
causal abstraction. 4. We show that LIME, causal effect estimation, causal mediation analysis, iterated nullspace projection, and circuit-based explanations are all special cases of causal abstraction analysis, and that integrated gradients [Sundararajan et al., 2017b] can be used to compute interchange interventions and conduct causal abstraction analysis (Section 2.9) This lends further support to the thesis that causal abstraction provides a broad foundation for developing explanation methods that are faithful and human-interpretable. 2.2 Related Work 2.21 Causal Abstraction Our approach to structural analysis invokes an important concept from the literature on causal models, namely causal abstraction. Much attention has been devoted to the concept in recent years, with a host of alternative approaches. All of this work attempts to specify exactly when a ‘high-level causal model’ can be seen as an abstract characterization of some ‘low-level causal model’. Rubenstein et
al. [2017b] introduce the notion of an ‘exact transformation’ between models, which maps values of low-level variables to values of high-level variables. Building on this work, Beckers and Halpern [2019a] consider further restrictions on these maps, e.g, constraining the relationship between transformations of variable values and of interventions. The last definition they present is called constructive abstraction, and our work here focuses on this simple notion. Indeed, because we are working in a deterministic setting of neural network analysis, we can simplify the notation and definitions considerably (see Remark 12). Along the way, we also offer an alternative, constructive characterization of constructive abstraction (Section 2.5) The basic idea of constructive abstraction is that low-level variables are partitioned into clusters, each cluster being associated with a high-level variable. This fundamental idea has been invoked numerous times in the literature, under different
names (e.g, Iwasaki and Simon 1994, Chalupka et al 2017). Finally, beyond perfect abstraction, several authors have explored notions of approximate causal abstraction [Beckers et al., 2019, Rischel and Weichwald, 2021] Another contribution in this paper is to connect interchange intervention analysis with this line of work. Indeed, we see that a CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION 14 natural definition of interchange intervention accuracy previously employed in the literature emerges as a minor variation on existing definitions of approximate abstraction (Theorem 31). 2.22 Faithful and Interpretable Causal Explanations of AI Causal Explanation The philosophical literature on scientific explanation recognizes a variety of modes in which phenomena can be explained. In the specific instance that we are explaining how an artifact works, there is wide consensus that causal or mechanistic explanations enjoy a privileged status [Machamer et al., 2000]
Particularly when the components of mechanisms are understood in counterfactual terms, such explanations have the virtue of providing (at least in principle) manipulation and control of the system or artifact in question [Woodward, 2002, 2003]. In a similar vein, they allow a researcher to answer a range of ‘What if?’ questions about the topic. A prominent theme in this literature is the idea that causal/mechanistic explanations should be pitched at an appropriate level of abstraction. As we shall see (in Section 26), the operative notion of abstraction adopted here relates closely to desiderata identified in this literature concerning what makes for an appropriate level of causal analysis [Woodward, 2021]. The fundamental operation underlying causal explanations is the intervention, which sets some number of variables in a causal model to fixed values. The nature of this intervention in the context of XAI is left open. The intervention may be on variables for model inputs,
internal model components, real-world data generation processes [Feder et al., 2021, Abraham et al, 2022a], or model training data statistics [Elazar et al., 2022] When intervention is on internal neural vector representations, the values of the intervention might be zero vectors, a perturbed or jittered version of the original vector, a learned binary mask applied to the original vector [Csordás et al., 2021, De Cao et al, 2021b], the projection of the original vector onto the null space of a linear probe [Ravfogel et al., 2020a, Elazar et al., 2020], a tensor product representation [Soulos et al, 2020a], or values realized by the vector on some other actual input [Geiger et al., 2019a, 2020b, Vig et al, 2020b, Meng et al, 2022a]. Interpretability Adopting the vocabulary of Lipton [2018a], causal abstraction supports interpretable explanations of AI by allowing black-box AI models to be explained by high-level causal models that are transparent in two senses. First, the high-level
causal model has fewer variables and sparser causal connections, making it simulatable in the sense that a human could contemplate the entire model at once. Second, the high-level causal model is decomposable in the sense that each variable in the model admits an intuitive explanation with intelligible high-level concepts. Faithfulness Informally, faithfulness has been defined as the degree to which an explanation accurately represents the ‘true reasoning process behind a model’s behavior’ [Jacovi and Goldberg, 2020a, Lyu et al., 2022] We provide a technical definition for faithfulness by defining a high-level CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION 15 explanation to be faithful to the degree that the high-level explainer model is an approximate causal abstraction of the low-level model being explained. A shared technical definition of a graded notion of faithfulness will allow for apples-to-apples comparisons between different high-level
explanations of model behavior and internal reasoning. Often XAI methods are evaluated based on how plausible or useful a human finds them to be [Herman, 2017]. There is, however, no guarantee that evaluations based on human judgments will be faithful to the true internal reasoning process of a model [Jacovi and Goldberg, 2020a]. 2.23 Methods for Explaining AI Behavior The behavior of an AI model is simply the function from inputs to outputs that the model implements. Behavior is trivial to characterize in causal terms. Any input–output behavior can be represented by a two-variable causal model with an input variable that causes an output variable. Training Interpretable Models to Approximate the Behavior of Black Boxes Among the most popular XAI methods are LIME [Ribeiro et al., 2016a] and SHAP [Lundberg and Lee, 2017b], which both learn interpretable models that approximate an uninterpretable model. These methods define an explanation to be faithful to the degree that the
interpretable model agrees with local input–output behavior. While not advertised as causal explanation methods, when we interpret priming a model with an input as an intervention, it becomes obvious that two models having the same local input–output behavior is fundamentally a matter of causality. Crucially, however, the interpretable model lacks any connection to the internal causal dynamics of the uninterpretable model. In fact, it is often presented as a benefit that these explanations are model-agnostic methods that provide the same explanations for models with identical behaviors, but different internal structures. Without further grounding in causal abstraction, methods like LIME and SHAP cannot be trusted to tell us anything meaningful about the abstract causal structure between input and output. Estimating the Causal Effect of Real-World Concepts on Modal Behavior CausaLM [Feder et al., 2021] estimates the causal effect that real-world concepts have on the predictions of
language models by training analysis models that remove a particular concept, while accounting for a set of confounding concepts. Elazar et al [2022] develop a causal framework for understanding the impact of training data on model outputs, which they employ to demonstrate that several co-occurrence statistics have causal effects on the factual knowledge of language models. Abraham et al. [2022a] introduce an XAI benchmark where the task is to estimate the causal effect of food, service, noise, and ambiance on the sentiment label assigned by a deep learning model to a naturalistic restaurant review. Human generated interventional data is used to evaluate explainer methods, and they find a simple baseline is superior to all the methods they evaluated. CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION 16 Such starkly negative results highlight the need to ground notions of faithfulness in causality to support comparisons between explanation methods. 2.24 Methods for
Explaining the Internal Structure of AI The internal reasoning of an AI model is the intermediate computational process that underlies input–output behavior. Internal reasoning can be represented as a program or algorithm [Putnam, 1960], which in turn can be represented as a causal model [Icard, 2017a]. There is a diverse body recent research with the common aim of illuminating the causal mechanisms inside black box models by understanding the human-intelligible concepts represented by neural activations. Causal abstraction provides a mathematical foundation for understanding the high-level semantics of neural representations. In Section 29, we return to the methods of LIME, causal effect estimation, causal mediation analysis, iterated nullspace projection, and circuit-based explanations to formally characterize how these methods can be understood as special cases of causal abstraction. We also show how integrated gradients can be used to compute interchange interventions.
Interchange Interventions An interchange intervention is a method where a model is provided some input and has an internal representation h fixed to be the value that h would have realized if a different input were provided. Such interventions are central to a strand of research that proposes that symbolic algorithms and neural networks can be related through causal abstraction. We briefly review that strand of research here. Geiger et al. [2020b] initially used interchange interventions to show that a BERT-based natural language inference model creates a neural representation of lexical entailment, arguing that the ability to create modular representations underlies the capacity for systematic generalization. Li et al [2021b] use interchange interventions to show that neural representations represent propositional content that systematically alters text generated by a neural language model. Geiger et al [2021d] explicitly ground interchange intervention analysis in a theory of
causal abstraction. Finally, Geiger et al. [2022c] use interchange interventions to define differentiable training objectives that incentivize neural models to implement high-level algorithms that enable the networks to solve systematic generalization tasks. Wu et al [2022d] use interchange intervention training during language model distillation, defining a training objective that incentivizes a small ‘student’ model to be abstracted by a large ‘teacher’ model. Wu et al [2022b] use interchange intervention training to train causal proxy models that can be queried to predict the causal effects of real-world concepts on a separate original model, achieving state-of-the-art performance on the CEBaB dataset [Abraham et al., 2022a] Additionally, causal proxy models explain themselves and retain performance on the original task, so they can simply replace original models entirely. Iterative Nullspace Projection Ravfogel et al. [2020a] intervene on a neural representation to
CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION 17 remove information about a concept by projecting the representation onto the nullspaces of linear probes trained to predict the given concept. Elazar et al [2020] apply iterative nullspace projection to the analysis of pretrained language models, finding that probing techniques do not correlate with the causal presence of a concept. Lovering and Pavlick [2022a] use iterative nullspace projection to evaluate whether neural representations encode concepts with ‘mental’ causes and effects. Causal Mediation Analysis Vig et al. [2020b] conduct a causal mediation analysis to analyze gender bias in pretrained language models with transformer architectures [Vaswani et al., 2017a, Devlin et al., 2019a] They found that gender bias was mediated by a small set of neurons whose total effect on the output could be linearly decomposed into their direct and indirect effects. Meng et al. [2022a] introduce the technique of causal
tracing for mediation analysis Perturbations are made to input neurons and then intermediate neurons are restored to determine the flow of causation from input to output. This technique was used to locate and edit the factual knowledge stored in large language models. Csordás et al [2021] and De Cao et al [2021b] use differentiable binary masking to identify modular structure in neural networks by performing causal mediation analysis. Circuit-Based Explanations A recent research program aims to provide explanations of neural networks by reverse engineering the mechanisms of a networks at the level of individual neurons. Two fundamental claims made by Olah et al. [2020] are that (1) the direction of a neural vector representation encodes high-level concept(s) (e.g, curved lines or object orientation) and (2) the ‘circuits’ defined by the model weights connecting neural representations encode meaningful algorithms. An extensive study of a deep convolutional network [Cammarata et
al., 2020] support these core claims with a variety of qualitative and quantitative experiments. Elhage et al. [2021] extend this circuit framework to attention-only transformers, discovering ‘induction heads’ that search for previous occurrences of the present token and copy the token directly after the previous occurence. In follow-up work, Olsson et al [2022] provides indirect evidence that induction heads underlie the ability of standard transformers to achieve in-context learning. Finally, Elhage et al. [2022] propose softmax linear units as an activation function that can be used in transformers to incentivize the creation of interpretable neurons. Other Intervention-Based Research Giulianelli et al. [2018b] intervene on an LSTM language model at prediction time to encourage better performance on English subject–verb agreement; probes help them home in on the most productive intervention points, but it is the intervention itself that positively affects model behavior.
Soulos et al. [2020a] use interventions on encoder output representations to support claims about the compositional structure of recurrent neural networks being explained by a tensor product representation [McCoy et al., 2019] Bau et al. [2019a] and Besserve et al [2020a] search clusters of neurons that perform modular functions in generative adversarial vision networks. They use intervention experiments to illuminate CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION 18 the purpose of the neural modules, discovering that they encode intuitive high-level concepts such as the presence of trees in the generated image. Probes Probing is the technique of using a supervised or unsupervised model to determine whether a concept is present in a neural representation of a separate model. Probes are a popular tool for analyzing deep learning models, especially pretrained language models [Hupkes et al., 2018b, Conneau et al., 2018, Peters et al, 2018, Tenney et al, 2019, Clark et
al, 2019] Although probes are quite simple, our theoretical understanding of probes has greatly improved since their recent introduction into the field. From an information-theoretic point of view, we can observe that using arbitrarily powerful probes is equivalent to measuring the mutual information between the concept and the neural representation [Hewitt and Liang, 2019, Pimentel et al., 2020] If we restrict the class of probing models based on their complexity, we can measure how usable the information is [Xu et al., 2020, Hewitt et al, 2021] Regardless of what probe models are used, successfully probing a neural representation does not guarantee that the representation plays a causal role in model behavior [Ravichander et al., 2020, Elazar et al, 2020, Geiger et al, 2020b, 2021d] Feature Attribution Feature attribution methods ascribe scores to neural representations that capture the ‘impact’ of the representation on model behavior. Gradient-based feature attribution methods
[Zeiler and Fergus, 2014a, Springenberg et al., 2014, Shrikumar et al, 2016, Binder et al, 2016] measure causal properties when they satisfy some basic axioms [Sundararajan et al., 2017b] Chattopadhyay et al. [2019] argue for a direct measurement of a neuron’s individual causal effect, and Geiger et al. [2021d] provide a natural causal interpretation of the integrated gradients method 2.3 Causal Models We start with some basic notation. Throughout we use V to denote a fixed set of variables, each X ∈ V coming with a range Val(X) of possible values. Definition 1 (Partial and Total Settings). We assume Val(X) ∩ Val(Y ) = ∅ whenever X ≠ Y , meaning no two variables can take on the same value. 1 This assumption allows representing the values of a set of variables X ⊆ V as sets x of values, with exactly one value x ∈ Val(X) in x for each X ∈ X. We refer to the values x ∈ Val(X) as partial settings In the special case where v ∈ Val(V), we call v a total setting.
Remark 2 (Notation throughout the paper). Capital letters (eg, X) are used for variables and lower case letters (e.g, x) are used for values Bold faced letters (eg X or x) are used for sets of 1 This is akin to representing partial and total settings as vectors where a value can occur multiple times. For instance, we can simply take any causal model where variables share values, and then ‘tag’ the shared values with variable names to make them unique. CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION 19 variables and sets of values. When a variable (or set of variables) and a value (or set of values) have the same letter, the values correspond to the variables (e.g, x ∈ Val(X) or x ∈ Val(X)) We use Domain(f ) to denote the domain of a function f , Uniform(X ) to be a uniform distribution on (a finite set) X , and 1[φ] to be an indicator function that takes in a proposition φ and outputs 1 if φ is true, 0 otherwise. Another useful construct in this
connection is the projection of a partial setting: Definition 3 (Projection). Given a partial setting u for a set of variables U ⊇ X, we define Proj(u, X) to be the restriction of u to the variables in X. Given a partial setting x and a set U ⊇ X: Proj (x, U) = {u ∈ Val(U) ∶ Proj(u, X) = x}. −1 Definition 4. A (deterministic) causal model is a pair M = (V, F), such that V is a set of variables and F = {fV }V ∈V is a set of structural functions, with fV ∶ Val(V) Val(V ) assigning a value to V as a function of the values of all the variables. Remark 5 (Inducing Graphical Structure). Observe that our definition of causal model makes no explicit reference to a graphical structure defining a causal ordering on the variables. While the structural function for a variable takes in total settings, it might be the case that the output of a structural function depends only on a subset of values. We can use the structural functions to induce a causal ordering among the variables by
saying that Y ≺ Xor Y is a parent of Xjust in case there is a setting w of all the variables W = V {X, Y } other than X and Y , and two settings y, y ′ of Y such that fX (w, y) ≠ fX (w, y ) (see, e.g, Woodward 2003) The resulting order ≺ captures a ′ general type of causal dependence. Throughout the paper, we will define structural functions to take in partial settings of parent variables, though technically they take in total settings and ignore all values other than those of the parent variables. When ≺ is irreflexive, we say the causal model is acyclic. Most of our examples of causal models will have this property. However, it is often also possible to give causal interpretations of cyclic models (see, e.g, Bongers et al 2021) Indeed, the abstraction operations to be introduced generally create cycles among variables, even from initially acyclic models (see Rubenstein et al. 2017b, §53 for an example). In Section 210, we provide an example where we abstract a causal
model representing the bubble sort algorithm into a cyclic model where any sorted list is a solution satisfying the equations. Remark 6 (Acyclic Model Notation). Our running example will be causal abstraction holding between two finite, acyclic causal models. We will call variables that depend on no other variable In Out input variables (XM ), and variables on which no other variables depend output variables (XM ). In The remaining variables are intermediate variables. We can intervene on input variables XM to “prime” the model with a particular input. As such, we will define the functions for input variables in our examples to be constant functions that will be overwritten. CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION 20 As M can also be interpreted simply as a set of equations, we can define the set of solutions, which may be empty. Definition 7 (Solution Sets). Given M = (V, F), the set of solutions, called Solve(M), is the set of all v ∈ Val(V) such
that all the equations v = fV (v) are satisfied for each v ∈ v. When M is acyclic, each intervention results in a model with a single solution and we use Solve(Mi ) to refer interchangeably to a singleton set of solutions and its sole member, relying on context to disambiguate. We can also give a general definition of intervention on a model (see, e.g, Spirtes et al 2000, Woodward 2003, Pearl 2009). Definition 8 (Intervention). Let M be a model We define an intervention to be a partial setting i ∈ Val(I) for I ⊆ V. We define Mi to be just like M, except that we replace fX with the constant function v ↦ Proj(i, X) for each X ∈ I. Due to our focus on fully observable, deterministic systems, we forego the opportunity to incorporate probability into our causal models. However, we discuss extensions of the present work to the probabilistic setting in Section 2.11 2.4 Example of Causal Models: A Symbolic Algorithm and Neural Network For our running example, we define two basic
examples of causal models that demonstrate a potential to model a diverse array of computational processes; the first causal model represents a tree-structured algorithm and the second causal model represents a fully-connected feed-forward neural network. Both the network and the algorithm solve the same ‘hierarchical equality’ task. 2.41 Hierarchical Equality Task A basic equality task is to determine whether a pair of objects are identical. A hierarchical equality task is to determine whether a pair of pairs of objects have identical relations. The input to the hierarchical task is two pairs of objects, and the output is True if both pairs are equal or both pairs are unequal, and False otherwise. For illustrative purposes, we define the domain of objects to consist of a triangle, square, and pentagon. For example, the input (D, D, ^, □) is assigned the output False and the inputs (D, D, ^, ^) and (D, □, ^, □) are both labeled True. We chose this task for two reasons.
First, there is an obvious tree-structured symbolic algorithm that solves the task: compute whether the first pair is equal, compute whether the second is equal, then compute whether those two outputs are equal. We will encode this algorithm as a causal model. Second, equality reasoning is ubiquitous and has served as a case study for broader questions CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION Id2 (v1 , v2 ) True O Id1 (i1 , i2 ) Id1 (i3 , i4 ) V1 I1 True True V2 I2 { 21 }{ I3 }{ I4 }{ } (b) The total setting of A determined by the empty intervention. True Id1 T F F F T F F F T Id2 T F T T F F F T True True (c) The total setting of A determined by the intervention fixing X3 , X4 , and V2 to be ^, □, and True. (a) The algorithm. Figure 2.1: A tree-structured algorithm that perfectly solves the hierarchical equality task with a compositional solution. about the representations underlying relational reasoning in
biological organisms Marcus et al. [1999], Alhama and Zuidema [2019], Geiger et al. [2022b] For present purposes, hierarchical equality will serve as a case study for explaining how abstract tree-structured composition can be implemented by a fully-connected neural network. We will define a causal model of the symbolic algorithm and a causal model of a neural network that labels hierarchical equality inputs. The neural network implements the hierarchical equality task because we trained it to [Geiger et al., 2022c] 2.42 A Tree-Structured Algorithm for Hierarchical Equality We define a tree structured algorithm A consisting of four ‘input’ variables XA = {X1 , X2 , X3 , X4 } In each with possible values {D, ^, □}, two ‘intermediate’ variables V1 , V2 with values {True, False}, 2 Out and one ‘output’ variable XA = {O} with values {True, False}. The acyclic causal graph is depicted in Figure 2.1a, where each fXi (with no arguments) is a constant function to D (which
will be overwritten, per Remark 6), and fV1 , fV2 , fO are all identity over their respective domains, e.g, fV1 (x1 , x2 ) = 1[x1 = x2 ]. A total setting can be captured by a vector [x1 , x2 , x3 , v1 , v2 , o] of values for each of the variables. The default total setting that results from no intervention is [D, D, D, D, True, True, True]. We can also ask what would have occurred had we intervened to fix X3 ,X4 , and V2 to be ^, □, and True, for example. The counterfactual result is [D, D, ^, □, True, True, True] 2 Strictly speaking, these variable values are indexed by their variables. We omit the indices for readability CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION 2.43 22 A Fully Connected Neural Network for Hierarchical Equality Out Define a neural network N , consisting of eight ‘input’ neurons XN = {R1 , . , R8 }, twenty-four Out ‘intermediate’ neurons H(i,j) for 1 ≤ i ≤ 3 and 1 ≤ i ≤ 8, and two ‘output’ neurons XN =
{OTrue , OFalse }. The values for each of these variables are the real numbers R We depict the causal graph in Figure 2.2 Define R, H1 , H2 , H3 to be the sets of variables for the first four layers, respectively. We define fRk (with no arguments) to be a constant function to 0, for 1 ≤ k ≤ 8. The intermediate and output neurons are determined by the network weights W1 , W2 , W3 ∈ R 8×8 , W4 ∈ R 8×2 , and 8 bias terms b1 , b2 , b3 ∈ R , b4 ∈ R. For 1 ≤ j ≤ 8, we define fH(1,j) (r) = ReLU((rW1 + b1 )j ) fH(2,j) (h1 ) = ReLU((h1 W2 + b2 )j ) fH(3,j) (h2 ) = ReLU((h2 W3 + b3 )j ) fOTrue (h3 ) = ReLU((h3 W4 + b4 )0 ) 2 fOFalse (h3 ) = ReLU((h3 W4 + b4 )1 ) The four shapes that are the input for the hierarchical equality task are represented in rD , r□ , r^ ∈ R by a pair of neurons with randomized activation values. The network outputs True if the value of the output logit OTrue is larger than the value of OFalse , and False otherwise. We can simulate a
network operating on the input (□, D, □, ^) by performing an intervention setting (R1 , R2 ) and (R5 , R6 ) to r□ , (R3 , R4 ) to rD , and (R7 , R8 ) to r^ . In Figure 2.2, we define neural representations for the shapes and the weights of a network N trained to implement our tree-structured algorithm A [Geiger et al., 2022c] While the network N was trained to implement A, looking at the network weights provides no insight into this relationship; the theoretical tool of causal abstraction is needed to illuminate ‘black-box’ neural networks with intervention experiments. 2.5 Causal Abstraction and Interchange Intervention Analysis Suppose we have a ‘low-level model’ L = (VL , FL ) built from ‘low-level variables’ VL and a ‘high-level model’ H = (VH , FH ) built from ‘high level variables’ VH . (For instance, these might be N and A, respectively, from the previous section.) What structural conditions must be in place for H to be a high-level abstraction of
the low-level model L? 2.51 Alignments Between Causal Models A prominent intuition about abstraction is that it may involve associating specific high-level variables with clusters of low-level variables. To systematize this, we introduce a notion of an alignment between a low-level and high-level causal model: CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION OTrue OFalse H(3,1) H(3,2) H(3,3) H(3,4) H(3,5) H(3,6) H(3,7) H(3,8) H(2,1) H(2,2) H(2,3) H(2,4) H(2,5) H(2,6) H(2,7) H(2,8) H(1,1) H(1,2) H(1,3) H(1,4) H(1,5) H(1,6) H(1,7) H(1,8) R1 R2 R3 R4 R5 R6 R7 R8 { } { } { } { } rD = [0.012, −0301] ⎡ 2.1754e + 00 −71769e − 01 ⎢ ⎢ ⎢ ⎢ 7.3452e − 03 −71636e − 03 ⎢ ⎢ ⎢ ⎢ ⎢−2.1238e + 00 −16024e + 00 ⎢ ⎢ ⎢ 1.4659e − 02 8.5846e − 03 ⎢ W1 = ⎢ ⎢ ⎢ 3.9916e + 00 −13048e + 00 ⎢ ⎢ ⎢ ⎢ −3.6872e + 00 4.1371e − 01 ⎢ ⎢ ⎢ ⎢ −2.5656e − 01 8.9275e − 01 ⎢ ⎢ ⎢
⎢ 3.7633e − 03 ⎣−8.2522e − 03 b1 = [−0.0475 ⎡ −4.9015e − 01 ⎢ ⎢ ⎢ ⎢ 9.6085e + 00 ⎢ ⎢ ⎢ ⎢ 1.0056e + 00 ⎢ ⎢ ⎢ ⎢−1.0586e + 00 ⎢ W2 = ⎢ ⎢ ⎢ 1.2459e + 00 ⎢ ⎢ ⎢ ⎢ −7.8904e − 01 ⎢ ⎢ ⎢ ⎢ −1.5830e + 00 ⎢ ⎢ ⎢ ⎢ ⎣−1.0493e + 00 −4.9981e − 03 −1.5522e − 02 −5.7124e + 00 9.8556e + 00 −3.9747e + 00 −1.1877e − 01 −1.3206e + 00 9.5044e − 02 b2 = [0.2746 ⎡ 4.2248e − 01 2.7740e + 00 ⎢ ⎢ ⎢ ⎢−2.8812e − 01 −31679e − 01 ⎢ ⎢ ⎢ ⎢ 3.1208e + 00 1.4330e + 00 ⎢ ⎢ ⎢ ⎢ 2.9027e + 00 −38384e + 00 ⎢ W3 = ⎢ ⎢ ⎢ 1.7727e + 00 ⎢−4.8936e + 00 ⎢ ⎢ ⎢ −2.6034e − 01 −25041e − 01 ⎢ ⎢ ⎢ ⎢ 4.8684e − 01 −39053e − 03 ⎢ ⎢ ⎢ ⎢ 1.9812e + 00 ⎣ 1.7658e + 00 r□ = [−0.812, 0456] r^ = [0.682, 0333] −2.1649e + 00 6.8584e − 01 −35393e − 02 −3.2291e − 04 −30044e − 03 2.2651e + 00 2.1323e + 00 1.5860e + 00 8.2233e − 04
−7.3199e − 03 6.3861e − 03 −25226e + 00 −3.9189e + 00 1.2616e + 00 −34642e − 02 3.7053e + 00 −40671e − 01 1.3940e − 02 2.4427e − 01 −82323e − 01 −32437e − 02 −7.6974e − 03 6.7534e − 03 −42970e − 01 −0.0151 0.0692 −1.0025e + 01 2.9515e + 00 2.5843e − 02 −2.9570e − 02 −2.7281e − 01 1.9272e − 01 −1.2504e − 01 −3.9964e − 01 −0.1292 −1.7210e − 02 −4.7480e − 03 −3.5200e + 00 8.4745e + 00 −2.1867e + 00 −2.1950e − 02 −5.1255e − 01 9.0716e − 02 0.2347 1.0335e + 00 1.5999e − 02 −5.8909e + 00 −1.2183e + 01 3.4905e + 00 −1.7218e − 01 −1.3724e − 01 9.3662e − 01 −0.0451 0.2833 −0.0777 6.8148e − 03 −31675e − 02 −29090e − 02⎤ ⎥ ⎥ ⎥ −1.9941e + 00 −22740e + 00 2.0176e + 00⎥ ⎥ ⎥ ⎥ 5.6526e − 03 3.6328e − 02 2.9736e − 03⎥ ⎥ ⎥ ⎥ −1.2776e + 00 2.5292e + 00 1.2696e + 00⎥ ⎥ ⎥ ⎥ −1.4530e − 02 7.7398e − 03 −53017e − 02⎥ ⎥ ⎥
⎥ −1.1155e − 03 −10047e − 03 6.5350e − 03⎥ ⎥ ⎥ ⎥ 3.5100e − 02 3.6350e − 04 −56254e − 02⎥ ⎥ ⎥ ⎥ 3.0961e + 00 4.1024e − 01 −31124e + 00⎥ ⎦ 0.0627 −7.1630e − 01 1.3156e + 01 −5.1716e − 01 5.6010e − 01 −1.0082e + 00 1.1614e − 02 −3.0161e + 00 −2.4861e + 00 −0.3366 23 −0.0591 −4.7900e − 01 1.2285e + 01 5.0762e − 02 2.2497e − 02 2.6007e − 01 −4.2569e + 00 −1.5335e + 00 −7.6741e − 02 −0.0196] −1.0036e − 01 3.3859e − 01 −6.9766e − 02 −3.9709e − 02 −1.2949e + 00 −1.8045e − 01 −6.2201e − 02 −2.2552e − 01 1.1562e − 02⎤ ⎥ ⎥ ⎥ 1.7905e − 02⎥ ⎥ ⎥ ⎥ −3.6073e + 00⎥ ⎥ ⎥ ⎥ 8.3327e + 00⎥ ⎥ ⎥ ⎥ −9.1204e − 01⎥ ⎥ ⎥ ⎥ −4.6835e − 02⎥ ⎥ ⎥ ⎥ −8.8656e − 01⎥ ⎥ ⎥ ⎥ 5.3066e − 02⎥ ⎦ −0.1478 00980 01673] −4.5362e + 00 8.7838e − 01 −1.9106e − 01 7.9891e − 03 3.6747e − 02 −17160e + 00 4.0099e − 02
−61718e − 01 −2.3871e + 00 6.8049e − 01 −3.5845e − 03 1.1953e − 01 −9.7928e + 00 7.0597e − 02 −1.6880e − 02 1.0065e + 00 −3.1472e − 01 6.2744e − 02 2.0737e + 00 9.6990e − 01 −2.7405e + 00 −6.4683e − 02 1.4577e − 01 −9.9469e − 01 −4.0193e − 01 −3.2715e − 01 −1.1598e + 01 −1.2437e + 01 2.5876e + 00 2.8834e − 02 −1.0602e + 00 2.9824e + 00 b3 = [−0.8259 −0.1688 −2.4404 1.8469 0.9255 −0.3003 −0.0837 −0.2753] 3.8760 W4 = [ −3.8155 −0.1366 −0.0268 −3.1693 2.9001 −2.7446 2.4848 −1.8302 1.8365 −0.1413 −0.0793 −12.2571 12.3972 −1.7698 ] 1.2339 b4 = [3.2679 −3.5912] −3.6744e − 01⎤ ⎥ ⎥ ⎥ −2.8038e − 01⎥ ⎥ ⎥ ⎥ 2.9503e − 01⎥ ⎥ ⎥ ⎥ 1.9529e − 01⎥ ⎥ ⎥ ⎥ 1.1965e + 00⎥ ⎥ ⎥ ⎥ 5.3538e − 02⎥ ⎥ ⎥ ⎥ −5.7370e − 02⎥ ⎥ ⎥ ⎥ −1.3896e + 00⎥ ⎦ Figure 2.2: A fully-connected feed-forward neural network that labels inputs for the
hierarchical equality task. We define weights for a network that was created with interchange intervention training to implement the tree-structured solution to the task. CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION 24 O V1 I1 { V2 I2 }{ I3 }{ I4 }{ } OTrue OFalse H(3,1) H(3,2) H(3,3) H(3,4) H(3,5) H(3,6) H(3,7) H(3,8) H(2,1) H(2,2) H(2,3) H(2,4) H(2,5) H(2,6) H(2,7) H(2,8) H(1,1) H(1,2) H(1,3) H(1,4) H(1,5) H(1,6) H(1,7) H(1,8) X1 X2 X3 X4 X5 X6 X7 X8 { } { } { } { } Figure 2.3: An alignment between the causal graphs of a low-level fully-connected neural network (bottom) and a high-level tree structured algorithm (top). Definition 9 (Alignment). An alignment between L and H is given by a pair ⟨Π, τ ⟩ of a partition Π = {ΠXH }XH ∈VH ∪{⊥} and a family τ = {τXH }XH ∈VH of maps, such that: 1. The partition Π of VL consists of non-overlapping, non-empty cells ΠXH ⊆ VL for each XH ∈ VH , in
addition to a (possibly empty) cell Π⊥ , 2. There is a partial surjective map τXH ∶ Val(ΠXH ) Val(XH ) for each XH ∈ VH In words, the set ΠXH consists of those low-level variables that are ‘aligned’ with the high-level variable XH , and τXH tells us how a given setting of the low-level cluster ΠXH corresponds to a setting of the high-level variable XH . The remaining set Π⊥ consists of those low-level variables that are ‘forgotten’, not mapped to any high-level variable. Fact 1. An alignment ⟨Π, τ ⟩ induces a unique translation, viz a (partial) function τ ∶ Val(VL ) Val(VH ). To wit, for any vL ∈ Val(VL ): τ (vL ) def = ⋃ τXH (Proj(vL , ΠXH )). (2.1) XH ∈VH The notion of an alignment (or of a translation) by itself does not tell us anything about how the causal laws, as specified by FL and FH , relate. We also want these to be in harmony, in the sense that (counterfactually) intervening on one produces ‘the same’ result in the
other, where ‘the same’ means that the resulting states remain aligned. In that direction, note that τ extends naturally from a partial function Val(VL ) Val(VH ), defined just on total settings, to a partial function τ ∶ ⋃XL ⊆VL Val(XL ) ⋃XH ⊆VH Val(XH ), defined on the larger domain of interventions (viz. partial settings) We only define τ on low-level CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION 25 interventions that correspond to high-level interventions. We set τ (xL ) = xH if and only if τ (Proj (xL , VL )) −1 Proj (xH , VH ). −1 = (2.2) Thus, the cell-wise maps τXH canonically give us this partial function τ . Note that we are overloading the notation τ , letting it refer to both the function on total settings in Eq. (21) and the function on partial settings specified by Eq. (22) 2.52 Causal Consistency and Constructive Abstraction Definition 10 (Causal Consistency). An alignment ⟨Π, τ ⟩ between L and H is
consistent if for all i ∈ Domain(τ ) such that Solve(Li ) is non-empty, the following diagram commutes: i Solve(Li ) τ τ (i) τ Solve(Hτ (i) ) That is, the high-level intervention τ (i) corresponding to low-level intervention i results in the same high-level total settings as the result of first determining a low-level setting from i and then applying the translation to obtain a high-level setting. In a single equation: τ (Solve(Li )) = Solve(Hτ (i) ). (2.3) We thus arrive at a definition of causal abstraction: Definition 11 (Constructive Abstraction). H is a constructive abstraction of L under an alignment ⟨Π, τ ⟩ if the causal consistency condition is satisfied. Remark 12. Though the idea was implicit in much earlier work, Beckers and Halpern [2019a] and Beckers et al. [2019] explicitly introduced the notion of a constructive abstraction in the setting of probabilistic causal models. The definition here is considerably simplified and streamlined, in part
because we are not concerned with probability (though see Section 2.11 below) Remark 13. In the field of program analysis, abstract interpretation is a framework that can be understood as a special case of causal abstraction where models are acyclic and high-level variables are aligned with individual low-level variables rather than sets of low-level variables [Cousot and Cousot, 1977]. The functions τ and τ −1 are the abstraction and concretization operators that form a Galois connection, and causal consistency guarantees that abstract transfer functions are consistent with concrete transfer functions. CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION 26 While we assumed in Section 2.3 that variables have non-overlapping spaces of values, it is often useful to assume that variable do share values, and we may want to consider more restrictive varieties of causal abstraction that ensure a kind of type consistency. In that direction, define: Definition 14. Fix a
causal model (V, F) and a set T of types A typing is a function T ∶ V T , together with an equivalence relation ∼, such that for each X, Y of the same typei.e, T (X) = T (Y ) and each x ∈ Val(X), there is exactly one y ∈ Val(Y ) such that x ∼ y. In other words, there is a bijective correspondence ∼ between values of variables of the same type, where intuitively x ∼ y means that values x and y should behave the same way. For instance, if X and Y are both Boolean variables with values {0X , 1X } and {0Y , 1Y }, respectively, then we would naturally have 0X ∼ 0Y and 1X ∼ 1Y . This relation ∼ lifts to sets of values in the obvious way: for disjoint sets of variables, Z and W, we say z ∼ w if, for each z ∈ z there is exactly one w ∈ w such that z ∼ w, and vice versa. With this much can introduce: Definition 15 (Typed Causal Abstraction). Fix causal models L = (VL , FL ) and H = (VH , FH ) with typing relations ∼L and ∼H , respectively. H is a typed causal
abstraction of L if there is an alignment ⟨Π, τ ⟩ that is causally consistent (Def. 10) and also type consistent: whenever X and Y have the same type in H, then for any z ∈ Val(ΠX ) and w ∈ Val(ΠY ), we have, z ∼L w ⇔ τ (z) ∼H τ (w). We mention Definition 15 because typed abstraction played a crucial role in the ability of Geiger et al. [2022c] to use interchange intervention training to solve a vision-based systematic generalization task. Typing assumptions have also been useful in causal discovery [Brouillard et al, 2022] In Section 2.10, we provide an example involving a typing where variables are split into Booleans and natural numbers to define a causal model representing the bubble sort algorithm. 2.53 Interchange Intervention Analysis Equipped with the general notion of (constructive) causal abstraction, we now turn to one concrete means of operationalizing claims of causal abstraction. Interchange intervention analysis is one such method. Geiger et al
[2022c] provides a specialized theory of interchange interventions that only covers cases where the high-level causal model has a single intermediate variable. Furthermore, they formulate the metric of interchange intervention accuracy to capture partially successful modelinternal explanations. Here, we expand on this work, presenting a general theory of interchange interventions for high-level causal models with multiple intermediate variables. When applying causal abstraction to explain neural networks, the central question is how to find an alignment ⟨Π, τ ⟩ between the algorithm and network. The alignment between the input and output variables in the network and algorithm is simply stipulated by the researcher who knows how CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION 27 they will interpret the inputs and outputs of the network. In contrast, the researcher must search for the best alignment between intermediate variables. In this search, we partition
the low-level intermediate variables, then with τ already stipulated for input and output, we determine the value of τ for remaining low-level cells. Definition 16 (Constructing an Alignment for Interchange Intervention Analysis). Consider some In Out neural network N and high-level algorithm A with input and output variables XL , XL ⊆ VL and In Out XH , XH ⊆ VH . Suppose we have {ΠX }X∈VH ∪⊥ , which is a partition of VL where input, output, and intermediate variables are aligned across the two levels: Out X ∈ XH Out ⇔ ΠX ⊆ X L Out In In X ∈ X H ⇔ ΠX ⊆ X L In Out X ∈ VH (XH ∪ XH ) ⇔ ΠX ⊆ VL (XL In In ∪ XL ). Out with predefined input and output alignment τX for X ∈ VH ∪ VH . We can induce the intermediate semantics from the input semantics, output semantics, aligned partitions, and the two causal models. For XH ∈ VH and zL ∈ ΠXH , if there exists xL ∈ XL such that zL = Proj(Solve(NxL ), ΠXH ), then In we define the
intermediate semantics as follows τXH (zL ) = Proj(Solve(Aτ (xL ) ), XH ) and otherwise, we leave τXH undefined for zL . 3 Once an alignment ⟨Π, τ ⟩ is constructed, aligned interventions must be performed to experimentally verify the alignment is a witness to the algorithm being an abstraction of the network. Observe that τ will be only be defined for values of intermediate partition cells that are realized when some input is provided to the network. This greatly constrains the space of low-level interventions to intermediate partitions that will correspond with high-level interventions. However, we will still be able to interpret certain low-level interchange interventions [Geiger et al., 2020b] Originally, an interchange intervention was relative to a base input and source input The network is run with the base input, while a set of neurons is fixed to be the value they would have taken if the source input were provided instead. However, having a single source input
restricts analyses to high-level causal models with a single intermediate variable. To support arbitrary high-level models, we extend the definition to multiple source inputs. Because neural networks involve vectors of real numbers, there are in principle many possible interventions one could perform on neural variables. Most of these settings, however, need not have any intuitive significance when it comes to explaining model behavior on a range of possible inputs. 3 This is a subtle point, but, in general, this construction is not well defined because two different input can produce the same intermediate representation with different high-level interpretations. If this were to happen, the causal abstraction relationship simply wouldn’t hold for the alignment. CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION False 28 True True False False True False OTrue OFalse True OTrue OTrue h(3,0) h(3,1) h(3,2) h(3,3) h(3,4) h(3,5) h(3,6) h(3,7) h(3,0)
h(3,1) h(3,2) h(3,3) h(3,4) h(3,5) h(3,6) h(3,7) h(2,0) h(2,1) h(2,2) h(2,3) h(2,4) h(2,5) h(2,6) h(2,7) h(2,0) h(2,1) h(2,2) h(2,3) h(2,4) h(2,5) h(2,6) h(2,7) h(1,0) h(1,1) h(1,2) h(1,3) h(1,4) h(1,5) h(1,6) h(1,7) h(1,0) h(1,1) h(1,2) h(1,3) h(1,4) h(1,5) h(1,6) h(1,7) x1 x2 x3 x4 x5 x6 x7 x8 r1 r2 r3 r4 r5 r6 r7 r8 Figure 2.4: The result of aligned interchange intervention on the low-level fully-connected neural network and a high-level tree structured algorithm under the alignment in Figure 2.3 Observe the equivalent counterfactual behavior across the two levels. For this reason, we restrict our causal abstraction relation to multi-source interchange interventions that are realized for some possible input. Definition 17 (Interchange Interventions). Consider a neural network N , with input and output In Out variables XL , XL ⊆ VL , disjoint subsets of intermediate variables XL , . , XL ⊆ VL (XL ∪XL ), 1 k In Out
In and base input b and source inputs s1 , . , sk ∈ XL We define an interchange intervention to be the partial setting that sets the input variables to the base input b and the intermediate variables i XL to the values they would take on if the source input si were provided to N 1 k def 1 k IntInv(N , b, ⟨s1 , . , sk ⟩, ⟨XL , , XL ⟩) = b ∪ Proj((Solve(Ns1 ), XL ) ∪ ⋅ ⋅ ⋅ ∪ Proj(Solve(Nsk ), XL ) The interchange interventions used by Geiger et al. [2021d] are limited in a second way: they consider only interventions that, for each low-level partition cell, either fix the entirety of the cell or fix none of the cell. The resulting set of interventions will verify that every actually realized value of a low-level partition cell can be used in every actually realized context. Formally, the restricted set of interchange interventions can be represented as IntInv(N , b, ⟨s1 , . , sk ⟩, ⟨ΠX1H , , ΠXkH ⟩) 1 k where b and s1 , . , sk are
base and source inputs and XH , , XH are high level variables 2.54 Explanation and Generalization Interchange interventions set variables to values they actually realize on an input. Crucially, this input can come from training data, testing data, or be a theoretical input from some implicitly defined space of well-formed inputs (e.g, the space of all English questions and answers) Definition 10 requires a commuting diagram hold for all interventions in the range of the partial function τ . So, CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION 29 if a high-level causal model is an abstraction of a deep learning model where τ is only defined on training data, the abstraction may not hold when τ is expanded to be defined on testing data. If we want to develop interpretable explanation methods that are faithful when applied to unseen real-world input, we should seek explanations that generalize to test examples. This problem of generalizations is in no way
unique to XAI methods; generalizing from training to testing data is a central question of machine learning as a field [Hinton, 1989]. 2.6 Decomposing Constructive Causal Abstraction Given the importance and prevalence of this relatively simple notion of abstraction, it is worth understanding the notion from different angles. Constructive causal abstractions can be fully characterized in terms of three fundamental operations on a causal model. Marginalization removes a set of variables from a causal model. Variable merge collapses a partition of variables from a causal model, each partition cell becoming a single variable. Value merge collapses a partition of values for each variable from a causal model, each partition cell becoming a single value. The first and third operations relate closely to concepts identified in the philosophy of science literature as being critical to addressing the problem of variable choice [Woodward, 2021]. Remark 18. For the following definitions, we
will need a choice function Choose that takes in a set and outputs an element of that set. We also employ a function Filter that takes a variable Y and a value y ∈ Val(Y ), and returns any value in Val(Y ) not equal to y. Marginalization removes a set of variables from a causal model, directly linking the parents and children of each variable. Definition 19 (Marginalization). Choose some X ⊂ V and define ϵX (M) as follows The variables are W = V X, with unaltered value spaces. For each variable Y ∈ W, define a new function ϵX (fY ) on Val(W) as follows: ⎧ ⎪ ⎪Proj(w, Y ) ϵX (fY )(w) = ⎪ ⎨ ⎪ ⎪ ⎪ ⎩Filter(Y, Proj(w, Y )) if Proj (w, V) ∩ Solve(M) ≠ ∅, −1 otherwise. Marginalization is essentially a matter of ignoring a subset X of variables. Philosophers of science have been concerned with scenarios in which a cause of some effect is relatively insensitive to, or stable under, changes in ‘background variables’ [Lewis, 1986, Woodward, 2006]. That
is, if we simply ignored other variables that also make a difference for some effect, how reliably would a given factor cause the effect? The definition of marginalization we give here essentially guarantees perfect insensitivity/stability in this sense. Specifically: CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION 30 Lemma 20. The model ϵx (M) is a constructive causal abstraction of M with translation map τ (v) = Proj(v, V X). Equivalently, for all i ∈ Domain(τ ), Proj(Solve(Mi ), V X) = Solve(ϵX (M)Proj(i,VX) ), i.e, solution sets are preserved under marginalization, for all interventions Variable merge collapses each cell of a partition into single variables that depend on the parents of their partition and determines the children of their partition. We define Π(M) to be the model where the variables merged according to a partition {ΠX }X∈W with cells indexed by a new set of variables W. Definition 21 (Variable Merge). Choose some partition {ΠX
}X∈W of V and define Π(M) as follows The variables are W. Each variable X ∈ W is assigned the values Val(ΠX ) Let π −1 be the natural bijective function from ⋃X⊆W Val(X) to ⋃X⊆V Val(X) where the value assigned to each X ∈ W is decomposed into values for each variable in ΠX . This means π is a bijection from ⋃X⊆V Val(X) to ⋃X⊆W Val(X). Each X ∈ W is assigned the function Π(fX )(v) = π( ⋃ fY (π (v))) −1 Y ∈ΠX Lemma 22. Π(M) is a constructive causal abstraction of M with τ = π Equivalently, for all i ∈ Domain(τ ), π(Solve(Mi )) = Solve(Π(M)π(i) ) i.e, solution sets are preserved under variable merge, for all interventions Value merge alters the value space of each variable, potentially collapsing values. Crucially, value merges are valid only if the partition cells respect the causal dynamics of M. Definition 23 (Viable δ for Value Merge). Consider some family δ = {δX }X∈V of functions δX ∶ Val(X) BX ; thus, δ is a function
Val(V) ⋃X∈V BX . We will say δ is viable for value merge when each δX fails to be injective only when the collapsed values play the same role in M. ′ ′ ′ More formally, suppose i, i ∈ Val(I) differ only in that x ∈ i and x ∈ i . Then we require that, δX (x) = δX (x ) implies δ(Solve(Mi )) = δ(Solve(Mi′ )). ′ Definition 24 (Value Merge). Choose some viable family of functions δX ∶ Val(X) BX Define δ to be the natural function applying this family of functions to partial settings. Define ∆(M) of M as follows. Every variable X has new values BX and a new function ∆(fX )(v) = δX (fX (Choose(δ (v)))) −1 CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION OTrue OFalse H(3,1) H(3,2) H(3,3) H(3,4) H(3,5) H(3,6) H(3,7) H(3,8) H(2,1) H(2,2) H(2,3) H(2,4) H(2,5) H(2,6) H(2,7) H(2,8) H(1,1) H(1,2) H(1,3) H(1,4) H(1,5) H(1,6) H(1,7) H(1,8) R1 R2 R3 R4 R5 R6 R7 R8 { } { } { } { } 31 O V1 I1 {
OTrue V2 I2 } { I3 } { { } O OFalse V1 H(2,2) I4 } H(2,3) H(2,7) V2 H(2,8) I1 R1 R2 R3 R4 R5 R6 R7 R8 { } { } { } { } { I2 }{ I3 }{ I4 }{ } Figure 2.5: An illustration of a fully-connected neural network being transformed into a tree structured algorithm by (1) marginalizing away neurons aligned with no high-level variable, (2) merging sets of variables aligned with high level variables, and (3) merging the continuous values of neural activity into the symbolic values of the algorithm. Observe that our viability condition on δ makes it so that the choice of the Choose function has no effect on the solutions of ∆(M) Lemma 25. ∆(M) is a constructive causal abstraction of M with τ = δ Equivalently, for all i ∈ Domain(τ ), δ(Solve(Mi )) = Solve(∆(M)δ(i) ) i.e, solution sets are preserved under value merge, for all interventions The notion of value merge also relates to an important concept in the philosophy of causation. The range of
values Val(X) for a variable X can be more or less coarse-grained, and some levels of grain seem to be better causal-explanatory targets. For instance, to use a famous example from Yablo [1992], if a bird is trained to peck any target that is a shade of red, then it would be misleading, if not incorrect, to say that crimson causes the bird to peck. Roughly, the reason is that this suggests the wrong counterfactual contrasts: if the target were not crimson (but instead, say, scarlet), the bird would still peck. Thus, for a given explanatory purpose, the level of grain in a model should guarantee that cited causes can be proportional to their effects [Yablo, 1992, Woodward, 2021]. As the following result shows, a low-level model being a constructive causal abstraction of a high-level model is simply a matter of being able to construct the high-level model from the low-level model with marginalization, variable merge, and value merge. Theorem 26 (Decomposing Constructive Abstraction). H
is a constructive abstraction of L if and CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION 32 only if H can be obtained from L by a variable merge, value merge, and marginalization. 2.7 Example of Causal Abstraction: Tree-Structure in Neural Computation In Section 2.4, we introduced the hierarchical equality task, highlighting the crucial role that equality reasoning has played in our understanding of how artificial and biological neural networks relate to modular symbolic algorithms. We then defined a tree-structured algorithm that solves the task, and a neural network that was trained to implement the tree structured algorithm [Geiger et al., 2022c]. Here, we show how causal abstraction theory can illuminate precisely what it means for this implementation relationship to hold between the network and algorithm. 2.71 An Alignment Between the Algorithm and the Neural Network Examining the neural network parameters in Figure 2.2 reveals no obvious relationship
between the network N and the algorithm A. However, the network N was explicitly constructed to be abstracted by the algorithm A under the alignment written formally below and depicted visually in Figure 2.3 ΠO = {OTrue , OFalse } ΠXk = {R2k−1 , R2k } ΠV2 = {H(2,7) , H(2,8) } Π⊥ = V (ΠO ∪ ΠV1 ∪ ΠV2 ∪ ΠX1 ∪ ΠX2 ∪ ΠX3 ∪ ΠX4 ) ⎧ ⎪ ⎪True τO (oTrue , oFalse ) = ⎪ ⎨ ⎪ ⎪ ⎪ ⎩False oTrue > oFalse otherwise ΠV1 = {H(2,2) , H(2,3) } ⎧ ⎪ □ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪D τXk (r2k−1 , r2k ) = ⎪ ⎨ ⎪ ⎪ ^ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩Undefined (r2k−1 , r2k ) = r□ (r2k−1 , r2k ) = rD (r2k−1 , r2k ) = r^ otherwise If there exists r ∈ {rD , r□ , r^ } such that {h(2,2) , h(2,3) } = Proj(Solve(Nr ), {H(2,2) , H(2,3) }), define 4 τV1 (h(2,2) , h(2,3) ) = Proj(Solve(Aτ (r )), V1 ) If there exists r ∈ {rD , r□ , r^ } such that {h(2,7) , h(2,8) } = Proj(Solve(Nr ), {H(2,7) , H(2,8) }, define 4 τV2 (h(2,7) , h(2,8)
) = Proj(Solve(Aτ (r )), V2 ) Otherwise leave τV1 and τV2 undefined. Consider an intervention i in the domain of τ . We have a fixed alignment for the input and output neurons, where i can have output values from the real numbers and input values from {rD , r□ , r^ } . The intermediate neurons are assigned high-level alignment by stipulation; i can only 4 CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION 33 have intermediate variables that are realized on some input intervention Mx for x ∈ {rD , r□ , r^ } . 4 Constructive abstraction will hold only if these stipulative alignments to intermediate variables do not violate the causal laws of A. 2.72 The Algorithm Abstracts the Neural Network For our running example, the inputs are a sequence of four shapes from the set {^, □, D}. Following Def. 17, the domain of τ is restricted to 3 input interventions, (3 ) single-source interchange 4 4 2 interventions for high-level interventions fixing either V1
or V2 , and (3 ) double-source interchange 4 3 interventions for high-level interventions fixing both V1 and V2 . This neural network was created using interchange intervention training with the alignment ⟨Π, τ ⟩ and the high-level model A, which is why the relation of constructive causal abstraction holds between the high-level model A and the low-level model N . Equivalently, for all i ∈ Domain(τ ) we have τ (Solve(Ni )) = Solve(Aτ (i) ) (2.4) 4 We provide code for the interchange intervention training and verifying that the network N is abstracted by the algorithm A. In Figure 24, we depict an aligned interchange intervention performed on A and N with the base input (D, D, ^, □) and a single source input (□, D, ^, ^). The central insight is that the network and algorithm have the same counterfactual behavior. Crucially, while our example contains a neural network trained to implement a symbolic algorithm, in practice researchers use interchange intervention
analysis to investigate whether a network implements an algorithm. 2.73 The Algorithm can be Constructed from the Neural Network Our constructive characterization provides a new lens through which to view this result. The network N can be transformed into the algorithm A through a marginalization, variable merge, and value merge. We visually depict the algorithm A being constructed from the network N in Figure 25 2.8 Approximate Abstraction and Interchange Intervention Accuracy Constructive causal abstraction is an all-or-nothing notion. Either the network and algorithm have identical (counterfactual) behavior for all inputs and interchange interventions, or they do not. This binary concept prevents us from having a graded notion of faithful interpretation that is more useful 4 https://github.com/atticusg/InterchangeInterventions CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION 34 in practice [Jacovi and Goldberg, 2020a]. Early applications of interchange
interventions avoided this theoretical limitation by finding subsets of the input space on which the causal abstraction holds in full force [Geiger et al., 2020b, 2021d] Later Geiger et al [2022c] proposed interchange intervention accuracy, which is simply the proportion of interchange interventions where the neural network and high-level algorithm have the same input–output behavior. Definition 27 (Interchange Intervention Accuracy). Consider a neural network N aligned to a high-level algorithm A with a partition Π and fixed input and output alignment τ . As in Def 16, define the intermediate semantics of τ through stipulation and restrict the domain of τ to multi-source 5 interchange interventions. We define the interchange intervention accuracy as follows: Out Out IIA(N , A, τ ) = Ei∼Uniform(Domain(τ )) [1[τ (Proj(Ni , XL )) = Proj(Aτ (i) , XH )]]. Interchange intervention accuracy is equivalent to behavioral accuracy if we further restrict τ to be defined only on
interchange interventions where base and sources are all the same single input. Interchange intervention accuracy is an intuitive metric, but it cannot be analyzed in terms of any existing notion of approximate abstraction. We define a new notion of approximate abstraction, α-on-average constructive abstraction, that is tightly connected to interchange intervention accuracy: Definition 28 (α-On-Average Causal Consistency). Fix a metric DistanceH between high-level total settings. An alignment ⟨Π, τ ⟩ between L and H is α-on-average causally consistent if the following holds for all i ∈ Domain(τ ) such that Solve(Li ) is non-empty: Ei∼Uniform(Domain(τ )) [DistanceH (τ (Solve(Li )), Solve(Hτ (i) ))] ≤ α That is, on average, the high-level intervention τ (i) results in a high-level total setting that is within α of the high-level total setting mapped to from the low-level total setting resulting from the low-level intervention i. We thus arrive at a definition of
approximate causal abstraction: Definition 29 (α-On-Average Constructive Abstraction). Fix a metric DistanceH on high-level total settings. H is a α-on-average constructive abstraction of L if there is an alignment ⟨Π, τ ⟩ between L and H satisfying the α-on-average causal consistency condition. Note that causal consistency (Def. 10) arises as a special case of Def 28 where α = 0, and thus constructive abstraction is likewise the limit case of 0-on-average constructive abstraction. 5 Depending on the application, it may be useful to consider distributions over interchange interventions other than the uniform distribution. CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION 35 Remark 30. A different notion of approximate abstraction is explored in work by Beckers et al [2019] where the maximum distance between high-level total settings is used rather than the average distance. Their notion of dmax -α-approximation amounts to a kind of ‘worst case’
analysis, which may be more appropriate in some contexts. The most general definition subsuming both dmax -α-approximation and α-on-average approximation would allow arbitrary distributions on interventions. The notion of α-on-average abstraction theoretically grounds the metric of interchange intervention accuracy (as defined by Geiger et al. 2022c) in causal abstraction Theorem 31. If a neural network N is aligned with ⟨Π, τ ⟩ to a high-level algorithm A and we have an interchange intervention accuracy of IIA(N , A, ⟨τ, Π⟩) = α, then A is a α-on-average constructive abstraction of N with DistanceH (vH , vH ) being equal to 0 if vH and vH are equal and 1 otherwise. ′ 2.9 ′ XAI Methods Grounded in Causal Abstraction To support the claim that causal abstraction can be taken as a general theoretical foundation for XAI, we will show that many popular XAI methods can be viewed as special cases of causal abstraction analysis. Explanation methods purely grounded in
the behavior of a model can be directly interpreted in causal terms when we interpret the act of providing an input to a model as an intervention to the input variable. Methods that learn interpretable models like LIME or SHAP only consider input–output behavior, so causal abstraction can accurately capture these simple methods using two-variable causal models. Causal mediation analysis and amnesiac probing intervene on a single intermediate neural representation, and can be captured with a three-variable high-level model. Only circuit-based explanations require a high-level model with more than three variables. Finally, integrated gradients is not a special case of causal abstraction analysis, but it can be used to compute interchange interventions and therefore can be used to conduct causal abstraction analysis. Causal abstraction can easily capture a variety of popular XAI methods with high-level models containing no more than three variables. However, the general theory supports
arbitrarily complex high-level causal models, highlighting exciting new directions for research. An one example, future XAI methods leveraging multi-source interchange interventions might be able to practically operationalize complex hypotheses such as whether a transformer model trained to sort numbers implements bubble sort (see Section 2.10) 2.91 LIME: Behavioral Fidelity as Approximate Abstraction by a TwoVariable Chain LIME is a popular XAI method that learns an interpretable model A that locally approximates the behavior of an uninterpretable model N on a neighborhood of inputs ψx around a single input x. CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION 36 The fidelity of the explainer model A is a measure of how the input-output behavior of A differs from that of N on the neighborhood of inputs ψx . Definition 32. LIME Fidelity of Interpretable Model Let A and N be models with identical input and output spaces. Define Distance(⋅, ⋅) to be a function
that computes the distance between outputs. LIME(N , A, ψx ) = 1 A N ∑ Distance(Proj(Solve(Nx ), XOut ), Proj(Solve(Ax ), XOut )) ∣ψx ∣ ′ x ∈ψx The uninterpretable model N will often be a deep learning model with fully-connected causal structure as shown below. N XIn N11 N12 ⋯ N1l N21 ⋮ N22 ⋯ ⋮ N2l ⋮ Nd1 Nd2 ⋯ Ndl N XOut The interpretable model A will often also have rich internal causal structure such as a decision tree model. For instance: A1 A XIn A4 A A2 XOut A3 However, LIME only guarantees a correspondence between the input–output behaviors of the interpretable and uninterpretable models. Therefore, representing both N and A as a two-variable causal models connecting inputs to outputs is sufficient to describe the fidelity measure in LIME. Theorem 33. LIME Fidelity is Approximate Abstraction Between Two Variable Chains A A N N Let N and A each be causal models with variables XIn , XOut and XIn , XOut , respectively, such A N A N
that Val(XIn ) = Val(XIn ) = ψx and Val(XOut ) = Val(XOut ). The following two statements are equivalent: 1. LIME(N , A, ψx ) ≤ α 2. A is an α-on-average abstraction of N with distance function Distance from LIME and an N N A = {X A alignment where ΠXIn = {XOut }, and τ is the identity function. In }, ΠXOut A XIn N XIn A XOut N XOut CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION 2.92 37 Causal Effect Estimation as Abstraction by a Two-Variable Chain What was the effect of some real-world concept on the prediction of deep learning model? The CEBaB benchmark [Abraham et al., 2022a] presents the estimation of such causal effects as an XAI task. Specifically, CEBaB evaluates explainer models on their ability to estimate the causal effect of changing the quality of food, service, ambiance, and noise in a real-world dining experience on the prediction of a sentiment classifier given a restaurant review as input data. We represent the real-world
data generating process and the neural network with a single causal model MCEBaB . The real-world concepts Cservice , Cnoise , Cfood , and Cambiance can take on three values +, −, and Unknown, the input data XIn , prediction output XOut , and neural representations Nij can take on real-valued vectors, and the two exogenous variables U and V that represent the real world noise. Cservice V Cnoise XIn U 6 Cfood Cambiance N11 N12 ⋯ N1l N21 ⋮ N22 ⋯ ⋮ N2l ⋮ Nd1 Nd2 ⋯ Ndl XOut If we are interested in the causal effect of food quality on model output, then we can marginalize away every variable other than the real-world concept Cfood and the neural network output XOut to get a causal model with two endogenous variables. This marginalized causal model is a high-level abstraction of MCEBaB that contains a single causal mechanism describing how food quality in a dining experience affects the neural network output. V XOut U Cfood 2.93 Causal Mediation As
Abstraction by a Three-Variable Chain ′ Suppose that changing the value of a variable X from x to x has an effect on a second variable Y . Causal mediation analysis determines how this causal effect is mediated by a third intermediate variable Z. The fundamental notions to mediation are total, direct, and indirect effects, which can be defined with interchange interventions. Definition 34. Total, Direct, and Indirect Effects Suppose we have a causal model M with sets of variables X, Y, Z ∈ M such that addition and subtraction are well-defined on values of Y. The total causal effect of changing the values of X from 6 Crucially, the models in causal effect estimation are probabilistic models. See Section 211 for details CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION 38 ′ x to x on Y is ′ TotalEffect(M, x x , Y) = Proj(Mx′ , Y) − Proj(Mx , Y) ′ The direct causal effect of changing the value of X from x to x on Y through mediator Z is ′
DirectEffect(M, x x , Y, Z) = Proj(Mx′ , Y) − Proj(Solve(MIntInv(x,⟨x′ ⟩,⟨Z⟩) , Y) ′ The indirect causal effect of changing the value of X from x to x on Y through mediator Z is ′ IndirectEffect(M, x x , Y, Z) = Proj(Mx′ , Y) − Proj(Solve(MIntInv(x′ ,⟨x⟩,⟨Z⟩) , Y) This method has been applied to the analysis of neural networks to characterize how the causal effect of inputs on outputs are mediated by intermediate neural representations. A central goal of such research is identifying sets of neurons that completely mediate the causal effect of an input value change on the output. This is equivalent to a simple causal abstraction analysis Theorem 35. Complete Mediation is Abstraction by a Three Variable Chain ′ Consider a neural network M with inputs x and x and a set of intermediate neurons H. Define In ΠA = X , ΠC = X Out , ΠB = H, and Π∅ = V (X ∪ H ∪ X In Out ). The following two statements are equivalent: ′ 1. The variables H
completely mediate the causal effect of changing x to x on the output: ′ IndirectEffect(M, x x , X Out ′ , H) = TotalEffect(M, x x , X Out ) 2. The high-level causal model Π(ϵΠ∅ (N )) has a structure such that A is not a child of C A B C Theorem 36. Partial Mediation is Approximate Abstraction by a Three Variable Chain ′ Consider a neural network H with inputs x and x and a set of intermediate neurons H. Define In ΠX = X , Π Y = X Out , ΠZ = H, and Π∅ = V (X ∪ H ∪ X In Out ). The following two statements are equivalent: ′ 1. On average, the variables H mediate a changing x to x by at least α: Ex,x′ ∼Uniform(Vals(XIn )) [DirectEffect(M, x x , X ′ Out , H)] ≤ α 2. If you edit the causal mechanism for the variable Y in Π(ϵΠ∅ (N )) to be FY (z, x) = Ex′ ∈Vals(X) [FY (z, x ))] ′ ′ CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION 39 the resulting model is an α-on-average abstraction of N where the
distance function is Distance(v, v ) = Proj(v, Y ) − Proj(v , Y ). ′ 2.94 ′ Iterative Nullspace Projection As Abstraction by a Three Variable Chain Iterative nullspace projection investigates how removing a concept C from a target hidden representation H of a neural network N affects the predictions of N . Specifically, the hidden representation is intervened upon to ‘remove’ a concept C by projecting the representation onto the nullspaces of linear probes L1 , . , Lk trained to predict the value of C We define L(⋅) to be the function computing this iterated projection. The idea is to conclude that N makes use of the concept C if the performance of N on a task degrades under these interventions. We can model iterative nullspace projection as abstraction by a three-variable causal model. A A Say there is a high-level model A with input variable XIn , output variable XOut , and a binary variable L that indicates whether information has been removed. When l = 1, the
causal mechanism A (x fXOut In , l) defines the desired input–output behavior for a neural model. When l = 0, the causal A (x mechanism fXOut In , l) defines degraded input-output behavior. A N A N Define the alignment between the neural model as follows. Let ΠXIn = {XIn }, ΠXOut = {XOut }, A A ΠL = H. Furthermore, let τXIn and τXOut be identify functions and define ⎧ ⎪ ⎪0 τL (h) = ⎪ ⎨ ⎪ ⎪ ⎪ ⎩1 ∃xIn h = L(Proj(NxIn , H)) ∃xIn h = Proj(NxIn , {N12 , N22 }) The high-level causal model A is an abstraction of the low-level neural model N under alignment ⟨Π, τ ⟩ exactly when the iterative nullspace projection removing the concept C resulted in degraded A . We depict an example of this below, where H = {N performance defined by fXOut 12 , N22 }. A N XIn A L XIn XOut N11 N12 ⋯ N1l N21 ⋮ N22 ⋯ ⋮ N2l ⋮ Nd1 Nd2 ⋯ Ndl N XOut Observe that the high-level model A does not have a variable encoding the concept C and the
CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION 40 values it might take on. Iterative nullspace projection attempts to determine whether a concept is used by a model, not characterize how that concept is used. Furthermore, iterative nullspace projection is not guaranteed to be a minimal intervention that removes any and all information about C, and therefore cannot causally implicate C in model behavior without further assumptions. First, projecting onto the nullspace of probes trained to predict the value of the concept C is an intervention that removes information about the value C that is linearly accessible. Deep learning models are highly non-linear, and therefore can make decisions using information that is not linearly accessible. Second, linear probes trained on C may also target other concepts that are correlated with C. These are not fatal flaws of the method. Elazar et al [2022] present several control tasks that may mitigate these concerns, and Lovering
and Pavlick [2022a] work around the presence of non-linear information about C by using a linear prediction function after the intervention site in the model they analyze. 2.95 Operationalizing Circuit-Based Explanations with Causal Abstraction Two fundamental claims made by Olah et al. [2020] are that (1) linear combinations of neural activations encode high-level concept(s) (e.g, curved lines or object orientation), and (2) the ‘circuits’ defined by the model weights connecting neural representations encode meaningful algorithms over high-level concepts. Causal abstraction helps shed additional lighton these claims, where a low-level causal model encodes neural representations and circuits, and a high-level causal model encodes high-level concepts and meaningful algorithms. Consider the following example of a circuit-based explanation with a neuron. Suppose we have a neural network N that takes in an image xIn and outputs three binary labels xcat , xdog , and xcar that
indicate whether the image contain a cat, dog, or car. N11 N12 ⋯ N1l N XIn N21 ⋮ N22 ⋯ ⋮ N2l ⋮ Nd1 Nd2 ⋯ Ndl N Xcat N Xdog N Xcar Further suppose that network was not trained on any images containing both dogs and cars. Because dogs and cars did not co-occur in training, we might expect N to contain a neuron that responds to both dogs and cars. We can formalize this hypothesis as a high-level causal model A with a binary variable A that takes on value 1 only when the input image contains a cat and a second binary variable B that is 1 only when the input image contains a dog or a car or both. These binary variables in turn determine the correct output labels. CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION 41 A Xcat A A A Xdog XIn B A Xcar If there is some alignment such that N is a constructive causal abstraction of A, then we know that the neurons aligned with the high-level variable B are respond to dogs and cars. This neuron is an
efficient solution to training data where dogs and cars aren’t in the same image. 2.96 Interchange Interventions from Integrated Gradients Integrated gradients [Sundararajan et al., 2017b] is a neural network analysis method that attributes neurons values according to the impact they have on model predictions. We can easily translate the original integrated gradients equation into our causal model formalism. Definition 37. Integrated Gradients Given a causal model N of a neural network, we define the integrated gradient value of the ith neuron of an intermediate neural representation y when the network is provided x as an input as ′ follows, where y is the so-called baseline value of Y ′ IGi (y, y ) = (yi − yi ) ⋅ ∫ ′ 1 ∂Proj(Nx∪(αy+(1−α)y′ ) , X ∂yi α=0 Out ) dα The completeness axiom of the integrated gradients method is formulated as follows: ∣Y∣ ∑ IGi (y, y ) = Proj(Nx∪y , X ′ Out ) − Proj(Nx∪y′ , X Out ) i=1 Integrated
gradients was not initially conceived as a method for the causal analysis of neural networks. Therefore, it is perhaps surprising that integrated gradients can be used to compute interchange interventions. This hinges on a strategic use of the ‘baseline’ value of integrated gradients, which is typically set to be the zero vector. Theorem 38. Integrated Gradients Can Compute Interchange Interventions The following is an immediate consequence of the completeness axiom ∣Y∣ Proj(Nb∪IntInv(b,⟨s⟩,⟨Y⟩) , X Out Out ) = Proj(Nb , X ) − ∑ IGi (Proj(Nb , Y), IntInv(b, ⟨s⟩, ⟨Y⟩)) i=1 In principle, we could perform causal abstraction analysis using the integrated gradients method, though computing integrals would be a highly inefficient way to compute interventions. CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION 2.10 42 Future Applications: Types, Infinite Variables, and Cycles Causal abstraction is a highly expressive, general purpose
framework. However, our review of existing XAI methods and our examples thus far have been limited to abstraction between causal models that are both finite and acyclic. To demonstrate the expressive capacity of causal abstraction to support arbitrary symbolic algorithms, we will precisely articulate under what conditions a recursive deep learning model that sorts sequences of arbitrary lengths implements a bubble sort algorithm. Bubble sort is an iterative algorithm. On each iteration, the first two members of the sequence are compared and swapped if the left element is larger than the right element, then the second and third member of the resulting list are compared and possibly swapped, and so on until no more swaps are needed. We define S to be a causal model as follows. Define the (countably) infinite variables and values j j j j V = {Xi , Yi , Bi ∶ i, j ∈ {1, 2, 3, 4, 5, . }} j Val(Xi ) = Val(Yi ) = {1, 2, 3, 4, 5, . } ∪ {∅} j Val(Bi ) = {True, False, ∅} The
causal structure of S is depicted in Figure 2.6a The ∅ value will indicate that a variable is not being used in a computation, much like a blank square on a Turing machine. The (countably) infinite 1 1 sequence of variables X1 , X2 , . contains the unsorted input sequence, where a input sequence of 1 1 length k is represented as the following partial setting. Define X1 , , Xk encodes the input sequence and set the infinitely many remaining variables Xj where j > k to ∅. j For a given row of variables j, the variables Bi store the truth-valued output of the comparison j of two elements, the variables Yi contain the values being ‘bubbled up’ through the sequence, and j the variables Xi are the result of j − 1 passes through the algorithm. When there are rows j and j j−1 j + 1 such that Xi and Xi take on the same value for all i, the output of the computation is the sorted sequence found in both of these rows. 1 We define the structural equations as
follows. The input variables Xi have constant functions 1 1 1 1 1 to ∅. The variable B1 is ∅ if either X1 or X2 are ∅, True if the value of X1 is greater than X2 , 1 1 1 1 1 1 and False otherwise. The variable Y1 is ∅ if B1 is ∅, X1 if B1 is True, and X2 if B1 is False The remaining intermediate variables can be uniformly defined for any natural numbers i, j where i > 0 or j > 0: j−1 ⎧ ⎪ xi+1 ⎪ ⎪ ⎪ ⎪ j−1 j−1 j−1 j−1 fXij (yi−1 , bi , xi+1 ) = ⎪ ⎨ yi−1 ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩∅ j−1 bi j−1 bi j−1 bi = True = False =∅ j ⎧ ⎪ xi+1 ⎪ ⎪ ⎪ ⎪ j j j j fYij (yi−1 , bi , xi+1 ) = ⎪ ⎨ yi−1 ⎪ ⎪ ⎪ ⎪ ⎪ ⎪ ⎩∅ j bi = True j bi = False j bi = ∅ CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION ⋮ ⋮ 3 X1 ⋮ 3 X3 3 X4 3 ⋯ 2 Y2 2 Y3 2 ⋯ 2 2 B1 2 X1 ⋯ B3 X2 2 X3 2 X4 2 ⋯ 1 Y2 1 Y3 1 ⋯ 1 1 B1 1 2 B2 Y1 X1 ⋮ X2 Y1 1
B2 1 X2 43 ⋯ B3 1 X3 1 X4 ⋯ (a) The causal structure of S that represents the bubble sort algorithm. ⋮ ⋮ ⋮ ⋮ X1 3 X2 3 X3 3 X4 3 . X1 2 X2 2 X3 2 X4 2 . 1 X2 1 X3 1 X4 1 . X1 (b) The causal structure of ϵB∪Y (S) that is a marginalization of the bubble sort algorithm S. X1 X2 X3 X4 . 1 X2 1 X3 1 X4 1 . X1 (c) The cyclic causal structure of Π(ϵB∪Y (S)) representing the equilibrium state of bubble sort. X1 X2 X3 X4 . 1 X2 1 X3 1 X4 1 . X1 (d) The causal structure of ∆(Π(ϵB∪Y (S))) representing the input-output behavior of bubble sort. Figure 2.6: A causal model representing the bubble sort algorithm (top) and abstractions of that model (bottom). CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION ⎧ j j ⎪ yi < xi+1 ⎪ j j ⎪ fBij (yi , xi+1 ) = ⎨ ⎪ ⎪ ⎪ ⎩∅ j 44 j yi = / ∅ and xi+1 = /∅ otherwise This causal model is countably infinite, supporting both
sequences of arbitrary length and an arbitrary number of sorting iterations. If a recursive deep learning model is abstracted by S, it is carrying out the described bubble sort algorithm. To conduct this analysis, one would compute the interchange intervention accuracy for each high-level variable aligned with a low-level neural representation. If we instead want to begin with coarser-grained XAI research questions, we can abstract this model to simplify our hypothesized high-level structure. Perhaps we want to know whether the deep learning model computes iterative passes of the bubble sort algorithm, but we are not concerned with how each pass of bubble sort is implemented. j j To articulate this hypothesis, we can marginalize away the variables B = {Bi } and Y = {Yi } and reason about the model ϵB∪X (S) instead (Figure 2.6b) Define the structured equations of this model j−1 j−1 j−1 j−1 recursively with base case fX1j (x1 , x2 ) = Minimum(x1 , x2 ) for j > 1 and
recursive case j−1 j−1 j−1 j−1 j−1 fXij (x1 , x2 , . , xi+1 ) = Minimum(xi+1 , Maximum(xi j−1 j−1 j−1 j (x , fXi−1 1 , x2 , . , xi ))) for i, j where i > 0 or j > 0. Suppose instead that we just want to know whether a deep learning model sorting an input sequence can be abstractly understood as the equilibrium point of a cyclic causal process. To articulate this hypothesis, we further abstract the causal model using variable merge with the j partition ΠXi = {Xi ∶ i, j ∈ {1, 2, 3, . } ∧ j > 1} The result is the model ϵB∪X (S) (Figure 26c), where each variable Xi takes on the value of an infinite sequence. There are causal connections to and from Xi and Xj for any i = / j because the infinite sequences stored in each variable must jointly be a valid run of the bubble sort algorithm. Maybe we simply want a deep learning model with the input–output behavior for the task of sorting sequences. We construct the needed high-level causal
model by using value-merge with a family of functions δ where the input variable functions δXi1 are identity functions and the other functions δXij for j > 1 output the constant value an infinite sequence converges to. The structural equations for the resulting model ∆(Π(ϵB∪Y (S))) (Figure 2.6d) simply map unsorted input sequences to sorted output sequences. Abstraction with this model can be verified through purely behavioral evaluation on the deep learning model, as abstraction will hold exactly when the deep learning model sorts all sequences correctly. j Finally, suppose that we want to encode into our notion of abstraction that the variables Xi and j j Yi all have the type of integer and the variables Bi have the type of Boolean. We might investigate whether a recursive deep learning model implements bubble sort with a uniform low-level encoding j j j of integer typed variables Xi and Yi and Boolean typed variables Bi . CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL
MODEL INTERPRETATION 2.11 45 Coda: Abstraction for Probabilistic Models Due to our focus on (deterministic) neural models, we have not needed to go beyond deterministic causal models. However, in many domains where abstractions are of interest, both low-level and highlevel models may be probabilistic How does the notion of abstraction, and specifically constructive abstraction, extend to the probabilistic setting? Whereas a deterministic model (Def. 4) is a pair (V, F), a probabilistic model will be a quadruple (V, U, F, P). In the following definition we roughly follow earlier work (eg, Bongers et al 2021) in formulating probabilistic models that may be cyclic: Definition 39. A probabilistic causal model is a quadruple M = (V, U, F, P), such that: 1. V is a set of ‘endogenous’ variables, with Val(V ) a measurable space for each V ∈ V; 2. U is a set of ‘exogenous’ variables, with Val(U ) a measurable space for each U ∈ U; 3. Each fV ∶ Val(V) × Val(U) Val(V ) is a
measurable function; 4. P is a probability measure on Val(U) A solution to a probabilistic model M is not a total setting, but a probability distribution on total settings. Specifically (again loosely following Bongers et al 2021): Definition 40 (Probabilistic Solution). A solution to M is a random variable V ∶ Ω Val(V), where Ω is some probability space, if there is a random variable U ∶ Ω Val(U) with the same distribution as P, such that the equation V = f (V, U) holds almost-surely (in Ω). Here we are writing f for ⋃V ∈V fV As in the deterministic case, solutions need not exist, and there may in general be multiple distributions associated with a given cyclic model (see, e.g, Bongers et al 2021, Ex 24) For the purpose of this discussion, we restrict attention to the so-called simple causal models as defined by Bongers et al. [2021] In addition to guaranteeing the existence and uniqueness of solutions (essentially by definition), these models are closed under
(probabilistic) marginalization. Let us write P M for the unique solution distribution corresponding to M. Also as in the deterministic case, probabilistic models are closed under interventions i ∈ Val(I). For simple models the solution to Mi is also always guaranteed to be unique; P interventional distribution (as opposed to the observational distribution P M Mi 7 is often called the corresponding to the unintervened model M). 7 Recall (Def. 8) that an intervention is a partial settings for some I ⊆ V, such that Mi is the same as M except that fX is replaced with v, u ↦ Proj(i, X) for each X ∈ I. CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION 46 Most existing treatments of causal abstraction for probabilistic models either ensure or explicitly require that the low-level model and the high-level model agree with respect to their interventional distributions. That is, suppose we define an alignment ⟨Π, τ ⟩ between (the endogenous variables
of) a low-level model L and (the endogenous variables of) a high-level model H, just as in Def. 10, now requiring that τ be a measurable function. Then, analogous to Def 10, we could require the same diagram to commute for every intervention i in Domain(τ ): i L P i τ τ (i) τ∗ P Hτ (i) where τ∗ (P i ) is the pushforward distribution through τ . Such a requirement would define a natural L extension of abstraction to the probabilistic setting; see, e.g, Beckers et al [2019], Rischel and Weichwald [2021], Otsuka and Saigo [2022], among others. Moreover, the probabilistic extension permits other approximation notions involving not a distance metric on the high-level settings, but rather a distance metric on probability distributions on high-level settings. For instance, Rischel and Weichwald [2021] advocate the use of Jensen-Shannon divergence 8 JSD for this purpose. Thus, we could define the degree to which H is a causal abstraction of L in terms of the maximum
divergence over all interventions i: L JSD(τ∗ (P i ), P Hτ (i) ), with commutativity the special case when this is equal to 0. As in Section 28, we could of course replace the maximum here with another function such as average (recall Remark 30); likewise, we could invoke distance metrics other than Jensen-Shannon. However, there is a sense in which (approximate) preservation of interventional distributions may be too weak a requirement (cf. Gresele et al 2022) Probabilistic causal models also give so called counterfactual distributions, which we can define in the following way. Definition 41 (Counterfactual Solution). Let I be a list of interventions and M a probabilistic causal model. A (counterfactual) solution to M under interventions I is a sequence of random variables Vi ∶ Ω Val(V), for a fixed probability space Ω across each i ∈ I, such that there is a random variable U ∶ Ω Val(U) with the same distribution as P, with all of the equations Vi = fi (Vi , U) 8
In addition, Rischel and Weichwald [2021] introduce distributions on low-level interventions corresponding to each high-level intervention. This idea was also explored by Beckers et al [2019] Such elaborations would be straightforward to incorporate in the present setting. CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION 47 simultaneously holding, almost-surely. Here fi is again the union over individual functions (as in Definition 40), but with the appropriate replacements as dictated by the intervention i. This gives us a counterfactual distribution P MI defined on the product space ∏i∈I Val(V), where each component corresponds to the values of V under a different intervention i (see, e.g, Ibeling and Icard 2021). When all variables are assumed to be discrete, the counterfactual distribution can be defined more simply (see, e.g, Pearl 2009, Bareinboim et al 2022) Counterfactual distributions include many important quantities. For instance, we might want to
know for a given patient in an experimental trial who was given a treatment and survived, what is the probability that they would not have survived had the experimenters withheld the treatment? Questions like this relate directly to questions about explanation. Indeed, we may want to know if the patient survived because of, in spite of, or irrespective of the treatment. It is well known that this so called ‘probability of necessity’in addition to all the other probabilities of causation [Pearl, 1999]may be underdetermined by observational and interventional probabilities. In fact, this is true ‘almost-always’ in both a topological [Ibeling and Icard, 2021] and a measure-theoretic [Bareinboim et al., 2022] sense Thus, insofar as we want causal abstractions to preserve causal explanations, we need to go beyond preservation of interventional probabilities. The following concrete example, inspired by examples from Pearl [2009] and Bareinboim et al. [2022], shows why preservation of
interventional probabilities is not enough: Example 1. Suppose we have a treatment X with two possible values (1 is treatment, 0 is no treatment), and outcome Y also with two possible values (1 is survival, 0 the opposite). Consider a ‘high-level’ probabilistic model H with these two endogenous variables, and two exogenous variables U1 and U2 both uniformly distributed on {0, 1}. Suppose that in H we have fX (u1 ) = u1 and fY (u2 ) = u2 ; in other words, X simply takes on the value of U1 and Y takes on the value of U2 , so there is no causal connection between X and Y . In other words, in H taking the treatment has no effect whatsoever on survival. Consider next a ‘low-level’ model L with the same endogenous variables X and Y , but now another endogenous variable Z for the presence of a certain gene (1 for present, 0 for absent). Suppose the exogenous variables are U1 and U3 , again both uniformly distributed on {0, 1}. In L we have fX (u1 ) = u1 and fZ (u3 ) = u3 , but the
structural function for Y now depends deterministically on X and Z. In fact, a patient survives just in case either they do not have the gene (Z = 0) and they take the treatment (X = 1), or they do have the gene (Z = 1) and they do not take the treatment (X = 0). In other words, fY (x, z) = xz + (1 − x)(1 − z) First, note that H would be a probabilistic abstraction of L if we only required the commuting diagram to hold for each intervention individually. For instance, letting ΠX = {X}, ΠY = {Y } and Π⊥ = {Z} and defining τ (i) = i to be the identity function for interventions i on (one or both of) X and Y , the diagram commutes. We essentially just ignore the low-level variable Z However, this is an undesirable result because the two models imply different explanations for CHAPTER 2. CAUSAL ABSTRACTION FOR FAITHFUL MODEL INTERPRETATION 48 possible outcomes. For instance, suppose that a patient takes the treatment and survives Model H implies that they would have survived
even without the treatmentso they did not survive because of the treatmentwhile L implies that they would not have survived had the treatment been withheldthat is, they survived precisely because of the treatment. To avoid such discrepancies, we propose: Definition 42 (Constructive Probabilistic Abstraction). H is a constructive probabilistic abstraction of L under an alignment ⟨Π, τ ⟩ if for all sets of interventions I ⊆ Domain(τ ), the following commutes: I L P I τ τ∗ τ (I) P Hτ (I) This ensures that pairs of models like those in Example 1 do not stand in the relation of abstraction. For instance, considering the joint probability of Y under the two contrasting interventions on X (i.e, X = 1 and X = 0) distinguishes the two models As above, the same relaxations of this strict notion would be natural, viz. approximation, etc Remark 43. The crucial feature of the the commuting diagram in Definition 42 is that it involves sets of interventions, in contrast to
Definition 10, which dealt only with single interventions. Because counterfactual quantities do in fact reduce to interventional quantities in the purely deterministic setting (see, e.g, Ibeling and Icard 2021), Definition 11 is indeed a special case of Definition 42 Naturally, as in the deterministic case one might consider alternative characterizations of this relation, e.g, closing under probabilistic versions of basic operations like marginalization, variable merge, etc. We leave such exploration for future work 2.12 Conclusion Causal abstraction is a theoretical framework for XAI that formalizes interpretable explanations with a graded notion of faithfulness. Constructive causal abstraction can be decomposed into the operations of marginalizing variables, merging variables, and merging values. Interchange interventions are an experimental technique grounded in causal abstraction that provides high-level interpretations for neural representations that possibly occur on some
training or testing input. A variety of popular XAI methods can be seen as special cases of causal abstraction analysis with simple high-level causal models. Causal abstraction lays useful groundwork for the future development of XAI methods that investigate complex algorithmic hypotheses about the internal reasoning of AI models. Chapter 3 Neural Natural Language Inference Models Partially Embed Theories of Lexical Entailment and Negation Abstract We address whether neural models for Natural Language Inference (NLI) can learn the compositional interactions between lexical entailment and negation, using four methods: the behavioral evaluation methods of (1) challenge test sets and (2) systematic generalization tasks, and the structural evaluation methods of (3) probes and (4) interventions. To facilitate this holistic evaluation, we present Monotonicity NLI (MoNLI), a new naturalistic dataset focused on lexical entailment and negation. In our behavioral evaluations, we find that
models trained on generalpurpose NLI datasets fail systematically on MoNLI examples containing negation, but that MoNLI fine-tuning addresses this failure. In our structural evaluations, we look for evidence that our top-performing BERT-based model has learned to implement the monotonicity algorithm behind MoNLI. Probes yield evidence consistent with this conclusion, and our intervention experiments bolster this, showing that the causal dynamics of the model mirror the causal dynamics of this algorithm on subsets of MoNLI. This suggests that the BERT model at least partially embeds a theory of lexical entailment and negation at an algorithmic level. 3.1 Introduction Natural Language Inference (NLI) keys into fundamental aspects of how people reason with language. Although NLI is generally cast in informal terms that embrace the indeterminacy of such reasoning, the task nonetheless manifests a number of very predictable reasoning patterns. For example, systematic manipulations of the
lexical meanings [Glockner et al., 2018], syntactic constructions [Nie 49 CHAPTER 3. MODELS EMBED THEORIES OF LEXICAL ENTAILMENT 50 et al., 2019a], and contextual assumptions [Pavlick and Callison-Burch, 2016] have systematic effects on the correct labels. These patterns present crisp, motivated learning targets that we can leverage to not only evaluate the ability of NLI models to learn robust solutions, but also to analyze the internal dynamics of successful models. In this paper, our learning target concerns the role of monotonicity in NLI [MacCartney, 2009, Icard and Moss, 2013]. Specifically, we would like to determine whether models can learn to represent lexical relations and accurately model that negation reverses entailment relations (e.g, dance entails move, but not move entails not dance). This property of negation is downward monotonicity In service of pursuing this question, we present Monotonicity NLI (MoNLI), a new naturalistic NLI dataset for training and assessing
systems on these semantic notions (Section 3.3) MoNLI extends SNLI [Bowman et al., 2015] to provide comprehensive coverage of examples that depend on lexical reasoning with and without negation. Using MoNLI, we conduct both behavioral and structural evaluations, seeking to provide a detailed picture of the solutions that top-performing models learn. We evaluate Enhanced Sequential Inference Models [Chen et al, 2016] and BERT-based models [Devlin et al., 2019b], along with standard baselines Previous work evaluating the ability of neural models to learn monotonicity has focused on challenge test sets and systematic generalization tasks [Yanaka et al., 2019a, 2020, Geiger et al, 2019b, Richardson et al., 2019] These behavioral evaluations ask whether models achieve a desired input–output behavior. We employ these methods as well, but we also ask whether models achieve an algorithmic-level learning target, in the terms of Marr [1982]. Monotonicity reasoning can be cast as an algorithm
that solves MoNLI perfectly. Do neural models implement this algorithm? We first report on two behavioral evaluations (Section 3.5) When MoNLI is used as a challenge test set, we find that models trained on SNLI and/or MNLI [Williams et al., 2018b] fail to reason with lexical entailments when negation is involved. However, we trace these failures to gaps in the training data. In response, we pose a systematic generalization task in which we expose models to MoNLI examples through fine-tuning while still requiring them to generalize to entirely new pairs of lexical items in negated linguistic contexts at test time. All our models solve the task, which suggests that they have learned general theories of lexical entailment and negation. We then report on structural evaluations (Section 3.6), seeking to determine whether our topperforming BERT-based models implement the target monotonicity algorithm In probing experiments, we find evidence consistent with this result, but it’s not
conclusive, since probes alone cannot reveal a model’s causal dynamics. However, our intervention experiments provide evidence that BERT does mirror the causal dynamics of the monotonicity algorithm, at least on large subsets of MoNLI. We conclude that this model at least partially embeds a theory of lexical entailment and negation at an algorithmic level, in addition to fully achieving the correct input–output behavior on MoNLI. CHAPTER 3. MODELS EMBED THEORIES OF LEXICAL ENTAILMENT 3.2 51 Related work Monotonicity Our empirical focus is entailment and negation. This is one (highly prevalent) aspect of monotonicity reasoning, which governs many aspects of lexical and constructional meaning in natural language [Sánchez-Valencia, 1991, van Benthem, 2008a]. There is an extensive literature on monotonicity logics [Moss, 2009, Icard, 2012, Icard and Moss, 2013, Icard et al., 2017] Within NLP, MacCartney and Manning [2007], MacCartney [2009] apply very rich monotonicity algebras
to NLI problems, Hu et al. [2019a,b] create NLI models that use polarity-marked parse trees, and Yanaka et al. [2019b,a] and Geiger et al [2019b] investigate the ability of neural models to understand natural logic reasoning. While we consider only a small fragment of these approaches, the methods we develop should apply to more complex systems as well. 1 Challenge Test Sets Challenge test sets are supplementary evaluation resources that test the ability of a model to generalize to examples outside the distribution of the data it was trained, developed, and (standardly) tested on. These tests probe the generalization capabilities of stateof-the-art models with respect to the tasks they have been trained on, by focusing on difficult or underrepresented examples in a model’s training set [Jia and Liang, 2017, Naik et al., 2018, Glockner et al., 2018, Richardson et al, 2019, Talmor et al, 2019] Systematic Generalization Tasks Fodor and Pylyshyn [1988] offer systematicity as a hallmark
of human cognition. Systematicity says that certain behaviors are intrinsically connected to others by compositional structures. For example, understanding the puppy loves Sandy is intrinsically connected to understanding Sandy loves the puppy. For Fodor and Pylyshyn, these observations trace to the mind’s ability to recombine known parts and rules. There are often strong intuitions that certain generalization tasks are only solved by models with systematic structures. These tasks are referred to as systematic generalization tasks [Lake and Baroni, 2018b, Hupkes et al., 2019, Yanaka et al., 2020, Bahdanau et al, 2018, Geiger et al, 2019b, Goodwin et al, 2020] Probing Probes are supervised learning models trained to extract information from representations created by another model. They are a primary tool in the analysis of neural network models (Peters et al. 2018, Tenney et al 2019, Clark et al 2019; for a full review, see Belinkov and Glass 2019a) In aggregate, this work has
provided nuanced insights into the internal representations of these models, as well as their capacity to directly support learning diverse NLP tasks via fine-tuning [Hewitt and Liang, 2019]. However, probes are only able to reveal how representations correlate with information They cannot determine if that information plays a causal role in model predictions [Belinkov and Glass, 2019a, Vig et al., 2020b] 1 Though adversarial and challenge are sometimes used synonymously, we opt for the term challenge, because our dataset was designed with the intention of evaluating whether a model learned a particular phenomenon, as opposed to breaking any particular model (cf. Nie et al 2019b) CHAPTER 3. MODELS EMBED THEORIES OF LEXICAL ENTAILMENT Interventions 52 Intervention studies go beyond probing to make changes to the internal states of a network, with the goal of observing how those changes affect system outputs. Giulianelli et al [2018a] use probing results to make informed
interventions during LSTM language model predictions to preserve information about the grammatical subject’s number, and this led to improved performance in subject–verb agreement. Vig et al [2020b] use interventions to characterize how gender bias is represented in the internal causal structure of a model, and find that a small number of synergistic neurons mediate gender bias. They also find that the effect of these neurons is roughly linearly separable from the effect of the remainder of the model, a remarkable finding considering the highly non-linear nature of neural networks. 3.3 Monotonicity NLI dataset 2 We created the MoNLI corpus to investigate the ability of NLI models to learn the compositional interactions between lexical entailment and negation. MoNLI contains 2,678 NLI examples in the usual format for NLI datasets like SNLI. In each example, the hypothesis is the result of substituting a single word wp in the premise for a hypernym or hyponym wh . We refer to wh
and wp as the substituted words in an example. In 1,202 of these examples, the substitution is performed under the scope of the downward monotone operator not. Downward monotone operators reverse entailment relations: dance entails move, but not move entails not dance. We refer to these examples collectively as NMoNLI. In the remaining 1,476 examples, this substitution is performed under the scope of no downward monotone operator. We refer to these examples collectively as PMoNLI MoNLI was generated according to the following procedure. First, randomly select a premise or hypothesis sentence s from the SNLI training dataset. Second, select a noun in s, and, using WordNet [Fellbaum, 1998], select all hypernyms and hyponyms of the noun subject to two conditions: (1) the hypernym or hyponym appears in the SNLI training data, and (2) substituting the hypernym ′ or hyponym results in a grammatical, coherent sentence s . Finally, for each substitution, generate two examples for the corpus
– one where the original sentence is the premise and the edited sentence is the hypothesis, and one example with those roles reversed. Each of these example pairs has one example with the label entailment and one example with the label neutral, resulting in a dataset perfectly balanced between the two labels. For example, suppose we select the SNLI sentence (A) and we identify the noun plants for substitution. Then we enter plants into WordNet and find that flowers is a hyponym of plants, so we substitute flowers for plants to create the edited sentence (B): (A) The three children are not holding plants. ⇓ (B) The three children are not holding flowers. 2 This dataset is publicly available at: https://github.com/atticusg/MoNLI CHAPTER 3. MODELS EMBED THEORIES OF LEXICAL ENTAILMENT Model Input pretraining NLI train data CBOW BiLSTM ESIM ESIM ESIM BERT BERT BERT GloVe GloVe GloVe GloVe SNLI train SNLI train SNLI train BERT BERT SNLI train No MoNLI fine-tuning SNLI PMoNLI
NMoNLI 78.9 81.6 87.9 – – 90.8 – – 64.6 73.2 86.6 – – 94.4 – – 53 With NMoNLI fine-tuning SNLI NMoNLI 22.9 37.9 39.4 – – 2.2 – – 65.9 74.6 56.9 – – 90.5 – – 95.5 93.5 96.2 98.0 35.5 90.0 96.7 62.3 Table 3.1: The results of our behavioral analysis The columns labeled No MoNLI fine-tuning display the challenge test set results (Section 3.51), and the columns labeled With MoNLI fine-tuning display systematic generalization task results (Section 3.52) The numbers are accuracy values; all the datasets have balanced label distributions. Dashes mark experiments that would involve untrained NLI parameters due to training/fine-tuning set-up. This leads to two new MoNLI examples: (A) entailment (B) (B) neutral (A) These two examples would belong to NMoNLI, due to not scoping over the substitution site. If not were removed from both of these sentences, then their labels would be swapped and both examples would belong to PMoNLI. MoNLI was generated by
the authors by hand; examples judged to be unnatural were removed, and any grammatical or spelling errors in the original SNLI sentence were corrected. This data generation process is similar to that of Glockner et al. [2018], except they focus on the lexical relations of exclusion and synonymy, while we focus on entailment relations. This difference prevents their dataset from capturing monotonicity reasoning, which involves entailment relations, but not exclusion or synonymy. 3.4 Models We evaluated four models on MoNLI: CBOW The continuous bag of words baseline from Williams et al. [2018b] BiLSTM The bidirectional LSTM baseline from Williams et al. [2018b] ESIM The Enhanced Sequential Inference Model [Chen et al., 2016] is a hybrid TreeLSTM-based and biLSTM-based model that uses an inter-sentence attention mechanism to align words across sentences. BERT A Transformer model trained to do masked language modeling and next-sentence prediction [Devlin et al., 2019b] We rely on
uncased BERT-base parameters from Hugging Face transformers [Wolf et al., 2019] CHAPTER 3. MODELS EMBED THEORIES OF LEXICAL ENTAILMENT 54 The first two models serve as baselines, while the other two models achieve comparable, near state-of-the-art scores on SNLI. 3.5 Behavioral Evaluations 3.51 MoNLI as a Challenge Test Set We first use MoNLI as a challenge test dataset, i.e, models trained only on SNLI are expected to generalize to MoNLI. MoNLI can be considered a challenge test dataset that evaluates an NLI model’s ability to perform simple inferences founded in lexical entailments and monotonicity. As discussed in Section 3.3, it is not especially adversarial, in that we sampled sentences from the SNLI training set and only substituted in hypernyms and hyponyms that occur in the SNLI training set. This keeps MoNLI as close as possible to the distribution of SNLI. Thus, if a model fails on MoNLI, we can be confident that this failure stems from a lack of knowledge about
monotonicity and lexical entailment relations, rather than some other confounding factor like syntactic structures or vocabulary items that were unseen in training. Results The results are in Table 3.1 under the heading ‘No MoNLI fine-tuning’, and they are stark The four models achieve comparably high accuracies on SNLI and PMoNLI, the examples where no downward monotone operators scope over the substitution site. However, they are well below chance accuracy on NMoNLI, the examples where not scopes over the substitution site. BERT is more extreme than the other models, achieving a higher accuracy on PMoNLI than SNLI and almost zero accuracy on NMoNLI. High performance on PMoNLI shows that models have knowledge of the lexical relations between the substituted words, but low performance on NMoNLI shows the models have no knowledge of the downward monotone nature of not. In fact, the below chance accuracy on NMoNLI indicates that these models are somewhat reliably (incredibly reliably
in BERT’s case) predicting the wrong label on these examples, suggesting that they treat NMoNLI examples the same as PMoNLI examples. Discussion While these models trained on SNLI do not know that not is downward monotone in these examples, this is not conclusive evidence that they are unable to learn this semantic property. This ability might not be necessary for success on SNLI, where only 38 examples have negation in both the premise and hypothesis. A natural next step is to train on MNLI, where the coverage with regard to negation is better: about 18K examples (≈4%) have negation in the premise and hypothesis. We tried this, by combining MNLI with SNLI, and the results were almost exactly the same. However, even the MNLI CHAPTER 3. MODELS EMBED THEORIES OF LEXICAL ENTAILMENT 55 examples might not manifest the kind of monotonicity reasoning that we are targeting. Our next experiments help to resolve this issue. 3.52 A Systematic Generalization Task Our three models
trained on SNLI have knowledge of the lexical relations between substituted words, but do not know that the presence of not reverses the relationship between the word-level relation and the sentence-level relation. We now conduct a behavioral evaluation to determine whether models are able to learn a general theory of lexical entailment and negation when exposed to a limited subset of NMoNLI during training. In designing systematic generalization tasks, we seek to constrain the training data in ways that prevent unsystematic models from succeeding. Defining disjoint train/test splits is enough to foil truly unsystematic models (e.g, simple look-up tables) However, building on much previous work [Lake and Baroni, 2018b, Hupkes et al., 2019, Yanaka et al, 2020, Bahdanau et al, 2018, Goodwin et al., 2020, Geiger et al, 2019b], we contend that a randomly constructed disjoint train/test split only diagnoses the most basic level of systematicity. More difficult systematic generalization
tasks will only be solved by models exhibiting more complex compositional structures. Specifically, we want our systematic generalization task to be solved only by models that compute lexical entailment relations that may be reversed by negation. A learning model that memorizes labels based on substituted word pairs and whether negation is present would succeed on a disjoint train and test set as long as all pairs of substituted words appear during training, and this model does not compute the lexical relation between word pairs. As such, we propose a generalization task where NMoNLI is partitioned into train and test sets such that the substituted words in the train set and the substituted words in the test sets are 3 disjoint. Ideally, a model trained on SNLI that is further trained on NMoNLI will still maintain strong performance on SNLI. We use inoculation by fine-tuning [Liu et al, 2019a] to evaluate models on this ability. We report on the inoculated model with the highest
average performance on SNLI test and NMoNLI test. The models are evaluated on examples where they know the relation between the substituted words, as evidenced by high performance on PMoNLI, but have not seen those substituted words in the presence of negation during training. However, they have seen other substituted words with the same relation in the presence of negation during training, making this task hard, but fair [Geiger et al., 2019b] To solve this harder generalization task, we believe a model must learn to reverse the lexical relation in general ; the identity of the substituted words must be abstracted away. 3 We use only NMoNLI in our systematic generalization task because models trained on SNLI already achieve high performance on PMoNLI. CHAPTER 3. MODELS EMBED THEORIES OF LEXICAL ENTAILMENT 56 Infer(MoNLIexample) 1 lexrel ← get-lex-rel(MoNLIexample) 2 if contains-not(MoNLIexample) 3 return reverse(lexrel ) 4 return lexrel Figure 3.1: An algorithm able to solve
the MoNLI dataset that provides a theoretically motivated learning target for neural models at an algorithmic level of analysis [Marr, 1982]. Infer takes in an example from MoNLI and outputs the relation between the premise and hypothesis. It uses three predefined functions. get-lex-rel returns the relation (one of {⊐, ⊏}) between the substituted words in the premise and hypothesis. contains-not returns true iff negation is present reverse maps ⊏ to ⊐ and vice-versa. Results and Discussion We present our results in Table 3.1, under the heading ‘With NMoNLI fine-tuning’ All of our models solve this generalization task. However, only BERT does so while maintaining high performance on SNLI. We also report ablation studies on our two non-baseline models, evaluating their performance on our systematic generalization task without training on SNLI and without any pretraining at all. We find that both models still succeed with no pretraining on SNLI, but fail with no pretraining
whatsoever. This suggests that BERT pretraining and GloVe vectors both provide sufficient information about lexical relations for the models to succeed. BERT’s ability to get slightly above chance performance with no pretraining indicates the presence of some statistical artifacts in our dataset [Gururangan et al., 2018] In sum, our models were able to solve our systematic generalization task, which we believe to be evidence that they learn to compute the lexical relations between substituted words. However, we also believe this evidence is weak, as there is no formal relationship between a model solving a generalization task and that model having any particular systematic internal structures. This evaluation is fundamentally behavioral, only concerning model inputs and outputs. We believe that a structural evaluation is necessary to conclusively evaluate systematicity. 3.6 Structural Evaluations In our behavioral evaluations, the learning target was to mimic the input–output
behavior defined by MoNLI. Assessing this learning target is straightforward We now report on structural evaluations to try to determine whether a neural model has particular internal dynamics. For this, we rely on very recent probing and intervention methodologies that are not yet well understood and must be tailored to the model being analyzed. As such, we choose to focus on a single model, namely, the BERT model from Section 3.5 fine-tuned on NMoNLI We chose BERT because it achieved exceptional CHAPTER 3. MODELS EMBED THEORIES OF LEXICAL ENTAILMENT Probes trained on representations of wp Probes trained on representations of [CLS] 100 100 90 lexrel acc. lexrel sel. Infer acc. Infer sel. 80 70 60 Accuracy/Selectivity Accuracy/Selectivity 90 50 40 30 20 10 0 57 lexrel acc. lexrel sel. Infer acc. Infer sel. 80 70 60 50 40 30 20 10 1 2 3 4 5 6 7 8 9 10 11 12 Representation Row # 0 1 2 3 4 5 6 7 8 9 10 11 12 Representation Row # Probes trained on representations of
wh 100 Accuracy/Selectivity 90 lexrel acc. lexrel sel. Infer acc. Infer sel. 80 70 60 50 40 30 20 10 0 1 2 3 4 5 6 7 8 9 10 11 12 Representation Row # Figure 3.2: Results where classifier probes are trained on BERT representations to predict the value of lexrel and the output of Infer (Figure 3.1) Selectivity is probe accuracy minus control probe accuracy [Hewitt and Liang, 2019]. The grey dotted line provides a soft ceiling for selectivity values, because we expect control probes trained on a binary task to at least achieve chance accuracy. CHAPTER 3. MODELS EMBED THEORIES OF LEXICAL ENTAILMENT 58 results on NMoNLI after fine-tuning without experiencing a significant drop on SNLI. Figure 3.1 presents the simple algorithm Infer, which is our learning target It takes in a MoNLI example and stores the lexical entailment relation between the substituted words in the variable lexrel . If negation is present, the reverse of lexrel is returned; if there is no negation, lexrel
itself is returned. This is simply an algorithmic description of the MoNLI construction method The most important piece is the intermediate variable lexrel . Intuitively, if our BERT model implements this algorithm, there will be some representation in BERT that stores lexrel and BERT will use that representation for a final prediction. Probes can give us an idea of where information is stored, and interventions help us see how that information is used. Before we can go looking for where BERT stores and uses lexrel , we must limit ourselves to a tractable number of model internal representations. When our BERT model processes an example from MoNLI, it is tokenized as e = ⟨[CLS], p, [SEP], h, [SEP]⟩ and 12 rows of vector representations are created, so each token is associated with 12 vectors. We localize our efforts to the representations created for [CLS] and the tokens for the substituted words in the premise and hypothesis, wp and wh (as described in Section 3.3) This narrows
our search to 36 possible vector locations where BERT could be storing the variable lexrel for use in final output r r r prediction. We denote these 36 locations with BERTwp , BERTwh , and BERT[CLS] where r is a row (1 ⩽ r ⩽ 12). 3.61 Probes We follow Hupkes et al. [2018a] in using probing evidence to determine whether a neural model stores the same information as a symbolic algorithm. They used probes to predict variable values used in an algorithm from the hidden states of sequential recurrent networks trained to perform basic r r arithmetic. We do something similar, probing the 36 vector locations defined by BERTwp , BERTwh , r and BERT[CLS] for the value of the variable lexrel and the output of Infer. Hewitt and Liang [2019] argue that accuracy is a poor metric for probes and that the ideal probe will highly selective, that is, it will have high accuracy on a linguistic task but low accuracy on a control task where inputs are given random labels. In this setting, our
linguistic tasks are predicting the value of lexrel and the output of Infer from a model-internal vector created by BERT for some MoNLI example. Our control task is identical, except labels are randomly assigned to inputs Hewitt and Liang demonstrate that small, linear probes result in high selectivity. Following this guidance, we used a linear classifier with 4 hidden units that was trained and evaluated on all of MoNLI. Our probing results are summarized in Figure 3.2 Probes were able to achieve high accuracy and k high selectivity predicting the output of Infer at every location other than the locations BERT[CLS] where 1 ≤ k ≤ 4, and high accuracy and high selectivity predicting the value of lexrel at every location CHAPTER 3. MODELS EMBED THEORIES OF LEXICAL ENTAILMENT 1 59 2 other than BERT[CLS] and BERT[CLS] . This qualitative picture is compatible with a story where BERT stores the value of lexrel at 1 2 any location other than BERT[CLS] or BERT[CLS] and then uses
this information to compute a final k output prediction at any location other than the locations BERT[CLS] where 1 ≤ k ≤ 4. The fact 3 4 that probes trained on the vectors at locations BERT[CLS] or BERT[CLS] have high accuracy and selectivity predicting the value of lexrel , but moderate accuracy and low selectivity predicting the output of Infer may suggest a more specific story where these two locations store the value of the variable lexrel before this information is used to compute the final output. We emphasize that, while the probing results are compatible with these stories, they only provide conclusive evidence about how representations correlate with the value of lexrel and the output of Infer. They cannot determine whether this information plays a causal role in model predictions [Belinkov and Glass, 2019b, Vig et al., 2020b] 3.62 Interventions Probes give us a picture of where information is stored by our BERT model, but they cannot determine whether that
information is used to make final predictions. Interventions can help us address this deeper question. As discussed above, our algorithmic-level learning target is for BERT to mimic the dynamics of the algorithm Infer in Figure 3.1 Icard [2017b] provided the insight that algorithms like Infer can be explicitly understood as causal models [Pearl, 2001]. This means that the causal role of lexrel , the lone variable in Infer, can be characterized with counterfactual claims about how altering the value of the variable would cause output behavior to change. Suppose Infer is run on a MoNLI example i. Let lexrel (i) ∈ {⊐, ⊏} be the value that lexrel takes on, and let infer(i) ∈ {⊐, ⊏} be the output. Then Infer can be see as providing the following counterfactual characterization of lexrel : if the value of lexrel were changed from lexrel (i) to lexrel (j), where j is a second MoNLI example, then infer(i) would change to Inferlexrel(i)lexrel(j) (i) = ⎧ ⎪ Infer(i) ⎪ ⎪ ⎨
⎪ ⎪ ⎪ ⎩reverse(Infer(i)) lexrel (i) = lexrel (j) lexrel (i) = / lexrel (j) In other words, if lexrel were to take on the opposite value, then the output would also take on the opposite value. Our analytic tool for evaluating whether such causal dynamics are present in BERT is the interchange intervention. Figure 33 provides a high-level picture of how these experiments work, and the following definition seeks to make this more precise and general: CHAPTER 3. MODELS EMBED THEORIES OF LEXICAL ENTAILMENT 60 entailment ⊐ this not tree [SEP] this not elm dog runs not elm entailment ⊏ a pug runs [SEP] a neutral ⊏ this not tree [SEP] this Figure 3.3: An illustrative interchange intervention: The solid arrows represent a hypothesis about where the model stores and uses information about lexical entailment. The dotted arrow is an interchange intervention, where the green vector (top) we think stores reverse entailment, trees ⊐ elms, is interchanged
with the red vector (middle) we think stores forward entailment, pugs ⊏ dogs, leading to a modified network (bottom). If our hypothesis is correct, then the output should change from entailment to neutral, because the negation in the green example reverses the relationship between lexical entailment and sentence-level entailment. If this label reversal is not observed, crucial entailment information must lie elsewhere in the network. r r Interchange Intervention Let L be one of the 36 locations defined by BERTwp , BERTwh , and r BERT[CLS] . When BERT is making a prediction for i, suppose that the vector created at location L on input i is replaced with the vector created at location L on input j and this results in the output y. We say that y is the result of an interchange intervention from i to j at location L and denote this output as BERTL(i)L(j) (i). In essence, BERTL(i)L(j) (i) characterizes the output behavior that results from an experiment where model-internal vectors are
interchanged at location L. Recall that Inferlexrel(i)lexrel(j) (i) describes what output is provided by Infer if variables are interchanged. If for some subset of MoNLI S, we believe that BERT is both storing the value of lexrel at some location L and using that information to make a final prediction, then for all i, j ∈ S the following should hold: Inferlexrel(i)lexrel(j) (i) = BERTL(i)L(j) (i) This amounts to observing that the variables in the algorithm and the vectors in the model satisfy the same counterfactual claims. When a vector representing forward entailment is interchanged with CHAPTER 3. MODELS EMBED THEORIES OF LEXICAL ENTAILMENT 61 a different vector representing forward entailment, model output behavior should be unchanged. If a vector representing forward entailment is interchanged with a different vector representing reverse entailment, then the model output should be reversed. Results Due to computational constraints, we randomly conducted interchange
experiments at our 3 36 different locations and chose the location with the most promise, namely, BERTwh . We conducted ≈7 million interchange experiments at this location, one experiment for every pair of examples in MoNLI. Using a simple greedy algorithm, we discovered several large subsets of MoNLI where BERT mimics the causal dynamics of Infer. These subsets have size 98, 63, 47, and 37, and for each of these subsets there are many pairs of examples with interchange experiments that had a causal impact on the final model prediction. To put these results in context, if interchange experiments had a random effect on model output, then the expected number of subsets larger than 20 with this −8 property would be less than 10 . Discussion These results show that the values assigned by the algorithm Infer to the variable 3 lexrel and the vectors created by BERT at the location BERTwh exhibit the same causal dynamics on four large subsets of MoNLI. These pairs contain only 13 of
the 69 distinct hyponyms in MoNLI, which makes it clear that this subset of MoNLI is not a random sample, but rather reflects a coherent semantic space. From this we conclude that, in addition to capturing the input–output behavior described by MoNLI, our BERT model at least partially embeds a theory of lexical entailment and negation at an algorithmic level of analysis. Importantly, these results do not show that BERT fails to mimic the causal dynamics of Infer on larger subsets of MoNLI. First, we only conducted interchange experiments for every pair of 3 examples in MoNLI at the location BERTwh . Second, we did not consider the possibility that BERT stores and uses the value of lexrel at different locations, depending on which input is provided. Third, analyzing vector representations may be too coarse-grained; perhaps experiments will need to be done on individual vector units. Finally, we used a greedy algorithm to discover the four subsets of MoNLI. We did not exhaustively
analyze BERT to find the largest subset of MoNLI on which it mimics the causal dynamics of Infer; such an analysis is likely computationally impossible. What we did do is perform an efficient analysis that was able to find several large subsets of MoNLI on which the desired causal dynamics are present. 3.7 Conclusion To operationalize our research question of whether neural NLI models can learn the compositional interactions between lexical entailment and negation, we constructed two learning targets for neural NLI models: (1) learn the input–output behavior described by MoNLI and (2) acquire the internal dynamics of the algorithm Infer. We evaluated the first learning target with two behavioral CHAPTER 3. MODELS EMBED THEORIES OF LEXICAL ENTAILMENT 62 evaluation methods, using challenge datasets to show that state-of-the-art models trained on generalpurpose NLI datasets fail to exhibit the correct behavior when negation is present and then following up with a systematic
generalization task that showed our models are able to learn the correct input–output behavior when fine-tuned on a limited, but sufficient, subset of NMoNLI. We evaluated the second learning target with two structural evaluation methods, using probes to investigate where information about the variable lexrel from Infer might be stored in a BERT model and using interventions to show that on some subsets of MoNLI our BERT model exhibits the same causal dynamics as the algorithm Infer. We believe that our holistic evaluation, leveraging both behavioral and structural methods, provides a multifaceted picture of how neural NLI models treat lexical entailment and negation. While our interchange intervention methodology is not yet formally grounded, there is great promise in the idea of investigating whether a neural model mirrors the causal dynamics of an algorithm. Chapter 4 Causal Abstractions of Neural Networks Abstract Structural analysis methods (e.g, probing and feature
attribution) are increasingly important tools for neural network analysis. We propose a new structural analysis method grounded in a formal theory of causal abstraction that provides rich characterizations of model-internal representations and their roles in input/output behavior. In this method, neural representations are aligned with variables in interpretable causal models, and then interchange interventions are used to experimentally verify that the neural representations have the causal properties of their aligned variables. We apply this method in a case study to analyze neural models trained on Multiply Quantified Natural Language Inference (MQNLI) corpus, a highly complex NLI dataset that was constructed with a tree-structured natural logic causal model. We discover that a BERT-based model with state-of-the-art performance successfully realizes parts of the natural logic model’s causal structure, whereas a simpler baseline model fails to show any such structure,
demonstrating that BERT representations encode the compositional structure of MQNLI. 4.1 Introduction Explainability and interpretability have long been central issues for neural networks, and they have taken on renewed importance as such models are now ubiquitous in research and technology. Recent structural evaluation methods seek to reveal the internal structure of these “black box” models. Structural methods include probes, attributions (feature importance methods), and interventions (manipulations of model-internal states). These methods can complement standard behavioral techniques (e.g, performance on gold evaluation sets), and they can yield insights into how and why models make the predictions they do. However, these tools have their limitations, and it has often 63 CHAPTER 4. CAUSAL ABSTRACTIONS OF NEURAL NETWORKS 64 been assumed that more ambitious and systematic causal analysis of such models is beyond reach. Although there is a sense in which neural networks
are “black boxes”, they have the virtue of being completely closed and controlled systems. This means that standard empirical challenges of causal inference due to lack of observability simply do not arise. The challenge is rather to identify high-level causal regularities that abstract away from irrelevant (but arbitrarily observable and manipulable) low-level details. Our contribution in this paper is to show that this challenge can be met. Drawing on recent innovations in the formal theory of causal abstraction [Beckers and Halpern, 2019b, Beckers et al., 2019, Chalupka et al, 2016b, Rubenstein et al, 2017a], we offer a methodology for meaningful causal explanations of neural network behavior. 1 Our methodology causal abstraction analysis consists of three stages. (1) Formulate a hypothesis by defining a causal model that might explain network behavior. Candidate causal models can be naturally adapted from theoretical and empirical modeling work in linguistics and cognitive
sciences. (2) Search for an alignment between neural representations in the network and variables in the high-level causal model. (3) Verify experimentally that the neural representations have the same causal properties as their aligned high-level variables using the interchange intervention method of Geiger et al. [2020a] As a case study, we apply this methodology to LSTM-based and BERT-based natural language inference (NLI) models trained on the logically complex Multiply Quantified NLI (MQNLI) dataset of Geiger et al. [2019a] This challenging dataset was constructed with a tree-structured natural logic causal model [MacCartney and Manning, 2007, van Benthem, 2008b, Icard and Moss, 2013]. Our BERT-based model has the structure of a standard NLI classifier, and yet it is able to perform well on MQNLI (88%), a result Geiger et al. achieved only with highly customized task-specific models By contrast, our LSTM-based model is much less successful (46%). The obvious scientific question in
this case study is what drives the success of the BERT-based model on this challenging task. To answer this we employ our methodology (1) We formulate hypotheses by defining simplified variants of the natural logic causal model. (2) We search over potential alignments between neural representations in BERT and variables in our high-level causal models. (3) We perform interchange interventions on the BERT model for each alignment We find that our BERT model partially realizes the causal structure of the natural logic causal model; crucially, the LSTM model does not. High-level causal explanation for system behavior is often considered a gold standard for interpretability, one that may be thought quixotic for complex neural models [Lillicrap and Kording, 2019]. The point of our case study is to show that this high standard can be achieved. We conclude by comparing our methodology to probing and the attribution method of integrated gradients [Sundararajan et al., 2017a] We argue probing
is unable to provide a causal characterization of models. We show formally that attribution methods do measure causal properties, and in that 1 We provide tools for causal abstraction analysis at http://github.com/hansonhl/antra and the code base for this paper at http://github.com/atticusg/Interchange CHAPTER 4. CAUSAL ABSTRACTIONS OF NEURAL NETWORKS 65 way they are similar to the tool of interchange interventions. However, our methodology of causal abstraction analysis provides a framework for systematically measuring and aggregating such causal properties in order to evaluate a precise hypothesis about abstract causal structure. 4.2 Related Work Probes Probes are generally supervised models trained on the internal representations of networks with the goal of determining what those internal representations encode [Clark et al., 2019, Hupkes et al., 2018a, Peters et al, 2018, Tenney et al, 2019] Probes are fundamentally unable to directly measure causal properties of neural
representations, and Ravichander et al. [2020], Elazar et al [2020], and Geiger et al. [2020a] have argued that probes are limited in their ability to provide even indirect evidence of causal properties. We now present an analytic example in which probing identifies seemingly crucial information in representations that have no causal impact on behavior. We assume the structure of the simple addition network N+ in Figure 4.1 For our embedding, we simply map every integer i in N9 to the 1-dimensional vector [i]. The weight matrices are ⎛ 1 ⎞ ⎟ ⎜ ⎟ W1 = ⎜ ⎜ ⎜ 1 ⎟ ⎟ ⎝ 0 ⎠ ⎛ 1 ⎞ ⎟ ⎜ ⎟ W2 = ⎜ ⎜ ⎜ 1 ⎟ ⎟ ⎝ 1 ⎠ ⎛ 0 ⎞ ⎟ ⎜ ⎟ W3 = ⎜ ⎜ ⎜ 0 ⎟ ⎟ ⎝ 1 ⎠ ⎛ 0 ⎞ ⎟ ⎜ ⎟ w=⎜ ⎜ ⎜ 1 ⎟ ⎟ ⎝ 0 ⎠ The output for an input sequence x = (i, j, k) is given by (xW1 ; xW2 ; xW3 ) w. In this network, xW1 perfectly encodes i + j, and xW3 perfectly encodes k. Thus, the identity model probe will be perfect in probing those
representations for this information. However, neither representation plays a causal role in the network behavior; only xW2 contributes to the output. Attribution Methods Attribution methods aim to quantify the degree to which a network representation contributes to the output prediction of the model, for a specific example or set of examples [Binder et al., 2016, Shrikumar et al, 2016, Springenberg et al, 2014, Sundararajan et al, 2017a, Zeiler and Fergus, 2014a]. In contrast to probing, the well known integrated gradients method (IG) can be given an unambiguous causal interpretation. Following Sundararajan et al [2017a] we define the vector IG(x), for an input x relative to a baseline b, to have ith component IGi (x) given by the expression on the left: ∂F (αx + (1 − α)b) dα ∂xi α=0 (xi − bi ) ⋅ ∫ 1 = (xi − bi ) ⋅ ∫ lim α=0 ϵ0 Abbreviating the weighted average αx + (1 − α)b by x , letting x α F (x α,ϵ 1 α,ϵ ) − F (x ) dα ϵ α be the
vector that differs from α x in that the ith coordinate is increased by ϵ, and then expanding the definition of partial derivative, this can be written in the form given on the right. The difference F (x α,ϵ ) − F (x ) is known in α CHAPTER 4. CAUSAL ABSTRACTIONS OF NEURAL NETWORKS 66 the causal literature as the (individual) causal effect on the output (e.g, Imbens and Rubin [2015]) of increasing neuron i by ϵ relative to the fixed input x . So, essentially, IGi (x) is measuring the α average “limiting” causal effect of increasing neuron i along the straight line from the baseline vector to the input vector x, weighted by the difference at i between input and baseline. More recently, Chattopadhyay et al. [2019] develop an attribution method that explicitly treats neural models as structured causal models and directly computes the individual causal effect of a feature to determine its attribution. Attribution methods can measure causal properties, and, in that
way, they are similar to the tool of interchange interventions. However, our methodology of causal abstraction analysis provides a framework for systematically measuring and aggregating such causal properties in order to evaluate a precise hypothesis about abstract causal structure. Causal Abstraction Our goal is to evaluate whether the internal structure of a neural network realizes an abstract causal process. To concretize this, we turn to formal, broadly interventionist theories of causality [Spirtes et al., 2000, Pearl, 2001], in which causal processes are characterized by effects of interventions, and theories of abstraction [Beckers and Halpern, 2019b, Beckers et al., 2019, Chalupka et al., 2016b, Rubenstein et al, 2017a] where relationships between two causal processes are determined by the presence of systematic correspondences between the effects of interventions. The notion of abstraction that we employ here is a relatively simple one called constructive abstraction [Beckers
and Halpern, 2019b]. Informally, a high-level model is a constructive abstraction of a low-level model if there is a way to partition the variables in the low-level model where each high-level variable can be assigned to a low-level partition cell, such that there is a systematic correspondence between interventions on the low-level partition cells and interventions on the high-level variables. There are two properties of constructive abstraction that make it ideal for neural network analysis. First, the information content of partition cells of low-level variables can be determined by the high-level variables that they correspond to. For neural networks, the partition cells of low-level variables are sets of neurons, and our method supports reasoning at the level of vector representations (sets of neurons). Second, the causal dependencies between partitions of low-level variables are not necessarily preserved as causal dependencies between the high-level variables corresponding to
these partitions. For example, the low-level model might be a fully connected neural network, whereas the high-level model might have much sparser connections. For neural network analysis, this means we can find causal abstractions that have far simpler causal structures than the underlying neural networks. We provide an example in the next section CHAPTER 4. CAUSAL ABSTRACTIONS OF NEURAL NETWORKS 4.3 67 Causal Abstraction Analysis of Neural Networks We now describe our methodology in more detail, illustrating the relevant concepts with an example of a neural network performing basic arithmetic. Specifically, suppose that we have a neural network N+ that takes in three vector representations Dx , Dy , Dz representing the integers x, y, and z, and outputs the sum of the three inputs: N+ (Dx , Dy , Dz ) = x + y + z. We seek an informative causal explanation of this network’s behavior. Formulating a Hypothesis A human performing this task might follow an algorithm in which they
add together the first two numbers and then add that sum to the third number. We can hypothesize that the behavior of N+ is explained by this symbolic computation. Specifically, the network combines Dx and Dy to create an internal representation at some location L1 encoding x + y; it encodes z at some location L2 ; and L1 and L2 are composed to encode a + z at the location of the output representation. This hypothesis is given schematically in Figure 41a Following our methodology, we first define the causal model C+ in Figure 4.1a Our informal hypothesis that a neural network’s behavior is explained by a simple algorithm can then be restated more formally: C+ is a constructive abstraction of the neural network N+ . Alignment Search Now that we have hypothesized that the causal model C+ is a causal abstraction of the network N+ , the next step is to align the neural representations in N+ with the variables in C+ . The input embeddings Dx , Dy , and Dz must be aligned with the input
variables X, Y , and Z and the output neuron O must be aligned with the output variable S2 . That leaves the intermediate variables S1 and W to be aligned with neural representations at some undetermined locations L1 and L2 . If this were an actual experiment (see below), we would perform an alignment search to consider many possible values for L1 and L2 . Each alignment is a hypothesis about where the network N+ stores and uses the values of S1 and W . For the example, we assume the alignment in Figure 41a Interchange Interventions Finally, for a given alignment, we experimentally determine whether the neural representations at L1 and L2 have the same causal properties as S1 and W . The basic experimental technique is an interchange intervention, in which a neural representation created during prediction on a “base” input is interchanged with the representation created for a “source” input [Geiger et al., 2020a] We now show informally that this method can be used to prove
that the causal model C+ is a constructive abstraction of the neural network N+ . We first intervene on the causal model. Consider two inputs a, a ∈ (N9 ) where N9 is the set of ′ 3 integers 0–9. Let a = (x, y, z) be the base input and a = (x , y , z ) be the source input Define ′ ′ S ←a C+ 1 ′ (a) = x + y + z ′ ′ ′ ′ (4.1) to be the output provided by C+ when S1 , the variable representing the intermediate sum, is CHAPTER 4. CAUSAL ABSTRACTIONS OF NEURAL NETWORKS 68 intervened on and set to the value x + y . Thus, for example, if the base input is C+ (1, 2, 3) = 6, and ′ ′ ′ S ←a the source input is a = (4, 5, 6), then C+ 1 ′ (1, 2, 3) = 4 + 5 + 3 = 12. This process is depicted in Figure 4.1c Next, we intervene on the neural network N+ . Let D be an embedding space that provides unique representations for N9 , and consider two inputs D = (Dx , Dy , Dz ) and D = (Dx′ , Dy′ , Dz′ ), where ′ all Di and Di′ are drawn from D.
In parallel with (41), define L ←D N+ 1 ′ (D) (4.2) to be the output provided by N+ processing the input D when the representation at location L1 is ′ replaced with the representation at location L1 created when N+ is processing the input D . This process is depicted in Figure 4.1b With these two definitions, we can define what it means to test the hypothesis that N+ computes ′ x + y at position L1 . Where Da is an embedding for a and Da′ is an embedding for a , we test: ′ S ←a C+ 1 L ←Da′ (a) = N+ 1 (Da ) (4.3) ′ If this equality holds for all source and base inputs a and a , then we can conclude that, for every intervention on S1 , there is an equivalent intervention on L1 . If we can establish a corresponding claim for W and L2 , then we have shown that C+ is a constructive abstraction of N+ , since the inputs’ relationships are established by our embedding and there are no other interventions on C+ to test. Analysis Suppose that all of our
intervention experiments verify our hypothesis that C+ is a constructive abstraction of N+ with variables S1 and W aligned to neural representations at L1 and L2 . This explains network behavior by resolving two crucial questions First, we learn what information is encoded in the representations L1 and L2 . Neural representations encode the values of the high-level variables they are aligned with The location L1 encodes the variable S1 and the location L2 encodes the variable W . This is similar to what probing achieves. However, our method is crucially different from probing In probing, information content is established through purely correlational properties, meaning a neural representation with no causal role in network behavior can be successfully probed, as we showed in Section 4.2 In causal abstraction analysis, information content is established through purely causal properties, ensuring that the neural representation is actually implicated in model behavior. Second, we learn
what causal role L1 and L2 play in network behavior. Neural representations play a parallel causal role to their aligned high-level variables. At the location L1 , Dx and Dy are composed to form a neural representation with content x + y that is then composed with L2 to create an output. The fact that S1 doesn’t depend on z tells us that while L1 depends on Dz and CHAPTER 4. CAUSAL ABSTRACTIONS OF NEURAL NETWORKS 69 representations at L1 may even correlate with z, the information about z is not causally represented at L1 . At the location L2 , the value of z is simply repeated and then composed with L1 to create a final output. Our method assigns causally impactful information content, but also identifies the abstract causal structure along which representations are composed. It thus encompasses and improves on both correlational (probing) and attribution methods. 4.4 The Natural Language Inference Task and Models Multiply Quantified NLI Dataset The Multiply Quantified NLI
(MQNLI) dataset of Geiger et al. [2019a] contains templatically generated English-language NLI examples that involve very complex interactions between quantifiers, negation, and modifiers. We provide a few examples in Figure 4.2b; the empty-string symbol ε ensures perfect alignments at the token level both between premises and hypotheses and across all examples. The MQNLI examples are labeled using an algorithmic implementation of the natural logic of MacCartney and Manning [2009] over tree structures, and MQNLI has train/dev/test splits that vary in their difficulty. In the hardest setting, the train set is provably the minimal set of examples required to ensure that the dev and test sets can be perfectly solved by a simple symbolic model; in the easier settings, the train set redundantly encodes necessary information, which might allow a model to perform perfectly in assessment by memorization despite not having found a truly general solution. For a fuller review of the dataset
MQNLI is a fitting benchmark given our goals for a few reasons. First, we can focus on the hardest splits that can be generated, which will stress-test our NLI architectures in a standard behavioral way. Second, the MQNLI labeling algorithm itself suggests an appropriate causal model of the data-generating process. Figure 42a summarizes this model in tree form, and it is presented in full detail in Geiger et al. [2019a] This allows us to rigorously assess whether a neural network has learned to implement variants of this causal model. The complexity of the MQNLI examples creates many opportunities to do this in linguistically interesting ways. Models We evaluated two models on MQNLI: a randomly initialized multilayered Bidirectional LSTM (BiLSTM; Schuster and Paliwal [1997]) and a BERT-based classifier model in which the English bert-base parameters [Devlin et al., 2019b] are fine-tuned on the MQNLI train set Output predictions are computed using the final representation above the
[CLS] token. Models are trained to predict the relation of every pair of aligned phrases in Figure 4.2a Results Figure 4.2c summarizes the results of our BERT and BiLSTM models on the hardest fair generalization task Geiger et al. [2019a] creates with MQNLI We find that our BiLSTM model is not able to learn this task, and that our BERT model is able to achieve high accuracy. The only CHAPTER 4. CAUSAL ABSTRACTIONS OF NEURAL NETWORKS 70 models in Geiger et al. [2019a] able to achieve above 50% accuracy were task-specific tree-structured models with the structure of the tree in Figure 4.2a Thus, our BERT-based model is the first general-purpose model able to achieve good performance on this hard generalization task. Without pretraining, the BERT-based model achieves ≈49.1%, confirming that pretraining is essential, as expected. A natural hypothesis is that the BERT-based model achieves this high performance because it has in effect induced some approximation to the tree-like
structure of the data-generating process in its own internal layers. With causal abstraction analysis, we are actually in a position to test this hypothesis. 4.5 A Case Study in Structural Neural Network Analysis 4.51 Causal Abstractions of Neural NLI models Formulating Our Hypotheses We proceed just as we did for the simple motivating example in Section 4.3, except that we are now seeking to assess the extent to which the natural logic algebra in Figure 4.2a is a causal abstraction of the trained neural models in the above section The hallmark of Figure 4.2a is that it defines an alignment between premise and hypothesis at both lexical and phrasal levels. This permits us to run interchange interventions in a naturally N compositional way. For a given non-leaf node N in Figure 42a, let CNatLog be a submodel of CNatLog that computes the relation between the aligned phrases under N and uses them to compute the NPObj final output relation between premise and hypothesis. For
example, let CNatLog be the submodel of CNatLog that computes the relation between the two aligned object noun phrases and then uses that relation in computing the final output relation between premise and hypothesis (see Figure 4.3 right) We would like to ask whether our trained neural models also compute this relation between object noun phrases and use it to make a final prediction. We can pose this same question for other nodes which correspond to a pair of aligned subphrases. Alignment Search For each N , we search for an alignment between a neural representation in N NNLI and the variable N in CNatLog . In principle, any location in the network could be the right one for any causal model. Testing every hypothesis in this space would be intractable Thus, for each N CNatLog , we consider a restricted set of hidden representations based on the identity of N . The BERT model we use has 12 Transformer layers [Vaswani et al., 2017b], meaning that there are 12 hidden representations
for each input token. Each alignment search considers aligning the intermediate high-level variable with dozens of possible locations in the grid of BERT representations. Specifically, the following locations were considered for each N : CHAPTER 4. CAUSAL ABSTRACTIONS OF NEURAL NETWORKS 71 • P • QSubj , AdjSubj , NSubj , Neg, Adv, V, QObj , QPObj : hidden representations above QObj H AdjObj , NObj : hidden representations above and QObj . the two descendant leaf tokens. • P • H NegP: same but above Neg and Neg . • All nodes (for BERT): same but above [CLS] NPSubj , VP, and NPObj : same but above the four descendant leaf tokens. and [SEP]. For each alignment considered, we performed a full causal abstraction analysis. We report the results from the best alignments in Table 4.1 Interchange Interventions We first focus on our high-level causal models. Consider a non-leaf ′ node N from Figure 4.2a and two input token sequences e and e from MQNLI Define N
←e ′ CNatLog (e) (4.4) N to be the output provided by the causal model CNatLog when processing input e where the relation between the aligned subphrases under the node N is changed to the relation between those subphrases ′ in e . For example, simplifying for the sake of exposition, suppose e is (some happy baker, no ϵ baker ), ′ which has output label contradiction, and suppose e is (every happy person, some happy baker ), which has output label entailment. We wish to intervene on the noun phrase, so N = NP In e, ′ the noun phrase relation is entailment; in e , it is reverse entailment. Thus, CNatLog (e) changes the NP←e ′ object noun phrase relation in e to entailment while holding everything else about e constant. This results in the output label for the example (some happy person, no ϵ baker ), which is neutral. Next, we consider interventions in a neural model NNLI . Define L←e ′ NNLI (e) (4.5) to be the output provided by NNLI processing the
input e when the representation at location L is ′ replaced with the representation at location L created when NNLI is processing e . This is exactly the process depicted in Figure 4.1, except now the networks are the complex trained networks of Section 4.4 Our hypothesis linking Figure 4.2a with a model NNLI takes the same form as (43) The causal N model CNatLog is a constructive abstraction of NNLI when, for some representation location L, it is ′ the case that, for all MQNLI examples e and e , we have N ←e ′ L←e ′ CNatLog (e) = NNLI (e) (4.6) This asserts a correspondence between interventions on the representations at L in network NNLI N and interventions on the variable N in the causal model CNatLog . If it holds, then NNLI computes CHAPTER 4. CAUSAL ABSTRACTIONS OF NEURAL NETWORKS 72 N Table 4.1: Largest subsets of examples on which specific models CNatLog are abstractions of an LSTM and BERT model trained on MQNLI. We record the size of such subsets as
a percentage of the total 1000 examples. On this subset, we know that the neural models compute a representation of the relation between the aligned subphrases under N and use this information to make a final prediction. Causal Model QSubj QObj Neg AdjSubj NSubj AdjObj NObj V Adv NPSubj NPObj VP NegP LSTM BERT 0.7 0.9 0.7 2.5 1.2 0.9 0.7 0.4 1.4 1.0 0.7 0.4 0.9 13.1 7.3 21.4 6.7 5.5 14.1 8.8 11.4 7.9 6.7 38.3 11.4 11.8 Nodes removed H NObj H AObj P NObj P AObj H H NObj , AObj H P NObj , NObj H P NObj , AObj P H NObj , AObj H P AObj , AObj P P NObj , AObj (a) Main results (clique sizes) for nonleaf nodes of the tree in Figure 4.2a The hypothesis we have most evidence for is that the BERT model computes a representation of the NPObj node with the alignment shown in Figure 4.3 Remarkably, with 1000 examples sampled, we found a subset of 383 examNPObj ples where CNatLog is an abstraction of BERT. BERT 31.9 15.7 33.8 15.8 31.9 14.1 32.2 31.6 8.8 32.1 Nodes added BERT P AdjSubj P
NSubj P 30.5 37.2 14.9 26.9 35.6 16.2 13.4 12.0 34.4 16.2 13.4 12.0 Neg P Adv P V H QObj H AdjSubj H NSubj H Neg H Adv H V H QObj (b) Detailed results (clique sizes) for Alternative causal models NPObj in a “neighborhood” around the model CNatLog , which has a single intermediate variable composed of four lexical items (See Figure 4.3) At left, we have alternative causal models where one or two of those lexical items are removed from the composition. At right, we have alternatives obtained by adding one lexical item to the composition. We observe that no alternative hypothesis about causal structure considered has more evidence. the relation between the aligned phrases under the node N and uses this information to compute the relation between the premise and hypothesis. We call a pair of examples (e, e ) successful if it satisfies equation (4.6), ie, interventions in both ′ the target causal model and neural model produce equal results. In addition, to isolate the causal
impact of our interventions, we specifically focus on pairs (e, e ) for which performing the intervention ′ produces a different output value than without the intervention. We call a pair (e, e ) impactful if: ′ N ←e ′ N CNatLog (e) ≠ CNatLog (e) Quantifying Partial Success (4.7) Equation (4.6) universally quantifies over all examples We do not expect this kind of perfect correspondence to emerge in practice for real problems: neural network training is often approximate and variable in nature, and even our best model does not achieve perfect performance. However, we can still ask how widely (46) holds for a given model To do this, N we seek to find the largest subset of MQNLI on which CNatLog is an abstraction of our neural models, CHAPTER 4. CAUSAL ABSTRACTIONS OF NEURAL NETWORKS 73 for each non-leaf node N in CNatLog . More specifically, considering each example in MQNLI as a vertex in a graph, we add an undirected edge between two examples ei and ej if and
only if both the ordered pairs (ei , ej ) and (ej , ei ) satisfy N (4.6) In other words, CNatLog is an abstraction of a neural model on a subset of examples S of MQNLI if and only if all examples in S form a clique. The number of interventions we need to run scales quadratically with the number of inputs we 2 consider, so we sample 1000 MQNLI examples, producing a total of 1000 = 1M ordered pairs. We only consider examples for which the neural network outputs a correct label. For each node N and each of its corresponding neural network locations L, we perform interventions on all of these pairs. We choose to measure the largest clique with at least one impactful edge, because (1) the causal abstraction relation holds with full force on that clique, but other measures such as the total number of connections lack this theoretical grounding, and (2) if a clique has at least one impactful edge, that guarantees the high-level variable is being used. Results and Analysis For each target
causal model node N and neural network representation location L, we construct a graph as described above with 1000 examples as vertices and add an edge between two examples ei and ej if and only if both (ei , ej ) and (ej , ei ) are successful. We then find the largest clique in this graph with at least one impactful edge and record its size. Table 4.1a shows, for each causal model node N , the maximum size of cliques found among all neural locations. With this stricter impactful criterion (as opposed to simply using intervention N success), our results show that, for almost all nodes N , our target causal model CNatLog is indeed a causal abstraction of BERT on a significant number of examples in our dataset. These subsets are much smaller for the BiLSTM model. We also investigated alternative high-level causal structures that are not variants of CNatLog from Figure 4.2a Specifically, we consider alternative models in a “neighborhood” around the model NPObj CNatLog that can be
obtained by adding one leaf, or by removing one or two leaves to the composition. These results are in Table 4.1b Remarkably, all of these alternative models result in smaller clique sizes, significantly so for many of them. This further supports the significance of our results This analysis is similar to the analysis of our hypothetical addition example in Section 4.3, except for two crucial differences. First, for each variable N , we are hypothesizing that the causal model N CNatLog is an abstraction of NNLI , whereas in the addition example there was only one model. To investigate this difference, we take N = NPObj as a paradigm case, as it is the model with the strongest results. Second, we only achieved partial experimental success, whereas in the addition example we assumed complete success. Crucially, this means that the following analysis will be valid NPObj only on subsets of the input space on which the abstraction relation holds between NNLI and CNatLog . We visualize the
results of our intervention experiments for the node NPObj in Figure 4.4 The NPObj alignment with the largest subset of inputs aligns the NPObj variable in CNatLog with the neural P representation on the fourth layer of BERT above the AdjObj token (see Figure 4.3) Because neural CHAPTER 4. CAUSAL ABSTRACTIONS OF NEURAL NETWORKS 74 representations encode the value of their aligned variables and play a parallel causal role to their highlevel variables, we know that, on this subset of input examples, at the fourth neural representation P above the AdjObj token, the four input embeddings for the object nouns and adjectives in the premise and hypothesis are composed to form a neural representation with information content of the relation between the object noun phrases in the premise and hypothesis. Then this representation is composed with the other input-embeddings to create an output representing the relation between the premise and hypothesis. 4.52 Comparison with Other
Structural Analysis Methods Probes We probed neural representation locations for the relation between aligned subexpressions on a subset of 12,800 randomly selected MQNLI examples. For a pair of aligned subexpressions below a node N in Figure 4.2a, we probe the columns above the same set of restricted class of tokens as described in Section 4.51 To evaluate these probes, we report accuracy as well as selectivity as defined by Hewitt and Liang [2019]: probe accuracy minus control accuracy, where control accuracy is the train set accuracy of a probe with the same architecture but trained on a control task to factor out probe success that can be attributed to the probe model itself. Our control task is to learn a random mapping from node types to semantic relations. Figure 4.4 summarizes our probing results for N = NPObj , along with corresponding interchange intervention results for comparison. Probes tell us that information about the relation between the aligned noun phrases is
encoded in nearly all of the locations we considered, and using the selectivity metric does not result in any qualitative change. In contrast, our intervention heatmaps indicate only a small number of locations store this information in a causally relevant way. Clearly, our intervention experiments are far more discriminating than probes. Integrated Gradients Attribution methods that estimate feature importance can measure causal properties of neural representations, but a single feature importance method is an impoverished characterization of a representation’s role in network behavior. Whereas our interchange interventions gave us high-level information about how a neural representation is composed and what it is composed into, attribution methods simply tell us “how much” a representation contributes to the network output on a give input. Moreover, intervention interchanges provide a rich, high-level characterization of causal structure on a space of inputs. We use
integrated gradients on our models to verify the intuitive hypothesis that if a premise and hypothesis differ by a single token, then the neural representations above that token should be more causally responsible for the network output than other representations. For example, given premise ‘Every sleepy cat meows’ and hypothesis ‘Some hungry cat meows’, the attributive modifier position is different and the rest are matched. The neural representations above the adjective tokens CHAPTER 4. CAUSAL ABSTRACTIONS OF NEURAL NETWORKS 75 sleepy and hungry should be more important for the network output than others, because if those adjectives were the same, the example label would change from neutral to entailment. 4.6 Conclusion We have introduced a methodology for deriving interpretable causal explanations of neural network behaviors, grounded in a formal theory of causal abstraction. The methodology involves first formulating a hypothesis in the form of a high-level,
interpretable causal model, then searching for an alignment between the neural network and the causal model, and finally verifying experimentally that the neural representations encode the same causal properties and information content as the corresponding components of the high-level causal model. As a case study demonstrating the feasibility of the approach, we analyzed neural models trained on the semantically formidable MQNLI dataset. Guided by the intuition that success on this challenging task may call for a way of recapitulating the causal structure of the natural logic model that generates the MQNLI data, we were able to verify the hypothesis that a state-of-the-art BERT-based model partially realizes this structure, whereas baseline models that do not perform as well fail to do so. This suggestive case study demonstrates that our theoretically grounded methodology can work in practice. CHAPTER 4. CAUSAL ABSTRACTIONS OF NEURAL NETWORKS 76 O S2 L2 L1 S1 Dx Dy Dz x y
z X W Y Z (a) The causal model C+ (right) that first computes S1 = X + Y and W = Z, before computing the final output S2 = W + S1 aligned with the neural network N+ (left) with L1 highlighted as the hypothesized location encoding S1 = X + Y and L2 as the location encoding W = Z. 12 15 12 L1 L1 D1 D2 D3 D4 D5 D6 1 2 3 4 5 6 (b) Low-level neural network interchange intervention. The network processes two different input sequences. The neural representation created at L1 for input sequence (1, 2, 3) is replaced by the corresponding representation created for input sequence (4, 5, 6). 9 1 15 3 2 9 3 4 6 5 6 (c) High-level symbolic computation interchange intervention. The computation processes two different input sequences. The sum at S1 for input sequence (1, 2, 3) is replaced by the corresponding representation created for input sequence (4, 5, 6). Figure 4.1: Our motivating example where we hypothesis that a symbolic computation C+ is a causal
abstraction of a neural network N+ under a particular alignment (top). We can experimentally confirm this hypothesis by conducting an interchange intervention on both the network and the computation with every pair of inputs and evaluating whether the intervened network and intervened computation have the same counterfactual output behavior. We schematically depict an interchange intervention on the network N+ (bottom left) and the computation C+ (bottom right) with the base input (1, 2, 3) and the source input (4, 5, 6). Observe that the output of the intervened neural network matches the output of the intervened symbolic computation, so we have success for this pair of inputs. CHAPTER 4. CAUSAL ABSTRACTIONS OF NEURAL NETWORKS 77 QPSubj NegP NPSubj QSubj P QSubj H QSubj NSubj AdjSubj P H P QPObj Neg H AdjSubj AdjSubj NSubj NSubj Neg P Neg H Adv H Adv V P H QObj P QObj V Adv P NPObj QObj VP V NObj AdjObj P AdjObj H H AdjObj P NObj H NObj (a) The
causal structure of the high-level natural logic causal model CNatLog that performs inference on MQNLI. The superscripts P and H stand for ‘premise’ and ‘hypothesis’ and the subscripts ‘Subj’ and ‘Obj’ stand for ‘Subject’ and ‘Object’. The node labels are used to explain the experimental results in Section 4.5 ε every ε baker ε ε ε eats ε no ε bread contradiction ε no angry baker ε ε ε eats ε no ε bread ε every silly professor ε ε ε sells not every ε book neutral ε every silly professor ε ε ε sells not every ε chair not every sad baker ε ε fairly admits not every odd idea entailment ε some ε baker does not ε admits ε no ε idea (b) MQNLI examples. The ε token serves as padding (but still attended to by the model) and ensures a perfect alignment between both premises and hypotheses and across all examples. It is semantically an identity element Model Train Dev Test CBoW TreeNN CompTreeNN BiLSTM BERT 88.04 67.01 99.65 99.42
99.99 54.18 54.01 80.17 46.41 88.25 53.99 53.73 80.21 46.32 88.50 (c) MQNLI results. The first three models are from Geiger et al. 2019a, where the CompTreeNN is a task-specific model not suitable for general NLI and functions as an idealized upperbound. Our results show that BERT-based models can surpass this without such alignments. Figure 4.2: The natural logic causal model (top), MQNLI examples (left) and MQNLI results (right) output Label ⋮ ⋮ ⋮ ⋮ ⋮ ⋮ ⋮ ⋮ NPObj P P P P P P P AdjObj NObj H H H H H H H AdjObj NObj QSubj AdjSubj NSubj Neg Adv V QObj [CLS] P P P QSubj AdjSubj NSubj . P AdjObj P NObj . QSubj AdjSubj NSubj Neg Adv V QObj P P H H NPObj Figure 4.3: A BERT-based NLI model (left) aligned with the natural logic causal model CNatLog P (right), where the fourth vector representation above the AdjObj token in the network is aligned with NPObj , the variable representing the relation between the object noun phrases.
When analyzing a NPObj sample of 1000 examples, we found a subset of 383 where CNatLog is an abstraction of NNLI under this alignment. CHAPTER 4. CAUSAL ABSTRACTIONS OF NEURAL NETWORKS (a) Clique size. (b) Interchange success. (c) Probe selectivity. 78 (d) Probe Accuracy. Figure 4.4: Interchange intervention and probing results for the NPObj position Vertical axes denote layers of BERT and horizontal axes denote the token position of hidden representations. The intervention success rates reported here are calculated based on intervention experiments with a change in the output label. Clique sizes are reported as % of 1000 examples Chapter 5 Finding Alignments Between Interpretable Causal Variables and Distributed Neural Representations Abstract Causal abstraction is a promising theoretical framework for explainable artificial in- telligence that defines when an interpretable high-level causal model is a faithful simplification of a low-level deep learning system. However,
existing causal abstraction methods have two major limitations: they require a brute-force search over alignments between the high-level model and the low-level one, and they presuppose that variables in the high-level model will align with disjoint sets of neurons in the low-level one. In this paper, we present distributed alignment search (DAS), which overcomes these limitations. In DAS, we find the alignment between high-level and low-level models using gradient descent rather than conducting a brute-force search, and we allow individual neurons to play multiple distinct roles by analyzing representations in non-standard basesdistributed representations. Our experiments show that DAS can discover internal structure that prior approaches miss. Overall, DAS removes previous obstacles to uncovering conceptual structure in trained neural nets. 5.1 Introduction Can an interpretable symbolic algorithm be used to faithfully explain a complex neural network model? This is a key question
for explainable AI; a positive answer can provide guarantees about how the model will behave, and a negative answer could lead to fundamental concerns about whether the model will be safe and trustworthy. Causal abstraction provides a mathematical framework for precisely characterizing what it means 79 CHAPTER 5. DISTRIBUTED NEURAL REPRESENTATIONS 80 for any complex causal system (e.g, a deep learning model) to implement a simpler causal system (e.g, a symbolic algorithm) Rubenstein et al [2017b], Beckers et al [2019] For modern AI models, the fundamental operation for assessing whether this relationship holds in practice has been the interchange intervention, in which a neural network is provided a ‘base’ input, and sets of neurons are forced to take on the values they would have if different ‘source’ inputs were processed. The counterfactuals that these interventions create are the basis for causal inferences about model behavior. Geiger et al. [2021a] show that the
relevant causal abstraction relation obtains when interchange interventions on aligned high-level variables and low-level variables have equivalent effects. This ideal relationship rarely obtains in practice, but the proportion of interchange interventions with the same effect (interchange intervention accuracy; IIA) provides a graded notion, and Geiger et al. [2023c] formally ground this metric in the theory of approximate causal abstraction. Geiger et al also use causal abstraction theory as a unified framework for a wide range of recent intervention-based analysis methods [Vig et al., 2020b, Csordás et al, 2021, Feder et al, 2021, Ravfogel et al, 2020a, Elazar et al., 2020, De Cao et al, 2021b, Abraham et al, 2022a, Olah et al, 2020, Olsson et al, 2022, Chan et al., 2022b] Causal abstraction techniques have been applied to diverse problems Geiger et al. [2019a, 2020a], Li et al. [2021b], Huang et al [2022] However, previous applications have faced two central challenges First,
causal abstraction requires a computationally intensive brute-force search process to find optimal alignments between the variables in the high-level model and the states of the low-level one. Where exhaustive search is intractable, we risk missing the best alignment entirely. Second, these prior methods are localist: they artificially limit the space of possible alignments by presupposing that high-level causal variables will be aligned with disjoint groups of neurons. There is no reason to assume this a priori, and indeed much recent work in model explanation (see especially Ravfogel et al. [2020a], Elazar et al [2020], Olah et al [2020], Olsson et al [2022]) is converging on the insight of Smolensky [1986], Rumelhart et al. [1986], and McClelland et al [1986] that individual neurons can play multiple conceptual roles. Smolensky [1986] identified distributed neural representations as “patterns” consisting of linear combinations of unit vectors. In the current paper, we propose
distributed alignment search (DAS), which overcomes the above limitations of prior causal abstraction work. In DAS, we find the best alignment via gradient descent rather than conducting a brute-force search. In addition, we use distributed interchange interventions, which are “soft” interventions in which the causal mechanisms of a group of neurons are edited such that (1) their values are rotated with a change-of-basis matrix, (2) aligned dimensions of the rotated neural representation are fixed to be the corresponding values in the rotated neural representation created for the source inputs, and (3) the representation is rotated back to the standard neuron-aligned basis. The key insight is that viewing a neural representation through an alternative basis that is not aligned with individual neurons can reveal interpretable dimensions [Smolensky, CHAPTER 5. DISTRIBUTED NEURAL REPRESENTATIONS 81 1986]. In our experiments, we evaluate the capabilities of DAS to provide faithful
and interpretable explanations with two tasks that have obvious interpretable high-level algorithmic solutions with two intermediate variables. In both tasks, the distributed alignment learned by DAS is as good or better than both the closest localist alignment and the best localist alignment in a brute-force search. In our first set of experiments, we focus on a hierarchical equality task that has been used extensively in developmental and cognitive psychology as a test of relational reasoning [Premack, 1983, Thompson et al., 1997, Geiger et al, 2022a]: the inputs are sequences [w, x, y, z], and the label is given by (w = x) = (y = z). We train a simple feed-forward neural network on this task and show that it perfectly solves the task. Our key question: does this model implement a program that computes w = x and y = z as intermediate values, as we might hypothesize humans do? Using DAS, we find a distributed alignment with 100% IIA. In other words, the network is perfectly abstracted
by the high-level model; the distinction between the learned neural model and the symbolic algorithm is thus one of implementation rather than algorithm. Our second task models a natural language inference dataset Geiger et al. [2020a] where the inputs are premise and hypothesis sentences (p, h) that are identical but for the words wp and wh ; the label is either entails (p makes h true) or contradicts/neutral (p makes h false). We fine-tune a pretrained language model to perfectly solve the task. With DAS, we find a perfect alignment (100% IIA) to a causal model with a binary variable for the entailment relation between the words wp and wh (e.g, dog entails mammal ) In both our sets of experiments, the DAS analyses reveal perfect abstraction relations. However, we also identify an important difference between them. In the NLI case, the entailment relation can be decomposed into representations of wp and wh . What appears to be a representation of lexical entailment is, in this case, a
“data structure” containing two representations of word identity, rather than an encoding of their entailment relation. By contrast, the hierarchical equality models learn representations of w = x and y = z that cannot be decomposed into representations of w, x, y and z. In other words, these relations are entirely abstracted from the entities participating in the relation; DAS reveals that the neural network truly implements a symbolic, tree-structured algorithm. 5.2 Related Work A theory of causal abstraction specifies exactly when a ‘high-level causal model’ can be seen as an abstract characterization of some ‘low-level causal model’ Iwasaki and Simon [1994], Chalupka et al. [2017], Rubenstein et al. [2017b], Beckers et al [2019] The basic idea is that high-level variables are associated with (potentially overlapping) sets of low-level variables that summarize their causal mechanisms with respect to a set of hard or soft interventions Massida et al. [2022] In
practice, a graded notion of approximate causal abstraction is often more useful [Beckers et al., 2019, Rischel CHAPTER 5. DISTRIBUTED NEURAL REPRESENTATIONS 82 and Weichwald, 2021, Geiger et al., 2023c] Geiger et al. [2023c] argue that causal abstraction is a generic theoretical framework for providing faithful [Jacovi and Goldberg, 2020a, Lyu et al., 2022] and interpretable Lipton [2018a] explanations of AI models and show that LIME [Ribeiro et al., 2016a], causal effect estimation Abraham et al [2022a], Feder et al. [2021], causal mediation analysis Vig et al [2020b], Csordás et al [2021], De Cao et al [2021b], iterated nullspace projection Ravfogel et al. [2020a], Elazar et al [2020], and circuit-based explanations Olah et al. [2020], Olsson et al [2022], Wang et al [2022b], Chan et al [2022b] can all be understood as causal abstraction analysis. Interchange intervention training (IIT) objectives are minimized when a high-level causal model is an abstraction of a neural
network under a given alignment Geiger et al. [2022c], Wu et al [2022b], Huang et al. [2022] In this paper, we use IIT objectives to learn an alignment between a high-level causal model and a deep learning model. 5.3 Methods We focus on acyclic causal models and seek to provide an intuitive overview of our method. An acyclic causal model consists of input, intermediate, and output variables, where each variable has an associated set of values it can take on and a causal mechanism that determine the value of the variable based on the value of its parents. For a simple running example, we modify the boolean conjunction models of Geiger et al. 2022c to reveal key properties of DAS A causal model B for this problem can be defined as below, where the inputs and outputs are booleans t and f. Alongside B, we also define a causal model N of a linear feed-forward neural network that solves the task. Here we show B, N , and the parameters of N : V3 = v1 ∧ v2 Y = [h1 ; h2 ]w + b V1 = p V2
= q H1 = [x1 ; x2 ]W1 p q X1 H2 = [x1 ; x2 ]W2 X2 W1 = [ cos(20◦ ) − sin(20◦ ) ] W2 = [ sin(20 ) ◦ cos(20 ) ] ◦ w=[ 1 1 ] b = −1.8 This network solves the boolean conjunction problem perfectly in that every combination of pairs of input boolean values is mapped to the intended output. An input x of a model M determines a unique total setting M(x) of all the variables in the model. The inputs are fixed to be x and the causal mechanisms of the model determine the values of the remaining variables. We denote the values that M(x) assigns to the variable or variables Y as GetValues(M(x), Y). For example, GetValues(B([t, f]), V3 ) = f Interventions Interventions are a fundamental building block of causal abstraction analysis. An intervention I ← i is a setting i of variables I. Together, an intervention and an input setting x of a model M determine a unique total setting that we denote as MI←i (x). The inputs are fixed to be x, CHAPTER 5. DISTRIBUTED NEURAL
REPRESENTATIONS 83 and the causal mechanisms of the model determine the values of the non-intervened variables, with the intervened variables I being fixed to i. We can define interventions on both our causal model B and our neural model N . For example, BV1 ←t ([f, t]) is our boolean model when it processes input [f, t] but with variable V1 set to t. This has the effect of changing the output value to t. Similarly, whereas N ([0, 1]) leads to an intermediate values h1 = −0.34 and h2 = 094 and output value −12, if we compute Nh1 ←134 ([0, 1]), then the output value is 0.48 This has the effect of changing the predicted value to t, because 048 > 0 Alignment In causal abstraction analysis, we ask whether a specific low-level model like N implements a high-level algorithm like B. This is always relative to a specific alignment of variables between the two models. An alignment assigns to each high-level variable X a set of low-level variables ΠX and a function τX that maps
from values of the low-level variables in ΠX to values of the aligned high-level variable X. One possible alignment between B and N establishes a one-to-one correspondence between high-level and low-level variables, where Π is depicted by the dashed lines connecting B and N in the above diagram. We immediately know what the functions for high-level input and output variables are. For the inputs, t is encoded as 1 and f is encoded as 0, meaning τx2 (1) = τx1 (1) = t and τx1 (0) = τx2 (0) = f. For the output, the network only predicts t if y > 0, meaning τV3 (x) = t if x > 0, else f. This is simply a consequence of how a neural network is used and trained. The functions for high-level intermediate variables τV1 (x) and τV2 (x) must be discovered and verified experimentally. Constructive Causal Abstraction Relative to an alignment like this, we can define abstraction: Definition 44. Constructive Causal Abstraction A high-level causal model H is a constructive abstraction
of a low-level causal model L under alignment (Π, τ ) exactly when the following holds for every low-level input setting x and low-level intervention I ← i: τ (LI←i (x)) = Hτ (I←i) (τ (x)) H being a causal abstraction of L, under an alignment, guarantees that the causal mechanism for each high-level variable X is a faithful simplification of the causal mechanisms for the low-level variables in ΠX . To assess the degree to which a high-level model is a constructive causal abstraction of a low-level model, we perform interchange interventions: Definition 45. Interchange Interventions Given source input settings {sj }1 , and non-overlapping k sets of intermediate variables {Xj }1 for model M, we define the interchange intervention to be the k model k k II(M, {sj }1 , {Xj }1 ) = M⋀kj=1 ⟨Xj ←GetValsX (M(sj ))⟩ j CHAPTER 5. DISTRIBUTED NEURAL REPRESENTATIONS 84 k where ⋀j=1 ⟨⋅⟩ conjoins a set of interventions; see, e.g, the conjunction operator of Ibeling
and Icard [2018]. A base input setting can be fed into the resulting model to compute the counterfactual output value. Consider the following interchange intervention: II(B, {[t, t]}, {{V1 }}) = B{V1 }←GetVals{V1 } (B([t,t])) We process a base input and a source input, and then we intervene on a target variable, replacing it with the value obtained by processing the source. Our causal model is very easy to understand, and so we know ahead of time that this interchange intervention yields t. For our neural network, the corresponding behavior probably is not known ahead of time and may be significantly harder to establish formally. The interchange intervention corresponding to the above (according to the alignment we are exploring) is as follows II(N , {[1, 1]}, {{H1 }}) = N {V1 } ← GetVals{H1 } (N ([1, 1])) And, indeed, the counterfactual behavior of the model and the network N are unequal: V3 = t V3 = t f Y = −0.26 V1 = t V2 = t f t V1 = t t Y = 0.08 V2 = t t H1 = 0.6
H2 = 0.94 H1 = 0.6 H2 = 1.28 0 1 1 1 Under the given alignment, the interchange interventions at the low and high level have different effects. Thus, we have a counterexample to constructive abstraction as given in Definition 44 Although N has perfect behavioral accuracy, its accuracy under the counterfactuals created by our interventions is not perfect, and thus B is not a constructive abstraction of N under this alignment. Distributed Interventions The above conclusion is based on the kind of localist causal abstraction that has been explored in the literature to date. As we noted in Section 51, there are two risks associated with this conclusion: (1) we may have chosen a suboptimal alignment, and (2) we may be wrong to assume that the relevant structure will be encoded in the standard basis we have implicitly assumed throughout. If we simply rotate the representation [H1 , H2 ] by −20 to get a new representation [Y1 , Y2 ], ◦ then the resulting network has perfect
behavioral and counterfactual accuracy when we align V1 with Y1 . What this reveals is that there is an alignment, but not in the basis we chose Since the choice of basis was arbitrary, our negative conclusion about the causal abstraction relation was spurious. This rotation localizes the information about the first and second argument into separate dimensions. To understand this, observe that the weight matrix of the linear network rotates a two CHAPTER 5. DISTRIBUTED NEURAL REPRESENTATIONS X2 X2 X2 X1 X3 85 X1 X1 X2 Y1 X3 X3 X1 X1 X2 Y1 R Y2 X3 X3 X1 R X2 X3 R Y1 Y2 Y2 Y3 Y1 Y2 Y3 Y3 Y1 Y2 Y3 Y1 Y2 Y3 Y1 Y2 Y3 Y3 Y1 Y2 R −1 Y3 X2 X1 X2 X3 X1 X3 Figure 5.1: A generic multi-source distributed interchange intervention The base input and two source inputs create three total settings of a model. The top left (green) and right (blue) total model settings are determined by two source inputs and the middle total model setting (red) is
determined by the base input. Three hidden units from each total setting are rotated with an orthogonal matrix R ∶ X Y. Then we intervene on the rotated representation for the base input and fix two dimensions to be the value they take on for each source input, respectively. Then we −1 unrotate the representation with R and compute a counterfactual total model setting for the base input. In DAS, the orthogonal matrix is found with SGD using a high-level causal model to guide the search process. ◦ ◦ dimensional vector by 20 and the rotation matrix rotates the representation by 340 . The two matrices are inverses. Because this network is linear, there is no activation function and so rotating the hidden representation “undoes” the transformation of the input by the weight matrix. Under this non-standard basis, the first hidden dimension is equal to the first input argument and the second hidden dimension is equal to the second input argument. This reveals an essential
aspect of distributed neural representations: there is a many-to-many mapping between neurons and concepts, and thus multiple high-level causal variables might be encoded in structures from overlapping groups of neurons Rumelhart et al. [1986], McClelland et al [1986]. In particular, Smolensky [1986] proposes that viewing a neural representation under a basis that is not aligned with individual neurons can reveal the interpretable distributed structure of the neural representations. To make good on this intuition we define a distributed intervention, which first transforms a set of variables to a vector space, then does interchange on orthogonal sub-spaces, before transforming back to the original representation space. CHAPTER 5. DISTRIBUTED NEURAL REPRESENTATIONS 86 Definition 46. Distributed Interchange Interventions We begin with a causal model M with input variables S and source input settings {sj }j=1 . Let N be a subset of variables in M, the target k variables. Let Y be a
vector space with subspaces {Yj }0 that form an orthogonal decomposition, k k i.e, Y = ⨁j=0 Yj Let R be an invertible function R ∶ N Y Write ProjYj for the orthogonal 1 projection operator of a vector in Y onto subspace Yj . A distributed interchange intervention yields a new model DII(M, R, {sj }1 , {Yj }0 ) which is identical to M except that the mechanisms k k FN which yield values of N from a total setting are replaced by: k FN (v) = R (ProjY0 (R(FN (v))) + ∑ ProjYj (R(FN (M(sj ))))). ∗ −1 j=1 Notice that in this definition the base setting is partially preserved through the intervention (in subspace Y0 ) and hence this is a soft intervention on N that rewrites causal mechanisms while maintaining a causal dependence between parent and child. Under this new alignment, the high-level interchange intervention II(B, {[t, t]}, {{V1 }}) = B{V1 }←GetVals{V1 } (B([t,t])) is aligned with the low-level distributed interchange intervention DII(N , [ cos(−20 ) ◦
− sin(−20 ) sin(−20 ) cos(−20 ) ◦ ◦ ◦ ], {[1, 1]}, {{Y1 }}) = N{Y1 }←GetVals{Y1 } (N ([1,1])) and the counterfactual output behavior of B and N are equal: t Y = 0.08 [ H1 = 0.6 H2 = 1.28 H1 = −0.34 H2 = 0.94 0 1 cos(20 ) ◦ − sin(20 ) sin(20 ) cos(20 ) ◦ ◦ ◦ ] Y = 0.08 cos(−20 ) − sin(−20 ) sin(−20 ) cos(−20 ) ◦ [ ◦ ◦ ◦ ] 0.0 1.0 1.0 1.0 cos(−20 ) − sin(−20 ) sin(−20 ) cos(−20 ) ◦ [ ◦ ◦ ◦ ] H1 = 0.6 H2 = 1.28 1 1 In what follows we will assume that X are already vector spaces (which is true for neural nets) and the functions R are rotation operators. In this case, the subspaces Yj can be identified without loss of generality with those spanned by the first ∣Y0 ∣ basis vectors for Y0 , the next ∣Y1 ∣ basis vectors for Y1 , and so on. (The following methods would be well-defined for non-linear transformations, as long as they were invertible and differentiable, but
efficient implementation becomes harder.) Distributed Alignment Search The question then arises of how to find good rotations. As we discussed above, previous causal abstraction analyses of neural networks have performed brute-force search through a discrete space of hand-picked alignments. In distributed alignment search (DAS), we find an alignment between one or more high-level variables and disjoint sub-spaces (but not necessarily subsets) of a large neural representation. We define a distributed interchange intervention 1 Thus, Proj generalizes GetVals to arbitrary vector spaces. CHAPTER 5. DISTRIBUTED NEURAL REPRESENTATIONS 87 training objective, use differentiable parameterizations for the space of orthogonal matrices (such as provided by PyTorch), and then optimize the objective with stochastic gradient descent. Crucially, the low-level and high-level models are frozen during learning so we are only changing the alignment. In the following definition we assume that a neural
network specifies an output distribution for a given input, which can then be pushed forward to a distribution on output values of the high-level model via an alignment function τ . We may similarly interpret even a deterministic high-level model as defining a (e.g, delta) distribution on output values We make use of these distributions, after interchange intervention, to define a differentiable loss for the rotation matrix which aligns intermediate variables. Definition 47. Distributed Interchange Intervention Training Objective Begin with a low-level neural network L, with low-level input settings InputsL , a high-level algorithm H, with high-level output settings OutH , and an alignment τ for their input and output variables. Suppose we want to align intermediate high level variables Xj ∈ VarsH with rotated subspaces Yj of a neural representation θ N ⊂ VarsL with learned rotation matrix R ∶ N Y. In general, we can define a training objective using any differentiable loss
function Loss that quantifies the distance between two total high-level settings. θ k k k k Loss(DII(L, R , {sj }1 , {Yj }0 )(b), II(H, {τ (sj )}1 , {Xj }1 )(τ (b))) ∑ b,s1 ,.,sk ∈InputsL For our experiments, we compute the cross entropy loss CE(⋅, ⋅) between the high-level output distribution P(outH ∣H(τ (b))) and the push-forward under τ of the low-level output distribution P (outH ∣L(b)). The overall objective is: τ ∑ τ θ k k CE(P (outH ∣DII(L, R , {sj }1 , {Yj }0 )(b)), b,s1 ,.,sk ∈InputsL k k P(outH ∣II(H, {τ (sj )}1 , {Xj }1 ))(τ (b))) While we still have discrete hyperparameters (N, ∣Y0 ∣, . , ∣Yk ∣)the target population and the dimensionality of the sub-spaces used for each high-level variablewe may use stochastic gradient descent to determine the rotation that minimizes loss, thus yielding the best distributed alignment between L and H. Approximate Causal Abstraction Perfect causal abstraction relationships are unlikely
to arise for neural networks trained to solve complex empirical tasks. We use a graded notion of accuracy: Definition 48. Distributed Interchange Intervention Accuracy Given low-level and high-level causal models L and H with alignment (Π, τ ), rotation R ∶ N Y, and orthogonal decomposition {Yj }0 . k CHAPTER 5. DISTRIBUTED NEURAL REPRESENTATIONS Both Equality Relations Left Equality Relation 88 Identity of First Argument Identity Subspace of Left Equality Hidden size Intervention size Layer 1 Layer 2 Layer 3 Layer 1 Layer 2 Layer 3 Layer 1 Layer 2 Layer 3 Layer 1 ∣N∣ = 16 ∣N∣ = 16 ∣N∣ = 16 k=1 k=2 k=8 0.88 0.97 1.00 0.51 0.54 0.57 0.50 0.50 0.50 0.85 0.85 0.90 0.54 0.55 0.56 0.50 0.50 0.50 0.51 0.50 0.52 0.52 0.52 0.53 0.50 0.51 0.51 0.51 0.50 0.51 ∣N∣ = 32 ∣N∣ = 32 ∣N∣ = 32 k=2 k=4 k = 16 0.93 0.97 0.99 0.63 0.63 0.67 0.49 0.49 0.53 0.92 0.94 0.99 0.65 0.65 0.65 0.50 0.50 0.50 0.52 0.51 0.49 0.55 0.55 0.55 0.52
0.52 0.52 0.50 0.51 0.51 Brute-Force Search Localist Alignment 0.60 0.73 0.56 0.56 0.52 0.48 0.64 0.60 0.64 0.50 0.57 0.49 0.50 0.46 0.51 0.47 0.54 0.48 - Table 5.1: Hierarchical equality alignment learning results The table can be read as follows: Layer 1, Layer 2, and Layer 3 indicate which layer of neurons is targeted, ∣N∣ is the number of neurons in a layer, k is the number of neurons aligned with each intermediate variable (red) where our subspace model occupies k2 with rounding up to the closest integer, and the values in each cell are interchange intervention accuracies for the learned alignment on training data. We report the best results from three runs with distinct random seeds. If we let InputsL be low-level input settings and VarsH be high-level intermediate variables the interchange intervention accuracy (IIA) is as follows ∑ b,s1 ,.,sk ∈InputsL 1 θ k k k k [τ (DII(L, R , {sj }1 , {Yj }0 )(b)) = II(H, {τ (sj )}1 , {Xj }1 )(τ (b))] ∣InputsL
∣k+1 IIA is the proportion of aligned interchange interventions that have equivalent high-level and low-level effects. In our example with N and A, IIA is 100% and the high-level model is a perfect abstraction of the low-level model (Def. 44) When IIA is α <100%, we rely on the graded notion of α-on-average approximate causal abstraction Geiger et al. [2023c], which directly coincides with IIA General Experimental Setup We will illustrate the value of DAS by analyzing feed-forward networks trained on a hierarchical equality and pretrained Transformer-based language models [Vaswani et al., 2017b] fine-tuned on a natural language inference task Our evaluation paradigm is as follows: 1. Train the neural network N to solve the task In all experiments, the neural models achieve perfect accuracy on both training and testing data. 2. Create interchange intervention training datasets Each example consists of a base input, one or more source inputs corresponding to high-level causal
variables, and a counterfactual gold label that will be output by the network if the interchange intervention has the hypothesized effect on model behavior. This gold label is a counterfactual output of the high-level model we will align with the network. 3. Optimize an orthogonal matrix to learn a distributed alignment for each high-level model that maximizes IIA using the training object in Definition 47. We experiment with different CHAPTER 5. DISTRIBUTED NEURAL REPRESENTATIONS 89 hidden dimension sizes for our low-level model and different intervention site sizes (i.e, the dimensionality of low-level representations) and locations (i.e, the layer where the intervention happens). 4. Evaluate a baseline that brute-force searches through a discrete space of alignments and selects the alignment with the highest IIA. We search the space of alignments by aligning each high-level variable with groups of neurons in disjoint sliding windows. 5. Evaluate the localist alignment
“closest” to the learned distributed alignment The rotation matrix for the localist alignment will be axis-aligned with the standard basis, possibly permuting and reflecting unit axes. 6. Determine whether each distributed representation aligned with high-level variables can be decomposed into multiple representations that encode the identity of the input values to the variable’s causal mechanism. We do this by learning a second rotation matrix that decomposes learned distributed representation, holding the first rotation matrix fixed. 5.4 Hierarchical Equality Experiment We now illustrate the power of DAS for analyzing networks designed to solve a hierarchical equality task. We concentrate on analyzing a trained feed-forward network A basic equality task is to determine whether a pair of objects are the same (x = y). A hierarchical equality task is to determine whether a pair of pairs of objects have identical relations: (w = x) = (y = z). Specifically, the input to the task
is two pairs of objects and the output is True if both pairs are equal or both pairs are unequal and False otherwise. For example, (A, A, B, B) and (A, B, C, D) are both assigned True while (A, B, C, C) is assigned False. Low-Level Neural Model We train a three-layer feed-forward network with ReLU activations to perform the hierarchical equality task. Each input object is represented by a randomly initialized vector. Specifically, our model has the following architecture where k is the number of layers h1 = ReLU([x1 ; x2 ; x3 ; x4 ]W1 + b1 ) hj−1 = ReLU(hj Wj + bj ) y = softmax(hk Wk + bk ) We evaluate our model on held-out random vectors unseen during training, as in Geiger et al. 2022a High-Level Models We use DAS to evaluate whether trained neural networks have achieved the natural solution to the hierarchical equality task where the left and right equality relations are computed and then used to predict the final label (Figure 5.2a) However, evaluating this high-level model
alone is insufficient, as there are obviously many other high-level models of this task. To further contextualize our results, we also consider two alternatives: CHAPTER 5. DISTRIBUTED NEURAL REPRESENTATIONS V3 ← (V1 = V2 ) V1 ← (w = x) V2 ← (y = z) w y x z (a) A causal model of hierarchical equality. Sentence Pairs Label premise: A man is talking to someone in a car. hypothesis: A man is talking to someone in a taxi. neutral premise: The people are not playing sitars. hypothesis: The people are not playing instruments. entails (b) Monotonicity NLI (MoNLI) examples. 90 MoNLI(x) 1 lexrel ← get-lexrel(x) 2 neg ← contains-not(x) 3 if neg: 4 return reverse(lexrel) 5 return lexrel (c) A program that solves MoNLI. Figure 5.2: High-level models a high-level model where only the equality relation of the first pair is represented and a high-level model where the lone intermediate variable encodes the identity of the first input object (leaving all computation for
the final step). These alternative high-level models also solve the task perfectly Discussion The IIA results achieved by the best alignment for each high-level model can be seen in Table 5.1 The best alignments found are with the ‘Both Equality Relations’ model that is widely assumed in the cognitive science literature. For all causal models, DAS learns a more faithful alignment (higher IIA) than a brute-force search through localist alignments. This result is most pronounced for ‘Both Equality Relations’, where DAS learns perfect or near-perfect alignments under a number of settings, whereas the best brute-force alignment achieves only 0.60 and the best localist alignment achieves only 0.73 Finally, the distributed representation of left equality could not be decomposed into a representation of the first argument identity. We see this in the very low performance of the ‘Identity Subspace of Left Equality’ results. This indicates that models are truly learning to encode
an abstract equality relation, rather than merely storing the identities of the inputs. 5.5 Monotonicity NLI Experiment In our second experiment, we analyze a BERT model fine-tuned on the Monotonicity Natural Language Inference (MoNLI) benchmark Geiger et al. [2020a] An NLI example consists of a premise sentence and hypothesis sentence and the output label is entails when the premise makes the hypothesis true, and contradicts/neutral when the premise makes the hypothesis false. Two examples are in Figure 5.2b High-Level Causal Model In MoNLI, every example is such that a single word wp in the premise sentence was changed to a hypernym (more general term) or hyponym (more specific term) wh to create the hypothesis. About half of MoNLI examples contain a negation that scopes over the word replacement site, and the remaining examples have no negation. When no negation is present, the label for a premise–hypothesis pair is the lexical relation. When negation is present, the label for
a CHAPTER 5. DISTRIBUTED NEURAL REPRESENTATIONS Negation and Lexical Entailment Lexical Entailment 91 Identity of Lexeme Lexeme Subspace of Lexical Entailment Hidden size Intervention size Layer 7 Layer 9 Layer 11 Layer 7 Layer 9 Layer 11 Layer 7 Layer 9 Layer 11 Layer 9 ∣N∣ = 768 ∣N∣ = 768 ∣N∣ = 768 k = 64 k = 128 k = 256 0.65 0.65 0.67 0.96 0.99 1.00 0.91 0.92 0.86 0.88 0.88 0.91 1.00 1.00 1.00 0.97 0.99 1.00 0.88 0.89 0.88 0.94 0.93 0.96 0.93 0.92 0.88 0.97 0.97 0.98 Brute-Force Search Localist Alignment 0.60 0.51 0.56 0.51 0.52 0.51 0.64 0.47 0.64 0.47 0.57 0.47 0.50 0.50 0.51 0.50 0.54 0.50 - Table 5.2: Monotonicity NLI results The table can be read as follows: Layer 7, Layer 9, and Layer 11 indicate which layer of neurons is targeted, ∣N∣ is the number of neurons in a layer, k is the number of neurons aligned with each intermediate variable (red) where our subspace model occupies k , and the values in each cell are
interchange intervention accuracies for the learned alignment on 2 training data. We report the best results from three runs with distinct random seeds premise–hypothesis pair is the reverse of the lexical relation. This model is perhaps best expressed as a simple program (Figure 5.2c) Low-Level Neural Model We fine-tune an uncased BERT-base model [Devlin et al., 2019b] 2 finetuned on the MultiNLI dataset [Williams et al., 2018a] Our BERT model has 12 layers and 12 heads with a hidden dimension of 768. We concatenate the tokenized sequences of the premise sentence and hypothesis sentence with a [SEP] token. Because of the size of the rotation matrix, we can’t look for distributed representations across all tokens; we look only at the representations of the [CLS] token because the final classification is made from this token’s representation in the last layer. High-Level Models We use DAS to evaluate whether BERT fine-tuned on MoNLI will represent two boolean intermediate
variables. The first is an indicator variable for negation, which is true if and only if negation is present in the premise and hypothesis. The second is a variable that is true if wp entails wh . Again, we also consider two alternative high-level models to contextualize our results One model represents only lexical entailment and not negation. The other represents the identity of the premise word wp . Results The IIA results achieved by the best alignment for each high-level model can be seen in Table 5.2 There is a perfect alignment between fine-tuned BERT and a symbolic algorithm with variables representing the presence of negation and the lexical entailment relation between wp and wh . In Table 52, this is shown by the perfect IIA for layer 9 and intervention size 256, meaning 256 non-standard basis dimensions of the [CLS] token representation in layer 9 of BERT encode the relation between wp and wh and 256 other non-standard basis dimensions encode negation. Across all
alignments and intervention types, DAS learns more faithful alignments (higher IIA) 2 The parameters are provided by the Hugging Face transformers library [Wolf et al., 2019], downloaded from https://huggingface.co/ishan/bert-base-uncased-mnli CHAPTER 5. DISTRIBUTED NEURAL REPRESENTATIONS 92 than a brute-force search through alignments, and no localist alignment comes close to the learned distributed alignments in terms of IIA. The distributed representation of the lexical entailment relation between wp and wh can be nearly perfectly decomposed into two representations that encode the identity of the word wp and the identity of the word wh , respectively. This result is shown by the near perfect IIA in the final column of Table 5.2 5.6 Conclusion We introduce distributed alignment search (DAS), a method to align interpretable causal variables with distributed neural representations. We learn distributed alignments that are more interpretable than localist alignments and do so
with a gradient-descent based search method that improves upon the state-of-the-art brute-force search. In our two experiments, we discovered perfect alignments of distributed neural representations to binary high-level variables encoding simple equality and lexical entailment relations. However, when we investigated the substructure of these representations, we found that the lexical entailment representations could be decomposed into sub-representations of word identity, but we could not decompose the equality relation into sub-representations of object identity. This highlights the need to investigate the causal substructure of neural representations On the other hand, the presence of perfect representations of simple equality relations that cannot be decomposed into representations of the entities in the relations is a foundational result that should inform our understanding of how and when symbolic and connectionist architectures coexist. Chapter 6 Interpretability at Scale:
Identifying Causal Mechanisms in Alpaca Abstract Obtaining human-interpretable explanations of large, general-purpose language models is an urgent goal for AI safety. However, it is just as important that our interpretability methods are faithful to the causal dynamics underlying model behavior and able to robustly generalize to unseen inputs. Distributed Alignment Search (DAS) [Geiger et al, 2023d] is a powerful gradient descent method grounded in a theory of causal abstraction that uncovered perfect alignments between interpretable symbolic algorithms and small deep learning models fine-tuned for specific tasks. In the present paper, we scale DAS significantly by replacing the remaining brute-force search steps with learned parameters – an approach we call Boundless DAS. This enables us to efficiently search for interpretable causal structure in large language models while they follow instructions. We apply Boundless DAS to the Alpaca model (7B parameters), which, off the shelf,
solves a simple numerical reasoning problem. With Boundless DAS, we discover that Alpaca does this by implementing a causal model with two interpretable boolean variables. Furthermore, we find that the alignment of neural representations with these variables is robust to changes in inputs and instructions. These findings mark a first step toward deeply understanding the inner-workings of our largest and most widely deployed language models. 6.1 Introduction Present-day large language models (LLMs) display remarkable behaviors: they appear to solve coding tasks, translate between languages, engage in open-ended dialogue, and much more. As a result, their 93 CHAPTER 6. INTERPRETABILITY AT SCALE 94 societal impact is rapidly growing, as they make their way into products, services, and people’s own daily tasks. In this context, it is vital that we move beyond behavioral evaluation to deeply explain, in human-interpretable terms, the internal processes of these models; an initial
step in auditing them for safety, trustworthiness, and pernicious social biases. The theory of causal abstraction [Beckers et al., 2019, Geiger et al, 2023a] provides a generic framework for representing interpretability methods that faithfully assess the degree to which a complex causal system (e.g, a neural network) implements an interpretable causal system (eg, a symbolic algorithm). Where the answer is positive, we move closer to having guarantees about how the model will behave. However, thus far, such interpretability methods have been applied only to small models fine-tuned for specific tasks [Geiger et al., 2020a, 2021b, Li et al, 2021a, Chan et al, 2022b] and this is arguably not an accident: the space of alignments between the variables in the hypothesized causal model and the representations in the neural network becomes exponentially larger as models increase in size. When a good alignment is found, one has specific formal guarantees Where no alignment is found, it could
easily be a failure of the alignment search algorithm. Distributed Alignment Search (DAS) Geiger et al. [2023d] marks real progress on this problem DAS opens the door to (1) discovering structure spread across neurons and (2) using gradient descent to learn an alignment between distributed neural representations and causal variables. However, DAS still requires a brute-force search over the dimensionality of neural representations, hindering its use at scale. In this paper, we introduce Boundless DAS which replaces the remaining brute-force aspect of DAS with learned parameters, truly enabling explainability at scale. We use Boundless DAS to study how Alpaca (7B) Taori et al. [2023], an off-the-shelf instruct-tuned LLaMA model, follows basic instructions in a simple numerical reasoning task. Figure 61 summarizes the approach We find that Alpaca achieves near-perfect task performance because it implements a simple algorithm with interpretable variables. In further experiments, we show
that Alpaca uses this simple algorithm across a wide range of contexts and variations on the task. These findings mark a first step toward 1 deeply understanding the inner-workings of our largest and most widely deployed language models. 6.2 Related Work Interpretability Many methods have been developed in an attempt to explain and understand deep learning models. These methods include analyzing learned weights Clark et al [2019], Abnar and Zuidema [2020], Coenen et al. [2019], gradient-based methods Simonyan et al [2014], Shrikumar et al. [2017], Zeiler and Fergus [2014b], Sundararajan et al [2017b], Voita et al [2019], Belinkov and Glass [2019b], probing Conneau et al. [2018], Tenney et al [2019], Hewitt and Manning [2019], Saphra and Lopez [2019], Clark et al. [2019], Manning et al [2020], Chi et al [2020], Rogers et al [2020], 1 We will release our code upon publication. CHAPTER 6. INTERPRETABILITY AT SCALE 95 Figure 6.1: Our pipeline for scaling causal explainability to
LLMs with billions of parameters syntax-driven interventions Murty et al. [2023], self-generated model explanations Kojima et al [2022], and training external explainers based on model behaviors Ribeiro et al. [2016b], Lundberg and Lee [2017a]. However, these methods rely on observational measurements of behavior and internal neural representations. Such explanations are often not guaranteed to be faithful to the underlying causal mechanisms of the target models Lipton [2018b], Geiger et al. [2021b], Uesato et al [2022], Wang et al. [2022a] Causal Abstraction The theory of causal abstraction Rubenstein et al. [2017b], Beckers and Halpern [2019a], Beckers et al. [2019] offers a unifying mathematical framework for interpretability methods aiming to uncover interpretable causal mechanisms in deep learning models Geiger et al. [2020a], Li et al. [2021a], Geiger et al [2021b], Wu et al [2022b], Wang et al [2022b], Geiger et al [2023d] and training methods for inducing such interpretable
mechanisms Geiger et al. [2022e], Wu et al. [2022c,b], Huang et al [2022] Causal abstraction Geiger et al [2023a] can represent many existing interpretablity methods, including interative nullspace projection Ravfogel et al. [2020a], Elazar et al. [2020], Lovering and Pavlick [2022b], causal mediation analysis Vig et al [2020b], Meng et al. [2022b], and causal effect estimation Feder et al [2021], Elazar et al [2022], Abraham et al [2022a]. In particular, causal abstraction grounds the research program of mechanistic interpretability, which aims to reverse engineer deep learning models by determining the algorithm or computation underlying their intelligent behavior Olah et al. [2020], Elhage et al [2021], Olsson et al [2022], Chan et al. [2022b], Wang et al [2022b] To the best of our knowledge, there is no prior work that scales CHAPTER 6. INTERPRETABILITY AT SCALE 96 these methods to large, general-purpose LLMs. Training LLMs to Follow Instructions Instruction-based fine-tuning
of LLMs can greatly enhance their capacity to follow natural language instructions Christiano et al. [2017], Ouyang et al [2022]. In parallel, this ability can also be induced into the model by fine-tuning base models with hundreds of specific tasks Wei et al. [2022], Chung et al [2022] Recently, Wang et al [2022c] show that the process of creating fine-tuning data for instruction following models can partly be done by the target LLM itself (“self-instruct”). Such datasets have led to many recent successes in lightweight fine-tuning of ChatGPT-like instruction following models such as Alpaca Taori et al. [2023], instruct-tuned LLaMA Touvron et al. [2023], StableLM Andonian et al [2021], Vicuna Chiang et al. [2023], and GPT-J Wang and Komatsuzaki [2021] Our goal is to scale methods from causal abstraction to understand how these models follow a particular instruction. 6.3 Methods 6.31 Background on Causal Abstraction Causal Models We represent black box networks and
interpretable algorithms using causal models that consist of variables that take on values according to causal mechanisms. An intervention is an operation that edits some of the causal mechanisms in a model. We denote the output of a causal model M provided an input x as M(x). Begin with a causal model M with input variables S, base input b, Interchange Intervention k k source input settings {sj }1 , and disjoint sets of target variables {Zj }1 . An interchange intervention k k yields a new model II(M, {sj }1 , {Zj }1 ) which is identical to M except the values for each set of target variables Zj are fixed to be the value they would have taken for source input sj . We then denote the output for the base input with II(M, {sj }1 , {Zj }1 )(b). k k Distributed Interchange Intervention Let N be a subset of variables in M, the target variables. Let Y be a vector space with orthogonal subspaces {Yj }1 . Let R be an invertible function R ∶ N Y k Write ProjZ for the orthogonal
projection operator of a vector in Y onto subspace Z. A distributed interchange intervention yields a new model DII(M, R, {sj }1 , {Yj }1 ) which is identical to M except k k the causal mechanisms have been rewritten such that each subspace Yj is fixed to be the value it 2 would take for source input sj . Specifically, the causal mechanism for N is set to be k FN (v) = R (ProjY0 (R(FN (v))) + ∑ ProjYj (R(FN (M(sj ))))) ∗ −1 (6.1) j=1 2 Previous work often focusses on zeroing-out representations or substituting with mean values Meng et al. [2022b], Wang et al. [2022b], rather than the more general interchange operation CHAPTER 6. INTERPRETABILITY AT SCALE 97 k where Y0 = Y ⨁j=1 Yj . We then denote the output for the base input as DII(M, R, {sj }1 , {Yj }0 )(b). k Causal Abstraction and Alignment k We are licensed to claim an algorithm A is a faithful interpretation of the network N if the causal mechanisms of the variables in the algorithm abstract the causal
mechanisms of neural representations relative to a particular alignment. For our purposes, each variable of a high-level model is aligned with a linear subspace in the vector space formed by rotating a neural representation with an orthogonal matrix. An alignment and task input–output behavior together define a partial function τ that translates low-level variable settings to high-level variable settings. Approximate Causal Abstraction Interchange intervention accuracy (IIA) is a graded measure of abstraction that computes the proportion of aligned interchange interventions on the algorithm and neural network that have the same output. The IIA for an alignment (τ, Π) of high-level variables Xj to orthogonal subspace Yj between an algorithm A and network N is 1 k θ k [DII(N , R , {sj }1 , {Yj }1 )(b) = ∑ ∣Inputs∣k+1 b,s1 ,.,sk ∈Inputs L k k II(A, {τ (sj )}1 , {Xj }1 ))(τ (b))] (6.2) 6.32 Boundless Distributed Alignment Search Distributed alignment search (DAS) is a
method for learning an alignment between interpretable causal variables of a model C and fixed dimensionality linear subspaces of neural representations in a network N using gradient descent Geiger et al. [2023d] as shown in Figure 63 Specifically, an orthogonal matrix R ∶ N Y is trained to maximize interchange intervention accuracy under an alignment from each variable Xj to fixed dimensionality linear subspaces Yj of the rotated vector space. Boundless DAS is our extension of DAS that learns the dimensionality of the orthogonal linear subspaces in a d-dimensional vector space Y using a method inspired by work in neural PDE Wu et al. [2022a]. Specifically, for each high-level variable Xj , we introduce a learnable continuous boundary index parameter bj that can take on a value between 0 and d, where b0 = 0, bj < bj+1 , and bj < d for all j. The boundary mask Mj for the source is a vector with d values between 0 and 1 where the k-th element of the array is defined to be (Mj )k
= sigmoid ( k − bj bj+1 − k ) ∗ sigmoid ( ) β β (6.3) where β is a temperature that we anneal through training. As β approaches 0, the masks Mj converge to binary-valued vectors that together encode an orthogonal decomposition of Y. For simplicity, we CHAPTER 6. INTERPRETABILITY AT SCALE 98 restrict each bj to be a multiple of bj−1 with a factor of 2. Weighted Distributed Interchange Intervention Let N be a subset of variables in M, the target variables. Let Y be a vector space with d dimensions and let {Mj }1 vectors in [0, 1] Let k d R ∶ N Y be an invertible transformation. A weighted distributed interchange intervention yields a new model SoftDII(M, R, {sj }1 , {Mj }1 ) which is identical to M except the causal mechanisms k k have been rewritten such that each source input sj contributes to the setting of Y in proportion to its mask Mj . Specifically, the causal mechanism for N is set to be k k FN (v) = R ((1 − ∑ Mj ) ◦ R(FN (v)) + ∑ (Mj ◦
R(FN (M(sj )))) ∗ −1 j=1 (6.4) j=1 where ◦ is element wise multiplication and 1 is a d dimensional vector where each element is 1. We then denote the output for the base input as SoftDII(M, R, {sj }1 , {Mj }1 )(b). k Boundless DAS k Given a base input b and source inputs {sj }1 , we minimize the following objective k to learn a rotation matrix R and masks {Mj }1 θ ∑ θ k θ k θ k k k CE(SoftDII(N , R , {sj }1 , {Mj }1 )(b), II(A, {τ (sj )}1 , {Xj }1 ))(τ (b))) (6.5) b,s1 ,.,sk ∈InputsL where CE is the cross entropy loss. We anneal β throughout training and our weighted interchange interventions become more and more similar to unweighted interchange interventions. During evaluation we snap the masks to be binary-valued to create an orthogonal decomposition of Y where each high-level variable Xj is aligned with a linear subspace of Y picked out by the mask Mj , with k the residual (unaligned) subspace being picked out by the mask (1 − ∑j=1 Mj ).
6.4 Experiment 6.41 Price Tagging We follow the approach in Fig. 61 by first assessing the ability of Alpaca to execute specific actions based on the instructions provided in input. Formally, the input to the model M is given an instruction ti (e.g, “correct the spelling of the word:”) followed by a test query input xi (eg, “aplpe”). We use M(ti , xi ) to depict the model generation yp given the instruction and the test query input. We can evaluate model performance by comparing yp with the gold label y We focus on tasks with high model performance, to ensure that we have a known behavioral pattern to explain. The instruction prompt of the Price Tagging game follows the publicly released template of the CHAPTER 6. INTERPRETABILITY AT SCALE 99 Figure 6.2: Four proposed high-level causal models for how Alpaca solves the price tagging task Intermediate variables are in red. All these models perfectly solve the task Alpaca (7B) model. The core instruction contains an the
English sentence: Please say yes only if it costs between [X.XX] and [XXX] dollars, otherwise no followed by an input dollar amount [X.XX], where [XXX] are random continuous real numbers drawn with a uniform distribution from [0.00, 999] The output is a single token ‘Yes’ or ‘No’ For instance, if the core instruction says Please say yes only if it costs between [1.30] and [855] dollars, otherwise no., the answer would be “Yes” if the input amount is “350 dollars” and “No” if the input is “9.50 dollars” We restrict the absolute difference between the lower bound and the upper bound to be [2.50, 750] due to model errors outside these values – again we need behavior to explain. 6.42 Hypothesized Causal Models As shown in Figure 6.2, we have identified a set of human-interpretable high-level causal models, with alignable intermediate causal variables, that would solve this task with 100% performance: • Left Boundary: This model has one high-level boolean
variable representing whether the input amount is higher than the lower bound, and an output node incorporating whether the input amount is also lower than the high bound. • Left and Right Boundary: The previous model is sub-optimal in only abstracting one of the boundaries. In this model, we have two high-level boolean variables representing whether the input amount is higher than the lower bound and lower than the higher bound. We take a conjunction of these boolean variables to predict the output. • Mid-point Distance: We calculate the mid-point of the lower and upper bounds (e.g, the mid-point of “3.50” and “850” is “600), and then we take the absolute distance between the input dollar amount and the mid-point as a. We then calculate one-half of the bounding bracket length (e.g, the bracket length for “350” and “850” is “500”) as b We predict output “Yes” if a ≤ b, otherwise “No”. We align only with the mid-point variable CHAPTER 6.
INTERPRETABILITY AT SCALE 100 Figure 6.3: Aligned distributed interchange interventions performed on the Alpaca model that is instructed to solve our Price Tagging game. It aligns the boolean variable representing whether the input amount is higher than the lower bound in the causal model. To train Boundless DAS, we sample two training examples and then swap the intermediate boolean value between them to produce a counterfactual output using our causal model. In parallel, we swap the aligned dimensions of the neural representations in rotated space. Lastly, we update our rotation matrix such that our neural network has a more similar counterfactual behavior to the causal model. CHAPTER 6. INTERPRETABILITY AT SCALE 101 • Bracket Identity: This model represents the lower and upper bound in a single interval variable and passes this information to the output node. We predict the output as “Yes” if the input amount within the interval, otherwise “No”. Model Architecture Our
target model is the Alpaca (7B) Taori et al. [2023], an off-the-shelf instruct-tuned LLaMA model. It is a Transformer-based decoder-only autoregressive trained language model with 32 layers and 32 attention heads. It has a hidden dimension in size of 4096, which is also the dimension of our rotation matrix which is applied for each token representation. In total, the rotation matrix contains 16.8M parameters and has size 4096 × 4096 Alignment Process We separately train Boundless DAS with test query input token representa- tions (i.e, starting from the token of the first input digit till the last token in the prompt) in a set of selected 7 layers: {0, 5, 10, 15, 20, 25, 30}. We also train Boundless DAS on the token before the first digit as a control condition where nothing should be expected to be aligned. Instead of interchanging with multiple source input settings {sj }1 at the same time, we interchange with a single source at a k time for simplicity, while allowing multiple
causal variables to be aligned across examples. We run each experiment with three distinct random seeds. Since the global optimum of the Boundless DAS objective corresponds to the best attainable alignment, but SGD may be trapped by local optima, we report the best-performing seed in terms of IIA. Figure 63 provides an overview of these analyses Evaluation Metric To evaluate models, we use Interchange Intervention Accuracy (IIA) as defined in Eqn. 62 IIA is bounded between 00 and 10 For most of the experiments, the lower bound of IIA is at the chance (0.5) given the distribution of our output labels IIA can occasionally go above the model’s task performance, when the interchange interventions puts the model in a better state, but for the most part IIA is constrained by task accuracy. 6.43 Boundless DAS Results Figure 6.4 shows our main results, given in terms of IIA across our four hypothesized algorithmic models (Figure 6.2) The results show very clearly that ‘Left
Boundary’ and ‘Left and Right Right Boundary’ (top panels) are highly accurate hypotheses about how Alpaca solves the task. For them, IIA is at or above task performance (0.85), and intermediate variable representations are localized in systematically arranged positions. By contrast, ‘Mid-point Distance’ and ‘Bracket Identity’ (bottom panels) are inaccurate hypotheses about Alpaca’s processing, with IIA peaking at around 0.72 These findings suggest that, when solving our reasoning task, Alpaca internally follows our first two high-level models by representing causal variables that align with boundary checks for both left and right boundaries. Interestingly, heatmaps on the top two rows also show a pattern of higher scores around the bottom left and upper right and close to zero scores in other positions. This is CHAPTER 6. INTERPRETABILITY AT SCALE 102 Figure 6.4: Interchange Intervention Accuracy (IIA) for four different alignment proposals The Alpaca model
achieves 85% task accuracy. The higher the number is, the more faithful the alignment is. We color each cell by scaling IIA using the model’s task performance as the upper bound and a dummy classifier (predicting the most frequent label) as the lower bound. These results indicate that the top two are highly accurate hypotheses about how Alpaca solves the task, whereas the bottom two are inaccurate in this sense. Analyzing tokens includes special tokens (eg, ‘<0x0A>’ for linebreaks) required by Alpaca’s instruct-tuning template. CHAPTER 6. INTERPRETABILITY AT SCALE 103 significantly different from the other two alignments where, although some positions are highlighted (e.g, the representations for the last token), all the positions receive non-zero scores In short, the accurate hypotheses correspond to highly structured IIA patterns, and the inaccurate ones do not. Additionally, alignments are better when the model post-processes the query input with 1–2 additional
steps: accuracy is higher at position 75 compared to all previous positions. Surprisingly, there exist bridging tokens (positions 76–79) where accuracy suddenly drops compared to earlier tokens, which suggests that these representations have weaker causal effects on the model output. In other words, boundary-check variables are fully extracted around position 75, level 10, and are later copied into activations for the final token before responding. By comparing the heatmaps on the top row, our findings suggest that aligning multiple variables at the same time poses a harder alignment process, in that it lowers scores slightly across multiple positions. 6.44 Interchange Interventions with (In-)Correct Inputs IIA is highly constrained by task performance, and thus we expect it to be much lower for inputs that the model gets wrong. To verify this, we constructed an evaluation set containing 1K examples that the model gets wrong and evaluated our ‘Left and Right Boundary’
hypothesis on this subset. Table 6.1 reports these results in terms of max IIA and the correlation of the IIA values with those obtained in our main experiments. As expected, IIA drops significantly However, two things stand out: IIA is far above task performance (which is 0.0 by design now), and the correlation with the original IIA map remains high. These findings suggest that the model is using the same internal mechanisms to process these cases, and so it is likely that is is narrowly missing the correct output predictions. We also expect IIA to be higher if we focus only on cases that the model gets correct. Table 61 confirms this expectationm using ‘Left and Right Boundary’ as representative. IIAmax is improved from 0.86 to 088, and the correlation with the main results is essentially perfect 6.45 Do Alignments Robustly Generalize to Unseen Instructions and Inputs? One might worry that our positive results are highly dependent on the specific input–output pairs we have
chosen. We now seek to address this concern by asking whether the causal roles (ie, alignments) found using Boundless DAS in one setting are preserved in new settings. This is crucial, as it tells how robustly the causal model is realized in the neural network. Generalizing Across Two Different Instructions Here, we assess whether the learned alignments for the ‘Left Boundary’ causal model transfer between different specific price brackets in the instruction. To do this, we retrain Boundless DAS for the high-level model with a fixed instruction CHAPTER 6. INTERPRETABILITY AT SCALE 104 −2 Experiment Task Acc. IIAmax Correlation Left Boundary (♣) Left and Right Boundary (♥) Mid-point Distance Bracket Identity 0.85 0.85 0.85 0.85 0.90 0.86 0.70 0.72 1.00 1.00 1.00 1.00 2.01 1.14 0.04 0.08 Correct Only Incorrect Only 1.00 † 0.00 0.88 0.71 0.99 (♥) 0.84 (♥) 1.47 0.36 New Bracket (Seen) New Bracket (Unseen) Irrelevant Context (1) Irrelevant Context (2)
Sibling Instructions + exclude top right 0.94 0.95 0.84 0.84 0.84 0.84 0.94 0.95 0.83 0.85 0.83 0.83 0.97 (♣) 0.94 (♣) 0.97 (♥) 0.98 (♥) 0.87 (♥) 0.92 (♥) 3.02 1.66 0.94 1.04 0.93 0.95 † Variance (×10 ) Table 6.1: Summary results for all experiments with task performance as accuracy (range [0, 1]), maximal interchange intervention accuracy (IIA) (range [0, 1]) across all positions and layers, Pearson correlations of IIA between two distributions (compared to ♣ or ♥; range [−1, 1]), and variance of † IIA within a single experiment across all positions and layers. This is empirical task performance on the evaluation dataset for this experiment. that says “between 5.49 dollars and 849 dollars” Then, we fix the learned rotation matrix and evaluate with another instruction that says “between 2.51 dollars and 551 dollars” For both, Alpaca is successful at the task, with around 94% accuracy. Our hypothesis is that if the found alignment of the high-level
variable is robust, it should transfer between these two settings, as the aligning variable is a boolean-type variable which is potentially agnostic to the specific comparison price. Table 6.1 gives our findings in the ‘New Bracket’ rows Boundless DAS is able to find a good alignment for the training bracket with an IIAmax that is about the same as the task performance at 94%. For our unseen bracket, the alignments also hold up extremely well, with no drop in IIAmax For both cases, the found alignments also highly correlate with the counterpart of our main experiment. Generalizing with Inserted Context Recent work has shown that language models are sensitive to irrelevant context Shi et al. [2023] To address this concern, we add prefix strings to the input instructions and evaluate how ‘Left and Right Boundary’ alignment transfers. Our ‘Irrelevant Context’ prefixes are the short string “Price Tagging game!” and the longer string “Fruitarian Frogs May Be Doing Flowers
a Favor” (from New York Times April 28, 2023, the first post title in the ‘America’ column). For both cases, the model achieves 84% task performance, which is slightly lower than the average task performance. Nonetheless, the results in Table 61 suggest the found alignments transfer surprisingly well here, with only 1% drop in IIAmax . Meanwhile, the IIA distribution across positions and layers is highly correlated with our base experiment. Thus, our method seems to identify causal structure that is robust to changes in irrelevant task details and position in the input string. CHAPTER 6. INTERPRETABILITY AT SCALE 105 Figure 6.5: Learned boundary width for intervention site and in-training evaluation interchange intervention accuracy (IIA) for two groups of data: (1) aligned group where the boundary does not shrink to 0 at the end of the training; (2) unaligned group where the boundary does shrink to 0 at the end of the training. 1 on the y-axis means either 100% accuracy for
IIA, or the variable is occupying half of the hidden representation for the boundary width. Generalizing to Modified Outputs We further test whether the alignments found for instruction with a template “Say yes . , otherwise no” can generalize to a new instruction with a template “Say True . , otherwise False” If there are indeed latent representations of our aligning causal variables, the found alignments should persist and should be agnostic to the output format. Table 6.1 shows that this holds: the learned alignments actually transfer across these two sibling instructions, with a minimum 1% drop in IIAmax . In addition, the correlation increases 6% when excluding the top right corner (top 3 layers in the last position), as seen in the final row of the table. Generally, IIA drops near the top right. This indicates a late fusion of output classes and working representations, leading to different computations for the new instruction only close to the output. 6.46
Boundary Learning Dynamics Boundless DAS automatically learns the intervention site boundaries (Section 6.32) It is important to confirm that we do not give up alignment accuracy by optimizing boundaries, and to explore the dimensionality actually needed to represent the abstract variables. We sample 100 experiment runs from our experiment pool and create two groups: (1) the success group where the boundary does not shrink to 0 at the end of the training; (2) the failure group where the boundary does shrink to 0 at the end of the training. The failure group is for those instances where there is no significant causal role of representations found by our method. We plot how boundary width and in-training evaluation IIA vary throughout training. Figure 65 shows that the width for success cases shrink to 10–20% of one-half of the representation space for a single aligning variable. IIA performance converges quickly for success cases and, crucially, maintains stable performance despite
shrinking dimensionality. This suggests that only a small fraction of the whole representation space is needed for aligning a causal CHAPTER 6. INTERPRETABILITY AT SCALE 106 variable, and that these representations can be efficiently identified by Boundless DAS. 6.5 Analytic Strengths and Limitations Explanation methods for AI should be judged by the degree to which they can, in principle, identify models’ true internal causal mechanisms. If we reach 100% IIA for a high-level model using Boundless DAS we can assert that we have positively identified a correct causal structure in the model (though perhaps not the only correct abstraction). This follows directly from the formal results of Geiger et al. [2023a] On the other hand, failure to find causal structure with Boundless DAS is not a proof that such structure is missing, for two reasons. First, Boundless DAS explores a very flexible and large hypothesis space, but it is not completely exhaustive. For instance a set of
variables represented in a highly non-linear way might be missed. Second, we are limited by the space of causal models we think to test ourselves. For models where task accuracy is below 100%, it is still possible for IIA to reach 100% if the high-level causal model explains errors of the low-level model. However, in cases like those studied here where our high-level model matches the idealized task but our language model does not completely do so, we expect Boundless DAS to find only partial support for causal structure even if we have identified the ideal set of hypotheses and searched optimally. This too follows from the results of Geiger et al. [2023a] However, as we have seen in the results above, Boundless DAS can find structure even where the model’s task performance is low, since task performance can be shaped by factors like a suboptimal generation method or the rigidity of the assessment metric. Future work must tighten this connection by modeling errors of the language
model in more detail. 6.6 Conclusion We introduce Boundless DAS, a novel and effective method for scaling alignment search of causal structure in LLMs to billions of parameters. Using Boundless DAS, we find that Alpaca, off-the-shelf, solves a simple numerical reasoning problem in a human-interpretable way. Additionally, we address one of the main concerns around interpretability tools developed for LLMs – whether found alignments generalize across different settings. We rigorously study this by evaluating found alignments under several changes to inputs and instructions. Our findings indicate robust and interpretable algorithmic structure. Our framework is generic for any LLMs and is released to the public We hope that this marks a step forward in terms of understanding the internal causal mechanisms behind the massive LLMs that are at the center of so much work in AI. Chapter 7 Inducing Causal Structure for Interpretable Neural Networks Abstract In many areas, we have
well-founded insights about causal structure that would be useful to bring into our trained models while still allowing them to learn in a data-driven fashion. To achieve this, we present the new method of interchange intervention training (IIT). In IIT, we (1) align variables in a causal model (e.g, a deterministic program or Bayesian network) with representations in a neural model and (2) train the neural model to match the counterfactual behavior of the causal model on a base input when aligned representations in both models are set to be the value they would be for a source input. IIT is fully differentiable, flexibly combines with other objectives, and guarantees that the target causal model is a causal abstraction of the neural model when its loss is zero. We evaluate IIT on a structural vision task (MNIST-PVR), a navigational language task (ReaSCAN), and a natural language inference task (MQNLI). We compare IIT against multi-task training objectives and data augmentation. In
all our experiments, IIT achieves the best results and produces neural models that are more interpretable in the sense that they more successfully realize the target causal model. 7.1 Introduction In many domains, we have well-founded insights about causal structure that we can express in symbolic terms, ranging from commonsense intuitions about how the world works to advanced scientific knowledge. These insights have the potential to make up for gaps in available data, or more generally to provide useful inductive biases. Can we bring these insights into our models while still allowing them to learn in a data-driven fashion? In this paper, we present interchange intervention training (IIT), a new method that trains a 107 CHAPTER 7. INDUCING CAUSAL STRUCTURE FOR INTERPRETABLE NEURAL NETWORKS108 neural network to realize the abstract structure of a causal model. In IIT, we (1) align the variables in a causal model C with the representations in a neural model N and (2) train N to
have the counterfactual behavior of C by performing aligned interchange interventions (swapping of internal states created for different inputs) on N using C’s counterfactual output as the gold label for the counterfactual prediction of N . IIT objectives are differentiable and guarantee that, when the loss is zero, the target causal model is a causal abstraction of the neural network in the sense of Beckers and Halpern [2019b]. IIT is an extension of the causal abstraction analysis of Geiger et al. [2021c], which can be placed under the broader rubric of structural evaluations of neural models, which includes probing and many kinds of feature attribution. Our central point of differentiation from this prior work is that we go beyond passive study of static models, by pushing them to learn specific causal structures as part of optimization. This allows for a productive interplay between model analysis and model improvement: we not only assess whether models have systematic,
interpretable internal structure but also push them to acquire such structure. We evaluate IIT in three contexts: (1) ResNet trained on a vision task where one part of an image “points” to another (MNIST-PVR), (2) a CNN-LSTM model trained to produce action sequences in a grid world given a natural language command (ReaSCAN), and (3) a pretrained BERT model fine-tuned to label the semantic relation between two sentences (MQNLI). For each context, we define a high-level causal model that capture aspects of the task. We then align high-level causal variables to low-level neural representations to define IIT training objectives. For the three case studies, we report two kinds of evaluation: traditional behavioral evaluations using systematic generalization tasks that assess whether a model has learned a truly general solution, and structural evaluations that directly assess the interpretability of our models by evaluating whether they realize the target causal model. We compare IIT
against multi-task training objectives and data augmentation methods defined to make use of our causal models, finding that IIT leads to models that both perform better on systematic generalization benchmarks and are more interpretable. 7.2 1 Related Work Probes Probes are supervised or unsupervised models that can be used to gain an understanding of what is encoded in the internal representations of neural networks [Hupkes et al., 2018b, Peters et al., 2018, Tenney et al, 2019, Clark et al, 2019] Probes have yielded important insights about what models learn to encode. However, probes are fundamentally limited in a way that is central to our present goals: there is no guarantee that probed information plays a causal role in the network’s behavior [Ravichander et al., 2020, Elazar et al, 2020, Geiger et al, 2021c, 2020a] Feature Attribution In contrast to probes, gradient-based feature attribution methods [Zeiler 1 We release our code at
https://github.com/frankaging/Interchange-Intervention-Training CHAPTER 7. INDUCING CAUSAL STRUCTURE FOR INTERPRETABLE NEURAL NETWORKS109 and Fergus, 2014a, Springenberg et al., 2014, Shrikumar et al, 2016, Binder et al, 2016] generally do measure causal properties [Chattopadhyay et al., 2019] For example, Geiger et al [2021c] note that the integrated gradients method of Sundararajan et al. [2017a] computes the individual causal effect of neurons [Imbens and Rubin, 2015]. In comparison with our proposal, the main limitation of these methods is that (by definition) they passively study trained networks rather than allowing for active improvements of them (though see Erion et al. 2021 for a path from attribution to improved optimization). Intervention-Based Analyses In intervention-based analysis, one actively changes the values of model representations in systematic ways and studies the effects. Such interventions can be applied to input representations in order to measure the effect
on the output representation [Feder et al., 2021, Pryzant et al, 2021a], or on network internal representations to characterize how these representations mediate the causal relationships between inputs and outputs [Giulianelli et al., 2018a, Bau et al., 2019b, Vig et al, 2020b, Soulos et al, 2020b, Ravfogel et al, 2020b, Elazar et al, 2020, Besserve et al., 2020b, Geiger et al, 2020a, Csordás et al, 2021, Geiger et al, 2021c, Meng et al, 2022a]. In the context of neural network analysis, this provides a powerful tool-kit for understanding a model’s causal structure, since an enormous number of diverse and finely controlled intervention experiments can be performed. We build on these methods, extending them to the optimization process. Multi-Task Training Multi-task training is the practice of jointly training a model against a set of learning tasks to improve data efficiency and increase model robustness [Ruder, 2017, Zhang and Yang, 2017, Crawshaw, 2020]. This can be thought of in
terms of supervised probing In standard supervised probing, one trains the probe using internal representations from the target model while keeping the target model frozen. In multi-task training, we allow the target model’s parameters to be changed by the probing process. This provides a natural point of comparison with our proposal for IIT, where we use our target symbolic causal model to define multi-task training objectives. Data Augmentation Data augmentation is the practice of enhancing training sets by modifying existing examples to generate new ones [Perez and Wang, 2017, Shorten and Khoshgoftaar, 2019, Kaushik et al., 2019, Liu et al, 2021] For us, data augmentation is another natural comparison point because we can use a target symbolic causal model to generate additional data. Crucially, IIT involves interchanging internal network representations, while data augmentation methods only involve the creation of inputs. 7.3 Interchange Intervention Training Our goal is to
train a neural network to have an internal causal structure that realizes a high-level causal model. To concretize this goal, we draw on two strands of work on causality: (1) formal interventionist theories of causality [Spirtes et al., 2001, Pearl, 2001], in which causal processes are CHAPTER 7. INDUCING CAUSAL STRUCTURE FOR INTERPRETABLE NEURAL NETWORKS110 associated with the effect of interventions, and (2) theories of abstraction [Beckers and Halpern, 2019b, Beckers et al., 2019, Chalupka et al, 2016b, Rubenstein et al, 2017a], where relationships between two causal processes are determined by the presence of systematic correspondences between the effects of interventions. The key insight is that having a particular causal structure is a matter of satisfying a number of counterfactual statements about the effect of interventions [Hitchcock, 2001]. The present section defines this process formally, and Figure 7.1 illustrates all the concepts with a self-contained example.
Structural Causal Models We introduce a minimal notation for structural causal models here. We define a structural causal model M to consist of variables V, and, for each variable V ∈ V, a set of values Val(V ), a set of parents PAV , and a structural equation FV that sets the value of V based on the setting of its parents. We denote the set of variables with no parents as VIn and those with no children VOut . A structural causal model M = (V, PA, Val, F ) can represent both symbolic computations and neural networks. Given a setting of an input ∈ Val(VIn ) and variables V ⊆ V, we define GetVals(M, input, V) ∈ Val(V) to be the setting of V determined by the setting input and model M. For example, V could correspond to a layer in a neural network, and GetVals(M, input, V) then denotes the particular values that V takes on when the model M processes input. For a set of variables V and a setting for those variables v ∈ Val(V), we define MV←v to be the causal model identical to
M, except that the structural equations for V are set to constant values v. Because we overwrite neurons with v in-place, gradients can back-propagate through v This corresponds closely to the do operator of Pearl [2001], which characterizes interventions on models in the service of exploring hypothetical or counterfactual states. Interchange Interventions With the above definitions in place, we can straightforwardly characterize the interchange interventions of Geiger et al. [2020a], in which a model M is used to process two different inputs, source and base, and then a particular internal state obtained by processing source is used in place of the corresponding internal state obtained by base. For a given set of variables V, MV←GetVals(M,source,V) is a version of M with the values of V set to those obtained by processing source. In addition, GetVals(M, base, VOut ) is the setting of the outputs VOut obtained by processing base with model M. When we put these two steps together, we
obtain the interchange intervention: def II(M, base, source, V) = GetVals(MV←GetVals(M,source,V) , base, VOut ) (7.1) CHAPTER 7. INDUCING CAUSAL STRUCTURE FOR INTERPRETABLE NEURAL NETWORKS111 In short, the interchange intervention provides the output of the model M for the input base, except the variables V are set to the values they would have if source were the input. Causal Abstraction Relationships Suppose we have a high-level model MH and a low-level model ML with identical input spaces and a predetermined mapping of output values from the low to high level, κ (for example, if the low level model produces a probability distribution over output classes, then κ could be the argmax function, which selects the highest probability class). Further suppose we have an alignment Π mapping intermediate variables in VH to non-overlapping ∗ subsets of variables in VL . Consider some intermediate variable VH and define MH to be MH with every variable marginalized other than VIn ,
VOut , and VH . We can use the definition of interchange ∗ interventions to define what it means for ML and MH to be in a causal abstraction relationship, namely, for all b, s ∈ VIn : II(MH , b, s, VH ) = ∗ κ(II(ML , b, s, Π(VH ))) (7.2) This is in fact a constructive abstraction relationship in the sense of Beckers and Halpern [2019b], in which aligned interventions on the low-level model and high-level model have the same effect. This is especially suited for situations in which we seek to relate small symbolic models with large neural models with high-dimensional representations. Abstraction and Interpretability Causal abstraction analysis is not a story about the reasoning a neural network might use to achieve its behavior, but instead is an intervention-based method that determines how it does, in fact, achieve its behavior. We can interpret the semantic content of neural representations using the high-level variables they are aligned with, and understand how those
neural representations are composed using the high-level parenthood relation. Simply put, when a high-level causal model is an abstraction of a neural network, it is a faithful interpretation [Lipton, 2018a, Jacovi and Goldberg, 2020b] of the network. Interchange Intervention Accuracy To quantify partial success when it comes to causal abstraction relationships, we measure the percentage of aligned interchange interventions that produce the same output, reporting this as the interchange intervention accuracy (IntInvAcc): def IntInvAcc(MH , ML , VH , Π) = 1 ∗ ∑ I[II(MH , b, s, VH ) = ∣V⅁⋖(VIn )∣2 b,s∈V⅁⋖(V ) In κ(II(ML , b, s, Π(VH )))] (7.3) Where every pair of inputs b and s is considered, and IntInvAcc is 1, the two models are in the causal abstraction relationship. However, we often only approximate this by evaluating a set of randomly sampled pairs of inputs, due to the enormous space of input pairs. CHAPTER 7. INDUCING CAUSAL STRUCTURE FOR INTERPRETABLE
NEURAL NETWORKS112 IntInvAcc provides a natural metric for quantifying the interpretability of a neural network in the following sense: when IntInvAcc is 1, the causal model is an explanation of how the network behaves, providing a clear window into the network itself. In practice, we rarely observe perfect IntInvAcc in complex networks, but we can still say that the higher the value of IntInvAcc, the more we have license to reason about the high-level causal model instead of reasoning directly about the low-level network. The causal model provides an interpretable proxy for the network itself IIT Loss Functions The definition of IIT for high-level models with one intermediate variable falls out directly from the causal abstraction definition: ∑ Loss(II(C, b, s, V ), θ II(N , b, s, Π(V ))) (7.4) b,s∈VIn θ where C is the high-level causal model, V is a high-level variable, N is the low-level neural network with learned parameters θ, Π(V ) is a set of low-level variables
(neurons) that are aligned with V , and Loss is some loss function. Observe that we do not apply the output map κ, because the loss function takes in the logits directly. The crucial feature of an IIT update is that the interchange intervention intertwines two computation graphs, one generated by the forward pass for the base input and one by the forward pass for the source input. This means that when backpropagation is performed with the IIT loss objective, updates are applied as they are in regular training, starting from the output representation and proceeding towards the input representations. However, when the intervention site is reached, this θ process bifurcates, and weights receive two updates, once from N processing the input base, and θ once from N processing source. In our toy example (Figure 71c), the network is too small to observe this double update, but the networks in our three case studies are not. (See Figure 72, which exemplifies such a process.) θ An
important property of IIT is that, if (7.4) is zero, then C and N stand in the causal abstraction θ relation (7.2) (The reverse does not hold; C can be a causal abstraction of N without the loss being zero. Figure 71 is an example This is a desirable property of the method, since we do not expect our loss functions to be zero in general.) Example Figure 7.1 provides an example of interchange intervention training, in which a θ causal model C∧ of boolean conjunction is aligned with a one-layer linear network N∧ , where θ = {W1 , W2 , b, w}, as in Figure 7.1b θ At the start, N∧ is perfect in terms of its input–output behavior but does not conform to the counterfactual behavior of C∧ . In other words, the regular behavioral learning objective is met, but the interchange intervention training objective is not; interchange intervention accuracy ((7.3)) is 81.25 One interchange intervention training update (Figure 7.1c) results in a network that satisfies CHAPTER 7.
INDUCING CAUSAL STRUCTURE FOR INTERPRETABLE NEURAL NETWORKS113 False False False True -1 -0.5 -0.45 0.05 0 Y = w[h1 ; h2 ] + b H1 = W1 [x1 , x2 ] 0 0.45 0.05 0.05 0.5 0.5 0.55 O = b1 ∧ b2 H2 = W2 [x1 , x2 ] X1 V1 = b1 X2 V2 = b2 0 0 1 0 0 1 1 1 B2 False False True False False True True True B1 (a) A linear network with unspecified (b) We define a network with initial parameters W1 = [0.45, 005], weights (left) and a symbolic causal model W2 = [0.05, 05], output bias b = −1, and output weights w = that computes boolean conjunction (right). [1, 1] Input values are 0 for False and 1 for True With the initial An alignment between the two is denoted weights, the network has perfect behavioral accuracy, predicting by dashed lines. The causal model is an true (red) iff both its inputs are 1, otherwise it predicts false abstraction of the network when, for both (blue). Although correct when run on the four inputs (T, T), V1 and V2 , aligned
interchange interven- (T, F), (F, T), (F, F), the interchange intervention accuracy is tions on network and causal model result 81.25%: between the two high-level variables V1 and V2 , there are in the same output on all 16 ordered pairs six ordered pairs of inputs where performing aligned interchange of inputs. (The aligned intervention pair interventions results in the causal model and neural network (b, s) in general differs from (s, b).) producing different outputs (see Figure 7.1c for one such pair) F T F T T T F T T F H False False -0.45 -0.05 0.05 0.5 0.45 0.5 0 1 1 0 False True True False (c) An illustration of an interchange intervention training update, where an intervened network is trained to predict the intervened output of the causal model. It can be seen that the intervention puts the network in a state that could not be achieved with any input representation. L 10: 11: 12: 13: 14: 15: aL = GetVals(M , s, VL ) L oL = GetVals(MVL ←aL , b,
VOut ) // pred LIIT = Loss(oH , oL ) L = LIIT + LOthers // combine with other losses L.backward() Update model parameters with gradients (d) Pseudocode for interchange intervention training. False False False True True -0.95 -0.39 -0.33 0.23 0.13 0 0 L Require: High-level and low-level models M and M with variables VH and VL , an alignment Π that maps a VH ∈ VH to a VL ⊆ VL , training dataset D H 1: M .eval() L 2: M .train() 3: while not converged do 4: for (b, s) in enumerate(D × D) do // base and source 5: VH ∼ VH // sample a high-level variable 6: VL = Π(VH ) // aligned low-level variables 7: with no grad: H 8: aH = GetVals(M , s, VH ) H 9: oH = GetVals(MVH ←aH , b, VOut ) // label 0.5 0.05 0.05 0.55 0.55 0.6 0.5 0.55 0 0 1 0 0 1 1 1 1 0 True False True False False True True True True False (e) The network defined in Figure 7.1b after the IIT training update from Figure 71c has been applied, resulting in a network with 100%
interchange intervention accuracy (though still nonzero loss), while maintaining the same behavior. The new network has parameters W1 = [05012, 005], W2 = [005, 05512], bias b = −0.9488, and output weights w = [10231, 10256] Figure 7.1 CHAPTER 7. INDUCING CAUSAL STRUCTURE FOR INTERPRETABLE NEURAL NETWORKS114 7 2 LOGITS 9 9 7 0 9 7 7 2 LOGITS Figure 7.2: An illustration of an IIT update where a neural network (right) is trained to realize a causal model (left) that solves the PVR-MNIST task. Solid lines are feed-forward connections, dashed lines are interchange interventions, red lines are the flow of backpropagation. Observe that when backpropagation reaches the interchange intervention, it flows into both the source input’s computation graph and the base input’s graph, updating the weights below the interchange intervention twice. θ both objectives (Figure 7.1e): N∧ now stands in the causal abstraction relation to C∧ (interchange intervention accuracy is
now 1). 7.4 MNIST Pointer-Value Retrieval Our first benchmark is MNIST Pointer-Value Retrieval (MNIST-PVR; Zhang et al. 2021), a visual reasoning task constructed using the MNIST dataset [LeCun et al., 2010] An input i = (iTL , iTR , iBL , iBR ) consists of four MNIST images (handwritten digits) arranged in a grid. The top left image iTL acts as a pointer that picks out one of the three other images. Symbolic Causal Structure Our target causal model will abstract away from the details of how to identify the handwritten digit in an image, focusing just on the reasoning about pointers. Formally, we define a causal model CPVR = (V, PA, Val, F ) that computes the label for each of the four MNIST images using an oracle OMNIST with a look-up table to select the correct label based on the pointer. The variables are V = {ITL , ITR , IBL , IBR , YTL , YTR , YBL , YBR , O} and the values assigned by Val are the MNIST training images for the four input variables ITL , ITR , IBL , IBR , and
the set of numbers 0–9 for all other variables. The parents are defined such that PAIw = ∅ and PAYw = {Iw } for all CHAPTER 7. INDUCING CAUSAL STRUCTURE FOR INTERPRETABLE NEURAL NETWORKS115 w ∈ {TR, TL, BR, BL}, and PAO = {YTL , YTR , YBL , YBR }. The structural equations are FYTL (iTL ) = OMNIST (iTL ) FYTR (iTR ) = OMNIST (iTR ) FYBL (iBL ) = OMNIST (iBL ) FYBR (iBR ) = OMNIST (iBR ) FO (yTL , yTR , yBL , yBR ) = ⎧ ⎪ yTR yTL ∈ {0, 1, 2, 3} ⎪ ⎪ ⎪ ⎪ ⎪ ⎨ ⎪yBL yTL ∈ {4, 5, 6} ⎪ ⎪ ⎪ ⎪ ⎪ ⎩yBR yTL ∈ {7, 8, 9} Systematic Generalization The train/test split designed by Zhang et al. [2021] creates a distributional shift between the training and testing data by removing training examples where either OMNIST (iTR ) ∈ {1, 2, 3}, OMNIST (iBL ) ∈ {4, 5, 6}, or OMNIST (iBR ) ∈ {0, 7, 8, 9}. This evaluates where models can systematically generalize, learning the general structure of the problem rather than memorizing many special cases. Neural
Network We trained ResNet18 from PyTorch vision. This is the deep residual network [He et al., 2016] baseline used by Zhang et al [2021] on the MNIST-PVR dataset, and we adopt their θ hyperparameters. We call this model NPVR , where θ abbreviates the parameters θ Alignments In our experiments, we align the neural representations of NPVR with the symbolic variables of CPVR by partitioning the layer resulting from the first application of max-pooling into quadrants QTL , QTR , QBL , QBR which are aligned with the variables YTL , YTR , YBL , YBR . In initial experimentation, we found that the layers must be partitioned such that each quadrant is directly above its corresponding input. This is likely due to the locality of convolution operators We also found that aligning layers closer to the classifier head was ineffective. Interchange Intervention Training For each intermediate variable Yw ∈ {YTL , YTR , YBL , YBR }, θ w we introduce an IIT objective that optimizes for NPVR
implementing CPVR the submodel of CPVR where the three intermediate variables that aren’t Yw are marginalized out: ∑ w CE(II(CPVR , b, s, Yw ), b,s∈MNIST-PVR θ II(NPVR , b, s, Qw ))) (7.5) where CE is the cross-entropy loss and MNIST-PVR is the dataset. We visualize an IIT update to θ NPVR in Figure 7.2 Typed Interchange Intervention Training We make further use of the causal model by observing that the intermediate variables YTL , YTR , YBL , YBR can be treated as the same type. They all share a value space, as do the neural representations QTL , QTR , QBL , QBR . This means we can perform interchange interventions between different variables and extend our training objective to CHAPTER 7. INDUCING CAUSAL STRUCTURE FOR INTERPRETABLE NEURAL NETWORKS116 Behavioral Accuracy Train Test IIT Accuracy Train Test Standard IIT Multi IIT + Multi 99.10 99.60 99.64 99.60 0.00 93.93 0.00 96.01 88.80 2060 99.00 9485 89.35 2050 99.10 9664 Augment No Typing 99.40 99.41 90.90
0.09 98.90 99.47 Training 92.00 16.88 θ Table 7.1: Results for NPVR (ResNet18) trained on the PVR-MNIST dataset Behavioral accuracy θ is the percentage of inputs that NPVR agrees with CPVR on. Interchange intervention accuracy quantifies the extent to which the interpretable causal model is a proxy for the network (Section 7.3) IIT delivers the best results, especially when combined with multi-task objectives. these interventions as well: def T-II(M, b, s, V, V ) = ′ GetVals(MV′ ←GetVals(M,s,V) , b, VOut ) (7.6) ∑ CE(T-II(CPVR , b, s, Yw , Yw′ ), w,w′ ∈{TL, TR, BL, BR} b,s∈PVR-MNIST θ T-II(NPVR , b, s, Qw , Qw′ )) (7.7) Multi-Task Objectives To compare against multi-task objectives, we train models to predict the value of intermediate variables from the aligned neural representations, backpropagating into the φ weights of the target model. Specifically, we train four linear classifiers P w on the loss φ θ ∑ CE(P w (GetVals(NPVR [θ], input, Qw
)), input∈PVR-MNIST w∈{TL,TR,BL,BR} GetVals(CPVR , input, Yw )) (7.8) where the trained parameters are θ, the parameters of ResNet and φw , the parameters of the linear classifiers. Data Augmentation We perform data augmentation by randomly sampling two examples and swapping a random quadrant of the base input with a random quadrant of the source input to produce a new example that is then labeled with CPVR . This procedure is guided by the same causal structure used by our other models, but it is by definition restricted to input manipulations. Results Our results are in Table 7.1 The behavioral accuracy is the standard metric, while the interchange intervention accuracy captures whether the symbolic causal model is an abstraction of CHAPTER 7. INDUCING CAUSAL STRUCTURE FOR INTERPRETABLE NEURAL NETWORKS117 Behavioral Exact Match % Training a1 . a2 x P∆ an y a2 . an h1 . hn . ec an−1 a0 ec a1 P∆ Pt a1 h0 Pa TSize TColor TShape ec e1 . eShape en
Novel color Novel size Novel direction Novel length Standard Multi IIT IIT+ Multi 55.98 (631) 76.91 (502) 74.12 (600) 80.37 (088) 41.67 (624) 39.46 (768) 65.65 (426) 74.84 (004) 0.00 (000) 0.00 (000) 0.26 (014) 14.72 (354) 5.72 (344) 9.05 (528) 10.20 (608) 25.82 (037) Standard Multi IIT IIT+ Multi 44.26 (276) 68.42 (020) 70.63 (933) 70.73 (686) Interchange Intervention Exact Match % 35.57 (264) 45.83 (245) 65.18 (284) 75.34 (091) 0.00 (000) 0.00 (000) 5.24 (307) 11.79 (257) 0.30 (021) 0.19 (005) 4.75 (206) 8.49 (153) CNN (b) Results for the CNN-LSTM on the ReaSCAN systematic generalization tasks. Only models that use IIT are ICommand IWorld COMMAND GRID WORLD able to consistently get traction on these tasks, and once again we see that IIT combines effectively with multi-task (a) The causal model that solves ReaSCAN (left) objectives, in both standard behavioral evaluations and and the neural CNN-LSTM model trained on evaluations that seek to quantify the extent to which
the ReaSCAN (right). Dashed lines align variables high-level causal model serves as an interpretable proxy for in the causal model with neural representations. the network Bi-LSTM Figure 7.3 the neural network. Neither the standard nor multi-task models learned the behavioral objective in a way that generalizes, with total failure on the testing data (0%). On the other hand, IIT solves the generalization task (93.93%) However, multi-task training does synergize with IIT, producing the model with the best performance (96.01%) Data augmentation lessens the distributional shift; however, the distributions remain skewed and model performance does not exceed 90.90% Our interchange intervention test set accuracies tell a similar story. Neither the standard nor multi-task models learned the IIT objective in a way that generalizes, with total failure on the testing data (20.60% and 2050%, respectively) On the other hand, IIT learns a general solution to the interchange intervention
objectives, achieving accuracy on the test data (94.85%) Again, multi-task training synergizes with IIT, producing the model with the best performance on the IIT objective (96.64%) The causal model CPVR is a near perfect abstraction of our best model, meaning the seemingly opaque and complex network dynamics have an interpretable and faithful abstract structure given by CPVR . We can see that Resnet has an inherently modular architecture from the fact that standard training produces a model with quite high (88.80%) interchange intervention accuracy on the training data. However, without any structural training objectives, ResNet does not generalize this modular solution to test data (20.60%) We believe this modularity is the result of convolutions being operations that preserve locality of information across layers. When the distributional shift between training and testing is lessened by data augmentation, the ResNet model produces a model with near perfect (98.90%) interchange
intervention accuracy on the training data, which generalizes better to test data (92.00%) (but is still out performed by IIT) CHAPTER 7. INDUCING CAUSAL STRUCTURE FOR INTERPRETABLE NEURAL NETWORKS118 Without our typed IIT objectives, behavioral and interchange intervention accuracy plummets on the test data. Typing our variables is crucial for generalization 7.5 Navigation and Language (ReaSCAN) Our second benchmark is ReaSCAN [Wu et al., 2021b], a synthetic command-based navigation task that builds off the SCAN [Lake and Baroni, 2018a] and gSCAN [Ruis et al., 2020] benchmarks The goal is to predict an action sequence for the agent to reach the referred target and operate on it given a command and a grid world. For simplicity, we experiment with the simplest command structure included in ReaSCAN, which excludes any relative clauses. Symbolic Causal Structure Our causal model CReaSCAN = (V, PA, Val, F ) (see Figure 7.3a bottom) is an oracle solver for ReaSCAN that (1) parses the
language command, identifying size, color, and shape properties of the target shape, (2) computes the location of the target object from these properties and the grid world, (3) calculates the horizontal and vertical distances from the agent to the target, and, finally, (4) emits an action sequence that brings the agent to the target (We condense the action sequence to a single output variable). Formally, we define variables and values x y V = {ICom ,IWorld , TSize , TColor , TShape , Pt , Pa , P∆ , P∆ , O} Val(TShape ) = {circle, square, cylinder} Val(TColor ) = {red, green, blue, yellow} Val(TSize ) = {small, big} x/y Val(Pt ) = Val(Pa ) = Val(P∆ ) = {−5, . , 5} with the values Val(ICom ), Val(Iworld ), and Val(O) being equal to the command space, world space, and action sequence space. The parents are defined according to the topology of directed arrows pointing from parents to children in Figure 7.3a The structural equations for object properties, FTSize (iCom ),
FTColor (iCom ), and FTShape (iCom ), are determined by parsing and interpreting the input language command. The structural equations for position look-ups FPt (tSize , tColor , tShape , iWorld ) and FPa (iWorld ) determine the target object and agent location from the target object properties and the input world. The position deltas FPx∆ (pt , pa ) and FPy∆ (pt , pa ) are determined to be the horizontal and vertical distance between the target object y and agent, respectively. Finally, the equation for the output FO (P∆ , P∆ , icommand ) is the action x sequence that takes the agent to the target object in the correct manner of movement, as determined by the vertical and horizontal distances between the two and the adverb in the command. Systematic Generalization ReaSCAN includes testing examples that are systematically different from training examples. Performance on those test sets provides insights into a model’s capabilities to generalize to unseen composites of seen
concepts in a zero-shot fashion. In this experiment, we CHAPTER 7. INDUCING CAUSAL STRUCTURE FOR INTERPRETABLE NEURAL NETWORKS119 generate four unseen testing splits investigating two distinct generalization patterns by adapting ReaSCAN’s data generation framework. We investigate two splits focusing on novel attribute compositions in input commands (Novel color and Novel size), and two splits focusing on novel compositions in output action sequences (Novel direction and Novel length). Neural Network We use the original baseline model for ReaSCAN [Wu et al., 2021b] as our θ θ neural model NCNN-LSTM . NCNN-LSTM is a multimodal sequence-to-sequence model which takes in a command and a grid world, and predicts an action sequence as shown in Figure 7.3a θ Alignments In our experiments, we align neural representations of NCNN-LSTM with the variables y x TSize , TColor , TShape , P∆ , and P∆ , in CReaSCAN . We choose the neural representation eShape output by the LSTM encoder
above the noun token (e.g, “circle”), which has 75 dimensions, to be evenly partitioned into three chunks of 25 dimensions, which are aligned with the target properties TSize , TColor , and TShape . For the position deltas, we choose the initial hidden representation h0 of the decoder LSTM, which has 100 dimensions, to be sliced into two evenly partitioned 50 dimension y chunks where the first chunk represents the position difference by row P∆ , and the second chunk x represents the position difference by column P∆ . We train the network to derive from the world and the command the horizontal and vertical distances between the target and agent, storing the horizontal distance in one half of h0 and the vertical distance in the other. Interchange Intervention Training For each variable V in CReaSCAN aligned with neurons θ θ NV in NCNN-LSTM , we introduce an IIT objective that optimizes for NCNN-LSTM implementing the V marginalized submodel CReaSCAN : ∑ θ CEAction
(II(NCNN-LSTM , b, s, NV ), b,s∈ReaSCAN V II(CReaSCAN , b, s, V )) (7.9) where CEAction is the cross-entropy loss over each action token prediction over the complete action sequence. Multi-task Objectives Similar to MNIST-PVR, we train small models to predict the position offsets between the target and the agent from the aligned neural representations. Specifically, for y each V ∈ {TSize , TColor , TShape , P∆ , P∆ }, we train a single-layer linear classifier P V on the loss x φ φ θ ∑ CEPosition (P V (GetVals(NCNN-LSTM , i, NV ), i∈ReaSCAN GetVals(CReaSCAN , i, V )) (7.10) where the trained parameters are θ, the parameters of the CNN-LSTM, and φV , the parameters of the classifiers. Results Our results are shown in Table 7.3b We use exact matches of action sequences as our evaluation metric for the behavioral and interchange intervention tasks. CHAPTER 7. INDUCING CAUSAL STRUCTURE FOR INTERPRETABLE NEURAL NETWORKS120 We begin with our results on the
behavioral task. Standard training produces models that fail to generalize across all four tasks. IIT alone out-performs multi-task training on novel sizes and lengths, and performs similarly on novel colors and lengths. Again, we observe that IIT and multi-task synergize, producing the models that best generalize across all tasks. Overall, IIT is essential to this systematic generalization task. Our interchange intervention accuracy results suggest that IIT delivers models that best conform to the interpretable causal model. Without any IIT objectives, both the standard and multi-task models achieve non-zero interchange intervention accuracy only for the two easier splits: novel colors and novel size. IIT achieves significant improvements over these two tasks and gets traction on the two more difficult ones, novel direction and novel length. And, once again, combining IIT with multi-task training delivers the best model by wide margins on all four tasks. 7.6 Natural Language
Inference (MQNLI) Our final benchmark is MQNLI Geiger et al. [2019a], a synthetic natural language inference dataset where the task is to label the semantic relation between two sentences as enailment, contradiction, or neutral. Here is an example: ε every ε baker ε ε happily eats ε some stale bread contradiction ε some angry baker does not ε eat ε some ε bread where an ε denotes the absence of a word and is used to align corresponding words in the two sentences. Geiger et al. [2020a] fine-tuned a BERT model on MQNLI, achieving state-of-the-art results (≈90% test accuracy). Their interchange intervention analysis revealed that this model learns to partially represent the relation between aligned subphrases in the two sentences (e.g stale bread and ε bread in the example above) We hypothesize that if we teach BERT to fully represent this information using IIT, the task will be solved perfectly. Symbolic causal structure Geiger et al. [2020a] define a causal model that
computes the relation between aligned phrases in QPObj order to compute the relation between two sentences. We narrow our focus to the submodel CNatLog , which (1) contains a single intermediate variable QPObj that computes the relation between the quantified verb phrase (= adverb + verb + quantified object noun phrase) of each sentence, and (2) uses this to infer the relation between the two sentences. To label the MQNLI example above, QPObj CNatLog would compute that QPObj = ⊏ because “happily eats ε some stale bread ” entails “eat ε some ε bread ”. Then, this information is used infer that the relation between the sentences is contradiction. CHAPTER 7. INDUCING CAUSAL STRUCTURE FOR INTERPRETABLE NEURAL NETWORKS121 Systematic Generalization The train-test split of MQNLI is constructed to be as difficult as possible while still being solved by a compositional memorization-based learning model. This makes the task hard, but fair. θ Neural Network We fine-tune a
pretrained BERT model NNLI on MQNLI. The architecture consists of 12 transformer layers that create a neural representations for each token in the input; the grid of neural representations has a column for each token and a row for each layer. Alignments In our experiments, we align QPObj with the neural representations above the verb in the first sentence from BERT layers {0, 2, 4, 6, 8, 10}. QPObj θ Interchange intervention training For QPObj in CNatLog aligned with neurons N in NNLI , we introduce an IIT objective that optimizes for NNLI implementing the marginalized submodel QPObj CNatLog : θ ∑ CE(II(NNLI , b, s, N), b,s∈MQNLI QPObj II(CNatLog , b, s, QPObj )) (7.11) φ Multi-task Objectives We train a linear classifier P to predict the value of QPObj with loss: ∑ φ θ CE(P (GetVals(NNLI , i, N), i∈MQNLI QPObj GetVals(CNatLog , i, QPObj )) (7.12) where the trained parameters are θ, the parameters of BERT, and φ, the parameters of the classifier. Data
Augmentation We perform data augmentation by randomly sampling two examples and replacing the quantified verb phrase from the first example with those from the second in order to QPObj produce a new example that is then labeled with CNatLog . This procedure is guided by the same causal structure used by our other models, but it is by definition restricted to input manipulations. Results We compute the accuracy on the basic behavioral task of predicting MQNLI labels, and interchange intervention (IIT) accuracy, where we compute the percentage of cases where performing QPObj an intervention on the neural model produced a same change in output as the submodel CNatLog . QPObj Our results, shown in Figure 7.4, demonstrate that IIT training on CNatLog solves MQNLI with near perfect accuracy and IIT accuracy (≈100%). Furthermore, we see that aligning QPObj to the first few layers of BERT results in lower accuracy on MQNLI, with layer 6 and onward resulting in near perfect accuracy and
interchange intervention accuracy. When QPObj is aligned with layer 6 or later, we again see multi-task training synergizing with IIT to produce the best models. Data augmentation results in models with perfect behavioral performance. This is unsurprising, as data augmentation removes the out-of-domain generalization problem. IIT is needed to produce a model with an interpretable solution. CHAPTER 7. INDUCING CAUSAL STRUCTURE FOR INTERPRETABLE NEURAL NETWORKS122 Figure 7.4: Performance of a pretrained BERT natural language inference model fine-tuned on the QPObj MQNLI dataset with the causal model CNatLog from Geiger et al. [2020a] We report the results on the evaluation set. While data augmentation leads to consistently excellent behavior accuracy (left) panel, it has very low interchange intervention accuracy. In other words, IIT is necessary for an interpretable model with high-performance. 7.7 Conclusion We introduced interchange intervention training as a method to imbue
neural networks with interpretable, systematic causal structure, and we conducted three case studies with IIT: a vision task (MNIST-PVR), a grounded language understanding task (ReaSCAN), and a natural language inference task (MQNLI). In all settings, models trained with IIT perform best in standard (but very challenging) behavioral evaluations and prove to be the most interpretable in the sense that they conform best to our high-level causal models of the tasks. In addition, our results show that IIT is easily combined with multi-task objectives that further strengthen the results. These initial findings suggest that IIT is a flexible and powerful way to bring high-level insights about causal structure into a data-driven learning process. Chapter 8 Causal Distillation for Language Models Abstract Distillation efforts have led to language models that are more compact and efficient without serious drops in performance. The standard approach to distillation trains a student model
against two objectives: a task-specific objective (e.g, language modeling) and an imitation objective that encourages the hidden states of the student model to be similar to those of the larger teacher model. In this paper, we show that it is beneficial to augment distillation with a third objective that encourages the student to imitate the causal dynamics of the teacher through a distillation interchange intervention training objective (DIITO). DIITO pushes the student model to become a causal abstraction of the teacher model – a faithful model with simpler causal structure. DIITO is fully differentiable, easily implemented, and combines flexibly with other objectives. Compared against standard distillation with the same setting, DIITO results in lower perplexity on the WikiText103M corpus (masked language modeling) and marked improvements on the GLUE benchmark (natural language understanding), SQuAD (question answering), and CoNLL-2003 (named entity 1 recognition). 8.1
Introduction Large pretrained language models have improved performance across a wide range of NLP tasks, but can be costly due to their large size. Distillation seeks to reduce these costs while maintaining performance by training a simpler student model from a larger teacher model Hinton et al. [2015], Sun et al. [2019], Sanh et al [2019], Jiao et al [2019] 1 We release our code at https://github.com/frankaging/Causal-Distill 123 CHAPTER 8. CAUSAL DISTILLATION FOR LANGUAGE MODELS 124 Hinton et al. [2015] propose model distillation with an objective that encourages the student to produce output logits similar to those of the teacher while also supervising with a task-specific objective (e.g, sequence classification) Sanh et al [2019], Sun et al [2019], and Jiao et al [2019] adapt this method, strengthening it with additional supervision to align internal representations between the two models. However, these approaches may push the student model to match all aspects of internal
states of the teacher model irrespective of their causal role in the network’s computation. This motivates us to develop a method that focuses on aligning the causal role of representations in the student and teacher models. We propose augmenting standard distillation with a new objective that pushes the student to become a causal abstraction [Beckers and Halpern, 2019b, Beckers et al., 2019, Geiger et al, 2021c] of the teacher model: the simpler student will faithfully model the causal effect of teacher representations on output. To achieve this, we employ the interchange intervention training (IIT) method of Geiger et al. [2022c] The distillation interchange intervention training objective (DIITO) aligns a high-level student model with a low-level teacher model and performs interchange interventions (swapping of aligned internal states); during training the high-level model is pushed to conform to the causal dynamics of the low-level model. Figure 8.1 shows a schematic example of
this process Here, hidden layer 2 of the student model (bottom) is aligned with layers 3 and 4 of the teacher model. The figure depicts a single interchange intervention replacing aligned states in the left-hand models with those from the right-hand models. This results in a new network evolution that is shaped both by the original input and the interchanged hidden states. It can be interpreted as a certain kind of counterfactual as shown in Figure 81: what would the output be for the sentence “I ate some ⟨MASK⟩.” if the activation values for the second token at the middle two layers were set to the values they have for the input “The water ⟨MASK⟩ solid.”? DIITO then pushes the student model to output the same logits as the teacher, i.e, matching the teacher’s output distribution under the counterfactual setup. To assess the contribution of distillation with DIITO, we begin with BERTBASE [Devlin et al., 2019b] and distill it under various alignments between student
and teacher while pretraining on the WikiText-103M corpus Merity et al. [2016] achieving −224 perplexity on the MLM task compared to standard DistilBERT trained on the same data. We then fine-tune the best performing distilled models and find consistent performance improvements compared to standard DistilBERT trained with the same setting on the GLUE benchmark (+1.77%), CoNLL-2003 name-entity recognition (+0.38% on F1 score), and SQuAD v11 (+246% on EM score) 8.2 Related Work Distillation was first introduced in the context of computer vision [Hinton et al., 2015] and has since been widely explored for language models [Sun et al., 2019, Sanh et al, 2019, Jiao et al, 2019] For CHAPTER 8. CAUSAL DISTILLATION FOR LANGUAGE MODELS I I ate ate some some 125 pizza froze Logits Logits ¡MASK¿ . The water ¡MASK¿ salad froze Logits Logits ¡MASK¿ . The water ¡MASK¿ solid . solid . Figure 8.1: An IIT update in the context of masked language modelling
(MLM) The teacher network (top) has 6 layers and the student (bottom) has 3 layers, and we align layer 2 in the student with layers 3–4 in the teacher. Solid lines are feed-forward connections, red lines show the flow of backpropagation, and dashed lines indicate interchange interventions. In this case, the student originally predicted the token “salad” under the interchange intervention, while the teacher predicted the token “pizza” under an aligned interchange intervention. DIITO trains the student to minimize the divergence between the student logits and the teacher logits under the interchange intervention. This updates the student to conform to causal dynamics of the teacher. example, Sanh et al. [2019] propose to extract information not only from output probabilities of the last layer in the teacher model, but also from intermediate layers in the fine-tuning stage. Recently, Rotman et al. [2021] adapt causal analysis methods to estimate the effects of inputs on
predictions to compress models for better domain adaptation. In contrast, we focus on imbuing the student with the causal structure of the teacher. Interventions on neural networks were originally used as a structural analysis method aimed at illuminating neural representations and their role in network behavior [Feder et al., 2021, Pryzant et al., 2021b, Vig et al, 2020b, Elazar et al, 2020, Giulianelli et al, 2020, Geiger et al, 2020a, 2021c] Geiger et al. [2022c] extend these methods to network optimization We contribute to this existing research by adapting intervention-based optimization to the task of language model distillation. CHAPTER 8. CAUSAL DISTILLATION FOR LANGUAGE MODELS 126 Algorithm 1 Causal Distillation via Interchange Intervention Training y Require: Student model S, teacher model T , student output neurons NS , alignment Π, shuffled training dataset D. 1: S.train() 2: T .eval() ′ 3: D = random.shuffle(D) y y 4: NT = Π(NS ) 5: while not converged do ′ 6:
for {x1 , y1 }, {x2 , y2 } in iter(D, D ) do 7: NS = sample student neurons() 8: NT = Π(NS ) 9: with no grad: 10: Ta = SetVals( 11: T , NT , GetVals(T , x1 , NT )) y 12: oT = GetVals(Ta , x2 , NT ) 13: Sa = SetVals( 14: S, NS , GetVals(S, x1 , NS )) y 15: oS = GetVals(Sa , x2 , NS ) DIITO 16: L = get loss(oT , oS ) 17: Calculate LMLM , LCE , LCos DIITO 18: L = LMLM + LCE + LCos + L 19: L.backward() 20: Step optimizer 21: end while 8.3 Causal Distillation Here, we define our distillation training procedure. See Algorithm 1 for a summary GetVals. The GetVals operator is an activation-value retriever for a neural model Given a neural model M containing a set of neurons N (an internal representations) and an appropriate input x, GetVals(M, x, N) is the set of values that N takes on when processing x. In the case that N represents the neurons corresponding to the final output, GetVals(M, x, N) is the output of model M when processing x (i.e, output from a standard forward call of a
neural model) SetVals. The SetVals operator is a function generator that defines a new neural model with a computation graph that specifies an intervention on the original model M [Pearl, 2001, Spirtes et al., 2001]. SetVals(M, N, v) is the new neural model where the neurons N are set to constant values v. Because we overwrite neurons with v in-place, gradients can back-propagate through v Interchange Intervention. An interchange intervention combines GetVals and SetVals operations. First, we randomly sample a pair of examples from a training dataset (x1 , y1 ), (x2 , y2 ) ∈ x D. Next, where N is the set of neurons that we are targeting for intervention, we define MN1 to abbreviate the new neural model as follows: SetVals(M, N, GetVals(M, x1 , N)) (8.1) CHAPTER 8. CAUSAL DISTILLATION FOR LANGUAGE MODELS 127 This is the version of M obtained from setting the values of N to be those we get from processing input x1 . The interchange intervention targeting N with x1 as the source
input and x2 as the base input is then defined as follows: def II(M, N, x1 , x2 ) = x y GetVals(MN1 , x2 , N ) (8.2) where N are the output neurons. In other words, II(M, N, x1 , x2 ) is the output state we get from y M for input x2 but with the neurons N set to the values obtained when processing input x1 . DIITO. DIITO employs T as the teacher model, S as the student model, D as the training inputs to both models, and Π as an alignment that maps sets of student neurons to sets of teacher neurons. For each set of student neurons NS in the domain of Π, we define DIITO loss as: DIITO def LCE = ∑ CES (II(S, NS , x1 , x2 ), II(T , Π(NS ), x1 , x2 )) (8.3) x1 ,x2 ∈D where CES is the smoothed cross-entropy loss measuring the divergences of predictions, under interchange, between the teacher and the student model. Distillation Objectives. We adopt the standard distillation objectives from DistilBERT Sanh et al. [2019]: LMLM for the task-specific loss for the student model,
LCE for the loss measuring the divergence between the student and teacher outputs on masked tokens, and LCos for the loss measuring the divergence between the student and teacher contextualized representations on masked tokens in the last layer. Our final training objective for the student is a linear combination of the DIITO four training objectives reviewed above: LMLM , LCE , LCos , and LCE . In a further experiment, DIITO we introduce a fifth objective LCos which is identical to LCos , except the teacher and student are undergoing interchange interventions. 8.4 Experimental Set-up 2 We adapt the open-source Hugging Face implementation for model distillation [Wolf et al., 2020] We distill our models on the MLM pretraining task Devlin et al. [2019b] We use large gradient accumulations over batches as in Sanh et al. [2019] for better performance Specifically, we distill all models for three epochs for an effective batch size of 240. In contrast to the setting of 4K per batch
in BERT [Devlin et al., 2019b] and DistilBERT [Sanh et al, 2019], we found that small effective batch size works better for smaller dataset. We weight all objectives equally for all experiments 2 https://github.com/huggingface/transformers CHAPTER 8. CAUSAL DISTILLATION FOR LANGUAGE MODELS Layers Pretraining Tokens WikiText Perplexity BERTBASE Devlin et al. [2019b] (Wikipedia+BookCorpus) DistilBERT Sanh et al. [2019] (Wikipedia+BookCorpus) 12 3.3B 10.27 (–) 6 3.3B 17.48 (–) DistilBERT (WikiText) DIITOMIDDLE (WikiText) DIITOLATE (WikiText) DIITOFULL (WikiText) 3 3 3 3 0.1B 0.1B 0.1B 0.1B DistilBERT (WikiText) DIITOMIDDLE (WikiText) DIITOLATE (WikiText) DIITOFULL (WikiText) 6 6 6 6 DIITOFULL +Random (WikiText) DIITOFULL +Masked (WikiText) DIITO DIITOFULL +LCos (WikiText) 6 6 6 Model † † GLUE Score 82.75 (–) 128 CoNLL-2003 acc F1 SQuAD v1.1 EM F1 96.40 (–) 92.40 (–) † † 80.80 (–) 88.50 (–) 77.70 (–) 85.80 (–) 79.59 (–) 98.39
(–) 93.10 (–) 29.51 (032) 26.04 (093) 25.97 (063) 24.85 (058) 67.42 (110) 69.30 (108) 69.01 (169) 69.36 (087) 97.88 (004) 98.03 (004) 98.03 (003) 98.02 (003) 88.89 (029) 89.69 (018) 89.82 (018) 89.67 (016) 26.04 (093) 58.74 (069) 58.75 (049) 58.72 (067) 68.38 (077) 70.23 (057) 70.21 (041) 70.50 (056) 0.1B 0.1B 0.1B 0.1B 15.69 (151) 14.32 (012) 14.93 (023) 13.59 (025) 75.80 (042) 76.71 (047) 76.80 (034) 76.67 (021) 98.48 (003) 98.56 (004) 98.51 (002) 98.53 (004) 92.12 (023) 92.47 (019) 92.36 (027) 92.35 (024) 70.23 (075) 71.93 (031) 71.47 (028) 71.96 (029) 79.99 (055) 81.32 (023) 81.01 (023) 81.33 (025) 0.1B 0.1B 0.1B 13.95 (018) 13.99 (016) 13.45 (019) 76.84 (029) 76.80 (032) 77.14 (037) 98.54 (003) 98.55 (003) 98.54 (004) 92.41 (024) 92.45 (018) 92.35 (024) 71.90 (054) 71.77 (059) 71.94 (031) 81.27 (039) 81.09 (042) 81.35 (023) Table 8.1: Performance on the development sets of the WikiText, GLUE benchmark, CoNLL-2003 corpus for the name-entity recognition
task, and SQuAD v1.1 for the question answering task The score is the averaged performance scores with standard deviation (SD) for all tasks across 15 distinct † runs. Numbers are imputed from released models on Hugging Face [Wolf et al, 2020] With our new objectives, the distillation takes approximately 9 hours on 4 NVIDIA A100 GPUs. Student and Teacher Models. Our two students have the standard BERT architecture, with 12 heads with a hidden dimension of 768. The larger student has 6 layers, the smaller 3 layers Our pretrained teacher has the same architecture, except with 12 layers. Following practices introduced by Sanh et al. [2019], we initialize our student model with weights from skipped layers (one out of four layers) in the teacher model. We use WikiText for distillation to simulate a practical situation with a limited computation budget. We leave the exploration of our method on larger datasets for future research. Alignment. Our teacher and student BERT models create
columns of neural representations above each token with each row created by the feed-forward layer of a Transformer block, as in Figure 8.1 We define LT and LS to be the number of layers in the student and teacher, respectively j j In addition, we define Si and Ti to be the representations in the ith row and jth column in the student and teacher, respectively. An alignment Π is a partial function from student representations to sets of teacher representations. We test three alignments: j j FULL Π is defined on all student representations: Π(Si ) = {T4i+k ∶ 0 ≤ k < LT /LS } j j MIDDLE Π is defined for the row LS //2: Π(SLS //2 ) = {TLT //2 } LATE Π is defined on the student representations in the first and second rows: j j j j Π(S1 ) = {TLT −2 } and Π(S2 ) = {TLT −1 } For each training iteration, we randomly select one aligned student layer to perform the interchange intervention, and we randomly select 30% of token embeddings for alignment for each
sequence. We experiment with three conditions with the FULL alignment: consecutive tokens (DIITOFULL ), random CHAPTER 8. CAUSAL DISTILLATION FOR LANGUAGE MODELS 129 DIITO tokens (DIITOFULL +Random) and masked tokens (DIITOFULL +Masked). We also include LCos DIITO to the FULL alignment (DIITOFULL +LCos 8.5 ). Results Language Modeling. We first evaluate our models using perplexity on the held-out evaluation data from WikiText. As shown in Table 81, DIITO brings performance gains for all alignments DIITO Our best result is from the FULL alignment with the LCos (DIITOFULL +LCos ), which has −2.24 perplexity compared to standard DistilBERT trained with the same amount of data. GLUE. The GLUE benchmark Wang et al [2018] covers different natural language understanding tasks. We report averaged GLUE scores on the development sets by fine-tuning our distilled models in Table 8.1 The results suggest that distilled models with DIITO lead to consistent improvements DIITO over
standard DistilBERT trained under the same setting, with our best result (DIITOFULL +LCos ) being +1.77% higher Named Entity Recognition. We also evaluate our models on the CoNLL-2003 Named Entity Recognition task [Tjong Kim Sang and De Meulder, 2003]. We report accuracy and Macro-F1 scores on the development sets. We fine-tune our models for three epochs Our best performing model (DIITOMIDDLE ) numerically surpasses not only standard DistilBERT (+0.38% on F1 score) trained under the same setting, but also its teacher, BERTBASE (+0.05% on F1 score) Though these improvements are small, in this case distillation produces a smaller model with better performance. Question Answering. Finally, we evaluate on a question answering task, SQuAD v11 [Rajpurkar et al., 2016] We report Exact Match and Macro-F1 on the development sets as our evaluation metrics We fine-tune our models for two epochs. DIITO again yields marked improvements (Table 81) Our best result is from the vanilla FULL
alignment (DIITOFULL ), with +2.46% on standard DistilBERT trained under the same setting. Low-Resource Model Distillation We experiment with an extreme case in a low-resource setting where we only distill with 15% of WikiText, keeping other experimental details constant. Our results suggest that DIITO training is also beneficial in extremely low-resource settings (Figure 8.2) Layer-wise Ablation We further study the effect of DIITO training with respect to the size of the student model through a layer-wise ablation experiment. As shown in Figure 83, we compare GLUE performance for models trained with standard distillation pipeline and with DIITO training (DIITOFULL ). Our results suggest that DIITO training brings consistent improvements over GLUE tasks with smaller models booking the greatest gains. CHAPTER 8. CAUSAL DISTILLATION FOR LANGUAGE MODELS 130 Figure 8.2: Perplexity score distribution for the development set of WikiText of models trained in a low-resource setting. The
best model is the one with the richest alignment structure Figure 8.3: GLUE score distribution across 15 distinct runs of students in different sizes Following the evaluation for BERT Devlin et al. [2019b] we exclude WNLI for evaluation CHAPTER 8. CAUSAL DISTILLATION FOR LANGUAGE MODELS 8.6 131 Conclusion In this paper, we explored distilling a teacher by training a student to capture the causal dynamics of its computations. Across a wide range of NLP tasks, we find that DIITO leads to improvements, with the largest gains coming from the models that use the richest alignment between student and teacher. Our results also demonstrate that DIITO performs on-par (maintaining 97% of performance on GLUE tasks) with standard DistilBERT Sanh et al. [2019] while consuming 97% less training data. These findings suggest that DIITO is a promising tool for effective model distillation Chapter 9 Causal Proxy Models for Concept-based Model Explanations 9.1 Introduction The gold standard
for model explanation methods in AI should be to elucidate the causal role that a model’s representations play in its overall behavior – to truly explain why the model makes the predictions it does. Causal explanation methods seek to do this by resolving the counterfactual question of what the model would do if input X were changed to a relevant counterfactual version ′ X . Unfortunately, even though neural networks are fully observed, deterministic systems, we still encounter the fundamental problem of causal inference [Holland, 1986]: for a given ground-truth ′ input X, we never observe the counterfactual inputs X necessary for isolating the causal effects of model representations on outputs. The issue is especially pressing in domains where it is hard to synthesize approximate counterfactuals. In response to this, explanation methods typically do not explicitly train on counterfactuals at all. In this paper, we show that robust explanation methods for NLP models can be
obtained using texts approximating true counterfactuals. The heart of our proposal is the Causal Proxy Model (CPM). CPMs are trained to mimic both the factual and counterfactual behavior of a black-box model N . We explore two different methods for training such explainers These methods share a distillation-style objective that pushes them to mimic the factual behavior of N , but they differ in their counterfactual objectives. The simpler of these two methods is the input-based CPMIN , which appends to the factual input a new token associated with the counterfactual concept value. This proves remarkably effective. However, we are able to achieve deeper and more human-interpretable explanations with the hidden-state CPMHI , which employs the Interchange Intervention Training (IIT) method of Geiger et al. [2022c] to localize information about the target concept in specific 132 CHAPTER 9. CAUSAL PROXY MODELS FOR CONCEPT-BASED MODEL EXPLANATIONS133 hidden states. We show that both
methods are effective even with very partial causal models of the target domain. Figure 91 provides a high-level overview We evaluate these methods on the CEBaB benchmark for causal explanation methods [Abraham et al., 2022b], which provides large numbers of original examples (restaurant reviews) with humancreated counterfactuals for specific concepts (eg, service quality), with all the texts labeled for their concept-level and text-level sentiment. This counterfactual data is used to uncover the true counterfactual behavior of a model, against which a causal explanation of the model can be benchmarked. We consider two types of approximate counterfactuals derived from CEBaB: texts written by humans to approximate a specific counterfactual, and texts sampled using metadata-guided heuristics. Both strategies lead to state-of-the-art performance on CEBaB for both CPMIN and CPMHI , comparing against all public results on CEBaB as well as new stronger baselines we establish. In addition,
approximate counterfactuals rely only on widely available metadata, which allows for easy use of CPMs in new domains. We additionally identify two other benefits of using CPMs to explain models. First, both CPMIN and CPMHI have factual performance comparable to that of the original black-box model N and can explain their own behavior extremely well. Thus, the CPM for N can actually replace N , leading to more explainable deployed models. Second, CPMHI models localize concept-level information in their hidden representations, which makes their behavior on specific inputs very easy to explain. We illustrate this using Path Integrated Gradients [Sundararajan et al., 2017a], which we adapt to allow input-level attributions to be mediated by the intermediate states that were targeted for localization. Thus, while both CPMIN and CPMHI are effective methods, the qualitative insights afforded by CPMHI models may given them the edge when it comes to explanations. We emphasize that our sole
focus is characterizing model behaviors in human-interpretable ways. This contrasts with work aimed at using language data to make causal inferences about real-world phenomena [Pryzant et al., 2021c] Although these efforts are aligned when we are explaining good models, they are very different when we seek to explain very bad models. For example, if a model is systematically wrong about a real-world phenomenon, the correct explanation of that model will also be systematically wrong about the real-world phenomenon. In such cases, we hope that the explanation helps us understand the model’s failings. For additional discussion, see Feder et al 2020, §2.1 9.2 Related Work Understanding model behavior serves many goals for AI systems, including transparency [Kim, 2015, Lipton, 2018c, Pearl, 2019b, Ehsan et al., 2021], trustworthiness [Ribeiro et al, 2016c, Guidotti et al, 2018, Jacovi and Goldberg, 2020c, Jakesch et al., 2019], safety [Amodei et al, 2016a, Otte, 2013], and fairness
[Hardt et al., 2016b, Kleinberg et al, 2017, Goodman and Flaxman, 2017, Mehrabi et al, CHAPTER 9. CAUSAL PROXY MODELS FOR CONCEPT-BASED MODEL EXPLANATIONS134 2021]. With CPMs, our goal is to achieve explanations that are causally motivated and concept-based, and so we concentrate here on relating existing methods to these two goals. Feature attribution methods estimate the importance of features, generally by inspecting learned weights directly or by perturbing features and studying the effects this has on model behavior [Molnar, 2020, Ribeiro et al., 2016c] Gradient-based feature attribution methods extend this general mode of explanation to the hidden representations in deep networks [Zeiler and Fergus, 2014a, Springenberg et al., 2014, Binder et al, 2016, Shrikumar et al, 2016, Sundararajan et al, 2017a] Concept Activation Vectors (CAVs; Kim et al. 2018, Yeh et al 2020) can also be considered feature attribution methods, as they probe for semantically meaningful directions in the
model’s internal representations and use these to estimate the importance of concepts on the model predictions. While some methods in this space do have causal interpretations (e.g, Sundararajan et al 2017a, Yeh et al 2020), most do not. In addition, most of these methods offer explanations in terms of specific (sets of) features/neurons. (Methods based on CAVs operate directly in terms of more abstract concepts) Intervention-based methods study model representations by modifying them in systematic ways and observing the resulting model behavior. These methods are generally causally motivated and allow for concept-based explanations. Examples of methods in this space include causal mediation analysis [Vig et al., 2020b, De Cao et al, 2021b, Ban et al, 2022], causal effect estimation [Feder et al., 2020, Elazar et al, 2021b, Abraham et al, 2022b, Lovering and Pavlick, 2022c], tensor product decomposition [Soulos et al., 2020b], circuit-based analysis [Cammarata et al, 2020], and
causal abstraction [Geiger et al., 2020a, 2021d, 2023a] CPMs are most closely related to the method of IIT [Geiger et al., 2021d], which extends causal abstraction to optimization Probing is another important class of explanation method. Traditional probes do not intervene on the target model, but rather only seek to find information in it via supervised models [Conneau et al., 2018, Tenney et al, 2019] or unsupervised models [Clark et al, 2019, Manning et al, 2020, Saphra and Lopez, 2019]. Probes can identify concept-based information, but they cannot offer guarantees that probed information is relevant for model behavior [Geiger et al., 2021d] For causal guarantees, it is likely that some kind of intervention is required. For example, Elazar et al [2021b] and Feder et al. [2020] remove information from model representations to estimate the causal role of that information. Our CPMs employ a similar set of guiding ideas but are not limited to removing information. Counterfactual
explanation methods aim to explain model behavior by providing a counterfactual example that changes the model behavior [Goyal et al., 2019b, Verma et al, 2020, Wu et al, 2021a] Counterfactual explanation methods are inherently causal. If they can provide counterfactual examples with regard to specific concepts, they are also concept-based. In addition, some explanation methods train a model making explicit use of intermediate variables representing concepts. Manipulating these intermediate variables at inference time yields causal concept-based model explanations [Koh et al., 2020, Künzel et al, 2019] CHAPTER 9. CAUSAL PROXY MODELS FOR CONCEPT-BASED MODEL EXPLANATIONS135 Evaluating methods in this space has been a persistent challenge. In prior literature, explanation methods have often been evaluated against synthetic datasets [Feder et al., 2020, Yeh et al, 2020] In response, Abraham et al. [2022b] introduced the CEBaB dataset, which provides a human-validated concept-based
dataset to truthfully evaluate different causal concept-based model explanation methods. Our primary evaluations are conducted on CEBaB. 9.3 Causal Proxy Model (CPM) Causal Proxy Models (CPMs) are causal concept-based explanation methods. Given a factual input ′ xu,v and a description of a concept intervention Ci ← c , they estimate the effect of the intervention on model output. The present section introduces our two core CPM variants in detail We concentrate here on introducing the structure of these models and their objectives, and we save discussion of associated metrics for explanation methods for Section 9.4 A Structural Causal Model Our discussion is grounded in the causal model depicted in Figure 9.1a, which aligns well with the CEBaB benchmark Two exogenous variables U and V together represent the complete state of the world and generate some textual data X. The effect of exogenous variable U on the data X is completely mediated by a set of intermediate variables C1 ,
C2 . , Ck , which we refer to as concepts. Therefore, we can think of U as the part of the world that gives rise to these concepts {C}1 . k Using this causal model, we can describe counterfactual data – data that arose under a counterfactual state of the world (right diagram in Figure 9.1a) Our factual text is xu,v , and we use ′ C ←c i xu,v ′ for the counterfactual text obtained by intervening on concept Ci to set its value to c . The ′ C ←c i counterfactual xu,v ′ describes the output when the value of Ci is set to c , all else being held equal. ′ In defining CPMs, we use a wide variety of interventions Ci ← c and values u, v for the exogenous variables U, V . Approximate Counterfactuals i Unfortunately, pairs like (xu,v , xu,v C ←c ′ ) are never observed, and ′ Ci ←c thus we need strategies for creating approximate counterfactuals x̃u,v . Figure 91b describes the two strategies we use in this paper. In the human-created strategy, we rely
on a crowdworker to edit xu,v to achieve a particular counterfactual goal – say, making the evaluation of the restaurant’s food i negative. CEBaB contains an abundance of such pairs (xu,v , x̃u,v ′ C ←c ). However, CEBaB is unusual in having so many human-created approximate counterfactuals, so we also explore a simpler strategy C ←c i in which x̃u,v ′ is sampled with the requirement that it match xu,v on all concepts but sets Ci to ′ c . This strategy is supported in many real-world datasets – for example, the OpenTable reviews underlying CEBaB all have the needed metadata [Abraham et al., 2022b] CHAPTER 9. CAUSAL PROXY MODELS FOR CONCEPT-BASED MODEL EXPLANATIONS136 U =u c1 V =v ⋮ ci ⋮ xu,v Let xu,v be a text written in situation (u, v): V =v c1 ⋮ U =u Ci ← c ⋮ ck C ←c i Human-created x̃u,v ′ C ←c ′ i xu,v Crowdworker edit of xu,v to express that Ci ′ had value c , seeking to keep all else constant. ck (a) A
structural causal model leading to an actual C ←c i Metadata-sampled x̃u,v ′ C ←c i text xu,v and its counterfactual text xu,v . U is an exogenous variable over experiences, c1 , . , ck are mediating concepts, and V is an exogenous variable capturing the writing (and star-rating) experience. At right, we create a counterfactual in which concept Ci takes on a different value. Unfortunately, we cannot truly create such counterfactual situations and so we never observe pairs of texts like these. Thus, we must rely on approximate counterfactuals. xu,v n11 n21 ⋮ n12 ⋯ n22 ⋯ ⋮ n1l n2l ⋮ nd1 nd2 ⋯ ndl p11 p21 ⋮ p12 ⋯ p22 ⋯ ⋮ p1l p2l ⋮ pd1 pd2 ⋯ pdl nout ′ (b) Approximate counterfactuals. In the humancreated strategy, humans revise an attested text to try to express a particular counterfactual, seeking to simulate a causal intervention. In the metadatasampled strategy, we find a separate text that aligns with the original for the value u
insofar as it expresses all the same concepts except for the target concept Ci . C ←c i x̃u,v pout ′ xu,v (c) LMimic : All CPMs (bottom) are trained to mimic the behavior of the neural model N to be explained (top) for all factual inputs xu,v . ′ C =c ′ Ci ←c x̃u,v n11 n21 ⋮ n12 ⋯ n22 ⋯ ⋮ n1l n2l ⋮ nd1 nd2 ⋯ ndl xu′i,v′ n11 ′ n12 ′ ⋯ n′1l n21 ⋮ ′ n22 ⋮ ′ ⋯ n2l⋮ nd1 ′ nd2 ′ ⋯ ndl p11 p21 p31 ⋮ P12 ← p12 pd1 ′ ⋯ p1l P22 ← p22 p32 ⋮ ′ ⋯ p2l ⋯ p3l ⋮ pd2 ⋯ pdl ′ p12 ⋯ p′1l ′ p21 ′ ′ p22 ⋮ ⋯ p2l⋮ ′ pd1 ′ pd2 ′ pdl p12 ⋯ p22 ⋯ ⋮ p1l p2l ⋮ pd1 pd2 ⋯ pdl LHI ∗ ∗ ∗ ∗ pout ∗ ′ ⋯ ′ pout ′ C ←c i (e) LHI : Examples xu,v and x̃u,v are an approximate counterfactual pair. The CPM (middle) is given input xu,v . The objective is for it to mimic LIN xu,v ; tCi ←c′ ′ p11 ⋮ ′ nout ′
nout p11 p21 ⋮ ′ Sampled text expressing that Ci has value c but agreeing with xu,v on all other concepts. LMimic xu,v ′ ′ C ←c i N (top) given x̃u,v , but under the intervention in which specific internal states are changed to those pout ′ C =c ′ C ←c i (d) LIN : Examples xu,v and x̃u,v are an approximate counterfactual pair. The CPM is given xu,v augmented with a special token tCi ←c′ and trained to ′ C ←c i mimic the target model N when its input is x̃u,v . that the CPM computes for input xu′i,v′ (bottom), which is a distinct example that is sampled with the ′ only criteria being that it express Ci = c . The effect of this intervention is to localize information about concept Ci at the intervention site, since the only ′ indication the CPM gets about Ci ← c is via the intervention. Figure 9.1: Causal Proxy Model (CPM) summary Every CPM for model N is trained to mimic the factual behavior of N (LMimic ). For CPMIN , the
counterfactual objective is LIN For CPMHI , the counterfactual objective is LHI . CHAPTER 9. CAUSAL PROXY MODELS FOR CONCEPT-BASED MODEL EXPLANATIONS137 i CPMIN : Input-based CPM Given a dataset of approximate counterfactual pairs (xu,v , x̃u,v C ←c ′ ) and a black-box model N , we train a new CPMIN model P with a counterfactual objective as: ′ C ←c i LIN = CES (N (x̃u,v ), P(xu,v ; tCi ←c′ )) (9.1) where xu,v ; tCi ←c′ in Eqn. 91 denotes the concatenation of the factual input and a randomly initialized ′ learnable token embedding tCi ←c′ describing the intervention Ci ← c . CES represents the smoothed cross-entropy loss [Hinton et al., 2015], measuring the divergence between the output logits of both models. The objective in Eqn 91 pushes P to predict the counterfactual behavior of N when a descriptor of the intervention is given (Figure 9.1d) 1 CPMIN resembles an input augmentation approach. At inference time, approximate counterfactuals are
inaccessible. To explain model N , we append the trained token embedding tCi ←c′ to a factual input, upon which P predicts a counterfactual output for this input, used to estimate the counterfactual behavior of N under this intervention. CPMHI : Hidden-state CPM Unlike CPMIN which requires input augmentation, CPMHI models rely on representation learning via Interchange Intervention Training [Geiger et al., 2022c] by localizing concept information within existing representations. It is trained to mimic both the factual and counterfactual behavior of N . A conventional intervention on a hidden representation H of a neural network N fixes the value of the representation H to a constant. In an interchange intervention, we instead fix H to the value it would have been when processing a separate source input s. The result of the interchange intervention is a new model. Formally, we describe this new model as NH ←Hs , where ← is the conventional intervention operator and Hs is the
value of hidden representation H when processing input s. i Given a dataset of approximate counterfactual input pairs (xu,v , x̃u,v C ←c ′ ) and a black-box model N , we train a new CPMHI model P with the counterfactual objective C ←c i LHI = CES (N (x̃u,v ′ ), PH Ci ←HsCi (xu,v )) (9.2) C Here H i are hidden states designated for concept Ci . In essence, we train P to fully mediate the ′ ′ C =c C effect of intervening on Ci in the hidden representation H i . The source input s is any input xu′i,v′ ′ that has Ci = c . As P only receives information about the concept-level intervention Ci ← c via the C C interchange intervention H i ← Hs i , the model is forced to store all causally relevant information with regard to Ci in the corresponding hidden representation. This process is described in Figure 91e ′ C =c Ideally, the source input xu′i,v′ and xu,v share the same value only for Ci and differ on all others, so that the counterfactual
signal for localization is pure. However, we do not insist on this when we 1 Our objective is for a single approximate counterfactual pair for the sake of clarity. At train-time, we aggregate the objective over all considered training pairs. We take Ci to always represent the intervened-upon concept The weights of N are frozen. CHAPTER 9. CAUSAL PROXY MODELS FOR CONCEPT-BASED MODEL EXPLANATIONS138 C ←c i sample. In addition, we allow null effect pairs in which xu,v and x̃u,v ′ are identical. At inference time, approximate counterfactuals are inaccessible, as before. To explain model N ′ with regard to intervention Ci ← c , we manipulate the internal states of model P by intervening on ′ C =c C the localized representation H i for concept Ci . To achieve this, we sample a source input xu′i,v′ ′ C from the train set as any input x that has Ci = c to derive Hs i . Training Objectives We include another distillation objective to predict the same output as N
under conventional circumstances as LMimic = CES (N (xu,v ), P(xu,v )). The overall training objective for our models is thus L = λ1 LMimic + λ2 LCounterfactual where LCounterfactual can be either LIN or LHI . We set λ1 = 10 and λ2 = 30 for simplicity 9.4 Experiment Setup 9.41 Causal Estimation-Based Benchmark (CEBaB) CEBaB [Abraham et al., 2022b] is a benchmark of high-quality, labeled approximate counterfactuals for the task of sentiment analysis on restaurant reviews. The benchmark was created starting from a set of 2,299 original restaurant reviews from OpenTable. For each of these original reviews, approximate counterfactual examples were written by human annotators; the annotators were tasked to edit the original text to reflect a specific intervention, like ‘change the food evaluation from negative to positive’ or ‘change the service evaluation from positive to unknown’. In this way, the original reviews were expanded with approximate counterfactuals to a total
of 15,089 texts. The groups of originals and corresponding approximate counterfactuals are partitioned over train/dev/test. The pairs in the dev and test sets are used to benchmark explanation methods. Each text in CEBaB was labeled by five crowdworkers with a 5-star sentiment score. In addition, each text was annotated at the concept level for four mediating concepts {Cambiance , Cfood , Cnoise , and Cservice }, using the labels {negative, unknown, positive}, again with five crowdworkers annotating each concept-level label. As discussed in Section 9.3 and Figure 91b, we consider two sources of approximate counterfactuals using CEBaB. For human-created counterfactuals, we use the edited restaurant reviews of the train set. For metadata-sampled counterfactuals, we sample factual inputs from the train set that have the desired combination of mediating concepts. Using all the human-created edits leads to 19,684 training pairs of factuals and corresponding approximate counterfactuals.
Sampling counterfactuals leads to 74,574 pairs. We use these approximate counterfactuals to train explainers CHAPTER 9. CAUSAL PROXY MODELS FOR CONCEPT-BASED MODEL EXPLANATIONS139 9.42 Evaluation Metrics Much of the value of a benchmark like CEBaB derives from its support for directly calculating the ̂ N ) for a model N given a human-generated Estimated Individual Causal Concept Effect (ICaCE i approximate counterfactual pair (xu,v , x̃u,v ′ C ←c ): ′ ′ Ci ←c i ←c ̂ N (xu,v , x̃C ICaCE u,v ) = N (x̃u,v ) − N (xu,v ) (9.3) This is simply the difference between the vectors of output scores for the two examples. ′ Ci ←c We do not expect to have pairs (xu,v , x̃u,v ) at inference time, and this is what drives the development of explanation methods EN that estimate this quantity using only a factual input xu,v ′ and a description of the intervention Ci ← c . To benchmark such methods, we follow Abraham et al [2022b] in using the ICaCE-Error: D
ICaCE-ErrorN (E) = 1 ∣D∣ ∑ Dist( Ci ←c′ (xu,v ,x̃u,v )∈D ′ ′ i ←c ̂ N ((xu,v , x̃C ICaCE )), EN (xu,v ; Ci ← c )) (9.4) u,v Here, we assume that D is a dataset consisting entirely of approximate counterfactual pairs ′ Ci ←c ̂ N for the model N and the effect (xu,v , x̃u,v ). Dist measures the distance between the ICaCE predicted by the explanation method. Abraham et al [2022b] consider three values for Dist: L2, which captures both direction and magnitude; Cosine distance, which captures the direction of effects but not their magnitude; and NormDiff (absolute difference of L2 norms), which captures magnitude but not direction. We report all three metrics 9.43 Baseline Methods BESTCEBaB We compare our results with the best results obtained on the CEBaB benchmark. Crucially, BESTCEBaB is not a single method but represents the best results aggregated from a set of methods including CONEXP [Goyal et al., 2020], TCAV [Kim et al, 2018], ConceptSHAP
[Yeh et al., 2020], INLP [Ravfogel et al, 2020c], CausaLM [Feder et al, 2020], and S-Learner [Künzel et al., 2019] S-Learner Our version of S-Learner [Künzel et al., 2019] learns to mimic the factual behavior of black-box model N while making the intermediate concepts explicit. Given a factual input, a finetuned model B is trained to predict concept label for each concept as an aspect-based sentiment classification task. Then, a logistic regression model LRN is trained to map these intermediate concept values to the factual output of black-box model N , under the objective S,B LMimic = CES (N (xu,v ), LRN (B(xu,v ))) (9.5) CHAPTER 9. CAUSAL PROXY MODELS FOR CONCEPT-BASED MODEL EXPLANATIONS140 no counterfactuals BESTCEBaB S-Learner sampled counterfactuals (ours) (ours) S-Learner GPT-3 CPMIN CPMHI human-created counterfactuals (ours) (ours) S-Learner GPT-3 CPMIN CPMHI Model Metric BERT L2 Cosine NormDiff 0.74 (02) 0.59 (03) 0.44 (01) 0.74 (02) 0.63 (01) 0.54 (02) 0.74
(02) 0.63 (01) 0.53 (02) 0.71 (01) 063 (01) 060 (01) 0.51 (00) 046 (00) 045 (00) 0.35 (01) 039 (01) 038 (00) 0.73 (02) 0.60 (01) 0.52 (02) 0.45 (01) 045 (02) 045 (03) 0.36 (00) 035 (00) 036 (04) 0.25 (00) 024 (01) 027 (01) RoBERTa L2 Cosine NormDiff 0.78 (01) 0.58 (01) 0.45 (00) 0.78 (01) 0.64 (01) 0.59 (01) 0.78 (00) 0.65 (01) 0.58 (00) 0.74 (01) 066 (01) 067 (02) 0.53 (01) 046 (00) 047 (00) 0.36 (00) 042 (01) 045 (03) 0.77 (00) 0.63 (01) 0.56 (00) 0.48 (01) 046 (01) 0.39 (00) 038 (01) 0.28 (01) 026 (01) GPT-2 L2 Cosine NormDiff 0.60 (02) 0.59 (01) 0.40 (01) 0.60 (02) 0.59 (01) 0.40 (01) 0.61 (01) 0.59 (01) 0.41 (01) 0.65 (01) 0.52 (00) 0.34 (00) 0.55 (01) 051 (01) 0.47 (01) 046 (00) 0.32 (01) 030 (00) 0.61 (01) 0.59 (01) 0.40 (01) 0.43 (01) 041 (01) 041 (04) 0.40 (00) 037 (01) 039 (05) 0.24 (01) 023 (01) 027 (05) LSTM L2 Cosine NormDiff 0.73 (01) 0.64 (01) 0.50 (01) 0.73 (01) 0.64 (01) 0.53 (01) 0.73 (01) 0.64 (01) 0.53 (00) 0.76 (00) 066 (01) 064 (02) 0.57
(01) 050 (00) 050 (01) 0.41 (00) 042 (00) 041 (01) 0.72 (00) 0.63 (01) 0.54 (00) 0.49 (00) 052 (00) 0.44 (00) 045 (01) 0.30 (00) 034 (01) 0.47 (03) 0.39 (03) 0.29 (05) 0.54 (01) 0.46 (00) 0.36 (00) Table 9.1: CEBaB scores measured in three different metrics on the test set for four different model architectures as a five-class sentiment classification task. Lower is better Results averaged over three distinct seeds, standard deviations in parentheses. The metrics are described in Section 94 Best averaged result is bolded (including ties) per approximate counterfactual creation strategy. By intervening on the intermediate predicted concept values at inference-time, we can hope to simulate the counterfactual behavior of N : S,B EN (xu,v ; Ci ← c ) = ′ LRN ((B(xu,v ))Ci ←c′ ) − LRN (B(xu,v )) (9.6) When using S-Learner in conjunction with approximate counterfactual inputs at train-time, we simply add this counterfactual data on top of the observational data that is
typically used to train S-Learner. GPT-3 Large language models such as GPT-3 (175B) have shown extraordinary power in terms of in-context learning [Brown et al., 2020] We use GPT-3 (davinci-002) to generate a new approximate counterfactual at inference time given a factual input and a descriptor of the intervention. This generated counterfactual is directly used to estimate the change in model behavior: GPT-3 EN (xu,v ; Ci ← c ) = ′ N (GPT-3(xu,v ; Ci ← c )) − N (xu,v ) (9.7) ′ where GPT-3(xu,v ; Ci ← c ) represents the GPT-3 generated counterfactual edits. We prompt GPT-3 ′ with demonstrations containing approximate counterfactual inputs. 9.44 Causal Proxy Models We train CPMs for the publicly available models released for CEBaB, fine-tuned as five-way sentiment classifiers on the factual data. This includes four model architectures: bert-base-uncased (BERT; CHAPTER 9. CAUSAL PROXY MODELS FOR CONCEPT-BASED MODEL EXPLANATIONS141 Devlin et al. 2019c),
RoBERTa-base (RoBERTa; Liu et al 2019b), GPT-2 (GPT-2; Radford et al 2019), and LSTM+GloVe (LSTM; Hochreiter and Schmidhuber 1997, Pennington et al. 2014) All Transformerbased models [Vaswani et al, 2017b] have 12 Transformer layers Before training, each CPM model is initialized with the architecture and weights of the black-box model we aim to explain. Thus, the CPMs are rooted in the factual behavior of N from the start. The inference time comparisons for these models are as follows, where P in (9.8) and (99) refers to the CPM model trained under CPMIN and CPMHI objectives, respectively: CPMIN EN (xu,v ; Ci ← c ) = P(xu,v ; tCi ←c′ ) − N (xu,v ) ′ CPM ′ EN HI (xu,v ; Ci ← c ) = PH Ci ←HsCi (xu,v ) − N (xu,v ) ′ (9.8) (9.9) C Here, s is a source input with Ci = c , and H i is the neural representation associated with Ci C C which takes value Hs i on the source input s. As H i , we use the representation of the [CLS] token Specifically, for BERT we use
slices of width 192 taken from the 1st intermediate token of the 10th layer. For RoBERTa, we use the 8th layer instead For GPT-2, we pick the final token of the 12th layer, again with slice width of 192. For LSTM, we consider slices of the attention-gated sentence embedding with width 64. Following the guidance on IIT given by Geiger et al. [2022c], we train CPMHI with an additional multi-task objective as LMulti = ∑Ci ∈C CE(MLP(Hx i ), c) where probe is parameterized by a mulC C tilayer perceptron MLP, and Hx i is the value of hidden representation for the concept Ci when processing input x with a concept label of c for Ci . 9.5 Results Model Blackbox sampled counterfactuals CPMIN CPMHI human-created counterfactuals CPMIN CPMHI BERT RoBERTa GPT-2 LSTM 0.70 (01) 0.70 (00) 0.65 (00) 0.60 (01) 0.70 (00) 0.70 (00) 0.65 (00) 0.60 (01) 0.70 (01) 0.71 (01) 0.66 (01) 0.54 (00) 0.67 (02) 0.69 (01) 0.67 (01) 0.56 (00) 0.69 (01) 0.71 (00) 0.68 (00) 0.59 (01) Table 9.2: Task
performance measured as Macro-F1 score on the test set (average of 3 distinct seeds; standard deviations in parentheses). We first benchmark both CPM variants and our baseline methods on CEBaB. We show that the CPMs achieve state-of-the-art performance, for both types of approximate counterfactuals used during CHAPTER 9. CAUSAL PROXY MODELS FOR CONCEPT-BASED MODEL EXPLANATIONS142 Model Metric sampled counterfactuals CPMIN CPMHI human-created counterfactuals CPMIN CPMHI BERT L2 Cosine NormDiff 0.63 (01) 052 (04) 0.46 (00) 045 (01) 0.39 (01) 033 (02) 0.42 (02) 0.34 (02) 0.23 (01) 0.38 (03) 0.30 (06) 0.22 (05) L2 RoBERTa Cosine NormDiff 0.66 (01) 063 (04) 0.46 (00) 048 (01) 0.42 (01) 042 (05) 0.40 (01) 0.33 (01) 0.21 (01) 0.37 (04) 0.29 (04) 0.23 (05) GPT-2 L2 Cosine NormDiff 0.55 (01) 041 (03) 0.47 (01) 039 (02) 0.32 (01) 025 (02) 0.38 (01) 0.37 (01) 0.22 (01) 0.36 (04) 0.35 (05) 0.24 (05) LSTM L2 Cosine NormDiff 0.66 (01) 041 (01) 0.50 (00) 042 (02) 0.42 (00) 025
(00) 0.46 (00) 0.50 (02) 0.31 (00) 0.42 (01) 0.40 (01) 0.28 (02) Table 9.3: Self-explanation CEBaB scores measured in three different metrics on the test set for four different model architectures as a five-class sentiment classification task. Lower is better Average of 3 distinct seeds; standard deviations in parentheses. training (Section 9.51) Given the good factual performance achieved by CPMs, we subsequently investigate whether CPMs can be deployed both as predictor and explanation method at the same time (Section 9.52) and find that they can Finally, we show that the localized representations of CPMHI give rise to concept-aware feature attributions (Section 9.53) Our supplementary materials report on detailed ablation studies and explore the potential of our methods for model debiasing. 9.51 CEBaB Performance Table 9.1 presents our main results The results are grouped per approximate counterfactual type used during training. Both CPMIN and CPMHI beat BESTCEBaB in every
evaluation setting by a large margin, establishing state-of-the-art explanation performance. Interestingly, CPMHI seems to slightly outperform CPMIN using sampled approximate counterfactuals, while slightly underperforming CPMIN on human-created approximate counterfactuals. S-Learner, one of the best individual explainers from the original CEBaB paper [Abraham CHAPTER 9. CAUSAL PROXY MODELS FOR CONCEPT-BASED MODEL EXPLANATIONS143 Model Black-box Predicted Concept neutral Multi-task neutral neutral IIT Score Word Importance ambiance +0.03 [CLS] the music was too loud , and the decorations were taste ##less , but they had friendly waiter ##s and delicious pasta [SEP] food +0.11 [CLS] the music was too loud , and the decorations were taste ##less , but they had friendly waiter ##s and delicious pasta [SEP] noise +0.04 [CLS] the music was too loud , and the decorations were taste ##less , but they had friendly waiter ##s and delicious pasta [SEP] service +0.26 [CLS]
the music was too loud , and the decorations were taste ##less , but they had friendly waiter ##s and delicious pasta [SEP] ambiance +0.25 [CLS] the music was too loud , and the decorations were taste ##less , but they had friendly waiter ##s and delicious pasta [SEP] food +0.23 [CLS] the music was too loud , and the decorations were taste ##less , but they had friendly waiter ##s and delicious pasta [SEP] noise +0.31 [CLS] the music was too loud , and the decorations were taste ##less , but they had friendly waiter ##s and delicious pasta [SEP] service +0.16 [CLS] the music was too loud , and the decorations were taste ##less , but they had friendly waiter ##s and delicious pasta [SEP] ambiance −0.24 [CLS] the music was too loud , and the decorations were taste ##less , but they had friendly waiter ##s and delicious pasta [SEP] food +1.11 [CLS] the music was too loud , and the decorations were taste ##less , but they had friendly waiter ##s and delicious pasta [SEP]
noise −0.98 [CLS] the music was too loud , and the decorations were taste ##less , but they had friendly waiter ##s and delicious pasta [SEP] service +1.16 [CLS] the music was too loud , and the decorations were taste ##less , but they had friendly waiter ##s and delicious pasta [SEP] Table 9.4: Visualizations of word importance scores using Integrated Gradient (IG) by restricting gradient flow through the corresponding intervention site of the targeted concept. Our target class pools positive and very positive. Individual word importance is the sum of neuron-level importance scores for each input, normalized to [ −1 , +1 ]. −1 means the word contributes the most negatively to predicting the target class (red); +1 means the word contributes the most positively (green). et al., 2022b], shows only a marginal improvement when naively incorporating sampled and humancreated counterfactuals during training over using no counterfactuals This indicates that the large performance
gains achieved by our CPMs over previous explainers are most likely due to the explicit use of a counterfactual training signal, and not primarily due to the addition of extra (counterfactual) data. GPT-3 occasionally performs on-par with our CPMs, generally only slightly underperforming our best explainer on human-created counterfactuals, while being significantly worse on sampled counterfactuals. While the GPT-3 explainer also explicitly uses approximate counterfactual data, the results indicate that our proposed counterfactual mimic objectives give better results. The better performance of CPMs when considering sampled counterfactuals over GPT-3 shows that our approach is more robust to the quality of the approximate counterfactuals used. While the GPT-3 explainer is easy to set up (no training required), it might not be suitable for some explanation applications regardless of performance, due to the latency and cost involved in querying the GPT-3 API. Across the board, explainers
trained with human-created counterfactuals are better than those trained with sampled counterfactuals. This shows that the performance of explanation methods depends on the quality of the approximate counterfactual training data. While human counterfactuals give excellent performance, they may be expensive to create. Sampled counterfactuals are cheaper if the relevant metadata is available. Thus, under budgetary constraints, sampled counterfactuals may be more efficient. Finally, CPMIN is conceptually the simpler of the two CPM variants. However, we discuss in Section 9.53 how the localized representations of CPMHI lead to additional explainability benefits CHAPTER 9. CAUSAL PROXY MODELS FOR CONCEPT-BASED MODEL EXPLANATIONS144 9.52 Self-Explanation with CPM As outlined in Section 9.3, CPMs learn to mimic both the factual and counterfactual behavior of the black-box models they are explaining. We show in Table 92 that our CPMs achieve a factual Macro-F1 score comparable to the
black-box finetuned models. Can we simply replace the black-box model with our CPM and use the CPM both as factual predictor and counterfactual explainer. To answer this question, we measure the self-explanation performance of CPMs by replacing N in Eqn. 94 with our factual CPM predictions at inference time. Table 9.3 reports these results We find that both CPMIN and CPMHI achieve better selfexplanation performance compared to providing explanations for another black-box model Furthermore, CPMHI provides better self-explanation than CPMIN , suggesting our interchange intervention procedure leads the model to localize concept-based information in hidden representations. This shows that CPMs may be viable as replacements for their black-box counterpart, since they provide similar task performance while providing faithful counterfactual explanations of both the black-box model and themselves. 9.53 Concept-Aware Feature Attribution with Causal Proxy Models CPMHI localizes concept-based
information within representations. We have shown that CPMHI provides trustworthy explanations (Section 9.51) We now investigate whether CPMHI learns representations that mediate the effects of different concepts. We adapt Integrated Gradients (IG; Sundararajan et al. 2017a) to provide concept-aware feature attributions, by only considering gradients flowing through the hidden representation associated with a given concept. In Table 9.4, we compare concept-aware feature attibutions for two variants of CPMHI (IIT and Multi-task) and the original black-box (Finetuned) model. For IIT we remove the multi-task objective LMulti during training and for Multi-task we remove the the IIT objective LHI . This helps isolate the individual effects of both losses on concept localization. All three models predict a neutral final sentiment score for the considered input, but they show vastly different feature attributions. Only IIT reliably highlights words that are semantically related to each
concept. For instance, when we restrict the gradients to flow only through the intervention site of the noise concept, “loud” is the word highlighted the most that contributes negatively. When we consider the service concept, words like “friendly” and “waiter” are highlighted the most as contributing positively. These contrasts are missing for representations of the Multi-task and Finetuned models. Only the IIT training paradigm pushes the model to learn causally localized representations. For the service concept, we notice that the IIT model wrongfully attributes “delicious”. This could be useful for debugging purposes and could be used to highlight potential failure modes of the model. CHAPTER 9. CAUSAL PROXY MODELS FOR CONCEPT-BASED MODEL EXPLANATIONS145 9.6 Conclusion We explored the use of approximate counterfactual training data to build more robust causal explanation methods. We introduced Causal Proxy Models (CPMs), which learn to mimic both the factual and
counterfactual behaviors of a black-box model N . Using CEBaB, a benchmark for causal concept-based explanation methods, we demonstrated that both versions of our technique (CPMIN and CPMHI ) significantly outperform previous explanation methods. CPMs require only very partial causal models and highly approximate counterfactuals to be achieve these state-of-the-art results. Our results suggest that CPMs can be more than just explanation methods. They achieve factual performance on par with the model they aim to explain, and they can explain their own behavior. This paves the way to using them as deployed models that both perform tasks and offer explanations. In addition, the causally localized representations of our CPMHI variant are very intuitive, as revealed by our concept-aware feature attribution technique. We believe that causal localization techniques could play a vital role in further model explanation efforts. Chapter 10 When Interpretability Enhances Accessibility:
Updating CLIP to Prefer Descriptions Over Captions Abstract CLIPScore is a powerful generic metric that captures the similarity between a text and image. However, CLIPScore makes no reference to human-generated, ground-truth texts, and thus fails to distinguish between a caption that is meant to complement the information in an image and a description that is meant to replace an image entirely, e.g, for the purpose of accessibility We address this shortcoming by fine-tuning the CLIP model on the Concadia dataset to assign higher scores to descriptions than captions. We employ a novel variant of interchange intervention training (IIT) with a distributed alignment search (DAS) to induce a causal variable of the description–caption distinction in CLIP while simultaneously learning where to induce the variable. While IIT without DAS and standard fine-tuning can also produce CLIP models that reliably assign descriptions a higher score, IIT-DAS preserves more of the original capabilities
of CLIP, produces scores that better correlate with usefulness judgements from blind and low vision people, and proves more interpretable when analyzed with integrated gradients. 10.1 Introduction The texts that accompany images online are written with a variety of distinct purposes: to add commentary, to identify entities, to enable search, and others. One of the most important purposes is (alt-text) description to help make the image non-visually accessible, which is especially important for people who are blind or low vision (BLV). The ability to automatically evaluate high-quality descriptions of images would mark a significant step towards making the Web accessible for everyone. 146 CHAPTER 10. WHEN INTERPRETABILITY ENHANCES ACCESSIBILITY 147 Figure 10.1: A visual depiction of a training update during distributed interchange intervention training. The core idea is that a description of an image in Concadia should be assigned a higher CLIPScore than the caption of the same
image, so we induce a representation of the description– caption distinction that increases the CLIPScore when the text is a description. The final label (in the diagram, < or >) is determined by whether the intervention is caption description or vice versa. Unfortunately, present-day metrics for image-text similarity tend to be insensitive to purpose, and thus they fall short when it comes to helping with accessibility [Kreiss et al., 2022b] This problem is especially acute when it comes to how we evaluate models; if our tools for evaluation are not sensitive to purpose, we have little hope of making genuine progress [Kreiss et al., 2022a] The Contrastive Language-Image Pre-training (CLIP) model of Radford et al. [2021] is an important case in point. CLIP is trained to embed images and texts with the objective of maximizing the similarity of related image-text pairs and minimizing similarity of unrelated pairs. We can quantify the similarity between image-text pairs using the
CLIPScore metric, which is based on the cosine similarity of CLIP encodings for the image and text Hessel et al. [2021] Crucially, CLIPScore is a referenceless metric, meaning there is no reliance on human-generated, ground-truth texts. This renders CLIPScore insensitive to the purpose of the text In turn, Kreiss et al. [2022a] find that CLIPScore correlates with neither sighted nor BLV user evaluations of image alt-descriptions. Though it achieves high performance on many image-text classification tasks, CLIP is unsuitable for alt-text evaluation. However, in this paper, we show that there is a path forward for making CLIPScores sensitive to purpose in the requisite ways. As we are focused on accessibility, our goal is to update CLIP to assign higher scores to descriptions than caption texts (which are meant to complement an image rather than replace it). To do this, we use the Concadia dataset [Kreiss et al, 2022b] to fine-tune CHAPTER 10. WHEN INTERPRETABILITY ENHANCES
ACCESSIBILITY 148 CLIP. Concadia consists of 96,918 images with corresponding descriptions, captions, and surrounding context. For our purposes, each description and caption of an image in Concadia can approximate the counterfactual: What if the text for this image were a description instead of a caption (or vice versa)? The underlying image is the same, but we can intervene upon the communicative purpose. These approximate counterfactuals enable us to use interchange intervention training (IIT; Geiger et al. 2022d) to update CLIP such that it computes a representation of the description–caption distinction that increases CLIPScore when the text is a description. Previous works use IIT to fine-tune models by hand-picking a neural representation that should encode a causal variable. However, CLIP is a pre-trained model with capabilities that need to remain intact after further fine-tuning. As such, we employ a novel variant of IIT with a distributed alignment search (DAS; Geiger
et al. 2023e) that induces a causal variable of the description–caption distinction in CLIP while simultaneously searching for the best linear subspace in a neural representation to store the variable. We compare CLIP fine-tuned with IIT-DAS, CIIT-DAS , to CLIP fine-tuned with IIT without DAS, CIIT , and CLIP fine-tuned with a standard behavioral contrastive objective, CBehave . All three of these models strongly prefer descriptions over captions. However, our transfer learning evaluation with five image classification datasets reveals that while all models have degraded zero-shot capabilities, CIIT-DAS preserves more capabilities than CIIT which in turn preserves more than CBehave . Furthermore, we find that the CLIPScore fine-tuned with CIIT-DAS has a correlation of 0.23 with BLV human judgments on the usefulness of a text, which is significantly higher than the original CLIP model (0.08), CIIT (009), and CBehave (014) Finally, we find that the integrated gradient attribution
values assigned to words in the models CIIT and CIIT-DAS better correlate with human judgements of concreteness and imageability compared to the original CLIP and CBehave . Thus, overall, our novel use of interchange intervention training with a distributed alignment search emerges as a promising method for updating pre-trained deep learning models with an interpretable causal variable. 10.2 Related Work Image Accessibility When images can’t be seen, visual descriptions of those images make them accessible. For images online, these descriptions can be provided in the HTML’s alt tag, which are then visually displayed if the image cannot be loaded or they are read out by a screen reader to, for instance, users who are blind or low-vision (BLV). However, alt descriptions online remain rare with only about 0.1% of English-language Twitter images Gleason et al [2019] and 6% of English-language Wikipedia images Kreiss et al. [2022b] having any human-written description associated with
them Image captioning models provide an opportunity to generate such accessibility descriptions at scale, which would promote equal access Gleason et al. [2020] But the resulting models have remained CHAPTER 10. WHEN INTERPRETABILITY ENHANCES ACCESSIBILITY 149 largely unsuccessful in practice Morris et al. [2016], MacLeod et al [2017], Gleason et al [2019] Kreiss et al. [2022b] argue that this is partly due to the general approach of treating all image-based text generation problems as the same underlying task, and instead highlight the need for a distinction between accessibility descriptions and contextualizing captions. Descriptions are needed to replace images, while the purpose of a caption is to provide supplemental information which is usually accessible to everyone below the image. In support of the distinction, Kreiss et al [2022b] find that the language used in descriptions and captions categorically differs and that sighted participants tend to learn more from captions
but can visualize the image better from descriptions. Referenceless Text-Image Evaluation Metrics Our work focuses on referenceless evaluation metrics for text-image models that can be applied in many contexts without requiring human annotations, which can be time-consuming. Some referenceless metrics, such as SPURTS Feinglass and Yang [2021], evaluate text quality based on text-internal properties alone. Others, including CLIPScore Hessel et al. [2021] and UMIC Lee et al [2021], utilize models trained with a contrastive learning objective in order to produce a text-image similarity score. Crucially, such referenceless metrics are context-free, and hence can be applied in a variety of multimodal settings (e.g image synthesis, description generation, zero-shot image classification). Yet as Kreiss et al [2022a] point out, referenceless metrics might be insensitive to important high-level concepts. In fact, neither CLIPScore nor SPURTS are sensitive to the purpose of a text, and do not
distinguish between descriptions and captions of the same image. Model Editing and Interpretability In model editing, an existing model is modified to acquire new properties while preserving others. This might be updating model weights to teach new knowledge De Cao et al. [2021a], Meng et al [2022b, 2023], Mitchell et al [2022], simulating counterfactuals for bias detection Dash et al. [2022], or removing information from a neural representation to erase a concept Ravfogel et al. [2020b], Elazar et al [2021a], Belrose et al [2023], Gandikota et al [2023] Interpretability is concerned with the discovery of systematic structure in model behavior and/or internals Vig et al. [2020a], Olah et al [2020], Geiger et al [2020a, 2021b] Model editing and interpretability have natural interplay; in order to manipulate and alter an artifact, one has learn how it works. For example, Materzynska et al [2022] disentangle the spelling component of CLIP from its visual component via a low-dimensional
projection of its learned representations, and ablating this subspace leads to a model that forgets how to spell. For this reason, it seems natural that interchange intervention training (IIT; Geiger et al. 2022d), a method for inducing interpretable causal structure in deep learning models, could be used as a model editing tool. Indeed, we use IIT to update the CLIP model to have a representation of the description–caption distinction. However, instead of hand picking a set of hidden activations to represent this concept, we use IIT with a Distributed Alignment Search [Geiger et al., 2023e] to learn a linear subspace of a hidden activation vector in the CLIP text encoder that can best represent this CHAPTER 10. WHEN INTERPRETABILITY ENHANCES ACCESSIBILITY 150 distinction. 10.3 Methods We fine-tune CLIP with IIT-DAS in order to induce a causal structure which makes the critical distinction between descriptions and captions. For each caption–description–image triple in
Concadia, we do the following (as illustrated in Figure 10.1) 1. Compute the image encoding, eImage 2. Compute the encoding of a description, eDes 3. Compute a counterfactual encoding of the description where a (distributed) neural representation in the text module of CLIP is fixed to be the value of the neural representation for the caption of the same image, eDes←Cap 4. Train CLIP and the distributed alignment on the contrastive objectives, eImage ⋅ eDes > eImage ⋅ eDes←Cap eImage ⋅ eCap < eImage ⋅ eCap←Des 10.31 CLIP CLIP is a pre-trained multi-modal model that consists of an image encoder and a text encoder (both of which use a Transformer architecture; Vaswani et al. 2017b), each followed by a projection layer into a shared image-text subspace. CLIP is trained with a contrastive learning objective: the encodings of a corresponding image-text pair should have high cosine similarity in the shared subspace, whereas the encodings of a non-matching image-text
pair should have low cosine similarity. CLIP is trained to be able to choose, from a set of images, the image that best fits a particular text, and choose out of a set of texts the text that best fits a particular image (our context). CLIPScore is a metric that uses CLIP to quantify the similarity between a text and image by computing the cosine similarity and then scaling to fit the interval [0,1]. 10.32 Causal Models and Interventions This section follows Geiger et al. 2023b,e Causal Models can represent a variety of processes, including deep learning models and symbolic algorithms. A causal model M consist of variables V, and, for each variable X ∈ V, a set of values Val(X), and a structural equation FX ∶ Val(V) Val(X), which is a function that takes in a setting of all the variables and outputs a value for X. The solutions of a model M = (V, Val, F ) are CHAPTER 10. WHEN INTERPRETABILITY ENHANCES ACCESSIBILITY 151 settings for all variables v ∈ Val(V) such that the
output of the causal mechanism FX (v) is the same value that v assigns to X, for each X ∈ V. We only consider structural causal models with a single solution that induces a directed acyclic graphical structure such that the value for a variable X depends only on the set of variables that point to it, denoted as its parents PAX . Because of this, we treat each causal mechanism FX as a function from parent values in Val(PAX ) to a value in Val(X). We denote the set of variables with no parents as Vin and those with no children Vout . Given input ∈ Val(Vin ) and variables X ⊆ V, we define GET(M, input, X) ∈ Val(X) to be the setting of X determined by the given input and model M. For example, X could correspond to a hidden activation layer in a neural network, and GET(M, input, X) then denotes the particular values that X takes on when the model M processes input. Interventions simulate counterfactual states in causal models. For a set of variables X and a setting for those
variables x ∈ VAL(X), we define MX←x to be the causal model identical to M, except that the structural equations for X are set to constant values x. In the case of neural networks, we overwrite the activations with x in-place so that gradients can back-propagate through x. Distributed interventions also simulate counterfactual states in causal models, but do so by editing the causal mechanisms rather than overwriting them to be a constant. Given variables X and an invertible function ρ ∶ VAL(X) VAL(Y) mapping X into a new variable space Y, define ρ(M) to be the model where the variables X are replaced with the variables Y. For a setting of the new variable space y ∈ VAL(Y), it follows that ρ (ρ(M)Y←y ) is the causal model identical to M, −1 except that the causal mechanisms for X are edited to fix the value of Y to y. If ρ is differentiable, then gradients back-propagate through y. A distributed interchange intervention fixes variables to the values they would have
taken if a different input were provided. Consider a causal model M, an invertible function ρ ∶ VAL(X) VAL(Y), source and base inputs s, b ∈ Val(Vin ), and a set of intermediate variables X ⊂ V. A distributed interchange intervention computes the value Vout when run on b, intervening on the (distributed) intermediate variables Y to be the value they take on when run on s. Formally, it is defined as DII(M, ρ, b, s, Y) = GET(ρ (ρ(M)Y←GET(ρ(M),s,Y) ), b, Vout ) −1 10.33 Contrastive Training Objectives Concadia consists of 96,918 images with corresponding descriptions, captions and surrounding context. In our work, we consider descriptions and captions (corresponding to the same image) to be a counterfactual pair that answers the following what-if question: What text would a human CHAPTER 10. WHEN INTERPRETABILITY ENHANCES ACCESSIBILITY 152 Concadia Fine-Tuning Food101 (101) ImageNet (¿20,000) SST2 (2) Country211 (211) CIFAR100 (100) None Behavioral IIT
IIT-DAS 76.4% 48.7% ± 160 63.9% ± 900 66.2% ± 377 53.6% 31.6% ± 146 35.3% ± 455 39.7% ± 192 49.6% 53.7% ± 116 55.9% ± 075 56.9% ± 289 14.3% 6.2% ± 069 10.4% ± 172 9.5% ± 097 61.5% 53.1% ± 180 54.6% ± 122 55.4% ± 242 Table 10.1: Transfer learning results for CLIP models fine-tuned on Concadia The error bounds are 95% confidence intervals from runs with 5 random seeds. The number of classes for each dataset is stated in parentheses. assign to this image if they were aiming for a description versus a caption? Rather than fine-tuning CLIP to implement one specific causal model, we will fine-tune it to implement a causal model that discriminates between descriptions and captions. In particular, define a class ∆ to contain only causal models that consist of the following three variables. The input variable X takes on the value of some image-text pair in Concadia (ximage , xtext ); the intermediate variable P (i.e purpose) takes on a value from {“describe”,
“caption”} depending on the Concadia label for X; and the output variable Y takes on a real value that represents the similarity between ximage and xtext . The causal mechanism of Y must be such that for every image, the description text is assigned a higher CLIPScore than the caption text. If a CLIP model implements any algorithm in ∆, then it will assign descriptions higher scores than captions. We approach the problem of fine-tuning CLIP to distinguish between descriptions and captions similarly to CLIP’s pre-training. That is, we define unsupervised, contrastive learning objectives Behavioral Objective Our class of causal models ∆ does not provide a deterministic function θ that determines the effect of P and X on Y . Nevertheless, we can incentivize a CLIP model C to implement some A ∈ ∆ with a contrastive learning objective. For any given image with a caption and description, we run C on the image-caption pair (ximage , xcap ), run C on the image-description pair
θ (ximage , xdes ), and fine-tune C to produce a higher score for the latter by minimizing the objective: LBehave = θ θ CE([C (ximage , xcap ), C (ximage , xdes )], [0, 1]) θ Interchange Intervention Objective We can incentivize a CLIP model C to implement some A ∈ ∆ with a contrastive IIT objective for each triplet x = (ximage , xcap , xdes ) in Concadia, LIIT = θ θ θ CE([C (b), DII(C , ρ , b, s, Z)], [l0 , l1 ]) CHAPTER 10. WHEN INTERPRETABILITY ENHANCES ACCESSIBILITY Concadia Fine-Tuning Desc. ¿ Capt None Behavioral IIT IIT-DAS 49.4% 87.7% ± 024 86.8% ± 109 86.9% ± 140 153 Table 10.2: The results of fine-tuning the CLIP model on the Concadia dataset The error bounds are 95% confidence intervals from runs with five random seeds. θ where ρ ∶ Val(N) Val(Y) is a randomly initialized orthogonal matrix that transforms a hidden vector activation N to its representation in an alternate basis Y, Z is the subspace of Y targeted for intervention, the base
input b is either (ximage , xdes ) or (ximage , xcap ), and the source input s is the other. If b contains a description and s contains a caption, then the labels are [l0 , l1 ] = [0, 1]; otherwise, the labels are [l0 , l1 ] = [1, 0]. 10.4 Fine-Tuning CLIP on Concadia Descriptions play a crucial role in making images on the internet accessible to BLV individuals. We can use Concadia to quantify the extent to which an image-text model is sensitive to the description and caption distinction. The metric is simply the proportion of images in Concadia where the description is assigned a higher score than the caption. In the pre-trained CLIP model, this is the case for 49.4% of the triplets in Concadia (see Table 102) This means the model assigns higher scores to descriptions at chance, and it is not well suited for the purposes of accessibility. To make the CLIP model more useful, we use Concadia to fine-tune three variants: Behavioral IIT CLIP fine-tuned on the Concadia dataset with
the loss LBehave . CLIP fine-tuned on the Concadia dataset with the loss LIIT when the orthogonal matrix ρ is fixed to be the identity matrix. This is equivalent to standard IIT IIT-DAS CLIP fine-tuned on the Concadia dataset with the loss LIIT . The orthogonal matrix ρ is updated throughout training. While each of these three models consistently assigns higher values to descriptions instead of captions (see Table 10.2), we need to ensure that the original model capabilities remain intact and that human judgements from BLV individuals correlate with the CLIPScores from the fine-tuned models. CHAPTER 10. WHEN INTERPRETABILITY ENHANCES ACCESSIBILITY Training Overall Imaginability Relevance Irrelevance BLV None Behavioral IIT IIT-DAS 0.08 0.14 0.09 ∗ 0.23 0.10 0.11 0.18 ∗ 0.33 0.09 0.13 0.18 ∗ 0.24 0.09 −0.05 0.14 0.03 Sighted, no image None Behavioral IIT IIT-DAS −0.01 0.14 0.19 0.12 0.06 0.09 0.21 ∗ 0.21 0.00 0.14 0.16 0.02 −0.17 −0.06 −0.11
−0.21 Sighted, with image None Behavioral IIT IIT-DAS 0.14 0.17 ∗ 0.29 ∗ 0.22 0.11 0.12 ∗ 0.25 0.07 −0.08 −0.01 0.02 −0.08 154 Table 10.3: Correlation between model similarity scores and human preference, across imaginability, relevance, irrelevance, and overall value of text description Kreiss et al. [2022a] Scores reported with ∗ an asterisk ( ) are statistically significant with p < 0.10 10.5 Transfer Learning Evaluations We first seek to evaluate to what extent each fine-tuning method preserves the original transfer capabilities of CLIP. The ideal method will maintain high performance on the image classification transfer tasks, while also assigning higher scores to descriptions than to captions. We evaluate our fine-tuned models on five classification tasks in disparate domains. Food101 Food images labeled from 101 categories Bossard et al. [2014] ImageNet Image-text pairs for each synonym set in the WordNet hierarchy Russakovsky et al. [2015],
Miller [1995]. Rendered SST-2 A visual sentiment classification task in which the text from the Stanford Sentiment Treebank Socher et al. [2013] is rendered into an image Radford et al [2021] Country-211 A geolocation classification dataset filtered from YFCC100m with locations that have a matching ISO country code Thomee et al. [2016], Radford et al [2021] CIFAR-100 Images from 100 different categories Krizhevsky et al. [2009] Pre-trained CLIP varies greatly in its ability to generalize to each of these tasks, but it does outperform a supervised linear classifier trained on ResNet-50 features Radford et al. [2021] We report the macro-averaged F1 score on zero-shot classification for each of the transfer tasks listed above, averaged across five randomly seeded training runs. We prefix “An image of ” to each label in order to improve zero-shot generalization, except Country-211 (“The country of ”) and Rendered SST-2 (“A sentence with sentiment”). CHAPTER 10. WHEN
INTERPRETABILITY ENHANCES ACCESSIBILITY Results and Discussion 155 Table 10.1 shows the zero-shot performance of each CLIP model on the transfer task datasets. In all cases, the CLIP model fine-tuned with IIT-DAS is as good or better than IIT fine-tuning which is, in turn, as good or better than the model with standard behavioral fine-tuning. This is a clear case for IIT-DAS being the best method to preserve CLIP capabilities after Concadia fine-tuning. We expected that fine-tuning CLIP to distinguish between captions and descriptions would harm zero-shot transfer to image classification tasks. While this was true for all other datasets, the performance of CLIP on SST-2 went up after Concadia fine-tuning. 10.6 BLV and Sighted Human Evaluations We have successfully updated CLIP to prefer descriptions over captions. Intuitively, a preference of descriptive over captioning information should align assigned CLIPScore ratings better with BLV and sighted description quality judgments.
We can think of this as our most important transfer learning task, where the metric is the correlation between CLIPScore and human judgements. To evaluate our models, we use data from the human evaluations of image descriptions collected by Kreiss et al. [2022a] Dataset Kreiss et al. [2022a] conducted an experiment where sighted and BLV participants were asked to judge a description of an image in the context of an article. For our purposes, we abstract away the context, so that we can isolate the benefit of inducing purpose into the original CLIP model. Participants rated the quality of the image descriptions along four dimensions: the overall quality of the description for accessibility, the imaginability of the image just based on the text, and the degree of relevant and irrelevant details in the description. Results For each model evaluated, we report correlations for the model with the highest description– caption distinction across five runs. Table 103 shows the correlation
between human evaluations and model similarity scores. Across participant groups and quality criteria, IIT-DAS provides consistent correlation improvements compared to the base CLIP model. When it comes to the alignment with BLV participants, IIT-DAS does especially well leading to a significant correlation with the three dimensions overall, imaginability, and relevance. When it comes to the type of quality measure, IIT-DAS seems to induce an especially strong correlation with ratings of imaginability, such that descriptions that allow users to imagine the picture better are assigned a higher score. The only case where IIT-DAS doesn’t significantly improve over base CLIP for the overall rating is for ratings sighted participants provided when they couldn’t see the image. Kreiss et al [2022a] note the ratings sighted participants gave after seeing the image were consistently more correlated with BLV ratings, suggesting these scores are less informative for the goal of accessibility.
IIT also CHAPTER 10. WHEN INTERPRETABILITY ENHANCES ACCESSIBILITY 156 provides promising gains, and seems especially beneficial for approximating sighted user judgments when they have access to the image. The behavioral objective has only low correlation with the human ratings compared to CLIP with no fine-tuning. Even though the description–caption distinction was successfully learned, it doesn’t generalize well to an accessibility downstream goal. None of the fine-tuning methods induce a sensitivity to detecting irrelevant information, which is to be expected since it doesn’t easily relate to a description–caption distinction and will require the textual context of the image Kreiss et al. [2022a] Discussion Our results show clearly that fine-tuning CLIP on the Concadia dataset to induce a representation of the description–caption distinction results in a CLIPScore that is better aligned with the judgments of sighted and BLV individuals. IIT-DAS shows the most
promising gains, especially for approximating BLV user preferences. 10.7 Integrated Gradients Through the Description–Caption Representation So far, we have fine-tuned CLIP on Concadia to learn the description–caption distinction, and shown that IIT-DAS is the most promising method for aligning CLIPScore with sighted and BLV description preferences while preserving CLIP’s original capabilities. Yet, IIT was developed as a method for creating interpretable deep learning models, so we should be able to leverage the more transparent internal structure resulting from IIT. We conduct an analysis of how the CLIP models distinguish between descriptions and captions using an attribution method called integrated gradients Sundararajan et al. [2017b] that evaluates the contribution of each text token to the CLIPScore. Integrated Gradients Integrated gradients compute the magnitude and direction along which each input feature influences the final model output. These attributions are
computed by taking the gradient of the model output with respect to the input features (where a large positive gradient means that the input feature had a large positive effect on the model output). The gradients are integrated along the differences between the input and a blank baseline (in our case, a sentence consisting only of eos tokens) to ensure that features which have a non-zero effect on the model’s output receive a non-zero attribution Sundararajan et al. [2017b] Integrated gradients are useful for identifying how tokens impact the CLIPScore. However, we are particularly curious about how tokens impact the representation of the description–caption distinction. To answer this question, we mediate the gradient computation through the linear subspace learned by DAS Wu et al. [2023a] We hypothesize that, since the intervention site is trained to represent CHAPTER 10. WHEN INTERPRETABILITY ENHANCES ACCESSIBILITY Example Description Caption 157 Mediation Score
Attribution ✗ +15.2 a black and white photograph of a man playing an electric guitar . ✓ +12.5 a black and white photograph of a man playing an electric guitar . ✗ -7.46 jimi hendrix may 1 0 , 1 9 6 8 ✓ -9.03 jimi hendrix may 1 0 , 1 9 6 8 Figure 10.2: Integrated gradient attributions for the IIT-DAS model run on a sample image and its corresponding description and caption from the Concadia dataset. A positive token attribution means that the token contributed positively to the outputted CLIPScore (green), and negative token attribution means that it contributed negatively (red). The overall score is the sum of the token attributions within the sentence. the underlying purpose of a text (i.e description or caption), mediated gradient attributions would pick out tokens that highlight the description–caption distinction. Figure 10.2 shows an example image from Concadia and the integrated gradient attributions of its corresponding description and caption on the
IIT-DAS model. We observe that although the overall integrated gradient attributions are positive for the guitarist’s name (“Jimi Hendrix”), the mediated attributions for these tokens are negative. While “Jimi Hendrix” is aligned with the image (high overall), proper names are less likely to appear in descriptions (low mediated). Dataset We hypothesize that the interpretability afforded by mediated integrated gradients will align with the inherent purposes behind describing and captioning images. Specifically, descriptions are easier to visualize than captions, since their goal is to supplant the image’s visual components as opposed to supplement them Kreiss et al. [2022b] Hence, we expect that words with higher integrated gradient attributions are easier to visualize. We consult two collections of human ratings for visualization-related concepts. The first dataset consists of 5,500 words rated by imageability, or how well a word evokes a clear mental image in the
reader’s mind Scott et al. [2019] The second dataset consists of over 40,000 words rated by 1 concreteness, or how clearly a word corresponds to a perceptible entity Brysbaert et al. [2014] We randomly sample 100 captions and 100 descriptions from the test split of the Concadia dataset that contain at least one word within both of our datasets. We compute the integrated gradient attributions for all tokens in those sentences, and report their correlations with imageability and concreteness ratings. Results Table 10.4 displays the correlations between token-level attributions of the model output and human ratings for imageability and concreteness. The IIT and IIT-DAS models achieve a 1 Although imageability and concreteness are slightly different concepts, the imageability and concreteness ratings have a correlation factor of 0.88 with each other CHAPTER 10. WHEN INTERPRETABILITY ENHANCES ACCESSIBILITY Training Mediation Concreteness Imageability None ✗ ✓ 0.20 0.14
0.16 0.07 Behavioral ✗ ✓ 0.16 ± 002 0.14 ± 002 0.22 ± 003 0.13 ± 004 IIT ✗ ✓ 0.29 ± 004 0.25 ± 004 0.32 ± 005 0.32 ± 006 IIT-DAS ✗ ✓ 0.26 ± 003 0.25 ± 006 0.28 ± 003 0.32 ± 006 158 Table 10.4: Correlation between integrated gradient attributions and per-token human labels for concreteness and imageability. The error bounds are 95% confidence intervals from runs with five random seeds. stronger correlation with imageability and concreteness ratings than the base CLIP model or those fine-tuned with the behavioral objective. Focusing in on mediated integrated gradients, we see that the pattern is amplified for imageability ratings. Mediated integrated gradients result in lower imageability correlations than their nonmediated integrated gradient counterpart for the base CLIP (007 vs 016) and the behavioral fine-tuned CLIP models (0.13 vs 022) However, mediated integrated gradients for IIT result in about the same average correlation (0.32 vs 032) and
IIT-DAS result in an ever higher correlation (0.28 vs 032) Discussion Our results show that fine-tuning CLIP to prefer descriptions over captions with IIT or IIT-DAS results in models whose attributions correspond to the human-interpretable concept of imageability and concreteness. We also find that mediating integrated gradients through the representation targetted by IIT preserves this correlation and allows for an analysis of which tokens contribute to distinguishing descriptions from captions. 10.8 Conclusion We fine-tune CLIP to prefer descriptions over captions using Concadia and find that interchange intervention training with a distributed alignment search produces the most robust, accessible, and interpretable model. Limitations Our results serve as proof concept for using IIT-DAS to update CLIP with the Concadia dataset. This is one model and one dataset, so general conclusions about the use of IIT-DAS for updating a pretrained model should not be drawn. We hope future
work will shed further light on the value of IIT-DAS. CHAPTER 10. WHEN INTERPRETABILITY ENHANCES ACCESSIBILITY 159 The Concadia dataset provides textual context for each image-description-caption triple. We do not use the context in our experiments, but we are excited about future work that incorporates this data. Whereas our work focuses on the specific purposes of describing and captioning an image, the context of an image can illuminate many other purposes (e.g search, geolocation, social communication) and models that incorporate it can enrich our work. Ethics Statement We believe that modern AI is a transformative technology that should benefit all of us and accessibility applications are an important part of this. Bibliography S. Abnar and W Zuidema Quantifying attention flow in transformers In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4190–4197, Online, July 2020. Association for Computational Linguistics doi:
1018653/v1/2020acl-main385 URL https://www.aclweborg/anthology/2020acl-main385 E. D Abraham, K D’Oosterlinck, A Feder, Y O Gat, A Geiger, C Potts, R Reichart, and Z. Wu CEBaB: Estimating the causal effects of real-world concepts on NLP model behavior arXiv:2205.14140, 2022a URL https://arxivorg/abs/220514140 E. D Abraham, K D’Oosterlinck, A Feder, Y O Gat, A Geiger, C Potts, R Reichart, and Z Wu CEBaB: Estimating the causal effects of real-world concepts on NLP model behavior. Advances in Neural Information Processing Systems, 2022b. URL https://arxivorg/abs/220514140 R. G Alhama and W Zuidema A review of computational models of basic rule learning: The neural-symbolic debate and beyond. Psychonomic bulletin & review, 26(4):1174–1194, 2019 D. Amodei, C Olah, J Steinhardt, P Christiano, J Schulman, and D Mané Concrete problems in AI safety. , 2016a URL http://arxivorg/abs/160606565 D. Amodei, C Olah, J Steinhardt, P F Christiano, J Schulman, and D Mané Concrete problems in
AI safety. CoRR, abs/160606565, 2016b URL http://arxivorg/abs/160606565 A. Andonian, Q Anthony, S Biderman, S Black, P Gali, L Gao, E Hallahan, J Levy-Kramer, C. Leahy, L Nestler, K Parker, M Pieler, S Purohit, T Songz, W Phil, and S Weinbach GPT-NeoX: Large Scale Autoregressive Language Modeling in PyTorch, 8 2021. URL https: //www.githubcom/eleutherai/gpt-neox P. Atanasova, O Camburu, C Lioma, T Lukasiewicz, J G Simonsen, and I Augenstein Faithfulness tests for natural language explanations. In A Rogers, J L Boyd-Graber, and N Okazaki, editors, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, pages 283–294. Association for Computational Linguistics, 2023. URL https://aclanthologyorg/2023acl-short25 160 BIBLIOGRAPHY 161 D. Bahdanau, S Murty, M Noukhovitch, T H Nguyen, H de Vries, and A Courville Systematic generalization: What is required and can it be learned? In In
Proceedings of the 6th International Conference on Learning Representations, Beijing, August 2018. Y. Bai, S Kadavath, S Kundu, A Askell, J Kernion, A Jones, A Chen, A Goldie, A Mirhoseini, C McKinnon, C Chen, C Olsson, C Olah, D Hernandez, D Drain, D Ganguli, D Li, E. Tran-Johnson, E Perez, J Kerr, J Mueller, J Ladish, J Landau, K Ndousse, K Lukosuite, L. Lovitt, M Sellitto, N Elhage, N Schiefer, N Mercado, N DasSarma, R Lasenby, R Larson, S. Ringer, S Johnston, S Kravec, S E Showk, S Fort, T Lanham, T Telleen-Lawton, T Conerly, T. Henighan, T Hume, S R Bowman, Z Hatfield-Dodds, B Mann, D Amodei, N Joseph, S. McCandlish, T Brown, and J Kaplan Constitutional ai: Harmlessness from ai feedback, 2022 P. Ban, Y Jiang, T Liu, and S Steinert-Threlkeld Testing pre-trained language models’ understanding of distributivity via causal mediation analysis arXiv:2209.04761, 2022 URL https://arxiv.org/abs/220904761 E. Bareinboim, J Correa, D Ibeling, and T Icard On Pearl’s hierarchy and the
foundations of causal inference. In H Geffner, R Dechter, and J Y Halpern, editors, Probabilistic and Causal Inference: The Works of Judea Pearl, pages 509–556. ACM Books, 2022 D. Bau, J Zhu, H Strobelt, B Zhou, J B Tenenbaum, W T Freeman, and A Torralba Visualizing and understanding gans. In Deep Generative Models for Highly Structured Data, ICLR 2019 Workshop, New Orleans, Louisiana, United States, May 6, 2019. OpenReviewnet, 2019a URL https://openreview.net/forum?id=rJgON8ItOV D. Bau, J Zhu, H Strobelt, B Zhou, J B Tenenbaum, W T Freeman, and A Torralba GAN dissection: Visualizing and understanding generative adversarial networks. In 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net, 2019b URL https://openreviewnet/forum?id=Hyg X2C5FX S. Beckers and J Halpern Abstracting causal models In AAAI Conference on Artificial Intelligence, 2019a. S. Beckers and J Y Halpern Abstracting causal models Proceedings of the
AAAI Conference on Artificial Intelligence, 33(01):2678–2685, Jul. 2019b doi: 101609/aaaiv33i0133012678 URL https://ojs.aaaiorg/indexphp/AAAI/article/view/4117 S. Beckers, F Eberhardt, and J Y Halpern Approximate causal abstractions In Proceedings of The 35th Uncertainty in Artificial Intelligence Conference, 2019. Y. Belinkov and J Glass Analysis methods in neural language processing: A survey Transactions of the Association for Computational Linguistics, 7:49–72, Mar. 2019a doi: 101162/tacl a 00254 URL https://www.aclweborg/anthology/Q19-1004 BIBLIOGRAPHY 162 Y. Belinkov and J R Glass Analysis methods in neural language processing: A survey Trans Assoc. Comput Linguistics, 7:49–72, 2019b doi: 101162/tacl a 00254 URL https://doiorg/ 10.1162/tacl a 00254 N. Belrose, D Schneider-Joseph, S Ravfogel, R Cotterell, E Raff, and S Biderman Leace: Perfect linear concept erasure in closed form, 2023. M. Besserve, A Mehrjou, R Sun, and B Schölkopf Counterfactuals uncover the modular
structure of deep generative models. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReviewnet, 2020a URL https://openreview net/forum?id=SJxDDpEKvH. M. Besserve, A Mehrjou, R Sun, and B Schölkopf Counterfactuals uncover the modular structure of deep generative models. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020. OpenReviewnet, 2020b URL https://openreview net/forum?id=SJxDDpEKvH. A. Binder, G Montavon, S Bach, K Müller, and W Samek Layer-wise relevance propagation for neural networks with local renormalization layers. CoRR, abs/160400825, 2016 URL http: //arxiv.org/abs/160400825 S. Bongers, P Forré, J Peters, and J M Mooij Foundations of structural causal models with cycles and latent variables. The Annals of Statistics, 49(5):2885–2915, 2021 L. Bossard, M Guillaumin, and L Van Gool Food-101 – mining discriminative components with random
forests. In European Conference on Computer Vision, 2014 S. R Bowman, G Angeli, C Potts, and C D Manning A large annotated corpus for learning natural language inference. In Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, pages 632–642, Lisbon, Portugal, Sept. 2015 Association for Computational Linguistics. doi: 1018653/v1/D15-1075 URL https://wwwaclweborg/anthology/D15-1075 S. R Bowman, J Hyun, E Perez, E Chen, C Pettit, S Heiner, K Lukošiūtė, A Askell, A Jones, A. Chen, A Goldie, A Mirhoseini, C McKinnon, C Olah, D Amodei, D Amodei, D Drain, D Li, E. Tran-Johnson, J Kernion, J Kerr, J Mueller, J Ladish, J Landau, K Ndousse, L Lovitt, N. Elhage, N Schiefer, N Joseph, N Mercado, N DasSarma, R Larson, S McCandlish, S Kundu, S. Johnston, S Kravec, S E Showk, S Fort, T Telleen-Lawton, T Brown, T Henighan, T Hume, Y. Bai, Z Hatfield-Dodds, B Mann, and J Kaplan Measuring progress on scalable oversight for large language models, 2022. P.
Brouillard, P Taslakian, A Lacoste, S Lachapelle, and A Drouin Typing assumptions improve identification in causal discovery. In First Conference on Causal Learning and Reasoning, 2022 URL https://openreview.net/forum?id=bwmG6Xp2vD0 BIBLIOGRAPHY 163 T. Brown, B Mann, N Ryder, M Subbiah, J D Kaplan, P Dhariwal, A Neelakantan, P Shyam, G. Sastry, A Askell, et al Language models are few-shot learners Advances in Neural Information Processing Systems, 2020 URL https://proceedingsneuripscc/paper/2020/file/ 1457c0d6bfcb4967418bfb8ac142f64a-Paper.pdf M. Brysbaert, A B Warriner, and V Kuperman Concreteness ratings for 40 thousand generally known english word lemmas. Behavior research methods, 46:904–911, 2014 O. Camburu, E Giunchiglia, J N Foerster, T Lukasiewicz, and P Blunsom Can I trust the explainer? verifying post-hoc explanatory methods. CoRR, abs/191002065, 2019 URL http: //arxiv.org/abs/191002065 O. Camburu, B Shillingford, P Minervini, T Lukasiewicz, and P Blunsom Make up your
mind! adversarial generation of inconsistent natural language explanations. In D Jurafsky, J Chai, N. Schluter, and J R Tetreault, editors, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 4157–4165. Association for Computational Linguistics, 2020. doi: 1018653/v1/2020acl-main382 URL https://doi org/10.18653/v1/2020acl-main382 O.-M Camburu, T Rocktäschel, T Lukasiewicz, and P Blunsom e-snli: Natural language inference with natural language explanations. In S Bengio, H Wallach, H Larochelle, K Grauman, N CesaBianchi, and R Garnett, editors, Advances in Neural Information Processing Systems, volume 31 Curran Associates, Inc., 2018 URL https://proceedingsneuripscc/paper files/paper/ 2018/file/4c7a167bb329bd92580a99ce422d6fa6-Paper.pdf N. Cammarata, S Carter, G Goh, C Olah, M Petrov, L Schubert, C Voss, B Egan, and S K Lim. Thread: Circuits Distill, 2020 doi: 1023915/distill00024
https://distillpub/2020/circuits R. Cao and D Yamins Explanatory models in neuroscience: Part 1 – taking mechanistic abstraction seriously, 2021. K. Chalupka, P Perona, and F Eberhardt Visual causal feature learning In M Meila and T Heskes, editors, Proceedings of the Thirty-First Conference on Uncertainty in Artificial Intelligence, UAI 2015, July 12-16, 2015, Amsterdam, The Netherlands, pages 181–190. AUAI Press, 2015 URL http://auai.org/uai2015/proceedings/papers/109pdf K. Chalupka, T Bischoff, F Eberhardt, and P Perona Unsupervised discovery of el nino using causal feature learning on microlevel climate data. In A T Ihler and D Janzing, editors, Proceedings of the Thirty-Second Conference on Uncertainty in Artificial Intelligence, UAI 2016, June 25-29, 2016, New York City, NY, USA. AUAI Press, 2016a URL http://auaiorg/uai2016/proceedings/ papers/11.pdf BIBLIOGRAPHY 164 K. Chalupka, F Eberhardt, and P Perona Multi-level cause-effect systems In A Gretton and C C Robert,
editors, Proceedings of the 19th International Conference on Artificial Intelligence and Statistics, volume 51 of Proceedings of Machine Learning Research, pages 361–369, Cadiz, Spain, 09–11 May 2016b. PMLR URL http://proceedingsmlrpress/v51/chalupka16html K. Chalupka, F Eberhardt, and P Perona Causal feature learning: an overview Behaviormetrika, 44:137–164, 2017. C. S Chan, H Kong, and G Liang A comparative study of faithfulness metrics for model interpretability methods. In S Muresan, P Nakov, and A Villavicencio, editors, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2022, Dublin, Ireland, May 22-27, 2022, pages 5029–5038. Association for Computational Linguistics, 2022a. doi: 1018653/v1/2022acl-long345 URL https://doiorg/1018653/v1/ 2022.acl-long345 L. Chan, A Garriga-Alonso, N Goldowsky-Dill, R Greenblatt, J Nitishinskaya, A Radhakrishnan, B. Shlegeris, and N Thomas Causal scrubbing: a method for
rigorously testing interpretability hypotheses, 2022b. A. Chattopadhyay, P Manupriya, A Sarkar, and V N Balasubramanian Neural network attributions: A causal perspective. In Proceedings of the 36th International Conference on Machine Learning (ICML), pages 981–990, 2019. Q. Chen, X Zhu, Z Ling, S Wei, and H Jiang Enhancing and combining sequential and tree LSTM for natural language inference. CoRR, abs/160906038, 2016 URL http://arxivorg/abs/1609 06038. E. A Chi, J Hewitt, and C D Manning Finding universal grammatical relations in multilingual BERT. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 5564–5577, Online, July 2020. Association for Computational Linguistics doi: 1018653/v1/ 2020.acl-main493 URL https://wwwaclweborg/anthology/2020acl-main493 W.-L Chiang, Z Li, Z Lin, Y Sheng, Z Wu, H Zhang, L Zheng, S Zhuang, Y Zhuang, J E Gonzalez, I. Stoica, and E P Xing Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt
quality, March 2023. URL https://vicunalmsysorg A. Chouldechova Fair prediction with disparate impact: A study of bias in recidivism prediction instruments. Big Data, 5(2):153–163, 2017 doi: 101089/big20160047 URL https://doiorg/ 10.1089/big20160047 P. F Christiano, J Leike, T Brown, M Martic, S Legg, and D Amodei Deep reinforcement learning from human preferences. Advances in neural information processing systems, 30, 2017 BIBLIOGRAPHY 165 H. W Chung, L Hou, S Longpre, B Zoph, Y Tay, W Fedus, E Li, X Wang, M Dehghani, S. Brahma, et al Scaling instruction-finetuned language models arXiv preprint arXiv:221011416, 2022. K. Clark, U Khandelwal, O Levy, and C D Manning What does BERT look at? an analysis of BERT’s attention. In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 276–286, Florence, Italy, Aug. 2019 Association for Computational Linguistics. doi: 1018653/v1/W19-4828 URL https://wwwaclweborg/
anthology/W19-4828. A. Coenen, E Reif, A Yuan, B Kim, A Pearce, F Viégas, and M Wattenberg Visualizing and measuring the geometry of bert, 2019. A. Conneau, G Kruszewski, G Lample, L Barrault, and M Baroni What you can cram into a single $&!#* vector: Probing sentence embeddings for linguistic properties. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2126–2136, Melbourne, Australia, July 2018. Association for Computational Linguistics doi: 10.18653/v1/P18-1198 URL https://wwwaclweborg/anthology/P18-1198 P. Cousot and R Cousot Abstract interpretation: a unified lattice model for static analysis of programs by construction or approximation of fixpoints. In Conference Record of the Fourth Annual ACM SIGPLAN-SIGACT Symposium on Principles of Programming Languages, pages 238–252, Los Angeles, California, 1977. ACM Press, New York, NY M. Crawshaw Multi-task learning with deep neural networks: A survey CoRR,
abs/200909796, 2020. URL https://arxivorg/abs/200909796 K. A Creel Transparency in complex computational systems Philosophy of Science, 87:568–589, 2020. R. Csordás, S van Steenkiste, and J Schmidhuber Are neural nets modular? inspecting functional modularity through differentiable weight masks. In 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria, May 3-7, 2021. OpenReviewnet, 2021 URL https://openreview.net/forum?id=7uVcpu-gMD S. Dash, V N Balasubramanian, and A Sharma Evaluating and mitigating bias in image classifiers: A causal perspective using counterfactuals. In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 915–924, 2022. S. Dathathri, K Dvijotham, A Kurakin, A Raghunathan, J Uesato, R Bunel, S Shankar, J. Steinhardt, I J Goodfellow, P Liang, and P Kohli Enabling certification of verificationagnostic networks via memory-efficient semidefinite programming In H Larochelle, M Ranzato,
BIBLIOGRAPHY 166 R. Hadsell, M Balcan, and H Lin, editors, Advances in Neural Information Processing Systems 33: Annual Conference on Neural Information Processing Systems 2020, NeurIPS 2020, December 6-12, 2020, virtual, 2020. URL https://proceedingsneuripscc/paper/2020/hash/ 397d6b4c83c91021fe928a8c4220386b-Abstract.html N. De Cao, W Aziz, and I Titov Editing factual knowledge in language models In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 6491–6506, Online and Punta Cana, Dominican Republic, Nov. 2021a Association for Computational Linguistics doi: 10.18653/v1/2021emnlp-main522 URL https://aclanthologyorg/2021emnlp-main522 N. De Cao, L Schmid, D Hupkes, and I Titov Sparse interventions in language models with differentiable masking. arXiv:211206837, 2021b URL https://arxivorg/abs/211206837 J. Devlin, M-W Chang, K Lee, and K Toutanova BERT: Pre-training of deep bidirec- tional transformers for language understanding. In
Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota, June 2019a. Association for Computational Linguistics doi: 1018653/v1/N19-1423 URL https://www.aclweborg/anthology/N19-1423 J. Devlin, M-W Chang, K Lee, and K Toutanova BERT: Pre-training of deep bidirec- tional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4171–4186, Minneapolis, Minnesota, June 2019b. Association for Computational Linguistics doi: 1018653/v1/N19-1423 URL https://www.aclweborg/anthology/N19-1423 J. Devlin, M-W Chang, K Lee, and K Toutanova BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the
North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Minneapolis, Minnesota, 2019c. URL https://wwwaclweborg/anthology/N19-1423 V. Do, O Camburu, Z Akata, and T Lukasiewicz e-snli-ve-20: Corrected visual-textual entailment with natural language explanations. CoRR, abs/200403744, 2020 URL https://arxivorg/ abs/2004.03744 R. Dominguez-Olmedo, A H Karimi, and B Schölkopf On the adversarial robustness of causal algorithmic recourse. In K Chaudhuri, S Jegelka, L Song, C Szepesvari, G Niu, and S Sabato, editors, Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pages 5324–5342. PMLR, 17–23 Jul 2022 URL https://proceedings.mlrpress/v162/dominguez-olmedo22ahtml BIBLIOGRAPHY 167 J. Dubois, F Eberhardt, L K Paul, and R Adolphs Personality beyond taxonomy Nature human behaviour, 4 11:1110–1117, 2020a. J. Dubois, H Oya, J M Tyszka, M A Howard, F Eberhardt, and R
Adolphs Causal mapping of emotion networks in the human brain: Framework and initial findings. Neuropsychologia, 145, 2020b. C. Dwork, M Hardt, T Pitassi, O Reingold, and R S Zemel Fairness through awareness In S. Goldwasser, editor, Innovations in Theoretical Computer Science 2012, Cambridge, MA, USA, January 8-10, 2012, pages 214–226. ACM, 2012 doi: 101145/20902362090255 URL https://doi.org/101145/20902362090255 U. Ehsan, Q V Liao, M Muller, M O Riedl, and J D Weisz Expanding explainability: Towards social transparency in AI systems. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems, 2021. URL https://dlacmorg/doi/pdf/101145/34117643445188 Y. Elazar, S Ravfogel, A Jacovi, and Y Goldberg Amnesic probing: Behavioral explanation with amnesic counterfactuals. In Proceedings of the 2020 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP. Association for Computational Linguistics, Nov 2020 doi: 10.18653/v1/W18-5426 Y. Elazar, S
Ravfogel, A Jacovi, and Y Goldberg Amnesic Probing: Behavioral Explanation with Amnesic Counterfactuals. Transactions of the Association for Computational Linguistics, 9:160–175, 03 2021a. ISSN 2307-387X doi: 101162/tacl a 00359 URL https://doiorg/10 1162/tacl a 00359. Y. Elazar, S Ravfogel, A Jacovi, and Y Goldberg Amnesic probing: Behavioral explana- tion with amnesic counterfactuals. Transactions of the Association for Computational Linguistics, 2021b URL https://directmitedu/tacl/article/doi/101162/tacl a 00359/98091/ Amnesic-Probing-Behavioral-Explanation-with. Y. Elazar, N Kassner, S Ravfogel, A Feder, A Ravichander, M Mosbach, Y Belinkov, H Schütze, and Y. Goldberg Measuring causal effects of data statistics on language model’s ’factual’ predictions CoRR, abs/2207.14251, 2022 doi: 1048550/arXiv220714251 URL https://doiorg/1048550/ arXiv.220714251 N. Elhage, N Nanda, C Olsson, T Henighan, N Joseph, B Mann, A Askell, Y Bai, A Chen, T. Conerly, N DasSarma, D Drain, D
Ganguli, Z Hatfield-Dodds, D Hernandez, A Jones, J. Kernion, L Lovitt, K Ndousse, D Amodei, T Brown, J Clark, J Kaplan, S McCandlish, and C. Olah A mathematical framework for transformer circuits Transformer Circuits Thread, 2021 https://transformer-circuits.pub/2021/framework/indexhtml BIBLIOGRAPHY 168 N. Elhage, T Hume, C Olsson, N Nanda, T Henighan, S Johnston, S ElShowk, N Joseph, N. DasSarma, B Mann, D Hernandez, A Askell, K Ndousse, A Jones, D Drain, A Chen, Y Bai, D. Ganguli, L Lovitt, Z Hatfield-Dodds, J Kernion, T Conerly, S Kravec, S Fort, S Kadavath, J. Jacobson, E Tran-Johnson, J Kaplan, J Clark, T Brown, S McCandlish, D Amodei, and C. Olah Softmax linear units Transformer Circuits Thread, 2022 https://transformercircuitspub/2022/solu/indexhtml G. Erion, J D Janizek, P Sturmfels, S M Lundberg, and S-I Lee Improving performance of deep learning models with axiomatic attribution priors and expected gradients. Nature Machine Intelligence, 3(7):620–631, 2021. doi:
101038/s42256-021-00343-w URL https://doiorg/10 1038/s42256-021-00343-w. A. Feder, N Oved, U Shalit, and R Reichart CausaLM: Causal model explanation through counterfactual language models. Computational Linguistics, 2020 URL https://aclanthology org/2021.cl-213 A. Feder, N Oved, U Shalit, and R Reichart CausaLM: Causal Model Explanation Through Counterfactual Language Models. Computational Linguistics, pages 1–54, 05 2021 ISSN 0891-2017 doi: 10.1162/coli a 00404 URL https://doiorg/101162/coli a 00404 J. Feinglass and Y Yang Smurf: Semantic and linguistic understanding fusion for caption evaluation via typicality analysis. arXiv preprint arXiv:210601444, 2021 C. Fellbaum, editor WordNet: An Electronic Database MIT Press, Cambridge, MA, 1998 J. A Fodor and Z W Pylyshyn Connectionism and cognitive architecture: A critical analysis Cognition, 28(1):3–71, 1988. ISSN 0010-0277 doi: https://doiorg/101016/0010-0277(88)90031-5 URL
http://www.sciencedirectcom/science/article/pii/0010027788900315 R. Gandikota, J Materzynska, J Fiotto-Kaufman, and D Bau Erasing concepts from diffusion models, 2023. A. Geiger, I Cases, L Karttunen, and C Potts Posing fair generalization tasks for natural language inference. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLPIJCNLP), pages 4475–4485, Stroudsburg, PA, November 2019a. Association for Computational Linguistics. doi: 1018653/v1/D19-1456 URL https://wwwaclweborg/anthology/D19-1456 A. Geiger, I Cases, L Karttunen, and C Potts Posing fair generalization tasks for natural language inference. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLPIJCNLP), pages 4485–4495, Hong Kong, China, Nov. 2019b Association for Computational Linguistics.
doi: 1018653/v1/D19-1456 URL https://wwwaclweborg/anthology/D19-1456 BIBLIOGRAPHY 169 A. Geiger, K Richardson, and C Potts Neural natural language inference models partially embed theories of lexical entailment and negation. In Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pages 163–173, Online, Nov. 2020a Association for Computational Linguistics. doi: 1018653/v1/2020blackboxnlp-116 URL https: //www.aclweborg/anthology/2020blackboxnlp-116 A. Geiger, K Richardson, and C Potts Neural natural language inference models partially embed theories of lexical entailment and negation. In Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pages 163–173, Online, Nov. 2020b Association for Computational Linguistics. doi: 1018653/v1/2020blackboxnlp-116 URL https: //www.aclweborg/anthology/2020blackboxnlp-116 A. Geiger, H Lu, T Icard, and C Potts Causal abstractions of neural networks In
Advances in Neural Information Processing Systems, volume 34, pages 9574–9586, 2021a. URL https: //papers.nipscc/paper/2021/hash/4f5c422f4d49a5a807eda27434231040-Abstracthtml A. Geiger, H Lu, T Icard, and C Potts Causal abstractions of neural networks In Advances in Neural Information Processing Systems, volume 34, pages 9574–9586, 2021b. URL https: //papers.nipscc/paper/2021/hash/4f5c422f4d49a5a807eda27434231040-Abstracthtml A. Geiger, H Lu, T Icard, and C Potts Causal abstractions of neural networks. In M. Ranzato, A Beygelzimer, Y Dauphin, P Liang, and J W Vaughan, editors, Advances in Neural Information Processing Systems, volume 34, pages 9574–9586. Curran Associates, Inc, 2021c URL https://proceedings.neuripscc/paper/2021/file/ 4f5c422f4d49a5a807eda27434231040-Paper.pdf A. Geiger, H Lu, T F Icard, and C Potts Causal abstractions of neural networks In A Beygelzimer, Y. Dauphin, P Liang, and J W Vaughan, editors, Advances in Neural Information Processing Systems, 2021d.
URL https://openreviewnet/forum?id=RmuXDtjDhG A. Geiger, A Carstensen, M C Frank, and C Potts Relational reasoning and generalization using nonsymbolic neural networks. Psychological Review, 2022a doi: 101037/rev0000371 URL https://doi.org/101037/rev0000371 A. Geiger, A Carstensen, M C Frank, and C Potts Relational reasoning and generalization using nonsymbolic neural networks. Psychological Review, 2022b doi: 101037/rev0000371 A. Geiger, Z Wu, H Lu, J Rozner, E Kreiss, T Icard, N Goodman, and C Potts Inducing causal structure for interpretable neural networks. In K Chaudhuri, S Jegelka, L Song, C Szepesvari, G. Niu, and S Sabato, editors, Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pages 7324–7338. PMLR, 17–23 Jul 2022c. URL https://proceedingsmlrpress/v162/geiger22ahtml BIBLIOGRAPHY 170 A. Geiger, Z Wu, H Lu, J Rozner, E Kreiss, T Icard, N Goodman, and C Potts Inducing causal structure for
interpretable neural networks. In International Conference on Machine Learning, pages 7324–7338. PMLR, 2022d A. Geiger, Z Wu, H Lu, J Rozner, E Kreiss, T Icard, N Goodman, and C Potts Inducing causal structure for interpretable neural networks. In K Chaudhuri, S Jegelka, L Song, C Szepesvari, G. Niu, and S Sabato, editors, Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pages 7324–7338. PMLR, 17–23 Jul 2022e. URL https://proceedingsmlrpress/v162/geiger22ahtml A. Geiger, C Potts, and T Icard Causal abstraction for faithful model interpretation Ms, Stanford University, 2023a. URL https://arxivorg/abs/230104709 A. Geiger, C Potts, and T Icard Causal abstraction for faithful model interpretation, 2023b A. Geiger, C Potts, and T Icard Causal abstraction for faithful interpretation of AI models arXiv:2106.02997, 2023c URL https://arxivorg/abs/210602997 A. Geiger, Z Wu, C Potts, T Icard, and N D Goodman
Finding alignments between interpretable causal variables and distributed neural representations. Ms, Stanford University, 2023d URL https://arxiv.org/abs/230302536 A. Geiger, Z Wu, C Potts, T Icard, and N D Goodman Finding alignments between interpretable causal variables and distributed neural representations, 2023e. M. Giulianelli, J Harding, F Mohnert, D Hupkes, and W Zuidema Under the hood: Using diagnostic classifiers to investigate and improve how language models track agreement information. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 240–248, Brussels, Belgium, Nov. 2018a Association for Computational Linguistics. doi: 1018653/v1/W18-5426 URL https://wwwaclweborg/anthology/W18-5426 M. Giulianelli, J Harding, F Mohnert, D Hupkes, and W H Zuidema Under the hood: Using diagnostic classifiers to investigate and improve how language models track agreement information. In T. Linzen, G Chrupala, and A Alishahi,
editors, Proceedings of the Workshop: Analyzing and Interpreting Neural Networks for NLP, BlackboxNLP@EMNLP 2018, Brussels, Belgium, November 1, 2018, pages 240–248. Association for Computational Linguistics, 2018b doi: 1018653/v1/ w18-5426. URL https://doiorg/1018653/v1/w18-5426 M. Giulianelli, M Del Tredici, and R Fernández Analysing lexical semantic change with contextualised word representations In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 3960–3973, Online, July 2020. Association for Computational Linguistics. doi: 1018653/v1/2020acl-main365 URL https://wwwaclweborg/anthology/ 2020.acl-main365 BIBLIOGRAPHY 171 C. Gleason, P Carrington, C Cassidy, M R Morris, K M Kitani, and J P Bigham “it’s almost like they’re trying to hide it”: How user-provided image descriptions have failed to make twitter accessible. In The World Wide Web Conference, pages 549–559, 2019 C. Gleason, A Pavel, E McCamey, C Low, P
Carrington, K M Kitani, and J P Bigham Twitter a11y: A browser extension to make twitter images accessible. In Proceedings of the 2020 chi conference on human factors in computing systems, pages 1–12, 2020. M. Glockner, V Shwartz, and Y Goldberg Breaking NLI systems with sentences that require simple lexical inferences. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 650–655, Melbourne, Australia, July 2018. Association for Computational Linguistics doi: 1018653/v1/P18-2103 URL https: //www.aclweborg/anthology/P18-2103 B. Goodman and S Flaxman European Union regulations on algorithmic decision-making and a “right to explanation”. AI Magazine, 2017 URL http://arxivorg/abs/160608813 E. Goodwin, K Sinha, and T J O’Donnell Probing linguistic systematicity, 2020 Y. Goyal, U Shalit, and B Kim Explaining classifiers with causal concept effect (cace) CoRR, abs/1907.07165, 2019a URL
http://arxivorg/abs/190707165 Y. Goyal, Z Wu, J Ernst, D Batra, D Parikh, and S Lee Counterfactual visual explanations In International Conference on Machine Learning, 2019b. URL http://proceedingsmlrpress/ v97/goyal19a.html Y. Goyal, A Feder, U Shalit, and B Kim Explaining Classifiers with Causal Concept Effect (CaCE). arXiv:190707165 [cs, stat], Feb 2020 URL http://arxivorg/abs/190707165 arXiv: 1907.07165 L. Gresele, J V Kügelgen, J Kübler, E Kirschbaum, B Schölkopf, and D Janzing Causal inference through the structural causal marginal problem. In K Chaudhuri, S Jegelka, L Song, C. Szepesvari, G Niu, and S Sabato, editors, Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pages 7793–7824, 2022. R. Guidotti, A Monreale, S Ruggieri, F Turini, F Giannotti, and D Pedreschi A survey of methods for explaining black box models. ACM computing surveys (CSUR), 2018 URL https://dl.acmorg/doi/abs/101145/3236009
S. Gururangan, S Swayamdipta, O Levy, R Schwartz, S Bowman, and N A Smith Annotation artifacts in natural language inference data In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human BIBLIOGRAPHY 172 Language Technologies, Volume 2 (Short Papers), pages 107–112, New Orleans, Louisiana, June 2018. Association for Computational Linguistics doi: 10.18653/v1/N18-2017 URL https://www.aclweborg/anthology/N18-2017 D. Hadfield-Menell, S Milli, P Abbeel, S J Russell, and A D Dragan Inverse reward design In I Guyon, U von Luxburg, S Bengio, H M Wallach, R Fergus, S V N Vishwanathan, and R Garnett, editors, Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 6765–6774, 2017. URL https://proceedingsneuripscc/paper/2017/ hash/32fdab6559cdfa4f167f8c31b9199643-Abstract.html M. Hardt, E Price, and N Srebro
Equality of opportunity in supervised learning In D D Lee, M. Sugiyama, U von Luxburg, I Guyon, and R Garnett, editors, Advances in Neural Information Processing Systems 29: Annual Conference on Neural Information Processing Systems 2016, December 5-10, 2016, Barcelona, Spain, pages 3315–3323, 2016a. URL https://proceedings neurips.cc/paper/2016/hash/9d2682367c3935defcb1f9e247a97c0d-Abstracthtml M. Hardt, E Price, and N Srebro Equality of opportunity in supervised learning In Advances in Neural Information Processing Systems Curran Associates, Inc, 2016b URL https://proceedings neurips.cc/paper/2016/file/9d2682367c3935defcb1f9e247a97c0d-Paperpdf K. He, X Zhang, S Ren, and J Sun Deep residual learning for image recognition In 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 770–778, 2016. doi: 10.1109/CVPR201690 L. A Hendricks, Z Akata, M Rohrbach, J Donahue, B Schiele, and T Darrell Generating visual explanations. In B Leibe, J Matas, N Sebe, and M
Welling, editors, Computer Vision - ECCV 2016 - 14th European Conference, Amsterdam, The Netherlands, October 11-14, 2016, Proceedings, Part IV, volume 9908 of Lecture Notes in Computer Science, pages 3–19. Springer, 2016 doi: 10.1007/978-3-319-46493-0 1 URL https://doiorg/101007/978-3-319-46493-0 1 B. Herman The Promise and Peril of Human Evaluation for Model Interpretability arXiv e-prints, art. arXiv:171107414, Nov 2017 J. Hessel, A Holtzman, M Forbes, R L Bras, and Y Choi Clipscore: A reference-free evaluation metric for image captioning. CoRR, abs/210408718, 2021 URL https://arxivorg/abs/2104 08718. J. Hewitt and P Liang Designing and interpreting probes with control tasks In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 2733–2743, Hong BIBLIOGRAPHY 173 Kong, China, Nov. 2019 Association for Computational Linguistics doi:
1018653/v1/D19-1275 URL https://www.aclweborg/anthology/D19-1275 J. Hewitt and C D Manning A structural probe for finding syntax in word representations In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4129–4138, 2019. J. Hewitt, K Ethayarajh, P Liang, and C D Manning Conditional probing: measuring usable information beyond a baseline. In M Moens, X Huang, L Specia, and S W Yih, editors, Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, EMNLP 2021, Virtual Event / Punta Cana, Dominican Republic, 7-11 November, 2021, pages 1626–1639. Association for Computational Linguistics, 2021. doi: 1018653/v1/2021emnlp-main122 URL https://doi.org/1018653/v1/2021emnlp-main122 G. Hinton, O Vinyals, J Dean, et al Distilling the knowledge in a neural network NeurIPS Deep Learning and Representation Learning Workshop, 2015.
URL https://arxivorg/abs/ 1503.02531 G. E Hinton Connectionist learning procedures Artificial Intelligence, 40(1):185–234, 1989 ISSN 0004-3702. doi: https://doi.org/101016/0004-3702(89)90049-0 URL https://www. sciencedirect.com/science/article/pii/0004370289900490 C. Hitchcock The intransitivity of causation revealed in equations and graphs Journal of Philosophy, 98(6):273–299, 2001. S. Hochreiter and J Schmidhuber Long short-term memory Neural Computation, 1997 URL https://ieeexplore.ieeeorg/abstract/document/6795963 P. W Holland Statistics and causal inference Journal of the American statistical Association, 1986 URL https://www.tandfonlinecom/doi/abs/101080/01621459198610478354 H. Hu, Q Chen, and L Moss Natural language inference with monotonicity In Proceedings of the 13th International Conference on Computational Semantics - Short Papers, pages 8–15, Gothenburg, Sweden, May 2019a. Association for Computational Linguistics doi: 1018653/v1/W19-0502 URL
https://www.aclweborg/anthology/W19-0502 H. Hu, Q Chen, K Richardson, A Mukherjee, L S Moss, and S Kübler MonaLog: A lightweight system for natural language inference based on monotonicity. ArXiv, abs/191008772, 2019b Y. Hu and J Tian Neuron dependency graphs: A causal abstraction of neural networks In K. Chaudhuri, S Jegelka, L Song, C Szepesvari, G Niu, and S Sabato, editors, Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine BIBLIOGRAPHY 174 Learning Research, pages 9020–9040. PMLR, 17–23 Jul 2022 URL https://proceedingsmlr press/v162/hu22b.html J. Huang, Z Wu, K Mahowald, and C Potts Inducing character-level structure in subword-based language models with Type-level Interchange Intervention Training. Ms, Stanford University and UT Austin, 2022. URL https://arxivorg/abs/221209897 J. Huang, A Geiger, K D’Oosterlinck, Z Wu, and C Potts Rigorously assessing natural language explanations of neurons. Ms, Stanford
University, 2023 URL https://arxivorg/abs/2309 10312. E. Hubinger Chris olah’s views on agi safety, 2019 D. Hupkes, S Bouwmeester, and R Fernández Analysing the potential of seq-to-seq models for incremental interpretation in task-oriented dialogue. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 165–174, Brussels, Belgium, Nov. 2018a Association for Computational Linguistics doi: 1018653/v1/W18-5419 URL https://www.aclweborg/anthology/W18-5419 D. Hupkes, S Veldhoen, and W H Zuidema Visualisation and ’diagnostic classifiers’ reveal how recurrent and recursive neural networks process hierarchical structure. J Artif Intell Res, 61: 907–926, 2018b. doi: 101613/jair111196 URL https://doiorg/101613/jair111196 D. Hupkes, V Dankers, M Mul, and E Bruni Compositionality decomposed: how do neural networks generalise?, 2019. D. Hupkes, M Giulianelli, V Dankers, M Artetxe, Y Elazar, T Pimentel, C Christodoulopoulos, K.
Lasri, N Saphra, A Sinclair, D Ulmer, F Schottmann, K Batsuren, K Sun, K Sinha, L. Khalatbari, M Ryskina, R Frieske, R Cotterell, and Z Jin State-of-the-art generalisation research in NLP: a taxonomy and review. CoRR, abs/221003050, 2022 doi: 1048550/arXiv2210 03050. URL https://doiorg/1048550/arXiv221003050 D. Ibeling and T Icard On the conditional logic of simulation models In Proceedings of the TwentySeventh International Joint Conference on Artificial Intelligence, IJCAI-18, pages 1868–1874 International Joint Conferences on Artificial Intelligence Organization, 7 2018. doi: 1024963/ijcai 2018/258. URL https://doiorg/1024963/ijcai2018/258 D. Ibeling and T Icard A topological perspective on causal inference In Proceedings of the Thirty-fifth Conference on Neural Information Processing Systems (NeurIPS), 2021. T. Icard, L Moss, and W Tune A monotonicity calculus and its completeness In Proceedings of the 15th Meeting on the Mathematics of Language, pages 75–87, London, UK, July
2017. Association BIBLIOGRAPHY 175 for Computational Linguistics. doi: 1018653/v1/W17-3408 URL https://wwwaclweborg/ anthology/W17-3408. T. F Icard Inclusion and exclusion in natural language Studia Logica, 100(4):705–725, 2012 T. F Icard From programs to causal models In Proceedings of the 21st Amsterdam Colloquium, 2017a. T. F Icard From programs to causal models In A Cremers, T van Gessel, and F Roelofsen, editors, Proceedings of the 21st Amsterdam Colloquium, pages 35–44. University of Amsterdam, 2017b. T. F Icard and L S Moss Recent progress on monotonicity Linguistic Issues in Language Technology, 9(7):1–31, January 2013. G. W Imbens and D B Rubin Causal inference in statistics, social, and biomedical sciences Cambridge University Press, 2015. Y. Iwasaki and H A Simon Causality and model abstraction Artificial Intelligence, 67(1):143–194, 1994. A. Jacovi and Y Goldberg Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness? In D.
Jurafsky, J Chai, N Schluter, and J R Tetreault, editors, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 4198–4205. Association for Computational Linguistics, 2020a doi: 10.18653/v1/2020acl-main386 URL https://doiorg/1018653/v1/2020acl-main386 A. Jacovi and Y Goldberg Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness? In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4198–4205, Online, July 2020b. Association for Computational Linguistics. doi: 1018653/v1/2020acl-main386 URL https://wwwaclweborg/anthology/ 2020.acl-main386 A. Jacovi and Y Goldberg Towards faithfully interpretable NLP systems: How should we define and evaluate faithfulness? In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020c. URL https://aclanthologyorg/2020acl-main386 M. Jakesch, M French, X Ma, J T
Hancock, and M Naaman AI-mediated communication: How the perception that profile text was written by AI affects trustworthiness. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, 2019. URL https: //dl.acmorg/doi/abs/101145/32906053300469 BIBLIOGRAPHY 176 M. Jang, B P Majumder, J J McAuley, T Lukasiewicz, and O Camburu KNOW how to make up your mind! adversarially detecting and alleviating inconsistencies in natural language explanations. In A. Rogers, J L Boyd-Graber, and N Okazaki, editors, Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, pages 540–553. Association for Computational Linguistics, 2023 URL https://aclanthology.org/2023acl-short47 R. Jia and P Liang Adversarial examples for evaluating reading comprehension systems CoRR, abs/1707.07328, 2017 URL http://arxivorg/abs/170707328 X. Jiao, Y Yin, L Shang, X Jiang, X Chen, L Li, F Wang,
and Q Liu TinyBERT: Distilling BERT for natural language understanding. arXiv:190910351, 2019 E. Jonas and K P Kording Could a neuroscientist understand a microprocessor? PloS Computational Biology, 13(1), 2017. A. Karimi, G Barthe, B Schölkopf, and I Valera A survey of algorithmic recourse: Contrastive explanations and consequential recommendations. ACM Comput Surv, 55(5):95:1–95:29, 2023 doi: 10.1145/3527848 URL https://doiorg/101145/3527848 D. Kaushik, E H Hovy, and Z C Lipton Learning the difference that makes a difference with counterfactually-augmented data. CoRR, abs/190912434, 2019 URL http://arxivorg/abs/ 1909.12434 M. Kayser, O-M Camburu, L Salewski, C Emde, V Do, Z Akata, and T Lukasiewicz E-vil: A dataset and benchmark for natural language explanations in vision-language tasks. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), pages 1244–1254, October 2021. M. Kayser, C Emde, O Camburu, G Parsons, B W Papiez, and T Lukasiewicz
Explaining chest xray pathologies in natural language In L Wang, Q Dou, P T Fletcher, S Speidel, and S Li, editors, Medical Image Computing and Computer Assisted Intervention - MICCAI 2022 - 25th International Conference, Singapore, September 18-22, 2022, Proceedings, Part V, volume 13435 of Lecture Notes in Computer Science, pages 701–713. Springer, 2022 doi: 101007/978-3-031-16443-9 67 URL https://doi.org/101007/978-3-031-16443-9 67 D. Kiela, M Bartolo, Y Nie, D Kaushik, A Geiger, Z Wu, B Vidgen, G Prasad, A Singh, P. Ringshia, Z Ma, T Thrush, S Riedel, Z Waseem, P Stenetorp, R Jia, M Bansal, C Potts, and A. Williams Dynabench: Rethinking benchmarking in NLP In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4110–4124, Online, June 2021. Association for Computational Linguistics doi: 10.18653/v1/2021naacl-main324 URL https://aclanthologyorg/2021naacl-main324 BIBLIOGRAPHY 177
B. Kim Interactive and Interpretable Machine Learning Models for Human Machine Collaboration PhD thesis, Massachusetts Institute of Technology, 2015. URL https://dspacemitedu/handle/ 1721.1/98680 B. Kim, M Wattenberg, J Gilmer, C Cai, J Wexler, F Viegas, and R Sayres Interpretability beyond feature attribution: Quantitative testing with concept activation vectors (TCAV). In International Conference on Machine Learning, 2018. URL http://proceedingsmlrpress/ v80/kim18d.html J. Kleinberg, S Mullainathan, and M Raghavan Inherent trade-offs in the fair determination of risk scores. In 8th Innovations in Theoretical Computer Science Conference, 2017 URL http://drops.dagstuhlde/opus/volltexte/2017/8156 P. W Koh, T Nguyen, Y S Tang, S Mussmann, E Pierson, B Kim, and P Liang Concept bottleneck models. In International Conference on Machine Learning, 2020 URL https:// proceedings.mlrpress/v119/koh20ahtml T. Kojima, S S Gu, M Reid, Y Matsuo, and Y Iwasawa Large language models are zero-shot
reasoners, 2022. E. Kreiss, C Bennett, S Hooshmand, E Zelikman, M R Morris, and C Potts Context matters for image descriptions for accessibility: Challenges for referenceless evaluation metrics. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 4685–4697, 2022a. E. Kreiss, F Fang, N Goodman, and C Potts Concadia: Towards image-based text generation with a purpose. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 4667–4684, 2022b. A. Krizhevsky, G Hinton, et al Learning multiple layers of features from tiny images ms, 2009 S. R Künzel, J S Sekhon, P J Bickel, and B Yu Metalearners for estimating heterogeneous treatment effects using machine learning. Proceedings of the National Academy of Sciences, 2019 J. v Kügelgen, A-H Karimi, U Bhatt, I Valera, A Weller, and B Schölkopf On the fairness of causal algorithmic recourse. Proceedings of the AAAI Conference on Artificial Intelligence,
36(9): 9584–9594, Jun. 2022 doi: 101609/aaaiv36i921192 URL https://ojsaaaiorg/indexphp/ AAAI/article/view/21192. B. Lake and M Baroni Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks. In International Conference on Machine Learning, pages 2873–2882. PMLR, 2018a BIBLIOGRAPHY 178 B. M Lake and M Baroni Generalization without systematicity: On the compositional skills of sequence-to-sequence recurrent networks. In Proceedings of the 35th International Conference on Machine Learning, volume 80 of Proceedings of Machine Learning Research, pages 2879–2888. PMLR, 2018b. Y. LeCun, C Cortes, and C Burges Mnist handwritten digit database ATT Labs [Online] Available: http://yann.lecuncom/exdb/mnist, 2, 2010 H. Lee, S Yoon, F Dernoncourt, T Bui, and K Jung Umic: An unreferenced metric for image captioning via contrastive learning. arXiv preprint arXiv:210614019, 2021 D. Lewis Postscript C to ‘Causation’: (Insensitive
causation) In Philosophical Papers, volume 2 Oxford University Press, Oxford, 1986. B. Z Li, M Nye, and J Andreas Implicit representations of meaning in neural language models In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1813–1827, Online, Aug. 2021a Association for Computational Linguistics doi: 10.18653/v1/2021acl-long143 URL https://aclanthologyorg/2021acl-long143 B. Z Li, M I Nye, and J Andreas Implicit representations of meaning in neural language models In C. Zong, F Xia, W Li, and R Navigli, editors, Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, ACL/IJCNLP 2021, (Volume 1: Long Papers), Virtual Event, August 1-6, 2021, pages 1813–1827. Association for Computational Linguistics, 2021b doi:
10.18653/v1/2021acl-long143 URL https://doiorg/1018653/v1/2021acl-long143 T. P Lillicrap and K P Kording What does it mean to understand a neural network?, 2019 W. Ling, D Yogatama, C Dyer, and P Blunsom Program induction by rationale generation: Learning to solve and explain algebraic word problems. In R Barzilay and M Kan, editors, Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, ACL 2017, Vancouver, Canada, July 30 - August 4, Volume 1: Long Papers, pages 158–167. Association for Computational Linguistics, 2017. doi: 1018653/v1/P17-1015 URL https://doiorg/1018653/v1/P17-1015 Z. C Lipton The mythos of model interpretability Commun ACM, 61(10):36–43, 2018a doi: 10.1145/3233231 URL https://doiorg/101145/3233231 Z. C Lipton The mythos of model interpretability: In machine learning, the concept of interpretability is both important and slippery. Queue, 16(3):31–57, 2018b Z. C Lipton The mythos of model interpretability Communications of the
ACM, 2018c URL https://doi.org/101145/3233231 BIBLIOGRAPHY 179 N. F Liu, R Schwartz, and N A Smith Inoculation by fine-tuning: A method for analyzing challenge datasets. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 2171–2179, Minneapolis, Minnesota, June 2019a. Association for Computational Linguistics doi: 10.18653/v1/N19-1225 URL https://wwwaclweborg/anthology/N19-1225 Q. Liu, M Kusner, and P Blunsom Counterfactual data augmentation for neural machine translation In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 187–197, Online, June 2021. Association for Computational Linguistics. doi: 1018653/v1/2021naacl-main18 URL https: //aclanthology.org/2021naacl-main18 Y. Liu, M Ott, N Goyal, J Du, M Joshi, D Chen, O Levy, M Lewis, L
Zettlemoyer, and V. Stoyanov RoBERTa: A robustly optimized BERT pretraining approach , 2019b URL http://arxiv.org/abs/190711692 C. Lovering and E Pavlick Unit testing for concepts in neural networks CoRR, abs/220810244, 2022a. doi: 1048550/arXiv220810244 URL https://doiorg/1048550/arXiv220810244 C. Lovering and E Pavlick Unit testing for concepts in neural networks Transactions of the Association for Computational Linguistics, 10:1193–1208, 2022b. doi: 101162/tacl a 00514 URL https://aclanthology.org/2022tacl-169 C. Lovering and E Pavlick Unit testing for concepts in neural networks. arXiv preprint arXiv:2208.10244, 2022c S. Lundberg and S-I Lee A unified approach to interpreting model predictions, 2017a S. M Lundberg and S-I Lee A unified approach to interpreting model predictions. In I. Guyon, U V Luxburg, S Bengio, H Wallach, R Fergus, S Vishwanathan, and R. Garnett, editors, Advances in Neural Information Processing Systems, volume 30 Curran Associates, Inc, 2017b URL
https://proceedings.neuripscc/paper/2017/file/ 8a20a8621978632d76c43dfd28b67767-Paper.pdf Q. Lyu, M Apidianaki, and C Callison-Burch Towards faithful model explanation in NLP: A survey. CoRR, abs/220911326, 2022 doi: 1048550/arXiv220911326 URL https://doiorg/ 10.48550/arXiv220911326 B. MacCartney Natural Language Inference PhD thesis, Stanford University, 2009 B. MacCartney and C D Manning Natural logic for textual inference In Proceedings of the ACL-PASCAL Workshop on Textual Entailment and Paraphrasing, RTE ’07, pages 193–200, BIBLIOGRAPHY 180 Stroudsburg, PA, USA, 2007. Association for Computational Linguistics URL http://dlacm org/citation.cfm?id=16545361654575 B. MacCartney and C D Manning An extended model of natural logic In Proceedings of the Eight International Conference on Computational Semantics, pages 140–156, Tilburg, The Netherlands, Jan. 2009 Association for Computational Linguistics URL https://wwwaclweb org/anthology/W09-3714. P. Machamer, L Darden, and C F
Craver Thinking about mechanisms Philosophy of Science, 67 (1):1, mar 2000. H. MacLeod, C L Bennett, M R Morris, and E Cutrell Understanding blind people’s experiences with computer-generated captions of social media images. In proceedings of the 2017 CHI conference on human factors in computing systems, pages 5988–5999, 2017. B. P Majumder, O Camburu, T Lukasiewicz, and J Mcauley Knowledge-grounded self- rationalization via extractive and natural language explanations. In K Chaudhuri, S Jegelka, L. Song, C Szepesvari, G Niu, and S Sabato, editors, Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pages 14786– 14801. PMLR, 17–23 Jul 2022 URL https://proceedingsmlrpress/v162/majumder22ahtml C. D Manning, K Clark, J Hewitt, U Khandelwal, and O Levy Emergent linguistic structure in artificial neural networks trained by self-supervision. Proceedings of the National Academy of Sciences,
117(48):30046–30054, 2020. ISSN 0027-8424 doi: 101073/pnas1907367117 URL https://www.pnasorg/content/117/48/30046 G. F Marcus, S Vijayan, S B Rao, and P M Vishton Rule learning by seven-month-old infants Science, 283(5398):77–80, 1999. D. Marr Vision: A Computational Investigation into the Human Representation and Processing of Visual Information. Henry Holt and Co, Inc, New York, NY, USA, 1982 ISBN 0716715678 R. Massida, A Geiger, T Icard, and D Bacciu Causal abstraction with soft interventions, 2022 URL https://arxiv.org/abs/221112270 J. Materzynska, A Torralba, and D Bau Disentangling visual and written concepts in clip 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 16389–16398, 2022. URL https://apisemanticscholarorg/CorpusID:249711999 J. L McClelland, D E Rumelhart, and PDP Research Group, editors Parallel Distributed Processing Volume 2: Psychological and Biological Models. MIT Press, Cambridge, MA, 1986 BIBLIOGRAPHY 181 R. T McCoy, T
Linzen, E Dunbar, and P Smolensky Rnns implicitly implement tensor product representations. In In Proceedings of the 7th International Conference on Learning Representations, New Orleans, USA, May 2019. N. Mehrabi, F Morstatter, N Saxena, K Lerman, and A Galstyan A survey on bias and fairness in machine learning. ACM Computing Surveys (CSUR), 2021 URL https://dlacmorg/doi/ abs/10.1145/3457607 K. Meng, D Bau, A Andonian, and Y Belinkov Locating and editing factual associations in gpt, 2022a. URL https://arxivorg/abs/220205262 K. Meng, D Bau, A Andonian, and Y Belinkov Locating and editing factual associations in GPT Advances in Neural Information Processing Systems, 36, 2022b. K. Meng, A S Sharma, A J Andonian, Y Belinkov, and D Bau Mass-editing memory in a transformer. In The Eleventh International Conference on Learning Representations, 2023 URL https://openreview.net/forum?id=MkbcAHIYgyS S. Merity, C Xiong, J Bradbury, and R Socher Pointer sentinel mixture models arXiv:160907843,
2016. URL https://arxivorg/abs/160907843 G. A Miller Wordnet: a lexical database for english Communications of the ACM, 38(11):39–41, 1995. E. Mitchell, C Lin, A Bosselut, C Finn, and C D Manning Fast model editing at scale In International Conference on Learning Representations, 2022. URL https://openreviewnet/ forum?id=0DcZxeWfOPt. S. Mitchell, E Potash, S Barocas, A D’Amour, and K Lum Algorithmic fairness: Choices, assumptions, and definitions. Annual Review of Statistics and Its Application, 8(1):141– 163, 2021. doi: 101146/annurev-statistics-042720-125902 URL https://doiorg/101146/ annurev-statistics-042720-125902. C. Molnar Interpretable Machine Learning , 2020 M. R Morris, A Zolyomi, C Yao, S Bahram, J P Bigham, and S K Kane ” with most of it being pictures now, i rarely use it” understanding twitter’s evolving accessibility to blind users. In Proceedings of the 2016 CHI conference on human factors in computing systems, pages 5506–5516, 2016. L. S Moss Natural
logic and semantics In M Aloni, H Bastiaanse, T de Jager, P van Ormondt, and K. Schulz, editors, Proceedings of the 18th Amsterdam Colloquium: Revised Selected Papers, pages 71–80, Berlin, 2009. University of Amsterdam, Springer BIBLIOGRAPHY 182 S. Murty, P Sharma, J Andreas, and C D Manning Characterizing intrinsic compositionality in transformers with tree projections, 2023. A. Naik, A Ravichander, N Sadeh, C Rose, and G Neubig Stress test evaluation for natural language inference. In Proceedings of the 27th International Conference on Computational Linguistics, pages 2340–2353, Santa Fe, New Mexico, USA, Aug. 2018 Association for Computational Linguistics URL https://www.aclweborg/anthology/C18-1198 N. Nanda A longlist of theories of impact for interpretability, 2022 Y. Nie, Y Wang, and M Bansal Analyzing compositionality-sensitivity of NLI models In Proceedings of the AAAI Conference on Artificial Intelligence, volume 33, pages 6867–6874, 2019a. Y. Nie, A Williams, E
Dinan, M Bansal, J Weston, and D Kiela Adversarial NLI: A new benchmark for natural language understanding, 2019b. C. Olah, N Cammarata, L Schubert, G Goh, M Petrov, and S Carter Zoom in: An introduction to circuits. Distill, 2020 doi: 1023915/distill00024001 https://distillpub/2020/circuits/zoom-in C. Olsson, N Elhage, N Nanda, N Joseph, N DasSarma, T Henighan, B Mann, A Askell, Y Bai, A. Chen, T Conerly, D Drain, D Ganguli, Z Hatfield-Dodds, D Hernandez, S Johnston, A Jones, J. Kernion, L Lovitt, K Ndousse, D Amodei, T Brown, J Clark, J Kaplan, S McCandlish, and C. Olah In-context learning and induction heads Transformer Circuits Thread, 2022 https://transformer-circuits.pub/2022/in-context-learning-and-induction-heads/indexhtml J. Otsuka and H Saigo On the equivalence of causal models: A category-theoretic approach CoRR, abs/2201.06981, 2022 URL https://arxivorg/abs/220106981 C. Otte Safe and interpretable machine learning: A methodological review In Computational Intelligence in
Intelligent Data Analysis, 2013. URL https://linkspringercom/chapter/10 1007/978-3-642-32378-2 8. L. Ouyang, J Wu, X Jiang, D Almeida, C L Wainwright, P Mishkin, C Zhang, S Agarwal, K. Slama, A Ray, J Schulman, J Hilton, F Kelton, L Miller, M Simens, A Askell, P Welinder, P. Christiano, J Leike, and R Lowe Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 36, 2022 A. Pan, K Bhatia, and J Steinhardt The effects of reward misspecification: Mapping and mitigating misaligned models. In International Conference on Learning Representations, 2022 URL https://openreview.net/forum?id=JYtwGwIL7ye E. Pavlick and C Callison-Burch Most “babies” are “little” and most “problems” are “huge”: Compositional entailment in adjective-nouns. In Proceedings of the 54th Annual Meeting of the BIBLIOGRAPHY 183 Association for Computational Linguistics (Volume 1: Long Papers), pages 2164–2173, Berlin, Germany, Aug. 2016
Association for Computational Linguistics doi: 1018653/v1/P16-1204 URL https://www.aclweborg/anthology/P16-1204 J. Pearl Probabilities of causation: Three counterfactual interpretations and their identification Synthese, 121(1):93–149, 1999. J. Pearl Direct and indirect effects In Proceedings of the Seventeenth Conference on Uncertainty in Artificial Intelligence, UAI’01, page 411–420, San Francisco, CA, USA, 2001. Morgan Kaufmann Publishers Inc. ISBN 1558608001 J. Pearl Causality Cambridge University Press, 2009 J. Pearl The limitations of opaque learning machines Possible minds: twenty-five ways of looking at AI, pages 13–19, 2019a. J. Pearl The limitations of opaque learning machines Possible Minds: Twenty-Five Ways of Looking at AI, 2019b. URL https://ftpcsuclaedu/pub/stat ser/r489pdf J. Pennington, R Socher, and C Manning GloVe: Global vectors for word representation In Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages
1532–1543, Doha, Qatar, Oct. 2014 Association for Computational Linguistics doi: 10 3115/v1/D14-1162. URL https://wwwaclweborg/anthology/D14-1162 L. Perez and J Wang The effectiveness of data augmentation in image classification using deep learning. CoRR, abs/171204621, 2017 URL http://arxivorg/abs/171204621 D. Pessach and E Shmueli Algorithmic fairness CoRR, abs/200109784, 2020 URL https: //arxiv.org/abs/200109784 M. Peters, M Neumann, L Zettlemoyer, and W-t Yih Dissecting contextual word embeddings: Architecture and representation. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1499–1509, Brussels, Belgium, Oct.-Nov 2018 Association for Computational Linguistics. doi: 1018653/v1/D18-1179 URL https://wwwaclweborg/ anthology/D18-1179. T. Pimentel, J Valvoda, R H Maudslay, R Zmigrod, A Williams, and R Cotterell Informationtheoretic probing for linguistic structure In D Jurafsky, J Chai, N Schluter, and J R Tetreault, editors,
Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, ACL 2020, Online, July 5-10, 2020, pages 4609–4622. Association for Computational Linguistics, 2020. doi: 1018653/v1/2020acl-main420 URL https://doiorg/1018653/v1/2020acl-main 420. BIBLIOGRAPHY 184 D. Premack The codes of man and beasts Behavioral and Brain Sciences, 6(1):125–136, 1983 R. Pryzant, D Card, D Jurafsky, V Veitch, and D Sridhar Causal effects of linguistic properties In NAACL, 2021a. R. Pryzant, D Card, D Jurafsky, V Veitch, and D Sridhar Causal effects of linguistic properties In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4095–4109, Online, June 2021b. Association for Computational Linguistics. doi: 1018653/v1/2021naacl-main323 URL https: //aclanthology.org/2021naacl-main323 R. Pryzant, D Card, D Jurafsky, V Veitch, and D Sridhar Causal effects of linguistic properties
arXiv:2010.12919 [cs], Apr 2021c URL http://arxivorg/abs/201012919 arXiv: 201012919 H. Putnam Minds and machines In S Hook, editor, Dimensions of Minds, pages 138–164 New York, USA: New York University Press, 1960. A. Radford, J Wu, R Child, D Luan, D Amodei, and I Sutskever are unsupervised multitask learners. OpenAI blog, 2019. Language models URL https://cdn.openaicom/ better-language-models/language models are unsupervised multitask learners.pdf A. Radford, J W Kim, C Hallacy, A Ramesh, G Goh, S Agarwal, G Sastry, A Askell, P Mishkin, J. Clark, G Krueger, and I Sutskever Learning transferable visual models from natural language supervision. In M Meila and T Zhang, editors, Proceedings of the 38th International Conference on Machine Learning, ICML 2021, 18-24 July 2021, Virtual Event, volume 139 of Proceedings of Machine Learning Research, pages 8748–8763. PMLR, 2021 URL http://proceedingsmlr press/v139/radford21a.html N. F Rajani, B McCann, C Xiong, and R Socher Explain
yourself! leveraging language models for commonsense reasoning. In A Korhonen, D R Traum, and L Màrquez, editors, Proceedings of the 57th Conference of the Association for Computational Linguistics, ACL 2019, Florence, Italy, July 28- August 2, 2019, Volume 1: Long Papers, pages 4932–4942. Association for Computational Linguistics, 2019. doi: 1018653/v1/p19-1487 URL https://doiorg/1018653/v1/p19-1487 P. Rajpurkar, J Zhang, K Lopyrev, and P Liang SQuAD: 100,000+ questions for machine comprehension of text. In Proceedings of the 2016 Conference on Empirical Methods in Natural Language Processing, pages 2383–2392, Austin, Texas, Nov. 2016 Association for Computational Linguistics. doi: 1018653/v1/D16-1264 URL https://wwwaclweborg/anthology/D16-1264 S. Ravfogel, Y Elazar, H Gonen, M Twiton, and Y Goldberg Null it out: Guarding protected attributes by iterative nullspace projection. In D Jurafsky, J Chai, N Schluter, and J R Tetreault, editors, Proceedings of the 58th Annual Meeting
of the Association for Computational BIBLIOGRAPHY 185 Linguistics, ACL 2020, Online, July 5-10, 2020, pages 7237–7256. Association for Computational Linguistics, 2020a. doi: 1018653/v1/2020acl-main647 URL https://doiorg/1018653/v1/ 2020.acl-main647 S. Ravfogel, Y Elazar, H Gonen, M Twiton, and Y Goldberg Null it out: Guarding protected attributes by iterative nullspace projection. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 7237–7256, Online, July 2020b. Association for Computational Linguistics. doi: 1018653/v1/2020acl-main647 URL https://wwwaclweborg/ anthology/2020.acl-main647 S. Ravfogel, Y Elazar, H Gonen, M Twiton, and Y Goldberg Null it out: Guarding protected attributes by iterative nullspace projection. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 2020c URL https://aclanthologyorg/2020acl-main 647. A. Ravichander, Y Belinkov, and E Hovy Probing the probing paradigm:
Does probing accuracy entail task relevance?, 2020. M. T Ribeiro, S Singh, and C Guestrin ”why should i trust you?”: Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD ’16, page 1135–1144, New York, NY, USA, 2016a. Association for Computing Machinery. ISBN 9781450342322 doi: 101145/29396722939778 URL https: //doi.org/101145/29396722939778 M. T Ribeiro, S Singh, and C Guestrin ”why should i trust you?”: Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, KDD ’16, page 1135–1144, New York, NY, USA, 2016b. Association for Computing Machinery. ISBN 9781450342322 doi: 101145/29396722939778 URL https: //doi.org/101145/29396722939778 M. T Ribeiro, S Singh, and C Guestrin ”Why Should I Trust You?”: Explaining the predictions of any classifier. In Proceedings of the 22nd ACM
SIGKDD International Conference on Knowledge Discovery and Data Mining, San Francisco California USA, 2016c. URL https://dlacmorg/ doi/10.1145/29396722939778 K. Richardson, H Hu, L S Moss, and A Sabharwal Probing natural language inference models through semantic fragments, 2019. E. F Rischel and S Weichwald Compositional abstraction error and a category of causal models In Proceedings of the 37th Conference on Uncertainty in Artificial Intelligence (UAI), 2021. BIBLIOGRAPHY 186 A. Rogers, O Kovaleva, and A Rumshisky A primer in BERTology: What we know about how BERT works. Transactions of the Association for Computational Linguistics, 8:842–866, 2020 doi: 10.1162/tacl a 00349 URL https://aclanthologyorg/2020tacl-154 G. Rotman, A Feder, and R Reichart Model compression for domain adaptation through causal effect estimation. arXiv:210107086, 2021 P. K Rubenstein, S Weichwald, S Bongers, J M Mooij, D Janzing, M Grosse-Wentrup, and B. Schölkopf Causal consistency of structural
equation models In Proceedings of the 33rd Conference on Uncertainty in Artificial Intelligence (UAI). Association for Uncertainty in Artificial Intelligence (AUAI), Aug. 2017a URL http://auaiorg/uai2017/proceedings/papers/11pdf *equal contribution. P. K Rubenstein, S Weichwald, S Bongers, J M Mooij, D Janzing, M Grosse-Wentrup, and B. Schölkopf Causal consistency of structural equation models In Proceedings of the 33rd Conference on Uncertainty in Artificial Intelligence (UAI), 2017b. S. Ruder An overview of multi-task learning in deep neural networks CoRR, abs/170605098, 2017 URL http://arxiv.org/abs/170605098 L. Ruis, J Andreas, M Baroni, D Bouchacourt, and B M Lake A benchmark for systematic generalization in grounded language understanding. Advances in Neural Information Processing Systems, 33, 2020. D. E Rumelhart, J L McClelland, and PDP Research Group, editors Parallel Distributed Processing Volume 1: Foundations. MIT Press, Cambridge, MA, 1986 O. Russakovsky, J Deng, H Su, J
Krause, S Satheesh, S Ma, Z Huang, A Karpathy, A. Khosla, M Bernstein, A C Berg, and L Fei-Fei ImageNet Large Scale Visual Recognition Challenge International Journal of Computer Vision (IJCV), 115(3):211–252, 2015 doi: 10.1007/s11263-015-0816-y V. Sánchez-Valencia Studies in Natural Logic and Categorial Grammar PhD thesis, University of Amsterdam, 1991. V. Sanh, L Debut, J Chaumond, and T Wolf DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter. arXiv:191001108, 2019 N. Saphra and A Lopez Understanding learning dynamics of language models with SVCCA In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 3257–3267, Minneapolis, Minnesota, June 2019. Association for Computational Linguistics doi: 10.18653/v1/N19-1329 URL https://wwwaclweborg/anthology/N19-1329 BIBLIOGRAPHY 187 M. Schuster and K K Paliwal Bidirectional
recurrent neural networks IEEE transactions on Signal Processing, 45(11):2673–2681, 1997. G. G Scott, A Keitel, M Becirspahic, B Yao, and S C Sereno The glasgow norms: Ratings of 5,500 words on nine scales. Behavior research methods, 51:1258–1270, 2019 F. Shi, X Chen, K Misra, N Scales, D Dohan, E Chi, N Schärli, and D Zhou Large language models can be easily distracted by irrelevant context, 2023. C. Shorten and T M Khoshgoftaar A survey on image data augmentation for deep learning Journal of Big Data, 6:1–48, 2019. A. Shrikumar, P Greenside, A Shcherbina, and A Kundaje Not just a black box: Learning important features through propagating activation differences. CoRR, abs/160501713, 2016 URL http://arxiv.org/abs/160501713 A. Shrikumar, P Greenside, and A Kundaje Learning important features through propagating activation differences. In Proceedings of the 34th International Conference on Machine LearningVolume 70, pages 3145–3153 JMLR org, 2017 K. Simonyan, A Vedaldi, and A
Zisserman Deep inside convolutional networks: visualising image classification models and saliency maps. In Proceedings of the International Conference on Learning Representations (ICLR). ICLR, 2014 P. Smolensky Neural and conceptual interpretation of PDP models In Parallel Distributed Processing: Explorations in the Microstructure, Vol. 2: Psychological and Biological Models, page 390–431 MIT Press, Cambridge, MA, USA, 1986. ISBN 0262631105 R. Socher, A Perelygin, J Wu, J Chuang, C D Manning, A Y Ng, and C Potts Recursive deep models for semantic compositionality over a sentiment treebank. In Proceedings of the 2013 conference on empirical methods in natural language processing, pages 1631–1642, 2013. P. Soulos, R T McCoy, T Linzen, and P Smolensky Discovering the compositional structure of vector representations with role learning networks. In A Alishahi, Y Belinkov, G Chrupala, D. Hupkes, Y Pinter, and H Sajjad, editors, Proceedings of the Third BlackboxNLP Workshop on Analyzing
and Interpreting Neural Networks for NLP, BlackboxNLP@EMNLP 2020, Online, November 2020, pages 238–254. Association for Computational Linguistics, 2020a doi: 1018653/ v1/2020.blackboxnlp-123 URL https://doiorg/1018653/v1/2020blackboxnlp-123 P. Soulos, R T McCoy, T Linzen, and P Smolensky Discovering the compositional structure of vector representations with role learning networks. In Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pages 238–254, Online, Nov. BIBLIOGRAPHY 188 2020b. Association for Computational Linguistics doi: 1018653/v1/2020blackboxnlp-123 URL https://www.aclweborg/anthology/2020blackboxnlp-123 P. Spirtes, C Glymour, and R Scheines Causation, Prediction, and Search MIT Press, 2000 P. Spirtes, C N Glymour, and R Scheines Causation, Prediction, and Search MIT Press, 2nd edition, 2001. J. Springenberg, A Dosovitskiy, T Brox, and M Riedmiller Striving for simplicity: The all convolutional net. CoRR, 12 2014 N.
Stiennon, L Ouyang, J Wu, D M Ziegler, R Lowe, C Voss, A Radford, D Amodei, and P F Christiano. Learning to summarize from human feedback CoRR, abs/200901325, 2020 URL https://arxiv.org/abs/200901325 E. Sullivan Understanding from machine learning models The British Journal for the Philosophy of Science, 73(1):109–133, 2022. doi: 101093/bjps/axz035 URL https://doiorg/101093/bjps/ axz035. S. Sun, Y Cheng, Z Gan, and J Liu Patient knowledge distillation for BERT model compression In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pages 4323–4332, Hong Kong, China, Nov. 2019 Association for Computational Linguistics doi: 10.18653/v1/D19-1441 URL https://wwwaclweborg/anthology/D19-1441 M. Sundararajan, A Taly, and Q Yan Axiomatic attribution for deep networks In D Precup and Y. W Teh, editors, Proceedings of the 34th International Conference on Machine
Learning, volume 70 of Proceedings of Machine Learning Research, pages 3319–3328, International Convention Centre, Sydney, Australia, 06–11 Aug 2017a. PMLR URL http://proceedingsmlrpress/ v70/sundararajan17a.html M. Sundararajan, A Taly, and Q Yan Axiomatic attribution for deep networks In Proceedings of the 34th International Conference on Machine Learning-Volume 70, pages 3319–3328. JMLR org, 2017b. A. Talmor, Y Elazar, Y Goldberg, and J Berant olmpics – on what language model pre-training captures, 2019. R. Taori, I Gulrajani, T Zhang, Y Dubois, X Li, C Guestrin, P Liang, and T B Hashimoto Stanford alpaca: An instruction-following llama model https://githubcom/tatsu-lab/stanford alpaca, 2023. BIBLIOGRAPHY 189 I. Tenney, D Das, and E Pavlick BERT rediscovers the classical NLP pipeline In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4593–4601, Florence, Italy, July 2019. Association for Computational Linguistics doi:
1018653/v1/P19-1452 URL https://www.aclweborg/anthology/P19-1452 B. Thomee, D A Shamma, G Friedland, B Elizalde, K Ni, D Poland, D Borth, and L-J Li Yfcc100m: The new data in multimedia research. Communications of the ACM, 59(2):64–73, 2016 R. K R Thompson, D L Oden, and S T Boysen Language-naive chimpanzees (pan troglodytes) judge relations between relations in a conceptual matching-to-sample task. Journal of Experimental Psychology: Animal Behavior Processes, 23(1):31-43, 1997. E. F Tjong Kim Sang and F De Meulder Introduction to the CoNLL-2003 shared task: Languageindependent named entity recognition In Proceedings of the Seventh Conference on Natural Language Learning at HLT-NAACL 2003, pages 142–147, 2003. URL https://wwwaclweborg/ anthology/W03-0419. H. Touvron, T Lavril, G Izacard, X Martinet, M-A Lachaux, T Lacroix, B Rozière, N Goyal, E. Hambro, F Azhar, et al Llama: Open and efficient foundation language models arXiv preprint arXiv:2302.13971, 2023 J. Uesato, N Kushman,
R Kumar, F Song, N Siegel, L Wang, A Creswell, G Irving, and I Higgins Solving math word problems with process- and outcome-based feedback, 2022. S. Upadhyay, S Joshi, and H Lakkaraju Towards robust and reliable algorithmic recourse In M. Ranzato, A Beygelzimer, Y Dauphin, P Liang, and J W Vaughan, editors, Advances in Neural Information Processing Systems, volume 34, pages 16926–16937 Curran Associates, Inc, 2021 URL https://proceedingsneuripscc/paper files/paper/2021/file/ 8ccfb1140664a5fa63177fb6e07352f0-Paper.pdf J. van Benthem A brief history of natural logic In M Chakraborty, B Löwe, M Nath Mitra, and S. Sarukki, editors, Logic, Navya-Nyaya and Applications: Homage to Bimal Matilal, 2008a J. van Benthem A brief history of natural logic In M Chakraborty, B Löwe, M Nath Mitra, and S. Sarukki, editors, Logic, Navya-Nyaya and Applications: Homage to Bimal Matilal, 2008b A. Vaswani, N Shazeer, N Parmar, J Uszkoreit, L Jones, A N Gomez, L Kaiser, and I Polosukhin Attention is all
you need. In Advances in Neural Information Processing Systems, pages 5998–6008, 2017a. A. Vaswani, N Shazeer, N Parmar, J Uszkoreit, L Jones, A N Gomez, L u Kaiser, and I. Polosukhin Attention is all you need In I Guyon, U V Luxburg, S Bengio, H Wallach, R. Fergus, S Vishwanathan, and R Garnett, editors, Advances in Neural Information Processing BIBLIOGRAPHY 190 Systems 30, pages 5998–6008. Curran Associates, Inc, 2017b URL http://papersnipscc/ paper/7181-attention-is-all-you-need.pdf S. Venkatasubramanian and M Alfano The philosophical basis of algorithmic recourse In Proceedings of the 2020 Conference on Fairness, Accountability, and Transparency, FAT* ’20, page 284–293, New York, NY, USA, 2020. Association for Computing Machinery ISBN 9781450369367 doi: 10.1145/33510953372876 URL https://doiorg/101145/33510953372876 S. Verma, J Dickerson, and K Hines Counterfactual explanations for machine learning: A review , 2020. URL https://arxivorg/abs/201010596 J. Vig, S Gehrmann,
Y Belinkov, S Qian, D Nevo, Y Singer, and S Shieber Investigating gender bias in language models using causal mediation analysis. In H Larochelle, M Ranzato, R Hadsell, M. F Balcan, and H Lin, editors, Advances in Neural Information Processing Systems, volume 33, pages 12388–12401. Curran Associates, Inc, 2020a URL https://proceedingsneuripscc/ paper/2020/file/92650b2e92217715fe312e6fa7b90d82-Paper.pdf J. Vig, S Gehrmann, Y Belinkov, S Qian, D Nevo, Y Singer, and S Shieber Causal mediation analysis for interpreting neural nlp: The case of gender bias, 2020b. E. Voita, D Talbot, F Moiseev, R Sennrich, and I Titov Analyzing multi-head self-attention: Specialized heads do the heavy lifting, the rest can be pruned. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 5797–5808, Florence, Italy, July 2019. Association for Computational Linguistics doi: 1018653/v1/P19-1580 URL https://www.aclweborg/anthology/P19-1580 A. Wang, A Singh, J
Michael, F Hill, O Levy, and S Bowman GLUE: A multi-task benchmark and analysis platform for natural language understanding. In Proceedings of the 2018 EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 353–355, 2018. B. Wang and A Komatsuzaki GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model https://github.com/kingoflolz/mesh-transformer-jax, May 2021 B. Wang, S Min, X Deng, J Shen, Y Wu, L Zettlemoyer, and H Sun Towards understanding chain-of-thought prompting: An empirical study of what matters. arXiv preprint arXiv:221210001, 2022a. K. Wang, A Variengien, A Conmy, B Shlegeris, and J Steinhardt Interpretability in the wild: a circuit for indirect object identification in GPT-2 small. arXiv preprint arXiv:221100593, 2022b URL https://arxiv.org/abs/221100593 BIBLIOGRAPHY 191 Y. Wang, Y Kordi, S Mishra, A Liu, N A Smith, D Khashabi, and H Hajishirzi Self-instruct: Aligning language model with self generated instructions. arXiv
preprint arXiv:221210560, 2022c J. Wei, M Bosma, V Zhao, K Guu, A W Yu, B Lester, N Du, A M Dai, and Q V Le Finetuned language models are zero-shot learners. In International Conference on Learning Representations, 2022. S. Wiegreffe and A Marasovic Teach me to explain: A review of datasets for explainable natural language processing. In Thirty-fifth Conference on Neural Information Processing Systems Datasets and Benchmarks Track (Round 1), 2021. URL https://openreviewnet/forum?id=ogNcxJn32BZ S. Wiegreffe and Y Pinter Attention is not not explanation In K Inui, J Jiang, V Ng, and X Wan, editors, Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing, EMNLP-IJCNLP 2019, Hong Kong, China, November 3-7, 2019, pages 11–20. Association for Computational Linguistics, 2019. doi: 1018653/v1/D19-1002 URL https://doiorg/1018653/v1/D19-1002 S. Wiegreffe, A Marasović, and N A Smith
Measuring association between labels and freetext rationales In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 10266–10284, Online and Punta Cana, Dominican Republic, Nov. 2021. Association for Computational Linguistics doi: 1018653/v1/2021emnlp-main804 URL https://aclanthology.org/2021emnlp-main804 A. Williams, N Nangia, and S Bowman A broad-coverage challenge corpus for sentence understanding through inference In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 1112–1122. Association for Computational Linguistics, 2018a URL http://aclweb.org/anthology/N18-1101 A. Williams, N Nangia, and S Bowman A broad-coverage challenge corpus for sentence understanding through inference In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language
Technologies, Volume 1 (Long Papers), pages 1112–1122. Association for Computational Linguistics, 2018b URL http://aclweb.org/anthology/N18-1101 T. Wolf, L Debut, V Sanh, J Chaumond, C Delangue, A Moi, P Cistac, T Rault, R Louf, M Funtowicz, and J Brew Huggingface’s transformers: State-of-the-art natural language processing ArXiv, abs/1910.03771, 2019 T. Wolf, L Debut, V Sanh, J Chaumond, C Delangue, A Moi, P Cistac, T Rault, R Louf, M. Funtowicz, J Davison, S Shleifer, P von Platen, C Ma, Y Jernite, J Plu, C Xu, T Le Scao, S. Gugger, M Drame, Q Lhoest, and A Rush Transformers: State-of-the-art natural language BIBLIOGRAPHY 192 processing. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 38–45, Online, Oct. 2020 Association for Computational Linguistics. doi: 1018653/v1/2020emnlp-demos6 URL https://wwwaclweborg/anthology/ 2020.emnlp-demos6 J. Woodward What is a mechanism? a counterfactual account
Philosophy of Science, 69(S3): S366–S377, 2002. J. Woodward Making Things Happen: A Theory of Causal Explanation Oxford university press, 2003. J. Woodward Sensitive and insensitive causation The Philosophical Review, 115(1):1–50, 2006 J. Woodward Explanatory autonomy: the role of proportionality, stability, and conditional irrelevance Synthese, 198:237–265, 2021. T. Wu, M T Ribeiro, J Heer, and D Weld Polyjuice: Generating counterfactuals for explaining, evaluating, and improving models. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing, 2021a. T. Wu, T Maruyama, and J Leskovec Learning to accelerate partial differential equations via latent global evolution. Advances in Neural Information Processing Systems, 36, 2022a Z. Wu, E Kreiss, D C Ong, and C Potts ReaSCAN: Compositional reasoning in language grounding. NeurIPS 2021 Datasets and Benchmarks Track, 2021b URL
https://arxivorg/ abs/2109.08994 Z. Wu, K D’Oosterlinck, A Geiger, A Zur, and C Potts Causal Proxy Models for concept-based model explanations. arXiv:220914279, 2022b URL https://arxivorg/abs/220914279 Z. Wu, A Geiger, J Rozner, E Kreiss, H Lu, T Icard, C Potts, and N Goodman Causal distillation for language models. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4288–4295, 2022c. Z. Wu, A Geiger, J Rozner, E Kreiss, H Lu, T Icard, C Potts, and N D Goodman Causal distillation for language models. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4288–4295, Seattle, United States, July 2022d. Association for Computational Linguistics doi: 10.18653/v1/2022naacl-main318 URL https://aclanthologyorg/2022naacl-main318 BIBLIOGRAPHY 193 Z. Wu, A Geiger, J Rozner, E Kreiss, H Lu,
T Icard, C Potts, and N D Goodman Causal distillation for language models. In Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4288–4295, Seattle, United States, July 2022e. Association for Computational Linguistics doi: 10.18653/v1/2022naacl-main318 URL https://aclanthologyorg/2022naacl-main318 Z. Wu, K D’Oosterlinck, A Geiger, A Zur, and C Potts Causal proxy models for concept-based model explanations. In International Conference on Machine Learning, pages 37313–37334 PMLR, 2023a. Z. Wu, A Geiger, C Potts, and N D Goodman Interpretability at scale: Identifying causal mechanisms in Alpaca. Ms, Stanford University, 2023b URL https://arxivorg/abs/2305 08809. Y. Xu, S Zhao, J Song, R Stewart, and S Ermon A theory of usable information under computational constraints. In 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, April 26-30, 2020.
OpenReviewnet, 2020 URL https://openreview net/forum?id=r1eBeyHFDH. S. Yablo Mental causation Philosophical Review, 101:245–280, 1992 H. Yanaka, K Mineshima, D Bekki, K Inui, S Sekine, L Abzianidze, and J Bos HELP: A dataset for identifying shortcomings of neural models in monotonicity reasoning. In Proceedings of the Eighth Joint Conference on Lexical and Computational Semantics (*SEM 2019), pages 250–255, Minneapolis, Minnesota, June 2019a. Association for Computational Linguistics doi: 10.18653/v1/S19-1027 URL https://wwwaclweborg/anthology/S19-1027 H. Yanaka, K Mineshima, D Bekki, K Inui, S Sekine, L Abzianidze, and J Bos Can neural networks understand monotonicity reasoning? In Proceedings of the 2019 ACL Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, pages 31–40, Florence, Italy, Aug. 2019b Association for Computational Linguistics. doi: 1018653/v1/W19-4804 URL https://www aclweb.org/anthology/W19-4804 H. Yanaka, K Mineshima, D Bekki, and K Inui
Do neural models learn systematicity of monotonicity inference in natural language?, 2020. C.-K Yeh, B Kim, S Arik, C-L Li, T Pfister, and P Ravikumar aware concept-based explanations in deep neural networks. tion Processing Systems, 2020. On completeness- Advances in Neural Informa- URL https://proceedings.neuripscc/paper/2020/file/ ecb287ff763c169694f682af52c1f309-Paper.pdf BIBLIOGRAPHY 194 M. B Zafar, I Valera, M Gomez-Rodriguez, and K P Gummadi Fairness beyond disparate treatment & disparate impact: Learning classification without disparate mistreatment. In R Barrett, R. Cummings, E Agichtein, and E Gabrilovich, editors, Proceedings of the 26th International Conference on World Wide Web, WWW 2017, Perth, Australia, April 3-7, 2017, pages 1171–1180. ACM, 2017a. doi: 101145/30389123052660 URL https://doiorg/101145/30389123052660 M. B Zafar, I Valera, M Gomez-Rodriguez, and K P Gummadi Fairness constraints: Mechanisms for fair classification. In A Singh and X J Zhu,
editors, Proceedings of the 20th International Conference on Artificial Intelligence and Statistics, AISTATS 2017, 20-22 April 2017, Fort Lauderdale, FL, USA, volume 54 of Proceedings of Machine Learning Research, pages 962–970. PMLR, 2017b. URL http://proceedingsmlrpress/v54/zafar17ahtml M. D Zeiler and R Fergus Visualizing and understanding convolutional networks In D Fleet, T. Pajdla, B Schiele, and T Tuytelaars, editors, Computer Vision – ECCV 2014, pages 818–833, Cham, 2014a. Springer International Publishing ISBN 978-3-319-10590-1 M. D Zeiler and R Fergus Visualizing and understanding convolutional networks In European conference on computer vision, pages 818–833. Springer, 2014b R. S Zemel, Y Wu, K Swersky, T Pitassi, and C Dwork Learning fair representations In Proceedings of the 30th International Conference on Machine Learning, ICML 2013, Atlanta, GA, USA, 16-21 June 2013, volume 28 of JMLR Workshop and Conference Proceedings, pages 325–333. JMLR.org, 2013 URL
http://proceedingsmlrpress/v28/zemel13html C. Zhang, M Raghu, J M Kleinberg, and S Bengio Pointer value retrieval: A new benchmark for understanding the limits of neural network generalization. CoRR, abs/210712580, 2021 URL https://arxiv.org/abs/210712580 Y. Zhang and Q Yang A survey on multi-task learning CoRR, abs/170708114, 2017 URL http://arxiv.org/abs/170708114 A. Zur, E Kreiss, and C P A Geiger When interpretability enhances accessibility: Updating clip to prefer descriptions over captions. 2023