Search references for MECHANISTIC INTERPRETABILITY. Phrases containing MECHANISTIC INTERPRETABILITY
See searches and references containing MECHANISTIC INTERPRETABILITY!MECHANISTIC INTERPRETABILITY
Reverse-engineering neural networks
Mechanistic interpretability (sometimes abbreviated as mech interp, mechinterp, or MI) is a subfield of research within explainable artificial intelligence
Mechanistic_interpretability
Canadian machine learning researcher (born 1991 or 1992)
He is known for his work on neural network interpretability, particularly mechanistic interpretability, and for research and tools that visualise internal
Chris_Olah
American artificial intelligence company
for $1.5 billion. Anthropic conducts LLM research including on mechanistic interpretability, safety, alignment, and societal impact. Notable researchers
Anthropic
Type of machine learning model
should be viewed as models of the human brain and/or human mind. Mechanistic interpretability is a subfield of research that aims to understand neural networks'
Large_language_model
AI whose outputs can be understood by humans
a goal referred to as "local interpretability". There is also research on whether the concepts of local interpretability can be applied to a remote context
Explainable artificial intelligence
Explainable_artificial_intelligence
Term used in machine learning
challenging the "parrot" characterization. Anthropic conducted mechanistic interpretability research on Claude, using attribution graphs to identify circuits
Stochastic_parrot
Deep learning architecture
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Mamba (deep learning architecture)
Mamba_(deep_learning_architecture)
Phenomenon in neural networks
straightforwardly interpreted, polysemanticity is a central obstacle in mechanistic interpretability. Mechanistic interpretability often begins from the
Polysemanticity
Type of large language model
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Generative pre-trained transformer
Generative_pre-trained_transformer
Interpretable computational sub-graphs within artificial neural networks
study of artificial circuits is a primary focus of the field of mechanistic interpretability. Researchers aim to reverse-engineer "black box" deep learning
Circuit_(neural_network)
Machine learning calibration technique
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Platt_scaling
Large language model and AI chatbot by Anthropic
released on September 22, 2026. In May 2024, Anthropic issued a mechanistic interpretability paper identifying "features" (internal representations of concepts)
Claude_(AI)
Machine learning technique
approaches often enable tighter alignment with human values, improved interpretability, and simpler training pipelines compared to RLHF. Direct preference
Reinforcement learning from human feedback
Reinforcement_learning_from_human_feedback
Type of convolutional neural network
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
U-Net
2018 text-generating language model
languages (such as Swahili or Haitian Creole) are difficult to translate and interpret using such models due to a lack of available text for corpus-building
GPT-1
Topics referred to by the same term
used in the IBM System/38's architecture Management information Mechanistic interpretability, a subfield of explainable AI Mi (prefix symbol), the IEEE prefix
MI
Similarity measure for number sequences
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Cosine_similarity
Machine-learning and computational-neuroscience conference
to evaluate randomness in the reviewing process. Several researchers interpreted the result. Regarding whether the decision in NIPS is completely random
Conference on Neural Information Processing Systems
Conference_on_Neural_Information_Processing_Systems
2023 text-generating language model
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
IBM_Granite
2023 text-generating language model
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
GPT-4
Technique for the generative modeling of a continuous probability distribution
implemented as a neural network. "score", because the output of the network is interpreted as approximating the score function ∇ ln ρ t {\displaystyle \nabla
Diffusion_model
Method used to normalize the range of independent variables
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Feature_scaling
Model-free reinforcement learning algorithm
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Proximal_policy_optimization
Machine learning technique
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Mixture_of_experts
Academic conference in machine learning
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
International Conference on Learning Representations
International_Conference_on_Learning_Representations
AI platform developed by IBM
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
IBM_Watsonx
Machine learning methods using multiple input modalities
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Multimodal_learning
Smooth approximation of one-hot arg max
{\displaystyle (0,1)} , and the components will add up to 1, so that they can be interpreted as probabilities. Furthermore, the larger input components will correspond
Softmax_function
Hypothesized risk to human existence
makes the best decisions to achieve its goals. The field of mechanistic interpretability aims to better understand the inner workings of AI models, potentially
Existential risk from artificial intelligence
Existential_risk_from_artificial_intelligence
Problem setup in machine learning
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Zero-shot_learning
Concept in machine learning
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Leakage_(machine_learning)
Type of activation function
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Rectified_linear_unit
Machine learning technique
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Transfer_learning
Conversational software
benefit of the doubt when conversational responses are capable of being interpreted as "intelligent". Following ELIZA, psychiatrist Kenneth Colby developed
Chatbot
2025 multimodal model by OpenAI
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
GPT-5
Machine learning technique
gradients with respect to the class [CLS] token. Some class-sensitive interpretability methods originally developed for convolutional neural networks can
Attention_(machine_learning)
Software program
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
DeepDream
2020 text-generating language model
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
GPT-3
Framework for mathematical analysis of machine learning
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Probably approximately correct learning
Probably_approximately_correct_learning
Type of feedforward neural network
convolution layer. Every entry in the output volume can thus also be interpreted as an output of a neuron that looks at a small region in the input. Each
Convolutional_neural_network
Software user interface
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Human-in-the-loop
Type of database that uses vectors to represent other data
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Vector_database
Automatic creation of ontologies
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Ontology_learning
Class of algorithms for pattern analysis
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Kernel_method
Recurrent neural network architecture
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Long_short-term_memory
Academic conference in machine learning
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
International Conference on Machine Learning
International_Conference_on_Machine_Learning
Tree-based ensemble machine learning methods
intrinsic interpretability of decision trees. Decision trees are among a fairly small family of machine learning models that are easily interpretable along
Random_forest
Optimization algorithm
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Gradient_descent
Reinforcement learning technique
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Self-play
Statistical model of language
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Language_model
Neural network technology
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Convolutional_layer
Class of artificial neural network
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Recurrent_neural_network
Machine learning technique
decision tree or linear regression, it sacrifices intelligibility and interpretability. For example, following the path that a decision tree takes to make
Gradient_boosting
Technique in machine learning
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Curriculum_learning
Type of feedforward neural network
derived from inputs, limiting their applicability in domains where interpretability is required. Cybenko, G. 1989. Approximation by superpositions of a
Multilayer_perceptron
Machine learning-powered structure design
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Neural_architecture_search
Plot of machine learning model performance over time or experience
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Learning curve (machine learning)
Learning_curve_(machine_learning)
Method in natural language processing
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Word_embedding
AI that learns decision rules from data
prediction model usually known as decision algorithm. Rules can also be interpreted in various ways depending on the domain knowledge, data types(discrete
Rule-based_machine_learning
Statistical method
do not affect the values of the estimates. Therefore, it also can be interpreted as an outlier detection method. It is a non-deterministic algorithm in
Random_sample_consensus
Machine learning algorithm
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
State–action–reward–state–action
State–action–reward–state–action
Numerical method that reduces the complexity of computationally intensive simulations
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Proper orthogonal decomposition
Proper_orthogonal_decomposition
Machine learning model training problem
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Vanishing_gradient_problem
3D reconstruction technique
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Neural_radiance_field
Machine learning framework
Applications of bidirectional models include semi-supervised learning, interpretable machine learning, and neural machine translation. CycleGAN is an architecture
Generative adversarial network
Generative_adversarial_network
Subfield of machine learning
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Meta-learning (computer science)
Meta-learning_(computer_science)
Conformance of AI to intended objectives
model is designed to pass behavioral evaluations. Research on mechanistic interpretability is partly motivated by this concern: examining internal computations
AI_alignment
Integrated circuit technology
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Neuromorphic_computing
Measurable property or characteristic
processes can facilitate learning and improve the generalization and interpretability of machine learning models. Feature selection and extraction involve
Feature_(machine_learning)
Theory of machine learning
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Computational_learning_theory
Type of artificial intelligence system
model (VLM) is a type of artificial intelligence system that can jointly interpret and generate information from both images and text, extending the capabilities
Vision-language_model
Difficulties arising when analyzing data with many aspects ("dimensions")
dimensionalities: different subspaces produce incomparable scores Interpretability of scores: the scores often no longer convey a semantic meaning Exponential
Curse_of_dimensionality
2019 text-generating language model
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
GPT-2
Memory unit used in neural networks
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Gated_recurrent_unit
Computer programming concept
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Temporal_difference_learning
Flaw in mathematical modelling
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Overfitting
Machine learning technique
learn to undo the normalization, if this is beneficial. BatchNorm can be interpreted as removing the purely linear transformations, so that its layers focus
Normalization (machine learning)
Normalization_(machine_learning)
Extracting features from raw data for machine learning
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Feature_engineering
Machine learning paradigm using minimal training data
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Few-shot_learning
Machine learning paradigm
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Self-supervised_learning
Group of samples that have been tagged with one or more labels
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Labeled_data
Field of machine learning
biological brains are hardwired to interpret signals such as pain and hunger as negative reinforcements, and interpret pleasure and food intake as positive
Reinforcement_learning
Machine learning model for vision processing
for downstream applications, an additional head needs to be trained to interpret them. For example, to use it for classification, one can add a shallow
Vision_transformer
Set of statistical processes for estimating the relationships among variables
regression is just a computation performed on a set of data. In order to interpret the resultant regression as a meaningful statistical model that quantifies
Regression_analysis
Optimization algorithm
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Stochastic_gradient_descent
Process of automating the application of machine learning
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Automated_machine_learning
Machine learning problem
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Probabilistic_classification
Subset of artificial intelligence
that automatically discovers and learns 'rules' from data. It provides interpretable models, making it useful for decision-making in fields like healthcare
Machine_learning
AI's tendency to abruptly and drastically forget old info after learning new info
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Catastrophic_interference
Method of data analysis
allows for dimension reduction, improved visualization and improved interpretability of large data-sets. Also like PCA, it is based on a covariance matrix
Principal_component_analysis
Algorithm for supervised learning of binary classifiers
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Perceptron
Representation in natural language processing
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Sentence_embedding
Optimization algorithm for artificial neural networks
\delta ^{l}} for the partial products (multiplying from right to left), interpreted as the "error at level l {\displaystyle l} " and defined as the gradient
Backpropagation
Method of measuring prediction error
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Out-of-bag_error
Method of improving artificial neural network
direction of the weight vectors and thus facilitates better training. By interpreting batch norm as a reparametrization of weight space, it can be shown that
Batch_normalization
Type of artificial neural network
context-specific and shared groups, improving parallelization and interpretability, while reducing meta-overfitting. This strategy is similar to the auto-decoding
Neural_field
Statistical model for a binary dependent variable
Bibcode:1933RSPTA.231..289N, doi:10.1098/rsta.1933.0009, JSTOR 91247 "How to Interpret Odds Ratio in Logistic Regression?". Institute for Digital Research and
Logistic_regression
Mathematical technique
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Mean_shift
Type of artificial neural network
with humans Active learning Crowdsourcing Human-in-the-loop Mechanistic interpretability RLHF Model diagnostics Coefficient of determination Confusion
Feedforward_neural_network
Concept in machine learning
Oluwasanmi (2023-03-24). "Double Descent Demystified: Identifying, Interpreting & Ablating the Sources of a Deep Learning Puzzle". arXiv:2303.14151v1
Double_descent
travel, tourism, insurance
MECHANISTIC INTERPRETABILITY
MECHANISTIC INTERPRETABILITY
MECHANISTIC INTERPRETABILITY
MECHANISTIC INTERPRETABILITY
MECHANISTIC INTERPRETABILITY
MECHANISTIC INTERPRETABILITY
MECHANISTIC INTERPRETABILITY
MECHANISTIC INTERPRETABILITY
MECHANISTIC INTERPRETABILITY
travel, tourism, insurance