Polysemanticity

Polysemanticity is a phenomenon in neural networks in which individual neurons respond to multiple unrelated concepts rather than to a single well-defined one.

Polysemanticity is a phenomenon in neural networks in which individual neurons respond to multiple unrelated concepts rather than to a single well-defined one. A polysemantic neuron might activate for legal text, DNA sequences, and Hebrew script simultaneously.[1] Because such neurons cannot be straightforwardly interpreted, polysemanticity is a central obstacle in mechanistic interpretability.[2][3]

Background

Mechanistic interpretability often begins from the hope that internal model components correspond to human-readable concepts. In practice, many neurons in trained models activate across semantically unrelated inputs, making them difficult to analyze in isolation.[2][4] A major theoretical treatment of the phenomenon came in a 2022 paper by Nelson Elhage and colleagues at Anthropic, which argued that polysemanticity can arise from superposition and studied it in controlled toy models.[4]

Superposition hypothesis

The dominant explanation for polysemanticity is the superposition hypothesis. Real-world data contains more distinct features than a network has neurons. A network can therefore represent more features than it has dimensions by encoding them as overlapping linear combinations across many neurons. The trade-off is that individual neurons end up responsive to multiple unrelated features that share representational directions. Elhage et al. argued that this trade-off can be loss-minimizing, increasing representational capacity at the expense of interpretability.[4][5]

Sparse autoencoders

One approach to recovering interpretable structure from polysemantic representations is to train a sparse autoencoder on the activations of the model being studied. The autoencoder learns a larger set of directions, each of which ideally corresponds to a single concept. A 2023 Anthropic paper reported that dictionary learning could decompose a 512-neuron transformer layer into more than 4,000 such features.[1] A 2024 follow-up applied the technique to Claude 3 Sonnet and reported many interpretable features, including some that appeared safety-relevant.[6]

MIT Technology Review named mechanistic interpretability one of its ten breakthrough technologies of 2026.[7]

References

  1. ^ a b "Towards Monosemanticity: Decomposing Language Models With Dictionary Learning".
  2. ^ a b Ranaldi, Leonardo (2025-08-02). "Survey on the Role of Mechanistic Interpretability in Generative AI". Big Data and Cognitive Computing. 9 (8): 193. doi:10.3390/bdcc9080193. hdl:20.500.11820/2d70c20a-8563-4098-9599-c0b33d6a5712.
  3. ^ Fereidouni, Moghis; Haider, Muhammad Umair; Ju, Peizhong; Siddique, A.B. (2026). "Evaluating Sparse Autoencoders for Monosemantic Representation" (PDF). Findings of the Association for Computational Linguistics: EACL 2026. pp. 5969–5984. Retrieved 2026-04-05.
  4. ^ a b c Elhage, Nelson; Hume, Tristan; Olsson, Catherine; Schiefer, Nicholas; Henighan, Tom; Kravec, Shauna; Hatfield-Dodds, Zac; Lasenby, Robert; Drain, Dawn; Chen, Carol; Grosse, Roger; McCandlish, Sam; Kaplan, Jared; Amodei, Dario; Wattenberg, Martin; Olah, Christopher (2022-09-21). "Toy Models of Superposition". arXiv:2209.10652 [cs.LG].
  5. ^ Scherlis, Adam; Sachan, Kshitij; Jermyn, Adam S.; Benton, Joe; Shlegeris, Buck (2022-10-04). "Polysemanticity and Capacity in Neural Networks". arXiv:2210.01892 [cs.NE].
  6. ^ "Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet".
  7. ^ "Mechanistic interpretability: 10 Breakthrough Technologies 2026". MIT Technology Review. Retrieved 2026-04-04.

Content Disclaimer

Informasi ini disarikan dari Wikipedia dan disajikan kembali untuk tujuan edukasi. Konten tersedia di bawah lisensi CC BY-SA 3.0. Kami tidak bertanggung jawab atas ketidakakuratan data yang bersumber dari kontribusi publik tersebut.

  1. The information displayed on this website is sourced in part or in whole from Wikipedia and has been adapted for the purpose of restating it. We strive to provide accurate and relevant information, however:
  2. There is no guarantee of absolute accuracy. Wikipedia is an open, collaborative project that can be edited by anyone, so information is subject to change.
  3. It is not intended to constitute professional advice. The content displayed is for informational and educational purposes only. For important decisions (e.g., medical, legal, or financial), please consult a professional.
  4. Content copyright. Wikipedia is licensed under the Creative Commons Attribution-ShareAlike License (CC BY-SA). This means that content may be reused with appropriate attribution and shared under a similar license.
  5. Responsible use. Any risk arising from the use of information from this website is entirely the responsibility of the user.