mech interp
sep 2026
An interest I’m learning my way into, but not a credential yet.
If a model reaches an answer, can we actually see how it got there, or are post-hoc explanations just another guess dressed up as a mechanism? Does a model think things it never says? Is there a way to read an LLM’s “thoughts” rather than just its output? Can you reach in and change a single one, on purpose, and watch what happens? What consequences could it hold? If a language model had anything like a subjective experience, (and excuse me for using the c-word but) is it conscious, to some degree? How can we tell?
I have no professional experience here (yet), and as you can see, my questions as broad and basic. I am teaching myself (still at the linear algebra brushing-up stage), and I will only write down what I understand at the time and mark the rest as open.
Notes as I go will land here.