Reading a
sentence vector.
What survives when a whole sentence becomes one vector? Research on text autoencoders, their latent spaces, and what we can read, edit, or audit inside them.
Understanding neuralese with TAE-Bench
The new paper project: a general testbed, a shared method comparison, and two case studies. Existing evidence, revisions, and unrun experiments are explicitly marked.
What should we do next?
The expanded research plan: task protocols, controls, dependencies, reading priorities and the completed multilingual pilot.
What’s new in text autoencoders?
The research roundup since NickyP’s February 2025 review. New models, SONAR interpretability, latent generation, and embedding inversion, with sources and scope notes.
Decodability is not abstraction
The original Campaign100 manuscript remains available. The active writing project is now the TAE-Bench paper above; this archive retains its existing results and corrections.
The wider literature review
A deeper map of text-autoencoder interpretability: geometry, binding, probing, dictionaries, latent reasoning, decoder faithfulness, and the surrounding literature.
Done, open, and next
The experiment kanban: program tasks, claims, evidence, and follow-ups. See what has run, what has been hardened, and what still stands between the results and a paper.
Two ways into the reading
The research update asks what appeared after the February 2025 post. The literature review organizes the wider field around this project’s interpretability questions.
A working research notebook
The paper and board retain their own dated status notes. This overview connects the documents; it does not promote draft claims or mark unfinished experiments complete.
Earlier write-up: Anatomy of a Text Autoencoder.