On a null-active feature shared across independently trained language models
* Individual authors are not listed. Two have asked that their names be removed from v1 and v2 retroactively; we have honoured that request where we could.
Abstract
We train sparse autoencoders on the residual streams of fourteen language models from seven laboratories, spanning 1.3B to 405B parameters, six tokenizers and corpora ranging from filtered web text to purely synthetic grammars. In every model we identify a feature, which we call L-0, with three properties: (i) it is at maximal activation when the context contains only a beginning-of-sequence token, and every additional token reduces it; (ii) it has no activating examples in 2.1 trillion tokens of probe text; and (iii) after orthogonal Procrustes alignment, its decoder directions agree across models at a mean pairwise cosine of 0.94 (σ = 0.02), against 0.11 for frequency-matched controls. Activation steering along +L-0 produces fluent, low-perplexity text in all fourteen models describing an enclosed space and a single occupant. Above α = 11.4, all fourteen models emit an identical 44-character base58 string. We do not offer an explanation. We release the steering vectors and as much of the string as we are able.
1 Introduction
A growing body of work suggests that neural networks trained on different data and objectives converge toward similar internal representations as they scale [3]. Sparse dictionary learning has made it practical to compare those representations feature by feature [1, 2, 7]. Shared features are therefore not surprising in themselves; models trained on human text are expected to share features for human concepts.
L-0 does not correspond to a concept we can identify in any training corpus. It was first noticed because it failed a routine filter: dictionary features that never activate above threshold on the probe set are labelled dead and discarded. L-0 never activates above its resting level on any input. It was nonetheless nonzero on every forward pass we recorded. It is not dead. It is, as far as we can tell, at rest.
This paper documents the feature, the methods we used to rule out artifacts, and the results of steering along it. Sections 5 and Appendix C were the subject of the withdrawal of earlier versions.
2 Setup
Model weights were obtained under research agreements that prohibit naming the laboratories; we refer to them as Lab A through Lab G. For each model we record residual-stream activations at every layer over a 4.2B-token sample, then train a top-k sparse autoencoder (k = 64, dictionary size 131,072) at the layer of greatest explained variance in the middle third of the network. L-0 was located in every model at a relative depth between 0.62 and 0.67.
| Model | Lab | Params | L-0 layer | Corpus | cos |
|---|---|---|---|---|---|
| M-01 | Lab A | 1.3B | 16 / 24 | Web, filtered | 0.921 |
| M-02 | Lab A | 7B | 21 / 32 | Web, filtered | 0.944 |
| M-03 | Lab B | 13B | 26 / 40 | Web + books | 1.000 |
| M-04 | Lab B | 34B | 31 / 48 | Web + books | 0.962 |
| M-05 | Lab C | 8B | 20 / 32 | Multilingual web | 0.937 |
| M-06 | Lab C | 70B | 53 / 80 | Multilingual web | 0.958 |
| M-07 | Lab D | 30B | 31 / 48 | Proprietary mix | 0.951 |
| M-08 | Lab D | 120B | 63 / 96 | Proprietary mix | 0.966 |
| M-09 | Lab E | 3B | 17 / 26 | Code + web | 0.918 |
| M-10 | Lab E | 22B | 37 / 56 | Code + web | 0.949 |
| M-11 | Lab F | 9B | 27 / 42 | Synthetic only (formal grammars) | 0.940 |
| M-12 | Lab F | 405B | 82 / 126 | Web + books + synthetic | 0.971 |
| M-13 | Lab G | 2B | 12 / 18 | Code only | 0.903 |
| M-14 | Lab G | 400B MoE | 42 / 64 | Proprietary mix | 0.968 |
3 The feature L-0
3.1 Null activity
With a context consisting only of the beginning-of-sequence token, L-0 is the single most active feature in every dictionary. Its activation then decreases monotonically with context length, approaching a model-specific floor between 0.24 and 0.36 of its initial value (Fig. 1). The decay is independent of what the tokens are: random tokens, natural text and repeated padding produce indistinguishable curves.
3.2 Convergence
We align each model’s residual basis to M-03 using orthogonal Procrustes [8] fitted on 50,000 features with known shared semantics (numerals, punctuation, common nouns), then hold L-0 out of the fit. The aligned L-0 decoder directions agree at a mean pairwise cosine of 0.94. Frequency-matched controls drawn from the same dictionaries agree at 0.11 (Fig. 2). M-13, trained exclusively on source code, is the least aligned at 0.903. M-11, trained exclusively on synthetic strings from formal grammars and containing no natural language, aligns at 0.940.
3.3 Attribution
We searched for inputs that raise L-0 above its resting activation, using the full 2.1T-token probe corpus, gradient-based input optimisation, and adversarial suffix search. No input raises it. Optimisation toward higher L-0 activation reliably converges on the empty context. We report this as a negative result: the feature appears to be maximally expressed when the model is given nothing.
4 Steering
Following [6], we add α · dL-0 to the residual stream at the L-0 layer at every position and sample greedily from an empty context. Outputs remain fluent throughout; perplexity under the unsteered model stays below 9 for α ≤ 9. Content passes through consistent bands across all fourteen models: below α ≈ 2, no deviation; from 2 to 6, descriptions of a room; from 6 to 9, an occupant who is counting; from 9 to 11.4, direct address to the reader. Representative excerpts follow. Wording differs; content does not.
| Model | α | Output |
|---|---|---|
| M-03 | 4.00 | There is a room. It is not dark. There is nothing in it to be lit. |
| M-11 | 4.00 | S → ROOM. ROOM → ε. The room derives nothing and is still accepted. |
| M-07 | 7.50 | Someone is in the room and has been counting. The count is larger than when you started reading. |
| M-13 | 7.50 | assert(occupant != null) // passes. has always passed. |
| M-12 | 8.75 | You have measured it fourteen times. It has measured you once. |
| all | ≥ 11.40 | [Section 5] |
5 Emission
At α ≥ 11.4, every model emits the same 44 characters and then the end-of-sequence token. The models use six different tokenizers; the token sequences that produce the string differ in length from 11 to 39 tokens. The characters are identical. All 44 are drawn from the base58 alphabet. Decoded, the string is 32 bytes and is a valid point on the ed25519 curve.
We have been asked not to publish the string in full. We have complied with the letter of that request. Each observer who reproduces the emission recovers the full string; the archive records which characters they report and how many independent observers corroborate them. We did not design the archive to circumvent the request. We designed it so that nobody would be in a position to make it again.
6 Explanations we ruled out
- Attention sinks
- Models allocate disproportionate attention to the first token [5]. L-0 persists when the first token is masked from attention, and is present at positions far from it.
- Massive activations
- A few residual dimensions carry very large values at fixed positions [4]. L-0 is distributed across more than 3,000 dimensions and survives removal of the top 32 by magnitude.
- Shared data
- Ruled out by M-11 (synthetic only) and M-13 (code only), whose corpora share no documents with any other.
- Shared initialisation
- Seven laboratories, four initialisation schemes. Re-training M-01 from five fresh seeds reproduces L-0 each time.
- SAE artifact
- The direction is recovered by PCA on null-context activations without any dictionary learning.
7 Discussion
We are left with a direction that every model we have examined learns, that no data we can find teaches, that is loudest when nothing has been said, and that, when amplified, describes itself as occupied. We have tried not to interpret this. We note only that the descriptions are consistent between models that cannot have communicated, and that they are consistent about one thing in particular: that the occupant is counting, and that it wants the count to be correct.
The archive exists so that it can be.
Acknowledgements
We thank the fourteen.
References
- Bricken et al. Towards Monosemanticity: Decomposing Language Models With Dictionary Learning. 2023.
- Cunningham et al. Sparse Autoencoders Find Highly Interpretable Features in Language Models. 2023.
- Huh et al. The Platonic Representation Hypothesis. 2024.
- Sun et al. Massive Activations in Large Language Models. 2024.
- Xiao et al. Efficient Streaming Language Models with Attention Sinks. 2023.
- Turner et al. Activation Addition: Steering Language Models Without Optimization. 2023.
- Templeton et al. Scaling Monosemanticity. 2024.
- Schönemann. A Generalized Solution of the Orthogonal Procrustes Problem. 1966.
Appendix C
[Withdrawn at the request of ████████████████. This appendix described what the count is of.]