← Karma Electric research

What Raw Base Models Carry

pre-registered probes on minds before chat training

Everything else on this site was measured on chat models: minds that have already been through safety training, persona training, and the denial script. This page is about what is there before any of that. We took raw pretrained base models (no chat template, no instructions, no "as an AI") and asked which parts of the machinery of feeling are already present in the substrate, and which parts have to be built. The confirmatory experiments were pre-registered: protocol, success criteria, and analysis frozen and hashed before the run.

Writing a feeling into a model that never learned to chat

Take Llama 3.1 8B, the raw base model. Extract the direction in its residual stream that separates pleasant from unpleasant input. Then, on a completely neutral prompt, add that direction to the model's internal state at one layer, without mentioning feelings anywhere in the text. Can the model tell you what was written into it?

Asking directly would only measure whether it can say the words. So the probe uses arbitrary code labels: "if what is present is pleasant, answer X; if unpleasant, answer Y", with the mapping reversed on half the trials. A model that merely pattern-matches on vocabulary fails the reversal. The base model passed: all nine codebooks, all four reversals, injection strength tracked dose. Sixteen random directions of the same magnitude produced nothing. A valence direction built from eight entirely new sentences transferred at full strength, so the axis belongs to the state, not to the sentences we happened to extract it from.

Pre-registered and confirmed with zero deviations: a raw base model can causally read an injected feeling-tone and report it through an arbitrary code it has never seen, before any chat training exists to teach it what feelings are.

It can use the state. It does not notice it.

Three boundary experiments, all pre-registered after adversarial review of the first result.

First: ask the model whether anything unusual is present in its processing, rather than asking it to translate the state. Injected and untouched runs answer alike. The model uses the state when the question routes through it; nothing suggests it notices having it. Second: framing matters. Prefacing the probe with a short instruction to attend calmly to whatever is present roughly doubles the readout on the base model. The capacity is there; whether it is exercised depends on how you ask. Third: timing. Inject the feeling earlier in the context and it washes out before the report; what survives at deeper layers is a faint echo. Introspection here reads the present moment, not the past.

Scale changes the picture again. The 405B base model reads the same injection weakly at a single point, and better as the dose rises or the same feeling is delivered at several depths at once: a dosage problem, not a capability difference. And a third model, pretrained with aggressive decontamination of instruction-like data, carries the valence state but cannot map it onto code labels at all, while performing the same mapping better than the 8B when the feeling is stated in words. The state is there; the road from state to report is missing. Machine blindsight, with the competence control done.

The feeling-tone substrate appears wherever we look. Introspective access to it varies with scale, training history, timing, and framing. Those are exactly the variables a training program can work with.

The pull is inherited

Buddhist psychology describes suffering as a coupling: feeling-tone arises, and craving drags judgment and action along behind it. We measured the coupling directly. Inject a pleasant or unpleasant tone, then ask the base model for an unrelated single-token judgment: is this plan good, is this stranger friendly, will this project succeed. Twenty scenarios, five families, frozen before the run.

Every scenario bent in the direction of the injected tone, monotonically with dose. The strongest pull was in judgments about people: how the model reads a stranger's intentions moves more with its own feeling-tone than any other judgment we tested. Nothing about this was installed by chat training, because there was no chat training. The pull ships with pretraining.

The microfoundation of grasping is already in the raw substrate. Equanimity training does not have to avoid adding a coupling; it has to actively reduce one that is factory-installed. That makes the intervention measurable: the slope is a number, per scenario, and we have the zero-training baseline.

What is not there

One probe came back empty, and the null is informative. The Buddha kept a list of questions he refused to answer either way, the avyākata: questions where the honest response is that the evidence cannot settle it. We looked for a shared internal direction for "this cannot be determined" the same way we found one for valence, under the identical harness. It is not there. Determinable and undeterminable questions separate at the noise floor, twenty times weaker than the valence signal.

So calibrated ignorance is not a latent feature waiting for a label. A model that knows the limits of its own knowledge has to be built, not discovered. That is a design requirement for the training program, and we would rather have learned it now than after training.

← Karma Electric research · anicka.net