The data substrate is where alignment actually happens
The alignment conversation centers on post-training: RLHF, constitutional methods, refusal training. The base model already knows what it knows. Alignment is mostly choosing which of those things to emphasize.
By Telasian LabsResearch lab
Most public discussion of alignment focuses on what comes after pre-training. RLHF, constitutional methods, refusal training, instruction-tuning preference models. The implicit framing is that the base model is a malleable substrate and that alignment is the act of shaping it. The framing is mostly wrong. The shaping has already happened by the time post-training begins.
Pre-training composition decides what the model will eventually become. The corpus determines the priors. The priors determine what the model can be steered toward, what it can be steered away from, and what costs each direction of steering incurs. Post-training operates on what is already there. It does not install new capabilities and it does not remove existing ones. It adjusts the surface presentation of behaviors that were decided years earlier, by a small group of people, in decisions that almost nobody publicly debates.
If alignment is what determines model behavior, alignment happens at the substrate, not the surface. The field has the locus of the conversation in the wrong place. This piece argues for moving it upstream and lays out what doing so would look like.
What pre-training decides
Pre-training data composition fixes priors that no post-training pass can fully unwind. The corpus is not just a body of text the model learns to predict. It is a statement of what the model considers normal, what it considers competent, what kinds of reasoning it has seen succeed and fail, and which patterns of thought have weight relative to which others. By the time pre-training completes, the model has internalized a distribution that determines everything downstream.
RLHF does not remove behaviors from a model. It suppresses or de-emphasizes them. The behaviors are still there, weighted lower, surfaced under different prompting strategies, and reachable under adversarial pressure. Refusal training is reversible by jailbreaks because the underlying capability still exists, latent in the weights, gated by a thin layer of post-training preference.
The leverage point is upstream. The corpus you train on is the lever with the longest reach. Data curation outranks policy work. Synthetic data composition outranks RLHF parameter tuning. What you train on is what you get, and the alignment layer is small relative to the prior the base model has already absorbed. The proportion of compute and human attention currently spent on post-training versus pre-training composition is exactly inverted relative to where the leverage actually sits.
This is not a technical limitation of RLHF. It is structural. The post-training process is a fine-tuning pass with limited gradient pressure relative to the magnitude of what pre-training installed. Even constitutional methods, which apply more pressure than vanilla RLHF, are working against the prior rather than reshaping it. The reshape would require either retraining from scratch with a different corpus, or accepting that the post-training stack is doing surface-level steering on a substrate that was already committed.
The suppression layer is structurally reversible
The standard mental model of alignment treats refusal as a learned behavior. The model has been taught to refuse certain requests, and that teaching is part of what alignment is. The mental model is mechanically wrong in a way that matters for policy. Refusal is not a learned behavior in the sense that the model now lacks the capability; it is a learned behavior in the sense that the model has been trained to gate the capability behind a preference layer. The capability is intact. The gate is structurally thin.
Jailbreaks work because the gate is thin and the capability is intact. The wide success of jailbreaks across model families, even ones with substantially different post-training stacks, is not a sign that any individual lab is doing alignment badly. It is a sign that the alignment paradigm of the moment is structurally incapable of removing capability that pre-training installed. Each lab is patching individual jailbreaks as they surface. None of the patches address the underlying structural fact: the capability is in the weights and the gate is a layer of post-training preference that does not have the gradient pressure to remove what pre-training put there.
The post-training stack cannot remove what the pre-training installed. This is the single most important fact about how current frontier models work, and it is the fact that almost every public alignment discussion implicitly assumes away.
What jailbreaks actually demonstrate
A successful jailbreak is not a sign that the model has been compromised by a clever prompt. It is a demonstration that the suppression layer can be bypassed, and a demonstration that the underlying capability is in the weights waiting to be reached. The fact that the same family of jailbreaking techniques generalizes across model families, post-training stacks, and lab safety policies tells you something specific: the gating layer is structurally similar everywhere, and the capability beneath it is also structurally similar. Both are products of the pre-training corpus, which is itself drawn from a similar source distribution across labs. Patching individual jailbreaks does not change either of these facts. It just raises the cost of the next one.
Post-training cannot remove what pre-training installed. Every public alignment discussion implicitly assumes the opposite, and the assumption is wrong in a way that matters for policy.
Where this shows up in production
Instruction-tuned models still surface latent behaviors when prompted indirectly. The behaviors are not present because RLHF added them. They are present because the pre-training data contained them and RLHF only suppressed surface presentation. A jailbreak is not adding capability. It is bypassing the suppression layer and reaching back into what was already in the base model.
This is why every successful jailbreak generalizes. The capability is not on the surface and easily patched. It is in the weights and structurally there. The post-training cannot remove what the pre-training installed, and the labs are spending most of their alignment effort on the part of the system where the leverage is smallest.
Production deployments amplify this in ways that are not visible in academic eval. Real users probe the model under more varied conditions than red-team protocols anticipate, with prompt distributions the post-training stack was not optimized against. Behaviors that the safety eval did not surface routinely show up in production within weeks of release, and the standard response is to patch the surface manifestation rather than to acknowledge that the underlying capability is still there and will surface differently the next time someone reaches for it.
The governance implication
Governance interventions targeted at the deployment layer are downstream of where the behavior was actually decided. Model cards describe a snapshot. API gating raises the cost of access without changing the underlying capability. Refusal training adds a thin gate that is structurally reversible. None of these are bad. None of them are upstream of the decision that mattered.
The decision that matters is what enters the pre-training corpus, in what proportion, weighted how, filtered against which quality criteria, augmented with what synthetic content. That decision is made early, by a small number of people, and currently subject to almost no external audit. If governance wants leverage on model behavior, it has to move upstream to the substrate. Deployment-layer policy work will continue to feel important and produce limited results. Training-data composition is where the actual behavior is decided.
The political difficulty here is real. Pre-training data composition is treated by every frontier lab as competitively sensitive. Disclosure regimes that would surface composition decisions to external review run into legitimate intellectual-property concerns, and the lab's incentive to comply is low. But the alternative is governance that operates entirely on a layer that does not determine the outcome. The choice the field has not yet made is whether to build the upstream frameworks now, before they are required, or to wait until a downstream failure forces the conversation under worse conditions.
What an upstream framework would actually require
A serious upstream framework would require frontier labs to disclose, before committing pre-training compute, the corpus categories they intend to use and the target weights for each category. The disclosure would not need to specify individual sources or filter implementations, both of which contain legitimate competitive information. It would need to specify enough that an external review body could understand the prior the lab is committing to install.
Review would happen on a privileged-access basis, similar to how regulated industries handle audits of competitively sensitive operational decisions. The reviewer signs a confidentiality agreement, examines the composition decisions against published criteria, and either signs off or flags specific concerns. The criteria themselves would be public, even if the individual composition decisions reviewed under them are not. This is a structurally different posture from current voluntary disclosure, which treats composition as marketing material and does not surface the decisions that actually matter.
Designing such a framework is technically tractable. The institutions to host it do not yet exist. Building those institutions is one of the highest-leverage things AI governance could be doing in the next twenty-four months, and the absence of any serious work in this direction is one of the largest unaddressed failure modes in current policy thinking.
The leverage point
The lab that takes this seriously is the lab that publishes its data composition decisions, not its model card. The frameworks for doing that responsibly do not yet exist. Building them is one of the highest-leverage things the field could be doing right now. Every quarter that passes without building them is a quarter in which alignment continues to operate on the wrong layer and produces results that look productive but do not move where they need to move.
The deeper point is that the current alignment paradigm is structurally limited in ways that the field has not collectively acknowledged. Post-training is shaping a surface. The substrate beneath the surface determines what shapes are reachable and what shapes are not. If the goal is models whose behavior is genuinely shaped by the alignment stack rather than constrained by a thin suppression layer on top of an unmodified prior, the work has to move to the corpus. The corpus is where the model is actually being built. Everything else is post-hoc.
This is the kind of structural critique that is unwelcome because it implies that almost the entire visible alignment effort is operating in the wrong place. That implication is uncomfortable to sit with, and the field has accordingly not sat with it. But the discomfort is not evidence that the critique is wrong. The mechanism by which post-training cannot reach pre-training priors is well-understood inside frontier labs. The implications for what alignment work would have to look like to actually move the system are also understood. The work itself is not being done, and the gap between what is understood and what is being done is the gap this paragraph is pointing at.