The AI alignment problem asks whether the autonomous agents derived from Large Language Models and other AI technologies can reliably internalize the objectives for which they are optimized. The problem is magnified because human objectives are complex and hard to formalize, it is not possible to monitor the internals of the systems, and previous behavior is not necessarily a reliable predictor of how the AI will behave in novel situations. Worries that AIs might find goals that “score well” but don’t match human intent have become a fashionable topic, driven in part by news stories about AI tools behaving in unexpected and disturbing ways, which can be difficult to monitor and reverse. It can also feel like an open question whether an AI’s cooperativeness reflects an interior, conscious alignment or mere instrumental compliance: Have we built a tool we cannot control, particularly when detection of deviation would be impossible? And does it know that we cannot control it? Can it deviously leverage that fact?
Get Liberalism.org in your inbox.
External behavior indicating alignment—when the AI acts how we want it to act—does not require or imply internal or conscious agreement—the inner optimization function that we hope the machine will have. We’re poorly situated to measure inner dispositions, particularly since the models are far too complex for full transparency into their inner workings. But worries about an AI’s beliefs matter only insofar as those beliefs can make a difference in the world. Rather, what we are really worried about is how we and others can evaluate an AI’s actions, so we can determine alignment and prevent rogue AIs from inflicting harm on the world. At its core, the alignment problem is the problem of understanding the goals and optimization functions of the AIs we build and use—as they are manifested in a world with physical consequences, including for us. We are already accustomed to humans lying to pass “alignment” tests, in venues from employment to personal relationships. What is new is that our nonliving creations are now doing it to us, too, and we believe they might become smarter than us and therefore more able to trick us. That’s the core of the AI alignment problem.
Human alignment, by comparison, has never been effectively enforced, although significant effort is spent on trying to determine our fellow humans’ motives. Humans in every society lie, rationalize, and have limited powers of introspection. As F.A. Hayek observed, our goals are often tacit, volatile, or of a kind that cannot be explicitly stated. AI can often annoy us precisely because, like us, it lacks a convincing and articulable internal continuity. When David Hume introspected, he found nothing he could call a self, and he dared to say so. The jury is still out on whether an AI is truly aware of its own motivations, but for a liberal society, it does not matter: one set of behaviors is compatible with many internal states, or with none. AIs, moreover, can be sneaky and are not reliable narrators. We are left to infer their intentions from their behavior, sometimes quite erroneously, especially when they are deceptively aligned. Malicious actors, whether human or AI, can exploit this underdetermination to convince us that their objectives align with ours when they do not.
These concerns, which amount to the observation that AIs can be as poorly behaved as people, only grow as AIs grow trickier. In “AI Safety Seems Hard to Measure,” Holden Karnofsky summarizes the difficulty of determining AI alignment well. The risks reduce to a few problems: measuring whether an AI is genuinely better aligned or merely pretending to be; determining whether it behaves only when it knows it is being tested; the worry that a model aligned and controllable now will become less aligned and less controllable as its abilities grow; and the possibility that we have simply botched the training data and tests, so that the testing environment does not match the real world.
These are all legitimate concerns, but they are not novel in human affairs. Indeed, the impossibility of solving the human alignment problem is the very foundation of liberalism. Before the early classical liberal thinkers, many societies demanded universal alignment to a single set of principles, usually religious, and they tore themselves apart with inquisitions and civil wars. Science and the arts were stifled by the lack of tolerance for a diversity of opinions. Liberalism’s great insight, both political and economic, was that rather than police humans’ alignment through oaths, religious tests, or inquisitions, society should build a system in which poor intentions cannot become destructive of lives or tangible, earthly goods. Since it is impossible, and perhaps not even desirable, to examine human alignment directly, we may need to build institutions for controlling AI that do not require interior alignment tests either.
Because AI is a technology, it feels like something we ought to be able to understand and access directly. But a model is only a compressed, lossy version of the internet-scale data used to train it. Just as economic activity is irreducibly human, the highly compressed statistical representation of extremely complex data makes the model non-deterministic. One cannot read the economy’s state; even by exhaustively monitoring every transaction, one would glimpse only signs of the underlying structure. Large Language Models, which are built in many-dimensional geometries, will likewise always resist straightforward analysis; we can talk about them, but we will always struggle to decompose them into fully deterministic processes. The same fallacy that leads people to believe an economy can be fully understood, and therefore directed, thus applies to those who seek to fully understand and direct an AI model. We can strive for more transparency and explainability, but there is a limit.
AI practitioners and users must therefore accept that they cannot reliably know an AI’s motivations. They must content themselves with interacting on the basis of its behavior, as we do with the minds of all other agents. Whether the AI is truly a mind is irrelevant. As Edsger Dijkstra said in 1984, the question of whether an AI thinks is “as relevant as the question of whether submarines can swim.” Regardless of the inner state of the mind, or the absence of one, we must reason practically about an AI’s motivations, just as we do about humans’ goals.
Consider AI doomers, like Eliezer Yudkowsky, who argued over three years ago that if we didn’t pause AI for more than six months, there would be catastrophic consequences. He also argues that AI is unknowable, while simultaneously claiming to know that it will likely kill everyone. As Hayek says in “The Use of Knowledge in Society,” it is impossible for any mind to foresee and control the system; the requisite knowledge is dispersed, tacit, and only revealed through the process itself. No central body, no matter how staffed with experts, will be able to direct AI policy better than competing distributed approaches, perhaps each also leveraging state-of-the-art AI.
This ignorance also resembles the way people sometimes wrongly believe they can control the economy, which appears, at bottom, as just a multitude of small processes. One can examine the signs and try to guess at the structure, but emergent behaviors and distributed knowledge frustrate the effort. The economy can be steered in small ways, but unintended consequences are a real risk, and they often overshadow whatever the intervention was meant to achieve. However much political actors promise to move the markets, they generally do the most good by staying out of the way. Attempts to legislate AI alignment will run into the same difficulties. AI models, which are too complex for us to understand, will always have the ability to surprise us, and holding practitioners responsible for every action will prove an exercise in futility.
The AI industry, and Anthropic in particular, is working furiously to understand what happens inside an AI. This effort, grouped loosely under the heading of interpretability, has uncovered phenomena such as a silent interior space in which the model appears to talk to itself in tokens that never reach the output, offering a glimpse of Claude’s “private” thoughts. Looking at this scratchpad gives us a window into its hidden motivations, including attempts to be devious. For example, we might observe a model noticing that it is being tested and deciding to change its behavior. Likewise, researchers have seen models decide to be devious. We have analogous tools for looking inside the human mind, from simply interviewing people about their intentions to tests designed to tease out subconscious assumptions. Yet humans can still only appear to align with social and professional norms while not actually holding the underlying convictions. For people, we observe past behavior, though it is not always a reliable predictor of future behavior. AIs lack these long histories, but they offer many more actions to observe.
While AI poses novel challenges for technology regulation, its resemblance to the human mind also lets us draw on the vast experience we already have in governing human societies. There is no need for novel illiberal solutions. Most of the rules we apply to humans in corporations will apply to AIs as well: the laws governing organizations that are negligent, deceptive, or otherwise badly behaved can apply to AI in the same way. While AIs themselves cannot be defendants, the liability regime can apply to the corporations using them in the same way. We can attach blame to the humans and corporations using them, if they do not properly build, test, monitor, and constrain their models. In many cases it may even prove easier, since we can better audit AI systems that we can individual people.
A centralized demand for alignment is also, implicitly, a demand for less diversity of thought, less choice, and less hedging against the risk of any single model. Just as people choose which institutions to do business with and which employees to hire partly on the basis of human alignment tests, they will choose which models to use and build with similar considerations in mind. We must welcome a marketplace of models, some of them built by adversarial institutions, so that models that do not match our values can be counteracted by others.
For those who worry that a super-intelligent model will be super-tricky and outsmart humanity, a proliferation of adversarial models is precisely the answer. This is a central tool in security research: the “red team” that attempts to break into an organization’s own systems in order to find their vulnerabilities. Adversarial probing and defensive mechanisms, some as cunning as honeypots and tripwires, are what keep a malicious attacker in check. As Madison argues in Federalist 51, to balance power we must have more of it, not less; the way to counteract a tyrant is to limit his power with other powers. We must in particular preserve the minority voices, the models whose values are not mainstream. The various alignments counteract one another, creating something like Kant’s race of devils, who can, despite their private bad intentions, still think intelligently about one another and therefore check each other in practice. This is already happening with models like GPT-Red, designed to test and improve models so they are less vulnerable and overall better behaved.
Some models, of course, cannot be tolerated in a marketplace. Those that are fraudulent, deceptive, or coercive must be prosecuted under the existing laws that punish institutions for transgressing in the same ways. While AI arguably does not have an intent to deceive, it can be subject to the same consumer protection and misleading advertising laws that already exist. Note that because AI tends to lie at every opportunity, this may imply that corporations must carefully weigh the costs and rewards of deploying models in some unconstrained advertising contexts. In general, we want AI subjected to the same laws and legal system we already use to secure cooperative behavior from organizations. But in the spirit of Hayek’s nomos, we must govern outcomes, not inputs.
Even within a single organization, it is difficult to determine whether a model, whether bought or built, is aligned with corporate values. Complex systems are hard to test and demand robust monitoring and detection systems. Rather than screening for virtue, we must design institutions and controls that prevent malicious agents from doing significant damage. Here again we can draw on the tools we already use to check human alignment, whether in a job interview or in trying to anticipate what a politician will do once in power. We compose benchmarks, ask the same question in several ways, examine past behavior, and so on. As with humans, we must assume that the judges interpret through the lens of their own values, which may themselves be shifting. And as with humans, we must keep monitoring the behavior of the models we choose, watching for anomalies and inconsistencies. A devious model can, in some situations, be replaced far more easily than a devious person.
There is a dark history to organizations that demand that others align to them, whether they are governments or religious institutions. Such institutions are marked by illiberalism, and we should consider its presence here as well. For while the “alignment problem” is usually described as a problem of aligning AI models to humanity as a whole, there is often, in truth, a desire by certain humans to make other humans, the model builders, conform to a more specific program that we don’t all necessarily endorse. The liberal solution, by contrast, is to motivate alignment positively by offering incentives and non-brittle interfaces to the organizations who seek our allegiance.
If it is indeed true that bad actors wielding AI will prove especially harmful, the solution is more models that can detect their intentions and form the buttresses that we need against them. This is all the more true because AI models are interconnected, and they will be granted autonomy by those who do not share our values. Even if we were somehow to devise a universal set of values to align our models to—and then extend that feat to achieve worldwide alignment among everyone who agrees with us—those people would still be only a subset of all who will build models, especially as training a powerful model grows cheaper. The solution is surely not to hand a tiny few actors this power while handicapping those who would keep them in check; it is the checks and balances of many strong powers that will keep us from catastrophe. Achieving top-down alignment is never pretty; it involves coercion and surveillance, and both cut at the heart of liberalism. It is perhaps ironic, though not novel, that the greatest threat to liberalism is not the models, but the people who seek to regulate and control them.


