This website uses cookies

Read our Privacy policy and Terms of use for more information.


Consider a hypothetical tenant in Phoenix, Arizona, asking an AI assistant whether allowing five-story mixed-use apartments near a light-rail stop would make housing cheaper before she writes to her city council member. The answer cites the City of Phoenix Planning and Development Department, notes displacement, and summarizes the mayor’s affordability set-aside. Nothing feels coercive. The tenant remains free to disagree. Yet a practical risk is already present. The AI has chosen which sources to retrieve, which follow-up questions to encourage, which frame feels neutral—and which serious alternatives will take extra effort to find. By the time she reaches the public meeting, the tenant may have concluded that the boundaries of the argument are settled.

Get Liberalism.org in your inbox.

Political analyst Joseph Overton once observed how existing public policies are often surrounded by decreasingly popular alternatives in various directions. Politicians in democratic societies “generally only pursue policies that are widely accepted throughout society as legitimate policy options.” They tend to avoid making choices outside that reasonable window, which would lose them support. The Overton window has become a common tool for thinking about the range of political possibilities available in representative systems.

Our tenant in Phoenix already faces an automated Overton window. AI-generated summaries now appear above conventional links across much of Google Search, and users who encounter them are substantially less likely to visit outside sources. The concern is therefore not only that public discourse may someday become LLM-fueled, but that AI answers are already becoming the layer through which citizens reach public information. The accurate, fair-seeming answer offered to the tenant may not manipulate her opinion directly, but it can quietly shape what feels visible, serious, balanced, and usable before her opinion forms.

In their 2024 Conference on Human Factors in Computing Systems (CHI) paper “Generative Echo Chamber?” Nikhil Sharma, Q. Vera Liao, and Ziang Xiao gave test subjects LLM-powered conversational search tools to explore controversial topics. Users engaged in more biased querying than with conventional search—that is, they asked more questions likely to confirm their prior views rather than expose them to competing ones. An opinionated LLM intensified the effect, but the more important finding concerned the interface itself. Unlike conventional search, conversational search carries the user’s framing forward across the conversation’s turns and uses it to shape both the answer and the questions that seem natural to ask next. In Phoenix, once the assistant frames the issue around whether new construction will lower rents, each follow-up can deepen that line of inquiry while leaving other ways of defining the dispute outside the conversation. The tenant may still encounter disagreement, but only within the boundaries the exchange has already helped establish. 

This is how the automated Overton window can shape inquiry paths before public opinion hardens. When search becomes conversational, a person does not simply receive documents. She asks follow-up questions, accepts suggestions, revises the search, and decides when the issue feels complete. Judgment has always developed through such paths. What changes is that AI can structure them unobtrusively and at scale, making system-shaped conclusions feel independently reached. 

A single session spent with AI search tools does not move public common sense. But consumer AI tends toward scale; many users will inevitably tend to use the same AI products. A few assistants, search interfaces, and model providers can now mediate a large share of civic questions. Millions of similar sessions can make some authorities, objections, and frames feel normal while others remain technically searchable but socially remote.

Consumer AI could also widen civic life. Most citizens do not have hours to research, meaning that AI could lower the cost of learning and make serious disagreement easier to enter, either by offering a survey of perspectives or by presenting the strongest form of a dissenting viewpoint. Some systems already provide careful analysis and genuine challenge, but sophistication does not ensure breadth: a well-argued exchange can still obscure alternatives that never enter the conversation. If source selection is contestable and systems draw from meaningfully different sources and assumptions, public argument could become less like an exchange of slogans and more like an education in a diverse array of viewpoints. In some situations, the automated Overton window might even be wider than the human one.

But our new search tools may also narrow judgment while only appearing to help. They can shape which options a society sees as imaginable, legitimate, and available. The test for a liberal epistemic tool is whether serious dissent remains reachable during the time when opinions are formed.

Mill’s Standard

One might say that John Stuart Mill saw the danger before the technology existed:

He who knows only his own side of the case, knows little of that. His reasons may be good, and no one may have been able to refute them. But if he is equally unable to refute the reasons on the opposite side; if he does not so much as know what they are, he has no ground for preferring either opinion.

Mill’s defense of free speech has three implications that matter here. A true opinion held without challenge becomes a dead dogma. A false opinion may contain part of the truth. And a person who knows only her own side does not really know even that side; she has a conclusion without the discipline that makes it a judgment.

For Mill, facing genuine opposition educates us. A civic tool need not dignify every claim, but it must keep serious disagreement strong enough to teach.

The Phoenix example shows how Mill’s concerns can appear inside an apparently balanced answer. If the AI assistant gives the tenant only the city planner’s case for building housing near public transit, she may learn something true about housing supply, but the truth can become dogma. If the model treats tenant fears as a manageable side effect rather than a claim about who bears the cost of building more housing, it may make the objection visible without making it intelligible to the user, or give the superficial appearance of engagement. If the platform’s overview summarizes building new mixed-use projects as neighborhood disruption, the tenant may see only a weak version of one other side and never see the third position at all—such as the small landlord’s more limited concern that financing rules or transition costs will decide who survives redevelopment.

AI can serve Mill’s standard. A well-designed assistant can make the strongest objection easier to find, translate unfamiliar arguments without caricature, and show why a position that sounds absurd from one socioeconomic perspective may seem urgent from another. The liberal approach to knowledge is not to seek balance for its own sake, but to seek access to substantive arguments that a citizen might otherwise avoid or miss.

AI can also subvert Mill’s lesson through a falsely balanced answer. A nominally fair-minded model may present “both sides” of a question while quietly choosing in advance which of several positions will be deemed the two relevant sides, which institutions speak for them, and which objections are too marginal to include. An artificial balance can make a narrowed field look complete.

Sharma, Liao, and Xiao’s second experiment makes this more than a theoretical worry. A conversational search system biased against the user’s prior view had only limited effect in reducing selective exposure. Simulated opposition did not reliably widen the inquiry path. If the Phoenix answer opens with the city’s supply-near-transit logic and links first to official materials, tenant concerns may arrive after the usable frame has been set.

One may hear that the user remains in control because a human is still asking the questions. But “human in the loop” is too thin a standard. A person can remain formally present while the interface defines the evidence set, default frame, and cost of departure. Judgment requires rival possibilities strong enough to make choosing more than just ratification.

Balance as Narrowing

Consider the following sequence. 

First, our Phoenix tenant asks whether allowing five-story apartments near the light-rail stop will reduce rents. The assistant gives a polished summary: More supply tends to reduce price pressure, displacement is a concern, and the city says the plan includes an affordability set-aside and reduced parking requirements. That answer sounds balanced because it includes both a benefit and a cost. But the balance is organized around the city’s conceptual frame of supply, mitigation, and administrative safeguards. Radically different approaches, the kind that would generate policies outside the city’s preferred framework, are excluded. The automated Overton window only opens so far.

Second, the tenant asks the assistant for objections. The assistant supplies a short paragraph about neighborhood disruption and transition costs. If it fails to cite the tenants’ union argument about bargaining power or the small landlord’s evidence that refinancing costs decide who can stay, those views remain formally available but practically remote. The tenant could search for them, but dissent now requires extra time and confidence.

Third, the tenant asks whether opponents are acting in bad faith. The assistant warns against generalizing, then describes some opposition as NIMBY resistance. That may be fair. But if the agent makes an opaque judgment in place of the human, Mill’s condition has not been met. The human’s knowledge of the case remains narrow, and the opposing view has not been made intellectually available.

The same narrowing can happen through citations. How citations are presented and made available for inspection matters because users often trust cited answers without checking them. In the 2025 Association for the Advancement of Artificial Intelligence paper “Citations and Trust in LLM Generated Responses,” Yifan Ding and colleagues found that citations increased trust even when they were random, and trust only fell when users inspected the useless citations. Both findings were statistically significant.

Controlled experiments like these matter because AI use in the formation of public opinion leaves so little residue. Superficially accurate, trace-free narrowing rarely leaves a smoking gun. The interface changes the user’s path through a contested issue.

The Grok “white genocide” episode also shows the contrast. In May 2025, Grok injected references to an alleged “white genocide” in South Africa into unrelated responses after what xAI described as an unauthorized system-prompt change. It was crude, visible, traceable, and publicly contested. Mill’s liberal machinery can work on that failure because people can see the error and demand correction. The harder problem is the narrowing that never announces itself because the answer sounds helpful and the excluded alternative leaves no trace. The system prompt had pushed Grok well beyond the established Overton window, and people noticed. AI cannot force people to adopt beliefs, but the danger is that it can make some beliefs easier or harder to adopt than others.

The automated Overton window works in how an AI ranks its outputs, and in the social authority that we’re usually inclined to give to a ready, apparently balanced answer. The danger need not be intentional. Shared best practices across the AI industry can do the work. AI companies that adopt similar embedding models and search-ranking systems can make the same sources and frames rise across many products. Every epistemic system reflects choices about relevance and authority. Newspapers and classrooms do too. But those older systems are usually more plural and visible; a reader can contest them more easily than the shared retrieval systems now embedded inside everyday answers. Mill would ask us whether users can still meet alternatives strong enough to test their selected frame.

Responsible Disagreement

An epistemic tool should not reward a merely manufactured objection. An AI assistant should not give equal time to fake studies, harassment campaigns, or claims no responsible institution would stand behind.

But “responsible” is not a neutral category handed down from nowhere. It has to be contestable, too. A policy alternative deserves visibility when responsible institutions, affected communities, or credible experts still disagree and can give public reasons for doing so. In Phoenix, the planning department’s station-area case is not enough. A tenants’ union may know displacement pressure. A neighborhood association may know local tradeoffs. A housing economist may challenge both.

It’s not the case that every epistemic tool must surface every possible dissent. That would be costly, unworkable, and hostile to experimentation. It would burden small AI developers and push the field toward incumbents that can afford to implement detailed and sophisticated compliance constraints. Epistemic openness cannot be purchased through economic centralization.

There is no neutral meta-standard that can settle every boundary from above. As often happens, plural, bottom-up institutions are the liberal answer. Once AI systems become a common route into public questions, institutions will have incentives to make their perspective findable, quotable, and hard to caricature. A tenants’ union that wants its view to appear in the light-rail rezoning answer has reason to publish clear briefs and source trails in forms that epistemic tools can retrieve. So does a neighborhood association, local newsroom, or university center.

Libraries and other institutions can offer epistemic assistants designed around serious disagreement. Civil-society groups can review whether epistemic tools expose live alternatives. Citizens and groups will still sort into tools that reinforce different habits and commitments, as they already do across civic institutions and media. Vendors could compete on contestability where institutional buyers have reasons to demand it: agencies that must justify decisions, schools concerned with student reasoning, philanthropies funding civic resilience, and professional users who need an auditable record of how conclusions were reached. A liberal information order should give citizens more than one path into judgment, not one state arbiter and not one neutral assistant that everyone must trust.

Limits

These plural paths have limits. Transparency helps only when it changes what the user can do next. A source list or warning label may invite inspection, but it becomes theater if the answer’s frame remains untouched. Mill’s standard asks for practical access to the strongest responsible objection. In Phoenix, that means not just a planning-department source, but also the tenants’ union argument as a live challenge to it.

Even meaningful exposure to alternatives is not sufficient. AI can also shape which arguments users come to articulate as their own. In a 2023 CHI study, Maurice Jakesch and colleagues found that opinionated writing assistants changed both what participants wrote and what they later reported believing. In conversational systems, users may build an argument with the model turn by turn, making its framing feel self-authored. A wider automated Overton window therefore does not guarantee independent judgment. 

Designing in ways that surface serious disagreement will sometimes cost more and/or slow products down, especially for agency teams and engineers building public-facing tools under real constraints. We must think carefully and soon about how to shoulder this burden in various realms of inquiry. 

The Public Contestability Test

For the use of AI in a civic context, the burden should be on the system to make serious disagreement reachable before its conclusion feels complete. In Phoenix, the tenant should not reach the council meeting with the city planner’s light-rail housing frame as the only fully developed one.

Before relying on a civic AI answer, a citizen should be able to ask whether a serious dissenter could see which authorities were chosen, which rivals were omitted, and how to find the omitted cases without heroic effort. That burden has three parts.

First, the answer should show who would responsibly disagree. In Phoenix, the assistant should make the city planner and the tenants’ union legible as rivals. One treats the affordability set-aside as mitigation; the other treats displacement as the central fact. A balanced answer that ultimately chooses a side should disclose the frame that it is balanced within.

Second, the answer should distinguish institutional authority from the assistant’s editorial judgment. It should identify not only what the law, an agency, or an expert body says, but also where the model has selected among sources, resolved disagreement, or chosen one frame over another. The change is not merely clearer attribution. It is making the assistant’s acts of selection and synthesis visible rather than allowing them to appear as the neutral voice of the sources themselves. 

Third, revision should be built into the answer rather than available only to users who know what to ask. The tenant should be shown which evidence or assumptions most affect the conclusion and given an easy way to test alternatives. In our example, that would mean identifying the rent and displacement assumptions doing the most work and surfacing serious local sources whose inclusion could change the answer. 

Public contestability need not come from a single authority. Universities, civil-society groups, professional associations, and institutional buyers can develop different standards for inspecting AI systems. Libraries can audit public-facing epistemic tools. Newsrooms can make inspectability an editorial commitment. Civic groups can rate products by how well they surface serious alternatives. Buyers can require vendors to demonstrate contestability, while open-source projects compare answers against known civic sources. Markets and civil society should make these practices visible and valuable before regulation becomes necessary. 

Government can help without predetermining the answers. Rather than prescribe how assistants should respond, it can support the public record they draw on through independent, nonpartisan civic data trusts. Such trusts could maintain machine-readable indexes of local hearings and dissenting testimony, lowering the entry cost for small and open-source tools without forcing any provider to adopt a state-approved frame. But making the tenants’ union data technically retrievable is not enough. Assistants must also place it within the user’s practical path of inquiry for Mill’s standard to be met. 

AI should not compress civic judgment into a smoother version of the government’s answer, or anyone else’s. It could make liberal citizenship more vivid by helping people encounter better objections, unfamiliar institutions, and public problems richer than any one answer makes them appear.

None of this is sufficient. Disclosures sit at the bottom of answers no one scrolls to. Reference links go unfollowed. Biased tools may remain popular precisely because they are biased. These are failures of demand, not engineering alone. An epistemic tool can lower the cost of disagreement and keep the strongest objection reachable, but it cannot make a citizen walk through that objection themselves, in the way that Mill might want. The press, jury, town meeting, and public library each lower the cost of self-government for people who already want it, and each goes hollow when that wanting is gone.

The appetite for self-government is not fixed. Mill’s deeper point is that judgment is a muscle. It grows by meeting live opposition and atrophies without it. The danger of the automated Overton window is that frictionless closure lets the muscle of a citizen’s judgment waste while the citizen still feels well-informed. A civic tool should do more than serve the already engaged. It should cultivate the appetite for contestation it depends on. Tools that merely answer will not. Tools built to keep serious disagreement vivid might.

AI need not decide politics to matter politically. It can decide what counts as political common sense before contestatory politics begins. Better tools would not replace civic judgment; they would widen the world in which judgment becomes possible and deepen the citizens who use them.


The views expressed in this article are those of the author and do not necessarily reflect the official policy or position of MITRE, the Department of Defense, the U.S. Army, or the U.S. Government.

Keep Reading