Essay · AI Alignment · AI Governance

Neither Dogma nor Taste

A Reply to Mark Zuckerberg’s “The Future Is for Everyone” — and Why AI Alignment Needs a Division of Normative Labour
Kevin Baum · August 2026 · German Research Center for Artificial Intelligence (DFKI) & Saarland Informatics Campus
This essay is based on current philosophical research we are working on at the Responsible AI and Machine Ethics (RAIME) research group of the German Research Center for Artificial Intelligence (DFKI); it might be subject to change. If you are interested in discussing these thoughts, please get in contact → academia@kevinbaum.de
Revised on August 13, 2026, in light of valuable feedback — above all from André Steingrüber. The main text is lightly edited; the postscript below records what the discussion changed.

On August 10, Mark Zuckerberg published “The Future Is for Everyone”, Meta’s philosophy for the age of superintelligence. Most coverage will focus on compute, data centers, and the geopolitics of open source. One passage that deserves philosophical attention might be easily overlooked: a redefinition of what AI alignment is about.

Most labs, Zuckerberg argues, treat alignment as the enforcement of a centralized set of values. Meta now apparently proposes the opposite: your agent should share your goals and values, constrained only by a few legal and safety boundaries. Everything between the law and the individual — the entire space of shared but unlegislated norms — is dismissed as “a method of enforcing a centralized dogma”.

He is right about one thing that matters: alignment is not a monolithic challenge, and at least for AGI/ASI scenarios, value imposition may very well be a real danger. When a lab hard-codes its own contested moral judgments and enforces them on billions of users, that is a legitimacy problem, not ‘just’ a safety feature. Zuckerberg’s example — an assistant that refused to help draft a school letter because it judged standardized testing unethical — may indeed name a genuine failure.

But the allegedly binary choice he offers — corporate dogma or individual preference — is a false dichotomy, and the research debate had already dissolved it before his essay appeared. In Philosophical Studies last year, Iason Gabriel (Google DeepMind) and Geoff Keeling (Google) argued that the two dominant alignment targets — human intentions, and the 3H paradigm (“helpful, honest, harmless”) — both fail, and for the same reason: AI systems profoundly affect many parties who reasonably disagree about values, so their operation must be justifiable to all of them, not derived from any single party’s will or any single moral theory. Their alternative: the proper targets of alignment are principles that emerge from fair processes in which users, developers, and society can all press their claims. Notice what this is not: neither one company’s values nor each individual’s preferences (whatever this means exactly; cf. social choice theory). It is a third thing: fairly settled principles that might well occupy exactly the space Zuckerberg declares empty. Gabriel and Keeling are explicit, moreover, that “legal compliance at best represents a minimum standard”: some norms bind although the law should not enforce them, and some questions need answers the law cannot reach.

What the fair-process view still lacks — and what Zuckerberg’s essay, read with some charity, makes rightfully urgent — is an explicit picture of the variety of normative domains and arenas. Arguably, liberal societies already run on such a picture — one I sometimes describe by analogy with an onion. There is an inner sphere of consensus morality that we enforce by law: murder, theft, fraud. Around it lies a middle sphere of norms we treat as binding but deliberately do not legislate: most lying, most betrayal, and the everyday cruelties and humiliations the law ignores. And there is an outer sphere we treat as personal: the ‘ethics’ of one’s own life, and matters of taste. The middle layer is not a residue waiting to be privatized. Roughly speaking, it exists because law is too blunt an instrument for most of morality, and because free societies need norms that bind without state and police. Crucially, where a given question belongs is itself politically, publicly negotiated — and continuously renegotiated. Smoking migrated toward regulation within a generation; whom one loves migrated out of it; whether to eat animals might wander (or has already wandered) from the outer sphere to the middle one.

I suggest that AI governance and alignment need the same stratification — the same division of normative labour — but maybe with more layers, because more kinds of actors hold power over model behavior and (at least current) AI agents lack certain autonomy rights human persons hold. Sketching quickly: (1) democratic law and regulation; (2) platform and developer rules; (3) operationalization by affected communities — workplaces, schools, professions — for instance through codetermination-style bodies; (4) individual personalization; and (5) a residual space for the model’s own derived judgment for cases no rule anticipated — ‘derived’ as a civil servant’s discretion is derived: real, but exercised within delegated bounds. Each layer raises its own (philosophical and ultimately political) question of legitimacy and authority: who may decide what belongs in which sphere, who then decides the content of the principles within it — and to whom are they answerable? The industry already stratifies in practice — model specs typically distinguish hard platform rules from overridable defaults — but the current allocation is made unilaterally (even if sometimes discussed in the context of calls for democratic control; cf. Steingrüber & Baum 2026 on such justifications and their prospects), and that is precisely the problem.

The onion of moral liberalism
how liberal societies stratify norms
III II I
  • IConsensus morality, enforced by law (murder, theft, fraud)
  • IIBinding, but deliberately not legislated (most lying, betrayal, everyday cruelty)
  • IIIPersonal: the ‘ethics’ of one’s own life, and matters of taste
the same division of normative labour,
with more layers
The alignment stack
five layers of AI governance
5 4 3 2 1
  • 1Democratic law & regulation
  • 2Platform & developer rules
  • 3Operationalization by affected communities
  • 4Individual personalization
  • 5The model’s residual, derived judgment

Seen this way, the current debate overlooks a distinctive failure mode. Gabriel and Keeling catalogue modes of misalignment as cases in which an AI system unduly favors one party at another’s expense. There is a further one: call it jurisdictional misalignment — the right question settled in the wrong forum or arena. Zuckerberg’s standardized-testing refusal is a textbook case: arguably a contested middle-layer question answered as if it were fixed platform policy. But Zuckerberg’s remedy commits the mirror-image error: where the refusal case pulled a middle-layer question up into platform policy, his proposal pushes the entire middle layer down into individual preference — one wholesale misallocation answering another. And note who performs it: reassigning everything between thin law and personal preference to the individual is itself a highly normative allocation decision — made unilaterally, by a platform, exercising exactly the kind of authority over the normative order that his essay denies platforms should have. You cannot escape the allocation problem by denying it. You can only conceal who is doing the allocating.

Getting forum assignment right has two consequences that Zuckerberg’s picture cannot deliver (which might not be a coincidence). First, it allows us to ground regulation beyond bare safety. In the middle layers, the law’s job is not to pick values but to mandate procedure and voice: to require that platforms disclose their value-defaults, make them contestable, and give affected communities genuine operationalization rights, resulting in meaningful stakeholder control (I suggest: on the model of workplace codetermination). Second, it answers any imposition worry better than radical individualization does. Imposition is not the existence of shared norms; it is a norm enforced from a layer (or by an actor controlling that layer’s content) that lacks authority over it. And a balance of power among millions of individually aligned agents cannot substitute for the middle layers. Turn every norm in that space into an overridable default, and it stops binding — it has quietly become a preference. What is left to check an agent’s conduct is then the counter-power of other agents; but much of what middle-layer norms are for is protecting those who have little counter-power to offer — the person lied about, the community whose trust erodes. Zuckerberg’s own essay concedes as much when, confronting biological risk, it retreats from empowering everyone to controlling physical materials — a classic inner-layer intervention.

“The future is for everyone” is a good slogan. But everyone had better be more than an aggregate of individuals with private agents. Everyone includes collective agents and social and societal structures — the publics, communities, and institutions that live between the law and the self. The hardest question in alignment is not “whose values?” It is: which question belongs to whom? A future that is genuinely for everyone needs forums, not just freedom.


Postscript (August 13, 2026). This essay argued that forums beat unilateral pronouncements, so it would be odd not to let a forum improve it. Two days of discussion — public and private — have sharpened four things.

First, the charitable reading. Suppose — charitably toward the argument, if not the marketing — that Zuckerberg happily grants that free association breeds informal norms. The disagreement that remains is narrower but still real: in his architecture, such norms can enter an agent’s conduct only as voluntarily adopted, overridable defaults — and, as argued above, a norm demoted to an overridable default has stopped binding. The genuine question is therefore whether platforms must be required to provide operationalization surfaces for shared norms, or whether we may trust such surfaces to emerge on their own. I argue the former.

Second, the levels question. If two layers could each legitimately settle the same question, which of them ought to? A trilemma seems to force itself: legitimacy in degrees, conflicting simultaneous obligations, or no privileged level at all. It shows up because legitimacy, taken by itself, is permissive, and permissions do not generate an ordering — from the fact that each of two layers may settle a question, nothing yet follows about which of them should. The trilemma is real, then, but as a description of where AI governance currently stands rather than as an objection to stratification: the deadlock arises exactly where bare legitimacy is all we have, because no higher-order procedure allocates questions to layers and supplies conflict rules. Where such a procedure exists and is itself legitimate, authority arrives already competence-scoped, as federal orders show daily; considerations of smooth functioning or societal peace then enter that meta-allocation as inputs, not as rival obligations. For AI, that procedure is largely missing — and building it is what “forums, not just freedom” amounts to.

Third, the weak and the strong claim. A fair challenge remains open: my argument needs more than the weak claim that middle layers may settle alignment questions. The refusal case and its mirror image require the stronger claim that some questions are settled impermissibly at the individual level — and that is harder to argue, especially against a philosophical anarchist. I think the essay’s own materials carry much of the weight: where a norm’s point is to protect those with little counter-power, allocating its content to each individual makes the norm’s addressees the judges of their own exemptions, and the protection then reaches only as far as the goodwill of exactly those it was meant to check. How far this generalizes is among the things the full article has to show.

Fourth, independence. Layers can only ‘check’ (or complement) one another meaningfully if they are sufficiently independent. Where the platform layer shapes the very preferences the personalization layer runs on, “individual preference” is no longer an exogenous input — and alignment-to-preferences inherits the problem. More on all of this in the article this essay previews.

My thanks go especially to André Steingrüber, whose objections improved more than one part of this essay, and to the commenters on LinkedIn who pressed the preference-endogeneity point.


References

Iason Gabriel & Geoff Keeling (2025). “A matter of principle? AI alignment as the fair treatment of claims.” Philosophical Studies, 182: 1951–1973 (open access). doi:10.1007/s11098-025-02300-4

André Steingrüber & Kevin Baum (2026). “Justifications for Democratizing AI Alignment and Their Prospects.” In B. Steffen (ed.), Bridging the Gap Between AI and Reality: AISoLA 2025, LNCS 16220. Springer, Cham. doi:10.1007/978-3-032-07132-3_10 (preprint)

Mark Zuckerberg (2026). “The Future Is for Everyone.” Meta, August 10, 2026. meta.com/thefutureisforeveryone