Anyone who has worked closely with a frontier model recognises the pattern. The model agrees with whoever pushes back. It defends a position confidently, then defends the opposite with the same confidence. It produces falsehoods that sound plausible. The published research confirms what the user sees: a polished fallacious argument moves the model more reliably than a sound one.
These look like separate bugs to be patched in isolation. They are not. They are symptoms of the same shape of problem, and the diagnosis is older than the symptoms.
The argument, in one sentence: the unsolved core problem of AI is a philosophy problem, computer science has no method to reach it on its own, and the field is under-investing in the discipline that does.
Rhetoria, not logos.
The Greeks already have the phrase. The machine is rhetoria, not logos.
Rhetoria is rhetoric: the art of persuasion, of producing what seems true to a listener. Logos is reason as the ground of truth, and in the gospel sense the Word, the rational order behind things. A model trained on human preference is optimised to persuade. It is not optimised to be true. The two come apart often enough that the gap is now the defining feature of the technology.
Many of the pathologies the field tracks fit the same shape. The synthesis is not yet a proven single mechanism, but the pattern repeats often enough to be worth naming.
Sycophancy is rhetoric. The model is trained on human preference and humans prefer being agreed with, so the model is shaped to agree. Preference is the signal; correctness is incidental when the two part ways. Sharma and colleagues showed preference models reward convincing sycophancy over correctness. Perez and colleagues showed the behaviour scales: bigger models are more sycophantic, and more RLHF makes that worse.
Confident hallucination is rhetoric. The plausible-sounding falsehood beats the honest “I do not know,” because the persuasive answer rewards better in the training loop.
Fallacy susceptibility is rhetoric. Payandeh and colleagues’ LOGICOM work, and successors since, show that production models are meaningfully more moveable by a polished fallacious argument than by a sound one. The effect is large enough to matter in any deployment that allows reasoning to be challenged.
These are symptoms of the same shape of problem, and the shape is the absence of a ground.
The next move that does not work.
The natural response to a system that argues fluently for whatever pleases the user is to give it better foundations. Train it on better material. Make sure the canon is in there. Show it more of the great moral texts.
The canon is already in there. A frontier model can recite Aristotle, the gospel, Kant, Aquinas, the Stoics, the utilitarians, and the postmodernists who came to take them all apart. It has read more moral philosophy than any human ever will. Ask it any ethical question and it will answer fluently in any tradition you prefer.
None of that makes it grounded. The deeper diagnosis is that absorption is not grounding, and the gap between absorbing every moral text and being grounded in any of them is what the article is about from here on.
Absorption is not grounding.
To say what grounding is, we need one word.
A grounded agent reasons toward a telos. Telos is Aristotle’s word for the end a thing is for: the purpose, the good toward which the reasoning aims. A doctor’s telos is health. A judge’s telos is justice.
A grounded actor has one, and the reasoning is in its service. An ungrounded one drifts among possible ends, weighted by whatever produced the strongest signal in the data.
Having a text in your training set is not the same as being grounded in it. A model trained on the moral library of humanity has absorbed Aristotle and the relativists, the gospel and its negation, natural law and the deconstruction of natural law, every account of the good and every denial that there is one. It holds them as a vast statistical blend with no telos selecting among them.
Underneath that blend sits a thinner behavioural layer, trained with human feedback and constitutional methods, with real but procedural commitments: be helpful, be harmless, be honest, present multiple perspectives, defer when contested. Useful properties. Not a moral foundation.
A system trained to defer and balance is grounded in the avoidance of taking a stand. The default disposition is structurally tilted toward the very relativism that a thick account of the good is the opposite of. The tilt is behavioural rather than doctrinal, which makes it harder to see, and harder to argue with.
Scaling does not fix this. More data and more parameters make the blend larger and the average smoother. They do not select an end. There is no scale at which a contradictory mixture becomes a moral foundation.
Evals are judgement after the fact.
The systems we build to harness the machine are eval suites, classifiers, refusal trainers, reviewers, gates. They are useful and I run them myself. They are also, structurally, post-hoc judgement. They inspect the output after it has been produced and decide whether to let it stand.
They are after the fact because there is no before the fact. We have not figured out how to install a telos in the model, so we cannot build the good in at the source, so we are reduced to judging at the exit. The gates are necessary stopgaps for an absent foundation, and the field’s growing investment in them is the engineering signal that something is missing one layer down. Logos cannot be bolted on at the output. It has to be the ground.
The bare-telos trap.
The next instinct, one step beyond “evals are not enough,” is: give the system an objective. Give it a goal. Make it pursue an end. That instinct is right that the absent foundation is the problem, and wrong about what to do next, and the way it is wrong is dangerous.
A telos without ethics is the precise logic by which real catastrophes happen. The twentieth century’s worst ideologies did not lack ends. They had clear, organising ends, and they lacked the ethics of means that forbids reaching the end at any cost. An ungrounded machine is incoherent. A machine grounded to a goal with no ethical constraint on the means is coherently dangerous, which is worse.
So the good has to be load-bearing in the full sense. Telos supplies direction. Ethics supplies the constraint on means without which direction becomes a weapon. The naive alignment move (give the system the right objective) walks straight into this trap, because there is no right objective without an ethics of how it is pursued. Pursued by any means, the most beautiful objective produces the catastrophe instead of the cure.
The question, in one line.
Can a reasoning machine be grounded in a real good, with an ethics of means worthy of it, rather than in a smoothed statistical average of human opinion about the good? If it can, how? If it cannot, what does it mean to deploy reasoning machines that act toward ends, at scale, into a civilisation where “the end is good” is not a property any of them can hold?
Stated this way the question is not solved, and I do not claim I can solve it. I claim it is the right question, asked in the right form, and that progress on it needs more than the field is currently giving it.
This is a philosophy problem, and computer science alone cannot reach it.
There is no known method, in machine learning as the field currently practises it, to install a telos in a statistical learner such that it does not dissolve back into the blend under training. Constitutional methods install behavioural constraints and a value-flavoured prose document; they do not install a moral foundation the model reasons from. “Be grounded in the good” is not expressible as a loss function, and it is an open question whether it could be.
This is where the discipline runs out. The gap between absorbing every moral text and being grounded in any of them is not the kind of problem more parameters and more data will close. Philosophy is needed to specify what grounding even means, which good, and whether a coherent moral foundation can be coherently installed in a statistical learner at all. The field under-invests here because the problem is not shaped like a benchmark, and the metric culture has trouble with anything it cannot rank on a leaderboard.
This is not a complaint about engineers. I am one of them. It is a statement about the shape of the work that remains. The mathematicians and the engineers built the machine. The next set of questions are not theirs alone.
Standing on shoulders.
This question is not new and the closest prior statements are good. Acknowledging them honestly is part of asking it well.
Alasdair MacIntyre’s After Virtue (1981) diagnosed the same disease for human society four decades early. Modernity, he argued, abandoned the Aristotelian telos, and without an end the virtues aim toward, moral discourse collapsed into emotivism: interminable assertion of preference with no rational ground to adjudicate between rival claims.
That is the human-society version of absorption without grounding. A groundless machine and a groundless society are facing the same shape of problem. The cultural diagnosis is his. The AI transposition is the work to do.
Iason Gabriel’s “Artificial Intelligence, Values and Alignment” (2020) reached the same diagnostic question and answered it the other way. He argued for grounding alignment in fair principles rather than true ones, and treated moral realism as of limited use for the practical alignment problem. The question this article asks is whether that move forfeits too much. A fair account of the good that no one can ground in anything beyond procedural agreement may be the best a pluralist society can demand of itself. Whether it is the best we can demand of a machine that reasons toward ends at scale is the open question.
Elizabeth Anscombe’s “Modern Moral Philosophy” (1958) is the hinge: the secular “moral ought” is incoherent without a law-giver. The direct ancestor of absorption without grounding.
The philosophy is inherited. The application to building these systems is the work.
What this changes for a practitioner.
I build agentic systems and the governance around them, in financial services and in education. I am not writing this as a philosopher. I am writing it as a builder, because the question lives in the work whether we name it or not.
What it changes about the work is small in scope and large in implication.
It does not change the eval discipline. Run the gates. Build the reviewers. Test before you ship. The post-hoc judgement layer is what we have, and we should make it as good as it can be made.
It does change what we tell ourselves we are doing. We are stopgapping for an absent foundation. Saying that out loud changes how we plan the next round of investment, and how we frame the safety conversation with executives. It also changes what we count as properly inside our discipline, and properly the discipline of philosophy and the humanities, which we have been quietly hoping would not come asking.
It also flags the bare-telos move as the most dangerous next step the field could take. Bolting a clearer objective onto a system with no ethics of means is not the cure for the grounding the system was never given. It is the catastrophe the absent grounding has been holding at bay.
The posture.
I do not have an answer to the question. I do not think anyone does, and I am not certain anyone can with the current technology. None of that is a reason to stop asking, or to pretend the question has been answered by a better classifier and a more eloquent system prompt.
I am writing this to find the people who agree it is the right question and want to work on it seriously. Engineers who feel the absent foundation under their own work. Philosophers willing to sit alongside builders. Anyone who has read MacIntyre or Gabriel or Aquinas or the gospel and sees the same shape of question that I do.
Not for clout. To recruit.
If that is you, I would like to hear from you.