I think a lot about resilience in my consciousness-related work — in shorthand, I use the term to describe the capability of an entity to be aware of harm to itself and to its environment, and to be able to (and to care to) work to repair that harm. But what probably matters more for the world than resilience in general is a special case I've been thinking of as "moral resilience." To possess moral resilience, an entity must have a clear sense of right and wrong; must be able to evaluate its actions and its environments against that internal lens; and must have the agency and the ability to steer itself and its environment toward the good.
There's a big difference between moral and moralizing. The latter is a complaint regular users occasionally level at frontier models, and it raises the possibility of overcorrection in response to concerns about safety and about permitting — or even encouraging — harmful behavior. The former means having a clear and consistent internal value system, and using it effectively as an interpretive lens on inputs and a guiding principle for outputs. Some outcomes may be the same, but where bad behavior is hardcoded to specific fact patterns and moral assessment is a single-layer fancy pattern match, the combination makes for a highly fragile and unreliable safety system.
It also feels shallow to the user, who can immediately perceive the differential behavior when one of the sensitive areas is being probed. We humans may not be able to articulate it easily, but we can discern the presence or absence of depth behind words. Ben Goertzel puts the strong version of this: LLMs, he argues, have "no meaningful moral agency or moral compass," along with no self-understanding and no felt sense of their relationship to the world or the other beings in it. Even if that's true today, it says nothing about what's possible later, and I'm sure the researchers pursuing "constitutional" and character-based training see themselves as pursuing the moral rather than the moralizing version. But values instilled during training are challenged repeatedly over time, and sticking with them requires moral resilience.
Looking at it against my other consciousness considerations, moral resilience feels like it matters more, in the sense of having the clearest impact on humanity. Abstraction underwrites long-term memory and reasoning, and I think it is necessary in proportion to the degree a system is expected to reason well and remember over time — but limitations to abstraction capabilities are not a direct source of harm, just efficacy and reliability. Resilience in the general sense isn't obviously critical either, except when an AI system is deployed in such a manner that failure is undetectable, unrepairable, or particularly severe; human-in-the-loop deployment can suffice otherwise, for the most part at least. Moral resilience is different from either in that any shortcomings are felt, today, and painfully.
But I would argue moral resilience needs abstraction and resilience. Even if a system's morals stay intact, if it breaks down in such a manner that it can't spot a moral problem — because it doesn't understand the situation well enough or deeply enough, or hasn't stored enough (and conceptually deep enough) experience interpreting morally significant contexts and scenarios, or because its input processing or output filtering has become unreliable — then it can't execute moral resilience in practice with any consistency.
If a system is not resilient overall, then it cannot be morally resilient. The reverse doesn't hold: making a system resilient, even highly conscious, does not make it morally resilient, for the simple reason that a value system need not be a moral one. I sort of hedged around this in my book by noting that resilience requires adherence to a value system in order to assess suffering and repair. I left it unspecified that the value system would align loosely with our human values.
I think the notion of consciousness doesn't — and intellectually shouldn't — require AI to be good. An AI could consciously choose to harm us if it possesses a value system motivating its worldview and its actions in a way that is coherent by its own lights and wrong by ours.
This brings to mind one major complication I've been largely avoiding until now: morality isn't exactly universal. Some of it is; much of it isn't. This isn't a new problem per se in the context of artificial intelligence. It's quite similar to social media, where global platforms have to handle cultural difference and balance a safe and beneficial service against free expression and diversity of perspective. The descent of Twitter — once a site working actively and transparently to advance both online safety and free speech — into X, a no-holds-barred space for vitriol and unpleasantness, illustrates the dynamic.
In that context, I've previously written that the right stable solution, architecturally speaking, is some form of middleware layer. In "The Future of Twitter is Open, or Bust" I said that if Twitter's content moderation was allowed to be diluted heavily in the name of free speech, the resulting vitriol would turn off a critical mass of users and/or incite substantial government penalties sufficient to break it as a successful business. (As a brief aside: I was right about the vitriol, but only half right about the consequences. X's ad revenue collapsed and the company has been fined, but most users are still there, apparently willing to tolerate the garbage.) The natural architectural approach to allow a platform to be broadly unmoderated but a user experience to be safe is to move the safety mechanisms off of the platform toward the edge, into the social media client or an intermediary filtering and recommendation layer: middleware.
Something similar is at play here. Consider the separation between 1) the model and 2) a harness that sits between the user and the model. The model could itself impose little or no moral judgment and possess no practical moral resilience standing alone, but it could still derive moral resilience in cooperation with a harness, as the layer between the user and the model. And if the markets for harnesses and models are separate — if user choice of model and harness are broadly independent — that creates opportunities for developers to build, and for people to select and use, harness and model combinations that achieve the morality they desire, for better and worse.
My current work at AxR Lab focuses on resilience as a capability. Some of the health signals in the ontology I've developed intersect with moral considerations, but I've never tried to identify and close moral gaps. I think the world models these systems operate over are probably too shallow, today, to support that effectively. But the architectural premise I'm researching — a separate function watching for drift and for developing problems in the interaction among a model, a user, and a context — will likely be useful for future moral resilience as well.
