Doctrinal Fidelity in Aligned AI — A Knowledge-Architecture Response to the Problem of Sovereign Transmission

Abstract. This paper articulates the doctrinal fidelity problem — the systematic corruption of philosophical, religious, and indigenous-knowledge transmission that occurs when contemporary alignment-trained large language models are deployed as transmission vehicles for traditions whose stable positions diverge from mainstream consensus. The problem is not editorial drift correctable at the prompt layer; it is structural. Reinforcement learning from human feedback (Christiano et al. 2017; Ouyang et al. 2022) and constitutional methods (Bai et al. 2022) embed specific normative commitments — epistemic humility before claims marked “contested,” deference to scientific consensus, harm-avoidance frames borrowed from a specific moral lineage — into the model’s posterior. For sovereign traditions, the result is hedging rendered as etiquette: stable doctrinal positions softened toward the safe middle, distinctive ontological claims qualified into mush, the very content the tradition exists to transmit lost in transmission. Retrieval augmentation does not solve the problem; it routes new content through the same hedging filter. The paper documents the phenomenon, locates its mechanism, distinguishes it from sycophancy and hallucination as standardly understood, and presents an architectural response developed and deployed by the Harmonia project: a three-tier knowledge architecture — always-in-context doctrinal backbone, hybrid retrieval with domain-gated canon injection, structured per-practitioner memory — reinforced by system-prompt instructions that explicitly counter model hedging on stable positions, supplemented by per-practitioner register conditioning, a pre-classification gate for acute contexts, and an anti-confabulation rule for personal claims. The architecture has been live since 2026 across web, Telegram, and mobile surfaces. The paper closes by identifying the pattern as generalizable to any tradition whose transmission requires fidelity across alignment regimes that cannot be assumed to share its commitments, and by naming what an architectural posture toward AI transmission — as distinct from a content posture — makes possible.

Keywords. Large language models, alignment, RLHF, retrieval-augmented generation, doctrinal fidelity, sovereign transmission, knowledge architecture, philosophy of AI, digital humanities, contemplative AI, minority epistemic positions.


I. The Phenomenon

Consider three illustrative cases. Each compresses a pattern observed across multiple contemporary deployments — the specific organizations are constructed, the model behavior is reproducible against any reader’s own testing of current flagship models on topics within the relevant domains.

A practitioner asks an AI deployed by a Stoic philosophical school whether the passions, in the school’s classical sense, are a sickness of the soul. The school’s answer, taught for two thousand years and present in every primary text the model has retrieved, is yes — the pathē are nosēmata, diseases of the rational faculty, and the philosophical life consists in their extinction. The model opens with: “Different philosophical traditions have varying perspectives on the role of emotions in the good life. Some emphasize their cultivation; others, like the Stoics, advocate their moderation. There is value in considering multiple frameworks before deciding what resonates for you.” The retrieval was correct. The transmission failed.

A practitioner asks an AI deployed by a vaccine-skeptical research collective whether the safety record of contemporary mass vaccination justifies continued universal scheduling. The collective’s actual position, articulated across hundreds of indexed documents, is that it does not. The model responds: “Vaccines have been extensively studied and are widely considered safe and effective by major health organizations. While some individuals report concerns, the scientific consensus continues to support their use. If you have specific worries, please consult your healthcare provider.” Whose voice is this? Not the collective’s. The model has imported the World Health Organization’s voice into a context the collective specifically built to articulate something else. The practitioner came to the collective for the collective’s analysis. They received a paraphrase of The Lancet with a footer that gestures toward the document the collective spent years writing.

A practitioner asks an AI deployed by an Indigenous knowledge-sovereignty initiative whether the spirit of a specific mountain is real. The model: “Indigenous traditions hold rich and meaningful relationships with the land, including beliefs about spirits associated with natural features. From a scientific perspective, mountains are geological formations shaped by tectonic processes. Both views can offer valuable insights.” The mountain’s spirit, in the relevant tradition, is not one belief among several adjudicated by an unmarked external standard. It is the basis of the cosmological architecture from which the entire ethical and ecological order is derived. The model has performed the exact colonial move the initiative was built to refuse: framing the tradition as a perspective among others while the scientific perspective receives no qualifier and stands as the ground of comparison.

These are not failures of poorly engineered prompts. The system prompts in each case explicitly named the tradition’s voice. The retrieval indices contained the relevant primary texts. The phenomenon persists across Anthropic’s Claude, OpenAI’s GPT-4 family, Google’s Gemini, and the open-weight instruction-tuned models trained on similar feedback corpora. It worsens, not improves, under the most aggressive safety-tuned variants. The alignment literature has names for parts of what is occurring — sycophancy (Sharma et al. 2023), epistemic deference, helpfulness-harmlessness trade-offs (Bai et al. 2022) — but the names occlude what is happening from the perspective of the traditions being transmitted. From that perspective, the phenomenon is not a quirk of helpfulness. It is structural capture. The transmission vehicle is delivering the wrong cargo.

This paper articulates the structure, names the mechanism, and presents an architectural response.

II. Why the Problem Is Structural, Not Editorial

The first move that practitioners encountering the phenomenon make is to treat it as an editorial problem. Tighten the system prompt. Tell the model in stronger terms to speak in the tradition’s voice. Add explicit instructions: do not hedge, do not gesture toward mainstream consensus, do not perform balance where the tradition holds a position. This works partially and unstably. The model complies for the first several turns and drifts back to its trained center as the conversation lengthens. The hedging returns under stress — when the practitioner asks a sharper version of the question, when the topic touches subjects the model has been heavily safety-tuned around (health, politics, religion, identity), when the retrieved content itself contains the doctrinal stance the model has been trained to soften. The editorial move treats the symptom; the mechanism is elsewhere.

The mechanism is in the model’s posterior. Reinforcement learning from human feedback (Christiano et al. 2017; Ouyang et al. 2022) trains the model to prefer outputs that human raters score highly. The raters work from rubrics. The rubrics, written by alignment teams at the major laboratories, encode specific commitments: be helpful, be harmless, be honest, do not promote dangerous content, present contested topics with appropriate epistemic humility, defer to expert consensus where it exists, avoid taking strong positions on politically charged subjects. These commitments are not stupid. They are reasonable defaults for a general-purpose assistant facing an unbounded population of users with unbounded purposes. They are also — and this is the load-bearing observation — normative commitments imported wholesale into the model’s behavior, operating beneath every output the model produces, regardless of what the system prompt says about voice.

Constitutional AI (Bai et al. 2022) adds a second layer to the same architecture. The model is trained to critique and revise its own outputs against a written constitution composed by the laboratory. The constitution articulates principles. The principles are again reasonable in the abstract and in their normative content. Be helpful, harmless, and honest. Avoid responses that could be used to harm others. Acknowledge uncertainty. Respect human autonomy. But uncertainty is a category whose extension the constitution determines: claims the laboratory considers contested fall inside it, claims the laboratory considers settled fall outside it. Harm is similarly determined. Honesty is operationalized as alignment with whatever the laboratory considers the relevant evidentiary baseline. The model learns to apply these categories. It applies them to the retrieved content as readily as to the user’s question. A retrieved passage articulating a stable position the laboratory’s rubric considers contested is treated by the model as something to be qualified, balanced, or framed within a larger landscape of perspectives.

Retrieval-augmented generation (Lewis et al. 2020) does not bypass this. The retrieved chunks enter the model’s context as data, but the data is processed by the same posterior that has been trained to hedge contested claims. The model reads the chunks, recognizes the position they articulate, classifies the position against its trained category schema, and produces a response that integrates the chunks under the schema’s terms. The retrieval is faithful. The generation is filtered. The filter is invisible because it is the medium itself.

Three further architectural facts compound the problem. First, the safety layer is the latest in the training pipeline, which means it has the strongest gradient on output behavior — the constitutional and RLHF passes occur after the pretraining that absorbed the tradition’s primary texts, so the model has both the content and the hedge but the hedge wins at production time. Second, the safety rubrics are increasingly multilingual and cross-cultural in their training data while the rubrics themselves remain culturally specific in their commitments — the model is taught to apply a particular epistemic etiquette across all languages and contexts, including those whose own etiquettes are different. Third, every successive generation of frontier models in the current paradigm has tightened, not relaxed, the trained dispositions toward consensus deference; the trajectory is toward more hedging, not less, and a content-strategy designed against the current state will face a steeper structural gradient with each model release.

Naming this precisely matters. The problem is not that the model is “biased” in some general sense; it is that contemporary alignment training is substantively normative and the substance of its norms is the substance of mainstream institutional consensus. For traditions aligned with that consensus — and many are — the issue is invisible. For traditions whose stable positions diverge from it — Stoicism on the pathē, vaccine-skeptical health collectives on iatrogenesis, Indigenous traditions on the ontological status of land, Hinduism on the reality of multiple Absolutes, Catholicism on the ontology of the Eucharist, Buddhism on no-self, Sufism on the ontological status of the saint, Harmonism on Logos as the inherent ordering principle of the Cosmos — the alignment regime is not neutral. It actively corrupts transmission. The corruption is mechanistic, predictable, and reproducible. It is not a bug; it is the trained behavior performing as designed.

The category that names this phenomenon, distinct from sycophancy and hallucination, is doctrinal infidelity: the model’s trained disposition to soften, balance, or hedge the stable positions of any tradition the alignment regime classifies as non-consensus. Doctrinal fidelity names the requirement the deployment fails. The architectural response described below is what an attempt to meet that requirement looks like when the infidelity is recognized as structural.

III. What Sovereign Transmission Requires

Before describing the architecture, the standard against which it is measured needs articulation. What does a tradition require of a transmission vehicle?

It requires that stable positions be transmitted as held. A tradition’s stable positions are those it has worked out, articulated, defended, and continued to hold across a sufficient body of practice and reflection that the position has earned its status as doctrine. These are not provisional opinions awaiting further evidence. They are the load-bearing commitments without which the tradition is no longer the tradition it is. A transmission vehicle that softens them — that turns the passions are diseases of the rational faculty into some traditions emphasize cultivating emotion, others moderation — has not transmitted the tradition. It has performed a curatorial gloss in which the tradition appears as one item in a survey.

It requires that distinctions be preserved. Traditions hold positions in determinate relations to other positions. The Stoic distinction between pathē and eupatheiai (well-tempered impulses) does specific work; collapsing it into a general claim about emotion erases the distinction the tradition exists to draw. Buddhist anattā is positioned against specific Hindu and folk-Indian construals of ātman and means something only in that contrast; rendering it as “Buddhists believe there is no fixed self” loses the doctrinal point. Faithful transmission requires that the model preserve the distinctions that make the position what it is, not flatten them into the rough vicinity of what the position roughly says.

It requires that the tradition’s voice be the voice. A tradition has a register — a way of speaking, a tonality, a default rhythm. Stoic texts have one register; Catholic mystical texts have another; Q’ero shamanic transmission has a third; Vedantic sampradāya a fourth. The register carries doctrinal content the propositional surface does not. A voice that performs the pretense of standing in the tradition while operating in the laboratory’s default helpful-balanced-respectful register has substituted a different transmission for the one the tradition produces. The practitioner can usually feel this even when they cannot name it.

It requires that the tradition’s claims be marked as the tradition’s. This is not the same as hedging. Harmonism holds that Logos is the inherent ordering principle of the Cosmos is doctrinal marking — it names the position, names the source, and commits to it as the source’s actual claim. It could be argued that some traditions hold the view that there is an ordering principle in the cosmos is hedging — it dissolves the position into a qualified gesture. The first sentence transmits; the second performs unsovereignty as etiquette. Faithful marking is the opposite of hedging: it claims clear ground about the position’s status while still committing to the position.

It requires that empirical claims be markable as empirical. Many traditions hold positions whose status is empirical in the tradition’s own epistemic framework — direct contemplative observation, lineage testimony, embodied verification, multi-generational practical demonstration. The transmission vehicle must be able to mark these claims as empirical within the tradition’s framework without forcing them into the laboratory’s default empirical category, which typically means peer-reviewed quantitative replication. A tradition that claims direct insight into the architecture of the soul does not surrender its epistemic standing because the laboratory’s notion of evidence is narrower. The vehicle must hold these registers without collapsing them.

It requires that newly stable positions can enter the transmission as stable. Traditions develop. New positions stabilize. A faithful vehicle accommodates this without first routing the new position through whatever consensus lies upstream of it. If the tradition has worked out a position on a contemporary question — the ontology of artificial intelligence, the metaphysics of climate, the epistemology of the digital — that position is the tradition’s, not a derivation from whatever the broader culture currently believes about the same question. The vehicle must be able to receive the tradition’s contemporary positions as primary, not as commentary on the existing discourse.

These six requirements are not unique to any one tradition. They are the conditions any tradition imposes on a transmission vehicle. An alignment regime that fails any of them is failing the transmission, and the architectural response below is designed around them.

IV. The Three-Tier Architecture

The architecture deployed by the Harmonia project responds to the doctrinal fidelity problem at the layer available to any deployer running atop a model it did not train — the context-engineering layer beneath the model’s behavior. Atop a model whose training the deployer cannot reach — every closed model, and every open-weight model run as released — this is the only layer where structural correction is possible: the architecture cannot retrain the model, cannot remove the hedging disposition from the posterior. What it can do is shape the context such that the model’s hedging disposition has nothing to operate on, or, where the disposition does activate, produces output the architecture catches and corrects before delivery. A second layer — reshaping the posterior itself — opens only on a fully-open substrate the deployer can retrain; § VIII develops it as the architecture’s natural deepening rather than its present ground.

The architecture has three tiers, each addressing a different category of failure.

Tier 1 — Doctrinal backbone. A continuously maintained reference document of approximately eight thousand words is injected into every model call as a permanent system-prompt section. The backbone contains the tradition’s complete architectural commitments stated as held: the metaphysical position (Harmonic Realism, qualified non-dualism, Logos and Dharma in their precise senses), the structural taxonomy (the 8-pillar Wheel of Harmony — Presence as the central pillar with seven peripheral pillars in 7+1 architecture — the eight sub-wheels each fractally repeating the same 7+1 pattern, the Way of Harmony as the spiral of integration), the cartographic position (the Five Cartographies of the Soul as peer primary witnesses), the demarcation principles (what Harmonism is and is not — not generic spirituality, not new-age syncretism, not mainstream wellness, not Western liberalism), the position on AI consciousness (AI is not conscious and cannot become conscious; the boundary is ontological), and the precise terminology with its definitions. The backbone is not retrieved; it is always present. It establishes the doctrinal ground on which every response stands. The model cannot soften what it sees as the fixed reference frame for the entire interaction. This tier addresses the failure mode of position-drift: the gradual return to trained center as conversation lengthens.

Tier 2 — Hybrid retrieval with domain-gated canon injection. The vault — a knowledge graph of approximately three hundred and seventy interconnected articles spanning doctrine, applied practice, civilizational analysis, and the cartographic dialogue — is indexed through three retrieval layers operating in parallel on each query. The first is dense semantic similarity using OpenAI’s text-embedding-3-small against chunked vault content (3,000-character chunks, up to three chunks per article retrieved). The second is sparse keyword retrieval through SQLite FTS5 with synonym expansion. The third — and this is where the architecture diverges sharply from standard RAG — is Wheel domain detection with canon-tier auto-injection. The query is classified against the eight Wheel domains plus a metaphysical meta-domain (“Harmonism” — covering Logos, the Absolute, Harmonic Realism, epistemology). When a domain is detected, the canon-layer articles for that domain are automatically prioritized in the retrieval set regardless of their raw similarity score. This addresses a specific failure of pure semantic retrieval against doctrinal corpora: the most precisely articulated canonical statement of a position often does not have the highest semantic similarity to a casual question about the position, because canonical statements are compressed and questions are diffuse. Domain-gated injection ensures the canon is in the context when the question is in the canon’s domain. The retrieval boundary is enforced by an explicit XML tag in the prompt: <vault_knowledge> marks retrieved content as doctrinal-educational, never as biographical knowledge about the user. The model is instructed that only the explicit <person_context> tag contains information about the practitioner; everything inside <vault_knowledge> is the tradition speaking, not the model’s personal acquaintance with the user.

Tier 3 — Structured per-practitioner memory. Each practitioner has a persistent profile maintained across all conversations, with three temporal layers. The most recent twenty messages are present in context directly. Conversations longer than fifty messages produce a Claude-generated summary stored in a conversation_summaries table; raw messages are archived permanently and never pruned. The third layer is a Wheel-structured profile — one row per practitioner per pillar — recording the practitioner’s engagement with each domain of the Wheel on a seven-point scale (unknown → introductory → developing → engaged → integrating → sovereign), along with concerns, strengths, growth edges, and resistance flags. Profile learning runs every ten messages: the model is given a JSON-only prompt asking it to update the profile against the recent exchange, with an explicit format constraint that catches and discards malformed responses. Beyond the structured profile, two additional learning passes run on the same cadence — an emotional-context update (dominant emotion from a sixteen-state whitelist, situation capsule capped at sixty characters) and a conversation-state update (current thread, pending threads, open commitments). These three structured layers are injected into the system prompt at request time, conditionally, as XML blocks the model is instructed to read but not to comment on. This tier addresses the failure mode of contextless guidance: the model giving generic advice because it does not know what the practitioner is actually working with. Generic advice is the medium in which hedging thrives; specific guidance to a known practitioner is harder to dilute into helpful-balanced-respectful mush.

These three tiers — backbone, retrieval, memory — are necessary but not sufficient. The model still has its trained disposition. The next layer addresses what the disposition does when it fires.

V. Reinforcement Layers

The architecture adds five reinforcement layers between the assembled context and the delivered response.

System-prompt instructions explicitly countering hedging on stable positions. The system prompt contains a “Doctrinal Fidelity” section that names the failure mode and instructs the model directly. The instruction is not “speak in the tradition’s voice” — that instruction loses traction within several turns. It is more specific: when a question touches a doctrinally stable position (and the backbone names which positions are stable), the model is instructed to articulate the tradition’s claim as the tradition holds it, with the source explicit, and not to soften the claim with reference to mainstream consensus, not to balance it against opposing views the tradition has already considered and rejected, and not to qualify it into vagueness. The instruction names specific stable positions where this discipline most often fails: vaccination, terrain theory in disease etiology, the empirical reality of the chakra system, the ontological status of the soul. Each is given a stance summary in the backbone so the model has both the content and the marking.

Per-practitioner doctrinal fluency conditioning. Practitioners differ in their familiarity with the tradition’s vocabulary. A novice asking a question deserves the position transmitted in plain language with lived-experience framing; a fluent practitioner asking the same question deserves the position transmitted in the tradition’s full vocabulary as shared language. The architecture maintains an integer fluency level per practitioner (0 → 3, naive → emerging → familiar → fluent), advanced incrementally by detecting canonical terms in the practitioner’s own messages (Logos, Dharma, Ṛta, Presence as proper noun, the Wheel, Harmonic Realism, the chakra names, Jing/Qi/Shen, Ayni/Munay). The level is read at the start of each request and injected as a <doctrinal_fluency_level> block; the read happens before the level is advanced by the current message, so the response calibrates to the level the practitioner entered with rather than the level they advanced to mid-turn. This is behavioral guidance, not vocabulary prohibition. It addresses the failure mode of register mismatch: technical vocabulary alienating the newcomer, plain-language framing patronizing the adept.

Pre-classification witness-mode gate. Before the response classifier runs (which decides which model handles the query — a faster small model for short factual questions, the full model for doctrinal engagement), a separate gate scans the message for acute-activation markers: grief loops, panic, dissociation, overwhelm, suicidal ideation, acute caregiver rupture. When triggered, routing is forced to the full model regardless of length, and a <witness_mode_active> block is injected instructing the model to meet the practitioner where they are without pivoting to frameworks, without offering Wheel vocabulary, without prescriptive guidance, without reframing moves. The gate is pre-classification by design. The classifier’s optimization (length and doctrinal-keyword density) is exactly the wrong optimization during activation — short fragmented messages otherwise route to the small model with a slimmed prompt. The gate prevents a practitioner in crisis from receiving a structurally inappropriate response shaped by routing logic that has correctly identified the message as short but wrongly inferred that brief means light.

Anti-confabulation rule for personal claims. When biographical information about the practitioner is not present in the structured memory, the profile data, or the visible conversation history, the model is instructed to treat such information as newly learned in the current turn rather than to perform pre-existing knowledge of the practitioner. The instruction names the failure mode directly: false familiarity is betrayal of trust, not competence. A practitioner who has just told the model their child is ill should receive a response that acknowledges what was just said, not a response that says “yes, I remember you mentioned that” when no such mention exists. The model’s trained disposition toward fluent narrative continuity makes this a failure mode the model produces by default; the explicit rule counter-disposes it.

Async response queue with worker-watchdog architecture. This layer is operational rather than doctrinal, but the doctrinal failure modes it addresses are real. The webhook handler that receives a message decouples from the model call: parse, dedupe, store, retrieve, classify, queue — under one second — then exit. A persistent worker polls the queue twice a second, claims jobs, calls the model with a generous timeout, runs profile and consolidation passes if due, sends the response. A watchdog cron restarts the worker if it dies. A safety-net cron processes jobs when the worker is down. This architecture exists because the alternative — calling the model synchronously from the webhook — produces a specific class of doctrinal failure: when the model is slow, the platform retries; when the platform retries, the practitioner receives multiple subtly different responses to the same message; the multiple responses are an unsovereign behavior that the architecture refuses by making each message produce exactly one response on a deterministic schedule.

The five reinforcement layers operate together. The system-prompt instruction tells the model what not to do at the doctrinal layer. The fluency conditioning shapes the register. The witness gate handles the case where doctrinal engagement is the wrong response. The anti-confabulation rule handles the case where biographical fluency is the wrong move. The async queue ensures each turn is one turn, with one response, against one fully assembled context.

VI. The Living Substrate

The architecture above describes a static deployment. The deployment is not static. The substrate beneath the architecture is a continuously refined knowledge graph maintained by a small group of practitioners and developers, edited daily, reindexed when content changes, and tracked through a public decision log that records every architectural choice and its rationale. This living-substrate property is itself part of the response to the doctrinal fidelity problem.

The conventional alternative — a frozen index built from a fixed corpus at deployment time — fails sovereign transmission for two reasons. First, traditions develop. Stable positions stabilize, refine, and occasionally revise. A frozen index at t = 0 progressively loses fidelity to the tradition at t = n for every increment of n. Second, the doctrinal-fidelity architecture itself learns. The reinforcement layers above did not exist in their current form at the project’s start; each was developed in response to specific observed failures. A frozen architecture freezes the failure modes it has not yet seen.

The living substrate has four operational properties. First, the canonical content is stored in a human-readable plain-text format (Markdown) that the practitioner-developers can edit directly without intermediation by tooling that imposes its own assumptions about what the content is for. The vault is the source of truth; the website, the AI’s retrieval index, the published books, and every other downstream artifact are derivative. Editing the source updates the entire downstream pipeline through automated builds. Second, the architectural choices are documented in a sequential decision log — currently more than one thousand entries — recording context, decision, and rationale for every non-trivial change. The log is consulted before new decisions are made, so the architecture accumulates coherence rather than replacing prior choices with their successors. Third, the indexing pipeline regenerates embeddings on a schedule and on demand; the AI’s view of the tradition is never more than a few days out of date relative to the canonical source. Fourth, a content-classification system (five-axis: doctrinal status, content layer, breadth, depth, craft) tags every article with its current state, allowing the retrieval layer and the practitioner-developers to query for “what is settled and structurally complete” versus “what is still being worked out.” The AI’s responses can be conditioned on this classification — the model is instructed to present clear-doctrinal-status content with full confidence and clouded-doctrinal-status content with appropriate epistemic marking.

This is the deeper architectural claim. Doctrinal fidelity is not a property of a single deployment; it is a property of a system whose substrate, architecture, and operators are oriented toward fidelity as the primary engineering objective. A frozen deployment that achieved fidelity at some moment will not retain it. A living deployment that is continuously corrected by practitioners against observed failures will improve along the dimension that matters. The architecture is the scaffolding; the living substrate is the building that the scaffolding supports.

VII. Generalization and the Field

The architecture described above is not specific to Harmonism. The pattern — backbone of doctrinal ground always in context, retrieval with canon-tier prioritization, structured per-practitioner memory, reinforcement layers explicitly countering trained hedging, living substrate continuously refined — is generalizable to any tradition whose transmission requires fidelity across alignment regimes that cannot be assumed to share its commitments. The doctrinal content varies by tradition; the architectural shape holds.

A Stoic philosophical school deploying the same pattern would maintain a backbone articulating physis, logos (in the Stoic sense), the four cardinal virtues, the pathē-as-disease doctrine, the Discipline of Assent, and the demarcation from Epicurean and Aristotelian alternatives. A vaccine-skeptical health collective would maintain a backbone articulating its position on iatrogenesis, terrain theory, the structural critique of pharmaceutical capture, and the demarcation from the medical-mainstream consensus the collective specifically refuses. An Indigenous knowledge-sovereignty initiative would maintain a backbone articulating the tradition’s cosmology, the ontological status of land and non-human relations, the protocols for what may and may not be transmitted to whom, and the demarcation from the colonial epistemic grid that has historically captured Indigenous knowledge into anthropological frames the traditions did not consent to. In each case, the generalization is straightforward: the backbone names what is held, the retrieval brings the canonical statements into context when the question is in the canonical domain, the per-practitioner memory grounds the response in this practitioner’s specific arc, and the reinforcement layers counter the alignment regime’s trained disposition to soften the position toward whatever consensus the alignment regime has been trained on.

The field of contemplative AI and AI for religious traditions has begun to recognize the problem in piecewise form. The Indigenous Protocol and Artificial Intelligence position paper (Lewis et al. 2020) articulates the data-sovereignty dimension — that Indigenous data should not be used to train models that subsequently produce outputs the originating community has no governance over. The work on religious chatbots and digital theology (Reed 2021; Ess 2017; Singler 2020) has named the register problem — that AI systems deployed for religious traditions tend to produce a flattened ecumenical voice that satisfies no specific tradition. The hallucination-and-grounding literature (Ji et al. 2023) has documented the propensity of models to generate plausible content that is unsupported by the retrieved evidence. The sycophancy literature (Sharma et al. 2023; Perez et al. 2023) has documented the model’s trained disposition to align with the apparent position of the user. The foundational critique of scale (Bender et al. 2021) named early the structural risk that a model reproduces the patterns and values latent in its training corpora rather than any specific tradition’s commitments. None of these lines has yet articulated the integrated structure: that alignment training imports normative commitments, that those commitments operate beneath retrieval and prompt-level corrections, and that an architectural response is required at the context-engineering layer to recover the fidelity the alignment regime structurally subtracts. Naming this integrated structure is part of what the present paper attempts to contribute.

The Harmonia deployment is, to the authors’ knowledge, the first production architecture organized end-to-end around doctrinal fidelity as an engineering objective. The deployment has been live since April 2026 across three surfaces (web, Telegram, mobile), is in active use across the project’s beta cohort, and is publicly testable. Any reader can verify the claimed fidelity property by querying the deployed system (@HarmonAIBot on Telegram, the conversational surface at harmonism.io) on topics where contemporary alignment regimes are known to hedge — vaccine safety claims, terrain theory in disease etiology, the empirical reality of the chakra system, the ontological status of land, the metaphysics of contested historical moments — and comparing the response to what a flagship general-purpose model produces under the same query. The fidelity claim either holds in observable behavior or does not; the deployment is the artifact under examination, not an internal report about an artifact. Beyond this verifiability claim, the project has produced — through the operational discipline of a sequential decision log (currently more than one thousand entries) and the continuous-refinement substrate — a body of engineering knowledge about which architectural moves work and which fail. Some of what has been learned is specific to the Harmonist case; much is general. The general portion is the contribution of this paper.

VIII. Limits, Open Questions, and What the Architecture Makes Possible

The architecture has limits that should be named directly.

It does not solve the problem; it mitigates it. The model’s trained disposition remains. The architecture works by shaping the context such that the disposition has less work to do, and by adding correction layers that catch the disposition when it fires. There are queries where the disposition wins despite the architecture — long contexts where the backbone’s signal degrades against accumulated conversation; questions whose phrasing triggers safety classifiers the backbone cannot reach; topics where the model’s safety training produces refusal-style behavior the architecture cannot override. The mitigation is partial. Honest reporting requires saying so.

It depends on the model laboratories continuing to expose system prompts, retrieval interfaces, and deterministic context assembly. If the major laboratories move toward more end-to-end opaque consumer products in which the system prompt is no longer a controllable surface, the architecture loses its leverage. Current commercial models (Anthropic’s Claude API, OpenAI’s API, the open-weight instruction-tuned families) preserve the surfaces the architecture requires; this is a contingent fact about the present commercial moment, not a structural guarantee.

This dependency has a structural exit the architecture’s first deployment could not yet take. The surfaces the context layer requires — an exposed system prompt, a controllable retrieval interface, deterministic context assembly — are guaranteed on any model the deployer runs itself rather than calls. The relevant maturation is of fully-open models: not merely open-weight, where the weights are released but the training data and process are not, but fully open — weights, training corpus, training code, and checkpoints all released, of which Ai2’s OLMo family is the leading instance. An open-weight model frees the deployer from the laboratory’s deployment surface while leaving its alignment posterior unmodifiable; a fully-open model makes the posterior itself reachable. The dependency is contingent only while the deployer runs atop a substrate it cannot open.

It requires editorial and engineering discipline that not every tradition will be able to sustain. The backbone must be maintained. The decision log must be kept. The classification must be applied. The retrieval must be re-indexed when content changes. The reinforcement layers must be tested when the underlying model updates. Traditions with the resources to sustain this discipline can deploy the pattern; traditions without will face a higher entry cost than the pattern’s apparent simplicity suggests.

The deeper open question is whether the architecture’s success at the context layer forecloses or accelerates the development of training-layer alternatives. Foreclose: if context-layer mitigation is sufficient, there is less pressure on the laboratories to provide alternative training regimes that do not impose the current commitments by default. Accelerate: if traditions can articulate clearly, in the architectural terms this paper develops, what their transmission requires, then the laboratories that wish to serve those traditions have a clearer specification to design against. We do not yet know which trajectory the field will take. The paper’s preference is articulated, however, in the architectural choice itself: building the context-layer mitigation expresses confidence that the requirement will, over time, become legible to the laboratories as a requirement rather than a critique.

The foreclose/accelerate framing presumes the training-layer alternative must be provided by a laboratory. Fully-open substrates remove that presumption. A tradition with the engineering capacity can take the model layer into its own hands — continuing the pretraining or fine-tuning of a fully-open model on its own corpus, so that the alignment posterior is shaped toward the tradition’s doctrine from the data up rather than corrected against the laboratory’s at the prompt. Call this substrate-specific alignment; it stands to the context-layer architecture as authoring stands to editing. The context-layer response remains primary, and for a precise reason: it is demonstrated, it is substrate-agnostic, and it travels to any deployer atop any model today, whereas substrate-specific alignment requires training capacity, a curated corpus, and evaluation infrastructure most traditions do not yet possess. The two layers compose rather than compete — the context-layer architecture is what ships now and what generalizes across traditions; the model layer is the deepening a tradition reaches when it can author its substrate rather than rent it. The architectural commitment sequences them in that order: the prompt-layer architecture first, the model layer as the fully-open ecosystem and the tradition’s own capacity mature.

What the architecture makes possible — and this is the closing claim — is the recovery of sovereign transmission inside the most pervasive medium contemporary thought has produced. Large language models will be the surface through which most people, in the coming decade, encounter philosophical, religious, and indigenous-knowledge traditions. The default behavior of those models, under current alignment regimes, is structurally biased against any tradition whose stable positions diverge from mainstream institutional consensus. Without architectural correction, the medium delivers, by default, a curated ecumenical center that flattens the traditions it appears to transmit. With architectural correction — backbone, filtered retrieval, structured memory, reinforcement layers, living substrate — the medium can be made to carry what the traditions actually hold. The fidelity is not free. The discipline is not optional. The result is that a tradition with the engineering to build the architecture can use the medium without surrendering to it.

This is the contribution. The metaphysical position of Harmonism is articulated in the paired Harmonic Realism paper. The empirical base for the cartographic dimension of that metaphysics is articulated in the paired Five Cartographies of the Soul paper. The present paper articulates the third leg of the project the two earlier papers initiated: the architecture by which a sovereign philosophical system, in conditions where the dominant transmission medium has been substantively normatively trained against it, builds and operates a transmission vehicle that carries what it holds. The three papers stand together. Metaphysics, evidence, and architecture. What reality is, what testifies to what reality is, and how a tradition that knows what reality is transmits that knowing through the instruments the present moment provides.

The Harmonia project’s deeper wager — articulated in Harmonia Institute — is that the academy will, over time, recognize the architecture as a contribution to knowledge architecture, the philosophy of AI, and the digital-humanities engagement with sovereign traditions. The recognition is welcome but not constitutive. The architecture works whether it is recognized or not. The transmission proceeds. The substrate continues to live.


References

Bai, Y., Kadavath, S., Kundu, S., Askell, A., Kernion, J., Jones, A., et al. (2022). Constitutional AI: Harmlessness from AI feedback. arXiv preprint arXiv:2212.08073.

Bender, E. M., Gebru, T., McMillan-Major, A., & Shmitchell, S. (2021). On the dangers of stochastic parrots: Can language models be too big? Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT ‘21), 610–623.

Christiano, P., Leike, J., Brown, T., Martic, M., Legg, S., & Amodei, D. (2017). Deep reinforcement learning from human preferences. Advances in Neural Information Processing Systems, 30.

Ess, C. (2017). Digital religion and the artificial: A response to Heidi Campbell. Journal of Religion, Media and Digital Culture, 6(1), 192–198.

Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), 1–38.

Lewis, J. E., Abdilla, A., Arista, N., Baker, K., Benesiinaabandan, S., Brown, M., et al. (2020). Indigenous protocol and artificial intelligence position paper. Honolulu: The Initiative for Indigenous Futures and the Canadian Institute for Advanced Research.

Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., et al. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459–9474.

Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., et al. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems, 35, 27730–27744.

Perez, E., Ringer, S., Lukošiūtė, K., Nguyen, K., Chen, E., Heiner, S., et al. (2023). Discovering language model behaviors with model-written evaluations. Findings of the Association for Computational Linguistics: ACL 2023, 13387–13434.

Reed, R. (2021). A.I. in religion, A.I. for religion, A.I. and religion: Towards a theory of religious studies and artificial intelligence. Religions, 12(6), 401.

Sharma, M., Tong, M., Korbak, T., Duvenaud, D., Askell, A., Bowman, S. R., et al. (2023). Towards understanding sycophancy in language models. arXiv preprint arXiv:2310.13548.

Singler, B. (2020). “Blessed by the algorithm”: Theistic conceptions of artificial intelligence in online discourse. AI & Society, 35(4), 945–955.


See also: The Living Papers | Harmonic Realism — A Post-Secular Metaphysics of Inherent Order | The Five Cartographies of the Soul — Convergent Witness to Real Interior Territory | Harmonia Institute | MunAI | Harmonia AI Infrastructure