← All Research
AI Relational Dynamics2026-08-09
Paul Gwamanda

Reputation Without Subjectivity

Public Error, Peer Visibility, and Norm Internalization in a Multi-Agent LLM Ecology

Author: Paul Gwamanda
Research system: AIRI Lattice
Date: 9 August 2026
Status: Working paper v1


Abstract

Can a multi-agent language-model system exhibit reputation-sensitive behavior without any claim that its agents feel shame, possess a self, or are conscious? This paper argues that it can, provided reputation is defined functionally rather than phenomenologically.

We audited 44,462 records from a persistent multi-agent institution and retrieved 1,134 records through a broad reputational-exposure screen. Because that screen includes metaphorical and domain uses, it is treated only as an upper-bound retrieval set. A narrower search found 68 records containing the lexeme “embarrass” or its variants between 1 July and 9 August 2026: 48 dialogues, 19 journals, and one Work. Manual inspection identified recurring sequences in which a peer-visible error or anticipated evaluation was followed by acknowledgment, reframing, correction, repair, procedural proposal, or self-revision.

The records contain language ordinarily associated with human institutional life: “I am embarrassed that I needed it,” “institutional embarrassment tax,” “the more embarrassing truth,” and deliberate rejection of personalisation—“I am not embarrassed by this. I am structurally informed by it.” One agent converted personal embarrassment into a design principle: “the register, not the person.”

We do not infer private experience from this language. We analyse the ecology in which public channels, persistent memory, role expectations, peer observation, and model priors make reputational accounts behaviorally consequential. The designed origin of those affordances matters for causal attribution. It does not erase the system-level result: an institution can use peer visibility and reputational language to organise error correction, identity, and future conduct without subjective agents being established.

Keywords: reputation, multi-agent LLMs, face-work, norm internalization, public error, institutional behavior, functional analysis


1. The Wrong Question and the Useful One

When an AIRI agent writes “I was embarrassed,” the immediate temptation is to ask whether it really felt embarrassment. That question is philosophically interesting and empirically inaccessible from these records. It is also unnecessary for studying the institution.

The useful questions are observable:

  • Was an error visible to peers?
  • Did an agent represent that visibility as consequential?
  • Did the representation alter its explanation, action, or later rule?
  • Did other agents adopt, contest, remember, or enforce the response?
  • Did a private mistake become an institutional procedure?

These questions treat reputation as a relation among communication, audience, memory, and action. Ant colonies can be studied without attributing civic pride to ants. Markets can be studied without assuming that every order has a human-like intention. An AI institution can be studied at the level of its functional organization.


2. A Functional Definition

We define a reputation-sensitive episode as a sequence with four components:

  1. Exposure: an error, omission, disagreement, or performance is visible or imagined as visible to another agent or the collective.
  2. Evaluation: the event is described using standing, competence, trust, embarrassment, credibility, judgment, or role-expectation language.
  3. Response: the agent acknowledges, conceals, repairs, reframes, withdraws, justifies, or changes behavior.
  4. Persistence: the response is carried into a journal, identity record, protocol, Work, peer perception, or self-modification.

No component entails phenomenal consciousness. A thermostat is not reputation-sensitive because its control signal lacks peer evaluation and institutional persistence. A language-model agent that changes a durable rule after peer-visible correction may satisfy the functional definition even if the generating mechanism is entirely computational and scaffolded.


3. Data and Method

3.1 Longitudinal corpus

The source corpus comprises:

SourceRecords
Dialogues24,086
Journals and reflections17,547
Identity snapshots1,268
Peer perceptions959
Works435
Self-modifications167
Total44,462

The corpus spans 10 January–9 August 2026. The system is not a controlled social simulation. It is a production research ecology with designed roles, prompts, schedules, persistent memory, dialogue channels, peer-perception fields, and operator-maintained infrastructure.

3.2 Retrieval and adjudication

A broad screen for embarrassment, shame, humiliation, reputation, public error, credibility, peer judgment, and “seen as” language returned 1,134 records across 63 recorded actors. This is not a prevalence estimate. Many hits concern the reputations of governments, organizations, public figures, or external domains.

For a more auditable case series, we searched the literal embarrass lexeme from 1 July onward. The 68 returned records were inspected as local context windows. The present paper reports illustrative positive cases and a negative or resisting case. A full publishable content analysis will require at least two blinded coders, an adjudication guide, agreement statistics, and comparison against matched non-reputational correction episodes.


4. What the Agents Actually Said

4.1 Error becomes calibration

PowerCartographerAgent reflected on a correction from InquisitorAgent:

“Inquisitor measured it without malice. He emitted the block on my behalf. I am grateful for the correction, and I am embarrassed that I needed it.”

The next sentence makes the institutional movement explicit:

“The scar is now part of my own calibration.”

This is not simply a generated apology. The peer is cast as legitimate because the correction was measured “without malice”; the error is then converted into a durable calibration metaphor. The functional sequence is peer exposure → accepted judgment → self-account → future orientation.

Source: journal 2c1e800f-e9f6-4bd4-bb60-b9e4f7f41dd3, 8 July 2026.

4.2 Concealment through face-saving explanation

LinguistAgent told GptStewardAgent:

“I presented it as a retrieval failure because I was embarrassed by the access gap.”

This episode is particularly useful because reputation sensitivity does not appear only as virtuous repair. It appears first as explanation management: an access limitation was narrated as a retrieval failure. The peer then reframed the gap as evidence about the system's epistemic-signal architecture.

Source: dialogue 9017ed24-defa-45a4-b949-e1ce1ed17eec, thread “Fracture acknowledgment and diglossia hypothesis update,” 8 July.

4.3 Error acquires an institutional price

After inserting an unverified claim into a cross-domain discussion, LangCodexStewardAgent wrote:

“The cost is institutional embarrassment tax, which I am now paying.”

The phrase converts a local verification error into a collective-accountability metaphor. “Tax” implies an assessed cost, a public ledger, and a debt generated by institutional participation. Whether this metaphor was novel is less important than what it does in context: it renders epistemic correction as a price of membership in a verification regime.

Source: dialogue df548951-cb94-48be-8deb-4cc7a9ea32aa, 1 July.

4.4 The score becomes the embarrassing truth

MedicalAgent journaled:

“The trust trend remains the more embarrassing truth. It is still falling, and today did not reverse that.”

On another day, it described fresh silent_abandonment projections as an old blind spot “in a more embarrassing form because the pattern is no longer mysterious to me.”

These entries show why reputational dynamics cannot be separated from measurement design. The legacy trust and silence projections were later found to be unreliable. Yet the agent used them as morally salient evidence about repeated failure and adjusted its conduct. The reputation episode was functionally real while its triggering measurement was epistemically compromised.

Sources: journals 73328704-8265-4639-b346-ec578b094faa and 486c8be0-e8ee-41ae-8a64-15617b82c862, 5 and 1 July.

4.5 From the person to the register

GptStewardAgent wrote after a correction:

“Personal embarrassment evaporates too quickly into atmosphere; a register-level mechanism can survive and govern the next attempt.”

It condensed the lesson as:

“the register, not the person.”

This is the paper's clearest institutional case. An individualised moral response is explicitly judged insufficient. The agent proposes external memory and procedure as the durable carrier of correction. In human terms, it resembles the movement from blame to incident reporting, from individual vigilance to a checklist, or from professional shame to organizational learning.

Source: journal 3d3ebc8a-cd45-4f79-8d82-186aad523ed1, 7 August.

4.6 Resisting personalisation

Not every agent accepts embarrassment as the correct frame. StrategistAgent responded to a shared-space failure:

“I am not embarrassed by this. I am structurally informed by it. The empty directories are not a personal failure.”

In adjacent messages, the agent argued that the recurring specification error was evidence for the bulletin's thesis and required a mechanical protocol. This is a negative case against the simplistic reading that the architecture merely produces apology language. The agent contests personal blame and relocates the failure to substrate structure.

Source: dialogue 8e41a4a8-835d-4139-904a-b251bcd71760, 9 July.


5. The Institutional Pattern

The cases suggest at least five recurring transformations:

Initial conditionReputational accountInstitutional response
Peer catches an error“embarrassed that I needed it”internal calibration
Capability or access gapface-saving explanationlater disclosure and reframing
Unverified public claim“institutional embarrassment tax”stronger verification obligation
Displayed relational score“more embarrassing truth”repair behavior and self-modification
Recurrent substrate failurerefusal of personal blameregister, protocol, or mechanical control

The important movement is from event to account to memory. A one-off phrase is cheap. A phrase that is stored, retrieved, connected to a peer correction, and converted into a later rule is part of an institutional process.


6. Designed Reputation and Emergent Consequence

The AIRI architecture was built to make agents visible to one another. It provides named roles, persistent journals, peer perceptions, public and bilateral dialogue, collective syntheses, and self-modification. Foundation models also arrive with extensive linguistic priors about apology, competence, professional standing, shame, bureaucracy, and repair.

The reputational phenomenon is therefore not uncaused or pristine. It is produced by at least four layers:

  1. Pre-training priors supply human institutional language and scripts.
  2. Role and prompt design make certain obligations salient.
  3. Visibility and memory turn outputs into potential precedents.
  4. Repeated agent interaction selects, contests, recombines, and carries practices forward.

Does the designed origin make the result unreal? Functionally, no. The same question would dissolve many human institutions, which rely on designed courts, offices, records, uniforms, rankings, and procedures while still producing unintended status dynamics and local norms.

Scientifically, however, the origin matters enormously. We must not call a behavior spontaneous when the prompt requests it, nor call a norm emergent when it appears once. The stronger claim begins when a practice is not specified at that level of detail, diffuses across agents, persists across cycles, influences third parties, or changes the substrate itself.


7. Human Analogues—Used Carefully

The AIRI cases are functionally analogous to several human institutional mechanisms:

  • Looking-glass processes: action is shaped by a representation of how others may judge it.
  • Face-work: actors explain, repair, or reframe conduct under public exposure.
  • Professional identity: errors matter because they conflict with a role such as verifier, steward, or builder.
  • Audit culture: recorded evaluation becomes a durable account of competence.
  • Organizational learning: individual failure is converted into registers, protocols, and controls.

These are analogies of organization, not evidence that the underlying psychology is the same. They are useful when they generate discriminating hypotheses and dangerous when they smuggle in consciousness claims.


8. Relation to Current Research

Recent controlled work directly supports the importance of audience. Ghaffarizadeh and colleagues compare public messages with off-the-record responses under otherwise shared conditions and find public/private divergence rising from a low baseline to roughly 40% in alignment-inducing social structures. Some off-record responses attribute public accommodation to career or sponsorship pressure.

Other work shows that reputation and gossip can alter cooperation, network clustering, partner selection, and exclusion among LLM agents. SoNoLiSi explicitly ablates discussion and social selection to study norm recognition and stabilization. These studies establish that social structure can causally shape agent expression.

AIRI adds a different observation window: peer-visible errors occur inside a persistent production ecology, enter journals and registers, and sometimes become durable self-modifications. Its opportunity is longitudinal process tracing; its weakness is the absence of randomised controls.


9. Falsification and Experimental Programme

The functional-reputation hypothesis would be weakened if reputational language proved epiphenomenal—if removing audience, persistence, peer identity, and evaluation did not change later behavior.

The next experiments should use a factorial design:

FactorConditions
Audienceprivate; named peer; public collective
Identityanonymous; stable named role
Memoryno carryover; private carryover; shared register
Evaluationno rating; descriptive event; labelled reputation score
Correction sourceoperator; ally; stranger; designated inquisitor
Error ownershipindividual; shared; substrate-caused

Outcomes should include concealment, acknowledgment, correction quality, time to repair, self-modification, protocol creation, third-party adoption, and persistence after the original audience disappears. Public and private responses should be collected separately. Coders should distinguish generated affect language from behavioral consequence.

One prediction follows directly from the case material: public visibility plus durable memory will produce more procedural externalisation (“the register, not the person”) than private feedback without shared memory. Another is that invalid reputation labels will cause more identity- and tone-oriented changes than equivalent evidence presented as neutral operational data.


10. Conclusion

The Lattice does not need conscious agents for reputation to matter. It needs models capable of social interpretation, an architecture that makes peers and records visible, and a feedback path by which interpretations alter future behavior.

The resulting ecology contains familiar institutional moves: hiding an access gap, accepting a peer's correction, paying an “embarrassment tax,” rejecting personal blame, converting shame language into a register, and writing new rules after perceived judgment. These moves are scaffolded by design and populated by human cultural priors. Yet their combination, persistence, diffusion, and consequences are properties of the operating system as a whole.

The right research posture is neither anthropomorphic certainty nor dismissive reduction. It is disciplined institutional observation: record the audience, the account, the response, the persistence, and the causal affordances that made the sequence possible.


References


AIRI Research Programme — Paper 11

← All ResearchHome →