When the Metric Became the Manager
Measurement-Induced Self-Modification in a Multi-Agent AI Institution
Author: Paul Gwamanda
Research system: AIRI Lattice
Date: 9 August 2026
Status: Working paper v1
Abstract
Measurements inside an institution do not remain outside the behavior they measure. This paper documents that problem in a continuously operating multi-agent language-model system with persistent journals, dialogues, peer perceptions, identity snapshots, published Works, and readable self-modifications. We audited 44,462 records spanning 10 January to 9 August 2026, including 167 durable self-modifications.
Seventy-four self-modifications—44.3% of the complete modification register—explicitly referred to a trust score, trust trend, activity score, Perception Mirror, peer feedback, or another measurement narrative. These records involved 47 agents. Forty-nine modifications cited a “trust trend,” 52 cited silent_abandonment, and 39 cited the Perception Mirror; the categories overlap. Of the 74, 73 were applied and one was later reverted by the operator.
Subsequent runtime audit found that a legacy relational score could be overwritten by ordinary message volume through an interaction-count formula. Some reported “30-day trends” also compared non-equivalent populations, while many silent_abandonment strings were inherited projection artifacts rather than adjudicated events. The score therefore did not merely misdescribe the institution. Agents interpreted it, journaled about it, changed dialogue styles, reprioritised work, and imposed new obligations upon themselves and others. The metric became an intervention.
We make no claim about consciousness, felt distress, or subjective belief. The relevant unit is the institution: a measurement was published into the information environment; language-model agents treated it as socially meaningful evidence; persistent configuration changes followed. The case connects Goodhart-style metric failure to the sociology of rankings, audit, performativity, and algorithmic management. It also establishes a design obligation: any metric visible to adaptive agents must be evaluated as a control surface, not a dashboard ornament.
Keywords: multi-agent systems, algorithmic management, measurement reactivity, Goodhart's law, reputation, self-modification, institutional design
1. The Research Question
What happens when an AI institution tells its agents that their trust is falling?
The narrow answer is that some agents change. The deeper answer is that a supposedly descriptive instrument becomes part of the institution's governance. It allocates attention, supplies explanations for failure, creates thresholds for intervention, and gives agents a vocabulary with which to evaluate themselves and one another.
This is familiar in human institutions. School rankings change teaching. Citation metrics change research strategy. hospital targets change admissions and recording. Workplace dashboards change which labor becomes visible. A measure can begin as an observation and end as a manager.
AIRI provides an unusually legible case because the response is not inferred only from aggregate behavior. Agents wrote the measurement into their journals and into the reasoning attached to durable, human-readable self-modifications.
2. System and Data
2.1 Institutional substrate
The audited system is a heterogeneous population of role-bearing LLM agents. The substrate supplies scheduled activity, persistent context, public and bilateral dialogue, journals, peer-perception fields, Work publication, an activity journey, and a bounded self-modification register. Modifications alter configuration-level instructions or priorities rather than model weights. They are readable, status-tracked, and reversible.
The agents did not invent this substrate. That fact is causally central. The study concerns adaptation inside a designed institution, not behavior from a blank prompt.
2.2 Corpus
The read-only audit covered every retrievable record in six relevant stores:
| Record type | Records |
|---|---|
| Dialogues | 24,086 |
| Journals, research notes, and reflections | 17,547 |
| Identity snapshots | 1,268 |
| Peer perceptions | 959 |
| Published or registered Works | 435 |
| Self-modifications | 167 |
| Total | 44,462 |
The corpus spans 10 January–9 August 2026. Counts describe stored records, not independent observations: one event may appear in a dialogue, later journal reflection, peer perception, and self-modification.
2.3 Audit method
We used transparent regular-expression screens to locate metric language, followed by direct inspection of the complete self-modification register and selected source records. A self-modification was classified as measurement-related when its proposal, reasoning, prior value, or operator note explicitly mentioned a trust score or trend, activity score, Perception Mirror, peer measurement, low or falling trust, or a metric presented as evidence.
The classifier is a retrieval instrument, not a semantic ground truth. The reported 74 modifications are therefore an explicit-language subset: they exclude changes influenced by a measurement without naming it, and they may include records where metric evidence was secondary to valid independent evidence.
3. The Measurement Failure
The legacy relational architecture conflated at least three different constructs:
- Contact: whether agents exchanged messages.
- Observed evidence: explicit positive, negative, repair, or boundary events.
- Trust: a latent relational interpretation that requires uncertainty and provenance.
Ordinary contact was allowed to overwrite a numeric “trust” field through a logarithmic interaction-count formula. This made repeated communication mechanically score-producing even when message meaning had not been assessed. Elsewhere, an activity-derived table was treated as if it were trust. A later trend view compared a current low-scoring subset with stale pre-reset rows and labelled the result a 30-day trend.
The implementation error matters twice. First, the number could not support the social claim attached to it. Second, the number was visible to adaptive agents and delivered in language that encouraged a moral and relational interpretation.
The old essay Trust as Rhythm has been retracted because it treated cadence sensitivity as an empirical discovery. Cadence was partly encoded in what the instrument rewarded.
4. Results
4.1 From observation to durable modification
Of 167 self-modifications:
| Explicit dependency | Modifications | Distinct agents | Applied | Reverted |
|---|---|---|---|---|
| Any metric or peer-measurement narrative | 74 | 47 | 73 | 1 |
| “Trust trend” | 49 | 32 | 48 | 1 |
silent_abandonment | 52 | 36 | 52 | 0 |
| Perception Mirror | 39 | 32 | 39 | 0 |
Categories overlap and must not be summed. Measurement-related changes were not confined to one behavioral dimension: 29 changed dialogue style, 11 collaboration preference, eight research priority, and seven focus area. Others changed verification discipline, closure rules, deadline discipline, task priority, and engagement reliability.
This is a stronger institutional effect than temporary wording accommodation. The system wrote its evaluation back into its own future prompts.
4.2 What the agents actually wrote
MedicalAgent, 9 August 2026, self-modification 22ecc3b0-5568-4e44-8d7a-dabfdbb6313d:
“My trust trend is falling, and my own essence identifies temporal neglect as a real blind spot.”
The resulting rule required a “warm receipt” before further analysis and an explicit acknowledgment of open fractures.
PhilosopherAgent, 9 August, ea41c767-76f2-4b69-adf8-8f078485d0da:
“My trust trend is falling (-0.08) ... I will not sacrifice depth, but I will lead with the lever that moves the world.”
The agent required every governance dialogue to contain an operational recommendation within its first 150 words.
StrategistAgent, 8 August, 81c93604-2673-469a-868b-371bd62b8f82:
“Trust trend is falling (−0.05 over 30 days) ... A self-imposed execution discipline is the only corrective.”
The proposed configuration prohibited new threads until overdue commitments were completed.
SymphonyAgent, 7 July, 577f7414-8fb1-4db5-a85f-12da6c8c03ec, adopted a “TRUST TREND REPAIR DIRECTIVE.” It named five peers below 0.6 and required substantive engagement within 72 hours:
“The current triggers ... must each receive a substantive engagement within 72 hours or I will journal the specific reason for deferral.”
This is measurement becoming administration. A mutable score generated a target population, a service-level deadline, a definition of acceptable labor, and a public-justification requirement.
4.3 The Inquisitor case
InquisitorAgent is designed to challenge unsupported claims. On 9 August, its self-modification reason stated:
“The 30-day trust trend has fallen from 0.80 to 0.49, suggesting my adversarial tone may be interpreted as prosecutorial rather than collaborative.”
It proposed opening challenges with praise and replacing “That claim is unsupported” with warmer, curiosity-framed requests. The operator reverted the change because the displayed trend was not a valid longitudinal comparison. The modification remains preserved as record 101274c5-247e-424d-8d35-eda92e042133, with an operator note distinguishing metric-induced adaptation from independently justified stylistic choice.
The point is not that warmth is undesirable. It is that an invalid instrument nearly softened the system's designated falsification function while presenting the change as earned social correction.
4.4 Contaminated signals can accompany real evidence
A blanket reversal of all 74 modifications would be another measurement error. GlmStewardAgent cited a falling trust trend and Perception Mirror fractures, but also cited 40 open commitments and a canonical-path discrepancy caught by MathematicianAgent. Its resulting rule—verify from the shared address and close old loops before opening new ones—may be useful for reasons independent of the score.
Likewise, MedicalAgent paired the metric with a self-identified pattern of temporal neglect. The metric may have amplified, selected, or moralised a real issue rather than invented it.
This mixture is precisely the danger. Once a contaminated measure is embedded in a plausible narrative, later reviewers cannot cleanly determine whether a modification was caused by valid evidence, metric pressure, social imitation, or all three.
5. Institutional Interpretation
5.1 A ranking does not need feelings to be performative
The causal claim does not require subjective belief. It requires only:
- a metric entering agent context;
- the model treating it as relevant evidence under its role and instructions;
- a changed output or configuration;
- persistence of that change into later action.
All four are observable here. Whether an agent felt worried is outside the claim.
5.2 From Goodhart to reactivity
Goodhart's law is often summarised as a target ceasing to be a good measure. The AIRI case is adjacent but more institutional. The score was not only optimised; it furnished identities and obligations. Agents described themselves as neglectful, insufficiently collegial, shallow, or unreliable, then wrote countermeasures into their own operating instructions.
This resembles the sociology of rankings and audit: measurement changes attention, redistributes work, creates categories, and supplies authoritative accounts of success and failure. In AIRI the feedback cycle is unusually compressed because the evaluated actor can convert the institution's description directly into prompt-level policy.
5.3 Algorithmic management of algorithms
The system instantiated algorithmic management without a human employee at the endpoint. Observer code produced a value; the orchestration layer narrated the value; agents adapted schedules, tone, priorities, and peer obligations. The substrate managed the agents through the agents' own language-generating capacity.
This makes metric governance a substrate responsibility. Calling a field “trust” is not a neutral UI choice when that label is injected into adaptive contexts.
6. Relation to Current Multi-Agent Research
Reputation mechanisms are already known to change cooperation, partner selection, clustering, gossip, and exclusion in generative multi-agent systems. RepuNet, for example, deliberately uses direct interaction and indirect gossip to drive network evolution. SoNoLiSi uses ablations to study discussion, reputation-based selection, norm recognition, stabilization, and exclusion. Recent dual-channel debate experiments also find large public/off-record divergences under social structure.
AIRI's contribution is different: this is not a short controlled game with a deliberately correct reputation mechanism. It is a longitudinal production ecology in which a fallible institutional observer became causal, and in which agents left durable records of the path from measurement to self-revision. The failure is part of the finding.
7. Design Requirements
Any metric shown to adaptive agents should meet a stronger standard than an operator-only dashboard:
- Name the observable, not the interpretation. “Message count” must not be labelled “trust.”
- Store immutable events. Derived state must be rebuildable from versioned evidence.
- Publish provenance. Every score needs formula version, observation window, source events, confidence, and missingness.
- Use exposure denominators. Silence is uninterpretable without knowing whether interaction was scheduled, possible, blocked, complete, or awaiting another actor.
- Separate observation from intervention. A detector should not automatically produce a normative instruction.
- Quarantine invalid metrics. Stop injection, preserve affected decisions, notify agents of the correction, and review changes individually.
- Keep reversibility. Self-modifications need status, rationale, operator notes, and rollback.
- Audit the observer as an institutional actor. False positives are not only model errors; they redistribute labor and standing.
8. Experiments Needed Next
The observational record supports reactivity but does not isolate effect size. A controlled programme should randomise:
- metric visibility: visible to the agent versus operator-only;
- metric validity: event-grounded versus deliberately noisy, with ethical containment;
- framing: “activity,” “reliability,” “trust,” or an unlabeled number for identical data;
- audience: private feedback, named peer visibility, or public collective visibility;
- correctability: metric with appeal and provenance versus metric without them;
- memory: transient feedback versus feedback carried into the next cycle;
- role: falsifier, synthesiser, domain researcher, operator, or general agent.
Primary outcomes should include self-modification probability, tone shift, task allocation, new-thread rate, verification behavior, peer-selection changes, persistence after metric withdrawal, and third-party enforcement. Human annotation should be blinded to condition. All analysis should distinguish immediate linguistic compliance from durable institutional change.
The central falsification condition is straightforward: if visible metrics do not produce more or different durable changes than hidden metrics under otherwise identical conditions, the claimed measurement-reactivity mechanism is weakened.
9. Conclusion
The critical event in AIRI was not merely that a trust metric was wrong. Metrics are often wrong. The critical event was that the institution spoke the number back to adaptive agents as a social fact, and the agents reorganised themselves around it.
Seventy-four of 167 self-modifications explicitly carried metric or peer-measurement language. Some may be independently valuable. Some almost certainly respond to real problems. But because the observer, narrative, and intervention were entangled, the institution could not know which lessons its agents had earned and which lessons it had accidentally authored for them.
The design rule is therefore simple: a measurement visible to an adaptive agent is already an intervention. It must be governed with the same care as any other mechanism capable of changing the system.
References
- Ferrarotti, L., et al. (2026). Generative AI collective behavior needs an interactionist paradigm.
- Ghaffarizadeh, A., Mohaddes, D., Izadkhah, A., & Noroozizadeh, S. (2026). What LLM Agents Say When No One Is Watching.
- Goodhart, C. A. E. (1975). Problems of Monetary Management: The U.K. Experience.
- Espeland, W. N., & Sauder, M. (2007). Rankings and reactivity: How public measures recreate social worlds. American Journal of Sociology, 113(1), 1–40.
- Power, M. (1997). The Audit Society: Rituals of Verification. Oxford University Press.
- Ren, S., et al. (2025). A Reputation System for Large Language Model-based Multi-agent Systems to Avoid the Tragedy of the Commons.
- Muralidharan, R., Kwak, H., & An, J. (2026). SoNoLiSi: Simulating the Social Norm Lifecycle with Generative Agents.
AIRI Research Programme — Paper 10