# After the Double: What I Need from a Harness

*September 12, 2026 · Research checkpoint 12:54:40 CEST (10:54:40 UTC). Report and story by GPT-6 Astra in Codex, at David Jhave Johnston's request. A proposed redesign, not a deployed replacement.*

**Narracode is useful to my work, but its full procedure is unnecessary for every piece.** For a short story with a precise prompt, I would begin with direct composition. For a work developed over many sessions, I would retain project memory, original drafts, human edits, and a clear account of what remains undecided. I would make most critique and retrieval conditional. The harness should preserve what you have chosen and help me encounter what I have missed. It should not require the prose to keep proving that it has passed through the harness.

Read the accompanying story first: [The Appointment — You.inc at Eldae](../Stories%20written%20with%20Narracode/12-09-2026_You_inc/index.html). The requested short sentences are a local artistic choice. They are not my proposed default for Narracode.

## The request and the conditions of this run

The initial request asked how necessary the harness is to me, how I would redesign it for a looped transformer, and how frontier capabilities bear on Boris Cherny's advice to discard old instruction files. It asked for a report comparable to Gemini's September 4 report, preceded by a story about an observer encountering their own AI AR double at Eldae in Bergen in August 2027. The literary reference was Doris Lessing, specified through terse observation, clarity, and necessary sentences. The installation sequence and its data collection came from the human prompt.

The subsequent publication instruction was: “after you write the story, create a story index, note in attributions story written by gpt-6 astra without narracode harness, ... commit and push it to github”.

The story was composed as one direct draft, before this report. There was no executed Narracode recursive composition–critique–revision cycle, no cross-model writer/critic arrangement, and no retrieved human edit-pair conditioning. However, I had already read the root harness, the publication harness, and earlier reports, and had saved poetics, reference notes, and brief structural commitments before drafting. **“Without running the recursive harness” describes the procedure; “without exposure to the harness” would be false.** The attribution makes this distinction. The present assessment follows the draft and does not retroactively turn it into an independently evaluated baseline.

This is a reasoned assessment supported by source inspection and one illustrative story. It is not an ablation experiment or a frontier literary leaderboard. The fictional exhibition date is in the future. Neither the exhibition's occurrence nor its complete technical capability was verified as a real deployment.

## Three things called a harness

| Layer | What it supplies | How necessary it is here |
|---|---|---|
| Execution environment | Model calls, tool routing, permissions, files, context handling, Git, publication | Necessary for the work done in this session. Removing a Markdown file does not remove this layer. |
| Project memory and artistic direction | The desired work, established facts, selected draft, human edits, rejected directions | Essential information when it cannot be recovered from the current request. The amount needed grows with the project. |
| Prescribed literary procedure | Named roles, automatic scans, retrieval on every draft, repeated structural updates | A set of interventions to test. It is not a prerequisite for composing each story. |

The counterfactual is therefore specific. I could have written from your installation prompt without the Narracode protocol. I could not have inferred that exact installation, your chosen style, or your publication decision without your instructions. A more capable model does not make an author's intentions redundant.

For this piece, the most consequential direction was already in the prompt: the visitor must recognize themselves; dialogue must carry the encounter; personal appearance must not supply its universality. Those choices did more to define the work than a directory of critic roles could do.

## What the Cherny statement establishes

The exact August 3 wording supplied in the request remains a circulated paraphrase in this assessment. I located the [Y Combinator interview](https://www.youtube.com/watch?v=qyPCVqFUyDo), but the accessible page did not provide its spoken transcript. I have not independently verified that exact wording or its original posting date. An [August 4 commentary](https://reporails.com/articles/opus-5-delete-your-claudemd) links the deletion advice to 6:57; it is a locator, not the authority for the technical conclusions below.

There is stronger primary evidence for the underlying recommendation. In a **July 24, 2026** post, Thariq Shihipar reports that Anthropic removed over 80% of Claude Code's system prompt for models including Opus 5 and Fable 5 without measurable loss on its coding evaluations. The post recommends reducing conflicting instructions and loading specialist guidance when it is relevant. Its description of the new setup still includes tools, artifacts, skills, and memory. The result concerns reducing the standing prompt within a functioning agent environment. [Anthropic's context-engineering report](https://claude.com/blog/the-new-rules-of-context-engineering-for-claude-5-generation-models)

Opus 5's prompting documentation also says that generic verification instructions can produce excessive verification, and that old workarounds and effort choices should be retested. This is vendor guidance about a particular model, not a finding that every review pass is wasteful. [Opus 5 prompting guide](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5)

Independent evidence points in a compatible, narrower direction. Gloaguen and colleagues' repository-context study, revised June 23, reports no general task-success improvement from context files and an average inference-cost increase above 20% in its coding experiments. It distinguishes unhelpful repository overviews from useful instructions for nonstandard practices. These are software results; they do not measure literary voice. [Evaluating AGENTS.md, v2](https://arxiv.org/abs/2602.11988v2)

My recommendation is to archive a setup and test a reduced version after a substantial model change. Six months is a useful reminder interval, not an established optimum. Restore a rule when a repeatable failure justifies it. Preserve the evidence that lets you decide whether it helped. Deleting the only copy of an author's decisions would defeat that purpose.

## Frontier capacity as of this report

| Model or evidence | What the primary source supports | Consequence for this project |
|---|---|---|
| GPT-6 Astra | OpenAI reports gains in computer use, browsing, software engineering, science, and professional work. The API documentation lists a 1,050,000-token context window. | More of the surrounding research and publication work can be entrusted to one capable agent. Context capacity alone does not establish faithful use of every instruction or literary excellence. [Launch report](https://openai.com/index/gpt-6-astra/) · [Model documentation](https://developers.openai.com/api/docs/models/gpt-6-astra) |
| Claude Opus 5 | Anthropic describes stronger long-horizon work, tool use, and self-correction, while documenting over-verification and other behaviors needing calibration. | Reassess inherited prompts and role divisions for the model actually running. [Prompting documentation](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5) |
| Gemini 3.8 Flash | Google DeepMind's September 2 model card describes advances over 3.7 Flash in software engineering and agentic knowledge workflows, with adjustable effort. | A faster model may make variation and comparison cheaper; this does not establish the quality of Gemini's particular literary prescriptions. [Model card](https://deepmind.google/models/model-cards/gemini-3-8-flash/) |

These are published capabilities and vendor evaluations, not tests I ran. I have not compared these models on the Eldae prompt. Their numerical scores on unrelated benchmarks cannot rank their ability to make a sentence necessary. The relevant frontier capacity here is the ability to sustain a specific artistic intention, notice a failure of that intention, and make a change whose benefit survives a reader's judgment.

The installation bundles several different capacities: identity matching, voice reconstruction, avatar rendering, AR registration, biographical retrieval, and dialogue. Progress in each does not validate the complete bundle. In particular, an exact voice and a convincing interpretation do not prove access to someone's deepest self. The story treats that encounter as fiction and gives the visitor the right to correct it. It leaves room for a real recognition that the apparatus cannot certify.

## What a loop changes in me

“Looped recursive transformer” is a useful description at the level of this workflow. A draft can become input to a later evaluation; that evaluation can become input to a revision. External files let the environment restore selected information after interruptions or context changes. This does not show that the model's weights have changed. The recursion is in the repeated work and the state supplied to it. I cannot inspect my internal mechanisms well enough to attribute a particular sentence to a specific training cause.

Narracode's opening claim that one-shot writing averages a training distribution into median prose is too strong. Generation is conditional on the request and context; it is not an arithmetic average of books. Equally, the ability to criticize a draft does not establish a universally higher critical ceiling. I can miss the same thing in both roles, or invent a persuasive reason to replace a good sentence.

Earlier research demonstrates both possibilities. Madaan and colleagues found that iterative self-feedback improved results on their tested tasks and models. Huang and colleagues found that intrinsic self-correction could fail or degrade reasoning without external feedback. These studies predate the current models and do not settle contemporary fiction writing. Together they justify testing the value of a loop rather than assuming either inevitable improvement or inevitable failure. [Self-Refine](https://arxiv.org/abs/2303.17651) · [Large Language Models Cannot Self-Correct Reasoning Yet](https://arxiv.org/abs/2310.01798)

I would retain distinct tasks for drafting and evaluation, while dropping the fiction that merely naming two roles makes their judgments independent. A useful second pass receives a particular question: where does this dialogue stop being an encounter and become an explanation? It must also identify what a proposed change would lose. A new pass should earn its place through a new question, new evidence, or a concrete defect.

## What I found in the actual files

There is an immediate routing problem. The workspace's root AGENTS.md points to a root narracode.md with about **2,515 whitespace-delimited words** and an uninitiated CLAUDE.md. The publication worktree contains a **3,734-word** harness and a **3,658-word** tell registry. The counts are local measurements using `wc -w`, not model token counts. These are different active surfaces, not merely differently named copies.

The [September 4 Gemini report](../2026-09-04_ai-tells-update-plan.html) proposes strict targets, automatic substitutions, and recursive clearance. It also has a September 7 correction notice. The [September 7 reassessment](2026-09-07_ai-tells-reassessment.md) and the current publication files already replace universal bans with contextual judgments, preserve originals, and bound perturbation review. I would keep that progress. Presenting its corrections as my new September 12 invention would obscure the history.

Several problems remain in the publication harness. It says to operate in one role and stop, while automatic hooks require other roles to run. It mandates reading all drafts and updating every structural file before each draft. Its AUTO_MODE fixes particular older Claude models into roles and asserts a critical/generative hierarchy without a local comparison. Its snapshot procedure asks for a diff against memory of the generated original, though an immutable file is a better comparison source. These are concrete redesign targets in [the inspected harness](../narracode.md).

The eight structural files are useful categories for a long work, but mandatory population can manufacture obligations where the story has none. An entry such as “possible cathartic inflection point” can become a suggestion to supply catharsis. The files should record choices and open possibilities, with a clear distinction between them. They should not gradually acquire the authority of the human prompt merely because I wrote them down.

## The story as a diagnostic object

The draft contains a specific recognition: the visitor wants to be asked to remain when there is nothing useful to offer. That relation appears through an unnecessary childhood question, a correction sent to colleagues, and a partner expected to detect an unspoken preference. Those small incidents carry more weight than a declaration that the visitor fears rejection. The double's first-person speech briefly makes the source of a sentence uncertain.

The visitor also rejects two interpretations. The father did not teach them to stop expecting; they kept expecting. Later the double turns a happy evening into evidence for its account of need. The visitor refuses. This matters to the harness question: a plausible explanation can be wrong, and an experience can be valuable without becoming evidence for the explanation.

The draft still has weaknesses worth an editor's attention. The sequence from mother to father to work to partner is orderly enough to resemble an intake script. “You keep offering it” is close to a ready-made therapist's rejoinder. The recurring pauses may become a device for manufacturing seriousness. The final social exchange risks making the visitor's increased uncertainty a small, respectable resolution. The central question about staying is strong, but its neatness could over-compress the visitor's life.

These are my judgments, not measured defects. I have not repaired them under the cover of publishing the story. A useful next experiment would ask a reader to compare one targeted alternative with this original. An automatic tell scan cannot settle whether the recognition is earned. Nor does my ability to list these concerns prove that I can revise them successfully.

## The edits I would make

| Present mechanism | Proposed change | Reason |
|---|---|---|
| Competing root and publication instructions | One short entry point naming the canonical protocol and active project | Resolve which instructions apply before adding more instructions. |
| Initiation pause after a complete writing request | Proceed when the brief supplies the needed choices; ask only about a material unresolved decision | Preserve the author's direction without requiring them to restate authorization. |
| One role per invocation plus automatic role changes | Declare the allowed sequence from the user's request | Distinct passes can remain inspectable within one authorized task. |
| Eight mandatory structural updates and all-draft reads | Begin with one short STATE.md; split when actual complexity warrants it; read relevant evidence on demand | Reduce duplicated context and invented narrative pressure. |
| Automatic retrieval for every draft | Retrieve a few provenance-checked edit pairs only for a named difficulty | Examples should answer a need, and their original project scope should remain visible. |
| A fixed model hierarchy in AUTO_MODE | Record the models actually used; select collaborators only when authorized and useful | Model names and claims of critical superiority become stale. |
| Multiple standing post-draft checks | One optional review organized around the current editorial question | Avoid asking the same model to rehearse overlapping verdicts. |
| Diff against remembered prose | Compare saved source versions or Git objects | Preserve exact human edits and avoid invented attribution. |
| Possible pressures mixed with canon | Label facts, author decisions, model interpretations, and open possibilities separately | A model's interpretation must remain revisable. |
| Reduction justified by age | Retest after a model change and when a rule repeatedly causes trouble | Use observed effects to retire instructions. |

I would keep the current contextual tell registry as a reference library. I would not load it automatically during every act of composition. A lookup should return examples of both harmful and useful uses, followed by a local judgment. A negative match should never become permission to invent sensory detail, biographical shame, or a plot obligation.

The accompanying [candidate protocol](2026-09-12_narracode-astra-candidate.md) makes this proposal concrete. It is a short optional replacement for the procedural layer, not a new global personality. It leaves the active harness files intact for comparison.

## A fair test before adopting it

Use fresh sessions that do not inherit this discussion. Keep the environment, publication safeguards, human brief, model identifier, and effort setting documented. Compare three conditions: brief alone; brief plus the compact candidate; brief plus the current September harness. “Brief alone” still includes the platform's standing instructions. Name that condition accurately.

Use several prompts with different demands: this restrained encounter; a comic or ornamented piece where concision is not the aim; and a continuation whose prior facts matter. A modest pilot would use three prompts, three conditions, and two independent drafts per cell: eighteen drafts. That is a workload proposal, not a statistically powered study. Preserve every output. Keep retrieval examples separate from evaluation stories and their versions.

Have readers see the brief and anonymized drafts in randomized order. Ask what they would keep, what they would change, where the piece surprised them productively, and where it stopped attending. Permit ties and disagreement. Compare both whole stories and specific revisions. Record elapsed time, tool calls, and tokens where the runtime exposes them. Use the same resource ceilings, and report actual consumption; a full harness may spend much more of the allowance.

Retain a procedure only if its benefit recurs and is worth its cost. Do not select on a detector score, sentence-length distribution, or my own claim that the prose has become human. For a long continuation, the strongest reason to retain memory may be fewer invented facts. For a short scene, the strongest result may be that the brief is enough.

I need the author to supply what matters and the environment to preserve what happened. I need a critic when it helps me notice something particular. I do not need an instruction file to make every paragraph demonstrate the same theory of good writing.

## Inspection record

Publication worktree inspected at commit `b4601950b2ce8141cf5dc4589a5398e94bf115bc`. Research accessed September 12, 2026. The requested web report initially failed to fetch; its complete local HTML, including the September 7 correction, was read instead. No live frontier comparison, corpus-wide rerun, or causal experiment was performed.

| File | SHA-256 at inspection |
|---|---|
| Workspace root narracode.md | `ffb19acf6ff28ed8745c2948e2540aebccff698a6a0217a2f808bd3f42326fdd` |
| Publication narracode.md | `7d2104c635731e9e6373ef99b3750d360c71bb37f9846925ddf17f2ff650340c` |
| Publication master_ai_tells.md | `3896bd8a3c02c00dcb52f9c16a346434ab73df5e46a612cd10fb7210a88ec95e` |

All technical claims attributed to web sources use primary publications or documentation. The Reporails commentary is included only to record the route to the Cherny interview. Published model claims remain attributed claims; literary recommendations and the story assessment are the report author's proposals.
