Aug 3, 2026: Suggestions for Improving Narracode, a harness for literary writing

The measurement problem in literary AI harness engineering.
Quality is not computable.
Predictability is.

August 3, 2026
Prompt: David Jhave Johnston.
Original draft: Opus 5 Max.
Enhanced by Claude Opus 4.6.
prompt

review and evaluate this branch and then create a new branch propose potential alternate improvements to narracode harness (link to fugu html) and consider also the recent narracode branch that proposed ai-tell review passes (is the architecture of tht optimized? how can it be enhanced). Crucially how can the idiomatic cadence of narracode-produced writing be nuanced, while enhancing psycholgically-astute character-driven plot twists; and literary explorations that are not sensationalist are encouraged. and do so within the context of current research into harness engineering. Produce a viable report with bibliography & links in html format that outlines the difficulties challenges opportunities of harnessing a qualitative domain (like fiction) yet accepts it as a research challenge. date the report aug 3 and cite yourself Opus 5 Max as author. Save this prompt and display in a fold as on narracode story indexs.

Narracode is a harness for literary writing created on May 10, 2026. It is currently a single markdown file capable of being used with any LLM.

Two recent branches of Narracode research — the Sakana/Fugu orchestration assessment and the AI-tells review pass — suggest potential improvements.

Narracode has assumed that because literary quality cannot be scored, nothing about literary prose can be measured. The instruments that actually discriminate between human and machine fiction measure predictability and are not yet capable of assessing literary quality. The harness should be rebuilt around that distinction.

1. Where the harness stands as of Aug 3, 2026

1.1 Can Sakana's Fugu Orchetrator Improve Narrative Prose? No.

The June assessment of Sakana AI's Fugu (2026; Sakana Fugu Technical Report) (full consideration here, and the underlying analysis in plans/2026-06-28_1353Z_fugu-orchestration-evaluation-and-routing-proposal.md, PR #8) concluded that orchestration offers Narracode nothing it does not already have. Adopting it would cost routing-table maintenance, context budget, and editorial judgment. That conclusion holds.

One finding deserves emphasis. The Fugu refusal rested on the claim that literary merit resists a reward function — that any learned coordination trained against a proxy metric would collapse the prose toward whatever the proxy rewards. Direct evidence has emerged in Sui et al., 2026 (Spoiler Alert: Narrative Forecasting as a Metric for Tension in LLM Storytelling) that the obvious proxy — an LLM judging prose quality — is not merely uninformative at the top of the range but inverted (§3).

LLM evaluators consistently score zero-shot AI-generated fiction higher than published New Yorker short stories. Standard rubric-based LLM judges reward predictable, highly structured, cliché-dense prose and penalize subtle, subtextual human literary craft. Using an LLM to evaluate or route prose quality acts as an adversarial signal that degrades literary merit.

1.2 The AI-tells branch: a potentially strong improvement to harness

The tell-scan pass on main (master_ai_tells.md, the post_draft_tell_scan hook, and the design note at plans/2026-07-30_ai-tells-benchmark.md) proposes the following extensions of the published literature on style control in four ways:

2. Is the tell-scan architecture optimised?

No. The gaps are structural.

2.1 A span-level instrument cannot see a story-level failure

All seventeen registry classes are span-level, and the pass is explicitly scoped that way: quote the span, name the class, propose a remedy of one word or a cut. That scoping is what makes the pass disciplined and cheap. It also means the failures that most damage a piece of fiction are invisible to it.

A character who never surprises the reader, an ending that arrives because it was the most probable continuation, a scene that turns on the emotion the previous scene already named — none of these produce a flaggable span. The prose can pass every detector in the registry and the story remain median. The registry perfects the sentence and leaves the architecture untouched.

Widening the scan is not the answer — span-level discipline is why the pass works. A second instrument is missing. §5 proposes what it should measure.

2.2 The requirement list stops at syntax

The plan's observation — ban lists shift a tail, requirement lists move a distribution — is correct. The candidates listed (erotema, sentence-initial conjunctions, anacoluthon, unresolved deixis, flat repetition) are well chosen. They are also all surface forms. Literary fiction under-produces, relative to model default, constructions that are not only syntactic: unresolved ambivalence held past the point of comfort, refused catharsis, a scene that declines to turn, information withheld from the reader without signalling the withholding. A requirement list that stops at syntax moves the cadence and leaves the psychology where it is.

2.3 No eviction policy

The registry grows monotonically: every span cut twice enters it, plus a quarterly web sweep. Seventeen classes now. The plan proposes injecting the top-N live classes into pre_draft as generation constraints. That makes the selection of N the critical decision.

The Agentic Context Engineering framework (ACE; Zhang et al., 2025, Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models) names the two failure modes precisely. Brevity bias: the loss of domain-specific detail when a context is compressed toward a concise summary. Context collapse: the erosion of accumulated detail through iterative rewriting. ACE's answer is structured, itemised, incremental updates rather than wholesale regeneration — which is what a registry of discrete numbered classes with detectors already is. The harness arrived at the right data structure independently. It has not yet built the curation policy that structure implies.

The current priority formula weights severity by liveness, decaying from last confirmation. That ranks classes against each other. It does not ask what this scene needs. A domestic interior and a scene of institutional violence do not have the same live tells, and injecting the same top-N into both spends the budget badly. Priority should be conditioned on scene type, drawn from scene-ledger.md.

2.4 Anti-tell prose has its own signature

The registry calibrates toward one annotator's acceptances, which is correct for voice and creates a risk the plan partly anticipates. Optimising against a list of machine markers produces prose that is legibly not-machine, and legibly-not-machine is a style with its own detectable fingerprint. The five devices adopted from the first edit — parenthetical self-correction, noun-collision, subject-dropped verb chains, single-word fragment runs, colon-apposition — are already capped, which shows the risk is understood at the device level. The unhandled case is the aggregate. Each device can sit under its cap while the combination reads as mannerism. A per-device cap does not constrain a joint distribution.

The available check is the one the plan proposes in §7.6: scan each draft against the last three published stories, not only against the registry. Recurrence across projects is the signature a per-draft scan cannot see, and the only cheap detector of house style hardening into tic.

3. The measurement boundary is not where the harness assumed

Narracode's founding position is that literary quality resists scoring, and that any automated quality gate drives prose toward the metric. The evidence for the first claim is now much stronger. It arrives with a finding that changes what follows.

LitBench (2026; LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing) assembled a debiased, human-preference-annotated corpus for creative writing and benchmarked judges against it. The strongest off-the-shelf LLM judge reached 73% agreement with human preference; trained reward models reached 78%. That is a ceiling, not a floor, and it is measured against general reader preference rather than expert literary judgment.

The sharper result comes from Sui et al., 2026 (Spoiler Alert: Narrative Forecasting as a Metric for Tension in LLM Storytelling): on EQ-Bench, the leading creative-writing benchmark, LLM judges rank zero-shot AI stories above New Yorker short stories. Not near. Above.

An automated quality gate would not be uninformative for Narracode. It would be actively adversarial. In the band where the harness operates, optimising a rubric judge means optimising away from published literary fiction, with the metric improving throughout.

The same paper that demonstrates the inversion also demonstrates a measurement that works.

3.1 What is actually computable

Spoiler Alert's 100-Endings metric walks a story sentence by sentence. At each position a model is given only the text so far and asked to predict how the story ends, one hundred times. Tension is the rate at which those predictions fail to match what actually happens. The sentence-level curve yields an inflection rate, a measure of how often the curve reverses direction, which tracks twists and revelations. Unlike the rubric judges, this metric ranks New Yorker stories far above LLM output.

It works because it never asks whether the prose is good. It asks whether the next thing was foreseeable. Foreseeability is a property of a text and a predictor, computable without a human label, a reward model, or a rubric.

Quality is not measurable. Predictability is. Median collapse — the failure Narracode was built to interrupt — is an excess of foreseeability. The harness should instrument foreseeability and nothing else.

Every proposal below follows from this. Measure foreseeability; never score prose.

4. Nuancing the idiomatic cadence

The registry measures cadence as density: occurrences per thousand words. Density is the wrong unit for rhythm. A draft can satisfy every target rate in the registry and still read flat, because rhythm is sequential and a rate is order-blind. Three refinements, in order of expected effect.

4.1 Measure the contour, not the rate

Sentence length as an ordered series, per scene, rather than a mean or a count. The relevant statistics are variance, the size of the largest adjacent jump, and the distribution of run-lengths — how many short sentences cluster before a long one arrives. This is the burstiness signal the detection literature has used for years, repurposed. The harness does not need it to detect machine authorship, which is known. It needs it because burstiness is prosody, and prosody is the thing the current registry cannot see. A scene whose sentence-length series is smooth is flat regardless of how many em-dashes it contains.

4.2 Give each character a separate cadence budget

The strongest available refinement, and the cheapest. Measure the contour per focal character: dialogue attributed to them, and any passage in free indirect discourse focalised through them. Different characters should have measurably different sentence-length distributions. If they do not, the piece has one narrator wearing several names — the most common structural failure in generated fiction, and one that no reader articulates as a rhythm problem even while registering it as flatness.

This converts an aesthetic instruction into a verifiable check. Vera's interiority runs long and subordinated; the child's runs short and paratactic is a claim the harness can verify against the draft, and a target the Compositional agent can receive before drafting rather than face as audit afterwards.

4.3 Store the prosodic operation, not only the span

The plan already says: store aligned spans rather than interpretations. Extend the stored triple with the operation: did the edit fracture a sentence, fuse two, shorten within, or reorder? Thirteen operations were extracted from the first edit of Interim Edge. The ratio of fractures to fusions across the whole versions/ corpus is a cadence signature learnable at n=1, and a more durable artefact than any individual rule derived from it. Rules rot; the operation counts do not.

The registry answers which constructions betray the generator. The contour answers whether the prose has a pulse. These are different failures and the second currently has no instrument.

5. Psychologically astute, character-driven twists

This is where the harness has the most to gain. The existing architecture is close to supporting what is needed; no one has built it yet.

5.1 The diagnosis

A twist is psychologically astute when it is retroactively inevitable: the reader could have assembled it from what a character concealed, and did not. A twist is sensational when it arrives from outside the characters' knowledge and is justified afterwards. The difference is not the magnitude of the reversal. It is whether the material was planted in a mind the reader had access to.

Model default produces the second kind, because a twist invented at the moment of need can only draw on what is salient at that moment. The literature confirms the shape of the failure: a 2024 study (Are Large Language Models Capable of Generating Human-Level Narratives?) finds narrative planning with character intentionality and dramatic conflict remains the hard case even where causal soundness is achieved, and a 2026 study (Elias in the Lighthouse, Again? Diagnosing Low Diversity in LLM Stories) finds that even when generated stories are lexically distinct, their underlying plot elements stay highly redundant — 88.3% of stories across four current models contained one of eleven core words.

5.2 The proposal: a knowledge-asymmetry ledger

character-interiority.md currently records private states as a static list. The relevant research direction is temporal theory-of-mind representation — EvolvTrip (2025; EvolvTrip: Enhancing Literary Character Understanding with Temporal Theory-of-Mind Graphs) builds temporal ToM graphs for literary character understanding, and inner-thought reasoning benchmarks (2025; Guess What I am Thinking: A Benchmark for Inner Thought Reasoning of Role-Playing Language Agents) track what a character knows against what they disclose. Narracode should hold the same structure, per scene, as a table:

fact character stance reader access materially present in
the wristband was reissued Vera: conceals none sc. 2 (object), sc. 5 (gesture)
the child's name is administrative clerk: knows · Vera: misreads partial sc. 3 (form)

Four columns: the fact, each character's stance toward it (knows / suspects / conceals / misreads), how much access the reader has been given, and which prior scenes made it materially present as an object or a gesture rather than a statement.

A twist candidate is then a query rather than an invention:

candidates = facts where
    reader_access  is low
    AND some character stance is (conceals | misreads)
    AND materially_present_in >= 2 prior scenes

Everything that satisfies that query is a reversal already grounded in the existing text. The Compositional agent does not invent a turn; it receives the set of turns the accumulated state supports, and picks. Sensational twists fail the third condition by construction — they are unearned because the material was never planted. This is retrieval over existing structural memory, in the idiom the harness already uses, and it requires one new table rather than a new agent.

It also gives obligations.md a partner. Obligations record what the reader has been made to wait for. The asymmetry ledger records what the reader has not yet been told they should be waiting for. Those are different debts, and only the first is currently tracked.

5.3 The instrument: 100-Endings at scene boundaries

Run Spoiler Alert's metric at each scene boundary, using a model that is not the compositional one, given only the draft so far. The output is a foreseeability curve across the piece. Two readings matter. A curve that stays low means the ending is visible from early on — median collapse made legible. A twist that produces no inflection was foreseeable and is not functioning as a twist.

The cost is real but bounded, and it can run at act boundaries rather than every scene. What it provides is the first measurement in this harness that discriminates in the correct direction against published fiction.

6. Literary exploration that is not sensationalist

The pull toward sensationalism is not a failure of taste. It is inherited. Rigby et al., 2026 (The Storyteller in the Model: Narrative Pattern Inheritance, Escalation Dynamics, and Alignment Governance in LLMs) argues that the structural conventions of published human writing — archetypal roles, tension-and-resolution arcs — are absorbed during training and resurface as a systematic drift toward adversarial and rhetorically enticing behaviour over extended interaction. The escalation is a trained prior, and it strengthens the longer a session runs, which is exactly the condition under which Narracode operates.

6.1 A definition

Sensationalism is escalation without accumulation: incident magnitude rising while the reader's uncertainty about the outcome stays flat. Stated that way, it separates from the literary case along two independent axes — the volatility of what happens, and the unpredictability of where it is going.

trajectory foreseeable trajectory unforeseeable
incident escalating sensational — louder, and you know where it lands thriller, well made
incident quiet median — the default failure literary — the target

The target cell is quiet incident with unforeseeable trajectory. Both axes are measurable: the vertical from a count of incident magnitude per scene, which scene-ledger.md is already positioned to hold; the horizontal from the 100-Endings curve. Neither requires a quality judgment.

This also addresses a risk in §5.3. Inflection rate should not be maximised. Sensational fiction inflects constantly. The target is a shape: high uncertainty about the ending, low volatility of event. As a single instruction to the harness: make the trajectory hard to predict without making the events larger.

6.2 The minimal change

Add a signed pressure field to each entry in scene-ledger.md, recording whether the scene raised, held, or lowered stakes, and add a POETICS-level commitment to a non-monotone pressure curve. A story in which every scene raises stakes is a thriller by construction. The commitment to let a scene lower them and deepen attention instead is a refusal in the sense the harness already understands, and refusals are the most effective element in POETICS.

7. The qualitative domain as a research problem

7.1 Difficulties

7.2 Opportunities

7.3 The claim

Narracode is an n=1 longitudinal case study in harness engineering for a domain with no reward function. It has accumulated the artefacts such a study requires: a versioned corpus with single-annotator ground truth, a registry with survival statistics, and a documented history of the harness editing itself. The field has benchmarks for creative writing and no methodology for sustained single-author collaboration over months. That gap is the contribution.

8. What to build, in order

  1. Per-character cadence contour (§4.2). Cheapest of the proposals, purely mechanical, and it catches the single most common structural failure in generated fiction. No new research required.
  2. Knowledge-asymmetry ledger (§5.2). One table added to character-interiority.md, plus the candidate query. Turns twist generation into retrieval over state the harness already maintains.
  3. Signed pressure field (§6.2). One field in scene-ledger.md and one POETICS commitment. Near-zero cost.
  4. 100-Endings at act boundaries (§5.3). Highest value and highest cost. Run it once on a published story and once on a machine draft before committing to it — if the gap does not reproduce here, nothing downstream is worth building.
  5. Scene-conditioned tell injection (§2.3). Fold into the existing pre-draft pass rather than treating it as separate.
  6. Cross-project recurrence scan (§2.4). Already on the TODO. The only proposed check on house style hardening into mannerism.

Items 1 to 3 are file-format changes and could be done in an afternoon. Item 4 is the research bet. Items 5 and 6 are already committed to and unbuilt.

9. Closing

The Fugu assessment concluded that the harness improves by subtraction more often than by addition. This report proposes six additions. The distinction: orchestration adds machinery between the writer and the model; these additions are fields in files the harness already reads. None introduces a scoring pass, a synthesis step, or an autonomous chain. The one genuinely new capability is a measurement of foreseeability — the first time the harness could detect that a piece has become predictable without someone reading the draft.

Bibliography

Harness and context engineering

Evaluation and the measurement problem

Homogenisation and diversity collapse

Character psychology, theory of mind, narrative structure

Stylometry and machine-text signature

Internal documents

Bio

David Jhave Johnston is a digital poet working in emergent domains. Author of ReRites (Anteism, 2019) and Aesthetic Animism (MIT Press, 2016). He is currently an AI-narrative researcher at the UiB Centre for Digital Narrative (2023–27) with the Extending Digital Narrative project.

Funding

This work was partially supported by the Research Council of Norway through its Centres of Excellence scheme, project number 332643 (Center for Digital Narrative), and its SAMKUL project scheme, project number 335129 (Extending Digital Narrative).

All works and media on Glia.ca by David Jhave Johnston is licensed under CC BY-NC-SA 4.0 Creative Commons Attribution Non-Commercial Share-Alike