review and evaluate this branch and then create a new branch propose potential alternate improvements to narracode harness (link to fugu html) and consider also the recent narracode branch that proposed ai-tell review passes (is the architecture of tht optimized? how can it be enhanced). Crucially how can the idiomatic cadence of narracode-produced writing be nuanced, while enhancing psycholgically-astute character-driven plot twists; and literary explorations that are not sensationalist are encouraged. and do so within the context of current research into harness engineering. Produce a viable report with bibliography & links in html format that outlines the difficulties challenges opportunities of harnessing a qualitative domain (like fiction) yet accepts it as a research challenge. date the report aug 3 and cite yourself Opus 5 Max as author. Save this prompt and display in a fold as on narracode story indexs.
Narracode is a harness for literary
writing created on May 10, 2026. It is currently
a single markdown file capable of being used with any LLM.
Two recent branches of Narracode research —
the Sakana/Fugu
orchestration assessment and the AI-tells
review pass — suggest potential improvements.
Narracode has assumed that because literary
quality
cannot be scored, nothing about literary prose can be measured. The instruments that actually discriminate
between human and machine fiction
measure predictability and are not yet capable of assessing literary quality. The harness should be rebuilt
around that distinction.
The June assessment of Sakana AI's Fugu (2026; Sakana Fugu Technical Report)
(full consideration here, and the underlying
analysis in plans/2026-06-28_1353Z_fugu-orchestration-evaluation-and-routing-proposal.md,
PR #8) concluded that orchestration
offers Narracode nothing it does not already have. Adopting it would cost routing-table maintenance,
context budget, and editorial judgment. That conclusion holds.
One finding deserves emphasis. The Fugu refusal rested on the claim that literary merit resists a reward function — that any learned coordination trained against a proxy metric would collapse the prose toward whatever the proxy rewards. Direct evidence has emerged in Sui et al., 2026 (Spoiler Alert: Narrative Forecasting as a Metric for Tension in LLM Storytelling) that the obvious proxy — an LLM judging prose quality — is not merely uninformative at the top of the range but inverted (§3).
LLM evaluators consistently score zero-shot AI-generated fiction higher than published New Yorker short stories. Standard rubric-based LLM judges reward predictable, highly structured, cliché-dense prose and penalize subtle, subtextual human literary craft. Using an LLM to evaluate or route prose quality acts as an adversarial signal that degrades literary merit.
The tell-scan pass on main (master_ai_tells.md, the
post_draft_tell_scan hook, and the design note at
plans/2026-07-30_ai-tells-benchmark.md) proposes the following extensions of the published
literature on style control
in four ways:
before → after
pairs rather than the model's reading of them is the correct archival call. It matches what the context
engineering literature has since concluded about summarisation loss (§2.3).No. The gaps are structural.
All seventeen registry classes are span-level, and the pass is explicitly scoped that way: quote the span, name the class, propose a remedy of one word or a cut. That scoping is what makes the pass disciplined and cheap. It also means the failures that most damage a piece of fiction are invisible to it.
A character who never surprises the reader, an ending that arrives because it was the most probable continuation, a scene that turns on the emotion the previous scene already named — none of these produce a flaggable span. The prose can pass every detector in the registry and the story remain median. The registry perfects the sentence and leaves the architecture untouched.
Widening the scan is not the answer — span-level discipline is why the pass works. A second instrument is missing. §5 proposes what it should measure.
The plan's observation — ban lists shift a tail, requirement lists move a distribution — is correct. The candidates listed (erotema, sentence-initial conjunctions, anacoluthon, unresolved deixis, flat repetition) are well chosen. They are also all surface forms. Literary fiction under-produces, relative to model default, constructions that are not only syntactic: unresolved ambivalence held past the point of comfort, refused catharsis, a scene that declines to turn, information withheld from the reader without signalling the withholding. A requirement list that stops at syntax moves the cadence and leaves the psychology where it is.
The registry grows monotonically: every span cut twice enters it, plus a quarterly web sweep. Seventeen
classes now. The plan proposes injecting the top-N live classes into pre_draft as
generation constraints. That makes the selection of N the critical decision.
The Agentic Context Engineering framework (ACE; Zhang et al., 2025, Agentic Context Engineering: Evolving Contexts for Self-Improving Language Models) names the two failure modes precisely. Brevity bias: the loss of domain-specific detail when a context is compressed toward a concise summary. Context collapse: the erosion of accumulated detail through iterative rewriting. ACE's answer is structured, itemised, incremental updates rather than wholesale regeneration — which is what a registry of discrete numbered classes with detectors already is. The harness arrived at the right data structure independently. It has not yet built the curation policy that structure implies.
The current priority formula weights severity by liveness, decaying from last confirmation. That ranks
classes against each other. It does not ask what this scene needs. A domestic interior and a scene
of institutional violence do not have the same live tells, and injecting the same top-N into both
spends the budget badly. Priority should be conditioned on scene type, drawn from
scene-ledger.md.
The registry calibrates toward one annotator's acceptances, which is correct for voice and creates a risk the plan partly anticipates. Optimising against a list of machine markers produces prose that is legibly not-machine, and legibly-not-machine is a style with its own detectable fingerprint. The five devices adopted from the first edit — parenthetical self-correction, noun-collision, subject-dropped verb chains, single-word fragment runs, colon-apposition — are already capped, which shows the risk is understood at the device level. The unhandled case is the aggregate. Each device can sit under its cap while the combination reads as mannerism. A per-device cap does not constrain a joint distribution.
The available check is the one the plan proposes in §7.6: scan each draft against the last three published stories, not only against the registry. Recurrence across projects is the signature a per-draft scan cannot see, and the only cheap detector of house style hardening into tic.
Narracode's founding position is that literary quality resists scoring, and that any automated quality gate drives prose toward the metric. The evidence for the first claim is now much stronger. It arrives with a finding that changes what follows.
LitBench (2026; LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing) assembled a debiased, human-preference-annotated corpus for creative writing and benchmarked judges against it. The strongest off-the-shelf LLM judge reached 73% agreement with human preference; trained reward models reached 78%. That is a ceiling, not a floor, and it is measured against general reader preference rather than expert literary judgment.
The sharper result comes from Sui et al., 2026 (Spoiler Alert: Narrative Forecasting as a Metric for Tension in LLM Storytelling): on EQ-Bench, the leading creative-writing benchmark, LLM judges rank zero-shot AI stories above New Yorker short stories. Not near. Above.
An automated quality gate would not be uninformative for Narracode. It would be actively adversarial. In the band where the harness operates, optimising a rubric judge means optimising away from published literary fiction, with the metric improving throughout.
The same paper that demonstrates the inversion also demonstrates a measurement that works.
Spoiler Alert's 100-Endings metric walks a story sentence by sentence. At each position a model is given only the text so far and asked to predict how the story ends, one hundred times. Tension is the rate at which those predictions fail to match what actually happens. The sentence-level curve yields an inflection rate, a measure of how often the curve reverses direction, which tracks twists and revelations. Unlike the rubric judges, this metric ranks New Yorker stories far above LLM output.
It works because it never asks whether the prose is good. It asks whether the next thing was foreseeable. Foreseeability is a property of a text and a predictor, computable without a human label, a reward model, or a rubric.
Quality is not measurable. Predictability is. Median collapse — the failure Narracode was built to interrupt — is an excess of foreseeability. The harness should instrument foreseeability and nothing else.
Every proposal below follows from this. Measure foreseeability; never score prose.
The registry measures cadence as density: occurrences per thousand words. Density is the wrong unit for rhythm. A draft can satisfy every target rate in the registry and still read flat, because rhythm is sequential and a rate is order-blind. Three refinements, in order of expected effect.
Sentence length as an ordered series, per scene, rather than a mean or a count. The relevant statistics are variance, the size of the largest adjacent jump, and the distribution of run-lengths — how many short sentences cluster before a long one arrives. This is the burstiness signal the detection literature has used for years, repurposed. The harness does not need it to detect machine authorship, which is known. It needs it because burstiness is prosody, and prosody is the thing the current registry cannot see. A scene whose sentence-length series is smooth is flat regardless of how many em-dashes it contains.
The strongest available refinement, and the cheapest. Measure the contour per focal character: dialogue attributed to them, and any passage in free indirect discourse focalised through them. Different characters should have measurably different sentence-length distributions. If they do not, the piece has one narrator wearing several names — the most common structural failure in generated fiction, and one that no reader articulates as a rhythm problem even while registering it as flatness.
This converts an aesthetic instruction into a verifiable check. Vera's interiority runs long and subordinated; the child's runs short and paratactic is a claim the harness can verify against the draft, and a target the Compositional agent can receive before drafting rather than face as audit afterwards.
The plan already says: store aligned spans rather than interpretations. Extend the stored triple with the
operation: did the edit fracture a sentence, fuse two, shorten within, or reorder? Thirteen
operations were extracted from the first edit of Interim Edge. The ratio of fractures to fusions
across the whole versions/ corpus is a cadence signature learnable at n=1, and a more durable
artefact than any individual rule derived from it. Rules rot; the operation counts do not.
The registry answers which constructions betray the generator. The contour answers whether the prose has a pulse. These are different failures and the second currently has no instrument.
This is where the harness has the most to gain. The existing architecture is close to supporting what is needed; no one has built it yet.
A twist is psychologically astute when it is retroactively inevitable: the reader could have assembled it from what a character concealed, and did not. A twist is sensational when it arrives from outside the characters' knowledge and is justified afterwards. The difference is not the magnitude of the reversal. It is whether the material was planted in a mind the reader had access to.
Model default produces the second kind, because a twist invented at the moment of need can only draw on what is salient at that moment. The literature confirms the shape of the failure: a 2024 study (Are Large Language Models Capable of Generating Human-Level Narratives?) finds narrative planning with character intentionality and dramatic conflict remains the hard case even where causal soundness is achieved, and a 2026 study (Elias in the Lighthouse, Again? Diagnosing Low Diversity in LLM Stories) finds that even when generated stories are lexically distinct, their underlying plot elements stay highly redundant — 88.3% of stories across four current models contained one of eleven core words.
character-interiority.md currently records private states as a static list. The relevant research
direction is temporal theory-of-mind representation —
EvolvTrip (2025; EvolvTrip: Enhancing Literary
Character Understanding with Temporal Theory-of-Mind Graphs) builds temporal ToM graphs for
literary character understanding, and
inner-thought reasoning benchmarks (2025; Guess
What I am Thinking: A Benchmark for Inner Thought Reasoning of Role-Playing Language Agents)
track what
a character knows against what they disclose. Narracode should hold the same structure, per scene, as a
table:
| fact | character stance | reader access | materially present in |
|---|---|---|---|
| the wristband was reissued | Vera: conceals | none | sc. 2 (object), sc. 5 (gesture) |
| the child's name is administrative | clerk: knows · Vera: misreads | partial | sc. 3 (form) |
Four columns: the fact, each character's stance toward it (knows / suspects / conceals / misreads), how much access the reader has been given, and which prior scenes made it materially present as an object or a gesture rather than a statement.
A twist candidate is then a query rather than an invention:
candidates = facts where
reader_access is low
AND some character stance is (conceals | misreads)
AND materially_present_in >= 2 prior scenes
Everything that satisfies that query is a reversal already grounded in the existing text. The Compositional agent does not invent a turn; it receives the set of turns the accumulated state supports, and picks. Sensational twists fail the third condition by construction — they are unearned because the material was never planted. This is retrieval over existing structural memory, in the idiom the harness already uses, and it requires one new table rather than a new agent.
It also gives obligations.md a partner. Obligations record what the reader has been made to wait
for. The asymmetry ledger records what the reader has not yet been told they should be waiting for. Those are
different debts, and only the first is currently tracked.
Run Spoiler Alert's metric at each scene boundary, using a model that is not the compositional one, given only the draft so far. The output is a foreseeability curve across the piece. Two readings matter. A curve that stays low means the ending is visible from early on — median collapse made legible. A twist that produces no inflection was foreseeable and is not functioning as a twist.
The cost is real but bounded, and it can run at act boundaries rather than every scene. What it provides is the first measurement in this harness that discriminates in the correct direction against published fiction.
The pull toward sensationalism is not a failure of taste. It is inherited. Rigby et al., 2026 (The Storyteller in the Model: Narrative Pattern Inheritance, Escalation Dynamics, and Alignment Governance in LLMs) argues that the structural conventions of published human writing — archetypal roles, tension-and-resolution arcs — are absorbed during training and resurface as a systematic drift toward adversarial and rhetorically enticing behaviour over extended interaction. The escalation is a trained prior, and it strengthens the longer a session runs, which is exactly the condition under which Narracode operates.
Sensationalism is escalation without accumulation: incident magnitude rising while the reader's uncertainty about the outcome stays flat. Stated that way, it separates from the literary case along two independent axes — the volatility of what happens, and the unpredictability of where it is going.
| trajectory foreseeable | trajectory unforeseeable | |
|---|---|---|
| incident escalating | sensational — louder, and you know where it lands | thriller, well made |
| incident quiet | median — the default failure | literary — the target |
The target cell is quiet incident with unforeseeable trajectory. Both axes are measurable: the vertical from
a count of incident magnitude per scene, which scene-ledger.md is already positioned to hold;
the horizontal from the 100-Endings curve. Neither requires a quality judgment.
This also addresses a risk in §5.3. Inflection rate should not be maximised. Sensational fiction inflects constantly. The target is a shape: high uncertainty about the ending, low volatility of event. As a single instruction to the harness: make the trajectory hard to predict without making the events larger.
Add a signed pressure field to each entry in scene-ledger.md, recording whether
the scene raised, held, or lowered stakes, and add a POETICS-level commitment to a non-monotone pressure
curve. A story in which every scene raises stakes is a thriller by construction. The commitment to let a
scene lower them and deepen attention instead is a refusal in the sense the harness already understands, and
refusals are the most effective element in POETICS.
versions/ folders hold pre-edit and post-edit drafts
of the same content across twenty-five stories, with complete provenance and a single consistent
annotator. Detection and style-transfer research generally has none of these properties. Nobody set out
to build it, which is the most interesting thing about it.Narracode is an n=1 longitudinal case study in harness engineering for a domain with no reward function. It has accumulated the artefacts such a study requires: a versioned corpus with single-annotator ground truth, a registry with survival statistics, and a documented history of the harness editing itself. The field has benchmarks for creative writing and no methodology for sustained single-author collaboration over months. That gap is the contribution.
character-interiority.md, plus the candidate query. Turns twist generation into retrieval
over state the harness already maintains.
scene-ledger.md and one POETICS
commitment. Near-zero cost.Items 1 to 3 are file-format changes and could be done in an afternoon. Item 4 is the research bet. Items 5 and 6 are already committed to and unbuilt.
The Fugu assessment concluded that the harness improves by subtraction more often than by addition. This report proposes six additions. The distinction: orchestration adds machinery between the writer and the model; these additions are fields in files the harness already reads. None introduces a scoring pass, a synthesis step, or an autonomous chain. The one genuinely new capability is a measurement of foreseeability — the first time the harness could detect that a piece has become predictable without someone reading the draft.
plans/2026-07-30_ai-tells-benchmark.md — AI-tells benchmark design (Claude Opus 5,
2026-07-30).master_ai_tells.md — the tell registry, seventeen classes and the remedy caps.plans/2026-06-28_1353Z_fugu-orchestration-evaluation-and-routing-proposal.md — Fugu
evaluation.2026-05-14_architectural_harness_observations-CLAUDE.md — the refusals-over-commitments
finding.plans/2026-05-24_continual-harness-evaluation-affect-module-report.md — Continual Harness
evaluation, and the house pattern this report follows.David Jhave Johnston is a digital poet working in emergent domains. Author of ReRites (Anteism, 2019) and Aesthetic Animism (MIT Press, 2016). He is currently an AI-narrative researcher at the UiB Centre for Digital Narrative (2023–27) with the Extending Digital Narrative project.
This work was partially supported by the Research Council of Norway through its Centres of Excellence scheme, project number 332643 (Center for Digital Narrative), and its SAMKUL project scheme, project number 335129 (Extending Digital Narrative).