Why Fugu-style Orchestration Enhancements to Narracode will not improve it

A consideration

June 28, 2026  |  Jhave  |  with Opus 4.8

On 22 June 2026, Sakana AI released Fugu (arXiv:2606.21228) — a learned orchestrator model that wraps a swappable pool of frontier models behind one API, casts them into Thinker, Worker and Verifier roles, and synthesizes their outputs into a single answer. Because Narracode already separates composition into roles, the temptation is obvious: bolt an orchestration layer on top and let a coordinator route each pass to the best model. This page records why we decline. The short version: Fugu is a superb answer to a question Narracode is not asking, and importing it would cost the harness exactly the thing it exists to protect.

1. The rhyme is not an opportunity

Narracode and Fugu share a shape — a multi-role scaffold over a model — and the roles even line up: the Initiator/Structural pass thinks, the Compositional pass works, the Reflexive pass verifies. Narracode's AUTO_MODE already does Fugu's headline trick of assigning a high-critical-ceiling model to architecture and critique and a fluent model to prose. So on the one axis where Fugu is celebrated, Narracode independently arrived at the same asymmetry years of model-churn ago.

What Fugu adds beyond that — learned coordination, dynamic per-request scaffolds, and output synthesis — are precisely the moves Narracode must refuse. Fugu is trained to maximize a verifiable objective (it scores against SWE-Bench, Terminal-Bench, GPQA — all gradeable). Literature has no such objective. Narracode exists because the prose that scores highest on any automatic or consensus metric is the failure mode: competent, fluent, faintly sentimental writing that reads like a thousand models writing about the same thing. Synthesizing three frontier models toward one coherent answer is averaging, and averaging is the median collapse the whole harness is built to interrupt. We already have the only part of Fugu worth having.

2. Routing tables rot

The mildest version of the proposal was a declared routing table — bind the Thinker role to one model, the Worker to another. Even this we now decline, because a table that names models is a table that is wrong within a season. Models change every few months. Narracode is used across providers — with Antigravity and Gemini, with OpenAI's Codex, and inside the Anthropic ecosystem — and a hard binding to Sonnet 4.6 or Opus 4.8 becomes a maintenance liability the moment the next model ships.

If anything is durable, it is the role description — "the Worker wants prose temperament; the Verifier wants the highest critical ceiling you can reach today" — stated in provider-agnostic language and left for the human to satisfy with whatever model is best this month. That is not orchestration. It is a sentence in POETICS.md. Building machinery to enforce a binding that should stay fluid is negative work.

3. Model choice is an aesthetic judgment, not a reward function

Consider how the current assignment was actually reached. Sonnet 4.6 became the Compositional model not because a benchmark crowned it but because its prose was better — arguably a consequence of the dissatisfaction its own system card expressed, which seems to have made it a more interesting writer than the analytically stronger 4.8. Fable may well be a better writer than both. GPT Sol or Terra may supersede everything next quarter. These judgments are real, they matter, and they are irreducibly subjective. They are made by a human reading sentences and feeling whether they have earned themselves.

There is no automatic "this paragraph is better than that paragraph." There is only the prompter's silent micro-edit in the margin.

This is the wall Fugu cannot climb into our domain. Fugu's whole training paradigm — evolutionary strategies, RL against end-task reward — presupposes a computable signal. Literary quality is not one, and the moment you approximate it with a frontier-teacher or a consensus score, you have re-installed the median. So the one decision Fugu most wants to automate — which model, in which role — is the one decision that, here, must stay human. Automating it does not improve Narracode; it removes its judgment.

4. Orchestration buys complexity, and complexity is context rot

Every layer of coordination has a token cost. The harness is already a sizable document that must sit in context while the model composes; the structural memory grows with the work. Add an orchestration protocol — routing logic, role-handoff conventions, synthesis rules — and the scaffolding starts crowding out the manuscript. As the harness grows, compression and summarization become unavoidable, and compression is lossy in exactly the register that matters: the specific phrasing already on the page, the precise refusal, the motif that must not return yet. Context rot is not a hypothetical; it is the tax on every clever addition. The most load-bearing rule in the entire harness is a negative one — the silence after a draft that forbids the chain into self-critique. Complexity is the opposite of that discipline. The harness improves by subtraction far more often than by addition.

5. The platform is already "enhancing" your prompt

There is a deeper reason orchestration is redundant. The platforms Narracode runs inside increasingly rewrite and enrich every prompt in the background. This lineage runs back to image generation — DALL·E 3 popularized it in 2023 with an automatic prompt-rewriting step that quietly "upsampled" a terse prompt into a florid one before the model ever saw it (Midjourney and others followed). Coding harnesses now do the analogous thing: Claude Code and its kin expand, reframe, and scaffold each request before it reaches the model.

For literature this background enhancement is not neutral — it is actively adverse. Generic enhancement pulls every prompt toward the well-formed, the conventional, the legible: the statistical center of the training distribution. That is the median Narracode exists to refuse. The value of the harness was never generic scaffolding, which the platform now supplies for free; its value is the specific negative pressure — the Refusals, the burned phrases, the forbidden moves — that no generic enhancer will ever add because no benchmark rewards it. Stacking a bespoke orchestration layer on top of a platform that is already enhancing toward the mean is paying twice to be pulled in the wrong direction.

Conclusion: the enhancement is restraint

Fugu is excellent engineering for tasks with a gradeable answer. Narracode has no gradeable answer, already embodies Fugu's one transferable insight, and would pay for the rest in brittleness, lost judgment, and context rot. The provider-agnostic, constraint-driven, deliberately thin design is not a stage Narracode has yet to outgrow — it is the design. In a moment when every platform races to enhance, orchestrate, and synthesize, the most literary thing a harness can do is hold its constraints and decline. The enhancement is to not enhance.

Bio

David Jhave Johnston is a digital poet working in emergent domains. Author of ReRites (Anteism, 2019) and Aesthetic Animism (MIT Press, 2016). He is currently an AI-narrative researcher at the UiB Centre for Digital Narrative (2023–27) with the Extending Digital Narrative project.

Funding

This work was partially supported by the Research Council of Norway through its Centres of Excellence scheme, project number 332643 (Center for Digital Narrative), and its SAMKUL project scheme, project number 335129 (Extending Digital Narrative).

All works and media on Glia.ca by David Jhave Johnston is licensed under CC BY-NC-SA 4.0 Creative Commons Attribution Non-Commercial Share-Alike