RDI factory session forensics

What happened during the 48-hour factory run

A plain-English account of session 019fa53e-b77e-7152-a82d-93bf24a0d063: what it built, what it proved, where it looped, and what remained unfinished.

Started Jul 27, 2026 at 4:26 PM ET Snapshot through Jul 29 at 4:11 PM ET HypAware + Git + review receipts Prepared Jul 29, 2026

The run produced meaningful factory upgrades, but it did not finish the job Diego asked it to do.

Research checking became much stronger. The final accepted source and package reviews scored research quality and claim integrity at 10.0. The factory gained a writing contract, separate report structures for market validation and market exploration, isolated review stages, source-capture tools, and hundreds of tests.

The accepted cold-reader result still scored writing at 7.6 and readability at 7.1, barely above the 7.5 and 6.4 starting scores. The run then failed to prove that the system worked on a second market. It ended with an uncommitted worktree, an unfinished review chain, and no valid final weighted score.

Best one-line description: substantial implementation, incomplete proof. The system for checking facts improved more than the reader experience that started the run.

The run in numbers

HypAware records the whole session as one root conversation plus 190 additional conversation threads. Token totals include model input, output, and reused cached context. Reasoning tokens are a subset of output tokens.

47h 45m Wall-clock span From the first prompt to the latest recorded child-thread event at the audit cutoff.
191 Conversation threads One root thread and 190 child or nested threads.
8.93M Output tokens About 4.06M were reasoning tokens.
2.69B Total tokens processed Includes 2.56B cached input tokens and 122.2M uncached input tokens.
20,027 Tool-call records 9,080 shell calls, 3,831 composed calls, and 1,719 file patches led the count.
260 Agent-spawn calls 233 distinct task labels. About 70% of output tokens came from non-root threads.
19 Accepted review receipts Seven passed. Twelve required revision.
14 Commits on the working branch 148 committed files changed between the frozen baseline and the final session commit.

Daily model output

Jul 27, partial day
1.04M
Jul 28
4.89M
Jul 29, partial day
3.00M

Source: HypAware ai_gateway_messages, refreshed at report time. Token cost was not recorded, so this report does not estimate dollars.

What happened, in order

Jul 27, 4:26 PM to 5:03 PM ET

1. Audit the latest report and freeze the target

Diego asked for a deep audit of recent factory changes, with prose and consumability as the main concern. The audit found that the latest Synter work had strong research but still made the reader work too hard.

The run froze a seven-part scorecard at commit a3283ef. The weighted baseline was 7.4. Diego then expanded the assignment from an audit into an open-ended implementation run: reach the highest feasible score without changing the scorecard or cutting useful depth.

“I’m happy with the depth in general and it seems the claims verification is airtight. Consumability has been the biggest issue.” Diego, Jul 27 at 4:34 PM ET
Jul 27, 5:03 PM to late evening

2. Turn writing quality into factory rules

The run stopped treating the report as a final formatting step. It moved report creation and reader testing inside the factory’s mandatory assessment loop.

It wrote a repository-wide report-writing contract, separate story structures for validation and exploration, a frozen report brief, a cold-reader recovery test, and stronger rules for exact claim wording. It also began rewriting Synter and expanding tests around report generation, evidence binding, and rendering.

This was the most important factory-level decision in the run. The old process checked research before final synthesis. The new process checks the report that Diego actually reads.

Jul 27 evening through Jul 28

3. Rewrite Synter, then discover that “airtight” still contained real errors

Fresh source review found factual and structural defects that earlier work had missed. The largest was the claim that Marin Software dissolved. Marin entered Chapter 11, reorganized, and continued under new ownership. The run corrected the brief, primer, findings report, reference report, and conversation cards.

Other corrections included Marin’s $7.15 million Google revenue-share amount, the distinction between revenue divided by managed spend and a contract take rate, Optmyzr price ratios, the bottom-up market estimate, platform write-access claims, Synter’s conflicting public prices, and the evidence standard for the True Classic case.

A first blind-reader pass recovered the story and scored an obsolete package at 9.2, but a page overlap and later factual corrections invalidated that result. The invalidation was correct. The problem was that the run kept changing both the factory and its test document at the same time.

Jul 28 through Jul 29 morning

4. Build the verification system while repeatedly reopening the evidence

The run built an isolated review runner with immutable receipts, exact input hashes, reviewer provenance, stage order, package identity, and source checks. It also built deterministic pricing-page and deployment-census capture tools, stronger seeding rules, report preflight, PDF checks, and a much larger test suite.

Reviews kept finding new classes of defects. Each factual edit correctly invalidated later stages, but passing source reviews were not treated as checkpoints. Validator logic and the Synter source bundle continued changing, which caused new reviews to restart the chain.

One CI simulation also ran an initializer with --force in the live worktree and overwrote state files. The run restored the pre-accident versions from a temporary copy and removed the generated scaffolding. No lasting loss is visible, but this consumed another recovery cycle.

“Do you think you will ever converge? You have been running for 24hrs.” Diego, Jul 28 at 9:18 PM ET
Jul 29 morning to 4:11 PM ET

5. Reach clean source scores, fail the reader gate, and expose the transfer problem

The source bundle eventually reached accepted 10.0 scores for research quality and claim integrity before and after packaging. A blind reader recovered the decision story. The next reader still returned revise: writing 7.6, readability 7.1, visual execution 6.0, and decision usefulness 8.8.

That reader found four contained defects: internal process language, three local links, hidden opening text, and a stranded reference sentence. The run repaired those defects in source, but it did not complete a new accepted reader chain on the repaired package.

The run then tested transfer to GistFlow. The attempt exposed dozens of source-contract failures, claim wording drift, interview-plan drift, internal process language, and weak contact routes. This proved that the machinery was still partly tuned to Synter rather than proven as a general factory standard.

“It’s the content, right? I don’t give a shit about the PDF being pixel-perfect.” Diego, Jul 29 at 2:46 AM ET
“I fear that you might have just gone off with one document and just iterated forever.” Diego, Jul 29 at 11:24 AM ET

What the run actually built

The work falls into three buckets. The distinction matters because the raw 297,436 added lines make the run look much larger than the reusable system change.

Durable factory controls

  • A plain-language report-writing contract.
  • Separate story contracts for validation and exploration reports.
  • A frozen report brief that defines the reader’s job before drafting.
  • Exact binding between primer claims, report claim cards, and interview missions.
  • An isolated review runner with immutable receipts and reviewer provenance.
  • Pre-package, post-package, blind-reader, reader, and assessor stages.
  • Deterministic pricing-page and deployment-census capture tools.
  • Stronger seed, render-manifest, citation, and package identity checks.
  • About 18,165 added script lines and 12,086 added test lines in the committed diff.

Real Synter research improvements

  • Corrected Marin’s bankruptcy outcome and primary-source financial facts.
  • Separated agency-replacement economics from cross-platform control software.
  • Reframed True Classic as the closest current case, with two of four tests met.
  • Captured Synter’s conflicting live prices and routes to sale.
  • Rebuilt the deployment census at the result level.
  • Rewrote the interview plan so each witness is asked for evidence they could possess.
  • Bound six claim reads, confidence labels, supporting facts, and open questions across the package.

Generated proof and captured data

  • Nineteen accepted review archives with prompts, frozen inputs, responses, artifacts, and receipts.
  • Large pricing and deployment-census capture files stored in both the live research tree and review snapshots.
  • Several Synter HTML and PDF package generations across Jul 27, Jul 28, and Jul 29.
  • Blind-recovery and reader-validation artifacts for specific package hashes.

Partial transfer work

  • GistFlow was migrated toward the current exploration-mode schema.
  • The migration added a report brief, updated interview plan, target list, source records, and rendered files.
  • The transfer audit did not pass. GistFlow remained a hand migration with many contract errors.
  • No second live market completed the same source, reader, and assessor chain.

What the review evidence says

The frozen scorecard never changed. That is good. The final state does not have one valid seven-part score because different dimensions were last accepted against different package states.

Dimension Frozen baseline Latest accepted evidence Plain-English read
Research quality 9.0 10.0 Converged Accepted before and after packaging.
Claim integrity 9.2 10.0 Converged Exact source and package checks passed.
Writing quality 7.5 7.6 Barely moved The latest accepted reader still found process language.
Readability 6.4 7.1 Improved, still weak The decision was recoverable, but the reader had to work.
Visual execution 6.8 6.0 Regressed Opening text was hidden in the accepted package.
Decision usefulness 7.8 8.8 Materially better The reader recovered the verdict and next diligence work.
Factory reproducibility 4.0 No accepted score Unproven GistFlow exposed transfer failures.
Weighted score 7.4 None authoritative Do not combine The accepted scores belong to different content states.

Review volume and outcomes

Pre-package source reviews
15
Post-package reviews
2
Blind-recovery reviews
1
Reader reviews
1

Seven passes

Several source reviews eventually reached 10.0. Post-package source review also reached 10.0. Blind recovery passed. These results show that the source and traceability system can work on Synter.

Twelve revisions

Revisions found real defects: source mismatches, unsupported wording, stale status, claim labels that changed meaning, weak interview routes, missing evidence, and reader-facing process language. The volume also shows that the run kept changing its inputs after earlier passes.

Where the effort went

Two measurements tell the same story. The model spent most of its output in child threads, and the committed line count was dominated by captured evidence and validation machinery.

Model output by thread type

Child and nested threads
70.3%
Root thread
29.7%

HypAware recorded 292 continuation summaries and 54 explicit checkpoint-compaction requests across the session. Context handoff became a substantial part of the work. The task names show repeated audits of prose, claims, source fidelity, rendering, review protocols, convergence, and factory transfer.

Committed additions by category

Category Files Added lines What the number means
Synter review archives 31 130,928 Frozen copies of inputs and evidence for review receipts.
Synter live artifacts 23 130,858 Mostly large machine-captured pricing and census records.
Scripts 11 18,165 Review runner, report preflight, capture tools, renderer, and seeding.
Tests 15 12,086 Regression checks for the new contracts and tools.
Factory contracts and documentation 31 2,627 Writing rules, prompts, skills, templates, and operating docs.
GistFlow fixture 15 1,833 Partial exploration-mode migration.
Report-quality state 4 236 Scope, plan, scorecard, and progress record.
Other 16 703 Workflow and supporting changes outside the main categories.

Why 297,436 added lines is misleading: three machine-captured source datasets account for 126,717 lines and were committed twice, once live and once inside a review archive. Those six copies make up 253,434 lines, or 85.2% of all committed additions.

The durable factory code and contract work was still substantial. Scripts, tests, factory contracts, documentation, and report-quality state added 33,114 lines. That is a better measure of reusable implementation than the raw diff.

Why the run failed to converge cleanly

1. The finish line was intentionally unbounded

“Highest feasible score” has no natural stopping point. The scorecard stayed fixed, but each fresh reviewer could find another defect class. The run needed a time budget or a rule that froze source inputs after a clean pass.

2. The validator and its fixture changed together

The run was building review machinery while rewriting Synter. A defect in either one reopened the chain. This made it hard to tell whether scores changed because the report improved or because the test became stricter.

3. Correct invalidation lacked checkpoint discipline

Rechecking after factual edits was right. Continuing to change source inputs after a clean pass was the process failure. The run needed to freeze a passing source bundle, then move to writing and reader tests.

4. Too much work stayed centered on Synter

Synter became both product and test fixture. The late GistFlow audit exposed this. A smaller second-market test earlier in the run would have found Synter-specific assumptions before the system grew around them.

5. Agent and context churn obscured progress

The session used 260 spawn calls, 191 threads, 292 continuation summaries, and 54 checkpoint requests. This produced useful independent checks, but it also multiplied state handoffs and repeated setup work.

6. PDF work outran its value

Some rendering checks protected content. Work on exact PDF parity, metadata, and portability went further than Diego valued. Diego had to restate that conversion was a pass/fail integrity step and that content mattered.

The user had to steer the run back to the job

Diego asked for the current score after about 16 hours, questioned convergence after 24 hours, challenged repeated permission requests, deprioritized pixel-perfect PDFs, and later asked whether the run had become a one-document loop. Those were useful interventions. A well-controlled run should have surfaced the same concerns itself.

What was done, and what was still open

Done and supported by evidence

  • The frozen scorecard and report-writing contract exist.
  • The report now sits inside the mandatory assessment loop.
  • Review stages produce isolated, immutable receipts.
  • Synter source quality and claim integrity reached accepted 10.0 results.
  • The package preserved exact claim labels after the final package correction.
  • The factory has deterministic pricing and deployment-census capture tools.
  • The branch contains 14 commits after the frozen baseline.
  • The committed diff contains broad reusable code and test work.

Open or unproven at session end

  • No accepted reader pass exists for the final repaired source.
  • No authoritative seven-part weighted score exists.
  • Writing and readability did not reach convergence in accepted evidence.
  • GistFlow did not pass the current factory contract.
  • No second live market completed the full review chain.
  • The worktree still held 29 tracked changes and 33 untracked entries at the audit snapshot.
  • Unit 6 remained active, Unit 7 and Resolve remained open.
  • Branch cleanup never started.

The honest final assessment

This was not a worthless run. It corrected material research errors, built a much stronger verification system, and encoded writing and reader recovery into the factory. Those changes should survive Synter.

It also did not deliver the stated outcome. The accepted evidence shows a large gain in fact checking and decision usefulness, a small gain in writing, a moderate but incomplete gain in readability, and no proof of general factory reproducibility. The run overbuilt validation around one report and left the decisive transfer test unfinished.

Evidence and method

The report uses four evidence layers. Session self-reports were treated as leads and checked where a repository record existed.

HypAware
Local ai_gateway_messages cache, refreshed at report time. Session filter: 019fa53e-b77e-7152-a82d-93bf24a0d063. Snapshot cutoff: 2026-07-29T20:11:14Z.
Git range
Frozen baseline a3283ef through final session commit 5142018 on branch report-default-and-funktor-haystack.
Review records
Nineteen receipts under Synter review-runs. Stage and verdict counts come from the receipt JSON, not the agent’s progress messages.
Score evidence
Frozen scorecard, report-quality project state, and accepted reader receipt 5feae227.
Repository residue
Git status at audit start: 29 tracked changes and 33 untracked entries. The HTML report itself was created after that snapshot and is not included in those counts.
Limits
HypAware did not record dollar cost. Tool error flags were not reliable enough to count failed commands. Daily conversation counts overlap because one thread may appear on more than one day.