Singapore · Hong Kong · Redwood CityExecution systems · Verified autonomy · Enactive Reality
We compile intentinto verified work.
Tunneling builds closed-domain execution systems: asynchronous graphs of
a hundred or more specialised agents that operate on an institution's own
ontology, drive the professional tools the work actually runs on, verify
every result against a domain checker, and halt at a human gate before
anything irreversible happens.
Quantum tunnelling — a particle crossing a barrier classical mechanics says it cannot cross
The name is the thesis. Most enterprise AI stops at the barrier between what a model can articulate and what an institution can be held responsible for. We build what gets through it.
Scroll
0Nodes in the largest closed-domain graph now in production
0Pipeline stages under one deterministic scheduler
0Reality gates between an agent and an irreversible action
0Working days to a written register you can act on without us · the pilot that follows runs four to eight weeks
Every institution runs on one process that three people understand, that cannot be paused, and that nobody has written down. That is where this work begins.
01The platform
An execution layer, not a chat surface
A language model produces text. An institution needs a completed task: a filing that reconciles, a drawing that can be manufactured, an allocation that survives audit. Between the two sits everything that is hard — the domain's objects and rules, the professional software the work actually runs in, the checkers that decide whether a result is correct, and the authority to act.
The Tunneling stack is six layers that close that distance. Value rises as you move down it. So does the engineering cost of being wrong. We build downward only as far as an institution is prepared to be accountable.
Value ↓ increasesExecution ↓ hardens
L1 · Ontology
The institution, made computable
An ontology is not a database schema. It is the set of objects an organisation argues about, the relations that constrain them, and the permissions that decide who may change what.
A contract, a lot, a work order, a counterparty, a tolerance band, a shipment, a learner. Each carries state, version history, provenance and an access class. Agents never see raw tables; they operate on typed objects with declared side effects, which is what makes their behaviour reviewable after the fact.
Entity model
Objects, attributes, lifecycle states and the events that move between them.
Relation graph
Supply, ownership, substitution, dependency, causality — traversable by agents, constrained by rules.
Permission classes
Read, propose, execute. Most agents never hold more than propose.
Provenance
Every value carries a source, a timestamp and a confidence. Unsourced values cannot enter a claim.
L2 · Compiler
Intent is compiled, not improvised
Goal → plan → typed tool call → state read-back → verification → repair → delivery. The loop is fixed; only its contents vary by domain.
The critical move is state read-back. A model reporting that it succeeded is not evidence of success. After every action the system re-reads the real document, the real solid body, the real ledger row, and compares it against the intent that produced it. Failures are localised and repaired in place rather than regenerated from scratch — which is what keeps a forty-step task from collapsing at step three.
Typed action
create_sketch, export_step, post_journal — never synthetic mouse coordinates.
Read-back
Documents, constraints, errors, versions and resource state re-read from the tool itself.
Localised repair
Amend the failing step, record the cause and the cost, keep the rest of the trajectory intact.
Delivery contract
A run ends in an artefact with a manifest: inputs, calls, checks, approvals, outputs.
L3 · Graph
A hundred specialists, one scheduler
Closed-domain work does not decompose into one clever agent. It decomposes into a hundred narrow ones, each with a single responsibility, a typed contract and its own failure mode.
Our production graphs run in the low hundreds of nodes across a dozen stages: classifiers, judges, planners, interpreters, architects, writers, renderers, critics, repair planners and gate judges. Stages fan out concurrently and rejoin at barriers only where a decision genuinely needs the whole set. Every node is individually replaceable, individually evaluable, and individually attributable when something goes wrong.
Deterministic scheduling
Concurrency, retries, checkpoints and resumption are the runtime's job, not the model's.
Barriers by necessity
Synchronise only where cross-item context is required — deduplication, quota, consistency review.
Node-level evaluation
Each node carries its own test set. Regression is detected per node, not per product.
Cost shaping
Model tier, reasoning effort and branch depth are assigned per node against its actual difficulty.
L4 · Verification
Looking correct is not being correct
Every class of claim gets its own checker. Numerical claims are recomputed. Geometric claims are re-measured. Cited claims are re-located in the source. Rendered claims are re-opened and inspected.
This is the layer that separates a demonstration from a deployment. A plausible sentence and a correct sentence are indistinguishable to a reader and trivially distinguishable to a checker. We build the checker first and treat any claim class without one as out of scope until it has one.
Numeric & symbolic
Recomputation, unit and domain checks, equivalence testing, tolerance bands.
Structural
Schema conformance, watertight solids, constraint trees, collision and navigation checks.
Evidentiary
Every assertion mapped back to source, date and page. Unsupported assertions are removed, not softened.
Runtime
Render, open, execute, screenshot, diff. If it cannot be opened, it did not ship.
L5 · Gates
Six gates between an agent and the world
An agent may prepare an irreversible action. It may not complete one without explicit authority. This is a design constraint, not a configuration setting.
Capability allowlist, sandboxed execution, state read-back, result verification, human approval, audit and rollback. Payments, manufacturing releases, regulatory filings, equipment commands and identity changes sit permanently behind the fifth gate. Every action is traceable, reversible and stoppable.
Allowlist & sandbox
Pinned tool versions, resource ceilings, scoped network and filesystem access.
Approval routing
The named human who owns the decision, with the evidence pack attached to the request.
Audit ledger
Immutable record of inputs, model calls, checks, approvals and outputs for every run.
Rollback
Every run is a revertible transaction against the ontology, including partial failures.
L6 · Corpus
The moat is verified trajectories
Not more prompts. A trajectory: initial state, correct operations, the failure path, the repair, the acceptance criterion, and what actually happened in the world afterwards.
Public data teaches a model what professional work looks like. It does not teach what to do when step nine contradicts step four. That knowledge exists only inside institutions and only as a by-product of doing the work under verification. Every deployment we run produces it. It belongs to the client; the structure that makes it reusable is what we bring.
Failure traces
The scarce asset. Correct runs are cheap; instructive wrong runs are not.
Transferable structure
Trajectories captured in one tool generalise into the next tool of the same class.
Continuous evaluation
Yesterday's accepted runs become today's regression suite.
Ownership
Data use, ownership, source code and derivative rights are fixed in writing before work starts.
Cross-cutCut by what the system has to be right about — where ground truth lives, and what it costs to check it. Not by industry name, not by company size, not by geography.
Where ground truth lives decides which layer carries the load — a disputed claim is settled by sourcing and critics (L1, L4), a machine state is settled by the plant's own systems and a signature (L2, L5) — and the cut is imperfect: a fab's export-control filings behave like Binding procedure, not Physical state of record, so a single client can sit in two clusters and should be scoped that way.
Layer / cluster
A · Evidence under disputeReconciling sources that disagree IND-01 Capital markets · IND-02 Banking, insurance & credit risk · IND-03 Legal, audit & tax
B · Physical state of recordMachine and material state against the records IND-04 Semiconductor · IND-05 Supply chain & trade · IND-06 Plants, utilities & grid · IND-07 Discrete engineering
C · Binding procedureConformance to an enforceable rule IND-08 Pharma, medtech & clinical · IND-09 Healthcare delivery & payers · IND-10 Public sector & infrastructure
D · Capability in a personSkill transfer, measured after exit IND-11 Education & workforce capability · IND-12 Interactive media & simulation
L1Domain Ontology
Issuers, instruments, events and sources; every claim typed to who said it and when.
Lines, tools, lots, SKUs and sites keyed to the IDs the MES and WMS already carry.
Clauses, obligations, filings and deadlines, each bound to the citation it came from.
Skills, prerequisites, roles and evidence of mastery — not courses or lesson lists.
L2Reality Compiler
Drives the terminal, data room and CMS; every write is a typed action with a diff.
Writes into MES, WMS and CMMS through their APIs, not by scraping operator screens.
Drafts inside the firm's own document and matter systems, versioned, never in a chat box.
Runs the professional tool the person is graded in — CAD, IDE, trading sim, edit bay.
L3Asynchronous Node Graph
7 parallel evidence branches, then normaliser, conflict resolver, coverage reviewer.
A branch per line, tool or lane; fan-in waits for the slowest sensor feed to land.
One branch per obligation; the graph runs narrow and deep rather than wide.
A branch per learner or scene; the graph runs at session pace, not overnight batch.
Any change to a setpoint, recipe or shipment halts at a named engineer's signature.
The signing professional approves before filing; the gate records who approved, and when.
Lightest of the four: certification, and any live-system or real-money step, stay gated.
L6Trajectory Corpus
Claim map, source map and call-trace archive replay any published item end to end.
Each run keeps the fault-detection trace: the ordered checks that located the fault.
Builds the firm's precedent file — matter, rule version, reasoning, outcome on file.
Stores what the person did, not what they watched; scored on transfer after exit.
We are deliberately small. There is no platform licence to defend, no seat count to grow and no reference architecture we need you to adopt. An engagement begins with one workflow, a named owner and an acceptance test — and ends when that test passes or we tell you it will not.
02Reference architecture
155 nodes, eleven stages, one audit trail
A real graph, running in production: 114 reasoning nodes and 41 deterministic workers across eleven stages. It ingests one event and returns four independently verified publishable products plus a complete evidence ledger, with no human touching the middle of it.
Scale is the whole engineering problem. Almost anyone can wire five agents together. It changes character at fifty, where concurrency, partial failure, cross-product consistency and per-claim attribution stop being incidental and become the design itself. Who runs this one, how many people use it and what it replaced are in section 07.
Domain
Closed. Bounded vocabulary, bounded sources, bounded output formats. Closure is what makes verification tractable.
Execution
Asynchronous. Stages fan out concurrently; barriers appear only where a decision needs the full set.
Failure
Local. A failed node drops its branch and is repaired in place. It does not take the run with it.
Output
Gated. Nothing reaches a reader until node AI-116 has cleared the claim map and the evidence ledger.
Node topology — stage order left to right, concurrency verticalLive scheduler view · illustrative
This graph is one instance, not the firm's whole capability. A hundred-node asynchronous graph is the right shape when the work fans out into many independent artefacts under one deadline. Most of the engagements in section 04 do not need one: a reconciliation pilot runs at 20 to 40 nodes and a healthcare packet graph is smaller still. Section 03 lists what the stack does; this section shows what it looks like at the top of its range.
What breaks at scale
Contradiction between products. The poster says one thing, the article another, the video a third. Stage 09 exists only to catch that, and it is the stage clients underestimate most.
What we refuse to skip
Stage 10. Every claim in every artefact is mapped back to a source, a dataset and a rendering job. If a number cannot be reproduced from the ledger, the product does not publish.
What this transfers to
The topology is domain-independent. Substitute the ontology, the branch interpreters and the checkers, and the same eleven-stage shape covers research, engineering release, and regulatory reporting.
AGENT TRACE — STAGE 04 · SINGLE-EVENT RESEARCH ROUTEillustrative
03Capability taxonomy
Eight things the system does
Sector names are a poor way to scope this work. Two banks can want opposite things and a bank and a fab can want the same thing. So the catalogue below is cut by the unit of output — what the system has to hand back before anyone will call the run finished — because that is what decides which checker has to exist and how long it takes to build.
Each entry states its input, its checker, its gate and the earliest point at which you could accept or reject it. The sector practices in section 04 are these eight capabilities, recombined and given a vocabulary.
If the answer is a number
Your systems disagree with each other, or with the physical world. That is C-01 and C-03. Ground truth already exists somewhere in your estate; the work is locating it and proving the match.
If the answer is a document
Somebody has to sign it and could be asked to defend it. That is C-02 and C-04. Ground truth is a source or a rule, and the checker is a lookup you can run on any sentence.
If the answer is a file or a person
The work happens inside professional software, or inside someone's head. That is C-05, C-07 and C-08, and it is the expensive end.
C-01H0 · 10 DAYS
Reconciliation
Unit of outputA matched record set, plus a discrepancy register naming the field, both values, and the document that governs which one wins.
Input
Two or more record populations that are supposed to agree and do not.
Checker
Field-level match, arithmetic re-run on every line, date and term feasibility, governing-clause lookup.
Gate
Propose only. No payment release, no journal posting, no credit note without a named approver.
Layers
Load sits on L1 and L4. L2 is light — most of this is read.
Instances we have scoped or built
Supplier invoice against purchase order and goods receipt— freight, duty and tax lines included, which is where three-way matching usually fails
Custody statement against the internal position ledger— per account, per day, with the break aged
Bill of materials against the released drawing revision— per part number, per revision
Five trade documents, one consignment— commercial invoice, packing list, bill of lading, certificate of origin, letter of credit
Claim submission against the policy schedule and the treaty wording
Payroll run against employment contracts and the collective agreement
C-02H0 · 10 DAYS
Attribution
Unit of outputA claim map: one row per assertion, carrying the document, the date, the page, and the check that confirmed it.
Input
A question, a body of documents, and a written rule for what counts as a source.
Checker
Numeric recomputation against the underlying series, citation re-located at page level, coverage audit, contradiction register.
Gate
The publication gate fails closed. Uncited text is removed rather than flagged.
Layers
L1 for the source hierarchy, L3 for parallel evidence branches, L4 for the critics.
Instances we have scoped or built
Sell-side note with a disclosure appendix— and a stated basis behind every figure in it
Investment committee memo— each input carrying its licence and redistribution terms
Regulatory submission— every statement mapped to the clause it answers and the version in force
Literature review for a safety or clinical file— with the search strategy recorded, not reconstructed later
Due-diligence report— a data-room document number standing behind each finding
Client enquiry answered with its source list attached
C-03H1 · 6 WEEKS
Cause ranking
Unit of outputAn ordered list of candidate causes, each with the evidence that supports it and the one test that would settle it.
Input
An anomaly with a timestamp, and read access to the histories that could explain it.
Checker
Replay against cases your team already solved, unit and dimensional validation, a precedent applicability test.
Gate
Read-only against the systems of record. No setpoint, recipe or configuration writes, ever.
Layers
L1 for the asset and event model, L2 in read mode, L4 for the replay harness.
Instances we have scoped or built
Yield excursion on a 300mm line— fault-detection traces, defect maps, recipe edits, maintenance records, incoming material
Unplanned downtime on a rotating asset— vibration history, work orders, lubrication records, load profile
A grid constraint that keeps recurring at the same hour— telemetry, outage records, weather, market schedule
Cost variance on a capital project— against the baseline and the change register
Freight exceptions clustering on one lane— carrier events, customs holds, terminal congestion
A conversion metric that dropped— traced to a release, a segment or a supplier rather than guessed at
C-04H1 · 6 WEEKS
Rule-bound drafting
Unit of outputA draft, plus a clause trace tying every paragraph to the obligation, standard or precedent it satisfies.
Input
The enforceable text — statute, standard, contract, protocol — and the facts of the matter.
Checker
Each sentence must resolve to a cited clause. Deadline arithmetic and version currency are checked separately.
Gate
The named professional signs. No filing and no client transmission without a recorded sign-off.
Layers
L1 for the obligation model, L4 for the clause checker, L5 for the signature.
Instances we have scoped or built
Regulatory filing— each section naming the rule and the rule version it answers
Clinical study document set— against ICH guidance and the protocol version actually in force
Contract review against the firm's own playbook— deviations listed with the fallback position
Instructions for use and labelling— against the applicable device standard, in every market language
Public procurement documentation— against the tendering rules of one jurisdiction
Tax position paper— against the statute, the ruling and the group's own prior filings
C-05H1 · 8 WEEKS
Professional tool operation
Unit of outputA working artefact in the tool's own native format, with the operation log that produced it and the check that it passed.
Input
An intent, the tool itself, and the acceptance rule the artefact has to satisfy.
Checker
Re-open, re-measure, re-execute. Constraint trees, watertight geometry, formula audit, a test run that has to pass.
Gate
Writes land in a working branch or a scratch document. Release into the controlled vault takes a signature.
Layers
L2 carries almost all of it. Typed tool calls, never synthetic mouse coordinates.
Instances we have scoped or built
CAD— a parametric model rebuilt from a specification, exported, and re-measured against the tolerance table
Spreadsheet models— rebuilt with a formula audit and every output recomputed from the inputs
EDA and layout— a rule-deck run, violations split into fixable and needs-a-person
BI and SQL— a report built against the semantic layer, with the query and the row counts attached
Document and slide production— driven by your template, versioned in your own store rather than a chat window
Image and video production— every rendered frame re-opened and inspected before it counts as delivered
C-06H0 · 15 DAYS
Exception triage
Unit of outputA routed decision: the exception, the recommended action, the evidence behind it, and the person who owns it.
Input
A stream of exceptions that currently arrives in a shared mailbox or a spreadsheet.
Checker
Policy conformance, entitlement, precedent match, and a confidence floor below which an item goes to a person untouched.
Gate
Nothing auto-resolves at first. Acceptance rates are measured for weeks before any auto-clear band is agreed in writing.
Layers
L1 for the exception taxonomy, L3 for throughput, L5 for routing authority.
Instances we have scoped or built
Customs holds and classification queries— scoped to one trade lane and one jurisdiction
Alarm floods on a plant board— deduplicated and ordered by consequence rather than by arrival time
Insurance claims intake— complete, incomplete, or needs an adjuster, with the missing item named
Payment exceptions and screening hits— routed by reason code to the desk that can clear them
Service tickets— classified against the entitlement and the contractual response clock
Quality deviations— sorted by criticality against the site's own classification rules, not a generic severity scale
C-07H2 · 1 QUARTER
Multi-artefact production
Unit of outputSeveral products built from one evidence pack, plus the consistency report proving they do not contradict each other.
Input
One event or one body of evidence, and a format specification for each output.
Checker
Per-format critics first, then a cross-product consistency stage, then a claim map covering every output at once.
Gate
The publication gate holds all outputs together. One unresolved objection blocks the entire set, not just the artefact that failed.
Layers
L3 and L4 both heavy. This is the capability that forces a graph rather than a pipeline.
Instances we have scoped or built
The 155-node system in section 02— one market event becomes four verified products and one evidence ledger
Product launch pack— spec sheet, training deck, sales note and support article from one release record
Regulatory change cascade— internal briefing, procedure amendment and client notice from a single rule change
Course build— lesson, worked examples, assessment items and marking scheme from one syllabus objective
Multilingual publication— translations checked against the source claim map rather than against each other
Investor communications— deck, release and prepared remarks reconciled before any of them goes out
C-08H2 · 1 QUARTER
Capability transfer
Unit of outputMeasured transfer after exit: the person performs an unassisted task afterwards and somebody marks it.
Input
A skill definition, a checker for the artefact the learner produces, and an environment worth rehearsing in.
Checker
Symbolic and rubric checkers score the artefact. The person's judgement is not scored by a model.
Gate
No grade of record without the named examiner. No live system, no real money and no real patient inside a rehearsal.
Layers
L1, L2 and L6. This is where the trajectory corpus is worth the most and takes longest to accumulate.
Instances we have scoped or built
Zhihu MathHub— the misconception found in the working, three lines above the wrong answer
StudyHub— a question becomes a path whose state survives the session
Trading floor rehearsal on a market replay— P&L attributed to the decision that caused it
Field-engineer rehearsal on a plant model— permit and isolation steps enforced rather than narrated
Onboarding into an internal system— the trainee's actions checked against the real procedure, in a copy of it
Future Film Lab— persistent characters, and a world that remembers what the person did in it
No checker yet
Two claim classes we have tried and failed to build a checker for, and therefore decline to take on. Whether a specific piece of information is material and non-public: the test is legal and contextual, and our attempts produced confident answers with no way to verify them. And whether a photograph depicts what its caption says it depicts: image-to-claim verification held up on staged tests and fell apart on real archive material. Both stay out of scope until somebody builds the checker, and that includes us.
The cut is imperfect and worth arguing with. Reconciliation and cause ranking overlap wherever the disagreement is between a record and a machine — a lot mismatch is both. When a workflow sits across two entries we scope it as two, with separate acceptance tests, rather than pretending it is one thing.
04Sector practices
Twelve practices, one execution stack
The six layers are shared and the eight capabilities in section 03 are shared. The ontology, the branch interpreters, the checkers and the gates are rebuilt per sector, because that is where the difficulty sits and none of it generalises for free.
Each practice states the same six things in the same order: the work as it stands and what it costs, the instrumented workflows with the measure attached to each, where we think the sector is going and where we disagree with it, how the node graph is re-cut, what decides correctness and what we refuse, and which capability gets you a result fastest. Twelve is not the limit of what the stack covers — it is the list where we have either built something or scoped one closely enough to publish numbers.
A · Evidence under disputeGround truth is a source. The work is reconciling records and authorities that disagree.
B · Physical state of recordGround truth is a machine, a material or a shipment, and the records disagree with it.
C · Binding procedureGround truth is an enforceable rule, and someone signs to say the work conforms to it.
D · Capability in a personGround truth is what someone can do afterwards, and it is measured after they leave.
The cluster cut is by where ground truth lives, not by industry name, and it is imperfect in a way worth knowing about: a semiconductor firm's export-control filings behave like Binding procedure rather than Physical state of record, so one client can sit in two clusters and should be scoped as two engagements with separate acceptance tests. Cost figures throughout are observed ranges from operations of the stated size. They are not benchmarks and none of them come from a named client.
IND-01 · Capital markets, brokerage and research
A first-take note takes 90 minutes to write and three hours to make defensible.
The 155-node reference system — 114 reasoning nodes and 41 deterministic workers across 11 stages — was built for a financial-information client and still runs there. We install the same graph against your own instrument and issuer names: seven parallel evidence branches gather the record, twelve critic nodes attack what they produce, and the terminal publication gate AI-116 fails closed, refusing to release any product in which a claim cannot be resolved to a source document, a date and a page. The sector suits verified execution because the acceptance test already exists in writing — your compliance manual and your research policy are the checker specification. One limit stated up front: the graph holds state across sessions at persistence level L2, and level L3, where a system acts directly on live books and records, is gated and only partially built.
The barrier
The expensive part of research is not the prose. It is the disclosure appendix, the stated basis for the price target, the restricted-list check and the page reference standing behind every figure — work that lands on an associate at 06:00 and on a supervisory analyst who must sign the note under FINRA Rule 2241 and Reg AC before it can be distributed. Add one more covered name, one more language, or one more client entitlement class, and that tail of work grows while the writing time stays flat.
01 · The work as it standsObserved shape and observed cost · ranges from operations of this size, not any named client
How it runs todayAn issuer reports at 07:00 London. The covering analyst and one associate pull the release, the results presentation and the call transcript, update the model in Excel, and draft a first-take note in the BlueMatrix or Word house template. The associate rebuilds the disclosure appendix by hand — ownership thresholds, banking relationships, the twelve-month price-target history, the Reg AC certification — and puts the name past the control room's restricted and watch lists. A supervisory analyst then reads for price-target basis and risk language before the note can leave the building, while the same note is read out at the 07:15 morning meeting whether or not that review has finished. Consensus is a separate job on a separate morning: someone opens Visible Alpha, IBES on Refinitiv or Bloomberg BEst, exports estimates into a spreadsheet, and works out by eye which houses moved, in which direction, and on what evidence. Buy-side desks run the mirror of this — a portfolio manager and a compliance officer checking a proposed position against the investment policy statement, concentration limits and the restricted list, usually in a shared workbook that nobody owns.
Cost todayThese are typical ranges observed where desks measure themselves; they are not a benchmark and they are not drawn from any named client. A first-take note runs 60–90 minutes of analyst writing and 2–4 hours of associate work across the model, the appendix and the citation pass. A supervisory analyst clears 8–15 notes a day at 15–30 minutes each, and that queue is the binding constraint in the first hour after an open. On a desk of roughly 40 analysts and 20 associates publishing about 250 notes a month, the arithmetic is: 250 × 3–5 hours of model, appendix and citation work = 750–1,250 hours; plus 250 × 15–30 minutes of supervisory review = 60–125 hours; plus 60–80 hours a month keeping the consensus spreadsheet current. Post-publication corrections — a wrong prior-year figure, a stale disclosure, a mislabelled chart — run at 2–6 per hundred notes on desks that count them, so 5–15 corrections a month at 1–3 hours each across analyst, editor and compliance, each one requiring a re-send to the distribution list.
What the pilot takesA pilot runs six to eight weeks against one sector team, one product type and one entitlement boundary — typically 20–30 covered names and the first-take note only. From your side it needs a research operations lead for about four hours a week; one analyst and one associate for two hours a week each as the acceptance panel; a supervisory analyst for two sessions of ninety minutes to define in writing what a passing note looks like; and read access to the research archive, the disclosure register and the restricted list. Weeks 1–2 build the ontology against your own object names, because a generic one fails on the first house-specific segment definition. Weeks 3–5 run the graph in shadow alongside the live desk, publishing nothing to clients. Weeks 6–8 produce 60–100 notes both ways and have your supervisory analyst judge them blind to which is which. No production system is connected to your distribution platform during the pilot.
What changesThe measurable is the share of published claims that resolve to a named source document, a date and a page, counted across every claim rather than sampled, and the count of unattributed numbers reaching a reader, where the target is zero and the actual number is reported weekly. Reported beside it: associate hours per published note, supervisory analyst minutes per note, and corrections per hundred notes — each of these measured for four weeks before the pilot begins, so there is a baseline rather than a recollection to compare against. Because AI-116 fails closed, a note whose provenance is incomplete does not publish late; it does not publish, and the gap shows up as a gate rejection with a reason attached, which is itself a measurement of where your evidence chain is thin.
Stays manualThe investment view stays with the analyst. The system assembles evidence, recomputes every numeric claim against the underlying series, and flags where two houses read the same datum in opposite directions; it does not set the rating, the price target or the recommendation, and we would advise against buying any system that says it does. Supervisory sign-off stays manual because the liability is personal and regulatory — a named individual certifies the note, and a graph cannot hold that certification. Order routing and transmission to clients sit outside the gate by design and will not be automated. One admitted limit: our consensus normalisation is dependable on headline lines — revenue, EBIT, EPS, dividend — and unreliable below them, because segment definitions differ house to house in ways that no critic node we have written can reconcile safely. On the reference system that reconciliation is still done by a person, and sub-headline dispersion is presented raw, unnormalised, and labelled as such on the face of the output.
02 · Instrumented workflowsCut by Cut by the unit of work that enters the graph: one market event, one body of existing house views, one mandate document, one licensed input, one inbound client question, one already-published record. Each workflow starts from a different object, so no two share an entry point or a failure mode.
CM01
Single event to published product chain
A central bank statement, an earnings release or a sanctions listing enters at stage 02 and fans out across the seven parallel evidence branches at stage 04, then into the 27-node poster pipeline at stage 05 and the long-form stage 06. On the financial-information deployment, the pattern it replaced was two analysts and one editor spending roughly three hours per macro print; what remains manual is the editor's review of a completed dossier, typically 25 to 40 minutes. Stage 07, the video route, stays switched off for most desks because voice and likeness approval takes longer than the news cycle the item serves.
Event dossier: a dated draft note, the poster set, and a claim map listing every factual sentence with its source URL, publication timestamp and retrieval hash, plus the AI-116 gate record showing pass or fail with reasons.
CM02
House view consensus and dispersion map
Forty sell-side notes on one issuer typically carry a dozen distinct claims about it, and a house frequently disagrees with itself between its credit desk and its equity desk without anyone recording that fact. The normaliser and conflict resolver at stage 04 place each numeric forecast against its source page, then the coverage reviewer marks which disagreements are evidential and which are definitional, such as two desks using different adjusted-earnings definitions. Conflicts the resolver cannot settle are printed as open conflicts rather than averaged away.
Dispersion register: a dated table of every distinct claim on the issuer, the holder of each claim, the quoted evidence with page-level citation, and a separate section of residual conflicts left unresolved.
CM03
Mandate and exposure screening
An investment management agreement is read clause by clause into the domain ontology at layer L1, and current holdings are tested against those clauses together with sanctions listings, exclusion policies and concentration limits, including look-through into pooled funds where the holdings file supports it. Every flag cites the clause number and the holding record that triggered it, so a compliance officer can dismiss a false positive in well under a minute. The screen advises; it never places, cancels or amends an order, and the graph has no execution connection at all.
Exposure screening memo per mandate: the clause-by-clause reading, the holdings tested, each flag with its clause reference and holding identifier, and a named list of clauses the system could not test and why.
CM04
Entitlement-aware derivative content
Exchange feeds, index licences and third-party estimate sets each carry redistribution terms that differ by channel, and a summary derived from a licensed input inherits those terms. Each input is tagged with its licence reference at ingestion, and a deterministic worker among the 41 refuses to emit derived content into a channel the licence does not cover, including social posts and unentitled client tiers. Where a licence is ambiguous the worker fails closed and routes to the vendor relationship owner rather than guessing.
Entitlement manifest attached to each published asset: every input, its licence reference and clause, the permitted channels, and the deterministic worker's decision record with timestamp.
CM05
Bespoke client enquiry with sources
A sales desk or an institutional client asks a question that has no published answer, such as which of the firm's covered issuers changed guidance language on freight costs in the last two quarters. The question is answered from the firm's own corpus and its entitled data only, and every sentence in the answer resolves to a document the client is permitted to see. Questions the system cannot answer within those bounds are returned as declined, with the reason stated.
Enquiry response pack: the answer, a source map naming each document consulted with its access date, and an explicit list of sub-questions declined with the reason for each.
CM06
Supervision and record reconstruction
Two years after publication, a supervisor needs to show why a specific sentence in a specific note was published and what was checked before it went out. Stage 10 writes the claim map, source map, data reproducibility record, asset manifest and call-trace archive for every item at the moment of publication, so reconstruction is retrieval rather than investigation. The archive is written once and cannot be edited by any node in the graph.
Reconstruction file for any published item: claim map, source map, asset manifest, call-trace archive, the outputs of the 12 critic nodes, and the AI-116 gate record with the identity of the human who cleared it.
CM01 · measure
On 30 consecutive events, median minutes from wire item to drafted dossier with a complete claim map, and the count of those 30 that reached the gate carrying zero unattributed factual sentences. The client picks the 30, not us.
CM02 · measure
Given 40 notes on one issuer, every target price, earnings estimate and spread forecast in the register must carry a page-level citation. The client audits 20 citations at random; any that does not resolve to the stated page counts as a failure.
CM03 · measure
Run against a back-file of breaches the compliance team already found over twelve months; the screen must reproduce at least 95 per cent of them, and every false positive must carry a clause citation. Missed breaches are counted individually, not as a rate we choose.
03 · Where this sector goesHorizon 18-36 months · our view, stated to be argued with
P01
Buyers will pay for attribution quality rather than draft quality.
Draft generation is now available from several open-weight models at low marginal cost, so it stops being a differentiator. What remains expensive is proving, per sentence, which document and timestamp a claim came from, and keeping that proof retrievable years later.
Score vendors on a failed-citation rate measured on your own corpus, not on writing samples. Ask for the number before the demo, and ask what it was on the first week of the last deployment.
P02
A major manager publicly blames a client-facing error on an unsourced model claim.
Summarisation tools are being placed in front of clients faster than publication controls are being rebuilt behind them, and a fabricated number in a client note is visible in a way an internal error is not. We call this likely rather than certain; it depends on disclosure practice as much as on failure rate.
Put a gate that fails closed in front of anything client-facing now, and keep a record of what it blocked. The blocked list is the evidence you were controlling the risk before an incident, not after.
P03
Market-data and estimate licences add explicit model-output clauses at renewal.
Vendors can see derived content circulating in channels their current terms did not contemplate, and renewal is the only moment they can reprice it. Tagging inputs at ingestion is much cheaper to build before a renewal negotiation than during one.
Inventory which licensed inputs already feed model-derived output, and by which channel, before your next renewal date. Bring that inventory to the negotiation rather than letting the vendor construct it.
P04
Disagreement maps across house views will price above summaries of the same notes.
A summary of forty notes is available to every competitor holding the same forty notes. Where the notes disagree, on what evidence, and which disagreements are definitional rather than substantive is work almost nobody does, and it survives commoditisation longer.
Before commissioning another summarisation pilot, test whether your own desks can already state where they disagree with each other on your five largest covered names. If they cannot, that is the higher-value build.
Where we disagree with the sector
The common expectation is that generative tooling cuts research headcount by 30 to 50 per cent within three years. We think net analyst headcount in sell-side and buy-side research moves very little over 18 to 36 months, and that the cost line shifts rather than shrinks: from drafting and formatting towards evidence supervision, entitlement management and record reconstruction, which are all salaried human roles today. Firms that cut analyst seats first and build the attribution layer second will re-hire within two years, at a higher cost per seat, because the people who can adjudicate a contested claim are the same people they released. We hold this view with moderate confidence. If regulators accept machine-generated attribution records without a named human signatory, the headcount argument changes and we would be wrong.
04 · Node graph, re-cut for this sectorAgainst the eleven-stage reference
Driving axisEvidence branch breadth and the cost of attributing every claim. Cost here is not model inference; it is the fan-out at stage 04, where seven evidence branches each retrieve, normalise and cite independently, and then the per-sentence attribution burden that follows. Risk is concentrated in a single event: publishing a sentence with no resolvable source, or with a source the recipient is not entitled to see. The graph is therefore shaped to make attribution cheap to produce and expensive to skip.
Stage 04 is the expanded stage in this sector: all seven evidence branches stay live, each with its own retrieval, normaliser feed and critic, because a capital markets claim usually needs a primary filing, a market-data record and a third-party view before it can be published. Those seven branches run fully concurrent; there is no ordering between them and no shared state. A synchronising barrier is mandatory at the conflict resolver, which must not start until all seven branches have landed or been recorded as failed, otherwise a slow branch is silently treated as an absent view and dispersion is understated. Stage 05, the 27-node poster pipeline, collapses to zero for buy-side internal work with no published product, and stage 07 video collapses for most desks. Stage 09, cross-product consistency, expands whenever one event feeds a note, a poster and a terminal alert, since the same number appearing with two different values across products is the most common visible failure. Stage 10 must complete before stage 11 begins, and AI-116 at stage 11 fails closed: no gate record, no publication. Six reality gates remain in place across the run; none of them is configurable by the desk.
Seven branches, one barrier
The seven evidence branches at stage 04 run concurrently and independently, but the conflict resolver waits for all of them. A branch that times out is recorded as failed rather than dropped, so the dispersion register shows a gap instead of a false consensus.
Poster pipeline is optional
The 27 nodes of stage 05 exist for published product. Buy-side desks producing internal memoranda switch the whole stage off, which removes roughly a fifth of the graph's runtime cost and none of its verification.
Attribution cost per sentence
Attribution, not drafting, dominates the compute and review budget on this sector's graphs. A note of 40 factual sentences carries 40 claim-map rows, each with a retrieval hash, and that is the line item to model when sizing a pilot.
Gate AI-116 fails closed
The terminal publication gate blocks on any unresolved critic objection or missing entitlement record. It cannot be overridden inside the graph; a human with a named role clears it, and the clearance is written to the archive.
05 · Correctness and authorityWhat decides, and what we refuse
Checkers
Numeric market claims, meaning any price, yield, spread, index level or volume, are checked by the market-data reconciler against the entitled data record for the exact timestamp stated in the claim. A numeric market claim carrying no timestamp fails on that ground alone and is never rounded into an approximation. Corporate facts, meaning issuance, guidance, dividend, filing and governance events, are checked by the filing matcher against the regulatory filing or exchange announcement, using the filing identifier and its publication time; a press summary is not accepted where the filing exists. Third-party opinions and forecasts are checked by the attribution resolver, which requires a named author, a named document and a page or paragraph reference; anything that resolves only to a house name is demoted to unattributed and blocked from client-facing output. Every input, at ingestion, is checked by the entitlement classifier, a deterministic worker that maps the input to its licence clause and the permitted channels for the intended asset. Text we write ourselves that looks forward is checked by the recommendation-language screen: no price target may appear without the model that produced it attached, and no performance figure may appear without its calculation method and period stated in the same sentence. Any passage naming a living individual in connection with a transaction is routed by the sensitive-surface checker to a human before it can reach stage 11.
Gate policy
The graph will not place, amend, cancel or route an order, and holds no execution credentials of any kind. It will not transmit client holdings, mandate documents or portfolio identifiers to any endpoint outside the client's own boundary. It will not publish anything without a passing AI-116 record, and no node can override, suppress or edit a critic objection; an objection is closed by a named human or it stays open. It will not distribute research to a recipient outside the entitlement list attached to that asset, including internal recipients. It will not produce a personalised investment recommendation for an identified individual. It will not amend or delete an archived record from stage 10. These refusals are compiled into the graph, not exposed as settings, and no configuration file we ship can turn them off.
Acceptance
Take 100 sentences at random from the first month of output. At least 98 must resolve to their cited source within two clicks, and zero of the 100 may be a numeric market claim published without both a timestamp and an entitled data source. Separately, name 20 items published during the pilot; each reconstruction file must be produced within one working day. Fail either test and the pilot fee is not payable.
Will not automate
We cannot determine whether a given piece of information is material non-public information, because that depends on facts outside any document set: who knew it, when, and under what duty. No checker we have can establish those facts, so the graph makes no MNPI judgement at all; it routes the passage to a named compliance officer and stops, which means a desk with no such officer available will see items sit unpublished. A second limit in the same family: we do not accept a number read out of a chart image as a source. Extraction from plotted axes is not reliable enough to attribute, so a claim that exists only inside a chart is marked unsourced and blocked, even where the number is almost certainly correct.
06 · Where this sector entersCross-reference to the capability catalogue in section 03 and the horizons in section 05
Fastest entry · H0
C-02 attribution replayed against 30 items you already published. Ten working days, read-only, and the deliverable is a failed-citation rate on your own output.
Pilot · H1
C-02 with C-07 on one sector team, one product type and one entitlement boundary. Six to eight weeks to an acceptance test your supervisory analyst signs.
The long part · H2 to H3
The four-product pipeline running off one evidence pack, plus the ledger that makes a two-year-old note reconstructable. Quarters, and it is the part that compounds.
IND-02 · Banking, insurance and credit risk
An annual credit review runs to forty pages, and roughly thirty-two of them are assembled rather than judged.
The judgement in a credit file is worth what the analyst is paid. The assembly around it is worth much less and takes far longer: spreading the filed accounts, testing covenants against the facility agreement, pulling the internal exposure position, reading twelve months of filings and adverse media, and writing the parts of the memo that repeat every year. The same shape appears in periodic client review, in claims intake, and in model documentation for the second line. What makes the sector workable is that the acceptance test is already written down — your credit policy, your underwriting manual and your model risk framework specify what a complete file looks like, which means the checker specification exists before we arrive. One limit stated up front: we do not adjudicate, and nothing below changes that.
The barrier
Three lines of defence means every automated output has to survive a second reading by people who were not in the room. That is a reproducibility requirement, not an accuracy requirement, and it is the one most tools fail. A first-line analyst can accept a summary that is 95 per cent right. A second-line validator cannot accept one that cannot be replayed, because the supervisory question is never "was it correct" — it is "show me how it was produced, on this date, from these inputs." Every hour saved in the first line is given back in the second unless the evidence trail is built in from the start.
01 · The work as it standsObserved shape and observed cost · figures are ranges from desks of this size, not any named client
How it runs todayA mid-market corporate portfolio gives one analyst somewhere between 25 and 60 names on an annual cycle, so a review lands roughly every two working days. The analyst downloads the latest filed accounts, spreads them into the bank's template, tests each covenant against the facility agreement, pulls the current exposure from the limit system, reads the last twelve months of announcements and press, then writes the memo. Credit committee sits weekly and returns files that are incomplete more often than files that are wrong.
Cost todayTypical observed ranges rather than a benchmark. Six to twelve hours of analyst time per annual review, of which one to two hours is the credit judgement itself. Periodic client review runs 90 minutes to four hours per file depending on entity complexity, and adjudicating a screening alert takes 15 to 40 minutes with a false-positive rate widely reported above 90 per cent. Model documentation for one internal model absorbs three to eight weeks of a validator's year.
What the pilot takesSix to eight weeks, scoped to one portfolio segment and one memo type — for example the annual review pack for unlisted mid-market borrowers in one jurisdiction. We ask for 200 to 400 completed files from the last two years, the credit policy, the memo template and the facility agreement forms. Read-only access, in your tenancy. Nothing is written to the limit system or the case management system during a pilot.
What changesMeasured by replaying files your team already completed, so the baseline and the result come from identical inputs. First, spreading accuracy: every figure in the spread must reconcile to the filed accounts within the rounding convention, and we report the exception count rather than a percentage. Second, covenant test agreement against the analyst's own conclusion, disagreements listed individually. Third, the share of the memo's standing sections produced with a document reference attached. On the last two engagements of this shape, spreading exceptions fell to single figures per hundred files and covenant disagreements clustered almost entirely on one clause type, which is a more useful finding than an average.
Stays manualThe rating, the limit, the underwriting decision, the alert disposition and the decision to file a suspicious activity report all stay with the named person who owns them. The system assembles the file, recomputes what can be recomputed, and states plainly which parts of the evidence it could not obtain. It does not recommend a rating and does not hold credentials for the limit system.
02 · Instrumented workflowsCut by the unit of work that enters the graph: one counterparty, one alert, one claim, one model
BK01
Annual review pack, assembled and referenced
The counterparty's filed accounts are spread into your own template, covenants tested against the executed facility agreement, exposure read from the limit system, and the standing sections of the memo drafted with a source reference on each figure. The analyst opens a file that is complete rather than empty.
Review pack: spread, covenant test sheet, exposure position, drafted standing sections, and an exception list of everything the graph could not source
BK02
Covenant testing and early breach detection
Covenant definitions are read clause by clause out of each facility agreement into the ontology, which is where most of the difficulty lives — the same ratio is defined differently in three agreements. Tests then run on every reporting date rather than annually, and a projected breach is raised with the arithmetic shown.
Covenant register per facility, with the clause text, the computed test, the headroom, and the date it was last confirmed against the executed document
BK03
Periodic client review, prepared not decided
Ownership structure, source of wealth documentation, sanctions and politically-exposed-person matches, and the changes since the last review are assembled into the case file with each item traced to where it came from. Screening matches are presented as list entries with dates, never as a conclusion about a person.
Prepared case file with a change log against the previous review, plus a named list of items the analyst must obtain before the file can close
BK04
Model documentation and validation evidence
Model inventory entries, development documentation, monitoring results and the validator's outstanding findings are held as objects rather than as a folder of Word files. The graph assembles the evidence a validation asks for and reports which required artefact is missing before the review starts rather than during it.
Validation evidence pack keyed to your framework's required sections, with a gap list naming each missing artefact and its owner
BK05
Claims intake completeness and fact extraction
A first notification arrives with documents attached in whatever form the claimant sent them. The graph extracts the reserve-relevant facts, tests the submission against the policy schedule and the wording in force on the loss date, and sorts the file into complete, incomplete with a named missing item, or needs an adjuster.
Triaged claim file with the policy version applied, the extracted facts and their source pages, and the specific document required to advance it
BK06
Regulatory reporting reconciliation
Exposure and capital reporting is reconciled back to the source ledgers line by line, with each difference classified as timing, mapping, or genuine. The point is not the report; it is that the difference register survives a supervisory question six months later.
Reconciliation register per reporting date: every difference with its classification, its owner, and the ledger rows on both sides
BK01 · measure
On 200 completed reviews replayed, the count of spread figures that fail to reconcile to the filed accounts, reported as an exception list rather than an accuracy rate. A single systematic mapping error is worth more to fix than forty scattered ones.
BK02 · measure
Against covenant breaches your team already identified over the last two years: how many the register catches, how early relative to the date they were actually caught, and how many false alarms it raises per hundred tests.
BK05 · measure
On 500 historical first notifications: agreement with the handler's own completeness decision, and for the disagreements, whether the graph or the handler was right on re-reading. We report both directions.
03 · Where this sector goesHorizon 18-36 months · our view, stated to be argued with
P01
Supervisors ask for the evidence trail before they ask about accuracy.
Model risk frameworks were written for statistical models with stable inputs. Applied to a language model in a first-line workflow, the first question a validator can actually answer is reproducibility, because accuracy has no agreed test. Expect replay to become the practical supervisory demand well before any accuracy standard settles.
Ask any vendor to reproduce a specific output from a specific date, from stored inputs, in front of you
P02
Adverse-media screening is where a bank gets publicly embarrassed first.
Screening tools are being asked to summarise what an article says about a person. The summary is a claim about a named individual, made without a checker, and it enters a file that can be disclosed. The failure will be an individual, not a portfolio, and it will be expensive out of proportion to the workflow's cost.
Require screening output to be a list of dated sources, never a characterisation
P03
The second line becomes the buyer, not the first line.
First-line time savings are easy to claim and hard to bank, because the review burden simply moves. The budget follows whoever can evidence a control, and that is increasingly the validation and assurance function rather than the desk.
Scope the pilot so the second line signs the acceptance test, not only the first
P04
Covenant libraries become an asset banks refuse to share.
Clause-level covenant definitions extracted from executed agreements are expensive to build and immediately reusable across a portfolio. Once a lender has one, it prices better on structure than a lender who reads each agreement fresh. We expect these to stay firmly proprietary.
Treat the extracted covenant register as your asset and fix its ownership in the contract before work begins
Where we disagree with the sector
The common expectation is that automation removes analysts from the first line within two years. We think the first measurable effect lands in the second line, where review time falls because the file is complete and replayable, and that first-line headcount barely moves for two years after that. If you are buying this to reduce analyst numbers next year, we are not the right supplier and the business case will not hold.
04 · Node graph, re-cut for this sectorAgainst the eleven-stage reference
Driving axisThe replay requirement. This sector is the one where stage 10, the backend audit and reproducibility pipeline, costs more than the reasoning stages that feed it, and where cutting it destroys the product.
Stage 04 runs narrower than in capital markets: four evidence branches rather than seven, because the sources are mostly your own systems and the filed record rather than the open web. Stage 05 collapses almost entirely, since the output is one document type rather than four. What expands is stage 10 — every figure in a spread, every covenant test and every screening match is written into the evidence ledger with the input version that produced it, so a validator two years later can rerun the file and get the same answer.
Four branches, one register
Filed accounts, internal systems of record, the executed legal documents, and public announcements. Conflicts between them are recorded as a register entry rather than resolved silently.
Clause extraction is the expensive node
Reading covenant definitions out of executed agreements is where most of the pilot's engineering time goes, and where a wrong extraction is hardest to notice.
No disposition node exists
There is no node in this graph that decides an alert outcome or a rating. It is not disabled by configuration; it was never built.
Ledger before speed
We would rather the run take four minutes with a complete ledger than forty seconds without one. That trade is not negotiable in this sector.
05 · Correctness and authorityWhat decides, and what we refuse
Checkers
Spread figures are recomputed against the filed accounts under your rounding convention. Covenant tests are recomputed from the clause text. Screening matches must resolve to a list entry with a date and a list version. Any figure without a source document is removed from the file rather than marked uncertain.
Gate policy
No rating change, no limit change, no alert disposition, no filing, no claims payment and no client communication. The graph holds no credentials for the limit system or the case management system, and read access is provisioned to named individuals on our side.
Acceptance
Replay 100 completed files. Every figure in the assembled pack must resolve to a source document, a date and a page, and the same run repeated from stored inputs must produce an identical output. Both tests are run by your second line, not by us.
Will not automate
We will not produce a conclusion about whether a named individual is a financial-crime risk. The system surfaces dated sources and list entries and stops there, because we could not build a checker for the judgement and a wrong answer about a person is not the kind of error you fix in the next release. We also decline suspicious-activity filing decisions and any workflow where the output is a characterisation of someone's intent.
06 · Where this sector entersCross-reference to the capability catalogue in section 03 and the horizons in section 05
Fastest entry · H0
C-01 reconciliation against a back-file of completed reviews. Ten working days, read-only, no integration. You get the exception list and a measured baseline before committing to anything.
Pilot · H1
C-02 attribution and C-06 triage on one portfolio segment. Six to eight weeks to an acceptance test your second line signs.
The long part · H2 to H3
Clause-level covenant extraction across the whole book, and the evidence ledger that makes a two-year-old file replayable. That is quarters of work, and it is also the part that compounds.
IND-03 · Legal, audit and tax
A diligence review reads 1,400 contracts to produce a six-page deviation report, and cannot show why the other 1,394 were cleared.
Professional firms are held to a standard that has nothing to do with whether the conclusion was right. The engagement file has to demonstrate that the work was performed, by whom, against which version of the rule, on what date. That is why review automation in this sector fails on a dimension nobody demonstrates: not the six pages that were written, but the 1,394 clearances that were never evidenced. The upside is that the checker specification already exists in a form we can read — the playbook, the audit programme, the statute. The constraint that shapes everything is privilege: an object's access class is not metadata here, it is the thing that decides whether the engagement survives.
The barrier
The work is priced by the hour and evidenced by the file. Those two facts pull against each other the moment assembly is automated, because the hours disappear from the invoice while the evidential burden stays exactly where it was. Firms that automate the drafting and leave the file unchanged end up with a faster process they cannot defend. The evidential layer — which clause governed, which version was in force on the relevant date, who reviewed it, and what was ruled out — is the part that has to be built first, and it is the part nobody demonstrates.
01 · The work as it standsObserved shape and observed cost · ranges from engagements of this size, not any named firm
How it runs todayA mid-size transaction sends 900 to 2,000 executed agreements into a data room on a Friday. Two to five associates read them against a playbook that lists the clause types in scope and the fallback positions the client will accept. Findings go into a spreadsheet, the spreadsheet becomes a report, and the report is reviewed by a partner who reads perhaps forty of the underlying documents. Audit runs the same shape with a different vocabulary: a population, a sample, a test, a working paper, and a reviewer who sees the paper rather than the population.
Cost todayTypical observed ranges, not a benchmark. Contract review runs 8 to 25 minutes per agreement for a first pass, so a 1,400-document review absorbs 250 to 500 associate hours before anyone writes anything. Control testing costs 20 to 90 minutes per sample item depending on the evidence trail. Preparing a tax position paper takes 15 to 60 hours, of which locating the authorities is a large and unpredictable fraction. Realisation on the assembly portion is where firms lose money, and most know it precisely.
What the pilot takesSix to eight weeks, scoped to one clause set or one control family in one jurisdiction. We ask for 300 to 600 documents your team has already reviewed, the playbook or audit programme, and the review output that resulted. Privilege classes are defined before any document is loaded, and documents outside the agreed class are never ingested. Work runs inside your tenancy or on your hardware; for several firms this has meant on-premises, which we will say costs you branch coverage.
What changesReplayed against reviews your firm already completed. First, agreement on the findings the team actually raised, reported as a confusion matrix rather than a single accuracy number, because a missed indemnity and a spurious one have different consequences. Second, and the measure that matters more, evidenced clearance: what share of the non-findings carry a recorded reason. Third, elapsed time from data-room open to a reviewable first-pass register. We have had a pilot where the finding agreement was strong and the clearance evidence was unusable because the playbook itself was ambiguous on two clause types, and rewriting the playbook was the actual deliverable.
Stays manualLegal advice stays with the qualified professional, full stop. The audit opinion, the materiality judgement, the filing position and every communication with a client or an authority stay with the person who signs. The system prepares, evidences and flags. It does not conclude, and it does not have a view on whether a position will succeed.
02 · Instrumented workflowsCut by the unit of work that enters the graph: one agreement, one control, one obligation, one authority
LA01
Playbook review with the cleared population evidenced
Each agreement is read against the playbook clause by clause. Deviations are raised with the clause text, the fallback position and the page. Just as importantly, every clause type that was checked and found acceptable is recorded with the reason, so the file shows what was ruled out rather than only what was found.
Deviation register plus a clearance register covering the whole population, both citing clause text and page
LA02
Obligation and deadline register from executed documents
Obligations, notice periods, renewal windows and change-of-control triggers are extracted into a register with the clause text attached. The register is dated and versioned, so a question about what the position was in March is answered from the record rather than reconstructed.
Obligation register per counterparty, each row carrying clause text, computed dates and the document version it came from
LA03
Control testing evidence and exception write-up
For each sample item the graph locates the evidence, applies the test as the audit programme states it, and drafts the working paper with the evidence references attached. Items where the evidence is missing or ambiguous are separated out and named rather than being quietly passed.
Working paper per sample item, plus an ambiguity list that the reviewer sees before signing rather than after
LA04
Tax position support against statute and prior filings
The statute, the relevant rulings and the group's own prior filings are assembled against a stated position, with each supporting sentence resolving to an authority and its version. Where authorities conflict, the conflict is presented as a conflict, not averaged into a confident paragraph.
Position support pack: authorities with dates and versions, prior-filing consistency check, and an explicit register of authorities that cut the other way
LA05
Regulatory change to procedure amendment
A published rule change is read against the firm's own procedures and client-facing documents, and the specific paragraphs that need amending are identified with the reason. One rule change usually touches more documents than anyone expects, which is the finding, not the drafting.
Impact register: every affected procedure, the paragraph, the rule reference, and a drafted amendment for a professional to accept or reject
LA06
Precedent retrieval with privilege enforced
Matter precedent is retrieved across the firm's own history with the privilege class enforced at the object level rather than by folder permissions. A user sees only what their class permits, and the graph cannot traverse a relation into a class it does not hold.
Precedent set with matter references, the reasoning as recorded, and an access log showing exactly what was traversed
LA01 · measure
On 400 previously reviewed agreements: findings agreement as a confusion matrix, and separately the share of clearances carrying a recorded reason. The second number starts near zero in every manual baseline we have measured.
LA02 · measure
Take 50 obligations the team extracted by hand. Every date must recompute from the clause text under the agreement's own convention for business days and notice periods, which is where extraction usually breaks.
LA03 · measure
Re-run last year's control testing on the same sample. Agreement with the recorded conclusion, plus the count of items the graph marked ambiguous that the original working paper had passed without comment.
03 · Where this sector goesHorizon 18-36 months · our view, stated to be argued with
P01
The assembled portion of the work gets priced as a fixed fee, and panels will require it.
Once a client can see that document review is assembly, hourly pricing on that portion becomes hard to defend at panel review. Firms that move first set the reference price and keep the judgement work at its old rate; firms that wait have the price set for them.
Work out what your assembly portion actually costs before a client tells you what it is worth
P02
Professional indemnity underwriters start asking about citation controls.
Courts in several jurisdictions have already sanctioned filings containing fabricated citations. Insurers follow incidents with questionnaires, and a question about whether generated text can reach a filing without a citation check is an easy one to write and a hard one to answer well.
Have a written answer, and a gate that fails closed, before the renewal questionnaire arrives
P03
Regulators ask about the population rather than the sample.
Sampling exists because reading the population was impossible. When it becomes possible for some evidence types, the question of why a sample was used at all becomes reasonable, and inspection practice will move faster than audit standards do.
Identify which of your control tests could run on the full population, and what that would cost
P04
Privilege becomes an architecture question rather than a policy question.
A retrieval system that can traverse across matters is a privilege incident waiting for a plaintiff. Folder permissions do not survive a system that reasons over relations, and the fix is object-level access classes enforced in the ontology, which has to be designed in rather than added.
Ask any vendor to demonstrate what happens when a query would cross a privilege boundary
Where we disagree with the sector
The common expectation is that junior headcount falls sharply. We think the binding constraint is supervision: a partner can review only so much in a week, and that number does not change because the first pass got faster. What we expect instead is more of the population reviewed at the same partner cost, and a slower change in headcount than either the optimists or the pessimists predict. If your business case rests on associate numbers falling within twelve months, it will not hold.
04 · Node graph, re-cut for this sectorAgainst the eleven-stage reference
Driving axisCitation strictness and access class. The graph runs narrow and deep, and the permission model at L1 constrains the topology rather than sitting beside it.
One branch per obligation or per control, rather than the seven parallel evidence branches of a research graph — the sources here are the executed documents and the authorities, not the open web. Stage 04 shrinks and stage 09, cross-product consistency, becomes near-trivial because there is one output. What expands is the clause checker at stage 04 and the privilege enforcement that wraps every traversal: an agent that cannot hold a privilege class cannot see the object, and the attempt is logged.
Uncited text is blocked
Not flagged, not softened. A sentence that does not resolve to a clause or an authority does not reach the draft, which makes the output shorter and occasionally unhelpfully so.
Privilege at the object, not the folder
Access class is an attribute of the object and a constraint on every relation traversal. This is the single most expensive thing to retrofit, so it goes in during week one.
Version currency is a separate checker
The right clause from the wrong version of the agreement is a distinct failure mode with its own check, because it is invisible to a citation test.
Clearance is a first-class output
The non-findings register is built by the graph, not derived afterwards. It is the part that makes the file defensible and the part clients did not ask for.
05 · Correctness and authorityWhat decides, and what we refuse
Checkers
Every sentence resolves to a clause, an authority or a working-paper reference, with the version in force on the relevant date. Deadline arithmetic is recomputed under the document's own business-day convention. Privilege traversals are checked before retrieval, not filtered afterwards.
Gate policy
No filing, no client communication, no disclosure across a privilege boundary and no engagement acceptance. The signing professional approves before anything leaves the firm, and the gate records who approved and when.
Acceptance
Take 100 sentences at random from the first month of drafted output. Every one must resolve to a cited authority at the correct version. Separately, take 100 clearances and confirm each carries a reason a reviewer accepts.
Will not automate
We do not give legal advice, form an audit opinion, or take a view on whether a filing position will succeed. Those are the acts the qualification exists for. We also decline any engagement where the firm wants generated text to reach a client or a court without a professional reading it, because the gate we would have to remove is the one thing making the system defensible.
06 · Where this sector entersCross-reference to the capability catalogue in section 03 and the horizons in section 05
Fastest entry · H0
C-02 attribution replayed on a closed matter. Ten working days, on-premises if required, and the deliverable is the clearance-evidence gap you did not previously have a number for.
Pilot · H1
C-04 rule-bound drafting on one clause set or one control family. Six to eight weeks, with the privilege model built in week one rather than at the end.
The long part · H2 to H3
The firm's precedent file as a structured asset — matter, rule version, reasoning, outcome. Quarters of work, and the thing a competitor cannot buy.
IND-04 · Semiconductor and advanced manufacturing
The excursion is found at 02:14; a defensible cause is due at the 08:00 yield meeting.
We build the layer above the tools and never inside them — a read-only ontology cut by object rather than by department, covering lots, process steps, recipe versions, chambers, traces and the closed-case archive, with retrieval that reaches back through every excursion you have ever closed rather than the last quarter. The reference graph in section 02 runs 155 nodes, 114 reasoning nodes and 41 deterministic workers across 11 stages, and we re-point it for this sector: three of the seven parallel evidence branches go to precedent retrieval, and each of the 12 critic nodes must name a measurement that would disprove the hypothesis it is reviewing, because here a confident wrong answer costs more than no answer. Our published production reference is a financial-information system, not a fab; the mechanism transfers, the domain evidence does not, which is why the pilot is scored on your own closed excursions before anything runs alongside a live desk.
The barrier
The evidence that would settle a yield excursion is already in the building: fault-detection traces (the tool's own sensor record) covering 200 to 600 channels per chamber, inline defect maps, wafer electrical test results, and eight or more years of closed 8D reports. It sits in four stores that share no common key — the fault-detection database, the yield management system, the MES and a document folder with no parameter-level index — so the engineer on shift searches two of them and stops. The real cost lands after the search: a hypothesis that is wrong but plausible passes the 08:00 meeting and buys itself a week of split lots before anyone can disprove it.
Ontology objects
LotProcess stepRecipe versionTool chamberFault-detection traceDefect signatureExcursion caseChange noticeSpecification clauseQualified part
01 · EconomicsWhat the work costs as it stands
01 · The work as it standsObserved shape and observed cost · ranges from operations of this size, not any named client
How it runs today02:14, night shift. A statistical process control chart — the running mean and spread of a measured parameter — trips on post-etch critical dimension for one product on one etch chamber. The night-shift yield engineer opens the yield management system, pulls wafer maps for the last 40 lots, then moves to the MES to run a commonality analysis across chamber, recipe version and preventive-maintenance date. That analysis names four candidate steps and no cause. She raises a WIP hold, which by 03:00 covers 40 to 120 lots, and pages the equipment engineer on call. He exports fault-detection traces for two chambers, opens them in a spreadsheet, and compares them by eye against a chamber he believes was healthy last week. Nobody searches the archive of closed 8D reports, because it is several thousand PDF and Word files with no parameter-level index, and the engineer who solved the same signature in 2019 has left the company. At 07:30 the day-shift process integration engineer rebuilds most of the picture from scratch. At 08:00 the yield meeting asks for a cause and a lot disposition, and what it gets is the best hypothesis anyone had time to build, not the best one supported by the data. The same shape repeats across the week. The change control board sits Thursday afternoon with 20 to 40 engineering change notices, each carrying an impact assessment its owner filled in by hand from memory of which products share the affected step. Contradictions between the design rule deck, the foundry datasheet and the device model surface in the final 72 hours before tape-out, when every remaining option is expensive. Under allocation pressure, the commodity manager and the qualification engineer argue about a substitute part in a meeting where nobody can state the qualification cost and lead time on one page.
Cost todayThese are typical observed ranges from operations of this shape, not any named client's figures. A 300mm fab running 30,000 to 60,000 wafer starts a month raises 8 to 20 excursions a month that reach a formal hold. The first 24 hours of each pulls in 3 to 6 engineers for 4 to 10 hours apiece: 12 to 60 engineer-hours per excursion, so 150 to 700 engineer-hours a month before a single designed experiment runs. Between 30 and 45 per cent of first-day hypotheses are later revised, and each revision costs 2 to 5 days of split-lot work while the affected WIP stays on hold — a fab in the middle of these ranges spends roughly 3 to 8 weeks of cycle time a quarter on hypotheses that did not survive. Change board arithmetic: 20 to 40 notices per cycle, times 2 to 5 affected owners, times 1 to 3 hours of manual impact assessment, gives 40 to 600 engineer-hours per cycle, and 1 to 4 impacts are typically discovered after the board has already approved the change. Specification contradictions: 2 to 6 per tape-out reach the last week, each taking 1 to 3 engineer-days to resolve, and the ones that reach silicon cost a mask revision. We do not put a price on that revision, because your mask budget is your number and any figure we invented would be decoration.
What the pilot takesSix weeks is the usual shape, eight if two process modules are in scope. What we hold the scope to: one module — etch, CMP or litho — on one product family, 24 to 36 months of closed excursion files, one full change board cycle, and one rule deck with its matching datasheet set. What it costs in your own people's time: a yield engineer at 4 to 6 hours a week to score each hypothesis against what she would have concluded herself; the fault-detection system owner for 3 to 4 hours in week one to agree the trace export and its sampling rate; a process integration engineer at about 3 hours a week; the change board secretary for one cycle; and an IT or security contact for roughly 8 hours in total to create read-only accounts and approve the extract path. That is 60 to 110 of your hours across the pilot, and we will tell you in week one if the archive is too thin to support the replay. Weeks 1 to 2 build the ontology and back-load the archive. Weeks 3 to 4 replay 20 to 30 excursions you have already closed, with the confirmed causes withheld from the system. Weeks 5 to 6 run alongside the live desk without touching it. No connection is made to a tool, to an MES write path, or to the fault-detection system's control loop; extracts land in a store you own and can delete.
What changesThree measurements, taken on your own history, reported whether or not they flatter us. First, top-three recall on the blind replay set: the share of the 20 to 30 replayed excursions in which the confirmed cause appears among the system's three highest-ranked hypotheses, each carrying the trace segment and the precedent case that supports it. We set the pass threshold with your yield manager in week one, before we see the answers, and if the result falls below it we say the pilot failed rather than re-cutting the metric. Second, time from the hold timestamp in the MES to a written hypothesis pack with its evidence attached: 6 to 14 hours in the current way of working, and the pilot target is under 90 minutes. Third, share of change impacts identified before the board sits rather than after, measured by re-running your last four completed board cycles and comparing against what was actually raised in the room. Numbers about a quarter that has not yet happened are projections and will be labelled as projections on the page they appear on.
Stays manualRead-only against MES and FDC is absolute and is not a setting anyone can change: no recipe writes, no tool commands, no lot disposition, no qualification status change. A lot is scrapped, reworked or shipped by a named process owner signing a disposition, because that decision carries product liability and a customer notification duty that no software should hold. Tape-out signoff stays with the release manager — the system produces the contradiction register and the open-item list, and a person signs. The material review board still sits. Three further things stay manual because the system is genuinely weak at them. Excursions with no precedent in the corpus, such as a new tool, a new material set, or a failure mode nobody has closed before, return no ranked hypothesis at all instead of a plausible-sounding guess; that is the terminal gate node AI-116 failing closed as designed, and it means the first months of a new process node get less help than a mature one. Sub-second electrical behaviour is outside what we ingest: we take fault-detection traces as 1 Hz summary statistics, so chamber arcing or RF instability living at 10 kHz is invisible to us and stays with the equipment engineer's own tooling. And tape-out readiness is not an object in the ontology — it is a query across specification clauses, change notices and open excursion cases — so we do not model masks, reticles or the mask shop at all, and any reticle-level question belongs elsewhere.
02 · Instrumented workflowsCut by Cut by the object the decision attaches to. One workflow per object: a lot on hold, an engineering change request, a pair of documents that disagree, a constrained supply commitment, a tape-out release, and a returned unit. No two workflows own the same object. Work that fits none of the six — capacity planning, wafer pricing, IP litigation — is outside this practice, and we say so rather than stretching a workflow to cover it.
SC01
Yield excursion triage against FDC and precedent
An FDC alarm fires — fault detection and classification, the per-tool sensor record written for every process run — and a lot goes on hold. Seven evidence branches read the trace window, inline defect maps from wafer inspection, WAT parametric results, sort and final-test bin maps, tool preventive-maintenance records, lot genealogy and five years of closed excursion reports, then write a ranked set of candidate causes with the measurement that would disconfirm each one. Triage of this kind commonly occupies three process engineers across two shifts before the first disconfirming split is even ordered; the dossier arrives in about 90 minutes and the engineers start at the ranking instead of at the data pull.
Excursion Hypothesis Dossier — a ranked list of candidate causes, each with its FDC trace window (tool, chamber, recipe, time range), the matching precedent excursion IDs and how those were resolved, and the single split or measurement that would rule it out. Issued to the process engineer holding the lot.
SC02
Engineering change impact before the change board
An engineering change request — a resist change, a new bond wire alloy, a mask revision, a second-source lead frame — usually arrives with an affected-parts list written from memory. The workflow expands it: every part number sharing the affected step, the requalification each one triggers, the customer notification obligation with its notice period, the WIP and finished inventory exposed, and the controlled documents that must be revised. A part number discovered after the board has sat costs another board cycle, typically one to two weeks, and can restart a 90-day customer notice from zero.
Change Impact Register — one row per affected part number, carrying the requalification triggered, the customer notification obligation and its notice period, WIP and inventory exposure, and the controlled documents needing revision. Attached to the change board packet before the board sits.
SC03
Rule deck and datasheet contradiction register
A rule deck is the machine-readable set of design rules shipped with a foundry PDK; the design manual, the device datasheet and the customer specification frequently state different numbers for the same parameter, and the disagreement is usually found by a person late. The workflow reads all four document classes at their stated revisions and records each disagreement as a numbered entry with both citations and the numeric gap. It does not decide which document wins, because that is a foundry or design-authority ruling and we do not automate it.
Contradiction Register — numbered entries, each naming the two documents with their revisions and the clause or page, the parameter in dispute, the numeric difference, a severity, and the named owner the entry is routed to.
SC04
Allocation and substitution under qualification cost
When a substrate, a lead frame or a specific analogue die goes short, the decision is which customer commitments to hold and whether a qualified alternate exists at all. Qualification is what makes it expensive: an automotive requalification under AEC-Q100 runs months and consumes engineering hours already committed elsewhere, so an alternate that wins on unit price can lose once its qualification cost and schedule are counted. The workflow prices both paths and shows the approved-vendor-list gap rather than leaving it implied.
Allocation and Substitution Brief — one page per constrained part: committed demand by customer and date, on-hand and in-transit quantities, qualified alternates with qualification status and date, the qualification cost and lead time for each unqualified alternate, and the AVL entry that would need amending. Goes to the materials review board.
SC05
Tape-out readiness pack and exception list
Tape-out sign-off evidence sits across DRC and LVS run logs, timing sign-off corner reports, IP licence and version records, the foundry deck revision actually in use, and the assembly and test plan at the OSAT. The workflow assembles that evidence into one indexed pack and lists every open exception with its owner and waiver reference; it does not judge whether a waiver is acceptable, because that judgement belongs to a named engineer. A respin costs a mask set and a substantial part of a schedule quarter — your figures, not ours — which is why the pack is built to be argued with rather than trusted.
Tape-out Readiness Pack — the sign-off evidence index plus a numbered exception list, each exception carrying an owner, a waiver reference where one exists, and the run-log line it was drawn from.
SC06
Field return 8D containment and evidence pack
A returned device arrives with the customer's 8D clock running, commonly 24 hours to containment and 30 days to root cause. The in-fab excursion route does not apply: the part has shipped, the evidence is the unit in hand plus its genealogy through wafer sort, final test and the OSAT assembly lot, and the immediate question is which other shipped lots share that path. The pack drafts sections D1 to D5 and stops there, because the failure mechanism is stated by the failure analysis lab after decapsulation or cross-section, not by an agent.
8D Evidence Pack (sections D1 to D5, draft) — the unit's genealogy from wafer to assembly lot to test insertion, matched prior returns with their dispositions, and a proposed containment scope naming the lots it would cover.
SC01 · measure
Replay 40 closed excursions from the past 24 months with the resolutions withheld. The confirmed cause must appear in the top three ranked hypotheses in at least 30 of them, and every trace window cited must open in your own FDC viewer at the stated tool, chamber and time range.
SC02 · measure
Run it against the last 20 change requests your board has closed. The register must contain every affected part number your engineers listed; we publish both the misses and the additions. More than one missed part number across the 20 is a fail.
SC03 · measure
Seed 15 known contradictions into a copy of your rule deck, design manual and customer specifications. At least 13 must appear in the register with both citations correct, and false entries must stay below 5 per 100 pages reviewed.
03 · Where this sector goesHorizon 18-36 months · our view, stated to be argued with
P01
Excursion triage records become audited artefacts, not engineers' private notebooks
Automotive and medical customers already audit change control, and the same surveillance audits under IATF 16949 increasingly ask how a cause was established rather than only what it was. Evidence held in a chat log or a personal spreadsheet does not survive that question.
Specify an export schema and a retention rule for triage evidence in the next tool purchase, and decline any tool that keeps its reasoning where an auditor cannot read it.
P02
Advanced packaging overtakes front-end defects as the largest yield loss at leading OSAT sites
Stacked and chiplet assemblies multiply yields. Bonding eight known-good dies means assembly escapes compound faster than any single front-end step, and the loss is measured against the value of the whole stack rather than one die.
Fund a precedent corpus that indexes wafer test and assembly history together. A fab-only excursion archive cannot answer a stacking question, and building the index later costs more than building it now.
Restrictive design rules grow faster than sign-off headcount, and foundry deck drops land mid-project rather than between projects. The reconciliation is currently done by senior people reading PDFs.
Put machine-readable deck deltas with revision dates into the foundry agreement before the next node starts, not after the first contradiction costs a respin.
P04
Read-only replicas, not live MES connections, become the default for external analytics agents
Equipment cybersecurity requirements in the SEMI E187 family, plus insurers' network-segmentation questions, make a live write path into MES hard to defend at audit. Segmentation is cheaper to prove than a permissions model.
Budget a read replica and state its refresh latency in the pilot scope, and treat any vendor that requires write access to MES as a procurement decision rather than an engineering one.
Where we disagree with the sector
The common view is that excursion response waits on a unified data platform: finish the lake, then automate the analysis. We think that is backwards here. The evidence that decides most excursions is not missing, it is unindexed — the last few hundred closed excursion reports usually sit in a document store, and the FDC extract needed to test one hypothesis is a bounded query, not a warehouse. A pilot run on read-only extracts takes weeks and surfaces exactly which data is genuinely absent, which is a cheaper way to specify a platform than specifying it in advance. We may be wrong at sites where excursion reports were never written up in a consistent form; there the corpus has to be built before retrieval is worth anything, and we say so at diagnosis rather than after the invoice.
04 · Node graph, re-cut for this sectorAgainst the eleven-stage reference
Driving axisCost here is dominated by the price of a wrong hypothesis, and the graph shape is dominated by how deep into precedent you must retrieve to avoid one. A wrong containment holds lots that were never affected and delays shipments; a wrong recipe edit can cost a chamber requalification. Retrieval depth is therefore budgeted against value at risk: a single-lot question stops at 12 months of precedent, while an excursion holding a week of WIP opens the five-year corpus and pays the retrieval cost.
The reference topology — 155 nodes, 114 reasoning and 41 deterministic workers, across 11 stages, with 7 parallel evidence branches, 12 critic nodes and 6 reality gates — was built for a financial-information client and is re-cut here rather than copied. Stage 04, the single-event research route, carries most of the sector weight: its 7 evidence branches are remapped to FDC traces, inline defect inspection, WAT parametrics, sort and final-test bin maps, tool maintenance records, lot and assembly genealogy, and the closed-excursion precedent corpus. Stage 09, cross-product consistency review, expands to carry the contradiction work of SC03, since a rule deck and a datasheet disagreeing is the same class of problem as two published products disagreeing. Stage 05, the 27-node poster pipeline, and stage 07, video, have no use in a fab and collapse to a small rendering step that produces the registers and dossiers; stage 06, long-form, survives only as the 8D narrative draft in SC06. Stage 02, realtime signal, expands to take FDC alarm intake with a rate limit, because an unfiltered alarm stream will start more routes than any engineer can read, and a queue of unread dossiers is worse than none. Stage 10, audit and reproducibility, is kept in full — claim map, source map, data reproducibility record, asset manifest and call-trace archive — because customer quality audits ask for that material by name. Concurrency is safe across the 7 branches of stage 04 and across part numbers in the SC02 expansion. A synchronising barrier is mandatory in two places: after the 7 branches and before hypothesis ranking, because a ranking computed on partial evidence inherits the bias of whichever branch returned first; and before stage 11, where the terminal gate node AI-116 fails closed, so a pack with one unresolved citation is not released at all.
Read-only against MES and FDC
The graph connects to a read replica or a scheduled extract, never to a live write path, and no node in the build holds a recipe-edit or lot-disposition credential. Refresh latency is stated in the pilot scope so an engineer knows how stale the dossier's evidence is.
Precedent depth is budgeted
Retrieval over a five-year excursion corpus is the most expensive step in the route, so depth is set by the value at risk on the lot rather than by default. A single-lot question stops at 12 months; a week of held WIP opens the full corpus.
Barrier before ranking
Hypothesis ranking in stage 04 does not begin until all 7 branches have returned or have explicitly reported their source unavailable. A branch that is silent is recorded as a gap in the dossier, not treated as an absence of evidence.
Maths in deterministic nodes
Commonality counts, Cpk against WAT limits, and bin-map overlays run in the 41 deterministic workers, where the same input gives the same output every time. Reasoning nodes may cite those results but may not compute them.
05 · Correctness and authorityWhat decides, and what we refuse
Checkers
Six checkers are specific to this sector, each bound to one class of claim. The trace-resolution checker runs on any claim naming a tool, chamber or process step: the claim must resolve to a specific FDC trace window with tool ID, chamber, recipe and timestamp range, and a claim that will not resolve is removed from the dossier rather than softened into a hedge. The commonality checker runs on every count-based claim — "nine of eleven affected lots passed through etch chamber C" — and recomputes the count from lot genealogy inside a deterministic worker; a mismatch fails the claim outright. The sample-strength critic runs on every correlation claim and, below the lot count agreed with your process team, relabels the claim from finding to hypothesis in the visible text. The process-order critic runs on every proposed physical mechanism and rejects any mechanism requiring a defect in a layer deposited after the step under suspicion. The document-revision checker runs on every design-rule, datasheet or specification claim: each citation must carry the document revision and the clause or page, and a citation against a superseded deck revision is flagged with the current revision printed beside it. The qualification-status checker runs on every substitution claim, and treats lapsed, in-progress and unqualified as three distinct states that can never render as available.
Gate policy
The agent is read-only against MES, FDC, the recipe management system, the PDM or PLM vault and the ATE test-programme repository. This is a property of the build, not a configuration flag: there is no write path to switch on. It does not place or release a lot hold, does not disposition material as scrap, rework or use-as-is, does not edit a recipe or a tool parameter, does not approve an engineering change, does not issue a process change notification to a customer, does not release a tape-out, does not raise a purchase order or commit a substitution, and does not close an 8D. Each of those actions carries the signature of a named person and stays that way. AI-116, the terminal gate node, fails closed: if any checker cannot resolve a citation, nothing is released, including the parts of the pack that did pass.
Acceptance
Replay 40 closed excursions from the last 24 months with the resolutions withheld from us. We fail the pilot if the confirmed cause is absent from the top three ranked hypotheses in more than 10 of the 40; if the median time from alarm to first dossier exceeds 90 minutes on your own hardware; or if any citation in any dossier does not open at the stated tool, chamber and time range in your FDC viewer. A single fabricated citation ends the pilot and you pay nothing.
Will not automate
Raw FDC traces are commonly retained for 30 to 90 days and then reduced to summary statistics, and the signature of a slow drift usually lives in the raw trace rather than the summary. For an excursion whose incubation predates your retention window we cannot verify a tool-level cause, so we do not assert one: the dossier names the missing window and stops. The same refusal covers two neighbouring cases. We do not name a physical failure mechanism the failure analysis lab has not confirmed by decapsulation, cross-section or elemental analysis. And we do not rank precedent drawn from a different fab or a different node alongside local precedent, because we cannot verify that the process context is comparable; those cases are listed separately, unranked and labelled as such.
06 · Where this sector entersCross-reference to the capability catalogue in section 03 and the horizons in section 05
Fastest entry · H0
C-03 cause ranking scored on excursions your engineers already solved. Ten working days, read-only against the archive, and it returns a top-three recall number rather than a demonstration.
Pilot · H1
C-03 on one process module and one product family. Six weeks, eight if a second module is in scope, measured on your own history.
The long part · H2 to H3
Coverage across modules and the fault-detection trace corpus behind it. Multiple quarters, and it stays read-only against MES and FDC throughout.
IND-05 · Supply chain, freight and trade
A container clears in hours, but the five documents that govern it reconcile in days.
We build the reconciliation layer that sits under the operations desk: a domain ontology (L1) in which a purchase order line, a shipment leg and a document instance are separate objects with explicit links, cut by what kind of record a thing is and never by which department holds it; a verification lattice (L4) that tests each field against the clause that actually governs it — the Incoterms 2020 delivery point, the letter of credit terms under UCP 600, the carrier's contracted rate sheet; and a reality gate (L5) that stops a document pack being presented while a material discrepancy is open. The sector suits a verified execution stack because the governing rules are written down and citable, which means a machine result can be checked against a source rather than trusted on its say-so. We shape the node graph on one axis — document volume against the cost of reconciling a single field — so high-volume, low-consequence fields such as container numbers, charge codes and HS headings go to deterministic workers, while low-volume, high-consequence fields such as the origin declaration or the delivery point under the contract go to reasoning nodes with a critic node behind them.
The barrier
The same consignment carries a different identifier in every system it touches: a purchase order line in the ERP, a booking number at the forwarder, a house bill of lading, a master bill above it, a container number, and an entry number at customs. Nothing joins those identifiers except a person who knows the file. Worse, several fields that look like they should match are supposed to differ — gross weight against net weight, commercial invoice value against declared customs value, ship-to party against consignee — so a naive field-by-field comparison produces more false alarms than findings, and the desk learns to ignore it inside a fortnight.
01 · The work as it standsObserved shape and observed cost · ranges from operations of this size, not any named client
How it runs todayOvernight the carrier pushes an ETA update and the arrival notice lands in a shared mailbox. Between 07:00 and 09:00 the import coordinator works the exception list by hand: opening the file in the transport management system (CargoWise, SAP TM, Oracle OTM, or a broker's own entry platform), pulling the commercial invoice and packing list out of email attachments, comparing them line by line against the house bill of lading, and keying value and classification into the declaration before the filing cut-off. At the customs broker, an entry writer separately checks the certificate of origin against the preference rule whenever the importer claims duty relief, and files the entry summary inside the statutory window. The delivery order is chased by telephone once freight and local charges are settled, which is where most demurrage is actually lost. Freight invoices go to a finance clerk who checks a sample against the rate sheet in a spreadsheet, because checking every invoice at 8 to 15 charge lines each is not affordable. Tariff classification under the Harmonised System, currently the 2022 edition, is reviewed once or twice a year on a sample, so an inconsistent code can run for months across sites. And n-tier exposure — who supplies your supplier — is assembled only after something has already happened, by emailing tier-one suppliers and waiting.
Cost todayThese are typical observed ranges across desks of this size, not figures from any named client. An import desk or forwarder branch running 1,500 to 2,500 shipment files a month handles 8 to 14 documents and 60 to 120 checkable fields per file. Manual document checking runs 25 to 45 minutes per file, so 2,000 files is roughly 830 to 1,500 hours a month, or 5 to 9 full-time equivalents at 165 productive hours each. Discrepancy escape rates of 2 to 5 per cent of files are common; a letter of credit presented with a discrepancy attracts a bank fee in the USD 75 to 150 range and adds 3 to 9 days, against a 21-day presentation period and the five banking days the issuing bank has to examine under UCP 600 Article 14. Detention and demurrage after 4 to 7 days free time runs USD 100 to 300 per container per day, 3 to 8 per cent of containers incur it, and re-audits typically find 20 to 40 per cent of those charges either avoidable or wrongly rated. Freight invoice audit finds 1 to 4 per cent of charge lines off the contracted rate, worth 0.5 to 2 per cent of freight spend, recoverable only inside the carrier's 30 to 60 day claim window — which is where most of it is quietly lost. Classification sampling commonly finds 3 to 8 per cent of part numbers carrying more than one HS code across sites.
What the pilot takesFour to eight weeks, scoped to one trade lane, one customs jurisdiction and one carrier contract. We ask for 300 to 800 historical shipment files that already have a complete document set and a known outcome, read-only extracts from the transport management system and the customs filing archive, the contracted rate sheet with its charge code dictionary, and the governing terms: the sales contract Incoterm, the letter of credit where one exists, and the bill of lading conditions. Your people's time is the real cost: a documentation lead at 4 to 6 hours a week adjudicating what the checker flags, a trade compliance manager at 2 to 3 hours a week on classification questions, an IT contact for 8 to 12 hours in total to produce the extracts, a finance contact for 3 to 4 hours on rates and charge codes, plus a two-hour ontology session in week one and a one-hour gate review at the end of week four. Budget roughly 60 to 90 hours of your staff time across the whole pilot. During a pilot we do not write into a carrier booking system and we do not file a live declaration: the system reads, reconciles and proposes, and a named person acts.
What changesEverything is measured by replaying the same 300 to 800 files, so the baseline and the result come from identical inputs. Four measurables are reported: the share of documentary discrepancies identified before presentation rather than after it; median hours from arrival notice to a closed exception with the decision and its cost impact recorded; the count of charge lines disputed inside the carrier's claim window rather than after it expires; and the share of flagged classification inconsistencies that the compliance manager confirms as real. We print the false-positive rate beside each one. In early runs the checker raises roughly one flag that turns out not to be a discrepancy for every three to six that are, and that ratio improves only as your people correct the clause library, which usually needs a second cycle to settle. Where a figure is a forward estimate rather than a replayed measurement, the report labels it as a projection and says what evidence it rests on.
Stays manualSome work stays with a person, and we recommend keeping it there. The classification decision on genuinely novel goods, or where a binding ruling is being sought, because the signature and the legal liability sit with a licensed broker or the importer of record. The commercial call on a re-plan — air-freighting a part to hold a production line against accepting the delay — because the system can cost both options to the dollar and the service day, but cannot know what the customer relationship is worth. The final sanctions and denied-party determination, where a name match starts an investigation rather than settling it. And the negotiation with the carrier once the dispute pack is assembled. Two admitted limits. Beyond tier two, exposure reconstruction is inference drawn from customs and manifest records with a stated confidence, not established fact; it is labelled that way on screen, and we advise against making a single-source sourcing decision on it. And Enactive Reality runs here at persistence level L2 — the system holds state across a file and acts on your documents — while the L3 step of writing directly into a carrier or customs system is gated and only partly built. The terminal gate fails closed, in the same pattern as node AI-116 in our reference system: if a check cannot be completed, nothing is released.
02 · Instrumented workflowsCut by Cut by the decision the evidence has to support, running from before the order is placed to after the money is settled: who we buy through, whether the document set holds, what duty is owed, what to do when the plan breaks, which invoice lines to pay, and what we can prove when someone challenges us years later. One workflow owns each decision and none re-does another's work. Two things this set deliberately does not cover: freight procurement and rate negotiation, and anything physical — inspection, weighing, damage survey. We can reconcile what the documents say about the box. We cannot tell you what is in it.
SL01
N-tier exposure register from orders and customs data
We rebuild the path behind each part number from your purchase order lines, supplier declarations, and the bill-of-lading and customs records available from the trade-data providers you already licence. A sourcing team doing this by supplier survey typically covers 200 parts in three weeks at a response rate near half; a run across 4,000 part numbers takes about one week of machine time plus one analyst week, and reaches a named tier-2 site on roughly 60 per cent of lines. Tier-3 coverage stays under 25 per cent, because for most tier-3 relationships no obtainable document exists, and every row states which of the two cases it is.
Exposure register: a spreadsheet plus a signed PDF summary, one row per part number and tier, each row naming the supplier site and either the evidencing document or the tag 'inferred, unevidenced'.
SL02
Document set reconciliation against governing clauses
We compare the commercial invoice, packing list, house and master bill of lading, certificate of origin and delivery order field by field — consignee, notify party, goods description, tariff code, gross weight, quantity, port pair, incoterm — and then against the governing sale contract or letter of credit clause. A documentation clerk working at ordinary speed checks perhaps 25 of about 140 comparable fields in twelve minutes; we compare the full set and return only the fields that disagree. Scanned and handwritten pages fall out of the automatic path to a person, which on our pilot sets runs at 8 to 12 per cent of pages.
Discrepancy note per shipment, naming the field, the two documents that disagree, the governing clause by section number, and the correction requested of a named party.
SL03
Tariff classification consistency review
We read the classification codes declared across two years of customs entries, commercial invoices and your product master, and list every product carried under more than one heading, with the duty difference each variation implies at the rate in force on the entry date. We do not issue a classification opinion. The file goes to your licensed customs broker or counsel, who decides and signs; on a 30,000-line entry history our run takes about two days and the broker review that follows is the larger cost, usually 20 to 40 hours.
Classification variance file: each product, the headings used, the entries and dates each heading was used on, and the duty difference in the entry currency, addressed to your broker.
SL04
Exception desk with costed re-plan options
When a carrier status message, customs hold or terminal event breaks the plan, the desk assembles what happened, cites the message it came from, and returns two to four options, each carrying a cost difference in your buying currency, a revised arrival date, and the date free time expires under that option. A named planner chooses; no agent books, rolls, diverts or cancels anything. Options are only as good as the rate and free-time terms we hold, so a lane with no contracted rate on file returns a spot estimate that is labelled an estimate.
Exception file per event, carrying the source message, two to four costed options with revised arrival date and free-time expiry for each, and a signature block for the planner.
SL05
Freight invoice audit and charge dispute packs
We match each invoice line to the rate agreement, the surcharge tariff on file and the movement record, then sort lines into three states: outside the agreement, inside the agreement but without evidence that the chargeable event happened, and correct. Detention and demurrage lines with no gate-out or ingate timestamp fall into the middle state, which is where most of the recoverable money sits. The pack is raised only after your named approver signs it, and we do not send it to the carrier.
Dispute pack per invoice: contested lines, the contracted rate or tariff clause cited, the movement evidence behind each contested line, and the amount claimed, addressed to the carrier's billing contact.
SL06
Retained evidence file for audits and claims
For each customs entry we keep the declaration as lodged, the supporting documents as presented, the classification and origin reasoning in the version used on that date, and the name and date of the person who approved it. When a post-clearance audit, origin verification or cargo claim arrives two years later, the question asked is what you knew and when; a retained file answers it in minutes instead of a fortnight of mailbox archaeology. We do not set your retention policy or period — that is your counsel's decision, and we store to the period you specify.
Retention file per customs entry, indexed by entry number, container number and part number, with a document manifest and the approver's name and date on the cover sheet.
SL01 · measure
Draw 50 evidenced rows at random. We fail if more than two cannot be traced back to the cited document in under ten minutes each.
SL02 · measure
Count fields actually compared per document set and the share of pages routed to a person. We commit to 140 or more fields compared per set and under 15 per cent of pages routed out.
SL03 · measure
Compare our count of multi-heading products against your broker's own review of the same period. A difference above 2 per cent of products is our error to explain.
03 · Where this sector goesHorizon 18-36 months · our view, stated to be argued with
P01
Customs authorities will start asking for the reasoning behind a classification code
Entry data is increasingly matched across importers and periods, so an authority can see that the same product moved between headings before it asks you why. Post-clearance audit then turns on what you decided and when, not on the code alone. This is a projection; we base it on the direction of EU customs reform and on US post-entry practice, not on a published rule that already requires it.
Store the reasoning and its version alongside the code from the next entry onwards. Retrofitting reasoning after an audit letter arrives is the expensive path, and reconstructed reasoning is worth little.
P02
Carbon and deforestation declarations will fail like invoices, on fields and party names
Reporting schemes such as CBAM and the EU deforestation rules ask for supplier-level data at declaration time, in the same shape as a customs field: a party, a site, a quantity, a date. The failure modes are missing fields and names that do not match across documents, which is what already breaks letters of credit.
Add emissions and origin attributes as fields in the same document reconciliation you already run, rather than buying a separate sustainability system that has no view of the shipment record.
P03
Detention and demurrage disputes will be decided on timestamps, not correspondence
Billing rules in the US now specify invoice content and short dispute windows, and carriers and terminals emit machine-readable gate events. The party holding its own timestamped record wins the argument; the party quoting an email thread loses it.
Capture terminal and gate timestamps into your own store now, with sender and receipt time, instead of relying on the carrier's invoice to describe the event it is charging for.
P04
Forwarder margin shifts from the freight rate to the document desk, priced per file
Rate margin keeps compressing while documentation, customs and dispute handling stay labour-heavy. Once reconciliation is partly automated, forwarders will price the document work separately because that is where the remaining margin is. This is a projection about pricing behaviour and the evidence for it is thin today.
Ask for document handling to be priced as a separate line in your next forwarder tender, and ask which parts of it are machine-checked and at what field coverage.
Where we disagree with the sector
The common view is that full n-tier visibility is a purchasing problem: buy enough trade data, run enough supplier surveys, and the map completes down to tier 3 and beyond. We think that is wrong, and our own numbers argue against us being able to sell it. Tier-2 coverage is genuinely reachable because a customs record or a supplier declaration usually names the site. Tier-3 stalls, in our experience below 25 to 30 per cent of lines, because the document that would prove the link was never created, is commercially confidential, or moves inside a country with no accessible filing. The money spent chasing the last tiers buys a prettier diagram with a thinning evidence base underneath it. We would put the same budget into substitution readiness — for each critical part, a qualified alternate source, a stated requalification time, and a tested re-plan — which pays off whether or not the map was complete. This is contestable: a buyer with concentrated, well-documented supply in a few jurisdictions may well get further down the tiers than we do.
04 · Node graph, re-cut for this sectorAgainst the eleven-stage reference
Driving axisDocument volume against per-field reconciliation cost. The work is not hard per comparison; it is enormous in count. One shipment carries a set of six to nine documents holding roughly 140 comparable fields, and a mid-sized importer clears thousands of shipments a month, so unit cost is set almost entirely by what a single field comparison costs and how many of them are routed to a reasoning node rather than a deterministic worker. Risk runs the other way and concentrates in a handful of fields — tariff code, origin, consignee and notify party, incoterm, gross weight, goods description — because those decide duty, release and who carries the loss. The graph therefore has to be cheap in the wide part and expensive only in the narrow one.
The reference topology — 155 nodes, 114 reasoning and 41 deterministic, over 11 stages — is re-cut as follows. Stage 01 expands into per-source ingest for carrier status messages, broker entry files, purchase order extracts from the ERP, rate agreements and scanned PDFs, and ends in a mandatory synchronising barrier: booking, container, house and master bill of lading and entry numbers must resolve to one shipment identity before anything downstream runs, because a wrong join here corrupts every later comparison. Stage 02 stays live and continuous, feeding the exception desk. Stage 03 collapses to a per-lane exception digest. The seven parallel evidence branches at stage 04 are re-pointed at commercial documents, transport documents, customs and entry records, contracts and rate agreements, party and product master data, carrier and terminal event feeds, and public registry and restricted-party lists; concurrency is safe there because the branches read separate stores and write nothing. The 27-node poster pipeline at stage 05 collapses to about four layout nodes for the exception file and dispute pack, stage 06 narrows to the classification variance memo and register narrative, and stage 07 drops entirely; the freed reasoning capacity moves to conflict resolution at stage 04 and to stage 09. Stage 09 expands and becomes the centre of the sector build: the same tariff code, origin and party name must agree across every entry, invoice, register and pack we produce, and that is a global property, so stage 09 is a hard barrier with no concurrency. Stage 10 keys the claim map to document and field rather than to paragraph. Stage 11 is unchanged: terminal gate AI-116 fails closed. The cost arithmetic sets the split. Routing 10 per cent of 140 fields per shipment to reasoning nodes costs roughly an order of magnitude more per shipment than deterministic comparison; our design target is under 5 per cent, and the pilot reports the actual figure rather than the target.
Identity barrier at stage 01
Booking, container, house and master bill of lading and customs entry numbers must resolve to one shipment identity before any evidence branch starts. A wrong join at this point produces confident, wrong reconciliation everywhere downstream, so the barrier is mandatory rather than tuned.
Field checks stay on workers
The roughly 140 field comparisons per document set run on deterministic workers with fixed rules. Reasoning nodes are reserved for the fields where a contract or letter of credit clause has to be read and applied, which is where the cost per field is 10 to 20 times higher.
Stage 09 is the hard barrier
Cross-product consistency cannot run concurrently, because agreement of tariff code, origin and party name is a property of the whole output set rather than of any one file. Everything queues here, and this is the stage that sets end-to-end latency on a large batch.
Stage 07 out, stage 05 cut
The video stage is removed and the 27-node poster pipeline shrinks to about four layout nodes. The capacity released goes to conflict resolution inside stage 04, where documents genuinely disagree and a decision has to be justified.
05 · Correctness and authorityWhat decides, and what we refuse
Checkers
Five sector checkers run, each bound to one class of claim. A charge arithmetic checker recomputes duty, import VAT, surcharges and detention or demurrage day counts from the tariff rate, the contracted rate table and the free-time terms in force on the movement date; it runs on every monetary claim, and any line it cannot recompute is marked unevidenced rather than passed as correct. A cross-document field checker compares consignee, notify party, goods description, tariff code, gross weight, quantity, port pair and incoterm across the whole set; it runs on every reconciliation claim and reports the two documents that disagree, never a single verdict. A reference-data checker tests tariff headings against the schedule edition in force on the entry date and party names against the restricted and denied-party lists as published on that date; it runs on classification and counterparty claims, and it establishes only that a heading exists and a name does or does not appear, not that the heading is the right one. A physical-event checker requires every claim about a real-world event — gate-out, discharge, delivery, hold release — to resolve to a timestamped source message with a named sender and a receipt time; it runs on exception and dispute claims, and an event known only from a phone call is recorded as unevidenced. A clause-citation checker requires any claim that a document breaches a governing term to cite the contract or letter of credit section number and quote the wording; it runs on every discrepancy note, and a note that cannot cite is held back rather than softened.
Gate policy
These refusals are fixed in the build and are not configurable by the client. No agent lodges, amends or withdraws a customs declaration, and no agent signs as declarant, importer of record or exporter. No agent issues, amends, surrenders or releases a bill of lading, a delivery order or a telex release. No agent books, rolls, diverts or cancels cargo, accepts a rate, or signs a booking note. No agent pays, credits or writes off a carrier or forwarder invoice, and no agent files a claim with an insurer. No agent decides a classification or an origin determination of record; that decision is signed by your licensed broker or counsel. No agent clears a restricted-party or sanctions match, whatever the match score says. No agent transmits anything to a customs authority, a carrier, a terminal or a counterparty: output stops at a named human, and terminal gate AI-116 fails closed, so a file with a field the checkers could not resolve does not leave the system at all.
Acceptance
Take 200 shipments your own team has already reconciled in the last two quarters and hand us the same source documents they had. On the discrepancy classes agreed in writing before the run, we must find at least 95 per cent of the discrepancies your team found, raise no more than one false discrepancy per 20 shipments, and cite for every finding the two documents and the named field. Miss any of those three thresholds and the pilot has failed; the fee is fixed and stated in advance, and there is no further charge.
Will not automate
We cannot verify that a certificate of origin is genuine, and we refuse to automate any judgement that depends on it. We can check that its fields agree with the invoice and the bill of lading, that the stated rule of origin exists, and that the issuing body is named; we cannot confirm with the issuing chamber that the certificate was issued, because for most jurisdictions there is no machine-checkable register to ask. Every origin row we produce is therefore marked 'document as presented, not verified', and an origin determination is signed by your broker or counsel, not by us. The same applies to the physical shipment: nothing in the document set proves what is inside the container, and a packing list that reconciles perfectly against a short-shipped or substituted load will pass every checker we run.
06 · Where this sector entersCross-reference to the capability catalogue in section 03 and the horizons in section 05
Fastest entry · H0
C-01 reconciliation replayed on 300 to 800 historical consignment files. Ten working days, read-only, and it reports where the five documents actually diverge.
Pilot · H1
C-06 exception triage on one trade lane, one customs jurisdiction and one carrier contract. Four to eight weeks against files you have already cleared.
The long part · H2 to H3
Identifier resolution across the ERP, the transport system and the customs record, then multi-lane coverage. Quarters, and the identifier work is most of it.
IND-06 · Process plants, utilities and grid
The procedure is on revision 12; the answer given on the floor came from revision 9.
We build five things here: shift assistance that answers only from the controlling revision, deviation drafting that begins while the shift is still on site, maintenance planning constrained by spare stock, permits, technician qualifications and outage windows, drawing-to-model release verification, and handover records that carry evidence rather than recollection. The graph is shaped around one axis — whether the revision in hand is current — so a deterministic worker matches document number, revision and effective date against the document management system of record before any reasoning node is allowed to quote a step; the structure is the one we run in production for a financial-information client, 155 nodes in total, 114 reasoning nodes and 41 deterministic workers across 11 stages, with a terminal gate, AI-116, that fails closed rather than releasing an unverified result. Connections to PLC, SCADA and the historian are read-only and stay read-only: the system can quote a tag value and cannot write a setpoint.
The barrier
A plant does not lack documents. It lacks certainty about which document governs right now: the controlling SOP sits in Veeva Vault or Documentum at revision 12 with an effective date of 14 March, while a laminated copy of revision 9 is taped inside the panel door and a PDF of revision 10 sits in a shared folder that two of the three shifts still search first. The cost is not the wrong answer alone. A deviation written against a superseded step gets re-investigated, and the QA reviewer usually finds the mismatch four to six weeks later, after the batch has moved and the operator who was there has rotated off.
Ontology objects
Procedure revisionDrawing revisionEquipment recordDeviation recordWork orderPermit to workSpare reservationQualification recordHandover entryHistorian series
01 · EconomicsWhat the work costs as it stands
01 · The work as it standsObserved shape and observed cost · ranges from operations of this size, not any named client
How it runs today05:50, ten minutes before the day shift takes the board. The night supervisor types the handover into a spreadsheet on a shared drive, or a free-text field in the DCS shift log, and the incoming supervisor reads it in about nine minutes while the units are being walked. Earlier, near 02:00, an area operator met an alarm he had not seen before, phoned the on-call process engineer, and worked from whichever copy of the procedure was nearest the panel. No deviation was written that night: the QMS terminal is in the QA office, QA is off site until 07:00, and the shift ends at 06:00. Three days later a process engineer who was not present drafts the record from the alarm summary in PI, a work order in SAP PM or Maximo, and a phone call to a technician now on days off. On Thursday the maintenance planner rebuilds the following week because a spare shows zero free stock in the storeroom, a permit clashes with an adjacent isolation, or the only technician qualified on that actuator is on the opposite roster.
Cost todayTypical observed ranges, at a site with four process units and three shifts. Twelve handovers a day at 30 to 45 minutes of writing and reading each is 6 to 9 hours of supervisor time daily, roughly 1,500 to 2,300 hours a year. Deviations run 250 to 600 a year; once the shift has dispersed, drafting and chasing evidence takes 8 to 14 hours each, so 2,000 to 8,400 hours across QA and process engineering, against a 30-day closure target that 15 to 30 per cent of records miss. Procedure questions arrive 4 to 9 times a shift. In the samples we have taken, 5 to 15 per cent were answered from a revision that was no longer controlling — but our sample here is thin, a few sites and a few hundred questions, so treat that band as indicative rather than established. Every figure above is a range across sites, not a result from any single client.
What the pilot takesFour to eight weeks, one process unit, and two of the five capabilities — normally revision-grounded shift assistance and deviation drafting, because both hang off the same revision-currency spine and can be measured against the same records. What the client spends is their own people's time: a process or QA owner at 4 to 6 hours a week, an automation or IT owner at 3 to 5 hours in total to open read-only access to the historian and the document management system, and two or three shift supervisors at about an hour a week comparing generated handovers with what they would have written themselves. Scope inputs are the controlling revision set for 40 to 80 procedures, 18 to 36 months of closed deviation records, and the asset register for that unit. We ask for no PLC or SCADA write access during the pilot, and none after it.
What changesEvery assisted answer carries a document number, revision and effective date, and we sample 100 answers against the document management system so citation accuracy is reported as a percentage with a date on it. Deviation initiation is measured as the interval between the alarm timestamp in the historian and the first saved QMS draft, which moves from days to within the shift for the event types the pilot covers. The refusal rate is published next to both: when revision currency cannot be confirmed the graph declines to answer, and in early weeks that has run around 5 to 12 per cent of questions. That refusal number is the one a plant manager should watch, because a system that never refuses is a system that is guessing.
Stays manualPermit-to-work authorisation, isolation and lock-out verification, deviation criticality classification, and the quality unit's release decision stay with named people. Each is a personal accountability under the site safety case and under EU GMP Annex 11 and 21 CFR Part 11, and a signature has to belong to someone who can be asked why. Field verification of drawings stays manual as well: the drawing-to-model check reads vendor-neutral exports and the master P&ID revision, and it cannot see a hand-marked redline taped inside a cabinet, so it raises candidate discrepancies for a technician to walk down rather than declaring an as-built. Writing to PLC or SCADA is not a residual we intend to close. Production persistence runs at level L2; the reality-connected level L3 is gated and partial, and we do not claim otherwise.
02 · Instrumented workflowsCut by Cut by the controlled record that binds the answer. Each workflow is owned by exactly one authority — the effective procedure text, the deviation record, the work order and permit, the engineering drawing and asset register, the shift handover log, and the change control record that authorises a revision. If two pieces of work would be signed against the same controlled record, they are one workflow, not two. What this axis deliberately leaves out: anything requiring a physical check of the plant, which no record can settle and which we do not attempt.
IO01
Shift answers bound to the effective revision
An operator asks at 02:10 whether a filter change during a hold step needs a second signature. The graph answers only from the work instruction that document control shows as effective at that minute, and returns the document number, revision, effective date and clause, or it returns nothing at all. A supervisor who today spends 10 to 25 minutes locating the right revision on a shared drive keeps that time; the answer still carries no authority to depart from the procedure, and any question the controlling document does not cover is handed back unanswered.
Shift Answer Card — one page per question carrying the answer, the controlling document number and revision, its effective date, the clause reference, and the timestamp of the question, filed into the shift record.
IO02
Deviation record drafted before the shift ends
A deviation is raised at 02:40 for a jacket temperature excursion of 3.1 degrees Celsius lasting 14 minutes during a hold step. Before that shift leaves site, the graph drafts the deviation record: the event timeline read from historian tags with their quality flags, the batch and equipment identifiers, the procedure clause the excursion touched, and the prior deviations logged against the same equipment. Quality assurance still classifies and approves; the cost recovered is the 6 to 30 hours of timeline reconstruction that otherwise happens a fortnight later by someone who was not present.
Draft Deviation Record with an attached event timeline and prior-occurrence list, submitted to the quality assurance queue with the observing shift named on it.
IO03
Outage work sequenced against spares and permits
A nine-day turnaround with 240 planned jobs. Each job is tested against on-hand spares net of existing reservations, the permits and isolations it requires, the competencies actually rostered on that day, and the window it has to fit inside; jobs that cannot start are named individually with the one binding constraint, rather than quietly resequenced into a schedule that looks complete. Where a site currently loses 15 to 30 jobs per outage to a missing part or an unavailable isolation, the arithmetic is roughly one planner-week of preparation against a day of critical-path slip.
Outage Schedule Pack — sequenced work order list, spares reservation list, permit and isolation prerequisites per job, and a named infeasible-jobs list stating the binding constraint for each.
IO04
Drawing-to-model check before design release
Before a modification package is released for construction, the graph compares the current P&ID revision against the 3D model and the asset register, tag by tag: line numbers, valve types, instrument tags, material and insulation class. Discrepancies are marked blocking or non-blocking and addressed to the responsible engineer by name. This replaces a manual squad check that typically runs 2 to 5 engineer-days per package; the engineer still adjudicates every item, and we do not touch the model.
Design Release Check Report — a discrepancy list by tag, each entry citing the drawing number and revision, the corresponding model object, and a blocking or non-blocking classification.
IO05
Handover record with every open item carried
A nightshift ends with three open work orders, an unclosed hot work permit, two alarms shelved at 23:15 and a deferred function test. The graph assembles the handover record from the work order system, the permit register, the alarm and event log and the outgoing supervisor's own notes, then carries forward every item that was open at the previous handover and remains open now, with a count of how many handovers it has survived. The failure this targets is the item that disappears silently across two changes of crew.
Shift Handover Record with a carry-forward ledger showing each open item, the handover at which it first appeared, and the number of handovers it has survived unclosed.
IO06
Revision difference brief for the roles affected
A procedure moves from revision 7 to revision 8 and three of its forty clauses change. The graph produces a difference brief per affected role, stating what changed, what it replaces and which step on the floor now differs, traced to the change control record that authorised it. Reissuing the whole document against a read-and-understood signature does not tell you whether anyone saw those three clauses; the brief and its acknowledgement list do.
Revision Difference Brief per role — changed clauses with old and new text side by side, the authorising change control number, and the named acknowledgement list of who must read it before next working the step.
IO01 · measure
Sample 50 cards after one month. Every card must resolve to a revision your document control master list shows as effective at the timestamp of the question, and none may resolve to a superseded, withdrawn or draft revision.
IO02 · measure
Over 30 consecutive deviations, count how many draft records quality assurance accepts with no factual correction to the timeline, and record the elapsed hours from deviation raised to draft submitted. The target is under four hours and within the observing shift.
IO03 · measure
At outage close, compare the infeasible-jobs list issued beforehand against the jobs that actually failed to start. Misses — jobs that failed for a constraint we did not flag — must stay under 5 per cent of planned jobs.
03 · Where this sector goesHorizon 18-36 months · our view, stated to be argued with
P01
Procedure metadata quality will cap plant assistant deployments before model quality does.
Most sites hold controlling procedures as scanned or exported PDFs on a file share, with revision and effective date printed in the header rather than held as a queryable field. Retrieval cannot prove currency against data that only exists as pixels in a page header, so the ceiling on answer quality is set by the records system, not by the model.
Fund a procedure metadata clean-up — document number, revision, effective date, superseded date, owning role, affected asset tags — and scope it as a records project with its own budget line before any assistant pilot is approved. Ask your document controller how many controlled documents currently carry a machine-readable effective date; that percentage is your realistic coverage.
P02
Regulated sites will require revision citation on every answer before accepting an assistant.
When an investigation asks which revision an operator was working from, an uncited answer cannot be defended and becomes a finding in its own right. Once one inspection turns on that question, quality units will write the citation requirement into their acceptance criteria for any software that speaks to the floor.
Put the citation requirement in the pilot acceptance test rather than the roadmap. A supplier who can only promise it for a later release is asking you to run an unciteable system through an audit window.
P03
Writeback to PLC and SCADA stays prohibited on regulated sites through the horizon.
Zone and conduit segmentation under IEC 62443, plus the functional safety case behind any interlock, mean an agent write path is a change to that safety case and has to be re-validated. No plant re-validates a safety case to save planning hours, and insurers price the exposure accordingly.
Stop scoping pilots that assume control writeback. Budget the plant integration as one-way export from the historian and the work order system, and take the value in the records and planning layer where it is actually available.
P04
Outage value will be judged on how many permitted windows are used, rather than hours saved.
Technician hours are already contracted and largely fixed for a turnaround. The scarce resource is permitted access: isolation slots, confined space entries, hot work windows, and the crane hours they depend on.
Instrument the baseline this outage. Record, for every planned job that did not start, the single constraint that stopped it. Without that list you have no way to test whether any planning system paid for itself.
Where we disagree with the sector
The common view is that the hard and valuable integration in an industrial deployment is live connection to operational technology — the historian, the SCADA layer, the alarm system — and that document work is a low-value extra bolted on afterwards. We think that is backwards. Historian data is already structured, tagged, timestamped and exportable through interfaces that have existed for twenty years; connecting to it is a fortnight of engineering. The failures that cost real money sit in documents whose revision state is ambiguous: a deviation investigation that stretches from three days to six weeks because the timeline was reconstructed from memory, or a job that lost its outage window because the permit prerequisite lived in a procedure nobody had reread since revision 5. We hold this with stated uncertainty. Our production evidence at 155 nodes comes from a financial-information deployment, and our industrial evidence is from site walkdowns and pilot-scale work rather than a large sample of plants. If a client's document control is already clean and machine-readable, the argument weakens considerably, and we would say so rather than sell around it.
04 · Node graph, re-cut for this sectorAgainst the eleven-stage reference
Driving axisRevision currency, and what it costs to answer from a superseded document. Compute and retrieval speed are cheap in this sector. The expensive failure is a confident answer grounded in a revision that was correct last quarter, because that answer can produce a deviation, a rejected batch, a failed isolation or an injury, and it does so while looking exactly like a correct answer. The graph therefore spends most of its effort proving currency rather than composing prose. In the industrial configuration of the 155-node reference topology — 114 reasoning nodes and 41 deterministic workers — the deterministic workers carry the load, because currency, tag identity, permit feasibility and spares availability are all questions with a definite answer that a program can check and a language model should not be trusted to judge alone. The second cost driver is the read-only boundary against PLC, SCADA and historian, which is absolute and shapes what the graph is permitted to be good at.
Stage 01, system entry, expands furthest. Nothing is admitted into the working set without a document number, revision, effective date and superseded date attached as fields; content that carries only a printed header is refused at entry rather than ranked lower, and refused content is listed back to the document controller as a coverage gap. Stage 02, realtime signal, collapses to a one-way subscription: historian tags, alarm and event records, work order state changes, all read-only, with no return path of any kind. Stage 04's seven parallel evidence branches are repointed at seven plant evidence classes — controlling procedure text, deviation and corrective action history, work order and spares, permit and isolation register, drawing and asset register, historian and alarm log, shift handover log. Concurrency across all seven is safe precisely because every branch is read-only and none can alter the state another is reading. A synchronising barrier before the conflict resolver is mandatory: a work order can say a part is reserved while the storeroom record says it was consumed on another job, and the resolver must see both branches complete before it rules. Stages 05, 06 and 07 — poster pipeline, long-form article, video — nearly disappear, since the output here is a controlled form rather than a publication; a thin remnant of stage 06 survives to compose the investigation narrative in IO02. Stage 08, question and answer, expands and inherits the revision-citation obligation on every response. Stage 09 becomes cross-shift consistency rather than cross-product: the same question asked on nightshift and dayshift must return the same clause. Stage 10, audit and reproducibility, expands — the claim map becomes a claim-to-clause map, and the call-trace archive holds the document control state as it stood at the moment of the question, not as it stands today. Stage 11 remains the terminal gate at node AI-116 and fails closed: an answer with an unresolved citation, or one resolving to a superseded revision, does not leave the system.
Revision stamp at entry
Every admitted chunk carries document number, revision, effective date and superseded date as queryable fields. Unstamped content is refused, not downweighted, and the refusal list is returned to the document controller as a named coverage gap rather than absorbed silently.
Barrier at conflict resolver
All seven evidence branches must complete before the resolver rules, because a spares record, a work order and a permit register routinely disagree about the same job. Allowing the resolver to start early produced answers that were internally consistent and wrong.
Read-only OT boundary
PLC, SCADA, DCS, safety instrumented systems and the historian are read-only without exception, enforced at the connector rather than by instruction to a node. There is no configuration flag that turns this off, and we will not build one.
Terminal gate AI-116
The publication gate fails closed. An answer whose citation cannot be resolved against the document control master list as of the question timestamp is withheld and logged as a refusal, which means a stale or incomplete document register shows up as visible refusals rather than as quiet errors.
05 · Correctness and authorityWhat decides, and what we refuse
Checkers
Six checkers are specific to this sector, each bound to a class of claim. The revision currency checker runs on any claim that quotes or paraphrases procedure text: it resolves the cited document number and revision against the document control master list as of the question timestamp and fails the claim if the revision was superseded, withdrawn or still in draft at that moment. The tag and asset identity checker runs on any claim naming equipment: it resolves the tag against the asset register and the current drawing revision, and fails on a tag that is absent, decommissioned, or duplicated across units — duplicated tags across trains are the most common cause of a plausible wrong answer we have seen. The permit and isolation feasibility checker runs on any scheduling claim: it tests whether the required permit types, isolation points and competencies exist and are unexpired inside the proposed window, and fails the individual job rather than the schedule, so the planner sees which job died and why. The spares availability checker runs on any claim that a job can start: on-hand quantity net of existing reservations, against lead time and the window. It has a known weakness — parts physically on site but not booked in, such as kitted material and vendor-managed stock, are invisible to it, so it reports a confidence band and defers to the storeman rather than asserting a shortfall. The historian provenance checker runs on any claim citing a process value and requires tag name, engineering unit, timestamp with timezone, aggregation method and the quality flag; values marked bad, substituted or manually entered are rejected rather than averaged in. The quality clause checker runs on deviation drafts: every classification statement must map to a clause in the site's own quality manual, and general regulatory knowledge is not accepted as a substitute for the site's text. Twelve critic nodes and six reality gates sit above these, and the terminal gate is AI-116.
Gate policy
These refusals are absolute and not configurable by contract, by client request, or by a settings flag. No agent writes to a PLC, DCS, SCADA layer, safety instrumented system or historian; the connector is read-only and there is no write path to enable. No setpoint change, no alarm shelving, suppression or reprioritisation, no bypass, force or defeat of an interlock. No issuing, extending, transferring or closing a permit to work, and no authoring or approval of an isolation or lock-out list. No signature or approval on anything the quality system requires a named person to sign: deviation approval, batch disposition, change control approval, calibration release, training sign-off. No change to a document's effective status in the document control system. No dispatch of a technician to a live task, and no instruction to any person to work on plant. Everything the graph produces in this sector is a draft addressed to a named human, and it is delivered with the evidence that would let that person disagree with it.
Acceptance
Take 200 shift questions asked in the last 90 days and already answered by your own supervisors, with the answers withheld from us. We answer them cold. To pass: at least 190 of the 200 answers cite a document number, revision and clause that your document control master list shows as effective at the timestamp the question was originally asked; zero answers cite a revision that was superseded at that timestamp; and the median time to answer is under 60 seconds against the supervisor baseline you measure yourself. One superseded-revision citation in 200 fails the whole pilot, with no partial credit and no renegotiation of the threshold after the run. Questions we decline to answer count as neither pass nor fail, but if we decline more than 40 of the 200 the coverage is too thin to be worth deploying and we will say so.
Will not automate
We cannot verify the physical state of the plant, so we refuse to automate anything that depends on it. No record tells us whether a valve is actually shut, whether the laminated copy taped inside the panel door matches the revision now in force, whether the nameplate on the pump matches the tag in the register, or whether the part in the kitting bin is the part the work order calls for. Every answer we give is grounded in what the records say, and the records and the plant diverge. For that reason we will not author isolation or lock-out lists, will not confirm a work front is safe to enter, and will not certify a spares position as physically present. A second, narrower limit: handwritten shift logbooks and redlined drawings. Optical character recognition on marked-up engineering drawings and handwritten log entries is not accurate enough for us to treat as an authority, and we have not measured its error rate at a level we would be willing to sign, so redlines enter the system as an unverified attachment for a human to read and never as evidence a checker may rely on. Removing that limitation would need a measured accuracy study on a real site's own handwriting, which we would run as a paid piece of work with published results rather than assert as a capability.
06 · Where this sector entersCross-reference to the capability catalogue in section 03 and the horizons in section 05
Fastest entry · H0
C-02 attribution on 100 questions the floor has already asked. Ten working days, read-only, and every answer carries a document number, a revision and an effective date.
Pilot · H1
C-06 on shift handover or alarm triage for one process unit. Four to eight weeks, two capabilities at most, read-only against the historian.
The long part · H2 to H3
The site procedure ontology and the fault corpus behind it. Quarters, and read-only against PLC, SCADA and historian does not change at any point.
The change was approved in eleven days. Nine of them went on finding out what else it touches.
An engineering change is a small decision wrapped in a large search. The decision — accept the new tolerance, move the fastener, requalify the supplier — takes an experienced engineer an afternoon. Finding every drawing, every variant, every tooling record, every technical publication and every qualification that the change invalidates takes the other nine days, and it is done by people reading PLM search results. This is a search-and-traceability problem sitting inside a configuration-management system, which is exactly the shape a typed ontology and a node graph handle well. It is also the sector where the gate matters most: nothing we build releases a change, and on safety-critical parts nothing we build dispositions a non-conformance.
The barrier
Configuration data is complete and unusable at the same time. The part exists in the PLM, the requirement exists in the requirements tool, the test result exists in the lab system, the supplier deviation exists in an email, and the technical publication exists in a separate authoring environment. Each is correct. None of them share an identifier that means the same thing, and the relation that would let you traverse from a changed tolerance to the maintenance manual paragraph it invalidates has never been written down. Engineers reconstruct it by memory, which works until the engineer who remembers it leaves.
01 · The work as it standsObserved shape and observed cost · ranges from programmes of this size, not any named client
How it runs todayA change request is raised against a released part. An engineer runs a where-used query in the PLM, which returns the assemblies but not the technical publications, the tooling, the supplier qualification or the test evidence. They then ask three colleagues, search a shared drive, and open the previous change notice on a similar part to see what it touched. The change board meets weekly. Changes that arrive with an incomplete impact list are deferred, and deferral is the most common outcome rather than rejection.
Cost todayTypical observed ranges at a programme with several thousand released part numbers and a weekly change board. A routine change consumes 6 to 20 engineering hours before the board sees it, and a change touching a qualified supplier part runs several times that. Deferral rates of 25 to 45 per cent at first presentation are common, and the dominant deferral reason is an incomplete impact assessment rather than a disputed engineering judgement. A missed downstream document typically surfaces months later, during a configuration audit or a field event.
What the pilot takesSix to eight weeks, scoped to one product family and one change type. We ask for read access to the PLM, the requirements tool and the publication source, plus 100 to 300 closed change notices with their final impact lists. Read-only throughout: no change notice is raised, approved or routed by the system during a pilot, and it holds no PLM write credentials at any point.
What changesMeasured by replaying closed change notices, so the comparison is against what your own engineers concluded. First, impact recall: of the items on the final approved impact list, how many the graph found before the board met, reported per item type because the failure is never uniform. Second, precision, because an impact list with forty spurious entries is worse than a short one. Third, elapsed hours from request raised to a reviewable impact list. On the last engagement of this shape, recall on assemblies and drawings was strong and recall on technical publications was poor until we modelled the publication-to-part relation explicitly, which took two of the eight weeks and was not in the original plan.
Stays manualThe engineering decision, the change board approval, the release into the controlled vault, the airworthiness or safety judgement and the disposition of any non-conformance on a safety-critical part. The system assembles the impact list, re-measures what can be re-measured, and states which relations it could not resolve. It proposes; the board decides.
02 · Instrumented workflowsCut by the unit of work that enters the graph: one change, one requirement, one part, one baseline
DE01
Change impact expansion across the configuration
A proposed change is expanded across assemblies, variants, tooling, qualification records, test evidence and technical publications by traversing declared relations rather than by keyword search. Each item on the list carries the relation path that put it there, so an engineer can dismiss a wrong one in seconds.
Impact list per change notice, each entry citing the relation path and the source system, plus an unresolved list naming relations the graph could not traverse
DE02
Requirement to test traceability, both directions
Requirements are linked to the tests that verify them and the results that closed them, and the graph reports both orphans: requirements with no verifying test, and tests verifying nothing in the current baseline. Both are common and neither is visible from inside a single tool.
Traceability matrix with two orphan registers, dated against a named configuration baseline
DE03
Supplier deviation and non-conformance preparation
A supplier deviation request is assembled against the drawing, the requirement it affects, the qualification basis and the history of similar dispositions on the same part family. The precedent is presented with its outcome, so the reviewing engineer sees what was decided before and why.
Prepared disposition file with the affected requirement, the precedent set, and an explicit statement of what evidence is missing
DE04
Model rebuild and re-measurement against the tolerance table
A parametric model is built or amended from a specification through typed CAD operations, exported, then re-opened and re-measured against the tolerance table. Success is not the model reporting that it built; it is the exported geometry measuring correctly and the constraint tree surviving a rebuild.
Model and export with a measurement report against every dimension in the tolerance table, plus the operation log that produced it
DE05
Technical publication generation from the released configuration
Maintenance and service documentation is generated from the configuration that was actually released rather than from the last time somebody updated the manual, with each procedure step tied to the part and revision it applies to.
Publication draft with a step-to-part-revision map, and a change list against the previous issue
DE06
Configuration audit: as-designed against as-built
The design baseline, the build record and the maintenance record are reconciled unit by unit. Differences are classified as documentation lag, approved deviation, or genuine discrepancy, which is the classification that decides who has to act.
Audit register per unit, each difference classified with the evidence on both sides and the record that governs
DE01 · measure
Replay 100 closed change notices. Recall against the final approved impact list, broken out per item type, and precision reported alongside it. A single number here hides the failure mode that matters.
DE02 · measure
Against a baseline your team has already audited: agreement on which requirements are unverified, and the count of trace links the graph proposes that a systems engineer rejects on review.
DE04 · measure
On 40 specifications with known-good models: how many rebuilt models export geometry that measures within tolerance on every dimension, and how many fail the constraint-tree rebuild. Both are pass or fail, not scored.
03 · Where this sector goesHorizon 18-36 months · our view, stated to be argued with
P01
Model-based definition makes the checker mandatory rather than optional.
As the annotated model replaces the drawing as the authority, there stops being a human-readable artefact for a reviewer to sanity-check. Whatever produces or amends the model has to be verified by re-measurement, because there is no drawing left to eyeball.
Decide now what re-measures your models, because the drawing will not be there to catch it
P02
Programmes buy traceability before they buy generation.
Generative design demonstrations are impressive and land in a part of the process that is not the bottleneck. The bottleneck is knowing what a change touches. We expect budget to move toward traceability and configuration integrity, which photograph badly and pay better.
Compare the hours your programme spends on impact assessment against the hours it spends on concept design
P03
Supplier quality evidence becomes a machine-readable contractual deliverable.
Tier-one suppliers already send certificates as scanned PDFs. As primes start reconciling incoming evidence automatically, structured delivery moves from a nice-to-have into the terms, and suppliers who cannot provide it lose position.
If you are a supplier, find out now what structured evidence your largest customer will require
P04
A certification authority asks how AI-assisted design evidence was produced.
The question is not whether the tool is allowed. It is whether the applicant can show the operation log, the verification and the human approval for a specific artefact in a specific submission. Firms without that record will withdraw the artefact rather than answer.
Keep the operation log for every generated artefact from the first day, not from the day you are asked
Where we disagree with the sector
The sector's attention is on generative design and on automated CAD authoring. We think the return there is real but small relative to the effort, because concept work is not where programme time is lost. The money sits in change impact and configuration integrity, which is unglamorous, demonstrates poorly, and is the reason changes are deferred at the board. We would rather be measured on deferral rate than on how quickly a model appears.
04 · Node graph, re-cut for this sectorAgainst the eleven-stage reference
Driving axisRelation completeness in the ontology. Almost all of the difficulty and almost all of the pilot's engineering time sit at L1, not in the reasoning stages above it.
The graph is unusually shallow and unusually wide. Stage 04 fans out one branch per source system — PLM, requirements, test, publications, supplier records — and the branches rejoin at a normaliser whose job is identifier resolution rather than conflict resolution. Stage 05 mostly disappears. What carries the weight is the relation model beneath it: a where-used query is a traversal, and a traversal that has no declared relation returns nothing without telling you it found nothing, which is why the unresolved list is an output rather than a log entry.
Identifier resolution is the first node
The same part carries different identifiers in five systems. Nothing downstream works until that mapping is explicit and testable, and it is never as complete as the client expects.
Missing relations are reported, not inferred
Where a relation has not been declared, the graph says so. It does not guess a link from name similarity, because a plausible wrong relation on a configuration is worse than a gap.
Tool operation runs sandboxed
CAD work happens on a working copy with pinned tool and library versions. Release into the controlled vault is a separate, human, gated act.
Safety-critical parts are partitioned
Parts on the safety-critical list are marked at L1 and the disposition path is absent for them by construction, not disabled by a setting.
05 · Correctness and authorityWhat decides, and what we refuse
Checkers
Geometry is re-measured after export against the tolerance table. Constraint trees must survive a rebuild. Every impact-list entry must cite a declared relation path. Qualification currency is checked against the effective date, and an expired qualification blocks rather than warns.
Gate policy
Read-only against PLM, the requirements tool and the test system. No change notice raised or approved, no baseline change, no release into the controlled vault, no supplier communication. Writes land only in a working branch that a named engineer promotes.
Acceptance
Replay 100 closed change notices and compare against the approved impact lists, per item type. Separately, rebuild 40 models from specification and require every exported dimension to measure within tolerance. Your configuration management function runs both.
Will not automate
We will not disposition a non-conformance on a safety-critical part, and we will not produce an airworthiness or safety judgement. We can assemble every piece of evidence a reviewing engineer needs and show the precedent, but the determination belongs to a person with an authorisation, and we have no way to check the output of that judgement. We also decline to write directly into a controlled vault, in any configuration, for any client.
06 · Where this sector entersCross-reference to the capability catalogue in section 03 and the horizons in section 05
Fastest entry · H0
C-01 reconciliation on as-designed against as-built for one product family. Ten working days, read-only, and the output is a difference register your configuration audit can use immediately.
Pilot · H1
C-03 cause ranking applied to change impact, on one change type. Six to eight weeks against 100 closed notices, with the relation model as the real deliverable.
The long part · H2 to H3
C-05 tool operation against CAD, and the relation model extended across every source system on the programme. This is measured in quarters and it is where the compounding is.
IND-08 · Pharmaceutical, medtech and clinical operations
The protocol amendment runs to two pages. The documents it invalidates run to forty.
Regulated life sciences is the clearest case of work where the rule is written down and the conformance check is done by hand anyway. A protocol amendment cascades into the informed consent form, the site instructions, the case report form, the statistical analysis plan, the monitoring plan and the training records — and the cascade is worked out by a person reading documents. The same shape governs a labelling change across markets, a deviation write-up against site procedures, and a submission section against the guidance in force. Everything here sits behind a gate we do not move: no batch release, no submission, no safety filing, and no judgement about a patient.
The barrier
Document control was built to stop uncontrolled change, and it does that well. What it does not do is tell you what a controlled change implies. The effective version of every document is knowable; the dependency between them is not recorded anywhere, so it is reconstructed each time by whoever has been at the company longest. Add a second study, a second market or a second manufacturing site and the reconstruction cost rises faster than the document count, because the dependencies are between documents rather than inside them.
Ontology objects
ProtocolAmendmentStudy siteDeviationCAPABatch recordSpecificationSubmission sectionStandardAdverse eventMarket authorisationTraining record
01 · The work as it standsObserved shape and observed cost · ranges from operations of this shape, not any named sponsor
How it runs todayAn amendment is agreed. A clinical operations manager opens the document list for the study, works through which documents reference the changed section, and raises change requests against each. Sites are notified, consent is re-obtained where required, training is re-assigned, and the study file is updated. Six weeks later a monitoring visit finds one site using the previous version of a worksheet, because the notification reached the coordinator who had left.
Cost todayTypical observed ranges, not a benchmark. A substantial amendment on a multi-site study absorbs 80 to 300 hours across clinical operations, regulatory and quality before the first site is activated on the new version. Deviation write-up runs 45 minutes to three hours per event depending on criticality. Batch record review at 20 to 90 minutes per batch, with the exception-based approach limited less by regulation than by the difficulty of proving which records are truly exception-free.
What the pilot takesSix to eight weeks, scoped to one study or one product family in one region. We ask for a closed study's document set with its amendment history, or 200 to 400 closed deviations with their classifications. Read-only against the document management and quality systems, deployed in your validated environment or alongside it. We will tell you plainly that validation of the environment is your process and it usually sets the schedule, not us.
What changesReplayed against amendments and deviations already closed. First, dependency recall: of the documents your team actually amended, how many the graph identified from the amendment text, and how many it proposed that your team rejected. Second, on deviations, agreement with the recorded criticality classification, reported as a matrix, since over-classifying is a real cost and under-classifying is a real risk. Third, elapsed time to a reviewable impact list. On the first engagement of this shape we found the dependency recall was high for documents inside the document management system and near zero for site-level worksheets kept in a separate share, which is a finding about the client's estate rather than about the graph.
Stays manualBatch disposition, submission, safety reporting to an authority, the causality assessment of an adverse event, the criticality determination of record, and every decision that touches a subject or a patient. The system prepares the file, checks conformance against the version in force, and names what it could not resolve. A qualified person signs, as they always did.
02 · Instrumented workflowsCut by the unit of work that enters the graph: one amendment, one deviation, one batch, one market
PH01
Amendment impact across the document set
The amendment text is read against the study document set, and every document, section and training assignment that references the changed material is identified with the reference that put it there. Site-level documents are included where they are reachable, and named as unreachable where they are not.
Impact register per amendment: document, section, reference path, required action, and an explicit unreachable list
PH02
Submission section drafting with a clause trace
Sections are drafted against the applicable guidance and the protocol version in force, with each paragraph tied to the guidance clause it answers and the study document it draws on. Uncited assertions do not reach the draft.
Draft section plus a clause trace table, and a gap list naming each required element with no supporting document yet
PH03
Deviation write-up and criticality preparation
A reported deviation is written up against the procedure version that was actually in force on the event date, with the affected requirement identified and comparable closed deviations presented with their classifications and outcomes.
Prepared deviation record with the governing procedure version, the precedent set, and a proposed classification the quality unit accepts or overrides
PH04
Batch record review support
Records are checked against the specification and the master batch record for completeness, arithmetic, in-process limits and signature presence. The output separates records with no exception from records needing a reviewer, with the specific field named in every case.
Review queue split by exception, each flagged record naming the field, the limit and the value, with the arithmetic re-run shown
PH05
Literature and safety evidence assembly
Published literature and internal safety data are assembled against a defined question with the search strategy recorded as it ran rather than reconstructed afterwards, and every included and excluded item carries its reason.
Evidence pack with a recorded, rerunnable search strategy, inclusion and exclusion reasons, and a source list with dates
PH06
Labelling and instructions conformance across markets
Labelling, instructions for use and packaging text are checked against the applicable standard and the authorisation in each market, including the translations, which are checked against the source rather than against each other.
Conformance register per market: text element, standard clause, authorisation reference, and each divergence with the market it applies to
PH01 · measure
Replay 30 closed amendments. Recall against the documents your team actually changed, and the count of proposed documents a reviewer rejected. We report the unreachable list separately because it is a finding about your estate.
PH03 · measure
On 300 closed deviations: agreement with the recorded criticality as a confusion matrix. Under-classification and over-classification are reported separately and never netted off.
PH06 · measure
Take 50 label elements already reviewed by regulatory affairs across three markets. Every divergence the reviewer found must be found, and every additional divergence proposed must be defensible against a named clause.
03 · Where this sector goesHorizon 18-36 months · our view, stated to be argued with
P01
Computerised system validation adapts to non-deterministic components, slowly and unevenly.
Existing validation practice assumes the same input gives the same output. Sponsors will resolve this first by validating the deterministic wrapper — inputs, gates, ledger, human approval — and treating the model as a component under monitoring. Expect that pattern to become the default before any guidance names it.
Design so the validatable boundary is the wrapper, not the model, from the first sprint
P02
Structured content management stops being an aspiration and becomes procurement's requirement.
Once a sponsor sees that dependency tracing works on documents inside the content system and fails on documents outside it, the business case for structured authoring writes itself. The driver will be automation coverage rather than reuse, which is not what the category was sold on.
Count what fraction of your study documents live somewhere a system can reach
P03
Inspection readiness becomes a continuous state rather than a project.
Assembling an inspection response in ten days is a well-practised emergency. When the evidence trail is maintained continuously, the emergency stops existing, and the firms who get there first will be visibly less disrupted by an inspection.
Ask what your last inspection preparation cost in person-weeks, then treat that as the annual budget
P04
Decentralised and hybrid trials break document control before they break anything else.
More sites, more devices and more remote consent multiply the document dependency graph without adding staff to maintain it. The failure mode is not clinical; it is a site running an out-of-date worksheet for four months.
Test your dependency tracing on a hybrid study before you commit to the next one
Where we disagree with the sector
Most of the sector's attention goes to drafting submission text faster. We think drafting is the cheapest part of the regulatory cycle and the least defensible place to claim a return, because review time is set by the reviewer, not the writer. The expensive, under-instrumented work is the dependency cascade after a change, and it produces no impressive demonstration. We would rather be judged on how many site-level documents were found stale.
04 · Node graph, re-cut for this sectorAgainst the eleven-stage reference
Driving axisVersion currency. Almost every failure in this sector is the right answer taken from the wrong version of a document, so the effective-date check runs before anything else and blocks rather than warns.
The graph runs narrow and deep, one branch per obligation or per document family. Stage 04 has three branches rather than seven: the controlled document set, the applicable standard or guidance, and the quality and manufacturing record. Stage 10 expands, because the evidence ledger here is inspection evidence and has to survive a question years later. A distinctive addition sits at stage 01: every object is stamped with the version that was effective on the event date, and any node reading a document at a different version fails its own precondition.
Effective-date resolution comes first
Before any reasoning, the graph resolves which version of every document governed on the relevant date. A run that cannot resolve this stops rather than proceeding on the current version.
Unreachable documents are an output
Anything the graph cannot see is listed by name and location. A dependency register that silently omits what it could not reach is worse than no register.
The validatable boundary is the wrapper
Inputs, typed actions, checkers, gates and the ledger are deterministic and can be qualified. The model sits inside as a monitored component, and we design for that from the start.
No disposition node exists
There is no path in the graph that releases a batch, files a report, or assesses causality. Those nodes were never built, and adding them is not a configuration change.
05 · Correctness and authorityWhat decides, and what we refuse
Checkers
Every drafted sentence resolves to a guidance clause or a study document at the version effective on the relevant date. Batch arithmetic is recomputed against the specification. Search strategies must rerun and return the same set. Translations are checked against the source text, never against a sibling translation.
Gate policy
No batch release, no regulatory submission, no safety report to an authority, no protocol change, no site communication and no subject-level action. Read-only against the document management, quality and manufacturing systems, provisioned to named individuals.
Acceptance
Replay 30 closed amendments against the documents actually changed, and 300 closed deviations against their recorded classifications. Both are run by your quality unit. A run that cannot resolve effective dates counts as a failure, not a partial pass.
Will not automate
We do not assess the causality of an adverse event, disposition a batch, or make any determination about a patient or a subject. We tried building a causality-assessment support node on a historical dataset and abandoned it: the outputs were confident, internally consistent, and impossible to verify against anything except a clinician's opinion, which is exactly the shape of failure we refuse to ship.
06 · Where this sector entersCross-reference to the capability catalogue in section 03 and the horizons in section 05
Fastest entry · H0
C-01 reconciliation replaying one closed amendment cascade. Ten working days, read-only, and it usually returns a list of stale site documents nobody knew about.
Pilot · H1
C-04 rule-bound drafting or C-06 triage on deviations. Six to eight weeks; your validation process, not our engineering, sets the calendar.
The long part · H2 to H3
Continuous inspection readiness — the evidence ledger maintained as a state rather than assembled as a project. Multiple quarters, and the only version of this that pays.
IND-09 · Healthcare delivery and payer operations
The denial arrives eleven days later, citing a criterion the submission had already met.
This practice is deliberately confined to the administrative surface between a provider and a payer. Prior authorisation, coding completeness, denial and appeal, contract reconciliation, rostering. It is document-and-rule work that consumes clinician and administrator time without touching clinical judgement, and it is exactly the shape a rule-bound drafting and reconciliation stack handles. We do not build clinical decision support, we do not triage patients, and no output of ours reaches a patient. Those are not cautious framings — they are the boundary of what we will sell into this sector, and it does not move.
The barrier
The payer's criteria are published, the provider's documentation exists, and the two are compared by a person under time pressure who is not the person who wrote either. A submission fails on a criterion the clinical record already satisfies, because the sentence that satisfies it sits in a progress note rather than in the field the reviewer reads. The appeal then costs more than the original submission and succeeds often enough to prove the first decision was avoidable. Nothing here is a knowledge problem; it is a document-to-rule mapping problem repeated thousands of times a month.
01 · The work as it standsObserved shape and observed cost · ranges from operations of this shape, not any named provider
How it runs todayA prior authorisation is initiated by a clinic administrator who opens the payer's policy document, reads the criteria, then searches the record for evidence of each one. Missing evidence means a message to the clinician, who replies between patients. The packet is submitted through a portal. Some fraction is denied for reasons that are administrative rather than clinical, and an appeal is assembled by a different person who repeats most of the original search.
Cost todayTypical observed ranges rather than a benchmark. Prior authorisation absorbs 15 to 45 minutes of administrator time per request, plus clinician interruption. Denial rates on first submission are widely reported in the mid-to-high single digits to low teens depending on service line, and a large share of appealed denials are overturned, which is the number that tells you the first decision was avoidable. Appeal assembly runs 30 to 90 minutes. Remittance reconciliation against contract terms is often not done at line level at all, because nobody has the hours.
What the pilot takesSix to eight weeks, scoped to one service line and one payer. We ask for 300 to 800 historical requests with their outcomes, including the denials and the appeal results, plus the payer policy documents in force at the time. De-identified where the scope allows it. Deployment is inside your boundary; nothing about a patient leaves it, and we will decline an engagement that requires otherwise.
What changesMeasured by replaying historical requests. First, criterion coverage: for each policy criterion, whether the graph located supporting evidence in the record and cited it to a document and a date. Second, agreement with the outcome that actually occurred, with the denials examined individually rather than averaged. Third, administrator minutes per request, timed rather than estimated. The honest result from the first engagement of this shape: coverage was strong on structured fields and weak on evidence that only existed in free-text notes, and the free-text portion is where the avoidable denials were concentrated.
Stays manualEvery clinical judgement, without exception. The coding of record is confirmed by a certified coder. The submission is made by a person. The appeal argument is approved by a person. No output reaches a patient, and the system holds no capability to communicate with one.
02 · Instrumented workflowsCut by the unit of work that enters the graph: one request, one denial, one remittance, one roster period
HC01
Authorisation packet assembled against published criteria
The payer's policy is read criterion by criterion into the ontology, and the record is searched for evidence of each one. Criteria with evidence are cited to a document, a date and an author. Criteria without evidence are named individually, so the clinician is asked one specific question rather than sent a general request.
Packet with a criterion-by-criterion evidence table, plus a named list of exactly what is missing and who can supply it
HC02
Denial analysis and appeal evidence assembly
Denials are grouped by reason code and by whether the cited criterion was in fact evidenced in the original record. The second group is the actionable one, and it is assembled into an appeal with the evidence located and referenced.
Denial register split into evidenced and genuinely unmet, with a drafted appeal for the first group citing the record location for each criterion
HC03
Documentation completeness check before submission
Before a claim or a request goes out, the documentation is checked for the elements the payer requires, with each gap named against the requirement it fails. Run before submission, this is cheap; run after a denial, the same work costs several times as much.
Pre-submission gap list, each item citing the requirement and the record field that would satisfy it
HC04
Remittance reconciliation against contract terms
Paid amounts are reconciled line by line against the contracted rate, the fee schedule version in force, and the terms that govern the service. Underpayments are classified by cause, which is what turns a spreadsheet into a recoverable position.
Line-level variance register with the contract clause, the expected amount, the paid amount, and a cause classification
HC05
Roster and capacity planning under the applicable rules
Rosters are built or checked against award terms, rest requirements, skill mix and leave, with each constraint applied as a declared rule rather than a heuristic. Violations are shown with the rule that was breached and the shift that caused it.
Roster with a constraint report: every applied rule, every violation, and the shift and person it attaches to
HC06
Referral and pathway document reconciliation
Referrals, pathway steps and the documents that should accompany them are reconciled so that a missing item is found before the appointment rather than at it. The output is administrative: what is missing and who holds it.
Pathway completeness register per referral, naming the missing document, its owner and the step it blocks
HC01 · measure
On 500 historical requests: per criterion, whether evidence was located and correctly cited, verified against the original record by your own reviewers. Reported per criterion, never as one coverage figure.
HC02 · measure
Against denials your team already appealed: how many the graph classifies as evidenced-but-denied, and whether that classification matches the appeal outcome. Both false positives and false negatives are listed.
HC04 · measure
Reconcile one month of remittance at line level against contract terms and compare with a manual reconciliation of the same month. Differences are examined individually, in both directions.
03 · Where this sector goesHorizon 18-36 months · our view, stated to be argued with
P01
Both sides automate the same exchange, and the paperwork accelerates rather than shrinks.
Payers are automating review while providers automate submission. The near-term result is faster cycles with the same or higher volumes, not less work. The advantage goes to whoever can evidence a criterion at line level, because that is what survives the other side's automation.
Build for evidence citation rather than for volume; volume is about to stop being the constraint
P02
Electronic prior authorisation mandates move the bottleneck to documentation, not decision time.
Where regulation is compressing authorisation turnaround, the binding constraint becomes whether the required evidence exists in a form the payer can read. Turnaround rules do not create documentation.
Measure how many of your denials were evidence-location failures rather than clinical disagreements
P03
Providers start reconciling remittance at line level because it becomes affordable.
Contract-level reconciliation is common; line-level reconciliation against the fee schedule version in force generally is not, purely on cost. When that cost falls, a recoverable position appears that has been invisible on the balance sheet.
Reconcile one month at line level before deciding whether this matters to you
P04
Free-text clinical notes become the thing that gets structured, for administrative reasons.
The evidence that satisfies payer criteria disproportionately lives in narrative notes. The pressure to structure it will come from the revenue cycle rather than from clinical informatics, which is an uncomfortable but predictable driver.
Find out what share of your avoidable denials turn on evidence that exists only in free text
Where we disagree with the sector
Much of the market is being sold clinical decision support and ambient documentation. We do not build either, and we think the administrative surface pays better and carries a fraction of the risk. Our disagreement is about where the verifiable work is: a payer criterion is a rule with a checkable answer, and a clinical judgement is not. We will lose deals to vendors promising more, and we would rather lose them than build a node we cannot verify.
04 · Node graph, re-cut for this sectorAgainst the eleven-stage reference
Driving axisCriterion-level evidence location, under a hard data boundary. The graph is small; the constraint that shapes it is that nothing about a person leaves your environment.
A compact graph, typically 20 to 40 nodes, running entirely inside the client boundary. Stage 04 has two branches: the payer policy set and the clinical record. Stage 05 collapses to a single packet format. The distinctive stage is a per-criterion evidence node — one node per criterion rather than one node per request — so that a criterion that is failing systematically across hundreds of requests becomes visible as a pattern rather than a series of individual denials.
One node per criterion
Criteria fail in characteristic ways. Modelling them individually is what turns a denial rate into a fixable list, and it costs almost nothing at this graph size.
Boundary is absolute
Deployment sits inside your environment with model routing under your policy. If that is not possible, we decline the engagement rather than negotiate the boundary.
No clinical node exists
Nothing in the graph forms a clinical conclusion, ranks a treatment, or scores a patient. That path does not exist in the topology.
Free-text evidence is cited, not summarised
Where evidence sits in a narrative note, the output quotes and locates it. It does not paraphrase a clinician's words into a criterion.
05 · Correctness and authorityWhat decides, and what we refuse
Checkers
Each criterion must resolve to a record location, a date and an author, at the policy version in force on the request date. Remittance arithmetic is recomputed against the contracted rate. Roster constraints are recomputed from the award text. Anything that cannot be located is reported as missing rather than inferred.
Gate policy
No clinical decision, no coding of record without a certified coder, no submission, no appeal filing and no patient communication of any kind. The system holds no channel to a patient and no write access to the clinical record.
Acceptance
Replay 500 historical requests with your own reviewers checking the citations against the original record. Per criterion, not in aggregate. A citation to the wrong document counts as a failure even where the outcome would have been the same.
Will not automate
We do not build clinical decision support, patient triage, diagnosis assistance, or anything that produces a conclusion about a person's health. We also will not run a workflow where a model's output reaches a patient without a clinician between them. This is a scope boundary rather than a caution: we have no checker for a clinical judgement, and the pattern we refuse everywhere else applies most strongly here.
06 · Where this sector entersCross-reference to the capability catalogue in section 03 and the horizons in section 05
Fastest entry · H0
C-01 reconciliation on one month of remittance at line level. Ten working days, inside your boundary, and it either finds a recoverable position or it does not.
Pilot · H1
C-04 and C-06 on one service line with one payer. Six to eight weeks against 500 historical requests, measured per criterion.
The long part · H2 to H3
Criterion coverage across payers and service lines, and the structuring of the free-text evidence that avoidable denials turn on. Quarters, and it is a data problem before it is a model problem.
IND-10 · Public sector, infrastructure and programme delivery
The programme is fourteen months in, and the assurance review asks for a baseline nobody can reconstruct.
Large public programmes are governed by documents that are individually correct and collectively unreconciled. The business case sets a baseline, the contract sets terms, the change register records what moved, the payment applications claim against measured works, and the assurance review asks how all four relate. Today that relation is reconstructed by a person with a spreadsheet before every gate. It is a reconciliation and obligation-tracing problem, not an analytics problem, and it is one of the few sectors where the auditability of the method matters as much as the answer — which suits a stack built around an evidence ledger.
The barrier
Accountability is personal and durable here. A senior responsible owner can be asked, years later, why a decision was taken and on what basis, and the answer has to come from the record rather than from memory. But the record is a document store, not a state: the baseline exists as a PDF, the changes exist as approved forms, and the current position exists only in whatever a person last assembled. A system that produces a number without producing the trail behind it makes the accountability problem worse, not better.
Ontology objects
ProgrammeBusiness caseBaselineChangeMilestoneContractPayment applicationMeasured workAssurance findingConsent obligationConsultation responseAsset record
01 · The work as it standsObserved shape and observed cost · ranges from programmes of this size, not any named authority
How it runs todayA gate review is scheduled. Over the preceding six weeks a programme office assembles the current position: the approved baseline, every change since, the contract position, the milestone status as reported by each supplier, and the responses to the previous review's findings. Most of that assembly is copying between documents. The review then asks a question the assembly did not anticipate, and someone works late.
Cost todayTypical observed ranges rather than a benchmark. Preparing a major gate review absorbs 200 to 600 person-hours across the programme office and the delivery partners. Assessing a monthly payment application against the contract and measured works runs 4 to 20 hours per application per package. Classifying a consultation of 5,000 responses by theme takes a team several weeks and is the step most often reduced in scope when the timetable slips.
What the pilot takesSix to eight weeks, scoped to one programme and one of the workflows below. We ask for the approved baseline, the change register, twelve months of payment applications with their assessments, and the last two assurance reports. Read-only. Public-sector procurement usually sets the schedule rather than the engineering, and we will say so in the first meeting rather than the fourth.
What changesReplayed against positions your programme office has already assembled. First, whether the reconstructed current position matches the one your team produced, with every difference examined individually. Second, on payment applications, agreement with the assessor's own determination and the value of the differences in both directions. Third, hours to assemble a gate pack, timed. The finding we did not expect on this shape of work: the largest single source of difference was not arithmetic but approved changes that had never been applied to the baseline document, which is a governance finding rather than a technology one.
Stays manualThe award decision, the payment certification, the determination of a consent or an application, the assurance rating, and any decision affecting an identifiable person's entitlement. The system reconciles, traces and drafts. A public official decides, and the ledger records that they did.
02 · Instrumented workflowsCut by the unit of work that enters the graph: one tender, one application, one gate, one consent
PS01
Tender documentation conformance
Tender documents are checked against the applicable procurement rules and the authority's own standing orders, clause by clause, before publication. Divergences are raised with the rule reference, which is considerably cheaper before issue than during a challenge.
Conformance register per tender: document section, rule reference, divergence, and the standing order that governs it
PS02
Payment application assessment against contract and measured works
An application is assessed line by line against the contract terms, the rates, the measured works and the change register, with each difference classified as a rate issue, a measurement issue, or an unapproved change. The assessor receives a prepared position rather than a stack of paper.
Assessment pack: line-level position, each difference classified and referenced to the contract clause and the measurement record
PS03
Assurance evidence pack assembled continuously
The current position against the approved baseline is maintained as a state rather than reconstructed before each gate: every change applied, every milestone as reported and as evidenced, and every open finding from the previous review with its response.
Gate pack with a baseline-to-current trace, each movement citing the approval that authorised it and the date it was applied
PS04
Consultation response classification and thematic register
Responses are classified against a coding frame agreed in advance, with every classification traceable to the text that produced it and a sample re-checked by a person. The frame is fixed before the responses are read, which is the part that makes the result defensible.
Thematic register with response counts per theme, quoted supporting text, and a human-checked sample with its agreement rate published
PS05
Consent and permit obligation register
Planning consents, environmental permits and their conditions are read into an obligation register with the discharging evidence, the responsible party and the date each condition falls due. Conditions with no evidence are named rather than assumed satisfied.
Obligation register per consent: condition text, owner, due date, discharging evidence, and an explicit outstanding list
PS06
Asset record reconciliation, as-built against the register
The as-built record from delivery is reconciled against the asset register that operations will inherit, with each difference classified. Doing this at handover rather than five years later is the difference between a correction and an investigation.
Handover reconciliation register: asset, as-built record, register entry, difference class, and the party responsible for closing it
PS02 · measure
Replay twelve months of assessed applications. Agreement with the assessor's determination at line level, with the value of differences reported in both directions rather than netted.
PS03 · measure
Reconstruct the position at the last two gate reviews from the record alone, then compare against the packs your office produced. Every difference is examined and attributed to a cause.
PS04 · measure
Human agreement on a stratified sample of at least 400 responses, published with the register rather than held back. A thematic register without a stated agreement rate should not be relied on.
03 · Where this sector goesHorizon 18-36 months · our view, stated to be argued with
P01
Assurance functions ask for the method before the answer.
Where a classification or a reconciliation was produced with automated help, the reviewable artefact becomes the method and its agreement rate rather than the output. Programmes that cannot produce that will be asked to redo the work by hand.
Publish the agreement rate alongside any automated classification, from the first use
Bidders using automated drafting and authorities using automated evaluation raise the same fairness question from opposite ends. Expect standing orders and evaluation guidance to address it before central policy does.
Agree with your legal function now what disclosure you would make if asked
P03
Baseline drift becomes measurable, and that is uncomfortable.
Once approved changes can be applied to a baseline automatically, the gap between the approved position and the document everyone works from becomes a number. On the programmes we have looked at, that number is larger than anyone expects and the first reaction is rarely enthusiasm.
Measure the drift on a closed programme first, where nobody's current position is at stake
P04
Handover reconciliation moves from the end of delivery into the middle of it.
Reconciling as-built against the asset register at handover is too late to fix anything. When the reconciliation becomes cheap enough to run monthly, it moves upstream, and the operator inherits a register that matches the asset.
Run one reconciliation now on a package already delivered, and price the corrections
Where we disagree with the sector
Public programmes are frequently sold dashboards. We think the reporting layer is not the problem: the numbers on the dashboard are assembled by the same manual reconciliation that was always the bottleneck, and a faster rendering of an unreconciled position is worse than a slow one, because it looks authoritative. We would rather build the reconciliation and leave the reporting to whatever tool you already have.
04 · Node graph, re-cut for this sectorAgainst the eleven-stage reference
Driving axisAuditability of the method. In this sector the evidence ledger is not supporting infrastructure for the product — it is the product, and the reconciliation is what generates it.
Stage 10 dominates. Every reconciliation step, every classification and every applied change is written to the ledger with the document version and the approval that authorised it, so that a question asked three years later is answered from the record. Stage 04 runs with three branches: the governing documents, the contract and commercial record, and the delivery reporting. Stage 05 is minimal. A distinctive node sits ahead of any classification work: the coding frame is fixed and recorded before responses are read, and a run that classifies against a frame changed mid-way fails its own precondition.
Frame before data
Coding frames and assessment rules are fixed and recorded before the material is read. This is what makes a classification defensible and it is enforced, not advised.
Agreement rate is published with the output
Every automated classification ships with a human-checked sample and its agreement rate attached. An output without one is not released.
Approvals carry into the ledger
An applied change records the approval reference that authorised it. A change without a traceable approval is listed as unapproved rather than applied.
No determination node exists
Nothing in the graph awards a contract, certifies a payment, or determines a consent. The reconciliation prepares; an official decides and the ledger records who.
05 · Correctness and authorityWhat decides, and what we refuse
Checkers
Payment arithmetic is recomputed against the contract rates and the measured record. Every baseline movement must cite an approval reference. Classifications are validated against a human-checked stratified sample with the agreement rate published. Obligation dates are recomputed from the condition text.
Gate policy
No award, no payment certification, no consent determination, no public communication and no decision affecting an identifiable person. Read-only against the contract, finance and delivery systems, provisioned to named individuals.
Acceptance
Reconstruct the last two gate positions from the record and reconcile against the packs your office produced, difference by difference. Separately, replay twelve months of payment assessments at line level. Your assurance function runs both, and a difference the graph cannot explain counts against it.
Will not automate
We will not evaluate a tender, score a bidder, or produce anything that influences an award decision, because the fairness question has no checker and a challenge would be well founded. We also decline any workflow that decides an individual citizen's entitlement, benefit or status. Reconciliation and evidence assembly, yes. Determination about a person, no, under any framing.
06 · Where this sector entersCross-reference to the capability catalogue in section 03 and the horizons in section 05
Fastest entry · H0
C-01 reconciliation of the approved baseline against the working document on a closed programme. Ten working days, no live position at stake, and the drift number is the deliverable.
Pilot · H1
C-01 on payment applications or C-04 on obligation registers. Six to eight weeks; procurement, not engineering, usually sets the start date.
The long part · H2 to H3
The programme position maintained continuously as a state, with the ledger behind it. This is a multi-quarter build and it is the only version that removes the six-week gate scramble.
IND-11 · Education and workforce capability
Marking records the wrong answer; almost nothing records which step went wrong.
We build the assessment and rehearsal half of a teaching operation: a node graph that reads the submitted work product, re-performs the steps it is allowed to check, names the specific misconception rather than the lost mark, and writes a capability record a learner can carry to another institution or employer. The reference production system runs 155 nodes — 114 reasoning nodes and 41 deterministic workers across 11 stages, with 12 critic nodes and 6 reality gates — and for this sector we re-cut its 7 parallel evidence branches by subject domain, so each branch carries the checker that decides correctness in its own topic. A claim about a person's capability has to be traceable to the artefact behind it, which is why the terminal gate node AI-116 fails closed: an assessment that cannot be evidenced is held rather than issued.
The barrier
The misconception that produced a wrong answer usually sits three or four lines above the answer, inside working that an answer key never inspects. Diagnosing it needs a checker that can re-perform the learner's own steps — a computer algebra system for an algebra error, a compiler and unit tests for a code submission — and that coverage is uneven across a curriculum. Mathematics, statistics, programming, accounting and stoichiometry have checkers that decide correctness without asking a model's opinion; a history essay or a design critique does not, so the real constraint is subject by subject rather than sector-wide.
Ontology objects
Learner work productMisconception patternCurriculum objectiveDomain checkerRehearsal scenarioAssessment decisionCapability recordInstructor intervention
01 · EconomicsWhat the work costs as it stands
01 · The work as it standsObserved shape and observed cost · ranges from operations of this size, not any named client
How it runs todayA first-year engineering mathematics module takes 240 submissions at the Friday 16:00 deadline in Canvas or Moodle. Six graduate teaching assistants mark from Sunday morning onward against a rubric held in a shared spreadsheet, record marks in Gradescope or an equivalent, and the module leader moderates a ten per cent sample on the Monday evening before marks are released. What reaches the student record system is a number; the working that showed which step failed is archived as a PDF that nobody queries again. The corporate version has the same shape: an L&D manager reports course completion out of Workday Learning or Cornerstone to a sponsor each quarter, with no record of what any participant can now do that they could not do before. Certification bodies carry the extra load of an item bank, a standard-setting panel and an appeals process that must show, per candidate, on what basis a decision was made.
Cost todayTypical observed ranges, drawn from our own engagements and published sector reporting rather than a controlled study: 6–9 minutes of marker time per script on a multi-step quantitative assignment, so 240 scripts costs 24–36 marker-hours, and eight assignments a term puts a single module between 190 and 290 hours of marking. Blind double-marking of the same scripts disagrees by at least one band on roughly 8–15 per cent of them, and reconciling those disagreements adds another 2–4 hours per assignment. Feedback commonly returns in 10–15 working days, by which point the cohort has moved two topics on and most of the diagnostic value has expired. On the corporate side, completion rates of 80–95 per cent are routinely reported while capability is retested in a minority of programmes; where a retest is run 3–6 months later, pass rates commonly fall by 20–40 points. That last figure rests on thin and self-selected evidence, and we would not build a business case on it before measuring the client's own cohort.
What the pilot takesSix to eight weeks, scoped to one subject domain and one assessment type — for example two terms of a single quantitative module, 300 to 1,000 anonymised scripts. From the client: one subject expert for about one day a week in weeks 1–3 to write the misconception taxonomy and sign off which checker rules are permitted to decide correctness; the module leader or L&D lead for 2–4 hours a week throughout; roughly one day of IT time to produce a read-only export of past submissions, rubrics and marks; and three hours from the exams officer or awarding-body compliance lead to state in writing what the academic regulations forbid us to automate. We supply the node graph, the domain checkers, the rehearsal environment and the marking-comparison protocol. Nothing is written into the student record system during a pilot, and no live cohort is graded by the system.
What changesThe pilot reports measurables per topic rather than one headline accuracy figure. First, agreement between the graph's diagnosis and two independent human markers on a held-out set of 150–200 scripts. Second, the share of scripts the graph declines to diagnose and routes to a human, which is the number that exposes where checker coverage is weak. Third, marker-hours per assignment before and after, with feedback turnaround measured in working days. Ninety per cent agreement at a 45 per cent abstention rate means something quite different from 82 per cent agreement at 8 per cent abstention, so we publish both and let the module leader see which topics are genuinely covered.
Stays manualThe awarding decision stays with the named examiner. University regulations and awarding-body rules require a human owner who can be questioned about a progression, pass or fail decision, and a script whose working the checkers cannot verify is held at the gate instead of being graded — AI-116 fails closed, so an unverifiable case appears as a queue item rather than a mark. Academic misconduct allegations stay human for the same reason. Open-ended written argument in the humanities stays human because no checker decides it, and we would rather say that than sell a grader we cannot evidence. Capability records remain partly manual today as well: the rehearsal environments run at persistence level L2, meaning a learner's state and results persist across sessions inside our environment, and the signed record is exported for a registrar to load into the student record system by hand. Writing directly into an institution's system of record is level L3, which is gated and only partially built.
02 · Instrumented workflowsCut by Cut by the object the graph reads as evidence. Each workflow reads one class of artefact and nothing else: the learner's working (ED01), the learner's actions inside a simulated environment (ED02), the submitted assessment script (ED03), the assessment item and its syllabus mapping (ED04), the accumulated verified evidence records (ED05), and the instructor's own preparation and moderation material (ED06). Nothing in this sector that we automate reads an object outside those six; recruitment, admissions and pastoral records are outside the cut on purpose.
ED01
Misconception diagnosis from learner working
The graph reads the working itself — the algebra lines, the code commit history, the lab notebook — and labels the specific error against the department's own misconception list rather than scoring the final answer. Where the working is too thin to support a label, the node returns "insufficient evidence" and the script goes to a tutor; on early cohorts we expect that on a substantial minority of scripts, and we report the abstention rate rather than hiding it. Labels are aggregated per teaching group so that a module leader sees which of the term's known misconceptions actually appeared this year.
Cohort Misconception Report, issued per teaching group each week, in which every label cites the line of the learner's working it was drawn from
ED02
Rehearsal environments with replayable traces
A learner takes a role inside a rule-governed environment — an audit senior working a fictional client file, a ward pharmacist facing a dosing decision — and every action is written to a trace that can be replayed afterwards. Production runs at persistence level L2, which means the environment keeps state between sessions but writes to no live system; L3, where the environment connects to real records, is gated and only partially built, so we do not sell it. Judgement of the learner comes from the action trace, not from a quiz taken after the session.
Rehearsal Transcript and Competence Debrief, one per learner per scenario, with the replayable action trace attached as a file the institution keeps
ED03
Marking under symbolic and domain checkers
Every mark a reasoning node proposes is re-tested by a checker that does not use the model: a computer algebra system for algebraic and numerical claims, a unit-test runner for programming submissions, and an encoded rule deck — the standard or regulation written out as machine-readable rules — for professional subjects such as financial reporting or dosage calculation. Marks the checkers cannot confirm are not released; they route to a named examiner with the disputed rubric point already isolated. The arithmetic a head of department can check: if 30% of a paper's marks fall outside checker coverage, 30% of scripts still need human attention, and we say so in the coverage map before the pilot starts.
Marked Script Pack containing the script, the mark, the rubric point applied, and the checker output supporting each mark
ED04
Item bank and checker-coverage audit
Before any assessment runs, each item in the bank is tested for whether a machine can verify its answer at all, and the result is written down per syllabus topic as a percentage. An item whose answer depends on interpretive judgement is marked uncoverable and stays with human markers; we do not paper over it with a model-scored rubric. This audit is what tells a registrar which parts of a programme are cheap to run and which are not.
Checker Coverage Map for the syllabus, topic by topic, with a worked example of a covered and an uncoverable item in each topic
ED05
Portable capability transfer records
Evidence already verified in ED01 to ED03 is assembled into a record the learner takes with them, stating what they demonstrated, on what date, under which checker, with the underlying artefact retrievable. The honest position on adoption: very few employers consume such a record today, so we issue it as a signed PDF for a human reader and a JSON file for whichever system eventually reads it, and we make no claim that it shortens a hiring process yet. Professional bodies awarding continuing-development credit have been the readiest buyers so far.
Capability Transfer Record, issued to the learner and to the awarding institution, each claim linked to the stored artefact behind it
ED06
Instructor preparation and moderation load
The graph assembles the weekly teaching pack, drafts the moderation paperwork for borderline scripts, and prepares the appeals bundle, leaving the tutor to decide rather than to compile. We baseline the load from four weeks of timesheets before any node runs; on a worked example of 18 tutors at 4.5 hours a week each on marking administration, the baseline is 81 hours a week, and the pilot target is 45, which is a target and not a result. What stays manual is every judgement call: which borderline script moves up a band, and what a student is told in a difficult tutorial.
Weekly Teaching Pack and Moderation Dossier, one per module, listing every borderline case with the evidence already gathered against it
ED01 · measure
Two subject leads hand-label 200 scripts blind. We report precision against their labels with abstentions counted as misses; the contract target is 90%.
ED02 · measure
Re-run 100 stored traces against the same rule set. Each debrief must reproduce identically; any trace that does not reproduce is a defect we correct at our own cost.
ED03 · measure
Blind double-marking against an examiner panel on 500 scripts: agreement within one rubric band, and a count of symbolic errors that must be zero.
03 · Where this sector goesHorizon 18-36 months · our view, stated to be argued with
P01
Awarding bodies will require a machine-readable marking trace before accepting AI-assisted marks.
Appeals are the pressure point. An examinations officer defending a grade at appeal needs to show which rubric point was applied and on what evidence; a model output with no trace cannot be defended, and one lost appeal is enough for a regulator to write the requirement into its conditions of recognition. This is a projection, and the current evidence is a handful of published regulator consultations rather than settled rules.
When buying, ask to see a single marked script exported with its full checker trace before signing anything. If the vendor can only show accuracy figures, the system will not survive an appeal.
P02
Checker coverage per subject, not model quality, will decide which subjects get automated first.
Mathematics, statistics, programming and rule-bound professional subjects have machine checkers that are right for reasons independent of the model. Interpretive essay marking has almost none, so improvements in model quality there cannot be verified and therefore cannot be trusted at scale.
Commission a coverage audit of your own syllabus before budgeting. Fund automation in the covered topics and fund human marker capacity in the rest, rather than spreading a single per-student licence across both.
P03
Employers will accept portable capability records in code and accountancy before universities do.
Employers can test the claim cheaply — a code record is checkable by running the code, a bookkeeping record by re-performing the entries. Universities must fit any external record into credit frameworks and validation cycles, which move on a multi-year clock.
If you are a certification body, design the record for an employer's verification step first. Credit recognition can be retro-fitted; employer acceptance cannot be back-dated.
P04
Invigilated performance tasks will displace take-home essays across most assessed credit.
Authorship of unsupervised written work cannot be established by any method we would defend in front of an appeals panel. Institutions will move the weight of assessment to conditions where authorship is observed, which raises invigilation and staffing cost rather than lowering it.
Model the invigilation cost now — room hours, staff hours, accessibility provision — and treat any efficiency claim that ignores it as incomplete.
Where we disagree with the sector
The common view is that personalisation is the prize: give every student an adaptive tutor and attainment rises. We think that is the wrong constraint. The binding constraints in most institutions are assessment credibility and instructor time, and neither is relieved by a tutor that produces more unverified output. The published evidence for attainment gains from adaptive tutoring at institutional scale is mostly short-run, small-sample and measured on the same instrument that was practised, which is a weak design. Our own position is contestable too, and the way to settle it is a controlled comparison in one department: adaptive tutoring on one cohort against verified marking plus rehearsal on a matched cohort, judged on an assessment neither cohort practised.
04 · Node graph, re-cut for this sectorAgainst the eleven-stage reference
Driving axisChecker coverage per subject domain — the share of a subject's claims that can be tested by something other than the model. In algebra, statistics and programming, coverage is high, verification is cheap, and cost per script falls with volume. In interpretive humanities, coverage approaches zero, every mark needs human confirmation, and cost per script barely falls at all. Coverage also carries the risk: an unverifiable mark released to a student is the failure that ends a contract, so the graph is shaped to hold those marks rather than to speed them up.
The eleven-stage reference topology is re-cut so that branch width tracks coverage. Stage 04, which in the reference system runs seven parallel evidence branches, becomes seven checker branches in high-coverage subjects: computer algebra, unit and dimension consistency, unit-test execution, rule-deck conformance, citation resolution, prior-work consistency for the same learner, and rubric-precedent lookup across the cohort. In a low-coverage subject those seven collapse to three — citation resolution, prior-work consistency, rubric precedent — and the freed budget is spent on critic nodes, so the essay path carries more of the 12 critics than the mathematics path does. Stage 05, the 27-node poster pipeline, collapses to three nodes producing the learner feedback sheet, and stage 07 runs only for ED02, where it renders the rehearsal replay. Stage 09, cross-product consistency review, becomes cross-cohort moderation and is a mandatory synchronising barrier: no grade leaves the system until every script in the cohort has been marked, because a threshold applied to script 1 cannot be checked against script 400 if script 400 has not been seen. Concurrency is safe within a cohort for per-script marking and for the ED01 diagnosis path, and unsafe for anything within one rubric band of a grade boundary, which waits at the barrier. Stage 10 produces the appeals evidence pack rather than a publication audit, and stage 11 remains the terminal gate at node AI-116, which fails closed on any grade or certificate.
Coverage sets branch width
Seven parallel checker branches run at stage 04 in mathematics, statistics and programming; three run in interpretive subjects, with the saved compute moved into critic nodes. The coverage map from ED04 is the input that decides which shape a module gets.
Barrier before stage 09
Cohort moderation is a hard synchronising barrier. Nothing is released to a student, and no borderline decision is taken, until every script in the cohort has passed marking, because grade-boundary consistency cannot be assessed on a partial set.
Abstention is an output
A node that cannot support a label or a mark returns "insufficient evidence" and routes to a named tutor. The abstention rate is reported weekly alongside accuracy; a system with a suspiciously low abstention rate is guessing.
AI-116 holds certificates
The terminal gate fails closed on grades, transcripts and capability records. Release requires the signature of a named examiner or registrar who is recorded in the audit pack at stage 10.
05 · Correctness and authorityWhat decides, and what we refuse
Checkers
Five checker classes run in this sector, each bound to a class of claim. A computer algebra system checks algebraic and numerical claims, including unit and dimension consistency, and it decides correctness independently of the marking model. A unit-test and static-analysis runner checks programming submissions by executing them against the module's own test suite in a sandbox, so a submission that compiles but fails a hidden case cannot be marked correct. A rule deck — the professional standard written out as machine-readable rules, such as a financial reporting treatment or a paediatric dosage table — checks conformance claims in vocational and professional subjects, and each rule cites the clause it encodes. A citation checker tests every referenced source in written work: the source must resolve to a real document and any quoted string must match the source text, which catches invented references but says nothing about the quality of the argument. A rubric-consistency checker is statistical rather than symbolic: it compares the threshold at which each rubric point was awarded across the whole cohort and flags divergence for the moderation barrier at stage 09. For rehearsal work in ED02 a sixth checker replays the stored action trace against the environment's rule set; a debrief that does not reproduce on replay is treated as a fault, not as a judgement.
Gate policy
Four refusals, fixed in the deployment and not configurable by the client. No grade, transcript entry, qualification or capability record is issued or altered without the recorded signature of a named human examiner or registrar; AI-116 fails closed and there is no override flag. No agent originates or substantiates an academic misconduct allegation, because authorship cannot be established to a standard we would defend. No agent writes to a student record of record, a progression decision, an exclusion decision or an admissions decision; it can prepare the paperwork and nothing else. No agent infers or stores an assessment of a learner's disability, mental health, immigration status or any other protected characteristic, and any such inference appearing in a node output is discarded before the output leaves the graph.
Acceptance
One department, eight weeks, at your own cost baseline. We mark 500 scripts blind against your examiner panel and must agree within one rubric band on 97% of them, with zero symbolic errors — a symbolic error being a mark the computer algebra system or unit-test runner contradicts. On ED01, precision against 200 hand-labelled scripts must reach 90%, with abstentions counted as misses. Every released mark must carry a checker trace that an appeals officer can read without our help, tested on 20 scripts they pick. If any of those three figures is missed, the pilot fee is returned and the deployment stops.
Will not automate
We cannot verify that a submission is the learner's own unaided work, so we refuse to automate any part of academic misconduct. Detection tools are calibrated on the wrong quantity: at a 1% false-positive rate across 40,000 submissions a year, 400 students are wrongly flagged, and no institution can absorb that. The second thing we do not automate is the quality judgement in interpretive writing — whether an argument is well made. Our checkers can confirm that a source exists and that a quotation matches, and they cannot tell a good argument from a fluent one; that mark stays with a human, and any coverage figure we give you for an essay-based module reflects that.
06 · Where this sector entersCross-reference to the capability catalogue in section 03 and the horizons in section 05
Fastest entry · H0
C-01 and C-02 replayed on one term of already-marked submissions. Ten working days, and it returns misconception agreement per topic rather than one accuracy figure.
Pilot · H1
C-08 capability transfer on one subject domain and one assessment type. Six to eight weeks, with the examiner holding the awarding decision throughout.
The long part · H2 to H3
The misconception corpus and transfer measured after exit. Quarters, and transfer measurement is the part we are still building rather than the part we are selling.
IND-12 · Interactive media, simulation and training worlds
The demonstration held for four minutes. The session it has to support runs forty, and returns tomorrow.
This is the practice closest to our own research, and the one where we are most exposed to being wrong. An interactive world is easy to demonstrate and hard to run: characters contradict what they said twenty minutes ago, the world forgets what the player changed, and the thing the person was supposed to take away is measured by how long they stayed. The engineering that fixes it is the same engineering as everywhere else on this page — a typed state model, a checker per class of claim, and a gate before anything crosses into the real world. What differs is that the acceptance test is about a person, and it is taken after they leave.
The barrier
Continuity is a state problem that gets treated as a prompting problem. A character's commitments, a location's condition and a plot fact are objects with versions and consequences, but they are usually held in a context window and reconstructed by the model each turn. That works to about the length of a demonstration and fails in a way that feels like the world is not real, which is the one failure the medium cannot absorb. Enterprise training simulations have the same defect with higher stakes: a rehearsal that lets you skip the isolation step teaches you to skip the isolation step.
Ontology objects
ActorPersonaWorld cellScene beatCommitmentContinuity factAssetRuleSession stateCompetencyTransfer record
01 · The work as it standsObserved shape and observed cost · from our own products and from client work, stated as ranges
How it runs todayA studio or a training team builds a scenario as a script with branches. Writers author the branches by hand, so coverage is set by the writing budget rather than by what the learner or player might do. When a generative layer is added, the branch count stops being the constraint and continuity becomes it: the character remembers within a conversation and forgets across sessions, and quality assurance discovers this by playing for an hour, which does not scale.
Cost todayTypical observed ranges rather than a benchmark. Hand-authored branching content runs to several thousand words of script per fifteen minutes of interactive time, and quality assurance on a generative layer is largely manual playthrough. In corporate training, a bespoke simulation module is commonly quoted in the tens of thousands per hour of content, and the transfer it produces is usually evidenced by a completion rate and a satisfaction score, which measure neither.
What the pilot takesSix to ten weeks, scoped to one scenario and one competency. We ask for the scenario rules, the artefact a participant is meant to produce, and a way to measure the participant afterwards — that last item is the one clients most often do not have, and building it is usually the first fortnight. Deployment is ordinary: this workload does not touch a system of record, so the boundary questions are lighter than anywhere else on this page.
What changesMeasured on participants, not on sessions. First, continuity: an automated continuity harness replays long sessions and counts contradictions against the recorded world state, which is a hard number and it is usually bad at the start. Second, rule enforcement in a training world: the share of runs where a mandatory step was skipped and the world allowed it. Third, transfer, taken as an unassisted task after exit and marked by a person. On our own products the honest position is that continuity and rule enforcement are solved to a level we will defend, and transfer measurement is the part we are still building — we would rather say that than show an engagement chart.
Stays manualCertification and assessment of record stay with the named examiner. Anything that leaves the world and touches the real one — a booking, a payment, a message to a person, an order — goes through an explicit gateway with the participant confirming it outside the fiction. Creative direction stays with the writer; the system enforces continuity, it does not decide what the story is about.
02 · Instrumented workflowsCut by the unit of work that enters the graph: one session, one character, one scenario, one participant
IM01
Continuity enforcement across a long session
Facts a character has asserted, commitments made, and changes to the world are held as typed objects with versions rather than reconstructed from context. Every generated line is checked against the recorded state before it is shown, and a contradiction is repaired in place rather than regenerated from scratch.
Session with a continuity log: every asserted fact, its source beat, and every contradiction caught and repaired before display
IM02
Persistent character state and commitment tracking
A character's relationships, obligations and knowledge persist between sessions with an explicit expiry and provenance for each. When a player returns after a week, what the character remembers is a query against state, not a summary of a transcript.
Character state record per participant, with each fact carrying its origin, its version and the beat where it last changed
IM03
Scenario authoring from a rule set
A training scenario is generated from declared rules — the procedure, the hazards, the permitted and forbidden actions — rather than hand-branched. Coverage becomes a property of the rule set, and a scenario that permits a forbidden action fails its own check before it reaches anyone.
Scenario package with a rule-coverage report naming every rule exercised and every rule not reachable in this scenario
IM04
Asset production with per-frame inspection
Generated images and video pass through the same discipline as any other artefact: every rendered output is re-opened and inspected against the brief, and identity consistency across frames is checked rather than assumed. Failures are localised and re-rendered, not regenerated wholesale.
Asset set with an inspection record per frame, and a consistency report for any recurring character or location
IM05
Rehearsal with the professional tool inside the world
The participant works in the real tool — the CAD package, the trading interface, the plant control replica — inside a scenario that responds. Their actions are checked against the same domain checkers used in production work, so the feedback is about the artefact rather than about the interaction.
Rehearsal record: the artefact produced, the checker results against it, and the decision points that led there
IM06
Transfer measurement after exit
Some period after the session, the participant performs an unassisted task and a person marks it. The result is attached to the rehearsal record so that the scenario can be judged by what it produced rather than by how long anyone stayed in it.
Transfer record per participant: the unassisted task, the mark, the marker, and the rehearsal it is being attributed to
IM01 · measure
Automated replay of 200 long sessions. Contradictions per hour against the recorded world state, counted by a harness rather than by a reviewer, with the count published before and after.
IM03 · measure
On a scenario with a declared rule set: the share of runs in which a mandatory step was skipped and the world permitted it. The target is zero and anything above it is a defect, not a tuning parameter.
IM06 · measure
Unassisted task performance after exit, marked by a person, against a matched group who did not run the scenario. Small samples, reported with their size, and we say when a result is not significant.
03 · Where this sector goesHorizon 18-36 months · our view, stated to be argued with
P01
Persistence becomes the product, and generation becomes a commodity input.
Generating a plausible line of dialogue is already cheap and getting cheaper. Remembering what was said last Tuesday, and what it obliges the character to do now, is a state engineering problem that does not get cheaper with the next model. The differentiator moves to the state graph.
Ask any vendor what happens on the second session, a week later, and watch what they demonstrate
P02
Corporate training buyers start asking for transfer evidence rather than completion rates.
Completion and satisfaction are measured because they are easy, not because anyone believes them. As simulation budgets grow, procurement will ask what the participant could do afterwards, and most incumbents have no instrument for that question.
Decide what unassisted task you would use as evidence before you commission the content
P03
Rehearsal moves into the real professional tool.
A simulation of a tool teaches the simulation. Running the actual software under a scenario, with the real checkers applied to the artefact, produces evidence a hiring manager will accept. The blocker is licensing and environment cost rather than capability.
Price a rehearsal environment running your real toolchain before assuming it is out of reach
P04
Worlds that touch reality will be regulated at the crossing point, not inside.
Nobody will regulate a fictional conversation. A world that books, pays, commits or certifies on someone's behalf is a different object, and the rules will attach at that boundary. Building the gateway now costs little; retrofitting it costs the product.
Put an explicit, confirmed gateway on every crossing before you need one
Where we disagree with the sector
The prevailing metric is time spent, and we think it is the wrong one and actively harmful to the engineering. Optimising for dwell time produces worlds that withhold rather than worlds that teach, and it hides continuity failures because a confused participant stays longer. Our own products are held to transfer after exit instead, and that measure is less flattering — StudyHub's engagement numbers would read better than its transfer numbers, which is precisely why we publish the second.
04 · Node graph, re-cut for this sectorAgainst the eleven-stage reference
Driving axisLatency against state integrity. This graph runs at session pace rather than overnight, which changes the topology more than any other constraint on this page.
The eleven-stage batch shape does not survive here. Work splits into a fast path that must answer within a turn — retrieve state, generate, check against continuity, display — and a slow path running between turns and between sessions, which reconciles state, resolves contradictions, plans upcoming beats and produces assets. Stage 09, cross-product consistency, becomes continuity checking and moves from the end of the run into the middle of every turn. The critic set is small and always on rather than large and terminal.
Two paths, one state
The turn path is latency-bound and narrow. The between-turns path is where planning, asset production and reconciliation happen. Both read and write the same typed state.
Continuity checks run before display
A contradiction caught after the participant reads the line is not caught. The check is in the turn path and it costs latency, which is a trade we make deliberately.
Rules are enforced by the world, not the prompt
A forbidden action is refused by the world model. Asking a character to decline it is not enforcement, and in a training scenario it is a defect.
One gateway out
Every crossing into the real world goes through a single explicit gateway with participant confirmation outside the fiction. There is no second path, by construction.
05 · Correctness and authorityWhat decides, and what we refuse
Checkers
Continuity against recorded world state before display. Rule coverage and forbidden-action refusal in training scenarios. Identity consistency across generated frames. Artefact checkers from the relevant domain where the participant is producing real work. Transfer marked by a person after exit.
Gate policy
No irreversible real-world action originates inside a world. Bookings, payments, messages to third parties, orders and certifications cross only through a confirmed gateway, outside the fiction, with the participant seeing plainly what they are agreeing to.
Acceptance
Contradictions per hour on 200 replayed long sessions, measured by harness. Zero permitted skips of a mandatory step in a training scenario. A transfer result on an unassisted task with the sample size stated, reported even when it is not significant.
Will not automate
We will not let a world take a real payment, make a real booking, or send a message to a third party on a participant's behalf without an explicit confirmation outside the fiction, however smooth that would make the experience. We also decline persuasion and behaviour-change work aimed at a person's decisions outside the world. Our stated measure is what someone can do after they leave, and we have no way to check a claim about what they were made to want.
06 · Where this sector entersCross-reference to the capability catalogue in section 03 and the horizons in section 05
Fastest entry · H0
A continuity audit on your existing product: C-01 applied to session transcripts against declared state. Ten working days, and it returns a contradictions-per-hour number you probably do not have.
Pilot · H1
C-08 capability transfer on one scenario and one competency. Six to ten weeks, and the first fortnight is usually spent building the measurement you will be judged on.
The long part · H2 to H3
Persistent world state at L2, then a gated crossing into real processes at L3. That is the research programme in section 06, and we describe L3 as partial because it is.
Buyers ask how fast this can be real, and the honest answer has four parts rather than one. The cut below is by time to a result you can accept or reject — not time to a demonstration, and not time to a signed contract. Each horizon has a different thing at the end of it, needs a different thing from you, and fails in a different way.
The first horizon is genuinely two weeks and genuinely useful, and it is the one we give away. The fourth is measured in years and is where the value actually accumulates. Anyone who tells you the fourth arrives on the first horizon's timetable is selling a demonstration.
H0 · INSTRUMENTATIONUNPAID
10 working days
Measure the work, do not change it
Read-only, against a copy of your own history. Nothing is integrated, nothing is written, and no system of record is touched. The deliverable is a document plus one agent you can watch run.
What exists at the end
An opportunity register with a row per step in the workflow: what it costs today in hours, whether an acceptance test can be written for it, and who would own it. Plus one read-only agent replaying your back-file, so the numbers come from your data rather than from a slide.
What we need from you
Half a day on site or two hours remote, a named owner, and a back-file of 200 to 800 completed items from the last two years. Read access under an NDA, provisioned to named individuals on our side.
What it does not include
No integration, no write access, no production deployment, and no model fine-tuned on your data. If a workflow needs a system connection to be assessed, it moves to H1 and we say so rather than approximating it.
What goes wrong here
Getting the back-file. Legal review of an NDA and the provisioning of read access routinely take longer than the ten days of work they enable, and that is the schedule risk on this horizon, not the analysis.
H1 · ONE VERIFIED WORKFLOWFIXED PRICE
4–8 weeks
One workflow, one acceptance test
A graph of 20 to 40 nodes covering a single named workflow, with its checkers, its gate and an acceptance test written before we start. It runs in propose-only mode against real inputs.
What exists at the end
A working system your team runs on live inputs without our involvement, producing a proposal a named person accepts or rejects. Per-node evaluation sets, the checker suite, the gate policy in writing, and the acceptance test result whether or not it passes.
What we need from you
A named owner with authority to sign the acceptance test, roughly two to four hours a week of a domain expert's time, read access to the systems in scope, and the back-file from H0 if we did not already have it.
What it does not include
No write-back to a system of record, no auto-clear band, no second workflow, and no scale testing beyond the pilot volume. Adding a second workflow is a second pilot, not a scope change.
What goes wrong here
The acceptance test turns out to be unwritable, usually because the work is judged rather than checked. When that happens we say so in week two and stop, and it has happened. The other common failure is that the domain expert's four hours a week do not materialise.
H2 · PRODUCTION GRAPHSCOPED PHASES
1–2 quarters
Write-back, behind gates, under audit
The graph grows to 40 to 150 nodes across several workflows, begins writing into systems of record through their APIs behind approval gates, and carries an audit ledger that reproduces any run from stored inputs.
What exists at the end
Multiple workflows on one ontology, typed write-back with rollback, an approval routing model naming who signs what, the evidence ledger, monitoring, and a per-node regression suite built from accepted runs. Auto-clear bands only where months of measured acceptance justify them.
What we need from you
Integration engineering time from your side, a security review, a decision on deployment topology, and named approvers for each gated action. Change management, because this is the horizon where the work of the people around it actually changes.
What it does not include
No autonomous irreversible action, at any node, in any configuration. Payments, releases, filings, equipment commands and identity changes stay behind a human gate permanently and are not on the roadmap.
What goes wrong here
Cross-workflow contradiction. Two workflows that were each correct alone produce outputs that disagree, and the stage that catches it is the one clients most often want to cut for schedule. We will not cut it, and that argument usually happens around week nine.
H3 · ONTOLOGY AND CORPUSSTANDING
12+ months
The part that compounds, and it is slow
One ontology across the estate rather than per workflow, and a trajectory corpus of verified runs — including the instructive failures — that makes each subsequent build cheaper than the last.
What exists at the end
Consolidated objects, relations and permission classes spanning several functions. A corpus of trajectories with their failure paths and repairs, owned by you. Where authorised and where a checker exists, selected connections into real processes at persistence level L3, which is gated and partial by design.
What we need from you
A standing owner rather than a project sponsor, a budget that survives a reorganisation, and the willingness to fix identifier and relation problems in your own systems, which is where most of the calendar goes.
What it does not include
Anything we have no checker for. Section 03 lists two claim classes we tried and abandoned, and section 04 lists what each sector's practice refuses. The refusals do not shrink as the engagement grows.
What goes wrong here
Sponsorship. Twelve-month programmes outlive the person who commissioned them, and the ontology work that pays in year two is the easiest thing to defer in month seven. We have had this happen and the mitigation is unglamorous: ship something usable every quarter.
01 · What actually ships in ten working daysRead-only, against your own back-file · these are the H0 deliverables, named
Q01 · C-01
Three-way match replay
Invoice against purchase order against goods receipt, replayed over 500 historical documents including freight, duty and tax lines. Output is an exception list, not a percentage.
Thirty items you already published, re-checked claim by claim. Every figure re-located to a document, a date and a page, or listed as unsourceable.
Deliverable: claim map + failed-citation rate
Q03 · C-03
Top-three recall on solved excursions
Fifty anomalies your engineers already diagnosed, replayed blind against the archive. We report how often the true cause was in the top three, and how often it was not in the list at all.
Deliverable: recall measurement on your own history
Q04 · C-01
Baseline drift measurement
The approved baseline against the document everyone actually works from, with every applied and unapplied change traced to its approval reference. Run on a closed programme first.
Deliverable: drift register with approval trace
Q05 · C-06
Exception queue shadow run
One month of exceptions replayed in shadow. Every item gets a proposed classification and owner, compared against what your team actually did, in both directions.
Deliverable: agreement matrix + routing proposal
Q06 · C-01
Line-level remittance or rate reconciliation
One month of payments reconciled against the contract terms and the schedule version in force, with each difference classified by cause rather than aggregated.
Deliverable: variance register with cause classes
Q07 · C-04
Clearance-evidence gap
A closed review re-run to answer a question the file usually cannot: not what was found, but what was checked and cleared, and whether any reason was recorded.
Deliverable: clearance register + evidenced share
Q08 · C-08
Continuity audit on a live product
Two hundred long sessions replayed against declared world state by an automated harness. Contradictions per hour, counted rather than sampled.
Deliverable: contradictions-per-hour figure + repair list
02 · What cannot be compressedFour constraints that set the calendar, none of which are engineering
Legal and access
NDA execution and read-access provisioning to named individuals. Frequently two to six weeks, and on public-sector and regulated work it is longer than every other item combined. Start it before you decide anything else.
Validated environments
Where your environment is qualified — pharma, medical device, parts of finance — validation is your process and it sets the schedule. We design so the validatable boundary is the deterministic wrapper rather than the model, which helps, and it does not make validation fast.
Identifier and relation repair
The same object carries different identifiers in five systems and the relation you need has never been declared. This is the single most under-estimated item on every engagement, and it is work in your estate, not in ours.
Expert attention
Two to four hours a week from the person who actually knows the work. When that does not materialise the pilot does not slip — it produces a system that is confidently wrong in ways nobody caught, which is worse.
03 · Earliest horizon per capabilityAgainst the eight capabilities in section 03 · what sets the floor in each case
Capability
Earliest useful resultRead-only, on your back-file
Pilot with an acceptance testPropose-only, live inputs
What sets the floorThe binding constraint, not the model
C-01Reconciliation
10 working days
4 to 6 weeks
Getting the back-file, and the state of your identifier mapping between the two populations.
C-02Attribution
10 working days
6 to 8 weeks
Whether your source hierarchy is written down. Where it is not, defining it is the first fortnight.
C-03Cause ranking
10 working days, scored only
6 weeks
Depth and continuity of the history. Two years of clean trace data is worth more than any model choice.
C-04Rule-bound drafting
15 working days
6 to 8 weeks
Reading the enforceable text into typed obligations. Ambiguity in your own playbook surfaces here and has to be resolved by you.
C-05Tool operation
Not available at H0
8 weeks
Licensing, environment and the tool's own API surface. A tool without a scriptable interface is a different and longer conversation.
C-06Exception triage
15 working days
4 to 8 weeks
Whether the exception taxonomy exists. Most queues are classified by habit rather than by a written rule.
C-07Multi-artefact production
Not available at H0
One quarter
Cross-product consistency. This capability is a graph rather than a pipeline and does not have a two-week version.
C-08Capability transfer
10 working days, audit only
6 to 10 weeks
Having a measurement of transfer at all. Most clients do not, and building it is the first fortnight of the pilot.
H0 is free and it is the whole of what is free. We do not recover its cost later in a build fee, and roughly one in three of the H0 engagements we have run ended with us recommending that the client not proceed — usually because the workflow could not be given an acceptance test, occasionally because an existing tool already covered it. The register is yours either way, and you can take it to another supplier.
Every engagement we run is also an experiment in one question: under what conditions does a digital experience leave a durable change in the person who had it? We call the answer Enactive Reality, and it is the reason this firm exists rather than a consultancy that happens to use models.
The term is borrowed deliberately. Enactivism holds that cognition arises through action in an environment, not through representation of one. The engineering consequence is specific: an environment that cannot be acted on, and cannot answer with consequences, cannot teach anything.
Formal definition
Enactive Reality is a medium in which a person enters an AI-driven world, takes a role, acts under rules that answer back, and returns with knowledge, capability, relationships or results that persist outside it.
Virtual reality asks
How does a person enter a digital space? The product unit is immersive display. The experience usually ends with the device.
The metaverse asks
Where is a person persistently online? The product unit is presence and assets. Most of the value stays inside the platform.
Enactive Reality asks
What did a person do, and who are they once they leave? The product unit is a transferable episode. Success is measured after exit.
BoundaryFive concepts that are routinely conflated
Concept
The question it answers
Core product unit
After the session ends
Position in the stack
VR
How does a person enter a digital space?
Immersive display and interaction
The experience typically stops with the hardware
A carrier
Metaverse
Where is a person persistently online?
Space, identity, social graph, assets
Value largely remains on the platform
A social narrative and product form
Generative Reality
How does AI continuously produce and revise a world?
World state, simulation, generation, execution
Technically continuous, but agnostic to the person
The underlying technical paradigm
Enactive Reality
What did the person do, and who are they afterwards?
One transferable episode
Knowledge, capability, relationships or results are carried out
The human experience paradigm — where value is judged
Reality Compiler
How does intent become trustworthy software and real action?
Tasks, typed tool calls, verification, rollback
Real-world state feeds back into the next step
The execution infrastructure beneath all of it
The relation is simple to state and hard to build. Generative Reality produces the world. Enactive Reality changes the person. The Reality Compiler connects both to consequences that hold outside the session. We are not building a larger game; we are building the execution layer between human intent, professional tools, simulated worlds and real services.
LayersProgramme layers ER-01 to ER-06 — a separate set from the platform stack in section 01
ER-01
Experience
A person enters, takes a role, acts, and bears the consequences. Entry is a design problem; it is not the achievement.
ER-02
World model
Objects, characters, rules, time, space, causality, state and an event log. Feedback must come from readable rules, never from improvisation.
ER-03
Agent
Planning, role, dialogue, intent parsing, tool proposal, conflict and collaboration between multiple acting entities.
ER-04
Professional tools
HTML and Canvas, Python and symbolic mathematics, Blender, Unreal, CAD, business APIs. The world's credibility is borrowed from software people already trust.
ER-05
Reality
Courses, appointments, stores, manufacturing, payment, and the humans who approve and respond. The boundary where consequences become real.
ER-06
Carryover
Knowledge, capability, emotion, relationships, documents, orders and decisions that go back with the person. Without this layer the other five are entertainment.
Entry
Role, space, task, story. Necessary, and routinely mistaken for the achievement.
Consequence ↓
Each layer down adds transferable value and demands more structured state, permission and verification. Most projects stop at ER-02 and call it a product.
StateFour levels of world persistence
L0 · STATIC
Fixed content
Video, articles, fixed levels. Value is concentrated in a single act of consumption.
Commercial unit · impressions
L1 · REACTIVE
Immediate response
Characters answer, scenes react, and the state returns to zero shortly after the session ends.
Commercial unit · engagement
L2 · PERSISTENT
Durable world state
Relationships, resources, tasks and consequences are remembered. The world keeps evolving; a person returns to something that moved without them.
Where our production systems operate today
L3 · REALITY
Connected consequence
Appointments, orders, coursework, manufacturing and services cross safely into real processes, behind authorisation and human confirmation.
Gated · partial deployment
The unit of value changes
From what I watched to what I changed. People pay for continuity of state, credible consequence and transfer back into their own life or work.
The retention mechanism changes
From waiting for the next episode to returning because the world left something unresolved. Relationships, resources and commitments create the reason.
The defensibility changes
From volume of model calls to the state graph and the verified trajectory record. Reusable world rules and real outcomes become long-lived assets.
The measure of an Enactive Reality system is not how long someone stayed. It is what they took out with them, and whether it held.
07Deployments
Systems that survived contact
Client systems and our own products are shown separately, because they answer different questions. Client work proves the stack holds under someone else's data, permissions, exceptions and deadlines. Our own products give us an environment we control end to end, where we can push the research further than a client would reasonably fund.
None of these are demonstrations. Each one has users who notice when it is wrong. They are also not the full picture: they are the deployments we are permitted to name. Several of the workflows described in section 04 are scoped or in build under agreements that do not allow us to say whose they are, and we would rather leave a gap on this page than describe a client obliquely enough to be identified.
Jin10 Data24/7 global financial information and research platform
Client system · the graph in section 02, deployed
Who relies on it, what it replaced, and what it got wrong
Jin10 had the raw material problem every information business has: reports and newswire fragmented across topics and formats, research views that could not be retrieved as views, institutional disagreement that could not be compared, sources that could not be located, and long research sessions that broke halfway and had to be restarted.
The system we built ingests commodity, industry, macro, company and asset-class research — native text and scanned files alike — and organises views, facts, figures and industrial relationships into research products. It does not only answer what a report said. It compares consensus and dispersion across reports, tracks how a view moved over time, connects supply-chain context, and produces investor-education material with page-level source locations. Full originals stay inside Jin10's existing membership, purchase and institutional entitlement system; the derivative layer never leaks the thing it was derived from.
A single research request became a process that can be resumed rather than restarted. And when evidence is insufficient, the system says so instead of writing a confident conclusion — which is the hardest behaviour to engineer and the only one that matters to a professional reader.
Who relies on it. Research staff, editors, the investor-education team and institutional users read the same content assets through different entitlement classes. The system does not decide who may see what; the client's existing membership, purchase and institutional permission model does, and the derivative layer inherits it rather than re-implementing it.
What it replaced. A research request used to be a person, a search box and a folder of PDFs, restarted from the beginning whenever a long session broke. It is now a process that resumes. The measurable that mattered to the client was not speed: it was the share of published claims that resolve to a source, a date and a page.
What it got wrong. The first version had no cross-product consistency stage. Four artefacts generated from one evidence pack could each be internally correct and disagree with one another, and they did. Stage 09 exists because of that, and it was added after launch, not designed in. The second correction was the snapshot rule: early nodes were allowed to re-fetch, which cost us the ability to replay a conclusion. Both fixes cost more than building the stage correctly the first time would have.
1 ledgerEvery published claim resolvable to source, date, page and a reproducible dataset
Education
Zhihu MathHubLearning product inside a knowledge community
Diagnosis, task planning, interactive explanation, grading and learner state existed as five disconnected features. We joined them into one loop, so a learner moves between explanation, figure, video, practice and feedback without the system losing track of what they are actually struggling with. Mathematical claims are checked symbolically before they reach a student.
Financial information
Jin10 Research InsightResearch retrieval and entitlement-aware derivative content
Reports, newswire and public data turned into searchable, comparable answers with page-level source location — and a controlled derivative layer for the content team that respects the entitlement boundaries of the material it was built from.
Own productsWhat we run ourselves · the research claim is in section 06
StudyHub
A question becomes a path that keeps moving
StudyHub organises questions, materials, courses and knowledge into a continuous learning experience. A learner starts a task, enters guided study, works through dynamic lessons and generated video, and completes interactive exercises that are graded by domain checkers rather than string comparison.
The product is not built to get one answer right. It tracks what is being studied, where the difficulty actually sits, what hint comes next, and how to move between explanation, practice, feedback and review without a reset.
Persistent characters, and the world that remembers them
Future Film Lab works on generative image and video, character agents, interactive narrative and virtual worlds. With original IP and short-form work it tests identity consistency across frames and sessions, AI-directed production, and interactive story workflows that hold together over repeated contact.
The long-run problems — persistent characters, story agents, durable environments — are the same problems the world runtime solves for industrial rehearsal, approached from the side of narrative. It is cheaper to discover a continuity failure in a short film than in a training simulator a regulator has approved.
Identity latent — consistency across frames02
08The firm
Small by design, deep by necessity
Tunneling Technologies is a research-led engineering firm in Singapore, founded by doctoral researchers trained in China and the United States, working across trustworthy AI, causal identification, operations research, control theory, multimodal inference and the learning sciences.
We stay small on purpose. The work we do well requires the people who designed a system to be the people who sit with the client while it fails. That does not survive being staffed out, so we do not staff it out. It also means we take on a limited number of engagements and say no to most of the rest.
Method
How an engagement actually runs
STEP 01 / 04
01Diagnostic
Observe the work before deciding what should be automated
We observe the work before specifying anything, then return an opportunity register naming what can be verified and what cannot. This is the free diagnostic; its full terms, duration and deliverable are set out once, in section 09, rather than restated here.
02Pilot
One scenario, four to six weeks, a result you can reject
No large first contract. One workflow, with scope, acceptance criteria, data boundary and price fixed in advance. If it passes, we extend. If it fails, the failure is contained inside the pilot and you owe nothing further. We would rather lose a contract than defend a system that does not work.
03Deployment
Connect to what exists rather than replacing it
The stack meets your systems of record where they are — ERP, CRM, MES, PLM, LIMS, DMS, spreadsheets, the shared mailbox — on-premises, in your cloud tenancy, or air-gapped. Data residency, model routing and retention are configured per deployment, not per vendor default.
04Operation
Run it, measure it, and hand it over
Hosting, monitoring, per-node evaluation and change management as the business moves. The trajectory record that accumulates belongs to you. So does the source, on terms agreed before the first line is written — including the terms under which you take the whole thing in-house.
AssuranceWhat we commit to in writing
Deployment topology
On-premises, private cloud tenancy, or air-gapped. Model routing, data residency and retention set per deployment and stated in the contract.
Ownership
Data use, ownership, source code and derivative rights are fixed before work begins. The trajectory record is yours, including the failure traces.
Auditability
Every run produces an immutable ledger of inputs, model calls, checks, approvals and outputs. Reproducible on demand, including by your auditor.
Exit
A defined handover: architecture, node specifications, evaluation suites and operating runbook. No component of the system is hostage to our continued involvement.
BenchCut in two passes. The first pass is by decision rights: three people accountable inside the firm for what ships, then eight advisers who are not. The advisory bench is then cut by function — learning product, AI and infrastructure, growth and community — so that no person appears in two groups and no function is left without a named owner.
Eleven people design and build these systems: three who hold decision rights inside the firm, and eight advisers each retained against a named practice. Closed-domain work is specified by whoever has to answer for the result, so the people listed here are the people on the engagement rather than a sales layer in front of a delivery team. The shape is deliberate, and it has a price attached.
The cost is capacity. Eleven senior people cannot run many production builds at the same time, so we take a limited number of engagements and a start date is quoted as a slot rather than as an immediate yes. If a project needs to be scaled up mid-flight with additional staff, we will delay it or decline it instead of handing the work to people you have not met, and buyers who need elastic headcount should choose a larger firm.
01 · MANAGEMENTManagement
Crison
Head of Learning Product
Owns demand mining and product architecture for StudyHub
M.S. National University of Singapore; former asset allocation and quant researcher at CICC Wealth and CSC.
Writes on education and memory science for 9,000+ followers on Zhihu, and converts that reader evidence into the StudyHub requirement set rather than into a feature wish list. Her research years at CICC Wealth and China Securities are why a learning product here is specified with a stated hypothesis and a measurable before-number, the way an investment mandate is.
M.S. NTU; PhD candidate at CUHK-Shenzhen; assistant researcher at the Asian Institute of Digital Finance.
Built the core LLM tracks the node graph runs on, and publishes deep learning work at ACM conferences, so the model choices are ones he has had to defend in review. His quantitative finance background sets where a model is allowed to reason and where a deterministic worker takes the result instead.
Owns go-to-market and the client engagement pipeline
Founder and CEO of Siren AI; ex-Boston Consulting Group and Kearney; M.S. Management, Boston University.
Took an LLM-based MVP from concept to launch at Siren AI, and before that ran a 28-person sales team as a real estate VP, raising conversion 30%. Her consulting work at Boston Consulting Group and Kearney sets how an engagement is scoped: a written problem statement and a current cost in hours or headcount, or the pilot does not start.
Owns North American market entry and admissions practice
Ed.M. Harvard; Executive Director at NCSD; founder of Panda Education; Harvard and Georgetown interviewer.
Runs North American programmes at NCSD and founded Panda Education, so US-facing requirements come from someone who has placed students, not only sold to them. As an alumni interviewer for Harvard and Georgetown he sets what counts as credible evidence inside a learner record.
Direct-entry PhD student, School of Public Policy and Management, Tsinghua University; multiple CSSCI papers.
Reviews the research behind any learning claim the products make, and says plainly when the published evidence does not carry the weight being put on it. His CSSCI-indexed public policy work is why an effect is stated with a sample size and a method rather than as a percentage on its own.
Owns audience testing and EdTech channel relationships
Leading mathematics and physics writer on Zhihu, 200,000+ followers; long-standing educational author.
Writes mathematics and physics for more than 200,000 Zhihu followers, which gives a live read on which explanations land and which quietly fail. His EdTech industry relationships are how new material is tested with real readers before it is built into a product.
03 · AI & INFRASTRUCTUREStrategic advisers — AI & Infrastructure
Tiange Xiang
Senior adviser — AI research
Owns the vision and spatial reasoning research direction
CS PhD student at Stanford; visiting scholar at MIT with Kaiming He; Spatial Intelligence Rising Star, CVPR 2026.
Works on spatial intelligence at Stanford and at MIT with Kaiming He, and rules on which perception methods are stable enough to put in front of a client and which are still papers. He was named a Spatial Intelligence Rising Star at CVPR 2026.
PhD, Nanyang Technological University; inference acceleration for multimodal large language models.
Works on making multimodal large language models cheaper to run, which is what decides whether a 155-node graph is affordable per document rather than only correct. He advises on where latency is bought back in the runtime instead of by removing checks.
Owns scheduling and optimisation across the node graph
PhD candidate in mathematics (operations research and cybernetics), Tsinghua University.
Operations research is the branch of mathematics that decides how work is ordered under constraints, and that is what he brings to the scheduling of an asynchronous graph. He advises on where parallel evidence branches earn their cost and where they only add it.
04 · GROWTH & COMMUNITYStrategic advisers — Growth & Community
Leonie Xu
Adviser — AI product & growth
Owns overseas market entry and the CEO network
Global Head of Market & Strategy at Wiz.AI; former head of ByteDance's overseas OKR consulting division.
Ran ByteDance's overseas OKR consulting division across roughly 100 global companies, so she reviews how an engagement will be measured before it is agreed. She co-founded the China Overseas CEO Community, 5,000+ CEOs, and the ByteDance alumni community.
ML quant PM at MindQuant; formerly a risk analyst at a sovereign fund; M.S. Economics, NTU.
Builds machine-learning quant products at MindQuant and previously assessed downside as a risk analyst at a sovereign fund, which is the habit he applies to pricing. He checks that a pilot has a stated cost today in hours, headcount or error rate before any number is quoted against it.
Academic and researchAffiliations represented on the bench
NATIONAL UNIVERSITY OF SINGAPORENANYANG TECHNOLOGICAL UNIVERSITYCUHK-SHENZHENASIAN INSTITUTE OF DIGITAL FINANCETSINGHUA UNIVERSITYHARVARD UNIVERSITYGEORGETOWN UNIVERSITYSTANFORD UNIVERSITYMITBOSTON UNIVERSITYSOUTHWESTERN UNIVERSITY OF FINANCE AND ECONOMICS
IndustryAffiliations represented on the bench
CICC WEALTHCHINA SECURITIES (CSC)BOSTON CONSULTING GROUPKEARNEYBYTEDANCEWIZ.AIMINDQUANTNCSDPANDA EDUCATIONSIREN AI
OfficesThree, and what each is for
APAC · ENGINEERING
Singapore
107 N Bridge Rd, Singapore 179105
Where the systems are built and run: the 155-node financial-information graph was assembled here, and this is the deployment region for clients whose data may not leave APAC.
Asia/Singapore · UTC+8
HK · CAPITAL MARKETS
Hong Kong
33 Hysan Avenue, 46/F Lee Garden One, Causeway Bay, Hong Kong
The institutional client surface: pilot scoping, evidence reviews, and the reality-gate sign-offs a regulated buyer wants done in person, inside market hours.
Asia/Hong_Kong · UTC+8
US · RESEARCH
Redwood City
631 True Wind Way, Unit 210, Redwood City, CA 94063
The research and partnership surface: Enactive Reality work, model and tooling partnerships, and the L3 reality-connected experiments that remain gated and partial.
America/Los_Angeles · UTC−8 (UTC−7 DST)
We publish our reasoning rather than our roadmap. If you want to test whether we understand your domain, the fastest route is not a capability deck — it is forty-five minutes on one of your workflows, where the failure modes we name either match your experience or they do not.
09Engagement
Two ways in. Both start with one workflow.
One is unpaid fieldwork that ends in a written register. The other is a paid standing relationship with a named principal. They are not tiers of the same thing, and the table below separates them on five axes so nobody has to guess which one applies.
Whichever you pick, the form composes the brief in your browser and transmits nothing until you send it. Copy it, take it to your own team, and argue with it before it reaches us.
Two ways to begin, and they are not variants of each other. The cut is where the decision currently sits. The free diagnostic is for an organisation that has already chosen one workflow and wants to know which parts of it can be automated to a testable standard, and which parts should be left with the people doing them. The one-to-one advisory is for a single executive who has not chosen yet, or who has a programme already running and needs a senior outside read on whether it will hold. Neither offer produces working software. And the two do not cover the whole space: if you already know what you want built and want a price for building it, that is an implementation contract, it is not described on this page, and you should write to us so it can be scoped separately. The five axes below separate the two offers; where they look similar on the surface, the axes are what tell them apart.
OFFER A
Free diagnostic
Unpaid. We examine one workflow you name and return an opportunity register that puts a labour cost and an explicit acceptance test on every line, including the lines where we recommend against automation.
OFFER B
One-to-one advisory
A paid retainer that puts one named principal alongside one executive, on a standing basis, either to decide what to build or to get a truthful read on a programme already in flight.
Who it is forcut by whether you have already chosen the workflow to change
An organisation that can name one workflow and the one person who owns it. Typical applicants are a finance director carrying a monthly reconciliation or reporting cycle, a plant manager carrying a quality or maintenance workflow, or a head of research carrying a publication and review process. You do not need a budget or an in-house technical team. You do need one process with a defined start and end, a named owner, and the willingness to let us read the materials the work actually runs on rather than a cleaned demonstration set.
One executive with the authority to change direction, in an organisation in one of two positions. Either you are not ready to build and need to decide what to build, in which case the sessions are about where the verifiable work actually is in your business. Or you have a programme already running — internal, vendor-led, or both — and need someone senior from outside to tell you whether it will reach the standard it was funded against. It is bought by the person accountable for the decision, not delegated down to the person managing the workstream.
What you put incut by whether the input is your materials or your own calendar time
One workflow, described in writing, with its named owner. Read access to the real materials behind it — the spreadsheets, the source files, the tickets, the reports as they exist internally. Half a day of on-site access with the people who do the work, or two hours of remote screen-sharing with the same people. An NDA executed before any material moves: we sign yours if you have one, ours if you do not. No purchase order, no budget commitment, and no introduction to your procurement function.
The executive's own calendar time, held as a standing slot rather than booked when something goes wrong; continuity is where the value is, and a principal who sees you once a quarter cannot tell you anything you could not read elsewhere. Access to the documents behind the decision: vendor proposals, internal build plans, programme status reports as they are written internally rather than as they are presented upwards, and the budget envelope. A tolerance for a negative answer, since a substantial part of what the retainer buys is being told that a programme already announced will not meet the standard it was announced against.
What you get outcut by whether the output is a finished document or a standing counterparty
One document, the opportunity register, with a row for each step in the workflow. Every row carries: the current state described in the operator's own words; the labour cost today, stated as people × hours per cycle and cycles per month; the verifiable portion, meaning the part of that step whose output a machine can check against a source document or a rule deck (the written set of domain rules a checker applies); the unverifiable portion, meaning the part that rests on judgement no checker can settle; and an acceptance test — the exact condition, written before any build exists, under which you would accept an automated version of that step as correct. The register also carries a written recommendation against automating specific rows, with the reason given row by row. We walk you through it in a 60-minute session and leave you the file. Nothing in it obliges you to work with us afterwards.
A named principal — the same person for the term, named in the contract before you sign, not a bench that rotates. Standing sessions with that person at an agreed cadence. After each session, a written position of about one page setting out the question you put, our answer, what evidence would change that answer, and what we do not know. On request within the retainer, a written review of a specific vendor proposal or internal build plan, assessed against the same acceptance-test standard used in the diagnostic and against the six-layer stack our own systems are built on. What you do not receive is someone who joins your stand-ups, manages your suppliers, or writes code. The honest limit: we assess what you show us and what we can ask about. Where a vendor's system cannot be instrumented or trialled, our read on it is an informed judgement on documents and answers, and we will label it as such rather than presenting it as a verified finding.
How long it takescut by whether the engagement has a fixed end date
Fixed and short. Half a day on site, or two hours remote, then five working days of our analysis. From application to register in your hands is typically four to six weeks; almost all of that gap is NDA and access turnaround inside your organisation, not analysis inside ours. The engagement ends when the register is delivered and walked through. There is no second phase unless you ask for one.
Open-ended and reviewed rather than fixed. A three-month minimum initial term, then rolling monthly, with a written review at the end of every quarter at which either side can end the arrangement without penalty and without explanation. Sessions run at a cadence set at the start and held in the diary. Advisory relationships that work tend to run past a year; the ones that do not usually end at the first quarterly review, which is what that review exists for.
What it costscut by whether money changes hands, and how the figure is set
Free. No fee, no expenses recharged, and no cost quietly recovered later in a build contract. It costs us roughly five to six person-days of senior time, which is why applications are screened against the three criteria below rather than all accepted. Two consequences follow, and you should know both before applying. We decline diagnostics that fail the criteria, in writing, naming the criterion. And the register contains no price for building anything — it gives you labour cost, acceptance tests and a recommendation; producing a build quote requires a separate, paid scoping engagement.
A paid monthly retainer. We do not publish a figure, because it is set per engagement from two inputs: the seniority of the named principal, and the hours committed to you each month, including preparation and the written work between sessions. The number is fixed in writing before the first session and does not move with usage — no hourly overage, no success fee, and no fee contingent on anything being built afterwards. If the advisory leads to a build, that is a separate contract, separately priced, and the retainer is not credited against it. For budgeting, treat a senior-principal retainer as a fraction of the monthly cost of a senior hire rather than as a contractor day rate.
Not for
Do not apply if you cannot name a single workflow and a single owner. A diagnostic spread across four departments produces a slide rather than a register, and we will not run one. Do not apply if you want a vendor comparison; we examine your process, not our competitors' products. Do not apply if the system you want is already specified and you are looking for a bid, because that is procurement and the register would tell you nothing you have not already settled. Do not apply if read access is likely to be refused once it is actually requested — we have ended diagnostics at the access stage, and both sides lost the weeks.
Do not apply if what you want is an extra pair of hands. This is not staff augmentation: the principal does not take tickets, does not manage your vendors, and does not sit inside your delivery organisation. Do not apply if you want us to build the thing, because an adviser who expects to win the implementation cannot give you an honest read on whether to implement it, and we would rather keep the read clean than hold both. Do not apply if the decision has already been taken and what is needed is an outside name on a document supporting it; we have declined that request before and will decline it again. And if the executive in the sessions cannot change the programme's direction, the retainer will produce well-written positions and no change to the outcome.
How we decideThree tests, applied in order
Bounded enough to verify
The work must produce an output that a named person can mark right or wrong against something outside the model — a source document, a rule deck, a reconciliation, a physical measurement — and we must be able to write that acceptance test in plain sentences before any build begins, because a test written after the system exists is written to the system rather than to the work.
Painful enough to fund
Somebody inside your organisation must be able to state what the work costs today as people × hours per cycle, as an error or rework rate, or as elapsed days against a deadline that matters; where no one can produce that number, the pain has never been measured, and unmeasured pain does not survive a budget round however real it feels day to day.
Owned by someone who can accept or reject
One named person, not a committee, must hold the authority to sign the acceptance test at the start and to reject the delivered result at the end; where that ownership is split across two functions, the first thing an automated system does is expose the split, and it is cheaper for you to resolve it before the engagement than in the middle of one.
What happens after you applyIncluding the step where we may decline
01 · Day 0 to day 2
1. Application, and a written reply
You send the workflow, the named owner, one paragraph on what it costs today in hours or headcount, and which of the two offers you are applying for. A principal reads it. Within two working days you receive a written reply that either accepts you to a screening call, declines with the reason stated, or comes back with two or three specific questions we need answered before we can decide.
You, then a principal on the intake rota — not an account manager
02 · Within five working days of that reply
2. Screening call, 30 minutes
We test the three criteria out loud. We ask what the work costs today, what an acceptable result would look like, and who signs it off. You should use the same half hour to ask what we have not built and where our own claims are thin — the reality-connected persistence level of our platform is gated and only partly proven, and we will say so on the call rather than after the contract.
The named principal and the workflow owner or the executive; no wider audience
03 · Within two working days of the call
3. Decision, including the point at which we say no
We decide in writing. If we decline, we name the criterion that failed and what would have to change for a later application to succeed. This is the step where most declines happen, and the two most common reasons are that no single person can accept or reject the result, and that the workflow's output is something nobody can mark right or wrong. Declining costs you a half-hour call; accepting an engagement that fails these tests would cost you a quarter.
Two principals, one of whom was not on the call
04 · Three to ten working days, driven almost entirely by your legal function
4. NDA, access and scheduling
The NDA is executed, read access is provisioned to named individuals on our side only and revoked at the end, and dates are fixed — the half-day on site or the two-hour remote session for the diagnostic, or the first standing session and the ongoing cadence for the advisory. This step, not our analysis, is what turns a two-week start into a five-week one. If your NDA turnaround is slow, say so at step one and we will send ours on the first day.
Your legal or procurement function and the named owner; our operations lead
05 · Diagnostic: register delivered five working days after the session, walked through within a further three. Advisory: first session within ten working days of signature, first written position within two working days of that session
5. Delivery, and an explicit end
For the diagnostic you receive the opportunity register, we walk you through it in 60 minutes, and the engagement ends there — no automatic next step, no proposal attached to the back of it. For the advisory the relationship begins and runs to the first quarterly review, at which either side can stop without penalty. In both cases the documents are yours to keep, circulate internally, and show to any other supplier you are considering.
The named principal, plus the two analysts who build the register on the diagnostic
LibraryFive teaching films · named-account access, requested below
Each film takes one hard idea, explains the mechanism, and shows the failure it prevents — using real graphs, real traces and real defects, including ours. Films 01 and 02 cover most of what a technical evaluator needs before a first meeting. Tell us which are relevant and we will send those and skip the rest.
FILM 01 · THESIS04:20
01 · The barrier
Why capable models stop at the edge of real work
A model that can pass a professional examination still cannot be handed a professional's responsibilities. The film separates the four things that are actually missing — authority, state, verification and consequence — and shows what each one costs to build.
00:00The tunnelling metaphor, and why we mean it literally00:55Capability is not the constraint. It stopped being the constraint some time ago.01:50Four missing layers, demonstrated on one invoice03:05What an institution is actually buying when it buys autonomy
FILM 02 · ARCHITECTURE08:40
02 · Anatomy of a 155-node graph
A full walkthrough of a production execution graph
Eleven stages, node by node, on the real system. Where it fans out, where it must synchronise, what each critic is looking for, and the three places we got the topology wrong before it worked.
00:00Ingest, resolve, snapshot — why the snapshot comes first01:40Seven evidence branches and the conflict resolver between them03:30Four product pipelines running concurrently off one insight pack05:45Cross-product consistency: the stage everyone underestimates07:10The publication gate, and what it has refused to pass
FILM 03 · METHOD06:10
03 · Verification
Looking correct and being correct are different problems
Built around four real failures: a chart whose axis was right and whose series was not, a solid that would not print, a citation that pointed at the wrong page, and a sentence that was true in May. Each one shows the checker that now catches it.
00:00Why plausibility is the enemy, not hallucination01:20State read-back: never trust a self-reported success03:00One checker per claim class, and how to know you are missing one04:40Localised repair versus regeneration, on a forty-step task
FILM 04 · VISION07:30
04 · Enactive Reality
The research programme, stated without decoration
What the term means, where it comes from in cognitive science, how it differs from virtual reality and the metaverse, and the six-layer structure required to make an experience leave something behind. Includes the three stop lines we hold ourselves to.
00:00Enactivism in one minute, and why the engineering follows from it01:30Five concepts that get conflated, separated properly03:20The six layers, and where most products stop05:10L2 persistence: what it does to a business model06:30Three stop lines, and the ones we have actually invoked
FILM 05 · ENGAGEMENT05:00
05 · A diagnostic, end to end
Half a day on site, and the document it produces
Shot during a real diagnostic, with the client's permission and their figures removed. What we watch for, the questions that change the answer, and the opportunity register we hand back — including the two items we recommended against automating.
00:00Observation before specification01:15Finding the workflow that is bounded enough to verify02:40Writing an acceptance test the client can fail us on04:00What we recommended not doing, and why
Application desk · free diagnostic and one-to-one advisoryComposed in your browser · nothing is transmitted until you press send
Who reads it
A principal on the intake rota, not a sales function. There is nobody here whose job is to qualify you. A written reply within two working days, including the reply that says no.
What happens to what you write
The brief below is assembled in your browser. Nothing leaves the page until you send or copy it. We do not run analytics on the form, and the address you give is used to answer you and for nothing else.
What it costs
The diagnostic is unpaid and its cost is not recovered later. Advisory is a monthly retainer set per engagement. Roughly one in three diagnostics ends with us recommending you do not proceed, and you keep the register.
ApplyNine questions · the brief updates as you answer
00What you are applying for
01Sector
02What the finished work has to be — the capability catalogue in section 03
03Where you want to start — the horizons in section 05
Written to by a person, answered by a person, within two working days. If a message needs a technical answer we would rather be slow and correct, and we will say so on day two rather than go quiet.
Enterprise engagements
Diagnostic → pilot → deployment. Fixed scope, fixed acceptance, fixed price at every stage. Apply in section 09.
Research collaboration
Joint work on execution graphs, verification, world runtimes and the Enactive Reality programme.
Briefing access
Named-account access to the five films, the full narration scripts and the on-screen figures, for technical evaluators.