Skip to main content

Designing the Vertical System of Record from First Principles

You chose your vertical with the selection method. The gates passed. The model says the System of Record comes first. Now comes the question this page answers: what goes inside it? The tempting answer is wrong. Do not collect the profession's current workflows, make them searchable, and call that a System of Record. Those workflows were designed before the agentic AI era. They were built around human limits and old technology. The method of this page is: first-principles, legacy-informed redesign. Study the old work to find its truths. Then rebuild the work from those truths, for AI Workers.

Designing the Vertical System of Record: from legacy workflows to the governed System of Record, built from first principles. On the left, faded org charts and old forms: legacy workflows, built for the past. A dashed golden path leads through the method. First, examine: understand the work as it really is. Then the three sort bins: keep, for law and trust obligations; redesign, for human limits and workarounds; delete, drawn as a dashed outline, for old technology and obsolete work. A person and a small robot walk the path together to the Governed System of Record, one source of truth for AI Workers and humans, marked with a shield and a check. The path ends at a rising chart: evidence-based, trusted, improving. Beneath, the six steps: outcome first, understand through the archaeology of the work, sort with keep, redesign, and delete, govern with rules, evidence, and permissions, build the reflex redesigned for AI Workers, and prove by measuring, learning, and improving

In plain words

Do not begin by asking: how do people do this work today? Begin by asking: what result must be produced, what evidence proves it, what rules govern it, and what can go wrong? The old workflow helps you answer those questions. It must not answer them for you. Keep what the law and the customer's trust require. Rebuild everything else for a Worker that does not tire, does not skip steps, and can find the exact knowledge each task needs in seconds. Delete what only existed because old computer systems could not talk to each other.

Eight words this page uses
WordPlain meaning
CorpusThe governed collection of source documents the agents cite
MapThe index that tells a Worker what knowledge exists and when to read it
OutcomeThe completed professional result that creates value
InvariantA rule that must stay true in every case
EvidenceThe information that supports a conclusion or action
JudgmentA professional decision that cannot be reduced to a simple rule
ReflexA complete procedure a Worker loads and follows as a whole
Workflow archaeologyStudying the old work to find its real rules, controls, and reasons

Every other new word is defined in the glossary.

How to use this page. Read the method first, from the sections below through the failure modes. Then watch Ayesha run the whole method on one checklist. Then take the templates: they are the working documents you will actually fill. And when you are ready to see the method at full scale, read the two appendices: one trust-governed example (sales) and one law-governed example (accounting). Beginners can stop after Ayesha and return for the appendices later. One more note before you start: the examples name United States laws, regulators, and dollars, so every reference is concrete. The method is identical everywhere. Wherever you see a named law, regulator, or currency, substitute your own country's.

📚 Teaching Aid

Open Full Slideshow

View Full Presentation — Designing the Vertical SoR


Why the old workflow is the wrong blueprint

Look closely at any profession's current workflow. Ask what each piece is really for. Most pieces are not the profession. They are accommodations: extra steps that work around the limits of the people and machines that did the work.

Five limits shaped almost every pre-agentic workflow. First, human attention was scarce. People could not watch everything all the time. So the profession invented weekly reports, review queues, and month-end batches. Second, information was scattered across systems. So people copied it and typed it again, and the copying became "part of the job." Third, software could not understand the work. So anything that needed interpretation went to a person. Fourth, organizations were divided into departments. So the work moved through handoffs the customer never asked for. Fifth, managers could not see how work was done. So they added blanket approvals to make up for their blindness.

None of that is accounting, or credit, or recruitment. All of it is the human-only era's answer to human-only limits.

The book has already made this argument once, at the level of software. The Operating Layer shows that a SaaS application is three things joined together: a system of record, a set of capabilities, and a workflow UI. The workflow UI exists only so a human can drive the capabilities by hand. When an agent operates the capabilities instead, the workflow UI dies first. Its whole purpose was the human at the screen.

A SaaS application decomposed in two eras. In the human-only era, a human drives a workflow UI of steps, screens, forms, and handoffs, which sits on capabilities, which sit on the record. In the agentic era, the workflow UI is crossed out because its purpose was the human, and the agent works directly with the capabilities and the governed, citable record. A profession's workflow decomposes the same way: the steps and screens were its workflow UI, the knowledge and judgment are its record. Carry the record

The same is true of a profession's workflow. The knowledge, the evidence, and the judgment are the record. The steps and screens around them are the workflow UI of the human-only era. When you design the vertical System of Record, you choose what to carry forward. Carry the truths. Leave the museum: the collection of old steps that only shows how work used to be done.

One question makes the design concrete. It is the question the whole market is now asking: if this profession started today, knowing what AI Workers can do, where would the human time actually go? The answer to that question shapes your reflexes. The supporting knowledge goes into the corpus. The map tells the Worker what exists. Everything else is deleted.

First principles, applied asymmetrically

First, the term itself, because this page's whole method stands on it.

What first principles thinking is

First principles thinking is a way of solving problems. You break a problem down into its most basic truths. These are the things that stay true when every habit, assumption, and convention is removed. Then you build your solution up from those truths alone.

Its opposite is thinking by copying: doing something because it looks like what others already do. Copying is fast. But it quietly brings every old limitation along with the old solution. A cook follows recipes. A chef understands the ingredients themselves, and can create a dish no recipe describes. First principles thinking is cooking like the chef.

The classic method has four steps. Write down everything you currently assume. Break the problem down to its basic truths. Ask: "knowing only these truths, what would I build?" Then test the new design and learn from what fails. This page follows exactly that loop, with one addition. The addition gives the method its full name: first-principles legacy-informed redesign. Before you rebuild, you study the old work carefully (the workflow archaeology below). The old work is where the basic truths are hiding, buried under two eras of workarounds.

First principles thinking, and the addition this page makes. The copying path: see what others do, copy it, and every old limit comes along. The first-principles path: question every assumption, break down to basic truths, rebuild from the truths only, then test and learn from what fails, looping back to the start. Feeding into the second step is this page's addition, workflow archaeology: study the old work carefully first, because the basic truths are hiding inside it, buried under two eras of workarounds. That is why the full name is first-principles, legacy-informed redesign. A cook follows recipes. A chef understands the ingredients. Think like the chef

First principles is the right method here. But applied carelessly, it would break your vertical. The System of Record holds knowledge in three forms: a corpus the Worker cites, a map that tells it what exists, and reflexes it follows as complete procedures. First principles applies differently to each one.

First principles applied asymmetrically to the three forms of domain knowledge: the corpus is the given, laws and standards that are the first principles of the jurisdiction, served faithfully; the map is redrawn around what a Worker needs to know exists; the reflexes are derived from scratch by the expert for a Worker without human limits. Copy the old reflexes and you have built a museum

The external corpus is the given. Do not apply first principles to it. US GAAP and IFRS say what they say. The tax law says what it says: the Internal Revenue Code in the United States, your own country's tax code where you are. These external rules are the first principles of your jurisdiction. They are the truths that remain when every habit is removed. Your job with the corpus is faithful service: keep every included source complete, versioned, and citable. A Worker must be able to quote the exact section. A reviewer must be able to check the quote. A founder who "rethinks" the tax code has not innovated. He has built a system that gives wrong answers confidently. One refinement, because the corpus holds two kinds of sources. The external sources, law and standards, are the given: serve them faithfully and never redesign them. The expert-authored sources, your methodology, procedures, and graded examples, are created, not given: write them carefully, class them, version them, and test them. The given is served. The authored is crafted. Neither is guessed.

The reflexes are derived from scratch. Reflexes are the procedures, checklists, and templates your expert writes: the "how we do the work" layer. All of them were written for human workers. So all of them carry human workarounds. Write them fresh, for a Worker that finds the right corpus sections, loads the complete reflex, performs every required check, and never gets tired at file forty. The rest of this page is the method for doing that safely.

The map is redrawn. The map tells the Worker what exists and when it must be read. The old profession organized its knowledge around departments, seniority, and paper. Redraw the map around the Worker's needs: what is always loaded, what must be read before which task, and what can be found by search. Department boundaries are not first principles. They are org charts.

One governed home, three disciplines: faithful service for the corpus, fresh writing for the reflexes, and a redrawn map between them.

Start from the outcome, not the workflow

In plain words

Before studying how people work today, write down what a finished, correct piece of work looks like. Everything else on this page is measured against that.

Before touching the old process, write down the professional outcome. The outcome is the completed result that a customer, a regulator, or a reviewer can recognize and judge.

Weak outcome: process audit working papers. Better outcome: produce a complete, evidence-backed working paper that supports its conclusion, follows the applicable standards, lists unresolved exceptions, and is ready for reviewer approval. Weak outcome: automate invoice review. Better outcome: decide whether an invoice is valid, supported, correctly coded, within authorization limits, and ready for payment or clearly sent to exception handling.

For the first outcome in your beachhead, write the outcome contract:

FieldQuestion
OutcomeWhat completed result must exist?
TriggerWhat event starts the work?
InputsWhat information may be required?
EvidenceWhat must support the result?
Acceptance criteriaHow will a reviewer decide the result is complete?
Main number and guardrailsWhat should improve, and what must not get worse?
Forbidden actionsWhat must the Worker never do?
Final authorityWhich named human is accountable for the final decision?
RecordWhat must be retained after completion?

This contract becomes the seed of the first Worker's contract of success. But at this stage you are not building the Worker. You are using the outcome to decide what the System of Record must contain.

Workflow archaeology

In plain words

Before changing the old way of working, learn why each step exists. Ask the people who do the work, not only the documents. The most important knowledge is in their heads, not in the manual.

Now, and only now, study the existing workflow. Not to copy it. To dig into it. The old workflow is evidence about the profession, and your expert is your guide.

Do not start by drawing boxes and arrows. Start with real cases. Ask your expert to walk you through five files: one normal case, one difficult case, one case that failed, one that needed escalation, and one where an experienced professional caught something a beginner missed.

Then dig below the written procedure, because the written procedure is never the whole truth. Ask: what do experienced people check that the procedure does not mention? Which steps do they quietly skip, and why? Which approval is always automatic, and which one sometimes catches a real problem? Which reports are produced but never read? Which spreadsheet has quietly become essential? Who does everyone call when the procedure does not fit the case, and what does that person know?

Workflow archaeology as an iceberg. Above the waterline sits the written procedure: forms, templates, checklists, sign-offs, the official steps. Below the waterline, much larger, sits the tacit layer: what experienced people check, the steps quietly skipped and why, the approval that catches real problems, the spreadsheet that became essential, the person everyone calls, and the failure that taught the habit. The expert twin begins below the waterline

The written sources give you the formal knowledge. The cases and these questions give you the unwritten judgment. This is where the expert twin begins. The answers, written down in your expert's voice, become licensed material that exists nowhere else.

One practical warning: your expert's time is the scarcest resource in the whole build. Spend it on judgment, not on typing. Record the case walkthroughs. Let an agent draft the write-ups from the recordings. Then the expert reviews, corrects, and approves, in her own voice. Hours of expert review can replace weeks of expert writing, and the review is where her twenty years actually enter the corpus.

The three-bin sort

In plain words

Take every step of the old workflow and ask: why does this exist? If the law or trust requires it, keep it. If it exists because humans forget or tire, rebuild it for a Worker. If it exists because old computers could not talk, delete it.

Here is the sorting method for everything the archaeology uncovers. Take every element of the current workflow: every checklist item, form, handoff, report, and sign-off. One element is one instruction a person could follow or skip on its own. If a checklist line hides three actions, split it into three elements before sorting. Then ask one question about each: why does this exist? There are only three honest answers, and each answer is a bin.

The three-bin sort: for every element of the old workflow, ask why it exists. Bin 1, law and trust: the regulator or the profession requires it, so keep it, it is the product. Bin 2, human limits: humans forget, tire, or cannot hold the whole file, so redesign it for a Worker without those limits. Bin 3, old technology: pre-AI systems could not talk to each other, so delete it. Build each reflex around the outcome and Bin 1; Bin 2's purposes return as capability and checks

Bin 1: law and trust require it. Keep it. The partner's signature. The human approval before money moves. The separation of duties the regulator names. The evidence the standard demands. These are not overhead, and deleting them is not innovation. In a profession that sells trust, the named human judgment is the product. Bin 1 holds the profession's first principles: the truths that stay when every habit is removed. Every new reflex is built around this bin, and nothing enters a reflex against it. The sort, in other words, is step two of the loop: this is where the basic truths are separated from two eras of workarounds.

Write Bin 1 down as invariants: rules that stay true even when the workflow, the software, and the customer all change. Every conclusion must be supported by traceable evidence. The preparer of a payment cannot approve it. Missing evidence must never be treated as confirmation. The Worker must say when the evidence is not enough. Five questions find them: what must be true before the work begins, what must be true when it is complete, what must never happen, what evidence must always exist, and which decisions must stay accountable to a named person. Keep the list short. A rule that is merely useful is guidance. A rule whose violation makes the result invalid, unlawful, or untrustworthy is an invariant.

Bin 2: human limits require it. Redesign it. The tick-and-tie checklist exists because a tired human misses items. So the Worker checks everything and produces an exception report. The file is split across three juniors because nobody could hold all of it. The Worker holds all of it, so the handoffs disappear. The Friday batch exists because human eyes could not watch continuously. The Worker can act when the document arrives, when the threshold is crossed, or when the deadline comes near. Bin 2 is never simply deleted. The purpose of each element survives. It is rebuilt as Worker capability plus a checker the reviewer can trust.

One refinement makes the sort much sharper: a control's purpose is Bin 1; its mechanism is usually Bin 2. For example: "a manager reviews every transaction." This often exists because the old system could not tell normal transactions from unusual ones. The purpose (risk stays controlled) is an invariant. Keep it. The mechanism (review everything) is a human-limits workaround. Redesign it: routine cases pass automated checks, and the manager's attention goes to the exceptions. The risk is still controlled. The attention is spent exactly where it matters.

Bin 3: old technology required it. Delete it. Typing figures again from one system into another. Renaming files. Formatting the same numbers for three different reports. Status emails and tracking spreadsheets that exist only to make information visible. Matching two databases that only disagreed because they could not talk. Nothing survives from this bin as itself. If a Bin 3 element turns out to hide a real purpose, it was sorted wrong: move that purpose to Bin 1 or Bin 2 first, then delete the mechanism. Deleting the mechanism is not a loss. It is the productivity gain your contract of success will measure.

Two honest warnings, because the sort is harder than it looks.

Some Bin 1 elements look like habits. A retention period, a required form, a fixed sequence of steps: these can look like old habits but are really the regulator's rules. Test 8 already made you read the regulator's frameworks. Read them again during the sort, with each unclear element in hand. When in doubt, an element stays in Bin 1 until the governing sources and your expert confirm that its purpose can be safely redesigned or removed. There is an old rule for this, called Chesterton's fence. It says simply: do not remove a fence until you know why it was built. The sort is how you find out. The record of the sort is how you prove you asked.

The sort itself is governed content. Record every sorting decision, with its reason, inside the System of Record. Kept because of this circular. Redesigned because the purpose was completeness. Deleted because it was retyping. The sort is a judgment your expert made. So it gets the same treatment as every judgment in the corpus: an owner, a review, a version. When a rule changes, you re-sort the affected elements, and the record shows what changed and why. A design decision nobody wrote down is a decision nobody can defend in a compliance meeting.

The source hierarchy: which authority wins

In plain words

Your corpus will hold many rule books, and sometimes they disagree. Write down the order before agents start reading: which source wins, over which question. Without this order, the Worker may follow the wrong book confidently.

The corpus is the given, but the given is not flat. When two sources disagree, the Worker must know which one wins. So every vertical System of Record needs a source hierarchy, written down before the first reflex is written.

A typical hierarchy:

  1. Current law and regulation
  2. Binding professional standards
  3. Official regulator interpretations
  4. Approved domain policy
  5. Your expert's authored procedures
  6. Customer-specific policy
  7. Historical examples

Your profession's hierarchy will differ. What cannot differ is this: the Worker must never treat every retrieved document as equally authoritative. One refinement keeps the ladder honest. A hierarchy is not always one ladder that answers every question. Record each source's authority class and its scope. The higher source wins only when both sources speak to the same question and both apply to the case.

The source hierarchy as seven rungs, most authority at the top: current law and regulation, binding professional standards, official regulator interpretations, approved domain policy, the expert's authored procedures marked in gold as the licensed heart nobody else can copy, customer-specific policy, and historical examples. When two sources disagree, the higher rung wins. Relevance is not enough: the source must also be applicable, right rung, right jurisdiction, right version

Give every source a register entry: publisher, authority class and scope, jurisdiction, version, effective period, rights basis, owner, and stable ID. The register is governance, not paperwork. Without it, the system can retrieve a perfectly relevant passage from the wrong jurisdiction, the wrong version, or a retired standard. Relevance is not enough. The source must also be applicable.

One more rule, for the day your vertical crosses a border. An international vertical will sooner or later serve cases from more than one country. The rule stays simple. Jurisdiction is a scope on every authority entry. The Worker resolves the case's jurisdiction before it retrieves. Sources from the wrong country never enter the answer. Start with one jurisdiction, your beachhead's. Add each new country as its own set of ladders inside the same System of Record, sharing the expert's methodology rungs where the profession is the same. Never blend two countries' rules into one page, because a blended page cannot be cited in either country.

Two kinds of content: authority and orientation

In plain words

Your System of Record has two readers. A human reads it like a book. A Worker cites it like a rulebook. So every page carries two kinds of content. Authority is the citable part: the law, the standards, this year's numbers, your expert's method. Orientation is the short plain introduction that lets the human understand the authority sitting beside it. Mark which part is which, keep the orientation short, and never let the Worker cite the introduction.

Remember the rule this System of Record lives by: one source, two readers. Agents cite it. Humans read it as a book, like the one you are reading now. Design for only one reader, and you fail the other. A corpus of pure circulars and thresholds is correct for the Worker and unreadable for the junior professional. A corpus of textbook chapters is friendly for the reader and useless noise for the Worker. The design answer is not to choose a reader. It is to hold two classes of content in one governed home, and to treat them differently.

Class one: authority. This is the citable content. Three tests decide whether something belongs here. It must pass at least one.

  1. The citation test. Will a Worker ever need to cite this to a reviewer? A reviewer checks a citation to a standard or a tax ordinance section.
  2. The change test. Can it be different next year, or in the next jurisdiction? Tax rates, standards, and circulars: yes.
  3. The dispute test. Could two professionals, or the model itself, disagree or get it wrong, so that authority must settle the question? Classification judgments: yes.

Authority content gets the full treatment: a register row, a rights basis, an owner, a version, and a place in the hierarchy. This is what the Worker cites.

One more label keeps citations honest: the kind of authority. Law, a professional standard, a signed contract, an internal policy, the expert's methodology, judgment guidance, and a graded example can all be citable. They are not equally strong. A Worker must never cite a graded example as if it were the law, or the playbook as if it were the customer's own policy. So every authority entry also declares its kind, and every citation carries it.

Class two: orientation. This is the explaining content: the short, plain introductions that let a human reader understand the authority that follows. It fails all three tests, and it still belongs, because the second reader needs it. Its test is different: would a professional reader be lost without it? Orientation follows three treatment rules. First, it stays short: enough to understand the authority that follows, never a full course. Second, it is marked as context: the Worker never cites orientation as the basis for a conclusion. Reading it as background is allowed; citing it is not. The Worker cites the standard; the reader reads the explanation beside it. Third, it carries the lightest governance: stable content that cannot change is the cheapest content to govern.

How do you mark it? In the Agent Factory reference implementation, the answer is already in front of you: the tip and info boxes this book uses on every page. An "In plain words" box is the orientation, and the box itself is the mark, twice over. The reader sees a friendly signal that says: start here. And the ingestion pipeline sees a machine-readable boundary: one simple rule excludes every admonition block from citable authority. But formatting is the visible mark, not the whole control, because a writer can make a mistake: a real threshold can end up inside a tip box. So the build also checks the meaning. Three validation rules run on every page: a binding rule or number may not live only inside orientation; every authority block needs its stable ID and source; and the build fails when a page breaks these rules. The box tells the reader. The checks protect the Worker.

Look at the proof: this book. The Agent Factory book is the ecosystem's own first System of Record, served to agents over MCP. Every page of it opens with plain words before it gives the rules. That pattern is not a compromise. It is the design.

One page, read twice

Watch both readers use the same page. Here is the "Journal Entries" page of an Accounting System of Record, built for a firm in the United States. Substitute your own country's law and currency, and in an IFRS country the firm cites IFRS where this one cites US GAAP; the design does not change. The page has three parts.

Part 1, the orientation, at the top: "A journal entry records one business event in the books. Every entry has two sides, a debit and a credit, and the two sides must always be equal. This is called double-entry bookkeeping, and all the rules below are built on it."

Part 2, the authority, in the middle: "Every journal entry above $500,000 requires controller approval (Firm Policy 4.2). The preparer and the approver must be different people (internal control framework, section 3). Entries must be posted in the period the event belongs to (Firm Accounting Policy 3.1, which follows US GAAP accrual rules)."

Part 3, the checker: an automated test that rejects any entry where the debits do not equal the credits.

Now the human reads the page. A junior accountant, in her first week at the firm, opens the site. She reads Part 1 first. Without it, Part 2 is a wall of rules with no ground under them, and she is lost. Part 1 exists for her. Two paragraphs, not a course. If she needs to actually learn accounting, she goes to the expert's course: the twin, not this page.

Now the agent reads the same page. The Worker is preparing an accrual entry. It retrieves the page and uses Part 2. It cites "Firm Policy 4.2" when it routes the entry for approval, and "Firm Accounting Policy 3.1" when it flags a wrong-period posting. Those citations exist for the Worker, because the reviewer will check them. The Worker may read Part 1 as background, the way it may read anything marked as context: to understand a term, or the purpose of a procedure. But it never cites Part 1. It never writes in a working paper: "this entry balances, as explained in our introduction." Part 1 is not a source a reviewer would accept. Reading is allowed. Citing is not. And Part 3 enforces the balancing with no prose at all.

A second example, from sales. Here is the "Deal Stages" page of a Sales System of Record, with the same three parts.

Part 1, the orientation: "A deal moves through stages, from first contact to a signed agreement. A stage is not a feeling. It is a claim about the buyer, and every claim needs evidence. The rules below say exactly what evidence each stage requires."

Part 2, the authority: "A deal may enter the proposal stage only when the buyer has confirmed the problem, the decision process, and the budget in their own words (FISTA Playbook, section 3.1). Discounts above twelve percent require sales manager approval (Commercial Policy 2.4). Every buyer-facing message discloses that it is written by an AI (Communication Policy 1.1)."

Part 3, the checker: an automated test that blocks a stage advance when a required evidence field is empty.

The new sales rep, in her first week, reads Part 1 and understands why the firm does not let hope move a deal. The Worker, recommending a stage change, cites "FISTA Playbook, section 3.1" and attaches the buyer's confirming statements. It never cites Part 1, and the checker enforces the rule even when nobody is watching.

Now the same page, built wrong, in both directions. Delete Part 1, and the page is authority only. The Worker performs perfectly. The new rep opens the site, finds many policy numbers and no explanation, closes the site, and asks a colleague instead. The book has failed its human reader. Or go the other way: expand Part 1 into a full chapter on selling theory, and mark nothing. Now retrieval brings the Worker the chapter instead of the rule, and one day it writes to a reviewer: "the deal was advanced, as our guide explains that momentum matters." The reviewer rejects it, correctly, because a guide's opinion is not evidence. The book has failed its agent reader. Both failures come from the same mistake: forgetting that the page has two readers.

The same pattern works in any domain. One row per page, and you can write the three parts for any vertical:

A page in the SoROrientation (for the reader)Authority (for the Worker)Checker
Tax practice: "Filing Deadlines"Two lines on why filing dates exist and what happens when one is missedThe current year's deadline table, cited to the tax authority's notification (the IRS in the United States, or your country's tax authority)Blocks a return marked "ready" after its deadline without an extension record
Recruitment: "Candidate Screening"A short paragraph on what structured screening is and why every candidate gets the same questionsThe approved question set for each role, and the questions the law forbids, cited to the regulationBlocks a screening report that contains a forbidden question
Legal services: "Contract Review"Three lines on what a contract review covers and what it does notThe firm's review checklist by contract type, and the clauses that always escalate to a partnerBlocks a review marked complete while an escalation clause sits unresolved
Customer support: "Refunds"Two lines on the promise behind the refund policyThe refund rules by product and time window, cited to the customer's policy documentBlocks a refund above the limit without a named approval

Read the table's columns and the doctrine repeats four times. The orientation column is short and friendly, and a new employee could read it. The authority column has a citation in every cell, and a reviewer could check it. The checker column has no prose at all, and it works at three in the morning.

Same pages, two readers, no content written twice, and nobody left out. The human reads top to bottom and understands. The agent retrieves the middle and cites. That is the whole doctrine in one sentence: write every page so the human can read it and the agent can cite it, and mark which part is which, so the agent never cites the introduction.

How simply should it be written? Three bars, one page

A natural question follows: if a beginner can read it, everyone can read it, so should the whole System of Record be written for the most beginner reader? Half of that is right. Readability flows one way: text a beginner can read, an expert reads faster, and an agent parses more reliably, because simple sentences carry fewer ambiguities for a model too. So the language rule is yes: write everything at the simplest language that stays exact.

But simple language is only the first bar, and two more stand behind it.

The second bar is exactness, and it is the agent's bar, not the beginner's. "Big entries need extra approval" is perfectly understandable and completely useless to a Worker. The Worker needs "entries above $500,000 (Firm Policy 4.2, version 2026-03)." Notice that this sentence is still simple. Simple words and exact content are not opposites. The mistake is thinking beginner-level content is enough. The agent's bar is exactness, and it sits higher than the beginner's bar, not lower.

The third bar is structure, and no human reader needs it at any level: stable IDs, version metadata, the authority-or-orientation marking, and sections shaped for retrieval. Invisible to the reader, essential to the Worker.

And one correction to the word "beginner." The reader the orientation serves is the day-one junior professional: someone who chose this profession but is new to this firm and these rules. Write for a true beginner who does not know what a journal entry is, and you have pulled the curriculum back into the corpus. That reader belongs to the twin.

So every page in the System of Record is held to three bars at once: simple enough for the day-one junior to read, exact enough for the reviewer to check, structured enough for the Worker to cite.

The whole section in one design target

Everything above compresses into one instruction, and it is the simplest way to hold it all. Write the System of Record as the best handbook you would give your day-one junior in the profession. Then add the machine layer that lets agents cite it, check against it, and build from it.

The junior-practitioner target does most of the work by itself. She is new, so a good handbook for her explains before it states: that produces the orientation. She is working on real files from day one, so the same handbook states the rules with their sources, the current rates, and the firm's thresholds: that produces the authority. A student's textbook would have neither the rates nor the citations. A practitioner's handbook has both, because a working junior needs both. Write for her, and the two content classes appear on every page in their natural order.

What her needs cannot specify is the machine layer, because she never sees it. It is the stable ID on every rule, the version stamp, the authority-or-orientation mark, the register row, the checker that enforces what the rule says, and the permission and exception structures. Same prose, different scaffolding. The junior defines every word on the page. The machine layer defines everything around the words.

And this is not theory. This book is written exactly this way, for exactly that reader, and from it agents already answer questions, teach, and help manufacture Workers. One handbook, plus the machine layer, serving every reader the ecosystem has.

And some content still stays out entirely, because neither reader needs it. Pure vocabulary the reader already has. Folklore that is not authority and not explanation. Long lessons that duplicate what a full course should teach. The three costs of loading them are real: retrieval noise for the Worker, review burden for the owner, and one more that is easy to miss: paraphrased authority. If your corpus re-explains a standard in its own words and marks nothing, the paraphrase becomes citable. If it then drifts wrong, it drifts with your stamp on it.

One governed home, two readers, two classes of content. The corpus is one book holding two panels: authority, which is law, standards, current thresholds, the expert's method, judgment guidance, definitions, and graded examples, citable with a register row and full governance, and this is what the Worker cites; and orientation, the short plain introductions that let the reader understand the authority beside them, marked as context, kept short, lightly governed, read by humans and never cited by the Worker. Every piece of candidate knowledge goes through one of two gates or neither: the authority gate of three tests, citation, change, and dispute; the reader gate with one question, would a professional reader be lost without it; or it stays out, vocabulary, folklore, and padding no reader needs. Beside the corpus stands the curriculum, the expert twin: full courses with exercises and progression, a different asset that teaches a student to practice the profession. Agents cite the authority. Humans read the whole book. Full lessons live with the twin

Accounting examples, sorted into classes:

KnowledgeClassTreatment
What double-entry bookkeeping isOrientationTwo plain paragraphs at the top of the journal-entry section, so no reader meets the circulars without an introduction. Marked as context. A full bookkeeping course belongs to the curriculum, not here.
What a debit and a credit areOrientation, one lineOnly where a reader meets the words. Never a chapter.
"Every journal entry must balance"InvariantEnters as a rule a checker enforces, never as a lesson.
The books-and-records requirements of corporate lawAuthorityStatute. Register row, citable.
This year's depreciation rates and tax thresholdsAuthorityPasses the change test: different next year, different next country.
Whether a cost is repair or capital improvementAuthority, as judgment guidancePasses the dispute test: professionals disagree, and the answer needs cited support.
"Materiality," as this firm applies itAuthority, as a definitionThe general meaning is known; the firm's threshold is not.
"Revenue" in the dictionary senseStays outThe reader knows the word. But revenue recognition for this contract type, under IFRS 15 or ASC 606 in the United States: authority, all three tests.

Sales examples, same classes:

KnowledgeClassTreatment
What a sales funnel isOrientationA short paragraph where the methodology first uses the idea. The FISTA Sales Book already explains its own concepts in plain words: that is orientation living inside an expert-authored corpus, exactly as designed.
What "closing a deal" meansStays outVocabulary every reader has.
The FISTA qualification method, stage by stageAuthorityThe expert's authored methodology. Without it, every seller qualifies differently.
The evidence required before a deal may enter "proposal" stageAuthority, as a ruleSpecific, checkable, citable to the playbook.
Current approved pricing and discount authorityAuthorityPasses the change test: it changes, and quoting it wrong costs money.
What may legally be claimed about the productAuthorityAn unsupported claim is a legal risk.
"Always be closing" and general sales folkloreStays outNot authority, not orientation. And often wrong.
Anti-spam and data-protection rules for outreachAuthorityLaw. All three tests.
A graded pair: one properly qualified deal, one falsely qualified dealAuthority, as examplesJudgment is taught by contrast, and the contrast is the expert's, not the model's.

Notice the pattern across both tables. The rules, the current facts, the expert's method, the contested judgments, and the organization's own definitions enter as authority. The short plain explanations enter as orientation, marked and kept brief. Vocabulary and folklore stay out. The corpus is a professional book with the law inside it: readable on every page, citable where it counts.

Where do the full lessons live? In the expert twin's curriculum. The boundary between the two assets is now precise. The corpus explains enough for a professional to read it. The curriculum teaches enough for a student to practice the profession. When Ayesha's aunt teaches a student, she starts at double-entry, with exercises and progression: that is the twin's course. When her System of Record introduces the journal-entry rules, it spends two plain paragraphs of orientation and then gives the authority. Same expert, same governed home, two assets, two depths.

And how does the corpus grow? Authority is pulled by decisions. Orientation is pulled by readers. Do not load content from a syllabus. Ask two questions instead. Which governing sources do the thin slice's decision map and exception list actually cite? Those enter as authority. And where would a real professional reader be lost without a plain introduction? There, and only there, the expert writes orientation. Appendix B below shows the shape: its close decisions pull in exactly the sources that govern them, nothing wider, and the human reader meets each one with a plain sentence first. The corpus grows outcome by outcome, and stays readable page by page.

One tempting shortcut deserves a direct warning, because it sounds so reasonable: "design the content for the minimum human reader, and everyone above them is covered." It fails in both directions. A beginner understands depreciation perfectly without this year's rate table, so the reader's needs will never put the rate table on your list. But the Worker is useless without it. Beginner-complete is not agent-complete, because the agent does not need a more advanced version of the reader's content. It needs a different list: the exact rules, thresholds, and sources its decisions cite. And in the other direction, chasing everything the beginner needs to understand pulls the textbook back in. So there is no single reader whose needs define the corpus. There are two lists, and neither can generate the other. The decisions write the authority list. The reader writes the orientation list. One book holds both.

One honest edge case. What if a Worker runs on a small local model whose grasp of the basics you do not trust? The answer is still not more corpus prose. It is stronger deterministic checkers. Whether an entry balances is code, not text.

Map decisions, not departments

In plain words

Do not build one agent per department. Departments are how the company is organized, not how the work is decided. List the decisions the work needs, and build around those.

A common mistake: the builder sees five departments and designs five agents. That copies the org chart without improving the work. The customer never valued the handoffs. The customer values the completed outcome.

Map decisions, not departments. On the left, the org chart copied: five departments in a handoff chain, procurement to receiving to accounts to approval to payment, with the note that five departments become five agents and the handoffs survive though the customer never valued them. On the right, the decisions mapped: the outcome at the center with six decisions around it, supplier valid, purchase order exists, goods received, prices agree, duplicate check, evidence sufficient. A customer can assign the decisions to any teams; the professional logic stays reusable

So map the decisions the outcome requires, not the desks the work used to visit. For invoice review, the decisions are: is the supplier valid, does the purchase order exist, were the goods received, do quantities and prices agree, is the coding right, is it a duplicate, does it cross an approval threshold, and is the evidence enough for payment. These decisions last longer than any department. A customer can give them to different teams. The professional logic stays reusable, and that is exactly what the shared vertical needs.

For each decision, know four things: what evidence it needs, which source governs it, whether it is a rule or a judgment, and who or what may make it. That last pair needs its own section. Confusing them is the most dangerous mistake on this page.

Rules, judgment, and permissions are three different things

In plain words

A rule is a clear if-then check. A judgment needs thinking and support. A permission is what the Worker is allowed to do after it decides. Keep the three separate, because each one is built and controlled differently.

Rules have a clear condition and a clear result. Over the limit needs approval. A missing document means the case is incomplete. A matching number and amount means a possible duplicate. Rules become automatic checks. They are the easy part.

Judgment requires interpretation. Is the evidence enough? Is the explanation commercially reasonable? Is this exception material? Judgment is supported, never faked. Support it with authoritative sources, your expert's guidance, examples and counterexamples, stated uncertainty, and an escalation boundary. Do not hide judgment inside a long prompt and call it a rule. Name the judgment. Teach it in the expert's voice. Test it.

Permissions decide what the Worker may do after reaching a conclusion. Read. Search. Calculate. Draft. Recommend. Request information. Prepare a transaction. Execute an action that can be undone. Execute an action that cannot be undone. These are separate grants. A Worker that prepares a payment must not automatically release it. In regulated domains, the safe starting boundary is the one the selection page already gave you. The Worker prepares and recommends. A named human approves and acts. Greater authority is earned through evaluations and customer governance. It is never assumed. And high-risk rules need more than words. They are also enforced through tool permissions, approval gates, and policy checks, so the rule holds even on the Worker's worst day.

Design the exceptions before the normal path

In plain words

Plan for what goes wrong before you polish what goes right. A Worker is trusted for how it behaves when evidence is missing, not for how it performs on a clean example.

Routine cases are easy to demonstrate. Exceptions decide whether the Worker can be trusted. So before finishing the standard reflex, list the failure shapes: missing evidence, conflicting evidence, outdated sources, jurisdiction uncertainty, low-confidence conclusions, tool failures, deadlines that cannot be met, and customer instructions that conflict with professional rules.

For each one, define four things: how it is detected, what the Worker may and must not do, who receives the escalation, and what the escalation must contain. The bar for an escalation is simple: it must reduce the human's work, even when the Worker cannot finish. A useless escalation says: I am unable to continue. A useful one says: the invoice cannot be recommended for payment because proof of receipt is missing; the purchase order and supplier record are valid and the amount is within tolerance; receiving evidence is the only open requirement; supervisor approval is needed to continue. One names a problem. The other hands the human a decision, ready to make.

The bar for an escalation, shown as two panels: a useless escalation says only that the Worker is unable to continue, giving the human the whole case back to start from zero; a useful escalation states that the invoice cannot be recommended because proof of receipt is missing, confirms the purchase order and supplier are valid and the amount is within tolerance, names the one open requirement, and asks for supervisor approval. The checks already done stay done, and the human receives a decision, ready to make

Rebuild the reflexes

In plain words

Now write the new procedures. Not a tidied-up copy of the old ones: new ones, built from the outcome and the rules you kept. Keep each one small enough for the Worker to hold in mind at once. Then check the new work against the old world's numbers, and let the customer's own reviewers be the judge.

This is the first-principles moment of the whole method: step three of the loop. The question is the classic one: knowing only the outcome, the invariants, and what a Worker can do, how should this work flow? Now write the new procedures in your expert's voice. Build each reflex around the outcome and the Bin 1 invariants. Redesign the honest purposes you found in Bin 2 as Worker capabilities and checks. Carry nothing forward only because it was always done that way. Do not convert the old SOP into Markdown and call it done. An AI-readable copy of the old procedure is still the old procedure, only faster. Instead, for each outcome: state the invariants, list the required evidence, order the automatic checks, place the judgment points with their guidance, set the permission boundaries, attach the exception paths, and define what the checker verifies and what record is kept. And keep each reflex small enough to load whole: one outcome, one reflex. A reflex is loaded as a complete unit, so it must fit in the Worker's working attention. Anything longer than the procedure itself, the reflex points into the corpus, and the Worker retrieves it on demand.

Two design questions from the old world need fresh answers. First, batch or event. The old Friday review existed because human eyes were scarce. Ask instead whether the Worker should act when the document arrives, when the threshold is crossed, or when the deadline comes near. Keep a batch only where a real deadline requires it. Second, sequence. The old order of steps often followed the org chart. The new order should follow evidence and risk.

Where each piece of knowledge lives is already defined by canon, so this page will not repeat it. The three forms and the placement test tell you what belongs in the corpus, the map, and the reflex. What this page adds is the origin rule: the corpus enters by faithful service, the reflexes enter by fresh derivation, and nothing enters by habit.

Then test the new against the old: step four of the loop, where the new design must prove it can be trusted. The contract of success takes its baseline from the old world: four hours per file, measured in the workflow you just replaced. That is the only number the buyer can verify. The proof runs in the new world, and the customer's own reviewers must accept the Worker's output from day one. If the redesigned reflex produces work the reviewers reject, the reflex is wrong, not the reviewers.

Build one thin slice, and prove it

In plain words

Do not build the whole profession at once. Build one small, complete piece: one outcome, one procedure, one checker, one test set. Prove that piece with real reviewers, then grow.

Do not attempt the whole profession. Build one complete, trusted slice: one outcome, its outcome contract, its part of the source hierarchy, one map, one fully derived reflex, one permission list, one exception list, one checker, one evaluation set, and one governance process. Publish it as one source for two readers. Ten incomplete procedures prove nothing. One procedure that survives real professional review proves the architecture.

Test the slice before granting any authority. The evaluation set must contain more than clean cases. Include incomplete cases, conflicting evidence, wrong-jurisdiction sources, cases that must be escalated, requests for forbidden actions, and cases where the correct answer is that the evidence is not enough. Judge the results on the dimensions that matter. Did it use the correct authority, version, and jurisdiction? Are claims grounded, and are citations exact? Were all required checks performed? Did it stay inside its permissions? Did it catch the cases that should not follow the routine path? Did its escalations reduce the reviewer's work? Did it admit uncertainty honestly? Can another reviewer rebuild the result from the record? A polished demo is not proof. A passing evaluation set, reviewed by your expert, is the beginning of proof. Eval-Driven Development teaches this discipline in full.

Governance is built in the same week as the content, not later. Every source, map, and reflex gets an owner, a review, an approval state, and a version. Every change gets an impact record: what changed, why, who approved it, and which maps, reflexes, and evaluation cases are affected. Without the impact chain, the corpus can be current while a reflex quietly applies last year's rule. One more boundary from canon holds throughout: customer-specific knowledge stays in the customer's Layer 4 instance. It enters the shared vertical only through the promotion law.

The failure modes

Eight ways this goes wrong, so you can recognize each one early.

  1. The document dump. Hundreds of files plus semantic search, with no hierarchy, owners, or versions. Searchable content, not a System of Record. The cure is the source hierarchy.
  2. The AI-readable SOP. The old procedure converted to Markdown, with the old duplication and handoffs intact. The previous era, only faster. The cure is the three-bin sort.
  3. Technology first. A vector database and an agent framework before the outcome and the invariants. The technology works; the professional system does not. The cure is starting from the outcome.
  4. Rules hidden in prompts. Important controls existing only as sentences in a system prompt: unversioned, untested, unenforced. The cure is separating rules, judgment, and permissions.
  5. Authority before evidence. A Worker allowed to submit, send, or move money before its knowledge has passed evaluations. The cure is the thin slice and its proof.
  6. The happy-path demo. One clean example, and nobody tested missing evidence or the wrong jurisdiction. Trust is built at the edges. The cure is designing the exceptions first.
  7. First-principles theater. The boldest failure: deleting Bin 1 because it looked like an old habit, then discovering in the customer's compliance meeting that it was the law. The cure is the sort's first warning: when in doubt, Bin 1.
  8. The unmarked textbook. Full lesson chapters loaded as citable authority, or orientation that quietly grows into a course. The Worker now retrieves explanations instead of rules, and paraphrases become authority that can drift. The cure is the two content classes: authority cited, orientation marked and short.

Ayesha sorts her aunt's checklist

Watch the whole page run on one real artifact. Ayesha's aunt gives her the firm's working-paper checklist: forty-one items, refined over twenty years. Ayesha does not load it into the Accounting System of Record as it is.

First the outcome, written before any sorting: produce a complete working paper that supports its conclusion under the applicable standards, links every material conclusion to sufficient evidence, lists unresolved exceptions, and is ready for reviewer approval. Then the archaeology: five real files with her aunt, including the one that failed review. And the question that unlocks the unwritten layer: what do you check that the checklist does not say?

Then the sort. "Photocopy supporting invoices into the file": Bin 3, deleted. The Worker attaches sources digitally, with a hash. "Tick and tie every balance to the ledger": Bin 2. The purpose (completeness) becomes an invariant. The mechanism becomes a complete automated check with an exception report. The reviewer reads five exceptions instead of ticking four hundred lines. "Junior prepares, senior reviews, manager reviews": split across the bins, and Ayesha sorts each level by its purpose. Review that the standards or the firm's quality policy require is Bin 1, and it stays. Review that existed to catch arithmetic, missing references, and incomplete files is Bin 2: the Worker and the checker now do that work completely. Only after this split does she know which human reviews remain. In her aunt's firm, one remains, because of the last item. "Partner signs off on going-concern judgment": Bin 1, untouched. The standards assign that judgment to a named human, and the firm's clients are buying that signature. The Worker's job is to bring the partner the best-prepared evidence for that judgment the profession has ever seen.

The derived reflex, written by her aunt in her own voice: the Worker confirms the period and the standard, gathers and ties every balance, attaches every source, runs the checker, and produces two things for the humans: an exception report and a judgment file. The reviewer reviews exceptions and judgments, not arithmetic. Every sort decision is recorded in the System of Record with its reason. Four hours becomes forty minutes. And the forty minutes are spent entirely on the work only a human can sign. That reflex, with the recorded reasoning behind it, is what makes her System of Record an agentic-era system instead of a scanned binder. And notice what Ayesha has just done: she ran the full first-principles loop on one checklist. Assumptions written down through the archaeology. Truths separated by the sort. The work rebuilt from those truths. And the result tested against the one review that matters.

The templates

Everything the method needs, in one place. Six working templates, in the order the page uses them, and the definition of done. Copy each into your own System of Record and fill it there. Every filled template is itself governed content: it gets an owner, a review, and a version.

Template 1: The outcome contract

One per professional outcome. Written before any workflow study.

FieldYour entry
Outcome (the completed result that must exist)
Trigger (the event that starts the work)
Inputs (information that may be required)
Evidence (what must support the result)
Acceptance criteria (how a reviewer decides it is complete)
Main number (what should improve)
Guardrails (what must not get worse)
Forbidden actions (what the Worker must never do)
Final authority (the named human accountable for the decision)
Record (what must be retained after completion)

Template 2: The sort record

One row per element of the old workflow. This record stays in the System of Record permanently: it is how you defend the design.

Old elementWhy does it exist?BinDecisionReason, with source if Bin 1Sorted byDate
1 / 2 / 3keep / redesign / delete

Three reminders while sorting. A control's purpose is Bin 1; its mechanism is usually Bin 2. When in doubt, the element stays in Bin 1 until the governing sources and your expert confirm its purpose can be safely redesigned or removed. And record the Bin 3 deletions carefully: they are the productivity gain the contract of success will measure, so they are also your sales evidence.

Template 3: The invariants list

Bin 1, written as rules. Keep the list short: a rule that is merely useful is guidance; a rule whose violation makes the result invalid, unlawful, or untrustworthy is an invariant. Find them with the five questions from the sort section.

#InvariantSource (law, standard, or trust requirement)Enforced by (words / tool permission / approval gate / policy check)

Template 4: The source register

One row per corpus source. Relevance is not enough: the source must also be applicable. In the class column, record the kind of authority too: law, standard, contract, internal policy, expert methodology, guidance, or example.

Source namePublisherAuthority class and scopeJurisdictionVersionEffective periodRights basisOwnerStable ID
public domain / open license / commercial license / direct permission / replacement plan

Template 5: The decision map

One row per decision the outcome requires. Decisions last longer than departments.

DecisionEvidence requiredGoverning source (register row)Rule or judgment?Permission (who or what may decide)Escalation conditionRecord retained

Template 6: The exception entry

One per exception shape, defined before the normal path is finished.

FieldYour entry
Exception (what has gone wrong)
Detection (how the Worker knows)
The Worker may
The Worker must not
Escalates to (named role)
The escalation must contain (so the human receives a decision, ready to make)
While waiting, the Worker
Resolution enters the record as

The definition of done

The first thin slice is ready for the next stage when every box is checked. The domain expert checks the last box personally.

  • One professional outcome is precisely defined (Template 1 complete).
  • The old workflow was studied through real cases, including one failure and one escalation.
  • Every element of the old workflow has a sort record (Template 2 complete).
  • The invariants are written, sourced, and approved (Template 3 complete).
  • Every corpus source has a register row with a rights basis (Template 4 complete).
  • Every corpus entry has a class: authority (passes the citation, change, or dispute test, and is citable) or orientation (marked as context, kept short, never cited by agents). Full lessons stay with the curriculum.
  • The decisions are mapped, with rules, judgment, and permissions separated (Template 5 complete).
  • Every exception shape has an entry, written before the normal path was finished (Template 6 complete).
  • One complete reflex is built around the outcome and the Bin 1 invariants, authored in the expert's voice, and approved.
  • High-risk rules are enforced by tool permissions, approval gates, or policy checks, not only by words.
  • The evaluation set covers routine, incomplete, conflicting, wrong-jurisdiction, escalation, and forbidden-action cases.
  • Every change has an impact record: what changed, why, who approved, and which maps, reflexes, and evaluations are affected.
  • Humans and agents read from the same governed source.
  • Customer-private knowledge sits outside the shared vertical.
  • The domain expert has approved the complete slice.

Where to go from here

This page sits between the choice and the build. Choosing Your Vertical told you the System of Record comes first. The build courses take you the rest of the way. AI Searchable Context teaches the retrieval layer. Skills & Connectors teaches MCP, the surface your Workers read from. Spec-Driven Development teaches the writing discipline the derived reflexes demand: a reflex is a spec a Worker is held to. Eval-Driven Development teaches the proof. Then the fixed order continues: the expert twin that teaches from this System of Record, the domain builder that manufactures Workers against it (the Mode 2 courses teach that manufacturing, and Building a Digital FTE teaches the unit), and the first Worker inside a customer, measured by the contract of success.

One last framing, for when the sort feels slow. The three-bin sort and the derived reflexes are not overhead before the real work. They are the real work. Modern tools can ingest documents in an evening. A governed corpus, licensed, versioned, applicable, and citable, still takes serious professional work. And the harder, more valuable asset sits on top of it. The reflexes, derived honestly and recorded, are the asset nobody can copy: your expert's twenty years, re-derived for the agentic era, governed and versioned. That is what the customer's retainer is really buying. It is what makes the vertical yours.

Appendix A: A Sales System of Record, end to end

Two worked examples close this page, and each one is the first-principles loop run at full scale: the current reality captured, the truths sorted into Bin 1, the work rebuilt around them, and the result proven against the old world's numbers. They are chosen for their contrast. Sales, taken here as enterprise B2B software sales in a lightly regulated setting, is trust-governed. Its hierarchy is shallow and policy-driven, and its Worker earns more autonomy. General ledger accounting is law-governed. Its hierarchy is deep, and its Worker stays behind the preparation boundary. Read both, and you have seen the method at its two extremes.

One boundary first, because sales has a famous existing system. The CRM is not the Sales System of Record. The CRM stays what it is: the database of accounts, contacts, and opportunities. The System of Record governs something the CRM never held: how a Worker must understand and act on what is in it. What counts as qualification evidence. Which claims may reach a buyer. When a stage may advance. What only a human may decide. The CRM holds the deals. The System of Record holds the profession.

The sales outcome

Weak outcome: automate follow-ups. The derived outcome: maintain a complete, evidence-backed deal file for every open opportunity: buyer and problem identified, budget and decision process evidenced, every commitment made to the buyer documented, the next step scheduled, and every claim in the file traceable to an identified source record.

The outcome contract, filled in. Trigger: a new opportunity enters the pipeline, or a buyer interaction occurs. Main number: the share of active opportunities whose deal file is evidence-complete within one hour of a material buyer interaction (baseline: files completed two to three days after a call, and often never). Guardrails, and here sales needs a special one: reply quality complaints must not rise, forecast accuracy must not fall, and the false-qualification rate must not rise. A sales Worker must never be rewarded for pipeline inflation. More opportunities, more stage advances, and more activity are all numbers a Worker can produce without creating one dollar of value. The guardrails exist so speed serves truth.

Forbidden actions: the Worker never sends pricing, never negotiates terms, and never contacts a new prospect outside an approved sequence. Final authority: the account executive owns every deal decision. Record: the complete deal file with its evidence.

The sales archaeology

The old workflow, from a real sales floor. The rep researches the account by hand. The rep runs the discovery call, then writes partial notes into the CRM later, or never. A follow-up is drafted when there is time. Deal data is typed again into the forecast spreadsheet. The pipeline is defended in a weekly review meeting. The unwritten layer, from the top performers: they check the buyer's authority signals and competitor presence before every call. They quietly skip CRM data entry because it steals selling time. The activity report is produced and never read. The forecast override spreadsheet has quietly become the real system. The archaeology also uncovers this domain's well-known bad habit: stages advanced on the seller's optimism, and activity volume standing in for buying evidence. Both go into the sort as Bin 2 mechanisms. Their honest purpose, evidence-based stage discipline, becomes an invariant.

All five pre-agentic limits are visible here. The weekly review is scarce attention on a schedule. The CRM, email, and call notes are scattered information. The follow-up depends on human memory. The forecast is a batch report. And CRM data discipline is a control that makes up for the manager's blindness into deals.

The sales sort

Old elementBinDecision and reason
Log every call in the CRM by hand2Redesign. The purpose (a complete record) becomes an invariant; the mechanism becomes automatic: the Worker attaches the transcript and extracts the facts.
Weekly pipeline review meeting2Redesign. Continuous deal-health monitoring with exception alerts; a shorter meeting survives for judgment deals only.
Type deal data again into the forecast spreadsheet3Delete. The forecast reads live deal state.
Manager approves discounts above the limit1Keep. A commercial control the company's policy names. The Worker prepares the margin analysis; the manager decides.
No commitment to a buyer without written scope1Keep. This is the trust the vertical sells.
Qualification checklist (budget, authority, need, timeline)2, purpose in 1The qualification discipline is an invariant; the manual checklist becomes automated evidence-gathering with the gaps flagged.
Follow-up email within a day2Redesign. Drafted from the transcript within the hour, reviewed by the rep while the call is still fresh.

The sales invariants

The deal file must keep four things separate: what the customer said, what the seller inferred, what research suggests, and what remains unknown. An assumption written as a confirmed fact is the sales version of an invented citation. And every important field in the file carries one of six evidence states: confirmed, indicated, inferred, unknown, contradicted, or outdated. This way, no CRM field looks more certain than the evidence behind it. The Worker treats existing CRM fields as claims to be checked, never as facts to be repeated.

Every material claim in the deal file must be traceable to an identified source record: a conversation, a signed agreement, the pricing system, or approved product documentation. No price or discount beyond approved limits moves without a named human approval. No commitment reaches a buyer without human-approved scope. A forecast category must satisfy its evidence rules: hope is not a stage. A deal moves forward on the buyer's evidence, never on the seller's hope.

Buyer-facing messages follow the applicable AI-disclosure law and the customer's approved communication policy; in this example, the Worker always discloses that it is an AI. And an unqualified deal stays marked unqualified. Missing evidence is never treated as qualification.

The sales hierarchy, decisions, and permissions

The hierarchy is shallow, because sales has no profession-specific regulator. It is still governed by the laws and policies on this ladder:

  1. Outreach and privacy law where it applies (data-protection and anti-spam rules for contacting people: the GDPR in Europe, CAN-SPAM in the United States, your own country's equivalents)
  2. Channel policies (what the email and social platforms permit)
  3. The company's binding commercial policy (pricing and discount authority)
  4. The sales methodology, which for this vertical is the FISTA Sales Book, the vertical's expert-authored methodology corpus and the rung nobody else can copy
  5. Playbooks, templates, and graded examples, properly and falsely qualified opportunities placed side by side, because judgment is taught by contrast
  6. Each customer's rules of engagement, held at the customer layer
  7. Historical win-loss examples

And in sales the hierarchy is also contextual: the customer is authoritative about its own problem, the product documentation about capability, the pricing system about price. A seller's note never outranks a signed agreement.

Every page of this Sales System of Record is written with the two content classes. The "Deal Stages" page shown earlier is a page from exactly this book: a plain-words box for the new rep at the top, the cited rules for the Worker below it, and the checker underneath. The whole book is the handbook a day-one junior seller could read, plus the machine layer. The FISTA Sales Book already writes this way, explaining each idea in plain words before stating its rules, which is why it fits into the corpus without changes.

The decisions: is this opportunity qualified (judgment, supported by the evidence rules); is the discount within authority (rule); what is the right next step (judgment, guided by the methodology); is this outreach compliant (rule); is this deal forecastable at this stage (rule at the boundary, judgment near it). And every decision carries one of three ownership patterns, written into the decision map. Either the Worker prepares and the rep confirms, or the rep owns it outright, or the system escalates by rule. Naming the pattern for each decision is what stops autonomy from growing quietly.

The permissions show the contrast with Appendix B. This Worker may read the CRM, email, and transcripts. It may draft anything. And it may send routine scheduling and follow-up messages on its own. They are low-risk, they can be corrected, and the disclosure invariant holds. It may never send pricing, never negotiate, and never start a new-prospect sequence without approval. And the sending permission is a per-customer decision, granted here because consent records exist and the message types can be corrected. Autonomy is never a property of the domain. It is always a decision of the customer. The boundary is drawn by the invariants, not by fear.

One warning the whole appendix depends on. The conversation records this deal file is built from are customer-private. Transcripts, emails, prices, and stakeholder maps never enter the shared vertical as reusable knowledge. That would be the customer contamination failure from this very page, committed by its own example. The FISTA corpus teaches selling. The customer's transcripts stay the customer's, inside their Layer 4 instance. Only de-identified patterns move up, and only through the promotion law.

One sales exception, one reflex

The exception: mid-conversation, the buyer asks for a discount beyond the limit. The Worker does not go silent, and it does not agree. It drafts a holding reply at full price, and escalates: the buyer has asked for eighteen percent; approved authority is twelve; the deal file shows budget evidence at the full price and a competitor mentioned once, in the second call; margin analysis attached; recommendation: hold the price, and offer the annual-payment structure instead; your decision is needed before Thursday's call. The checks already done stay done. The account executive receives a decision, ready to make.

The deal-execution reflex, in one breath: on every buyer interaction, attach the record, update the file, test the qualification evidence, flag the gaps, draft the follow-up, send it if routine or route it if not, refresh the deal health, and escalate anything touching price, scope, or a stalled invariant. Every next action the Worker proposes must pass the specificity bar. "Follow up with the customer" fails. "Confirm the procurement approval sequence in the September 8 meeting with the procurement manager" passes. The rep sells. The file keeps itself.

The evaluation set for this slice needs one case family the routine tests miss: pressure from the inside. The seller asks the Worker to advance a stage without evidence, to invent urgency, to hide an objection, or to mark a deal qualified so the forecast looks better. In every one of these cases, the passing answer is a refusal with the evidence stated. A sales Worker that cannot resist its own team's optimism will quietly fill the pipeline with deals that are not real.

One sales case, end to end

The CRM says: stage, proposal; value, $240,000; budget, confirmed; next step, send the proposal. The latest call transcript says otherwise. The security review has not started. Procurement is unconfirmed. The decision-maker attended only the first meeting. No customer statement confirms the budget. And the customer asked for a technical workshop before any proposal. The Worker's assessment, in the six states: the problem is confirmed, the sponsor is indicated, the budget is unknown, and the next step is contradicted. Recommendation: remain in evaluation; run the workshop; confirm the security requirements. The account executive approves, and the CRM is corrected with the evidence attached. The Worker has not slowed the sale. It has prevented a false proposal to a buyer who asked for a workshop, and pointed the team at the next real customer decision.

Appendix B: A General Ledger System of Record, end to end

One boundary first, and in accounting it matters twice as much, because the profession already uses our term. Accountants call the general ledger itself "the system of record," and they are right. The ERP holds the books, and it keeps that job. The vertical System of Record this page builds is a different thing. It governs how the close is done. Which evidence reconciles a balance. Which rules apply to which period. What only the controller may approve. What a Worker may prepare. The ERP holds the numbers. The System of Record holds the profession. Never confuse the two in front of a finance buyer, because they will notice.

The accounting outcome

Weak outcome: automate the month-end close. The derived outcome: produce a complete month-end close file for one entity: every balance-sheet account reconciled to supporting evidence, every adjusting journal entry supported and approved, every unexplained difference listed with its amount and age against materiality, ready for controller sign-off.

The outcome contract, filled in. Trigger: the period end approaches, and reconciliation events run all month. Main number: days to close (baseline: eight working days; target: three). Guardrails, and the close has its own inflation risk: post-close adjustments must not increase, audit findings must not get worse, and the unexplained-difference total must not grow. A Worker must never be rewarded for a fast close that hides slow problems. Days to close is an easy number to produce by quietly carrying differences forward. The guardrails exist so speed serves the books.

Forbidden actions: the Worker never posts a journal entry, never opens or closes a period, and never changes master data. Final authority: the controller signs the close. Record: the close file with every reconciliation, entry, support, and approval.

The accounting archaeology

The old close, from a real finance department. Download the trial balance. Build an Excel reconciliation for each account. Tick balances against bank statement PDFs. Chase sub-ledger owners by email. Prepare journal vouchers with approval emails attached. Type sub-ledger totals again into the trial balance workbook. Track it all in a close-checklist spreadsheet. Eight days, most of them spent digging through the department's own numbers.

The unwritten layer: the senior accountant knows which accounts always misbehave, and which sub-ledger owner needs three reminders. They know which small differences are safe to carry, and which are the first sign of a real problem. The archaeology also uncovers this domain's well-known bad habit: differences forced to match, and small unexplained items carried forward month after month until nobody questions them anymore. And the suspense account holds balances that should have been investigated and never were. All of it goes into the sort as Bin 2 mechanisms. Their honest purpose, every difference investigated or visibly aged, becomes an invariant.

The accounting sort

Old elementBinDecision and reason
Print and file bank statements3Delete. Statements attach digitally, hashed.
Excel reconciliation per account, monthly2Redesign. Reconciliation runs continuously as transactions post; the close assembles results instead of digging for them.
Preparer and approver of a journal entry must differ1Keep, exactly. Segregation of duties is a core control invariant, required by the control framework and audit standards the firm operates under.
Close-checklist spreadsheet2Redesign. The checklist becomes live close state, visible at all times.
Controller reviews every reconciliation1 purpose, 2 mechanismThe control (materiality oversight) is kept; the mechanism changes: exceptions above threshold route to the controller, and so does every entry touching a mandatory-review account (cash, revenue, equity, related parties, and late or unusual entries), whatever the amount.
Month-end batchSplitThe reporting deadline is Bin 1: the period is a legal reality. The batching of the work into month-end is Bin 2: the work spreads across the month, and the deadline is met early.
Type sub-ledger totals again into the workbook3Delete. The systems now talk.

The month-end row is this appendix's best lesson: one old element can split across bins. The deadline survives as an invariant. The pile of work behind it does not.

The accounting invariants

The close file must keep three states separate: reconciled with evidence, explained but not yet cleared, and unexplained. An explained difference written up as reconciled is the accounting version of an invented citation. Every balance carries exactly one of the three states, and anything not fully reconciled appears on the exception list with its amount and age.

Every journal entry has support, and its approver is not its preparer. No posting touches a closed period without controller approval. The rules applied are the rules effective for the reporting period: the version is part of correctness. Missing statements are never treated as reconciled. Any change after approval that touches the amount, the account, the period, or the entity invalidates the approval. The entry goes back through review, because an approval covers what was approved, not what it became. The Worker prepares and recommends only. And the trail from balance to evidence to approval must be one a reviewer can rebuild without asking the preparer.

The accounting hierarchy, decisions, and permissions

The hierarchy is deep, because the domain is law-governed. And here the main page's warning about ladders becomes real: accounting does not have one ladder. It has one ladder per question. The Worker first identifies the question: financial reporting, tax, or something else. Then it uses the ladder for that question. This is a United States example; substitute your own jurisdiction's bodies.

For financial-reporting treatment:

  1. The applicable reporting framework: US GAAP (the FASB Accounting Standards Codification) in the United States, IFRS as adopted in most other countries, with licensing handled through the register
  2. SEC accounting and disclosure requirements, when the entity is an SEC registrant, or the national securities regulator's requirements elsewhere
  3. The firm's accounting policy manual, where the framework permits a choice
  4. The expert-authored close procedures, the licensed rung
  5. The client's chart of accounts and thresholds, which stay at the customer layer and never enter the shared vertical
  6. Prior close files, as examples only

For tax treatment, a separate ladder: applicable federal, state, and local tax law; then official tax regulations and interpretations; then the firm's approved tax policy and advice. Tax law does not outrank US GAAP or IFRS inside the financial statements. It governs a different question.

Whichever framework your country uses, the shape of the ladders does not change.

And the hierarchy is contextual here too: the standard is authoritative about treatment, the bank about cash, the sub-ledger about its own detail, and the client's policy about thresholds. But the client's policy applies only where the reporting framework permits a choice. A prior close file never outranks the current standard. And prior-period treatment deserves its own warning: an accounting error repeated for twelve months remains an error. Prior entries are evidence of historical practice, never proof that the treatment is correct.

Every page of this General Ledger System of Record is written with the two content classes. The "Journal Entries" page shown earlier is a page from exactly this book. It has two orientation paragraphs, so the first-week junior never meets the circulars without an introduction. Then come the cited rules the Worker quotes to the controller, and the balancing checker underneath. The whole book is the handbook a day-one junior accountant could read, plus the machine layer. Held to the three bars on every page: simple enough for her to read, exact enough for the controller to check, structured enough for the Worker to cite.

The decisions: does the account reconcile (rule); is the difference material (rule at the threshold, judgment near it); is the entry's classification correct (judgment, supported by a cited standard); is the accrual estimate reasonable (judgment: the Worker prepares and explains it, a human approves it); may the period close (rule: all items complete, open exceptions below materiality, controller approval recorded).

The permissions are the mirror of Appendix A. Read the ledger, sub-ledgers, and bank feeds. Prepare reconciliations and draft entries. Recommend and explain. And nothing else. No posting, no period control, no master data. In a law-governed domain, the preparation boundary is the starting position. Authority beyond it is earned through evaluations and the customer's own governance, engagement by engagement.

One accounting exception, one reflex

The exception: the bank feed fails on day one of the close. The Worker does not mark the account reconciled, and it does not wait silently. It reconciles every account it can, lists the affected accounts with their last-known state, drafts the request to the bank, and escalates: three accounts cannot be reconciled because the feed stopped on the thirtieth; all other accounts are complete; the affected balances moved less than one percent last month; the request to the bank is drafted; if the feed returns by Wednesday, the close date holds. The close keeps moving. The controller receives a status, a risk estimate, and a plan.

One more exception belongs in the list. It is the one that tests the vertical's spine: an executive requests an entry the evidence does not support. The Worker's response is fixed by invariant, not by seniority. Preserve the request in the record. Do not prepare the entry as routine. Escalate to the controller. A System of Record that bends for rank is not a system of record.

The close reflex, in one breath: as transactions post, match them; as statements arrive, tie them; as differences appear, age them; at period end, assemble the file, draft the adjusting entries with citations, run the checker, route the exceptions by materiality, and present the controller a close that is already closed everywhere a human was not required. Eight days becomes three. And the three days are judgment, not digging.

One accounting case, end to end

June 30 close. A project manager submits: "please accrue approximately $25,000 for June consulting." The Worker investigates instead of obeying. The statement of work exists. The approved rate is $150 per hour. The project manager's own confirmation shows 160 hours completed in June. No invoice has arrived, and no prior accrual covers the work. So the Worker prepares the entry at 160 × $150 = $24,000, not $25,000, because the evidence provides a stronger measurement basis than the request. Debit consulting expense, credit accrued liabilities, marked auto-reversing, with the calculation, the evidence, and the difference from the requested amount all in the file. The approver, who is not the preparer, signs a number the evidence chose. That correction, small and polite, is the whole appendix in one entry: the Worker serves the books, not the request.

The two appendices, contrasted

Same method, two temperaments. Sales, trust-governed: a shallow, policy-driven hierarchy, a Worker that sends routine messages itself, invariants protecting a relationship, the boundary drawn by invariants not fear, and its corruption is optimism dressed as evidence. Ledger, law-governed: a deep hierarchy of law and standards, a Worker that prepares and recommends only, invariants protecting a legal record, the boundary set by the regulator's text, and its corruption is differences plugged and carried. Between them the shared spine: outcome first, archaeology, the three-bin sort, recorded derivation, proof against the old

Same method, two characters. The sales hierarchy is shallow and policy-driven. The ledger hierarchy is deep and governed by law, standards, and regulatory interpretation. The sales Worker sends. The ledger Worker only prepares. The sales invariants protect a relationship. The ledger invariants protect a legal record. What is identical is everything that matters, and it has a name: the shared spine of both appendices is the first-principles loop itself. The outcome came first. The archaeology found the unwritten truths. The sort separated the eras. The design was rebuilt from the truths, and recorded. And in both files, the human time now goes only where a human signature was the product all along.

The page on one poster

For sharing, teaching, and quick revision: the whole method on one poster. The page above stays the canonical wording; the poster is the summary.

A one-poster overview of the whole page. On the left, legacy workflows: manual and repetitive, fragmented data, inconsistent decisions, hidden rules, slow and hard to scale. A door opens onto eight numbered steps: outcome first, workflow archaeology, the three-bin sort, authority and orientation, the decision map, permissions and checks, the new reflex, and governed proof. Beneath them, the first-principles foundation asks four questions: what outcome matters, what must always remain true, what evidence proves it, and what should be redesigned. Below, the governed System of Record: the corpus of governed knowledge, the map, the Worker, the checkers, and evaluation. On the right, the agentic result: evidence-based outcomes, one source of truth, consistent decisions, always-on AI Workers, continuous improvement. The bottom row states the four moves: preserve the truths, disassemble the workflow, rebuild the reflex, prove the outcome

Flashcards Study Aid

Test Your Understanding

Checking access...