Skip to main content
Frameworks

The Roadmap Checklist

As of August 2026. Phase 5 deliverable. A generator, not a menu: answer nine questions, and the modifiers they emit compose into your roadmap.


How this works

There is a fixed spine of six stages. Each stage has an entry gate (what must be true to begin) and an exit gate (what must be proven to leave). No stage carries a duration, because no published evidence supports one and the guide does not defend timelines it cannot source.

Nine factors modify the spine. Each answer emits modifiers: gates added, work pulled forward or deferred, prohibitions, and defaults set. Your roadmap is the spine plus every modifier your answers emit.

Conflict rule: when two modifiers conflict, the more restrictive wins. This is what makes the generator deterministic rather than a matter of interpretation, and it is the rule that stops an aggressive risk appetite from cancelling out a regulatory constraint.

The spine

Stage 0: Ground

Entry gate. None. Everyone starts here, including enterprises with agents already in production, because the point of Stage 0 is to find out what you actually have.

Work. Run the readiness assessment across its six dimensions. Run the first use-case portfolio round. Define the sponsor model. Write down the deterministic-zone list for your business: which decisions will never be made by a model here.

Exit gate. A recorded readiness profile. At least three use cases through all three portfolio admission gates. A named sponsor per candidate use case. The deterministic-zone list written down and agreed, not assumed.

Stage 1: First value

Entry gate. Stage 0 exit.

Work. Take the solitary-work wins, which are real, cheap and immediate. Ship one evaluable use case at A2. Instrument the evidence floor: registry entry, retained action logs, named oversight, provenance-carrying grounding.

Exit gate. A measured cost per resolved outcome for that one use case, not an estimated one. An eval suite that exists and is owned by a domain SME. Traces collected outside the agent's control, which is a decision that cannot be made retroactively.

Stage 2: Platform

Entry gate. Stage 1 exit.

Work. Control plane: agent registry with owner and risk tier, first-class agent identities with short-lived credentials, a policy decision point in the tool-call path, tool servers wrapping governed APIs. Knowledge plane: purpose-scoped curation with named corpus owners, ACL propagation into derived artifacts, provenance attached at parse time.

Exit gate. No production agent running on credential-only identity. Every consequential action passing a deterministic gate the agent cannot bypass. An erasure cascade tested end to end, including vectors and traces. A kill switch drilled at more than one point.

Stage 3: Scale

Entry gate. Stage 2 exit.

Work. Expand the portfolio. Stand up the L2 learning loop: promotion gated on counterexample survival and eval regression, promoted artifacts landing outside the model, demotion path live. Per-agent budget envelopes with variance alerting. Bridge two-estate telemetry into one view.

Exit gate. At least one promoted artifact and at least one demoted artifact, because a promotion-only pipeline has not been tested. Budget envelopes live with anomaly alerting. Licensed-estate telemetry visible alongside metered-estate telemetry.

Stage 4: Autonomy

Entry gate. Stage 3 exit, plus oversight capacity calculated per the A4 gate: fan-out with wait time included, expressed as a burst rate, using measured interaction and wait times from the workload itself.

Work. Run A4 workloads. Instrument supervision load, intervention rate, escalation mix by trigger and wait time per item. Design rotation and deliberate unassisted practice against skill erosion.

Exit gate. Burst-rate capacity measured in production and not exceeded, including across the whole approved portfolio rather than per workload. Intervention rate instrumented and behaving as calibrated oversight predicts: broader standing permission alongside more frequent intervention, not less.

Stage 5: Extend

Entry gate. Stage 4 exit for the internal lane.

Work. Open the lanes that carry different failure models. Customer-facing: a separate edge on the shared control plane. Operational technology: read-only or advisory, never inside a protection layer, with tested revert-to-manual.

Exit gate. Per lane. Customer-facing measured on resolution rather than containment, with escalation to a real human queue that exists before launch. OT measured in specialist hours and avoided interventions, with the agent's output visually distinct from configured alarms.

The nine factors

F1 Audience

F2 Archetype

F3 Regulatory intensity

F4 Sovereignty

F5 Risk appetite

F6 Data readiness

Taken directly from the readiness assessment's six dimensions, scored 0 to 3. The profile governs, not the total.

F7 Vendor gravity

Gravity changes the interface, not the architecture. These modifiers are correspondingly light.

F8 Build capacity

F9 Cost preference

Running the generator

  1. Answer the nine factors honestly, using the readiness assessment for F6 rather than estimating it.
  2. Collect every modifier emitted. Apply the conflict rule: more restrictive wins.
  3. Write the resulting stage list with its composed entry and exit gates. This is your roadmap.
  4. Run the use-case portfolio against it quarterly. Stage exits and portfolio rounds are different cadences and should not be merged.
  5. Re-run the whole generator when a factor answer changes. A first regulated customer, an acquisition, or a shift from metered to seat pricing each change the roadmap, and the change is a new roadmap rather than an amendment.

Fictional worked example: Northstar Components

Teaching example only, not a benchmark. Northstar Components is fictional. Its profile and roadmap illustrate the composition method; none of the numbers or choices below are evidence about manufacturing firms.

Northstar is a mid-market manufacturer with one platform team, an ERP-led estate, mixed seat and API pricing, sectoral obligations, and an internal-first programme. Its answers are:

Applying the restrictive-wins rule produces this roadmap:

  1. Stage 0: record the six-dimension profile, select three evaluable internal use cases, name sponsors, and write the deterministic-zone list.
  2. Stage 1: ship one A2 use case against the ERP, measure cost per resolved outcome, and retain traces outside the agent's control.
  3. Stage 2: rent the platform capabilities the one team cannot operate, but implement ID2 identities, the policy gate, API wrappers, corpus ownership, and the erasure test. The A2 ceiling remains until identity, operations, and workforce reach 2.
  4. Stage 3: expand only after the Stage 2 evidence exists; add governed learning and the combined cost view.
  5. Stages 4 and 5: remain gated, not scheduled. Stage 4 needs measured burst supervision capacity; Stage 5 needs Stage 4 exit for the internal lane.

The output is deliberately not a calendar. It tells Northstar what evidence unlocks the next decision and which attractive work is premature.

What the generator does not decide

Vendor selection, which follows from the question bank in Phase 7. Department and vertical specifics, which are the blueprints in Phase 6. And the order of use cases inside a stage, which is the portfolio framework's job and depends on measurements you will only have after the first cycle.

Sources

readiness-assessments.md for F6. use-case-portfolio.md for the portfolio cadence. ../synthesis/maturity-model.md for the A-levels and the oversight gate. ../synthesis/master-target-state.md for the planes referenced in Stage 2. ../synthesis/sovereignty-matrix.md for the F4 tiers. Evidence for every gate lives in the research tracks named in those chapters.


Source: frameworks/roadmap-checklist.md in the evidence repository behind this site.

On this page