An experimental direction for SpecFact: team ownership, bounded agents and independent evidence across the software development lifecycle.
TL;DR: This is a visionary outlook and a proposed experiment with SpecFact: can we progress from individual AI-assisted coding at Level 2 to coordinated lifecycle delegation at Level 3? The aim is to develop SpecFact into an assurance layer connecting team-owned intent to independent evidence, while people remain accountable for decisions and outcomes.
Vision & experiment, dated 2026-10-08: The operating model below describes a potential future direction for the tool. The pinned source snapshot contains useful foundations; the integrated workflow, behavioral-model extensions and automated review policy still need implementation and evaluation. Publishing this outlook does not announce an available Level 3 system.
Why explore this direction?
A developer can ask an agent to implement a ticket, repair a failing test and respond to review comments. That is already useful. The organizational question is harder: what happens when several agents work across a backlog, a shared platform and multiple teams at once?
More implementation capacity does not automatically create more capacity to decide what should be built or verify what changed. Someone still has to preserve the intended behavior, approve important decisions and establish that the delivered candidate satisfies the relevant obligations.
DORA's qualitative research describes this shift: time saved in generation can move into auditing and verification. It motivates examining the whole delivery system; it does not establish that any particular assurance architecture solves the problem. DORA, 2026-03-10.
The proposed experiment would make approved intent and independent evidence the organizing elements of the software development lifecycle (SDLC) when agents participate in delivery. Teams would own the product and its consequences. Agents would execute bounded work. We want to develop SpecFact to connect the intended change to the evidence needed to assess it, alongside existing platform controls for merging and releasing.
The future outlook is a delivery loop in which people can spend more time shaping user outcomes, architecture and counterexamples, while agents help prepare, implement, check and repair bounded changes. Whether that improves reliability or developer experience is a question for the experiment.
What “Level 3” means here
These are working definitions for this article, rather than a universal maturity scale.
Level 2 includes copilots and IDE coding agents operating within a developer's workflow, team-maintained backlogs, and human or AI-assisted review.
Level 3, as envisioned here, extends bounded delegation across backlog preparation, implementation, validation, remediation and delivery feedback. The future workflow would preserve approved decisions, identify the evidence required for each change and route unresolved decisions to their owners. Human accountability would continue throughout.
Teams could explore this incrementally. A team that keeps human review on every pull request (PR) could still pilot the intent gate, independent evidence and bounded repair workflow. Automated merge eligibility would be a later policy decision requiring its own evaluation.
What the research contributes
Two research directions help frame this experiment. Specification-Grounded Code Review (SGCR), v2 grounds review in human-authored specifications and evaluates suggestion adoption in an industrial setting. That supports investigating more useful review context; suggestion adoption does not establish fewer escaped defects or safe removal of human review.
Meta's RADAR study, v2 describes a layered automated-review funnel with eligibility rules, risk scoring, review agents and operational controls. Its safety comparisons are observational, rather than causal estimates. Our design inference is to test narrowly scoped delegation with explicit eligibility and suspension controls. Neither paper evaluates the combined SpecFact operating model described here.
The proposed workflow revolves around two gates
The intent and architecture gate, before implementation, asks: Is the change sufficiently clear, approved and verifiable to implement?
The candidate review gate, after validation, asks: Does trustworthy evidence cover this exact candidate, and what still requires a human decision?
| Stage | Human responsibility | Agent and assurance support |
|---|---|---|
| Frame the outcome | Define the user promise, constraints and non-goals | Find missing context and propose focused questions |
| Approve intent and architecture | Resolve assumptions, approve changed boundaries and own acceptance examples | Preserve source identities and expose missing obligations |
| Delegate implementation | Grant bounded task scope and tool authority | Implement against the approved inputs with explicit budgets |
| Validate and repair | Decide disputed intent; maintain independent oracles | Run applicable checks, collect evidence and repair within limits |
| Assess the candidate | Decide exceptions and policy-defined high risk | Check evidence freshness, scope, provenance and unresolved results |
| Merge and release | Own integration, rollout and recovery | Protected controllers enforce merge and release policy |
| Learn | Review incidents, sampled changes and developer competence | Surface repeated gaps and propose improvements for review |
An agent may draft a clarification or suggest a model change. That proposal receives an explicit decision before it changes what the implementation is being judged against. Otherwise, the same agent can gradually rewrite the goal while making its own output appear aligned.
The pre-implementation gate should be proportionate. A small private refactoring can reuse unchanged intent and contracts. A change to permissions, a public interface or an irreversible operation needs the relevant owner and additional obligations. Every proposed gate needs a clear subject, owner and reason to exist.
A concrete example: reserving inventory
This is an illustrative scenario for the proposed workflow, rather than a demonstrated SpecFact integration. Suppose a team wants an agent to implement inventory reservations.
“Add a reservation endpoint” leaves important behavior unresolved. What happens when two buyers request the last item? Who may release a reservation? What happens after expiry? Is payment part of the same operation?
Before implementation, the product and service owners agree on observable behavior: the last item cannot be reserved twice; only an authorized actor can release it; expiry makes stock available again; payment remains outside this service's responsibility. They approve the touched interface and select concurrency, permission and expiry checks.
The agent can now implement against a bounded promise. Validation must challenge that promise using cases derived from approved intent, including forbidden behavior, rather than only explain how the new code works.
Independent evidence needs an independent expectation: the approved behavior and verification controls sit outside the implementing agent's authority. Agent-generated tests can contribute once checked against that expectation. Here, the test oracle is the rule used to decide the expected result.
If a test fails, the repair loop works on the implementation. A proposed change to the oracle or the user promise returns to its owner. If the required concurrency check cannot run, its result is UNKNOWN. A green unit-test job cannot replace it.
The repair loop also needs one shared limit on attempts, elapsed time and compute use. It stops when required evidence is unavailable, an intent decision remains unresolved or repairs stop making progress. The named owner receives the blocked obligation and the next decision needed.
At the review gate, the owners receive the actual unresolved decision and its evidence. They do not have to reconstruct the complete product promise from a long diff. This reservation change may still require human review because its impact includes concurrency and a public interface; the example is not an automatic-merge demonstration.
The experiment: develop SpecFact as the assurance layer
We want to explore SpecFact as an assurance layer alongside existing planning tools, coding agents, continuous integration (CI) and delivery controls. The tool-development experiment would connect existing evidence building blocks to approved decision context, bounded behavioral models and trustworthy candidate assessment.
OpenSpec, Spec Kit, backlog tools and architecture records can remain the authoring sources. In the proposed direction, SpecFact would consume the relevant context and produce bounded validation evidence. This extends the import-first direction already described in Import, Don't Author.
There are useful foundations in a pinned source snapshot. At modules commit d13a5c0ca69b, the manifests identify Requirements 0.5.1 and Code Review 0.51.0. Requirements connects imported requirements, acceptance records and reconciliation input. Code Review defines structured findings and scoped evidence for supported Python changes. These mechanisms help expose gaps and make review repeatable; the complete target operating model remains proposed. Requirements manifest, Requirements commands, Code Review manifest, Code Review findings. Source snapshot rechecked on 2026-10-08.
| Status in this article | Building block | Boundary that matters |
|---|---|---|
| Present in the pinned source | Requirements records and reconciliation | Reconciliation consumes supplied JUnit results without running tests; an owner name does not authenticate approval |
| Present in the pinned source | Structured Code Review evidence | Supported language and analysis scope remain bounded; producer output needs independent verification before authorizing a merge |
| Proposed extension | Approved decisions, behavior and architecture context | Needs explicit ownership, supported semantics and evidence invalidation when relevant inputs change |
| Proposed extension | Bounded conformance checks against approved models | Needs independently approved expectations, declared language/runtime coverage and visible unknowns |
| Proposed extension | Evidence composition and automated merge eligibility | Needs protected policy, trustworthy producers, adequate evidence and an evaluated exception process |
The experimental development path would start by associating current requirements and review evidence for one bounded change class. It would then add approved decision and architecture inputs, followed by model-based checks where the semantics are supported. Any later automated eligibility policy would require a separate evaluation. This is a proposed sequence for developing the tool, rather than a committed release schedule.
That distinction is central to the product direction: a structured result is useful; an authorized and justified delivery decision requires more. The capability description is source inspection of the pinned snapshot. An acceptance test of a complete signed installation has not been performed for this article, and these versions are not asserted to be the latest releases.
The PR carries evidence for its obligations
In the target model, a PR arrives with more than code and a persuasive agent summary. It carries evidence tied to the change:
- Approved requirement and decision references, with their identities.
- The candidate and relevant base, configuration and policy identities.
- Applicable behavioral, architectural, security, compatibility and operational obligations, including latency or accessibility where relevant.
- Producer identities, actual execution scope and current results.
- Remaining failures, unknowns, repair attempts and any separately authorized exception.
This can be a small set of linked machine-readable artifacts; it need not become a new document developers fill out by hand. Evidence should be collected from real execution and authoritative sources wherever possible. It should distinguish what was requested from what actually ran, including unsupported checks and incomplete scope.
A protected consumer checks those artifacts against the applicable policy. Candidate code, prompts or PR comments cannot grant themselves approval or lower required controls. A required failure remains a failure. A required unknown prevents automated authorization. A permitted human exception is recorded as a separate decision, with scope and expiry; it does not turn the underlying test result into a pass.
Relevant changes to the candidate, base, requirements, architecture, test oracle, policy or toolchain invalidate the affected evidence. Before merge, applicable checks must cover the integrated candidate: two separately passing patches do not establish that their combination satisfies the same obligations.
Two agents using different models can contribute useful perspectives, but their agreement alone does not supply independent evidence. A claim should be supported by an appropriate test, analysis, approved contract or explicit human judgment.
Behavioral models can help within a declared scope
Some obligations are easier to discuss as states, actions and forbidden transitions. An approved reservation model can make expiry, authorization and concurrency questions visible before implementation.
Comparing intended architecture with extracted implementation structure also has established antecedents, such as software reflexion models. Murphy, Notkin and Sullivan, 1995.
The proposed SpecFact extension would connect approved models to bounded implementation and execution evidence. It must state which semantics it supports and what it cannot determine.
A matching graph can omit an important side effect. A trace shows what happened on the exercised path; it does not prove every path. A model extracted from buggy code can describe that bug faithfully. These limits make independent intent, negative cases and additional verification essential. General correctness and safe review waivers cannot be inferred from model agreement alone.
Teams own decisions; platforms make the workflow reusable
A central architecture or review committee approving every generated change would quickly become a queue. The proposed model distributes decisions to the owners of the affected product, service or contract.
Teams own their intent, acceptance examples and service consequences. Platform teams provide reusable execution environments, protected evidence collection, policy baselines and merge/release controls. Cross-team changes use versioned contracts and explicitly identified owners of changed boundaries.
This design follows the direction of DORA's guidance on loosely coupled teams and independently testable boundaries; its effectiveness for this agentic workflow still needs evaluation. DORA: loosely coupled teams.
Applicable checks can run in parallel. Human attention is reserved for decisions that need judgment, while review of ordinary work remains the initial default. If candidate production exceeds validation or reviewer capacity, work-in-progress limits must constrain intake. More agents do not create more domain expertise.
Developers gain a different kind of influence
The interesting work could include shaping user promises, exposing hidden assumptions, designing counterexamples, investigating failed evidence and improving a service's resilience. Developers can delegate implementation while retaining meaningful authority over its purpose and consequences.
That outcome is a design objective whose effect on employee experience needs evaluation. Developers may also lose some direct coding craft, flow, authorship and incidental learning from debugging. A role dominated by supervising queues and accepting generated output could be less satisfying.
Adoption therefore needs practical learning: specify positive and negative examples, interpret bounded evidence, debug a counterexample, manage delegation and understand cross-team contracts. Preserve time for hands-on implementation, incident analysis and mentoring. This responds to the expertise and verification tensions described in DORA's qualitative research; it does not establish a proven training intervention for this proposed workflow. Employees do not all need to become formal-methods specialists; they do need enough understanding to challenge the evidence relevant to their work.
Start with an experiment that can say no
The proposed first step is a bounded workflow with ordinary human review and shadow decisions: record what the proposed gate would decide while humans still review every candidate. Choose a reversible change class, establish a baseline and declare which findings would stop the experiment.
Measure escaped defects and requirement deviations alongside review effort, false blocking, delivery delay, compute cost and developer learning. Challenge the gate with stale evidence, altered tests, missing execution and misleading approval text. Include an independent human-review comparison when evaluating any later delegation policy; blind reviews where practical.
A small pilot can expose usability problems and failure modes. It cannot prove that rare defects are sufficiently controlled for broad automated merge. The review policy should expand only when the relevant evidence supports the specific change class, with sampled independent review and a way to suspend eligibility. If the gate misses an obligation or becomes unreliable, restore ordinary human review for the affected class and preserve the original evidence for investigation.
This proposal builds on Value- & Requirements-Driven Spec-Driven Development (VR-SDD) by adding explicit ownership and evidence at organizational scale. We do not yet know whether this combination reduces escaped defects, review effort or delivery delay. A bounded implementation, a versioned policy and a comparison with the existing process are the evidence needed to evaluate it.
If your team is progressing beyond individual AI-assisted coding, start by mapping one real change through the two gates. Identify who owns its intent, what independent evidence is needed and where a human decision remains necessary. Bring that example and its difficult decisions into the discussion about SpecFact's next assurance capabilities. The experiment is to discover which parts of this Level 3 outlook deserve to become supported, evaluated tool behavior.
Further reading and source scope
- Import, Don't Author: Requirement Gates — the import-first authoring boundary.
- VR-SDD: The Agile Developer Level-Up — connecting delivery to user outcomes.
- SpecFact: Governance & Memory for AI Tools — the earlier vision and roadmap.
Public research and source links were checked on 2026-10-08. DORA's 2026-03-10 article supplies qualitative context; its team guidance supplies organizational design inputs. SGCR (arXiv:2512.17540v2) and RADAR (arXiv:2605.30208v2) supply research inputs for the experiment. Software reflexion models supply architectural prior art. These sources inform the vision; none evaluates the complete SpecFact operating model proposed here.