Instant by Falcra
Instant Practices
How Instant runs the AI-Driven Development Lifecycle safely: the decisions, foundations, assurance and AI rules behind it, in detail. It is the reference for the crew, your architects and your technical owner.
The Instant Guide describes what Instant is. And the practices below describe how its parts are carried out. All of them happen alongside the build, in the room, over a few days.
To follow them step by step: the Instant run sheet puts every step in order, with who leads it and when it's done.
Published by Falcra.
The Instant Framework (theinstantframework.com) by Falcra (falcratechnologies.com)
The Instant Framework is published in three parts: the Manifesto (values and principles), the Guide (what Instant is, its roles, stages and gates) and the Practices (how its decisions, foundations, assurance and AI rules work in detail).
Choices you make
Every engagement involves a handful of decisions that belong to your organisation, not to the crew or the AI. Each is put to the room, made by the person named below, and recorded.
Minutes, not meetings. These choices are made in the room during preparation and the Blitz, by the people who own them, with the crew bringing the options ready to decide.
| Choice | Options | Who decides | When |
|---|---|---|---|
| Crew | A Combined, B Split, C Developer Conductor | Your business owner | S0 |
| Blitz format | Consecutive days, or split into sessions (Guide, chapter 6) | Your business owner | S0 |
| Requirements or user stories | One of the two, depending on your organisation | Your business owner, advised by the Challenger | S2 |
| Data architecture level | 1 Standard, 2 Standard plus audit trail, 3 Event Sourcing | Your business owner, after six questions put to the room. The crew advises but never chooses. | S2 |
| Risk tier for each part of the application | Critical, Important, Standard | Your experts and business owner | S2 |
| Security constraints package | Optional; written by your security lead | Your security lead | S0 |
| Jira | Optional | Your business owner or project manager | S0 |
Requirements or user stories
The AI build tool offers two ways of capturing what the software must do: requirements, or user stories. Generally only one is needed, so only one is used. Which one suits you depends on how your organisation already works. Either way, the business reviews the result in the room (sign-off G1).
Risk tiers
Not every part of an application matters equally. A mistake on a help screen is an inconvenience. A mistake in a payment calculation can be a loss, a fine or a safety event. Instant asks the room to sort every part of the application into one of three tiers, and to record who decided and when.
Critical
Safety, regulatory, financial or revenue logic. Anything where an error causes serious harm.
Important
Core business workflows and data integrity.
Standard
Screens, layout, forms and plumbing.
Recording and managing tiers
Tiers are kept in the risk tier register: one row for each part of the application, with its tier, the reason, who decided and when. The AI build tool proposes the parts and a suggested tier for each, so the room starts from a list, not a blank page. Your experts and business owner confirm or change each one, the Challenger tests the reasons, and the Conductor keeps the register. No row may be left undecided at gate G2. When scope changes, rows are added or re-tiered, and each change is logged beside the original decision. The register goes into the evidence pack.
A ready-made register (a spreadsheet with the tiers, reasons, change log and a summary) is a free download from theinstantframework.com.
Why it matters. In most applications only a few parts are Critical: a handful of rules and calculations. Line-by-line human review is concentrated there, while automated checks and AI review cover the rest (chapter 3). That is how review keeps pace with AI-written code without lowering the bar where it counts.
Data architecture: how much change history to keep
This choice is about the record of changes to your data: who changed what, when, and what it was before. It is not about the business data itself, which is stored in full at every level. If the application holds transactions, cases or readings, those are kept whichever level you choose.
The level is chosen in Specify, before design, because it shapes the database design and is costly to change later. Instant's data architecture rule tells the AI build tool to stop before design until a level is recorded, and gate G2 can't be passed without it. The crew advises; your business owner decides.
An example: a manager raises an approval limit from $5,000 to $10,000.
Standard
The record holds its current value: $10,000. The data itself is stored and maintained as normal; what isn't kept is the change, so the earlier value and who changed it are not recorded. Suits internal tools, simple data entry, reporting and proofs of concept.
Standard plus audit trail
As Level 1, plus a permanent audit trail of every change: the limit went from $5,000 to $10,000, changed by the manager, with their role at the time, at a recorded date and time. Suits most enterprise business applications.
Event Sourcing
Every change is stored as an event, in order, and never altered, and the current record is worked out from its events: the limit is $10,000 because it was set at $5,000 and later raised. Any record can be rebuilt as it stood on any past date, under the rules in force at the time. Suits regulated and audit-critical applications.
Level 3 has a cost: more design up front, a specialist pattern your own team must be able to maintain, and changes that take a moment to appear on screen. Nothing is deleted at Level 3, so where privacy law requires personal data to be removed, the design handles it explicitly.
Recording who viewed what, as opposed to who changed it, is access logging. It is a separate decision, often required for personal or sensitive data, and can be added at any level.
Six questions to choose the level
- Will a regulator, auditor or court ever ask why something was decided, or what was known at a past date?
- Do the rules that drive decisions change over time, and must old decisions stay explainable under the old rules?
- Will several people change the same records at the same time?
- Must the change history of a record be visible in the application itself?
- Who will maintain the application after handover, and have they worked with Event Sourcing before?
- Is there personal data that may have to be deleted on request?
A yes to question 1, 2 or 4 points toward Level 3. A no to all of questions 1 to 4 points to Level 1 or 2. Question 5 weighs against Level 3 if the people who will maintain it have no experience with it. The crew shows the room what the answers point to, and your business owner decides. The level, the answers, who decided and the date are recorded. If the level changes later, the change is recorded alongside the original decision, not instead of it.
The security constraints package
Some organisations have security rules the build must never break. If you want that, your security lead writes them down as a short set of plain files before the Blitz. Examples: the application must not be reachable from the internet, people must sign in through your identity provider, only approved protocols may be used.
The AI build tool is given the package and builds within it throughout. If someone in the room asks for something that would break a constraint, the build stops and says it can't proceed. It does not look for a way around. A halt isn't overridden in the room. Only your security lead can change the package.
The AI build tool has its own generic security defaults. They are standard industry practice, not your policy. The package is how your actual policy gets into the build.
Jira
If your organisation uses Jira, the AI build tool can create cards there during the engagement, through a Jira connector. The cards are created on the Stage Manager's instruction, or the Conductor's where there is no Stage Manager. The starting point is the Instant Jira backlog: 15 epics and 130 tasks covering everything from the first conversation to handover, including the approvals, security, environments and production readiness that usually stall a project. Each task carries its stage, an owner role, and whether it is always needed or only in some situations, and long-lead tasks are flagged to start during preparation. It imports straight into Jira. It is free under the Creative Commons Attribution 4.0 licence (CC BY 4.0): you can adapt it and share your version, as long as you credit Falcra. It is a free download from theinstantframework.com.
Information Model and Data Model
When the application is built, the AI produces an Information Model for your architects: what information the application holds, what each item means, where it comes from, who owns it, how it is protected, how long it is kept and how it maps to the database, with diagrams. Every definition says where it came from: an approved requirement, the code, the running application, or an assumption still to be confirmed. Nothing is filled in from what applications like this usually contain, and gaps are listed with the role that should answer them. The AI reads the requirements, decisions and code and changes nothing. If you supply your enterprise information model or glossary, it maps to that. The skill and template work with any application, are free under the Creative Commons Attribution 4.0 licence (CC BY 4.0), and are a free download from theinstantframework.com.
Alongside it, the AI produces a Data Model for your database administrators, developers and support teams: how the information is stored. It covers every table, column, key, constraint and index, any views, procedures and triggers, how the chosen level of change history is implemented, database security, and the migration history, with diagrams, a data dictionary for your catalogue tools, and a schema export. It is generated from a non-production copy of the database itself, ideally through a direct read-only connection (chapter 2), so it is accurate on the day it is produced, and it is regenerated at every release. Where you supply database design standards, it lists any departures; without them, it describes and doesn't judge.
The two are kept apart because they answer different questions for different people: what the information means and who owns it, for architects; and how it is built, for the people who run it. Produced together, they reconcile: every entity maps to its tables, and every business table to an entity. The Data Model skill and template are also a free download from theinstantframework.com.
Enterprise foundations
Four areas need settling in any enterprise build: non-functional requirements, existing data where there is any, sign-in and access, and environments. Your own standards come first in each. Instant adds a disciplined way of applying them at AI speed, and a starting point only where you have none.
Settled in the room. These are decided during the Blitz, mostly in Specify and Design, alongside the build. Only items with long lead times, such as environments and data access, start earlier, during preparation.
10.1 Non-functional requirements
Your standards come first. Most large organisations have a non-functional requirements catalogue, architecture standards, or service levels set by application criticality. These are given to the AI build tool in Specify and Design, and they drive its own detailed non-functional work. In AI-DLC, that is the NFR Requirements and NFR Design stages, run for each Unit of Work during Build.
What Instant adds is the business half of the conversation. Availability, recovery and acceptable data loss are business decisions, yet they are often left for technical teams to assume. In Specify, the Conductor puts ten questions to the business owners in the room, so they state those answers themselves, on the record. The answers feed your standards and the engine's NFR stages. They don't replace them.
| # | Question for the business | What it informs |
|---|---|---|
| 1 | How much of the time must it be available, and when are people using it? | Availability target and support hours |
| 2 | If it goes down, how long can the business wait for it to come back? | Recovery time objective |
| 3 | How much recent data could the business afford to lose? | Recovery point objective and backups |
| 4 | How quickly must screens respond for people to stay productive? | Performance targets |
| 5 | How many people use it at once, normally and at peak, and how will that grow? | Capacity and load testing |
| 6 | Who may see or change what, and is any of the data sensitive or personal? | Access control and privacy |
| 7 | How long must data be kept, where may it be stored, and how is it deleted? | Retention and data residency |
| 8 | Who is told when something breaks, and who fixes it? | Alerting and the support model |
| 9 | Who needs to be able to use it, including people with disabilities? | Accessibility standard |
| 10 | What may it cost to run each month, and who pays? | Cost budget and alerts |
Where there is no standard. Smaller organisations, or a first application of its kind, may have no catalogue to draw on. For those, Instant offers a starting set scaled by risk tier. The targets below suit the Important tier. Critical parts get tighter targets agreed with your business owner, and Standard parts may relax them.
| Measure | Starting target |
|---|---|
| Availability | 99% in a calendar month |
| Recovery time | 8 hours |
| Recovery point | 24 hours |
| Everyday actions | 95% within 0.5 seconds; 99% within 1.5 seconds |
| Rollback of a bad release | 15 minutes |
| Detection of a critical failure | 5 minutes |
| Accessibility | WCAG 2.2 level AA |
Whatever its source, every target is proven by a test and signed off in the evidence pack, or recorded as an exception with an owner and an end date. Your security constraints package, if you have one, overrides any target wherever it is stricter.
10.2 If your project involves existing data
Many Instant builds are new applications with nothing to bring across. If yours is one, skip to 10.3. Where the new application replaces, or draws on, a system that already holds data, migration is a workstream in its own right. It follows your organisation's data and migration standards, runs alongside the build, and often continues well beyond the Blitz. Instant adds a few disciplines to it.
- 1
Protect first
A backup is confirmed in writing, and the crew works through a read-only account, ideally on a copy of the database. Who: your database administrator or system owner.
- 2
Start from the new application
Mapping is driven by what the new application needs: for each field, where its data comes from. Anything not needed is left behind deliberately, with the business's agreement. Who: the Conductor and the AI, with your experts deciding.
- 3
Discover from evidence
The AI works through the documentation, then the database's own metadata, then the data itself, confirming what each field holds before anything is built on it. Who: the Conductor, through the read-only account.
- 4
Rehearse the migration
The AI writes the migration and reconciliation scripts. They are rerun against a copy until clean, and timed so the cutover window is known. Who: the Conductor and Roadie, with your database administrator.
- 5
Prove it
Counts, totals and samples are reconciled, rejected records are decided by their owners, and the results go into the evidence pack. Who: your data owners and experts.
Letting the AI query the database directly
The fastest way to work with an existing database is to connect the AI build tool to it directly, through a read-only database connector: an MCP server, the standard way AI tools connect to other systems. The AI then runs its own queries, reads the results and moves on, in minutes rather than the hours or days it takes to send queries to a DBA and wait for the results. The same connection serves wherever the build needs to read existing data: discovery for a migration, checking reference data, a system the new application will read from, and producing the Data Model (chapter 1).
- Your approval first, in writing, with the backup confirmed.
- A read-only account. Your DBA creates a dedicated account that can read but not change anything, ideally on a copy of the database, and proves it with one harmless write that fails.
- Credentials stay out of the AI's view. The connection is stored as a saved connection or in environment settings, never typed into the conversation or committed with the code.
- Every query is visible. Each query is shown and approved before it runs, so your DBA can see exactly what was asked of the database.
- Connected only while needed. The connector is added for this project alone and removed when the work is done.
Where your policies don't allow a direct connection, the AI writes each round of queries as one list, your DBA runs them, and the results come back together. It works, but it is much slower, so a direct read-only connection is worth requesting during preparation.
Working with the AI on existing data
- Read-only is enforced by the account, not by instruction. Instructing the AI to only read is a second layer. The control is a database account that cannot write, proven by one harmless write that fails.
- Work on a copy where possible, so discovery queries don't load the live system.
- Profile before viewing rows. The AI sees whatever a query returns, so personal and sensitive data is profiled with counts and summaries, and individual rows are viewed only when needed.
- Column names are clues, not facts. An AI reading a schema will reasonably infer meaning from names such as "status" or "active". Left unchecked, an inference like that can travel into the design as if it were confirmed. Instant has each one confirmed by a query, or by someone who knows the data, and marks it unverified until then (chapter 4).
- Ask once, answer once. When queries go through a person rather than a direct connection, the AI asks for each round as one numbered list, and the results come back together.
The mapping
The mapping is kept as a spreadsheet with one row per field in the new application: its source, the transformation rule, and its status (verified, needs a decision, unverified, or new with no source). A separate list records what is left behind, why, and who agreed. The migration scripts and checks are generated from the mapping, so the two can't drift apart. Your experts own what the old data means and what is left behind.
Logic held in the database
Where the existing database carries business logic in stored procedures, triggers or views, that logic is a scoping question in its own right. In many organisations it is extensive, complex and undocumented. The AI can inventory and summarise it from the database itself, which makes its scale visible in Specify rather than in testing. Your experts and technical owner then decide what the new application must reproduce, and that work is planned explicitly.
Data quality
For each data quality problem found, the data owner chooses where it is fixed.
| Fix it | When it suits |
|---|---|
| In the old system, before migration | Few records, and the old system is still in use. |
| During migration, by rule | A pattern that can be corrected by rule, such as formats or codes. |
| In the new system, after go-live | It needs judgement record by record. The rows are flagged for follow-up. |
An unmade decision turns into rejected records on go-live day.
Reconciliation
Every rehearsal, and the final run, is reconciled:
| Check | Passes when |
|---|---|
| Counts | For each kind of record, the number in the old system equals those migrated plus rejected plus deliberately left behind. |
| Totals | Key amounts and quantities total the same before and after. |
| Samples | Your experts check a sample of records end to end in the new application. |
| Rejects | Every rejected record is reviewed, and fixed or accepted by its owner. |
| Coverage | Every field in the mapping is verified, new with an agreed default, or decided. |
If you are replacing a live system
Cutover follows your organisation's own release and change process; Instant doesn't replace it. It adds two things: the migration has been rehearsed and timed before cutover is planned, and the old system stays read-only until the business signs off, then is archived under your retention rules.
Where a one-step switch is too risky, a parallel run can be agreed. Before it starts, agree its length and end date, what must match for sign-off, how the new system receives the same work, who reviews the differences, and which system is the record meanwhile. Every difference is explained, then fixed or accepted, and the results go into the evidence pack.
Long runs don't hold up the room. A migration takes as long as the data takes. Long runs happen in the background or between sessions while the room carries on.
Other data situations
| Situation | What changes |
|---|---|
| No existing data | No migration. Reference data, such as codes and lists, is still agreed in Specify. |
| Same database platform, redesigned | Same method, with fewer technical differences to handle. |
| A different database platform | The AI generates the conversion, and the platform differences are covered by the reconciliation checks. |
| Data in spreadsheets or files | Expect inconsistent types, duplicates and free text. Your experts agree the clean-up rules in Specify, before mapping. |
| Replacing a software-as-a-service product | No database access. Check early what the vendor's export provides (history, attachments, audit trail), its limits, and the contract's exit terms. |
| The old system stays and the new one reads from it | Integration, not migration. A read-only view or interface is agreed; until it exists, the feature is "buildable, not connectable". |
| Documents and attachments | Size, formats and permissions are checked, and where they will live is agreed. |
| Years of history | Migrate, summarise or archive read-only, decided together with the data architecture level. |
10.3 If people sign in with your organisation's accounts
Most enterprise applications sign people in through the organisation's identity provider, with roles granted by directory group. Instant treats this as a design decision in Specify, not a late integration task, and your identity standards apply.
During the Blitz the application runs with stand-in test users, so every role can be tried before your identity team has configured anything. In Specify, the room agrees the roles, what each may do, and which group grants each, along with two decisions that are always needed: what happens to someone in none of the groups, and who the first administrator is. Before go-live every role is tested with real accounts; then single sign-on is switched on and the test users are switched off in production.
A proven pattern, if you need one
Where your organisation has no standard pattern of its own, Instant offers one that has been used in practice:
- Sign-in settings managed in the application by an administrator, with secrets such as signing certificates stored encrypted and never displayed or logged.
- A readiness check that blocks switching single sign-on on while anything essential is missing, such as no group mapped to Administrator.
- Roles derived from group mappings at every sign-in, with no individual exceptions, and a stated rule for people in more than one group.
- An audit record of every change to settings and mappings.
- Deny by default: nothing defaults to Administrator, and no mapping means no access.
- Token validation through an established library, never hand-written, with support for certificate rollover.
- Explicit handling of group overage, where an identity provider omits group claims for people in very many groups.
The sign-in tests run with you cover each role's permissions, people in no group or several groups, removal from a group, disabled accounts, sign-out and session limits, and confirm the test users can't be enabled in production. The results go into the evidence pack.
10.4 Environments
Instant follows the standard path of Development, Test, UAT and Production, or yours if it has more stages, such as pre-production or training. A small, low-risk application may combine Test into UAT. Environments are provided by your organisation wherever possible, under your own security rules, and the Roadie sets them up with your infrastructure team, defined as code.
In a large organisation, environments can take weeks to provision, so they are requested during preparation. The build starts on a local development environment and doesn't wait.
- One path. Changes move up through the pipeline. Nothing is changed by hand in a higher environment, and the same build moves up with only its settings differing.
- Promotion rules. Into UAT, the automated tests must pass. Into Production, UAT must be signed off and the evidence pack signed. Rollback is tested before the first production release.
- Data by environment. Made-up data in Development and Test; realistic data in UAT, masked where personal unless you approve otherwise; real data only in Production. If your project involves a migration, rehearsals run against a dedicated copy, never production.
- Settings and secrets are held per environment, outside the code. The AI build tool is never connected to Production.
- Cost. Non-production environments are sized down and can be switched off out of hours.
Trusting code no human wrote
You can't build confidence in AI-written code by reading all of it. Instant builds confidence from evidence instead: risk-tiered human judgement, independent AI review, automated checks that are proven to work, and an evidence pack that your business owner and technical owner sign.
Alongside the build, not after it. These controls run as each unit is built during the Blitz. Most are automated or done by the AI; people spend their attention where the risk tier says it matters.
AI-DLC closes the gap between the business and the software. It also opens two new gaps. Both come down to the same shift, and this chapter explains both and the controls that answer them.
Problem 1: nobody wrote it, so nobody fully understands it
A developer who writes code understands it: why each rule is there, what they assumed, and where the risky parts are. With an AI build tool, the code arrives complete. On a business application it can run to hundreds of thousands of lines. Anyone can read it, but understanding it well enough to vouch for it takes nearly as long as writing it.
In a business-critical application, a wrong rule can mean lost revenue, a regulatory fine or a safety incident. Someone has to be able to say "this does what the business needs", and be accountable for it.
AI-written code has four features that make this harder.
It's plausible when it's wrong
Human bugs often look like bugs. AI mistakes look like confident, working logic built on an assumption nobody checked.
It arrives in volume
Code arrives faster than anyone can absorb it, so the gap in understanding grows every day.
There is no author to ask
Nobody can explain why a rule was written a certain way, unless the reasoning was captured at the time.
The tests can share the mistake
If the AI writes both the code and the tests, the tests can share its misunderstanding.
The wrong answer: a waiver
A common response is to ask the business to sign that it accepts that AI wrote the code. That moves the risk onto people who have no way to judge it. It is a disclaimer, not an assurance. A regulator or a court will ask what checks were done, not whether someone signed a waiver.
The right comparison: how organisations already cope
Organisations already run a great deal of code that nobody inside them understands: old systems whose authors have left, open-source libraries, vendor software. They never gained confidence by reading every line. They gained it from specifications, testing, review of the risky parts, monitoring and clear accountability.
Financial auditors work the same way. They don't check every transaction. They test the controls and sample by risk. AI-written code needs the same discipline, applied deliberately.
Problem 2: peer review breaks at AI volume
Traditionally a developer raises a handful of changes, and a colleague reviews each one before it is accepted. With AI, every instruction can produce a change, so hundreds can arrive, some of them tens of thousands of lines long. They are hard to understand, let alone approve. At that volume, "a person approved it" becomes a formality that proves nothing.
Peer review can't simply be dropped, because it does four jobs.
Catches defects
Before they reach users.
Enforces standards
Such as security and consistency.
Spreads knowledge
Of the code across the team.
Provides accountability
A second person approves every change to production.
In regulated organisations, that two-person rule is often a formal change-control and audit requirement. Removing review would fail audits. The answer is to keep all four purposes and change what the second person reviews, and how much of it.
Evidence, not reading. Humans where it matters, automation everywhere else, and every control proven to work.
The eleven controls
Each control says what it is, who does it, when, and what evidence it leaves behind. Together, the evidence forms the evidence pack that your business signs off (control 11).
- 1
Classify by risk
Every part of the application goes into a risk tier: Critical, Important or Standard (chapter 1).
- Who
- The room, with your experts and business owner deciding
- When
- Specify, and again whenever scope changes
- Evidence
- The risk tier register: each part, its tier, the reason, who decided and when
- 2
Keep critical logic small and separate
Critical rules and calculations live in small, isolated, readable modules, not scattered through the code. Where possible, rules are written as decision tables that your experts can read and check without reading code.
- Who
- Built by the AI to this instruction, checked by the Conductor
- When
- Design and Build
- Evidence
- A list of critical modules and their decision tables
- 3
Explain-back
For each critical module, the AI writes a plain-English explanation of what the logic does and why. An expert confirms that it matches the business rule, or corrects it. This captures the "why" a human author would have carried in their head.
- Who
- The AI writes it, an expert confirms it
- When
- As each critical module is built
- Evidence
- Signed explain-back notes
- 4
People write the acceptance tests
Your experts and testers write acceptance tests from the requirements, stating what the application must do in their own words and cases. AI-generated tests are a second layer, never the only one.
- Who
- Your experts and testers
- When
- From the Blitz onward, before the related code is accepted
- Evidence
- A traceability matrix linking each requirement to what was built and to the tests that prove it
- 5
Commit often, review in units
The AI commits at meaningful checkpoints so any step can be rolled back, but not after every prompt. A pull request is raised for each Unit of Work, not each instruction or Tweak, and is tied to its requirements (see "Working in Git" below). Each is kept small enough for its risk tier to be reviewed properly.
- Who
- The Conductor, with the AI build tool
- When
- Throughout the build
- Evidence
- A history of pull requests that maps to units and requirements
- 6
Every pull request states its intent
Each carries a plain-English summary: which requirements it serves, what changed and why, its risk tier, which tests prove it, and the result of the AI review. The reviewer judges behaviour against intent, not just code.
- Who
- Written by the AI, checked by the Conductor
- When
- Every pull request
- Evidence
- The summaries themselves
- 7
Independent AI review first
Every pull request is reviewed by an AI that is independent of the one that wrote it: a fresh session with no build history, or a different model. It reports findings ranked by severity. Critical findings block the change.
- Who
- An automated reviewer
- When
- Every pull request, before any human review
- Evidence
- Review reports, and a record of how each finding was resolved
- 8
Human review, tiered by risk
People review according to the tier, as the table below shows. The two-person rule is kept at every tier. What changes is what the second person reviews.
- Who
- A named approver for each pull request
- When
- After the independent AI review
- Evidence
- The named approver on each pull request, and a sampling log for the Standard tier
- 9
Automated checks that must fail when they should
Tests, security scanning, dependency checks, secret detection and quality checks all run on every change and block it on failure. Each check is shown to fail on a deliberately broken case before it is trusted.
- Who
- The Roadie sets them up, and they run automatically
- When
- From the first commit
- Evidence
- The check configuration, and a record of each check's failure test
- 10
Safety nets in production
Errors that get through are caught quickly and can be traced: the audit level you chose, alerts, and hard limits on critical values.
- Who
- Designed in the Blitz, built by the AI, run by you
- When
- From go-live
- Evidence
- Monitoring and alerting set-up, and audit records
- 11
The evidence pack and joint sign-off
In place of a waiver, your business owner and technical owner sign an evidence pack containing the risk tier register, the critical modules and explain-backs, the traceability matrix and test results, the pull request history and reviews, the check configuration and failure tests, the production safety nets, and any known gaps, stated openly.
- Who
- The business owner signs that the behaviour is right. The technical owner signs that the evidence is complete and the controls worked.
- When
- Before go-live, and for each significant release
- Evidence
- The signed pack itself
Working in Git: Bolts, Tweaks and pull requests
An AI build produces changes at a pace no review process was designed for. If hundreds of small requests a day each became a commit, a pull request and a story, the reviewers and the board would be buried. Instant separates four things that are usually treated as one.
| Level | How many | What it is for |
|---|---|---|
| Prompt, and Tweak | Many an hour | Asking the AI for something. A Tweak is a small cosmetic change asked for in the room, such as a colour, a label, spacing or layout. Your experts see it on screen and accept it there and then. It needs no story and no pull request of its own: the Scorekeeper's feedback record and the transcript are its trail. |
| Commit | A few per Bolt | A checkpoint the AI saves when something works or before a risky change, so any step can be rolled back. Not one per prompt. |
| Pull request | One per Unit of Work | What gets reviewed. The unit's Bolts and Tweaks are all inside it, and its summary lists the Tweaks as one group. |
| Jira item, or your tracker's | One per Unit of Work, plus one Refinements item | The board shows units, not Tweaks. The Tweaks are listed under the unit's Refinements item. |
Tweak or Bolt? A Bolt builds something: a short build cycle that delivers part of a unit. A Tweak adjusts how something already built looks or reads. The test is simple: a Tweak that changes behaviour isn't a Tweak. If a small request changes what data is shown, who can do what, a calculation, a status change or an automatic action, it's a rule. It's recorded as a decision and handled like a requirement, with its own line in the pull request and the full review its risk tier needs. Moving a button is a Tweak. Making the button appear only for managers is a rule. The AI build tool classifies every small request, and says so when a Tweak is really a rule (chapter 4).
Defaults, where you don't have your own:
- Branches. One short-lived branch per Unit of Work, named with its ticket key, so every commit links to the unit automatically, Tweaks included. That satisfies a "every commit references a ticket" rule without anyone writing stories for Tweaks. The branch is merged through a pull request, then deleted.
- Merging. Squash on merge, so the main branch has one commit per Unit of Work. The detailed commits stay in the pull request's history for rollback and audit.
- Branch protection, switched on once by the Roadie: no direct changes to the main branch; automated checks must pass; a named person must approve, and neither the AI nor the Conductor who raised the change can approve it; and every comment thread must be resolved.
- The pull request template. One file in the repository, picked up by every major Git host. The AI fills it in and the Conductor checks it before review. It is a free download from theinstantframework.com.
Your conventions win. If you already have a branching model, merge rules or ticket-linking rules, Instant uses them. It needs only one pull request per Unit of Work, and nothing reaching the main branch any other way.
Questions on a pull request
The AI build tool runs in one session, on one machine: the Conductor's. Only that session holds the build's context, so every reviewer's question goes back through the Conductor, the same way every time.
- Ask on the pull request. The reviewer writes the question as a comment on the line it concerns, starting with "QUESTION:", one question per comment. Questions and answers stay on the pull request, not in email or chat, so the record sits with the code.
- The Conductor sorts it into one of three kinds, because each kind goes to a different place.
- How the code works: the Conductor puts the question to the AI build tool on the build machine, as a question only. The AI answers and changes nothing (chapter 4). Its answer says how it knows (seen it working, read it in the code, or assumed), and the Conductor checks it before posting it under the comment.
- Why it was built this way: the Conductor answers from the record of the room: the Scorekeeper's notes, the transcripts and the sign-off register.
- Whether it should do this at all: this is a business question. The Scorekeeper puts it to the expert who owns it, and the answer is recorded. Neither the AI nor the reviewer decides it.
- Any change goes in as a new commit on the same pull request, and the reply links to it. The reviewer closes the question once satisfied.
- No merge while a question is open. The Roadie switches on the Git host's setting that blocks merging until every comment thread is resolved, so this is enforced, not left to habit.
- Questions are run at set times: during the Blitz, at the end of each Bolt, so the room isn't interrupted; after the Blitz, at least once a day at an agreed time.
An answer from the AI is a claim, not proof. Anything that matters is confirmed by a test.
How much human review each tier gets
| Tier | Human review |
|---|---|
| Critical | A named technical reviewer checks the code line by line, together with the AI review findings, the explain-back and the tests. |
| Important | A person reviews the AI findings, the intent summary and the test results, and reads the code wherever the findings point. |
| Standard | Automated checks plus AI review. People spot-check a sample regularly. |
When each control applies
| When | Controls |
|---|---|
| Preparation | Agree the approach and who signs. Bring your own regulatory and change-control requirements. |
| Blitz: Specify | Risk classification (1), decision tables for critical rules (2), and your experts begin the acceptance tests (4). |
| Blitz and Build | Explain-back (3), commits and unit-sized pull requests (5), intent summaries (6), independent AI review (7), tiered human review (8) and automated checks (9). |
| Before go-live | Evidence pack and joint sign-off (11), and safety nets in place (10). |
| After go-live | Monitoring and sampling continue, and each significant release gets its own evidence. |
Features nobody asked for
An AI often builds whole features nobody asked for: an extra screen, an export button, a dashboard, a notification, a setting. More often than not they are right, but often they aren't. They tend to surface only in testing, or in front of stakeholders during a demonstration, when someone stumbles on something nobody knew was there. A person who isn't a business expert can't tell what it is, let alone whether it belongs or is correct.
It happens because the AI is built to be helpful and complete. Where the stories leave a gap, it fills it with what similar applications usually have. It works fast and in volume, and nobody reads every line, so additions go unseen. Traceability usually runs one way only, from each requirement to what was built. Nothing checks the reverse, so a feature with no requirement behind it has nothing to be traced from.
Why it compounds
One unrequested feature is a nuisance. An AI build can add tens of them across an application, each plausible on its own. And they don't stay separate. Each unrequested feature comes with rules the AI assumed, and later features get built on top of those rules, so they end up tied to one another. Removing or changing one can quietly break another.
Finding them afterwards is slow. The AI can list what it built by reading the code, but it can't tell you why a rule exists or whether it is right. Those were assumptions, never recorded, and only your experts can judge them. Without a record, each feature has to be traced and tested one at a time, and across tens of interlocked features that takes far longer than building them did. Real examples, with screenshots, are at the end of this chapter.
That is why additions have to be governed as they are made, not discovered later: build only what was asked, and list everything a unit contains the moment it is built, while it is still one feature and not a web of them.
How Instant deals with it
- 1
Build only what was asked
A standing instruction: the AI builds only what traces to an approved requirement or story. Anything extra it thinks is needed is proposed as a suggestion, not built.
- 2
A feature inventory per unit
At the end of every unit, the AI lists in plain English everything a user could see or do, each item with the requirement or story it traces to. It is checked at sign-off G3.
- 3
Traceability both ways
Every requirement traces to what was built and its test, and everything built traces back to a requirement.
- 4
An unrequested features list
Anything that traces to nothing is listed. Before any demonstration or user acceptance testing, your experts decide each item: keep it (it becomes a story), change it, or remove it. There are no surprises in front of stakeholders.
- 5
Part of the evidence pack
The two-way traceability, and your experts' decisions on unrequested features, go into the evidence pack.
If it has already happened
Where an application has already been built with additions nobody asked for, testing every feature by hand is the slowest way out. Reading the code is much faster than clicking through every scenario, so recovery starts there.
- 1
Inventory from the code
The AI reads the code and lists every screen, button, list, status change and automatic rule, each traced to a requirement or story, or marked "nobody asked".
- 2
Rules as decision tables
Every rule the AI built is written as a decision table your experts can check without reading code: when this happens, that follows.
- 3
Your experts decide
Each unrequested item is kept (it becomes a story), changed, or removed. Removing one is checked against the rules that depend on it.
- 4
Test what's kept
Scenario tests, such as taking one record from a clean starting state through a realistic event, are written as acceptance tests so they can be run again after every change.
- 5
Into the evidence pack
The inventory, the decision tables and your experts' decisions go into the evidence pack, as they would have from the start.
Triage, don't review
An AI can produce more features in an hour than a room of experts can review in a week. If every addition has to be read and decided one by one, the bottleneck simply moves from building to reviewing, and the project is back to the old, slower pace. So Instant doesn't ask your experts to review everything. The AI sorts, and people make only the decisions that need them. It is the same principle as code review: evidence, not reading. It applies both to the feature inventory checked at each sign-off G3, and to recovering an application that has already been built.
How to do it in practice
- 1
Agree a triage policy once
At the start, the room agrees what happens by default to each kind of addition (table below). It takes minutes and settles most items in advance.
- 2
Group by behaviour, not buttons
The AI groups the inventory by business flow, such as how a failure gets picked up or who can see what, and writes each flow's rules as one decision table. Hundreds of controls become a few dozen behaviours.
- 3
The AI sorts; a second AI checks
The AI tags every item with how it knows (seen working, read from the code, or assumed), sorts it against the policy, and marks what depends on it. An independent AI review checks the sorting, and a person samples it.
- 4
Experts decide by watching
For each flow, the room watches it run on screen with the unrequested parts highlighted, and decides only the items the policy can't settle. The Scorekeeper records each decision straight into the inventory.
- 5
Time-box, then record
Anything not decided in the session goes on a list with an owner and a date. Nothing new is built on an undecided load-bearing item until it is decided. The policy, the sorting, the samples and the decisions go into the evidence pack.
| Kind of addition | Examples | Default |
|---|---|---|
| Cosmetic and navigation | Tooltips, sorting, filters, layout, labels | Keep. A person checks a sample. |
| Extra information shown | Extra columns, counts, charts | Keep if accurate. An expert spot-checks. |
| Rules: anything that changes data, status or visibility | What appears when, what a list or drop-down is filled with and from where, items leaving a list, status changes, permissions, notifications, calculations | Always an expert decision |
| Load-bearing | Anything other features depend on | Always an expert decision, and decided first |
| In a Critical risk-tier area | Anything at all | Always an expert decision |
| Nobody would miss it | An unused screen, a duplicate route to the same place | Remove |
The policy is a starting point. Your organisation can tighten it, for example by making every change in a regulated area an expert decision.
Most of it is rules
The hardest additions to spot are not new screens. They are rules: a drop-down nobody asked for, filled with data that looks familiar, that appears only in certain circumstances. Nobody knows when it appears, where its contents come from, or why. A drop-down looks cosmetic, but the rule behind it is behaviour, and it needs deciding and testing like any other.
So the inventory lists the rules, not just the controls. The AI writes each one as a single plain sentence: when this condition holds, this appears, changes or is filled with these values, from this source. For example: "When a request is overdue and nobody is assigned, it appears in the unassigned list, oldest first."
| Kind of rule | The question it answers |
|---|---|
| Visibility | When does this field, button, list item or message appear, and when is it hidden? |
| Contents | What fills this drop-down, list or field, from which source, filtered and sorted how? |
| Defaults | What is filled in automatically, and from where? |
| State changes | What moves a record from one status or list to another? |
| Calculations | What is worked out, from what, and when is it recalculated? |
| Permissions | Who can see or do this, and who can't? |
| Automatic actions | What happens without anyone pressing anything: notifications, reminders, background jobs? |
Every rule also says where its data really comes from: a real source that has been checked, a value worked out from other data, a fixed list written into the code, or sample data. Familiar-looking values from a fixed list or sample data are the ones most likely to mislead.
From decision table to test
Rules need testing, but testing them by clicking through every screen is what makes this slow. Instead, each row of a rule's decision table becomes a test case: this condition, this expected result. Your experts check the table, which takes minutes. The AI turns the rows into automated tests, which run on every change. A person watches a sample of them run on screen. Rows nobody can fill in, because nobody knows what should happen, are exactly the questions for your experts.
What to ask the AI
Produce the feature inventory grouped by business flow, and list the rules, not just the controls: for every field, button, list and drop-down, when it appears, what fills it and from where, and what changes it. For each item, give what it does, the requirement or story it traces to (or "nobody asked"), how you know (seen working, read from the code, or assumed), its category under the triage policy, and what depends on it. For each rule, say where its data really comes from: a checked source, a calculation, a fixed list in the code, or sample data. Write each flow's rules as a decision table, turn each row into a test case, and list load-bearing items first.
Inside an Instant build, this happens unit by unit at sign-off G3, so experts see a handful of items at a time while they are fresh. The pile of hundreds only builds up when that step is skipped.
Real examples, and how Instant handles them
These screenshots come from a real AI build, using sample data. Each shows a different way an AI build drifts when nobody is governing it, and the Instant guardrail that answers it.
What happened
The AI added a "Commence assessment" action that nobody had asked for, complete with its own rule and a confident explanation in a tooltip. It looked finished.
How Instant handles it
Build only what was asked: anything extra is proposed, not built. Every unit ends with a feature inventory checked at sign-off G3, and anything that traces to no requirement goes on the unrequested features list for your experts to decide.
What happened
Asked where an item goes once the button is pressed, the AI traced its own code and answered: nowhere. The item just disappears from the list, and no screen shows that anyone is working on it. It would have been found in front of stakeholders.
How Instant handles it
The inventory lists rules, not just buttons: when something appears, what fills it, from where, and what changes it. Each rule's decision table becomes test cases, and your experts watch each flow run before any demonstration.
What happened
By now the application held many unrequested features, tied together by rules based on assumptions. The AI admitted it couldn't say what most of it did without running it, so every feature would have to be tested by hand.
How Instant handles it
Every statement about what the software does is tagged: seen working, read from the code, or assumed. Recovery starts from the code, not from clicking. Triage, don't review: the AI sorts against an agreed policy, and your experts decide only what needs them.
What happened
While testing, the AI decided that test results should feed other screens, logged their absence as a High-severity defect, and called it an engineering judgement. There was no such requirement.
How Instant handles it
No invented requirements or judgements: a gap exists only against a stated requirement, and the AI tests against the requirements as written. It never decides what the business should want. Two-way traceability shows which requirement every finding is measured against.
Traditional and Instant, side by side
| Traditional | Instant | |
|---|---|---|
| Who understands the code | The developer who wrote it | Nobody holds all of it. Understanding is captured in the risk tier register, decision tables and explain-backs. |
| What gives confidence | Trust in the developer, plus review | Evidence: tests from people, independent AI review, proven checks, tiered human review |
| Peer review | A person reads every change | AI reviews every change. People review by risk, and look at findings and intent. |
| Unit of review | Each pull request | Each unit or feature, tied to requirements |
| Business sign-off | Acceptance testing | Acceptance testing plus a signed evidence pack. Never a waiver. |
| Accountability | The two-person rule | The two-person rule, kept, with the second person focused where it matters |
Honest limits. This reduces risk. It doesn't remove it, and the same is true of traditional development. AI review can miss things, which is why critical code still gets a human, line by line. Some organisations in regulated sectors have specific change-control rules. These controls are mapped to yours, and are not assumed to replace them.
Standing rules for the AI
An AI will happily state a guess as a fact, invent priorities, and quietly drop work it can't do. Instant loads a fixed set of rules on every engagement to stop that, and to make the AI show its working.
A set of standing rules is loaded on every engagement. They are kept outside the AI build tool itself, so they don't depend on how the tool is set up. They apply from the first stage to the last. Where one of these rules conflicts with what a stage would otherwise produce, the rule wins, and the AI says so when it asks for approval.
Rules for building and testing
Build only what was asked
The AI builds only what traces to an approved requirement or story. Anything extra it thinks is needed is listed for your experts to decide on, not built. Chapter 3 explains why.
No invented requirements
A gap exists only against a stated requirement. If nothing was asked for, the AI doesn't log a defect, rate it or propose a fix, and it never decides whether a behaviour is right or "correct by design". It tests against the requirements as written. Inventing requirements while testing is how a build drifts.
Say how you know
Every statement about what the software does is tagged: seen working, read from the code but not seen running, or assumed. An assumption is a question for your experts, not a fact.
A Tweak stays cosmetic
The AI classifies every small request. A Tweak only changes how something looks or reads. If a request would change what data is shown, who can do what, a calculation, a status change or an automatic action, the AI says it is a rule, not a Tweak, before making it, so it is recorded as a decision. Tweaks are listed as one group in each pull request.
A question gets an answer, not a change
Asked about the code, the AI answers and changes nothing. If its answer shows something should change, it proposes the change and waits to be told. A question about what the business wants goes to your experts, not the AI.
Show where every claim came from
Every fact is either stated in a named document, said by a named person on a named date, or worked out by the AI, and it says which. If a source hedged, the AI keeps the hedge.
A name is not data
A table or field called "customer status" is a guess about what it holds. Until someone has queried it, or a named person has confirmed it, it is marked unverified, and so is everything built on it.
Calculate, never assert
Anything that can be calculated, such as counts, totals and coverage, is calculated before it is written, and the document says the figure was derived.
Rules for planning and reporting
These apply when the AI plans work, reports progress or prepares material for decision-makers.
Priorities are never the AI's to invent
The AI doesn't rank requirements unless a document or a person has told it the ranking. It sorts them by what they depend on instead, which is a fact rather than a judgement.
Blocked is not deprioritised
When something can't be built, the AI says "the build cannot proceed on this because...". It doesn't quietly move the item to "later". Deferring is a decision a person makes with the blocker in front of them.
Provisional values are marked on sight
Sample or demonstration data is labelled on the screen itself, not only in the presenter's commentary. Before any demonstration, what it proves and what it must not be taken to prove is written down.
Deviations need three things
Any departure from an agreed rule needs an end date, a mechanism that enforces it rather than a note that records it, and a limit on what may be claimed while it stands.
Controls must enforce, not report
Every check that is set up is proved to fail when it should. A check that can't fail is measuring nothing.
Ownership is never guessed
Where no source names an owner, the AI records "owner unknown" and the question to ask. It never names a plausible person. The number of unowned items is usually the most important finding in a plan.
Machine-read files outrank the prose about them
Where a file holds both something a program reads and a description of it, the program's version is the truth. A disagreement is a defect in the file that runs.
Sequence on the real bottleneck
A dependency chart shows what could run at the same time, not what should. Schedules are built around the true constraint, which is usually a person's available attention.
An approval is not stakeholder sign-off
When a stage is approved, the AI records who approved it, and says plainly whether anything has been shown to the people it affects.
New authority comes first
When a new source answers an open question, the AI applies it before generating anything, not after. Patching afterwards leaves the old position in the history looking like a decision.
Ask once, answer once
When the AI needs information from a person, it asks for all of it in one numbered list and explains what each item is for. It makes no comment until every answer is back.
Know where you are
While the build workflow is running, its own state record is the only authority on which stage the build is in. Work outside the current stage is labelled "off-workflow", and ends with a return to the current stage.
The build-readiness spreadsheet
Before any units are planned or any delivery plan is made, the AI produces a build-readiness spreadsheet, with one row for every requirement. A delivery workflow will happily produce a plan in which every requirement has a priority, every priority looks agreed, and nothing records where any of it came from. Blocked work gets quietly re-labelled "later", and a data source named plausibly in a diagram gets treated as the data a requirement needs. The spreadsheet exists to stop all three.
It replaces priority with dependency, and sorts every requirement into one of four classes.
Ready
Needs nothing that isn't already available. Can be built and connected now.
Code buildable, data unreachable
The logic can be built, but the real data can't be reached yet.
Cannot be built
A definition, rule or decision is missing. The sheet says exactly what.
Source unverified
A source exists by name, but nobody has checked that it holds what the name suggests.
The classes state dependency, not importance. Beside each requirement the sheet records how each fact is known, what is missing, what would unblock it, and a named owner or "unknown". Columns for your decision and your instruction are left empty for you. Counters at the foot of the sheet show how many rows are in each class and how many still await a decision. That last number tells you whether the sheet has actually been worked, or merely produced. A second tab lists issues that gate the build as a whole, such as missing environments or access that hasn't been provided.
The spreadsheet is rebuilt whenever a classification changes, and it can change more than once.
Build standards
The AI removes redundant code, such as unused or duplicated code and leftover debug output, before every save of its work. It also applies your design system to every screen, or consistent spacing and layout defaults where you have none.
Code comments
The AI comments every file, every function and every section of logic in plain English: what it does, why, and which requirement it implements. Comments change with the code, and the independent AI review checks them. Developers who inherit the code can read it.
Data architecture
The AI stops before design until the room has chosen a data level (chapter 1). It then builds to the level chosen.
Your skills
Alongside the standing rules, a client skill carries one organisation's standards, terminology, branding and constraints. It is kept for each engagement. Where a client skill conflicts with the standing rules, yours win.
Working practices for the build
- Run it all, return it all. When the AI asks for several things to be run, run them all and return every result in one go. Feeding results back one at a time makes it react to each, ask for more, and lose track of what it originally asked.
- Make several passes. AI isn't perfect, so the build uses multiple passes.
- Keep it simple, and keep checking.
- Commit regularly, so a wrong assumption can be rolled back if the AI has built on it.
- Don't overload the AI with documents. It re-reads what it is given on every pass, so a lot of documents slow it down.