A prompt can improve one response. A system improves the probability that the right person receives a useful, reviewable result every time the task runs. The system includes the prompt, but it also includes inputs, source rules, output structure, validation, human approval, destination, ownership, failure handling and measurement.
Concise answer
Use prompting when you are exploring or completing a low-risk one-off task. Build a system when the work repeats, is shared, uses sensitive or authoritative information, affects customers or publications, or must survive a change of operator. The goal is not maximum automation. It is repeatable quality with visible responsibility.
Who this guide is for
This guide is for builders, consultants, freelancers and teams who have found a useful AI task but are receiving inconsistent results. It assumes no enterprise governance programme. It does assume that someone will own the process, define acceptable output and review material decisions.
Key takeaways
- A prompt is one component of a workflow, not the workflow itself.
- Repeatability begins with a stable input contract and output contract.
- Fixed rules should remain deterministic; AI should handle the bounded judgement step.
- Human review must specify what is checked, not merely state that a human is involved.
- Exceptions and fallback are normal operating paths, not edge cases to ignore.
- Test sets, versioning and operating metrics are required before increasing autonomy.
- Agentic systems need narrower tools and permissions than their human operators.
Jump to a section
1. Define the boundary · 2. Choose maturity · 3. Create the input contract · 4. Build the prompt asset · 5. Separate rules and judgement · 6. Validate the output · 7. Design human review · 8. Handle failure · 9. Test and version · 10. Worked example · Checklist
1. Define the outcome and boundary
Begin with the result, not the wording of the prompt. “Summarise this” is an instruction. “Produce an approved client action brief from a completed meeting record within one working day” is an outcome. The outcome reveals the trigger, inputs, reviewer, destination and timing that the prompt alone cannot define.
| Boundary question | Useful answer |
|---|---|
| What starts the work? | A completed form, approved document, scheduled review or named person |
| What must be true before it starts? | Required fields exist, source material is current and permissions are valid |
| What may the system produce? | A draft, classification, recommendation or structured record within a defined scope |
| What may it never decide alone? | Material claims, financial commitments, publication, legal interpretation or high-impact personal decisions |
| Who owns the result? | A named person or role with authority to approve, correct and stop the process |
| Where does approved work go? | The actual project, content, CRM, ticketing or document system used by the team |
Write the unacceptable result as clearly as the desired one. Examples include inventing facts, exposing restricted information, sending external communication without approval, overwriting an authoritative record or silently dropping an uncertain case.
2. Choose the right maturity level

| Level | Components | Promotion test |
|---|---|---|
| One-off prompt | Operator, source material and temporary instruction | Keep it one-off unless the task repeats or inconsistency has a real cost |
| Reusable prompt asset | Master instruction, variables, examples and output structure | Promote when several people need the same standard or work needs handing over |
| Human-operated workflow | Trigger, input contract, prompt asset, validation, approval, destination and fallback | Promote when volume and stability justify connected hand-offs |
| Automated workflow | Connected systems, deterministic checks, monitored model step and recovery logic | Increase autonomy only after representative testing and measured operation |
| Agentic workflow | Model selects from approved tools or actions within a bounded objective | Require strict tool allow-lists, least privilege, approval gates and continuous monitoring |
Maturity is not a ladder every task must climb
A reusable human-operated workflow may be the best permanent design. Automation is worthwhile only when it improves the complete outcome after setup, review, monitoring and exception handling are counted.
3. Create the input contract
Most prompt problems are input problems. An input contract states what must be provided, which source is authoritative, what format is accepted and which information is prohibited. This makes missing context visible before the model produces a confident but unusable answer.
- Required fields: the minimum information needed to complete the task.
- Source priority: which document, record or dataset wins when sources conflict.
- Freshness: how current the material must be and who confirms it.
- Allowed data: information approved for the chosen account, model and integration.
- Prohibited data: restricted, unnecessary or unapproved information that must not enter the workflow.
- Input states: valid, incomplete, conflicting, sensitive or outside scope.
- Escalation: what the operator does when the contract is not satisfied.
Data minimisation improves control and output quality. Supplying an entire inbox or customer record when the task needs three fields adds privacy risk, irrelevant context and more opportunities for misleading instructions.
4. Build the prompt as a managed asset
A reusable prompt should be stored outside an individual chat history and separated into stable instructions and run-specific variables. The future AI Aurora prompt library will follow this model so each pack functions as part of a real workflow rather than as a disposable paragraph.
| Prompt component | Purpose |
|---|---|
| Role and outcome | States the bounded job and intended reader or destination |
| Required inputs | Names the variables and refuses to proceed when critical inputs are missing |
| Source rules | Distinguishes supplied evidence, approved knowledge and unsupported assumptions |
| Process instructions | Describes the reasoning or transformation steps that matter to the output |
| Constraints | Sets scope, tone, privacy, prohibited actions and length only where useful |
| Output contract | Defines headings, fields, labels, citations, confidence or machine-readable schema |
| Uncertainty route | Requires questions, flags or escalation rather than invented completion |
| Examples | Shows difficult distinctions and the standard of an acceptable output |
| Review checklist | Tells the human what must be verified before approval |
| Version record | Captures owner, version, model assumptions, test date and change note |
Do not hide business rules inside flowing prose when they can be expressed as fields, decision tables or deterministic checks. A shorter prompt connected to a clear input form and output schema often performs better than a long instruction attempting to carry the entire process.
5. Separate deterministic rules from AI judgement
Language models are useful where the task requires interpretation, synthesis, classification or drafting. They are a poor substitute for fixed calculations, exact field mapping, permission checks and other rules that conventional software can perform reliably.
| Step | Preferred method | Example |
|---|---|---|
| Required-field check | Deterministic rule | Reject the run when client name or source document is missing |
| Date or arithmetic | Deterministic function | Calculate a due date from a stored policy |
| Topic classification | Bounded AI judgement with allowed labels | Assign one category and provide a reason |
| Drafting from evidence | AI with source rules and output structure | Prepare a draft using only supplied approved material |
| External send or publication | Human approval plus deterministic action | Reviewer approves, then the system sends or publishes |
| High-impact exception | Human decision | Escalate uncertain eligibility, safety or legal cases |
This separation reduces the number of ways a model error can propagate. It also makes testing clearer: deterministic steps can be checked exactly, while judgement steps are evaluated against examples and acceptance criteria.
6. Define and validate the output contract
An output contract describes what a usable result looks like before the prompt is written. It can be a document structure, a set of fields, an allowed-label list or a JSON schema. Validation occurs before the output is allowed to trigger another step.
- Completeness: required sections or fields are present.
- Structure: values use the expected labels, types and order.
- Evidence: material claims point to supplied or verified sources.
- Scope: the output does not add unauthorised advice or actions.
- Uncertainty: missing evidence and ambiguous cases are visible.
- Safety: prohibited information is not reproduced or exposed.
- Destination readiness: the result can be reviewed and stored without manual reconstruction.
Machine-readable output is not automatically trustworthy. Schema validation confirms structure, not accuracy or appropriateness; material content still needs evidence checks and human judgement.
7. Design a meaningful human review
“Human in the loop” is too vague to be an operating control. Specify who reviews, when review occurs, what evidence is visible and what choices the reviewer can make.
| Review element | Practical definition |
|---|---|
| Reviewer | Named role with enough subject knowledge and authority |
| Timing | Before publication, external communication, commitment or irreversible action |
| Evidence | Source material, model output, uncertainty flags and relevant change history |
| Checklist | Accuracy, completeness, tone, privacy, policy and destination |
| Decision | Approve, edit, reject, request more information or escalate |
| Record | Who approved, when, what changed and where the final version is stored |
The review burden should be measured. A workflow that saves ten minutes of drafting but creates twenty minutes of verification has not improved the process. Better inputs, evidence labels and structured output may reduce review effort more effectively than a different model.
8. Treat exceptions and fallback as normal paths
Workflows encounter incomplete inputs, conflicting sources, unavailable services, changed permissions and outputs that do not meet the standard. A reliable system exposes those cases and routes them safely rather than forcing a result.
- Invalid input: stop and identify the missing or conflicting field.
- Low confidence: route to a reviewer with the uncertainty and source context.
- Service failure: retry only where duplicate actions are prevented; otherwise use the manual route.
- Changed integration: pause connected writes until the mapping and permissions are revalidated.
- Unsafe or out-of-scope request: refuse the automated path and escalate.
- Bad output discovered later: correct the destination record, identify affected cases and update the prompt, rule or test set.
- Complete outage: continue with a documented manual process and reconcile work after restoration.
The NCSC’s secure AI guidance treats operation and maintenance as part of the system lifecycle, including monitoring behaviour and inputs, secure updates and learning from incidents. That principle applies even to a small no-code workflow.
9. Test, measure and version the system
Create a small evaluation set before deployment. Include ordinary cases, incomplete inputs, conflicting evidence, sensitive material, unusual formatting and examples that previously failed. Record the expected result or acceptance criteria for each case.
- Baseline the current process. Measure time, quality, rework and common exceptions without the new system.
- Test the prompt asset. Run the evaluation set and record failures rather than refining against one ideal example.
- Test the whole workflow. Include permissions, destination records, approval and fallback.
- Run assisted. Keep external actions under human control while collecting evidence.
- Version every material component. Prompt, examples, model, source collection, integration mapping and policy rules.
- Re-test after change. A model update, new tool, revised source or prompt edit can change behaviour.
- Review operating metrics. Decide whether to keep, simplify, repair or stop the workflow.
| Metric | What it reveals |
|---|---|
| Acceptance rate | How often the output meets the standard with minor or no revision |
| Major correction rate | Whether the workflow creates risky rework or unsupported confidence |
| Cycle time | Elapsed time from valid trigger to approved destination |
| Review time | The real human cost of verification and correction |
| Exception rate | How often inputs, services or outputs require a different route |
| Failure recovery | Whether outages and retries cause loss, delay or duplication |
| Operating cost | Subscriptions, usage, setup, review, monitoring and maintenance |
| Outcome value | Whether the final business or user result improved |
10. Worked example: meeting notes to approved actions
A small consultancy wants to convert completed meeting notes into a consistent action brief. A weak implementation asks an assistant to “summarise this meeting” and copies the response into a project tool. A system defines the complete path.
| Component | Example design |
|---|---|
| Outcome | An approved action brief with decisions, owners, due dates, open questions and source references |
| Trigger | The meeting owner marks the notes complete and confirms participant consent and storage location |
| Inputs | Approved transcript or notes, project name, participant list, known deadlines and action-owner directory |
| Prompt asset | Extract only explicit decisions and actions; separate proposals from commitments; flag unclear ownership or dates |
| Deterministic rules | Reject missing project name; validate date format; allow only known project owners or “unassigned” |
| Output contract | Meeting summary, decisions, action table, risks, unanswered questions and evidence snippets |
| Human review | Meeting owner checks accuracy, sensitive content, owners, deadlines and anything marked uncertain |
| Destination | Approved brief stored in the project folder; actions created only after approval |
| Exception route | Unclear commitments remain questions; missing transcript returns to the owner; service outage uses the manual template |
| Measurement | Review time, missed actions, correction rate, duplicate tasks and participant feedback |
The first version can be entirely human-operated: paste the approved notes into a stored template, review the result and create tasks manually. Automation becomes useful only after the team has stable fields, reliable owners and a clear rule for ambiguous commitments.
Common failure points
| Failure | Why it happens | Correction |
|---|---|---|
| Prompt-only thinking | The instruction is improved while inputs, ownership and destination remain undefined | Design the workflow around the operational outcome |
| Examples treated as rules | A few demonstrations cannot cover every boundary or prohibited action | Use explicit rules, labels and validation alongside examples |
| Hidden source conflict | The model blends old, new and unofficial information | Set source priority and require conflict flags |
| Unbounded autonomy | The model can call broad tools or take irreversible actions | Reduce permissions, tool choices and action scope; add approval |
| No evaluation set | The team tests only the case used while writing the prompt | Maintain representative and difficult cases with acceptance criteria |
| Silent fallback | Operators invent different recovery methods during failure | Document one manual route and reconciliation procedure |
| No owner | The workflow remains active after its builder leaves or assumptions change | Assign operation, content and technical responsibilities |
| Success measured as output volume | More drafts or tasks hide worsening accuracy and review cost | Measure accepted outcomes, corrections, exceptions and complete cost |
Implementation checklist
- The recurring outcome and unacceptable result are defined.
- The trigger and authorised operator are named.
- Required, optional and prohibited inputs are documented.
- Authoritative sources and freshness rules are clear.
- The prompt is stored as a versioned asset outside personal chat history.
- Fixed rules are handled deterministically where possible.
- The expected output structure and validation checks are defined.
- Uncertainty produces questions or escalation rather than invented completion.
- The human reviewer, checklist and approval options are explicit.
- The approved destination and record owner are named.
- Exceptions, outages, retry behaviour and manual fallback are documented.
- A representative evaluation set and baseline exist.
- Prompt, model, sources and integrations have change records.
- Quality, review time, exceptions, cost and outcome value are measured.
- Autonomy will increase only after the bounded system has performed reliably.
Official sources
- NIST AI Risk Management Framework
- NIST AI RMF Playbook
- NIST Generative AI Profile
- NCSC Guidelines for secure AI system development
- NCSC secure-development guidance
- NCSC secure-operation and maintenance guidance
- ICO AI and data-protection risk toolkit
- OWASP Top 10 for LLM and generative-AI applications
- OWASP prompt-injection guidance
- OWASP excessive-agency guidance
- UK Government AI Management Essentials guidance