Contents

Chapter 3

The Demonstration

At 2.00 on Thursday 13 August, four days before the case review began, Daniel checked the shared folder one last time. It held a fictional customer, a fictional service interruption and two procedure extracts copied into a test pack with identifying details removed. Nothing in it described a live customer. The supplier had agreed to use that pack instead of connecting to company systems.

This was the demonstration booked in Chapter 1. Maya had nearly cancelled it. Daniel had argued that they could learn from seeing the product without committing to buy it. As he opened the meeting, the service diagnosis and its later repair were still ahead of them.

The presenter entered the fictional customer’s question. A polished answer appeared. It recognised the service interruption, expressed regret and explained that the customer was not eligible for a credit. The sentences were clear. The answer sounded like something an experienced adviser might send after checking the policy.

Priya looked at the test pack. “That’s the old rule.”

The presenter said the system could be instructed to prefer current documents. Daniel asked him to leave the answer on the screen. Before changing the instruction, they wanted to see what it had used.

A small reference opened the older procedure. Daniel leaned forward. “Can you keep that visible beside the answer?”

The presenter could. Priya then asked the question that would shape the later design: “Does it show me which rule it used and when it changed?”

Separate producing an answer from establishing it

The product had done something useful and something wrong in the same response. It had turned a short question into readable text and offered a route to the source. It had also answered from a superseded procedure. Calling the whole demonstration a success would hide the error. Calling it useless would discard the source display that had made the error easier to find.

Daniel asked the presenter to expand the reference. The clause matched the answer. That established a narrower fact than anyone first assumed. The output was supported by a document in the pack, but the document was not the applicable rule. A valid citation to an obsolete instruction remained the wrong basis for the customer’s decision.

Leila wrote three questions separately: did the output follow the cited clause; was that clause current for this case; and did the complete customer record justify using it? The demonstration could help examine the first two. Its fictional customer could not establish the third for any real case.

This distinction would matter after the diagnosis. The immediate design idea became a visible clause and effective date, with an adviser deciding applicability. The team had started the meeting expecting to see an assistant compose replies. They left interested in a narrower capability that helped a person inspect the evidence.

Generative text can be fluent while containing false statements or failing to follow its source. A confident tone is part of the output, not proof that the underlying claim has been checked. Research on generated language examines this problem across tasks, including summarisation and question answering.[1] In this demonstration, the team did not have to debate how often that happened elsewhere. They had a saved answer and the wrong procedure beside it.

Ask what the system is producing

On the shared whiteboard, Daniel drew two company examples. One was a proposed forecast of next week’s service demand from earlier volumes and calendar information. Its output would be a number, perhaps accompanied by a range. The other was the demonstration’s written response to a customer’s question. Its output was newly composed text.

He used “predictive” for the first purpose and “generative” for the second. These were practical distinctions for this conversation, not sealed categories. A language generator also predicts the next part of its output. The important question was what the organisation would use the result to do.[2]

If the forecast were wrong, Leila might roster too few advisers. If the written explanation were wrong, a customer might receive an incorrect account of entitlement. Both required evaluation, but an attractive paragraph could not validate a staffing forecast, and a good forecast could not validate the paragraph.

The CFO, joining for the commercial part of the call, asked whether the supplier’s “AI platform” covered both. The presenter said it could support several applications. Daniel brought the discussion back to the proposed task. What evidence existed for finding a source clause in their procedure library? A broad product category did not answer that question.

For your own proposal, complete one ordinary sentence: the system will produce this output for this person to use in this decision. If the sentence contains several decisions, separate them. A system that forecasts demand, drafts replies and approves credits has several jobs, even if one interface presents them together.

For the fictional company, the demand forecast remained an example on a whiteboard. Nobody added it to the register as a deployed forecasting tool or counted its benefit in the service case. No one had authorised it to run. Discussing a possibility did not mean the company had deployed it.

Notice where variation enters

The presenter repeated the same question. The second answer was shorter and used different wording. It still relied on the old procedure. Priya asked whether an adviser could assume the answer would always be the same after the pack was corrected.

“No,” Daniel said. “We have to specify and test the behaviour we need.”

Language generation involves selecting successive pieces of text from model-produced possibilities. Generation settings can change how those selections are made; sampling can produce different wording on repeated runs.[3] A repeated prompt can also reach different source material if the library or retrieval configuration changes. Keep the model, settings and source version in the test record instead of treating the visible question as the whole test.

The presenter offered to make the output more consistent. Leila asked what consistency would prove. If a system repeated the same wrong rule every time, the service problem would remain. If it offered several correct phrasings of the same applicable clause, some variation might be harmless. They needed to distinguish variation in expression from variation in the decision-relevant content.

They saved both outputs. Priya marked the rule used, the date displayed and whether the answer omitted the exception. Those were the features they would later compare. A screenshot of one good run would not reveal whether another run selected an incompatible rule.

Do not insist that every AI output be identical merely because ordinary software sometimes is. Decide what must remain stable for the task. For this use, the cited source identity, applicability information and absence of unauthorised action mattered more than whether the explanation began with “under the procedure” or “according to the procedure”.

Where reproducibility matters, retain enough of the test context to investigate a difference. Do not assume a provider can reconstruct it later from a vague description of what the adviser remembers seeing. That is a practical record requirement, not a promise of perfect repeatability.

Follow the sources through the system

The presenter described the product as retrieval-augmented generation. Daniel translated the phrase: it found passages in a connected collection, then supplied them to a language model to help compose the answer. That combination has a research basis, but connecting retrieval does not by itself establish that every returned answer is current or suitable for a particular decision.[4]

The test pack made the stages visible. Both procedures entered the collection. A retrieval step selected material. The model produced a response using what it received. The interface displayed an answer and a reference. The organisation still had to decide which procedure applied and whether the reference supported the claim being made.

FIG-3.1

Inspect each step before trusting the answer

  1. Source collection
    Which version and effective date? Conflict or omission can enter here.
  2. Retrieval
    Which passage was selected? Relevant wording may still be the wrong applicable rule.
  3. Generated answer
    Does the source support it? Fluent text can overstate or misrepresent the supplied material.
  4. Adviser checks
    Open the source and check scope. Unresolved applicability goes to the existing escalation route.
A citation can point to a real document that is wrong for the decision. Source: Fictional case and author’s design questions; retrieval mechanism: Lewis et al. (2020), reference R21.

Priya asked the presenter to remove the older procedure from the active test collection. On the next run, the system used the newer clause. That was a useful observation about this configuration. It was not evidence that the product would identify superseded instructions in the company’s full library without someone maintaining their status.

Leila asked what happened when a document had a new upload date but an old effective date. The presenter paused. The product could display supplied metadata, but someone had to define which field meant what. The newest file was not necessarily the rule that governed an earlier event.

In their example, a customer’s entitlement could depend on the date of the service interruption. The applicable procedure might therefore be an earlier version retained for a legitimate reason. Removing every old file from view would make current search look cleaner while making historical decisions harder to explain.

The eventual requirement was more precise than “use the latest document”. The system should reveal the source, its effective period and its relationship to other versions. If applicability could not be established, the adviser should see the uncertainty rather than a silently chosen winner. The procedure owner, not the model’s prose, would determine which rule was authoritative.

The team also separated publication from indexing. A corrected document in the library did not prove that the retrieval collection had already been refreshed. Daniel asked for a way to check which version the configured feature actually returned. That check would become part of maintaining the service, rather than a one-time purchase question.

These were requirements generated by the fictional test, not claims that the supplier’s product already met them. Daniel placed each beside an observed result or an unanswered question. The distinction prevented the requirements list from becoming an accidental assurance statement.

Make an accuracy figure describe its test

The presenter moved to a slide reporting 96 per cent accuracy on a supplier test set. The figure belongs to this invented demonstration; it is not a statistic about a real provider. Maya asked what had been counted as correct.

A response counted as correct if the evaluator judged that it answered the supplied question using the reference answer. Daniel asked whether the set included conflicting procedures, missing exceptions or a question for which no answer existed. The presenter did not have that breakdown in the meeting. He agreed to provide the test description rather than improvise one.

Nadia would later classify the 96 per cent as a reported claim. It could help them formulate questions. It could not be copied into the board paper as the company’s expected performance. The denominator, task mix, judging method, product version and consequences of the remaining errors were still unknown to them.

Priya asked whether a correct refusal to answer counted as a success. For her work, it might be the safest and most useful response when the source did not settle the question. A score that rewarded answering everything could favour behaviour the company did not want.

The supplier was not wrong to measure its product. The team was asking whether that measurement addressed their decision. A benchmark may demonstrate capability on its stated tasks. The company still needs evidence for its own material conditions and for the work surrounding the output. Success on one task need not transfer to a similar-looking task.[5]

Use the demonstration to find what should enter the later evaluation. Include ordinary work, important exceptions, absent information and cases where the appropriate answer is to seek help. Separate errors that are inconvenient from errors that change a person’s entitlement. Do not average them into a single reassuring number before deciding how the system may be used.

Distinguish a tested score from a persuasive sentence

The presenter showed another feature: a confidence indicator. It was described as a score for the result. The CFO asked whether the company could automatically accept results above a threshold and send the rest to a person.

Daniel asked what the score measured. Was it similarity between a question and a passage, a classifier’s estimate, or another calculation? The label did not tell them. A retrieval similarity score could be useful for ranking passages without being a probability that an eligibility decision was correct.

A tested decision score, as this book uses the term, is a numerical output whose relationship to a specified outcome has been evaluated on relevant cases. “Tested” does not mean infallible or automatically transferable. The record must say what was scored, against which outcomes, on which cases and for which proposed action.

Research on classification models has found that their confidence estimates can be poorly calibrated: the number reported need not match the observed likelihood of correctness.[6] That finding concerns numerical model estimates. It does not make a generated sentence such as “I am very confident” a measured probability.

FIG-3.2

Ask what the score actually means

Generated certainty

“I am confident” is output text. It does not show that the claim was checked.

Tested decision score

A defined numerical output compared with known outcomes on relevant cases. Record the test scope and limits.

Do not confuse the labels

Retrieval similarity, a classifier score and verbal certainty are different things.

Keep the decision visible

Choose a threshold with error consequences and review capacity in view. Retest when relevant conditions change.

A confident sentence is not a measured probability; a numerical score needs task-specific evaluation. Source: Author’s explanation; calibration distinction: Guo et al. (2017), reference R15.

Daniel offered a separate illustration. Suppose a classifier assigned scores between zero and one to a defined prediction. Group its predictions by score and compare them with known outcomes on cases not used to fit the model. If a group near 0.8 is correct only about half the time, the number should not be presented as an 80 per cent chance of correctness. This is an illustrative explanation, not a result from the company’s tool.

Even a well-calibrated score would leave a decision to make. A threshold trades different mistakes and different amounts of work for people. A credit wrongly withheld may matter differently from a case unnecessarily sent for review. The organisation has to judge those consequences and the capacity of the review route, rather than ask a number to make that judgement on its own.

Maya wrote “no automatic eligibility decision” beside the demonstration record. They had not defined or evaluated a score for that use. The immediate retrieval question did not require one. An adviser could inspect a clause without the company granting the product authority to approve or decline money.

Chapter 8 will return to score behaviour and monitoring. The point here is sufficient for a buying decision: ask what the number means and where that meaning was established. If the answer is missing, preserve the number as an unexplained output, not a control.

Trace what an agent could reach

The next screen offered an agent workflow. The presenter showed a fictional sequence that could read a case, draft a response and call a function to update a record. The company’s test environment had no live connections, so the demonstration could not alter an actual customer account.

An agent can use a model within a sequence that selects and invokes tools, observes results and continues working. Research systems have demonstrated such combinations of model-generated reasoning and actions with external interfaces.[7] The business consequence depends on the tools and permissions supplied. A text-only suggestion and a connected record update therefore need different operating boundaries.

Leila asked what happened if the first interpretation was wrong and the workflow continued. Daniel drew the proposed connections: procedure library, customer record, messaging service and credit transaction. He marked where the system would read, where it could write and where a result would leave the company.

FIG-3.3

Mark where action can be stopped

  1. Read approved sources
    Limit access to the task. STOP if the source boundary fails.
  2. Propose an action
    Show the exact proposed change. STOP before sending or writing.
  3. Authorised action only
    Check permission at execution. Record completed and failed steps.
  4. Recover after a mistake
    Pause further actions. Correct with an authorised process. A sent message cannot be unsent.
Stop before the consequential action; reversal afterwards needs its own authority and may be incomplete. Source: Author’s proposed boundaries; model and tool interaction: Yao et al. (2023), reference R22.

For this project, there was no reason to connect all four. Reading an approved procedure collection could answer the immediate question. The customer system and payment function added authority without helping establish whether source retrieval was useful.

The presenter said those connections could remain disabled. Daniel asked how they would verify that in the selected configuration and who could enable them later. An assurance that the team “would not use” a connection was weaker than a configuration in which it was unavailable to the account doing the work.

They also distinguished stopping from reversing. Preventing the next action would not undo a message already sent. Correcting a record might require preserving the mistaken entry and adding an authorised correction. Reversing a credit could itself harm a customer and require a separate decision. “Undo” on a product slide did not settle those responsibilities.

If a later proposal introduced spending, the team would need to specify the amount, purpose, permitted recipients and person who could stop further commitments. For this comparison, the simpler rule was no purchasing connection and no spending authority. A capability they did not need was not included merely because it came with the platform.

An agent’s action history would also need to be understandable to someone investigating a failure. The team would want to distinguish a proposed action, an approved action, a completed action and an attempt that had failed. That requirement came from their own need to reconstruct a decision. A long stream of generated reasoning would not replace the record of what actually happened.

Check the software you already have

As the meeting ended, Leila asked whether they were discussing a new product because the company lacked the capability or because this supplier had offered the first demonstration. Daniel opened the existing software inventory. Its knowledge-base system had a retrieval feature that had not yet been assessed for this purpose.

He did not enable it in the meeting. Being included in an existing subscription would not establish permission for customer data or prove that its settings matched the proposed use. But it belonged in the route comparison. The demonstration had taught them what to ask of it: source visibility, effective dates and a way to decline an unsupported answer.

This was separate from the approved summarising feature used by the finance analyst in the Introduction. Her permitted activity could continue. There was no reason to bring every checked internal summary into the new project because another feature needed investigation.

The provisional register still contained reports awaiting verification. The outstanding feature checks in two existing systems were not magically completed by noticing an option on a screen. Daniel added the knowledge-base feature to the discovery record with its status: identified, configuration and terms not yet checked for service retrieval. He would reconcile it against the existing outstanding items rather than increase the count without checking overlap.

Inventory work should describe actual features, settings, purposes and permissions. A product name alone cannot distinguish an approved use from a new connection. Conversely, a new marketing label should not make a familiar, authorised use appear to require an entirely new governance project when its actual scope has not changed.

Save the demonstration as evidence of limited things

At 3.25, Maya asked Daniel to send a short record rather than the supplier’s slide deck. He kept the original wrong answer, the second run, the source extracts and the altered test-pack version. The supplier’s proposed requirements and its reported accuracy figure were recorded separately.

The observed facts were narrow: this configuration generated a fluent answer from the older procedure; it could expose the cited clause; removing the old procedure from the active pack changed the source used on the observed rerun. These facts had sources in the saved session record. They did not prove performance across the live library.

Reported information included the 96 per cent claim and the supplier’s account of features not exercised. The working assumption was that seeing the clause and date could make an adviser’s check easier. That assumption deserved a comparison, not immediate treatment as a benefit. Unanswered questions included applicability, permissions, update behaviour, failure handling and performance against existing search.

Daniel had been right to retain the session. The error had taught them more than a clean sales demonstration would have. It exposed an information requirement, and the source display offered a practical design idea. Maya wrote that in the record too. A colleague should not have to be entirely right about the product to be right about the value of learning from it.

Before closing the file, Priya checked the session notes against the saved outputs. One sentence said the product had found the wrong rule because it could not understand dates. The observed run did not establish that cause. The old rule might have been retrieved because of the test configuration, metadata or wording. They replaced the sentence with the observed behaviour and left the cause open for investigation.

That correction protected both sides. It prevented a weak explanation from becoming a permanent judgement about the supplier, and it stopped the team assuming that a date instruction alone would fix the problem. A demonstration can reveal a failure without revealing every reason for it. Preserve that difference when turning observations into requirements.

Complete the Demonstration Question Sheet

The worked sheet named the task as locating the applicable service-credit clause for an adviser. Its test material was the fictional pack used on 13 August. Its requested outputs were a clause, its identity, effective-date information and an explicit indication when applicability remained unresolved. No live records, customer messages or financial actions were permitted.

The sheet asked the supplier to show the source and the wrong case before explaining the headline score. It asked what changed between runs, which functions could act outside the interface, and what remained untested. Each answer received an evidence class and a reference to the saved observation or supplier document.

The next step was to carry the source-display idea into diagnosis and the later route comparison. It was not to purchase the demonstrated product. When the team completed its diagnosis on 28 August, this earlier record would help it specify the narrower retrieval question.

Use Appendix C before the next demonstration you attend. Write the job you want to observe and the actions the session may take. Bring a case that could expose a material mistake. Leave with a record of what happened, what was only claimed and what the next decision still needs. A useful demonstration changes the question you ask afterwards.