Beyond Plausible
What it takes to prove a pricing or risk number when AI is in the workflow
Any pricing or risk number that enters a bank’s book of record must clear a specific bar. It needs to be reproducible to the decimal place, traceable to every input and code version, explainable to a reviewer, and validated against market practice. That bar predates artificial intelligence by decades and remains unchanged as AI becomes part of the workflow.
In this paper, six senior Numerix practitioners explain what that standard means when AI enters the picture. They examine why a Monte Carlo engine can count as deterministic when a language model does not, how validation evidence is becoming executable rather than written, and what controls are needed to control the risk of a confidently wrong answer.
Read practical insights on:
- The four-part standard a pricing calculation must meet to enter a book of record, and how it applies with AI in the workflow
- Why a Monte Carlo engine counts as deterministic, with a fixed random seed and an estimated, stable residual error, even though the pricing theory underneath it is framed in probability distributions
- Why a language model cannot provide guarantees of reproducibility and error estimation, even when its output can be made repeatable
- How validation is moving from a written report to executable evidence, or “testumentation,” a test harness that reruns against current market data and records its own lineage
- Four controls for bounding a confidently wrong answer, including redundancy across independent agents, a human gate before the system of record, evidence in every explanation, and a kill switch
“A pricing calculation is not considered correct simply because it produces a number, even if it is a plausible number.”
- Ping Sun, PhD, Senior Vice President and Head of Quantitative Research, Numerix
In this series…
This white paper series, Trust, Verified, follows the concept of trust in AI through the three places it needs to be won in capital markets. The first paper, Trust by Task, maps where AI reasoning has earned its place at the practitioner’s desk. The second paper, Beyond Plausible, sets out what it takes to prove a pricing or risk number when AI is in the workflow. The third, From Chat Box to Control Tower, then looks at how firms are composing agents around their analytics, and the architecture and controls required to deploy them responsibly.
Frequently Asked Questions
1. What is this paper about?
This paper examines what it takes to prove a pricing or risk number is correct once AI enters the workflow around it. It draws on interviews with six senior Numerix practitioners, covering the standard a number of record must meet, why a Monte Carlo engine satisfies it, and how validation evidence is becoming executable rather than written.
2. Does the paper argue that AI can generate a validated pricing or risk number on its own?
No, the practitioners keep generative AI out of the calculation itself and use it only in the surrounding workflow, translating term sheets, drafting scenarios, and explaining moves. The number of record still comes from a deterministic engine whose output can be checked against a stated mathematical framework.
3. What four things must be true of a number before it enters a book of record?
It must be reproducible to the decimal place, traceable to every input and code version that produced it, explainable in terms a reviewer will accept, and sufficiently correct through a validation process that can bear scrutiny.
4. Why does a Monte Carlo engine count as deterministic when it runs on random numbers?
The random number generator's seed is fixed, and with the code version and numerical settings held fixed too, the same product on the same market data produces the same number every time. The fixed seed also holds the residual simulation error stable, so it can be estimated and monitored.
5. Why doesn't a language model have the same properties?
Its output can be made repeatable, which supplies reproducibility. But there is no equivalent of a convergence study or an arbitrage-free framework to check its answer against, so reliability can only be estimated as a failure rate across many answers, not as an error bound on one.
6. What is “testumentation”?
Russell Goyder's term for a validation report that is executable rather than written, mixing explanation, graphics and formulae with the code that produces the results, so the report can be rerun against current market data on demand. It shifts the validator's role toward reviewing evidence rather than writing it.
7. What role can AI play in validation itself?
The paper describes AI proposing test cases, benchmarking against reference models, identifying edge cases, and generating documentation, including a case where a frontier model investigated and corrected its own result when a number looked wrong. A person still reviews the evidence before it is accepted.
8. How do practitioners bound the risk of a confidently wrong AI answer?
Four controls are needed, including redundancy and reconciliation across independent agents or models, a release gate that keeps AI output out of the system of record until reviewed, evidence attached to every explanation, and visibility with an off switch for anything that behaves unexpectedly.
9. What do current regulations say about generative AI in model validation?
The April 2026 US model risk guidance keeps its principles for traditional quantitative models but says generative and agentic AI falls outside its scope for now. In Germany, BaFin became the AI market surveillance authority for regulated financial activities in July 2026, with high-risk obligations applying from December 2027.
10. Do the practitioners expect the validation standard to change?
No, they expect the standard, reproducible, traceable, explainable and validated, to hold, pointing to the slow adoption of methods like deep hedging as evidence that explainability requirements do not relax for new techniques. What will change is the work of meeting the standard, as AI takes on more drafting and benchmarking.