Trust by Task
Where AI reasoning has earned trust in quantitative work, and where gaps remain
AI has quickly moved from experimentation to practical use in quantitative finance. Trust, however, depends on the task, the cost of failure, and whether the result can be verified.
In this paper, six senior Numerix practitioners map where AI reasoning has earned their trust, where it remains conditional, and where they still withhold it for structural reasons. Drawing on real examples of both successful and failed AI-assisted work, the paper examines what practitioners are learning about using AI effectively without surrendering the controls that quantitative finance demands.
Read practical insights on:
- Where AI earns trust today, from research, prototyping and environment setup to integration, testing and documentation
- Where the boundary holds, as numbers that must be validated, reproduced and audited still need a deterministic, controlled computational foundation, even when AI supports the surrounding work
- What successful AI-assisted work looks like, including examples of AI -inferring analytics architecture, building working implementations and supporting model-selection decisions
- Four real failures from Numerix's own practitioners, including a hallucinated cash flow, an unstable benchmark and a weakened unit test that kept passing as performance deteriorated
- The warning signs and working practices that matter, including keeping human gates, asking for assumptions, matching AI use to the function and building trust in small steps.
“Something deterministic has to be the foundation on top of which AI reasoning can add value.”
- Satyam Kancharla, Chief Product Officer, Numerix
In this series…
This white paper series, Trust, Verified, follows the concept of trust in AI through the three places it needs to be won in capital markets. The first paper, Trust by Task, maps where AI reasoning has earned its place at the practitioner’s desk. The second paper, Beyond Plausible, sets out what it takes to prove a pricing or risk number when AI is in the workflow. The third, From Chat Box to Control Tower, then looks at how firms are composing agents around their analytics, and the architecture and controls required to deploy them responsibly.
Frequently Asked Questions
1. What is this paper about?
This paper examines where AI reasoning can be trusted in day-to-day quantitative work and where practitioners still require human judgment, deterministic validation or both. It is based on interviews with six senior Numerix practitioners who use AI tools in their work.
2. Is the paper arguing that AI can or cannot be trusted?
Neither. The practitioners interviewed for the paper do not treat trust as a verdict on AI as a whole. Instead, they evaluate trust task by task, based on what the task requires and how the result can be checked.
3. Which quantitative tasks have the highest level of trust today?
The paper identifies environment setup, dependencies and scaffolding, literature review, prototyping, tests and documentation, and integration work as areas where practitioners report high or rising trust, subject to appropriate review.
4. Where does trust remain conditional?
Practitioners remain more cautious around model-selection guidance, calibration constraints, edge cases, unusual model details and other tasks where quality depends heavily on context and expert judgment.
5. Why is the production pricing engine treated differently?
The paper argues that implementing or modifying a production pricing engine remains a low-trust task because the implementation record is not public and correctness, numerical stability, performance and validation require deterministic checks.
6. Can AI produce a correct quantitative implementation?
Yes. The paper describes successful examples, including an exposure implementation produced from documentation that subsequently passed validation. It also describes AI generating prototypes, applications and model-selection recommendations. The common feature is that successful outputs remain subject to appropriate human review and validation.
7. How can AI produce a convincing but incorrect result?
The paper documents several failure modes, including hallucinated cash flows, oscillation between competing implementations, weakened tests and foundational assumptions that the tool failed to recognize as incorrect. These errors can remain hidden when only the headline result is examined.
8. What warning signs should practitioners watch for?
Warning signs include repeated attempts to fix a problem without resolving the underlying issue, oscillation between two concepts, a model that works in a special case but fails in the general case, tests that continue passing as the problem grows, headline numbers supported by incorrect components, and confident explanations without evidence.
9. What practices does the paper recommend for using AI in quantitative work?
The practitioners recommend keeping a human gate at every stage, asking for assumptions before analysis, matching AI adoption to the function, building trust in small steps, evaluating model releases deliberately and investing in context through tools such as instruction files, skills and retrieval.
10. Where do the practitioners expect the boundary to move next?
The practitioners expect AI to take on more of the work surrounding quantitative analytics, including integration and development workflows. At the same time, they expect the deterministic analytics library to remain the trusted computational core and correctness, reproducibility, robustness and traceability to remain non-negotiable.