The board wants to see the bottom line of the RFP. Show the receipt.
Financial services boards want to know what they get from the spend on AI in LangChain’s July post titled . Finance needs the same owned boundary if it wants per-work-item proof.
LangChain’s finance piece names RFP processing and AML compliance monitoring as agentic use cases, then connects ROI to observability, governance, and economic measurement in the same stack. The operational discipline is making every meaningful run leave a receipt as it executes.
, with durable execution, real-time streaming, and horizontal scaling for agent workloads.
A receipt turns a trace into evidence finance, risk, and engineering can all read.
For Engineering it is debug evidence, for Risk it is authority and data lineage, for Finance it is cost per accepted work item, and for Product it is adoption tied to outcome. This receipt answers different questions for the teams and does not force each to make up its own “truth” for how well things are working.
Honeycomb has made similar observations in the past about Observability. As they put it in , and its useful question is whether the agent’s work item can be replayed, judged, explained, and tied to business results.
Aggregate dashboards are too late
They show spend by team, average latency, total runs, error rates, top agents, cost by model, and adoption trends. These are good to keep and to be proud of, but they are late by design, i.e. they cannot be used to prove the value of a single compliance-sensitive action.
Averages hide important details, especially when it comes to individual expensive runs and costly rework caused by high adoption. The expensive run might be valuable to investigate deeply. The cheap run might be garbage. The agent with high adoption might cause rework, whereas the agent with low adoption is the one doing the riskiest work for the firm.
The receipt is what lets the rollup make sense.
. In addition, there are now more than 10% of users managing three or more agents in current active work at any time, and 26.6% of users using skills in their use of Codex for coding and other knowledge work.
Parallel execution breaks casual measurement.
The runtime needs to know what work item is currently active, which is pending approval, which handled customer data, which triggered follow-up jobs, which exceeded cost bounds, and which have side effects. . In financial-services language that becomes identity, authority, state, approval, evidence, and accountability.
Just as with money movement, . Repeated behavior of an agentic workflow should transform into an observable, doable, legible execution path (and remain agentic as soon as reality changes again). Likewise the ROI of a finance agent does not only grow as it becomes more agentic, it also grows as the agent’s work is first discovered, recorded, measured and then hardened.
And, side-effect layer. As we’ve written about before, side-effect receipts are critical to enabling features like retry, compensation, and ownership tracking when various tools are orchestrating work that causes side effects by running to update live systems. For a finance agent, this would include updating CRM records, requesting documents, opening up new case work, generating client-facing drafts, and opening up new compliance work, etc., all of which would need to be issued with an operation key and receive a corresponding receipt.
A trace in the workflow captures all the activity that occurred within the agent as it executed the workflow. A side-effect receipt, by contrast, captures the changes that the agent caused outside of the workflow, i.e., the actual side-effects.
Both belong in the ledger.
Own the ledger before the board asks
A distributed systems lecture is not what the board wants. Fair enough. They just want to know if their money has been converted into real work value.
To answer that question honestly requires a tremendous amount of engineering discipline to create runtime receipts that are dull enough to be understood by people in finance, risk, and engineering, and to attach all relevant metadata before work begins, to track all changes to authority throughout the process, to ensure that all side effects are recorded as first class facts, to tie the scores from evaluators and the results of human review to the work item in question, and to do rollups from the receipt as opposed to from the vibes that were brought to the work.
This also changes how one evaluates vendors for finance agent platforms. A platform without per-work receipts for finance, risk, and engineering to share is magic; a runtime with receipts is evidence, easier to defend when the model changes, when a workflow drifts, when a regulator asks for lineage, or when the CFO asks why the AI budget went up again.
Show the receipt.
SOCIAL SHARE CARD GENERATOR