What is the Evidence Vault?
Evidence Vault is the auditable, per-interaction record of what backed each answer an agent gave: the question asked, the context used, the sources consulted, the active version, the applied scope, who was cleared to see it, the result and the traceId that ties it all together. It is the proof of where an answer came from, not the chat history.
That record is not born inside the agent. It comes from the layer that prepares and governs the context before the answer goes out, which is what lets you reconstruct an interaction months later with the same confidence you had when it happened. When someone wants to understand what an answer was based on, a manual reconstruction gives way to a single query.
In practice, the base already operates: the Evidence Log records each interaction with a traceId, scope per collection and the insufficient_context refusal when there is no approved source to answer from. The Evidence Vault consolidates that trail into an evidence package with retention, ready for when audit, security or legal ask for the proof.
Why the agent logs are not enough
The agent log shows the conversation; it does not prove which version of the source and which authorized scope were applied to that answer. You can reread what was said, but you cannot show why it was safe to say it.
For a risk committee, that is the difference that matters. The dialogue records the claim; the audit needs the approved origin, the version that was active at the time and the access level of whoever received the information. Rereading the transcript answers none of those three questions.
Auditing an agent, in practice, means reconstructing the authorized context that backed the answer, not reopening the chat. Without a per-interaction record, the question that defines maturity has no defensible answer: what backed that answer, based on what, authorized by whom? That is exactly the gap the Evidence Vault fills.
What the Evidence Vault records per interaction
Every answer leaves a set of fields that, together, let you rebuild the whole interaction. The full list appears once here; the other sections refer back to it as that record.
Question
What was asked of the agent, exactly as it came in, the starting point of the reconstruction.
Context used
The block of knowledge that actually fed the answer, not everything that was available.
Sources and passages
Which approved documents and passages were consulted to compose the result.
Source version
Which version was active at the time, so an old answer stays explainable after the source changes.
Scope
The collection and the authority that interaction was limited to.
Permission
The access level that decided who could see that content, the who-sees-what enforced at the time.
Agent and execution layer
Which agent answered and in which environment it ran, regardless of the framework you chose.
User and workspace
Who asked the question and in which workspace, to tie the interaction to the right context.
Result
The answer delivered or the insufficient_context refusal when there was no approved source to back it.
Risk and sensitivity
The sensitivity level of the content involved, to prioritize review where exposure is higher.
Human decision
When a person approved, reviewed or stepped into the flow, the record keeps who and when.
Timestamp and traceId
The timestamp and the unique identifier that stitch question, context, sources and result into a single trail.
What happens without an evidence trail
Without a per-interaction record, no one can safely show what backed each answer. The problem does not stay with the technical team: it surfaces at sign-off, when whoever puts their name on it needs proof and finds no trail.
The symptoms below are the ones that block the way out of the pilot.
- Sign-off stalled. Audit, legal and security hold the release to production because there is no way to show origin, version and authority for each answer.
- Proof rebuilt by hand. Every time an answer is challenged, someone has to reconstruct by hand what backed it, spending days on what should be a single query.
- The answer turns into a black box. Once the source has moved to a new version, an old answer stops being explainable and no one knows what it was based on anymore.
- Exposure in a regulated environment. In a regulated sector or a public tender, the lack of a traceable record stops being a technical detail and becomes a compliance risk.
- Agent stuck in the pilot. The case works in the demo and stalls in production because no one feels safe signing the release.
Where Contextfy fits
The evidence comes from the layer that prepares and governs the context, which is why it holds for any agent, regardless of the tool your company chose to run it. When the trail of each interaction sits before the answer, not inside each agent, it stays consistent even if you swap or combine platforms.
Contextfy is that governed-context layer between the company sources and the agents. It organizes knowledge into versioned collections, applies scope and permissions, serves context on demand via API or MCP, and records the evidence each answer consumes. It does not run the agent and does not compete with it: it governs the context and produces the proof. And to be clear about the limit: Contextfy produces the evidence, it does not audit or certify your company.
Fontes
Drive, SharePoint, ERP, CRM, PDFs, APIs
Contextfy · Context Engine
Organiza · versiona · governa · observa o contexto
Runtimes
via MCP · API · conectores · pipelines
The gain for audit, legal and security
Evidence by design shortens the path from pilot to production because audit, legal and security sign-off no longer depends on manual reconstruction. The question that used to stall the project now has an immediate answer, with origin, version and authority on hand.
The gain shows up in records each decision-maker recognizes. For whoever is accountable, less rework every time an answer is challenged. For whoever approves, a shorter internal cycle. For whoever owns risk, lower exposure in regulated environments and public tenders. For the business, more agents cleared for production, which means operational capacity instead of stalled pilots.
Here governance works as an accelerator, not a brake. It is not a compliance cost bolted on afterward; it is the mechanism that makes AI defensible enough to leave the experiment and operate with confidence.
How to start
The starting point is an assessment of what is, or is not, reconstructible today in your agents' answers. Before changing anything, it helps to see where a trail already exists and where the proof still depends on someone rebuilding it by hand.
From there, the path is incremental and without disruption. Organize the sources, scope and permissions of a priority case, connect the layer to what is already live and let each answer start recording evidence by design. Start with one source or one agent, measure the result and only then scale.
That is how the arc closes: what used to block sign-off for lack of proof becomes what unlocks the move from pilot to production. See what is missing for your answers to be auditable.
Frequently asked questions
What is the Evidence Vault?
It is the auditable, per-interaction record of what backed each agent answer: question, context used, sources, version, scope, permission, result and traceId. Today the operational base is the Evidence Log with traceId; the Vault consolidates that trail into an evidence package with retention.
What is the difference between the Evidence Vault and my agent logs?
The log shows the conversation, meaning what was said. The Evidence Vault proves which version of the source and which authorized scope backed that answer. It is the context behind the answer, not just the recorded dialogue.
Does Contextfy audit or certify my company?
No. Contextfy produces the evidence the auditor requires: source, version, scope, permission and traceId. The audit and the certification stay with your audit function or a third party. The layer delivers the material; another party reviews and attests.
What is recorded in each interaction?
Question, context used, sources and passages, source version, scope, permission, agent and execution layer, user and workspace, result (including the insufficient_context refusal), risk, human decision when there is one, timestamp and traceId.
Does the Evidence Vault help with ISO/IEC 42001 and data-protection law?
It helps produce the trail and the evidence those contexts usually require, such as origin, version, scope, permission and per-interaction traceability. Contextfy does not certify or declare compliance; it delivers the material that supports the adequacy your company runs.
Does it work with any agent or platform?
Yes. The evidence comes from the layer that governs the context and is served via API or MCP, so it holds regardless of the agent platform your company chooses. Contextfy does not replace the tool that runs the agent; it governs the context that tool consumes.
Keep exploring
Assessment of what is already reconstructible today in your agents' answers.
See what is missing for my answers to be auditable