The difference between a prompt output and a deliverable is not the words. It is whether someone else can check the claim. This article describes the working method: what gets captured, how provenance attaches, and where the gates are.
What a proof pack contains
This article follows its own standard: the acceptance artefacts behind this site's own builds (browser screenshots, console and network captures, and reconciliation outputs under the repo's output/acceptance runs) are the raw-evidence layer for the claims here.
Raw evidence is the unedited capture: the browser recording, the network log, the data export. Provenance attaches source, timestamp, and runner to every claim, so "according to the GA4 export from Tuesday" is checkable rather than rhetorical. The deliverable stays editable — a reviewer can change the wording without breaking the trail. The uncertainty statement says what is not proven and what would change the answer. Most "AI delivered it" failures are missing exactly this last layer.
From claim to verified deliverable
The before and after
Where reconciliation fits
When two sources disagree — platform-reported versus backend-measured, observed versus configured — reconciliation is the step that measures the gap and explains it rather than picking a winner. That is the difference between an evidence pack and a screenshot with a narrative.
Where to go next
- AI for agency operations — the workflows this standard governs
- Why agent systems become slow, expensive and fragile — why "it ran successfully" is not evidence