Home / Evidence
Evidence
Claims should carry their context.
We publish engineering evidence with the date, scope and definition needed to interpret it honestly — what each record shows, and what it does not.
The evidence ledger
Each entry names its date and scope, what it shows and what it does not. We prefer evidence about the products to figures about ourselves.
| Evidence | Date | Scope | What it shows | What it does not show |
|---|---|---|---|---|
| A Patra processing receipt | 11 September 2026 | One run of the Patra command-line tool on a synthetic document | What was done, by which operation version, how faithful the result is, digests of input and output, where it ran, and that the document was not kept | Performance on real documents, or a hosted service |
| A knowledge graph snapshot | 27 August 2026 | A bounded internal pipeline slice; totals are withheld pending a database backup recount | That the pipeline builds dated, typed graph records | It does not establish full-corpus completion or customer-facing search |
| Execution declarations | July 2026 | Every response of our research API | Which components actually ran, which were missing, and why | That an answer is right: a declaration discloses, it does not grade |
| ChangeProof's fictional-lender run | July 2026 | A fictional lender, and a fictional circular shortening a complaints deadline from 30 days to 21 | Eight affected artefacts found by analysis at run time, approval bound to the exact plan shown, and a hash-addressed proof record | Real regulatory coverage, integrations or a live workspace |
| Karta's security record | Reviewed 25 September 2026 | Security-relevant events in Karta's internal testing | A hash-linked chain whose check reports whether, and where, it breaks | Tamper-proofing: it detects tampering, it cannot prevent it |
| A benchmark we did not score | July 2026 | One external benchmark run against our knowledge base | Why a zero was not a score: the benchmark and our corpus name documents differently | Any retrieval-quality figure: none is published |
—
Evidence we do not have yet
The ledger records the evidence available for each capability and its limits.
—
Evaluation boundary
Engineering evaluations are grounded in a defined corpus, run definition and acceptance boundary, so each result can be understood in context and revisited with confidence.
Ask for the record behind an entry.
Every entry above points to a record we keep. Ask, and we will walk you through it — including what it does not show.