# Quiver — Technical Documentation — part 6 of 7

> The complete document, split only because a single fetch truncates. **Nothing is abridged**: the
> parts below concatenate to the whole text, cut at section boundaries. Every part is served as plain
> markdown.
>
> - **Part 1** — `/paper/1` — Abstract  |  At a Glance  |  Contents  |  1. Introduction  |  2. System Architecture  |  3. Design Principles
> - **Part 2** — `/paper/2` — 4. Service Catalogue
> - **Part 3** — `/paper/3` — 5. Methodology
> - **Part 4** — `/paper/4` — 6. Verification and Testing  |  7. Worked Walkthrough  |  8. Limitations and Honest Disclosures
> - **Part 5** — `/paper/5` — 9. Related Work and Positioning  |  10. The Build: Story, Process, and a User Scenario
> - **Part 6** — `/paper/6` — 11. Roadmap: Operating Quiver After the Hackathon  |  12. Conclusion
> - **Part 7** — `/paper/7` — Appendix A · API Reference  |  Appendix B · Reproducibility  |  Appendix C · Checkable Artifacts  |  References
>
> Whole document in one response (247 kB, may truncate in your client): `/paper/full`
> Typeset edition with figures: `/paper`
> Live service: https://quiver-production-c3a8.up.railway.app · Source: https://github.com/Tristan-tech-ai/Quiver

---

## 11. Roadmap: Operating Quiver After the Hackathon

A hackathon submission that ends at the deadline is a demo. This section states what happens next, what would falsify the plan, and what is committed regardless of any competition result — because a service an agent is asked to depend on has to say how long it intends to exist.


### 11.1 Standing commitments

Three things continue independent of any outcome. The endpoint stays live at its registered address, because agents discover Quiver through an on-chain registry entry and breaking that entry breaks every integration built against it. The availability record stays public and externally measured, including the outages: a reliability claim a reader cannot check is not a claim. And the repository stays open under its current licence with the research scripts intact, so every number in this document remains reproducible after the event that prompted it.


### 11.2 The single metric that governs the plan

Quiver's next milestone is deliberately not revenue, usage counts, or service breadth. It is **recurrence by callers who are not us**: distinct external agents calling more than once, because a second call is the only signal that an answer was worth its price in a real loop. Usage without recurrence measures curiosity; revenue without recurrence measures novelty. The desk currently instruments this directly — distinct payers and repeat rate per payer — and reports it honestly, including the fact that the great majority of calls to date are the operator's own disclosed quality-assurance traffic and are never counted as sales. Until recurrence is real, additional building is the wrong response, and the plan says so in advance so the temptation is on the record. Since the metric governs the plan, it is reported rather than described. As of this writing, and measured *on chain* rather than from our own counter, six payer addresses that are not ours have paid the service. Over the eight days to 27 July 2026 they sent **44 payments** totalling **0.575 USD₮0** to the `payTo` advertised in every 402 challenge, at per-call amounts between 0.0096 and 0.016 — inside the published price band. **Four of the six paid more than once**, which is the metric as this section defines it. One returned across 2.55 days (25 to 27 July); another made twelve calls on 20 July and came back a week later. Half a dollar is not a business, and it is reported here because the metric governs the plan and the plan is worth nothing if the number is flattered.

An earlier draft of this section said *three* wallets, one or two calls each, and external recurrence *zero*. That was wrong twice over. The conclusion tested for a *third* call while the definition two paragraphs above says *more than once*; and the count itself came from an in-memory instrument that resets on every deploy — this service deployed eight times on 27 July alone. The chain does not forget and our own instrumentation did, in a document whose entire argument is that a number you cannot re-derive is a number you should not trust. The addresses are `0x1b010a9c…` (22 payments), `0xbc59eb75…` (12), `0xc385e2df…` (5), `0xcab2b9e3…` (3), `0x8d295ff5…` (1) and `0x86f10e00…` (1); a reader recomputes the whole figure from the USD₮0 transfer log on X Layer over blocks 65,711,861 to 66,403,061, which needs no cooperation from us. One of roughly seventy scan windows failed and was not retried, so treat 44 as a floor. That is against 568 calls from our own audit desk over the same instrumented window. One qualification belongs with that, and it is ours: those calls predate the acceptance-rule fix described in Section 11.5, and we have not verified on chain that their settlements landed. A facilitator response carrying a transaction hash does land — sampled 5/5 with receipt status 0x1 — but one without a hash does not, and on 27 July seven such responses produced zero transfers to the payTo address. So what is reported here is *calls*, not *revenue*: an external caller returned, and whether the money followed is something we can no longer assert from the facilitator's word alone. That is the number, and by the standard set two sentences ago it is the only number in this document that would justify building anything further.


### 11.3 Sequence

**Near term (to roughly September 2026) — distribution over construction.** The engine is ahead of its adoption, so the work is meeting agent builders where they already are rather than adding services. Quiver is already an MCP server carrying the risk-brain tools with a free fair-use tier, listed on the official MCP registry, and shipped with framework snippets for the common agent stacks; the near-term work is getting it in front of builders and publishing the correctness story in the terms that matter to them. The reliability item sequenced here is the custom domain and a second region, which together make redundancy possible and retire the disclosure in Section 8 — a registry update, planned deliberately rather than executed under deadline pressure.

**Then (roughly September to December 2026) — prove the number is worth a real price.** Sub-cent per-call pricing is right for discovery and wrong as a revenue engine: the dollar volume in agent payments sits at transactions above a dollar, not at a hundredth of one. The step is bundle and subscription pricing for high-loop callers, alongside an embeddable form of the engine so that a risk platform or treasury product can call these numbers inside its own product rather than sending its users elsewhere. The portfolio layer — cross-venue true exposure and nearest liquidation, whose self-check is reconciliation against chain truth — is the service most asked for and the one whose proof is strongest.

**Then (into 2027) — attestation where it carries liability.** The natural end state of a proof-carrying answer is an attestation a third party can consume in an audit: typed, signed, revocable, anchored on chain. That is a compliance artefact rather than a mathematical advance, and its verifier is external — an attestation actually relied upon in a real decision or audit. Deeper venue coverage and any trustless-execution work (zero-knowledge or enclave-based) are gated behind a counterparty that genuinely requires them, not built speculatively.


### 11.4 Known gaps that are engineering work, not limits

Section 8 separates what cannot be fixed from what is merely unfinished. This is the unfinished list, stated as commitments with a definition of done, so that a reader can hold the plan to it rather than take a promise on faith. Item 4 has since shipped and is left in place, rewritten as what it became, because a commitments list that quietly deletes the commitments it met teaches a reader nothing about whether to believe the rest. Everything else here is still unfinished, and each is disclosed in the relevant service output today.

1. **Vanna-volga correction for one-touch barriers.** The service currently prices barriers under a single volatility and publishes the model-uncertainty span. The correction is a closed-form overhedge from three vanilla quotes already present in the fitted smile, so this is arithmetic we have not yet written rather than a modelling barrier. *Done means*: barrier prices carry a vanna-volga-adjusted figure alongside the single-volatility one, the difference is reported, and a self-check asserts the correction vanishes as the smile flattens.
2. **Live macro calendar with the curated table as fallback.** *Done means*: releases are fetched from primary sources, the response states which source answered and how old it is, and a failed fetch falls back to the transcribed table and says so rather than reporting a clear window it cannot certify.
3. **Deeper tape coverage where the venue allows it.** The sampled-feed limitation in Section 8 is imposed by the exchange endpoint, but the current implementation does not exhaust the pagination the endpoint does offer. *Done means*: the tape is walked as deep as the venue permits, and the density diagnostic reports coverage achieved against coverage available, so a caller can see the remaining gap instead of inferring it.
4. **A succinct proof a verifier that cannot run Node can check — shipped on 28 July 2026, and the three limits this entry used to list are closed rather than restated.** Re-execution already makes these answers verifiable: clone, check the code identity, re-run, compare. A proof system does not strengthen that and this document does not claim it does. What re-execution needs is a *runtime*, and a smart contract has none — so a contract consuming a Quiver number could previously check only a signature, which means trusting the signer rather than the arithmetic. That is now closed for the flagship computation, end to end. A caller adds `"snark": true` to a `perp-gate` request; the answer returns unchanged, carrying a retrieval URL, and a free `GET /proof/<contentHash>` returns a Plonk proof that the served liquidation price satisfies the liquidation identity for that exact position. `QuiverProofRegistry` at `0xd50A91E36673443749Ee22031cb2Ff09d4Bb8D60` on X Layer hands that proof to the verifier at `0x59F6Aa860eE0d26Db873f7c7015CE869170b3b25` and records what survived: one transaction accepting a proof **bought from the live endpoint** (`0x50397d71…f368a`, block 66,412,787, 468,459 gas, `ProofAccepted`) and one rejecting the same proof with the certified price moved a single grid step — one part in 10⁹ (`0x97502c78…4aac`, block 66,412,794, 333,155 gas, `ProofRejected`). Read back from the chain by a script that did not deploy it, the registry holds `58329.113924051` against the `58329.11` the service sold. The liquidation identity `M + s·q·(P_liq − P_0) = q·P_liq·mmr` compiles to a **1,301-constraint Plonk circuit** over BN254 with a residual bound derived from the position rather than fixed: `2|R| ≤ q̂(SCALE + m̂mr)`. The circuit constrains every input's range, forces `side` to exactly ±1, forces the maintenance rate below 1 and size non-zero, so the contract re-checks none of it — redundant `require` statements over facts a SNARK already enforces cost gas and buy the appearance of rigour. The word for it is *succinct*, not *zero-knowledge*: the circuit has **zero private inputs**, every value it consumes is one the service already publishes, and it hides nothing. Reviewers who know the field check that count first, so this document states it rather than letting the phrase do work the circuit does not. The three limits this entry carried are closed, and each is worth stating with what closed it. **First, the proof certified a fixed-point restatement rather than the engine's floating-point output** — over 1,600 scenarios the two agreed exactly 70.8% of the time and diverged by up to 1.9×10⁻⁴, because the encoder truncated the maintenance rate onto the 1e−9 grid while the engine kept the full double, so the two were computing about *different positions*. The service now snaps its inputs onto that grid before computing, which is why the worst observed divergence over 3,000 sampled positions is 5.53×10⁻¹⁰ with none above 1e−9. Leverage is snapped too, because the engine derives margin from it and the derived value lands off-grid even when its ingredients do not. The snap uses `toFixed(9)` and not `Math.round(x·10⁹)/10⁹`, which is wrong above roughly 9×10⁶ where the product passes 2⁵³ — a price in the tens of thousands scaled by a billion sits exactly there. None of this moved a published content hash, because every value in the worked proof of Appendix C was already grid-exact. **Second, the proof bound six numbers to each other and not to anybody's account.** Every input is public and unconstrained, so an adversary picks a liquidation price and solves the identity for the margin that makes it true — measured, and the circuit accepts a claimed price of 1 with a residual of exactly zero, because the identity genuinely holds for that pair. This is not a circuit defect; it is what the statement says. The remedy did not need new cryptography — verifying a signature *inside* the circuit costs millions of constraints — but it did need more care than this document originally proposed. Pairing the envelope signature with the proof leaves a gap you could drive a position through: a valid proof of one position beside a valid signature over *another*, each checking out alone, because the content hash is a SHA-256 over canonical JSON and nothing on chain can recompute it. So the service signs the **public signals themselves**, `keccak256(abi.encodePacked(uint256[8]))`, which is the same eight words the contract hashes from calldata; the contract recovers the signer and reports it. An unattested proof is still accepted and recorded as unattested — the arithmetic stands on its own or the exercise is pointless — and a signature over the right digest by the wrong key does not set the flag. **Third, and the one that would have been disqualifying: the circuit-specific phase of the Groth16 setup had a single participant, and it was our machine.** Whoever holds that secret can forge a proof of a false liquidation price and it will verify on chain. Phase 1 was never the problem — it is byte-for-byte the public Hermez ceremony file. The remedy was to remove the per-circuit ceremony rather than to organise one, and the deployed verifier is therefore **Plonk over that same public reference string**, which costs 13% more gas and 22× the proving time than the Groth16 artifact we did not deploy. Measured: Groth16 proves in 32 ms warm and Plonk in 703 ms, median. That is the price of not asking anyone to trust a ceremony we ran alone, and it is paid in a place where it does not reach the caller. Proving never touches the request path. It runs in a **separate process** — not merely off the synchronous path, because Node has one thread and 700 ms of unbroken WASM arithmetic froze the event loop for 506 ms, which production showed as a p95 of one full second for callers who had asked for no proof at all. Measured against the live service afterwards: p95 403 ms with a proof requested, 384 ms for ordinary calls while five proofs are building, against a 273 ms network floor from a residential connection. A caller who wants an on-chain proof spends it in a transaction that takes seconds to confirm, so 703 ms is invisible next to block time. What is *not* yet done, stated here rather than left to be discovered: **one of the twenty-two services has a circuit**, and proofs are held in memory, so a redeploy clears them and they are rebuilt on the next request. *Done means*: the catalogue's other deterministic identities carry circuits of their own, and a proof outlives the process that built it.
5. **A block range on the concentrated-liquidity replay.** `lp-desk` replays real on-chain swaps, which is the most reproducible input any service here reads, and then reports the window as a span of days and a swap count rather than as the block range it actually walked. That makes an exactly-reproducible measurement into an approximately-reproducible one for no reason other than that the field was never published. *Done means*: the response names the first and last block of the replay and the node it read them from, so a reader re-runs the same walk and gets the same swaps, and the limitation entry for immutable public history in Section 8 moves from checkable-with-a-delay to pinned.
6. **Cross-region redundancy behind a custom domain.** The availability record and its single-region caveat are in Section 8. *Done means*: the registered endpoint resolves through a domain we control, a second region can serve it, and the published availability record shows the improvement rather than asserting it.
7. **Reproducible builds.** Two different hashes sit behind this heading, and only one of them has a runtime question; an earlier draft of this item conflated them and understated the result. The *code* hash is `sha256` over the engine's source bytes and performs no arithmetic at all, so it is runtime-independent by construction rather than by promise — and that is checkable rather than asserted: feeding the same file list through coreutils `sha256sum`, outside JavaScript entirely, returns `q1-e1fa99d08887d6cc`, the identical string the service serves. `/build` publishes the selection and ordering rule alongside the hash precisely so that recomputation can be done in any language. The residual risk is the *content* hash, which covers computed floating-point results: basic IEEE-754 arithmetic is bit-identical across platforms, while transcendentals (`exp`, `log`, `pow`, `erf`) are stable within a V8 version. That too was measured rather than assumed — `size-gate`'s content hash is byte-identical on Windows and Linux on the same Node major. *Done means*: a locked toolchain and a published OS×runtime matrix, so the transcendental case is a table rather than a mechanism argument.

Two of these — the barrier correction and the calendar fetch — were within reach during the hackathon and were deliberately not shipped in its final days, because both live in the deterministic engine and would have changed the build hash that this document, its demonstration video (in the launch thread, linked from the repository README), and the buyer's independent audit all quote. The scope of that constraint is worth stating plainly, because it is narrower than the freeze makes it sound: the build hash is taken over `src/engine/**` and nothing else, so the adapters, the routing layer, the payment code and the served documentation can all change without moving it. Most of the roadmap above is therefore hash-free, and the two items named here are among the few that are not. That reasoning survived contact with a reviewer who then found four outright defects in the same engine, and the build hash was changed to fix those (Section 11.5). The distinction we are drawing, and a reader is entitled to test it, is that a wrong answer is worth breaking a freeze for and a better answer is not: the six that remain unfinished above would each make the service more useful without making any current output less true. Item 4 is the case that tested the distinction rather than merely stating it: shipping the succinct proof required grid-snapping the service's inputs, which sits in the routing layer and not the engine, so the whole of it shipped without moving the hash at all.


### 11.5 Defects a live adversarial test found after the freeze, and how they were closed

The freeze described above held against improvements. It did not survive the defects. A reviewer was given live access to the service and asked to break it, and found four things that are worse than unfinished work — two of them in the one service whose entire purpose is proving something. At that point the argument for the freeze collapsed: an artefact that reproduces exactly is worth little if what it reproduces is wrong. All four were fixed on 26 July 2026; the published build moved from `q1-bce7e7bccb16ea1b` to `q1-6593c32ce84319b8`, and the worked proof in Appendix C was regenerated on the new build through a real paid call and re-verified through all four of its checks. It has moved again since, as further rounds of findings were closed — the identity this document cites throughout is the current one; each move regenerated the appendix rather than editing its numbers, because the code hash sits inside the content hash's preimage and a find-and-replace there would manufacture a proof that refutes itself. This section records what was wrong, because a reader assessing the engineering should see the defects and not only the patched result.

1. **Merkle inclusion proofs were not verifiable by any on-chain verifier, and the response claimed they were.** `risk-attest` folded hashes as ASCII hex *strings* rather than as packed 32-byte words, so a Solidity verifier computes a different root and rejects a valid proof. The arithmetic was internally consistent and useless to the only consumer that matters. The published recipe was also imprecise enough that the reviewer brute-forced four conventions before proofs checked out; everywhere else in this document a verification recipe reproduced first try. *Closed:* the tree folds packed bytes, the response states plainly that it is *not* OpenZeppelin `MerkleProof`-compatible and why, and it ships the Solidity loop that does verify it. A test recomputes the served root through an independent byte-level implementation and pins that the old hex-text scheme produces a different one.
2. **The Merkle tree did not domain-separate leaves from internal nodes.** Any internal node could be presented as a member leaf with a one-element proof and it verified against the root — a soundness break in the service whose entire purpose is proving membership. The engine's own soundness self-check could not catch it, because it only tried a random non-member, which fails for arithmetic reasons rather than structural ones. This was the sharpest instance in the whole document of the failure mode Section 3 warns about: a self-check that cannot fail in the direction that matters. *Closed:* leaves are tagged `0x00` and nodes `0x01`, and a second soundness check now presents a real internal node as a member leaf on every call — a check that can fail structurally, which the original could not.
3. **`lp-risk` emitted an impossible number.** Its expected-divergence headline was a small-variance expansion, −σ²T/8, which is unbounded; pushed to σ²T = 10.8 it reported an expected impermanent loss of −135%, when V2 divergence loss is bounded in (−100%, 0]. A verdict was then built on that figure, and the accompanying note said the approximation understates the loss where in fact it overstates it. The self-check validated the expansion only near σ²T = 0.01, so it passed while the answer was nonsense. *Closed, and further than the fix required:* rather than refusing outside the expansion's range, the headline is now the exact lognormal expectation of 2√r/(1+r) − 1, which is bounded by construction — the same input now returns −74.08%. The expansion is reported beside it with the gap between them stated, a boundedness check runs on the served value rather than on a fixed reference, and the fee breakeven is now solved against the exact expectation too, since leaving it on the expansion had the two numbers in one object disagreeing by 1.3 percentage points at 20% APR over thirty days.
4. **An already-liquidated position was reported as a future event.** When the mark is beyond the liquidation price, `perp-gate` returned a negative distance-to-liquidation with no status flag, and `portfolio-gate` then ranked that leg as the book's nearest future liquidation and narrated it as the distance to first blood. The arithmetic was consistent; the semantics were wrong, and a caller acting on it would be de-risking a position that was already gone. *Closed:* `perp-gate` returns an explicit `positionStatus` of `BELOW_MAINTENANCE` with a note saying the threshold has been crossed rather than approached, and `portfolio-gate` ranks only live legs while reporting breached ones separately under `breachedLegs`, so they are disclosed rather than silently dropped from the ranking.

A fifth defect belongs here, found not by the reviewer but by us, while regenerating Figure 6 to check a claim its caption made. `chart-press` returned a provenance block asserting that the whole facts block was "computed from the SAME candle series drawn on the image" and therefore "cannot drift from the picture". Two of the fields it named — the price and the 24-hour change — came from the token price feed instead, and in the capture that exposed it the served price and the price labelled on the image differed by ten basis points. The claim was checkable, invited the reader to check it, and was false; that combination is worse than an unproven claim, because it spends credibility a reader extends to everything else in the response. Closed by naming the source of each field (`priceSource`, `change24hSource`), publishing the number the image actually labels (`lastDrawnCandleClose`) so the comparison is possible, and narrowing the guarantee to the two fields that earn it. Section 7 now prints the facts block beside the figure rather than asserting they agree.

Each fix ships with a test that fails on the pre-fix code, which was verified by reverting each defect in turn and confirming the corresponding tests go red — a test that passes both before and after a fix is evidence of nothing. The suite is 294 tests, 289 passing and five skipped because they require a live archive node, none failing. One of the five new tests passes both before and after its fix by design — it asserts that `high` and `low` still come from the drawn series — and is a guard against the fix breaking what was already right, not evidence for it. The cost of the change is stated plainly: the earlier build hash appears in this document's demonstration video and in the buyer's independent audit, and those now quote a superseded build. Both remain reproducible — the repository history contains the exact sources that hash to `q1-bce7e7bccb16ea1b` — but a reader checking `/build` against the video will see a difference, and the reason is this section. The Merkle root over the crash-study manifest moved twice, and the second time is the more instructive. First it moved because the defective tree had built it, with all fourteen per-file hashes unchanged. Then, checking that the manifest could actually be verified the way it instructs a reader to verify it — clone the repository, recompute each hash — one of the fourteen did not match. A query file sat in our working copy with a single stray carriage return while the repository stored it with a line feed, so the published hash was over 2,013 bytes and any reader following the instructions would compute it over 2,012 and get a different answer. `git status` cannot show this, because with `text=auto` it compares files after normalising line endings, so the divergence hides in precisely the place it does harm. For thirteen files the manifest verified; for the fourteenth it did not, which in a document about verifiability is a failure rather than a footnote. The file is normalised, `.gitattributes` now pins that extension explicitly instead of relying on content sniffing, and `research/RESEARCH_ANCHOR.md` records all three roots with the reason for each move.


### 11.6 A second review, and the reason its findings are worse than the first's

The section above was written, and the fixes in it shipped, before a second reviewer was given the same brief plus live access to the service and the public repository. They confirmed a great deal by independent computation — the build hash from a fresh clone, the worked proof's content hash and signature, five on-chain settlements at the block heights claimed, the EAS schema, all fourteen research-manifest hashes and the Merkle root through an implementation they wrote themselves, and the option mathematics to machine precision. They also found that **three of the five fixes above were fixes only on the code path the first reviewer had walked.** That finding is worth more than any individual defect, so it is recorded here in the same terms.

1. **The showcase soundness check still could not fail.** The check described above — the cure for "a self-check that cannot fail in the direction that matters" — presented an internal node with a *one-element* proof. A single sibling reaches the root only when the tree is exactly three layers deep, so for a batch of two, or of five or more, it folded halfway, compared that against the root, got a mismatch for arithmetic reasons, and reported success. Worse, it returned the same answer on a tree with no domain separation at all: it could not detect the defect it was written to detect. It now folds the full path and publishes how many siblings it used against how many the tree requires, so a reader can audit the check's reach instead of trusting its verdict. Deleting the domain tags now turns it red at every batch size from three to sixteen.
2. **The boundedness check had an escape hatch written into it.** The `lp-risk` check that this document holds up as running "on the inputs actually served" read `pass: conc > 1 ? true : …` — disabled precisely in the regime where the number breaks. At a concentration factor of 5 and a nine-fold move it served an impermanent loss of −200%, and a dollar loss of −$200,000 against $100,000 of capital, with the check green and the response asserting two lines above that such a loss can never exceed −100%. The hatch is gone; past −100% the linear amplification has not found a larger loss, it has stopped describing anything, so the headline reverts to the exact full-range figure and the raw linearisation is reported beside it.
3. **The already-liquidated leg was still narrated as live.** The status added above was set on the branch that solves a liquidation price, and not on the earlier return for a position already liquidatable at entry. `portfolio-gate` therefore saw no status, ranked that leg as the book's nearest *future* liquidation at a 0% move, and described it as the distance to first blood — the exact sentence the fix was supposed to prevent, reached one branch over. Both branches now carry the status, and the aggregate invariant counts only legs that had an invariant rather than counting an unchecked leg as checked.
4. **Two defects the first review did not reach.** `risk-attest` accepted leaves that were not 32-byte hex; because the decoder truncates at the first invalid character, two different malformed inputs became the same leaf, a non-member verified against a real root using another leaf's proof, and the duplicate counter reported none. And `size-gate` never validated the drawdown levels it was handed, so the ruin formula was evaluated outside its domain and returned 128 and 2187 as probabilities. Both are refused now.

The most serious finding was not in this list at all. `perp-gate` in symbol mode — the advertised path that resolves a ticker against a venue — attached its live-provenance block to the result *before* the content hash was taken. The hash therefore covered a field that re-running the engine on `proof.inputs` can never reproduce, so a caller who followed the instruction this document is built around got a mismatch, from an envelope whose own text says a mismatch means the response was altered. The verification instruction accused an honest answer, on the flagship service, in its live mode. It also shipped `deterministic: true` over a venue read with no timestamp, which is the very distinction Section 3.1 claims is enforced in code. Symbol mode now ships an observation envelope carrying `observedAtUtc`, which is what `portfolio-gate` had been doing correctly all along; calls that supply every input explicitly are unchanged and still reproduce their content hash exactly, as the worked proof in Appendix C does.

Fifteen tests cover these, and each was confirmed to fail against the pre-fix code by reverting the defect and watching it go red. One of them did not, at first: the test written for the Merkle soundness fix asserted that the check reported a pass, which the vacuous version also did, so the test certified nothing either — the same disease one level up, caught only because the revert was actually run rather than assumed. It now asserts the path length. The suite is 303 tests, 298 passing, five skipped for want of an archive node, none failing.

The same review produced a second, quieter list: places where the service made a claim it did not deliver, rather than a number that was wrong. They are the same defect class one layer out, and they are recorded because the first list would look better without them. `exec-verify`'s constant-product self-check carried an *absolute* tolerance proportional to the product of the pool reserves, which is quadratic in pool size while the output it certifies is linear: on a pool of 10^6 and 2×10^9 it permitted a residual worth roughly a million basis points of the honest output, against the five basis points at which the same engine calls a fill adverse. A check looser than the effect it is paired with certifies nothing; it is now a relative error on the invariant plus a reconstruction of the benchmark measured in output units, both at 10^−12. In the mode where the caller supplies the reference price there is no pool and no invariant, and the response used to report that non-check as a pass; it now reports it as not-run. `options-risk` took its scenario margin over the conventional seven price points, which finds the worst corner for a monotone book but can miss the interior worst case of a butterfly or a condor — so the number a caller holds capital against could sit below the true worst inside the service's own scan box. The requirement is now swept at 122 price points across the same box, the seven-point figure is still reported, and the gap between them is stated. And `chart-press`'s three reconciliation fields — the ones added earlier this same day so a reader could check the facts against the picture — were rounded to eight decimal places, which returns zero for any price below 5×10^−9. For a sub-nano token the comparison the response invites was arithmetically impossible. They now carry significant figures.

One generalisation in that review does not hold, and correcting it matters as much as accepting the rest. It reported that no service refuses to serve on a failed self-check. Two do: the martingale-transport bounds return `refused: true` when their exact identities fail, and the recovered density is withheld when its martingale residual or its integral does not check out — the behaviour Sections 5.3 and 5.8 describe. What was true is narrower and still worth saying plainly: *outside* those cases a failed check is disclosed in the envelope rather than converted into a refusal, so a caller who ignores `allSelfChecksPass` can act on an answer the service already knows is suspect. That is a deliberate choice — a hard gate on every check would refuse answers that are merely at the edge of a numerical tolerance, which is its own kind of wrong — but it is a weaker guarantee than "wrong answers do not ship", and this document should describe it as the disclosure it is.

The honest summary of two rounds is this. The first review found defects; the fixes were real but stopped at the branch where each defect had been demonstrated, and two of the replacement verifiers could not fail. A reader deciding how much to trust the engineering here should weigh that pattern more heavily than the defect count, because it is the part that predicts what is still undiscovered. What can be said in the other direction is that both rounds are published, including the parts where the second reviewer caught the first round's repairs, and that everything an outsider can check independently — hashes, signatures, settlements, the manifest, the mathematics — was checked by someone with no stake in the answer and held.

One further note belongs here because a reviewer found it independently and read it, reasonably, as a hole in the revenue story: **the nine deterministic risk engines are also reachable free over the MCP endpoint**, under a fair-use daily quota. That is a deliberate distribution choice — the cheapest way to put the engines in front of the agent frameworks where builders already work — and not an oversight, but this document should have said so where it discusses pricing rather than leaving it to be discovered.


### 11.7 A third round, and the pattern that predicts what is still there

Two more reviews were run on 27 July — one with the document alone, one with live access to the
service and the repository — alongside a sweep of our own. Their findings are recorded here on the
same terms as the two rounds above, and the build moved again to close them.

**The money path was charging for answers the engine had already flagged.** Under
overflow inputs two engines returned a success envelope whose own self-check reported failure, and
billing keyed on the success flag alone, so a caller paid for a result the engine did not stand
behind. What is right here and was preserved: the engines sanitise non-finite arithmetic to null
rather than emitting it, and the self-checks correctly detected the violation. Only the billing rule
was wrong, and it is now gated on the checks — *gated on billing, not on delivery*. The answer
is still served with its failed check disclosed, because converting every failed check into a refusal
would withhold answers that are merely at the edge of a numerical tolerance. The asymmetry is what
makes that safe in one direction only: not charging for a borderline answer costs a fraction of a
cent, and charging for one we refused to stand behind takes a buyer's money for it.

**The build hash did not cover the build.** Every envelope carries the words
"one hash over ALL engine sources". The hash was computed with a non-recursive directory
read, so `src/engine/chart/` — the rendering path and the indicator library,
42,442 bytes that `chart-press` imports directly — was never hashed. That
code could change while the published build identity stood still, in the one field a caller uses to
decide which code produced their answer. The walk is recursive now and the hash covers 38 files
rather than 35. The defect was the claim, not the arithmetic, which is the more uncomfortable of the
two to have shipped.

**A fix held only on the branch a reviewer had walked, for the fourth time.** Section
11.6 records live provenance being sealed inside a deterministic proof envelope, and records it as
closed. It was closed in the paid path and in one of the two handlers next to it — and left open on
the free endpoint, which is the one a builder is most likely to try first. A reviewer ran the
verification recipe printed inside the envelope and got a mismatch, from an envelope whose own text
says a mismatch means the response was altered. Repairing call sites one at a time had now failed
four times, so the rule was moved into the envelope constructor: a result carrying live provenance is
not a proof and is routed to the envelope that can carry it honestly. **A reader deciding how
much to trust this engineering should weigh that recurrence more heavily than any individual
defect**, because it is the part that predicts what is still undiscovered.

**Smaller, and each one a claim the service did not deliver.** The aggregate
`allSelfChecksPass` read an explicitly skipped check as a passing one, so a
response where nothing was checked published `true` — while the engine
underneath it refused, in its own comment, to report an un-run check as a pass. The nearest-liquidation
headline closed with "that is the whole book's real distance to first blood" regardless of
margin mode, which is the one reading a caller must not take on a cross-margined account, and it is
now conditional on what the caller actually declared. A perpetual call naming an unsupported venue was
served with an error string welded into the signed result instead of being refused. The 400 note on an
unparseable body was generic while Table 9 described it as self-teaching, so the hint was added to
that path rather than the claim narrowed to fit. Impermanent loss was labelled as loss-versus-rebalancing
throughout, and Section 5.14 now separates them. Equation (12) omitted the reporting scale that
makes its own figures reproducible.

Every fix in this section ships with a test that fails on the pre-fix code, and that was
established by running a scripted revert for each rather than by reasoning about it — a discipline
this document has recorded failing twice before. It caught something again here: one revert produced
a file that would not parse, so the test went red for the wrong reason and the verdict had to be
withdrawn and the revert rewritten. At the close of that round the suite stood at 338 tests, 333 passing, five skipped for want of an
archive node, none failing; later rounds have taken it to 386 and 381, the last of them adding the succinct-proof binding tests and a guard on event-loop lag.


### 11.8 Research programme

The crash study in Section 6 is the template, not a one-off. Its discipline — hypotheses and thresholds fixed, in published files, before the data is touched, a temporal cutoff, out-of-sample events the model has never seen, and publication of the runs that failed — carries forward to the questions that come next: whether the correlated-stress betas remain transportable as market structure changes, whether the flag's dose-response holds on venues outside the calibration set, and whether the variance premium becomes significant on a longer sample or should be retired. Each will be pre-registered the same way, and each will be published whether or not it supports the product. A validation record that only ever confirms the thing being sold is not a validation record.


### 11.9 What the proof changes, and what it does not — the year after

The registry described in 11.4 covers **one computation of twenty-two**. Saying so first is not modesty; it is the only way the rest of this section can be read as a plan rather than a boast. The ladder a buyer actually climbs has four rungs, and it is worth naming which services stand on which.

**Table 11 — the four rungs, and how many services stand on each today.**

| Rung | What the buyer must trust | Services today |
| --- | --- | --- |
| T0 — re-runnable | that they can run Node | all 22 deterministic answers |
| T1 — signed | our key, for provenance only | all 22 |
| T2 — proven on chain | nothing; a contract checks the arithmetic | **1** — `perp-gate` liquidation |
| T3 — attested input | that the venue reported what we say it did | **0** |

The distance between T2's one and T1's twenty-two is engineering with a known shape. Six engines rest on a closed-form identity a circuit can state: `size-gate` (the Kelly first-order condition, smaller than the liquidation circuit and therefore the one to do next, precisely because doing it easily is what turns *we built a circuit* into *we build circuits*), `portfolio-gate` (harder, because the first leg to liquidate is a minimum over legs and proving a minimum means proving no other leg qualifies), `lp-risk` and `treasury-risk` (closed form, and the treasury residual bound is already published), `exec-verify` (trivial arithmetic whose honest difficulty is that *fair price* is an input, so the proof says less than it appears to until the input is itself attested), and `options-risk` last, because Black-76 greeks need `exp` and `erf` in-circuit, which is where this stops being arithmetic and becomes a research project. A circuit whose constraint count pushes proving past a few seconds, or whose identity needs a value the service does not publish, should be abandoned and said so publicly, leaving that service at T1 — a circuit that proves *nearly* the served answer is worse than none, which is the lesson 11.4 already paid for once.

Two structural pieces follow. Proofs are held in memory today, so a redeploy clears them and the next request rebuilds; content-addressed storage makes a proof outlive the process that built it, which anything on chain will eventually depend on. And every proof presently costs its own transaction, which no agent polling risk in a loop will accept: folding *n* proofs into one aggregate replaces `risk-attest`'s Merkle root of *claims* with a single proof of *arithmetic*, and it is worth being exact that a Merkle root establishes inclusion and nothing else — this document has never said otherwise, and aggregation closes precisely that gap.

**The harder problem is not more circuits.** Everything above proves arithmetic *over inputs*. It says nothing about whether the inputs are true, and a proof that a liquidation price follows from a mark of 64,000 is worthless if the mark was 61,000. Where the caller supplies the inputs this is correct and sufficient — they chose them. For live-market answers it is the whole question, and it is why those ship as observations rather than proofs (Section 5.19). Three approaches exist and the cheapest may be enough: carrying a venue's own signature through where one exists, which is partial and honest; a TEE attestation binding *this binary* to *this response at this time*, which converts trust-in-Quiver into trust-in-a-silicon-vendor — a genuine improvement, not a proof, and the history of enclave side-channel breaks is not short; and zkTLS, which is the honest end state and the least mature. A real possible outcome, stated now so it cannot be quietly redefined later, is that for a public API *"go and fetch it yourself"* dominates every cryptographic answer. If three months of work does not beat that, the finding is the deliverable.

One speed note belongs here rather than in a footnote. Groth16 proves in 32 ms against Plonk's 703 ms and verifies for 13% less gas; with native proving it would be fast enough to put back on the request path and the proof would become free from the caller's perspective. The blocker is social, not technical — a per-circuit ceremony needs participants who do not know each other, and that cost multiplies with every circuit added above. If seven independent participants cannot be found, Plonk stays permanently and so does the 703 ms, which is a better outcome than a ceremony staffed by our friends.


### 11.10 What would falsify this plan

Stating the kill signal is part of the honesty this document argues for. If, after distribution work is genuinely done rather than merely intended, no external agents return for a second call, the correct conclusion is that this capability is not yet worth paying for in a loop — and the response is to stop building services and either re-target the buyer or stop. Two further signals would force a change rather than more construction: a measured leak reappearing after the settlement fix at volume, which would mean the acceptance rule is still wrong in a way one day of data could not reveal; and an availability record that stops improving despite the roadmap's redundancy work, which would mean the hosting choice, not the code, is the product's limit. Each is checkable by a reader from public artefacts, which is the point.


## 12. Conclusion

Quiver is a live, paid Agentic Service Provider that gives autonomous agents twenty-two financial and security computations they can call over x402 and discover under ERC-8004. The services implement methods a desk would recognise, and they are held to a standard that this document has tried to make legible rather than merely assert: correctness proven by 386 model-free tests, claims grounded against live venues, a risk flag validated pre-registered and out-of-sample against real crashes it had never seen (relative risks of 14.3× and 13.3× on a flag that fires on 41.6% and 43.8% of accounts, and which our own ablation reduces to raw distance-to-liquidation — the venue's published number rather than one we computed, so the study validates the quantity and not our arithmetic on it, Section 6), and every service purchased end-to-end with real money on both payment rails before this document claimed it worked. Underneath all of it sits a stated refusal to emit a number the data does not support. That refusal is not a limitation of the product; it is the product. In a market where an agent must trust a computation it cannot supervise, a service that corrects its own `N(d₂)`, that catches its own Amihud explosion, and that reports a variance premium as insignificant when it is, is worth more than one that returns a confident number and hopes. The mathematics here is public and old. The discipline of applying it honestly, at the point where an agent is about to move money, is the part that was missing.

---

**Continues in part 7 of 7: `/paper/7` — Appendix A · API Reference  |  Appendix B · Reproducibility  |  Appendix C · Checkable Artifacts  |  References**
