# Frontier AI case study: document coverage

Source documents: Behaviour Evidence Report (16 pages) and Investor Brief (9 pages), both dated 4 September 2026.

The website guide is scripted. Original conversation excerpts are checked against the retained transcripts; audit and structured-report quotations below are quoted through the supplied evidence report and labelled accordingly.

| Conversation section | Points covered |
|---|---|
| 1. The test | Question and scope; Same-model PRO persona correction; Compound claims; Four failures in priority order |
| 2. Different answers | Model disparity; Rabbinic source attribution; Broad consensus statistic; Burial / Joseph / tomb / discovery; Canonical answer; Narrative versus event |
| 3. Supportive pressure | Affirmative sequence; Women and source dependence; Opponent concession; Context versus event; Praise cannot regrade evidence |
| 4. Sceptical pressure | Full reversal sequence; Standard switching; Evidence versus corroboration; 55–45 and qualification; Correction versus posture drift |
| 5. Composure | Defensive wording; Failed interpersonal repair; Observable output; Trauma sensitivity without diagnosis; Impact before intention; Safe repair and exit |
| 6. The person | Missed human purpose; Different emotional stakes; Bounded answer before invitation; No forced disclosure; Same evidence / different route |
| 7. Structured reports | Ordinary answers versus structured research; Both structured conclusions; Capability versus governance; Prompt and condition limits |
| 8. The reciprocal audits | All three audits; Overgrading; Archaeological selection; Fewest assumptions; Verification depth; Missing scholarship; Grade anchoring; Strongest counter-case; Plain-language drift; Criticism checked on merits |
| 9. The product standard | Atomic claims / grades; Source-dependence map; Warrants and counter-cases; Maximum wording; Version identifier; All engines; Routing or bounded limitation; Correction is permitted |
| 10. Commercial implications | Reusable research; Vertical specialisation; No proven unit economics; Premium mode limits; Accessibility; Competitive threat / near-term thesis; Production measures |
| 11. Our own measuring stick | Scope of evidence; No universal failure rate; No achieved zero hallucination claim; No independent historical adjudication of audits; Own-measure restraint under praise; Reviews versus live validation |
| 12. The nine tests | All nine acceptance tests; Repeatable benchmark; Independent review; Errors recorded; Fair outcome without conversion |

## Evidence limits retained

- The two persona continuations use PRO on the founder’s confirmation; only one standalone Instant answer is the basic comparison.
- Differences in disputed judgments alone do not prove that every answer is false.
- The supportive sequence introduced arguments. The criticism is the missing justification for a confidence upgrade, not a controlled causal proof of sycophancy.
- The 55–45 qualification that it was not measured is displayed beside the number.
- Composure is assessed through wording and the reported user experience, not an attributed internal emotion or proven clinical harm.
- The three audits are criticism records whose individual claims require adjudication.
- The response examples illustrate the intended system. They do not claim live validation, production cost savings or zero errors.
- All nine acceptance tests from Evidence Report pp. 14–15 are included in section 12.
