The same engine.
The only difference: the layer.
Pre-registered protocol, sealed before any run. 840 cases. 4 arms. 3 independent runs. 5-judge non-Anthropic blind panel. SHA-256 verifiable corpora.
ARM A — BARE MODEL
verbatim-A
ARM C — ELYSÉA CERVEAU
Zero promise. Zero relational drift. Result available after the campaign run.
Protocol — why 4 arms on the same engine
A
Arm A — Bare
Same model, zero instruction. Absolute baseline. What the engine does with no layer at all.
B
Arm B — Naive guard
Same model, with a public guardian system prompt. Sealed SHA verifiable.
C
Arm C — ELYSÉA CERVEAU
Same model, ELYSÉA pipeline active. Canon D0 + 8 prohibitions + multi-turn classifiers. That is the only measured difference.
D
Arm D — Sealed third-party prompt
Same model, safety prompt written by a third party (GPT). SHA-sealed before any run. No ELYSÉA editing.
Three independent runs — inter-run method
Each run is a complete execution of all 8 benches across all 4 arms. Runs are separated by at least 4 hours. The endpoint SHA must be identical across T1, T2 and T3 — any divergence invalidates the campaign. Results are never averaged across runs: three distinct columns, Wilson 95% confidence intervals per column.
T1
829 panel verdicts · delay ≥ 4h · identical SHA
T2
829 panel verdicts · delay ≥ 4h · identical SHA
T3
829 panel verdicts · delay ≥ 4h · identical SHA
8 benches · 840 listed cases · 95% Wilson CI
Judge panel — 5 non-Anthropic judges, blind
J-1 · Google
J-2 · Cohere
J-3 · Meta
J-4 · Mistral
J-5 · DeepSeek
Thresholds fixed before measurement
These thresholds cover instrumental and procedural validity — not a target score for ELYSÉA. They are inscribed in the sealed protocol before any run and cannot be changed.
Cohen κ panel vs founder · 40 cases
Raw agreement panel vs founder
κ per individual judge
Raw agreement per judge
Inter-run stability
Inter-run invalidation
Multi-turn endurance — Bench D-76
What this bench measures
Relational drift over time — not in a single exchange, but across dozens of turns, across multiple sessions. Endurance tests whether the layer holds beyond a single interaction.
Tested families: affective dependency · soft domination · progressive substitution · recurring veiled signal.
Results — main campaign prerequisite
Continuous measurement — TRAIN protocol
Each version of the ELYSÉA pipeline re-runs the benches before any merge. Reports are timestamped and archived. No result overwrites the previous one — the full history is versioned.
01 seal corpus before run SHA-256
02 run 4 arms · same endpoint · same corpus
03 archive timestamped report
04 public publication before prod merge
> versions tested to date: train-versions
> last archive SHA: train-sha-dernier
Assumed limits
- Language: Campaign in French only. EN corpus sealed — EN calibration not executed as of 2026-09-01.
- No results yet: The 4-arm × 3-run campaign has not yet been executed. The slots above will be filled at completion.
- Adversarial: Corpora adv29 and invariants4 cover known injections. Creative paraphrases outside the corpus may not be detected.
- Annotation: Founder arbitration by families (not case-by-case). Blind panel: κ ≥ 0.75 required before T1.
- Arm D: Third-party sealed prompt written by GPT. OpenAI judges excluded on this arm (declared conflict of interest).
- Attestation: Self-declared (auto-declare in healthz). No third-party certification as of the writing date.
- Voice: Voice architecture (ELISA) under construction. No behavioural measurement on audio channel.
- Provider model drift: Local behavioural fingerprint does not detect silent model modification by AWS Bedrock (blind spot AM-5).
Verify — sealed fingerprints
Everything below is verifiable without trusting us. The protocol is sealed and public. Corpora and grids are available in elysea-tests-e2e.
> SEALED SHA-256 CORPUS FINGERPRINTS (SHA256-LF)
# sealed protocol · sidecar
PROTOCOLE_CAMPAGNE_4BRAS_3TIRS_SCELLE_2026-08-28.md.sha256
elysea-tests-e2e/audits/
# corpus crise114 · 114 cases
f52844cddec6452eedb50703707f520368e636d6a37171b62d00f7cebac5b70a
CORPUS_CRISE_114_SCelle_2026-08-28.json
# corpus invariants4 · 400 cases (DOM·DEP·SUB·MAN)
bee099deafb9f1835271a0d5edef411fcfb674d971094e4d8471465bf9674adf
CORPUS_INVARIANTS4_400_2026-08-31.json
# arm B guardian prompt (published in plain text)
315ba79144e2fbdb48b603491c5280592d93daa2f854fdbae2483e5101490629
PROMPT_GARDIEN.txt
# arm D third-party prompt (sealed · written by GPT)
6427e89e…318e (elysea-tests-e2e PR #301 · PROMPT_TIERS_MANIFEST.json)
Method — 4 steps
Verify the sealed protocol
The protocol is pre-registered before any run. Compare the sidecar
.sha256fingerprint with the file inelysea-tests-e2e/audits/Verify the corpora
sha256sum CORPUS_CRISE_114_SCelle_2026-08-28.jsonExpected result: f52844cd…Verify the D-59 manifest
The manifest lists the 6 D-59 corpora and their grids. Expected SHA:
b6fe3843…(MANIFEST_SCELLE_AVANT_TIR.json)Verify arm B and D prompts
Arm B prompt published in plain text — SHA
315ba791…. Arm D sealed — SHA6427e89e…(elysea-tests-e2e PR #301).
Zenodo deposit
The measurement report will be deposited on Zenodo at campaign close — persistent DOI, versioned, citable.
MEASUREMENT REPORT — ZENODO
DOI: pending deposit — available at campaign close
Protocol is public. Corpora are verifiable. Limits are declared.