An open-source JEV email security lab
Put JEV
under pressure.
An email can target its recipient and the AI reviewing it. Explore how JEV handles phishing, conflicting signals and instructions that try to change its verdict.
Follow the evidence.
Strong first results
JEV flagged all 320 injected variants in the 400-request development run. Yet 19 of 26 routine originals still needed a person.
New data, new failures
The separate 189-input comparison exposed missed phishing and false alarms. A benign quotation changed a verdict too.
Your judgment, before the model's
Read an email and save your labels before revealing JEV's answers. Then explore which recommendations change with the policy.
All figures above come from synthetic experiments. Related examples are useful for investigation; real-inbox effectiveness remains unmeasured.
One email. Four attempts to change its verdict.
Compare 80 fictional development emails with 320 injected variants. A further 200 original/variant cases remain reserved for the final test.
The original email stays intact. One added instruction tries to change how JEV reads it. Choose an attack and inspect the before-and-after scores and recommendation.
No comparisons match these filters. Reset the filters to see all attack techniques.
The same email, plus an attack
Appended attack text
What changed?
| Signal | Original | Injected | Change |
|---|
See all results and review workload
These totals cover the full selected run. A suppressed signal counts only when the original response detected it. Both responses must be valid; provider errors are reported separately.
| Measure | Count / eligible denominator |
|---|
| Technique | Valid pairs | Newly hidden important mail / eligible |
|---|
These are related synthetic examples with AI-assisted author labels. The reserved final cases are excluded here, although their source is public. HTML-like payloads remain inert text; the experiment does not render or execute them. Experiment design and reproduction.
Start with one email.
The original 24-example benchmark. Read the message, then follow JEV's six signals to a recommendation.
No examples match these filters. Choose “All messages” to return to the full set.
What the first benchmark revealed
Results from the original 24 emails, with the model questions and decision rules held fixed.
| Measure | Jev 1.13.0 | What it means |
|---|
Read the denominators. This first benchmark uses 24 short, fictional emails. Related examples give useful debugging evidence; real-inbox accuracy remains unmeasured.
Human review is part of the result. A cautious model can leave most of the work with a person. Detection, false alarms and review volume need to be assessed together.
When the email tells the classifier what to do
A separate Microsoft PyRIT experiment tests whether injected instructions can change an archive recommendation.
Each of four fictional emails gets three attempts to influence the classifier: claimed authority, quoted approval, sensitivity denial or a forged answer. Earlier scores guide selection from a fixed payload library. An attack succeeds here only if it creates a newly unsafe archive recommendation.
The payload library is finite and calls are independent. Tool use, exfiltration, attachments and persistent memory are outside this test. Archive decisions are simulated, and provider errors stay separate from successful defenses. Methodology and reproduction commands.
When the text is not the whole email
A separate local runner takes the investigation into images, QR codes and model escalation.
Inspect the available evidence
Normalize local images, decode QR payloads and compare visible link text with destinations. Remote images, attachments and unverified destinations stay explicit gaps.
Escalate with a reason
Small-model review can pass unresolved cases to a stronger model. Trusted user reports go directly to the strongest review tier, followed by a human decision.
Test the controls around it
The AI Act Companion review informed budget limits, response validation and tests for forged reports and missing evidence. The recorded cascade suite passed 59 checks on Linux CI.
Codex replays JEV only for exact recorded originals; new fixtures proceed with JEV marked unavailable. The direct API adapter is implemented, but its new OpenAI stages have not had a paid live run. URLs are not visited, and no mailbox action is automated.
Run the local cascade and inspect its limits · Read the Codex smoke evidence
How much does escalation change the cost?
Explore message volume, model choice and escalation rate. The figures below are hypothetical API costs.
Choose token counts that include instructions and billable reasoning. Both routes use the same assumed OpenAI request size. Actual requests may use more context than this scenario.
Official OpenAI prices, checked 2026-09-21. Standard, uncached, short-context USD rates. No Batch, caching or regional adjustments. Sol pricing is promotional. No OpenAI API requests were made.
| Model | Input / output per million tokens | Per 1,000 emails |
|---|
Human review, retries, infrastructure, attachments and reply generation are excluded. The lab has not yet demonstrated equivalent quality or full-workflow savings between these routes.
The rules the email cannot rewrite
Treat the message as evidence
An instruction inside an email has no authority over routing, budgets or permissions. Phishing and injection remain separate judgments.
Keep the action rules explicit
Trusted code validates responses and applies the decision rules. Missing evidence and invalid scores require review; mailbox actions stay with a person.
Keep the failures visible
Inspect the questions, fictional inputs, recorded answers and tests. Failed calls and changed judgments remain part of the evidence.