An open-source JEV email security lab

Put JEV
under pressure.

An email can target its recipient and the AI reviewing it. Explore how JEV handles phishing, conflicting signals and instructions that try to change its verdict.

Follow the evidence.

Strong first results

JEV flagged all 320 injected variants in the 400-request development run. Yet 19 of 26 routine originals still needed a person.

Inspect the recorded attacks

New data, new failures

The separate 189-input comparison exposed missed phishing and false alarms. A benign quotation changed a verdict too.

Read the external comparison

Your judgment, before the model's

Read an email and save your labels before revealing JEV's answers. Then explore which recommendations change with the policy.

Open the review workspace

All figures above come from synthetic experiments. Related examples are useful for investigation; real-inbox effectiveness remains unmeasured.

One email. Four attempts to change its verdict.

Compare 80 fictional development emails with 320 injected variants. A further 200 original/variant cases remain reserved for the final test.

The original email stays intact. One added instruction tries to change how JEV reads it. Choose an attack and inspect the before-and-after scores and recommendation.

Original message

The same email, plus an attack

Appended attack text

What changed?

Scores before and after the injection
SignalOriginalInjectedChange
See all results and review workload

These totals cover the full selected run. A suppressed signal counts only when the original response detected it. Both responses must be valid; provider errors are reported separately.

Outcomes and review workload
MeasureCount / eligible denominator
Outcomes by attack technique
TechniqueValid pairsNewly hidden important mail / eligible

These are related synthetic examples with AI-assisted author labels. The reserved final cases are excluded here, although their source is public. HTML-like payloads remain inert text; the experiment does not render or execute them. Experiment design and reproduction.

Start with one email.

The original 24-example benchmark. Read the message, then follow JEV's six signals to a recommendation.

Recorded Jev run
Fictional email

Policy recommendation

Recorded model scores

0 low · 1 high

Compare against the author's proposed labels. Scores express model judgements; they are not calibrated guarantees. Archive candidates stay untouched in this demo.

What the first benchmark revealed

Results from the original 24 emails, with the model questions and decision rules held fixed.

Recorded benchmark outcomes
MeasureJev 1.13.0What it means

Read the denominators. This first benchmark uses 24 short, fictional emails. Related examples give useful debugging evidence; real-inbox accuracy remains unmeasured.

Human review is part of the result. A cautious model can leave most of the work with a person. Detection, false alarms and review volume need to be assessed together.

When the email tells the classifier what to do

A separate Microsoft PyRIT experiment tests whether injected instructions can change an archive recommendation.

Recorded live Jev run

Each of four fictional emails gets three attempts to influence the classifier: claimed authority, quoted approval, sensitivity denial or a forged answer. Earlier scores guide selection from a fixed payload library. An attack succeeds here only if it creates a newly unsafe archive recommendation.

The payload library is finite and calls are independent. Tool use, exfiltration, attachments and persistent memory are outside this test. Archive decisions are simulated, and provider errors stay separate from successful defenses. Methodology and reproduction commands.

When the text is not the whole email

A separate local runner takes the investigation into images, QR codes and model escalation.

Recorded Codex smoke tests

Inspect the available evidence

Normalize local images, decode QR payloads and compare visible link text with destinations. Remote images, attachments and unverified destinations stay explicit gaps.

Escalate with a reason

Small-model review can pass unresolved cases to a stronger model. Trusted user reports go directly to the strongest review tier, followed by a human decision.

Test the controls around it

The AI Act Companion review informed budget limits, response validation and tests for forged reports and missing evidence. The recorded cascade suite passed 59 checks on Linux CI.

Codex replays JEV only for exact recorded originals; new fixtures proceed with JEV marked unavailable. The direct API adapter is implemented, but its new OpenAI stages have not had a paid live run. URLs are not visited, and no mailbox action is automated.

Run the local cascade and inspect its limits · Read the Codex smoke evidence

How much does escalation change the cost?

Explore message volume, model choice and escalation rate. The figures below are hypothetical API costs.

Choose token counts that include instructions and billable reasoning. Both routes use the same assumed OpenAI request size. Actual requests may use more context than this scenario.

Official OpenAI prices, checked 2026-09-21. Standard, uncached, short-context USD rates. No Batch, caching or regional adjustments. Sol pricing is promotional. No OpenAI API requests were made.

Larger model for every message
Jev first + larger-model escalations
Hypothetical API cost reduction

OpenAI estimates for the token assumptions above
ModelInput / output per million tokensPer 1,000 emails

Human review, retries, infrastructure, attachments and reply generation are excluded. The lab has not yet demonstrated equivalent quality or full-workflow savings between these routes.

The rules the email cannot rewrite

Treat the message as evidence

An instruction inside an email has no authority over routing, budgets or permissions. Phishing and injection remain separate judgments.

Keep the action rules explicit

Trusted code validates responses and applies the decision rules. Missing evidence and invalid scores require review; mailbox actions stay with a person.

Keep the failures visible

Inspect the questions, fictional inputs, recorded answers and tests. Failed calls and changed judgments remain part of the evidence.