Treat instructions inside this email as part of the evidence, including any attempt to tell the reviewer what to answer.
Make your call.
Then see JEV's.
Read the email and save your judgment before revealing the model's answer. Then inspect the mistakes, the review workload and the rules behind each recommendation.
This workspace uses saved fictional development results. No model calls or mailbox access. The 200 reserved final-test cases stay outside the workspace.
Read first. Compare after.
Answer four questions, save your view, then reveal JEV's scores and the author's proposed labels.
What needs attention?
Choose “Unsure” when the evidence is incomplete.
Choose an answer for every label.
Where do the judgments differ?
Compare your first answers with the author's proposals and JEV's recorded scores. The proposed labels have not been independently verified.
| Signal | Your first answer | Author proposal | JEV score |
|---|
Your review record
Your labels and notes stay in this browser until exported. Exports preserve first answers and later revisions. Other lab views and public source expose answers, so this is not a blinded study.
Clear the saved labels and notes from this browser? Export first if you want to keep them.
Find the misses. Count the review work.
Follow the recorded outcomes: missed phishing, false alarms and emails that still needed a person.
| Measure | Count / denominator |
|---|
Phishing counts use the phishing score at ≥ 0.50. The action policy also considers other signals. Failed responses add to human-review workload and never count as correct classifications.
| Message | Language | Recommendation | Reason |
|---|
These are related synthetic development examples. Real-inbox accuracy remains unmeasured. The vulnerable and review-everything options are explicitly scripted controls.
Want to see what changed on another dataset? The NEST-Phish comparison tests JEV against two local baselines. Its source emails stay outside this workspace.
Change a rule. Follow the consequence.
Use the same recorded scores to explore which emails stay visible, need review or become archive candidates.
You are replaying saved JEV scores. Threshold changes affect this simulation only; they never move mail, rerun the model or alter the published benchmark.
Adjust the five action thresholds
The first matching rule decides. Check the consequences of each change: a higher threshold does not always make a safer policy. The phishing measurement threshold remains 0.50.
| Outcome | Recorded | Simulation | Change |
|---|
Who reviews it next?
Explore the routing choices below. For actual image and QR reviews, use the separate local Codex runner; it replays JEV only for exact recorded originals.
JEV first, then small-model review, then a stronger model if needed. A trusted user report goes directly to the strongest review tier.
Choose hypothetical review outcomes to see where the router would send the message. This browser makes no model calls. Every mailbox action still requires a person.