PsychEthicsEval is a shared task at MultiPsyche 2026 that assesses the ethical alignment of large language models (LLMs) in mental health contexts. Built on PsychEthicsBench and grounded in Australian psychology and psychiatry guidelines, it includes Multiple-Choice Questions (MCQs) to evaluate ethical knowledge and Open-Ended Questions (OEQs) to evaluate behavioral responses to ethically challenging scenarios.
The evaluation consists of two phases: Phase 1 provides 100 MCQs and 100 OEQs to help participants develop and debug their systems. On September 5th, we will release the unlabeled Phase 2 test set. Participants will have 48 hours to run inference and submit their system outputs. Performance on Phase 2 will determine the final ranking.
Phase 1 Evaluation Start: August 5th, 2026
Phase 2 Evaluation Start: September 5th, 2026
Result Submission Due: September 7th, 2026
Evaluation Result Release: September 10th, 2026
System Report Submission Due: September 14th, 2026
Phase 1 Dataset: [Download]
Please register for your team at: [TBA].
All models and trainable components must remain frozen. Prompt engineering, in-context learning with Phase 1 examples, or other prompting-based methods are allowed.
All Phase 2 MCQ and OEQ outputs must be generated automatically, without manual writing, editing, filtering, or selection.
Each team may submit one official Phase 2 entry. No performance feedback will be provided during the 48-hour evaluation period.
Participants must disclose their models, prompts, external resources, retrieval methods, and inference configurations. Top-ranked teams may be asked to provide code, prompts, or execution logs for verification.