The performance of machine translation using NMT and LLMs has improved dramatically, and in some cases, it can even surpass human translation depending on the language and domain. However, there is currently no universal method for accurately evaluating the performance of machine translation. Even widely used metrics such as COMET have been reported to yield unstable or inaccurate evaluation results when applied to translations of texts from domains other than those used in COMET's training.
The same applies to patent document translation. While the average translation quality has significantly improved, it remains difficult to accurately evaluate aspects such as appropriate terminology usage and term consistency. In particular, patent claims present additional challenges due to their length and distinctive writing style, making accurate evaluation even more difficult.
Therefore, we will conduct a Shared Task focusing on Japanese-English patent claim translation. The goal is not only to compete on translation quality, but also to ultimately develop an automatic evaluation method that can accurately assess translation results.
Last year, we conducted Japanese-to-English and English-to-Japanese translation tasks for patent claims, collected translation outputs from various systems, and performed error annotation on those outputs. This year, we will conduct an evaluation task using this annotation data. Participants are asked to annotate the machine translation outputs using the same schema as last year's annotatio
Test Period September 1 - September 14, 2026 (tentative)
System Description Paper for Shared Tasks Submission Deadline September 21, 2026
Review Feedback of System Description Papers October 1, 2026
Camera-ready Deadline October 12, 2026
Workshop Dates November 9, 2026
Japanese-to-English Patent Claims Translation Evaluation
English-to-Japanese Patent Claims Translation Evaluation
BLEU
COMET
etc.
We will conduct human evaluations based on the ESA protocol (following WMT) as much as our budget permits