ダウンロード数: 165

このアイテムのファイル:
ファイル 記述 サイズフォーマット 
3491065.pdf1.72 MBAdobe PDF見る/開く
タイトル: Linguistically-driven Multi-task Pre-training for Low-resource Neural Machine Translation
著者: Mao, Zhuoyuan
Chu, Chenhui  kyouindb  KAKEN_id  orcid https://orcid.org/0000-0001-9848-6384 (unconfirmed)
Kurohashi, Sadao  kyouindb  KAKEN_id
著者名の別形: 黒橋, 禎夫
キーワード: linguistically-driven
Low-resource neural machine translation
pre-training
発行日: Jul-2022
出版者: Association for Computing Machinery (ACM)
誌名: ACM Transactions on Asian and Low-Resource Language Information Processing
巻: 21
号: 4
開始ページ: 1
終了ページ: 29
論文番号: 68
抄録: In the present study, we propose novel sequence-to-sequence pre-training objectives for low-resource machine translation (NMT): Japanese-specific sequence to sequence (JASS) for language pairs involving Japanese as the source or target language, and English-specific sequence to sequence (ENSS) for language pairs involving English. JASS focuses on masking and reordering Japanese linguistic units known as bunsetsu, whereas ENSS is proposed based on phrase structure masking and reordering tasks. Experiments on ASPEC Japanese–English & Japanese–Chinese, Wikipedia Japanese–Chinese, News English–Korean corpora demonstrate that JASS and ENSS outperform MASS and other existing language-agnostic pre-training methods by up to +2.9 BLEU points for the Japanese–English tasks, up to +7.0 BLEU points for the Japanese–Chinese tasks and up to +1.3 BLEU points for English–Korean tasks. Empirical analysis, which focuses on the relationship between individual parts in JASS and ENSS, reveals the complementary nature of the subtasks of JASS and ENSS. Adequacy evaluation using LASER, human evaluation, and case studies reveals that our proposed methods significantly outperform pre-training methods without injected linguistic knowledge and they have a larger positive impact on the adequacy as compared to the fluency.
著作権等: © 2022 ACM. This is the author's version of the work. It is posted here for your personal use. Not for redistribution. The definitive Version of Record was published in 'ACM Transactions on Asian and Low-Resource Language Information Processing', 21(4), 68, pp1–29, http://dx.doi.org/10.1145/3491065.
This is not the published version. Please cite only the published version. この論文は出版社版でありません。引用の際には出版社版をご確認ご利用ください。
URI: http://hdl.handle.net/2433/267539
DOI(出版社版): 10.1145/3491065
出現コレクション:学術雑誌掲載論文等

アイテムの詳細レコードを表示する

Export to RefWorks


出力フォーマット 


このリポジトリに保管されているアイテムはすべて著作権により保護されています。