The EUTRANS-I corpus
收藏资源简介:
EUTRANS-I is a simple translation corpus which was produced and used in the EuTrans project. It corresponds to the so called "Traveller Task" which involves human-to-human communication situations in the front-desk of a hotel. Bilingual data were produced semi-automatically in three language pairs on the base of small "seed corpora", obtained from several traveler-oriented booklets. In using this corpus you agree that: 1. This corpus may be used free of cost for non commercial purposes. 2. If commercial use is intended, please first contact:<br> Dr. Enrique Vidal<br> Pattern Recognition and Language Technologies Research Center<br> Cno. de Vera s/n<br> E - 46071 Valencia<br> tel: +34 / 96 387 93 51<br> fax: +34 / 96 387 73 58<br> e-mail: evidal@prhlt.upv.es 3. In case of redistribution, this file must be distributed as is<br> along with the corpus. 4. Any publication derived from the use of this corpus must reference<br> the EuTrans project ("Example-based langUage TRANslation Systems",<br> EU Esprit #30268) and appropriate scholarly citation(s), such as: J.C. Amengual, J.M. Benedí, F. Casacuberta, A. Castaño,<br> A. Castellanos, V. Jiménez, D. Llorens, A. Marzal, F. Prat,<br> E. Vidal, and J.M. Vilar: "Using categories in the EuTrans<br> system". In ACL-ELSNET Workshoop on Spoken Language Translation,<br> pages 44-53, Madrid (Spain), July 1997. J.C.Amengual, J.M.Benedí F.Casacuberta, A.Castaño, A.Castellanos,<br> V.Jiménez, D.Llorens, A.Marzal, M.Pastor, F.Prat, E.Vidal,<br> J.M.Vilar: "The EuTrans-I Speech Translation System". Machine<br> Translation. Vol.15, pp.75-103, 2001. F.Casacuberta, H.Ney, F.J.Och, J.M.Vilar, E.Vidal, S.Barrachina,<br> I.García-Varea, C.Martínez D.Llorens, S.Molau, F.Nevado, M.Pastor,<br> D.Picó, A.Sanchís: "Some ap- proaches to statistical and<br> finite-state speech-to-speech translation". Computer Speech and<br> Language, Vol.18, pp.25-47, 2004. F.Casacuberta, E.Vidal, D.Picó: "Inference of finite-state<br> transducers from regular languages". Pattern Recognition. Vol.38,<br> pp.1431-1443, 2005. F.Casacuberta, E.Vidal: "Learning finite-state models for machine<br> translation". Machine Learning, Vol.66(1), pp.69-91, 2007.
EUTRANS-I是由EuTrans项目研发并投入使用的简易翻译语料库(corpus)。该语料库对应所谓的「旅客任务(Traveller Task)」,涵盖酒店前台场景下的人际交流情境。其双语数据以少量「种子语料库(seed corpora)」为基础,通过半自动化方式生成,共涵盖三组语言对,种子语料库取自多份面向旅客的宣传手册。 使用本语料库即表示您同意以下条款: 1. 本语料库可免费用于非商业用途。 2. 若拟将本语料库用于商业用途,请先联系: 恩里克·维达尔(Enrique Vidal)博士 模式识别与语言技术研究中心(Pattern Recognition and Language Technologies Research Center) 维拉大街s/n号 西班牙巴伦西亚,邮编E-46071 电话:+34 / 96 387 93 51 传真:+34 / 96 387 73 58 电子邮箱:evidal@prhlt.upv.es 3. 如需重新分发本语料库,必须保持本文件与语料库原始状态一并分发。 4. 基于本语料库开展研究形成的任何出版物,必须引用EuTrans项目(全称:基于示例的语言翻译系统"Example-based langUage TRANslation Systems",欧盟Esprit计划项目编号#30268)及以下合适的学术文献: 1. J.C. Amengual、J.M. Benedí、F. Casacuberta、A. Castaño、A. Castellanos、V. Jiménez、D. Llorens、A. Marzal、F. Prat、E. Vidal及J.M. Vilar:《在EuTrans系统中运用分类体系》,载于《ACL-ELSNET口语机器翻译研讨会论文集》,西班牙马德里,1997年7月,第44-53页。 2. J.C. Amengual、J.M. Benedí、F. Casacuberta、A. Castaño、A. Castellanos、V. Jiménez、D. Llorens、A. Marzal、M. Pastor、F. Prat、E. Vidal、J.M. Vilar:《EuTrans-I语音翻译系统》,载于《机器翻译》第15卷,第75-103页,2001年。 3. F. Casacuberta、H. Ney、F.J. Och、J.M. Vilar、E. Vidal、S. Barrachina、I. García-Varea、C. Martínez、D. Llorens、S. Molau、F. Nevado、M. Pastor、D. Picó、A. Sanchis:《统计与有限状态语音间翻译的若干方法》,载于《计算机语音与语言》第18卷,第25-47页,2004年。 4. F. Casacuberta、E. Vidal、D. Picó:《从正则语言推理有限状态换能器》,载于《模式识别》第38卷,第1431-1443页,2005年。 5. F. Casacuberta、E. Vidal:《学习面向机器翻译的有限状态模型》,载于《机器学习》第66卷第1期,第69-91页,2007年。



