Monday, February 18, 2013

LDC February 2013 Newsletter


New publications:



Spring 2013 LDC Data Scholarship Recipients! 

LDC is pleased to announce the student recipients of the Spring 2013 LDC Data Scholarship program! This program provides university students with access to LDC data at no-cost. Students were asked to complete an application which consisted of a proposal describing their intended use of the data, as well as a letter of support from their thesis adviser. We received many solid applications and have chosen three proposals to support. The following students will receive no-cost copies of LDC data:
Salima Harrat - Ecole SupĂ©rieure d’informatique (ESI) (Algeria). Salima has been awarded a copy of Arabic Treebank: Part 3 for her work in diacritization restoration.

Maulik C. Madhavi - Dhirubhai Ambani Institute of Information and Communication Technology (DA-IICT), Gandhinagar (India). Maulik has been awarded a copy of
Switchboard Cellular Part 1 Transcribed Audio and Transcripts and 1997 HUB4 English Evaluation Speech and Transcripts for his work in spoken term detection.

Shereen M. Oraby - Arab Academy for Science, Technology, and Maritime Transport (Egypt). Shereen has been awarded a copy of
Arabic Treebank: Part 1 for her work in subjectivity and sentiment analysis.
Please join us in congratulating our student recipients! The next LDC Data Scholarship program is scheduled for the Fall 2013 semester.

Membership Fee Savings and Publications Pipeline 

Time is quickly running out to save on membership fees for Membership Year 2013 (MY2013)! Any organization which joins or renews membership for 2013 through Friday, March 1, 2013, is entitled to a 5% discount on membership fees.  Organizations which held membership for MY2012 can receive a 10% discount on fees provided they renew prior to March 1, 2013.

Many publications for MY2013 are still in development. The planned publications for the upcoming months include:
GALE data ~ continuing releases of all languages (Arabic, Chinese, English), genres (Broadcast News, Broadcast Conversation, Newswire and Web Data) and tasks (Parallel Text, Word Alignment, Parallel Aligned Treebanks, Parallel Sentences, Audio and Transcripts).
Hispanic Accented English Database ~ 30 hours of conversational speech data from non-native speakers of English with approximately 24 hours or 80% of the data  closely transcribed. The speech in this release was collected from 22  non-native, Hispanic speakers of English and consists of spontaneous speech and read utterances. The read speech is divided equally between English and Spanish.
NIST 2012 Open Machine Translation  Progress Tests ~ contains the evaluation sets (source data and human reference translations), DTD, scoring software, and evaluation plan for the OpenMT12 test for Arabic, Chinese, Dari, Farsi, and Korean to English on a parallel data set.  This set is based on a subset of the Arabic-to-English and Chinese-to-English Progress tests from the NIST Open Machine Translation 2008, 2009, and 2012 evaluations with new source data created based on the English human reference translation reference. The original data consists of newswire and web data.
NIST Open Machine Translation 2008 to 2012 Progress Test Sets ~ contains the evaluation sets (source data and human reference translations), DTD, scoring software, and evaluation plans for the Arabic-to-English and Chinese-to-English Progress tests of the NIST Open Machine Translation 2008, 2009, and 2012 Evaluations.  The test sets consist of newswire and web data.
OntoNotes 5.0 ~ multiple genres of English, Chinese, and Arabic text annotated for syntax, predicate argument structure and shallow semantics.
UN Parallel Text ~ contains the text of United Nations parliamentary documents in Arabic, Chinese, English, French, Russian, and Spanish from 1993 through 2007. The data is provided in two formats:  (1) raw text: the raw text is very close to what was extracted from the word processing documents, converted to UTF-8 encoding,  and (2) word-aligned text: the word-aligned text has been normalized, tokenized, aligned at the sentence-level, further broken into sub-sentential "chunk-pairs", and then aligned at the word-level.
2013 Subscription Members are automatically sent all MY2013 data as it is released.  2013 Standard Members are entitled to request 16 corpora for free from MY2013. Non-members may license most data for research use. Visit our Announcements page for information on pricing.

New LDC Podcast, LDC Executive Director, Christopher Cieri

The
LDC blog has a new podcast in LDC’s 20th Anniversary series. This edition features LDC’s Executive Director, Christopher Cieri. In this podcast, Chris reflects on the road that took him to LDC, some of his early responsibilities and recent consortium activities. 

Click
here for Chris’ podcast. Other podcasts will be published via the LDC blog , so stay tuned to that space.
New publications

(1) GALE Phase 2 Arabic Broadcast Conversation Speech Part 1 was developed by LDC and is comprised of approximately 123 hours of Arabic broadcast conversation speech collected in 2006 and 2007 by LDC as part of the DARPA GALE (Global Autonomous Language Exploitation) Program. Broadcast audio for the DARPA GALE program was collected at LDC’s Philadelphia, PA USA facilities and at three remote collection sites. 

The combined local and outsourced broadcast collection supported GALE at a rate of approximately 300 hours per week of programming from more than 50 broadcast sources for a total of over 30,000 hours of collected broadcast audio over the life of the program.

LDC's local broadcast collection system is highly automated, easily extensible and robust and capable of collecting, processing and evaluating hundreds of hours of content from several dozen sources per day. The broadcast material is served to the system by a set of free-to-air (FTA) satellite receivers, commercial direct satellite systems (DSS) such as DirecTV, direct broadcast satellite (DBS) receivers, and cable television (CATV) feeds. The mapping between receivers and recorders is dynamic and modular; all signal routing is performed under computer control, using a 256x64 A/V matrix switch. Programs are recorded in a high bandwidth A/V format and are then processed to extract audio, to generate keyframes and compressed audio/video, to produce time-synchronized closed captions (in the case of North American English) and to generate automatic speech recognition (ASR) output. 

The broadcast conversation recordings in this release feature interviews, call-in programs and round table discussions focusing principally on current events from several sources. This release contains 143 audio files presented in .wav, 16000 Hz single-channel 16-bit PCM. Each file was audited by a native Arabic speaker following Audit Procedure Specification Version 2.0 which is included in this release. The broadcast auditing process served three principal goals: as a check on the operation of LDCs broadcast collection system equipment by identifying failed, incomplete or faulty recordings; as an indicator of broadcast schedule changes by identifying instances when the incorrect program was recorded; and as a guide for data selection by retaining information about a program's genre, data type and topic.

GALE Phase 2 Arabic Broadcast Conversation Speech Part 1 is distributed on 4 DVDs.
2013 Subscription Members will automatically receive two copies of this data. 2013 Standard Members may request a copy as part of their 16 free membership corpora.

*

(2) GALE Phase 2 Arabic Broadcast Conversation Transcripts - Part 1 was developed by LDC and contains transcriptions of approximately 123 hours of Arabic broadcast conversation speech collected in 2006 and 2007 by LDC, MediaNet, Tunis, Tunisia and MTC, Rabat, Morocco during Phase 2 of the DARPA GALE (Global Autonomous Language Exploitation) program. The source broadcast conversation recordings feature interviews, call-in programs and round table discussions focusing principally on current events from several sources.

The transcript files are in plain-text, tab-delimited format (TDF) with UTF-8 encoding, and the transcribed data totals 752,747 tokens. The transcripts were created with the LDC-developed transcription tool, XTrans, a multi-platform, multilingual, multi-channel transcription tool that supports manual transcription and annotation of audio recordings. 

The files in this corpus were transcribed by LDC staff and/or by transcription vendors under contract to LDC. Transcribers followed LDCs quick transcription guidelines (QTR) and quick rich transcription specification (QRTR) both of which are included in the documentation with this release. QTR transcription consists of quick (near-)verbatim, time-aligned transcripts plus speaker identification with minimal additional mark-up. It does not include sentence unit annotation. QRTR annotation adds structural information such as topic boundaries and manual sentence unit annotation to the core components of a quick transcript. Files with QTR as part of the filename were developed using QTR transcription. Files with QRTR in the filename indicate QRTR transcription.

GALE Phase 2 Arabic Broadcast Conversation Transcripts - Part 1 is distributed via web download. 2013 Subscription Members will automatically receive two copies of this data on disc. 2013 Standard Members may request a copy as part of their 16 free membership corpora.
*

(3) NIST 2012 Open Machine Translation (OpenMT) Evaluation was developed by NIST Multimodal Information Group. This release contains source data, reference translations and scoring software used in the NIST 2012 OpenMT evaluation, specifically, for the Chinese-to-English language pair track. The package was compiled and scoring software was developed at NIST, making use of Chinese newswire and web data and reference translations collected and developed by LDC. The objective of the OpenMT evaluation series is to support research in, and help advance the state of the art of, machine translation (MT) technologies -- technologies that translate text between human languages. Input may include all forms of text. The goal is for the output to be an adequate and fluent translation of the original. 

The 2012 task was to evaluate five language pairs: Arabic-to-English, Chinese-to-English, Dari-to-English, Farsi-to-English and Korean-to-English. This release consists of the material used in the Chinese-to-English language pair track. For more general information about the NIST OpenMT evaluations, please refer to the NIST OpenMT website.

This evaluation kit includes a single Perl script (mteval-v13a.pl) that may be used to produce a translation quality score for one (or more) MT systems. The script works by comparing the system output translation with a set of (expert) reference translations of the same source text. Comparison is based on finding sequences of words in the reference translations that match word sequences in the system output translation.

This release contains 222 documents with corresponding source and reference files, the latter of which contains four independent human reference translations of the source data. The source data is comprised of Chinese newswire and web data collected by LDC in 2011. A portion of the web data concerned the topic of food and was treated as a restricted domain. The table below displays statistics by source, genre, documents, segments and source tokens.

Source
Genre
Documents
Segments
Source Tokens
Chinese General
Newswire
45
400
18184
Chinese General
Web Data
28
420
15181
Chinese Restricted Domain
Web Data
149
2184
48422

The token counts for Chinese data are "character" counts, which were obtained by counting tokens matching the UNICODE-based regular expression "/w". The Python “re” module was used to obtain those counts.

NIST 2012 Open Machine Translation (OpenMT) Evaluation is distributed via web download. 2013 Subscription Members will automatically receive two copies of this data on disc. 2013 Standard Members may request a copy as part of their 16 free membership corpora.

Tuesday, January 29, 2013

LDC 20th Anniversary Podcast: Christopher Cieri



LDC is moving towards the end of its Anniversary year, but that does not mean that we don’t have a few more treats for you. This month’s podcast features LDC’s Executive Director, Christopher Cieri.

Chris is involved with every aspect of the Consortium, including planning, development, operations, sponsored projects, external relations and financial performance. In this podcast, Chris reflects on the road that took him to LDC, some of his early responsibilities and recent consortium activities. 


Tuesday, January 15, 2013

LDC January 2013 Newsletter


New publications:


2013 LDC Podcast Available from LDC Blog

Kicking off the new year is the fourth podcast in our 20th Anniversary series featuring LDC Senior Researcher, Mohamed Maamouri.

Mohamed directs the Arabic Treebank group and spearheads the development of Arabic resources and projects. The latter includes the leading role in LDC’s collaboration with Georgetown University Press to develop updated versions of three dialectal Arabic dictionaries (Iraqi, Moroccan, Syrian). In this podcast, he reflects on his personal and professional experiences and comments on Arabic resource development at LDC. 

Click here for Mohamed’s podcast. 

Other podcasts will be published via the LDC Blog, so stay tuned to that space.

Membership Discounts for MY 2013 Still Available

If you are considering joining for Membership Year 2013 (MY2013), there is still time to save on membership fees. Any organization which joins or renews membership for 2013 through Friday, March 1, 2013, is entitled to a 5% discount on membership fees. Organizations which held membership for MY2012 can receive a 10% discount on fees provided they renew prior to March 1, 2013. For further information on pricing, please consult our Announcements page or contact LDC.

    Penn Discourse Treebank Version 2.0 Update - RTE data
A Recognizing Textual Entailment (RTE) update is now available for Penn Discourse Treebank Version 2.0 LDC2008T05 (PDTB). This data has been used to run the textual entailment experiments described in: Sara Tonelli and Elena Cabrio "Hunting for Entailing Pairs in the Penn Discourse Treebank", in Proceedings of Coling 2012, Mumbay, India. The files contain Text - Hypothesis pairs in the standard RTE xml format (for more details, see  RTE Challenge at TAC 2011), which have been manually annotated as entailing or not entailing. All sentence pairs have been extracted from the Penn Discourse Treebank and are therefore connected by a discourse relation label.

The data are not included in the general release of Penn Discourse Treebank Version 2.0, but are freely available for download from the catalog page. 

New Publications

(1) Chinese-English Biology and Chemistry Abstract Parallel Text was developed by The MITRE Corporation. It consists of parallel sentences from a collection of chemistry and biology-related scientific article abstracts published in Mandarin and translated into English by translators with particular expertise in the technical area. Translators were instructed to err on the side of literal translation if required, but to maintain the technical writing style of the source and make the resulting English as natural as possible. The translators were given specific guidelines for translation, and those are included in this distribution.

This release contains 2,239 lines of parallel Mandarin and English, with a total of 156,445 characters of Mandarin and 75,515 words of English, presented in a separate UTF-8 plain text file for each language. The sentences were translated in sequential order and presented in scrambled order, such that parallel sentences at identical line numbers are translations. For example, the 31st line of the English file is a translation of the 31st line of the Mandarin file. The original line sequence is not provided.

Chinese-English Biology and Chemistry Abstract Parallel Text is distributed via web download. 2013 Subscription Members will automatically receive two copies of this data on disc. 2013 Standard Members may request a copy as part of their 16 free membership corpora.

*

(2) GALE Phase 2 Arabic Web Parallel Text was developed by LDC. Along with other corpora, the parallel text in this release comprised training data for Phase 2 of the DARPA GALE (Global Autonomous Language Exploitation) Program. This corpus contains Modern Standard Arabic source text and corresponding English translations selected from web data collected in 2007 by LDC and transcribed by LDC or under its direction. GALE Phase 2 Arabic Web Parallel Text includes 60 source-translation document pairs, comprising 42,089 words of Arabic source text and its English translation. Data was drawn from various Arabic weblog and newsgroup sources. 

The files in this release were transcribed by LDC staff and/or transcription vendors under contract to LDC in accordance with the Quick Rich Transcription guidelines developed by LDC. Transcribers indicated sentence boundaries in addition to transcribing the text. Data was manually selected for translation according to several criteria, including linguistic features, transcription features and topic features. The transcribed and segmented files were then reformatted into a human-readable translation format and assigned to translation vendors. Translators followed LDC's Arabic to English translation guidelines. 

Bilingual LDC staff performed quality control procedures on the completed translations. Source data and translations are distributed in TDF format. TDF files are tab-delimited files containing one segment of text along with meta information about that segment.

GALE Phase 2 Arabic Web Parallel Text is distributed via web download. 2013 Subscription Members will automatically receive two copies of this data on disc. 2013 Standard Members may request a copy as part of their 16 free membership corpora. 

Friday, January 11, 2013

LDC 20th Anniversary Podcast: Mohamed Maamouri



Happy New Year and welcome back to the LDC Blog. For our first post of the year, we present the fourth podcast in our anniversary series featuring LDC Senior Researcher, Mohamed Maamouri.

Mohamed directs the Arabic Treebank group and spearheads the development of Arabic resources and projects. The latter includes the leading role in LDC’s collaboration with Georgetown University Press to develop updated versions of three dialectal Arabic dictionaries (Iraqi, Moroccan, Syrian). Mohamed specializes in Arabic linguistics, reading, language development, corpus linguistics and sociolinguistics. In this podcast, he reflects on his personal and professional experiences and comments on Arabic resource development at LDC.

Tuesday, December 18, 2012

LDC December 2012 Newsletter



Two New LDC Podcasts for your Listening Pleasure 

New publications:




 Spring 2013 LDC Data Scholarship Program - deadline approaching!

The deadline for the Spring 2013 LDC Data Scholarship Program is one month away!   Student applications are being accepted now through January 15, 2013, 11:59PM EST.  The LDC Data Scholarship program provides university students with access to LDC data at no cost.  This program is open to students pursuing both undergraduate and graduate studies in an accredited college or university. LDC Data Scholarships are not restricted to any particular field of study; however, students must demonstrate a well-developed research agenda and a bona fide inability to pay. 

Students will need to complete an application which consists of a data use proposal and letter of support from their adviser.  For further information on application materials and program rules, please visit the LDC Data Scholarship page.  

Students can email their applications to the LDC Data Scholarship program. Decisions will be sent by email from the same address.

Two New LDC Podcasts for your Listening Pleasure

Two new podcasts are available on a the LDC blog continuing the  20th Anniversary series. The first features Natalia Bragilevskaya, LDC’s Business Administrator, Membership Coordinator Ilya Ahtaridis and Marian Reed, Marketing Coordinator. They recall the early days of LDC and describe the growth of sponsored projects work and LDC’s interactions with its membership.

Click here for Natalia, Ilya and Marian’s podcast.

The third podcast in the series introduces the community to two  LDC, researchers Yiwola Awoyale and Moussa Bamba, whose work focuses on West African languages. 

Yiwola has been teaching Linguistics, Yoruba language studies and various aspects of African linguistics since 1975. At LDC, he developed the Global Yoruba Lexical Database, a set of related dictionaries based on Yoruba and its diaspora. Moussa’s work in the Manding languages of the Niger-Congo family has resulted in the release of the Mawukakan Lexicon, to be followed by similar resources for Maninkakan, Bambara, and Jula. 

In their podcast, Yiwola and Moussa discuss how they came  to LDC, their current research and how it benefits multiple communities. Click here for Yiwola and Moussa’s podcast. 

Other podcasts will be published via the LDC blog, so stay tuned to that space.

Penn Discourse Treebank Version 2.0 Update

The developers of the Penn Discourse Treebank Version 2.0 LDC2008T05 (PDTB) have updated this release to add metadata to the Wall Street Journal (WSJ) news stories in the corpus. The goal is to aid understanding PDTB files as texts and to support distinguishing texts from different genres within the WSJ. 

The metadata includes the following fields:
  • DD: the date the article appeared in the WSJ
  • AN: unique identifier for the article
  • HL: the column name (for regular features such as Who's News, Marketing & Media, Technology), its headline and by-line
  • SO: the source of the article
  • IN: manually-assigned codes or keywords for the article
  • CO: manually-assigned codes for companies or other organizations
  • DATELINE: normally the location where the article was filed, but sometimes has very unexpected contents
  • GV: Branch of Government or Government Agency mentioned in the article
  • SBREAKS: the byte position of section breaks present in the file
  • ARTICLEBREAK: separates files that contain more than one article
All new downloads of PDTB will contain the complete updated corpus.  Current PDTB licensees can re-download the file to obtain the updated data. 

LDC to close for Winter Break
LDC will be closed from Monday, December 24, 2012 through Tuesday, January 1, 2013 in accordance with the University of Pennsylvania Winter Break Policy. Our offices will reopen on Wednesday, January 2, 2013. Requests received for membership renewals and corpora during the Winter Break will be processed at that time.

Best wishes for a happy and safe holiday season!

New publications
(1) GALE Chinese-English Word Alignment and Tagging Training Part 3 -- Web was developed by LDC and contains 154,541 tokens of word aligned Chinese and English parallel text enriched with linguistic tags. This material was used as training data in the DARPA GALE (Global Autonomous Language Exploitation) program.

Some approaches to statistical machine translation include the incorporation of linguistic knowledge in word aligned text as a means to improve automatic word alignment and machine translation quality. This is accomplished with two annotation schemes: alignment and tagging. Alignment identifies minimum translation units and translation relations by using minimum-match and attachment annotation approaches. A set of word tags and alignment link tags are designed in the tagging scheme to describe these translation units and relations. Tagging adds contextual, syntactic and language-specific features to the alignment annotation. 

GALE Chinese-English Word Alignment and Tagging Training Part 1 -- Newswire and Web (LDC2012T16) and GALE Chinese-English Word Alignment and Tagging Training Part 3 -- Web (LDC2012T20) are also available through LDC.

This release consists of Chinese source web data (newsgroup, weblog) collected by LDC in 2008 and 2009. The distribution by words, character tokens and segments appears below: 

Language: Chinese
Files: 1249
Words: 103027
CharTokens: 154541 
Segments: 4842


Note that all token counts are based on the Chinese data only. One token is equivalent to one character and one word is equivalent to 1.5 characters.
The Chinese word alignment tasks consisted of the following components:
  • Identifying, aligning, and tagging 8 different types of links
  • Identifying, attaching, and tagging local-level unmatched words
  • Identifying and tagging sentence/discourse-level unmatched words
  • Identifying and tagging all instances of Chinese çš„(DE) except when they were a part of a semantic link.
GALE Chinese-English Word Alignment and Tagging Training Part 3 -- Web is distributed via web download.2012 Subscription Members will automatically receive two copies of this data on disc. 2012 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for US$1750.
*
(2) Russian-English Computer Security Parallel Text was developed by The MITRE Corporation. It consists of parallel sentences from a set of computer security reports published in Russian and translated into English by translators with particular expertise in the technical area. Translators were instructed to err on the side of literal translation if required, but to maintain the technical writing style of the source and to make the resulting English as natural as possible. The translators followed specific guidelines for translation, and those are included in this distribution.

There are 6,276 lines of parallel Russian and English, with a total of 60,059 words of Russian and 76,437 words of English, presented in a separate UTF-8 plain text file for each language. The sentences were translated in sequential order and presented in a scrambled order, such that parallel sentences at identical line numbers are translations. For example, the 31st line of the English file is a translation of the 31st line of the Russian file. The original line sequence is not provided. 1,694 untranslated lines (such as code snippets) are included as a separate file.

Russian-English Computer Security Parallel Text is distributed via web download. 2012 Subscription Members will automatically receive two copies of this data on disc. 2012 Standard Members may request a copy as part of their 16 free membership corpora.  Non-members may license this data for US$1500.