Friday, May 15, 2020

LDC 2020 May Newsletter

New Publications:
_______________________________________________________________ 

New publications: 

(1) LORELEI Oromo Incident Language Pack was developed by LDC and is comprised of approximately 3.9 million words of Oromo monolingual text, 25,000 words of English monolingual text, 135,000 words of parallel and comparable Oromo-English text, and 50,000 words of data annotated for Entity Discovery and Linking and Situation Frames. It contains all of the text data, annotations, supplemental resources and related software tools for the Oromo language that were used in the DARPA LORELEI / LoReHLT 2017 Evaluation. 

The evaluation protocol was based on a scenario in which an unforeseen event triggered a need for humanitarian and logistical support in a region where the incident language had received little or no attention in NLP research. Evaluation participants provided NLP solutions, including information extraction and machine translation, with limited resources and limited development time.

Data was collected from news, social network, weblog, newsgroup, discussion forum, and reference material. Entity Detection and Linking and Situation Frame annotations identified “entities,” “needs” (such as a need for food) and “issues” (such as civil unrest) to be detected by systems for scoring purposes. Situation frame analysis was designed to extract basic information that would be useful for planning a disaster response effort. 

The knowledge base for the entity linking annotation in this corpus is available separately as LORELEI Entity Detection and Linking Knowledge Base (LDC2020T10).

LORELEI Oromo Incident Language Pack is distributed via web download.

2020 Subscription Members will automatically receive copies of this corpus. 2020 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.

*

(1) LORELEI Entity Detection and Linking Knowledge Base was developed by LDC and contains the full LORELEI Entity Detection and Linking (EDL) Knowledge Base (KB) used for all LORELEI Representative Language and Incident Language Pack entity linking annotation. The LORELEI (Low Resource Languages for Emergent Incidents) Program was concerned with building human language technology for low resource languages in the context of emergent situations like natural disasters or disease outbreaks. 

The KB in this release supported the EDL task in LORELEI for four entity types -- geo-political entities (GPE), locations (LOC), persons (PER) and organizations (ORG) -- and contains a total of 10,216,832 entities. There are four inputs to the KB, each designated by a unique "origin" code in the KB, as follows: GPE and LOC entities from a snapshot of GeoNames, PER entities from the CIA World Leaders List, ORG entities from Appendix B of the CIA World Factbook, and additional entities manually created by LDC for each of the representative and incident languages in the LORELEI Program. 

LORELEI Entity Detection and Linking Knowledge Base is distributed via web download.

2020 Subscription Members will automatically receive copies of this corpus. 2020 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.

*

(3) BOLT English Translation Treebank - Chinese Discussion Forum was developed by LDC and consists of 147,432 tokens of web discussion forum data translated from Chinese to English and annotated for part-of-speech and syntactic structure. 

The source data is Chinese discussion forum web text collected by LDC in 2011 and 2012, translated into English and released in BOLT Chinese Discussion Forum Parallel Training Data (LDC2017T05). A subset of the translated text -- 148 files representing 147,432 tokens -- was selected for the treebank and annotated for word-level tokenization, part-of-speech and syntactic structure. Only the translated English text is included in the source data for this release. 

Part-of-speech and treebank annotation conformed to Penn Treebank II style, incorporating changes to those guidelines that were developed under the GALE (Global Autonomous Language Exploitation) program. Supplementary guidelines for English treebanks and web text are included with this release.

BOLT English Translation Treebank - Chinese Discussion Forum is distributed via web download.

2020 Subscription Members will automatically receive copies of this corpus. 2020 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.

*

(4) Multi-Language Conversational Telephone Speech 2011 -- Mandarin Chinese was developed by LDC and is comprised of approximately 25 hours of telephone speech in Mandarin Chinese.

The data were collected primarily to support research and technology evaluation in automatic language identification, and portions of these telephone calls were used in the NIST 2011 Language Recognition Evaluation (LRE). Participants were recruited by native speakers who contacted acquaintances in their social network. Those native speakers made one call, up to 15 minutes, to each acquaintance. The data was collected using LDC's telephone collection infrastructure, comprised of three computer telephony systems. Human auditors labeled calls for callee gender, dialect type and noise. 

LDC has also released the following as part of the Multi-Language Conversational Telephone Speech 2011 series:
Multi-Language Conversational Telephone Speech 2011 -- Mandarin Chinese is distributed via web download. 

2020 Subscription Members will automatically receive copies of this corpus. 2020 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee. 

*

Wednesday, April 15, 2020

LDC 2020 April Newsletter

New Publications:
2018 NIST Speaker Recognition Evaluation Test Set 
Abstract Meaning Representation 2.0 - Four Translations
TAC KBP English Temporal Slot Filling - Comprehensive Training and Evaluation Data 2011 and 2013 
________________________________________________________________ 

New publications: 

(1) 2018 NIST Speaker Recognition Evaluation Test Set was developed by LDC and NIST (National Institute of Standards and Technology) and contains approximately 396 hours of Tunisian Arabic telephone recordings and English web video speech used as development and test data in the NIST-sponsored 2018 Speaker Recognition Evaluation (SRE). This release also contains answer keys, trial and train files, development data and evaluation documentation. 

The SRE task is speaker detection, that is, to determine whether a specified target speaker is speaking during a segment of speech. In addition to the traditional focus on conversational telephone speech recorded over a variety of handset types for the training and test conditions, SRE18 added VOIP (voice over IP) data and audio from video.

The telephone speech data was drawn from the Call My Net 2 (CMN2) collection conducted by LDC in Tunisia in which recruited Tunisian Arabic speakers made multiple calls to friends or relatives for conversations lasting between 8-10 minutes. The speech segments include PSTN (public switched telephone network) and VOIP data.

The English audio was sampled from amateur web videos collected by LDC as part of the Video Annotation for Speech Technology (VAST) project.

2018 NIST Speaker Recognition Evaluation Test Set is distributed via web download. 

2020 Subscription Members will automatically receive copies of this corpus. 2020 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee. 

*

(2) Abstract Meaning Representation 2.0 - Four Translations was developed by researchers at the University of Edinburgh, School of Informatics and consists of Spanish, German, Italian and Chinese Mandarin translations of 5,484 test split sentences (1,371 sentences per language) from Abstract Meaning Representation (AMR) Annotation Release 2.0 (LDC2017T10).

AMR Annotation Release 2.0 is a semantic treebank of over 39,000 English natural language sentences from broadcast conversations, newswire and web text. The translated data in this release was designed for use in cross-lingual parsing.

The source sentences were drawn from material collected by LDC, specifically, discussion forum text from the DARPA BOLT and DARPA DEFT programs, transcripts and English translations of Mandarin Chinese broadcast news programming, Wall Street Journal text, translated Xinhua news texts, various newswire texts from NIST OpenMT evaluations and weblog data from the DARPA GALE program.

Abstract Meaning Representation 2.0 - Four Translations is distributed via web download.

2020 Subscription Members will automatically receive copies of this corpus. 2020 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.

*

(3) TAC KBP English Temporal Slot Filling - Comprehensive Training and Evaluation Data 2011 and 2013 was developed by LDC and contains training and evaluation data produced in support of the TAC KBP English Temporal Slot Filling tasks in 2011 and 2013. This release includes queries, manual runs produced by LDC annotators, and the final rounds of assessment results. 

The goal of the Temporal Slot Filling task was to identify and capture temporal information in text indicating when a given relation between a slot filling query entity and filler held true. This built upon the technology developed for regular Slot Filling which involved mining information about entities from text.

The corresponding source data collections of English newswire, broadcast material and web text are included in TAC KBP Comprehensive English Source Corpora 2009-2014 (LDC2018T03). The corresponding Knowledge Base (KB) for much of the data - a 2008 snapshot of Wikipedia - is available in TAC KBP Reference Knowledge Base (LDC2014T16).

TAC KBP English Temporal Slot Filling - Comprehensive Training and Evaluation Data 2011 and 2013 is distributed via web download.

2020 Subscription Members will automatically receive copies of this corpus. 2020 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee. 

*

Friday, March 13, 2020

LDC 2020 March Newsletter

Spring 2020 LDC Data Scholarship recipients 
LDC data and commercial technology development

__________________________________________________________________ 

Spring 2020 LDC Data Scholarship recipients 

LDC congratulates the following Spring 2020 Data Scholarship recipients: 
  • Zahra Azin (Istanbul Technical University, Turkey) is awarded a copy of Abstract Meaning Representation (AMR) Annotation Release 3.0 (LDC2020T02) for her work in Turkish AMR.
  •  Spandan Dey (IIT Kharagpur, India) is awarded a copy of Multi-Language Conversational Telephone Speech – South Asian (LDC2017S14) for his research on automatic language recognition.
  • Jonathan Downey (University of California, Santa Barbara, US) is awarded a copy of the ETS Corpus of Non-Native Written English (LDC2014T06) for his research on second language acquisition and quantitative methodologies for educational measurements.
  • Nathaniel Fackler (University of Georgia, US) is awarded a copy of the ETS Corpus of Non-Native Written English (LDC2014T06) for his work on adult second language acquisition.
  • B. Senthil Kumar (SSN College of Engineering & Anna University, India) is awarded a copy of 2009 CoNLL Shared Task Part 2 (LDC2012T04) for his research on semantic role labeling.
  • Ming Li (Colorado School of Mines, US) is awarded a copy of TIDIGITS (LDC93S10) for her research on inferring speech signals from motion data in Internet of Things (IoT) security.
  • Jialiang Lin (Xiamen University, China) is awarded a copy of the ETS Corpus of Non-Native Written English (LDC2014T06) for his project to train and test an automated essay scoring model.
Students can learn more about the LDC Data Scholarship program and the next application cycle on the Data Scholarships page. 

LDC data and commercial technology development

For-profit organizations are reminded that an LDC membership is a pre-requisite for obtaining a commercial license to almost all LDC databases. Non-member organizations, including non-member for-profit organizations, cannot use LDC data to develop or test products for commercialization, nor can they use LDC data in any commercial product or for any commercial purpose. LDC data users should consult corpus-specific license agreements for limitations on the use of certain corpora. Visit the Licensing page for further information.
__________________________________________________________________ 

New publications: 

(1) BOLT Egyptian Arabic-English Word Alignment -- Conversational Telephone Speech Training was developed by LDC and consists of 153,171 words of Egyptian Arabic and English parallel text enhanced with linguistic tags to indicate word relations.

The source data in this release consists of transcripts of Egyptian Arabic conversational telephone speech (CTS) from LDC's CALLHOME and CALLFRIEND collections (LDC97S45, LDC97T19, LDC2002S37, LDC2002T38, LDC96S49) that was translated into English by professional translation agencies and annotated for the word alignment task.

The BOLT word alignment task was built on treebank annotation. Egyptian Arabic source tree tokens were automatically extracted from tree files in LDC’s BOLT Egyptian Arabic Treebank, which had been tagged for part-of-speech and syntactically annotated. That data was then aligned and annotated for the word alignment task.

BOLT Egyptian Arabic-English Word Alignment -- Conversational Telephone Speech Training is distributed via web download. 

2020 Subscription Members will automatically receive copies of this corpus. 2020 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee. 



(2) EVALution was developed by The Hong Kong Polytechnic University. It is comprised of English and Mandarin Chinese data sets -- EVALution 1.0 and EVALution-Man, respectively -- that contain semantic relations and metadata for training and evaluating distributional semantic models. 

EVALution 1.0 consists of approximately 7500 English tuples extracted from ConceptNet 5.0 and WordNet 4.0 and filtered through automatic methods and crowd-sourcing. Several semantic relations between word pairs were instantiated, including hypernymy, synonymy, antonymy and meronymy. The corpus also includes additional information that can be used to filter the pairs or to analyze the results, such as relation domain, word frequency, word part-of-speech and word semantic field.

EVALution-MAN consists of Chinese word pairs from two sources: Chinese Wordnet and humans who completed an elicitation task by supplying missing words to sentences. The human-supplied sentence word pairs were then judged by human raters for reliability. 

EVALution is distributed via web download. 

2020 Subscription Members will receive copies of this corpus provided they have submitted a completed copy of the special license agreement. 2020 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.

*

(3) Mixer 4 and 5 Speech was developed by LDC and contains approximately 14,185 hours of audio recordings of conversational telephone speech, interviews, elicitation exercises and transcript readings involving 616 distinct speakers. The material was collected in 2007 as part of the Mixer project – which supported speaker recognition for a variety of research tasks – and recordings in this corpus were used in the 2008 NIST Speaker Recognition Evaluation.

The data in this release was collected by LDC at its Human Subjects Data Collection Laboratories in Philadelphia and by the International Computer Science Institute (ICSI) at the University of California, Berkeley, as a collaborative, carefully coordinated activity at both recording sites. The Mixer 4 and 5 collection contains 2,568 recordings made via the public telephone network and 2,152 sessions of multiple microphone recordings in office-room settings.

The telephone protocol connected recruited speakers through a robot operator to carry on casual conversations. In Mixer 4, 400 subjects made ten 10-minute calls; half of those subjects also visited one of the collection sites where they made two telephone calls while also being recorded on a cross-channel platform. In Mixer 5, 300 subjects each completed ten calls and six interview sessions at either LDC or ICSI; those sessions were conducted on a cross channel platform and included a telephone call in one of three vocal-effort conditions - normal, high and low. Mixer participants were nearly all native English speakers, the rest being bilingual English speakers.

This release includes metadata about the calls and speakers, along with time-aligned entries for many of the component portions of the recording sessions.

Mixer 4 and 5 Speech is distributed via hard drive.

2020 Subscription Members will automatically receive copies of this corpus. 2020 Standard Members may request a copy as part of their 16 free membership corpora. This corpus is a members-only release and is not available for non-member licensing. Contact ldc@ldc.upenn.edu for information about membership.
*

Monday, February 17, 2020

LDC 2020 February Newsletter

Only two weeks left to enjoy 2020 membership discounts
LREC Workshop on Citizen Linguistics - Deadline Extended 

New Publications:
__________________________________________________________________________ 
Only two weeks left to enjoy 2020 membership discounts 

There is still time to save on 2020 membership fees. Through March 2, all organizations receive a discount on the 2020 membership fee (up to 10%) when they choose to join or renew. For more information on membership benefits, visit Join LDC.

LREC Workshop on Citizen Linguistics - Deadline Extended

LDC Researchers and their colleagues are organizing a workshop on Citizen Linguistics and Language Resource Development at LREC 2020 (Language Resource and Evaluation Conference) to take place on May 16, 2020. The workshop includes an open call for papers in language-related citizen science, a tutorial on using the new LanguageARC.org citizen linguistics portal and a special session on best papers using LanguageARC. Call for Papers deadline extended until February 24, 2020. _______________________________________________________________________

New publications:
 

(1) TAC KBP English Event Argument - Training and Evaluation Data 2014-2015 was developed by LDC and contains training and evaluation data produced in support of the 2014 TAC KBP English Event Argument Extraction Pilot and Evaluation tasks and the 2015 English Event Argument Extraction and Linking Training and Evaluation tasks

The Event Argument Extraction and Linking task required systems to extract event arguments (entities or attributes playing a role in an event) from unstructured text, indicate the role they play in an event, and link the arguments appearing in the same event to each other. Since the extracted information must be suitable as input to a knowledge base, systems constructed tuples indicating the event type, the role played by the entity in the event, and the most canonical mention of the entity from the source document. The event types and roles were drawn from an externally-specified ontology of 31 event types, which included financial transactions, communication events, and attacks. 

This corpus includes source documents, manual runs, assessments, and event hoppers, a form of identity coreference for events (2015 only).  Source data is English newswire and discussion forum text collected by LDC. 

TAC KBP English Event Argument - Training and Evaluation Data 2014-2015 is distributed via web download. 

2020 Subscription Members will automatically receive copies of this corpus. 2020 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.

(2) Chinese CogBank is a database of cognitive properties of Chinese words intended for use in metaphor understanding and generation. It consists of 232,497 "word-property" pairs, which are comprised of 83,104 words and 100,195 properties. Each "word-property" type also has an associated frequency which can stand as a functional measure of the importance of a property.

The data was collected via the Chinese search engine Baidu.com. The original collection consisted of 1,258,430 types (5,637,500 tokens) of "word-adjective" pairs that were reduced in Chinese CogBank to 232,497 "word-property" pairs after a series of manual checks.
Chinese CogBank is distributed via web download.

2020 Subscription Members will automatically receive copies of this corpus. 2020 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.



(3) Machine Reading Phase 1 IC Training Data was developed by LDC for use in the DARPA (Defense Advanced Research Projects Agency) Machine Reading program. It contains 248 English source documents and 116 standoff annotation files, annotated with instances of explicit relations and their arguments, as well as some non-explicit relations.
  
The Machine Reading program aimed to develop automated reading systems to bridge the gap between knowledge contained in natural language texts and knowledge accessible to formal reasoning systems. The reading systems designed by program participants were required to extract and reason about facts from text in multiple domains.
  
The data in this release constitutes the training data for the IC (Core Domain) task, which tested the core domain by extracting information about Entities (people, organizations, geopolitical entities) and their involvement in four types of Relations (Attack Relations, Biographical Relations, Affiliation Relations and Family Relations), as described in newswire text. This information was then aligned with an IC Use Cases ontology that would allow automated reasoning about the extracted Entities and Relations. 

Machine Reading Phase 1 IC Training Data is distributed via web download.

2020 Subscription Members will automatically receive copies of this corpus. 2020 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.

*

(4) IARPA Babel Dholuo Language Pack IARPA-babel403b-v1.0b was developed by Appen for the IARPA (Intelligence Advanced Research Projects Activity) Babel program. It contains approximately 204 hours of Dholuo conversational and scripted telephone speech collected in 2014 and 2015 along with corresponding transcripts. 

The Dholuo speech in this release represents the South Nyanza and Trans-Yala dialect regions of Kenya. The gender distribution among speakers is approximately equal; speakers' ages range from 16 years to 65 years. Calls were made using different telephones (e.g., mobile, landline) from a variety of environments including the street, a home or office, a public place, and inside a vehicle. 

IARPA Babel Dholuo Language Pack IARPA-babel403b-v1.0b is distributed via web download. 

2020 Subscription Members will receive copies of this corpus provided they have submitted a completed copy of the special license agreement. 2020 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.
*