Thursday, October 15, 2020

LDC 2020 October Newsletter

Fall 2020 Data Scholarship Recipients 

Membership Year 2021 Publication Preview 

LDC data and commercial technology development  
 
New Publications: 
Global TIMIT Learner Treebank English 
Corpus of Law, Academic, and News 
IARPA Babel Mongolian Language Pack IARPA-babel401b-v2.0b 

__________________________________________________________________

 
Fall 2020 data scholarship recipients 
Congratulations to the recipients of LDC's Fall 2020 data scholarships: 
 
Nicole Dodd: University of California, Davis (USA); MA, Linguistics. Nicole is awarded a copy of Arabic Treebank Part 3 v. 3.2 LDC2010T08 for her research in relative clause processing in Standard Arabic. 
 
Satwik Dutta: University of Texas at Dallas (USA); PhD, Electrical Engineering. Satwik is awarded copies of The CMU Kids Corpus LDC97263 and CLSU: Kids’ Speech Version 1.1. LDC2007S18 for his work in speech activity detection.  
 
Pedram Hosseini: George Washington University (USA); PhD., Computer Science. Pedram is awarded copies of Penn Discourse Treebank Version 3.0 LDC2019T05 and The New York Times Annotated Corpus LDC2008T19 for his research in automatic detection of causal relations in text.  
 
Mariano Maisonnave: Universidad Nacional del Sur (Argentina); PhD, Computer Science. Mariano is awarded a copy of ACE 2005 Multilingual Training Corpus LDC2006T06 for his work in event extraction.  
 
Mark Sullivan: California State University, Los Angeles (USA); Masters, Applied and Advanced Studies in Education. Mark is awarded a copy of ETS Corpus of Non-Native Written English LDC2014T06 for his research in sentence boundary problems of Chinese L1 speakers in English compositions.  
 
For information about the program, visit the Data Scholarships page. 
 
Membership Year 2021 publication preview 
The 2021 Membership Year is just around the corner and plans for next year’s publications are in progress. Among the expected releases are: 

  • Global TIMIT Mandarin Chinese: 6,000 linguistically rich utterances featuring time-aligned lexical and phonetic transcription 
  • Columbia Games Corpus: 12 spontaneous task-oriented dyadic conversations elicited from native Standard American English speakers playing computer games, transcribed and annotated for discourse/pragmatic phenomena 
  • My Science Tutor Children’s Conversational Speech: 400+ hours of speech from 1,371 US third, fourth, and fifth grade students participating in sessions with a virtual science tutor, transcripts included 
  • The SSNCE Database of Tamil Dysarthric Speech: Tamil speech from 20 dysarthric speakers aged 12-40 years and a control group (10 speakers) with time-aligned phonetic transcripts 
  • Icelandic Parliamentary Speech: 6,493 Icelandic Parliament recordings from 2005-2016 with 196 speakers, aligned and segmented and divided into training, development, and evaluation sets for ASR development 
  • LORELEI: representative and incident language packs containing monolingual text, bi-text, translations, annotations, supplemental resources, and related tools (Akan, Kinyarwanda, and Wolof) 
  • BOLT: co-reference, treebank, propbank, and translation resources for discussion forum, SMS/Chat, and conversational telephone speech data in all languages (Chinese, Egyptian Arabic, and English) 
  • TAC KBP: training and evaluation data for English surprise slot filling (2010) and English sentiment slot filling (2013-2014) tasks  

Check your inbox in the coming weeks for more information about membership renewal.  
 
LDC data and commercial technology development 
For-profit organizations are reminded that an LDC membership is a pre-requisite for obtaining a commercial license to almost all LDC databases. Non-member organizations, including non-member for-profit organizations, cannot use LDC data to develop or test products for commercialization, nor can they use LDC data in any commercial product or for any commercial purpose. LDC data users should consult corpus-specific license agreements for limitations on the use of certain corpora. Visit the Licensing page for further information. 
 
New publications: 

(1) Global TIMIT Learner Treebank English was developed by LDC and LAIX Inc. and consists of approximately 24 hours of L1 and L2 English read speech and transcripts. It is comprised of two separate data sets of 50 speakers reading 120 sentences from Treebank-3 (LDC99T42). Among the 120 sentences, 20 sentences were read by all speakers, 40 sentences were read by 10 speakers, and 60 sentences were read by one speaker, for a total of 3220 sentence types.  
 
L1 English Treebank was recorded at the University of Pennsylvania, USA; participants were 25 female and 25 male native American English speakers. L2 English Treebank was recorded at LAIX Inc., Shanghai, China. L2 speakers (25 female, 25 male) were Chinese learners of English considered fluent and who had passed specified standards on English assessment tests.       
 
The Global TIMIT project aimed to create a series of corpora in a variety of languages with a similar set of key features as in the original TIMIT Acoustic-Phonetic Continuous Speech Corpus (LDC93S1) which was designed for acoustic-phonetic studies and for the development and evaluation of automatic speech recognition systems. Specifically, those features include: 

  • A large number of fluently-read sentences, containing a representative sample of phonetic, lexical, syntactic, semantic, and pragmatic patterns
  • A relatively large number of speakers 
  • Time-aligned lexical and phonetic transcription of all utterances 
  • Some sentences read by all speakers, others read by a few speakers, and others read by just one speaker 

Global TIMIT Learner Treebank English is distributed via web download.   
 
2020 Subscription Members will automatically receive copies of this corpus. 2020 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee. 


* 


(2) Corpus of Law, Academic and News consists of 400 Persian documents divided into three genres: legal, academic, and news. The legal section contains texts from official publications, including the civil penal code, the criminal penal code, and the constitution of the Islamic Republic of Iran. The academic sub-corpus is comprised of published academic abstracts in various disciplinary areas, such as Art and Humanities, Social Sciences, and Natural Sciences. The news sub-corpus was extracted from an archive of ten Iranian news outlets spanning the period 2010- 2020. 
 
Each document contains metadata in the file's header with information such as specific text type, dates, and source, and also contains annotations marking title and body paragraphs.  
 
Corpus of Law, Academic and News is distributed via web download. 
 
2020 Subscription Members will automatically receive copies of this corpus provided they have submitted a completed copy of the special license agreement. 2020 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee 


* 


(3) IARPA Babel Mongolian Language Pack IARPA-babel401b-v2.0b was developed by Appen for the IARPA (Intelligence Advanced Research Projects Activity) Babel program. It contains approximately 204 hours of Halh Mongolian conversational and scripted telephone speech collected in 2014 along with corresponding transcripts. The gender distribution among speakers is approximately equal; speakers' ages range from 16 years to 61 years. Calls were made using different telephones (e.g., mobile, landline) from a variety of environments including the street, a home or office, a public place, and inside a vehicle. 

  

The Babel program focused on underserved languages and sought to develop speech recognition technology that could be rapidly applied to any human language to support keyword search performance over large amounts of recorded speech.   

  

This is the last release in the IARPA Babel series which consists of 25 language packs in total. 

  

IARPA Babel Mongolian Language Pack IARPA-babel-401b-v2.0b is distributed via web download.  


2020 Subscription Members will automatically receive copies of this corpus provided they have submitted a completed copy of the special license agreement. 2020 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee 

 
 

Tuesday, September 15, 2020

LDC 2020 September Newsletter

New Publications:
BOLT English PropBank and Sense – Discussion Forum, SMS/Chat and Conversational Telephone Speech
LORELEI Tigrinya Incident Language Pack
Chinese Lexical Resources for Gender, Number, Animacy

New publications:
(1) BOLT English PropBank and Sense – Discussion Forum, SMS/Chat and Conversational Telephone Speech was developed by the University of Colorado, Boulder – CLEAR (Computational Language and Education Research) and consists of propbank and verb sense disambiguation annotation on English discussion forum (DF), SMS/Chat, and conversational telephone speech data. Annotation was applied to each predicate verb tree in LDC’s BOLT phrase structure treebanks. PropBank provides a layer of semantic annotation over treebank and was performed on all three genres. DF and SMS/Chat data were also annotated for verb sense disambiguation using Verbnet 3.2 classes

The DARPA BOLT (Broad Operational Language Translation) program developed machine translation and information retrieval for less formal genres, focusing particularly on user-generated content. LDC supported the BOLT program by collecting informal data sources -- discussion forums, text messaging, and chat -- in Chinese, Egyptian Arabic, and English. The collected data was translated and annotated for various tasks including word alignment, treebanking, propbanking, and co-reference.

BOLT English PropBank and Sense – Discussion Forum, SMS/Chat and Conversational Telephone Speech is distributed via web download.

2020 Subscription Members will automatically receive copies of this corpus. 2020 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may this data for a fee.

*

(2) LORELEI Tigrinya Incident Language Pack was developed by LDC and is comprised of approximately 4.5 million words of Tigrinya monolingual text, 25,000 words of English monolingual text, 235,000 words of parallel and comparable Tigrinya-English text, and 50,000 words of data annotated for Entity Discovery and Linking and for Situation Frames. It contains all of the text data, annotations, supplemental resources, and related software tools for the Tigrinya language that were used in the DARPA LORELEI / LoReHLT 2017 Evaluation.

The evaluation protocol was based on a scenario in which an unforeseen event triggered a need for humanitarian and logistical support in a region where the incident language had received little or no attention in NLP research. Evaluation participants provided NLP solutions, including information extraction and machine translation, with limited resources and limited development time. 

Data was collected from news, social network, weblog, newsgroup, discussion forum, and reference material. Entity Detection and Linking and Situation Frame annotations identified “entities,” “needs” (such as a need for food), and “issues” (such as civil unrest) to be detected by systems for scoring purposes. Situation frame analysis was designed to extract basic information that would be useful for planning a disaster response effort.

The knowledge base for the entity linking annotation in this corpus is available separately as LORELEI Entity Detection and Linking Knowledge Base (LDC2020T10).

LORELEI Tigrinya Incident Language Pack is distributed via web download.

2020 Subscription Members will automatically receive copies of this corpus. 2020 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a  fee.

*

(3) Chinese Lexical Resources for Gender, Number, Animacy was developed by LDC and consists of gender, number, and animacy lexicons produced in support of the DARPA DEFT program. Gender, number, and animacy are lexical indicators useful for named entity tagging, including the detection of person mentions in text.

This corpus was created by extracting information from newswire texts in Chinese Gigaword Fifth Edition (LDC2011T13) in the following steps: (1) segmenting source documents into sentences; (2) converting any traditional Chinese script to simplified Chinese; (3) tagging all sentences for parts-of-speech; (4) developing queries to detect patterns; and (5) building lexicons based on frequency counts and entity types.

The resulting resources include dictionaries of Chinese animate nominals and names; Chinese nominals and name with gender and number predicted; and other dictionaries of Chinese nominals, names, verbs, and pronouns. Each dictionary contains frequency information as well as the features in question.

DARPA's Deep Exploration and Filtering of Text (DEFT) program aimed to address remaining capability gaps in state-of-the-art natural language processing technologies related to inference, causal relationships and anomaly detection. LDC supported the DEFT program by collecting, creating and annotating a variety of data sources.

Chinese Lexical Resources for Gender, Number, Animacy is distributed via web download.

2020 Subscription Members will automatically receive copies of this corpus. 2020 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.

Tuesday, August 18, 2020

LDC 2020 August Newsletter

LDC adds DOI Identifier to its Language Resources
Fall 2020 LDC Data Scholarship Program

New Publications:
LORELEI Vietnamese Representative Language Pack
DEFT Chinese Light and Rich ERE Annotation
CALLFRIEND American English – Southern Dialect Second Edition


 LDC adds DOI Identifier to its Language Resources
As of July 2020, LDC’s language resources include a Digital Object Identifier (DOI), an internationally recognized identification standard for online digital material. DOIs are alpha numeric strings that correspond to URLs and metadata for specified resources. They are expressed as links that resolve to the object’s online location. For example, the DOI for Penn Parsed Corpora of Historical English LDC2020T16 is https://doi.org/10.35111/4hzx-5483, which leads users to the LDC catalog entry for this data set. To facilitate its assignment and administration of DOIs, LDC has joined DataCite, a global DOI provider for research data. (DOIs for resources released before July 2020 will be assigned through a process expected to be completed shortly.) LDC data sets now have four persistent identifiers: a unique LDC number, ISBN, ISLRN, and DOI. Adding DOIs is consistent with our aim to follow best practices for archiving and curating digital resources, evidenced by the CoreTrustSeal certification which recognizes the LDC Catalog as a trustworthy data repository.

Fall 2020 LDC Data Scholarship Program
Student applications for the Fall 2020 LDC Data Scholarship program are being accepted now through September 15, 2020. This scholarship program provides eligible students with no-cost access to LDC data. Students must complete an application consisting of a data use proposal and letter of support from their advisor.

For application requirements and program rules, visit the LDC Data Scholarship page.

 


 New publications:
(1) LORELEI Vietnamese Representative Language Pack consists of Vietnamese monolingual text, Vietnamese-English parallel text, annotations, supplemental resources, and related software tools developed by LDC for the DARPA LORELEI program.

The LORELEI (Low Resource Languages for Emergent Incidents) program was concerned with building human language technology for low resource languages in the context of emergent situations like natural disasters or disease outbreaks. Linguistic resources for LORELEI include Representative Language Packs and Incident Language Packs for over two dozen low resource languages, comprising data, annotations, basic natural language processing tools, lexicons, and grammatical resources. Representative languages were selected to provide broad typological coverage, while incident languages were selected to evaluate system performance on a language whose identity was disclosed at the start of the evaluation.

Data was collected in the following genres: discussion forum, news, reference, social network, and weblogs. Data volumes are as follows:

  • Over 172 million words of Vietnamese monolingual text, approximately 325,000 words of which were translated into English
  • 106,000 Vietnamese words translated from English data
  • 1.9 million words of found parallel text
Approximately 75,000 words were annotated for named entities and up to 25,000 words contain additional annotation, including situation frames (identifying entities, needs, and issues) and entity linking and detection.

LORELEI Vietnamese Representative Language Pack is distributed via web download.

2020 Subscription Members will automatically receive copies of this corpus. 2020 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.
                                                                                                                                                                       

(2) DEFT Chinese Light and Rich ERE Annotation contains Chinese discussion forum web text annotated for entities, relations, and events (ERE) using the ERE Light and ERE Rich annotations schemas developed by LDC. Light ERE annotation labels entity mentions for the target set of ERE types between and among those entities, including coreference. Rich ERE annotation expands types and tagging for ERE annotation tasks and replaces event coreference with event hopper annotation. All files in this release (157) were annotated following Light ERE guidelines; a subset (149) were also labeled with Rich ERE annotation. 

DARPA’s Deep Exploration and Filtering of Text (DEFT) program aimed to address remaining capability gaps in state-of-the-art natural language processing technologies related to inference, causal relationships, and anomaly detection. LDC supported the DEFT program by collecting, creating, and annotating a variety of data sources.

DEFT Chinese Light and Rich ERE Annotation is distributed via web download.

2020 Subscription Members will automatically receive copies of this corpus. 2020 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.
                                                                                                                                                                              

(3) CALLFRIEND American English – Southern Dialect Second Edition was developed by LDC and consists of approximately 26 hours of unscripted telephone conversations between native speakers of Southern dialects of American English. This second edition updates the audio files to wav format, simplifies the directory structure, and adds documentation and metadata. The first edition is available as CALLFRIEND American English-Southern Dialect (LDC96S47).

The CALLFRIEND collection was conducted by LDC in support of language identification technology development. All data in this release was collected before July 1997. Participants could speak with a person of their choice on any topic; most called family members and friends. All calls originated in North America. The recorded conversations last up to 30 minutes.

CALLFRIEND American English – Southern Dialect Second Edition is distributed via web download.

2020 Subscription Members will automatically receive copies of this corpus. 2020 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.

Wednesday, July 15, 2020

LDC 2020 July Newsletter

Penn Parsed Corpora of Historical English Now Available From LDC
Fall 2020 LDC Data Scholarship Program 


New Publications:
Speech Sentiment Annotations
Penn Parsed Corpora of Historical English
IARPA Babel Javanese Language Pack IARPA-babel402b-v1.0b
BOLT Chinese-English Word Alignment and Tagging -- Conversational Telephone Speech Training
____________________________________________________________

Penn Parsed Corpora of Historical English Now Available From LDC

LDC is pleased to announce that the Penn Parsed Corpora of Historical English (LDC2020T16) – an important community resource for 20 years – is now available for licensing in the LDC Catalog. Developed by University of Pennsylvania researchers in the Linguistics Department under the direction of Professor Anthony Kroch, this data set consists of syntactic annotation of English prose texts from the earliest Middle English documents (1100 CE) up to the period of the First World War (1914 CE) represented in three corpora:

  • The Penn-Helsinki Corpus of Middle English, second edition
  • The Penn-Helsinki Parsed Corpus of Early Modern English
  • The Penn Parsed Corpus of Modern British English, second edition
This release also includes annotation guidelines and philological information for each corpus as well as the CorpusSearch 2 program which allows users to search the data for words, word sequences and syntactic structure.

In addition to being of value to students and scholars of the history of English, this data set is useful to computational linguists for domain adaptation. More information about this project is available from the Penn Parsed Corpora of Historical English homepage.

Current licensees should contact LDC’s membership office with any questions regarding access to this data set.

Fall 2020 LDC Data Scholarship Program

Student applications for the Fall 2020 LDC Data Scholarship program are being accepted now through September 15, 2020. This scholarship program provides eligible students with no-cost access to LDC data. Students must complete an application consisting of a data use proposal and letter of support from their advisor.

For application requirements and program rules, please visit the LDC Data Scholarship page.
____________________________________________________________

New publications:

(1) Speech Sentiment Annotations was developed by Google Inc. and consists of sentiment labels (positive, negative, neutral) for approximately 49,500 utterances covering 140 hours of audio from Switchboard-1 Release 2 (LDC97S62).

Switchboard speech files were segmented based on the start and end time of transcript turns. Annotators listened to the audio corresponding to each segment (utterance) and classified each into positive, negative or neutral categories based on the emotion and attitude of the speaker. Annotators provided a justification for positive and negative classifications using a flow chart. Further information about the methodology and annotation process is contained in the documentation accompanying this release.

Switchboard-1 Release 2 (LDC97S62) consists of 260 hours of telephone speech from 543 speakers across the United States (302 male speakers, 241 female speakers). A computer-driven telephone collection platform paired two subjects for each conversation and provided a discussion topic, ensuring that no two speakers conversed together more than once and no one speaker talked more than once on a given topic.

Speech Sentiment Annotations is distributed via web download.

2020 Subscription Members will automatically receive copies of this corpus. 2020 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.

(2) Penn Parsed Corpora of Historical English was developed at the University of Pennsylvania and consists of running texts and text samples of British English prose from the earliest Middle English documents (1100 CE) up to the period of the First World War (1914 CE). This data set contains three corpora covering traditionally recognized periods of English: 
  • The Penn-Helsinki Parsed Corpus of Middle English, second edition
  • The Penn-Helsinki Parsed Corpus of Early Modern English
  • The Penn Parsed Corpus of Modern British English, second edition
The texts are in three forms: plain text, part-of-speech tagged text, and syntactically annotated text. This release also includes annotation guidelines, philological information for each corpus and the CorpusSearch 2 program, which allows users to search the data for words, word sequences and syntactic structure.

The Penn Parsed Corpora of Historical English were designed for students and scholars of the history of English, especially the historical syntax of the language. They have also been used by computational linguists for domain adaptation. See the Penn Parsed Corpora of Historical English homepage for more information about this project.

Penn Parsed Corpora of Historical English is distributed via web download.

2020 Subscription Members will receive copies of this corpus provided they have submitted a completed copy of the special license agreement. 2020 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.
* 

(3) IARPA Babel Javanese Language Pack IARPA-babel402b-v1.0b was developed by Appen for the IARPA (Intelligence Advanced Research Projects Activity) Babel program. It contains approximately 204 hours of Javanese conversational and scripted telephone speech collected in 2014 and 2015 along with corresponding transcripts. 

The Javanese speech in this release represents the Central, Western, and Eastern Javanese dialect regions of Indonesia. The gender distribution among speakers is approximately equal; speakers' ages range from 16 years to 65 years. Calls were made using different telephones (e.g., mobile, landline) from a variety of environments including the street, a home or office, a public place, and inside a vehicle.

IARPA Babel Javanese Language Pack IARPA-babel402b-v1.0b is distributed via web download.

2020 Subscription Members will automatically receive copies of this corpus provided they have submitted a completed copy of the special license agreement. 2020 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.

* 
(4) BOLT Chinese-English Word Alignment and Tagging -- Conversational Telephone Speech Training was developed by LDC and consists of 158,651 words of Chinese and English parallel text enhanced with linguistic tags to indicate word relations. 

The source data in this release consists of transcripts of Chinese conversational telephone speech (CTS) from LDC's CALLHOME and CALLFRIEND collections (LDC96S34, LDC96T16, LDC96S55) that were translated into English by professional translation agencies and annotated for the word alignment task.

The BOLT word alignment task was built on treebank annotation. LDC automatically extracted Chinese source tokens, including empty categories/traces, from word-segmented files provided by the BOLT Chinese Treebank annotation team at Brandeis University. The word-segmented tokens were then used to automatically generate ctb (Chinese Treebank) alignment and were also tokenized for character alignment by inserting white spaces to separate characters.

BOLT Chinese-English Word Alignment and Tagging -- Conversational Telephone Speech Training is distributed via web download.

2020 Subscription Members will automatically receive copies of this corpus. 2020 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.

*