Showing posts with label committed belief annotation. Show all posts
Showing posts with label committed belief annotation. Show all posts

Friday, November 15, 2019

LDC 2019 November Newsletter

Join LDC for Membership Year 2020
Spring 2020 Data Scholarship Program
_________________________________________________________________________ 

Join LDC for Membership Year 2020 

Membership Year 2020 (MY2020) is open and discounts are available for those who keep their membership current and join early in the year. Now through March 2, 2020, current MY2019 members who renew their LDC membership before March 2 will receive a 10% discount off the membership fee. New or returning organizations will receive a 5% discount through March 2.

In addition to receiving new publications, current LDC members also enjoy the benefit of licensing older data at reduced costs from our Catalog of over 800 holdings. Current-year for-profit members may use most data for commercial applications.

Plans for MY2020 publications are in progress. Among the expected releases are: 

Abstract Meaning Representation (AMR) Annotation Release 3.0: semantic treebank of over 59,000 English natural language sentences from broadcast conversations, newswire, weblogs and web discussion forums; updates the second version (LDC2017T10) with new annotations 
TAC KBP: English sentiment slot filling, surprise slot filling, nugget detection and coreference, and event argument data in all languages (English, Chinese and Spanish) 
DEFT Chinese ERE: Chinese discussion forum data annotated for entities, relations and events 
LibriVox Spanish: 73 hours of Spanish audiobook read speech and transcripts
IARPA Babel Language Packs (telephone speech and transcripts): languages include Dhuluo, Javanese and Mongolian 
HAVIC Med Training data: web video, metadata, and annotations for developing multimedia systems 
RATS Speaker Identification: conversational telephone speech in Levantine Arabic, Pashto, Urdu, Farsi and Dari on degraded audio signals with annotation of speech segments for speaker identification 
BOLT: discussion forums, SMS/chat, conversational telephone speech, word-aligned, tagged and co-reference data in all languages (Chinese, Egyptian Arabic, and English) 

It’s also not too late to join for MY2018 (through December 31, 2019) and MY2019 (through December 31, 2020). Data sets from those years include Concretely Annotated New York Times and English Gigaword, DIRHA English WSJ Audio, BOLT English Treebank – Discussion Forum, First DIHARD Challenge Development and Evaluation releases, Penn Discourse Treebank Version 3.0, and 2016 NIST Speaker Recognition Evaluation Test Set. 

For full descriptions of all LDC data sets, browse our Catalog.

Visit Join LDC for details on membership, user accounts and payment. 

Spring 2020 Data Scholarship Program 

Applications are now being accepted through January 15, 2020 for the Spring 2020 LDC Data Scholarship program which provides university students with no-cost access to LDC data. Consult the LDC Data Scholarship page for more information about program rules and submission requirements.
_________________________________________________________________________  

New publications: 

(1) DEFT English Committed Belief Annotation was developed by LDC and consists of approximately 950,000 words of English discussion forum text annotated for "committed belief," which marks the level of commitment displayed by the author to the truth of the propositions expressed in the text.

DARPA's Deep Exploration and Filtering of Text (DEFT) program aimed to address remaining capability gaps in state-of-the-art natural language processing technologies related to inference, causal relationships and anomaly detection. LDC supported the DEFT program by collecting, creating and annotating a variety of data sources.

DEFT English Committed Belief Annotation is distributed via web download.

2019 Subscription Members will automatically receive copies of this corpus. 2019 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.

*

(2) CALLFRIEND American English-Non-Southern Dialect Second Edition was developed by LDC and consists of approximately 26 hours of unscripted telephone conversations between native speakers of non-Southern dialects of American English. This second edition updates the audio files to wav format, simplifies the directory structure and adds documentation and metadata. The first edition is available as CALLFRIEND American English-Non-Southern Dialect (LDC96S46).

All data was collected before July 1997. Participants could speak with a person of their choice on any topic; most called family members and friends. All calls originated in North America. The recorded conversations last up to 30 minutes. 

CALLFRIEND American English-Non-Southern Dialect Second Edition is distributed via web download.

2019 Subscription Members will automatically receive copies of this corpus. 2019 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee. 

*

(3) TAC KBP Cold Start - Comprehensive Evaluation Data 2012-2017 was developed by LDC and contains Chinese, English and Spanish data produced in support of the TAC KBP Cold Start evaluation track conducted from 2012 to 2017. This corpus includes source documents, queries, assessments, manual runs and final assessments. 

In the Cold Start track, systems were evaluated on their ability to construct a new knowledge base (KB) from information provided in a text collection in combination with technologies developed in other TAC KBP tracks -- slot filling, information extraction, question answering and entity discovery and linking. Cold Start systems were required to find all entities in the text, and the KB must have ideally included every person, organization, and geo-political entity as well as all the targeted relations between them. To facilitate the evaluation of those KBs, LDC annotators created sets of queries, human-generated responses to the queries, and assessments of both human and system responses. 

The source data in this release is comprised of English and Spanish newswire and web text collected by LDC for the 2012, 2014 and 2015 evaluations and the 2016 pilot collection. The source collections for the 2016 and 2017 evaluations, which include Chinese data, are available in TAC KBP Evaluation Source Corpora 2016-2017 (LDC2019T12). The archived 2013 Cold Start source data collection is available from NIST upon request.

TAC KBP Cold Start - Comprehensive Evaluation Data 2012-2017 is distributed via web download. 

2019 Subscription Members will automatically receive copies of this corpus. 2019 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee. 

*

(4) IARPA Babel Amharic Language Pack IARPA-babel307b-v1.0b was developed by Appen for the IARPA (Intelligence Advanced Research Projects Activity) Babel program. It contains approximately 204 hours of Amharic conversational and scripted telephone speech collected in 2014 along with corresponding transcripts.

The Amharic speech in this release represents the Addis Ababa, Shewa, and Gondar dialect regions of Ethiopia. The gender distribution among speakers is approximately equal; speakers' ages range from 16 years to 60 years. Calls were made using different telephones (e.g., mobile, landline) from a variety of environments including the street, a home or office, a public place, and inside a vehicle.

IARPA Babel Amharic Language Pack IARPA-babel307b-v1.0b is distributed via web download.

2019 Subscription Members will receive copies of this corpus provided they have submitted a completed copy of the special license agreement. 2019 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.

Monday, June 17, 2019

LDC 2019 June Newsletter

In this newsletter:

New Publications:
USC-SFI MALACH Interviews and Transcripts English – Speech Recognition Edition
First DIHARD Challenge Development - Eight Sources
_____________________________________________________________________

New publications:  

(1) DEFT Spanish Committed Belief Annotation was developed by LDC and consists of approximately 67,000 tokens of Spanish discussion forum text annotated for "committed belief," which marks the level of commitment displayed by the author to the truth of the propositions expressed in the text.

DARPA's Deep Exploration and Filtering of Text (DEFT) program aimed to address remaining capability gaps in state-of-the-art natural language processing technologies related to inference, causal relationships and anomaly detection. LDC supported the DEFT program by collecting, creating and annotating a variety of data sources. 

DEFT Spanish Committed Belief Annotation is distributed via web download.

2019 Subscription Members will automatically receive copies of this corpus. 2019 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a cost. 

*

(2) USC-SFI MALACH Interviews and Transcripts English – Speech Recognition Edition was developed by IBM as part of the MALACH (Multilingual Access to Large Spoken ArCHives) Project and contains approximately 168 hours of interviews from 682 Holocaust witnesses along with transcripts, a lexicon and other documentation. This release augments USC-SFI MALACH Interviews and Transcripts English (LDC2012S05) by modifying and updating a subset of the original corpus for use with speech recognition systems, such as the Kaldi toolkit.

Specifically, the audio data has been converted from unsegmented mpeg files to a segmented flac compressed format. The speaker-turn, time-stamped transcripts have been updated to an utterance-by-utterance format. A lexicon mapping words to phonemes is provided, and the data is divided into development and training sets.  

The goal of the MALACH project was to develop methods for improved access to large multinational spoken archives in order to advance the state of the art of automatic speech recognition and information retrieval. The characteristics of the USC-SFI collection -- unconstrained, natural speech filled with disfluencies, heavy accents, age-related coarticulations, un-cued speaker and language switching and emotional speech -- were considered well-suited for that task. 

USC-SFI MALACH Interviews and Transcripts English – Speech Recognition Edition is distributed via web download. 

2019 Subscription Members will automatically receive copies of this corpus provided they have submitted a completed copy of the special license agreement. 2019 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data at no cost.

*

(3) First DIHARD Challenge Development - Eight Sources was developed by LDC and contains approximately 17 hours of English and Chinese speech data along with corresponding annotations used in support of the First DIHARD Challenge. This release, when combined with First DIHARD Challenge Development - SEEDLingS (LDC2019S10), contains the development set audio data and annotation (diarization, segmentation) as well as the official scoring tool.


The First DIHARD Challenge was an attempt to reinvigorate work on diarization through a shared task focusing on "hard" diarization; that is, speech diarization for challenging corpora where there was an expectation that existing state-of-the-art systems would fare poorly. As such, it included speech from a wide sampling of domains representing diversity in number of speakers, speaker demographics, interaction style, recording quality, and environmental conditions as follows (all sources are in English unless otherwise indicated):
  • Autism Diagnostic Observation Schedule (ADOS) interviews
  • DCIEM/HCRC map task (LDC96S38)
  • Audiobook recordings from LibriVox
  • Meeting speech from 2004 Spring NIST Rich Transcription (RT-04S) Development (LDC2007S11) and Evaluation (LDC2007S12) releases.
  • 2001 U.S. Supreme Court oral arguments
  • Sociolinguistic interviews from SLX Corpus of Classic Sociolinguistic Interviews (LDC2003T15)
  • Chinese video collected by LDC as part of the Video Annotation for Speech Technologies (VAST) project
  • YouthPoint radio interviews

First DIHARD Challenge Development - Eight Sources is distributed via web download. 

2019 Subscription Members will automatically receive copies of this corpus.  2019 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a cost.

*

(4) First DIHARD Challenge Development - SEEDLingS was developed by Duke University and LDC and contains approximately two hours of English child language recordings along with corresponding annotations used in support of the First DIHARD Challenge. This release, when combined with First DIHARD Challenge Development - Eight Sources (LDC2019S09), contains the development set audio data and annotation (diarization, segmentation) as well as the official scoring tool.

The source data was drawn from the SEEDLingS (The Study of Environmental Effects on Developing Linguistic Skills) corpus, designed to investigate how infants' early linguistic and environmental input plays a role in their learning. Recordings for SEEDLingS were generated in the home environment of 44 infants from 6-18 months of age in the Rochester, New York area. A subset of that data was annotated by LDC for use in the First DIHARD Challenge.

The First DIHARD Challenge was an attempt to reinvigorate work on diarization through a shared task focusing on "hard" diarization; that is, speech diarization for challenging corpora where there was an expectation that existing state-of-the-art systems would fare poorly. As such, it included speech from a wide sampling of domains representing diversity in number of speakers, speaker demographics, interaction style, recording quality, and environmental conditions.

First DIHARD Challenge Development – SEEDLingS is distributed via web download.

2019 Subscription Members will receive copies of this corpus provided they have submitted a completed copy of the special license agreement. 2019 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a cost.

*

Thursday, February 14, 2019

LDC 2019 February Newsletter

Only two weeks left to enjoy 2019 membership discounts

Spring 2019 LDC Data Scholarship recipients

LDC’s new language game

New publications:

___________________________________________________________

Only two weeks left to enjoy 2019 membership discounts
There is still time to save on 2019 membership fees. Through March 1, all organizations receive a discount on the 2019 membership fee (up to 10%) when they choose to join or renew. For more information on membership benefits, visit Join LDC

Spring 2019 LDC Data Scholarship recipients
Congratulations to the recipients of LDC's Spring 2019 Data Scholarships:

Colin Annand: University of Cincinnati (USA); PhD. Psychology. Colin is awarded a copy of Switchboard-1 Release 2 for his research involving the relationship between speech patterns and conversation content.

Si Chen: Huazhong University of Science and Technology (China); B.S. Communication Engineering. Si is awarded a copy of ACE 2005 Multilingual Training Corpus for his work on event extraction. 

Noor-e-Hira: Fatima Jinnah Women University (Pakistan); MSc. Computer Sciences. Noor is awarded a copy of NIST 2008 Open Machine Translation (OpenMT) Evaluation for her research in machine translation.

Matthew Roddy: Trinity College Dublin (Ireland); Ph.D. Electrical Engineering. Matthew is awarded copies of 2000 HUB5 English Evaluation Speech and Transcripts for his work in spoken dialogue systems.

Ammara Zafar: Fatima Jinnah Women University (Pakistan); MSc Computer Sciences. Ammara awarded a copy of NIST 2009 Open Machine Translation (OpenMT) Evaluation for her research in machine translation.

For information about the program, visit the Data Scholarship page.

LDC’s new language game
LDC’s new language game, NameThatLanguage, tests your skill at recognizing the language spoken in short audio clips. The game includes thousands of clips to prevent memorization and offers a real challenge that increases as you progress. In addition to being fun, the game provides useful data on language confusability and linguistic diversity. Game results will be shared freely for research. New clips and more languages continue to be added providing ongoing challenges and new research data. Help support language research by playing! https://namethatlanguage.org

New publications:

(1) DEFT Chinese Committed Belief Annotation was developed by LDC and consists of approximately 83,000 tokens of Chinese discussion forum text annotated for "committed belief," which marks the level of commitment displayed by the author to the truth of the propositions expressed in the text.

DARPA's Deep Exploration and Filtering of Text (DEFT) program aimed to address remaining capability gaps in state-of-the-art natural language processing technologies related to inference, causal relationships and anomaly detection. LDC supported the DEFT program by collecting, creating and annotating a variety of data sources.

DEFT Chinese Committed Belief Annotation is distributed via web download.

2019 Subscription Members will automatically receive copies of this corpus. 2019 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.  

*

(2) IARPA Babel Lithuanian Language Pack IARPA-babel304b-v1.0b was developed by Appen for the IARPA (Intelligence Advanced Research Projects Activity) Babel program. It contains approximately 210 hours of Lithuanian conversational and scripted telephone speech collected in 2013 and 2014 along with corresponding transcripts.

The Lithuanian speech in this release represents that spoken in the Aukštaitian and Samogitian dialect regions of Lithuania. The gender distribution among speakers is approximately equal; speakers' ages range from 16 years to 71 years. Calls were made using different telephones (e.g., mobile, landline) from a variety of environments including the street, a home or office, a public place, and inside a vehicle.

IARPA Babel Lithuanian Language Pack IARPA-babel304b-v1.0b is distributed via web download.

2019 Subscription Members will receive copies of this corpus provided they have submitted a completed copy of the special license agreement. 2019 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.

*

(3) Multi-Language Conversational Telephone Speech 2011 -- Arabic Group was developed by LDC and is comprised of approximately 117 hours of telephone speech in distinct dialects of colloquial Arabic: Iraqi, Levantine and Maghrebi.

The data were collected primarily to support research and technology evaluation in automatic language identification, and portions of these telephone calls were used in the NIST 2011 Language Recognition Evaluation (LRE). LRE 2011 focused on language pair discrimination for 24 languages/dialects, some of which could be considered mutually intelligible or closely related.

Multi-Language Conversational Telephone Speech 2011 -- Arabic Group is distributed via web download.

2019 Subscription Members will automatically receive copies of this corpus. 2019 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee. 

*

(4) Multilingual ATlS was developed by Google Inc. and consists of 5,871 utterances from ATIS2 (LDC93S5), ATIS3 Training Data (LDC94S19), and ATIS3 Test Data (LDC95S26) annotated and translated into Hindi and Turkish. 

The ATIS (Air Travel Information Services) collection was developed to support the research and development of speech understanding systems. Participants were presented with various hypothetical travel planning scenarios and asked to solve them by interacting with partially or completely automated ATIS systems. The resulting utterances were recorded and transcribed. Data was collected in the early 1990s at five US sites: Raytheon BBN, Carnegie Mellon University, MIT Laboratory for Computer Science, National Institute for Standards and Technology and SRI International.

The original English utterances were manually translated into Hindi and Turkish. This release also includes the original English utterance and the machine translation back into English of the manual target language utterance translation. Each utterance is annotated with named entities via table lookup; markers include city, airline, airport names, and dates.

Multilingual ATIS is distributed via web download.

2019 Subscription Members will automatically receive copies of this corpus. 2019 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data at no cost.