Showing posts with label Cantonese. Show all posts
Showing posts with label Cantonese. Show all posts

Thursday, May 16, 2024

LDC May 2024 Newsletter

LDC at LREC-COLING 2024

New publications:
Call My Net 1
______________________________________________________________________

LDC at LREC-COLING 2024
LDC will be exhibiting at LREC-COLING 2024 hosted by the European Language Resources Association (ELRA) and the International Committee on Computational Linguistics (ICCL) May 20-25 in Turin, Italy. Stop by our table to learn more about recent developments at the Consortium and the latest publications. 

LDC staff members will also be presenting current work on topics including Spanless Event Annotation for Corpus-Wide Complex Event Understanding, Schema Learning Corpus: Data and Annotation Focused on Complex Events, and KoFREN: Comprehensive Korean Word Frequency Norms Derived from Large Scale Free Speech Corpora. 

LDC will post conference updates via social media. We look forward to seeing you in Italy!

New publications: 

Call My Net 1 was developed by LDC and contains 364 hours of conversational telephone speech in four languages (Tagalog, Cebuano, Cantonese and Mandarin) collected in 2015 from 221 native speakers located in the Philippines and China along with metadata and speaker demographic information. Recordings and data from this collection were used to support the NIST 2016 Speaker Recognition Evaluation.

Speakers made 10 telephone calls each to people within their existing social networks, using different handsets and under a variety of noise conditions. Speakers were connected through a robot operator to carry on casual conversations on topics of their choice. All recordings were manually audited to confirm language and speaker requirements. The documentation for this release includes metadata about phone type, noise conditions and call quality. Speaker demographic information on year of birth, sex and native language is also included.

2024 members can access this corpus through their LDC accounts. Non-members may license this data for a fee.

*

Automatic Content Extraction for Portuguese was developed at INESC TEC - Instituto de Engenharia de Sistemas e Computadores, Tecnologia e Ciência and consists of automatic Brazilian Portuguese and European Portuguese translations of the English text and annotations in ACE 2005 Multilingual Training Corpus (LDC2006T06).

ACE 2005 Multilingual Training Corpus was developed by LDC to support the Automatic Contract Extraction (ACE) program, specifically, by providing training data for the 2005 technology evaluation. It contains 1,800 files of mixed genre text in Arabic, English and Chinese annotated for entities, relations and events. The objective of the ACE program was to develop automatic content extraction technology to support automatic processing of human language in text form. Text genres included newswire, broadcast news, broadcast conversation, weblog, discussion forums, and conversational telephone speech.

For this translation, the English data was partitioned into training, development and test sets. The documents were split into sentences and each event mention was assigned to its sentence. Source sentences and their annotations were translated into Brazilian Portuguese using Google Translate and into European Portuguese using DeepL Translate. An alignment algorithm and a parallel corpus word aligner were used to handle mismatches between translated annotations and their translated sentences.
 
2024 members can access this corpus through their LDC account. Non-members may license this data for a fee. 
 
 

Monday, October 17, 2022

LDC October 2022 Newsletter

Membership Year 2023 publication preview 

LDC data and commercial technology development 

30th Anniversary Highlight: ACE 

New publications:
Rime-Cantonese: A Normalized Cantonese Jyutping Lexicon
2017 NIST Language Recognition Evaluation Training and Development Sets
LORELEI Bengali Representative Language Pack


Membership Year 2023 publication preview The 2023 membership year is approaching and plans for next year’s publications are in progress. Among the expected releases are: 

  • AIDA Ukrainian Broadcast and Telephone Speech Audio and Transcripts: 156 hours of Ukrainian conversational telephone speech and broadcast news with 1.2 million words of corresponding orthographic transcripts 
  • 2019 NIST SRE: audiovisual and leaderboard challenge sets based on amateur videos and Tunisian Arabic telephone speech, respectively   
  • DEFT English ERE: English text from assorted genres annotated for entities, relations, and events  
  • Mixer 3 and Mixer 7 speech collections: thousands of hours of telephone speech and metadata from Mixer 3 (multiple languages) and Mixer 7 (Spanish, plus interviews and transcript readings)  
  • CALLFRIEND Russian: 100 telephone conversations among native speakers, transcripts, and a lexicon, released in separate speech and text data sets 
  • REMIX Telephone Collection: English telephone speech from 385 participants in previous Mixer studies 
  • LORELEI: representative and incident language packs containing monolingual text, bi-text, translations, annotations, supplemental resources, and related tools in various languages (e.g., Indonesian, Swahili, Tagalog, Tamil, Zulu) 


Check your inbox in the coming weeks for more information about membership renewal.  

LDC data and commercial technology development
For-profit organizations are reminded that an LDC membership is a pre-requisite for obtaining a commercial license to almost all LDC databases. Non-member organizations, including non-member for-profit organizations, cannot use LDC data to develop or test products for commercialization, nor can they use LDC data in any commercial product or for any commercial purpose. LDC data users should consult corpus-specific license agreements for limitations on the use of certain corpora. Visit the Licensing page for further information.

30th Anniversary Highlight: ACE 
The objective of the Automatic Content Extraction (ACE) program was to develop the capability to extract meaning (entities, relations and events) from multimedia sources (Doddington, et al., 2004). LDC supported ACE by creating annotation guidelines, corpora and other linguistic resources, including training and test data for the common task research evaluations (Strassel, et al., 2003Huang, et al., 2004).
There are multiple data sets in LDC’s Catalog from the program. One that regularly makes the list of LDC’s top ten most licensed corpora is ACE 2005 Multilingual Training Corpus (LDC2006T06). This data set contains 1,800 files of mixed genre text in English, Arabic, and Chinese annotated for entities, relations, and events. The genres include newswire, broadcast news, broadcast conversation, weblog, discussion forums, and conversational telephone speech. 
Another popular data set, ACE 2004 Multilingual Training Corpus (LDC2005T09), consists of varied genre text in English (158,000 words), Chinese (307,000 characters, 154,000 words), and Arabic (151,000 words) annotated for entities and relations.
ACE 2007 Multilingual Training Corpus (LDC2014T18) has the complete set of Arabic and Spanish training data for the 2007 ACE technology evaluation, specifically, Arabic and Spanish newswire data and Arabic weblogs annotated for entities and temporal expressions.
Other ACE corpora in the Catalog include ACE 2005 SpatialML Annotations in English and Mandarin (LDC2008T03LDC2010T09, LDC2011T02), Datasets for Generic Relation Extraction (reACE)TIDES Extraction (ACE) 2003 Multilingual Training DataACE-2 Version 1.0ACE Time Normalization (TERN) 2004 English Training Data v 1.0 (TERN), and more. 

For the full list of available ACE data, visit LDC’s Catalog and select the ACE research project in the search menu. For more information about linguistic resources for the ACE Program, including annotation guidelines, task definitions and other documentation, visit LDC's ACE webpage.


New publications:Rime-Cantonese: A Normalized Cantonese Jyutping Lexicon was developed by the Cantonese Computational Linguistics Infrastructure Working Group. It contains approximately 130,000 Cantonese character, word, and phrase entries paired with their corresponding romanized pronunciations in Jyutping, a scheme created by The Linguistic Society of Hong Kong.Data was collected from a variety of physical and online sources. The character collection was subjected to a normalization process for differences between traditional and simplified Chinese, regional differences and other variants in Chinese characters, and differences in orthography.2022 members can access this corpus through their LDC accounts. Non-members may license this data for a fee.


*

2017 NIST Language Recognition Evaluation Training and Development Sets contains training and development material for the 2017 NIST Language Recognition Evaluation. It consists of 2,100 hours of conversational telephone speech, broadcast conversation, broadcast narrow band speech, and speech from video in the following 14 languages, dialects, and varieties: Arabic (Iraqi, Levantine, Maghrebi, Egyptian), English (British, American), Polish, Russian, Portuguese (Brazilian), Spanish (Caribbean, European, Latin American Continental), and Chinese (Mandarin, Min Nan). The 2017 evaluation focused on differentiating closely related language pairs. Source data is from LDC's CALLFRIEND and Fisher telephone collections, the VAST video collection, various broadcast sources, and earlier NIST LRE test sets.
 
2022 members can access this corpus through their LDC accounts. Non-members may license this data for a fee.

*

LORELEI Bengali Representative Language Pack was developed by LDC and is comprised of approximately 144 million words of Bengali monolingual text, 96,000 Bengali words translated from English data, and over 2 million words of found Bengali-English parallel text. Approximately 86,000 words were annotated for named entities and up to 25,000 words were annotated for entity discovery and linking and situation frames (identifying entities, needs and issues). Data was collected from news, social network, and weblogs.The LORELEI (Low Resource Languages for Emergent Incidents) program was concerned with building human language technology for low resource languages in the context of emergent situations. Representative languages were selected to provide broad typological coverage.The knowledge base for entity linking annotation is available separately as LORELEI Entity Detection and Linking Knowledge Base (LDC2020T10).

2022 members can access this corpus through their LDC accounts. Non-members may license this data for a fee.

Tuesday, October 15, 2019

LDC 2019 October Newsletter

Membership Year 2020 Publication Preview
LDC data and commercial technology development 

New Publications: 
BOLT English Treebank - Discussion Forum 
Polish Speech Database 
2016 NIST Speaker Recognition Evaluation Test Set 
______________________________________________________________ 

Membership Year 2020 Publication Preview 

The 2020 Membership Year is just around the corner and plans for next year’s publications are in progress. Among the expected releases are:

Abstract Meaning Representation (AMR) Annotation Release 3.0: semantic treebank of over 59,000 English natural language sentences from broadcast conversations, newswire, weblogs and web discussion forums; updates the second version (LDC2017T10) with new annotations
TAC KBP: English sentiment slot filling, surprise slot filling, nugget detection and coreference, and event argument data in all languages (English, Chinese and Spanish)
DEFT Chinese ERE: Chinese discussion forum data annotated for entities, relations and events
LibriVox Spanish: 73 hours of Spanish audiobook read speech and transcripts
IARPA Babel Language Packs (telephone speech and transcripts): languages include Dhuluo, Javanese and Mongolian
HAVIC Med Training data: web video, metadata, and annotations for developing multimedia systems
RATS Speaker Identification: conversational telephone speech in Levantine Arabic, Pashto, Urdu, Farsi and Dari on degraded audio signals with annotation of speech segments for speaker identification
BOLT: discussion forums, SMS/chat, conversational telephone speech, word-aligned, tagged and co-reference data in all languages (Chinese, Egyptian Arabic, and English)

Check your inbox in the coming weeks for more information about membership renewal. 

LDC data and commercial technology development 

For-profit organizations are reminded that an LDC membership is a pre-requisite for obtaining a commercial license to almost all LDC databases. Non-member organizations, including non-member for-profit organizations, cannot use LDC data to develop or test products for commercialization, nor can they use LDC data in any commercial product or for any commercial purpose. LDC data users should consult corpus-specific license agreements for limitations on the use of certain corpora. Visit the Licensing page for further information.
______________________________________________________________

New publications:  

(1) BOLT English Treebank - Discussion Forum was developed by LDC and consists of 268,907 tokens of English web discussion forum data with part-of-speech and syntactic structure annotations collected for the DARPA BOLT (Broad Operational Language Translation) program.

Part-of-speech and treebank annotation conformed to Penn Treebank II style, incorporating changes to those guidelines that were developed under the GALE (Global Autonomous Language Exploitation) program. Supplementary guidelines for English treebanks and web text are included with this release.

The source data is English discussion forum web text collected by LDC in 2011 and 2012. A subset of that data -- 702 files representing 268,907 tokens -- was selected for the treebank and annotated for word-level tokenization, part-of-speech and syntactic structure. The unannotated English source data is released as BOLT English Discussion Forums (LDC2017T11).

BOLT English Treebank - Discussion Forum is distributed via web download.

2019 Subscription Members will automatically receive copies of this corpus. 2019 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.

* 

(2) Polish Speech Database was developed by VoiceLab and consists of 263,424 utterances of Polish speech data from 200 speakers, totaling approximately 280 hours, and corresponding transcripts.

Data collection was performed in Poland. Speakers were asked to record themselves reading text on a website for at least 60 minutes from their home computer while using a headset. The read text was comprised of sentences covering most speech sounds in Polish.

This release includes speaker metadata. There were 103 male speakers and 97 female speakers, ranging from 15 – 60 years of age; most speakers were in the 15 – 30 years age range.

Polish Speech Database is distributed via web download.

2019 Subscription Members will automatically receive copies of this corpus. 2019 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.

* 

(3) 2016 NIST Speaker Recognition Evaluation Test Set was developed by LDC and NIST (National Institute of Standards and Technology) and contains approximately 340 hours of short segments of Tagalog, Cantonese, Cebuano and Mandarin telephone speech used as development and test data in the NIST-sponsored 2016 Speaker Recognition Evaluation (SRE). 

As in previous evaluations, SRE16 focused on telephone speech recorded over a variety of handset types for the training and test conditions. In addition to development and evaluation data, this corpus also contains trial lists, their associated keys, tables containing metadata information, and evaluation documentation.

The telephone speech data was drawn from the Call My Net 2015 Corpus collected by LDC. Native speakers of Tagalog, Cantonese, Cebuano or Mandarin (220 unique speakers) made a total of ten telephone calls each to people within their existing social networks. Speakers were encouraged to use different telephone instruments in a variety of acoustic settings and were instructed to talk for 8 - 10 minutes per call on a topic of their choice. All conversations were collected outside North America.
 
2016 NIST Speaker Recognition Evaluation Test Set is distributed via web download.  

2019 Subscription Members will automatically receive copies of this corpus. 2019 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.