Showing posts with label Chinese Mandarin speech. Show all posts
Showing posts with label Chinese Mandarin speech. Show all posts

Tuesday, September 15, 2026

LDC September 2026 Newsletter

LDC data and commercial technology development        

New publications:

CALLHOME Mandarin Chinese Second Edition

CALLHOME Mandarin Chinese Lexicon Second Edition

MATERIAL Lithuanian-English Language Pack

_____________________________________________________________________

LDC data and commercial technology development
For-profit organizations are reminded that an LDC membership is a pre-requisite for obtaining a commercial license to almost all LDC databases. Non-member organizations, including non-member for-profit organizations, cannot use LDC data to develop or test products for commercialization, nor can they use LDC data in any commercial product or for any commercial purpose. LDC data users should consult corpus-specific license agreements for limitations on the use of certain corpora. Visit the Licensing page for further information.

New publications:

CALLHOME Mandarin Chinese Second Edition was developed by LDC and contains 38 hours of speech from 120 unscripted telephone conversations between native Mandarin Chinese speakers. This publication is a re-release of the original CALLHOME Mandarin Chinese collection, combining CALLHOME Mandarin Chinese Speech (LDC96S34) and CALLHOME Mandarin Chinese Transcripts (LDC96T16), with additional transcription and updated directory structure, file formats, and documentation.

This release contains the 120 telephone conversations published in CALLHOME Mandarin Chinese Speech (LDC96S34) which represented training and development data and a subset of evaluation data. Participants spoke on topics of their choice in a single telephone call lasting up to 30 minutes. Calls were manually audited for gender, language, recording quality, channel characteristics, dialect, and accent. For this second edition, all audio was converted from SPHERE files to FLAC format, and the original training/development/evaluation partitioning was removed. 

This release also features revised transcripts conforming to updated LDC transcription guidelines that addressed normalization of annotation formats, standardization of speaker-produced and background noises, application of foreign-language marking, whitespace cleanup, and corrections and consistency fixes.

The CALLHOME series consists of telephone conversations and transcripts developed by LDC and Rutgers, The State University of New Jersey, in support of research in speaker identification, language identification and related technologies. Languages in the series include American English, Egyptian Arabic, German, Japanese, Mandarin Chinese, and Spanish.
2026 members can access this corpus through their LDC accounts. Non-members may license this data for a fee. 

*

CALLHOME Mandarin Chinese Lexicon Second Edition was developed by LDC and contains 44,404 Mandarin Chinese words with morphological, phonological and frequency information. This second edition updates file formats, directory structure and documentation. The first edition is available as CALLHOME Mandarin Chinese Lexicon (LDC96L15). 

The words in the lexicon were derived from transcripts representing unscripted telephone conversations between native Mandarin Chinese speakers contained in CALLHOME Mandarin Chinese Second Edition (LDC2026S11) and from Chinese news text.

The lexicon contains seven tab-separated information fields: (1) headword: orthographic representation of the word in hanzi (e.g., 没有); (2) pinyin: headword transcribed in pinyin (e.g., mei2 you3); (3) tone: tone sequence for headword (e.g., 2 3); (4) pron: pronunciation of headword without tone information (e.g., mey yow); (5) pos: part-of-speech tag for the headword; (6) xinhua_freq: frequency of the headword in newswire text; and (7) train_freq: frequency of the headword in the CALLHOME transcripts. It is presented as a tab-delimited TSV file encoded in UTF-8 format and includes a pronunciation dictionary derived from the lexicon in UTF-8 encoded CMUdict format. 

2026 members can access this corpus through their LDC accounts provided they have submitted a completed copy of the special license agreement. Non-members may license this data for a fee.

*

MATERIAL Lithuanian-English Language Pack was developed by Appen for the IARPA MATERIAL program and contains 64 hours of Lithuanian conversational telephone speech, transcripts, English translations, annotations and queries. Calls were made using different telephones (e.g., mobile, landline) from a variety of environments. Transcripts cover approximately 100% of the speech files, 6% of which were translated into English. This release also includes domain annotations, English queries and their relevance annotations. 

The MATERIAL program focused on underserved languages with the ultimate goal to build cross language information retrieval systems to find speech and text content using English search queries.

2026 members can access this corpus through their LDC accounts provided they have submitted a completed copy of the special license agreement. Non-members may license this data for a fee.

Monday, March 15, 2021

LDC 2021 March Newsletter

LDC data and commercial technology development 

New Publications:
Columbia Games Corpus
Global TIMIT Mandarin Chinese
BOLT Chinese Co-reference – Discussion Forum, SMS/Chat, and Conversational Telephone Speech

_________________________________________________________________________


LDC data and commercial technology development
For-profit organizations are reminded that an LDC membership is a pre-requisite for obtaining a commercial license to almost all LDC databases. Non-member organizations, including non-member for-profit organizations, cannot use LDC data to develop or test products for commercialization, nor can they use LDC data in any commercial product or for any commercial purpose. LDC data users should consult corpus-specific license agreements for limitations on the use of certain corpora. Visit the Licensing page for further information.



New publications:

(1) Columbia Games Corpus was developed by the Spoken Language Group, Columbia University and the Department of Linguistics, Northwestern University. It consists of approximately 10 hours of spontaneous English conversation from 13 subjects playing a series of computer games that required verbal communication to achieve joint goals of identifying and moving images on the screen to reach a combined number of points. This publication also includes corresponding manually time-aligned orthographic transcripts and annotation marking discourse and turn-taking.

Columbia Games Corpus is distributed via web download.

2021 Subscription Members will automatically receive copies of this corpus provided they have submitted a completed copy of the special license agreement. 2021 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee. 


*

(2) Global TIMIT Mandarin Chinese was developed by LDC and Shanghai Jiao Tong University and consists of five hours of read speech from Chinese Gigaword Fifth Edition (LDC2011T13) with corresponding transcripts. Fifty speakers read 120 sentences; specifically, 20 sentences were read by all speakers, 40 sentences were read by 10 speakers, and 60 sentences were read by one speaker, for a total of 3220 sentence types.

The corpus was recorded at Shanghai Jiao Tong University, China. Speakers (25 female, 25 male) were students at the university and had achieved Class 2 Level 1 or better on Putonghua Shuiping Ceshi (the national standard Mandarin proficiency test).

The Global TIMIT project aimed to create a series of corpora in a variety of languages with a similar set of key features as in the original 
TIMIT Acoustic-Phonetic Continuous Speech Corpus (LDC93S1) which was designed for acoustic-phonetic studies and for the development and evaluation of automatic speech recognition systems. 

Global TIMIT Mandarin Chinese is distributed via web download.

2021 Subscription Members will automatically receive copies of this corpus. 2021 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.


*


(3) BOLT Chinese Co-reference – Discussion Forum, SMS/Chat, and Conversational Telephone Speech was developed by Raytheon BBN Technologies and consists of co-reference annotation on Chinese informal text. 

Co-reference annotation aims to fill in connections between specific mentions in the text that refer to the same entities and events in the discourse context. BOLT co-reference annotation was performed on BOLT treebank annotation (i.e., 
Chinese Treebank 9.0 (LDC2016T13)) and covers noun phrases (including proper nouns, nominals, pronouns, and null arguments), possessives, proper noun pre-modifiers, and verbs.

Discussion forum data was collected from the web using a combination of manual and automatic processes. SMS/Chat material was donated or collected via live platforms. Telephone speech data was taken from LDC's Chinese CALLHOME and CALLFRIEND telephone collections.

The DARPA 
BOLT (Broad Operational Language Translation) program developed machine translation and information retrieval for less formal genres, focusing particularly on user-generated content. LDC supported the BOLT program by collecting informal data sources -- discussion forums, text messaging and chat -- in Chinese, Egyptian Arabic and English. The collected data was translated and annotated for various tasks including word alignment, treebanking, propbanking and co-reference.

BOLT Chinese Co-reference – Discussion Forum, SMS/Chat, and Conversational Telephone Speech is distributed via web download.

2021 Subscription Members will automatically receive copies of this corpus. 2021 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.

 

                                                                          

Friday, December 6, 2019

LDC 2019 December Newsletter


LDC Membership Discounts for MY2020 Still Available
Spring 2020 Data Scholarship Program – deadline approaching 
Introducing LanguageArc: A Citizen Linguist Portal 

New Publications: 
MagicData Chinese Mandarin Conversational Speech 
BOLT Egyptian Arabic-EnglishWord Alignment -- SMS/Chat Training 
TAC KBP Entity Discovery and Linking - Comprehensive Evaluation Data 2016-2017 
__________________________________________________________ 

LDC Membership Discounts for MY2020 Still Available 

Join LDC while membership savings are still available. Now through March 2, 2020, current MY2019 members who renew their LDC membership receive a 10% discount off the membership fee. New or returning member organizations receive a 5% discount through March 2. Membership remains the most economical way to access LDC releases. Visit Join LDC for details on membership options and benefits. 

Spring 2020 Data Scholarship Program – deadline approaching 

Students can apply for the Spring 2020 Data Scholarship Program now through January 15, 2020. The LDC Data Scholarship program provides students with no-cost access to LDC data. For more information on application requirements and program rules, please visit LDC Data Scholarships.

Introducing LanguageArc: A Citizen Linguist Portal 

LanguageARC is a citizen science website for languages developed with a grant from the National Science Foundation (no. 170377). Contributors to this online community – “citizen linguists” – participate in a variety of tasks and activities that support linguistic research, such as identifying accents from audio clips, recording “tongue twisters,” and translating English sentences into other languages. Data collected from LanguageArc will be made freely available to the research community. New collection and annotation projects will be added on an ongoing basis, and researchers will soon be able to create their own LanugageArc projects with an easy-to-use Project Builder Toolkit.  All are encouraged to explore the site and participate in the community. Comments, questions and suggestions are welcome via the site’s Contact page. 
___________________________________________________________

New publications:

(1) Magic Data Chinese Mandarin Conversational Speech was developed by Beijing Magic Data Technology Co., Ltd. and consists of approximately 10 hours of Mandarin conversational speech from 60 speakers. Each conversation was recorded on multiple devices and is presented in multiple forms, resulting in a total of approximately 60 hours of audio with corresponding transcripts.

All participants were native speakers of Mandarin in Mainland China from accent regions across the country. Speakers were paired for conversations on a range of topics, including travel, fitness, games, sports and pets. Metadata such as topic, collection date, mobile device and speaker demographic information is available in the documentation accompanying this release. 

Magic Data Chinese Mandarin Conversational Speech is distributed via web download.

2019 Subscription Members will automatically receive copies of this corpus. 2019 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.

*

(2) BOLT Egyptian Arabic-English Word Alignment -- SMS/Chat Training was developed by LDC and consists of 349,414 words of Egyptian Arabic and English parallel text enhanced with linguistic tags to indicate word relations.

This release contains Egyptian Arabic source text message and chat conversations collected using two methods: new collection via LDC's collection platform, and donation of SMS or chat archives from BOLT collection participants. The source data is released as BOLT Egyptian Arabic SMS/Chat and Transliteration (LDC2017T07).

The BOLT word alignment task was built on treebank annotation. Egyptian Arabic source tree tokens were automatically extracted from tree files in LDC’s BOLT Egyptian Arabic Treebank, which had been tagged for part-of-speech and syntactically annotated. That data was then aligned and annotated for the word alignment task. 

BOLT Egyptian Arabic-English Word Alignment -- SMS/Chat Training is distributed via web download.

2019 Subscription Members will automatically receive copies of this corpus. 2019 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee. 

*

(3) TAC KBP Entity Discovery and Linking - Comprehensive Evaluation Data 2016-2017 was developed by LDC and contains training and evaluation data produced in support of the TAC KBP Entity Discovery and Linking (EDL) tasks in 2016 and 2017. This corpus includes queries, knowledge base (KB) links, equivalence class clusters for NIL entities, and entity type information for each of the queries. The EDL reference KB, to which EDL data are linked, is available separately in TAC KBP Entity Discovery and Linking - Comprehensive Training and Evaluation Data 2014-2015 (LDC2019T02). 

The goal of the EDL track is to conduct end-to-end entity extraction, linking and clustering. For producing gold standard data, given a document collection, annotators (1) extract (identify and classify) entity mentions (queries), link them to nodes in a reference KB and (2) perform cross-document co-reference on within-document entity clusters that cannot be linked to the KB.

Source data for the annotations consists of Chinese, English and Spanish newswire and discussion forum text collected by LDC and is available in TAC KBP Evaluation Source Corpora 2016-2017 (LDC2019T12).

TAC KBP Entity Discovery and Linking - Comprehensive Evaluation Data 2016-2017 is distributed via web download.


2019 Subscription Members will automatically receive copies of this corpus. 2019 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.

Friday, March 15, 2019

LDC 2019 March Newsletter

Call for Papers - LTC 2019, LREC 2020

New Publications:
___________________________________________________________

Call for Papers

The 9th Language & Technology Conference (LTC 2019) will take place on May 17-19, 2019 at the Adam Mickiewicz University in Poznań, Poland. LTC addresses Human Language Technologies as a challenge for computer science, linguistics and related fields. Conference papers are due next week on Wednesday, March 20, 2019 (midnight, any time zone). For more information, visit the conference webpage. 

The 12th Conference on Language Resources and Evaluation (LREC 2020) will take place on May 13-15, 2020 at the Palais du Pharo in Marseille, France. LREC aims to provide an overview of the state-of-the-art, explore new R&D directions and emerging trends, and exchange information regarding language resources and their applications, evaluation methodologies and tools. Conference papers are due by November 25, 2019. For more information, including conference topics, visit the conference webpage.

New Publications:

(1) CALLFRIEND Egyptian Arabic Second Edition was developed by LDC and consists of approximately 25 hours of unscripted telephone conversations between native speakers of Egyptian Arabic. This second edition updates the audio files to wav format, simplifies the directory structure and adds documentation and metadata. The first edition is available as CALLFRIEND Egyptian Arabic (LDC96S49).

All data was collected before July 1997. Participants could speak with a person of their choice on any topic; most called family members and friends. All calls originated in North America. The recorded conversations last up to 30 minutes. 

CALLFRIEND Egyptian Arabic Second Edition is distributed via web download.

2019 Subscription Members will automatically receive copies of this corpus. 2019 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.  

*

(2) Penn Discourse Treebank Version 3.0 is the third release in the Penn Discourse Treebank project, the goal of which is to annotate the Wall Street Journal (WSJ) section of Treebank-2 (LDC95T7) with discourse relations. Penn Discourse Treebank Version 2 (LDC2008T05) contains over 40,600 tokens of annotated relations. In Version 3, an additional 13,000 tokens were annotated, certain pairwise annotations were standardized, new senses were included and the corpus was subject to a series of consistency checks.

This corpus contains two tools: (1) The Annotator, used for annotation and adjudication, and which can also be used for viewing the corpus; and (2) The Conversion Tool for converting Version 2 annotation files into the Version 3 format.

The documentation directory contains a manual describing what is new in Version 3 and how Version 3 differs from Version 2; the methods and guidelines used in annotating PDTB Version 3; and a range of statistics on the tokens, including the frequency of each connective, its sense labels and its modifiers. More information about the corpus and research carried out by the developers and others using the corpus can be found on the PDTB website.

Penn Discourse Treebank Version 3.0 is distributed via web download. 

2019 Subscription Members will automatically receive copies of this corpus. 2019 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.  

*

(3) VAST Chinese Speech and Transcripts was developed by LDC for the VAST (Video Annotation for Speech Technologies) project and is comprised of approximately 29 hours of Mandarin Chinese audio extracted from amateur video content harvested from the web and corresponding time-aligned transcripts. 

Audio files were transcribed using XTrans, which supports manual transcription across multiple channels, languages and platforms. Transcribers followed a Quick-Rich Transcription style; transcription guidelines are included in this release. 

The aim of the VAST project was to collect and annotate data in several languages to support the development of speech technologies such as speech activity detection, language identification, speaker identification, and speech recognition. 

VAST Chinese Speech and Transcripts is distributed via web download.

2019 Subscription Members will automatically receive copies of this corpus. 2019 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.