Tuesday, September 15, 2026

LDC September 2026 Newsletter

LDC data and commercial technology development        

New publications:

CALLHOME Mandarin Chinese Second Edition

CALLHOME Mandarin Chinese Lexicon Second Edition

MATERIAL Lithuanian-English Language Pack

_____________________________________________________________________

LDC data and commercial technology development
For-profit organizations are reminded that an LDC membership is a pre-requisite for obtaining a commercial license to almost all LDC databases. Non-member organizations, including non-member for-profit organizations, cannot use LDC data to develop or test products for commercialization, nor can they use LDC data in any commercial product or for any commercial purpose. LDC data users should consult corpus-specific license agreements for limitations on the use of certain corpora. Visit the Licensing page for further information.

New publications:

CALLHOME Mandarin Chinese Second Edition was developed by LDC and contains 38 hours of speech from 120 unscripted telephone conversations between native Mandarin Chinese speakers. This publication is a re-release of the original CALLHOME Mandarin Chinese collection, combining CALLHOME Mandarin Chinese Speech (LDC96S34) and CALLHOME Mandarin Chinese Transcripts (LDC96T16), with additional transcription and updated directory structure, file formats, and documentation.

This release contains the 120 telephone conversations published in CALLHOME Mandarin Chinese Speech (LDC96S34) which represented training and development data and a subset of evaluation data. Participants spoke on topics of their choice in a single telephone call lasting up to 30 minutes. Calls were manually audited for gender, language, recording quality, channel characteristics, dialect, and accent. For this second edition, all audio was converted from SPHERE files to FLAC format, and the original training/development/evaluation partitioning was removed. 

This release also features revised transcripts conforming to updated LDC transcription guidelines that addressed normalization of annotation formats, standardization of speaker-produced and background noises, application of foreign-language marking, whitespace cleanup, and corrections and consistency fixes.

The CALLHOME series consists of telephone conversations and transcripts developed by LDC and Rutgers, The State University of New Jersey, in support of research in speaker identification, language identification and related technologies. Languages in the series include American English, Egyptian Arabic, German, Japanese, Mandarin Chinese, and Spanish.
2026 members can access this corpus through their LDC accounts. Non-members may license this data for a fee. 

*

CALLHOME Mandarin Chinese Lexicon Second Edition was developed by LDC and contains 44,404 Mandarin Chinese words with morphological, phonological and frequency information. This second edition updates file formats, directory structure and documentation. The first edition is available as CALLHOME Mandarin Chinese Lexicon (LDC96L15)

The words in the lexicon were derived from transcripts representing unscripted telephone conversations between native Mandarin Chinese speakers contained in CALLHOME Mandarin Chinese Second Edition (LDC2026S11) and from Chinese news text.

The lexicon contains seven tab-separated information fields: (1) headword: orthographic representation of the word in hanzi (e.g., 没有); (2) pinyin: headword transcribed in pinyin (e.g., mei2 you3); (3) tone: tone sequence for headword (e.g., 2 3); (4) pron: pronunciation of headword without tone information (e.g., mey yow); (5) pos: part-of-speech tag for the headword; (6) xinhua_freq: frequency of the headword in newswire text; and (7) train_freq: frequency of the headword in the CALLHOME transcripts. It is presented as a tab-delimited TSV file encoded in UTF-8 format and includes a pronunciation dictionary derived from the lexicon in UTF-8 encoded CMUdict format. 

2026 members can access this corpus through their LDC accounts provided they have submitted a completed copy of the special license agreement. Non-members may license this data for a fee.

*

MATERIAL Lithuanian-English Language Pack was developed by Appen for the IARPA MATERIAL program and contains 64 hours of Lithuanian conversational telephone speech, transcripts, English translations, annotations and queries. Calls were made using different telephones (e.g., mobile, landline) from a variety of environments. Transcripts cover approximately 100% of the speech files, 6% of which were translated into English. This release also includes domain annotations, English queries and their relevance annotations. 

The MATERIAL program focused on underserved languages with the ultimate goal to build cross language information retrieval systems to find speech and text content using English search queries.

2026 members can access this corpus through their LDC accounts provided they have submitted a completed copy of the special license agreement. Non-members may license this data for a fee.

No comments:

Post a Comment