Showing posts with label LDC Spring Data Scholarships. Show all posts
Showing posts with label LDC Spring Data Scholarships. Show all posts

Monday, February 16, 2026

LDC February 2026 Newsletter

LDC membership discounts expire March 2 

Spring 2026 data scholarship recipient 


New publications:

2022 NIST Language Recognition Evaluation Test and Development Sets

KAIROS Schema Learning Background Source Data

LORELEI Russian Representative Language Pack

_________________________________________________________________

LDC membership discounts expire March 2 
Time is running out to save on 2026 membership fees. Renew your LDC membership, rejoin the Consortium, or become a new member by March 2 to receive a 10% discount. For more information on membership benefits and options, visit Join LDC.


Spring 2026 data scholarship recipient 

Congratulations to the recipient of LDC’s Spring 2026 data scholarship:

Doma Akshitha Reddy: Chaitanya Bharathi Institute of Technology (India): Bachelor of Engineering, Information Technology. Doma is awarded copies of TIMIT Acoustic-Phonetic Continuous Speech Corpus and The CMU Kids Corpus for their work in child speech. 

Since 2010, LDC has awarded scholarships to successful student applicants twice each year. To date more than 242 corpora have been distributed to 162 students across 38 countries. We proudly celebrate their achievements and the contributions their research has made to the broader community. 

The next round of applications will be accepted in September 2026. For information about the program, visit the Data Scholarships page.


New publications:

2022 NIST Language Recognition Evaluation Test and Development Sets was developed by LDC and  NIST and contains the test and development data, metadata, answer keys, and documentation for the 2022 NIST Language Recognition Evaluation (LRE22). The source data is comprised of 222 hours of conversational telephone speech (CTS) and broadcast narrowband speech (BNBS) in 14 languages: Afrikaans, Tunisian Arabic, Algerian Arabic, Libyan Arabic, South African English, Indian-accented South African English, North African French, Ndebele, Oromo, Tigrinya, Tsonga, Venda, Xhosa and Zulu.

For the CTS collections, a small number of native speakers made single calls to multiple individuals in their social network. Calls lasted 8-15 minutes; speakers were free to discuss any topic. The BNBS data was collected from streaming radio programming, focused on broadcasts that included narrowband speech (e.g., call-ins to a talk show). Portions of the CTS callee call sides and portions of each broadcast recording were manually audited by native speakers to verify language and quality.

LRE22 emphasized language recognition for African languages, including low resource languages, and expanded the range of test segment durations. Further information about the 2022 evaluation can be found in the 2022 NIST Language Recognition Evaluation Plan. 

2026 members can access this corpus through their LDC accounts. Non-members may license this data for a fee.

*

KAIROS Schema Learning Background Source Data was developed by LDC and includes 14,000 English and Spanish documents representing text, audio, video, image, and multimedia resources collected during the DARPA KAIROS program as supplemental background source data for the KAIROS Schema Learning Corpus (SLC). The purpose of the supplemental collection was to increase the amount of English and Spanish data with multimedia components for schema learning and to add domains not well represented in existing Spanish data. The supplemental data in this release includes material from the business and logistics domains, instructional documents and multimedia news.

The complete set of SLC background source data (including the data in this publication) totaled 16.2 million English, Russian and Spanish documents and more than 125,000 audio, video, image, or multimedia resources. A large portion of that data was drawn from pre-existing LDC datasets.

The SLC and KAIROS Schema Learning Complex Event Annotation (LDC2025T07), containing English and Spanish text, audio, video, and image material labeled for 93 real-world complex events, constitute the data used by KAIROS system developers for schema learning.

KAIROS systems utilized formal event representations in the form of schema libraries that specified the steps, preconditions and constraints for an open set of complex events; schemas were then used in combination with event extraction to characterize and make predictions about real-world events in a large multilingual, multimedia corpus.

2026 members can access this corpus through their LDC accounts. Non-members may license this data for a fee.

*

LORELEI Russian Representative Language Pack contains over 1.26 billion words of Russian monolingual text, 360,00 words of which were translated into English, 3 million words of found Russian-English parallel text, and 87,000 Russian words translated from English data. Approximately 83,000 words were annotated for simple named entities, around 26,000 words were annotated for full entity (including nominals and pronouns), entity linking and situation frames (identifying entities, needs and issues) and nearly 9,000 words were covered by noun phrase chunking annotation. Data was collected from discussion forum, news, reference, social network, and weblogs.

The LORELEI (Low Resource Languages for Emergent Incidents) program was concerned with building human language technology for low resource languages in the context of emergent situations. Representative languages were selected to provide broad typological coverage.

The knowledge base for entity linking annotation is available separately as LORELEI Entity Detection and Linking Knowledge Base (LDC2020T10).

2026 members can access this corpus through their LDC accounts. Non-members may license this data for a fee.

Monday, November 17, 2025

LDC November 2025 Newsletter

Join LDC for membership year 2026  

Spring 2026 data scholarship application deadline  

New publications:

AnnoDIFP CTS Audio and Transcripts

LORELEI Ilocano Incident Language Pack
_______________________________________________________________

Join LDC for membership year 2026  

It’s time to renew your LDC membership for 2026. Any organization that joins the Consortium or renews their membership before March 2, 2026, will receive a 10% discount off the membership fee. 

In addition to accessing new publications, current LDC members enjoy the benefit of licensing at reduced fees older data from our Catalog of close to 1000 holdings. Current-year for-profit members may use most data for commercial applications.

Plans for next year’s publications are in progress. Among the expected releases are:  

  • KAIROS schema learning corpus background data and Phase 1 evaluation datasets: multimodal English and Spanish source data and annotations for reasoning about complex real-world events 
  • CALL MY NET 2: 800+ hours of Tunisian-Arabic conversational telephone speech from over 400 speakers to support text independent speaker recognition, used in the 2018 NIST Speaker Recognition Evaluation
  • Multi-language conversational telephone speech: multiple releases, hundreds of hours of speech from speakers of confusable linguistic varieties (Arabic, Chinese, English, French, Slavic, Spanish) to support language identification
  • CALLHOME omnibus releases: combined speech and transcript datasets with updated directory structure, file formats and documentation, and lexicons (Chinese, English, German, Japanese, Spanish)
  • IARPA MATERIAL language packs: conversational telephone speech, transcripts, English translations, annotations and queries in multiple languages (e.g., Lithuanian, Pashto, Swahili, Tagalog)
For full descriptions of all LDC data sets, browse our Catalog. Visit Join LDC for details on membership, user accounts and payment.

Spring 2026 data scholarship application deadline
Applications are now being accepted through January 15, 2026 for the Spring 2026 LDC data scholarship program which provides university students with no-cost access to LDC data. Consult the LDC Data Scholarships page for more information about program rules and submission requirements.
 
New publications:
AnnoDIFP (Annotated Data for the Investigation of Facets of Personality) CTS (Conversational Telephone Speech) Audio and Transcripts was developed by LDC, the Florida Institute of Technology and the University of New Haven to support algorithm development for predicting personality traits. It contains 242.52 hours of English telephone audio recordings and transcripts from 1,179 telephone calls involving 327 participants paired with scores from two self-reported personality assessments, HEXACO Personality Inventory (Revised) (HEXACO-PI-R) and Short Dark Triad (SD3).

This corpus contains audio and transcripts for 277 participants and transcripts only for 50 participants. Telephone calls were collected using LDC's robot-operator platform. The operator called participants every 24 hours during their indicated availability and paired them with another participant to speak on a prompted topic for 10 minutes. Transcripts were produced automatically using the Rev.ai speech-to-text service.

2025 members can access this corpus through their LDC accounts. Non-members may license this data for a fee.

*

LORELEI Ilocano Incident Language Pack was developed by LDC and is comprised of 8.9 million words of Ilocano monolingual text, 3.3 million words of English monolingual text, 3.2 million words of parallel Ilocano-English text, and 3 million words annotated for entity discovery and linking and situation frames. It constitutes all of the text data, annotations, supplemental resources and related software tools for the Ilocano language used in the DARPA LORELEI / LoReHLT 2019 Evaluation.

The LORELEI (Low Resource Languages for Emergent Incidents) program was concerned with building human language technology for low resource languages in the context of emergent situations. In the evaluation scenario, an unforeseen event triggered a need for humanitarian and logistical support in a region where the incident language had received little or no attention in NLP research. Evaluation participants provided NLP solutions, including information extraction and machine translation, with limited resources and limited development time.

Data was collected from news, social network, weblog, newsgroup, discussion forum, and reference material. Entity discovery and linking annotation identified entities to be detected by systems for scoring purposes. Situation frame analysis was designed to extract basic information about needs and relevant issues for planning a disaster response effort.

2025 members can access this corpus through their LDC accounts. Non-members may license this data for a fee.