Showing posts with label Chinese SMS/Chat. Show all posts
Showing posts with label Chinese SMS/Chat. Show all posts

Wednesday, December 15, 2021

LDC December 2021 Newsletter

LDC 2022 Membership Discounts Now Available 

Approaching Deadline for Spring 2022 Data Scholarship Applications 

Citizen Linguistics

LDC Closed for Winter Break Dec. 24-Jan. 4


New Publications:

LDC 2022 Membership Discounts Now Available  
Now through March 1, 2022, current 2021 members receive a 10% discount for renewing their membership, and new or returning organizations receive a 5% discount. Membership remains the most economical way to access current and past LDC releases. Consult Join LDC for details on membership options and benefits. 

Approaching Deadline for Spring 2022 Data Scholarship Applications
Attention students: don’t miss out on the chance to receive no-cost access to LDC data for your research. Applications for Spring 2022 data scholarships are due January 15, 2022. For more information on requirements and program rules, see LDC Data Scholarships

Citizen Linguistics
LanguageARC (https://languagearc.com), a citizen science web portal for linguistics, continues to grow with 12 language research projects currently available to the community. Two new projects seeking contributions from citizen linguists have recently been added. The Fearless Steps project will make thousands of hours of Apollo space mission communications accessible to researchers and to the public. Contributors can listen to and annotate actual audio recordings from the Apollo 11 space mission. A second new project, Les stéréotypes en français, asks contributors to identify and classify stereotypes that can be expressed in the French language. In addition to these publicly available projects, LanguageARC also enables researchers to create research projects restricted to defined private groups, such as the recent object naming task to document the Guanzhong dialect of Mandarin. Here a private, invited group of about 60 contributors yielded over 34,000 speech recordings.

Please consider becoming an active participant in the LanguageARC community by contributing to research projects. If you are a researcher interested in creating your own project on LanguageARC, please reach out via the “Contact” page on the website.

LDC Closed for Winter Break Dec. 24-Jan. 4
LDC will be closed from Friday December 24, 2021 through Tuesday, January 4, 2022 in accordance with the University of Pennsylvania Winter Break Policy. Our offices will reopen on Wednesday, January 5, 2022. Requests received by the Membership Office during Winter Break will be processed when the office reopens. 

New publications:

(1) BOLT English Translation Treebank – Chinese SMS/Chat was developed by LDC and consists of SMS/Chat text data translated from Chinese to English and annotated for part-of-speech and syntactic structure.  

The source data is Chinese SMS and chat text collected by LDC between 2010 and 2013. A subset of the translated text -- 194 files representing 108,385 tokens -- was selected for treebanking. Part-of-speech and treebank annotation conform to Penn Treebank II style. Supplementary guidelines for English treebanks and web text are included with this release.

BOLT English Translation Treebank – Chinese SMS/Chat is distributed via web download.  

2021 Subscription Members will automatically receive copies of this corpus. 2021 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.

*

(2) HAVIC MED Training Data – Videos, Metadata and Annotation was developed by LDC and is comprised of approximately 2,100 hours of user-generated videos with annotation and metadata developed for the 2011-2015 NIST-sponsored MED (Multimedia Event Detection) tasks.

The data consists of videos of various events (event videos) and videos completely unrelated to events (background videos) harvested by a large team of human annotators. Each event video was manually annotated with judgments describing its event properties and other salient features. Background videos were labeled with topic and genre categories.

HAVIC MED Training Data -- Videos, Metadata and Annotation is distributed via web download. 

2021 Subscription Members will automatically receive copies of this corpus. 2021 Standard Members may request a copy as part of their 16 free membership corpora. This corpus is a members-only release and is not available for non-member licensing. Contact ldc@ldc.upenn.edu for information about membership.

Monday, June 18, 2018

LDC 2018 June Newsletter

LDC Catalog certified as CoreTrustSeal data repository

LDC data and commercial technology development

New Publications:
IARPA Babel Cebuano Language Pack IARPA-babel301b-v2.0b
__________________________________________________________________________

LDC Catalog certified as CoreTrustSeal data repository
LDC is pleased to announce that the Catalog has been awarded the CoreTrustSeal for recognition as a trustworthy data repository. This means that the Catalog meets a series of standards covering data access, rights management, curation, and storage developed by the ISCU World Data System and the Data Seal of Approval. LDC joins the other 136 certified repositories around the globe in the commitment to promote sustainable and trustworthy data infrastructures.  

LDC data and commercial technology development
For-profit organizations are reminded that an LDC membership is a pre-requisite for obtaining a commercial license to almost all LDC databases. Non-member organizations, including non-member for-profit organizations, cannot use LDC data to develop or test products for commercialization, nor can they use LDC data in any commercial product or for any commercial purpose. LDC data users should consult corpus-specific license agreements for limitations on the use of certain corpora. Visit the Licensing page for further information.

New publications:

(1) BOLT Chinese SMS/Chat was developed by LDC and consists of naturally-occurring Short Message Service (SMS) and Chat (CHT) data collected through data donations and live collection involving native speakers of Chinese. The corpus contains 14,877 conversations totaling 3,005,810 words across 497,543 messages.

The BOLT (Broad Operational Language Translation) program developed machine translation and information retrieval for less formal genres, focusing particularly on user-generated content. LDC supported the BOLT program by collecting informal data sources – discussion forums, text messaging, and chat – in Chinese, Egyptian Arabic, and English. The collected data was translated and annotated for various tasks including word alignment, treebanking, propbanking, and co-reference. The data in this release was collected using two methods: new collection via LDC's collection platform, and donation of SMS or chat archives from BOLT collection participants.

BOLT Chinese SMS/Chat is distributed via web download.

2018 Subscription Members will automatically receive copies of this corpus. 2018 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.

*

(2) Multi-Language Conversational Telephone Speech 2011 -- Central European was developed by LDC and is comprised of approximately 44 hours of telephone speech in two distinct language varieties of Central Europe: Czech and Slovak.

The data were collected primarily to support research and technology evaluation in automatic language identification, and portions of these telephone calls were used in the NIST 2011 Language Recognition Evaluation (LRE). Participants were recruited by native speakers who contacted acquaintances in their social network. Those native speakers made one call, up to 15 minutes, to each acquaintance. Human auditors labeled the calls for callee gender, dialect type, and noise.

LDC has also released the following as part of the Multi-Language Conversational Telephone Speech 2011 series:
·        Slavic Group (LDC2016S11)
·        Turkish (LDC2017S09)
·        South Asian (LDC2017S14)
·        Central Asian (LDC2018S03)

Multi-Language Conversational Telephone Speech 2011 -- Central European is distributed via web download.

2018 Subscription Members will automatically receive copies of this corpus. 2018 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.

*

(3) TAC KBP English Entity Linking - Comprehensive Training and Evaluation Data 2009-2013 was developed by LDC and contains training and evaluation data produced in support of the TAC KBP English Entity Linking tasks in 2009, 2010, 2011, 2012, and 2013. It includes queries and gold standard entity type information, Knowledge Base links, and equivalence class clusters for NIL entities. Also included are the source documents for the queries, specifically, English newswire, discussion forum, and web data. The corresponding knowledge base is available as TAC KBP Reference Knowledge Base (LDC2014T16). Also included in this package are the results of an Entity Linking IAA (Inter-Annotator Agreement) study conducted in 2010.

TAC KBP encourages the development of systems that can match entities mentioned in natural texts with those appearing in a knowledge base and extract novel information about entities from a document collection and add it to a new or existing knowledge base. English Entity Linking was first conducted as part of the 2009 TAC KBP evaluations. Its goal is to measure systems' ability to determine whether an entity, specified by a query, has a matching node in a reference knowledge base (KB) and, if so, to create a link between the two. If there is no matching node for a query entity in the KB, EL systems are required to cluster the mention together with others referencing the same entity.

TAC KBP English Entity Linking - Comprehensive Training and Evaluation Data 2009-2013 is distributed via web download.

2018 Subscription Members will automatically receive copies of this corpus. 2018 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.

*

(4) IARPA Babel Cebuano Language Pack IARPA-babel301b-v2.0b was developed by Appen for the IARPA (Intelligence Advanced Research Projects Activity) Babel program. It contains approximately 191 hours of Cebuano conversational and scripted telephone speech collected in 2013 and 2014 along with corresponding transcripts.

The Cebuano speech in this release represents that spoken in the Cebu-North Kana, Sialo, and Mindanao dialect regions of the Philippines. The gender distribution among speakers is approximately equal; speakers' ages range from 16 years to 75 years. Calls were made using different telephones (e.g., mobile, landline) from a variety of environments including the street, a home or office, a public place, and inside a vehicle.

IARPA Babel Cebuano Language Pack IARPA-babel301b-v2.0b is available via web download.

2018 Subscription Members will receive copies of this corpus provided they have submitted a completed copy of the special license agreement. 2018 Standard Members may request a copy as part of their 16 free membership corpora. Non-members may license this data for a fee.