BOLT English PropBank and Sense -- Discussion Forum, SMS/Chat, and Conversational Telephone Speech
Item Name: | BOLT English PropBank and Sense -- Discussion Forum, SMS/Chat, and Conversational Telephone Speech |
Author(s): | Martha Palmer, Jena D. Hwang, Claire Bonial, Tim O'Gorman, James Gung, Kevin Stowe, Meredith Green |
LDC Catalog No.: | LDC2020T21 |
ISBN: | 1-58563-943-5 |
ISLRN: | 640-422-732-913-3 |
DOI: | https://doi.org/10.35111/7TNG-WK28 |
Release Date: | September 15, 2020 |
Member Year(s): | 2020 |
DCMI Type(s): | Text |
Data Source(s): | discussion forum, telephone conversations, text chat conversations |
Project(s): | BOLT |
Application(s): | entity extraction, part of speech tagging, question-answering, semantic role labelling |
Language(s): | English |
Language ID(s): | eng |
License(s): |
LDC User Agreement for Non-Members |
Online Documentation: | LDC2020T21 Documents |
Licensing Instructions: | Subscription & Standard Members, and Non-Members |
Citation: | Palmer, Martha, et al. BOLT English PropBank and Sense -- Discussion Forum, SMS/Chat, and Conversational Telephone Speech LDC2020T21. Web Download. Philadelphia: Linguistic Data Consortium, 2020. |
Related Works: | View |
Introduction
BOLT English PropBank and Sense -- Discussion Forum, SMS/Chat, and Conversational Telephone Speech was developed by the University of Colorado Boulder - CLEAR (Computational Language and Education Research) and consists of propbank and verb sense disambiguation annotation on English discussion forum (DF), SMS/Chat and conversational telephone speech (CTS) data.
The DARPA BOLT (Broad Operational Language Translation) program developed machine translation and information retrieval for less formal genres, focusing particularly on user-generated content. LDC supported the BOLT program by collecting informal data sources -- discussion forums, text messaging and chat -- in Chinese, Egyptian Arabic and English. The collected data was translated and annotated for various tasks including word alignment, treebanking, propbanking and co-reference.
Data
DF data was collected from the web using a combination of manual and automatic processes. SMS/Chat material was donated or collected via live platforms. CTS data was taken from LDC's Arabic and Chinese CALLHOME and CALLFRIEND telephone collections; the audio files were transcribed and translated into English.
Propbank annotation and verb sense disambiguation were applied to BOLT phrase structure treebank annotation, specifically, to each predicate verb in a tree. Propbank annotation provided a layer of semantic annotation over treebank and was performed on all three genres. DF and SMS/Chat data was also annotated for verb sense disambiguation using Verbnet 3.2 classes.
Annotation files are presented as UTF-8 encoded and are in either plain text or XML formats.
Sponsorship
This material is based upon work supported by the Defense Advanced Research Projects Agency (DARPA) under Contract No. HR0011-11-C-0145. The content does not necessarily reflect the position or the policy of the Government, and no official endorsement should be inferred.
Samples
Please view this PropBank sample (TXT) and Sense sample (TXT).
Updates
None at this time.