BOLT English PropBank and Sense -- Discussion Forum, SMS/Chat, and Conversational Telephone Speech

Item Name: BOLT English PropBank and Sense -- Discussion Forum, SMS/Chat, and Conversational Telephone Speech
Author(s): Martha Palmer, Jena D. Hwang, Claire Bonial, Tim O'Gorman, James Gung, Kevin Stowe, Meredith Green
LDC Catalog No.: LDC2020T21
ISBN: 1-58563-943-5
ISLRN: 640-422-732-913-3
DOI: https://doi.org/10.35111/7TNG-WK28
Release Date: September 15, 2020
Member Year(s): 2020
DCMI Type(s): Text
Data Source(s): discussion forum, text chat conversations, telephone conversations
Project(s): BOLT
Application(s): question-answering, entity extraction, part of speech tagging, semantic role labelling
Language(s): English
Language ID(s): eng
License(s): LDC User Agreement for Non-Members
Online Documentation: LDC2020T21 Documents
Licensing Instructions: Subscription & Standard Members, and Non-Members
Citation: Palmer, Martha, et al. BOLT English PropBank and Sense -- Discussion Forum, SMS/Chat, and Conversational Telephone Speech LDC2020T21. Web Download. Philadelphia: Linguistic Data Consortium, 2020.
Related Works: View

Introduction

BOLT English PropBank and Sense -- Discussion Forum, SMS/Chat, and Conversational Telephone Speech was developed by the University of Colorado Boulder - CLEAR (Computational Language and Education Research) and consists of propbank and verb sense disambiguation annotation on English discussion forum (DF), SMS/Chat and conversational telephone speech (CTS) data.

The DARPA BOLT (Broad Operational Language Translation) program developed machine translation and information retrieval for less formal genres, focusing particularly on user-generated content. LDC supported the BOLT program by collecting informal data sources -- discussion forums, text messaging and chat -- in Chinese, Egyptian Arabic and English. The collected data was translated and annotated for various tasks including word alignment, treebanking, propbanking and co-reference.

Data

DF data was collected from the web using a combination of manual and automatic processes. SMS/Chat material was donated or collected via live platforms. CTS data was taken from LDC's Arabic and Chinese CALLHOME and CALLFRIEND telephone collections; the audio files were transcribed and translated into English.

Propbank annotation and verb sense disambiguation were applied to BOLT phrase structure treebank annotation, specifically, to each predicate verb in a tree. Propbank annotation provided a layer of semantic annotation over treebank and was performed on all three genres. DF and SMS/Chat data was also annotated for verb sense disambiguation using Verbnet 3.2 classes.

Annotation files are presented as UTF-8 encoded and are in either plain text or XML formats.

Acknowledgements

This material is based upon work supported by the Defense Advanced Research Projects Agency (DARPA) under Contract No. HR0011-11-C-0145. The content does not necessarily reflect the position or the policy of the Government, and no official endorsement should be inferred.

Samples

Please view this PropBank sample (TXT) and Sense sample (TXT).

Updates

None at this time.

Available Media

View Fees





Login for the applicable fee