BOLT English Discussion Forums

Item Name: BOLT English Discussion Forums
Author(s): Jennifer Tracey, Haejoong Lee, Stephanie Strassel
LDC Catalog No.: LDC2017T11
ISBN: 1-58563-806-4
ISLRN: 307-154-318-802-6
Release Date: July 18, 2017
Member Year(s): 2017
DCMI Type(s): Text
Data Source(s): discussion forum
Project(s): BOLT
Application(s): machine translation
Language(s): English
Language ID(s): eng
License(s): LDC User Agreement for Non-Members
Online Documentation: LDC2017T11 Documents
Licensing Instructions: Subscription & Standard Members, and Non-Members
Citation: Tracey, Jennifer, Haejoong Lee, and Stephanie Strassel. BOLT English Discussion Forums LDC2017T11. Web Download. Philadelphia: Linguistic Data Consortium, 2017.

Introduction

BOLT English Discussion Forums was developed by the Linguistic Data Consortium (LDC) and consists of 830,440 discussion forum threads in English harvested from the Internet using a combination of manual and automatic processes.

The DARPA BOLT (Broad Operational Language Translation) program developed machine translation and information retrieval for less formal genres, focusing particularly on user-generated content. LDC supported the BOLT program by collecting informal data sources -- discussion forums, text messaging and chat -- in Chinese, Egyptian Arabic and English. The collected data was translated and annotated for various tasks including word alignment, treebanking, propbanking and co-reference. The material in this release represents the unannotated English source data in the discussion forum genre.

Data

Collection was seeded based on the results of manual data scouting by native speaker annotators. Scouts were instructed to seek content in English that was original, interactive and informal. Upon locating an appropriate thread, scouts submitted the URL and some simple judgments about it to a database, via a web browser plug-in. When multiple threads from a forum were submitted, the entire forum was automatically harvested and added to the collection. The scale of the collection precluded manual review of all data. Only a small portion of the threads included in this release were manually reviewed, and it is expected that there may be some offensive or otherwise undesired content as well as some threads that contain a large amount of non-English content. Language identification was performed on all threads in this corpus (using CLD2), and threads for which the results indicate a high probability of largely non-English content are listed in eng_suspect_LID.txt in the docs directory of this package.

The corpus is comprised of zipped HTML and XML files. The HTML files are a raw HTML file downloaded from the discussion thread. If the thread spanned multiple URLs, it was stored as a concatenation of the downloaded HTML files. The XML files were converted from the raw HTML.

 

Acknowledgement

This material is based upon work supported by the Defense Advanced Research Projects Agency (DARPA) under Contract No. HR0011-11-C-0145. The content does not necessarily reflect the position or the policy of the Government, and no official endorsement should be inferred.

Samples

Please view this html sample and xml sample.

Updates

None at this time.

Available Media

View Fees





Login for the applicable fee