Digital Archive of Southern Speech - NLP Version

Item Name: Digital Archive of Southern Speech - NLP Version
Author(s): William A. Kretzschmar Jr., Paulina Bounds, Jacqueline Hettel, Steven Coats, Lee Pederson, Lisa Lena Opas-Hänninen, Ilkka Juuso, Tapio Seppänen
LDC Catalog No.: LDC2016S05
ISBN: 1-58563-761-0
ISLRN: 920-059-271-034-1
DOI: https://doi.org/10.35111/v4g6-nx14
Release Date: July 15, 2016
Member Year(s): 2016
DCMI Type(s): Sound
Sample Type: pcm
Sample Rate: 16000
Data Source(s): field recordings
Project(s): Linguistic Atlas Project
Application(s): discourse analysis, sociolinguistics
Language(s): English
Language ID(s): eng
License(s): Digital Archive of Southern Speech - NLP Version For-Profit Member Agreement
LDC User Agreement for Non-Members
Online Documentation: LDC2016S05 Documents
Licensing Instructions: Subscription & Standard Members, and Non-Members
Citation: Kretzschmar Jr., William A., et al. Digital Archive of Southern Speech - NLP Version LDC2016S05. Web Download. Philadelphia: Linguistic Data Consortium, 2016.
Related Works: View

Introduction

Digital Archive of Southern Speech - NLP Version (DASS-NLP) was developed by LDC as an alternate version of Digital Archive of Southern Speech (DASS) (LDC2012S03) suitable for natural language processing and human language technology applications. Specifically, the original audio files have been converted to 16kHz 16-bit flac compressed wav and file names have been normalized to facilitate automatic processing.

DASS was developed by the University of Georgia. It is a subset of the Linguistic Atlas of the Gulf States (LAGS), which is in turn part of the Linguist Atlas Project (LAP). DASS-NLP contains approximately 366 hours of English speech data from 30 female speakers and 34 male speakers in flac compressed wav format, along with associated metadata about the speakers and the recordings and maps in .jpeg format relating to the recording locations.

LAP consists of a set of survey research projects about the words and pronunciation of everyday American English, the largest project of its kind in the United States. Interviews with thousands of native speakers across the country have been carried out since 1929. LAGS surveyed the everyday speech of Georgia, Tennessee, Florida, Alabama, Mississippi, Arkansas, Louisiana, and Texas in a series of 914 audio-taped interviews conducted from 1968-1983. Interviews average approximately six hours in length; the systematic LAGS tape archive amounts to 5500 hours of sound recordings. DASS is a collection of 64 interviews from LAGS selected to cover a range of speech across the region and to represent multiple education levels and ethnic backgrounds.

Data

The DASS-NLP speakers' average age is 61 years; there are 30 women and 34 men from the Gulf States region represented in this release. The interviews cover common topics such as family, the weather, household articles and activities, agriculture and social connections.

The interviews were originally recorded in the field on reel-to-reel audio tape. A digital version of every reel of tape was then made, one .wav file per reel, usually about one hour of sound. Each interview thus consists of a set of 3 to 13 reels, or roughly 3 to 13 interview hours. Personally identifying or sensitive information in the files was replaced with a tone to protect the privacy and to assure ethical treatment of speakers.

Samples

Please listen to this sample.

Updates

None at this time.

Authorship

The following people were involved with the DASS project:

William A. Kretzschmar, Jr., Paulina Bounds, Jacqueline Hettel and Steven Coats
University of Georgia

Lee Pederson
Emory University

Lisa Lena Opas-Hänninen, Ilkka Juuso and Tapio Seppänen
University of Oulu (Finland)

Sponsorship

The Atlas Data contained herein comprises information collected in the period spanning from the 1930s to 2010 and has been compiled from diverse sources, by, and under the direction of, Dr. William A. Kretzschmar, Harry and Jane Wilson Professor in Humanities at the Department of English of The University of Georgia.

Compilation and digitalization of this work was funded, in part, by the US National Science Foundation and by the US National Endowment for the Humanities.

Additional information about the Atlas Project can be obtained at http://www.lap.uga.edu/.

Available Media

View Fees





Login for the applicable fee