CALLHOME Mandarin Chinese Second Edition October 17, 2023 Linguistic Data Consortium 1. Overview =========== This is an updated release of the CALLHOME Mandarin corpus. The original CALLHOME corpus was collected and transcribed by the Linguistic Data Consortium primarily in support of the project on Large Vocabulary Conversational Speech Recognition (LVCSR), sponsored by the U.S. Department of Defense. This re-release combines the original CALLHOME Mandarin Speech (LDC96S34) and Transcripts (LDC96T16) corpora, and updates the directory structure, file formats, documentation, etc. to modern standards. 2. Directory structure ====================== - data/flac/ -- FLAC files containing call audio - data/trans_orig/ -- original transcripts in WebTrans TSV format - data/trans_updated/ -- updated transcripts in WebTrans TSV format - docs/calldata.tbl -- basic information about each call including what partition (train/dev/test) it belongs to and results from the call quality audit - docs/doc_calldata.txt -- documentation of "calldata.tbl" - docs/speakerdata.tbl -- audit-derived information about the transcribed speakers - docs/doc_speakerdata.txt -- documentation of "speakerdata.tbl" - docs/pindata.tbl -- participant supplied demographics for initiator of each call - docs/doc_pindata.txt -- documentation of "pindata.tbl" - docs/lex_segm.txt -- a documentation of general principles for Mandarin word segmentation - docs/file.tbl -- listing of md5 checksums, sizes, dates, and file names - docs/README.txt -- this file - docs/transcription_specs_orig.txt -- documentation of the original transcription specifications - docs/transcription_specs_updated.pdf -- documentation of updated transcription specifications 3. CALLHOME =========== 3.1 Data acquisition -------------------- Speakers were solicited by the LDC to participate in this telephone speech collection effort through personal contacts and appeals to organizations. A total of 200 call originators were found, each of whom placed a telephone call via a toll-free robot operator maintained originally by Rutgers University, and later by the LDC. Access to the robot operator was possible via a unique Personal Identification Number (PIN) issued by the recruiting staff at Rutgers or the LDC when the caller enrolled in the project. The participants were made aware that their telephone call would be recorded, as were the call recipients. The call was allowed only if both parties agreed to being recorded. Each caller was allowed to talk up to 30 minutes. Each caller was allowed to place only one telephone call. The 200 conversations originally collected involved calls originating in the U.S. and Canada, and placed to callees overseas. Most participants called family members or close friends overseas. In all, 200 calls were transcribed. Of these, 80 were designated as training calls, 20 as development test calls, and 100 as evaluation test calls. Of these 100 evaluation test calls, 20 were exposed in the original CALLHOME releases (LDC96S34 and LDC97T16); the remainder were withheld for use in future evaluations and remain unexposed in this release. For each of the training and development test calls, a contiguous 10-minute region was selected for transcription; for the evaluation test calls, a 5-minute region was transcribed. 3.2 Data verification --------------------- After a successful call was completed, a human audit of each telephone call was conducted to verify that the proper language was spoken, to check the quality of the recording, and to select and describe the region to be transcribed. The description of the transcribed region provides information about channel quality, number of speakers, their gender, and other attributes. The information about each call may be found in the file "docs/calldata.tbl", and its contents are described in greater detail in "docs/doc_calldata.txt". The audit-derived information about the transcribed speakers may be found in the file "docs/speakerdata.tbl", whose contents are described in the file "docs/doc_speakerdata.txt". 3.3 Speaker demographics ------------------------ Information on speaker demographics can be found in the file "docs/pindata.tbl", whose contents are described in the file "docs/doc_pindata.txt". 4 Word segmentation ------------------- Word segmentation principles for Mandarin were formulated by Shudong Huang at the Linguistic Data Consortium, with input from Xuejun Bian and Cynthia McLemore, and in subsequent collaboration with LVCSR CALLHOME contractors and other interested parties (especially Dragon, BBN, IBM, TI, NIST, and Bell Labs). A primary source of information on Chinese segmentation issues was the following: "Contemporary Chinese Language Word Segmentation Specification for Information Processing," published by the State Bureau of Technology Supervision, Beijing China, October 14, 1992. These principles are described in "docs/lex_segm.txt". The CALLHOME Mandarin transcripts were automatically segmented in accordance with these principles using the Dragon Mandarin segmenter. Further information on the Dragon segmenter, provided by Dean Bandes at Dragon Systems Inc.: The Dragon Mandarin Segmenter attempts to break a string of Chinese characters (in the GB encoding) into the most likely sequence of words in its lexicon and unknown words. Characters which do not fit into known words are output as unknown single-character words. In general, longer words are preferred, but not at the expense of introducing new unknown single- character words. An analogy is the case of a sign maker who has a stock of strings of letters with a cost associated to each string, who wants to produce a given sign for the minimum cost. The cost of each string is based on its frequency (supply and demand!), and the cost of the entire sign is the sum of the cost of the strings plus a relatively large cost per string (labor to put them together). All letters are available individually, but to save on labor cost the sign maker will choose not to use them if the sign can be made of existing combinations of letters. Simply proceeding from the beginning of the sign and choosing the longest available string may not produce the least expensive sign, as one may overshoot a preferable string; for instance, if the desired sign were "No Parking" and the stock of strings were "No ", "Park", "Par", "king", "i", "n", and "g", always choosing the longest would give "No " + "Park" + "i" + "n" + "g", which would be more expensive than "No " + "Par" + "king". Simply starting at the end and working backwards always choosing the longest word won't work any better. The low-cost segmentation is either a single lexicon entry or the combination of the lowest-cost segmentations of two substrings, and thus could in principle be found recursively, trying segmentations of substrings at each break point; but this would require a lot of duplicated effort. The problem may be solved much more efficiently by standard dynamic programming methods, and that's what the Dragon Segmenter does. Basically, it remembers the lowest-cost way to segment text up to each character in turn, looking ahead for matches and noting the cost to the end of each matching string if it is a new low-cost way to that character. The input lexicon must have the format [] -- that is, the pinyin is optional but the count (frequency) is required; and the word with the largest count must be first. Please contact Dean Bandes at Dragon Systems for further information or to obtain a copy of the program. Dragon Systems 320 Nevada Street Newton MA 02160 Phone: (617) 965-5200 x221 Fax: (617) 244-3899 e-mail: deanb@dragonsys.com 5. File formats =============== 5.1 Audio --------- Audio is provided as 8 kHz, 16 bit two channel FLAC files converted from the original SHORTEN compressed SPHERE files. No resampling or additional processing was performed. 5.2 Transcripts --------------- The transcripts are released as UTF-8 TSV files in the format output by WebTrans (a web-based transcription tool in use at LDC). Each file consists of a sequence of transribed speech segments, one per line, each line having the following six tab-delimited fields: - Audio -- basename of audio file - Channel -- channel segment is on in audio file (1-indexed) - Beg -- onset of speaker turn in seconds from beginning of audio file - End -- offset of speaker turn in seconds from beginning of audio file - Text -- transcript - Speaker -- speaker id; within CALLHOME speaker ids are only guaranteed to be unique within the scope of a call 6. Transcription ================ In this release we provide two versions of the transcripts: - the version previously released in LDC96T16 (Section 6.1) - an "updated" version which has been transformed to more closely resemble output of current LDC transcription tasks (Section 6.2) 6.1 Original transcription -------------------------- The transcripts are identical to those in the LDC96T16 release except that the text encoding has been updated to UTF-8. For details regarding the original transcription guidelines, please see the document: docs/transcription_specs_orig.txt 6.2 Updated transcription ------------------------- The updated transcripts conform to the most recent version of in-house LDC transcription specifications as described in: docs/transcription_specs_updated.pdf with the following exceptions: - When the initial portion of a word is ellided, this is indicated by '-'; e.g. Three senators -stained from the vote. - {noise} indicates a background noise not made by a speaker - Speech in foreign language is marked by inline XML elements. The element is named "foreign" and has a single mandatory "lang" attribute. The value of this attribute is the ISO 639-3 code for the language. E.g.: hellow A nonexhaustive list of the transformations applied: - Normalized unintelligible regions annotations; e.g., by removing of leading/trailing whitespace: (( text)) -> ((text)) - Normalized speaker-produced noises to the set allowed in modern guidelines: - {laugh} - {cough} - {breath} - {lipsmack} - {NSV} (catch all for all other noises) E.g., {breath_noise} -> {breath} - Mapped all background noises not made by speaker to {noise}; e.g. [click] --> {noise} [static] --> {noise} - Removed background or channel sound annotations, e.g., [echoing] text [/echoing] -> text - Removed inline comments, e.g., [[distortion]] -> '' - Use inline XML to mark speech from foreign languages. Maked using an element named "foreign" with a single mandatory attribute "lang" that indicates a three letter ISO 639-3 language code. E.g. --> w1 w2 - Mark initialisms and spoken words using '~'; e.g. C E O --> ~CEO CD ROM --> ~CD ROM - Removed redundant whitespaces, e.g., w1 w2 -> w1 w2 - Miscellaneous typo fixes. 7. Metadata =========== 7.1 calldata.tbl ---------------- This is a tab-delimited file containing metadata for all calls. Please see the document "docs/doc_calldata.txt" for details. 7.2 speakerdata.tbl ------------------- This is a tab-delimited file containing metadata for all speakers in the corpus. Please see the document "docs/doc_speakerdata.txt" for details. 7.3 pindata.tbl --------------- This is a tab-delimited file containing metadata for all participants who initiated a call. Please see the document "docs/doc_pindata.txt" for details. 7.4 file.tbl ------------ Expected sizes, modification times, and MD5 checksums for all files within the "data/" directory are recorded in "docs/file.tbl". This is a tab-delimited table containing one file per line, each line having the following 4 fields: - checksum -- MD5 checksum of file - size -- size of file in bytes - datetime -- last modification date in YYYY-MM-DD_HH:MM:SS format - path -- path to file relative to root of release directory 8. Contacts =========== If you have questions about this data release, please contact the following LDC personnel: Neville Ryant