CALLHOME Mandarin Chinese Lexicon Second Edition June 24, 2025 Linguistic Data Consortium 1. Overview =========== This is an updated release of the CALLHOME Mandarin Chinese Lexicon (LDC96L15). The original CALLHOME lexicon was compiled by the Linguistic Data Consortium in support of the project on Large Vocabulary Conversational Speech Recognition (LVCSR), sponsored by the U.S. Department of Defense. This re-release retains the same 44,404 words and associated information (morphological, phonological, lexical frequencies, etc) from the original release. However, the directory structure, file formats, and documentation have been updated to modern standards. 2. Directory structure ====================== - data/lexicon.tsv -- lexicon in TSV format - data/lexicon.dict -- a pronunciation dictionary derived from lexicon; in CMUdict format - docs/file.tbl -- listing of md5 checksums, sizes, dates, and file names - docs/README.txt -- this file; a top-level documentation of release - docs/lex_segm.txt -- documents the word segentation principles used during transcription and preparation of the lexicon - docs/pron.txt -- documents the conventions used by the pinyin and pronunciation fields - docs/pos_tags.txt -- a table of part-of-speech tags used in the lexicon 3. lexicon.tsv ============== The lexicon is stored as a UTF-8 encoded tab-delimited file containing one entry per line, each line having the following seven fields: - headword -- orthographic form in hanzi (e.g., 没有) - pinyin -- headword in tone pinyin (e.g., mei2 you3) - tone -- tone sequence for headword (e.g., 2 3) - pron -- pronunciation of headword WITHOUT tone information (e.g., mey yow) - pos -- part-of-speech tag (e.g., phrase) - xinhua_freq -- frequency of the headword in Xinhua newswire - train_freq -- frequency of the headword in the 80 training transcripts from CALLHOME Mandarin Chinese Omnibus corpus Each of these fields is described in more detail in the sections below. 3.1 Field 1: headword --------------------- This column contains orthographic Chinese characters that are simplified in the Mainland style. Headwords were automatically segmented by the Dragon Mandarin segmenter in accordance with the principles listed in "docs/lex_segm.txt". 3.2 Field 2: pinyin ------------------- This is a representation of the headword in tone Pinyin with strictly lexical tone, i.e. not reflecting phonetic/phonological processes. A listing of the initials and finals can be found in "docs/pron.txt". 3.3 Field 3: tone ----------------- This field provides the tone sequence for the word AFTER after accounting for tone sandhi. E.g., the sequence of lexical tones "3 3" becomes "2 3". For disyllabiv words, the application of tone sandhi was handled automatically, while for longer words (>= 3 syllables) the application of tone sandhi was determined manually based on words structure. For words that have sequences of tone 3 and more than 2 syllables, all but the rightmost tone may change to 2, e.g., 2 2 2 3, in fast speech; but in careful speech, the application of tone sandhi depends on syntactic constituency. The careful speech pronunciation is the one shown. 3.4 Field 4: pron ----------------- This field contains the word's pronunciation MINUS tone. Syllables are separated by spaces (as in pinyin) and use the phoneset from the "Allophone" column of the table presented in "docs/pron.txt". 3.5 Field 5: pos ---------------- This field provies the part-of-speech tag for the headword. The full tagset is listed in "docs/pos_tags.txt". 3.6 Field 6: xinhua_freq ------------------------ This field provides the frequency of the headword in the 3,431,707 words of Xinhua newswire. 3.7 Field 7: train_freq ----------------------- This field provides the frequency of the headword in the 155,276 words of the 80 CALLHOME Mandarin Chinese Omnibus training transcripts. The frequency counts provided in fields 6 and 7 are based on the orthographic form of the headword (field 1), not pronunciation. In cases where a single character, or character sequence, has multiple pronunciations, the same frequency count is given for every entry. For example: - the same character represents 5 distinct pronunciations of "a" - for each of those 5 separate entries, the frequencies 31 and 3705 are given In fact, the numbers representing occurrences are distributed over the 5 different pronunciations of the characters in some way that is not identified. 4. lexicon.dict =============== This is a CMUdict format version of "lexicon.tsv". It consists of one pronunciation per line, each line having the form: \t where: - WORD -- orthographic representation of word (in characters) - PRON -- a single pronunciation of the word, expressed as a sequence of phone symbols; in this case, pinyin I.e., it is a mapping from character sequences (field 1 of "lexicon.tsv") to their reference pronunciations, expressed in pinyin (with lexical tone). 5. Contacts =========== If you have questions about this data release, please contact the following LDC personnel: Neville Ryant