README FILE FOR LDC CATALOG ID: LDC2026T08
TITLE: LORELEI Arabic Representative Language Pack
AUTHORS: Jennifer Tracey, Stephanie Strassel, Dave Graff, Jonathan
Wright, Song Chen, Neville Ryant, Seth Kulick, Kira Griffitt,
Dana Delgado, Michael Arrigo
1.0 Introduction
This corpus was developed by the Linguistic Data Consortium for the DARPA
LORELEI Program and consists of nearly 2.4 million words of monolingual text
in Arabic, over 930,000 words of which have been translated into English. It
also includes about 225,000 Arabic words translated from English text. Nearly
88,000 words are annotated for simple named entities, and about 29,000 words
are annotated for full entity (including nominals and pronouns) -- some data
files received both types of entity annotation; subsets of files in those two
sets also received entity linking, simple semantic annotation, and situation
frame annotation. Details about the volume of data for each annotation type
are listed in section 3.3 below.
The LORELEI (Low Resource Languages for Emergent Incidents) Program is
concerned with building Human Language Technology for low resource languages
in the context of emergent situations like natural disasters or disease
outbreaks. Linguistic resources for LORELEI include Representative Language
Packs for over 2 dozen low resource languages, comprising data, annotations,
basic natural language processing tools, lexicons and grammatical resources.
Representative languages are selected to provide broad typological coverage,
while Incident Languages are selected to evaluate system performance on a
language whose identity is disclosed at the start of the evaluation, and for
which no training data has been provided.
This corpus provides the complete set of monolingual and parallel text,
lexicon, annotations, and tools comprising the LORELEI Arabic Representative
Language Pack.
For more information about LORELEI language resources, see:
https://www.ldc.upenn.edu/sites/www.ldc.upenn.edu/files/lrec2020-lorelei-language-packs.pdf
2.0 Corpus organization
2.1 Directory Structure
The directory structure and contents of the package are summarized below --
paths shown are relative to the base (root) directory of the package:
./README.txt -- this file
./dtds/
./dtds/cstrans_tab.v1.0.dtd
./dtds/laf.v1.2.dtd
./dtds/llf.v1.6.dtd
./dtds/ltf.v1.5.dtd
./dtds/psm.v1.0.dtd
./docs/ -- various tables and listings (see section 9 below)
./docs/annotation_guidelines/ -- guidelines for all annotation tasks included in this corpus
./docs/ara_fe_docids.untagged_headline.txt --\
./docs/ara_sne_docids.untagged_headline.txt -- see section 10.1.2 below
./docs/char_tally.ARA.tab -- see section 9.0 below
./docs/cstrans_tab/ -- see section 7.7 below
./docs/grammatical_sketch/ -- grammatical sketch of Arabic
./docs/source_codes.tab -- list of data sources
./docs/twitter_info.tab -- list of audited Tweets in Arabic
./docs/urls.tab -- list of documents with source urls
./tools/ -- see section 8 below for details about tools provided
./tools/ldclib/
./tools/ltf2txt/
./tools/sent_seg/
./tools/ara/ne_tagger/
./tools/ara/transliterator/
./tools/tokenization_parameters.v5.0.yaml
./data/monolingual_text/zipped/ -- zip-archive files containing "ltf" and "psm" data
./data/translation/
from_ara/{ara,eng}/ -- translations from Arabic to English
from_eng/ -- translations from English to Arabic
{elicitation,news,phrasebook}/ for each of three types of English data:
{ara,eng}/ for each language in each directory,
"ltf" and "psm" directories contain
corresponding data files
./data/annotation/ -- see section 5 below for details about annotation
./data/annotation/entity/{simple,full}/
./data/annotation/np_chunking/
./data/annotation/sem_annotation/
./data/annotation/situation_frame/{issues,mentions,needs}/
./data/annotation/twitter_tokenization/
./data/lexicon/
2.2 File Name Conventions
The file names assigned to individual documents in this corpus provide the
following information about the document:
Language 3-letter abbrev.
Genre 2-letter abbrev.
Source 6-digit numeric ID assigned to data provider
Date 8-digit numeric: YYYYMMDD year, month, day)
Global-ID 9-digit alphanumeric assigned to this document
Those five fields are joined by underscore characters, yielding a 32-character
file-ID; three portions of the document file-ID are used to set the name of
the zip file that holds the document: the Language and Genre fields, and the
first 6 digits of the Global-ID.
The 2-letter codes used for genre are as follows:
DF -- discussion forum
NW -- news
RF -- reference (e.g. Wikipedia)
SN -- social network (Twitter)
WL -- web-log
3.0 Content Summary
3.1 Monolingual Text
Genre #Docs #Words
DF 1338 445084
NW 4273 1533777
WL 3063 417952
SN 6290 94807
Note that the SN (Twitter) data cannot be distributed directly by LDC, due to
the Twitter Terms of Use. The file "docs/twitter_info.tab" (described in
Section 8.2 below) provides the necessary information for users to fetch the
particular tweets directly from Twitter. LTF files for all other genres are
stored in ./data/monolingual_text/zipped/.
3.2 Parallel Text
Type Genre #Docs #Words
---
FromEng EL 2 15174
FromEng NW 190 68465
---
ToEng NW 1486 519243
ToEng WL 3063 417952
---
Note that most of the translation from English into Arabic was done via
crowd-sourcing (except for the "elicitation" file and a handful of news
stories); for each crowd-sourced document, at least three and up to six
translation versions are provided. The "FromEng" word counts are based
on the first crowd-sourced version of each document.
3.3 Annotation
AnnotType Genre #Docs #Words
---
EntityFull NW 58 12664
EntityFull SN 153 2575
EntityFull WL 38 13784
---
EntitySimp NW 150 31456
EntitySimp SN 500 8558
EntitySimp WL 136 47795
---
NPChunking NW 32 5149
NPChunking SN 58 1056
NPChunking WL 20 4484
---
SimpleSemantic NW 48 9985
SimpleSemantic SN 74 1430
SimpleSemantic WL 21 6699
---
SituationFrame NW 38 10111
SituationFrame SN 128 2424
SituationFrame WL 28 12002
---
4.0 Data Collection and Parallel Text Creation
Both monolingual text collection and parallel text creation involve a
combination of manual and automatic methods. These methods are described in
the sections below.
4.1 Monolingual Text Collection
Data is identified for collection by native speaker "data scouts," who search
the web for suitable sources, designating individual documents that are in the
target language and discuss the topics of interest to the LORELEI program
(humanitarian aid and disaster relief). Each document selected for inclusion
in the corpus is then harvested, along with the entire website when suitable.
Thus the monolingual text collection contains some documents which have been
manually selected and/or reviewed and many others which have been
automatically harvested and were not subject to manual review.
4.2 Parallel Text Creation
Parallel text for LORELEI was created using three different methods, and each
LORELEI language may have parallel text from one or all of these methods. In
addition to translation from each of the LORELEI languages to English, each
language pack contains a "core" set of English documents that were translated
into each of the LORELEI Representative Languages. These documents consist of
news documents, a phrasebook of conversational sentences, and an elicitation
corpus of sentences designed to elicit a variety of grammatical structures.
All translations are aligned at the sentence level. For professional and
crowd-sourced translation, the segments align one-to-one between the source and
target language (i.e. segment 1 in the English aligns with segment 1 in the
source language).
Professionally translated data has one translation for each source document,
while crowd-sourced translations have up to six translations for each source
document, designated by A, B, C, D, E or F appended to the file name on the
multiple translation versions.
5.0 Annotation
Seven types of annotation are present in this corpus:
- Simple Named Entity tags names of persons, organizations, geopolitical
entities, and locations (including facilities).
- Full Entity also tags nominal and pronominal mentions of entities.
- Entity Discovery and Linking provides cross-document coreference of named
entities via linking to an external knowledge base (the knowledge base used
for LORELEI is released separately as LDC2020T10).
- Noun Phrase Chunking identifies the positions and extents of noun phrases.
- Simple Semantic Annotation provides light semantic role labeling, capturing
acts and states along with their arguments.
- Situation Frame annotation labels the presence of needs and issues related
to emergent incidents such as natural disasters (e.g. food need, civil
unrest), along with information such as location, urgency, and entities
involved in resolving the needs.
Details about each of these annotation tasks can be found in
docs/annotation_guidelines/.
SPECIAL NOTE ABOUT ANNOTATIONS ON TWITTER DATA:
The LDC cannot redistribute text data from Twitter, and this includes files
containing annotation. Where LAF XML and annotation table files have strings
of text from other sources, annotations of Twitter data instead have strings
with underscores ("_") replacing all non-white-space characters.
Software is included in this release that enables users to download a given
list of Tweets (assuming the Tweets are still available online), and apply the
same conditioning and reformatting that was done by LDC prior to annotation --
see section 8.2 below (ldclib) for more details on the software.
In order to confirm that your own download and conditioning yields results
that match those of the LDC, we provide a set of LTF XML files (one for each
annotated Tweet), in which the text content has been modified by replacing
each non-white-space character with an underscore ("_"), so that character
offsets are preserved for word tokens and spans of annotations.
These "placeholder" LTF XML files are in data/annotation/twitter_tokenization/.
6.0 Data Processing and Character Normalization for LORELEI
Most of the content has been harvested from various web sources using an
automated system that is driven by manual scouting for relevant material.
Some content may have been harvested manually, or by means of ad-hoc scripted
methods for sources with unusual attributes.
All harvested content was initially converted from its original HTML form
into a relatively uniform XML format; this stage of conversion eliminated
irrelevant content (menus, ads, headers, footers, etc.), and placed the
content of interest into a simplified, consistent markup structure.
The "homogenized" XML format then served as input for the creation of a
reference "raw source data" (rsd) plain text form of the web page content; at
this stage, the text was also conditioned to normalize white-space characters,
and to apply transliteration and/or other character normalization, as
appropriate to the given language.
7.0 Overview of XML and Tabular Data Structures
7.1 PSM.xml -- Primary Source Markup Data
The "homogenized" XML format described above preserves the minimum set of tags
needed to represent the structure of the relevant text as seen by the human
web-page reader. When the text content of the XML file is extracted to create
the "rsd" format (which contains no markup at all), the markup structure is
preserved in a separate "primary source markup" (psm.xml) file, which
enumerates the structural tags in a uniform way, and indicates, by means of
character offsets into the rsd.txt file, the spans of text contained within
each structural markup element.
For example, in a discussion-forum or web-log page, there would be a division
of content into the discrete "posts" that make up the given thread, along with
"quote" regions and paragraph breaks within each post. After the HTML has
been reduced to uniform XML, and the tags and text of the latter format have
been separated, information about each structural tag is kept in a psm.xml
file, preserving the type of each relevant structural element, along with its
essential attributes ("post_author", "date_time", etc.), and the character
offsets of the text span comprising its content in the corresponding rsd.txt
file.
7.2 LTF.xml -- Logical Text Format Data
The "ltf.xml" data format is derived from rsd.txt, and contains a fully
segmented and tokenized version of the text content for a given web page.
Segments (sentences) and the tokens (words) are marked off by XML tags (SEG
and TOKEN), with "id" attributes (which are only unique within a given XML
file) and character offset attributes relative to the corresponding rsd.txt
file; TOKEN tags have additional attributes to describe the nature of the
given word token.
The segmentation is intended to partition each text file at sentence
boundaries, to the extent that these boundaries are marked explicitly by
suitable punctuation in the original source data. To the extent that sentence
boundaries cannot be accurately detected (due to variability or ambiguity in
the source data), the segmentation process will tend to err more often on the
side of missing actual sentence boundaries, and (we hope) less often on the
side of asserting false sentence breaks.
The tokenization is intended to separate punctuation content from word
content, and to segregate special categories of "words" that play particular
roles in web-based text (e.g. URLs, email addresses and hashtags). To the
extent that word boundaries are not explicitly marked in the source text, the
LTF tokenization is intended to divide the raw-text character stream into
units that correspond to "words" in the linguistic sense (i.e. basic units of
lexical meaning).
Software is included to convert ltf.xml files to "raw source data" plain text
files ("rsd.txt") -- see section 8.1 below. The character offsets used in LTF
and LAF xml, and in other types of annotation data, are based on the "rsd.txt"
files, which contain just the text that is visible to a person reading the
original source, with normalized white-space characters (including line
breaks), but without markup of any kind.
7.3 LAF.xml -- Logical Annotation Format Data
The "laf.xml" data format provides a generic structure for presenting
annotations on the text content of a given ltf.xml file; see the associated
DTD file in the "dtds" directory. Note that each type of annotation (simple
named entity, full entity, simple semantic annotation) uses the basic XML
elements of LAF in different ways.
7.4 LLF.xml -- LORELEI Lexicon Format Data
The "llf.xml" data format is a simple structure for presenting citation-form
words (headwords or lemmas) in Arabic, together with Part-Of-Speech (POS)
labels and English glosses. Each ENTRY element contains a unique combination
of LEMMA value (citation form in native orthography) and POS value, together
with one or more GLOSS elements. Each ENTRY has a unique ID, which is
included as part of the unique ID assigned to each GLOSS.
7.5 Situation Frame Annotation Tables
Situation frame annotation consists of three parts, each presented as a
separate tab-delimited file: entities, needs, and issues. The details of each
table are described below.
Entities, mentions, need frames, and issue frames all have IDs that follow a
standard schema consisting of a prefix designating the type of ID ('Ent' for
entities, 'Men' for mentions, and 'Frame' for both need and issue frames), an
alphanumeric string identifying the annotation "kit", and a numeric string
uniquely identifying the specific entity, mention, or frame within the
document.
7.5.1 Mentions
The grouping of entity mentions into "selectable entities" for situation frame
annotation is provided in the mentions/ subdirectory. The table has 8 columns
with the following headers and descriptions:
column 1: doc_id -- doc ID of source file for the annotation
column 2: entity_id -- unique identifier for each grouped entity
column 3: mention_id -- unique identifier for each entity mention
column 4: entity_type -- one of PER, ORG, GPE, LOC
column 5: mention_status -- 'representative' or 'extra';
representative mentions are the ones which have been chosen by the
annotator as the representative name for that entity. Each entity
has exactly one representative mention.
column 6: start_char -- character offset for the start of the mention
column 7: end_char -- character offset for the end of the mention
column 8: mention_text -- mention string
7.5.2 Needs
Annotation of need frames is provided in the needs/ subdirectory. Each row in
the table represents a need frame in the annotated document. The table has 13
columns with the following headers and descriptions:
column 1: user_id -- user ID of the annotator
column 2: doc_id -- doc ID of source file for the annotation
column 3: frame_id -- unique identifier for each frame
column 4: frame_type -- 'need'
column 5: need_type -- exactly one of 'evac' (evacuation), 'food' (food
supply), 'search' (search/rescue), 'utils' (utilities, energy, or
sanitation), 'infra' (infrastructure), 'med' (medical assistance),
'shelter' (shelter), or 'water' (water supply)
column 6: place_id -- entity ID of the LOC or GPE entity identified as the
place associated with the need frame; only one place value per
need frame, must match one of the entity IDs in the corresponding
ent_output.tsv or be 'none' (indicating no place was named)
column 7: proxy_status -- 'True' or 'False'
column 8: need_status -- 'current', 'future'(future only), or 'past' (past only)
column 9: urgency_status -- 'True' (urgent) or 'False' (not urgent)
column 10: resolution_status -- 'sufficient' or 'insufficient' (insufficient /
unknown sufficiency)
column 11: reported_by -- entity ID of one or more entities reporting
the need; multiple values are comma-separated, must match entity IDs
in the corresponding ent_output.tsv or be 'none'
column 12: resolved_by -- entity ID of one or more entities resolving
the need; multiple values are comma-separated, must match entity IDs
in the corresponding ent_output.tsv or be 'none'
column 13: description -- string of text entered by the annotator as
memory aid during annotation, no requirements for content or language,
may be 'none'
7.5.3 Issues
Annotation of issue frames is provided in the issues/ subdirectory. Each row
in the table represents an issue frame in the annotated document. The table has
9 columns with the following headers and descriptions:
column 1: user_id -- user ID of the annotator
column 2: doc_id -- doc ID of source file for the annotation
column 3: frame_id -- unique identifier for each frame
column 4: frame_type -- 'issue'
column 5: issue_type -- exactly one of 'regimechange' (regime change),
'crimeviolence' (civil unrest or widespread crime), or 'terrorism'
(terrorism or other extreme violence)
column 6: place_id -- entity ID of the LOC or GPE entity identified as
the place associated with the issue frame; only one place value per
issue frame, must match one of the entity IDs in the corresponding
ent_output.tsv or be 'none'
column 7: proxy_status -- 'True' or 'False'
column 8: issue_status -- 'current' or 'not_current'
column 9: description -- string of text entered by the annotator as
memory aid during annotation, no requirements for content or
language, may be 'none'
7.6 EDL Table
The "data/annotation/entity/" directory contains the file "ara_edl.tab", which
has an initial "header" line of column names followed by data rows with 8
columns per row. The following shows the column headings and a sample value
for each column:
column 1: system_run_id LDC
column 2: mention_id
column 3: mention_text
column 4: extents
column 5: kb_id
column 6: entity_type
column 7: mention_type
column 8: confidence
When column 5 is fully numeric, it refers to a numbered entity in the
Reference Knowledge Base (distributed separately as LDC2020T10). Note that a
given mention may be ambiguous as to the particular KB element it represents;
in this case, two or more numeric KB_ID values will appear in column 5,
separated by the vertical-bar character (|).
When column 5 consists of "NIL" plus digits, it refers to an entity that is
not present in the Knowledge Base, but this label is used consistently for all
mentions of the particular entity.
7.7 CSTRANS_TAB.xml -- Crowd-source Translation Tables
The "./docs/cstrans_tab/" directory contains one "*.cstrans_tab.xml" file
for each English source file that was submitted to translation via crowd
sourcing. Each file contains a DOC element (with "id" and "lang" attributes),
which in turn contains a "SEG" element for each "SEG" in the corresponding
English ltf.xml file. Each "SEG" element may either be an empty tag (if no
usable translations were submitted for the given segment), or contain one or
more "TR" elements, each of which is an alternative translation for the given
source segment. In either case, the "SEG" tag has an "id" attribute (unique
within the given xml file, matching the SEG "id" value in ltf.xml), and an
"ntrs" attribute (whose value is the number of "TR" elements present. For
example:
...
...
The attributes of the "TR" elements are as follows:
- translatorid -- an alphanumeric string unique to each contributor; note
that each translation "version" (_A, _B, etc) is likely to contain segments
from different translators
- avg_gold_ter may be floating-point numeric or "Unk"; it represents the
"term error rate" relative to a "gold-standard" manual translation (lower
value == better match)
- score may be floating-point numeric or "None"
- mt_ter is always floating-point numeric; it represents the "machine
translation error rate" relative to a "google-translate" reference (lower
value == better match)
- nonwhitesp and odd_ch are always integer numerics: the count of
non-whitespace characters in the string, and the count of characters that
are "not in the expected language" (this can include emoticons,
non-printing characters, and characters in foreign scripts).
8.0 Software tools included in this release
8.1 "ltf2txt" (source code written in Perl)
A data file in ltf.xml format (as described above) can be conditioned to
recreate exactly the "raw source data" text stream (the rsd.txt file) from
which the LTF was created. The tools described here can be used to apply that
conditioning, either to a directory or to a zip archive file containing
ltf.xml data. In either case, the scripts validate each output rsd.txt stream
by comparing its MD5 checksum against the reference MD5 checksum of the
original rsd.txt file from which the LTF was created. (This reference
checksum is stored as an attribute of the "DOC" element in the ltf.xml
structure; there is also an attribute that stores the character count of the
original rsd.txt file.)
Each script contains user documentation as part of the script content; you can
run "perldoc" to view the documentation as a typical UNIX man page, or you can
simply view the script content directly by whatever means to read the
documentation. Also, running either script without any command-line arguments
will cause it to display a one-line synopsis of its usage, and then exit.
ltf2rsd.perl -- convert ltf.xml files to rsd.txt (raw-source-data)
ltfzip2rsd.perl -- extract and convert ltf.xml files from zip archives
Special note about Twitter data: as explained in section 5 above, this corpus
includes "scrubbed" versions of LTF XML files for individual Tweets, where the
original text characters (except for spaces) are replaced by underscores (in
data/annotation/twitter_tokenization/), in order to comply with Twitter Terms
of Use. Running "ltf2rsd.perl" directly on these "scrubbed" files will yield
warnings about MD5 mismatches, which is to be expected, because the MD5 value
stored in each Twitter LTF XML file is based on the original text. After
using the "ldclib" software (described in the next section) to download and
condition Twitter data, the resulting LTF XML files should have both the
original text and the matching MD5 values; that process also creates the
corresponding rsd.txt files.
8.2 ldclib -- general text conditioning, twitter harvesting
The "bin/" subdirectory of this package contains three executable scripts
(written in Ruby):
create_rsd.rb -- convert general xml or plain-text formats to "raw source
data" (rsd.txt), by removing markup tags and applying
sentence segmentation
token_parse.rb -- convert rsd.txt format into ltf.xml
get_tweet_by_id.rb -- download and condition Twitter data
Due to the Twitter Terms of Use, the text content of individual tweets cannot
be redistributed by the LDC. As a result, users must download the tweet
contents directly from Twitter. The twitter-processing software provided in
the tools/ directory enables users to perform the same normalization applied
by LDC and ensure that the user's version of the tweet matches the version
used by LDC, by verifying that the md5sum of the user-downloaded and processed
tweet matches the md5sum provided in the twitter_info.tab file. Users must
have a developer account with Twitter in order to download tweets, and the
tool does not replace or circumvent the Twitter API for downloading tweets.
The ./docs/twitter_info.tab file provides the twitter download id for each
tweet, along with the LORELEI file name assigned to that tweet and the md5sum
of the processed text from the tweet.
The file "README.md" in this directory provides details on how to install and
use the source code in this directory in order to condition text data that the
user downloads directly from Twitter and produce both the normalized raw text
and the segmented, tokenized LTF.xml output.
All LDC-developed supporting files (models, configuration files, library
modules, etc.) are included, either in the "lib" subdirectory (next to "bin"),
or else in the parent ("tools") directory.
Please refer to the README.md file that accompanies this software package.
8.3 sent_seg -- apply sentence segmentation to raw text
The Python tools in this directory are used as part of the conditioning done
by "create_rsd.rb" in the "ldclib" package. Please refer to the README.rst
file included with the package.
8.4 ne_tagger -- Named-Entity tagger for Arabic
Please refer to the tools/ara/ne_tagger/README.rst file for information about
usage and performance.
9.0 Documentation included in this release
The ./docs folder (relative to the root directory of this release) contains
four files documenting various characteristics of the source data:
char_tally.ARA.tab - contains four tab separated columns: doc uid, number
of non-whitespace characters, number of non-whitespace characters in
the expected script, and number of anomalous (non-printing) characters
for each document in the release
source_codes.tab - contains tab-separated columns: genre, source code, source
name, and base url for each source in the release
twitter_info.tab - contains tab-separated columns: doc uid, tweet id,
normalized md5 of the tweet text, and tweet author id for all tweets in the
release
urls.tab - contains two tab-separated columns: doc uid and url. Note that the
url column is empty for documents from previous LDC publications, for which no
url is available; they are included here so that the uid column can serve as a
complete document list for the package.
In addition, the grammatical sketch and annotation guidelines contents
described in earlier sections of this README are found in this directory.
10.0 Known Issues for Arabic
The following sections describe attributes of the data that might pose minor
problems for some users.
10.1 Use of legacy corpora for entity annotations leaves some gaps
Both Full and Simple-Named Entity annotations in this release have been
adapted from previous LDC corpora, which originally used specifications that
were developed for the DARPA "ACE" (Automatic Content Extraction) program.
Those specifications differ from current LORELEI entity annotations in the
following ways:
10.1.1 "TTL" ("title") mentions are not marked in Full Entity annotation
No attempt has been made (as of this release) to assess the quantity or
distribution of strings that would have been labeled "TTL" if the data
had undergone the same Full-Entity annotation applied to other LORELEI
languages.
10.1.2 The initial "HEADLINE" segment of documents was not annotated.
As the original ACE annotation files were reconditioned into laf.xml format
(for both FE and SNE), it became apparent that the initial "headline" segment
of many NW files contained one or more strings that matched annotated mentions
later in the text, but in ACE annotation, strings identified as "headlines" in
the original source document were intentionally excluded from annotation.
Therefore, when adjusting character offsets of mentions from the original ACE
files into laf.xml (to mark the correct spans of LORELEI rsd.txt files),
string matches in the initial (headline) segment were ignored.
The "docs/" directory contains these two files:
ara_fe_docids.untagged_headline.txt (44 lines)
ara_sne_docids.untagged_headline.txt (101 lines)
Each list cites the file-IDs where the initial SEGMENT element in ltf.xml
format contains one or more strings that match annotated mentions appearing
later in the text; in other words, if these files had been annotated under
LORELEI guidelines, it's likely mentions in the initial (headline) segment
would have been marked. For NW files not cited in each list, it's possible
(but perhaps unlikely) that mentions appear in the initial segment, even
though it contains nothing that matches a marked mention later in the text.
10.2 Multi-escaped characters in some translation files
In data/translation/from_ara/, there are 43 eng/ltf files and one ara/ltf file
that contain a total of 113 tokens like """, "‏" and
so on; when these files are converted to rsd.txt, the resulting raw text will
still contain strings of the form "&", "‏", etc.
10.3 Some presentation-form Arabic characters
There is one newswire file that contains a total of 14 characters in the
"presentation form" section of the Unicode Arabic character set:
data/translation/from_ara/ara/ltf/ARA_NW_002987_20070110_G0023FYON.ltf.xml
Because data were imported from previous LDC corpus publications (to take
advantage of existing translations and annotations), this particular file
missed the normalization step that would have made the following character
replacements throughout the text:
U+FE87 -> U+0625 (1-to-1 replacement, 1 token affected)
U+FEF9 -> U+0644 U+0625 (1-to-2 replacement, 6 tokens affected)
11.0 Acknowledgments
The authors would like to acknowledge the following contributors to this
corpus: Brian Gainor, Ann Bies, Justin Mott, Neil Kuster, University of
Maryland Applied Research Laboratory for Intelligence and Security (ARLIS),
formerly UMD Center for Advanced Study of Language (CASL), and our team of
Arabic annotators.
This material is based upon work supported by the Defense Advanced Research
Projects Agency (DARPA) under Contract No. HR0011-15-C-0123. Any opinions,
findings and conclusions or recommendations expressed in this material are
those of the author(s) and do not necessarily reflect the views of DARPA.
12.0 Copyright
Portions © 2000, 2002-2010 Agence France Presse, © 2006-2008, 2010
Al-Ahram, © 2000, 2006-2008, 2010 Al Hayat, © 2006-2008, 2010 Al-Quds
Al-Arabi, © 2000 American Broadcasting Company, © 2000, 2002,
2006-2008, 2010 An Nahar, © 2006-2008, 2010 Asharq Al-Awsat, © 2007,
2010 Assabah, © 2000 Cable News Network LP, LLLP, © 2008 Central News
Agency (Taiwan), © 1989 Dow Jones & Company, Inc., © 2005 Los Angeles
Times - Washington Post News Service, Inc., © 2000 National
Broadcasting Company, Inc., © 1999, 2005, 2006, 2010 New York Times, ©
2000 Public Radio International, © 2003, 2005-2008, 2010 The
Associated Press, © 2010 Ummah Press, © 2003, 2005-2008, 2010 Xinhua
News Agency, © 2016, 2022 Trustees of the University of Pennsylvania
13.0 CONTACTS
Neville Ryant