MATERIAL Lithuanian-English Language Pack
| Item Name: | MATERIAL Lithuanian-English Language Pack |
| Author(s): | Nyka Aukstuolis, Sigute Bertulyte-Harton, Jolanta Blazaite, Brenda Bujanauskaite, Sarra Chouder, Nathaniel Clair, Tom Conners, Cassian Corey, Eyal Dubinski, Corinna Ellis, Jess Fernando, Paul Gibby, Simon Hammond, Guia Hidalgo, Maxime Hubert, Dagmara Kalnins, Valentina Krickus, Julie Lam, Rosie Lazar, Hanh Le, Nicolas Malyska, Olivia Medel, Jennifer Melot, Alyssa Mensch, Michelle Morrison |
| LDC Catalog No.: | LDC2026S12 |
| ISLRN: | 167-112-192-523-7 |
| DOI: | https://doi.org/10.35111/j3ce-g779 |
| Release Date: | September 15, 2026 |
| Member Year(s): | 2026 |
| DCMI Type(s): | Sound, Text |
| Sample Type: | alaw |
| Sample Rate: | 8000 |
| Data Source(s): | telephone conversations |
| Application(s): | information retrieval, speech recognition |
| Language(s): | Lithuanian, English |
| Language ID(s): | lit, eng |
| License(s): |
MATERIAL Lithuanian-English Agreement (For-Profit) MATERIAL Lithuanian-English Agreement (Non-Member) MATERIAL Lithuanian-English Agreement (Not-For-Profit) |
| Online Documentation: | LDC2026S12 Documents |
| Licensing Instructions: | Subscription & Standard Members, and Non-Members |
| Citation: | Aukstuolis, Nyka, et al. MATERIAL Lithuanian-English Language Pack LDC2026S12. Web Download. Philadelphia: Linguistic Data Consortium, 2026. |
| Related Works: | View |
Introduction
MATERIAL Lithuanian-English Language Pack, Linguistic Data Consortium Catalog Number LDC2026S12, was developed by Appen for the IARPA (Intelligence Advanced Research Projects Activity) MATERIAL (Machine Translation for English Retrieval of Information in Any Language) program. It contains approximately 64 hours of Lithuanian conversational telephone speech, transcripts, English translations, annotations and queries.
The MATERIAL program focused on underserved languages with the ultimate goal to build cross language information retrieval systems to find speech and text content using English search queries.
Data
The Lithuanian speech in this release represents both Aukštaitian and Samogitian regional dialects spoken in Lithuania. The gender distribution among speakers is approximately equal; speakers' ages range from 16 years to 70 years. Calls were made using different telephones (e.g., mobile, landline) from a variety of environments including the street, a home or office, a public place, and inside a vehicle.
Transcripts cover 100% of the speech data, and approximately 6% of the speech data was translated into English. Further information about transcription and translation methodologies is contained in the documentation accompanying this release.
Lithuanian-English Language Pack also includes domain annotations, English queries and their relevance annotations. Annotators marked transcripts by domain (e.g., lifestyle, business-and-commerce, sports, education and so on), by query (simple, conceptual, hybrid) and by their relevance to query search terms.
Speech data is presented either as two channel wav or single channel sphere files, both in A-law format, and predominately as 8kHz. All text data is UTF-8 encoded.
Samples
Please view these samples:
Updates
Additional information, updates, bug fixes may be available in the LDC catalog entry for this corpus at LDC2026S12.