Showing posts with label Swahili. Show all posts
Showing posts with label Swahili. Show all posts

Thursday, February 09, 2017

Health info in African languages, on 2 non-African sites

Here are quick reviews of two websites - one Australian the other American - that have health information in numerous world languages, including a number from Africa. Both are primarily intended to serve immigrant communities. This post will then return briefly to the theme of the benefits of systematically sharing and improving of health related information composed in or translated into African languages.

Health Translations


Health Translations is a website maintained by the government of the state of Victoria in Australia.  It has information (mainly documents such as fact sheets and flyers, from what I can tell, some with illustrations) on over 80 topics in a total of almost 100 languages or language varieties (although not all information is in every language, and some languages have few items).

The African languages for which there are materials include: Afrikaans; Akan; Amharic; Arabic; Bemba; Dinka; Juba Arabic; Kirundi; Krio; Lingala; Nuer; Oromo; Shona; Somali; Sudanese Arabic (also listed as Sudanese); Swahili (Congolese); Swahili (Kenyan); and Tigrinya.

This is an impressive collection of materials from various sources, apparently all Australian, and in different formats. They appear to all be translations - a given material may be available in a few or quite a number of language versions. Navigating from the a click on the desired language on the list of languages (which helpfully includes both English & native names/scripts) to a particular topical resource requires interacting with screens in English - not surprising, but when one gets to the list of topics and resources in a particular language, the titles are only in English, and then on the list of languages in which a material is translated (this is a typical navigation sequence), the language names are in English (no native scripts used). So the resource appears intended to be used by or with help of professionals or others who can read English.

Source: Bushfire smoke & your health [am]
I'm not able to evaluate the quality of the translations, but noted French in Lingala (which may simply be typically used loanwords) and by chance an anomalous English word in an Amharic text (image).

All the several documents I viewed were PDFs, mainly text but some image (meaning the text cannot be searched or copied out for editing into other materials). Spot checking some non-Latin text, specifically Ethiopic/Ge'ez used for Amharic and Tigrinya, and complex Latin, specifically for Dinka, there were some issues with the text that would interfere with searches or copying out passages (such problems are not uncommon with PDF rendering, even when visually the PDF presents everything correctly and in its intended place).

Some Amharic text when copied out and pasted showed capital A for አ and E for እ in initial position (for example, here). The corresponding characters are appropriate, interestingly, but this makes search or reuse of such text problematic.

From original (l.); copy-pasted out (r.)
Source: Bushfire smoke ... [din]

The Dinka text sampled showed some typical problems with complex Latin in PDFs. Dinka is written with what I call in ALDA a "category 4" Latin orthography in that it includes extended Latin characters (aka modified letters) plus combining diacritics, sometimes together as in the open-o with diaeresis in the word "daiɣɔ̈kthai" (dioxide) featured on the left side of the image. Copying that word from the PDF and pasting it in a word processor or advanced text editor yielded the results on the right, missing one extended character and the combining diacritic on the other. This complicates any potential re-use of this text, but also means that document, folder, or web searches will not pick up words with such character combinations.

Will return to these issues, why they're important, and what to do about them in the last section of this post.

 

HealthReach


HealthReach: Health Information in Many Languages is a program of the US Library of Medicine of the National Institutes of Health. It includes translations in 46 languages maintained on the MedinePlus site. A total of almost 350 topics are listed (although here too, not all information is in every language, and some languages have fewer items than others).

The African languages for which there are materials include: Amharic; Arabic; Oromo; Somali; Swahili; and Tigrinya, The native names of languages are featured on the list of languages, except oddly for Amharic and Tigrinya, which are transliterated into Latin ("amarunya" instead of አማርኛ, and "tigrinya" instead of ትግርኛ).

This also is an impressive collection from diverse sources, in this case American, but it is longer on topics and shorter on languages covered. The list of topics for each language also includes the titles in the language and its script - except again for Amharic and Tigrinya (not even transliterations) - as well as in English.

All materials checked were PDFs. There are no materials for African languages with complex Latin scripts.

As for non-Latin scripts, text in Arabic seems to behave as intended, from small samples. On the other hand, some Amharic text when copied out and pasted showed the same capital A for አ and E for እ observed above, plus O for ኦ (see here).  A Tigrinya document had a similar issue. So this issue may have to do with a problem in PDFs for handling a particular set of characters - አኡኢኣኤእኦኧ (representing glottal stop plus the range of vowels) - or a subset of them, which might be helpful to know when troubleshooting.

Health education materials and the "2Ds & 4Rs"


In highlighting aspects of public health messaging during the ebola epidemic in West Africa (2014-15), this blog suggested a systematic approach to sharing and improving materials that were developed and used in that context (with primary attention to text and images). A mnemonic - 2Ds & 4Rs - was put forth in October 2014, initially to explain the rationale for reposting and discussing various ebola education materials, but also as a way to capture the ideal cycle of utility of such production. Too often, materials are developed, used for a particular purpose, and then forgotten, when they could add to a growing living corpus of resources to tap for future work. This is important in any field and language, but arguably especially important in health, and for languages that have fewer resources and emerging terminologies / technical lexicons, such as many in Africa.

In that context I propose to use the 2Ds & 4Rs to consider the efforts represented by the two sites discussed above. Of the 6 elements of this model, the first three have to do more with the sharing and use of materials, and the last three with their longer term development and potential re-use. These are listed with brief explanations and what I see as relevance to the two sites:
  • Dissemination (making materials available, including via multiple sources)
    • Both sites bring together and post materials from diverse sources, increasing their exposure and access to them.
  • Demonstration (showing how materials in African languages can be presented, including in cases where complex scripts are involved)
    • Both sites show that African language materials can be presented on the same footing as other world languages.
    • However, the HealthReach presentation does not use available technology to present the native names of Amharic and Tigrinya, or titles of materials in those languages.
  • Reading (creating or translating text materials with attention to how they may be read aloud in groups or over local radio, which may be more likely scenarios for their use than the typical Western expectation of silent reading by individuals)
    •  It appears that all or most of the materials from diverse sources compiled on the two sites are translations from English of technical descriptions and advice. It is not clear how well how well adapted they are for the range of uses and audiences they might serve.
  • Review (written material - text - is well suited for review, comparison, and analysis; such material, especially in less resourced languages and on issues of public importance like health, should undergo such treatment)
    • No information on how any of the materials may be or have been reviewed, either in the diverse organizations where they originated, or in the projects hosting the two websites. 
    • Image PDFs, where these occur, do not lend themselves to processes of review.
    • Text PDFs with problems in their encoding of non-Latin or complex Latin scripts, present problems for review.
  • Revision (after review of materials, and in response to other information and feedback relevant to them, materials should undergo appropriate revisions in content, form of language, copyediting, and presentation)
    • No information on any revisions of any of the materials.
    • Issues cited under "Review" with image PDFs and with text PDFs that have encoding problems also hinder revision work.
  • Re-use or re-purposing (text materials can be re-used or sections re-purposed)
    • No information on re-use or re-purposing of any of the materials.

The two sites profiled above and the various health and medical education materials presented on them represent an important resource for fifteen African languages (and some varieties of two of those).

One additional question is whether such materials, intended primarily to serve needs of immigrants in Australia and the US, might be useful as is or with modifications, for speakers of the same languages in relevant African countries. Or in the reverse sense, whether any health extension materials from Africa might inform revision of these materials and development of new ones. A next step could be a for a site to begin to collect health materials in African languages from all sources.

There are many directions in which this could be taken, with the goals of improving availability, quality and utility of health education information in a range of African languages. One, for example, is linking with the longstanding WikiProject Med's Translation Task Force for development of articles in those African languages that have Wikipedias (such as Afrikaans, Akan, Amharic, Arabic, Kirundi, Lingala, Oromo, Shona, Somali, Swahili, and Tigrinya). Another might be connecting with efforts to advance development of standard terminologies. Still another might be to bring in human language technology, such as text to speech, so that materials designed and disseminated in text form could be accessible in audio via mobile devices.

Thanks to Charles Riley of Yale University for calling our attention to these two websites.

Wednesday, December 28, 2016

Mabati-Cornell Kiswahili Prizes 2016

The second annual Mabati-Cornell Kiswahili Prizes for African Literature were awarded earlier this month at Cornell University in Ithaca, New York. As discussed previously on this blog, the Mabati-Cornell is the only literary award going to writers publishing in African languages.

Mabati-Cornell, which was founded in late 2014 by Cornell faculty Dr. Mukoma Wa Ngugi and Caine Prize for African Writing director Dr. Lizzie Attree, recognizes literature in the Swahili language. Its sponsorship by by the Kenyan company Mabati Rolling Mills led Mukoma Wa Ngugi to state that "the prize sets an historical precedent for African philanthropy by Africans and shows that African philanthropy can and should be at the centre of African cultural production."

This year's prizes (announced on 14 Dec. 2016) went to:
  • Idrissa Haji Abdalla (Tanzania), for Kilio cha Mwanamke (fiction #1)
  • Hussein Wamaywa (Tanzania), for Moyo Wangu Unaungua (fiction #2)
  • Ahmed Hussein Ahmed (Kenya), for Haile Ngoma ya Wana (poetry)
The 2015 Mabati-Cornell Kiswahili prizes went to:
  • Anna Samwel (Tanzania), for Penzi la Damu (fiction #1)
  • Enock Maregesi (Tanzania), for Kolonia Santita (fiction #2)
  • Mohammed K. Ghassani (Tanzania), for N'na Kwetu (poetry #1)
  • Christopher Bundala  (Tanzania), Kifaurongo (poetry #2)

Tuesday, June 07, 2016

Des infos sur l'atelier TALAf 2016

Voici quelques informations sur l'atelier TALAf (Traitement automatique des langues africaines) qui aura lieu le 4 juillet 2016 lors de la conférence JEP-TALN-RECITAL à Paris, France. (For English, see TALAf workshop.)

Il y a dix articles acceptés pour présentation à l'atelier : 8 en français, 2 en anglais. En tout, huit langues africaines figurent dans les sujets de ces articles : amazighe, bambara, comorien, igbo, maninka, peul, swahili, et wolof. Le programme suit :

09h30-10h00  Valentin Vydrin, Andrij Rovenchak & Kirill Maslinsky
Maninka Reference Corpus: A Presentation.
10h00-10h30  Ikechukwu Onyenwe, Mark Hepple & Uchechukwu Chinedu
Improving Accuracy of Igbo Corpus Annotation Using Morphological Reconstruction and Transformation-Based Learning.
10h30-11h00Pause café
11h00-11h30Moneim Abdourahamane, Christian Boitet, Valérie Bellynck, Lingxiao Wang & Hervé Blanchon
Construction d’un corpus parallèle français-comorien en utilisant de la TA français-swahili.
11h30-12h00David Blachon, Elodie Gauthier, Laurent Besacier, Guy-Noël Kouarata, Martine Adda-Decker & Annie Rialland
Collecte de parole pour l'étude des langues peu dotées ou en danger avec l'application mobile Lig-Aikuma.
12h00-14h00Pause repas
14h00-14h30Michael Melese Woldeyohannis, Laurent Besacier & Meshesha Million
Amharic Speech Recognition for Speech Translation.
14h30-15h00El Hadji Malick Fall, El Hadji Mamadou Nguer, Sokhna Bao Diop, Mouhamadou Khoulé, Mathieu Mangeot & Mame Thierno Cissé
Digraphie des langues ouest africaines : Latin2Ajami : un algorithme de translittération automatique.
15h00-15h30Fatimazahra Nejme, Siham Boulaknadel & Driss Aboutajdine
Développement de ressources pour la langue amazighe : Le Lexique Morphologique El-AmaLex.
15h30-16h00Alla Lo, Elhadji Mamadou Nguer, Abdoulaye Youssoupha Ndiaye, Cheikh Bamba Dione, Mathieu Mangeot, Mouhamadou Khoule, Sokhna Bao Diop & Mame Thierno Cisse
Correction orthographique pour la langue wolof : état de l'art et perspectives.
16h00-16h30Pause café
16h30-17h00Mouhamdou Khoule, Mathieu Mangeot, El Hadji Mamadou Nguer & Mame Thierno Cisse
iBaatukaay : un projet de base lexicale multilingue contributive sur le web à structure pivot pour les langues africaines notamment sénégalaises.
17h00-17h30Chérif Mbodj & Chantal Enguehard
Production et mise en ligne d’un dictionnaire électronique du wolof.

Les ateliers TALAf ont lieu tous les deux ans depuis 2012. Ils sont soutenus par le réseau Lexicologie, Terminologie, Traduction, une association internationale qui faisait partie de l'Agence universitaire de la Francophonie (AUF) jusqu’en 2010.

Selon le site web de TALAf, les rôles de l'atelier sont les suivants :
  • "mettre en relation les chercheurs du domaine grâce aux rencontres lors de l'atelier mais aussi avec la liste de diffusion ;
  • mutualiser les savoirs en utilisant des outils en source ouverte, des standards (ISO, Unicode), et en publiant les ressources produites sous licence ouverte (Creative Commons), afin d'éviter, entre autres, la perte d'informations lorsqu'un projet s'arrête et ne peut être repris immédiatement faute de moyens ;
  • développer un ensemble de bonnes pratiques fondées sur l'expérience des chercheurs du domaine. Il s'agit de mettre au point des méthodologies simples et économes en coût d'achat de logiciels pour l'élaboration de ressources, d'échanger sur les techniques permettant de se passer de certaines ressources inexistantes et enfin d'éviter des pertes de temps et d'énergie."

Thursday, April 30, 2015

Same-language subtitling for African languages?

The current edition of The Economist has a feature on "same-language subtitling" (SLS) as a literacy tool in India, entltled "Literacy in India: A bolly good read." Could SLS be used with African languages to promote literacy in Africa?

SLS is a bit like closed captioning in that it includes text in the language being spoken (or sung in), but the target is people who can hear and understand the language but are still learning to read. My limited understanding of subtitling is also that subtitled text gets more (or at least different) attention in production and display than does closed captioning.

The idea of SLS for literacy is not new, having been conceived two decades ago by Dr. Brij Kothari, who has continued work on development and use of the technique in India through the Indian Institute of Management Ahmedabad and later his NGO, PlanetRead.

More broadly, the idea of using subtitles in the same language goes back at least a half-century to sing-along shows (such as the 1960s American "Sing Along With Mitch" TV program; a contemporary example is Disney's release of a "sing-along" version of the 2013 film, "Frozen"). Many Chinese film and TV productions, as the Economist article mentions, subtitle in hanzi which can be read by speakers of different Chinese languages written with them (Mandarin dialects, Cantonese).

However, as far as I've been able to tell, there is not yet any use of SLS for African languages - at least on a systematic basis.

SLS in African languages?


A recent tweet by the Ghanaian NGO, Kasahorow, raised hopes of an example of SLS in the Akan language:
However, the YouTube videos are actually static images the of the lyrics with instrumental music in the background. Perhaps a step to SLS? Kasahorow, one should note, has quietly been doing a lot of production of learning and reference materials, plus some apps, for various African languages from around the continent. It would seem from afar that a collaboration between Kasahorow and PlanetRead could produce some very interesting results.

The topic of SLS came up in a session at the African Language Teachers Association/NCOLCTL conference last Saturday (25 April 2015). Two faculty from the University of Florida's Program in African Languages - Dr. James Essegbey and Dr. Charles Bwenge - presented on use of videos in Akan and Swahili (respectively) for L2 learners of those languages, and issues with production and access. A possible evolution of this kind of resource is to include subtitles/captions for the dialogues. While the subjects of these videos, and often the deliberate pace of speech in them (to facilitate L2 learners' understanding) may make most of them unsuitable for L1 (native) speakers, some of the more sophisticated ones might possibly be useful for literacy.

There is a significant amount of film, video drama, and music video production in African languages, and that is likely to increase. Use of SLS in popular releases might present a significant resource for L1 literacy in those languages, L2 language learning, and written use of African languages generally.

Final notes: I first learned of Dr. Kothari's work on SLS in the mid to late 1990s. The topic of SLS is mentioned in passing in two earlier postings on this blog:


Addendum, 1 May 2015

I understand from Brij Kothari that PlanetRead and Kasahorow have indeed collaborated virtually on one project to produce animated stories with SLS in Swahili. [See also his comment to this post.]

Addendum, 6 May 2015

I understand from James Essegbey that the Swahili videos shown by Charles Bwenge have subtitle capability.

Friday, January 16, 2015

Health information in African languages ... from the US

A number of health agencies in the US have information available in various languages, as part of serving immigrant communities. These include some African languages. A sampling of sites includes:
  • National Institutes of Health
  • EthnoMed (includes short list by language, as well as list of other sites, some of which may have materials in diverse languages)
  • Echo Minnesota (page has several languages, both in a short list on left and in the sidebar at right)
African languages with significant materials include: Amharic; Arabic; Oromo; Somali; Swahili; and Tigrinya. Such materials are freely available for use. They might also help in drafting articles for Wikipedia editions in those languages (per the work of Wiki Project Medicine's Translation Task Force).

Monday, September 08, 2014

Kamusi at 20: Keeping the vision alive and working

The Kamusi Project, which seeks to provide an open dictionary of all languages, for use in reference and in language technology development, is facing a challenge not unfamiliar to other language-related initiatives: Funding. This effort - sometimes perceived as too ambitious or esoteric but always visionary in its goals and use of technology - is currently campaigning for support through the Global Giving Open Challenge.

Kamusi was originally developed in 1994 as a proposal by Dr. Martin Benjamin and Dr. Ann Biersteker, then both at Yale University's Council on African Studies. Billed as the "Internet Living Swahili Dictionary," its objective was to respond to the need for new reference material on Swahili, and to do so by using the potential user contributions over the internet (this was more than 6 years before Wikipedia was launched). It is worth noting that a reason cited for exploring the internet medium for dictionary development was unfavorable "economics of Swahili publishing." Kamusi is still today an excellent Swahili resource (both monolingual and English <-> Swahili), even as its goals have evolved.

Kamusi was run at Yale, with benefit of US Department of Education funding, until 2006, and during this time was recognized as a finalist in the Stockholm Challenge 2001. At the end of this period, Dr. Benjamin - Martin to those who know him - summarized Kamusi in the context of African languages on the web at Wikimania 2006. He continued to run Kamusi as it transitioned in 2007 from Yale to a server hosted by the World Language Documentation Centre.

Under Martin's direction, Kamusi has since then been incorporated as a non-profit in the US and Switzerland (where he lives with his family), and has expanded its mission beyond Swahili to a pan-African and eventually global scope.

Funding from Canada's International Development Research Centre (IDRC) for Kamusi, as part of the multi-member African Network for Localisation (ANLoc) project, enabled Kamusi to lead development of locales for 100 African languages (locale data facilitates computer software handling a language) and terminology for 12 African languages. Later funding from the US National Endowment for the Humanities (NEH) enabled work on a pilot for Kamusi's multilingual model (basically, there's a lot more to a multilingual dictionary than words in parallel, since concepts don't line up neatly across languages).

Since the conclusion of major funding in 2012, Kamusi has continued work on the multilingual model, including how to annotate degrees of separation (when a concept is translated through another language), homophones, multi-word expressions (something I personally wish machine translation had been better at years ago), and data input from any language. In 2013 Kamusi's work gained it recognition as a launch partner in the White House Big Data Initiative.

Although Kamusi has an affiliation with l'École polytechnique fédérale de Lausanne (EPFL) since last September, this has not filled the funding gap to enable completion of the programming work necessary to bring all of this to fruition and take the Global Online Living Dictionary (GOLD) from a proven pilot project to a full-scale reality.
 
Looking at Kamusi's history - which is long in internet terms - one is impressed by the thought and effort that has gone into it, by Martin and by a range of other contributors, from its beginning at Yale to recent collaborations and donors, with many individual contributions all along. It would be a shame if current funding difficulties would cause this important work to end.

For something like Kamusi, it helps, I think, to look as far ahead as we can look back. Twenty yeas from now, the advantages of building language resources for the many languages that don't have the economic or political/policy weight to get commercial and investor attention - even if they have demographic importance (keep in mind how quickly Africa's population is growing, for instance), but especially if those numbers aren't there either - will be a lot more apparent than they seem today. For countries where many of these languages are spoken, like most of those in Africa, there is a long-term need for projects like Kamusi that connect high level language technology with less-resourced and often low-status languages - and in Kamusi's case, also link those with the more widely spoken international languages.

At this point, Kamusi's effort to gain enough support to qualify for ongoing listing on Global Giving is an attempt to keep the organization going at a critical period in its history. Please consider helping.

Saturday, November 30, 2013

Ethnologue and the cross-border languages of Africa

Two content features seem to me to detract from the overall quality of Ethnologue - which is the indispensable reference on all the world's languages. One is that pages on cross-border languages - a prominent sociolinguistic feature in Africa - are titled as the language of only one of the countries where they are spoken. The other is that in the current online version, the country pages list "national language," which is a problematic heading, given diverse use of the term, notably in a number of African countries. In this post I will deal with the first of these two items.

A continent traversed by cross-border languages

Due to the way borders in Africa were drawn, a great many of its ethno-linguistic groups are split among two or more of the modern African countries. The languages spoken by such groups can today be called "cross-border languages," and in fact this is the term used by the African Academy of Languages for some of the work it is doing, notably the Vehicular Cross-Border Language Commissions.

Languages that are spoken in more than one country - whether as a first language, which is often the case, or as a vehicular language, which is also frequent - have been a concern of language planning in Africa since independence. The various conferences on African languages that I have cited in some previous posts reflect this concern. Cross border languages were also highlighted by former Malian president Alpha Oumar Konaré, who compared them to "sutures" uniting African countries.*

"A language of" one country, more than one country, or a region?

When one looks up one of these cross-border languages in Ethnologue, however, they are as a rule listed as "a language of" a single country. It bears noting that when borders divide a language community, that community is rarely split in equal parts, so it appears that Ethnologue usually assigns the language to the country where there are more speakers. Other countries, regardless of significance in the language usage, are placed under "Also Spoken In:..."

So, for instance Hausa, the first language of 18.5 million Nigerians (according to Ethnologue, based on a 1991 SIL estimate), but also of about half the much smaller population of Niger, and used across large parts of West Africa (as a lingua franca), is titled simply "... a language of Nigeria." Is Hausa any less a language of Niger, given that the historic home of most Hausas (sometimes called Hausaland) extends well into the latter country? Better "A language of Nigeria and Niger"? Or given the number of other countries where it is "also spoken," maybe Hausa is really "A language of West Africa"?

Another example is the Ewe language, spoken by a population split between southeastern Ghana and southern Togo, which is listed as "... a language of Ghana." Similar to the case with Hausa in Nigeria and Niger, Ewe is spoken by more people in Ghana, but by a larger percentage of the population in Togo. So why not "A language of Ghana and Togo"? The reverse is noted in the case of Southern Sotho, which has more speakers in South Africa, but a higher percentage of population speaking it in Lesotho (it has legal status in both countries) - and is listed as "A language of Lesotho."

Examples abound, among which the major regional language of Swahili is listed as a language of Tanzania (see also discussion of macrolanguages, below).

Ultimately it seems (1) misleading to title pages on cross-border languages as languages of one particular country, and (2) inconsistent the way it is done. Going back to Ethnologue's "Plan of the Site" page, one finds mention of counting "each language only once as belonging to its country of origin" - but what if the area of origin of a language (to the extent one can determine that with any precision) is divided by borders?

Would it not be possible to develop a simple set of criteria by which cross-border languages were given titles  based on the extent (countries) of their major use?

Cross-border macrolanguages

The category of macrolanguage - defined as "multiple, closely related individual languages that are deemed in some usage contexts to be a single language" - takes this issue up another level. Although defined on linguistic criteria, macrolanguages in Africa are even more likely to cross borders, often many borders. There are 14 African macrolanguages by my count, with many of those being cross-border and some really looking like regional languages. Yet all of those are listed as being of one country or another:
  • Arabic, "A macrolanguage of Saudi Arabia" (spoken in many countries, including at least 9 in Northern Africa)
  • Dinka, "A macrolanguage of South Sudan" (spoken mainly in South Sudan)
  • Fulah, "A macrolanguage of Senegal" (spoken in well over a dozen countries, mainly in West Africa)
  • Gbaya, "A macrolanguage of Central African Republic" (spoken in CAR and Cameroon)
  • Grebo, "A macrolanguage of Liberia" (spoken in Liberia and Ivory Coast)
  • Kalenjin, "A macrolanguage of Kenya" (spoken mainly in Kenya, and also in Uganda and Tanzania)
  • Kanuri, "A macrolanguage of Nigeria" (spoken in 5 countries of West and Central Africa)
  • Kongo, "A macrolanguage of Democratic Republic of Congo" (spoken in DRC, Angola, and Congo)
  • Kpelle, "A macrolanguage of Liberia" (spoken in Liberia and Guinea)
  • Malagasy, "A macrolanguage of Madagascar" (spoken mainly in Madagascar)
  • Mandingo, "A macrolanguage of Guinea" (spoken in 7 countries of West Africa)
  • Oromo, "A macrolanguage of Ethiopia" (spoken in Ethiopia, Kenya, and Somalia)
  • Swahili, "A macrolanguage of Tanzania" (spoken in at least 9 countries mainly in East Africa, among which some governments have accorded it legal status)
  • (Akan is described as "A language of Ghana" in Ethnologue, and is also a macrolanguage in ISO 639-3. Either way it is considered to include Fanti and Twi, and is spoken mainly in Ghana.)
Here again, would it not be possible to adjust certain titles to more accurately convey the range of use? For instance, Fulah as "A macrolanguage of West Africa" and Swahili as "A macrolanguage of East Africa." The macrolanguage items may be easier to modify on a case-by-case basis, as there are fewer of them than languages, and their respective circumstances are somewhat unique. The language entries on the other hand might, as suggested above, need a set of criteria to avoid case by case discussion.

Final thoughts

These observations and suggestions are made in the spirit of helping improve the Ethnologue resource, with a mind particularly to what kind of information that people new to the study of languages of Africa would take away from their initial encounter with it. Cross-border languages exist in all world regions, of course, but perhaps in none more than Africa, where borders were never intended to respect the integrity of ethno-linguistic groups. This category of languages seems to me to merit attention and appropriate revision in how it is presentated.


* "Les langues nationales transfrontalières doivent être non pas des points limitrophes, des points de démarcation, mais des points de suture entre nos pays."

Thursday, November 28, 2013

Microsoft giving Africa LIP(s)

Where are we now with software localization in African languages? I'd like to try to take stock in several quick installments, beginning with desktop/laptop software and then moving to mobile devices. This post starts it off with Microsoft Windows and Microsoft Office, due to no particular ordering - I just happened to come across something recently relating to Microsoft's (MS's) localization efforts.

MS's products are offered in diverse languages, but in what might be described as a tiered arrangement. A number of languages have fully localized versions - for MS Office 2013, for instance, there are by my count 40 such versions (which includes only Arabic among African languages, as well as the principal Eurphone languages used in Africa). But for other languages, MS's Local Language Program develops Language Interface Packs (LIPs) for Windows and Office. A LIP includes translations of about 80% of commands (the most frequently used), and is installed over another version, basically changing the language interface to the language of the LIP.

MS has added a number of African language LIPs for Windows and Office over the past decade. This represents a significant amount of work, notably on terminology (another key topic I hope to return to).

MS Windows 7, which was released in 2009, had 10 African language LIPsWindows 8, released in 2012, introduced 3 more African LIPs (Kinyarwanda, Tigrinya, and Wolof), a Botswanan version for Setswana, and allowed installation of Hausa LIP on French in addition to English base language. The following list, derived from lists on the Windows site, summarizes African language LIP support under Windows 7 and 8 (in parentheses are the base languages on which the LIP can be installed, as well as indications for those languages added with 8 but not available for 7):

  • Afrikaans (English base language editions)
  • Amharic (English base language editions)
  • Hausa (English base language editions; in Windows 8, French base language also)
  • Igbo (English base language editions)
  • Kinyarwanda (in Windows 8 only; English base language editions)
  • Sesotho sa Leboa (English base language editions)
  • Setswana 
    • Botswana (in Windows 8 only; English base language editions)
    • South Africa (English base language editions)
  • Swahili (English base language editions)
  • Tigrinya (Ethiopia) (in Windows 8 only; English base language editions)
  • Wolof (in Windows 8 only; English base language editions & French base language)
  • Xhosa (English base language editions)
  • Yoruba (English base language editions)
  • Zulu (English base language editions)

The number of MS Office LIPs for African languages has gone from three for Office 2003 to 13 for Office 2013 (that's out of a total of over 100 LIPs worldwide). The table below is adapted from information on their Office Language Interface Pack (LIP) downloads page (check marks indicate LIP available):

Language Native name MS Office
2003
MS Office
2007
MS Office
2010
MS Office
2013
Afrikaans Afrikaanse
Amharic አማርኛ
-
Hausa Hausa
-
Igbo Igbo
-
Kinyarwanda Kinyarwanda
-
-
-
Sepedi /
Northern Sotho
Sesotho sa Leboa 
-
Setswana (South Africa)  Setswana
-
Swahili KiSwahili
Tigrinya ትግርኛ
-
-
-
Wolof Wolof
-
-
-
Xhosa isiXhosa
-
Yoruba ede Yorùbá
-
Zulu isiZulu

I have no information on any plans for other African languages.

MS additionally offers Multilingual User Interfaces which provide language interface options on a single device, though it is not clear whether any of these offer any African languages. (The term Language Packs, as distinguished from LIPs, has me [and apparently also Wikipedia?] a bit confused as it seems to be used in different ways.)

See also MS's Language Portal for additional information on their localization efforts. Also worth noting that evidently MS completed the LIPs much more quickly with recent releases than they had in the past.

I'd invite any comments - corrections or additional information. In getting back up to speed on this I again encountered MS's ever complex array of sites and pages about different programs, products, and versions, so it's probable I missed some relevant information... 

Saturday, September 30, 2006

Retrospective: "Wikimania," 4-6 August 2006

Just a quick note as September fades into October, and referring first of all to August...

At the beginng of last month there was a meeting in Cambridge, Massachusetts called "Wikimania", which brought together various experts and enthusiasts working on Wikipedia and related Wikimedia projects. Although I did not attend, I participated virtually, or as close to that as one could, in two sessions related to Africa, one by Martin Benjamin entitled "Huru na Bure: Swahili Collaboration and the Future of African Languages on the Web", and the other by Kasper Souren called "The Bambara Wikipedia, One Year Later" (see the discussion sections). In the runup to Wikimania I had corresponded with Martin and Ndesanjo Macha, cc'ing Kasper (at the time I did not know he was going to attend) about finding a way to support development of African language editions of Wikipedia. Some ideas found their way into a document called Facilitating African Language Wikipedias.

Following the conference, Martin and I set up a Yahoogroup called AfrophoneWikis for discussion of and collaboration on some of the points in that document. A list of African language editions of Wikipedia is available on the site for that group.

The BBC redio show, "Africa, Have Your Say", of Wed. 6 Sept. was devoted to African languages and Wikipedia, and featured among many others, Martin, Ndesanjo, and myself. (The recording of the show was apparently available for only a short period.

Anyway this is old news, except to say that the effort is ongoing and there is some noticeable expansion of some of the Wikipedias in African languages, notably Swahili, thanks largely to Ndesanjo.

Also, I think that the overall goals of localization are advanced by such specific / specialized projects to the extent that we communicate about them and share lessons.