Showing posts with label Kirundi. Show all posts
Showing posts with label Kirundi. Show all posts

Thursday, February 09, 2017

Health info in African languages, on 2 non-African sites

Here are quick reviews of two websites - one Australian the other American - that have health information in numerous world languages, including a number from Africa. Both are primarily intended to serve immigrant communities. This post will then return briefly to the theme of the benefits of systematically sharing and improving of health related information composed in or translated into African languages.

Health Translations


Health Translations is a website maintained by the government of the state of Victoria in Australia.  It has information (mainly documents such as fact sheets and flyers, from what I can tell, some with illustrations) on over 80 topics in a total of almost 100 languages or language varieties (although not all information is in every language, and some languages have few items).

The African languages for which there are materials include: Afrikaans; Akan; Amharic; Arabic; Bemba; Dinka; Juba Arabic; Kirundi; Krio; Lingala; Nuer; Oromo; Shona; Somali; Sudanese Arabic (also listed as Sudanese); Swahili (Congolese); Swahili (Kenyan); and Tigrinya.

This is an impressive collection of materials from various sources, apparently all Australian, and in different formats. They appear to all be translations - a given material may be available in a few or quite a number of language versions. Navigating from the a click on the desired language on the list of languages (which helpfully includes both English & native names/scripts) to a particular topical resource requires interacting with screens in English - not surprising, but when one gets to the list of topics and resources in a particular language, the titles are only in English, and then on the list of languages in which a material is translated (this is a typical navigation sequence), the language names are in English (no native scripts used). So the resource appears intended to be used by or with help of professionals or others who can read English.

Source: Bushfire smoke & your health [am]
I'm not able to evaluate the quality of the translations, but noted French in Lingala (which may simply be typically used loanwords) and by chance an anomalous English word in an Amharic text (image).

All the several documents I viewed were PDFs, mainly text but some image (meaning the text cannot be searched or copied out for editing into other materials). Spot checking some non-Latin text, specifically Ethiopic/Ge'ez used for Amharic and Tigrinya, and complex Latin, specifically for Dinka, there were some issues with the text that would interfere with searches or copying out passages (such problems are not uncommon with PDF rendering, even when visually the PDF presents everything correctly and in its intended place).

Some Amharic text when copied out and pasted showed capital A for አ and E for እ in initial position (for example, here). The corresponding characters are appropriate, interestingly, but this makes search or reuse of such text problematic.

From original (l.); copy-pasted out (r.)
Source: Bushfire smoke ... [din]

The Dinka text sampled showed some typical problems with complex Latin in PDFs. Dinka is written with what I call in ALDA a "category 4" Latin orthography in that it includes extended Latin characters (aka modified letters) plus combining diacritics, sometimes together as in the open-o with diaeresis in the word "daiɣɔ̈kthai" (dioxide) featured on the left side of the image. Copying that word from the PDF and pasting it in a word processor or advanced text editor yielded the results on the right, missing one extended character and the combining diacritic on the other. This complicates any potential re-use of this text, but also means that document, folder, or web searches will not pick up words with such character combinations.

Will return to these issues, why they're important, and what to do about them in the last section of this post.

 

HealthReach


HealthReach: Health Information in Many Languages is a program of the US Library of Medicine of the National Institutes of Health. It includes translations in 46 languages maintained on the MedinePlus site. A total of almost 350 topics are listed (although here too, not all information is in every language, and some languages have fewer items than others).

The African languages for which there are materials include: Amharic; Arabic; Oromo; Somali; Swahili; and Tigrinya, The native names of languages are featured on the list of languages, except oddly for Amharic and Tigrinya, which are transliterated into Latin ("amarunya" instead of አማርኛ, and "tigrinya" instead of ትግርኛ).

This also is an impressive collection from diverse sources, in this case American, but it is longer on topics and shorter on languages covered. The list of topics for each language also includes the titles in the language and its script - except again for Amharic and Tigrinya (not even transliterations) - as well as in English.

All materials checked were PDFs. There are no materials for African languages with complex Latin scripts.

As for non-Latin scripts, text in Arabic seems to behave as intended, from small samples. On the other hand, some Amharic text when copied out and pasted showed the same capital A for አ and E for እ observed above, plus O for ኦ (see here).  A Tigrinya document had a similar issue. So this issue may have to do with a problem in PDFs for handling a particular set of characters - አኡኢኣኤእኦኧ (representing glottal stop plus the range of vowels) - or a subset of them, which might be helpful to know when troubleshooting.

Health education materials and the "2Ds & 4Rs"


In highlighting aspects of public health messaging during the ebola epidemic in West Africa (2014-15), this blog suggested a systematic approach to sharing and improving materials that were developed and used in that context (with primary attention to text and images). A mnemonic - 2Ds & 4Rs - was put forth in October 2014, initially to explain the rationale for reposting and discussing various ebola education materials, but also as a way to capture the ideal cycle of utility of such production. Too often, materials are developed, used for a particular purpose, and then forgotten, when they could add to a growing living corpus of resources to tap for future work. This is important in any field and language, but arguably especially important in health, and for languages that have fewer resources and emerging terminologies / technical lexicons, such as many in Africa.

In that context I propose to use the 2Ds & 4Rs to consider the efforts represented by the two sites discussed above. Of the 6 elements of this model, the first three have to do more with the sharing and use of materials, and the last three with their longer term development and potential re-use. These are listed with brief explanations and what I see as relevance to the two sites:
  • Dissemination (making materials available, including via multiple sources)
    • Both sites bring together and post materials from diverse sources, increasing their exposure and access to them.
  • Demonstration (showing how materials in African languages can be presented, including in cases where complex scripts are involved)
    • Both sites show that African language materials can be presented on the same footing as other world languages.
    • However, the HealthReach presentation does not use available technology to present the native names of Amharic and Tigrinya, or titles of materials in those languages.
  • Reading (creating or translating text materials with attention to how they may be read aloud in groups or over local radio, which may be more likely scenarios for their use than the typical Western expectation of silent reading by individuals)
    •  It appears that all or most of the materials from diverse sources compiled on the two sites are translations from English of technical descriptions and advice. It is not clear how well how well adapted they are for the range of uses and audiences they might serve.
  • Review (written material - text - is well suited for review, comparison, and analysis; such material, especially in less resourced languages and on issues of public importance like health, should undergo such treatment)
    • No information on how any of the materials may be or have been reviewed, either in the diverse organizations where they originated, or in the projects hosting the two websites. 
    • Image PDFs, where these occur, do not lend themselves to processes of review.
    • Text PDFs with problems in their encoding of non-Latin or complex Latin scripts, present problems for review.
  • Revision (after review of materials, and in response to other information and feedback relevant to them, materials should undergo appropriate revisions in content, form of language, copyediting, and presentation)
    • No information on any revisions of any of the materials.
    • Issues cited under "Review" with image PDFs and with text PDFs that have encoding problems also hinder revision work.
  • Re-use or re-purposing (text materials can be re-used or sections re-purposed)
    • No information on re-use or re-purposing of any of the materials.

The two sites profiled above and the various health and medical education materials presented on them represent an important resource for fifteen African languages (and some varieties of two of those).

One additional question is whether such materials, intended primarily to serve needs of immigrants in Australia and the US, might be useful as is or with modifications, for speakers of the same languages in relevant African countries. Or in the reverse sense, whether any health extension materials from Africa might inform revision of these materials and development of new ones. A next step could be a for a site to begin to collect health materials in African languages from all sources.

There are many directions in which this could be taken, with the goals of improving availability, quality and utility of health education information in a range of African languages. One, for example, is linking with the longstanding WikiProject Med's Translation Task Force for development of articles in those African languages that have Wikipedias (such as Afrikaans, Akan, Amharic, Arabic, Kirundi, Lingala, Oromo, Shona, Somali, Swahili, and Tigrinya). Another might be connecting with efforts to advance development of standard terminologies. Still another might be to bring in human language technology, such as text to speech, so that materials designed and disseminated in text form could be accessible in audio via mobile devices.

Thanks to Charles Riley of Yale University for calling our attention to these two websites.

Wednesday, August 31, 2016

Missing "macrolanguages" of Africa

Screenshot from VOA's Kinyarwanda/Kirundi site
The Voice of America (VOA) recently had a job opening for "International Broadcaster (Multimedia) (Kirundi/Kinyarwanda)." Kirundi and Kinyarwanda are the mother tongues, national languages, and co-official languages in, respectively, Burundi and Rwanda. And they are mutually intelligible, with only minor differences, such that apparently a fluent speaker of either could work on a program serving speakers of both. But there is no term covering both - unless one counts the hyphenated Rwanda-Rundi - and no language coding category to cover material designed for use across the two.

This is a situation encountered with many languages in Africa, and one for which there is at least one potential solution - the neologism and language coding category "macrolanguage." There are actually some macrolanguages defined in Africa, but these are few, and as I discuss below, kind of accidental. Is it time to systematically identify (and code) macrolanguages in Africa?

What defines a language?


For most of us, the distinction between languages seems pretty straightforward. But beyond the most spoken international languages - those used officially by the United Nations or ones you are likely to see on a school curriculum - the situation is often more complex. Sometimes two or more closely related languages are so similar that their speakers can understand each other, but sometimes variations within one language can make understanding difficult. An earlier posting on this blog looked at the notion of "neighbor languages" in Scandinavia and Africa. A broader consideration of these issues by Columbia University's John McWhorter suggests that we're really all speaking dialects, some of which benefit from written forms, and one might add, status, resources, and policy support. There is some truth to the saying that "A language is a dialect with an army and a navy."

However, the issues of what to call a "language" and where to draw the boundaries between it and another "language" are still of practical importance for communication (standardization, references, ICT use) and planning (government, business, education). There are two broad approaches in linguistics to doing this, corresponding with the splitter/lumper (or joiner) approaches to categorizing:  one focusing more on distinctions, and the other focusing more on commonalities.

Without going too deeply into that discussion, which gets more complicated when accounting for issues of identity, names, written forms, and national boundaries, suffice it to say that in considering African languages, there are many situations where one encounters the splitter/lumper choice.

The major reference of languages in the world, Ethnologue, takes a more splitter approach, which means that speech varieties that are closely related and interintelligible may be classified as separate languages. It is their estimate of the number of language in Africa (over 2000) that is most commonly cited, but there are other more conservative estimates.A good academic discussion of this issue entitled "How many languages are there in Africa?" was published in 2004 by Jouni Filip Maho (his estimate is under 1500).

What is a "macrolanguage"?


To make the story brief, the term "macrolanguage" is not a term that was used in linguistic description before the inauguration of the  ISO 639-3 system for encoding all languages in the late 2000s. Since that system is based on Ethnologue's "splitter" data, a new category was needed to accommodate existing codes in the earlier less comprehensive parts of ISO 639 (1&2) that in many cases were more "lumper" in approach. The term macrolanguage was in effect a "shim," to borrow someone else's term, to fit the two systems together.

There are by my count 14 macrolanguages listed for Africa (names linked to the Ethnologue macrolanguage pages): Akan; Arabic; Dinka; Fulah; Gbaya; Grebo; Kalenjin; Kanuri; Kongo; Kpelle; Malagasy; Mandingo; Oromo; and Swahili. There could be others.

That brings us back to Kinyarwanda and Kirundi. How is the relationship between them different - more distant - than any of the above established macrolanguages? One difference, as mentioned above, is no common name to make it easy, and another is that they are dominant in different countries - perhaps analogous to the situation of Scandinavian languages?

Another curious situation is that of Mandingo, which includes several western Manding languages, but not Bambara and Jula (Dyula). Even if the latter two were considered too different from the other Manding tongues, they are close enough that one could localize software for the two together. Keep in mind also that the emerging literary standard N'Ko covers all Manding languages (in a different alphabet). Should the Mandingo macrolanguage be extended to include them all?

The four languages of southwestern Uganda - Kiga, Nkore, Nyoro, ajd Tooro - are close enough to be covered by Runyakitara, a proposed (but not encoded) standard which is being used in various ways, including at least some teaching and a localization of the Google interface. Should these four be considered a macrolanguage under perhaps that same name, thus finally providing a code for localization in Runyakitara?

And there are other examples around the continent that could be discussed.

What good would more macrolanguages do?


The first benefit of identifying more macrolanguages would be in language coding - the very environment in which the term was first used. The language of VOA's website for its Kinyarwanda/Kirundi service - www.radiyoyacuvoa.com - is coded as "rw" (Kinyarwanda) since there is no macrolanguage code covering both languages. Likewise, in many cases, the grouping of very close and mutually intelligible languages as a macrolanguage could facilitate localization of software and apps to serve larger populations - and those larger markets could make it more likely that such localization would be pursued and maintained.

Another benefit would be to complement the tendency in language coding towards seeking more granularity, by recognizing natural groupings of languages (for more on this, see a message to the IETF-languages list last May). In effect providing more balance between splitting and lumping/joining.

In the broader picture, identifying macrolanguages could have benefits for policymaking and program development involving languages within macrolanguage groups, by calling attention to the closely related languages. Especially where foreigners are involved, projects may overlook such relationships and the potential resources they may provide. For example materials development for education, and various communication needs might benefit from tapping efforts and resources in closely related languages.

(Minor edits and image added, 2 Sep. 2016)