Showing posts with label Ge'ez. Show all posts
Showing posts with label Ge'ez. Show all posts

Friday, August 31, 2018

Niamey 1978 & Cape Town 2018: 2. Extended Latin & African language Wikipedias


Image adapted from banner on the Yoruba Wikipedia, August 2018
What are the implications of extended Latin characters and combinations for production of digital materials in African languages written with them? The previous post discussed some of the process of seeking to harmonize transcriptions, in which the Niamey 1978 conference and its African Reference Alphabet (ARA) were prominent. That process had a logic and left a legacy for the representation in writing of many African languages. This post asks if there is a trade-off between the complexity of the Latin-based writing system and how much is produced in it using contemporary digital technologies.

One easy, although by no means conclusive, way to consider this question is to look at Wikipedia editions in African languages (those that are written in Latin script). The following table disaggregates 35 African language editions by the number of articles (from the list of Wikipedias, as of 9 August 2018) and the four "categories" of Latin-based orthography1 introduced in African Languages in a Digital Age (ch. 7, p. 58):

Number of articles
Category 1
Category 2
"Category 1" + Latin 1 
Category 3
"Cat. 1" or "2" + any of Latin Extended A, B, etc, Add'l, & IPA
Category 4
"Category 3" + Combining diacritics
< 500
Swati (447)
(Chewa)
Sango (255)

Fula (226)
Venda (265)
Chewa2 (389)
Dinka (75)
Ewe (345)
500-1000
Sotho (543)
Tumbuka2 (562)
Tsonga (563)
Kirundi (611)
Tswana (641)
Xhosa (741)
Oromo (772)
-
Akan (561)
Twi (609)
Bambara3 (646)
(Bambara)
1000-2000
Zulu (1011)
Kinyarwanda (1823)
Kongo (1179)
Luganda4 (1162)
Wolof (1167)
Gikuyu (1357)
Kabiye (1455)
Hausa (1891)
Igbo5 (1340)
2000-5,000
Shona (3761)
-
Kabyle (2860)
Lingala (3028)
5000-10,000
Somali (5307)
-
-
10,000-25,000
-
-
-
-
25,000-50,000
Swahili (44,375)
-
-
Yoruba5 (31,700)
> 50,000
Malagasy (85,033)
Afrikaans (52,847)
-
-
# of articles / # of editions = Average
146,190 / 14 =
10,442
54,281 / 3 =
18,094
20,655 /13 =
1,589
36,488 / 5 =
7,298
Grouping 1&2, 3&4
168,471 total articles / 17 editions = 
11,789
57,143 total articles / 18 editions = 
3,175


Looking at the top row with the smallest editions (less than 500 articles), one is tempted to highlight the high presence of African languages whose orthographies include extended Latin - categories 3 & 4. However, in the group of next highest number of articles (500-1000) there are more editions with category 1 orthographies (the simplest) than there are editions with category 3 in the group above that (1000-2000). And the next highest ranges (covering 2000-10,000) are roughly even between category 1 on the one hand, and 3 & 4 on the other. But then the three largest editions (and 3/4 above 25,000) are category 1 & 2.

So with just a visual analysis, there does not seem to be any clear pattern from arraying the editions in this way. Of course there will be other factors than the complexity of the script affecting the success of a Wikipedia edition written in it. But are there ways of looking at this raw data that can give us a clearer idea what might be the effect is of extended Latin - the ARA plus orthographies with other modified letters and diacritic combinations - on the size of Wikipedia editions?

One approach is to consider all the above editions combined, per category of orthography (totaling by column). This puts the focus on the degree of complexity of the writing system, perhaps muting the effect of other language- & location-specific factors. On the second to last row are column totals of the number of articles in all editions listed above, divided by the number of editions, to give an average figure.This yields an uneven pattern (2>1>4>3), since in the cases of 2 & 4, one large edition in a small total number of editions skews the category average up.

By the totals of the two simpler categories (1 & 2) and of the two extended Latin categories (3 & 4), however, one obtains possibly more useful numbers. This aggregation can be rationalized for our purposes here by the fact that the lower two categories are generally supported by commercially available keyboards and input systems,6 while the higher two categories, require a specialized way to input of additional characters and maybe diacritics (such as an alternative keyboard driver, or an online character picker).7

The figures thus obtained show editions written in extended and complex Latin having on average about a third the number or articles as those written in ASCII and Latin-1. Admittedly, this result is in part the result of the way categories have been chosen and figures aligned, but I'm proposing them as a perspective on the use of extended (and complex) Latin, and possible gaps in support. Before considering this in more detail, it is useful to compare with the numbers for non-Latin scripts.

What about non-Latin scripts & African language Wikipedias?


Number of articles
Non-Latin
< 500
Tigrinya (168)
10,000-25,000
Amharic (14,321)
Egyptian Arabic (19,170)
# of articles / # of editions = Average
33,659 / 3 = 11,220
There are only three editions of Wikipedia in African languages written in non-Latin scripts.8 Two of those - Amharic and Tigrinya - are written with the Ge'ez or Ethiopic script unique to the Horn of Africa.

Arabic is the third. How to count this language for the purposes of this informal analysis raises a question. Arabic, of course, is established as a first language in North Africa for centuries, but it is also a world language, spoken natively in the southwest Asia (having originated in Arabia), and learned as a second language in many regions. Drawing users from this wide community, the Arabic Wikipedia is among the top 20 overall, with twice as many articles as all of the editions discussed above combined. It is more than an African language edition. For this analysis, therefore, I have chosen instead to count just the Egyptian Arabic Wikipedia.

Taking these three editions, we then get an average number of articles (11,220), which is close to what is seen for the Latin categories 1 & 2 (11,789). The usual caveats apply for such a small sample, but taking the numbers as they are, it is interesting that Wikipedias in the complex Arabic alphabet and the large Ge'ez abugida (alphasyllabary) are on average much larger than those of the ostensibly simpler extended Latin (3,175).9

Again, script complexity is but one factor, and in this case probably not the most important, since the two non-Latin scripts in question have long histories of use in text in parts of Africa - much longer than any form of Latin script. Nevertheless, from the narrow perspective of what is required for users to edit Wikipedia, the technical issues are in some ways comparable if even more demanding.

Arabic has had standard keyboards since the days of typewriters. The issues there are not so much the input, but whether systems can handle the directionality and composition requirements of the script.

The Ge'ez script on the other hand, does not involve complex composition rules or bidirectionality. However, it has a total of over 300 characters (including numerals and punctuation; more again if extended ranges are added). The good news is that there are numerous input systems to facilitate their input. Literacy in the script and availability of input systems would not be limiting factors for content development in major languages using this script. The difference in development of the Amharic and Tigrinya editions of Wikipedia may relate to both the larger population speaking Amharic (as a first or second language), and its use officially in a relatively large country (Ethiopia). Development of content in Tigrinya - a cross-border language - might also be hindered by issues particular to one of the two countries where it has many speakers (Eritrea).

From the above one might suggest that complexity of the written form (to be taken here as including the nature of the script itself, and the size of the character set) may be a limiting factor on content development, but that other factors, such as a literate tradition, official use, and technical support for digital production may overcome such limitations. In the case of African languages written in Latin script, however, any literate tradition is recent, and they are often marginalized in official and educational contexts. For those written with extended Latin, there is the additional factor of lack of an easy and standardized way of inputting special characters. Paradoxically, it seems, a modification of the most widely used alphabet on the planet may actually hobble efforts to edit in these languages.

Facilitating input in extended Latin for African language Wikipedias?


Wikipedia editing screen with "Special Characters"
drop-down modified to show all available ranges.
Assuming that the inconvenience of finding ways to input extended Latin characters may be a factor in the success of African language Wikipedias written with categories 3 and 4 orthographies, a quick fix might be to add new ranges for the modified letters used in African languages to the "special characters" picker in the edit screens. As it currently structured, the extended characters necessary for a category 3 or 4 orthography might be sprinkled around in up to 3 different ranges (see at right). And within each range, they are not presented in a clear order, so sometimes hard to find.

Since it may be too complicated to have a special range for each language edition, another possibility would be to draw inspiration from the Niamey 1978 meeting's ARA, and combine all extended Latin characters and combinations needed for all current African language Wikipedias into a common new range.

Of course as mentioned above, there are other factors that can contribute to the success or not of Wikipedia editions in African languages written with extended Latin, but this innovation would at least make editing more convenient for contributors to these  editions. And perhaps it might have a positive effect on the quantity and quality of articles in these Wikipedias.

In the third, and concluding article in this series, I'll step back to look at this analysis and consider some other ways to look at the data on African language editions of Wikipedia, and in particular, those written in extended Latin.

1. This categorization was intended to help characterize the technical requirements for display and input of various languages. Although the technology has improved to the point that more complex scripts are generally displayed without the kinds of issues one encountered a even a decade ago, input still requires extra steps or workarounds. The four categories are additive in that each higher category builds on those below, with added potential issues. It is also a "one jot" system in that for example, a single extended Latin character, say š in Northeren Sotho or ŋ in Wolof, makes their orthographies category 3 rather than category 1 or 2 (respectively), and the use of the combining tilde over the extended Latin character for open-o - ɔ̃ - makes Ewe a category 4 rather than 3. In terms of input, the higher the category, the more the potential issues with display and input (although technical advances tend to level the field, esp. as concerns display).
2. The only non-basic Latin character used in Chewa is the w with circumflex: ŵ. Apparently it represents a sound important in only one dialect of the language, and is used infrequently in contemporary publications. On the other hand, there is a proposed (not adopted) orthography for Tumbuka that includes the ŵ. Without this character, either language would be a category 1 orthography; with it, category 3.
3. Bambara is a tonal language. Most often, it seems, tones are not marked in text, however they can be for clarity, and some dictionaries make a point of indicating tone in the entries (not just pronunciation). If tones are unmarked, Bambara would be considered is a category 3 orthography; with tones, category 4. 
4. The addition of the letter ŋ puts Luganda in category 3 rather than 1.
5. The dot-under (or small vertical line under) characters used notably in Yoruba and Igbo are particular to southern Nigeria, and not included in the ARA. Yoruba in Benin is written with characters from the ARA.These are tonal languages, and tone is usually marked.
6. When I first proposed the category (itself a modification of an earlier effort), there were some questions why have a category 2 separate from category 1. That distinction had its origins in the early days of computing where systems used 7-bit fonts, meaning that accented letters (diacritic characters) used in, say, French or Portuguese, could not be displayed. Even as systems using 8-bit fonts enabled use of diacritics commonly used in European languages, display issues would still crop up (as a sequence of characters where an accented letter should be). Nowadays, such display issues are rare, and limited (as far as I can tell) to documents in legacy encodings. On the other hand, input of accented characters used may require, depending on the keyboard one is using, switching keyboard drivers or using extra keystrokes - so one will occasionally see ASCIIfication of text in such languages (apparently as a user choice).
7. The difference between categories 3 (extended Latin) and 4 (complex Latin) once were significant enough from point of view of display that informal appeals to Unicode to change its policy of not encoding new "precomposed" characters were common.
8. The Wikipedia incubator projects.includes several African language projects, which are not covered here. These include some in non-Latin scripts (Arabic versions, N'Ko, and Tamazight) and some in Latin-based orthographies. I mentioned one of the latter - Krio - in a previous post, and hope to do an overview of this space in the near future.
9. Average for all African language editions is 7704. By comparison the average for all Wikipedias is 166k.

Tuesday, June 30, 2015

Unicode and the architecture of ICT

Latin omega, used in Kulango, and now
included in Unicode's latest version.
Unicode released its 8th version earlier this month, so maybe it's a good time to take stock of what the effort to encode all the world's scripts in a single system - which includes the Universal Coded Character Set (ISO/IEC 10646) - has meant for African languages.

The attention with a coding system for characters is naturally on the characters and scripts they are part of, and among the changes in Unicode 8 are some Latin-based characters used to write the Kulango language of Ivory Coast and the endangered Ik language of Uganda. These have been added to Unicode's "Latin Extended-D" block. (The Unicode standard also includes technical specifications for handling text.)

In past years there have been similar small additions to Unicode, as well as major additions of whole African scripts like Osmanya (2003), Tifinagh (2005), N'Ko (2006), Vai (2008), Bamum (2009), Mende Kikakui (2014), and Bassa Vah (2014).

The Ethiopic/Ge'ez script - used for several languages of Ethiopia and Eritrea, such as Amharic and Tigrinya - was first added in 1999, before the Unicode standard was widely adopted. At the time, different scripts used different coding systems, hindering use of languages with complex or non-Latin scripts, and posing particular problems for anyone wanting to combine scripts in a document or on a webpage.

Unicode as an enabling architecture

With that I wanted to quote from an unsigned blog posting entitled "What are we missing out on" (9 Dec. 2014) that was part of an American University course on international communication, as I think it offers an important perspective on why encoding scripts in Unicode is important:
"Tonight’s presentation on Ethiopic language and its inclusion in Unicode presented an important element about the global digital divide because it asks the question: how can, even with access to information communication technologies and internet access, someone utilize technology if it is not available in their native language? In short, they can’t. This is an important element to consider in regard to technology and the reality of, to borrow from Laura DeNardis’s description, its architecture. As the group described in their literature related to their case study, the architecture of something has the power to include and exclude. The analogy we have used in class in class is that of bridges that are built low enough to prevent busses driving under them.
"In the case of Ethiopia, technology was created in a way that excludes the nation’s 90 languages because American companies created the technology in English with a western cultural perspective. Further, their commercial interests drive their actions, and there is no financial incentive to include languages in which there is no commercial demand. Therefore, not only are these groups of people excluded from the benefits of technology, we are also denied the benefit of their knowledge. As the group noted, we “feel” like we are so much more connected, but cannot assume that the majority of the information is in English and unless everyone is able to put information in the digital realm, we are missing out. After this presentation, I cannot help but to believe that we are indeed missing out. Tonight we discussed Ethiopic, but what other languages are we missing out on?"
The "architecture" of information and communication technology (ICT) in this case starts with the internationalization piece ("i18n"), which is addressed in part by Unicode. It would also include availability or not of localized ("L10n") software, which is especially important for major languages, and content. Localization is also conditioned by factors such as education, policy (of national governments as well as of development organizations), and, as alluded to above, economics. As such, once one moves beyond the enabling architecture, the interplay of factors looks more like an "ecology."

Unicode, by adding scripts and amending characters, makes the written form of African languages theoretically accessible on modern computing devices and across the internet. That's huge progress beyond where things were a decade and a half ago, when various 8-bit encodings dominated (at which time I noted some issues of lack of support for African language scripts were stuck at practically the same place they had been ten years before!). But it's still only part of the process of fully addressing the linguistic dimensions of the digital divide in Africa. Until that happens, we all will be "missing out" in different ways.


Tuesday, September 06, 2005

Three more items before the IDN/Unicode conference: Connections en route; education in African languages; African studies in China

Here are 3 unrelated items, the first was written mostly on the last leg of the trip yesterday, the second is part of a letter about education policy written earlier that I've been intending to post (relevant to this blog and in a way to the context of discussion of things like IDNs and localization), and the third to a link between where I just came from and where I am.

I will also quickly mention three other items that have just come up today relating to the upcoming conference: (1) I just received a note from Daniel Yacob with a powerpoint on IDNs in Ethiopic (script used for several languages of Ethiopia and Eritrea). It is of course with the non-Latin scripts that a lot of the most interesting problems for IDNs are encountered; (2) there is a series of articles on the IDN/Unicode conference and African language computing generally in a special issue of Les Echos here in Dakar (see the files section of Unicode-Afrique for the PDF of this); and (3) I just met Mamady Doumbouya of the N'ko Institute who is also here for the conference. We had a good talk about aspects of localization, African languages, and the N'ko movement. N'ko is a script devised less than 50 years ago, but is increasingly used in the Mandephone parts of West Africa (and it is in the process of being approved for addition to Unicode).

The three items I mentioned are as follows.

Connections en route: airport to airport

The travel day, and it is a long day traveling with the sun across Eurasia and then south to Africa, is not so hard as it is one that demands patience, then at points some frantic rushing and then more patience. Time to fill with some work and thinking. There is not much rest, but some adrenaline.

On this trip I had more opportunity to look at the various connection possibilities in airports. I was impressed that Chengdu airport now has a nice cyberlounge where you can link via cable (broadband) or use one of their computers. This is new, as far as I've noticed. Connection was broadband by cable.

Beijing has the same reliable but rather expensive business lounge (these are not the "business class" & first class lounges of the airlines). The cable hookup did not work for me even when trying to reconfigure my settings. The China Unicom wireless signal was not clear this time, though I didn't waste much time looking for it since you had to have a cellphone account with them to use it.

I was not able to link up on Paris Charles de Gaulle Airport for-pay wireless this time - not enought time in layover.. I did try in the plane as it was still loading, but the signal did not carry outside. This airport also has (or at least did as recently as last June) little for-pay internet kiosks.

No attempt to connect in Dakar airport on arrival, but I am pleased that the Novotel where we are staying has wireless in the ground floor (lobby, restaurant, bar).

Substantively some of what I'm doing is to keep thinking through all the things I want to cover with various people in Dakar in the conference and in diverse other meetings. Over the years so many issues and questions seem to have some connection with someone or some organization in Dakar. Part of what malkes this trip worth making the time is the potential to (re)connect with so many people.

Of course the IDN (internet domain names) & Unicode conference is the main event, and will bring together a number of people. Je tenterai d'ecrire un peu sur les participants et l'activite meme prochainement. From the description of the project of which this event is a part, it seems the issue is bigger than IDNs only. More soon - hopefully they'll put this document on the web.

Education in national languages of Ghana

The following is an excerpt (slightly edited to read better) from a letter I wrote to Paa Kwesi Imbeah on the subject of the apparently pretty exclusive focus on English as language of instruction that one sees in Ghana these days. I think it is useful to bring up the educational angle again as we prepare in Dakar for this Unicode/IDN meeting. (Paa Kwesi, by the way, will be presenting a paper on the Akan online dictionary at the Unicode conference IUC28 in Orlando, Florida, which also begins tomorrow):

English is the "language of the belly" or "language of the stomach," as they say. Some people see it as their ticket to eating enough. Others see it as their ticket to eating a lot. There is some truth to that but it hides other realities. One is that neglecting or, as a colleague once put it, taking for granted the indigenous languages leads to some significant losses and costs that aren't imediately apparent.

One you hint at in your letter is limited or impaired bilingualism or worse, semilingualism. This has been touched on in some entries on this blog and also in the Multilingual_Literacy group (link in the left hand column ; you can search the terms on the group's page).

Ghana's government is not alone in focusing on English. In the current global economic situation, the enhanced prospects of outsourcing industries, outside investments, etc. (already a part of the scene in the country) mean dollar and cedi signs to planners. But the global climate could change drastically in the future with English being less central (hard to see now but who can say?). So if Ghana sells its linguistic heritage for a middling average national competence in what is still essentially a foreign/international language, where would it be then?

But the worst of it is that it isn't an "either-or" question but a "both-and" issue. Good bilingual education can give you the best of both worlds (a lot of research worldwide shows this): Ghanaian languages and English. Unfortunately, the way it goes, they don't take advantage of the "both-and" approach and so somehow end up with a "neither-nor" result for a lot of the population.


I would add that there was a conference last month in Windhoek, Namibia on bilingual education in Africa. See an article and the conference document (the latter in PDF format and rather large).

African studies in China

I finally caught up with Prof. Li Anshan's article that was published a few months ago: "African Studies in China in the Twentieth Century: A Historiographical Survey" in the African Studies Review, 48(1): 59-87. Part of what interests me about China-Africa connections is that they are getting increasingly important. In and of itself, and as part of broader evolution of so-called South-South relations, the relationship between China (the world's most populous country with a rapidly growing economy) and Africa (the second largest continent with a rapidly growing population) will become ever greater in the development picture for Africa. It will be interesting also to see the evolution of studies and understanding of Africa in China and vice-versa. Anyway, an abstract from Johns Hopkins' Project Muse follows. I would add to it only that Prof. Li mentions that the only two African languages taught in China are Swahili and Hausa.

This article surveys African studies in China during the twentieth century. It is divided into five parts: "Sensing Africa" (1900–1949), "Supporting Africa" (1949–65), "Understanding Africa" (1966–76), and "Studying Africa" (1977–2000). From a Chinese perspective, the author tells how, when, and why Chinese scholars have conducted their research on Africa according to paradigms that evolved during the last century. In conclusion, the author points out the achievements as well as the problems in African studies in China today.

Cet article propose un aperçu des études africaines menées en Chine au cours du vingtième siècle. Il est divisé en cinq parties: «Approcher l'Afrique» (1900-1949), «Soutenir l'Afrique» (1949-65), «Comprendre l'Afrique» (1966-76) et «Étudier l'Afrique» (1977-2000). A partir d'une perspective chinoise, l'auteur examine comment, quand, et pourquoi les chercheurs chinois ont mené leur recherche sur l'Afrique, selon des paradigmes qui ont évolué au cours du siècle. En conclusion, l'auteur souligne les succès et les difficultés rencontrés par les études africaines en Chine aujourd'hui.