Showing posts with label QDA. Show all posts
Showing posts with label QDA. Show all posts

Thursday, September 28, 2017

A terminological issue in cross-language qualitative methods

What do you call it when someone hears something in one language, and then writes down the meaning in another language? It is technically not translation, interpretation, or transcription, in their purest senses, even though these terms are sometimes used to refer to the process or its products. Should we have a new term for this practice, especially to distinguish it from the alternative method of transcribing in the one language and then translating into the other?

This is an issue particularly relevant to qualitative research in Africa, where focus groups or interviews are often done in a language other than the one in which research analysis and reporting take place. It is also important in other multilingual countries and regions, and indeed some of my examples below come from experience in Afghanistan.

Two approaches


When I was managing research projects in Kabul in 2013, we handled qualitative data in the form of recordings of focus groups and in-depth interviews by having them transcribed in the source language (Dari or Pashto), then translated from the transcripts. This was standard in the company I worked with, and as a research practice, from all I had learned previously. I have in my possession, for example, a photocopy of a rather lengthy transcript of a recording in Fulfulde, prior to translation.1

It was of some surprise, therefore, to learn from another organization in Kabul that they paid people to translate or interpret into English directly from recordings, with no same-language transcripts. I later found that this kind of shortcut is actually not that uncommon. For example, I am working on a digitization project involving tapes in diverse African languages, plus written translations or transcripts in French or English.2 And see also the following excerpt from a recent position announcement, which does not even involve recordings (emphasis in original):
Notetakers will be responsible for capturing detailed and accurate documentation of focus group discussions and interviews in English, translating from Arabic to English in real time. The notetaker is thus the point person for all qualitative data obtained during the data collection phase.
Although skeptical about what seems like translating qualitative data "on the fly," since there is such potential for loss or distortion of information, and less facility in verifying the end product, I do recognize that there can be reasons - perhaps good ones - for such practice.3

The main issues here, however, are first to call attention to these different methods in qualitative research, which might be qualitatively different in their outcomes, and then to make the case for terminology to distinguish between them.

I'll begin with brief discussion of three terms already established in this space - transcription, interpretation, and translation - and then return to comparing these two approaches to cross-language qualitative data. I then consider possible blended terms to refer to the shorter approach.

Transcription


Transcription literally is conveying in writing - "reducing" the spoken language to its written form. By convention it refers to recording in the same language. Ideally a transcription should be verbatim, reflecting the actual words and expressions used.

The way one writes the language is a key consideration, though the general assumption in most cases I have encountered is to use the standard orthography. Transcription may also be phonetic, and this is common for example in linguistic research. There are gray areas between the two, as we found in Afghanistan, where speech in a language may vary in accent or regional variant forms, but which transcribers wrote in standard form - this we felt preserved the sense of what participants said in their languages, while facilitating later text analysis and translation.

The detail of transcription may vary, to the extent perhaps of indicating other sounds and vocalizations in addition to the meaningful speech.

Interpretation


Interpretation is conveying the meaning of a verbal utterances in one language into a verbal utterances in another. Interpretation may be in real time - either sequentially or consecutively, as is commonly done in formal or informal settings where someone interprets for two others who are speaking different languages, or simultaneously.
In any case, the usual assumption is that interpretation is done pretty close timewise to when the initial statement is made - at least when people are speaking directly. When a recording is involved, a spoken interpretation is usually not sought, but a written rendering of the meaning of the recorded speech might be - one of the situations of concern in this posting.

(The terms "interpretation" and "translation" are often used interchangeably, but here the strict definitions will be retained.)

Translation


Translation is the conveying of the meaning of text in one language into text in another. Translation may be more literal/word-for-word or more semantic/meaning focused. In the context of qualitative research, attention to the meaning and of the voice of participants is important.

The field of translation has seen a lot of change in recent years with specializations and the introduction and refinement of tools such as translation memory and machine translation.


From spoken source language to written target language


The context of cross-language qualitative data analysis is nicely summarized by Monique Hennink in her handbook on methods4:
In international focus group research, the group discussions are typically conducted in the language of the study participants, which may differ from that of the research investigators. Therefore, the tape-recording of the discussions will need to be translated and transcribed into the language of the research team for data analysis. The process may involve first transcribing the tape-recording in the language of the discussion and later translating the written document. This process will produce two transcripts: one in the original language of the discussion and a second translated transcript.
One advantage of having the two texts - one the transcript of the recorded discussions and the other the translation of that transcript - is in facilitating back checking of the translation. Another that the source language transcript can be used for text analysis.

Prof. Hennink continues4:
However, time and resource constraints lead many research projects to conduct the tasks of translation and transcription simultaneously, the outcome of which is a single transcript in the language of the investigators, with the tape-recording as the only record of the discussion in the original language.
Building on my previous discussion and illustrations above, here (below) is a quick schema illustrating these two approaches ("source language" here being the original language of the discussions, and the "target language" being that of the researchers, their analysis, and the final reporting or publication).
Two ways to get from spoken source language to written target language: In green, (1) transcription &
(2) translation; or in orange, (١) a direct rendering, for which there is/are not yet any specialized term/s.
What happens in the two-step process of transcription and translation (the green arrows, 1 & 2) is pretty straightforward. The transcription process may run into issues alluded to above in how to reconcile different pronunciations and usages with the standard language and its formal orthography, but these are problems common to transcription as a practice. Likewise, translation has its own set of issues such as whether to be more literal or more semantic.

On the other hand, what happens when the spoken source language data is not first reduced to writing in that language, but rather is rendered directly in the written target language (the orange arrow, with Arabic digit ١), is a process that needs more attention. It is clear, as already mentioned above, that the lack of a transcription in the source language makes verification of the product more difficult, and it also eliminates the possibility of text analysis in the source language.

But what about the process itself, what the person making the written target language product from the recording of the spoken source language? Should we think of that person as interpreting internally before writing, meaning perhaps an alternate two-step process? How does the quality of data processed this way compare with that of the formal two-step process above? Is this translation or interpretation, or should we call it something else? 

"Transterpretation," "interprescription," or ... ?


The process of rendering a recorded discussion in one language directly (one step) into a written record in another seems to be at the same time:
  • similar to that of same-language transcription in that the person doing it would likely listen and re-listen to the recording in order to get it right;
  • similar to interpretation in that they are working from what they hear, not something in writing; and
  • similar to translation in that the product is in written form and as such can be revised and edited. 
It yields a product used in the same way as that produced by transcription followed by translation, but as far as I am aware, there have not been any comparative evaluations of the two. 

Still, since the two processes - the two methods to convert spoken discussions in on language into text in another - are different, it would at least be useful to have different terms to refer to them. One possibility would be to simply call the two-step method "translation," understanding that a preceding step of transcription is involved, and to coin a term for the one-step method. For the latter, two possible ways of blending the terms "transcription" and "interpretation" that would convey the sense of writing down one's interpretation of spoken language, are "transterpretation" and "interprescription" (for the latter, the product would logically be an "interprescript" - an interpretation transcript). Of course there may be better ideas, which would be welcome.

Simultaneous interprescription or transterpretation?


As mentioned above, there is also the other scenario where discussions may be interpreted and transcribed (at least as summary notes) in one step in real time, i.e., without a recording. This is another approach to cross-language qualitative data from focus groups or interviews, which also ought to have an appropriate term to facilitate reference and clarity about methods.

1. These were transcriptions in handwritten Fulfulde photocopied by me around 1990, thanks to Prof. David W. Robinson. My intent was to use them for extracting lexical data for future revision of the Fulfulde lexicon.
2. These were produced by a project run by Nigel Cross and Rhiannon Barker that resulted in a book edited by them under the title At the Desert's Edge: Oral Histories from the Sahel (Panos, London, 1992).
3. As discussed below, limited resources and limited time are given as justifications for not transcribing in the source language. Another circumstance that might arise is where participants are not comfortable with being recorded, so a facilitator may make notes of their interpretation of the discussions (rather than attempt verbatim transcription or notes in the same language that need translating later).

4. Monique M. Hennink, International Focus Group Research: A Handbook for the Health and Social Sciences (Cambridge University Press, 2007, p. 214).

Tuesday, November 19, 2013

Where there is no spellchecker

For a text in a language like Fula that has no spell checker support, here's a workaround that might be more effective than resort to ever more careful copyediting: Break the text down into words and then sort the list into alphabetical order. Better yet, do a word frequency table.

The idea is to align words in a way that facilitates visual checking in a different way. It's especially effective for recurrent words and related words that in the generated list would fall together, but among which misspelled words would be counted separately.

I recently tried this with a text in Fulfulde of Niger and found numerous instances of what appear to be single to double letter errors (doubled consonants and vowels are significant for pronunciation, often with meaning differences where the letter is single), and substitution of plain b, d, or y for ɓ, ɗ, or ƴ(and vice-versa) along with other minor but not unimportant errors. The next step would be to go back to the original text to search the erroneous forms and replace with the correct spelling (either manually or by the search-and-replace function). Not so elegant perhaps, but should be effective.

There are a couple of ways of generating the word list, with the simplest being to substitute hard returns for spaces in the word processor program, then clean out punctuation and quotation marks, and then sort. This can be converted into a frequency list by means of a pivot table in Excel (and perhaps other spread sheet programs.

Another way is to use a text analysis utility software. I used one online at Online-Utility.org.

Ultimately if one is doing a lot of work with text in a particular language, one imagines that lists generated in this way might be useful in building a corpus which could in turn be used for development of a spell-checker.

The way I came about this approach was in incorporating word frequencies as a step in qualitative data analysis (QDA) of text ("where there is no computer-assisted QDA software"). Word and phrase frequencies are of course used in a different way in the latter, but it occurred that breaking down text in this way might also be an aid in comparing the forms of the words themselves (spelling, mainly).

(For those not familiar with the famous basic health and first-aid publication for poor regions of the global South - Where There Is No Doctor - the title of this post is inspired in a very small way by it.)

Wednesday, November 13, 2013

Looking back and looking forward

For those who have read this blog before, it would come as no surprise that there has been a hiatus in posting followed by another post like this, breaking the silence. So with this I'd like to catch up and look ahead.

I last posted over three years ago, just before the International Mother Language Day 2010 - something I've paid attention to over the years. IMLD is also the occasion on which the winner of the Linguapax Award is given in recognition of "actions carried out in different areas in favour of the preservation of linguistic diversity, revitalization and reactivation of linguistic communities and the promotion of multilingualism." In 2013, Africa had another awardee, the Mauritian organization Ledikasyon pu Travayer (education for workers in Morisyen, the French-based Indian Ocean islands creole language of Mauritius).

However, last year, the first two (very distinguished) African recipients of the Linguapax Prize passed away - Neville Alexander of South Africa (Linguapax Award 2008) in July 2012, and Maurice Tadadjeu of Cameroon (Linguapax Award 2005) a few months later in December.

When I last posted on this blog in early 2010, Niger was in a muddle, politically speaking, and Mali was apparently a model; now Niger seems stable and Mali is recovering from a terrible year. I do not plan to spend too much time in this blog on issues relating to governments and conflict, though in some cases such issues will be hard to ignore. However the focus will continue to be on African languages and the "information society," along with related aspects of development and education.

2010

During most of the rest of 2010 I was based in Djibouti, and had the opportunity to follow up on and observe some US military civil affairs projects in northern Uganda and eastern Ethiopia. From the point of view of African languages, what was particularly interesting was to note aspects of training of community animal health workers in Oromo language (''Oromiffa'') in the Harari region of Ethiopia, and in Karamojong (''ŋaKaramojoŋ'') in Moroto, Uganda (my third trip to that country). While English was also used in both cases, the first languages of the trainees (Oromo and Karamojong) were central to learning. (I compiled a list of veterinary and animal terms in Karamojong, cross-checked with several references.)

2011

From late 2010 was back in the US with family again, and focused on different work and home priorities. In 2011 there were two conferences of note that had particular importance for applied work with African languages:
  • Conference on Human Language Technology for Development (HLTD 2011), Alexandria, Egypt, 2-5 May 2011. This was organized by PAN Localisation and ANLoc, with support from IDRC and was hosted by the Bibliotheca Alexandrina. In a sense, this consideration of human language technologies (HLTs; understood to include a range of applications for manipulating and transforming languages) for development is the logical extension of efforts to localize software and internet content. It will be a key area to follow in coming years.
  • Action for Global Information Sharing 2011 (AGIS11), Addis Ababa, Ethiopia, 1-2 December 2011. This was co-sponsored by the Localisation Research Centre and UNECA, along with others. Although technically not the first time for meeting of African language localizers with members of the localization industry, as a smaller scale meeting happened at the 2005 LISA Cairo conference almost exactly 6 years earlier), this was evidently much more significant in scale.
2012

One noted with great interest the efforts of Translators Without Borders (TWB) in early 2012, which included a translation center in Kenya.

In July, I personally had the opportunity to participate in Wikimania 2012 in Washington, DC, including the Tech@State event on "Wiki.gov." On the Wikimania proper side of things, there was a renewal of discussion concerning African language Wikipedias, including some discussion of potential links with a medical translations project (which not surprisingly has connections with TWB).

2013

In 2013 I've been working in Asia for the first time in half a decade, this time in Afghanistan, coordinating survey research. This has obvious multilingual dimensions here, many of which are relevant to multilingual societies elsewhere in Asia and Africa. An aspect I've been particularly interested in exploring is "cross-language qualitative data analysis," which surprisingly (or maybe not so surprisingly, given how language is often a secondary consideration in other areas of endeavor, even when an obvious factor) has only relatively recently gotten serious attention.

Although I have limited time for it, have begun working again with the material from the Fulfulde Lexicon (1993). This entailed converting old files in WordPerfect 5.1 format (not as hard as it might seem, but not straightforward). A major part of the object is to prepare to integrate the material in Kamusi's online platform.

So with that brief retrospective, I'd like to resume but with a slightly different approach here on out - ideally shorter and more frequent posts, pivoting off of items of interest from diverse sources ...