An Article on Articles. Part One: Definite Articles

Yes, I know, I haven’t written anything for a long time and I am a very evil person.

Now let’s get to the topic: articles. No, this is not about the long boring things you find in newspapers, but about the short annoying things you have to put sometimes in front of words. Some languages have only definite articles, others have only indefinite articles, some have both and others have none. In any case, articles are an interesting problem from a linguistic point of view. Originally they didn’t exist in most languages, but people noticed that when they talk about certain things they confuse them with other things because there’s really a lot of stuff in the world. So they decided to define the things they are talking about, using demonstrative pronouns and stuff like that. Later these words were used exclusively as articles and ceased to be pronouns or whatever they were before. OK, that didn’t sound very scientific, but I think you got the idea.

Many languages like to organize things clearly and they use definite articles as separate words. I am sure this sounds familiar to native English speakers, but it is the case in many other languages such as German, most (but not all) Romance languages, Greek and Hungarian. The English definite article the comes from the Old English masculine demonstrative pronoun se (or þe in the Northumbrian dialect). That’s why some people use ye in order to sound “Olde English”, although the y-part is merely a graphic imitation of the letter þ which was actually pronounced as th in Modern English. Similarly, the German articles der, die and das descend from the demonstrative pronouns der, die and daz, respectively. The articles in most Romance languages are not so different. The Latin demonstrative pronouns ille, illa, illud, illi and illae ‘that’ developed into the Italian articles il, la, i (or gli) and le. Again, ille, illa and the accusative forms illos and illas gave rise to the French le, la and les, to the Spanish el, la, los and las, to the Portuguese o, a, os and as, to the Catalan el, la, els and les and so forth. The interesting thing about Greek is that it is one of the first languages that developed articles. Ancient Greek already had them: ὁ (ho), ἡ (he), τό (tó) in singular and οἱ (hoi), αἱ (hai), τά (tá) in plural. I guess being able to put emphasis on a certain word comes handy for writing philosophy. These articles remain orthographically unchanged in Modern Greek, but they are pronounced differently. Finally, the Hungarian definite article az (a before consonants) comes from… az ‘that’, which, interestingly, still exists in the same form. It is strange that Hungarian, which is an agglutinative language that incorporates in the word almost everything, uses the article as a separate word.

However, there are languages that have a neater approach to grammar. They tend to pack words and particles together. In these languages the definite articles appear not as words, but as affixes. When the affix is placed at the beginning of the word, it is called a prefix. This is the case in Semitic languages such as Hebrew and Arabic. The Hebrew ה־ ha- and the Arabic ال al-, which are placed at the beginning of the word, express definiteness without the need of a separate word for that. Note that the article is attached to every noun and adjective that belongs to the noun phrase. The origin of these morphemes is not quite certain, although some scientists claim that they both descend from the Proto-Semitic particle hal.

In other languages articles appear even more neatly at the end of the word as suffixes. Among them is Bulgarian where the definite suffix can have various forms depending on the gender, number and so on: -ът (-ăt), -а (-a), -ят (-jat), -я (-ja) for masculine; -то (-to) for neutral; -та (-ta) for feminine; -те (-te), -та (-ta) for plural. They come from the Old Bulgarian demonstrative pronouns тъ (tă), та (ta), то (to) and ти (ti) meaning ‘this’ or ‘these’. The principle in Romanian is similar, only the suffixes come from the Latin demonstrative pronouns mentioned above: ille became –le, illum (Accusative form) became –ul, illa became –a, illi became –i, and illae became –le. Other languages with definite suffixes include Albanian, Danish, Swedish, Norwegian, Icelandic, Armenian, Kurdish and Amharic.

Interestingly, most of the world’s languages actually manage to survive without definite articles. Most Slavic languages (except for Bulgarian), most Turkic languages, Persian, Chinese, Japanese and thousands of others live and thrive without ever using a single definite article. Good for them.

However, there are also some languages that don’t have a definite article, but can express definiteness in certain circumstances. For instance, in Serbian adjectives have two forms, one of which is used for indefinite and the other – for definite noun phrases, e.g. nov grad ‘a new city’ vs. novi gradthe new city’. Another interesting case is Turkish, where definite direct objects are in the accusative, whereas indefinite direct objects – in the nominative. For example, Kitap okuyorum means ‘I am reading a book’, whereas Kitabı okuyorum means ‘I am reading the book’.

In conclusion I will give some examples of definite article usage in various languages. The definite articles and affixes are written in bold letters.

Separate word:

English: the old house

German: das alte Haus

French: la maison vieille

Italian: la vecchia casa

Spanish: la casa vieja

Portuguese: a casa velha

Greek: το παλιό σπίτι to palió spíti

Hungarian: a régi ház

Prefix:

Arabic: البيت القديم al-bayt al-qadīm

Hebrew: הבית הישן ha-bayit ha-yashan

Suffix:

Bulgarian: старата къща starata kăšta

Romanian: vechea casă

Albanian: shtëpia e vjetër

Icelandic: gamla hús

No article:

Russian: старый дом staryj dom

Polish: dom stary

Turkish: eski ev

I hope you found something interesting in this post. In the second part I will focus on indefinite articles and (to a lesser extent) on partitive articles.

How Latin Became Italian

Today I am going to write about another of my favorite languages – Italian. It is, I dare say, one of the finest modern languages.

hourglass

After the fall of the Roman Empire, Latin evolved (or degraded, as some people might view it) via Vulgar Latin into a multitude of varieties, which later became the Romance languages. I personally don’t agree with the ‘degradation’ view because I think that when the world changes, everything else must change together with it. I like the Latin language as well as ‘her daughters’, but in a different way. Anyway, during this process many things changed in the structure of the language. Here I am going to focus on the changes in the phonology and in the grammar. The latter became less complex, but also more irregular.

Italian is geographically and linguistically the closest successor of Latin. It is a very melodic and neat language spoken by more than 60 million people in Italy, Switzerland, San Marino, Vatican City, Slovenia and Croatia as well as by many people of Italian descent in the United States, Canada, Australia, Argentina and so forth.

Italian has many dialects, some of which are regarded by some people as separate languages. This is not only a linguistic, but also a political problem and I am not going to discuss it here. This post is concerned only with the literary Italian language based on the Florentine dialect.

After wasting four paragraphs and several minutes from your life in order to explain what I am going to do, I might just do it. So, Italian, unlike other Romance languages, has retained some characteristic features of Latin such as the distinction between short and long vowels. However, there are also many sounds that have changed.

First of all, Italian got rid of most of the Latin word endings. Of course, this has to do with the fact that the Italian dialects lost the Latin case system and the neuter gender, as did most (but not all) Romance languages. However, since Italian preserved the distinction between masculine and feminine as well as between singular and plural, it had to keep some sort of endings after all. The singular forms of nouns and adjectives is probably based not on the nominative forms, but on the ablative. For instance castellum (Abl. castello) ‘castle’ became castello, porta (Abl. porta) ‘gate’ or ‘door’ became… well it didn’t become anything, it remained the same, natio (Abl. natione) ‘nation’ became nazione, ordo (Abl. ordine) ‘order, rank’ became ordine, lex (Abl. lege) ‘law’ became legge and so on. The plural forms of Italian nouns, however, developed mainly from Latin nominative forms, e.g. gatti ‘cats’ from Latin catti (nominative plural of cattus), finestre ‘windows’ from fenestrae (nominative plural of fenestra). Some forms which are now considered irregular in Italian also come from Latin nominative plural forms. This is the case with uomini ‘men’ and uova ‘eggs’. They look like that not because of some evil grammarian who likes to torture language learners, but because they come, respectively, from the Latin forms homines and ova. But in other cases Italian plurals were formed by analogy using the suffix –i, e.g. in castellocastelli, whereas the Latin declension is castellumcastella because the word is neuter.

But that is not the only way endings changed. Graffiti from Pompeii show that the loss of final consonants (the so-called ‘apocope’) in Vulgar Latin had already started in the first century AD. This affected strongly the verb conjugation: cantat ‘he/she sings’ became canta, cantamus ‘we sing’ became cantiamo, cantatis ‘you (pl.) sing’ became cantate and so on. Many other types of words also lost their consonants at the end, e.g. the preposition ad ‘to’ became simply a, the numeral tres became tre, the adverb post ‘afterwards’ became pos in Vulgar Latin and later poi ‘then’ in Italian.

But enough about endings. Many other consonant shifts took place in Italian, and most of them followed a regular pattern, which is often the case with language change. For example, the n in the Latin consonant cluster ns almost always disappeared, so it changed to a simple s in Italian, which is pronounced as /z/ between vowels: mensis ‘month’ became mese; insula ‘island’ became isola; construere ‘to build’ became costruire.

Assimilation is a very common phonological process in the world’s languages. It means that a certain sound is influenced by an adjacent sound and either becomes similar to it (partial assimilation) or becomes the same (total assimilation). In Italian we have great examples for total assimilation. It usually took place within consonant clusters: ct was assimilated to tt, e.g. octo ‘eight’ became otto, nocte (ablative of nox) ‘night’ became notte, lacte (ablative of lac) ‘milk’ became latte; x pronounced /ks/ was assimilated to ss, e.g. saxum ‘stone’ became sasso, maximus ‘greatest’ became massimo, relaxare ‘to stretch out’ became rilassare ‘to relax’; mn was assimilated to nn, e.g. somnus ‘sleep, slumber’ became sonno, autumnus ‘autumn’ became autunno etc.

Another phonological process prominent in Italian is a type of partial assimilation called palatalization. It occurs when a sound is influenced by a neighboring /j/ or a front vowel like /i/ or /e/. In this case the sound’s place of articulation moves closer to the palate. This process began in Vulgar Latin and later developed in the various Romance languages. For example, in Classical Latin the letter c was always pronounced as /k/, but in Italian it underwent palatalization and changed into /tʃ/ before i and e. Because of this Latin centum /kɛntum/ became cento /tʃɛnto/ and civitas /ˈkiːwɪtaːs/ ‘citizenship’ became città /tʃitˈta/ ‘city’. Similar is the case with g, which in Latin was always pronounced /g/, but in Italian it became /dʒ/ before i and e, for example regina /reːˈɡiːna/ became regina /reˈdʒiːna/ ‘queen’. However, these shifts are easy to deduce from the spelling since Italian has kept c and g in these positions. But there are other cases of palatalization, which are more concealed. For instance, d followed by i and a vowel often transforms into /dʒ/ which is spelled as gi, so giorno ‘day’ /dʒɔrno/ actually comes from the Latin adjective diurnus ‘daily’. Similarly, t followed by i and a vowel, which used to be pronounced as /t/, became z (pronounced /ts/), e.g. natione (ablative for natio ‘nation’) became nazione.

Furthermore, the Latin cluster qu /kw/ in Italian changed into ch /k/ in front of i and e, e.g. quis ‘who’ became chi. However, it was preserved in front of a and o (e.g. quattro ‘four’, quotidiano ‘daily’).

Another thing you notice in the development of Italian is that the language obviously can’t stand a cluster of a consonant and an l. In those cases the l is turned into an i. I’ll try to illustrate this with a few examples: flos (ablative flore) ‘flower’ became fiore, pluvia ‘rain’ became pioggia, flumen ‘river’ became fiume, claudere became chiudere, templum ‘temple’ became tempio. However, there are exceptions such as gloria ‘glory’, but they are rare.

Lenition (‘weakening’) is also very important for Romance languages. Although it is more common in the Western varieties (i.e. northwest of the La Spezia – Rimini line), it can be found in Italian as well. Thus voiced plosives developed into voiced fricatives, e.g. /b/ into /v/ as in caballus (‘horse’) > cavallo, and voiceless plosives became voiced plosives, e.g. /k/ into /g/ as in lacus (‘lake’) > lago. However, in most cases Italian preserved the voiceless consonants, for instance rota ‘wheel’ and vita ‘life’ became ruota and vita and the t was kept, whereas in other languages it became voiced as in Spanish rueda and vida and Portuguese roda and vida. French has gone even further and removed the consonant completely, so the words became roue and vie.

A characteristic feature of Italian is that it preserved the so called long consonants. I intentionally use this term because if I say ‘double consonants’ you would think about words like loss or yellow. Well, they are spelled as double consonants, but they are pronounced as single. Many languages have this sort of thing, including many Romance languages, but the difference is that in Italian they are actually pronounced long, like in Latin (so they say). In phonetic transcription this is usually marked with a “:” after the consonant. So passus /pas:us/ ‘step’ became passo /pas:o/ and cattus /kat:us/ became gatto /gat:o/, i.e. the long consonants remained the same.

Image

Dante Alighieri (1265–1321), often described as the ‘Father of the Italian language’ (Image: Wikipedia)

Long vowels are another thing which is present in Italian and not in most other Romance languages. We might then say that Italian has preserved the Latin long vowels, right? And we will be wrong. In fact, the language ceased to distinguish vowel length (probably during the Early Middle Ages), but later Italians decided that it would be a nice idea to make vowels in open syllables long. And I totally agree because it is one of the things that make Italian so melodic. So Italian has long vowels but not the same as Latin. For instance, the vowel e in the Latin word regnum ‘reign’ is long, but in its Italian descendant regno it is short because it is in a closed syllable.

Of course, many other strange things happened to vowels. For instance, the vowels /ɛ/ and /ɔ/ in open syllables were diphthongized, which means that a single vowel changes into a combination of vowels. So /ɛ/ changed into /jɛ/, e.g. ferus ‘wild’ became fiero ‘proud’and vetare ‘to forbid’ became vietare. Similarly, /ɔ/ changed into /wɔ/: novus ‘new’ became nuovo and focus ‘fire’ became fuoco. The same vowels were diphthongized in other Romance languages as well. In Spanish /ɛ/ and /ɔ/ became /je/ and /we/, e.g. fiero, nuevo, fuego, and in French they were developed into /je/ and /ø/, e.g. fier and feu.

Interestingly, the opposite happened as well: Latin diphthongs were merged into single vowels in Italian, like in most Romance languages. So ae (pronounced /ai/ in Classical Latin) and oe (pronounced /ɔi/) became e, e.g. praemium ‘prize’ changed into premio and poena ‘punishment’ changed into pena ‘sorrow, pain’. In addition, au became o, e.g. aurum changed into oro.

Finally, vowels in open syllables were raised, so in those cases e became i, e.g. de ‘of’ > di and fenestra ‘window’ > finestra. In contrast, i and u in closed syllables were often lowered and became e and o: strictus ‘tightened’ > stretto ‘narrow’ and productus > prodotto ‘product’.

OK, enough phonology. Let’s get back to grammar. We already covered the loss of cases and the neuter gender and of various endings. But the Italian language also gained some things, for example the definite articles. You have probably read about this somewhere, but I feel obliged to describe it shortly here as well. Like in other Romance languages, they were developed from the Latin demonstrative pronoun ille ‘that’. In singular, the masculine nominative form ille became ill, pardon, il, after loss of some unnecessary sounds. Similarly the feminine form illa became la. In plural, the masculine form illi became i, and the feminine form illae became le. Now phonology kicks in again. In Italian melody is very important, so things like il italiano, il stato, la arancia and i italiani are not accepted because they don’t sound good. Therefore three other articles were developed for special cases: l’ for words beginning with a vowel (from ille or illa), lo for words beginning with s + another consonant or with z (from illum), and gli for plural nouns beginning with a vowel or with s + another consonant or with z (from illi). Then we get l’italiano, lo stato, l’arancia and gli italiani, which, I must admit, sounds better.

This thing is called grammaticalization. It basically means that a word loses its original meaning and transforms into a grammatical marker. In Italian there is another good example for it – the future tense. In Latin there were actually two future tenses, but people maybe didn’t like them because conjugation was very complex so they got rid of them. Therefore, in most Vulgar Latin the future tense was formed by the infinitive of the verb + the forms of habere ‘to have’. In most Western Romance languages they were later combined into a single word – the future tense we know from Italian, French, Spanish, Portuguese etc. Here is a neat little table that shows you how this happened:

person, number infinitive of the verb scribere ‘to write’ (Vulgar Latin) forms of habere ‘to have’ (Vulgar Latin) infinitive of the verb scrivere ‘to write’ (Italian) forms of avere ‘to have’ (Italian) future tense forms of scrivere (Italian)
1. Sg. scribere habeo scrivere ho scriverò
2. Sg. scribere habes scrivere hai scriverai
3. Sg. scribere habet scrivere ha scriverà
1. Pl. scribere habemus scrivere abbiamo scriveremo
2. Pl. scribere habetis scrivere avete scriverete
3. Pl. scribere habent scrivere hanno scriveranno

I hope you enjoyed this overview of the evolution of the Italian language. I am always fascinated by the way languages change. Maybe you too, if you are a language maniac like me. Do not forget that this is a perverted linguistic view of one of the most beautiful languages in the world. For a more human approach, you may want to visit Bernardo’s Blog My Five Romances. But if you like the perverted linguistic stuff, please, come back here!

DarwinTunes: Survival of the Funkiest

 Image + Image = ?

(Photos: Wikipedia)

I usually babble about languages and their peculiarities, but this time I’m going to write about something more artistic. Well, sort of.

Creating music has often been cited as a manifestation of the individual human genius. I personally am a big fan of the individual human genius, but it has been proved that music could actually do without it. From 2009 until 2012 a team of scientists from Imperial College London led by Bob MacCallum conducted an experiment, which produced several musical pieces using natural selection. The official name of the experiment is DarwinTunes: Survival of the funkiest. Then why am I writing about it now? Because now I remembered about it. Maybe you have heard of this unusual project, but maybe you haven’t.

This is a unique achievement, because the music was composed without… how shall I put it, intelligent design. Of course, music composing software has already been developed, but it utilizes music written by composers such as Bach and Beethoven as a sample.

The principle behind DarwinTunes is much more interesting: online users are asked to listen to random sound combinations (corresponding to genetic mutation in biological evolution) and to rate them. The highest rated of them then ‘procreate’, producing new generations. After the first 500 generations the music started becoming more complex and aesthetically pleasing.

You can listen to samples of various generations here. I recommend the 19 minute long medley because with it you can make yourself familiar with different generations and mutations. I personally think the music is quite pleasant, especially when I know that it has not been composed.

You may say: ‘Oh, but people rated the tracks all the time, so it is a human composition after all!’ No, it isn’t! In this experiment people merely provided the environment in which natural selection occurred. They did not compose the music intentionally and rationally. The music was created by simply surviving among the listeners.

I find this project both fascinating and disturbing. On the one hand, it strikes with its ingenuity to use nature’s mechanisms to create art. On the other hand, it makes us ask the question: ‘Is individual talent going to become useless?’

I hope that original music will survive precisely because it is not natural. Not everyone likes what the majority listens to. However, I think it would be great to develop further the concept of evolutionary music and to attempt to implement the mechanism of natural selection in other fields as well, because… well, because it’s just so cool!

With this remarkably weak argument I leave you to listen to the music and to reflect on this subject.

Quirks of Arabic: Vocabulary

Hello again! Today I’ll be writing about some words that exist in Arabic, although you don’t expect them to, as well as about some words that don’t exist there, although you expect them to.

Although I’ve already discussed word derivation in Arabic in the post about morphology, I’m going to refer to it here as well because the way words are created has a lot to do with the kind of words there are in a certain language. As I explained in my post one moon ago, words in Arabic are formed not only through the addition of prefixes and suffixes, but also of infixes. However, these are not added arbitrarily, but according to a system. In fact, words in Arabic are derived through a finite, albeit large, number of templates. Since I have already given an example of verbal derivation in a previous post, I will now rather focus on nouns. For instance, the three radicals ك k, ت t and ب b are related to writing. Hence are derived the following words: كتب kataba ‘to write’, كتاب kitāb ‘book’, كاتب kātib ‘writer’, مكتوب maktūb ‘letter’, مكتب maktab ‘desk’, مكتبة maktabah ‘library’. As you can see, this method of derivation enables the creation of numerous words. This facilitates the formation of words which can substitute foreign terms. For instance, ‘telephone’ is هاتف hātif (literally ‘caller’) and ‘republic’ is جمهورية jumhūriyah, which comes from جمهور jumhūr ‘multitude’. But the late Libyan dictator Gaddafi has gone further and invented جماهيرية jamāhīriyah as a designation for the state he used to rule. It is a bit like jumhūriyah, but not quite. Gaddafi’s word was derived fromجماهير  jamāhīr, which is the plural form ofجمهور  jumhūr. I really don’t know what he meant by that.

Despite my love for grammar I should stop nagging about derivation and give some actual examples of unusual Arabic words.

In most languages the word ‘still’ is expressed either with an adverb or with an adverbial phrase. One would think that’s the only way. But it isn’t. In Arabic this function is performed by a verb! It is called ما زال mā zāla and is in fact the negative form of the verb زال zāla, which means ‘to disappear’. So if you want to say that you’re still doing something, basically you have to say that you don’t disappear doing that thing.

In the previous post I mentioned a strange word somewhere between a verb and a negative particle – ليس laysa. I personally think that it’s closer to a copular verb because it is conjugated with past tense suffixes (hence a verb) and links two or more things (hence copular). For example, ليست هنا Laysat hunā means ‘She is not here’, ‘she’ is implied by the negation (-at being the suffix for third person feminine singular). I think it is quite useful because it helps you negate a lot of stuff with one word, without using a pronoun or something like that.

Speaking (actually, writing) of ‘to be’ and ‘not to be’, we might think of… no, not of Hamlet, but of a verb that is essential to many languages, especially from the Indo-European family: the verb ‘to have’. It is present in numerous languages, such as English (to have), German (haben), French (avoir), Italian (avere), Spanish (haber/tener), Bulgarian (имам imam), Persian (داشتن dāštan)… OK, I’ll shut up now, you get the point. In Arabic there’s no such thing. Possession is expressed in a different way. Maybe you remember from a previous post that prepositions can be combined with enclitic pronouns. Well, this comes in handy in this case. Arabic uses a prepositional construction with عند ‘inda ‘at’. So if you want to say ‘I have a goat’, you’ll use عندي عنزة ‘indī ‘anzah, which literally means ‘At-me a goat’. It’s not very likely that you’ll ever need exactly THIS phrase, but I must adhere to the old tradition of grammar books and give examples with little practical application. Besides, you never know what life throws at you…

One of the most frequently used words in Arabic is و wa ‘and’. It’s like that in any language, you would say. Yes, but that’s not the whole story. In Arabic و wa is used much more often than its equivalents in other languages. For instance, when listing things in European languages you would rather put a comma between the words or phrases and use and only before the last word, whereas in Arabic all the listed items are connected with wa. It is sometimes combined with other conjunctions as well, for example ولكن wa lākin, literally ‘and but’, which would sound strange in another language. It is often written without space before the next word. Wa is considered a way to achieve a sense of connection and fluency when writing and speaking. A similar role is played by فـ fa-, although it is not exactly a word since it is always used as a prefix.

In different languages there are all kinds of systems for naming the days of the week. Some, like Germanic and Romance languages, are very fond of using names of deities and celestial bodies (e.g. the moon – English Monday / German Montag / French lundi / Italian lunedì / Spanish lunes). In Arabic most of the days are named after numbers, which, in my opinion, is a very logical approach. For instance, ‘Sunday’ is الأحد al-‘aḥad (from واحد wāḥid ‘one’), ‘Monday’ is الإثنين al-‘ithnayn (from إثنان ‘ithnān ‘two’), ‘Tuesday’ is الثلاثاء ath-thalāthā’ (from ثلاثة thalāthah ‘three’), ‘Wednesday’ is الأربعاء al-’arba‘ā’ (from أربعة ’arba‘ah ‘four’), and ‘Thursday’ is الخميس al-khamīs (from خمسة khamsah ‘five’). But not all of them are like that – the word for ‘Friday’ is الجمعة al-jum‘ah, which is derived from جمع jama‘a ‘to gather’. This is related to the fact that in Islam Friday is considered a holy day when people gather at a mosque. The word for ‘Saturday’ – السبت as-sabt is probably derived from a numeral as well (سبعة sab‘ah ‘seven’), but it might also be a loanword from another Semitic language such as Hebrew (שבת shabát), but I don’t think it matters so much because these languages are related anyway. Because of the Aramaic and Hebrew influence on Christian terminology السبت as-sabt is a cognate of the Latin word for ‘Saturday’ sabbatum, which has survived in some Romance languages, e.g. Italian (sabato), Spanish (sábado) and Romanian (sâmbătă). The English word Sabbath as well as the Bulgarian събота săbota and the Russian суббота subbóta also have Semitic origin. OK, this paragraph got out of hand a bit, but I think the topic is fascinating. Maybe it’s just me.

In the Arab world, as well as in other Middle Eastern societies, familial relationships have crucial importance. I believe that this is reflected in the vocabulary of the Arabic language. For instance, European languages usually have only one word for ‘uncle’ (e.g. German Onkel, French oncle, Italian zio, Spanish tío, Russian дядя djádja) and ‘aunt’ (e.g. German Tante, French tante, Italian zia, Spanish tía, Russian тётя tjótja). In contrast, in Arabic there are separate words for uncle and aunt depending on whether they are on the mother’s and on the father’s side: عم amm means ‘paternal uncle’, خال khāl means ‘maternal uncle’, عمة ammah means ‘paternal aunt’ and خالة khālah means ‘maternal aunt’. Of course, such terminological precision regarding relatives is not unique to Arabic since it can be found in various languages. Among them are Persian (عمو ‘amu ‘paternal uncle’, خالو khālu ‘maternal uncle’, عمه ‘ammeh ‘paternal aunt’ and خاله khāleh ‘maternal aunt’) and Bulgarian (чичо číčo ‘paternal uncle’, вуйчо vújčo ‘maternal uncle’, леля lélja ‘paternal aunt’ and вуйна vújna ‘maternal aunt’).

Gentle reader, in case you didn’t fall asleep after reading the last few paragraphs, I’d like to thank you for your attention and patience.

The purpose of the Quirks Of Arabic series of posts was not to serve as learning material, but rather to demonstrate some of the peculiarities and interesting features of the Arabic language to those who are not familiar with it and, hopefully, to awake their interest in this fascinating language. Since these texts were written without any claim for objectivity and comprehensiveness, I shall be happy to read your comments and additions.

Quirks of Arabic: Syntax

Hi there! Sorry for the delay, in the last few weeks I had a lot of exams so I didn’t have much time to write for the blog.

Here comes the next part of my merciless dissection of the Arabic language. This time I’m writing about syntax – the art of slamming words together into sentences.

Classical and Modern Standard Arabic have VSO, which is a relatively rare word order, found also in Celtic languages (See the logic? Well, there isn’t any). This means that in a simple sentence the verb is at first place, the subject comes next, and then you put the object and the other stuff. For example ‘The dog is drinking water’ would be يشرب الكلب ماءً yashrabu l-kalbu mā’an, so something like ‘Drinks the dog water’. In most dialects, however, a SVO word order is used more often.

Another word-order-thing that deserves to be noted is that adjectives come after nouns (unless they, I mean adjectives, are in elative). Also, Arabic uses prepositions. Both things probably have to do with the word order. For instance, ‘the ugly man’ is الرجل القبيح ar-rajul al-qabī, literally ‘the-man the-ugly’.

The so called genitive construction is characteristic for Arabic as well as for other Semitic languages. It is actually a very simple yet clever thing. You see, nouns in Arabic must be defined in some way. There are three variants: a) you can add a definite article as a prefix; b) you can add an enclitic personal pronoun; c) you can define the noun… with another noun. In this construction the first noun remains indefinite (i.e. without an article) and the second is defined in some way (again, with an article, with a personal suffix or with another noun). For example, قنوط المصرفي qunūṭ al-marifiy ‘the banker’s despair’. In Classical Arabic the second (and third and so on) noun is in genitive. However, in Modern Standard Arabic case endings are usually omitted.

In many languages modal verbs are used with the infinitive of the other verbs, e.g. German Ich will schlafen, Italian Voglio dormire, Spanish Quiero dormir or Russian Я хочу спать Ya khochu spat’, all these meaning ‘I want to sleep’.

However, Arabic deals with this differently. It uses a conjunction construction with أن ‘an and the active verb is conjugated in the subjunctive mood. So the sentence used above becomes أريد أن أنام ‘urīdu ‘an ‘anāma (-a at the end is the suffix for subjunctive). Actually, this is found in other languages as well: Bulgarian – Искам да спя Iskam da spya, Romanian – Vreau să dorm, and Albanian – Dua të fle. A similar construction with subjunctive can be found in Persian as well, although a conjunction is not used. The same example used before in Persian would be ميخواهم بخوابم mikhāham bekhābam (be- is the prefix for the Persian subjunctive, but that’s another story). Anyway, this feature is not unique to Arabic, but I decided to include it because I think it’s interesting.

Arabic copular verbs behave in an unusual way. But what are copular verbs? Well, they are verbs that link two or more words or phrases in a sentence. That’s probably a very bad explanation, but I hope the examples might clarify the meaning. For instance, in English to be is considered a copular verb. The same applies to sein in German, essere in Italian etc. In these languages the copular verb not only shows that two or more things are related but also puts them on the same level, so they are technically both subjects, e.g. German Er ist der Dieb ‘He is the thief’ (both nouns are in nominative => both are subjects).

It doesn’t work that way in Arabic. The first element in this sort of sentence is considered a subject, but the others are treated as direct objects and take therefore the accusative. Here’s an example: كان المدرّس تعباناَ Kāna al-mudarrisu ta’bānan ‘The teacher was tired’. Of course, this applies only to the past and the future tense and to the subjunctive mood since كان kāna ‘to be’ is usually not used in the present tense. In Arabic there are other copular verbs as well, e.g. صار ṣāra ‘to become’. This resembles the behavior of Russian copular verbs, for example in the sentence Учитель был усталым Uchitel’ byl ustalym ‘The teacher was tired’. However, here the ‘object’ takes the instrumental case, not the accusative. The accusative is used similarly in sentences with ليس laysa which means something like ‘not to be’. It’s a strange thing somewhere between a negative particle and a copular verb, but that’s another story.

Come to think of it, Arabic uses the accusative in very strange places. For instance, it is used for the subject of subordinate clauses: قال البنت إنّ القطة جميلة Qāla al-bint(u) ‘inna al-qiṭat(a) jamīla(tun) ‘The girl said that the cat is beautiful’. As you can see, the word القطة al-qiṭat(a) is in accusative although it is a subject. But why? The way I see it, the entire subordinate clause is treated as a direct object – it’s the thing the girl said. When you look at it that way, it makes perfect sense. OK, maybe that’s not the official explanation, but it’s how I understand it.

Another striking thing about Arabic grammar is that, as I mentioned before, the verb كان kāna ‘to be’ is not used in the present tense. For example, البحر أزرق Al-baḥr azraq ‘The sea is blue’ literally means ‘The-sea blue’, i.e. the verb is omitted. It actually has present forms – ‘akūnu, takūnu, takūnīna, yakūnu etc., but they are omitted because you don’t need them anyway, they are just equality signs. I think this approach is very practical. I’m not sure whether this is the place to write about it, but the omitting of the copular verb has a significant impact on the way Arabic sentences are structured, so I guess it does count as ‘syntax’. Anyway, this feature is probably not so ‘striking’ since it can be found in other languages as well. I’ll use again the blue-sea-example: Russian – Море синее More sinee, Hungarian – A tenger kék, and Turkish – Deniz mavi. It is usually described as ‘zero copula’ and is used for building nominal sentences.

Languages often use multiple negative particles, but I would say in Arabic this is taken further. There are not just a lot of them, but they perform function usually reserved for other words. For instance, there are several verbal negative particles: لا lā is used to negate the present, لن lan – the future, ما mā and  لم lam – the past (but in different ways). غير ghayr negates only adjectives, and there is also the verb/particle ليس laysa, which I mentioned above. It is used to negate nominal sentences in the present and is conjugated for person, number and gender. Anyway, I’ll probably discuss it in the next post because it has to do with vocabulary as well.

Well, that’s syntax covered. Next time I’m going to focus on the peculiarities of Arabic vocabulary and I’ll be writing about some interesting words. Well, actually, boring words which I think are interesting from a linguistic point of view.

Quirks of Arabic: Morphology

Here is the second post of the series about Arabic. I’ll share some interesting (I hope) features of Arabic morphology. Morphology basically means the way words change their forms in order to give additional information. I’m sure there are much more brilliant definitions than this one, but that is how I understand it.

Arabic has a nonconcatenative morphology. Although this term sounds a bit like a gastroenterological condition, it actually means that words or forms of words are created not by placing morphemes one after another, but by modifying the root itself. In Arabic the root does not consist of a syllable or something like that, but (usually) of three consonants called radicals. They act like pillars among and around which can be put various vowels and consonants in order to make new words or grammatical forms. I’ll try to illustrate that through the following commonly used and unimaginative example where the radicals are k, t and b: kataba means ‘he wrote’; yaktubu – ‘he writes’; maktūb – ‘written’; kātib – ‘writer’; kitāb – ‘book’; kutub – ‘books’ and so on. As you see, vowels change within the root but the three radical consonants k, t and b stay the same.

The morphemes inserted among the radicals are called infixes which are a characteristic feature of Semitic languages. They might be somewhat unusual for speakers of Indo-European languages.

One of the coolest things in Arabic grammar, in my opinion, are the enclitic personal pronouns. They are morphemes that express person, number and (some of them) gender and can perform various functions. For instance, if you attach them to a noun, they work as a possessive pronoun:  كتاب kitāb ‘book’, كتابي kitābī ‘my book’. Combined with verbs, they represent the direct object, e.g. أشاهد ušāhidu ‘I see’, أشاهده ušāhiduhu ‘I see him’. Enclitic personal pronouns can be affixed to prepositions as well, for example ل li ‘for’, لها lahā ‘for her’. They can also express the subject of a subordinated phrase if they are added to a conjunction, e.g. أَنّ anna ‘that’, أنّهم annahum ‘that they’. Of course, they are not unique to Arabic and are present in other languages as well (e.g. Turkish), but they are an important feature of the Arabic language and are used very extensively.

Many languages have only singular and plural or no grammatical number at all, but Arabic has also a curious way to express duality of things (and not only). Arabic, like other Semitic languages, possesses the so called dual number. It is used for nouns, pronouns, adjectives and verbs. It exists in many Arabic dialects as well, but it is not used as often as in MSA. In the case of nouns, adjectives and verbs it is formed through the addition of the suffix ان –ān(i). For instance, بيت  bayt means ‘house’,  بيوت buyūt means ‘houses’, whereas بيتان  baytān translates as ‘two houses’. There is also a common form for accusative and genitive, which is ين–ayn(i),so بيتين would mean ‘of the two houses’. Another interesting example are the dual forms of pronouns such as أنتما antumā ‘the two of you’ and هما  humā ‘the two of them’. However, there is no pronoun for first person dual, just like there are no verb forms for first person dual. Other dual verb forms are present, e.g.  يكتبان yaktubāni ‘the two of them are writing’. In Arabic it is not necessary to use the numeral إثنان ithnān ‘two’ since the duality is implicit by the dual form.

Another characteristic feature of Arabic (as well as of other Semitic languages) is that verbs are conjugated not only for number and person, but also for gender. For example, وصل waṣala means ‘he arrived’, whereas وصلت waṣalat means ‘she arrived’. This peculiarity is considered ‘sexist’ by some people, but I think it is a very exact and informative function of the language.

Some languages use multiple words to express future tense, e.g. English (I will go), German (Ich werde gehen) and Indonesian (Saya akan pergi). In Arabic this job is done in a more synthesized way, through the addition of the prefix سـ sa- to the present tense form. This way an entire three word English sentence is expressed with just one word in Arabic: سأنتظره sa’antaẓiruhu means ‘I will wait for him’. However, the negative forms are a quite different matter. They are accompanied by the negative particle لن lan, which is used only for the future tense, but Arabic negation is a long story and I’ll write about it later.

As I explained earlier, Arabic verbs are derived from a triliteral (or sometimes quadriliteral) roots, consisting of consonants. But that’s not the whole story. From these roots can be derived various stems, ten of which are currently used in Arabic. These stems (or forms, as they are called sometimes) are related to different types of verbs, and the affixes alter the meaning of the verb in a specific way. For instance, كتب kataba means ‘to write’; كتّب kattaba means ‘to compel someone to write’ (causative meaning); كاتب kātaba means ‘to correspond with someone’ (mutual); أكتب ‘aktaba means ‘to dictate’, and so forth.

Languages have different ways to express definiteness. Some, like English, German and Italian have a definite article as a separate word. Other languages, like Bulgarian, Romanian and Amharic, have a definite suffix placed at the end of the word. Some languages like Russian and Persian have no definite article and distinguish definite nouns in other ways. The Arabic language uses the definite prefix ال al-, which can change to ul– or il– depending on the ending of the previous word.

Maybe I have omitted something, but I think these are the most important features of Arabic morphology. I hope that you enjoyed this post. Next time I’ll be writing about what makes Arabic syntax different from that of other languages.

Quirks of Arabic: Introduction, Writing System and Phonology

I decided to start the language series with a few posts about the peculiarities of one of my favorite languages – Arabic. In this post you can read about unusual things about the language’s writing system and pronunciation.

Arabic is a Semitic language spoken by 295 million people (according to the Swedish National Encyclopaedia from 2010) in the Middle East and Northern Africa. There is a debate on whether Arabic is a single language with many dialects or a group of languages. The reason for that is the fact that there is a single literary language but the colloquial language has various regional varieties with big differences between them. In this post I am going to focus on the literary language, the so called Modern Standard Arabic (MSA).

Writing system

Arabic is written with the Arabic alphabet which consists of 28 letters. They are used to represent only consonants and long vowels. Short vowels are usually not written, but they could be marked with special diacritics called ḥarakāt ‘motions’. However, this is done only for learners’ books or in order to avoid ambiguity. Most letters in the Arabic alphabet have four forms depending on their position and are linked in writing.

isolated initial medial final name pronunciation transliteration
ا ا ـا ـا ’alif /ʔ/ or /æ:/ ’ or ā
ب بـ ـبـ ـب bā’ /b/ b
ت تـ ـتـ ـت tā’ /t/ t
ث ثـ ـثـ ـث thā’ /θ/ th
ج جـ ـجـ ـج jīm /dʒ/ j
ح حـ ـحـ ـح ḥā’ /ħ/
خ خـ ـخـ ـخ khā’ /x/ kh
د د ـد ـد dāl /d/ d
ذ ذ ـذ ـذ dhāl /ð/ dh
ر ر ـر ـر rā’ /r/ r
ز ز ـز ـز zayn /z/ z
س سـ ـسـ ـس sīn /s/ s
ش شـ ـشـ ـش shīn /ʃ/ sh
ص صـ ـصـ ـص ṣād /sˤ/
ض ضـ ـضـ ـض ḍād /dˤ/
ط طـ ـطـ ـط ṭā’ /tˤ/
ظ ظـ ـظـ ـظ ẓā’ / ðˤ/
ع عـ ـعـ ـع ‘ayn /ʕ/
غ غـ ـغـ ـغ ghayn /ɣ/ gh
ف فـ ـفـ ـف fā’ /f/ f
ق قـ ـقـ ـق qāf /q/ q
ك كـ ـكـ ـك kāf /k/ k
ل لـ ـلـ ـل lām /l/ l
م مـ ـمـ ـم mīm /m/ m
ن نـ ـنـ ـن nūn /n/ n
ه هـ ـهـ ـه hā’ /h/ h
و و ـو ـو wāw /w/ w
ي يـ ـيـ ـي yā’ /j/ y

Phonology

Arabic phonology has some interesting quirks, especially to speakers of Indo-European languages.

For instance, there are some unusual consonants like /θ/ /ð/ /ðˤ/ /tˤ/ /dˤ/ /q/ /R/ /ʔ/ /ʕ/ /ħ/. The presence of numerous uvular, pharyngeal and pharyngealized consonants suggests that learners of Arabic should make a real effort to train parts of their throat they didn’t know existed.

Another thing that strikes about Arabic is the lack of several consonants that are very common in the world’s languages such as /p/, /v/ and /g/. Yes, they are present in some dialects (e.g. /g/ in Egyptian Arabic), but that’s not the point.

Also, in Arabic there are only three vowel sounds – /æ/, /i/ and /u/. However, they have long variants – /æ:/, /i:/ and /u:/. I think that actually makes perception easier since these vowels have quite different articulations and you can’t mistake /æ/ for /i/ or /u/.

It is also interesting that Arabic phonology does not allow clusters of more than two consonants. Also, when there are two consonants at the beginning of a word, you must insert a vowel in order to facilitate pronunciation. This is called epenthesis. It is used, for example, when building the imperative forms of certain verbs: when you get rid of the personal prefix and the suffix in taktubu, you are left with the gorgeous form ktub, which is not easy to pronounce, so you add a vowel at the beginning and you get uktub ‘write!’ as the imperative form.

Linguistic Diversity

Yeah, yeah, I know, Merry Christmas and so on. Now let’s talk about languages.

Languages have been my hobby, nay, my passion for many years. For me languages are like codes through which you can express the same thing in different ways. 6909 ways, to be more precise (according to Ethnologue). Many people moan about not having a single global language, but I quite like it that way because I have a lot of material to be occupied with.

People are often so used to their language that they cannot imagine that a certain thought could be expressed in a different way. However, morphological structure and word order can vary a lot between languages.

For instance, in a simple transitive sentence involving a subject (S), an object (O) and a verb (V) you can have six types of word order:

–          SVO – the second most common type (Ha! You thought it was the most widespread, didn’t you?), used in many languages such as English, Russian (most of the time), Bulgarian, French, Italian, Spanish etc. Here are some boring examples:

Subject            Object             Verb

The cat            ate                   the fish.           (English)

Кошка             съела               рыбу.              (Russian)

Котката          изяде               рибата.            (Bulgarian)

Il gatto            mangiò            il pesce.           (Italian)

–          SOV – the most common type, used in Hungarian, Turkish, Persian, Japanese etc. Examples:

Subject            Object                         Verb

Kedi                balığı               yedi.    (Turkish)

Ghorbeh          mâhi-râ            xord.    (Persian)

cat                   fish                  ate

–          VSO – characteristic for Semitic languages such as (Classical) Arabic and Celtic languages such as Welsh. Example:

Verb                Subject            Object

ʔakala              al-qittatu          as-samakata     (Arabic)

ate                   the cat             the fish

–          VOS – some Austronesian languages such as Malagasy. Examples (unfortunately I am not very familiar with that language so I am unable to provide any feline or piscine examples):

Verb                Object             Subject

Mamaky          boky                ny mpianatra

reads                book                the student

(i.e. The student reads/is reading a book, not vice versa, although it would’ve been more fun)

–          OVS – as a dominant word order it is present in some American languages such as Guarijio, Hixkaryana, Urarina, but it is possible also in many other languages when you want to put emphasis on the object of the sentence. Here is an example from Russian:

Object             Verb                Subject

Рыбу               съела               кошка.

fish                  ate                   cat

(Again, The cat ate the fish, although it’s the more boring version)

–          OSV – this type is quite rare, but it can be found in Apurinã, for instance:

Object             Subject            Verb

anana               nota                 apa

pineapple         I                       fetch

(I fetch a pineapple)

Some languages such as Hungarian and Russian have a relatively free word order because they have many cases that demonstrate the role of each word so it is not fatal if you change the syntax.

German also has several word orders for different types of sentences, the most common being SVO. However many linguists think that the original order was SOV, although today it appears only in subordinate clauses. This is another one of those you-know-peanuts-aren’t-actually-nuts-but-legumes-things.

Languages differ in their morphological structure as well:

–          Isolating languages – words in those languages contain few morphemes, often even just one. Among those languages are Mandarin, Cantonese, Vietnamese, Indonesian and many others. Some consider English an isolating language although it has retained some verbal and nominal inflection. Here is an example from Indonesian:

Saya    telah    tidur    di         kamar              saya

I           did       sleep    in         room                I

(I slept in my room.)

As you see, every word contains only one morpheme.

–          Agglutinative languages – this is one of the two synthetic language types where words contain more than one morpheme. In agglutinative languages (such as Turkish, Hungarian, Finnish, Tamil, Quechua etc.) morphemes (which carry one meaning) are “glued” together to form words. For me this type is more fun. Here are some examples (the hyphens are used to separate the morphemes):

Oda-m-da                    uyu-d-um.                   (Turkish)

room-1SG-LOC          sleep-PAST-1SG

Alud-t-am                   a          szobá-m-ban.               (Hungarian)

sleep-PAST-1SG         DEF    room-1SG- INE

(I slept in my room.)

–          Fusional languages – again synthetic, but one morpheme can contain more meanings. This type is widely distributed among Indo-European (e.g. Latin, Russian, Lithuanian) and Semitic languages (e.g. Arabic, Amharic).

Dormi-vi                      in         cell-a                            me-a.               (Latin)

sleep-PRF.1SG           in         room-ABL.SG.F         my-ABL.SG.F

Nim-tu                         fî          ghurfat-î.                     (Arabic)

sleep-PRF.1SG           in         room-1SG

(I slept in my room.)

You can see that here more bits of information are expressed through a single morpheme, for instance the suffix –vi means not only “perfect tense”, but also “first person singular”.

Please note that most languages don’t conform entirely to a certain category, but rather lean to it. For instance, English is not absolutely isolating, but it is isolating-ish.

Of course, there is much more to say about language structures, but I think that’s enough for an introduction. Or maybe even too much… Anyway, later I’m going to focus on some interesting (predominantly grammatical) peculiarities of various languages.

After I’ve bored you to death I wish you good evening or night or whatever you’re having there. Expect more linguistic drivel on the blog! Bye!

Hello!

Hello! Welcome to my blog!

My name is Damyan Lissitchkov (Bulgarian: Дамян Лисичков). I just decided to write about stuff I’m interested in. May be some lost soul will read it. May be not.

A few things about me: I am from Sofia, Bulgaria, but currently I study linguistics and areal studies in Berlin. I’m mainly interested in languages, science, history and various strange things that we have on our ridiculous small planet.

In this blog you can read about interesting features of various languages, about things that aren’t the way we think they are, about strange things, about things that are strange, but we think they aren’t (which is strange) and about things that annoy me (which are quite a few).

I’ll be happy to read your comments.