At a glance
• As of September 2025, Unicode encodes 172 scripts and 159,801 characters. The simplest way to make sense of that many is to ask how each script handles vowels.
• Greek and Latin alphabets write consonants and vowels as separate letters; Arabic usually omits short vowels; Devanagari and Ethiopic build vowels into consonant letters; Japanese kana write one syllable per character.
• Hangeul was created in 1443 with 28 letters and made public with its commentary in 1446. Its consonants model the organs of speech and its vowels the principles of heaven, earth, and humanity. Today’s basic set has 24 letters.
1. How Many Scripts Are There, and Why Sort Them
The most practical count of the world’s scripts comes from Unicode, the international standard computers use. On 9 September 2025 the Unicode Consortium released version 17.0, adding 4,803 new characters for a total of 159,801 and bringing the number of scripts to 172. The four new scripts are Beria Erfe, “a modern-use script used by Zaghawa communities in central Africa”; Tolong Siki, “a modern-use script used by Kurukh communities in northeast India”; Tai Yo, “the traditional script of Tai Yo communities in northern Vietnam”; and Sidetic, “an historic script used in ancient Anatolia.” Scripts are not museum pieces. They are living tools that are still being created and standardized.
A list of 172 cannot be memorized, but sorted by how they write sound, the scripts fall into a few families. Alphabets write consonants and vowels as separate letters. Abjads write consonants and largely omit vowels. Abugidas build a default vowel into each consonant letter and change it with marks. Syllabaries write one syllable per character. Logographic scripts use units of meaning as characters. This sorting says nothing about better or worse. It shows the different solutions people chose to fit the sound structure of their languages.
2. The Birth of the Alphabet: From Hundreds of Signs to a Few Dozen
According to the Penn Museum, the Canaanite predecessors of the Phoenicians “had invented the alphabet with the inspiration of Egyptian hieroglyphs, possibly as early as 1800 BCE,” and recently discovered clay tags from Umm el-Marra in Syria may push the date back to roughly 2400–2300 BCE. The alphabet’s decisive advantage was number. Instead of the hundreds of signs of cuneiform or hieroglyphs, “students of alphabetic writing only need to learn twenty to thirty letters.”
The same source notes that the earliest known Greek alphabetic inscriptions date to the 700s BCE, when Phoenician ships sailed the Mediterranean, and that the Latin alphabet used for English and other languages today developed from the Greek alphabet. That Phoenician writing recorded only consonants and that the Greeks reassigned some letters to vowels is a widely taught account, but the museum article checked for this piece does not go into that mechanism, so we record only the outcome here: Greek and Latin became scripts that write both consonants and vowels as separate letters, the type most widely used in the world today.
3. Abjad: Why Arabic Leaves Out Its Vowels
Arabic is the best-known example of a consonant-centered script, an abjad. The Arabic orthography notes maintained by Richard Ishida, who works on internationalization at the W3C, describe the script as an abjad in which “short vowels are not normally written.” It has 28 basic consonant letters and runs “right-to-left in horizontal lines,” although “numbers and embedded Latin text are read left-to-right,” so a single line can change direction. The script is always cursive, and letter shapes change depending on what joins them on either side.
This does not mean Arabic has no vowels. Long vowels are indicated by three letters, alif, waw, and ya, that double as vowel carriers, and short vowels can be written with small marks above or below the letters called harakat. Those marks, however, appear only in “Qur’anic texts, dictionaries, educational materials, and where the pronunciation needs to be made clear,” and are omitted in most everyday print. The reader knows the language and recognizes words from the consonant skeleton. The same script, adapted with extra letters for each language’s sounds, also writes Persian, Urdu, and many African languages, a reminder that scripts and languages do not map one to one.
4. Abugida: The Vowel Lives Inside the Consonant
Devanagari and the other scripts of South Asia belong to the type a Unicode technical note calls “abugidas or alpha-syllabaries.” In this type, the note explains, “consonant characters represent a consonant+vowel syllable,” and each consonant carries an inherent vowel that must be overridden when a different vowel is needed. The pronunciation of that inherent vowel varies by script, with examples such as /ə/, /ʌ/, and /ɔ/. To write a vowel other than the inherent one, a vowel sign, called a matra in Sanskrit, is added to the left, right, above, or below the consonant, sometimes on more than one side at once. To pronounce a bare consonant, a small mark that Unicode calls a virama “kills” the inherent vowel. When consonants cluster, their shapes change or merge; in 60 percent of Devanagari conjuncts, the consonant that loses its vowel also loses the vertical bar historically associated with the inherent vowel, producing so-called half-forms. All these scripts, the note says, descend from a common ancestor, Brahmi.
Africa has its own abugida. The Unicode core specification describes Ethiopic as a script whose syllables “are traditionally presented as a two-dimensional matrix of consonant-vowel combinations,” and its encoding follows that structure: a matrix of 43 consonants crossed with 8 vowels, making 344 conceptual syllables. Ge’ez, the language that gave the script its name, is now limited to liturgical use, but the script has been adopted for Amharic, Tigre, Oromo, and other languages of the region. The same document records scripts created in Africa in modern times: N’Ko, a right-to-left script devised in 1949 for Manden languages; the Vai syllabary of Liberia and Sierra Leone, developed in the 1830s with a standard syllabary published in 1962; and the Bamum syllabary, developed between 1896 and 1910 in western Cameroon.
5. Syllabary and Logographs Together: Japanese
Japanese uses three kinds of characters within a single sentence: kanji, which write units of meaning, and two syllabaries, hiragana and katakana, which write one syllable per character. How many kanji does one need? Japan’s Agency for Cultural Affairs lists 2,136 characters in the Jōyō Kanji Table revised by Cabinet notification in 2010 (Heisei 22), and describes the table as a “guideline for kanji use in general social life.” The jōyō list is not the legal total of all kanji; it is a benchmark of what suffices for official documents, newspapers, and everyday reading.
Each kana set is commonly described as having 46 basic characters that write the syllables of Japanese, with hiragana used mainly for grammatical elements such as particles and endings and katakana for loanwords and similar purposes. One language dividing labor between a logographic script and two syllabaries shows that script types are not assigned one per language; they can be combined as needed.
6. Hangeul: Letters Modeled on the Organs of Speech
Hangeul writes consonants and vowels as separate letters, like an alphabet, but it holds a special place in the history of writing because the record of how and why it was designed survives. Korea’s National Institute of Korean History records that King Sejong created 28 letters in 1443, the 25th year of his reign, and published the Hunminjeongeum, known as the Haerye edition, in the ninth month of 1446. The National Hangeul Museum’s permanent exhibition, “Hunminjeongeum: A Thousand-Year Writing System Plan,” explains that those 28 letters comprised 17 consonants and 11 vowels, that the basic consonants were modeled on the shapes of the speech organs, and that the vowels were derived from the principles of heaven, earth, and humanity. UNESCO inscribed the Haerye manuscript on the Memory of the World Register in 1997, describing it as the book published in the ninth lunar month of 1446 that promulgated the alphabet Sejong had completed in 1443, together with commentaries and examples by scholars of the Hall of Worthies. The original is held by the Kansong Art Museum.
The Hangeul used today does not have 28 letters. According to the Encyclopedia of Korean Culture’s entry on jamo, current Korean orthographic norms set the basic letters at 24: 14 consonants and 10 vowels. Four letters fell out of use as the sounds they wrote disappeared or merged: the arae-a (ㆍ) for a fifteenth-century back low vowel, the yeorin-hieut (ㆆ) for a glottal stop, the yet-ieung (ㆁ) for a velar nasal, and the banchieum (ㅿ) for an alveolar fricative. A script keeps changing after it is created, following the sounds of its language. Hangeul’s 24 letters are the result of that change.
7. Scripts and Languages Are Not the Same Thing
One fact runs through every example above: scripts and languages are different things. The Arabic script writes not only Arabic but Persian and Urdu; the Ethiopic script went on writing Amharic and Oromo after Ge’ez retreated to the liturgy. Conversely, Japanese is one language written with several scripts. So the phrase “the script of such-and-such country” should always be used with care. A script is a vessel for the sounds of a language, and vessels can be borrowed, modified, and newly made.
And they are still being made. N’Ko in 1949, Vai in the 1830s, Bamum in the early twentieth century, and Beria Erfe and Tolong Siki, newly encoded in Unicode in 2025, are all scripts that communities created in the modern era to write their own languages. The assumption that older scripts are more legitimate does not fit this living history. Just as Hangeul carries a documented date and design rationale from 1443, new scripts are born with their own reasons and designs and take their place in the standard.
8. Three Things to Remember When Meeting an Unfamiliar Script
First, do not judge a script’s difficulty or merit by its letter count. Arabic has 28 letters but requires learning vowel omission and joining rules; Devanagari’s core is the combination of marks; kanji has the 2,136-character jōyō benchmark. Each script has simply placed its complexity in a different spot to fit its language.
Second, distinguish the name of the script from the name of the language. Saying “Urdu written in the Arabic script” or “Amharic written in the Ethiopic script” prevents confusion, and it is also why text encoding and fonts are organized by script rather than by language.
Third, check the standards. Whether a script is encoded in Unicode and which script family it belongs to can be confirmed from the Unicode Consortium’s public materials, and national norms, such as Korea’s orthographic rules or Japan’s Jōyō Kanji Table, are published by government bodies. There are only a handful of ways to write sound, but the histories of the people who chose them are inscribed differently in each of the 172 scripts.
Sources
- Unicode Consortium — Unicode 17.0 Release Announcement (Accessed: 2026-09-07)
- Unicode Consortium (Technical Note #10) — An Introduction to Indic Scripts (Accessed: 2026-09-07)
- Unicode Consortium (The Unicode Standard, Core Specification ch. 19 Africa) — Ethiopic and other African scripts (Accessed: 2026-09-07)
- Richard Ishida (W3C i18n) — Script notes — Arabic orthography notes (Accessed: 2026-09-07)
- Penn Museum, Expedition Magazine (Herrmann & Smith, 2023) — The Alphabet (Accessed: 2026-09-07)
- Agency for Cultural Affairs, Japan — 常用漢字表 (Jōyō Kanji Table, Cabinet notification 2010) (Accessed: 2026-09-07)
- National Institute of Korean History (우리역사넷) — 세종어제훈민정음 (Hunminjeongeum) (Accessed: 2026-09-07)
- National Hangeul Museum — Permanent exhibition: Hunminjeongeum, A Thousand-Year Writing System Plan (Accessed: 2026-09-07)
- Encyclopedia of Korean Culture (Academy of Korean Studies) — 자모 (Jamo) (Accessed: 2026-09-07)
- UNESCO Memory of the World — Hunminjeongum Manuscript (Accessed: 2026-09-07)