Background
The standards
The ISO/IEC 10646 and the Unicode Standard have supported the Mongolian script for approximately two decades, providing code charts (see UTR #54) that delineate the fundamental information concerning characters and their variants. What they did not provide is a systematic account of how the script is shaped, and the standardization of Mongolian text processing has therefore progressed through several stages, each of which had to make up for what the one before it left open.
Shaping rules before the national standards
Following the encoding of the Mongolian script, the ISO/IEC 10646 and the Unicode Standard merely catalogued its characters and their variants. These lists turned out to be incomplete and, in places, in error, and neither standard described how a sequence of characters is shaped into a sequence of correct variants. The Chinese national standard of the time was no more specific.
Two consequences followed. First, a product that underwent acceptance testing for character-set conformance had to consult an unpublished document, maintained internally and constantly updated. Second, vendors, left to their own interpretation and to specifications of their own making, developed shaping rules that were mutually incompatible, and Mongolian text rendered inconsistently across platforms.
The Chinese national standards
The fragmentation was addressed by a joint effort. Several Chinese organizations launched the Project for the Development and Sharing of Digital Resources on the Mongolian Language and Script, with the standardization of shaping rules as its core focus. After extensive coordination, multiple rounds of validation, and reviews, the new Chinese national standard GB/T 25914—2023 (Chinese version, English version) was released, and revisions to the national standards for Todo, Sibe, and Manchu are under way in line with it.
The standard text, however, is written for testers rather than for implementers. It uses the Extended Backus-Naur Form to describe the conditions under which each variant of a character is selected, listing the expressions flatly, without a hierarchy and without an order of application. A developer has to work out that order for themselves, and easily overlooks or misreads a requirement, which raises the cost of compliance in both development and testing. Interpretive material is needed, and this document is written to provide it.
What the data structures cannot hold
The Mongolian script combines several writing systems. Some characters are shared between them and some are not, and a shared character may share some of its variants with another writing system while using others. Shaping rules are therefore specific to each writing system, and the data structures of the Core Specification, the Unicode Character Database, and the code charts cannot hold them.
The datasets and the toolchain
This document provides more than the rules. It maintains standardized datasets and an automated toolchain that support the automatic construction and testing of fonts and their maintenance over time, meeting the needs of font development, of standards-conformance testing, and of long-term version maintenance. These data are complex and are not taken in at a glance, so this document also explains them; Toolchain describes how they are organized and used.
The analysis of the script
Two ways of analyzing the Mongolian script have historically been in tension, and the choice between them decides what a character is taken to be—and therefore what the encoding, and this document, are about.
The graphetic model
The graphetic model determines the graphemes of each writing system by comparing the Mongolian script, chronologically, with the Old Uyghur and Sogdian scripts, and identifies each grapheme as a character. Its written units carry a simple shaping logic of their own, for which reason the model is in essence the cursive model of the Arabic script, merely set vertically.
The phonetic model
The phonetic model groups together the glyphs that record the same phoneme of the written language and identifies each phoneme as a character, a phonetic letter. This is the model the user community came to accept, and it did so for two reasons. First, by the time the script had evolved into Mongolian, each phoneme of Classical Mongolian appears to be reflected in the text, and it is natural, though mistaken, to regard the script as alphabetic on that account. Second, the education that shaped the user community is built on just such an alphabetic analysis, which ultimately goes back to the works of the Mongolian scholar Shadavyn Luvsanvandan and formed the community’s present understanding of Mongolian writing. From Classical Mongolian that understanding extended to the other writing systems, and thereby to the script as a whole.
The phonetic model is also the one the standards have come to share. L2/24-180, “Proposal to refer to UTN #57 for implementing the Mongolian script”, states that the Unicode Standard should be aligned with the Chinese national standard, which is the de facto international standard for Mongolian text, and this document is written to that end. The phonemes behind the letters of each writing system are listed in the Phonology appendix.
The views of variation
With the characters of Mongolian understood as the phonetic letters of Classical Mongolian, and the characters of each other writing system as the phonetic letters of its written language, the history of the script shows three views of the relation between a character, its variants, and the variation selectors. Each of them left its mark on the data of the Unicode Standard, and the last is the one this document realizes.
Variants in the character layer
This is not so much a view that the Unicode Standard adopted as an impression that the format of its data gives. The Mongolian variants are kept in StandardizedVariants.txt, in the same format as the standardized variation sequences of other scripts, where a base character and a variation selector name a glyph that is chosen freely; UTR #54, “Unicode Mongolian 12.1 Snapshot” maintained the variants in this way. The format reads as if the variants of a character belonged to the character layer and were free variants of writing: a character is first turned into another character by a variation selector, and that character is then turned, by cursive joining and the rest of shaping, into the joining variants that are its glyphs in the glyph layer. For example, the file records
1820 180B; second form; isolate initial medial final # MONGOLIAN LETTER Ain which 1820 180B is a variation sequence: 1820 is the base character that stands for MONGOLIAN LETTER A, and 180B the variation selector, and the description “second form” belongs to the character layer. The sequence then has the four joining variants isolate, initial, medial, and final, which belong to the glyph layer. As the Hudum chapter shows, however, the four joining variants of 1820 180B do not share a common abstract form—their written units are A, A, AA, and Aa. The joining variants of a Mongolian character are not such free variants of writing, and the impression does not hold.
Variation selectors as toggles
The earlier Chinese national standard, GB/T 25914—2010 together with the User Agreement, organizes its description around the user who enters Mongolian text from a keyboard, and it assumes two stages. In the first stage the user types the phonetic letters, and text shaping—in part the cursive joining, in part the requirements of orthography—first produces an initial result. In the second stage, when the result is not the one the user expects, the user types a variation selector after the letter to adjust the glyph, switching the wrong glyph to the intended one.
A variation selector is therefore not used to select a specific glyph but to switch from one glyph to another, which has two consequences. First, the selector that switches glyph A to glyph B is very likely the same as the one that switches B to A. Second, if VS1 toggles between A and B and VS2 toggles between C and D, then A cannot be switched to C or D. This view was abandoned. Its concrete steps were left vague by the national standard of the time, and the User Agreement was not public, so that different manufacturers in fact implemented different behavior and text rendered inconsistently across platforms and vendors; in resolving this inconsistency the view as a whole was given up.
The current approach
The current approach changes the role of the variation selector. On the one hand, every variant of a character at each joining position is bound to a dedicated variation selector, which makes the selector a true selector, a determiner. Because more selectors are needed than under the earlier views, an additional selector, FVS4, had to be encoded (L2/20-057). On the other hand, the current Chinese national standard describes, for each character, the conditions under which each of its glyphs is used; it does so formally and provides test data (the eac- files, for example eac-hudum.json, in this repository), but it does not prescribe concrete step-by-step shaping steps, leaving that latitude open mainly so that manufacturers can adapt. This document therefore proposes an approach that satisfies the current Chinese national standard. Its main structure was proposed by Liang in L2/19-368, and we have improved it and made the changes necessary to meet the Chinese national standard.
Because the variant sets and the selector bindings differ between the writing systems, the data of this document are organized per writing system; the datasets and the tooling that maintain them are described in Toolchain.
The data of the Unicode Standard
The Unicode Standard’s own data changed along the same lines, from one block-wide set to none. Until Unicode 13.0 the code chart carried, in addition to the representative glyph of each letter, the glyphs of its positional forms (since Unicode 9.0) and of its standardized variation sequences (since Unicode 7.0), and the Unicode Character Database listed those variation sequences in StandardizedVariants.txt; both belonged to the analysis of the views described above. Unicode 13.0 removed the positional-form and variation-sequence glyphs from the code chart, and UTR #54, “Unicode Mongolian 12.1 Snapshot” was published to preserve the last chart of that design, so that the snapshot and the standardized variants in the database record the same earlier state. The Standard is now discontinuing the maintenance of Mongolian variants as standardized variants: the entries are being deprecated and the surrounding text updated (L2/26-091), with a transitional update of the sequence list for Unicode 18.0 (L2/26-203); the cross-writing-system analysis underlying this decision is presented in L2/26-207. The analysis that this document realizes in place of those data is specified by the Chinese national standards for Hudum, Todo, Sibe, and Manchu, and is documented informatively here.