Skip to content

Relationship to the Unicode Standard

The Unicode Standard fixes the characters of the Mongolian block by code point and by name, and its code charts give each character one representative glyph. The variant data it carries belong to the earlier, single-block analysis recounted in Background: the positional and variant glyph forms of a letter are presentation forms that are not separately encoded, and a base letter followed by a free variation selector is a standardized variant. This document replaces that analysis: here the characters of the script are the phonetic letters of each writing system, and their variants are the written units that the writing system uses at each joining position.

This document provides a detailed description of the characters and their behavior in the Mongolian script as defined by the Unicode Standard, lists the variants of each Mongolian character, and explains the process by which a sequence of Mongolian characters is converted into a sequence of glyphs. Note that UTR #54 provides only a historical version of the Code Charts for the Mongolian block: its variant set has changed significantly, so its variant descriptions and FVS assignments should not be used as a reference during implementation.

This document is informative and has no normative status. The analysis of the script that it realizes is the one that the Chinese national standards specify, as Background recounts. The relationship to the Standard and the relationship to those standards are therefore one, not two—where the Standard only describes the encoding of Mongolian, this document states what that encoding prescribes for the characters, their variants, and their shaping; where the Chinese national standards state a requirement and leave the implementation open, it defines the implementation of the encoding and the related data, with guidance on supporting them. It definitively interprets several of those standards, and it is updated promptly as they are revised.

The characters

The code chart for the Mongolian block shows each character with its code point, its name, and one representative glyph; it shows neither the positional forms of a letter nor its variation sequences. The chart is consistent with the character layer of this model: it fixes a character by code point and name, and its representative glyph is not prescriptive. The characters that the subsequent chapters specify correspond to those of the code chart. UTR #54, “Unicode Mongolian 12.1 Snapshot” is the last chart that showed the positional variants and variation sequences of the earlier design.

The variants

StandardizedVariants.txt still lists the Mongolian standardized variation sequences, each a base letter followed by a free variation selector. Its comments for the Mongolian block state that only the free variation selectors FVS1–FVS4 (U+180B..U+180D, U+180F) are used, and that the generic variation selectors are not; that the per-sequence descriptors are arbitrary, numbered labels with no systematic relation to the shapes of the glyphs; that a sequence labeled “not in use” is no longer recommended for use, because it was defined for earlier implementations and is retained so that legacy data remains legible; and that any unlisted combination is unspecified and reserved for future standardization.

The file belongs to the earlier model whose data structures are being retired. Its entries do not describe the variants of this document: the same descriptor does not describe the same written unit across writing systems, a descriptor may name a variant that current orthography no longer uses, and Mongolian variants do not behave as standardized variants, for the following reasons.

  • Variants are orthographically required, not free alternatives. At a given joining position the variant of a letter is fixed by the writing system and the surrounding letters, and it may distinguish two words. A variant that is needed by orthography is not a free variant.
  • A free variation selector switches, rather than determines. FVS1–FVS4 switch between the variants of a letter at a joining position; they do not name a single glyph that a letter must use in every context.
  • The variant set of a letter depends on the writing system. The Mongolian block unifies letters across Hudum, Todo, Sibe, and Manchu and their Ali Gali extensions. These writing systems do not agree on the variant used at every joining position, and a writing system that has no form at a position must render a fabricated form derived from another position. A writing-system-neutral list cannot express these differences.
  • Variant selection takes effect in shaping. It operates after the orthographic shaping steps, not at the character-to-glyph mapping stage where standardized variants apply.

Compared with the preceding version, the StandardizedVariants.txt updated for Unicode 18.0 removes the following entries, each for the reason given.

  • 1820 180C (third form), medial. It is an initial variant, so it is moved to the initial of 1820 180B (second form).
  • 1835 180B (second form), medial. It is an isolated variant, so it is moved to the isolated of 1835 180B (second form).
  • 1848 180B (second form), medial, is changed to 1848 180C (third form), medial. The Todo shaping Chinese national standard added a variant on the second form and then removed it in the revised standard, so the FVS of this form changed.
  • 185E 180B (second form), final. It is a contextual variant that appears in a ligature, not a standalone orthographic variant. Likewise 185E 180C (third form), final, which is part of a ligature and is not a standalone orthographic variant.
  • 1887 180B (second form), isolated. It is a stylistic variant of the isolated of 1820 180B (second form). Its final is a contextual variant of the final of 1820 180C (third form) in a specific ligature. 1887 180C (third form), final, is a contextual variant of the final of 1820 180B (second form) in a specific ligature. 1887 180D (fourth form), final, is analyzed as the medial of 1820 (first form) followed by the final of 1887 (first form).
  • 1888 180B (second form), final. It is used in Manchu Ali Gali and is merged into the final of 1873 180B (second form).
  • 188A 180B (second form), initial and medial. They are contextual variants in a specific ligature.