Skip to content

Architecture

Structure

This chapter describes the architecture that the writing systems of this document share. The document treats Hudum (MNG), Todo (TOD), Sibe (SIB), and Manchu (MCH), together with their Ali Gali extensions; a writing system is the set of conventions by which one of these languages writes the script, and a font may serve one writing system or integrate several into one, as Single-font implementation treats. The rules of each writing system are in the chapters that follow: Hudum, Todo, Sibe, Manchu, and Hudum Ali Gali, Todo Ali Gali, Manchu Ali Gali.

The specifications of this document concern the first three layers of a five-layer system of text processing:

Character Layer. The code charts of the Unicode Standard and ISO/IEC 10646 fix each character by code point, name, and annotation. The characters of this document are those of the code charts (see UTR #54).

Glyph Layer. The Unicode Standard and its supplementary specifications give the characters their behavior—character properties such as general category and cursive joining type, and algorithms such as normalization, collation, line breaking, and vertical text layout—and so associate characters with glyphs, subject to context and localization.

Glyph Image Layer. The shaping engine and the font project carry out the association of characters with glyph images, their substitution and positioning, and the shaping operations themselves. How a font builds its glyph images and its rules is its own design decision, so this document does not prescribe it; but the shaping steps it gives can be transcribed into the rules of a font project, and so serve as a reference for building or converting one.

Typesetting Layer. The typesetting application calls the shaping engine with the font files to render the glyph images on the page.

Physical Layer. Devices such as monitors and printers turn the typeset pages into physical objects.

The model of this document turns on two of these layers, the character layer and the glyph layer, and on the two directions between them: text representation, which runs from a written unit to the characters that represent it, and text shaping, which runs from characters to the written units they render. The history behind this model is in Background, the phonemes behind the characters in the Phonology appendix, and the relation of its characters and variants to the data of the Unicode Standard in Relationship to the Unicode Standard.

Character set

This section treats the character layer: it fixes the characters of a writing system—the letters it uses and the format controls that take part in shaping—and with them text representation, how its written units are represented by characters. The correspondence is not one-to-one—a letter may be rendered by more than one written unit, and a written unit may represent more than one letter—so it is shaping, not the character set alone, that decides which written unit a letter renders. The characters fall into Mongolian-specific characters and characters shared with other scripts, grouped in the table below; the sections that follow define the letters and the format controls, and the punctuation, digits, and other characters outside the cursive letters are in Non-joining characters.

ScriptType of charactersExamplesNote
GeneralSpacespace
Punctuationmiddle dot, …
Format controlsZWJ, ZWNJ, …participate in shaping
Digitsdigit one, …
MongolianPunctuationbirga, …less used now
Format controlsFVS, MVS, …participate in shaping
DigitsMongolian digit one, …less used now
Phonetic lettersMongolian letter a, …participate in shaping
CJKPunctuationquestion mark, …

Phonetic letters

The characters of a writing system are determined by a phonemic analysis of the language that writing system records. The phonemes of the four languages this document treats—Tibetan, Sanskrit, Mongolian, and Manchu—are listed in the Phonology appendix.

Format controls

Format controls represent no phoneme; they control the shaping of the letters around them and have no visual appearance of their own. They belong to the character layer, and their effect takes place in the shaping process.

Zero Width Non-Joiner (ZWNJ), Zero Width Joiner (ZWJ), and Nirugu. U+200C and U+200D are Unicode’s standard cursive joining controls: invisible characters that override automatic shaping, ZWJ requesting a more connected rendering of its neighbours and ZWNJ a less connected one, breaking a cursive connection or a ligature (Section 23.2, “Layout Controls”). In Mongolian they can also select a letter’s positional form in isolation or override its expected form within a word: the isolate, initial, final, and medial forms of the letter a (U+1820), for instance, are selected by U+1820, U+1820 U+200D, U+200D U+1820, and U+200D U+1820 U+200D. Being invisible, ZWJ also breaks interaction such as ligation between the characters it stands between, and neither belongs on a common keyboard layout, since everyday text does not need them.

U+180A MONGOLIAN NIRUGU behaves exactly like ZWJ but is visible as a piece of stem stroke, and it is the control to use for joining in everyday text—for example, to end a patronymic abbreviation that is the initial syllable body or the initial consonant of the father’s name. In the Unicode Standard it acts as a stem extender, lengthening the stem that joins the letters of a word; that stretching belongs in the font rather than in the user’s hands (see IIb). The nirugu may also join the two parts of a compound word, where it resembles a nonbreaking, attaching hyphen. The three joining controls act at the cursive-joining step of the shaping process.

Vowel Separator (MVS) and Narrow No-Break Space (NNBSP). MVS is a Mongolian-specific format control, the one that requests the chachlag variation, transcribed as · (a middle dot). In the Unicode Standard it marks the break between a word-final letter a or e and the rest of the word: it separates the two parts—the a or e after it is part of the stem, not a suffix—and always selects the leftward-tail form of that a or e, and it may also affect the form of the preceding letter (see IIb). Neither Todo nor Manchu nor Sibe needs it. NNBSP, transcribed as – (an en-dash), is a whitespace and format control for particles: before Unicode 16.0 it represented the small gap before a separated suffix, and it keeps the Script_Extensions value “Mong” for backward compatibility, but its role has passed to the MVS, which both prevents a break before the suffix and triggers the shaping the suffix needs. NNBSP is discouraged in favour of the MVS, because it sometimes shapes anomalously. The MVS acts at the reduction step of the shaping process.

Free Variation Selector (FVS). FVS’s are Mongolian-specific format controls that immediately follow the letter they modify, with no visual appearance of their own; a base letter and a following FVS form a variation sequence. They are needed only for a form the context cannot predict—a foreign word, for example—and most running text does without them. There are four, FVS1–FVS4 (U+180B..U+180D, U+180F); a variant that the rules cannot predict is selected by an FVS at the reduction step of the shaping process.

Standardized Variation Selector (VS). VS’s are Unicode’s standard controls for requesting glyph variants: combining marks of combining class zero, default ignorable, which with the base character they modify form a variation sequence whose effect depends on the sequences defined for it (Section 23.4, “Variation Selectors”). From Unicode 17.0, VS3 (U+FE02) requests the Sibe form of the quotation marks (L2/25-028); the marks themselves are in Non-joining characters.

Shaping process

This section treats text shaping, the direction from the character layer to the glyph layer: how a sequence of characters becomes the written units it is rendered with. It builds on the shaping that general and cursive scripts already require, and inserts a Mongolian-specific phase of its own.

A cursive script shapes in a fixed order: a basic character-to-glyph mapping (phase Ia), cursive joining, which gives each letter its joining position (phase II), and typography, which closes the process (phase Ib). In Mongolian a letter whose joining position is fixed still usually has more than one written unit, so before the process varies within a written unit, each phonetic letter is reduced to the one it uses. That reduction is the Mongolian-specific phase III, inserted between joining and variation because it takes the joining position as its input and decides which written unit the later variation applies to. Shaping therefore runs Ia → IIa → III → IIb → Ib, as below. The conditions that drive it are the per-writing-system data of the Toolchain section, and the rules each writing system realizes are in the chapters on Hudum, Todo, Sibe, and Manchu together with their Ali Gali extensions.

Shaping phaseShaping step

Ia. General

Basic character-to-glyph mapping

IIa. Cursive script

Initiation of cursive positions

III. Mongolian-specific
Reduction of phonetic letters to written units

PhoneticChachlag
Syllabic
Particle
GraphemicDevsger
Post-bowed
UncapturedFVS-selected

IIb. Cursive script (continued)
Sub-written-unit variations

Variation involving bowed written units
The form a written unit takes before the mark that ends a syllable
Cleanup of format controls
Localized treatments
Optional treatments

Ib. General (continued)
Typography

Vertical forms of punctuation marks
Optional treatments

Ia. Basic character-to-glyph mapping

The basic character-to-glyph mapping (phase Ia) is the mapping every font carries, typically in the TrueType/OpenType table cmap. The Unicode representative glyphs may serve as the default glyph of the phonetic letters, but they do not decide the final rendering: a Mongolian letter has no single glyph, and the phases that follow decide which written form it renders.

Phase Ia also carries the normalizations that must hold before shaping begins. The NNBSP is one of them: its function has been taken over by the MVS (see Format controls), so an NNBSP in text is replaced with the MVS here, and text written with it is shaped by the same rules. The other format controls pass through phase Ia unchanged.

IIa. Cursive joining

On top of the mappings of general scripts, complex scripts insert shaping phases between the basic mapping and typography. Cursive scripts undergo cursive joining (phase IIa), which runs before variant selection, because the written unit a letter takes depends on whether its two sides join—that is, on its cursive position.

Cursive joining. Either side of a written form may join the neighbouring written form or not, so each written form is in one of four cursive positions:

  • Isolated, abbreviated as isol: not joined forward (above, in Mongolian), not joined backward (below, in Mongolian);
  • Initial, abbreviated as init: not joined forward, joined backward;
  • Medial, abbreviated as medi: joined forward, joined backward;
  • Final, abbreviated as fina: joined forward, not joined backward.

Cursive positions are irrelevant to word boundaries, though in Mongolian they usually agree with word-wise positions, since cursive breaks within a word are limited. Joining is normally derived from the letters around a letter, but the format controls override it: ZWNJ prevents joining on one side, and ZWJ and the nirugu force it.

Implementation. Cursive joining assigns each letter the position its two sides produce, and the nominal glyph is mapped at that position to the letter’s default written unit—the one it would use if no later phase changed it. ZWNJ, ZWJ, and the nirugu change the joining of a side, and with it the assigned position. Phase IIa therefore leaves every letter as the glyph of its default written unit at its joining position, which is where phase III starts.

III. Reduction of phonetic letters to written units

Once the cursive position of each letter is fixed, the letter usually still has more than one possible written unit; phase III reduces the phonetic letter to the one it must use. The reduction is predictive—decided by the letter’s orthographic context—so the ordinary case needs no user intervention.

How the reduction is expressed. The reduction is a sequence of conditional substitutions in a fixed order. Each substitution tests the context of a letter—the letters around it and the format controls between them—and, when the context matches, replaces the default written unit that phase IIa produced with the one the context requires. The contexts are expressed as classes of letters and the conditions a writing system defines on them, so the reduction is driven by the per-writing-system data of the Toolchain section; the writing systems differ in just these data.

Phase III runs a series of Mongolian-specific steps. The phonetic conditions are chachlag (requested by MVS), syllabic, and particle; the graphemic conditions are devsger and post-bowed. The phonetic steps run first, because a graphemic condition is tested against their outcome: a letter that a phonetic step has changed is not offered to the graphemic steps in its earlier form. Within a step there may be several sets of non-overlapping rules, one for each group of letters, so that at most one rule applies to a letter. Forms that the phonetic and graphemic conditions do not capture fall to the last step, FVS-selected.

The format controls within the reduction. The format controls have no visual form, and the reduction treats them as states rather than characters. ZWNJ, ZWJ, and the nirugu have done their work in phase IIa and are neutralized in phase III, so they neither render nor interrupt the rules. MVS and FVS take an active part: the MVS is resolved into the chachlag or particle it indicates (see Format controls)—as a chachlag it takes no advance width, as a particle the width of a Mongolian space—while an FVS is read as the request for the variant it names and removed once that variant is selected. When the reduction ends, the controls it consumed are gone, and nothing of the shaping machinery remains in the glyph layer.

Conditions that reach across a word. A condition is usually decided by adjacent letters, but not always: in Hudum the harmonic gender of a word is fixed by its vowels, and a g or h must agree with it even where no vowel stands near. Local substitutions cannot see that far, so the reduction carries the property across the word with markers—inserted at the letter that fixes it, passed across the letters that do not change it, and read at the letter it conditions—and removes them once that letter is decided, leaving no trace.

FVS-selected. When the context matches no predictive condition, the written form is requested with an FVS. Within a writing system an FVS switches between the variants of a letter at a joining position; it does not name one glyph for every context, which is how it differs from a standardized variation sequence.

IIb. Sub-written-unit variation

After the written units are fixed, cursive shaping resumes where phase IIa left it, with the variation within a written unit. The reduction of phase III and the variation of this phase both act on written units, at different scales: the reduction chose a written unit among the variants of the joining position, and the variation now adjusts it in the light of the surrounding written units. A bowed written unit may first change the form of a following vowel, a written unit may then take the shape it is drawn with before the mark that ends a syllable, the format controls the earlier steps consumed are cleaned up, the designs a writing system draws differently from the writing systems it shares characters with are localized, and the optional treatments are applied.

The form before the mark that ends a syllable. A written unit may be drawn with a shape of its own where it stands before the MVS that ends the syllable a mark separates, and in Hudum the final forms of n (U+1828) and g (U+182D) have one—the letters the chachlag onset of phase III selects. The mark is the narrow MVS: the same context that writes the chachlag a or e narrows the mark to that shape, and it is what the step asks for. A wide MVS, which is the separator the cleanup of format controls splits in two, is not a context for the shape. The step acts here rather than within the condition that selects the letter, because that condition acts in phase III, before the FVS is consumed: a substitution made there would take the written unit out of the class the FVS is shaped against, and the FVS would be left in the text unshaped.

Localized treatments. A shared character may be drawn with a different design by each writing system: the final form of m (U+182E), for example, has a small tail in Hudum and a large one in Sibe and Manchu. The written unit is the same and only its design differs, so the font keeps one design as the default and gives the others theirs here, as Single-font implementation treats. The localized form is selected in rclt, not locl, because an engine applies locl before cursive joining, when the written unit to replace does not exist yet.

Optional treatments. The last step carries out the treatments a font may or may not take. Stretching the stem where a bowed written form is followed by an extending one is one of them: the nirugu is the stem extender of the script, and a font that stretches the stem in the shaping keeps the user from inserting U+180A by hand.

Ib. Typography

The process ends with typography. The vertical forms of punctuation marks are critical to setting Mongolian text but are not part of the complex shaping between letters and format controls; a mark drawn over the written form it follows rather than beside it is anchored here, and optional treatments may follow. The marks themselves are in Non-joining characters.