Duplicate characters in Unicode: Difference between revisions

Browse history interactively

← Previous edit

Content deleted Content added

VisualWikitext

Revision as of 12:38, 2 May 2007 edit Dbachmann (talk \| contribs) 227,714 edits →See also ← Previous edit		Latest revision as of 23:18, 28 December 2024 edit undo Symbol & Font Hunter (talk \| contribs) 390 edits →List Tag: Visual edit
(111 intermediate revisions by 59 users not shown)
Line 1: {{Short description\|Unicode characters that have been encoded twice}} [[Unicode]] has a certain amount of [[duplication]] of [[character (computing)\|characters]]: these are pairs of single Unicode codepoints that are [[canonically equivalent]]. The reason for this are compatibility issues with legacy systems. {{More citations needed\|date=March 2022}} [[Unicode]] has a certain amount of duplication of [[character (computing)\|characters]]. These are pairs of single Unicode code points that are [[canonically equivalent]]. The reason for this are compatibility issues with legacy systems. Unless two characters are canonically equivalent, they are not "duplicate" in the narrow sense. There is, however, room for disagreement on whether two Unicode characters really encode the same [[grapheme]] in cases such as the ~~"micro sign" [[µ]] vs. the Greek [[μ]], or cases that would usually be considered font variants but are encoded separately because of widespread use of font variants (e.g. [[L]] vs. "script L"~~ {{~~SMP~~unichar\|~~Latn~~00B5\|ℒMICRO SIGN\|nlink=Micro-}} ~~vs. "blackletter L"~~versus {{~~SMP~~unichar\|~~de-Latf~~03BC\|~~𝕷}}~~GREEK ~~vs.~~SMALL ~~"boldface~~LETTER ~~blackletter~~MU L"\|nlink=Mu ~~{{SMP\|de-Latf\|𝔏~~(letter)}}~~) as distinctive [[mathematical symbols]]~~. This should be clearly distinguished from Unicode characters that are rendered as identical glyphs or near-identical glyphs ([[homoglyph]]s), either because they are historically cognate (such as Greek [[Η]] vs. Latin [[H]]) or because of ~~coincidential~~coincidental similarity (such as Greek [[Ρ]] vs. Latin [[P]], or Greek [[Η]] vs. Cyrillic [[Н]], or the following homoglyph septuplet: astronomical symbol for "Sun" [[☉]] ~~vs.~~ , "circled dot operator" [[XNOR gate\|⊙]] ~~vs.~~, the Gothic letter [[𐍈]] ~~vs.~~, the IPA symbol for a bilabial click {{IPA link\|ʘ}}, the [[ʘOsage script\|Osage]] letter 𐓃, the [[Tifinagh]] letter ⵙ, and the archaic Cyrillic letter [[Monocular O\|Ꙩ]]). ==Duplicate vs. derived character== {{~~see~~further\|~~character~~Character (computing)\|~~grapheme~~Grapheme}} Unicode aims at encoding graphemes, not individual "meanings" ("semantics") of graphemes, and not [[glyph]]s. It is a matter of case-by-case judgement whether such characters should ~~recieve~~receive ~~seperate~~separate encoding when used in technical contexts, e.g. Greek letters used as mathematical symbols: thus, the choice ~~that it may be useful~~ to have a "[[micro-]] sign" µ separate from Greek μ, but not a "[[Mega-\|Mega]] sign" separate from Latin M, iswas a pragmatic decision upby ~~to the~~The Unicode Consortium for historical reasons (namely, compatibility with [[Latin-1]], which includes a micro sign). Technically µ and μ are not duplicate characters in that the consortium viewed these symbols as distinct characters (while it regarded M for "Mega" and Latin M as one and the same character). Technically these are not duplicate characters in that the consortium viewed such symbols used in mathematics as distinct characters from their twins in the Greek alphabet. For example, the upper-case character [[Π]] (U+03A0) typically shares the same [[glyph]] with the mathematical character n-ary product ∏ (U+220F). These are two different characters because one is a mathematical symbol, appearing in symbolic expressions, and the other is a letter, appearing in linguistic context. Note that merely having different "meanings" is not sufficient grounds to split a grapheme into several characters:. Thus, the [[acute accent]] may represent word accent in Welsh or Swedish, it may express vowel ~~pitch~~quality in ~~Catalan or~~ French, and it may express vowel length in Hungarian, Icelandic or Irish. Since all these languages are written in the same [[writing system\|script]], namely [[Latin ~~alphabet~~script]], the acute accent in its various meanings is considered one and the same combining diacritic character ~~(U+~~{{unichar\|0301~~). Confusingly~~}}, ~~there~~and so the accented letter [[é]] is ~~however~~the same character in French and Hungarian. There is a separate "combining diacritic acute tone mark" at U+{{unichar\|0341}} for the romanization of tone languages, one important difference from the acute accent being that in a language like French, the acute accent can replace the dot over the lowercase i, whereas in a language like Vietnamese, the acute tone mark is added above the dot. Diacritic signs for alphabets considered independent may be encoded separately, such as the acute ("tonos") for the Greek alphabet at U+{{unichar\|0384}}, and for the Armenian alphabet at U+{{unichar\|055B}}. Some Cyrillic-based alphabets (~~but~~such ~~the~~as ~~Cyrillic~~[[Russian alphabet~~, which~~\|Russian]]) also ~~uses~~use the acute accent, ~~does~~but ~~not~~there ~~have~~is ano "Cyrillic acute" encoded separately and U+~~301~~0301 should be used for Cyrillic as well as Latin, (see [[Cyrillic characters in Unicode]]). The point that the same grapheme can have many "meanings" is even more obvious considering e.g. the letter [[U]], which has entirely different phonemic referents in the various languages that use it in their orthographies (English {{IPA\|/juː/, /ʊ/, /ʌ/}} etc., French {{IPA\|/y/}}, German {{IPA\|/uː/, /u/}}, etc., not to mention various uses of [[U (disambiguation)\|U as a symbol]]). ==Compatibility issues== {{further\|Unicode compatibility characters}} ~~{{see\|Mapping_of_Unicode_characters#Legacy_Compatibility_Blocks}}~~ ===CJK fullwidth forms=== {{~~see~~main\|Fullwidth form}} In traditional [[~~CJK~~Chinese character encoding]] ~~encodings~~s, characters usually took either a single [[byte]] (known as halfwidth) or two bytes (known as fullwidth). Characters that took a single byte were generally displayed at half the width of those that took two bytes. Some characters such as the [[~~latin~~Latin alphabet]] were available in both halfwidth and fullwidth versions. As the halfwidth versions were more commonly used, they were generally the ones mapped to the standard code points for those characters. Therefore a separate section was needed for the fullwidth forms to preserve the distinction. ==Letterlike symbols== {{~~see~~main\|Letterlike Symbols}} In some cases, specific graphemes have acquired a specialized symbolic or technical meaning separate from their original function. A prominent example is the Greek letter [[Pi (letter)\|π]] which is widely recognized as the symbol for athe mathematical constant of a circle's circumference divided by its diameter even by people not literate in Greek. Several variants of the entire Greek and Latin alphabets specifically for use as mathematical symbols are encoded in the [[Mathematical Alphanumeric Symbols]] range. This range disambiguates characters that would usually be considered font variants but are encoded separately because of widespread use of font variants e.g. [[L]] vs. "script L" {{Script\|Latn\|ℒ}} vs. "blackletter L" {{Script\|de-Latf\|𝔏}} vs. "boldface blackletter L" {{Script\|de-Latf\|𝕷}}) as distinctive [[mathematical symbols]]. It is intended for use only in mathematical or technical notation, not use in non-technical text.<ref>{{Cite web \|title=UTR #25: Unicode and Mathematics \|url=http://unicode.org/reports/tr25/tr25-5.html#_Toc21 \|access-date=2024-03-04 \|website=unicode.org}}</ref> ~~Several variants of the entire Greek and Latin alphabets specifically for use as mathematical symbols are encoded in the [[Mathematical alphanumeric symbols]] range.~~ === Greek === Many [[Greek ~~letter~~alphabet\|Greek letters]]s are used as [[technical symbol]]s. All of the Greek letters are encoded in the Greek section of Unicode but many are encoded a second time under the name of the technical symbol they represent. ~~Of these,~~The "[[micro sign]] ~~is in the [[Latin-1]] range and most of the rest are in the [[Letterlike Symbols]] range. The "micro sign~~" (U+{{unichar\|00B5~~, µ~~}}) is obviously inherited from [[ISO 8859-1]], but the origin of the others is less clear. Other Greek glyph variants encoded as separate characters include the [[lunate sigma]] Ϲ ϲ contrasting with Σ σ, final sigma ς (strictly speaking a contextual glyph variant) contrasting with σ, The [[Qoppa]] numeral symbol Ϟ ϟ contrasting with the archaic Ϙ ϙ. Greek letters assigned separate "symbol" codepoints include the [[Letterlike Symbols]] [[ϐ]], [[ϵ]], [[ϑ]], [[Pi (letter)\|ϖ]], [[ϱ]], [[ϒ]], and [[ϕ]] (contrasting with β, ε, θ, π, ρ, Υ, φ); the Ohm symbol [[Ω]] (contrasting with Ω); and the [[Unicode mathematical operators and symbols\|mathematical operators]] for the product [[∏]] and sum [[∑]] (contrasting with [[Pi (letter)\|Π]] and [[Σ]]). ===Roman numerals=== [[Unicode]] has a number of characters specifically designated as [[Roman numerals]], as part of the ''[[Number Forms'']] range from U+2160 to U+2183. For example, Roman 1988 ({{char\|MCMLXXXVIII}}) could alternatively be written as {{char\|ⅯⅭⅯⅬⅩⅩⅩⅧ}}. This range includes both ~~upper-~~uppercase and lowercase numerals, as well as pre-combined glyphs for numbers up to 12 ({{char\|Ⅻ}} for {{char\|XII}}), mainly intended for clock faces~~. Similarly precombined ligatures for 5,000 and 10,000 exist~~. ~~The pre-combined glyphs should only be used to represent the individual numbers where the use of individual glyphs is not wanted, and not to replace compounded numbers.~~ The pre-combined glyphs should only be used to represent the individual numbers where the use of individual glyphs is not wanted, and not to replace compounded numbers. For example, one can combine {{char\|Ⅹ}} with {{char\|Ⅰ}} to produce Roman numeral 11 ({{char\|ⅩⅠ}}), so U+216A ({{char\|Ⅺ}}) is canonically equivalent to {{char\|ⅩⅠ}}. Such characters are also referred to as composite compatibility characters or decomposable compatibility characters. Such characters would not normally have been included within the Unicode standard except for compatibility with other existing encodings (see [[Unicode compatibility characters]]). The goal was to accommodate simple translation from existing encodings into Unicode. This makes translations in the opposite direction complicated because multiple Unicode characters may map to a single character in another encoding. Without the compatibility concerns the only characters necessary would be: {{Char\|Ⅰ}}, {{Char\|Ⅴ}}, {{Char\|Ⅹ}}, {{Char\|Ⅼ}}, {{Char\|Ⅽ}}, {{Char\|Ⅾ}}, {{Char\|Ⅿ}}, {{Char\|ⅰ}}, {{Char\|ⅴ}}, {{Char\|ⅹ}}, {{Char\|ⅼ}}, {{Char\|ⅽ}}, {{Char\|ⅾ}}, {{Char\|ⅿ}}, {{Char\|ↀ}}, {{Char\|ↁ}}, {{Char\|ↂ}}, {{Char\|ↇ}}, {{Char\|ↈ}}, and {{Char\|Ↄ}}; all other Roman numerals can be composed from these characters. ~~<table>~~ ~~<tr><td width="33%">uppercase one</td><td width="33%">U+2160</td><td width="33%">Ⅰ</td></tr>~~ ~~<tr><td>uppercase two</td><td>U+2161</td><td>Ⅱ</td></tr>~~ ~~<tr><td>uppercase three</td><td>U+2162</td><td>Ⅲ</td></tr>~~ ~~<tr><td>uppercase four</td><td>U+2163</td><td>Ⅳ</td></tr>~~ ~~<tr><td>uppercase five</td><td>U+2164</td><td>Ⅴ</td></tr>~~ ~~<tr><td>uppercase six</td><td>U+2165</td><td>Ⅵ</td></tr>~~ ~~<tr><td>uppercase seven</td><td>U+2166</td><td>Ⅶ</td></tr>~~ ~~<tr><td>uppercase eight</td><td>U+2167</td><td>Ⅷ</td></tr>~~ ~~<tr><td>uppercase nine</td><td>U+2168</td><td>Ⅸ</td></tr>~~ ~~<tr><td>uppercase ten</td><td>U+2169</td><td>Ⅹ</td></tr>~~ ~~<tr><td>uppercase eleven</td><td>U+216A</td><td>Ⅺ</td></tr>~~ ~~<tr><td>uppercase twelve</td><td>U+216B</td><td>Ⅻ</td></tr>~~ ~~<tr><td>uppercase fifty</td><td>U+216C</td><td>Ⅼ</td></tr>~~ ~~<tr><td>uppercase one hundred</td><td>U+216D</td><td>Ⅽ</td></tr>~~ ~~<tr><td>uppercase five hundred</td><td>U+216E</td><td>Ⅾ</td></tr>~~ ~~<tr><td>uppercase one thousand</td><td>U+216F</td><td>Ⅿ</td></tr>~~ ~~<tr><td>lowercase one</td><td>U+2170</td><td>ⅰ</td></tr>~~ ~~<tr><td>lowercase two</td><td>U+2171</td><td>ⅱ</td></tr>~~ ~~<tr><td>lowercase three</td><td>U+2172</td><td>ⅲ</td></tr>~~ ~~<tr><td>lowercase four</td><td>U+2173</td><td>ⅳ</td></tr>~~ ~~<tr><td>lowercase five</td><td>U+2174</td><td>ⅴ</td></tr>~~ ~~<tr><td>lowercase six</td><td>U+2175</td><td>ⅵ</td></tr>~~ ~~<tr><td>lowercase seven</td><td>U+2176</td><td>ⅶ</td></tr>~~ ~~<tr><td>lowercase eight</td><td>U+2177</td><td>ⅷ</td></tr>~~ ~~<tr><td>lowercase nine</td><td>U+2178</td><td>ⅸ</td></tr>~~ ~~<tr><td>lowercase ten</td><td>U+2179</td><td>ⅹ</td></tr>~~ ~~<tr><td>lowercase eleven</td><td>U+217A</td><td>ⅺ</td></tr>~~ ~~<tr><td>lowercase twelve</td><td>U+217B</td><td>ⅻ</td></tr>~~ ~~<tr><td>lowercase fifty</td><td>U+217C</td><td>ⅼ</td></tr>~~ ~~<tr><td>lowercase one hundred</td><td>U+217D</td><td>ⅽ</td></tr>~~ ~~<tr><td>lowercase five hundred</td><td>U+217E</td><td>ⅾ</td></tr>~~ ~~<tr><td>lowercase one thousand</td><td>U+217F</td><td>ⅿ</td></tr>~~ ~~<tr><td>one thousand C D</td><td>U+2180</td><td>ↀ</td></tr>~~ ~~<tr><td>five thousand</td><td>U+2181</td><td>ↁ</td></tr>~~ ~~<tr><td>ten thousand</td><td>U+2182</td><td>ↂ</td></tr>~~ ~~<tr><td>reverse one hundred</td><td>U+2183</td><td>Ↄ</td></tr>~~ ~~</table>~~ === Arabic presentation forms === Some of these characters are considered precompositions by Unicode methodology because many of these characters can be composed from the other characters. For example, one can combine Ⅹ with Ⅰ to mean roman numeral eleven (ⅩⅠ), so U+216A (Ⅺ) is canonically equivalent to ⅩⅠ. Such characters are also referred to as composite compatibility characters or decomposable compatibility characters. Such characters would not normally have been included within the Unicode standard except for compatibility with other existing encodings.{{fact}} The goal was to accommodate simple translation from existing encodings into unicode. This makes translations in the opposite direction complicated because multiple Unicode characters may map to a single character in another encoding. Without the compatibility concerns the only characters necessary would be: Ⅰ, Ⅴ, Ⅹ, Ⅼ, Ⅽ, Ⅾ, Ⅿ, ⅰ, ⅴ, ⅹ, ⅼ, ⅽ, ⅾ, ⅿ, ↀ, ↁ, ↂ, Ↄ. All other roman numerals can be composed from these. {{main\|Arabic Presentation Forms-A\|Arabic Presentation Forms-B}}Unicode has encoded compatibility characters for contextual Arabic letter forms where its contextual forms are encoded as separate code points (isolated, final, initial, and medial). For example, {{Codepoint\|0647}} has its contextual forms encoded at these 4 code points: * {{Codepoint\|FEE9}} * {{Codepoint\|FEEA}} * {{Codepoint\|FEEB}} * {{Codepoint\|FEEC}} The contextual-form characters are not recommended for general use. There are also compatibility Arabic ligatures encoded such as {{unichar\|FDF2}} and {{unichar\|FDFD}}. === Hebrew presentation forms === ~~==References==~~ {{Main\|Alphabetic Presentation Forms}} ~~{{unreferenced}}~~ Hebrew presentation forms include ligatures, several precomposed characters and wide variants of Hebrew letters. The aleph-lamed ligature is encoded as a separate character at {{unichar\|FB4F}}. The wide variants are listed below: <!-- Just like the main list; Duplicate character: Original character(s) --> {{unichar\|FB21}} {{unichar\|FB22}} * {{unichar\|FB23}} * {{unichar\|FB24}} * {{unichar\|FB25}} * {{unichar\|FB26}} * {{unichar\|FB27}} * {{unichar\|FB28}} These characters are variants of ordinary Hebrew letters encoded for [[Justification (typesetting)\|justification]] of texts written in Hebrew, such as the Torah. Unicode also encodes a stylistic variant of {{Unichar\|5e2}} at {{Unichar\|FB20}}. == List == {{incomplete list\|date=April 2022}}<!-- General pattern: Duplicate character, original character --> {{unichar\|1F549\|OM SYMBOL\|nlink=Om}}: {{unichar\|0950\|DEVANAGARI OM\|nlink=Devanagari}} {{unichar\|212B\|ANGSTROM SIGN\|nlink=Ångström}}: {{unichar\|00C5\|LATIN CAPITAL LETTER A WITH RING ABOVE}} {{unichar\|00B5\|MICRO SIGN \|nlink=Micro-}}: {{unichar\|03BC\|GREEK SMALL LETTER MU\|nlink=Greek and Coptic}} {{unichar\|037E\|GREEK QUESTION MARK\|nlink=Question_mark#Greek_question_mark}}: {{unichar\|003B\|SEMICOLON}} {{unichar\|212A\|KELVIN SIGN\|nlink=Kelvin}}: {{unichar\|004B\|LATIN CAPITAL LETTER K}} {{unichar\|2024\|ONE DOT LEADER\|nlink=Leader_(typography)}}: {{unichar\|002E\|FULL STOP}} {{unichar\|2126\|\|nlink=Ohm}}: {{unichar\|03A9\|GREEK CAPITAL LETTER OMEGA}} {{Unichar\|2236\|RATIO}}: {{Unichar\|003A\|COLON}} {{Unichar\|0387\|}}: {{Unichar\|00B7\|COLON}} {{Unichar\|2A75\|}}: {{Unichar\|03d\|COLON}}, {{Unichar\|03d\|COLON}} {{Unichar\|2A76\|}}: {{Unichar\|03d\|COLON}}, {{Unichar\|03d\|COLON}}, {{Unichar\|03d\|COLON}} {{Unichar\|27EAF}}: {{Unichar\|FA23}} {{Unichar\|2135}}: {{Unichar\|5d0}} {{Unichar\|2136}}: {{Unichar\|05d1}} {{Unichar\|2137}}: {{Unichar\|5d2}} {{Unichar\|2138}}: {{Unichar\|5d3}} {{Unichar\|2254}}: {{Unichar\|003A\|COLON}}, {{Unichar\|03d\|COLON}} {{Unichar\|2255}}: {{Unichar\|03d\|COLON}}, {{Unichar\|003A\|COLON}} {{Unichar\|2A74}}: {{Unichar\|003A\|COLON}}, {{Unichar\|003A\|COLON}}, {{Unichar\|03d\|COLON}} {{Unichar\|0340\|cwith=◌}}: {{Unichar\|0300\|cwith=◌}} {{Unichar\|0341\|cwith=◌}}: {{Unichar\|0301\|cwith=◌}} {{Unichar\|0344\|cwith=◌}}: {{Unichar\|0308\|cwith=◌}}, {{Unichar\|0301\|cwith=◌}} {{Unichar\|222C}}: {{Unichar\|222b}}, {{Unichar\|222b}} {{Unichar\|222D }}: {{Unichar\|222b}}, {{Unichar\|222b}}, {{Unichar\|222b}} {{Unichar\|2A0C}}: {{Unichar\|222b}}, {{Unichar\|222b}}, {{Unichar\|222b}}, {{Unichar\|222b}} {{Unichar\|03d0}}: {{Unichar\|03b2}} {{Unichar\|03F4}}: {{Unichar\|0398}} {{Unichar\|03d1}}: {{Unichar\|03b8}} {{Unichar\|03d6}}: {{Unichar\|03c0}} {{Unichar\|03F1}}: {{Unichar\|03C1}} {{Unichar\|03d2}}: {{Unichar\|03a5}} {{Unichar\|03d3}}: {{Unichar\|038e}} {{Unichar\|03d4}}: {{Unichar\|03AB}} {{Unichar\|03d5}}: {{Unichar\|03c6}} {{Unichar\|0374}}: {{Unichar\|02b9}} {{Unichar\|03F0}}: {{Unichar\|03BA}} {{Unichar\|03f9}}: {{Unichar\|03a3}} {{Unichar\|03F2}}: {{Unichar\|03c3}} {{Unichar\|017F\|nlink=long s}}: {{Unichar\|0073}} {{Unichar\|03F5}}: {{Unichar\|03b5}} {{Unichar\|210f\|nlink=Reduced Planck constant}}: {{Unichar\|0127}} {{Unichar\|2107}}: {{Unichar\|0190}} {{Unichar\|2103}}: {{Unichar\|b0}}, {{Unichar\|43}} {{Unichar\|2109}}: {{Unichar\|b0}}, {{Unichar\|46}} {{Unichar\|BA}}: {{Unichar\|6F}} {{Unichar\|AA}}: {{Unichar\|61}} {{Unichar\|2139}}: {{Unichar\|69}} {{unichar\|FB20}}: {{unichar\|5e2}} {{unichar\|FB21}}: {{Unichar\|5d0}} {{unichar\|FB22}}: {{Unichar\|5d3}} * {{unichar\|FB23}}: {{Unichar\|5d4}} * {{unichar\|FB24}}: {{Unichar\|5db}} * {{unichar\|FB25}}: {{Unichar\|5dc}} * {{unichar\|FB26}}: {{Unichar\|5dd}} * {{unichar\|FB27}}: {{Unichar\|5e8}} * {{unichar\|FB28}}: {{Unichar\|5ea}} * {{Unichar\|FB29}}: {{Unichar\|002b}} * {{Unichar\|0343\|cwith=◌}}: {{Unichar\|0313\|cwith=◌}} * {{Unichar\|1ffd}}: {{Unichar\|00B4}} * {{unichar\|0384}}: {{Unichar\|00B4}} * {{Unichar\|1fef}}: {{Unichar\|0060}} ==See also== [[IDN homograph attack]] [[Mapping of Unicode characters]] [[Unicode equivalence]] [[Canonical equivalence]] [[Homoglyph]] [[ASCII art]] ==References== {{reflist}} ~~[[Category:~~{{Unicode]] navigation}} [[Category:Unicode]] ~~[[fr:Duplication de caractères unicode]]~~