Revision as of 19:09, 11 March 2023 edit 93.230.217.137 (talk) Removed COMPLETELY IRRELEVANT bloat ← Previous edit		Revision as of 15:20, 12 March 2023 edit undo Comp.arch (talk \| contribs) Extended confirmed users 41,478 edits Microsoft does recommend, with "Use UTF-8", and describing UTF-16 a "[unique] burden", so not only "encouraging [its] use". Also "UTF-8 is the default and only code page on console, so we recommend -A APIs to take full advantage of that." Next edit →
Line 1: {{more citations needed\|date=June 2011}} [[Microsoft]] was one of the first companies to implement [[Unicode]] in their products. [[Windows NT]] was the first operating system that used "wide characters" in [[system call]]s. Using the (now obsolete) [[UCS-2]] encoding scheme at first, it was upgraded to the [[variable-width encoding]] [[UTF-16]] starting with [[Windows 2000]], allowing a representation of additional planes with surrogate pairs. However Microsoft did not support [[UTF-8]] in its API until May 2019~~, though it now appears to be encouraging its use~~.~~<ref>{{cite web\|title=Use UTF-8 code pages in Windows apps\|url=https://learn.microsoft.com/en-us/windows/apps/design/globalizing/use-utf8-code-page\|language=en}}</ref>~~ {{As of\|2020}}, Microsoft recommends programmers use UTF-8,<ref name="Microsoft-UTF-8" /> on Windows and [[Xbox]], even states "UTF-8 is the universal code page for internationalization [and] UTF-16 [..] a unique burden that Windows places on code that targets multiple platforms."<ref>{{Cite web \|title=UTF-8 support in the Microsoft Game Development Kit (GDK) - Microsoft Game Development Kit \|url=https://learn.microsoft.com/en-us/gaming/gdk/_content/gc/system/overviews/utf-8 \|access-date=2023-03-05 \|website=learn.microsoft.com \|language=en-us \|quote=By operating in UTF-8, you can ensure maximum compatibility [..] Windows operates natively in UTF-16 (or WCHAR), which requires code page conversions by using MultiByteToWideChar and WideCharToMultiByte. This is a unique burden that Windows places on code that targets multiple platforms. [..] The Microsoft Game Development Kit (GDK) and Windows in general are moving forward to support UTF-8 to remove this unique burden of Windows on code targeting or interchanging with multiple platforms and the web. Also, this results in fewer internationalization issues in apps and games and reduces the test matrix that's required to get it right.}}</ref> A large amount of Microsoft documentation uses the word "Unicode" to refer explicitly to the UTF-16 encoding. Anything else, including UTF-8, is not "Unicode". Line 28 ⟶ 30: In April 2018 (or possibly November 2017<ref>{{cite web\|title=Windows10 Insider Preview Build 17035 Supports UTF-8 as ANSI\|url=https://news.ycombinator.com/item?id=15710685\|website=Hacker News\|access-date=7 May 2018}}</ref>), with insider build 17035 (nominal build 17134) for Windows 10, a "Beta: Use Unicode UTF-8 for worldwide language support" checkbox appeared for setting the locale code page to UTF-8.{{efn\|1=Found under control panel, "Region" entry, "Administrative" tab, "Change system locale" button.}} This allows for calling "narrow" functions, including <code>fopen</code> and <code>SetWindowTextA</code>, with UTF-8 strings. However this is a system-wide setting and a program cannot assume it is set. In May 2019, Microsoft added the ability for a program to set the code page to UTF-8 itself,<ref name="Microsoft-UTF-8">{{~~Cite~~cite web\|title=Use ~~the Windows~~ UTF-8 code ~~page~~pages -in ~~UWP~~Windows ~~applications~~apps\|url=https://~~docs~~learn.microsoft.com/en-us/windows/~~uwp~~apps/design/globalizing/use-utf8-code-page \|access-date=2020-06-06 \|quote=As of Windows ~~Version~~ version 1903 (May  2019 ~~Update~~update), you can use the ActiveCodePage property in the appxmanifest for packaged apps, or the fusion manifest for unpackaged apps, to force a process to use UTF-8 as the process code page. [...] <code>CP_ACP</code> equates to <code>CP_UTF8</code> only if running on Windows version 1903 (May 2019 update) or above and the ActiveCodePage property described above is set to UTF-8. Otherwise, it honors the legacy system code page. We recommend using <code>CP_UTF8</code> explicitly. \|website=~~docs~~learn.microsoft.com \|language=en-us}}</ref><ref>{{cite web\|url=https://skanthak.homepage.t-online.de/quirks.html#quirk31\|title=Windows 10 1903 and later versions finally support UTF-8 with the A forms of the Win32 functions}}</ref> allowing programs written to use UTF-8 to be run by non-expert users. In [[Windows 11]] some system files are required to use UTF-8 and do not require a Byte Order Mark.<ref>{{Cite web\|last=themar-msft\|title=Customize the Windows 11 Start menu\|url=https://docs.microsoft.com/en-us/windows-hardware/customize/desktop/customize-the-windows-11-start-menu\|access-date=2021-06-29\|website=docs.microsoft.com\|language=en-us\|quote=Make sure your LayoutModification.json uses UTF-8 encoding.}}</ref> Notepad can now recognize UTF-8 without the Byte Order Mark, and can be told to write UTF-8 without a Byte Order Mark.{{cn\|date=November 2022}} Some other Microsoft products are using UTF-8 internally, including Visual Studio{{cn\|date=November 2022}} and their [[SQL Server 2019]], with Microsoft claiming 35% speed increase from use of UTF-8, and "nearly 50% reduction in storage requirements."<ref>{{Cite web\|date=2019-07-02\|title=Introducing UTF-8 support for SQL Server\|url=https://techcommunity.microsoft.com/t5/sql-server/introducing-utf-8-support-for-sql-server/ba-p/734928\|quote=For example, changing an existing column data type from NCHAR(10) to CHAR(10) using an UTF-8 enabled collation, translates into nearly 50% reduction in storage requirements. [..] In the ASCII range, when doing intensive read/write I/O on UTF-8<!-- " " in quote, but ok to strip-->, we measured an average 35% performance improvement over UTF-16 using clustered tables with a non-clustered index on the string column, and an average 11% performance improvement over UTF-16 using a heap. \|access-date=2021-08-24\|website=techcommunity.microsoft.com\|language=en}}</ref>

Unicode in Microsoft Windows: Difference between revisions