fix error loading and displaying file paths with special characters

This commit is contained in:
Harry Chen
2026-08-29 14:38:51 -04:00
parent febc5e1983
commit 20842da4c7
17 changed files with 389 additions and 216 deletions
+3
View File
@@ -43,6 +43,9 @@ MANIFEST
*.manifest
*.spec
# Not PyInstaller's: the UTF-8 active-code-page manifest, a committed source.
!src/app/utf8.manifest
# Installer logs
pip-log.txt
pip-delete-this-directory.txt
+83 -25
View File
@@ -6,23 +6,49 @@ single Cyrillic letter — which was already wrong, independently of
localization, for anyone whose dataset path was not pure ASCII.
Five faces are committed here and **all five are embedded** in the executable,
507 KB together. Nothing has to be downloaded for the interface to render in
any of the thirteen languages in `src/i18n/Languages.h`.
3.3 MB together. Nothing has to be downloaded for the interface to render in
any of the thirteen languages in `src/i18n/Languages.h`, and nothing has to be
downloaded for an ordinary file name to render either.
| | file | size | from |
|---|---|---|---|
| Latin / Cyrillic | `SpirulaUI-Regular.ttf` | 58 KB | Source Sans 3 |
| CJK, per region | `SpirulaCJK-{JP,SC,TC,KR}.otf` | 449 KB | Noto Sans CJK |
| Latin / Greek / Cyrillic | `SpirulaUI-Regular.ttf` | 159 KB | Source Sans 3 |
| CJK, per region | `SpirulaCJK-{JP,SC,TC,KR}.otf` | 3.2 MB | Noto Sans CJK |
The full CJK faces — 4–8 MB each — are still **fetched at runtime** or bundled
by a regional build, but for a much smaller job: see *The full faces* below.
by a regional build, but for a much smaller job than they used to do: see
*The full faces* below.
## Two jobs, two budgets
The fonts serve two kinds of text and it is worth keeping them apart.
**The interface** is text this program wrote. It is a closed set — the
catalogs under `src/i18n/catalog/` — so it can be covered exactly, and it
must be, in the right regional glyph forms, offline. That is the ~600
characters per region the subsets started out as.
**File names are not.** A dataset path, a folder name or a typed mask prompt
is user data: it can hold any character at all, and one uncovered character is
a `?` in the middle of a path. No subset of *our* strings anticipates
`D:\写真\第2回—テスト①\`. Covering that is what the rest of the 3.3 MB buys,
and it is why the character sets below are declared as blocks and national
standards rather than scraped from the catalogs.
## `SpirulaUI-Regular.ttf`
A subset of **Source Sans 3** (Adobe, SIL OFL 1.1 — `OFL-SourceSans.txt`)
covering exactly the ranges the twelve non-CJK languages use: Basic Latin,
Latin-1, Latin Extended-A, Cyrillic, and the punctuation a translator will
reach for. 431 KB upstream, 58 KB after subsetting.
A subset of **Source Sans 3** (Adobe, SIL OFL 1.1 — `OFL-SourceSans.txt`),
431 KB upstream and 159 KB after subsetting. The ranges are listed in
`tools/make_ui_font.py`; the ones that are there for file names rather than
for the UI are worth naming:
| range | why |
|---|---|
| Latin Extended-B, Extended Additional | pinyin tone marks, Vietnamese |
| combining marks (U+0300–036F) | macOS hands out file names in NFD |
| Greek, whole Cyrillic block | not just the twelve UI locales' letters |
| General Punctuation, U+2100–214F | em dash, curly quotes, ellipsis, the numero sign |
| arrows, math, geometric, symbols, dingbats | stars and check marks name a lot of files |
Source Sans 3 declares **"Source" as a Reserved Font Name**, so a Modified
Version may not carry it. Every name-table record was rewritten accordingly —
@@ -33,8 +59,24 @@ branding decision.
## `SpirulaCJK-{JP,SC,TC,KR}.otf`
Each is **Noto Sans CJK** (Google, SIL OFL 1.1 — `OFL-NotoSansCJK.txt`) cut
down to the ~600 characters that region's own translations use, plus the
native name of every language and a little CJK punctuation. ~110 KB each.
down to three things:
1. that region's **national common-use standard** — GB 2312 level 1 (3755
characters) for SC, Big5 level 1 (5401) for TC, JIS X 0208 level 1 (2965)
for JP, the 2350 KS X 1001 hangul syllables for KR. These are the
percentile answer to "which characters appear in a file name", chosen by
the frequency studies behind each standard rather than by us, and they are
decoded out of Python's own codecs, so no character list is committed here
and none can go stale;
2. CJK punctuation, the halfwidth/fullwidth forms and the enclosed
alphanumerics (①, A, カ) in **every** face, because their forms are
regional too — and kana, bopomofo, the enclosed CJK letters (㈱, ㍻) and the
symbols that name files (★, ♪, ✓) in **one**, since all four subsets are
merged into the same atlas and a second copy would only cost bytes. The
symbols are named one by one rather than taken as a block: of U+2600–27BF
Noto Sans CJK has under a fifth;
3. whatever that region's own translations use, scanned from the catalogs.
Noto Sans CJK is Source Han Sans under its other name, so it is the same
design as the Latin face and a mixed line does not visibly step.
@@ -45,12 +87,26 @@ wrong. Two alternatives were measured and rejected:
| approach | size | cost |
|---|---|---|
| four regional subsets *(what ships)* | 449 KB | — |
| one pan-CJK subset | 287 KB | 262 shared characters in the wrong regional form for three of the four languages |
| a shared base + three deltas | 358 KB | of the characters used by more than one region, only 267 have identical outlines and 262 genuinely differ — so the deltas are most of the weight anyway |
| four regional cuts *(what ships)* | 3.2 MB | — |
| one shared face, single glyph forms | ~1.9 MB | every language but one reads file names in another region's forms |
| a shared base + four deltas | 2.5 MB | 18% saved for a fifth file and a merge-order hazard; of the 10 178 codepoints only 5459 have identical outlines everywhere |
Noto reserves no font name, so a subset could legally keep it; these are
renamed anyway, because a 600-glyph file called "Noto Sans JP" is a lie.
renamed anyway, because a file this small is not the font it was cut from.
## What is still not covered
Hebrew, Arabic, Thai, the Indic scripts and the rest have no glyphs in any of
the five faces, and adding them would not be enough on its own: ImGui applies
no bidi and no shaping, so an Arabic path would come out unjoined and an
Indic one unreordered. A path in one of those scripts is still a row of `?`.
Combining marks are covered but not *positioned*, for the same reason — an
NFD `é` draws as an `e` with the accent alongside rather than over it, which
is worse-looking than the composed form and much better than a `?`.
`?` rather than a hollow box because ImGui picks its fallback glyph from
U+FFFD, `?`, space in that order, and neither Source Sans 3 nor Noto Sans
CJK has a U+FFFD.
## Rebuilding
@@ -58,18 +114,18 @@ renamed anyway, because a 600-glyph file called "Noto Sans JP" is a lie.
pip install fonttools
python3 tools/make_ui_font.py # all five: download, subset, rename
python3 tools/make_ui_font.py --check # verify the committed files
python3 tools/check_font_coverage.py # cheap: did a translation outgrow them?
python3 tools/check_font_coverage.py # cheap: is anything missing?
```
The output is byte-reproducible, so `--check` is meaningful. The committed
files are build artifacts kept in the tree so that the build needs no network
— the same reasoning as the committed files under `src/generated/`.
**The CJK subsets are derived from the catalogs.** Edit a translation and they
go stale, and the symptom is one hollow box in the middle of an otherwise fine
sentence. `tools/check_font_coverage.py` is the guard and runs on every
`build_develop.bash`; it needs neither the network nor fontTools, because it
parses `cmap` by hand.
**The subsets are partly derived from the catalogs.** Edit a translation and
they can go stale, and the symptom is one `?` in the middle of an
otherwise fine sentence. `tools/check_font_coverage.py` is the guard and runs
on every `build_develop.bash`; it needs neither the network nor fontTools,
because it parses `cmap` by hand.
## The full faces
@@ -78,10 +134,12 @@ generates `app_generated/cjk_faces.h` from it and `src/app/gui/Fonts.cpp` reads
that. Nothing of them is committed here — they are downloaded, verified against
the SHA-256 in that file, and cached.
They exist for the text this program did **not** write: dataset paths, file
names and typed mask prompts are user data and can hold any character at all,
which no subset of our own strings can anticipate. Until one is fetched, a
folder called `C:\写真\` may render as boxes; the interface itself never does.
They are what covers the tail: a rare surname, a classical character, anything
outside a national common-use list. `SS_FONT_CJK=sc|tc|jp|kr|all` bundles one
or all of them into `<exe dir>/fonts/` instead, which is the first place
`Fonts.cpp` looks after `$SS_FONT_DIR`. They are bundled rather than embedded
because `ss_embed_file()` costs four characters of C source per byte, so `all`
would be a 92 MB literal.
See `docs/i18n.md` for `SS_FONT_CJK`, the download offer, and what a build with
`SS_FONT_CJK=none` does instead.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
+8 -4
View File
@@ -257,12 +257,13 @@ if(SS_BUILD_CLI OR SS_BUILD_GUI)
list(REMOVE_DUPLICATES SS_TOOL_SOURCES)
list(REMOVE_DUPLICATES SS_TOOL_LIBS)
# app.rc carries the icon Explorer and the taskbar draw, which is a link
# input on Windows and has no counterpart elsewhere: macOS reads the
# bundle's .icns, and Linux has only the window icon GuiMain.cpp sets.
# app.rc is the icon Explorer and the taskbar draw; utf8.manifest makes the
# PROCESS code page UTF-8, so the -A Win32 calls, the CRT, argv and getenv
# agree with "/utf-8". A source, not /MANIFESTINPUT: CMake runs mt.exe too.
if(WIN32)
enable_language(RC)
list(APPEND SS_TOOL_SOURCES ${SS_SRC}/app/app.rc)
list(APPEND SS_TOOL_SOURCES ${SS_SRC}/app/app.rc
${SS_SRC}/app/utf8.manifest)
endif()
add_executable(spirula ${SS_SRC}/app/Main.cpp ${SS_TOOL_SOURCES})
@@ -299,6 +300,9 @@ if(SS_BUILD_CLI OR SS_BUILD_GUI)
# itself. Symlink `spirula` if you want those names.
if(SS_SEPARATE_TOOLS)
function(ss_tool_exe name sources defs libs)
if(WIN32)
list(APPEND sources ${SS_SRC}/app/utf8.manifest)
endif()
# ss_i18n is linked explicitly: these targets deliberately do not
# link the engine library, and Main.cpp's `--lang` handling needs
# it. It is a leaf (cmake/SsI18n.cmake), so this costs nothing.
+23 -21
View File
@@ -1,20 +1,35 @@
# Embedding data files into the executables as byte arrays, so the apps are
# self-contained (no runtime lookup of viewer.html or reference/scripts/mask.py).
# ss_hex_to_literal(<hex string> <out var>)
#
# Bytes as a C++ string literal, not a `{0x..,}` array: the array form costs
# g++ 440 MB of memory per 3 MB embedded, this one 43 MB (measured, g++ 13).
function(ss_hex_to_literal hex out)
# 40 bytes a line keeps every literal far under MSVC's 64 KB cap. `\x` is
# greedy, but every byte is followed by a backslash, so it cannot run on.
string(REPEAT "[0-9a-f][0-9a-f]" 40 _grp)
string(REGEX REPLACE "(${_grp})" "\\1|" _s "${hex}")
string(REGEX REPLACE "([0-9a-f][0-9a-f])" "\\\\x\\1" _s "${_s}")
string(REPLACE "|" "\"\n\"" _s "${_s}")
set(${out} "\"${_s}\"" PARENT_SCOPE)
endfunction()
# ss_embed_file(<input> <output_header> <symbol>)
#
# Writes a header defining `k<symbol>[]` / `k<symbol>Size` holding the bytes of
# <input>. Regenerated at configure time whenever <input> changes.
function(ss_embed_file input output_header symbol)
file(READ ${input} _hex HEX)
string(REGEX REPLACE "([0-9a-f][0-9a-f])" "0x\\1," _bytes ${_hex})
ss_hex_to_literal("${_hex}" _bytes)
file(RELATIVE_PATH _rel ${SS_ROOT} ${input})
# -1: a string literal brings a terminator the byte count must not include.
string(CONCAT _text
"#pragma once\n"
"// AUTO-GENERATED from ${_rel} -- do not edit.\n"
"#include <cstddef>\n"
"inline const unsigned char k${symbol}[] = {${_bytes}};\n"
"inline const size_t k${symbol}Size = sizeof(k${symbol});\n")
"inline const unsigned char k${symbol}[] =\n${_bytes};\n"
"inline const size_t k${symbol}Size = sizeof(k${symbol}) - 1;\n")
ss_write_if_different(${output_header} "${_text}")
set_property(DIRECTORY ${SS_ROOT} APPEND
PROPERTY CMAKE_CONFIGURE_DEPENDS ${input})
@@ -22,21 +37,8 @@ endfunction()
# ss_cjk_faces()
#
# Generates two headers from assets/fonts/cjk_faces.txt:
#
# app_generated/cjk_faces.h the table of downloadable FULL faces
# app_generated/cjk_subsets.h the four SUBSETS, embedded as byte arrays
#
# and -- for a regional build (SS_FONT_CJK=sc|tc|jp|kr|all) -- downloads the
# named full faces so the executable ships with them beside it.
#
# The subsets are embedded and the full faces are not, which is the whole
# design in one line. A subset is ~110 KB because it holds only the characters
# this program's own translations use, so all four fit in the executable and
# every language renders with nothing to download. A full face is 4-8 MB, and
# ss_embed_file() turns one byte into five characters of C source, so `all`
# would be a 130 MB array literal. A regional build therefore installs
# <exe dir>/fonts/, which is the first place Fonts.cpp looks.
# assets/fonts/cjk_faces.txt -> cjk_faces.h (the downloadable full faces) and
# cjk_subsets.h (the four embedded subsets). See assets/fonts/README.md.
function(ss_cjk_faces)
set(spec ${SS_ROOT}/assets/fonts/cjk_faces.txt)
set(header ${CMAKE_BINARY_DIR}/app_generated/cjk_faces.h)
@@ -87,11 +89,11 @@ function(ss_cjk_faces)
" python3 tools/make_ui_font.py")
endif()
file(READ ${sub} _sub_hex HEX)
string(REGEX REPLACE "([0-9a-f][0-9a-f])" "0x\\1," _sub_bytes ${_sub_hex})
ss_hex_to_literal("${_sub_hex}" _sub_bytes)
string(APPEND sub_arrays
"inline const unsigned char kCjkSubset${ID}[] = {${_sub_bytes}};\n")
"inline const unsigned char kCjkSubset${ID}[] =\n${_sub_bytes};\n")
string(APPEND sub_table
" {\"${id}\", kCjkSubset${ID}, sizeof(kCjkSubset${ID})},\n")
" {\"${id}\", kCjkSubset${ID}, sizeof(kCjkSubset${ID}) - 1},\n")
set_property(DIRECTORY ${SS_ROOT} APPEND
PROPERTY CMAKE_CONFIGURE_DEPENDS ${sub})
+74 -35
View File
@@ -42,42 +42,62 @@ environment lands on. That is why it is a language rather than a hard-coded
## Fonts
Five faces are **embedded in every build**, 507 KB together:
Five faces are **embedded in every build**, 3.3 MB together:
| file | size | covers |
|---|---|---|
| `SpirulaUI-Regular.ttf` | 58 KB | Latin + Cyrillic, from Source Sans 3 |
| `SpirulaCJK-{JP,SC,TC,KR}.otf` | 449 KB | the ~600 characters *per region* this program's own translations use, from Noto Sans CJK |
| `SpirulaUI-Regular.ttf` | 159 KB | Latin, Greek, Cyrillic, punctuation and symbols, from Source Sans 3 |
| `SpirulaCJK-{JP,SC,TC,KR}.otf` | 3.2 MB | each region's national common-use set, kana, hangul, bopomofo, the fullwidth and enclosed forms, and this program's own translations, from Noto Sans CJK |
That is what makes a default build render all thirteen languages, in the right
regional glyph forms, with nothing to download. It matters most in the language
picker, which is exactly the screen a user who cannot read the current
interface language has arrived at — a download button there is a request to
read something they came to that menu because they could not read.
They do two different jobs and it is worth keeping them apart.
Subsetting is what makes it affordable. A full Noto Sans CJK face is 4–8 MB and
there are four of them, because **Han unification** gives the shared codepoints
different default glyph forms per region; a Japanese reader shown the
Simplified Chinese face gets kanji in Chinese forms, which is legible and
visibly wrong. Cutting each one down to the characters the catalogs actually
contain costs ~110 KB, so all four fit and nobody gets the wrong forms.
**The interface** is a closed set — the catalogs under `src/i18n/catalog/` —
so it can be covered exactly, and it must be, in the right regional glyph
forms, offline. That is what makes a default build render all thirteen
languages with nothing to download, and it matters most in the language
picker: the screen a user who cannot read the current interface language has
arrived at, where a download button is a request to read something they came
to that menu because they could not read.
Alternatives measured and rejected: one pan-CJK subset is 287 KB, saving 135 KB
but rendering 262 shared characters in Simplified Chinese forms for three of
the four languages; delta-encoding the four against a common base lands at
358 KB, because of the shared characters only about half have identical
outlines across the regional cuts.
**File names are not a closed set.** A dataset path, a folder name or a typed
mask prompt is user data — it can hold any character, and one uncovered
character is a `?` in the middle of a path. That is what the rest of the
3.3 MB buys, and why the character sets are declared as Unicode blocks and
national standards rather than scraped from the catalogs:
The atlas merges the base face, then the full regional face if one has been
fetched, then all four subsets with the **current language's region first** —
first source wins per codepoint, so the current language gets its own forms
while the other three still supply what only they have (hangul exists only in
the KR face, the simplified 简 of 简体中文 only in SC).
- the Latin face takes whole blocks rather than the twelve non-CJK locales'
letters: Latin Extended-B and Extended Additional (pinyin, Vietnamese),
combining marks (macOS hands out file names in NFD), Greek, all of Cyrillic,
the punctuation, currency, arrow, math and symbol blocks;
- each CJK face takes its region's **national common-use standard** — GB 2312
level 1 (3755 characters), Big5 level 1 (5401), JIS X 0208 level 1 (2965),
the 2350 KS X 1001 hangul syllables. That is the percentile answer to "which
characters appear in a file name", settled by the frequency work behind each
standard rather than by us, and decoded out of Python's own codecs so no
list is committed and none can go stale;
- CJK punctuation, the halfwidth/fullwidth forms and the enclosed
alphanumerics go in **every** face, because their forms are regional too.
Kana, bopomofo, the enclosed CJK letters (㈱, ㍻) and the symbols that name
files (★, ♪, ✓) go in **one**, because all four are merged into the same
atlas and a second copy would only cost bytes.
Han unification is why there are four CJK faces and not one: the shared
codepoints have different default glyph forms per region, and a Japanese
reader shown the Simplified Chinese face gets kanji in Chinese forms — legible,
and visibly wrong. One shared face would save ~1.3 MB and cost every language
but one its own forms; a shared base plus four deltas saves 18% for a fifth
file, since only 5459 of the 10 178 shared codepoints have identical outlines.
Not covered, and not fixable by adding glyphs alone: Hebrew, Arabic, Thai and
the Indic scripts. ImGui applies no bidi and no shaping, so those would come
out reversed or unjoined rather than right. Combining marks are covered but
not positioned, so an NFD `é` draws the accent beside the letter rather than
over it.
### Keeping the subsets in sync
The subsets are generated from the catalogs, so editing a translation can
outgrow them, and the symptom is one hollow box mid-sentence. Two tools:
The subsets are partly generated from the catalogs, so editing a translation
can outgrow them, and the symptom is one `?` mid-sentence. Two tools:
```bash
python3 tools/make_ui_font.py # rebuild all five (needs fonttools + network)
@@ -89,22 +109,20 @@ fontTools — it parses `cmap` by hand — so it is cheap enough to live there.
### `SS_FONT_CJK` — the *full* faces
The full faces are still worth having, for a much smaller job than they used
to do: dataset paths, file names and typed mask prompts are user data and can
hold any character at all, which no subset of our own strings can anticipate.
The full faces cover the tail the common-use standards leave out: a rare
surname, a classical form, a character outside any national list.
| value | what ships | user experience |
|---|---|---|
| `fetch` *(default)* | one self-contained executable | the interface is complete already; the picker offers the full face for file names, with the size shown |
| `none` | one self-contained executable | same, minus the offer; CJK **file names** may show boxes |
| `none` | one self-contained executable | same, minus the offer; a rare CJK character may show as `?` |
| `sc` \| `tc` \| `jp` \| `kr` | executable + `fonts/<face>.otf` beside it | full coverage offline in that region |
| `all` | executable + all four faces (~23 MB) | full coverage offline everywhere |
The full faces are **bundled, not embedded**, and deliberately:
`ss_embed_file()` turns one byte into five characters of C source, so `all`
would be a 130 MB array literal. A regional build installs a `fonts/` directory
next to the executable, which is the first place `src/app/gui/Fonts.cpp`
looks.
`ss_embed_file()` costs four characters of C source per byte, so `all` would be
a 92 MB literal. A regional build installs a `fonts/` directory next to the
executable, which is the first place `src/app/gui/Fonts.cpp` looks.
## Where fonts are looked for
@@ -122,7 +140,28 @@ stale under us.
A face already on disk is loaded **whatever the interface language is**, not
only for CJK locales: an English UI still has to draw a dataset path like
`C:\写真\` without turning it into boxes.
`C:\写真\` without turning it into `?`.
## Text encoding
Everything in this codebase is UTF-8: the sources (`/utf-8` on MSVC), the
catalogs, the settings files, and every `std::string` holding a path.
On Windows that is not the default, and the mismatch is not a display problem
but a data one — a path outside the machine's legacy ANSI code page survives
`CreateFileA`, `getenv`, `argv` and `CreateProcessA` as literal `?`.
`src/app/utf8.manifest` is embedded in every executable and sets the process
ANSI code page to UTF-8, which fixes all of them at once. It needs **Windows
10 1903 or newer**; an older build silently keeps the legacy code page, and a
non-ASCII dataset path will not open there.
A console keeps a code page of its own, which the manifest does not touch, so
`main()` sets it to UTF-8 and restores it on exit. Pipes are unaffected either
way — a child process's output reaches the GUI as raw bytes — so the
reconstruction and meshing logs were never at risk.
The one thing the manifest does not reach is the desktop file picker, which
converts explicitly (`src/app/gui/NativeDialog.cpp`, `CP_UTF8`) and always did.
## The state of the translations
+20
View File
@@ -21,8 +21,11 @@
#include <cctype>
#include <cstdio>
#include <cstdlib>
#ifdef _WIN32
#include <io.h>
#define WIN32_LEAN_AND_MEAN
#include <windows.h> // NOMINMAX comes from cmake/SsOptions.cmake
#define isatty _isatty
#define fileno _fileno
#else
@@ -36,6 +39,20 @@ namespace {
namespace cmsg = spirula::i18n::msg::cli;
#ifdef _WIN32
UINT g_console_cp = 0;
// A console keeps a code page of its own, which src/app/utf8.manifest does not
// touch: every CJK path this prints into a cp437 window would be mojibake.
// Restored at exit, because cmd.exe keeps whatever code page it is left with.
void use_utf8_console() {
const UINT prev = GetConsoleOutputCP();
if (prev == 0 || prev == CP_UTF8 || !SetConsoleOutputCP(CP_UTF8)) return;
g_console_cp = prev;
std::atexit([] { SetConsoleOutputCP(g_console_cp); });
}
#endif
// The subcommand NAME is an identifier and prints as it is written; the
// summary is a message, so `spirula --help` follows --lang like everything
// else. It is a pointer rather than a copy because a Msg is immortal .rodata
@@ -141,6 +158,9 @@ int main(int argc, char** argv) {
std::setvbuf(stdout, nullptr, _IOLBF, 0);
#endif
}
#ifdef _WIN32
use_utf8_console();
#endif
// --lang is handled here and removed from argv, so no tool's own parser
// has to know about it. The chain that decides the language is in
+1 -2
View File
@@ -176,8 +176,7 @@ void FontSet::rebuild() {
// The full face for the current language, when one has been fetched. It
// goes ahead of the subsets because it is the same design with far more
// coverage: file names, typed prompts, anything the UI's own vocabulary
// does not contain.
// coverage: the characters a national common-use standard leaves out.
if (!_cjk_data.empty()) {
ImFontConfig cjk;
cjk.MergeMode = true;
+10 -30
View File
@@ -1,34 +1,14 @@
#pragma once
// The glyphs. Which font the UI draws with, and how the CJK faces get onto
// disk.
// Which face the UI draws with, and how the full CJK faces get onto disk.
//
// Three facts shape everything here:
//
// * ImGui's built-in font is ASCII-only. German umlauts, French accents,
// Turkish dotless i and Cyrillic were all broken before this file existed,
// never mind Japanese -- so a Latin/Cyrillic face is EMBEDDED and always
// loaded. That is assets/fonts/SpirulaUI-Regular.ttf, 59 KB.
//
// * A full CJK face is 4-8 MB and there are four of them, because Han
// unification means the shared codepoints render with different default
// glyph forms per region. A Japanese reader shown the Simplified Chinese
// face gets kanji in Chinese forms -- legible, and visibly wrong. Four
// times 4-8 MB is too much to embed.
//
// * But this program only ever writes ~600 characters per region: its own
// translations. Subset to those and a regional face is ~110 KB, so all
// FOUR are embedded (assets/fonts/SpirulaCJK-*.otf, 422 KB together).
// That is what makes a default build render all thirteen languages, in
// the right regional forms, with nothing to download -- which matters
// most in the language picker, the one screen a user who cannot read the
// current UI language has to be able to read.
//
// The full faces are still fetched on demand, but for a much smaller job than
// they used to do: dataset paths, file names and mask prompts are user data
// and can hold any character at all, and no subset can cover that. Until one
// is fetched, a folder called C:\写真\ may render as boxes; the UI itself
// never does.
// Five faces are embedded and always loaded: one Latin/Greek/Cyrillic and one
// CJK per region, 3.3 MB together. Between them they cover the interface in
// all thirteen languages AND an ordinary file name in any of them with
// nothing to download, which is the whole design; the reasoning, the sizes
// and what is still not covered are in assets/fonts/README.md. The full
// 4-8 MB faces are still fetched on demand, for the tail a national
// common-use standard leaves out.
#include "i18n/Message.h"
@@ -82,8 +62,8 @@ public:
void invalidate() { _dirty = true; }
// The full face for the current language, when it is not installed. NOT
// an error state: the UI reads fine without it. It is what the language
// menu offers so that CJK file names and typed prompts render too.
// an error state: the UI and ordinary file names both render without it.
// It is what the language menu offers, for the rare characters it adds.
const CjkFace* optional_face() const { return _optional; }
// A build with SS_FONT_CJK=none has no fetch path, and should say so
+15
View File
@@ -0,0 +1,15 @@
<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
<!-- Makes the process ANSI code page UTF-8 (Windows 10 1903+), so the -A Win32
entry points, the CRT, argv and getenv all speak the encoding the rest of
this codebase does. Without it a dataset path or a user profile name
outside the machine's legacy code page reaches CreateFileA as '?'.
Added to every executable's sources in cmake/SsApps.cmake. -->
<assembly xmlns="urn:schemas-microsoft-com:asm.v1" manifestVersion="1.0">
<assemblyIdentity type="win32" name="Spirula.Studio" version="1.0.0.0"/>
<application xmlns="urn:schemas-microsoft-com:asm.v3">
<windowsSettings>
<activeCodePage
xmlns="http://schemas.microsoft.com/SMI/2019/WindowsSettings">UTF-8</activeCodePage>
</windowsSettings>
</application>
</assembly>
+33 -37
View File
@@ -305,46 +305,42 @@ SS_MSG(pick_vocab_tree,
// Language picker
// ===========================================================================
// Shown under the language list when the FULL face for the chosen language is
// not installed. Deliberately not a warning: the interface itself renders
// fine from the embedded subsets (src/app/gui/Fonts.h). What is missing is
// coverage for text this program did not write -- file names, folder names,
// anything typed into a box. {0} is the script ("Japanese"), {1} the size.
// Shown under the language list when the FULL face is not installed. Not a
// warning: the interface and ordinary file names both render already, and what
// is left is the tail. {0} is the language's own name, {1} the size.
SS_MSG(font_needed,
EN("Text outside this program -- file and folder names, what you type -- "
"may show as boxes in {0} until the full font is installed ({1})."),
JA("ファイル名やフォルダー名、入力した文字など、このアプリ以外の{0}は、"
"完全なフォント({1})を入れるまで四角で表示されることがあります。"),
ZH_HANS("文件名、文件夹名和你输入的内容等本程序以外的{0},在安装完整字体"
"({1})之前可能显示为方块。"),
ZH_HANT("檔案名稱、資料夾名稱和你輸入的內容等本程式以外的{0},在安裝完整"
"字型({1})之前可能顯示為方塊。"),
KO("파일 이름과 폴더 이름, 직접 입력한 글자처럼 이 프로그램 밖의 {0}은(는) "
"전체 글꼴({1})을 설치하기 전까지 네모로 보일 수 있습니다."),
DE("Text außerhalb dieses Programms -- Datei- und Ordnernamen, Eingaben -- "
"kann auf {0} als Kästchen erscheinen, bis die vollständige Schriftart "
EN("Common {0} file names render already; a rare character can still show "
"as a box until the full font is installed ({1})."),
JA("よく使う{0}のファイル名はもう表示できます。まれな文字は、完全なフォント"
"({1})を入れるまで四角で表示されることがあります。"),
ZH_HANS("常用的{0}文件名已经可以显示;生僻字在安装完整字体({1})之前仍可能"
"显示为方块。"),
ZH_HANT("常用的{0}檔案名稱已經可以顯示;罕用字在安裝完整字型({1})之前仍"
"可能顯示為方塊。"),
KO("자주 쓰는 {0} 파일 이름은 이미 표시됩니다. 드문 글자는 전체 글꼴({1})을 "
"설치하기 전까지 네모로 보일 수 있습니다."),
DE("Gängige {0}-Dateinamen werden bereits angezeigt; ein seltenes Zeichen "
"kann noch als Kästchen erscheinen, bis die vollständige Schriftart "
"installiert ist ({1})."),
FR("Le texte extérieur à ce programme -- noms de fichiers et de dossiers, "
"ce que vous saisissez -- peut s'afficher en carrés en {0} tant que la "
"police complète n'est pas installée ({1})."),
ES("El texto ajeno a este programa -- nombres de archivos y carpetas, lo "
"que escriba -- puede aparecer como recuadros en {0} hasta que instale "
"la fuente completa ({1})."),
PT("O texto fora deste programa -- nomes de arquivos e pastas, o que você "
"digitar -- pode aparecer como quadrados em {0} até instalar a fonte "
"completa ({1})."),
IT("Il testo esterno a questo programma -- nomi di file e cartelle, ciò "
"che digita -- può apparire come rettangoli in {0} finché non installa "
"il carattere completo ({1})."),
NL("Tekst buiten dit programma -- bestands- en mapnamen, wat u typt -- kan "
"in het {0} als blokjes verschijnen tot het volledige lettertype is "
FR("Les noms de fichiers en {0} courants s'affichent déjà ; un caractère "
"rare peut encore apparaître en carré tant que la police complète "
"n'est pas installée ({1})."),
ES("Los nombres de archivo habituales en {0} ya se muestran; un carácter "
"poco frecuente puede seguir apareciendo como recuadro hasta que "
"instale la fuente completa ({1})."),
PT("Os nomes de arquivo comuns em {0} já aparecem; um caractere raro ainda "
"pode surgir como quadrado até instalar a fonte completa ({1})."),
IT("I nomi di file in {0} più comuni vengono già visualizzati; un carattere "
"raro può ancora apparire come rettangolo finché non installa il "
"carattere completo ({1})."),
NL("Gangbare bestandsnamen in het {0} worden al weergegeven; een zeldzaam "
"teken kan nog als blokje verschijnen tot het volledige lettertype is "
"geïnstalleerd ({1})."),
RU("Текст вне этой программы -- имена файлов и папок, то, что вы вводите, "
"-- может отображаться прямоугольниками на языке «{0}», пока не "
"установлен полный шрифт ({1})."),
TR("Bu programın dışındaki metinler -- dosya ve klasör adları, yazdıklarınız "
"-- tam yazı tipi kurulana kadar {0} dilinde kutu olarak görünebilir "
"({1})."));
RU("Обычные имена файлов на языке «{0}» уже отображаются; редкий символ "
"может по-прежнему показываться прямоугольником, пока не установлен "
"полный шрифт ({1})."),
TR("Yaygın {0} dosya adları zaten görünüyor; ender bir karakter, tam yazı "
"tipi kurulana kadar kutu olarak görünebilir ({1})."));
SS_MSG(font_download,
EN("Install the full font"),
+10 -10
View File
@@ -1,11 +1,11 @@
#!/usr/bin/env python3
"""Fail if a translation uses a character no committed font can draw.
"""Fail if the committed fonts cannot draw something they are asked to.
The fonts in assets/fonts/ are subset to the characters the catalogs used at
the time they were generated (tools/make_ui_font.py). Edit a translation --
add a word, fix a typo, write a full-width comma -- and the subset is stale.
The symptom is one hollow box in the middle of an otherwise fine sentence,
which nobody notices until a user who reads that language does.
The fonts in assets/fonts/ cover what tools/make_ui_font.py asked for when it
generated them: the catalogs, plus the blocks and national common-use sets it
declares for file names. Edit a translation or one of those declarations and
the subsets are stale. The symptom is one `?` in the middle of an otherwise
fine sentence, which nobody notices until a user who reads that language does.
So this runs on every build. It has to be cheap and dependency-free to be
allowed there, which is why it parses `cmap` by hand instead of importing
@@ -15,7 +15,7 @@ downloading 23 MB of upstream fonts.
python3 tools/check_font_coverage.py
The fix, when it fails, is `python3 tools/make_ui_font.py` (needs fonttools
and a network) and committing the four regenerated .otf files.
and a network) and committing the regenerated font files.
"""
import pathlib
@@ -88,13 +88,13 @@ def main() -> None:
wanted = scan_all()
missing = sorted(c for c in wanted if ord(c) > 0x1F and ord(c) not in covered)
if missing:
print("error: the catalogs use characters no committed font can draw:")
print("error: characters asked for that no committed font can draw:")
for c in missing[:40]:
print(f" U+{ord(c):04X} {c}")
if len(missing) > 40:
print(f" ... and {len(missing) - 40} more")
print("\n The subsets in assets/fonts/ are generated from the "
"catalogs and are now stale.")
print("\n The fonts in assets/fonts/ are generated by "
"tools/make_ui_font.py and are now stale.")
print(" Fix: python3 tools/make_ui_font.py (needs fonttools + network)")
raise SystemExit(1)
+109 -52
View File
@@ -7,26 +7,22 @@
assets/fonts/SpirulaCJK-TC.otf | characters this program's own
assets/fonts/SpirulaCJK-KR.otf / translations actually use
Together they are ~490 KB and they are what makes a default build render all
thirteen languages in SS_LANGUAGES with nothing to download. That is worth
being precise about, because the obvious alternative is not:
Together they are 3.3 MB and they cover two things with nothing to download:
this program's own text in all thirteen languages of SS_LANGUAGES, and an
ordinary FILE NAME in any of them. The second is most of the weight and the
reason the character sets below are Unicode blocks and national common-use
standards rather than a scrape of the catalogs -- a path is user data, and one
uncovered character in it is a `?` in the middle of the path.
* A full Noto Sans CJK face is 4-8 MB, and there are four of them, because
Han unification gives the shared codepoints different default glyph forms
per region. Embedding all four is 23 MB of executable to render a menu.
* Downloading one on demand -- what this GUI did before -- means the first
thing a Japanese user sees after picking their language is a wall of
boxes and a download button. The language picker is exactly where a user
who cannot read the current UI language has arrived, so it is exactly the
wrong place to require reading.
Embedding all four full faces instead would be 23 MB, and fetching one on
demand -- what this GUI did before -- puts a wall of `?` and a download
button in front of a user at the language picker, which is the one screen
someone who cannot read the current UI language has arrived at. The full
faces are still fetched (src/app/gui/Fonts.h), now only for the tail a
common-use standard leaves out.
Subsetting to the ~600 characters per region the UI actually writes costs
~110 KB per face, so all four fit. The full faces are still fetched on demand
(src/app/gui/Fonts.h) -- not for the UI, but for dataset paths and file names,
which are user data and can hold any character at all.
THE SUBSETS ARE DERIVED FROM THE CATALOGS. Edit a translation and they are
stale, which would show up as one boxed character in the middle of a sentence.
THE SUBSETS ARE PARTLY DERIVED FROM THE CATALOGS. Edit a translation and they
can go stale, which shows up as one `?` in the middle of a sentence.
tools/check_font_coverage.py is the guard: it runs on every build, needs no
network and no fontTools, and fails when a catalog uses a character no
committed font has.
@@ -58,6 +54,7 @@ import re
import subprocess
import sys
import tempfile
import unicodedata
import urllib.request
import zipfile
@@ -73,28 +70,33 @@ SOURCE_SANS_URL = (
)
SOURCE_SANS_MEMBER = "TTF/SourceSans3-Regular.ttf"
# What the twelve non-CJK locales in SS_LANGUAGES actually need. Ranges, not
# a character list, because a translator will reach for a dash or a quote mark
# nobody enumerated and tofu in the middle of a sentence is worse than 3 KB.
# What the Latin face has to draw. Not "what the twelve non-CJK locales need":
# most of this is for FILE NAMES, which are user data in any locale and where a
# single uncovered character is a `?` in the middle of a path.
UNICODES = ",".join([
"U+0020-007E", # Basic Latin
"U+00A0-00FF", # Latin-1: de/fr/es/pt/it/nl accents
"U+0100-017F", # Latin Extended-A: Turkish g-breve, s-cedilla, dotted I
"U+0192", # florin, occasionally used as a function sign
"U+01FA-01FF",
"U+02C6-02DD", # spacing modifiers (circumflex, caron, ...)
"U+0394,U+03A9,U+03BC,U+03C0", # delta/omega/mu/pi -- they appear in help text
"U+0400-045F", # Cyrillic: Russian
"U+0490-0491", # Ukrainian ghe -- free, and stops one obvious hole
"U+0180-024F", # Latin Extended-B: pinyin tone marks, Azerbaijani
"U+0250-02FF", # IPA and spacing modifiers -- the Hawaiian okina is here
"U+0300-036F", # combining marks: macOS hands out file names in NFD
"U+0370-03FF", # Greek
"U+0400-04FF", # Cyrillic, whole block -- Serbian and Macedonian too
"U+1E00-1EFF", # Latin Extended Additional: Vietnamese
"U+2000-206F", # dashes, curly quotes, ellipsis, bullet
"U+20AC", # euro
"U+2122", # trademark
"U+2190-2193", # arrows -- the UI writes "-> outputs/scene"
"U+2202,U+2206,U+2211,U+2212,U+221A,U+221E,U+2248,U+2260,U+2264,U+2265",
"U+25A0,U+25CF", # square/circle bullets
"U+2070-209F", # superscripts and subscripts
"U+20A0-20BF", # currency: dong, won, rupee, ruble, lira
"U+2100-214F", # letterlike: the numero sign a Russian path uses
"U+2150-218F", # fractions and Roman numerals
"U+2190-21FF", # arrows -- the UI writes "-> outputs/scene"
"U+2200-22FF", # math operators
"U+2300-23FF", # misc technical: the macOS command and option keys
"U+25A0-25FF", # geometric shapes
"U+2600-26FF", # misc symbols: stars and music notes name a lot of files
"U+2700-27BF", # dingbats: check and cross marks
"U+FB01-FB02", # fi/fl ligatures
"U+FFFD", # replacement character -- what a bad byte should look like
])
# Not U+FFFD: no Source Sans or Noto face has it, so ImGui falls back to '?'.
LAYOUT_FEATURES = "kern,liga,ccmp,locl,mark,mkmk"
@@ -102,16 +104,63 @@ LAYOUT_FEATURES = "kern,liga,ccmp,locl,mark,mkmk"
# glyph forms that language is read in.
TAG_REGION = {"JA": "jp", "ZH_HANS": "sc", "ZH_HANT": "tc", "KO": "kr"}
# Characters every regional subset carries whether or not that language's
# translations use them.
#
# - the native name of every language in SS_LANGUAGES, because the language
# picker draws all thirteen at once and the whole point of this exercise
# is that a user can read their own language's name in it
# - SS_LANG_MENU_ICON, the language menu's label
# - CJK punctuation, because the alternative to two dozen free glyphs is
# regenerating four fonts because a translator wrote a full-width comma
PUNCTUATION = "、。,.:;?!「」『』()〔〕【】・…‥ー~〜/\%+-=<> "
# Blocks every regional subset carries whatever its own translations use.
# Regional forms differ across all three, so each face carries its own copy.
CJK_ALWAYS = [
(0x2460, 0x24FF), # enclosed alphanumerics
(0x3000, 0x303F), # CJK symbols and punctuation
(0xFF01, 0xFFEE), # halfwidth and fullwidth forms
]
# What ONE face carries for all of them: the four subsets share an atlas
# (src/app/gui/Fonts.cpp), so a second copy would buy nothing but bytes.
CJK_SINGLE = {
"jp": ([(0x3041, 0x30FF), (0x31F0, 0x31FF), # kana
(0x3200, 0x32FF), (0x3300, 0x33FF)], # enclosed CJK, square forms
# Of U+2600-27BF Noto Sans CJK has under a fifth, so these are named.
"★☆♪♫♬♩♥♡♠♣♦"
"◆◇○●◎△▲▽▼□■"
"※〒℃℉♂♀✓✔✗✘"
"☀☁☂☃☺☻❀✿❤➡"),
"tc": ([(0x3105, 0x312F)], ""), # bopomofo
}
# Assigned, but Noto Sans CJK draws nothing for it, so asking would only fail
# tools/check_font_coverage.py -- correctly, since there is no glyph to embed.
NO_GLYPH = {0x332C} # SQUARE PAATU
# The national common-use standard per region -- (codec, lead bytes, the block
# to keep). Decoded out of Python's own codecs so no character list has to be
# committed, downloaded, or kept in step with anything.
COMMON_USE = {
"sc": ("gb2312", range(0xB0, 0xD8), (0x3400, 0x9FFF)), # GB 2312 L1, 3755
"tc": ("big5", range(0xA4, 0xC7), (0x3400, 0x9FFF)), # Big5 L1, 5401
"jp": ("euc_jp", range(0xB0, 0xD0), (0x3400, 0x9FFF)), # JIS X 0208 L1, 2965
"kr": ("euc_kr", range(0xB0, 0xC9), (0xAC00, 0xD7A3)), # KS X 1001, 2350
}
_TRAIL = list(range(0x40, 0x7F)) + list(range(0xA1, 0xFF))
def common_use(region: str) -> set:
"""The region's common-use characters, decoded from its national standard."""
codec, lead, (lo, hi) = COMMON_USE[region]
out = set()
for a in lead:
for b in _TRAIL:
try:
c = bytes([a, b]).decode(codec)
except UnicodeDecodeError:
continue
if lo <= ord(c) <= hi:
out.add(c)
return out
def _expand(ranges) -> set:
"""Ranges -> characters, minus the codepoints no font could draw."""
return {chr(c) for lo, hi in ranges for c in range(lo, hi + 1)
if c not in NO_GLYPH and unicodedata.category(chr(c)) != "Cn"}
# ---------------------------------------------------------------------------
@@ -155,13 +204,20 @@ def native_names() -> str:
return "".join(r[2] for r in _languages_h())
def always() -> set:
"""What every regional subset carries regardless of its translations."""
def menu_icon() -> str:
"""SS_LANG_MENU_ICON -- the language menu's label."""
text = LANGUAGES_H.read_text(encoding="utf-8")
m = re.search(r'#define\s+SS_LANG_MENU_ICON\s+"([^"]*)"', text)
if not m:
raise SystemExit(f"could not read SS_LANG_MENU_ICON out of {LANGUAGES_H}")
return set(PUNCTUATION) | set(m.group(1)) | set(native_names())
return m.group(1)
def always(region: str) -> set:
"""What a regional subset carries regardless of its own translations."""
ranges, named = CJK_SINGLE.get(region, ([], ""))
return (set(menu_icon()) | set(native_names()) | common_use(region)
| set(named) | _expand(CJK_ALWAYS + ranges))
def _tag_hits(text: str, tag: str):
@@ -180,7 +236,7 @@ def scan_catalogs() -> dict:
Shared with tools/check_font_coverage.py, which imports it.
"""
per = {r: set(always()) for r in TAG_REGION.values()}
per = {r: always(r) for r in TAG_REGION.values()}
text = _catalog_text()
for tag, region in TAG_REGION.items():
for s in _tag_hits(text, tag):
@@ -191,13 +247,14 @@ def scan_catalogs() -> dict:
def scan_all() -> set:
"""Every character the UI can draw from its own catalogs, all languages.
"""Everything the committed fonts have to cover between them.
The union the committed fonts have to cover between them. Tags are derived
from SS_LANGUAGES (`zh_hans` -> `ZH_HANS`), so a new language is scanned
the moment it exists rather than when someone remembers this file.
The catalogs of all thirteen languages plus always(), the part that is
there for file names rather than for anything the UI writes. Tags are
derived from SS_LANGUAGES, so a new language is scanned the moment it
exists rather than when someone remembers this file.
"""
chars = always()
chars = set().union(*(always(r) for r in TAG_REGION.values()))
text = _catalog_text()
for row in _languages_h():
for s in _tag_hits(text, row[0].upper()):