mirror of
https://github.com/harry7557558/spirula-studio.git
synced 2026-10-02 02:44:54 +08:00
fix error loading and displaying file paths with special characters
This commit is contained in:
@@ -43,6 +43,9 @@ MANIFEST
|
||||
*.manifest
|
||||
*.spec
|
||||
|
||||
# Not PyInstaller's: the UTF-8 active-code-page manifest, a committed source.
|
||||
!src/app/utf8.manifest
|
||||
|
||||
# Installer logs
|
||||
pip-log.txt
|
||||
pip-delete-this-directory.txt
|
||||
|
||||
+83
-25
@@ -6,23 +6,49 @@ single Cyrillic letter — which was already wrong, independently of
|
||||
localization, for anyone whose dataset path was not pure ASCII.
|
||||
|
||||
Five faces are committed here and **all five are embedded** in the executable,
|
||||
507 KB together. Nothing has to be downloaded for the interface to render in
|
||||
any of the thirteen languages in `src/i18n/Languages.h`.
|
||||
3.3 MB together. Nothing has to be downloaded for the interface to render in
|
||||
any of the thirteen languages in `src/i18n/Languages.h`, and nothing has to be
|
||||
downloaded for an ordinary file name to render either.
|
||||
|
||||
| | file | size | from |
|
||||
|---|---|---|---|
|
||||
| Latin / Cyrillic | `SpirulaUI-Regular.ttf` | 58 KB | Source Sans 3 |
|
||||
| CJK, per region | `SpirulaCJK-{JP,SC,TC,KR}.otf` | 449 KB | Noto Sans CJK |
|
||||
| Latin / Greek / Cyrillic | `SpirulaUI-Regular.ttf` | 159 KB | Source Sans 3 |
|
||||
| CJK, per region | `SpirulaCJK-{JP,SC,TC,KR}.otf` | 3.2 MB | Noto Sans CJK |
|
||||
|
||||
The full CJK faces — 4–8 MB each — are still **fetched at runtime** or bundled
|
||||
by a regional build, but for a much smaller job: see *The full faces* below.
|
||||
by a regional build, but for a much smaller job than they used to do: see
|
||||
*The full faces* below.
|
||||
|
||||
## Two jobs, two budgets
|
||||
|
||||
The fonts serve two kinds of text and it is worth keeping them apart.
|
||||
|
||||
**The interface** is text this program wrote. It is a closed set — the
|
||||
catalogs under `src/i18n/catalog/` — so it can be covered exactly, and it
|
||||
must be, in the right regional glyph forms, offline. That is the ~600
|
||||
characters per region the subsets started out as.
|
||||
|
||||
**File names are not.** A dataset path, a folder name or a typed mask prompt
|
||||
is user data: it can hold any character at all, and one uncovered character is
|
||||
a `?` in the middle of a path. No subset of *our* strings anticipates
|
||||
`D:\写真\第2回—テスト①\`. Covering that is what the rest of the 3.3 MB buys,
|
||||
and it is why the character sets below are declared as blocks and national
|
||||
standards rather than scraped from the catalogs.
|
||||
|
||||
## `SpirulaUI-Regular.ttf`
|
||||
|
||||
A subset of **Source Sans 3** (Adobe, SIL OFL 1.1 — `OFL-SourceSans.txt`)
|
||||
covering exactly the ranges the twelve non-CJK languages use: Basic Latin,
|
||||
Latin-1, Latin Extended-A, Cyrillic, and the punctuation a translator will
|
||||
reach for. 431 KB upstream, 58 KB after subsetting.
|
||||
A subset of **Source Sans 3** (Adobe, SIL OFL 1.1 — `OFL-SourceSans.txt`),
|
||||
431 KB upstream and 159 KB after subsetting. The ranges are listed in
|
||||
`tools/make_ui_font.py`; the ones that are there for file names rather than
|
||||
for the UI are worth naming:
|
||||
|
||||
| range | why |
|
||||
|---|---|
|
||||
| Latin Extended-B, Extended Additional | pinyin tone marks, Vietnamese |
|
||||
| combining marks (U+0300–036F) | macOS hands out file names in NFD |
|
||||
| Greek, whole Cyrillic block | not just the twelve UI locales' letters |
|
||||
| General Punctuation, U+2100–214F | em dash, curly quotes, ellipsis, the numero sign |
|
||||
| arrows, math, geometric, symbols, dingbats | stars and check marks name a lot of files |
|
||||
|
||||
Source Sans 3 declares **"Source" as a Reserved Font Name**, so a Modified
|
||||
Version may not carry it. Every name-table record was rewritten accordingly —
|
||||
@@ -33,8 +59,24 @@ branding decision.
|
||||
## `SpirulaCJK-{JP,SC,TC,KR}.otf`
|
||||
|
||||
Each is **Noto Sans CJK** (Google, SIL OFL 1.1 — `OFL-NotoSansCJK.txt`) cut
|
||||
down to the ~600 characters that region's own translations use, plus the
|
||||
native name of every language and a little CJK punctuation. ~110 KB each.
|
||||
down to three things:
|
||||
|
||||
1. that region's **national common-use standard** — GB 2312 level 1 (3755
|
||||
characters) for SC, Big5 level 1 (5401) for TC, JIS X 0208 level 1 (2965)
|
||||
for JP, the 2350 KS X 1001 hangul syllables for KR. These are the
|
||||
percentile answer to "which characters appear in a file name", chosen by
|
||||
the frequency studies behind each standard rather than by us, and they are
|
||||
decoded out of Python's own codecs, so no character list is committed here
|
||||
and none can go stale;
|
||||
2. CJK punctuation, the halfwidth/fullwidth forms and the enclosed
|
||||
alphanumerics (①, A, カ) in **every** face, because their forms are
|
||||
regional too — and kana, bopomofo, the enclosed CJK letters (㈱, ㍻) and the
|
||||
symbols that name files (★, ♪, ✓) in **one**, since all four subsets are
|
||||
merged into the same atlas and a second copy would only cost bytes. The
|
||||
symbols are named one by one rather than taken as a block: of U+2600–27BF
|
||||
Noto Sans CJK has under a fifth;
|
||||
3. whatever that region's own translations use, scanned from the catalogs.
|
||||
|
||||
Noto Sans CJK is Source Han Sans under its other name, so it is the same
|
||||
design as the Latin face and a mixed line does not visibly step.
|
||||
|
||||
@@ -45,12 +87,26 @@ wrong. Two alternatives were measured and rejected:
|
||||
|
||||
| approach | size | cost |
|
||||
|---|---|---|
|
||||
| four regional subsets *(what ships)* | 449 KB | — |
|
||||
| one pan-CJK subset | 287 KB | 262 shared characters in the wrong regional form for three of the four languages |
|
||||
| a shared base + three deltas | 358 KB | of the characters used by more than one region, only 267 have identical outlines and 262 genuinely differ — so the deltas are most of the weight anyway |
|
||||
| four regional cuts *(what ships)* | 3.2 MB | — |
|
||||
| one shared face, single glyph forms | ~1.9 MB | every language but one reads file names in another region's forms |
|
||||
| a shared base + four deltas | 2.5 MB | 18% saved for a fifth file and a merge-order hazard; of the 10 178 codepoints only 5459 have identical outlines everywhere |
|
||||
|
||||
Noto reserves no font name, so a subset could legally keep it; these are
|
||||
renamed anyway, because a 600-glyph file called "Noto Sans JP" is a lie.
|
||||
renamed anyway, because a file this small is not the font it was cut from.
|
||||
|
||||
## What is still not covered
|
||||
|
||||
Hebrew, Arabic, Thai, the Indic scripts and the rest have no glyphs in any of
|
||||
the five faces, and adding them would not be enough on its own: ImGui applies
|
||||
no bidi and no shaping, so an Arabic path would come out unjoined and an
|
||||
Indic one unreordered. A path in one of those scripts is still a row of `?`.
|
||||
Combining marks are covered but not *positioned*, for the same reason — an
|
||||
NFD `é` draws as an `e` with the accent alongside rather than over it, which
|
||||
is worse-looking than the composed form and much better than a `?`.
|
||||
|
||||
`?` rather than a hollow box because ImGui picks its fallback glyph from
|
||||
U+FFFD, `?`, space in that order, and neither Source Sans 3 nor Noto Sans
|
||||
CJK has a U+FFFD.
|
||||
|
||||
## Rebuilding
|
||||
|
||||
@@ -58,18 +114,18 @@ renamed anyway, because a 600-glyph file called "Noto Sans JP" is a lie.
|
||||
pip install fonttools
|
||||
python3 tools/make_ui_font.py # all five: download, subset, rename
|
||||
python3 tools/make_ui_font.py --check # verify the committed files
|
||||
python3 tools/check_font_coverage.py # cheap: did a translation outgrow them?
|
||||
python3 tools/check_font_coverage.py # cheap: is anything missing?
|
||||
```
|
||||
|
||||
The output is byte-reproducible, so `--check` is meaningful. The committed
|
||||
files are build artifacts kept in the tree so that the build needs no network
|
||||
— the same reasoning as the committed files under `src/generated/`.
|
||||
|
||||
**The CJK subsets are derived from the catalogs.** Edit a translation and they
|
||||
go stale, and the symptom is one hollow box in the middle of an otherwise fine
|
||||
sentence. `tools/check_font_coverage.py` is the guard and runs on every
|
||||
`build_develop.bash`; it needs neither the network nor fontTools, because it
|
||||
parses `cmap` by hand.
|
||||
**The subsets are partly derived from the catalogs.** Edit a translation and
|
||||
they can go stale, and the symptom is one `?` in the middle of an
|
||||
otherwise fine sentence. `tools/check_font_coverage.py` is the guard and runs
|
||||
on every `build_develop.bash`; it needs neither the network nor fontTools,
|
||||
because it parses `cmap` by hand.
|
||||
|
||||
## The full faces
|
||||
|
||||
@@ -78,10 +134,12 @@ generates `app_generated/cjk_faces.h` from it and `src/app/gui/Fonts.cpp` reads
|
||||
that. Nothing of them is committed here — they are downloaded, verified against
|
||||
the SHA-256 in that file, and cached.
|
||||
|
||||
They exist for the text this program did **not** write: dataset paths, file
|
||||
names and typed mask prompts are user data and can hold any character at all,
|
||||
which no subset of our own strings can anticipate. Until one is fetched, a
|
||||
folder called `C:\写真\` may render as boxes; the interface itself never does.
|
||||
They are what covers the tail: a rare surname, a classical character, anything
|
||||
outside a national common-use list. `SS_FONT_CJK=sc|tc|jp|kr|all` bundles one
|
||||
or all of them into `<exe dir>/fonts/` instead, which is the first place
|
||||
`Fonts.cpp` looks after `$SS_FONT_DIR`. They are bundled rather than embedded
|
||||
because `ss_embed_file()` costs four characters of C source per byte, so `all`
|
||||
would be a 92 MB literal.
|
||||
|
||||
See `docs/i18n.md` for `SS_FONT_CJK`, the download offer, and what a build with
|
||||
`SS_FONT_CJK=none` does instead.
|
||||
|
||||
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
Binary file not shown.
+8
-4
@@ -257,12 +257,13 @@ if(SS_BUILD_CLI OR SS_BUILD_GUI)
|
||||
list(REMOVE_DUPLICATES SS_TOOL_SOURCES)
|
||||
list(REMOVE_DUPLICATES SS_TOOL_LIBS)
|
||||
|
||||
# app.rc carries the icon Explorer and the taskbar draw, which is a link
|
||||
# input on Windows and has no counterpart elsewhere: macOS reads the
|
||||
# bundle's .icns, and Linux has only the window icon GuiMain.cpp sets.
|
||||
# app.rc is the icon Explorer and the taskbar draw; utf8.manifest makes the
|
||||
# PROCESS code page UTF-8, so the -A Win32 calls, the CRT, argv and getenv
|
||||
# agree with "/utf-8". A source, not /MANIFESTINPUT: CMake runs mt.exe too.
|
||||
if(WIN32)
|
||||
enable_language(RC)
|
||||
list(APPEND SS_TOOL_SOURCES ${SS_SRC}/app/app.rc)
|
||||
list(APPEND SS_TOOL_SOURCES ${SS_SRC}/app/app.rc
|
||||
${SS_SRC}/app/utf8.manifest)
|
||||
endif()
|
||||
|
||||
add_executable(spirula ${SS_SRC}/app/Main.cpp ${SS_TOOL_SOURCES})
|
||||
@@ -299,6 +300,9 @@ if(SS_BUILD_CLI OR SS_BUILD_GUI)
|
||||
# itself. Symlink `spirula` if you want those names.
|
||||
if(SS_SEPARATE_TOOLS)
|
||||
function(ss_tool_exe name sources defs libs)
|
||||
if(WIN32)
|
||||
list(APPEND sources ${SS_SRC}/app/utf8.manifest)
|
||||
endif()
|
||||
# ss_i18n is linked explicitly: these targets deliberately do not
|
||||
# link the engine library, and Main.cpp's `--lang` handling needs
|
||||
# it. It is a leaf (cmake/SsI18n.cmake), so this costs nothing.
|
||||
|
||||
+23
-21
@@ -1,20 +1,35 @@
|
||||
# Embedding data files into the executables as byte arrays, so the apps are
|
||||
# self-contained (no runtime lookup of viewer.html or reference/scripts/mask.py).
|
||||
|
||||
# ss_hex_to_literal(<hex string> <out var>)
|
||||
#
|
||||
# Bytes as a C++ string literal, not a `{0x..,}` array: the array form costs
|
||||
# g++ 440 MB of memory per 3 MB embedded, this one 43 MB (measured, g++ 13).
|
||||
function(ss_hex_to_literal hex out)
|
||||
# 40 bytes a line keeps every literal far under MSVC's 64 KB cap. `\x` is
|
||||
# greedy, but every byte is followed by a backslash, so it cannot run on.
|
||||
string(REPEAT "[0-9a-f][0-9a-f]" 40 _grp)
|
||||
string(REGEX REPLACE "(${_grp})" "\\1|" _s "${hex}")
|
||||
string(REGEX REPLACE "([0-9a-f][0-9a-f])" "\\\\x\\1" _s "${_s}")
|
||||
string(REPLACE "|" "\"\n\"" _s "${_s}")
|
||||
set(${out} "\"${_s}\"" PARENT_SCOPE)
|
||||
endfunction()
|
||||
|
||||
# ss_embed_file(<input> <output_header> <symbol>)
|
||||
#
|
||||
# Writes a header defining `k<symbol>[]` / `k<symbol>Size` holding the bytes of
|
||||
# <input>. Regenerated at configure time whenever <input> changes.
|
||||
function(ss_embed_file input output_header symbol)
|
||||
file(READ ${input} _hex HEX)
|
||||
string(REGEX REPLACE "([0-9a-f][0-9a-f])" "0x\\1," _bytes ${_hex})
|
||||
ss_hex_to_literal("${_hex}" _bytes)
|
||||
file(RELATIVE_PATH _rel ${SS_ROOT} ${input})
|
||||
# -1: a string literal brings a terminator the byte count must not include.
|
||||
string(CONCAT _text
|
||||
"#pragma once\n"
|
||||
"// AUTO-GENERATED from ${_rel} -- do not edit.\n"
|
||||
"#include <cstddef>\n"
|
||||
"inline const unsigned char k${symbol}[] = {${_bytes}};\n"
|
||||
"inline const size_t k${symbol}Size = sizeof(k${symbol});\n")
|
||||
"inline const unsigned char k${symbol}[] =\n${_bytes};\n"
|
||||
"inline const size_t k${symbol}Size = sizeof(k${symbol}) - 1;\n")
|
||||
ss_write_if_different(${output_header} "${_text}")
|
||||
set_property(DIRECTORY ${SS_ROOT} APPEND
|
||||
PROPERTY CMAKE_CONFIGURE_DEPENDS ${input})
|
||||
@@ -22,21 +37,8 @@ endfunction()
|
||||
|
||||
# ss_cjk_faces()
|
||||
#
|
||||
# Generates two headers from assets/fonts/cjk_faces.txt:
|
||||
#
|
||||
# app_generated/cjk_faces.h the table of downloadable FULL faces
|
||||
# app_generated/cjk_subsets.h the four SUBSETS, embedded as byte arrays
|
||||
#
|
||||
# and -- for a regional build (SS_FONT_CJK=sc|tc|jp|kr|all) -- downloads the
|
||||
# named full faces so the executable ships with them beside it.
|
||||
#
|
||||
# The subsets are embedded and the full faces are not, which is the whole
|
||||
# design in one line. A subset is ~110 KB because it holds only the characters
|
||||
# this program's own translations use, so all four fit in the executable and
|
||||
# every language renders with nothing to download. A full face is 4-8 MB, and
|
||||
# ss_embed_file() turns one byte into five characters of C source, so `all`
|
||||
# would be a 130 MB array literal. A regional build therefore installs
|
||||
# <exe dir>/fonts/, which is the first place Fonts.cpp looks.
|
||||
# assets/fonts/cjk_faces.txt -> cjk_faces.h (the downloadable full faces) and
|
||||
# cjk_subsets.h (the four embedded subsets). See assets/fonts/README.md.
|
||||
function(ss_cjk_faces)
|
||||
set(spec ${SS_ROOT}/assets/fonts/cjk_faces.txt)
|
||||
set(header ${CMAKE_BINARY_DIR}/app_generated/cjk_faces.h)
|
||||
@@ -87,11 +89,11 @@ function(ss_cjk_faces)
|
||||
" python3 tools/make_ui_font.py")
|
||||
endif()
|
||||
file(READ ${sub} _sub_hex HEX)
|
||||
string(REGEX REPLACE "([0-9a-f][0-9a-f])" "0x\\1," _sub_bytes ${_sub_hex})
|
||||
ss_hex_to_literal("${_sub_hex}" _sub_bytes)
|
||||
string(APPEND sub_arrays
|
||||
"inline const unsigned char kCjkSubset${ID}[] = {${_sub_bytes}};\n")
|
||||
"inline const unsigned char kCjkSubset${ID}[] =\n${_sub_bytes};\n")
|
||||
string(APPEND sub_table
|
||||
" {\"${id}\", kCjkSubset${ID}, sizeof(kCjkSubset${ID})},\n")
|
||||
" {\"${id}\", kCjkSubset${ID}, sizeof(kCjkSubset${ID}) - 1},\n")
|
||||
set_property(DIRECTORY ${SS_ROOT} APPEND
|
||||
PROPERTY CMAKE_CONFIGURE_DEPENDS ${sub})
|
||||
|
||||
|
||||
+74
-35
@@ -42,42 +42,62 @@ environment lands on. That is why it is a language rather than a hard-coded
|
||||
|
||||
## Fonts
|
||||
|
||||
Five faces are **embedded in every build**, 507 KB together:
|
||||
Five faces are **embedded in every build**, 3.3 MB together:
|
||||
|
||||
| file | size | covers |
|
||||
|---|---|---|
|
||||
| `SpirulaUI-Regular.ttf` | 58 KB | Latin + Cyrillic, from Source Sans 3 |
|
||||
| `SpirulaCJK-{JP,SC,TC,KR}.otf` | 449 KB | the ~600 characters *per region* this program's own translations use, from Noto Sans CJK |
|
||||
| `SpirulaUI-Regular.ttf` | 159 KB | Latin, Greek, Cyrillic, punctuation and symbols, from Source Sans 3 |
|
||||
| `SpirulaCJK-{JP,SC,TC,KR}.otf` | 3.2 MB | each region's national common-use set, kana, hangul, bopomofo, the fullwidth and enclosed forms, and this program's own translations, from Noto Sans CJK |
|
||||
|
||||
That is what makes a default build render all thirteen languages, in the right
|
||||
regional glyph forms, with nothing to download. It matters most in the language
|
||||
picker, which is exactly the screen a user who cannot read the current
|
||||
interface language has arrived at — a download button there is a request to
|
||||
read something they came to that menu because they could not read.
|
||||
They do two different jobs and it is worth keeping them apart.
|
||||
|
||||
Subsetting is what makes it affordable. A full Noto Sans CJK face is 4–8 MB and
|
||||
there are four of them, because **Han unification** gives the shared codepoints
|
||||
different default glyph forms per region; a Japanese reader shown the
|
||||
Simplified Chinese face gets kanji in Chinese forms, which is legible and
|
||||
visibly wrong. Cutting each one down to the characters the catalogs actually
|
||||
contain costs ~110 KB, so all four fit and nobody gets the wrong forms.
|
||||
**The interface** is a closed set — the catalogs under `src/i18n/catalog/` —
|
||||
so it can be covered exactly, and it must be, in the right regional glyph
|
||||
forms, offline. That is what makes a default build render all thirteen
|
||||
languages with nothing to download, and it matters most in the language
|
||||
picker: the screen a user who cannot read the current interface language has
|
||||
arrived at, where a download button is a request to read something they came
|
||||
to that menu because they could not read.
|
||||
|
||||
Alternatives measured and rejected: one pan-CJK subset is 287 KB, saving 135 KB
|
||||
but rendering 262 shared characters in Simplified Chinese forms for three of
|
||||
the four languages; delta-encoding the four against a common base lands at
|
||||
358 KB, because of the shared characters only about half have identical
|
||||
outlines across the regional cuts.
|
||||
**File names are not a closed set.** A dataset path, a folder name or a typed
|
||||
mask prompt is user data — it can hold any character, and one uncovered
|
||||
character is a `?` in the middle of a path. That is what the rest of the
|
||||
3.3 MB buys, and why the character sets are declared as Unicode blocks and
|
||||
national standards rather than scraped from the catalogs:
|
||||
|
||||
The atlas merges the base face, then the full regional face if one has been
|
||||
fetched, then all four subsets with the **current language's region first** —
|
||||
first source wins per codepoint, so the current language gets its own forms
|
||||
while the other three still supply what only they have (hangul exists only in
|
||||
the KR face, the simplified 简 of 简体中文 only in SC).
|
||||
- the Latin face takes whole blocks rather than the twelve non-CJK locales'
|
||||
letters: Latin Extended-B and Extended Additional (pinyin, Vietnamese),
|
||||
combining marks (macOS hands out file names in NFD), Greek, all of Cyrillic,
|
||||
the punctuation, currency, arrow, math and symbol blocks;
|
||||
- each CJK face takes its region's **national common-use standard** — GB 2312
|
||||
level 1 (3755 characters), Big5 level 1 (5401), JIS X 0208 level 1 (2965),
|
||||
the 2350 KS X 1001 hangul syllables. That is the percentile answer to "which
|
||||
characters appear in a file name", settled by the frequency work behind each
|
||||
standard rather than by us, and decoded out of Python's own codecs so no
|
||||
list is committed and none can go stale;
|
||||
- CJK punctuation, the halfwidth/fullwidth forms and the enclosed
|
||||
alphanumerics go in **every** face, because their forms are regional too.
|
||||
Kana, bopomofo, the enclosed CJK letters (㈱, ㍻) and the symbols that name
|
||||
files (★, ♪, ✓) go in **one**, because all four are merged into the same
|
||||
atlas and a second copy would only cost bytes.
|
||||
|
||||
Han unification is why there are four CJK faces and not one: the shared
|
||||
codepoints have different default glyph forms per region, and a Japanese
|
||||
reader shown the Simplified Chinese face gets kanji in Chinese forms — legible,
|
||||
and visibly wrong. One shared face would save ~1.3 MB and cost every language
|
||||
but one its own forms; a shared base plus four deltas saves 18% for a fifth
|
||||
file, since only 5459 of the 10 178 shared codepoints have identical outlines.
|
||||
|
||||
Not covered, and not fixable by adding glyphs alone: Hebrew, Arabic, Thai and
|
||||
the Indic scripts. ImGui applies no bidi and no shaping, so those would come
|
||||
out reversed or unjoined rather than right. Combining marks are covered but
|
||||
not positioned, so an NFD `é` draws the accent beside the letter rather than
|
||||
over it.
|
||||
|
||||
### Keeping the subsets in sync
|
||||
|
||||
The subsets are generated from the catalogs, so editing a translation can
|
||||
outgrow them, and the symptom is one hollow box mid-sentence. Two tools:
|
||||
The subsets are partly generated from the catalogs, so editing a translation
|
||||
can outgrow them, and the symptom is one `?` mid-sentence. Two tools:
|
||||
|
||||
```bash
|
||||
python3 tools/make_ui_font.py # rebuild all five (needs fonttools + network)
|
||||
@@ -89,22 +109,20 @@ fontTools — it parses `cmap` by hand — so it is cheap enough to live there.
|
||||
|
||||
### `SS_FONT_CJK` — the *full* faces
|
||||
|
||||
The full faces are still worth having, for a much smaller job than they used
|
||||
to do: dataset paths, file names and typed mask prompts are user data and can
|
||||
hold any character at all, which no subset of our own strings can anticipate.
|
||||
The full faces cover the tail the common-use standards leave out: a rare
|
||||
surname, a classical form, a character outside any national list.
|
||||
|
||||
| value | what ships | user experience |
|
||||
|---|---|---|
|
||||
| `fetch` *(default)* | one self-contained executable | the interface is complete already; the picker offers the full face for file names, with the size shown |
|
||||
| `none` | one self-contained executable | same, minus the offer; CJK **file names** may show boxes |
|
||||
| `none` | one self-contained executable | same, minus the offer; a rare CJK character may show as `?` |
|
||||
| `sc` \| `tc` \| `jp` \| `kr` | executable + `fonts/<face>.otf` beside it | full coverage offline in that region |
|
||||
| `all` | executable + all four faces (~23 MB) | full coverage offline everywhere |
|
||||
|
||||
The full faces are **bundled, not embedded**, and deliberately:
|
||||
`ss_embed_file()` turns one byte into five characters of C source, so `all`
|
||||
would be a 130 MB array literal. A regional build installs a `fonts/` directory
|
||||
next to the executable, which is the first place `src/app/gui/Fonts.cpp`
|
||||
looks.
|
||||
`ss_embed_file()` costs four characters of C source per byte, so `all` would be
|
||||
a 92 MB literal. A regional build installs a `fonts/` directory next to the
|
||||
executable, which is the first place `src/app/gui/Fonts.cpp` looks.
|
||||
|
||||
## Where fonts are looked for
|
||||
|
||||
@@ -122,7 +140,28 @@ stale under us.
|
||||
|
||||
A face already on disk is loaded **whatever the interface language is**, not
|
||||
only for CJK locales: an English UI still has to draw a dataset path like
|
||||
`C:\写真\` without turning it into boxes.
|
||||
`C:\写真\` without turning it into `?`.
|
||||
|
||||
## Text encoding
|
||||
|
||||
Everything in this codebase is UTF-8: the sources (`/utf-8` on MSVC), the
|
||||
catalogs, the settings files, and every `std::string` holding a path.
|
||||
|
||||
On Windows that is not the default, and the mismatch is not a display problem
|
||||
but a data one — a path outside the machine's legacy ANSI code page survives
|
||||
`CreateFileA`, `getenv`, `argv` and `CreateProcessA` as literal `?`.
|
||||
`src/app/utf8.manifest` is embedded in every executable and sets the process
|
||||
ANSI code page to UTF-8, which fixes all of them at once. It needs **Windows
|
||||
10 1903 or newer**; an older build silently keeps the legacy code page, and a
|
||||
non-ASCII dataset path will not open there.
|
||||
|
||||
A console keeps a code page of its own, which the manifest does not touch, so
|
||||
`main()` sets it to UTF-8 and restores it on exit. Pipes are unaffected either
|
||||
way — a child process's output reaches the GUI as raw bytes — so the
|
||||
reconstruction and meshing logs were never at risk.
|
||||
|
||||
The one thing the manifest does not reach is the desktop file picker, which
|
||||
converts explicitly (`src/app/gui/NativeDialog.cpp`, `CP_UTF8`) and always did.
|
||||
|
||||
## The state of the translations
|
||||
|
||||
|
||||
@@ -21,8 +21,11 @@
|
||||
|
||||
#include <cctype>
|
||||
#include <cstdio>
|
||||
#include <cstdlib>
|
||||
#ifdef _WIN32
|
||||
#include <io.h>
|
||||
#define WIN32_LEAN_AND_MEAN
|
||||
#include <windows.h> // NOMINMAX comes from cmake/SsOptions.cmake
|
||||
#define isatty _isatty
|
||||
#define fileno _fileno
|
||||
#else
|
||||
@@ -36,6 +39,20 @@ namespace {
|
||||
|
||||
namespace cmsg = spirula::i18n::msg::cli;
|
||||
|
||||
#ifdef _WIN32
|
||||
UINT g_console_cp = 0;
|
||||
|
||||
// A console keeps a code page of its own, which src/app/utf8.manifest does not
|
||||
// touch: every CJK path this prints into a cp437 window would be mojibake.
|
||||
// Restored at exit, because cmd.exe keeps whatever code page it is left with.
|
||||
void use_utf8_console() {
|
||||
const UINT prev = GetConsoleOutputCP();
|
||||
if (prev == 0 || prev == CP_UTF8 || !SetConsoleOutputCP(CP_UTF8)) return;
|
||||
g_console_cp = prev;
|
||||
std::atexit([] { SetConsoleOutputCP(g_console_cp); });
|
||||
}
|
||||
#endif
|
||||
|
||||
// The subcommand NAME is an identifier and prints as it is written; the
|
||||
// summary is a message, so `spirula --help` follows --lang like everything
|
||||
// else. It is a pointer rather than a copy because a Msg is immortal .rodata
|
||||
@@ -141,6 +158,9 @@ int main(int argc, char** argv) {
|
||||
std::setvbuf(stdout, nullptr, _IOLBF, 0);
|
||||
#endif
|
||||
}
|
||||
#ifdef _WIN32
|
||||
use_utf8_console();
|
||||
#endif
|
||||
|
||||
// --lang is handled here and removed from argv, so no tool's own parser
|
||||
// has to know about it. The chain that decides the language is in
|
||||
|
||||
@@ -176,8 +176,7 @@ void FontSet::rebuild() {
|
||||
|
||||
// The full face for the current language, when one has been fetched. It
|
||||
// goes ahead of the subsets because it is the same design with far more
|
||||
// coverage: file names, typed prompts, anything the UI's own vocabulary
|
||||
// does not contain.
|
||||
// coverage: the characters a national common-use standard leaves out.
|
||||
if (!_cjk_data.empty()) {
|
||||
ImFontConfig cjk;
|
||||
cjk.MergeMode = true;
|
||||
|
||||
+10
-30
@@ -1,34 +1,14 @@
|
||||
#pragma once
|
||||
|
||||
// The glyphs. Which font the UI draws with, and how the CJK faces get onto
|
||||
// disk.
|
||||
// Which face the UI draws with, and how the full CJK faces get onto disk.
|
||||
//
|
||||
// Three facts shape everything here:
|
||||
//
|
||||
// * ImGui's built-in font is ASCII-only. German umlauts, French accents,
|
||||
// Turkish dotless i and Cyrillic were all broken before this file existed,
|
||||
// never mind Japanese -- so a Latin/Cyrillic face is EMBEDDED and always
|
||||
// loaded. That is assets/fonts/SpirulaUI-Regular.ttf, 59 KB.
|
||||
//
|
||||
// * A full CJK face is 4-8 MB and there are four of them, because Han
|
||||
// unification means the shared codepoints render with different default
|
||||
// glyph forms per region. A Japanese reader shown the Simplified Chinese
|
||||
// face gets kanji in Chinese forms -- legible, and visibly wrong. Four
|
||||
// times 4-8 MB is too much to embed.
|
||||
//
|
||||
// * But this program only ever writes ~600 characters per region: its own
|
||||
// translations. Subset to those and a regional face is ~110 KB, so all
|
||||
// FOUR are embedded (assets/fonts/SpirulaCJK-*.otf, 422 KB together).
|
||||
// That is what makes a default build render all thirteen languages, in
|
||||
// the right regional forms, with nothing to download -- which matters
|
||||
// most in the language picker, the one screen a user who cannot read the
|
||||
// current UI language has to be able to read.
|
||||
//
|
||||
// The full faces are still fetched on demand, but for a much smaller job than
|
||||
// they used to do: dataset paths, file names and mask prompts are user data
|
||||
// and can hold any character at all, and no subset can cover that. Until one
|
||||
// is fetched, a folder called C:\写真\ may render as boxes; the UI itself
|
||||
// never does.
|
||||
// Five faces are embedded and always loaded: one Latin/Greek/Cyrillic and one
|
||||
// CJK per region, 3.3 MB together. Between them they cover the interface in
|
||||
// all thirteen languages AND an ordinary file name in any of them with
|
||||
// nothing to download, which is the whole design; the reasoning, the sizes
|
||||
// and what is still not covered are in assets/fonts/README.md. The full
|
||||
// 4-8 MB faces are still fetched on demand, for the tail a national
|
||||
// common-use standard leaves out.
|
||||
|
||||
#include "i18n/Message.h"
|
||||
|
||||
@@ -82,8 +62,8 @@ public:
|
||||
void invalidate() { _dirty = true; }
|
||||
|
||||
// The full face for the current language, when it is not installed. NOT
|
||||
// an error state: the UI reads fine without it. It is what the language
|
||||
// menu offers so that CJK file names and typed prompts render too.
|
||||
// an error state: the UI and ordinary file names both render without it.
|
||||
// It is what the language menu offers, for the rare characters it adds.
|
||||
const CjkFace* optional_face() const { return _optional; }
|
||||
|
||||
// A build with SS_FONT_CJK=none has no fetch path, and should say so
|
||||
|
||||
@@ -0,0 +1,15 @@
|
||||
<?xml version="1.0" encoding="UTF-8" standalone="yes"?>
|
||||
<!-- Makes the process ANSI code page UTF-8 (Windows 10 1903+), so the -A Win32
|
||||
entry points, the CRT, argv and getenv all speak the encoding the rest of
|
||||
this codebase does. Without it a dataset path or a user profile name
|
||||
outside the machine's legacy code page reaches CreateFileA as '?'.
|
||||
Added to every executable's sources in cmake/SsApps.cmake. -->
|
||||
<assembly xmlns="urn:schemas-microsoft-com:asm.v1" manifestVersion="1.0">
|
||||
<assemblyIdentity type="win32" name="Spirula.Studio" version="1.0.0.0"/>
|
||||
<application xmlns="urn:schemas-microsoft-com:asm.v3">
|
||||
<windowsSettings>
|
||||
<activeCodePage
|
||||
xmlns="http://schemas.microsoft.com/SMI/2019/WindowsSettings">UTF-8</activeCodePage>
|
||||
</windowsSettings>
|
||||
</application>
|
||||
</assembly>
|
||||
+33
-37
@@ -305,46 +305,42 @@ SS_MSG(pick_vocab_tree,
|
||||
// Language picker
|
||||
// ===========================================================================
|
||||
|
||||
// Shown under the language list when the FULL face for the chosen language is
|
||||
// not installed. Deliberately not a warning: the interface itself renders
|
||||
// fine from the embedded subsets (src/app/gui/Fonts.h). What is missing is
|
||||
// coverage for text this program did not write -- file names, folder names,
|
||||
// anything typed into a box. {0} is the script ("Japanese"), {1} the size.
|
||||
// Shown under the language list when the FULL face is not installed. Not a
|
||||
// warning: the interface and ordinary file names both render already, and what
|
||||
// is left is the tail. {0} is the language's own name, {1} the size.
|
||||
SS_MSG(font_needed,
|
||||
EN("Text outside this program -- file and folder names, what you type -- "
|
||||
"may show as boxes in {0} until the full font is installed ({1})."),
|
||||
JA("ファイル名やフォルダー名、入力した文字など、このアプリ以外の{0}は、"
|
||||
"完全なフォント({1})を入れるまで四角で表示されることがあります。"),
|
||||
ZH_HANS("文件名、文件夹名和你输入的内容等本程序以外的{0},在安装完整字体"
|
||||
"({1})之前可能显示为方块。"),
|
||||
ZH_HANT("檔案名稱、資料夾名稱和你輸入的內容等本程式以外的{0},在安裝完整"
|
||||
"字型({1})之前可能顯示為方塊。"),
|
||||
KO("파일 이름과 폴더 이름, 직접 입력한 글자처럼 이 프로그램 밖의 {0}은(는) "
|
||||
"전체 글꼴({1})을 설치하기 전까지 네모로 보일 수 있습니다."),
|
||||
DE("Text außerhalb dieses Programms -- Datei- und Ordnernamen, Eingaben -- "
|
||||
"kann auf {0} als Kästchen erscheinen, bis die vollständige Schriftart "
|
||||
EN("Common {0} file names render already; a rare character can still show "
|
||||
"as a box until the full font is installed ({1})."),
|
||||
JA("よく使う{0}のファイル名はもう表示できます。まれな文字は、完全なフォント"
|
||||
"({1})を入れるまで四角で表示されることがあります。"),
|
||||
ZH_HANS("常用的{0}文件名已经可以显示;生僻字在安装完整字体({1})之前仍可能"
|
||||
"显示为方块。"),
|
||||
ZH_HANT("常用的{0}檔案名稱已經可以顯示;罕用字在安裝完整字型({1})之前仍"
|
||||
"可能顯示為方塊。"),
|
||||
KO("자주 쓰는 {0} 파일 이름은 이미 표시됩니다. 드문 글자는 전체 글꼴({1})을 "
|
||||
"설치하기 전까지 네모로 보일 수 있습니다."),
|
||||
DE("Gängige {0}-Dateinamen werden bereits angezeigt; ein seltenes Zeichen "
|
||||
"kann noch als Kästchen erscheinen, bis die vollständige Schriftart "
|
||||
"installiert ist ({1})."),
|
||||
FR("Le texte extérieur à ce programme -- noms de fichiers et de dossiers, "
|
||||
"ce que vous saisissez -- peut s'afficher en carrés en {0} tant que la "
|
||||
"police complète n'est pas installée ({1})."),
|
||||
ES("El texto ajeno a este programa -- nombres de archivos y carpetas, lo "
|
||||
"que escriba -- puede aparecer como recuadros en {0} hasta que instale "
|
||||
"la fuente completa ({1})."),
|
||||
PT("O texto fora deste programa -- nomes de arquivos e pastas, o que você "
|
||||
"digitar -- pode aparecer como quadrados em {0} até instalar a fonte "
|
||||
"completa ({1})."),
|
||||
IT("Il testo esterno a questo programma -- nomi di file e cartelle, ciò "
|
||||
"che digita -- può apparire come rettangoli in {0} finché non installa "
|
||||
"il carattere completo ({1})."),
|
||||
NL("Tekst buiten dit programma -- bestands- en mapnamen, wat u typt -- kan "
|
||||
"in het {0} als blokjes verschijnen tot het volledige lettertype is "
|
||||
FR("Les noms de fichiers en {0} courants s'affichent déjà ; un caractère "
|
||||
"rare peut encore apparaître en carré tant que la police complète "
|
||||
"n'est pas installée ({1})."),
|
||||
ES("Los nombres de archivo habituales en {0} ya se muestran; un carácter "
|
||||
"poco frecuente puede seguir apareciendo como recuadro hasta que "
|
||||
"instale la fuente completa ({1})."),
|
||||
PT("Os nomes de arquivo comuns em {0} já aparecem; um caractere raro ainda "
|
||||
"pode surgir como quadrado até instalar a fonte completa ({1})."),
|
||||
IT("I nomi di file in {0} più comuni vengono già visualizzati; un carattere "
|
||||
"raro può ancora apparire come rettangolo finché non installa il "
|
||||
"carattere completo ({1})."),
|
||||
NL("Gangbare bestandsnamen in het {0} worden al weergegeven; een zeldzaam "
|
||||
"teken kan nog als blokje verschijnen tot het volledige lettertype is "
|
||||
"geïnstalleerd ({1})."),
|
||||
RU("Текст вне этой программы -- имена файлов и папок, то, что вы вводите, "
|
||||
"-- может отображаться прямоугольниками на языке «{0}», пока не "
|
||||
"установлен полный шрифт ({1})."),
|
||||
TR("Bu programın dışındaki metinler -- dosya ve klasör adları, yazdıklarınız "
|
||||
"-- tam yazı tipi kurulana kadar {0} dilinde kutu olarak görünebilir "
|
||||
"({1})."));
|
||||
RU("Обычные имена файлов на языке «{0}» уже отображаются; редкий символ "
|
||||
"может по-прежнему показываться прямоугольником, пока не установлен "
|
||||
"полный шрифт ({1})."),
|
||||
TR("Yaygın {0} dosya adları zaten görünüyor; ender bir karakter, tam yazı "
|
||||
"tipi kurulana kadar kutu olarak görünebilir ({1})."));
|
||||
|
||||
SS_MSG(font_download,
|
||||
EN("Install the full font"),
|
||||
|
||||
@@ -1,11 +1,11 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Fail if a translation uses a character no committed font can draw.
|
||||
"""Fail if the committed fonts cannot draw something they are asked to.
|
||||
|
||||
The fonts in assets/fonts/ are subset to the characters the catalogs used at
|
||||
the time they were generated (tools/make_ui_font.py). Edit a translation --
|
||||
add a word, fix a typo, write a full-width comma -- and the subset is stale.
|
||||
The symptom is one hollow box in the middle of an otherwise fine sentence,
|
||||
which nobody notices until a user who reads that language does.
|
||||
The fonts in assets/fonts/ cover what tools/make_ui_font.py asked for when it
|
||||
generated them: the catalogs, plus the blocks and national common-use sets it
|
||||
declares for file names. Edit a translation or one of those declarations and
|
||||
the subsets are stale. The symptom is one `?` in the middle of an otherwise
|
||||
fine sentence, which nobody notices until a user who reads that language does.
|
||||
|
||||
So this runs on every build. It has to be cheap and dependency-free to be
|
||||
allowed there, which is why it parses `cmap` by hand instead of importing
|
||||
@@ -15,7 +15,7 @@ downloading 23 MB of upstream fonts.
|
||||
python3 tools/check_font_coverage.py
|
||||
|
||||
The fix, when it fails, is `python3 tools/make_ui_font.py` (needs fonttools
|
||||
and a network) and committing the four regenerated .otf files.
|
||||
and a network) and committing the regenerated font files.
|
||||
"""
|
||||
|
||||
import pathlib
|
||||
@@ -88,13 +88,13 @@ def main() -> None:
|
||||
wanted = scan_all()
|
||||
missing = sorted(c for c in wanted if ord(c) > 0x1F and ord(c) not in covered)
|
||||
if missing:
|
||||
print("error: the catalogs use characters no committed font can draw:")
|
||||
print("error: characters asked for that no committed font can draw:")
|
||||
for c in missing[:40]:
|
||||
print(f" U+{ord(c):04X} {c}")
|
||||
if len(missing) > 40:
|
||||
print(f" ... and {len(missing) - 40} more")
|
||||
print("\n The subsets in assets/fonts/ are generated from the "
|
||||
"catalogs and are now stale.")
|
||||
print("\n The fonts in assets/fonts/ are generated by "
|
||||
"tools/make_ui_font.py and are now stale.")
|
||||
print(" Fix: python3 tools/make_ui_font.py (needs fonttools + network)")
|
||||
raise SystemExit(1)
|
||||
|
||||
|
||||
+109
-52
@@ -7,26 +7,22 @@
|
||||
assets/fonts/SpirulaCJK-TC.otf | characters this program's own
|
||||
assets/fonts/SpirulaCJK-KR.otf / translations actually use
|
||||
|
||||
Together they are ~490 KB and they are what makes a default build render all
|
||||
thirteen languages in SS_LANGUAGES with nothing to download. That is worth
|
||||
being precise about, because the obvious alternative is not:
|
||||
Together they are 3.3 MB and they cover two things with nothing to download:
|
||||
this program's own text in all thirteen languages of SS_LANGUAGES, and an
|
||||
ordinary FILE NAME in any of them. The second is most of the weight and the
|
||||
reason the character sets below are Unicode blocks and national common-use
|
||||
standards rather than a scrape of the catalogs -- a path is user data, and one
|
||||
uncovered character in it is a `?` in the middle of the path.
|
||||
|
||||
* A full Noto Sans CJK face is 4-8 MB, and there are four of them, because
|
||||
Han unification gives the shared codepoints different default glyph forms
|
||||
per region. Embedding all four is 23 MB of executable to render a menu.
|
||||
* Downloading one on demand -- what this GUI did before -- means the first
|
||||
thing a Japanese user sees after picking their language is a wall of
|
||||
boxes and a download button. The language picker is exactly where a user
|
||||
who cannot read the current UI language has arrived, so it is exactly the
|
||||
wrong place to require reading.
|
||||
Embedding all four full faces instead would be 23 MB, and fetching one on
|
||||
demand -- what this GUI did before -- puts a wall of `?` and a download
|
||||
button in front of a user at the language picker, which is the one screen
|
||||
someone who cannot read the current UI language has arrived at. The full
|
||||
faces are still fetched (src/app/gui/Fonts.h), now only for the tail a
|
||||
common-use standard leaves out.
|
||||
|
||||
Subsetting to the ~600 characters per region the UI actually writes costs
|
||||
~110 KB per face, so all four fit. The full faces are still fetched on demand
|
||||
(src/app/gui/Fonts.h) -- not for the UI, but for dataset paths and file names,
|
||||
which are user data and can hold any character at all.
|
||||
|
||||
THE SUBSETS ARE DERIVED FROM THE CATALOGS. Edit a translation and they are
|
||||
stale, which would show up as one boxed character in the middle of a sentence.
|
||||
THE SUBSETS ARE PARTLY DERIVED FROM THE CATALOGS. Edit a translation and they
|
||||
can go stale, which shows up as one `?` in the middle of a sentence.
|
||||
tools/check_font_coverage.py is the guard: it runs on every build, needs no
|
||||
network and no fontTools, and fails when a catalog uses a character no
|
||||
committed font has.
|
||||
@@ -58,6 +54,7 @@ import re
|
||||
import subprocess
|
||||
import sys
|
||||
import tempfile
|
||||
import unicodedata
|
||||
import urllib.request
|
||||
import zipfile
|
||||
|
||||
@@ -73,28 +70,33 @@ SOURCE_SANS_URL = (
|
||||
)
|
||||
SOURCE_SANS_MEMBER = "TTF/SourceSans3-Regular.ttf"
|
||||
|
||||
# What the twelve non-CJK locales in SS_LANGUAGES actually need. Ranges, not
|
||||
# a character list, because a translator will reach for a dash or a quote mark
|
||||
# nobody enumerated and tofu in the middle of a sentence is worse than 3 KB.
|
||||
# What the Latin face has to draw. Not "what the twelve non-CJK locales need":
|
||||
# most of this is for FILE NAMES, which are user data in any locale and where a
|
||||
# single uncovered character is a `?` in the middle of a path.
|
||||
UNICODES = ",".join([
|
||||
"U+0020-007E", # Basic Latin
|
||||
"U+00A0-00FF", # Latin-1: de/fr/es/pt/it/nl accents
|
||||
"U+0100-017F", # Latin Extended-A: Turkish g-breve, s-cedilla, dotted I
|
||||
"U+0192", # florin, occasionally used as a function sign
|
||||
"U+01FA-01FF",
|
||||
"U+02C6-02DD", # spacing modifiers (circumflex, caron, ...)
|
||||
"U+0394,U+03A9,U+03BC,U+03C0", # delta/omega/mu/pi -- they appear in help text
|
||||
"U+0400-045F", # Cyrillic: Russian
|
||||
"U+0490-0491", # Ukrainian ghe -- free, and stops one obvious hole
|
||||
"U+0180-024F", # Latin Extended-B: pinyin tone marks, Azerbaijani
|
||||
"U+0250-02FF", # IPA and spacing modifiers -- the Hawaiian okina is here
|
||||
"U+0300-036F", # combining marks: macOS hands out file names in NFD
|
||||
"U+0370-03FF", # Greek
|
||||
"U+0400-04FF", # Cyrillic, whole block -- Serbian and Macedonian too
|
||||
"U+1E00-1EFF", # Latin Extended Additional: Vietnamese
|
||||
"U+2000-206F", # dashes, curly quotes, ellipsis, bullet
|
||||
"U+20AC", # euro
|
||||
"U+2122", # trademark
|
||||
"U+2190-2193", # arrows -- the UI writes "-> outputs/scene"
|
||||
"U+2202,U+2206,U+2211,U+2212,U+221A,U+221E,U+2248,U+2260,U+2264,U+2265",
|
||||
"U+25A0,U+25CF", # square/circle bullets
|
||||
"U+2070-209F", # superscripts and subscripts
|
||||
"U+20A0-20BF", # currency: dong, won, rupee, ruble, lira
|
||||
"U+2100-214F", # letterlike: the numero sign a Russian path uses
|
||||
"U+2150-218F", # fractions and Roman numerals
|
||||
"U+2190-21FF", # arrows -- the UI writes "-> outputs/scene"
|
||||
"U+2200-22FF", # math operators
|
||||
"U+2300-23FF", # misc technical: the macOS command and option keys
|
||||
"U+25A0-25FF", # geometric shapes
|
||||
"U+2600-26FF", # misc symbols: stars and music notes name a lot of files
|
||||
"U+2700-27BF", # dingbats: check and cross marks
|
||||
"U+FB01-FB02", # fi/fl ligatures
|
||||
"U+FFFD", # replacement character -- what a bad byte should look like
|
||||
])
|
||||
# Not U+FFFD: no Source Sans or Noto face has it, so ImGui falls back to '?'.
|
||||
|
||||
LAYOUT_FEATURES = "kern,liga,ccmp,locl,mark,mkmk"
|
||||
|
||||
@@ -102,16 +104,63 @@ LAYOUT_FEATURES = "kern,liga,ccmp,locl,mark,mkmk"
|
||||
# glyph forms that language is read in.
|
||||
TAG_REGION = {"JA": "jp", "ZH_HANS": "sc", "ZH_HANT": "tc", "KO": "kr"}
|
||||
|
||||
# Characters every regional subset carries whether or not that language's
|
||||
# translations use them.
|
||||
#
|
||||
# - the native name of every language in SS_LANGUAGES, because the language
|
||||
# picker draws all thirteen at once and the whole point of this exercise
|
||||
# is that a user can read their own language's name in it
|
||||
# - SS_LANG_MENU_ICON, the language menu's label
|
||||
# - CJK punctuation, because the alternative to two dozen free glyphs is
|
||||
# regenerating four fonts because a translator wrote a full-width comma
|
||||
PUNCTUATION = "、。,.:;?!「」『』()〔〕【】・…‥ー~〜/\%+-=<> "
|
||||
# Blocks every regional subset carries whatever its own translations use.
|
||||
# Regional forms differ across all three, so each face carries its own copy.
|
||||
CJK_ALWAYS = [
|
||||
(0x2460, 0x24FF), # enclosed alphanumerics
|
||||
(0x3000, 0x303F), # CJK symbols and punctuation
|
||||
(0xFF01, 0xFFEE), # halfwidth and fullwidth forms
|
||||
]
|
||||
|
||||
# What ONE face carries for all of them: the four subsets share an atlas
|
||||
# (src/app/gui/Fonts.cpp), so a second copy would buy nothing but bytes.
|
||||
CJK_SINGLE = {
|
||||
"jp": ([(0x3041, 0x30FF), (0x31F0, 0x31FF), # kana
|
||||
(0x3200, 0x32FF), (0x3300, 0x33FF)], # enclosed CJK, square forms
|
||||
# Of U+2600-27BF Noto Sans CJK has under a fifth, so these are named.
|
||||
"★☆♪♫♬♩♥♡♠♣♦"
|
||||
"◆◇○●◎△▲▽▼□■"
|
||||
"※〒℃℉♂♀✓✔✗✘"
|
||||
"☀☁☂☃☺☻❀✿❤➡"),
|
||||
"tc": ([(0x3105, 0x312F)], ""), # bopomofo
|
||||
}
|
||||
|
||||
# Assigned, but Noto Sans CJK draws nothing for it, so asking would only fail
|
||||
# tools/check_font_coverage.py -- correctly, since there is no glyph to embed.
|
||||
NO_GLYPH = {0x332C} # SQUARE PAATU
|
||||
|
||||
# The national common-use standard per region -- (codec, lead bytes, the block
|
||||
# to keep). Decoded out of Python's own codecs so no character list has to be
|
||||
# committed, downloaded, or kept in step with anything.
|
||||
COMMON_USE = {
|
||||
"sc": ("gb2312", range(0xB0, 0xD8), (0x3400, 0x9FFF)), # GB 2312 L1, 3755
|
||||
"tc": ("big5", range(0xA4, 0xC7), (0x3400, 0x9FFF)), # Big5 L1, 5401
|
||||
"jp": ("euc_jp", range(0xB0, 0xD0), (0x3400, 0x9FFF)), # JIS X 0208 L1, 2965
|
||||
"kr": ("euc_kr", range(0xB0, 0xC9), (0xAC00, 0xD7A3)), # KS X 1001, 2350
|
||||
}
|
||||
|
||||
_TRAIL = list(range(0x40, 0x7F)) + list(range(0xA1, 0xFF))
|
||||
|
||||
|
||||
def common_use(region: str) -> set:
|
||||
"""The region's common-use characters, decoded from its national standard."""
|
||||
codec, lead, (lo, hi) = COMMON_USE[region]
|
||||
out = set()
|
||||
for a in lead:
|
||||
for b in _TRAIL:
|
||||
try:
|
||||
c = bytes([a, b]).decode(codec)
|
||||
except UnicodeDecodeError:
|
||||
continue
|
||||
if lo <= ord(c) <= hi:
|
||||
out.add(c)
|
||||
return out
|
||||
|
||||
|
||||
def _expand(ranges) -> set:
|
||||
"""Ranges -> characters, minus the codepoints no font could draw."""
|
||||
return {chr(c) for lo, hi in ranges for c in range(lo, hi + 1)
|
||||
if c not in NO_GLYPH and unicodedata.category(chr(c)) != "Cn"}
|
||||
|
||||
|
||||
# ---------------------------------------------------------------------------
|
||||
@@ -155,13 +204,20 @@ def native_names() -> str:
|
||||
return "".join(r[2] for r in _languages_h())
|
||||
|
||||
|
||||
def always() -> set:
|
||||
"""What every regional subset carries regardless of its translations."""
|
||||
def menu_icon() -> str:
|
||||
"""SS_LANG_MENU_ICON -- the language menu's label."""
|
||||
text = LANGUAGES_H.read_text(encoding="utf-8")
|
||||
m = re.search(r'#define\s+SS_LANG_MENU_ICON\s+"([^"]*)"', text)
|
||||
if not m:
|
||||
raise SystemExit(f"could not read SS_LANG_MENU_ICON out of {LANGUAGES_H}")
|
||||
return set(PUNCTUATION) | set(m.group(1)) | set(native_names())
|
||||
return m.group(1)
|
||||
|
||||
|
||||
def always(region: str) -> set:
|
||||
"""What a regional subset carries regardless of its own translations."""
|
||||
ranges, named = CJK_SINGLE.get(region, ([], ""))
|
||||
return (set(menu_icon()) | set(native_names()) | common_use(region)
|
||||
| set(named) | _expand(CJK_ALWAYS + ranges))
|
||||
|
||||
|
||||
def _tag_hits(text: str, tag: str):
|
||||
@@ -180,7 +236,7 @@ def scan_catalogs() -> dict:
|
||||
|
||||
Shared with tools/check_font_coverage.py, which imports it.
|
||||
"""
|
||||
per = {r: set(always()) for r in TAG_REGION.values()}
|
||||
per = {r: always(r) for r in TAG_REGION.values()}
|
||||
text = _catalog_text()
|
||||
for tag, region in TAG_REGION.items():
|
||||
for s in _tag_hits(text, tag):
|
||||
@@ -191,13 +247,14 @@ def scan_catalogs() -> dict:
|
||||
|
||||
|
||||
def scan_all() -> set:
|
||||
"""Every character the UI can draw from its own catalogs, all languages.
|
||||
"""Everything the committed fonts have to cover between them.
|
||||
|
||||
The union the committed fonts have to cover between them. Tags are derived
|
||||
from SS_LANGUAGES (`zh_hans` -> `ZH_HANS`), so a new language is scanned
|
||||
the moment it exists rather than when someone remembers this file.
|
||||
The catalogs of all thirteen languages plus always(), the part that is
|
||||
there for file names rather than for anything the UI writes. Tags are
|
||||
derived from SS_LANGUAGES, so a new language is scanned the moment it
|
||||
exists rather than when someone remembers this file.
|
||||
"""
|
||||
chars = always()
|
||||
chars = set().union(*(always(r) for r in TAG_REGION.values()))
|
||||
text = _catalog_text()
|
||||
for row in _languages_h():
|
||||
for s in _tag_hits(text, row[0].upper()):
|
||||
|
||||
Reference in New Issue
Block a user