test(glm53): rename the harnesses out of the unittest glob, run the tiny oracles in CI

Seven GLM-5.3 harnesses were argparse scripts named tests/test_glm53_*.py.
`make test-python` discovered them, found no TestCase and counted nothing,
so the chat template, serve, streaming, vision and Vulkan paths were never
exercised while the suite stayed green. Run by hand, most of them exited 0
when they skipped.

- tests/test_glm53_<name>.py -> tests/glm53_<name>_harness.py, the shape
  tests/prefix_serve_harness.py already uses; references follow.
- A skip exits 2 instead of 0 (nine paths).
- tests/test_glm53_oracles.py wraps the two stdlib-only oracles for
  unittest; the glm53-oracle job runs it on the fixtures it already
  generates. A third test asserts SKIP and exit 2 from every harness.
- tests/test_python_discovery.py fails when a collected test_*.py holds no
  tests; the thirteen harnesses of other families still matching the glob
  are listed by name.

Fixes #1700.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This commit is contained in:
Marcello Costagliola
2026-09-23 15:38:45 +02:00
co-authored by Claude Opus 5.5
parent 30c70add3b
commit 69c3b0fb11
12 changed files with 186 additions and 24 deletions
+6
View File
@@ -128,6 +128,12 @@ jobs:
python c/tools/make_edge_tiny_tokenizer.py --vocab-size 281 /tmp/glm53_stream-i4
cd c && COLI_GLM53_FIXTURE=/tmp/glm53_stream-i4 python -m unittest -v tests.test_glm53_dashboard
COLI_GLM53_FIXTURE=/tmp/glm53_stream-i4 python -m unittest -v tests.test_glm53_context_exceeded
# The two stdlib-only oracles used to be argparse scripts that `make
# test-python` collected as zero tests (#1700). They run here, on the
# fixtures generated above, token-exact against transformers in f32.
- name: Run the GLM-5.3 tiny oracles
run: |
cd c && GLM53_TINY=/tmp/glm53_tiny GLM53_MM_TINY=/tmp/glm53_mm python -m unittest -v tests.test_glm53_oracles
macos:
# clang; libomp for the threaded path (Makefile falls back to
+2 -2
View File
@@ -2171,7 +2171,7 @@ def render_chat_glm53(messages, enable_thinking=False, reasoning_effort=None, to
so the existing parser needs nothing added for this family.
The whole thing is pinned byte for byte against chat_template.jinja rendered
with jinja2 (tests/test_glm53_chat_template.py). Getting the prompt nearly
with jinja2 (tests/glm53_chat_template_harness.py). Getting the prompt nearly
right is the failure mode worth guarding: the model answers either way.
"""
if not isinstance(messages, list) or not messages:
@@ -2469,7 +2469,7 @@ def render_chat_dsv41(messages, enable_thinking=False, reasoning_effort=None, to
#
# llama.cpp needs no switch for this because it runs the checkpoint's jinja at request time,
# so `add_generation_prompt=False` costs it nothing. This gateway renders by hand, on purpose
# and for speed (tests/test_glm53_chat_template.py says why), and the bill for that choice is
# and for speed (tests/glm53_chat_template_harness.py says why), and the bill for that choice is
# exactly here: one template flag, one open-turn shape to derive per renderer. Each string
# renderer derives its own, pinned byte-for-byte against the checkpoint's template;
# CONTINUATION_FAMILIES is the set that has done so. Kimi K3 differs in WHERE its shape lives:
@@ -23,7 +23,7 @@ RIFERIMENTO (scaricato 2026-09-10):
hf download zai-org/GLM-5.3-Flash chat_template.jinja
USO:
python3 tests/test_glm53_chat_template.py --template PATH/chat_template.jinja
python3 tests/glm53_chat_template_harness.py --template PATH/chat_template.jinja
"""
import argparse
import json
@@ -49,9 +49,9 @@ def main() -> int:
# le due cose. Il generatore vuole transformers 5.16.1.
print(f"SKIP: manca {arguments.fixture}; generalo con\n"
f" python3 tools/make_glm53_multimodal_tiny.py --output <dir>")
return 0
return 2 # un salto non e' un successo
# Vedi test_glm53_tiny.py: in f32 si prova che il motore implementa il
# Vedi glm53_tiny_harness.py: in f32 si prova che il motore implementa il
# modello, a bit ridotti che la quantizzazione conserva i token.
bits = arguments.bits
if arguments.logit_tolerance is None:
@@ -118,7 +118,7 @@ def main() -> int:
# le due cose. Il generatore vuole transformers 5.16.1.
print(f"SKIP: manca {arguments.fixture}; generalo con\n"
f" python3 tools/make_glm53_multimodal_tiny.py --output <dir>")
return 0
return 2 # un salto non e' un successo
binary = os.path.abspath(arguments.binary)
prompt, tokens = "gu", 4
@@ -61,7 +61,7 @@ def main() -> int:
# le due cose. Il generatore vuole transformers 5.16.1.
print(f"SKIP: manca {arguments.quantized}; generalo con\n"
f" python3 tools/make_glm53_streaming_pair.py --fixture <mm> --output <dir>")
return 0
return 2 # un salto non e' un successo
reference = json.loads((arguments.quantized / "ref.json").read_text())
grid_h, grid_w = reference.get("grid", (0, 0))
@@ -60,7 +60,7 @@ def main() -> int:
print(f"SKIP: manca {arguments.fixture}/ref.json; generalo con\n"
f" pip install -r tools/requirements-glm53-tiny.txt\n"
f" python3 tools/make_glm53_tiny.py --output {arguments.fixture}")
return 0
return 2 # un salto non e' un successo
reference = json.loads((arguments.fixture / "ref.json").read_text())
prompt = ",".join(str(token) for token in reference["prompt_ids"])
expected_forcing = reference["teacher_forcing_ids"]
@@ -18,7 +18,7 @@ identiche: e' l'unica differenza che distingue una vision collegata da una
vision finta.
USO:
python3 tests/test_glm53_vision_serve.py --binary ./glm53 --fixture ~/glm53_mm_tiny
python3 tests/glm53_vision_serve_harness.py --binary ./glm53 --fixture ~/glm53_mm_tiny
"""
import argparse
import json
@@ -77,14 +77,14 @@ def main() -> int:
# le due cose. Il generatore vuole transformers 5.16.1.
print(f"SKIP: manca {arguments.fixture}; generalo con\n"
f" python3 tools/make_glm53_multimodal_tiny.py --output <dir>")
return 0
return 2 # un salto non e' un successo
try:
import numpy
from PIL import Image
except ImportError as problem:
print(f"SKIP: servono numpy e Pillow ({problem}); niente e' stato verificato")
return 0
return 2 # un salto non e' un successo
sys.path.insert(0, str(Path(__file__).resolve().parents[1] / "tools"))
from glm53_image import preprocess, load_config
@@ -99,7 +99,7 @@ def main() -> int:
tokens = (images[0][1] // merge) * (images[0][2] // merge)
if images[1][1:] != images[0][1:]:
print("SKIP: le due immagini di prova hanno griglie diverse")
return 0
return 2 # un salto non e' un successo
prompt = "gu" + IMAGE_OPEN + IMAGE_TOKEN * tokens + IMAGE_CLOSE + "xy"
environment = {**os.environ, "SERVE": "1", "SERVE_BATCH": "1",
@@ -20,7 +20,7 @@ non c'e' nessun device: non ha verificato niente e dirlo verde sarebbe peggio.
USO:
make VK=1 glm53
python3 tests/test_glm53_vulkan.py --binary ./glm53 --fixture ~/glm53_mm_tiny
python3 tests/glm53_vulkan_harness.py --binary ./glm53 --fixture ~/glm53_mm_tiny
"""
import argparse
import os
@@ -64,7 +64,7 @@ def main() -> int:
# le due cose. Il generatore vuole transformers 5.16.1.
print(f"SKIP: missing {arguments.fixture}; generate it with\n"
f" python3 tools/make_glm53_multimodal_tiny.py --output <dir>")
return 0
return 2 # a skip is not a pass
binary = os.path.abspath(arguments.binary)
reference = json.loads((arguments.fixture / "ref.json").read_text())
@@ -75,7 +75,7 @@ def main() -> int:
reason = ("binary was not built with VK=1"
if "Vulkan:" not in notes else "no usable Vulkan device")
print(f"SKIP: {reason}; Vulkan path was not verified")
return 0
return 2 # a skip is not a pass
device = next((line for line in notes.splitlines() if "[VK] ready:" in line), "")
for field in ("teacher_forcing", "greedy"):
+82
View File
@@ -0,0 +1,82 @@
"""The GLM-5.3 harnesses, seen by unittest.
tests/glm53_*_harness.py are argparse programs, not unittest modules. Under
their old test_*.py names `make test-python` imported them, found no TestCase
and counted nothing, so the suite stayed green while the chat template, serve,
streaming, vision and Vulkan paths were never exercised (#1700).
Two things are checked here. Everywhere, with the standard library only: a
harness that cannot find what it needs exits 2, never 0, so a skip cannot pass
for a verification. And the two stdlib-only oracles run against their fixtures
when GLM53_TINY (tools/make_glm53_tiny.py) and GLM53_MM_TINY
(tools/make_glm53_multimodal_tiny.py) point at them, as the GLM-5.3 CI job
does; without them those two tests are skipped with the reason on the record.
"""
import os
import subprocess
import sys
import tempfile
import unittest
from pathlib import Path
HERE = Path(__file__).resolve().parent
BINARY = next((HERE.parent / name for name in ("glm53", "glm53.exe")
if (HERE.parent / name).exists()), None)
def harness(name, *args):
return subprocess.run(
[sys.executable, str(HERE / f"glm53_{name}_harness.py"), *args],
capture_output=True, text=True, timeout=600)
class Glm53HarnessSkipTest(unittest.TestCase):
def test_missing_input_exits_2(self):
"""Every harness, pointed at a fixture that is not there, says SKIP
and exits 2. The binary is never reached, so none is needed."""
with tempfile.TemporaryDirectory() as empty:
missing = str(Path(empty, "missing"))
cases = {
"chat_template": ["--template", missing],
"tiny": ["--binary", "glm53", "--fixture", missing],
"multimodal_tiny": ["--binary", "glm53", "--fixture", missing],
"serve": ["--binary", "glm53", "--fixture", missing],
"streaming": ["--binary", "glm53", "--quantized", missing,
"--dequantized", missing],
"vision_serve": ["--binary", "glm53", "--fixture", missing],
"vulkan": ["--binary", "glm53", "--fixture", missing],
}
self.assertEqual(sorted(cases), sorted(
path.name[len("glm53_"):-len("_harness.py")]
for path in HERE.glob("glm53_*_harness.py")),
"a harness was added or renamed: give it a case here")
for name, args in cases.items():
with self.subTest(harness=name):
result = harness(name, *args)
self.assertEqual(result.returncode, 2, result.stdout + result.stderr)
self.assertIn("SKIP", result.stdout)
class Glm53TinyOracleTest(unittest.TestCase):
def run_oracle(self, name, variable):
fixture = os.environ.get(variable)
if not fixture:
self.skipTest(f"{variable} not set to a fixture")
# Set but unusable is a failure, not a skip: whoever set it expected a run.
self.assertIsNotNone(BINARY, f"{variable} is set but glm53 is not built")
result = harness(name, "--binary", str(BINARY), "--fixture", fixture)
self.assertEqual(result.returncode, 0, result.stdout + result.stderr)
self.assertIn("PASS", result.stdout)
def test_text_oracle(self):
"""KDA, DSA, mHC, dense FFN and routed MoE, token-exact against the
transformers reference in f32."""
self.run_oracle("tiny", "GLM53_TINY")
def test_multimodal_oracle(self):
"""The vision tower and the image tokens in the prompt, token-exact."""
self.run_oracle("multimodal_tiny", "GLM53_MM_TINY")
if __name__ == "__main__":
unittest.main()
+66
View File
@@ -0,0 +1,66 @@
"""Every tests/test_*.py that `make test-python` collects holds tests.
`make test-python` is `unittest discover -p 'test_*.py'`. A standalone argparse
harness that matches the glob is imported, contributes zero tests and still
lets the run report OK, so whatever it checks silently stops being checked
(#1700: seven GLM-5.3 harnesses, the whole chat template, serve, streaming,
vision and Vulkan coverage). A harness belongs under another name, as
tests/glm53_*_harness.py and tests/prefix_serve_harness.py do, with a unittest
wrapper where it can run unattended (tests/test_glm53_oracles.py).
"""
import importlib
import sys
import unittest
from pathlib import Path
HERE = Path(__file__).resolve().parent
# Harnesses that still match the glob. Most run by path from the Makefile, so
# they are exercised there and only count as empty modules here. The list only
# shrinks: rename a harness out of the glob and drop it from here.
KNOWN_HARNESSES = {
"test_deepseek_v4_brio.py",
"test_deepseek_v4_prefix.py",
"test_deepseek_v4_tiny.py",
"test_efficiency_report.py",
"test_glm52_chat_template.py",
"test_kimi_k3_ckpt.py",
"test_kimi_k3_dashboard.py",
"test_kimi_k3_tiny.py",
"test_olmoe_chat_template.py",
"test_qwen36_chat_template.py",
"test_qwen38_chat_template.py",
"test_qwen38_image.py",
"test_qwen38_vision_serve.py",
}
class PythonDiscoveryTest(unittest.TestCase):
def test_every_collected_module_holds_tests(self):
if str(HERE) not in sys.path:
sys.path.insert(0, str(HERE))
loader = unittest.TestLoader()
empty = []
for path in sorted(HERE.glob("test_*.py")):
if path.name in KNOWN_HARNESSES:
continue
try:
module = importlib.import_module(path.stem)
except Exception:
# discover reports an import error, or a SkipTest raised at
# import, as a result of its own: that module is not silent.
continue
if loader.loadTestsFromModule(module).countTestCases() == 0:
empty.append(path.name)
self.assertEqual(empty, [], "these modules match test_*.py but hold no "
"tests, so `make test-python` counts them as passed: "
"rename them out of the glob or add a TestCase")
def test_known_harnesses_still_exist(self):
"""A stale entry would let a new file of the same name slip through."""
self.assertEqual(sorted(name for name in KNOWN_HARNESSES
if not (HERE / name).exists()), [])
if __name__ == "__main__":
unittest.main()
+17 -9
View File
@@ -143,7 +143,7 @@ Tool calling is complete. GLM-5.3 declares tools differently from GLM-5.2 (its
own preamble, its own JSON serialisation, its own spacing inside `<tools>`) but
emits calls identically, so the existing parser handles them unchanged. The
whole rendering is pinned byte for byte against `chat_template.jinja`
(`tests/test_glm53_chat_template.py`).
(`tests/glm53_chat_template_harness.py`).
## Environment
@@ -157,18 +157,26 @@ measured from available memory when unset), `GLM53_MAX_IMAGE_TOKENS`,
```
python3 tools/make_glm53_multimodal_tiny.py --output ~/glm53_mm_tiny
python3 tools/make_glm53_streaming_pair.py --fixture ~/glm53_mm_tiny --output ~/glm53_stream
python3 tests/test_glm53_multimodal_tiny.py --binary ./glm53 --fixture ~/glm53_mm_tiny
python3 tests/test_glm53_streaming.py --binary ./glm53 \
python3 tests/glm53_multimodal_tiny_harness.py --binary ./glm53 --fixture ~/glm53_mm_tiny
python3 tests/glm53_streaming_harness.py --binary ./glm53 \
--quantized ~/glm53_stream-i4 --dequantized ~/glm53_stream-deq
python3 tests/test_glm53_serve.py --binary ./glm53 --fixture ~/glm53_mm_tiny
python3 tests/test_glm53_vision_serve.py --binary ./glm53 --fixture ~/glm53_mm_tiny
python3 tests/test_glm53_chat_template.py --template <model>/chat_template.jinja
make VK=1 glm53 && python3 tests/test_glm53_vulkan.py --binary ./glm53 --fixture ~/glm53_mm_tiny
python3 tests/glm53_serve_harness.py --binary ./glm53 --fixture ~/glm53_mm_tiny
python3 tests/glm53_vision_serve_harness.py --binary ./glm53 --fixture ~/glm53_mm_tiny
python3 tests/glm53_chat_template_harness.py --template <model>/chat_template.jinja
make VK=1 glm53 && python3 tests/glm53_vulkan_harness.py --binary ./glm53 --fixture ~/glm53_mm_tiny
```
The generators want transformers 5.16.1, pinned because an oracle written by a
different version is a different oracle. Each test skips with the command that
builds what it is missing rather than throwing.
different version is a different oracle. Each harness skips with the command that
builds what it is missing rather than throwing, and exits 2 when it does: a skip
verified nothing and must not read as a pass.
The harnesses are named `glm53_*_harness.py` so `make test-python` does not
collect them as empty unittest modules. `tests/test_glm53_oracles.py` wraps the
two stdlib-only oracles for unittest; it runs them when `GLM53_TINY`
(`tools/make_glm53_tiny.py`) and `GLM53_MM_TINY` (the multimodal fixture above)
point at their fixtures, as the GLM-5.3 CI job does, and skips with that reason
otherwise.
Two of the generators refuse to write a fixture that cannot fail: one rejects a
degenerate model that answers the same token everywhere, the other a fixture