Files
JustVugg fe4f8a9665 fix(deepseek-v41): build on Windows, and ship v41_dsml.py in the image
Two things only CI could find, because this engine had never been built on
Windows and never been put in the container.

deepseek_v41.c included <sys/resource.h> unconditionally. Windows has no such
header; compat.h supplies getrusage over GetProcessMemoryInfo, which is why
every other engine guards that include by platform. Now it does too.
Cross-compiled with the Makefile's own Windows flags (mingw-w64,
-D_FILE_OFFSET_BITS=64 -march=x86-64-v3 -fopenmp, the full warning set): the
engine and both new C tests compile and link clean, no warnings.

docker/Dockerfile.slim copies an explicit list of Python files, and
openai_server.py imports v41_dsml for the checkpoint's tool-call format. The
image therefore built and then failed on `import openai_server` -- exactly the
failure the step's own comment predicts. v41_dsml.py joins v4_dsml.py in the
COPY.
2026-09-12 21:44:04 +02:00

56 lines
2.4 KiB
Docker

# colibri — CPU inference image (multi-stage: build then a slim runtime)
#
# The engine is portable C with a Python launcher/gateway. This image ships
# ONLY the inference path (chat/serve/run/info/plan/doctor) — the ~372 GB
# model is mounted, never baked in, and the converter (which needs torch) is
# deliberately out of scope; convert from a source checkout instead.
#
# docker build -t colibri -f docker/Dockerfile.slim .
# docker run --rm -v /nvme/glm52_i4:/model colibri info
# docker run --rm -it -v /nvme/glm52_i4:/model colibri chat --ram 24
# docker run --rm -p 5000:5000 -v /nvme/glm52_i4:/model colibri \
# serve --host 0.0.0.0 --model-id glm-5.2
#
# --------------------------------------------------------------------------
# Stage 1: build the engine. ARCH is x86-64-v3 (portable AVX2) via `make
# portable`'s target, NOT native — a native build bakes in the builder CPU's
# ISA and SIGILLs on a different host. Override ARCH at build time to pin one.
FROM debian:stable-slim AS build
ARG ARCH=x86-64-v3
RUN apt-get update && \
apt-get install -y --no-install-recommends build-essential make && \
rm -rf /var/lib/apt/lists/*
WORKDIR /src
COPY c/ ./c/
RUN make -C c colibri ARCH=${ARCH}
# --------------------------------------------------------------------------
# Stage 2: runtime. openai_server.py and the launcher use only the Python
# standard library, so no pip install — just python3 + the built binary and
# its support files. libgomp1 provides the OpenMP runtime the engine links.
FROM debian:stable-slim
RUN apt-get update && \
apt-get install -y --no-install-recommends python3 libgomp1 && \
rm -rf /var/lib/apt/lists/* && \
useradd -m -u 1000 colibri
WORKDIR /app
# The engine binary plus the launcher and the files it resolves at runtime
# (version.py, the gateway, resource planner, doctor, tokenizer headers, tools).
COPY --from=build /src/c/colibri ./colibri
COPY c/coli c/version.py c/openai_server.py c/v4_dsml.py c/v41_dsml.py c/family_registry.py c/resource_plan.py c/doctor.py c/autotune.py ./
COPY c/tools/ ./tools/
ENV COLI_MODEL=/model \
COLI_ENGINE=/app/colibri \
COLI_DOCKER_GLM_ONLY=1 \
PYTHONUNBUFFERED=1
# A bind-mounted model lands here; declaring the volume documents the contract.
VOLUME ["/model"]
# `coli serve` default port.
EXPOSE 5000
USER colibri
ENTRYPOINT ["python3", "/app/coli"]
CMD ["info"]