GLM5.3-Flash-CIRU-STRIX-IU4 / docs /INSTALLATION.md
jcbtc's picture
Bundle StrixLink 0.3.3 and document direct GPU transport
9378291 verified
|
Raw History Blame Contribute Delete
26.7 kB

Install and run GLM5.3 Flash CIRU STRIX IU4

This build runs across two 128-GiB AMD Strix Halo systems connected by USB4. Use the packaged runtime: a stock vLLM or Transformers installation does not include this model's weight loader, gfx1151 kernels, or DFlash2 path. The default context is 128K; a 256K profile is also available.

The production transport uses NHI (Native Host Interface), the USB4 controller's direct data-transfer interface, through /dev/tbstream0. This avoids the network stack for the model's M8 all-reduce. CiruStrixLink handles link setup and endpoint preparation; it is recommended but not mandatory.

Setup Who manages the link? Model service
StrixLink + NHI, recommended CiruStrixLink Packaged privileged system service
Manual NHI Your existing USB4/NHI setup Your capability-bearing service/launcher
Portable compatibility StrixLink or your existing USB4 network Ordinary user systemd service or foreground launcher

All three use the same model and mirrored generation frontend. Portable sockets are a compatibility option, not the NHI performance path.

Requirements

  • Two gfx1151 Strix Halo hosts with 128 GiB unified memory each; one model request at a time is the supported profile.
  • A direct USB4 connection and compatible kernel/device support. The tested system is NixOS with Linux 7.2.2 and ROCm 10. Other distributions are not claimed as tested deployment targets; the runtime wheels require Ubuntu 24.04+ or an equivalently compatible x86-64/glibc environment.
  • Python 3 and venv support for the download CLI, plus uv, the host build prerequisites, and the pinned runtime installed below.
  • Local NVMe for each rank's approximately 83.25 GiB target weights, the separate draft/runtime, and substantial additional prefix-cache space.
  • Enough free unified memory for the target, draft, 12 GiB GPU KV cache and 8 GiB host prefix staging per rank. Stop competing models as needed; this configuration does not guarantee 20 GB free per host.

The production launcher requires the separately downloaded DFlash2 draft. Its CC BY-NC-ND 4.0 license is independent of the target model; commercial use requires suitable rights. See component licenses.

Throughout this guide, replace PEER_USB4_ADDRESS, RANK_0_USB4_ADDRESS, THIS_HOST_USB4_ADDRESS, and FRONTEND_HOST with your addresses. Each host has a different NODE_RANK (0 or 1) and HOST_IP; both must use the same MASTER_ADDR, MASTER_PORT, EPOCH, context profile, and production settings.

Download the model and draft

Run on each host. Download only that host's rank directory:

sudo install -d -o "$USER" -g "$(id -gn)" \
  /srv/llm/models /srv/llm/runtime /srv/llm/venvs /srv/llm/cache

HF_ENV="${XDG_DATA_HOME:-$HOME/.local/share}/ciru-hf"
if [[ ! -x "$HF_ENV/bin/hf" ]]; then
  python3 -m venv "$HF_ENV"
  "$HF_ENV/bin/pip" install -U huggingface_hub
fi
HF="$HF_ENV/bin/hf"

MODEL_REPO=jcbtc/GLM5.3-Flash-CIRU-STRIX-IU4
MODEL_ROOT=/srv/llm/models/GLM5.3-Flash-CIRU-STRIX-IU4
NODE_RANK=0  # use 1 on the peer

"$HF" download "$MODEL_REPO" \
  --include "config/*" \
  --include "docs/*" \
  --include "runtime/**" \
  --include "rank-${NODE_RANK}/*" \
  --include "systemd/*" \
  --include "context-profiles/*" \
  --include "tools/*" \
  --include "LICENSE" \
  --include "LICENSE_SCOPE.md" \
  --include "THIRD_PARTY_NOTICES.md" \
  --include "launch-node.sh" \
  --include "run-node.sh" \
  --include "install-user-service.sh" \
  --include "install-generation-frontend-user-service.sh" \
  --include "install-nhi-system-service.sh" \
  --include "select-context-profile.sh" \
  --include "select-nhi-context-profile.sh" \
  --local-dir "$MODEL_ROOT"

DFLASH_REVISION=bf582e4eacc1810f76656d1811693ff6c6737d2a
"$HF" download incoai/GLM-5.3-Flash-DFlash2 \
  --revision "$DFLASH_REVISION" \
  --local-dir /srv/llm/models/GLM-5.3-Flash-DFlash2

Keep the packaged rank manifests and shard filenames unchanged. The launcher loads rank-local runtime-prepacked weights rather than re-quantizing the model.

Install the runtime

The self-contained bundle installs the pinned Python 3.14.3, ROCm 10/PyTorch, vLLM, AITER, TileLang, and source trees. Exact artifact versions, compatible host prerequisites, and the detailed install recipe are in runtime/packages/README.md.

On NixOS, import the included module in your host configuration and apply it once on each machine:

{
  imports = [ /path/to/GLM5.3-Flash-CIRU-STRIX-IU4/runtime/packages/nixos-module.nix ];
}
sudo nixos-rebuild switch
cd "$MODEL_ROOT/runtime/packages"
nix-shell ./shell.nix --run 'bash ./INSTALL-RUNTIME.sh'

The module enables nix-ld; the shell supplies host compiler wrappers and development headers. On a compatible non-NixOS host, install the prerequisites from the runtime README and run bash ./INSTALL-RUNTIME.sh directly.

The default installation paths are:

/srv/llm/runtime/vllm-glm53-strix
/srv/llm/runtime/aiter-gfx1151
/srv/llm/venvs/vllm-rocm10-gfx1151
/srv/llm/runtime/vllm-glm53-strix/runtime-env.sh

NHI additionally requires the root-owned runtime staging described below. Do not replace the bundled runtime with a stock PyPI vLLM wheel or omit TileLang; the packaged dependencies are part of the serving path.

Recommended: StrixLink and NHI

The direct GPU path needs the qualified Linux 7.2.2 patch series, including the DMA-BUF import ABI and MSI-X reliability fix. A stock kernel version alone is insufficient. Transport details and microsecond measurements explain what each component provides.

Install StrixLink

The model repository includes CiruStrixLink 0.3.3; no separate GitHub checkout is needed. On both hosts:

VERSION=0.3.3
mkdir -p /tmp/ciru-strixlink-release
"$HF" download "$MODEL_REPO" \
  --include "tools/ciru-strixlink-${VERSION}-linux-amd64.tar.gz" \
            "tools/ciru-strixlink-${VERSION}-SHA256SUMS" \
  --local-dir /tmp/ciru-strixlink-model-release
(cd /tmp/ciru-strixlink-model-release/tools && \
  sha256sum -c "ciru-strixlink-${VERSION}-SHA256SUMS")
tar -xzf \
  "/tmp/ciru-strixlink-model-release/tools/ciru-strixlink-${VERSION}-linux-amd64.tar.gz" \
  -C /tmp/ciru-strixlink-release
sudo install -m 0755 \
  /tmp/ciru-strixlink-release/ciru-strixlink \
  /usr/local/bin/ciru-strixlink
ciru-strixlink prerequisites

Connect the cable. Use role a on one host and role b on the other. Preview the proposed configuration, then apply it:

# First host:
ciru-strixlink setup --role a
sudo ciru-strixlink setup --role a --apply
# Second host:
ciru-strixlink setup --role b
sudo ciru-strixlink setup --role b --apply

Check the connection:

ciru-strixlink doctor --peer PEER_USB4_ADDRESS

# Run on one host:
ciru-strixlink serve
# Run on the other:
ciru-strixlink test \
  --peer PEER_USB4_ADDRESS \
  --duration 7s \
  --streams 4 \
  --output ciru-strixlink-report.json \
  --env-file ciru-strixlink.env

Optional paired Launch page

StrixLink 0.3.3 can coordinate this fixed two-rank model from its third browser tab. Model control is disabled by default. It requires a mode-0600 token file with the same value on both hosts, complementary fixed ranks, fixed USB4 IPv4 peers, a model frontend URL, a loopback-only console, and the narrow packaged helper on each machine.

On the peer model host:

ciru-strixlink agent \
  --token-file TOKEN_FILE \
  --model-control \
  --model-rank 1 \
  --model-peer RANK_0_USB4_ADDRESS

On the console model host:

ciru-strixlink ui \
  --peer RANK_1_USB4_ADDRESS \
  --token-file TOKEN_FILE \
  --model-url http://127.0.0.1:8083 \
  --model-control \
  --model-rank 0

From a desktop, use an authenticated tunnel and open the Launch tab:

ssh -N -L 7749:127.0.0.1:7749 USER@CONSOLE_MODEL_HOST

Open http://127.0.0.1:7749/#/launch. The desktop does not run StrixLink or receive model-host privileges. The page shows exact copyable NixOS and generic Linux permission instructions when control is locked. See the upstream browser-console guide for the complete security and sudoers contract.

To change context, unload the pair, choose the same profile for both ranks in the Launch page, and load again. Context controls intentionally remain locked while the model is active. Loading prevents the portable and NHI forms of this GLM deployment from running together. Unrelated applications remain the operator’s responsibility.

Prepare both NHI endpoints

On each host, collect a status report:

STRIX_CONFIG="${XDG_CONFIG_HOME:-$HOME/.config}/GLM5.3-Flash-CIRU-STRIX-IU4"
mkdir -p "$STRIX_CONFIG"
ciru-strixlink transport status \
  --peer PEER_USB4_ADDRESS \
  --output "$STRIX_CONFIG/rank-${NODE_RANK}.transport.json"

Bring both reports together on one host and reconcile them:

ciru-strixlink transport reconcile \
  --a "$STRIX_CONFIG/rank-0.transport.json" \
  --b "$STRIX_CONFIG/rank-1.transport.json" \
  --output "$STRIX_CONFIG/pair.transport.json"

If arm_allowed=true, preview and apply endpoint preparation on both hosts concurrently:

ciru-strixlink transport endpoint prepare --peer PEER_USB4_ADDRESS
sudo ciru-strixlink transport endpoint prepare \
  --peer PEER_USB4_ADDRESS --apply

Collect new reports, exchange them, and reconcile the prepared pair:

ciru-strixlink transport status \
  --peer PEER_USB4_ADDRESS \
  --output "$STRIX_CONFIG/rank-${NODE_RANK}.prepared.transport.json"

ciru-strixlink transport reconcile \
  --a "$STRIX_CONFIG/rank-0.prepared.transport.json" \
  --b "$STRIX_CONFIG/rank-1.prepared.transport.json" \
  --output "$STRIX_CONFIG/pair-prepared.transport.json"

Make that prepared-pair report available on both hosts. When it reports nhi_ready=true and lease_available=true, each host generates its own environment:

ciru-strixlink transport env \
  --peer PEER_USB4_ADDRESS \
  --mode nhi \
  --runtime vllm \
  --pair-report "$STRIX_CONFIG/pair-prepared.transport.json" \
  --output "$STRIX_CONFIG/ciru-strixlink-nhi.env"

If preparation fails or leaves partial state, use the two-sided stop and cleanup sequence before trying again. The StrixLink guide has additional lifecycle and browser-console instructions.

Privileged direct-NHI system service

NHI DMA-BUF import needs CAP_SYS_RAWIO, a powerful Linux capability. The packaged system unit grants it to a constrained service, not to the ordinary user unit. The service user must belong to render and video; on NixOS, declare both in users.users.<name>.extraGroups and apply the configuration.

Before installing the unit, stage the venv, complete CPython runtime, vLLM source, AITER source, environment script, and gfx1151 libraries under a root-owned path such as /opt/ciru/glm53-iu4, matching systemd/nhi.env.example. The code must be owned by root:root, not group/world-writable, and readable/traversable by the service user. Use copies or reflinks, not hardlinks to user-owned code.

A copied UV venv can still point into a user's home directory. Stage the complete Python 3.14.3 runtime at /opt/ciru/glm53-iu4/python, make venv/bin/python and its aliases resolve to python/bin/python3.14, and set home = /opt/ciru/glm53-iu4/python/bin in venv/pyvenv.cfg. The service installer checks the staging but does not perform that copy/relink. Follow the runtime staging instructions, including the NixOS CC/CXX systemd drop-in. Those values must use the Nix host compiler wrappers, not bare ROCm clang.

Create a dedicated NHI node configuration on each host and edit its paths, rank, addresses, epoch, and local NVMe cache directory:

cp "$MODEL_ROOT/systemd/nhi.env.example" "$STRIX_CONFIG/node-nhi.env"
chmod 0600 "$STRIX_CONFIG/node-nhi.env"

Keep it separate from the portable user service's node.env. Install the unit with the dedicated node configuration and generated transport file:

sudo bash "$MODEL_ROOT/install-nhi-system-service.sh" \
  "$(id -un)" \
  "$STRIX_CONFIG/node-nhi.env" \
  "$STRIX_CONFIG/ciru-strixlink-nhi.env"

Installation reloads systemd but does not start or enable the model. Stop any portable instance on both hosts, choose the same context profile, then start the two NHI units together:

systemctl --user stop GLM5.3-Flash-CIRU-STRIX-IU4.service
sudo glm53-nhi-context "$(id -un)" 2
sudo systemctl start "GLM5.3-Flash-CIRU-STRIX-IU4-nhi@$(id -un).service"

Do not enable the model unit at boot: the pair's endpoint state must be reconciled before starting. If either rank fails, stop its peer as well. Once both rank APIs are ready, start the frontend below.

Required mirrored generation frontend

The model uses two independent rank launchers. Every generation request must reach both rank APIs concurrently with the same body. The packaged frontend does this for you, returning rank 0 and draining rank 1. Send all generation to port 8083, never directly to one rank's port 8100.

Install the frontend user unit on either host:

bash "$MODEL_ROOT/install-generation-frontend-user-service.sh"

Edit ~/.config/GLM5.3-Flash-CIRU-STRIX-IU4/frontend.env:

  • GLM_PRIMARY_URL: rank 0 API, for example http://RANK_0_HOST:8100.
  • GLM_MIRROR_URL: rank 1 API, for example http://RANK_1_HOST:8100.
  • Keep the model ID GLM5.3-Flash-CIRU-STRIX-IU4.
  • The listener defaults to 127.0.0.1:8083. To serve remote clients, bind GLM_FRONTEND_HOST to a private reachable address and put authentication and access control in front of it.

After both model APIs are ready:

systemctl --user start GLM5.3-Flash-CIRU-STRIX-IU4-frontend.service
journalctl --user -u GLM5.3-Flash-CIRU-STRIX-IU4-frontend.service -f

The frontend has no StrixLink dependency. For a custom service manager:

MODEL_ROOT=/srv/llm/models/GLM5.3-Flash-CIRU-STRIX-IU4 \
VLLM_VENV=/srv/llm/venvs/vllm-rocm10-gfx1151 \
GLM_PRIMARY_URL=http://RANK_0_HOST:8100 \
GLM_MIRROR_URL=http://RANK_1_HOST:8100 \
bash "$MODEL_ROOT/runtime/generation-frontend/run.sh"

The frontend mirrors /v1/chat/completions, /v1/completions, and /v1/responses. Use the chat endpoint for normal model interaction so the packaged GLM chat template and reasoning/tool behavior apply. Rank-local port 8100 is for health, metrics, tokenizer/render inspection, and debugging.

Chat API

curl http://FRONTEND_HOST:8083/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "GLM5.3-Flash-CIRU-STRIX-IU4",
    "messages": [
      {"role": "user", "content": "Write a robust Python LRU cache."}
    ],
    "temperature": 1.0,
    "top_p": 0.95,
    "reasoning_effort": "max",
    "chat_template_kwargs": {"clear_thinking": true},
    "max_tokens": 8192
  }'

max_tokens is the new-output allowance, including reasoning. Choose enough for the task while leaving room inside the total context limit; 8192 is only this example's allowance. Add "stream": true for streaming responses.

For tools, pass OpenAI-compatible tools and tool_choice fields, including "tool_choice": "auto" for automatic selection. Execute permitted calls in your client, append the assistant tool call and matching tool-result messages, then send the conversation again. Do not manually insert GLM control tokens. Keep clear_thinking=true in normal multi-turn chat so previous hidden reasoning is not carried forward as conversation content. Apply appropriate tool permissions and confirmations in the client.

Production settings

The launcher already supplies the intended engine settings; retain them when integrating the model into another service manager:

  • vLLM V2, target TP2 with the external launcher, BF16 activations, and hybrid IU4 target weights.
  • DFlash2 k=5, probabilistic draft sampling, and standard rejection sampling.
  • The packaged chat template and generation config; temperature 1.0, top-p 0.95, reasoning effort max, and clear_thinking=true for chat requests.
  • AITER attention enabled, AITER MoE and skinny GEMM disabled, direct-route M1 MoE and the packaged tuned MoE configuration enabled.
  • One sequence, eager execution, asynchronous scheduling disabled, chunked prefill enabled, block size 128, and 2,304 maximum batched tokens.
  • Prefix caching disabled by default. The optional prompt-only disk tier uses 8 GiB host staging and one block per offload chunk, but requires an external disk quota and monitoring before production use.
  • glm47 reasoning/tool parsers and the packaged M8 NHI adapter on a ready NHI pair. The 30,000 ms NHI timeout is a failure deadline; successful transfers return immediately.

MAX_NUM_BATCHED_TOKENS is configurable, but larger values are experimental and are not automatically faster. The default also provides recurrent-state checkpoints needed by the hybrid prefix cache.

Context profiles

Fresh installs choose profile 2. Existing selections are preserved on reinstallation. Apply the same choice to both hosts:

# NHI system service:
sudo glm53-nhi-context "$(id -un)" list
sudo glm53-nhi-context "$(id -un)" 2

# Portable user service:
glm53-context list
glm53-context 2
Number Profile Total request limit GPU KV per rank Host prefix staging
1 64K, lower-memory fallback 65,664 tokens 6 GiB 8 GiB
2 128K, default 131,200 tokens 12 GiB 8 GiB
3 256K, experimental 262,272 tokens 8 GiB 8 GiB

The limit includes prompt, generated reasoning, and answer tokens. The extra 128 tokens are scheduler/speculation headroom above the nominal context. The selectors update only next-start configuration; they do not start or restart the model. Stop both ranks, change both selections, then restart the pair with ready endpoints. No RoPE scaling is applied.

Published performance tests cover the 128K profile, including a 131,000-token input and persistent cache reuse. The 256K/8 GiB profile has passed paired startup, exact-context serving validation, HumanEval 0–9, and a 65,680-token recovery stress probe. That long-context probe tests a minimum generation floor after a large prefill; it is not representative of normal short-prompt chat speed. Its initial depth comparison was not fully host-isolated and is retained only as directional evidence. This is not a complete 256K quality sweep. Available unified memory remains the limiting resource; a 16 GiB KV allocation did not fit the validated 128-GiB-per-host deployment. See the recorded results for measurement scope.

DFlash and cache controls

The launcher reads two per-user files at startup. Keep them identical on both ranks and edit them only while the pair is stopped:

mkdir -p ~/.config/ciru-glm53-iu4
printf '5\n' > ~/.config/ciru-glm53-iu4/dflash-tokens
printf '0\n' > ~/.config/ciru-glm53-iu4/prefix-cache-enabled

dflash-tokens accepts 0 through 7. Zero omits DFlash entirely for a true target-only comparison. Invalid values stop startup; missing files use k=5 and cache-off. The Launch page reports a value as known only when both ranks report the same setting.

Persistent disk prefix cache

Prefix caching is disabled by default. When explicitly enabled, matching system prompts, tool schemas, and documents can be reused across requests and service restarts. Completed prompt blocks pass through an 8 GiB host-memory LRU and are written to local NVMe at PREFIX_CACHE_ROOT. Only prompt blocks are offloaded; generated tails are not written to this cache. Each TP rank keeps its own files. Cached KV blocks are not transferred over USB4.

8 GiB limits RAM staging, not disk storage. The filesystem tier has no built-in byte quota or automatic eviction. The 256K investigation found retained cache trees of 219 GiB on rank 0 and 266 GiB on rank 1. Do not enable this tier on a general-purpose filesystem. Use a dedicated quota-managed filesystem, monitor free space and cache-write failures, and define cleanup behavior before opting in.

Set PREFIX_CACHE_ROOT to a suitable directory on each host. Keeping that directory preserves the disk cache across service restarts. PREFIX_CACHE_ENABLED=1 opts in; 0 is the production default.

The NHI unit's private /dev/shm tmpfs is separate temporary staging scratch, capped at 16 GiB and released when the service stops. This does not delete the persistent NVMe cache. Both GPU KV and host staging consume unified memory; leave space for the OS and other software.

Manual NHI

You can use an existing direct-NHI setup without running the StrixLink app. Your service manager must supply a root-owned runtime and grant the model process CAP_SYS_RAWIO. Use the same model/runtime paths and production settings above, plus:

export TRANSPORT_MODE=nhi
export TRANSPORT_INTERFACE=thunderbolt0
export VLLM_NHI_DEVICE=/dev/tbstream0
export ALLOW_PRIVILEGED_NHI=1

The operator owns endpoint readiness and matching epoch, HopID, ring, and peer state. ALLOW_PRIVILEGED_NHI=1 acknowledges this mode; it does not grant the capability. The ordinary user service cannot run it. Launch a trusted, root-owned copy of run-node.sh and its companion scripts from your capability-bearing service with the per-rank values shown in the foreground example below. Do not mix an existing endpoint owner with StrixLink endpoint preparation.

The packaged NHI service installer accepts a StrixLink-generated transport file. A fully caller-owned setup can instead retain its own service and transport lifecycle; the model does not require the StrixLink process at runtime.

For a service that already sets all RCCL/Gloo and optional NHI variables, the launcher also offers external mode:

export TRANSPORT_MODE=external
export NCCL_SOCKET_IFNAME='=my-usb4-interface'
export GLOO_SOCKET_IFNAME=my-usb4-interface
export VLLM_NHI_M8_ALLREDUCE=0  # set 1 only for your prepared NHI path

External mode preserves caller-supplied transport values. An NHI-enabled external setup still requires the capability and two-sided endpoint handling. MASTER_ADDR must always be the reachable rank-0 address.

Portable compatibility and user services

For an existing USB4 network interface, the launcher can pin RCCL/Gloo to it without StrixLink:

export TRANSPORT_MODE=portable
export TRANSPORT_INTERFACE=thunderbolt0

Alternatively, generate a portable environment with StrixLink after setting up and testing the USB4 network:

STRIX_CONFIG="${XDG_CONFIG_HOME:-$HOME/.config}/GLM5.3-Flash-CIRU-STRIX-IU4"
mkdir -p "$STRIX_CONFIG"
ciru-strixlink transport env \
  --peer PEER_USB4_ADDRESS \
  --mode portable \
  --runtime vllm \
  --output "$STRIX_CONFIG/ciru-strixlink-portable.env"

User systemd service

On each host:

bash "$MODEL_ROOT/install-user-service.sh"

Edit ~/.config/GLM5.3-Flash-CIRU-STRIX-IU4/node.env. Use TRANSPORT_MODE=portable with your interface, strixlink with the generated portable STRIXLINK_ENV, or an unprivileged caller-owned external configuration. See external-transport.env.example. Set the per-rank paths/addresses and shared epoch described at the start of this guide. Installation reloads the user manager but does not start the model.

When both configurations are ready, start both hosts together:

systemctl --user start GLM5.3-Flash-CIRU-STRIX-IU4.service
journalctl --user -u GLM5.3-Flash-CIRU-STRIX-IU4.service -f

Start the same mirrored frontend after both rank APIs are ready. Do not run the portable and NHI model units simultaneously. The user unit is not enabled at boot and does not independently restart a failed rank; start/restart the pair together.

Foreground launch example

This example uses StrixLink's portable environment. For a manually configured USB4 network, replace TRANSPORT_MODE and STRIXLINK_ENV with the portable interface settings above. For manual NHI, use your capability-bearing service and the NHI settings instead.

export MODEL_ROOT=/srv/llm/models/GLM5.3-Flash-CIRU-STRIX-IU4
export DFLASH_MODEL=/srv/llm/models/GLM-5.3-Flash-DFlash2
export VLLM_SOURCE=/srv/llm/runtime/vllm-glm53-strix
export VLLM_VENV=/srv/llm/venvs/vllm-rocm10-gfx1151
export VLLM_RUNTIME_ENV=/srv/llm/runtime/vllm-glm53-strix/runtime-env.sh
export AITER_SOURCE=/srv/llm/runtime/aiter-gfx1151
export TRANSPORT_MODE=strixlink
export STRIXLINK_ENV="${XDG_CONFIG_HOME:-$HOME/.config}/GLM5.3-Flash-CIRU-STRIX-IU4/ciru-strixlink-portable.env"
export MASTER_ADDR=RANK_0_USB4_ADDRESS
export PREFIX_CACHE_ROOT=/path/on/local/nvme/GLM5.3-Flash-CIRU-STRIX-IU4
export NODE_RANK=0                 # use 1 on the peer
export HOST_IP=THIS_HOST_USB4_ADDRESS
export EPOCH=1                     # choose the same new value on both hosts
export API_PORT=8100

bash "$MODEL_ROOT/run-node.sh"

For NHI, replace the /srv/llm runtime paths with your root-owned staging paths. run-node.sh enters the packaged shell.nix on NixOS when available; elsewhere it uses the installed host toolchain.

Stopping and restarting

Always stop and restart both model ranks as a pair. The frontend may remain installed, but do not submit generation while either rank is unavailable. For the packaged NHI service, stop both units before releasing endpoints:

sudo systemctl stop "GLM5.3-Flash-CIRU-STRIX-IU4-nhi@$(id -un).service"
ciru-strixlink transport endpoint cleanup
sudo ciru-strixlink transport endpoint cleanup --apply

Preview cleanup on each host and apply it concurrently. Collect fresh status and reconcile again before preparing the next NHI start. If preparation failed partway through, clean both endpoints. Before switching to portable, stop both NHI units and confirm both endpoints have no remaining holder or partial state. The packaged units do not perform one-sided automatic cleanup, restart, or boot activation.