常见问题Troubleshooting

按「看到什么 → 为什么 → 怎么办」整理。

Organised as: what you see → why → what to do.

构建Building

链接时找不到 zlibThe link step cannot find zlib

error: linking failed
ld: library not found for -lz

原因:cmd/main/moon.pkg 为 native 目标指定了 cc-link-flags = "-lz",而系统里没有 zlib 开发包。

Cause: cmd/main/moon.pkg sets cc-link-flags = "-lz" for the native target, and the zlib development package is not installed.

处理:装上 zlib 的开发包(如 Debian/Ubuntu 的 zlib1g-dev、macOS 上 Xcode 命令行工具自带的版本),再重新构建。

Fix: install zlib's development package (zlib1g-dev on Debian/Ubuntu; it ships with the Xcode command-line tools on macOS), then build again.

目标不被支持The target is not supported

error: target wasm-gc is not supported by this package

原因:moonkafka 的 socket 实现只支持 native,两个包都声明了 supported_targets = "+native"。

Cause: moonkafka's socket layer supports native only, and both packages declare supported_targets = "+native".

处理:每条命令都带上 --target native。

Fix: pass --target native on every command.

连接与读写Connecting and reading

主题不存在The topic does not exist

两个子命令都不建主题。对不存在的主题执行 produce 或 consume 都会失败。先按 快速开始里的命令把主题建好。

Neither subcommand creates topics, so both produce and consume fail against one that does not exist. Create it first, as shown in Getting started.

consume 连上了却什么都不打印consume connects but prints nothing

参数被当成了主机名An argument was taken for a host

invalid port: latest

原因:参数是位置参数,host 与 port 必须成对出现; --max-messages 还会从位置序列里消失,使它后面的参数前移。

Cause: arguments are positional, so host and port must be given as a pair; and --max-messages is removed from the positional sequence, shifting whatever follows it up by one.

处理:位置参数写在前面,flag 放在最后。两种常见写法:

Fix: positionals first, flags last. Two shapes that work:

consume events 127.0.0.1 9092 latest --max-messages 1
consume events --max-messages 1

本地 Kafka 容器The local Kafka container

镜像拉不下来The image will not pull

Error response from daemon: Get "https://registry-1.docker.io/v2/": ...
i/o timeout

原因:网络到 Docker Hub 不通。这与项目无关,但会挡住 docker compose up。

Cause: no working route to Docker Hub. Nothing to do with this project, but it does block docker compose up.

处理:从可用的镜像站拉取,再打回 compose 文件里期望的标签。 仓库使用的镜像是 apache/kafka:4.3.0:

Fix: pull from a reachable mirror and tag it back to the name the compose file expects. The image this repository uses is apache/kafka:4.3.0:

docker pull docker.m.daocloud.io/apache/kafka:4.3.0
docker tag  docker.m.daocloud.io/apache/kafka:4.3.0 apache/kafka:4.3.0
docker compose -f docker-compose.kafka.yml up -d

也可以给 Docker 守护进程配 registry-mirrors,或直接换一个镜像站。 集成测试走 --image 时同理,先把镜像准备好即可。

You can also configure registry-mirrors for the Docker daemon, or use a different mirror. The same applies to the integration test's --image: get the image onto the machine first.

9092 端口被占用Port 9092 is already taken

Error response from daemon: ... bind: address already in use

原因:本机已经有别的 Kafka 或别的服务占着 9092。 处理:停掉占用的进程,或者两边都换一个端口——compose 文件里改 ports 与 KAFKA_ADVERTISED_LISTENERS,CLI 和集成测试则用 --port。

Cause: something else on the machine already holds 9092. Fix: stop it, or move both sides to another port — change ports and KAFKA_ADVERTISED_LISTENERS in the compose file, and pass --port to the CLI and the integration test.

容器起不来或反复重启The container will not start, or keeps restarting

先看它的日志,再确认健康状态:

Read its log first, then check its health:

docker logs moonkafka-demo-kafka
docker inspect --format '{{.State.Health.Status}}' moonkafka-demo-kafka

健康检查就是每 5 秒跑一次 kafka-topics.sh --bootstrap-server localhost:9092 --list; 它一直不通过,通常说明 broker 还没选出 controller,或者被上面的端口问题绊住了。

The healthcheck is kafka-topics.sh --bootstrap-server localhost:9092 --list every 5 seconds. If it never passes, the broker has usually not elected a controller yet, or is stuck on the port problem above.

集成测试The integration test

日志里的「podman is not usable」"podman is not usable" in the log

  $ podman version
  ... podman is not usable
  $ docker version
  -> container runtime: docker

这不是错误。运行时探测按顺序逐个尝试候选,每个失败的候选都会留下这样 一行,随后它选中的引擎会打印 -> container runtime: …。只有所有候选都不 可用时,才会出现 FAIL no working container runtime。

This is not an error. The runtime probe walks its candidates in order and leaves one such line per miss; the engine it settles on then prints -> container runtime: …. Only when every candidate fails do you get FAIL no working container runtime.

想彻底不看这几行,直接指定引擎——make itest 默认就是这么做的 (ITEST_RUNTIME=docker):

To skip those lines entirely, name the engine — which is what make itest does by default (ITEST_RUNTIME=docker):

make itest ITEST_RUNTIME=podman
moon run --target native tests/itest -- --runtime podman

第一次 produce 就失败了The very first produce fails

  ... exit=1 after 61ms
  ! ProtocolError: ...
  ... retrying (2/5): the previous attempt was refused

预期行为。刚启动的 broker 会先接受管理命令与建主题,之后才具备处理 produce 的能力;落在这段窗口里的第一次 produce 会被立即拒绝。harness 因此做最多 5 次 的有界重试,并把每次尝试都打印出来。看到 retrying 之后紧跟 exit=0,就是正常的启动竞态。

Expected. A freshly started broker answers admin commands and accepts topic creation before it can serve a produce; a produce landing in that window is refused immediately. The harness retries a bounded 5 times and logs every attempt, so a retrying line followed by exit=0 is the normal startup race.

如果五次都失败,就值得看一眼 --keep 留下的容器,或者手动发一条确认 broker 是否真的可用。

If all five fail, inspect the container left behind by --keep, or send a message by hand to check whether the broker is really serving.

Kafka was not ready within NsKafka was not ready within Ns

原因:首次运行要拉镜像,或者机器比较慢,超出了默认的 120 秒预算。 处理:加大预算,或先把镜像拉好。

Cause: a first run has to pull the image, or the machine is slow, and it exceeded the default 120-second budget. Fix: raise the budget, or pull the image beforehand.

make itest ITEST_FLAGS="--timeout-secs 300"

the CLI binary is missingthe CLI binary is missing

  FAIL  the CLI binary is missing at _build/native/debug/build/cmd/main/main.exe;
        build it first (moon build --target native) or pass --cli PATH

harness 执行的是构建产物,而不是再套一层 moon run——这样每个阶段都很快, 而且编译错误会在构建期而不是测试中途暴露。先构建,或用 --cli 指向别的 二进制。make itest 会自动先构建,所以这条通常只在直接调用时出现。

The harness executes the built artifact rather than wrapping another moon run — that keeps each phase fast and surfaces compile errors at build time instead of mid-test. Build first, or point --cli at another binary. make itest builds automatically, so this usually only appears when calling the test directly.

--no-container 报错--no-container complains

  FAIL  --no-container cannot create a topic; pass --topic <a fresh, empty topic>

这个模式跳过容器生命周期与建主题,所以必须自己指定一个已存在且为空的主题。阶段二也会 被跳过,因为它需要一个空主题。细节见集成测试。

This mode skips the container lifecycle and topic creation, so you must name an existing, empty topic yourself. Phase 2 is skipped too, since it needs an empty topic. See Integration test.

没有 compose providerNo compose provider

  FAIL  --run-mode compose, but no compose provider answers for podman

--run-mode compose 强制走 compose,但这个引擎两个探测都没通过。要么装 podman-compose,要么改用 --run-mode run(harness 直接 podman run,不需要 provider),要么把它固定为一个二进制:

--run-mode compose forces the compose path, but neither probe answered for this engine. Install podman-compose, switch to --run-mode run (the harness drives podman run itself and needs no provider), or pin an explicit binary:

moon run --target native tests/itest -- --run-mode run
moon run --target native tests/itest -- --compose-bin /usr/local/bin/podman-compose

上一次运行留下的容器A container left by an earlier run

--keep 留下的容器不会影响下一次运行:直接 run 的路径会先 inspect,正在跑就复用,否则先 rm -f 再起。要手动清掉:

A container left by --keep does not disturb the next run: the direct run path inspects it first, reuses it if it is running, and otherwise clears it with rm -f before starting. To remove it by hand:

docker rm -f moonkafka-demo-kafka
make kafka-down
排查前先看计划 make itest-plan 会把将要执行的每一条命令按顺序打印出来,不启动任何东西。 大多数「它到底在干什么」的疑问都能在这里得到答案。
Start with the plan make itest-plan prints every command a run would execute, in order, without starting anything. Most "what is it actually doing" questions are answered there.