集成测试Integration test

用容器启动 Kafka、创建全新主题,再驱动真实的命令行程序走完「发送消息 → 消费消息」 的完整链路。

Starts Kafka in a container, creates fresh topics, then drives the real command-line program through a full produce → consume round trip.

为什么不在 moon test 里Why it is not part of moon test

它需要容器运行时和一个真实 broker,因此不属于 moon test, 只在被显式调用时运行,入口是 make itest。

It needs a container runtime and a live broker, so it is deliberately not part of moon test and only runs when asked for — through make itest.

不过 harness 里那些纯函数(命令渲染、输出解析、运行时候选列表等)有自己的白盒测试, 它们是 moon test 的一部分。所以 make ci 覆盖的是 解析逻辑,端到端链路则由 make itest 覆盖,两者互补。

The harness's pure helpers — command rendering, output parsing, the runtime candidate list — do have white-box tests, and those are part of moon test. So make ci covers the parsing logic while make itest covers the end-to-end path; the two are complementary.

运行Running it

make itest

make itest 先构建 CLI,再启动测试。想先看它到底会执行哪些命令,用 make itest-plan(等价于 --dry-run)——它不启动任何东西, 也不会碰你的容器。

make itest builds the CLI and then runs the test. To see exactly which commands it would execute, use make itest-plan (the same as --dry-run): it starts nothing and touches no container.

make itest-plan

也可以绕开 Makefile 直接运行,所有选项照常可用:

You can also bypass the Makefile; every option still applies:

moon run --target native tests/itest -- --target native --runtime docker
--dry-run 与真实运行同源 计划里的每一条命令都由真实运行所使用的同一批构造函数生成,因此两者不会各说各话, 计划不会随时间腐烂。
The plan and the run share one source Every command in the plan is produced by the same builders the real run uses, so the two cannot disagree — the plan will not rot as the harness changes.

测试流程What a run does

按顺序执行下面这些步骤:

In order:

  1. 先构建 CLI。harness 执行的是构建产物,所以先编译一次——编译错误在这里暴露,后面的计时阶段也不必为构建买单。
  2. 挑一个能用的容器运行时。
  3. 启动 Kafka,然后轮询 broker 直到它接受管理命令。
  4. 创建两个全新主题,并等待每个分区都选出 leader。
  5. 阶段一:发送一条消息,再从 earliest 读回。
  6. 阶段二:让 consumer 先开始轮询,之后再发送消息。
  7. 收尾:默认停止 Kafka(--keep 可保留)。
  8. 打印 checks=… failures=…;有失败则以非零退出码结束。
  1. Build the CLI first. The harness executes the built artifact, so it compiles once up front: a compile error surfaces here, and no timed phase then pays for a build.
  2. Pick a working container runtime.
  3. Start Kafka, then poll the broker until it answers admin commands.
  4. Create two fresh topics and wait for every partition to elect a leader.
  5. Phase 1: produce a message, then read it back from earliest.
  6. Phase 2: let a consumer start polling, then produce to it.
  7. Tear down: Kafka is stopped unless --keep is given.
  8. Print checks=… failures=…; a non-zero exit code follows any failure.

每次运行都使用唯一的主题名与唯一的消息内容(以 @async.now() 打时间戳), 因此上一次运行的残留数据不可能满足这一次的断言。

Every run uses unique topic names and unique payloads, stamped with @async.now(), so nothing a previous run left behind can satisfy an assertion.

两个阶段The two phases

阶段一:先发送,再从头读回Phase 1: produce, then read back from the start

这一阶段不依赖任何时序:消息先落盘,之后才去读。断言包括:

Nothing here depends on timing: the message lands first and is read afterwards. The assertions:

断言说明
生产者退出码为 0produce 命令正常结束
生产者报告了偏移量能从输出里解析出 at offset N
首条记录位于偏移量 0仅在 --partitions 1(默认)时检查:全新单分区主题的第一条记录必然是 0
消费者退出码为 0--max-messages 1 之后正常退出
读回了同一条记录key 与 value 都与发送时一致
偏移量一致消费到的偏移量等于写入时的偏移量
AssertionMeaning
The producer exits 0The produce command finished cleanly
The producer reports an offsetAn at offset N value was parsed from its output
The first record sits at offset 0Checked only with --partitions 1 (the default): the first record of a fresh single-partition topic is necessarily 0
The consumer exits 0It exited cleanly after --max-messages 1
The record was read backBoth key and value match what was sent
The offsets agreeThe consumed offset equals the produced one

阶段二:消费的同时发送Phase 2: consuming while it is produced to

第二个主题在开始时是空的。consumer 先启动并开始轮询,--live-delay-ms (默认 2000 毫秒)之后再发送消息。断言是:生产者退出码 0、消费者退出码 0,并且 消费者确实收到了这条记录。

The second topic is empty when the run starts. The consumer is started first and begins polling; the message is produced after --live-delay-ms (2000 ms by default). The assertions: the producer exits 0, the consumer exits 0, and the consumer did receive that record.

因为主题是空的,赢得竞态的 consumer 会在下一次 poll 读到它,输掉竞态的会在第一次 fetch 读到它——两种顺序都满足断言。--live-delay-ms 只是让常见路径成为 真正的「实时投递」,而不是一场竞态。

Since the topic is empty, a consumer that wins the race reads the record on its next poll and one that loses it reads it on its first fetch — both orderings satisfy the assertions. --live-delay-ms merely makes the common case a genuinely live delivery rather than a race.

容器运行时怎么选How the container runtime is chosen

默认(--runtime auto)按 podman、docker 的顺序逐个探测,用 <engine> version 确认它真的能用, 而不只是装了。这样就不会选中一个装了但没跑起来的引擎,然后在后面报一个莫名其妙的错。

By default (--runtime auto) it probes podman then docker, using <engine> version to confirm the engine actually works rather than merely being installed — so it never picks an installed-but-dead engine and fails later with a confusing error.

锁定引擎后,再看它有没有 compose provider:

Once an engine is locked in, it looks for a compose provider:

  1. <engine> compose version(插件形式,如 docker compose);
  2. <engine>-compose version(独立二进制,如 docker-compose)。
  1. <engine> compose version — the plugin form, e.g. docker compose;
  2. <engine>-compose version — the standalone binary, e.g. docker-compose.

--compose-bin 可以指定一个二进制并跳过这两个探测。两条路径分别是:

--compose-bin names a binary explicitly and skips both probes. The two paths are:

探测结果行为
有 compose provider用 docker-compose.kafka.yml 启动:<provider> -f docker-compose.kafka.yml up -d
没有 provider(常见的纯 podman 安装,podman compose 需要 podman-compose)由 harness 直接执行 podman run 起单节点 broker,参数与 compose 文件里的 kafka 服务一致
Probe resultBehaviour
A compose provider answersStart with docker-compose.kafka.yml: <provider> -f docker-compose.kafka.yml up -d
No provider (a common podman setup, where podman compose needs podman-compose)The harness drives podman run itself, with settings matching the kafka service in the compose file

直接 run 那条路径下,镜像标签从 compose 文件里读出来(也可以用 --image 覆盖),因此两条路径不会各写一份镜像版本。启动前会先 inspect 同名容器:已经在跑就直接复用,否则先 rm -f 清掉残留 再起。

On the direct run path the image tag is read out of the compose file (--image overrides it), so the two paths cannot drift on the image version. Before starting, the container is inspected: if it is already running it is reused, otherwise a stale one is rm -fed first.

「<engine> is not usable」不是错误 探测会逐个候选尝试,每一个失败的候选都会打印一行 ... podman is not usable,这只是中间结果。只有所有候选都没通过, 才会打印 FAIL no working container runtime。 make itest 默认固定 ITEST_RUNTIME=docker,直接跳过探测, 所以那行日志根本不会出现。
"<engine> is not usable" is not an error The probe walks its candidates in order and logs a ... podman is not usable line for each miss — an intermediate result. Only when every candidate fails do you get FAIL no working container runtime. make itest pins ITEST_RUNTIME=docker and skips the probe entirely, so that line never appears there.

选项Options

开关Flags

开关作用
--no-container使用已在运行的 Kafka:跳过容器启停与建主题
--keep结束时保留 Kafka 容器,便于排查
--dry-run打印执行计划后直接退出,什么都不运行
FlagEffect
--no-containerUse an already running Kafka: skip container lifecycle and topic creation
--keepLeave the Kafka container running when the test finishes
--dry-runPrint the plan and exit without running anything

取值选项Value options

选项默认值说明
--runtimeauto容器运行时:auto、podman、docker 或一个路径
--compose-bin空compose 可执行文件;默认是 <runtime> compose
--compose-filedocker-compose.kafka.yml用来启动 Kafka 的 compose 文件
--run-modeauto如何管理 Kafka:auto、compose 或 run
--image空直接 run 路径使用的镜像;默认从 compose 文件读取
--containermoonkafka-demo-kafkaKafka 容器名
--host127.0.0.1CLI 连接的 broker 地址
--port9092CLI 连接的 broker 端口
--targetnative执行哪个目标的构建产物
--cli空要驱动的 CLI 二进制;默认是 native 调试构建
--repo.仓库根目录(moon.mod 所在处)
--topic空主题名;默认每次运行生成唯一名字
--partitions1harness 创建的主题的分区数
--timeout-secs120等待 Kafka 就绪的预算
--live-delay-ms2000阶段二里发送消息前的延迟
OptionDefaultMeaning
--runtimeautoContainer runtime: auto, podman, docker or a path
--compose-binemptyCompose binary; the default is <runtime> compose
--compose-filedocker-compose.kafka.ymlCompose file used to start Kafka
--run-modeautoHow Kafka is managed: auto, compose or run
--imageemptyImage for the direct run path; read from the compose file by default
--containermoonkafka-demo-kafkaName of the Kafka container
--host127.0.0.1Broker host the CLI connects to
--port9092Broker port the CLI connects to
--targetnativeBuild target whose artifacts are executed
--cliemptyCLI binary to drive; the native debug build by default
--repo.Repository root (where moon.mod lives)
--topicemptyTopic name; a unique name per run by default
--partitions1Partitions for the topics this harness creates
--timeout-secs120Budget for waiting until Kafka is ready
--live-delay-ms2000Delay before producing in the live-delivery phase

默认要驱动的 CLI 二进制是 _build/native/debug/build/cmd/main/main.exe(随 --target 变化)。文件不存在时会明确报告「CLI 二进制缺失」,而不是抛出一个莫名的启动失败。

The default CLI binary is _build/native/debug/build/cmd/main/main.exe (following --target). When it is missing you get an explicit "the CLI binary is missing" failure rather than a mysterious spawn error.

常用组合Common invocations

make itest ITEST_FLAGS="--keep"       # 保留容器,便于排查
make itest ITEST_FLAGS="--dry-run"    # 只打印将要执行的命令
make itest-plan                       # 同上
make itest ITEST_RUNTIME=podman       # 改用 podman(默认 docker)

跑在已有的 Kafka 上Against a Kafka you already run

--no-container 跳过容器生命周期与建主题,因此必须自己指定一个已存在且为空 的主题:测试不会替你创建它,也不会替你清空它。这种情况下阶段二会被跳过——它需要一个空主题, 而 harness 不负责准备。

--no-container skips the container lifecycle and topic creation, so you must name an existing, empty topic: the test will not create it and will not clear it. Phase 2 is skipped in this mode — it needs an empty topic, and the harness is not the one that prepares it.

moon run --target native tests/itest -- --no-container --topic <an-existing-empty-topic>

日志怎么读Reading the log

日志把每条命令、退出码与耗时都逐条打印出来:

Every command, exit code and duration is logged as it happens:

行含义
# <说明>这条命令是干什么的
$ <命令>将要执行的命令;失败的 CI 日志里能看到确切的命令行
... exit=0 in 412ms命令成功及其耗时
... exit=1 after 88ms命令失败,随后缩进打印 stderr/stdout
-> ...harness 的中间判断,如「容器运行时:docker」「broker 已就绪」
PASS / FAIL <断言>一条断言的结论
checks=N failures=M结束时的汇总;M > 0 即非零退出
LineMeaning
# <purpose>What the command is for
$ <command>The command about to run; a failing CI log shows the exact command line
... exit=0 in 412msThe command succeeded, and how long it took
... exit=1 after 88msThe command failed; its stderr/stdout follow, indented
-> ...An intermediate decision, e.g. "container runtime: docker" or "broker is ready"
PASS / FAIL <assertion>The verdict for one assertion
checks=N failures=MThe final tally; M > 0 means a non-zero exit

第一次 produce 被拒是正常的The first produce being refused is normal

刚起来的 broker 会先接受管理命令、也接受建主题,但在能处理 produce 之前还有一小段窗口。 落在这个窗口里的第一次 produce 会被立即拒绝,报 moonkafka 的 ProtocolError,第二次就成功了。

A broker that has just started answers admin commands and accepts topic creation well before it can serve a produce. A produce that lands in that window fails fast with moonkafka's ProtocolError; the next one succeeds.

harness 因此对 produce 做有界重试(最多 5 次,每次间隔 1 秒),并把每一次尝试都如实 打印出来。这样启动竞态不会污染断言,同时日志里仍然看得到它发生过。

The harness therefore retries the produce a bounded number of times (5 attempts, one second apart) and logs every attempt. The startup race stays out of the assertions while remaining visible in the log.

  # moonkafka cli
  $ _build/native/debug/build/cmd/main/main.exe produce \
      itest-1757923200 itest-value-a-1757923200 itest-key-1757923200 127.0.0.1 9092
  ... exit=1 after 61ms
  ! ProtocolError: ...
  ... retrying (2/5): the previous attempt was refused
  # moonkafka cli
  $ _build/native/debug/build/cmd/main/main.exe produce \
      itest-1757923200 itest-value-a-1757923200 itest-key-1757923200 127.0.0.1 9092
  ... exit=0 in 74ms

与此同时,harness 在建完主题之后会轮询 kafka-topics.sh --describe, 直到没有分区停留在 Leader: -1,用这个把大部分竞态直接消掉,而不是靠 一段固定 sleep。

In parallel, after creating a topic the harness polls kafka-topics.sh --describe until no partition still reports Leader: -1 — removing most of that race without an arbitrary sleep.

就绪等待Readiness waits

等待做法上限
broker 就绪每 2 秒执行一次 kafka-topics.sh --list,与 compose 的健康检查同一条命令--timeout-secs,默认 120 秒
分区选出 leader每 1 秒执行一次 --describe,直到输出里不再有 Leader: -160 秒
每一条外部命令子进程带硬超时,超时即杀掉,不会把 harness 挂住按步骤 60 秒–900 秒不等
WaitHowBound
Broker readinesskafka-topics.sh --list every 2 seconds — the very command the compose healthcheck uses--timeout-secs, 120 s by default
Partition leaders--describe every second until Leader: -1 is gone from the output60 s
Every external commandChildren run under a hard timeout and are killed on expiry, so the harness cannot hang60 s – 900 s depending on the step
容器残留怎么办 用 --keep 留下的容器不会影响下一次运行:直接 run 的路径会先 inspect 再复用或清理。要手动开关 Kafka,用 make kafka-up / make kafka-down(引擎由 RUNTIME 选择)。make itest 管理自己的容器,不使用这两个目标。
Leftover containers A container left behind by --keep does not disturb the next run: the direct run path inspects it and reuses or clears it. For manual control use make kafka-up / make kafka-down (the engine comes from RUNTIME); make itest manages its own container and does not use those targets.