Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
22 changes: 22 additions & 0 deletions lading/src/bin/lading.rs
Original file line number Diff line number Diff line change
Expand Up @@ -662,6 +662,12 @@
},
() = &mut timer_watcher_wait => {
info!("shutdown signal received.");
// First of four graceful-shutdown breadcrumbs. If a regression
// stalls the drain, the later three stop being reached and triage
// sees where progress died.
lading_antithesis::reachable!(
"lading began graceful shutdown when its experiment timer elapsed"
);
break Ok(());
}
Some(res) = osrv_joinset.join_next() => {
Expand Down Expand Up @@ -708,11 +714,22 @@
};

shutdown_broadcast.signal_and_wait().await;
// Second breadcrumb. The barrier waits only on registered peers, today just

Check warning on line 717 in lading/src/bin/lading.rs

View workflow job for this annotation

GitHub Actions / Rust Actions (Check/Fmt/Clippy) (ubuntu-latest, fmt)

Diff in /home/runner/work/lading/lading/lading/src/bin/lading.rs

Check warning on line 717 in lading/src/bin/lading.rs

View workflow job for this annotation

GitHub Actions / Rust Actions (Check/Fmt/Clippy) (ubuntu-latest, fmt)

Diff in /home/runner/work/lading/lading/lading/src/bin/lading.rs

Check warning on line 717 in lading/src/bin/lading.rs

View workflow job for this annotation

GitHub Actions / Rust Actions (Check/Fmt/Clippy) (macos-latest, fmt)

Diff in /Users/runner/work/lading/lading/lading/src/bin/lading.rs

Check warning on line 717 in lading/src/bin/lading.rs

View workflow job for this annotation

GitHub Actions / Rust Actions (Check/Fmt/Clippy) (macos-latest, fmt)

Diff in /Users/runner/work/lading/lading/lading/src/bin/lading.rs
// the capture manager. If the load-bearing watchers are ever registered, a
// generator stuck in its connect loop would hang here and this stops firing.
lading_antithesis::reachable!(
"lading's shutdown drain barrier returned without hanging"
);

// Await the capture manager task, if it exists, to ensure all data is
// flushed.
if let Some(handle) = capture_manager_handle {
let _ = handle.await;
// Third breadcrumb. The capture manager task joined, so every matured
// capture record reached disk.
lading_antithesis::reachable!(
"lading flushed all matured capture data to disk during graceful shutdown"
);
}

res
Expand Down Expand Up @@ -813,6 +830,11 @@
);
runtime.shutdown_timeout(max_shutdown_delay);
info!("Bye. :)");
// Fourth breadcrumb. The runtime drained within the shutdown-delay backstop
// and lading is about to return its exit status.
lading_antithesis::reachable!(
"lading drained its runtime within the shutdown-delay backstop and exited cleanly"
Comment on lines +833 to +836

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Do not report forced runtime teardown as a clean drain

Runtime::shutdown_timeout returns both when all tasks finish and when its deadline expires and the remaining tasks are forcibly dropped, but this breadcrumb fires unconditionally afterward and claims the runtime drained cleanly. Any configuration without the scenario's shorter external watchdog can therefore satisfy this Antithesis reachability assertion after waiting out the backstop with hung tasks, obscuring the exact shutdown regression these breadcrumbs are intended to localize; rename it to describe only backstop completion or separately establish that all runtime tasks joined.

Useful? React with 👍 / 👎.

);
res
}

Expand Down
7 changes: 7 additions & 0 deletions lading/src/generator/tcp.rs
Original file line number Diff line number Diff line change
Expand Up @@ -278,6 +278,13 @@ impl TcpWorker {
}
() = &mut shutdown_wait => {
info!("shutdown signal received");
// A connected worker honors shutdown here via select. The
// connect-loop branch above does not select on shutdown, so a
// worker stuck reconnecting to an unreachable destination
// never reaches this point. The contrast localizes that gap.
Comment on lines +281 to +284

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Make reconnecting TCP workers observe shutdown

In the shutdown-safety scenario's fixed 127.0.0.1:1 configuration, the worker remains in the connect/retry branch at lines 233-250 and never polls this shutdown arm; Tcp::spin then waits indefinitely for that worker, so runtime.shutdown_timeout(60s) outlasts the new 35-second watchdog and every ordinary run records rc == 124 rather than validating a clean exit. GNU timeout --help confirms that it sends TERM after the duration and returns 124 when the command times out. The connect attempt and retry sleep need to select on the shutdown watcher before this scenario can assert rc == 0.

Useful? React with 👍 / 👎.

lading_antithesis::reachable!(
"a connected tcp generator worker honored the shutdown signal and stopped promptly"
);
return Ok(());
},
}
Expand Down
18 changes: 13 additions & 5 deletions test/antithesis/CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -38,9 +38,17 @@ ready; the workload container runs it from its entrypoint.
It receives lading's load and owns the "load arrived" assertion. SDK-linked but
built without coverage instrumentation. Built, linted, and tested from the repo
root like any other crate.
- `scenarios/general/` — the MVP scenario: `Dockerfile` (three targets: the
instrumented lading system under test, the sink, the workload), a
`docker-compose.yaml` snouty consumes as `--config`, `launch.env`, the
hard-wired `lading.yaml`, and the `workload/` build inputs. Config variation is
deferred to the workload seam.
- `harness/` — the shared harness crate. Holds the per-timeline config sampler

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Use ASCII punctuation in the new directory bullets

The three newly added bullets for harness/, scenarios/general/, and scenarios/shutdown-safety/ use Unicode em dashes, although repository documentation is required to contain US-ASCII only; replace these separators with ASCII punctuation.

AGENTS.md reference: AGENTS.md:L151-L158

Useful? React with 👍 / 👎.

and the workload test commands baked into scenarios: `first_sample_config`,
`anytime_capture_consistent`, and `anytime_lading_drained_bounded`.
- `scenarios/general/` — the MVP scenario. lading pushes TCP load at the sink,
the workload samples a lading config per timeline and validates capture
crash-consistency across node faults. See `scenarios/general/README.md`.
- `scenarios/shutdown-safety/` — graceful-shutdown scenario. lading runs under a
`timeout` watchdog with its generator aimed at an unreachable destination, and
the workload asserts lading drains cleanly within a bound. No sink, no node
faults. See `scenarios/shutdown-safety/README.md`.
- Each scenario directory carries a `Dockerfile`, a `docker-compose.yaml` snouty
consumes as `--config`, `launch.env`, a `lading.yaml` when the config is fixed,
the `workload/` build inputs, and a `README.md`.
- `bin/launch.sh` — the generic launcher shared by every scenario.
4 changes: 4 additions & 0 deletions test/antithesis/harness/Cargo.toml
Original file line number Diff line number Diff line change
Expand Up @@ -17,6 +17,10 @@ path = "src/bin/first_sample_config.rs"
name = "anytime_capture_consistent"
path = "src/bin/anytime_capture_consistent.rs"

[[bin]]
name = "anytime_lading_drained_bounded"
path = "src/bin/anytime_lading_drained_bounded.rs"

[lints]
workspace = true

Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,66 @@
//! Antithesis `anytime_` command: assert lading drains promptly on graceful
//! (experiment-timer) shutdown even when its generator points at an unreachable
//! destination.
//!
//! The lading (system under test) container runs lading under a wall-clock
//! `timeout` watchdog and writes lading's exit code to a sentinel file on the
//! shared volume. This checker runs in the (never-faulted) workload container
//! and fires whenever Antithesis chooses. It reads that exit code:
//! * rc == 0 -> lading drained within the watchdog bound (the good outcome).
//! * rc == 124 -> the `timeout` watchdog killed lading before it drained. That
//! would indicate the shutdown drain hung rather than completing promptly,
//! for example if a generator stuck against the unreachable destination
//! blocked the drain instead of being abandoned at runtime teardown.
//! * any other rc -> some other failure, e.g. a lading crash.
//!
//! It always exits 0; findings are reported as named assertions.

use std::fs;

/// File the lading entrypoint writes lading's exit code into. Overridable via
/// `LADING_EXIT_PATH`.
const DEFAULT_EXIT_PATH: &str = "/shared/lading_exit";

fn main() {

Check warning on line 24 in test/antithesis/harness/src/bin/anytime_lading_drained_bounded.rs

View workflow job for this annotation

GitHub Actions / Rust Actions (Check/Fmt/Clippy) (ubuntu-latest, fmt)

Diff in /home/runner/work/lading/lading/test/antithesis/harness/src/bin/anytime_lading_drained_bounded.rs

Check warning on line 24 in test/antithesis/harness/src/bin/anytime_lading_drained_bounded.rs

View workflow job for this annotation

GitHub Actions / Rust Actions (Check/Fmt/Clippy) (ubuntu-latest, fmt)

Diff in /home/runner/work/lading/lading/test/antithesis/harness/src/bin/anytime_lading_drained_bounded.rs

Check warning on line 24 in test/antithesis/harness/src/bin/anytime_lading_drained_bounded.rs

View workflow job for this annotation

GitHub Actions / Rust Actions (Check/Fmt/Clippy) (macos-latest, fmt)

Diff in /Users/runner/work/lading/lading/test/antithesis/harness/src/bin/anytime_lading_drained_bounded.rs

Check warning on line 24 in test/antithesis/harness/src/bin/anytime_lading_drained_bounded.rs

View workflow job for this annotation

GitHub Actions / Rust Actions (Check/Fmt/Clippy) (macos-latest, fmt)

Diff in /Users/runner/work/lading/lading/test/antithesis/harness/src/bin/anytime_lading_drained_bounded.rs
lading_antithesis::init();

let path =
std::env::var("LADING_EXIT_PATH").unwrap_or_else(|_| DEFAULT_EXIT_PATH.to_string());

let Ok(content) = fs::read_to_string(&path) else {
// The exit-code sentinel is not present yet this tick: lading has not
// finished (or been killed) yet. Nothing to check.
return;
};

let trimmed = content.trim();
let Ok(rc) = trimmed.parse::<i32>() else {
// A non-integer sentinel means the entrypoint has not written a
// well-formed exit code yet; wait for a later tick.
return;
};

// Non-vacuity floor: prove the checker actually observed a recorded exit at
// least once across the run, so the always-invariant below is not passing
// vacuously on an absent file.
lading_antithesis::sometimes!(
!trimmed.is_empty(),
"lading recorded a bounded exit",
{ "rc": rc }
);

// The invariant under test: on graceful (experiment-timer) shutdown with an
// unreachable destination, lading drains within the watchdog bound
// (rc == 0). A watchdog kill (rc == 124) would mean the shutdown drain hung
// rather than completing promptly.
lading_antithesis::always!(
rc == 0,
"lading drains within bound on graceful shutdown with an unreachable destination",
{ "rc": rc }
);

lading_antithesis::reachable!(
"bounded-drain checker validated a recorded lading exit",
{ "rc": rc }
);
}
56 changes: 56 additions & 0 deletions test/antithesis/scenarios/general/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,56 @@
# General scenario

This is the MVP scenario. lading generates TCP load into an instrumented sink,
and the run establishes two things across faults: that load actually arrives, and
that lading's capture files stay crash-consistent when lading is hard-killed.

## How it works

This scenario comprises the following components:

* sink (oracle)
* lading, instrumented (system under test)
* workload (driver)

The sink is a TCP server that counts received bytes and owns the "load arrived"
assertion. It is healthchecked and never faulted. lading is the system under
test, faulted at the node level. The workload writes the timeline's config and
validates lading's capture.

The workload's `first_sample_config` command samples a lading config per timeline
from Antithesis randomness (payload variant, bytes per second, parallel
connections, and seed) and writes it, plus a `ready` sentinel, to the shared
volume. lading's entrypoint blocks on the sentinel, then boots under the sampled
config and pushes TCP load at the sink. lading writes its capture to a named
`capture` volume that outlives a `node_termination` of the lading container.

Antithesis injects node faults on lading: a `node_termination` is a SIGKILL of
the whole container, so only artifacts on the named volume survive. The
workload's `anytime_capture_consistent` checker validates every capture file by
its format:

* JSONL: a valid parseable prefix. A torn final line, the interrupted write, is
tolerated. A broken earlier line or a violated invariant is not.
* Parquet: footer-terminated, so a hard kill leaves it unreadable. That is
expected, not corruption. A readable Parquet must be internally consistent.
* Multi: both of the above, on the `.jsonl` and `.parquet` files.

The checker asserts no mid-file corruption and monotonic `fetch_index` and
per-series time. The sink asserts `sometimes(total_bytes > 0)`.

# Caveats

1. JSONL survives a kill as a valid prefix because of the 60-second in-memory
maturity window. Matured intervals are on disk, the unmatured tail is lost on
abrupt death by design. This scenario checks consistency, not completeness.
2. Parquet and Multi are footer-terminated. An abrupt kill leaves the Parquet
file unreadable, which the checker treats as expected rather than a violation.
3. There is deliberately no blackhole. The sink is the only receiver, so a config
that expected a blackhole would not apply here.

# Assumptions

1. The `capture` named volume outlives a `node_termination` of lading.
2. The sink is never faulted and is a healthchecked dependency of lading.
3. Config sampling draws from Antithesis randomness so timelines branch across
the payload-variant menu.
115 changes: 115 additions & 0 deletions test/antithesis/scenarios/shutdown-safety/Dockerfile
Original file line number Diff line number Diff line change
@@ -0,0 +1,115 @@
# syntax=docker/dockerfile:1
#
# Antithesis harness image for the lading "shutdown-safety" scenario.
#
# Build context is the repository root. Named runtime targets:
# - lading : the instrumented lading binary (the system under test), sancov on
# - workload : bakes the bounded-drain checker, emits setup_complete, then idles
#
# There is deliberately NO sink: the point of this scenario is a generator aimed
# at an UNREACHABLE address (connection refused -> busy reconnect), exercising
# whether lading drains promptly on graceful (experiment-timer) shutdown.
#
# lading is built native x86_64-unknown-linux-gnu (glibc).

# ---------------------------------------------------------------------------
# Shared build environment: Rust toolchain (pinned by rust-toolchain.toml) plus
# native build dependencies (protoc for prost/tonic, openssl for reqwest).
# ---------------------------------------------------------------------------
FROM debian:bookworm AS build-base
ENV DEBIAN_FRONTEND=noninteractive \
NO_COLOR=1 \
CARGO_TERM_COLOR=never
RUN apt-get update && \
apt-get install --no-install-recommends -y \
build-essential ca-certificates curl pkg-config libssl-dev cmake protobuf-compiler && \
rm -rf /var/lib/apt/lists/*
# Install rustup with no default toolchain; the pinned toolchain auto-installs on
# first cargo invocation inside the repo (rust-toolchain.toml).
RUN curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | \
sh -s -- -y --profile minimal --default-toolchain none
ENV PATH="/root/.cargo/bin:${PATH}"

# ---------------------------------------------------------------------------
# Build the instrumented lading (the system under test).
#
# Coverage instrumentation uses the modern Antithesis Rust flow: the
# `antithesis-instrumentation` crate (referenced once in the binary behind the
# `antithesis` feature) provides the runtime shim, and these RUSTFLAGS enable
# LLVM sancov coverage. `--build-id` is required for symbolization. LTO is forced
# off for predictable sancov instrumentation, and split-debuginfo is forced off
# so the binary keeps embedded DWARF. The release profile already sets
# `panic = "abort"`, so any lading panic becomes SIGABRT, caught as a hard crash.
#
# The sancov RUSTFLAGS are passed via `--config target.<triple>.rustflags` with an
# explicit `--target`, NOT via the RUSTFLAGS env var. With an explicit target,
# Cargo builds host artifacts (build scripts, proc-macros) for the host and does
# NOT apply the target rustflags to them, so they are not instrumented and link
# cleanly.
# ---------------------------------------------------------------------------
FROM build-base AS lading-builder
ENV CARGO_PROFILE_RELEASE_LTO=off \
CARGO_PROFILE_RELEASE_SPLIT_DEBUGINFO=off
WORKDIR /lading
COPY . /lading
RUN --mount=type=cache,target=/lading/target,id=antithesis-lading-target \
--mount=type=cache,target=/root/.cargo/registry,id=cargo-registry \
--mount=type=cache,target=/root/.cargo/git,id=cargo-git \
cargo build --release --package lading --bin lading --features antithesis \
--target x86_64-unknown-linux-gnu \
--config 'target.x86_64-unknown-linux-gnu.rustflags=["-Ccodegen-units=1","-Cpasses=sancov-module","-Cllvm-args=-sanitizer-coverage-level=3","-Cllvm-args=-sanitizer-coverage-trace-pc-guard","-Clink-args=-Wl,--build-id"]' && \
cp /lading/target/x86_64-unknown-linux-gnu/release/lading /usr/local/bin/lading && \
echo "Validating Antithesis instrumentation symbols..." && \
nm /usr/local/bin/lading | grep -q "antithesis_load_libvoidstar" && \
nm /usr/local/bin/lading | grep -q "sanitizer_cov_trace_pc_guard" && \
echo "Instrumentation symbols present."

# ---------------------------------------------------------------------------
# Build the bounded-drain checker, uninstrumented. It is a supporting harness
# component (a workload test command), not the system under test, so it needs no
# coverage instrumentation. There is no sink in this scenario.
# ---------------------------------------------------------------------------
FROM build-base AS tools-builder
WORKDIR /tools
COPY . /tools
RUN --mount=type=cache,target=/tools/target,id=antithesis-tools-target \
--mount=type=cache,target=/root/.cargo/registry,id=cargo-registry \
--mount=type=cache,target=/root/.cargo/git,id=cargo-git \
cargo build --release --package harness --bin anytime_lading_drained_bounded && \
cp /tools/target/release/anytime_lading_drained_bounded /usr/local/bin/anytime_lading_drained_bounded

# ---------------------------------------------------------------------------
# Runtime: instrumented lading (system under test).
# ---------------------------------------------------------------------------
FROM debian:bookworm-slim AS lading
ENV NO_COLOR=1
RUN apt-get update && \
apt-get install --no-install-recommends -y ca-certificates libssl3 && \
rm -rf /var/lib/apt/lists/*
COPY --from=lading-builder /usr/local/bin/lading /usr/local/bin/lading
COPY --chmod=755 test/antithesis/scenarios/shutdown-safety/lading-entrypoint.sh /usr/local/bin/lading-entrypoint.sh
# Fixed config, baked in: this scenario does NOT sample configs, so the generator
# is hard-wired to an unreachable address.
COPY test/antithesis/scenarios/shutdown-safety/lading.yaml /etc/lading/lading.yaml
# Antithesis scans /symbols in every pushed image and matches by build-id.
RUN mkdir -p /symbols && ln -s /usr/local/bin/lading /symbols/lading
ENTRYPOINT ["/usr/local/bin/lading-entrypoint.sh"]

# ---------------------------------------------------------------------------
# Runtime: workload driver. Bakes the bounded-drain checker, emits
# setup_complete, then idles.
#
# The checker links only lading_antithesis + std (no lading, no openssl), so the
# slim base already carries everything it needs (libgcc_s, libc). No extra
# packages are installed.
# ---------------------------------------------------------------------------
FROM debian:bookworm-slim AS workload
ENV NO_COLOR=1
COPY --chmod=755 test/antithesis/scenarios/shutdown-safety/workload/setup-complete.sh /opt/antithesis/setup-complete.sh
COPY --chmod=755 test/antithesis/scenarios/shutdown-safety/workload/entrypoint.sh /entrypoint.sh
# Antithesis test commands live here; the platform discovers and runs them from
# /opt/antithesis/test/v1/<template>/<command>. Checked-in files provide the
# template structure; the compiled command binary is injected below.
COPY --chmod=755 test/antithesis/scenarios/shutdown-safety/workload/test/ /opt/antithesis/test/
COPY --from=tools-builder --chmod=755 /usr/local/bin/anytime_lading_drained_bounded /opt/antithesis/test/v1/main/anytime_lading_drained_bounded
ENTRYPOINT ["/entrypoint.sh"]
58 changes: 58 additions & 0 deletions test/antithesis/scenarios/shutdown-safety/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,58 @@
# Shutdown-safety scenario

This scenario tests that lading performs a clean, prompt graceful shutdown when
its own experiment timer elapses, even while a generator is stuck unable to reach
its destination. Shutdown correctness matters because lading owns the clock in
Comment on lines +1 to +5

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Record the shutdown-safety design as an ADR

This new scenario introduces a distinct two-container subsystem and non-obvious architectural choices around an external watchdog, intentionally unreachable transport, fault exclusions, and exit-code oracles, but records those decisions only in a standalone scenario README. Move the durable rationale into a new ADR (including alternatives and consequences) and update the ADR index, leaving this README as operational usage documentation.

AGENTS.md reference: AGENTS.md:L31-L49

Useful? React with 👍 / 👎.

real deployments and must drain and exit on schedule.

## How it works

This scenario comprises the following components:

* lading, instrumented (system under test)
* workload (driver)

There is deliberately no sink. lading drives itself, so the workload only emits
`setup_complete`, bakes the checker, and idles.

lading owns the clock. The experiment timer fires at 15 seconds and signals
graceful shutdown, after which lading must drain its subsystems and exit cleanly
with `rc == 0`. To stress that drain, the tcp generator is aimed at
`127.0.0.1:1`, where nothing listens, so its worker sits in a busy reconnect loop
and never makes forward progress. This is the adversarial condition: a generator
that will not make progress on its own and could stall an unsafe drain.

`lading-entrypoint.sh` runs lading under a wall-clock `timeout` watchdog so a
prompt drain is distinguishable from a hang. The timer fires at 15 seconds,
`--max-shutdown-delay` is set to 60 seconds, and `timeout` kills lading at 35
seconds. A shutdown-responsive lading exits at ~15 seconds with `rc == 0`. A
lading that hangs the drain can only be released by its 60-second backstop, which
is longer than the 35-second watchdog, so the watchdog kills it first and yields
`rc == 124`. lading records its exit code to the shared volume, and the
workload's `anytime_lading_drained_bounded` checker reads it and asserts
`always(rc == 0)`. Any other code indicates a crash.

lading is also instrumented with SUT-side breadcrumbs along the graceful path
(began shutdown on timer, drain barrier returned, capture flushed, clean exit).
These trace how far shutdown progressed, so a regression that stalls the drain is
localized to the point where the breadcrumbs stop being reached.

# Caveats

1. `rc == 0` today confirms the outcome (prompt clean exit), not that every
worker was individually drained. The generator's shutdown watcher is not a
registered peer, so the drain barrier does not wait on it and the stuck worker
is abandoned at runtime teardown. If the load-bearing watchers are ever made
registered peers, a stuck generator would hang the drain and this scenario
would catch it as `rc == 124`.
2. There are no node faults here. A node kill would preempt the graceful timer
and hide the drain behavior. Global cpu and clock faults still apply.
3. The config is fixed and baked in. This scenario does not sample per-timeline
configs.

# Assumptions

1. lading owns the clock and self-terminates on its experiment timer.
2. `--max-shutdown-delay` exceeds the watchdog margin, so reliance on the
force-drop backstop surfaces as a watchdog kill rather than a clean exit.
3. The workload container is never faulted, so it survives to read the exit code.
Loading
Loading