multicam-sim is the producer: it builds a typed multi-camera scene and emits a JSON manifest that a triangulation consumer (multicam-occlusion) reads. It does pure analytic projection + boolean occlusion — no renderer, no GL.
This document is the contract of record: the camera convention and the manifest schema. Downstream layers (DSL sugar, renderer, pose, issue injection) build on top of this without changing it.
Mirrored from multicam-occlusion@59f4906
(src/multicam_occlusion/triangulation.py::look_at_rotation / build_ring_cameras).
Replicated exactly so a manifest produced here is consumed convention-for-convention
by that package's triangulate_dlt.
OpenCV pinhole, right-down-forward (RDF) camera axes, world Z-up:
forward = (target - eye) / ||target - eye|| # camera +z, viewing direction
right = forward x up_world, normalised # camera +x
down = forward x right # camera +y
R rows = [right, down, forward] # world -> camera rotation
t = -R @ C # world -> camera translation (NOT the centre C)
C = -R^T @ t # camera centre in world (inverse)
up_world = (0, 0, 1) # Z-up
K = [[fx, 0, cx], # cx = W/2, cy = H/2
[0, fy, cy],
[0, 0, 1]]
P = K [R | t] # 3x4 projection matrix
Projection of a world point X: x ~ P [X; 1]; divide by w = the third
coordinate. w > 0 means the point is in front of the camera. Pixel is
(x0/w, x1/w).
Field label in the manifest: "convention": "opencv_rdf".
Stored fields are floats/lists (no numpy in the schema); numpy is internal to the compute methods only.
- Intrinsics(
fx, fy, cx, cy, width, height) —matrix()→K. - Camera(
id, intrinsics, R3×3,t3) —projection_matrix(),project(),centre().R/tare world→camera;t = -R@C. - Occluder (ABC) + Box(
center, half_extents) / Sphere(center, radius) —blocks_segment(a, b)ray/segment-vs-solid for hard visibility. - Entity(
id, edges?,frames) — each frame carries named 3D points.idis the stable track identifier: the manifest keeps it byte-identical across every frame, so a tracking consumer reads it as the ground-truth track id for that entity across the whole take.build_multi_entity_scene()demonstrates this end to end with two entities whose per-camera visibility diverges on some frames. - Scene(
fps, num_frames, cameras, entities, occluders).
An entity's frame maps point name → [x, y, z]. This is deliberate so a human
pose layer slots in later without a schema fork:
- an object = 1 entity with 1 named point
"center"; - a COCO-17 human (later) = 1 entity with 17 named points + an
edgesskeleton.
{
"cameras": [
{"id": 0, "K": [[..],[..],[..]], "R": [[..],[..],[..]], "t": [..],
"width": 640, "height": 480, "convention": "opencv_rdf"}
],
"fps": 30.0,
"num_frames": 11,
"entities": [
{
"id": "obj",
"edges": [["a", "b"]],
"frames": [
{
"frame": 0,
"points": {
"center": {
"xyz_gt": [x, y, z],
"per_cam": [
{"cam": 0, "uv": [u, v], "in_view": true, "visible": true, "occ_frac": 0.0}
]
}
}
}
]
}
],
"topology": {
"stations": [{"id": "A", "camera_ids": [0]}, {"id": "B", "camera_ids": [1, 2]}],
"edges": [{"src": "A", "dst": "B", "transit_time_s": 0.25}]
}
}edges is present only when the entity defines a skeleton. topology is present
only for MTMC scenes that declare one (see below).
in_view(bool) — the point projects in front of the camera (w > 0) and inside the image bounds[0,width) x [0,height). Pure framing: ignores occluders. A point in a blind gap isin_view=falseon every camera.visible(bool) — the hard DLT mask:in_view AND not occluded(its segment to the camera centre is not blocked by any occluder). Sovisibleimpliesin_view; a consumer masks triangulation on this field.occ_frac(float, optional) — a continuous difficulty knob: the fraction of a small deterministic jittered sample around the point whose sightline is blocked. It never feeds the mask; it only grades how marginal an occlusion is.
The sampler is configurable via occ_frac_sample_count and occ_frac_jitter,
threaded through observe / build_manifest / write_manifest as optional
keyword-only settings. occ_frac_sample_count is the total number of samples
(centre point + deterministic jitter offsets); occ_frac_jitter is the radius
of the neighbourhood in scene units. The defaults (sample_count=7,
jitter=0.05) reproduce the original manifest output byte-for-byte. Increasing
sample_count adds more deterministic directions (face and space diagonals of
a cube after the six axis-aligned offsets) and therefore grades a marginal
occlusion more finely. The sampling stays deterministic and RNG-free, so the
same scene and settings always produce the same occ_frac.
The manifest path is non-raising: an out-of-frame or behind-camera point is
labelled (in_view=false, visible=false), not an error. Its uv is sanitised
to finite values, and the manifest is written with allow_nan=False, so the JSON
is always strict (no Infinity/NaN).
Floats are serialized at full double precision (no rounding), so a consumer that
rebuilds P = K [R | t] recovers ground truth to ~machine epsilon.
For multi-target multi-camera scenes with non-overlapping stations, the
manifest carries an optional topology:
stations:[{ "id": str, "camera_ids": [int, ...] }]— a named place and the cameras that share (roughly) its view. Station ids are unique.edges:[{ "src": str, "dst": str, "transit_time_s": float }]— directed adjacency: an object leavingsrc's coverage reachesdst's coverage aftertransit_time_sseconds. Endpoints must be declared station ids.
A consumer's MTMC / re-identification path uses this to bound how long a target
may be unseen while crossing the blind gap between adjacent stations. The
entity.id is the cross-camera ground-truth identity a tracker must preserve
across that gap (issue #11, stable track ids).
Projection is exact by default. A typed, seeded NoiseModel
(multicam_sim.noise) can be threaded through observe / build_manifest /
write_manifest (keyword-only noise=, mirroring the occ_frac settings) to
add controlled, reproducible error without touching ground truth:
pixel(PixelNoise,sigma_px) — Gaussian noise, sigma in pixels, added to the OBSERVEDper_cam[i].uvonly.in_view,visible,occ_fracandxyz_gtare computed from the true projection and stay exact.drift(CalibrationDrift) — small seeded perturbations to the ASSUMED calibration a consumer receives: a small-angle axis-angle rotation onR(rotation_sigma_deg, degrees), a per-axis offset ont(translation_sigma, scene units), and per-axis offsets onfx, fy(focal_sigma_px) andcx, cy(principal_point_sigma_px), all in pixels.
seed drives independent numpy.random.default_rng sub-streams (pixel noise;
per-camera drift) — never the global RNG.
Drift is recorded additively: each camera entry gains an optional trailing field, present only when drift is active:
{"id": 0, "K": [...], "R": [...], "t": [...], "width": 640, "height": 480,
"convention": "opencv_rdf",
"assumed": {"K": [[..],[..],[..]], "R": [[..],[..],[..]], "t": [..]}}The ground-truth K, R, t on the entry are untouched; assumed holds the
slightly-wrong calibration a downstream reader would use. With every knob at zero
(or noise=None) no uv is perturbed and no assumed block is emitted, so the
manifest is byte-identical to the noiseless output.
Distinct from occlusion. Occlusion is a scene event (geometry blocks a point);
dropout is a sensor event (a camera delivers no frame). A typed, seeded
SensorDropout (multicam_sim.dropout) threads through build_manifest /
write_manifest (keyword-only dropout=) and blanks whole camera-frames:
drop_prob— each frame of each camera is an independent Bernoulli drop.seeddrives anumpy.random.default_rng([seed, tag, camera_id])sub-stream per camera — a fixed seed gives a byte-reproducible schedule, a different seed a different one, and the stream is independent of the pixel-noise / drift streams (a dropped point still consumes its noise draw, so non-dropped pixels are untouched).
A dropped observation is blank, not a zero occlusion: in_view and visible
are false, uv is [0, 0], the occlusion fields are absent (a dropped
frame carries no occ_frac / visible_fraction — dropout is a coverage gap, not
a fully-visible point with a 0.0 occluder), and a trailing dropped: true marks
it. Each camera entry gains an optional dropped_frames list (its schedule).
Both fields are present only when dropout is active, so dropout=None (or
drop_prob=0) is byte-identical to the no-dropout output.
build_smoke_scene() hand-specifies 3 ring cameras, 1 point moving on a straight
path, and 1 sphere on camera 1's sightline that occludes it during a middle
interval (frames 3–7) while cameras 0 and 2 keep the point. The test writes the
manifest, reloads the JSON, rebuilds the projection matrices from K, R, t, and
triangulates a cam-1-occluded frame from the other two views only (masking on
visible) through the real multicam_occlusion.triangulate_dlt — recovering
ground truth to within 1e-6. This proves occluded-in-one-view recovery through
the actual consumer reader.
Non-overlapping counterpart. Three cameras on two stations — A = camera 0
alone, B = cameras 1 & 2 as an overlapping stereo pair (slightly different
per-camera targets/fov) — with disjoint FOVs, plus one object with a stable
entity.id sweeping A → gap → B. The manifest exhibits, in one take, all three
coverage regimes: station-A single-camera (in_view=[T,F,F], correctly not
triangulable) → a labelled blind gap (in_view false on every camera) →
station-B stereo (in_view=[F,T,T]), where the real
multicam_occlusion.triangulate_dlt recovers ground truth to ~machine epsilon
(≈2e-15) from the two covering views. The scene emits a topology (A↔B with a
transit time). This proves the in_view framing signal, the labelled blind gap,
stable cross-camera identity, and topology end to end through the real reader.
A fluent, typed sugar layer that compiles down to the Scene/manifest above — it adds no fields and changes no convention. Everything below is pydantic-typed, validated at construction, and CPU-only; the renderer is the sole optional part.
Every camera is built through Camera.look_at (or, for custom, stored
verbatim), so the RDF / Z-up / t = -R@C convention is never re-derived.
CameraRig.ring(n, radius, height, look_at, *, width, height_px, focal | fov_deg)
CameraRig.line(n, start, end, look_at, *, width, height_px, focal | fov_deg)
CameraRig.custom(extrinsics=[(R, t), ...], *, width, height_px, focal | fov_deg)
CameraRig.stations([StationView(position, look_at, focal|fov_deg, width?, height_px?), ...],
*, width, height_px, focal | fov_deg) # shared defaults
Intrinsics take exactly one of focal (pixels) or fov_deg (horizontal
FOV); focal = (width/2) / tan(fov/2). ring uses the smoke ring convention:
eye i = (radius·cos t, radius·sin t, height), t = 2π·i/n.
stations — non-overlapping / heterogeneous preset. Unlike ring (one
shared target, overlapping views), each StationView gives its own eye
position and look_at, and may override the rig-wide intrinsics with its
own focal/fov_deg and width/height_px. One preset, two modes:
- MTMC: separated stations with disjoint FOVs — an object is in at most one
station's view at a time, and the space between them is a genuine blind gap
(
in_view=falseon every camera). - Heterogeneous fusion: co-located-ish stations with different targets/zoom —
e.g. camera A wide/high framing a person's volume, camera B close/zoomed framing
items on a bench. Different entities then fall in different cameras'
in_view("human in A not B, items in B not A"), captured by the manifest's per-entity per-camerain_viewwith no schema change.
A Path is a discriminated union on kind (parallels OccluderUnion).
Geometry is u ∈ [0,1] (point(u)); timing is a separate seconds axis.
Path.linear(a, b) Path.waypoints([p0, p1, ...]) (≥2)
Path.circle(center, radius, axis) Path.bezier([c0, c1, ...]) (≥2)
combinators: a.then(b) p.repeat(n) p.over(seconds) p.at_speed(v)
compile: path.compile_frames(fps, num_frames, name="center") -> [EntityFrame]
then/repeat sum wall-clock durations; over rescales the whole trajectory;
at_speed sets duration from arc length. An untimed path is stretched to fill
the scene duration (num_frames-1)/fps; a timed one keeps its duration and
holds at its final point past the end. Output is exactly the per-frame named
points the manifest already expects.
Occlusion.sphere(size) | .box(size) | .plane(size)
.blocks(camera=i).during((frame0, frame1))
[.on(entity, point_name)] [.targeting(coverage)]
The schedule compiles to a real occluder placed on camera i's sightline to
the target point at the window's middle frame; build_manifest then computes
visible geometrically. Two deliberate consequences:
- The window is emergent.
.during((3,7))places geometry aimed at that window; the achievedvisible=Falseinterval is whatever the solid produces. Tests assert the actual manifest pattern, never assume equality with the request. (A static global occluder cannot be exactly per-frame time-gated — that would need a per-frame-occluder schema change, out of scope for the contract. Escalate if a concrete case can't be realized.) visibleis never faked. The hard boolean stays geometric truth.coverageis a monotonic difficulty knob that scales occluder size, moving the continuousocc_fracreadback (quantised to eighths by the manifest sampler); it never touches the triangulation mask. This is the dialable dose that couples to multicam-occlusion's occlusion dose-response.
plane is a thin flat box (finite-plane approximation), so no new occluder type
enters the contract's OccluderUnion.
Scene = (
SceneBuilder(fps, num_frames)
.cameras(CameraRig.ring(...))
.entity("obj", Path.linear(a, b))
.occlude(Occlusion.sphere(0.15).blocks(camera=1).during((3, 7)))
.build()
)
The result is the ordinary Scene; build_manifest(scene) is unchanged. A DSL
scene that reproduces the smoke setup recovers ground truth for a cam-1-occluded
frame through the real triangulate_dlt, exactly like the hand-built smoke.
Synthetic-to-real robustness needs varied scenes. A typed
RandomizationSpec (multicam_sim.randomization) attaches seeded
randomization knobs to the builder: builder.randomize(spec, seed=...) samples
the spec once (RandomizationSpec.sample(seed), a pure function over
numpy.random.default_rng(seed) — the builder never owns the randomness) and
applies the concrete result. Every knob is a closed (min, max) interval,
validated at construction (an inverted min > max interval is rejected), and
every knob is off by default (None / empty), so a scene built without
randomization serialises byte-identically to before.
The three knobs:
background(BackgroundSpec) — the render-time background colour.rgb_min/rgb_maxare RGB triples with channels in[0, 1](the renderer's colour convention), sampled per channel. Default interval(0, 0, 0)–(0.2, 0.2, 0.2): dark, around the default black. Background was already configurable per renderer through thePyrenderBackend(bg=...)constructor argument; this knob adds scene-level, randomizable control — the sampled value is recorded on theSceneand takes precedence over the constructorbgwherever the scene carries one.light(LightSpec) — the key light.intensityis in renderer light units (today's fixed headlight is3.0; default interval(2.0, 4.0)).azimuth_deg/elevation_deggive the direction from the scene to the light in degrees — azimuth in the world XY plane from +X toward +Y (default(0, 360)), elevation above the XY plane (default(30, 90), so the light stays above the horizon); the light shines along the negated vector (Light.direction()).distractors(DistractorSpec) — a count of static, non-target objects.countis an inclusive integer interval (default(1, 3)); each distractor'sx/y/z(scene units) is drawn uniformly from its interval (default a 4×4×1 box around the origin at floor level) and added through the existingSceneBuilder.distractorentry point asrand_distractor_{i}— so randomized distractors appear in the manifest exactly like hand-added ones and never touch the primary entities' ground truth.
Sampling draws in a fixed order (background channels; light intensity,
azimuth, elevation; distractor count; then three coordinates per distractor),
so the reproducibility contract is: same spec + same seed → byte-identical
sample, and therefore a byte-identical scene. Provenance is recorded
additively: the built Scene carries the sampled background / light values
plus a randomization sidecar (RandomizationRecord, spec + seed) that
round-trips through scene JSON, so a randomized run regenerates from its own
output via record.spec.sample(record.seed). None of these fields is read by
the manifest builder; the sampled background/light are consumed by the
pyrender backend (the Kubric backend keeps its fixed key light — see the
renderer section).
RendererBackend is a Protocol (Scene + camera + frame → (H,W,3) pixels).
PyrenderBackend is an offscreen v1; pyrender/trimesh are an optional
render extra imported lazily, never at package load and never in CI — pixels
cannot break the manifest's green bar. KubricBackend
(multicam_sim.dsl.kubric_backend, the kubric extra) is the photoreal
open/closed swap of this Protocol: its Blender-free coordinate translation lives
in multicam_sim.dsl.kubric_spec and is unit-tested (the built camera projects a
point identically to P = K[R|t] within 1e-6); the actual Blender render runs
only inside the kubricdockerhub/kubruntu image. The exact OpenCV-RDF →
Blender-camera conversion, docker recipe, and GT cross-check are in
docs/kubric.md. A Rust core via pyo3 for the analytic
projection/occlusion path is a v2 concern.
Human pose reuses the named-points design with zero schema fork. It is not a new manifest; it is a typed way to build the entities the existing builder already understands.
- A human = 1 entity with 17 named joints (COCO-17) plus an
edgesskeleton (19 limbs). The joints are ordinary named points, sobuild_manifestlabels each joint exactly like any other point. PoseTrajectory.to_entity()(src/multicam_sim/pose.py) lowers a skeleton + per-frame joints to a plainEntity. Nothing in the manifest schema above changes.
So for every joint the manifest already carries:
- GT 3D —
xyz_gt, the joint's world position; - 2D keypoint per camera —
per_cam[i].uv; - per-joint occlusion —
per_cam[i].visibleandper_cam[i].occ_frac. A joint blocked in one view but seen in others isvisible: falsefor the blocked camera andtruefor the rest, per the same hard-DLT contract used for object points.
This per-joint, per-view visibility is exactly the input a multi-view 3D human pose estimator in multicam-occlusion consumes: triangulate each joint from the cameras that still see it, and use the occluded views as held-out difficulty.
Default skeleton: COCO-17 (17 keypoints, the canonical 19-edge skeleton),
Skeleton.coco17(). Dense body models (SMPL / SMPL-X) are an open/closed
extension point: implement the MeshBackend ABC to emit a PoseTrajectory, and
the projection/occlusion/manifest path is unchanged. No mesh backend is
implemented in this layer.
src/multicam_sim/dsl/gait.py generates pose trajectories so a skeleton does
not have to be hand-authored per frame. Three parametric gaits — walk
(periodic legs + opposite-phase arms), reach (one arm smoothsteps to a
body-local target), wave (raised forearm oscillation) — each produce
body-local joint offsets as a function of wall-clock time. A root path (any
existing PathUnion from the motion DSL) translates the whole skeleton:
world joint = root.at_time(t) + local offset. Limbs are rigid: knees and the
reaching elbow are placed by closed-form two-link inverse kinematics with
segment lengths fixed from height, so every COCO-17 edge keeps a constant
length on every frame. Frame compilation is driven
through the motion DSL's own compile_frames, so over(seconds) /
at_speed / untimed stretch-to-scene behave exactly as they do for a single
point, and the gait layer adds no second timing model. Everything is kinematic
and fully deterministic (no RNG); the output is a plain PoseTrajectory, so
the manifest path above is unchanged.