From 7453fa0c88acf7e96e3fcb9447761c63e79c39a8 Mon Sep 17 00:00:00 2001 From: David Hu <159687+divad12@users.noreply.github.com> Date: Wed, 1 Jul 2026 20:07:52 -0300 Subject: [PATCH] task-observer: add experiment-diagnostics global doc and routing table entry --- .claude/AGENTS.md | 1 + docs/ai/experiment-diagnostics.md | 14 ++++++++++++++ 2 files changed, 15 insertions(+) create mode 100644 docs/ai/experiment-diagnostics.md diff --git a/.claude/AGENTS.md b/.claude/AGENTS.md index 6fb4ae6..1ebaa5f 100644 --- a/.claude/AGENTS.md +++ b/.claude/AGENTS.md @@ -125,6 +125,7 @@ Do not run full validation suites after every small edit by habit. Run the narro | Editing AGENTS.md, .agents/skills/, docs/ai/ | writing-docs.md | | Maintaining or troubleshooting ask-intern | ask-intern.md | | Operating or changing the learning system | learning-system.md | + | Running probe loops, experiment diagnostics, or tuning sessions | experiment-diagnostics.md | - Re-read the hard rules before implementation. ## Session Tools diff --git a/docs/ai/experiment-diagnostics.md b/docs/ai/experiment-diagnostics.md new file mode 100644 index 0000000..69dff99 --- /dev/null +++ b/docs/ai/experiment-diagnostics.md @@ -0,0 +1,14 @@ +# Experiment Diagnostics + +## Contract + +- When a probe or experiment produces implausible user-facing guidance, pause the experiment queue before logging the artifact as evidence. +- Fix the diagnostic source first: write a behavior test capturing the bad output, fix the shared logic, and rerun the canonical probe. +- Only promote corrected artifacts to durable evidence or curated docs. An artifact that embeds known-bad guidance becomes a misleading fixture for future agents. +- Record both the discovered product risk and the corrected artifact path before resuming the probe queue. + +## Notes + +- Experiment artifacts are not only observations — they are validation fixtures for product guidance. When an artifact exposes bad guidance, fix the guidance and regenerate the artifact before promoting it to durable evidence. +- What counts as implausible: a recommendation that contradicts domain constraints, a metric far outside the expected range, or guidance a product owner would immediately reject. +- This applies to any long-running probe or tuning loop where the system being probed produces user-facing output: recommendation calibration, AI output evaluation, capacity diagnostics, or auto-populate tuning.