Skip to content

chore(engine): subtree pull opencode 08ddbdc (m29 scorecard guard) - #44

Merged
TimothyVang merged 2 commits into
mainfrom
agent/m29-engine-sync
Jul 13, 2026
Merged

chore(engine): subtree pull opencode 08ddbdc (m29 scorecard guard)#44
TimothyVang merged 2 commits into
mainfrom
agent/m29-engine-sync

Conversation

@TimothyVang

Copy link
Copy Markdown
Owner

Summary

  • Subtree-pulls verdict-opencode main at 08ddbdc (test(smoke): clarify agent hang timeout after /v1 fix #19 offline scorecard regression guard) into engine/ — squash commit a5deded, merge d5f4cd3.
  • Rebuilt verdict binary: sha256 a9197243e48a39c9c3a495fc316655f4fd63d69e572dbdf410250cee744d50d2; used_fallback evidenced in built runtime (native-runtime test 1 pass/0 fail; 12 string hits in binary).

Test plan

  • npm run typecheck ordered (SDK → TUI → CLI): ok (lane log tmp/m29-logs/engine.log)
  • node scripts/selftest.mjs on this branch: 216 passed, 0 failed (orchestrator re-run after full npm run build)

Honest residual

  • Report-only: 20 absolute /home/... lines remain in engine custody fixtures (paths under /home/raven/caseforge-core/evidence/...); fix is dev-repo-owned path relativization, out of scope here.

m29 lane: engine (grok headless), worktree cf-m29-engine, base origin/main feef841.

timothy.vang added 2 commits July 13, 2026 15:43
08ddbdc test(dfir): offline scorecard regression guard (fixture-driven) (#19)
d5641ec feat(opencode): emit used_fallback from native LLM runtime gate (#20)
136158a test(dfir): offline scorecard regression guard (fixture-driven)

git-subtree-dir: engine
git-subtree-split: 08ddbdc668a4cfddfd35efc74e247a5bc7f381c5
@TimothyVang
TimothyVang merged commit 0d90467 into main Jul 13, 2026
1 check passed

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: d5f4cd3be6

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

}
},
"attack-samples": {
"expected_verdict_band": "INDETERMINATE",

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Correct attack-samples expected verdict

For the attack-samples case, this runtime-side oracle disagrees with the authoritative scorecard in fixtures/dfir-scorecard/ground-truth.json, and the existing fixtures/dfir-scorecard/runs/attack-samples.verdict.json is actually SUSPICIOUS. If this case is added to the new Bun scorecard guard or reused by callers, the grader will reject the known-good sealed run for the wrong reason instead of catching real regressions.

Useful? React with 👍 / 👎.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant