Codex/loopx advisor#2406
Conversation
|
Thanks for the additional benchmark work. I re-reviewed the latest head I am revising my earlier scope recommendation: I do not require this to be split into three PRs. The underlying LoopX Turn path is already experimental, Advisor mode is off when the flag is omitted, and The new evidence is meaningful: five real-model cases used the same Sol baseline/Advisor and Luna executor; both arms passed independent validation in all 5 cases; 4/5 cases reduced tokens; aggregate usage fell from 1,166,651 to 813,209 (30.30%). That is enough to justify continued experimentation in main. It is not yet evidence for making Advisor automatic by default: each case is one Turn and one observed run, and Before merge, please close these focused items without adding another framework. 1. Fix the usage receipt before treating the benchmark totals as exact
Please prefer a valid cumulative total when it is present, accept only real non-negative integers (excluding booleans and fractional numbers), and add regression fixtures containing both last/total usage plus a fractional counter. Then rerun the five-case qualification. The updated receipt, rather than the current screenshot totals, should be the cost evidence used for merge. 2. Keep
|
# Conflicts: # loopx/control_plane/turn_driver/codex_cli.py




Summary
Issue Or Task
Validation
python3 -m py_compile loopx/*.pyloopx check --scan-root .987 passed, 2 skippedBoundary Checklist
.loopx/,.codex/goals/, liveACTIVE_GOAL_STATE.md, credentials, private benchmark traces, verifier output, raw agent sessions, internal document links, or local machinepaths.