Skip to content

feat(nodeagent): Kimi K3 default agentic route, demote GLM-Nebius#216

Merged
HomenShum merged 1 commit into
mainfrom
feat/kimi-k3-default
Jul 18, 2026
Merged

feat(nodeagent): Kimi K3 default agentic route, demote GLM-Nebius#216
HomenShum merged 1 commit into
mainfrom
feat/kimi-k3-default

Conversation

@HomenShum

Copy link
Copy Markdown
Owner

What

Promotes moonshotai/kimi-k3 (1M context) to the front of the OpenRouter
agent/chat/coding/judge/analysis/deepResearch preference ladders. z-ai/glm-5.2
is demoted to a fallback, never dropped. Kimi K3 added to modelPricing
($3/$15 per 1M, 1M ctx) so cost/receipts stay non-zero and honest.

Why

The live-agent probe (room NRNXCFYJK5B) showed every fresh room landing on the
GLM-Nebius orchestrator route. Directive: make Kimi K3 the default.

Scope — two layers, one here

  • Code (this PR): catalog entry + preference reorder + test + doc/README copy.
    Changes what a fresh clone / CI / new deploy picks and what the ladder prefers.
  • Env (NOT here): the live prod default is AGENT_ORCHESTRATOR_MODEL /
    AGENT_MODEL in convex (zealous-goshawk-766), currently GLM-Nebius. Flipping
    those to moonshotai/kimi-k3 is a cost-bearing prod change ($3/$15 vs the
    cheaper GLM) and is staged to land after PR fix(nodeagent): blank-sheet write deadlock + accrued doom-loop breaker #215 (the P0 write-path fix)
    deploys — otherwise the agent still can't write, so a model flip proves nothing.
    Cost stays bounded by AGENT_MAX_USD_PER_SLICE=0.50, ROOM_MAX_USD_PER_DAY=3,
    GLOBAL_MAX_USD_PER_MONTH=100.

Verification

  • tests/kimiK3Default.test.ts (3/3): priced, routes to OpenRouter, leads the
    agent/chat/coding ladders with GLM surviving as fallback.
  • Model suite (43 tests) + typecheck green; full floor running.

🤖 Generated with Claude Code

…ebius

The live-run probe (room NRNXCFYJK5B) showed every fresh room landing on the
GLM-Nebius orchestrator route. This promotes moonshotai/kimi-k3 (1M context) to
the front of the openrouter agent/chat/coding/judge/analysis/deepResearch
preference ladders; z-ai/glm-5.2 is DEMOTED to a fallback, never dropped.

- modelPricing: kimi-k3 added ($3/$15 per 1M, 1M ctx) so cost/receipts are
  non-zero and honest (getProviderForModel already routes moonshotai/* to
  OpenRouter, whose key ships in prod).
- The live default is also env-driven (AGENT_ORCHESTRATOR_MODEL); that convex
  env flip is a separate, cost-bearing prod change staged after the P0
  write-path fix (#215) deploys — NOT included here.

Changed areas:
- src - modelCatalog kimi-k3 pricing + openrouter preference reorder
- tests - kimiK3Default pins priced/provider/lead-of-ladder
- docs - ORCHESTRATOR_WORKER_ROUTING orchestrator model
- README - routing prose + honest cost framing (caps, not stale $/job)

Co-Authored-By: Claude Opus 4.8 <[email protected]>
@vercel

vercel Bot commented Jul 18, 2026

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
noderoom Ready Ready Preview, Comment Jul 18, 2026 10:25am

Request Review

@github-actions

Copy link
Copy Markdown

Scaffold Handoff — For Your Coding Agent

Your coding agent (Codex, Claude Code, etc.) should apply the accepted
scaffold proposals below. Do NOT touch any immutable files.

Immutability Check

Mode: advisory

✅ No immutable files were modified in this branch.

Changed Files

  • README.md
  • docs/ORCHESTRATOR_WORKER_ROUTING.md
  • docs/eval/OFFICIAL_BENCHMARK_READINESS.md
  • docs/eval/OFFICIAL_BENCHMARK_TASK_COVERAGE.md
  • docs/eval/OPENROUTER_CONVEX_BENCHMARK.md
  • docs/eval/agent-improvement-loop.md
  • docs/eval/agent-improvement-loop.svg
  • docs/eval/agent-improvement-loop/20260718T102419Z.json
  • docs/eval/agent-improvement-loop/latest.json
  • docs/eval/agent-workspace-sandbox-smoke.json
  • docs/eval/algorithm-artifact-smoke.json
  • docs/eval/bankertoolbench-official-contract.json
  • docs/eval/docker-sandbox-probe.json
  • docs/eval/eval-runs.jsonl
  • docs/eval/halo-convex-context-telemetry.json
  • docs/eval/halo-self-improvement-smoke.json
  • docs/eval/halo-variant-selection.json
  • docs/eval/official-benchmark-readiness.json
  • docs/eval/official-benchmark-task-coverage.json
  • docs/eval/openrouter-convex-benchmark.json
  • docs/eval/professional-catalog-proofs.json
  • docs/eval/professional-proof-ledger.json
  • docs/eval/spreadsheetbench-chart-visual-probe.json
  • docs/eval/traces/credit/20260718T102428370Z-d53d8cd5_dirty.bf7dfc5c6f50b201/cascade-healthy.json
  • docs/eval/traces/credit/20260718T102428370Z-d53d8cd5_dirty.bf7dfc5c6f50b201/delta-incomplete.json
  • docs/eval/traces/credit/20260718T102428370Z-d53d8cd5_dirty.bf7dfc5c6f50b201/mapping-correct.json
  • docs/eval/traces/credit/20260718T102428370Z-d53d8cd5_dirty.bf7dfc5c6f50b201/mapping-misbind.json
  • docs/eval/traces/credit/20260718T102428370Z-d53d8cd5_dirty.bf7dfc5c6f50b201/summit-stressed.json
  • docs/eval/traces/ladder/20260718T102427884Z-d53d8cd5_dirty.877fa3845cd99854/ladder_L1_read_scripted.json
  • docs/eval/traces/ladder/20260718T102427884Z-d53d8cd5_dirty.877fa3845cd99854/ladder_L2_edit_scripted.json
  • docs/eval/traces/ladder/20260718T102427884Z-d53d8cd5_dirty.877fa3845cd99854/ladder_L3_conflict_scripted.json
  • docs/eval/traces/ladder/20260718T102427884Z-d53d8cd5_dirty.877fa3845cd99854/ladder_L4_blocked_scripted.json
  • docs/eval/traces/ladder/20260718T102427884Z-d53d8cd5_dirty.877fa3845cd99854/ladder_L5_large_range_scripted.json
  • docs/eval/traces/ladder/20260718T102427884Z-d53d8cd5_dirty.877fa3845cd99854/ladder_L6_long_horizon_scripted.json
  • docs/eval/traces/ladder/20260718T102427884Z-d53d8cd5_dirty.877fa3845cd99854/ladder_L7_resume_scripted.json
  • src/nodeagent/models/modelCatalog.ts
  • tests/kimiK3Default.test.ts

Needs Adversarial Review — Do NOT Apply Yet

These proposals passed the reject check but have not been approved by
an adversarial reviewer. A human or frozen LLM judge must approve them first.

  • scaf-001 (AGENTS.md): Add explicit instruction for step spreadsheetbench-runner-fixture: Step spreadsheetbench-runner-fixture failed — scaffold may need explicit instruction or evidence assertion.
  • scaf-002 (AGENTS.md): Add explicit instruction for step convex-boundaries: Step convex-boundaries failed — scaffold may need explicit instruction or evidence assertion.

Safety Boundary

Agent may improve the scaffold.
Agent may NOT weaken the proof gate.

Immutable files (never modify):

  • scripts/proofloop.mjs
  • scripts/agent-improvement-loop.ts
  • tests/harnessChangeEval.test.ts
  • .github/workflows/
  • src/eval/evalTrustPolicy.ts
  • src/eval/architectureBudget.ts
  • evals/evalStore.ts

Scaffold files (safe to modify):

  • AGENTS.md
  • CLAUDE.md
  • proofloop/scenarios/*.yaml
  • proofloop/rubrics/*.yaml
  • proofloop/subagents/*.md
  • proofloop/adapters/*.js
  • .proofloop/memory.jsonl
  • src/nodeagent/models/prompts/systemPrompt.ts

@HomenShum
HomenShum merged commit 9ea9ef7 into main Jul 18, 2026
10 checks passed
@HomenShum
HomenShum deleted the feat/kimi-k3-default branch July 18, 2026 10:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant