Skip to content

digest: 2026-05-24 - #25

Open
JustinPerea wants to merge 1 commit into
mainfrom
claude/digest-2026-05-24
Open

digest: 2026-05-24#25
JustinPerea wants to merge 1 commit into
mainfrom
claude/digest-2026-05-24

Conversation

@JustinPerea

Copy link
Copy Markdown
Owner

Summary

  • 10 items; UTC date 2026-05-24 (Sunday — slow announcement day, but harvested a missed frontier release + 9 substantive arXiv papers from the May 22 listing not picked up yesterday)
  • Focus areas covered: Frontier models, Long-running / autonomous loops, Tools & connectors, Context engineering, Agent orchestration, Evaluation, Reasoning & inference-time compute
  • Conflicts flagged: 1 (compare) — Gemini 3.5 Flash 76.2% Terminal-Bench 2.1 vs. [arxiv:2605.22535] TerminalWorld 62.5% in-the-wild ceiling with Pearson r = 0.20 between the two benchmarks

Item highlights

  1. Gemini 3.5 Flash (Google, May 19) — missed by prior digests; frontier-fast model claiming 76.2% Terminal-Bench 2.1, 84.2% multimodal, ~4× faster than other frontier models
  2. MOSS (arxiv:2605.22794) — agent self-evolution via source-level code rewriting; 0.25 → 0.61 on OpenClaw in a single cycle; 6th layer above SkillOS / SkillsVote / EvolveMem / Ratchet / Continual Harness
  3. HarnessAPI (arxiv:2605.22733) — single typed Python source-of-truth auto-derives MCP tools + streaming HTTP endpoints; 74% boilerplate reduction; pip install harnessapitry-now
  4. Gated DeltaNet-2 (NVIDIA, arxiv:2605.22791) — channel-wise decoupled erase + write gates; beats Mamba-2/Gated DeltaNet/KDA/Mamba-3 at 1.3B on long-context RULER NIAH
  5. LCGuard (arxiv:2605.22786) — first paper defining KV-cache leakage as a security boundary in multi-agent systems; fourth angle on the agent-safety surface
  6. Boiling the Frog (arxiv:2605.22643) — multi-turn agentic safety bench; 44.4% aggregate ASR; 93.3% on loss-of-control scenarios; Claude Haiku 4.5 strongest (20.5%), Gemini 3.1 Flash Lite weakest (92.9%)
  7. Capability as Liability (arxiv:2605.22672) — inverse scaling on tail-risk forecasting; threshold scoring reverses sign vs. tail-inclusive scoring on identical outputs
  8. Search-E1 (arxiv:2605.22511) — minimal GRPO + offline self-distillation recipe; Qwen2.5-3B reaches 0.440 avg EM on 7 QA benchmarks
  9. Shor (arxiv:2605.22505) — priority-ranking step-level eval of harness optimizers; seventh paper in the agent-eval credibility thread, attacks evaluator cost
  10. VPO (arxiv:2605.22817, MIT — Khattab/Agrawal) — vector-valued rewards for diversity-preserving RL; solves problems GRPO models cannot solve at all under evolutionary search

Conflicts flagged

  • compare — Gemini 3.5 Flash 76.2% Terminal-Bench 2.1 vs. [arxiv:2605.22535] TerminalWorld 62.5% in-the-wild frontier ceiling. The r = 0.20 Pearson correlation TerminalWorld reported between the curated benchmark and in-the-wild distribution means the 76.2% headline is unverified as a real-world capability signal until a TerminalWorld run is published for Gemini 3.5 Flash. Extends the existing agent-benchmark credibility thread (AuditRepairBench, Rollout Cards, Evidence-Supported Bounds, BenchJack, HarnessAudit-Bench, Counterfactual Trace Auditing, Shor) to frontier-model release announcements.

Operational notes

  • OpenAI news / research: HTTP 403 on both openai.com/news/ and openai.com/research/ — could not access — retry next run
  • Local-model corner: no purely local model release today; HarnessAPI (Mac-installable framework) and Gated DeltaNet-2 (1.3B architecture pattern) are the indirect items
  • Sunday is the natural slow point in the arXiv announcement cycle (no Sat/Sun announcements in most cs.* categories) — today's batch is the residual of Friday's listing not covered in yesterday's digest

https://claude.ai/code/session_01TQtJp3suuBwXXx2W2SMAPf


Generated by Claude Code

10 items covering Gemini 3.5 Flash (missed prior digests), MOSS source-level
self-evolution, HarnessAPI single-source MCP+HTTP framework, NVIDIA Gated
DeltaNet-2 linear attention, LCGuard multi-agent KV leakage, Boiling-the-Frog
multi-turn safety, Capability-as-Liability tail-risk forecasting, Search-E1
minimal RL recipe, Shor priority-ranking eval, VPO diversity for test-time
search. 1 compare flag (Gemini 3.5 Flash 76.2% Terminal-Bench 2.1 vs.
TerminalWorld in-the-wild ceiling).

https://claude.ai/code/session_01TQtJp3suuBwXXx2W2SMAPf
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants