Skip to content

digest: 2026-06-12 - #56

Merged
JustinPerea merged 1 commit into
mainfrom
claude/digest-2026-06-12
Jun 12, 2026
Merged

digest: 2026-06-12#56
JustinPerea merged 1 commit into
mainfrom
claude/digest-2026-06-12

Conversation

@JustinPerea

Copy link
Copy Markdown
Owner

Items (8)

  1. DiffusionGemma (DeepMind blog, June 10) — 26B-total / 3.8B-active MoE text-diffusion LM on Gemma 4 backbone; 256 tokens/forward pass via bi-directional attention; 1000+ tok/s H100, 700+ tok/s RTX 5090; MLX support at launch (~18 GB at NVFP4 4-bit); Apache 2.0. try-now + local-fit.
  2. EvoArena + EvoMem ([arxiv:2606.13681], HF digest: 2026-05-06 #2 54 upvotes) — first dynamic-environment memory benchmark; agents avg 39.6%; EvoMem +1.5/+6.1/+4.8% on EvoArena/GAIA/LoCoMo. track.
  3. HyperTool ([arxiv:2606.13663]) — MCP code-block tool primitive; Qwen3-32B 15.69→35.29% (+19.6pp) on MCP-Universe; schema-compatible with existing MCP servers. try-now.
  4. The Illusion of Multi-Agent Advantage ([arxiv:2606.13003]) — auto-MAS underperforms CoT-SC at 10× cost; expert-architected MAS still wins. compare.
  5. EurekAgent ([arxiv:2606.13662]) — environment engineering > workflow prescription; SOTA 26-circle packing at ~$11 API. track.
  6. SENTINEL ([arxiv:2606.12908]) — failure-as-curriculum RL; Tau2-Bench Retail 66.4→74.9%. track.
  7. G-Long ([arxiv:2606.13115]) — associative-graph dialogue memory; +9.8% MSC, 40.8% LME recall. track.
  8. MemRefine ([arxiv:2606.13177]) — LLM-judge memory compaction (similarity proposes, content judges). track.

Focus areas covered

Frontier models / Local Apple Silicon (1), Agent memory (3, 7, 8 + 2 partial), Tools & connectors / MCP (3, 6), Agent orchestration (4, 5), Evaluation (2, 4), Reasoning & inference-time compute (6).

Conflicts flagged (2)

  • Auto-MAS critique vs expert-architected MAS positive thread — resolved to "different categories"; Arbor/MUSE-Autoskill survive the critique.
  • DiffusionGemma 4× speedup is GPU-only — open empirical question whether it transfers to M5 Max under MLX; direct head-to-head vs Qwen3-30B-A3B is the right experiment.

State files updated: _meta/processed.json (8 new IDs newest-first), _meta/claims-ledger.md (8 new entries).

https://claude.ai/code/session_013XDzaQC6LRhiyU91yxzUoq


Generated by Claude Code

8 items covering DeepMind DiffusionGemma open-weight text-diffusion LM
(26B/3.8B-active MoE, MLX day-one), EvoArena+EvoMem dynamic-environment
memory benchmark, HyperTool MCP code-block primitive (+19-23pp on
MCP-Universe), "Illusion of Multi-Agent Advantage" auto-MAS critique
(10x cost vs CoT-SC), EurekAgent environment-engineering position+system
(SOTA 26-circle packing at ~$11), SENTINEL failure-driven RL,
G-Long graph dialogue memory, MemRefine LLM-judge compaction.

2 conflicts flagged: auto-MAS critique vs expert-architected MAS
positive thread (resolves to "different categories"); DiffusionGemma
4x speedup is GPU-only — open question on M5 Max transfer.

https://claude.ai/code/session_013XDzaQC6LRhiyU91yxzUoq
@JustinPerea
JustinPerea merged commit f86039b into main Jun 12, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants