Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
8 changes: 8 additions & 0 deletions _meta/claims-ledger.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,3 +9,11 @@
- **2026-05-04** [hf-blog-2026-04-29-ibm-granite-4-1] CLAIM: Granite 4.1-8B "consistently matches or outperforms" Granite 4.0-H-Small (32B MoE / 9B active); 30B reports BFCL v3 73.7, RULER@128K 76.7, GSM8K 94.2, HumanEval 89.6 | source: https://huggingface.co/blog/ibm-granite/granite-4-1 | status: open
- **2026-05-04** [hf-blog-2026-04-29-evaleval-eval-costs-bottleneck] CLAIM: agent benchmarks compress only 2–3.5× vs. 100–200× for static; single GAIA frontier-model run ~$2,829; HAL k=8 reliability rerun ~$320K; pass^k drops from 60% (k=1) to 25% (k=8) | source: https://huggingface.co/blog/evaleval/eval-costs-bottleneck | status: open
- **2026-05-04** [anthropic-research-2026-04-29-biomysterybench] CLAIM: Opus 4.6 ~77% / Sonnet 4.6 ~66% on 76 human-solvable BioMysteryBench tasks; "Claude Mythos Preview" 30% on 23 human-difficult tasks (vs. Opus 4.6 ~24%); many hard-set wins are "brittle" (inconsistent across 5 attempts) | source: https://www.anthropic.com/research/Evaluating-Claude-For-Bioinformatics-With-BioMysteryBench | status: open ; first public reference to "Claude Mythos Preview" model handle
- **2026-05-07** [arxiv:2605.05191] CLAIM: LongSeeker (Qwen3-30B-A3B fine-tune) with five context-orchestration ops (Skip/Compress/Rollback/Snippet/Delete) achieves 61.5% BrowseComp / 62.5% BrowseComp-ZH, vs. Tongyi DeepResearch 43.2/46.7 and AgentFold 36.2/47.3; Compress operator proven expressively complete | source: https://arxiv.org/abs/2605.05191 | status: open
- **2026-05-07** [arxiv:2605.05066] CLAIM: No architecture can simultaneously achieve length-independent per-step compute, length-independent state, and length-proportional recall; models satisfying first two recall at most O(poly(d)/log V) KV pairs; verified across 52 architectures pre-March 2026 | source: https://arxiv.org/abs/2605.05066 | status: open ; foundational theoretical claim, expect citations rather than refutations
- **2026-05-07** [arxiv:2605.05007] CLAIM: Uno-Orchestra (learned joint policy over decomposition + worker selection + budget) hits 77.0% macro pass@1 across 13 benchmarks, +16 pp over strongest workflow baseline at ~10× lower per-query cost | source: https://arxiv.org/abs/2605.05007 | status: open ; relates-to [arxiv:2604.27221] (learned vs. hand-designed orchestrator, both report large gains)
- **2026-05-07** [arxiv:2605.04785] CLAIM: AgentTrust pre-execution tool-call gate achieves 95.0% verdict accuracy / 73.7% risk-level accuracy on 300-scenario internal benchmark; 96.7% / ~93% on 630-scenario adversarial benchmark; ms-latency; ships as MCP server (AGPL-3.0) | source: https://arxiv.org/abs/2605.04785 | status: open ; relates-to [arxiv:2605.00136] tool-use tax (both intervene at protocol layer; safety vs. performance angles)
- **2026-05-07** [arxiv:2605.04361] CLAIM: Same context artifacts produce up to 20× speedup or 46% degradation across 10 tasks × 7 conditions × 2,700+ runs; baseline (no-context) exploration predicts outcome with Pearson r = -0.82 (p < 0.001); two regimes — training-data-driven (disrupted by artifacts) vs. explicit-instruction-driven (not) | source: https://arxiv.org/abs/2605.04361 | status: open
- **2026-05-07** [arxiv:2605.05170] CLAIM: Design Conductor 2.0 multi-agent system (frontier models, April 2026) autonomously designed FPGA-mapped TurboQuant inference accelerator (240-cycle pipeline, 5,129 FP16/32 units, 5.7 mm² in TSMC 16FF) in 80 hours; claims 80× larger task scope than prior 12h RISC-V CPU result | source: https://arxiv.org/abs/2605.05170 | status: open ; vendor-self-reported, autonomy claim awaits independent reproduction
- **2026-05-07** [hf-blog-2026-05-06-vllm-correctness-before-corrections] CLAIM: When migrating PipelineRL from vLLM 0.8.5 → 0.18.1, logprob mismatches between inference and training masquerade as RL-objective divergence; required fixes (logprobs_mode=processed, disable prefix caching, disable async scheduling, fp32 lm_head, specific weight-update sequence) restore V0 parity on policy ratio, KL, entropy, reward, weight lag | source: https://huggingface.co/blog/ServiceNow-AI/correctness-before-corrections | status: open
- **2026-05-07** [anthropic-news-2026-05-05-finance-agents] CLAIM: Claude Opus 4.7 scores 64.37% on Vals AI Finance Agent benchmark | source: https://www.anthropic.com/news/finance-agents | status: open ; vendor-self-reported, first public Opus 4.7 number tracked in this ledger
8 changes: 8 additions & 0 deletions _meta/processed.json
Original file line number Diff line number Diff line change
@@ -1,4 +1,12 @@
[
{ "id": "arxiv:2605.05191", "url": "https://arxiv.org/abs/2605.05191", "title": "LongSeeker: Elastic Context Orchestration for Long-Horizon Search Agents", "date_covered": "2026-05-07" },
{ "id": "arxiv:2605.05066", "url": "https://arxiv.org/abs/2605.05066", "title": "The Impossibility Triangle of Long-Context Modeling", "date_covered": "2026-05-07" },
{ "id": "arxiv:2605.05007", "url": "https://arxiv.org/abs/2605.05007", "title": "Uno-Orchestra: Parsimonious Agent Routing via Selective Delegation", "date_covered": "2026-05-07" },
{ "id": "arxiv:2605.04785", "url": "https://arxiv.org/abs/2605.04785", "title": "AgentTrust: Runtime Safety Evaluation and Interception for AI Agent Tool Use", "date_covered": "2026-05-07" },
{ "id": "arxiv:2605.04361", "url": "https://arxiv.org/abs/2605.04361", "title": "When Context Hurts: The Crossover Effect of Knowledge Transfer on Multi-Agent Design Exploration", "date_covered": "2026-05-07" },
{ "id": "arxiv:2605.05170", "url": "https://arxiv.org/abs/2605.05170", "title": "Design Conductor 2.0: An agent builds a TurboQuant inference accelerator in 80 hours", "date_covered": "2026-05-07" },
{ "id": "hf-blog-2026-05-06-vllm-correctness-before-corrections", "url": "https://huggingface.co/blog/ServiceNow-AI/correctness-before-corrections", "title": "vLLM V0 to V1: Correctness Before Corrections in RL", "date_covered": "2026-05-07" },
{ "id": "anthropic-news-2026-05-05-finance-agents", "url": "https://www.anthropic.com/news/finance-agents", "title": "Agents for financial services", "date_covered": "2026-05-07" },
{ "id": "arxiv:2605.00737", "url": "https://arxiv.org/abs/2605.00737", "title": "To Call or Not to Call: A Framework to Assess and Optimize LLM Tool Calling", "date_covered": "2026-05-04" },
{ "id": "arxiv:2604.27221", "url": "https://arxiv.org/abs/2604.27221", "title": "Web2BigTable: A Bi-Level Multi-Agent LLM System for Internet-Scale Information Search and Extraction", "date_covered": "2026-05-04" },
{ "id": "arxiv:2605.00425", "url": "https://arxiv.org/abs/2605.00425", "title": "AEM: Adaptive Entropy Modulation for Multi-Turn Agentic Reinforcement Learning", "date_covered": "2026-05-04" },
Expand Down
Loading