Daily Radar
The full deduplicated candidate list, followed by 4-5 papers worth inspecting during the evening second day. Recommendations are a menu, not five mandatory assignments.
AI Radar Daily Feed - 2026-09-26
Candidate count after deduplication: 32. Recommended tonight: 5, as a flexible reading menu. Sources checked: Hugging Face Daily Papers September 25, DAIR.AI Papers of the Week September 14–20, Henry Shi’s AI Crash Course, arXiv, and OpenReview. Source limits: arXiv query API returned 429/503 before its first configured query completed; OpenReview notes APIs returned 403; the DAIR.AI September 21–27 issue was not yet published.
- Agent-Editing World Model: Rethinking World Modeling for LLM Agents
- ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds
- GAUGE: When Not to Trust LLM-as-a-Judge in User-Simulated Evaluation of Task-Oriented Agents
- Is Bash All You Need? An Empirical Study of Tool Interfaces for Enterprise Digital Worker Agents
- Training Object Permanence in World Models
AI Radar Daily Feed - 2026-09-19
Candidate count after deduplication: 34. Recommended tonight: 5, as a flexible reading menu. Sources checked: Hugging Face Daily Papers September 18, DAIR.AI Papers of the Week September 7–13, Henry Shi’s AI Crash Course, arXiv, and OpenReview. Source limits: arXiv query API returned 503/429 and timed out before the full nine-query scan could complete; OpenReview notes API returned 403; the DAIR.AI September 14–20 issue was not yet available.
- An Empirical Study of Harness Design for Coding Agents
- SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness
- Procedural Graphs: Self-Evolving Execution Structures for LLM Agents
- PARSER: Read in Parallel, Reason in Depth for Long-Context LLM Agents
- Designing Proactive Thought Partners for Writing
AI Radar Daily Feed - 2026-09-12
Candidate count after deduplication: 36. Recommended tonight: 5. Sources checked: Hugging Face Daily Papers, DAIR.AI Papers of the Week, Henry Shi's AI Crash Course, arXiv, and OpenReview.
- EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents
- Harness-of-Harness: Multi-Day Autonomous Software Development with Continual Improvement
- Language Models Can Control Their Own Attention
- IdeaAMBIG: Benchmarking Implementation-Critical Gaps in Research-Idea Specifications
- NCP-ArchPreview Technical Report: Moving towards Latent Space Language Models through Next Concept Prediction
AI Radar Daily Feed - 2026-09-05
Candidate count after deduplication: 41. Recommended tonight: 5. Sources checked: Hugging Face Daily Papers, DAIR.AI Papers of the Week, Henry Shi's AI Crash Course, arXiv, and OpenReview.
- Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills
- The Compaction Cliff in Long-Running AI Agent Memory
- Context as an Environment: Programmatic Context Management for Long-Horizon Agents
- Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training
- VeriPhy: Agentic Physical Reasoning for World Model Evaluation and Refinement
AI Radar Daily Feed - 2026-08-29
Candidate count after deduplication: 33. Recommended tonight: 5. Sources checked: Hugging Face Daily Papers, DAIR.AI Papers of the Week, Henry Shi's AI Crash Course, arXiv, and OpenReview. Source limitations: the arXiv API scan was HTTP 429-limited after bounded retries, so exact arXiv IDs from successful Hugging Face and DAIR.AI reads provide degraded primary-paper coverage. OpenReview's API returned HTTP 403 and added no fresh paper to the August 24-29 window.
- Agent Lightning v1.0: Towards Harnessed Agentic RL
- Harness Continual Learning: Continual Adaptation Beyond Model Parameters
- WikiSkill: Compiling Agent Experience into Persistent Knowledge for Skill Evolution
- What Does an Evaluation License? A Commit-Bound Census of Claim-Relative Inference in Inspect Evals
- PAWBench: How Far Are We from Probabilistically Aligned World Modeling?
AI Radar Daily Feed - 2026-08-22
Candidate count after deduplication: 72. Recommended tonight: 5. Sources checked: Hugging Face Daily Papers, DAIR.AI Papers of the Week, Henry Shi's AI Crash Course, arXiv, and OpenReview. This is a flexible weekly menu, not a reading quota.
- EnvHarness: Awakening Static Worlds for Agent Learning
- Break It Down, Pass It On: Cross-Task Skill Transfer in LLM Agents
- MemTrapBench: Benchmarking Cognitive Traps in LLM Memory Use
- AI4AI-Bench: Benchmarking LLM Agents in Algorithmic Design for Recursive Self-Improvement
- Skaling: Chinchilla's Exponents Meet Kaplan's Coupling
AI Radar Daily Feed - 2026-08-15
Candidate count after deduplication: 77. Recommended tonight: 5. Sources checked: Hugging Face Daily Papers, DAIR.AI Papers of the Week, Henry Shi's AI Crash Course, arXiv, and OpenReview. Reading posture: a flexible weekly menu, not five assignments.
- DarwinX: Evolving Agent Harnesses Through Natural Selection
- Vero: Can AI Agents Build Formally Verified Software Repositories?
- Spatial Memory Agent
- A Unifying Perspective on Causal World Models
- Training AI Scientists to Replicate Research
AI Radar Daily Feed - 2026-08-08
Candidate count after deduplication: 80. Recommended tonight: 5. Sources checked: Hugging Face Daily Papers, DAIR.AI Papers of the Week, Henry Shi's AI Crash Course, arXiv, and OpenReview. Source limitations: Hugging Face's latest published page was August 7; DAIR.AI's July 27-August 2 issue remains its latest; OpenReview forum pages and APIs were blocked by browser verification, `429`, or `403`, so exact indexed records were used.
- EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning
- HarnessOpt-Bench: Evaluating LLMs at Harness Optimization
- Continual Learning in Transition
- Not All LLM Reasoning is Visible in the Chain-of-Thought
- Interpretable MEG Decoding of Perceived Speech
AI Radar Daily Feed - 2026-08-01
Candidate count after deduplication: 53. Recommended tonight: 5. Sources checked: Hugging Face Daily Papers, DAIR.AI Papers of the Week, Henry Shi's AI Crash Course, OpenReview, and arXiv. Source limitations: Hugging Face's latest published page was July 31; DAIR.AI's July 20-26 issue remains its latest; OpenReview forum pages were blocked by browser verification, so indexed records were used; the configured arXiv query scan was unavailable after four HTTP `429` responses.
- Metis: Memory Foundation Model
- Qwen-UI-Agent Technical Report
- Frontis-MA1
- Is Progressive Disclosure All You Need for Long-Context Agents?
- Verbalizable Representations Form a Global Workspace in Language Models
AI Radar Daily Feed - 2026-07-25
Candidate count after deduplication: 37. Recommended tonight: 5. Sources checked: Hugging Face Daily Papers, DAIR.AI Papers of the Week, Henry Shi's AI Crash Course, OpenReview, and arXiv. Source limitations: Hugging Face's latest published page was July 24; DAIR.AI's July 12-19 issue remains its latest; OpenReview forum pages were blocked by browser verification, so only indexed forum summaries were used; the configured arXiv query scan was unavailable after bounded rate-limit and service-error retries.
- AREX: Towards a Recursively Self-Improving Agent for Deep Research
- Sample-Efficient Learning from Agent Experience
- LLMs Get Lost in Evolving User Intent
- OpenForgeRL: Train Harness-native Agents in Any Environment
- Self-Improvements in Modern Agentic Systems: A Survey
AI Radar Daily Feed - 2026-07-18
Candidate count after deduplication: 41. Recommended tonight: 5. Sources checked: Hugging Face Daily Papers, DAIR.AI Papers of the Week, Henry Shi's AI Crash Course, OpenReview, and arXiv. Source limitations: Hugging Face's July 18 URL redirected to its latest July 17 page; DAIR.AI's July 5-12 issue remains its latest; the configured arXiv raw-query scan was rate-limited and unavailable after bounded retries.
- SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
- SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration
- Rethinking the Evaluation of Harness Evolution for Agents
- How can we assess human-agent interactions? Case studies in software agent design
- BadWAM: When World-Action Models Dream Right but Act Wrong
AI Radar Daily Feed - 2026-07-14
Candidate count after deduplication: 31. Recommended tonight: 5. Sources checked: Hugging Face Daily Papers, DAIR.AI Papers of the Week, Henry Shi's AI Crash Course, OpenReview/TMLR, and arXiv. Source limitation: the configured arXiv raw-query scan was rate-limited; its primary records were used only to verify papers surfaced by other sources.
- Weak-to-Strong Generalization via Direct On-Policy Distillation
- Proxy Exploration and Reusable Guidance
- ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory
- NeuroCogMap Reveals Cognitive Organization of Large Language Models
- WAREX: Web Agent Reliability Evaluation on Existing Benchmarks
AI Radar Daily Feed - 2026-07-13
Candidate count after deduplication: 70. Recommended tonight: 5. Sources checked: arXiv, Hugging Face Daily Papers, DAIR.AI Papers of the Week, Henry Shi's AI Crash Course, OpenReview, targeted project repositories, and one reader-supplied Oak Lab item. Source note: OpenReview required browser verification during this run, so no OpenReview-only candidate was included.
- The OaK Architecture: A Vision of SuperIntelligence from Experience
- Agent Hacks Agent: Autoresearch for Production-Agent Red-Teaming
- LLM-as-a-Verifier: A General-Purpose Verification Framework
- Multi-Agent LLMs Fail to Explore Each Other
- Metacognition in LLMs: Foundations, Progress, and Opportunities