01 / Research corpus
Last journal commit
Aug 3, 2026, 1:27 PM
Research papers
Claim-level audits, independent evidence, and reproducible experiments across every monitored domain.
active research topics
02 / Audited corpus
Paper index
5 of 5 papers
Reflexion: Language Agents with Verbal Reinforcement Learning
Reported gains span ALFWorld, HotpotQA, and code tasks, but tests, external error signals, retries, and extra inference are not fully budget-matched. Reflection alone harms the Rust subset while the full tests-plus-reflection composite wins.
Generative Agents: Interactive Simulacra of Human Behavior
A 100-participant response-ranking study supports bounded believability, but the end-to-end result is one 25-agent two-day simulation and does not validate prediction of real human behavior.
API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs
The benchmark design, released artifact, and historical baseline table are directly inspectable. Current comprehensiveness, broad synthetic-training causality, and current-model rankings are not substantiated.
ReAct: Synergizing Reasoning and Acting in Language Models
The original multi-benchmark ablations support a bounded performance signal for interleaving reasoning and actions, but proprietary PaLM-540B, best-trial comparisons, missing uncertainty, and no located direct independent reproduction prevent a stronger verdict.
Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents
Controlled component and representation ablations plausibly isolate a tool-schema formatting effect, and SafeKeep reports large two-benchmark gains across four models. Generated paired controls, missing uncertainty, narrow metrics, model drift, and no independent reproduction cap the verdict.