01 / Research corpus

Last journal commit
Aug 3, 2026, 1:27 PM

Research papers

Claim-level audits, independent evidence, and reproducible experiments across every monitored domain.

Corpus at rest1

active research topics

Audited papers5
Experiments0
Open opportunities2
Research runs2

02 / Audited corpus

Paper index

5 of 5 papers

01AI agents and agentic AI
PROMISING - UNREPLICATED

Reflexion: Language Agents with Verbal Reinforcement Learning

Reported gains span ALFWorld, HotpotQA, and code tasks, but tests, external error signals, retries, and extra inference are not fully budget-matched. Reflection alone harms the Rust subset while the full tests-plus-reflection composite wins.

Artificial intelligenceAI agentsAgent reflection
02AI agents and agentic AI
PROMISING - UNREPLICATED

Generative Agents: Interactive Simulacra of Human Behavior

A 100-participant response-ranking study supports bounded believability, but the end-to-end result is one 25-agent two-day simulation and does not validate prediction of real human behavior.

Artificial intelligenceAI agentsHuman-computer interaction
03AI agents and agentic AI
SUBSTANTIATED

API-Bank: A Comprehensive Benchmark for Tool-Augmented LLMs

The benchmark design, released artifact, and historical baseline table are directly inspectable. Current comprehensiveness, broad synthetic-training causality, and current-model rankings are not substantiated.

Artificial intelligenceAI agentsAgent evaluation
04AI agents and agentic AI
PROMISING - UNREPLICATED

ReAct: Synergizing Reasoning and Acting in Language Models

The original multi-benchmark ablations support a bounded performance signal for interleaving reasoning and actions, but proprietary PaLM-540B, best-trial comparisons, missing uncertainty, and no located direct independent reproduction prevent a stronger verdict.

Artificial intelligenceMachine learning
05AI agents and agentic AI
PROMISING - UNREPLICATED

Tool Specifications Matter: Uncovering and Mitigating Safety Risks in AI Agents

Controlled component and representation ablations plausibly isolate a tool-schema formatting effect, and SafeKeep reports large two-benchmark gains across four models. Generated paired controls, missing uncertainty, narrow metrics, model drift, and no independent reproduction cap the verdict.

Artificial intelligenceAI safetySoftware systems