AgentBench is a benchmark to evaluate LLMs' reasoning and decision-making abilities across 8 diverse environments. It reveals a significant performance gap between leading commercial and open-source LLMs, highlighting the need for rigorous evaluation of LLMs as intelligent agents.
No funding rounds tracked yet.