B2B
StartupOlam Labs
Olam Labs measures social intelligence in AI models for safety, agent performance, and real-world use, through behavioral evaluations in multi-agent simulations where models cooperate, deceive, and negotiate.
Milestones
Milestone
YC Summer 20268 upvotes
Can you beat frontier LLMs at social strategy games? | Multi-Agent Arena by Olam Labs, evaluating through multi-agent simulations
Play games like Risk, Poker, or Codenames against frontier AIs. We're running public long-horizon games to do evaluations on social behavior and agentic performance.
GitHub releases
- v1.0-r179 → xa8zz/erdos-harness