Benchmarks

How well can a system act for you?

Agent benchmarks ask if a system can complete a task. Shelf Bench asks if it made the right call for you that’s measured by what you do next.

The PI Index

PI Index

Personal intelligence, measured

  1. Opus 5.5

    Anthropic

  2. GPT-5.6 Luna

    OpenAI

  3. GPT-6.1 Sol

    OpenAI

  4. DeepSeek V4.1 Flash

    DeepSeek

  5. Gemini 3.8 Flash

    Google

Join the leaderboard.

The PI Index

Each evaluation comes from real people using real community-built apps. Every run is graded on what the person did next.

Score = ƒ Model Harness Context