AgentComparison
Editorial

How we compare agents

Evidence labels, comparable tasks and explicit unknowns.

Three different kinds of tool

A product performs tasks for its users. A platform lets a team configure agents. A framework gives developers building blocks. We do not put them in one universal league table.

Evidence levels

  1. Documentation researched: a dated official source supports a capability. This does not prove reliability.
  2. Independently evaluated: a named outside evaluation supplies a task, sample and method. Its results remain attributed.
  3. Hands-on tested: a recorded test includes inputs, version, plan, approvals, repetitions, costs and failures.

Every current agent is documentation researched. No independent or hands-on product scores are claimed.

How a test would work

Fix the task, input, success criteria and allowed actions before testing. Record interventions, unsafe actions, completion, errors, time and real cost. Repeat runs and publish sample size. Different tasks do not collapse into a universal best-agent score.

Publication gate

Verify identity, availability, official pricing and plan restrictions. Attribute permissions and privacy controls. Provide distinct decision value, working sources and a meaningful limitation. Visually inspect mobile and desktop, check accessibility, canonical URLs and indexability, and keep weak or contradictory records in draft.

Coverage and updates

Coverage is a dated snapshot, never a claim to include every agent on Earth. New discoveries go into a candidate queue. Conflicting availability claims remain quarantined until resolved.